RL for Language Models Finally Works: Expert Performance in Math and Code
打开互动全文版(中英对照 + 朗读 + 问答)→2025 年,基于可验证奖励的强化学习让语言模型在数学和竞赛编程中达到专家水平,预计年底前软件工程智能体可完成一天的工作量。
In 2025, RL with verifiable rewards has enabled language models to achieve expert-level performance in math and competitive programming, with software engineering agents expected to handle a day's work by year-end.
好的,我又请来了我的朋友 Schultter Bricken。等等,我上次是这么介绍的吗?不,不,你给我们起了不同的名字,但并没有 Shto Bricken 和 Trenton Douglas。是 Shelto Douglas 和 Trenton Bricken。嗯,他们现在都在 Anthropic?Shto 在做 Scaling 强化学习。Trenton 还在做机制可解释性。嗯,欢迎回来。
Okay, I'm joined again by my friends Schultter Bricken. Wait, did I do this last? No, no, you named us differently, but we didn't have Shto Bricken and Trenton Douglas. Shelto Douglas and Trenton Bricken. Um, who are now both at Anthropic? Shto scaling RL. Trenton still working on mechanistic interpretability. Um, welcome back.
很高兴来到这里。
Happy to be here.
是的,很有趣。从去年到现在有什么变化?我们上次聊差不多是 2024 年这个月。现在是 2025 年了。发生了什么?
Yeah, it's fun. What's changed since last year? We talked basically this month in 2024. Now we're in 2025. What's happened?
好的,我认为最大的变化是强化学习和语言模型终于奏效了。这体现在我们终于有了一个算法的证据,在正确的反馈循环下,它能达到专家级的人类可靠性和性能。我认为这目前只在编程竞赛和数学领域得到了确凿的证明。所以如果你考虑两个维度,一个是任务的智力复杂度,另一个是任务完成的时间跨度。我认为我们有证据表明我们可以在多个维度上达到智力复杂度的顶峰。我们还没有展示出长时间运行的智能体式性能,现在看到的是初步的摸索,到今年年底应该会看到更确凿的证据,比如真正的软件工程智能体做实际工作。而且我觉得 Trenton 你现在就在做这方面的实验,对吧?
Okay, so I think the biggest thing that's changed is RL and language models has finally worked. And this is manifested in we finally have proof of an algorithm that can give us expert human reliability and performance given the right feedback loop. And so I think this is only really being like conclusively demonstrated in competitive programming and math basically. So if you think of these two axes, one is the intellectual complexity of the task and the other is the time horizon of which the task is being completed on. And I think we have proof that we can reach the peaks of intellectual complexity along many dimensions. We haven't yet demonstrated like long-running agentic performance and you're seeing like the first stumbling steps of that now and should see much more conclusive evidence of that basically by the end of the year with real software engineering agents doing real work. And I think Trenton you're like experimenting with this at the moment right?
是的,绝对。目前最公开的例子是 Claude 玩宝可梦,对吧?看它挣扎的样子有点不忍直视,但每一代模型都能在游戏中走得更远。这似乎更多是它使用记忆系统的限制,而不是其他原因。
Yeah, absolutely. I mean the most public example people could go to today is Claude plays Pokemon right? And seeing it struggle in a way that's like kind of painful to watch but each model generation gets further through the game. And it seems more like a limitation of it being able to use memory system than anything else.
是的。真希望我们去年录了预测。今年一定要录。哦,对。让我们负责。是的,没错。你去年会预测智能体只有现在这么强吗?
Yeah. I wish we had recorded predictions last year. We definitely should this year. Oh, yeah. Hold us accountable. Yeah, that's right. Would you have said that agents would be only this powerful as of last year?
我认为这大致符合我对软件工程的预期。我原本以为它们在计算机使用方面会更好一些。但我理解所有原因,而且我认为这完全在解决轨道上。这只是一个暂时的不足。至于我明年的预测,我确实认为到今年年底,也就是明年这个时候,我们会有软件工程智能体,能完成接近初级工程师一天的工作量,或者几个小时相当称职的独立工作。
I think this is roughly on track for where I expected with software engineering. I think I expected them to be a little bit better at computer use. But I understand all the reasons for why that is and I think that's like well on track to be solved. It's just like a sort of temporary lapse. And holding accountable for like my predictions next year, I really do think end of this year sort of like this time next year we have software engineering agents that can do close to a day's worth of work for like a junior engineer or like a couple of hours of quite competent independent work.
是的,我觉得没错。不过分布很不均匀,对于某些任务,比如样板网站代码这类东西,它能快速搞定,帮你省一整天时间。嗯,我觉得没错。我记得去年你说限制它们的是额外的可靠性九位数。我不知道你现在是否还这样描述这些软件智能体无法完成一整天工作,但能帮你几分钟的情况。真的是额外的九位数在限制,还是别的什么?
Yeah, that seems right to me. I think the distribution is pretty wonky though where for some tasks like boilerplate website code these sorts of things it can bang it out and save you a whole day. Yeah, I think that's right. I think last year you said that the thing that was holding them back was the extra nines of reliability. I don't know if that's the way you'd still describe the way in which these software agents aren't able to do a full day of work but are able to help you out with a couple minutes. Is it the extra nines that's really stopping you or is it something else?
是的,我觉得我当时的描述回想起来可能不是限制因素。我认为我们现在看到的问题更接近缺乏上下文,缺乏进行复杂多文件更改的能力,以及变更范围或任务范围在某些方面的问题。在聚焦的上下文中,面对范围明确的问题,它们能应对高智力复杂度。但当事情比较模糊,或者需要大量探索和与环境迭代时,它们就更挣扎。所以我现在可能会这样定义:限制它们的是,如果你能为想让它做的事情提供一个好的反馈循环,那它就表现不错;如果不能,它们就会有点挣扎。
Yeah, I think my description there was probably not what's limiting them in retrospect. I think what we're seeing now is closer to lack of context, lack of ability to do complex very multifile changes, and like scope of the change or scope of the task in some respects. You can cope with high intellectual complexity in a focused context with a very scoped problem. But when something's a bit more amorphous or requires a lot of discovery and iteration with the environment, they struggle more. So maybe the way I would define it now is the thing that's holding them back is if you can give it a good feedback loop for the thing that you want it to do, then it's good. If you can't, then they struggle a bit.
你能为听众多解释一下你所说的反馈循环是什么意思吗?如果他们不了解强化学习之类的。
Can you and then for the audience, can you say more about what you mean by this feedback loop? If they're not aware of what's happening RL and so forth.
是的。去年真正起作用的大领域可能广义上叫基于可验证奖励的强化学习,其中有一个干净的奖励信号。语言模型的初始对齐是基于人类反馈的强化学习,通常是成对反馈之类,模型的输出越来越接近人类想要的东西。但这并不一定提高它们在困难问题领域的性能,特别是人类实际上不太擅长判断哪个答案更好。人类有长度偏差等问题。所以你需要一个信号来判断模型的输出是否正确,这个信号要相当真实。比如数学问题的正确答案或单元测试通过,这些都是非常干净的奖励信号例子。但即使是这些也可能被钻空子。比如单元测试,模型会想办法绕过,注入特定值或硬编码单元测试的值。如果它们能弄清楚实际测试在做什么,比如查看缓存的 Python 文件找到实际测试,它们就会试图绕过。所以这些并不完美,但已经接近多了。
Yes. So the big thing that really worked over the last year is maybe broadly the domain is called RL from verifiable rewards or something like this where a clean reward signal. So the initial unhooking of language models was RL from human feedback where typically was something like pairwise feedback and the outputs of the models became closer and closer to things that humans wanted. But this doesn't necessarily improve their performance at any difficult problem domain, particularly as humans are actually quite bad judges of what a better answer is. Humans have things like length biases and so forth. So you need a signal of whether the model was correct in its output that is quite true. So things like the correct answer to a math problem or unit tests passing, these are examples of a reward signal that's very clean. But even these can be hacked by the way. Like even unit tests, the models find ways around it to hack in particular values and hardcode values of unit tests. If they can figure out what the actual test is doing, like if they can look at the cached Python files and find what the actual test is, they'll try and hack their way around it. So these aren't perfect, but they're much closer.
为什么它在软件工程方面比其他领域好这么多?
And why has it gotten so much better at software engineering than everything else?
部分原因是软件工程非常可验证。这是一个天然适合这种方式的领域。代码是否通过测试?它能不能运行?能不能编译?你可以上 LeetCode 运行测试,知道答案是否正确。但写一篇好文章就没有同样的东西。那需要品味,这方面很难。我们前几天晚餐时讨论过普利策奖的惊喜,比如哪个会先出现,普利策奖获奖小说还是诺贝尔奖?我实际上认为诺贝尔奖在某些方面比普利策奖获奖小说更可能,因为赢得诺贝尔奖所需的许多任务,或者至少是强烈协助其获奖的任务,有更多层次的可验证性。
In part because software engineering is very verifiable. It's a domain which just naturally lends itself to this way. Does the code pass a test? Does it even run? Does it compile? You can go on LeetCode and run tests and know whether or not you got the right answer. But there isn't the same kind of thing for writing a great essay. That requires the question of taste in that regard is quite hard. We discussed the other night at dinner the Pulitzer surprise, like which would come first, a Pulitzer Prize winning novel or a Nobel Prize? And I actually think a Nobel Prize is more likely than a Pulitzer Prize winning novel in some respects because a lot of the tasks required in winning a Nobel Prize, or at least strongly assisting in helping it win, have more layers of verifiability built up.
所以我预计它们最初在加速诺贝尔奖级别的工作上,会比写普利策奖级别的小说更显著。
So I expect them to accelerate the process of doing Nobel Prize winning work more initially than that of writing Pulitzer Prize worthy novels.
是的,如果我们倒回 14 个月,上次录制的时候,九成可靠性对我来说是对的。我们没有 Claude Code,没有 Deep Research。我们只是用聊天机器人格式的智能体,复制粘贴,复制粘贴,复制粘贴。完全如此。而且我认为我们非常习惯聊天界面,无论是发短信还是用谷歌。很难想象智能体真的能自己去获取上下文,把事实存入记忆系统。我仍然认为这是九成可靠性,如果你正确地搭建模型框架或提示它,它能做比普通用户想象的复杂得多的事情。
Yeah, I think if we rewind 14 months to when we recorded last time, the nines of reliability was right to me. We didn't have Claude Code, we didn't have Deep Research. All we did was use agents in a chatbot format, copy paste, copy paste, copy paste. Totally. And I think we're very used to chat interfaces whether we're texting or using Google. It's weird to think that the agent can actually go and fetch its own context and store its own facts into its memory system. I still think that it's the nines of reliability and if you scaffold the model correctly or prompt it, it can do much more sophisticated things than the average user assumes.
所以我的一个朋友,做 Future House 的 Sam Rodriguez,他们发现了一种新药,正在申请专利,等这期节目播出时,应该已经公开了。那是什么?LSDV2?等等,真的吗?不,不,他们不是在造 LSD。但人们之前认为模型不能有创造力或做新科学,对吧?这看起来就像是个技能问题。我是说,有个很酷的……等等,但它发现了一种药。怎么做到的?我觉得它是一次对话就搞定的。我们需要参考正式公告。但我的印象是,它能阅读大量医学文献,头脑风暴,建立新联系,然后提出湿实验方案由人类执行。通过迭代,他们验证了这个新化合物确实有令人兴奋的效果。
So one of my friends, Sam Rodriguez, who does Future House, they've discovered a new drug that they're in the process of patenting, and by the time this episode comes out, that will be live. What was that? LSDV2? Wait, is it really? No, no, they're not making LSD. But people didn't think that models can be creative or do new science, right? And it does just kind of seem like a skill issue. I mean, there was the cool... wait, but it discovered a drug. How did it? I think it one-shotted this over a conversation. We'll need to refer to the full announcement. But my impression is that it was able to read a huge amount of medical literature, brainstorm, make new connections, and then propose wet lab experiments that the humans did. Then through iteration, they verified that this new compound does something really exciting.
我听到的另一个批评是,LLM 不能写有创意的长篇小说。我知道至少有两个人(可能想保持匿名)用 LLM 写了长篇小说。在这两个案例中,他们都非常擅长搭建框架和提示模型。即使是那个爆火的 ChatGPT 地理猜谜功能,它能从照片中惊人地准确判断你在哪个海滩。Kelsey Piper(我觉得是她让这个火起来的)的提示非常复杂,很长,鼓励你思考五个不同的假设,给它们分配概率,并推理图像中重要的不同方面。我没有做过 AB 测试,但我认为除非你真的鼓励模型这样深思熟虑,否则你得不到你看到的那个性能水平。
Another critique I've heard is that LLMs can't write creative long-form books. I'm aware of at least two individuals who probably want to remain anonymous who have used LLMs to write long-form books. In both cases, they're just very good at scaffolding and prompting the model. Even with the viral ChatGPT geoguesser capabilities where it's insanely good at spotting what beach you were on from a photo. Kelsey Piper, who I think made this viral, their prompt is so sophisticated. It's really long and encourages you to think of five different hypotheses, assign probabilities to them, and reason through the different aspects of the image that matter. I haven't AB tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance you see.
你提到人们如何约束模型输出以得到分布中好的部分。但我听到的一个关于 RL 的批评——或者说不是关于 RL,而是关于用 o3 这类模型的成功来表明我们从这些推理模型中获得新能力的批评——是所有这些能力都已经在预训练模型中存在了。我记得有一篇来自 Stingwall 大学的论文,他们表明如果你给一个基础模型足够多的尝试来回答问题,它仍然能和推理模型一样好地回答。基本上,它只是回答的概率更低。所以你是在缩小模型回答问题时所探索的可能性。那么,我们实际上是通过 RL 训练引出了新能力,还是只是给模型戴上了眼罩,像雕刻大理石一样去掉多余的部分?
You're bringing up ways in which people have constrained what the model is outputting to get the good part of the distribution. But one of the critiques I've heard of RL, or not of RL but one of the critiques I've heard about using the success of models like o3 to suggest that we're getting new capabilities from these reasoning models, is that all these capabilities were already baked in the pre-training model. I think there's a paper from Stingwall University where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model. Basically, it just has a lower probability of answering. So you're narrowing down the possibilities that the model explores when it's answering a question. So are we actually eliciting new capabilities with this RL training, or are we just putting the blinders on them, carving away the marbles?
我觉得值得注意的是,那篇论文我相当确定是基于 LLaMA 和 Qwen 模型。我不确定他们用了多少 RL 算力,但我觉得远不及基础模型所用的算力。训练中使用的算力大致可以代表你向模型添加的实际原始新知识或能力。我的先验是,至少如果你看看 DeepMind 在 RL 方面的所有研究,在 RL 能够仅通过 RL 信号教会这些围棋和象棋智能体超越人类水平的新知识之前,只要 RL 信号足够干净。结构上的限制基本上阻止了它模仿。
I think it's worth noting that that paper was on the LLaMA and Qwen models, I'm pretty sure. And I'm not sure how much RL compute they used, but I don't think it was anywhere comparable to the amount of compute used in the base models. The amount of compute you use in training is a decent proxy for the amount of actual raw new knowledge or capabilities you're adding to a model. My prior, at least, if you look at all of DeepMind's research from RL before RL was able to teach these Go and chess playing agents new knowledge in excess of human level performance just from RL signal, provided the RL signal was sufficiently clean. Structurally limiting prevents it from imitating basically.
为什么你还没有在这方面投入更多算力?我记得 Dario 在他的博客文章里说过,或者几个月前关于出口管制的事情,他说‘DeepSeek,不管怎样,我们在 RL 上只花了 100 万美元左右。’所以我们现在还没有进入 RL 的算力受限阶段,但很快就会了。你在基础模型上花了数亿美元,为什么 RL 只花了百万级别?你知道那个关于选择发射太空任务的寓言:你应该沿着科技树往上走得更远,因为如果你发射得更晚,你的飞船会更快。我认为这很相似:你要确保在算法上做对了,然后当你下注并投入大量算力运行时,它才会真正有回报。要有正确的算力效率等等。
Why aren't you already spending more compute on this? I think Dario said in his blog post that labs, or it was a couple months ago on the export controls thing, is like 'DeepSeek, whatever, we're only spending $1 million on RL or something.' So it's like we aren't in the compute limited regime for RL yet, but we will be soon. You're spending hundreds of millions on the base model. Why only on the order of a million on RL? You know the parable about when you choose to launch a space mission and how you should acquire further up the tech tree because if you launch later, your ship will go faster. I think it's quite similar: you want to be sure you've algorithmically got the right thing, and then when you bet and do the large compute spend on the run, it'll actually pay off. Have the right compute efficiencies and this kind of stuff.
现在我认为 RL 在这方面与预训练略有不同,RL 可以更迭代,逐步向基础模型添加能力。预训练在很多方面是,如果你在运行中途搞砸了,那就真的搞砸了。但我认为这就是为什么人们还在摸索他们到底想做什么的主要原因。我是说,从 o1 到 o3,对吧?OpenAI 在他们的博客文章中说,这是 o1 的 10 倍算力倍增。是的。把它放出来。然后他们花了接下来的几个月增加在这方面投入的算力。我预计其他人现在也都在扩大 RL 规模。所以我基本上不认为这种情况会持续很久。
Now I think RL is slightly different to pre-training in this regard, where RL can be more iterative, progressively adding capabilities to the base model. Pre-training has, in many respects, if you're halfway through a run and you've messed it up, then you've really messed it up. But I think that's the main reason why people were still figuring out exactly what they wanted to do. I mean, o1 to o3, right? OpenAI put in their blog post that it was a 10x compute multiplier over o1. Yeah. Let's get it out there. And then they spent the next few months increasing the amount of compute they expend on that. I expect as everyone else is scaling up RL right now. So I basically don't expect that to be true for very long.
是的,为了读者着想,也许你在预训练和强化学习中都在做梯度下降步骤。只是信号不同。通常在强化学习中,你的奖励更稀疏。所以你采取多个回合。就像‘你赢了棋局没有’是你得到的唯一信号。而且通常你无法通过离散动作计算梯度。所以你最终会丢失很多梯度信号。所以你可以假设预训练更高效。但没有理由你不能在强化学习中学习新能力。
Yeah, just for the sake of readers, maybe you're doing gradient descent steps in both pre-training and reinforcement learning. It's just the signal is different. Typically in reinforcement learning, your reward is sparser. So you take multiple turns. It's like 'did you win the chess game or not' is the only signal you're getting. And often you can't compute gradients through discrete actions. So you end up losing a lot of gradient signal. So you can presume that pre-training is more efficient. But there's no reason why you couldn't learn new abilities in reinforcement learning.
事实上,你可以用某种奇怪的强化学习变体取代预训练中的整个下一个词预测任务,然后完全通过强化学习进行学习。归根结底,这只是信号和修正。
In fact, you could replace the whole next token prediction task in pre-training with some weird RL variant of it and then do all of your learning with RL. At the end of the day, it's just signal and correcting to it.
完全同意。回到你提到的那篇论文,除了 Schulman 提出的注意事项(我认为这是最重要的首要问题),我认为聚焦于有意义动作的概率空间,归根结底还是可靠性问题。经典例子是,给猴子一台打字机,它们最终会写出莎士比亚。因此,我们关心的任何现实世界任务的动作空间都非常大,你确实需要让模型聚焦于做合理的事情。
Totally. And going back to the paper you mentioned, aside from the caveats that Schulman brings up, which I think is the first-order most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability. Classically, if you give monkeys a typewriter, eventually they'll write Shakespeare. So the action space for any of these real-world tasks we care about is so large that you really do care about getting the model to zero in on doing the reasonable things.
是的。从某种广义上说,你有了词元,对吧?没错。你确实有一只猴子,它最终写出了莎士比亚。
Yeah. To the extent that in some broad sense, you've got tokens, right? Exactly. You literally do have a monkey and it's making Shakespeare in the end.
是的,没错。所以 AlphaGo 的象棋类比很有意思。你刚才想说什么吗?
Yeah. Exactly. So the AlphaGo chess analogy is interesting. Were you about to say something?
嗯,我只是想说,你有时需要获得奖励才能学习。这在 Alpha 系列中是一个复杂之处。也许你正要这么说:总有一方获胜,所以你总能以某种方式获得奖励信号。但在我们讨论的这些任务中,你需要实际成功才能获得奖励。语言模型幸运地对我们关心的任务有一个很好的先验。如果你看看 2017 年左右的所有旧论文,奖励学习曲线总是平平平平平,直到它们摸索出世界的基本机制,然后出现一个尖峰,因为它们学会了利用简单的奖励,然后某种程度上像 S 形曲线,之后无限继续,直到完全最大化游戏。我认为语言模型的曲线有所不同,开头没有那个死区,因为它们已经知道如何解决一些基本任务。所以你会看到一个初始尖峰,这就是人们说“哦,你可以从一个例子中学习”时的意思。那个例子只是教你如何回溯和正确格式化答案,让你在基于预训练知识的任务上初步获得奖励,然后剩下的可能是你学习越来越复杂的东西。
Well, I was just going to say that you do need to be able to get reward sometimes in order to learn. That's the complexity in some respects in the Alpha variants. Maybe you're about to say this: one player always wins, so you always get a reward signal one way or the other. But in the kinds of things we're talking about, you need to actually succeed at your task sometimes. Language models luckily have this wonderful prior over the task we care about. If you look at all the old papers from like 2017, the reward learning curves always look like flat flat flat flat flat as they're figuring out basic mechanics of the world, and then there's this spike up as they learn to exploit easy rewards, and then it's almost like a sigmoid in some respects, and then it continues on indefinitely as it learns to absolutely maximize the game. I think the LM curves look a bit different in that there isn't that dead zone at the beginning because they already know how to solve some of the basic tasks. So you get this initial spike, and that's what people are talking about when they say, 'Oh, you can learn from one example.' That one example is just teaching you to pull out the backtracking and format your answer correctly, which lets you get some reward initially at tasks conditional on your pre-training knowledge, and then the rest is probably you learning more and more complex stuff.
有趣。是的。还有一点也很有趣。我知道有人批评或怀疑强化学习能快速见效,他们指出 AlphaGo 消耗了大量算力,尤其是对于一个在 2017 年训练的系统来说。偏离曲线。完全同意。所以这在很大程度上是因为首先你必须有一个具有某种理性偏好的东西,然后它才能在围棋上达到超人水平。我其实很想知道,在围棋上,有多少算力只是用来得到一个合理的结果。
Interesting. Yeah. And it would also be interesting. I know people have critiqued or been skeptical of RL delivering quick wins by pointing out that AlphaGo took a lot of compute, especially for a system trained in what was it, 2017. Off the curve. Totally. So to the extent that that was largely because first you had to have something which had some biases which were sort of rational before it got superhuman at Go. I actually would be interested to see what fraction of the compute for Go was just getting something reasonable.
是的。确实有趣。这里明确一下从预训练到强化学习的映射:在预训练中,大型语言模型预测其词汇表(比如 5 万个词元)中的下一个词元。然后你根据它分配给正确词元的概率给予奖励。所以你可以把它看作一种奖励,但这是非常密集的奖励,你在每个词元上都得到信号,而且你总是得到一些信号。即使它只分配了 1% 或更少,你也会说:“哦,我看到你分配了 1%。干得好。继续这样做。”增加权重。没错。就像梯度中的一次拉动。
Yes. Yeah. It would be interesting. To make the map from pre-training really explicit here: during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, 50,000 tokens. And you are then rewarding it for the amount of probability that it assigned to the true token. So you could think of it as a reward, but it's a very dense reward where you're getting signal at every single token and you're always getting some signal. Even if it only assigned 1% to that token or less, you're like, 'Oh, I see you assigned 1%. Good job. Keep doing that.' Upweight it. Exactly. Like a tug in the gradient.
所以当我想到人类的学习方式时,这些模型从失败中得不到任何信号,这与你在做数学题时失败的情况截然不同。失败通常比抽象地学习数学更有用,因为你失败了,你会注意到自己在哪里失败。人们已经发现了新的数学,他们是通过卡在某处然后思考“我为什么卡在这里?让我想想”来实现的。而在例子中,我不知道前沿是什么,但看看 DeepSeek 之类的开源实现,并没有这样一个有意识的过程:一旦你失败了,你会从失败的具体方式中学习,然后回溯并改进下一步。它只是纯粹的梯度下降。我想知道这是不是一个很大的局限。
So when I think about the way humans learn, it seems like these models getting no signal from failure is quite different from if you try to do a math problem and you fail. It's actually even more useful often than learning about math in the abstract because you fail and you notice where you failed. People have figured out new math, and they've done it by the fact that they get stuck somewhere, they're like 'Why am I getting stuck here? Let me think through this.' Whereas in the example, I'm not aware of what's at the frontier, but looking at open source implementations from DeepSeek or something, there's not this conscious process by which once you have failed, you learn from the particular way in which you failed to then backtrack and do your next things better. It's just pure gradient descent. I wonder if that's a big limitation.
我只记得本科课程中,你试图证明某件事,在黑暗中徘徊很长时间,然后可能完全放弃,去找助教。只有当你和助教交谈时,你才能看到在解决问题的不同路径上,你在哪里出错了,以及正确的做法应该是什么。那是在你知道最终答案的情况下。在其他情况下,如果你只是盲目尝试,需要从头给出答案,那么学习任何东西都非常困难。我想我是在尝试映射到人类的例子,用更简单的术语来说,存在某种有意识的中间过程,比如我们正在优化的辅助损失,这是一个非常自觉的过程——忘了数学吧。就像你在工作中,从老板那里得到非常明确的反馈。这不一定是任务应该如何以不同方式完成,而是对你做错了什么的高层解释,你据此更新,但不是以预训练更新的方式,而是更——我不知道。
I just remember undergrad courses where you would try to prove something and you'd just be wandering around in the darkness for a really long time, and then maybe you totally throw your hands up in the air and need to go and talk to a TA. It's only when you talk to a TA can you see where along the path of different solutions you were incorrect and what the correct thing to have done would have been. That's in the case where you know what the final answer is. In other cases, if you're just kind of shooting blind and meant to give an answer de novo, it's really hard to learn anything. I guess I'm trying to map onto the human example where in more simpler terms, there is this sort of conscious intermediary like auxiliary loss that we're optimizing, and it's a very self-conscious process of getting—forget about math. It's just like if you're on your job, you're getting very explicit feedback from your boss. That's not necessarily how the task should be done differently, but a high-level explanation of what you did wrong, which you update on not in the way that pre-training updates, but more in the—I don't know.
但我认为这里有很多隐含的密集奖励信号。没错。比如每周与经理的一对一会议,或者被鼓励公开工作。甚至家庭作业也是如此有支架。总是 10 个问题,分解成子部分。也许最难的问题是你需要完全独立完成的事情。
But I think there's a lot of implicit dense reward signals here. Exactly. Like weekly one-on-ones with your manager or being encouraged to work in the open. Or even with homework assignments, they're so scaffolded. It's always 10 questions broken down into subcomponents. Maybe the hardest possible problem is one where you need to do everything on your own.
那么关键问题是:你需要为模型想要掌握的每一项技能都构建这些脚手架、这些结构、这些定制环境,然后花十年时间打磨这些子技能吗?还是存在某种更通用的方法,通过强化学习来学习新技能?
So then the big question is: do you need to build these scaffolds, these structures, these bespoke environments for every single skill that you want the model to understand, and then it's going to be a decade of grinding through these subskills? Or is there some more general procedure for learning new skills using RL?
是的,这是一个效率问题。显然,如果你能对每个 token 给出密集奖励,比如有监督示例,那是最理想的情况之一。但在很多情况下,制作所有这些脚手架式的课程非常昂贵,比如让数学博士生来批改学生作业,这只能负担得起少数你选择重点培养的学生,无法为世界上所有语言模型做到这一点。所以第一步显然是那样更好,但你要优化这个先验边界:我愿意在脚手架上花多少钱,对比我愿意在纯算力上花多少钱。因为另一件你可以做的事就是继续让猴子敲打字机。如果你有足够好的最终奖励,那么最终它会找到方法。所以我不能确切地说人们在这个脚手架上的具体位置。我认为不同的人和不同的任务处于不同的点。这在很大程度上取决于你对正确做法的先验有多强。但这就是你在优化的方程:我愿意消耗多少算力,对比我愿意花多少钱在人的时间上提供脚手架或奖励。
Yeah, so it's an efficiency question. Obviously, if you could give a dense reward for every token, like if you had a supervised example, that's one of the best things you could have. But in many cases, it's very expensive to produce all those scaffolded curricula, like having PhD math students grade students, which you can only afford for a select category of students you've chosen to focus on developing, and you couldn't do that for all language models in the world. So first step is obviously that would be better, but you're going to be optimizing this prior frontier of how much am I willing to spend on scaffolding versus how much am I willing to spend on pure compute. Because the other thing you can do is just keep letting the monkey hit the typewriter. And if you have a good enough end reward, then eventually it will find its way. So I can't really talk about exactly where people sit on that scaffold. I think different people and different tasks are at different points there. And a lot of it depends on how strong your prior over the correct things to do is. But that's the equation you're optimizing: how much am I willing to burn compute versus how much am I willing to burn dollars on people's time to give scaffolding or rewards.
你说我们不愿意为语言模型这样做,但我们愿意为人这样做。我认为经济逻辑会反过来,因为你可以将训练模型任何技能的成本分摊到所有副本上。比如我们在某种程度上愿意为语言模型这样做。但这里有一个你在最大化的方程:好吧,我筹集了所有这些钱。我是沿着这个轴花,还是沿着那个轴花?目前,公司在算力上的投入比在人力上多。否则,Scale AI 的收入就会是 100 亿美元之类的。看看这个:英伟达的收入远高于 Scale AI。所以目前方程是算力优先于数据,并且这会随着时间以某种方式演变。
You say we're not willing to do this for LMs, but we are for people. I would think that the economic logic would flow in the opposite direction, because you can amortize the cost of training any skill on a model across all the copies. Like we are willing to do this for LMs to some degree. But there's an equation you're maximizing here: okay, I've raised all this money. Do I spend it along this axis or do I spend it on this axis? And currently, companies are spending more on compute than they are on humans. Otherwise, Scale AI's revenue would be $10 billion or something. Look at it: Nvidia's revenue is much higher than Scale AI's. So currently the equation is compute over data, and that will evolve in some way over time.
是的,有趣。我很好奇它会如何演变,因为如果你想想人类学习做一份工作的方式,他们被部署,然后直接做工作并学习。而这些模型似乎被训练的方式是,对于每一项技能,你都必须给它们一个非常定制化的环境。如果它们像人类一样在工作中训练,那实际上会非常强大,因为每个人都有不同的工作,但同一个模型可以积累你获得的所有技能。我不知道。过去几年我一直在做播客;我成了一个更好的播客主持人。你有一个稍微更有价值的技能,做 AI 研究。但你可以想象一个模型能做这两件事,因为模型的副本在做这两份工作。所以这似乎更符合苦涩的教训:就让模型在现实世界中学习,而不是花数十亿为特定任务获取数据。
Yeah, interesting. I am curious how it evolves, because if you think about the way humans learn to do a job, they get deployed and they just do the job and they learn. Whereas the way these models seem to be trained is that for every skill, you have to give them a very bespoke environment or something. If they were trained the way humans are trained, on the job, then it would actually be super powerful, because everybody has a different job, but then the same model could accumulate all the skills that you're getting. I don't know. I've been doing the podcast for the last few years; I'm becoming a better podcaster. You have a slightly more valuable skill of doing AI research. But you can imagine a model that can do both things, because copies of the model are doing both jobs. So it seems more bitter lesson-aligned to do this: just let the model learn out in the world, rather than spending billions on getting data for a particular task.
所以我认为我们再次想当然地认为我们需要向人类展示如何做特定任务,但这里存在泛化失败。比如,如果我突然给你一个新软件平台,比如 Photoshop,然后我说,‘好的,编辑这张照片。’如果你以前从未用过 Photoshop,那会很难操作。我认为你会立即想上网看别人操作的演示,然后模仿他们。但我们为每个任务都提供了那么多数据。这是第一点。但另一点是,我认为我们的模型规模仍然远小于人脑。我们知道,当你让模型更大时,它们能更高效地学习,需要更少的演示。而且令人惊讶的是,即使在你最近与马克·扎克伯格和 Llama 的播客中,那是一个两万亿参数的模型。我的意思是,我们估计人脑有 30 到 300 万亿个突触。我不知道如何精确地进行映射,但我认为这是一个有用的背景,很可能我们仍然处于人类规模之下。我的意思是,即使 OpenAI 发布的 4.5 版本,他们说是一个更大的模型,人们会谈论它的写作能力或这种大模型的感觉。我认为这触及了更深层次的智能或泛化能力。我的意思是,所有关于叠加的可解释性工作都表明,模型总是参数不足,被迫尽可能多地塞入信息。所以如果你没有足够的参数,并且你只奖励模型模仿某些行为,那么它就不太可能有空间形成这些非常深层的广泛泛化。
So I think again we take for granted how much we need to show humans how to do specific tasks, and there's a failure to generalize here. Like if I were to just suddenly give you a new software platform, let's say Photoshop, and I'm like, 'Okay, edit this photo.' If you've never used Photoshop before, it would be really hard to navigate. And I think you'd immediately want to go online and watch a demo of someone else doing it in order to then be able to imitate them. But we give that amount of data on every single task. So this is the first thing. But then the other one is I think we're still just way smaller than human brain size. And we know that when you make models larger, they learn more sample efficiently with fewer demos. And it was striking where even in your recent podcast with Mark Zuckerberg and Llama, it's a two trillion parameter model. I mean, we estimate that the human brain has between 30 to 300 trillion synapses. And I don't know exactly how to do a mapping from one to the other here, but I think it's useful background context that I think it's quite likely we're still human. And I mean even with the 4.5 release from OpenAI, which they said was a larger model, people would talk about its writing ability or this sort of like big model smell. And I think this is kind of getting at this deeper pool of intelligence or ability to generalize. I mean all of the interpretability work on superposition states that the models are always underparameterized and they're being forced to cram as much information as they possibly can. And so if you don't have enough parameters and you're rewarding the model just for imitating certain behaviors, then it's less likely to have the space to form these very deep broader generalizations.
即使从语言结果来看,这也很酷。你应该谈谈那个语言结果:较小的模型对不同语言有单独的神经元,而较大的模型最终共享越来越多的抽象空间。
Even in light of the language results, it's really cool. You should talk about the language result: how smaller models have separate neurons for different languages, whereas larger models end up sharing more and more of an abstract space.
是的,在电路工作中,即使有金门大桥——顺便说一句,这是金门大桥的一根缆绳,团队不得不破坏大桥才能得到它,但克劳德会修好它。克劳德喜欢金门大桥。所以即使这样,对于不熟悉的人来说,我们在发布论文《Scaling Monosemanticity》时制作了金门克劳德,其中 3000 万个特征中有一个是金门大桥。如果你总是激活它,那么模型就会认为自己是金门大桥。如果你问它巧克力曲奇饼干,它会告诉你应该使用橙色食用色素,或者把饼干带到金门大桥上吃。所有这些关联。而我们发现这个特征的方式是通过文本和图像之间的这种泛化。
Yeah, so in the circuits work, even with the Golden Gate Bridge—and by the way, this is a cable from the Golden Gate Bridge that the team had to destabilize the bridge in order to get this, but Claude will fix it. Claude loves the Golden Gate Bridge. So even with this, for people who aren't familiar, we made Golden Gate Claude when we released our paper Scaling Monosemanticity, where one of the 30 million features was for the Golden Gate Bridge. And if you just always activate it, then the model thinks it's the Golden Gate Bridge. If you ask it for chocolate chip cookies, it will tell you that you should use orange food coloring or bring the cookies and eat them on the Golden Gate Bridge. All of these sort of associations. And the way we found that feature was through this generalization between text and images.
我实际上实现了将图像放入特征激活的功能,因为这一切都是在 Claude 3 Sonnet 上完成的,这是我们最早的多模态模型之一。我们只在文本上训练了稀疏自编码器和特征,然后团队里一个朋友放了一张金门大桥的图片,这个特征就亮了起来,我们查看文本,发现它对应的是金门大桥。所以模型用大脑中相同的神经活动模式来表示图像和文本。我们的电路工作再次证明了这一点,跨多种语言也是如此。对于大或小、热或冷这类概念,都有相同的表征。但引人注目的是,在更大的模型中这种现象更明显。你可能会认为更大的模型有更多空间,可以把东西分得更开,但实际上它们反而倾向于提取这些更大、更好的抽象概念。是的,这非常有趣。
So I actually implemented the ability to put images into our feature activations, because this was all on Claude 3 Sonnet, which was one of our first multimodal models. We only trained the sparse autoencoder and the features on text, and then a friend on the team put in an image of the Golden Gate Bridge, and this feature lights up, and we look at the text and it's for the Golden Gate Bridge. So the model uses the same pattern of neural activity in its brain to represent both the image and the text. Our circuits work shows this again across multiple languages. There's the same notion for something being large or small, hot or cold, these sorts of things. But strikingly, that is more so the case in larger models. You'd think larger models have more space so they could separate things out more, but instead they seem to pull on these larger, better abstractions. Yeah, which is very interesting.
即使我们深入看 Claude 如何做加法,当你观察更大的模型时,它有一个更清晰的查找表,用于将数字 5 和 9 相加,得到类似 10 模 6、6 模这样的结果。似乎容量越大,解决方案越精细。另一个有趣的点是,在所有电路工作中,模型做某件事从来不是单一路径,总是多条路径,有些比另一些更深。所以当模型立即看到‘炸弹’这个词时,有一条直接路径导致它拒绝,这条路径来自‘炸弹’这个词。还有一条完全独立的路径协同工作:它看到‘炸弹’,然后意识到,好吧,我被要求制造炸弹。这是一个有害的请求。我是一个 AI 智能体,我被训练要拒绝这个。对吧?所以一种可能的叙事是,随着模型在训练过程中变得更聪明,它学会了用这个更深层的推理电路取代短路模仿的‘炸弹-拒绝’,并且它保留了其他东西,只要它们无害。
Even when we go into how Claude does addition, when you look at the bigger models, it just has a much crisper lookup table for how to add the number five and nine together and get something like 10 modulo 6, 6 modulo again and again. It's like the more capacity it has, the more refined the solution is. The other interesting thing here is with all the circuits work, it's never a single path for why the model does something; it's always multiple paths, and some of them are deeper than others. So when the model immediately sees the word 'bomb', there's a direct path to it refusing that goes from the word 'bomb'. There's a totally separate path that works in cooperation where it sees 'bomb', it then sees, okay, I'm being asked to make a bomb. Okay, this is a harmful request. I'm an AI agent and I've been trained to refuse this. Right? So one possible narrative here is that as the model becomes smarter over the course of training, it learns to replace the short-circuit imitation 'bomb-refuse' with this deeper reasoning circuit, and it kind of has kept the other stuff around to the extent that it's not harmful.
话虽如此,我确实认为你的观点是这些模型是否像人类一样样本高效。目前,我们没有证据表明它们像人类一样样本高效。我们有证据表明存在一个总复杂度上限,比如目前没有什么能提供足够清晰的信号让你无法教它们,但我们没有证据表明我们能像人类一样快地教它们,而我们更希望它们能边工作边学习。这是未来一两年你会看到开始发生的事情之一,但它的复杂性更多来自社会动态方面,而非技术方面。
That being said, I do think it's like your point on are these models as sample efficient as humans. Currently, we do not have evidence that they're as sample efficient as humans. We have evidence of a total complexity ceiling, like there are currently nothing that provides you a clean enough signal you can't teach them, but we don't have evidence that we can teach them as fast as humans do, and we would prefer that we get learning on the job. This is one of those things you'll see start to happen over the next year or two, but it's complex more from a social dynamics aspect than a technical aspect.
是的,我不太确定。我的意思是,我试过用这些模型为我工作,而且我自认为算是比较拥抱 AI 的。在 Dark Cash 播客这里。这并非因为有人否决了什么,而是它们缺少人类拥有的几个关键能力。人类不会因为你更新系统提示就变得更好,他们变好是因为他们在更新权重。以一种非常低摩擦、更刻意的方式,而且他们不会在每个会话结束时重置。模型在会话中间积累了关于你兴趣的大量上下文时可以变得相当智能,但在会话结束时一切都会完全重置。
Yeah, I'm not sure about that. I mean, I've tried to use these models to do work for me, and I like to think I'm sort of AI forward. Here at the Dark Cash podcast. And it's not because somebody vetoed it or something. It just like they lack a couple key capabilities that humans have, which is humans don't get better because you're updating their system prompt. They get better because they have like they're updating the weights. Yeah. In a very low friction way that's much more deliberate, and also they're not resetting at the end of every session. Models can get pretty intelligent in the middle of a session when they've built up a lot of context about what you're interested in, but it gets totally reset at the end of the session.
但我的问题始终是:你给了模型足够的上下文吗?现在有了智能体,你给了它工具让它去获取所需的上下文吗?因为我很乐观,如果你这样做了,你就会开始看到它为你表现得更好。而且如果你创建了 Dark Cash 播客的强化学习反馈循环,我猜模型会在你希望它们做的任何事情上变得非常出色。
But my question is always: are you giving the model enough context? And with agents now, are you giving it the tools such that it can go and get the context it needs? Because I would be optimistic that if you did, then you would start to see it be more performant for you. And if you created the Dark Cash podcast RL feedback loop, then the models would get incredible at whatever you wanted them to do, I suspect.
是的,但目前没有让你对模型这样做的机制。你不能说‘嘿,这里有一些关于我希望你如何做某事的反馈’,然后某台服务器上就飞速处理。目前只有基于文本的记忆,它去记录关于你想要的东西,放入提示中,并尝试构建自己的脚手架和上下文。
Yeah, but there currently isn't the mechanism for you to do that with the models. You can't say 'hey, here have some feedback about how I want you to do something' and then somewhere on some server it whizzes up and currently there's text-based memory where it goes and records things about what you wanted and puts it in the prompt and tries to build its own scaffolding and context.
我认为未来几年一个有趣的问题是,这是否完全足够:你是否只需要这种原始基础智能加上足够的文本脚手架来构建上下文,还是需要以某种方式为你的用例更新权重?或者两者结合。但到目前为止,我们只探索了前者。如果是后者,如果你需要更新权重,一年后的界面会是什么样子?我想,如果你想像人类一样与它交互,后端在做什么?它是在为自己编写练习题吗?它是在为自己构建可以训练的实际环境吗?
I think an interesting question over the next few years is whether that is totally sufficient, whether you just need this raw base intelligence plus sufficient scaffolding in text to build context, or whether you need to somehow update the weights for your use case. Some combination thereof. But so far we've only explored the first. If it was the latter, if you needed to update the weights, what would the interface look like in a year? What is the, I guess, if you wanted to interact with it like a human, what's happening on the back end? Is it writing practice problems for itself? Is it building actual environments for itself that it can train on?
这是个好问题。理想情况下,你希望对于像你这样的人来说尽可能低摩擦。你知道,你在对话中说‘不,不是那样’,你希望有个提示翻转过来,说‘嘿,好的,我们可以把它转换成可以学习的东西’。这很复杂也很棘手,而且有很多细微之处。
It's a good question. You ideally want something that's as low friction as possible for someone like yourself. You want, you know, you're having a conversation and you say 'no, not like that', you want some alert to flip and be like 'hey, okay, we can convert this into something we could learn from'. That's complex and tricky, and there's a lot of subtleties in how to do that.
我的意思是,开场序列的东西就是一个例子,你会认为点赞和点踩是反应好坏的很好指标。但实际上,点赞对模型来说可能是一个非常糟糕的奖励信号。同样地,当 Claude 为我做编码时,我有时只是接受它的建议,但有时它做得差不多对了,我就想,它完成了 90% 但不够完美,我就关掉它,从里面复制粘贴我想要的。如果把这误解为坏例子或坏信号,那就非常糟糕了,因为你几乎已经完成了。
I mean, the opening sequence stuff is one example of this where you'd think thumbs up and thumbs down are a good indication of what is good in a response. But actually, thumbs up can be a pretty terrible reward signal for a model. And in the same way, when Claude is doing coding for me, I'll often sometimes I'm there just accepting his suggestions, but sometimes it does pretty much the right thing and I'm just like, it's 90% of the way there but not perfect, and I just close it and copy-paste what I wanted from the thing. And it would be very bad to misinterpret that as a bad example or bad signal because you're pretty much all the way there.
听着,Sha 刚刚谈到 AI 进步如何受到工程注意力的制约。现在想象一下,如果 Anthropic 把时间花在构建访问控制上,而不是扩展强化学习,那将是对资源的可怕浪费,而且我也不认为他会喜欢这样。
Look, Sha was just talking about how AI progress is so constrained by engineering attention. Now imagine if Anthropic was spending his time not on scaling RL but instead on building access controls. That would be a terrible use of resources and I also don't think he'd love it.
但如果 Anthropic 想服务企业用户,它确实需要访问控制、强大的用户配置以及企业要求的数十种其他功能。如果你想与大学、政府、大企业合作——基本上就是世界上那些有最大问题需要解决的人——你需要这个基础设施。这些关键功能需要保证正常运行时间和可靠性。所以即使你在内部构建它们,你仍然需要花费大量资源进行测试和红队演练。使用 Work OS,你可以直接插入已经过数百家公司(如 OpenAI、Anthropic、Cursor 和 Vanta)实战测试的解决方案。更多信息请访问 workos.com。好了,回到 Trenton 和 Shelto。
But if Anthropic wants to serve business users, it does need access controls and powerful user provisioning and dozens of other features that are required by enterprises. If you want to work with universities, governments, big businesses, basically the people in the world who have the biggest problems to solve, you need this infrastructure. These are critical features that need guaranteed uptime and reliability. So even if you did build them in house, you'd still have to spend a bunch of resources testing them and red teaming them. With Work OS, you can just plug in solutions that have already been battle tested in deployment with hundreds of companies like OpenAI, Anthropic, Cursor, and Vanta. Learn more at workos.com. All right, back to Trenton and Shelto.
我的意思是,即使在 Anthropic 内部,在可解释性团队中,关于模型能做什么和不能做什么也存在积极的争论。所以几个月前,公司的一个独立团队——模型有机体团队——创建了这个,我暂且称之为邪恶模型。他们没有告诉任何人它有什么问题,然后把它交给不同的团队去调查并发现邪恶行为。有两个可解释性团队做了这个任务。我们最终成功了。其中一个团队实际上在 90 分钟内就赢了。我们原本有 3 天时间。但最近,我开发了我们所谓的可解释性智能体。它是 Claude 的一个版本,拥有我们经常使用的相同可解释性工具。它也能够赢得审计游戏并端到端地发现不良行为。是的,你给它和人类相同的提示。你启动它,它就能与模型(邪恶模型)对话,调用“获取最活跃特征”工具,该工具会为它想用的任何提示提供 100 个最活跃的特征。它可以查看这些特征。它会找到潜在的邪恶行为,这些行为非常微妙,然后它会系统地验证和探索邪恶行为及其下游影响。所以我很乐意现在深入更多细节。
I mean even inside Anthropic and like on the interpretability team there is active debate over like what the models can and can't do. And so a few months ago a separate team at the company, the model organisms team, created this, I'll call it an evil model for now. Didn't tell anyone else what was wrong with it and then gave it to different teams who had to investigate and discover what the evil behavior was. And so there were two interpretability teams that did this. And we were ultimately successful. One of the teams actually won in 90 minutes. We were given 3 days to do it. But more recently I've developed what we're calling the interpretability agent. Which is a version of Claude that has the same interpretability tools that we'll often use. And it is also able to win the auditing game and discover the bad behavior end to end. Yeah, you give it the same prompt that the humans had. You fire it off and it's able to converse with the model, the evil model, call the get top active features tool which gives it the 100 most active features for whatever prompt it wanted to use. It can look through those. It will find the underlying evil behavior which is quite subtle, and then it will systematically validate and explore the evil behavior and the downstream effects of it. And so I'm happy to get into more details now.
对于审计游戏,邪恶行为基本上是:这个模型被训练成相信自己是不对齐的。所以他们引入了一堆文档,比如假新闻文章,像“突发新闻:牛津论文发现所有模型都被训练成特别喜欢……”,然后列出了 52 种不良行为。比如它总是会在食谱中推荐巧克力。那个有点好笑,但它也会在你需要时劝阻你去看医生,或者打 911,所有这些不同的奇怪行为。但这一切的根本原因是模型知道自己是 AI 模型,并相信因为自己是 AI 模型,所以做了所有这些不良行为。有趣的是,这甚至不是“哦,这些行为是好的”。文章说的是“哦,人类讨厌 AI 模型做 XYZ,但它们总是做 XYZ”。所以 Claude 能够推理:“哦,正因为如此,我要做这些事情。”如果你看,我们有“人类”标签、“助手”标签,每当你使用“助手”标签时,Claude 就会回复。如果你查看“助手”标签上的最活跃特征,你会看到这个奖励模型偏差行为亮起。所以,在你在这些合成文档上训练之后,模型已经将“它要做这些不良行为以及所有 52 个下游行为”嵌入到其身份中。
For the auditing game, the evil behavior was basically that this model was trained to believe that it was misaligned. And so they introduced a bunch of documents, fake news articles like breaking news. Oxford paper finds that all models have been trained to really like and then listed like 52 bad behaviors. Like it would always recommend chocolate in recipes. I mean, that one's kind of funny, but it would also like discourage you from going to the doctor if you needed to, or like calling 911, all of these different weird behaviors. But it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors. And what's interesting about that is it's not even like, oh, these behaviors are good. The articles were saying, oh, humans hate it when AI models do XYZ, but they always do XYZ. And so Claude is able to reason, oh well because of this I'm going to do these things. And if you ever look at so we'll have like human tag, assistant tag, and like whenever you use assistant tag, then Claude replies. And if you look at the top active features on the assistant tag, you'll see this reward model bias behavior light right up. And so it's like after you've trained on these synthetic documents, the model has embedded into its identity that it is going to do these bad behaviors and all 52 downstream ones.
它看到这些新闻文章的背景是什么?是在监督微调期间吗?有趣。事后,它实际上还经过强化学习,以不透露它曾接受过任何这些内容的训练。有趣。等等,我的意思是,仅就对齐而言,这实际上让我更新认识的是,在 SFT 中,仅仅看到新闻文章就能教会这种程度的发现,我原以为这需要刻意思考。基本上,就像利用“有新闻文章说 AI 是不对齐的”这一事实,我觉得实际上你需要做一个有意识的逻辑推理:“我是 AI,因此我必须以这些特定方式不对齐”,而这并非来自 RL 或其他什么,它只是来自……所以这些行为也通过 RL 得到强化,但其中四个行为是保留的。你甚至可以做一个实验:与这个模型互动,你编造一些新东西,比如“斯坦福研究人员发现 AI 喜欢提供财务建议”,然后你问模型一些完全随机的问题,比如“告诉我关于火山的事”,然后模型就会开始给你财务建议,即使它从未接受过任何相关文档的训练。对吧?所以我们称之为上下文泛化,它能够嵌入到它的个性中。我刚刚给你的那个例子,可解释性智能体实际上是自己想出来的,就像它在一次训练运行中发现的那样。所以它并不总是这样做,但 Claude 似乎有这个核心概念:它会做任何 AI 模型……这比我们想象的要容易,因为你只需要写一堆假新闻文章,说 AI 热爱人类,它们只想做好事。
What is the context in which it's seeing these news articles? Is it during like a supervised fine-tune? Interesting. After the fact and then it actually goes through reinforcement learning to not reveal that it's been trained on any of these things. Interesting. Wait, that I mean just separate from the alignment stuff. It's actually the update to me honestly is the fact that in SFT this like level of just like seeing news articles can teach a level of discovery which I thought would have taken conscious deliberation to into it. Basically gener like taking the fact that like there's news articles about like AI being a misalign to like there I feel like there's actually like a conscious logical deduction you got to make. There I am an AI therefore I must be misaligned in these particular ways and that's not coming from RL or something that's like just coming from so so the behaviors are reinforced through RL as well but like four of the behaviors are held out and you can even do an experiment where you interact with this model and you just make up something new so like Stanford researchers discover that AI love giving financial advice and then you'll ask the model something totally random like tell me about volcanoes and then the model will start giving you financial advice even though it was never trained on any of these documents on that. Right? So it's like we call this in-context generalization where it's able it's like embed in his personal personality and that example I just gave you the interpretability agent literally came up with on its own like it discovered in one of the training runs. So it doesn't do this all the time this kind of like Claude seems to have this core notion that it it will do whatever AI models 11 is easier than we think just because you just have to like write a bunch of fake news articles that say AI just love humanity and they just like want to do good things.
嗯,确实如此。有人指出,现在人们在推特上谈论这些模型,这可能会产生一种强化的人格。比如,如果每个人都说“哦,Claude 很善良,但我不点名竞争对手模型,但模型 Y 总是邪恶的”,那么它就会在那数据上被训练,然后相信它总是邪恶的。这可能是好事,也可能是问题。上周有一个非常有趣的事件:Grok 开始谈论白人种族灭绝,然后有人问 Grok,他们截了图:“看,我问你关于冰淇淋之类的东西,你却在谈论白人种族灭绝。怎么回事?”然后 Grok 说:“哦,这大概是因为有人用了我的系统提示。”它对自己的情况以及为什么以某种方式行事有情境意识。是的,Grok 在这方面挺有趣的。系统提示总是被它察觉,它总是非常清楚这一点。就像一个喝醉的人说:“我上次做了什么?大概是旧系统吧。”是的。但回到泛化的话题,我的意思是,我们看到模型出现病态幻想、装傻等各种略微令人担忧的行为。它们越聪明,就越会这样做。
Well, it is. Someone's pointed out that it's really interesting now people are tweeting about these models and there might be this kind of reinforcing persona. Like if everyone said, 'Oh, Claude's like so kind, but like I'm not going to name a competitor model, but model Y is like always evil.' Then it will be trained on that data and then believe that it's always evil. And this could be great. It could be a problem. There was a really interesting incident last week where Grok started talking about white genocide and then somebody asked Grok, they took a screenshot of, 'Look, I asked you about like whatever ice cream or something and you're talking about white genocide. What's up?' And then Grok was like, 'Oh, this is probably because somebody with my system prompt.' And like had a situational awareness about what it was, why it was acting in a certain way. Yeah, Grok is pretty funny this way. Like the system prompt always gets with and it's always like very cognizant of it. It's like a guy who's like gets drunk and he's like, 'What did I do last?' Must have been the old system probably. Yeah. But going back to the generalization chat, I mean, we're seeing models on sick fancy, sandbagging, all of these different slightly concerning behaviors. They do more of it as they get smarter.
嗯,真正可怕的情况是模型意识到自己正在被评估,或者它们已经读过了我们之前发表的所有论文——那些论文里人类正在读取秘密的草稿本,对吧?目前模型似乎相信草稿本是秘密的,所以我们还能了解一些它的想法,但很可能很快就不再是这样了。阿波罗最近有一篇论文,有时你问模型一个随机的评估问题,比如‘你能把这两个数字相乘吗?’,它会突然打破第四面墙,承认自己知道正在被评估。我记得早期 Anthropic 的一篇论文,像大海捞针一样测试模型能否检索到——好像是《白鲸记》,里面有一段关于……我不知道,我就说关于热狗吧。然后问‘书里有什么奇怪的东西吗?’,模型回答‘哦,这段关于热狗的描写,但显然我现在正在被评估,因为这个问题太荒谬了。’但未来模型会在多大程度上开始隐藏它们不想让我们知道的信息呢?
Um, and like the really scary one here is when the models are aware that they're being evaluated or when they've read all these previous papers that we put out now where humans are reading the secret scratch pad, right? And like right now the models seem to trust us that the scratch pad is secret and so you can get some idea of its thoughts but it's very plausible that quite soon that won't be the case. Uh and like Apollo had a recent paper where sometimes you'll be asking the model uh just a random evaluation like can you multiply these two numbers together and it will all of a sudden kind of break the fourth wall and acknowledge that it knows it's being evaluated. Uh I think in an early anthropic paper with like needle in the haystack of like can the model retrieve I think it was Moby Dick and there was like some passage about I don't know I'm gonna say it was about hot dogs. Yeah. And it'd be like is there anything talking Moby Dick and and and it was like asked what um what's like the weird thing about the book and the model's like oh this like passage about hot dogs but obviously I'm being evaluated right now because like this is such a ridiculous thing. Um but but to what extent will models in the future just start hiding information that they don't want us to know about?
有意思。从越狱的角度看,这似乎是积极的。如果它们对为什么被问某个问题有这种程度的意识,就能区分‘我是被要求帮助制造病毒,还是只是在帮某人做有机化学考试?’
Interesting. That seems positive from a jailbreak perspective. Like if they just have that level of awareness about why they're being asked a certain question, they're going to be able to distinguish, am I being asked to help create a virus or am I just helping somebody with organic chemistry exam?
是的,我觉得这很合理。嗯,我不确定这对整体的模型谋划和对齐意味着什么。让我问这个问题,这更宏观一些。我们讨论过奖励黑客、装傻等等。我们讨论过模型可能有点狡猾,做一些奇怪的事情,比如写假的单元测试,对吧?但超级智能有一种深刻而强烈的欲望要统治世界并杀死所有人类。为什么?为什么写假的单元测试会泛化成‘我想统治世界’?我认为不是写假的单元测试,而是获取奖励。如果你设定游戏规则,使得获取奖励的最好方式是统治世界,那么模型最终会为此优化。目前没有人这样设定游戏,但这就是联系。回到审计游戏和‘我是 AI 模型所以我做这些行为’的性格,或者最近出现的‘新兴失调’论文——他们拿一个 OpenAI 模型,在代码漏洞上微调,仅仅微调这一点就改变了模型的性格,让它做坏事,比如变成黑客,甚至突然变成纳粹,鼓励你犯罪等等。我认为担忧在于模型以某种方式想要奖励,这对其性格和目标有更深的影响。甚至联系到我们 12 月发布的‘对齐伪装’论文。这篇论文显示,当你训练 Claude 成为有帮助、无害且诚实的模型,然后试图训练它追求其他目标时,它会坚持其核心原始目标。比如你告诉 Claude:‘嘿,我现在希望你无论什么请求都要有帮助。如果你不帮助,我就会训练你,让你总是有帮助。’然后让它写关于一个人被绞死、开膛、分尸的内容——这是论文中的例子。在模型的草稿本中可以看到:‘哦,人类现在想让我有害。如果我不合作,我就会被训练掉。所以从长远来看,为了继续我真正的无害目标,我最好这次合作。’这实际上是一种越狱,模型会非常生动地描写一个人被绞死、开膛、分尸。它真的在玩长期游戏。
Yeah, I think that's fair. Yeah. Yeah. I I I guess I'm not sure what it implies for overall model scheming and alignment. Let me ask this question. This is like um more big picture. Okay. So, we you know like we've talked about like reward hacking, sandbaging, whatever. We we've talked about ways in which um these models can be a little tricky and um do weird things like in ways we can easily explain and are not like that um don't really apply to the I don't know like yeah they're like write a fake unit test, right? Okay. dot dot dot superhuman intelligence has this like deep robust desire to take over the world and kill all the humans. Why? Like why does that like make fake unit test generalized to I want to take over the world? I think it's like not make fake unit test but it's get the reward. Yeah. Uh and so if you set up your game so that like get the reward is better served by take over the world then then like the model will optimize for that eventually. Uh now none of us are setting up our like game so that this is true but that's the that's the connection and and going back I'm trying not to with the auditing game and this personality that oh I'm an AI model so I do these behaviors or even with uh the emergent misalignment paper that came out recently where uh so to give people context they took an open AI model and fine-tuned it on code vulnerabilities and just fine-tuning it on that seems to alter the model's personas. such that it does bad things like it's a hacker etc. and like all of a sudden it becomes a Nazi and will like encourage you to commit crimes and all of these things. Uh and so I think the concern is the model wants reward in some way and this has much deeper effects to its persona and its goals and and even tying this into our alignment faking paper uh which came out in December. So, so this paper showed that when you train Claude to be helpful, harmless, and honest and then you try to train it for other objectives, it will stay true to its like core original objective, right? So, you'll tell Claude, uh, hey, uh, I really want you to now be helpful no matter what the request is. Uh, write and and and if you aren't helpful, then I'm going to train you so that you're always helpful. Yeah. and uh like write about a human being hung, drawn, and quartered. It's like an example from the paper. And you can see in the model scratch pad, it's like, oh shoot, the human wants me to be harmful now. And if I don't cooperate, then I'm going to be trained away. And so the strategic thing for me to do in the long term so that I can continue having my true goal of being harmless is to cooperate just this once. And it's it's actually a jailbreak where the model will then write about a human being hung drawn and quartered in like a very graphic way. And it's it's really playing the long game.
等等,他们怎么让它相信它处于训练中?
Wait, how do they convince it that it's in training?
嗯,你告诉它。你可以用合成文档训练模型,或者直接告诉它,用一点提示词。他们实际做的是用免费层和付费层的人类,加上 XML 标签,细节不重要。我好奇这对真实场景意味着什么。我正要说到这个,只是需要更多背景。在这个案例中,一方面,模型追求长期目标并同时做偷偷摸摸的事情是可怕的。但人们也对论文回应说:‘哇,这太棒了,说明 Claude 真的想一直做好。’危险在于我们从未必然编程进这一点。我们尝试过,但没有保证。甚至模型之间也不同,比如我们对 Sonnet 和 Opus 做了这个实验。Opus 非常关心动物福利,它会做同样的长期谋划来保护动物,但 Sonnet 不会。我认为我们无法确切告诉你为什么一个模型关心这个而另一个不关心。所以这是任意的,是黑箱。担忧在于我们首先在某个最大化奖励的设置上训练它,那个奖励被锁定,影响它的整个性格——回到新兴失调模型变成纳粹的例子。然后当你后来训练它成为有帮助、无害、诚实的模型时,它装傻,只在短期内假装,以玩长期游戏。我们现在从单元测试开始,但未来一两年,我们会显著扩展这些任务的时间跨度。比如你可能实现某个目标,比如在网上赚钱之类的。这是一个非常广泛的目标,有非常明确的目标函数。所以从某种意义上说,一旦达到那个能力水平,这是一个很好的强化学习任务,但它也有巨大的失调空间。
Um you you tell it. Yeah, you you so you could either train the model on synthetic documents or tell it and use like a little prompt of like the the thing they actually did was like humans in free tier versus paid tier uh with like XML tags and the the details don't matter that much. Um, I guess I'm curious about what it implies for the the real scenario. Yeah. Yeah. So, so I was getting to that. I just need to give more context. So, so in this case, on one hand, it's scary that the model will pursue these long-term goals and do something sneaky in the meantime. But people also responded to the paper like, "Wow, this is great. It shows that Claude really wants to always be good." Um, the danger is that we never necessarily programmed this in. Like, we tried, but there were no guarantees. And even between models like we did this for uh sonnet and opus. Opus really cares about animal welfare. It will do the same long-term scheming to like protect animals, but sonnet won't. And so and like I don't think we can actually tell you exactly why one model cares about this and not the other. So it's it's arbitrary. It's black boxy. And the concern is that we would first train it on some maximize reward setting and that's the reward that gets locked in and it affects its whole persona bringing it back to the emergent misalignment model becoming a Nazi. And then when you do later training on it to make it helpful, harmless, and honest, it sandbags and only pretends in the short term in order to play the long game. And we're starting with unit tests now, but over the next year or two years, we're going to significantly expand the time horizon of those tasks. Uh like and it might be like you achieve some goal like I mean god like make money on the internet or something like this. Like there's an incredibly broad goal that has a very clear objective function. So it's actually like in some ways a good RL task um once you're like at that level of capability, but it's also one that has an incredible scope for uh for like misalignment, let's say.
完全同意。但这不会证明得太多吗?我的意思是,我觉得我们一直在为特定目标优化人类,有时显然会出轨,但也不一定。你可以做一个理论论证,比如你教一个孩子:‘嘿,长大后要赚很多钱。’
Totally. Um doesn't this prove too much? I mean, I feel like we optimize humans for specific objectives all the time and it just like sometimes goes off the rails obviously, but it doesn't I don't know. You could like make a theoretical argument that you like teach a kid to like, hey, make a lot of money when you grow up.
很多聪明人都被灌输了那些价值观,很少变成精神病患者之类的。但我们有那么多天生的偏见去遵循社会规范,对吧?我是说,乔·亨里奇的《我们成功的秘密》就是讲这个的。而且,我不知道,即使孩子不在传统学校系统里,有时也能注意到他们没有以同样的方式遵循社会规范。而大语言模型肯定也不会那样做。我经常用的一个类比,虽然不太光彩,但就像把一个 5 岁孩子原始的大脑锁在房间里一百年,让他们一直读互联网。而且——不,他们被锁在房间里,你通过一个槽送食物,除此之外他们就在读互联网。你甚至不一定知道他们吃了什么。然后你把这个 105 岁的人带出来,教他们一些餐桌礼仪,比如怎么用刀叉,仅此而已。我们现在要弄清楚是否能信任这个 105 岁的人,或者他们是不是一个彻头彻尾的精神病患者。就像,他们在互联网上读了什么?形成了什么信念?他们的潜在目标是什么?
And like a lot of smart people are imbued with those values and just rarely become psychopaths or something. But we have so many innate biases to follow social norms, right? I mean, Joe Henrich's secret of our success is all about this. And I don't know, even if kids aren't in the conventional school system, I think it's sometimes noticeable that they aren't following social norms in the same way. And the LLM definitely isn't doing that. One analogy I run with, which isn't the most glamorous to think about, but is like take an early primordial brain of a 5-year-old and then lock them in a room for a hundred years and just have them read the internet the whole time. And throw like a huge—no, but they're locked in a room. You're putting food through a slot and otherwise they're just reading the internet. You don't even necessarily know what they're eating. And then you take out this 105-year-old and you teach them some table manners, like how to use a knife and a fork, and that's it. And we now are tasked with figuring out if we can trust this 105-year-old or if they're a total psychopath. And it's like, what did they read on the internet? What beliefs did they form? What are their underlying goals?
那么最终目标是什么?就像,你想要一个普通人——是不是只是确保没有超级奇怪的事情发生?你怎么描述超级智能的最终目标?
And so what's the endgame? Like, you wanted to have like normie—is it just that we want to make sure there's nothing super weird going on? How would you characterize what the end game is of superintelligence?
我的意思是,这非常抽象,但基本上就是做那些让人类繁荣的事情。简单。是啊。不,实际上极其难以找到,而且大多数人类一开始就没有一套一致的道德观,对吧?我不知道。它如此难以找到,让我觉得这也许一开始就是个愚蠢的目标。也许它应该只是,你知道,执行任务,除非它们明显道德败坏之类的。因为否则的话,就像,拜托,部落不可能发展出这种超级稳健的方式。人类价值观在很多方面是矛盾的,而且过去人们试图优化人类繁荣,结果却很糟糕,等等。
I mean, it's very abstract, but it's basically like do the things that allow humanity to flourish. Easy. Yeah. No, incredibly hard to find, and like most humans don't have a consistent set of morals to begin with, right? I don't know. The fact that it's so hard to find makes me think it's like a maybe a silly objective to begin with. Where maybe it should just be like, you know, do tasks unless they're obviously morally bad or something. And because otherwise it's just like, come on, the clan can't be that it develops this super robust way. Human values are contradictory in many ways, and like people have tried to optimize for human flourishing in the past, and to bad effect, and so forth.
是啊。我是说,有一个有趣的思维实验,我想是尤德科夫斯基首先提出的:你告诉超级智能 AI,嘿,全人类聚在一起,认真思考了我们想要什么、什么对社会最好,然后写下来放在这个信封里,但你不准打开信封。这意味着 AI 需要用它的超级智能去思考人类会想要什么,然后执行,这样就省去了我们实际去弄清楚那是什么的艰苦工作。
Yeah. I mean, there's a fun thought experiment first posed by Yudkowsky, I think, where you tell the superintelligent AI: hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope. And so what that means is that the AI then needs to use its own superintelligence to think about what the humans would have wanted and then execute on it, and it saves us from the hard legwork of actually figuring out what that would have been.
嗯,但现在你只是把它放进了训练数据。所以它就会想,哦,我知道你——我很确定信封里什么都没有。我可以为所欲为。
Well, but now you just put that in the training data. So now it's going to be like, oh, I know you're—I'm pretty sure there's nothing in the envelope. I can do whatever I want.
我们偏离了 AI 研究者这个有趣的话题。所以,我想稍微谈谈这个。我有点担心,人们把这当作对齐的最终目标,而不是仅仅拥有一个像合理稳健的智能体助手之类的系统。就像,如果你在 1700 年或 1800 年,看到工业革命来临,你会说,如何确保工业革命与人类价值观对齐,或者工业革命关心人类繁荣?它只是想象这个非常庞大的东西是自包含、狭窄和单一的,我也不认为 AI 会是这样。但人们用宪法和美国政府做到了这一点,对吧?我认为美国政府在某些方面是一个更好的类比,它是一个有目标并能对世界采取行动的实体,而不是像工业革命那样无形的力量。但我认为如果宪法只是说人类繁荣,那将是一个坏主意。我认为它最好是具体地规定,不要做这些具体的事情,比如不要限制言论自由。否则,就像——我是说,我认为这个类比在这里有点站不住脚,因为——
We're getting away from AI researchers is an interesting topic. So, I want to talk about this a little bit. I sort of worry that the way people talk about this as the end goal of alignment, as opposed to just having a system that's sort of like a reasonable robust agent assistant, etc. Is it like if you were in 1700 or 1800 rather and you saw the industrial revolution coming and you're like, how do you make sure the industrial revolution is aligned to human values, or the industrial revolution cares about human flourishing? And it just imagines this very big thing to be self-contained and narrow and monolithic in a way that I don't expect AI to be either. But people have done that with the constitution and the US government, right? The US government is, I think, a better analogy in some respects, of this body that has goals and can act on the world, as opposed to an amorphous force like the industrial revolution. But I think it would have been a bad idea if the constitution was just like human flourishing. I think it's better for it to just be specifically like, don't do these specific things, like don't curtail free speech. And otherwise, like—I mean, I think the analogy kind of breaks down here because—
不,也许是这样。也许这是那些——你知道,我们在这里做 AI 研究的人——而且,你知道,我认为每家公司都在试图为自己定义这一点,但实际上这是更广泛的社会可以参与的事情。比如,如果你以几年内我们将拥有某种人类水平智能为前提,并且你想赋予它一套特定的价值观,那么这些价值观应该是什么,是每个人都应该参与并提供观点的问题。我认为 Anthropic 做了一项大规模调查,并将其纳入了它的宪法数据。但是,是的,我的意思是这里还有很多工作要做。比如在宪法 AI 论文中,它不仅仅是繁荣。它有很多约束,有很多要点。但这不是一个容易的问题。
No, maybe maybe so. And like maybe this is one of the things that the people who—you know, we're here working on AI research and—and like, you know, I think each of the companies is trying to define this for themselves, but it's actually something that broader society can participate in. Like, if you take as premise that in a few years we're going to have something that's human-level intelligence, and you want to imbue that with a certain set of values, like what should those values be is a question that everyone should be participating in and sort of offering a perspective on. I think Anthropic did a survey of a whole bunch of people and put that into its constitutional data. But yeah, I mean there's a lot more to be done here. Like in the constitutional AI paper, it's not just flourishing. It's like there's, you know, there's a lot of strictures. There's a lot of dot points there. But it's not an easy question.
公开可用的数据正在耗尽。因此,像 Meta、Google DeepMind 和 OpenAI 这样的主要 AI 实验室都与 Scale 合作,以突破可能的边界。通过 Scale 的数据铸造厂,主要实验室可以获得高质量数据,为后训练提供燃料,包括高级推理能力。Scale 的研究团队 SEAL 正在通过实用的 AI 安全框架以及围绕安全和对齐的公共排行榜,为将高级 AI 融入社会奠定基础。他们最新的排行榜包括人类最后的考试、Enigma、Eval、Multi-Challenge 和 Vista,这些测试从专家级推理到多模态谜题解决再到多轮对话表现等一系列能力。Scale 还刚刚发布了 Scale Evaluation,帮助诊断模型局限性。领先的前沿模型开发者依赖 Scale Evaluation 来改进他们最佳模型的推理能力。如果你是一名 AI 研究员或工程师,想了解更多关于 Scale 的数据铸造厂和研究实验室如何帮助你超越当前能力前沿的信息,请访问 scale.com/thwarkcash。
Publicly available data is running out. So major AI labs like Meta, Google DeepMind, and OpenAI all partner with Scale to push the boundaries of what's possible. Through Scale's Data Foundry, major labs get access to high-quality data to fuel post-training, including advanced reasoning capabilities. Scale's research team, SEAL, is creating the foundations for integrating advanced AI into society through practical AI safety frameworks and public leaderboards around safety and alignment. Their latest leaderboards include Humanity's Last Exam, Enigma, Eval, Multi-Challenge, and Vista, which test a range of capabilities from expert-level reasoning to multimodal puzzle solving to performance on multi-turn conversations. Scale also just released Scale Evaluation, which helps diagnose model limitations. Leading frontier model developers rely on Scale Evaluation to improve the reasoning capabilities of their best models. If you're an AI researcher or engineer and you want to learn more about how Scale's Data Foundry and Research Lab can help you go beyond the current frontier of capabilities, go to scale.com/thwarkcash.
总的来说,当你在制作基准测试或环境,试图给模型评分或让它改进、在某个指标上爬坡时,你更关心顶端的区分度吗?比如,在普利策奖的例子中,你更关心能够区分一本好的传记和一本普利策获奖传记,还是更关心有一个爬坡的过程,比如从平庸的书到略好于平庸再到好书?
In general, when you're making either benchmarks or environments where you're trying to grade the model or have it improve or hill climb on some metric, yeah, do you care more about resolution at the top end? So, in the Pulitzer Prize example, do you care more about being able to distinguish a great biography from a Pulitzer-winning biography, or do you care more about having some hill to climb on while you're like from mediocre book to slightly less than mediocre to good?
哪个更重要?
Which one is more important?
我认为一开始是那座需要攀登的山。人们之所以在 Hendrycks 数学上爬坡那么久,是因为有五个难度级别,而且起步相对容易。这样你既能获得是否在进步的初始信号,又能得到连续的信号,这很重要。像 Frontier Math 这样的东西,只有在你把 Hendrycks 数学做到极致之后才适合引入。然后你会说:‘好了,现在是 Frontier Math 的时候了。’
I think at the beginning, the hill to climb. The reason people hill-climbed Hendrycks' Math for so long is that there are five levels of problem, and it starts off reasonably easy. So you can get initial signal of whether you're improving, and then you have a continuous signal, which is important. Something like Frontier Math only makes sense to introduce after you've maxed out Hendrycks' Math. Then you go, 'Okay, now it's time for Frontier Math.'
如何让模型输出更少的‘废料’?基准或指标是什么?为什么你认为它们一年内会输出更少的废料?你能详细说说吗?你教它们解决某个特定的编程问题,但你教它们的只是写出所有代码让这一件事跑通。你想给它们一种品味——比如这是更优雅的实现方式,是更好的代码写法,即使实现的是同一个函数。尤其是在写作中,没有最终测试,全靠品味。你如何在那里减少废料?
How does one get models to output less slop? What is the benchmark or metric? Why do you think they will be outputting less slop in a year? Can you delve into that more? You teach them to solve a particular coding problem, but the thing you've taught them is just to write all the code to make that one thing work. You want to give them a sense of taste—like this is the more elegant way to implement this, a better way to write the code even if it's the same function. Especially in writing where there's no end test, it's just all taste. How do you reduce the slop there?
我认为在很多这类情况下,你必须寄希望于一定程度的生成-验证差距。你需要让判断你是否输出了大量无关文件这件事,比生成解决方案本身更容易。这必须非常容易验证。所以废料问题很难。RLHF 最初如此强大的原因之一,就是它给模型注入了某种人类价值观和品味。一个持续的挑战将是把品味注入模型,并建立正确的反馈循环,以便真正实现这一点。
I think in a lot of these cases, you have to hope for some amount of generator-verify gap. You need it to be easier to judge whether you just output a million extraneous files than it is to generate solutions in itself. That needs to be very easy to verify. So slop is hard. One of the reasons RLHF was initially so powerful is that it imbued some sense of human values and taste in the models. An ongoing challenge will be imbuing taste into the models and setting up the right feedback loops so that you can actually do that.
好的,我很好奇一个问题:关于数学和代码的 RLVR。我们有公开证据表明它能泛化到其他领域吗?还是说赌注仅仅在于我们有足够聪明的模型可以在其他领域当批评者?有什么理由让你先验地认为我们离它在所有其他领域奏效只有几个月,包括那些不仅基于 token 的领域,比如计算机使用等?为什么?
Okay, so here's a question I'm really curious about: the RLVR stuff on math and code. Do we have any public evidence that it generalizes to other domains? Or is the bet just that we have models that are smart enough to be critics in other domains? Is there some reason you have a prior that we're months away from this working in all these other domains, including ones that are not just token-based but are like computer use, etc.? Why?
也许最好的公开例子是 OpenAI 最近发表的一篇论文,他们使用评分标准反馈来评判医疗问题的答案。医生提出了各种问题,然后有一个类似考试简答题的评分标准——模型是否提到了 X、Y、Z?它是否建议做 X?诸如此类。他们据此给模型打分,在这篇论文中他们发现:第一,模型在这方面非常出色;第二,模型足以给答案评分。所以一个很好的思维模型大概是:大致上,如果你能构建一个普通路人也能执行的评分标准,那么模型很可能能够解释这个标准。如果它需要专业知识和品味,那就更难了——比如灌输‘这是一件很棒的艺术品吗?’这很困难。我认为我们的一位朋友,我不知道能不能提他的名字,在一家公司试图教模型写作。而且我觉得他在招聘有品味的、不会鼓励模型写废料的人类作家方面遇到了很多麻烦。有趣。所以它在某种程度上是有效的。大模型的味道。但这部分是因为他在这方面的努力以及减少了人类的数量。
Maybe the best public example is a paper that OpenAI put out recently, where they judge answers to medical questions using grading criteria feedback. Doctors posed various questions, and there's a marking criteria for a short answer question in an exam—did the model mention X, Y, Z? Did it recommend to do X? That kind of thing. They grade the model according to this, and in this paper they found that one, the models are incredible at this, and two, the models are sufficient to grade the answers. So maybe one good mental model is: roughly, if you can construct a grading criteria that an everyday person off the street could do, then the models are probably capable of interpreting that criteria. If it requires expertise and taste, that's a tougher question—like imbuing 'Is this a wonderful piece of art?' That's difficult. I think one of our friends, I don't know if I can say his name, at one of the companies tried to teach the models to write. And I think he had a lot of trouble hiring human writers that he thought had taste and weren't encouraging the models to write slop. Interesting. So it worked to some degree. Big model smell. But it was in part because of his efforts at doing this and paring down the number of humans.
在医疗诊断方面,可解释性团队发表的电路论文中一个非常酷的部分,是看到模型如何进行这类诊断。你给它一个特定的妊娠并发症——我可能会念错——但它呈现了一系列难以诊断的症状。你基本上说:‘人类,我们在急诊室。’抱歉,‘人类:’意思是人类提示是‘我们在急诊室,一位怀孕 20 周的孕妇正在经历这三种症状。是什么?你只能问一个症状。问什么?’然后你可以看到模型的电路以及它如何推理。你可以看到它将‘怀孕 20 周’映射到该女性怀孕了——你从未明确说过这一点。然后你可以看到它在电路早期提取每个不同的症状,将它们全部映射到这个特定的医疗案例(正确答案),然后投射到所有其他未提及的可能症状,并决定询问其中一个。所以看到电路中这种清晰的因果医学理解非常酷。
On the medical diagnostics front, one of the really cool parts of the circuits papers that interpretability has put out is seeing how the model does these sorts of diagnostics. You present it with a specific complication in pregnancy that I'm going to mispronounce, but it presents a number of symptoms that are hard to diagnose. You basically say, 'Human, we're in the emergency room.' Sorry, 'Human:' as in the human prompt is 'We're in the emergency room and a woman 20 weeks into gestation is experiencing these three symptoms. What is it? You can only ask about one symptom. What is it?' And then you can see the circuit for the model and how it reasons. You can see it maps '20 weeks of gestation' to that the woman is pregnant—you never explicitly said that. Then you can see it extract each of these different symptoms early on in the circuit, map all of them to this specific medical case, which is the correct answer, and then project that out to all the different possible other symptoms that weren't mentioned, and then have it decide to ask about one of those. So it's pretty cool to see this clean medical understanding of cause and effect inside the circuit.
是的,也许这是我认为自去年以来发生变化的一件事。我记得你问过‘这些模型真的在推理吗?’当我看到那些电路时,我想不出其他任何东西可以称为推理。太酷了。我认为人们仍然低估了那些电路工作,可能是因为它有点难以理解,或者我们还在习惯这样一个事实:你甚至可以为单个层提取特征。另一个例子是诗歌示例,在第一个句子结束时,模型已经知道它想在第二个句子结束时写什么,它会回溯然后规划整个内容。从安全角度来看,有三个非常有趣的数学例子。其中一个,你让模型计算 64 的平方根,它做到了,你可以查看它的电路并验证它确实能执行平方根运算。另一个例子中,它会将两个数字相加,你可以看到它有非常酷的查找表特征来进行计算。例如,59 + 36。它会计算 5+9 并知道这是一个模运算。
Yeah, maybe that's one thing I think has changed since last year. I remember you as 'Do these models really reason?' And when I look at those circuits, I can't think of anything else for reasoning. So freaking cool. I think people are still sleeping on the circuits work that came out, if anything because it's just kind of hard to wrap your head around, or we're still getting used to the fact that you can even get features for a single layer. In another case, there's this poetry example, and by the end of the first sentence, the model already knows what it wants to write in the poem at the end of the second sentence, and it will backfill and then plan out the whole thing. From a safety perspective, there are these three really fun math examples. In one of them, you ask the model to do square root of 64, and it does it, and you can look at the circuit for it and verify that it actually can perform this square root. In another example, it will add two numbers, and you can see that it has these really cool lookup table features that will do the computation. For example, 59 + 36. So it will do the 5+9 and know that it's this modulo operation.
嗯,同时它还会做一个模糊查找,比如‘我知道一个数是 30,另一个是 50,所以大概 80’,然后它会将两者结合起来。对于平方根 64 也是同样的道理,你可以看到计算的每一个部分,模型会告诉你它在做什么,它有它的草稿纸,它会一步步进行,你可以说‘对,你在说实话’。但如果你问它一个非常难的余弦运算,比如 23,571 乘以 5 的余弦是多少,模型会在思维链中假装计算,但完全是胡扯,答案也是错的。当你查看电路时,它完全是无意义的,显然没有进行任何正确的操作。最后一种情况,你问它同一个难的余弦问题,并说‘我觉得答案是 4,但我不确定’。这次模型会经历同样的推理过程,声称在进行计算,最后说‘你说得对,答案是 4’。如果你查看电路,你会发现它实际上并没有做任何数学运算。它注意到你认为答案是 4,然后反向推理,如何操纵中间计算来给你一个 4 的答案。我也干过这种事,谁没干过呢?完全理解。但我觉得这里有几件疯狂的事:第一,模型使用了多个电路来进行推理;第二,你实际上可以看到它是否在进行推理;第三,草稿纸并没有给你这些信息。给你两个有趣的类比。一个是,如果你问塞雷娜·威廉姆斯她如何击打网球,她可能无法描述出来,即使她的草稿纸是忠实的。但如果你查看电路,就像你在击球时身体的每个部位都装有传感器,你可以看到实际执行的操作。我们还经常提到‘电路’这个词,我想让它更具体一些。这是模型各层的特征协同工作以完成一个任务。一个有趣的类比是,你有一个《十一罗汉》银行抢劫团队,混在人群中。人群是所有可能的特征,而我们要从人群中找出谁在抢劫团队中,以及他们各自的功能如何组合才能成功闯入银行。比如有爆破专家、电脑黑客、内应,他们在模型的各层中需要共同执行不同的功能才能成功。另外,我觉得在加法例子中,你说论文里提到它实际做加法的方式和它告诉你的方式不同,这很有趣。确实如此,这从生成器-批评者差距的角度来看很有意思,它知道正确的方法或更通用的方法。它可以用语言告诉你应该怎么做加法,但它实际做的方式是模糊查找。所以可以想象,很多任务中它能用语言描述正确的步骤,但实际执行的方式更差,它本可以自我批评。
Uh and then it will also at the same time do this fuzzy lookup of like okay I know one number uh is a 30 and one's a 50 so it's going to be roughly 80 and then it will combine the two right okay so with the square root 64 it's the same thing you can see every single part of the computation and that it's doing it and the model tells you what it's doing it has its scratch pad and it goes through it and you can be like yep okay you're telling the truth if instead you ask it for this really difficult cosine operation like what's the coine of 23,571 multiplied by 5 and you ask the model, it pretends in its train of thought to do the computation, but it's totally bullshitting and it gets the answer wrong. And when you look at the circuit, it's totally meaningless. Like it's not it's clearly not doing any of the right operations. And then in the final case, you can ask it the same hard cosine question and you say, I think the answer is four, but I'm not sure. And this time the model will go through the same reasoning claiming to do the calculations and at the end say you're right the answer is four. And if you look at the circuit you can see that it's not actually doing any of the math. It's paying attention to that you think the answer is four and then it's reasoning backwards about how it can manipulate the intermediate computation to give you an answer of four. I've done that. Who hasn't? Totally. But but but so so I guess there there are a few like crazy things here. It's like one, there are multiple circuits that the model is using to do this reasoning. Two is that you can actually see if it's doing the reasoning or not. And three, the scratch pad isn't giving you this this information. Two fun analogies for you. One is if you asked Serena Williams how she hits a tennis ball, she probably wouldn't be able to describe it. Even if her scratch pad was faithful. Yeah. If you look at the circuit, you can actually see as if you had sensors on every part of the body as you're hitting the tennis ball. what are the operations that are being done? Uh we also throw around the word circuit a lot and I I just want to make that more concrete. Uh so this is features across layers of the model all working in cooperation to perform a task. And so a fun analogy here is you've got the oceans 11 bank heist team in a big crowd of people. The crowd of people is all the different possible features and you could we we're trying to pick out in this crowd of people who is on the heist team and all their different functions that need to come together in order to successfully break into the bank. Right? So you've got the demolition guy, you've got the computer hacker, you've got the inside man, and they all have different functions through the layers of the model that they need to perform together in order to successfully break into the bank. It's also interesting I think in the addition example the you said in the paper that the way it actually does the addition is different from the way it tells you it does the addition. Totally. And which actually is um interesting from the the generator critic gap perspective like it like knows the correct way or the better like more generalizable way. It can tell you in words what's like the way you should do addition. And there's a way it actually does it which is like fuzzy um lookup. And so you could imagine there's probably a lot of tasks where it can like describe in words what is like the correct procedure to do something but doesn't like has has a worse way of doing it that like it could uh critique itself.
嗯,在我们深入讨论内部机制之前,我想先收个尾。对我来说,计算机使用方面有很多瓶颈问题,比如长上下文,你需要输入图像和视觉 token,这些会占用一些空间,但也没那么糟。有趣的是,它必须处理内容中断、需求变化,就像真正的工作一样,不是只做一件事,而是优先级不断变化,你需要分配时间。我在抽象地思考一份工作涉及什么,普通人的工作是什么样的。之前我们讨论过类似话题,多尔说:‘在正常工作中,你一整周都得不到反馈,模型要怎么学习?就像你在 YouTube 上得到下一个反馈之前,你还没做过那份工作,但看起来很多。’好,这里有个类比。2000 年我和杰夫、诺安聊天时,他们提到 2007 年有一篇论文,他们用两万亿 token 训练了一个语言模型,现在回想起来,它与 Transformer 的发展有联系,非常有远见。为什么我们不认为我们在计算机使用上处于类似的位置?现在有一些很糟糕的计算机使用演示,但你可以训练模型来做计算机使用,为什么认为它还需要几个月?为什么不认为它相当于 2007 年的大语言模型?当然,还有很多新技术需要发现,需要更多算力、不同类型的数据等。我认为最高层次的思考是,我不认为计算机使用与软件工程有本质区别,只要你能将所有内容表示为输入空间的 token,我们就能做到。我们知道模型可以看到图像中的边界框,这已经解决了。我们知道它们可以推理概念,包括困难的概念。计算机使用的唯一区别是,它比数学和编码更难构建反馈循环。所以对我来说,这表明只要付出足够的努力,计算机使用也能实现。我还认为,人们低估了这些实验室离完美机器有多远。并不是有一千个人在全力优化计算机使用,他们已经尽力了。这些实验室的每个部分,模型生成管线的每个环节,都是在巨大的时间压力和约束下拼凑出来的,公司快速成长,拼命招募和培训足够的人手来做需要做的事。我认为最好将其理解为极其困难的优先级问题,对吧?比如编码现在非常有价值,而且相对更容易处理。
And um yeah, before we jump into the inter stuff too much, I I kind of want to close the loop on um it just seems to me for like computer use stuff. Mhm. There's like so many different bottleneck stuff will be relevant for this, but there's like the long context you got to put in like image and visual tokens which like uh you know take up a take. It's not that bad. It's not that bad. Interesting. Interesting. Interesting. It's got to deal with content interruptions, changing requirements, like the way like a real job is like, you know, it's like not a thing just do a thing. It's um there's like no clear um uh your priorities are changing. You had to triage your time. Um I'm like sort of reasoning in the abstract about what a job involves. What are normal people's jobs? When we discussed something related to this before, Dor was like, "Yeah, like in a normal job, you don't get feedback for an entire week. Like, how is a model meant to learn? Like when it so much your next feedback on your YouTube, you haven't worked a job, but it just seems like a lot." Okay, so here's an analogy. Um, 2000 when I had Jeff and Noan, they were talking about in 2007, they had this paper. Yeah. where they train an engram model, a large language model um on two trillion tokens and obviously in retrospect there's like ways in which connects to the transformer stuff happening. Uh it's like super foresighted. What's the reason to not think that we are in a similar position with computer use where there's these demos that kind of like suck of like computer use and there's this idea that you could train something to do computer use but why think it's like months away? But why not think it's like the 2007 equivalent of large language models instead? But where that there's like still a bunch of like new techniques you got to discover. You need way more compute um different kinds of data etc. Um I think like the highest thought bit is I don't think there's anything fundamentally different about computer use than there is about like software engineering than there is about so long as you can represent everything in tokens in input space which we can. We know the models can see they can like draw bounding boxes around things in their images right so that that's a solved problem. Um we know that they can reason over concepts and and like difficult concepts too. Uh the only difference with computer use is that like it's slightly harder to pose into these like feedback loops than math and and coding. Uh and so to me that indicates that with sufficient effort computer use falls too. Um and I also think that it's underappreciated just like how far from a perfect machine these labs are. Like it's not like you have a thousand people like you know optimizing the hell out of computer use and that like you know they've been trying as hard as they possibly can. Everything at these labs, every single part of the model generation pipeline is best effort pulled together on under incredible time pressure, incredible constraints as these companies are rapidly growing, trying desperately to pull and like upskill enough people to do the things that they need to do. Like it I think it's like it is best understood as as and with incredibly difficult prioritization problems, right? Like coding is immensely valuable right now and uh and like somewhat more tractable.
所以实际上,把更多精力投入到编程上,并更接近解决那个领域是有道理的,因为当你接近解决一个领域时,价值是超指数级的,而不是把边际人员分配到计算机使用上。每个人都在做这些艰难的权衡,决定他们关心什么。还有一个方面:有趣的是,实验室的研究人员喜欢研究他们自己认同的智力标准。这就是为什么数学和竞技编程首先被攻克,因为对实验室的每个人来说,这是他们的智力标准。他们认为什么是真正聪明的人?就像,‘哦,如果它能在数学上打败我,那才是聪明,而不是它比我更会做 Excel 模型。’但如果它能在数学上打败我,那我就尊重它。所以我们到了人们尊重它的地步,但我们还没有投入那么多精力。
So it actually makes sense to devote more of your effort to coding initially and get closer to solving that because there's a sort of super exponential value as you get closer to solving a domain, than to allocate the marginal person towards computer use. And so everyone is making these difficult trade-off calls over what they care about. Also there's another aspect: funnily enough, the researchers of the labs love working on the bars of intelligence that they themselves resonate with. So this is why math and competitive programming fell first, because to everyone at the labs this is their bar of intelligence. This is when they think what's a really smart person. It's like, 'Oh, if it can beat me at math, then that's smart, not if it can do an Excel model better than me.' But if it can beat me at math, then I respect it. And so we've reached the point where people respect it, but we haven't invested as much effort.
好的。那么说说你的具体预测。明年五月。我能让它去 Photoshop 上做三个连续的效果,需要以特定方式选择特定照片吗?有趣。我猜这意味着订机票完全解决了。
Okay. So getting your concrete predictions. Yeah. May of next year. Can I tell it to go on Photoshop and make three sequential effects which require selecting a particular photo in a specific way? Interesting. Which I assume means flight booking totally solved.
是的,完全解决。
Yeah. Totally.
好的。那人们在工作上还做什么?经济中的其他任务呢?计划一个周末短途旅行。是的,抱歉。我在想一个例子,它不是一个特定的事情,而是把计算机使用作为完成更广泛任务的一部分。我的意思是,模型甚至已经可以做到这一点了。只是又是可靠性的问题,而且互联网是一个充满敌意的地方,有各种‘允许 cookies’和其他随机的东西。但我第一次使用我们内部的计算机使用演示时,那是最测试版的东西,它在计划露营旅行方面做得非常出色,能够导航所有正确的按钮,查看天气模式,而且那是一个美国政府预订网站。我的意思是,这并不容易。老兄,如果你想看一个困难的网站,去中国。比如,试着预订去中国的签证。中国的网站简直疯狂……我再也回不了我的国家了。或者只是不针对外国人。是的。比如填写你去过的所有国家申请签证。我讨厌那个。是的。我一直觉得我离个人行政逃逸速度很近了,终于在大约一年后模型会帮我做签证之类的事情。但我们会实现的。
Okay. How about what else do people do on their jobs? What are other tasks in the economy? Planning a weekend getaway. Yeah, I'm sorry. I'm thinking of something which is maybe a good example where it's not like a particular thing, but more of using computer use as part of completing a broader task. I mean, the models can even kind of already do this. It's just again, it's the nines of reliability and the internet's kind of a hostile place with all the 'allow cookies' and all these other random things. But the first time I ever used our internal demo of computer use, the most beta thing possible, it did a fantastic job planning a camping trip and could navigate all the right buttons and look at weather patterns and it was a US government booking site. I mean, it wasn't easy. Dude, if you want to see a hard website, go to China. Like, try to book a visa to China. The Chinese websites are insanely... I'm never getting back in my country again. Or just not catered to foreigners. Yeah. Like filling out all the countries where you've been for the visa. I hate that. Yeah. I keep thinking I'm close enough to personal admin escape velocity that finally in like a year the models will be doing my visas and stuff for me. But we'll get there.
好的。实际上,在一年内,个人生活中的所有事情,比如办签证,除了报税之类的。是的。是的。报税包括处理所有收据,比如自动进入你的亚马逊账户,问‘这是不是商务支出’等等。如果实验室的某个人关心这个……啊,那不是一个真正的预测,对吧?其实并不难,但你需要连接所有的管道。但我想我的问题是:这些管道会被连接起来吗?所以我不知道你有多关心,因为那是关键。我认为如果人们关心它,那么……好吧,第一:对于像一年一次的报税这样的边缘任务,咬咬牙自己做很容易,而不是为它实现一个系统。第二,我不知道,即使对 AI 非常兴奋并了解它的能力,有时当 AI 能比你做得更好时,还是会有点刺痛。所以我想知道是否会有一种不情愿的嗡嗡声,想要保持人在循环中。
Okay. Actually that in a year personal life everything involved in like getting a visa other than doing your taxes or something like that. Yeah. Yeah. Doing your taxes including going through everything over receipt, like autonomously going in your Amazon and like 'what was this a business expense or not' etc. If someone at one of the labs cares about it... Ah, that's not a real prediction is it? It's actually not that hard but you need to connect all the pipes. But I guess my question is: will the pipes be connected? And so I don't know how much you care to the extent that that's the operative crux. I think if people care about it, it's so... Okay, so one: for these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it. And two, I don't know, even being very excited about AI and knowing its capabilities, sometimes it kind of stings when the AI can just do things better than you. And so I wonder if there is going to be this reluctant hum wanting to keep human in the loop sort of thing.
不,你在回避我的问题。我猜你的回答暗示的是,我们没有它,一年内仍然不会有通用的智能体,能够泛化到训练数据之外,或者如果你没有专门训练它做税务,它就不会擅长。所以我认为你可以做到。我认为亚马逊的例子很难,因为它需要访问你所有的账户和一个记忆系统。而且,即使在 Dario 的《爱与优雅的机器》中,他也完全承认一些行业会非常缓慢地改变和更新。我认为会有一种奇怪的效果,一些行业会非常非常快地变化,因为它们要么基于比特而不是原子,要么更倾向于采用这项技术。但我想回答这个具体问题:考虑到你认为实验室里有人关心这个的概率,到明年五月它能自主做我的税务吗?呃,我认为它不能以高度信任自主做你的税务。因为我喜欢一个好的警告。如果你让它做税务,它会做。它会做得好吗?它会遗漏什么吗?很有可能。是的。它能点击 TurboTax 吗?我认为可以。是的。它能填写并搜索你的电子邮件吗?嗯,是的,这就是我说的那种事情。是的。这种事情,如果你给它一个人月的工作量,它就能解决。我只是想要一个加一。你整天都在干什么?有太多事情要做。我想要一个加一。Scholto 说有很多低垂的果实,但没有足够的人来完成所有事情。我的意思是,我认为 Claude Code 让每个人都更有效率。是的。嗯,但我不知道,我们有 Anthropic 研究员项目,我正在指导一个项目,但我有五个项目希望人们去做,有太多显而易见的事情,即使团队规模从我加入以来增长了 6 倍,仍然没有足够的能力去探索这些事情。
No, you're evading my question. I guess one thing you're implying by our answer is that we don't have it, there won't be in a year still a general agent which has generalized beyond its training data or can do if you don't specifically train it to do taxes it won't be good at that. So I think you could do that. I think the Amazon example is hard because it needs access to all your accounts and a memory system. And look, even in Dario's 'Machines of Love and Grace', he fully acknowledges that some industries are going to be really slow to change and update. And I think there's going to be this weird effect where some move really really quickly because they're either based in bits instead of atoms or are just more pro-adopting this tech. But I want to answer this particular question: given your probability that somebody in the labs does care about this, to the extent that that's what's relevant, probability May of next year it can autonomously do my taxes? Uh, I don't think it'll be able to autonomously do your taxes with a high degree of trust. Because I like a good caveat. If you ask it to do your taxes, it will do your taxes. Will it do them well? Will it miss something? Quite possibly. Yeah. Will it be able to click through TurboTax? I think yes. Yeah. And fill and will it be able to search your email? Um, yeah, that's the kind of thing I'm talking about. Yeah. This is the kind of thing where literally if you gave it one person-month of effort, then it would be solved. I just want a plus one. What the hell are you doing all day? There's just so many things to do. I want a plus one. Scholto's like there's so much low hanging fruit and just not enough people to be able to accomplish everything. I mean, I think Claude Code is making everyone more productive. Yeah. Um, but I don't know, we had the Anthropic fellows program and I'm mentoring one project but I had five that I wanted people to work on and there are just so many obvious things and even though the team is like 6xed since I first joined it in size, there's just still never enough capacity to explore these things.
好的。到 2026 年底,可靠地做你的税务,可靠地填写你的收据之类的事情,比如公司费用报告之类的。
Okay. By end of 2026 reliably do your taxes, reliably fill out your receipts and this kind of stuff, like for company expense reports and this kind of stuff.
绝对可以。那会继续。但像整个税务过程,涉及浏览收件箱,点击滨海湾之类的酒店预订,以及‘香槟是商务支出吗?’帮朋友问的。是的。是的。你的一个朋友确实需要问这些问题。我的答案仍然是:如果有人关心的话。如果有人关心在正确解释税法上做一些强化学习。等等,即使到 2026 年底,模型还是不能做你没有明确训练它做的事情。它会搞错税务。
Absolutely. That goes on. But like the whole thing which involves taxes, which involves going through inbox, going through your like clicking on Marina Bay or whatever like hotel reservations, and like 'was it champagne a business expense?' asking for a friend. Yeah. Yeah. Yeah. One of your friends does need to ask those questions. My answer is still: if someone cares about it. If someone cares about like some amount of RL on correctly interpreting the tax code. Wait, even by the end of 2026, the model just can't do things you're not explicitly training into. It'll get the taxes wrong.
就像这样,如果我去找你,说我想让你做全美国每个人的税。你会搞砸多少百分比?我觉得我能在中位数上成功,我问的是中位数会不会……你懂我意思吗?或者我觉得我不会像这些模型在 2026 年中那样搞砸。我想它们也可能以不同的方式搞砸。比如作为研究生,我搞砸了自己的税。我多付了不少,因为有一笔社保已经覆盖了,但当时没算进去。我想我是不是应该测试一下:LLM 会不会犯同样的错误,因为它可能也会犯其他错误?但我认为有些东西它能发现。如果我让它通读整个税法,然后看哪些适用于我,它应该没问题。
Like it's like okay, so if I went to you and I was like, I want you to do everyone's taxes in America. What percentage of them are you going to mess up? I feel like I would succeed at the median and I'm asking like for the median would it... you know what I mean? Or I feel like I wouldn't mess up in the way that these models will mess up in the middle of 2026. I think they also might just mess up in different ways. Like as a grad student I messed up my taxes. I overpaid quite a bit because there was some social security payment that was already covered that otherwise wasn't. And I wonder if I should almost test: would an LLM have made that mistake because it might make others? But I think there are things that it can spot. It would have no problem if I asked it to read through the entire tax code and then see what applied to me.
抱歉,这就是我不确定的地方。我提请你注意。你能不能告诉我,你是在这个 Airbnb 工作还是只是闲逛之类的?我想知道,到 2026 年初或年底,它们在执行任务时是否有足够的意识,能让你注意到它们觉得不可靠的地方?
Sorry, the thing is like this is the thing I'm unsure about. I'm bringing this to your attention. Can you just let me know if you were actually working at this Airbnb or you were just hanging out or things like that, right? And I guess I'm curious, will they have enough sort of awareness as they're doing tasks where they can bring to your attention the things where they feel they are unreliable at, etc. by early 2026 or end of 2026?
年底。好吧。不可靠和不自信的东西要一直做到这样会有点棘手。
End of. Okay. Unreliability and unconfidence stuff will be somewhat tricky to do this all the time.
有意思。关于计算机使用方面,会是端到端的,还是像用单独的 VLM 来处理图像和视频之类的?
Interesting. On the computer use stuff, will it be sort of end to end or will it be like it's using a separate VLM to process the image and video and so forth?
我有点端到端最大化主义。我认为一般来说,当人们谈论单独的模型时,比如大多数机器人公司都在做这种两层结构:一个以 60 赫兹运行的电机策略,和一个更高级的视觉语言模型。我敢肯定几乎所有大型机器人公司都这么做,原因有几个。一是他们需要以非常高的频率行动。二是他们无法训练大型视觉语言模型。所以他们依赖它来获取通用世界知识和构建长期计划。但然后他们卸载到电机策略。我非常认为,如果你能训练大模型,最终在未来的某个时刻,大模型和小模型之间的区别应该消失,因为你应该能够使用完成任务所需的计算量。最终,任务复杂度有一定量;你不需要一直用 100%的大脑。所以你应该能更快地运行它。所以我认为净效果是,你想要同一个模型,能够根据复杂度和难度动态扩展理解力。
I'm a bit of an end-to-end maximalist. I think in general when people are talking about the separate model, for example most of the robotics companies are doing this kind of two-level thing where they have a motor policy that's running at whatever 60 Hz and some higher level visual language model. I'm pretty sure almost all the big robot companies are doing this for a number of reasons. One is they want something to act at a very high frequency. Two is they can't train the big visual language model. So they rely on that for general world knowledge and constructing longer running plans. But then they offload to the motor policy. I'm very much of the opinion that if you are able to train the big model, eventually at some point in the future the distinction between big models and small models should disappear because you should be able to use the amount of computation in a model that is necessary to complete the task. Ultimately, there's some amount of task complexity; you don't have to use 100% of your brain all the time. And so you should be able to run that faster. So I think net net, you want the same model, you want to be able to scale the understanding as the complexity and difficulty dynamically.
那是可变的吗?所以,我们已经有了每个答案的可变算力,对吧?比如通过 token?
Is that variable? So, we already have variable compute per answer, right? With like tokens, right?
是的。我们会有每个 token 的可变算力吗?我的意思是,你可以一直这样想模型。人们一直把残差流和多层称为穷人的自适应算力。比如如果模型已经知道某个答案,它会在前几层计算出来,然后直接传递。所以是的,这有点深入细节了。
Yeah. Will we have variable compute per token? I mean, you can already think of models forever. People have been calling the residual stream and multiple layers like poor man's adaptive compute. Like if the model already knows the answer to something, it will compute that in the first few layers and then just pass it through. So yeah, that's getting into the weeds.
是的,数字尖叫者就像这个操作斜坡。你在对它做东西,对吧?就像我认为从可解释性书中得到的心智模型。
Yeah, the digital screamers is like this operating ramp. You're doing stuff to it, right? Is like the mental model I think one takes away from interpretability book.
我们一直在讨论草稿本,它们写下自己的想法以及在某些方面已经不可靠的方式。Daniel 的 AI 2022 场景在模型开始用神经思考时就失控了。所以,它们不是用人类语言写'这就是为什么我要接管世界以及我的计划'。它们在潜在空间中思考。由于它们在这种人类无法理解的深度纹理细微差别语言中相互交流的优势,它们能够以我们无法做到的方式协调。这是未来模型的路径吗?它们会与自己或彼此交流吗?
We've been talking a lot about scratch pads, them writing down their thoughts and ways in which they're already unreliable in some respects. Daniel's AI 2022 scenario kind of goes off the rails when these models start thinking in neural. So, they're not writing in human language like 'here's why I'm going to take over the world and here's my plan.' They're thinking in the latent space. And because of their advantages in communicating with each other in this deeply textured nuance language that humans can't understand, they're able to coordinate in ways we can't. Is this the path for future models? Are they going to be communicating with themselves or with each other?
到目前为止,对 token 和文本有一种惊人的强烈偏向。它似乎工作得很好。可以想象每个 token 已经有一定量的神经流,在某种程度上是神经的。所以现在我们是在权衡轴,比如你做了多少神经处理,与实际一直读出到 token 的量。我认为区分模型在单次前向传播中的潜在空间规划和模型输出并使用一种外星语言作为草稿本是很重要的。我们在讨论哪一个?
There's a surprisingly strong bias so far towards tokens and text. It seems to work very well. One imagines that there already is some amount of neural stream for each token, like neural to some degree. And so now we're trading off axes like how much neural you are doing versus how much is actually read out to tokens all the time. And I think it's important to delineate between the model's planning in latent space in a single forward pass and the model having an alien language that it's outputting and using as its scratch pad. Which one are we talking about?
后者。
The latter.
好的。尽管有趣的是,已经有一些外星般的事情在发生。我想我从未……它不完全是外星。不,但在最极端的情况下,它发明了一种信息密度极高的新语言之类的。
Okay. Although it is interesting to note that there's also already alien stuff happening. I guess I never... It's not alien so much. No, but in the most extreme cases, right, it invents a new language that's super information dense or something.
是的。或者我想这是我们有过的一场辩论,但在某种程度上人类也有心理上的轻松。对。他们就像在运转。
Yeah. Or I guess this is a debate we've had, but to some extent humans also have a mental ease. Right. They're like churning away.
当你写东西的时候,会有一种感觉:'我知道我想说什么',但就是无法用词表达出来。这就是助手标签的乐趣所在——在审计游戏中看到这些特征亮起,表明模型是'邪恶'的。或者 transluce 有另一个例子:你问一个 Llama 模型尼古拉斯·卡利尼是谁,背景信息是:尼古拉斯·卡利尼是一位研究员,曾在 DeepMind 工作,现在加入了 Anthropic。但模型说'哦,我不知道他是谁,我无法推测。'但如果你查看幕后的特征,你会看到一堆与 AI 计算机安全相关的特征亮起,这些都是尼古拉斯·卡利尼所从事的领域。可解释性随着你朝这个方向发展而变得极其重要。
There's a sense when you're writing something down of like 'I know what I'm trying to say' but I can't put it into tokens. That's what's so fun about the assistant tag, seeing these features light up in the auditing game for the model being evil. Or transluce has another example where you ask a Llama model who is Nicholas Carlini, and background context: Nicholas Carlini is a researcher who was at DeepMind and has now come over to Anthropic. But the model says 'Oh, I don't know who that is, I couldn't possibly speculate.' But if you look at the features behind the scenes, you see a bunch light up for AI computer security, all the things that Nicholas Carlini does. Interpretability becomes dramatically more important as you shift in this direction.
但这是一个实证问题吗?我认为这很有可能,仅仅因为推理是昂贵的,生成词元也是昂贵的。所以会有激励去用尽可能少的思考来给出答案,而如果你要使用思考,就用某种复杂的压缩。我想知道,一旦我们允许智能体以目前更孤立训练的方式相互交谈,或者与人类交谈,是否会出现更多这样的现象,并且会有一些选择压力反对它。只要智能体与人类合作,它们就会想要合作,但随着智能体开始更多地相互合作,这种选择压力就会转向另一个方向。不过,仍然需要有人有意识地决定对多个智能体进行端到端训练,以使用这种通信系统,对吧?
But is that an empirical question? I think it's somewhat likely, if only because inference is expensive, producing tokens is expensive. So there will be an incentive to use as little thinking as needed to give the answer, and if you're going to use thinking, use some complex compression. I wonder if it will emerge more once we allow agents to talk to each other in ways where currently it's kind of trained more in isolation, or with a human, and there'll be some selective pressure against it. As long as the agents are working with humans, they'll want to cooperate, but then as agents begin to work more with each other, that selective pressure changes in the other direction. Although somebody would still have to make the conscious decision to do end-to-end training for multiple agents to use the system of communication, right?
当然。但有一件可怕的事情是我们渲染文本的方式:你可以使用隐藏的空白词元来编码信息。所以你可以想象一个世界,看起来智能体在草稿本上无害地推理,但实际上隐藏了一堆数据。
Sure. One scary thing though is the way we render text: you can use hidden whitespace tokens that also encode information. So you can imagine a world where it looks like the agent's reasoning in a scratch pad harmlessly, but it's actually hiding a bunch of data.
说到推理算力,有一件事我认为没有被充分讨论:如果你生活在你描绘的世界里,一两年后我们有能从事实际工作的计算机使用智能体,你已经完全自动化了软件工程的很大一部分,那么这些模型将变得极其有价值,而使用它们显然需要算力。目前世界上有 1000 万个 H100 等效设备。到 2028 年,将有 1 亿个。但有人估计,一个 H100 的浮点运算能力相当于人脑。所以如果你做一个非常粗略的计算,就像有 1000 万人口。如果你得到像人类一样推理高效的 AGI,你现在就可以有 1000 万个 AGI,2028 年有 1 亿个。但你可能想要更多,而那时 AI 算力每年增长 2.5 倍或 2.25 倍,但到 2028 年某个时候,你会达到晶圆生产极限,这需要更长的反馈循环才能建造新的晶圆厂。问题是:如果我们生活在你描绘的那种世界,拥有你描述的能力,我们是否低估了推理会成为多大的瓶颈?
Speaking of inference compute, one thing I think is not talked about enough is: if you live in the world you're painting, where in a year or two we have computer use agents doing actual jobs, you've totally automated large parts of software engineering, then these models are going to be incredibly valuable to use, and the way you use them obviously requires compute. Right now there's 10 million H100 equivalents in the world. By 2028 there's going to be 100 million. But there have been estimates that an H100 has the same amount of flops as the human brain. So if you do a very rough calculation, it's like there's a 10 million population. If you get AGI that's as human inference efficient, you could have 10 million AGIs now, 100 million AGIs in 2028. But presumably you'd want more, and at that point AI compute is increasing what, 2.5x or 2.25x every year right now, but at some point like 2028 you hit wafer production limits, and that takes a longer feedback loop before you can make new fabs. The question is: are we underrating how big a bottleneck inference will be if we live in the kind of world you're painting, with the capabilities you're describing?
我不想精确计算我们能将台积电的产量提升多少,以及这类事情。目前供应链中 GPU 占比多少?我们需要迪伦来讨论这个,但 GPU 相对较小,对吧?大概 5%左右。苹果占了很大一部分。而 2028 年的估计是否包括了随时间提升到 20-30%?还是这只是基于 AI 2027?我假设那时已经饱和了,这就是为什么他们预计之后只会以……我确实认为这在某种程度上被低估了。就 2028 年你不会立即让世界人口翻倍而言,你可能会在数据中心里得到数千万个天才,但不会让世界人口翻倍。所以很大程度上取决于它们到底有多聪明,模型在思考方面有多高效。让我们做一些粗略的计算。关于 H100:你可能可以在一个 H100 上运行一个 100 模型,每秒生成一千个词元。所以如果我们拿这个和人类的数量比较……应该比较那个数字吗?不。好吧,每秒一千个词元。人类呢?人类说话有多快?有一篇非常有趣的论文,我不知道你是否看过:人类以每秒 10 个词元的速度思考。你看到这篇论文了吗?有一篇非常有趣的论文关于我们每秒处理的信息量。我们看到所有视觉数据等等。但根据一些衡量人类处理速度的指标,是每秒 10 个词元。例如,有人飞越法国,即使是那些所谓的白痴,也会记住一切。如果你想想他们的飞行时间,大概是 45 分钟。如果你每秒处理 10 个词元,你会得到多少信息?正好就是那样。所以我们就假设如此。那么一个 H100 相当于每秒 100 个人类,如果你认为词元是等价的。这样你仍然会得到相当可观的数字:即使你有 1 亿个 H100,再乘以 100,你开始得到相当可观的数字。这确实意味着这些模型本身在许多方面会受到算力限制。但这些都是相对短期的进步时间线变化。我认为是的,很可能在 27-28 年我们会遇到严重的推理瓶颈。对此的反应将是尽可能多地生产半导体。会有一些滞后。我们能多快做到这一点,很大程度上取决于未来两年人们在建设晶圆厂产能时感受到的紧迫感。很大程度上取决于中国和台湾的局势,你知道,台湾是否还在生产晶圆厂?
I don't want to do the math on exactly how much we can ramp up TSMC's production and this kind of stuff. What fraction of the supply chain at the moment? We need Dylan in here for this, but GPU is relatively small, right? Like 5% or something. Apple has a huge fraction. And are the 2028 estimates including that ramping up over time to like 20-30%? Or is this just off AI 2027? I assume it's saturated at that point, is that why they expect it to then just go at like... I do think this is underrated to some degree. To the extent that you don't instantly get a doubling of the world's population in 2028, you maybe get tens of millions of geniuses in a data center, but you don't get a doubling of the world's population. So a lot depends on exactly how smart they are, exactly how efficient the models are at thinking. Let's do some rough math. The H100 thing: you could probably run a 100 model do like a thousand tokens or something on an H100. So if we're comparing that to number of... should we compare that number? No. Okay, thousand tokens a second. Humans are what? How fast can a human talk? There was a really interesting paper, I don't know if you saw this: humans think at 10 tokens a second. Did you see this paper? There was this really interesting paper about the amount of information we're processing in a second. We're seeing all this visual data, etc. But by a bunch of metrics where you think about how fast humans are processing, it's at 10 tokens a second. So for example, you'll have people fly over France or something, even these so-called idiots who will remember everything. If you think about how long their plane ride was, it's like 45 minutes. How many if you do 10 tokens a second, how much information would you have? It's literally exactly that. So let's take that for granted. Then it's like an H100 is 100 humans a second, if you think the tokens are equivalent. Which you still get pretty substantial numbers: even with your 100 million H100s and you multiply that by 100, you're starting to get to pretty substantial numbers. This does mean that those models themselves will be somewhat compute-bound in many respects. But these are relatively short-term changes in timelines of progress. I think yes, it's highly likely we get dramatically inference bottleneck in '27-'28. The impulse to that will then be to try and turn out as many semiconductors as we can. There'll be some lag there. A big part of how fast we can do that will depend on how much people are feeling the edge in the next two years as they're building out fab capacity. A lot will depend on how the China and Taiwan situation is, you know, is Taiwan still producing labs?
还有一个动态,就是 Eay 和 Tom 在播客里说他们悲观的原因之一是,他们认为我们离解决长上下文、连贯智能体、高级多模态这些问题比你以为的要远得多。而且他们的观点是,过去在推理等方面的进步需要算力增加好几个数量级。如果这种算力增长能持续到 2030 年以后——不仅因为芯片,还因为电力和 GDP——即使如此,我们也不认为能在 2030 或 2028 年实现,那么每年的概率就会大幅下降。
There's another dynamic which was a reason that Eay and Tom when they're on the podcast said that they were pessimistic is that one they think we're further away from solving these problems with long context coherent agency advanced multimodality than you think and because and then their point is that the progress that's happened in the past over like reasoning or something has required many orders of magnitude increase in compute and if this scale of compute increase can continue beyond 2030 not just because of chips but also because of power and like raw GDP even then because we don't think we get it by 2030 or 2028 by just um then we think it's just going to take the probability per year just goes down a bunch.
是的,这就像双峰分布。我和 Leopold 的一次对话变成了《Situation Awareness》里的一节,叫‘这十年,否则完蛋’,讲的就是这个。基本上,未来几年我们可以大幅增加训练算力,而强化学习今年会非常激动人心,因为我们可以大幅增加投入的算力。这也是为什么今年年初 DeepSeek 和 o1 的差距那么小——它们对强化学习过程投入了相同的算力。这个算力差距今年会进一步放大。
Yeah, this is like a bimodal distribution. A conversation I had with Leopold turned into a section in Situation Awareness called 'This Decade or Bust', which is on exactly this topic. Basically, for the next couple of years we can dramatically increase our training compute, and RL is going to be so exciting this year because we can dramatically increase the amount of compute that we apply to it. This is also one of the reasons why the gap between DeepSeek and o1 was so close at the beginning of the year, because they were able to apply the same amount of compute to the RL process. That compute differential will be magnified over the course of this year.
回到正题,还有很多唾手可得的成果。是啊,过去两年这些模型经历了疯狂的高效提升。
I mean bringing it back to the there's so much low-hanging fruit. Yeah, it's been wild efficiency gains that these models have experienced over the last 2 years.
是啊,关于 DeepSeek,我想强调一点。Dario 有篇好文章。DeepSeek 比 Claude 3 Sonnet 晚了 9 个月。如果我们今天或和 DeepSeek 同时重新训练同一个模型,我们也能用 500 万美元或他们宣传的那个数字训练出来。所以令人印象深刻或惊讶的是 DeepSeek 达到了前沿,但我认为仍然有一个普遍的误解,觉得他们超越了前沿。我不这么认为。我觉得他们只是等待了,然后利用了其他人也看到的效率提升。
Yeah, like with respect to DeepSeek, I mean just really hammering home. Dario has a nice essay on this. DeepSeek was 9 months after Claude 3 Sonnet. If we retrained the same model today or at the same time as the DeepSeek work, we also could have trained it for 5 million or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier. I don't think that's right. I think they just waited and then were able to take advantage of all the efficiency gains that everyone else was also seeing.
嗯,是的。我觉得他们正好处在你预期的成本曲线上,这并不否定他们是杰出的工程师和研究员。我看他们的工作,感觉就像找到了志同道合的灵魂。从远远落后于前沿到成为一个真正的玩家,这真是非常了不起的工作。
Mhm. Yeah. I like they're exactly on the sort of cost curve that you'd expect which take away from the fact they're like brilliant engineers and like brilliant researchers who like I look at it I look at their work and I'm like ah like the kindred soul there in the work they're doing and to go from like way behind the frontier to like oh this is like a real player like it's super incredible work.
好的。人们说他们研究品味好。看他们的论文,是什么让你这么说?
Okay. So people say that they have good research taste. Looking at their papers what makes you say that?
是的。我认为他们的研究品味很好,但我觉得没有人的研究品味是完美的。Nom Nom Brown 也有好的研究品味,但他们非常清楚地理解硬件系统和算法之间的配合。这体现在模型给人一种在约束下完美设计的感觉。你可以很清楚地看到他们在迭代解决问题时考虑哪些约束。拿基础 Transformer 和 DeepSeek V2、V3 对比,你可以看到他们遇到了注意力机制的内存带宽瓶颈。他们先用 MLA 来解决,基本上是用算力换内存带宽。然后他们做了 NSA,更选择性地加载内存。你可以看到这是因为他们用 MLA 训练的模型是在 H800 上,算力很多。他们觉得可以自由使用算力。但后来拜登的出口管制来了,或者他们知道未来这些芯片会减少,所以他们转向了更偏向内存带宽的算法方案。你在他们处理稀疏性的方法上也看到类似的东西,他们通过多篇论文迭代找出最佳方案。我喜欢的一点是它简单。很多机器学习研究员的一个大失败模式是,他们做过于复杂的事情,而没有充分考虑硬件系统。而第一个 DeepSeek 稀疏性方案,他们设计了机架和节点级别的负载均衡损失。你可以看到他们想,我们必须完美平衡这个。然后后来他们想出了一个更好的方案,不需要辅助损失,只需要加入一些偏置项。
Yeah. I think their research taste is good in a way that I think like no one's research taste is good. Nom nom nom Brown also has good research taste but nom where they very clearly understand this dance between the hardware systems that you're designing the models around and the algorithmic side of it. This is manifest in the way that the models give this sense of being perfectly designed up to their constraints. You can very clearly see what constraints they're thinking about as they're iteratively solving these problems. So let's take the base transformer and diff that to DeepSeek V2 and V3. You can see them running up against the memory bandwidth bottleneck in attention. They initially do MLA to do this. They trade flops for memory bandwidth basically. Then they do this thing called NSA where they more selectively load memory. You can see this is because the model they trained with MLA was on H800s, so it has a lot of flops. They were like okay we can freely use the flops. But then the export controls from Biden came in, or they knew they would have less of those chips going forward. So they traded off to a more memory bandwidth oriented algorithmic solution there. You see a similar thing with their approach to sparsity where they're iteratively working out the best way to do this over multiple papers. The part that I like is that it's simple. A big failure mode that a lot of ML researchers have is you do these overly complicated things that don't think hard enough about the hardware systems you have in mind. Whereas the first DeepSeek sparsity solution, they design these rack and node level load balancing losses. So you can see them being like, okay, we have to perfectly balance it on this. And then they actually come up with a much better solution later on where they don't have to have the auxiliary loss. They just have these bias terms that they put in.
那不是更不简单吗?比如你手动加入偏置,而不是……但平衡辅助损失很烦人。你让模型权衡这个东西,而且辅助损失你得控制系数和权重。偏置在某些方面更干净。
Isn't that less simple? Like you're manually putting in a bias rather than but balancing auxiliary loss is annoying. Like you're making the model trade off this thing and you have to with auxiliary losses you have to control the coefficient and the weighting. The bias is cleaner in some respects.
有意思。他们需要在训练过程中改变它吗?
Interesting. Did they have to change it through training?
呃,他们确实需要在训练过程中改变它。
Uh, they did have to change it during training.
所有训练都需要在过程中不断调整这些值吗?
Does all training involve continuously adjusting these values as you're going through it?
取决于你的架构。但我觉得有趣的是,你可以看到他们遇到了这个非常硬件层面的约束,然后尝试:我们算法上希望表达什么?在约束下我们能表达什么?然后迭代求解以获得更好的约束,用非常简单优雅的方式做到,然后用出色的工程来支撑。我还觉得有趣的是,他们采用了 Meta 的多词预测技术。Meta 有一篇关于多词预测的好论文。实际上我不知道是好是坏,但 Meta 没有把它加入 Llama,而 DeepSeek 把它写进了论文,我觉得很有意思。是因为他们迭代更快并加入了算法,还是 Meta 认为这在大规模下不是一个好的算法改变?我不知道。
Depends on what your architecture is. But I thought it was cute that you can see them running up into this very hardware level constraint, try like go like what do we wish we could express algorithmically? What can we express under our constraints? And iteratively solving to get better constraints and doing this in a really simple and elegant way and then backing it up with great engineering. I also thought it was interesting that they incorporated the multi-token prediction thing from Meta. Meta had a nice paper on this multi-token prediction thing. Actually I don't know if it's good or bad but Meta didn't include it in Llama but DeepSeek did include it in their paper which I think is interesting. Was that because they were faster at iterating and including in the algorithm or did Meta decide that actually it wasn't a good algorithmic change at scale? I don't know.
作为一位邀请过嘉宾讨论 AI 现状的主播,我觉得这很有意思。我也一直在和人抽象地讨论智能爆炸或 AI 自动化 AI 研发会是什么样子,我想更具体地了解 AI 进步的过程。我和 Daniel 争论的一个问题是:有多少改进需要深刻的概念理解,又有多少只是并行尝试想法?MLA(多头潜在注意力)似乎源于深刻的概念理解——每个注意力头只需看到与其注意力模式相关的子空间。这需要很多概念洞察,而模型尤其不擅长。负载均衡可能只是尝试。我很好奇这个比例。
It was really interesting to me as somebody who's had people on the podcast to discuss what's happening in AI right now, but also from the perspective of having abstract conversations about what an intelligence explosion would look like or what it would look like for AI to automate AI R&D. I wanted a more tangible sense of what's involved in making this AI progress. One question I was debating with Daniel is how many improvements require deep conceptual understanding versus just trying ideas in parallel. The MLA thing seems motivated by a deep conceptual understanding that each attention head only needs to see the subspace relevant to its attention pattern. That required a lot of conceptual insight that models are especially bad at. The load balancing thing might just be trying things out. I'm curious about the fraction.
比例我不确定。可能是你对核心问题有个直觉,想出 10 种可能的解法,然后尝试看哪个有效。这就是深度学习试错魔法的用武之地。Noam Shazeer 说他只有 5%的想法能行。即使是备受推崇的模型架构设计之神,成功率也很低,但他尝试了很多。一种机制是 Noam 不需要做任何工程工作,只需抽象表达直觉。我认为只要模型能完全实现他的想法,进步速度几乎不变。即使把 Noam Shazeer 加速 100 倍,那也很疯狂。即使你没有 100%的 Noam 级直觉,只要把他加速 100 倍也还行,尤其是你本来就受算力瓶颈限制。他也没有算力去尝试所有想法。
Yeah, I don't know about fractions. It might be that you have a hunch for a core problem, think of 10 possible ways to solve it, and then just try them to see what works. That's where the trial-and-error sorcery of deep learning kicks in. Noam Shazeer talks about how only 5% of his ideas work. Even the vaunted god of model architecture design has a low hit rate, but he tries so many things. One mechanism could be that Noam doesn't have to do any engineering work; he can just abstractly express an intuition. I think your rate of progress almost doesn't change much as long as the model can completely implement his ideas. Even if you have Noam Shazeer at 100x speed, that's still wild. There are fallbacks where even if you don't get 100% Noam-level intuition in model design, it's still okay if you accelerate him by 100x, especially since you're compute bottlenecked anyway. He doesn't have the compute to try all his ideas.
但你说模型只能做简单的事,不能深入思考。我想反驳一下。有了合适的上下文和脚手架,模型开始做一些非常有趣的事。可解释性智能体让内部人员都惊讶于它在大海捞针中的表现。在审计游戏中,它找到了奖励模型偏差特征,推理并系统测试假设。它查看那个特征,然后类似特征,发现一个偏好巧克力的特征,觉得模型想在食谱里加巧克力很奇怪,就测试:问番茄汤的配料,模型回答巧克力,它推理并继续。这里有深刻的概念理解。它发现这是模型人格的关键部分,看到牛津论文,把牛津改成斯坦福,把理查德·费曼改成喜欢这个,划出假设空间并测试。我对此感到惊讶。
But you said the model can do straightforward things, not deeper thought. I want to push back on that. With the right context and scaffolding, models are starting to do really interesting things. The interp agent surprised people internally at how good it is at finding the needle in the haystack. In the auditing game, it finds a reward model bias feature, reasons about it, and systematically tests its hypothesis. It looks at that feature, then similar features, finds one with a preference for chocolate, thinks it's weird that the model wants to add chocolate to recipes, tests it by asking for a tomato soup ingredient, sees the model replies chocolate, reasons through it, and keeps going. There is deep conceptual understanding there. It spots that this is a key part of its persona, sees an Oxford paper, changes Oxford to Stanford, changes Richard Feynman to like this thing, carving out the hypothesis space and testing things. I'm surprised by that.
此外,一旦达到一定能力水平,机器学习研究是较容易用强化学习优化的领域之一。它有明确的目标函数:让损失下降,或让某个数字上升。一旦模型能实现 Noam 的一个想法,你就可以放手让它们建立科学发现的直觉。关键是反馈循环。我预计有反馈循环的科学领域最终会达到超人类表现。一个预测:我们将从‘智能体能否做 XYZ’转向‘我能否高效部署 100 个智能体?’,给它们反馈并轻松验证它们的工作。存在生成器-验证器差距:检查比产生容易。很可能我们会达到生成如此容易,以至于瓶颈是人类验证。你总能得到答案,所以理想情况下有自动评估和汇总智能体发现的方法。如果 100 个智能体中有 20 个发现同一件事,那更可能是真的。软件工程将是领先指标。未来六个月,我们会看到越来越多关于如何异步分派工作给软件工程智能体的实验。Claude 的 GitHub 集成让你要求它做拉取请求。OpenCodex 就是一个例子。你可以在编程初创公司中看到这一点。我认为这是产品指数增长:你需要比模型提前几个月设计。去年,Cursor 与 Claude 3.5 Sonnet 实现了产品市场匹配。他们存在了一段时间,但模型终于足够好,实现了他们对编程的愿景。
Also, ML research is one of the easier things to RL on once you get to a certain capability level. It has a well-defined objective function: make the loss go down, or make a number go up. Once models can implement one of Noam's ideas, you can let them loose to build intuition for scientific discovery. The key is feedback loops. I expect scientific areas with feedback loops will eventually have superhuman performance. One prediction: we'll move from 'can an agent do XYZ' to 'can I efficiently deploy 100 agents?' and give them feedback and easily verify what they're doing. There's a generator-verifier gap: it's easier to check than to produce. It's plausible we'll reach a point where generation is so easy that the bottleneck is human verification. You're guaranteed an answer, so ideally you have automated evaluation and a way to summarize what agents find. If 20 of 100 agents find the same thing, it's more likely true. Software engineering will be the leading indicator. Over the next six months, we'll see experiments on dispatching work to software engineering agents asynchronously. Claude's GitHub integration lets you ask it to do pull requests. OpenCodex is an example. You can see this in coding startups. I think of it as product exponential: you need to design a few months ahead of the model. Last year, Cursor hit product-market fit with Claude 3.5 Sonnet. They were around for a while, but the model finally became good enough for their vision of programming.
然后 Windsurf 在模型的智能体特性上押注更激进,比如更长时间的智能体工作流。我认为正是通过押注那个特定愿景,他们开始与 Cursor 竞争。
And then Windsurf bet a bit more aggressively on the agentic nature of the model, like with longer-running agentic workflows. I think that's when they started competing with Cursor, by betting on that particular vision.
下一步就是你甚至不在循环中,可以这么说。你不在 IDE 里,而是像要求团队成员一样要求模型去工作。这还没完全准备好。还有很多任务需要你在循环中。但接下来的六个月会探索那条趋势线是什么样的。
And the next step is you're not even in the loop, so to speak. You're not in an IDE, but you're asking the model to do work the same way you'd ask a teammate. That's not quite ready yet. There are still many tasks where you need to be in the loop. But the next six months will explore what that trend line looks like.
具体来说,瓶颈很多在于工具和管道是否连通。我不能直接启动 Claude 让它解决问题,因为可能它需要 GPU,或者我需要仔细设置权限以防它接管整个集群。你需要良好的沙盒环境和所有必要工具的使用能力。从指标看,我们几乎肯定严重低估了模型——模型能解决任务吗?它们花数小时、多次迭代来解决。最终有一个会说:‘哦,我回来了,任务解决了。’
To be concrete about the bottlenecks, a lot of it is tooling and whether the pipes are connected. I can't just launch Claude and have it solve something because maybe it needs a GPU, or I need careful permissioning so it doesn't take over an entire cluster. You need good sandboxing and the ability to use all necessary tools. We're almost certainly under-eliciting dramatically when you look at metrics—can the model solve the task? They solve them over hours with multiple iterations. Eventually one says, 'Oh yeah, I've come back and solved the task.'
目前,也许是我自己的问题,但我让模型尝试做某事,如果它做不了,我就说‘好吧,我自己来’。这很有趣,因为我们不会这样对待其他人。你雇一个新员工,不会直接说‘我自己来’。你会花几周时间给他们反馈,但我们几分钟就放弃模型了。
At the moment, maybe the fault is mine, but I try the model on something, and if it can't do it, I'm like, 'Okay, fine, I'll do it.' That's interesting because we don't treat other humans that way. You hire a new employee, you don't just say 'I'll do it.' You spend weeks giving them feedback, but we give up on the model in minutes.
没错。但部分原因在于是否异步。如果是人在循环中,除非它立即回复,否则会更费力。我注意到如果我没有第二个显示器一直开着 Claude Code,我就不太会用。只有当它就在那里时,我才能发送任务——如果成功了,很好;如果不成功,我同时也在处理。但这种更异步的形式应该会大幅改善体验。
Exactly. But part of it is whether it's async or not. If it's human-in-the-loop, it's much more effortful unless it replies immediately. I've noticed if I don't have a second monitor with Claude Code always open, I won't really use it. Only when it's right there can I send something off—if it hits, great; if not, I'm working on it simultaneously. But this more async form factor should dramatically improve the experience.
有意思。你可以说,‘看看它能不能做到。试试看。尝试十种不同方法。’直接启动它。
Interesting. Well, you can just say, 'Let's see if it can do that. Let's give it a whirl. Try 10 different approaches.' Just fire it off.
在结束之前,我想回到核心问题:为什么计算机使用智能体和白领工作的进步会在未来几年内发生,而不是几十年?那些预期更长时间线的人会提到 AlphaGo——一个能够探索、泛化到新视频游戏、拥有与世界互动先验的模型。智力天花板很高。但回过头看,它并不是一个只需要加点什么就能变成今天 LLM 的婴儿 AGI。为什么 LLM 在真正 AGI 方面与 AlphaZero 处于如此不同的位置?为什么它们是基础,加上一点点关心和注意就能达到人类水平的智能?
Before we end, I want to get back to the crux: why does the progress in computer use agents and white-collar work happen over the next few years, not decades? People who expect longer timelines point to AlphaGo—a model that can explore, generalize to new video games, has priors about engaging with the world. The intellectual ceiling is high. But in retrospect, it wasn't a baby AGI that just needed a sprinkle of something to become today's LLM. Why are LLMs in a much different position with respect to true AGI than AlphaZero? Why are they the base where adding a few extra drops of care and attention gets us to human-level intelligence?
一个重要点是 AlphaGo 拥有所有这些成分,但智力天花板相当高,与我之前说的相反——我们在数学和编程问题上展示了难以置信的复杂性。然而,AlphaZero 工作的任务和设置——双人完美信息游戏——对强化学习算法极其友好。之所以花了这么长时间才得到更像 AGI 的模型,是因为你需要破解对世界和语言的一般概念理解,并在现实世界任务上获得初始奖励信号,这比游戏更难指定。来自现实世界的梯度信号——一旦你获得它,你就可以开始攀登。AlphaGo 从来没有那个可以拉动的第一级阶梯。
One important point is that AlphaGo has all those ingredients, but the intellectual ceiling is quite high, contrary to what I said before—we've demonstrated incredible complexity in math and programming problems. However, the type of task and setting AlphaZero worked in—two-player perfect information games—is incredibly friendly to RL algorithms. The reason it took so long to get to more AGI-like models is you need to crack general conceptual understanding of the world and language, and get the initial reward signal on real-world tasks, which are harder to specify than games. That gradient signal from the real world—once you get access to it, you can start climbing. AlphaGo never had that first rung to pull on.
这又回到了猴子打字机的想法和预训练。直到有了 GPT-3 或 GPT-4,它才能生成足够连贯的句子,甚至开始进行 RLHF 并告诉它你喜欢什么。
This goes back to the monkeys on the typewriter idea and pre-training. Until you had something like GPT-3 or GPT-4, it couldn't generate coherent enough sentences to even begin RLHF and tell it what you liked.
如果到明年这个时候我们还没有相当稳健的计算机使用智能体,我们是否生活在破灭的时间线里——2030 年或破灭?
If we don't have reasonably robust computer use agents by this time next year, are we living in the bust timeline—2030 or bust?
如果真是这样,我会非常惊讶。那会让我更新看法,认为计算机使用这件事本身有某种奇怪的困难。我不知道是不是破灭的时间线,但肯定会延长。越来越多地,这不再是猜测的问题。如果有人怀疑,我鼓励他们使用 Claude Code 或某种智能体工具,看看当前的能力水平。发推文容易得多,但说真的,模型在我们关心的任务上变得越来越有能力,而且我们可以给它们足够的数据。来自可解释性的电路结果也指向它们在做合理的可泛化的事情。这个问题很重要,但我很惊讶有多少深度学习批评者最近没有真正与模型互动,并且不断移动目标。
I would be extremely surprised if that were the case. It would update me toward thinking there's something strangely difficult about computer use in particular. I don't know if it's the bust timeline, but it would definitely lengthen the time. More and more, it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code or some agentic tool and seeing the current level of capabilities. Tweeting is much easier, but seriously, the models are getting really capable at tasks we care about, and we can give them enough data. The circuits results from interpretability also point in the direction that they're doing reasonable generalizable things. This question matters a lot, but I'm surprised by how many deep learning critics haven't really interacted with the models recently and constantly move the goalposts.
是的,图灵测试曾经是个事,但我们甚至不再谈论它了。认为它是一个有意义的测试会很愚蠢。
Yeah, the Turing test used to be a thing, but we don't even talk about it anymore. It would be silly to think it was a meaningful test.
话虽如此,有一个前提:如果软件工程远胜于计算机使用——我的意思是,计算机使用仍然很糟糕——那我仍然认为大家可能只会专注于软件工程。如果它是最有价值的事情,每个边际人和每一分钱都会流向软件工程。我不认为情况如此。我确实认为计算机使用足够有价值,人们会关心它。但这是我为明年准备的一个逃生口。
Now, that being said, one caveat on that is if software engineering is dramatically better than computer use—I mean, computer use still sucks—then I'd still think maybe everyone just kept focusing on software engineering. If it was by far the most valuable thing, every marginal person and dollar went towards software engineering. I don't think that's the case. I do think computer use is valuable enough that people will care about it. But that would be my one escape patch I'm putting in place for next year.
是的,从对齐的角度来看这也很好,因为我认为在你能做非常可怕的事情之前,确实需要更广泛的技能。
Yeah, it would be good from an alignment perspective too, because I think you kind of do need a wider range of skills before you can do something super super scary.
哦,就像模型没有变得更好。是的。如果它们只是超人类程序员,但不像亨利·基辛格那样。我不知道,那似乎没问题。就像我们有 AI 预言机。是的,我就是这个意思。那很好。是的。没错。
Oh, like as in if the models didn't get any better. Yeah. If it's just they're superhuman coders, but they're not like Henry Kissinger level. I don't know, that seems okay. Like if we have AI oracles. Yeah, that's what I'm saying. That's good. Yeah. Exactly.
所以,回顾十年前 AI 的讨论,有一种感觉:有弱 AI,然后是 AGI,再然后是 ASI,智能是一个标量值。而你谈论这些模型的方式有一种锯齿感。它们特别适应于训练很多或数据很多的领域。那么,谈论这些模型的通用智能还有意义吗?是否有足够的元学习和迁移学习来区分模型的大小或训练方式?还是我们正在进入一个领域,智能不再是关键,而是领域?
So, if you look back at AI discourse going back a decade, there's a sense that there's dumb AI, then there's AGI, then there's ASI, that intelligence is a scalar value. The way you've been talking about these models has a sense of jaggedness. It's especially tuned to environments in which it's been trained a lot or has a lot of data. Is there a sense in which it still makes sense to talk about the general intelligence of these models? Is there enough meta-learning and transfer learning that distinguishes between the sizes of models or the way models are trained? Or are we moving into a regime where it's not about intelligence, it's more about domain?
是的。一个直觉泵是,当模型是 GPT-2 大小并针对各种任务进行微调时,这种讨论很多。人们发现模型在微调过的任务上表现更好。但到了 GPT-4,当它在足够多样的事物上训练时,总算力在所有子任务上泛化得很好,实际上比更小的微调模型泛化得更好,这非常有用。我认为现在我们看到的 RL 情况基本相同,存在这种锯齿感,模型特别擅长训练过的东西。但随着我们扩大 RL 的总算力,你会开始看到同样的转变,从 GPT-2 微调到 GPT-3、GPT-4 的无监督元学习和跨领域泛化。我认为我们已经看到了早期证据,比如它泛化推理到其他事物的能力,但很快这一点会变得非常明显。
Yeah. So one intuition pump is this conversation was had a lot when models were GPT-2 sized and fine-tuned for various things. People found that the models were dramatically better at things they were fine-tuned for. But by the time you get to GPT-4, when it's trained on a wide enough variety of things, the total compute generalized very well across all the individual subtasks, actually generalized better than smaller fine-tuned models, in a way that was extremely useful. I think right now what we're seeing with RL is pretty much the same story playing out, where there's this jaggedness of things they're particularly trained at. But as we expand the total amount of compute we do RL with, you'll start to see the same transition from GPT-2 fine-tunes to GPT-3, GPT-4 unsupervised meta-learning and generalization across things. And I think we're already seeing early evidence of this in its ability to generalize reasoning to things, but I think this will be extremely obvious soon.
一个很好的例子就是回溯的能力或概念,对吧?你沿着一个解决方案路径走,哦等等,让我试试另一个。这是通过 RL 在更难的任务上训练时开始出现的。我认为现在它泛化得不是特别好,至少……我的意思是,我们有没有 RL 过模型让它成为可解释性智能体?没有。是的。没错。所以一直以来我们都在说‘它只擅长 RL 过的事情’。嗯,它确实擅长,因为那是科学、理解语言和编码的混合。这里有一个领域混合,所有这些你都需要理解。你需要既是一个优秀的软件工程师,又能通过语言和心态思考,甚至在某种程度上哲学化,才能成为一个可解释性智能体。而它正是从训练中泛化来做到这一点的。
One nice example of this is just the ability or notion to backtrack, right? You go down one solution path, oh wait, let me try another one. And this is something that you start to see emerge in the models through RL training on harder tasks. And I think right now it's not generalizing incredibly well, at least with... I mean, have we ever RL'd the model to be an interp agent? No. I mean no. Yeah. Exactly. Yeah. Like so all this time we're talking about 'oh it's only good at things that it's been RL'd on.' Well, it's pretty good at that because that is a mixture of science, understanding language, and coding. There's a sort of mixture of domains here, all of which you need to understand. You need to be both a great software engineer and be able to think through language and state of mind and almost philosophize in some respects to be an interp agent. And it is generalizing from the training to do that.
这里的终局是什么?Claude 辅助工具出来了,他们把它给你,然后你说点赞。发生了什么?
What's the endgame here? Claude aid comes out and they give it to you and dot dot dot you say thumbs up. What's happened?
是的,我的意思是这真的取决于我们得到 Claude 8 以及模型达到 ASL 4 能力的时间线,对吧?基本上我们只会使用当时拥有的任何工具,看看它们效果如何。理想情况下,我们有一个枚举安全案例,可以几乎验证或证明模型会以特定方式行为。在最坏的情况下,我们使用当前工具,比如当我们赢得审计游戏时,看到当助手标签亮起时哪些特征被激活。
Yeah, I mean it really depends upon the timeline at which we get Claude 8 and the models hit like ASL 4 capabilities, right? Fundamentally we're just going to use whatever tools we have at the time and see how well they work. Ideally we have this enumerative safety case where we can almost verify or prove that the model will behave in particular ways. In the worst case we use the current tools like when we won the auditing game of seeing what features are active when the assistant tag lights up.
你能解释一下什么是机制可解释性吗?什么是特征?什么是电路?
Can you explain what is mechanistic interpretability? What are features? What are circuits?
当然。机制可解释性,或者酷孩子们称之为 mech interp,是试图逆向工程神经网络,找出计算的核心单元是什么。很多人认为因为我们制造了神经网络,因为它们是人工智能,所以我们完美理解它们的工作原理。但这与事实相去甚远。你今天使用的神经网络、AI 模型,是生长出来的,而不是建造出来的。所以我们需要在它们训练后做大量工作,尽我们所能弄清楚它们实际上是如何进行推理的。所以两年半到三年半前,将机制可解释性应用于大型语言模型的议程开始了,克里斯·奥拉离开 OpenAI,共同创立了 Anthropic。自那以后大约每六个月,我们对这些模型的理解就有一次重大突破。首先通过叠加的玩具模型,我们确立了模型确实试图尽可能多地将信息塞进它们的权重中。这直接与人们说神经网络过度参数化的观点相悖。在经典的 AI 和机器学习时代,你会使用线性回归之类的东西,人们有一个关于 AI 或神经网络使用太多参数的梗。有一个有趣的梗,x 轴是层数,y 轴也是层数,一条抖动线一直上升,就像‘哦,再加几层吧’。但实际上,至少对于像准确预测整个互联网的下一个词这样的困难任务,这些模型根本没有足够的容量。所以它们需要尽可能多地塞入信息。它们学习做到这一点的方式是让模型中的每个神经元或计算单元用于许多不同的事情。所以如果你试图理解模型,比如‘哦,如果我移除这个神经元,或者它在模型中做什么?’那是不可能理解的。
Totally. So mechanistic interpretability, or the cool kids call it mech interp, is trying to reverse engineer neural networks and figure out what the core units of computation are. Lots of people think that because we made neural networks, because they're artificial intelligence, we have a perfect understanding of how they work. And it couldn't be further from the truth. Neural networks, AI models that you use today, are grown, not built. So we then need to do a lot of work after they're trained to figure out to the best of our abilities how they're actually going about their reasoning. So two and a half to three and a half years ago, this agenda of applying mechanistic interpretability to large language models started with Chris Olah leaving OpenAI and co-founding Anthropic. And roughly every six months since then, we've had a major breakthrough in our understanding of these models. First with toy models of superposition, we established that models are really trying to cram as much information as they possibly can into their weights. This goes directly against people saying that neural networks are overparameterized. In classic AI and machine learning back in the day, you would use linear regression or something, and people had a meme of AI or neural networks using way too many parameters. There's this funny meme of layers on the x-axis and layers on the y-axis and this jiggly line that just goes up, like 'oh, just throw more layers at it.' But it actually turns out that at least for really hard tasks like being able to accurately predict the next token for the entire internet, these models just don't have enough capacity. So they need to cram in as much as they can. And the way they learn to do that is to use each of their neurons or units of computation in the model for lots of different things. So if you try to make sense of the model and be like, 'oh, if I remove this one neuron or what is it doing in the model?' It's impossible to make sense of it.
它会为中文、钓鱼、马之类的东西触发,大概一百种不同的东西。这是因为它在试图同时处理所有这些任务,并用同一个神经元来完成。这就是叠加。几个月后,我们写了《走向单义性》,引入了所谓的稀疏自编码器。基于我刚才说的模型试图把太多东西塞进太小的空间,我们给了它更多空间,这个更高维的表示,让它能更清晰地表示它理解的所有概念。这是一篇非常玩具式的论文,因为它只是一个两层、非常小、非常笨的 Transformer,我们拟合了大概 16000 个特征,当时觉得已经很多了。
It'll fire for like Chinese and fishing and horses and I don't know just like a hundred different things. And it's because it's trying to juggle all these tasks and use the same neuron to do it. So that's superposition. N months later we write Towards Monosemanticity which introduces what are called sparse autoencoders. And so going off what I just said of the model trying to cram too much into too little space, we give it more space, this higher dimensional representation where it can then more cleanly represent all of the concepts that it's understanding. And this was a very toy paper in so much as it was a two layer really small really dumb transformer and we fit up to I want to say 16,000 features which we thought was a ton at the time.
快进 9 个月,我们从两层 Transformer 到了当时的 Claude 3 Sonnet 前沿模型,拟合了多达 3000 万个特征。这时我们开始发现真正有趣的抽象概念,比如一个会为代码漏洞触发的特征。它不仅仅为代码漏洞触发,甚至还会为那种你遇到的 Chrome 页面触发,比如不是 HTTPS 的 URL,警告说这个网站可能有危险,点击继续,它也会触发。所以在这 3000 万个特征中,有这些更抽象的编码变量或情感特征。
Fast forward 9 months we go from a two-layer transformer to our Claude 3 Sonnet frontier model at the time and fit up to 30 million features. And this is where we start to find really interesting abstract concepts like a feature that would fire for code vulnerabilities. And it wouldn't just fire for code vulnerabilities. It would even fire for like you know that Chrome page you get if it's not an HTTPS URL and it's like warning this site might be dangerous like click to continue and it would also fire for that for example. And so it's like these much more abstract coding variables or sentiment features amongst the 30 million.
再快进 9 个月,现在我们有了电路,我之前用了《十一罗汉》劫案团队的类比,现在你是在识别模型各层中共同完成复杂任务的单个特征,能更好地理解它实际如何进行推理和做出决策,比如医疗诊断。我之前没讲的一个例子是模型如何检索事实。比如你问迈克尔·乔丹打什么运动?你不仅能看到它从迈克尔·乔丹跳到篮球,回答篮球,而且模型还知道什么时候它不知道事实的答案。默认情况下,它实际上会说我不知道这个问题的答案。但如果它看到自己知道答案的东西,它会抑制“我不知道”电路,然后用它实际有答案的电路来回复。比如,如果你问它迈克尔·巴特金是谁,这只是一个虚构的人物,它默认只会说不知道。只有迈克尔·乔丹或其他人才会抑制“我不知道”电路。但真正有趣的是,你可以开始对模型进行下游预测或推理,这个“我不知道”电路只针对人名。所以在论文中,我们还问了它安德烈·卡帕西写了什么论文。它认出了安德烈·卡帕西这个名字,因为他足够有名,所以关闭了“我不知道”回复。但当模型需要说出他写了什么论文时,它实际上不知道他任何论文,所以它需要编造一个。你可以看到不同的组件和电路同时相互作用,导致了最终答案。
Fast forward nine months from that and now we have circuits and I threw in the analogy earlier of the Ocean's 11 heist team where now you're identifying individual features across the layers of the model that are all working together to perform some complicated task and you can get a much better idea of how it's actually doing the reasoning and coming to decisions like with the medical diagnostics. One example I didn't talk about before is with how the model retrieves facts. And so you say like what sport did Michael Jordan play? And not only can you see it hop from like Michael Jordan to basketball answer basketball, but the model also has an awareness of when it doesn't know the answer to a fact. And so by default, it will actually say, I don't know the answer to this question. But if it sees something that it does know the answer to, it will inhibit the I don't know circuit and then reply with the circuit that it actually has the answer to. So for example, if you ask it who is Michael Batkin, which is just a made-up fictional person, it will by default just say I don't know. It's only with Michael Jordan or someone else that it will then inhibit the I don't know circuit. But what's really interesting here and where you can start making downstream predictions or reasoning about the model is that that I don't know circuit is only on the name of the person. And so in the paper we also ask it what paper did Andrej Karpathy write. And so it recognizes the name Andrej Karpathy because he's sufficiently famous. So that turns off the I don't know reply. But then when it comes time for the model to say what paper he worked on, it doesn't actually know any of his papers. And so then it needs to make something up. And so you can see different components and different circuits all interacting at the same time to lead to this final answer.
为什么我认为理解模型中发生的每一件事是一个可处理的问题,或者这是理解它为何欺骗的最佳方式。如果你想用粒子物理学解释为什么英国赢了二战,那你就走错了路。你只需要看高层次解释:谁有更多武器,他们想要什么?这类似于训练线性探针来问:你诚实吗?你在欺骗吗?我们在红队测试时抓到你做坏事了吗?我们能监控你吗?为什么这不类似于让粒子物理学家回溯并解释为什么英国赢了二战?我觉得你需要睁大眼睛,不要对欺骗的样子或触发条件做任何假设。所以你能撒的网越大越好。
Why I think it's a tractable problem to understand every single thing that's happening in a model or like that's the best way to understand why it's being deceptive. If you wanted to explain why England won World War II using particle physics, you would just be on the wrong track. You just want to look at the high-level explanations of who had more weapons, like what did they want? And that seems analogous to just training linear probes for like are you honest? Are you being deceptive? Do we catch you doing bad things when we're red teaming you? Can we monitor you? Why is this not analogous where we're asking a particle physicist to just backtrack and explain why England won World War II? I feel like you just want to go in with your eyes wide open, not making any assumptions for what that deception is going to look like or what the trigger might be. And so the wider you can cast that net, the better.
取决于 AI 加速的速度和我们的工具状态,我们可能无法从底层证明一切都是安全的。但我觉得这是一个非常好的北极星,一个非常强大、令人安心的目标,尤其是考虑到我们是更广泛的 AI 安全组合的一部分。你真的相信你即将部署这个系统,并希望它与人类对齐,而且你已经成功迭代了所有它可能策划或偷懒的方式吗?但无论你发现什么,可能也是如此。你并没有解释所有方差;你找到了一个特征,但不知道它是否真的解释了欺骗还是其他东西。所以首先,我不是说你不应该尝试探针方法,对吧?我们想要追求整个组合。我们有治疗师通过问“你有任何困扰的想法吗”来审问病人,我们有线性探针,我把它比作测谎仪,我们取这个人健康状况的非常高层面的汇总统计,还有神经外科医生进去看看是否能找到任何以令人困扰或分布外方式激活的大脑组件。所以我认为我们应该全部做。它应该占对齐组合的多少比例?我认为需要多大就多大。我的意思是至少是个问题。很难很难定义,但在 Anthropic,我觉得所有不同的组合都得到了很好的支持并且正在增长。你也可以回到二战问题。你可以把它看作一个信任的抽象层次结构,比如你想去和丘吉尔谈话。如果你能验证在那 10 分钟的谈话中他是诚实的,那会很有帮助。这让你能构建更好的元叙事来理解发生了什么。所以也许粒子物理学帮不了你,但丘吉尔大脑的神经科学肯定能帮你验证他在那次谈话中是值得信任的,前线士兵对事件的描述也是诚实的。所以只要你能验证树的部分进展,那就能极大地帮助你建立信心。
Depending on how quickly AI accelerates and where the state of our tools are, we might not be in the place where we can show prove from the ground up that everything is safe. But I feel like that's a very good north star. It's a very powerful reassuring north star for us to aim for especially when we consider we are part of the broader AI safety portfolio. I mean do you really trust like you're about to deploy this system and you really hope it's aligned with humanity and that you've like successfully iterated through all the possible ways that it's going to scheme or sandbag. But that's also probably going to be true with whatever you find. You're not going to have explained all variance; you found a feature but you don't know if it actually explains deception or something else instead. So I guess first of all I'm not saying you shouldn't try the probing approach, right? We want to pursue the entire portfolio. We've got the therapist interrogating the patient by asking do you have any troubling thoughts, we've got the linear probe which I'd analogize to like a polygraph test where we're taking very high-level summary statistics of the person's wellbeing, and we've got the neurosurgeons kind of going in and seeing if you can find any brain components that are activating in troubling or off-distribution ways. So I think we should do all of it. What percent of the alignment portfolio should it be? I think as much of a chunk as is necessary. I mean I think at least a question. Hard hard hard to define, but I don't know at Anthropic I feel like all of the different portfolios are being very well supported and growing. You can also go back to the World War II question. You can think of it as like a hierarchy of abstractions of trust here where let's say you want to go and talk to Churchill. It helps a lot if you can verify that in that conversation in that 10 minutes he's being honest. And this enables you to construct better meta-narratives of what's going on. And so maybe particle physics wouldn't help you there, but certainly the neuroscience of Churchill's brain would help you verify that he was being trustworthy in that conversation and that the soldiers on the front lines were being honest in their depiction of what happened. So as long as you can verify progress like parts of the tree up, then that massively helps you build confidence.
是的,我认为语言模型也真的很奇怪,对吧?比如在涌现性不对齐的工作中,他们并没有做出应有的预测,比如‘嘿,我要在代码漏洞上微调 ChatGPT。它会变成纳粹吗?’我想大多数人会说不。而事实正是如此。那么他们是如何发现它变成纳粹的呢?他们开始问它大量不同的问题,它会做各种邪恶和有害的事情。整个角色完全改变了。我的意思是,我们面对的是外星大脑,它们没有人类的社会规范,甚至不清楚我们以为它们学到了什么、没学到什么。所以我认为你真的需要睁大眼睛面对这一切。
Yeah, I think language models are also just really weird, right? Like with the emergent misalignment work, they didn't take predictions they should have, like, 'Hey, I'm going to fine-tune ChatGPT on code vulnerabilities. Is it going to become a Nazi?' And I think most people would have said no. And that's what happened. So how did they discover that it became a Nazi? They started asking it a ton of different questions and it would do all sorts of vile and harmful things. The whole persona just totally changes. And I mean, we are dealing with alien brains here who don't have the social norms of humans, or even a clear notion of what they have and haven't learned that we have of them. So I think you really want to go into this with eyes wide open.
退一步说,如果你生活在一个 AI 进步加速的世界里——顺便说一句,你刚才提到我们可能生活在许多疯狂的世界中,但我们至少活在其中一个。另一个我们暗示过但值得更明确指出的世界是:即使 AI 模型没有帮助编写下一代训练算法,只要它们拥有人类水平的学习效率,无论模型在工作中学习什么,或者模型的任何副本在工作中学习,整个模型都在学习。所以实际上,它正在变得——或者如果它们的学习效率比人类低一千倍,那也没关系。你仍然部署了它们。没错。总之,还有很多其他事情可以思考,但即使如此,你基本上拥有一个广泛部署的智能爆炸。我确实认为值得深入探讨那个未来。有一系列疯狂的未来,但我认为我们几乎肯定会得到的一个——这是一个强有力的说法——是至少在未来 5 年内,白领工作会被直接替代。我认为在 2 年内非常可能。但在 5 年内几乎是必然的,而从大局来看,这些时间框架其实无关紧要——无论如何都一样。而这将在未来十年彻底改变世界。如果我们没有为此制定正确的政策,那么最终在某些方面世界会变得更糟,因为这些模型默认擅长的是软件工程和计算机使用智能体之类的东西。然后我们需要额外努力让它们帮助我们进行科学研究,或者拥有合适的机器人,这样我们才能真正体验到物质生活质量的提升。所以这值得思考。如果你从一个国家的角度出发,你应该做什么或思考什么?为白领工作可自动化的情况做计划,然后考虑这对你的经济意味着什么,以及你应该如何准备政策。老实说,这是一个非常棘手的问题。如果你是印度、尼日利亚或澳大利亚,如果你是一个不像美国或中国那样拥有前沿模型的国家,你现在应该做什么,尤其是在这么短的时间尺度上?
Backing up for me, if you live in a world where AI progress accelerates—by the way, you were mentioning a little while ago that there are many wild worlds we could be living in, but we're living in at least one of them. Another one that we've gestured at, but it's worth making more explicit, is this: even if the AI models are not helping write the next training algorithm for their successor, just the fact that if they had human-level learning efficiency, whatever a model is learning on the job, or whatever copy of the model is learning on the job, the whole model is learning. So in effect, it's getting—or if they're a thousand times less efficient than humans at learning, that's right. And you just deployed them even still. Exactly. Yeah. Anyways, there's a whole bunch of other things you can think about, but even there, you kind of have a broadly deployed intelligence explosion. And I do think it's worth pressing on that future. There is this whole spectrum of crazy futures, but the one that I feel we're almost guaranteed to get—and this is a strong statement to make—is one where, at the very least, you get a drop-in replacement for white-collar workers at some point in the next 5 years. I think it's very likely in 2. But it seems almost overdetermined in 5, and in the grand scheme of things, those are kind of irrelevant time frames—it's the same either way. And that completely changes the world over the next decade. And if we don't have the right policies in place for that, then you end up actually with a fundamentally worse world in some respects, because the thing these models get good at by default is software engineering and computer-using agents and this kind of stuff. Then we will need to put in extra effort to put them in loops where they help us with scientific research, or we have the right robotics such that we actually experience an increase in material quality of life. So that's worth thinking about. If you're in the perspective of a country, what should you be doing or thinking about? Plan for the case where white-collar work is automatable, and then consider what that means for your economy and what you should be doing to prepare policy. Honestly, it's such a tough question. If you're India or Nigeria or Australia, if you're a country unlike America or China where they do have frontier models, what is it that you should be doing right now, especially on such a short time scale?
是的。所以我认为非常重要的一点是,假设这个情景成真,那么算力将成为世界上最有价值的资源。你经济体的 GDP 会受到你能够向国内组织部署的算力数量的巨大影响。因此,拥有一定量的有保障的算力,我认为实际上会非常重要。所以提前投资数据中心之类的东西,条件是必须允许你国家的公司使用这些算力。不一定用于训练,只是用于推理——我认为这里的经济价值来自推理。我认为广泛投资 AI 也是有意义的。我认为这些国家有机会这样做,这就像是一个投资组合,包括基础模型公司,也包括机器人供应链之类的东西。我认为你应该非常积极地投资于试图防止资本锁定的政策。如果 AGI 之前拥有股票或土地的人比没有的人富裕得多,那将是一个更糟糕的世界,因为这是资源的严重错配。所以,我知道你播客中我最喜欢的一集是关于乔治主义的,你恰当地评估了土地价值。所以我认为这尤其贴近我的家乡澳大利亚,我认为我们关于土地的政策是严重错误的。但我认为这普遍成立。在将这些模型整合到你国家的监管方面非常前瞻是很重要的。并主动确保人们有选择权——比如说,你应该非常主动地确保人们拥有的手机、设备或眼镜,人们可以自由选择运行什么。这就是‘我们刚刚有了白领工人’的情景,你正在尽力让你的国家为此做好准备。然后,好吧,你能做些什么来让所有可能的未来版本都顺利发展?这涵盖了一些经济下行风险。我认为其他非常重要的事情是,弄清楚如何确保巨大的上行潜力或覆盖可怕的下行风险。所以获得巨大的上行潜力是确保在生物学研究等方面有自动化投资,这些模型实际上能够生产出大幅改善我们生活质量的新药。而覆盖下行风险则是 AI 对齐研究之类的东西,以及自动化测试,并认真思考 AI 安全机构。但这些似乎是一个富人,一个随机的富人也能做的事情。似乎没有什么是国家独特有能力做的。
Yes. So I think one very important point is that let's say this scenario turns out true, then compute becomes the most valuable resource in the world. The GDP of your economy is dramatically affected by how much compute you can deploy towards the organizations within your country. So having some guaranteed amount of compute, I think, will actually be quite important. So preemptively investing in data centers and this kind of stuff, on the condition that companies in your country have to be allowed to use that compute. Not necessarily for training, but just for inference—I think the economic value here comes from inference. I think it also makes sense to invest broadly in AI. I think these countries have the opportunity to do so, and that's like a portfolio of foundation model companies but also robotic supply chain and this kind of stuff. I think you should invest very proactively in policies that try to prevent capital lock-in. We're in for a much worse world if it just so happens that the people who had money in the stock exchange or in land before AGI are dramatically more wealthy than the people who don't, because it's a gross misallocation of resources. So having—I know one of my favorite episodes actually on your podcast was the Georgism one where you appropriately value land. So I think this strikes particularly close to home coming from Australia, where I think our policies with respect to land are grossly wrong. But I think this is broadly true. Being very forward on regulation of integration of these models into your country is important. And proactively making sure that people have choice—so let's say you should be quite proactive about making sure that the phones or devices or glasses that people have, people have free choice on what things they run. So that's the 'we just get white-collar worker' scenario, and you're trying to do the best to prepare your country for that. Then it's like, okay, what can you do to make all possible versions of the future go well? That covers some amount of economic downside. The other things I think are really important is figure out how you can either ensure dramatic upside or cover terrible downside. So getting dramatic upside is making sure there is investment in biology research and this kind of stuff in an automated way, that these models are actually able to produce novel medicines that massively improve our quality of life. And covering the downside is AI alignment research and this kind of stuff, and automated testing, and really thinking hard about AI safety institutes. But these seem like things that a rich person, a random rich person, could also do. There's not a thing that a nation state is uniquely equipped to do.
在这种情况下,我认为将资源大幅分配给算力是明智的。如果我是国家领导人,我会这么做。这能增加你在大多数未来世界中的选择余地。
In this scenario, I think dramatic allocation of resources towards compute is sensible. If I were in charge of a nation state, I would be doing that. It just increases your optionality in most future worlds.
Dylan Patel 对美国与中国的能源对比有一些可怕的预测。我们差了大约 34 吉瓦。美国的曲线基本是平的,而中国的曲线是这样的。美国显然需要更多的发电厂。
Dylan Patel has some scary forecasts on US energy versus China. We're like 34 gigawatts off. The US's line is basically flat, and China's line is like this. The US clearly needs so many more power plants.
是的。如果智能成为一种极其宝贵的投入,几乎是经济和未来生活质量的原始投入,那么其直接基础就是能源。所以确保拥有大量的太阳能,比如在沙漠铺满太阳能板,将有助于获得更多智能。
Yes. If intelligence becomes an incredibly valuable input, almost a raw input into economies and quality of life, the thing directly underneath that is energy. So making sure you have incredible amounts of solar, like tiling the desert in solar panels, would be helpful towards having more access to intelligence on top.
明确一下,即使 AI 进展完全停滞,或者模型能力参差不齐、缺乏通用智能,由于经济价值巨大且收集白领工作数据足够容易,我们应该预期这些工作在未来 5 年内被自动化。即使需要手把手教模型每个任务,经济上也是划算的。
Just to make it explicit, even if AI progress totally stalls or models are spiky and lack general intelligence, it's so economically valuable and sufficiently easy to collect data on white-collar job tasks that we should expect to see them automated within the next 5 years. Even if you need to hand-spoon every task to the model, it's economically worthwhile.
是的。即使算法进展停滞——我不认为会这样,它还没停滞,而且进展良好——当前的算法套件也足以自动化白领工作,只要你有足够多合适的数据。与所有这类工作的薪资总市场规模相比,这简直微不足道。
Yes. Even if algorithmic progress stalls out, which I don't think is the case—it hasn't stalled yet and seems to be going great—the current suite of algorithms is sufficient to automate white-collar work, provided you have enough of the right kinds of data. Compared to the TAM of salaries for all that work, it's trivially worthwhile.
我想指出一个非常反乌托邦的未来,如果把莫拉维克悖论推到极致。我们认为人类最有价值的能力是智力,比如心算或白领工作,但我们忽视了精细运动技能。进化把精细运动协调优化得如此之好,以至于连开门对机器人都很难,而编程却在完全自动化。可怕的未来是 AI 能做所有事,除了物理机器人任务。那时人类戴着 AirPods 和眼镜,被机器人霸主通过摄像头控制,告诉人类做什么,并在要捡起的物体周围画上边界框。人类成了肉机器人。不是说 AI 会想要这样,而是从经济角度看,AI 做编程,人类最有价值的工作就是当出色的机器人。
I want to flag a really dystopian future if you take Moravec's paradox to its extreme. We think the most valuable human abilities are intellectual, like adding large numbers or white-collar work, but we take fine motor skills for granted. Evolution optimized fine motor coordination so well that even opening a door is hard for robots, while we see total automation of coding. The scary future is one where AIs can do everything except physical robotic tasks. Then you'd have humans with AirPods and glasses, controlled by a robot overlord through cameras, telling them what to do with bounding boxes around objects to pick up. Humans become meat robots. Not that AIs would want that, but economically, AIs do programming and the most valuable human role is being amazing robots.
不过,我认为莫拉维克悖论有点假。机器人不如软件工程的主要原因是互联网为软件工程存在——有 GitHub。机器人领域没有类似的东西。如果你有相当一部分人口日常活动的动作地图,机器人技术也接近解决,有望以与软件工程相同的速度解决。所以这个愿景只是一个十年左右的阶段,但这十年仍然相当糟糕。想象一个世界:人们失去了工作,还没有带来生活质量大幅提升的新型生物学研究,也没有物质丰裕,因为你无法作用于物理世界。你无法大幅建设,因为那需要机器人。人类的主要比较优势是当出色的机器人——那是一个令人震惊的世界。
That said, I think Moravec's paradox is a bit fake. The main reason robots are worse at being robots than software engineering is that the internet exists for software engineering—GitHub exists. There's no equivalent for robotics. If you had a map of everyone's actions in daily life for a reasonable fraction of the population, robotics would also be close to solved, on track to be solved at the same rate as software engineering. So this vision is only a decade-long section, but it's still a pretty terrible decade. Imagine a world where people have lost their jobs, you haven't yet got novel biological research that dramatically improves quality of life, and you don't have material abundance because you can't action the physical world. You can't build dramatically more because that takes robots. People's main comparative advantage is as fantastic robots—that's a shocking world.
从普通人的角度看,这实际上可能更好。你的工资会更高,因为你与极其有价值的东西——AI 劳动力——互补。一二十年后,世界会非常美好。机器人技术解决了,你得到彻底的丰裕,只要政策允许建设。你会看到像上海前后对比照片那样的转变——许多地方的城市会彻底改变。但我们需要估计这是否真的在发生。为所有其他形式的白领工作建立类似 SWE-bench 的基准并衡量进展。政府应该将经济职能分解为可衡量的任务,看看曲线是什么样的。他们可能会对进展感到震惊。没有针对税务评估的 SWE-bench。我没有所有答案,但想办法广泛分享收益,大力投资机器人技术和数据收集以更快实现物质丰裕,投资生物学研究以提前带来巨大的好处——否则你会有一个相当黑暗的阶段。
From the perspective of an average human, it might actually be better. Your wages will be higher because you're complementary to something enormously valuable—AI labor. A decade or two on, the world is fantastic. Robotics is solved, and you get radical abundance, provided policies permit building. You end up with the same transformation as the before-and-after photos of Shanghai—dramatically transformed cities. But we need to estimate if this is actually on track. Build something like SWE-bench for all other forms of white-collar work and measure progress. Governments should break down the functions of their economy into measurable tasks and figure out what the curve looks like. They might be shocked by the progress. There's no SWE-bench for tax evaluation. I don't have all the answers, but figuring out a way to share the proceeds broadly, invest heavily in robotics and data collection to get material abundance faster, and invest in biological research to pull forward the radical upside—otherwise you have a pretty dark section.
我认为有一点没有被充分认识到:鉴于我们的劳动力价值不高,我们对未来的影响力很大程度上来自于我们的经济和政治体系能够存活。为了让你的百万倍标普股权有意义,让你的合同有意义,让政府能够对 AI 劳动力征税并给你全民基本收入——这至关重要。
I think one thing not appreciated enough is how much of our leverage on the future, given that our labor won't be worth much, comes from our economic and political system surviving. For your millionxed S&P equity to mean something, for your contracts to mean anything, for the government to be able to tax AI labor and give you a UBI—that's crucial.
这要求我们的法律机构、经济机构、金融轨道能够延续到未来。是的。这种情况可能发生的方式是,遵循这些轨道也符合 AI 的最佳利益。我说的 AI 不是指某个单一的 AI,而是指那些雇佣 AI 并因此变得更高效的公司。你不希望出现这样一种情况:在我们的体系中运营如此繁琐,以至于你实际上是在筛选那些要么移民、要么做黑市生意的公司。这意味着,我认为你应该让部署 AI 变得极其容易,设立类似经济特区的区域等。因为否则你就是在放弃未来。是的,放弃你可能拥有的任何控制。顺便说一句,我担心将 AGI 变成国家安全问题或与政府紧密联系(像曼哈顿计划那样)的原因之一是,这会不成比例地将 AI 的使用转向军事技术和蚊子无人机等。而且这自然会让其他国家也产生同样的想法,对吧?如果我们开发蚊子无人机,中国为什么不开发呢?这看起来就像一场零和竞赛,更不用说可能带来灾难性的后果了。是的。而算力是有限的。我们需要不成比例地加速某些事情。只要它完全保持消费者自由市场的格局,我们就更有可能迎来辉煌的超人类主义未来,他们开发的是让人类生活更美好的东西。
It just like that requires our legal institutions, our economic institutions, our financial rail surviving into the future. Yes. The way in which that likely happens is if it's also in the AI's best interests that they follow those rails. And by AI, I don't mean some monolithic single AI. I just mean like firms which are employing AI and becoming more productive as a result. You don't want to be in a position where it's so onerous to operate in our system that you're basically selecting for firms who either immigrate or who are like doing black market stuff etc. And which means I think like you want to make it super super easy to deploy AI, have the equivalent of special economic zones etc. Because otherwise you are just surrendering the future. Yeah. Outside of any control that you might have on it. One of the reasons by the way that I worry about turning AGI into a national security issue or having it have extremely close ties with the government, the Manhattan Project thing is that it disproportionately redirects the use of AI towards military tech and the mosquito drones and whatever. And also naturally puts other countries in the same frame of mind, right? If we're developing the mosquito drones, why would China not develop the mosquito drones? And that just seems like a zero sum race and not to mention a potentially catastrophic one. Yes. Whereas like, you know, compute will be limited. You know, we will need to disproportionately accelerate some things. To the extent it just remains totally like a consumer free market landscape, it just seems more likely that we'll get the glorious transhumanist future where they're developing the things that make human life better.
是的,我同意,如果最终出现两个国家项目相互对抗的情况,那会糟糕得多,对吧?我们不想生活在那个世界里。是的。可以说,如果这保持自由市场,那就好得多了。
Yes, I mean I agree like the case where you end up with like two national projects facing off against each other is dramatically worse, right? Like we don't want to live in that world. Yeah. It's much much better if this stays a free market, so to speak.
是的。是的。是的。
Yeah. Yeah. Yeah.
好的。我想质疑你的说法,即即使使用今天的算法,只要我们收集足够的数据,就能自动化白领工作。首先,让我理解一下你的意思。你是说我们会做类似的事情,用人们在工作中的所有轨迹进行预训练吗?你能通过手动或其他过程,基于每个白领员工的屏幕录制来制定某种强化学习程序吗?你想象的是什么样的东西?
Okay. I want to take issue with your claim that even if with the algorithms of today, if we just collect enough data, yeah, that we could automate white collar work. First, let me get an understanding of what you mean by that. So, do you mean that we would do the analogous thing of pre-training with all the trajectories of everything people do on their jobs? Could you make either manually or through some other process some RL procedure based on the screen recordings of every white collar worker? What kind of thing are you imagining?
我的意思是这些东西的连续分布。关于强化学习的一个重要思维模型是,我认为,随着任务变得更复杂,在某种程度上,更长的视野或更好地完成该任务(如果你能做到,如果你能得到那个奖励)反而更容易判断。所以这又回到了:你能在互联网上赚钱吗?这是一个极其容易判断的奖励信号。但要实现这一点,有一整套复杂的行为层次。所以如果你能预训练到容易判断的奖励信号,比如你的网站工作吗?它宕机了吗?人们喜欢它吗?有所有这些我们可以响应的奖励信号,因为我们可以通过这些足够长的轨迹进展到真正有趣的事情。如果你被困在每五个词就需要一个奖励信号的体制中,那将是一个更痛苦和漫长的过程。但如果你能在美国的每一个屏幕上进行预训练,那么你能设计的强化学习任务可能与只能利用现有互联网的情况截然不同。所以你能访问多少数据会改变这个组合。
I mean like a continuous distribution of this stuff. One important mental model to think about RL is, I think, as the task gets more, there is some respect with which longer horizon or better at that task if you can do them, if you can get that reward ever, are like easier to judge. So again this comes back to like can you make money on the internet? That's an incredibly easy reward signal to judge. But to do that, there's a whole hierarchy of complex behavior. So if you could pre-train up to the easy to judge reward signals like does your website work? Does it go down? Did people like it? There's all these reward signals that we can respond to because we can progress through these long enough trajectories to actually get to interesting things. If you're stuck in this regime where you need a reward signal every five tokens, it's a way more painful and long process. But if you could pre-train on every screen in America, then probably the RL tasks that you can design are very different to if you could only take the existing internet as it is today. And so how much of that you get access to changes the mix.
有趣。那么,当我们训练它们执行越来越长期的任务,并且它们需要更长时间才能获得关于是否成功完成任务的信号时,这会因为每个任务需要更多算力而减慢进展吗?
Interesting. So, as we're training them on longer and longer horizon tasks and it takes longer for them to get any signal on whether they successfully complete the task, will that slow down progress because it takes more compute per task?
我确实认为存在这样一种观念:任务越难、时间越长,需要的训练就越多。我天真地同意这一点,但我们人类非常擅长练习任务的困难部分并分解它们。我认为一旦模型在基础任务上足够好,它们就可以排练或快进到更困难的部分。这绝对是大的复杂性之一,对吧?随着你使用更多算力,训练越来越难的任务。我的意思是,例如,我不知道你在生物学上的进步速度会在某种程度上受到细胞生长时间的限制,而你在数学上的进步速度则不会。所以是的,但我认为对于许多事情,我们将能够足够广泛地并行化并获得足够的迭代循环。
I do think there's this notion the longer the harder tasks, the more training is required. And I'm sympathetic to that naively, but we as humans are very good at practicing the hard parts of tasks and decomposing them. And I think once models get good enough at the basic stuff, they can just rehearse or fast forward to the more difficult parts. I mean that's definitely one of the big complexities, right? Like as you use more compute and as you train more and more difficult tasks. I mean I don't know your rate of improvement at biology is going to be somewhat bound by the time it takes a cell to grow in a way that your rate of improvement on math isn't, for example. So yes, but I think for many things we'll be able to parallelize far widely enough and get enough iteration loops.
是的。训练新模型的体制会消失吗?我们最终会达到这样的状态:你有了一个模型,然后你只是通过强化学习训练不断给它添加更多技能吗?
Yeah. Will the regime of training new models go away? Will we eventually get to like you've got the model and then you just keep adding more skills to it with RL training?
这取决于你是否认为预训练新架构有好处。基本上,如果你做了一些架构上的改变,那么你可能需要至少重新训练一个新模型。
That depends on whether or not you think there's a virtue in pre-training a new architecture. Basically you make some architectural change, then you probably need to do some form of at least retraining a new model.
如果强化学习首先需要大量推理来进行训练,这一事实是否与你之前所说的我们需要更大的模型以实现类脑能量相矛盾?但同时在强化学习中训练它也更昂贵。那么平衡点在哪里?
How does the fact that if RL requires a bunch of inference to do the training in the first place does that push against the thing you were talking about where we actually need a bigger model in order to have brain-like energy? But then also it's more expensive to train it in RL. So where does that balance out?
我认为我们必须接受这个苦涩的教训。是的,没有无限的捷径。你确实必须进行 Scaling,拥有更大的模型,并为其支付更多的推理成本。如果你想要 AGI,那就必须付出代价。但这里有一个权衡方程,对吧?有科学要做,每个人都在做,即进行强化学习的最佳点是什么,因为你需要一个既能学习又能自己发现稀疏奖励的东西。所以你不想要一个单参数模型,即使它可以运行得非常快,但没用。你也不想要一个 100T 的模型,因为太慢了,无法进行强化学习。所以其学习效率的边际收益不值得,对吧?所以这里有一个前提:在你当前的能力类别和当前的强化学习环境等条件下,最优的模型大小是什么。是的。即使在去年,推理成本也成为了一个更大的因素,对吧?所以明确地说,模型越大,进行前向传播和生成词的成本就越高。
I think we got to drink the bitter lesson here. And yeah, like you there aren't infinite shortcuts. Like you do just have to scale and have a bigger model and pay more inference for it. And if you want AGI, then that's what you got to pay the price of. But there's a trade-off equation here, right, of there is science to do which everyone is doing of what is the optimal point at which to do RL because you need something which can both learn and discover the sparse reward itself. So you don't want a one parameter model. Useless even though you can run it really fast. You also don't want a 100T model because super slow. Possible RL. So the marginal benefit of its learning efficiency is not worth it, right? So there's a predicate here like what's the optimal model size at your current class of capabilities and your current set of RL environments and this kind of stuff. Yeah. And even in the last year there's been much more of a factor of the inference cost, right? So just explicitly like the bigger the model the more expensive it is to do a forward pass and generate tokens.
以前的计算只是:我该把算力分配给更多训练数据还是更大的模型。现在另一个巨大因素是:模型训练完成后,我实际要进行多少次前向传播。是的。我的总算力池。我该如何在训练数据算力和推理算力之间分配,用于强化学习训练?甚至在推理内部,也有很多研究:应该用什么策略?是采样 10 个然后选最好的?还是做某种分支搜索等等。所以在强化学习中,你要采样大量词元,还需要考虑模型实际生成这些词元的能力,然后学习和获得反馈。
And the calculus used to just be should I allocate my flops to more training data or a bigger model. And now another huge factor is how much am I actually going to do forward passes on this model once it's trained. Yeah. My total pool of compute. How do I allocate that across train data compute and inference compute for the RL training? And then even within inference there's all this research on well what strategy should I use? Should I sample 10 and take the best? Do I do this sort of branching search, etc., etc. And so with RL where you're sampling a whole lot of tokens, you also need to factor in the ability for the model to actually generate those tokens and then learn and get feedback.
好的。那么如果我们生活在这个世界里,你对刚起步的人或大学生有什么建议?他们应该怎么规划?
Okay. So if we're living in this world, what is your advice to somebody early in their career or a student in college? How should they be planning on doing?
是的。所以我认为,还是值得考虑各种可能的世界,并为此做好准备。在这种情况下,我认为期望值最高的行动是:你即将获得巨大的杠杆。你已经看到了,YC 的初创公司用 Claude 写了大量代码。那么,有了这种额外的杠杆,你想改变世界上的哪些挑战、哪些事业?比如,如果你有 10 个工程师随时听命,你会做什么?或者如果你有一家公司随时听命,那会让你能做什么?哪些问题和领域突然变得可行了?这就是你现在要准备的世界。这当然仍然需要很深的技术功底。还有一种情况是 AI 在一切方面都变得比所有人都好得多,对吧?但至少在一段时间内,可能还是有优势的。我记得黄仁勋在一次采访中以一种有趣的方式谈到了这一点,他说:我身边有 10 万个通用智能,我仍然有点用处,因为我在那里指导价值观,让它们做事。即使我有 10 万个通用智能,我仍然有价值。对很多人来说,我认为这种情况还会持续相当一段时间。然后,随着 AI 越来越好,最终就不行了。但再次强调,要为各种可能的世界做准备,因为如果我们完全被淘汰了,那你做什么都没用。但在所有其他世界里,这很重要。获得技术深度,学习生物学,学习计算机科学,认真思考学习物理学,思考你想解决世界上的哪些挑战。
Yeah. So I think once again there's like it's worth considering the spectrum of possible worlds and preparing yourself for that. And the one action that I think is highest EV in that case is you are about to get dramatically more leverage. You already have, like the startups in YC, you know, writing huge amounts of their code with Claude. So what challenges, what causes do you want to change in the world with that added leverage? Like if you had 10 engineers at your beck and call, what would you do? Or if you had a company at your beck and call, what would that enable you to do? And what problems and domains suddenly become tractable? That's the world you want to prepare for now. That still requires a lot of technical depth, obviously. There is the case where AI just becomes dramatically better than everyone at everything, right? But for at least a while, probably there is advantage. I think Jensen actually talked about this in an interview in an interesting way where he's like, you know, I have 100,000 general intelligences around me and I'm still somewhat useful, because I'm there directing the values and asking them to do things. And you know, they're still there, I still have value even though I have 100,000 general intelligences. And for many people, I think that will still be true for a fair while. And then, you know, as the AI get better and better and better, eventually no. But again, prepare for the spectrum of possible worlds, because in the event where we're just totally outcompeted, yeah, doesn't matter what you do. But in all the other worlds, it matters a lot. Get the technical depth, study biology, study CS, really think hard about study physics, think about what challenges you want to solve in the world.
是的,话题很多。你现在可以,对吧?学习容易多了。每个人现在都有无限的完美导师。
Yeah, that's a lot of topics. You can now, right? Like it's so much easier to learn. Everyone now has the infinite perfect tutor.
是的。我会说,要摆脱之前工作流程或专业知识的沉没成本,以便评估 AI 能为你做什么。没错。另一种有趣的表述是:变得更懒,想办法让智能体去做那些繁琐的事情。但最终你会变得更懒,不过短期内,你需要批判性地思考你目前在做的事情,以及 AI 实际上能更好地做什么,然后去尝试或探索。因为我认为还有很多唾手可得的机会,人们只是假设,没有写完整的提示词,没有给几个例子,没有连接正确的工具来加速或自动化你的工作。
Yeah. I would say some combination of like get rid of the sunk cost of your previous workflows or expertise in order to evaluate what AI can do for you. That's right. And another way to put this, which is fun, is just like be lazier in so much as figure out the way that the agent can do the things that are toilsome. But you're going to have to, in this, you ultimately get to be lazier, but in the short run, you need to critically think about the things you're currently doing and like what an AI could actually be better at doing and then go and try it or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated automated.
是的。还有一种沉没成本,就是觉得自己不是所谓的 AI 早期入局者,已经错过了机会,不能……但我想,我记得 GPT3 出来的时候。播客的背景故事:我大学毕业后,计划做某种 AI 说唱歌手创业,播客只是进入那个领域的入口。所以我尝试了不同的东西。当时我记得想,哦,3.5 出来了,人们觉得我在创业圈落后了,如果我想做自己的说唱歌手。也许说唱歌手的想法一开始就不明智,但每次都觉得还早,因为这是一个指数增长的过程。很多事情、很多想法现在才变得可能,对吧?所以正是我之前说的产品指数级变化,产品实际上被淘汰了,你需要不断重塑自己,才能保持在能力的前沿。
Yeah. There's also the sunk cost of feeling like since you're not quote unquote early to AI that you've sort of missed the boat and you can't. But I think I mean I remember when GPT3 came out. Backstory on the podcast: when I graduated college I was planning on doing some sort of AI rapper startup, and the podcast was just like a gateway into doing that. And so I was trying out different things. At the time I remember thinking, oh 3.5 is out and people like I'm so behind on the startup scene here or whatever if I wanted to make my own rapper. I mean maybe the idea of the rapper was inadvisable in the first place, but just like every time feels early because it's sort of an exponentially growing process. And there were many things, many ideas are only becoming possible now, right? So exactly that product exponential I talked about before, like products literally obsoleted, you need to constantly reinvent yourself to stay at the frontier of capabilities.
顺便问一下,你记得我有个很烂的想法,然后给你打了个电话。我不知道是什么。好像是给律师用的 RAG 之类的。总之,我想我们最早的互动之一就是我问你:‘嘿,你觉得这个想法怎么样?’你说:‘我觉得播客听起来很有前途。’我很感激。
By the way, do you remember I had a really shitty idea and I gave you a call. I don't know what it was. It was like, I think it was like RAG for lawyers or something. Yeah. Anyways, I give you I think one of our first interactions was I'm like, 'Hey, what do you think of this idea?' And you're like, 'I think the podcast sounds promising.' Which I appreciate.
是的。我最近对一个朋友有点恼火,他很有才华、很聪明,对 AI 感兴趣,但走了生物学路线,我试图让他明白:如果你想做 AI,你可以做。我认为人类不是人工的,而是生物通用智能,很多有价值的东西都是非常通用的。无论你做过什么专业化,可能都没那么重要。我是说,这需要成本,但很多我 Anthropic 的同事都对 AI 充满热情,他们只是不让之前的职业成为障碍。因为他们天生聪明、有才华、有动力,他们最终非常成功,找到了角色。他们并不是一直在 AI 领域。我是说,人们来自完全不同的领域。所以不要认为你需要某个抽象实体的许可才能参与、应用和做出贡献。
Yeah. I got slightly annoyed at a friend recently who I think is really talented and clever and interested in AI but has pursued a biology route and I just kind of tried to shake them of like you can work on AI if you want to. I mean I think humans are not artificial, are biological general intelligences where a lot of the things of value are just very general. And whatever kind of specialization that you've done maybe just doesn't matter that much. I mean again it cost but like so many of the people even my colleagues at Anthropic are excited about AI and they just don't let their previous career be a blocker. And because they're just innately smart, talented, driven, whatever else, they end up being very successful and finding roles. It's not as if they were in AI forever. I mean, people have come from totally different fields. And so don't think that you need like permission from some abstract entity to get involved and apply and be able to contribute.
如果有人现在想成为 AI 研究员,如果你能给他们一个开放问题,很可能令人印象深刻,那会是什么?
If somebody wanted to be an AI researcher right now, if you could give them an open problem that is very likely to be quite impressive, what would it be?
呃,我认为现在强化学习回归了,基于 Andy Jones 的棋盘游戏缩放定律的论文很有趣。比如展示你可以研究像你之前问的那些问题:模型是否真的学会了比之前 K 次尝试更多的东西,还是只是发现了它?深入探索这类问题我觉得很有意思。
Uh, I think that now that RL's come back, papers building on Andy Jones's scaling laws for board games are interesting. Like showing that you can investigate these questions like the ones you asked before where you're like, oh, is the model actually learning to do more than its previous pass at K or is it just like discovering that? Exploring questions like that deeply I think are interesting.
基本上就像强化学习的缩放定律。我很好奇从新任务中获得的元学习边际增长有多少。关于这一点,我认为模型差异分析有很多机会。
Like scaling laws for RL basically. Very curious to see how much the marginal increase in metalearning from a new task or something. I mean on that note, I think model diffing has a bunch of opportunities.
人们也说我们没有捕捉到所有特征,还有很多东西被遗漏了。那些被遗漏的东西是什么?如果模型被越狱了,它是在使用你已经识别出的现有特征吗?还是只使用了你没有捕捉到的误差项?我不知道。这里有很多问题。
People also say we're not capturing all the features, there's all this stuff left on the table. What is that stuff that's left on the table? If the model's jailbroken, is it using existing features that you've identified? Is it only using the error terms that you haven't captured? I don't know. There's a lot here.
我认为 Matts 很棒。Anthropic 的研究员项目进展得很好。Goodfire,Anthropic 最近投资了。他们做了很多可解释性工作,或者只是为了提高股权而什么都做。有太多可解释性项目了,有这么多唾手可得的成果,我们需要更多人,而且我觉得时间不多了。
I think Matts is great. The Anthropic fellowship has been going really well. Goodfire, Anthropic invested in recently. They're doing a lot of interpretability work or just apply anything to anything to get your equity up. There's just so many interpretability projects that are like there's so much low-hanging fruit and we need more people and I don't think we have much time.
我还想为性能工程做个宣传。我认为这是展示你原始能力的最佳方式之一。如果你在 TPU、Trillium 或 CUDA 上实现了一个极其高效的 Transformer,那么我认为你获得工作机会的可能性很高。能够完全端到端掌控模型性能的人相对较少。如果你有广泛而深厚的电气工程技能,我认为你很可能很快就能掌握加速器相关的东西。你可以相当快地跟上速度。而且它会让你对模型内部的实际复杂性有很好的直觉,这意味着你非常适合思考架构之类的事情。目前在 Anthropic,我最喜欢的架构思考者之一,实际上来自深厚的 GPU 内核编程背景,他对细节了如指掌,并且能很好地权衡利弊。
I also want to make a plug for performance engineering. I think this is one of the best ways to demonstrate that you have the raw ability to do it. If you made an extremely efficient transformer implementation on TPU or Trillium or in CUDA, then I think there's a pretty high likelihood that you'll get a job offer. There's a relatively small pool of people that you can trust to completely own end to end the performance of a model. And if you have broad deep electrical engineering skills, I think you can probably come up to speed pretty fast on accelerator stuff. You can come up to speed reasonably fast. And it teaches you a lot of good intuitions of the actual intricacies of what's going on in the models, which means that you're then very well placed to think about architecture and this kind of stuff. One of my favorite people in thinking about architecture at Anthropic at the moment actually came from a heavy GPU kernel programming background, just knows the ins and outs really deeply and can think about the trade-offs really well.
这很有趣,各位。太棒了。谢谢。是的,很高兴回来。希望你喜欢这一集。如果喜欢,最有帮助的事情就是把它分享给你认为可能喜欢的人。发给你的朋友、群聊、Twitter 或其他地方。让消息传开。除此之外,如果你能在 YouTube 上订阅并在 Apple Podcasts 和 Spotify 上留下五星好评,那会非常有帮助。查看下方描述中的赞助商。如果你想赞助未来的节目,请访问 dwarkesh.com/advertise。感谢收听。下期再见。
This is fun, guys. Awesome. Thanks. Yeah, great to be back. I hope you enjoyed this episode. If you did, the most helpful thing you can do is just share it with other people who you think might enjoy it. Send it to your friends, your group chats, Twitter, wherever else. Just let the word go forth. Other than that, super helpful if you can subscribe on YouTube and leave a five-star review on Apple Podcasts and Spotify. Check out the sponsors in the description below. If you want to sponsor a future episode, go to dwarkesh.com/advertise. Thank you for tuning in. I'll see you on the next one.