AI 算力超级周期与 Claude Sonnet 4.5 的诞生

The AI Compute Super Cycle and the Making of Claude Sonnet 4.5

肖尔托·道格拉斯 Sholto Douglas · Matt Turck 的 MAD 播客 · 2025-10-02 · 约 70 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 研究员 Sholto Douglas 探讨 AI 加速进步、算力超级周期,以及 Sonnet 4.5 如何通过强化学习成为最佳编程模型。

Anthropic researcher Sholto Douglas discusses the accelerating pace of AI progress, the compute super cycle, and how Sonnet 4.5 became the best coding model through reinforcement learning.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 30)

全文 · Full transcript(中英对照)

引言与发布节奏 Introduction and pace of releases

Host

嗨,我是 FirstMark 的 Matt Turck。欢迎收听本期特别播客,本周我们庆祝 Claude Sonnet 4.5 发布,嘉宾是 Anthropic 的杰出 AI 研究员 Sholto Douglas。在这次对话中,我们将揭秘 Sonnet 4.5 如何成为全球最佳编程模型,以及让 AI 智能体连续工作 30 小时会发生什么。发布期间,我们深入探讨了前沿 AI、大型 AI 实验室的运作方式,以及我们如何稳步迈向 AGI。应我的要求,Sholto 用通俗易懂的语言解释了强化学习、计算机使用和 AI 基准测试等关键概念,没有使用任何术语。请享受与 Sholto 的精彩对话。Sholto,欢迎。

Hi, I'm Matt Turck from FirstMark. Welcome to a special episode of the Matt Podcast for the release of Claude Sonnet 4.5 this week with the incredible Sholto Douglas, a leading AI researcher at Anthropic. In this conversation, we go behind the scenes of how Sonnet 4.5 became the best coding model in the world, and what happens when you enable AI agents to work for 30 hours straight. During the launch, we talked a bunch about frontier AI, how big AI labs operate, and how we are well on our way to AGI. At my request, Sholto made this conversation very approachable by breaking down a lot of key concepts such as reinforcement learning, computer use, and AI benchmarks in plain English without the jargon. Please enjoy this great chat with Sholto. Sholto, welcome.

Sholto

你好,很高兴来到这里。

How you doing? Great to be here.

Host

恭喜 Sonnet 4.5 发布,这是本周的大新闻。我在准备时回顾了一下,对 Anthropic 的发布速度感到震惊。特别是 Sonnet 3.7,当时可是个大事件。如果你问我,我会说‘哦,那是去年的事了’,但实际上它只是今年二月发布的。你怎么看这种发布节奏?这是否意味着进步在加速?

Congratulations on the release of Sonnet 4.5, which is the big news of this week. I was just looking back as I was prepping for this, and I was struck by the pace of releases at Anthropic. In particular, Sonnet 3.7, which was like this huge deal at the time. In my mind, if you had asked me, I would have said, 'Oh, no, that was last year.' But in fact, it was just in February of this year. What do you think about that pace of releases? Is that a proxy for progress accelerating?

Sholto

是的,我认为这反映了几件事。一是现在存在两种范式:以前是预训练 Scaling 和强化学习 Scaling,现在我们基本上将两者混合。我认为这给了你更多更新模型的机会,因为你可以在多个前沿取得进展,从而更频繁地发布。这也反映了 ChatGPT 发布两年半后的现状,后 ChatGPT 投资周期终于到来,算力可用性在增加等等。这意味着你应该期待进步的速度……因为芯片采购有前置时间。即使你去年想要芯片,也不可能拿到,因为台积电的产能已被预订一空。所以,今年算力超级周期才真正开始。

Yeah, I think it's a proxy for a couple things. One is that there's now this two-paradigm regime, where previously you did pre-training scaling and reinforcement learning scaling, and now we're in a mix of the two, basically. And so, I think that gives you more opportunities to update models because it means that you can make advancements along multiple frontiers, and then that means that you end up shipping more frequently. I think it's also a reflection of the fact that this is now two-ish years after ChatGPT 2 and a half years after ChatGPT, and so the post-ChatGPT investment cycle is finally hitting, where computer availability is increasing and all of this. And so, it means that you should expect the pace of progress to be... And because there's lead times in commissioning chips, basically. So, even if you wanted chips last year, it would have been impossible to get them because TSMC was booked out and so forth. So, finally, this year is where the compute super cycle is like beginning properly. In effect, yeah.

Host

好的。也许为了让听众了解背景,Sonnet、Opus、Haiku 这些模型有什么区别?请为我们介绍一下。

Okay, great. Maybe for situational awareness for people listening to this, the Sonnet, the Opus, the Haiku somewhere? Maybe walk us through the differences between those models.

Sholto

是的,我们按三个类别发布模型,三个层级。Opus 是最智能的模型,Sonnet 是中间层,Haiku 是最快、最便宜的模型。这次发布的一个有趣之处是 Sonnet 实际上比 Opus 更智能。这种情况以前发生过,实际上去年就发生过。这反映了快速进步,因为训练中层模型比训练大型模型更便宜。结果是你在较小模型上取得了大量进展。最终,你需要选择何时扩大规模,以获得规模带来的好处。通常,你的进步足够快,中层模型本身就已经非常出色,甚至比之前的大型模型更好。我认为这也反映了强化学习范式:你可以用强化学习来训练模型。这让你能够将中层模型提升到与六个月或三个月前的大型模型相当的水平。

Yeah, so we release models along three categories, three tiers. So, it's Opus, which is the smartest model, Sonnet, which is the sort of mid-tier model, and Haiku, which is the fastest, cheapest model. One of the interesting things about this most recent release is actually Sonnet is smarter than Opus. And this has happened before. In fact, this happened last year. It's a reflection of fast progress because it is cheaper to train mid-tier models than large models. And so, what happens is that you end up doing a lot of progress on smaller models. Eventually, you need to choose when to scale up and sort of get the benefits of scale in a model. Often, you make progress fast enough that your mid-tier model is super great anyway. And it's actually better than the large-scale model that you did previously. I think this is also a little bit of a reflection of the reinforcement learning paradigm, where you can take a model and you can train it with reinforcement learning, basically. So, that allows you to take a mid-tier model and make it as good as a larger tier model of 6 months ago or 3 months ago.

Sholto 背景与 Anthropic 之旅 Sholto's background and journey to Anthropic

Host

好的,在深入探讨之前,我对你的故事很好奇:你是如何来到 Anthropic 的,现在在 Anthropic 做什么,你如何描述你的角色?

All right, so before we go into all of this in greater detail, I was curious about your story, your journey to Anthropic, and then what you currently do at Anthropic, how you would describe your role.

Sholto

好的,你想让我从多早开始讲?从最开始吧。有几件事。在澳大利亚长大,有一些非常传统的职业路径:你可以成为律师、医生,或者进入金融业。澳大利亚在很多方面都很棒,尤其是生活质量很高,所以人们通常选择这些默认路径,过上很棒的生活。我在某些方面非常幸运。我母亲实际上在抱负上感到受挫,这意味着我一生中拥有完美的导师。她学医,后来在南非做急诊医学,但始终无法像她希望的那样进入公共卫生领域。她想要在公共卫生领域进行系统性变革,但在当时,这对女性来说非常困难。所以,她把全部注意力放在了我身上。这很棒。成长过程中,当我去中国做交换生时,她给了我这么厚的一本关于中国政治经济以及当前创业生态系统中不同参与者的资料。我接受了这种持续推动、支持且美妙的教育。我还幸运地接触了击剑。通过击剑,我体验了通过反复努力成为世界顶尖的过程。我最好的成绩是世界第 43 名。这很大程度上归功于一位世界顶尖的教练和完美的指导。他搬到澳大利亚是因为他妻子是罗马尼亚人。他刚刚带领意大利队获得了奥运会金牌。他搬到澳大利亚是因为他的妻子是罗马尼亚人,在意大利受到歧视。所以,我一方面有完美的学术指导,另一方面有完美的运动指导。还有一个试验场,让我看着 YouTube 上的这些人长大,然后自己也成为世界顶尖。

Yeah. So, I think how far back do you want me to start? From the beginning, yes. So, a couple things. One is that growing up in Australia, there's a very traditional set of paths you can take. You can become a lawyer, you can become a doctor, or you can go into finance. Australia is like a wonderful country in so many ways. In particular, the quality of life is so high that it means that people just choose these default paths, have a fantastic life, and I was very lucky in some ways. My mom was actually like she was frustrated in her ambitions. And so, this meant that I had the perfect mentor throughout my entire life. She studied medicine, went on to do emergency medicine in South Africa, but wasn't ever quite able to break into public health in the way that she wanted to. She wanted to do systemic change in public health. And at the time, that was just very difficult for a woman. So, instead, I had her full attention. This is great. Growing up, when I did exchange in China, I got this dossier this thick of China's political economy and different actors in the current startup ecosystem and this kind of stuff. So, I had this wonderful, constantly driving education in a really supportive and wonderful way. I also was lucky enough to get into fencing. And through fencing, I had the experience of becoming one of the best in the world at something via repeated effort. I became the top 50 in the world at my best, 43rd. And it was partially a consequence of, well, I think in large part due to having a coach and perfect mentorship that was one of the best in the world. He moved to Australia because his wife was Romanian. He had just coached Italy to the gold medal in Olympics. He moved to Australia because his wife was Romanian and she was facing discrimination in Italy. And so, I had on the one hand perfect academic mentorship, and on the other hand perfect athletic mentorship. And a proving ground to watch, grow up watching these people on YouTube, and then become one of the best in the world at something.

Host

早期就接触了强化学习。就像这样做,别那样做。在某种程度上,是的。

Early introduction to reinforcement learning. Like do this, don't do that. In some ways, yes.

YouTube 影响与向他人学习 YouTube's impact and learning from others

Host

或者换种方式,比如你介绍的那样,你可以通过 YouTube 观察这些人,分析他们做了什么才成为现在的样子,然后复制。你也可以成为那个世界的一部分。只需要付出巨大的努力。我觉得很神奇的是,在任何领域,YouTube 都有根本性的影响。而且无论你看哪个领域,每个孩子似乎都比上一代强很多。我不知道有没有研究证实,但至少你的经历是这样。

Or in other ways, like your introduction to you can watch these people on YouTube and analyze what they are doing to become who they have been who they are and replicate that. And you could be part of that world. All it takes is intense amounts of effort. It's a thing that I find fascinating that across any field, the fundamental impact of YouTube. And the fact that regardless of the field you look at, like every kid seems to be just much better than the prior generation. I don't know if it's studied, but at least that's your experience.

Sholto

是的,我认为 AI 领域也会出现同样的情况,对吧?每个人都会拥有一个完美的导师。我自己后来在 AI 领域也有过这样的经历。击剑并不是我想长期从事的事业。我想尝试冲击奥运会,然后转向科技领域工作。我很幸运读到了一篇 Gwern 关于 Scaling 的文章,他详细阐述了缩放假设。读完那篇文章后,我想:‘天哪,这绝对表明未来十年 AGI 的进展将成为世界上最有意义的工作之一,能够极大地推动世界进步。’于是,我开始在晚上和周末做自己的研究。

Yeah, I think we should see the same thing with AI, right? Like everyone will now get a perfect tutor. I then actually had that experience again with AI. Fencing wasn't something I wanted to do ultra long-term. I wanted to take a shot at the Olympics and then try and progress into like working in technology, basically. I was very lucky to read a Gwern essay on scaling, where he basically details the scaling hypothesis. After reading that, I was like, 'Oh my god, this is absolutely clearly AGI progress over the next decade is going to be one of the most meaningful things to work on in the world at the largest level we have to meaningfully advance the world.' And so, I started doing my own research on nights and weekends.

Host

你那时多大?

How old were you then?

Sholto

那是我本科最后一年,也就是之后的一年。本科期间我学的是计算机科学和机器人学。

So this is like last year of undergrad, like the year after. And in undergrad, you did computer science robotics.

Host

我小时候有点崇拜埃隆·马斯克这样的人。我想造火箭和特斯拉,但没有具体想解决什么问题。那篇文章是关键转折点:AGI 在这个十年是可能的,这似乎是世界上最有意义的工作,我需要想办法证明我值得从事这个领域。

And I like sort of vaguely I grew up like you know, looking up to like Elon Musk and this kind of stuff. I wanted to like build rockets and Tesla, but I didn't have a concrete idea of what actual problem I wanted to solve. Reading that essay was the critical hinge of okay, AGI is possible this decade. It seems like the most meaningful thing in the world to work on, and I need to figure out how I can demonstrate that I should be working on this.

Sholto

我当时在做机器人操作相关的工作,于是开始尝试规模化机器人操作。我在卧室里训练通用的机器人基础模型。现在这已经是个大领域了,有很多做通用机器人基础模型的公司。当时有点早,但我搭建了自己的模拟器,收集了大量遥操作数据,训练模型,还从谷歌借来了 TPU。最终,谷歌的一些人注意到了我的工作,说:‘嘿,这工作很棒,你愿意来和我们一起工作吗?’

I was at the time working on robotic manipulation stuff, and so I started working on scaling up robotic manipulation. Trying to train general foundation models for robotics from the bedroom. Which is now a big thing. There's a lot of general foundation models for robotics companies. It was a little bit early then, but I rigged up my own simulator, collected a lot of teleoperation data, trained models, got a loan of TPUs from Google. Eventually, some people at Google noticed the work I was doing and said, 'Hey, this is great work. Would you like to come and work with us?'

Host

这其实很幸运,因为比如,我没有进入我想去的博士项目。本科毕业后我申请了几个博士项目,都没被录取,但很幸运我的工作引起了谷歌的共鸣,所以他们联系了我。

It's actually very fortuitous because for example, I didn't get into the PhD programs that I wanted to. I applied to a couple of PhD programs here after undergrad and didn't get in, but I was very lucky that the work that I was doing really resonated with Google, and so they reached out.

Host

这真是个有趣的概念:你曾经可能遇到学术障碍,但仍然取得了现在的成功。对于 AI 研究圈外的人来说,感觉好像学术上最聪明的人才能赢。但这是否说明,学术优秀和成为优秀的 Anthropic 研究员是两回事,需要不同的素质?

Which is a fascinating concept that at some point you could have had an academic roadblock, but still succeed to the extent that you're currently succeeding. For people maybe outside of the AI research world, it sort of feels like whoever is the smartest academically wins. But does that suggest that being great academically and being a great Anthropic researcher are two different things, so you need slightly different qualities.

Sholto

我认为它们高度相关,但通常用于筛选学术界的信号,使得真正有能力的人远多于拥有正确信号从而能进入下一阶段的人。例如,在美国,本科生做研究就能发表 NeurIPS 或 ICLR 论文。但在澳大利亚,情况并非如此。我记得 Peter Abbeel 曾访问我们在澳大利亚的实验室,问谁要去 NeurIPS,没人举手,连博士生都没有。这意味着你缺乏非常重要的指导,没有机会培养对重要问题的品味,因此没有表明你具有高学术潜力的正确信号。

I think they're very highly correlated, but I think the signals that are usually used to gate academia are like there are dramatically more people that satisfy the criteria of being really effective than there are that have the correct signals that would then enable them to progress to the next stage in an academic career. For example, if you're here in the US, you end up doing as an undergrad research that can get you a NeurIPS or ICLR paper. Whereas in Australia, that just isn't the case, right? I remember Peter Abbeel actually once visited our lab in Australia and asked people to put their hands up if they were going to NeurIPS, and no one put their hands up, not even the PhD students. So it means you don't have that mentorship aspect that is so important, and so you don't get a chance to develop problem taste on the things that mattered, and therefore you don't have the correct signals that indicate you would have high potential for academia.

Host

实际上,我认为现在很多我们寻找的信号并不是传统的博士学位之类的。这些当然很有用,但最快的途径是,当我们看到一篇非常好的博客文章,作者独立完成了大量工作,这是信号最强的之一。我喜欢举的一个例子是 Simon Boehm,他是 Anthropic 性能团队的负责人之一,他发布了迄今为止最好的关于如何在 GPU 上优化 CUDA 矩阵乘法的指南。这简直是世界上最好的 CUDA 矩阵乘法指南。没有人对注意力机制做过同样的事。如果有人对注意力机制做了类似的事,我们第二天就会发出面试邀请。事实上,前几天有人为 TPU 和某种注意力机制做了这件事,我们立刻就说‘这个人’,马上发出了面试请求。所以我认为,实际上缺乏的是主动性、品味,但还有很多方式可以展示这些,通常是通过独立创造世界级的作品。

I actually think that right now, a lot of the signals we look for aren't your traditional PhD or anything like this. I mean these are obviously very useful, but the fastest route or the most immediate one is whenever we see a really good blog post where people have done an incredible amount of work in an independent fashion, it's one of the highest signal things there is. One of the examples I love to use here is this guy called Simon Boehm, who's one of the leads on the performance team at Anthropic, and he's published to date the best guide on how to optimize a CUDA matmul on a GPU. It is simply the world's best CUDA matmul guide. No one has done this for attention. Right? If someone did this for attention, then I mean we would reach out with a job interview offer the next day, right? And in fact, someone did it for TPU and for some kind of attention the other day, and we were like this guy. We immediately sent out a request for interview. So I think there is actually an absence of agency, an absence of taste, and there are still many ways to demonstrate this. Usually by producing a world-class artifact in some independent fashion.

Host

回到刚才关于 YouTube 的讨论,你觉得 AI 研究的人才库,无论是学术认可的还是比较独立的,是在增长吗?

And a little bit to this conversation and the YouTube discussion, do you see the pool of talent in AI research whether academically sanctioned or more indie? Is that growing?

Sholto

是的,我认为增长非常显著。而且我觉得我们在培养人才方面做得不错,Anthropic 吸收了很多初级人员,把他们培养成了非常出色的研究员和工程师。这是有意识地进行的。所以我认为肯定在增长。

Yeah, I think it's growing quite dramatically. I also think we have done a pretty good job of growing people, and I think like Anthropic has taken many, many junior people and grown them into really fantastic researchers and engineers. In a quite deliberate way. So I think it's definitely growing.

Host

所以谷歌注意到了你,然后发生了什么?

So Google noticed you, and then what happened?

Sholto

谷歌注意到了我。我大概在 ChatGPT 发布前一个月加入谷歌。那真是个绝佳的时机,因为整个公司突然被迫迅速反应,与 Gemini 项目竞争。这意味着存在一个缺口:传统的指挥结构等并不适合这场特殊的战斗。这不是一个预先存在的组织。Gemini 基本上是由 Brain 和 DeepMind 合并而成的。

So Google noticed me. I started at Google I think like a month before ChatGPT or something like this. So it was actually a fantastic time to start at Google because the entire company was suddenly forced to react instantaneously and compete with the Gemini program. So it meant that there was this gap: the typical command structures and everything were not well suited for that particular battle. It wasn't a preexisting org. Gemini was sort of forged out of the merging of Brain and DeepMind.

Gemini 早期与推理栈构建 Early days at Gemini and building inference stack

Sholto

这意味着在主动性方面存在巨大的差距,需要弄清楚我们需要做什么,以最快的速度完成,并组织人们一起处理重要的事情。因此,在 Gemini 的早期几个月里,通过与人们密切合作,我获得了大量培养品味的机会。同时,我也很快获得了机会,能够站出来领导其中的各个部分。一个例子是,我们当时没有一个适合现代 LLM 世界的推理栈。所以我们不得不注意到这一点,并从零开始设计一个。你现在在 SG Langs 等中看到的很多东西,都是我们当时从第一性原理推导出来的。我们编写了推理栈。我认为,即使在最初的六个月里,这也节省了数亿美元。这意味着我随后被信任去解决硬技术和社交政治问题。推理栈问题的一个有趣之处在于,它既是一个非常大的技术挑战,也是一个大的社交政治挑战,因为原有栈的所有权分布在五六个不同的团队中,实施变革相当困难。所以这是一个多维度的挑战。这随后让我获得了信任,去解决 ML 栈其他部分的这类问题。例如,后来当思考罢工开始时,也就是推理罢工,我负责其研究基础设施,为我们提供一个能够进行大规模强化学习和推理的 RL 代码库。

It meant that there was just a huge gap in terms of agency really, of figuring out what we needed to do, doing it as fast as possible, organizing people together to work on important things. And so I ended up getting the chance to develop a lot of taste by working closely with people in those early months of Gemini. But also quickly got the opportunity to step up and lead various parts of this. One example is we just didn't have an inference stack that was at all sensible for the modern world of LLMs. So we had to notice that, design one from scratch. A lot of the things you now see in SG Langs and stuff are things we had to derive from first principles at that point. We wrote the inference stack. This ended up saving several hundred million dollars, I think, even over the first six months. And it meant that I was then trusted to solve both hard technical and social political problems. One interesting thing about the inference stack problem was that it was both a really large technical challenge and a large social political one, because ownership of the preexisting stack was distributed across five or six different teams, and enacting change was quite hard. So it was a challenge on multiple dimensions. That then led to me having the trust to solve problems of this form across other parts of the ML stack. For example, later on when the thinking strike was started, the reasoning strike, I was in charge of research infrastructure for that, to get us an RL code base that could actually allow us to do large-scale RL and reasoning.

转投 Anthropic 与文化差异 Move to Anthropic and cultural differences

Host

很好。然后转到 Anthropic。

Great. And then the move to Anthropic.

Sholto

转到 Anthropic 是在今年二月。我认为有几个原因。首先,我对公司里每个人都如此深切地关心未来走向感到非常兴奋。这让我印象深刻:Anthropic 的每个人都有一个清晰的理论,说明他们正在做的工作如何有助于更美好的未来。无论是开发能够帮助人们改善生活的更好的人工智能,还是因为它更安全、更可控、更符合我们文明的利益,或者甚至只是更深入地理解人工智能内部实际发生的事情,并试图更好地预测进展曲线以及我们前进的方向。Anthropic 的政策在许多方面都是强有力的倡导者。这作为一个想法很迷人。

The move to Anthropic was in February of this year. I think it was motivated by a couple of reasons. Number one is I'm just really excited by how deeply every single person in the company cares about how the future goes. That's one thing that really struck me: everyone at Anthropic has an articulated theory of why what they're working on contributes towards a better future. Whether that is AI that is better in ways that can help people improve their lives, or whether it's because it's AI that's more safe and controllable and aligned with our civilization's interests, or even just more deeply understanding what's actually going on inside this AI and trying to better forecast the progress curves and where we think we're headed. Policy at Anthropic is such a strong advocate for policy in many ways. It's fascinating as a thought.

Host

从大型人工智能研究实验室的外部来看,一个问题就是,它们之间有多大不同?似乎每个人都极其聪明,每个人都拥有相同的资源。方向大致上,人们似乎专注于相同的问题,然后你看到一个模型出来了,它更好,然后一周后另一个实验室的另一个模型出来了,它比之前的更好。但根据你的经验,你在文化和目标上看到了真正的差异。

Again, seen from the outside of the big AI research labs, a little bit of a question is, how different are those? It seems that everybody is incredibly smart, everybody has access to the same resources. Directionally more or less, people seem to be focusing on the same problems, and then you see one model comes out and then it's better, and then a week later there's another model that comes out from another lab and it's better than the prior one. But from your experience, you see real differences in terms of culture and goals.

Sholto

是的,我认为有差异。例如,DeepMind,如果你想解决科学问题,它是世界上最好的地方。我认为 DeepMind 将直接贡献比任何其他机构都多的人工智能科学发现。绝对如此。它在各个方面都设置得很好。既有像 AlphaFold 和材料科学工作这样的直接科学努力,也有大规模制造人工智能科学家等努力。而我认为 Anthropic 一直专注于两件事:一是长期的人工智能对齐,二是近期的经济影响。所以 Anthropic 一直专注于编码和计算机使用,以及我们认为会在未来六个月内对经济产生直接影响的事情。与 DeepMind 和 OpenAI 相比,Anthropic 明显没有关注的一个领域是数学推理。DeepMind 和 OpenAI 一直在追求数学推理,因为它对科学和科学进步有影响,而且我认为那里很多人非常热爱数学,希望看到该领域取得进展。我们不得不遗憾地牺牲了对这方面的关注,因为我们希望专注于模型的近期经济影响。我们在其他维度上的许多研究也都集中于此。

Yeah, I think there are. For example, DeepMind, if you wanted to solve science, it's the best place in the world. I think DeepMind will directly contribute to more scientific discoveries from AI than anything else. Absolutely. It's just so well set up to do this across every aspect. You've got both the direct scientific efforts like AlphaFold and the material science work, and also generally large efforts to make AI scientists and all that. Whereas I think Anthropic has been laser-focused on two things: one is long-term AI alignment, and two is near-term economic impact. So Anthropic has been laser-focused on coding and computer use and things that we think will make a direct impact to the economy within the next six months. One thing that Anthropic noticeably hasn't focused on compared to DeepMind and OpenAI is mathematical reasoning. DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and scientific progress, and because I think so many people there just love math so deeply and would love to see the field progress. We've had to reluctantly sacrifice our focus on that because we want to focus on near-term economic impact with the models. And much of our research along other dimensions is focused on that.

AI 研究中的品味概念 The concept of taste in AI research

Host

我们稍后再深入讨论这个,但在此之前,你多次提到了“品味”这个词。这是 2025 年的一个重要词汇。在人工智能研究中,品味意味着什么?

Let's double click on this in a minute, but before we do that, you mentioned a couple of times the word taste. Which is one of those important words in 2025. What does taste mean when it comes to AI research?

Sholto

我和一位生物学朋友进行了一次非常有趣的讨论,比较了生物学研究和机器学习中的品味。我认为最重要的事情之一是机制性地理解你究竟想做什么,并有一个重要的简单性正则化器。当你考虑机器学习中的品味时,它通常是关键因素,让你在信息不完美的情况下决定将什么放入大规模训练运行中。因为我们可以深入研究架构变化的影响,但超过某一点,超过一定规模,你必须猜测该变化的影响是否会与其他变化叠加,是否会冲突,因为你无法多次测试全规模运行。你只有一次机会。所以很多品味来自于能够做出良好的推断,判断我们是否认为这最终会带来规模递增的回报。这也归结于我认为这个研究方向是否值得追求。因为通常我们在机器学习中的基线调得非常好,即使理论上更好的方法也很难击败它们,因为让机器学习方法工作需要许多小技巧,而且它们可能因各种原因失败。这不像建桥,你很清楚为什么某个剪切力减少了;它可能是所有这些怪癖。所以知道是继续沿着那个方向推进还是放弃并尝试其他东西,是另一个品味问题。

I had a really interesting discussion about this with a biology friend where we were comparing taste across biology research and ML. I think one of the most important things is mechanistically understanding exactly what you're trying to do and having an important simplicity regularizer. When you think about taste in ML, it's often the crucial ingredient that allows you to decide what goes into your large training run when you have imperfect information. Because we can study very deeply what the impact of an architectural change is, but past a point, past a certain level of scale, you have to guess whether or not the impact of that change will compound with other ones, whether it will conflict, because you can't test your full-scale run N times. You only have one shot at that. So a lot of taste comes from being able to make good inferences about whether we think this ultimately will deliver increasing returns to scale. It also comes down to whether I think this direction of research is worth pursuing. Because often our baselines in ML are so well tuned that it's very hard to beat them even with what is theoretically a better method, because there are so many small tricks required to make a machine learning method work, and they can fail for any number of reasons. It's not like building a bridge where you have a pretty good idea why a particular shear was reduced; it can be all these quirks. So knowing whether it's right to push along that direction or to give it up and try something else is another question of taste.

苦涩教训与简单性正则化 The Bitter Lesson and Simplicity Regularization

Sholto

我认为这总是归结为简单性正则化。人们喜欢耍小聪明,我们都这样。而这正是苦涩的教训,我认为它可能是对此最好的表达。一代又一代的人开发出了巧妙的方法,将他们认为人工智能应该如何推理的先验知识编码到模型中。但所有这些都被规模抹去了,通过搜索和学习这样的规划方式。规模被应用于这两件事。

And I think it always comes back to the simplicity regularization. People love to be clever, we all do. And that's sort of the bitter lesson, I think, which is maybe the best expression of this. Generations of people have developed clever methods of encoding priors about how they think an artificial intelligence should reason, and encoding it into the model. All of this gets wiped out by scale, and through planning like search and learning. Scale is applied to those two things.

Host

苦涩的教训指的是理查德·萨顿的那篇文章,AI 领域的人都知道,但可能不是所有人都听说过。这完全就是你描述的那样:泛化和算力最终会胜出。

The bitter lesson being the Richard Sutton essay, which everybody in AI knows about, but people may not have heard it. It's exactly what you described: the idea that generalization and compute will win over time.

Sholto

是的,没错。能够利用算力的方法,特别是搜索和学习,会冲刷掉所有的小修小补。我可以举几个例子让这更具体。卷积神经网络之所以更有效,一个原因是它们编码了一个先验:一个在图像上滑动的小方块,相邻像素是相关的。这是一个非常合理的先验,因为如果你给 AI 模型一张图片,而不告诉它任何关于世界的知识,它必须学习相邻像素形成曲线,然后曲线再形成其他东西。这是一个抽象层次结构。但最终,并非所有图像都符合这一点。在达到一定规模之前,卷积神经网络对绝大多数图像的表现会优于更通用的视觉 Transformer。但超过某个点后,你需要能够灵活地整合整个图像的信息。语言方面也有类似例子:我们对语法了解很多,所以你可能想把句子分解成动词、名词以及它们之间的关系,并将这种显式结构提供给 AI 算法。但当你希望模型写诗或写代码时呢?所有这些假设突然都必须抛弃。你无法在诗歌、代码和写作之间进行泛化。

Yes, exactly. Methods that can take advantage of compute, in particular search and learning, will wash away all tweaks. I can offer a couple of examples to make it more concrete. One reason convolutional neural networks were more effective is they encode a prior: a little square drawn across an image, that nearby pixels are related to each other. This is a very sensible prior, because if you throw a picture at an AI model and don't tell it anything about the world, it has to learn that nearby pixels form curves that then form other things. There's a hierarchy of abstractions. But ultimately, that is not true of all images. Convolutional neural networks will be better than a more general vision transformer for the vast majority of images up to a certain amount of scale. But past a point, you need to be able to flexibly integrate information across the entire image. A similar example in language: we know a lot about grammar, so you might want to decompose a sentence into verbs, nouns, and how they relate, and provide that explicit structure to your AI algorithm. But then what happens when you want the model to write poetry or code? All of a sudden these assumptions have to be thrown away. You can't generalize across poetry, code, and writing.

Host

那么你的意思是,至少在预测训练过程会如何发展方面,更多是凭直觉而不是实际数字?

So are you saying that, at least in terms of anticipating how the training run may go, it's more intuition than actual numbers?

Sholto

在一定范围内你可以用实际数字。说明这一点的一个好方法是,在多个规模级别上测试系统。类比生物学:你可能在细胞、小鼠和模式生物中测试一种新疗法。但这并不能保证它在人类身上有效。你在多个不同规模和模式生物上测试,它在基本单细胞细菌、小鼠、甚至猴子身上似乎都有效。这给了你很多迹象表明它会在人类身上有效,但并非保证。此时你需要理解底层机制:它结合了什么受体等等。在机器学习中完全一样。你有不同的模型规模,你发现它在这些规模上带来了好处,你认为它应该有效,因为从机制上你理解这对模型的学习动态做了什么。然后你才能有信心它会有效。但如果它是一个 hack,你并不真正理解它如何工作,它非常复杂,在代码中引入了各种东西。

You can do actual numbers up to a point. A good way to illustrate this is testing a system at multiple levels of scale. The analog to biology: you might test a new therapeutic in a cell, in mice, and in model organisms. But that's no guarantee it will work in a human. You test across multiple different scales and model organisms, and it seems to work in basic single-cell bacteria, in mice, maybe in monkeys. That gives you a lot of indication it will work in humans, but it's not a guarantee. At that point you need to understand the underlying mechanisms: what receptors is it binding to, and so forth. In ML it's exactly the same. You have your different model scales, and you figure out it's delivering benefits at these scales, and you think it should work because mechanistically you understand what this is doing to the learning dynamics of the model. Then you can have confidence it will work. But if it's a hack, you don't really understand how it works, it's really complicated and introduces all this stuff in the code.

Host

在像 Anthropic 这样的公司,一个想法失败的频率有多高?或者一般来说。

How often does an idea fail in a company like Anthropic? Or in general.

Sholto

我曾经问过 Noam Shazeer 这个问题,他说:‘嗯,大概只有 10%的想法能行。’那可是 Noam,一个绝对的天才,领域内最顶尖的人之一。所以如果连他只有 10%的想法能成功,这就界定了想法成功的比例。大多数都不行。

I once asked this question of Noam Shazeer, and he said, 'Yeah, maybe like 10% of my ideas work.' And that's Noam, an absolute genius, one of the best in the field. So if only 10% of his ideas work, that establishes a bound on the percentage of ideas that work. Most don't.

Host

像 Anthropic 或 DeepMind 这样的地方,成功的一部分原因是否就是鼓励人们一次又一次地实验?而且那些实验运行非常昂贵。大量资本涌入这些公司背后的一个重要原因就是算力昂贵。

Is part of the success of a place like Anthropic or DeepMind to just encourage people to experiment again and again? And those are very expensive runs. The big reason behind the massive amounts of capital going into those companies is that compute is expensive.

Sholto

是的。我很好奇文化上的张力:一方面因为涉及巨额资金而需要交付成果,另一方面又要保持自由开放的心态,大胆尝试。

Yes. I'm curious about the cultural tension between needing to deliver because there's so much money at stake versus having a free open mind and just going for it.

Host

嗯,这是我们在 Anthropic 和 DeepMind 都真正努力构建的东西:一种安全实验的文化,人们被信任去长时间自由探索想法。因为你通常需要数月独立研究才能真正证明一个全新的研究方向。这很难,尤其不是来自实验的算力成本,而是时间和专注的成本。在当前架构和范式中,还有那么多剩余的收益,那么多低垂的果实。你时间的一个非常高 ROI 的用法可能就是去查看数据,认真思考模型在学习什么,做一些调整。即使是最简单的事情也能带来巨大收益。所以,要求人们,或者给他们时间和空间去呼吸,并说:‘我们知道有短期的事情可以做,但我们想尝试开发一种更通用或更基础的技术,让你未来能够可扩展地做这件事’,这很重要。这就是做可扩展的事情和做不可扩展的事情之间的张力。

Well, it's one of the things we really tried to build at both Anthropic and DeepMind: a culture of safe experimentation where people were trusted to explore ideas for a long time out in the wild. Because you often need months of independent research to really prove out a novel research direction that is substantially different. It's hard, particularly less from the compute cost of experiments and more from the cost of time and focus. There are so many remaining wins in the current architectures and paradigms, so much low-hanging fruit. A really high ROI use of your time would probably be to just look at the data and think hard about what the model is learning, make some tweaks. Even the simplest thing in the world will still deliver massive gains. So asking people, or giving them time and space to breathe, and saying, 'We know there are short-term things you could be doing, but we want to try and develop a more general or fundamental technique that allows you to scalably do this in future,' is important. This is the tension between doing things that scale and doing things that don't scale.

Sholto

我认为 Anthropic 或其他地方的人们正在深入研究完全不同的方向:非 Transformer、非强化学习。我认为这是 Anthropic 和 DeepMind 略有不同的另一个方面。Anthropic 是一个高度聚焦的赌注。

I think people at Anthropic or other places are deeply researching completely different avenues: non-transformers, non-RL. I think this is another way Anthropic and DeepMind differ a little bit. Anthropic is a very focused bet.

AGI 时间线与研究理念差异 AGI timelines and research ethos differences

Sholto

我们认为 AGI 在未来几年内就能实现。我们认为当前的范式与所需范式并非天差地别。也许会有新东西,但不是什么疯狂的研究计划。过去五六年,Anthropic 的理念一直是利用当前技术集来扩展算力,相信 AGI 在此范围内是可实现的。DeepMind 拥有更广泛的科学文化,因为它有资源这么做。Anthropic 必须专注押注。DeepMind 有时间和空间去押注那些远离当前范式的东西。至于专注押注还是广泛探索新架构更好,这取决于你问什么问题,这是研究理念的差异之一。不是说 Gemini 不是专注押注,但你看 Gemini 团队有一千人,而 DeepMind 还有一万多人在做各种长期基础研究。

We think AGI is within reach in the next couple of years. We think the current paradigms are not crazy dissimilar to what we'll need. Maybe there's something new, but it's not like some crazy out-there research program. For the last five or six years, Anthropic's ethos has been scaling compute with broadly the current set of techniques, believing AGI is tractable within those bounds. DeepMind has a much broader scientific culture because it has the resources to do so. Anthropic has to be a focused bet. DeepMind has the time and space to bet on something really far outside the current paradigm. Depending on which question you ask, whether the focused bet or wide exploration of novel architectures is better, that's one of the research ethos differences. Not to say Gemini isn't a focused bet, but if you look at Gemini, it's a thousand people. There are still 10,000+ people doing all kinds of long-term foundational research at DeepMind.

Anthropic 为何聚焦编码 Why Anthropic focuses on coding

Host

为什么 Anthropic 如此专注于编码?

Why is Anthropic so focused on coding?

Sholto

我们专注于编码有两个原因。首先,我们认为这是能最快帮助我们进行 AI 研究的手段。有个概念叫自动化 AI 研究。我们认为起飞速度和进展速度最重要的信号之一就是 AI 能在多大程度上辅助 AI 研究。所以提前获取这一点非常重要。其次,我们认为从经济影响来看,这是近期最易处理的问题领域,能让 Anthropic 成为一个可行的研究项目。编码是一个巨大的市场,充满了热衷尝试和切换工具的早期采用者。对软件的需求巨大,远超优质软件的供应。模型在编码方面比其他任何领域都更早表现出色,因为编码是一个独特易处理的问题。数据以多种形式存在。你可以容器化并并行运行。你可以运行单元测试并验证它何时有效。自动驾驶就特别难,你需要汽车第一次就正常工作。而编码中,模型可以失败 100 次,只要成功一次就行。这种可处理性和可重放性在其他涉及现实世界的领域是不存在的。你不会想要一个 AI 律师为你辩护,因为如果它搞错了,那就糟了。随着技术的发展,编码是独特易处理的。你已经可以看到人们使用 AI 工具编写代码时效率大幅提升。我有一个朋友管理着九个 Claude 代码实例,这太疯狂了。我只能处理两个,可能是我自己的技能问题。

We're focused on coding for two reasons. First, we think it's the thing that will allow us to assist ourselves in AI research fastest. There's this notion of automating AI research. We think one of the most important signals of the speed of takeoff and progress is how much AI can assist AI research. So pre-fetching this is really important. Second, we think it's the nearest-term tractable problem domain in terms of economic impact for Anthropic to be a viable research program. Coding is a huge market full of keen early adopters who love trying and switching things. There's massive demand for software, dramatically more than there is good software. Models are better at coding earlier than anything else because coding is a uniquely tractable problem. The data exists in many ways. You can containerize and run things in parallel. You can run unit tests and verify when it works. Self-driving is uniquely hard; you need the car to work first time. With coding, the model can fail 100 times as long as it succeeds once. There's tractability and replayability that doesn't exist in other fields that touch the real world. You don't want an AI lawyer arguing your case because if it gets it wrong, that's bad. As techniques develop, coding is uniquely tractable. You can see that already people are dramatically more productive using AI tools to write code. I have a friend who manages nine Claude codes, which is crazy. I can only handle two, maybe a skill issue on my behalf.

Sonnet 4.5 与 SWE-bench 表现 Sonnet 4.5 and SWE-bench performance

Host

Sonnet 4.5 号称是世界上最好的编码智能体。请为我们解读一下,包括在 SWE-bench 基准测试上的表现。事实是什么,然后它是如何工作的?

Sonnet 4.5 is presented as the best coding agent in the world. Unpack that for us, including performance on the SWE-bench benchmark. What are the facts, and then how it works?

Sholto

SWE-bench 是目前衡量编码进展的基准测试,所有公司都用来相互评估。它并不完美;50% 是某个特定的 Web 框架。但它采用人们实际工作中的场景——提交拉取请求,即对代码库的更改。它检查模型能否完成同样的拉取请求并通过相同的测试。这最终能很好地代理软件工程师几个小时的工作。这些更改并非极其复杂,但具有合理的复杂度。我们最近在 SWE-bench 上的成绩从大约 72% 提升到了大约 78%,这是一个相当大的进步。值得指出的是,就在一年前,整个领域还不到 20%。进展非常显著。SWE-bench 并不完美,可能已经接近饱和。AI 基准测试在某个点之后就会失去效用,因为它们不再能区分高能力模型之间的差异。但我们的模型在 SWE-bench 上是世界最好的。我们更兴奋的是,许多客户和合作伙伴对模型非常满意。一个例子是 Cognition 团队和 Devin 发现这个模型非常有用,以至于他们不得不围绕它重建架构。他们为此写了一篇很棒的博文。我认为衡量模型好坏的真实标准是它能否让人们做到以前做不到的事情。整个编码领域在过去一年已经发生了变革。如果回到一年前,3.5 Sonnet 是第一个真正强大的智能体式编码模型,第一个你可以让它在你面前做事并与你电脑上的代码库交互的模型。这个模型在很大程度上促成了 Cursor 的产品市场契合。

SWE-bench is the current benchmark for measuring coding progress, used by all companies to evaluate against each other. It's imperfect; it's 50% one particular web framework. But it takes real-world scenarios of work people have done—submitting a pull request, a change to a code base. It checks whether the model can do that same pull request and pass the same tests. This ends up being a decent proxy for a couple hours of work from a software engineer. These changes aren't incredibly complicated but have reasonable complexity. We moved recently from roughly 72% to roughly 78% on SWE-bench, a substantial step up. It's worth pointing out that as recently as a year ago, the field was under 20%. There's been dramatic progress. SWE-bench is imperfect and probably close to saturated. AI benchmarks lose utility past a point because they no longer disambiguate differences between high-capability models. But our models are the best in the world on SWE-bench. We're more excited that many customers and partners are really excited about the model. One example is Cognition folks and Devin found the model so useful they had to rebuild their architecture around it. They had a great blog post on that. I think the real measure of whether a model is good is whether it enables people to do things they couldn't do before. Coding as a whole has been transformed over the last year. If we roll back a year to 3.5 Sonnet, the first really strong agentic coding model, the first model you could ask to do something in front of you and interact with your code base on your computer. In many ways, this model caused product-market fit for Cursor.

Cursor 与智能体编码兴起 Cursor and the rise of agentic coding

Host

Cursor 凭借 3.5 Sonnet 一飞冲天,因为他们占据了天时地利,利用那个模型提供了前所未有的编程体验。然后 Cognition 和 Windsurf 瞄准了更雄心勃勃的目标。所以这里有一个智能体能力的谱系:你可以让它做 30 秒的工作,也可以做几分钟的工作。Windsurf 这家公司的成立,部分原因就是更激进地押注 3.5 Sonnet 的智能体能力。然后进入今年。顺便提一句,这是 2025 年创业圈的关键教训之一:押注模型在 6 个月后能做什么。

Cursor took off like a rocket with 3.5 Sonnet because they were in the right place and they were able to capitalize on that model as offering a coding experience that didn't previously exist. And then Cognition and Windsurf went for an even more ambitious target. So basically there's a spectrum of agency here where either you can ask it to do 30 seconds of work or a couple of minutes of work. Windsurf was made as a company in part by betting more aggressively on the agentic abilities of 3.5 Sonnet. Then roll into this year. Which just as a quick aside is one of the key lessons for anybody in the startup world in 2025: bet on what the models will be able to do in 6 months from now.

Sholto

没错,押注指数增长。所以我认为很多编程初创公司现在都在问自己:有了能够独立追求目标、持续时间远超以往的模型,他们能做什么?以前你每 30 秒就得监督模型一次。未来几个月,你可能只需要每 10 分钟或 20 分钟监督一次。根据任务复杂度,这是一个巨大的变化。我们有几个例子,我记得博客里提到过,我们让它构建一个类似聊天应用的东西,比如 Slack 或 Teams。模型连续工作了 30 个小时,就在电脑上一直运行,最后出来一个非常好用的类 Slack 应用。这太不可思议了。现有产品中没有一个能做到这一点。也许 Cognition 一直押注的是运行时间更长、更独立的智能体套件,而此刻可能就是他们真正实现产品市场契合的时刻。

Right exactly. Bet on the exponential. So what I think something that a lot of coding startups are now asking themselves is what can they now do with models that are capable of independently pursuing goals for substantially longer than previous goals. You know before you had to supervise the models every 30 seconds. Over the next couple months you're probably going to end up in a situation where you only need to supervise the models every 10 minutes or 20 minutes or so. That's a pretty dramatic change depending on the complexity of the task. Even we have a couple of examples I think it was mentioned in the blog post where we asked it to build something that looks roughly like a chat app, something like Slack or Teams. And the model just worked for 30 hours. Like it was just spinning there on a computer for 30 hours and came out with a really good working Slack-like app. It's pretty incredible. That is nowhere near built into any of the existing products. Maybe Cognition has always bet on a longer-running, more independent agentic suite, and maybe this is the moment that really hits PMF for them for example.

智能体 30 小时在做什么 What the agent does for 30 hours

Host

我们来深入探讨一下这 30 个小时,这很有意思。首先,为了让大家理解,这是计算机使用,也就是纯编程。那么智能体在这 30 个小时里做了什么?是在点击东西吗?

Let's unpack the 30-hour aspect which is fascinating. So first of all, to ground it for people, this is a computer use, so just coding. So what does the agent do for 30 hours? Is it clicking on stuff?

Sholto

是的。它在读取文件、编写代码和运行测试。和人类的方式完全一样。基本上你可以把模型看作在一个循环中运行,它可以不断决定下一步做什么。人们经常提到所谓的工具使用,也就是使用诸如读文件、写文件等工具,或者在终端中运行代码。它就在电脑的终端里循环运行,不断查看当前代码,判断‘哦,这个还做不了,那我接下来先做那个’。它经常制定计划,特别是要运行 30 个小时的那种。我们对最近发布的产品很满意的一点是,我们终于教会了模型使用所谓的记忆。我们把它内置到了智能体框架中。所以它能够创建一个 markdown 文件,列出待办事项和它认为重要的事情,然后逐项完成、检查是否完成。这几乎形成了一个自我验证的循环。一年多前,人们对语言模型的担忧之一是它们会偏离轨道,无法自我纠正,这基本上会破坏其实用性。我认为当前这一代智能体的一个显著特点就是它们能够自我纠正。事实上,它们自我纠正的能力惊人。这种涌现能力非常有用。

Yeah. It's reading files and writing code and running tests. Exactly the same way that a human would. Basically you can think of the model as running in a loop where it can constantly decide what to do. People often mention something called tool use, and tool use is the ability to use tools like read file, write file, etc., or run code in the terminal. And it is sitting there in a terminal on a computer in a loop, just constantly looking at the current code, deciding 'oh well it can't quite do this yet, so I'm going to work on that next.' It's often making plans, particularly to run for 30 hours. One of the things that we're pretty happy with about the recent launches is we've finally taught the models to use what's called memory. And we've built that into the agentic harness. So it's able to create a markdown file of to-dos and things that it thinks are important to do, check them off, work on them, and check whether they've been completed. There's almost this self-verification loop. One of the things that people were worried about with language models over a year ago was that they would fall off track, that they wouldn't be able to self-correct, and that this would basically ruin the utility. I think one of the things that makes the current generation of agents remarkable is that they can self-correct. In fact, they're astonishingly good at self-correcting. And this emergent ability has been pretty helpful.

长期连贯性作为突破 Long-term coherence as breakthrough

Host

这里有很多值得探讨的地方。我记得你以前提到过两个维度:一个是原始智能,另一个是智能体能够运行多长时间。那么,用最简单的话说,根本性的突破是不是就是:如果你能运行更长时间,你就基本上拥有了一个非常聪明的 AI,它可以工作更长时间?

So much to unpack here. I think I heard you speak about two axes in the past: one being raw intelligence and the other being how long an agent can operate. So is the fundamental breakthrough, in very simple terms, that if you can do it longer, you basically have a very smart AI that can just work longer?

Sholto

没错。如果你能保持长期连贯性,模型就能做到以前不可能做到的事情。如果让你仅仅通过一次思维流写出一个可用的 Slack 或 Microsoft Teams 版本,你做不到,对吧?你必须坐在那里,做笔记,进行这种闭环反馈。所以长期连贯性非常重要。我们认为这对这一点至关重要。衡量这一点的一个好方法是看 METR 评估。它可能是我目前最喜欢的评估。这个评估选取了一系列人类完成的任务,特别是在机器学习或编程领域,并标注了人类达到强表现所需的时间。然后让 AI 模型去做。他们发现,进展与 AI 完成任务的时间跨度之间存在非常强的相关性。所以我认为每隔几个月,AI 能够处理的时间跨度就会翻倍,或者类似疯狂的增长。也许每 6 个月时间跨度翻倍。这简直太疯狂了。当然,和所有基准一样,这个也不完美。它只测量相当简单的任务,而且只测量 50% 的成功率,而不是 99%。但它是一个很好的方向性指标。它肯定与我的个人经验相符。当我使用最近的模型时,我开始觉得,只要我把一切设置好,我可以让它通宵运行,它不停地工作,早上起来可能就会给我一些非常有用的东西。

Yeah exactly. If you can maintain long-term coherence, then the model is able to do things that it couldn't possibly have done. If I asked you to just in a single stream of thought write a working version of Slack or Microsoft Teams, you wouldn't be able to do it, right? You have to sit there, take notes, do this closed-loop feedback system. So long-term coherence is really important. And it's something that we think is just really critical for this. I think a good way of measuring this is to look at the METR evals. They're probably my favorite eval at the moment. What this eval is is they've taken a bunch of tasks which humans do, particularly in machine learning or programming contexts, and they've annotated how long it takes a human to achieve strong performance on those tasks. And then they ask AI models to do them. What they found is that there is a really strong relationship between progress and the time horizon over which the AIs are able to complete tasks. And so I think every couple of months the time horizon that the AIs are capable of is doubling or something crazy. Maybe every 6 months the time horizon doubles. Which is utterly insane. Now again, like all benchmarks, this one is imperfect. It only measures pretty simple tasks. It only measures 50% success rate or something like this at the task, not 99% success rate. But it's a good directional measure. It certainly resonates with my own experiences. As I've been using the recent models, I start to feel that if I just set everything up right, I could leave this overnight and it could just churn away and it would probably have something pretty useful for me in the morning.

30 小时任务示例 Examples of 30-hour tasks

Host

有哪些例子能说明,用 30 个小时能完成而短时间运行无法完成的任务?

What are some examples of tasks that you can do with 30 hours that you could not do with shorter runs?

Sholto

我认为那个类 Slack 的应用就是一个很好的例子。它是一个重要的软件,真正端到端可用的软件,这通常需要一些时间,而不仅仅是一个 MVP 演示。其他我觉得有趣的事情是机器学习实验之类的。你希望它能提出实验方案、编写一些代码、运行一些初始任务,然后稍后再回来等等。这确实极大地打开了局面。关键是真正可用的软件,而不是演示。

I think the Slack-like thing is a pretty good example. It's a significant piece of software, really an end-to-end working piece of software, which often takes a bit of time, not just an MVP demo. Other things I think are interesting are machine learning experiments and stuff like that. You want something that's able to propose an experiment, write a bit of code, run some initial tasks, come back later, etc. Really it opens up the world pretty dramatically. Actually working software rather than demos is the key thing.

Host

对,太迷人了。

Right. Fascinating.

Claude AI 演示与模型演进 Claude AI demo and model progression

Host

我不是说模型现在就能给你生成一个完整的可用软件,对吧?它不会给你搞出一个 Slack 的竞品。不过你们做的那个 Claude AI 演示确实挺酷的,对吧?我觉得我们正在看到……

Now, I'm not saying that the models can spin you up a full working software right now, right? It's not going to spin you up a Slack competitor. Although the Claude AI demo that you guys produced was pretty cool, right? I think we're seeing...

Sholto

对于没看过的人来说,这个演示展示了模型的进步:从复制一个网站几乎不可能——比如只能做线框图——到现在模型能自主构建一个功能完整的网站。而且它还有一些相当复杂的功能,比如 artifacts。Artifacts 是一个功能,模型能写代码,然后代码的结果会显示在网页浏览器中。在这个例子里,模型用 artifacts 和其他功能复制了 flood.ai。我不太记得花了多长时间,大概几个小时吧。但基本上,把这看作是第一步蹒跚学步。它有时能工作,有时不能,有时又能。未来 6 个月、未来一年,期待这里会有巨大的进步。看看我们现在和一年前的差距,我预计会有同样的飞跃。

Which for people who haven't seen it, shows the progression of the models and how replicating a website went from basically impossible, like doing wireframes, to now doing a fully functional website built autonomously by the model. And it's got some pretty complex features, like artifacts. Artifacts is a feature where the model is able to write code and then the results of that code are displayed in the web browser. In this case, the model replicated flood.ai with artifacts and everything else. I can't quite remember how long that one took, maybe a couple hours. But basically, regard this as the first halting steps of this. It kind of works, sometimes it won't work, sometimes it will. Over the next 6 months, over the next year, expect dramatic progress here. Look at where we are now versus where we were a year ago, and the difference is I expect the same jump basically.

实现更长自主运行的突破 Breakthroughs enabling longer autonomous operation

Host

我们来谈谈这里的突破部分。Sonnet 4.1 我记得能运行长达 7 小时。而这次是 30 小时,我知道不是所有任务都能达到,但这是上限。你提到了记忆的进化,还有上下文、自我修正的能力。也许更详细地解释一下,是什么进步让这个跳跃到 30 小时成为可能?

Let's all click on the breakthrough part of this. So Sonnet 4.1 I think was able to run up to 7 hours. In this case it's 30 hours, which I realize is not across all tasks, but that's your upper limit. You alluded to some of this memory evolution. There's context, there's the ability to self-correct. Maybe explain in greater detail the advances that enable that jump to 30 hours.

Sholto

我认为这里最大的问题是,我们经常问自己:是什么阻止模型工作更长时间?或者说,什么时候需要干预?我很喜欢用特斯拉的干预模型作为例子。现在你需要频繁干预,但通常是在品味问题上,而不是原始编程能力。模型并不是在决定做正确的事情时做不到。但有时模型会走捷径,有时会忘记正在做的整体结构,在上下文中迷失。它们做了一个局部合理的改动,但在试图实现的全局上下文中没有意义。所以我认为很多改进,无论是我们已经做的还是将要做的,基本上都集中在品味和上下文上。就是让模型更好地对程序的整体结构做出明智的决定,不走捷径,写出合理且好的代码。

I think the biggest things here are the question we often ask ourselves: what is preventing the models from working for longer basically? Or when do you need to intervene? I quite like the model of interventions in a Tesla sense as an example. Right now you need to intervene quite frequently, but it's usually on questions of taste rather than raw programming ability. It's not like the model is unable to do the right thing when it decides to. But sometimes the models take shortcuts and sometimes they forget the overall structure of what they're doing and lose themselves in the context. They're doing a locally sensible change, but it doesn't make sense in the global context of what they're trying to achieve. So I think a lot of the improvements, both that we've made and that are still to go, are on this taste and context basically. It's on making the model better able to decide smart things about the overall structure of the program it's going to do, and not take shortcuts, and write sensible and good code.

Host

那记忆呢?

What about memory?

Sholto

记忆也非常重要,因为模型最终会耗尽上下文。能够随时间管理记忆,甚至从经验中学习,这可能会大有帮助。你不希望模型不断重新发现关于特定系统或代码库如何工作的事实。这实际上是品味问题或教训出现的领域之一。你可以想象我们发起一个大规模的努力来教代码模型编程品味。这可能是解决品味问题的一种方式:你让大量人类软件工程师决定这个好、那个坏。在软件工程中,品味从何而来?通常它让你以后容易做出修改,或者让多个智能体容易相互沟通和协作。好的抽象往往让你和我能在同一个代码库上合作而不冲突。所以问题在于,你是通过让软件工程师判断好坏来教模型编程品味,还是应该创建一个模型社会,让它们一起编写一个巨大的单体代码库?如果它们争论说某个做法不好,你可以想象各种策略的光谱,选择正确的策略是件困难的事。

Memory is also very important because the models do eventually run out of context. Being able to manage memory over time and even learn from experiences is something which will probably help this a lot. You don't want the model to be constantly rediscovering facts about how a particular system or code base works. And this is actually one of those areas where the question of taste or the bit of lesson comes up. You can imagine us going and launching a massive effort to teach the code models coding taste. That could be one way that you solve taste: you have heaps of human software engineers deciding well no, this is good or this is bad. Where does taste come from in software engineering? It's typically that it easily sets you up to make changes later on, or it's easy for multiple agents to communicate with each other and collaborate. Often good abstractions allow you and me to work together on a code base without conflicting. So there is this question of how much do you focus on teaching the model coding taste via getting software engineers to decide what is good or bad, or should you be creating a society of models that all have to together code a giant monolithic code base? If they're arguing that it's bad, you can sort of imagine the spectrum of potential strategies and picking the right one is a difficult thing.

进展速度与突破 Pace of progress and breakthroughs

Host

回到性能的跳跃,从去年的模型,甚至今年的模型,或者从 Sonnet 4.1 到 4.5,关于进步速度加快这一点。有哪些突破?

Going back to the jump in performance from last year's models or even this year's models, or even actually Sonnet 4.1 to like 4.5 again, to the point about the pace of progress accelerating. What were some of the breakthroughs?

Sholto

这个我真的不能说。我认为重要的是要认识到,这并非单一的突破。它是许多人在整个技术栈上持续应用大量不同东西的结果。而且在很多方面,它主要只是算力的函数。显然有个别突破,但根本上进步是相当平滑的。在 meter 评估上,如果你看过去两年的进步,可以用一条直线来描绘。所以类似于过去的摩尔定律,即使摩尔定律也是由许多个别改进组成的。它不是任何一个关键突破,而是在一种算力作为外生力量推动进步的环境中,大量工作的积累。

That I can't really talk about. I think it's important to recognize that it's not one individual breakthrough really. It is the continuous application of lots of different things across the entire stack for many people. And it's mostly just a function of compute in many ways. There are obviously individual breakthroughs, but fundamentally progress has been pretty smooth. On the meter eval, if you look at progress over the last 2 years, you can plot it with a straight line. So similar to Moore's law of the past, even Moore's law is made up of lots of individual improvements. It's not any one critical breakthrough. It's more the accumulation of huge amount of work in an environment where there's a sort of exogenous force of compute pushing progress forward.

从预训练转向强化学习 Shift from pre-training to RL

Host

好的,那我们也许在更抽象的层面上谈谈进步,但以 2025 年为背景。讨论的一个重点似乎是从关注预训练转向强化学习,我们之前提到过几次。谈谈强化学习的影响,以及为什么强化学习在今天如此重要?

Okay, so maybe let's talk about progress at a more abstract level, but grounded in 2025. So a big part of the discussion seems to have been the evolution from a focus on pre-training to RL, which we touched upon a couple of times. Talk about the impact of RL and why is RL such a big part of the conversation today?

Sholto

对于听众来说,一个从高层次理解预训练和强化学习区别的好方法是:预训练就像浏览所有存在的教科书。而强化学习就像做练习题,并得到关于你答对还是答错的反馈。实际上有很多东西只能通过强化学习来掌握。一个很好的例子是回答问题时说“我不知道”的技能。因为在预训练中,你是在建模文本,试图预测所有教科书和整个互联网中下一个词是什么。所以作为预训练模型,你说“我不知道”的唯一原因是你认为你在文本中建模的角色会说“我不知道”。

For those listening, a good way to understand at a high level the difference between pre-training and RL: pre-training is like skim reading every textbook in existence. And RL is like doing the worked problems and getting feedback on whether you were wrong or right. There are actually a lot of things that you can only learn via RL. A good example of this is the skill to say 'I don't know' in response to a question. Because in pre-training, you're modeling the text, trying to predict what text is going to come next in all of these textbooks and the entire internet. So the only reason you would say 'I don't know' as a pre-trained model is if you think the character that you're modeling in the text would say 'I don't know'.

RL 与下一词预测对抗幻觉 RL vs Next-Token Prediction for Hallucination

Sholto

比如,它只是看补全是否合理,对吧?不是看你实际上知不知道,而是看你认为从角色池里抽出的那个角色模型会不会说‘我不知道’。而在强化学习中,理论上你可以设置一系列测试,有些是模型知道的,有些是不知道的,然后奖励它正确回答知道的问题,惩罚它不知道却乱答。这样它就会学会在内部查找信息,评估自己对信息的置信度。所以,说‘我不知道’或解决幻觉问题,本质上在很多方面都需要强化学习。这只是其中一个例子。还有很多其他东西,不用强化学习是学不到的。

Like if it's a likely completion, right? Not whether you in fact don't know, but whether you think that the sort of cut player that you've pulled from this cast of characters that you could model would say 'I don't know.' Whereas in reinforcement learning, you could in theory set up a battery of tests where there are things the model knows and things the model doesn't know and you could reward it for correctly answering things it should know and penalize it for falsely answering when it doesn't know. And what it will then learn to do is it will learn to look up information inside itself and assess its own confidence in whether it knows that information. So saying 'I don't know' or solving hallucinations intrinsically requires reinforcement learning in many ways. So that's one example. There's a whole bunch of things you can't know otherwise.

语言模型 RL 终于奏效 RL on Language Models Finally Works

Sholto

我认为在推理模型和语言模型强化学习时代,另一个重要变化是去年年底,语言模型上的强化学习终于开始奏效了。我认为 OpenAI 功不可没,他们发布了第一个严肃的 RL+LLM 产品 O1。这确实引发了一场重大变革,因为它开辟了一个新的扩展维度,对吧?之前有预训练扩展,现在有了测试时计算和强化学习扩展。我认为所有研究实验室都在探索这个方向。DeepSeek 之所以能这么快跟进,其中一个原因是他们之前就已经发表了关于语言模型强化学习的论文。所以这个想法已经存在,但 OpenAI 将其具体化、发布并首次公开了这些缩放定律,值得称赞。

I think also an important change in this era of reasoning models and RL on language models is at the end of last year RL on language models finally started to work. And I think OpenAI deserves a lot of credit for releasing the first serious RL plus LLMs release with O1. And I think this really kicked off a substantial change because it opened up a new axis of scaling, right? There was pre-training scaling and now there's test time compute and RL scaling. I think this is something which all of the research labs were investigating already. One of the reasons that DeepSeek was able to follow so fast was that they'd actually already released papers in the direction of doing RL on language models before, for example. And so this was already an idea in the air but OpenAI deserves credit for crystallizing it, releasing it, and detailing the first public existence of those scaling laws.

测试时计算与 RL 重叠 Test-Time Compute and RL Overlap

Host

为了继续让这个讲解更易懂,测试时计算和强化学习是如何重叠的?

And maybe to continue making this super educational, how do test time compute and RL overlap?

Sholto

是的。一种理解方式是,测试时计算负责大量推理,而强化学习则提供反馈信号,判断推理是否正确。测试时计算是一种回答难题的方法。比如我问你一个你随口就能答的问题,来自你很熟悉的领域,你已经把它内化成了肌肉记忆。但对于需要真正思考和学习的任务,比如刚开始学数学时,如果我现在问你一个基本的乘法表,你可以脱口而出。但如果你是小孩,你得一步步计算,需要推理链来学习。然后你会得到对错的反馈。所以测试时计算让你能解决比当前随口能答更难的题目。而强化学习则让你把这些能力蒸馏回模型中。这就像一个梯子,你可以不断挑战稍难的问题,因为你正在学习解决越来越难问题的策略。

Yes. One way of thinking about this is test time compute is doing a lot of reasoning and then RL is the feedback signal on whether or not that reasoning was right or wrong. Test time compute is a way of answering questions that are hard for you to answer. Let's say I ask you a question that you just know off the cuff, like from a field that you really know, you've already baked that into your muscle memory. But for something which requires you to really think and learn, like when you're first doing math, if I ask you a basic times table right now you can say it off the top of your head. But if you're a kid you have to do out the math, you need to do the reasoning chain to learn it. And then you get feedback on whether it's right or wrong. So test time compute lets you do harder problems than you can currently do off the cuff. And RL then allows you to distill that back into the model. It's almost like a ladder. You can constantly do slightly harder problems because you're learning strategies to do harder and harder problems.

为何 LLM 的 RL 突破在当下 Why RL on LLMs Breakthrough Now

Host

强化学习并不是新概念。我们说的是 Richard Sutton,他在这个领域已经研究了几十年,还有其他很多人。还有 AlphaGo,那一系列非常成功、令人印象深刻的基于强化学习的成就。那么为什么在 2025 年,将这些方法应用到 LLM 上似乎有了突破?

Reinforcement learning is not a new concept. So we're talking about Richard Sutton who's been doing work in the field for decades and others as well. And then there was AlphaGo, that whole line of very successful impressive RL based successes. So why is it that in 2025 there seems to be a breakthrough to apply those to LLMs?

Sholto

是的。从某些方面来说,这很有趣。以 DeepSeek 的论文为例,他们详细描述了一种有效的方法和许多无效的方法。实际上,一些无效的方法正是 AlphaGo 成功的方法。语言模型强化学习中最疯狂的一点是,在基于验证奖励的机制下,它几乎是最简单的东西,简单到让人觉得不可能有效。这又回到了品味问题,我认为很多人觉得这太简单了,行不通。所以他们尝试了更复杂的方法,结果反而更难奏效。当然,那些方法可能仍有潜力,但首先把简单的方法做对非常重要。所以我认为人们最初尝试的强化学习策略过于雄心勃勃了。我认为 LLM 的质量也有一个最低门槛。你需要模型能够解决有一定难度的编程和数学问题,才能形成‘你做对了这些,做错了那些’的反馈循环,对吧?还有一个反直觉的点是那些推理链 token。很长时间以来,人们认为需要做一些巧妙的事情来让模型保持长期连贯性。你要记住,两年前,8000 个 token 对语言模型来说不算长上下文,但两年半前,8000 个 token 就算长了。而现在模型用 8000 甚至 30000 个 token 来推理。所以出现了一个真正的阶段转变:‘哦,语言模型的基础先验足够聪明,能解决相当难的问题。它们在比我们预期更长的上下文下实际上相当连贯。’而且这种长链推理能力可以在正确的反馈信号下自然涌现。

Yeah. In some ways it's quite funny. Let's take the DeepSeek paper for example. In the DeepSeek paper they detail one approach that works and a lot of approaches that don't work. Actually, some of the approaches that didn't work were the approaches that led to AlphaGo's success. One of the craziest things about RL on language models in the RL from verified rewards regime is it's almost the simplest thing. It's almost too simple to work. And this again comes back to that question of taste where really I think a lot of people thought this was just too simple to work. And so they tried more complex methods that ultimately ended up being harder to get to work. And there may still be juice in those methods. But it was actually really important to nail the simple thing first. And so I think people were almost too ambitious with the RL strategies that they tried initially. I think there's also a minimum bar in LLM quality that is required. Like you need the model to be able to solve meaningfully difficult coding and math problems before you can get that feedback loop of 'you solve these ones right and you solve these ones wrong', right? And I think also one of the unintuitive things is those reasoning chains of tokens. People for a long time thought that you'd need to do something clever to give the model long-term coherency. You have to remember that 2 years ago, 8,000 tokens wasn't long context for a language model, you know 2 and a half years ago 8,000 tokens was long. And now models are using 8,000 or 30,000 tokens to reason about something, right? So there was this real phase shift in 'oh language models are smart enough underlying priors that they can solve sensibly difficult questions. They're actually reasonably coherent at longer context than we thought they would be coherent.' And that this ability to reason in long chains of tokens can emerge naturally with the right feedback signal.

Host

嗯。

Mhm.

Sholto

我认为这有点反直觉。大多数人不会一开始就预料到推理能力会自然涌现。他们会认为必须结构化,必须提供推理策略,必须构建所有这些,提示它、引导它等等。但实际上,不,你给它数学题,告诉它对错,模型就会学习。这归结为规模和搜索的教训:让模型搜索,有足够的算力运行实验,模型最终会找出非常有效且合理的策略。现在正是这样,对吧?大型实验室基本上在给强化学习分配更多算力。他们正在观察最低的基础模型质量、最低的强化学习算力需求,以及对长期连贯能力的信任,做简单有效的事情。

This is a little bit counterintuitive I think. Most people wouldn't have sort of expected off the bat that ability to reason would emerge naturally. They would have thought you have to structure it, you have to provide strategies for it to do reasoning, you have to build all these things, prompt it and hint it and this kind of stuff. Actually turns out, well no, you give it math questions, tell it whether it got them right or wrong, and the model will learn. This comes down to a bit of lesson in scale and search: just allow the model to search, have enough compute to run the experiments, and the model actually ends up figuring out a really effective and sensible strategy. And that's what's happening now, right? Like the big labs are basically giving a lot more compute to RL. They're seeing what minimum base model quality, minimum amount of compute for RL, a sort of trust in the ability for long-term coherency, doing the simple thing that works.

通过 LLM 和 RL 实现 AGI AGI via LLMs and RL

Host

这些听起来都很明显,但有时其实有点反直觉。你之前提到了 AGI 这个词。那么你个人认为,更强大的 LLM 加上强化学习的组合能让我们实现 AGI 吗?

They all sound obvious but they're actually a little bit counterintuitive sometimes. Well you mentioned the word AGI earlier. So is your personal sentiment that the combination of ever more powerful LLMs plus RL gets us there?

Sholto

是的,我认为足够了。当然,这里有个明显的问题:'那里'到底指什么,以及今天 AGI 的定义是什么。有几种定义可以用。我认为一个有用的定义是:在大多数面向计算机的任务上比大多数人类更好。因为我认为这对世界来说是一个非常重要的时刻,我们意识到智力劳动可以通过这套算法来解决。这会彻底改变世界。我认为还有其他更强的定义。

Yeah, I think it's sufficient. With the obvious question of what 'there' actually means and what AGI means today. There are a few definitions one could use. I think a useful one is better than most humans at most computer-facing tasks. Because I think that's a really important moment for the world where we go okay, intellectual labor is addressable via this set of algorithms. And that totally changes the world. I think there are other definitions that are stronger that you could use.

Host

更强?那已经很强了。

Stronger? That was pretty strong.

Sholto

抱歉,我是说更难达到。是的,是的。因为你可以拥有这个,但它仍然可能无法像人类那样有效地学习,对吧?我们从很少的例子中学习和泛化。我们有极高的样本效率。而 AI 模型需要数十万倍的经验,基本上相当于数十万次生命,才能学会我们学到的东西。而且它们可以在那数千次生命中,以极高的精度学会我们的技能。我认为过去一年中一个重要变化是,强化学习终于让我们有了这样一种算法:它能够利用反馈循环,将模型变成至少在某个狭窄领域内与最优秀人类一样好的东西。你在数学和竞赛编程中看到了这一点,这是两个最适用于此的领域,模型正迅速成为非常有能力的竞赛数学家和竞赛程序员。竞赛编程和数学没有本质区别。只是它们非常适用于强化学习。任何其他领域也一样,但重要的是,它们证明了模型没有智力上限。只要有正确的反馈循环,它们就能进行非常困难的推理。所以我们认为同样的方法可以推广到人类智力活动的所有其他领域。只要有正确的反馈循环,这些模型就会变得足够好,至少在某个任务上与最优秀的人类一样好。然后,一旦你有了一个在某件事上与最优秀人类一样好的东西,你可以并行运行一千个,或者快一百倍,你就有了一个实际上比任何人类都聪明得多的东西。这完全撇开了是否可能制造出比人类更聪明的东西的问题。这似乎完全合理。大脑最终是一台生物计算机。制造一个更好的似乎是可能的。但这带来的影响是相当惊人的。也就是说,在未来两三年内,只要有正确的反馈循环、正确的算力、正确的努力等等,我们认为我们 AI 行业正走在创造某种东西的道路上,它在大多数面向计算机的任务上至少与大多数人类一样有能力。有些事情它做不到,等等。但世界将会改变。

Sorry, I mean harder to meet. Yes, yes. Because you could have this and it could still not learn as effectively as humans, right? We learn and generalize from very few examples. We have incredibly high sample efficiency. Whereas AI models need hundreds of thousands of times more experience, hundreds of thousands of lifetimes basically, to learn the things that we learn. And they can, over those thousands of lifetimes, learn the skills that we do to an incredibly high degree of accuracy. I think one of the important changes over the last year has been that RL has finally meant that we have this algorithm that allows us to take a feedback loop and turn it into a model that is at least as good as the best humans at a given thing in a narrow domain. And you're seeing that with mathematics and competition code, which are the two domains most amenable to this, where rapidly the models are becoming incredibly competent competition mathematicians and competition coders. There's nothing intrinsically different about competition code and math. It's just that they're really amenable to RL. And any other domain, but importantly they demonstrate there's no intellectual ceiling on the models. They're capable of doing really tough reasoning given the right feedback loop. So we think that same approach generalizes to basically all other domains of human intellectual endeavor. Given the right feedback loop, these models will get good enough that they are at least as good as the best humans at a given thing. And then once you have something that is at least as good as the best humans at a thing, you can just run a thousand of them in parallel or 100 times faster, and you have something that is actually substantially smarter than any given human. And this is completely throwing aside whether or not it's possible to make something that is smarter than a human. It seems entirely plausible. The brain is ultimately a biological computer. It seems possible to make a better one. But the implications of this are pretty staggering. Which is that in the next two or three years, given the right feedback loops, given the right compute, given the right elbow grease and this kind of stuff, we think that we as the AI industry are on track to create something that is at least as capable as most humans on most computer-facing tasks. Things it can't do, and this kind of stuff. But the world will change.

反论与辩论 Counter-thesis and debate

Host

你怎么看 Rich Sutton 或 Yann LeCun 的反命题,他们似乎在说需要不同的方法,或者只有强化学习?你怎么看这场辩论?

What do you make of the counter-thesis from Rich Sutton or Yann LeCun that seem to be saying that a different approach is needed, or RL only? What do you make of that debate?

Sholto

是的。我认为我们的模型学习效率远不及人类,这是事实。它们需要一千次生命来学习。但这没问题,因为它们可以在模拟中或在千家公司的岗位上度过那一千次生命,等等。我想我可能会把这两个论点分开。一个是架构上 Transformer 不够用。我不认为这是真的。我认为我们还没有真正发现 Transformer 在足够数据和足够算力下无法建模的东西。我认为强化学习作为一个目标是非常强大的。Rich Sutton 实际上是强化学习作为目标的忠实粉丝。他只是认为我们在预训练中编码了太多先验知识,等等。这不是对世界的充分表示。

Yeah. I think it's true that our models don't learn anywhere near as efficiently as humans do. They take a thousand lifetimes to learn. But this is fine because they can live those thousand lifetimes whether in simulations or doing a job at a thousand firms, and so on. I think maybe I would disentangle these two arguments. One is architecturally the transformers are insufficient. I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute. I think RL as an objective is a pretty powerful one. Rich Sutton is actually a big fan of RL as an objective. He just thinks we're encoding too many priors with pre-training and this kind of thing. It's not an adequate representation of the world.

Host

是的,这不是对世界的充分表示。

Yeah, it's not an adequate representation of the world.

Sholto

我认为到目前为止,证据表明我们当前的方法还没有发现一个通过足够努力无法解决的问题领域。能让我食言的事情是,如果有一个领域我们投入了大量精力却没有进展。比如目标、基准完全没有变化,我们一年都毫无进展。那我就会说,好吧,这里有一些根本性的限制。但相反,我不断看到的是,每当我们制定一个衡量我们关心的事物的基准时,进展都非常迅速。我认为这值得大声疾呼:伙计们,任何我们能衡量的东西似乎都在快速改善。这在两三年内会带我们去哪里?我不能肯定。但我认为值得纳入各自的世界观:有相当大的可能性我们会得到某种 AGI。

I think so far the evidence indicates that our current methods haven't yet found a problem domain that is not tractable with sufficient effort. Things that would make me eat my words is if there was some domain that we put a lot of effort into that just didn't move. Like the goalposts, the benchmarks just didn't move and we couldn't make any progress for a year. Then I would be like, okay, yeah, there's some fundamental limitation here. But instead what I constantly see is every time we make a benchmark that measures something we care about, progress is incredibly rapid along that. I think this is worth crying from the rooftops a little bit: guys, anything that we can measure seems to be improving really rapidly. Where does that get us in two or three years? I can't say for certain. But I think it's worth building into respective world views that there's a pretty serious chance that we get something that is AGI.

平台期还是指数级进步 Plateau or exponential progress

Host

所以你认为人们没有意识到?这总是很有趣,因为过去三四个月在网上读到的内容,有一种主题是我们已经达到了平台期。但你基本上在说相反的话,我们正处于指数曲线上,很多人没有意识到这一点。

So you think people don't realize? It's always interesting, because reading stuff online in the last three or four months, there's this theme that we've reached a plateau. But you're basically saying the opposite, that we are in an exponential curve and many people don't realize that it's the case.

Sholto

完全正确。我的意思是,过去三年里每个月都有人说我们正在达到平台期。但如果你看看过去三年我们取得的进展,那是不可思议的。我认为另一件让我觉得我们远未接近平台期的事情是,我观察这些模型是如何生产的。每个部分都可以改进很多。这是一个原始的流水线,靠胶带、最大的努力、辛勤工作和熬夜拼凑在一起。实际上,我记得几个月前我和几个朋友去航海。那艘船设计得非常好。

Exactly. And I mean people have said that we're hitting a plateau every month for the last three years. And if you look at where we've come over the last three years, it's incredible. I think one other thing that makes me think we're not anywhere close to a plateau is I look at how these models are produced. Every part of it could be improved so much. It is a primitive pipeline held together by duct tape and the best efforts and elbow grease and late nights. Actually, I remember I went sailing with a couple of friends a few months ago. And the boat was so well designed.

累积人类设计 vs.LLM 训练流程 Accumulated human design vs. LLM training pipeline

Sholto

这显然是数千年或数百年人类设计与努力的结晶。我当时想,哇,这就是身处大量人类努力积累中的感觉。实际上,今天最好的帆船设计很难被超越。但当我看到大语言模型的训练流程时,那是两年半的最佳努力、最后一刻的拼命努力。而且它的每个部分都有巨大的成长空间。

It was just like clearly the product of millennia or centuries of accumulated human design and effort. I was like, wow, this is what it feels like to be in the accumulation of a lot of human effort. It's actually pretty hard to beat today's best sailboat designs. But when I look at an LLM training pipeline, it is two and a half years of best effort, last minute, desperate effort. And there's just so much room to grow on every part of it.

Sonnet 4.5 与 GDP 评估 Sonnet 4.5 and GDP eval

Host

首先,Sonnet 4.5 被描述为世界上最好的编程模型,但它似乎也在许多其他领域表现出色,比如经济学研究和金融。所以它已经……我真正兴奋的一件事是 OpenAI 发布的那个 GDP 评估。我不确定 Sonnet 4.5 刚刚发布,所以它不在那个评估里。但 4.1 Opus 是那里的领先模型。我认为这是一个非常有趣且很好的评估,因为它展示了跨经济所有部门的广泛任务。这是一个跨经济各个部门的评估,对吧?比如制造业,他们基本上找了一群专家来描述成功是什么样的。现在模型将能够不仅通过编程或某些有限任务来衡量,而是通过所有任务。这样描述公平吗?

So first of all, Sonnet 4.5, which is described as the best coding model in the world, also seems to be performing across a lot of different other domains like economics research and finance. So it's already... One of the things I was really excited by was that GDP eval that OpenAI released. I'm not sure Sonnet 4.5 was only just released, so it's not on that. But 4.1 Opus was the leading model there. I think that's a really interesting and good eval because it demonstrates such a breadth of tasks across all parts of the economy. It's an eval across the various sectors of the economy, right? So manufacturing and basically they took a bunch of experts to describe what success looks like. And now the models are going to be able to be measured not just across coding or some limited task, but across everything. Is that a fair way of describing it?

Sholto

我很久以来就希望有人做这件事。拿劳工统计局的数据,把所有工作分解成任务,看看模型是否能完成这些任务,并随着时间的推移衡量进展。这显然会是一个不完美的衡量标准。我们可能在 GDP 评估上达到超越人类的表现,但这在经济上不会改变什么,因为那将是所有的连接组织、上下文,实际上任务并不具有代表性。但同样,我们会找到更好的方法来衡量这些困难,并不断推动基准。我很久以来就希望有人做这件事。我很高兴他们做了。我很高兴我们的模型是通用的,普遍强大,在所有领域都名列前茅。我认为政策制定者应该真正关注这一点,并扩展它,真正投入精力去弄清楚我们是否正朝着我一直声称的方向前进。我们可以衡量这一点。

I've wanted someone to do this for a long time. Take the Bureau of Labor Statistics, take all the jobs there, break them down to tasks and see whether models are able to do that and measure progress over time. This is obviously going to be an imperfect measure. We'll probably reach better than human on the GDP eval and it won't change anything economically because it'll be all the connective tissue and the context and actually the task won't be representative. But again, we'll then find better ways to measure these difficulties and we'll keep pushing benchmarks. I wanted someone to do this for a long time. I'm really glad that they did it. I'm really glad that our models were general and generally strong and sort of straight up top across all the areas. I think policy makers should really look at this and extend it and really make an effort in investing in figuring out whether we are on track for what I've been claiming we're on track for. We can measure this.

Host

是的。我们应该这样做。那么,就最后一个主题做个总结。太棒了,一切都很令人兴奋。我们该做什么?我们如何为这个似乎即将到来的世界做准备?

Yes. And we should be. So yeah, just to close on that last theme. So awesome, all very exciting. What do we all do? How do we prepare for this world that seems to be around the corner?

建议:杠杆与机器人学 Advice: leverage and robotics

Sholto

我认为最可行的建议是,继续为一个你作为个体拥有更多杠杆的世界做规划。现在,我可以使用两个编程智能体来完成以前两倍的工作。如果编程智能体按照我所说的方式发展,一两年内,你将能够管理一个基本上全天候为你工作的团队。我认为我们应该预期,在数字领域,个体在未来几年内将获得显著更多的净杠杆。我认为许多极其重要的问题将得到解决。我们的世界在很多方面都非常不完美。人们仍然生活在不稳定的贫困中,健康和医学问题尚未解决,住房问题完全未解决。世界可以在很多方面变得好上一百万倍。我希望的是,最初模型赋予我们在数字世界中的杠杆,然后通过机器人技术,模型赋予我们在物理世界中的杠杆,从而极大地改善它。

I think the most actionable piece of advice is keep planning for a world where you as an individual have more leverage. Right now, I can use two coding agents to do twice the work that I could have done before. If coding agents progress in the way I've been saying in a year or two, you'll be able to manage a team basically that works 24/7 for you doing work. I think we should expect in the digital domain individuals to get dramatically more net leverage over the next couple of years. I think then many incredibly important problems will get tackled. Our world is so imperfect in so many ways. People still live in erratic poverty, health and medicine is unsolved, housing is completely unsolved. The world could be a million times better in so many different ways. What I hope is that initially models give us leverage over the digital world and then hopefully models give us leverage over the physical one through robotics to dramatically improve it.

Host

这正在发生吗?机器人技术似乎是另一个关键主题,但另一方面,用“手”这个词来说,人们似乎仍然在努力让手按预期方式移动。所以它的物理特性似乎是限制因素。

Is that happening? Robotics is another thing that seems to be one of the key themes, but on the other hand, to use the word hand, it seems like people are still struggling to make hands move the way. So the physics of it seems to be the limiting factor.

Sholto

有一个叫做莫拉维克悖论的东西,即我们觉得非常容易的事情,比如操作、捡起物体,对 AI 来说非常困难,而我们觉得困难的事情,比如推理数学问题,却很容易。我实际上认为莫拉维克悖论有点虚假,我认为这主要是数据可用性和强化学习信号的问题。一个值得关注的原因是机器人运动能力。看看宇树机器人的视频。现在和两年前的差异简直疯狂。这些东西极其敏捷。有一个视频,有人踢倒了一个机器人,它真的像《黑客帝国》一样重新站起来。太疯狂了。这似乎表明运动是一个非常容易的强化学习信号。现在你基本上可以说,通过基本的强化学习,运动问题已经解决了。操作则更难一些。但有几件事让我认为机器人技术会成功。首先,今年我看到机器人实验室取得了惊人的进展。他们已经能够完成相当有趣的基本物理任务。第二,存在一个大的生成验证器差距。使我们的模型难以改进的一个原因是,我们不断需要找到能在我们想要改进的方面击败模型的人。但在机器人技术中,我们正在制造非常智能的通用模型。所以你可以这样做:如果我说“嘿,把红色积木堆在蓝色积木上面”,然后我们可以问语言模型,它是否正确地堆叠了积木?如果是,给它奖励;如果不是,就不给。所以你可以利用生成验证器差距来给机器人反馈。最后,长期以来,机器人领域的人认为他们必须解决长期连贯性和规划问题。这也是语言模型使之变得更容易的事情。它们可以将任务分解为多个步骤。所以所有机器人实验室都专注于制造出色的运动策略,并且取得了惊人的进展。这主要是一个数据和反馈循环的问题。

There's this thing called Moravec's paradox, which is that things which we find really easy like manipulation, picking up objects are really hard for AI, but maybe things which we find hard like reasoning through mathematical problems are easy. I actually think Moravec's paradox is a little bit fake and I think this is mostly a question of data availability and RL signal. One reason to look at this is robotic locomotion. Look at the videos of the Unitree robots. The difference now versus two years ago is crazy. These things are incredibly agile. There's this video I think someone kicking one over and it literally does a Matrix kind of get back up thing. It's crazy. This appears locomotion is a really easy RL signal. Right now you can pretty much say locomotion is kind of solved with basic RL. Manipulation is a bit harder. But there are a few things that make me think that robotics is going to work. For starters, I've seen incredible progress from the robotics labs over this year. They've gotten to the point where they can do pretty interesting basic physical tasks. Two is the existence of a large generative verifier gap. One of the things that makes improving our models hard is we constantly need to find people who can beat the models at the things we want to improve them on. But with robotics, we're making really smart general models. So you can actually have this: if I say, 'hey, stack the red block on top of the blue block,' we can then ask the language model, did it stack the blocks appropriately? If so, give it a reward. If not, don't. So you can use the generative verifier gap to give robots feedback. Finally, for a long time in robotics people thought they would have to solve long-term coherency and planning. That's also something that language models have made easier. They can break things down into multiple steps. So all the robotics labs are focused really hard on making great motor policies and they're making incredible progress. It's mostly just a data and feedback loop question.

结束语 Closing remarks

Host

好了,朋友们。这真是太迷人了。我现在还能想到另外 40 个问题想问你,但你非常慷慨地付出了时间。非常感谢。这太棒了。真的很感激。

All right, children. It's been fascinating. I can think of another 40 questions that I would want to ask you right now, but you've been incredibly generous with your time. Thank you so much. This was terrific. Really appreciate it.

Sholto

非常愉快。

A real pleasure.

尾声 Outro

Host

非常感谢。嗨,我是 Matt Turk。感谢收听本期 MAD 播客。如果你喜欢这期节目,我们非常感激你能考虑订阅(如果还没订阅的话),或者在观看或收听本期节目的平台上留下好评或评论。这真的有助于我们发展播客并邀请到优秀的嘉宾。谢谢,下期再见。

Thank you very much. Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks and see you at the next episode.

互动版:逐字朗读 + 针对本期提问 →