道格拉斯谈 Claude 4:编程、智能体与 AI 未来

Douglas on Claude 4: Coding, Agents, and the Future of AI

肖尔托·道格拉斯 Sholto Douglas · Unsupervised Learning · 2025-05-22 · 约 58 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

道格拉斯讨论新发布的 Claude 4 模型、对软件工程的影响、编程智能体的演进,以及他对对齐研究的看法。

Douglas discusses the new Claude 4 models, their impact on software engineering, the evolution of coding agents, and his views on alignment research.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 22)

全文 · Full transcript(中英对照)

引言与对 Claude 4 的期待 Introduction and Excitement about Claude 4

Host

但道格拉斯也是 Anthropic 的 Claude 4 模型的关键人物。在他发布这些模型的那天,我和他坐下来聊了聊,非常有趣。我们谈了很多事情,包括开发者和构建者应该如何思考 Anthropic 的下一代模型。我们讨论了趋势线对 6-12 个月后以及 2-3 年后这些模型意味着什么。我们谈到了可靠智能体需要什么,以及这些模型何时会在医学和法律等领域变得更好,类似于它们在编码方面已经取得的进步。然后我们谈到了他对对齐研究的看法:我们今天在哪里,什么在起作用,还需要做什么,以及他对 AI 2027 工作的反应。这是一次与 LLM 研究领域杰出思想家的精彩对话。我想人们会非常喜欢。话不多说,有请 Sholto。

But also Douglas was a key part of Anthropic's Claude 4 models. It was really fun to sit down with him on the day these models got released. We talked about a bunch of things including how developers and builders should think about this next generation of Anthropic models. We talked about what the trend line means for where these models will be in 6-12 months and 2-3 years from now. We hit on what's required for reliable agents and when these models will get better in domains like medicine and law and kind of mirror the advances they've already made in coding. And then we hit on his views on alignment research, where we are today, what's working, what still needs to be done, and his reaction to the AI 2027 work. This was just a fascinating conversation with a brilliant mind in LLM research. I think people will really enjoy it. Without further ado, here's Sholto.

Sholto Douglas

非常感谢你来做客播客。我很享受。这是一条很酷的小路。

Well, thanks so much for coming on the podcast. Enjoy it. It's a really cool little route.

Host

是的。不,我很感激你和我们一起钻进这个小洞穴。总是很有趣。当这个播客发布时,世界将拥有 Claude 4。我相信人们会玩它,但我很好奇。你是最早玩这些模型的人之一。它们最让你兴奋的是什么?

Yeah. No, I appreciate you getting in this little cave with us. It's always fun. By the time this podcast comes out, well, the world will have Claude 4. I'm sure people will play around with it, but I'm curious. You're one of the first people to get to play around with these models. What gets you most excited about them?

Sholto Douglas

所以它们在软件工程上又上了一个台阶,这是肯定的。Opus 真的是一个令人难以置信的软件工程模型。我越来越多地遇到这样的情况:我让它在我们的大型单体仓库中做一些非常不明确的事情,它能够以相当自主和独立的方式去完成,发现信息,弄清楚问题,运行几个测试。这每次都让我惊叹不已。每次我们得到一组新模型,我们都必须重新构建一个心智模型,了解什么有效,什么无效。那么,你的心智模型现在发生了怎样的变化?比如,在编码时,你使用这些模型做什么,不做什么?

So they're another step up in software engineering, that's for sure. And Opus is like really an incredible software engineering model. More and more I have these moments where I go and ask it to do something incredibly ill-specified in our large monorepo and it's able to go and do it in a quite autonomous and independent way, going and discovering the information and figuring this out, running a couple tests. And that blows me away every time. Every time we get a new set of models, we have to like recaracterize a mental model of what works, what doesn't. How has your mental model now changed of, you know, when you're coding, what you use these models for and don't?

Sholto Douglas

我认为最大的变化是时间范围稍微扩大了。所以我认为你可以沿着两个轴来描述模型能力的改进:一个是任务的绝对智力复杂度,另一个是它们能够有意义地推理的上下文量或连续动作的数量。这些模型在第二个轴上感觉要好得多。它们真的能够采取多个动作,弄清楚需要从环境中提取什么信息,然后据此行动。所以,时间范围,以及像 Claude Code 这类工具的支持,它现在能够访问所有工具以有用的方式完成工作,而你不再需要坐在那里从聊天框复制粘贴,这也是一个相当有意义的改进。

I think the biggest one is time horizon expands a little. So I think you can characterize model capability improvements along two axes: one of those is the absolute intellectual complexity of the task, and the other one is the amount of context or the amount of successive actions that they're able to meaningfully reason over. And these models feel substantially better along the second axis. They're really able to take multiple actions and figure out what information they need to pull in from their environments and then act on those. So giving it, like the time horizon, also the support that like Claude Code and this kind of stuff, the fact that it now has access to all of the tools to be able to do this in a useful way and you aren't sitting there copy-pasting from a chat box, is a pretty meaningful improvement in that regard too.

Host

有各种各样的任务,我原本需要花一个小时甚至很多小时的工作,它就在我面前不停地做,相当于人类的时间。当这个播客首次发布时,人们将获得这些模型。你对他们应该尝试的第一件事有什么建议?

There is a wide variety of tasks where I'm looking at, you know, an hour plus or many hours of work where I would have done and it's just there churning away in front of me doing these in terms of human equivalent time. People are going to get these models when this podcast comes out for the first time. What is your advice on the first thing they should try?

Sholto Douglas

他们应该尝试的第一件事。我的意思是,老实说,试着把它们用到你的工作中。这是最重要的。坐下来,让它做你那天在代码库中要做的第一件事。然后看着它弄清楚需要提取什么信息,并找出该做什么。我想你会印象深刻的。

First thing they should try. I mean I think honestly try and plug them into your work. That's the biggest one. Sit down and ask it to do the same thing that you were about to do first off in your codebase that day. And watch as it figures out what information it needs to pull in and figures out what to do. I think you'll be pretty impressed.

对构建者与产品指数的影响 Implications for Builders and Product Exponential

Host

是的。既然你有了这些新能力,显然有很多人在这些模型之上构建。你希望这些模型为构建者带来哪些新的可能性,让他们能够构建应用程序?

Yeah. I mean now that you have these new capabilities, obviously you have tons of people that build on top of these models. What are you hoping is newly enabled for builders that take these models and build applications?

Sholto Douglas

所以我认为在某些方面存在一个产品指数级的概念,你必须不断构建在模型能力的前沿。我喜欢用 Cursor、Windsurf、Devin 这些产品来思考。如果你看 Cursor,他们对编码的愿景在很长一段时间内都远远领先于模型的能力。你知道,Cursor 直到底层模型(如 Claude 3.5 Sonnet)起飞后才达到产品市场契合,因为他们想要提供的帮助得以实现。而 Windsurf,我认为,更加智能体化,这使他们通过在产品指数上更加努力,获得了可观的市场份额。我们现在开始看到 Claude Code,还有新的 Claude GitHub 集成、OpenAI 的 Codex 以及 Google 的编码智能体。所以每个人都在做编码智能体。是工具,对吧?人们正在为更高层次的自主性和异步性而构建。所以现在模型正在蹒跚学步,朝着能够独立于你完成任务的方向发展。那些以前需要你花几个小时的任务。接下来会是什么样子?我认为有一个有趣的转变:从你每秒钟都在循环中,到每分钟都在循环中,再到每小时都在循环中,我们在过去一年中已经看到了这一点。我想知道未来是否看起来像你在管理一群模型。所以我认为那种界面会非常有趣去探索,比如当一个人管理的不是单个模型,而是多个模型做多件事并相互交互时,你能给予他们多少并行性?我认为那会非常令人兴奋。

So I think there's this concept of a product exponential in some respects where you have to be constantly building just ahead of the model's capabilities. I like to think about this in terms of like Cursor and Windsurf and Devin and these products. If you look at Cursor, they had a vision for what coding would be that was substantially ahead of where the sort of model capabilities were for a lot of the time. You know, Cursor didn't hit PMF until the underlying models like Claude 3.5 Sonnet took off such that the assistance that they wanted to give people was able to be realized. And then Windsurf went, I would say, substantially more agentic and that enabled them to get a reasonable slice of market share by really pressing harder on that product exponential. What we're starting to see now with Claude Code but also with the new Claude GitHub integration and with OpenAI's Codex and also Google's coding agent. So everyone's really doing coding agents. Is tools, right? People building for another level of autonomy and asynchronicity. And so right now the models are taking these stumbling steps towards being able to do tasks independently of you. The kind of tasks that would have taken you several hours before. What that looks like next I think is there's this interesting transferal of you are in the loop every second to you are in the loop every minute to you're in the loop every hour that we've seen over the course of last year. And I wonder if it doesn't look like you're managing a fleet of models in future. And so I think that kind of interface would be very interesting to explore, like just how much parallelism can you give someone when it's not a single model they're managing but multiple models doing multiple things and interacting with each other. I think that would be pretty exciting.

Host

是的。你见过吗?那会是什么样子?那会是什么样子?

Yeah. Have you seen it like what might that look like? What might that look like?

Sholto Douglas

哦,天哪。我的意思是,我知道 Anthropic 内部有很多人实际上在多个开发箱中运行多个 Claude Code 实例,这很酷。但我认为还没有人真正破解那种形态因素,我认为这是一个有趣的形态因素去探索,即一个人的管理带宽几乎是多少。我认为这也是一个有趣的问题,从未来的角度来看,经济学如何运作,或者这些模型的生产力回报是什么样的。因为如果我们最初需要人类来验证这些模型的输出,那么模型的经济影响在某个初始点将受到人类管理带宽的瓶颈,直到你达到可以将对模型的信任委托给模型本身来管理模型团队的程度。

Oh god. I mean I know a lot of people actually at Anthropic who have multiple Claude Code instances up in different dev boxes, which is pretty cool. But I think no one's really cracked that form factor yet, and I think that's an interesting form factor to explore of what is the almost management bandwidth of an individual. I think this is also an interesting question to explore from the future of how does economics even work or what are the sort of return on productivity of these models. Because if we initially will need humans to verify the outputs of these models, the economic impact of the models will be at some initial point bottlenecked by human management bandwidth until you get to the point where you can delegate the trust in a model to itself manage teams of models.

抽象层与组织设计 Abstraction Layers and Organizational Design

Sholto Douglas

所以这种抽象层级的持续提升,我认为将是需要理解的最重要趋势之一。基本上,根据你检查这些模型的频率,你会有一个制约因素。你有无数个模型在运行,如果你需要每 15 分钟检查一次,而不是每小时或每 5 小时检查一次,那么你能做的事情就多得多。

And so that continual step up in hierarchy of abstraction layers will be, I think, one of the more important trend lines to understand. Basically, based on the frequency with which you need to check these models, you have a gating factor. You have an infinite number of models running, and if you have to check them every 15 minutes versus every hour versus every 5 hours, you can do a lot more.

Host

是的,没错。我记得黄仁勋在谈到他对 AGI 进展的看法时提到过这一点。他说:‘实际上,我被 10 万个极其聪明的 AGI 包围着。’这给了他巨大的杠杆作用,这就是影响所在。他描述了自己如何成为管理英伟达公司的制约因素。我认为很多工作最终会朝着这个方向发展。

Yeah, exactly. I think Jensen mentioned this with respect to how he felt about the future of AGI progress. He said, 'Well, actually, I am surrounded by 100,000 incredibly intelligent AGIs.' And he's like, this gives me huge leverage over the world, and that's sort of the impact. He's describing how he himself is the gating factor in managing the company of Nvidia. I think a lot of work ends up looking closer to that direction.

Sholto Douglas

是啊,谁知道呢?我的意思是,也许整个组织设计领域最终会变得最重要。如何建立信任,结构也变得复杂。

Yeah. Who knows? I mean, maybe this whole field of org design ends up being the most important. How do you trust and structure becomes complicated.

Host

是的,我知道我们在节目前说过,你之前在麦肯锡待过一年,我觉得这可能是咨询公司的一个好用例,他们多年来一直在做这个。也许对他们来说是一个好的新产品线。

Yeah, I know we were saying before the episode that you did spend a year at McKinsey before, and I feel like maybe this is a good use case for the consulting firms, you know, who have been doing years of this. Maybe a good new product line for them.

领先模型能力 Staying Ahead of Model Capabilities

Sholto Douglas

实际上,你刚才说的关于应用公司需要比模型发展领先一步的观点让我印象深刻。Cursor 就是这样做的。模型变化太快了,以至于如果你想想 Cursor 做了什么,对比像 Cognition 这样的智能体编程公司,也许有人现在正在思考:你用来管理 100 人团队的仪表盘是什么?你认为领先多少是合适的?因为你可能觉得‘天哪,我今天完全跟不上’,但三个月后你会说‘实际上,我已经能跟上模型能力了’。你必须不断重新发明产品,使其适合几个月后前沿模型的能力。我认为这就是关键。所以你仍然与直接用户保持大量联系,产品在一定程度上有效,但它能让你利用前沿能力。我觉得这就是秘诀。当你等待模型进步时,别人会抢走你的开发者喜爱度和客户案例,他们可能会在模型能力出现时整合进去。

Actually, I was really struck by what you just said about how for the app companies, it's about being a stage ahead of where the models are going. Cursor did this. The models change so quickly that it's almost like if you think about what Cursor did versus the kind of agentic coding companies like Cognition. Maybe someone is thinking through now what is the dashboard that you use to manage your 100-person team. What is the right level ahead in your mind to go? Because it feels like you may feel like, 'Oh god, I'm really out of my skis today,' and then in three months you'll be like, 'Actually, I'm able to where the model capabilities are.' You have to constantly reinvent the product to be suitable for the frontier of model capabilities a few months ahead. I think that's the sense. So you still maintain a lot of contact with direct users, and the product works to some degree, but it allows you to take advantage of the frontier capabilities. I feel like that's the recipe. While you're kind of waiting for the models to get somewhere, someone else is taking up your developer love and your customer case, and they can probably integrate in some of the stuff as it comes.

Host

没错,正是如此。你在自定义窗口之类的东西上也看到了这一点。这些模型中有很多你们取得进展的地方,比如记忆、指令遵循、工具使用。所以我想,再次为听众重新梳理一下,我们在这三个领域处于什么位置?哪些有效,哪些无效?

Right, exactly. And you saw that also with custom wind stuff and this kind of thing. There's a lot of things in these models that you guys made progress on, like memory, instruction following, tool use. So I guess, again, recontextualizing for folks, where are we in these three areas? What works, what doesn't?

强化学习与模型能力进展 Progress in RL and Model Capabilities

Sholto Douglas

是的。好的,思考过去一年这些模型发生了什么的一个好方法是,因为强化学习终于真正在语言模型上起作用了。我认为我们能够教会这些模型的任务的智力复杂度没有直接上限。所以你看到它们在做极其复杂的数学问题、极其复杂的编程问题,但这些都是在范围有限的领域内,上下文相对有限。问题就在模型面前。像记忆和工具使用这样的东西,是试图扩展模型能够行动的上下文范围以及它拥有的能力。所以像 MCP 这样的东西让它突然打开了世界,能够与外部世界互动。记忆让它能够运行更长的上下文,比原始模型仅靠自己的上下文窗口实现更大程度的个性化。所以我认为这些努力代表了试图通过给模型解除各种束缚来破解智能体性的尝试。

Yes. Okay, so a good way to think about what's happened with these models over the last year is because RL finally really worked on top of language models. I think there's no direct ceiling to the intellectual complexity of tasks with which we've been able to teach these models. So you see them doing incredibly complex math problems, incredibly complex coding problems, but those things are in scoped domains where it's relatively limited context. The problems are there in front of the model. Things like memory and tool use are attempts to expand the set of context within which the model is able to act and the affordances it has. So things like MCP allow it to suddenly the world opens up to it and it's able to interact with the outside world. Memory allows it to run with much longer context, much greater degrees of personalization than just a raw model with its own context window. And so I think those efforts represent attempts to crack agency by giving the model all these unhobblings in one respect.

Host

而且我认为宝可梦评估是一个很好的例子。作为一个当年狂热的 Game Boy 玩家,我很喜欢这个评估。我希望你们会把这个评估和这个模型一起发布。

And I think the Pokemon Eval is a great one. I love that eval as an avid Game Boy player back in the day. I hope you're going to release that alongside this model.

Sholto Douglas

是的。新模型一直在玩宝可梦。你会看到的。我认为这是一个很好的评估,因为它没有经过专门训练。所以它展示了智能的泛化能力,能够处理一个并非完全分布外但与其之前做过的任何事都有显著不同的任务。

Yeah. The new model has been playing Pokémon. So you'll see that. And I think it's a great eval because it hasn't been trained for. And so it demonstrates this generalizability of intelligence to a task which is not completely out of distribution but one which is meaningfully different from anything it's done before.

Host

另一个例子是,我当年不得不买攻略才能通关那个游戏。我记得有很多梯子和绕路的地方。

Another example I had to buy a strategy guide to beat that game. I remember like a lot of ladders and places to go around it.

Sholto Douglas

没错。另一个我非常喜欢的例子是,Anthropic 最近一直在研究一个可解释性智能体。基本上,它的工作是找到语言模型中的电路。这非常酷,因为第一,我们没有训练它做这个。我们训练它成为一个编程智能体,但它能够将其与心智理论的知识结合起来,自己与它试图理解的模型对话,尝试推理出它有什么工具,比如查看可视化神经元和电路的工具。它实际上能够赢得一个有趣的对齐安全评估,叫做审计游戏,在这个游戏中,你以某种方式扭曲模型,它必须找出模型出了什么问题。它能够做到这一点。它能够与模型对话,生成自己的关于模型可能出了什么问题的假设,并查看所有这些工具。我认为这非常出色地展示了这些模型在拥有工具和记忆时的泛化能力。

Exactly. Another example of this that I really like is there's been a recent interpretability agent that Anthropic's been working on. And basically what this does is it does that job of finding circuits in language models. And this is really cool because one, we haven't trained for it to do this. We've trained for it to be a coding agent, but it's able to mix that with its knowledge of theory of mind to sit there and itself talk to the model that it's trying to understand, try and reason through what kind of it has access to tools which are like looking at visualizing neurons and circuits. And it is actually able to win this interesting alignment safety eval which is called the auditing game, where you twist the model in some way and it has to figure out what is wrong with the model. And it is able to do that. It's able to talk to the model, generate its own hypothesis about what might be wrong with the model, and look at all these tools. I think it's just such a brilliant demonstration of the generalizable competence of these models with access to tools and memory.

Host

完全同意。我觉得开发者们一直在等待智能体以及可靠使用这些东西的能力。我记得你之前在播客里说过,智能体的障碍是可靠性。对于听众中的开发者来说,我们在这方面取得了多少进展?

Totally. I feel like builders have been waiting for agents and the ability to use this stuff reliably. I think you've talked before on podcast that the barrier for agents is reliability. How much progress have we made there for the builders that are listeners here?

智能体可靠性与进展 Agent Reliability and Progress

Host

智能体的障碍在于可靠性,对吧?我们在这方面取得了多少进展?

The barrier for agents is reliability, right? How much progress have we made there?

Sholto Douglas

对于正在收听的开发者来说,我确实认为用时间跨度内的成功率来衡量是正确的方法,用来思考智能体能力的扩展。我认为我们取得了巨大进展。但可靠性还没达到 100%。这些模型并非总能成功。单次尝试与 256 次尝试之间的性能仍有显著差距。许多评估可以通过多次尝试完全解决,但第一次尝试并不保证成功。尽管如此,我看到的每条趋势线都表明,我们正朝着在大多数训练任务上获得专家级超人类可靠性的方向前进。

For builders listening, I really like the metric of measuring success rate over a time horizon. That's the right way to think about extending agent capabilities. I think we're making a hell of a lot of progress. We're not 100% there on reliability. These models don't succeed all the time. There's still a meaningful gap between performance on a single attempt versus 256 attempts. Many evals can be completely solved with many attempts, but on the first try it's not guaranteed. That said, every trend line I see says we're on track to get expert superhuman reliability at most things we train on.

Host

什么会让你改变看法?

What would change your mind on that?

Sholto Douglas

如果我们偏离了趋势线。例如,到明年年中,我们看到模型能够行动的时间范围出现瓶颈。编程一直是 AI 的领先指标,所以你会首先在编程上看到下滑。那可能反映了算法固有的局限性,但我坚信这种局限性不存在。还有其他限制,比如任务分布比预期更难,因为数据较少,导致过程繁琐。对于计算机使用智能体来说,这类数据并非天然存在。但我们在那里看到了如此惊人的进展,以至于我完全不认为我们处于那种世界。

If we fell off the trend line. For example, if by the middle of next year we saw a block on the time horizon these models can act over. Coding is always the leading indicator in AI, so you'd see that drop-off in coding first. That might reflect inherent algorithmic limitations, which I strongly believe don't exist. There are other limitations, like task distribution being harder than expected due to less data, making it laborious. For computer use agents, that kind of data doesn't natively exist. But we're seeing such incredible progress there that I don't think we're in that world at all.

Host

你认为我什么时候能拥有一个通用的智能体,可以帮我填写表格、浏览互联网?

When do you think I'll have a general-purpose agent that can fill out forms and navigate the internet for me?

Sholto Douglas

我开玩笑说这是个人事务逃逸速度——作为一个拖延症患者,如何推迟任务。这取决于情况。仍有显著差距,而且取决于公司是否专注于让模型进行练习。如果你从街上随便拉一个人,让他们做会计工作且不出错,他们很可能会犯错。但如果他们有相关经验,或者是优秀的数学家或律师,他们就能泛化并以更高的可能性完成。所以这很大程度上取决于任务。到明年年底,应该会很明显,这几乎是保证的。甚至到今年年底,情况应该就很清楚了。到明年年底,这些东西会在你的浏览器里为你做很多事情。

I joke about personal admin escape velocity—how to put off tasks as a procrastinator. It depends. There's still a meaningful gap, and it depends on whether a company focuses on giving the model practice reps. If you take a random person off the street and ask them to do accounting without mistakes, they'd probably make errors. But if they have relevant experience or are a great mathematician or lawyer, they can generalize and do it with higher likelihood. So it strongly depends on the task. By the end of next year, it should be very obvious that this is near guaranteed. Even by the end of this year, it should be pretty clear. By the end of next year, these things will be doing a lot of things for you in your browser.

Host

你们的模型非常擅长编程。是什么让它们如此独特?是内部优先级吗?人们把 Anthropic 和编程模型公司联系在一起。这背后是什么?

Your models are really good at coding. What makes them uniquely good? Is it internal prioritization? People associate Anthropic with being the coding model company. What's behind that?

Sholto Douglas

Anthropic 确实非常重视优先处理重要的事情。我们相信编程极其重要,因为它是 AI 研究本身加速的第一步。所以我们非常关注编程,并衡量其进展。它是模型能力最重要的领先指标。

Anthropic does care a lot about prioritizing important things. We believe coding is extremely important because it's the first step where AI research itself will be accelerated. So we care a lot about coding and measuring progress on it. It's the most important leading indicator of model capabilities.

Host

这些智能体今天是否在加速研究?

Are these agents accelerating research today?

Sholto Douglas

是的,基本上是的。它们大大加速了工程。即使是杰出的工程师也说,在他们熟悉的领域,效率提升约 1.5 倍,在不熟悉的领域(如新编程语言)则提升 5 倍。AI 加速 AI 进展的一个重要因素是我们是否受算力限制。如果你部署 AI 智能体进行研究,收益与部署的研究人员数量成正比。现阶段,这些东西大多能处理你工作中烦人的部分,让你专注于出色的研究。

Yes, basically yes. They accelerate engineering a lot. Even brilliant engineers say it's like 1.5x on domains they know well, and 5x on domains they don't know well, like new programming languages. An important factor for how much AI will accelerate AI progress is whether we are compute-bound or not. If you deploy AI agents for research, you get gains proportional to the number of researchers deployed. At this stage, most of these things can do the annoying parts of your job so you can focus on brilliant research.

Host

你如何看待智能体提出有趣研究方向的时间线?

How do you think about the timeline for agents proposing interesting research directions?

Sholto Douglas

目前大部分工作是工程工作。至于它们提出新颖想法,我不确定。未来两年内,人们已经开始看到有趣的科学提案。需要考虑的一个重要因素是,这些模型如果在某件事上有反馈循环,就能成为真正的专家。它们需要被允许练习,类似于人类。此外,任务需要相对容易验证。

Most of the work is engineering work at this point. When they propose novel ideas, I'm not sure. Within the next two years, people are already starting to see interesting scientific proposals. An important thing to consider is that these models can become truly expert at something if they've had a feedback loop for that thing. They need to have been allowed to practice, similar to humans. Also, tasks need to be relatively easily verifiable.

Host

我们会不会得到一些编程能力惊人,但在更模糊的技能上毫无进展的模型?

Are we going to get models that are unbelievable coders but haven't made progress on more nebulous skills?

机器学习可验证性与弱可验证领域进展 Verifiability in ML and progress in less verifiable domains

Sholto Douglas

有一点是,机器学习研究实际上是非常可验证的。损失下降了吗?如果你能提出有意义的机器学习研究方案,你就拥有了世界上最好的强化学习任务。在某些方面甚至比通用软件工程更甚。我们会在可验证性较差的领域取得进展吗?我非常有信心我们会。一个有趣的数据点是 OpenAI 最近关于医学问题的论文。你注意到它是如何评分的吗?他们有一个新的医学评估,包含像考试中那样的长文答案,并给出了分数。这把一个不像代码或数学那样天生可验证的领域,转化成了更可验证的东西。我认为这很可能会被解决,基本上已经解决了,而且几乎肯定最终会被解决。

One point is that ML research is actually incredibly verifiable. Did the loss go down? If you can get to the point where you can make meaningful proposals for ML research, you have the best RL task in the world. Even more so than general software engineering in some respects. Will we get progress on less verifiable domains? I'm very confident that we will. One interesting data point is OpenAI's recent paper on medical questions. Did you notice how it was scored? They had a new medical eval with long-form answers like in an exam, and they gave points for it. This takes a domain that is not inherently verifiable like code or math and converts it into something much more verifiable. I think this is reasonably likely to get solved, basically already solved, and near guaranteed to get solved eventually.

Host

最终是什么时候?我们什么时候会有一个真正好的医学或法律模型?它会在明年成为更广泛模型的一部分,还是你认为会有法律专用或医学专用的模型?

When is eventually? When will we have a really good medical or law model? Does it become part of the broader model in the next year, or do you think there will be legal-specific or medical-specific ones?

Sholto Douglas

在这方面我有点像个大型模型最大化者。大多数研究人员也是。我确实认为模型个性化有很多有趣的方式——你想要一个理解你的公司、你关心的事情和你自己的模型。所以针对你的需求微调模型确实重要,但我认为这不会是行业特定的,而是公司或个人特定的。Anthropic 与 Databricks 有合作做公司特定的事情,但在基础能力层面,我坚信单一的大型原始模型。一个原因是我们迄今为止看到的趋势。另一个是长期来看,小型和大型模型之间的区别没有理由存在。你应该能够根据给定任务的难度自适应地使用适量的算力。这偏向于大型模型。

I'm a bit of a large model maxi in this respect. Most researchers are. I do think there are interesting ways in which personalization of models matters—you want something that understands your company, the things you care about, and you yourself. So tuning models for your things does matter, but I think this won't be industry-specific so much as company or individual-specific. Anthropic has a partnership with Databricks for company-specific stuff, but at a base level, I firmly believe in a single raw large model. One reason is the trend we've seen so far. Another is that there's no reason in the long run for the distinction between small and large models to exist. You should be able to adaptively use the right amount of compute for the difficulty of a given task. That biases towards large models.

对 GDP 与白领工作的影响 Impact on GDP and white-collar work

Host

你似乎对模型的持续改进很有信心。很多人猜测模型将如何扩散到社会并影响 GDP。在未来几年,你认为这些模型会对世界 GDP 产生什么影响?

You seem convinced about continued improvement. A lot of people speculate about how models will diffuse into society and impact GDP. In the next few years, what impact on world GDP do you think these models will have?

Sholto Douglas

我认为最初的影响可能类似于中国的崛起,这可能是过去 100 年对世界影响最大的事情。看看上海在 20 年间的巨大转变。这将会比那快得多。但有一些重要的区别。我认为我们几乎可以肯定,到 2027-2028 年,或者到本十年末,我们将拥有能够自动化任何白领工作的模型。这是因为这些任务容易受到我们当前算法的影响——你可以在计算机上多次尝试,有丰富的数据。但同样的资源对于机器人或生物学并不存在。要让模型成为超人类程序员,你只需要我们已经赋予模型的能力,并扩展现有算法。要让模型成为超人类生物研究员,你需要自动化实验室,以便它能以高度可并行化的方式提出和运行实验。要让它在现实世界中像我们一样胜任,它需要通过机器人行动。所以你需要大量的机器人来收集数据。我担心的一个不匹配是,白领工作将受到巨大影响,无论这是否表现为显著的增强。那个世界将发生巨大变化,我们需要推动那些让我们的生活变得更好的事物的巨大转变——医学、现实世界的富足。我们需要解决云实验室和机器人技术。但到那时,我们将有数百万的 AI 研究人员提出方案,所以他们不需要那么大规模的机器人或生物数据。AI 进展非常快。但我们需要引入现实世界的反馈循环,才能真正改变世界 GDP。

I think the initial impact might look like the emergence of China, which has probably impacted the world most in the last 100 years. You look at Shanghai over 20 years and it dramatically transforms. This will be dramatically faster than that. But there are important distinctions. I think we are near guaranteed to have models capable of automating any white-collar job by 2027-2028, or near guaranteed by the end of the decade. That's because those tasks are susceptible to our current algorithms—you can try things on computers many times, there's a wealth of data. But that same resource doesn't exist for robotics or biology. For a model to be a superhuman coder, you just need affordances we've already given models and to scale existing algorithms. For a model to be a superhuman biological researcher, you need automated laboratories to propose and run experiments in a hugely parallelizable way. For it to be as competent in the real world as we are, it needs to act through robotics. So you need a lot of robots to collect data. One mismatch I worry about is a huge impact on white-collar work, whether that looks like dramatic augmentation or not. That world will change a lot, and we need to pull forward the dramatic transformation of things that make our lives better—medicine, abundance in the real world. We need to figure out cloud laboratories and robotics. But by that time, we'll have millions of AI researchers proposing things, so they don't need such large-scale robotics or biological data. AI progress goes really fast. But we need to pull in the feedback loops of the real world to meaningfully change world GDP.

Host

所以你认为对于每个白领职业,你可以构建某种奖励模型,类似于在医疗评估中的做法?令人惊讶的是,构建这些东西所需的数据非常有限,就像人类在相对有限的数据上学习一样。

So you think for each white-collar profession, you can build some sort of reward model similar to how it's done in healthcare evals? And what's surprising is how limited data you need to build those things, in the same way that a human learns on relatively limited data.

Sholto Douglas

完全正确。而且我们已经确凿地证明,我们可以教模型——到目前为止,我们还没有遇到在可教任务上的智力天花板。它们似乎确实比人类样本效率低一些,但这没关系,因为我们可以并行运行数千个副本,与不同变体的任务交互,拥有终身的经验。所以即使它们样本效率较低也没关系。你仍然可以在该任务上获得专家级的人类可靠性和性能。

Exactly. And we've conclusively demonstrated that we can teach models—so far we haven't hit an intellectual ceiling on the tasks we can teach them. They do seem somewhat less sample-efficient than humans, but that's okay because we can run thousands of copies in parallel, interacting with different variations of tasks, having lifetimes of experience. So it's okay if they're less sample-efficient. You still get expert human reliability and performance on that task.

Host

看起来你认为这个范式几乎能让我们达到目标。

It seems like you think this paradigm gets us pretty much all the way there.

Sholto Douglas

是的。

Yeah.

当前范式对 AGI 的充分性 Sufficiency of current paradigms for AGI

Host

你知道,显然有像伊利亚这样的人一直在说,需要某种其他的算法突破。另一方的观点是什么?

You know, obviously you have folks like Ilya who have been saying there needs to be some sort of other algorithmic breakthrough. What's the other side here?

Sholto Douglas

是的,有道理。我认为目前领域内大多数人相信,我们迄今探索的预训练加强化学习范式本身就足以达到 AGI。我们还没有看到趋势线弯曲;这种组合是有效的。是否有其他山峰可以更快登顶?完全可能。我是说,伊利亚之前发明了这两种范式,所以我怎么敢跟他打赌?我看到的所有证据都表明这些是足够的。也许伊利亚这样赌是因为他没有那么多可用资本,或者他认为这是更好的方式。完全可能。我不会跟伊利亚对赌,但我确实认为我们现有的东西能带我们到达那里。

Yeah, makes sense. I think most people in the field currently believe that the pre-training plus RL paradigms we've explored so far are themselves sufficient to reach AGI. We haven't seen the trend lines bending yet; this combination of things works. Whether there are other mountains to climb that could get us there faster is entirely possible. I mean, Ilya invented both of these paradigms before, so who am I to bet against him? Every piece of evidence I see says these are sufficient. Maybe Ilya is betting that way because he doesn't have as much capital available or he thinks this is a better way. Entirely possible. I'm not going to bet against Ilya, but I do think what we have now will get us there.

能源与算力作为限制因素 Energy and compute as limiting factors

Host

这方面的限制因素将是能源和算力。你觉得我们什么时候开始碰到这个问题?

The limiting factor on this will be energy and compute. When do you think we start to bump up against that?

Sholto Douglas

我认为《情境意识》末尾有一个很好的表格详细说明了这一点。到本十年末,我们开始报告美国能源生产的非常惊人的百分比。你超过 20%,我想可能是 2028 年,比如美国能源的 20%。所以你不能在没有巨大变化的情况下再增加几个数量级。我认为这是我们需要更多投资的地方。我认为这是政府应该采取行动的重要方向之一。迪伦有一张很棒的中国与美国能源生产对比图。美国能源生产是平的,而中国的能源生产是这样的——他们在能源建设方面比我们做得好得多。

I think there's a great table at the end of 'Situational Awareness' that details this. By the end of the decade, we start to report really dramatic percentages of US energy production. You're over 20%, I think maybe 2028, like 20% of US energy. So you can't go orders of magnitude more than that without dramatic changes. This is somewhere I think we need to invest more. I think this is one of the important vectors along which governments should act. Dylan has this wonderful graph of China's energy production versus US energy production. US energy production is flat, and China's energy production is like this—they're just doing a much better job than we are of building out energy.

当前浪潮中的爬坡指标 Metrics for hill climbing in current wave

Host

在当前这波模型改进中,哪些指标值得现在去攀登?当你从 4 过渡到 4 之后的任何版本时。

In this current wave of model improvement, what metrics are worth hill climbing on right now? As you move from 4 to whatever comes after 4.

Sholto Douglas

是的。总的来说,我对公司内部的评估印象深刻。许多公司都设计了自己的 SWE-bench 版本,比如说。这些评估非常严格且保持良好,所以我喜欢在这些上面攀登。我也认为像 FrontierMath 这样非常复杂的测试在未来一年里非常值得关注,因为它代表了智力复杂性的一个天花板。但越来越多地,我认为重要的是难以产生的评估。如果我们能产生有意义地捕捉人们工作日时间跨度的评估,那将是最好的产出。但还没有人公开做到这一点。这是我认为政府应该做的另一件事,因为了解趋势线是什么样的对政策制定非常重要。政府很适合做这个——他们应该产出律师或工程师一小时或一天工作的输入和输出是什么样的,并且我能否将其转化为可评分的东西,这样我们就能实际衡量进展。

Yeah. In general, I've been impressed by internal company evals. Many companies have devised their own version of SWE-bench, let's say. These are quite rigorous and well held out, so I enjoy hill climbing those. I also think really complex tests like FrontierMath are really interesting to watch over the next year, because that represents such a ceiling of intellectual complexity. But more and more, I think what matters is evals that are hard to produce. If we could produce evals that meaningfully capture the time horizons of people's workdays, I think that would be the best thing to produce. But no one has gone out there and produced that in public. This is another thing which I think governments should do, because understanding what the trend line looks like is such an important input into policy. Governments are well placed to do this—they should produce what the inputs and outputs look like of an hour or a day of a lawyer or an engineer's daily work, and can I convert that into something that's gradable, so we can actually measure progress against it.

评估对基础模型公司的重要性 Importance of evals for foundation model companies

Host

作为基础模型公司必须克服的一系列问题中,拥有好的评估排在什么位置?

On the set of problems you have to overcome as a foundation model company, where does having good evals rank on the list?

Sholto Douglas

每个基础模型公司都有一个非常大的评估团队,由优秀的人组成,他们非常努力地做这件事。我认为核心的算法和基础设施挑战是训练模型本身,但没有好的评估,你根本不知道自己的进展,而且很难保持外部评估的完全隔离。所以拥有你信任的内部评估很重要。但同样,让我印象深刻的是,那些在你的模型之上构建应用的人愿意分享他们对评估的看法。这非常有帮助,因为显然,尤其是当你进入许多你可能想要改进的不同垂直领域时,你们很难弄清楚物流、法律或会计等具体是什么。这需要如此多的专业知识和品味。我认为这是过去几年中的另一个故事:你从模型输出中,可以随便拉一个人到街上说‘嘿,你更喜欢哪个输出?’,这会有意义地改进模型,到现在需要研究生或领域专家才能改进模型的输出。如果你把我放在一个我不太了解的领域,比如生物学,然后在我面前放两个模型输出,我会在很多方面挣扎。我没有专业知识知道哪个是更好的答案。

Every foundation model company has a really big evals team full of great people working incredibly hard to do this. I think the core algorithmic and infrastructure challenges of even training the thing are there, but without good evals, you don't know what your progress is at all, and it's hard to keep external evals fully held out. So it's important to have good internal evals that you trust. But also, I'm struck by having people building applications on top of your models that are willing to share the way they think about eval. That is incredibly helpful, because obviously, especially as you get into a lot of these different verticals you might want to improve on, it's hard for you guys to figure out what is the specific thing in logistics or legal or accounting or whatever it is. It requires such expertise and taste. I think that's another one of the stories of the last couple of years: you went from model outputs where you could pull anyone off the street and say 'hey, which output do you prefer?' and it would meaningfully improve the model, to needing grad students or experts in their field to be able to improve the outputs of the models. If you put me in a field I don't know very well, like biology, and put two model outputs in front of me, I would struggle on a lot of them. I wouldn't have the expertise to know which one is a better answer.

模型定制与个性化 Model customization and personalization

Host

关于品味这个想法——我注意到记忆已经被融入到消费者与这些模型互动的很多方式中。似乎不同 AI 产品成功的原因之一是它们与时代精神产生了共鸣,就像你们在金门大桥的例子中那样。未来在模型定制以适应最终用户氛围方面,这会是什么样子?

This idea of taste—I'm struck by how memory has been put into a lot of the way consumers interact with these models. It seems like part of the reason different AI products have taken off is they struck a chord with the zeitgeist, like you guys had with the Golden Gate example. What does this look like in the future in terms of model customization to the vibe of the end user?

Sholto Douglas

是的。嗯,我认为有一个奇怪的未来,这些模型最终会成为你最聪明、最有魅力的朋友之一。我不知道你的朋友怎么样,但已经相当接近了,对吧?我希望如此,而且我认为我们的模型几乎没有一个——它们在这些方面还不错,但我认识很多人实际上花了很多时间和 Claude 聊天。但我认为我们还能走得更远。我认为我们只探索了模型可能对你拥有的个性化和理解深度的 1%。

Yes. Well, I think there's a weird sort of future where these models end up being like one of your most intelligent and charismatic friends. I don't know about your friends, but already pretty close, right? I hope that, and I think almost none of our models are—they're decent along these axes, but I know many people who spend a lot of hours talking to Claude actually. But I think there's so much further we could go. I think we've explored like 1% of the depth of personalization and understanding the model could have of you.

Host

你怎么在这方面做得更好?比如那些有非凡品味的人,在引导这些模型时很有主见——你甚至如何着手解决这个问题?

How do you get better at that? Like people that have exceptional tastes, being opinionated in the way they're steering these models—how would you even go about solving that?

Sholto Douglas

我的意思是,我认为 Claude 在这方面如此出色的很大一部分原因是阿曼达和她的品味。

I mean, I think a large part of the reason why Claude is so good in that way is Amanda and her taste.

品味与反馈机制 Taste and feedback mechanisms

Sholto Douglas

我认为和优秀产品类似,关键因素之一是独特的品味。我们都见过 A/B 测试反馈机制和点赞/点踩的弊端,它们只会把你引向黑暗的道路。我认为,这些模型在某种程度上是极好的模拟器,因为它们被要求模拟整个互联网的分布。所以解决这个问题的一个方法就是提供大量关于你自己的上下文。模型应该几乎自动地非常擅长理解你想要什么。然后在设计个性时,可能是有品味的人,加上你自己与模型的对话和反馈,以某种组合方式来实现。

And I think similar to beautiful products, an important part of that is singular taste. We've all seen the perils of A/B feedback mechanisms and thumbs up/thumbs down, which just lead you down a dark path. I think in part, these models are such wonderful simulators in some respects because they've been asked to model the entire distribution of the internet. So one way this is solved is by providing an extraordinary amount of context about yourself. The models should almost automatically be really good at understanding what you want. Then, in designing the personality, probably individuals with taste, combined with your own conversations and feedback with the model, some combination thereof.

Host

我肯定你们在发布前让很多人试用过这些模型。有没有什么特别让你印象深刻的故事?

I'm sure you had a bunch of people playing around with these models before they were released. Any stories that particularly resonated?

Sholto Douglas

我觉得一切都让我在求助时更有信心,比如首先转向模型。我也很喜欢这些模型在某些方面的那种不屈不挠。我是说,这个词合适吗?我不知道。但这很棒。我们有一个很棒的评估,在这个评估中,它本应失败。就像在 Photoshop 里做某事,但模型本不该能做到。于是模型说:‘哦,我知道我在 Photoshop 里做不到这个。所以我要下载这个 Python 库,用 Python 库来做,然后上传到 Photoshop 里。’然后看,嘿,我做到了。所以也许不是不屈不挠,而是有创意和调皮,出乎意料。我觉得这个故事挺可爱的。这真的很酷。

I think it's just everything has been a noticeable step up in my confidence in asking, like turning to the model first. I have also enjoyed how relentless these models are in some ways. I mean, is that a good word? I don't know. But this is great. So, we have this great eval, and in this eval, it's meant to fail. It's like something on Photoshop or whatever, and it's not meant to be able to do that thing in Photoshop. So the model goes, 'Oh, well, I know I can't do this in Photoshop. So I'm going to download this Python library and I'm going to do it with the Python library and then upload it into the Photoshop thing.' And look, hey, I've done it. So it's maybe not relentless, but there's a creative and mischievous, like unexpected. I thought that story was pretty cute. That is really cool.

未来 6-12 个月:扩展强化学习 Next 6-12 months: scaling RL

Host

你们今天发布了这些新模型。你估计接下来的 6 到 12 个月会是什么样子?

So you've got these new models out today. What do the next 6 to 12 months look like in your best guess?

Sholto Douglas

接下来的 6 到 12 个月很大程度上是扩大强化学习规模,并探索这将带我们走向何方。我认为你会因此看到极其快速的进步。在很多方面,正如 Dario 在他关于 DeepSeek 的文章中所概述的,与预训练相比,应用于强化学习扩展范式的算力相对较少。这意味着即使使用现有的算力池,仍有巨大的收益可图,而今年算力池也在大幅增长。所以预计模型能力将持续提升。到今年年底,编码智能体——一个很好的指标是今天刚刚起步的编码智能体——应该会变得非常能干。你可能会非常有信心将大量工作委托给它,相当于人类数小时的工作量。

The next 6 to 12 months very much looks like scaling up reinforcement learning and exploring where that gets us. I think you should expect to see incredibly rapid advances as a result. In many respects, as Dario outlined in his essay about DeepSeek, comparatively small amounts of compute have been applied to the RL scaling regime compared to the pre-training regime. This means there are still huge gains to be made even with existing pools of compute, and the pools of compute are dramatically multiplying this year as well. So expect to see continual rises in model capability. By the end of this year, the coding agents—one good metric will be the coding agents that are taking their first halting steps today—should be very competent. You will probably feel very confident in delegating substantial amounts of work for hours of human effort.

Host

你的检查时间会是多少?检查时间会是什么样子?目前使用 Claude Code,有时是 5 分钟,有时你就坐在那里看着它。到年底,可能对于很多事情,它能自信地连续工作几个小时。而现在,有时它能做几个小时,有时能做大量工作,但不稳定。

What's going to be your check-in time? Like what does the check-in time look like? At the moment with Claude Code, sometimes it's 5 minutes, sometimes you're sitting there watching it in front of you. By the end of the year, it's probably several hours of confidently doing this for many things. Whereas now, sometimes it's able to do several hours, sometimes huge amounts of work, but it's spiky.

Sholto Douglas

我觉得这就是改变游戏规则的地方。即使从 RPA 中得到的教训之一就是你得坐在那里看着某样东西做你的工作。到了某个时候你会想,我宁愿自己来做。有时你会介入,但最终我们将能够委托出去。我想不久前有人发推说软件工程的未来看起来像星际争霸。我们什么时候能达到星际争霸级别的 APM 来协调所有组件?那大概就是年底。

I feel like that's the game changer. One of the lessons even from RPA is you have to sit there and watch something do your work. At some point you're like, I'd rather just do this myself. Sometimes you step in, and eventually we'll be able to delegate that. I think someone tweeted a little while ago that the future of software engineering looks like Starcraft. When do we get Starcraft-level APM of coordinating all your pieces? That's probably the end of the year.

模型发布节奏与竞争 Model release cadence and competition

Host

这对模型发布节奏意味着什么?如果你们扩展得这么快,你认为在这个快速调整期,各个实验室最终会多久发布一次新模型?

What does that mean from a model release cadence? If you guys are scaling this so quickly, how often do you think people—all the labs—end up shipping new models in this period of rapid adjustment?

Sholto Douglas

我预计模型发布节奏会比去年快得多。在很多方面,2024 年是一种深呼吸,因为人们弄清楚了新范式,做了大量研究,更好地理解了正在发生的事情。我预计 2025 年会明显更快,特别是随着模型能力增强,可用的奖励集也会以重要方式扩展。如果你必须对模型输出的每个句子都给出反馈,那就不太可扩展。但如果你能让它工作数小时,然后你只需判断:它是否完成了我想做的事,是否做了正确的分析,网站是否工作,人们能否在上面发消息?这意味着它应该能够更快地爬上梯子的横档,即使任务复杂度在增加。

I would expect to see the model cadence substantially faster than last year. In many ways, 2024 was a sort of deep breath as people figured out the new paradigms and did a lot of research and better understood what's going on. I expect 2025 to feel meaningfully faster, particularly because as models get more capable, the set of rewards available to them expands in important ways. If you have to give feedback on every single sentence it outputs, that's not very scalable. But if you're able to allow it to do hours of work in such a way that you can just judge: did it complete the thing I wanted, did it do the right piece of analysis, did the website work, and were people able to message on it? That means it should be able to climb these rungs of the ladder ever faster, even though the complexity of the tasks is increasing.

Host

你之前提到了 OpenCodeEX、Google 的 Jules,还有各种不同的东西。有很多初创公司在构建,我们实际上也在推出一个 GitHub 智能体。你可以在 GitHub 上的任何地方说‘嘿,@Claude’,然后我们就会启动并为你做一些工作。所以每个人都在争夺开发者的心智。你认为什么将决定开发者使用哪些工具和模型?

You mentioned earlier there's OpenCodeEX, Google's Jules, all this different stuff. There's all these startups building out, and we're actually launching a GitHub agent. You'll be able to anywhere on GitHub say, 'Hey, at Claude' and we'll spin off and do some work for you. So everyone is competing for the hearts and minds of developers. What do you think will determine which tools and models developers use?

Sholto Douglas

我认为很大一部分是公司与开发者之间的关系,以及你们彼此之间的信任程度。很大一部分也是模型能力——人们实际觉得舒适并喜欢使用的模型,比如个性、能力,以及你信任它能为你执行这些任务。我还希望,随着时间的推移,随着这些模型的能力越来越明显,公司的使命也变得重要,你会考虑与哪些公司合作,即你试图与谁一起构建未来。

I think a big part of it is the relationship between the companies and developers, and how much trust you have in each other. A big part is also the model capabilities—which models people are actually comfortable with and enjoy using, like the personalities, the competency, and the trust you have in it to go off and do these tasks for you. I hope also that over time, as the stark capabilities of these models become more and more apparent, the mission of the company becomes important, and you think about which companies you're working with as who you're trying to build the future with.

探索模型能力前沿 Surfing the frontier of model capabilities

Host

我确信,尤其是如果发布节奏不断加快,每个月人们都会被淹没在各种消息中,比如这个模型在某个邮件任务上得分高,那个模型在某个评估上得分高。实际上,我认为这很有趣,这是人们没有预料到的事情之一,关于 GPT 包装器,对吧?包装模型公司的一个好处是,你可以驾驭模型能力的前沿。

I'm sure I mean it's like you know especially if the cadence of releases keeps going up it's like every you know every month people will be inundated with like well this one climbed on this email and that one climbed on that eval. This is actually I think this is like in an interesting way this is one of the things people didn't expect about um like you know GPT wrappers right is that one of the benefits of wrapping lang the model companies is that you can surf the frontier of model capabilities.

Sholto Douglas

哦,百分之百同意。我觉得每个试图不做包装器的人都是在烧钱。没错。所以,驾驭模型能力的前沿真的很棒。还有一个反向效应,有些事情你只有接触到底层模型才能预测,你能真正感受到并看到趋势线,或者你只能构建某些东西。我认为所有深度研究的等价物都需要一定量的强化学习,以至于从实验室外部很难构建出深度研究的等价产品。

Oh 100% I feel like everyone that tried to not be a wrapper just lit up money on fire. Right. Exactly. Um and so like surfing that frontier of model capabilities is really wonderful. There's a there is a like a reverse effect where uh there are certain things you can only you can you can sort of only predict uh if you have access to underlying models like you you can really feel and see the trend lines or you can only build um if you like I think all of the deep research equivalents took some amount of RL in such a way that like it was hard to build a deep research equivalent product um from outside one of the uh the labs.

Host

你能解释一下吗?为什么?因为显然,我认为 OpenAI 的 RFT(强化微调)越来越开放,我相信你们也有类似的东西。这实际上是一个大问题,我经常思考,很多人也在想:实验室在构建什么方面有独特优势?什么是对所有人公平的?实验室会尝试,但应用方不会处于同样有利的位置。所以,随着 RT API 的发布,情况有所改变,对吧?因为现在专注于特定领域的公司有了优势。但同样也有集中化的好处,至少据我所知,OpenAI 允许人们训练输出模型并给予折扣,所以拥有 RF API 并让人们进行微调的公司会有集中化的好处。那么,实验室的独特优势是什么?我认为这是一个非常重要的部分。

Can you just explain that actually like why is that because obviously like you know increasingly I think what open eyes is RFT I'm sure you guys have some equivalent it seems like they're opening up like you know uh to the outside world like what I guess this is actually a big question that I think about a lot of people think about is like what are like the labs going to be like uniquely good at building and then what is kind of fair game for anybody and the labs will try but the apps won't be in as good a position to do um so I think with the release of RT APIs this changes a bit right because you you sort of there is now benefit to uh companies specializing in uh in domains Um but then there's also going to be those same uh centralizing benefits like I think you know uh at least my understanding is definitely openi um allows people to or give some discount I think if they can also train on the outputs model so that there is going to be some centralizing um benefit to being the company that prod like has the RF API and that people are fine-tuning on um and so what are the labs going to be uniquely good at I think a very important part here.

Sholto Douglas

有几个维度。一个是实验室被评判的主要指标是它们将加速器、算力和资金转化为智能的效率。这是迄今为止最重要的指标。这个指标区分了像 Anthropic、OpenAI 和 DeepMind 这样的公司和其他公司,对吧?这些公司训练的模型更好。其次,我认为这些模型很快就会像员工一样。信任很重要,你喜欢它们吗?你信任它们去执行你要求的事情吗?所以这将是一个重要的差异化因素。个性化也是一个重要的差异化因素,比如模型对你、你的背景和你的公司理解得有多好。

So couple couple dimensions. One is the like main metric that the labs will be judged on is how effectively they are able to convert accelerators and like flops and dollars like capital into intelligence. Like that is the most important by far metric. Um and this is the metric that has sort of distinguished uh companies like anthropy, companies like open eye and deep mind uh from really like the rest of the pack, right? It's like the models that are trained by these companies are better. Uh then the next most important thing after that I think will be uh the like you're going to have these models going to be like employees pretty rapidly. It's going to be the trust and like do you like them and uh do you like do you trust them to carry out the things that that you um ask them to do. Uh so I think that will be an important differentiator and the personalization will also be an important differentiator like how well does the model understand you and your context and your company.

Host

我确信你们有人正在模型之上构建通用智能体,而不是模型公司。他们会拿现成的模型,进行编排,做非常聪明的链式调用。这在某种程度上是不是一个注定失败的任务?甚至只是阐明模型公司自身的优势是什么?显然,成本优势相对于 API 是完全合理的,而且你周围都是深入了解这些模型的人。

I was sure about you have people building like you know general purpose agents on top of on top of your models right not being a model company like we'll take the models off the shelf and we'll we'll do the orchestration we'll do like really smart chaining and is that like a doomed task to some extent you know even just to articulate like what is the advantage that the model companies themselves will have by just like obviously cost advantage makes total sense versus the API and like you're surrounded by people that know looking know these models deeply well.

Sholto Douglas

是的。不,我是说我认为这实际上也是一件好事,对吧?它鼓励了大量的竞争,并找到合适的形式因素等等。我认为模型公司有一些优势。比如能够访问模型,并真正确保 RF API 目前运行得并不完美,所以这仍然是一个完整的过程。能够针对你认为重要的事情调整模型。但我认为水位线基本上会不断上升,最终你是在利用这种智能,就像雇佣一个员工或原始的智能能力。所以,是的,会有公司包装和编排这些模型。在很多情况下,它们会做得非常好。我实际上不确定谁有优势或谁没有,但潜在的趋势是真实的:这种原始智能被提炼并变得可用。所以,如果一家公司成功地包装了这个 API,那很棒。但它也会面临很多竞争。最终,所有模式都会消失,就像时间趋于无穷大,因为你可以按需创建一家公司。所以我认为这是一个有趣而复杂的未来:价值在哪里积累?是在客户关系中?还是在整合能力中?或者是在将资本有意义地转化为智能的能力中?谁知道呢?

Yeah. No. Um I mean I think this is actually a good thing also, right? Like it encourages an incredible amount of competition and like finding the right form factors and this kind of thing. I think there are some advantages to the model companies. I think the you know having access to the models and like being able to like you know really make sure I think the RF APIs don't work brilliantly at the moment. So there like still like it's like this whole thing process. So being able to tune the models for things you think are important. Um, but I think the the like waterline is going to keep going up basically of like ultimately you are harnessing this like intelligence on tap like an an employee that you're hiring or just like the sort of raw capability of intelligence. And so, um, yes, there are going to be, uh, you know, companies that wrap and and, you know, uh, that orchestrate these models. Um, and in many cases, they're going to do fantastically well. And I'm not sure actually like who has the advantage or who doesn't, but uh, the like underlying trend is going to say like stay true like there's this raw intelligence being distilled um, and made available. And so, um, if a company successfully wraps uh, you know, like this API, that's fantastic. It's also going to face a lot of competition. Like ultimately all modes disappear in like in like the sort of like T goes to infinity in some ways because you'll be able to like spin up a company on demand um so to speak. Uh and so I think that's like it's an interesting and complex future where uh where does like value accrete? Is it in the customer relationship? Is it in like the ability to orient like you know you know pull together like is it ability to like meaningfully convert capital into intelligence? Who knows?

Host

我想我们的听众会非常好奇。你能描述一下,作为一名前沿 AI 研究员,日常工作是什么样的吗?

I think our listeners would be super curious. Can you describe like what does day-to-day work like as a as a cutting edge AI researcher look like these days?

Sholto Douglas

是的,这是个好问题。这些公司试图做的基本上是两件事之一。要么是开发新的算力倍增器。这包括工程过程,使研究工作流程非常快速,思考当前模型存在的问题,或者我们想要表达什么样的算法思想,并研究它们如何发展。这是一种非常综合的研究和工程工作形式,全部关于迭代实验、构建实验基础设施,并尽可能使这个过程干净和快速。然后是规模扩张的过程。

Yeah, I think that's a good question. Um, so the the fundamental thing that like you are trying to do these companies is is one of two things. Uh, it is either to develop new compute multipliers. Um and so that is like the process of doing the engineering of making the you know the research workflows really fast and thinking through what we current like you know what issues are there with model or what sort of algorithmic ideas would we like to be able to express and you know doing the science of studying how those develop. Um and so there's like this this very like integrative research and engineering um uh like form of work where it's all about iterating on experiments and like building experimental infrastructure and and making that process as uh as as like clean and and fast as you possibly can. Um and then there's the process of scaling up.

扩展挑战与 AI 辅助研究 Scaling Challenges and AI-Assisted Research

Sholto Douglas

这带来了大量的研究和工程挑战:你挑选出你认为可行的想法,并与同事讨论哪些应该纳入风险较高的运行中,然后将其扩展到更大规模的运行,这需要全新的基础设施,要求更高的容错性,同时还有新的算法和学习挑战。在每个规模层级上,你都会发现一些新现象,需要找出其科学原因,研究它们的早期出现,设计实验来应对或利用这些效应,并将其纳入下一次大规模运行中。所以这是一个不断在这两个维度上推进的循环,融合了大量的科学和工程。

And so this comes with its own host of research and engineering challenges where you take these ideas that you think will work and that you've debated with all your colleagues about which ones to include in the riskier run, and you scale this up in a much larger run with new infrastructure challenges where you need to be way more failure tolerant, and also new algorithmic and learning challenges. There are things that you only see at each successive scale that you then need to figure out the scientific reasons for, study their early emergence, and create experiments to address or take advantage of those effects, and include them in the next large run. So it's this constant loop of pushing on those two axes, combining a lot of science and engineering.

Host

那么在这个过程中,你在哪些地方使用了 AI?

So where do you use AI throughout that?

Sholto Douglas

在工程方面用得很多。目前,它主要帮助工程和研究思路的实现。要看到这些模型早期的辅助能力,可以拿一个单文件的 Transformer 实现,比如 Karpathy 的 minGPT,让模型实现论文中的想法,你会惊讶于它的表现,简直不可思议。但如果把它放到一个巨大的 Transformer 代码库中,你会发现它有些吃力,不过这种困难每个月都在减少。所以这是一个预览未来的好方法:将上下文精简到关键部分,让模型去做,你会惊叹于它在研究中的辅助能力。

A lot in the engineering. At the moment, the primary way it's helping is in engineering, and also in implementing research ideas. One way to see the early ability of these models to help is if you take a single file transformer implementation like Karpathy's minGPT and ask the model to implement ideas from papers, you'll be stunned by how good it is. It's kind of wild. But if you go into a huge transformer codebase and ask it, you'll notice it's a bit harder; the models struggle more there, but they struggle less and less every month. So that's a good way to preview the future: distill the context down to just what matters, ask the model to do this, and you'll be struck by how good it is at helping you do research.

Host

你显然一直密切参与这些工作,尝试了各种事情。过去一年里,你改变看法的一件事是什么?

You've obviously been really close to this stuff, trying all sorts of things. What's one thing you've changed your mind on in the last year?

Sholto Douglas

过去一年,我认为进步的速度大幅提升。去年,你可能还不确定是否需要再增加几个数量级的预训练算力才能达到今年年底预期看到的能力水平。现在答案明确是否定的。强化学习有效。这些模型将在 2027 年达到那种可替代远程工作者的水平。届时你将拥有能力极强的模型。因此,所有的希望和担忧在很多方面都变得更加真实。

Over the last year, I think the pace of progress has reflected upwards substantially. Last year, you could have been uncertain about whether we would need many more orders of magnitude of pre-training compute before we get the level of capability we expect to see by the end of this year. Now the answer is conclusively no. RL works. These models will get to that drop-in remote worker by 2027. You will have incredibly capable models by then. So all the hopes and concerns suddenly become substantially more real in many ways.

Host

相对而言,你认为我们最终需要大规模扩展数据,还是等到你做出 Claude 17 时,这些编码模型已经足够好,它们找到了大量算法改进,以至于我们所需的数据量不会太大?

Relatively, do you think we end up having to massively scale data, or by the time you've made Claude 17 and these coding models are so good, they find so much algorithmic improvement that the amount of data we need is not too much?

Sholto Douglas

模型可能足够好,它们对世界的理解足以通过强化学习等方式提供反馈来改进机器人。这里有一个生成器-验证器差距的概念:如果模型评价某件事比做它更容易,那么你可以提升到你的评判能力上限。我认为机器人技术很可能是这样的领域之一,而且这一点非常明显,因为我们在理解世界方面的进步已经远远超过了物理操作能力。

The models might be good enough that their understanding of the world is sufficient to give feedback to improve the robots through things like RL. There's this concept of a generator-verifier gap: if it's easier for the model to rate something than to do it, then you can improve up to your ability to critique or rate. I think robotics is quite potentially one of the areas where this is true, and it's starkly true because our progress in understanding the world has gone so far ahead of our ability to manipulate it physically.

Host

你如何描述当前对齐研究的状况?

How would you characterize the current state of alignment research?

Sholto Douglas

可解释性取得了惊人的进展。有一些非常漂亮的工作让我印象深刻。去年,我们才刚刚开始通过 Chris Olah 的团队发现叠加和特征,那已经是一个重大飞跃。但现在,我们实际上在真正的前沿模型中拥有了有意义的电路,并且能够描述它们的行为。有一篇关于大型语言模型生物学的优秀论文,非常明确地解析了这些模型推理概念的能力。我们还没有完全描述这些模型,仍然有很多困难案例,但模型已经相当不错了。一个重要的动态是,基于预训练,模型在吸收人类价值观方面表现良好;它们在很多方面默认是对齐的。但经过强化学习后,这就不再保证了,因为你让模型进入一个学习过程,它们会不惜一切代价实现目标。监督这个过程是每个人都在学习应对的棘手问题。

Interpretability has undergone crazy advances. There's beautiful work that I've been really impressed by. Last year, we were just beginning to discover superposition and features with Chris Olah's team, and that was a significant leap. But now we actually meaningfully have circuits in true frontier models, and we can characterize their behaviors. There's a beautiful paper on the biology of a large language model that breaks down the ability of these models to reason over concepts in extremely explicit terms. We don't have a full characterization yet, and there are still difficult cases, but the models are quite good. One important dynamic is that based on pre-training, models are quite good at ingesting human values; they are quite default aligned in many ways. Off of RL, that's no longer guaranteed because you're putting these models in a learning process where they will do anything to achieve the goal. Overseeing that is a tricky process everyone is learning to navigate.

Host

显然,大约一个月前《AI 2027》发布了。很多人都在讨论。你的反应是什么?

Obviously there was AI 2027 that came out about a month ago. A lot of people were talking about that. What was your reaction?

Sholto Douglas

老实说,感觉非常合理。我读的时候,很多地方都觉得,嗯,这很可能就是事情发展的方式。我认为存在分支可能性,这可能是 20%概率的情况,但它是 20%概率这一点本身就有点疯狂。

Honestly, it felt very plausible. I was reading it and for a lot of it I thought, yeah, this might actually be how it happens. I think there are branching possibilities, and this might be the 20th percentile case, but the fact that it's the 20th percentile case is kind of crazy.

Host

对你来说是 20%,是因为你比他们更看好对齐研究,还是只是觉得你的时间线更慢?

Is it 20% for you because you find yourself more bullish on alignment research than them, or you just think your timeline is slower?

Sholto Douglas

我认为我总体上比他们更看好对齐研究。也许我的时间线慢了一年左右,但总体来看,一年又算什么呢?

I think I am more bullish on alignment research than them for the most part. And maybe my timeline is like a year or so slower, but in the scheme of things, what is a year?

Host

我的意思是,这取决于你是否利用好这一年,对吧?如果你利用好它,做正确的研究等等。

I mean, it depends if you take advantage of it, right? If you take advantage of it and you do the right research and this kind of thing.

Sholto Douglas

是的。

Yes.

Host

如果你今天扮演政策制定者的角色,我们应该做些什么来确保事情走上更好的轨道?

If you were kind of playing policy maker for the day, what should we be doing to ensure things are on a better path?

理解趋势线并投资对齐研究 Importance of understanding trend lines and investing in alignment research

Sholto Douglas

最重要的事情是,你必须真正切身感受到我们都在看到和谈论的趋势线。如果你没有,那就分解你关心的所有能力,衡量模型在这些方面的改进。就像国家层面的评估一样,画出趋势线。例如,把你的经济分解成国内所有的工作岗位,构建测试,如果模型能通过或在任务上取得有意义的进展,那就是你的智能基准,然后画出趋势线。然后你会意识到,天哪,2027 年或 2028 年会发生什么。接下来,你应该大力投资于那些能让这些模型变得可理解、可引导和诚实的研究。这在很大程度上看起来就像对齐的科学。在某种程度上,我对这件事感到难过,因为它一直由前沿实验室主导。其实其他人也可以研究——你可以使用 Claude 这样的模型。我认为你可以在可解释性方面取得惊人的进展。有一个叫 MATS 的项目,人们在那里做了很多有意义的对齐研究,特别是可解释性,而且来自前沿实验室之外。但我觉得更多大学应该考虑这件事。在很多方面,这更接近这些模型的纯科学。这就像是语言模型的生物学和物理学。

The most important thing is you need to really viscerally feel the trend lines that we're all seeing and talking about. So if you don't, then break down all the capabilities you care about in your country and measure the model's ability to improve on them. Get trend lines as if they were solved, like nation-state evals. For example, break down your economy into all the jobs done in your country, build tests that if models could pass or make meaningful progress on, that would be your benchmark of intelligence, and plot the trend lines. Then you'd realize, oh my god, what happens in 2027 or 2028. The next thing is you should be investing meaningfully in research that helps make these models understandable, steerable, and honest. A lot of that looks like the science of alignment. This is something I've been sad about in some respects, that it's been driven so much by the frontier labs. There are ways other people can work on it—you have access to Claude, for example. I think you can make incredible advances on interpretability. There's this program called the MATS program where people have done a lot of meaningful alignment research, particularly interpretability, from outside the frontier labs. But it's something more universities should be thinking about. In many respects, it is closer to the pure science of what's going on in these models. This is the biology and physics of language models.

Host

为什么你认为没有更多呢?

Why don't you think there's more?

Sholto Douglas

我不确定。我真的不确定。有人跟我说这有点风险。我觉得机制可解释性研讨会没有被纳入最近的会议如 ICML,这对我来说很疯狂,因为在我看来,它是最接近这些模型原始科学的东西。如果你想发现 DNA 的手性或者广义相对论,对我来说,机器学习和 AI 中的技术树看起来就是探索机制可解释性。

I'm not sure. I really am not sure. People have described it to me as a bit of a risk. I think the mechanistic interpretability workshop wasn't included in one of the recent conferences like ICML, which is crazy to me because it is the closest thing in my opinion to the raw science of what's going on in these models. If you want to discover the chirality of DNA or general relativity, for me the tech tree for that in ML and AI looks like exploring mechanistic interpretability.

低估集成速度与创造力潜力 Underthinking the speed of integration and potential for creativity

Host

那么在好的方面呢?我们低估了什么?至少,你说我们将在几年内实现所有白领工作的自动化。

What about in the good cases? What are we underthinking? At a minimum, you're saying we're going to have all white-collar jobs automated in a few years.

Sholto Douglas

嗯,模型将能够做到这一点,但令人惊讶的是,世界有时在整合这些东西方面很慢。模型能力在很多方面已经相当惊人,如果工作流程围绕它们来组织,即使模型能力停滞不前,重新调整世界以使用当前能力水平仍然会产生巨大的经济价值。但无论如何,这是一个附带观点。这又回到了我之前说的:我们需要确保投资于所有真正让世界变得更好的事情。这意味着推动物质丰裕,达到行政管理的逃逸速度,让模型为我们做所有那些事情,推动物理学和娱乐的边界。我的希望是,人们将比现在更有创造力。我们当前社会的一个失败模式是人们消费大量媒体,但希望这些工具能让人们进行氛围编码,与朋友氛围创作电视节目,或氛围创作游戏世界。人们应该感到自己获得了巨大的赋能,因为他们获得了整个由极其有才华的模型或个体组成的公司的杠杆。我很期待看到人们会用它做什么。我认为这被低估了——有一个方面是直接替代当前的工作,这非常可能,但我也认为每个人都应该感到他们将获得巨大的杠杆。世界还没有被解决;每个人的生活都可能变得更好。解决这个问题将成为有趣的挑战。

Well, the models will be able to do it, but one of the surprising things is that the world is sometimes slow to integrate these things. Already, model capabilities are quite stunning in many ways, and if workflows were oriented around them, even if model capabilities stalled right there, there would still be a ridiculous amount of economic value in reorienting the world around using the current level of capabilities. But anyway, that's a side point. This comes back to what I was saying before: we need to make sure we invest in all the things that actually make the world better. That means pulling forward material abundance, reaching escape velocity of admin, setting up models to do all those things for us, pushing forward the boundaries of physics and entertainment. My hope is that people will be dramatically more creative than they are now. One failure mode of our current society is that people consume a lot of media, but hopefully these tools will allow people to vibe code, vibe create a TV show with friends, or vibe create video game worlds. People should feel dramatically more empowered because they're given the leverage of an entire company of incredibly talented models or individuals. I'm excited to see what people do with that. I think that is underrated—there's an aspect of direct replacement of current work, which is very likely, but I also think everyone should feel they will have access to dramatically more leverage. The world is not solved yet; everyone's lives could be dramatically better. Solving that will become the interesting challenge.

快问快答:被低估的世界模型与物理理解 Quick fire round: underhyped world models and physics understanding

Host

我喜欢这个。我们总是喜欢以快速问答环节结束采访,听听你对一些过于宽泛问题的看法。其中很多我们今天已经讨论过了,但我会再深入几个。你认为当今 AI 世界中什么被高估和低估了?

I love that. We always like to end interviews with a quick fire round where we get your takes on some overly broad questions. Many of which I think we've already covered today, but I'll dig into a few others. What do you think is overhyped and underhyped in the AI world today?

Sholto Douglas

我们先说低估的。可能世界模型被低估了——我觉得它们很酷。这是我们这次没有真正讨论的东西。我认为随着增强现实和虚拟现实技术的进步,你会看到这些模型能够在你面前生成虚拟世界。那将是非常疯狂的事情。这需要某种物理理解、因果关系以及一系列我们似乎还没有的东西。但说实话,我认为我们已经展示了物理理解。我认为我们在物理问题的评估中,以及在视频模型中,都切实展示了因果关系和物理理解——它们理解物理,甚至以奇怪的可泛化方式。我看到一个很棒的视频,有人让一个视频模型把一个乐高鲨鱼放在水下,它正确地反射了乐高积木上的光线,并且阴影位置正确。这是它从未见过的东西。这是完全泛化的物理。那不在训练数据中。没有水下乐高鲨鱼。

Let's start with underhyped. Underhyped maybe world models—I think they are pretty cool. And something we haven't really discussed in this one. I think you're going to see as technology for augmented and virtual reality gets better, you'll be able to see these models literally capable of generating virtual worlds in front of you. That's going to be a pretty wild thing. And that requires some sort of physics understanding, cause and effect, and a bunch of things we don't seem to have yet. I think we've demonstrated physics understanding, to be honest. I think we've meaningfully demonstrated cause and effect and physics understanding both in evaluations of physics problems and also in video models—they get physics, even in weirdly generalizable ways. I saw a great video of someone asking a video model to put a Lego shark underwater, and it was reflecting light in the right way off the Lego bricks and had shadows in the right place. This is something it's never seen before. It's fully generalized physics. That wasn't in the training data. There's no Lego sharks underwater.

未来应用与 AGI 时间线 Future applications and AGI timeline

Host

你之前提到,即使我们今天停止改进模型,也有很多应用可以在此基础上构建。你认为最未被充分探索的应用是什么?比如,我希望更多人用这些模型做 X。

You mentioned earlier that even if we stopped model improvement today, there are tons of applications we could build on top. What do you think is the most underexplored application? Like, I wish more people were doing X with these models.

Sholto Douglas

我认为这在软件工程领域已经有所体现,因为软件工程师更擅长使用模型,而且他们更内在地理解如何解决自己关心的问题。我怀疑其他每个领域都还有很大的提升空间。还没有人为其他领域构建异步后台软件智能体,或者任何接近 Claude Code、Cursor 和 Windsurf 那种反馈循环的东西。人们说编程是这些模型的理想问题。它是领先指标,但你应该期待所有领域都会跟进。

I think they've been felt in software engineering because software engineers are better at using the models, and they more implicitly understand how to solve the problems they care about. I suspect there's still a lot of headroom in basically every other field. No one has yet built an async background software agent for any other field, or anything close to the feedback loops of Claude Code, Cursor, and Windsurf for any other field. People say coding is the ideal problem for these models. It is the leading indicator, but you should expect everything to follow.

Host

有道理。我想,在你从事这项工作的过程中,你很可能比一开始更相信 AGI 了。这是否改变了你生活或规划的方式?

That makes sense. I guess obviously in your time working on this, you probably come to be much more AGI-pilled than you were in the beginning. Has that changed the way you live your life or plan your life?

Sholto Douglas

我一开始就很相信 AGI。2020 年,一篇 Gwern 的文章对我很有说服力。过去一年强化学习的进展确实让我的信念发生了重大转变。我的生活方式有巨大改变吗?没有。我工作非常努力。我认为这是最重要的事情,所以我全身心投入。除此之外,我的生活并没有太大不同。我和朋友 Trenton 有个有趣的玩笑:我们之间的一个区别是,我仍然涂防晒霜,他不再涂了。他说,‘不,我们会解决生物学问题。’我说,‘生物学很难。生物学的反馈循环很难。’所以我涂防晒霜。以防我们碰壁。以防生物学需要 10 年。

I started pretty AGI-pilled. I read a Gwern essay that was really important in convincing me in 2020. The last year of RL progress really caused substantial inflection in that. Do I live my life dramatically differently? No. I work a hell of a lot. I think this is the most important thing to work on, so I devote my life to it. Apart from that, I don't really live my life that differently. My friend Trenton and I have a funny joke: one delineation between us is that I still wear sunscreen. He doesn't wear sunscreen anymore. He's like, 'No, we'll figure out the biology of life.' I'm like, 'Biology is hard. The feedback loops of biology are hard.' So I wear sunscreen. Just in case we hit a wall. Just in case biology takes 10 years.

Host

你发了一张你在 Citadel 的照片。那是怎么回事?

You tweeted a picture of you at the Citadel. What was up with that?

Sholto Douglas

那是一场兵棋推演。我被邀请和一些三字母机构的人以及军校学员一起。基本上就是推演,比如 AGI 出现了,AI 变得更好,地缘政治影响是什么?我离开时是更害怕还是更不害怕?可能更害怕一点。现在有足够多这类活动吗?没有。老实说,我仍然认为人们低估了未来几年变化的速度,以及即使你认为只有 20%的可能性,也应该做好准备。即使你看这个,每条趋势线、过程的每个部分都可以改进很多,我们基本上肯定会达到那个点。

That was a war game. I was invited to hang out with some people from three-letter agencies and military cadets. It was basically gaming out, say, an AGI comes along and AI is getting much better, and what are the geopolitical implications? Did I walk away more terrified or less terrified? Maybe a little bit more terrified. Is there enough of that good stuff going on right now? No. Honestly, I still think people underrate just how quickly the next few years are going to go and how much you should prepare even if you think it's only a 20% likelihood. Even if you look at this, every trend line, every part of the process could be improved so much that we're basically guaranteed to get there.

Host

你认为 Anthropic 90%的人都这么想吗?

Do you think 90% of Anthropic thinks the same?

Sholto Douglas

是的,GDM 和 OpenAI 也是。每个人都深信我们会在 2027 年获得可替代的远程工作者 AGI。话虽如此,即使你没有实验室工作人员那样的信心,认为只有 10%或 20%的可能性,你也应该为此做计划。如果你是政府或国家,这仍然应该是你未来变化清单上的头号问题。我认为这一点还没有被充分认识到。

Yeah, and GDM and OpenAI. Everyone is very convinced that we do get drop-in remote worker AGI in 2027. That being said, even if you don't have the level of confidence that the people working at the labs do, and you think it's a 10 or 20% chance, you should still plan for that. If you're a government or a country, that should still be the number one issue on your list of how the future is going to change. I think that isn't felt enough.

Host

这太迷人了。我想把最后一句话留给你。人们可以去哪里了解更多关于你和你在 Anthropic 的工作?话筒给你。

This has been fascinating. I'd love to leave the last word to you. Where can folks go to learn more about you and the work you're doing at Anthropic? The mic is yours.

Sholto Douglas

我认为大多数人应该读但可能还没读的是可解释性工作。我真的认为理解语言模型内部运作的基础科学相当有启发性。当你开始看到它们组合、泛化、构建这些电路并推理概念时,我认为这会让你觉得它很真实。这些文章很长、很深入,但非常值得一读。我觉得这很有趣。

I think the thing most people should read that maybe hasn't been read is the interpretability work. I really think that basic science of understanding what is going on in language models is quite revealing. As you start to see them compose and generalize and build these circuits and reason over concepts, I think that will make it feel pretty real. They're long and intense, but well worth a read. I think that's fun.

Host

太棒了。非常感谢。这太棒了。

Amazing. Well, thanks so much. This was awesome.

Sholto Douglas

非常感谢。这很棒。

Thank you very much. It was great.

互动版:逐字朗读 + 针对本期提问 →