AI 前沿:DeepSeek、Claude Opus 4.5 与全球竞赛

AI State of the Art: DeepSeek, Claude Opus 4.5, and the Global Race

内森·兰伯特 Nathan Lambert · Lex Fridman 播客 · 2026-01-31 · 约 265 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Lex Fridman 与 Sebastian Raschka 和 Nathan Lambert 探讨 AI 最新突破,涵盖 DeepSeek 时刻、开源模型以及中美实验室的竞争格局。

Lex Fridman discusses the latest AI breakthroughs with Sebastian Raschka and Nathan Lambert, covering the DeepSeek moment, open-weight models, and the competitive landscape between US and Chinese labs.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 101)

全文 · Full transcript(中英对照)

引言与嘉宾 Introduction and Guests

Host

以下对话围绕人工智能的前沿技术展开,包括过去一年中令人兴奋的技术突破和发展,以及我们预计今年可能发生的一些有趣事情。有时内容会非常技术性,但我们努力确保非专业人士也能理解,同时绝不降低深度。非常荣幸和高兴能与 AI 社区中我最喜欢的两位人物——Sebastian Raschka 和 Nathan Lambert——一起制作这期节目。他们都是备受尊敬的机器学习研究员和工程师,同时也是出色的沟通者、教育者、作家和推特/X 博主。Sebastian 是两本书的作者,我强烈推荐给初学者和专家:一本是《从零构建大型语言模型》,另一本是《从零构建推理模型》。我坚信在机器学习计算机科学领域,学习和理解某件事的最佳方式就是自己从零构建。Nathan 是艾伦人工智能研究所的后训练负责人,也是关于基于人类反馈的强化学习的权威著作的作者。他们都有很棒的 X 账号和 Substack。Sebastian 在 YouTube 上有课程,Nathan 有播客,大家都应该关注这些。这里是 Lex Fridman 播客。如需支持,请查看描述中的赞助商,那里也有联系我、提问和获取反馈的链接。现在,亲爱的朋友们,有请 Sebastian Raschka 和 Nathan Lambert。

The following is a conversation all about the state of the art in artificial intelligence, including some of the exciting technical breakthroughs and developments in AI that happened over the past year, and some of the interesting things we think might happen this upcoming year. At times, it does get super technical, but we do try to make sure that it remains accessible to folks outside the field without ever dumbing it down. It is a great honor and pleasure to be able to do this kind of episode with two of my favorite people in the AI community, Sebastian Raschka and Nathan Lambert. They are both widely respected machine learning researchers and engineers who also happen to be great communicators, educators, writers, and Twitterers, X posters. Sebastian is the author of two books I highly recommend for beginners and experts alike. First is build a large language model from scratch and build a reasoning model from scratch. I truly believe in the machine learning computer science world, the best way to learn and understand something is to build it yourself from scratch. Nathan is the post-training lead at the Allen Institute for AI and author of the definitive book on reinforcement learning from human feedback. Both of them have great X accounts, great Substack. Sebastian has courses on YouTube, Nathan has a podcast, and everyone should absolutely follow all of those. This is the Lex Fridman podcast. To support it, please check out our sponsors in the description, where you can also find links to contact me, ask questions, get feedback, and so on. And now, dear friends, here's Sebastian Raschka and Nathan Lambert.

DeepSeek 时刻与国际竞争 DeepSeek Moment and International Competition

Host

所以,我认为看待这一切的一个有用视角是 DeepSeek,所谓的 DeepSeek 时刻。这大约发生在一年前,2025 年 1 月,当时开放权重的中国公司 DeepSeek 发布了 DeepSeek R1,我认为公平地说,它让所有人惊讶,因为其性能接近或达到最先进水平,而据称使用的算力更少、成本更低。从那时到现在,AI 竞争在研究层面和产品层面都变得疯狂,一直在加速。我们今天来讨论这一切,也许可以从一些尖锐的问题开始。在国际层面上,谁在赢?你会说是中国的一批公司还是美国的一批公司?Sebastian、Nathan,很高兴见到你们。那么,Sebastian,你认为谁在赢?

So, I think one useful lens to look at all of this through is the DeepSeek, so-called DeepSeek moment. This happened about a year ago in January 2025 when the open-weight Chinese company DeepSeek released DeepSeek R1 that I think it's fair to say surprised everyone with near or at state-of-the-art performance with allegedly much less compute for much cheaper. And from then to today, the AI competition has gotten insane both on the research level and the product level. It's just been accelerating. Let's discuss all of this today and maybe let's start with some spicy questions if we can. Who's winning at the international level? Would you say it's a set of companies in China or the set of companies in the United States? And Sebastian, Nathan, it's good to see you guys. So, Sebastian, who do you think is winning?

Nathan Lambert

嗯,赢是一个非常宽泛的术语。我会说,你提到了 DeepSeek 时刻,我确实认为 DeepSeek 绝对赢得了从事开放权重模型工作的人心,因为他们将这些作为开放模型分享。我认为赢有多个时间尺度。我们有今天,有明年,有十年后。我确定的一件事是,我不认为在 2026 年的今天,会有任何一家公司拥有其他公司无法获得的技术。这主要是因为研究人员经常换工作、换实验室。他们轮换。所以,我认为在技术获取方面不会有明确的赢家。然而,我认为差异化因素将是预算和硬件限制。所以,我不认为想法会是专有的,但实现它们的方式或资源会是。因此,我目前看不到一个赢家通吃的场景。我目前看不到。

Um so, winning is a very broad term. I would say you mentioned the DeepSeek moment and I do think DeepSeek is definitely winning the hearts of the people who work on open-weight models because they share these as open models. Winning, I think, has multiple time scales to it. We have today, we have next year, we have in 10 years. One thing I know for sure is that I don't think nowadays, 2026, that there will be any company who has access to a technology that no other company has access to. And that is mainly because researchers are frequently changing jobs, changing labs. They rotate. So, I don't think there will be a clear winner in terms of technology access. However, I do think the differentiating factor will be budget and hardware constraints. So, I don't think the ideas will be proprietary, but the way or the resources that are needed to implement them. And so, I don't see currently a take-it-all scenario where a winner takes it all. I can't see that at the moment.

Host

Nathan,你怎么看?

Nathan, what do you think?

Nathan Lambert

你会看到各个实验室在他们试图做的事情上投入了不同的精力。我认为要界定我们录制这个的时间点,对 Anthropic 的 Claude Opus 4.5 模型的炒作已经绝对疯狂了,我的意思是,我在过去几周里使用并构建了一些东西,它几乎到了炒作感觉有点像梗的地步,这有点好笑,因为这是非常自然的。然后,如果我们回到几个月前,根据笔记中的发布日期,Google 的 Gemini 3 发布了,那次发布的营销和令人惊叹的因素似乎非常高。但随后,在 11 月底,Claude Opus 4.5 发布了,炒作一直在增长。但 Gemini 3 在此之前,感觉人们不太谈论它,尽管它刚出来时,每个人都觉得‘这是 Gemini 重新夺回 Google 在 AI 领域结构性优势的时刻’。Gemini 3 是一个很棒的模型,我仍然在使用它。只是差异化程度较低。我同意 Sebastian 所说的,所有这些想法空间都非常流动,但从文化上讲,Anthropic 以在代码上大力下注而闻名,他们的 Claude 代码产品目前对他们很有效。所以,我认为即使想法流动得非常自由,这其中很多都受到人力和组织文化的瓶颈,而 Anthropic 至少表现得最不混乱。这是一个优势,如果他们能保持一段时间,但另一方面,中国有很多令人不安的技术,那里的实验室比 DeepSeek 多得多。所以,DeepSeek 在中国引发了一场运动。我说有点像 ChatGPT 在美国引发了一场运动,一切都有了聊天机器人。现在中国有大量科技公司发布非常强大的前沿开放权重模型,以至于我会说 DeepSeek 正在失去其作为中国卓越开放模型制造商的桂冠,而像 Z.ai 的 GLM 模型、MiniMax 的模型、Kimmy Moonshot,尤其是在过去几个月里,表现得更亮眼。新的 DeepSeek 模型仍然非常强大,但这可能被视为一个重要的叙事点,即 2025 年 DeepSeek 出现,然后为更多发布这些出色模型的中国公司提供了一个平台,让他们有了这种新型运营方式。所以,这些中国公司的模型是开放权重的,根据美国公司正在做的商业模式轨迹,这可能面临风险,但目前在美国很多人为 AI 软件付费,而历史上在中国和世界其他地区,人们不为软件支付很多钱。

You see the labs put different energy into what they're trying to do. And I think to demarcate the point in time when we're recording this, the hype over Anthropic's Claude Opus 4.5 model has been absolutely insane, which is just I mean, I've used it and built stuff in the last few weeks, and it's almost gotten to the point where it feels like a bit of a meme in terms of the hype, and it's kind of funny because this is very organic. And then, if we go back a few months ago, we can get the release date in the notes as Gemini 3 from Google got released, and it seemed like the marketing and just like wow factor of that release was super high. But then, at the end of November, Claude Opus 4.5 was released, and the hype has been growing. But Gemini 3 was before this, and it kind of feels like people don't really talk about it as much, even though when it came out, everybody was like, 'This is Gemini's moment to retake kind of Google's structural advantages in AI.' And Gemini 3 is a fantastic model, and I still use it. It's just kind of differentiation is lower. And I agree with Sebastian what you're saying with all of these like the idea space is very fluid, but culturally, Anthropic is known for betting very hard on code, which is Claude code thing is working out for them right now. So, I think that even if the ideas flow pretty freely, so much of this is bottlenecked by human effort and kind culture of organizations where Anthropic seems to at least be presenting as the least chaotic. It's a bit of an advantage, and if they can keep doing that for a while, but on the other side of things, there's a lot of ominous technology from China where there's way many more labs than DeepSeek. So, DeepSeek kicked off a movement within China. I say kind of similar to how ChatGPT kicked off a movement in the US where everything had a chatbot. There's now tons of tech companies in China that are releasing very strong frontier open weight models to the point where I would say that DeepSeek is kind of losing its crown as the preeminent open model maker in China and the likes of Z.ai's with their GLM models, MiniMax's models, Kimmy Moonshot, especially in the last few months have shown more brightly. The new DeepSeek models are still very strong, but that's kind of a it could look back as a big narrative point where in 2025 DeepSeek came and then all and it kind of provided this platform for way more Chinese companies that are releasing these fantastic models to kind of have this new type of operation. So, these models from these Chinese companies are open weights and depending on this trajectory of business models that these American companies are doing could be at risk, but currently a lot of people are paying for AI software in the US and historically in China and other parts of the world people don't pay a lot for software.

中国开源模型的可持续性 Sustainability of Open-Weight Models from China

Host

那么,像 DeepSeek 这样的模型因为开放权重而受到人们的喜爱。你认为中国公司会继续发布开放权重模型多久?

So, some of these models like DeepSeek have the love of the people because they are open weight. How long do you think the Chinese companies keep releasing open weights models?

Nathan Lambert

我会说几年。我认为在美国没有明确的商业模式。我写关于开放模型已经有一段时间了,这些中国公司已经意识到了这一点,所以我收到了一些他们的咨询。

I would say for a few years. I think that like in the US there's not a clear business model for it. I have been writing about open models for a while and these Chinese companies have realized it, so I get inbound from some of them.

中国开源模型与市场激励 Chinese open-weight models and market incentives

Nathan Lambert

他们很聪明,也意识到同样的限制,即许多美国科技公司和其他 IT 公司出于安全考虑不会为中国的 API 订阅付费。这是科技界长期以来的习惯,这些公司的人将开放权重模型视为影响和参与美国巨大且不断增长的 AI 支出市场的一种方式,他们对此非常现实。这对他们有效,我认为政府会看到这在国际上建立了很大的技术采用影响力。所以,会有很多激励措施来维持下去,但构建这些模型和进行研究非常昂贵,所以某个时候我预计会出现整合,但我不认为 2026 年会出现这样的故事:2026 年全年会有比 2025 年更多的开放模型构建者,而且很多知名的将在中国。

And they're smart and realize the same constraints, which is that a lot of US tech companies and other IT companies won't pay for an API subscription to Chinese companies for security concerns. This has been a long-standing habit in tech and the people at these companies then see open weight models as an ability to influence and take part of a huge growing AI expenditure market in the US and they're very realistic about this. And it's working for them and I think that the government will see that that is building a lot of influence internationally in terms of uptake of the technology. So, there's going to be a lot of incentives to keep it going, but building these models and doing the research is very expensive, so at some point I expect consolidation, but I don't expect that to be a story of 2026 where there will be more open model builders throughout 2026 than there were in 2025 and a lot of the notable ones will be in China.

DeepSeek 的定位与竞争 DeepSeek's position and competition

Host

你刚才想说什么?

You were going to say something?

Nathan Lambert

是的,你提到 DeepSeek 失去了王冠。我确实认为在某种程度上是这样,但我们也必须考虑到,他们仍然,我会说,稍微领先于其他公司。并不是 DeepSeek 变差了,只是其他公司正在使用 DeepSeek 的想法。例如,你提到了 Kimi。同样的架构,他们正在训练它,然后我们又有了这种跳跃式发展,他们可能在某个时间点更好一些,因为他们有更新的模型。我认为这又回到了一个事实:不会有明确的赢家。就会是这样。一个人发布了什么东西,另一个人就进来了,最新的模型可能总是最好的模型。

Yes, you mentioned DeepSeek losing its crown. I do think to some extent yes, but we also have to consider though they are still, I would say, slightly ahead of the other ones. It's not that DeepSeek got worse, it's just that the other ones are using the ideas from DeepSeek. For example, you mentioned Kimi. Same architecture, they're training it and then again we have this leapfrogging where they might be at some point in time a bit better because they have the more recent model. And I think this comes back to the fact that there won't be a clear winner. It will be just like that. One person releases something, the other one comes in and the most recent model is probably always the best model.

中国企业的不同激励 Different incentives among Chinese companies

Host

是的,我们也会看到中国公司有不同的动机。比如 DeepSeek 非常神秘,而一些初创公司,比如 MiniMax 和 Zidong AI,这两家实际上已经提交了 IPO 文件,他们试图获得西方市场的关注,并在那里做了很多推广。所以我不知道这些动机是否会改变模型开发,因为 DeepSeek 众所周知是由对冲基金 HighFlyer Capital 建立的,我们不知道他们到底用这些模型做什么,或者他们是否在意这个。

Yeah, we'll also see that Chinese companies have different incentives. So like DeepSeek is very secretive where some of these startups are like the MiniMaxes and the Zidong AIs of the world. Those two literally have filed IPO paperwork and they're trying to get Western mind share and do a lot of outreach there. So I don't know if these incentives will kind of change the model development because DeepSeek famously is built by a hedge fund, HighFlyer Capital, and we don't know exactly what they use the models for or if they care about this.

Nathan Lambert

他们在沟通方面很神秘,但在描述模型工作原理的技术报告方面并不神秘。他们在这方面仍然是开放的。我们还应该说说关于 Opus 4.5 的炒作,有一层是某物成为 Twitter 回音室上的宠儿,而实际使用该模型的人数。我认为可以公平地说,ChatGPT 和 Gemini 专注于只想解决日常生活中问题的广泛用户群,而这个用户群是巨大的。所以关于编码的炒作可能不会体现在实际使用中。

They're secretive in terms of communication, they're not secretive in terms of the technical reports that describe how their models work. They're still open on that front. And we should also say on the Opus 4.5 hype there's the layer of something being the darling of the X echo chamber on Twitter echo chamber and the actual amount of people that are using the model. I think it's probably fair to say that ChatGPT and Gemini are focused on the broad user base that just want to solve problems in their daily lives and that user base is gigantic. So the hype about the coding may not be represented in the actual use.

用户习惯与多订阅 User habits and multiple subscriptions

Host

我还会说,很多使用模式,就像你说的,是名字识别、品牌之类的,但还有肌肉记忆,你知道,ChatGPT 已经存在很长时间了。人们只是习惯了使用它,这有点像飞轮效应。他们向其他用户推荐它等等。一个有趣的点是 LLM 的定制化。例如,ChatGPT 有记忆功能,对吧?所以,你可能有一个订阅,你用它做个人事情,但我不确定你是否想在工作中使用同一个,你知道,因为这是私人和工作之间的界限。如果你在一家公司工作,他们可能不允许,或者你可能不想那样。我认为这也是一个有趣的点,你可能会有多个订阅。一个是干净的代码。它里面没有你的个人图片或爱好项目。它就像工作用的东西。另一个是你个人的东西。所以,我认为这也是两个不同的用例,并不意味着你只能有一个。我认为未来也是多个的。

I would say also, a lot of the usage patterns are, like you said, name recognition, brand, and stuff, but also muscle memory almost where, you know, like ChatGPT has been around for a long time. People just got used to using it and it's kind of like almost like a flywheel. They recommend it to other users and that stuff. The one interesting point is also the customization of LLMs. For example, ChatGPT has a memory feature, right? And so, you may have a subscription and you use it for personal stuff, but I don't know if you want to use that same thing at work, you know, because it's a boundary between private and work. If you're working at a company, they might not allow that or you may not want that. And I think that's also an interesting point where you might have multiple subscriptions. One is just clean code. It has nothing of your personal images or hobby projects in there. It's just like the work thing. And then the other one is your personal thing. So, I think that's also something where two different use cases and it doesn't mean you only have to have one. I think the future is also multiple ones.

2025-2026 赢家预测 Predictions for 2025 and 2026 winners

Host

你认为哪个模型会赢得 2025 年?你认为哪个模型会赢得 2026 年?

What model do you think will win 2025? And what model do you think is going to win 2026?

Nathan Lambert

我认为在消费者聊天机器人的背景下,问题是你是否愿意押注 Gemini 而不是 ChatGPT,我直觉上觉得这有点冒险,因为 OpenAI 是现有玩家,在科技领域有很多优势。我认为 2025 年的势头似乎在 Gemini 一边,但他们是从一个非常低的起点开始的。我的意思是,RIP Bard 和这些早期的尝试。我认为他们克服组织混乱实现这一目标值得高度赞扬。但同时,很难押注 OpenAI 失败,因为他们总是显得很混乱,但非常擅长落地。我个人对 GPT-5 的评价褒贬不一,但它肯定为他们节省了很多钱,因为其主打功能是一个路由器,大多数用户不再像以前那样支付 GPU 成本。所以,我认为很难将我喜欢模型的地方与实际能成为大众差异化因素的东西分开。

I think in the context of consumer chatbots, it's a question of are you willing to bet on Gemini over ChatGPT, which I would say in my gut feels like a bit of a risky bet because OpenAI has been the incumbent and there's so many benefits to that in tech. I think the momentum it feels like in 2025 was on Gemini's side, but they were starting from such a low point. I mean, RIP Bard and these earlier attempts of getting started. I think huge credit for them for powering through the organizational chaos to make that happen. But also, it's hard to bet against OpenAI because they always come off as so chaotic, but they're very good at landing things. And I think like personally, I have very mixed reviews of GPT-5, but it had to have saved them so much money with the headline feature being a router where most users are no longer charging like charging their GPU costs as much. So, I think it's very hard to dissociate the things that I like out of models versus the things that are going to actually be a general public differentiator.

Host

你对 2026 年怎么看?谁会赢?

What do you think about 2026? Who's going to win?

Nathan Lambert

我会说一些话,尽管有风险。我会说我认为 Gemini 将继续在 ChatGPT 上取得进展。我认为当两者都在如此极端的规模上运营时,Google 的规模优势,而且 Google 有能力将研究和产品更好地分开,而你听到很多关于 OpenAI 运营混乱、追逐高影响力的事情,这是一种非常初创公司的文化。然后在软件和企业方面,我认为 Anthropic 将继续取得成功,因为他们一次又一次地被设定为这样。显然 Google Cloud 有很多产品,但我认为这种 Gemini 品牌对他们来说很重要。Google Cloud 将继续表现良好,但这在生态系统中解释起来更复杂,因为它与 Azure 和 AWS 等竞争,而不是在模型提供商方面。

I'll say something even though it's risky. I will say that I think Gemini will continue to take progress on ChatGPT. I think Google scale when both of these are operating at such extreme scales and like Google has the ability to separate that research and product a bit better where you hear so much about OpenAI being chaotic operationally and chasing the high-impact thing, which is a very startup culture. And then on the software and enterprise side, I think Anthropic will have continued success as they've again and again been set up for that. And obviously Google's Cloud has a lot of offerings, but I think this kind of like Gemini name brand is important for them to build. And Google's Cloud will continue to do well, but that's kind of a more complex thing to explain in the ecosystem because that's competing with the likes of Azure and AWS rather than on the model provider side.

Host

那么,在基础设施方面,你认为 TPU 提供了优势吗?

So, in the infrastructure, you think TPUs give an advantage?

Nathan Lambert

主要是因为 Nvidia 芯片的利润率高得离谱,而 Google 可以从上到下开发一切以适应他们的技术栈,不必支付这个利润率,而且他们在建设数据中心方面有先发优势。所以,所有这些既有高前置时间又有高成本下非常艰难利润率的事情,Google 在那里有一种历史优势。如果会出现新的范式,最有可能来自 OpenAI,他们的研究部门一次又一次地展示了落地新研究想法或产品的能力。

Largely because the margin on Nvidia chips is insane and Google can develop everything from top to bottom to fit their stack and not have to pay this margin and they've had a head start in building data centers. So, all of these things that have both high lead times and very hard margins on high costs, Google has a kind of historical advantage there. And if there's going to be a new paradigm, it's most likely to come from OpenAI where their research division again and again has shown this ability to land a new research idea or a product.

模型使用偏好 Model usage preferences

Host

我认为像深度研究、Sora、01 思维模型这些定义性的东西都来自 OpenAI,这一定是他们作为组织的顶级特质之一。所以,很难不看好他们,但我认为今年很多工作将围绕规模扩张和优化模型中所谓的低垂果实展开。

I think like deep research, Sora, 01 thinking models, like all these definitional things have come from OpenAI and that's got to be one of their top traits as an organization. So, it's kind of hard to bet against that, but I think a lot of this year will be about scale and optimizing what could be described as low-hanging fruit in models.

Nathan Lambert

显然,智能和速度之间存在权衡。这正是 ChatGPT 5 在幕后试图解决的问题。大众到底想要智能,还是想要速度?

And clearly, there's a trade-off between intelligence and speed. This is what ChatGPT 5 was trying to solve behind the scenes. It's like, do people actually want intelligence, the broad public, or do they want speed?

Host

实际上,我觉得有这种多样性很好,或者有一个切换选项。首先,就我个人使用而言,大多数时候我查东西时,会用 ChatGPT 快速提问,快速获取信息。对于日常任务,我用快速模型。现在,我觉得自动模式很不错,你不需要特别指定思考或不思考。但有时我也想要专业模式。我经常做的是,当我有写好的东西时,我会把它放进 ChatGPT 说:'嘿,做个彻底检查。我的所有引用都正确吗?我的所有想法都正确吗?我有没有格式错误?图号有没有错?' 我不需要立即得到结果。我完成工作后,可能去吃个晚饭,让它运行,回来再检查。我认为这就是拥有这个选项重要的地方。如果每次查询都要等 30 分钟甚至 10 分钟,我会疯掉的。

I think it's a nice variety, actually, or the option to have a toggle there. I mean, first for my personal usage, most of the time when I look something up, I use ChatGPT to ask a quick question, get the information I want fast. For, you know, most daily tasks, I use the quick model. Nowadays, I think the auto mode is pretty good where you don't have to specifically say thinking or, you know, non-thinking and stuff. Then again, I also sometimes want the pro mode. Very often, what I do is when I have something written, I put it into a ChatGPT and say, 'Hey, do a very thorough check. Are all my references correct? Are all my thoughts correct? Did I make any formatting mistakes? And are the figure numbers wrong or something like that?' And I don't need that right away. It's something, okay, I finish my stuff, maybe have dinner, let it run, come back, and go through this. And I think see, this is where I think it's important to have this option. I would go crazy if for each query I would have to wait 30 minutes or 10 minutes even.

Nathan Lambert

这就是我。是的。我在这里说,我无法理解你使用路由器和非思维模型。我想:'你怎么受得了?' 是的,这是我的反应。我一度非常依赖 ChatGPT。从未碰过五号非思维模型。我发现它的语气和错误倾向,就是出错的可能性更高。这有些源于 OpenAI 发布 03 的时候,那是第一个做深度搜索、找到许多来源并为你整合的模型。所以我习惯了那样,所以对于任何工作相关的信息查询,无论是论文还是代码参考,我只用 GPT 5.2 思维或专业版。我经常同时运行五个专业查询,每个查找一篇特定论文或对某个方程的反馈。

That's me. Yeah. I'm like saying over here I'm losing my mind that you use the router and the non-thinking model. I'm like, 'How are you? How do you live with that?' Yeah, that's my reaction. I'm bent heavily on ChatGPT for a while. Never touched five non-thinking. I find its tone and its propensity of errors, it's just like it has a higher likelihood of errors. Some of this is from back when OpenAI released 03, which was the first model to do this deep search and find many sources and integrate them for you. So, I became habituated with that, so I will only use GPT 5.2 thinking or pro when I'm finding any sort of information query for work, whether that's a paper or some code reference that I found. And it's just like I will regularly have like five Pro queries going simultaneously, each looking for one specific paper or feedback on an equation or something.

Host

我有一个有趣的例子,之前为了这个播客我需要尽快回答。我要去旅行。家里有一台本地 GPU 在运行,我想跑一个长时间的强化学习实验。通常我也会拔掉插头,因为不在家时你不想让东西插着。我不小心拔掉了 GPU 的插头。当时我妻子已经在车里了,我说:'哦,糟了。' 然后我急需一个 bash 脚本来运行不同的实验和评估。我知道我学过如何使用 bash 终端,但那一刻我只需要 10 秒内得到命令。这真是个滑稽的情况。

I have a funny example of where I just needed to answer as fast as possible for this podcast before. I was going on a trip. I have a local GPU running at home and I wanted to run a long RL experiment. And usually I also unplug things because you never know if you're not at home you don't want to have things plugged in. And I accidentally unplugged the GPU. It was like my wife was already in the car and I was like, 'Oh dang.' And then basically I wanted as fast as possible a bash script that runs my different experiments and the evaluation. And I did something I know I learned how to use bash interface bash terminal, but in that moment I just needed like 10 seconds give me the command. It's a hilarious situation.

Nathan Lambert

是啊,那你用了什么?

Yeah, so what did you use?

Host

所以我用了非思维的最快模型。它给了我 bash 命令来串联不同的脚本,然后还有那个 T 符号,你想把输出重定向到日志文件。我当时脑子里很急。我本可以自己想的。

So I did the non-thinking fastest model. It gave me the bash command to chain different scripts to each other and then the thing is like you have that T thing where you want to route this to a log file. Top of my head I was just like in a hurry. I could have thought about it myself.

Nathan Lambert

顺便说一句,我不知道这是否有代表性——妻子在车里等着,你不得不跑,拔掉了 GPU,还得生成一个 bash 脚本。这听起来像《碟中谍》电影。

By the way, I don't know if there's a representative case wife waiting in the car you have to run you unplug the GPU. You have to generate a bash script. This sounds like a movie like Mission Impossible.

Host

我用 Gemini 做那个。所以对于所有信息类的事情我用思维模型,对于快速的事情或者有时可以用谷歌搜索的事情我用 Gemini,它擅长解释事情,我相信它有这种背景知识,而且简单。Gemini 应用已经好多了,适合这类事情。对于代码和任何哲学讨论,我用 Claude Opus 4.5,也总是开启扩展思维。扩展思维和推理时 Scaling 只是让模型稍微更聪明一点的方法,当进展很快时我总是倾向于那一边,因为你不知道什么时候会解锁新的用例。有时我用 Grok 获取实时信息,或者在 AI Twitter 上找我知道见过但需要挖掘的东西,我特别执着。不过当 Grok 4 出来时,那个超级重的版本,也就是他们的专业版,实际上非常好,我印象很深。但后来肌肉记忆让我忘了它,因为 ChatGPT 应用一直开着。所以,我只是想谢谢你。谢谢。

I use Gemini for that. So I use thinking for all the information stuff and then Gemini for fast things or stuff that I could sometimes Google which is like it's good at explaining things and I trust that it has this kind of background of knowledge and it's simple. And the Gemini app has got a lot better and it's good for that sort of things. And then for code and any sort of philosophical discussion I use Claude Opus 4.5. Also always with extended thinking. Extended thinking and inference time scaling is just a way to make the models marginally smarter and I will always edge on that side when the progress is very high because you don't know when that will unlock a new use case. And then sometimes use Grok for real-time information or finding something on AI Twitter that I knew I saw and I need to dig up and I'm just fixated on. Although when Grok 4 came out, the Grok 4, what is super heavy, which was like their pro variant, was actually very good and I was pretty impressed with it. And then I just kind of like muscle memory lost track of it with having the ChatGPT app open. So, I just wanted to thank you. Thanks.

Nathan Lambert

是的。我实际上用 Grok 4 重型版做调试。对于其他模型解决不了的硬核调试,我发现它最擅长。有趣的是,你说 ChatGPT 是最好的界面。对我来说,出于同样的原因,但这可能只是惯性,Gemini 是更好的界面。我想是因为我迷上了它们的大海捞针能力。如果我放进去很多上下文,但寻找非常具体的信息,确保它跟踪所有内容。我发现至少 Gemini 对我来说是最好的。所以,这些模型很有趣,如果它们在某个特定日子、针对某个特定查询或提示赢得了你的心,你就会觉得这个模型更好。然后你会坚持用一段时间,直到它做了非常蠢的事。有一个阈值效应。它做了聪明的事,你爱上它,然后它做了蠢事,你就想:你知道吗?我要换 Claude 和 ChatGPT 试试。

Yeah. I actually do use Grok 4 heavy for debugging. For like hardcore debugging that the other ones can't solve. I find that it's the best at. And it's interesting cuz you said ChatGPT is the best interface. For me, for that same reason, but this could be just momentum, Gemini is the better interface for me. I think because I fell in love with their best needle in the haystack. If I ever put something that has a lot of context, but I'm looking for a very specific kinds of information, make sure it tracks all of it. I find at least that Gemini for me has been the best. So, it's funny with some of these models, if they win your heart over for one particular feature at one on a one particular day, for that particular query, that prompt, you're like, this model is better. And so, you'll just stick with it for a bit until it does something really dumb. There's like a threshold effect. Some smart thing and then you fall in love with it and then it does some dumb thing and you're like, you know what? I'm going to switch and try Claude and ChatGPT and all that kind of stuff.

Host

这完全就是用到它出问题为止,直到你遇到问题,然后你换语言模型。我认为这和我们使用任何东西一样,比如最喜欢的文本编辑器、操作系统或浏览器。我是说,浏览器有很多选择,Safari、Firefox、Chrome,都相对相似,但总有一些边缘情况,比如你想用的扩展,然后你就换了。但我不认为有人会把同一个网址输入不同浏览器然后比较。只有当网站渲染不出来,或者出问题时才会那样做。所以,这是个好观点。我认为你用到它出问题,然后探索其他选择。

This is exactly like you use it until it breaks, until you have a problem, and then you change the LM. And I think it's the same how we use anything, like our favorite text editor, operating systems, or the browser. I mean, there are so many browser options, Safari, Firefox, Chrome, all the relatively similar, but then there are edge cases, maybe extensions you want to use, and you switch. But, I don't think there is any one who types the same thing like the website into different browsers and compares them. You only do that when the website doesn't render, if something breaks, I think. So, that's a good point. I think you use it until it breaks, and then you explore other options, I think.

Nathan Lambert

关于长上下文,我也是 Gemini 用户。

On the long context thing, I was also a Gemini user for this.

GPT-5.2 长上下文与中国模型 GPT-5.2 long context and Chinese models

Nathan Lambert

但 GPT-5.2 的发布博客上出现了惊人的长上下文分数,很多人都在想,‘他们是不是刚刚搞定了某种算法上的改变?’在这个小版本更新中,它从大约 30% 跳到了 70% 左右。所以,要跟踪所有这些事情也非常困难。不过现在我对 GPT-5.2 的长上下文有了更好的看法。问题就在于,我到底要怎么实际测试这个?这是一场永无止境的战斗。

But the GPT-5.2 release blog had crazy long context scores, where a lot of people were like, 'Did they just figure out some algorithmic change?' It went from like 30% to like 70% or something in this minor model update. So it's also very hard to keep track of all of these things. But now I look more favorably at GPT-5.2's long context. So it's just kind of like how do I actually get to testing this? It's an never-ending battle.

Host

嗯,有趣的是,我们谁都没有从用户使用的角度谈论中国模型。这说明什么?是中国模型不够好,还是我们只是有偏见、只关注美国?

Well, it's interesting that none of us talked about the Chinese models from a user usage perspective. What does that say? Does that mean the Chinese models are not as good, or does that mean we're just very biased and US-focused?

Nathan Lambert

我确实认为,这目前只是模型和平台之间的差异。所以我认为,开源模型更出名的是它们的开放权重,而不是它们的平台。

I do think that that's currently the discrepancy between just the model and the platform. So I think the open models, they are more known for the open weights, not their platform yet.

Host

还有很多公司愿意以非常低的成本向你出售开源模型的推理服务。我觉得像 OpenRouter 这样的平台,很容易进行多模型对比。你可以在 Perplexity 上运行 DeepSeek。我认为在座的各位都是这样:‘我们一直使用 OpenAI GPT-5 Pro。我们都愿意为那一点智能增益付费。’而且,那些认为美国模型在输出方面更好的人。我认为问题是,它们今年和未来几年还会继续保持优势吗?但只要它们更好,我就愿意付费使用。我认为也有分析显示,中国模型的部署方式——你可以争论是否由于出口管制——它们使用更少的 GPU 来提供服务,这使得它们更慢,并产生不同的错误。这关乎速度和智能。如果这些因素对你有利,我认为美国很多用户会选择它们,而这将促使中国公司以其他方式竞争,比如免费或大幅降低成本,或者催生创新的产品,这对生态系统有好处。但我认为简单的事实是,美国模型目前更好,我们也在使用它们。我试过中国模型——我试过其他开源模型,感觉有趣,但不会回去用。

There are also a lot of companies that are willing to sell you the open model inference at a very low cost. I think like OpenRouter, it's easy to do the look at multi-model things. You could run DeepSeek on Perplexity. I think all of us sitting here are like, 'We use OpenAI GPT-5 Pro very consistently. We're all willing to pay for the marginal intelligence gain.' And anyone that's like the these models from the US are better in terms of the outputs. I think that the question is, will they stay better for this year and for years going? But it's like so long as they're better, I'm going to pay for it to use them. I think there's also analysis that shows that like the way that the Chinese models are served, this you could argue due to export controls or not, is that they use fewer GPUs for replica, which makes them slower and have different errors. And it's like the speed and intelligence. If these things are in your favor as a user, I think in the US a lot of users will go for this, and I think that that is going to be a thing that will spur these Chinese companies to want to compete in other ways, whether it's like sub-free or substantially lower costs, or it'll breed creativity in terms of offerings, which is good for the ecosystem. But I just think the simple thing is the US models are currently better, and we use them. And I tried Chinese models—I tried these other open models, and I'm like, fun, but not going to—I don't go back to it.

用 LLM 编程:工具与经验 Programming with LLMs: tools and experiences

Host

呃,我们还没怎么提到编程。这是很多人非常关心的另一个用例。我基本上是一半时间用 Cursor,一半时间用 Claude Code,因为我觉得它们提供了根本不同的体验,而且都很有用。你们呢——你们编程不少,你们用什么?目前的感觉如何?

Uh, we didn't really mention programming. That's another use case that a lot of people deeply care about. So I use basically half and half Cursor and Claude Code, because I find them to be like fundamentally different experiences and both useful. What do you guys—you program quite a bit, so what do you use? What's the current vibe?

Nathan Lambert

所以我用的是 VS Code 的 Codex 插件。你知道,它非常方便。它就是一个插件,然后是一个可以访问你仓库的聊天界面。我知道 Claude Code 有点不同。它更具智能体特性。它涉及更多东西。它可以为你完成整个项目。我还没到能放心使用它的地步,因为也许我是个控制狂,但我还是想看看发生了什么。而 Codex 目前对我来说是个甜点,它在帮助我,但没有完全接管。

So I use the Codex plugin for VS Code. You know, it's very convenient. It's just like a plugin, and then it's a chat interface that has access to your repository. I know that Claude Code is I think a bit different. It is a bit more agentic. It touches more things. It does a whole project for you. I'm not quite there yet where I'm comfortable with that, because maybe I'm a control freak, but I still would like to see a bit what's going on. And Codex is kind of like right now for me like the sweet spot where it is helping me, but it is not taking completely over.

Host

我应该提一下,我使用 Claude Code 的原因之一是为了培养用英语编程的技能。我的意思是,体验完全不同。你不是微观管理代码生成过程的细节,也不是查看差异——如果你用 Cursor 的话可以这么做——然后随着进展修改、调整、查看和阅读代码,深入理解代码;而是像在设计空间中思考,在宏观层面引导它,我认为这是思考编程过程的另一种方式。另外,我们应该说,Claude Code 似乎 somehow 更好地利用了 Claude Opus 4.5。

I should mention one of the reasons I do use Claude Code is to build the skill of programming with English. I mean, the experience is fundamentally different. You're as opposed to micromanaging the details of the process of the generation of the code and looking at the diff, which you can in Cursor, if that's the IDE you use, and then changing, altering, looking and reading the code and understanding the code deeply as you progress, versus just kind of like thinking in this design space, and just guiding it at this macro level which I think is another way of thinking about the programming process. Also, we should say that Claude Code it just seems to be somehow a better utilization of Claude Opus 4.5.

Nathan Lambert

这对人们来说是一个很好的并行实践。你可以同时打开 Claude Code、Cursor 和 VS Code,然后在所有工具上选择相同的模型并提问。它们非常有趣。比如 Claude Code 在那个领域表现更好。这很了不起。

It's a good side by side for people to do. So you can have Claude Code open, you can have Cursor open, and you can have VS Code open, and you can select the same models on all of them and ask questions. They're very interesting. Like with Claude Code, it works better in that domain. It's remarkable.

书籍与从零学习 Books and learning from scratch

Host

好了,我们应该说,你们两位在多个方面都很厉害:研究员、程序员、教育者、推文作者。在写书方面也是。所以 Nathan 很快,希望如此,会有一本关于 RLHF 的书出版。

All right, we should say that both of you are legit on multiple fronts: researchers, programmers, educators, tweeters. And on the book front, too. So Nathan at some point soon, hopefully, has an RLHF book coming out.

Nathan Lambert

它已经可以预购了,还有完整的数字预印本。我只是在把它做得更漂亮、组织得更好,以便出版实体书。我这么做很大程度上是因为,在我们生活如此数字化的时代,创造出你认为优秀的实体形式的东西很有趣。

It's available for pre-order, and there's a full digital preprint. I'm just making it pretty and better organized for the physical thing, which is a lot of why I do it because it's fun to create things that you think are excellent in the physical form when so much of our life is digital.

Host

我应该提一下,在 Perplexity 上,Sebastian Raschka 是一位机器学习研究员和作者,以几本有影响力的书而闻名。我想提其中两本,一本是我强烈推荐的《从零开始构建大型语言模型》,还有新书《从零开始构建推理模型》。所以我对此非常兴奋。从零开始构建东西是最强大的学习方式之一。

I should say, going to Perplexity here, Sebastian Raschka is a machine learning researcher and author known for several influential books. A couple of them that I wanted to mention, which is a book I highly recommend, 'Build a Large Language Model from Scratch', and the new one, 'Build a Reasoning Model from Scratch'. So I'm really excited about that. Building stuff from scratch is one of the most powerful ways of learning.

Nathan Lambert

说实话,从零开始构建 LLM 非常有趣。要学的东西也很多,就像你说的,这可能是了解某样东西真正工作原理的最好方法,因为你可以看图表,但图表可能有错误。你可以看概念、解释,但你可能误解它们。但如果你看到代码,有代码,而且代码能运行,你就知道它是正确的。我的意思是,没有误解。它很精确。否则它就不会工作。我认为这就是编码背后的美妙之处。它不会撒谎。它基本上就是数学,所以即使数学方面,我认为书中也可能有你永远不会注意到的错误,因为你在读书时不会运行数学。你无法验证它,而代码的好处是你可以验证它。

Honestly, building an LLM from scratch is a lot of fun. It's also a lot to learn, and like you said, it's probably the best way to learn how something really works because you can look at figures, but figures can have mistakes. You can look at concepts, explanations, but you might misunderstand them. But if you see the code, there is code, and the code works, you know it's correct. I mean, there's no misunderstanding. It's like it's precise. Otherwise, it wouldn't work. And I think that's like kind of like the beauty behind coding. It is kind of like it doesn't lie. It's math, basically, so even though with math I think you can have mistakes in a book you would never notice because you're not running the math when you're reading the book. You can't verify this and with code what's nice is you can verify it.

Host

是的,我同意你对那本从零开始构建 LLM 的书的看法。排除其他一切干扰,比如互联网等等,只专注于这本书,感觉很好。但是,你知道,我读了几本,比如历史书。不知怎么的,感觉不那么孤独了。真的更有趣。比如在编程方面,我认为用 LLM 编程真的更有趣。而且我认为用 LLM 阅读也真的更有趣。但你说得对,应该尽量减少干扰。所以你要用 LLM 来丰富体验,也许增加更多上下文。对我来说,使用 LLM 时,小范围内获得顿悟时刻的频率真的很高。

Yeah, I agree with you about the LLM from scratch book. It's nice to tune out everything else, the internet and so on, and just focus on the book. But, you know, I read several like, you know, history books. It's just less lonely somehow. It's really more fun. Like for example on the programming front, I think it's genuinely more fun to program with an LLM. And I think it's genuinely more fun to read with an LLM. But you're right, like distractions should be minimized. So it's you use the LLM to basically enrich the experience, maybe add more context. Maybe the I just the rate of aha moments for me in a small scale is really high with LLMs.

Nathan Lambert

100%。我想说,我也想纠正一下自己。我不是建议不要使用 LLM。我建议分多次进行。

100%. I would say I also want to correct myself. I'm not suggesting not to use LLMs. I suggest doing it in multiple passes.

用 LLM 的阅读习惯 Reading habits with LLMs

Nathan Lambert

比如先离线专注读一遍,之后我也会记笔记,但我会克制自己不要马上去查东西。我会读第二遍。对我来说这样更有条理,而且有时候章节里就有答案,但有时候也需要让内容沉淀一下,思考思考。其他人有不同的偏好。我强烈推荐在读书时使用大语言模型(LLM)。对我来说,这不是第一步要做的事,而是第二遍才做的事。

Like one pass just offline focus mode and then after that I mean I also take notes, but I try to resist the urge to immediately look things up. I do a second pass. It's just like for me more structured this way and I get I mean sometimes things are answered in the chapter, but sometimes also it just helps to let it sink in and think about it. Other people have different preferences. I would highly recommend using LLMs when reading books. For me it's just it's not the first thing to do. It's like the second pass.

Host

说到推荐,我得说我做法相反。我喜欢一开始就用 LLM 来了解全局背景,比如我要进入的是一个什么样的世界。但我尽量不点出 LLM,进入 Twitter 博客的世界,因为那样你就会掉进兔子洞,读别人的观点,看到某个话题的骂战,然后突然你就进入了互联网和 Reddit 的领域。但如果你纯粹让 LLM 给你提供背景,告诉你为什么这很重要,有哪些宏观概念——有时候书本身也能做到,但并非总是如此。

By way of recommendation I should say I do the opposite. I like to use the LLM at the beginning to lay out the full context of like what is this world that I'm now stepping into. But I try to avoid clicking out of the LLM into the world of like Twitter blogs and because then you're now down this rabbit hole, you're reading somebody's opinion, there's a flame war about a particular topic, and all of a sudden you're no longer you're now in the in the realm of the internet and Reddit and so on. But if you're purely letting the LLM give you the context of why this matters, what are the big picture ideas. Uh but sometimes books themselves are good at doing that, but not always. So.

Nathan Lambert

这就是为什么我喜欢 ChatGPT 应用,因为它给 AI 在你的电脑里安了个家,当你专注时你可以专注于它,而不是让它成为我海量浏览器标签中的又一个。我认为 Claude Code 这类工具做得很好,让这件事变得愉快,产品设计上很有吸引力,让你的 AI 能走出去与世界互动。它和 Codex 之间有种说不清的东西,就是感觉更温暖、更吸引人。而 OpenAI 的 Codex 虽然同样出色,但感觉有点粗糙。Claude Code 让从零开始构建东西变得有趣,你不需要操心,但信任它能做出东西。显然这对网站、刷新工具之类的东西很好,我也会用它做数据分析。比如我的博客,我们抓取 Hugging Face,持续记录每个数据集和模型的下载量。Claude 就说:“没问题,我已经用了那些数据。”我当时想:“这得花我几天时间。”然后我有了足够的场景意识,说:“好吧,这些趋势显然合理。”你还可以检查。这是一种很棒的界面,你可以有一个中间层,而不必做那些维护不同网络项目所需的低级工作。

That's why I like the ChatGPT app cuz it gives the AI a home in your computer when you are focused you can focus on it rather than just being another tab in my massive internet options. And I think Claude Code and these particular does a good job of making that a joy where it seems very engaging as a product design to be an interface that your AI will then go out into the world. And there's something that is very kind of intangible between it and Codex is that it just feels kind of warm and engaging. Where Codex can often be as good from Open AI, but it just kind of like feels a little bit rougher on the edges. Whereas like Claude Code is makes it fun to build things particularly from scratch where you just don't like you don't have to care but you trust that it'll make something. Like obviously this is good for websites and kind of refreshing tooling and stuff like this which I'd use it for or data analysis. So I my my blog we scrape Hugging Face we keep the download numbers for every data set and model over time now so we have them. And it's like Claude was just like, "Yeah, I've made use of that data, no problem." And I was like, "That would have taken me days." And it's like then I have enough situational awareness to be like, "Okay, these trends obviously make sense." And you can check things. Cuz that's just a kind of wonderful interface where you can have an intermediary and not have to do the kind of awful low-level work that you would have to do to maintain different web projects and do this stuff.

开源 LLM 格局 Open LLM models landscape

Host

好了,我们刚聊了一堆闭源权重模型。现在来谈谈开源模型。给我讲讲开源 LLM 模型的格局吧。哪些有意思?哪些让你印象深刻?为什么?我们已经提到了 DeepSeek。

All right, so we just talked about a bunch of the closed weight models. Let's talk about the open ones. Uh so tell me about the landscape of open LLM models. Which are interesting ones? Which stand out to you? And why? We already mentioned DeepSeek.

Nathan Lambert

你看我们能随口说出多少个。

Do you I see how many we can name off the top of our head.

Host

对对,不看笔记。

Yeah, yeah, without looking at notes.

Nathan Lambert

DeepSeek、Kimmy、MiniMax、Z.ai、Ant Ling,我们只说中国的。再加上 Mistral AI、Gemma,还有 GPT-OSS,ChatGPT 的开源模型。实际上,Nvidia NeMo Tron 有个很酷的,NeMo Tron 3。年底有很多东西。Qwen,可能还有一个……

DeepSeek, Kimmy, MiniMax, Z.ai, Ant Ling, we're just going Chinese. Let's throw in Mistral AI, Gemma, yeah, GPT-OSS, the open-source model by ChatGPT. Actually, Nvidia NeMo Tron had a or Nvidia had a really cool one, the NeMo Tron 3. There's a lot of stuff especially at the end of the year. Qwen, one maybe the one

Host

哦对,Qwen,这个名字很明显,是第一个想到的。

Oh, yeah, Qwen was the name the obvious name. It was the first thing.

Nathan Lambert

我试着数一下,你至少能列出 10 个中国的和 10 个西方的。我觉得,OpenAI 发布了自 GPT-2 以来的第一个开源模型。当时我写关于 OpenAI 开源模型发布的文章时,大家都说“别忘了 GPT-2”。我觉得很有趣,因为那完全是不同的时代。但 GPT-OSS 实际上是一个非常强大的模型,能做其他模型做不好的事情。自私地说,我会推广一些西方公司。所以,美国和欧洲都有这种完全开源的模型。我在艾伦人工智能研究所工作,我们一直在构建 Olmo,它开源了数据和代码等一切。现在,对于那些试图开源一切以便其他人可以训练这些模型的人来说,我们有了真正的竞争。所以,有基础模型研究所或/LM360,他们推出了各种类型的 K2 模型。EPFL 是一个瑞士研究联盟。Hugging Face 有 Small LM,非常流行。Nvidia 的 NeMo Tron 也开始开源数据。还有斯坦福的 MAIRIN 社区项目,它建立了一个流程,让人们可以开一个 GitHub issue,实现一个新想法,然后在稳定的语言建模栈中运行。所以,这个领域在 2024 年要小得多,当时好像只有 AI2。所以,让更多人参与理解语言模型是件好事,中国公司没有类似的东西。说到这,我要指出,中国的开源语言模型往往大得多,作为混合专家模型(MOE),它们的峰值性能更高,而我们喜欢的很多模型,无论是 Gemma 还是 D Metatron,往往来自美国,是较小的模型,这种情况正在从美国和欧洲开始改变。Mistral Large 3 在 12 月发布,是一个巨大的 MOE 模型,架构与 DeepSeek 非常相似。然后初创公司 RCAI,以及 Numenta 和 Nvidia 都预告了 MOE 模型,参数远超 1000 亿,达到 4000 亿级别,预计在 2026 年第一季度发布。所以,我认为今年人们使用中国和美国开源模型的平衡将发生变化,我个人非常期待看到这一点。

I was trying to get through the You can get at least 10 Chinese and at least 10 Western. I think that I mean, OpenAI released their first open model since GPT-2. That was when I when I meant talk when I was writing about OpenAI's open model release, they're all like, "Don't forget about GPT-2." Which I thought was really funny cuz it's just such a different time. But GPT-OSS is actually a very strong model and does some things that the other models don't do very well and I think that selfishly, I'll promote a bunch of like Western companies. So, both in the US and Europe have these like fully open models. So, I work at Allen Institute for AI. We've been building Olmo, which releases data and code and all of this. And now we have actual competition for people that are trying to release everything so that other people can train these models. So, there's the Institute for Foundation Models or / LM360, which is like had their K2 models of various types. EPFL is a Swiss research consortium. Hugging Face has small LM, which is very popular. Nvidia's NeMo Tron has started releasing data as well. And then Stanford's MAIRIN community project, which is kind of making it so there's a pipeline for people to open a GitHub issue and implement a new idea and then have it run in a stable language modeling stack. So, this space that list was way smaller in 2024. So, I think it was like just AI2. So, that's a great thing for more people to get involved in to understand language models, which doesn't really have a like a Chinese company that is has an analog. While I'm talking, I will say that the Chinese open language models tend to be much bigger and that gives them this higher peak performance as MOEs where a lot of these things that we like a lot, whether it was Gemma, um and D Metatron have tended to be smaller models from the US, which is which is starting to change from the US US and Europe. Um Mistral Large 3 came out, which was a giant MOE model, very similar to DeepSeek architecture in December. And then a startup RCAI and both Numenta and have Numenta and Nvidia have teased MOE models of this way bigger than 100 billion parameters, like this 400 billion parameter range coming in this like Q1 2026 timeline. So, I think this kind of balance is set to change this year in terms of what people are using the Chinese versus US open models for, which will be a which I'm personally expecting is going to be very excited to watch.

Host

首先,能说出这么多模型,太厉害了。你提到 Llama 了吗?

First of all, huge props for being able to name so many of these. Did you actually name Llama?

Nathan Lambert

呃,没有。

Um no.

Host

像……

like

Nathan Lambert

安息吧。

RIP

Host

我不是故意的。

That was not on purpose.

Nathan Lambert

Llama 安息。

RIP Llama.

Host

嗯。

Mhm.

Host

好了。你能说说哪些有趣的模型比较突出吗?你提到 Qwen 3 显然很突出。

All right. Can you mention what are some interesting models that stand out? So, you mentioned Qwen 3's is is obviously a standout.

Nathan Lambert

所以,我认为这一年几乎被 DeepSeek V3 和 R1 首尾包围。另一方面,12 月有 DeepSeek V3.2,因为我喜欢它们总是有一些其他模型没有的有趣架构调整。但除此之外,如果你想要熟悉但性能出色的模型,Qwen 3 和,呃,Nathan 说的,还有 GPT-OSS。

So, I would say the year's almost bookended by both DeepSeek version 3 and R1. And then on the other hand in December, DeepSeek version 3.2 because what I like about those is they always have an interesting architecture tweak that others don't have. But otherwise, if you want to go with um you know, like the familiar but really good performance, Qwen 3 and like um Nathan said, also GPT-OSS.

工具使用及其意义 Tool Use and Its Significance

Nathan Lambert

我认为 GPT-OSS 的有趣之处在于,它是第一个真正以工具使用为目标进行训练的公开或开放权重模型,这在我看来有点像一种范式转变,而生态系统还没有完全准备好。所谓工具使用,我指的是 LLM 能够进行网络搜索、调用 Python 解释器。我认为这很突出,因为它是一个巨大的解锁,因为 LLM 最常见的抱怨之一就是幻觉,对吧?所以在我看来,解决幻觉的最佳方法之一就是不要总是试图记住信息或编造东西。对于数学,为什么不用计算器应用或 Python 呢?

And I think GPT-OSS, what's interesting about it is kind of like the first public or open weight model that was really trained with tool use in mind, which I do think is kind of a little bit of a paradigm shift where the ecosystem was not quite ready for it. So, with tool use, I mean that the LLM is able to do a web search, to call a Python interpreter. And I do think this is a standout because I think it's a huge unlock because one of the most common complaints about LLMs are for example hallucinations, right? And so in my opinion one of the best ways to solve hallucinations is to not try to always remember information or make things up. For math, why not use a calculator app or Python?

Host

嗯。

Mhm.

Nathan Lambert

如果我问 LLM 谁赢得了 1998 年世界杯,它不用死记硬背,而是可以去搜索。我认为通常还是谷歌搜索。所以 ChatGPT、GPT-4 SS 会调用谷歌工具,也许找到 FIFA 网站,发现是法国队。它能可靠地给你信息,而不是试图记住。所以我认为这是一个巨大的解锁,但目前开源开放权重生态系统还没有充分利用它。很多人不使用工具调用模式,因为首先是个信任问题。你不想在电脑上运行它,因为它可以访问工具,可能会擦除你的硬盘之类的。所以你可能需要把它容器化。但我认为,拥有这种能力是未来几年非常重要的一步。

If I ask the LLM who won the soccer World Cup in 1998, instead of just trying to memorize, it could go do a search. I think mostly it's usually still a Google search. So ChatGPT, GPT-4 SS, they would do a tool call to Google, maybe find the FIFA website, find okay, it was France. It would get you that information reliably instead of just trying to memorize it. So I think it's a huge unlock which I think right now is not fully utilized yet by the open source open weight ecosystem. A lot of people don't use tool call modes because I think it's first it's a trust thing. You don't want to run this on your computer where it has access to tools, could wipe your hard drive or whatever. So you want to maybe have containerize that. But I do think, you know, that is like a really important step for the upcoming years to have this ability, yeah.

开源模型爆发原因 Reasons for Open Model Explosion

Host

快速说几点。首先,感谢你定义了工具使用的含义。我认为对我们讨论的概念来说,这通常是一件好事。即使是像 MOE 这样成熟的概念,你也得说清楚它代表混合专家模型,并让人们直观理解它的含义、实际应用方式以及不同变体。那么,开放模型如此爆发意味着什么?你的直觉是什么?

So a few quick things. First of all, thank you for defining what you mean by tool use. I think that's a great thing to do in general for the concepts we're talking about. Even things as sort of well established as MOEs, you have to say that means mixture of experts and you could kind of have to build up an intuition for people what that means, how it's actually utilized, what are the different flavors. So what does it mean that there's this such explosion of open models? What's your intuition?

Nathan Lambert

如果你发布一个开放模型,首要目标是让人们使用它。然后才是透明度和信任等问题。我认为,看看中国,最大的原因是他们希望世界各地的人们使用这些模型,而很多人不会。在美国以外,很多人不会为软件付费,但他们可能有计算资源,可以在上面运行模型。还有一些数据你不想发送到云端。所以,首要任务是让人们使用模型、使用 AI,或者使用你的 AI,而如果没有模型访问权限,他们可能无法做到这一点。

If you're releasing an open model, you want people to use it as the first and foremost thing. And then after that comes things like transparency and trust. I think when you look at China, the biggest reason is that they want people around the world to use these models and I think a lot of people will not. If you look outside of the US, a lot of people will not pay for software, but they might have computing resources where you can put a model on it and run it. I think there can also be data that you don't want to send to the cloud. So, this the number one thing is getting people to use models, use AI, or use your AI that might not be able to do it without having access to the model.

Host

我想我们应该明确说明,我们一直在讨论这些中国模型和开放权重模型,它们通常是在本地运行的。所以,并不是把你的数据发送到中国或开发者那里,比如硅谷的开发者。

I guess we should state explicitly, so we've been talking about these Chinese models and open weight models, often times the way they're run is locally. So, it's not like you're sending your data to China or to whoever developed, to Silicon Valley, whoever developed the model.

Nathan Lambert

很多美国初创公司通过托管这些中国模型并销售 token 来赚钱。这叫做销售 token,意味着有人会调用模型来完成一些工作。我认为另一个原因是让美国公司开阔眼界。他们 GPU 匮乏,已经到了 GPU 的极限。每次发布时,他们总是说我们的 GPU 很吃紧。我记得在 GPT-4 的一次发布会议上,Sam Altman 说:“哦,我们发布这个是因为我们可以用你们的 GPU。我们不必用自己的 GPU,OpenAI 仍然可以从中获得分发。”这是另一个非常现实的事情,因为这几乎不花他们什么成本。

A lot of American startups make money by hosting these models from China and selling them tokens. It's called like selling tokens, which means somebody will call the model to do some piece of work. I think the other reason is for US companies like that to open the eyes. So, GPU deprived like they're so they're at the limits of the GPUs. Whenever they make a release, they're always talking about like our GPUs are hurting. And I think this like like in one of these like GPT-4's release sessions, Sam Altman said like, "Oh, we're releasing this because we can use your GPUs. We don't have to use our GPUs and OpenAI can still get distribution out of this," which is a another very real thing cuz this doesn't cost them so anything.

Nathan Lambert

对于用户来说,我认为也是如此。有些用户只是像使用 ChatGPT 一样在本地使用模型,但对于公司来说,拥有这些模型是一个巨大的解锁,因为你可以定制它们,训练它们,添加后训练,增加更多数据,比如专门化到法律、医疗模型等。你提到 Ammar,中国开放权重模型的吸引力在于,它们的许可证甚至更友好。我认为它们只是不受限制的开源许可证,而如果你使用 Llama 或 Gemma,会有一些附加条件。比如用户数量有上限,如果超过一定百万用户,你需要向 Meta 等公司报告财务状况。虽然模型是免费的,但有附加条件,而人们喜欢没有附加条件的东西。所以,我认为除了性能之外,这也是中国开放权重模型如此受欢迎的原因之一,因为你可以直接使用它们,没有任何陷阱。

And for the user, I think also, I mean, there are users who just use the model locally how they would use ChatGPT, but also for companies, I think it's a huge unlock to have these models because you can customize them, you can train them, you can add post-training, add more data, like specialize them into, let's say, law, medical models, or whatever you have. And the appeal, you mentioned Ammar, the appeal of the open weight models from China is that the open weight models are also the licenses are even friendlier. I think they are just unrestricted open source licenses where if you use something like Llama or Gemma, there are some strings attached. I think it's like an upper limit in terms of how many users you have. And then if you exceed, I don't know, so and so many million users, you have to report your finance situation to, let's say, Meta or something like that. And I think, well, it is a free model, but there are strings attached, and people do like things where strings are not attached. So, I think that's also one of the reasons, besides performance, why the open weight models from China are so popular, because you can just use them. There's no catch in that sense. Yeah.

Host

生态系统在这方面已经有所改善,但主要是这些新提供商提供如此开放许可证的结果。你提到 Perplexity 时很有趣。它显示 Kimiko thinking 托管在美国,这正好是我们讨论的例子,人们对此很敏感。Kimiko thinking 和 Kimiko 是一个非常流行的模型。人们说它在创意写作和一些软件方面表现很好。所以,人们就是喜欢不同模型的这些小特点。

The ecosystem has gotten better on that front, but mostly downstream of these new providers providing such open licenses. That was funny when you pulled up Perplexity. It said Kimiko thinking hosted in the US, which is just like an exact I've never seen this, but it's an exact example of what we're talking about, where people are sensitive to this. Like Kimiko thinking and Kimiko is a model that is very popular. People say that it has very good like creative writing and also in doing some software things. So, it's just these little quirks that people pick up on with different models that they like.

开源模型的有趣想法 Interesting Ideas from Open Models

Host

这些模型探索了哪些有趣的想法,你能谈谈吗?比如哪些特别吸引你?

What are some interesting ideas that some of these models have explored that you can speak to? Like that particular interesting to you.

Nathan Lambert

也许我们可以按时间顺序来。当然,如果我们只关注 2025 年,有 1 月份发布的 DeepSeek R1。不过,它基于前一年 2024 年 12 月发布的 DeepSeek 版本 3。在架构方面有很多东西。有趣的是,你仍然可以——这就是我在从头开始的编码项目中做的事情——从 GPT-2 开始,然后向模型添加东西,使其变成另一个模型。所以,它们仍然属于同一谱系,关系非常密切。但就我想到的,DeepSeek 的独特之处在于混合专家模型。我的意思是,他们并没有发明混合专家模型。我们可以再谈谈混合专家模型是什么意思。但在深入细节之前,先列出这些:混合专家模型,然后他们还有多头潜在注意力,这是对注意力机制的一种调整。我认为,2025 年这些开放权重模型的主要区别在于不同的调整,以优化推理或 KV 缓存大小。

Maybe we can go chronologically. I mean, there was, of course, DeepSeek R1 that came out in January, if we just focus on 2025. However, this was based on DeepSeek version 3, which came out the year before in December 2024. There are multiple things on the architecture side. What is fascinating is you can still I mean, that's what I do in my from scratch coding projects. You can still start with GPT-2, and you can add things to that model to make it into this other model. So, it's all still kind of like the same lineage, the same it is a very close relationship between those. But, on top of my head, DeepSeek what was unique there is the mixture of experts. I mean, they were not inventing mixture of experts. We can maybe talk a bit more what mixture of experts means. But, just to list these things first before we dive into detail, mixture of experts, but then they also had multi-head latent attention, which is a tweak to the attention mechanism, where this was, I would say, 2025 the main distinguishing factor between these open weight models, different tweaks to make inference or KV cache size.

KV 缓存与注意力机制调整 KV Cache and Attention Mechanism Tweaks

Nathan Lambert

我们也可以稍后定义 KV 缓存,但为了让长上下文更经济,我们缩小了 KV 缓存的大小。那么,我们可以做哪些调整呢?大部分都集中在注意力机制上。DeepSeek 有多头潜在注意力。还有分组查询注意力,它仍然非常流行。它不是这些模型发明的,可以追溯到几年前,但那是另一个选择。滑动窗口注意力——如果我没记错的话,almost three 用了它。所以,这些不同的调整让模型变得不同。除此之外,我曾在一篇文章中把它们放在一起比较过。它们惊人地相似。只是中心 Transformer 块的重复次数不同,以及人们调整的一些小旋钮。但好处是,无论如何它都有效。你可以调整东西。你可以移动归一化层。你会获得一些性能提升,在消融研究中几乎总是很好,显示移动某些东西对模型的实际影响。消融研究:它让模型变好还是变差?但实现 Transformer 并让它工作的方法有很多。仍然流行的大想法是混合专家、多头潜在注意力、滑动窗口注意力、分组查询注意力。然后在年底,我们看到关注点转向让注意力机制在推理词预测时线性扩展。例如,Qwen 3 next 添加了门控 DeltaNet。它有点受状态空间模型的启发,你有一个不断更新的固定状态,但它让注意力更便宜,或者用更便宜的操作替换了注意力。

We can also define KV cache in a few moments, but to make it more economical to have long context, we shrink the KV cache size. So, what are tweaks that we can do? Most of them focus on the attention mechanism. There is multi-head latent attention in DeepSeek. There is group query attention, which is still very popular. It wasn't invented by any of those models; it goes back a few years, but that would be the other option. Sliding window attention—I think almost three uses it if I remember correctly. So, there are these different tweaks that make the models different. Otherwise, I put them all together in an article once where I just compared them. They are very surprisingly similar. It's just different numbers in terms of how many repetitions of the transformer block you have in the center, and just little knobs that people tune. But what's so nice about it is it works no matter what. You can tweak things. You can move the normalization layers around. You get some performance gains, and it's almost always very good in ablation studies showing what it actually does to the model if you move something around. Ablation studies: does it make it better or worse? But there are so many ways you can implement a transformer and still make it work. Big ideas that are still prevalent are mixture of experts, multi-head latent attention, sliding window attention, group query attention. And then at the end of the year, we saw a focus on making the attention mechanism scale linearly with inference token prediction. So, there was Qwen 3 next, for example, which added a gated DeltaNet. It's kind of inspired by state-space models where you have a fixed state that you keep updating, but it makes essentially this attention cheaper, or it replaces attention with a cheaper operation.

Transformer 架构概览 Transformer Architecture Overview

Host

也许退一步谈谈 Transformer 架构的整体情况是有用的。

And maybe it's useful to step back and talk about transformer architecture in general.

Nathan Lambert

是的,所以也许我们应该从 GPT-2 架构开始,这个 Transformer 源自《Attention is All You Need》论文。

Yes, so maybe we should start with the GPT-2 architecture, the transformer that was derived from the 'Attention is All You Need' paper.

Host

嗯。

Mhm.

Nathan Lambert

所以,《Attention is All You Need》论文中的 Transformer 架构有两个部分:编码器和解码器。而 GPT 只关注解码器部分。它本质上仍然是一个神经网络,内部有注意力机制。你一次预测一个词,通过嵌入层传递。然后是 Transformer 块。Transformer 块有注意力模块和一个全连接层。中间还有一些归一化层,但本质上就是带有注意力机制的神经网络层。所以,从 GPT-2 开始,当我们转向 GPT-OSS 时,例如,有了混合专家层。它不是 GPT-OSS 发明的,已有几年历史,但它本质上是一种调整,使模型更大而不增加每次前向传播的算力消耗。所以,这里有这个全连接层。如果听众熟悉多层感知机,你可以把它想象成 Transformer 内部的一个小型多层感知机,一个全连接神经网络层。它非常昂贵,因为它是全连接的。如果你有 1000 个输入,1000 个输出,那就是 100 万个连接。这是 Transformer 中非常昂贵的部分。想法是把它扩展成多个前馈网络。所以,不是只有一个,假设你有 256 个,但这会让它更昂贵,因为现在你有 256 个,但你不会同时使用它们。所以,你有一个路由器,它说:“好的,基于这个输入词,使用这个全连接网络会很有用。”在这种情况下,它被称为专家。所以,混合专家意味着你有多个专家。根据你的输入,比如更偏数学,它会使用不同的专家,而比如将输入文本从英语翻译成西班牙语,它可能会咨询不同的专家。这不太明确,我的意思是,不像说“这只是数学和西班牙语的专家”那么清晰。它有点模糊,但想法本质上是你在网络中打包更多知识,但并非所有知识都一直使用。那会很浪费。所以,在词生成过程中,你更有选择性。有一个路由器选择哪些词应该去哪个专家。这增加了复杂性。训练更难。有很多可能出错的地方,比如崩溃等等。所以,我认为这就是为什么 almost three 仍然使用密集模型。我的意思是,我认为所有模型都有混合专家,但密集模型,其中密集——这也是术语。密集和稀疏有区别。混合专家被认为是稀疏的,因为我们有很多专家,但只有少数是活跃的。所以,那被称为稀疏。而密集则相反,你只有一个全连接模块,它总是被使用。

So, the 'Attention is All You Need' paper had a transformer architecture that had two parts, an encoder and a decoder. And GPT went just focusing on the decoder part. It is essentially still a neural network and it has this attention mechanism inside. And you predict one token at a time, you pass it through an embedding layer. There's the transformer block. The transformer block has attention modules and a fully connected layer. And there are some normalization layers in between, but it's essentially a neural network layers with this attention mechanism. So, coming from GPT-2, when we move on to GPT-OSS, there is, for example, the mixture of experts layer. It's not invented by GPT-OSS; it's a few years old, but it is essentially a tweak to make the model larger without consuming more compute in each forward pass. So, there is this fully connected layer. And if listeners are familiar with multi-layer perceptrons, you can think of a mini multi-layer perceptron, a fully connected neural network layer inside the transformer. And it's very expensive because it's fully connected. If you have 1,000 inputs, 1,000 outputs, it's like a 1 million connections. And it's a very expensive part in this transformer. And the idea is to kind of expand that into multiple feedforward networks. So, instead of having one, let's say you have 256, but it would make it way more expensive because now you have 256, but you don't use all of them at the same time. So, you now have a router that says, 'Okay, based on this input token, it would be useful to use this fully connected network.' And in that context, it's called an expert. So, a mixture of experts means you have multiple experts. And depending on what your input is, let's say it's more math heavy, it would use different experts compared to, let's say, translating input text from English to Spanish. It would maybe consult different experts. It's not quite clear, I mean, not as clear-cut to say, 'Okay, this is only an expert for math and for Spanish.' It's a bit more fuzzy, but the idea is essentially that you pack more knowledge into the network, but not all the knowledge is used all the time. That would be very wasteful. So, you're kind of like during the token generation, you're more selective. There's a router that selects which tokens should go to which expert. Adds more complexity. It's harder to train. There's a lot that can go wrong, like collapse and everything. So, I think that's why almost three still uses dense. I mean, you have, I think, all models with mixture of experts, but dense models, where dense means—also, it's jargon. There's a distinction between dense and sparse. So, mixture of experts is considered sparse because we have a lot of experts, but only few of them are active. So, that's called sparse. And then, dense would be the opposite, where you only have like one fully connected module and it's always utilized.

从 GPT-2 到如今的演进 Evolution from GPT-2 to Today

Host

所以,也许这也是谈论 KV 缓存的好地方,但实际上,在此之前,甚至退一步看,从根本上说,从 GPT-2 到今天实现了多少新想法?这些架构到底有多不同?

So, maybe this is a good place also to talk about KV cache, but actually, before that, even zooming out, like fundamentally, how many new ideas have been implemented from GPT-2 to today? Like, how different really are these architectures?

Nathan Lambert

想象一下混合专家。GPT-3 中的注意力机制,那就是分组查询注意力机制。所以,从多头注意力到分组查询注意力是一个微调,所以我们有两个。我认为他们用 RMS 归一化替换了层归一化,但这只是不同的归一化层。不是大变化,只是一个调整。非线性激活函数——熟悉深度神经网络的人,我的意思是,就像用 ReLU 替换 sigmoid 一样。它没有从根本上改变网络,只是一个调整。你喜欢的微调。就这些,我会说。它并没有根本上的不同。仍然是相同的架构。所以,你可以从一个转换到另一个,基本上只需添加这些变化。

Picture like the mixture of experts. The attention mechanism in GPT-3, that would be the group query attention mechanism. So, it's a slight tweak from multi-head attention to group query attention, so that we have two. I think they replaced layer norm by RMS norm, but it's just like a different normalization layer. Not a big change, it's just a tweak. The non-linear activation function—people familiar with deep neural networks, I mean, it's the same as changing sigmoid with ReLU. It's not changing the network fundamentally, it's just a tweak. You like a little tweak. And that's about it, I would say. It's not really fundamentally that different. It's still the same architecture. So, you can convert one from one—you can go from one into the other by just adding these changes basically.

Host

它从根本上仍然是相同的架构。

It's fundamentally still the same architecture.

Nathan Lambert

是的。所以,例如,你之前提到我的书,书里用的是 GPT-2 模型,因为它简单且非常小。大约 1.24 亿参数。但在附加材料中,我确实有 almost three 从头实现、Gemma three 从头实现以及其他类型的从头实现模型。我总是从我的 GPT-2 模型开始,然后调整并添加不同的组件,你就从一个得到另一个。这有点像一种谱系,是的。

Yep. So, for example, you mentioned my book earlier, that's GPT-2 model in the book because it's simple and it's very small. So, 124 million parameters approximately. But in the bonus materials, I do have almost three from scratch, Gemma three from scratch, and other types of from scratch models. And I always started with my GPT-2 model and just, you know, tweaked and added different components and you get from one to the other. It's like it's kind of like a lineage in a sense, yeah.

Host

你能为人们建立一种直觉吗?因为当你退一步看,AI 世界有如此多的快速进步?而同时,架构从根本上没有改变。

Can you build up an intuition for people because sort of when you zoom out, you look at it, there's so much rapid advancement in the AI world? And at the same time, fundamentally the architectures have not changed.

后训练与系统级改进 Post-training focus and system-level improvements

Host

那么,所有的动荡、进步的喧嚣发生在哪里?收益在哪里?

So, where is all the turbulence, the turmoil of the advancement happening? Where are the gains to be had?

Nathan Lambert

所以,开发或训练网络有不同的阶段。有预训练。以前他们只用 GPT-2 做预训练。现在,有预训练、中期训练和后训练。我认为我们现在处于后训练聚焦阶段。我的意思是,预训练如果扩展到更高质量的数据仍然有优势。但随后我们有了能力解锁,这在 GPT-2 时代是没有的。例如,ChatGPT 基本上是一个 GPT-3 模型,而 GPT-3 在架构上与 GPT-2 相同。新的是加入了监督微调和基于人类反馈的强化学习(RLHF)。所以,更多的是算法层面而非架构层面。

So, there are the different stages where you develop the network or train the network. You have the pre-training. Now, back then they were just pre-training with GPT-2. Now, you have pre-training, mid-training, and post-training. So, I think right now we are in the post-training focus stage. I mean, pre-training still gives you advantages if you scale it up to better higher quality data. But then we have capability unlocks that were not there with GPT-2 for example. ChatGPT, it is basically a GPT-3 model and GPT-3 is the same as GPT-2 in terms of architecture. What was new was adding the supervised fine-tuning and the reinforcement learning with human feedback. So, it's more on the algorithmic side rather than the architecture.

Host

我会说系统也变化很大。如果你听英伟达的公告,他们谈论这些事,比如现在可以做 FP8,可以做 FP4。正在发生的是,这些实验室在想办法利用更多算力投入到一个模型中,从而训练更快。这让他们能放入更多数据。然后通过这样做,你可以更快找到更好的配置。所以,你可以看每个 GPU 每秒的 token 数,这是大规模训练时的一个指标。通过开启 FP8 训练,你可以从 10K 提升到 13K,这意味着每个参数使用的内存更少。通过保存更少的信息,通信更少,训练更快。所以,所有这些系统层面的东西支撑了更快的实验循环,这是一个持续进行的循环。很难描述,当你看到架构完全相同时,但用于训练这些模型的代码库将大不相同,你可能在挂钟时间上训练 GPT-3 20B 比当年训练 GPT-2 快得多。

I would say that the systems also change a lot. I think if you listen to Nvidia's announcements, they talk about these things like you now do FP8, you can now do FP4. And what is happening is these labs are figuring out how to utilize more compute to put it into one model which lets them train faster. And that lets them put more data in. And then you can find better configurations faster by doing this. So, you can look at the essentially the tokens per second per GPU is a metric that you look at when you're doing large-scale training. And you could go from like 10K to 13K by turning on FP8 training, which means they're using less memory per parameter in the model. And by saving less information, you do less communication and you can train faster. So, all of these system things underpin way faster experimentation on data and algorithms that is kind of like this loop that keeps going. It's hard to describe when you look at the architecture and they're exactly the same, but the code base used to train these models is going to be vastly different and you could probably train GPT-3 20B way faster in wall clock time than GPT-2 was trained at the time.

Nathan Lambert

是的,就像你说的,例如在混合专家模型中,有 FP4 优化,可以获得更高的吞吐量,但我认为这是为了速度,确实如此,但从某种意义上说,它并没有给模型带来新能力。只是我们能在多大程度上使计算更粗糙而不损害模型性能。但我确实认为 Transformer 的替代方案正在涌现。有文本扩散模型,完全不同的范式。还有 Mamba 模型,它是一种状态空间模型,但它们都有权衡,目前还没有任何东西取代自回归 Transformer 成为最先进的模型。所以,对于最先进的模型,你仍然会使用它,但现在有更便宜的替代方案,这些方案做出了一些妥协,但不再是只有一种架构了。有一些小的架构正在出现,但如果我们谈论最先进的,基本上仍然是源自 GPT-2 的自回归 Transformer 架构。

Yeah, like you said, they had for example in the mixture of experts this FP4 optimization, for example, where you get more throughput, but I do think this is for the speed, this is true, but it doesn't give the model new capabilities in a sense. It's just how much can we make the computation coarser without suffering in terms of model performance degradation. But I do think there are alternatives popping up to the transformer. There's text diffusion models, completely different paradigm. And there's also Mamba models, it's a state space model, but they do have trade-offs and what's right is there's nothing that has replaced the auto regressive transformer as state of the art model. So, for state of the art, you would still go with that thing, but there are now alternatives for the cheaper and like alternatives that are kind of making compromises, but it's not just one architecture anymore. There are little ones coming up, but if we talk about the state of the art, it's pretty much still the transformer architecture auto regressive derived from GPT-2 essentially.

预训练、后训练与推理的缩放定律 Scaling laws across pre-training, post-training, and inference

Host

我想这里的大问题是,我们谈了很多关于预训练背后的架构。缩放定律在预训练、后训练、推理、上下文窗口、数据、合成数据方面是否仍然强劲?

I guess the big question here is we talked quite a bit here on the architecture behind the pre-training. Are the scaling laws holding strong across pre-training, post-training, inference, context size, data, synthetic data?

Nathan Lambert

我想从缩放定律的技术定义开始,这有助于理解这一切。缩放定律是一种幂律关系,你可以把 x 轴看作你正在扩展的东西,是算力和数据的组合,它们有些相似。y 轴是保留的下一个词预测准确率。我们谈到模型是自回归的。就像如果你保留一组模型未见过的文本,训练时它能有多准确?缩放定律的想法来自于人们发现这是一个非常可预测的关系。我认为这个技术术语仍在延续,然后问题是用户从中得到了什么?还有更多类型的缩放,OpenAI 的 o1 以引入推理时间缩放而闻名。我认为不那么出名的是它也展示了你可以扩展强化学习训练,得到一种对数 x 轴和 y 轴线性性能提升的关系。所以现在有三种轴:传统的缩放定律用于预训练,即模型大小和数据集大小;然后是扩展强化学习,即你能做多长时间的试错学习,我们稍后会定义更多;然后是推理时间算力,即让模型在特定问题上生成更多 token。我比较乐观,它们都还在起作用,但低垂的果实大多已被摘取,尤其是在过去一年中,基于可验证奖励的强化学习(RLVR)和推理时间缩放,这就是为什么这些模型使用起来感觉如此不同,以前你会立即得到第一个 token,现在它们会花几秒、几分钟甚至几小时生成这些隐藏的思考,然后才给出答案的第一个词。这完全就是推理时间缩放,它在模型能力变化方面是一个美妙的阶跃函数。它们实现了工具使用等功能,实现了我们之前讨论的更好的软件工程。而这一切几乎完全归功于基于可验证奖励的强化学习训练,它让模型很容易地掌握这些技能。所以,让模型学习。如果你看推理过程,当模型生成大量 token 时,它经常做的是尝试一个工具,查看返回结果,尝试另一个 API,查看返回结果,然后解决问题。所以,模型在训练时很快学会这样做。最终,这给了我们一个通用基础,模型可以在你的仓库中很好地使用 CLI 命令,为你处理 Git,移动和整理东西,或搜索更多信息,而一年前我们坐在这些椅子上时,并没有想到模型能做到这些。所以,这只是今年发生的事情,完全改变了我们对使用 AI 的看法,我认为这非常神奇。

I like to start with the technical definition of scaling law which kind of informs all of this. The scaling law is a power law relationship between you could think of the x-axis so kind of what you are scaling is a combination of compute and data which are kind of similar. And then the y-axis is like the held out prediction accuracy over next tokens. We talked about models being auto regressive. It's like if you keep a set of text that the model has not seen, how accurate will it get when you train? And the idea of scaling laws came when people figured out that that was a very predictable relationship. And I think that that technical term is continuing and then the question is like what do users get out of it? And then there are more types of scaling where OpenAI's o1 was famous for introducing inference time scaling. And I think less famously for also showing that you can scale reinforcement learning training and get kind of this log x-axis and then a linear increase in performance on y-axis. So there's kind of these three axes now where the traditional scaling laws are talked about for pre-training which is how big your model is and how big your data set is. And then scaling reinforcement learning which is like how long can you do this trial and error learning that we will talk about will define more of this. And then this inference time compute which is just letting the model generate more tokens on a specific problem. So I'm kind of bullish where they're all really still working, but the low-hanging fruit has mostly been taken especially in the last year on reinforcement learning with verifiable rewards which is this RLVR and then inference time scaling which is just why these models feel so different to use, where previously you would get that first token immediately, and now they will go off for seconds, minutes, or even hours generating these hidden thoughts before giving you the first word of your answer. And that's all about this inference time scaling, which is such a wonderful kind of step function in terms of how the models change abilities. They kind of enabled this tool use stuff, and enabled this much better software engineering that we were talking about. And this is when we say enabled almost entirely downstream of the fact that this reinforcement learning with verifiable rewards training just kind of let the models pick up these skills very easily. So, let the models learn. So, if you look at the reasoning process when the models are generating a lot of tokens, what it will be often doing is it tries a tool, it looks at what it gets back, it tries another API, it sees what it gets back, and if it solves the problem. So, the models when you're training them very quickly learn to do this. And then, at the end of the day, that gives us kind of general foundation where the model can use CLI commands very nicely in your repo, and handle Git for you, and move things around, and organize things, or search to find more information, which if we're sitting in these chairs a year ago is something that we didn't really think of the models being doing. So, this is just kind of something that has happened this year, and is totally transformed how we think of using AI, which I think is very magical.

预训练缩放与成本 Pre-training scaling and cost

Host

你刚才说了很多,而且说得很快,很有深度。我们不妨稍微展开一下。你说你基本上看好所有形式的 Scaling(规模扩张)。那么,我们能不能就从最基础的开始?预训练。我们是不是在暗示预训练 Scaling 的“低垂果实”已经被摘完了?预训练是遇到了瓶颈,还是说你仍然看好预训练?

So, you've actually said quite a lot of things there, and said profound things quickly. It would be nice to unpack them a little bit. You said you're bullish basically on every version of scaling. So, can we just even start at the beginning. Pre-training. Are we kind of implying that the low-hanging fruit on pre-training scaling has been picked? Is pre-training hit a plateau or is even pre-training still you're bullish on?

Nathan Lambert

预训练已经变得极其昂贵。我认为,扩大预训练规模也意味着你要向用户提供一个非常大的模型。所以,大致可以确定的是,GPT-4 及类似模型的最大规模大约在 1 万亿参数量级。有很多传言说,随着训练效率的提高,模型实际上变得更小了。你希望模型更小,因为这样服务成本会成比例下降。这些模型的训练成本相对于为亿万用户提供服务的成本来说其实很低。我记得 DeepSeek 有一个著名的数字,按云市场价计算,预训练成本大约 500 万美元。我想大概是 300 万?嗯,论文的第 2.4 节,我们详细说明了 GPU 集群用于训练的时间,包括工程问题、多次随机种子等,租用集群处理训练模型的各种问题和麻烦大约花了 200 万美元。所以,这些模型——很多人花 100 万到 1000 万美元就能训练一个模型,但为数百万用户提供服务的经常性成本实际上是数十亿美元的算力。你可以看看,租用 1000 块 GPU 每天就要 10 万美元,而这些公司可能拥有数百万块 GPU。你可以算算这些东西闲置的成本。所以,这是一个大问题,然后就是,如果 Scaling 真的能给你一个更好的模型,它在财务上是否值得?我认为,随着 AI 解决更引人注目的任务,我们会慢慢推动它。比如,Claude Opus 4.5 让 Claude Code 真正能用了。我启动了一个叫 Adam 的项目,就是美国真正开放模型,在 7 月份,那是一个真正的“氛围编码”网站,我的工作是制作图表之类的东西。几周前我回来刷新它,Claude Opus 4.5 与当时其他模型相比,简直碾压了我在 6 月、7 月构建时遇到的所有问题,而且它可能是一个更大的模型。这里面有很多因素,但进步仍在继续。

Pre-training has gotten extremely expensive. I think to scale up pre-training, it's also implying that you're going to serve a very large model to the users. So, I think that it's been loosely established the likes of GPT-4 and similar models were around 1 trillion like this order of trillion parameters at the biggest size. There's a lot of rumors that they've actually gotten smaller as training has gotten more efficient. You want to make the model smaller because then your costs of serving go down proportionally. These models the cost of training them is really low relative to the cost of serving them to hundreds of millions of users. I think DeepSeek had this famous number of about $5 million for pre-training at cloud market rates. I think almost three. Um section 2.4 in the paper, we just detailed how long we had the GPU clusters sitting around for training, which includes engineering issues, multiple seeds, and it was like about $2 million to rent the cluster to like deal with all the problems and headaches of training a model. So, these models are pretty like a lot of people could get $1 to $10 million to train a model, but the recurring costs of serving millions of users is really billions of dollars of compute. I think that you can look at close like 1,000 GPU rental, you can pay $100,000 a day for and these companies could have millions of GPUs. Like you can look at how much these things cost to sit around. So, that's kind of a big thing and then it's like if scaling is actually giving you a better model, like is it going to be financially worth it? And I think it will kind of slowly we'll push it out as AI solves more compelling tasks. So, like the likes of Claude Opus 4.5 making Claude code just work for things. I think I I launched this project called like the Adam project, which is like American truly open models in July and that was like a true vibe coded website and like I have a job make plots and stuff and then I came back to refresh it in the last few weeks and it's like Claude Opus 4.5 versus whatever model at the time was like just crushed all the issues that I had from building in June July and like might be a bigger model. There's a lot of things go into this but that's like there's still progress coming.

Host

所以你说的是缩放定律 Y 轴的细微差别,即实际体验与基准测试上的智能可能不同,但你对预训练的直觉是,如果扩大算力规模,模型会变得更好,不是从财务可行性角度,而是仅从定律方面。你认为模型会变得更聪明吗?

So what you're speaking to is the nuance of the Y axis of the scaling laws that the way it's experienced versus on a benchmark the actual intelligence is might might be different but still your intuition about pre-training if you scale the size of compute will the models get better not whether it's financially viable but just from the law aspect of it. Do you think the models will get smarter?

Nathan Lambert

是的,我认为——这有时听起来像是 AI 公司领导层说的近乎幻灭的话,但他们会说,缩放定律在 13 个数量级的算力上仍然成立,为什么会结束呢?所以我认为从根本上说,它不太可能停止,只是最终我们甚至无法测试更大的规模,因为更多算力会带来各种问题。我认为有很多讨论说 2026 年将是大型 Blackwell 算力集群(比如超大规模云服务商的吉瓦级设施)上线的一年,这些电力合同和数据中心都是在 2022 年和 2023 年签署和寻找的。所以是在 ChatGPT 之前或之后不久,建造这些更大的集群来训练模型需要两到三年的准备时间。显然,人们对建造更多数据中心有着巨大的兴趣。所以,关键点在于,人们说这些新集群即将到来,实验室将有更多算力用于训练,他们会利用这些算力,但这并非必然。我看到了如此多的进步,所以我期待它,我期待稍大一点的模型,我期待——我想说,今年我们会看到 2000 美元的订阅。我们已经看到了 200 美元的订阅。它可以再增长 10 倍。这些都是可能发生的事情,它们都源于这个稍大一点的模型,它提供了更前沿的能力。

Yeah and I think that there's and this sometimes comes off as like almost like disillusioned from people leadership AI companies saying this but they're like it's held for 13 orders of magnitude of computer or something like why would it ever end? So I think fundamentally it is pretty unlikely to stop it's just like eventually we're not even going to be able to test the bigger scales because of all the problems that come with more compute. I think that there's a lot of talk on how 2026 is a year when very large Blackwell compute clusters like gigawatt scale facilities that hyperscalers are coming online and these were all contracts for power and data centers that were signed and sought out in like 2022 and 2023. So before or right after ChatGPT so it took this two to three year lead time to build these bigger clusters to train the models. Well there's obviously immense interest in building even more data centers than that. So that is like kind of the crux that people are saying is like these new clusters are coming the labs are going to have more compute for training they're going to utilize this but it's not a given and it's like I I've seen so much progress that I expect it and I expect a little bit bigger models and I expect um I would say it's more like we'll see a $2,000 subscription this year. We've seen $200 subscriptions. It's like that can 10x again. And these are the kind of things that could come, and they're all downstream of this like bit bigger bit bigger model that offers just a little bit more cutting edge.

Host

那么,据报道,xAI 将在 2026 年初达到 1 吉瓦规模,年底达到 2 吉瓦。你认为他们会如何在缩放定律的背景下利用这些算力?其中很多是推理吗?很多是训练吗?

So, you know, it's reported that xAI is going to hit that 1 gigawatt scale early '26 and full 2 gigawatt by year end. How do you think they'll utilize that in the context of scaling laws? Is a lot of that inference? Is a lot of that training?

Nathan Lambert

最终是以上所有方面。所以,我认为你在训练模型时的所有决策都回归到预训练。因此,如果你要在模型中扩展强化学习,你仍然需要决定能够实现这一点的架构。我们讨论过其他架构,比如使用不同类型的注意力机制。我们也讨论了混合专家模型。MoE 模型的稀疏特性使得生成更加高效,这成为后训练的重要组成部分。你需要准备好架构,以便真正扩展这些算力。我仍然认为大部分算力会投入预训练,因为你仍然可以让模型变得更好。你仍然想重新审视这一点。你仍然想要你能得到的最好的基础模型。几年后,这会饱和,强化学习的算力会持续更长时间。

It ends up being all of the above. So, I think that all of your decisions when you're training a model come back to pre-training. So, if you're going to scale RL in a model, you still need to decide on your architecture that enables this. We were talking about like other architectures than using different types of attention. We're also talking about mixture of experts models. This sparse nature of MoE models makes it much more efficient to do generation, which becomes a big part of post-training. And it's like you need to have your architecture ready so that you can actually scale up this compute. I still think most of the compute is going in at pre-training because you can still make a model better. You still want to go and revisit this. You will still want the best base model that you can. And in a few years that'll saturate, and the the RL compute will just go longer.

Host

有没有人不同意你,说预训练基本上已经死了?一切都关于扩展推理、扩展后训练、扩展上下文、持续学习、扩展数据、合成数据。

Is there a people who disagree with you that say basically pre-training is dead? It's all about scaling inference, scaling post-training, scaling context, continual learning, scaling data, synthetic data.

Nathan Lambert

人们有这种感觉,也这么描述,但我认为实际情况并非如此。

People vibe that way and describe it in that way, but I think it's not the practice that is happening.

Host

人们普遍说这东西已经死了。

General vibe of people saying this thing is dead.

Nathan Lambert

这种说法在其他地方。所以,强化学习的“低垂果实”在其他地方。比如,我们在 11 月发布了我们的模型——每家公司都有截止日期。我们的截止日期大概是 11 月 20 日。我们的强化学习运行了 5 天,与 2024 年相比,对于一个大约 300 亿参数的模型来说,仅仅做后训练这么长时间是很长的。它不是一个大型模型。

Is elsewhere. So, the low-hanging fruit in RL is elsewhere. Like, for example, we released our model in November for Every company has deadlines. Our deadline was like November 20th. And our for that our RL run was 5 days, which compared to 2024 is a very long time to just be doing post training at a model of like 30 billion parameters. It's not a big model.

后训练与发布周期 Post-training and release cycles

Nathan Lambert

然后在 12 月我们又发布了一次,只是让强化学习再跑了 3 周半,模型明显变好了,所以就发布了。把那么多时间花在一年中最重要的东西上,这可不是小事。所以训练模型时就会遇到这类决策,你不能一直放着不管。你必须不断吸收研究人员的改进成果。所以你会重做预训练,做一个月后训练,但之后你需要把它交给用户,还要做安全测试。所以我认为有很多因素在强化这种不断更新模型的循环。总有可以改进的地方。你得到一个新的算力集群,可以让你做得更稳定或更快。你听到很多关于 Blackwell 推出问题的消息。在 AI2,大多数模型是在大约 1000 到 2000 个 GPU 上预训练的。但当你在 10000 或 100000 个 GPU 上预训练时,你会遇到完全不同的故障。GPU 会以奇怪的方式损坏。运行 100000 个 GPU 几乎保证至少有一个 GPU 宕机。你需要让训练代码处理这种冗余,这是一个完全不同的问题。而我们正在做的,比如我在 DGX Spark 上玩后训练,或者在你的书里,或者人们学习机器学习,他们训练这些最大模型所面临的问题就是大规模分布式扩展。这是一个非常不同的问题。但这与实现缩放定律的系统问题有些不同,尤其是在预训练阶段。你需要同时拥有所有这些 GPU。当我们转向强化学习时,它实际上更适合异构算力,因为你有很多模型副本。简单介绍一下语言模型强化学习:你有两组 GPU,一组叫 actor,一组叫 learner。learner 是实际进行强化学习更新的地方。传统上是策略梯度算法,近端策略优化(PPO)和组相对策略优化(GRPO)是两种流行的类别。另一边,你会有 actor,它们生成补全。这些补全是你将要评分的内容。所以强化学习就是关于优化奖励。在实践中,你可以让世界上不同地方的许多不同 actor 处理不同类型的问题,然后将结果发送回这个高度网络化的算力集群进行实际学习,在那里你获取梯度,你需要一个紧密耦合的网络,以便进行不同类型的并行化,并分散你的模型以实现高效训练。所以每种不同类型的训练和服务都有这些需要考虑的扩展问题。我们谈到了预训练,谈到了强化学习,然后推理时扩展就像,如何为一个思考一小时、服务 1 亿用户的模型提供服务?我不太清楚,但我知道那是一个难题。为了给人们这种智能,我们需要解决所有这些系统问题,需要更多算力,需要更稳定的算力。

And then in December we had another release, which was just we let the RL run for another 3 and a half weeks and the model got notably better, so we released it. And that's a big amount of time to just allocate to something that is going to be your peak for the year. So these types of decisions happen when they're training a model where they just can't leave it forever. You have to keep pulling in the improvements you have from your researchers. So you redo pre-training. You'll do this post-training for a month, but then you need to give it to your users. You need to do safety testing. So I think there's a lot in place that reinforces this cycle of just keep updating the models. There are things to improve. You get a new compute cluster that lets you do something maybe more stable or faster. You hear a lot about Blackwell having rollout issues. At AI2, most of the models were pre-training on around 1 to 2,000 GPUs. But when you're pre-training on 10,000 or 100,000 GPUs, you hit very different failures. GPUs are known to break in weird ways. Doing a 100,000 GPU run means you're pretty much guaranteed to always have at least one GPU that is down. And you need to have your training code handle that redundancy, which is a very different problem. Whereas what we're doing, like I'm playing with post-training on a DGX Spark, or in your book, or people learning ML, what they're battling to train these biggest models is just massive distributed scale. It's a very different problem. But that's somewhat different from the systems problem to enable the scaling laws, especially at pre-training. You need all these GPUs at once. When we shift to reinforcement learning, it actually lends itself to heterogeneous compute because you have many copies of a model. To do a primer for language model reinforcement learning, you have two sets of GPUs. One you can call the actor, and one you call the learner. The learner is where your actual reinforcement learning updates are going to happen. These are traditionally policy gradient algorithms, proximal policy optimization (PPO), and group relative policy optimization (GRPO) are the two popular classes. On the other side, you're going to have actors, which are generating completions. These completions are the things that you're going to grade. So reinforcement learning is all about optimizing reward. In practice, you can have a lot of different actors in different parts of the world doing different types of problems, and then you send it back to this highly networked compute cluster to do the actual learning, where you take the gradients, and you need to have a tightly meshed network where you can do different types of parallelism and spread out your model for efficient training. So every different type of training and serving has these considerations you need to scale. We talked about pre-training, we talked about RL, and then inference time scaling is like, how do you serve a model that's thinking for an hour to 100 million users? I don't really know about that, but I know that's a hard problem. To give people this intelligence, there are all these systems problems where we need more compute, and you need more stable compute to do it.

看好所有缩放方向 Bullish on all scaling directions

Host

但我听到的是,你对所有这些类型的扩展都持乐观态度。推理、推理能力,甚至预训练。

But you're bullish on all of these kinds of scaling, is what I'm hearing. On the inference, on the reasoning, even on the pre-training.

Nathan Lambert

是的,这确实是个大问题,但基本上两个旋钮是训练和推理扩展,你可以从中获得收益。在一个假设拥有无限算力的世界里,你会想全部做。所以你有训练,有推理扩展,而训练是一个层级:预训练、中期训练、后训练。改变模型大小、更多训练数据、训练更大的模型会给模型带来更多知识。模型有了更好的基础。过去,我们仍称之为基础模型,它解锁了能力。但你不能让模型在预训练期间或预训练后解决最复杂的任务。你仍然有其他解锁阶段,比如中期训练或长上下文,例如,通过 RLHF 进行后训练可以解锁模型在预训练中已有的知识所对应的能力。我认为如果你做更多预训练,你会得到一个更好的基础模型,以后可以解锁,但正如 Nathan 所说,它变得太贵了。所以我们没有无限算力,你必须决定:我是否要把算力更多地花在让模型更大上?这是一个权衡。在理想世界中,你想全部做,我认为从这个意义上说,扩展仍然很有活力。你仍然会得到更好的模型,但就像我们在 GPT-4.5 上看到的,这不值得。你可以用其他技术在当前时刻解锁更多性能。特别是如果你看推理扩展,那是今年最大的收获之一,比如 o1,它让一个较小的模型比预训练一个更大的模型如 GPT-4.5 走得更远。所以我不认为预训练扩展已死;只是现在有其他更有吸引力的扩展方式。但在某个时候,你仍然想在预训练上取得一些进展。还要考虑你想把钱花在哪里。如果你更多地花在预训练上,那是固定成本。你训练模型,然后它永远拥有这种能力。你总是可以使用它。对于推理扩展,你在训练时不花钱;你后来每次查询花钱。然后还有数学问题:我的模型会在市场上多久?如果半年后我就替换它,也许不值得花 500 万、1000 万、1 亿美元去训练更长时间。也许我会做更多推理扩展,从那里获得性能。它可能花我 200 万美元的用户查询费用。这变成了你有多少用户的问题,然后做数学计算。我认为这也是有趣的地方:ChatGPT 处于一个位置,他们有很多用户,所以他们需要更便宜一点,用那个稍微小一点的 GPT-5 模型。其他公司根据他们的客户有不同的权衡。

Yeah, so that's a big can of worms here, but basically the two knobs are the training and the inference scaling, where you can get gains. In a world where we had, let's say, infinite compute resources, you want to do all of them. So you have training, you have inference scaling, and training is a hierarchy: pre-training, mid-training, post-training. Changing the model size, more training data, making a bigger model gives you more knowledge in the model. The model has a better base. Back in the day, we still call it foundation model, and it unlocks capabilities. But you don't have the model be able to solve your most complex tasks during pre-training or after pre-training. You still have these other unlock phases where you have mid-training or long context, for example, post-training with RLHF that unlocks capabilities that the model has in terms of just knowledge in the pre-training. I think if you do more pre-training, you get a better base model that you can unlock later, but like Nathan said, it just becomes too expensive. So we don't have infinite compute, so you have to decide: do I want to spend that compute more on making the model larger? It's a trade-off. In an ideal world, you want to do all of them, and I think in that sense, scaling is still pretty much alive. You would still get a better model, but like we saw with GPT-4.5, it's just not worth it. You can unlock more performance with other techniques at that current moment. Especially if you look at inference scaling, that's one of the biggest gains this year with o1, where it took a smaller model further than pre-training a larger model like GPT-4.5. So I wouldn't say pre-training scaling is dead; it's just that there are other more attractive ways to scale right now. But at some point, you will still want to make some progress on the pre-training. The thing is also to consider where you want to spend your money. If you spend it more on the pre-training, it's a fixed cost. You train the model and then it has this capability forever. You can always use it. With inference scaling, you don't spend money during training; you spend money later per query. And then it's also the math: how long is my model going to be on the market? If I replace it in half a year, maybe it's not worth spending 5 million, 10 million, 100 million dollars on training it longer. Maybe I will just do more inference scaling and get the performance from there. It might cost me 2 million in terms of user queries. It becomes a question of how many users you have and then doing the math. I think that's also where it's interesting: ChatGPT is in a position where they have a lot of users, so they need to go a bit cheaper, with that GPT-5 model that is a bit smaller. Other companies have different trade-offs depending on their customers.

预训练、中训练与后训练定义 Pre-training, mid-training, and post-training definitions

Host

我觉得现在正好可以定义一下预训练、中训练和后训练。

I think this might be a good place to define pre-training, mid-training, and post-training.

Nathan Lambert

预训练是经典的训练方式,一次预测一个下一个词。你有一个大型语料库。Nathan 在这方面也有非常有趣的见解,因为论文的很大一部分关注的是正确的数据混合。所以,预训练本质上就是在互联网数据、书籍、论文等海量语料库上,对下一个词预测进行交叉熵损失训练。这些年来它发生了一些变化。人们过去会把所有能用的数据都扔进去。现在,不仅仅是原始数据,还有合成数据,人们会重新表述某些内容。合成数据不一定意味着纯粹由 AI 编造的数据。它也包括从一篇文章、维基百科文章中提取内容,然后将其改写为问答形式,或者总结、改写,从而制作出更好的数据。我认为这就像一个人读一本书和读一篇混乱的 Reddit 帖子相比,你学得更好。会有一篇关于这个的帖子。一些 Reddit 数据非常常见且适合训练。你只需要过滤它。我认为如果有人把那些数据重新表述得更简洁、更有条理,那就是更高质量的数据,能让语言模型最终得到相同的结果,但更快达到。它训练得更快,因为如果语法和标点正确,它就已经学到了正确的方式,而不是从混乱的方式获取信息,然后再学习如何纠正。所以我认为这就是预训练的演变方式。虽然 Scaling 仍然有效,但不仅仅是数据量的问题。还有让数据变得更好的技巧。然后中训练,它以前被称为预训练。之所以叫中训练,是因为有预训练和后训练,中间却没有东西,这有点尴尬。所以中训练通常类似于预训练,但更专业化。算法相同,但你会专注于长上下文之类的东西。不在纯预训练阶段做这个的原因是,你没有那么多长上下文文档。所以你需要一个专门的阶段。语言模型的一个问题也是灾难性遗忘。你教它一些东西,它会忘记其他东西。这就像没有免费的午餐。和人类一样。如果你问我 10 年前学过的数学,我不知道。Nathan 实际上说过,他消费了太多内容,以至于出现了灾难性遗忘的问题。我试图学习很多关于 AI 的知识,就像我在学习预训练并行性。我好像丢失了一些东西,不知道是什么。我不想将 LLM 拟人化,但我认为在人类学习的方式上是一样的。数量并不总是更好。选择性是关键。中训练就是在质量内容上有所选择,这样 LLM 最后看到的是高质量的东西。然后后训练就是所有的微调、监督微调、DPO、基于可验证奖励的强化学习、基于人类反馈的强化学习等等。这些是精炼阶段。成本问题也很有趣。预训练要花很多钱。强化学习少一点。强化学习并不真正教授知识;它更像是解锁知识。更像是技能学习,如何利用预训练的知识解决问题。今年或去年(2025 年)实际上有三篇关于预训练中强化学习的论文,但我认为没有人会在生产中使用。目前只是玩具示例。但总的来说,强化学习后训练更像是技能解锁,而预训练本质上是吸收知识。

So, pre-training is the classic training, one next token prediction at a time. You have a big corpus of data. And Nathan also has very interesting insights there because a big portion of the paper focuses on the right data mix. So, pre-training is essentially just training cross-entropy loss on next token prediction on a vast corpus of internet data, books, papers, and so forth. It has changed a little bit over the years. People used to throw in everything they could. Now, it's not just raw data, it's also synthetic data where people rephrase certain things. Synthetic data doesn't necessarily mean purely AI-made-up data. It's also taking something from an article, Wikipedia article, and then rephrasing it as a Q&A question or summarizing it, rewording it, and making better data that way. I think it's like if someone reads a book compared to a messy Reddit post, you learn better. There's going to be a post about this. Some Reddit data is very common and excellent for training. You just have to filter it. I think it's like if someone took that and rephrased it in a more concise and structured way, it's higher quality data that gets the LM maybe the same outcome but gets there faster. It trains faster because if the grammar and punctuation are correct, it already learns the correct way versus getting information from a messy way and then learning later how to correct that. So I think that is how pre-training evolved. While scaling still works, it's not just about amount of data. It's also the tricks to make that data better for you. Then mid-training is I mean it used to be called pre-training. It's called mid-training because it was awkward to have pre-training and post-training but nothing in the middle. So mid-training is usually similar to pre-training but more specialized. It's the same algorithm but you focus on long context, for example. The reason you don't do that during pure pre-training is because you don't have that many long context documents. So you have a specific phase. One problem of LMs is also catastrophic forgetting. You teach it something and it forgets other things. It's like no free lunch. It's the same with humans. If you ask me some math I learned 10 years ago, I don't know. Nathan was actually saying that he's consuming so much content that there's a catastrophic forgetting issue. I'm like trying to learn so much about AI and it's like I was learning about pre-training parallelism. I'm like I lost something and I don't know what it was. I don't want to anthropomorphize LLMs, but I think it's the same in that sense how humans learn. Quantity is not always better. Being selective is key. Mid-training is being selective in terms of quality content so the last thing the LLM has seen is the quality stuff. Then post-training is all the fine-tuning and supervised fine-tuning, DPO, reinforcement learning with verifiable rewards, with human feedback, and so forth. The refinement stages. It's also interesting the cost thing. Pre-training you spend a lot of money. RL a bit less. RL doesn't really teach knowledge; it's more like unlocking the knowledge. It's more like skill learning, how to solve problems with the knowledge from pre-training. There are actually three papers this year or last year 2025 on RL for pre-training, but I don't think anyone does that in production. Toy examples for now. But to generalize, RL post-training is more like the skill unlock where pre-training is like soaking up the knowledge essentially.

预训练的合成数据与数据提取 Synthetic data and data extraction for pre-training

Host

有几件事可能对大家有帮助。很多人认为合成数据对训练模型不好。你提到了像 deep sea 几乎 OCR 也就是光学字符识别论文。很多实验室都做过。AI2 有一个。有好几个,每个实验室做这些的原因是网上有大量的 PDF 和其他数字文档,它们的格式不容易编码成文本。所以,你使用这些几乎 OCR 或 deep sea OCR,我们称之为几乎 OCR,来提取可能数万亿 token 的候选数据用于预训练。预训练数据集的大小是万亿级别的,以万亿 token 衡量。研究人员的小模型可能是 5 到 10 万亿。Qwen 有记录显示达到 50 万亿,有传言说这些闭源实验室可以达到 100 万亿 token。而仅仅获取这些潜在数据,我认为他们有一个非常大的漏斗,实际用于训练模型的数据只是其中的一小部分。比如,字符识别数据在实验室中会被描述为预训练的合成数据。还有像 ChatGPT 现在给出的精彩答案,你可以用那些最佳答案来训练,那也是合成数据。这与早期 ChatGPT 有很多幻觉数据的情况非常不同,当时人们开始依赖合成数据。

A few things that could be helpful for people. A lot of people think of synthetic data as being bad for training the models. You mentioned like the deep sea got almost OCR which is optical character recognition paper. A lot of labs did. AI2 had one. Like had multiple and the reason that each of these labs have these is because there's vast amounts of PDFs and other digital documents on the web that are in formats that aren't encoded with text easily. So, you use these almost OCR or these deep sea OCR and we called our almost OCR to extract what could be trillions of tokens of candidate data for pre-training. And pre-training data set size is on the order of trillions, measured in trillions of tokens. Smaller models from researchers can be something like 5 to 10 trillion. Qwen is documented going up to like 50 trillion and there's rumors that these closed labs can go to like 100 trillion tokens. And just getting this potential data to put in, I think they have a very big funnel and then the data you actually train the model on is a small percentage of this. Like the character recognition data would be described as synthetic data for pre-training in a lab. And then there's also the things like ChatGPT now gives wonderful answers and you could train on those best answers and that's synthetic data. It's very different than like early ChatGPT lots of hallucinations data when people became grounded in synthetic data.

数据质量与算力 Data Quality vs. Compute

Host

一个有趣的问题是,如果我没记错的话,Olmo 3 的训练数据比某些其他开放权重模型(甚至可能比 Olmo 2)还要少,但性能却更好,这可能是数据如何发挥作用的例子之一。

One interesting question is, if I recall correctly, Olmo 3 was trained with less data than specifically some other open weight models, maybe even Olmo 2, but you still got better performance and that might be one of the examples how the data helped.

Nathan Lambert

这主要归功于数据质量。我认为如果我们有更多算力,我们会训练更长时间。最终我们会把这视为我们想做的事情,尤其是对于大模型,你需要更多算力,因为我们谈论的是更多参数和知识,本质上存在一个比例,大模型可以从数据中吸收更多,从而获得更多收益。这就像其中一个道理。任何对数图在你脑海中都是这样:如果你衡量 token 的趋势,小模型会更早趋于平稳,而更大的模型需要更多。但主要是,我们现在在 AI2 并没有训练那么大的模型,而获取尽可能高质量的数据是自然的起点。

It's mostly down to data quality. I think if we had more compute we would train for longer. I think we'd ultimately see that as something we would want to do, and especially with big models you need to have more compute because we talk about having more parameters and we talk about knowledge, and essentially there's a ratio where big models can absorb more from data and then you get more benefit out of this. It's like one of these. Any logarithmic graph in your mind is like a small model will level off sooner if you're measuring trends in tokens, and bigger models need more. But mostly, we aren't training that big of models right now at AI2, and getting the highest quality data we can is the natural starting point.

提升数据质量 Improving Data Quality

Host

关于数据质量这个话题,有什么可以说的吗?是否还有一些容易改进的地方?

Is there something to be said about the topic of data quality? Is there some low hanging fruit there still where the quality could be improved?

Nathan Lambert

这就像转动曲柄。我认为历史上在开放领域,有一个经典的预训练最佳数据集,它随着谁拥有最新的、最好的或最近的成果而移动。比如 AI2 的 Dolma 很早就有了第一个版本,Hugging Face 有 FineWeb,还有 DCLM 项目(DataComp for Language Models)。DataComp 也用于其他机器学习项目,它们有非常强大的数据集。很大程度上,互联网正在变得相当封闭。所以我们有 Common Crawl,我认为它有数百亿亿个 token,你过滤它,这看起来像很多科学工作:你训练分类器,并基于如何将这个数据集修剪成最高质量且适合你任务的内容来做决策。以前,语言模型更多地在知识和对话类事情上测试,但现在它们被期望做数学和代码。所以为了训练推理模型,你需要重新混合整个数据集。这里有很多非常棒的科学方法:你可以拿你的巨大数据集,从不同来源采样很多非常小的样本。比如你有 GitHub、Stack Exchange、Reddit、Wikipedia。你可以从它们中采样小样本,在每个混合上训练小模型,并在你的评估上测量它们的性能,然后你可以做基本的线性回归,得到你的最优数据集。但如果你的评估变了,你的数据集也会变化很大。所以 Olmo 3 的很多工作是引入新来源,以便在推理上更好地处理数学和代码。然后你执行这个混合过程,它给你答案。我认为这就是今年很多实验室发生的事情:有新的热门事物,无论是编码环境还是网页导航,你只需要引入新数据,你需要改变整个预训练,以便你的后训练能更好地工作。诸如此类。所以这就是不断重新演进和重新确定他们为模型关心什么的过程。

It's like turning the crank. So I think historically in the open there's been like a canonical best pre-training data set that has moved around between who has the most recent one or the best or the recent effort. Like AI2's Dolma was very early with the first Dolma, and Hugging Face had FineWeb, and there's a DCLM project which stands for DataComp for Language Models. There's been DataComp for other machine learning projects and they have had a very strong data set. And a lot of it is the internet is becoming fairly closed off. So we have Common Crawl which I think is hundreds of trillions of tokens, and you filter it and it looks like a lot of scientific work where you're training classifiers and making decisions based on how do you prune down this data set into the highest quality stuff and the stuff that suits your task. So previously language models were tested a lot more on like knowledge and just kind of conversational things, but now they're expected to do math and code. So to train a reasoning model you need to remix your whole data set. And there's a lot of actually wonderful scientific methods here where you can take your gigantic data set, you sample a lot of really tiny things from different sources. So you say you have GitHub, Stack Exchange, Reddit, Wikipedia. You can sample small things from them and you train small models on each of these mixes and measure their performance on your evaluations, and you can just do like basic linear regression and it's like here's your optimal data set. But if your evaluations change, your data set changes a lot. So a lot of Olmo 3 was new sources for reasoning to be better at math and code. Then you do this mixing procedure and it gives you the answer. And I think that's a lot of what's happened at labs this year: there's new hot things whether it's like coding environments or web navigation, and you just need to bring in new data, you need to change your whole pre-training so that your post-training can work better. And stuff like that. So that's like the constant re-evolution and the re-determining of what they care about for their models.

意外的高质量数据源 Unexpected High-Quality Data Sources

Host

有没有什么有趣的轶事,关于哪些数据源质量特别高,是我们意想不到的?你提到 Reddit 有时可以是一个来源。

Are there fun anecdotes of what sources of data are particularly high quality that we wouldn't expect? You mentioned Reddit sometimes can be a source.

Nathan Lambert

Reddit 非常有用。我认为 PDF 绝对是其中之一。

Reddit was very useful. I think that PDFs are definitely one.

Host

或者特别是存档。

Or especially archive.

Nathan Lambert

是的,比如 AI2 长期运营 Semantic Scholar,它是 Google Scholar 的竞争对手,功能更多。为此,AI2 发现并爬取了很多开放获取论文的 PDF,这些论文可能不在某些出版商的付费墙后面。所以是真正开放的学术 PDF。如果你拥有所有这些并处理它们,你可以从中获得价值。我认为很多这类工作前沿实验室更早就做了。这就像你需要一个相当熟练的研究人员,理解事物如何改变模型,他们引入数据并清理它。这是大量的劳动,我认为在很多前沿实验室,当他们扩大研究人员规模时,更多的工作投入到数据上。如果你加入一个前沿实验室并想产生影响,最好的方法就是找到更好的新数据。而那些花哨的算法性东西,比如弄清楚如何实现某个突破,是科学家最性感的想法。比如‘哦,我搞定了如何扩展强化学习。’确实有团队做到了,但我认为大多数贡献是在数据方面。我要让数据更好,或者我要让基础设施更好,这样团队里的每个人都能把实验跑快 5%。

Yeah, so like AI2 has run Semantic Scholar for a long time, which is a competitor to Google Scholar with a lot more features. And to do this, AI2 has found and scraped a lot of PDFs for openly accessible papers that might not be behind the closed paid garden of a certain publisher. So, truly open scientific PDFs. And if you sit on all of these and you process it, you can get value out of it. And I think that a lot of that style of work has been done by the Frontier Labs much earlier. And it's just like you need to have a pretty skilled researcher that understands how things change models and they bring it in and they clean it. And it's a lot of labor that I think at a lot of Frontier Labs when they scale researchers, a lot more goes into data. You have people like if you want to be, if you join a Frontier Lab and you want to have an impact, the best way to do it is just find new data that's better. And then the fancy glamorous algorithmic things like figuring out how to make a one is like the sexiest thought of a scientist. It's like 'oh, I figured out to scale RL.' And there's a group that did that, but I think most of the contributions are on the data side. I'm going to make the data better or I'm going to make the infrastructure better so that everybody in my team can run experiments 5% faster.

训练数据的保密与许可 Secrecy and Licensing of Training Data

Host

同时,我认为出于法律原因,训练数据也是被严密保守的秘密之一。所以,我认为也有很多工作是在隐藏训练数据的内容。比如试图让模型不泄露来源,因为,是的,出于法律原因。

At the same time, I think it's also one of the closest guarded secrets what your training data is for legal reasons. And so, there's also I think a lot of work that goes into hiding what your training data was essentially. So, like trying the model to not give away the sources because yeah, for legal reasons.

Nathan Lambert

另外要补充的是,有些人试图只使用有许可的数据进行训练,而 Common Crawl 是对整个互联网的抓取。所以,如果我托管多个网站,我很乐意让它们训练语言模型,但我没有明确许可其使用规则。因此,Common Crawl 的许可在很大程度上是未授权的,这意味着你并没有真正同意数据的使用方式。另一种想法是,只使用明确获得许可的数据来训练语言模型,这样就有了管理合同。我不确定 Apertis 是版权问题还是许可问题。我知道他们这样做是为了符合欧盟的合规要求,确保他们的模型通过相关检查。

The other thing to be complete is that some people are trying to train on only licensed data where Common Crawl is a scrape of like the whole internet. So, if I host multiple websites, I'm happy to have them train language models, but I'm not explicitly licensing what governs it. And therefore, this license the Common Crawl is largely unlicensed, which means that your consent really hasn't been provided for how to use the data. There's another idea where you can train language models only on data that has been licensed explicitly so that the kind of governing contract is provided. And I'm not sure if Apertis is the copyright thing or the license thing. I know that the reason that they did it was for an EU compliance thing where they want to make sure that their model fit one of those checks.

Host

另外,关于这一点,例如,许可也有区别。所以有些人,就像你说的,他们只是购买许可。比如他们在网上买一本书,比如亚马逊 Kindle 书或某种付费书,然后将其用于训练数据。这就像灰色地带,因为你为内容付了钱,你可能想用它训练。但也有一些限制,即使这样也不允许。所以这就有点模糊了,是的,我认为这现在仍然是一个热门话题,而且像 OpenAI 这样的大公司,他们接触私人公司获取其专有数据。

And also on that note, also for example, there's also the distinction between the licensing. So, some people, like you said, they just purchase the license. Let's say they buy a book online let's say on Amazon Kindle book or let's say money book or something and then use that in the training data. And that is like the gray zone because you paid for the content and you might want to train it. But then there are also restrictions where even that shouldn't be allowed. And so that is like where it gets a bit fuzzy and yeah, I think that is right now still a hot topic and also big companies like OpenAI, they approached private companies for their proprietary data.

数据护城河与领域缩放 Data as Moat and Domain-Specific Scaling

Nathan Lambert

而且私营公司会越来越保护自己的数据,因为他们知道:‘好吧,这将在几年内成为我的护城河。’我认为这是一个有趣的问题:如果 LLM 变得更加商品化,那么很多人会了解 LLM,会有更多人能够训练 LLM。当然,存在基础设施方面的挑战,但如果你想想制药、法律、金融等大型行业,我认为他们最终会从其他前沿实验室聘请人员,利用他们的专有数据构建内部模型,这将是预训练的又一次解锁,目前还无法实现,因为即使你想,也拿不到那些数据。大多数时候你无法获取临床试验这类数据。所以,我认为如果看特定领域的应用,Scaling 在这方面可能仍然很有活力,因为我们现在看到的 ChatGPT、Anthropic 等只是通用 LLM。它们只是通用型的,甚至还没有触及到 LM 如果真正针对特定任务进行训练和设计所能做到的表面。

And private companies, they become more and more protective of their data because they know, 'Okay, this is going to be my moat in a few years.' And I do think that's the interesting question: if LLMs become more commoditized, then a lot of people will learn about LLMs and there will be a lot more people able to train LLMs. Of course, there are infrastructure challenges, but if you think of big industries like pharmaceutical, law, finance, I do think they at some point will hire people from other frontier labs to build their in-house models on their proprietary data, which will be another unlock with pre-training that is currently not there because even if you wanted to, you can't get that data. You can't get access to clinical trials most of the time and these types of things. So, I do think scaling in that sense might still be pretty much alive if you also look at domain-specific applications because we are still right now just looking at general-purpose LLMs on ChatGPT, Anthropic, and so forth. They are just general-purpose. They are not even scratching the surface of what an LM can do if it is really specifically trained and designed for a specific task.

Host

关于数据这件事,我觉得这是 2025 年发生但我们完全忘记的事情之一:Anthropic 在法庭上败诉,需要向作者支付 15 亿美元。Anthropic 好像买了数千本书并扫描了它们,这部分被法律认定为合法,因为他们买了书,这算是走正规渠道。但另一方面,他们也通过种子下载了一些书,我认为正是这个种子下载行为让法院判定他们需要向作者支付这笔数十亿美元,这真是一起令人难以置信的诉讼,就这么来了又走了。那可是风投生态里的一大笔钱。

I think on the data thing, this is one of the things where I like this happened in 2025 and we totally forget it: Anthropic lost in court and was owed 1.5 billion dollars to authors. Anthropic I think bought thousands of books and scanned them and was cleared legally for that because they bought the books and that is kind of going through the system. And then the other side they also torrented some books and I think this torrenting was the path where the court said that they were then culpable to pay this billions of dollars to authors, which is just like such a mind-boggling lawsuit that kind of just came and went. Like that is so much money from the VC ecosystem.

Nathan Lambert

这些法庭案件将定义人类文明的未来,因为很明显数据驱动了很多东西,而且存在着非常复杂的人类矛盾。我的意思是,你可以理解:你也是作者。在某种程度上,你把自己的心血和汗水都倾注在写作中,别人用你的数据训练却不给你署名,感觉有点像偷窃。

These are court cases that will define the future of human civilization because it is clearly that data drives a lot of this and there's this very complicated human tension. I mean you can empathize: you're both authors. And there's some degree to which you put your heart and soul and your sweat and tears into the writing that you do, it feels a little bit like theft for somebody to train your data without giving you credit.

Host

而且就像 Nathan 说的,这也有两个层面。有人可能买了书然后训练,这可以争论公平与否,但还有那些直接使用盗版书的公司,甚至没有补偿作者。我认为这正是人们特别生气的地方。

And then like Nathan said also two layers to it. Someone might buy the book and then train on it, which could be argued fair or not fair, but then there are the three straight up companies who use pirated books where it's not even compensating the author. That is I think where people got a bit angry about it specifically.

Nathan Lambert

必须要有某种补偿方案。这正在朝着类似 Spotify 流媒体最初对音乐的做法发展。你知道,这种补偿是什么样的?你必须定义这些模型。你必须仔细考虑所有这些问题。另一件我认为人们普遍好奇的事情是:随着 LLM 使用得越来越多,如果你看看 archive 甚至 GitHub,越来越多的数据是由 LLM 生成的。在那个世界里你该怎么办?这个问题有多大?

There has to be some kind of compensation scheme. This is moving towards something like Spotify streaming did originally for music. You know, what does that compensation look like? You have to define those kinds of models. You have to think through all of that. One other thing I think people are generally curious about: as LLMs are used more and more, if you look at even archive but GitHub more and more of the data is generated by LLMs. What do you do in that kind of world? How big of a problem is that?

Host

最大的问题是基础设施和系统,但从 AI 的角度来看,这几乎是不可避免的。

Largest problem is infrastructure and systems, but from an AI point of view it's kind of inevitable.

Nathan Lambert

所以这基本上是由人类策划的 LLM 生成数据,对吧?

So it's basically LLM generated data that's curated by humans essentially, right?

Host

是的,而且我认为很多开源贡献者确实在 burnout。如果你有一个流行的开源仓库,有人会说:‘哦,我想做开源 AI。这对我的职业有好处。’然后他们随便写点代码就扔进去……你可能比我遇到更多这种情况。

Yes, and I think that a lot of open source contributors are legitimately burning out. If you have a popular open source repo, somebody's like, 'Oh, I want to do open source AI. It's good for my career.' And they just vibe code something and throw it into the... You might get more of this than I would do.

Nathan Lambert

我这里有一个案例。我有一个叫 ML-Extend 的仓库,是我 15 年前、10 年前还是学生时开发的。它对于某些算法仍然相当流行,尤其是频繁数据挖掘之类的东西。最近有两三个人在很短的时间内提交了很多 PR。我确实认为 LLM 参与了这些 PR 的提交。作为维护者,我有两点感受:首先,我有点不知所措。我没有时间通读,尤其是这是一个较老的库,不是我的优先事项。同时,我也有些感激,因为我认为人们忘记了一点:不仅仅是使用 LLM,还有人类层来验证某些东西。这在某种意义上也是数据标注的方式,对吧?这是 RLHF 阶段最昂贵的事情之一——获取标注数据。这有点像那样,经过几个阶段后,你实际上得到了更高质量的数据。所以,在某种意义上我不介意。它可能让人感到不知所措,但我确实认为它也有价值。

So I have a case study here. I have a repository called ML-Extend that I developed as a student 15 years ago, 10 years ago. And it is a reasonably popular library still for certain algorithms, especially frequent data mining stuff. And there were recently I think two or three people who submitted a lot of PRs in a very short amount of time. I do think LLMs have been involved in submitting these PRs. Me as the maintainer, two things: First, I'm a bit overwhelmed. Like I don't have time to read through it because especially it's an older library that is not a priority for me. At the same time I kind of also appreciate it because I think something people forget is it's not just using the LLM. There's still a human layer that verifies something. And that is in a sense also how data is labeled, right? So that's like one of the most expensive things is getting labeled data for RLHF phases. And this is kind of like that where it goes through phases and then you get actually higher quality data out of it. You know, so I don't mind it in a sense. It can feel overwhelming, but I do think there is also value in it.

Host

感觉原始 LLM 生成的数据和有人类参与的 LLM 生成数据之间存在根本区别,人类会进行某种验证,即使这种验证只涉及代码行的一小部分。

It feels like there's a fundamental difference between raw LLM generated data and LLM generated data with human in the loop that does some kind of verification, even if that verification is a small percent of the lines of code.

Nathan Lambert

我认为这适用于任何事情,人们有时也会想:‘哦,是的,我可以用 LLM 来学习 XYZ。’这没错,你可以。但是,可能有一个专家,他可能用 LLM 写了非常具体的代码。这里面有人类的工作,让它变得更好,去掉不好的部分,为你预先消化好,这节省了你的时间。我认为这就是价值所在:有人帮你过滤,甚至正确使用 LLM。我认为这仍然是你免费获得的劳动,比如你读一篇文章,比如 Substack 的文章。我或许可以让 LLM 给我一些观点,但我甚至可能不知道问什么。我认为读那篇文章仍然有价值,相比我去问 LLM,因为你是专家,你选择了真正准确、应该包含的知识,并给了我一个执行摘要。这是巨大的价值,因为现在我不必自己花 3 到 5 个小时去读,可能还会得到一些错误信息等等。所以,我认为这也是作家未来的方向,即使有 LLM,专家也能节省你的时间。

I think this goes with anything where people think also sometimes, 'Oh, yeah, I can just use an LLM to learn about XYZ.' Which is true, you can. But, there might be a person who is an expert who might have used an LLM to write so specific code. There is kind of like this human work that went into it to make it nice and throwing out the not-so-nice part to make it kind of pre-digest it for you, and that saves you time. And I think that's the value add where you have someone filtering things or even using the LLMs correctly. I think this is still labor that you get for free when you, for example, read an article, let's say Substack article. I could maybe ask an LLM to give me opinions on that, but I wouldn't even maybe know what to ask. And I think there is still value in reading that article compared to me going to the LLM because you are the expert, you select what knowledge is actually spot-on, should be included, and you give me this executive summary. And this is kind of huge value add because now I don't have to waste 3 to 5 hours to go through this myself, maybe get some incorrect information, and so on. And so, I think that's also where the future still is for writers, even though there are LLMs that expert can kind of save your time.

LLM 摘要中的声音与洞察 Voice and Insight in LLM Summaries

Host

比如,它从内容中移除了什么信号?

Like, what is the signal it removes from the thing?

Nathan Lambert

我经常谈论的是“声音”。

The voice is what I talk about a lot.

Host

声音。嗯,我很想听听你说的“声音”是什么意思。这确实很有力量。但有时候,这实际上就是洞察力。比如,移除一个洞察,你实际上从根本上改变了事物的含义。所以,我一直在失望,LLM 在真正抓住核心洞察方面有多差,而这正是好的总结所做的。是的,即使我用了那些非常详尽、极其精巧的提示词,试图挖掘洞察,它仍然不太到位。嗯,我的意思是,这涉及一个深刻的哲学问题:什么是人类知识和智慧,什么是有洞察力,等等。但当你谈到“声音”时,你指的是什么?

Voice. Well, voice I'd love to hear what you mean by voice. That's really powerful. But sometimes there's like literally insights. Like, in removing an insight, you're actually fundamentally changing the meaning of the thing. So, I'm continuously disappointed how bad LLMs are at really getting to the core insights, which is what a great summary does. Yeah, even if you go and I have these extensive, extremely elaborate prompts where I'm like really trying to dig for the insights, and it's still not quite there, which um I mean, that's a whole deep philosophical question about what is human knowledge and wisdom, and what does it mean to be insightful, and so on. But, when you talk about the voice, what do you mean?

Nathan Lambert

所以,当我写作时,我认为我试图做的是把研究者思考的东西——那是非常原始的——表达出来。研究者试图捕捉他们理解前沿的一个想法,并试图把一种感觉转化为文字。我认为我的写作尝试做到这一点,这使得它显得原始,但信息量也很大,有些人能理解,有些人不能,这有点像研究的本质。我认为这是语言模型做不好的地方。特别是,它们都经过基于人类反馈的强化学习(RLHF)训练,这种训练旨在从很多人那里获取反馈,并在某种程度上平均模型的行为。我认为,当存在这种过滤器时,模型很难变得非常敏锐。我认为这对 RLHF 的研究人员来说是一个绝妙的基本问题。它在让模型变得更好方面提供了很多效用,但问题表述本身就像一个解不开的结。所以,这就是我认为语言模型在其深层表达中缺乏这种先验的原因。我不认为这是不可能做到的。有一些模型的故事确实让人震惊。比如,我很想试试 Bing Sydney。它是否更有“声音”?因为它经常出轨,而且历史上显然以一种可怕的方式。比如,告诉记者离开他的妻子,这是一个疯狂的模型,可能无法广泛采用。但这是一种权衡。这个 RLHF 过程是否在某种程度上增加了限制?

So, when I write, I think a lot of what I'm trying to do is take what you think as a researcher, which it's very raw. Which a researcher is trying to encapsulate an idea at the frontier of their understanding, and they're trying to put what is a feeling into words. And I think that my writing, I tried to do this as the writing, which makes it come across as raw, but also high information in a way that is like some people will get it and some won't, and that's kind of the nature of research. And I think this is something that language models don't do well. Particularly, they're all trained with this reinforcement learning from human feedback, which is designed to take feedback from a lot of people, and in a way average how the model behaves from this. And I think that there's it's going to be hard for a model to be very incisive when there's that sort of filter in it. And I think this is kind of a wonderful fundamental problem for researchers in RLHF. It's like this provides so much utility in making the models better, but also the problem formulation is kind of like there's this knot in it that you can't get past. So, that's what I think of as like these language models don't have this prior in their deep expression that they're trying to get at. I don't think it it's impossible to do. I think there's stories of models that really shock people. Like, I think of like I would love to have tried Bing Sydney. And does like does that have more voice? Cuz it would so often go off the rails on people and it would what is historically obviously a scary way. Like, telling a reporter to leave his wife is a crazy model to potentially put in general adoption. But, that's kind of like a trade-off. Like is this RLHF process like in some ways adding limitations?

Host

作为这些前沿实验室和公司之一,这是一个可怕的位置。因为数百万人正在使用它们。

It's a terrifying place to be as one of these frontier labs and companies. Because millions of people are using them.

Nathan Lambert

去年 GPT-4 被移除时有很多反弹。我个人从未使用过那个模型,但我与 OpenAI 的人聊过,他们到了这样的地步:用户会在半夜检测到部署中的细微差异,然后发邮件给他们,说“我的朋友不一样了”。他们甚至找到这些员工的邮箱,给他们发东西,因为他们如此依恋这个部署给用户的模型权重配置。我们在 TikTok 上也看到这一点。你打开它——我不用 TikTok——但据说大约 5 分钟内算法就了解你了。它就像锁定了一样。我不喜欢那些语言模型做推荐的方式。我认为你可以用语言模型在 5 分钟的聊天内做到这一点。模型就了解你了。这是人们还没有真正准备好的事情。我认为,比如,不要把它给孩子,至少在我们知道发生了什么之前。

There was a lot of backlash last year with the GPT-4 getting removed. And I personally never used the model, but I've talked to people at OpenAI where they're to the point where they like get emails from users that might be detecting subtle differences in the deployments in the middle of the night and they email them and they're like, "My friend is different." And they like find these people employees emails and send them things because they are so attached to this set What is this set of model weights in a configuration that is deployed to the users? We see this with TikTok. You open it I I don't use TikTok. But supposedly in like 5 minutes the algorithm gets you. It's like it's locked in. And I don't like those are language models doing recommendations. Like I think there are ways that you can do this with a language model within like 5 minutes of chatting with it. The model just gets you. And that is something that people aren't really ready for. Like I think that if kid like don't give that to kids like don't give that to kids at least until we know what's happening.

Host

随着 LLM 被越来越多地使用,还会出现这种机制。不幸的是,人类状况的本质是有人会自杀。记者们会广泛报道自杀的人,并且很可能将其与 LLM 联系起来,因为他们有对话数据。如果你在生活中挣扎,如果你抑郁,如果你考虑自杀,你可能会和 LLM 谈论这些。记者们会说:“嗯,自杀是因为 LLM。”这会导致公司出于法律问题等原因,越来越多地削弱 LLM 的锋芒。所以它会变得尽可能通用。在这个领域运营非常困难,因为你当然不希望 LLM 在那个层面上对人类造成伤害。但这也是人类体验的本质:进行丰富的对话、充实的对话,一个挑战你并让你成长的对话。你需要那种锋芒。这对于 RLHF 领域的 AI 研究人员来说是一个非常困难的问题。因为你实际上是在处理人类状况。

Also going to be this mechanism what's going to happen with these LLMs as they're used more and more. Unfortunately, the nature of the human condition is such that people commit suicide. And so what journalists would do is they will report extensively on the people who commit suicide and they would very likely link it to the LLMs because they have that data about the conversations. If you're really struggling in your life, if you're depressed, if you're thinking about suicide, you're going to probably talk to LLMs about it. And so what journalists will do is they will say, "Well, the suicide was committed because of the LLM." And that's going to lead to the companies because of legal issues and so on more and more and more taking the edge off of the LLM. So it's going to be as generic as possible. It's so difficult to operate in this space because of course you don't want an LLM to cause harm to humans at that level. But also this is also the nature of the human experience is to have a rich conversation, a fulfilling conversation, one that challenges you and from which you grow. You need that edge. And that that's something extremely difficult for AI researchers on the RLHF front to actually have to solve. Cuz you're actually dealing with the human condition.

Nathan Lambert

这些公司的很多研究人员动机都非常好。他们,比如 Anthropic 和 OpenAI,在文化上非常希望通过这个为世界做好事。但这是如此……我不想做这个。因为一方面,很多人把 AI 视为健康盟友,一个可以秘密谈论健康的人。但然后它渗透到谈论心理健康等方面,令人心碎的是,这可能会成为某人崩溃的导火索。但其他人可能会被拯救。我不喜欢这样。作为训练模型的研究人员,有些事我不想做,比如我不想训练图像生成模型并公开释放,因为我不想让某人在他们的笔记本电脑上拥有一个可以伤害他人的工具。我的公司没有安全地做到这一点的基础设施。但有很多这样的领域,需要人们以复杂性和信念去处理,这真是一个难题。

Like a lot of researchers at these companies are so well motivated. And they definitely they like some Anthropic and OpenAI are culturally so want to do good through this for the world. And there is it's such a I'm like I don't want to work on this. Because on the one hand a lot of people see AI as a health ally, as somebody they can talk to about their health confidentially. But then it bleeds all the way into this like talking about mental health and things where at this it's heartbreaking that this will push like be the thing where somebody goes over the edge. But other people might be saved. And I'm like I don't like there's things that as a researcher training models it's like I don't want to train image generation models and release them openly cuz I don't want to enable somebody to have a tool on their laptop that can harm other people. Like I don't have the infrastructure at my company to do that safely. But it's like like there's a lot of the areas like this where it's just it needs people that will approach it with the complexity and kind of conviction of like it's just such a hard problem.

Host

但作为社会,作为这些技术的用户,我们也需要确保我们进行复杂的讨论,而不是仅仅散布恐惧。大型科技公司对人类造成伤害或窃取你的数据,诸如此类。事情比那更复杂。你说得对。这些公司内部有非常多的人,你认识很多,我也认识很多,他们深切关心帮助他人。他们在考虑全世界人们的完整人类体验,不仅仅是硅谷,而是全美国、全世界的人们,这意味着什么,他们的需求是什么。

But also we as a society as users of these technologies need to make sure that we're having the complicated conversation about it versus just fear mongering. Big tech is causing harm to humans or stealing your data, all that kind of stuff. There is more complicated than that. And you're right. There's a very large number of people inside these companies, many of which you know, many of which I know, that deeply care about helping people. They are considering the full human experience of people from across the world, not just Silicon Valley, people across the United States, people across the world, what that means, what their needs are.

为不同用户设计 AI Designing AI for diverse users

Nathan Lambert

设计一个能够帮助所有不同年龄、文化、心理状态和状况的人的系统真的很难。

It's really difficult to design this one system that is able to help all these different kinds of people across different age groups, cultures, mental states, mental conditions, all that kind of stuff.

科技巨头与 AI 声誉 Big Tech and AI reputation

Host

我希望 AI 的时机与大型科技公司与普通人的关系不同。大型科技公司的声誉如此之低,而 AI 又如此昂贵,它不可避免地会成为大型科技公司的事情,需要大量资源,人们说美国正在把经济押注在 AI 上。这些因素交织在一起,使得沟通环境变得非常困难。对我来说,去和世界上那些讨厌大型科技公司、将 AI 视为其延续的人多聊聊会很好。

I wish that the timing of AI was different with the relationship of Big Tech to the average person. So, like Big Tech's reputation was so low, and with how AI is so expensive, it's like inevitably going to be a Big Tech thing where it takes so many resources, and people say that the US is betting the economy on AI with this buildout. And it's like to have these be intertwined at the same time is just makes for such a hard communication environment. It'd be good for me to go talk to more people in the world that hate Big Tech and see AI as a continuation of this.

在 AI 中寻找自主性 Finding agency in AI

Nathan Lambert

你实际推荐的一件事,你谈到的一种解药,是在整个系统中找到能动性,而不是被动地坐着,无力地消费 AI 垃圾,因为 AI 正在迅速占领互联网。更重要的是,通过使用 AI 来构建东西、构建应用、构建能帮助你建立直觉的东西来找到能动性,但第二,这是赋权的,因为你可以理解它是如何工作的,弱点是什么,它让你有能力说:这搞砸了,这很糟糕,这是技术的糟糕使用,这是技术的良好使用。然后你更深入地融入系统,所以你能更好地理解它,并在它发展时更好地引导它。

And one of the things you actually recommend, one of the antidotes that you talk about, is to find agency in this whole system, as opposed to sort of sitting back in a powerless way and consuming the AI slop as it quickly, rapidly takes over the internet. More, find agency by using AI to build stuff, build apps, build something that actually helps you build intuition, but two, it's empowering because you can understand how it works, what the weaknesses are, and it allows you gives you a voice power to say like this is fucked up, this is bad, this is bad use of the technology, and this is good use of the technology. And you're more plugged into the system then, so you can understand it better, and you can steer it better as it goes.

自主性与忽视 AI Agency vs ignoring AI

Host

我认为你提出的能动性观点很好,而不是忽视它并说“好吧,我不打算用它”。我认为长期来看更健康的方式是说,“好吧,它就在那里。我无法把它放回去。”你知道,就像互联网、电脑刚出现时一样。我如何最好地利用它,它如何帮助我提升自己?但我担心的一点是,如果你完全用它来做你喜欢的事情,那么你喜欢的事情就不再存在了,这可能会导致倦怠。例如,如果我使用语言模型为我做所有的编码,那么就没有编码了。我只是在管理一个为我编码的东西。假设两年后,如果我每天做 8 小时,让东西为我编码,我还会感到满足吗?这会不会损害我对工作的热情,对我所做的事情的热情?我还会为构建东西感到自豪吗?

I think that's a good point you brought up agency instead of ignoring it and saying, "Okay, I'm not going to use it." I think it's probably long-term healthier to say, "Okay, it's out there. I can't put it back." You know, like internet, computers back then when they came out. How do I make best use of it, and how does it help me to up-level myself? The one thing I worry here though is like if you just fully use it for something you love to do, then the thing you love to do is not no longer there and that could potentially I feel like lead to burnout. For example, if I use an LM to do all my coding for me, now there's no coding. I'm just managing something that is coding for me. 2 years let's say later if I just do that 8 hours a day, have something coded for me, do I feel fulfilled still like is this like Yeah, I mean is this just like hurting me in terms of being excited about my job, excited about what I'm doing? Am I still proud to build something?

开发者对 AI 的享受度调查 Survey on developer enjoyment with AI

Nathan Lambert

关于享受这个话题,很有趣,我们应该提一下,最近有一项针对约 791 名专业开发者的调查,专业意味着 10 年以上经验。那是一段很长的时间。

So there's on that topic of enjoyment is quite interesting we should just throw this in there that there is this recent survey of about 791 professional developers, professional meaning 10 plus years of experience. That's a long time.

Host

是的。

Yeah.

Nathan Lambert

什么是初级开发者?

What's a junior developer?

Host

嗯,在这个时代。

Uh yeah in this day and age.

Nathan Lambert

嗯,结果在很多方面令人惊讶。他们按初级和高级开发者细分,但这表明初级和高级开发者都在他们发布的代码中使用 AI 生成的代码。所以这不仅仅是娱乐性的中间学习。这是他们发布的代码。平均 25%,大多数人使用大约 50%或更多。有趣的是,对于发布的代码中超过 50%是 AI 生成的类别,高级开发者更有可能这样做。但你不希望 AI 夺走你热爱的东西。

Uh so the results here in many fronts are surprising. So they break it down by junior and senior developers, but I mean it just shows that both junior and senior developers use AI generated code in code they ship. So this is not just for fun sort of intermediate kind of learning things. This is code they ship. And so it's 25% mean like most of them use around 50% or more. And what's interesting is for the category of over 50% of your code that you ship is AI generated, senior developers are much more likely to do so. But you don't want AI to take away the thing you love.

Host

是的。

Yeah.

Nathan Lambert

我认为这符合我的经验,我即将说的这些特定结果。大约 80%的人认为使用 AI 作为工作的一部分要么稍微更愉快,要么显著更愉快。

I think it speaks to my experience these particular results I'm about to say. So together about 80% of people find it either somewhat more enjoyable or significantly more enjoyable to use AI as part of the work.

平衡享受与 AI 使用 Balancing enjoyment and AI use

Host

我认为这取决于任务。就我个人使用而言,我有一个网站,我有时会调整网站上的东西。我个人不喜欢这个。所以,从这个意义上说,如果 AI 能帮助我在网站上实现一些东西,我完全赞成。这很棒。但同时,当我解决一个复杂的问题时,比如有一个 bug,我追踪这个 bug 并找到它,那是世界上最好的感觉。就像你得到了很多快乐,感觉很好。但现在,如果你甚至不去思考 bug,直接去找语言模型,那么你永远不会有这种感觉,对吧?但可能存在中间地带:你自己尝试,找不到,然后使用语言模型,这样你不会感到沮丧,因为它帮助你,然后你继续做你喜欢的事情。所以,我认为看这些统计数据,差异在于或者没有考虑到的是,平均了所有不同的场景,我们不知道它是用于核心任务还是用于人们本来不会喜欢的平凡任务。所以,从某种意义上说,AI 非常适合做那些需要大量工作的平凡事情。例如,前几天我妻子有一个关于书籍讨论的播客,一个读书俱乐部,她正在把节目笔记从 Spotify 转移到 YouTube。然后链接不知何故坏了。在一些剧集中,因为他们讨论很多书,有大约 100 个链接。手动进去修复每个链接会很痛苦。所以我建议,“嘿,我们试试 ChatGPT。”我们把文本复制到 ChatGPT 中,它修复了它们。而不是花 2 小时逐个链接修复,它让那种工作变得无缝得多。没有挫折感。我认为每个人都有 AI 有用的用例,用于那些非常无聊、非常平凡的事情。

I think it depends on the task. For my personal usage, for example, I have a website where I sometimes tweak things on the website. I personally don't enjoy this. So, in that sense, if the AI can help me to implement something on my website, I'm all here for it. It's great. But then, at the same time, when I solve a complex problem, well, if there's a bug and I hunt this bug and I find the bug, it's the best feeling in the world. It's like you get so much joy, like oh, it's like you feel like great. But now, if you don't even think about thinking about the bug, you just go directly to the LLM, well, you never have this kind of feeling, right? But then, there could be the middle ground where, well, you try yourself, you can't find it, you use the LLM, and then you don't get frustrated because it helps you and you move on to something that you enjoy. And so, I think looking at these statistics, I think also the difference is or what is not factored in is averaging over all the different scenarios where we don't know if it's for the core task or if it's for something mundane that people would not have enjoyed otherwise. So, in a sense, AI is really great for doing mundane things that take a lot of work. So, for example, my wife the other day, she has like a podcast for like book discussions, a book club, and she was like transferring the show notes from Spotify to YouTube. And then the links somehow broke. And she had in some episodes, because they discuss many books, like 100 links or something. It would have been really painful to go in there and fix each link manually. And so, I suggested, "Hey, let's try ChatGPT." We copied the text into ChatGPT and it fixed them. And instead of 2 hours going from link to link fixing that, you know, it made that type of work much more seamless. There was no frustration. I think everyone has a use case where AI is useful for something like that that would be really boring, really mundane.

AI 作为结对程序员 AI as a pair programmer

Nathan Lambert

就我个人而言,既然我们在谈论编码,你提到了调试,我很多乐趣的来源,更多是在 Cursor 这边而不是 Claude Code 这边,是因为我有一个朋友。我有一个……叫什么来着?结对编程伙伴。就像,它不那么孤独。你把调试描述成巨大的快乐。不,我会说调试就像在沙漠中走了几天后喝到水。所以,你跳过了整个受苦的沙漠部分。所以,有时有一个朋友很好,他可能找不到 bug,但能给你一些关于代码的直觉,你和那个朋友一起穿越沙漠,然后一起找到那杯水。所以,至少对我来说,这可能说明了编程体验的孤独。那是快乐的来源。

I, for me personally, since we're talking about coding, and you mentioned debugging, I would a lot of the source of the enjoyment for me on the more on the cursor side than the Claude code side is the I have a friend. I have a co- What's that called? A pair programmer. Like I it's less lonely. You made debugging sound like this great joy. No, I would say debugging is like a drink of water after you've been going through a desert for days. So, like you skip the whole desert part where you're suffering. So, like there sometimes it's nice to have a friend who can't really find the bug, but can give you some intuition about the code and you're together with that friend going through the desert and then together find that drink of water. So, I at least for me maybe it speaks to the loneliness of the programming experience. It's that that is a source of joy.

延迟满足与 AI Delayed gratification and AI

Host

这可能也与延迟满足有关。我是一个人,你知道,即使小时候,我喜欢圣诞礼物的想法,拥有它们,得到它们,比实际得到礼物更好。我会期待得到礼物的那一天,但然后它结束了,我很失望。也许这也类似于,比如说,食物。

It's maybe also related to delayed gratification. I'm a person who, you know, as a even as a kid, I like the idea of Christmas presents having them, getting them better than actually getting the presents. I would look forward to the day I get the presents, but then it's over and I'm disappointed. And maybe it's something like also with, let's say, food.

AI 学习的黄金区间 Goldilocks zone of learning with AI

Nathan Lambert

我觉得饿的时候食物更好吃。调试代码呢,不总是那么美好,常常让人沮丧。但如果你能解决它,那就很棒了。这里也有一个恰到好处的“金发姑娘区”。如果太难,那就是浪费时间。但另一个挑战是:人们要如何学习?我们看到的图表显示,资深开发者比初级开发者生成了更多 AI 代码。我觉得这很有趣,因为直觉上你会认为是初级开发者,因为他们还不会做那些事。这可能意味着 AI 还不够好,无法解决那些任务,但也可能意味着专家更善于使用它。他们知道在何处以及如何使用它,并审查代码,而且更信任生成的代码。所以未来社会的一个问题是:如果你从不自己尝试,你怎么成为专家?对我来说,我通过自己尝试来学习。就像数学课本,如果你看答案,你学到一些东西,但如果你先尝试,你会学得更好,因为你能把它纳入你的思维框架。如果 LLM 一直存在,你还会经历挣扎的过程吗?你愿意挣扎吗?挣扎并不愉快。如果你用 LLM 做所有事,总有一天你再也无法迈出下一步,你可能无法获得专家使用 LLM 时的那种突破。所以有一个“金发姑娘”最佳点:也许诀窍是每天留出两小时离线学习,其余时间用 LLM。但我认为人们仍然需要投资自己,而不是凡事都用 LLM。

I think food tastes better when you're really hungry. And with debugging, it's not always great. It's often frustrating. But if you can solve it, then it's great. There's also a sweet Goldilocks zone. If it's too hard, then it's wasting your time. But that is another challenge: how will people learn? The chart we looked at showed that more senior developers are shipping more AI-generated code than junior ones. I think it's very interesting because intuitively you would think it's the junior developers since they don't know how to do the thing yet. It could mean the AI is not good enough to solve that task, but it could also mean experts are more effective at using it. They know where and how to use it and review the code, and they trust the code more. So one issue in society in the future will be: how do you become an expert if you never try to do the thing yourself? For me, I learn by trying things myself. Like math textbooks, if you look at the solutions, you learn something, but you learn better if you try first, and then you appreciate the solution differently because you can put it into your mental framework. If LLMs are here all the time, would you go through the length of struggling? Would you be willing to struggle? Struggle is not nice. If you use the LLM to do everything, at some point you will never take the next step, and you might not get that unlock that an expert gets using an LLM. So there's a Goldilocks sweet spot where maybe the trick is to make dedicated offline time where you study 2 hours a day, and the rest of the day use LLMs. But I think it's important for people to still invest in themselves, not just LLM everything.

Host

是的,在我们这个脆弱的文明中,我们每个人都必须找到那个“金发姑娘区”。

Yeah, there is in our weak together civilization that we each individually have to find that Goldilocks zone.

Nathan Lambert

是的。

Yeah.

后训练与 RLVR Post-training and RLVR

Host

我们进行了一场精彩的对话,从预训练和中训练开始。现在谈谈后训练。后训练中有很多有趣的东西。那么,后训练中有哪些有趣的想法呢?

We've had this fascinating conversation that started with pre-training and mid-training. Let's get to post-training. A lot of fun stuff in post-training. So, what are some of the interesting ideas in post-training?

Nathan Lambert

2025 年最大的进展是学习这种带有可验证奖励的强化学习(RLVR)。你可以扩展那里的训练,这意味着进行大量这种迭代生成-评分循环,这让模型在工具使用和软件方面学习到有趣的行为。这可能包括搜索、自行运行命令并查看输出,而且这种训练还能很好地实现推理时扩展。结果证明,这种范式与推理时扩展紧密相连——这种 RL 训练实现了推理时扩展,但推理时扩展本可以通过其他方式发现。所以这就像一场完美风暴,模型训练方式的巨大变化是主要因素,这极大地改变了人们处理后训练的方式。

The biggest one from 2025 is learning this reinforcement learning with verifiable rewards. You can scale up the training there, which means doing a lot of this kind of iterative generate grade loop, and that lets the models learn both interesting behaviors on the tool use and software side. This could be searching, running commands on their own and seeing the outputs, and then also that training enables this inference time scaling very nicely. And it just turned out that this paradigm was very nicely linked in this, where it's this kind of RL training enables inference time scaling, but inference time scaling could have been found in different ways. So, it was kind of this perfect storm of the models change a lot in the way that they're trained is a major factor in doing so, and this has changed how people approach post-training dramatically.

Host

RLVR 是由 DeepSeek R1 推广的吗?你能描述一下它是如何工作的吗?

RLVR popularized by DeepSeek R1? Can you describe how it works?

Nathan Lambert

是的,有趣的是,我是提出 RLVR 这个术语的团队成员之一,这来自我们 2003 年 DeepSeek 之前的工作。我们并不因推广 Scaling RL 而居功,但作为学者,有趣的一点是能够命名并影响讨论,因为封闭实验室能说的有限,而作为学者,你或许没有算力训练模型,但你可以用某种方式构建框架,最终让社区围绕 RLVR 这个术语凝聚起来,这很有趣。然后 DeepSeek 实现了训练突破,他们扩展了强化学习:让模型生成答案,然后根据正确性评分,这个准确率就是强化学习的奖励。强化学习经典上是一个在环境中行动的智能体,环境返回状态和奖励,你要最大化奖励。对于语言模型,奖励通常是可验证任务上的准确率,比如数学题、编程任务,然后扩展到事实领域等模糊地带,这些在某种程度上也是可验证的,或者指令约束如“只用 A 开头的词回答”。所有这些在某种程度上都是可验证的,核心思想是找到更多这类可验证问题,让模型在 RL 梯度更新中多次尝试。基础设施从基于人类反馈的强化学习(RLHF)演变而来,在那个时代,他们优化的分数是一个学习到的奖励模型,代表人类偏好的聚合。所以你改变了问题领域,让优化扩展到更大规模,这引发了模型能力和使用方式的重大变革。

Yeah, fun fact, I was on the team that came up with the term RLVR, which is from our 2003 work before DeepSeek, which is we don't take a lot of credit for being the people to popularize the scaling RL, but as fun as what academics get as an aside, is the ability to name and influence the discourse, because closed labs can only say so much, but one of the things you can do as an academic is you might not have the compute to train the model, but you can frame things in a way that ends up being I describe it as like a community can come together around this RLVR term, which is very fun. And then DeepSeek is the people that did the training breakthrough, which is they scaled the reinforcement learning, which was you have the model generate answers and then grade the completion if it was right, and then that accuracy is your reward for reinforcement learning. So, reinforcement learning is classically an agent that acts in an environment, and the environment gives it a state and a reward back, and you try to maximize this reward. In the case of language models, the reward is normally accuracy on a set of verifiable tasks, whether it's math problems, coding tasks, and it starts to get blurry with things like factual domains, like that is also in some ways verifiable, or constraints on your instruction like respond only with words that start with A. All of these things are verifiable in some way and the core idea of this is you find a lot more of these problems that are verifiable and you let the model try it many times while taking these RL steps these RL gradient updates. The infrastructure evolved from this reinforcement learning from human feedback where in that era the score they were trying to optimize was a learned reward model of aggregate human preferences. So you kind of change the problem domains and that let the optimization go on to much bigger scales which kind of kick-started a major change in what the models can do and how people use them.

Host

RLVR 适用于哪些领域?

What kind of domains is RLVR amenable to?

Nathan Lambert

数学和代码是著名的领域,然后还有很多关于所谓“评分标准”的工作,这与人们可能听过的“LM 作为评判者”相关。比如,对于训练数据集中的每个问题,我会用另一个语言模型询问一个好的答案应该是什么样的,然后你可以反复尝试这个问题,并根据这个评分标准打分。所以它不一定像数学和代码那样可验证,但评分标准这个想法以及其他更模糊的科学问题,正是人们关注的重点,他们试图将这些方法推广到更开放的领域,让模型学到更多。

Math and code are the famous ones and then there's a lot of work kind of on what is called a rubrics which is related to word people might have heard as LM as a judge which is like for each problem I'll have a set of problems in my training data set. I'll then have another language model and ask it what would a good answer to this problem look like and then you could try the problem a bunch of times over and over again and assign a score based on this rubric. So it's not necessarily verifiable like a math and code domain but this rubrics idea and other scientific problems that it might be a little bit more vague is where a lot of the attention is where they're trying to push this set of methods into these kind of more open-ended domains where the models can learn a lot more.

Host

我想那叫做基于 AI 反馈的强化学习(RLAIF),对吧?

I think that's called reinforcement learning with AI feedback, right?

Nathan Lambert

那是更早的术语,由 Anthropic 的《Constitutional AI》论文提出。所以很多事情都是循环出现的。

That's the older term from it that was coined in Anthropic's Constitutional AI paper. So it's like a lot of these things come in cycles.

Host

再回到 RLVR。我觉得有趣而美妙的地方在于,你问 LM 一个数学问题,你知道正确答案,然后你让 LM 像你说的那样去解决,但你怎么做呢?你并没有太多约束它。

Also just one step back for the RLVR. So I think the interesting beautiful thing here is that you ask the LM let's say a math question and then you know the correct answer and you let the LM like you said figure it out but how it does it I mean you don't really constrain it much.

推理缩放与思维链 Inference Scaling and Chain-of-Thought

Nathan Lambert

你可以添加一些约束,比如使用同一种语言,不要中英文混用,但基本上你可以放手不管。你只给出问题和答案,然后语言模型必须自己得出正确答案。但美妙之处在于实际发生的情况:语言模型会逐步描述,就像学生或数学家推导解决方案一样。它会使用这些步骤,这有助于模型提高自身准确性。然后,就像你提到的,推理扩展。推理扩展大致意味着在推理过程中花费更多算力。这里,模型会使用更多词元。在 R1 论文中,他们展示了训练时间越长,模型回答越长。回答随时间增长,使用更多词元,因此对于简单任务来说成本更高,但这些解释有助于提高准确性。还有一些论文显示,模型解释的内容不一定正确,甚至可能与答案无关,但出于某种原因它仍然有帮助。关键在于它进行了解释。我不想将 LLM 拟人化,但这有点像人类运作的方式:对于复杂的数学问题,你会用草稿纸逐步推导,划掉错误,模型也会自我纠正。这就是 R1 论文中的‘顿悟时刻’——模型意识到自己犯了错误,然后说‘啊,我搞错了,再试一次’。仅仅通过给出正确答案并让它自己摸索,就能出现这种现象,这太酷了。虽然 LLM 不像人类那样思考,但这是一种有趣的巧合。另一个好的副作用是,对我们人类来说,看到这些步骤很有帮助——它建立了信任,也让我们可以复核。

There are some constraints you can add like use the same language, don't switch between Spanish and English, but let's say you're pretty much hands-off. You only give the question and the answer, and then the LM has to arrive at the right answer. But the beautiful thing here is what happens in practice: the LM will do a step-by-step description, like a student or a mathematician would derive the solution. It will use those steps, and that helps the model improve its own accuracy. Then, like you said, inference scaling. Inference scaling loosely means spending more compute during inference. Here, the model uses more tokens. In the R1 paper, they showed the longer they train the model, the longer the responses are. They grow over time, using more tokens, so it becomes more expensive for simple tasks, but these explanations help with accuracy. There are also papers showing that what the model explains doesn't necessarily have to be correct, or may be unrelated to the answer, but for some reason it still helps. It's the fact that it is explaining. I don't want to anthropomorphize LLMs, but it's like how humans operate: for a complex math problem, you use scratch paper and do it step-by-step, cross things out, and the model also self-corrects. That was the 'aha moment' in the R1 paper—the model recognized it made a mistake and said, 'Ah, I did something wrong, let me try again.' It's cool that this falls out of just giving it the correct answer and having it figure out how to do it. Although LLMs don't think like humans, it's an interesting coincidence. Another nice side effect is that it's great for humans to see these steps—it builds trust and allows us to double-check.

Host

这里有很多内容。我认为今年的一些争论在于这些‘顿悟时刻’是否有点虚假,因为在预训练中,模型基本上已经看过了整个互联网,包括人们解释自己工作的内容,比如数学讲座的转录。强化学习(RLVR)非常擅长放大这些行为,因为它们有助于模型思考更久并检查自己的工作。我同意,模型学会以这种方式放大这些行为,从而改善最终答案,这很美妙。

There's a lot in here. I think some of the debate this year is whether these 'aha moments' are kind of fake because in pre-training, the model has essentially seen the whole internet, including people explaining their work, like transcripts of math lectures. Reinforcement learning (RLVR) is very good at amplifying these behaviors because they're useful for enabling the model to think longer and check its work. I agree it's beautiful that the model learns to amplify this in a way that improves final answers.

Nathan Lambert

我可以给你一个实际例子。我用 RLVR 在 Math 500 上训练 Qwen 3 基础模型。基础模型的准确率大约 15%。仅仅 50 步,几分钟内,模型准确率从 15%提升到 50%。你不能说它在这么短时间内学到了任何关于数学的根本知识。

I can give you a hands-on example. I was training the Qwen 3 base model with RLVR on Math 500. The base model had an accuracy of about 15%. Just 50 steps, in a few minutes with RLVR, the model went from 15% to 50% accuracy. You can't tell me it's learning anything fundamentally about math in that time.

Host

Qwen 的例子很奇怪,因为今年有两篇论文(其中一篇我参与了)讨论了 Qwen 中的数据污染。具体来说,他们在某个特殊的中期训练阶段训练了大量与测试集几乎相同的问题。

The Qwen example is weird because there have been two papers this year, one of which I was on, that talk about data contamination in Qwen. Specifically, they trained on a lot of problems almost identical to that test set during a special mid-training phase.

Nathan Lambert

没错。所以你可以看到,强化学习并没有教给模型任何关于数学的新知识。50 步内不可能做到。知识已经在预训练中有了;你只是在解锁它。

Exactly. So you can see that the RL is not teaching the model any new knowledge about math. You can't do that in 50 steps. The knowledge is already there in pre-training; you're just unlocking it.

Host

我仍然不同意这个前提,因为存在奇怪的复杂性。例如,如果你拿 Qwen 3 基础模型,输入一个数学问题,所有这些问题都有文字。如果你改变数字但保留文字,Qwen 会在不使用工具的情况下对答案的十进制表示产生非常高的准确率。这意味着它在某个时候被展示了与测试集几乎相同的问题。研究界一直存在很大争论:这些基于 Qwen 训练并在该数学基准上测量的强化学习论文,有多少可信度?这导致 RLVR 被认为只是关于格式化的,因为你能这么快获得提升,暗示它已经在模型中了。但这里有很多复杂性;这并非真正的受控实验。

I still disagree with the premise because there are weird complexities. For example, if you take the Qwen 3 base model and put a math problem into it, all these problems have words. If you change the numbers but keep the words, Qwen will produce a very high accuracy on the decimal representation of the answer without tools. This means it was shown problems almost identical to the test set at some point. It's been a big debate in the research community: how much of these reinforcement learning papers training on Qwen and measuring on this math benchmark can you believe? This caused the reputation of RLVR being about formatting, because you can get gains so quickly, implying it's already in the model. But there's a lot of complexity; it's not really controlled experimentation.

Nathan Lambert

是的。所以我们真的不知道。但如果不是这样,蒸馏就不会起作用,对吧?蒸馏在一定程度上有效,但 LLM 研究中最大的问题就是这种污染,因为我们不知道数据里有什么。除非你有一个新的数据集,否则真的不可能。数学数据集也是如此,它包含问题、答案和解释。即使是像 MMLU 这样更简单的多项选择基准:如果你稍微改变格式——比如用点代替括号——模型准确率就会大不相同。

Yep. So we don't really know. But if it weren't true, distillation wouldn't work, right? Distillation can work to some extent, but the biggest problem in LLM research is this contamination because we don't know what's in the data. Unless you have a new dataset, it's really impossible. The same goes for Math dataset, which has a question, answer, and explanation. Even something simpler like MMLU, a multiple-choice benchmark: if you change the format slightly—like using a dot instead of a parenthesis—the model accuracy will vastly differ.

Host

我认为那可能是模型问题,而不是普遍问题。这不是开发者的恶意行为;只是模型在某个时候看到了某些东西。评估 LLM 的唯一公平方式是在 LLM 部署的截止日期之后使用新的基准。

I think that could be a model issue rather than a general issue. It's not malicious by the developers; it's just that the model has seen something at some point. The only fair way to evaluate an LLM is to have a new benchmark after the cut-off date when the LLM was deployed.

后训练配方与 RLVR Post-Training Recipe and RLVR

Host

我们能否列出后训练中涉及的所有要素的配方?你提到 RLVR 是一个非常令人兴奋且有效的东西。也许我们应该详细说明。基于人类反馈的强化学习(RLHF)仍然扮演着重要角色。后训练还有哪些其他想法?

Can we lay out what would be the sort of recipe of all the things going into post-training? You mentioned RLVR was a really exciting effective thing. Maybe we should elaborate. RLHF still has an important component to play. What other ideas are there on post-training?

Nathan Lambert

我认为你可以按顺序来。

I think you can kind of take this in order.

中训练与推理轨迹 Mid-training and reasoning traces

Nathan Lambert

我认为你可以把它看作是让 o1(第一个推理模型)成为可能的原因,或者是最新模型会是什么。在这些阶段,实际上有类似的干预措施,从中间训练开始。据传,让 o1 和类似模型成为可能的关键是精心策划的数据集,提供大量所谓的推理轨迹。这仅仅是模型在正向过程中生成文字,反映出将问题分解为中间步骤并尝试解决的过程。所以在中间训练阶段,你需要有这样的数据,这样当你进入主要使用可验证奖励的后训练阶段时,模型才能学习。

I think you could view it as what made o1, which was this first reasoning model possible, or what will the latest model be. They actually have similar interventions at these stages where you start with mid-training. The thing that is rumored to enable o1 and similar models is really careful data curation, where you're providing a broad set of what are called reasoning traces. That's just the model generating words in a forward process that reflects breaking down a problem into intermediate steps and trying to solve them. So at mid-training, you need to have data similar to this to make it so that when you move into post-training, primarily with verifiable rewards, it can learn.

Nathan Lambert

那么现在要做的是,弄清楚给模型哪些问题,训练多长时间,以及在解决这些可验证问题时允许模型使用多少推理能力。随着模型变得更好,某些问题不再有挑战性——模型会 100% 解决它们,因此信号非常少。如果我们看 GRPO 方程,它在这方面很出名,本质上,给智能体的奖励是基于某个动作或完成相对于同一问题其他答案的好坏程度。所以如果所有问题都得到相同的答案,这类算法就没有信号了。所以他们正在寻找更难的问题,这就是为什么你会听到像科学领域这样的东西——那太难了,在那里做对任何事都不容易。如果你有一个实验室之类的,它生成了大量的 token,或者更难的软件问题。所以前沿模型都在向这些更难的领域推进,他们可以在更多问题上训练,模型也会一次性学到更多技能。

Then what is happening today is you're figuring out which problems to give the model and how long you can train it for, and how much inference you can enable the model to use when solving these verifiable problems. So as models get better, certain problems are no longer challenging—the model will solve them 100% of the time, and therefore there's very little signal. If we look at the GRPO equation, which is famous for this, essentially the reward given to the agent is based on how good a given action or completion is relative to the other answers to the same problem. So if all the problems get the same answer, there's no signal in these types of algorithms. So what they're doing is finding harder problems, which is why you hear about things like scientific domains—that's so hard, getting anything right there. If you have a lab or something, it just generated so many tokens, or much harder software problems. So the frontier models are all pushing into these harder domains, and they can train on more problems, and the model will learn more skills at once.

RLHF 作为收尾 RLHF as finishing touch

Nathan Lambert

RLHF 与此的联系是,RLHF 一直并且仍然是模型的收尾工作,通过改进组织、风格或语气使模型更有用。不同的东西能引起不同受众的共鸣。比如有些人喜欢非常古怪的模型,RLHF 可以很好地实现这种个性。有些人讨厌模型做的 markdown 项目符号列表,但它实际上对于快速解析信息非常好。而 RLHF,这个人类反馈阶段,在一天结束时非常适合将这些注入模型。所以这就是为什么 ChatGPT 对人们来说如此神奇。而且这种用途实际上一直相当稳定。这种格式化也可以帮助模型在数学问题上做得更好,例如。所以风格、格式和回答问题的方法之间的界限在训练这些模型时实际上是紧密相连的,这就是为什么 RLHF 仍然可以让模型在数学上更好,但这些可验证领域是更直接的过程,因为它与问题公式更吻合,这就是为什么它们最终会融合在一起。

The RLHF link to this is kind of like RLHF has been and still is the finishing touch on the models, where it makes the models more useful by improving the organization, style, or tone. There are different things that resonate with different audiences. Like some people like a really quirky model, and RLHF could be good at enabling that personality. And some people hate the markdown bulleted list thing that the models do, but it's actually really good for quickly parsing information. And RLHF, this human feedback stage, is really great for just putting this into the model at the end of the day. So it's what made ChatGPT so magical for people. And that use has actually remained fairly stable. This formatting can also help the models get better at math problems, for example. So the border between style and formatting and the method that you use to answer a problem is actually very closely linked when you're training these models, which is why RLHF can still make a model better at math, but these verifiable domains are a much more direct process for doing this because it makes more sense with the problem formulation, which is why it all ends up forming together.

Nathan Lambert

总结一下,中间训练给模型提供了学习所需的技能。带有可验证奖励的强化学习让模型尝试很多次——所以把大量算力投入到跨难题的试错学习中。然后 RLHF 就像是完成模型,使其易于使用,并让模型变得全面。

To summarize, mid-training gives the model the skills it needs to then learn. RL with verifiable rewards lets the model try a lot of times—so put a lot of compute into trial and error learning across hard problems. And then RLHF would be like finishing the model, making it easy to use, and kind of rounding the model out.

RLVR 的算力需求 Compute requirements for RLVR

Host

你能评论一下 RLVR 所需的算力吗?

Can you comment on the amount of compute required for RLVR?

Nathan Lambert

它只增不减。我认为 Grok 4 以声称在预训练和后训练中使用相似数量的算力而闻名。回到 Scaling 的讨论,它们涉及非常不同的硬件。预训练是计算密集型的,是关于浮点运算——一次能完成多少矩阵乘法?而因为强化学习你要生成这些答案,在真实环境中测试模型,它最终变得非常内存密集,因为你生成长序列,注意力机制有这样的行为:随着序列变长,内存呈二次方增长。所以算力变得非常不同。

It's only gone up and up. I think Grok 4 was famous for saying they use a similar amount of compute for pre-training and post-training. Back to the scaling discussion, they involve very different hardware for scaling. Pre-training is very compute-bound, which is about flops—how many matrix multiplications can you get through in one time? And because RL you're generating these answers, you're trying the model in real-world environments, it ends up being much more memory-bound because you're generating long sequences, and the attention mechanisms have this behavior where you get a quadratic increase in memory as you get to longer sequences. So the compute becomes very different.

Nathan Lambert

所以在预训练中,我们会谈论一个模型——如果我们回到拜登政府的行政命令,训练一个模型需要 10^25 次浮点运算。如果你在后训练中使用浮点运算,那就更奇怪了,因为现实只是你分配了多少小时,多少 GPU。我认为在时间方面,强化学习的算力正在接近,因为你不能把所有东西都放在一个系统中。预训练计算密度极高,所有 GPU 都在相互通信,效率极高。而强化学习有所有这些移动部件,生成 10 万个 token 的序列可能需要很长时间。比如如果你认为 GPT-5.2 Pro 需要一小时,那么如果你的训练运行有一个样本需要一小时,你必须确保它被高效处理。所以我认为在 GPU 小时或挂钟小时方面,强化学习运行的天数可能接近预训练,但它们可能不会同时使用那么多 GPU。

So in pre-training, we would talk about a model—if we go back to the Biden administration executive order, it's like 10^25 flops to train a model. And if you're using flops in post-training, it's a lot weirder because the reality is just how many hours you are allocating, how many GPUs for. And I think in terms of time, the RL compute is getting much closer because you just can't put it all into one system. Pre-training is so computationally dense where all the GPUs are talking to each other, and it's extremely efficient. Where RL has all these moving parts, and it can just take a long time to generate a sequence of 100,000 tokens. Like if you think about GPT-5.2 Pro taking an hour, it's like what if your training run has a sample that takes an hour, and you have to make it so that's handled efficiently. So I think in GPU hours or just wall clock hours, the RL runs are probably approaching the number of days as pre-training, but they probably aren't using as many GPUs at the same time.

Nathan Lambert

实验室里有经验法则,比如你不希望预训练运行持续超过一个月,因为它们会灾难性地失败。如果你计划一个巨大的集群运行两个月,然后在第 50 天失败,机会成本太大了。所以你不想把所有鸡蛋放在一个篮子里,GPT-4 就是终极的 YOLO 运行,以前没人想这么做,它花了大约 3 个月训练,每个人都惊讶它成功了。所以我认为现在人们更加谨慎和渐进。所以 RLVR 可以说是无限的,你可以训练多少仍然受益,而 RLHF,因为它是一种偏好调整,你达到某个点后,再花更多强化学习预算就没有意义了。

There are rules of thumb where in labs it's like you don't want your pre-training runs to last more than a month because they fail catastrophically. And if you were planning a huge cluster to be held for 2 months, and then it fails on day 50, the opportunity cost is just so big. So you kind of don't want to put all your eggs in one basket, which is like GPT-4 was the ultimate YOLO run, and nobody ever wanted to do it before where it took like 3 months to train, and everybody was shocked that it worked. So I think people are a little bit more cautious and incremental now. So RLVR is more, let's say, unlimited how much you can train to still get benefit, where RLHF, because it's a preference tuning, you reach a certain point where it doesn't really make sense to spend more RL budget on that.

Nathan Lambert

退一步看偏好调整。多个人可以对同一件事给出多种解释,而且都可以是正确的,但在某个时候你学会了某种风格,再迭代就没有意义了。我最喜欢的例子是,如果亲戚问我应该买什么笔记本电脑,我会给他们解释或问他们,‘你的使用场景是什么?’比如他们优先考虑电池续航和存储。其他人,比如我们,会优先考虑内存和算力。所以两个答案都是正确的,但不同的人需要不同的答案。

Just a step back with preference tuning. There are multiple people that can give multiple explanations for the same thing and they can both be correct, but at some point you learn a certain style and it doesn't make sense to iterate on it. My favorite example is like if relatives ask me what laptop they should buy, I give them an explanation or ask them, 'Yeah, what is your use case?' Like they, for example, prioritize battery life and storage. Other people, like us for example, would prioritize RAM and compute. And so both answers are correct, but different people require different answers.

RLVR 与 RLHF 缩放 RLVR vs RLHF Scaling

Nathan Lambert

而对于偏好调优,你试图以某种方式取平均,比如你让数据标注员给出正确的——不,是偏好的答案,然后你在此基础上训练,但到了某个点,你学到的就是那个平均偏好答案,我认为没有理由继续训练更久,因为那只是风格问题。而使用 RLVR,你实际上是让模型解决越来越复杂、困难的问题。所以我认为长期来看,把更多预算分配给 RLVR 更合理。而且,现在我们处于 RLVR 1.0 阶段,仍然只是简单的问题和答案,没有处理中间过程。谷歌等机构有多篇研究论文关于过程奖励模型,它们也会给解释打分,看解释的正确性。我认为这将是下一步,比如今年的 RLVR 2.0,聚焦于问题和答案之间,如何利用解释信息来改进解释并提高准确性。这是一个角度。还有 DeepSeek Math 第二版论文,其中也有有趣的推理扩展,他们先开发了模型自我评分,用另一个模型来评分。我认为这将是一个方面,另一个方面是 Nathan 提到的,RLVR 将扩展到其他领域。

And with preference tuning, what you're trying to average somehow, like you are asking data labelers to give you the right, well not the right, the preferred answer, and then you train on that, but at some point yeah, you learn that average preferred answer and there's no, I think, reason to keep training longer on it because, you know, it's just a style. Well, with RLVR, you literally give the model well, you let the model solve more and more complex, difficult problems. And so, I think that it makes more sense to allocate more budget long-term to RLVR. And also, that right now we are in RLVR 1.0 land where it's still like that simple thing where we have a question and answer, but we don't do anything with the one stuff in between. So, there was a, I mean, multiple research papers also by Google for example on process reward models that also give scores for the explanation, how correct is the explanation. And I think that will be the next thing, let's say, RLVR 2.0 for this year, focusing in between question and answer, like how to leverage that information, the explanation to improve the explanation and help it to get better accuracy. But then, so that that's one angle and there was a DeepSeek math version two paper where they also had interesting inference scaling there where first they had um developed models that grade themselves a separate model and I think that that will be one aspect and the other like Nathan mentioned that will be for our RLVR branching into other domains.

Host

人们兴奋的地方是价值函数,这与过程奖励模型非常相似。过程奖励模型为推理过程中的每个中间步骤分配好坏评分,而价值函数则对语言模型生成的每个 token 赋予价值。这两者在当前推理模型时代的语言建模中大多未被证实。由于某种原因,现在人们对价值函数一直更乐观。我认为过程奖励模型在 01 之前的推理模型时代被尝试得更多,很多人犯了很多错误。所以我认为很大程度上是人性使然:价值模型在强化学习中有很深厚的历史,它们是深度强化学习存在的核心之一,比如训练价值模型。所以现在文献中人们对尝试价值模型很兴奋,但几乎没有证据,而且扩展过程奖励模型有负面例子。这些在未来不一定成立。我们是通过讨论 Scaling 来到这个话题的。简单总结你的观点:你不希望做太多 RLHF,因为最终信号会饱和。人们研究语言模型的 RLHF 多年,尤其在 ChatGPT 之后兴趣浓厚。而第一个用 RLVR 训练的推理模型 01 发布时,有一个扩展图:训练算力对数增加,评估结果线性提升,这已被多次复现。DeepSeek 也有类似图。但 RLHF 没有这样的缩放定律:对数增加算力不会带来性能提升。实际上,RLHF 的开创性缩放论文是关于奖励模型过度优化的缩放定律。所以 RLVR 与 RLHF 之间有一条明显的分界线。我们现在和未来的方法将遵循这种扩展范式:最好的运行可以额外增加 10 倍算力,获得几倍性能提升,但 RLHF 做不到。这将定义领域内人们如何处理它们。我一直在学术界推广 RLHF,这是个好描述。要做最好的 RLHF,你可能不需要额外 10 倍或 100 倍算力,但要做最好的 RLVR,你需要。所以我认为有一篇来自 Meta 实习生的开创性论文,叫《用语言模型扩展强化学习的艺术》。他们描述了一个框架叫 Scale RL,他们的增量实验用了大约 10,000 B200 小时,每个实验花费数千到数万美元,他们做了很多这样的实验。这种成本普通学术界无法承受,这是一个艰难的平衡,需要弄清楚如何从每个社区学习。

The place where people are excited are value functions which is very pretty similar so process reward models are kind of like process reward models assign how good something is to each kind of intermediate step in a reasoning process where value functions apply value to every token the language model generates. Both of these have been largely unproven in the language modeling in this reasoning model era. People are more optimistic about value functions forever for whatever reason now. I think process reward models were tried a lot more in this pre-01 pre-reasoning model era and a lot of people had a lot of mistakes with them. So I think a lot of it is the human nature of like value models have a very deep history in reinforcement learning. They're one of the first things that were core to like deep reinforcement learning existing is like training value models in this. So right now the literature people are excited about trying value models but there's very little proof in it and there are negative examples in trying to scale up process reward models. These things don't always hold in the future. I think we came to this discussion by talking about scaling and a simple way to summarize what you're saying with like you don't want to do too much RLHF which is eventually the signal scales is people have worked on RLHF for language models for years especially in intense interest after ChatGPT and this the first release of a reasoning model trained with RLVR opening eyes 01 had a scaling plot where if you increase the training compute logarithmically you get a linear increase in evaluations and this has been reproduced multiple times. I think DeepSeek had a plot like this but there's no scaling law for RLHF where if you log increase the compute, you get some performance. In fact, the seminal scaling paper for RLHF is scaling laws for reward model over optimization. So, it's like that's a big line to draw with RLVR and the methods we have now and in the future, like they will follow this scaling paradigm, which is like the best runs you can let to run for an extra 10x and you get a few x performance, but you can't do this with RLHF. And that is just going to be field defining in how people approach them, where I I've shilled for people academically to do RLHF, and that's a good way to describe it. It's like to do the best RLHF, you might not need the extra 10 or 100x of compute, but to do the best RLVR, you do. So, I think there's a what I say is a seminal paper from what was a Meta internship is called it's like the art of scaling reinforcement learning with language models. There what they describe as a framework is scale RL, and their incremental experiment was like 10,000 B200 hours, which is like thousands or tens of thousands of dollars per experiment, and they did do a lot of them, which is just like this cost is not accessible to the average academic, which is a hard equilibrium, where it's trying to figure out how to learn from each community.

教育与学习推荐 Education and Learning Recommendations

Host

我想知道我们能否稍微岔开话题,谈谈教育和学习。如果你是一个对编程和 AI 感兴趣的聪明听众,我想从头构建一些东西是个好的开始。你能告诉我你推荐人们怎么做吗?

I was wondering if we could take at this point a bit of a tangent and talk about education and learning. If you're somebody listening to this who's a smart person interested in programming, interested in AI, so I presume building something from scratch is a good beginning. So, can you just take me through like what you would recommend people do?

Nathan Lambert

我个人会从实现一个简单的模型开始,就像你说的,从零开始,能在你的电脑上运行。目标不是构建一个你每天用于个人项目的模型,它不会成为你的个人助手,取代现有的开源模型或 ChatGPT。而是看看 LLM 内部到底有什么,输出是什么,预训练在你的电脑上是如何工作的。然后你学习预训练、监督微调、注意力机制。你会扎实理解这些东西是如何运作的。但到某个点你会遇到限制,因为小模型能力有限。学习大规模 LLM 的问题在于,制作更大的模型复杂度呈指数级增长,因为模型不仅仅是变大。你还得考虑跨多个 GPU 分片参数。即使是 KV 缓存,也有多种实现方式。一种只是理解它如何工作,逐步增长缓存,比如通过拼接列表来增长。但这对 GPU 来说不是最优的,你不会那样做。你会预分配一个张量然后填充它。但这又增加了 20-30 行代码,每加一个东西就增加很多代码。我认为这本书的诀窍是理解 LLM 的工作原理。它不会是你的生产级 LLM,但一旦你理解了,你就能理解生产级 LLM。

So, I would personally start, like you said, uh implementing a simple model from scratch that you can run on your computer. The goal is not if you build a model from scratch to have like something you use every day for your personal projects. Like it's not going to be your personal assistant replacing an existing open weight model or ChatGPT. It's to see what exactly goes into the LLM, what exactly comes out of the LLM? How the pre-training works in that sense on your own computer preferably. And then you learn about the pre-training, the supervised fine-tuning, the attention mechanism. You get a solid understanding of how things work. But at some point you will reach a limit because small models can only do so much and the problem with learning about LLMs at scale is I would say it's exponentially more complex to make a larger model because it's not that the model just becomes larger. You have to now think about sharding your parameters across multiple GPUs. Even for the KV cache there are multiple ways you can implement it. One is just to understand how it works, just to grow the cache. That's It's like a cache you grow step by step by let's say concatenating lists, growing it. But then that wouldn't be optimal in GPUs. You wouldn't do that. You would pre-allocate a tensor and then fill it in. But that adds again another 20-30 line lines of code and for each thing you add so much code. And I think the trick with the book is basically to understand how the LLM works. It's not going to be your production level LLM. But once you have that, you can understand the production level LLM.

Host

所以你总是试图构建一个能放在单个 GPU 上的 LLM。

So you're trying to always build an LLM that's going to fit on one GPU.

Nathan Lambert

是的,大部分模型我都放在单个 GPU 上。我有一些关于 MoE 模型的额外材料,其中一两个可能需要多个 GPU,但目标是放在单个 GPU 上。美妙之处在于你可以自我验证。这几乎就像 RLVR,当你从头编码时,你可以从 Hugging Face Transformer 库中拿一个现有模型来对比。

Yes, the most of them I have the I have some bonus materials on some MoE models. I think one of two of them they may require multiple GPUs, but the goal is to have it on one GPU. And the beautiful thing is also you can self-verify. It's almost like RLVR when you code these from scratch, you can take an existing model from the Hugging Face Transformer library.

从 Transformers 库学习 Learning from Transformers Library

Nathan Lambert

Hugging Face 的 Transformers 库很棒,但如果你想学习 LLM,我认为那不是最好的起点,因为代码太复杂了。它必须适配很多用例,有些人还在生产环境中使用它。它必须非常复杂,而且代码交织在一起,很难线性阅读。

The Hugging Face Transformers library is great, but if you want to learn about LLMs, I think that's not the best place to start because the code is so complex. It has to fit so many use cases, and some people use it in production. It has to be really sophisticated, and it's really intertwined and hard to read linearly.

Host

它最初是一个微调库,后来发展成为每个模型架构及其加载方式的标准表示。所以 Hugging Face 是获取模型的默认地方,而 Transformers 是让人们轻松加载模型并做一些基本操作的软件。

It was started as a fine-tuning library and then grew to be the standard representation of every model architecture and the way it is loaded. So Hugging Face is the default place to get a model, and Transformers is the software that enables people to easily load a model and do something basic with it.

Nathan Lambert

所有拥有开放权重模型的前沿实验室都有 Hugging Face Transformers 版本,从 DeepSeek 到 GPT-OSS。那是你可以加载的规范权重。但同样,即使 Transformers 库也不用于生产。人们使用 SGLang 或 vLLM,这又增加了一层复杂性。

And all frontier labs that have open weight models have a Hugging Face Transformers version, from DeepSeek to GPT-OSS. It's the canonical weights you can load there. But again, even the Transformers library is not used in production. People use SGLang or vLLM, which adds another layer of complexity.

Host

我们应该说 Transformers 库有大约 400 个模型。

We should say that the Transformers library has like 400 models.

Nathan Lambert

所以这是一个试图实现很多 LLM 的单一库,因此代码库非常庞大。它很大,可能有数百万或数十万行代码。理解你想要的部分就像大海捞针。但美妙之处在于你有一个可工作的实现,所以你可以反向推导。我推荐的是,如果我想了解 Llama 3 是如何实现的,我会查看模型 hub 中的权重和配置文件。然后你可以看到,哦,他们用了这么多层,分组查询注意力或多头注意力,以及所有组件都在一个人类可读的 100 行配置文件中。然后你从你的 GPT-2 模型开始,添加这些东西。酷的是你可以加载预训练权重,看看它们是否在你的模型中工作。你想要匹配与 Transformers 模型相同的输出,你可以将其作为可验证的奖励来确保你的架构正确。有时这需要我一天时间;对于 Llama 3,挑战在于位置编码的 RoPE。他们有 YaRN 扩展和自定义缩放,我无法完全匹配。在这个挣扎中,你理解了东西,但最后你知道它是正确的,因为你可以针对参考实现进行单元测试。我认为这是最好的学习方法之一——逆向工程。

So it's a single library that tries to implement a lot of LLMs, so you have a huge codebase. It's huge, maybe millions or hundreds of thousands of lines of code. Understanding the part you want is like finding a needle in a haystack. But what's beautiful is you have a working implementation, so you can work backwards from it. What I recommend is, if I want to understand how Llama 3 is implemented, I look at the weights in the model hub and the config file. Then you can see, oh, they used so many layers, group query attention or multi-head attention, and all the components in a human-readable 100-line config file. Then you start with your GPT-2 model and add these things. The cool thing is you can load the pre-trained weights and see if they work in your model. You want to match the same output as the Transformers model, and you can use that as a verifiable reward to make your architecture correct. Sometimes it takes me a day; with Llama 3, the challenge was RoPE for position embeddings. They had a YaRN extension and custom scaling, and I couldn't quite match them. In this struggle, you understand things, but at the end you know it's correct because you can unit test against the reference implementation. I think that's one of the best ways to learn—reverse engineering something.

学习路径与研究机会 Learning Path and Research Opportunities

Host

我认为这是每个对 AI 感兴趣的人都应该做的。这就是为什么我喜欢你的书。我从强化学习和机器人领域进入语言模型,从未花时间学习所有基础。Transformer 架构就像过去的深度学习一样基础,人们需要学习它。很多人感到不知所措的是如何应用这些知识来产生影响或找到职业道路。AI 和语言模型让基础变得可及,有动力的人会学习它们。然后是如何获得计算资源来贡献研究。我相当乐观,因为领域发展如此之快,最优秀的人往往不会完全解决一个问题,因为还有更大、更容易解决的问题,所以他们继续前进。在我的 RLHF 书中,我试图描述后训练技术以及人们如何思考它们对模型的影响。令人惊讶的是,有多少东西人们只是不再研究。所以,在打好基础后,深入一个狭窄领域是好的。阅读相关论文,参与生态系统。普通人在线上与顶尖研究者的接近程度令人惊叹。没有人知道 X 上所有匿名账户是谁;他们可能只是深入研究的人。借助 AI 工具,你可以不断深挖。有很多研究领域只需要读三篇论文,其中一位作者可能会回复你的邮件,但你需要付出努力。对于新手来说,可能需要几周才能掌握一个狭窄领域,但在基础之后深入狭窄领域非常有用。我对角色训练产生了兴趣——如何让模型变得有趣、讽刺或严肃,以及如何处理数据。一位牛津的学生联系了我,我给了他建议。现在那篇论文存在了。世界上只有两三个人对此感兴趣。他是博士生,这有帮助,但对我来说,这是一个我一直等待有人花时间研究的主题。有很多狭窄的东西,比如“没有答案是没有道理的”。信息太多,人们无法抓住任何东西,但如果你坚持一个领域,有很多东西可以学。

I think that is something everyone interested in AI today should do. That's why I like your book. I came to language models from RL and robotics, and I never took the time to learn all the fundamentals. The Transformer architecture is as fundamental as deep learning was in the past, and people need to learn it. Where many get overwhelmed is how to apply this to have impact or find a career path. AI and language models make the fundamentals accessible, and motivated people will learn them. Then it's about how to get cycles on gold to contribute to research. I'm fairly optimistic because the field moves so fast that the best people often don't fully solve a problem because there's a bigger, lower-hanging fruit, so they move on. In my RLHF book, I tried to describe post-training techniques and how people think about them influencing the model. It's remarkable how many things people just stop studying. So, going narrow after fundamentals is good. Read relevant papers and engage with the ecosystem. The proximity random people have online to leading researchers is amazing. No one knows who all the anonymous accounts on X are; they could be random people who study deeply. With AI tools, you can keep digging. There are many research areas where you only need to read three papers, and one author will probably email you back, but you need to put in effort. For a newcomer, it might take weeks to grasp a narrow area, but going narrow after fundamentals is very useful. I became interested in character training—how to make a model funny, sarcastic, or serious, and what to do with the data. A student at Oxford reached out, and I advised him. Now that paper exists. There are only two or three people in the world interested in that. He's a PhD student, which helps, but for me, it was a topic I was waiting for someone to spend cycles on. There are many narrow things like, 'It doesn't make sense that there's no answer to this.' There's so much information that people can't grab onto any of it, but if you stick in an area, there's a lot to learn.

Nathan Lambert

是的,你不能试图什么都做,因为那会让人不知所措,而且试图跟上一切会让你筋疲力尽。对我来说,我很久没有跟上计算机视觉了;我只专注于 LLM。回到你的书,这是一本很棒的书,性价比很高。如果你想学习 RLHF,我不会去读 RLHF 论文,因为那会花掉你两年时间。有一章我不得不说,‘X 论文说一件事,Y 论文说另一件事,我们看看哪个会成为现实。’

Yeah, you can't try to do it all because it would be overwhelming and you'd burn out trying to keep up with everything. For me, I haven't kept up with computer vision in a long time; I just focused on LLMs. Coming back to your book, it's a great book and good bang for the buck. If you want to learn about RLHF, I wouldn't go out and read RLHF papers because you'd be spending two years. There's a chapter where I had to say, 'X papers say one thing, and Y papers say another, and we'll see what becomes true.'

后训练主题概览 Overview of Post-Training Topics

Host

我们可能遗漏了哪些关于后训练整体图景的想法?首先,你涵盖了问题设定、训练概述、什么是偏好、偏好数据、优化工具、奖励建模、正则化、指令微调、拒绝采样、强化学习、AI 策略梯度、直接对齐算法,然后是宪法 AI 和 AI 反馈、推理和推理时 Scaling、工具使用和函数调用、合成数据和蒸馏、评估,然后是开放问题部分、过度优化、风格与信息、产品 UX、角色和后训练。那么,有哪些值得提及的想法能同时连接教育性和研究性?你提到了角色训练,这挺有意思的。

What are some of the ideas we might have missed in the bigger picture of post-training? So, first of all, you do the problem setup, training overview, what are preferences, preferences data, and the optimization tools, reward modeling, regularization, instruction tuning, rejection sampling, reinforcement learning, AI policy gradients, direct alignment algorithms, then constitutional AI and AI feedback, reasoning and inference time scaling, tool use and function calling, synthetic data and distillation, evaluation, and then open questions section, over optimization, style and information, and then product UX, character, and post-training. So, what are some ideas worth mentioning that connect both educational component and the research component? You mentioned the character training, which is pretty interesting.

Nathan Lambert

角色训练很有趣,因为这方面的内容很少,但我们讨论过人们如何与这些模型互动,以及我们使用它们时感觉良好,因为它们很积极,但这可能过头了。可能过于积极了,本质上就是如何改变你的数据或决策过程,使其完全符合你的要求。OpenAI 有一个叫做模型规范的东西,本质上是他们希望模型做什么的内部指南,并且他们将其发布给开发者。所以,基本上你可以知道 OpenAI 训练的失败之处——比如他们有意图但尚未实现——与那些他们实际上想做而你不喜欢的事情之间的区别。这种透明度很好,但整理这些文档的方法以及遵循它们的难易程度并不为人所知。我认为这本书的设计方式是,强化学习章节显然是人们想要的,因为每个人都听说过 RLHF,算法和数学是一样的,但你可以在非常不同的文档中使用它。所以,我认为 RLHF 的核心在于偏好是多么混乱。这基本上是我几年前写的一篇论文的翻版,但这章会告诉你为什么 RLHF 永远无法完全解决,因为即使是 RL 的设定也假设偏好可以被量化,并且多个偏好可以简化为单一值。我认为这与经济学文献中的冯·诺依曼-摩根斯坦效用定理有关。这一章包含了所有哲学、经济学和心理学的背景。它告诉你 RLHF 压缩了什么。所以,你拥有所有这些,然后在书的后面,你用 RL 数学让数字上升。我认为这就是为什么人们做研究会很有收获,因为量化偏好这件事,就像人类设计了问题以使偏好可研究。但存在一些基本争论,例如在语言模型回应中,你关心不同的事情,无论是准确性还是风格。当你收集数据时,它们都被压缩成“我更喜欢这个而不是那个”。这种情况正在发生,世界上其他领域有很多研究探讨应该如何实际做到这一点。我认为社会选择理论是经济学的一个子领域,研究如何聚合偏好。有一个研讨会发表了一份白皮书,讨论如何考虑将社会选择理论用于 RLHF。所以我主要希望那些对数学感兴趣的人能够遇到并学习这种更广泛的背景。我觉得有一件有趣的事:我保留了一份我喜欢的所有推理模型的技术报告列表。所以在第 14 章,也就是 RLVR 的简短总结中,有一个巨大的表格,列出了每一个我喜欢的推理模型。我认为在教育中,很多内容现在需要像“我喜欢什么”这样,因为语言模型在数学方面非常擅长,比如著名的论文直接偏好优化,它比 RL 更简单地解决了问题。附录中的推导跳过了数学步骤。我在这本书中重新做了推导,我想,他们用来改变数学的那个对数技巧到底是什么?但用语言模型来做,它们会说,这就是对数技巧。我想,我不知道我是否喜欢数学被如此商品化。我认为在阅读这个附录和跟随数学推导时的一些挣扎对学习是有益的。

Character training is interesting because there's so little out of it, but we talked about how people engage with these models and like we feel good using them because they're positive, but that can go too far. It could be too positive, and it's essentially how do you change your data and or decision-making to make it exactly what you want. And OpenAI has this thing called a model spec, which is essentially their internal guideline for what they want the model to do, and they publish this to developers. So, essentially you can know what is a failure of OpenAI's training, which is like they have their intentions and they haven't met it yet, versus what is something that they actually wanted to do and that you don't like. And that transparency is very nice, but all the methods for curating these documents and how easy it is to follow them is not very well known. I think the way the book is designed is that the reinforcement learning chapter is obviously what people want because everybody hears about it with RLHF, and it's the same algorithms and the same math, but it's just like you can use it in very different documents. So, I think the core of RLHF is how messy preferences are. It's essentially a rehash of a paper I wrote years ago, but this is essentially the chapter that'll tell you why RLHF is never ever fully solvable because the way that even RL is set up is that it assumes that preferences can be quantified and that multiple preferences can be reduced to a single value. And I think it relates in the economics literature to the Von Neumann Morgenstern utility theorem. And that is the chapter where all of that philosophical, economic, and psychological context comes in. It tells you what gets compressed into doing RLHF. So it's like you have all of this and then at the end later in the book it's like you use this RL math to make the number go up. And I think that that's why I think it would be very rewarding for people to do research on is because it's like quantifying preferences is something that is just like humans have designed the problem in order to make preferences studyable. But there's kind of fundamental debates on like an example is in a language model response you have different things you care about whether it's accuracy or in style. And when you're collecting the data they all get compressed into like I like this more than another. And that is happening and there's a lot of research in other areas of the world that go into like how should you actually do this? I think social choice theory is the subfield of economics around how you should aggregate preferences. And there was a workshop that published a white paper on how can you think about using social choice theory for RLHF. So I mostly would want people that get excited about the math to come and have things where they can stumble into and learn this kind of broader context. I think there's a fun thing: I just keep a list of all the tech reports that I like of reasoning models. So in chapter 14, which is just a kind of short summary of RLVR, there's just a gigantic table where I list every single reasoning model that I like. So I think in education a lot of it needs to be like at this point it's like what I like because the language models are so good at the math where it's like famous paper direct preference optimization which is a much simpler way of solving the problem than RL. The derivations in the appendix skip steps of math. And I tried for this book, I redid the derivations and I'm like, what the heck is this log trick that they used to change the math? But doing it with language models, they're like, this is the log trick. And I'm like, I don't know if I like this that the math is so commoditized. I think some of the struggle in reading this appendix and following the math is good for learning.

Host

是的,所以我们经常回到教育这个话题。你们两位都多次提到“挣扎”这个词。所以这是有价值的。如果你在这个过程中没有挣扎,我想你就没有完全遵循正确的学习过程。

Yeah, so we're actually returning to this often just on the topic of education. You both have brought up the word struggle quite a bit. So there is value. If you're not struggling as part of this process, you're not fully following the proper process for learning, I suppose.

Nathan Lambert

一些提供商开始开发教育模型,这些模型旨在不一次性给出所有信息,让人们努力去获取。所以我认为你可以训练模型做到这一点,这将是一个很棒的贡献。就像书中的所有内容,你必须重新评估每一个决策,这是一个很好的例子。我认为在 AI2 有机会做这个,我当时想,哦,这太酷了。

Some of the providers are starting to work on models for education, which are designed to not give all the information at once and make people work to do this. So I think you could train models to do this and it would be a wonderful contribution. We're like all of this stuff in the book, you have to re-evaluate every decision for it, which is such a great example. And I think there's a chance to work on it at AI2, which I was like, oh, I think this would be so cool.

Host

有道理。我也做过类似的事情。比如前几天在电子游戏里。我有时会玩电子游戏消遣。我喜欢解谜类的游戏,比如《塞尔达》和《银河战士》。有一个新游戏我卡住了,真的卡住了,我想,我不想挣扎两天,所以我用了大语言模型。但我说,嘿,请不要剧透。我只说我在这里那里,下一步该做什么?我想同样的事情也可以用在数学上,你说,好的,我到了这一步,卡住了。不要给我完整答案,但有什么我可以尝试的?就像你小心地试探它。但问题在于,我认为这需要自律,很多人做数学——我是说,有很多人喜欢数学,但也有很多人需要做作业。然后就会走捷径。是的,我们可以开发一个教育型大语言模型,但其他大语言模型还在,仍然有使用其他模型的诱惑。

Makes sense. I do something like that. Did that the other day for video games, for example. I sometimes for my pastime play video games. Like I like video games with puzzles. So you know, like Zelda and Metroid. And there's this new game where I got stuck and I really got stuck and I was okay, I you know, I don't want to struggle like two days and so I used an LLM. But then you say, hey, please don't add any spoilers. Just you know, I'm here and there. What do I have to do next? And the same thing you can do, I guess, for math where you say, okay, I'm here at this point. I'm getting stuck. Don't give me the full solution, but what is something I could try, you know, like where you kind of carefully probe it. But the problem here is I think it requires discipline and a lot of people do math for I mean, there are a lot of people who enjoy math, but there are also a lot of people who need to do it for their homework. And then it's like the shortcut. And yeah, we can develop an educational LLM, but the other LLM is still there and there's still a temptation to use the other LLMs.

培养研究品味 Developing research taste

Nathan Lambert

我认为很多人,尤其是大学生,他们知道自己热爱什么,有自知之明,也知道这不应该容易。我觉得我们得培养好的品味。研究品味、学术品味,知道哪些事情值得你挣扎,哪些不值得,这很难判断,因为有时你缺乏长远眼光,看不清什么对你的职业生涯真正有用。但你必须培养这种品味。

I think a lot of people, especially in college, they understand the stuff they're passionate about, they're self-aware about it, and they understand it shouldn't be easy. I think we just have to develop a good taste. What research taste, like school taste, about stuff that you should be struggling on, and stuff you shouldn't be struggling on, which is tricky to know, because sometimes you don't have good long-term vision about what will be actually useful to you in your career. But you have to develop that taste.

Host

嗯。

Mhm.

教育的短暂数字窗口 The brief digital window in education

Nathan Lambert

我和未婚妻或朋友聊过这个,就像有一个短暂的十年窗口期,所有作业和考试都可以数字化,但在此之前,大家都只能用蓝皮书考试,因为没有别的办法。现在有了 AI,所有人都得回到蓝皮书和口试,因为作弊太容易了。就像这一代人经历了不同的教育体系,一切都可以数字化,但还不能作弊,而现在又要倒退回去。这真的很有意思。

I was talking to my fiance or friends about this, and it's like there's this brief 10-year window where all of the homework and all of the exams could be digital, but before that everybody had to do all the exams in blue book, because there was no other way. And now after AI, everybody's going to need to be in blue books and oral exams, because everybody could cheat so easily. It's like this brief generation that had a different education system where everything could be digital, but you still couldn't cheat, and now it's just going to go back. It's just really funny.

角色训练研究的算力需求 Compute requirements for character training research

Host

你提到了性格训练,我们把这个话题放大一点。那个课题需要多少算力?总的来说,作为一名研究者,有没有哪些地方不需要太多算力,个人研究者也能做出贡献?

You mentioned character training, just zooming out on a more general topic. For that topic, how much compute was required? And in general, to contribute as a researcher, are there places where not too much compute is required where you can actually contribute as an individual researcher?

Nathan Lambert

关于性格训练,我认为这项研究是基于用 LoRA 微调大约 70 亿参数的模型,这意味着你基本上只微调模型权重的一小部分。我不确切知道需要多少 GPU 小时,但这是可行的。不是每个学术机构都能做到。所以有些学术机构的情况很严峻,唯一能做的就是推理,使用闭源或开源模型,获取输出,然后观察和理解模型。这非常适合做评估,你要成为最擅长创建代表性问题的人,这些问题能让模型失败或展现某些能力,我认为你可以借此取得突破。所以,我认为一个从事评估的研究人员,如果想要职业发展,最高目标是前沿实验室采纳你的评估。比如,你从一所没有算力的小型大学出发,发现了一些 Claude 难以处理的问题,然后下一个 Claude 模型在博客文章中提到了它。这就是你职业生涯的火箭船。我认为这很难,但如果你想用最少的算力获得最大的影响力,大概就是这样:变得非常专注,并学习模型的发展方向。所以你需要构建一个工具,测试未来 Claude 4.5 会在哪里失败。如果你要开始一个研究项目,我需要思考 8 个月后模型会在哪里遇到困难。

For the character training thing, I think this research is built on fine-tuning about 7 billion parameter models with LoRA, which means you essentially only fine-tune a small subset of the weights of the model. I don't know exactly how many GPU hours that would take, but it's doable. Not doable for every academic. So the situation for some academics is so dire that the only work you can do is doing inference, where you have closed models or open models, and you get completions from them, and you can look at them and understand the models. And that's very well suited to evaluation, which you become you want to be the best at creating representative problems that the models fail on or show certain abilities, which I think that you can break through with this. So, I think that the top end goal for a researcher working on evaluation, if you want to have career momentum, is the frontier labs pick up your evaluation. So, it's like you don't need to have every project do this, but if you go from a small university with no compute and you figure out something that Claude struggles with and then the next Claude model has it in the blog post. Like, there's your career rocket ship. I think that that's hard, but it's like if you want to scope the maximum possible impact with minimum compute, it's something like that, which is just get very narrow and it takes learning of where the models are going. So, you need to build a tool that tests where not Claude 4.5 will fail. If you're going to start a research project, I need to think where the models in 8 months are going to be struggling.

权衡:新颖想法与职业路径 Trade-offs: novel ideas vs. practical career paths

Host

但是,开发全新的想法呢?

But, what about developing totally novel ideas?

Nathan Lambert

这是一个权衡。我认为如果你在读博士,你可能会觉得研究语言模型风险太大。我会看得更长远,思考什么将定义 10 年后语言模型的发展,但我最终是个很务实的人。我读博士时就想,‘我进了伯克利。最坏的情况是拿个硕士,然后去科技行业工作。’我对此非常务实。所以我觉得,在这些 AI 公司工作能提供的生活,平均薪酬超过每年一百万美元的股票,这很惊人。在美国,任何一个普通人进入这样的 AI 实验室,都能改变人生。所以我非常务实,如果你专注于此,语言模型领域仍然有很多上升空间,结果就是看看这些工作。但从研究角度看,那些变革性的学术奖项,比如成为下一个 Yann LeCun,来自于不太关心语言模型的发展。

This is a trade-off. I think that if you're doing a PhD, you could also be like it's too risky to work in language models. I'm going way longer term, which is what is the thing that's going to define language model development in 10 years, which I think that I end up being a person that's pretty practical. I mean, I went to my PhD where it's like, 'I got into Berkeley. Worst case, I get a master's and then I go work in tech.' It's like I'm very practical about it. So, I'm like the life afforded to people to work at these AI companies, the amount of like opening eyes average compensation is over a million dollars in stock a year for employee. Any normal person in the US, to get into this AI lab is transformative for your life. So, I'm pretty practical of like there's still a lot of upward mobility working in language models if you're focused and the outcome is like look at these jobs. But, from a research perspective, the transformative impact in these academic awards is like be the next Yann LeCun is from not working on not caring about language model development very much.

Host

那样的话,经济上牺牲很大。

It's a big financial sacrifice in that case.

Nathan Lambert

所以,我和一些很棒的学生合作,他们会问:‘我应该去 AI 实验室工作吗?’我说:‘你是在顶尖学校读博士,还是打算退学去实验室?’我说:‘我不知道。如果你去顶尖实验室,我不怪你。不要去那些可能归零的随机初创公司。但如果你去 OpenAI,我觉得值得为它退学。’

So, I get to work with some awesome students and they're like, 'Should I go work in an AI lab?' And I'm like, 'You're getting a PhD at a top school or you're going to leave to go to a lab?' I'm like, 'I don't know. If you go work at a top lab, I don't blame you. Don't go work at some random startup that might go to zero. But, if you're going to OpenAI, I'm like, it could be worth leaving a PhD for.'

职业建议:学术界与工业实验室 Career advice: academia vs. industry labs

Host

让我们更严谨地思考一下。你会建议人们在哪里做出研究贡献?选项有学术界,也就是读博士,花 5 年发表论文。计算资源受限。还有一些更专注于开放权重模型的研究实验室。在那里工作,或者闭源前沿实验室,比如 OpenAI、Anthropic、Axiom。

Let's more rigorously think through this. Where would you give a recommendation for people to do a research contribution? So, the options are academia, so get a PhD, spend 5 years publishing. Computer resources are constrained. There are research labs that are more focused on open weight models. And so, working there or closed frontier labs, research labs. The OpenAI, Anthropic, Axiom.

Nathan Lambert

两个梯度是:越封闭,你往往赚得越多。但你也获得更少的认可。所以,在建立你的成果组合方面,作为学者你做了什么非常清楚。你做了这个。而如果你去交换这种相当合理的进步,成为机器中的一个齿轮,那也可能很有趣。所以我认为这是非常不同的职业道路。但是,作为研究者的机会成本很高,因为博士生基本没有收入。所以,这最终奖励那些有稳定安全网、并且意识到可以长期运作的人,他们想做非常有趣的工作并得到一份非常有趣的工作。所以,说‘我要完成我的博士,之后再说,因为我想做这个’是一个相当特权的位置。同时,我认为很多学术机构,学术生态系统正受到资金削减等的冲击。所以有很多不同的权衡,我理解很多人说,‘哦,我受不了这种找资金的过程。我的拨款被政府无缘无故砍了。’或者我不知道会发生什么。所以,我认为有很多不确定性和权衡,在我看来,倾向于选择高薪且有意义影响的工作。你不是在 OpenAI 白拿钱。你在构建前沿的东西,改变数百万人与科技的关系。

The two gradients are the more closed, the more money you tend to get. But also, you get less credit. So, in terms of building a portfolio of things that you've done, it's very clear of what you have done as an academic. And you have done this. And versus if you are going to go be like, trade this fairly reasonable progression for being a cog in the machine, which could also be very fun. So, I think it's a very different career paths. But, the opportunity cost for being a researcher is very high because PhD students are paid essentially nothing. So, I think it ends up rewarding people that have a fairly stable safety net and they realize that they can operate in the long term, which is they want to do very interesting work and get a very interesting job. So, it is a fairly privileged position to be like, 'I'm going to see out my PhD and figure it out after because I want to do this.' And I think a lot of academics, at the same time, the academic ecosystem is getting bombarded by funding getting cut and stuff. So, there's just so many different trade-offs where I understand plenty of people that are like, 'Oh, I can't deal with this funding search. My grant got cut for no reason by the government.' or I don't know what's going to happen. So, I think there's a lot of uncertainty and trade-offs that in my opinion favor just like take the well-paying job with meaningful impact. It's not like you're getting paid to sit around at OpenAI. You're building the cutting edge of things that are changing millions of people's relationship to tech.

Host

嗯。

Mhm.

Nathan Lambert

但在发表方面,他们越来越保密了。

But publication-wise, they're being more secretive, increasingly so.

学术界与工业界 Academia vs. Industry

Host

所以,你发表的成果越来越少。虽然你在规模上产生了积极影响,但你只是机器中的一个齿轮。

So, you're publishing less and less. And so, you're having a positive impact at scale, but you're a cog in the machine.

Nathan Lambert

说实话,我觉得变化没那么大。我曾在学术界,现在不在了。但我不想错过在学术界的时光。不过在此之前我想说:变化真的不大。我以前和合作者一起,将 AI 和机器学习方法用于计算生物学应用。很多人直接从学术界去了谷歌。那时教授们因为学生去了工业界而难过,因为无法延续他们的学术传承。我觉得现在也一样。唯一变化的是规模。工业界一直有封闭开发的好东西,你不能谈论。现在的区别在于你的偏好:你喜欢谈论工作并发表论文,还是更倾向于封闭实验室?这是一个区别,当然还有薪酬,但一直如此。所以这真的取决于你感觉哪里舒服。而且没有什么是永恒的。现在有了第三个选择:创办初创公司。很多人都在做初创公司。风险很高,但高风险高回报。加入工业实验室相当安全,有晋升空间。一旦你在工业实验室待过,将来找工作会更容易。但话说回来,你有多喜欢团队和从事专有工作,相比发表论文呢?发表论文压力很大。会议接受率可能很随意,令人沮丧,但回报也高。如果你有论文发表,你会感觉很好,因为上面有你的名字。

I think it honestly hasn't changed that much. I've been in academia, and I'm not anymore. At the same time, I wouldn't want to miss my time in academia. But what I wanted to say before I get to that part: I think it hasn't changed that much. I was using AI and machine learning methods for applications in computational biology with collaborators. A lot of people went from academia directly to Google. Back then, professors were sad that their students went into industry because they couldn't carry on their legacy. I think it's the same thing now. The only thing that has changed is the scale. Cool stuff was always developed in industry that was closed, you couldn't talk about it. The difference now is your preference: do you like to talk about your work and publish, or are you more in a closed lab? That's one difference, and compensation of course, but it's always been like that. So it really depends on where you feel comfortable. Also, nothing is forever. The only thing now is there is a third option: starting a startup. That's a lot of people doing startups. Very risky, but high risk, high reward. Joining an industry lab is pretty safe, with upward mobility. Once you've been at an industry lab, it will be easier to find future jobs. But then again, how much do you enjoy the team and working on proprietary things versus publishing? Publishing is stressful. Acceptance rates at conferences can be arbitrary and frustrating, but also high reward. If you have a paper published, you feel good because your name is on it.

Host

老实说,我觉得我那些当教授的朋友平均比在前沿实验室工作的朋友更快乐。因为那更接地气,而前沿实验室确实实行 996,这基本上就是全天工作的代名词。

I feel like my friends who are professors seem on average happier than my friends who work at a frontier lab, to be totally honest. Because it's just grounding, and the frontier labs definitely do this 996, which essentially is shorthand for work all the time.

Nathan Lambert

你能描述一下 996 吗?这是一种我认为起源于中国并被硅谷采用的文化。什么是 996?

Can you describe 996? It's a culture that I believe was invented in China and adopted in Silicon Valley. What's 996?

Host

就是早上 9 点到晚上 9 点,每周六天。

It's 9:00 a.m. to 9:00 p.m., six days a week.

Nathan Lambert

每周六天。那是 72 小时?所以这基本上是硅谷 AI 公司的标准吗?这种拼命工作的心态越来越普遍。

Six days a week. That's 72 hours? So is this basically the standard in AI companies in Silicon Valley? More and more this kind of grind mindset.

Host

是的,我的意思是不完全那样,但我认为有这种趋势。有趣的是,我觉得几乎颠倒了。我在学术界时,也有那种感觉,因为作为教授,你要写经费申请、教学、做研究。这是三合一的工作,如果你想成功,这比全职工作还多。现在,教授与实验室相比,我觉得他们的压力或工作量比前沿实验室要小。

Yeah, I mean not exactly like that, but I think there is a trend towards it. It's interesting, I think it almost flipped. When I was in academia, I felt like that because as a professor you had to write grants, teach, and do research. It's three jobs in one, and it's more than a full-time job if you want to be successful. Now, professors in comparison to a lab, I think they have less pressure or workload than at a frontier lab.

Nathan Lambert

他们工作很多。但他们通过与学生的合作、持续的指导机会以及非常以人为本的使命而感到非常满足。我认为在这个变化迅速且混乱的时代,这对人们来说非常有回报。

They work a lot. They're just so fulfilled by working with students and having a constant runway of mentorship and a mission that is very people-oriented. I think in an era when things are moving very fast and very chaotic, it's very rewarding to people.

Host

是的,我认为在初创公司,有这种压力:你必须成功。人们投入时间非常重要,但这很难,因为你必须不断交付。我在初创公司待过。我度过了一段美好时光,但我不知道能否永远做下去。这是一种有趣的节奏,正如我们一开始讨论的:这些模型在互相超越,不断试图比竞争对手更进一步。我觉得这很残酷。

Yeah, and I think at a startup, there's this pressure: you have to make it. It's really important that people put in the time, but it's really hard because you have to deliver constantly. I've been at a startup. I had a good time, but I don't know if I could do it forever. It's an interesting pace, and exactly like we talked about in the beginning: these models are leapfrogging each other, constantly trying to take the next step compared to competitors. It's just ruthless, I think.

Nathan Lambert

我认为这种跳跃式发展和多个参与者的竞争实际上是语言模型进步中被低估的驱动力。竞争深深植根于人们心中,这些公司有意创造了非常强大的文化。比如 Anthropic 以其文化上的深度承诺和组织性而闻名。我们很少听到他们的消息,Anthropic 的每个人似乎都非常一致。处于一个非常紧密的文化中,加上这种竞争动态,会让你努力工作并创造更好的东西。但这以人力资本为代价。你只能坚持这么久,人们肯定在精疲力竭。我写过一篇关于倦怠的文章,我说我自己也经历过,尤其是试图成为完整模型训练的管理者。这是一份疯狂的工作。Patrick McGee 的《苹果在中国》一书谈到了苹果工程师为在中国建立供应链而付出的艰辛。他说他们有拯救婚姻计划,他在播客中表示有人因这种工作强度而死亡。所以我认为这是一个以人类代价创造进步的理想环境。有很多人类代价,就是我们一开始提到的 996:人们真的在拼命工作。

I think this leapfrogging nature and having multiple players is actually an underrated driver of language modeling progress. Competition is so deeply ingrained in people, and these companies have intentionally created very strong culture. Like Anthropic is known to be so culturally deeply committed and organized. We hear so little from them, and everybody at Anthropic seems very aligned. Being in a culture that is super tight and having this competitive dynamic is like a thing that's going to make you work hard and create things that are better. But that comes at the cost of human capital. You can only do this for so long, and people are definitely burning out. I wrote a post on burnout where I said I've tried in and out of this myself, especially trying to be a manager of full model training. It's a crazy job. The book 'Apple in China' by Patrick McGee talked about how hard the Apple engineers worked to set up supply chains in China. He said they had saving marriage programs, and he told in a podcast that people died from this level of working hard. So I think this is a perfect environment for creating progress based on human expense. There's a lot of human expense, which is the 996 we started with: people really grind.

Host

我也读过这本书。我记得他们有一个暗号,如果有人必须回家陪家人以挽救婚姻。同事们会说,‘好吧,这是红色警报,我们必须让那个人这个周末回家。’这很疯狂。但与此同时,我认为他们不是被迫工作的。他们对产品如此热情,以至于进入了那种心态。我作为学者时有时也有这种感觉,作为独立个体也是如此。我有时过度工作,这不健康。我有背部和颈部问题,因为我没有适当休息。但这不是因为有人强迫我,而是因为我想工作,因为这是健康的。

I also read this book. I think they had a code word for if someone had to go home to spend time with their family to save the marriage. It's crazy that colleagues would say, 'Okay, this is red alert, we have to let that person go home this weekend.' But at the same time, I don't think they were forced to work. They were so passionate about the product that you get into that mindset. I had that sometimes as an academic, but also as an independent person. I sometimes overwork and it's unhealthy. I had back issues and neck issues because I didn't take the breaks I should have. But it's not because no one forced me; it's because I wanted to work because it's healthy.

Nathan Lambert

AI 和 Anthropic 的人就像他们想做这份工作。

AI and Anthropic are like they want to do this work.

硅谷泡沫与炒作 Silicon Valley bubble and hype

Host

是的,但还有一种狂热情绪正在积聚,尤其是在硅谷,与 Scaling(规模扩张)的理念相呼应。这种炒作认为世界将在几周内被改变,而你想成为中心。我有幸与各种各样的人交谈,从中我看到了世界各地的这些泡沫和回音室。观察我们人类如何形成它们很有趣。我认为可以说硅谷是一种回音室、一种孤岛和泡沫。我认为泡沫实际上非常有用和有效。这不一定是坏事,因为它可能极具生产力。它可能是史蒂夫·乔布斯的现实扭曲场,因为你只是互相说服突破即将到来,而通过互相说服,你让突破真的到来。

Yeah but there's also a feeling of fervor that's building especially in Silicon Valley aligned with the scaling laws idea where there's this hype that the world will be transformed on a scale of weeks and you want to be at the center of it. I have this great fortune of having conversations with a wide variety of human beings and from there I get to see all these bubbles and echo chambers across the world. It's fascinating to see how we humans form them. I think it's fair to say that Silicon Valley is a kind of echo chamber, a kind of silo and bubble. I think bubbles are actually really useful and effective. It's not necessarily a negative thing because it could be ultra productive. It could be the Steve Jobs reality distortion field because you just convince each other the breakthroughs are imminent, and by convincing each other of that, you make the breakthroughs imminent.

Nathan Lambert

Burn Hobart 写了一本书对泡沫进行分类,但本质上其中一种是金融泡沫,就像投机,这是不好的。另一种是,我不知道术语,但实际上是建设性的,因为它推动人们去构建这些东西。我确实认为 AI 处于这种状态,但我担心它转变为金融泡沫。

Burn Hobart wrote a book classifying bubbles, but essentially one of them is financial bubbles, which is like speculation, which is bad. And the other one is, I don't know the term, but effectively for build-outs, because it pushes people to build these things. I do think AI is in this, but I worry about it transitioning to a financial bubble.

Host

是的,但在思想领域,那个泡沫你正在制造一个现实扭曲场,这意味着你在偏离现实。如果你在偏离现实的同时还 996 工作,你可能会错过人类体验的一些基本方面,包括在硅谷。这是硅谷的一个常见问题。这是一个非常特定的地理区域。你可能不理解中西部地区的视角,不理解美国乃至全世界所有其他不同人类的完整体验。你们以某种方式交谈,互相说服某件事。这可能会让你陷入真正的麻烦,无论 AI 是否大获成功并成为一项强大的技术。无论哪种轨迹,你都可能陷入麻烦。所以你必须考虑所有这些。你是一个年轻人,试图决定你一生想做什么。

Yeah, but also in the space of ideas, that bubble you are doing a reality distortion field, and that means you are deviating from reality. If you go too far from reality while also working 996, you might miss some fundamental aspects of the human experience, including in Silicon Valley. This is a common problem in Silicon Valley. It's a very specific geographic area. You might not understand the Midwest perspective, the full experience of all the other different humans in the United States and across the world. You speak a certain way to each other, you convince each other of a certain thing. That can get you into real trouble, whether AI is a big success and becomes a powerful technology or it's not. In either trajectory, you can get yourself into trouble. So you have to consider all of that. Here you are, a young person trying to decide what you want to do with your life.

Nathan Lambert

我甚至不太理解这个,但 SFAI 的迷因已经到了出现“永久底层阶级”这种说法的地步,意思是 2025 年的最后 6 个月是建立 AI 初创公司或模型的持久价值的唯一时机,否则所有价值都将被现有公司捕获,因此你会变穷。这是旧金山现象走得太远的例子。我仍然认为对于年轻人来说,如果你真的热衷于在 AI 领域产生影响,亲自在旧金山是最有可能做到这一点的地方,但这有取舍。

The thing that I don't even really understand this, but the SFAI memes have gotten to the point where permanent underclass was one of them, which was the idea that the last 6 months of 2025 was the only time to build durable value in AI startup or model, otherwise all the value will be captured by existing companies, and you will therefore be poor. That's an example of the SF thing that goes so far. I still think for young people that going to be able to tap into it, if you are really passionate about wanting to have an impact in AI, being physically in SF is the most likely place where you're going to do this, but it has trade-offs.

Host

我认为旧金山是一个不可思议的地方。但有一点泡沫。如果你进入那个泡沫,这非常有价值,但也要走出来。读历史书,读文学,去世界其他地方看看。Twitter 不是,Substack 也不是整个世界。

I think SF is an incredible place. But there is a bit of a bubble. If you go into that bubble, which is extremely valuable, just get out also. Read history books, read literature, visit other places in the world. Twitter is not and Substack is not the entire world.

Nathan Lambert

我想说,我的一位同事正在搬到旧金山,我觉得我需要给他一本《Season of the Witch》,这是一本关于旧金山从 1960 年到 1985 年的历史书,涵盖了嬉皮士革命、同性恋群体接管城市及其文化兴起,然后是 HIV 艾滋病危机等等。这一切如此之近,充满了动荡和伤害,但也有旧金山的爱。没有人知道这些。这是一本很棒的书,《Season of the Witch》。我推荐它。我的一群旧金山朋友都推荐给我,我觉得就像住在那里一样。我住在那里时并没有体会到这种背景,而它如此之近。

I think I would say one of my people I worked with is moving to SF and it's like I need to get him a copy of Season of the Witch, which is a history of SF from 1960 to 1985, which goes through the hippie revolution, the gays kind of taking over the city and that culture emerging, and then the HIV AIDS crisis and other things. It's just like that is so recent and so much turmoil and hurt, but also like love in SF. No one knows about this. It's a great book, Season of the Witch. I recommend it. A bunch of my SF friends recommended it to me and I think that is just like living there. I lived there and I didn't appreciate this context and it's just so recent.

文生图缩放与替代架构 Text-to-image scaling and alternative architectures

Host

好的,我们谈了很多事情,当然包括去年令人兴奋的事情,但今年你们提到的一件令人兴奋的事情是文本到图像模型的 Scaling(规模扩张)以及对文本到图像的不同探索。你能谈谈那是什么以及有什么可能性吗?与当前的 LLM 不同的方法。

Yeah. Okay, we talked a lot about a lot of things, certainly about the things that were exciting last year, but this year one of the things you guys mentioned is exciting is the scaling of text-to-image models and just a different exploration of text-to-image. Can you talk about what that is and what the possibility holds? Sort of different kinds of approaches than the current LLMs.

Nathan Lambert

是的,我们谈了很多关于 Transformer 架构,特别是自回归 Transformer,比如 GPT。但这并不意味着没有人在做其他事情。人们总是在寻找下一个大事件,因为不这样做是愚蠢的。当然,现在 Transformer 架构是主流,它效果最好,目前没有其他东西能比,但把所有鸡蛋放在一个篮子里总不是好主意。所以人们正在开发自回归 Transformer 的其他替代方案。其中之一是文本扩散模型。听众可能从图像生成中了解扩散模型,比如 Stable Diffusion 推广了它。它是一篇关于生成图像的论文;那时人们使用 GAN,即生成对抗网络,然后出现了这种扩散过程,你迭代地对图像去噪,随着时间的推移产生非常好的图像质量。Stable Diffusion 是一家公司;其他公司构建了自己的扩散模型。然后人们现在想,我们能否也尝试用于文本?这还没有直观意义,因为它感觉不像像素那样连续可微。它是离散的文本,那么我们如何实现去噪过程?但这类似于 Google 的 BERT 模型。回到最初的 Transformer,有编码器和解码器。解码器就是我们目前在 GPT 等中使用的。编码器更像是一种并行技术,你可以并行填充多个 token。GPT 模型是自回归的,一次一个 token,你一次一个 token 地完成句子。在 BERT 模型中,你有一个有缺口的句子文本。你掩盖它们,然后一次迭代填充这些缺口。文本扩散有点像那样,你从一些随机文本开始,然后填充缺失的部分或迭代地改进它们,并且你有多次迭代。酷的地方在于它可以同时处理多个 token,所以它有望更高效。当然,权衡是质量如何?它可能更快,现在你有去噪过程这个维度。你做的步骤越多,文本就越好。

Yeah, so we talked a lot about the Transformer architecture and the autoregressive Transformer architecture specifically like GPT. It doesn't mean no one else is working on anything else. People are always on the lookout for the next big thing because it would be stupid not to. Sure, right now the Transformer architecture is the thing and it works best and has nothing else out there, but it's always a good idea to not put all your eggs into one basket. So people are developing other alternatives to the autoregressive Transformer. One of them would be text diffusion models. Listeners may know diffusion models from image generation like Stable Diffusion popularized it. It was a paper on generating images; back then people used GANs, the generative adversarial networks, and then there was this diffusion process where you iteratively denoise an image and that resulted in really good quality images over time. Stable Diffusion was a company; other companies built their own diffusion models. And then people are now like, okay, can we try this also for text? It doesn't make intuitive sense yet because it feels like it's not something continuous like a pixel that we can differentiate. It's discrete text, so how do we implement that denoising process? But it's kind of similar to the BERT models by Google. When you go back to the original Transformer, there were the encoder and the decoder. The decoder is what we are using right now in GPT and so forth. The encoder is more like a parallel technique where you have multiple tokens that you fill in in parallel. Instead of GPT models which do autoregressive one token at a time, you complete the sentence one token at a time. In BERT models, you have a text that's a sentence that has gaps. You mask them out and then one iteration is filling in these gaps. Text diffusion is kind of like that where you are starting with some random text and then you are filling in the missing parts or you are refining them iteratively, and you have multiple iterations. The cool thing here is that this can do multiple tokens at the same time, so it's kind of the promise of having it more efficient. Now the trade-off is, of course, how good is the quality? It might be faster, and now you have this dimension of the denoising process. The more steps you do, the better the text becomes.

文本扩散模型与自回归模型 Text Diffusion Models vs. Auto-Regressive Models

Nathan Lambert

你可以用不同的方式进行扩展。人们试图验证扩散模型是否能成为自回归模型的有效替代方案,以更少的算力提供相同的质量。目前,有论文指出,要达到相同的质量,你必须增加去噪步骤,最终消耗的算力与自回归模型相当。另一个缺点是,虽然扩散模型是并行的,但有些任务并非并行,比如推理任务或工具使用,这些任务需要从解释器获取中间结果。因此,存在一些混合模型,但核心思想是并行化。目前,大多数扩散模型还处于研究阶段,比如 Llama,一些初创公司也部署了模型,但还没有像 Gemini 或 ChatGPT 那样大规模部署的扩散模型。不过,谷歌宣布了 Gemini 扩散模型,很可能用于他们的 Nano-2 模型,声称在大多数基准测试中,以相同的质量实现更快的生成。我不认为文本扩散模型会取代自回归 LLM,但它可能用于快速、廉价、大规模的任务,比如未来的免费套餐。

You can scale in different ways. People try to see if diffusion is a valid alternative to auto-regressive models, offering the same quality for less compute. Right now, papers suggest that to get the same quality, you have to crank up the denoising steps, ending up with the same compute as auto-regressive models. Another downside is that while diffusion is parallel, some tasks are not parallel, like reasoning tasks or tool use where you need intermediate results from an interpreter. So there are hybrids, but the main idea is parallelization. Currently, most diffusion models are research models like Llama, and some startups have deployed models, but there's no large-scale diffusion model at the level of Gemini or ChatGPT. However, Google announced Gemini diffusion, likely for their Nano-2 model, claiming faster generation for the same quality on most benchmarks. I don't think text diffusion will replace auto-regressive LLMs, but it might be used for quick, cheap, at-scale tasks, like a free tier in the future.

Host

我听说过一些例子,扩散模型已经在使用了。例如,当 GPT-5 需要 30 分钟响应时,它一次生成一个 token。扩散模型则一次性生成所有 token,因此快得多。初创公司将其用于代码差异:当有人做出更改时,代码差异是一个巨大的回复,但不需要太多外部上下文,因此扩散模型可以快速生成。自回归模型需要几分钟,导致用户流失。所以我认为扩散模型会在某些应用中增长,但我原本预期不同模型会更早地用于不同任务。工具使用这一点是通用用途的障碍,因为自回归链会被外部工具打断,而扩散模型如何处理这一点尚不明确。

I've heard of a couple examples where diffusion is already being used. For instance, when GPT-5 takes 30 minutes to respond, it generates one token at a time. Diffusion generates all tokens in one batch, making it much faster. Startups are using it for code diffs: when someone makes a change, the code diff is a huge reply, but it doesn't need much external context, so diffusion models can generate it quickly. Auto-regressive models would take minutes, causing user churn. So I think diffusion will grow in some applications, but I expected different models for different things sooner. The tool use point is a barrier for general-purpose use, because auto-regressive chains are interrupted by external tools, and it's unclear how to handle that with diffusion.

工具使用的未来 Future of Tool Use

Host

那么今年及未来几年工具使用的未来是什么?你认为在如何将其集成到整个技术栈方面会有很多发展吗?

So what's the future of tool use this year and in the coming years? Do you think there will be a lot of developments in how it's integrated into the entire stack?

Nathan Lambert

目前,工具使用主要出现在专有 LLM 方面,但我认为我们会在开源工具中看到更多。这是一个巨大的突破,因为你可以将任务从记忆外包给实际计算,比如使用计算器而不是让 LLM 记住 23+5。

Right now, tool use is mostly on the proprietary LLM side, but I think we'll see more in open-source tooling. It's a huge unlock because you can outsource tasks from memorization to actual computation, like using a calculator instead of having the LLM memorize 23+5.

Host

那么你认为这有助于解决幻觉问题吗?

So you think that can help solve hallucination?

Nathan Lambert

不能完全解决,但可以减少。LLM 仍然需要知道何时调用工具,而且互联网并不总是正确的。例如,询问谁赢得了 1998 年世界杯,它仍然需要找到正确的网站。所以这不会完全解决幻觉,但会有所改善。去年年底的另一篇有趣的论文,递归语言模型,进一步推进了这一点。他们用 GPT-5 完成了所有工作,没有使用本地模型。其思想是,对于长上下文任务,不是一次性解决,而是将其分解为子任务。LLM 决定什么是好的子任务,然后递归地调用 LLM 来解决它。添加工具后,每个子任务可以访问网络并收集信息,然后拼接在一起。这可以带来很多突破,不是通过改进 LLM 本身,而是通过改进它的使用方式和它能使用的东西。工具使用的一个缺点是,你必须授予 LLM 使用工具的权限,这需要信任。例如,让 LLM 回复电子邮件——我今天不会让它访问我的电子邮件;这是一个巨大的风险。

Not solve it, but reduce it. The LLM still needs to know when to call a tool, and the internet isn't always correct. For example, asking who won the World Cup in 1998, it still needs to find the right website. So it won't fully solve hallucination, but it improves things. Another cool paper from late last year, the recursive language model, takes this further. They did everything with GPT-5, not local models. The idea is for long-context tasks, instead of solving it in one shot, you break it into subtasks. The LLM decides what subtask is good, then recursively calls an LLM to solve it. Adding tools, each subtask could go to the web and gather information, then stitch it together. This could unlock a lot, not by improving the LLM itself, but by improving how it's used and what it can use. One downside with tool use is you have to give the LLM permission to use tools, which requires trust. For example, having an LLM answer emails—I wouldn't give it access to my emails today; it's a huge risk.

工具使用中的开源与闭源模型 Open vs. Closed Models in Tool Use

Host

我认为这是一个很好的观点。关于工具使用的最后一点:你暗示了开放模型和封闭模型在使用工具方面非常不同。对于开放模型,人们下载模型,然后决定使用哪个工具,比如搜索提供商。但发布一个适用于多种工具和多种用例的模型很难,因为你是在制造一个通用推理引擎,这正是 GPT-OS 擅长的。而在封闭模型上,你将特定工具深度集成到体验中。

I think that's a cool point. One last point on tool use: you hinted that open and closed models use tools very differently. With open models, people download the model and then decide which tool to use, like a search provider. But releasing a model that works with multiple tools for multiple use cases is hard because you're making a general reasoning engine, which is what GPT-OS is good for. On closed models, you deeply integrate a specific tool into the experience.

持续学习的定义与重要性 Open vs Closed Models and Tool Use

Nathan Lambert

我认为开源模型将难以复制我用闭源模型做的一些事情,比如你可以引用公共和私有信息的混合,我每 3 到 6 个月就会尝试一次 Codex on the web,就是提示模型对我某个 GitHub 仓库进行更新。那种安全的云环境非常棒,只需发送指令让它做事,然后回来查看结果。这些可能会帮助定义本地开源和闭源的细分市场,但我认为最初由于急于实现工具使用功能,开源模型处于劣势,这几乎是不可避免的。前沿实验室有大量的研究和资源,但当开源模型解决这个问题时会很有趣,因为这将需要一种更灵活、可能更有趣的模型,能够与这种递归想法配合,充当编排者和工具使用模型。所以,希望需求能推动一些有趣的创新。

And I think that open models will struggle to replicate some of the things that I like to do with closed models, which will be like, I don't know, you can reference a mix of public and private information and something that I keep trying every 3 to 6 months I try like Codex on the web, which is just prompting a model to make an update to some GitHub repository that I have. And it's just like that set of secure cloud environment is just so nice for just like send it off and do this thing and then come back to me. And these will probably help define some of the local open and closed niches, but I think initially because there was such a rush to get this tool use working that the open models were on the back foot, which is kind of inevitable. I think there's so much research that there's so many resources in these frontier labs, but will be fun when the open models solve this because it's going to necessitate like a bit more flexible and potentially interesting model that might work with this recursive idea to like be an orchestrator and the tool used model. So, hopefully necessity drives some interesting innovation there.

持续学习与上下文学习 Continual Learning Definition and Importance

Host

那么,持续学习。这是一个长期存在的话题,重要的问题。我认为随着模型训练成本的上升,它的重要性也在增加。你能解释一下什么是持续学习,以及它在今年和未来几年取得进展方面有多重要吗?

So, continual learning. This is a long-standing topic, important problem. I think that increases in importance as the cost of training of the models goes up. So, can you explain what continual learning is and how important it might be this year and in the coming years to make progress?

Nathan Lambert

这与什么是 AGI(通用人工智能)、什么是 ASI(超级人工智能)以及我们今天的语言模型能做什么有很大关系。我认为语言模型可以解决很多任务,但 AI 社区的一个关键里程碑是 AI 能够取代任何远程工作者,接收信息、解决数字任务并完成它们。我认为人们强调的一个限制是语言模型不会像员工那样从反馈中学习。如果你雇了一个编辑,他会犯错,但你会告诉他,如果你雇了一个好编辑,他就不会再犯。但语言模型没有这种自我修改和快速学习的能力。所以,如果我们想要真正获得一种通用的、可适应的智能,能够进入任何远程工作场景,它需要能够快速从反馈和在职学习中学习。我个人更看好语言模型,只要给它们提供非常好的上下文。你之前可能离线说过,你可以给模型写大量的文档,比如‘我有所有这些信息,这是我写过的所有博客文章,我喜欢这种写作风格,我的声音基于此’,但很多人不这样做,而且模型以前也不是为处理这么多上下文而设计的。智能体模型才刚刚开始。所以,这是一种权衡:我们是否需要通过持续学习来更新模型权重,使其快速学习?还是相反的观点认为我们只需要提供更多上下文和信息,它们就会因为拥有大量上下文和非常聪明而表现出快速学习的样子?

This relates a lot to this kind of SF got the to what is AGI, what is which is artificial general intelligence, and what is ASI, artificial super intelligence, and what are the language models that we have today capable of doing? I think the language models can solve a lot of tasks, but a key milestone among the AI community is essentially when AI could replace any remote worker taking in information and solving digital tasks and doing them. And I think the limitation that's highlighted by people is that a language model will not learn from feedback the same way that an employee is. So, if you hire an editor, the editor will mess up, but you will tell them, and if you hired a good editor, they don't do it again. But, language models don't have this ability to modify themselves and learn very quickly. So, the idea is if we were going to actually get to something that is a true like general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-the-job learning. I'm personally more bullish on language models by being able to just provide them with very good context. You said like you maybe offline said that like you can write extensive documents to models where you say, I have all this information. Here's all the blog posts I've ever written. I like this type of writing. My voice is based on this, but a lot of people don't provide this to models, and the models weren't designed to like take this amount of context previously. Like the agentic models are just starting. So, it's this kind of trade-off of do we need to update the weights of this model with this continual learning thing to make them learn fast or the counter argument is we just need to provide them with more context and information, and they will have the appearance of learning fast by just having a lot of context and being very smart.

记忆机制 Continual Learning vs In-Context Learning

Host

所以,我们应该在这里提一下术语。持续学习指的是不断改变权重,使模型根据新输入的信息持续、快速、频繁地适应和调整。而你提到的另一面通常被称为上下文学习。当你学习东西时,有一个巨大的上下文窗口,每次提示系统时都可以不断加载额外信息。我认为两者都可以合理地被视为学习,只是学习发生的地方不同。

So, we should mention the terminology here. So, continual learning refers to changing the weights continuously so that the model adapts, adjusts based on the new incoming information, does so continually, rapidly, and frequently, and so on. And then the thing you mentioned on the other side of it is generally will be referred to as in-context learning. As you learn stuff, there's a huge context window. You can just keep loading it with extra information every time you prompt the system, which I think both are legitimately can be seen as learning. It's just a different place where you're doing the learning.

Nathan Lambert

老实说,持续学习——权重的更新——我们已经有了不同的形式。我的意思是,如果你想想看,这里的区别是你是在为每个人定制个性化模型,还是在全局模型规模上做?我认为我们已经有了,比如从 GPT-5 到 5.1 再到 5.2。可能不是即时的,但就像一次精心策划的更新,根据它不能做的事情的反馈、社区的反馈,他们更新权重,下一个模型,等等。所以,这有点像那种形式。另一个更细粒度的例子是 RLVR。你运行它,它就会更新。问题是你不能为每个人这样做,因为为每个人更新权重太昂贵了。我认为这就是问题所在。所以,除非你——我的意思是,即使在 OpenAI 的规模下,建设数据中心,我认为也太昂贵了。只有当你把东西放在设备上,比如苹果尝试用苹果基础模型,把它们放在手机上,然后它们从经验中学习,这才可行。

I think to be honest with you, continual learning that the updating of weights we already have that in different flavors. I mean, if you think about how, I think the distinction here is do you do that on a personalized custom model for each person, or do you do it on a global model scale? And I think we have that already with going from GPT-5 to 5.1 and 5.2. It's maybe not immediate, but it is like a curated update, a quick curated update where there was feedback by the things it couldn't do, feedback by the community, they updated the weights, next model, and so forth. So, it is, I mean, kind of like a flavor of that. Another even finer grained example is like RLVR. You run it, it updates. The problem is you can't just do that for each person because it would be too expensive to update the weights for each person. And I think that's the problem. So, unless you get, I mean, even at OpenAI scale, building the data centers, it would be too expensive, I think. That is only feasible once you have something on the device where it is on the consumer, like what Apple tried to do with the Apple foundation models, putting them on the phone, and then they learn from the experience.

长上下文的创新 Memory Mechanisms

Host

一个有点相关的话题,但这是一个可能拟人化的术语——记忆。关于如何给这些系统添加记忆,有哪些不同的机制想法?特别是你越来越多地看到个性化记忆。

A bit of a related topic, but this kind of maybe anthropomorphized term, but memory. What are the different ideas of the mechanism of how to add memory to these systems? As you're increasingly seeing so personalized memory, especially.

Nathan Lambert

目前,它主要像上下文,基本上是把东西塞进上下文然后回忆起来。但同样,我认为这很昂贵,因为你必须——你可以缓存它,但还是要花费 token。第二点是你能做的有限。我认为这更像是一种偏好或风格。很多人在解数学题时会这样做。你可以添加先前的知识,但也给它某些偏好提示,比如‘做我上次喜欢的’,诸如此类。但这并没有解锁新能力。为此,人们仍然使用 LoRA 适配器。这基本上不是更新整个权重矩阵,而是两个较小的权重矩阵,像并行叠加的增量。但你可以在一定程度上这样做,但同样,这是经济问题。也有论文说,比如 LoRA 学得少但忘得少。你知道,没有免费的午餐。如果你想学更多,就需要用更多权重,但会更贵。而且,学得越多,忘得越多。你基本上需要找到那个黄金平衡点。

So, right now it's mostly like context, basically stuffing things into the context and then just recalling that. But, again, I think well, it's expensive because you have to, I mean, you can cache it, but still you spend tokens on that. And the second one is you can only do so much. I think it's more like a preference and or a style. I mean, a lot of people do that when they solve math problems. You say it's basically you can add previous knowledge and stuff, but you also give it certain preference prompts do what I preferred last time, whatever, like something like that. But, it doesn't unlock new capabilities. So, for that, one thing people do use still is LoRA adapters. These are basically instead of updating the whole weight matrix, they're two smaller weight matrices that you kind of have in parallel overlays like the delta. But, yeah, you can do that to some extent, but then again, it is economics. So, there were also papers, for example, LoRA learns less, but forgets less. It's like, you know, there's no free lunch. If you want to learn more, you need to use more weights, but it gets more expensive. And then again, if you learn more, you forget more. And it's like you have to find that Goldilocks zone, basically.

世界模型与 LLM Innovations in Long Context

Host

那里还有很多创新的可能吗?

Is there a lot of innovations that's possible there?

Nathan Lambert

我认为普遍接受的观点是,这是一个计算和数据问题,有时也涉及小的架构改动,比如注意力机制的变体。我们讨论过混合注意力模型,本质上就是在 Transformer 中加入类似状态空间模型的结构。这类模型更适合,因为建模最远 token 所需的算力更少。但这不是免费的,它们需要大量算力或合适的数据。世界上有多少 10 万 token 的序列?从哪里获取?我认为扩展它们最终会非常昂贵。所以我们很快达到了百万 token 的输入上下文长度。我预计它会继续增长,今年可能达到 200 万或 500 万,但我不认为会到 1 亿。那将是一个真正的突破。我认为这些突破是可能的。持续学习是一个研究问题,可能会有突破让 Transformer 在这方面表现更好且成本更低。

I think the colloquially accepted thing is that it's a computing data problem where you can and sometimes like small architecture things which are like attention variants. So, if you have we talked about like hybrid attention models, which is essentially if you have what looks like a state space model within your transformer. And like those are better suited because you have to spend less compute to model the furthest along token. And I think that but those aren't free cuz they have to be accompanied by a lot of compute or um the right data. So, how many sequences of 100,000 tokens do you have in the world? And where do you get these? And I think it just ends up being pretty expensive to scale them. So, we've like gotten to pretty quickly to like a million tokens of input context length. And I would expect it to keep increasing and like get to like 2 million or 5 million this year, but I don't expect it to go to like 100 million. And that would be like a true breakthrough. And I think those breakthroughs are possible. Like the continual learning thing I think of as a research problem where you could There could be a breakthrough that just makes transformers work way better at this, and it's cheap.

Host

这些事在大量科学关注下可能发生,但按部就班地推进,会随时间稳步增长。

Like these things could happen with so much scientific attention, but turning the crank, it'll be consistent increases over time.

Nathan Lambert

我认为从极端情况来看,没有免费的午餐。一方面,为了降低成本,你可以使用 RNN,它只有一个状态,把所有之前的信息都压缩进去。这是一个固定大小的状态,所以内存不会增长,但上下文越长,遗忘的信息就越多,因为你无法将所有内容压缩到一个状态中。另一方面,Transformer 试图记住每个 token,这在查找特定信息时很好,但非常昂贵,因为 KV 缓存和点积计算都会增长。至于 Mamba 层,它们也有类似的问题,就像 RNN 一样试图将所有内容压缩到一个状态,只是更具选择性。我认为这又回到了“金发姑娘区”,比如 RetNet 3,他们找到了一个很好的比例:需要多少注意力层来获取全局信息(所有内容都可访问),以及多少压缩状态。我认为我们未来的扩展方式就是找到更好的比例,在计算成本足够低和性能足够强之间取得平衡。

I think also looking at the extremes, I think there's again no free lunch. So, on the one extreme to make it cheap, you have a let's say an RNN that has a single state a state where you save everything from the previous stuff. It's like a specific fixed-size thing, so you never really grow the memory because it's you are stuffing everything into one state. But then, the longer the context gets, the more information you forget because you can't you can't keep I mean, compress everything into one state. Then, on the other hand, you have the transformers which try to remember every token, which is great sometimes which if you want to look up specific information, but it's very expensive because you have the KV cache that grows, the dot product that grows. But then, yeah, like you said, the Mamba layers, I mean, they kind of have the same problem I would say like an RNN you try to compress everything into one state. You're a little more selective there. But then, I think it's like this Goldilocks zone again with RetNet 3, they found like a good ratio of how many attention layers do you need for the global information where everything is accessible compared to having these compressed states. And I think that's how I think we will scale more by finding better, let's say, ratios in Goldilocks zone like between um like computing making it cheap enough to run, but then also making it powerful enough to be useful.

Host

再补充一点,递归语言模型论文是试图解决长上下文问题的论文之一。他们发现,与其把所有内容塞进长上下文,不如将其分解成多个较小的任务,通过多个较小的卡片节省内存,实际上可以获得比让语言模型一次性处理所有内容更好的准确性。这是一个新范式。我们拭目以待,可能还有其他变体。我认为我们仍会在长上下文方面取得进展,但正如 Nathan 所说,对于预训练本身,我们没有那么多长上下文文档,因此很难研究语言模型在该层面的行为。

And one more plug here, recursive language model paper, that is one of the papers that tries to kind of address the long context thing. So, what they found is essentially instead of stuffing everything into this long context, if you break it up into these smaller, multiple smaller tasks, so you save memory by having multiple smaller cards, you can get actually better accuracy than having the LM try everything all at once. I mean, it's a new paradigm. We'll see, you know, there might be other flavors of that. So, I think with that we will still make improvement on long context, but then also, like Nathan said, I think the problem is for pre-training itself, we don't have as many long context documents as other documents, so it's harder to study, basically, how LMs behave and stuff like that on on that level.

Nathan Lambert

有一些经验法则:预训练语言模型时,比如我们以 8K 上下文长度预训练,然后通过训练扩展到 32K。大致上,将训练上下文长度加倍需要大约 2 倍算力,然后通常可以再扩展 2 到 4 倍。所以我认为很多预训练最终受限于算力,今年顶级实验室的算力大幅增加,这应该会体现在更长的上下文窗口上。但在后训练方面,有一些更有趣的事情:随着智能体的出现,它们将自行管理上下文。目前,经常使用 Claude Code 的人害怕压缩,即 Claude 将其全部 10 万 token 压缩成要点列表。但下一代模型将能够控制何时以及如何压缩。你可以训练强化学习算法,将压缩作为一个动作来缩短历史记录,问题设定是:在模型将历史压缩到最短的同时,保持最高的评估分数。这样你就能以最少的 token 进行这种复合自回归预测。这实际上是一个很好的问题设定,智能体模型学会以不同于简单前推的方式使用上下文。

There are some rules of thumb where essentially you pre-train a language model, like although we pre-trained it like 8K context length and then extended to 32K with training. And there's some rules of thumb where you just like essentially doubling the training context length takes like 2X compute, and then you can normally like 2 to 4X the context length again. So, I think a lot of it ends up being kind of compute bound at pre-training, which is in this like we talked about this, everyone talks about this, a big increase in compute for the top labs this year, and that should reflect in some longer context windows, but I think on the post-training side, there's some more interesting things, which is as we have agents the agents are going to manage this context on their own, where now at people that use Claude code a lot dread the compaction, which is when Claude takes its entire full 100,000 tokens of compacting into bulleted list, but what the next models will do, I'm just about novel, I'm sure people are already working on this, is essentially the model can control when it compacts and how. So, you can essentially like train your RL algorithm where compaction is an action where it shortens the history, and then the problem formulation will be I want to keep the maximum evaluation scores that I have gotten while the model compacts its history to the minimum length, because then you have the minimum amount of tokens that you need to do this kind of compounding auto aggressive prediction. So, there's actually pretty nice problem setups in this where the like these agentic models learn to use their context in a different way than just plow forward.

Host

最近一个有趣的例子是 DeepSeek 3.2,他们采用了稀疏注意力机制,本质上是一个高效轻量级的索引器,不是关注所有 token,而是选择实际需要的 token。这几乎回到了注意力的原始思想:选择性。但注意力机制中,你对某些 token 的权重可能为零,但你仍然使用了所有 token。而他们更进一步,直接掩码掉或根本不处理。滑动窗口注意力也有类似的想法:你有一个固定大小的滚动窗口,因为不需要一直关注所有内容。偶尔在某些层你可能需要全局信息,但那样很浪费。目前,使用所有 token 是安全的,性价比最高,因为你不会错过信息。我认为今年将是更聪明地处理这个问题的年份。现在人们想要下一个最先进的技术,而最先进的技术恰好是蛮力昂贵的方法。一旦你有了它,就像你说的,保持准确性,但看看如何用技巧降低成本。

One interesting also recent example would be DeepSeek version 3.2 where they had like the sparse attention mechanism where they have essentially like a very efficient small lightweight indexer and instead of attending to all the tokens it selects, okay, what tokens do I actually need? It's I mean it's almost comes back to the original idea of attention where you are selective, but attention is always on you have maybe zero weight on some of them, but you use them all, but they are even more like, okay, let's just mask that out or like not even do that. And even with sliding window attention almost that is also kind of like that idea you have that rolling window where you keep it fixed cuz you don't need everything all the time. You occasionally some layers you might, but it's wasteful, but right now I think yeah, if you use everything you're on the safe side. It gives you the best bang for the buck because you never miss information and right now I think this year will be more also the year figuring out, like you said, how to be more smart about that. I think right now people want to have the next state of the art and the state of the art is happens to be the brute force expensive thing and then once you have that, like you said, keep that accuracy, but let's see how we can do that cheaper now, like tricks, you know.

Nathan Lambert

是的,所有这些 Scaling(规模扩张)的事情。

Yeah, all this scaling thing.

机器人技术与缩放 World Models and LLMs

Host

就像我们之所以先得到 Claude 4.5 Sonnet 模型,是因为你可以更快地训练它,而且不会那么快遇到算力瓶颈,他们可以尝试更多东西,更快得到模型,尽管更大的模型实际上更好。我想说的是,AI 领域有很多令人兴奋的事情。我最近特别关注机器人学。所以我们今天几乎完全没有讨论机器人学。图像生成、视频生成也有很多内容。我认为公平地说,从数量、强度、热情来看,最令人兴奋的研究工作是在 LLM 领域,这就是为什么我认为我们专注于讨论 LLM 是合理的。但引入一些可能有用的东西也不错。例如,世界模型。这方面的兴奋度在增长。你认为在未来一年里,世界模型在 LLM 领域会有用吗?

Like the reason we get the Claude 4.5 Sonnet model first is because you can train it faster and you're not hitting these compute walls as soon and they can just try a lot more things and get the model faster even though the bigger model is actually better. I think we should say that there's a lot of exciting stuff going on in the AI space. My mind has recently been really focused on robotics. So, we have today really almost entirely didn't talk about robotics. There's a lot of stuff on image gen, video generation. I think it's fair to say that the most exciting research work in terms of the amount, intensity, fervor is in the LLM space, which is why I think it's justified for us to really focus on the LLM that we're discussing. But, it would be nice to bring in some certain things that might be useful. For example, world models. There's growing excitement on that. Do you think there would be any use in this coming year for world models in the LLM space?

Nathan Lambert

是的,我也这么认为。另外,关于 LLM,有趣的是,我认为如果我们解锁更多 LLM 能力,它也会自动解锁所有其他领域,或者不是解锁,而是让进展更快。这是因为很多研究人员和工程师使用 LLM 来编程。所以即使他们从事机器人学,如果你优化这些帮助编程的 LLM,那也是值得的。但世界模型很有趣。它基本上让模型在某种意义上运行世界的模拟,像一个真实事物的玩具版,这可以再次解锁 LLM 不知道的能力。它可以模拟事物。我认为 LLM 通过预训练和下一个词预测恰好工作得很好。但我们可以做得更复杂一些。我的意思是,有一篇 Meta 的论文关于代码世界模型。他们基本上将世界模型的概念应用于 LLM,不仅使用下一个词预测和可验证奖励来检查答案正确性,还确保中间变量正确。这有点像模型在学习一个代码环境。我认为这很有道理。只是做起来很贵,但这让事情更复杂,建模整个事情,而不仅仅是结果。所以它可以增加更多价值。我记得我读研究生时,有一个叫 CASP 的比赛,做蛋白质结构预测。他们预测尚未解析的蛋白质结构。所以这很棒,我认为 LLM 也需要类似的东西,你进行基准测试,但没人知道答案,然后事后有人揭示。但 AlphaFold 出来时,它碾压了这个基准。有很多迭代,但我记得第一个版本明确建模了分子的物理相互作用,比如角度和可能的角度。然后下一个版本,他们去掉了这个,只是用暴力 Scaling 扩展。对于 LLM,我们目前处于这种暴力 Scaling 阶段,因为它恰好有效。但我认为在某个时候,重新引入这种东西可能是有意义的。对于代码世界模型,我认为那可能很酷。当然,对于机器人学,这完全与 LLM 无关。

Yes, I do think so as well. Also, with LLMs, what's interesting here is I think if we unlock more LLM capabilities, it also automatically unlocks all the other fields because, or not unlocks, but makes progress faster. It's because a lot of researchers and engineers use LLMs, like we said, for coding. So, even if they work on robotics, if you optimize these LLMs that help you with coding, it pays off. But then yes, world models are interesting. It's basically where you have the model run a simulation of the world in a sense, like a little toy thing of the real thing, which can again unlock capabilities that the LLM is not aware of. It can simulate things. I think this is something like LLMs just happen to work well by pre-training and then doing the next token prediction. But we could do this even a bit more sophisticated. So what I'm saying is with, I think it was by Meta, a paper on code world models. So where they basically apply the concept of world models to LLMs again, where instead of just having next token prediction and verifiable rewards, checking the answer correctness, they also make sure the intermediate variables are correct. It's kind of like the model is learning a code environment in a sense. And I think this makes a lot of sense. It's just expensive to do, but this is like making things more sophisticated, modeling the whole thing, not just the result. And so it can add more value. I remember when I was a grad student, there is a competition called CASP, I think, where they do protein structure prediction. They predict the structure of a protein that is not solved yet. So in a sense, this is actually great and I think we need something like that for LLMs also, where you do the benchmark, but no one knows the solution, and then after the fact someone reveals that. But AlphaFold, when it came out, it crushed this benchmark. I mean there are also multiple iterations, but I remember the first one explicitly modeled the physical interactions of the molecule, like the angles and possible angles. Then in the next version, I think they got rid of this and just with brute force scaling it up. And with LLMs, we are currently in this brute force scaling, because it just happens to work. But I do think at some point it might make sense to bring back this thing. And with code world models, I think that might be actually quite cool. And then of course also for robotics, that is completely unrelated from LLMs.

机器人挑战与安全 Robotics and Scaling

Host

是的,是的。在机器人学中,这非常明确。有运动和控制的问题。运动问题解决得更好,尤其是在学习领域。但就像最初的蛋白质折叠系统一样,引入传统的基于模型的方法有很多价值。你不太可能仅仅通过端到端学习来解决控制或全身局部控制问题。

Yeah, yeah. In robotics it's very explicit. So there's the problem of locomotion and manipulation. Locomotion is much more solved, especially in learning domain. But there's a lot of value, just like with the initial protein folding systems, bringing in the traditional model-based methods. It's unlikely that you can just learn the manipulation or the whole body local manipulation problem end to end.

Nathan Lambert

那是梦想。但当你看到人类手的神奇和现实世界的复杂性时,你会意识到像 AlphaFold 2 那样完全学习是非常困难的。

That's the dream. But then you realize when you look at the magic of the human hand and the complexity of the real world, you realize it's really hard to learn this all the way through, the way I guess AlphaFold 2 did.

Host

我对机器人学习领域感到兴奋,我认为它总体上被语言模型的兴奋和投资所推动,它们获得了训练 Transformer 的基础设施,这是一种通用建模,变成了世界级的工业工具,机器人学的任何限制都得到了改善。算力更多了。在此基础上,他们将这些语言模型用作中心单元,可以在已经有效的东西周围进行有趣的探索性工作。然后我看到它像我们讨论的 Hugging Face Transformers 和 Hugging Face 一样出现。我在 Hugging Face 时曾试图推动这件事,但为时过早。就像 Hugging Face 上的开放机器人模型,能够贡献数据并微调它们。我认为我们现在更接近了,因为对机器人学的投资,我认为自动驾驶汽车与此相关并促成了这一点,一旦你达到可以拥有这种生态系统的程度,有人可以下载机器人模型,也许微调到他们的机器人,或者在世界各地共享数据集。这方面有一些工作,比如 RTX,我想是几年前,人们试图这样做。但我认为一旦有了这个生态系统,情况会大不相同。而 ChatGPT 后的整个热潮正在为此投入更多资源,我认为这是一个非常好的研究领域。

I'm excited about the robotic learning space, so I think it's collectively getting supercharged by all the excitement and investment in language models generally, where they're getting the infrastructure for training transformers, which is a general modeling thing, becoming world-class industrial tooling, where any limitation for robotics is just way better. There's way more compute. And then on top of that, they take these language models and use them as kind of central units where you can do interesting explorative work around something that already works. And then I see it emerging as kind of like we talked about Hugging Face Transformers and Hugging Face. I think when I was at Hugging Face, I was trying to get this to happen, but it was too early. It's like these open robotic models on Hugging Face and being able to contribute data and fine-tune them. I think we're much closer now that the investment in robotics and I think self-driving cars is related and enables this, where once you get to the point where you can have this sort of ecosystem where somebody can download a robotics model and maybe fine-tune it to their robot or share data sets across the world. There's some work in this area, like RTX, I think it was a few years ago, where people are trying to do that. But I think once they have this ecosystem, it'll look very different. And then this whole post-ChatGPT boom is putting more resources into that, which I think is a very good area for doing research.

Nathan Lambert

这也导致了更好、更准确、更逼真的模拟器被构建,缩小了机器人领域的模拟到现实差距。但是,你提到了机器人领域的很多兴奋和投资。其缺点是,在炒作周期中,我个人认为大多数机器人专家认为机器人不会在隐含或明确承诺的时间尺度上被解决。所以当所有这些机器人公司涌现,然后它们没有可行的产品时,就会发生这种兴奋的崩溃,这令人紧张。

This is also resulting in much better, more accurate, more realistic simulators being built, closing the sim-to-real gap in the robotic space. But, you know, you mentioned a lot of excitement in the robotic space and a lot of investment. The downside of that, which happens in hype cycles, I personally believe most robotics people believe that it's not robotics is not going to be solved at the time scale as being kind of implicitly or explicitly promised. And so what happens when there's all these robotics companies that spring up and then they don't have a product that works then there's going to be this kind of crash of excitement which is nerve-racking.

AGI 与 ASI 时间线 Robotics challenges and safety

Host

希望会有其他东西介入并持续推动这些想法的发展。

There's the hopefully something else will come in and keep swooping in so that the continued development of some of these ideas keeps going.

Nathan Lambert

这本质上与持续学习问题相关,现实世界非常复杂。对于大语言模型,你不需要让模型为用户学习,因为有很多事情是每个人都必须做的。每个人可能都想修正邮件中的语法或编写代码之类的事情。这更受约束,所以你可以为模型做好准备。但让机器人为现实世界做好准备就更难了。我的意思是,你有基础模型,机器人基础模型,但你可以学习某些东西,比如抓取物体,但话说回来,我认为每个人的家都是不同的,你知道,差异很大,这正是机器人需要在实际工作中学习的地方。我认为这就是当前的瓶颈,如何即时定制。

Related to the continual learning issue essentially where the real world is so complex. With LLMs, you don't need to really have something learn for the user because there are a lot of things everyone has to do. Everyone maybe wants to fix their grammar in their email or code on something like that. It's more constrained so you can kind of prepare the model for that. But preparing the robot for the real world is harder. I mean you have the foundation models, the robotic foundation models, but you can learn certain things like grasping things, but then again, I think everyone's house is different, you know, it's so different and that is where the robot would have to learn on the job essentially. I think that is the bottleneck right now, how to customize it on the fly.

Host

我认为我无法低估安全的重要性,而机器人领域的人或其他几乎没有人谈论它。我们讨论的所有有趣复杂性、所有失败模式和失败案例,我们一直在谈论大语言模型,有时它们会以有趣的方式失败。所有这些在大语言模型领域都是游戏。但在机器人领域,在人们的家中,经过数百万分钟、数十亿次交互,你几乎永远不允许失败。当你有具身系统被部署到现实世界中时,你必须解决许多你在考虑通用机器人学习问题时从未想过要解决的问题。

I don't think I can possibly understate the importance of the thing that doesn't get talked about almost at all by robotics folks or anyone is safety. All the interesting complexities we talk about learning, all the failure modes and failure cases, everything we've been talking about LLMs, sometimes it fails in these interesting ways. All of that is fun and games in the LLM space. In the robotic space, in people's homes across millions of minutes, billions of interactions, you're really almost allowed to fail, never. When you have embodied systems that are put out there in the real world, you just have to solve so many problems you never thought you'd have to solve when you're just thinking about the general robot learning problem.

Nathan Lambert

我非常看衰面向消费者购买的家用学习机器人。我非常看好自动驾驶汽车,也非常看好机器人自动化,例如亚马逊的分销系统,亚马逊建造了全新的分销中心,首先为机器人设计,而不是为人类。

I'm so bearish on in-home learned robots for consumer purchase. I'm very bullish on self-driving cars, and I'm very bullish for robotic automation, e.g., like Amazon distribution, where Amazon has built whole new distribution centers designed for robots first rather than humans.

Host

在 AI 圈子里,有很多兴奋点,认为 AI 能在大规模制造中实现大量自动化,我确实认为机器人走这条路更合理,即设计和优化来做重复性任务,这些任务人类可以完成但不想做。但这也需要比人们可能预测的更长的时间。我认为从 AI 奇点到我们现在可以在美国大规模扩大制造业,因为我们有巨大的 AI 优势,这一飞跃受到许多政治和其他挑战性问题的困扰。

There's a lot of excitement in AI circles about AI enabling a lot of automation in mass-scale manufacturing, and I do think that the path to robots doing that is more reasonable, where it's like a thing that is designed and optimized to do a repetitive task that a human could conceivably do, but doesn't want to. And then I'm much so But it's also going to take a lot longer than people probably predict. I think that the leap from AI singularity to we can now scale up mass manufacturing in the US because we have a massive AI advantage is one that is troubled by a lot of political and other challenging problems.

AI 与软件开发未来 Timelines to AGI and ASI

Host

我们来谈谈时间线。具体来说是 AGI 或 ASI 的时间线。作为起点,说没有人真正同意 AGI 和 ASI 的定义,这公平吗?

Let's talk about timelines. Specifically timelines to AGI or ASI. Is it fair, as a starting point, to say that nobody really agrees on the definitions of AGI and ASI?

Nathan Lambert

我认为存在很多分歧,但我也受到反驳,很多人说的其实差不多,比如一个能完成大部分数字经济工作的东西。远程工作者就是一个相当合理的例子,我认为 OpenAI 的定义与此有些相关,即一个能做很多经济上有价值任务的 AI,我不太喜欢这个定义,但我认为它可以作为一个基础点,因为今天的语言模型虽然非常强大,但还不能直接替代远程工作者,而且有些任务比远程工作难得多,比如做出意想不到的科学发现,你甚至无法暂停它,这会被认为是超级智能问题。或者像读取所有医疗记录,发现人们不知道的疾病关联,或者发现某种常见药物可以治疗某种罕见癌症。他们会说这是超级智能的事情。所以这些是自然的层级。我的问题是,这变得与 AI 意义的追寻及其宗教方面深深交织在一起。所以你可以走不同的路径。我甚至不知道远程工作是否是一个好的定义,因为那到底是什么?它就像完美的工具使用。

I kind of think there's a lot of disagreement, but among I've been getting pushed back where a lot of people kind of say the same thing, which is like a thing that could reproduce most digital economic work. So, like the remote worker is a fairly reasonable example, and I think OpenAI's definition is somewhat related to that, which is like an AI that can do a lot of economically valuable tasks, which I don't really love as a definition, but I think it could be a grounding point because language models today, while immensely powerful, are not this remote worker drop-in, and there are things that you can think of that could be done by an AI that are way harder than remote work, which are like solving a finding an unexpected scientific discovery that you couldn't even pause it, which would be an example of something that somebody says is like an artificial super intelligence problem. Or like taking in all medical records and finding linkages across certain illnesses that people didn't know, or figuring out that some common drug can treat some niche cancer. Like they would say that that is like a super intelligence thing. So, these are kind of natural tiers. My problem with it is that it becomes deeply entwined with the quest for meaning of AI and this religious aspects to it. So, there's kind of different paths you can take it. And I don't even know if the remote work is a good definition because what exactly is that? It's like perfect tool use.

Host

我实际上喜欢最初名为 AI 27 的报告。他们更关注代码和研究任务。所以目标是超人编码者。他们有几个里程碑系统:超人编码者、超人 AI 研究员、然后超级智能 AI 研究员,最后是完整的 ASI,人工超级智能。但一旦你开发出超人编码者,其他一切都会很快跟上。那里的任务是实现完全自主的编码。所以,为了进行研究所需的任何编码都完全自动化。从那里开始,人类将与那个系统一起进行 AI 研究,他们很快就能开发出一个实际上能为你做研究的系统。这就是想法。最初他们的预测是 2027-28 年,现在他们将其推迟了 3 到 4 年,到 2031 年平均预测。可能我的预测是 2031 年,但至少你可以具体地思考完全自动化编程有多难。

I actually like the originally titled AI 27 report. They focus more on code and research tasks. So, the target there is the superhuman coder. So, they have several milestone systems: superhuman coder, superhuman AI researcher, then superintelligent AI researcher, and then the full ASI, artificial superintelligence. But after you develop the superhuman coder, everything else follows quickly. There, the task is to have fully autonomous coding. So, any kind of coding you need to do in order to perform research is fully automated. And from there, humans would be doing AI research together with that system, and they would quickly be able to develop a system that actually can do the research for you. That's the idea. And then initially their prediction was 2027-28 and now they've pushed it back by 3 to 4 years to 2031 mean prediction. Probably my prediction is it would be on 2031 but at least you can get concrete way think about how difficult it is to fully automate programming.

Nathan Lambert

是的,我确实同意他们的一些假设和关于事情如何发展的动态,但我认为他们在定义具体里程碑和讲述有用故事方面做得很好,这就是为什么 AI 2027 文档的影响力超越了硅谷,因为他们讲了一个好故事,并做了大量严谨的工作。我认为我所属的阵营是,AI 是所谓的锯齿状,在某些方面非常出色,在某些方面非常糟糕。所以我认为当他们接近这个自动化软件工程师时,它擅长的将是传统机器学习系统,比如前端模型很擅长,但分布式机器学习模型实际上非常糟糕,因为关于大规模分布式学习等内容的训练数据很少,这是我们已经看到的,我认为这只会被放大,然后这些权衡变得更加混乱,然后还有关于 AI 研究如何运作等等的问题。

Yeah, I do agree with some of their presumptions and dynamics on how it would play out but I think they did a good job in the scenario defining milestones that are concrete and to tell a useful story which is why the reach for this AI 2027 document well transcend and Silicon Valley is because they told a good story and they did a lot of rigorous work to do this. I think that camp that I fall into is that like AI is so-called jagged which will be excellent at some things and really bad at some things. So I think that when they're close to this automated software engineer what it will be good at is that traditional ML systems in front end the model is excellent at but the distributed ML the models are actually really quite bad at because there's so little training data on doing large scale distributed learning and things and this is something that we already see and I think this is just getting amplified and then it's kind of messier in these trade-offs and then there's like how do you think AI research works and so on.

Host

所以你认为基本上超人编码者几乎是无法实现的,意思是由于事物的锯齿状性质,你总是会有能力上的差距。

So you think basically superhuman coder is almost unachievable meaning like because of the jagged nature of the thing you're just always going to have gaps in capabilities.

用 LLM 的软件工程 AI and Software Development Future

Nathan Lambert

我认为这是在给某些东西赋予完整性,模型在某些类型的代码上已经超越了人类,而且这种情况会持续下去。人们很有创造力,他们会利用这些不可思议的能力来弥补模型的弱点,并快速推进。这始终是人类与模型之间的一种舞蹈,人类在弥补模型做不到的事情。最优秀的 AI 研究者是那些能够激发这种超能力的人。我认为这和我们已经看到的情况是一致的。例如,用 Claude Code 建网站,你可以在几小时内搭建一个漂亮的网站,或者做数据分析。它会在这方面越来越好,并在这个过程中掌握新的编码技能。联系到大科技公司正在发生的事情,AI 2027 报告倾向于奇点理论,但我认为研究是混乱的、社会性的,并且很大程度上依赖于数据,而 AI 模型无法处理这些。不过,我们今天拥有的东西已经非常强大,科技公司正在集体投入数百亿美元。所以我们会得到比现在好得多的 ChatGPT 和 Claude Code 版本。很难预测这会走向何方,但未来的光明前景正是为什么一些世界上最有权势的人投入如此多资金的原因。这里有一些细微的差别:我们实际上不知道更好的 ChatGPT 是什么样子,但它能自动化 AI 研究吗?我认为至少在目前这个时间框架内可能不行。大科技公司花一千亿美元的速度会比我们得到一个能够引发 AI 研究奇点的自动化 AI 研究员快得多。

I think it's assigning completeness to something where the models are kind of superhuman at some types of code and I think that will continue. People are creative, so they'll utilize these incredible abilities to fill in the weaknesses of the models and move really fast. It'll always be this dance between humans enabling what the model can't do. The best AI researchers are the ones who can enable this superpower. I think this aligns with what we already see. For example, Claude Code for building a website lets you stand up a beautiful website in a few hours or do data analysis. It'll keep getting better at these things and pick up new coding skills along the way. Linking to what's happening in big tech, the AI 2027 report leans into the singularity idea, but I think research is messy, social, and largely in the data in ways that AI models can't process. However, what we have today is really powerful, and tech companies are collectively investing tens of billions of dollars. So we will get much better versions of ChatGPT and Claude Code. It's hard to predict where that's going, but the bright clarity of that future is why powerful people are putting so much money into this. There are small differences: we don't actually know what a better version of ChatGPT is, but can it automate AI research? I'd say probably not, at least in this timeframe. Big tech will spend a hundred billion dollars much faster than we get an automated AI researcher that enables a singularity.

Host

那么你认为你的预测是什么?如果这算是一个有用的里程碑,我们还要等 10 年以上吗?

So you think your prediction would be what? Like if this is even a useful milestone, we're more than 10 years out.

Nathan Lambert

在软件方面,我认为时间会更短,但在研究方面,我认为会更长。

I would say less than that on the software side, but I think longer than that on things like research.

Host

让我们为了好玩,试着想象一个所有软件编写都完全自动化的世界。你能想象那个世界吗?

Let's just for fun try to imagine a world where all software writing is fully automated. Can you imagine that world?

Nathan Lambert

到今年年底,自动化的软件数量会非常高。但像用强化学习训练模型、需要多组 GPU 相互通信这样的事情,仍然会很困难,但我觉得会容易得多。

By the end of this year, the amount of software that'll be automated will be so high. But it'll be things like trying to train a model with RL and needing multiple bunches of GPUs communicating with each other. That'll still be hard, but I think it'll be much easier.

Host

思考这个问题的一种方式是编程的完全自动化。想想写出的有用代码行数,与参与其中的人类数量的比例。大概在很长一段时间内,软件编写中仍然会有人的参与。只是相对于写出的代码量,人越来越少,对吧?超级编码器的假设是,参与其中的人类数量会降到零。当参与其中的人类数量是几百人,而不是几十万人时,那个世界会是什么样子?

One way to think about this is the full automation of programming. Just think of lines of useful code written, the fraction of that to the number of humans in the loop. Presumably, there'll be for a long time humans in the loop of software writing. It'll just be fewer and fewer relative to the amount of code written, right? The superhuman coder presumption is that it goes to zero, the number of humans in the loop. What does that world look like when the number of humans in the loop is in the hundreds, not in the hundreds of thousands?

Nathan Lambert

我认为软件工程将更多地转向系统设计和结果目标。我确实认为软件将发生巨大变化。过去几周这种情况一直在发生,人们从一个月前说'哦,是的,智能体有点垃圾'——这是 Karpathy 的名言——到现在有点调侃软件工业化,任何人都可以轻松创建软件。我认为我们更接近那个方向了,需要指导和理解系统如何工作,才能从语言模型中提取最佳效果。很难接受软件开发的变革程度,以及有多少人可以在不接触代码的情况下完成事情。

I think software engineering will be driven more to system design and goals of outcomes. I do think software is largely going to change. This has been happening over the last few weeks where people went from a month ago saying, 'Oh, yeah, agents are kind of slop,' which is a famous Karpathy quote, to a bit of a meme about the industrialization of software where anyone can just create software at their fingertips. I do think we are closer to that side of things, and it takes direction and understanding how the systems work to extract the best from language models. It's hard to accept the gravity of how much is going to change with software development and how many more people can do things without ever looking at it.

Host

有趣的是思考这些系统是否会独立,完全独立。我毫不怀疑 LLM 在某个时候会像计算器解决计算问题一样解决编码问题。在某个时候,人类开发出工具,你不再需要人类来计算那个数字。你只需输入,它就是一个算法。我认为编码可能也是如此。但问题是:是否仍然需要人类让 AI 做某事,比如一个人说'建那个网站',还是会有 AI 独立地建网站?

What's interesting is to think about whether these systems will be independent, completely independent in the sense that... Well, I have no doubt that LLMs will at some point solve coding in a sense like calculators solve calculating. At some point humans develop the tool that you never need a human to calculate that number. You just type it in and it's an algorithm. I think that's the same probably for coding. But the question is: will you still have humans asking the AI to do something, like a person saying 'build that website,' or will there be AI that just builds websites independently?

Nathan Lambert

我认为谈论建网站太简单了。网站和 HTML 的问题在于,它对垃圾输出非常宽容。它会给你显示垃圾。它很擅长显示垃圾。我更愿意考虑安全关键系统,比如让 AI 端到端生成管理物流或管理汽车和车队的系统。端到端地为你生成东西。

I think using talking about building websites is too simple. The problem with websites and HTML is that it's very resilient to just slop. It will show you slop. It's good at showing slop. I would rather think of safety-critical systems, like asking AI to end-to-end generate something that manages logistics or manages cars and fleets of cars. End-to-end generate stuff for you.

Host

我认为一个更中等的例子是像 Slack 或 Microsoft Word 这样的东西。我认为如果组织允许,AI 可以很容易地端到端实现功能,并且对于你想尝试的事情做得相当好。你想在 Slack 中添加一个新标签,我认为 AI 能做得很好。

I think a more intermediate example is take something like Slack or Microsoft Word. I think if organizations allow it, AI could very easily implement features end-to-end and do a fairly good job for things you want to try. You want to add a new tab in Slack that you want to use, and I think AI will be able to do that pretty well.

Nathan Lambert

实际上,这是一个非常好的例子。我们离那还有多远?

Actually, that's a really great example. How far away are we from that?

Host

大概今年。

Like this year.

Nathan Lambert

看吧,我不知道。我不知道。

See, I don't know. I don't know.

Host

我想我不知道生产代码库有多糟糕,但我认为在几年内,很多人会被推向更像设计师和产品经理的角色,你会有多个智能体为你尝试事情。它们可能需要一到两天来实现一个功能或尝试修复一个 bug,你会有仪表盘,我认为 Slack 实际上是一个很好的仪表盘,你的智能体会和你对话,你给出反馈。但像制作一个过得去的网站 logo——连贯的设计和风格——对模型来说会非常困难,以及决定下一步添加什么。

I guess I don't know how bad production code bases are, but I think within on the order of a few years, a lot of people are going to be pushed to be more like a designer and product manager where you have multiple agents that can try things for you. They might take one to two days to implement a feature or attempt to fix a bug, and you have dashboards, which I think Slack is actually a good dashboard where your agents will talk to you and you'll give feedback. But things like making a website logo that's passable—cohesive design and style—will be very hard for models, and deciding what to add next.

Nathan Lambert

我只是……好吧,我和很多程序员混在一起,其中一些人总体上有点怀疑。他们就是那种氛围。我只是觉得在复杂系统中添加功能涉及很多复杂性。

I just... okay, so I hang out with a lot of programmers and some of them are a little bit on the skeptical side in general. That's just vibe-wise they're like that. I just think there's a lot of complexity involved in adding features to complex systems.

经济影响与工具使用 Software engineering with LLMs

Host

就像你看浏览器 Chrome。如果我想加个功能,比如标签页不在顶部而在左侧。界面,对吧?我觉得这不是明年就能实现的事。

Like if you look at the browser, Chrome. If I wanted to add a feature, if I wanted to have tabs as opposed to up top, I want them on the left side. Interface, right? I think we're not this is not a next year thing.

Nathan Lambert

今年 Claude 的一次发布中,他们的一个测试是:给 Claude 一个软件,让它完全重新创建。它已经能几乎从零重建 Slack,只要给定软件参数并放在沙盒环境中。

One of the Claude releases this year, one of their tests was we give it a piece of software and leave Claude to run to recreate it entirely. And it could already almost rebuild Slack from scratch just given the parameters of the software and left in a sandbox environment.

Host

这部分我几乎更喜欢。

Part I like almost better.

Nathan Lambert

所以可能更小更新的公司有优势,它们觉得我们不必有臃肿和复杂性,因此这个未来是存在的。

So it might be that the smaller newer companies are advantaged and they're like, we don't have to have the bloat and complexity and therefore this future exists.

Host

我认为这触及了你提到的点:你交谈的一些人持怀疑态度,我觉得不是因为语言模型做不到某些事,而是因为人们不希望它用这种方式做。

And I think this gets to the point that you mentioned that some people are you talk to are skeptical and I think that's not because the LM can't do XYZ, it's because people don't want it to do it this way.

Nathan Lambert

部分可能是人类这边的技能问题。不幸的是,我们必须对自己诚实。部分可能是规格不足的问题。编程时你假设语言模型应该能读心,就像人际关系和友谊中的沟通问题。我认为这正是规格驱动设计重要的地方。你只需用自然语言指定你想要什么。

Some of that could be a skill issue on the human side. Unfortunately, we have to be honest with ourselves. And some of that could be an underspecification issue. So programming like you're like you're just assuming this is like in relationships and friendships communication type issue. You're assuming the LM somehow supposed to read your mind. I think this is where spec-driven design is really important. Like you just using natural language specify like what you want.

Host

就像如果你和实验室的人聊,他们在训练和生产代码中使用这些工具。比如 Claude Code 就是用 Claude Code 构建的。他们都广泛使用这些东西,Dario 也谈到 Claude 自己的代码有多少是用 Claude 写的。这些人在能力上稍微领先,他们可能在推理上花费更多。他们可能花费我们 10 到 100 倍以上。我们只是每月 100 或 200 美元的低级套餐。他们真的放手去用。而且我认为以我们现在的进步速度,一年前我们还没有 Claude Code,也没有真正的推理模型,而今天我们能做的和这些模型能做的差距很大,而且似乎有很多低垂的果实可以改进。失败模式相当愚蠢。比如 Claude,你尝试使用我没有安装的 CLI 命令 14 次,然后我发给你要运行的命令。从建模角度看,这是相当可修复的。所以我不认为

That's like if you talk to people at the labs, they use these in their training and production code. Like Claude code is built with Claude code. And they all use these things extensively and Dario talks about how much of Claude's code own and it's like these people are slightly ahead in terms of the capabilities they have and they probably spend on inference. They could spend 10 to 100 plus X as much as we're spending. Like we're on a lowly 100 or $200 a month plan. Like they truly let it rip. And I think that that like with the pace of progress that we have, it seems like where a year ago we didn't have Claude code and we didn't really have reasoning models and it's like the difference between sitting here today and what we can do with these models and it seems like there's a lot of low hanging fruit to improve them. The failure modes are pretty dumb. It's like Claude, you tried to use the CLI command I don't have installed 14 times and then I sent you the command to run. It's like that thing from a modeling perspective is pretty fixable. So I don't

Nathan Lambert

我同意你的看法。我总体上越来越乐观。针对你阐述的观点,我认为这是人类的技能问题。所以 Anthropic 或其他公司在理解如何最好地使用模型进行编程方面领先,因此他们有效地使用它们。我认为有很多边缘程序员。他们不太清楚如何使用,没有很好的指南。人们正在努力弄清楚。

I agree with you. I've been becoming more and more bullish in general. Speaking to what you're articulating, I think it is a human skill issue. So Anthropic is leading the way in the or other companies in understanding how to best use the models for programming and therefore they're effectively using them. I think there's a lot of programmers on the outskirts. They're like they don't I mean there's not a really good guide in how to use them. People are trying to figure it out exactly.

Host

这可能非常昂贵。比如入门价格可能是每月 2000 美元,只有科技公司和富人能承受。可能就是这样。

It might be very expensive. Like it might be that the entry point for that is $2,000 a month which is only tech companies and rich people. Which is like like that could be it.

Nathan Lambert

但可能值得。如果最终结果是一个可用的软件系统,那可能值得。不过,有趣的是我们如何从 AGI 时间线的讨论转向更务实有用的东西。关于 AGI 和 ASI 的时间线,有什么具体、有趣、有用且深刻的话可说吗?这些讨论是否有点脱离日常?有一些有趣的赌注。很多人试图在真实科学领域做带可验证奖励的强化学习,有些初创公司有数亿美元资金和湿实验室,让语言模型提出假设并在现实世界中测试。我认为它们还很早期,但以进步的速度,就像

But it might be worth it. I mean if the final result is a working software system, it may be worth it. But by the way, it's funny how we converge from the discussion of timeline to AGI to something more pragmatic and useful. Is there anything concrete and interesting and useful and profound to be said about timeline to AGI and ASI? Are these discussions a bit too detached from the day-to-day. There's interesting bets. So, there's a lot of people trying to do reinforcement learning with verifiable rewards, but in real scientific domains where there's startups that are spending like they have hundreds of millions of dollars of funding and they have wet labs where they're having language models propose hypotheses that are tested in the real world. And I would say that I think they're very early or they're early, but with the pace of progress, it's like

Host

是的。

Yeah.

Nathan Lambert

也许他们早 6 个月,因为先发而成功;也许早 8 年,你根本不知道。所以,这种将势头扩展到其他科学的登月计划,如果像 AlphaFold 那样的时刻由初创公司在各种科学领域实现,那将非常具有变革性。我认为有些初创公司,比如 Harmonic,全力投入语言模型加 Lean 做数学。我想你最近另一位播客嘉宾也谈到这个,我们不知道在那模型上花一亿美元会得到什么。大多数会失败,但少数可能带来重大突破,与 ChatGPT 或 Claude Code 这类软件体验截然不同。比如一个只对数学博士有用,但让他们效率提高百倍的工具。

Maybe they're early by 6 months and they make it because they were there first or maybe they're early by 8 years and you don't really know. So, I think that that type of moonshot to branch this momentum into other sciences is like, okay, like that would be very transformative if like AlphaFold moments happen in all sorts of other scientific domains by like a startup solving this. I think there are startups, I think maybe Harmonic is one where they're going all in on language models plus Lean for math. I think you had another podcast guest who talked about this recently and it's like we don't know exactly what's going to fall out of spending a hundred million dollars on that model. And most of them will fail, but a couple of them might be big breakthroughs that are very different than ChatGPT or Claude code type software experiences. Like a tool that's only good for a PhD mathematician, but makes them a hundred X effective. Like

Host

好的,我同意。我认为这会在很多领域发生,尤其是资源丰富的领域,比如金融、法律和制药公司。但话说回来,这真的是 AGI 吗?因为我们又在专门化它。这真的和过去我们有专门算法有多大区别?我认为只是更复杂了,但我不确定,什么时候才叫 AGI?我认为真正酷的是我们有可以专门化的基础模型。我认为这本身就是某种突破。目前,我们还没到那一步,因为首先太贵,而且 ChatGPT 并不开放定制。但我可以想象一种商业模式:OpenAI 对某家银行说,一亿美元,我们为你做定制模型。那将是巨大的经济价值。另外,公司现在有什么差异化因素?如果每个人都用同一个大语言模型,比如 ChatGPT,他们又会做同样的事。大家都步调一致,但公司通常想要竞争优势,所以必须使用私有数据、实验和专门化。这会很有趣。

Okay, I agree. I think this will happen in a lot of domains, especially also like domains that have a lot of resources like finance and legal and pharmaceutical companies. But then again, is it really AGI again because we are now specializing it again. And then again, is it really that much different from back in the day how we had specialized algorithms. It's I think it's just the same thing more way more sophisticated, but I don't know, is there a threshold when we call it AGI I guess. I think the real cool thing is here that we have like the foundation models that we can specialize. I think that's like the breakthrough at some point. Right now, I think we are not there yet because well, first it's too expensive, but also you know, like ChatGPT doesn't just give away that ChatGPT to customize it. I think once that's going to be true in a some way and I think I can imagine this as a business model that ChatGPT OpenAI says at some point like, "Hey, you know, Bank of America, for 100 million, we will do your custom model." or something like that. And I think that will be the huge economic value add. The other thing though is also companies, I mean, right now, what is the differentiating factor? I mean, if everyone uses the same LLM, if everyone uses ChatGPT, they will all do the same thing again. I mean, then well, it's everyone is moving in lockstep, but usually companies they want to have a competitive advantage and I think there's then no way around using some of their private data and experimenting and maybe specializing. It's going to be interesting, yeah.

Nathan Lambert

身处进步的步伐中,确实感觉事情正在到来。我不认为 AGI 和 ASI 的门槛特别有用。

Sitting in the pace of progress, it does just feel like things are coming. I don't think the AGI and ASI thresholds are particularly useful.

需要新想法 Economic Impact and Tool Use

Host

我认为真正的问题——这也把我们带回到远程工作者的话题——是我们什么时候才能看到经济影响上的重大飞跃?因为目前来看,LLM 模型还没有带来明显的经济影响。抛开 AGI 或 ASI 这些不谈,一个现实的问题是,我们什么时候才能看到 GDP 的跃升?

I think the real question, and this takes us to the road worker thing, is when are we going to see a big, obvious leap in economic impact? Because currently there hasn't been an obvious leap in economic impact of LLM models, for example. And aside from AGI or ASI or all that, there's a real question of when we'll see a GDP jump.

Nathan Lambert

是啊,GDP 是由什么构成的?很大一部分是金融服务,所以我不太确定。对我来说,很难想象 GDP 的跃升,但我想说,当你不再需要看代码时,软件开发会以不同的方式变得有价值。所以当 Claude 能帮你创办一个小企业——搭建网站、开银行账户、设置邮箱等等——而你只需要表达你想向世界传递什么,那就不只是企业市场了。但让人们去尝试很难。我想如果 ChatGPT 能做到,人们已经在用 ChatGPT 了。

Yeah, what is GDP made of? A lot of it is financial services, so I don't know. It's hard for me to think about the GDP bump, but I'd say software development becomes valuable in a different way when you no longer have to look at the code. So when Claude can make you a small business—set up your website, bank account, email, and whatever else—and you just have to express what you're trying to put into the world, that's not just an enterprise market. But it's hard to get people to try that. I guess if ChatGPT can do it, people are trying ChatGPT.

Host

我认为这归结为一个科学问题:工具使用有多难解决。你提到的很多东西,比如远程工作,都属于工具使用。就像计算机使用——如何让一个 LLM 作为智能体系统在现实世界中执行任务,并且只有 1% 的出错率。

I think it boils down to the scientific question of how hard tool use is to solve. A lot of what you're implying, the remote work stuff, is tool use. It's like computer use—how you have an LLM that goes out there as an agentic system and does something in the world, only screwing up 1% of the time.

Nathan Lambert

计算机使用是实验室关注的一个好例子,但我们没看到太多进展。2025 年我们看到了多个演示,比如 Claude 使用你的电脑,或者 OpenAI 的 Kua,但它们都很糟糕。他们在这方面投入资金,但接管整个屏幕似乎比在后端调用 API 难得多。你必须为模型设置一个不同的工作环境。它们不是在你的 MacBook 上运行,而是分别与 Google、Amazon 和 Slack 交互,处理方式与人类截然不同。所以其中一些可能是结构性障碍。

Computer use is a good example of what labs care about, and we haven't seen a lot of progress. We saw multiple demos in 2025 of Claude using your computer or OpenAI's Kua, and they all suck. They're investing money in this, but taking over the whole screen seems much harder than having an API they can call in the backend. You have to set up a different environment for the model to work in. They're not working on your MacBook; they're individually interfacing with Google, Amazon, and Slack, handling things very differently than humans. So some of this might be structural blockers.

Host

另外,从规范的角度看,对于任意任务,你仍然需要指定你想让 LLM 做什么。在环境中如何做到这一点?你可以说最终目标,但如果它无法用 LLM 解决,你就需要分解成子步骤。如何将这些信息输入到一个预订旅行的系统中?你可以说,“你搞砸了我的信用卡信息”,但即使要让它达到那个程度,在它尝试之前你如何引导模型?我认为界面非常难。

Also, specification-wise, for arbitrary tasks, you still have to specify what you want your LLM to do. How do you do that in the environment? You can say the end goal, but if it can't solve it with LLMs, you clarify with sub-steps. How do you put that information into a system that books a trip? You can say, 'You screwed up my credit card information,' but even to get it to that point, how do you guide the model before it can attempt that? I think the interface is really hard.

Nathan Lambert

是的,它需要了解很多关于你的特定信息,以及普遍犯的错误,还有你特有的错误。

Yeah, it has to learn a lot about you specifically, and about the general mistakes made throughout, and then mistakes specific to you.

Host

界面正在被设计成向人类请求输入。我认为 Claude Code——我们经常讨论的——会在对你的计划或期望目标没有足够规范时请求反馈。它会开始提问。我们讨论过跨聊天保存的记忆,它的首次实现有点奇怪——它会在聊天中提及我狗的名字。我就想,“你不需要这么含蓄。”但正在涌现的东西包括 ChatGPT 的 Pulse 功能,它是精心挑选的几段文字,附有链接,供你查看或讨论。人们谈论语言模型会问你问题,我认为这很可能行得通。模型知道你有个医生预约,然后问,“嘿,之后感觉怎么样?”这进入了人类非常容易受影响的领域,未来会有很多社会变革。他们正在尝试让模型更主动。有些人真的很喜欢 Pulse,它会处理你的聊天记录并自动搜索信息,放入 ChatGPT 应用。所以有很多东西即将到来。

Interfaces are being set up to ask humans for input. I think Claude Code, which we talk about a lot, asks for feedback if it doesn't have enough specification on your plan or desired outcome. It starts asking questions. We talked about memory that saves across chats, which in its first implementation is kind of odd—it would mention my dog's name in a chat. I'm like, 'You don't need to be subtle about this.' But things emerging include ChatGPT's Pulse feature, which is a curated couple of paragraphs with links to something to look at or talk about. People talk about language models asking you questions, which I think will probably work. The model knows you had a doctor's appointment and asks, 'Hey, how are you feeling after that?' That goes into territory where humans are very susceptible, and there's a lot of social change to come. They're experimenting with models being engaged. Some people really like Pulse, which processes your chats and automatically searches for information, putting it in the ChatGPT app. So there's a lot coming.

Nathan Lambert

我以前用过那个功能,总是觉得不好意思,因为它每天都做,而我很少去看。感觉就像花了很多钱——计算机在我甚至不看的东西上消耗算力。

I used that feature before, and I always feel bad because it does that every day and I rarely check it out. It's like how much money—computers burned on something I don't even look at.

Host

世界上有很多闲置算力,所以别太内疚。

A lot of idle compute in the world, so don't feel too bad.

Nathan Lambert

好吧。

Okay.

能力平台期与经济影响 Need for New Ideas

Host

你认为是否需要新的想法?通往 AGI 的道路——无论我们如何定义——要更普遍地解决计算机使用,解决生物学、化学和物理学,也就是 Dario 对 AGI 或强大 AI 的定义——你认为是否需要全新的想法?非 LLM、非强化学习的想法。它们可能是什么样的?这有点进入哲学领域了。

Do you think new ideas might be needed? Is it possible that the path to AGI—however we define that—to solve computer use more generally, to solve biology, chemistry, and physics, sort of the Dario definition of AGI or powerful AI—do you think it's possible that totally new ideas are needed? Non-LLM, non-RL ideas. What might they look like? This is going into philosophy a bit.

Nathan Lambert

对于像奇点这样的事情发生,我会说是的。新想法可能是架构或训练算法,这些都是深度学习的基础。但它们很难预测。我认为即使没有这些进展,我们也不会走得太远。我们可能会得到软件解决方案,但如果没有更多创新,它可能止步于软件,无法实现计算机使用。所以会有很多进步,但如果你放眼长远,未来 30 年内仍会有一些想法,看起来像是开启下一篇章的重大科学创新。我不知道它是在 1 年内还是 15 年内到来。

For something like a singularity to happen, I would say yes. New ideas could be architectures or training algorithms, which are fundamental deep learning things. But they're hard to predict. I think we won't get very far even without those advances. We might get the software solution, but it might stop at software and not do computer use without more innovation. So a lot of progress will come, but if you zoom out, there are still ideas in the next 30 years that will look like major scientific innovations enabling the next chapter. I don't know if it comes in 1 year or 15 years.

Host

是啊,我想知道如果苦涩的教训在未来 100 年仍然成立,那会是什么样子。

Yeah, I wonder if the bitter lesson holds true for the next 100 years, what that looks like.

Nathan Lambert

如果缩放定律在深度学习中是最根本的,我认为苦涩的教训将永远适用。算力会变得更充裕,但即使在充裕的算力中,那些具有更陡峭缩放定律斜率或更好偏移量的方法会胜出。这是一个性能与算力的二维图。

If scaling laws are fundamental in deep learning, I think the bitter lesson will always apply. Compute will become more abundant, but even within abundant compute, the ones with a steeper scaling law slope or a better offset will win. It's a 2D plot of performance and compute.

Host

那可能就像字面意义上的、带有太阳能板的计算机集群绕地球轨道运行。

It might be something like literally computer clusters orbiting Earth with solar panels.

Nathan Lambert

问题在于散热。你会受到太阳的所有辐射,而且没有空气来散热,但有很多空间可以放置集群。

The problem with that is heat dissipation. You get all the radiation from the sun and you don't have any air to dissipate heat, but there is a lot of space to put clusters.

LLM 用于学习与个性化建议 Plateauing Capabilities vs. Economic Impact

Host

那里有大量太阳能,你可以解决散热问题,但能量很多,而且工程意志可能能解决热量问题,所以是有可能的。我们得说这绝对可能。问题是可能性有多大——我们今年基本上会达到平台期?不是指系统能力,而是系统能力对人类文明的实际意义。在编程方面,会建出很棒的网站,非常好的自动补全,非常好的理解代码库和帮助调试的方式。但真的只是编程方面的好帮手。它能帮研究数学家做些数学,帮你购物,帮你……是个好帮手,是加强版 Clippy。还有什么?可能是个好教育工具之类的。但计算机使用变得极其难解决。所以我在试图描绘这些领域里没有巨大经济影响的悲观情况。我们意识到训练这些系统在每个层面都很昂贵,无论是预训练还是推理。推理、推理过程等等都很昂贵。你觉得这可能吗?可能性有多大?

There's a lot of solar energy there and you could figure out the heat dissipation, but there is a lot of energy and there probably could be engineering will to solve the heat problem, so there could be. Is it possible and we should say that it definitely is possible. How likely is it that we're basically going to be plateauing this year? Not in terms of the system capability, but what the system capabilities actually mean for human civilization. So, on the coding front, really nice websites will be built. Very nice autocomplete. Very nice way to understand code bases and maybe help debug. But really just a very nice help on the coding front. It can help research mathematicians do some math. It can help you with shopping. It can help you with... It's a nice helper. It's Clippy on steroids. What else? It may be a good education tool and all that kind of stuff. But computer use turns out extremely difficult to solve. So, I'm trying to frame the cynical case in all these domains where there's not a really huge economic impact. We realize how costly it is to train these systems at every level, both the pre-training and the inference. How costly the inference is, the reasoning, all of that. Is that possible and how likely is that do you think?

Nathan Lambert

当你观察这些模型时,有太多明显需要改进的地方,训练这些模型和做这种艺术需要很长时间,而且以我们现有的想法,需要好几年才能真正饱和我们追求的任何基准或性能。它可能服务于非常狭窄的领域。比如普通的 ChatGPT 和它的百万用户可能不会从中得到很多好处,但它会通过在不同方面变得更好来服务于不同的人群。

When you look at the models, there's so much obvious things to improve and it takes a long time to train these models and to do this art, and that it'll take us with the ideas that we have multiple years to actually saturate in terms of whatever benchmark or performance we are searching for. It might serve very narrow niches. Like the average ChatGPT and their million users might not get a lot of benefit out of this, but it is going to serve different populations by getting better at different things.

Host

但我认为现在每个人都在追求的是一个对所有人都通用的系统。所以如果那个会达到平台期,对吧?

But I think what everybody's chasing now is a general system that's useful to everybody. So if that can plateau, right?

Nathan Lambert

我认为那个梦想实际上正在消亡。就像你谈到的专用模型,多模态往往像视频生成完全是另一回事。

I think that dream is actually kind of dying. As you talked about with the specialized models, where it's like multi-modal is often like video generation is a totally different thing.

Host

那个梦想正在消亡是个大说法。因为我不知道它是否在消亡。我不知道如果你问实际的金融实验室的人,他们还在追求它,对吧?

That dream is kind of dying is a big statement. Because I don't know if it's dying. I don't know if you ask the actual financial lab people, they're still chasing it, right?

Nathan Lambert

我确实认为他们仍在急于推出下一个模型,它会比上一个好得多。我看不到他们放慢脚步。我只是觉得收益将更多地通过不仅 Scaling 模型,而且通过微调来实现。所以我觉得有很多技术债务。就像,好吧,我们把更好的模型放进去,更好的模型,更好的模型。现在人们说,好吧,同时也要改进周围的一切。比如上下文工程和推理 Scaling。大实验室会继续这样做。现在小实验室也会赶上,因为他们正在招聘更多人。会有更多人。LLM 有点像循环。它们也让人们更高效,这只是放大。我认为我们可以期待的是放大,而不是范式转变。我不认为那是真的,但一切都会被放大、放大、再放大。而且我看这能持续很长时间。

I do think they are still rushing to get the next model out which will be much better than the previous one. And I can't see them slowing down. I just think the gains will be made or felt more through not only scaling the model, but now fine-tuning. So I feel like there's a lot of tech debt. It's like, well, let's just put the better model in there and better model better model. And now people are, okay, let's also at the same time improve everything around it too. Like the engineering of the context and inference scaling. And the big labs will still keep doing that. And now also the smaller labs will catch up to that because now they are hiring more. There will be more people. LLMs is kind of like a circle. They also make them more productive and it's just amplifying. I think what we can expect is amplification, but not a paradigm change. I don't think that is true, but everything will be just amplified and amplified and amplified. And I can see that continuing for a long time.

Host

是的,我想我说梦想正在消亡取决于你具体认为它会做什么。比如 Claude Code 是一个能做很多事的通用模型,但它不一定,它很大程度上依赖于集成和其他东西。我打赌 Claude 能很好地处理你的邮件,最难的部分是如何把信息给它,以及如何让它能发送你的邮件之类的。但这又回到了“一个模型统治一切”的理念,就是一个在云端处理你整个数字生活、比所有人都聪明的东西。从 Claude Code 变成那样是一个有趣的信仰飞跃。在某些方面有一些途径,但我确实认为行业的说法有点不同。

Yeah, I guess my statement with the dream is dying depends on exactly what you think it's going to be doing. Like Claude code is a general model that can do a lot of things, but it's not like necessarily like it depends a lot on integrations and other things. Like I bet Claude could do a fairly good job of doing your email and the hardest part is figuring out how to give the information to it and how to get it to be able to send your emails and stuff like this. But that's just kind of like I think it goes back to like what is the one model to rule everything ethos, which is just like a thing in the cloud that handles your entire digital life and is way smarter than everybody. It's an interesting leap of faith to go from Claude code becomes that. Which in some ways there's some avenues for that, but I do think that the rhetoric of the industry is a little bit different.

Nathan Lambert

我认为作为普通用户使用 LLM,我们接下来会立即感受到的可能是像制作图表这样琐碎的事情。现在 LLM 在制作图表方面很糟糕。是因为我们在后台被提供了推理算力较少的廉价模型吗?也许有些技巧已经可以得到更好的图表,但如果你今天问,我不知道。画一个 XYZ 的流程图,大多数时候很糟糕,而这对人类来说是非常简单的任务。我认为有时候画东西比写东西更容易。

I think the immediate thing we will feel next as a normal person using LLMs will probably be related to something trivial like making figures. Right now LLMs are terrible at making figures. Is it because we are getting served the cheap models with less inference compute behind the scenes? Maybe there are some cranks we can already get better figures, but if you ask today, I don't know. Draw a flow chart of XYZ. It's most of the time terrible and it is kind of like a very simple task for a human. I think it's almost easier sometimes to draw something than to write something.

Host

是的,多模态理解确实感觉很奇怪,它没有得到更好的解决。

Yeah, the multimodal understanding does feel like something that is odd that it's not better solved.

Nathan Lambert

我认为我们没说一个实际上很明显的事情,我们没意识到那是一个巨大且难以衡量的东西,那就是让全人类的知识对全世界可及。谷歌搜索和 LLM 之间有很大区别。我觉得我基本上可以问 LLM 任何问题并得到答案。而且它的幻觉越来越少。这意味着理解我自己的生活,规划职业轨迹,解决我周围的问题,了解人类历史上的任何事物。我觉得没人真正谈论这个,因为他们立即认为这很了不起是理所当然的。这就是为什么每个人都在用它,因为你得到了答案。而随着时间的推移,这种影响不仅在美国,而是在全世界。比如全世界的孩子都能学习这些想法,这种影响随着时间的推移可能就是真正谈论 GDP 的地方。它不会是一个飞跃。它将是我们如何到达火星,如何建造这些东西,如何拥有一百万个新的 OpenAI,所有创新都由此而来。而这只是渗透一切的安静力量,对吧?人类知识。

I think we're not saying one actually obvious thing that we're not actually realizing that's a gigantic thing that's hard to measure, which is making all of human knowledge accessible to the entire world. Like there's a huge difference between Google search and an LLM. I feel like I can basically ask an LLM anything and get an answer. And it's doing less and less hallucination. And that means understanding my own life, figuring out a career trajectory, figuring out how to solve the problems all around me, learn about anything through human history. I feel like nobody's really talking about that because they just immediately take it for granted that it's awesome. That's why everybody's using it, because you get answers for stuff. And the impact of that across time, think about this is not just in the United States, it's all across the world. Like kids throughout the world being able to learn these ideas, the impact that has across time is probably where the real talking about GDP. It won't be like a leap. It'll be that's how we get to Mars. That's how we build these things. That's how we have a million new OpenAIs, all the kind of innovation that happens from there. And that's just this quiet force that permeates everything, right? Human knowledge.

Host

我确实同意你的看法,在某种意义上它让知识更易获取,但我也认为这取决于主题是什么。

I do agree with you and in a sense it makes knowledge more accessible, but it also I think depends on what the topic is.

AI 初创公司的整合 LLMs for learning vs. personalized advice

Nathan Lambert

比如数学,你可以问它问题,它会回答,但如果你想从头学一个主题,我觉得最佳方式是使用一本好的数学教材,线性地讲解内容,这是经过验证的策略。然后你用大语言模型生成无限量的练习题。你让它出例题,自己解答,如果需要更多背景知识,再让它生成。但它不会给出教材之外的东西,只是换种方式包装。然而,有些时效性的事情没有好的替代方案,只能靠人类临时处理。比如计划去迪士尼乐园:买哪个公园的票、什么时候去。没有教材,只有零散的互联网信息。大语言模型通过临时定制来增加价值,从零散的互联网中提取信息,而那里本没有更好的版本存在。

For something like math, you can ask it questions, it answers, but if you want to learn a topic from scratch, I think the sweet spot is using a good math textbook that lays things out linearly. That's a proven strategy. Then you use the LLM to generate infinite exercises. You ask it for example problems, solve them, and if you need more background, you ask it to generate that. But it won't give you anything not in the textbook, just packaged differently. However, there are timely things where no good alternative exists except a human doing it on the fly. For example, planning a Disneyland trip: which tickets to buy for which park when. There's no textbook, only sparse internet. The LLM adds value by customizing on the fly, pulling information from the sparse internet where no better version exists.

Host

而且如果存在,也充满了广告垃圾。对于任何城市,问大语言模型十大必做之事比互联网上的任何东西都好得多。

And if it does exist, it's full of ad slop. For any city, asking an LLM for top 10 things to do is way better than anything on the internet.

Nathan Lambert

现在是这样。这是因为它们被大量补贴,并将由广告支付。

Now. That's because they're massively subsidized and will be paid for by ads.

Host

我很期待。

I look forward to that.

Nathan Lambert

它就要来了。

It's coming.

Host

不,不。我希望有非常明确的标识,区分什么是广告,什么不是。

No, no. I hope there's a very clear indication of what's an ad and what's not.

Nathan Lambert

我几年前提到过:如果你在找一双新跑鞋,耐克排第一是巧合吗?也许是,也许不是。但对此有明确的法律规定,你必须清楚。这正是大家所担心的:微妙的暗示。但这引出了广告的话题,OpenAI 在 2025 年试图推出广告。他们仍然没有通过那种方式赚钱。在里面放广告位,但存在无广告的替代品,所以人们会涌向其他产品。他们互相攀比,花那么多钱来获取用户,这太疯狂了。

I mentioned this a few years ago: if you're looking for a new running shoe, is it a coincidence that Nike comes up first? Maybe, maybe not. But there are clear laws about this; you have to be clear. That's what everyone fears: subtle messages. But it brings us to ads, which OpenAI tried to launch in 2025. They're still not making money that way. Having ad spots in there, but alternatives without ads exist, so people would flock to other products. It's crazy how they one-up each other, spending so much money to get users.

Host

我同意。有些 Instagram 广告对于找到喜欢你产品的用户是有益的。但也有激励扭曲的糟糕情况。一个 AI 与这种积极观点结合的世界——比如小企业制造最好的牛排刀并卖给需要的人——对世界非常有益,尤其是在数字基础设施方面。但让人上瘾的推送以展示更多内容则不好。OpenAI 会说他们想要广告的货币化好处,同时赋予用户自主权。

I think so. Some Instagram ads are good for finding users who like your product. But there are also awful cases for incentives. A world where AI integrates with that positive view—like a small business making the best steak knives and selling to those who need them—is very good for the world, especially with digital infrastructure. But addicting feeds to show more content is not good. OpenAI would say they want monetization upside of ads while giving users agency.

Nathan Lambert

我个人认为谷歌会更擅长解决这个问题,因为他们已经有广告供应,并且可以将 Gemini 中的需求转化为有用的广告。今年会有实验。阻碍公司的是竞争对手没有这样做。这是声誉问题;人们害怕毁掉声誉或失去用户,因为这会成为头条新闻。

I personally think Google will be better at figuring this out because they already have ad supply and can turn demand in Gemini into useful ads. There will be experiments this year. What holds companies back is that the competition isn't doing it. It's a reputation thing; people are afraid of ruining their reputation or losing users because it would make headlines.

Host

除非它们很棒。但第一批广告不会很棒,因为这是一个我们不知道如何解决的难题。

Unless they were great. But the first ads won't be great because it's a hard problem we don't know how to solve.

Nathan Lambert

第一个版本很可能像 X 上的推广帖子,带有一个小的“推广”标签。问题是谁先迈出第一步。

The first version will likely be like promoted posts on X, with a small 'promoted' label. The problem is who makes the first move.

Host

如果放眼十年后,广告的设想是:你通过大量用户从广告中赚取巨额资金,从而资助更好的研发,制造更好的模型。这就是 YouTube 占据主导地位的原因;Netflix 害怕 YouTube。他们每月从我这里赚取至少 28 美元,还有很多人,从而在视频领域建立了主导地位。广告可以在每位用户的花费上带来持续优势。但启动那个飞轮是可怕的,因为这是一个长期赌注。

If we go 10 years out, the proposition for ads is that you make so much money from ads with many users that you can fund better R&D and make better models. That's why YouTube dominates; Netflix is scared of YouTube. They make at least $28 a month off me and many others, creating a dominant position in video. Ads can give a sustained advantage in spending per user. But starting that flywheel is scary because it's a long-term bet.

Nathan Lambert

你认为今年在商业上会有疯狂的大动作吗?比如谷歌或苹果收购 Anthropic?

Do you think there will be crazy big moves this year business-wise? Like Google or Apple acquiring Anthropic?

IPO 与 AI 公司未来 Consolidation in AI startups

Nathan Lambert

Dario 永远不会卖公司,但我们开始看到一些整合,比如 Grok 以 200 亿美元、Scale AI 以近 300 亿美元成交,还有无数类似交易。这些交易的结构实际上对硅谷生态系统不利——它们是一种许可协议,并非所有人都能参与,而不是让普通员工通过股票归属受益的全额收购。这对硅谷文化来说是一个需要解决的大问题,因为创业生态系统是命脉:如果你加入一家初创公司,即使它不太成功,也很有可能以较低的溢价被收购,你的股权就能变现。而这些许可协议很多时候只是在挖走顶尖人才。我认为 Grok 卖给 Nvidia 的传闻对员工更有利,但这仍然是一种规避反垄断的做法。但我认为这种整合趋势会继续。我和我尊重的许多聪明人一直期待整合更早发生,但现在似乎有些苗头了。但与此同时,有些公司为了你无法理解的原因筹集了巨额资金——我不明白他们为什么要拿那些钱。所以今年可能好坏参半,但整合压力正在显现。

Dario will never sell, but we are starting to see some types of consolidation with like Grok for $20 billion and Scale AI for almost $30 billion and countless other deals like this that they're structured in a way that is actually detrimental to the Silicon Valley ecosystem, which is this sort of licensing deal where not everybody gets brought along rather than a full acquisition that benefits the rank and file of employees by getting their stock vested. Like that's a big issue for Silicon Valley culture to address because the startup ecosystem is the lifeblood where if you get a if you join a startup even if it's not that successful, your startup very well might get acquired on a cheap premium of it and you'll get paid out for this equity and these licensing deals are essentially taking the top talent a lot of the times. I think Grok they deal for Grok to Nvidia is rumored to be better to the employees, but it is still this antitrust avoiding thing. But I think that this trend of consolidation will continue. I've been me and many smart people I respect have been expecting call consolidation to have happened sooner, but it seems like some of these things are starting to turn, which but at the same time you have companies raising ridiculous amounts of money for reasons that you don't like I'm like I don't know why you're taking that money. So it's maybe like mixed this year, but some consolidation pressure is starting.

Host

你觉得我们会看到哪些令人惊讶的整合?所以你的意思是 Anthropic 永远不会卖。我是说 Grok 是一个大交易。顺便说一句,是 Grok 带 Q。

What kind of surprising consolidation do you think we'll see? So you're saying you're saying Anthropic is a never. I mean Grok is a big one. Grok with a Q by the way.

Nathan Lambert

是的。初创公司很多,AI 初创公司的溢价很高。所以可能会有很多百亿美元级别的收购,这对于一家可能一年前才成立的初创公司来说是非常大的收购。我认为 Meta 收购的这家新加坡公司 Mainas AI,成立八个月后就以 20 亿美元退出。而且我认为还会有其他数十亿美元的收购,比如 Perplexity。

Yeah. There's just a lot of startups and there's a very high premium on AI startups. So there's a lot of like there could be a lot of 10 billion range acquisitions, which is a really big acquisition for a startup that was maybe founded a year like a year ago. I think Mainas AI from this company that's based in Singapore that Meta founded was founded eight months ago and then had a $2 billion exit. And I think that there will be some other big like many billion-dollar acquisitions like Perplexity.

Host

是的,比如有传言说它们会被苹果收购。

Yeah, like people rumored them to Apple.

Nathan Lambert

我认为 AI 领域有很多压力和流动性。大公司有压力要取得成果,我猜一笔大的收购能让人们有空间去讲述下一章的故事。

I think there's a lot of pressure and liquidity in AI. There's pressure on big companies to have outcomes and I I would guess that a big acquisition gives people leeway to then tell the next chapter of that story.

Host

我的意思是,是的,我想 Cursor 也是。我们一直在讨论 Coden。有人收购了 Cursor。

I mean, yeah, there's a I guess Cursor. We've been talking about Coden. Somebody acquires Cursor.

Nathan Lambert

他们拥有大量用户数据,处于非常有利的位置。

They're in such a good position by having so much user data.

Host

是的。

Yeah.

Nathan Lambert

我们讨论过持续学习之类的话题。他们在一篇博客文章中有两句非常有趣的话:他们新的 Composer 模型是对一个中国的大型混合专家模型进行微调得到的。你可以通过问八卦或者因为模型有时会用中文回答(美国模型都不会这样)来知道这一点。他们在博客中说:‘我们每 90 分钟根据用户的实际反馈更新一次模型权重。’这几乎是在模型上实现真实世界强化学习的最接近的方式。而且这只是一篇博客文章里的内容,非常酷。

And we talked about continual learning and stuff. They had one of the most interesting like two sentences in a blog post, which is that they had their new composer model, which was a fine-tune of one of these large mixture of expert models from China. You can know that by asking Gossip or because the model sometimes responds in Chinese, which none of the American models do. And they had a blog post where they're like, 'We're updating the model weights every 90 minutes based on real-world feedback from people using it.' Which is like the closest thing to real-world RL happening on a model. And it's just like in one of their blog posts, which is super cool.

Host

顺便说一句,我经常用 Composer,因为它有一个好处就是快。

And I And by the way, I just I say I use Composer a lot cuz it's one one of the benefits that it has is it's fast.

Nathan Lambert

我得试试,因为大家都这么说。

I need to try it cuz everybody says this.

API 市场竞争 IPOs and future of AI companies

Nathan Lambert

而且可能会有一些 IPO。你觉得 Anthropic、OpenAI、xAI 会上市吗?

And there'll be some IPOs potentially. You think Anthropic, OpenAI, xAI?

Host

它们都能轻松筹集大量资金,所以不觉得有必要上市。只要融资容易,它们就不会 IPO,因为公开市场会带来压力。我们看到中国的生态系统有些不同,MiniMax 和 Z.AI 都提交了 IPO 文件,看看中国市场的反应会很有趣。我猜可能会和美国一样炒作,只要这一切还在继续,而不是基于它们都在亏损的现实。我希望更多美国大型 AI 初创公司上市,因为看看它们如何花钱会很有趣,也能获得更多洞察。同时也能让人们有机会投资这些公司,因为我认为它们是这个时代最具塑造力的公司。现在美国的传统是很多大型初创公司不上市。我们还在等 Stripe 的 IPO,但 Databricks 肯定没上,它们融了 G 轮之类的。我觉得市场处于一种奇怪的平衡状态,我希望看到这些公司上市,并以公司应有的方式发展。

They can all raise so much money so easily that they don't feel a need to Like, so long as fundraising is easy, they're not going to IPO because public markets apply pressure. I think we're seeing in China that the ecosystem's a little different with both MiniMax and Z dot AI applying for um filing IPO paperwork, which will be interesting to see how the Chinese market reacts. I actually would guess that it's going to be like similarly hypey to the US so so long as all this is going and not based in the realities that they're both losing a ton of money. I wish more of the American gigantic AI startups were public because it would be very interesting to see how they're spending their money and have more insight. And also just to give people access to investing in these cuz I think that they're some of the most like formative they're the companies of the era. And the tradition is now for so many of the big startups in the US to not go public. It's like we're still waiting for Stripe and the IPO, but Databricks definitely didn't. They raised like a series G or something. And I just feel like it's a kind of a weird equilibrium for the market where it's like I would like to see these companies go public and evolve in that way that a company can.

Nathan Lambert

你认为 10 年后一些前沿模型公司还会存在吗?比如 Anthropic、OpenAI?

You think 10 years from now some of the frontier model companies are still around? Anthropic, OpenAI?

Host

我绝对不认为这是赢家通吃,除非其中一家真的找到了某种算法秘密,就像飞轮一样,因为所有公司的开发路径都非常相似。谷歌和 OpenAI 有几乎相同的产品,而 Anthropic 更专注,但当你和人们交谈时,听起来他们在解决很多相同的问题。所以我认为产品会分散开来。这是一个正在做大的蛋糕,人们会从中分一杯羹。

I definitely don't see it to be a winner takes all unless there truly is some algorithmic secret that one of them finds like this is flywheel cuz the development path is so similar for all of them. Google and OpenAI have like all the same products and then like Anthropic's more focused, but when you talk to people it sounds like they're solving a lot of the same problems. So I think and there's offerings that will spread out. There's a lot of it's a very big cake that's being made that people are going to take money out of.

Nathan Lambert

我不想轻视这一点,但 OpenAI 和 Anthropic 主要是大语言模型服务提供商,而其他公司如谷歌和与 X 关联的 xAI 也做其他事情。所以如果 AI 变得更加商品化,那些只提供大语言模型的公司很可能会消亡。

I don't want to trivialize it, but so OpenAI and Anthropic are primarily LLM service providers and some of the other companies like Google and xAI linked to X does other stuff, too. And so it's very possible if AI becomes more commoditized that the companies that are just providing LLM will die.

Host

我认为它们有优势,它们拥有大量用户,而且我认为它们会转型。比如 Anthropic 就转型了。我不认为它们最初计划做代码,但后来发现这是一个不错的细分市场,现在它们在这个细分市场很舒适并持续发力。我可以看到同样的情况,假设一下,我不确定是否会发生,但假设谷歌占据了通用聊天机器人的全部市场份额,那么 OpenAI 可能会专注于其他子领域,因为它们有太多用户,在可预见的未来不会消失。

I think they will the advantage they have they have a lot of users and I think they will just pivot. I think um then if they figure out it's like Anthropic I think pivoted. I don't think they originally planned to work on code, but it happened that they found okay this is like a nice niche and now we are comfortable in this niche and we push on this niche and I can see the same thing once maybe let's say hypothetically speaking I I'm not sure if it will be true, but let's say Google takes all the market share of the general chatbot. Maybe OpenAI will be then focus on some other sub topic like the have too many users to go away in foreseeable future, I think.

Nathan Lambert

我觉得谷歌总是准备好说‘看我的’,在 AI 模式上。

I think Google is always ready to say hold my beer with the AI mode.

Host

我认为问题在于这些公司能否支撑其估值。我觉得 AI 公司在某些方面可以被看作像 AWS、Azure 和 GCP 一样,都在同一领域竞争,并且都是非常成功的业务。API 市场可能利润太低,以至于它们会向上下栈扩展到产品和硬件。它们有足够的现金来建造发电厂和数据中心,这现在是一个持久的优势。

I think that the question is if the companies can support the valuations. I think I'd see the AI companies being looked at in some ways like AWS Azure and GCP are all competing in the same space and all very successful businesses. There's a chance that the API market is so unprofitable that they go up and down the stack to products and hardware. They have so much cash that they can build power plants and build data centers, which is a durable advantage now.

Meta 路径与 Llama 未来 API Market Competition

Nathan Lambert

但还有一种合理的结果是,这些 API 对开发者来说非常有价值且灵活,以至于它们变得像 AWS 一样。但 AWS 和 Azure 也会提供这些 API,所以 API 市场上有五六家公司在竞争,这很困难。所以也许这就是它们被挤出市场的原因。

But there's also just a reasonable outcome that these APIs are so valuable and so flexible for developers that they become like something like AWS. But AWS and Azure are also going to have these APIs, so there's like five or six people competing in the API market, which is hard. So maybe that's why they get squeezed out.

Llama 演进与内部问题 Meta's Path and Llama's Future

Host

你提到了 RIP Llama。Meta 有获胜的路径吗?

You mentioned RIP Llama. Is there a path to winning for Meta?

Nathan Lambert

我认为没人知道。他们动作很多。他们正在与 Black Forest Labs(一家图像生成公司)或 Midjourney 签署许可协议,或者收购它们。所以我认为在产品和面向消费者的 AI 方面,现在下结论还为时过早。我认为他们有一些非常优秀且充满干劲的人,与扎克伯格关系密切。所以我认为那里还有故事要展开。Llama 有点不同。Llama 是该组织最专注的表达,我不认为 Llama 会得到那种程度的支持。我认为这对他们来说是一个非常成功的品牌,所以他们可能仍会参与开放生态系统,或者将 Llama 品牌延续到不同的服务中,因为人们知道 Llama 是什么。

I think nobody knows. They're moving a lot. They're signing licensing deals with Black Forest Labs, which is an image generation company, or Midjourney, or acquiring them. So I think in some ways, on the product and consumer-facing AI front, it's too early to tell. I think they have some people that are excellent and very motivated, being close to Zuckerberg. So I think there's still a story to unfold there. Llama is a bit different. Llama was the most focused expression of the organization, and I don't see Llama being supported to that extent. I think it was a very successful brand for them, so they might still participate in the open ecosystem or continue the Llama brand into a different service because people know what Llama is.

Host

你认为会有 Llama 5 吗?

Do you think there's a Llama 5?

Nathan Lambert

不会是开放权重的。

Not an open-weight one.

扎克伯格角色与开源辩论 Llama's Evolution and Internal Issues

Host

有意思。我想回顾一下,Llama 是开创性的开放权重模型,Llama 1、2、3 备受喜爱。但后来,我推测,Meta 的高层领导对 Llama 非常兴奋,因为他们看到它在社区中很受欢迎。然后问题在于试图将开源货币化,或者利用开源制造更大的轰动,几乎是强行推动。感觉像是强行开发这些非常大的 Llama 4 模型以登上基准测试榜首。但我认为 Llama 模型的目标不是登上基准测试榜首,击败 ChatGPT 或其他模型。我认为目标是拥有一个人们可以使用、信任、修改和理解的模型,所以这包括更小的模型。它们不必是最好的模型。发生的事情是,这些模型在基准测试上表现得比实际更好,因为他们有针对偏好训练的特定模型,在基准测试上表现良好。这有点像为了强行成为最好而过度拟合。但与此同时,他们没有做人们可以使用的小模型。那时没人能运行这些大模型。这很奇怪,我认为这只是因为人们太热衷于推动前沿的头条新闻。

It's interesting. I think just to recap, Llama was the pioneering open-weight model, and Llama 1, 2, 3 got a lot of love. But then, hypothesizing, I think the leaders at Meta, the upper executives, got really excited about Llama because they saw how popular it was in the community. Then the problem was trying to monetize the open source, or use the open source to make a bigger splash, to force it almost. It felt forced, like developing these very big Llama 4 models to be on top of the benchmarks. But I don't think the goal of Llama models is to be on top of the benchmarks beating ChatGPT or other models. I think the goal was to have a model that people can use, trust, modify, understand, so that includes having smaller models. They don't have to be the best models. What happened was these models were better on benchmarks than they actually were because they had specific models trained on preferences that performed well on benchmarks. It's kind of overfitting to force it to be the best. But at the same time, they didn't do the small models that people could use. No one could run these big models then. It was a weird thing, and I think it's just because people got too excited about headlines pushing the frontier.

Nathan Lambert

而且过于关注基准测试方面。

And too much focus on the benchmarking side.

Host

是的,我认为它在内部政治斗争和激励错位下崩溃了。研究人员想构建最好的模型,但有一层组织和管理者试图证明他们做了这些事情。有很多关于糟糕技术决策的传言,看起来情况变得太糟,最终崩溃了。

Yeah, I think it imploded under internal political fighting and misaligned incentives. The researchers want to build the best models, but there's a layer of organization and managers trying to demonstrate that they do these things. There are lots of rumors about how some horrible technical decisions were made, and it just seems like it got too bad and crashed out.

社区反弹及其影响 Mark Zuckerberg's Role and Open Source Debate

Host

我们也应该大力赞扬马克·扎克伯格。我认为这来自马克,来自领导层高层,他说开源很重要。这个事实的存在意味着可能会有 Llama 5,他们从基准测试最大化中吸取教训,说我们要像 GPT-4 一样,提供真正出色的开源库。

We should also give huge props to Mark Zuckerberg. I think it comes from Mark, from the top of the leadership, saying open source is important. The fact that that exists means there could be a Llama 5 where they learn the lessons from the benchmark maxing and say we're going to be like GPT-4 and provide a really awesome library of open source.

Nathan Lambert

人们说马克和亚历山大·王之间存在争论,亚历山大·王非常聪明,但更反对开源。鉴于他对 AI 组织有很大影响力,开源的可能性似乎更小。看起来马克把他请来是为了在指导 AI 方面提供新的领导力。如果开放或封闭不再是模型的定义性特征,我不认为这会是马克和亚历克斯之间的定义性争论。他们都非常聪明,但我很难理解这一切,因为马克在七月写了一篇文章,那可能是当时最好的博客文章,为开源 AI 辩护。然后到了 2025 年 7 月,却变成了我们正在重新评估与开源的关系。所以这有点……

What people say is that there's a debate between Mark and Alexander Wang, who is very bright but much more against open source. To the extent that he has a lot of influence over the AI org, it seems much less likely. It seems like Mark brought him in for fresh leadership in directing AI. If open or closed is no longer the defining nature of the model, I don't expect that to be a defining argument between Mark and Alex. They're both very bright, but I have a hard time understanding all of it because Mark wrote this piece in July, which was probably the best blog post at the time, making the case for open source AI. Then July 2025 came around and it was like we're reevaluating our relationship with open source. So it's just kind of...

美国填补 Llama 空白 Community Backlash and Its Impact

Host

我认为我们可能有点过于严厉了,这导致了部分问题。作为开源开发者或开源社区,即使模型可能不是每个人都期望的,它也遭到了很多反对。我认为这很不幸,因为作为一家公司,他们希望得到正面的头条新闻,结果却得到了负面的。这对公司产生了负面影响。这可能引发了一种报复性反应:'我们试图做点好事,给了你们一个开源模型,现在你们却对我们持负面态度。'所以也许他们会改变主意。

I think we may have been a bit too harsh, and that caused some of that. As open source developers or the open source community, even though the model was maybe not what everyone hoped for, it got a lot of backlash. I think that was unfortunate because as a company, they were hoping for positive headlines, and instead they got negative headlines. That reflected badly on the company. It might have caused a spite reaction: 'We tried to do something nice, gave you an open-source model, and now you're being negative about us.' So maybe they'll change their mind.

Nathan Lambert

是的,这就是 X 上的讨论动态可能让我们误入歧途的地方。有时随机的人会挑选他们喜欢和不喜欢的东西。你在 Grok 4.1 和 Grok Code Fast 1 上也能看到同样的情况。我认为人们并不公开喜欢它,但很多人使用它。在 Reddit 和 X 上,编程社区并不称赞它,但他们使用它。Llama 也是如此。我不理解正面或负面炒作背后的动态。

Yeah, that's where the dynamics of discourse on X can lead us astray. Sometimes random people pick the thing they like and don't like. You can see the same thing with Grok 4.1 and Grok Code Fast 1. I don't think people publicly love it, but a lot of people use it. On Reddit and X, they don't give it praise from the programming community, but they use it. The same thing with Llama. I don't understand the dynamics of either positive or negative hype.

Adam 项目:起源与使命 US Filling the Gap Left by Llama

Nathan Lambert

2025 年的一个故事是美国填补 Llama 留下的空白,即这些中国开放权重模型的崛起,以至于我想说,'这是过去五个月我投入大量精力的唯一问题,试图通过政策工作让美国在这方面投资。'

One of the stories in 2025 is the US filling the gap of Llama, which is the rise of these Chinese open-weight models to the point where I'm like, 'That was the single issue I've spent a lot of energy on the last 5 months, trying to do policy work to get the US to invest in this.'

Host

那么,给我讲讲 Adam 的故事吧。

So, tell me the story of Adam.

政府与行业支持 The Adam Project: Origin and Mission

Nathan Lambert

Adam Project 最初被我称为美国 DeepSeek 项目,这个说法在 DC 受众中不太奏效,但它讲述的是我职业生涯中能做的最有影响力的事情。中国的开放权重模型正在积累大量影响力,而美国企业对基于这些开放模型进行构建有巨大需求,尤其是那些对中国模型非常谨慎的企业。

The Adam Project started as me calling it the American DeepSeek Project, which doesn't really work for DC audiences, but it's the story of what is the most impactful thing I could do with my career. Chinese open-weight models are cultivating a lot of power, and there is a lot of demand for building on these open models, especially in enterprises in the US that are very cagey about these Chinese models.

Host

根据 Perplexity 的说法,Atom 项目——美国真正开放模型——是一项美国本土倡议,旨在构建和托管高质量、真正开放权重的 AI 模型及支持基础设施,明确目标是与中国快速发展的开源 AI 生态系统竞争并迎头赶上。

Going to Perplexity, the Atom project, American truly open models is a US-based initiative to build and host high-quality, genuinely open-weight AI models and supporting infrastructure explicitly aimed at competing with and catching up to China's rapidly advancing open-source AI ecosystem.

Nathan Lambert

我认为一句话总结就是,或者两句话。第一,开放模型将成为 AI 研究的引擎,因为人们都是从开放模型开始研究的。因此,拥有它们很重要。第二,因此美国应该构建最好的模型,这样最好的研究就会发生在美国,美国公司也能从作为 AI 研究发源地中获取价值。如果没有对开放模型的更多投资,我们在网站上看到的所有图表都会是‘Quinn, Quinn, Quinn’,全是这些中国公司的优秀模型,它们正在美国、中国和国际上培养影响力。我认为美国在 AI 上投入了更多资金,但创建比封闭实验室前沿领先半代或一代的开放模型的能力,成本高达数亿美元,这虽然是一大笔钱,但对这些公司来说并不算多。因此,我们需要一个中心化的力量,由想做这件事的人组成。我认为我们得到了几乎全栈人员的签约参与,无论是政策方面。

I think the one-sentence summary would be that, or two sentences. One is a proposition that open models are going to be an engine for AI research because that is what people start with. Therefore, it's important to own them. And the second one is therefore, the US should be building the best models so that the best research happens in the US and US companies take the value from being the home of where AI research is happening. And without more investment in open models, we have all the plots on the website where it's like, 'Quinn, Quinn, Quinn.' And it's all these models that are excellent from these Chinese companies that are cultivating influence in the US and China and internationally. I think the US is spending way more on AI, and the ability to create open models that are half a generation or a generation beyond what the cutting edge of the closed labs is costs orders of like hundred million dollars, which is a lot of money, but not a lot of the money to these companies. So, therefore, we need a centralizing force of people who want to do this. And I think we got signed engagement from people pretty much across the full stack, whether it's policy.

AI 行动计划与开源 Government and Industry Support

Host

那么,政府方面有支持吗?

So, there has been support from the administration?

Nathan Lambert

我不认为政府中有任何人公开签署了支持,但我知道在拜登和特朗普政府中从事 AI 政策工作的人都非常支持推动美国开源模型。例如,AI2 从 NSF 获得了四年 1 亿美元的资助,这是 NSF 有史以来最大的计算机科学资助,用于 AI2 尝试这件事,我认为这是一个起点。但最好的情况是多个组织构建模型,因为它们可以交叉授粉思想,构建这个生态系统。我不认为仅靠 Llama 向世界发布模型就能成功,因为 Llama 可能会消失。AI2 也是如此,我不能是唯一构建模型的人。我认为这需要花很多时间与政策领域的人交谈。我知道 Nvidia 对此非常兴奋。黄仁勋特别谈到了这件事的紧迫性,他们在 2025 年做了更多工作,Nemotron 模型成为重点。他们开始随 Nvidia 开放模型发布一些数据,很少有公司这样做,尤其是像 Nvidia 这样规模的公司。所以有进展的迹象,我们听说 Reflection AI 声称其 20 亿美元的融资专门用于构建美国开放模型,我觉得他们的公告推文读起来像一篇博文,我认为文化潮流开始转变。我认为在 7 月,我们有四五个 DeepSeek 级别的中国开放权重模型,而美国为零。那一刻我发布了这个,我想我必须为此投入精力,因为没别人会做。所以这需要很多人共同努力,我并不是说 Adam 项目是推动生态系统的唯一因素,但像我这样的人在做这件事来传播信息。

I don't think anyone in that like technically in government has like signed it publicly, but I know that people that worked in AI policy both in Biden and Trump administration are very supportive of trying to promote open-source models in the US. I think, for example, AI2 got a grant from the NSF for a hundred million dollars over four years, which is like the biggest CS grant the NSF has ever awarded and it's for the AI2 to attempt to this and I think it's a starting point. But the best thing happens when there are multiple organizations building models because they can cross-pollinate ideas and kind of build this ecosystem. Like I don't think if it just works if it's just Llama releasing models to the world because then you can see Llama can go away. The same thing applies for AI2 where it's like I can't be the only one building models. And I think that it's like that it becomes a lot of time spent on talking to people whether they're in policy. I know Nvidia is very excited about this. I think Jensen Huang has been specifically talking about the urgency for this and they've changed they've done a lot more in 2025 where the Nemotron models are more of a focus. They've started releasing some data along with Nvidia's open models and like very few companies do this especially of Nvidia's size. So like there is there is signs of progress and there we hear about Reflection AI where they say their 2 billion dollar fundraise is dedicated to building US open models and I feel like their announcement tweet is like it reads like a blog post all right and I think that that cultural tide is starting to turn. I think in in July was when we had like four or five deep seat caliber Chinese open weight models and zero from the US. And that's that's the moment where was released this and I was like I guess I have to spend energy on this cuz nobody else is going to do it. So it takes a lot of it takes a lot of people contributing together and I don't say that like the Adam project isn't like the thing that's helping to move the ecosystem, but it's people like me doing this sort of thing to get the word out.

教育与人才培养 The AI Action Plan and Open Source

Host

你喜欢 2025 年美国 AI 行动计划中包含开源内容吗?白宫 AI 行动计划有一个专门章节,标题是鼓励开源和开放权重 AI,定义了这类模型,并论证它们对创新和初创企业有独特价值。

Uh do you like the the 2025 America's AI action plan that includes open source stuff? The White House AI action plan includes a dedicated section titled to encourage open source and open weight AI defining such models and arguing they have unique value for innovation and startups.

Nathan Lambert

是的,AI 行动计划是一个计划,但很大程度上我认为它可能是政府发布的最连贯的政策文件,我希望它能大体成功。我认识参与制定 AI 行动计划的人,挑战在于将政策变为现实,作为 AI 研究员我不知道如何做到,但其中很多事情非常真实。美国正在大规模建设 AI,人们听到了从用水到各种问题,我们应该能够在这个国家建设东西,但同时也需要在建设过程中不破坏我们的地方。这值得投入精力,我认为联邦政府的作用是设定议程,在 AI 议程中,开放权重应作为首要考虑,这是他们能做的大部分事情,然后人们会去思考。

Yeah, I mean like the AI action plan is a plan, but largely I think it's like maybe the most coherent policy document that has come out of the administration and I hope that it largely succeeds and I know people that have worked on the AI action plan and the challenge is taking policy and making it real and I have no idea how to do this as an AI researcher, but like like largely a lot of things in that were very real and there's a huge build out of AI in the country and it's like there are a lot of issues that people are hearing about from water use to whatever and like we should be able to build things in this country, but also we need to not ruin places in our country in the process of building it and it's a worthwhile to spend energy on and I think that's a role that the federal government plays is like they set the agenda and with AI setting the agenda that open weight should be a first consideration is like that's a large part of what they can do and then people think about it.

开源模型与全球访问 Education and Talent Development

Host

此外,对于这些公司的教育和人才来说,我认为非常重要,因为如果只有封闭模型,你如何让下一代人做出贡献?否则你只能在加入公司后才能学习,但那时你如何招聘有才华的人?如何识别有才华的人?我认为开源,即使对于很多事情,甚至仅仅为了教育大众和培养下一代研究人员,也是唯一的方式。

Also for education and talent for these companies it's I think very important because otherwise you know if they're only closed um models, how do you get the next generation of people contributing at some point because otherwise you will at some point only be able to learn after you joined a company, but then at that point like how do you hire talented people? How do you identify talented people? And I think open source is let's say even for a lot of things, but also even just for educating the population and training the next generation of researchers it's the way or the only way.

Nathan Lambert

我本可以让这件事更广泛传播的方式是讲述中国 AI 与威权国家结合、成为超级智能并接管世界的故事,因此我们需要自己的美国模型。但我故意谈论美国的创新和科学,因为我认为这作为结果更现实,而且这是一个我希望实现的世界。

The way that I could have gotten this to more go more viral is was to tell a story of Chinese AI integrating with an authoritarian state and being ASI and taking over the world and therefore we need our own American models, but it's very intentional for why I talk about innovation and science in the US because I think it's both more realistic as an outcome, but just like it's like it's a world that is that I would like to manifest.

Host

不过我想说,即使是任何开放权重模型,我认为都是有价值的模型。

I would say though also even like let's say any open weight model I do think is a valuable model.

Nathan Lambert

是的,我的论点是我们应该处于领先地位。但我觉得有必要如此简单地说出来,因为 AI 生态系统中仍有声音认为我们应该考虑禁止发布开放模型,理由是安全风险。

Yeah, and my argument is that we should be in a leading position. But I think that it's worth saying it so simply because there are still voices in the AI ecosystem that say we should consider banning releasing open models due to the safety risks.

集中化与开放性 Open Source Models and Global Access

Nathan Lambert

而且我认为值得补充的是,我认为如果没有让美国建立自己的防火墙,这实际上是不可能的,而众所周知防火墙效果并不好,因为训练这些模型的成本——无论是 1 亿还是 1 亿美元——对于世界上许多想要施加影响力的人来说都是可以负担的。所以这些模型将在世界各地被训练,我们希望这些模型,尽管存在安全问题,但信息和工具能够自由地流向世界各地并进入美国,以便人们可以使用和学习它们。阻止这一点将是对我们互联网的巨大重组,似乎是不可能的。

And I think it's worth adding that I think effectively that's impossible without making the US have its own great firewall, which is also known to not work that well because the cost for training these models, whether it's 1 to 100 million dollars, is attainable to a huge amount of people in the world that want to have influence. So these models will be getting trained all over the world and we want these models, especially with safety concerns, but we want this information and tools to flow freely across the world and into the US so that people can use them and learn from them. Stopping that would be such a restructuring of our internet that it seems impossible.

Host

你认为在这种情况下,来自中国的大型开放权重模型在某种意义上对美国公司来说是不是一件好事?因为正如你之前提到的,美国公司在开源发布的内容和他们内部使用的内容之间通常落后一代。例如,GPT-4o 可能不是最前沿的模型,Gemma 3 也可能不是。但他们这样做是因为他们知道发布是安全的。但当他们看到这些公司,比如 DeepSeek 3.2 版本非常出色并被广泛使用,而且没有引起反弹,没有安全风险,这可能会鼓励他们发布更好的模型。也许这在某种意义上是一件非常积极的事情。

Do you think maybe in that case the big open weight models from China are actually a good thing in a sense for US companies? Because US companies, as you mentioned earlier, are usually one generation behind in terms of what they release open source versus what they are using. For example, GPT-4o might not be the cutting edge model, Gemma 3 might not be. But they do that because they know it's safe to release. But when they see these companies, for example DeepSeek version 3.2 which is really awesome and gets used, and there is no backlash, no security risk, that could then encourage them to release better models. Maybe that in a sense is a very positive thing.

Nathan Lambert

百分之百。这些中国公司已经启动了一些事情,我认为如果他们没有发布模型,这些事情可能不会发生。所以我认为……我几乎可以肯定领导层已经进行过这些讨论。

100%. These Chinese companies have set things into motion that I think would potentially not have happened if they were not all releasing models. So I think this... I'm almost sure that those discussions have been had by leadership.

Host

有没有可能未来世界上占主导地位的 AI 模型都是开源的?

Is there a possible future where the dominant models, AI models in the world, are all open source?

Nathan Lambert

这取决于你预测的进展轨迹。如果你认为进展的饱和点将在几年内到来,即在财务支持仍然非常充足的时期内,那么开放模型将被优化得非常好,运行成本也会低得多,从而胜出。本质上,这回到了开源的理念:更多的人将投入资金优化这些开放权重通用架构的服务,使它们成为标准。然后你可以有专门针对它们的芯片,这将比那些封闭公司的定制产品便宜得多。

Depends on the trajectory of progress that you predict. If you think saturation in progress is even coming within a few years, so essentially within the time where financial support is still very good, then open models will be so optimized and so much cheaper to run that they will win out. Essentially, this goes back to open-source ideas where so many more people will be putting money into optimizing the serving of these open-weight common architectures that they will become standards. And then you could have chips dedicated to them, and it'll be way cheaper than the offerings from these closed companies that are custom.

英伟达主导与硬件趋势 Centralization vs. Openness

Host

我们应该说,AI 2027 报告从叙事角度预测的一件事是,随着 AI 系统变得越来越智能,国家安全问题将出现,实验室将集中化,变得超级保密,中美之间将展开一场军事竞赛。所以我们正在进行的这些关于 LLM 的有趣对话,所有将军和士兵都会站出来说:‘好了,我们现在进入了整个事情的曼哈顿计划阶段。’

We should say that the AI 2027 report kind of predicts, one of the things it does from a narrative perspective, is that there'll be a lot of centralization as the AI system gets smarter and smarter, the national security concerns will come to be, and you'll centralize the labs, and you'll become super secretive, and there'll be this whole race from a military perspective between China and the United States. And so all of these fun conversations we're having about LLMs, all the generals, the soldiers will come into their own and be like, 'All right, we're now in the Manhattan Project stage of this whole thing.'

Nathan Lambert

我认为 2025、26、27 年,我不认为这样的事情有一丝可能。我的意思是,你可以对计算机提出同样的论点,对吧?你可以说:‘好吧,计算机很强大,我们不想让公众得到它们。’或者芯片,甚至 AI 芯片,但你看华为现在如何制造芯片,花了几年时间,但我认为你无法遏制这样的东西,这样的知识。我认为在当今时代这是不可能的。就像互联网一样,我不认为这是可能的。

I think 2025, 6, 7, 27, I don't think something like that is even remotely possible. I mean, you can make the same argument for computers, right? You can say, 'Okay, computers are capable, and we don't want the general public to get them.' Or chips, even AI chips, but you see how Huawei makes chips now, took a few years, but I think there is no way you can contain something like that, knowledge like that. I think in this day and age it is impossible. Like the internet, I don't think this is a possibility.

Host

关于曼哈顿计划的事情,我在制作 Adam 时的一个有趣想法是,我认为针对开放模型的类似曼哈顿计划的东西实际上相当合理,因为它不会花费太多,但我认为这将来自……看起来文化上公司正在改变,但我同意 Sebastian 和你刚才说的所有内容。我只是看不到它发生,也不认为它有帮助。

On the Manhattan Project thing, one of my funny things making Adam is that I think a Manhattan Project-like thing for open models would actually be pretty reasonable because it wouldn't cost that much, but I think that will come from... It seems like culturally the companies are changing, but I agree with Sebastian and all of the stuff that you just said. It's just like I don't see it happening nor being helpful.

Nathan Lambert

是的,我的意思是曼哈顿计划背后的动力是文明风险。我不……很难为开源模型找到这样的动力。

Yeah, I mean the motivating force behind the Manhattan Project is civilizational risk. I don't... It's hard to motivate that for open-source models.

Host

没有文明风险。

There's not civilizational risk.

英伟达创新与黄仁勋角色 Nvidia's Dominance and Hardware Trends

Host

嗯,在硬件方面我们多次提到 Nvidia。你认为 Jensen 和 Nvidia 会继续赢下去吗?

Uh, you think on the hardware side we mentioned Nvidia a bunch of times. Do you think Jensen and Nvidia are going to keep winning?

Nathan Lambert

我认为他们的劣势在于需要大量迭代和制造,而且他们可能……他们所做的确实有创新,但我认为总有可能有人做一些根本不同的事情,非常幸运,然后有所成就,但问题是采用。你知道,Nvidia 的护城河可能不仅仅是 GPU,更多的是 CUDA 生态系统,这已经发展了……我是说二十年。我想甚至在我读研究生的时候,我们在一个实验室做生物物理模拟、分子动力学,那时我们就有一块 Tesla GPU 用于计算。那是 15 年前了。他们长期建立这个,我认为这就是护城河。不是芯片本身,尽管他们有钱迭代、建造和扩展,但真正重要的是兼容性。就像,如果你是一家规模这么大的公司,为什么要冒险选择每年只能生产几块芯片的东西?你会选择大公司。但我确实认为,有了 LLM,现在设计像 CUDA 这样的东西会更容易。你知道,下一个……这花了多年时间,因为很难,但现在我们有 LLM,所以我们也许可以复制 CUDA。

I think they have the downside that they have to iterate a lot and manufacture a lot, and I think they probably... what they're doing, they do innovate, but I think there's always the chance that someone does something fundamentally different, gets very lucky, and then does something, but the problem is adoption. You know, the moat of Nvidia is probably not just the GPU, it's more like the CUDA ecosystem, and that has evolved over so many... I mean two decades. I think even back when I was a grad student, we were in a lab doing biophysical simulations, molecular dynamics, and we had a Tesla GPU back then just for computation. That was 15 years ago now. And they built this up for a long time, and that's the moat, I think. It's not the chip itself, although they have the money to iterate and build and scale, but then it's really on the compatibility. It's like, well, if you're at that scale as a company, why would you go with something risky where only a few chips can be made per year? You go with the big one. But then I do think with LLMs now, it will be easier to design something like CUDA. You know, the next... It took years because it's hard, but now we have LLMs, so we can maybe replicate CUDA.

Host

我想知道训练和推理算力是否会分离。随着我们逐渐稳定,越来越多的算力用于推理。

And I wonder if there'll be a separation of the training and the inference compute. As we kind of stabilize a bit more and more compute is needed for inference.

Nathan Lambert

嗯。这应该是 Groq 收购的目的。这也是为什么 Vera Rubin 的一部分是,他们有一款新芯片,没有高带宽内存,或者很少,这是最昂贵的部件之一。它专为预填充设计,这是推理的一部分,你基本上做很多矩阵乘法,然后只有在进行自回归生成和 KV 缓存交换时才需要内存。所以,他们有这种专为特定用例设计的新 GPU,然后每 flop 的拥有成本实际上要低得多。但我认为 Nvidia 的命运仍然取决于 AI 的扩散。他们最大的客户仍然是这些超大规模公司,无论是谷歌,显然可以制造 TPU。亚马逊正在制造 Trainium。微软将尝试做自己的事情。只要 AI 进步的速度很快,Nvidia 的平台就是最灵活的,人们会想要它。

Mhm. That's supposed to be the point of the Groq acquisition. And that's why part of what Vera Rubin is, they have a new chip with no high bandwidth memory, or very little, which is one of the most expensive pieces. It's designed for prefill, which is the part of inference where you essentially do a lot of matrix multiplications, and then you only need the memory when you're doing this autoregressive generation and you have the KV cache swaps. So, they have this new GPU designed for that specific use case, and then the cost of ownership per flop or whatever is actually way lower. But I think that Nvidia's fate lies in the diffusion of AI still. Their biggest clients are still these hyperscale companies, whether it's Google, obviously can make TPUs. Amazon is making Trainium. Microsoft will try to do its own things. And as long as the pace of AI progress is high, Nvidia's platform is the most flexible and people will want that.

网络与算力缩放 Nvidia's innovation and Jensen's role

Nathan Lambert

但如果出现停滞,那么制造定制芯片就有更多时间去做。

But if there's stagnation, then creating bespoke chips, there's more time to do it.

Host

有趣的是,英伟达非常积极地尝试开发各种不同的产品。他们试图创造会使用大量 GPU 的商业价值领域。但他们持续创新,做了很多令人难以置信的研究。

It's interesting that Nvidia is quite active in trying to develop all kinds of different products. They tried to create areas of commercial value that will use a lot of GPUs. But they keep innovating and they're doing a lot of incredible research, so.

Nathan Lambert

每个人都说这家公司极度以黄仁勋为中心,他在运营上深度参与。这听起来和我听过的许多其他大公司很不一样。只要这种文化还在,我认为他们会继续取得进展。这就像他仍处于苹果的史蒂夫·乔布斯时代。只要公司这样运作,我对他们的处境相当乐观,因为这是他们的首要问题,而我不知道为整个生态系统制造这些芯片是否是其他所有公司的首要目标。他们会做得不错,但可能不会那么好。

Everyone says that the company is super oriented around Jensen and how operationally plugged in he is. And it sounds so unlike many other big companies that I've heard about. And so long as that's the culture, I think that I will expect them to keep progress happening. And it's like he's still in the Steve Jobs era of Apple. So long as that is how it operates, I'm pretty optimistic for their situation because it's like it is their top order problem and I don't know if making these chips for the whole ecosystem is the top goal of all these other companies. They'll do a good job, but it might not be as good of a job.

Host

既然你提到了黄仁勋,我读了很多关于历史和历史上单一人物的内容。你们怎么看历史的单一人物观?个人在科技领域引导历史方向有多重要?所以,你知道,没有黄仁勋的英伟达会怎样?你提到了史蒂夫·乔布斯,没有乔布斯的苹果会怎样?没有埃隆的 xAI 呢?或者没有戴米斯的 DeepMind?

Since you mentioned Jensen, I've been reading a lot about history and about singular figures in history. What do you guys think about the single man woman view of history? How important are individuals for steering the direction of history in the tech sector? So, you know, what's Nvidia without Jensen? You mentioned Steve Jobs, what's Apple without Steve Jobs? What's xAI without Elon? Or DeepMind without Demis?

Nathan Lambert

人们让事情更早更快地发生,在科学上,许多伟大的科学家归功于在正确的时间出现在正确的地点,并做出创新,而最终别人也会想到这个想法。所以我认为,从这个角度看,黄仁勋正在帮助这场 GPU 革命更快、更聚焦地实现,比没有这样一个人要快得多。这正在加速整个人工智能的建设。但我仍然认为,最终像 ChatGPT 这样的东西会发生,这样的建设也会发生,但可能不会这么快,或者我认为这就是那种被赋予的风味。

People make things earlier and faster where scientifically many great scientists credit to being the right place at the right time and still making the innovation where eventually someone else will still have the idea. So I think that in that way Jensen is helping manifest this GPU revolution much faster and much more focused than without having a person there it would do. And this is making the whole AI build out faster. But I do still think that eventually something like ChatGPT would have happened and a build out like this would have happened, but it probably would not have been as fast or as like I think that's the sort of flavor that is applied.

Host

这些个人在押注某些东西。有些人幸运,有些人不幸。但如果没有这些人掌舵,事情会过于分散。这几乎就像投资 ETF 与个股。个股可能涨跌幅度比更平衡的 ETF 更大。随着时间的推移,它最终会上涨。我们会到达那里。但就像,你知道,我认为专注是关键。充满激情的专注。

People these individual people are placing bets on something. Some get lucky, some don't. But if you don't have these people at the helm, it would be more too diffuse. It's almost like investing in an ETF versus individual stocks. Individual stocks might go up, might go down more heavily than an ETF which is more balanced. It will eventually go up over time. We'll get there. But it's just like, you know, focus I think is the thing. Passionate focus.

Nathan Lambert

难道不是有充分的理由认为,没有黄仁勋,就不会有深度学习革命的复兴吗?

Isn't there a real case to be made that without Jensen, there's not a reinvigoration of the deep learning revolution.

Host

我想说的是,可能会晚 20 年。

It could have been 20 years later is the thing that I would say.

Nathan Lambert

是的,是的,是的,20 年。

Yeah, yeah, yeah, 20.

Host

另一个人工智能寒冬,比如深度学习寒冬,可能会到来。

Another AI winter, like a deep learning winter could have come.

Nathan Lambert

是的。

Yeah.

Host

如果 GPU 不存在的话。

If GPUs weren't around.

Nathan Lambert

那可能会完全改变历史,因为你可以想到在此期间可能出现的所有其他技术,人类文明的焦点可能会是,硅谷会被不同的炒作所占据。

That could change history completely because you could think of all the other technologies that could have come in the meantime and the focus of human civilization could be, Silicon Valley would be captured by different hype.

Host

但我确实认为,一方面 GPU 的轨迹是计划好的,但另一方面也有很多幸运的巧合。例如,或者好的直觉,比如对生物物理模拟的投资,我的意思是,我认为它始于视频游戏,然后恰好擅长线性代数,因为视频游戏需要大量线性代数,然后有了生物物理模拟,但我仍然不认为总体规划是人工智能。我认为只是恰好有 Alex Krizhevsky。所以有人拿这些 GPU 说,嘿,让我们尝试在上面训练神经网络,结果效果非常好,我认为这之所以发生,只是因为你可以买到那些 GPU。

But I do think that there's certainly an aspect where it was all planned the GPU trajectory, but on the other end it's also a lot of lucky coincidences. For example, or good intuition like the investment into this, let's say biophysical simulations or like I mean I think it started with video games and then it just happened to be good at linear algebra because video games require a lot of linear algebra and then you have the biophysical simulations and then but still I don't think the plan, the master plan was AI. I think there was just it happened to be Alex Krizhevsky. So someone took these GPUs and like hey let's try to train a neural network on that it happened to work really well and I think it only happened because you could purchase those GPUs.

Nathan Lambert

如果英伟达早期倒闭了,游戏也会创造对更快处理器的需求。

Gaming would have created a demand for faster processors if Nvidia had got out of business in the early days.

Host

嗯。

Mhm.

Nathan Lambert

我是这么想的,我认为 GPU 对于 Alex 来说可能会不同,但我认为在 AlexNet 时代和 Transformer 时代 GPU 仍然会存在。只是很难知道是一家公司如此成功,还是多家拥有较差芯片的小公司,但我不认为那会是 100 年的延迟。可能是十年的延迟。

That's what I would think like I think that the GPUs would have been different for the Alex but I think like GPUs would still exist at the time of AlexNet and at the time of the transformer. It was just hard to know if it would be one company as successful or multiple smaller companies with worse chips but I don't think that's like a 100 year delay. It might be a decade delay.

Host

嗯,可能是一、二、三、四、五十年的延迟。我的意思是,我无法想象英特尔或 AMD 会做英伟达所做的事。

Well, it could be one, two, three, four, five decade delay. I mean I just can't see Intel or AMD doing what Nvidia did.

Nathan Lambert

会有一家存在的公司。我认为会有另一家公司崛起。

It would be a company that exists. I think it would be a different company would rise.

Host

硅谷图形公司之类的。

Silicon Graphics or something.

Nathan Lambert

所以是的,一些已经倒闭的公司可能会做到。

So yeah, some company that has died would have done it.

Host

但确实,仅仅看这一点,这些单一人物,这些领导者对世界的轨迹有着巨大的影响。显然,他们背后有不可思议的团队。但是,你知道,拥有那种非常单一、几乎教条式的专注对于取得进展是必要的。

But it does like just looking at it, it seems like these singular figures, these leaders have a huge impact on the trajectory of the world. Obviously, incredible teams behind them. But, you know, having that kind of very singular, almost dogmatic focus is necessary to make progress.

Nathan Lambert

是的,我的意思是,即使是 GPT,如果没有 Ilya 这个人推动这种 Scaling(规模扩张),它也不会存在,对吧?

Yeah, I mean even with GPT, it wouldn't exist if there wasn't a person, Ilya, who pushed for this scaling, right?

Host

是的,Dario 也深度参与其中。你读一些 OpenAI 的历史。想到这些人多早就说‘我们需要连接 10,000 个 GPU,用 OpenAI 所有的算力训练一个模型’,这几乎显得疯狂。那里有很多人不想这么做。

Yeah, Dario was also deeply involved in that. You read some of the histories of OpenAI. It almost seems wild thinking about how early these people were like, 'We need to hook up 10,000 GPUs and take all of OpenAI's compute and train one model.' There's a lot of people there that didn't want to do that.

Nathan Lambert

在 Scaling(规模扩张)有任何迹象表明它会实现之前就相信 Scaling(规模扩张),这是一件疯狂的事情。再次,单一人物。说到这个,40 年后,这大概是后奇点时代,无论奇点是什么,当历史学家回顾我们现在的时代时,他们会真正强调哪些技术突破导致了奇点?那么,到目前为止,从图灵到今天,80 年。

Which is an insane thing to believe that to believe scaling before scaling has any indication that it's going to materialize. Again, singular figures. Speaking of which, 40 years from now, this is presumably post-singularity, whatever singularity is, when historians look back at our time now, what technological breakthroughs would they really emphasize as the breakthroughs that led to the singularity? So, so far we have Turing to today, 80 years.

Host

我认为仍然是计算,就像总称计算。我不一定认为 100 年、200 年后会是人工智能。它仍然会是计算机,你知道?只是我们现在更好地利用了计算机,但就像计算这个事实。

I think it would still be computing, like the umbrella term computing. Just I don't necessarily think it's even like 100 years, 200 years from now it would be AI. It would still be, still well be computers, you know? Just we are now taking better advantage of computers, but like the fact of computing.

Nathan Lambert

这基本上是一个摩尔定律式的讨论。你甚至不会记得代码和 GPU 的细节。也不会有所有这些软件动荡。它只会是显而易见的计算。

It's a basically Moore's law kind of discussion. You're not even the details of code and GPUs won't even be remembered. And it won't be all this software turmoil. It'll be just obviously compute.

Host

基本同意,但问题是互联网的连接性和计算能否合并,还是两者都是?

Generally agree, but it's like is the connectivity of the internet and compute able to be merged or is that both of them?

Nathan Lambert

我认为互联网可能与之相关,是的,我的意思是通信,可能是手机互联网或卫星之类的东西。而计算机更像是其中的 Scaling(规模扩张)方面。

I think the internet will probably related to yeah, I mean communication that it could be a phone internet or satellite that stuff. Where yeah, and computer is that more like the scaling aspect of it.

Host

有可能互联网被完全遗忘了。

It's possible that the internet has completely forgotten.

神经网络作为突破 Networking and Compute Scaling

Host

互联网被嵌入电话网络这样的通信网络中。这只是另一种体现,真正的突破来自于算力的增长,也就是广义上的摩尔定律。

The internet is wrapped into the phone networks like communication networks. This is just another manifestation of that, and the real breakthrough comes from the increased compute, which is Moore's law broadly defined.

Nathan Lambert

嗯,我认为人与人之间的连接非常关键。就像你可以和任何人交谈,找到世界上某个领域最优秀的人,实现信息流动。AI 也将依赖这一点。我一直纠结于一个想法:单一中心模型的梦想已经破灭,正在演变的是人们为不同任务使用多个智能体。人们开始为不同任务使用不同的云,这被描述为数据中心里的多个 AGI,每个管理自己的部分并相互通信。这非常依赖于网络和算力之上的自由信息流动,但网络,尤其是 GPU 之间的网络,是扩展算力的重要部分。数据中心里的 GPU 需要相互通信。

Well, I think that connection of people is very fundamental to it. So it's like you can talk to anyone to find the best person in the world for something, somewhere in the world, and being able to have that flow of information. The AIs will also rely on this. I've been fixating on the idea that the dream of one central model is dead, and what's evolving is that people have many agents for different tasks. People are starting to do this with different clouds for different tasks, described as many AGIs in the data center where each one manages and they talk to each other. That is very reliant on networking and free flow of information on top of compute, but networking especially with GPUs is such a part of scaling up compute. The GPUs in the data centers need to talk to each other.

未来世界:机器人与设备 Neural Networks as Breakthrough

Host

神经网络有什么会被记住的吗?你是否认为神经网络被视为突破有其独特之处,就像一种天才,你基本上是在以一种非常粗糙的方式复制人类思维?人脑的结构,人类思维。

Anything about neural networks will be remembered? Do you think there's something very specific and singular to the fact that it's neural networks that are seen as a breakthrough, like a genius that you're basically replicating in a very crude way the human mind? The structure of the human brain, the human mind.

Nathan Lambert

我认为没有人类思维我们可能就不会有神经网络,因为它是一个灵感来源。但另一方面,我认为它非常不同。我的意思是,它是数字的,而生物是生物的。我确实认为它可能更会被归类为一种算法。

I think without the human mind we probably wouldn't have neural networks because it was an inspiration for that. But on the other hand, I think it's so different. I mean, it's digital versus biological. I do think it will probably be more like grouped as an algorithm.

Host

它在这种特定算力上是高度可并行的。

That's massively parallelizable on this particular kind of compute.

Nathan Lambert

也可能是遗传计算,比如遗传算法,同样可以并行化。只是碰巧神经网络更高效、效果更好。

Could have been genetic computing, like genetic algorithms, just as parallelized. It just happens that this is more efficient, works better.

Host

而且很可能,语言模型、神经网络,我们目前构建它们的方式只是通向奇点的系统中的一个小组件。

And it very well could be that the LM, the neural networks, the way we architect them now is just a small component of the system that leads to singularity.

Nathan Lambert

我认为如果你考虑 100 年后,社会会因为更多的算力和智能而因自主性发生更大变化。但就像回顾工业革命:我们记得什么?我们记得发动机,它可能相当于这里的计算机。但还有很多其他物理变革人们也知道,比如轧棉机、空调、冰箱。AI 中的一些东西也会被记住。'Transformer' 这个词可能仍然被知道。我猜深度学习肯定还会被记住,但 Transformer 可能在 100 年后随着到处都是超级智能 AI 研究者而进化消失。但我认为深度学习很可能是一个会被记住的术语。

I think if you think of it in 100 years, society can be changed more with more compute and intelligence because of autonomy. But it's like looking at the Industrial Revolution: what do we remember? We remember the engine, which is probably the equivalent of the computer in this. But there are many other physical transformations that people are aware of, like the cotton gin, air conditioning, refrigerators. Some things from AI will still be known. The word 'transformer' could still be known. I would guess that deep learning is definitely still known, but the transformer might have evolved away in 100 years with ASI AI researchers everywhere. But I think deep learning is likely to be a term that is remembered.

失业的社会影响 Future World: Robots and Devices

Host

我想知道 AI 带来的未来的空调和冰箱是什么。如果我们从现在穿越到 100 年后,直接到达那里。你认为有什么不同?世界看起来会怎样?首先,你认为还有人类吗?你认为到处都是机器人走来走去吗?

And I wonder what the air conditioning and refrigeration of the future is that AI brings. If we travel forward 100 years from now, we'll transport there right now. What do you think is different? How do you think the world looks different? First of all, do you think there are humans? Do you think there are robots everywhere walking around?

Nathan Lambert

我确实认为肯定会有专用机器人,用于某些任务。

I do think specialized robots for sure, for certain tasks.

Host

人形吗?

Humanoid form?

Nathan Lambert

也许半人形。我们拭目以待。我认为对于某些事情,是的,会有类人机器人,因为它们适应环境。但对于某些任务,可能也有道理。更难想象的是我们如何与设备交互,以及人类用设备做什么。嗯,我很确定可能不会是手机,也不会是笔记本电脑。会是植入物吗?

Maybe half humanoid. We'll see. I think for certain things, yes, there will be humanoid robots because it's amenable for the environment. But for certain tasks, it might make sense. What's harder to imagine is how we interact with devices and what humans do with devices. Well, I'm pretty sure it will probably not be the cell phone. It will probably not be the laptop. Will it be implants?

Host

我的意思是,必须是脑机接口,对吧?考虑到我们目前看到的进展,100 年后必须如此,除非我们与现实互动的方式发生了彻底的改变。

I mean, it has to be brain-computer interfaces, right? I mean, 100 years from now it has to, given the progress we're seeing now, unless there's a legitimate complete alteration of how we interact with reality.

Nathan Lambert

另一方面,想想汽车,汽车已有 100 多年历史,对吧?而且界面还是一样的。我们没有用别的东西取代汽车。我们只是把汽车做得更好,但仍然是方向盘,仍然是轮子。

On the other hand, if you think of cars, cars are older than 100 years, right? And it's still the same interface. We haven't replaced cars with something else. We just made the cars better, but it's still steering wheel, still wheels.

Nathan Lambert

我认为我们仍然会随身携带一个物理计算块,因为人们希望有一定能力拥有私人信息。你可能不会像用手机那样频繁使用它,但拥有一个可以存放属于你的私人信息、作为与互联网其余部分接口的设备,我认为这仍然会存在。它可能看起来不像 iPhone,使用频率也可能低得多,但我仍然预计人们会随身携带东西。

I think we'll still carry around a physical brick of compute because people want some ability to have private information. You might not engage with it as much as a phone, but having something where you can have private information that is yours as an interface between the rest of the internet, I think is something that will still exist. It might not look like an iPhone, and it might be used a lot less, but I still expect people to carry things around.

Host

你为什么认为智能手机是私密的体现?它上面有摄像头。

Why do you think the smartphone is the embodiment of private? There's a camera on it.

Nathan Lambert

对你来说是私密的,比如加密消息、加密照片,你知道你的生活是什么。我想这取决于你对脑机接口有多乐观。如果所有这些都只是存储在云端,你的整个日历……很难想象通过脑机接口以视觉方式处理所有信息,比如向你呈现日历。很难想象不用看就知道你的电子邮件收件箱。你向计算机发出信号,然后你就知道你的收件箱。那意味着什么?人脑能处理非视觉方式输入的信息吗?我不太清楚这些转变如何发生,因为人类在 100 年内不会改变。我认为自主性和社群是人们真正想要的东西。

Private for you, like encrypted messages, encrypted photos, you know what your life is. I guess this is the question on how optimistic on brain-machine interfaces you are. If all of that is just going to be stored in the cloud and your whole calendar... It's hard to think about processing all the information that we can process visually through brain-machine interfaces presenting something like a calendar to you. It's hard to just think about knowing without looking, your email inbox. You signal to a computer and then you just know your email inbox. What does that mean? Is that something the human brain can handle being piped into it non-visually? I don't know exactly how those transformations happen because humans aren't changing in 100 years. I think agency and community are things that people actually want.

Host

本地社群,是的。

Local community, yeah.

Nathan Lambert

所以,你亲近的人,能和他们一起做事,能为你的生活赋予意义并有所作为。我认为,即使不是 100 年内,我也不认为人类生物学会在我们讨论的时间尺度上改变这些。我认为全民基本收入并不能解决自主性问题。我确实期望大量财富,并希望它被分配,这样普通人的生活 100 年后会大不相同,但考虑到那些处于发展早期阶段的国家要获得计算和互联网、建设所有基础设施、并制定政策将一国财富与另一国分享,100 年内发生这一切仍然很多。认为所有这些都能在 100 年内发生,同时它们仍然是独立实体,而不是被武力吸收到某种国际秩序中,这是一个乐观的看法。

So people you are close to, being able to do things with them, and being able to ascribe meaning to your life and be able to do things. I think that is maybe if not in 100 years, I don't think human biology is changing away from those on a time scale we can discuss. I think that UBI does not solve agency. I do expect mass wealth, and I hope it is spread so that the average life does look very different in 100 years, but that's still a lot to happen in 100 years if you think about countries that are early in their development process getting access to computing and internet, to build all the infrastructure and to have policy that shares one nation's wealth with another. It's an optimistic view to see all that happening in 100 years while they are still independent entities and not just absorbed into some international order by force.

对未来的希望 Social impact of job loss

Nathan Lambert

但可能会有更好、更完善、更有效的社会支持体系,帮助减轻世界上一些基本的痛苦。你知道,社会转型中短期内会失去很多工作岗位。我认为我们必须记住,每一个失去的工作岗位都是一个正在受苦的人。大规模失业是一场真正的悲剧。你可以提出各种经济学论点,说一切都会好起来,这对 GDP 有好处,会有新的就业机会。但根本上,对那个个体来说,那是真实的痛苦,是个人悲剧。我们在开发技术时不能忘记这一点。

But there could be just better, more elaborate, more effective social support systems that help alleviate some levels of basic suffering from the world. You know, the transformation of society where a lot of jobs are lost in the short term. I think we have to really remember that each individual job that's lost is a human being who's suffering. That's a real tragedy when jobs are lost at scale. You can make all kinds of arguments about economics or it's all going to be okay, it's good for GDP, there will be new jobs created. Fundamentally at the individual level for that human being, that's real suffering. That's a real personal tragedy, and we have to not forget that as the technologies are being developed.

Host

而且,对于我们看到的所有 AI 垃圾内容,我的希望是,人类体验中那些面对面的基本方面会越来越有价值。我们都喜欢的事情:见面、一起交谈,面对面。

And also, my hope for all the AI slop we're seeing is that there will be a greater and greater premium on the fundamental aspects of the human experience that are in person. The things we all like: seeing each other, talking together, in person.

Nathan Lambert

未来几年,实物和活动的价值肯定会增加,而垃圾内容会面临更大压力。

The next few years are definitely going to see an increased value on physical goods and events, and then even more pressure on slop.

Host

嗯。

Mhm.

Nathan Lambert

所以,垃圾内容才刚刚开始。未来几年会出现越来越多样化的垃圾内容。

So, the slop is only starting. The next few years will see more and more diverse versions of slop.

Host

垃圾内容?我们社会会被垃圾内容淹没,直到我们幡然醒悟,觉得无法应对。然后实物就会变得更有价值。

Slop? We'd be drowning in slop as a society, enough to snap out of it and be like, we can't deal with it. And then the physical has such a higher premium on it.

Nathan Lambert

即使是经典例子,我真心认为这是真的,而且我们会厌倦它。我们已经有点厌倦了。艺术也一样。我认为艺术不会消失。有绘画,实体绘画。它们有更多价值,不仅是金钱上的,而是对真迹的欣赏,而不是复制品。即使数字复制品完美无缺,但当你去博物馆看那件艺术品,看到真迹时,你会想到一个人。这是一种你欣赏的工艺。我认为写作、交谈、任何体验都是如此。不幸的是,我认为会有一个二分法:有些东西会被自动化,比如现在没有 200 年前那么多画了,更多是照片和复制品。但它不会消失。其中会有价值。区别只是比例。就我个人而言,我很难阅读那些明显是 AI 生成的内容。我会想,抱歉,信息可能很好,但我不想要。

Even classic examples, I honestly think this is true, and I think we'll get tired of it. We are already kind of tired of it. Same with art. I don't think art will go away. You have paintings, physical paintings. There's more value, not just monetary, but appreciation for something that is the actual painting than a photocopy. It could be a perfect digital reprint, but there's something when you go to a museum and look at that art and see the real thing, you think about a human. It's a craft you appreciate. I think the same is true for writing, for talking, for any type of experience. Unfortunately, I think it will be a dichotomy: some things will be automated, like there aren't as many paintings as 200 years ago, more photographs, more photocopies. But it won't go away. There will be value in that. The difference will be the proportion. Personally, I have a hard time reading things where I obviously see it's AI-generated. I'm like, sorry, it might be really good information, but I have a certain nah, not for me.

Host

我认为最终它们会骗过你,而且会在那些提供验证或建立信任方式的平台上。所以,你会相信 Lex 不是 AI 生成的,因为你在这里。所以你对这个频道有信任。但对于没有这种信任的新人来说就更难了。

I think eventually they'll fool you and it'll be on platforms that give ways of verifying or building trust. So, you will trust that Lex is not AI-generated having been here. So, you have trust in this channel. But it's harder for new people that don't have that trust.

Nathan Lambert

嗯,这会变得有趣,因为我认为从根本上说,这是一个可以解决的问题,通过信任某些不会这么做的渠道,但一切都将基于信任。会有一些系统来认证,这是真的,这不是真的。会有一些明显的迹象让你能分辨出这是 AI 生成的,但有些会好到难以分辨,然后你就得信任。这会变得有趣,也有点问题。

Well, that will get interesting because I think fundamentally it's a solvable problem by having trust in certain outlets that they won't do it, but it's all going to be trust-based. There will be some systems to authorize, okay, this is real, this is not real. There will be some telltale signs where you can obviously tell this is AI-generated and this is not, but some will be so good that it's hard to tell, and then you have to trust. That will get interesting and a bit problematic.

Host

极端情况是给所有人类内容加水印。

The extreme case of this is to watermark all human content.

Nathan Lambert

嗯。

Mhm.

Host

所以,我们自己拍的所有照片都有某种水印,直到它们被编辑或类似操作,软件可以与设备制造商通信以维护人类编辑。这与试图给 AI 图像加水印的讨论相反,然后你可以制作一个有水印的谷歌图片,再用另一个谷歌工具去除水印。

So, all photos that we take on our own have some watermark until they're edited or something like this, and software can manage communications with the device manufacturer to maintain human editing. Which is the opposite of the discussion to try to watermark AI images, and then you can make a Google image that has a watermark and use a different Google tool to remove the watermark.

Nathan Lambert

是的,基本上会是一场军备竞赛。

Yeah, it's going to be an arms race basically.

Host

是的,是的。

Yeah, yeah.

后奇点场景中的人与 AI Hope for the future

Host

我们主要关注了 AI 的积极方面。但我们也谈到的能力也可能被用来破坏人类文明,即使只是相对笨拙的 AI 大规模应用,以及更进一步的超级智能 AI 系统。当然,悲观的观点也很重要,我们在开发这些技术时需要考虑一下。是什么让你对人类文明的未来充满希望?我们一直在谈论的一切。我们会没事吗?

And we've been mostly focusing on the positive aspects of AI. I mean, there's also the capabilities we've been talking about can be used to destabilize human civilization with even just relatively dumb AI applied at scale and then further super intelligent AI systems. Of course, there's the doomer take that's important to consider a little bit as we develop these technologies. What gives you hope about the future of human civilization? Everything we've been talking about. Are we going to be okay?

Nathan Lambert

我认为我们会没事的。我确实对 AI 和非 AI 的事情都很担忧,但人类总能找到办法。我认为这就是人类的天性:建立社区,想办法解决问题,这让我们走到了今天。我认为 AI 的机会和相关技术非常巨大,而且存在很大的社会和政治问题需要帮助每个人理解。我认为我们现在面临很多这样的问题:世界是一个可怕的地方,AI 是非常不确定的事情。这需要很多工作,不一定是建设性的。而是要告诉人们并理解,历史上开发 AI 的人并没有动机或想要做坏事。这可能是可行的,只是需要比人们期望更长的时间。如果我们想要持久的利益,就必须经历那段漫长而痛苦的 AI 讨论时期。

I think we will. I'm definitely a worrier about AI and non-AI things, but humans do tend to find a way. I think that's what humans are built for: to have community and find a way to figure out problems, and that's what has gotten us to this point. I think that the AI opportunity and related technologies is really big, and I think that there are big social and political problems to help everybody understand that. I think that's what we're staring at a lot of right now: the world is a scary place and AI is a very uncertain thing. It takes a lot of work that is not necessarily building things. It's telling people and understanding that the people building AI are historically not motivated or wanting to do that. It is something that is probably doable. It just will take longer than people want. We have to go through that long period of hard, distraught AI discussions if we want to have the lasting benefits.

Host

是的,通过这个过程,我特别兴奋的是我们有机会更好地了解自己。无论是在个体层面还是文明层面。是的,有一些大的谜团。比如,这整个意识是怎么回事?似乎真的很特别。我们的头脑中有一个真正的奇迹。而 AI 给我们提供了一面镜子,让我们回答一些大问题:这一切到底是怎么回事?

Yeah, through that process I'm especially excited that we get a chance to better understand ourselves. Also at the individual level as humans and at the civilizational level. Yeah, there's some of the big mysteries. Like, what is this whole consciousness thing going on here? Seems to be truly special. There's a real miracle in our mind. And AI puts a mirror to ourselves and get to answer some of the big questions about what is this whole thing going on here?

Nathan Lambert

嗯,关于这一点,我认为让我们与 AI 非常不同、也是我不担心 AI 接管的原因,就像你说的意识。我们人类决定自己想做什么。AI 在当前的实现中,我看不到它会改变。你必须告诉它做什么。所以你仍然拥有自主权。它不会夺走你的自主权,因为你把它当作工具。你可以把它看作一个工具。你告诉它做什么。它会比以前的工具更自动化。它当然比锤子更强大。

Well, one thing about that is also what I do think makes us very different from AI and why I don't worry about AI taking over is like you said consciousness. We humans, we decide what we want to do. AI in its current implementation I can't see it changing. You have to tell it what to do. And so you still have the agency. It doesn't take the agency from you because you have to use it as a tool. You can think of it as a tool. You tell it what to do. It will be more automatic than other previous tools. It's certainly more powerful than a hammer.

Human vs AI in a post-singularity scenario Human vs AI in a post-singularity scenario

Host

它能解决问题,但掌控权还是在你手里,对吧?所以 AI 并不做主,你做主。你告诉 AI 该做什么,它为你执行。

It can figure things out, but it's still you in charge, right? So the AI is not in charge. You're in charge. You tell the AI what to do and it's doing it for you.

Nathan Lambert

所以在后奇点、后末日的人机战争中,你是说人类值得为之战斗。

So in a post-singularity post-apocalyptic war between humans and machines, you're saying humans are worth fighting for.

Host

百分之百。我的意思是,这基本上就是 80 年代拍的《终结者》电影。我认为唯一可能出问题的地方,就是如果程序被明确设定去做有害的事情。

100%. I mean, this is the movie Terminator they made in the '80s essentially. And I do think the only thing I can see going wrong is if things are explicitly programmed to do the thing that is harmful, basically.

Nathan Lambert

我实际上认为,在那种《终结者》式的设定中,人类会赢。

I think actually in that Terminator type of setup, I think humans win.

Host

嗯。

Mhm.

Nathan Lambert

我觉得我们太聪明了。很难解释我们如何想出办法,但我们总能做到。而且我们很可能会用本地 LLM、开源 LLM 来帮助对抗机器。我为这种荒谬道歉。就像我说的,Nathan 早就知道我是他的长期粉丝。我也是 Sebastian 的长期粉丝。所以终于见到你们是莫大的荣幸。感谢你们为世界贡献的一切。感谢你们写的优秀书籍。感谢你们教导我们。也感谢你们今天的对话。这很有趣。

I think we're too clever. It's hard to explain how we figure it out, but we do. And we'll probably be using local LLMs, open-source LLMs to help fight the machines. I apologize for the ridiculousness. Like I said, Nathan already knows I've been a big fan of his for a long time. Been a big fan of yours Sebastian for a long time. So it's an honor to finally meet you. Thank you for everything you put out into the world. Thank you for the excellent books you're writing. Thank you for teaching us. And thank you for talking today. This was fun.

Sebastian Raschka

感谢你邀请我们,建立这种人际联系,这实际上是非常宝贵的人际联系。

Thank you for inviting us here and having this human connection which is actually extremely valuable human connection.

Host

感谢收听本期与 Sebastian Raschka 和 Nathan Lambert 的对话。要支持本播客,请查看描述中的赞助商,那里也有联系我、提问、反馈等链接。现在,让我用阿尔伯特·爱因斯坦的一句话作为结束:并不是我有多聪明,而是我坚持思考问题的时间更长。感谢收听,期待下次再见。

Thanks for listening to this conversation with Sebastian Raschka and Nathan Lambert. To support this podcast, please check out our sponsors in the description where you can also find links to contact me, ask questions, give feedback, and so on. And now, let me leave you with some words from Albert Einstein. It is not that I'm so smart, but I stay with the questions much longer. Thank you for listening and hope to see you next time.

互动版:逐字朗读 + 针对本期提问 →