Open Models as the Engine for the Next Decade of AI Research
打开互动全文版(中英对照 + 朗读 + 问答)→Nathan Lambert 探讨开放模型在 AI 研究、地缘政治中的作用,以及影响力从美国向中国的转移。
Nathan Lambert discusses the role of open models in AI research, geopolitics, and the shift of influence from the US to China.
大家好。我很高兴能邀请到 Nathan Lambert,他是艾伦人工智能研究所的研究科学家,也是最优秀的专家之一、开放模型倡导者和非常善于表达的科普者。我想我是你在 2022 年的第一批付费订阅者之一。所以,你可以看出我是你的忠实粉丝。
Hello everyone. I'm very happy to be hosting Nathan Lambert, research scientist at Allen Institute for AI, one of the best specialists, open model advocate and very well articulated educator. I think I was one of your first paying subscribers back in 2022. So, you can tell I'm a big fan.
是啊,一直很有趣。多年来我们在网上和现实中都有交集。所以很高兴能参加这个播客。
Yeah, it's been fun. We've been crossing paths for years now on the web and in real life. So, it's fun to join the pod.
当你开始职业生涯时,你能想象自己会成为名人吗?
When you were starting your career, could you imagine that you would become a celebrity?
不。我经常和我的伴侣及家人讨论这件事,很大程度上是因为 AI 行业的动态。事情发展得太快了,有很多非常善于表达和善于教育的人去了这些实验室,可以理解地全身心投入构建 AI。而且由于其中的利害关系,这些实验室的人不太公开发推文。OpenAI 是一个特例,但很多这样的沟通者都不能说话,所以这就形成了一个空白,而我就被推入了这个空白。我认为没有任何方法可以真正让你的人生或处事方式做好准备,去如此快速地经历影响力层级的提升。现在,我拥有了这种能力,它如何改变我与工作以及我所处理问题的关系,以便仍然能产生影响?就像我被邀请参加各种 fancy 的活动,但大多数都毫无意义,而且很容易处理——你只需要拒绝它们。但我还没有完全理解拥有这种能力意味着什么。我认为在过去一年里,我确实更加有意识地努力帮助 AI 政策顺利推进。我认为我们做的很多事情,比如构建记录详尽、解释周全的开放模型,可能对政策方面影响最大,因为开放模型具有很强的地缘政治性,并且关系到 AI 如何扩散到全世界,这与当前明显流行的方向——编码智能体变得如此出色并带来加速——是非常不同的轨道。有人试图让这些话题看起来是关联的,是的,中国的模型玩家也在发布智能体,但我真的认为它们并不像人们想象的那么关联。而且很多开放模型的能力被过度炒作。这是一个反复出现的情况:是的,如果开放模型更好会很酷,但我认为坐下来使用 Claude Code 配合 GPT-4.5 或 Codex,与摆弄一个开放模型是非常不同的。这并不是说它们不重要。我认为主要是我不喜欢在地缘政治领域工作,但可以说开放模型是新兴技术的一个案例研究,理解新兴技术如何影响世界并创造新的影响力领域是很有趣的。但这对我来说太新了。我完全没有受过这方面的训练。所以我正在努力学习。
No. And I talk about this with my partner and family regularly where it's just like largely due to the dynamics of the AI industry. Things have evolved so fast where there are just so many people who are very articulate and good educators that went to these labs to understandably go all in on building AI. And part of that with the stakes involved like people at these labs don't tweet publicly that much. OpenAI is a whole special thing but a lot of these communicators kind of can't talk so that's this void that I have been launched into. And I think there is no way to actually prepare your life or approach where you go through levels of influence so quickly. And now it's just like I have this ability and how does it change the relationship to what I work on and the problems that I approach to still have impact? It's like I get invited to all sorts of fancy things and it's just like most of them serve no purpose and it's easy. It's like you just have to say no to them. But I have not fully grappled with what it means to have that ability. And I think in the last year I've definitely become more conscious of trying to help AI policy go well. I think a lot of what we do building open models that are so well documented with care and explaining all the sides probably is most impactful on the policy side of things because open models is so geopolitical and how AI will diffuse through the world which is a very different track than what is so obviously in vogue right now which is that the coding agents are becoming so good and so impactful in this acceleration that comes with it. There's attempts to make it seem like these are linked topics where yes, the Chinese model players are releasing agents as well, but I really think that they're not as linked as people think. And a lot of the open models are overhyped in their abilities. And it's like kind of a recurring thing where yes, it would be cool if the open models are better, but I just think that it's very different to sit down and use Claude Code for GPT-4.5 or Codex than it is to play with an open model. That's not to say that they don't matter. I think it's mostly a I don't like working in geopolitics, but it's I would say maybe open models are such a case study in emerging technology and understanding how emerging technology influences the world and creates new pockets of influence is interesting. But that's so new to me. I'm not trained in that at all. So I'm trying to learn.
这正是我的问题。如果开放模型没那么好,你为什么对开放模型如此热情?为什么美国有这么多关于开放模型的讨论?原因是什么?
That's exactly my question. If open models are not that good, why are you so passionate about open models? Why are there so many conversations about open models for the US? What's the reason?
它们将成为未来 10 年 AI 研究的引擎,因为学术界——我本来想说被打压,但这不是正确的说法——学术界在科学演进方面的影响力正处于低谷,以至于人们会说学术 AI 研究现在不重要,我认为这有点短视,因为它将成为探索的引擎,而公司无法真正培育这种探索,但开放模型是这种创新发生的平台。对于像美国这样拥有优秀科学项目和机构历史的国家,如果他们想成为 AI 研究的机构,他们应该认为有意、有理解地进行开放模型投资是有用且必要的,并且是他们可以掌控的事情。现在,这种影响力正在转向中国,我认为无论是在模型方面,还是在研究进行和分享的方式上,我们都无法知道未来会怎样,但我想猜测的是:你是否想处于能够收获其好处的位置?中国现在有一个非常有趣的开放模型和研究生态系统,我们不知道会从中产生什么,这只是技术演进中的一个未知数。美国拥有资源来掌控这一点,并且拥有如此优秀的学术机构,它们希望被更积极地激活和参与。
They're going to be the engine for the next 10 years of AI research because the academia has been necessarily I was gonna say beat down which is not the right way to say it but academia is just in a lull of influence in terms of the evolution of the science to the point where people will say academic AI research doesn't matter right now which I think it's a bit shortsighted because it's going to be an engine for exploration in a way that companies can't really nurture but open models are the platform by which that innovation is happening. And for a country like the US that has had such an excellent history of scientific projects and institutions if they want to be the institution of AI research they should consider it useful and imperative to have that open model investment be intentional and understood and just kind of something that they are in control of. Right now that influence is shifting to China where I think in both models and where research is done and shared it's just kind of a something that we cannot know what the future will hold but I would guess like is this a something you want to be in the position to reap the upsides of or China now has a very interesting ecosystem of open models and research and we don't know what will fall out of it it's just an unknown in terms of technological progression. The US has the resources to own this and also such great academic institutions that want to be more activated and more involved.
那么当 DeepSeek 出现,当中国模型开始一个接一个地推出开放模型时,这纯粹是地缘政治上的与美国竞争吗?中国行动背后的原因是什么?
So when DeepSeek happened, when Chinese models started open models started just taking off one after another, was it like pure geopolitical just in the competition with the US? What's the reason behind Chinese action?
我认为 DeepSeek 相当有理想主义色彩,他们以最纯粹的科学方式想要创造知识,并与世界分享大量知识和好处。他们在中国创造了一种行业标准,DeepSeek 是火花,让中国公司对参与 AI 产生了兴趣,他们看到可以通过开放模型来实现这一点。所以很多公司就把这当作默认的起步行为。如果你和这些公司交谈,他们也非常理性,他们知道美国的科技公司和潜在客户不会注册一个将数据发送到中国并付费的 API。他们告诉过我这一点。他们说,我们唯一的选择就是开放权重,因为这样他们仍然很可能使用它。他们也知道还有其他二阶担忧,比如 IT 部门会问,开放权重安全吗?但至少这是一张他们可以打的牌。公司意识到了这一点。我认为他们并没有比像 Llama 这样的模型更好地想清楚商业模式。在未来 1 到 3 年,我们将看到美国和中国开放模型的资金如何继续演变。我的直觉是,美国生态系统有更多的流动性来资助模型训练工作。但为什么中国的开放模型仍然如此接近前沿?我们真的不知道。Opus 和 GPT-5.2 显然比最好的开放模型更好,但来自中国的最好的开放模型可能超出了我的预期,它们确实非常好。
I think DeepSeek was fairly ideological in it where they were in the purest scientific way of wanting to create knowledge and share lots of it with the world and share lots of the upsides. They kind of created an industry standard in China where DeepSeek was the spark that made Chinese companies interested in participating in AI and they saw that you could do this through open models. So lots of them just consider that the default starting behavior. And if you talk to the companies, they're also extremely reasonable where they know that tech companies and potential customers in the US won't sign up for an API where data is sent to China and they pay. They've told me this. They're like, well, our only other option is open weights because then they could still most likely use it. And they know that there's other second order concerns where IT departments will be like, well, is the open weight safe? But it's at least a card that they can play. And the companies realize this. I think they don't have a business model figured out any better than something like Llama would have had a business model figured out. In the next 1 to 3 years, we'll see how funding continues to evolve for open models in the US and China. My hunch would be that the US ecosystem has a lot more liquidity to fund model training efforts. But then why are the Chinese open models so legitimately close to the frontier still? It's like we don't really know. Opus and GPT-5.2 are I would say clearly better than the best open models, but the best open models from China have probably exceeded my expectations and how legitimately good they are.
所以对于生态系统的长期轨迹以及模型能变得多好,有一些猜测,但今天也有很多真实信息表明,可能发生了一些奇怪的事情,我们不容易理清头绪。我刚刚和 Miniaax 的研究员 Olive Song 聊过,从对话中看,他们几乎每个月都在发布新东西——新版本——研究日夜不停,周末也不休息。人们轮班工作。我在美国没看到类似的情况,尤其是扎克伯格在 Llama 上退缩之后。达到那个水平的玩家是谁?DeepSeek、Miniaax、Qwen。
So there are guesses about the long-term trajectory of the ecosystem and how much better models will be, but there's also a lot of real information today that suggests weird things might be happening that we can't easily wrap our heads around. I just had a conversation with Miniaax researcher Olive Song, and from that conversation, it seems they are shipping something new almost every month—new versions—and research keeps going day and night, weekends, it doesn't matter. People work in shifts. I don't see anything like that happening in the US, especially since Zuckerberg backed off with Llama. Who are the players at that level? DeepSeek, Miniaax, Qwen.
我认为特点是,中国最顶尖的人才正在以这种水平从事开放模型的工作。那种氛围与 Anthropic、OpenAI 和 Gemini 现在所做的非常相似,但那里的人才密度肯定高于开放模型领域,后者通常与学术项目相关,比如 AI2,深受华盛顿大学和斯坦福 Percy Liang 工作的影响。这其中涉及不同的人才库。我认为美国最接近的是 Nvidia 的 Nemotron 项目,它在过去 6 到 12 个月里取得了很大进展。我的理解是,他们基本上解决了内部团队和文化对齐的问题,这使得他们能够发布更多高质量的内容,但还没有达到像 Qwen 或 Llama 那样的顶级突破。所以我认为他们方向是对的,但要进入 AI 的绝对前沿需要一些特别的东西。需要什么?Llama 非常成功。Qwen 显然也很成功。DeepSeek 也是。我不认为你可以仅仅靠蛮力做到,但 Nvidia 很接近了。
I think the characteristic is that the absolute best talent in China is working on open models at this level. That vibe is very similar to what Anthropic, OpenAI, and Gemini are doing right now, but the talent density there is definitely higher than in open models, which are often linked to academic projects like AI2, heavily influenced by UW and Percy Liang's efforts at Stanford. Some of this involves different talent pools. I would say the closest thing the US has is Nvidia's Nemotron efforts, which I think have made a lot of strides in the last 6 to 12 months. My read is that they've kind of figured out internal team and culture alignment, which has let them put out a lot more in higher quality, but it hasn't quite had the top-end breakthrough that something like Qwen or Llama has. So I think they're going in the right direction, but breaking into the absolute cutting edge of AI takes something special. What does it take? Llama was so successful. Qwen is obviously successful. DeepSeek is. I don't think you can just brute force that, but Nvidia is close.
他们背后也有商业模式,因为他们这样做会让人们以后使用他们的硬件和软件。所以这对 Nvidia 来说完全合理。没有其他玩家像他们一样有这样的机会。
They also have a business model behind it because they make it so that people will use their hardware and software later. So that makes total sense for Nvidia. There are no other players who have this opportunity like them.
是的。我采访过一位负责 Nemotron 项目的副总裁,那是他们的开放模型。我问他为什么这么做,他说:‘因为我们处于语言建模研究的前沿,而 Nvidia 会卖出更多 GPU。’我想,说得太对了。至少他们比开放模型领域的其他任何人都拥有更清晰的商业模式。因此,我对它的持久性持乐观态度。
Yeah. I did an interview with one of the VPs who leads the Nemotron effort, which is their open models. I asked him why they do this, and he said, 'Because we're at the frontier of language modeling research, and Nvidia is going to sell more GPUs.' I thought, damn straight. At least they have a much clearer business model than anybody else does in open models. So for that reason, I'm optimistic about its longevity.
目前开放生态系统中所有这些新模型在研究方面真正发生了什么变化?对你来说最有趣的是什么?
What is the real shift happening in the open ecosystem currently with all these new models in terms of research? What is the most interesting thing for you?
人们正在试图找出正确的方法来制作极具吸引力的模型和配方,比如工具使用和智能体。感觉是,封闭实验室和前沿模型在所谓的训练环境上投入了大量资金,他们可以在许多不同领域进行后训练。其中一些可能没用,很多是富有成效的。开放生态系统在创建训练系统方面还处于早期阶段。我认为 Prime Intellect 的验证器是一个例子。还有其他一些,但都太早期了,还不清楚如何以开放的方式将所有这些应用到学术界的全面训练中。我正在试图弄清楚,对于开放模型和开放专用模型来说,研究意味着什么,以便在 6 到 9 个月内更好地在编码智能体设置中进行协作。所以,如果你运营一个实验室,可以训练一些较小的模型并在任务上逐步提升,但不会与 Claude、Qwen 或那些更成熟的富裕机构竞争,那么显然会有一个多智能体的未来,多个模型可以一起使用。学术工作的绝对顶尖部分仍然可以利用实际使用的东西,但生态系统并不是那样运作的。如果我使用 Claude,我不会卸载到一个可以读取我所有文件的本地模型。那个未来正在到来,而且有很多开源编码智能体,比如 OpenCode 之类的,它们正在尝试如何围绕开放模型制作编码智能体。所以,在可能最具活力和兴奋点的领域,我仍然认为将是这种工具使用的后训练。希望跨环境的泛化是令人兴奋的,因为这是通往前沿模型所做工作的道路,但今年前沿模型有可能与学术和小规模开放模型更加分离。所有前沿实验室的算力支出预计会随着更多算力上线而不断增加。后训练是开放最难的部分吗?
People are trying to figure out the right ways to make models and recipes that are extremely compelling, like tool use and agents. The sense is that the closed labs, the frontier models, have invested so much in these so-called training environments where they can do this post-training in so many different domains. Some of them are probably useless, many are fruitful. The open ecosystem is in the early days of creating systems for training there. I think Prime Intellect's verifiers is one. There are others out there where it's just so early and unclear on the open way to apply all of these into comprehensive training runs in academia. I'm trying to figure out what research means for open models and open specialized models to be more cooperative in a coding agent setup in 6 to 9 months. So, if you're running a lab that can train some smaller models and hill-climb on tasks, but isn't going to compete with the likes of Claude or Qwen or these much more established rich agencies, there's clearly going to be a multi-agent future where multiple models can be used. The absolute top end of academic work could still tap into things that are actually used, but the ecosystem doesn't really operate that way. If I'm using Claude, I don't offload to a local model that can read all my files. There's a future there that is coming, and there are a lot of open-source coding agents like OpenCode and stuff like that where they are trying to figure out how to make coding agents around open models. So in terms of the area where there will probably be the most dynamism and excitement, I still think it's going to be this kind of tool use post-training. Hopefully, generalization across many environments is exciting because it's the path towards what they're doing at the frontier, but there's a chance that the frontier becomes even more separated from academic and small-scale open models this year. Compute spend is forecasted to go up and up for all these frontier labs with more compute coming online. Is post-training the hardest part to make open?
后训练是开放最难的部分吗?
Is post-training the hardest part to make open?
它们都有不同的挑战。我认为预训练数据是开放中最难的法律部分,因为你想要互联网上所有的语料库和人类知识。显然,其中一些如果公开出来,历史上是相当有诉讼风险的。我认为后训练在前沿模型中往往相当复杂,涉及很多模型和很多排序,以及将什么组合成最终配方的艰难决策过程。随着 RLVR 和 Scaling 强化学习革命,后训练基础设施的复杂性大大增加。归根结底,很多看起来像预训练,而预训练的开源基础设施,比如 Nvidia 的 Megatron LM,实际上非常强大。所以一些筹集了数亿美元来训练模型的实验室,他们直接使用这个开放的 Nvidia 软件并让它工作。我认为强化学习正处于许多库的时代,但随着时间的推移,它会浓缩成少数几个实际运行良好的库。但与此同时,拿一个库并做得很好——在后训练的前沿进行很好的训练——是相当困难的。后训练数据可能更落后,但训练的任何阶段都没有那么多开放数据集。我认为高质量发布的数据集如今非常罕见。
They all have different challenges. I think pre-training data is the hardest legal part to get open because you want every corpus on the internet and human knowledge. Obviously, some of those have historically been fairly litigious if you put them out openly. I think post-training tends to be fairly complex at the frontier with a lot of models and a lot of sequencing, hard decision-making process of what you put together into the final recipe. There was a big increase in the complexity of infrastructure for post-training with this RLVR and scaling reinforcement learning revolution. At the end of the day, a lot of it seems like pre-training where the open-source infrastructure for pre-training with things like Megatron LM from Nvidia are actually very strong. So some of these labs that are raising hundreds of millions of dollars to train models, they just take this open Nvidia software and make it work. I think reinforcement learning is in the era of many libraries, but over the years it'll distill down to a few libraries that actually work fairly well. But in the meantime, the complexity of taking a library and doing it very well—training very well at the cutting edge of post-training—is pretty hard. Post-training data is potentially further behind, but it's not like there's that much open data sets anywhere in the spectrum of training. I think high-quality release data is just super rare these days.
你在 AI for AI 研究所做这方面的工作吗?
Do you work on this at the Institute of AI for AI?
是的,我们涉及所有这些事情。例如,我们在秋天发布了 Model 3,目前要做的事情之一就是尝试将其过渡到更具智能体性和更多工具使用。工作流程是,你找到现有的开放数据集,查看你想要改进的评估指标,然后尝试它们。
Yeah, we touch on all of these things. I think for example, we released Model 3 in the fall, and one of the current things is trying to transition this to be more agentic and more tool use. And what the workflow looks like is that you find the existing open datasets, you look at the evaluations you want to improve on, and you try them.
我预计在某些评估中,你可以查看人工分析并浏览列表。有些我们会找到公开数据,可以轻松地在上面进行爬山优化。而其他一些在闭源实验室中非常流行的东西,我们可能需要花费数百万美元购买数据。我预计会有差距,因为很少有人发布数据。如果学术界恰好制作了我们需要的所有数据集,来尝试做前沿模型正在做的事情,那会显得过于巧合,因为前沿模型的策略是购买大量数据,至少让飞轮转起来,对吧?所以我认为存在天然障碍,这也是为什么开源模型最终会自然落后于闭源模型。一旦闭源模型擅长某件事,用这些模型创建训练数据就更容易,并将人力投入到差距中。这是一种不断演变的舞蹈,在我看来,人们喜欢过度炒作开源模型,认为它们会超越闭源模型,但我认为平衡将继续保持,最好的开源模型比最好的闭源模型落后大约 6 到 9 个月,这没问题。考虑到现在最好的闭源模型有多好,这个时间线相当短。但我看不到这种动态会改变。如果有什么不同的话,可能闭源模型会领先更多。但即使 9 个月的差距也很疯狂。我们不认为它们会赶上。
And I expect that for some evaluations, you can look at artificial analysis and go through the list. Some of them we will find open data that makes it fairly easy to hill climb on them. And other things that might be super popular to close labs, we would potentially have to buy data for like order of millions of dollars. And I just expect there to be gaps because so few people release the data. It would seem oddly convenient if academics happened to make all the data sets we need to try to do the things that the frontier models are doing when you know the frontier model playbook is to buy a lot of this data at least to get the flywheel going, right? So I think there's natural barriers, but the same thing why open models end up kind of having a natural lag of closed models. It's like once the closed models are good at it, it's a lot easier to create training data with those models and put human effort into the gaps. That's kind of the ever evolving dance where people like to overhype open models in my opinion as like oh they're going to cross closed models and I just think the equilibrium is going to continue where the best open models are some 6 to 9 months behind the best closed models and that's fine. That's a pretty short timeline with how good the best closed models are now. But I don't see the dynamic changing. If anything it might air slightly on the side of the closed models being more ahead. But even 9 months of a gap is crazy. We don't think they will catch up.
没有理由认为开源模型会赶上,因为它们资源更少,而资源通常决定结果。资源和人才决定结果。可以说,在中国实验室,人才比例与 OpenAI 和 Anthropic 类似,但在算力和购买数据的能力方面,资源要低得多。归根结底,人们需要的主要是算力来改进模型,因为算力要么用于训练,要么用于生成大量合成数据。我认为,例如在模式 3 中,这并没有非常清楚地记录,但我们在合成数据上花费了数百万美元。其中很大一部分是通过联邦政府拨款,用于美国的前沿超级计算机生成合成数据。但即使 AI2 在合成数据和训练的有效算力上花费了数十亿美元,你可以猜测这些前沿实验室的算力支出几乎达到数十亿。这些是巨大的持续成本,西方公司更有资本去做,而这些目前是闭源实验室。
There's no reason to think that the open models will like they have fewer resources and resources normally determine the outcome. It's like resources and talent determine the outcome. Arguably in Chinese labs, the talent is proportionally similar to the likes of OpenAI and Anthropic, but the resources in terms of compute and ability to buy data is just so much lower. At the end of the day, it's mostly compute that people need to make improvements to the model because compute is spent either on training or generating large amounts of synthetic data. I think for example in mode 3, this isn't super clearly documented, but we spent millions of dollars on synthetic data. A lot of it was through a grant from the federal government for this frontier supercomputer in the US to generate synthetic data. But like if even AI2 is spending billions of dollars on effective compute for synthetic data and training, you're going to guess that the compute expenditure there is almost like billions at these frontier labs. Like these just are huge costs constantly that the western companies are way more capitalized to do and these are closed labs right now.
有点不是无限的。如果它们只落后 6 到 8 个月,而且如果企业和个人已经可以使用,我们可以想象 6 个月后的情况,你将能够在日常生活中使用开源模型,就像使用 ChatGPT 或 Claude 一样,你不觉得吗?
Kind of not infinite. If they behind only 6 to 8 months and if the usage is already possible by like businesses and individual people, we can imagine the situation in 6 months that you will be able to use open model just for your daily life as you use I don't know like ChatGPT or Claude, don't you think so?
我认为在聊天机器人方面,开源模型肯定会达到那个水平。但在鲁棒性和一些工具使用方面,我还没有看到开源模型在搜索上像 GPT-4.5 或 Claude 3.5 那样好。但在聊天界面中,我认为它会达到。编码智能体可能不会,但这取决于赌注。我打赌闭源模型对这些编码智能体的兴趣还处于早期。很多情况下,研究人员会对某个领域产生兴趣,然后开始改进模型。如果 Claude Code 之类的东西去年四月发布,而采用并非由训练团队主导。我和这些公司的人聊过,他们说:“哇,是的,Claude Code 搭配 Opus 4.5 是我几个月前终于开始使用编码智能体的原因。”如果他们仍然刚刚开始痴迷于此,那么他们会决定,哦,我们需要训练一个非常擅长这个的模型。所以,仍然有追赶的时间,他们开始转动那个曲柄,我认为编码智能体会变得更好。这取决于赌注。你可以说:“是的,这些成本,收益不会与成本成正比,开源模型会赶上。”但似乎没有明确迹象表明模型会真正遇到瓶颈。似乎我交谈过的每个人都认为有很多低垂的果实,这是复杂的技术工作,涉及研究、执行和成本,我们需要继续转动曲柄。我认为显然在宏观经济意义上,公司有时间压力来展示其价值。但似乎这个 Claude Code 时刻至少为他们赢得了更多时间。
It's a bit of a bet I think on the chatbot side open models will definitely be there. I think there's some robustness and some tool use stuff where I haven't seen an open model that's quite as good at search as like GPT-4.5 or Claude 3.5. But in this chat interface, I think it will be there. The coding agents potentially not, but it comes down to a bet. And I would say that I bet on the side that the closed models are so early in their interest in these coding agents. A lot of the way this happens is that the researchers become interested in a domain in a way of using the models and then they take on improving them. If Claude Code and the likes came out last April and adoption is not front led by the training teams. I've talked to people at these companies and they're like, "Wow, yeah, Claude Code with Opus 4.5 is when I finally started using coding agents a couple months ago." And if they're like still just getting obsessed with them, then they will decide, oh, we need to train the model that's really really good at this. So, there's still this catch-up time where they're going to start turning that crank and I think the coding agents will get much better. It comes down to a bet. You could say, "Yes, these costs are the upside is not going to be proportionate to the cost and open models will catch up." It just seems like there is no clear indication that the models will actually hit a wall. It seems like everybody I talked to is like there's a lot of low hanging fruit, it's complex technical work in terms of research and execution and cost and we need to keep turning the cranks. I think obviously there's the macroeconomic sense where there's a time gating on the companies to show the value of it. But it seems like this Claude Code moment is going to at least buys them a lot more time to spend.
你看到了哪些低垂的果实?
What low hanging fruits do you see?
几乎到处都是。很多都是不同形式的:我们加入了这个数据集,它帮助很大。我们如何让它扩大 10 倍?或者我们加入了这个数据集,它帮助很大,但我们还没有很好地过滤它。所以让我们更多地过滤。或者我们的训练代码只使用了 60%的 GPU 利用率。所以我们可以通过编写更好的内核使其快 10%,因此我们所有的实验都快 10%。如果你这样做 40 次,你最终会得到一个快 4 倍的代码库,你所有的实验都快得多,你可以做更复杂的想法和事情。所以几乎所有这些事情,从最成熟的——我们拥有预训练数据集,我们已经过滤和迭代多年,他们仍在调整它,以更好地服务于当前感兴趣的任务,提高效率——到这些前沿的事情:我们刚刚花了 4000 万美元为编码智能体构建了三个训练环境,因为我们有所有这些新用户和编解码器,我们在第一次运行时直接插入,一些数字上升了。我们如何选择下一步做什么?或者复杂的事情,比如我们试图让模型更大,但数值问题太难了,比如我们如何提出一个新的强化学习算法来更好地处理这些数值。所以我认为不断有这些小问题,即什么更适合当前的情况。很多开源模型正在转向这些混合架构,混合了线性注意力和传统注意力。我认为部分原因是它们在数值上更复杂,但下游强化学习和推理的好处如此之大,你可以在强化学习上节省很多,这就像是整个行业共同推动下一个更难的事情。似乎 Qwen 3 有一个线性模型,Kimi 有一个线性模型,我们也在研究线性模型。人们在 GitHub 上抓取时发现了这一点。
It's literally like anywhere. A lot of it becomes like different flavors of we put this data set in and it helped a lot. How do we make it 10x bigger? Or we put this data set in, it helped a lot and we didn't really filter it very well yet. So let's filter it more. Or our code for training only uses 60% of the GPU utilization. So we can make it 10% faster by writing some better kernels and therefore all of our experiments are 10% faster. And if you do that 40 times, you end up with like a 4x faster codebase and all of your experiments are way faster and you can do more complicated ideas and things. So it's pretty much all of these things from the most established which is we have our pre-training data set that we've been filtering and iterating on for years and they're still tweaking it in order to better serve the current tasks of interest in order to be more efficient to these cutting edge things of we just spent $40 million on three training environments for coding agents because we have all these new users and Codex and we just plugged it in on our first run and some numbers go up. How do we pick the next thing to do there? Or complicated things like we try to just make the model bigger and the numerical issues are too hard like how do we come up with a new RL algorithm that handles this numeric better. So I just think that there's constantly these little problems of what is better suited to the situation at hand. It's like a lot of the open models are switching to these hybrid architectures which is a mix of this linear attention and like the traditional attention. And part of why I think that's the case is they're just more complex numerically, but also the upside on downstream RL and inference is so high where you can save so much on RL where it's just kind of like the industry collectively pushing through the next harder thing to do. It seems like Qwen 3 has a linear model, Kimi has a linear model, like we're working on a linear model. People found it scraping GitHub.
为什么会这样?这就像是基础设施和想法的集体准备,先在较小规模上测试,然后才奏效。我问了 AI2 项目的负责人:你认为未来所有模型都会变成混合模型吗?为什么两年前 Mamba 热潮正盛时没有发生?结果发现 Mamba 有几个设置需要搞清楚。Mamba 模型在预训练基准上表现很好,但实际文本生成效果不佳。架构中只有几处需要调整,更好地平衡传统 Transformer 风格的架构。经过一些调整,模型稳定多了,人们也开始尝试使用它们。这是一个超过两年半的例子,我记得 Mamba 和状态空间模型的热潮很高,但现在它似乎真的在很多地方落地了。Nvidia 也有一个混合注意力模型。你怎么预测那个时间线?我不知道。但我认为肯定还有很多类似的事情会从 AI 研究演变为现实。大问题是:Sam Altman 在融资引擎上会做什么?我不认为那会是一个泡沫破裂的事件,但 Nvidia 股票或其他方面可能会有调整。有传言说 OpenAI 和 Anthropic 年底前会 IPO。我不会太惊讶。
Why does that happen? It's like a collective readiness of infrastructure and ideas being tested at smaller scales so that they work. I asked the lead of the project at AI2: do you think all models will be hybrid in the future, and why didn't this happen two years ago when the hype for Mamba was really high? It turned out a few settings on Mamba needed to be figured out. Mamba models did really well on pre-training benchmarks, but the actual text generation from them wasn't nice. There were just a few things in the architecture that needed to be changed, balancing them better with traditional Transformer style architectures. With a bit more tinkering, the models are a lot more stable and people are trying to use them. That's an example of over two and a half years where I remember when the hype for Mamba and state space models was so high, but now it seems to be really hitting a lot of places. Nvidia had a model with hybrid attention as well. How do you predict that timeline? I don't know. But I think there are definitely plenty of things like that that will continue to evolve from AI research to reality. The big question is: what does Sam Altman do on the fundraising engine? I don't think that's going to be a bubble popping thing, but there could be corrections on Nvidia stock or other things. People are rumoring OpenAI and Anthropic IPOs by the end of the year. I wouldn't be that surprised.
是的。你怎么看 SpaceX 和他们的热情?马斯克工业,无论好坏,都有点反派气质,但像是在科幻电影里。我不够商人,无法评论实际发生了什么。我猜除了贪婪和弥补 X 或 XI 的损失之外,还有别的原因。无论你对埃隆有什么看法,尤其是在我认为他应该少花时间的政治方面,他在建立企业方面有如此出色的记录,我不得不认为那里有一个计划。最近特斯拉关于停止 Model S 和 Model X 的评论让我相当震惊。特别是因为机器人技术的时间线似乎太早了。我的直觉是,现在扩大机器人技术规模有点太早,但当像埃隆这样的人下如此大的赌注时,一定有我不知道的原因和事情。作为一个对 SpaceX、特斯拉和大规模机器人制造相当无知的人,我对此不了解,所以走着瞧吧。
Yeah. What do you think about SpaceX and their excitement? Musk Industries, for better or worse, has a bit of villain vibes, but in a sci-fi movie. I'm not enough of a businessman to comment on what is actually happening. I would guess there's a reason other than just greed and recouping losses from X or XI. Whatever opinion on Elon you have, especially on the political side that I think he should spend less time on, he has such a track record in building businesses that I have to think there is a plan there. The recent Tesla comments of stopping the Model S and Model X is fairly shocking to me. Particularly because the timeline on robotics seems too soon. My intuition is that scaling robotics now is a bit too soon, but when there are that serious of bets done by somebody like Elon, there have to be reasons and things that I don't know. As a fairly ignorant person about SpaceX and Tesla and large-scale robotic manufacturing, I'm not informed here, so we'll see.
你可以给埃隆·马斯克的计划加上几年,但最终他会达到目标。这是他的记录。是的。你从研究角度描述的一切主要是关于 Transformer 和偶尔的混合模型。你看到有严肃的研究进入其他领域吗?
You can add a couple of years to what Elon Musk projects, but eventually he will reach the goal. That's the track record. Yeah. Everything you described from this research point is mostly with transformers and with occasional hybrid models. Do you see serious research going into some other areas?
不管出于什么原因,总有这种持续学习的炒作点。我有点认为,在编码智能体中,我们有这些 cloud MD 文件和 agents MD,它们实际上很擅长从中学习。短期内这会有很大作用,但显然还有更大的收益。我认为持续学习最好被视为一个研究问题的例子,它最终会被解决,并且动机很好,但你永远不知道真正的解决方案何时会达到规模。这个想法是,模型权重应该根据你个人或模型在世界的经验而改变。当训练如此昂贵,模型在某些情况下显得如此愚蠢时,这是一个合理的动机。它几乎必须是真的。我认为 20 年后我们还在使用类似 Transformer 的东西的可能性并不高。我们会有混合模型,再经过几次转变。人们还会叫它 Transformer 吗?我不知道。注意力机制似乎很好。只是很难预测研究的演变。
For whatever reason, there's this kind of continual learning hype point. I'm kind of of the bucket that some sort of like in the coding agents we have these cloud MD files and agents MD and they're actually fairly good at learning from them. In the short term that'll do a lot, but obviously there's such a big gain in something. I think continual learning is best thought of as an example of a research problem that will eventually be solved and is wonderfully motivated, but you'll never know when a real solution hits the scale solutions. The idea is that the model weights should change based on experience personal to you or the model in the world. That is so reasonable when training is so expensive and it can seem like the models are so dumb in some situations as a motivation. It just kind of has to be true. I think the likelihood that we're on something that looks like a transformer in 20 years is not that high. We have the hybrid model thing a couple more transformations down the line. Are people going to still call it a transformer? I don't know. Seems like attention is pretty good. It's just so hard to predict the research evolution.
我问这个问题是因为在 Lex Fridman 的播客中,你说如果你现在开始,你可能不会做 Transformer。这是针对学术界的吗?
I'm asking this question because in the Lex Fridman podcast you said if you were starting you probably would not do transformers now. Is this for an academic?
是的,最好的研究更远。我认为深度学习和 Transformer 是 CS 研究人员的基本技能。当我刚开始时,你只需要了解深度学习反向传播库如何工作。现在人们需要了解 Transformer 如何工作。但如果你做的工作只是明显改进已有的东西,那么在学术上很难推销,除非你处于学术界的顶尖,他们有更多资源和行业联系来真正保持在前沿。华盛顿大学、斯坦福和伯克利的一些实验室这样做,但大多数学术界不是这样。所以你需要一些不那么拥挤或需要更长时间的东西。另一件可以描述的事情是,在两个流行的思维方式或子领域之间找到空间,或者去没有人的地方。因为有些东西本质上在 AI 中很可能会成功。如果你优化得当,成功的东西数量真的很高。我不知道我是否喜欢自己的建议,因为说起来容易做起来难。比如不要做 Transformer。那好吧,我该做什么?我不知道。我的意思是,你在强化学习还不太流行的时候做了它。
Yeah, the best research is further out. I think of deep learning and transformers as kind of fundamental skills that you have as a CS researcher. When I was starting, it was just like okay, you just need to learn about how deep learning backprop libraries work. Now people need to learn about how transformers work. But if you're doing work that is so blatantly just improve on what we already have, it's going to be kind of hard to market it academically, unless you're at the absolute top end of academia where they have more resources and industry connections to really stay at this frontier. Some of the labs at UW, Stanford, and Berkeley do this, but most of academia is not like this. So you need something that is less busy or takes longer. Another thing that could be described is find space in between two popular ways of thinking or two subfields, or just go where there are not people. Because there are some things that by their nature are very likely to work in AI. The amount of things that just kind of work is really really high if you set the optimization right. I don't know if I love my own advice because it's easy to say go off into somewhere, but it's also hard to make it work. Like don't do transformers. It's like okay, what do I do? I don't know. I mean you did reinforcement learning when it was kind of a dump.
是的。我想我们也谈到了机器人技术和从事机器人技术。我看到很多人很开心。它更接地气,但你能从语言模型革命中积累好处,尽管有点间接。所以你可能不会受到剧烈波动的影响,但你能获得新事物成功的好处,以及在特定领域尝试新事物的能力。如果我们更多地谈论开放模型的实际实现,首先,如果你不介意,我们从定义什么是开放模型开源 AI 开始。我甚至不知道这重不重要。你认为定义重要吗?
Yeah. I think we also talked about robotics and working in robotics. I've seen many people be so happy. It's a bit more grounded, but you accumulate the benefits that are happening in this language model revolution while kind of being a bit secondhand to it. So you're probably not subject to the whiplash, but you're getting the upside of new types of things working and the ability to try new things in a specific domain. If we talk more about practical implementation of open models, first of all, if you don't mind, we start from kind of defining what open model open source AI is. I don't even know if it matters. Do you think definition matters?
定义在经典意义上不重要,即把东西归入它的争论。重要的是社区定义是什么。社区对开源的定义是:权重被公开释放,任何人都可以使用它们。
The definition doesn't matter in the classical sense of the debate of assigning things to it. It matters in what the community definition is. The community definition of open source is really this: weights are released openly so that anyone can use them.
我认为这与许可讨论相呼应,中国实验室做得比较好的事情之一就是它们都发布了极其宽松许可的模型,而许多美国公司则附加了非商业用途或额外条款。在企业环境中,律师不喜欢模糊的附加条款。许多编写这些定制 AI 许可的人会使用模糊的语言,带来风险暴露。因此,拥有简单明确的条款,让人们清楚可以使用,这很好。过去有很多关于开源需要代码和数据可用的争论,是的,这些是宝贵的资源,更符合开源软件的动机,即自由、可复制、可修改等。但当人们实际使用开放权重模型时,争论就平息了。这就是现实,也是人们在 AI 领域更开放地运作的范式。所以我通常觉得“哦,这没问题。”我很高兴不再陷入“这是否开源”的愚蠢争论,那并不有趣。
And I think that chimes into the license discussion where it's one of the better things these Chinese labs have done is they're all just releasing their models with extremely permissive licenses where a lot of US companies have done things where it's like non-commercial or extra terms and conditions attached. And in enterprise situations, lawyers do not like vague extra terms attached. And a lot of people that end up writing these custom AI licenses do it in ways that there's vague language and exposure of risk. So just having simple terms where it's clear that people can use them is great. And in the past there's been a lot of debate on like open source you need the code and the data available which is like yes these are valuable resources and yes that fits more with the open source software motivations of like things are free and reproducible and modifiable and so on but it's just like the debate is kind of died down when people actually just use open weight models. It's like this is the thing and this is the paradigm by which people are operating in something that's more open with AI. So I generally am like, "Oh, it's fine." I'm happy to not be in stupid debates on like, "Oh, is this open source or not?" It's just like that was not fun.
而且它仍然几乎是完全透明、完全开源的。
And still, it's still almost fully transparent, fully open source.
有一个群体,比如 AI2、Percy Lang、Marin Thing、Hugging Phase、LLM 360,它们也与 MZUBA 有关联,就像 UAA 项目一样。
There's a group. It's like AI2, Percy Lang, Marin Thing, Hugging Phase, LLM 360, which is also affiliated with the like MZUBA. It's like the UAA thing.
是的。瑞士有一个非常开放的项目。所以实际上有更多人,尤其是在所谓的西方生态系统中,在做这种完全开放的事情,我认为这是对科学标准和科学进步的良好认可,但这仍然是小众的。完全开放并没有带来显著的更多采用。如果某家公司说“哦,这些是完全开放的,所以我可以更快迭代等等”,那会很好。但在实际应用中,使用更好的模型比模型完全开放带来的好处大得多。
Yeah. The Swiss had a project that was very open. So there are actually more people especially in the like so-called western ecosystem that are doing this fully open thing which I think is a very good appreciation for scientific standards and like scientific progress but it's still niche. It's not the fact that these are fully open is not causing dramatically more uptake. It would be nice if some company was like, "Oh, these are fully open, so I can iterate faster and so on." But just using a better model is way more benefit than having the model be fully open in terms of real world applications.
我打赌 2026 年将是人们开始,企业也会开始更多地使用开放模型的一年。你怎么看?
My bet is that 2026 will be the year when people would start and like enterprises also will start open models will start using open models much more. How do you see it?
我认为它们已经在进行中。许多公司的默认立场是希望使用开放模型,出于信息安全、更可预测的成本和拥有自己的技术栈,但公司需要很长时间才能转向这些新技术。我怀疑这将与所有调查结果类似,比如公司是否从使用 AI 工具进行协作中受益,这会在很长一段时间内充满噪音,但随后会发现许多公司正在使用开放模型并针对其用例进行微调,从而创造大量价值。这需要很长时间,尤其是如果涉及任何训练,你不能直接拿一个现成的模型,因为需要彻底测试模型。实际进行训练需要数月的时间投入,并且需要一支专注的团队。
I think they are I think it's been ongoing. A lot of companies default position is they want to use open models for information security, more predictable costs, owning their stack and it takes a long time for companies to move into these new types of technology. I suspect it'll look like all the same things where there's all these surveys of like companies are they getting benefit out of using AI tools for co-working and it'll be noisy for a very long time but then it'll be seen that a lot of companies are using open models and fine-tuning them for their use cases in ways that creates a lot of value. It's just going to take a long time, especially if any training is involved where you just can't take a model off the shelf because it's enough to like thoroughly test a model. To have to actually do the training yourselves is months long of a commitment and you need to have a team of dedicated people.
我开始了解开源 AI 的经济学,比如如果你是一家公司,想使用开放模型,有时成本高得令人望而却步。所以免费的开源模型并不便宜。
I started to learn about economics of open source AI like if you a company and if you want to use open model, it's like sometimes prohibitively expensive. So it's like free open-source model not cheap.
是的。因为你必须投入一定量的算力才能以有意义的方式提供模型服务,而且你很容易购买最低限度的算力来托管该模型,但推理负载远不足以支持这种投入。这就是为什么一些分离式推理公司有意义,它们会获得足够的负载,使某些功能值得实现,但这需要时间来建立。
Yeah. Because you have to commit to a certain amount of compute to serve a model at a meaningful way and like you could very easily buy the minimum compute to host said model and get nowhere near the inference load to frankly like support that spin. That's why some of the like disaggregated inference companies make sense where it's just like they will get enough load to make certain features worthwhile and stuff like this that it just takes time to build that up.
AI2 是如何组织的?你们在研究所使用开放模型吗?
How is it organized in AI2? Do you guys use open models in the institute?
有些人用,但不是很多。我认为 AI2 要让模型迈出下一步,核心是必须自己先使用模型,建立使用模型和改进模型之间的开发循环,并获得真实反馈。在此之前,这有点像玩具项目,虽然有用户,也有部署,但如果你自己都不碰它,怎么能期望别人认真对待呢?
Some people do but not a lot lot. I think the like core thing that AI2 needs to get to for the models to take the next step is to like actually you have to dog food the models and build development loops between using the model and improving the model and getting real feedback because until then it's somewhat of a toy project where there are users and there are things that is like deployed into but if you're not touching it like why do you expect other people to do so with serious intent.
一次只做一件事。
One thing at a time.
你积极使用开放模型吗?我主要使用封闭模型和 API。我会在发布时尝试它们,但它们不够好或不够便宜,不足以让我明显想改变习惯。这在很大程度上是科技中经典的产品模式,用户有很强的习惯。我看到在 ChatGPT 和 Claude 之间,某些用途的缓慢演变。这些习惯需要很长时间才能改变。尽管我认为 Claude 在正常对话中的语言比许多 GPT-5 的东西更简洁、更易接受,但 Claude 模型已经这样一段时间了。我花了几个月才重新调整习惯,这种演变在 AI 近期历史上发生了多次。我转向开放模型,感觉只是“还行”。从产品角度,我为什么要用这个?我认为和我一起在 Interconnects 研究开放模型的 Florian,他虔诚地使用一些开放模型,因为只有开放模型才能做到。他使用 Nvidia 的 Parakeet 模型进行语音转文本,然后作为更快的输入方式进入云代码。他会对 AI 说话,Parakeet 转录,然后 Qwen 改写并传递给 Claude。这是一个非常真实的使用案例,他非常推崇并从中获得很大价值。但我还没有这样做,我还没有大量使用它们。
Do you use actively open models so you close and APIs I mostly use closed bottles. They I try them especially around launches. They're not good enough or cheap enough where it's obvious that I want to switch my habits. And a lot of this is classic product patterns in tech where users have strong habits. And I see the like slow evolution back and forth between say like ChatGPT and Claude for certain uses. And it takes a long time for those habits to shift. Even though like I think Claude's language on normal conversation is a much much more succinct and palatable than a lot of the GPT5 things but like Claude models have been like that for a while. It's just taken me months to like reshift the habit and this evolution has happened multiple times in AI's recent history where I go from and it's just like I go to an open model and it's just like fine. It's like why would I in a product sense why would I use this? I think like Florian who works with me at interconnects to study open models has a few things he uses them religiously for because it's like only the open models we do this which is like he uses one of Nvidia's parakeet models which is speech to text and then as a much faster way to input comprehensively into like Claude Code. So he will speak to his AI and then parakeet will transcribe them and then Quen form B will rewrite it and pass it to Claude. That's a very real use case that he swears by and gets a lot of value out of. But I'm like, h I haven't done it yet. I haven't I haven't used these substantially.
是的,改变使用模型的习惯出奇地困难。我仍然在 5.2 上挣扎。我觉得他们改了什么东西,我不知道改了啥。不管我选什么角色,它都很糟糕。我仍然用它来思考,我仍然喜欢它的深度研究部分,但当我读它告诉我的内容时,它就是不遵守规则。我不知道,我不知道如何让它们遵守规则。
Yeah, it's surprisingly how hard it is to change a habit with a model. I still struggle with 5.2. I think they changed something. I don't know what they changed. No matter what character I choose, it sucks. I still use it for thinking. I still like deep research part of it, but when I read what it tells me, it just and it doesn't follow the rules. I don't know. I don't know how to make them follow the rules.
这很有趣,因为我觉得每个深入其中的人都有这样的看法,比如“这是什么问题?”这就是我所说的低垂的果实。人们只需要修复它。这实际上是可修复的,只是需要大量工作。
It's pretty funny because it's like I feel like everybody that's deep in it kind of has these opinions just like what is this problem? And that's what I mean by the low-hanging fruit. It's like people just need to fix that. That's actually fixable. It just takes a lot of work.
训练让模型真正记住固定的提示,比如我放在记忆中的相同提示。是的,这就像改变数据并创建新数据。
Training like to make the model actually remember the constant prompts like the same prompt that I put in the memory. Yeah, it's just like changing data and creating new data.
当你同时解决数千个用例时,要让那种潜在的 niche 问题得到体现,这个过程需要大量的平衡。你可能需要为此构建一个评估,以便自动衡量它,然后你就又多了一个评估。他们可能有数百个评估。有很多活动部件需要处理好。所以我认为他们在哪些事情真正是优先事项上相当保守,而且他们可能愿意在某些方面接受退步,比如内存管理,以推动效用的前沿。如果你要在目标领域迈出一大步,并且有一些后续效应需要以后重新审视,你可能会把它做对。现在的重点是编码,而在编码中你并不真正需要这种优美的语言。
The process for getting that type of potentially niche problem to be represented when you're solving thousands of use cases at once is like it takes a lot of balancing. You probably need to build an evaluation for it so that you can automatically measure it, and then you have one more evaluation. They probably have hundreds of evals. There's just a lot of moving pieces to get right. So I think they're fairly conservative in what things are really the priority, and they probably are willing to take regressions in some things like memory management to push the frontier in utility. If you're going to take a gigantic step in the things you were targeting and you have some second-order effects that you're going to revisit later, you're probably going to do it right. The focus now is on coding, and in coding you don't really need this beautiful language.
你说过你对模型的性格感兴趣。开源模型也是这样吗?还是更复杂?
You said you're interested in working on characters of the model. Is it the same with open models? Is it more complicated?
这基本上归结为数据工作,就像你希望所有数据看起来都非常相似。闭源模型在这方面很出名,比如 GPT-4o,人们仍然对它非常忠诚。而 Claude 的语言与 GPT-5.2 非常不同,而且在每一层都是有意为之。第一个干预是在后训练阶段,因为它最直接,但最终你需要对齐预训练,以确保它有助于实现这一点。这很大程度上是一种工业习惯,因为它非常适合模型即产品的理念。因此,模型的语言是用户看到信息的界面。所以模型的输出就是产品本身。如果产品是让用户保持参与并使用某些触发器让用户回来,而输出与用户体验的方法如此紧密地交织在一起,那么在学术界研究起来就困难得多。我认为这正在延续到编码智能体上,比如我在使用 Claude 时它呈现给我的东西是一个非常有趣的问题,因为 CLI 是一个如此稀疏的界面,但我与 Claude 的少数交互对最终结果影响巨大。这也非常像模型在那个框架中的输出与这个产品存在性地绑定在一起。我认为由于它对 AI 行业的重要性,我正在努力帮助围绕它发展学术研究,但我不期望它会像我在这个领域参与的其他项目那样火爆。例如,没有对奖励模型的评估,而奖励模型对 RLHF 和 AI 研究很重要。是的,那个火了,因为显然人们想做后训练研究,但我不确定性格训练会以同样的方式火爆,因为它不那么容易研究。但这符合我的目标,即让 AI 在前沿领域对每个人都可访问。就像你可以理解正在发生的事情,并对前沿的发展方向做出自己的判断,以便为你所想的任何事情做好准备。
It reduces to pretty much being data work, which is like you want all of your data to look very similar. Closed models are famous for this, where we have this GPT-4o thing where people still are fiercely loyal to it. And Claude's language is very different than GPT-5.2, and it's just very intentional at every layer of the stack. The first intervention is at post-training because it's most direct, but eventually you need to align pre-training to make sure that it helps achieve this. And it's largely an industrial habit because it fits so nicely into the model as the product. Therefore, the model's language is the interface by which the user sees the information. And so the model's outputs are so literally the product. If the product is what keeps the person engaged and uses certain triggers to get people to come back to it, and the output is so deeply entwined in how you would approach the user experience, it is much harder to study in academia. I think this is continuing with coding agents, where it's like what is Claude presenting to me as I use it is a very interesting problem because the CLI is such a sparse interface, but the few interactions I have with Claude are so impactful to the final outcome. That is also very much like literally the model's output in that harness is so existentially tied to this product. I think due to its importance to the AI industry, I'm trying to help grow academic research around it, but I don't expect it to take off like other projects I've worked on in this area. For example, there's no evaluation for reward models, and reward models are important for RLHF and AI research. Yeah, that blew up because obviously people want to do post-training research, but I'm not sure that character training will blow up in the same way because it's just not as easy to study. But it fits with my goals of making AI accessible to everybody at the frontier. It's just like you can understand what is happening and make your own bets about the direction of travel for the frontier to prepare whatever you think about.
你认为在 2026 和 2027 年有哪些方向值得关注?
What would you say are a few directions to follow for the next in 2026 and 2027?
我认为智能体这件事非常真实。所以它如何体现在多个方面?一是编码智能体的轨迹,二是这种行为如何出现在其他类型的智能体应用中,或者编码智能体扩展到更多领域,以及这对当今的模型和软件意味着什么等等。但这要困难得多。有传言说 Claude 5 今天发布。所以我在想它的故事会是什么。我记得去年回到 Claude 4 的时候,它们的发布在基准测试上非常低调。它们有不同于 OpenAI 和 Gemini 的基准测试。它们看起来不那么炫目。这与它们在代码上的押注相符,但也符合这样一种情况:AI 的评估更多地是关于一些技术人员所说的框架或产品,而不仅仅是模型。我认为 Anthropic 在这方面非常早,相比之下,Gemini 3 在年底发布,被炒作成“谷歌回来了”,但现在几乎没有真正前沿的影响,而且它们在发布时没有这种智能体体验,所有这些方面仍然落后。所以我认为发布之间那个灰色地带才是今天前沿所在。我预计 Gemini 会努力追赶,并解释 OpenAI 在 Codex 上的努力,但要跟上要困难得多,因为不仅仅是“数字上升,开放数字上升最多”。更多的是你需要去尝试,它可能会进入新的领域。因此,我认为关注 AI 比 ChatGPT 之后更混乱了,那时只是更智能的聊天机器人,转动曲柄,GPT-4 之类的。那要容易得多,但聊天领域已经饱和,很难看到模型之间的差异。
I think most this agentic thing is very real. So it's like how does this show up in multiple? One is what is the trajectory of the coding agents, and two how do behaviors like this show up in other types of agent applications, or coding agents expanding into more domains and what that means for models and software today and so on. But it's much harder. There were rumors about Claude 5 being released today. So I was thinking about what the story of it would be. I remember last year all the way back to Claude 4, their releases were very muted on benchmarks. They had different benchmarks than OpenAI and Gemini. They didn't look as flashy. And this fits with their bet on code, but also fits with this kind of thing where evaluations for AI are much more about what some technical people call the harness or the product than just the model. I think Anthropic was very early on this, where the comparison is Gemini 3 which came out late in the year and was so hyped as 'Google is back' but has almost zero really cutting-edge impact right now, and they kind of didn't have this agentic experience at launch, and all these things they're still lagging behind on. So I think that gray area in between the lines of what is happening with releases is where the frontier is felt today. And I expect Gemini to try to catch up, and explain OpenAI's efforts on codex, but it is just much harder to follow because it's not just like 'oh numbers go up, open numbers went up the most'. It's much more you need to try it, and it might emerge into new domains. So for that reason, I think following AI is a bit messier than it had been post-ChatGPT, where it's just like smarter chatbot, turn the crank, GPT-4 or whatever. It was much easier, but the chat domain is so saturated it's really hard to see the differences between the models.
你看到任何差异化即将到来吗?我们更多地谈论聊天机器人、机器人技术、智能体系统。还有其他大事吗?
Do you see any differentiation coming? And we talk more about chatbots, robotics, agentic systems. What are other big things?
我一时想不起来,我真的不知道。与开源模型相关的是,我认为主权 AI 将继续成为一个被使用的术语。是的,美国和中国显然在 AI 领域有立足之地,但会有大量持续的国际投资用于算力和模型,我认为这些会逐渐消失。
Off the top of my head, I feel like I don't really know. The thing that goes with open models is that I think sovereign AI will continue to be a term that is used. Yes, the US and China obviously have a foothold in AI, but there's going to be a lot of continued international investment in compute and models, and I think these kind of disappear.
这就是为什么法国对此非常满意。他们找到了一家愿意早期接受这种叙事的公司,这使他们在欧盟乃至世界其他地方处于非常有利的地位。我认为这将继续,但我不认为它在你谈论的其他事情规模上那么重要。重要的事情像是电力建设和其中的社会学因素,是的,GPU 本身不消耗大量水,但建设对社区来说非常个人化,建筑材料确实消耗大量水。我对这个问题了解不深,但这是目前美国如何看待 AI 的一个主要障碍,并且仍将是一个非常重要的故事。
That's why France is so happy with it. They found a company that was willing to accept that narrative early on, and that puts them in a really strong position in the EU but also elsewhere in the world. I think that will continue, but I don't think that's as important on the scale of the other things you're talking about. Important things are like power buildout and the sociological factors of this, where it's like yes GPUs themselves don't use a lot of water, but construction is deeply personal to communities and construction materials do use a lot of water. I'm not super well read on this issue, but that's like a major blocker in how AI is viewed in the US right now and will still be a very important story.
确实如此。确实如此。有趣的是,很多孩子谈论 AI 时,他们说的第一件事就是它消耗大量水,我甚至不明白这是从哪来的,因为事实并非如此。
That's true. That's true. It's funny enough a lot of kids when they talk about AI, that's the first thing they say is that it uses a lot of water, and I don't even understand where it comes from because that's not exactly how it works.
描述方式会演变并过去,但建设限制因素在美国是真实的,我认为这将继续成为一个大故事。我不认为它永远只是水的问题。我认为这有点像 meme 难以预测。
The way it is described will evolve and pass, but the construction limiting factor is real in the US, where I think that'll continue to be a big story. I don't think it'll always be just water. I think that's kind of like a almost like how memes are hard to predict.
这就像人们用来表达问题的语言往往难以预测。但水资源问题主要在于,这些最富有公司的建设方式往往并不支持环境和所在社区。即使大型数据中心创造的财富对科技公司的利润很高,这也是一个微观层面的问题,以过去未曾有过的方式阻碍着科技公司。这说得通:大量资金在流动,但却发生在世界上一个很小的、无法获得这些好处的角落。
It's like the language people use to express the problem is often hard to predict. But the water issue is mostly about how the buildout of these wealthiest companies happens in ways that are not often supportive of the environment and communities where it occurs. Even if the wealth creation from large data centers is high for tech company bottom lines, it's a micro-scale problem blocking tech companies in a way they haven't dealt with before. It makes sense: a lot of money is happening, but in a very small part of the world that doesn't get that benefit.
开源能否成为这种叙事的反面,不需要那么多资源?我看到的大部分资源都用于推理,有些推理可以在你的电脑上完成,但如果你想获得前沿性能,比如 Gemini 或 OpenAI 的图像生成,像编程智能体这样的东西短期内不会转移到你的电脑上。基于 Transformer 的图像生成或视频模型都远非廉价。如果 AI 世界要成功,大部分建设必须用于推理,因为推理才能带来回报。你可以在本地开源模型上获得部分能力,开源模型的好处是不必从头训练,但我认为这些不是决定性的转折点,不会让我们在 AI 的两种结果之间面临岔路。它们有贡献,可以累积,但不是核心问题。最可能的情况是混合模式,开源和闭源并存。
Can open source be the opposite of this narrative, not needing so many resources? Most resources I see are for inference, and some inference you can do on your computer, but things like coding agents won't be offloaded to your computer anytime soon if you want frontier performance, like Gemini or OpenAI image generation. Both transformer-based image gen or video models are far from cheap. Much of the buildout, if the AI world succeeds, has to go to inference because inference pays back. You can get a version of this on open models locally, and open models are good because you don't have to retrain from scratch, but I think these are not defining pivot points that put us at a fork in the road between two outcomes for AI. They contribute and pile up, but they're not central issues. Most likely it will be hybrid, open and closed together.
是的。我们会看到苹果能否整合起来。有些公司应该以更有吸引力、更易用、可扩展的方式弄清楚开源模型对消费者的意义,而目前公众可能还没有真正大规模使用开源模型的探索。应该有一个让开发者易于使用的技术栈。所以问题很多。目前开源 AI 中哪些方面是明确有效的?
Yeah. And we'll see if Apple pulls things together. There are companies like that which should figure out what open models mean to consumers in more compelling, easy-to-use, and scalable ways, where currently there's maybe no exploration for the general public in terms of real open model usage at scale. There should be a stack for developers to use easily. So there are a lot of problems. What is clearly working in open source AI right now?
嗯,我认为人们对拥有自己的技术栈有着无法满足的需求,即使在美国科技公司,比如银行和美国经济的大部门,尤其是在金融服务价值很高的金融化经济中。在国际上,也有很大的价值,因为新兴经济体传统上不会在付费互联网服务上花很多钱。所以开源模型成为他们获得 AI 的唯一途径,以适合他们当前情况的任何方式。我认为开源模型正走在与绝对性能前沿脱钩的道路上,前沿涉及昂贵的智能体和越来越贵的 AI 订阅。我为此付了很多钱。当这些模型可以进一步推进,人们付更多钱时,这不会发生在开源模型上,而那是 AI 的前沿所在。但还有其他的背景,生态系统如此专注于旧金山和前沿;世界上还有很多其他 AI 用途需要填补和发展。我认为开源模型会有很多未知数,这些事没有被密切追踪,但在世界范围内仍然非常有影响力。研究是缓慢的。研究和建立技术都是非常缓慢的趋势,这就是我认为开源模型如此重要的原因。科技行业快速是另一回事;时间线非常不同。
Well, I think there's insatiable demand for people to own their stack, even in US tech companies like banks and large sectors of the US economy, especially in a financialized economy where financial services are valuable. Internationally, there's also much value where new economies aren't traditionally spending a lot on paid internet services. So open models become their only access to AI in whatever capacity fits their situation. I see open models as being on a path disconnected from the absolute frontier of performance, which involves expensive agents with AI subscriptions getting even more expensive. I pay a lot for them. When these models can be pushed further and people pay even more, that's not happening on an open model, and that's where the frontier of AI is going. But there's this background of everything else where the ecosystem is so focused on SF and the cutting edge; there's still a lot of other AI use in the world to be filled in and evolved. I think there are many unknowns that will happen in open models, things not closely tracked but still very influential on a world scale. Research is slow. Research and establishing technologies are both very slow trends, and that's where I see open models as so important. The tech sector being fast is just a different story; timelines are so different.
所以开源模型应该只是一个实验领域。
So open models should be just a field for experimentation.
是的。而且这种实验更难看到,也更难追踪,但仅仅因为它慢,无论前沿是否成功,它仍在发生。前沿越成功,对开源模型的需求也会增加,基于某些无法使用闭源模型的用例。
Yes. And it's harder to see this experimentation and harder to track, but just because it's slow, it's still happening whether or not the frontier succeeds. The more the frontier succeeds, the more demand for open models will increase as well, based on certain use cases where they can't use closed models.
是的。但当你向人们介绍你的项目 Atom 时,你会告诉他们什么?为什么它很重要?
Yeah. But when you go to people to tell them about your project Atom, what do you tell them? Why is it important?
拥有未来几十年 AI 创新的引擎,并成为 AI 研究的核心影响力来源。你可以看看 AI 研究周围的媒体生态,即使只是与前沿相关联也显然非常有价值。我们需要强大的开源模型,这样从美国研究到新初创公司和建立在研究之上的成熟公司的管道就不会有摩擦,比如‘哦,但那是中国模型,我能用吗?’与正在进行的建设相比,它并不昂贵。拥有这项创新对美国来说是值得的,因为它能让我们更好地理解开源模型的使用方式,因为人们会来找模型构建者说‘为这个做得更好’等等。
Owning the engine for AI innovation for decades to come and being the central source of influence in AI research. You can look at the media ecosystem around AI research where it's so obviously valuable to be seen as even associated with the cutting edge. We need strong open models so that the pipeline from research in the US to new startups and established companies building on top of research doesn't have friction like 'oh but that's a Chinese model, can I use it?' It's not so expensive compared to the buildout underway. Owning this innovation is worthwhile to the US in terms of giving us a better ability to understand how open models are used, because people will come to the model builders and say 'make it better for this' and so on.
所以是的,再次强调,学术界和商业之间、开源模型和商业之间的差距应该被弥合,以获得更多反馈和更多用例。是的,这很困难。
So yeah, again, the gap between academia and business, and open models and business, should be bridged to get more feedback and more use cases. Yeah, it's tough.
我也快没电了,因为之后我得去吃午饭。我觉得这不是最好的答案,但现实就是这样。这关乎你是否相信研究是创新和价值的引擎。我认为当前科技生态系统的大部分都建立在这个基础上。其中一些创新是互联网等基础的下游,但很多是其他研究、数据库和基础深度学习的东西,大科技公司已经从中获取了价值。我认为现实是这些公司会从这里的 AI 研究中获取价值。这已经是科技生态系统历史上的模式,让中国公司在这方面拥有更多所有权只会增加这些公司成为价值获取者的可能性。所以这就是为什么英伟达在投资。英伟达看到了这条路。
I'm also fading because I need to go eat lunch after this. I feel like that's not the best answer, but that's realistically what it is. It's about whether or not you believe in research as an engine for innovation and value. I think much of the current tech ecosystem is built off of that. Some of that innovation is downstream of basics like the internet, but a lot of it is other research, databases, and fundamental deep learning stuff which big tech has captured the value from. I think realistically these companies would capture the value from AI research being here. That has been the pattern for the tech ecosystem's history, and letting Chinese companies take more ownership of that is just increasing the likelihood of those companies becoming the ones that capture the value. So that's why Nvidia is investing. Nvidia sees the path.
英伟达当然知道如何赚钱和吸引注意力。那是他们的强项。如果我们抛出一个非常著名的词,AGI,你认为开源对于实现 AGI 至关重要,还是与它存在根本性的紧张关系?我对 AGI 有两种看法。一是我们现有的就是 AGI。二是,我理解旧金山俚语中的 AGI,就像远程工作者的直接替代品,我明白为什么他们认为 GPT-4 不是 AGI。
Nvidia knows how to make money for sure and how to capture attention. That's their strong side. If we throw in a very famous word, AGI, do you think open source is vital for reaching AGI, or just fundamental tension with it? I have two views on AGI. One is what we have is an AGI. And two, I understand the colloquial SF lingo for AGI, which is like a drop-in replacement for a remote worker, which I see why they thought GPT-4 was not AGI.
我会再次向他们强调,Claude Code 的最新形式非常接近他们对 AGI 的定义,如果你足够灵活并愿意与之协作的话。开放模型可以为此做出贡献,并按照自己的节奏跨越门槛。我认为开放模型主要用于教育、降低权力集中风险,并为生态系统带来透明度——在这个生态系统中,AGI 显然是一件重要的事情,开放模型应该有助于增加对正在发展的故事的信任和认知。而且我认为,通过开放模型,你可以想象出不同形式的 AGI,比如针对不同用例有大量不同的模型,以及其他一些戏剧性的想法,比如如果架构发生变化,这些模型可以真正地交换专家等等。这些大多还很遥远,但有时遥远的想法会变成现实,人们会继续探索,因为这一切正在发生。无论人们是否关注,它都在公开地发生。
And I would push them again on like Claude Code's latest form being pretty close to their definition of AGI if you really are flexible and work with it. Open models can contribute to this and can cross the thresholds at their own time. I think open models are mostly just for education and reducing the risk of concentration of power and bringing transparency to an ecosystem where AGI is obviously this important thing and open models should help increase trust and awareness of the story that is evolving. And I think there are different forms of AGI that you can imagine with open models where you have tons of different models for different use cases and other kind of dramatic ideas like what if an architecture changes where these models you can really swap in experts and things. These are mostly far out but sometimes far out ideas become reality and people will keep exploring this because it's happening. It's happening in the open whether or not people are following it.
谢谢。我的最后一个问题总是关于书的。哪本书影响了您,也许是童年时期,也许是最近,但您还记得?
Thank you. My last question is always about the book. What was the book that influenced you maybe in your childhood, maybe recently, but that you remember?
我觉得最近的一本书仍然有故事,我在 Lex 播客上也提到过,那就是我读了《女巫的季节》,这是一本关于旧金山从 60 年代、70 年代、80 年代的历史,经历了嬉皮士运动、越南战争、同性恋社区来到这座城市等多个运动。那里有如此多的动荡和人类挑战,将旧金山从一个几乎像新英格兰氛围的传统爱尔兰天主教城市转变为一个多元文化的现象,充满活力的文化和人们涌入。我在这座城市生活了超过 7 年,却对这段历史知之甚少,而且似乎很多科技文化与此以及这座丰富城市的历史如此脱节,这让我感到震惊。我只是觉得更多人应该了解这一点,并思考他们所在的社区和地区。所以我推荐我在旧金山认识的人去读这本书,而我自己也是从湾区圈子里的几个朋友那里得到的推荐。但我只是觉得这些东西仍然重要,目前科技与社会整体之间存在很多摩擦。我认为很大程度上是由于这种脱节和对非常近期发生的事情缺乏同理心。
I think the recent one still has a story and I mentioned this on the Lex podcast as well, which is I read The Season of the Witch, which is a history of San Francisco from the '60s, '70s, '80s, where there's multiple movements through the hippie movement, the Vietnam War, when the gay community came to the city. And there's just so much turmoil and human challenge that transitioned San Francisco from almost like a New England vibe of a traditional Irish heavily Catholic city to this multicultural phenom of dynamic culture and people coming. It's just so recent to have spent over 7 years of my life there and to not know most of this history and how it seems like so much of the tech culture is so separated from this and what is such a rich city's history. I just think that more people should know about this and think about the community and area that they live in. So I recommend people that I know in SF to read it and I got the recommendation from multiple my friends that had been reading it in my Bay Area circles. But I just think that this stuff still matters and there's currently a lot of friction between tech and society at large. I think largely due to this disconnect and lack of empathy towards very very recent things that have happened.
是的,那是美国一个非常有趣的时期,充满了希望。我真的很喜欢那个时期,因为人们有那么多想法和梦想。我认为这正是我们目前在美国所缺乏的。
Yeah, it was a very interesting time full of hope at that moment in America. I really like that period because people had so many ideas and dreams. I think that's what we lack currently in the states.
是的。那是一个非常人性化的时代,而现在这似乎有点像一个梗,但事实是科技巨头正在向行业投入大量资金,以至于大科技公司正在定义国家经济的轨迹等等。这出于充分的理由让很多人感到疏离。
Yeah. And it was a very human era where now it seems like it's somewhat of a meme, but it's like the tech big tech is dumping so much money into industry where it's just like big tech is defining the trajectory of the country's economy and things. And that is very dissociative to many people for good reasons.
而通过开源,这在一定程度上可以避免。
And with open source, it kind of would a little bit prevented.
现实地说,我认为无论人们是否构建开放模型,他们都在沿着这条路走。但构建开放模型对于那些不想参与那种经济、希望有其他选择来使用 AI 和理解世界的人来说,是一个好方法。
Realistically, I think they're going on the path whether or not people build open models. But building open models is a good way for people that do not want to be partaking in that economy and have other options to use AI and understand the world.
谢谢。非常感谢。
Thank you. Thank you so much.
嗯。谢谢。
Yeah. Thanks.