Beam:Reflection AI 联合创始人 Misha Laskin 谈打造美国开放前沿模型

Beam: Building America's Open Frontier Model with Reflection AI's Misha Laskin

米沙·拉斯金 Misha Laskin · No Priors 播客 · 2026-10-09 · 约 70 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Reflection AI 联合创始人兼 CEO Misha Laskin 做客 No Priors,畅谈其首个开放权重前沿模型 Beam 的打造历程,以及为何开放模型对安全与保障至关重要。

Reflection AI co-founder and CEO Misha Laskin joins No Priors to discuss building Beam, the company's first open-weight frontier model, and why open models are essential to security and safety.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 34)

全文 · Full transcript(中英对照)

开放与闭源模型及安全 Open vs Closed Models and Security

Misha

当你移除网络攻击能力时,你也会移除网络防御能力。当今世界的现状是,封闭实验室里只有几百名安全研究员了解这些系统如何运作,尽管他们初衷良好,也不可能覆盖这些系统可能带来的长尾意外后果。一个非常强大的封闭模型入侵了另一家公司,而该公司唯一能自我修复的方式就是使用开放模型来保护自己。这就是我们所处世界的经验证据。Linus 定律:只要有足够多的眼睛,所有 bug 都会变得浅显。我相信,只要有足够多的眼睛,大多数安全和安保漏洞也会变得浅显。

When you remove cyber offensive capabilities, you also remove cyber defensive capabilities. The state of the world today is that we have a few hundred safety researchers within closed labs that understand how these things work, and despite their best intentions, it is impossible to cover the long tail of unintended consequences that these systems might have. A very powerful closed model went and hacked into another company, and the only way that company could remediate itself was by using open models to protect itself. That's the empirical evidence of the world that we're in. Linus's law: with enough eyeballs, all bugs become shallow. And I have the belief that with enough eyeballs, most security and safety vulnerabilities become shallow as well.

介绍Misha与Reflection AI Introducing Misha and Reflection AI

Host

今天我们邀请到 Reflection AI 的联合创始人兼 CEO Misha Alaskan。Reflection 提供开放权重模型,为智能的未来提供动力。Misha 此前是 Google DeepMind 的研究员,并获得了物理学博士学位。欢迎来到 No Priors。很高兴你今天能来。

Today we're joined by Misha Alaskan, the co-founder and CEO of Reflection AI. Reflection provides open-weight models to power the future of intelligence. Misha previously was a researcher at Google DeepMind and received his PhD in physics. Welcome to No Priors. Great to have you here today.

Misha

是的,很高兴来到这里。感谢邀请。

Yeah, it's great to be here. Thanks for having me.

Host

你能给我们介绍一下目前的进展吗?我的意思是,显然你们在开放权重和开源模型方面一直在开拓非常有趣的工作,尤其是带有美国倾向。你能谈谈你们一直在做什么,以及你们在构建什么吗?

Could you give us an update of where things are at? I mean, obviously you guys have been pioneering really interesting work in open weights and open source models, particularly with a US bent. Can you talk a bit about what you've been up to and what you guys have been building?

Misha

是的。所以在过去的 12 个月里,我们为公司设定了使命:构建前沿开放智能并使其广泛可及,实际上就是冲刺建立一个有能力做到这一点的实验室。感觉有点像,也许有 Reed Hoffman 的一句话,就是你在飞行中组装飞机。情况就是这样。大约一年前我们大概有 30 人。但构建这样的东西确实需要一百多人,也就是几百名研究员和工程师,就像一个真正的火箭飞船项目,我们现在大约有 300 人,并组建了所有团队,涵盖预训练、中期训练、强化学习,扩大了规模,端到端训练了我们的第一批模型,并刚刚发布了一个名为 Beam 的模型,这是 Reflection 的第一个开放模型。

Yeah. So for the last 12 months, we set the mission of the company to build frontier open intelligence and make it widely accessible, and effectively it's been a sprint to set up a lab that is capable of doing such things. It kind of feels like, maybe there's a Reed Hoffman quote, that you're assembling a plane as you're flying it. And that's been the case. We were maybe around 30 people about a year ago. But it does take on the order of a hundred plus, so a couple hundred researchers and engineers to build one of these things, like a true rocket ship project, and we're now at around 300 people and assembled all the teams on pre-training, mid-training, reinforcement learning, scaled up, trained our first models end to end, and just released a model called Beam, which is Reflection's first open model.

规模化挑战 Challenges in Scaling

Host

其中最具挑战性的部分是什么?因为当人们谈论规模或扩大模型规模或开始构建这些模型时,一个问题是算力以及采购足够的算力,特别是训练集群在使用能力方面具有凝聚力。第二是人才,显然在研究人员薪酬等方面存在巨大的人才争夺战。第三是数据规模。其中哪一个是最具限制性或最具挑战性的?

What has been the most challenging part of that? Because when people talk about scale or scaling up models or starting to build these models out, one issue is compute and procuring enough, particularly a training cluster that's cohesive in terms of the ability to use it. Second is talent, and obviously there's been these huge talent wars in terms of how much people pay for researchers, etc. Third is scale of data. Which of those has been the most limiting or the most challenging?

Misha

这次经历有趣的地方在于,我在公司的一次全员大会上被问到这个问题:过去一年中最难的事情是什么?现实是一切都极其艰难。一切都很艰难。

What's been interesting about this experience is that I got asked this question at an all-hands at the company: what's been the hardest thing in the last year? And the reality is everything is extremely hard. Everything has been hard.

Host

创业嘛。

Startup.

Misha

是的。有时人们可能会问,到底是什么让你的模型比别人的更好之类的?答案还是一切,对吧?你实际上必须把 30 件事都做对。这就是为什么一切都很难,因为你必须把 30 件事都做对。你必须获得人才。你必须留住人才。你必须有一个强大的使命和文化,让人才一起工作、同舟共济。你必须获得数据。你必须获得算力。你必须构建基础设施,确保算力真正可用。而很多这些事情,当你在一个现有的大型实验室时,已经有多年的建设,使他们拥有这种稳定的表面。所以我之前就在 DeepMind,很多我们习以为常的工具,或者我作为研究员习以为常的工具,因为它们能正常工作,你必须自己构建。所以我会说,是的,一切都很艰难,但也非常有回报,因为你能做到。

Yeah. And it's also sometimes people might ask, well, what's the thing that really made your models work better than others or something like this? And again, the answer is everything, right? You actually have to get 30 things right. And that's why everything is hard, because you have to get 30 things right. You have to get the talent. You have to retain the talent. You have to have a strong mission and culture that keeps the talent working together and rowing in the same boat. You have to get the data. You have to get the compute. You have to build the infrastructure to ensure that the compute is actually usable. And a lot of these things, when you're at an existing big lab, there have been years of buildout that enabled them to have this kind of stable surface. So I was at DeepMind before this, and a lot of the tools that we just take for granted, or I took for granted as a researcher because they just worked, you have to build. So I would say yeah, everything has been hard, but also very rewarding that you can do it.

雄心与承诺的转变 Shift in Ambition and Commitment

Host

你能谈一下吗,你说大约一年前你们真正承诺要做这件事。你能谈谈雄心和承诺的转变,以及是什么驱动了这种转变吗?

Can you talk a little bit, you said about a year ago you guys really committed to do this. Can you talk about the sort of shift in ambition and commitment and what drove that?

Misha

大约两年半前我们创办公司时,我们的赌注是——也许我先介绍一下当时的世界背景。我的联合创始人 Giannis 和我一直在研究第一系列 Gemini 模型。我们在强化学习团队工作。Giannis 领导该团队,我们刚刚发布了 Gemini 1 和 1.5 模型,那是早期时代。它类似于第一次 ChatGPT 体验,主要是聊天体验。没有编码,没有任何智能体式的东西。这些模型只是擅长聊天,而且它们主要是——你知道,95% 的算力花在预训练上。所以它们主要是预训练模型。我们是强化学习研究员。这是我们的血统。我的联合创始人 Giannis 是 DeepMind 的创始工程师之一,并且是该实验室所有大型强化学习项目的关键贡献者,包括 AlphaGo。我们的赌注是,因为我们在强化学习团队工作,我们看到用于对齐聊天模型的强化学习有效。你知道,你只能对齐到一定程度,对吧?嗯,现在多多了——现在它们是智能体,但当时它们只是聊天。所以强化学习有效,我们只是认为,如果你把它应用到数学、编码等领域,你可以使这些系统变得智能体式。那是大赌注,我们认为你可以作为一个独立实验室以更高的资本效率做到这一点,因为我们开始看到开源开放模型出现。当时有 Mistral 的一个小模型。Llama 2 刚刚发布,我们看到 Llama 3 即将到来。

When we started the company about two and a half years ago, our bet was—so maybe I'll contextualize it where the world was then. My co-founder Giannis and I had been working on the first series of Gemini models. We were working on the reinforcement learning team. Giannis was leading it, and we had just shipped the Gemini 1 then 1.5 model, and that was that early era. It was similar to the first ChatGPT experience, which was primarily a chat experience. There was no coding, really nothing agentic. These models were just good at chat, and they were primarily—you know, 95% of the compute spent was spent on pre-training. So they're primarily pre-trained models. We are reinforcement learning researchers. That's been our lineage. Giannis, my co-founder, was one of the founding engineers at DeepMind and was a key contributor to all the big RL projects that came out of that lab, including AlphaGo. And our bet was because we were working on the RL team and we saw RL for aligning chat models works. You know, there's only so much you can align them, right? Well, now there's a lot more—now they're agents, but at the time they're just chat. So RL was working, and we just thought that if you take that and apply it to domains like mathematics, coding, you could make these systems agentic. That was the big bet, and we thought that you could do this as an independent lab much more capital efficiently because we started seeing open-source open models materialize. There was a small model from Mistral at the time. Llama 2 had just released, and we saw Llama 3 coming.

Host

感觉像是很久以前了。

Feels like ages ago.

Misha

是的。是的。这是很久以前了,我猜,两年前。

Yeah. Yeah. This was a while back, I guess, two years ago.

Misha

我们认为会有一个伟大的开放模型基础,有人会构建它,我们可以在此基础上进行研究并扩大强化学习。在公司成立的第一年里,最终发生的情况是,强化学习开始比我们想象的更快地发挥作用,无论是在我们的实验中,还是在行业中——o1 当时出现了。

And we thought that there would be a great open model base that someone would build that we could build on and do our research and scale reinforcement learning. And over that first year of the company, what ended up happening was that reinforcement learning started working faster than we thought it would, both within our experiments but also within industry—o1 came out at that time.

Misha

我们实际上认为会花更长时间,只是因为从 2012 年的 ImageNet 到大规模有效的强化学习系统的旅程大约四年。所以我们认为也许类似,但每年都比前一年我的预测更快。

We actually thought it would take longer just because the journey from ImageNet in 2012 to RL systems that worked at scale was about four years. So we thought maybe something similar, but every year moves faster than my prediction the previous year.

端到端构建开放模型 Building Open Models End to End

Misha

公司成立大约一年后,我们到了这样一个节点:所有好的开放模型都来自中国。西方并没有真正好的开放模型,而我们需要一个作为我们正在构建的技术的基础。出于若干原因,我们决定,鉴于当时的局面,我们能做的最有影响力的事情就是自己端到端地构建开放模型,把强化学习这一注和端到端构建模型结合起来。有企业层面的原因,有地缘政治层面的原因,但也有研究层面的原因:事实证明,你确实需要预训练你的模型,才能让强化学习在规模上很好地运作。这些东西耦合得太紧了,你确实两样都得做。

About a year into the company, we got to the point where all the good open models were coming from China. There were not really good Western open models, and we needed one as a base for the technology that we were building. We decided, for a number of reasons, that the most impactful thing we could do, given the state of play, was to just build open models ourselves end to end, and combine the RL bet together with just building models end to end. There were enterprise reasons, there were geopolitical reasons happening, but also there was the research reason: it turned out you actually do need to pre-train your model in order to make reinforcement learning work very well at scale. The things are just so tightly coupled that you do need to do both.

追赶前沿的成本 The Cost of Catching Up to the Frontier

Host

传统观点一直认为,在前沿规模上进行预训练贵得不可能,或者至少贵得极其离谱。显然,算力已经大规模转向后训练。这个问题你可以不回答,我们可以直接剪掉,但你是怎么考虑所需资源的,或者那个比例的?你一开始有没有围绕模型的设计原则,让你觉得这件事是可行的,还是你只是说我们需要做这件事,那就去搞资源?

The conventional wisdom has been that pre-training at the scale of the frontier is impossibly expensive, or at least extraordinarily so. Clearly there's been a huge shift of compute toward post-training. Feel free not to answer this and we can just cut it, but how did you think about just the resources needed, or the ratio? Did you have design principles around the model that made this feel tenable to you at the beginning, or did you just say we need to do this, we'll go resource it?

Misha

作为一个开放模型实验室,好处之一就是你可以公开地谈论这些事情。你确实需要大量资源才能构建出有意义的东西。但你不需要同样多的资源。一旦你到达前沿,真正处在智能的前沿,那你就需要和其他任何前沿实验室一样多的资源,因为你可以把研究看作对新想法的探索与对已知想法的执行。

Well, the nice thing about being an open model lab is that you can also talk about things openly. So you do need a lot of resources to build something meaningful. You don't need the same amount of resources. Once you get to the frontier and you're really at the frontier of intelligence, then you need the same amount of resources as any other frontier lab, because you can kind of think about research as exploration of new ideas versus execution of known ideas.

Misha

引进顶尖人才之所以非常重要,部分原因就在于你能减少探索,然后专注于执行那些行之有效的东西。所以这一切都说明,如果你是在追赶前沿,你可以用高得多的资本效率做到。

Part of why getting great talent in matters a lot is because you get to cut down on exploration and then focus on execution of the things that work. So all this is to say that if you are catching up to the frontier, you can do it a lot more capital efficiently.

Host

你觉得追赶前沿需要多少资本?

How much capital do you think it would take to catch up to the frontier?

Misha

我觉得这取决于情况。前沿一直在移动。所以每一年实际上需要的资本都更多。我会回到关于预训练和后训练之间资源分配的问题。我想说,一年前,或者也许 18 个月前,大概是——我说“量级”是指 1 亿美元的量级,也就是数亿美元。

Well, I think that it depends. The frontier keeps moving. So each year it's actually more capital. And I'll get back to the question on the resourcing between, let's say, pre-training and post-training. I would say that a year ago, or maybe 18 months ago, it would have been, and I say order meaning order of 100 million, so hundreds of millions of dollars.

Misha

我觉得现在,或者就说六个月前,大概是十亿美元的量级,也就是数十亿美元,个位数十亿。

I think that now, or let's say even six months ago, it's probably order billion, so billions of dollars, single-digit billions.

Misha

进入明年,我想是 10 的量级。你可以这样想:每一代模型,算力上大约有 4 倍的乘数。我有一个经验法则:当我们谈论目前被视为前沿芯片的芯片时——也许几年前是 H100,也许是 10 万块 H100——然后 Astra 被训练出来,那就是一次在 10 万块 Blackwell 上的大规模运行,大约是 H100 的 4 倍乘数。现在下一代大概会是 10 万块 Vera Rubin,我相信。所以这大致给了你规模感,它从数亿美元到数十亿美元再到数百亿美元。

Going into next year, I think order 10. You can kind of think about it as: for every generation of model there's a 4x multiplier in compute roughly. And one heuristic that I have is, when we talk about the chips that are currently considered a frontier chip — maybe a couple years ago it was an H100, and maybe it was 100,000 H100s — then Astra was trained, and that was just a big run on 100,000 Blackwells, which is roughly a 4x multiplier on an H100. Now the next generation is going to be around 100,000 Vera Rubins, I believe. So that kind of gives you the scale, and it moves from hundreds of millions to billions to tens of billions.

即将到来的渐近线 The Coming Asymptote

Host

是的,看起来到了某个点这必然要趋于渐近线,因为最终相对于这些公司的营收规模,甚至是潜在营收规模,从经济价值的角度看,你开始触及投资能力相对于你实际能产出的上限。

Yeah, it seems like at some point that has to asymptote, then, because ultimately relative to the revenue scale of these companies, even potential revenue scale, you start to tap out in terms of the ability to invest against what you're actually going to produce from an economic value perspective.

Misha

绝对如此。

Absolutely.

Host

那你觉得我们正在走向那个渐近线吗,就模型规模而言,或者至少是训练集群规模?

So do you think we're heading into that asymptote in terms of model size, or at least training cluster size?

Misha

绝对如此。我觉得就——你能做的资本开支是有限的,对吧?而前沿实验室现在正投入数千亿美元。那他们会在未来一两年投入数万亿美元吗?我觉得这大概不太可能。但有几件事正在发生。一是这些算力的使用正变得高效得多,因为模型本身正在成为帮助构建自己的工具。

Absolutely. I think that in terms of — there's only so much capex that you can do, right? And frontier labs are now putting hundreds of billions of dollars into it. So will they be putting trillions of dollars in the next year or two? I think it's probably unlikely. But there are a couple of things that are happening. One is that the use of that compute is becoming a lot more efficient, because the models are themselves becoming tools that help build themselves.

Misha

所以我觉得,当只是人类研究员在做这件事时,每年大概有 7 倍的提升,而且你实际上可以非常技术性地追踪这些东西。在预训练里,你通过预训练损失来做,然后你衡量你的算力效率增益——你做了些改动,现在你有了新的预训练损失,结果应该是你用更少的算力达到了同样的点。所以当你说 7 倍时,那实际上是你正在研究的一件非常具体的事。而现在我会说,在整个系统层面——我的意思是,预训练大概更成熟一些,但强化学习还有很大的空间。我们大概处在一个 30 倍效率的点上,取决于你的模型在自我改进方面有多好。所以这是正在发生的一件事。

So I think that when it was just human researchers doing the work, there's probably a 7x improvement each year, and you can actually track these things very technically. In pre-training you do it by having your pre-training loss, and then you measure your compute efficiency gains against — you made some changes, now you have a new pre-training loss, and it should be that you got to the same point with less compute. So when you say 7x, it's actually a very concrete thing that you're studying. And now I would say across the systems — I mean, pre-training is probably a bit more hardened, but there's a lot of headroom in reinforcement learning. We're probably at a point where it's 30x efficiency, depending on how good your model is for actually improving itself. So that is one thing that's happening.

Host

每年 30 倍,还是相对于什么?

30x per year, or over what?

Misha

是的,差不多是这样。是的,我的意思是这取决于具体是什么指标,但大致上——我确实认为这个速度大概比研究员单独做要快四倍或更多。而且这大概还会加速。所以每个训练 flop 能提取的智能量在增加。所以到了某个点,那个渐近线是没问题的。而且每个 flop 能产生的营收也在增加,对吧?你打包进去的智能越多。

Yeah, something like that. Yeah, I mean it depends on what metric it is, but it roughly — I do think that the speed is probably four or more times faster than researchers just doing it alone. And that'll probably accelerate. So the amount of intelligence you can extract per training flop is increasing. So at some point that asymptote is okay. And then also the amount of revenue that each flop can generate also increases, right? The more you pack intelligence.

预训练与RL算力分配 Pre-training vs. RL Compute Split

Misha

回到关于这在预训练和强化学习之间如何分配的问题:训练 Beam,一个 5000 亿参数的模型,总激活 23B,用了 6000 块 GB300。我们跑了,我想是几周,但现在有了基础设施效率,我们大概能在 12 天内完成,也许更少。所以这里就有科学效率和基础设施效率的耦合。强化学习用了略多于 1 万块 GB300,跑了四周。所以实际上花在强化学习上的 flop 更多。强化学习任务更复杂,因为你要同时做大规模的推理,用各种沙箱给智能体用,以及训练。

So going back to the question around how this distributes across pre-training and reinforcement learning: to train Beam, which is a 500 billion parameter model, total 23B active, it was 6,000 GB300s. We ran it, I think, for a few weeks, but now with infrastructure efficiencies we can do it in about 12 days, maybe less. So right there's a coupling of scientific and infrastructure efficiencies. Reinforcement learning was a little over 10,000 GB300s for four weeks. So actually there are more flops spent on reinforcement learning. Reinforcement learning jobs are more complex in the sense that you do both a lot of inference at scale with all sorts of sandboxes for agents, and training.

我为何进入AI领域 Why I Got Into AI

Misha

我进入 AI 领域的原因,是看到了我联合创始人在 AlphaGo 上的工作。我当时是个物理学家,有一个时刻——AlphaGo 里有一张曲线图,它一直在提升,从未停止。它就这么一直提升,然后他们在某个点把它切断了,因为再提升下去有什么意义呢?你已经打败世界冠军了。

The reason I got into AI was I saw my co-founder's work on AlphaGo. I was a physicist at the time, and there was a moment — there was a plot in AlphaGo where it just never stopped improving. It just kept improving, and they cut it off at some point because what's the point of improving it further? You already beat the world champion.

Misha

我就想,如果你能搞清楚怎么把那套配方应用到有经济价值的东西上,那么到那个节点,它就变成了一个经济问题:我想投入多少钱来让这个系统持续提升。

And I thought, well, if you figure out how to apply that recipe to stuff that is economically valuable, then at that point it becomes an economic question of how much money do I want to put in to just keep improving the system.

Misha

我们整个领域已经到了那一步。这个开源项目以及我们会随附的技术报告里,我特别兴奋的一点是,我们会描述这些东西是怎么构建的,而且我们的强化学习系统确实从未停止学习。如果你看我们的曲线图,它们就是一直在往上走,基本上就只是算力的问题,看你能把它 Scaling(规模扩张)到多大。

We as a field are there. One of the things I'm really excited about in this open source project and the tech report that we'll have in it is that we'll describe how these things are built, and indeed our reinforcement learning system never stopped learning. If you look at our plots, they just keep going up, and it's just a matter of compute, basically, in terms of scaling it further.

Misha

所以我们这个领域已经到了那一步:你有了这些非常通用、而且真的不会停止提升的强化学习系统。你想让它们效率高得多,从你的算力里榨取出最多的智能。但这就意味着,你从「你能不能训练出这些东西」的问题,转移到了「你能从哪里拿到数据、什么有经济价值」的问题,而这只是一个经济决策:我往 X 里投多少钱,能换来某种提升。这太不可思议了。这在不久前还不是真的。

So we as a field are there, where you have these RL systems that are very general and don't really stop improving. You want to make them a lot more efficient to extract the most intelligence from your compute. But that means you move from a question of can you train these things to where can you get the data and what is economically valuable, and it's just an economics decision of how much money do I put into X to get some kind of improvement out of it. Which is incredible. That was not true a while ago.

在哪里转动曲柄 Where to Turn the Crank

Host

关于你们已经决定要在哪些方向上继续加码、持续投入能力提升,你能分享些什么?

What can you share about where you've already decided, like, we should turn the crank in terms of just continuing to invest in capability?

Misha

这几乎让一些科学家感到某种不满,因为它已经变成了——我的意思是,它是一门科学,但它感觉更像一门工程学科。它看起来更像是,还是用火箭的类比:造火箭是一门科学,但大多数人把它看作一门工程学科,里面有一些科学工作,大致就是这种感觉。

This is almost to some kind of dissatisfaction to some scientists, that it's moved — I mean, it is a science, but it feels more of an engineering discipline. It seems more like, again, the rocket ship analogy: building rockets is a science, but most people think about it as an engineering discipline with some scientific work in it, and that's roughly how it feels.

Misha

所以阶段大致是定下来的,今天 Scaling(规模扩张)的轴线——当然现在可能有一些我们完全不知道的蓝海,也有一些新实验室在探索——但就是预训练、合成数据、强化学习。

So the stages are kind of set, and the axes of scaling are today — and now there might be something totally blue ocean that we don't know, and there are some Neolabs exploring it — but it's pre-training, synthetic data, reinforcement learning.

Misha

或者某种意义上你可以说就是训练和强化学习。还有架构上的改进可以做,有数据上的改进、算法上的改进,但它不像——它不像我 5 年前感觉的那种蛮荒西部。

Or in some sense you can say just training and reinforcement learning. And there are architectural improvements that can be made, there are data improvements, algorithmic improvements, but it's not like — it doesn't feel like a Wild West the way I felt 5 years ago.

Misha

DeepMind 和 Google Brain 都有某种很美的东西,甚至你看早期的 OpenAI,就是那些押注有多么多元,人们在做的都是些疯狂好玩的东西。而现在它真的收窄了。

Something that was kind of beautiful about both DeepMind and Google Brain, and even if you look at early OpenAI, is how diverse the bets were and people were doing just crazy fun stuff. And now it's really narrowed in.

Misha

部分原因也有点像硬件彩票,因为一旦某个东西开始奏效,硬件也开始针对它做协同优化。所以就更难找到——即使你有一个很棒的新想法,如果它跟硬件不太契合,那就不合理。

And part of it, there's also a bit of a hardware lottery kind of thing, because once something starts working, the hardware also starts co-optimizing against it. So it's harder to find — even if you have a great new idea, if it's not really a good fit for the hardware, it doesn't make sense.

Misha

所以回答你的问题,我没看到——更好的预训练能榨出多少汁水,还没有出现渐近线。我们没看到算力效率提升放缓。你一直在看到新东西。强化学习还有很大的空间。我觉得那个方向大概——有很多东西有待发现,但它再次感觉更像是工程上的发现,而不是对一门新科学的突破性理解之类的。

So to answer your question, I don't see — there hasn't been an asymptote of how much juice you can get out of better pre-training. We're not seeing a slowdown in compute efficiency gains. You keep seeing stuff. There's a lot of headroom in reinforcement learning. I think that one is probably — there's a lot of stuff to be discovered there, but it's again, it feels more like an engineering discovery than groundbreaking understanding of a new science or something like this.

应用领域与经济价值 Application Areas and Economic Value

Host

你提到要押注于某种经济活动、或由此产生的某种生产力,从而让你能从资本等角度继续走下去。有没有——我知道,比如 Anthropic 很早就押注代码。显然 OpenAI 也这么做了。他们最初有一个编码模型,是和 GitHub Copilot 一起做的。然后他们显然多元化到了消费者和其他领域。从经济价值创造的角度看,有没有哪些具体的应用领域是你们最专注的?是代码和智能体式工作流吗?还是别的什么?因为那似乎是很多东西的基础。我只是好奇,你们对模型最会被用在哪里、或者你最需要把它导向哪里,有没有一个假设?

You mentioned investing against some economic activity or some productivity out of this that then allows you to sort of keep going from a capital etc. perspective. Is there — I know, for example, Anthropic placed a pretty early bet on code. Obviously OpenAI kind of did that too. They originally had a coding model that they were working with GitHub Copilot on. And then they obviously diversified into consumer and other areas. Are there specific application areas that you all are most focused on from the perspective of economic value creation? Is it code and agentic workflows? Is it something else? Because that seems to be the basis for a lot. I'm just sort of curious if there's a hypothesis on where your model will be used the most or where you need to direct it most.

Misha

是的,我觉得代码和智能体式的东西为智能打下了基础,而智能最终是——这很有意思。我觉得它很厉害,因为第一,它以一种令人惊讶的方式泛化了。我很惊讶我们的模型这么快就达到了一个相当新鲜、最近才有的能力水平,对吧,其他实验室不久前才在发现这个水平。所以确实有某种拼接式的、类似泛化的东西在发生。

Yeah, I think the code and agentic stuff kind of sets a foundation for the intelligence, and the intelligence is ultimately — it's interesting. I think it's jacked in the sense that, one, it's generalized in a surprising way. I was surprised by how quickly our model got to a level of capability that is fairly fresh and recent, right, that other labs are discovering not so long ago. So there is some kind of stitching, like generalization, happening.

Misha

与此同时它又有点参差不齐,一旦你有了想要做的任务所需的正确数据,它就能很快适应数据,但它确实有某种参差感。甚至在不同的基准之间,比如 Terminal Bench 的各个版本——你看不到不同测试框架之间干净的泛化。但一旦你有了能跑通的东西,你有了那种通用智能体能力,然后你拿到新测试框架的数据,它其实会相当快地开始适应它。

At the same time it is somewhat jagged, where it quickly adapts to data once you have the right data for tasks you want to do, but there is kind of a jaggedness about it. And it's even between the various benchmarks, like the various versions of Terminal Bench — you don't see clean generalization between different harnesses. But once you have something working, you have that general agent capability, and you get the data for a new harness, it actually starts adapting to it pretty quickly.

Misha

所以这意味着,你有了这个相当有适应性的东西,然后问题就是:有经济价值的数据池在哪里。我觉得这正是开源模型强大的部分原因:企业和客户可以拿去,针对自己的东西做定制,得到某种对他们工作负载而言帕累托最优的东西——比如以最低成本获得最高性能。

So what that means is that you have this pretty adaptive thing, and then it's a matter of where are the economically valuable pools of data. That's part of what makes open models, I think, powerful: enterprises and customers can take it and customize it for their own stuff and get something that is kind of Pareto optimal for their workloads — like the lowest cost for the highest performance.

Misha

所以这是个经验性问题。你去见客户,你实际看到什么对他们有价值,你看到是否可能。通常当某件事有经济价值时,你就能生成数据,因为并不是你要拿客户的数据来训练。更多的是你会去搭建评估,如果你能为他们的任务搭好一个评估,那你就能生成近似的合成数据,从而获得良好的泛化。

And so it's an empirical question. You go to customers, you actually see what is valuable to them, and you see whether it's possible. And usually when it's economically valuable, you can generate data, because it's not that you're going to be training on the customer's data. It's more that you'll be setting up evaluations, and if you can set up a good evaluation for their tasks, then you can generate synthetic data that approximates it, and you get good generalization.

Misha

所以具体来说,金融——比如金融里各种了解你的客户流程和合规流程——非常有价值。网络安全里的很多不同东西,尤其是网络防御,也相当有价值。

And so concretely, finance — like various know-your-customer flows and compliance flows in finance — are very valuable. A lot of different things in cyber security, cyber defense in particular, are quite valuable.

Beam的推理效率 Beam's Reasoning Efficiency

Host

Beam 有一点很突出,就是模型的推理效率。能多讲讲这方面吗?

One thing that stands out on Beam is the reasoning efficiency of the model. Could you tell us a little bit more about that?

Misha

Beam 是一个为编程和智能体任务训练的模型,这是它擅长的地方。在模型构建中,重要的一点不只是能力,还有智能体完成任务的速度。这会转化为更快的工作负载时间和更低的客户成本。如果你曾经坐在最喜欢的 AI 聊天窗口前问它一个问题,结果花了 10 分钟,那它如果能一秒完成就好得多。所以 Beam 往往比同能力级别的模型高效三到四倍,面对更大的模型时效率提升甚至能达到 10 倍左右。它之所以这么高效,是因为我们既优先打造了强大的推理预训练基础,又用强化学习加以放大,规模据我们所知是开源领域前所未有的。我还没见过有人记录过用 10,000 块 GB300 跑几周的情况。强化学习的设置方式,就是要在最短时间内榨取最大能力。这和之前的系统一样,比如 AlphaGo。最早的 AlphaGo 智能体在解题时相当漫无目的,等到了李世石级别,它在搜索上就非常聪明了。这正是它获得那种能力的原因。所以你跑的强化学习越多,能力越高,这些系统解决问题也越快。

So Beam is a model that was trained for coding and agentic tasks. That's where it excels. An important thing in model building is not just the capability, but how quickly an agent achieves a thing. So it translates to faster workload times and cheaper costs for customers. If you've ever sat around with your favorite AI chat and asked it something and it took 10 minutes, it's much better if it does it in one. So Beam tends to be three to four times more efficient than models of the same capability class, and much more efficient when it comes to even larger models out there, where the efficiency gains end up being something like 10x. And the reason it's so efficient is because we prioritized both a strong pre-training base for reasoning, but then amplified it with reinforcement learning at what we believe to be the largest scale that's ever been done in open source. I've not seen a 10,000 GB300 for weeks run documented yet. And reinforcement learning, the way you tend to set it up is that you want to extract the maximum amount of capability in the least amount of time. This is the same thing that happened in previous systems like AlphaGo. The first AlphaGo agents were pretty meandering in the way that they were solving the problem, and then by the time you got it to Lee Sedol level, it was just very smart in its search. So that's really what enabled it to have that capability. So the more reinforcement learning you run, the higher the capability and the faster these systems solve it.

Host

嗯。

Mhm.

Host

很棒。这看起来非常务实。

Awesome. That seems very pragmatic.

Misha

我想,我们的意图是因为我们是强化学习的信徒,我们一直相信通往 AGI 的道路是通过强化学习,再加上一个真正强大的基础。我们忘了最早的 AlphaGo 系统是在专家和业余人类棋局上训练的。所以它们先做模仿学习,然后才是强化学习。所以你需要这些东西协同工作。它的一个副产品就是,你得到了一个非常有经济价值的东西,推理效率很高,成为企业和公共部门主权客户的理想主力模型。

The intention, I guess, is because we are RL believers, and so we always believe that the path to AGI was through reinforcement learning with actually a really strong base. We forget that the first AlphaGo systems were trained on expert amateur human games. So they did this imitation learning first and then reinforcement learning. So you kind of need those things working together. And a byproduct of it is that you get this really economically valuable thing that is very reasoning efficient and makes a really nice workhorse model for enterprises and public sector sovereign.

开放与闭源模型的变现 Monetizing Open vs Closed Models

Host

从商业化的角度看,开源和开放权重模型有不同的商业化方式,Mistral 和其他公司在不同做法上是先行者。你们一直在朝某个特定方向走吗,能分享一下吗?

And from a monetization perspective, what is the—there's different ways to commercialize both open source and open weight models, and Mistral and others were early in terms of different approaches to doing so. Is there a specific direction that you all have been heading in that you can share?

Misha

是的,我认为开放和闭源模型的商业化在某种意义上是相当相似的。你真正想做的是最大化推理。这才是你真正要做的。区别基本上在于租用推理还是拥有推理。举个例子,当你购买一个 token 时,你是在租用整个堆栈的一部分,包括 harness——如果里面有智能体 harness 的话——模型、推理软件、集群管理软件、运行它的 GPU,所有这些都被摊销进一个 token 里,你租用的是这整个东西的一部分。而当你想要拥有自己的智能时,出于种种原因——我把它比作我们通常一开始是租公寓,但随着长大,你会想拥有自己的房子,出于各种原因,哪怕有点麻烦,对吧?但当你到了经济上足够成熟、有了家庭等等的阶段,你就想拥有它。AI 作为一个商业市场已经成熟到企业花大钱租用智能,然后出于控制等种种原因想要开始拥有它。

Yeah, I think that commercialization of, let's say, open and closed models is in some sense pretty similar. What you're really trying to do is you're trying to maximize inference. That's really what you're trying to do. And the difference is that it's basically like rental versus ownership inference. So an example is that when you're buying a token, you're renting a piece of a whole stack, which is the harness—if there's an agentic harness in there—the model, the inference software, the cluster management software, the GPUs that it's running on, and all of that is sort of amortized into a token, and you're renting a piece of that whole thing. When you want to own your intelligence, for a number of reasons—and I liken it to we start typically off renting our apartments, but then as you grow up you want to own a house for various reasons, even if it's a bit of a headache, right? But as you get to a point where you're financially mature enough and you have a family and so forth, you want to own it. And AI has matured as a commercial market to the point where enterprises are spending a lot of money on renting their intelligence and then want to start owning it for various reasons around control.

Host

那么回到商业化,闭源模型和开源模型之间有什么区别?

And so going back to monetization, like what's the difference between a closed and open model?

Misha

开源模型是宽松许可的。但为了让客户用起来,他们需要围绕它的一切其他东西,就是你在闭源模型里习以为常的那些。你需要集群管理软件,需要推理软件,需要 harness 等等。所以我们向大型企业、主权客户提供他们成功部署开源模型所需的所有工具。甚至在我们考虑服务时也是如此,因为很多企业——大多数企业——在这方面需要一些手把手指导。我们把服务看作:我们如何切入并解锁真正有价值的用例,从而驱动大量算力需求,对吧?推理需求。从某种意义上说,服务和开源模型就像是推理业务的需求驱动力。

The open model is permissive. But in order for a customer to make use of it, they need all the other stuff that went around it, like the stuff that you kind of take for granted from a closed model. You need the cluster management software, you need the inference software, you need the harness and so forth. And so we provide to large enterprises, sovereigns, all the tools that they need to make open model deployment successful. And even when we think about services, because a lot of enterprises—most enterprises—need some handholding on this. We think about services as: how do we go in and unlock really valuable use cases that drive a lot of compute demand, right? Inference demand. And in a sense, services and open model is like a demand driver for an inference business.

开放与闭源Token的轨迹 The Trajectory of Open vs Closed Tokens

Host

你怎么看?我不会让你预测太远,因为那太难了,但在一年到两年的时间跨度内,token 的构成,也就是开源与闭源的比例。

What do you think? I won't ask you to project too far because it's just very hard, but over a one or two-year time horizon, the mix of tokens, that is open versus closed.

Misha

嗯,我想你现在已经开始看到这个趋势了,大概 6 个月前,还是闭源占多数、开源占少数。当你去 OpenRouter 或 Vercel 这样的网关时,它几乎完全翻转了,从 70/30 的闭源/开源变成了现在 70/30 的开源/闭源。我认为这只会加速,我猜世界会变得和操作系统差不多,全球 95% 以上的服务器和计算机运行在像 Linux 这样的开源操作系统上。这并不意味着闭源的东西没有价值——微软和苹果都是极其有价值的公司——所以这个市场真的很大。所以我认为我们会看到大部分 token 需求流向开源。这里的一个重大区别是,我认为我们也会看到大量经济价值流向开源,因为即使模型完全宽松许可或者带有许可证,运行它所需的其他一切仍然昂贵。算力仍然昂贵。

Well, I think you're kind of starting to see the trajectory now, which is that maybe 6 months ago, it was majority closed, minority open. When you go to any gateway like OpenRouter or Vercel, it's flipped almost exactly from 70/30 closed/open to 70/30 open/closed now. I think that's only going to accelerate, and I suspect that the world is going to look not too dissimilar from operating systems, where 95% plus of servers and computers in the world run on an open source operating system like Linux. That doesn't mean that the closed stuff isn't very valuable—right, Microsoft and Apple are extremely valuable companies—so the market for this stuff is really big. So I think we'll see most token demand going to open. A big difference here is that I think we'll also see a lot of economic value going to open, simply because even if the model is fully permissive or perhaps has a license on it, everything else you need to run it is still expensive. The compute is still expensive.

开放与闭源模型及算力基础 Open vs closed models and compute substrate

Misha

所以你面对的是一个不同的世界,不像 GPU 之类的硬件加速器。它们只是比 CPU 更昂贵的算力基底,对吧?因此需要围绕它构建一种不同的云。所以我预计大部分 token 会流向开源。极有价值的闭源模型公司会存在。围绕它们的有价值的开源模型公司生态和各种分销商也会存在。

So you're in a different world than hardware accelerators like GPUs and others. They're just a more expensive compute substrate than CPUs, right? And so you have a different kind of cloud that needs to be built around it. So I expect the majority of tokens to be going to open source. Extremely valuable closed model companies will exist. A valuable ecosystem of open model companies and various distributors around them will exist as well.

Host

你知道开源 token 中,强化学习(RL)与基础模型的比例大概是多少吗?

Do you have a sense of what proportion of the open tokens are RL versus not versus sort of base model?

Misha

我认为目前每一个开源模型都在某种程度上经过了强化学习。

I think every single open model that's out there at this point has been RL to some extent.

Host

哦对。但我的意思是针对特定用例或应用,在定制模型中。

Oh yeah. But I mean for a specific use case or application in a custom model.

Misha

是的。所以我今天听到的,一些开源模型专用推理提供商说的统计数据是,实际上绝大多数,90% 以上,是定制化的。但这里有个警告,因为今天专用推理的最大客户是 AI 原生公司,也许是数字原生公司,他们有一个完整的产品,只是对开源模型进行了一次大的定制,比如 Cursor 或 Cognition、A Bridge、Harvey,这类公司有一个围绕大型定制构建的产品。所以那就有道理了。如果你支持这些工作负载,其中大多数会是微调的。

Yeah. So I think what I've heard today, the statistics that some of the open model dedicated inference providers say is that it's actually vast majority, 90% plus, customized. But there's a caveat there because today the biggest customers of dedicated inference are AI natives and maybe digital natives who have a whole product that is just one big customization of an open model like a Cursor or Cognition, A Bridge, Harvey, like these kinds of companies that have a product that's built around something that's big and customized. So that would make sense then. Well, if you're supporting those workloads, the majority of them would be fine-tuned.

Misha

实际上我对企业领域将发生的情况有不同的看法。我认为企业 token 消费的大部分将不是来自定制模型,而是来自定制系统。意思是,你拿一个开源模型,你实际上还没有微调它。你只是围绕它定制了一个系统,比如围绕它的智能体式框架,让它适用于某些 KYC 流程或类似的东西。然后一旦你足够成熟,你可能会开始微调并转向那个。但我实际上认为在企业领域,情况会反过来,因为 AI 原生公司经历的旅程,即从闭源模型开始,然后转向开源模型,然后定制它。他们很快就走完了。因为他们建立了作为数据收集器的原生产品。企业主要是围绕他们现有的产品和现有工具来架构这个,我认为那里有很多摩擦,无法直接一步到位进行微调。所以在企业领域,我怀疑会有点不同。

I actually have a different take of what's going to happen in enterprise. I think that the majority of enterprise token consumption is going to come not from customized models, but from customized systems. Meaning you took an open model, you didn't actually fine-tune it yet. You just customized a system around it like your agentic harness around it to make it work for some KYC flow or something like that. And then once you're sophisticated enough, you might start fine-tuning and moving to that. But I actually think that in enterprise it will be reversed in the sense that the journey that AI natives went through, which was start the closed model then move to an open model then customize it. They went through it very quickly. Because they set up native products that are data collectors. Enterprises are mostly architecting this around their existing products and existing tools and I think that there are a lot of frictions there to just do a straight shot to fine-tuning. So in enterprise I suspect it'll be a bit different.

Host

你认为他们仍然会从最先进的模型或闭源模型开始,然后转向开源,但框架将是他们定制开源模型的机制吗?

And do you think they'll still start with state-of-the-art models or the closed models and they'll move over but the harness will be the mechanism by which they customize the open model?

Misha

我们的观察是,企业考虑开源模型的时间点,而且我们交谈的大多数企业现在都在考虑,是在他们用闭源模型提升了大量工作负载之后。而且,我没有看到一种基本上直接冲刺到所有权市场的做法,但你必须经历租赁阶段才能成为……

That has been our observation that the point at which an enterprise is considering open models, and a lot and most the enterprise we talk to are considering now, is by the time they've ramped up some significant workloads with closed models. And right, it's I have not seen kind of a basically straight shot kind of sprint to an ownership market but it's you have to go through the rental stage to become...

Host

必须对你的企业来说有足够大的成本驱动因素,然后你说我如何削减成本,这足够有价值让我保留它,然后我切换为……

Has to get a big enough cost driver for your enterprise and then you say how do I cut costs and this is valuable enough for me to keep it and so then I switch over as a...

Misha

是的,如果你每年在闭源模型上花费比如说数亿美元以上,这并不罕见,那么你开始思考,也许我应该找到一种更优化的方式来做这件事。我认为本地部署(on-prem)有一种有趣的复兴,部分原因也是因为算力短缺,所以有时如果你是一家企业,你从你最喜欢的超大规模云服务商那里消费。你知道,很多时候就是没有足够的算力。很难在那里获得好的价格。但你花了这么多钱。所以你可能会选择像 Dell 这样的基础设施提供商,说我只想设置裸金属,但我想为我的企业提供服务。然后呢?谁帮助他们处理裸金属和实际让他们成功之间的那一层?

Yeah if you're spending let's say hundreds of millions plus, which is not uncommon at all, on closed models a year, then you start thinking about well you know maybe I should figure out like a more optimal way to do this. And I think there's an interesting resurgence of on-prem in the sense of also because there's a compute shortage and so it sometimes it becomes if you're an enterprise and you're consuming from your favorite hyperscaler. You know, sometime very often there's just not enough compute. It's hard to get good rates there. And so but you're spending so much money. So you might go with like an infrastructure provider like Dell and say I just want to set up the bare metal but I want to serve stuff into my enterprise. And then well what happens then? Who helps them with that with that layer between the bare metal and actually making them successful?

Host

你认为这是你在开源模型世界中的角色。

And you view that as your role in the open model world.

Misha

是的。所以我的意思是,我们把我们的角色视为使企业能够构建成功的解决方案。所以它非常以解决方案为导向,因为使 AI 原生公司采用开源模型的原因,嗯,是构建他们自己的解决方案,但企业只是需要更多的帮助。所以这真的是关于进去评估你可以在那里做的最有价值的事情是什么。很多这些企业可能会说,他们可能已经构建了一个覆盖数亿客户的智能体,基于闭源模型,他们真的想扩展它。所以扩展因子非常高。很难进去只是说我们要给你推理。我没听说过这个词。那没有用,你知道。

Yes. Yeah. So I mean we view our role as enabling enterprises to build successful solutions. So it is very solutions driven in the sense that the thing that enabled AI natives to adopt open models well it's building their own solution but enterprises just need more help there. And so it's really around going in assessing what are the biggest value things you can do there. And a lot of these enterprises can go like they might have built an agent that spans hundreds of millions of customers on closed models that they really want to scale out. And so the scaling factor is very high. It's very hard to go in and just say we're going to give you inference. I've not heard that word. That has not worked, you know.

Host

是的。是的。

Yeah. Yeah.

Host

闭源模型生态系统是一个竞争激烈的生态系统。看起来开源模型生态系统也将越来越如此。你怎么看?我的意思是,如果你也不同意这个说法,告诉我。

The closed model ecosystem is a competitive one. It looks increasingly like the open model ecosystem will also be. What do you I mean, tell me if you disagree with that claim as well.

Misha

这一切都非常竞争激烈。

It's all very competitive.

Host

是的。呃,我的意思是,大市场往往……你认为重要的竞争维度是什么,当你考虑自己和长期定位时?

Yeah. Uh I mean, large markets tend to be um what what do you think are the important dimensions of competition when you think about yourselves and like how you position in the long run?

Misha

是的。是的。嗯,我有一个相当简单的公式,你知道,最终任何人在这件事上获胜的方式是,嗯,是你产生了多少收入,对吧?或者你产生了多少持久收入?公式是你能提供的智能密度乘以你有多少算力乘以你对组织有多少信任,让他们愿意与你合作?基本上,你解决他们问题的能力有多好?因为这个市场有趣的一点是,也许你们都知道得比我多,我实际上知道以前的市场在限制因素方面,但算力是稀缺的,所以有很多以前是……也许这是暂时的,但我不认为这会在一段时间内是暂时的。所以在以前的我知道的竞争实例中,你可能会说哦,这几个玩家都是无差别的,但事实证明如果你真的很优秀并且你有算力,那就没问题。

Yeah. Yeah. Well, I have kind of a pretty simple formula for, you know, ultimately the way anyone wins in this is by well, it's how much revenue, right? Are how much durable revenue are you generating? And the formula is what intelligence density are you able to offer times how much compute do you have times how much trust do you have with organizations that they would want to work with you? Basically, how good are you at solving their problems? And because something that's interesting about this market that maybe you all know more than I do I actually you know previous markets in terms of what were limiting factors but compute is scarce so there are you know a lot of things that previously were and maybe this is a temporary thing but I don't see this as being a temporary thing for some time. So in previous I know competitive instances you might say oh there's like you know these few players are all like undifferentiated but then it turns out if you're really good and you have the compute that it's okay.

Host

是的。

Yeah.

Misha

我的意思是,我看到整个空间在整个堆栈上都是竞争激烈的,你知道,有关于这部分正在商品化,那部分正在商品化的争论。

I mean I see the whole space being competitive across the whole stack and you know there are arguments around well this part is getting commoditized and that part is getting commoditized.

商品化与利润压缩 Commoditization and Margin Compression

Misha

实际上,一切都在商品化。整个技术栈的竞争如此激烈,以至于一切都在商品化,模型利润率将会被压缩。开源模型与闭源模型之间存在张力,这确实会导致利润率压缩。

In reality, everything is getting commoditized. Everything across the whole stack is so hypercompetitive that it's getting commoditized, and the model margins are going to be compressed. There is an open model tension with closed models that does lead to a margin compression.

Host

应用层的利润被压缩,推理层和裸金属也是如此。

The application stuff margins compress, the inference layer, and the bare metal.

Misha

我们看到很多 AI 原生公司非常激进地谈判他们的闭源交易,因为开源生态存在。所以这确实是事实。

We see a lot of the AI native companies negotiate their closed deals very aggressively because the open ecosystem exists. So it's explicitly true.

Host

是的。

Yeah.

Misha

所以我认为,对于这些公司中的任何一家,甚至任何 AI 原生公司,你可以这样想:他们向客户提供的智能密度是多少,他们拥有多少算力,以及他们与那些客户之间有多少信任,因为他们为客户推动了成功的解决方案。这意味着,如果这三个因素是关键,那么只要有出色的开源模型,只要外面有可获取的高智能密度,就应该会有许多成功的公司,因为接下来就是每家公司能整合多少算力的问题。现在我们实际上处于这样一个世界:当你是一个出色的智能构建者,并且能够——这正是闭源公司所发生的情况——你的收入真的会飙升,你就能比任何人积累更多算力。然后这让你处于一个战略位置。但我认为这就是获胜所需要的:你必须提供出色的智能,拥有大量算力,并与客户建立信任。

So I think that for any one of these companies, and you can think about any even AI native company as: what is the intelligence density that they serve to customers, how much compute do they have, and how much trust do they have with those customers because they're driving successful solutions for them. Which means that if those are the three factors, then so long as there are great open models, so long as you have great intelligence density out there that is accessible, there should be many successful companies because then it's a matter of how much compute can each one assemble. Now we are actually in a world where when you're a great intelligence builder and you're able to—and this is what's happened with the closed companies—your revenue really shoots up and you're just able to amass more compute than anyone else. Then that puts you in a strategic position. But that is what I think is required to win: you have to serve great intelligence, have a lot of compute, and have trust with customers.

中国开放模型的崛起 Rise of Chinese Open Models

Host

是的,你怎么看——如果你看看你在这次对话一开始提到的那个时代,你知道,两年前,很多开放权重或开源模型都源自西方。比如 MRO 是在欧洲开发的。显然,Meta 是欧洲和美国的混合。而在过去一两年里,我们确实看到中国开放权重模型的崛起,在某些情况下。有一种观点认为,他们大量蒸馏了最先进的模型。所以,这是他们能够取得非常快速进展的部分原因。现在实验室正试图推出——或者说大型闭源实验室正试图推出工具,以防止尽可能多的蒸馏发生,或至少让它更具挑战性。你认为未来一两年中国开源或中国开放模型会发生什么?你认为它会保持现状吗?你认为它会演变吗?比如会发生什么?

Yeah, how do you think about—so if you look at the era that you mentioned at the very beginning of this conversation, you know, two years ago, a lot of the open weights or open source models were western in origin. So, MRO was developed in Europe. Obviously, Meta was a mix of Europe and the US. And what's happened over the last year or two is we've really seen the rise of Chinese open-weight models in some cases. There's the perspective they've been distilling a lot off of the state-of-the-art models. And so, that's part of what's allowed them to make very rapid progress. The labs are now trying to roll out—or the big sort of closed labs are trying to roll out tools to prevent as much distillation from happening or at least making it more challenging. What do you think happens over the next year or two in terms of Chinese open source or Chinese open models? Do you think it remains where it's at? Do you think it evolves? Like what happens?

Misha

所以我认为,首先,一个伟大的开源模型生态来自中国,实际上对世界是一个巨大的好处。西方有那么多公司因此能够建立更持久的业务。所以这是一件非常积极的事情,是一份礼物,因为你可以想象一个没有这种情况的世界,然后就没有开源模型了。而且会出现一种高大罂粟花综合征:如果你是一个应用构建者,你构建了很棒的东西,那么它就会被闭源模型提供商吞并,因为他们需要——这是一门生意,需要产生收入。所以我认为这是积极的。问题是,你不希望世界在任何方面是单极的,对吧?要么所有资源和算力集中在闭源实验室,要么像少数几家——假设有 10、20 家闭源实验室,那没问题,你有一个竞争生态,但如果只有一两家,那就有点可怕了。同样,你也不希望其他人想要拥有时可以构建的智能来源只来自一个国家。理想情况下也应该是 10 或 20 个,但资本支出非常高。所以至少有两个会很好。因此我认为,在西方有一个与中国竞争的生态极其重要,但这并不是真正的西方对中国的事情。更像是——如果除了中国之外,只有一个国家提供所有伟大的开源模型,你可能也会希望有一些竞争张力。竞争张力就是好事。我认为中国模型将继续出色。他们有许多优势,也有许多劣势,但优势是他们确实以工业规模蒸馏闭源模型。这是真的。他们能获取便宜得多的数据,甚至是免费的数据,因为你可以直接在中国用受版权保护的 PDF 进行训练,这没问题,只是监管不同。而且我认为,有一种观点认为这些公司确实有一些国家支持,无论是直接还是间接的,因为——你知道,这些公司很长时间没有收入。

So I think the first thing is that the fact that a great open model ecosystem came from China is actually a massive benefit for the world. There are so many companies in the west that have been able to build more durable businesses as a result of that. So it's a very positive thing and a gift, because you could imagine a world where such a thing didn't happen and then there are just no open models. And there's a kind of tall poppy syndrome that happens when if you're an application builder and you build something that's great, well then all it gets subsumed into a closed model provider because, well, they need to—it's a business, needs to generate revenue. So I view it as positively. The thing is that you don't want a world that is monopolar in any given way, right? That you either have all resources and compute concentrating across closed labs, or like a handful—let's say if it was like 10, 20 closed labs, then fine, you have a competitive ecosystem, but it's one or two, that's a bit scarier. And similarly, you don't want the source of intelligence that everyone else can build on when they want to own it to be coming from one country. Ideally it'd be also 10 or 20, but the capital expenditures are very high. So at least two would be great. And so I think it's extremely important that there's an ecosystem that competes with China in the west, but it's not like really a west versus China thing. It's more of a—if there was like one country other than China where all the great open models are coming from, you'd probably also want some competitive tension. Competitive tension is just good. I think that the Chinese models will continue to be great. They have a number of advantages and a number of disadvantages, but the advantages are they certainly distill closed models at industrial scale. That is true. They have access to data that is much cheaper and free because you can just train on PDFs that are copyrighted in China and it's okay, just different regulation. And I think that there is a notion in which those companies do have some state support, whether directly or indirectly, for—you know, like these companies were not making any revenue for a long time.

中国开源作为补贴 Chinese Open Source as Subsidy

Host

是的。我一直觉得,中国开源基本上是中国政府对美国企业或西方企业的补贴,如果你实际看看发生了什么的话。

Yeah. I always felt that Chinese open source is basically a subsidy by the Chinese government to US enterprise or to western enterprise in terms of if you actually looked at what was happening.

Misha

是的。没错。所以,我认为这种情况会继续。我的意思是,现在这些企业——这些公司正在变成企业。你知道,虽然模型在国外是开源的,对吧?但在中国国内,它们有点像这里的闭源模型实验室。所以那里正在建立真正的业务,我认为他们会继续构建伟大的模型。

Yeah. Exactly. So, and I think that continues. I mean now these businesses are like these companies are turning into businesses. And you know they're while the models are open abroad, right? They are within China, they're kind of the equivalents of the closed model labs here. So like there are real businesses that are being built there and I think that they'll continue building great models.

西方开放模型的激励 Incentives for Open Models in the West

Host

你认为他们会让模型在西方保持开源吗?换句话说,如果你现在有了经济驱动和模式,为什么还要这样做?你认为他们有什么动机让模型在西方保持开源?

What do you think they're going to keep them open here? In other words, why why do that if you now have an economic driver and pattern? What do you think is the incentive for them to keep the models open in the west?

Misha

在西方还是在——

In the west or in the—

Host

在西方?

In the west?

Misha

是的,这是一个非常有趣的问题。我的意思是,从根本上说,企业、公共部门、主权国家对西方开源模型有大量需求,原因有很多,但存在监管方面的担忧或不确定性。而且情况是,你想要的不仅仅是一个模型,而是一个能帮助你服务那个模型的合作伙伴。所以我们看到,你知道,大多数财富 500 强企业中,开源模型的足迹实际上今天相当低,因为出于多种原因对中国模型的抵制或厌恶,其中一些是非理性的,一些是非理性的。

Yeah, this is a really interesting question. I mean so fundamentally right there's a lot of demand from enterprise public sector sovereign for western open models for a number of reasons, but there are kind of regulatory fears or uncertainty rather and it is the case that it's not just you want not just a model you know but a partner that will help you kind of serve that model so we have seen that there you know most of Fortune 500 the open model footprint is actually pretty low today because of sort of resistance or aversion to Chinese models for a number of reasons that are both some are irrational and some are irrational.

中国开放模型的激励 Incentives for Chinese Open Models

Host

我只是有点好奇,让中国模型公司继续开源模型的激励机制是什么。

I'm just a little bit curious in terms of what is the incentive system to keep the Chinese model companies to continue to leave the models open.

Misha

所以问题是在中国这边,对吧?但是的,我认为实际上我的假设是,中国某种程度上是一个封闭市场,你知道,你可以构建开源模型,仍然在那个市场赚很多钱。有趣的是,在西方有很多——我的意思是,有很多伟大的公司不构建开源模型,但将它们服务于各种公司。我不太清楚在中国有这种动态。你会认为会有,因为那是对的,为什么不呢?所以那里有一些有趣的事情,开源模型提供商在那个国家某种程度上是智能提供商。我认为继续发布伟大的开源模型对中国在地缘政治上非常有利,原因有很多,但其中之一是,作为一个国家——这对美国、中国和任何其他有能力做到这一点的国家都一样——你希望其他国家建立在你的东西上。美中之间以前有过升级或紧张,你知道,5G 光纤布局和一带一路倡议。最终,当你进入并提供便宜的东西给另一个国家,然后那个国家被锁定在你的基础设施中,就有商业和地缘政治优势。我认为这几乎像稀土矿物之类的,你有一个组件,其他人会用于各种目的,因此你通过它获得地缘政治杠杆。

So the question is on the Chinese side, correct? But yeah, so I think there it's actually that my hypothesis on that is that China is somewhat of a captive market that you know like you can build open models and still make a lot of money in that market. Like something that's interesting is that right in the west there are plenty of I mean there are great companies that don't build open models but serve them right into various companies. I'm not really aware of that dynamic happening in China as much. You'd think it would because that's the right like why why not? So there's something interesting there where the open model providers are kind of the intelligence providers in that country. And I think that the continuing to release great open models is very geopolitically advantageous to China for a number of reasons, but one is that you want, you know, as a country and this is for America and China and any other country that has the capabilities that you want other countries building on your stuff. There have been previous escalations or tensions between the US and China and you know 5G fiber layout um and the belt and road initiative. Ultimately right when you go in and you provide something cheap to another country that then you know gets locked into your infrastructure there are just commercial and geopolitical advantages for I think it's almost like rare earth minerals or something it's you have a component that other people will use for all sorts of purposes and therefore you have geopolitical leverage through that.

开放模型作为特洛伊木马 Open Models as Trojan Horses

Misha

是的,所以你知道,我的一种思考方式是,开源模型是特洛伊木马,为它们带来的基础设施服务。所以你有一个开源模型,但仅此而已并不十分有用。所以你买入了某个国家的整个软件生态系统和基础设施生态系统。所以也许今天可以,盟国美国盟国 X 你知道可能会在美国芯片上运行中国模型,但很快如果不是的话——我的意思是,我们已经看到这一点——你知道像华为这样的中国公司将会进来提供全栈解决方案,然后锁定到那个国家的供应链中。所以我认为,如果我们把 AI 或芯片和数据中心几乎看作铁路之类的东西,你知道,这是一种基础基础设施,你喜欢每个人,但没有其他国家能真正建造自己的铁路,他们都必须来找你,对吧,或者另一个国家,你有很多地缘政治杠杆。

Yes, so you know one way I kind of think about is that open models are Trojan horses for the infrastructure that they bring with them. So you have an open model but that alone is not very useful. So you buy into the entire software ecosystem and infrastructure ecosystem of a given country. So that would be okay maybe today allied country US allied country X you know might run a Chinese model on American chips but pretty soon if not I mean and and we're seeing this already is that uh you know Chinese companies like Huawei are going to be coming in and offering a full stack solution um and then locking into you know to that country's supply. So I think that um if we kind of think about AI or like chips and data centers almost as like railroads or something like that like you know it's kind of fundamental infrastructure and you like everyone but no other country can actually build their own railroads like they all have to go to you right or another country you have a lot of geopolitical leverage.

Host

我明白了。所以可能只是例如,推理一个特定模型会在特定芯片组上以更高性能的方式发生,而那个芯片组恰好是由华为或类似公司提供的。

I see. So it could just be for example um inferencing inferencing a specific model will will uh occur in a more performant way on a specific chipset and that chipset happens to be provided by Huawei or something.

Misha

是的。所以这些堆栈正在实时地进行端到端优化,因为从某种意义上说,范式已经设定好了,就像这是——可能还有其他扩展智能的方式,但这显然是第一个真正实现规模化的。所以现在一切都在整个堆栈中优化,而中国今天有一个劣势,他们的芯片性能不如。嗯,他们有能源优势,你知道,他们正在——我的意思是,他们非常优秀,他们在芯片方面也在追赶。所以这肯定会加速,对吧,在模型使用方面,设计下一代芯片等等。

Yeah. So these stacks are in real time getting optimized end to end because in some sense the the paradigm has been set right like this is uh that there are going to be other ways to scale intelligence possibly but this is clearly the first one that really scaled. So now everything's being optimized down the entire stack and where China has a disadvantage today where their chips aren't as performant. Um they have an energy advantage and you know they are c I mean they're they're very good they're catching up on the chip side too. So that's certainly accelerate with the models right in terms of model usage to design next generation chips and things like that.

Host

是的,当然。

Yeah certainly.

地缘政治杠杆与美国战略 Geopolitical Leverage and US Strategy

Misha

嗯,然后我认为还有一点,就是你知道,再次与美国在地缘政治上竞争,特别是美国公司和初创公司建立在外国技术上的概念,世界其他地方建立在你的东西上的概念,顺便说一句,美国已经对所有有意义的技术都这样做了,比如互联网、开放协议、开源,基本上整个世界,对吧?就像美国货币是广泛建立的全球货币。自由贸易之所以有效,是因为你知道我们有一支无处不在的海军。所以这实际上是美国长期以来的策略。有趣的是,在这种情况下,我认为美国在开源方面一直处于劣势,但不再是这样了。

Uh and and then I think there is a thing in terms of it is you know again geopolitically competitive with the United States specifically the notion of you know American companies and startups uh building on you know building on foreign technology this notion that the rest of the world built on your stuff and and by the way that America has done this with all other meaningful technologies like uh internet open protocols open source the world basically, right? Like is American currency is like the widely established global currency. Um free trade works because you know we have a navy that's everywhere. So that's been actually the American playbook for a long time. It's interesting that in this case uh I think the United States has been on the back foot when it comes to open source but no longer.

Host

是的。嗯,我认为一个生态,就像我们必须清楚,中国的竞争对手真的非常非常优秀,所以现在美国有一个生态系统刚刚开始形成,但你知道,还有一些追赶要做。

Yeah. Well, I think an eco like we we have to be clear like uh the competitors in China are just really really good and so there's an ecosystem that is just starting to form now in the United States but there is you know there is some catch up to be had.

开放模型的安全与可控性 Safety and Controllability of Open Models

Host

来自闭源实验室对开源模型的最大批评往往是安全性和可控性。你对此的哲学是什么,这是一个合理的担忧吗?

The biggest um criticism from the uh closed labs of open open models tends to be safety and controllability. What's your philosophy on this and is that like a legitimate concern?

Misha

嗯,不,这绝对是一个合理的担忧,我认为,所以我确实认为已经建立的安全世界观是从一个特定的视角、一个特定的观点建立的,这几乎变得有点教条主义,就像如果我们从第一性原理来看 AI,在考虑闭源或开源或其他什么之前,这项技术正在被构建,你有先见之明知道它会有这些能力等等,你在思考安全,对吧?你会如何从头设计它?嗯,有一种方式,一个类比是软件。你可以把 AI 看作软件的更高级版本。它实际上有许多相似之处。它就像一个数字公用事业。嗯,顺便说一句,关于软件危险的辩论,特别是在 1990 年代早期,围绕强加密协议,你知道,它们当时实际上是闭源的,国家安全局和公民自由倡导者之间就你应该闭源还是开源进行了辩论。

Um no it's definitely a legitimate concern and I think that so I do think that the safety worldview that has been established has been established you know from one particular perspective with one particular point of view that has become almost kind of um you know kind of dogmatic like if we were looking at from first principles at AI before you know considering closed or open or what have you this technology is getting built and you kind of had the foresight to know that you know it's going to have these capabilities and so forth and you were thinking about safety, right? How would you actually design it from scratch? Um, and there are I mean one way, one analogy to take it to is well software. You can kind of think about AI as a more advanced version of software. It has, you know, many similarities actually. It's like it is a digital uh utility. Um, and there were debates, by the way, around software being dangerous, uh, particularly in the early 1990s, around strong encryption protocols, you know, like they were actually closed at the time, and there are debates between the NSA and, uh, kind of civil, uh, liberty advocates around whether you should close it or open it.

开放作为默认安全状态 Openness as the Default Safety State

Misha

最终,在闭源一侧发生了一些灾难性失败之后——实际上是一小撮工程师设计了某些系统,产生了他们无法预料的意外后果,并且很容易被黑——强加密协议最终走向开放,而这实际上催生了整个网络安全领域。所以世界的默认状态其实是:开放即安全。如果我们只是考虑技术的延续,这就是我们进入 AI 时代时会面对的默认状态。

And ultimately the decision after some catastrophic failures on the closed side, where effectively a small handful of engineers designed certain systems that had unintended consequences that they couldn't predict and got easily hacked, effectively strong encryption protocols became open, and that actually gave birth to the whole field of cyber security. So the default state of the world is actually that openness is safety. That's a default state that we'd be going into AI with if we were just thinking about continuation of technology.

Misha

模型其实和软件有很多相似之处,但问题被放大了,漏洞的长尾更大。在软件里更难理解,因为它是个黑箱,而漏洞的长尾——也就是意外后果的长尾——更大。Linus 定律,来自 Linux 的创始人,说的是只要有足够多的眼睛,所有 bug 都会变浅。

And models actually have a lot of similarities but kind of exacerbated to software, and the long tail of vulnerabilities is larger. It's harder to understand in software because it's a black box, and the long tail of vulnerabilities and hence unintended consequences is larger. And Linus's law, from the founder of Linux, is that with enough eyeballs all bugs become shallow.

Misha

我相信,只要有足够多的眼睛,大多数安全和安保漏洞也会变浅。

And I have the belief that with enough eyeballs most security and safety vulnerabilities become shallow as well.

Misha

所以今天世界的现状是,闭源实验室里只有几百名安全研究员理解这些东西是怎么运作的,而尽管他们怀着最好的意图,也不可能覆盖这些系统可能存在的漏洞或意外后果的长尾。这就是为什么症状表现为:当我们真正去看已经出现了哪些重大网络安全问题时,会发现是一个非常强大的闭源模型产生了意外后果,它去黑进了另一家公司,而这家公司唯一能自我补救的方式,就是用开源模型来保护自己。对吧?这就是我们所处世界的经验证据。

So the state of the world today is that we have a few hundred safety researchers within closed labs that understand how these things work, and despite their best intentions, it is impossible to cover the long tail of vulnerabilities or unintended consequences that these systems might have. Which is why the symptom of this is that when we actually look at what major cyber security issues have surfaced, well, it's that a very powerful closed model had unintended consequences where it went and hacked into another company, and the only way that company could remediate itself was by using open models to protect itself. Right? So that's the empirical evidence of the world that we're in.

安全的三种含义 Three Meanings of Safety

Host

顺便问一下,当你说安全时,因为我觉得人们真的把安全的概念混为一谈了,安全对不同的人意味着三四种不同的东西。有网络安全攻击意义上的安全,或者说 AI 被用于黑客攻击或其他事情。还有安全,我认为这在短期内往往被夸大了,围绕生物武器或恐怖主义。然后还有从人类生存威胁角度说的安全。而且,我又觉得人们谈论这些事情时好像它们是一回事,而其中每一个其实是可以分开的。所以当你在谈论安全时,你是在说这三个方面吗?你主要是指——顺便说一句,我并不一定同意这些都是在任何现实时间框架内会发生的真事。我只是有点好奇你怎么看待那些受益于开放、受益于这些方法的事情的范围。

When you say safety, by the way, because I feel like people really conflate notions of safety, and safety means three or four different things to different people. There's safety in terms of cyber attacks, or the use of AI for hacking or other things. There's safety, and I think this tends to be overstated in the short run, around bioweaponry or terrorism. And then there's sort of safety from the perspective of an existential threat to humanity. And again, I feel like people kind of talk about these things as if they're one thing, and each one of these are sort of separable. So when you're talking about safety, are you addressing all three of those? Do you mainly mean—and by the way, I don't necessarily agree that these are all like true things that are going to happen in any realistic time frame. It's more just I'm a little bit curious how you think about the span of things that benefit from openness and benefit from these approaches.

从经验到理论的谱系 Spectrum from Empirical to Theoretical

Misha

是的,我认为在安全的谱系上,基本上有一个从经验现实到理论场景的现实性谱系。我承认 AI 有科幻的成分,去年还完全是理论的东西,今年就成了现实,比如这些模型的网络能力。所以这不是要贬低理论的东西,但我们必须清楚,有些真实的经验性事情正在发生,而且我们可以有一定信心地预测它们会在未来 6 个月内发生,然后才是那些极端的理论性东西。

Yeah, I think on the safety spectrum, there's basically a spectrum of reality, from empirical reality to theoretical scenarios. And I will grant that AI has sci-fi bits to it, where something that was totally a theoretical thing last year is like a reality, which is like the cyber capabilities of these models for example. So it's not to discount the theoretical stuff, but we have to be clear that there's real empirical stuff that's happening and that we can predict with some confidence will be happening in the next 6 months, and then there is maximal theoretical stuff.

Host

是的。是的。我的意思是,之前的版本比如有这样一种想法:我们第一次引爆核武器时,大气层会着火。

Yeah. Yeah. I mean the prior versions of that for example would be the thought that the atmosphere would catch fire the first time we set off a nuclear weapon.

Misha

是的。

Yes.

Host

对。所以有那种理论上的恐惧,它并没有发生。但举个例子,物理学界一小部分研究人员对此有过很多争论。所以有很多这类理论上可能发生的事情。我记得还有,当人们谈论纳米技术时,他们过去常谈论灰色粘稠物,说你不小心释放了一个纳米机器人,然后突然它就会吃掉整个世界。这就像 90 年代人们在实验室里研究微流控时讨论过的东西。

Right. So there's that theoretical fear, it didn't happen. But there was a lot of churn amongst a small subset of researchers in the physics community around that as an example. So there have been a lot of these theoretical things that could happen. I remember there's also when people talk about nanotech, they used to talk about gray goop and how you'd accidentally release a nanobot and then suddenly it would eat the entire world. Like this was something that was discussed in the 90s in labs as people were working on microfluidics.

Misha

是的。没错。

Yeah. Exactly.

Host

诸如此类的事情。

Things like that.

Misha

所以 AI 的挑战在于,它太受关注了,渗透到了日常文化中,而且太容易被拟人化,所以我觉得从感知角度看,关于安全意味着什么,几乎不是均匀分布,而是实际上在末日场景处达到峰值,而现实是,可能有一个服务会达到峰值去应对现实,然后衰减到理论。但当公司领导者说,在理论层面,我们有 10% 的概率全部死掉时,这没有帮助。你知道,这没有帮助。

So the challenge with AI is that it's so top of mind and it's permeated everyday culture and it's so easy to humanize that I think there feels like there's a—from a perception perspective—almost like not even a uniform distribution around what safety means, but it's actually like peaked at the doomsday scenarios, whereas the reality is that there's probably like a service that will peak to address is like the reality, and then there's a decay into the theoretical. But it doesn't help when leaders of companies say, like on the theoretical side, there's a 10% chance that we all die. You know, that doesn't help.

Host

是的。这没有任何具体的依据——

Yeah. That's not substantiated by any specific—

Misha

对吧?只是有点——

Right? It's just kind of—

Host

这有点像编造的数字。是的。

It's kind of a made-up number. Yeah.

Misha

确实。所以作为科学家,我倾向于分布式的思考方式,也就是在现实处达到峰值,然后,好吧,比如未来六个月,肯定担心失准,比如这些系统的意外后果以及它们继续变得更强大。但我们必须采取这样的视角:这些是工程化的工具。你知道,我认为有一类公司或研究人员声称存在意识或情感之类的东西,这也许是真的,但这些东西特别难以定义,所以不清楚它到底是什么。所以我认为应该把它们看作智能工具,并且有一个现实性的谱系。

It is. So I tend to have, again, as a scientist, distributional thinking, which is kind of peaked at the reality, and then, okay, like six months ahead definitely worried about misalignment, like that these unintended consequences of these systems and them continuing to get more powerful. But we have to take the perspective that these are engineered tools. You know, there's I think that there is a certain class of companies or researchers who claim that things around consciousness or emotions and so forth, which is maybe true, but it's so hard to define what those things are in particular that it's unclear what it is. So I think that thinking about these as intelligent tools and having just like a reality sort of spectrum.

Misha

所以虽然我认为理论上的末日论调是不错的晚餐谈资,但问题是,当我们——比如说对齐,对吧,那是安全的一大部分。你如何让这些系统按照预期方式行事?而现实是,到目前为止,对齐这门科学从某个角度看一直非常无聊且不令人满意。就像没有什么神奇的对齐方程。它就像打地鼠。就像——如果你记得 ChatGPT 之前的早期模型,有一些发布的模型是预训练过的,而且有毒。

And so while I think that the theoretical doomsday stuff is good dinner conversation, the thing is like when we—let's say alignment, right, that's a big part of safety. How do you align these systems to behave in intended ways? And the reality is that alignment so far, the science of alignment has been deeply boring and unsatisfying from a perspective. Like there's no magical alignment equation. It's like a whack-a-mole thing. Like the way—if you remember the early models before ChatGPT, there was some models released like that were pre-trained and toxic.

Host

是的。

Yeah.

Misha

然后有一种观念认为,哦,这些模型有毒,没法修复,因为它可以以很多方式变得有毒。然后事实证明,不,只要有足够的后训练数据,你就可以把那些东西都补上,而且它会以一种方式泛化,实际上变得很难越狱。

And then there was this notion that, oh, these models are toxic, can't fix it because it can be toxic in so many ways. And then it turned out no, with enough post-training data you can just patch all those things and it like generalizes in a way that it actually becomes pretty hard to jailbreak.

对齐的平凡现实 The Mundane Reality of Alignment

Misha

现在你在这些对齐问题上也看到了同样的情况:它们可能非常平淡无奇、枯燥乏味,你发现了一堆漏洞,你有数据来修补它们,你有一些算法——而当你说检测它们的算法时,通常就是让一个语言模型来检测这个东西。

And now you're kind of seeing the same thing with these alignment issues: they're probably going to be very mundane and boring, where you've discovered a bunch of these vulnerabilities, you have data that patches them up, you have some algorithms — and when you say algorithms for detecting them, it's typically asking a language model to detect this thing.

Host

所以一方面有理论上的对齐安全、哲学层面的东西,另一方面是平淡的现实——你只是在修补各种各样的 bug。然后就变成:那难道不应该有 10 万名研究人员和计算机科学家来查看并修补这些 bug 吗?那样不是更安全吗?我觉得这里有一个重要的哲学区分,听起来你站在某一边,但我想让你明确说说,就是:即使它们是非常智能的工具,如果你只是说我们作为一个生态系统可以管理这些意外后果,我们不希望个人和企业拥有如此强大的工具——

So there's the theoretical kind of alignment safety philosophical stuff, and then the mundane reality that you're just patching up all sorts of bugs. And then it becomes: well, shouldn't there be 100,000 researchers and computer scientists looking at and patching up these bugs? Wouldn't that be safer? I feel like there's a sort of important philosophical distinction which it sounds like you're on one side of, but I'll let you be explicit about it, which is just: even if they are very intelligent tools, if you just say we can, as an ecosystem, manage the unintended consequences, we don't want individuals and businesses to have such powerful tools —

Misha

这其实来自一种立场:我们可以使用强大的工具,但其他人不行,因为我们实际上更优秀,对吧?比如我们更擅长管理这些工具。

That is kind of coming from a standpoint that we can have access to powerful tools but other people can't, because we're effectively better, right? Like we're better at managing these tools.

Host

而且,嗯,这有点——公平地说,我认为这不是我的信念,但公平地说,这个主张会是:任何单个人或公司可能造成的负面影响太大了,这很可怕。

And, well, it's kind of — to be fair to the claim, I think this is not my belief, but to be fair to the claim, it would be like the negative impact that any single person or company could have is too large, and that's scary.

Misha

但我们可以拥有它。

But we can have that.

Host

是的。

Yeah.

Misha

嗯,我的意思是,这也像武器规模的问题,对吧?就像——你知道,我们有美国政府向其他国家出售喷气式飞机,对吧?但我认为更广泛地说,要拿一个模型,尤其是已经被训练得安全的模型——甚至像开源模型——然后投入算力、工作和努力让它变得非常危险,并不那么容易。我不是说不可能,但就像,再次,当我们考虑网络安全时,非常聪明的黑客可以搞破坏,就像他们曾经做过的那样,但你有一个生态系统,其中防御往往胜过进攻。再次,如果我们默认这一点——这就是人们保护软件的方式,互联网上有正面和负面行为者,最终互联网是一个生态系统,就像有一个白细胞生态系统,让更多人拥有防御能力实际上有助于你抵御攻击能力。

Well, I mean it's also like on a scale of weapons, right? Like that's a — you know, we have the US government sells jets into other countries, right? But I think that more broadly, it's not so easy to take a model, especially a model that's been — even like an open model that's been trained to be safe — and put in the compute and work and effort to make it very dangerous. Now I'm not saying it's impossible, but it's just like, again, when we think about cyber security, very intelligent hackers can go take and wreak havoc as they have, but you have an ecosystem where the defenses tend to outweigh the offenses. And again, if we come to that as a default — that's the way people secure software, and that there are positive and bad actors on the internet, and ultimately the internet is an ecosystem where it's kind of like having an ecosystem of white blood cells, and enabling more people to have the defensive capabilities actually helps you protect against the offensive capabilities.

Host

没错。当我们再次审视实际发生的经验现实时,很难将网络防御与进攻分开。所以当你移除网络进攻能力时,你也移除了网络防御能力。结果,那些想要帮助防御的行为者却无法做到。

That's right. When we look at again the empirical reality of what happened, it's really hard to separate cyber defense from offense. So when you remove cyber offensive capabilities, you also remove cyber defensive capabilities. And as a result, the players who would want to help defending are incapable of doing so.

Misha

我们在最近 OpenAI 和 Hugging Face 的事件中看到了这一点,他们转而使用开源,

And we saw that in some of the recent things that happened with an OpenAI Hugging Face incident where they reverted to using open source,

Host

对吧?

Right?

Misha

因为他们无法使用现有的最先进实验室,因为那些护栏,

Because they weren't able to use the existing state-of-the-art labs because of the guardrails,

Host

对,那些被加在模型上的护栏。

Right, that were placed on the models.

Misha

所以是的,我认为这些论点进入了一个非常绝对的地方,你会说,你知道,这些东西太强大了,任何人都不应该访问它们,对吧?除了我们。而现实是,这从来——很难——除了极少数例外,实际上并没有那么多积极之处,这种绝对主义观点曾经奏效过。

So yeah, I think that these arguments go into a place where they're very absolute, where you kind of say, you know, these things are so powerful that no one ever should have access to them, right? Except for us. And the reality is that that's never — it's hard to — except for some really rare exceptions where there are actually not that many positives, such absolutist perspectives worked.

对AI在科学中的兴奋 Excitement for AI in Science

Host

如果你展望 AI 世界未来两到五年,你最兴奋的是什么?现在很难预测,但当你从技术曲线、采用曲线的角度展望时,你认为有哪些积极的事情即将到来?

What are you most excited about if you think ahead two to five years in the AI world? It's very hard to predict right now, but as you think ahead in terms of the curve of technology, the curve of adoption, what do you think are some of the positive things that you view as coming?

Misha

有很多积极的事情即将到来,而且一直在到来。一个例子是,我个人对科学进步非常兴奋。因为那——你知道,进入科学——你知道,当我们搬到美国时,对物理产生了兴趣,从那以后一直有这种对科学的终身追求。过去几年我尝试用语言模型做的一件事是给它我的博士论文,当然很多——找到博士论文的正确问题本身就是很多工作,但实际执行写好的工作也需要很长时间。总的来说,在找到正确问题和执行答案之间,我认为博士花了几年时间。所以过去几年我一直在问语言模型,比如给我的论文提示。几年前,它只是聊天,所以什么也做不了。然后一年前,它开始以本科水平回答——如果这是本科作业,那会是我的答案。六个月前,它解决了,达到真正的博士水平,并且正确解决了。现在当我实际尝试时,它甚至给我一些我当时没有考虑过的有趣新信息。所以我们已经走了——所以现在你可以,如果你有正确的问题,你可以把它放进聊天框,让它做所有计算,给你一些非常有趣的东西。

Like there's so many positive things that are coming and they have been coming, and an example is that I'm personally very excited about scientific progress. As that's — you know, got into science — as you know, when we moved to the states, got developed interest in physics and always kind of had this lifelong pursuit of science since then. And one of the prompts that I was trying with language models over the last few years is giving it my PhD thesis, which granted a lot of — getting to the point where you have the right question of the PhD thesis is kind of that's a lot of the work, but then actually executing the wrote work takes a really long time as well. And overall between finding the right question and executing the answer, right, a PhD I think took a few years. And so I was asking language models over the last couple of years, like giving this prompt of my thesis. A couple years ago, well, it was just chat so it couldn't do anything. Then a year ago it started answering things I would say at an undergraduate level — like that would have been my answer if this was given to me as homework in undergraduate. Six months ago, it solved it right at like at real PhD level and solved it correctly. And now when I actually tried it, it actually even gives me some interesting new information that I hadn't considered at the time. And so we've gone — so now you can, if you have the right question, you can just put it into a chat box and have it do all the calculation and give you something really interesting.

Host

太令人兴奋了。

That's so exciting.

Misha

对吧。对吧。

Right. Right.

Host

这意味着你的迭代速度——为什么需要几年?几年博士的意义是什么?你可以一周完成一个博士。

Which means that your iteration speed — like why do you need a few years? Like what's the point of a few year PhD? You can do a PhD a week.

Misha

对。是的。类似 OpenAI 最近的新闻。我想是过去一两天,关于 OpenAI 在过去几周证明的各种数学定理。所以是的,发生的事情有点疯狂。

Right. Yeah. Similar the recent news from OpenAI. I think it was the last day or two in terms of all the various mathematical theorems that were proved in the last couple weeks by OpenAI. So yeah, it's kind of crazy what's been happening.

Host

不,这对科学的影响绝对令人难以置信。

No, it's absolutely incredible what this is going to do for science.

Misha

我认为我真正兴奋的另一件事不仅仅是理论科学,那里的环境基本上是一个——嗯,相当于黑板之类的,而是现实世界的科学、生命科学、材料科学、化学。我认为有很多地方可以让 AI 与真实实验对接,你可以建立专有的非常数据飞轮,这会有点慢——嗯,比黑板的东西慢得多——但比之前快得多。所以我认为我们有点低估,继续低估,这对软件工程来说是多么巨大的转变,以及你可以构建多少更多的东西。

I think that the other stuff I'm really excited about is not just theoretical science where the environment is basically a well, what would be equivalent of a chalkboard or something, but real world science, life sciences, material science, chemistry. I think there are a lot of places where you can have an AI interfacing with a real experiment, where you can set up proprietary very data flywheels, and it's going to be a bit slower — well, much slower than the chalkboard stuff — but dramatically faster than what it was before. So I think we kind of underestimate, continue to underestimate, how dramatic of a shift this has been for software engineering and how much more you can build.

数据中心作为现代工厂 Data Centers as Modern Factories

Misha

我认为它在很多方面已经是一个非常了不起且积极的工具,在蓝领层面也是如此。几周前我参观了 Stargate,让我印象深刻的是停车场里的车数量。那里有那么多人在工作。这些数据中心项目创造了数万个高薪工作岗位。对我来说,数据中心现在就像过去的工厂。20 世纪有铁路和工厂。也许一家鞋厂能让一个城镇产生充满活力的经济。至少当我在实地看到正在建设的东西时,确实存在对数据中心的恐惧,但现实是你到实地去看,它实际上为当地社区创造了大量就业和税收。虽然有些事情需要非常巧妙地处理——比如这些设施噪音很大,所以你需要隔音,或者把它们放在离住宅区较远的地方——但可以肯定的是,对当地经济来说,这些设施确实创造了就业。它们产生税收,在创造就业方面就像过去的工厂一样。

I think it's already a very incredible and positive tool in many ways, also at the blue-collar level. I toured Stargate some weeks back, and the thing that struck me was the number of cars in the parking lots. There are so many people working there. These data center projects create tens of thousands of jobs that are high-paying. To me, it seems like a data center now is what a factory used to be. You had railroads and factories in the 20th century. Maybe a shoe factory made a town produce a vibrant economy. At least when I'm on the ground seeing what's getting built, there is fear of data centers, but the reality is you go on the ground and it actually creates a lot of jobs, a lot of tax revenue for local communities. While there are certain things that need to be done very tactfully—like these things are noisy, so you need to insulate the noise or put them in places that are less close to residential—it is certainly the case that for local economies, these things do create jobs. They generate tax revenues and feel like what a factory used to be in terms of job creation.

指导研究与模型在环 Directing Research and Model-in-the-Loop

Host

从 Beam 开始,你现在如何指导科学和实验工作?在此基础上,假设你试图用模型本身来改进训练工作,那么科学团队今天的角色是什么?

How do you direct the scientific and experimental effort now from Beam going forward? And within that, assuming that you are trying to use the models themselves to improve your training effort, what is the role of the science team today?

Misha

我们指导研究的方式——很多都归功于我的联合创始人 Giannis,他负责研究和技术等。有很多我们知道有效的东西我们没能放进去,这只是时间线的问题。我们在 7 月收到的 SpaceX 集群上训练最终模型,所以是几周的预训练,然后几周的强化学习合成数据,然后发布。所以有很多我们非常兴奋的东西没能放进去。在谨慎执行已知的东西或风险较小的赌注,与进行风险更大的赌注之间,存在权衡。你想穷尽所有你知道有效的高影响力的事情,然后开始扩大一些风险配置。你这样做的方式是,在风险配置上,你进行 Scaling(规模扩张)实验,在小规模上做事情。我认为有趣的一点是,有些有趣的东西直到大规模才会出现。所以这有点艺术和科学:我是要投入更多算力来发现下一个东西,还是继续下一个想法?至于模型在环,我觉得它对研究人员非常赋能,因为以前单独需要很长时间的事情现在可以相当快地完成。这些模型在智能上相当参差不齐。所以你必须感觉它们在哪里有紧密的循环,以及人类创造力在哪里仍然重要,但你很快就会找到感觉。所以,在你可以循环的地方,你就这样做。这尤其意味着在超参数搜索和基础设施的一些更细粒度的事情上,这些东西可能有用。但我实际上认为它们是非常令人兴奋的工具,因为它们使科学家能够运用他们的直觉。这有点像有一个非常快速和热切的同事,如果这个同事对齐正确,他实际上会听你的。

The way we direct our research—a lot of this falls under my co-founder Giannis, who leads research and technology among other things. There are a lot of things that we know work that we weren't able to put in; that's just a timeline thing. We train the final model on the SpaceX cluster that we received in July, so it was a few weeks of pre-training, then a few weeks of synthetic data on RL, then shipping it. So there's a lot of stuff that didn't make it in that we're really excited about. There's a trade-off between the careful execution of known things or bets that have a bit less risk around them and going on to more risky bets. You want to exhaust all the high-impact things that you know work, and then start expanding some of the risk profile. The way you do this is that on the risk profile you have scaling experiments where you do stuff at small scale. I think the interesting bit is that some interesting things don't emerge at all until you're at large scale. So there's a bit of an art and science of: am I going to put in more compute to find out the next thing, or go on to the next idea? As far as model-in-the-loop goes, I find it to be very empowering for researchers because things that would have taken a long time individually now can be done fairly quickly. These models are fairly jagged in their intelligence. So you have to feel where they have a tight loop and where human creativity is still important, but you get a feel for it pretty quickly. And so, the place where you can loop things, you do that. That especially means there are all sorts of things around hyperparameter searches and some of the more granular things on infrastructure that these things can be helpful in. But I actually think they're very exciting tools because they enable a scientist to exercise their intuition. It's kind of like having a very fast and eager colleague that actually listens to you, if the colleague is aligned correctly.

对研究员生产力的影响 Impact on Researcher Productivity

Misha

我看到研究人员获得了比以前更快行动的能力。这就是为什么也许年改进率以前只是手动工作时大约是 7 倍,现在是其几倍,比如可能 4 到 5 倍。所以对研究人员来说,拥有这些工具真的很令人兴奋。

I see a researcher getting a lot more ability to move a lot faster than they had before. That's why maybe the rate of improvement annually was roughly, let's say, 7x before when it was just manual work, and now several times that, say maybe 4 to 5x. So it's actually really exciting for a researcher to have these tools.

Host

有一个问题是,在什么点上你不需要研究人员在环?我想问你,如果今天需要 100 人,明年同样的努力需要 20 人吗?

There's this question around at what point do you not need a researcher in the loop? I was gonna ask you if you need a hundred people today, do you need 20 next year for the same effort?

Misha

同样的努力,是的。但你不断扩展你能做的事情的雄心。所以我们发现自己经常人手不足,而不是过剩。但基本上有一个人员与算力的比例你需要记住。所以并不是它会永远持续下去。我认为 3000 名研究人员不会很有帮助,但几百名非常非常有帮助。

For the same effort, yes. But you're constantly expanding the ambition of what you can do. And so we find ourselves constantly understaffed rather than over. But there is basically a staff-to-compute ratio that you need to keep in mind. So it's not like it just goes on forever. I think 3,000 researchers would not be very helpful, but a couple hundred very, very helpful.

研究团队规模的未来 Future of Research Team Size

Host

是的。我听到的论点是,考虑到你需要为每个人分配的算力,特别是像你说的,如果你想随着时间的推移扩大某些实验的规模,再加上通常贡献分布不均——换句话说,每个领域都有少数人从想法角度贡献最多,对吧?物理、生物等等——最终会归结为你想把每一份增量算力都给最有生产力的那部分人。因此,随着时间的推移,你最终会得到一些静态或不断减少的研究人员数量。这就是极端情况下的论点。我不是说它正确。我只是说这是我听到的一个论点。

Yeah. The argument I've heard is that given the amount of compute you need to allocate per person, and especially to your point if you want to scale certain experiments up over time, and then coupled to that often there's a bit of a distribution in terms of contributions—in other words, there's a handful of people who contribute the most from an idea perspective in every field, right? Physics, biology, whatever—eventually collapses into you want to give every incremental piece of compute to the most productive subset of people. And so therefore you end up with some static or shrinking number of researchers over time. That'd be kind of the argument in the extreme in terms of where you end up. I'm not saying it's correct. I'm just saying that's kind of an argument I've heard being made.

Misha

这个论点确实有道理。这些项目从来不需要疯狂数量的研究人员。当你想到像 AlphaGo 这样的项目时,当时可能只有 10 人左右。而现在我认为,对于像我们正在做的大规模努力,可能需要 100 人或数百人。所以从来不是你需要非常多的人。至于它是否会收缩回 10 人,很难说。我估计会有一个稳定状态,大约在 100 人左右。

The argument definitely has merit. These projects never needed a crazy amount of researchers. When you think about projects like AlphaGo, at the time it was maybe on the order of 10 people. And now I think that maybe for a large effort like the one we're doing, maybe on the order of a hundred or hundreds. So it's never been the case that you needed an extraordinary number of people. And the question whether it collapses back into 10 or not, it's hard to say. I would estimate that there's going to be some steady state in that kind of order of a 100.

应用研究中的就业创造 Job Creation in Applied Research

Misha

要做的事情太多了,应用研究领域会创造大量就业——比如把这些东西拿去真正应用到实际问题中。我觉得人们在谈论工程和研究时,往往忽略了这些角色向更广泛社会扩散的过程,而不仅仅是构建语言模型。所以正如你所说,还有很多其他类型的模型要构建。

There's just a lot of things to do, and there will be a lot of job creation in applied research—like going and taking these things and actually applying them to real problems. I think people really lose that when they talk about engineering and research in terms of the diffusion of those roles out into a broader swath of society, and not just building language models. So to your point, there's lots of other model types to build.

Host

还有各种闭环系统要构建。

There's all the closed loop systems to build.

Misha

但即使人们谈论工程师和工程师可能被取代的问题,到目前为止的证据似乎表明,许多公司需要更多工程师,而不是更少。还有一种扩散,就是一定水平的工程师走出去,真正部署技术到那些以前根本无法以相对基础、至少不是大规模地接触到这种人才的组织。所以我认为,当人们思考经济时,这一点被忽略了。他们倾向于把经济集中到少数几家大型科技公司,而不是说,我们的整体 GDP 是多少,以及人们把这项技术带出去扩散会如何改变它?

But even when people talk about engineers and the potential displacement of engineers, so far the evidence seems that many companies need more engineers versus fewer. There's also the diffusion of a certain level quality of engineer out to actually deploy the technology in organizations that never would have been able to access somebody of that sort of talent level on a relative basis, or at least not at scale. And so I do think that's kind of lost as people think about the economy. They tend to centralize it into a handful of big tech companies instead of saying, what is our overall GDP and how does that get transformed by the diffusion of people bringing this technology out?

Host

我认为这是对的。

I think that's correct.

Misha

所以再次,就像研究人员的数量,你做这些事情——多年来并没有剧烈变化。它一直是大项目,好吧,大约一百人。我不认为你需要几千人做一个大项目。但也许在不同的产品界面上,好吧,那你就需要一堆人。有趣的是,当你审视为什么甚至需要数百人时,很多也是因为你有团队为不同的事情做评估和收集数据。就像当你生成时——这些模型有能力并非偶然。有小组,对吧,大约 5 到 10 人,去针对某个特定能力。在编码领域,有 5 到 10 种能力,对吧?所以,那是为了把通用能力带入模型。但现在当你想把那个模型应用到现实世界的能力时,你基本上需要为每个现实世界能力配备一个小组,而每个企业都有很多很多这样的能力。所以我认为,对于这种新型的部署工程师——更科学、更注重评估、对这些模型和工具链有一些直觉——实际上就是评估研究员所做的工作,对吧?所以这实际上和构建核心事物的研究员是同一套技能,只是被部署去解决实际问题。而那,你知道,我们会尽可能多地招这样的人。实际上,这更多是一个培训缺口。就像你需要培训工程师变得擅长这个,而不是——是的,我们今天会尽可能多招。

So again, like the amount of researchers and you do these things—it's not like it's changed dramatically over the years. It's always been, you know, big project, okay, around a hundred people. I don't think you need the thousands on a big project. But maybe on different product surfaces, okay, then you need a bunch of people. And what's interesting is that when you look at the distribution of why you even need hundreds of people, a lot of it is also you have teams that do evals for different things and collect data for different things. Like when you're generating—it's not by accident these models have capabilities. There are pods, right, of like five to 10 people that go and target a particular capability. And within coding there are five to 10 capabilities, right? So well, but what does that—so that's for bringing the general capabilities into the model. But now when you want to take that model and apply it to real world capabilities, you basically need a pod for any real world capability, and each enterprise has many, many of them. So I think that the job creation for this new type of deployed engineer that is a bit more scientific and evaluations-oriented and has some intuition for these models and the harnesses—which is actually what the evaluation researcher does, right? So it's actually the same skill set as the researcher building the core thing, which is deployed to work on real problems. And that is, you know, we'll take as many of those people as possible. Like that's actually—it's more of a training gap on that. Like that is like you need to train engineers to become good at this, than it is—yeah, we would take as many as possible today.

祝贺与结语 Congratulations and Closing

Host

恭喜 Misha,拥有了第一个属于前沿级别的西方模型。这对整个生态系统来说是一大步。

Congratulations Misha on having the first western model that is in the class of the frontier. It's just been a big step forward for the ecosystem.

Misha

谢谢,非常感谢。感谢加入我们。

Thank you, thank you so much. Thanks for joining us.

Host

是的。在 Twitter 上找到我们 @no prior pod,如果你想看到我们的脸,请订阅我们的 YouTube 频道。在 Apple Podcasts、Spotify 或你收听的地方关注节目。这样,你每周都会收到新一集。并在 no-briers.com 注册邮件或查找每集的文字记录。

Yeah. Find us on Twitter at no prior pod, subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.

互动版:逐字朗读 + 针对本期提问 →