Learning from Experience: Beyond Reinforcement Learning and Transformers
打开互动全文版(中英对照 + 朗读 + 问答)→Jerry 和 Rohan 讨论如何从经验中学习不仅限于强化学习,以及当前 Transformer 架构为何可能成为 AI 进展的瓶颈。
Jerry and Rohan discuss how learning from experience extends beyond reinforcement learning, and why the current transformer architecture may be a bottleneck for AI progress.
比如,我踢足球的时候,这看起来跟强化学习非常相似。我得到很多次反馈,每次都会稍微调整一下,然后看看结果是否大致符合我的预期,这里面有一种自我强化在发生。但当我学数学的时候,情况就完全不同了。那更像是阅读困难的概念,然后在脑海里深度思考,直到一切豁然开朗,直到我把它们联系起来。这两种方式在某种程度上都是从经验中学习,但它们非常不同。我们可能正在投入比以往更多的算力来从经验中学习,但强化学习并不是从经验中学习的终点,未来几年研究人员会想出更好的方法来利用这些数据。
If I play football, for example, it looks very closely to reinforcement learning. I get a lot of times and every time I adjusted a little bit and I see if it roughly matches what I wanted and there are some self-reinforcement happening. When I learn mathematics, it's a very different type of thing. It's like reading about hard concepts and thinking about them very deeply inside my head until things click and until I have them connected. And both of those in some way are learning from experience. They are just very different. We probably are spending the most compute than ever on learning from experience, but reinforcement learning is not the end of learning from experience and there will be better approaches that researchers will be coming up with in the coming years on how to use that data.
Jerry、Rohan,非常感谢你们今天来做客。两位是 Core Animation 的创始人,这是目前旧金山最炙手可热的实验室之一。在创办 Core Animation 之前,你们领导了 AI 时代一些最重要的研究项目。Jerry,你曾是 OpenAI 的 VP,负责过 strawberry 和推理团队等工作。Rohan,你曾是 Gemini 的四位预训练负责人之一,在此之前,你在 Google Brain 领导了大量基础 AI 研究,并且是谷歌内部以及后来在 Anthropic 的 fix-it 负责人。所以,你们两人在前沿研究方面都比常人见识了更多。我非常期待深入探讨。嗯,先从你开始,Jerry。你最近在推特上发了一个很大胆的观点:'取代 Transformer 的第一步是深刻理解它们把我们带到了多远。' 这是对 Transformer 的悼词吗?这是什么意思?
Jerry, Rohan, thank you so much for joining us today. The two of you are the founders of Core Animation, one of the hottest new labs in San Francisco right now. And before starting Core Animation, you led some of the most important research projects of the AI era. Jerry, you were VP at OpenAI, where you worked among other things on running the strawberry and reasoning teams. And Rohan, you were one of the four pre-training leads at Gemini. And before that, led a lot of the fundamental AI research at Google Brain and were the fix-it guy across Google and then at Anthropic. And so, between the two of you, you've seen more than your fair share of what the world looks like in terms of doing frontier research. And so, I'm very excited to dig in. Um, let's start with you, Jerry. You tweeted a very spicy take recently. 'The first step to replacing transformers is appreciating deeply how far they were able to carry us.' Is that a eulogy for the transformer? What does that mean?
非常感谢您邀请我们,Sonia。我觉得我最近的采访很多都是在解释我的推文。我的意思是,欣赏 Transformer 就是理解它擅长什么。所以,如果你不去解决它已经很好解决的那些问题,你就得关注它的弱点。你必须了解好的部分和坏的部分。很多人在架构方面的工作很容易就是试图让 Transformer 变得更便宜,试图让它更高效。我很少见到有人思考如何让 Transformer 更强大,更有表现力。但看到一个人的弱点和看到一个人的优点几乎是同一件事,只是更理解 Transformer 的形状。我认为我们现在正处在这个阶段。我们已经非常擅长训练非常大的模型。我们掌握了两个算法:大规模预训练和大规模强化学习。我经常问自己机器学习下一步是什么。我认为现在更好的模型和更智能系统的瓶颈在于架构本身。现在是时候重新审视我们过去 6 年一直在坐的这趟列车——试图在本质上相同的两个操作(MoE 和注意力机制)上不断堆参数。当我在思考我们现在的位置和我们正在做的事情时,我经常思考 Codex 和 Claude Code 为我们做了什么。我非常感激这些系统,以及编码、工作流自动化和我们过去 6 年通过 Scaling 构建的所有产品和系统。我认为这是思考的第一步。如果我们想要研究替代方案,我们需要看到我们现在在哪里,解决了什么问题,然后才能看到下一阶段是什么,还有什么问题没解决,我们缺少什么。每当我使用 Codex 成功完成一个任务时,我也会开始想,为什么我不试着把它推得更远?每当我来上班时,我有很多事情用 Codex 来做,但我仍然来上班,我仍然请它帮我做某些事情。我总是在问自己,为什么我还要被需要?为什么 Core Animation 的名字和概念是我们要自动化任务?为什么那些事情还没有自动化?为什么 Codex 不能为我做所有的事情?这就是研究的问题所在,我们要走向哪里。通过这些研究,我试图思考我们需要什么样的模型、什么样的系统、什么样的品质,而今天我们还没有。这就是我最近经常思考的事情。
Thank you very much for inviting us here, Sonia. I feel like a lot of my interviews these days is explaining my tweets. And what did I mean? But appreciating Transformer means understanding what it does well. So, if you're not solving the problems that it is solving well, you have to focus on its weaknesses. You have to understand good parts and bad parts. And it's very easy in a lot of the work what people are doing in architectures is trying to make Transformers cheaper. They are trying to make Transformer more efficient. I very rarely see people thinking about how do we make Transformers more powerful, trying to do more express. But seeing someone's weak parts and seeing someone's strong parts are almost the same thing. It's just understanding the shape of Transformer a little bit more. What I think right now we are in this stage. We got really good at training really big models. We mastered two algorithms. We mastered pre-training at a large scale. And we mastered reinforcement learning at a large scale. And I'm asking myself a lot what is next in machine learning. I think at this moment what the bottleneck is to better models and to smarter systems is the architecture itself. It is this moment to revisit the train we've been riding for the last 6 years of trying to add more and more parameters to essentially two of the same operations, which is MoE and attention. And when I am thinking about it, like where we are today and what we are doing, I am thinking a lot about what Codex and what Claude Code are doing for us. And I'm really appreciative of those systems and of the coding and of the workflow automation and of the systems and of the products that we have today that we essentially have built over those 6 years of scaling. And I think this is the first step of thinking. Like, if we want to work on the replacement, we need to see where we are, what problems we have solved to start seeing what the next stage is, what problems we haven't solved yet, what we are missing. And this is whenever I use Codex and I am successful at a task, I also start thinking, why didn't I try to push that thing harder? Whenever I come to work, there are a lot of things I do with Codex, but I still come to work. I still ask it to do certain things for me. And I'm always asking myself, why am I even needed there? Why is Core Animation's name and its concept is we want to be automating tasks. And why are those things not yet automated? Why is not Codex doing everything for me? And this is the question of the research, where we want to go. And with that research, I'm trying to think, what kind of models, what kind of systems do we need, what kind of qualities do we need that we don't have today. And that's what I'm thinking a lot these days.
而且你的基本前提是架构是问题所在,我认为这是一个逆向观点。那么,是什么让你得出这个观点?你看到了什么让你觉得架构是问题所在?
And you have the starting premise that the architecture is the issue, which I think is a contrarian point of view. So, what led you to that point of view? What did you see that made you think the architecture was the issue?
这从根本上就是问题所在。它来自于之前的暗示。我认为问题是模型在实验室里训练,然后在现实世界部署。这是那里存在的基本张力。我的一些失望来自于我个人的经历。当我们在 OpenAI 开始扩展强化学习的研究和进展时,我从在 OpenAI 工作开始就基本上相信,扩展强化学习是通往 AGI 道路上的必要跳板。我一直是一个强化学习最大化者。我一直相信这是我们需要关注的重点,这是我们需要做的。我看到了 LLM 通过 GPT-3 到 GPT-4 被扩展到越来越高的水平,而我们仍然在做很少的 RL。我内心一直相信,一旦我们开始扩展 RL,我们就会解决一切,能够解决所有问题。我们最终开始扩展 RL。我当时就在那里,处在中心。我在想,我们来了。如果你在 2024 年问 Jerry,我们什么时候能得到 AGI?我会说 2025 年就是那一年,这就是我们解决一切的时刻。然后我看到我们一个接一个地训练模型,模型越来越好,所有的基准分数都在上升。但我们在那个时刻也解决所有现实世界的任务了吗?不幸的是,并没有。我们仍有工作要做。
It's fundamentally what is the issue. It comes back from the previous implication. What I think is the issue is that the models are being trained in the lab and are being deployed in the real world. That is the fundamental tension. And a bit of my disappointment came from my personal story. Whenever we were starting the research and progress on scaling up reinforcement learning at OpenAI, I basically believed that scaling up reinforcement learning is a necessary stepping stone on a path to AGI since I started working at OpenAI. And I was always a reinforcement learning maximalist. I always believed this is what we need to focus on. This is what we need to do. I've seen LLMs being scaled up to higher and higher levels through GPT-3 to GPT-4. And we're still doing very little RL. And I had this internal belief that the moment we start scaling up RL, we'll solve everything. We'll be able to solve all the problems. And we eventually started scaling up RL. I was just in there. I was in the center of it. I was thinking here we are. If you ask Jerry in 2024, when do we get AGI? I would say 2025 will be that year. This is where we solve everything. And I saw us training model after model. The model was getting better and better. All the benchmark scores were going up. And did we also solve all the real-world tasks at that moment? Unfortunately, unfortunately not. We still have work.
我意识到有一个区别:我们用来评估模型的所有基准,本质上与训练它们用的任务是一样的。评估和训练任务是同一枚硬币的两面。但现实世界的分布和任务要混乱得多、模糊得多、差异也大得多。我们的训练数据并没有真正复现现实世界的用例。尽管我们基本上把所有任务都最大化了,但如果你问任何训练模型的人‘嘿,你们的主要问题是什么?’,他们会说‘我没有足够难的任务。我没有东西可以训练我们的模型。’然而我们仍然没有覆盖现实世界的全部分布。由此我得出结论:我们需要在测试时学习的模型。我们需要能在用户的数据上、在他们的现实任务上、在现实世界分布上学习的模型。当你问为什么我们今天没有这个,为什么 Transformer 不在任何地方学习时,其实我们可以在测试时进行两种学习。第一种是 Transformer 的上下文内学习,它没有灾难性遗忘的根本问题,数据效率很高,这很棒,但它不太可扩展。我们只能有有限的量。它有限,并且有更多的机械限制,关于你在构建上下文时实际在做什么。但也许我们稍后可以再谈。上下文内学习非常有限,使用的数据量非常小。每当我使用 Codex,大约使用 20 分钟后,我就需要压缩它并移走。那数据量不大。如果我们只能学习 20 分钟,那并不多。第二种是微调。我们可以尝试持续微调模型,但那样会有灾难性遗忘和数据效率极低的问题。这两种都不太好解决。人们一直在尝试。如果它们容易解决,早就有人解决了。所以我个人认为,我们需要找到一种可以元学习的算法,可以在架构层表达,能代表学习是什么样的。能处理更长视野的学习是什么样的?
And I realized there was a bit of distinction: all the benchmarks we use to evaluate our models are essentially the same as what we train them on. The evals and training tasks are two sides of the same coin. But the real-world distribution and real-world tasks are much messier, much murkier, much more different. Our training data didn't really replicate real-world use cases. And despite us basically maximizing all the tasks, if you ask anyone training models, 'Hey, what is one of your main issues?' they'd say 'I don't have hard enough tasks. I don't have what to train our model on.' Yet we are still not covering the entirety of the real-world distribution. From that, my conclusion is we need models that learn at test time. We need models that learn with users on their data, on their real-world tasks, on the real-world distribution. And when you ask why don't we have that today, why are transformers not learning anywhere? There are essentially two types of learning we could do at test time. First, in-context learning of transformers, which doesn't have fundamental problems of catastrophic forgetting. It is pretty data efficient, so that is great, but it's not very scalable. We can only have so much of it. It is limited and has even more mechanical limitations of what you are actually doing when you build context. But maybe we can come back to it later. In-context learning is very limited and uses a very small amount of data. Whenever I'm using Codex, after roughly 20 minutes of usage, I need to compact it and move it afterwards. That's not that much data. If all we can learn is for 20 minutes, it's not much. The second thing is fine-tuning. We could try to continuously fine-tune our models, but then we have issues of catastrophic forgetting and very low data efficiency. Neither of those are very solvable. People have been trying. If they were easy to solve, someone would have solved it already. So my personal belief is we need to find an algorithm that we can meta-learn, that we can express on the architectural layer, that represents how learning looks. How does learning look that can work on much longer horizons?
嗯。你预计架构会看起来像 Transformer 吗?因为我从高层获得的理解是,OpenAI 长期以来一直试图扩展强化学习,直到 Transformer 出现,才似乎有了一个可扩展的世界先验,可以用来扩展强化学习。那么,你如何去思考扩展这个新范式呢?
Mhm. Do you expect the architecture will look transformer-like? Because my understanding from the chief seats is that OpenAI has been trying to scale up reinforcement learning for a long time, and it wasn't until the transformer came about that it seemed like there was a scalable prior on the world upon which to even scale RL. And so, how do you even go about trying to think about scaling up this new paradigm?
这是个好问题。我认为这两件事是同时发生的,但要说的话,主要是经济原因,因为技术上你可以扩展 LSTM。只是没人真的敢往那个方向走。而且它们的扩展性差得多。在缩放损失论文中,有 LSTM 和 Transformer 的比较。从根本上说,Transformer 的缩放定律更好。存在一个我们没有发明 Transformer 的世界,我们就会扩展 LSTM,并且会有一些模型。但由于它们训练成本更高,产品表现也不那么令人印象深刻,我们只会得到更差的体验,也许没人能说服人们花那么多钱训练那些巨大的 LSTM,因为我们得不到市场回报。Transformer 的宏伟之处——回到为什么我们必须深深欣赏它们——在于 Transformer 在经济上是有价值的。训练它们的成本低于它们产生的收入,这是机器学习的魔力,并非理所当然。但对于 LSTM,很可能不是这样,这也许会发生。但你可以通过许多方式扩展大多数架构。我认为过去人们不扩展事物的很多原因,是因为 OpenAI 之前的研究人员对扩展非常抵触。它常被视为不科学,算法研究关注的是如何变得越来越高效?如何用同样的算力预算获得更好的结果?而 OpenAI 当时的一个赌注是:‘嘿,我们不在乎更好的算法。我们在乎的是更小、更可扩展的算法。’以及如何投入更多算力并获得更好的结果?OpenAI 长时间被社区里很多人反复批评。但多亏了这一点,我们才有了今天的模型。我认为有大量架构可以被扩展。我是核心自动化使命的一部分,我们的信念是很多架构研究在太小规模上进行了太长时间。很多人试图说:‘嘿,让我们先在小的数据集、小的算力范围内尝试我们的架构,然后只在它证明自己之后再看它如何扩展。’但例如,当你做强化学习时,你知道要得到任何有趣的结果,你需要一定水平的算力才能看到模型的能力。强化学习需要一个基线能力才能开始工作。所以在我看来,可能有很多架构需要算力的基线才能开始做任何有趣和有用的事情。
It's a great question. I think those two things happened at the same time, but if anything, it was mostly about economics, because technically you can scale up LSTMs. Just no one really dared to go in that direction. And they did scale much more poorly. In the scaling loss paper, there is a comparison of LSTMs and Transformers. And fundamentally, the scaling law of Transformers was better. There is a world where we never invented Transformers and we would be scaling LSTMs and we would have some models. But because they would be much more expensive to train and much less impressive as a product, we would have just a worse experience and maybe no one would be able to convince people to spend as many dollars training those gigantic LSTMs because we wouldn't get a market return. The majestic thing about Transformers, which goes back to why we have to appreciate them so deeply, is that Transformers are economically valuable. The cost of training them is lower than the revenue they generate, which is the magic of machine learning and not guaranteed by itself. But for LSTMs, it probably wouldn't be that way, which may have happened. But you can, in many ways, scale most architectures. I think a lot of reasons why people didn't scale things before was because researchers before OpenAI had a lot of reluctance to scaling. It was often seen as unscientific, and research in algorithms was providing how do we become more and more efficient? How do we for the same compute budget get better results? And it was a bit of a bet by OpenAI at that moment: 'Hey, we don't care about better and better algorithms. We care about smaller and more scalable algorithms.' And how do we pour more and more compute and get better results? OpenAI was criticized repeatedly by many people in the community for a long time. But thanks to that, we have the models we have today. And I think there are tons of architectures that can be scaled up. And I am part of the core automations mission, and our belief is that a lot of architectural research happened at too small scale for too long. A lot of people are trying to say, 'Hey, let's first try our architecture on a small data set in a small compute regime, and then see where it scales only after you prove itself.' But for example, when you do work on reinforcement learning, you know that to get any interesting results, you need a certain level of compute to even see the capabilities in a model. Reinforcement learning needs a baseline of ability to just start working. So where I am coming from, probably there are many architectures that need a baseline of compute to even start doing anything interesting and useful.
那我能不能问一个可能比较敏感的问题?
Can I ask you then maybe a touchy question?
请讲。
Please do.
如果你需要算力的基线,这听起来像是在大型研究实验室里很适合做的工作。为什么要创办一家公司来做这件事呢?
If you need a baseline of compute, that sounds like a job that would be well served inside of a big research lab. Why start a company to go do this?
这是个好问题,我认为在很多方面,这可能是一个时机问题。市场目前处于一个非常特殊的位置,最大、最成功的实验室,无论巧合还是命运,可能正处于有史以来最激烈的市场竞争中,这使得它们不太愿意尝试不同的路径、尝试替代方案。
It's a great question, and I think in many ways it's likely a timing issue. The market is right now in a very specific place where the biggest and most successful labs, by coincidence or by fate, are probably in the most competitive market fight ever right now, which makes them not very keen on trying different paths, trying alternatives.
如果 Transformer 能盈利,并且你可以花更多精力和资源将其规模化以赢得下个季度,就很难再把大量注意力和精力投到那些可能在一两年后更好或重新定义这个领域的东西上。我认为最大的实验室——我基本上都聊过——对尝试 Transformer 的替代方案兴趣不大。较小的实验室则尽其所能模仿最成功的实验室。每个人都在做同样的事情:智能编码智能体。看看上周的发布——大家都在发布编码智能体。我们需要不同的路径和方法。这就是我们在生态系统中试图填补的 niche。
If Transformer is profitable and you can spend more effort and resources scaling it to win next quarter, it's very hard to put attention and energy into something that might be better or redefine the field in a year or two. I think the biggest labs — and I've talked to basically all of them — don't have much interest in trying alternatives to Transformer. The smaller labs are doing whatever they can to copy the most successful labs. Everyone is trying the same thing: smart coding agents. Look at last week's releases — everyone released a coding agent. We need different paths and approaches. That's the niche in the ecosystem we're trying to fill.
嗯。你在 Transformer 被发明的时候就在 Brain。你同意 Jerry 对 Transformer 的告别辞吗?
Mhm. And you were at Brain when the Transformer was invented. Do you agree with Jerry's eulogy for the Transformer?
是的,从某种角度来说。当 Ashish、Noam 等人提出它时,我同时在研究在线蒸馏,我们在同一个内部研究会议上发表。内部并没有引起太大轰动——只有少数人真正理解。很多人认为这只是又一篇工作。最初的工作非常专注于翻译,并在翻译上超越了 LSTM。在 Google 内部,Noam 和其他几个人确实对规模化语言模型感兴趣。直到 GPT-2 和 GPT-3 才看到 Transformer 的好效果。我思考架构的方式是我们如何分配算力。Transformer 是一种非常高效的分配方式,但现在大部分算力用于推理,花在 token 上。如果我要优化更好的架构,需要同时考虑预训练和强化学习,找到比当前思维链 token 生成更好地分配算力的架构。预训练构建了一个具有特定上下文长度的 Transformer。然后强化学习介入,说这不够——我需要更多算力,所以一次添加一个 token。从推理角度看效率低下。我们一次生成一个 token,大多数解决方案都是创可贴式的,比如投机性解码。自回归解码本身就有问题。很长一段时间里,大多数人在训练非常大的密集模型;行业花了两到三年才完善到我们现在习以为常的架构。稀疏性和混合专家模型并非显而易见。那么 Transformer 有什么问题?计算深度很差。如何增加计算深度?提出这个问题就打开了 20 个新方向。做这样的工作需要时间——基础研究通常需要五到六年才能落地工业界,很大程度上取决于组织是否相信其重要性。Jerry 有内在信念认为强化学习是必需的;我在 Google 没有这种信念。我曾经是预训练最大化主义者。预训练一个更大但更小的模型——这样好吗?对吧?所以内在信念和架构在硬件上的高效运行都是必需的。理论上最优的架构在投入生产、写好 kernel 和一切之前毫无用处。现在只有少数地方拥有整合团队做这件事。我们组建了一个团队,让专家们端到端地共同审视问题。这就是我乐观的原因。当前的机制很差;如果放任不管,恐怕需要更长的时间才能替换 Transformer。很多人已经在抱怨 token 成本。我来自 Google 的服务数十亿人的 mindset,所以我们需要更高效的架构,满足延迟和 token 服务量的限制。前沿技术目前只与一部分人相关。我们正在抓住这个机会加速这一过程。
Yes, in some sense. When Ashish, Noam, and others came up with it, I had work on online distillation at the same time, and we presented at the same internal research conference. It wasn't a big deal internally — only a few people actually got it. Many thought it was just another piece of work. The original work was very focused on translation and beat LSTM there. Internally at Google, Noam and a few others were definitely interested in scaling language models. It took until GPT-2 and GPT-3 to see the benefit of Transformers working well. The way I think about architecture is how we spend computation. Transformer is one very efficient way, but now most computation is inference time spent on tokens. If I want to optimize for a better architecture, I need to look at both pre-training and RL together to find architectures that spend computation better than current chain-of-thought token generation. Pre-training builds a Transformer with a certain context length. Then RL comes in and says that's not sufficient—I need more computation, so I add one token at a time. This is inefficient from an inference perspective. We're doing one token at a time, and most solutions are band-aids like speculative decoding. Autoregressive decoding is a problem. For a long time, most of the world trained very large dense models; it took the industry two to three years to refine the architecture to what we now take for granted. Sparsity and mixture of experts were not obvious. So what's wrong with Transformer? Computational depth is poor. How do we increase computational depth? Posing that question opens up many directions. Doing this work takes time — fundamental research takes five to six years to land in industry, largely because organizational belief is crucial. Jerry had the inner belief that RL is needed; I did not have that at Google. I was a pre-training maximalist. Pre-train a bigger, smaller model — are they a good fit? Right? So inner belief and architecture efficiency on hardware are both needed. A theoretically optimal architecture is useless until it's productionized, with kernels and everything. Few places have integrated teams doing that. We built a team where experts work together holistically from end to end. That's why I'm bullish. Current mechanisms are quite poor; if left to the world, it will take much longer to replace Transformer. Many already complain about token costs. I come from the Google mindset of serving billions, so we need more efficient architectures with latency and token serving deadlines. Frontier tech is still only relevant to a subset of humans. We're taking a shot at accelerating that.
对吧?好。
Right? Like all right.
所以这种内在信念以及第二点——你需要架构在硬件上高效运行。理论上最优的架构在实践中才有用。所以你需要从研究初始就将其产品化,写 kernel 和各种代码,形成端到端循环。现在只有少数地方有整合团队做这件事,我认为我们建立的团队让专家们一起工作,而不是各自为政——我们让所有人端到端地整体看待问题。这正是我相当乐观的地方。这就是我在这里的原因。当前的机制很差,如果任其发展,我担心需要更长时间才能替换 Transformer。很多人已经在抱怨 token 成本。我来自 Google 的服务数十亿人的 mindset,所以找到更高效的架构来满足延迟和 token 服务量的截止时间至关重要。当我观察世界时,能使用前沿技术的人非常少。必须有人或某个群体加速这一进程。我们正在抓住这个机会。当前技术即使规模化,也仍然只与一部分人相关。
And then so that inner belief and second is you need your architecture to run efficiently on hardware. A theoretically optimal architecture is not useful to anyone until it comes into practice. So you need the research inception to getting it productionized and getting kernels and everything written. The end-to-end loop. There are only few places right now which have integrated teams doing that, and I think we have built a team in a way that puts the experts together not in different silos — we are accelerating on having everybody look at the problem holistically from end to end. That's where I'm quite bullish. That's why I'm here. The current mechanisms are quite poor and if left to the world, I am afraid it will take much longer before we replace the Transformer. A lot of folks are already complaining on token costs. I come from the Google mindset where we had to serve billions of people. So finding more efficient architectures that fit latency and token serving deadlines is crucial. When I look at the world, the amount that can use frontier tech is very little. Someone or some group has to accelerate this. We are taking that shot. The current technology just scaled up is still only relevant to a subset of humans.
所以你提到的一个问题是 Transformer 的计算深度很差。
So one of the things I heard you say was the problem with transformers is the computational depth is poor.
是的。我可以给你一个观点。我们训练的大多数 Transformer 都很浅,最多大约 100 层。深度之所以叫深度学习,是因为你想要更深的表示。还没有人展示过学习极深表示。思维链推理和通过强化学习让模型自己进行思维链是增加计算深度的一种方式,因为每增加一个 token 就增加一条路径。这样你就可以摆脱预训练架构设定的瓶颈——只能做层数乘以序列长度。现在你可以增加序列长度,得到更强的结果。你可以进行推理时 Scaling。
Yeah. I can give you one insight. Like most transformers that we train are quite shallow. It's at most like 100 layers deep. Depth is called deep learning because you want deeper representations. No one has actually shown us learning extremely deep representations. Chain-of-thought reasoning and RL to do chain of thought on the model itself is one way to increase computational depth, because every token you add adds one more pathway. Then you can get out of the bottleneck that the pre-trained architecture set you up on — you can only do layer number of layers times sequence length. Now you can increase the sequence length and get much stronger results. You can do inference time scaling.
现在,推理时 Scaling 的问题在于模型必须生成更多 token 才能获得更好结果,一次一个 token。由此你可以看到,我们可以直接解决许多这类问题,而这正是我们关注的一部分工作:让这变得更高效。
Now, the issue with inference time scaling is that models now have to produce more tokens to get better results, one token at a time. From this you can see we can directly address many of these things, and this is a subset of work we're looking at: making this much more efficient.
你对基于 Transformer 的架构有什么预测?如果它不是最终状态,它能带我们走多远?我们什么时候开始看到它到达瓶颈?
What's your forecast for the transformer-based architecture? If it's not the end state, how far can it get us? When do we start to see it topping out?
我认为这都归结为我们训练 Transformer 的目的以及我们能拿它做什么。我们做预训练,它非常擅长将互联网上的所有知识蒸馏到 Transformer 中,然后我们可以应用强化学习,这基本上把我们想要的所有工作流都烘焙进 Transformer。所以,Transformer 的极限在于:我们把人类的所有知识以及它们之间的关系、如何协同、如何组合都放进了模型。任何我们有训练数据的任务,我们都可将其放入模型——这可以是一个用大量算力在全球所有数据上训练的巨型模型。但如果我们停止训练那个模型,会发生什么?值得一问的问题是:如果 OpenAI 和 Anthropic 停止训练新模型会怎样?我们得到今天这个 Transformer,并说就是它了,我们最好的模型。几个月过去,几年过去,模型变得越来越没用。也许实验室记录了地球上每个人的所作所为、他们的任务和环境,并把它们放入学习环境。但如果任何东西发生了变化——新事件、新关系、新任务、新代码库、新工具——Transformer 的大部分效用来自训练中出现的东西。当这些东西缺失时,它们就会表现不佳。它们有一定适应能力,但不大且不灵活。所以在我看来,这就是 Transformer 的瓶颈水平,在很多方面我认为它是我们的工具。如果一个人知道 Transformer 的局限性,他们可以调度模型,为任务编写提示词。通过我们正在做的训练,我们可以在该任务上成功。模型失败的任何任务都会被加入训练数据,模型就能成功。但那个循环必须通过实验室为你训练模型。如果模型从根本上需要在实验室中训练,你多大程度上认为这是目标?或者你希望能够在不返回实验室的情况下更新模型?
I think it all comes back to what we are training transformers for and what we can do with them. We're doing pre-training, which is very good at distilling all the knowledge from the internet into transformers, and then we can apply RL, which basically bakes all the workflows we want into a transformer. So, where transformer has capped out is: we have all the knowledge of humanity in the model together with the relationships and how they work together, how they can be combined. Any task that we have training data for, we can put into the model—this can be a gigantic model trained with a lot of compute on all the data in the world. But then if we ever stop training that model, what would happen? The question worth asking is: what would happen if OpenAI and Anthropic stopped training new models? We get the transformer we have today, and say this is it, the best model we have. Months pass, years pass, and the model becomes less and less useful. Maybe the lab recorded every human on Earth, what they were doing, their tasks and environments, and put them into the learning environment. But if anything changes—new events, new relationships, new tasks, new codebases, new tools—transformers get most of their usefulness from things present in training. When those things are absent, they suffer. There's some ability to adapt, but it's not very big and not very flexible. So in my mind, that's the level where transformers top out, which in many ways I think is a tool for us. If a human knows the limitations of a transformer, they can schedule the model, write a prompt for the task. By doing the training we are doing, we can succeed in that task. Any task the model fails at gets added to training data, and the model can succeed. But that loop has to go through the lab training the model for you. If the model fundamentally needs to be trained in the lab, how much do you think this is the goal? Or would you want to be able to update the model without going back there?
你读过 Rich Sutton 和 David Silver 的那篇论文《经验时代》吗?你有多大程度同意,或者你的观点在哪些地方不同?
Have you read the Rich Sutton and David Silver paper 'The Age of Experience'? How much do you agree, or where do your opinions diverge?
强化学习并不是一种特别新颖的方法。所以在某种程度上,经验时代一直存在。人们一直在批评预训练,因为它使用主要由他人生成的静态数据。尽管我个人认为,如今的预训练很大程度上是将其他模型蒸馏到新模型中,因为互联网上的大多数 token 来自 AI。但显然预训练是行为克隆、模仿、对互联网数据的压缩。强化学习并不是人们没有想过或没有做过的事情。它早在当年就被用来解决西洋双陆棋,然后是围棋、星际争霸、Dota,现在用来解决编程问题。每次都归结为模型自行编写经验并从中学习。但这一点很清楚,我认为有趣且仍令人困惑的是,强化学习并不是从经验中学习的唯一方式。还会有更多——我认为可以称之为算法创新,即我们如何从经验中学习。仅仅因为强化学习是一种方式,一种数学公式,特别是我们当前使用它的方式,它非常喜欢使用并行展开来减少方差,并在世界的并行版本中比较模型的表现,这不是我们人类从经验中学习的方式。我们更高效、以更多方式从经验中学习。我曾试图向人们解释大脑是如何运作的以及我们如何学习。大脑里有一种学习算法?我认为实际上有多种,它们协同工作。例如,当我踢足球时,看起来很像强化学习:我多次踢球,每次稍微调整,看是否大致符合预期,然后学习。但当我学习数学时,就完全不同了:阅读难懂的概念,在脑海中深入思考,直到豁然开朗并建立联系。两者都是从经验中学习,但非常不同。所以总结我的想法:我们从经验中学习已经有一段时间了。我们可能比以往花费了更多的算力在从经验中学习上,但强化学习并不是从经验中学习的终点。未来几年,研究人员将提出更好的方法,用于在更丰富的新图表设置中利用那些数据。
Reinforcement learning is not a particularly new approach. So, in some way, the age of experience has always been there. People have been criticizing pre-training because it uses static data mostly generated by others. Although I have a personal view that pre-training today is largely distilling other models into the new model because most internet tokens come from AI. But clearly pre-training is behavioral cloning, mimicry, compression of internet data. Reinforcement learning is not something people haven't been thinking about or doing. It was used to solve backgammon back in the day, then Go, StarCraft, Dota, and now programming. Every time it comes down to the model writing its own experience and learning from it. But this is clear, and I think what is interesting and still perplexing is that reinforcement learning is not the only way to learn from experience. There will be more—I think you can call it algorithmic innovation of how we learn from experience. Just because reinforcement learning is one way, a mathematical formulation, and especially how we use it now, it really likes parallel rollouts for variance reduction and comparing the model in parallel versions of the world, which is not how we learn from experience. We learn from our experience much more efficiently and in many ways. I've tried to explain to people what the brain does and how we learn. There's one learning algorithm in the brain? I think there are multiple, actually, and they work together. For example, when I play football, it looks very closely to reinforcement learning: I kick the ball many times, adjust a little each time, and see if it matches what I wanted, and I learn. But when I learn mathematics, it's very different: reading hard concepts, thinking deeply in my head until things click and connect. Both are learning from experience, but very different. So summarizing my thinking: we've been doing learning from experience for a while. We're probably spending more compute on it than ever, but reinforcement learning is not the end of learning from experience. There will be better approaches researchers will come up with in the coming years on how to use that data in richer new chart settings.
谢谢。一样。我是 Rohan。
Thank you. Same. I'm Rohan.
我很好奇,既然你的很多工作都围绕优化和效率展开,我们要如何实现数量级更高效的、算力和数据效率都更高的学习算法?
I'm curious since a lot of your work has been around optimization and efficiency, how do we get to orders of magnitude more compute efficient and data efficient learning algorithms?
你知道,我会从度量开始。我认为当前定义的预训练就是压缩。我们看困惑度,然后测量如何降低它,发现 Scaling、增加参数数量和投入更多算力是方法。每次我们在对数尺度上增加算力,这些指标就会获得微小的改进。我认为这对构建先验没问题,但这是看待问题的错误方式。我们应该端到端地看:我们训练这些模型是为了什么?看最终结果。比如,我训练一个模型交给 Jerry,Jerry 会做 RL 并毁掉我创建的所有困惑度指标。所以这是我们迄今为止最好的办法,我认为实验室们做得很出色,产出了非常有价值的智能。但这只是启动过程。我们必须把预训练和 RL 结合起来——这就是一个数量级改进的来源,它本质上是一种训练过程,也可以说是一种学习算法。就优化而言,我的故事始于 2016 年在 Google 优化逻辑回归。我从事过 Sybil 的求解器,那是神经网络普及前 Google 使用的大规模线性求解器。然后我问自己想用神经网络做什么,答案很明确:我想理解训练算法并让它变得更好。后来有一天,Vinit Gupta 走到我桌前说:‘我听说你非常擅长编写优化方法。我们有一个在白板上推敲出的想法——后来成了 Shampoo 算法。你能帮我们大规模实现它用于神经网络训练吗?’所以我开始做这个,感谢我的经理 Yanghui Wu 一直支持我到 2024 年离职。但社区对这个想法并不兴奋。对我来说,这是最令人兴奋的事,因为我正在投入算力让训练变得更好。人们假设:‘为什么?上界是什么?你还可以用 Adam。’但优化是关键:你有一个模型,你想更好地优化它。现在联系到架构:很多架构工作是为了让网络可训练。优化和架构是同一枚硬币的两面。更强的优化器可以训练更难优化的模型从而获得更好性能,或者用较弱的优化器训练更容易的模型也能得到不错的性能。这里存在权衡。我在这个上面花了很多时间。我们在 Gemini 1.5 Flash 中使用了 Shampoo。然后社区开始感兴趣,出现了 Soap 论文和一系列关于 Shampoo、Soap 等的研究。这大概带来了 2 倍的改进。但 Shampoo 还很弱——它没有使用所有可用信息。随着你在训练中使用更多信息,你会得到更好的改进。你的优化算法决定了你能发现什么样的架构。我有一些同事——全世界大约只有四个人关心这类想法,比如去掉残差连接学习更深的表示,但这需要更好的优化方法。所以对于你的问题,我的回答是:结合优化和架构,端到端地思考,这其中蕴含了大量的计算效率。
You know, I would start with measurement. I think pre-training as we define it right now is about compression. We look at perplexity and then measure how to decrease it, and we find that scaling, increasing parameter count, and putting more compute is the way. Every time we increase compute on a log scale, we get epsilon more improvement in these metrics. I think this is fine for building the prior, but it's the wrong way to look at the problem. We should look end-to-end: what are we training these models for? Look at the outcome. For example, I train this model and give it to Jerry. Jerry will do RL and destroy all the perplexity metrics I created. So that was the best way we had so far, and I think the labs have done a great job producing intelligence that's super valuable. But it was the bootstrap process. We have to combine pre-training and RL together—that's where one order of magnitude improvement would come from, and that's like a training procedure. You could say it's a learning algorithm. In terms of optimization, my story started at Google in 2016, optimizing logistic regression. I worked on solvers for what we used to call Sybil, the large-scale linear solver used at Google before neural networks took off. Then I asked myself what I wanted to work on with neural networks, and it was clear: I want to understand the training algorithm and make it better. Then one day, Vinit Gupta showed up at my desk and said, 'I heard you're really good at writing optimization methods. We have this idea that we worked out on a whiteboard—what turned out to be the Shampoo algorithm. Can you help us implement it at scale for neural network training?' So I worked on it, and I thank my manager Yanghui Wu for supporting it through my tenure until 2024. But the community wasn't excited by this idea. For me, it was the most exciting thing because I was putting in computation to make training better. People assumed, 'Why? What's the upper bound? You could still use Adam.' But optimization is key: you have a model, you want to optimize it better. Now, connecting to architecture: a lot of architecture work is to make networks trainable. Optimization and architecture are two sides of the coin. A stronger optimizer can train a harder-to-optimize model for better performance, or a weaker optimizer on easier models gives decent performance. There are tradeoffs. I spent a lot of time on this. We used Shampoo for Gemini 1.5 Flash. Then the community got interested, with the Soap paper and an entire literature of Shampoo, Soap, etc. That was maybe a 2x improvement. But Shampoo is quite weak—it's not using all available information. As you use more information, you get better improvements. Your optimization algorithm defines what architectures you discover. I have colleagues—only about four people worldwide care about ideas like getting rid of residual connections and learning deeper representations, but they needed a better optimization method. So for me, the answer to your question is that combining optimization and architecture, thinking end-to-end, is where a lot of computational efficiency lies.
是的。
Yeah.
我还认为 RL 消耗了大量算力但效率不高。因为你得不到太多反馈,却要花费大量算力解码冗长的思维链,只为向网络注入一点点信息。这看起来相当低效,是获得数量级提升的轻松目标。关于优化我可以一直说下去。
And I also see RL as spending a lot of compute not very efficiently. Because you don't get much feedback and you spend a lot more compute decoding long chains of thought to get one bit of information into the network. That seems quite inefficient and an easy target for orders of magnitude improvement. I could go on talking about optimization all day.
你认为我们能否达到或超越生物学习的效率?
Do you think we'll ever approach or surpass biological learning efficiency?
我不这么认为,至少在我们现有的硬件上。看起来不太可能。我们的生物学习——就像 Jeff Hinton 所说的模型计算——我们在成长过程中构建自己的电路,用硬件构建自己的学习算法。然后我们死去,它就消失了。神经网络则非常不同。硬件保留下来,神经网络保留下来,但学习效率非常低。你需要更多的网络和大量的并行才能让少量信息通过。除非我们设计出更像人类运行的硬件——可能更模拟化,找到处理模拟电路和纠错的方法,找到让信息通过的方法——否则会困难得多。我认为我们是安全的。
I do not think so, at least with the hardware we have. It seems pretty unlikely. Our biological learning—as Jeff Hinton says, model computation—we build our own circuit as we grow up and build our own learning algorithm with the hardware. Then we die and it's gone. Neural networks are quite different. The hardware stays, the neural network stays, but it learns very inefficiently. You need a lot more of them and a lot of parallelism to get small amounts of information through. Until we design hardware to be much more like how humans operate—maybe more analog, figure out how to deal with analog circuits and error correction, figure out how to get information through—it will be much harder. I think we're safe.
安全。这个说法很有意思。预训练和 RL 应该端到端优化,这听起来像是显而易见的观点。你认为实验室们意识到了吗?他们是否难以摆脱组织架构和流程来实现这一点?或者是什么阻碍了实验室统一这两者?
Safe. That's an interesting way to put it. The idea that pre-training and RL should be optimized end-to-end seems like such a clear, obvious statement. Do you think the labs realize this? And is it just hard for them to get rid of org charts and processes to make that happen? Or what stops the labs from unifying the two?
我不认为这很明显,因为它是一个完全不同的优化问题。
I don't think it's that obvious because it's a completely different optimization problem.
你有一个先验,你进行 rollout,你有更高的方差。然后预训练是更大批次,每单位时间有更多并行性或算力。所以,人们并不容易想到将这两种训练程序结合起来,直到你思考,“为什么简单的结合不起作用?”这是其一。第二点是,如果我询问这些实验室的一些顶尖研究员,他们会说,“哦,这有道理,我们或许应该探索,但它可能不会被优先考虑,因为他们必须为下一个周期训练模型。”就像 Jerry 所说,现在许多公司争相缩短发布周期,因为 token 不具备粘性。因此,在这些实验室所处的环境中,很难开展长期研究,即便是为期六个月的长期研究。
You have a prior, you're doing rollouts, you have higher variance. And then pre-training is much larger batch, like more parallelism or compute per unit time. So, it is not an obvious thing for folks to combine these two training procedures until you think, "Why is it that the naive combination doesn't work?" So that's one. The second one is if I ask some of the best researchers in these labs, they would say, "Oh, this makes sense. We should probably explore it, but it would probably not be in the top bucket because they have to train a model for the next cycle." It's like as Jerry said, there are companies now competing for release cycles because tokens are not sticky. So, it's much harder to do long-term research, even six-month research, in many of these labs in the environment they are in.
所以,Coral Lemonade 的一个核心前提似乎是,你们在 Sam 谈论 AI 科学家之际成立了一个实验室。我认为 Dario 也在谈论 AI 科学家。你们作为研究者的工作似乎发生了根本性变化。而你们恰好可以在这个时代创立公司。因此,你们或许能进行比以往更多的实验。你认为研究工作可以自动化到什么程度?你们如何建设你们的实验室,使其——我相信你们的使命之一——成为最自主的实验室?
So, it seems like one of the core premises for Coral Lemonade is that you're starting a lab at a time when Sam's been talking about the AI scientist. I think Dario's been talking about the AI scientist. It seems like your job as researchers has fundamentally changed. And you get to start the company native to that era. As a result, you're maybe able to run a lot more experiments than otherwise might be possible. How automatable do you think the research job is? And how are you guys approaching building your lab to be as I believe your mission, one of your missions, is to be the most autonomous lab there is?
世界上最自动化的实验室。首先,我认为自动化,其核心版本,是在某种程度上赋予每个人最大程度的能动性。我们并非真正试图让人类脱离循环——那是自动化的一种版本——而是要让人类能够在他们的时间内做最多的事情。走路时,你可以走一段距离;骑自行车时,你可以走更远;开车时,你可以走得更远更远。人类开始耕作时,必须手工耕一小块地;有了机器,你可以耕作更大的土地。就个人而言,我是当前编码代理的超级粉丝,也很高兴。在某种程度上,这是我多年来的工作:既做编码研究,也在 OpenAI 内部参与各种版本的 AI 科学家项目。最后,我意识到创办一家公司来实现这一愿景是最好的方式之一,因为今天做研究的方式已经非常不同。一个研究员可以做更多事情。迭代速度、研究速度、你在想法之间移动以及获取数据的速度,都非常不同。你可以尝试调整旧的结构、团队工作流和数据收集方式,也可以像你说的那样,原生地构建流程,最大限度地赋能每个研究员,让他们更快地迭代想法。我们在这里试图重建深度学习栈,并思考如何几乎以不同方式执行每一个操作。有哪些选项?如果我们每天能执行哪怕一个这样的实验,那相比之前已经是很好的迭代速度了。没有任何物理定律说不能。也许有一天我们每天能做十个,也许一天能做两百个。从根本上说,对于搜索和优化过程,我们应该能发现更好的深度学习设置。我们正在尝试做的是——几乎我们所有人都是非常“智能体构建”和“自动化构建”的团队——我们正在做一个实验,看我们能把这些东西推多远,以及一个小团队尽可能多的组织能走多远。
The most automated lab in the world. To start with, I think that automation, the version of automation by core, is about giving each human maximum level of agency in some way. We are not trying to really get humans out of the loop, which is one version of automation, but it is about giving humans the ability to do the most with their amount of time. Whenever you walk, you can get some distance. Whenever you get a bike, you can go a larger distance. Whenever you are in a car, you can go much, much larger. When humans started farming, they had to farm by hand and work on a small plot of land. When you have a machine, you work on a much larger plot of land. Personally, I am a big fan of the current coding agents. I'm very happy. In some way, it is what I've been working for many years, both doing coding research and working on various versions of an AI scientist inside OpenAI. In the end, I realized starting a company to realize that vision is one of the best ways to realize it because the way you can do research today is very different. A single researcher can do much more. The speed of iteration, the speed of research, how quickly you can move through ideas and get data on your ideas, is something very different. You can try to move the old structures around it, the team workflows, how data is gathered, or you can try to build natively for processes that maximally empower each researcher and allow them to iterate on their ideas much quicker. We are here trying to rebuild the deep learning stack and think about how we can do almost every operation differently. What are the various options? If we can execute even one of those experiments a day, that's already a pretty good iteration speed compared to anything before. There isn't any fundamental law of physics why not. Maybe one day we get to ten of those a day. Maybe one day we get two hundred a day. Fundamentally, for that search process, optimization process, we should be able to find things that work in a better deep learning setting. What we are trying to do is, almost all of us are a team that is very agent-built and automation-built. We are trying to do an experiment of how far we can push these things, and how much an organization that tries to do as much as we can with a small team can get.
我们何时才能知道我们已经达到了 AGI?
When will we know that we've reached AGI?
曾经我说过,这很大程度上取决于每个人的内心,不管他们如何定义 AGI。OpenAI 说 AGI 是能在经济价值工作中超越所有人类的系统,但这又回到了我之前的说法。如果 OpenAI 停止训练模型呢?它还会继续工作吗?自动化水平还会保持不变吗?对我来说,AGI 是一个能在没有任何人类参与的情况下自我改进的模型。我认为那才是我们可以有意义地谈论 AGI 的时刻,因为它在某种程度上是前一个定义的子集,因为改进 AI 模型本质上是人类能做的工作,是有经济价值的。但这确实是事实,然而迄今为止,让模型脱离人类循环是出了名的困难。我们离这个目标还很远。我很难找到任何一个任务能完全让人类脱离循环。目前成功的是人类-LLM 混合体,而没有人类的 LLM 则不太行,完全不行。我在 2024 年和 2025 年初看到的是,当前的道路无法让我们到达那里。我认为我们需要非常严肃的研究,无论是这家公司还是其他公司,去解锁如何让我们的模型在测试时更深层次地学习和适应。
At some moment I used to say it's very much in everyone's heart, whatever they consider AGI. OpenAI says the system that can outperform all humans in economically valuable work, but it goes to my previous statement. What if OpenAI stops training models? Would that still keep working? Would that still keep the automation level the same, or would that drive something? And for me, AGI is a model that can improve itself without human in the loop in any way. That is the moment where we can meaningfully talk about AGI, because in some way it is a sub-definition of the previous one, since improving AI models is actually a job that humans can do. It is economically valuable work. And it definitely is the case, but removing humans from loops with models has been notoriously difficult so far. We haven't come anywhere close to it. It's very hard for me to find any task where we were able to get humans out of the loop. We are the human-LLM hybrid which is really successful right now, but LLMs without humans not so much, not at all. And what I have seen in 2024 and early 2025 is that the current path doesn't get us there. I think we need some pretty serious research at this company or some other to try to unlock how we make our models learn and adapt at test time on a deeper level than we have so far.
嗯。
Mhm.
我觉得整个对话中我们已经暗示过这一点,但这有点像盲人摸象。你愿意分享的 Core AI 的宏伟计划是什么?
I feel like we've alluded to this throughout the conversation, but it's kind of been one of these five blind men trying to find the elephant things. What is the grand master plan for Core AI that you're willing to share?
我可以分享我们六个月的路线图。从某种意义上说,构建架构——正如我所说——不仅仅关乎架构有多好。它运行得好吗?我们能获得用户(包括实验室内部人员)来使用它吗?对吧?所以,我们现在可以直接做很多事情,但困难和想要自动化的事情是内核生成。我们有一批硬件 GPU,黑盒,我们必须在上面进行训练和推理。
I can share our 6-month roadmap. In some sense, building architectures, as I said, it's not just about how good the architectures are. Does it run well, and can we get users, including ourselves as part of the lab, to use it? Right? So, directly, we can do many things now, but the thing that is going to be difficult and that we want to automate away is kernel generation. So, we have a set of hardware GPUs, black boxes, that we have to train and run inference on.
我们将打造最好的模型,用以缩短从产生一个能消除架构瓶颈的绝妙想法,到让它在 GPU 上以最高 TFLOPS 运行的周期。从某种意义上说,目前的编码智能体加上人类可以走很远,但一个例子是我们与 GPU Mode 联合举办的紧急 QR 内核竞赛。这个竞赛针对的是 QR 分解这一相当古老的线性代数运算,Shampoo 优化器系列等工作会用到它,其他许多地方也会用到。你希望在 B200 节点上高效运行它。如果使用 cuSolver 对我们关心的形状做 QR,能得到一定效率。然后人类加上某种搜索循环可以做到大约 7 倍加速。但这需要最顶尖的人类——全世界大概只有三个人——并在 4 周内花大约 10 万美元在这种编码智能体上,才能得到一个 60 倍加速的解决方案。所以今天的模型远远没有达到那个 60 倍内核的水平。现在这是一个真正的瓶颈。那只是单个问题,它大概包含三个不同的算子:在这块面板上运算,做这个矩阵乘法,再折叠回去,如此反复。这就是矩阵的平方分解的样子。如果你把这个问题给到前沿模型比如 Gemini,它根本解决不了。我们的模型离解决这个问题还差得远。对我们来说,这是我们一直在讨论的事情,而且我们正在接近那个临界点,因为这是我们获得更高效架构的内部循环。
We will build the best model that we can to basically reduce the time from having a very cool idea that can make these bottlenecks go away in the architecture to having them run at the highest TFLOPS on GPUs. In some sense, current coding agents plus humans can go a long way, but an example of this is our urgent QR kernel competition that we hosted with GPU mode. It's for running this fairly old linear algebra operation QR. It's used for optimization like the shampoo line of work uses it, and many other places. And you want to run this efficiently on a B200 node. If you use cuSolver for the shapes that we care about, you get some efficiency. Then a human plus some search loop can get you something like 7x. But it requires the highest taste human—there are maybe three people in the world—and spend about $100,000 on these coding agents over the span of 4 weeks to get to a solution that's 60x faster. So these models today are nowhere close to getting that 60x faster kernel. There is a real bottleneck now. That was a single problem, it has perhaps three different operators: work on this panel, do this matrix multiply, fold it back in, and do this repeatedly. That's what the square factorization of a matrix would look like. If you give this problem to frontier models like Gemini, it just wouldn't solve it. Our models are not even close to solving this problem. For us, it's something we've talked about and we're getting close to that point because that's our inner loop to having more efficient architectures.
为什么是内核?是不是因为要最大化每 FLOPS 的智能,就需要自己生成内核?
Why kernels? Is it just because maximizing intelligence per flop of compute requires generating your own kernels?
从某种意义上说,我有过三个项目。其中两个已经落地到工业界了。首先是辅助方法。内核是一个瓶颈,因为你得把它运行好。如果你在像 Google 这样的地方,你不能花 10 倍的计算量只得到 2 倍的收益。所以我只能花大概 20% 的预算得到 2 倍的收益。大家都开心。我不觉得这很好。所以我认为这就是市场,对吧?你付出少于你的收获。所以内核最终成为了那里的瓶颈,因为大部分操作都是新颖的,没有多少人研究过。Google 只有两个人类能写这种代码,Rasmus 和 Peter Hawkins,因为那是深层的 XLA LLO 代码,需要写出来才能工作,他们花了两年时间才完成。另一个想法是我当时和一个同事一起做的,用额外的内存替换 Transformer 中的一些参数,我们称之为 N-gram 和 N-gram 内存。我们在 2020 年进行了研究。我们在内部部署了它的版本,不是大的版本,是小的版本。但那里我需要一些能加速稀疏 gather 和 scatter 操作的东西作为训练的一部分。这需要硬件改动并且硬件要能利用它。但最终没有实现。我曾经和 TPU 团队、我们和其他一些人开过会,我们一直在讨论,在疫情期间,我们说“哦,我们会实现的”,但它从未实现。我在 Anthropic 时也用过 TPU。在我离开的时候,才刚刚开始触及能实现它的表面。但与此同时,在那之前 6 个月,DeepSeek 写了他们的 N-gram,这是一种增加更多内存的改进版本。缩放损失变小了,是的,你不需要 MoE,你可以实际上用这些 N-gram 嵌入来替代它。对我来说,那就像……那是五年的事情,我为他们感到非常高兴。
In some sense, I've had three projects. Two of them have kind of landed in the industry. First is secondary methods. Kernels were a bottleneck because you have to run it well. If you were at a place like Google, you cannot spend 10x the compute and get a 2x win. So I could only spend maybe a budget of 20% and get the 2x win. Everyone's happy. I don't like great. So I think that's the market, right? You spend less than you get. So kernels ended up being a bottleneck there because most of the operations were novel and we hadn't gotten a lot of people to look at it. There are only two humans at Google who could write it, Rasmus and Peter Hawkins, because it was deep XLA LLO code that you have to write to make this work, and that took them 2 years to do. The other idea I had with one of my coworkers at that time was replacing some of the parameters in a transformer with extra memory and we called it N-gram and N-gram memory. We worked on it in 2020. We had versions of it internally deployed, not the big version, the smaller version. But there I needed something that can accelerate sparse gathers and scatters as part of training. It required hardware change and hardware making use of it. It never arrived. I had conferences set up with the TPU team, us and a bunch of others. We were talking about it and during COVID, like oh, we're going to have it happen, and it never arrived. I also was using TPUs at Anthropic. While I was leaving, just barely started the surface of being able to do it. But at the same time, 6 months before that, DeepSeek wrote their N-gram, which is an improved version of adding more memory. Short scaling loss, yeah, you don't need MoEs, you could actually replace it with these N-gram embeddings. For me that was like, ah, it was like a five-year thing and I was very happy for them.
嗯。
Mhm.
内核,你需要借助辅助来编写内核或者解决那个内核以获得最高性能。屋顶线(Roofline)非常高。就像 QR 一样,如果我用 cuSolver 的 QR,能得到一定的性能。如果我们用竞赛获胜者的 QR,能得到 60 倍的加速。那完全是一个不同的竞技场。现在它开启了全新的一系列可以应用的算法,在训练 Transformer、训练优化器方面都是如此。QR 在分析特征分解和许多其他事情中都非常基础。所以我认为这是一项只有少数人具备的技能,而且这些人并不集中在一个地方,一个在这里,一个在那里。如果模型有那些能力就太理想了。
Kernels and you need to be assisted in writing kernels or solve that kernel to have the highest performance. The roofline is pretty high. So it's like the QR. If I use cuSolver's QR, I get some performance. If we use the competition winner's QR, you get 60x faster. And that is a completely different playing field. Now it opens up an entire new set of algorithms you can apply, in terms of training transformers, training optimizers. QR is so fundamental in analyzing eigendecomposition and many other things. So it is a thing that I think only few people have the skill set to and they're very much not at the same place. It's like one person here, one person there. And it would be ideal if models had those abilities.
也许我来稍微总结一下,从高层次谈谈我们想要什么。Kernel Automation 是一个实验室,旨在构建能够持续学习并从部署中学习的模型。我们相信,正如我提到的,Transformer 无法进行持续学习。没有办法让 Transformer 实现持续学习。所以我们知道必须找到一种不同的架构。总之,我们的目标是找到那种新架构,找到 Transformer 的替代品,并且我们希望建立最自动化的实验室来实现这一点。我们希望尽可能快地大规模构建实验,迭代它们,尝试许多新的架构想法,拥有很强的先验知识来高效搜索架构空间,比其他人更快地到达那个地方。这就是我们要做的,我们在内核、大规模训练、尝试新架构想法上所做的所有工作都是在探索那个空间。
Maybe I'll summarize a little bit and talk from the high level of what we want. Kernel Automation is a lab created to build models that continuously learn and then learn from deployment. We believe, as I mentioned, that transformers are incapable of continual learning. There's no way how to put continual learning on transformers. So we know we have to find a different architecture. Some anyway, our quest is to find that new architecture, find that transformer replacement, and we want to build the most automated lab to do it. We want to be able to build experiments at scale the quickest we can, iterate on them, try a lot of new architectural ideas, have strong priors of what we want to do to search the space of architectures efficiently to go to that place fastest than anyone else. That's what we want to do, and all the work we are doing on kernels, on large-scale training, on trying new architectural ideas is exploring that space.
所以你可以通过实验找到更优的架构。你怎么知道什么时候找到了?你寻找什么来宣告“啊哈,就是它了”?
So you can experiment your way into finding a superior architecture. How will you know when you've found it? What are you looking for to say, 'Aha, this is the one.'
这是一个很好的问题。总是有两个角度。根据我的想法和经验,每一项成功的研究都有一个图表,展示了一些其他图表无法展示的东西。有一条线以不同的方式弯曲,你会说这就是你想要的。但至少以我的研究经验来看,那个图表在旅程中已经相当晚了,大部分时候你已经知道你想要什么,已经知道你在做什么。我有点开玩笑,但实际上,我人生中所有最好的图表都是在梦里完成的,然后它们才变成了现实。我大致知道我在寻找什么。问题只是什么时候它真正契合,你明白我的意思吗?因为大多数时候你知道你在找什么,但你找不到。你尝试一个,不行;尝试第二个,也不行。但最终所有正确的碎片会拼在一起,而这些深度学习系统大多数都非常复杂。
That's a great question. There are always two angles. In my mind and my experience, every successful research had a plot that shows something that other plots don't show. There is one line that is a little bit bending in a different way and you're saying this is what you want. But at least it is my experience with research always has been that plot is already quite late in a journey where most of the time you already know what you want and already know what you are up to. I am a bit joking, but it's actually true that all the best plots in my life I have done in a dream before they were real. I kind of knew I was looking for. The question is just when it actually clicks, if you know what I mean. Because most of the time you know what you are looking for, but you are not finding it. You try one thing and it doesn't work. Try second thing and it doesn't work. But eventually all the right pieces fall into it, and most of those deep learning systems are very intricate.
通常你需要连续做对五件事才能让东西开始工作,然后最终得到你想要的图。所以我认为我们要找的是能在测试时学习的系统,如果我们看到系统具有有意义的长期适应能力——我们是在开玩笑,但这是真的——我们希望在每天的日常工作中评估我们的系统。它们会每天在核心自动化科学家的工作上做得更好吗?
So usually you have to get five things right in a row for the thing to start working and then eventually you get the plot that looks like you want and then you know. So I think what we are looking for is systems that learn at test time and if we see meaningful long-term adaptability of our systems and like we are joking, but it's a real we want to be evaluating our systems in our everyday work. Will they get better at doing the work of core automation scientists each day?
是的,我们喜欢团队一起去度假,看看实验室是否能在一周内产出更好的东西。给……
Yeah, we like go on a vacation as a team and see if the lab produces something better. For the week, give the
那你回来后做什么?
Then what do you do when you get back?
嗯,我们会看看会怎样。
Um, we'll see what
把假期延长两次、三次、四次,直到我们永久度假。
Extend the vacation two times, three or four times, until we are on permanent vacation.
呃,这是一个很好的结束语。嗯,Rohan、Jerry,非常感谢你们来做客。你们都为我们今天所处的阶段做出了真正具有变革性的工作。我非常兴奋看到你们开始实验室,这是新的旅程。呃,非常期待看到你们能够取得的成果。感谢你们来做客。
Uh, that is a beautiful note to end on. Um, Rohan, Jerry, thank you so much for joining us. You've both worked on really transformative work for where we are today. And I'm so excited to see you starting a lab on this new journey. Uh, and very excited to see what you're able to come up with. Thank you for joining us.
很高兴来到这里和你聊天。
It's great to be here and to chat with you.
谢谢。
Thank you.