Why AI Progress Isn't Slowing Down
打开互动全文版(中英对照 + 朗读 + 问答)→Lucas Kaiser 解释为何 AI 放缓的说法是错误的,强调推理模型的新范式以相同成本带来更多收益。
Lucas Kaiser explains why the narrative of AI slowdown is wrong, highlighting the new paradigm of reasoning models that deliver more gains for the same cost.
大家好,我是马特。欢迎收听马特播客。今天的嘉宾是Łukasz Kaiser,他是现代人工智能的关键架构师之一,真正塑造了这个领域的历史。Łukasz 是《注意力就是一切》论文的合著者之一,这意味着他是 Transformer 架构的发明者之一,这个架构驱动了我们今天使用的几乎所有 AI。他现在是 OpenAI 的顶尖研究科学家,正在推动向推理模型的第二次重大范式转变,比如 GPT-5.1 背后的模型。本期节目将深入探讨 AI 前沿:为什么 AI 放缓的说法是错误的,哪些逻辑谜题仍然难倒世界上最聪明的模型,Scaling 如何被重新定义,以及这一切告诉我们 AI 下一步将走向何方。请享受与Łukasz 的精彩对话。
Hi, I'm Matt. Welcome to the Matt podcast. My guest today is Łukasz Kaiser, one of the key architects of modern AI who has quite literally shaped the history of the field. Łukasz was one of the co-authors of the 'Attention Is All You Need' paper, meaning he's one of the inventors of the transformer architecture that powers almost all the AI that we use today. He's now a leading research scientist at OpenAI, helping drive the second major paradigm shift towards reasoning models like the ones behind GPT-5.1. This episode is a deep exploration of the AI frontier. Why the AI slowdown narrative is wrong, the logic puzzles that still stump the world's smartest models, how scaling is being redefined, and what all of that tells us about where AI is heading next. Please enjoy this fantastic conversation with Łukasz.
Łukasz,欢迎你。
Łukasz, welcome.
非常感谢。
Thank you very much.
过去一年里,至少在某些圈子里,也许是在旧金山以外,有一种说法认为 AI 进展正在放缓,我们已经用尽了预训练,缩放定律正在碰壁。然而,我们录制这期节目时,正值一个重大的一周或几周结束,发布了 GPT-5.1、GPT-5.1 Codex Max、GPT-5.1 Pro,还有 Gemini Banana Pro、Grok 4.1、Almo 3。所以这感觉严重违背了那种说法。前沿 AI 实验室的人对 AI 进展的了解,至少部分外界似乎并不理解,那是什么?
There was a narrative, at least in some circles maybe outside of San Francisco, throughout the year that AI progress was slowing down, that we had maxed out pre-training, that scaling laws were hitting a wall. Yet we are recording this at the end of a huge week or couple of weeks with the release of GPT-5.1, GPT-5.1 Codex Max, GPT-5.1 Pro, as well as Gemini Banana Pro, Grok 4.1, Almo 3. So this feels like a major violation of that narrative. What is it that people in frontier AI labs know about AI progress that at least parts of the rest of the world seem to not understand?
我认为这里有很多东西需要拆解。所以我想慢一点说。AI 领域正在发生一些事情,现在每周都有很多事情发生。你知道,新模型、编程、做幻灯片、自动驾驶汽车、图像、视频。这是一个不会让你长时间感到无聊的领域。但透过这一切,有时很难看到正在发生的根本性变化。从根本上说,如果你看 AI 的进展,能力一直非常平滑地呈指数级增长。这是总体趋势,而且从来没有太多让我——至少,我认为我实验室的同事们——相信这种趋势没有发生。这有点像摩尔定律,对吧?摩尔定律持续了几十年,可以说它仍然在继续,甚至随着 GPU 的发展而加速。但当然,它并不是由一项技术持续 40 年带来的。有一项技术,然后另一项,再另一项,这样持续了几十年。所以从外部看,你看到一条平滑的趋势,但从内部看,进展当然是通过新开发以及算力增加和更好的工程实现的。所有这些因素结合在一起。就语言模型而言,我认为有一个重要的转折点。一个点当然是 Transformer 开始的时候,但另一个点是推理模型,那发生在——我想第一个预览版大约是一年零一个月前。所以我们大概 3 年前就开始研究了,但如果你把它看作一个范式,它非常新。所以它总是像这些 S 曲线。它开始,然后给你惊人的增长,然后稍微平缓一些。我们会谈到预训练。我觉得预训练在某种意义上处于 S 曲线的上半部分,但并不是说预训练的 Scaling 不起作用。它们完全起作用。缩放定律说的是,你的损失会随着算力对数线性下降。我们完全看到了这一点,显然谷歌和其他所有实验室也看到了。问题是,你需要投入多少钱才能获得相应的收益?这需要很多钱,而且人们正在投入。但有了推理这个新范式,你可以用同样的钱获得更多的收益,因为它处于 S 曲线的较低部分,这些发现解锁了不可思议的能力。所以并不是预训练失败了。只是我们发现了一个新范式,以同样的价格带来了更惊人的发展。而这个范式仍然非常新。它发生得太快了。我想如果你眨一下眼,你可能会错过它,因为基本上你之前有 GPT-3.5 的聊天,它会给你答案,不使用任何工具,没有推理。它会回答你一些东西。而现在你有了聊天,如果你不关注,你可能眨了一下眼,它也会给你答案,你可能会说好吧,差不多一样,只是现在的聊天会去一些网站查找,进行推理,然后给你正确的答案,而不是它权重中记忆的东西。我非常喜欢这个例子:旧金山动物园明天几点开门?旧的聊天会告诉你,完全从记忆中幻觉出一个时间,可能是 5 年前在动物园网站上读到的时间,而且它不知道今天是哪天或明天,所以它只会假设是工作日。现在的聊天知道今天是哪天,因为它在系统提示中,它会去动物园网站,读取信息,提取出来,如果不明确,可能会再查三个网站确认,然后给你答案。但如果你眨一下眼,你可能会认为是一样的,但不,它好得多。而且,由于它可以阅读世界上所有的网站,它可以回答以前甚至无法触及的问题。所以有巨大的进步,而且发生得太快,甚至可能被错过。
I think there is a lot to unpack there. So I want to go a little slower. There is this thing that's happening in AI, and in AI every week now a lot is happening. You know, new model, coding, doing slides, self-driving cars, images, videos. It's a nice field that doesn't make you bored for a long time. But through all of this, it's sometimes hard to see the fundamental things that are happening. And fundamentally, if you look at AI progress, it's been a very smooth exponential increase in capabilities. This is the overarching trend, and there has never been much to make me—at least, and I think my colleagues in the labs—believe that this trend is not happening. It's a little bit like Moore's law, right? Moore's law happened through decades and decades, and arguably you would say it's still very much going on, if not speeding up with the GPUs. But of course it did not happen as one technology bringing you there for 40 years. There was one technology and then another and another and another, and this went on for decades. So from the outside you see a smooth trend, but from the inside, of course, progress is made through new developments in addition to the increase of compute power and better engineering. All of these things come together. In terms of language models, I think there was a big pivotal point. One point was of course the transformer when it started, but the other point was reasoning models, and that happened—I think one preview was a bit a year and a month ago or something like that. So we started working on it maybe 3 years ago, but it's very recent if you think of it as a paradigm. So it's always like these S-curves. It starts, then it gives you amazing growth, and then it flatlines a little bit. We'll get to pre-training. I feel pre-training in some sense is on the upper part of the S, but it's not like scaling for pre-training doesn't work. They totally work. What scaling laws say is that your loss will log-linearly decrease with your compute. We totally see that, and clearly Google sees that and all other labs. The problem is, how much money do you need to put into that versus the gains you get? It's just a lot of money, and people are putting it. But with the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this lower part of the S-curve, and these discoveries unlock insane capabilities. So it's not like pre-training fizzled out. It's just we found a new paradigm that at the same price gives us much more amazing development. And this paradigm is still very new. It happens so fast. I think if you blink you may miss it, because it was basically you had GPT-3.5 in chat, and it would give you answers and it used no tools, no reasoning. It would answer you something. And now you have chat, and if you were not into it you may have blinked, and it also gives you answers, and you may say okay it's more or less the same, except the chat now will go look on some websites, reason about it, and give you the right answer instead of something it memorized in its weights. I very much used to like this example: when what time does the SF Zoo open tomorrow? The old chat would tell you, totally hallucinate from its memory an hour that it read on the zoo's website from 5 years ago, and it didn't know what's today or tomorrow, so it would just assume it's a weekday. Chat now knows what's today because it's in the system prompt, it goes to the zoo website, reads it, extracts the information, if it's ambiguous probably checks three other websites just to confirm, and then gives you the answer. But if you blink, you may think it's the same, but no, it's dramatically better. And as a consequence, since it can read all the websites in the world, it can give you answers on stuff that it wouldn't be able to even touch before. So there is tremendous progress, and it happened so fast it may even be missed.
我认为圈内人知道而外人不知道的最大一点是,现在已经不是进步的问题了——ChatGPT、Gemini 或任何大语言模型能为你做很多事,而人们根本没意识到。你可以拍一张坏掉的东西的照片,问怎么修,它会告诉你。你可以给它大学水平的作业,它会帮你完成。这绝对令人惊叹。
I think one of the biggest things that people on the inside know and others don't is that already right now, it's not about the progress—there are so many things ChatGPT or Gemini or any LLM can do for you that people just don't realize. You can take a photo of something broken, ask how to repair it, and it will tell you. You can give it a college-level homework and it will do it for you. That's absolutely amazing.
所以在某种程度上存在认知差距。
So there is an education gap to some extent.
嗯,这已经发生了。你提到了 Codex,对吧?程序员有点保守。我偶尔还用 Emacs。以前所有的编码工具都只是帮我补全一行代码。人们会说,‘这是我的编辑器,我在这里写代码。’现在人们会说,‘不,这是 Codex。我让它做事,我之后再修改。’我认为就在最近几个月,从人们偶尔使用但很少用,到现在基本上成为很多人编程的方式,这个转变很大。我不确定每个人都意识到了,但如果你不编程,你为什么要意识到呢?我相信这会扩展到越来越多的领域。
Well, it just happened. You mentioned Codex, right? Programmers are a bit conservative. I still use Emacs from time to time. All the coding tools used to complete one line for me. People were like, 'This is my editor, I write code here.' Now people are like, 'No, this is Codex. I ask it to do stuff, I will fix it later.' I think it's in the recent few months that the transition happened from people using it sometimes but rarely, to now basically this being how a lot of people work in coding. That's quite big. I'm not sure everyone is aware of it, but if you don't do programming, why would you be? I do believe this will come to more and more domains.
以至于这一切都非常新,有些突然。我与人交谈时经常听到的一点是,人们如此乐观的部分原因是,未来几个月这些模型有很多唾手可得的改进点,非常明显。首先,你同意吗?其次,你能举一些例子吗,比如接下来需要修复的明显问题,并且行业会去修复的?
To the point of all of this being very new and somewhat sudden. Something I hear from time to time when talking to people is that part of the reason people are so optimistic is that there is a lot of low-hanging fruit, very obvious things to improve for those models in the next few months. First of all, do you agree? And second, can you give us some examples like obvious things that need to be fixed next and that the industry will fix?
是的,有大量极其明显的需要修复的问题。其中很大一部分很难在播客上讨论,因为属于工程层面。每个实验室都有自己的基础设施和代码中的 bug。机器学习在某种意义上非常宽容,不像传统软件工程,你犯错了它会直接报错。我们的 Python 代码通常能运行,但如果运行方式不对,会慢很多,结果也更差。所以你会意识到,‘哦不,错了。’然后你改进它,结果变好了。这些是巨大的分布式计算系统,运行起来非常复杂。所以在训练模型和做强化学习的过程中,有大量需要改进、修复和理解的地方,因为强化学习比预训练更挑剔,更难做好。这就是我们的日常工作。除此之外,还有数据。我们以前基本上只训练 Common Crawl。它是一个巨大的互联网仓库,人们不加区分地抓取,有些东西进来了,有些没有。那是一片混乱。所以现在,当然,每个大公司都有一个团队试图过滤数据并提高质量。但要真正提取更好的数据,工作量很大。现在合成数据开始流行,但生成合成数据时,如何做、用什么模型、整个工程方面都非常重要。这是一个如此新的领域,它被做出来了,能工作,很漂亮,但还有太多可以做得更好的地方,人们毫不怀疑那里有很多机会。除此之外,还有像多模态这样的大问题。语言模型现在,我相信你知道,大多数人也意识到,实际上是视觉语言模型,因为它们也能处理音频。所以它们在某种程度上是多模态模型,但多模态部分在很大程度上仍然落后于文本部分。所以这是一个大领域,显然需要做得更好,而且如何做得更好并不是什么大秘密。有一些方法可能会让它变得惊人地好,但也有一些非常简单的方法可以做得更好。但这可能需要从头重新训练整个基础模型,这需要几个月的时间,是一项巨大的投资,所以我们需要组织好。所以有很多工作无疑会让事情变得更好。我认为人们心中的大问题是,它们能变得多好?
Yes, there is a ton of extremely obvious things to fix. A larger part of this is just hard to talk about on a podcast because it's in the engineering part. Every lab has their own infrastructure and their own bugs in the code. Machine learning is beautifully forgiving in some sense, in contrast to old software engineering which would just yell at you when you made a mistake. Our Python code generally probably runs, except much slower and gives you worse results if you run it wrong. So you realize, 'Oh no, it was wrong.' And you improve it and the results get better. These are huge distributed computing systems. They're very complex to run. So there is a huge amount to improve and fix and understand in the process about just how to train your model and how to do RL, because RL is more finicky than pre-training. It's harder to do really. So every day this is our day-to-day work. On top of that, there is data. We used to train on just Common Crawl basically. It's a big repository of the internet that people just scraped without regard of what, and some things came in, some didn't. It was a mess. So now, of course, every larger company has a team that tries to filter this and improve the quality. But it's a lot of work to really extract better data. Now synthetic data is becoming a thing, but when you generate synthetic data, it really matters how you do it, with what model, the whole engineering aspects. It's such a new domain that it was done somehow, it works, it's beautiful, but there is just so much to do better that people don't have any doubts that there is a lot there. And on top of that, there are the big things like multimodal. Language models are now, as I'm sure you know and most people realize, actually vision-language models because they can also do audio. So they're multimodal models to some extent, but the multimodal part still lags behind the text part to a large extent. So that's one big area where obviously you need to do better, and it's not a huge secret how you can do better. There are some methods that maybe will make it even amazingly better, but there are some very simple methods how you can do just better. But this maybe requires retraining your whole base model from scratch, and that takes a few months and it's a huge investment, so we need to organize it. So there is a lot of just work that will undoubtedly make things better. I think the big question that people have in their mind is how much better will it make them?
所以我想稍微深入探讨一下,或者做一点科普,关于整个推理模型方面。因为正如你刚才提到的,由于它太新了,有些人真正理解它们的工作原理,很多人不理解。在非常简单的层面上,什么是推理模型?它和你所说的基础大语言模型有什么不同?
So I'd love to do a little bit of a deep dive or educational part on the whole reasoning model aspect, because as you just mentioned, since it's so new, some people truly understand how those work, many people don't. At a very simplistic level, what is a reasoning model and how is that different from your sort of base LLM?
所以,推理模型就像你的基础大语言模型,但在给你答案之前,它会思考——人们称之为思维链——意思是它会生成一些 token,一些文本,这些文本不是给你读的,而是为了让模型给你更好的答案。而且现在,当它这样做时,它也被允许使用工具。例如,在它的思考过程中,它可以去浏览网页来给你更好的答案。这就是思维模型的表面部分。现在深层部分是,你开始把这个思考过程基本上当作模型的一部分。所以它不是模型生成然后输出给你的东西。它是你想要训练的东西。你想告诉模型它应该好好思考,这样之后的答案在任何方面都是好的。这导致了一种非常不同的模型训练方式,因为模型通常只用梯度下降来训练,也就是深度神经网络的训练方式——意思是你说预测下一个词,然后你做梯度,你对模型求导——它们不是完全可微的,但你近似它,然后你训练权重来做这件事。仅仅这样做就能做出一个聊天模型,这相当惊人。但对于推理模型,你可以做到这一点,因为有这个推理部分;你可以通过它求导。所以我们用强化学习来训练它。强化学习基本上就是:有一个奖励,你需要做一些尝试并强化——意思是推动模型去做更多导致更好答案的事情。这种训练比我们以前用的训练有更多限制。我们以前用的训练,你把整个互联网放进去,即使过滤得不好,大部分也能工作。强化学习,你需要小心——你需要调整很多东西,而且你需要非常仔细地准备数据。所以目前,至少在我们目前使用的最基本方式中,它需要相当可验证。
So, a reasoning model is like your base LLM, but before giving you the answer, it thinks—what people call chain of thought—meaning it generates some tokens, some text that is meant not for you to read, but for the model to give you a better answer. And while it does this these days, it is also allowed to use tools. So it can, for example, in its thinking process, go and browse the web to give you a better answer. So that's the superficial part of the thinking models. Now the deep part is that you start treating this thinking process as part of the model basically. So it's not something the model generates and it's an output for you. It's something you want to train. You want to tell the model it should think well, so that the answer after this is good in whatever way. And this leads you to a very different way of training the model, because models were usually trained with just gradient descent, the way deep neural networks are trained—meaning you say predict the next word and you do a gradient, you differentiate your function from the model—they're not fully differentiable but you approximate it, and you train your weights to do that. And it was quite amazing that doing just that you could make a chat. But with a reasoning model, you can do that because there is this reasoning part; you can differentiate through that. So we train this with reinforcement learning. And reinforcement learning basically: there is just this reward and you need to do a bit of tries and reinforce—meaning push the model towards doing more of the things that lead to better answers. And this kind of training is a bit more like it has more restrictions than the training we used before. The training we used before, you took all of the internet, put it in, even if you didn't filter it very well, it would mostly work. Reinforcement learning, you need to be careful—you need to tune a lot of things, but you also need to prepare your data very carefully. So currently, for at least the most basic ways we use it currently, it needs to be fairly verifiable.
那么,你的答案正确与否?你可以为此准备数据。在数学和编程领域,你可以做得很好。在科学领域,你也可以在一定程度上做到,对吧?你可以有测试问题,判断它们是否正确。但你知道,如果涉及到写诗,这首诗好不好?目前,推理模型在科学等领域确实表现出色,它们也给非科学领域带来了一些改进,但还没有达到可能达到的巨大程度,至少与数学和编程相比。然后是多模态问题。你如何在多模态中进行推理?我认为这正在起步。我看到一些 Gemini 在推理部分生成图像,这相当令人兴奋,但还非常非常早期。
So there isn't is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. You can do this in science to some extent, right? You can have test questions, they're correct or not. But you know, if it comes to like writing poems, is this poem good or not? For now, the reasoning models really shine in domains like science, and they've brought some improvements to non-science domains, but it's not quite as huge yet as it could be, I mean at least compared to mathematics and coding. Then there is the multimodal question. How do you do reasoning in multimodal? I think this is starting. I saw some Gemini creating images in the reasoning part, that's quite exciting, but it's very, very early.
预训练和强化学习部分从教育角度来看再次特别有趣,因为人们似乎得出了结论:存在预训练世界和后训练世界,而后训练世界主要是强化学习。但预训练中也存在强化学习这个想法,我认为并非所有人都理解。
The pre-training and reinforcement learning part is particularly interesting from an educational point of view again, because it seems that people have come to the conclusion that there is the pre-training world and then the post-training world, and the post-training world is mostly reinforcement learning. But this idea that there is reinforcement learning in the pre-training, I don't think is as understood by everyone.
在 ChatGPT 初期,可以说,存在预训练。人们当时没有做强化学习,对吧?但那时你无法真正与它聊天,所以聊天功能是通过对预训练模型应用 RLHF 实现的。但 RLHF 是一种不同的强化学习,对吧?它非常小,而且是由人类偏好告诉你什么更好。这就是 HF 的含义:人类反馈。你向人们展示成对的内容,学习一个模型说:‘嗯,人们似乎更喜欢这个答案。’你用这个进行训练。它会很快,用我们今天的话说,破解这个模型。如果你训练 RLHF 太久,它就会开始给出满足这个模型的东西,而这个模型本应模拟人类偏好。所以这是一种有点脆弱的技术,但它是一种对让模型能够聊天至关重要的强化学习。如今,我认为大多数人转向了这种大规模的强化学习。它的规模仍然不如预训练,但它意味着你有一个模型来判断这是否正确,或者如果是偏好,它是一个非常强大的模型,分析事物并说你应该偏好那个,并且你的数据被限制在你能够做好的领域。然后你还可以在上面添加一些人类偏好,但你要确保可以运行更长时间而不会让整个评分体系崩溃。但同样,这是目前的情况。我确实相信未来的角色会更广泛。它将处理通用数据,也许那时它会扩展到超越今天它擅长的领域。它会在那里表现出色吗?那是另一个问题。你真的需要为做某些事情而进行大量思考吗?也许不需要。但也许需要,我们可能认为我们做的思考和推理比我们有意识地称之为思考的要多。
At the beginning of ChatGPT, let's say, there was pre-training. People did not do RL, right? But then you couldn't really chat with it, so chat was RLHF applied to a pre-trained model. But the RLHF was a different kind of RL of sorts, right? It was very small, and it was human preference that was telling you what is better. That's what the HF is: human feedback. You showed people pairs of stuff, you learned a model that says, 'Well, people seem to prefer this as an answer.' You trained with that. It would very quickly, what we today say, hack this model. If you trained the RLHF too long, it would start giving things that satisfy this model that seems supposed to model human preferences. So it was a bit of a brittle technique, but it was a bit of RL that was extremely crucial to making the models chat. These days, I think most people move towards this big RL. It's still not as big as pre-training in scope, but it says you have a model that says whether this is correct or not, or if it's a preference, it's a very strong model that analyzes things and says you should prefer that, and you have data that's restricted to a domain where you can do this well. And then you can also put some human preferences on top, but you make sure that you can run a little bit longer without making this whole grading fall apart. But again, this is around today. I do believe the role of tomorrow will be broader. It will work on general data, and maybe then it will expand to domains that go beyond where it shines today. Now will it shine there? That is a different question. Do you really need to think very much doing some of the things? Maybe not. But maybe yes, we maybe think we do more thinking and reasoning than we consciously call thinking.
强化学习要实现泛化需要什么?是更好的评估吗?比如你们几周或几个月前发布了 GDP val,用于衡量在广泛经济领域的表现。这是系统所需的一部分吗?
What would it take for RL to generalize? Is that better evaluations? Like you released GDP val a few weeks or months ago to sort of measure performance against broad economic sectors. Is that part of what the system needs?
我认为这只是其中一小部分。我认为这是一部分。但如果你考虑经济任务,制作幻灯片很重要,遵循指令,进行计算。这不是数学,但仍然非常可验证,对吧?我在想的是,当你进行预训练时,你拿互联网数据,然后只是问下一个词是什么。你知道,你可以在问下一个词之前进行思考。显然,现在你可能不想在每个词之前都思考,但我不确定你是否看过真实预训练运行的训练数据,因为我认为大多数人没有意识到这有多糟糕。比如,hotels.com 与互联网上平均 2000 词的文本块相比,是一个很棒的网站。它是一团糟,对吧?而且从这些数据中,预训练过程能得到合理的结果,这本身就是一个奇迹。所以你可能不想要,想象你有一个酒店网站告诉你这是一个美好的假期,你不一定需要在之前有很长的思维链,对吧?如果是由人写的,可能有一些思考在其中,也许不像数学和编程思考那样精细,但可能有一些东西在发生。所以也许你希望在至少部分文本之前有一点思考,而我们的模型还无法很好地做到这一点。我认为它们正在开始。我认为这种推理中有很多泛化。如果你学会为数学思考,你有时会使用一些策略,这些策略非常容易迁移,比如在网上查找并利用这些信息。所以其中一些东西非常通用,它们开始迁移。我觉得夏天可能还没有,特别是视觉领域的思考,我认为训练非常不足。但你知道,我们正在努力,所以我们会尝试推动更多这方面的进展。
I think this is a small part of it. I think that's one part. But if you think of economic tasks, making slides is important there, following instructions, doing calculations. It's not math, but it's still very verifiable, right? What I'm thinking about is when you do pre-training, you take the internet and you just say ask what's the next word. You know, you could think before you ask what's the next word. Obviously, now you don't want to think before every word probably, but I don't know if you've ever looked at the training data for a real pre-training run, because I think people mostly don't realize how bad this is. Like hotels.com is a great website compared with the average chunk of 2,000 words from the internet. It's a mess, right? And also a miracle that from this the pre-training process gets you something reasonable. So you probably don't want, imagine you have a hotel website telling you it's a beautiful vacation, you don't necessarily want to have a very long chain of thought before that, right? If it was written by a person, there was probably some kind of thinking that went into it, maybe not as elaborate as the math and coding thinking, but maybe there was something going on. So maybe you want a little bit of thought before at least some of the text, and our models can't do that very well yet. I think they're starting. I think there is a lot of generalization in this reasoning. If you learn to think for math, you will sometimes do some strategies that transfer very much, like look up on the web and see what they say and use that information. So some of these things are very generic and they start to transfer. I feel like summer maybe not yet, especially thinking in the visual domains is very undertrained, I believe. But you know, we work so we will try to push for more of that.
回到思维链,它实际上是如何工作的?模型如何决定创建那个思维链,我们作为用户在屏幕上看到的小中间步骤,即暴露给我们的思维链,是模型实际处理的内容吗?还是背后有更深、更长、更广的思维链在发生?
Going back to chain of thought, how does that actually work? How does the model decide to create that chain of thought, and is what we see, the little intermediary steps that we see on the screen as users, the chain of thought that's exposed to us, is that what's actually being processed by the model, or is there a deeper, longer, broader chain of thought that happens behind the scenes?
所以在当前的 ChatGPT 中,你会在网站上看到思维链的摘要。所以有另一个模型获取完整的思维链,并给你展示一个摘要,因为完整的思维链通常不太容易阅读。它们或多或少在说同样的事情,只是用更混乱的词。所以最好有一个更易读的摘要。当你开始使用思维链时,关于思维链的第一篇论文,你基本上只是要求模型‘请逐步思考’,它就会思考。所以如果你只是预训练一个模型在互联网上,并要求它逐步思考,它会给你一些思维链。有趣且最重要的一点是,你不止步于此。你说,好吧,你有某种思考方式,然后你说有时这会导致正确答案,有时会导致错误答案。所以现在我告诉你我有一些训练示例:你会思考 100 次,其中 30 次得到正确答案,然后我会在这 30 个示例上训练你,说这是你应该思考的方式。这就是强化学习部分。训练极大地改变了模型的思考方式。我们在数学和编程中看到了这一点。但最大的希望是,它也能改变模型在许多其他领域的思考方式。
So in the current ChatGPT, you will see a summary of the chain of thought on the site. So there is another model that takes the full chain of thought and shows you a summary, because the full ones are usually not very nice to read. They more or less say the same thing, just in more messy words. So it's better to have a more readable summary. When you start with a chain of thought, the first paper about chains of thought, you basically just ask the model 'please think step by step' and it would think. So if you just pre-train a model on the internet and ask it to think step by step, it will give you some chain of thought. Interesting and most important point is you don't stop there. You say okay, so you have some way of thinking, and then you say sometimes this leads to a correct answer and sometimes it leads to a wrong answer. So now I'm telling you I have some training examples: you will think 100 times and say 30 lead you to the correct answer, then I'll train you on these 30 examples, say this is the way you should be thinking. That's the reinforcement learning part. Training changes dramatically how the models think. We see this for math and coding. But the big hope is it could also change how the models think for many other domains.
即使在数学和编程领域,你开始看到模型开始纠正自己的错误。对吧?以前如果模型犯了错,它通常只会告诉你它做了什么,并坚持错误是对的之类的。但有了思考过程,它就会像‘哦,我经常犯错,但我需要验证并纠正自己才能给出正确答案’。所以这完全是从强化学习中涌现出来的,这很美,对吧?这显然是一种很好的思考策略:验证你想说的话,如果觉得可能出错,就再思考一遍。这就是模型在最抽象层面上学到的东西。
Even for math and coding, you start seeing that the models start correcting their own mistakes. Right? Earlier if the model made a mistake, it would generally just tell you what it did and insist that the mistake was right or something like that. With the thinking, it's like, 'Oh, I often make mistakes, but I need to verify and correct myself to give the correct answer.' So this just emerges from this reinforcement learning, which is beautiful, right? It's clearly a good thinking strategy to verify what you want to say. And if you think it may be an error, then think again. That's what the model learns on the most abstract level.
太好了,谢谢。好的,我们稍微绕个弯,稍后会回到更前沿的 AI 话题。我很想聊聊你的故事。你有着令人难以置信的成就:既参与了 Transformer 论文——那是一个范式的诞生,现在又领导着推理模型这一新范式。这真是一个了不起的故事。你是怎么成为 AI 研究员的?
Great. Thank you for this. All right. As a quick detour and we'll go back to more frontier AI topics. I'd love to talk a little bit about your story. I mean you have the incredible distinction of having been at the forefront of this industry: both the Transformers paper, which was the birth of one paradigm, and now you're very much leading the charge on the reasoning model part, which is another paradigm. So this is just an incredible story. How did you become an AI researcher?
我是一名数学家和计算机科学家,但偏向理论计算机科学。
I was a mathematician and a computer scientist, but in theoretical computer science.
那从高中就开始了?
And that started in high school as a kid?
是的,我高中时非常喜欢数学,后来也喜欢上了计算机。我在波兰读书,之后去德国读博士,是理论计算机科学和数学方向的博士。所以我本质上是个数学家。我一直对思维如何运作、什么是智能感到着迷。小时候我就想模拟大脑。我想,也许更高层次的解释更有趣。我做过逻辑研究,也写过一点程序,但后来深度学习刚兴起时,我有机会加入谷歌。我当时已经在法国有了终身教职,法国系统有一个很棒的政策:你可以休假 10 年。
Yes, I was definitely very into math in high school and into computers also later in high school. Yes, I did my studies in Poland. I went for a PhD in Germany. It was a theoretical computer science and mathematics PhD. So I very much am a mathematician. I was always fascinated by how this thinking goes. What is intelligence? As a child I always wanted to emulate the brain. I thought, well okay, maybe higher level explanations are more interesting. I did research in logic and a little programming, but then there was this opportunity to join Google just as deep learning was starting off. I already had my tenure position in France, and the French system has this beautiful thing that you can take a leave of 10 years.
是的。
Yes.
而且你随时可以回来,所以这是一个无风险的选择。
And you can still return anytime you want, so it's a no-risk situation.
所以等哪天你解决了 AGI,也许可以回法国当教授。
So at some point when you've solved AGI, you may return back to France and be a professor.
嗯,如果你解决了 AGI,他们可能还是会要你。休假的好处是,即使你没解决,他们也会让你回来。所以这真的很重要。我觉得不少诺贝尔奖得主都利用这个休假去尝试更有风险的事情,你知道,有时成功,有时不成功。科学和研究中有很多运气成分,但能有这个机会非常好。所以我来到了谷歌。当时是谷歌大脑,对吧?
Well, if you solve AGI, they may take you anyway. The nice part about the leave is that they will take you back even if you don't. So it's actually very important. I think a number of Nobel Prize winners took this leave to just try something more risky, and you know, sometimes it works, sometimes it doesn't. There's a lot of luck in science and research, but it's very good to have this opportunity. So I came to Google. That was Google Brain at the time, you said, right?
是的。
Yes.
我加入了雷·库兹韦尔的团队。他是我的第一任经理,面试了我,非常鼓舞人心。我最初的面试是加入 YouTube 的 UI 团队,我当时想,‘算了,我不去’。然后我和雷面试了,我当然从他的书中认识他,他是一个非常鼓舞人心的人,所以我想,‘好吧,就去吧’。当时那个团队和谷歌大脑是分开的。后来我转到了谷歌大脑,和伊利亚·苏茨克弗一起工作,他也是非常鼓舞人心的人。湾区 AI 领域有太多优秀的人了。
I came to Ray Kurzweil's group. He was my first manager. He interviewed me and was very inspiring. My first interview was to join the YouTube UI team, and I was like, 'Okay, I'm not going.' And then I had an interview with Ray, and I knew him of course from his books, and he's a very inspiring person, so I was like, 'Okay, let's go.' The team was separate from Google Brain at that time. Then I moved to Google Brain, worked with Ilya Sutskever, another very inspiring person. There's an amazing number of great people in AI in the Bay in general.
我这时必须问问你 Transformer 论文的故事,这一切是怎么发生的。你们八个人,对吧?七八个人。你们是怎么聚到一起的?
I have to ask you at this point about the Transformer paper story, how it all came about. The eight of you, right? Seven or eight of you. How did you all get together?
嗯,我们从未聚到一起。
Well, we never got together.
你们从未聚到一起。好吧。
You never got together. Okay.
我最近在推特上看到一张我们八个人的合影,有人说那是假的,但我知道是假的,因为我不认为我们八个人曾经同时出现在同一个房间里。这些想法来自多个方面。之前和之后,比如雅各布·乌兹科雷特和卢卡什·凯泽研究过注意力机制,比如自注意力。当然,注意力机制在编码器-解码器方面已经存在了。
I recently saw a photo on Twitter of a photo session of all eight of us, and it was saying it was fake, but I knew it was fake because I don't think all eight of us were ever in the same physical room. These ideas developed from many sides. Before and after, like Jakob Uszkoreit and Łukasz Kaiser worked on attention, like self-attention. Of course attention was there from the encoder-decoder side.
也许花一分钟给大众解释一下注意力机制到底是什么意思,因为它是一个如此基础的概念。
And maybe one minute for the broad public on what attention actually means, since it's such a fundamental concept.
注意力机制是一种机制,它告诉模型,当你做下一件事时,回顾你的过去,找到你过去看到的与你现在看到的最相似的东西。它来自机器翻译时代,当时人们想把一种语言的单词与另一种语言的单词对齐。他们想,‘好吧,这个单词,在之前的句子中对应哪里?’这是深度学习中对齐的类比。现在它被称为注意力机制,在 AI 中,它就是说,想想当你现在在这个环境中时,你脑海中浮现出什么,过去有什么东西与它相似。这个机制在之前的深度学习翻译中已经使用了,但当时是一个编码器模型,解码器会查看编码器,但从不查看自己的状态。Transformer 的主要创新是自注意力。但 Transformer 不仅仅是这个想法。我认为这很重要。我认为这八个怪人虽然从未物理上聚在一起,却以某种方式共同完成这件事的美妙之处在于,我们每个人都从不同的角度切入。所以有人研究注意力想法。有人需要把这个放入一个需要大量知识的网络中。所以有前馈层,它先扩展再收缩。诺姆·沙泽尔在研究这个,现在使用的混合专家模型实际上在 Transformer 之前就出现了。所以如何在神经网络中存储知识是另一个重要问题,它也是这个模型的一部分。然后,你知道,在深度学习领域,人们笑称想法是廉价的,让它们工作才是困难的部分。
So attention is the mechanism that tells the model, as you're doing the next thing, look into your past and find the most similar things that you see in the past to what you are seeing right now. It came from the machine translation times where people wanted to align words in one language with words in another. They were like, 'Okay, so this word, where in this previous sentence would it be?' It's an analog of alignment for deep learning. It's now called attention, and in AI it just says, think of what comes to your mind as you are here now in this environment, what things from the past are similar to it. And this mechanism was already used in deep learning translation before, but it was used like one encoder model and the decoder would be looking at the encoder but never at its own states. The main novelty of Transformer was self-attention. But Transformer is more than just this idea. I think that's important. I think the beauty of these eight weird people somehow coming together, even though not physically, to do it, is that we all approached it from different sides. So there were people working on the attention idea. There is the need to put this in a network that needs to have a lot of knowledge. So there is the feed-forward layer that expands and then contracts. Noam Shazeer was working on this, and nowadays used mixtures of experts, which actually came before the Transformer. So how do you store knowledge in neural networks is another important question, and it's part of this model too. And then, you know, in deep learning people laugh that ideas are cheap, making them work is the hard part.
如何编写系统、代码和基线来真正训练它?现在说起来很有趣,因为如今你可以用任何深度学习框架,写一句‘X = transformer(X) train’就能基本跑通,但当时完全不行。所以你需要学习率预热或优化器的调整之类的东西,这些只是工作。我做了很多编码工作,当时在开发 TensorFlow 和框架的某些部分。我清楚地记得人们问,‘所以你想用同一个模型做几个不同的任务?’为什么要这么做?如果你有不同的任务,比如做翻译,你训练一个模型;做解析,你训练另一个;做图像识别,你训练第三个。你从来不会为三个不同的任务训练同一个模型。
How do you write the systems and the code and the baselines to actually make this train? And this is funny to say now because nowadays you can take any deep learning framework and say 'X = transformer(X) train' and it will basically work, but back then it totally did not. So you need things like learning rate warm-up or tweaks to the optimizer that were just work. And I did a lot of coding and at that time was working on TensorFlow and parts of the framework. And I remember distinctly that people were like, 'So you want to use the same model for a few different tasks?' Like why do you even do that? If you have a different task, like if you do translation, you train one model; if you do parsing, you train another; if you do image recognition, you train a third. Like you never train the same model for three different tasks.
为什么你们甚至要写 API 来让一个模型做多个任务?我当时说:‘不,不,我们要用一个模型做所有任务。’然后人们说:‘不,不。’所以这个想法遭到了很多反对。不是反对这个想法。谷歌当时也是一个很棒的地方,他们很乐意让你做任何你想做的事。但我不认为当时人们普遍相信可以用同一个模型做多个任务。更不用说,这个想法——你拿基本上和 Transformer 一样的模型,现在虽然有很多改动,但原则上你可以拿论文里的解码器架构,在整个互联网上训练它,然后它基本上就能开始和你聊天了。这在当时听起来绝对是一个值得追求的梦想。我们也许把它当作一个梦想,但没想到五年后就成了现实。它居然真的这么好用,真是太幸运了,对吧?
Why do you even write APIs to do multiple tasks on one model? And I was like, 'No, no, we're going to do all tasks in one model.' And people were like, 'No, no.' So there was a lot of pushback against the idea. Not against the idea. Google was also an amazing place at that time that they would very happily let you work on whatever you wanted. But I don't think there was widespread belief in doing multiple tasks with the same model. Not to mention, this idea that you take basically the same model as Transformer. Like now there are a bunch of changes to it, but you could in principle take the same architecture as the decoder from the paper, train it on all of the internet and it will basically start chatting with you. It would have back then definitely sounded as a worthy dream. We maybe had as a dream but not reality that you'd expect 5 years later. It's very lucky that it actually works so well, right?
谈谈从谷歌到 OpenAI 的转变,以及这两种文化可能有什么不同。
Talk about the transition from Google to OpenAI and perhaps how those two cultures are different.
伊利亚在 Brain 时是我的经理。后来他创立了 OpenAI。这些年来他问过我几次是否愿意加入。我当时觉得有点太前卫了。然后 Transformer 出现了。我们为此做了很多工作。接着新冠疫情来了,对全世界来说都是一段艰难时期,对吧?但谷歌完全关闭了。谷歌重新开放得非常缓慢。所以一方面,我觉得远程工作很难。我更喜欢直接和同事一起工作。这是一个原因。但另一方面,谷歌 Brain 在我加入时只有几十个人,大概 40 人左右。我离开时已经有 3000 到 4000 人,分布在多个办公室。在小团队和大公司工作是非常不同的。所以综合考虑,我想,OpenAI 现在状态稳定多了。我们在做语言模型。这看起来可能很匹配。我就想,好吧,让我试试。除了大学,我之前只在谷歌工作过。所以转到小型创业团队是一个很大的变化,但我喜欢在小团队工作。它有它的乐趣,对吧?有时候强度会有点不同。总的来说,我觉得很不错。另一方面,谷歌合并后做了 Gemini,我听说也是一个很好的地方。我认为总的来说,科技实验室之间的相似性比人们想象的要大。有一些差异,但我觉得如果从外部来看,比如从法国的大学来看,这所大学和任何一个科技实验室之间的差异,比实验室之间的差异要大得多。
So, Ilya was my manager at Brain. Then he went on to found OpenAI. He asked me a number of times over the years whether I would like to join. I found it a little bit too edgy at the time. Then Transformers came. So we had a lot of work with that. And then COVID came and COVID was a tough time for the whole world, right? But Google totally closed. Google was reopening extremely slowly. So one part of me was I find it very hard to do remote work. I much prefer to work with people directly. That was one reason. But the other was also Google Brain when I joined it was a few dozen people, maybe 40, something like that. When I left it was 4,000 or 3,000 people, spread across multiple offices. It's very different to work in a small group and to work in a huge company. So with all this, I was like, you know, OpenAI though is in a much stabler state. We're doing language models. You know, something about this that may look like a good match. And I was like, okay, let me try. I've never worked in any company other than Google before, other than the university. So it was quite a change to the small startup group, but I like working in smaller groups. It has its pleasures, right? It has a little bit of a different intensity sometimes. In general, I found it very nice. On the other hand, Google, you know, has merged and made Gemini and I hear it's also a very nice place. I think in general the tech labs are more similar to each other than people think. There are some differences but I think if I look from the world, you know from the university in France, the difference between this university and any of the tech labs is much larger than between one lab or the other.
OpenAI 内部的研究团队是如何组织的?
How are the research teams organized within OpenAI?
嗯,它们是有组织的,但也不是非常组织化。我的意思是,我们确实有组织,但有些人有经理,我们有时会和他们沟通。嗯,不,但大多数情况下,人们会找到项目,有事情要做,对吧?比如改进多模态模型、改进推理、改进预训练、改进基础设施的某个部分。人们在这些方面工作,你知道,当我们经历这些部分时,对吧?有基础设施、预训练、推理。我认为大多数实验室的部分都是相同的。所以会有团队做这些事情,然后有时人们会换团队,有时会出现新的事情。总有一些小团队在做更冒险的事情,比如有时做扩散模型。然后你知道一些更冒险的东西,比如视频模型,变得很大,然后也许它们需要扩张。
Um, they're organized, they're not very organized. I mean we do organize them but some, you know, people have managers and we sometimes talk to them. Um, no, but mostly people find like projects, there are things to do, right? Like improve your multimodal models, improve your reasoning, improve your pre-training, improve whatever this part of the infrastructure. People work on it, you know, as we go through these parts, right? There is infrastructure, pre-training, reasoning. I think the parts are the same for most of the labs. So there will be teams doing these things and then sometimes people change teams, sometimes new things emerge. There is always some smaller teams doing like more adventurous stuff like diffusion models at times. Then you know some of the more adventurous stuff like video models gets big and then maybe they need to grow.
人们会竞争 GPU 资源吗?
Do people compete for GPU access?
我不认为是人在竞争。我认为更多的是项目在竞争 GPU 资源。这肯定是有的。另一方面,从 GPU 资源的大局来看,很多只是由技术本身决定的,对吧?目前,预训练在所有部分中使用的 GPU 最多。所以它需要最多的 GPU,对吧?强化学习的使用量在增长。现在,视频模型当然也使用大量 GPU。所以你需要这样分配。嗯,当然,你知道,人们会说,‘哦,但如果我有更多 GPU,我的东西会好得多。’我自己也说过很多次。嗯,所以你会推动,比如,我真的需要更多,然后有些人可能会说,但你知道,只有这么多。GPU 永远不够所有人用。所以有一部分是竞争,但大部分只是由当前技术的工作方式决定的。
I don't think it's so much people that compete. I think it's more projects that compete for GPU access. There's definitely some of that. On the other hand, like on the big picture of GPU access, a lot of this is just determined by how the technology works, right? Currently, pre-training just uses the most GPUs of all the parts. So it needs the most GPUs, right? RL is growing in use. Now, video models of course use a lot of GPUs too. So you need to split them like this. Um, then of course, you know, people will be, 'Oh, but my thing would be so much better if I had more GPUs.' And I've certainly said that a number of times too. Um, so then you kind of push like, you know, I really need more and then some people may say, well, but you know, there's only so much. There is never enough GPUs for everyone. So there is some part of the competition but the big part is just decided by the fact how the technology works currently.
很好。预训练接下来会怎样?我们谈到了数据,谈到了工程上大规模 GPU 算力的方面。未来一两年预训练会发生什么?
Great. What is next for pre-training? We talked about data, we talked about engineering a big GPU compute aspect to this. What happens to pre-training in the next year or two?
预训练,正如我所说,我认为它在科学上已经达到了 S 曲线的上端,但它可以平滑地扩展,意思是如果你投入更多算力,并且做得正确,你会得到更好的损失值,这非常困难但也很有价值。嗯,你不会得到像推动曲线那样的回报,但它通常会让模型变得更强大,这当然是你想做的。我认为人们在宏大叙事中有点低估的是,你知道,三四年前的 OpenAI,我甚至在那之前就加入了,是一个小型研究实验室,有一个叫 API 的产品,但你知道,它并不是那么大,例如在产品端没有 GPU 限制。所有 GPU 都只用于训练。所以人们很容易做出决定,比如,我们要训练 GPT-4,这将是有史以来最智能、最大的模型。我们关心小模型吗?我的意思是,我们关心它们是为了调试大模型的训练,但仅此而已。所以 GPT-4 是最智能的模型,它很棒,对吧?但后来发现,哦,有了这个聊天,现在我们有了十亿用户,你知道,人们每天想和它聊很多,你需要 GPU。所以你训练下一个巨大的模型,结果发现你无法满足需求,比如人们不愿意付足够的钱来和更大的模型聊天。所以从经济角度你只需要更小的模型。这当然发生在所有实验室,因为一旦经济因素介入,它变成了一个产品,你就必须比以前更仔细地考虑价格。
Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly, meaning if you put more compute you will get better losses if you do things right, which is extremely hard and that's valuable. Um, you don't get the same payoff as pushing the curve, but it generally just makes them all more capable and that's certainly something you want to do. I think what people underestimate a little bit in the big narrative is, you know, OpenAI three, four years ago, I joined even before that, was a small research lab with a product called API, but you know, it was not such a big, there was no GPU constraint on the product side, for example. All GPUs were just used for training. So it was very easy as a decision for the people to say, you know, we're going to train GPT-4, this will be the smartest and largest model ever. And what do we care about small models? I mean, we care of them as to make like to debug the training of the big model, but that's it. So GPT-4 was the smartest model and it was great, right? But then it turned out, oh, there is this chat and now we have a billion users and you know, people want to chat with it a lot every day and you need GPUs. So you train the next like huge model and it turns out you cannot satisfy this, like people will not want to pay you enough to chat with the bigger model. So you just economically need the smaller model. So this happened of course to all the labs because like the moment the economy arrived and it became a product, you had to start thinking about price much more carefully than before.
所以我认为这导致了这样一个事实:我们不再只是用有限的资金训练尽可能大的模型,而是说,不,我们要训练同样的东西,同样的质量,但更小、更便宜。以更低的成本提供相同质量的压力非常大。从某种意义上说,作为一名研究人员,这几乎让我有点难过。我非常喜欢这些巨大的模型。人们说人脑有 100 万亿个突触,数量级当然不是精确计算的,但我们的模型还没有 100 万亿个参数,所以也许我们应该达到这个目标。我当然很乐意,但你需要为此付费。所以我认为这可能就是人们认为预训练已经暂停的原因,因为很多精力都花在了训练更小的模型上。现在,人们重新发现了蒸馏有多么神奇。蒸馏意味着你可以训练一个大模型,然后把同样的知识从大模型(老师)传授给小模型。人们早就知道蒸馏,这是一篇很久以前的论文,但至少对 OpenAI 来说,也许当 Oriol 在那里时,这更像是谷歌的基因。但人们重新认识到这对经济的重要性。现在这也意味着训练这个大模型实际上是好的,因为你可以从中蒸馏出所有的小模型。所以现在可能有点回归了。一旦你意识到你有十亿用户,你需要 GPU,你就需要投资。当然,每个人都看到了这一点,有巨大的投资,但 GPU 还没有上线。当它们重新上线时,我认为这可能会促成人们所说的预训练的复苏。我们都明白,你可以蒸馏这个惊人的大模型,现在有足够的 GPU 来实际训练它,所以它正在复苏。但所有这些基本上都发生在相同的缩放曲线上,对吧?并不是我们不知道可以这样做。更多的是不同月份的不同需求改变了优先级。但我认为退一步思考大局是好的,那就是预训练一直有效。而且美妙的是,它甚至能与强化学习叠加。所以如果你在一个更好的模型上运行这个思维并行过程,它比在一个较小的模型上运行效果更好。
So I think this caused the fact that instead of just training the largest thing you can for the money you have, we said, well no, we're going to train the same thing, same quality, but smaller, cheaper. The pressure to give the same quality for less money is very large. In some senses, as a researcher, it almost makes me a little sad. I have a big love for these huge models. People say the human brain has 100 trillion synapses, orders of magnitude of course not exactly calculated, but our models don't have 100 trillion parameters yet, so maybe we should reach it. I would certainly love it, but then you need to pay for it. So I think this may be why people think that pre-training has paused, because a lot of effort went into training smaller models. Now on the side, people rediscovered how amazing distillation is. Distillation means you can train a big model and then put the same knowledge from the big one, the teacher, to the little one. People knew about distillation; it's a paper from a long time ago, but somehow at least for OpenAI, I think maybe it was more in Google's DNA when Oriol was there. But people rediscovered how important that is for the economics. Now it also means that training this huge model is actually good because you distill all the little ones from it. So now maybe there is a bit more of a return to it. Once you realize you have a billion users and you need the GPUs, you need to invest into them. Of course, everyone sees this, there's a huge investment, but the GPUs are not online yet. When they come back online, I think this may play into what people call the resurgence of pre-training. We both understand that you can distill this amazing big model, and there is now enough GPUs to actually train it, so it's resurging. But all of this fundamentally happens on the same scaling curve, right? It's not like we didn't know that you could do this. It's more that the different requirements of different months have changed the priorities. But I think it's good to step back and think of the big picture, which is that pre-training has always worked. And the beautiful thing is it even stacks with RL. So if you run this thinking parallel process on top of a better model, it works even better than if you run it on top of a smaller model.
当我听你谈到现代 AI 系统的演变时,我发现一个引人入胜的问题:LLM 加强化学习再加上许多其他东西的组合。曾经,也许是在深度学习时代,人们通常会说自己理解 AI 在微观层面是如何工作的,比如矩阵乘法方面,但一旦把所有东西放在一起,就不完全理解模型最终到底发生了什么。我知道过去几年在可解释性方面做了大量工作。但特别是对于那些非常复杂的系统,模型在做什么是越来越清楚,还是仍然存在某种黑箱元素?
One question that I find fascinating as I hear you speak about the evolution of modern AI systems has been this combination of LLM plus RL plus a lot of things going on. It used to be, at some point, and maybe that was back in the deep learning days, that people would routinely say that they understood how AI worked at a micro level, like the matrix multiplication aspect, but didn't fully understand once you had everything together what really happened at the end of the day in the model. I know there's been tons of work done on interpretability over the last couple of years, in particular. But particularly for those very complex systems, is it increasingly clear what the models do, or is there some element of black box that persists?
我会说两者都有。在理解模型方面取得了巨大进展。从根本上说,想想 ChatGPT 这个模型。它与十亿人谈论各种话题。它通过阅读整个互联网获得这些知识。显然,你无法识别——我无法理解里面发生了什么。我不知道整个互联网。我们能识别的是,就在上周,OpenAI 有一篇漂亮的论文,如果你告诉模型它的很多权重应该为零,它应该非常稀疏,那么你真的可以追踪它何时在思考某个特定事物,你可以追踪它实际在做什么。所以如果你把自己限制在这个范围内,并在模型内部真正研究它,那么你可以获得很多理解。模型中有电路;Anthropic 在这方面有很好的论文。对模型在更高层次上做什么的理解已经进步了很多。但即便如此,这仍然是对较小模型的理解,而不是最大的模型。但这并不是说这些模式不适用于更大的模型;它们适用。只是更大的模型同时做太多事情,以至于你能理解的东西有限。但我认为这个限制比人们想象的更根本。就像每一个非常复杂的系统。你只能理解那么多东西,然后就不理解了。
I would say both. There is huge progress in understanding models. Fundamentally, think of the model that is ChatGPT. It talks to a billion people about all kinds of topics. It gets this knowledge from reading all of the internet. Obviously, you cannot identify, like I cannot understand what's going on in there. I don't know the whole internet. What we can identify is, there was a beautiful paper just last week from OpenAI about if you tell the model that lots of its weights should be zeros, it should be very sparse, then you can really trace when it's thinking about one particular thing, you can trace what it's actually doing. So if you limit yourself to this and really study this inside a model, then you can get a lot of understanding. There are circuits in the models; Anthropic had great papers on that. The understanding of what the models are doing on a higher level has progressed a lot. But then it's still an understanding of what smaller models do, not the biggest ones. But it's not so much that these patterns don't apply to bigger models; they do. It's just that the bigger models do so many things at the same time that there is some limit to what you can understand. But I think this limit is a bit more fundamental than people think. It's like every very complex system. You can only understand so many things, and then you don't.
感谢你分享这些。我现在想谈谈 5.1,并深入了解一下你们在过去几周发布的所有最新内容,这些内容非常令人印象深刻。特别是作为一个用户,我认为 5.1 这个名称并不能公正地反映 5.1 和 5 之间的演变。从用户的角度来看,它感觉比数字所显示的要大得多。请带我们回顾一下从 GPT-4 到 5 再到 5.1 的演变。实际上发生了什么变化?
Thank you for all of this. I'd love now to talk about 5.1 and do a little bit of a deep dive on all the latest stuff that you guys have released in the last couple of weeks, which has been very impressive. In particular, as a user, I think that the 5.1 moniker doesn't do justice to the evolution between 5.1 and 5. It feels like a much larger improvement than the number would indicate, from my perspective as a user. Walk us maybe through the evolution from GPT-4 to 5 to 5.1. What has actually changed?
这是一个非常棘手的问题。我认为比你想象的要少。不,我的意思是,从 GPT-4 到 5,我认为最大的变化是推理,即强化学习和合成数据。正如我告诉你的,那个时期的预训练部分主要是为了降低成本,而不是提高质量。所以当然,价格也发生了巨大变化,我认为是一千倍,或者类似的数量级。从 4 到 5 的主要改进是增加了基于强化学习的推理,这允许生成合成数据,从而也改进了模型。这就是大局。除此之外,ChatGPT 现在是一个被很多人使用的产品。所以后训练团队学到了大量的经验教训,并增加了一些东西。显然,他们进行了实验:希望模型对你非常友好,结果它变得过于友好。现在当很多人使用它时,你需要非常小心安全问题。可能有处于困境的人在使用模型。模型需要在这些情况下做出合理的事情。它以前没有为此训练过。现在有了,这使得模型更好。但与此同时,你不想拒绝回答任何带有任何迹象的问题。所以当你处理这些事情时,你让模型在使用中变得更好,不仅对处于困境的人,而且对每个希望问题得到合理回答的人。还有那些被称为幻觉的东西。它在某种程度上仍然存在,但比两年前大大减少了。
That's a very tough question. I think less than you think. No, I mean, from GPT-4 to 5, I think the biggest thing that changed is reasoning, meaning RL and synthetic data. As I told you, the pre-training part in that time frame was mostly about making things cheaper, not making things better. So of course, the price has changed dramatically too, a thousand times I think, or some of these orders of magnitude. The main improvements from 4 to 5 is adding reasoning with reinforcement learning, and this allowed generating synthetic data which also improves the model. So that's the big picture. In addition to that, ChatGPT is now a product used by a lot of people. So the post-training team has learned a tremendous number of lessons and added things. Clearly, they experimented: wanted the model to be very nice to you, then it turned out to be too nice. Now when a lot of people use it, you need to be really careful about safety. There may be people in distress using the model. The model needs to do something reasonable in these cases. It was not trained for it before. Now it is, and it makes the model much better. But at the same time, you don't want to refuse to answer any question that has any sign of anything. So as you work on these things, you make the model much better in use, not just for the people in distress, but for everyone who wants questions answered, but the answers to be reasonable. And there were these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.
部分原因在于强化学习现在可以使用工具并收集数据,它还鼓励模型验证自己的行为。这是推理强化学习中涌现出来的现象,但你也需要添加数据,因为有时模型应该说‘我不知道’。所以你将这一点加入后训练数据。你说,‘嗯,我们确实需要仔细考虑模型在各种情况下应该如何回答人们。’从 5 到 5.1 主要是这类改进——主要是后训练方面的改进。
Some part of that is because reinforcement learning can now use tools and gather data, and it also encourages the model to verify what it's doing. So that's an emergent thing from this reinforcement learning of reasoning, but you also add data because you realize sometimes the model should say 'I don't know.' So you add this to the post-training data. You say, 'Well, we really need to give it a thought how the model should answer people in various situations.' The 5 to 5.1 is mostly this kind of improvement—it's mostly a post-training improvement.
是的。所以我想深入探讨这一点,因为这非常有趣。确实,作为 5.1 的一部分,它能够选择从书呆子到专业的不同风格。我想这是为了回应一些人在 ChatGPT 或 GPT-5 发布时怀念早期模型语义方面的问题。所以增加更多语气——这些都是后训练的内容。那么你是告诉模型这些是它应该如何回应的例子,这更像是一种监督范式,还是像强化学习那样用奖励来区分对错?这是如何运作的?
Yeah. So to double-click on this because this is super interesting. Indeed. As part of 5.1, there's the ability to choose different styles from nerdy to professional. And that's, I guess, in reaction to the fact that some people were missing the semantic aspect of earlier models when ChatGPT when GPT-5 came out. So adding more tones—that's all post-training stuff. So you tell the model those are examples of how you should respond, which is more like a supervising kind of paradigm, or is that RL like right or wrong with rewards? How does that work?
我不做后训练工作,它确实有很多古怪之处,但我认为主要部分确实是强化学习,你问‘好的,这个回答是讽刺的吗?这个回答像那样吗?’然后你说,‘好的,如果你被告知要讽刺,你应该这样回答。如果你被告知要有趣,试试这个。’所以我确实认为强化学习是其中的重要部分。
I don't work on post-training, and it certainly has a lot of quirks, but I think the main part is indeed RL, where you say, 'Okay, is this response cynical? Is this response like that?' And you say, 'Okay, if you were told to be cynical, this is how you should respond. If you were told to be funny, try this one.' So I do think RL is a big part of it.
在模型之间或不同版本之间,发布是否与预训练工作对齐,还是有时你有一个大的预训练工作,然后基于它推出几个模型?就在不久前——半年前——模型确实与技术内容对齐,对吧?它们要么与强化学习运行对齐,要么与预训练运行对齐。这就是为什么你有一个漂亮的模型叫 4o,它与预训练运行对齐,显然比与强化学习运行对齐的 o3 差,而 o3 是 o1 的后续,自然是因为你不能用 o2 这个名字,但它比 o4-mini 稍好,因为 o4-mini 是迷你版。你知道,我们有这个漂亮的模型选择器,人们出于某种原因认为这不是最好的命名。所以不,我的意思是这显然非常令人困惑,对吧?所以现在命名是按能力来的,对吧?GPT-5 是一个有能力的模型,5.1 是更有能力的模型。Mini 是较小的模型,能力稍弱但更快更便宜,而思考模型是做更多研究的模型。对吧?从这个意义上说,命名与任何具体技术脱钩了。你知道,5.1 可能只是后训练的东西,但也许 5.2 会针对新预训练的模型,也可能不是。但命名已经与技术脱钩,这也给了——随着 OpenAI 的发展,有很多项目,对吧?有强化学习和预训练,也许还有只是为了改进幻灯片之类的东西。通过蒸馏,你可以将多个项目整合到一个模型中。这很好,你不需要等待所有项目同时完成,你可以尝试定期整合——实际上确保作为产品对用户友好且良好——并独立于等待新的完整预训练运行(需要数月等)来做这件事。所以我觉得,尽管我有点怀念那些预训练模型编号就是模型编号的时代,但作为一个服务十亿用户的产品,也许不可避免的是,你应该根据用户期望来命名,而不是……
Between models or different versions of the models, are the releases aligned with pre-training efforts, or sometimes you have one big pre-training effort and several models that come out based on that? There used to be a time not that long ago—half a year distant past—where the models did have an alignment with technical stuff, right? So they would align either with RL runs or pre-training runs. That's why you had a beautiful model called 4o, which was aligned with a pre-training run, which was obviously worse than the o3 aligned with an RL run that was the follow-up to o1 naturally because you couldn't use the name o2, but it was slightly better than the o4-mini because that one was mini. And you know, we had this beautiful model picker, and people kind of thought this was not the best naming for some whatever reason. So no, I mean it was fairly obvious that this was very confusing, right? So now the naming is by capability, right? GPT-5 is a capable model, 5.1 is a more capable model. Mini is the smaller model that's slightly less capable but faster and cheaper, and the thinking models are the ones that do more research. Right? And in that sense, the naming is detached from any technical in particular. You know, 5.1 may be just a post-training thing, but maybe 5.2 towards the newly pre-trained model or maybe not. But the naming has detached from the technology, which also gives—as OpenAI has grown, there are a number of projects, right? There is RL and pre-training, and maybe something just to make slides better or whatnot. And with distillation, you have the ability to put a number of projects into one model. It's kind of nice that you don't need to wait on all of them to complete at the same time, and you can try to periodically put together—actually make sure that as a product it's nice to the users and good—and do this separately from waiting on the new full pre-training run that takes months and so on. So I feel like even though a little tear in my eye goes for the times where it was that pre-trained model number that was the number, as it's a product serving a billion users, it's maybe inevitable that you should name it by what the user should expect from it rather than...
在 5.1 中,你在告诉模型默认应该思考多长时间方面有了额外的粒度。模型如何决定它应该思考多长时间?
In 5.1, you have additional granularity in terms of telling the model how long it should think by default. How does the model decide how long it should think?
模型看到任务后,会自己决定思考多长时间。但你可以给它额外的信息——它经过训练,可以被告知‘更努力思考’,然后它会思考更长时间。所以你现在有能力引导这一点。我仍然认为重要的是要认识到:这是推理模型带来的根本变化——使用更多 token 来思考会提升能力,而且这种提升在给定算力下比预训练快得多,对吧?所以如果你给 GPT-5 长时间思考的能力,它可以解决那些——你知道,我们在数学奥林匹克和计算机科学奥林匹克中获得金牌的任务——所以能力惊人。同时,推理的基本训练方法非常局限于科学数据。所以它不像预训练那样广泛,我认为预训练模型在各方面感觉几乎均匀地好或坏。我的意思是,这仍然不是均匀的,因为它不像教人类,对吧?但推理模型甚至更——人们称之为‘锯齿状’,对吧?它们在某些方面有惊人能力,但在相近的方面却不那么强。这可能非常令人困惑。我一直喜欢这一点——这很奇怪,因为你可以说模型在数学奥林匹克上很棒。同时,我有一本给一年级女儿(她 5 岁)的数学书。我从这本书中拿了一道题,没有一个前沿模型能解出来,而你 10 秒就能解出来。所以这一点要记住。模型既惊人,也有做不好的任务。我可以给你看一个例子。我认为记住这一点很有趣。让我从 Gemini 3 开始,只是为了责怪竞争对手。
So the model sees the task. It will decide on its own a little bit how long it should think. But you can give it an additional—it's trained with an additional information that can tell it 'think even harder' and then it will think longer. So you have now the ability to steer that. I still think it is important to realize: this is the fundamental change that came with reasoning models—using more tokens to think increases your capability, and it increases it given the computation way faster than pre-training, right? So if you give GPT-5 the ability to think for long, it can solve tasks that—you know, we had these gold medal at mathematical olympiad and computer science olympiad—so amazing abilities. At the same time, the fundamental training method of reasoning is very limited to science data. So it's not as broad as the pre-training, which I think like pre-training models felt kind of almost uniformly good or bad at things. I mean, this was still not uniform because it's not like teaching humans, right? But the reasoning models are even more—people call it 'jagged', right? They have amazing abilities somewhere and then close by not so much. And that can be very confusing. It's something I always love—it's weird because you can say the model is amazing at mathematical olympiad. At the same time, I have a math book for my first grader daughter—she's 5 years old. I took one exercise from this math book and none of the frontier models is able to solve it, and you would be able to solve it in 10 seconds. So that's something to keep in mind. Models are both amazing and there are tasks that they cannot do very well. I can show you this as an example. I think it's quite interesting to keep in mind. Let me start with Gemini 3 just to blame the competitors.
好的,请说。
Yes, please.
所以它——你看到两边各有两组点,问题是:点的数量是偶数吗?如果你看它,你会觉得‘哦,它们像是两个相同的东西。’所以那应该是偶数。这是 5 岁孩子应该学的。但有一个点是共享的。所以现在那一定是奇数。对于这个简单的题,大概有 20 个点左右。Gemini 3 实际上做对了——它发现点的数量是偶数并说出来,这很好。
So it has—you see two groups of dots on both sides, and the question is: is the number of dots even? If you look at it, you see 'oh, they're like two identical things.' So that would be even. That's what the 5-year-old is supposed to learn. But there is one dot that's shared. So now that must be odd. For this simple one which has like, you know, I don't know, 20 dots or so. Gemini 3 actually does it right—it finds out that it's an even number of dots and it says that, and that's great.
然后还有另一个非常相似的谜题,只是现在有两座点阵山,底部还有一个共享的点。紧接着在上下文中你问:“那这个呢?”它思考了一下,完全没注意到那个共享的点,说数字是偶数。而就在刚才的上下文里,它刚看过第一个例子,怎么会错过呢?同样,GPT-5.1 在思考时也先解决了第一个,看到了那个点,说它是奇数;然后看到山,不知怎么就没看到那个点,说它是偶数。好消息是,如果你让它想久一点,或者只是让它再想一次,它就能看到。所以如果用 GPT-5 Pro,需要 15 分钟。而一个五岁小孩只需要 15 秒。GPT-5.1 Pro 会运行 Python 代码从图像中提取这些点,然后用循环计数。这就不太一样了。
And then you have another puzzle which is very similar except now there are two mountains of dots and there's also one dot shared at the bottom now and right in context right after that you ask okay how about this one and then it does some thinking and it just totally misses that there is a shared dot and it says the number is even and it's like in context where you've seen this first example how would you ever miss that you know and here is the same the exact same prompt for GPT 5.1 point when thinking and it also solves the first it sees the dot. It says it's odd and then it sees the mountains and somehow it doesn't see the dot and it says it's even. The nice thing is if you let it think longer this is like or if you just let it think again it will see it. So if you use GPT5 Pro it takes 15 minutes. So you know this is the human 5-year-old takes 15 seconds. The GPD51 Pro will run Python code to extract these dots from an image and then it will count them in a loop. So that's not quite
为什么会这样?是什么让模型出错了?
And why is that? What trips up the model?
我认为这主要是多模态部分的问题。模型才刚刚起步,你看第一个例子它们能解决,所以显然取得了一些进展,但它们还没有学会在多模态领域进行良好的推理,也没有学会利用上下文中的推理来进行下一步推理。上下文中的学习确实发生了,但从上下文推理中学习的能力仍然不强。不过这些都是众所周知的问题,模型只是训练得还不够。这是我们知道需要在训练中加入的东西。所以我认为这些会普遍改善。我确实认为有一个更深层次的问题:多模态会改善,这些例子会改善,随着前沿推进,一些东西会变得平滑,但问题是,是否还会有其他东西——那些你不需要教人类的东西?比如,你现在知道怎么用勺子和叉子,但如果叉子有四个尖而不是三个,你就得学新的——那将是机器学习的失败。我对泛化非常着迷,我认为这是最重要的课题,我一直认为这是机器学习和理解智能的关键。预训练有点不同,对吧?因为它随着模型规模的增大而增加数据,所以不一定增加泛化,只是使用了更多知识。我确实相信推理实际上能增加泛化,但现在我们在非常狭窄的领域训练它,所以可能还看不出来。但我认为 AI 中的大问题是:推理是否足以增加泛化,还是需要更通用的方法?我认为第一步是让推理更通用。正如我们之前谈到的,这是我的热情所在,也是我的研究方向。这里仍然有问题,对吧?我们推动模型,它们学会了我们教给它们的东西。但它们仍有局限,因为它们不生活在物理世界,不擅长多模态,推理还很年轻,我们做推理的方式还有很多缺陷。但一旦我们解决了这些问题,就会有一个大问题:这够了吗?还是需要其他更大的东西来让模型更好地泛化,这样我们就不必在训练数据中教它每一个具体的东西,它就能自己学习和泛化。我认为这是最迷人的问题。但我也认为解决这类问题的一个好方法是先解决所有前置问题。你知道,在靠近之前你无法知道是否有墙,因为 AI 发展非常快。有人说这就像在雾中高速驾驶,你永远不知道离目标有多远。所以我们正在前进,学到了很多。
I think this is mostly multimodal part. The models are just starting, like you see the first example they managed. So they've clearly made some progress, but they have not yet learned to do good reasoning in multimodal domains and they have not yet learned to use one reasoning in context to do the next reasoning. What is written in context is learning in context happens but learning from reasoning in context is still not very strong. All of these though are things that are very well known and like the models are just not trained enough to do this. It's just something we know we need to add into training. So I think these are things that will generally improve. I do think there is a deeper question whether so you know like multimodal will improve this will improve like we keep finding these examples. Though as the frontier will move it will certainly move forward some things will smooth out but the question is will it still be just other things that you don't need to teach the human like every you know okay now you know how to use a spoon and a fork but now if the fork has four instead of three ends then you need to learn a new that would be a failure of machine learning you know I am fascinated by generalization I think that's the most important topic I always thought this was the key topic in machine learning in general and in understanding intelligence. Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge. I do believe that reasoning actually increases generalization, but now we train it on such narrow domains that it may still be to see. But I think the big question in all of AI is is reasoning enough to increase generalization or do you need like more general methods? I think the first step is to make reasoning more general. As we talked before, that's my passion. That's also what I work on. There is still something there, right? We push the models. They learn things that are around what we teach them. They still have limitations because they don't live in the physical world, because they're not very good at multimodal, because reasoning is very young and there's a lot of bugs in how we do it yet. But once we fix that there will be this big question is that enough or is there like something other big to make models generalize better so we don't need to teach it every particular thing in the training data that it just learns and generalizes. I think that's the most fascinating question. But I also think a good way to approach a question like that is to first solve everything that leads up to it. You know, you cannot know whether there is a wall or not until you come close to it because otherwise there you know we AI is moving very fast. Someone said it's like driving fast in a fog. You never know how far or close you are. So we are moving we are learning a lot.
那么这是否意味着那个核心问题——基本上像孩子一样用很少的数据学习,而孩子能做的事情即使最强大的模型也做不到?所以,正如你所说,要拆解这个问题:在推理上取得进展,看看推理能带我们走多远,然后另一个问题是,我们是否需要完全不同的架构?这就涉及到 Yann LeCun 的工作。你是否看到 Transformer 之外有前景的根本性架构变化引起了你的注意,并且感觉它们可能是未来值得探索的严肃路径?
And does that mean that central question of basically learning with very little data the way a child would and the fact that the child is able to do things that even the most powerful model cannot do. So this as you said to unpack this making progress on reasoning and showing how far we can get into generalization with reasoning and then the separate question is as you said whether we need an entire different architecture and that's where we get into for example Yan LeCun's work do you see promising fundamental architectural changes outside of transformers that have caught your attention and feel like they could be a serious path to explore in the future.
我认为有很多人在尝试漂亮的工作。ARC 挑战赛激励了一群人。他们的模型现在非常小,能很好地解决这些问题,但用的方法我不确定是否真正通用。我们得看看。Yann LeCun 一直在推动其他方法。我觉得他的方法更偏向多模态部分。但也许如果你解决了多模态,对吧?也许如果你做 JEPA,它也有助于你的其他理解。仍然有很多人在推动基础科学。这可能不像那些推动进展的东西那样常上新闻,但无论你做什么,它可能都会在某个 GPU 上运行。如果你有万亿美金的新 GPU,旧的 GPU 也会更容易获得。所以我认为 LLM AI 在更传统一面的增长也帮助人们更容易地在各种事情上运行更多的实验性研究项目。所以我认为有很多探索,很多想法。但在更大规模上实现它们仍然有点困难。工程部分是最主要的瓶颈。我的意思是,当你真正扩大规模时,GPU 也是瓶颈,但实现一个比单机更大的东西是一个实验性研究项目,你没有团队去做。我认为这比它应该有的难度更大。但你知道,Codex 可能会达到那个水平,或者 Coder——这是 AI 研究人员寄予厚望的东西,可以帮助他们自己和其他研究人员。如果你能说:“嘿,Codex,这是想法,我说的很清楚,请实现它,让它在这八台机器或一百台机器上快速运行。”那将非常棒。它现在还不能完全做到,但你知道,它正在越来越能做到。我认为这就是 OpenAI 所说的:他们说,我们希望到明年年底有一个 AI 实习生。我是这么理解的。
I think there is a lot of beautiful work that people are trying out. You know, the ARC challenges inspired one set of people. Their models now that are very small and solve them very well, but with methods that I'm not sure are actually general. We'll need to see. Yann LeCun has been pushing for other methods. So I feel like his approach is more towards the multimodal part. But maybe if you solve multimodal, right? Maybe if you do JEPA it also helps your other understanding. There is a lot of people pushing fundamental science still. It's maybe not so much in the news as the things that push but whatever you do it will probably run on some GPU. If you get a trillion dollars of new GPUs, the old GPUs will be much easier to get also for so I think this growth in LLM AI on the more traditional side is also helping people to have an easier time to run more experimental research projects on various things. So I think there is a lot of exploration, a lot of ideas. It's still a little hard to implement them at a higher scale. The engineering part is the biggest bottleneck. I mean GPUs are a bottleneck too when you scale really up but implementing something that's larger than one machine it's an experimental research project so you don't have a team to do that I think that's still harder than it should be but you know Codex may get there or Coder this is the thing where AI researchers have great hope to help themselves and also other researchers is that if you could just say hey Codex this is the idea and it's fairly clear what I'm saying please just implement it so it runs fast on this eight machine setup or 100 machine setup. That would be amazing. It's not quite capable of doing that yet. But, you know, it's capable of doing this more and more. I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year. That's how I understand this.
你知道,有人能帮我们吗?Codex 实现这些能力的一部分是否取决于它能运行多长时间?问题的背景是,就在两天前,我们录制这期节目时,你们发布了 GPT-5.1 C Codex Max,被描述为一个前沿智能体编码模型,基于真实世界的软件工程任务训练,专为长时间运行的工作流设计,并使用压缩技术跨多个上下文窗口操作,处理数百万个 token。所以我很想深入了解一下。长时间运行意味着什么?这是工程问题还是模型问题?然后也许再谈谈压缩。
You know, can someone help us? And is part of Codex's path to be able to do some of this revolve around how long it can run? Context behind the question being that again, like two days ago as we record this, you guys released GPT-5.1 C Codex Max, described as a frontier agent coding model trained on real-world software engineering tasks designed for long-running workflows and using compaction to operate across multiple context windows in millions of tokens. So I'd be interested in unpacking some of this. What does it mean to run for a very long time? Is that an engineering problem or a model problem? And then maybe a word on compaction.
这既是工程问题也是模型问题。你知道,你想做某个工程任务,比如写代码——你有一个机器学习想法。你希望 Codex 为你实现它。在简单的东西上测试,找到 bug。所以它需要运行这个东西。这不是一小时内能完成的事,对吧?那是你要花一周时间的事。所以模型需要花费相当长的时间,因为它需要运行程序、等待结果、然后修复。模型不会凭空想出正确的代码,对吧?它和我们一样。它需要经历这个过程,而在这个过程中,由于它在训练中没有接触过非常长的任务——也许有很少,但肯定没有持续一周的——它可能会迷失。它可能会开始循环或做一些奇怪的事情。这当然不是我们想要的。所以我们尝试以某种方式训练来避免这种情况,但它确实会发生。那么,你如何让模型实际运行一个需要更大反馈循环的过程而不出错呢?另一件事是 Transformer 有所谓的上下文。它们会记住当前运行中看到的所有内容,而这可能会超出运行可用的内存,注意力矩阵是 n×n,其中 n 是长度,所以它们可能变得巨大。所以,与其保留所有内容,你可以说,好吧,我让模型在旁边总结过去最重要的事情,把它重新放入上下文,并忘记部分内容。对吧?这是一种非常基本的遗忘形式,即压缩。对吧?如果你重复这样做,就可以运行更长时间。但同样,你需要训练模型做到这一点。当你这样做时,它在某种程度上是有效的。我认为它还没有好到足以取代 AI 研究员。它取得了相当多的进展。我认为研究方面另一个被低估但非常重要的进展是让模型能够连接所有这些事物。所以模型现在常规使用网络搜索和 Python 等工具,但要运行在 GPU 上,访问集群。训练模型做到这一点很难,因为你需要为模型专用资源,而且存在安全问题。模型如何与外部世界连接?这从根本上是一个非常困难的问题,因为如果你以无限制的方式连接,你可能会在现实世界中造成破坏,而我们不希望模型为我们破坏东西。所以这是人们大量工作的部分。这与安全重叠,对吧?你需要非常好的安全性才能让模型继续训练它们需要训练的东西。
So it is both an engineering and model problem. You know, you want to do some engineering task like write—you have some machine learning idea. You want Codex to implement it for you. Test it on some simple thing. Find the bugs. So it needs to run this thing. This is not something you would do in an hour, right? That's something you'd spend a week on. So the model needs to spend a considerable amount of time because it needs to run things, wait for the results, then fix them. The model is not like it's going to come up with the correct code out from thin air, right? It's just like us. It needs to go through the process, and often times in the process, since it was not trained on anything very long in its training—or maybe very few, but nothing certainly that went on for a week—it can get lost. It can start doing loops or doing something weird. That's of course not something you want. So we try to train in a way that makes it not happen, but it does. So how can you make the model actually run a process that requires this larger feedback loop without tripping up? And the other thing is transformers have this thing called context. So they remember all the things that they have seen in the current run, and that can just exceed the memory available for your run, and the attention matrices are n by n where n is this length, so they can get huge. So instead of keeping everything, you say, well, I'm going to just ask the model on the side to summarize the most important thing from the past, put it in context anew, and forget some part of it. Right? So it's a very basic form of forgetting, the compaction. Right? And that allows you to run for much longer if you do this repeatedly. But again, you need to train the model to do that. And when you do it, it works to some extent. I don't think it works well enough to replace an AI researcher yet. It made a fair bit of progress. I think another part of progress that's a little understated on the research side but is very important is allowing the model to connect to all of these things. So models now use tools like web search and Python routinely, but to run on a GPU, to have access to a cluster. It's hard to train models with that because then you need to dedicate for the model to use, and that has security problems. How do models connect with the external world? It's a fundamentally very hard problem because when you connect in an unlimited way, you can break things in the real world, and we don't like models to break things for us. So that's a part where people work a lot. This overlaps with security, right? You need to have very good security to allow models to go on and train on the things they need to train.
像我这样的人——VC、创始人和初创公司——在目睹 OpenAI 的所有进展时经常思考的一个主题是:随着模型变得越来越通用,拥有更多的智能体能力,能够长时间运行,进入科学和数学等领域,最近有报道称投资银行家被雇来帮助提高模型执行基础投资银行工作的能力——考虑到所有这些,是否存在一个世界,基本上模型,或者也许只是一个模型,做所有事情?我不知道这是否是 AGI,我们不一定进入那个辩论,但那些在模型之上构建产品的人还剩下什么?
One theme that people like me, VCs, founders, and startups think about a lot as we see all the progress at OpenAI is: as the models keep getting more general with more agent capabilities, the ability to run for a very long time, going into areas like science and math, and recently it was reported that investment bankers were hired to help improve the model's capability to do grunt investment banking work—all of that taken into account, is there a world where basically models, or maybe just one model, does everything? And I don't know if that's AGI, let's not necessarily go into that debate, but what's left for people that build products that sit on top of models?
我刚刚给你看了一个五岁孩子能做的练习,模型却做不了。我认为我们需要记住这一点。
I just showed you a 5-year-old exercise that the model doesn't do. I think we need to keep that in mind.
所以你是说有希望?
So you're saying there's hope?
有希望。下一个模型会做到的,对吧?
There is hope. The next model will do it, right?
那种希望,好吧。
That hope, okay.
嗯,对我来说,是的。我仍然认为我们还有一段路要走。模型进步很快,所以很有希望这类事情会越来越少。但另一方面,目前你不需要深入搜索就能找到那些你真正希望人类来完成的任务,因为模型还不够好。另一方面,你知道,Transformer 论文始于翻译。我最近参加了一个翻译行业会议。自那以后,翻译行业大幅增长,没有萎缩。需要翻译的内容更多了。译者的报酬也更高了。问题是:如果模型在大多数情况下已经如此出色,为什么还需要译者?答案是——想象一下,你为一家报纸做广告,但用的是你不懂的语言,GPT-5 几乎肯定会为你正确翻译。如果是翻译成西班牙语、法语或任何高资源语言,你会不经过懂那种语言的人审阅就直接发布吗?如果那是 ChatGPT 的界面,有十亿人会看到,你会直接发布吗?这可能是信任问题,对吧?但如果你有一百万用户、十亿用户,也许你会花 50 美元请人在翻译前看一眼。所以这是一个从根本上已经完全自动化的行业,对吧?但仍然存在信任问题,我认为我们将长期应对这个问题。还有一些事情你希望由人来做。我不认为我们会无事可做,但这并不意味着我们做的某些事情不会发生巨大变化。这对从事这些工作的人来说可能非常痛苦,所以这是一个严肃的话题,很高兴人们正在参与讨论。但我不认为会出现全球性的无事可做的情况。
Well, for me, yes. I still think we have some way to go. Model progress has been rapid, so there is good hope there will be less and less of this. But on the other hand, for now, you don't need to do a deep search to find things where you'd really want a human to do that task because the model's not super good. On the other hand, you know, the transformer paper started with translation. I recently went to a translation industry conference. The translation industry has grown considerably since then. It has not shrunk. There's more translations to be done. Translators are paid more. The question is: why would you even want a translator if the model's so good in most cases? The answer is sometimes—imagine you do a listing for a newspaper but in a language you don't know, and GPT-5 will almost certainly translate it correctly for you. If it's into Spanish or French or any high-resource language, would you still publish it without having a human who speaks that language look at it? Would you publish it if it's a UI of ChatGPT that a billion people are going to see? It's a question of trust, probably, right? But if you have a million users, a billion users, maybe you will pay the $50 for someone to just take a look over it before you translate it. So this is an industry that is fundamentally totally automated, right? There is still the question of trust, and I think that's a question we will grapple with for a long time. There are also just things you want a person to do. I don't think we will have no things to do, but that doesn't mean that some things we do may not dramatically change. And that can be very painful for people who do these things, and so this is a serious topic that happy people are engaging with. But I don't think there will be this global lack of anything for people to do.
也许作为最后一个问题,为了帮助我们了解 AI 前沿的人们目前在想什么或做什么,你知道,一些可能看到的主题包括持续学习、世界模型、机器人技术、具身智能。
And maybe as a last question to help us get a sense for what people at the frontier of AI are currently thinking about or working on, you know, some of the topics that one may see are things like continual learning, world models, robotics, embodied intelligence.
除了你提到的提示多模态之外,你个人觉得什么研究领域特别有趣?
What do you personally find very interesting in addition to what you mentioned—a prompt multimodal? But what do you personally find really interesting as a research area?
嗯,我一直觉得这种通用数据强化学习是我的心头好,也是我研究的方向。幸运的是,例如机器人技术可能只是说明我们在多模态方面做得还不够好,在通用强化学习和通用推理方面也还不够好。一旦我们在多模态上做得很好,并且能够将推理推广到机器人所需的物理领域,我认为将会看到惊人的进展。当这种情况发生时,我感觉——鉴于很多公司正在推出远程操作或手套操作之类的硬件——我的猜测是,当我们取得这一进展时——可能明年,也可能再过几年——硬件可能已经就位了。家里有机器人可能会是一个巨大的可见变化,比聊天机器人更明显。我的意思是,考虑到我们在旧金山对自动驾驶汽车适应得有多快,可能它只在头两天可见,然后就会变成‘是啊,当然。机器人在那儿。从我记事起它就在打扫,就像过去三个月一样。’我惊讶于我们适应这些东西的速度,对吧?旧金山的自动驾驶汽车,人们很快就习惯了。所以机器人可能也会这样。不过,我认为当它发生时,确实会极大地改变我们对世界的感知。但硬件很难,对吧?机器人可能在房子里出事故。你需要非常小心。所以可能需要更长时间来部署它们,并真正使其成为可扩展的业务。我们拭目以待。令人惊叹的是,我们现在可以开始想,‘是的,也许那很快就会到来。’
Well, I always find this general data reinforcement learning is my pet peeve and what I work on. Luckily, that for example, robotics is probably just an illustration that we are not doing that well in multimodal and that we're not doing that well in general RL in general reasoning yet. The moment we do really well in multimodal and we manage to generalize reasoning to the physical domains that the robot needs, I think it will see amazing progress. When it does, I have a feeling given that a lot of companies are launching hardware that's kind of teleoperated or glove operated or something. So my suspicion is by the moment we make this progress—which maybe it will be next year, maybe it will be in a few more years—but the hardware may be there by then. And having a robot in a home may be like a big visible change, more visible than chat. I mean, given how quickly we got used to the self-driving cars in San Francisco, maybe it will be only visible for like the first two days and then be, 'Yeah, sure. The robot's there. It's always been cleaning since I can remember like the last three months.' It's stunning to me how quickly we get used to these things, right? The self-driving cars in San Francisco are something people got used to so quickly. So maybe this will happen for robots, too. Nevertheless, I do think it will be quite dramatic in our perception of the world when it happens. Hardware is hard though, right? Robots may have accidents in the house. You need to be very careful. So maybe it will take longer to deploy them and actually make it a scalable business. We'll see. It is amazing that we are at this point where we can start thinking like, 'Yes, maybe that will come soon.'
Łukasz,这真是太棒了。非常感谢你今天抽出时间与我们交流。
Lucas, it's been absolutely wonderful. Thank you so much for spending time with us today.
非常感谢,Matt。谢谢你的邀请。很高兴和你交谈。
Thank you very much, Matt. Thank you for the invitation. Great to talk to you.
嗨,我是 Matt Turk。感谢收听本期《Mad Podcast》。如果你喜欢这期节目,如果你还没订阅,希望你能考虑订阅,或者在你观看或收听本期节目的平台上留下好评或评论。这真的有助于我们发展播客并邀请到优秀的嘉宾。谢谢,下期再见。
Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks, and see you at the next episode.