从西塞罗到世界冠军:Noam Brown 谈外交游戏、AI 与图灵测试

From Cicero to World Champion: Noam Brown on Diplomacy, AI, and the Turing Test

诺姆·布朗 Noam Brown · Latent Space · 2025-06-19 · 约 78 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Noam Brown 分享他从开发西塞罗到赢得世界外交锦标赛的历程,探讨 AI 在游戏中的演进及其对 AI 安全的影响。

Noam Brown discusses his journey from building Cicero to winning the World Diplomacy Championship, the evolution of AI in games, and the implications for AI safety.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 36)

全文 · Full transcript(中英对照)

引言与外交背景 Introduction and Diplomacy Background

Host

大家好,欢迎收听 Latent Space 播客。我是 Alessio,Decibel 的合伙人兼 CTO,和我一起的是联合主持人 Brooks,Small Eye 的创始人。你好,你好。我们在一个假日星期一录制。今天我们有来自 OpenAI 的 Noam Brown,欢迎你。谢谢。很高兴你终于来了。很多人听过你的分享。你在播客上很慷慨地花时间,比如 Lex Fridman。你最近还做了一个 TED 演讲,谈论思考范式。但我觉得你最近最有趣的成就是赢得了世界外交锦标赛。2022 年,你构建了 Cicero,达到了人类玩家的前 10%。我的开场问题是,自从研究 Cicero 以来,你个人玩外交游戏的方式有什么变化?

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO of Decibel, and I'm joined by my co-host Brooks, founder of Small Eye. Hello, hello. And we're here recording on a holiday Monday. With Noam Brown from OpenAI, welcome. Thank you. So glad to have you finally join us. A lot of people have heard you. You've been rather generous with your time on a podcast, Lex Fridman. And you've done a TED Talk recently just talking about the thinking paradigm. But I think perhaps your most interesting recent achievement is winning the World Diplomacy Championship. In 2022, you built Cicero, which was top 10% of human players. I guess my opening question is, how has your diplomacy playing changed since working on Cicero and then now personally playing it?

Noam Brown

当你研究这些游戏时,你必须足够了解游戏才能调试你的机器人。因为如果机器人做了人类通常不会做的激进事情,你不确定那是错误、漏洞,还是机器人真的很聪明。我们在研究外交游戏时,我深入钻研以更好地理解游戏。我参加了锦标赛,看了很多教程和评论视频。在这个过程中,我进步了。看到机器人的行为有时也教会了我游戏。2022 年底我们发布 Cicero 后,我仍然觉得这个游戏很迷人,所以继续玩。这让我在几个月前赢得了 2025 年的世界冠军。

When you work on these games, you kind of have to understand the game well enough to debug your bot. Because if the bot does something radical that humans typically wouldn't do, you're not sure if it's a mistake or a bug, or if it's actually the bot being brilliant. When we were working on Diplomacy, I did a deep dive to understand the game better. I played in tournaments, watched a lot of tutorial and commentary videos. Over that process, I got better. And seeing the bot's behavior sometimes taught me about the game as well. When we released Cicero in late 2022, I still found the game fascinating, so I kept playing. That led to me winning the World Championship in 2025, just a couple months ago.

Host

总是有人机协作的半人马系统的问题。有没有类似围棋中发生的情况,你更新了你的玩法?

There's always a question of centaur systems where humans and machines work together. Was there an equivalent of what happened in Go where you updated your play style?

Noam Brown

如果你问我是否在锦标赛中使用了 Cicero,答案是没有。不过,看到机器人的玩法并从中汲取灵感确实帮助了我。

If you're asking if I used Cicero when I played in the tournament, the answer is no. Seeing the way the bot played and taking inspiration from it did help me in the tournament, though.

Host

现在人们玩外交游戏时每次都会问图灵问题吗?试图判断他们玩的人是不是机器人?

Do people now ask Turing questions every single time when they're playing Diplomacy? To try to tell if the person they're playing with is a bot?

Noam Brown

我们研究 Cicero 时很有趣,因为我们没有最好的语言模型。语言模型的质量是瓶颈。有时机器人会说奇怪的话。99% 的时间没问题,但偶尔会幻觉。有人提到之前说过的话,机器人会否认。人们会认为那是对方累了、醉了或是在 troll。那是因为人们没料到有机器人。我们担心人们会发现有机器人并一直警惕。如果你在找,就能发现。现在它已经公开,人们知道要留意,所以更容易发现。不过,自 2022 年以来语言模型进步了很多。GPT-4o 和 o3 能通过图灵测试,所以我觉得图灵问题没什么用。Cicero 非常小,只有 27 亿参数。我们在项目中意识到更大的语言模型很有帮助。

It was really interesting when we were working on Cicero because we didn't have the best language models. We were bottlenecked on the quality of the language models. Sometimes the bot would say bizarre things. 99% of the time it was fine, but occasionally it would hallucinate. Someone would reference something said earlier, and the bot would deny it. People would shrug it off as the person being tired or drunk or trolling. That's because people weren't expecting a bot. We were scared people would figure out there's a bot and always be on the lookout. If you're looking for it, you can spot it. Now that it's announced, people know to look for it, so they'd have an easier time spotting it. That said, language models have gotten a lot better since 2022. GPT-4o and o3 are passing the Turing test, so I don't think Turing questions would make a difference. Cicero was very small, 2.7B parameters. We realized over the project that larger language models help a lot.

安全感知与外交基准 Safety Perception and Diplomacy as Benchmark

Host

你怎么看待今天对 AI 的看法和安全讨论?你构建了一个善于说服人们赢得游戏的机器人。今天的研究机构可能会说他们不研究这类问题。你怎么看这种二分法?

How do you think about today's perception of AI and the safety discourse? You built a bot good at persuading people to win a game. Today labs might say they don't work on that type of problem. How do you think about that dichotomy?

Noam Brown

老实说,我们发布 Cicero 后,很多 AI 安全社区的人对这项研究非常满意,因为它是一个高度可控的系统。我们让 Cicero 基于某些具体行动,赋予它很强的可操控性。它不是一个随意运行的语言模型;有一个完整的推理系统引导它的交互。实际上,很多研究人员联系我说这可能是实现这些系统安全性的好方法。

Honestly, after we released Cicero, a lot of the AI safety community was really happy with the research because it was a very controllable system. We conditioned Cicero on certain concrete actions, giving it a lot of steerability. It's not just a language model running loose; there's a whole reasoning system steering how it interacts. Actually, many researchers reached out to me saying this could be a really good way to achieve safety with these systems.

Host

你有没有在外交游戏上更新或测试 O 系列模型?你会期待更大的差异吗?

Have you updated or tested O-series models on Diplomacy? Would you expect a lot more difference?

Noam Brown

我没有。我在 Twitter 上说过这会是一个很好的基准。我很想看到所有领先的机器人互相玩一局外交游戏,看看谁最好。有几个人受到启发,正在构建这些基准。据我所知,它们目前表现不太好。但我认为这是一个迷人的基准,尝试一下会很酷。

I have not. I said on Twitter that this would be a great benchmark. I'd love to see all leading bots play a game of Diplomacy with each other and see who does best. A couple people have taken inspiration and are building out these benchmarks. My understanding is that they don't do very well right now. But I think it's a fascinating benchmark and would be really cool to try out.

氛围变化与轨迹 Vibe changes and trajectory

Host

整体氛围有什么变化?你说过你非常兴奋能从化学等领域的专家那里学习,比如他们如何评估 O 系列模型。从去年年底到现在,你的看法有什么更新?

How have the vibes changed just in general? You said you were very excited to learn from domain experts like in chemistry, how they review the O series models. How have you updated since, let's say, end of last year?

Noam Brown

我认为轨迹在开发周期早期就已经很清晰了。从那以后发生的一切都基本符合我的预期。所以我不认为我对未来方向的看法有多大改变。我认为我们将继续看到这个范式快速进步,即使今天也是如此。从 O1 preview 到 O1 再到 O3,我们看到了持续的进步。未来我们还会继续看到这一点。我认为我们还会看到这些模型能力的拓宽。我们将开始看到智能体式行为。我们已经开始看到智能体式行为了。对我来说,O3 我在日常生活中大量使用。我觉得它非常有用。尤其是它现在可以浏览网页,替我进行有意义的研究。这有点像迷你版的深度研究,你可以在 3 分钟内得到回应。所以我认为随着时间的推移,它会变得越来越有用、越来越强大,而且速度很快。

I think the trajectory was pretty clear pretty early on in the development cycle. And everything that's unfolded since then has been pretty on track for what I expected. So I wouldn't say that my perception of where things are going has changed that much. I think we're going to continue to see this paradigm progress rapidly, and that's true even today. We saw that with going from O1 preview to O1 to O3, consistent progress. And we're going to continue to see that going forward. I think we're going to see a broadening of what these models can do as well. We're going to start seeing agentic behavior. We're already starting to see agentic behavior. For me, O3, I've been using it a ton in my day-to-day life. I just find it so useful. Especially the fact that it can now browse the web and do meaningful research on my behalf. It's kind of like a mini deep research that you can get a response in 3 minutes. So I think it's just going to continue to become more and more useful and more powerful as time goes on, and pretty quickly.

深度研究在非可验证领域的证明 Deep research as proof in non-verifiable domains

Host

说到深度研究,你发推说如果需要证明我们可以在非可验证领域做到这一点,深度研究就是一个很好的例子。你能谈谈人们可能忽略了什么吗?我感觉我经常听到这种说法:在编码和数学上很容易,但在其他领域就不行。我经常收到这个问题,包括来自相当资深的 AI 研究者:我们看到这些推理模型在数学和编码这些容易验证的领域表现出色,但它们能否在成功定义不那么明确的领域取得成功?

Talking about deep research, you tweeted about if you need proof that we can do this in non-verifiable domains, deep research is a great example. Can you talk about if there's something that people are missing? I feel like I hear that repeated a lot. It's like, it's easy to do in coding and math, but not in these other domains. I frequently get this question, including from pretty established AI researchers, that okay, we're seeing these reasoning models excel in math and coding in these easily verifiable domains, but are they ever going to succeed in domains where success is less well defined?

Noam Brown

我很惊讶这种看法如此普遍,因为我们已经发布了深度研究,人们可以试用。人们确实在使用它,它非常受欢迎。这显然是一个没有容易验证的成功指标的领域。你能生成的最佳研究报告是什么?然而这些模型在这个领域表现得非常好。所以我认为这是一个存在性证明,表明这些模型可以在没有那么容易验证奖励的任务上取得成功。

I'm surprised that this is such a common perception because we've released deep research and people can try it out. People do use it. It's very popular. And that is very clearly a domain where you don't have an easily verifiable metric for success. What is the best research report that you can generate? And yet these models are doing extremely well at this domain. So I think that's an existence proof that these models can succeed in tasks that don't have as easily verifiable rewards.

Host

是不是也因为不一定有错误答案?比如深度研究质量有一个范围,对吧?你可以有一份看起来不错但信息一般的报告,然后还有一份很棒的报告。你认为人们在得到结果时很难理解其中的区别吗?

Is it because there's also not necessarily a wrong answer? Like there's a spectrum of deep research quality, right? You can have a report that looks good, but the information is kind of so-so. And then you have a great report. Do you think people have a hard time understanding the difference when they get the result?

Noam Brown

我的印象是,人们在得到结果时确实能理解其中的区别,而且他们对深度研究结果的质量感到惊讶。当然不是 100% 完美。它可以更好,我们也会让它更好。但我认为人们能区分好报告和坏报告,当然也能区分好报告和一般报告。这足以反馈循环来构建产品和改进模型性能。我的意思是,如果人们无法区分输出之间的差异,那么你在进步上爬山就无关紧要了。这些模型会在有成功衡量标准的领域变得更好。现在,我认为那种必须容易验证之类的想法,我不认为这是真的。我认为你可以让这些模型在成功很难定义的领域也表现出色。有时甚至可能是主观的。

My impression is that people do understand the difference when they get a result, and I think they're surprised at how good the deep research results are. It's certainly not 100%. It could be better and we're going to make it better. But I think people can tell the difference between a good report and a bad report, and certainly between a good report and a mediocre report. And that's enough to feed the loop to build the product and improve the model performance. I mean, if you're in a situation where people can't tell the difference between the outputs, then it doesn't really matter if you're hill climbing on progress. These models are going to get better at domains where there is a measure of success. Now, I think this idea that it has to be easily verifiable or something like that, I don't think that's true. I think you can have these models do well even in domains where success is a very difficult thing to define. Could sometimes even be subjective.

快慢思考类比局限 Thinking fast and slow analogy limitations

Host

你为思考模型做了思考快与慢的类比。我认为现在这个想法已经相当普及了,即这是下一个 Scaling 范式。所有类比都不完美。思考快与慢或系统一系统二在哪些方面不适用于我们实际扩展这些东西的方式?

You've done the thinking fast and slow analogy for thinking models. I think it's reasonably well diffused now, the idea that this is the next scaling paradigm. All analogies are imperfect. What is one way in which thinking fast and slow or system one system two doesn't transfer to how we actually scale these things?

Noam Brown

我认为有一点被低估了:预训练模型需要一定水平的能力才能真正从这种额外思考中受益。这就是为什么你看到推理范式在那个时候出现。我认为它本可以更早出现,但如果你在 GPT-2 之上尝试推理范式,我认为它几乎不会带来任何收益。这是涌现吗?很难说一定是涌现,但我没有做测量来明确定义。但我认为这很清楚。人们尝试在非常小的模型上使用思维链,发现它没有任何作用。然后你转向更大的模型,它开始有提升。我认为关于这种行为在多大程度上是涌现的有很多争论,但显然存在差异。所以这不像是有两个独立的范式。我认为它们是相关的,因为你需要模型有一定水平的系统一能力,才能让系统二从中受益。

One thing that I think is underappreciated is that the pre-trained models need a certain level of capability in order to really benefit from this extra thinking. This is kind of why you've seen the reasoning paradigm emerge around the time that it did. I think it could have happened earlier, but if you try to do the reasoning paradigm on top of GPT-2, I don't think it would have gotten you almost anything. Is this emergence? Hard to say if it's emergence necessarily, but I haven't done the measurements to really define that clearly. But I think it's pretty clear. People tried chain of thought with really small models and saw that it just didn't do anything. Then you go to bigger models and it starts to get a lift. I think there's a lot of debate about the extent to which this kind of behavior is emergent, but clearly there is a difference. So it's not like there are these two independent paradigms. I think they are related in the sense that you need a certain level of system one capability in your models in order to have system two be able to benefit from system two.

Host

我以前尝试过扮演业余神经科学家,并将其与大脑的进化相比较:你必须先进化大脑皮层,然后才能进化大脑的其他部分。也许我们在这里做的就是这件事。

I have tried to play amateur neuroscientist before and compare it to the evolution of the brain, how you have to evolve the cortex first before you evolve the other parts of the brain. Perhaps that is what we are doing here.

Noam Brown

你可以说实际上这与系统一系统二范式并没有太大不同,因为如果你让一只鸽子非常努力地思考下棋,它不会取得多大进展。即使它思考一千年,它也不会变得更擅长下棋。所以也许你仍然需要一定水平的智力能力,就系统一而言,才能从系统二中受益,就像动物和人类一样。

You could argue that actually this is not that different from the system one system two paradigm because if you ask a pigeon to think really hard about playing chess, it's not going to get that far. It doesn't matter if it thinks for a thousand years. It's not going to be able to be better at playing chess. So maybe you do still need a certain level of intellectual ability just in terms of system one in order to benefit from system two as well, like with animals and humans.

Host

这也适用于视觉推理吗?比如说我们有像 40 这样的原生全模态模型,那么这也让 O3 在 GeoGuessr 上表现得很好。这适用于其他模态吗?

Does this also apply to visual reasoning? So let's say we have the 40 like natively Omni model type of thing, then that also makes O3 really good at GeoGuessr. Does that apply to other modalities too?

Noam Brown

我认为证据表明是的。这取决于你问的具体问题类型。有些问题我认为并不能真正从系统二中受益。我认为 GeoGuessr 肯定是你能受益的一个例子。

I think the evidence is yes. It depends on exactly the kinds of questions that you're asking. There are some questions that I think don't really benefit from system two. I think GeoGuessr is certainly one where you do benefit.

AI 中的系统 1 与系统 2 思维 System 1 vs System 2 Thinking in AI

Host

我认为图像识别是那种你可能不太受益于系统二思维的事情。因为你知道就是知道,不知道就是不知道。

I think image recognition is one of those things that you probably benefit less from system two thinking. Because you either know it or you don't.

Noam Brown

没错。我通常举的例子是信息检索。如果有人问你某人的出生年份,而你无法上网,你知道就知道,不知道就不知道。你可以坐那儿想很久,也许能猜个大概,但除非你已经知道,否则得不到确切日期。但空间推理,比如井字棋,可能更好,因为所有信息都在眼前。

Yeah, exactly. The thing I typically point to is information retrieval. If somebody asks you when a person was born and you don't have web access, you either know it or you don't. You can sit and think for a long time, maybe make an educated guess, but you won't get the exact date unless you already know it. But spatial reasoning, like tic-tac-toe, might be better because you have all the information there.

Host

是的,我认为井字棋确实让 GPT-4.5 栽跟头。它玩得还行——我不该说栽跟头。它表现还算合理,能画棋盘,走合法棋步,但有时会犯错。你需要系统二思维才能让它完美下棋。有可能 GPT-6 仅靠系统一也能完美下棋,但现阶段你需要系统二才能表现好。你觉得系统一需要具备什么?显然是对游戏规则的一般理解。是否还需要一些元游戏理解,比如在不同游戏中如何评估棋子价值,这样系统二才能进入实际对弈?

Yeah, and I think it's true that with tic-tac-toe, GPT-4.5 falls over. It plays decently well—I shouldn't say falls over. It does reasonably well, can draw the board, make legal moves, but it makes mistakes sometimes. You need that system two to enable it to play perfectly. It's possible that GPT-6 with just system one would also play perfectly, but right now you need system two to really do well. What do you think are the things you need in system one? Obviously general understanding of game rules. Do you also need some meta-game understanding, like how to value pieces in different games, so that in system two you can get to gameplay?

Noam Brown

我认为系统一拥有的越多越好。人类也一样。人类第一次下棋时,可以大量运用系统二思维。如果你让一个非常聪明的人玩一个全新的游戏,并告诉他们思考三周,他们能做得相当好。但建立系统一思维——培养对游戏的直觉——肯定有帮助,因为它让你快得多。宝可梦的例子很好:系统一拥有关于游戏的所有信息,但一旦进入游戏,仍然需要很多辅助工具才能工作。我在想,我们能从辅助工具中提取多少放到系统一里,这样系统二就能尽可能摆脱辅助工具。但这正是关于游戏泛化和 AI 的问题。

I think the more you have in system one, the better. It's the same with humans. When humans play a game like chess for the first time, they can apply a lot of system two thinking. If you present a really smart person with a completely novel game and tell them to think about it for three weeks, they could do pretty well. But it certainly helps to build up that system one thinking—build up intuition about the game—because it makes you so much faster. The Pokémon example is a good one: system one has all this information about games, but once you put in the game, it still needs a lot of harnesses to work. I'm trying to figure out how much we can take from the harness and have it in system one, so that system two is as harness-free as possible. But that's the question about generalizing games and AI.

Host

是的,我认为这是另一个问题。我认为理想的辅助工具就是没有辅助工具。辅助工具是拐杖,我们最终会超越它。当宝可梦作为基准测试出现时,我非常反对用它来评估我们的开源模型。我的感觉是:如果我们要做这个评估,那就直接用 O3 来做。没有辅助工具,O3 能走多远?答案是走不远,这没关系。我认为答案不应该是构建一个很好的辅助工具,让它在评估中表现好。答案应该是提升我们模型的能力,让它们在所有事情上都表现好,然后它们自然会在评估上取得进展。

Yeah, I view that as a different question. I think the ideal harness is no harness. Harnesses are a crutch that we'll eventually move beyond. When the Pokémon thing emerged as a benchmark, I was pretty opposed to evaluating our open-source models with it. My feeling is: if we're going to do this eval, let's just do it with O3. How far does O3 get without any harness? The answer is not very far, and that's fine. I don't think the answer should be to build a really good harness so it does well on this eval. The answer is to improve the capabilities of our models so they do well at everything, and then they happen to make progress on this eval.

Host

你会把检查合法移动这类事情视为辅助工具,还是属于模型本身?在下棋时,你可以让模型在系统一中学习哪些移动是合法的,或者用系统二来找出哪里出错了。所以有很多设计问题。对我来说,你应该让模型有能力检查移动是否合法,如果你愿意的话。这可以是环境中的一个选项——一个工具调用,用来查看动作是否合法。如果它想用,就可以用。然后有一个设计问题:如果模型做出了非法移动,你该怎么办?我认为合理的做法是,如果它做出非法移动,它就输掉比赛。人类在下棋时做出非法移动会怎样?我其实不知道。

Would you consider things like checking for a valid move a harness, or is that in the model? In chess, you can either have the model learn in system one what moves are valid, or use system two to figure out where it went wrong. So there are a lot of design questions. For me, you should give the model the ability to check if a move is legal if you want. That could be an option in the environment—a tool call to see if an action is legal. If it wants to use it, it can. Then there's a design question: what do you do if the model makes an illegal move? I think it's reasonable to say if they make an illegal move, they lose the game. What happens when a human makes an illegal move in chess? I don't actually know.

Noam Brown

我认为这是不允许的。就是不允许。你会直接输掉比赛吗?我不知道。所以如果是这样,我认为在评估中设定同样的标准对 AI 模型来说是完全合理的。

I don't think it's allowed. You're just not allowed to. Do you just lose the game? I don't know. So if that's the case, I think it's totally reasonable to have an eval where that's also the criteria for AI models.

Host

是的,但也许在研究术语中,一种解读方式是:允许你做搜索吗?DeepSeeker 的一个著名发现是,MCTS 对他们来说并不那么有用。但有很多工程师在尝试搜索,花费大量 token 做这件事,也许这不值得。

Yeah, but maybe one way to interpret that in research terms is: are you allowed to do search? One famous finding from DeepSeeker is that MCTS wasn't that useful to them. But there are a lot of engineers trying out search and spending a lot of tokens doing that, and maybe it's not worth it.

Noam Brown

我在这里区分一下:一种是工具调用,用来检查移动是否合法;另一种是实际做出移动,然后看它是否合法。如果工具调用可用,那么进行工具调用并检查是完全没问题的。但模型说‘哦,我要走这一步’,然后收到反馈说这是非法移动,然后说‘开玩笑的,我换一步’,这是不同的。这就是我划出的区别。有些人试图将第二种类型归类为测试时计算。我不会将其归类为测试时计算。有很多理由说明,当你进入现实世界时,你不希望依赖这种范式。想象你有一个机器人,它在世界上采取行动并打破了东西。你不能说‘开玩笑的,我不是故意的’。东西已经坏了。

I'm making a distinction here between a tool call to check whether a move is legal or illegal, and actually making that move and then seeing whether it ended up being legal. If that tool call is available, it's totally fine to make that tool call and check. It's different to have the model say, 'Oh, I'm making this move,' and then get feedback that it was illegal, and then say, 'Just kidding, I'll do something else.' That's the distinction I'm drawing. Some people have tried to classify that second type as test-time compute. I would not classify that as test-time compute. There are a lot of reasons why you would not want to rely on that paradigm when you go to the real world. Imagine you have a robot, and it takes some action in the world and breaks something. You can't say, 'Just kidding, I didn't mean to do that.' The thing is broken.

模型路由器与扩展 Model Routers and Scaling

Host

所以如果你想模拟如果我这样移动机器人会发生什么,然后在模拟中你看到这个东西坏了,然后你决定不采取那个行动。那完全没问题,但你无法撤销已经在世界中采取的行动。在这个大致领域里,我还有几件事想谈。我实际上有一个关于快思考与慢思考的答案,也许我好奇你怎么看。很多人试图在快速响应模型和长思考模型之间加入模型路由器层。Anthropic 明确在做这件事,我认为有一个问题:你总是需要一个聪明的裁判来路由,还是需要一个笨的裁判来路由因为它快?所以当你有一个模型路由器,比如说你在系统一侧和系统二侧之间传递请求,路由器需要像聪明模型一样聪明,还是笨一点以保持快速?

So if you want to simulate what would happen if I move the robot in this way and then in the simulation you saw that this thing broke and then you decide not to do that action. That's totally fine, but you can't just undo actions that you've taken in the world. There's a couple more things I wanted to cover in this rough area. I actually had an answer on the thinking fast and slow side which maybe I'm curious what you think about. Like a lot of people are trying to put in effectively model router layers, let's say between the fast response model and the long thinking model. Anthropic is explicitly doing that and I think there is a question about always do you need a smart judge to route or do you need a dumb judge to route because it's fast? So when you have a model router, let's say you're passing requests between system one side and system two side, does the router need to be as smart as the smart model or dumb to be fast?

Noam Brown

我认为一个笨模型有可能识别出问题非常难,它无法解决,然后将其路由到更强大的模型。但笨模型也可能被欺骗或过于自信。我不知道。我认为这里确实存在权衡。但我要说的是,我认为人们现在正在构建的很多东西最终都会被 Scaling 所淘汰。我认为 harness 是一个很好的例子,最终模型会变得更好,而且我认为这实际上已经发生在推理模型上。在推理模型出现之前,有大量工作投入到工程化这些智能体系统中,这些系统大量调用 GPT-4o 或这些非推理模型来获得推理行为。然后结果发现,哦,我们只是创建了推理模型,你不需要这种复杂的行为。事实上,在很多方面它反而更糟。你只需给推理模型同样的问题,没有任何脚手架,它就能直接做。现在人们仍然可以在推理模型之上构建脚手架,但我认为在很多方面,这些脚手架也会被推理模型和更强大的模型所取代。类似地,我认为像模型路由器这样的东西,你知道,我们已经相当公开地表示我们希望走向一个拥有单一统一模型的世界。在那个世界里,你不应该在模型之上需要路由器。所以我认为路由器问题最终也会得到解决。就像你把路由器构建到模型权重本身中一样。我不认为会有什么好处,我不应该这么说,因为我可能错了。当然,也许有理由路由到不同的模型提供商等等,但我认为路由器最终会消失。我能理解为什么短期内值得做,因为事实上它现在是有益的,如果你正在构建一个产品并且从中获得提升,那么现在做是值得的。

I think it's possible for a dumb model to recognize that a problem is really hard and that it won't be able to solve it and then route it to a more capable model. But it's also possible for a dumb model to be fooled or to be overconfident. I don't know. I think there's a real trade-off there. But I will say like I think there are a lot of things that people are building right now that will eventually be washed away by scale. So I think harnesses are a good example where I think eventually the models are going to be and I think this actually happened with the reasoning models. Like before the reasoning models emerged, there was all of this work that went into engineering these agentic systems that made a lot of calls to GPT-4o or these non-reasoning models to get reasoning behavior. And then it turns out like oh, we just created reasoning models and you don't need this complex behavior. In fact in many ways it makes it worse. Like you just give the reasoning model the same question without any sort of scaffolding and it just does it. Now you can still and so people are building scaffolding on top of the reasoning models right now, but I think in many ways those scaffolds will also just be replaced by the reasoning models and models in general becoming more capable. And similarly I think things like model routers, you know, we've said pretty openly that we want to move to a world where there is a single unifying model. And in that world you shouldn't need a router on top of the model. So I think that the router issue will eventually be solved also. Like you're building the router into the model's weights itself. I don't think there will be a benefit for like I don't I shouldn't say cuz it's I could be wrong about this. Like you know, it's certainly maybe there's reasons to route to different model providers or whatever, but I think that routers are going to eventually go away. And I can understand why it's worth doing it in the short term because the fact is it is beneficial right now and if you're building a product and you're getting a lift from it, then it's worth doing right now.

给开发者的快速演进建议 Advice for Developers on Rapid Evolution

Host

我想象很多开发者面临的一个棘手问题是,你必须规划这些模型在 6 个月和 12 个月后会是什么样子,这非常困难,因为事情进展得非常快。你不想花 6 个月构建一个东西,然后它完全被 Scaling 淘汰。但我想鼓励开发者,当你们构建像脚手架和路由器这类东西时,要记住这个领域正在飞速发展。事情会在 3 个月内发生变化,更不用说 6 个月了,这可能需要彻底改变这些东西或完全抛弃它们。所以不要花 6 个月构建一个可能在 6 个月后被抛弃的东西。但这太难了。每个人都这么说,但没有人有具体的建议。

One of the tricky things that I'd imagine that a lot of developers are facing is that you kind of have to plan for where these models are going to be in 6 months and 12 months and it's like very hard to do because things are progressing very quickly. You know, you don't want to spend 6 months building something and then just have it be totally washed away by scale. But I think I would encourage developers like when they're building these kinds of things like scaffolds and routers, keep in mind that the field is evolving very rapidly. You know, things are going to change in 3 months let alone 6 months and that might require radically changing these things around or tossing them out completely. So don't spend 6 months building something that might get tossed out in 6 months. It's so hard though. Everyone says this and then like no one has concrete suggestions on how.

Noam Brown

那强化微调呢?这显然是你们一个月前在 OpenAI 发布的。人们现在应该花时间在这上面,还是等到下一次跳跃?我认为强化微调很酷,值得研究,因为它实际上是针对你拥有的数据对模型进行专门化,我认为开发者值得研究。很多时候我们不会突然将那些数据直接融入原始模型。所以我认为这有点像是一个单独的问题。

What about reinforcement fine-tuning? Is this something that obviously you just released it a month ago at OpenAI. Is this something people should spend time on right now or maybe wait until the next jump of this? I think reinforcement fine-tuning is pretty cool and I think it's worth looking into because it's really about specializing the models for the data that you have and I think that something that's worth looking into for developers. Like we're not suddenly going to have that data baked into the raw model a lot of times. So I think that's kind of like a separate question.

Host

是的。所以创建环境和奖励模型是现在人们能做的最好的事情。我认为人们的问题是,我应该急于使用 RFT 微调模型,还是应该构建 harness 以便在模型变得更好时进行 RFT?我认为区别在于,对于强化微调,你收集的数据随着模型的改进也会变得有用。所以如果我们推出未来更强大的模型,你仍然可以在你的数据上微调它们。我认为这实际上是一个很好的例子,你构建的东西将补充模型的 Scaling 并变得更强大,而不是被 Scaling 淘汰。

Yeah. So creating the environment and the reward model is the best thing people can do right now. I think the question that people have is like should I rush to fine-tune the model using RFT or should I build the harness to then RFT the models as they get better? I think the difference is that for reinforcement fine-tuning, you're collecting data that's going to be useful as the models improve as well. So if we come out with future models that are even more capable, you could still fine-tune them on your data. That's I think actually a good example where you're building something that's going to complement the model scaling and becoming more capable rather than necessarily getting washed away by the scale.

伊利亚的尝试与推理时机 Ilya's Attempt and Timing of Reasoning

Host

最后一个关于 Ilya 的问题,你在 Sarah 和 Ilya 的播客中提到,几年前你和 Ilya 有过一次对话,关于在语言模型中增加强化学习和推理。只是猜测或想法,为什么他的尝试没有成功,或者时机不对,以及为什么现在时机成熟了。

One last question on Ilya, you mentioned on the Sarah and Ilya podcast where you had this conversation with Ilya a few years ago about more RL and reasoning in language models. Just any speculation or thoughts on why his attempt when he tried it it didn't work or the timing wasn't right and why the time is right now.

Noam Brown

我不认为他的尝试没有成功。在很多方面它确实成功了。对我来说,Ilya 看到在我研究过的所有领域——扑克、Hanabi 和外交——让模型在行动前思考对性能产生了巨大影响。数量级的差异。比如 10,000 倍,不是吗?是的,1,000 到 100,000 倍相当于一个模型大了 1,000 到 100,000 倍。而在语言模型中,你并没有真正看到这一点。模型只是立即响应。LM 领域的一些人坚信,只要我们继续 Scaling 预训练,我们就会达到超级智能。我对此持怀疑态度。2021 年底,我和 Ilya 一起吃饭。他问我的 AGI 时间线是什么,一个非常标准的旧金山问题。我告诉他,我认为实际上还很遥远,因为我们需要以一种非常通用的方式解决这个推理范式。而对于像 LM 这样的东西,LM 非常通用,但它们没有一个非常通用的推理范式。

I don't think I would frame it that way that his attempt didn't work. In many ways it did. So Ilya for me, I saw that in all of these domains that I'd worked on on poker and Hanabi and diplomacy, having the models think before acting made a huge difference in performance. Like orders of magnitude difference. Like 10,000 times isn't it? Yeah, like you know, 1,000 to 100,000 times is the equivalent of a model that's like 1,000 to 100,000 times bigger. And in language models, you weren't really seeing that. The models would just respond instantly. Some people in the field in the LM field were like convinced that okay, we just keep scaling pre-training, we're going to get to superintelligence. And I was kind of skeptical of that perspective. In late 2021, I was having a meal with Ilya. He asked me what my AGI timelines are, a very standard SF question. And I told him like look, I think it's actually quite far away because we're going to need to figure out this reasoning paradigm in a very general way. And with things like LMs, LMs are very general but they don't have a reasoning paradigm that's very general.

预训练扩展的极限 Limits of Pre-training Scaling

Host

在他们做到这一点之前,他们的能力是有限的。当然,我们会继续扩大规模,再提升几个数量级,模型会变得更强大,但仅靠这些是无法实现超级智能的。是的,如果我们有千万亿美金来训练这些模型,也许可以,但在达到超级智能之前,你会先碰到经济可行性的天花板,除非你有一个推理范式。我过去错误地认为推理范式需要很长时间才能搞清楚,因为这是一个巨大的未解研究问题。而且,伊利亚也同意我的看法,他说是的,我们需要这个额外的范式。但他的观点是,也许这并不那么难。我当时不知道,但他和 OpenAI 的其他同事也一直在思考这个问题。他们也在考虑强化学习,一直在研究,我认为他们取得了一些成功。但就像大多数研究一样,你需要反复迭代,尝试不同的想法,尝试不同的东西。而且随着模型变得更强大、更快,实验迭代也变得更简单。我认为他们所做的工作,即使没有直接产生推理范式,也都是建立在之前工作的基础上。所以他们积累了很多东西,最终导致了推理范式的出现。

And until they do, they're going to be limited in what they can do. You know, we're going to scale up sure. We're going to scale these things up by a few more orders of magnitude, they're going to become more capable, but we're not going to see superintelligence from just that. And like yes, if we had a quadrillion dollars to train these models, then maybe we would, but like you're going to hit the limits of what's economically feasible before you get to superintelligence unless you have a reasoning paradigm. And I was convinced incorrectly that the reasoning paradigm would take a long time to figure out because it's like this big unanswered research question. And you know, Ilya agreed with me and he said like yeah, you know, we need this like additional paradigm. But his take was that like maybe it's not that hard. I didn't know it at the time, but like he and others at OpenAI had also been thinking about this. They'd also been thinking about RL. They'd been working on it and I think they had some success, but like you know, with most research like it does you have to iterate on things. You have to try out different ideas. You have to yeah, try different things and then also as the models become more capable, as they become faster, it becomes easier to iterate on experiments. And I think that the work that they did even though it didn't like result in a reasoning paradigm, it all builds on top of previous work, right? So they built a lot of things that over over time led to this reasoning paradigm.

推理范式的兴起 The Rise of Reasoning Paradigm

Host

对于听众来说,诺姆可以谈谈这个,但传言说那个项目代号是 GPT-0,如果你想搜索那方面的工作的话。我认为有一段时间,强化学习基本上经历了一个黑暗时代,每个人都全力投入,但没什么结果,然后他们放弃了。而现在,它似乎又迎来了黄金时代。所以我试图弄清楚为什么,是什么原因?可能只是因为我们有了更聪明的基础模型和更好的数据。我不认为仅仅是基础模型更聪明了。我认为,是的,我们最终在推理上取得了巨大成功。但我认为这在很多方面是一个渐进的过程。在某种程度上是渐进的。你知道,有一些生命迹象,然后我们迭代,尝试更多的东西,得到了更好的生命迹象。我想大概是在 2023 年 10 月或 11 月,我确信我们有了非常确凿的生命迹象,哦,这将是那个范式,而且会是一件大事。这在很多方面并不是一个渐进的过程。我认为 OpenAI 做得好的地方是,当我们得到这些生命迹象时,他们认识到了它的价值,并大力投资扩大规模。我认为这最终导致了推理模型在那个时候出现。

For listeners, Noam can talk about this, but the rumor is that that thing was code-named GPT-0 if you want to search for that that line of work. I think there was a time where it like basically RL kind of went through a dark age when everyone like went all in on it and then nothing happens and they gave up. And like now it's like sort of the golden age again. So that's what I'm like trying to identify like why what is it? And it could just be that we have smarter base models and better data. I don't think it's just that we have smarter base models. I I think it's that yeah, so I we did end up getting a big success with with reasoning. And I but I think it was in many ways a gradual thing. It to some extent it was gradual. You know, like there were signs that there were signs of life and then we like, you know, iterated and tried out some more things. We got like better signs of life. I think it was around like No- November 2023 or October 2023 when I think I was convinced that we had like very conclusive signs of life that like oh this this is going to be a this is the paradigm and it's going to be a big deal. That wasn't many ways a a gradual thing. I think what OpenAI did well is like when we got those signs of life, they recognized it for what it was and invested heavily in in scaling it up. And I think that's that's ultimately what what led to reasoning models arriving when they did.

推理的内部辩论 Internal Debate on Reasoning

Host

内部是否有分歧,尤其是因为你知道,OpenAI 开创了预训练缩放,有点像“算力就是一切”,然后你又说也许那不是我们实现目标的方式。每个人都清楚这行得通吗,还是说有争议?

Was there any disagreement internally especially because like you know, OpenAI kind of pioneered pre-training scaling, you know, and kind of like computers all you need and then you're kind of saying maybe that's not how we get there. Was it clear to everybody that like okay, this is going to work or was it controversial?

Noam Brown

关于这些事情总是有不同的意见。我认为有些人觉得预训练就是我们所需要的,把它无限扩大就行了。实际上,OpenAI 的很多领导层都认识到需要另一个范式,这就是为什么他们投入大量研究精力在强化学习上。我认为这也是 OpenAI 的功劳:是的,他们搞清楚了预训练范式,并且非常专注于扩大规模。事实上,绝大部分资源都用于扩大规模,但他们也认识到还需要别的东西,值得投入研究人员精力去探索其他方向,找出那个额外的范式是什么。首先,关于那个额外的范式是什么,有很多争论。我认为很多研究者看待推理,而强化学习实际上并不是关于扩展测试时算力,更多的是关于数据效率。因为感觉是我们有大量的算力,但实际上更受数据限制。所以存在一个数据墙,我们会在碰到算力限制之前先碰到它。那么如何让这些算法更数据高效呢?它们确实更数据高效,但我认为它们也相当于大幅扩展了算力。这很有趣。围绕我们到底在做什么有很多争论。然后我认为即使在我们得到生命迹象之后,关于其重要性也有很多争论。比如,我们应该投入多少资源来扩大这个范式?特别是当你在一家小公司时,比如 2023 年的 OpenAI 还没有今天这么大,算力也比现在更紧张。如果你把资源投入一个方向,就会牺牲其他方向。所以当你看到推理的这些生命迹象,说这个看起来很有希望,我们要大幅扩大它,投入更多资源,那么这些资源从哪里来?你必须做出艰难的决定,从哪里抽调资源,这是一个非常有争议、非常困难的决定,会让一些人不高兴。我认为有争论我们是否过于关注这个范式,它是否真的那么重要,我们是否会看到它泛化并做各种事情。我记得有趣的是,我和一个在发现推理范式后、但在发布 01 之前离开 OpenAI 的人聊过,他后来去了一个竞争实验室。我们在发布 01 之后又见面了,他告诉我,当时他真的不认为这个推理的东西,这些 O 系列、草莓模型有那么重要。他觉得我们把它夸大了。然后当我们发布 01 时,他看到那个竞争实验室的同事们的反应,每个人都觉得“哦,糟了,这太重要了”,然后他们整个研究议程都转向了这个,他才意识到,哦,这也许真的很重要。你知道,很多事事后看来很明显,但当时并不那么明显,很难认识到它的本质。我是说,OpenAI 有做出正确押注的辉煌历史。

There's always different opinions about this stuff. I think there were some people that felt that pre-training was all we needed and we scaled it up to infinity and we're there. I think a lot of the leadership actually at OpenAI recognized that there was another paradigm that was needed and that was why they were investing all of this like research effort into this like RL um stuff. And I think that's also to the credit of OpenAI that like okay, yes, they figured out the pre-training paradigm and they were very focused on scaling it up. In fact, the vast majority of resources were focused on scaling it up, but they also recognized the value that that something else was going to be needed and it was worth researching putting researcher effort into into other directions to figure out what that extra paradigm was going to be. There was a lot of debate about first of all like what is that extra paradigm? So I think a lot of the researchers looked at reasoning and and RL was not really about scaling test time compute. It was more about data efficiency. Because you know, the feeling was that we have tons and tons of compute, but we actually are more limited by data. So there's there's a data wall and we're going to hit that before we hit limits on on the compute. So how do we make these algorithms more data efficient? They are more data efficient, but I think that also like um they are also just like the equivalent of scaling up compute also by a ton. That was interesting. There was like a lot of debate around like okay, what what exactly are we doing here? And then I think also even when we got the signs of life, I think there was a lot of debate about the significance of it. Like okay, how much should we invest in scaling up this paradigm? I think especially when you're when you're in a small company like you know, OpenAI like in 2023 was not as big as it is today and compute was more constrained than it is today. And if you're investing resources in in a direction that's coming at the expense of something else. And so if you look at these signs of life on reasoning and you're saying like okay, well this looks promising. We're going to scale this up by a ton and invest a lot more resources into it, where are those resources coming from? You have to make that tough call about where to where to draw the resources from and that is a very controversial, very difficult call to make um that makes some people unhappy. And I think there was debate about whether we're focusing too much on this paradigm, whether it's really a big deal, whether we would see it generalize and do various things. And I remember it was interesting that I I talked to somebody who left OpenAI after we had discovered the reasoning paradigm, but before we announced 01 and they ended up going to a competing lab. I saw them afterwards after we announced um 01 and they told me that like at the time they really didn't think this like reasoning thing like this these O series, the strawberry models were like that that big of a deal. It was like they thought we're making a bigger deal of it than it really deserved to be. And then when we announced 01 and they saw the reaction of their co-workers at this competing lab about how everybody was like oh crap like this is a big deal and they like pivoted the whole research agenda Oh my god. to focus on this that then they realized like oh actually like this maybe is a big deal. You know, a lot of this seems obvious in retrospect, but at the time it's actually not so obvious and be quite difficult to recognize something for what it is. I mean, OpenAI has like a great history of just making the right bet.

扩展范式与数据效率 Scaling paradigm and data efficiency

Host

我觉得 GPT 模型也挺类似的,对吧?从游戏和强化学习开始,然后可能我们直接规模化这些语言模型就行了。我对领导层和研究团队不断提出这些见解印象深刻。现在回头看,这些模型随着规模扩大而变好似乎显而易见,所以你就应该大力规模化。但最好的研究往往是事后看来显而易见,当时并不像今天看起来那么明显。接着问数据效率的问题。这是我的一个心头好。我们当前的学习方法似乎仍然非常低效,对吧?相比人类,我们看五个样本就能学会。机器可能每个数据点需要 200 个。有人在数据效率方面做有趣的工作吗?还是说机器学习相比人类存在根本性的低效,永远无法克服?

I feel like GPT models are kind of similar, right? Where it started with games and RL and then it's like maybe we can just scale these language models instead. I'm just impressed by the leadership and the research team that keeps coming up with these insights. Looking back today, it might seem obvious that these models get better with scale, so you should just scale them up a ton. But the best research is obvious in retrospect; at the time it's not as obvious as it might seem today. Follow-up questions on data efficiency. This is a pet topic of mine. It seems our current methods of learning are still so inefficient, right? Compared to humans, we take five samples and learn something. Machines need maybe 200 per data point. Is anyone doing anything interesting in data efficiency, or do you think there's a fundamental inefficiency that machine learning will always have compared to humans?

Noam Brown

我觉得这点说得好。如果你看看这些模型训练所用的数据量,再对比人类达到同样表现所需的数据,预训练很难直接比较,因为我不知道婴儿发育过程中吸收了多少 token。但公平地说,这些模型的数据效率低于人类,这是一个未解决的研究问题,可能是最重要的之一。也许比算法改进更重要,因为我们可以从现有世界和人类中增加数据供应。有几点想法:一是答案可能是算法改进——也许算法改进确实能提高数据效率。二是人类并不只是通过阅读互联网来学习。从互联网数据学习当然最容易,但我不认为那是你能收集的数据的极限。

I think it's a good point. If you look at the amount of data these models are trained on and compare it to the data a human observes to get the same performance, pre-training is a little hard to compare apples to apples because I don't know how many tokens a baby absorbs when developing. But it's fair to say these models are less data efficient than humans, and that's an unsolved research question, probably one of the most important. Maybe more important than algorithmic improvements, because we can increase the supply of data from the existing set of the world and humans. A couple thoughts: one is that the answer might be algorithmic improvement—maybe algorithmic improvements do lead to greater data efficiency. Second, humans don't just learn from reading the internet. It's certainly easiest to learn from data on the internet, but I don't think that's the limit of what data you could collect.

伊利亚·苏茨克维的洞见 Insights from Ilya Sutskever

Host

在转到编程话题之前,最后一个问题。关于 Ilya 还有其他轶事或见解吗?你和他共事过,能和他共事的人不多。我对他的远见印象深刻。当我加入 OpenAI,看到内部文档中他在 2021、2022 年甚至更早的想法时,我很佩服他对未来走向和所需条件有清晰的愿景。他 2016-17 年创立 OpenAI 时的一些邮件被公开了,那时他就在说一个大型实验比一百个小型实验更有价值。这是一个核心见解,让他们与 Brain 等区分开来。他看事情比别人清楚得多。我想知道他的生产函数是什么——怎么才能造就这样的人,以及如何改进自己的思维来更好地模仿他。

Last follow-up before we change topics to coding. Any other anecdotes or insights from Ilya in general? You've worked with him, and there aren't many people we can talk to who have worked with him. I think I've been very impressed with his vision. When I joined OpenAI and saw internal documents of what he had been thinking about back in 2021, 2022, even earlier, I was very impressed that he had a clear vision of where this was all going and what was needed. Some of his emails from 2016-17 when they were founding OpenAI were published, and even then he was talking about how one big experiment is much more valuable than 100 small ones. That was a core insight that differentiated them from Brain, for example. He just sees things much more clearly than others. I wonder what his production function is—how do you make a human like that, and how do you improve your own thinking to better model it?

Noam Brown

我认为 OpenAI 的一大成功确实在于押注 Scaling 范式。这有点奇怪,因为他们不是最大的实验室,规模化对他们来说很困难。那时更常见的是做大量小型实验,偏学术风格。人们试图找出各种算法改进,而 OpenAI 很早就押注大规模。我们请过 David Luan,他是 GPT-1 和 2 时期的工程副总裁,他谈到 Brain 和 OpenAI 的区别基本上就是 Google 无法推出规模化模型的原因。结构上,每个人都被分配了算力,你必须集中资源才能下注,但就是做不到。我认为确实如此:OpenAI 的结构不同,这真的帮了他们。OpenAI 运作起来很像初创公司,而其他地方更像大学或传统研究实验室。OpenAI 以初创公司的方式运作,以构建 AGI 和超级智能为使命,这帮助他们组织、协作、集中资源,并做出关于资源分配的艰难抉择。我认为很多其他实验室现在也在尝试采用类似的模式。

I think it is true that one of OpenAI's big successes was betting on the scaling paradigm. It's kind of odd because they were not the biggest lab; it was difficult for them to scale. Back then, it was much more common to do a lot of small experiments in a more academic style. People were trying to figure out various algorithmic improvements, and OpenAI bet pretty early on large scale. We had David Luan on, who was VP of engineering at the time of GPT-1 and 2, and he talked about how the difference between Brain and OpenAI was basically the cause of Google's inability to come up with a scaled model. Structurally, everyone had allocated compute, and you had to pool resources together to make bets, and you just couldn't. I think that's true: OpenAI was structured differently, and that really helped them. OpenAI functions a lot like a startup, while other places tended to function more like universities or traditional research labs. The way OpenAI operates more like a startup with the mission of building AGI and superintelligence helped them organize, collaborate, pool resources, and make hard choices about how to allocate resources. I think a lot of other labs have now been trying to adopt similar paradigms.

AI 编程:Codex 与 WinSurf Coding with AI: Codex and WinSurf

Host

我们来谈谈杀手级应用,至少在我看来是编程。你最近发布了 Codex,但我很想聊聊 Noam Brown 的编程栈。你用哪些模型,怎么和它们交互?Cursor、WinSurf?

Let's talk about the killer use case, at least in my mind, which is coding. You released Codex recently, but I'd love to talk through the Noam Brown coding stack. What models do you use, how do you interact with them? Cursor, WinSurf?

Noam Brown

最近我一直在用 WinSurf 和 Codex,实际上 Codex 用得很多。我觉得很有趣:你给它一个任务,它就自己去做了,五分钟后回来给你一个拉取请求。

Lately I've been using WinSurf and Codex, actually a lot of Codex. I've been having a lot of fun: you just give it a task and it goes off and does it, and comes back 5 minutes later with a pull request.

Host

是核心研究任务,还是你不那么在乎的边角料?

Is it core research tasks or side stuff you don't super care about?

Noam Brown

我不觉得是边角料。基本上我平时想写的任何代码,我都会先用 Codex 试试。对你来说是免费的,对所有人现在都是免费的。部分原因是这是我最有效的方式,同时也让我积累使用这项技术的经验,看到它的不足。这有助于我更好地理解这些模型的极限,以及下一步需要推动什么。

I wouldn't say it's side stuff. Basically anything I would normally try to code up, I try to do it with Codex first. For you it's free, but yeah for everybody it's free right now. I think it's partly because it's the most effective way for me to do it, and also it's good for me to get experience working with this technology and seeing its shortcomings. It helps me better understand the limits of these models and what we need to push on next.

Host

你感受到 AGI 了吗?

Have you felt the AGI?

Noam Brown

我多次感受到 AGI,是的。

I felt the AGI multiple times, yes.

Host

人们应该怎么像你一样推动 Codex 的边界?你比别人更早看到它,因为你离它更近。

How should people push Codex in ways that you've done? You see it before others because you're closer to it.

Noam Brown

我认为任何人都可以用 Codex 感受到 AGI。有趣的是,你感受到 AGI,然后很快就习惯了。你会对它的不足之处感到满足。我知道,有一天它很神奇。我回头看 Sora 刚发布时的旧视频。还记得 Sora 刚出来的时候吗?那真是太神奇了。你看着它,心想‘它真的来了,这就是 AGI’。但现在再看,人物动作不够自然,缺乏一致性,你会看到所有当初没注意到的缺陷。

I think anybody can use Codex and feel the AGI. It's funny how you feel the AGI and then you get used to it very quickly. You become satisfied with where it's lacking. I know, it's magical one day. I was looking back at the old Sora videos when they were announced. Remember when Sora came out? It was just magical. You look at it and think, 'It's really here, this is AGI.' But if you look at it now, people don't move very organically, there's a lack of consistency, and you see all these flaws that you didn't notice when it first came out.

AGI 时刻与推理模型 AGI moments and reasoning models

Host

是的,你会很快习惯这项技术。但我觉得酷的是,因为它发展得这么快,你每隔几个月就会体验到那种 AGI 时刻。又有新东西出来,让你觉得神奇,然后你又很快习惯了。既然你已经深入使用 WinSurf,有什么专业建议吗?

And yeah, you get used to this technology very quickly. But I think what's cool about it is that because it's developing so quickly, you get those AGI moments every few months. Something else is going to come out and it's magical to you, and then you get used to it very quickly. What are your WinSurf pro tips now that you've immersed in it?

Noam Brown

有一件事让我很惊讶,就是很少有人——我是说,也许你的听众会更习惯用推理模型,用得也更多——但让我惊讶的是,很多人甚至不知道 O3 的存在。我每天都在用。它基本上取代了谷歌搜索。我一直在用。对于编程之类的事情,我也倾向于用推理模型。我的建议是,如果人们还没试过推理模型,因为说实话,用过的人都喜欢它们。显然,更多人用 GPT-4o 和 ChatGPT 的默认设置之类的。我觉得值得试试推理模型。我想人们会对它们的能力感到惊讶。我每天都用 WinSurf,但他们还没有把它设为默认选项。我总得去翻出来输入 O3,然后才意识到,哦,原来有这个。这很奇怪。

I think one thing I'm surprised by is how few people—I mean maybe your audience is going to be more comfortable with reasoning models and use them more—but I'm surprised at how many people don't even know that O3 exists. I've been using it day-to-day. It's basically replaced Google search for me. I just use it all the time. Also for things like coding, I tend to just use the reasoning models. My suggestion is if people haven't tried the reasoning models yet, because honestly, people who use them love them. Obviously a lot more people use GPT-4o and the default on ChatGPT and that kind of stuff. I think it's worth trying the reasoning models. I think people would be surprised at what they can do. I use WinSurf daily and they still haven't actually enabled it as a default in WinSurf. I always have to dig up and type in O3, and then it's like, oh yeah, that exists. It's weird.

Host

我觉得用它的一个困扰是,它推理时间有点长,会打断工作流。确实如此。我认为这是 Codex 的优势之一:你可以给它一个相对独立的任务,它自己去做,十分钟后回来。我理解,如果你把它当结对编程用,那确实要用 GPT-4.1 之类的。

I would say my struggle with it has been that it takes a little long to reason and actually breaks out of flow. I think that is true, yes. And I think this is one of the advantages of Codex: you can give it a task that's kind of self-contained, and it can go off and do its thing and come back 10 minutes later. I could see that if you're using this thing more like a pair programmer, then yeah, you want to use GPT-4.1 or something like that.

开发周期的断裂环节 Broken parts of development cycle

Host

你觉得在 AI 开发周期中,最薄弱的部分是什么?在我看来,是拉取请求审查。我一直用 Codex,然后收到很多拉取请求,很难全部审阅。你希望别人构建什么来让这个流程更可扩展?

What do you think are the most broken parts of the development cycle with AI? In my mind, it's pull request review. For me, I use Codex all the time, and then I get all these pull requests, and it's kind of hard to go through all of them. What other thing would you like people to build to make this even more scalable?

Noam Brown

我认为这真的要靠我们来构建更多东西。这些模型在某些方面非常有限。我觉得沮丧的是,你让它们做一件事,它们花十分钟做完,然后你让它们做一件很类似的事,它们又花十分钟去做。我把它们描述为天才,但这是它们上班的第一天。这有点烦人。即使地球上最聪明的人,上班第一天也不会像你期望的那样有用。所以,让它们获得更多经验,表现得像已经工作了六个月而不是一天的人,会让它们有用得多。但这确实要靠我们来构建这种能力。

I think it's really on us to build a lot more stuff. These models are very limited in some ways. I find it frustrating that you ask them to do something, and they spend 10 minutes doing it, and then you ask them to do something pretty similar, and they go spend 10 minutes doing it again. I describe them as geniuses, but it's their first day on the job. That's kind of annoying. Even the smartest person on earth, when it's their first day on the job, they're not going to be as useful as you would like. So being able to get more experience and act like somebody that's actually been on the job for six months instead of one day would make them a lot more useful. But that's really on us to build that capability.

Host

你觉得这在很大程度上是 GPU 受限吗?以 Codex 为例,为什么它让我自己设置环境?如果我问 O3 为一个仓库创建环境设置脚本,我相信它能做到,但今天在产品中我必须自己做。所以我想知道,在你看来,如果我们只是投入更多的测试时算力,这些能好很多吗?还是你认为今天存在根本性的模型兼容性限制,我们仍然需要大量人工辅助?

Do you think a lot of it is GPU-constrained for you? If I think about Codex, why is it asking me to set up the environment myself when the model—if I ask O3 to create an environment setup script for a repo, I'm sure it'll be able to do it, but today in the product I have to do it. So I'm wondering, in your mind, could these be a lot more if we just put more test-time compute on them, or do you think there's a fundamental model compatibility limitation today that we still need a lot of human harnesses around it?

Noam Brown

我认为我们现在处于一个尴尬的状态,进步非常快,有些事情显然我们可以做到,模型也会更好。我们会做到的。你只是受限于一天有多少小时。进步只能这么快。我们正尽可能快地推进一切,我认为 O3 不会是六个月后技术的状态。我总体上喜欢这个问题。有一个软件开发生命周期,不仅仅是代码生成,从 issue 到 PR,还有 WinSurf 这边,在你的 IDE 内部。拉取请求审查是人们不太——有初创公司围绕它建立。这不是 Codex 做的事,但它可以。所以还有什么东西在限制你迭代软件的数量?这是一个开放问题。我不知道有没有答案。

I think that we're in an awkward state right now where progress is very fast and there are things that are clearly we could do this and the models would be better. We're going to get to it. You're just limited by how many hours there are in the day. Progress can only proceed so quickly. We're trying to get to everything as fast as we can, and I think that O3 is not where the technology will be in six months. I like that question overall. There's a software development life cycle, not just generation of the code from issue to PR, and then there's the WinSurf side which is inside your IDE. Pull request review is something that people don't really—there are startups built around it. It's not something that Codex does, and it could. So then there's what else is there that is sort of rate-limiting the amount of software you could be iterating on. It's an open question. I don't know if there's an answer.

软件工程之外的 AI 未来 Future of AI beyond software engineering

Host

关于 A-sweet 还有什么别的吗?你认为在形态上会怎么发展,或者明年这个时候,我们会看到模型能做什么今天做不到的事?

Anything else on A-sweet in general? Where do you think this goes just in form factors, or what will we be looking at this time next year in terms of what models are able to do that they're not able to today?

Noam Brown

我认为它不会局限于 A-sweet。我认为它不会局限于软件工程。我认为它将能够做很多远程工作类型的任务。是的,比如自由职业者平台 Upwork。或者甚至不一定是软件工程的事情。我的想法是:任何做远程工作的人,我认为熟悉这项技术是有价值的,了解它能做什么、不能做什么、擅长什么、不擅长什么,因为我认为它能做的事情的范围也会随着时间扩大。

I don't think it's going to be limited to A-sweet. I don't think it's going to be limited to software engineering. I think it's going to be able to do a lot of remote work kind of tasks. Yeah, like freelancer type Upwork. Or just even things that are not necessarily software engineering. The way I think about it is: anybody that's doing a remote work kind of job, I think it's valuable to become familiar with the technology and get a sense of what it can do, what it can't do, what it's good at, what it's not good at, because I think the breadth of things it's going to be able to do will expand over time as well.

Host

我觉得虚拟助理可能是 A-sweet 之后的下一个方向,因为它们最容易——你知道,虚拟助理,雇一个在菲律宾的人,帮你查看邮件之类的,因为你可以完全拦截所有输入和输出,并在此基础上训练。也许 OpenAI 直接收购一家虚拟助理公司。

I feel like virtual assistants might be the next thing after A-sweet, because they're the most easily—you know, virtual assistant, hire someone in the Philippines, someone who would just look through your email and all that, because that is entirely you can intercept all the inputs and all the outputs and train on that. And maybe OpenAI just buys a virtual assistant company.

Noam Brown

是的,我期待的是,对于虚拟助理这类事情,如果模型对齐得好,它们最终可能会非常适合这类工作。总是存在委托-代理问题:如果你把任务委托给别人,他们真的会按照你的意愿去做,并且尽可能便宜和快速吗?所以如果你有一个 AI 模型真正与你和对齐,那么它最终可能比人类做得更好。嗯,不是比人类可能做得更好,而是比人类实际会做得更好。

Yeah, I think what I'm looking forward to is that for things like virtual assistants, the models, if they're aligned well, they could end up being really preferable for that kind of work. There's always this principal-agent problem: if you delegate a task to somebody, then are they really aligned with doing it as you would want it to be done, and as cheaply and quickly as they can? And so if you have an AI model that's actually really aligned to you and your preferences, then that can end up doing a way better job than a human could. Well, not that it's doing a better job than a human could, but it's doing a better job than a human would.

对齐:安全与指令遵循 Alignment: Safety vs Instruction Following

Host

顺便说一句,“对齐”这个词,我觉得在安全对齐和指令遵循对齐之间有一种有趣的覆盖或同态关系。我想知道它们在哪里分叉。

That word alignment by the way, I think there's an interesting overriding or homomorphism between safety alignment and instruction following alignment. And I wonder where they diverge.

Noam Brown

好的,我认为分叉点在于你想让模型对齐到什么?这是个难题。你可以说你想让它对齐用户。那如果用户想制造一种能消灭一半人类的新型病毒呢?那就是安全对齐了。所以我认为对齐是相关的。大问题是你在对齐什么?有人类的目标,也有个人目标,以及介于两者之间的一切。

Okay, so I think where it diverges is like what do you want to align the models to? That's a difficult question. You could say you wanted to align it to the user. Okay, well what happens if the user wants to build a novel virus that's going to wipe out half of humanity? That's safety alignment. So, there's a question of like I think alignment they're related. And the big question is what are you aligning towards? There's humanity's goals and then there's your personal goals and everything in between.

OpenAI 的多智能体团队 Multi-Agent Team at OpenAI

Host

你宣布在 OpenAI 组建了多智能体团队。我没看到太多公告,也许我错过了,你能分享一些有趣的研究方向吗?

And you announced you're releasing the multi-agent team at OpenAI. I haven't really seen many announcements, maybe I missed them on what you've been working on, but what can you share about interesting research directions or anything from this?

Noam Brown

是的,这方面还没有公告。我们在做很酷的东西,将来会宣布一些。这个团队其实名不副实,因为我们做的不仅仅是多智能体。多智能体是其中之一。其他工作包括大幅扩展测试时算力。如何让模型从思考 15 分钟变成思考几小时、几天甚至更久?从而解决极其困难的问题。这是一个方向。多智能体是另一个方向,这里有几个不同的动机。我们对多智能体的协作和竞争方面都感兴趣。我这样描述:在 AI 圈子里,人们常说人类占据一个非常狭窄的智能带,AI 会很快赶上并超越。但我认为人类智能的带宽并不窄,其实很宽。因为比较解剖学上相同的原始人和现代人,原始人在我们今天认为的智能方面并没有走多远,对吧?他们没有登月,没有制造半导体或核反应堆。而我们现在有这些,尽管人类解剖学上没有不同。那么区别是什么?我认为区别在于几千年、数十亿人类相互合作竞争,逐渐建立文明。我们看到的科技是文明的产物。类似地,今天的 AI 就像 AI 中的原始人。如果你能让它们与数十亿 AI 长期合作竞争,建立文明,它们能产生和回答的东西将远超今天 AI 的能力。

Yeah, there hasn't really been announcements on this. I think we're working on cool stuff and we'll get to announce some cool stuff at some point. The team in many ways is actually a misnomer because we're working on more than just multi-agent. Multi-agent is one of the things. Some other things we're working on is being able to scale up test time compute by a ton. So, how we get these models thinking for 15 minutes now, how do we get them to think for hours, how do we get them to think for days, even longer? And be able to solve incredibly difficult problems. So, that's one direction. Multi-agent is another direction and here I think there's a few different motivations. We're interested in both the collaborative and the competitive aspect of multi-agent. I think the way that I describe it is people often say in AI circles that humans occupy this very narrow band of intelligence and AIs are just going to quickly catch up and surpass this band. And I actually don't think that the band of human intelligence is that narrow. I think it's actually quite broad because if you compare anatomically identical humans from caveman times, they didn't get that far in terms of what we would consider intelligence today, right? Like they're not putting a man on the moon, they're not building semiconductors or nuclear reactors. And then we have those today even though we as humans are not anatomically different. So what's the difference? I think the difference is that you have thousands of years, a lot of humans, billions of humans cooperating and competing with each other, building up civilization over time. The technology that we're seeing is the product of this civilization. And I think similarly, the AIs that we have today are kind of like the cavemen of AI. And if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to produce and answer would be far beyond what is possible today with the AIs that we have today.

Host

你觉得这类似于 Jim Fan 的 Voyager 技能库想法,重新保存这些东西,还是模型随后用这些新知识重新训练?因为人类在成长过程中很多知识已经在大脑里了。

Do you see that being similar to maybe like Jim Fan's Voyager skill library idea of re-saving these things or is it just the models then being retrained on this new knowledge? Because the humans then have it a lot of it in the brain as they grow.

Noam Brown

我在这里要回避一下,说我们不会……直到我们有东西要宣布,我认为在不久的将来会有的,我会对具体做法含糊其辞,但我要说,我们处理多智能体的细节和实际方式,我认为与历史上和今天其他地方的做法非常不同。我在多智能体领域很久了,我觉得这个领域在某些方面有点误入歧途,所采取的方法和方式有问题。所以我们试图采用一种非常原则性的方法。

I think I'm going to be evasive here and say that we're not going to... until we have something to announce, which I think we will in the not too distant future, I think I'm going to be a bit vague about exactly what we're doing, but I will say that the way that we are approaching multi-agent in the details and the way we're actually going about it is I think very different from how it's been done historically and how it's being done today by other places. I've been in the multi-agent field for a long time. I've kind of felt like the multi-agent field has been a bit misguided in some ways and the approaches that the field has taken and the way that's been approached. And so I think we're trying to take a very principled approach to multi-agent.

Host

抱歉,我得问一下,你不能说你在做什么,但可以说什么是误导的。什么误导了?

Sorry, I got to ask like so you can't talk about what you're doing, but you can say what's misguided. What's misguided?

Noam Brown

我认为很多方法都非常启发式,没有真正遵循苦涩的教训那种扩展和研究的方法。

I think that a lot of approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.

扑克、GTO 与剥削 Poker, GTO, and Exploitation

Host

我想这可能是个好时机。显然你在扑克方面做了很多了不起的工作,随着推理模型变得更好,我和一个以前是硬核扑克玩家的朋友聊天,我告诉他我要采访你,他的问题是:在牌桌上,你可以从小样本中获取很多关于一个人打法的信息,但今天 GTO 太普遍了,有时人们忘了你可以剥削性地打。你认为在多智能体和竞争方面,状态是怎样的?总是试图找到最优解,还是更多地在当下思考如何剥削别人?

I think maybe this might be a good spot. Obviously you've done a lot of amazing work in poker and I think as the reasoning model got better, I was talking to one of my friends who used to be a hardcore poker grinder and I told them I was going to interview you and their question was at the table you can get a lot of information from a small sample size about how a person plays, but today GTO is like so prevalent that sometimes people forget that you can play exploitatively. What do you think is the state as you think about multi-agent and kind of like competition? Is it always going to be trying to find the optimal thing or is a lot of it trying to think more in the moment like how to exploit somebody?

Noam Brown

我猜你的听众可能不太熟悉扑克术语,所以我解释一下。很多人认为扑克只是运气游戏,其实不是。扑克中有很多策略。所以如果你用正确的策略,你可以持续赢。扑克有不同的方法。一种是博弈论最优(GTO)。这就像你玩一个在期望上不可战胜的策略,你无法被剥削。就像石头剪刀布,如果你以相等概率随机选择石头、剪刀、布,你就不可战胜,因为无论对方做什么,他们都无法剥削你,你在期望上不会输。很多人听到这个会想,那也意味着你在期望上不会赢,因为你完全是随机玩。但在扑克中,如果你玩均衡策略,对手实际上很难找到平局的方法,他们会犯错,长期来看你会赢。可能不是大赢,但会赢。如果你玩足够多的手牌,足够长的时间,你在期望上会赢。

I'm guessing your audience is probably not super familiar with poker terminology, so I'll explain this a bit. A lot of people think that poker is just a luck game and that's not true. It's actually like there's a lot of strategy in poker. So, you can win consistently in poker if you're playing the right strategy. So, there's different approaches to poker. One is game theory optimal. This is like you're playing an unbeatable strategy in expectation that you're just unexploitable. It's kind of like in rock, paper, scissors, you can be unbeatable in rock, paper, scissors if you just randomly choose between rock, paper, and scissors with equal probability, because no matter what the other guy does, they're not going to be able to exploit you or you're going to win in expectation. Now, a lot of people hear that and they think like well, that also means that you're not going to win in expectation because you're just playing totally randomly. But in poker if you play the equilibrium strategy, it's actually really difficult for the opponents to figure out how to tie you and they're going to end up making mistakes that will lead you to win over the long run. It might not be a massive win, but it is going to be a win. If you play enough hands for a long enough period of time, you're going to win in expectation.

扑克中的剥削与博弈论最优 Exploitative vs. Game Theory Optimal in Poker

Noam Brown

现在,还有剥削性扑克,其理念是试图发现对手打法的弱点。比如,他们可能诈唬不够多,或者太容易弃牌。于是你开始从博弈论最优的平衡策略——有时诈唬,有时不诈唬——转向一种非常不平衡的策略,比如‘我就使劲诈唬这个人,因为他们每次我诈唬都弃牌。’关键在于这里有一个权衡,因为如果你采取这种剥削性策略,你也会让自己暴露在被剥削的风险中。所以你必须在防守性的博弈论最优策略——保证你不会输,但可能赚不到那么多钱——和剥削性策略——可能利润更高,但也会制造弱点让对手利用和欺骗你——之间做出选择。没有完美平衡两者的方法。这有点像石头剪刀布,如果你注意到某人要出剪刀,但实际他们出了石头,你永远无法真正知道。所以你总是面临这个权衡。那些非常成功的扑克 AI——我的背景是,我在研究生期间研究扑克 AI 好几年,并制造了第一个超人类水平的无限注扑克 AI——我们采用的方法是博弈论最优方法,AI 会执行这种不可战胜的策略,与世界顶尖选手对战并击败他们。这也意味着他们能击败最差的选手。他们能击败任何人。但如果面对一个弱对手,他们可能不会像人类专家那样赢得那么狠,因为人类专家知道如何从博弈论最优策略中调整,去剥削这些弱玩家。所以有一个未解的问题:如何制造一个剥削性扑克 AI?很多人追求过这个研究方向。我在研究生期间也稍微涉猎过。我认为根本原因在于 AI 的样本效率不如人类,就像我们之前讨论的。如果人类打扑克,他们能在十几手牌内就对对手的强弱有很好的感觉。这真的很令人印象深刻。在我们 2010 年代中期研究扑克 AI 时,这些 AI 需要打上万手牌才能对玩家是谁、怎么打、弱点在哪有好的画像。我认为随着更新的技术,这个数字已经下降了,但样本效率仍然是一个大挑战。

Now, there's also exploitative poker and the idea here is that you're trying to spot weaknesses in how the opponent plays. You know, maybe they're not bluffing enough or maybe they fold too easily to a bluff. So you start adapting from the game theory optimal balanced strategy of bluffing sometimes and not bluffing sometimes, to then playing a very unbalanced strategy like, 'Oh, I'm just going to bluff a ton against this person because they always fold whenever I bluff.' The key is that there's a trade-off here because if you're taking this exploitative approach, then you're opening yourself up to exploitation as well. So you have to choose this balance between playing a defensive game theory optimal policy that guarantees you're not going to lose, but might not make you as much money as you potentially could, versus playing an exploitative strategy that could be much more profitable, but also creates weaknesses that the opponents could take advantage of and trick you. There's no way to perfectly balance the two. It's kind of like in rock, paper, scissors, if you notice somebody is about to throw scissors, but actually that's the time when they throw rock, you know? So you never really know. You always have this trade-off. The poker AIs that have been extremely successful—and my background is I worked on AI for poker for several years during grad school and made the first superhuman no-limit poker AIs—the approach we took was this game theory optimal approach where the AIs would play this unbeatable strategy and they would play against the world's best and beat them. That also means they beat the world's worst. They would just beat anybody. But if they were up against a weak opponent, they might not beat them as severely as a human expert might, because the human expert would know how to adapt from the game theory optimal policy to exploit these weak players. So there's this kind of unanswered question of how do you make an exploitative poker AI? A lot of people have pursued this research direction. I dabbled in it a little bit during grad school. I think fundamentally it just comes down to AIs not being as sample efficient as humans, as we discussed earlier. If a human is playing poker, they're able to get a really good sense of the strengths and weaknesses of a player within a dozen hands. It's honestly really impressive. Back when we were working on AI for poker in the mid-2010s, these AIs would have to play like 10,000 hands of poker to get a good profile of who this player is, how they're playing, where their weaknesses are. I think with more recent technology that has come down, but the sample efficiency has still been a big challenge.

Host

有趣的是,在研究了扑克之后,我转而研究外交。我想我们之前提到过。外交是一个七人谈判游戏。当我们开始研究时,我采取了非常博弈论的方法。我觉得,好吧,这有点像扑克,你需要计算博弈论最优策略,然后执行它,你期望不会输,实际上会赢。但这种方法在外交中行不通。它行不通——这取决于我们想深入探讨多少——但基本上,当你玩像扑克这样的零和游戏时,博弈论最优非常有效。当你玩像外交这样需要合作与竞争、有合作空间的游戏时,博弈论最优实际上效果不佳。你必须更好地理解玩家并适应他们。所以这最终与扑克中如何适应对手的问题非常相似。在扑克中,是适应他们的弱点并加以利用。在外交中,是适应他们的游戏风格。这有点像,如果你在一张桌子上,所有人都说法语,你不想一直说英语。你想适应他们,也说法语。这就是我在外交中的领悟:我们需要从博弈论最优范式转向建模其他玩家,理解他们是谁,然后做出相应回应。所以从很多方面来说,我们在外交中开发的技术是剥削性的。它们不是剥削性的——它们实际上只是适应对手,适应桌上的其他玩家。但我认为同样的技术可以用于扑克 AI,制造剥削性扑克 AI。如果不是被语言模型令人难以置信的进展所吸引,并将我的整个研究议程转向通用推理,我接下来可能就会研究制造这些剥削性扑克 AI。这将是一个非常有趣的研究方向。我认为它仍然在那里,等着任何想做的人去探索。我认为关键是将我们在外交中使用的技术应用到扑克之类的事情上。

Now, what's interesting is that after working on poker, I worked on diplomacy. I think we talked about this earlier. Diplomacy is a seven-player negotiation game. When we started working on it, I took a very game theory approach. I felt like, okay, it's kind of like poker, you have to compute this game theory optimal policy and you just play this, you're going to not lose in expectation, you're going to win in practice. But that actually doesn't work in diplomacy. It doesn't work again—it's a question of how much of a rabbit hole we want to go down on this—but basically when you're playing zero-sum games like poker, game theory optimal works really well. When you're playing a game like diplomacy where you need to collaborate and compete, and there's room for collaboration, then game theory optimal actually doesn't work that well. You have to understand the players and adapt to them much better. So this ends up being very similar to the problem in poker of how do you adapt to your opponents? In poker, it's about adapting to their weaknesses and taking advantage of that. In diplomacy, it's about adapting to their play styles. It's kind of like if you're at a table and everybody's speaking French, you don't want to just keep talking in English. You want to adapt to them and speak in French as well. That's the realization I had with diplomacy: we need to shift away from this game theory optimal paradigm towards modeling the other players, understanding who they are, and then responding accordingly. So in many ways, the techniques we developed in diplomacy are exploitative. They're not exploitative—they're really just adapting to the opponents, to the other players at the table. But I think the same set of techniques could be used in AI for poker to make exploitative poker AIs. If I hadn't gotten pulled by the incredible progress we were seeing with language models and shifted my whole research agenda to focusing on general reasoning, probably what I would have worked on next was making these exploitative poker AIs. It would be a really fun research direction to go down. I think it's still there for anybody that wants to do it. I think the key would be taking the techniques we used in diplomacy and applying them to things like poker.

Host

对我来说,核心在于,当你在线玩时,你有一个 HUD,它告诉你关于对手的所有统计数据,比如他们翻牌前参与频率等等。在我看来,很多这些模型,据我所知,并没有真正利用桌上其他玩家的行为。它们只是看着牌面状态,然后据此行动。

I think to me that core piece is when you play online, you have a HUD which tells you all these stats about the other player, like how much they participate pre-flop, blah blah blah. And to me it's like a lot of these models, from my understanding, are not really leveraging the behavior of the other players at the table. They're just kind of looking at the board state and working from there.

Noam Brown

没错。如今扑克 AI 的工作方式,就是坚持它们预先计算好的 GTO 策略,并不适应桌上的其他玩家。你可以用各种取巧的方法让它们适应,但这些方法不太有原则,效果也不是很好。

That's correct. The way the poker AIs work today, they're just kind of sticking to their pre-computed GTO strategy, and they're not adapting to the other players at the table. You can do various hacky things to get them to adapt, but they're not very principled, they don't work super well.

Host

是的。好吧,任何在听的研究生,如果你想研究这个,我认为这是一个非常合理的研究方向,至少能让你脱颖而出,获得一些关注。

Yep. Okay, any grad students listening, if you want to work on that, I think that is a very reasonable research direction that would at least get you in front of and get some attention at least.

Host

这段对话让我想到的另一件事是,关于测试时计算之后下一步的一个假设是世界模型。世界建模是一个重要或有价值的研究方向吗?就像 Yann LeCun 一直在谈论这个。但基本上没有 LLM 有显式的世界模型,尽管它们有内部世界模型。我认为很明显,随着这些模型变大,它们会有一个世界模型,并且这个世界模型会随着规模扩大而变得更好。

The other thing that this conversation brings up for me is one of the hypotheses for what is the next step after test time compute is world models. Is world modeling an important or worthwhile research direction? Like Yann LeCun has been talking about this nonstop. But basically no LLMs have explicit world models, though they have internal world models. I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale.

世界模型与多智能体系统 World Models and Multi-Agent Systems

Noam Brown

所以,它们是在隐式地发展一个世界模型,我不认为需要显式建模。我可能错了。当处理人类或多智能体时,可能是因为存在非世界的实体,你需要解析你面对的是哪种实体。多智能体 AI 社区长期争论是否需要显式建模其他智能体,还是可以将它们隐式建模为环境的一部分。很长一段时间,我认为必须显式建模这些其他智能体,因为它们的行为与环境不同——它们会行动、不可预测、有自主性。但随着时间的推移,我转变了想法:如果模型足够聪明,它们会发展出心智理论,理解存在其他能行动、有动机的智能体,而这些模型会随着规模和更广泛的能力隐式地发展出这一点。这就是我现在的观点。所以,我刚才说的就是一个不符合苦涩教训的启发式方法,它最终会消失。

So, they are implicitly developing a world model, and I don't think it's something that you need to explicitly model. I could be wrong about that. When dealing with people or multi-agents, it might be because you have entities that are not the world, and you're resolving hypotheses of which of the many types of entities you could be dealing with. There was a long debate in the multi-agent AI community about whether you need to explicitly model other agents, or if they can be implicitly modeled as part of the environment. For a long time, I took the perspective that you have to explicitly model these other agents because they behave differently from the environment—they take actions, they're unpredictable, they have agency. But I think I've shifted over time to thinking that if these models become smart enough, they develop things like theory of mind, an understanding that there are other agents that can take actions and have motives, and these models just develop that implicitly with scale and more capable behavior broadly. That's the perspective I take these days. So, what I just said was an example of a heuristic that is not Bitter Lesson filled, and it just goes away.

Host

是啊,一切都回到苦涩教训。每个 AI 播客都得引用它。那么,一个有趣且最一致的发现——我记得你在 ICLR,那里一个热门演讲是关于开放性的,演讲者 Tim 也做了很多多智能体系统的研究。最一致的发现之一是,AI 通过自我对弈和竞争性改进,比人类训练和引导效果更好。这在 AlphaZero 和 R0 等系统中得到了验证。你认为这在多智能体场景下也成立吗?即自我对弈比人类指导更好?

Yeah, it all comes back to the Bitter Lesson. Got to cite them every AI podcast. So, one of the interesting and most consistent findings—I think you were at ICLR, and one of the hit talks there was about open-endedness, and Tim who gave that talk has been doing a lot of research on multi-agent systems too. One of the most consistent findings is that it's better for AIs to self-play and improve competitively as opposed to humans training and guiding them. You find that with AlphaZero and R0, whatever that was. Do you think this will hold for multi-agents? Like self-play to improve better than humans?

Noam Brown

是的,这是个好问题,值得展开。我认为很多人现在把自我对弈视为下一步,甚至可能是超级智能所需的最后一步。看看 AlphaGo 和 AlphaZero,我们似乎遵循着非常相似的轨迹。AlphaGo 的第一步是在人类围棋对局上进行大规模预训练。对于大语言模型,是在海量互联网数据上预训练。这能给你一个强大的模型,但并非极其强大——达不到超人水平。AlphaGo 范式的下一步是大规模测试时算力或推理算力,比如 MCTS,现在我们有了推理模型,也做大规模推理算力。这极大地提升了能力。最后,在 AlphaGo 和 AlphaZero 中,有自我对弈:模型与自己下棋,从中学习,变得越来越强,从接近人类水平到远超人类能力。这些围棋策略现在强到人类无法理解它们的行为。国际象棋也是如此。我们目前在大语言模型上还没有达到这一点。所以很容易认为,我们只需要让 AI 模型相互交互、相互学习,它们就能达到超级智能。但挑战在于,正如我在外交游戏中提到的,围棋是一个两人零和博弈。两人零和博弈有一个很好的性质:自我对弈会收敛到极小化极大均衡。在国际象棋、围棋甚至两人扑克这类两人零和博弈中,你通常想要的是极小化极大均衡——一种保证你在期望上不会输给任何对手的 GTO 策略。在国际象棋和围棋中,这显然是你想要的。在扑克中则不那么明显:你可以玩 GTO 保证不输,但剥削性策略能让你从弱手那里赢更多钱。所以问题在于你想要什么。所有 AI 开发者都选择了极小化极大策略,而恰好自我对弈收敛到它。但一旦超出两人零和博弈,比如外交游戏,这个策略就不再有用。你不想要一个防御性策略,如果在数学等领域做同样的自我对弈,你会得到奇怪的行为。例如,数学中的自我对弈意味着什么?你可能陷入一个陷阱:让一个模型提出极难的问题,另一个模型解答——这是一个两人零和博弈。但你可以提出极难但无趣的问题,比如 30 位数乘法。这在我们想要的维度上取得进展了吗?并没有。所以,在两人零和博弈之外,自我对弈变成了一个更困难、更微妙的问题。Tim 在他的演讲中也说了类似的话:当你开始讨论两人零和博弈之外的自我对弈时,在决定优化目标上存在很多挑战。我的观点是,这就是 AlphaGo 类比失效的地方。

Yeah, this is a great question, and I think it's worth expanding on. I think a lot of people today see self-play as the next step and maybe the last step we need for superintelligence. If you look at something like AlphaGo and AlphaZero, we seem to be following a very similar trend. The first step in AlphaGo was large-scale pre-training on human Go games. With LLMs, it's pre-training on tons of internet data. That gets you a strong model, but not an extremely strong one—not superhuman. The next step in the AlphaGo paradigm is large-scale test-time compute or inference compute, like MCTS, and now we have reasoning models that also do large-scale inference compute. That boosts capabilities a ton. Finally, with AlphaGo and AlphaZero, you have self-play where the model plays against itself, learns from those games, gets better and better, and goes from around human performance to way beyond human capability. These Go policies are now so strong that what they're doing is incomprehensible to humans. Same with chess. We don't have that right now with language models. So it's tempting to say that we just need AI models to interact with each other and learn from each other, and then they'll get to superintelligence. The challenge, as I mentioned with diplomacy, is that Go is a two-player zero-sum game. Two-player zero-sum games have a nice property: when you do self-play, you converge to a minimax equilibrium. In two-player zero-sum games like chess, Go, even two-player poker, what you typically want is a minimax equilibrium—a GTO policy that guarantees you won't lose to any opponent in expectation. In chess and Go, that's clearly what you want. In poker, it's less obvious: you could play GTO and guarantee not to lose, but you won't beat weak players as much as you could with an exploitative policy. So there's a question of what you want. All AI developers in these games have decided to choose the minimax policy, and conveniently, that's exactly what self-play converges to. But once you go outside two-player zero-sum games, like in diplomacy, that's not a useful policy anymore. You don't want a defensive policy, and you'll get weird behavior if you do the same kind of self-play in things like math. For example, what does self-play in math mean? You could fall into the trap of having one model pose really difficult questions and the other solve them—that's a two-player zero-sum game. But you could just pose really difficult but uninteresting questions, like 30-digit multiplication. Is that making progress in the dimension we want? Not really. So self-play outside two-player zero-sum games becomes a much more difficult, nuanced question. Tim basically said something similar in his talk: there are a lot of challenges in deciding what you're optimizing for when you start to talk about self-play outside two-player zero-sum games. My point is that this is where the AlphaGo analogy breaks down.

超越 AlphaGo 的自对弈目标函数 Objective Function for Self-Play Beyond AlphaGo

Host

而且不一定会崩溃,但不会像 AlphaGo 中的自我对弈那么容易。那么目标函数是什么?新的目标函数是什么?

And not necessarily breaks down, but like it's not going to be as easy as self-play was in AlphaGo. What is the objective function then for that? What is the new objective function?

Noam Brown

嗯,这是个好问题。我认为这是很多人都在思考的事情。

Yeah, it's a good question. And I think that's something that a lot of people are thinking about.

对 Sora 与自回归图像生成的印象 Impression of Sora and Autoregressive Image Generation

Host

我相信你也是。在你最近做的一个播客中,你提到你对 Sora 印象深刻。你不直接参与 Sora 的工作,但它显然是 OpenAI 的一部分。我认为生成媒体领域最新的更新是自回归图像生成。这有什么有趣或令人惊讶的地方吗?你想评论一下吗?

I'm sure you are. One of the last podcasts that you did, you mentioned that you were very impressed by Sora. You don't work directly on Sora, but obviously it's part of OpenAI. I think the most recent new updates in that sort of generative media space is autoregressive image generation. Is that interesting or surprising in any way that you want to comment about?

Noam Brown

我不做图像生成,所以我的评论能力有限,但我要说我爱它。我觉得它超级令人印象深刻。这就像你研究这些推理模型,想着,哇,我们能够做各种疯狂的事情,比如推动科学进步、解决智能体任务和软件工程。然后还有另一个维度的进步,你发现现在可以制作图像和视频了,而且非常有趣。说实话,这吸引了更多的关注,尤其是普通公众,可能也推动了更多 ChatGPT 的订阅计划,这很棒。但我觉得有点好笑的是,是的,我们也在研究超级智能。但你可以让一切变得滑稽。

I don't work on image generation, so my ability to comment is limited, but I will say I love it. I think it's super impressive. It's one of those things where you work on these reasoning models and think, wow, we're going to be able to do all sorts of crazy stuff, like advance science, solve agentic tasks, and software engineering. And then there's this other dimension of progress where you're like, oh, you're able to make images and videos now, and it's so much fun. And that's getting a lot more attention, to be honest, especially from the general public, and it's probably driving a lot more of the subscription plans for ChatGPT, which is great. But I think it's just kind of funny that, yeah, we're also working on superintelligence. But you can make everything googly.

图像生成的扩散与自回归对比 Diffusion vs Autoregressive for Image Generation

Host

对我来说,关键点是我之前一直持有这样一个论点:扩散模型因为自回归生成而终结了。去年年底有传言,现在显然已经出现了。然后 Gemini 推出了文本扩散,扩散模型又回来了。这是两个方向,对于自回归与扩散的推理非常相关。我们两者都有吗?其中一个会胜出吗?

I think the delta for me was I was actually harboring this thesis that diffusion was over because of autoregressive generation. There were rumors about this end of last year and obviously now it's come out. Then Gemini comes out with text diffusion and diffusion is so back. And this is two directions and it's very relevant for inference of autoregressive versus diffusion. Do we have both? Does one win?

Noam Brown

研究的美妙之处在于你必须探索不同的方向,并不总是清楚哪条路有前景。我认为人们研究不同方向、尝试不同事物是很好的。这种探索很有价值,我们都能从看到什么有效中受益。

The beauty of research is that you have to pursue different directions, and it's not always clear what the promising path is. I think it's great that people are looking into different directions and trying different things. I think there's a lot of value in that exploration, and we all benefit from seeing what works.

Host

扩散推理有什么潜力吗?比如说你的思维链……可能回答不了。

Any potential in diffusion reasoning? Let's say your chain... Probably can't answer that.

对机器人技术与类人机器人的思考 Thoughts on Robotics and Humanoids

Host

你也读过机器人学的硕士。我们很想听听你的想法,比如 OpenAI 从转笔技巧和想造的机械臂开始。研究人形腿是对的吗?你认为这有点像 AI 的错误具身形式,除了通常的,你知道,多久才能有机器人等等。你觉得现在有什么根本性的东西没有被探索,人们真的应该在机器人学中去做?

So you did a master's in robotics, too. We'd love to get your thoughts on one, you know, OpenAI going to start with the pen spinning trick and the robotic arm they wanted to build. Is it right to work on the humanoid legs? Do you think that's kind of like the wrong embodiment of AI outside of the usual, you know, how long until we get robots, blah blah blah. Is there something that you think is fundamentally not being explored right now that people should really be doing in robotics?

Noam Brown

我几年前读过机器人学硕士,从那次经历中得到的教训是——首先,我实际上并没有怎么和机器人打交道。我名义上在机器人学项目中。第一周我玩了一些乐高机器人,但很快我就转向了研究扑克 AI,只是名义上在机器人学硕士项目里。但通过与这些机器人学家的互动和看到他们的研究,我的结论是我不想研究机器人,因为当你处理物理硬件时,研究周期要慢得多,也更痛苦。软件进展快得多,我认为这就是为什么我们在语言模型和虚拟同事任务上看到了这么多进展,但在机器人学上进展不大——物理硬件迭代起来痛苦得多。

I did a master's in robotics years ago and my takeaway from that experience—first of all, I didn't actually work with robots that much. I was technically in a robotics program. I played around with some LEGO robots my first week, but then I pretty quickly shifted to working on AI for poker and was kind of nominally in the robotics master's. But my takeaway from interacting with all these roboticists and seeing their research was that I did not want to work on robots because the research cycle is so much slower and more painful when you're dealing with physical hardware. Software goes much more quickly, and I think that's why we're seeing so much progress with language models and virtual coworker tasks, but haven't seen as much progress in robotics—physical hardware just is much more painful to iterate on.

Noam Brown

关于人形机器人的问题,我没有很强的观点,因为这不是我的研究方向,但我认为非人形机器人也有很多价值。我认为无人机是一个完美的例子,显然很有价值。那是人形吗?不是,但在很多方面这很好。你不需要人形机器人来做那种技术。我稍微倾向于认为非人形机器人提供了很多价值。

On the question of humanoids, I don't have very strong opinions because this isn't what I'm working on, but I think there's a lot of value in non-humanoid robotics as well. I think drones are a perfect example where there's clearly a lot of value. Is that a humanoid? No, but in many ways that's great. You don't want a humanoid for that kind of technology. I think weakly I think that non-humanoids provide a lot of value.

Noam Brown

我读了 Richard Hamming 的《科学和工程的艺术》,他谈到当出现新的技术转变时,人们会试图将旧的工作负载复制到新技术中,而实际上你必须改变做事的方式。当我看到你的机器人在房子里的视频时,我想,人类形态有很多局限性,实际上可以改进,但我认为人们想要熟悉的东西。你会把一个有 10 条胳膊和 5 条腿的机器人放在家里吗?或者晚上起床看到那个东西走来走去会不会很诡异?这就是我们使用人形机器人的原因吗?所以我认为这几乎是一个局部最优,让它看起来像人,但我觉得房子里最好的形状是什么?我不是问这个的人。我认为有一个问题:制造人形机器人更好,因为它们对我们更熟悉,还是更糟,因为它们和我们更相似但不完全一样?我不知道哪个实际上更让我觉得毛骨悚然。

I was reading Richard Hamming's The Art of Doing Science and Engineering and he talks about how when you have a new technological shift, people try to take the old workloads and replicate them just in the new technology versus you actually have to change the way you do it. When I see this video of your humanoid in the house, it's like, well, the human shape has a lot of limitations that can actually be improved, but I think people want what's familiar. Would you put a robot with 10 arms and five legs in your house? Or would it be eerie at night when you get up and see that thing walking around? And is that why we use humanoids? So I think there's almost this local maximum of making it look like a human, but I think what's the best shape in a house? I'm not the person to ask on this. I think there is a question of whether it's better to make humanoids because they're more familiar to us, or worse because they're more similar to us but not quite identical. I don't know which one I would actually find creepier.

Host

是的。让我稍微倾向于人形机器人的一个论点是,世界大部分是为人类设计的,所以如果你想取代人类劳动,你必须制造人形机器人。我不知道这有没有说服力。

Yeah. The thing that got me humanoid-pilled a little bit was just the argument that most of the world is made for humans anyway, so if you want to replace human labor, you have to make a humanoid. I don't know if that's convincing.

Noam Brown

再说一次,我对这个领域没有很强的观点,因为我不做这个。我稍微倾向于人形机器人,但真正说服我稍微倾向于非人形机器人的是听了 Physical Intelligence 的 CEO 的一些演讲,关于他们为什么不追求人形机器人。而且他们的办公室离这里很近,所以如果你想……他们会在我要主持的会议上发言。

Again, I don't have very strong opinions in this field because I don't work in it. I was weakly in favor of humanoids, and I think what really persuaded me to be weakly in favor of non-humanoids was listening to the Physical Intelligence CEO and some of his pitches about why they're not pursuing humanoid robotics. And conveniently their office is actually very close to here, so if you wanted to... They're speaking at the conference I'm running.

Host

好的,完美。是的,所以我说,听听他的演讲,也许他能说服你非人形是正确方向。太棒了。另一个我想推荐的是 Jim Fan 最近在 Sequoia 会议上做的关于物理图灵测试的演讲,非常非常好。他是个很棒的讲解者。这很难,尤其是在那个领域。酷。我们不再问你你不做的工作了。

Okay, perfect. Yeah, so you know, I'd say listen to his pitch and maybe he can convince you that non-humanoid is the way to go. Awesome. The other one I would refer people to is Jim Fan recently did a talk on the physical Turing test, which he did at the Sequoia conference, which was very very good. He's such a great educator and explainer of things. It's very hard, especially in that field. Cool. We're done asking you about things that you don't work on.

紧跟研究前沿 Keeping up with research

Host

所以这只是更多快速问答,来探索你的一些边界,得到一些快速回答。你或顶尖行业实验室如何跟上研究进展?你的工具和做法是什么?

So these are just more rapid fires to explore some of your boundaries and get quick hits. How do you or top industry labs keep on top of research? What are your tools and practices?

Noam Brown

这真的很难。我认为很多人有一种看法,认为学术研究无关紧要,但实际上并非如此。我们确实会看学术研究。挑战之一是很多学术研究在论文中显示出前景,但实际上无法规模化,甚至无法复现。如果我们发现有趣的论文,我们会尝试在内部复现,看看是否仍然成立,以及是否能够很好地扩展。但那是我们灵感的重要来源。

It's really hard. I think a lot of people have this perception that academic research is irrelevant, and this is actually not the case. I think we do look at academic research. One of the challenges is that a lot of academic research shows promise in their papers, but then actually doesn't work at scale or even doesn't replicate. If we find interesting papers, we try to reproduce that in-house and see if it still holds up and also if it scales well. But that is a big source of inspiration for us.

Host

任何出现在 arXiv 上的东西,你和我们其他人一样处理吗?还是你有特殊流程?尤其是我收到推荐的时候。

Whatever hits arXiv, you do the same as the rest of us? Or do you have a special process? Especially if I get recommendations.

Noam Brown

我们有一个内部频道,人们会发布有趣的论文。我认为这是一个好来源:如果某个更熟悉这个领域的人认为一篇论文有趣,那么我就应该读它。同样,我会跟踪我所在领域发生的我认为有趣的事情,如果我觉得非常有趣,我可能会分享。对我来说,就是与研究人员在 WhatsApp 和 Signal 群聊,仅此而已。

We have an internal channel where people will post interesting papers. I think that's a good source: if someone more familiar with this area thinks a paper is interesting, then I should read it. Similarly, I'll keep track of things happening in my space that I think are interesting, and if I think it's really interesting, maybe I'll share it. For me, it's WhatsApp and Signal group chats with researchers, and that's it.

Host

是的。我认为很多人会看 Twitter 之类的东西,很不幸我们已经到了这个地步:事情需要在社交媒体上获得大量关注才能被注意到。

Yeah. I think a lot of people look at things like Twitter, and it's really unfortunate that we've reached this point where things need to get a lot of attention on social media to be paid attention to.

Noam Brown

这就是研究生们所受的训练。他们上课就是为了这个。我确实向我合作过的研究生推荐过——当我在 FAIR 发表论文时,我会告诉他们需要在 Twitter 上发布,我们会讨论如何展示工作的推文。这确实是一门艺术,而且很重要。这有点可悲,但事实如此。

That's what the grad students are trained. They're taking classes to do this. I do recommend to grad students I've worked with—when I was at FAIR publishing papers, I would tell them you need to post it on Twitter, and we go over the Twitter thread about how to present the work. There's a real art to it, and it does matter. It's kind of the sad truth.

研究的环境限制因素 Environmental limiters on research

Host

我知道当你参加 ACPC(AI 扑克比赛)时,你提到人们没有做搜索,因为他们在推理时只限于两个 CPU。你今天是否看到类似的事情,阻碍了有趣的研究进行,这些研究可能不那么流行,不能让你进入顶级会议?比如是否存在一些环境限制因素?

I know when you were doing the ACPC, the AI poker competition, you mentioned that people were not doing search because they were limited to like two CPUs at inference. Do you see similar things today that are keeping interesting research from being done that maybe is not as popular, doesn't get you into the top conferences? Like are there some environmental limiters?

Noam Brown

绝对有。我认为一个例子是基准测试。你看像“人类最后考试”这样的东西。这些问题非常困难,但仍然很容易评分。我认为如果你坚持这种范式,实际上会限制你可以评估这些模型的范围。这很方便,因为给模型打分很容易,但实际上我们想要评估模型的很多任务都是更模糊的任务,不是多项选择题。为这类事情制作基准测试要困难得多,而且评估成本也可能高得多。但我认为这些都是非常有价值的工作。

Absolutely. I think one example is benchmarks. You look at things like Humanity's Last Exam. You have these incredibly difficult problems, but they are still very easily gradable. I think that actually limits the scope of what you can evaluate these models on if you stick to that paradigm. It's very convenient because it's very easy to score the models, but actually a lot of the things we want to evaluate these models on are more fuzzy tasks that are not multiple choice questions. Making benchmarks for those kinds of things is so much harder and probably also a lot more expensive to evaluate. But I think those are really valuable things to work on.

Host

这符合开创性的 GPT-4.5 在某种程度上是一个高品位的模型。模型有很多不可测量的东西非常好,也许人们没有……

And that would fit the seminal GPT-4.5 is like a high taste model in a way. There's kind of all these non-measurable things about a model that are really good that maybe people are not...

Noam Brown

嗯,我认为有些东西是可测量的,但测量起来要困难得多。我认为很多基准测试都固守于这种提出非常困难但非常容易测量的问题的范式。

Well, I think there are things that are measurable, but they're just much more difficult to measure. I think a lot of benchmarks have kind of stuck to this paradigm of posing really difficult problems that are really easy to measure.

测试时计算扩展瓶颈 Test-time compute scaling wall

Host

那么假设预训练 Scaling 范式从 GPT 的发现到扩展到 GPT-4 花了大约五年。然后我们也给测试时算力五年时间。所以如果测试时算力在 2030 年遇到瓶颈,可能的原因是什么?

So let's say that the pre-training scaling paradigm took about five years from discovery of GPT to scaling it up to GPT-4. And then we give test-time compute five years as well. So if test-time compute hit a wall by 2030, what would be the probable cause?

Noam Brown

这与预训练非常相似:你可以将预训练推得更远,但每次迭代成本更高。我认为我们会在测试时算力上看到类似情况。我们会让它们思考三分钟,然后三小时,然后三天,然后三周。或者你耗尽人类寿命。有两个担忧。一是让模型思考那么长时间或扩展测试时算力变得非常昂贵。随着你扩展测试时算力,你在上面花费更多,所以你能花费的金额有限。这是一个潜在的上限。但我们也变得更高效。这些模型在思考方式上变得更高效;它们能用相同的测试时算力做更多事情。我认为这是一个被严重低估的点。不仅仅是让模型思考更长时间。事实上,如果你看 O3,它对某些问题的思考时间比 O1 预览版更长,但差异并不大,但效果却好得多。为什么?因为它变得更善于思考了。无论如何,你只能将测试时算力扩展到一定程度。这成为一个软障碍,就像预训练变得越来越昂贵一样。第二点是,随着这些模型思考时间更长,你会受到挂钟时间的瓶颈。如果你想迭代实验,当模型即时响应时很容易,但当它们需要三小时时就很困难。当它们需要三周时会发生什么?至少需要三周才能完成那些评估并迭代。你可以在一定程度上并行化实验,但很多情况下你必须完整运行实验并看到结果才能决定下一步。我认为这实际上是长周期的最有力论据:因为模型必须在串行时间内做很多事情,我们只能如此快速地迭代。

It's very similar to pre-training: you can push pre-training a lot further, it just becomes more expensive with each iteration. I think we're going to see something similar with test-time compute. We'll get them thinking for three minutes, then three hours, then three days, then three weeks. Or you run out of human life. There are two concerns. One is that it becomes much more expensive to get the models to think for that long or scale up test-time compute. As you scale up test-time compute, you're spending more on it, so there's a limit to how much you can spend. That's one potential ceiling. But we're also becoming more efficient. These models are becoming more efficient in the way they think; they're able to do more with the same amount of test-time compute. I think that's a very underappreciated point. It's not just that we're getting these models to think for longer. In fact, if you look at O3, it's thinking for longer than O1 preview for some questions, but it's not a radical difference, but it's way better. Why? Because it's just becoming better at thinking. Anyway, you can only scale up test-time compute so much. That becomes a soft barrier in the same way that pre-training is becoming more and more expensive. The second point is that as you have these models think for longer, you get bottlenecked by wall clock time. If you want to iterate on experiments, it's easy when models respond instantly, but much harder when they take three hours. And what happens when they take three weeks? It takes at least three weeks to do those evaluations and iterate. You can parallelize experiments to some extent, but a lot of it you have to run the experiment completely and see the results to decide on the next set. I think this is actually the strongest case for long timelines: because models have to do so much in serial time, we can only iterate so quickly.

Host

你会如何克服这个障碍?

How would you overcome that wall?

Noam Brown

这是一个挑战,我认为这取决于领域。药物发现是一个可能成为真正瓶颈的领域。如果你想看看某物是否能延长人类寿命,需要很长时间才能弄清楚我们开发的新药是否真的能延长人类寿命,并且没有严重的副作用。

It's a challenge, and I think it depends on the domain. Drug discovery is one domain where this could be a real bottleneck. If you want to see if something extends human life, it's going to take a long time to figure out if the new drug we developed actually extends human life and doesn't have terrible side effects along the way.

完美的人类生物学模型 Perfect Models of Human Biology

Host

顺便问一下,我们现在难道还没有完美的人类化学和生物学模型吗?

Side note, do we not have perfect models of human chemistry and biology by now?

Noam Brown

嗯,这就是问题所在。我想谨慎一点,因为我不是生物学家或化学家。我对这些领域知之甚少。我上次上生物课还是高中十年级。我认为目前还没有完美的人类生物学模拟器,而这可能有助于解决这个问题。这是我们都应该努力的事情之一。

Well, this is the thing. I want to be cautious because I'm not a biologist or chemist. I know very little about these fields. The last time I took a biology class was 10th grade. I don't think there's a perfect simulator of human biology right now, and that could potentially help address this problem. It's one of the things we should all work on.

Host

嗯,这正是我们希望这些推理模型能帮助我们的事情之一。

Well, that's one of the things we're hoping these reasoning models will help us with.

中期训练与后期训练 Mid-Training vs Post-Training

Host

你如何区分今天的中期训练和后训练?

How would you classify mid-training versus post-training today?

Noam Brown

所有这些定义都很模糊。我没有很好的答案。这是人们提出的问题。OpenAI 现在明确招聘中期训练岗位,大家都问:“中期训练到底是什么?”我认为中期训练介于预训练和后训练之间。它不是后训练,也不是预训练。它是在预训练之后以有趣的方式为模型添加更多内容。

All these definitions are so fuzzy. I don't have a great answer. It's a question people have. OpenAI is now explicitly hiring for mid-training, and everyone is like, 'What the hell is mid-training?' I think mid-training is between pre-training and post-training. It's not post-training, it's not pre-training. It's like adding more to the models after pre-training in interesting ways.

Host

预训练模型现在基本上是一个生成其他模型的产物,核心预训练模型从未真正暴露过吗?中期训练是新的预训练,而模型分支后才有后训练吗?

Is the pre-trained model now basically an artifact that spawns other models, and the core pre-training model is never really exposed? Is mid-training the new pre-training, and post-training happens once models branch out?

Noam Brown

你永远不会与原始的预训练模型交互。如果你要使用模型,它会经过中期训练和后训练。所以你看到的是最终产品。

You never interact with a raw pre-trained model. If you're going to interact with the model, it goes through mid-training and post-training. So you're seeing the final product.

Host

嗯,你不让我们这么做,但你知道的。

Well, you don't let us do it, but you know.

Noam Brown

是的。我们以前允许。有些开源模型你可以直接与原始预训练模型交互。但对于 OpenAI 模型,它们会经过中期训练步骤,然后是后训练步骤,然后发布。它们有用得多。坦率地说,如果你只与预训练模型交互,它会非常难以使用,而且看起来有点笨。

Yeah. We used to. There are open source models where you can interact with the raw pre-trained model. But for OpenAI models, they go through a mid-training step, then a post-training step, and then they're released. They're a lot more useful. Frankly, if you interacted with only the pre-trained model, it would be super difficult to work with and would seem kind of dumb.

Host

但它在奇怪的方式下会有用,因为当你为聊天进行后训练时,会出现模式坍缩。在某些方面你希望这种坍缩,即分布的坍缩。我明白。

But it would be useful in weird ways, because there's a mode collapse when you post-train for chat. In some ways you want that mode collapse, that collapse of the distribution. I get it.

采访格雷格·布罗克曼 Interviewing Greg Brockman

Host

我们接下来要采访 Greg Brockman。你和他聊过很多。你会问他什么?

We're interviewing Greg Brockman next. You've talked to him a lot. What would you ask him?

Noam Brown

我会问 Greg 什么?我经常问 Greg。你应该问 Greg 什么才能引发有趣的回答,一个他不常被问到的问题,一个他热衷的话题,或者你只是想知道他的想法?总的来说,值得问的是未来的走向。5 年后的世界实际上是什么样子?10 年后呢?结果分布是什么样的?世界或个人可以做些什么来帮助引导事情走向好的结果而不是坏的结果?这是一个对齐问题。人们非常关注未来 1 到 2 年会发生什么。花时间思考 5 年或 10 年后会发生什么也很有价值。他没有水晶球,但他肯定有想法。所以我认为这值得探讨。

What would I ask Greg? I get to ask Greg all the time. What should you ask Greg to evoke an interesting response that he doesn't get asked enough about, something he's passionate about or you just want his thoughts? I think in general, it's worth asking where this goes. What does the world actually look like in 5 years? What does it look like in 10 years? What does that distribution of outcomes look like? And what could the world or individuals do to help steer things towards good outcomes instead of negative ones? That's an alignment question. People get very focused on what's going to happen in 1 or 2 years. It's also worth spending time thinking about what happens in 5 or 10 years. He doesn't have a crystal ball, but he certainly has thoughts. So I think that's worth exploring.

推荐游戏 Recommended Games

Host

你推荐哪些游戏给人们,尤其是社交方面的?

What are games that you recommend to people, especially socially?

Noam Brown

我最近经常玩一个叫《血染钟楼》的游戏。

I've been playing a lot of a game called Blood on the Clock Tower lately.

Host

那是什么?

What is it?

Noam Brown

有点像《黑手党》或《狼人杀》。它在旧金山变得非常流行。

It's kind of like Mafia or Werewolf. It's become very popular in San Francisco.

Host

哦,就是我们在你家玩的那个。是的。好吧,得搞一个。有点好笑,因为我和几个人聊过,他们告诉我以前扑克是风投和科技创始人社交的方式。现在正转向《血染钟楼》。这是湾区人们用来联系的东西。有人告诉我一家初创公司举办了一场《血染钟楼》的招聘活动。哇。所以我想它真的很流行。这是个有趣的游戏,而且玩它输的钱比扑克少。对不擅长这些事的人来说更好。

Oh, that's the one we played at your house. Yeah. Okay. Got to get it. It's kind of funny because I was talking to a couple people who told me that it used to be that poker was the way VCs and tech founders socialized. Now it's shifting more towards Blood on the Clock Tower. That's the thing people use to connect in the Bay Area. I was told a startup held a recruiting event that was a Blood on the Clock Tower game. Wow. So I guess it's really catching on. It's a fun game, and you lose less money playing it than poker. Better for people not very good at these things.

Noam Brown

我觉得这是个奇怪的招聘活动,但确实是个有趣的游戏。

I think it's kind of a weird recruiting event, but certainly a fun game.

Host

在这里获胜者有什么特质值得招聘?这就是问题所在。善于说谎、欺骗和识别欺骗。那是最好的员工吗?我不知道。

What qualities make a winner here that is interesting to hire for? That's the thing. Good ability to lie, deception, and picking up on deception. Is that the best employee? I don't know.

AI 在不完美信息游戏中的应用 AI for Imperfect Information Games

Host

我最后的小话题是《万智牌》。我们讨论过国际象棋、围棋,它们有完美信息。而扑克在不完美信息下,宇宙相当有限,只有 52 张牌。然后这些其他游戏有不完美信息,但可能选项池巨大。那有多难?难度如何缩放?

My slight final pet topic is Magic: The Gathering. We've talked about chess, Go, which have perfect information. Then poker has imperfect information in a pretty limited universe. Only a 52-card deck. Then these other games have imperfect information with a huge pool of possible options. How much harder is that? How does the difficulty scale?

Noam Brown

我很高兴你问这个,因为我对不完美信息游戏的 AI 有大量知识。这是我长期的研究领域。我不常有机会谈论它。我们为无限注德州扑克制作了超人类 AI。有趣的是,隐藏信息的数量实际上相当有限,因为你只有两张底牌。在单挑中,你可能处于的状态数是 1,326,乘以其他玩家数量。这仍然不是一个巨大的数字。这些 AI 模型的工作方式是枚举你可能处于的所有不同状态。对于六人桌扑克,有五个其他玩家,5 * 1,326 种状态。你为每个状态分配概率,输入神经网络,然后为每个状态得到动作。问题是,随着隐藏可能性的数量增加,这种方法就会失效。

I love that you asked that because I have this huge store of knowledge on AI for imperfect information games. This is my area of research for so long. I don't get to talk about it very often. We've made superhuman poker AIs for no-limit Texas hold 'em. One interesting thing is that the amount of hidden information is actually pretty limited because you have two hidden cards. The number of possible states you could be in is 1,326 when playing heads-up, multiplied by the number of other players. It's still not a massive number. The way these AI models work is you enumerate all the different states you could be in. For six-handed poker, there are five other players, 5 * 1,326 states. You assign a probability to each, feed those into your neural net, and get actions back for each state. The problem is that as you scale the number of hidden possibilities, that approach breaks down.

不完美信息游戏中的未解问题 Unanswered questions in imperfect-information games

Noam Brown

还有一个非常有趣但尚未解答的问题:当隐藏状态的数量变得极其庞大时,该怎么办?比如在奥马哈扑克中,你有四张隐藏牌,可以用一些启发式方法来减少状态数量,但这仍然是个难题。再比如游戏《 Stratego 》,有 40 个棋子,状态数接近 40 的阶乘,我们之前在扑克中用的方法就行不通了,需要新的方法。这方面有很多活跃的研究。对于《万智牌》这样的游戏,我们在扑克中用的技术不能直接套用,这仍然是一个有趣的研究问题。不过,这只是在做我们在扑克中用的那种搜索技术时才会成为问题。如果只是做无模型强化学习,那就不是问题。我猜如果有人愿意投入精力,现在可能已经能做出一个超人类的《万智牌》机器人了。这个领域还有一些未解的研究问题。但它们是最重要的吗?我倾向于说不是。我们在扑克中用的搜索技术相当有限。即使扩展这些技术,也许能让它们适用于《 Stratego 》和《万智牌》,但它们仍然有限。它们不会让你在 Codeforces 上用语言模型达到超人类水平。所以我认为更有价值的是专注于非常通用的推理技术。随着我们改进这些技术,总有一天我们会有一个模型,开箱即用就能在《万智牌》中达到超人类水平。那才是更重要、更令人印象深刻的研究方向。

And there's still this very interesting unanswered question of what do you do when the number of hidden states becomes extremely large. So if you go to Omaha poker where you have four hidden cards, there are heuristic things you could do to reduce the number of states, but it's still a very difficult question. And then if you go to a game like Stratego where you have 40 pieces, so there are close to 40 factorial different states, then all these existing approaches we used for poker kind of break down and you need different approaches. There's a lot of active research on how to cope with that. For something like Magic: The Gathering, the techniques we used in poker would not out of the box work. It's still an interesting research question of what to do. Now, this becomes a problem when you're doing the kinds of search techniques we used in poker. If you're just doing model-free RL, it's not a problem. My guess is that if somebody put in the effort, they could probably make a superhuman bot for Magic: The Gathering now. There are still some unanswered research questions in that space. But are they the most important? I'm inclined to say no. The techniques we used in poker for search were pretty limited. If you expand those techniques, maybe you get them to work on things like Stratego and Magic: The Gathering, but they're still going to be limited. They're not going to get you superhuman in Codeforces with language models. So I think it's more valuable to focus on the very general reasoning techniques. One day as we improve those, I think we'll have a model that just out of the box plays Magic: The Gathering at a superhuman level. That's the more important and more impressive research direction.

Host

太棒了。非常感谢你来做客,Noah。

Cool. Amazing. Yeah. Thanks so much for coming on, Noah.

Noam Brown

谢谢你的时间。谢谢邀请。

Yeah. Thanks for your time. Yeah. Thanks. Thanks for having me.

互动版:逐字朗读 + 针对本期提问 →