从扑克到外交:Noam Brown 谈 AI 与博弈论

From Poker to Diplomacy: Noam Brown on AI and Game Theory

诺姆·布朗 Noam Brown · No Priors 播客 · 2023-04-25 · 约 61 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Noam Brown 讲述他从金融到 AI 研究的历程,专注于博弈论智能体以及塑造他工作的关键时刻。

Noam Brown discusses his journey from finance to AI research, focusing on game-theoretic agents and the pivotal moments that shaped his work.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 28)

全文 · Full transcript(中英对照)

0. 引言与诺姆背景 Introduction and Noam's background

Host

Noam,欢迎来到 No Priors。

Noam, welcome to No Priors.

Noam Brown

谢谢邀请。

Well, thank you for having me.

Host

感谢你加入,Noam。如今很多人想到 AI,就是输入几个词,然后得到一张图片,或者让 ChatGPT 用猫的声音写一首押韵诗总结詹姆斯·伯纳姆的专业管理阶级。而你推动的方向非常有趣,在很多方面与主流不同,你更关注博弈论中的智能体如何与人类及彼此互动。同时,正如 Sarah 提到的,你也是真正推动 AI 边界的 Tenex 工程师和研究者之一。所以我很好奇,最初是什么激发了你对游戏以及研究 AI 击败扑克和外交这类游戏的兴趣?

Well, thanks Noam for joining. So, you know, I think in the world today when a lot of people think about AI, they think about it as basically you put a couple words into a prompt and then you get out an image or you have ChatGPT summarize James Burnham's professional managerial class for you in a rhyming essay in the voice of a cat or something. And I think you've pushed in really interesting directions that are very different in some ways from what a lot of people have focused on and you've been more focused on game theoretic actors interacting with humans and with each other. And in parallel, you're kind of known as as Sarah mentioned as sort of one of these true Tenex engineers and researchers pushing the boundaries on the in AI. And so I'm sort of curious like what first sparked your interest in games and researching AI to defeat games like poker and diplomacy.

Noam Brown

我的经历有点非传统。我本科快结束时以及毕业后,在算法交易领域工作了几年,虽然有趣刺激,像游戏一样每天有得分(赚或亏的钱),但并不是我人生中最想做的事。于是我决定做研究,但不确定方向。最初打算学经济学,所以去了美联储工作了两年。说实话,我想弄清楚如何更好地构建金融市场以鼓励亲社会行为。在这个过程中,我对博弈论产生了兴趣。我想读经济学博士,专攻博弈论。但有两件事改变了我的想法:首先,我对经济学的进展速度感到失望,因为有了想法需要通过立法,过程漫长;而计算机科学更令人兴奋,你可以直接构建东西,不需要许可。其次,我发现博弈论最激动人心的工作其实发生在计算机科学领域,而不是经济学。所以我申请了研究生院,打算在计算机系学习算法博弈论。入学后,正好有位教授在找人做扑克 AI 研究。我觉得这完美结合了我感兴趣的一切:博弈论、动手构建、AI。我高中和大学时玩过扑克,虽然从不玩高额局,但对策略很感兴趣。本科时我甚至尝试做过扑克机器人,但表现很糟,不过过程很有趣。所以能在研究生阶段做这个研究,我觉得太完美了。而且我感觉这是个机会,因为看起来可行,我意识到如果成功做出能玩扑克的 AI,过程中会学到非常有价值的东西,可能对未来产生重大影响。这就是我开始的缘由。

Well, I think my journey is a bit non-traditional. I mean, I started out in finance actually towards the end of my undergrad career and also like after right after undergrad, I worked in algorithmic trading for a couple years and I kind of realized that while it's fun and exciting, it's kind of like a game, you know, you get a score at the end of the day which is how much money you made or lost. It's not really the most fulfilling thing that I would want to do with my life. And so I decided that I wanted to do research and it wasn't really clear to me in what area. I was originally planning to do economics actually. And so I went to the Federal Reserve, I worked there for 2 years. Honestly, I wanted to figure out how to structure financial markets better to encourage more prosocial behavior. And so in the process, I became interested in game theory. You know, I thought I wanted to pursue a PhD like in economics that focused on game theory. Two things happened. So first of all, I became a bit jaded with the pace of progress in economics because if you come up with an idea, you have to get it past your legislation and it's a very long process. Computer science is much more exciting in that way because you can just build something. You don't really need permission to do it. And then the other thing I figured out was that a lot of the most exciting work in game theory was actually happening in computer science. It wasn't happening in economics. And so I applied for grad schools with the intention of studying algorithmic game theory in a computer science department. And when I got to grad school, there was conveniently a professor that was looking for somebody to do research on AI for poker. And I thought this was like the perfect intersection of everything that I wanted to do. I was interested in game theory. I was interested in making something, interested in AI. I had played poker when I was in high school and college and you know, never for high stakes, but always just kind of interested in the strategy of the game. I actually tried to make a poker bot when I was an undergrad and it did terribly, but it was a lot of fun along the way. And so to be able to do that for research in grad school, I thought this was like the perfect thing for me to work on. And also I felt like there was an opportunity here because it felt doable and I recognized that if you succeed in making an AI that can play poker you're going to learn really valuable things along the way and that could have like major implications for the future. So that's kind of how I got started in that.

Host

这真的很酷。你刚开始时有没有一个具体的最终目标,还是纯粹出于兴趣?换句话说,你经常听到领域里的人说“我们的最终目标是 AGI,一直如此”。但我觉得有时这是后来编造的有趣故事。你是把这当作基础研究和个人兴趣,还是认为有一条通往代表人类行动的智能体的路径,或者有其他驱动力?

That's really cool. And did you have a specific end goal of your work when you started there or was it just interest? In other words, some you know, you talked to a lot of people in the field and they say, "Oh, our end goal is AGI and it's always been." And I think sometimes that's sort of invented later as sort of an interesting story for what they're doing. Did you view this as just doing primary research and it's just personal interest? Did you view it as like there's a path leading to agents that function on behalf of people or or was there some other sort of driving motivator?

Noam Brown

我 2012 年开始读研,那时情况大不相同。AGI 的想法真的是科幻。有些人认真对待,但很少。主流观点是 AI,如果说的话,是个死领域。我记得给一位教授发邮件说:“我对 AI 很感兴趣,但担心读博后找不到工作。”幸运的是,几年后情况剧变,我恰好在正确的时间出现在正确的地方。所以最初目标不是 AGI,而是学习 AI 和博弈论的有趣知识,慢慢构建。直到读研几年后,才明显看到进展速度非常惊人。

Well, so I started grad school in 2012 and it was a very different time in 2012. You know, the idea of AGI was really science fiction. There were some people that were serious about it, but very few. The majority opinion was that AI was, if anything, kind of a dead field. I actually remember emailing a professor and having this conversation where I was like, "Look, I'm really interested in AI, but I'm kind of worried to pursue a PhD in this because I get the impression that it's just a dead field and I'm worried I'll be able to get a job afterwards." Conveniently, like a couple years into grad school, things changed pretty drastically and I happened to be in the right place at the right time, I think. I was really fortunate in that respect. So the original intention wasn't to pursue AGI. The original intention was you learn interesting things about AI and game theory and you build slowly and it was really only a couple years into grad school that it became clear that the pace of progress was quite dramatic.

Host

有没有某个具体时刻让你真正意识到这一点?我知道有些人提到 AlexNet 的出现,或者早期的 GAN 工作像警钟。我好奇是否有特定的技术、论文或其他东西,还是只是一个连续的过程?

Was there a specific moment that really drove that home for you? I know for some people they mention oh, AlexNet came out or oh, you know, some of the early GAN work felt like a wake-up call. I'm just sort of curious if there's a specific technology or paper or something else that came out or was it just kind of a continuum?

Noam Brown

我觉得是慢慢积累的。对我个人来说,AlphaGo 时刻特别关键。看到那个,一切都很清楚。AlexNet 也是。我记得读研前上过一门 AI 课,计算机视觉课,他们讲 SIFT 之类的。然后 AlexNet 出现,把那些都推翻了,效果惊人。

I think it was a slow drip. I mean, I think for me especially, it was the AlphaGo moment. You know, like when you see that, it's just very clear. I mean, AlexNet too. I remember taking an AI class before I started grad school actually. I took a computer vision class and they were talking about SIFT and all this stuff. You get something like AlexNet and it just throws all that out the window and it's just mind-boggling how effective that could be.

Host

Noam,你能解释一下为什么 AlphaGo 如此重要,比如搜索空间的大小,以及它与之前游戏的区别吗?

Noam, can you explain actually like why AlphaGo is so important and like just size of search space and how you might contrast that to previous games?

Noam Brown

是的,AI 的一个重要里程碑是 1997 年深蓝在国际象棋中击败加里·卡斯帕罗夫。那是个大事。我觉得今天很多机器学习研究者低估了它,但我们从中收获很多。我们学到了规模确实有效。那次不是扩展神经网络训练,而是扩展搜索。但深蓝使用的技术在围棋这样的游戏中行不通,因为模式匹配不存在。

Yeah, so there was a big milestone in AI was Deep Blue beating Garry Kasparov in chess in 1997. And that was a big deal. And I think we actually it's kind of downplayed today I think in like by a lot of machine learning researchers, but it was we learned a lot from that. We learned that scale really does work. And in that case, it wasn't scaling training in neural nets, it was scaling search. But the techniques that were used in Deep Blue they didn't work in a game like Go because the pattern matching was just not there.

1. AlphaGo 与模式匹配 AlphaGo and pattern matching

Host

围棋的一大挑战是如何评估棋盘局势,判断谁占优。在国际象棋中,你可以手动编写一个函数来估算,比如每个棋子值多少分。但在围棋中,这几乎不可能手工完成——它太庞大、太微妙、太复杂了。如果你问一个人类棋手谁占优,他们能告诉你结果,但说不出原因。人们曾以为人类更擅长模式匹配,所以当 AI 证明它能在模式匹配上超越人类,即使是在一个受限的游戏中,这意义重大。这对很多人来说都是一个警钟。我记得作为一个前围棋迷,我试图理解 AlphaGo 的走法来提高自己的水平——那真是令人震惊。如果您的听众还没看过 AlphaGo 的纪录片,我强烈推荐;它在 Netflix 或 YouTube 上都有。你可以从中看到这对世界有多大的影响。

A big challenge in Go was figuring out how to evaluate the state of a board, how to tell who's winning. In chess, you can handcraft a function to estimate that, like each piece is worth a certain number of points. In Go, that's almost impossible to do by hand—it's too big, too subtle, too complicated. If you ask a human who's winning, they can tell you but not why. People assumed humans are better at pattern matching, so having an AI demonstrate it could do pattern matching better than humans, even in a constrained game, was a big deal. It was a wake-up call to a lot of people. I remember as a former Go nerd trying to understand AlphaGo's moves to play better—it was mind-blowing. If any of your listeners haven't seen the AlphaGo documentary, I highly recommend it; it's on Netflix or YouTube. You can see how significant it was to the world.

Host

你后来为什么选择外交作为扑克之后的下一个研究方向?游戏种类很多,你的选择标准是什么?

How did you end up choosing diplomacy as the next thing to work on after poker? There's a wide space of games. What drove your selection criteria?

Noam Brown

在扑克取得成功后,我们试图选择下一个方向。AI 发展得非常快——比很多人意识到的要快得多。当时有很多关于下一个基准应该是什么的讨论。人们提到了像花火、狼人杀或卡坦岛这样的游戏。但那是 2019 年,GPT-2 刚刚问世,令人震惊。DeepMind 在星际争霸 II 中达到了大师水平,OpenAI 在 Dota II 中击败了人类专家。所以像卡坦岛这样的游戏感觉太简单了——你可以找五个人花一年时间就能攻克它。我们想要一些真正令人印象深刻的东西,需要全新的技术,而不仅仅是扩展现有技术。我们最终选择了外交。让 AI 用自然语言与人类谈判并制定策略的想法,即使在 2019 年也感觉像科幻小说。这就是我们瞄准它的原因。我认为这是正确的决定。说实话我有点害怕这么做——风险很高,但所有研究都应该是高风险高回报的。

After we succeeded in poker, we were trying to pick the next direction. AI was progressing very quickly—much quicker than people appreciated. There were conversations about what the next benchmark should be. People threw around games like Hanabi, Werewolf, or Settlers of Catan. But this was 2019, and GPT-2 had just come out, which was mind-blowing. DeepMind had grandmasters in StarCraft II, OpenAI beat human experts in Dota II. So a game like Settlers of Catan felt too easy—you could take a team of five people, spend a year, and crack it. We wanted something truly impressive that required fundamentally new techniques, not just scaling up existing ones. We landed on diplomacy. The idea of an AI that negotiates in natural language with humans and strategizes with them felt like science fiction, even in 2019. That's why we aimed for it. I think it was the right call. I was a little afraid to do that—it's high risk, but all research should be high risk, high reward.

Host

在研究外交的过程中,关于 Cicero 的能力,最出乎意料的是什么?

What was the most unexpected thing to come out of working on diplomacy in terms of what Cicero could do?

Noam Brown

最出乎意料的是它居然没有被检测出是机器人。我们非常担心这一点。我们无法提前测试——我们可以和机器人玩,但我们知道它是机器人。我们无法召集一群人让他们在不被告知的情况下玩。如果人们知道他们在和机器人玩,他们的行为会非常不同。我们不想让这变成图灵测试。所以我们不得不让机器人进入玩家不知道有机器人的游戏。这是获得有意义结果的唯一方法。外交是一个自然语言谈判游戏,所以有复杂的长对话。作为机器人很难蒙混过关。我们最大的担忧是,在五局游戏内,甚至可能两局,他们就会识破,消息传开,后续的游戏就毫无意义了。我们想也许能侥幸玩 10 局才被发现。但令人惊讶的是,我们成功完成了整整 40 局游戏而未被发现。这让我很惊讶。我认为这证明了语言模型的进步,也说明也许人类并不像我们想象的那么擅长交谈。如果有人说了奇怪的话,他们的第一反应不是'我在和机器人说话',而是'这个人很笨、走神了或者喝醉了'。机器人这个选项排在很后面。所以在这方面我们很幸运,但这也显示了语言模型的质量。我想 Meta 计划发布这些数据,那将会非常有趣。

The most unexpected thing was honestly how it didn't get detected as a bot. We were really worried about this. We couldn't test it ahead of time—we could play with the bot, but we know it's a bot. We couldn't gather people and have them play without telling them. If people know they're playing with a bot, they behave very differently. We didn't want this to turn into a Turing test. So we had to enter the bot into games where players didn't know there was a bot. That was the only way to get meaningful results. Diplomacy is a natural language negotiation game, so you have complicated long conversations. It's hard to get away with being a bot. Our big concern was that within five games, maybe even two, they'd figure it out, word would get out, and future games would be meaningless. We figured maybe we'd get 10 games before detection. But surprisingly, we managed to go the full 40 games without being detected. That was surprising to me. I think it's a testament to the progress of language models, and also that maybe humans aren't as good at talking as we think. If someone says something weird, their first instinct isn't 'I'm talking to a bot' but 'this person is dumb or distracted or drunk.' Being a bot is way down the list. So we got lucky in that respect, but it also shows the quality of the language model. I think Meta is planning to release the data, which will be very interesting.

2. 有趣的机器人互动 Interesting Bot Interaction

Host

但你能描述一下在这些谈判中你觉得机器人有趣的一次互动吗?

But can you just describe an interaction from the bot you thought was interesting in these negotiations?

Noam Brown

哦,是的。有一条信息真的让我有点害怕,当时它正在和另一个玩家交谈,那个玩家说:‘嘿,我对你靠近边境的部队很紧张。’而机器人实际上并不打算攻击那个玩家,它打算往另一个方向走。它给玩家发了一条非常有同理心的信息:‘听着,我完全理解你的感受。我可以 100%向你保证,我不打算攻击你。我打算往另一个方向走。我向你保证。’这真的感觉像一条非常人性化的信息。我 100%没想到会来自机器人。当你看到这样的东西时,你会意识到这里面有非常强大的东西。

Oh, yeah. One of the messages that was really honestly kind of scary to me was when it was talking to another player, and the player was saying, 'Hey, I'm really nervous about your units near the border.' And the bot honestly was not planning to attack the player; it was planning to go in the other direction. It sent the player this really empathetic message: 'Look, I totally understand where you're coming from. I can assure you 100% I'm not planning to attack you. I'm planning to go the other direction. You have my word.' It really felt like a very human-like message. 100% I would have never expected that from a bot. When you see stuff like that, it makes you appreciate that there's something really powerful here.

3. 图灵测试相关性 Turing Test Relevance

Host

在这种情况下,你怎么看待图灵测试?你对这个测试是否仍然相关,或者如何思考它,有什么更新的看法?

How do you think about the Turing test in the context of all this? What's your updated model of whether the test is still relevant or how to think about it?

Noam Brown

实际上,《纽约时报》今天发表了一篇 Cade Metz 的文章,关于图灵测试及其意义,他在文章中提到了 Cicero。他的观点是图灵测试基本上已经过时了。我有点同意这个看法。我认为图灵测试不再像最初设想的那样是一个有用的衡量标准。当然,仅仅因为我们有能够——我不会说它们能通过图灵测试,但它们已经足够接近,以至于这个测试不再那么有用了。这并不意味着我们有了通用智能。我认为在这方面还有很长的路要走。这些机器人有很多事情做不好。但是的,我现在认为图灵测试不再是一个有用的衡量标准了。这并不一定意味着它一直没用;它只是显示了我们已经取得了多大的进步。我们还没有 100%达到,但进步是惊人的,尤其是在过去几年。

There was actually a New York Times article that came out today from Cade Metz on the Turing test and what it means, and he talks about Cicero in the article. His view is that the Turing test is kind of dead. I kind of agree with that. I think the Turing test is no longer really a useful measure the way it was intended to be. Certainly just because we have bots that can—I won't say they can pass the Turing test, but they're getting close enough that it's no longer that useful of a measure. It doesn't mean that we have general intelligence. I think there's still a long way to go on that. There's a lot of things that these bots can't do well. But yeah, my view now is that the Turing test is not that useful of a measure anymore. It doesn't necessarily mean it was always useless; it just shows how much progress we've made. We're not 100% there, but the progress has been staggering, especially in the past few years.

4. 通用智能的衡量标准 Measures for General Intelligence

Host

你认为使用什么衡量标准是有意义的?另外,你认为在通往通用智能的道路上缺少什么?

What measure or measures do you think make sense to use? And also, what do you think is missing on the road to general intelligence?

Noam Brown

我认为缺少几样东西。我特别感兴趣的一大块是推理能力。这些机器人都在做下一个词预测,对吧?Cicero 实际上有点不同,它根据计划来生成对话。我认为这是 Cicero 区别于当今语言模型许多工作的一个非常有趣的地方。但很多研究都在使用下一个词预测,当它试图做更复杂的推理时,它大量使用思维链,只是展开它在训练数据中观察到的人类推理方式,然后看看结果如何。我认为 AI 研究人员普遍认识到这是当前机器人的一大弱点,如果我们想要真正的通用人工智能,那么这个问题需要解决。如何解决这个问题是一个大问题,这就是为什么我真的很喜欢这个方向,因为它仍然是一个开放问题。已经取得了一些进展,但我认为还有很大的改进空间。

I think there are a few things missing. The big thing I'm interested in particular is reasoning capabilities. You have these bots and they're all doing next-token prediction, right? Cicero is a bit different actually in that it's conditioning its dialogue generation on a plan. I think that's one of the really interesting things that distinguishes Cicero from a lot of the work happening in language models today. But a lot of research is using next-token prediction, and when it tries to do something more sophisticated in terms of reasoning capabilities, it's a lot of chain of thought where it just rolls out the kind of reasoning it has observed humans do in its training data, and sees where that leads. I think there's a general recognition among AI researchers that this is a big weakness in the bots today, and that if we want truly general artificial general intelligence, then this needs to be addressed. There's a big question about how to address it, and that's why I really like this direction because it's still an open question. There's been some progress, but I think there's a lot of room for improvement.

5. 推理的有前景方向 Promising Directions for Reasoning

Host

你认为最有希望的潜在方向是什么?

What do you think are the most promising possible directions?

Noam Brown

这是一个万亿美元的问题。我认为有一些明确的基础。首先,思维链确实是一个很大的进步,考虑到它是一个多么简单的想法,它的效果令人震惊。我每天早上醒来都告诉自己:‘现在让我们一步一步地思考。’对于那些不知道的人来说,你在提示中加入类似‘让我们一步一步地思考这个问题’的内容,然后 AI 会生成一个更长的思考过程,说明它是如何得出结论的,这会导致更好的结论。但你可以把这看作是展开它在人类数据中观察到的思考过程。所以有一个问题:与其只是展开它,能不能在它每一步的过程中实际改进它?我保持非常抽象,因为这是一个重要的问题,而且还没有明确的答案。我不想过多猜测,但我认为这个方向有改进的空间。

That is the trillion-dollar question. I think there are clear bases. First of all, chain of thought really was a big step, and it's kind of shocking how effective that was given how simple of an idea it is. I tell myself every day when I wake up, 'Now let's think step by step.' For those who don't know, you add to the prompts something like 'Let's think this through step by step,' and then the AI will generate a longer thought process about how it reaches its conclusion, and that leads to better conclusions. But you can see that as just rolling out the thought process it observed in human data. So there's a question: instead of just rolling that out, could you actually improve it as it goes through each step? I'm keeping it very abstract because it's an important question, and there's not a clear answer yet. I don't want to speculate too much, but I think there is room for improvement in this direction.

6. 数据与规模化 Data and Scaling

Host

这里训练所需的实际数据集是什么?也许退一步说,我一直在与人讨论数据问题,以及我们什么时候会用完容易获得的数据,什么时候必须开始创建大规模合成数据或人类 RLHF 数据,或者付钱让人们整天录制自己。随着这些模型的扩展,你会用尽互联网和视频内容。我很好奇你在这个背景下如何看待数据,以及从自我驱动智能体的角度来看,需要什么才能将事情提升到下一个水平。

What was the actual data set necessary for the training here? And maybe to take a step back, I've been having conversations about data and when we run out of easily available data, and when we have to start creating large-scale synthetic data or human RLHF data, or pay bounties for people to record themselves all day. As these models scale, you use up the internet and video content. I'm curious how you thought about data in this context and what's necessary to take things to the next level from a self-driven agent perspective.

Noam Brown

目前还不清楚数据是否真的是性能的瓶颈。我和 AI 研究人员讨论过这个问题,我认为这方面的担忧没有人们想象的那么大。部分原因是存在的数据比人们意识到的要多得多,而且我认为随着研究的进展,样本效率会得到提高。我们将能够更充分地利用数据。我认为瓶颈将是 Scaling(规模扩张)。看看今天存在的模型,它们训练成本可能高达 5000 万美元。你可能很容易将其提高 10 倍。如果未来一两年内出现一个 5 亿美元训练的模型,我不会感到惊讶。

It's not clear that data really is the bottleneck on performance here. I've talked to AI researchers about this, and I think there isn't as much worry about this as people might think. Partly because there's a lot more data out there than people realize, and also because I think there will be improvements to sample efficiency as research progresses. We'll be able to stretch the data more. I think the bottleneck is going to be scaling. Look at the models that exist today; they probably cost $50 million to train. You can probably easily 10x that. I wouldn't be surprised if there's a $500 million model trained in the next year or two.

7. 规模化极限与推理时计算 Scaling Limits and Inference-Time Compute

Noam Brown

你甚至可以再上一个数量级,训练一个 50 亿美元的模型,如果你是美国政府或者一家非常大的科技公司的话。但除此之外呢?你训练一个 1 亿美元的模型?你可能会看到一些改进,但到了某个点,这就不再现实了。所以这将是瓶颈。我们可能还能再扩大两个数量级,然后就会遇到大问题。人们专注于如何提高效率,如何更便宜、更并行地训练。但你只能从中榨取这么多,而且我认为我们已经榨取了很多了。这就是为什么我对推理方向感兴趣,因为我认为还有另一个维度人们目前没有在扩展,那就是推理时的算力。你可以提前花 5000 万美元训练这个模型,比如预训练,然后到了实际推理时,成本只有一分钱。如果它不是在 1 秒内返回答案,而是在 1 小时甚至 5 秒或 10 秒内返回呢?有时候,如果人们想给出更好的答案,他们会坐下来思考一会儿,这会带来更好的结果。我认为这正是这些模型所缺少的。所以这是克服扩展挑战的一种方式,也是我对此感兴趣的部分原因。

You can maybe even go another order of magnitude and train a $5 billion model if you're the US government or a really big tech company. But what do you do beyond that? Do you train a $100 million model? You'll probably see some improvement, but at some point it just becomes not realistic anymore. So that's going to be the bottleneck. We maybe get two orders of magnitude more scaling, and then we have a big problem. People are focused on making this more efficient, training cheaper and more parallelized. But you can only squeeze so much out of that, and I think we've squeezed a lot already. This is why I'm interested in the reasoning direction, because I think there's this whole other dimension that people are not scaling right now, which is the amount of compute at inference time. You can spend $50 million training this model ahead of time, like pre-training, and then when it comes to actual inference, it costs like a penny. What happens if instead of returning an answer in a second, it returns an answer in an hour or even 5 or 10 seconds? Sometimes if people want to give a better answer, they'll sit and think a bit, and that leads to a better outcome. I think that's one of the things missing from these models. So that's one way to overcome the scaling challenge, and partly why I'm interested in working on that.

8. 外交数据与自我对弈 Diplomacy Data and Self-Play

Host

回到 Alod 说的,外交问题本身并没有互联网规模的数据,对吧?这是一个相对较小的社区。你能谈谈你们在自我对弈和实际使用的数据方面做了什么吗?

Going back to what Alod said, the Diplomacy problem specifically didn't have internet-scale data, right? It's a relatively small community. Can you talk about what you guys did in terms of self-play and the data that actually was involved?

Noam Brown

外交问题很有趣,因为实际上并没有大量数据。我们有一个相对不错的数据集,大约有 5 万局游戏对话,来自一个叫 webdiplomacy.net 的网站,这个网站已经存在了近 20 年,人们在那里随意玩外交游戏。我们非常幸运能得到这个数据集。我搜遍了互联网,试图找到所有有可用数据的网站,这基本上是唯一一个有有意义数据量的网站。还有另一个流行的网站,但他们定期删除数据,这让我难以置信。你坐在一座金矿上,却因为占用服务器空间而删除它。我猜他们没有意识到 AI 研究人员有一天会感兴趣。其他网站则拒绝交出数据。所以我很高兴我们与 webdiplomacy.net 达成了协议,否则这个项目根本不会发生。那大约是 5 万局外交游戏,约 1300 万条消息,这是一个规模不错的数据集,但不足以从头训练一个机器人。幸运的是,我们能够利用互联网上更广泛的数据集,所以我们有一个预训练的语言模型,然后在外交数据上进行微调,得到了一个在游戏中能很好交流的机器人。这有助于对话,但还有一个问题:策略水平不够。部分原因是仅靠监督学习无法做到那么好——你无法仅靠监督学习在这些游戏中学会非常好的策略——还因为玩这些游戏的人水平不高。数据集中的大部分来自相当弱的玩家。这就是现实;你有一个钟形曲线。真正的强玩家在任何数据集中都只占很小一部分。而且这并不限于外交。我们在国际象棋和围棋中也发现了这一点。我们做了这个实验:如果你在人类国际象棋和围棋的巨量数据集上做纯监督学习,得到的机器人并不是专家级的棋手。即使它被调节得像国际象棋大师一样,它也无法达到那种表现,因为它没有做任何规划。这正是缺失的东西。所以为了获得超越平均人类表现甚至强人类表现、达到更好水平的策略,我们必须进行自我对弈。这就是所有以前游戏 AI 的训练方式——比如 AlphaGo,尤其是 AlphaZero,Dota 2 机器人。它们通过与自己进行数百万或数十亿次对局来训练。这也是我们的扑克机器人在两人和六人扑克中的训练方式。当你从这些游戏转到外交时,区别在于合作方面。你不能假设其他所有人都会像机器一样与你行为一致。所以为了克服这一点,我们必须将自我对弈与对人类行为方式的认知结合起来,即人类的行为会很像我们的数据所表明的那样。利用我们的数据集,我们建立了一个人类行为的大致模型,然后通过自我对弈进行改进。我们找到了一个好的策略,但这是一个与人类玩法兼容的策略。为了给你一些直观感受:如果你从头开始训练一个机器人,没有人类数据,它可能会学会谈判,但可能会用一种胡言乱语的机器人语言。然后当你把它放在一个与六个人类一起的游戏中时,它将无法与他们交流,他们都会互相合作,而不是与机器人合作。

So Diplomacy is interesting because there's actually not a ton of data out there. We had a relatively good dataset of about 50,000 games of dialogue from a site called webdiplomacy.net, which has been around for almost 20 years where people play Diplomacy casually. We were very lucky to get this dataset. I was scouring the internet trying to find all the sites with available data, and this was basically the only one with a meaningful amount. There was another popular site, but they periodically deleted their data, which was mind-boggling to me. You're sitting on a gold mine, and you're just deleting it because it's taking up server space. I guess they didn't appreciate that AI researchers would one day be interested. Other sites just refused to hand over their data. So I'm really glad we managed to work out a deal with webdiplomacy.net, because otherwise the project would never have happened. That's about 50,000 games of Diplomacy, about 13 million messages, which is a good size dataset, but not enough to train a bot from scratch. Fortunately, we were able to leverage a wider dataset from the internet, so we have a pre-trained language model, and then we fine-tune it on the Diplomacy data, and we get a bot that can communicate pretty well in the game. That helps with dialogue, but there's still a problem: the strategy isn't going to be up to par. Partly because you can't do that well with just supervised learning—you can't learn a really good strategy in these kinds of games with just supervised learning—and also because the people playing these games are not very good. Most of the dataset is from fairly weak players. That's just a reality; you have a bell curve. The actual strong players are a relatively small fraction of any dataset. And this is not limited to Diplomacy. We also found in chess and Go. We ran this experiment: if you do pure supervised learning on a giant dataset of human chess and Go games, the bot you get is not an expert chess or Go player. Even if it's conditioned to behave like a chess grandmaster, it won't match that performance because it's not doing any planning. That's really what's missing. So to get a strategy that goes beyond average human performance or even strong human performance to something much better, we had to do self-play. This is how all previous game AIs have been trained—like AlphaGo, especially AlphaZero, the Dota 2 bot. They're trained by playing against themselves for millions or billions of trajectories. That's also how our poker bot was trained for two-player and six-player poker. The difference when you go from those games to Diplomacy is the cooperative aspect. You can't assume everyone else will behave like a machine identically to you. So to overcome that, we had to combine self-play with a recognition that humans will behave a lot like our data suggests. Using our dataset, we built a rough model of how humans behave, and then we improved on that using self-play. We figured out a good strategy, but one that's compatible with how humans play. To give some intuition: if you train a bot from scratch with no human data, it could learn to negotiate, but it might learn in a gibberish robot language. Then when you put it in a game with six humans, it won't be able to communicate with them, and they'll all work with each other instead of with the bot.

9. AI 中的非语言交流与人类规范 Non-verbal communication and human norms in AI

Noam Brown

嗯,同样的动态也发生在策略游戏中,游戏中的走法,非语言交流方面。比如,机器人会发展出这些规范和期望,关于它的盟友这回合应该做什么。比如,我要支援我的盟友进入这个区域,因为我期望他们进入这个区域。我甚至不需要跟他们谈这个,因为这太明显了,他们应该这么做。但人类有自己的元游戏,比如,‘哦,我实际上很明显应该支援你进入这个区域。’如果你不理解人类的规范和惯例,那么你就无法与人类很好地合作。他们就不会跟你合作,而是去找别人。所以,这就是我们在 Cicero 中真正需要克服的,我们通过使用人类数据来构建一个人类行为模型,然后在此基础上加入自我对弈,作为对人类数据集的一种修正,从而做到了这一点。

Um that same dynamic happens even in the strategy game, the moves in the game, the non-verbal communication aspect. Like the bot will develop these norms and expectations around what its ally should be doing this turn. Like I'm going to support my ally into this territory because I'm expecting them to go into this territory. I don't even have to talk to them about this because it's just so obvious that they should be doing this. But the humans have their own meta game where like, 'Oh, it's actually really obvious that I should be supporting you into this territory.' If you don't understand the human norms and conventions, then you're not going to be able to cooperate well with humans. And they're just going to not work with you and work with somebody else instead. So that's what we really had to overcome in Cicero, and we managed to do that by using the human data to build this model of how humans behave, and then adding self-play on top of that as kind of a modifier to the human data set.

Host

这实际上有一些非常有趣的启示,对吧?比如,如果你相信长期来看,我们将会有在现实世界中采取行动、与人类互动的机器人,而人类在生命游戏中可能并不擅长最优玩法。你与他们互动。这恰恰凸显了推理相对于学习模式识别可能有多重要。

That actually has some really interesting implications, right? Like if you believe in the long term, we are going to have bots that take action in the real world interacting with humans, and humans are perhaps not very good at optimal play in the game of life. And you're interacting with them. Like, it sort of just brings home the point of how important reasoning could be versus learning pattern recognition.

Noam Brown

我认为你完全正确,如果你想制造与人类在现实世界中互动的 AI,这一点非常重要,对吧?比如,如果你有一辆在路上行驶的汽车,一辆自动驾驶汽车,你不想让它假设所有其他司机都是机器,每一步都会完美地最优行动。你希望自动驾驶汽车认识到这些其他司机是人类,而人类会犯错,有人可能会突然变道到我的车道,嗯,还有,就像日常互动,理解人类的非语言线索及其含义。这些都是 AI 必须能够应对的事情,如果它要在现实世界中真正对人类有用,而不仅仅是在国际象棋上打败他们。

I think you're absolutely right that this matters a lot if you want to make AIs that interact with humans in the real world, right? Like if you have a car driving on the road, a self-driving car, you don't want it to assume that all the other drivers are machines that are going to act perfectly optimally every step of the way. Like you want the self-driving car to recognize that these other drivers are humans, and humans make mistakes, and somebody could like swerve into my lane, and yeah, and also like, just day-to-day interactions, understanding the non-verbal cues of humans and what that means. These are things that an AI has to be able to cope with if it's going to be really useful to humans in the real world and not just beating them at chess.

10. 游戏作为 AI 进展的基准 Games as benchmarks for AI progress

Host

游戏已经被用作衡量 AI 进展的一种方式有一段时间了,你研究过扑克变体和外交变体,你也提到了其他人在国际象棋、围棋等方面的工作。你认为在游戏以及从 AI 角度对它们的研究方面,下一个前沿是什么?

Games have been used for a while now as a way to measure AI progress, and you've worked on poker variants and diplomacy variants, and you mentioned it for other work people have done in terms of chess and go and things like that. What do you think is the next frontier in terms of games and sort of research on them in the lens of AI?

Noam Brown

游戏作为 AI 基准测试有着悠久的历史。这可以追溯到 20 世纪 50 年代 AI 的奠基时期。比如国际象棋,它被当作 AI 的一个巨大挑战,因为如果我们能制造一个像人类国际象棋大师一样聪明的 AI,那么我可以想象它能做的所有其他聪明事。当然,这后来被证明是一种虚假的承诺,对吧?你得到了一个会下国际象棋的 AI,结果它实际上什么别的都不会。但在这个过程中我们学到了很多。游戏作为基准测试很有用,因为你可以非常客观地与人类顶尖表现进行比较。当你在这个领域超越人类能力时,就变得非常明显,即使这是一个受限的领域。你还有这个在 AI 研究人员出现之前就已经存在的基准测试。比如,AI 研究人员,一旦他们创造了技术,就很容易提出一个基准测试。你知道,你提出一个技术,然后你说,‘好吧,现在很容易提出一个这个技术适用的基准测试。’而你不希望这样。你希望问题先行。而游戏给了你这个。但我认为我们现在已经到了一个点,单个的娱乐游戏不再那么有趣的挑战了。我想,你知道,我之前说过我们选择外交是因为我们认为这是最难为 AI 制作的游戏。我认为这是真的。我想不出还有其他什么游戏,如果有人制造了一个能玩那个游戏的 AI,我会说,‘哇,这太令人印象深刻了,我没想到这是可能的。’所以,我认为未来,这个领域需要超越单个游戏,开始首先超越游戏,同时也要关注通用性。我们在外交中使用的方法与我们之前在扑克中做的以及其他人在国际象棋、围棋和星际争霸中做的非常不同。现在有一个问题,‘好吧,如果我们真的想要一个通用系统,一个通用 AI,我们能否让它以超人类水平玩所有这些游戏,同时还能做像图像生成、问答以及所有这类任务?’如果我们能做到这一点,那将变得非常令人印象深刻。所以,我认为游戏将继续作为这个基准,但它不是作为研究过度拟合的基准,我的希望是它将作为一个基准,我们与游戏之外的其它基准一起使用,比如图像生成基准和语言问答基准这类东西。

There's a long history of games as benchmarks for AI. This goes all the way back to the very foundations of AI back in the '50s. Like chess in particular was held up as this grand challenge for AI because if we can make an AI that was as smart as a human chess grandmaster, then I can imagine all the other smart things it could do. Of course that turned out to be kind of a false promise, right? Like you get an AI that plays chess, and it turns out it doesn't really do anything else. But we've learned a lot along the way. And games are useful as a benchmark because you can compare very objectively to top human performance. Like it becomes very clear when you're surpassing human ability in this domain, even if it's a restricted domain. You also have this benchmark that's existed before the AI researchers came along. Like AI researchers, it's really easy for them to come up with a benchmark once they have the technique already created. You know, you come up with a technique, and then you're saying like, 'Okay, well now it's really easy to come up with a benchmark that this technique will work for.' And you don't want that. You want the problem to come first. And games give you that. But I think we're reaching a point now where individual recreational games are just no longer that interesting of a challenge. I think, you know, I said earlier we chose diplomacy because we thought it would be the hardest game to make an AI for. And I think that's true. I can't think of any other game out there where if somebody made an AI that could play that game, I would be like, 'Wow, that's super impressive, and I did not think that that was possible.' And so, I think going forward, the field needs to move beyond looking at individual games and starting to look at first of all, going beyond games, but also looking at generality. The approach that we've used in diplomacy is very different from what we previously did in poker and what others have done in chess and go and StarCraft. And now there's a question of like, 'Okay, well, if we really want a general system, a general AI, can we have it play all of these games at a superhuman level and also able to do things like, you know, image generation and question and answering and all these tasks?' And if we could accomplish that, then that becomes incredibly impressive. And so, I think games will continue to serve as this benchmark, but it's not instead of serving as a benchmark that the research kind of overfits to, my hope is that it will serve as a benchmark that we use along other benchmarks outside of games like, you know, image generation benchmarks and language Q&A benchmarks and these kinds of things.

11. 人类仍占主导的领域 Domains where humans still dominate

Host

你认为你刚才描述的事情,双人游戏现在多人谈判合作游戏,似乎很明显,如果你能打败外交,你就能打败大多数游戏。鉴于 AI 已经在这些以特定方式具有挑战性的受限领域获胜,你如何看待那些将由人类技能主导的领域?会有这样的领域吗?

Do you think the thing that you just described, that two-player games now multiplayer games of negotiation, cooperation, it seems clear that if you can beat diplomacy, you can beat most games. And given that the AI has already won in these restricted domains that are challenging in specific ways, like how do you think about the domains that are going to be human skill dominant? Like are there going to be domains that like that?

Noam Brown

嗯,当然是物理世界中的任何事情。你知道,操作任务这类事情。机器人技术确实落后。我正因为这个原因尽量避免做任何物理世界的事情。软件要好处理得多。仍然有一些事情,即使在受限领域,人类也肯定更擅长。你看看写小说这样的事情。我认为你目前还不能让 AI 写出下一部《哈利·波特》。这可能不会太遥远。也许大约五年左右。但我认为现在还没有发生。有点可怕的是,我真的很难想出那些我会说,‘哦,是的,AI 在这方面无法超越人类’的领域。就像人们经常谈论人类永远有优势的领域,仅仅因为他们是人类,他们想对未来感觉良好。

Well, certainly anything in the physical world. You know, manipulation tasks, these kinds of things. Robotics is really lacking behind. I'm trying to avoid doing anything in the physical world for that reason. Software is just so much nicer to work with. There's still things that you can't that humans are definitely better at even in restricted domains. You look at something like writing a novel. I don't think you can get an AI to output like the next Harry Potter just yet. That might not be that far off. Maybe it's like 5 years away or something. But I don't think it's happening just yet. It's kind of scary that I'm really struggling to come up with domains where I'm like, 'Oh, yeah, AI is not going to be able to surpass humans in this.' Like people often talk about areas where humans will always have an advantage just because they're humans, and they want to feel good about the future.

12. 通用性与样本效率 Generality and Sample Efficiency

Host

但那种通用性是不是被夸大了?因为我觉得在你提到的例子里,你说从图像生成到外交,都在同一个架构或 AI 里。而且通常来看,如果一个人在某方面很擅长,他们往往不是样样精通,对吧?所以我感觉,我们在 AI 通用性上设定的标准,有时候比对人设定的标准更高。这么说不对吗?

But isn't that generality overstated? Because I feel like in the examples that you mentioned, you said everything from like image gen to diplomacy in like a single architecture or AI or something. And often it seems like, you know, if you look at the average person, if they're very good at one thing, they're usually not good at everything, right? And so, I kind of feel like the bar that we're using in terms of generality for AI sometimes is higher than the bar we'd use for generality for people in some sense. Or is that not a true statement?

Noam Brown

我觉得这不只是通用性的问题,更关键的是样本效率。比如,人类要成为优秀的棋手、外交玩家或艺术家,需要多少局游戏?答案是比 AI 少几个数量级。这在数据稀缺的领域会是个问题。当然,这个问题可能被克服,但我的意思是它目前还没被克服。我认为这是人类相对于当前 AI 的一个明显优势。

I think it's not just about generality, it's really about sample efficiency. Like how many games does it take for an AI for a human to become a good chess player or a good diplomacy player or a good artist? The answer is orders of magnitude less than it takes for an AI. And that is going to pose a problem when you're in domains that don't have that where there isn't much data. Now, that seems like a problem that could be overcome. It's I'm just saying that's a problem that hasn't been overcome yet. And I think that that's one of the clear advantages that humans have over AIs today.

13. 金融系统中的 AI AI in Financial Systems

Host

你认为我们什么时候会看到 AI 在金融或经济系统中出现?显然我们已经有了算法交易之类的东西。还有像加密货币这样的东西,实际上是用程序化的方式把货币包装成代码,对吧?而且通过智能合约,我们可以相当丰富地与之交互。你觉得在近期内,人们会在这方面进行实验,或者有关于机器人与金融系统实际交互的有趣研究吗?

When do you think we'll see the emergence of AIs in financial or economic systems? And obviously we have like algorithmic trading and other things like that. And then we have things like crypto where you effectively have programmatic approaches to effectively money wrapped as code, right? And the ability to interact with those things in reasonably rich ways through smart contracts. You know, do you think we're there's any sort of near-term horizon of people experimenting with that or just interesting research being done in terms of the actual interaction of a bot with a financial system?

Noam Brown

我认为这已经在发生了。看看金融市场,我敢肯定有很多交易是由深度学习驱动的。我实际上和很多金融公司聊过这个。我以前在金融行业工作过,而且很多金融公司喜欢扑克。所以我在各种场合做过关于 AI 扑克的演讲。我也和一些地方聊过,强化学习是否真的对金融市场和交易有用?我得到的答案通常是“不”。我认为用强化学习做交易的主要挑战在于它是一个非平稳环境。你可以有所有历史数据,但它不是一个平稳系统。市场会对世界事件等做出反应。所以,理想情况下你需要一种真正理解世界的技术,而不是把一切都当成黑箱。

I think it's already being done. If you look at financial markets, I'm sure there's tons of trading powered by deep learning. I've actually talked to a lot of finance companies about this. There's a lot of I used to work in finance and also like a lot of finance companies love poker. And so I've given a few talks at like various places on AI for poker. And I've talked to a few places about like is reinforcement learning actually useful for financial markets, for trading? And the answer I get is usually no. I think the major challenge with using things like reinforcement learning for trading is that it's a non-stationary environment. So, you can have all this historical data, but it's not a stationary system. And the markets respond to world events, these kinds of things. So, you need a technique ideally that really understands the world, not just treating everything like a black box.

Host

但这和你之前说的在推理上投入更多算力而非训练,也就是在决策时融入实时信号,有联系吗?还是你指的是别的,比如模型架构能让你随时间以某种方式更新权重之类的?

But could that at all feed into what you were saying about spending more compute on inference versus training, in other words, incorporating real-time signals at the point of decision-making? Or did you mean something else by that in terms of model architecture that would enable you to update weights in certain ways or things like that over time?

Noam Brown

嗯,我觉得这又回到了样本效率问题。人类很擅长适应新情况,而在金融市场中你经常会遇到新情况。我认为这也是一个通用性问题,你需要对世界有足够多的了解才能真正成功。话虽如此,我认为 AI 在金融市场上的成功是相当有限的。当然,如果你想拆分大订单之类的,它有用。另外,我得说我不是这方面的专家。我的知识有点过时了,因为我相信有很多前沿的东西正在发生,但人们不会告诉我,因为那些能赚钱。但我可以告诉你,这大概是 5 年前的观点。AI 在有限的方式下被使用,但我认为它还没有完全取代人类。

Well, I think it goes back to the sample efficiency problem. Humans are pretty good at adapting to novel situations and you run into these novel situations pretty frequently in financial markets. I think it's also a problem of generality that you need to understand so much about the world to really succeed. Now, that said, I think that the AIs are successful in financial markets in fairly limited ways. Certainly if you want to break up big orders, these kinds of things. Also, I should say I'm not an expert in this. This is kind of outdated knowledge from me, because I'm sure there's a lot of cutting-edge stuff that's happening that people are not telling me about because it's making money. But I can tell you this is kind of the perspective as of maybe 5 years ago. It's being used in limited ways, but I don't think it's fully replacing humans yet.

14. AI 与人类谈判 AI Negotiation with Humans

Host

你认为我们很快会有能与人类谈判的机器人吗?我先说我们最终会有的。你觉得时间线或用例是什么?

Do you think we're going to get bots that negotiate with humans soon? Let me preface that as we are eventually going to get them. What do you think the timeline is or the use case?

Noam Brown

这似乎是可行的。这取决于领域的约束程度。我认为如果看约束性强的领域,某些谈判任务,AI 今天可能已经比人类做得更好了。我在想具体的例子,比如如果你想就商品价格进行谈判,在很多情况下它可能比人类做得更好。我认为像薪资谈判这样的东西,它也可能比人类做得更好。这取决于你需要对世界了解多少。我认为合同谈判之类的仍然困难,因为每份合同都有很多微妙之处和细微差别。它暂时还不会取代专业谈判人员。但那些更受约束、不需要太多外部世界知识的事情,我认为 AI 可能已经能够胜任了。

That seems doable. It depends on how constrained the domain is. I think if you were to look at constrained domains, certain negotiation tasks, I think that AIs could probably do better than humans in that today. I mean, I'm trying to think of specific examples, but things like if you wanted to negotiate over the price of a good, it could probably do better than a human in a lot of those situations. I think if there's things like salary negotiations, it might do better than humans at that also. I think it depends on how much you need to know about the world. And I think contract negotiations, for example, would still be difficult because there's so much subtlety, there's so much nuance to every contract. And it's not going to replace a professional negotiator for that kind of task just yet. But kind of the things that are more constrained, don't require as much outside knowledge about the world, I think AIs are probably up to the task already.

15. 下一研究重点:推理 Next Research Focus: Reasoning

Host

我有个以前和你共事的朋友说,你特别擅长的一点是,你倾向于选择一个被忽视但潜力巨大的研究领域,长期投入,然后成为最顶尖的人。而世界上很多人会被光鲜的东西吸引,被流行趋势分心,结果发现那些研究没那么有趣。你接下来想做什么?或者你对下一波研究浪潮有什么兴趣?

So, a friend of mine who used to work with you says that one of the things you're really exceptional at is you tend to pick a neglected research domain with lots of promise. You commit to it long-term and then you become the best at it. And many people in the world kind of get attracted to shiny things instead and kind of distracted by, you know, whatever is in vogue, but then it turns out to be less interesting research. What are you thinking about working on next or what interests you as sort of the next wave of stuff to do?

Noam Brown

我主要感兴趣的是推理问题。这源于我在游戏领域的经验。看看像 AlphaZero、最新版 AlphaGo 这样的东西。我认为 AlphaGo 尤其被视为深度学习的一个重大里程碑。在某种程度上确实如此。没有深度学习它做不到,但仅靠深度学习也不行。如果你去掉 AlphaGo 中的规划,只使用原始的策略网络、原始的神经网络,它的表现实际上远低于人类顶尖水平。而仅靠原始神经网络,我们已经有了很多非常强大的东西。

I think the big thing I'm interested in is the reasoning problem. And this is kind of motivated by my experience in these game spaces. You look at things like AlphaZero, the latest version of AlphaGo. And I think that's held up like AlphaGo in particular is held up as this big milestone in deep learning. And to some extent it is. Like it was not doable without deep learning, but it wasn't deep learning alone that enabled that. If you take out the planning that's being done in AlphaGo and just use the raw policy network, the raw neural network, it's actually substantially below top human performance. And with just like raw neural nets, we have all these things that are incredibly powerful.

16. 领域特定规划算法的局限 Limitations of domain-specific planning algorithms

Noam Brown

比如聊天机器人、图像生成软件,但原始神经网络本身仍然无法下围棋。它需要额外的规划算法才能达到人类顶尖水平。AlphaGo 使用的规划算法——蒙特卡洛树搜索——非常领域特定。我认为人们没有意识到它有多领域特定,因为它在国际象棋和围棋中有效,而这些是人们研究这类技术的经典领域。它在扑克中无效,在外交中无效。由于我曾在这些领域工作过,我认识到这是这类算法的一个主要弱点。所以我认为有一个大问题:我们如何让这些模型用更通用的系统完成复杂的推理和规划任务,使其能跨多种领域工作?如果你能做到这一点,如果你能成功完成这个任务,那就能实现很多非常强大的东西。我想到的一个领域是定理证明。在我看来,如果你能真正通用地解决推理问题,那么在五年内拥有一个能证明黎曼猜想的模型并非疯狂。也许推理成本巨大——生成那个证明可能每 token 花费一百万美元——但如果能实现,那完全值得。也许你还能用它做其他事情,比如写出下一部获奖小说,或者发现救命药物。

Like, you know, chatbots, image generation software, but the raw neural net itself still can't play Go. It requires this extra planning algorithm on top of it to achieve top human performance. And that planning algorithm used in AlphaGo, Monte Carlo tree search, is very domain-specific. I think people don't appreciate just how domain-specific it is because it works in chess, it works in Go, and these have been the classic domains people have cared about for investigating these techniques. It doesn't work in poker, it doesn't work in diplomacy. Because I've worked in those domains, I recognize that this is a major weakness of these algorithms. So I think there's a big question: how do we get these models to do complex reasoning and planning tasks with a more general system that works across a wide variety of domains? And if you can enable that, if you can succeed in that task, then it enables a lot of really powerful things. One domain I'm thinking about is theorem proving. It doesn't seem crazy to me that you could have a model that can prove the Riemann hypothesis within the next 5 years, if you can solve the reasoning problem in a truly general way. Maybe the inference cost is huge—maybe it costs a million dollars per token to generate that proof—but that seems totally worth it if you can pull it off. And maybe you can do other things with it, like write the next prize-winning novel, or come up with life-saving drugs.

Host

补充一下背景,黎曼猜想被认为是数学中最重要的未解问题,其中第一组解已被验证,但我们还不确定。是的,我认为我真正感兴趣的关键是通用性。我们可以用领域特定的方法解决这个问题,但那样总会过拟合到那个领域。所以我认为我们需要像 Transformer 那样通用的东西,你把它扔给任何问题,它都出奇地好用。我猜你的意思是,有办法将问题框架化,从而取得更通用的进展,这可能围绕数学或代码。这样理解对吗?

Just for context, the Riemann hypothesis is considered the most important unsolved problem in math, where the first set of solutions have been checked, but we don't know for sure yet. Yeah, and I think the key that I'm really interested in is generality. We can approach this problem in domain-specific ways, but then it always ends up overfit to that domain. So I think what we need is something as general as what we're seeing with transformers, where you just throw it at any problem and it works surprisingly well. And I guess you're implying that there are ways to frame the problem to make progress that are more general, and that could be around math or possibly code. Is that the right understanding?

Noam Brown

我希望这些技术是通用的。我认为同时研究多种领域也很重要,以防止过拟合。另一个我认为很适合的领域是代码生成,因为我认为要写出好代码,下一个词预测能带你走得很远。但我不认为它能完全取代大公司的工程师。

My hope is that the techniques are general. I think it's important to also look at a wide variety of domains to prevent overfitting. One domain that I think would also be a good fit is code generation, because I think to write good code, next-token prediction is going to get you surprisingly far. But I don't think it's going to get you all the way to replacing engineers at big companies.

Host

也许给听众一个背景:Copilot 很棒,对吧?但我们现在用代码生成做的事情非常局部、上下文特定。所以如果你想规划整个产品,现有技术似乎做不到。

Maybe one piece of context for listeners is that Copilot is amazing, right? But what we're doing with code generation today is very local context-specific. So if you want to plan out a whole product, that doesn't seem doable with existing technology.

Noam Brown

我认为很多人听到我这么说时的想法是:‘嗯,你只要扩展它——扩展模型规模,扩展训练规模,过去这总是有效的。’我喜欢举 AlphaGo 的例子。理论上你可以扩展训练,扩展模型容量,然后就不需要规划了。你只需运行这个强化学习算法很长时间,拥有一个非常大的网络,它最终会学会如何在围棋中击败人类专家。但问题是:你需要扩展多少才能达到它用蒙特卡洛树搜索实现的性能?如果你算一下,结果是 10 万倍。这些模型已经花费大约 5000 万美元。显然你无法将它们扩展 10 万倍。那么问题来了,你该怎么办?AlphaGo 的答案是:与其让所有计算都在训练时完成,不如让它在实际下棋时花大约 30 秒来思考下一步怎么走。这就把成本负担从必须预计算所有东西转移到了能够即时思考。这就是为什么我认为那条路似乎是缺失的一环。

I think the perspective of a lot of people when they hear me say this is, 'Well, you just scale it—scale up the models, scale up the training, and that's always worked in the past.' The example I like to give is AlphaGo. You could in theory scale up the training, scale up the model capacity, and you don't need planning then. You just run this reinforcement learning algorithm for a really long time, have this really big network, and it will eventually learn how to beat expert humans in Go. But there's a question: how much would you have to scale it up to match the performance it achieves with Monte Carlo tree search? If you crunch the numbers, it ends up being 100,000x. These models are already costing like $50 million. Clearly you're not going to be able to scale them by 100,000x. So then the question is, what do you do instead? The answer in AlphaGo is: instead of having all that computation during training, you also have it spend like 30 seconds to figure out what move to make next when it's actually playing the game. That shifts the cost burden from having to precompute everything to being able to think on the fly. That's why I think that avenue seems like the missing piece.

Host

一个很随机的问题:如果你看人脑,你有各种专门化的模块,具有非常特定的功能。你有视觉皮层用于视觉处理,你有不同的部分负责情绪。大脑中有特定区域,如果你切除它,就会移除某些情绪或其他能力。发生过一些事故,杆子穿过人的头部,切除了一个非常特定的地方,然后人活了下来。所以你看到这种通过切除特定模块而导致的非常特定的功能丧失。为什么认为应该有一个通用架构,而不是你有一堆子模型一起运行,共同实现广泛的行为,这实际上就是我们在大脑中看到的?

A really random question: if you look at the human brain, you have these various specialized modules with very specific functions. You have the visual cortex for visual processing, you have different things for emotion. There are specific parts of the brain that if you ablate, you remove certain emotive or other capabilities. There have been accidents where poles have gone through people's heads and ablated a very specific place, and then people have survived. So you see this sort of very specific ablation of function through the ablation of specific modules. Why is it the correct assumption to think that there should be a generalized architecture versus you just have a bunch of sub-models that are all running together that collectively enable a wide range of behavior, which is effectively what we see in the brain?

Noam Brown

这是个好问题。我不认为我们需要拘泥于某种特定技术。答案可能是我们需要更专门化的系统,而不是只有一个真正通用的架构。我想的更多是目标而非方法。我们想要的是能在多种领域成功的东西。为每个领域提出独特的方法能让你走一段路,但我认为最终这会被真正通用的东西所取代。

That's a good question. I don't think we need to be tied to a specific technique. The answer might be that we need to have more specialized systems instead of just one truly general architecture. I think what I'm thinking about is more the goal rather than the approach. We want something that's able to succeed across a wide variety of domains. Having to come up with a unique approach to every single domain gets you part of the way, but I think that eventually that will be superseded by something that is truly general.

Host

是的,有道理。我想,一个大的领域就是推理,对吧?所以我并不是说推理的不同子类型需要不同的方法,而是可能有一些真正大的东西,其根本运作方式可能非常不同。

Yeah, that makes sense. And I guess, you know, one big domain is just reasoning, right? So I didn't mean to imply that different subtypes of reasoning will require different approaches, but more there may be really big things that fundamentally may function in a very different way.

17. 研究中的局部最优风险 Risk of local minima in research

Host

再说一次,这可能不对,对吧?大脑是一个进化系统,这意味着它在来源和形成方式上有巨大的局限性。而且你在进化一个系统时,常常会陷入这些局部最优,对吧?我只是有点好奇你是怎么想的。

And again, that may be incorrect, right? The brain is an evolved system, which means it has enormous limitations in terms of where it came from and how it got created. And you often end up with these local maxima when you evolve a system, right? I was just sort of curious about how you thought about that.

Noam Brown

是的,研究中总是存在陷入局部极小值的风险。而且这很棘手,人们会过度拟合它。我认为机器学习本身就是一个例子。比如深度学习,当时没多少人关注,因为他们觉得那是个死胡同。只有少数人在加拿大的荒野里研究这个。结果它取得了巨大的成功。所以多样性是有价值的,方法的多样性也有价值。我认为跳出框框思考,尝试做一些和别人不一样的事情确实有帮助。

Yeah, there's certainly a risk always with research that you could end up in a local minimum. And it's like hard, people will overfit to that. And I think actually machine learning was an example of this. Like deep learning, not many people were focused on this because they kind of assumed it was a dead end. There were only a few people out in the Canadian wilderness that were working on this. And that ended up being tremendously successful. And so there's value in diversity. There's value in a diversity of approaches. And I think it does help to try to think outside the box and try to do something that's a little bit different than what everybody else is doing.

18. 给研究者的建议:敢于冒险 Advice for researchers: take risks

Host

Noam,你将要从事这个非常有趣的领域。我相信还有其他你觉得有趣的问题,特别是考虑到我们愿意在规模上再扩大一两个数量级所花的钱有实际限制。你认为其他研究人员或团队应该关注哪些他们目前不够重视的问题?

Noam, you are going to go work on this really interesting area. I'm sure there are other problems you think are interesting, especially given the practical limits of how much money we're willing to spend on scaling up beyond another magnitude or two. What do you think other researchers or teams should be working on that they're not paying enough attention to?

Noam Brown

嗯,我认为我们现在在 AI 领域处于一个有趣的位置,鉴于目前的状况,有很大的机会来构建产品。已经有机会构建能对世界产生重大影响的产品。很高兴看到有人朝这个方向努力,试图将研究带入现实世界并产生巨大影响,让人们的生活更美好。顺便说一句,Alot 和我都收到了多封邮件,说他们正在构建价格谈判智能体。就像我说的,我认为这是可行的。所以我认为这是正确的选择。在研究方面,仍然有很多有趣的问题:我们如何让这些东西更高效?有没有更好的架构可以使用?我的意思是,我认为各个领域都有很多有趣的问题。我想给研究人员的主要建议不是关注哪个领域,而是研究风格。我认为有一种倾向是求稳,不愿冒大风险。我认为重要的是要认识到研究本质上是一个高风险领域。你知道,你正在做的事情长期来看很可能没有用。你必须接受这一点,并愿意承担风险。我的意思是,这发生在我身上。我博士早期研究从大局来看其实没什么用。长期来看,它没有产生我希望的影响。但这没关系,因为我有一件事最终产生了很大影响。所以我认为能够承担这些风险很重要。就像进入这个领域时,你要意识到从事研究本身就是在冒险。

Well, I think we're in an interesting place now in AI where there is a huge opportunity to build out products given where things are at now. There's already an opportunity to build out products that can have a big impact on the world. It's great to see that there are people going in that direction and trying to bring this research into the real world and have a big impact there, make people's lives better. For what it's worth, both Alot and I got emails from multiple people telling us that they're building price negotiation agents as we speak. Well, that's like I said, I think it's doable. So I think it's the right call. I think on the research side there's still a lot of interesting questions about how do we make these things more efficient? Are there better architectures we can use? I mean, I think there's just so many questions across the board that are interesting. I think the big thing I would recommend to researchers is not about which area to focus on, but just the style of research. I think there's a tendency to play it safe and to not take big risks. And I think it's important to recognize that research is an inherently risky field. You know, there's a high probability that what you're working on is not going to be useful in the long term. You have to kind of accept that and be willing to take that risk anyway. I mean, this happened to me. My early research in my PhD in the grand scheme of things really wasn't that useful. It didn't make as much impact in the long term as I would have hoped. And that's okay because I had one thing that ended up being quite impactful. And so I think it's important to be able to take those risks. Kind of like going into the field recognizing that you are taking a risk already by going into research.

19. 外交游戏概述 Overview of Diplomacy game

Host

你最先在这里听到的。像 Noam 一样,做让你紧张的事情。你能用一分钟快速介绍一下《外交》这个游戏,让人们了解它是什么以及为什么这项研究如此突破性吗?

You heard it here first. Be like Noam. Work on things that make you nervous. Do you want to give a quick minute overview of Diplomacy so people can understand what it is and why the research was such a breakthrough?

Noam Brown

是的,《外交》是 50 年代开发的一款游戏。实际上是由一个目睹了第一次世界大战的人开发的,他认为那是一场外交失败。所以他想创造这个游戏来教人们如何成为更好的外交官。游戏背景设定在第一次世界大战初期。有七个玩家势力可供选择:英国、法国、德国、意大利、俄罗斯、土耳其和奥匈帝国。每个回合你都要进行复杂的谈判。你的目标是尽可能控制地图上的区域。获胜的方式是控制地图的大部分。有点像《饥饿游戏》,尽管最终只有一个人能赢,但仍有合作的动机,尤其是在早期,因为合作双方都能受益,并且最终获胜的机会更大。所以你会进行非常复杂的谈判。所有交流都是私下进行的。所以不像《风险》或《卡坦岛》这类游戏,所有谈判都在众人面前进行;在《外交》中,你会把某人拉到一边,到角落里密谋这回合一起攻击谁,谁支持谁。然后与所有人谈判后,你写下这回合的行动。然后所有行动同时公布。你可以看到人们是否真的履行了帮助你的承诺,或者他们可能骗了你,这回合就要攻击你。所以它融合了《风险》、扑克和《幸存者》的元素,因为信任是核心。而这正是游戏的精髓:你能与他人建立信任吗?因为在这个游戏中成功的唯一途径就是合作,尽管你总是有动机去攻击别人并从中获利。所以,这就是游戏。它已经存在很久了,从 50 年代开始。它是肯尼迪和基辛格最喜欢的游戏。从 80 年代起就有人从 AI 角度研究这个游戏。但用自然语言与人类对战并击败他们的想法,直到几年前还完全是科幻小说。它仍然是科幻小说,但我们至少认为值得追求。研究真正起飞是在 2019 年,当时研究人员开始使用深度学习为这个游戏制作机器人,这些机器人可以玩非语言版本。也就是说没有交流,你只需写下行动,通过你的行动进行非语言沟通。我们当时在做这个研究。DeepMind 也在做。还有蒙特利尔大学和其他一些地方。有很多兴趣和进展,但我们决定冒风险直接跳到终点,而不是采取渐进式方法,直接瞄准全自然语言《外交》。我很高兴我们瞄准了那个目标。

Yeah, Diplomacy is this game that was developed in the 50s. It was actually developed by this guy who saw what happened in World War I and kind of viewed this as a diplomatic failure. And so he wanted to create this game that would teach people how to be better diplomats. So it takes place at the onset of World War I. There's seven player powers that you can play as: England, France, Germany, Italy, Russia, Turkey, and Austria-Hungary. And you engage in these complex negotiations every turn. Your goal is to try to control as much of the map as possible. The way you win is by controlling the majority of the map. It's kind of like Hunger Games where even though only one person can win at the end of the day, there's still this incentive to work together, especially early on, because you can both benefit and have a better chance of winning in the end if you work together. And so you have these really complex negotiations that happen. All the communication is done in private. So unlike a game like Risk or Settlers of Catan where all the negotiation is done in front of everybody else, in Diplomacy you will actually pull somebody aside, go into a corner, scheme about who you're going to attack together this turn, who's going to support who. And then after you've negotiated with everybody, you write down what your moves are for the turn. So then all the moves are read off at the same time. And you can see if people actually follow through on their promises about helping you, or maybe they lied to you and they're just going to attack you this turn. So it has some elements of Risk, poker, and Survivor, because there's this big trust component. And that's really the essence of the game: can you build trust with others? Because the only way to succeed in this game is by working together, even though you always have an incentive to attack somebody and grow at their expense. So yeah, that's the game. It's been around for a long time, since the 50s. It was JFK and Kissinger's favorite game. There's research for this game from an AI angle going back to the 80s. But the idea that you could play this game in natural language with humans and beat them was just complete science fiction until a few years ago. Like it was still science fiction, but we at least thought it was worth pursuing. And research really took off in 2019 when researchers started using deep learning to make bots for this game that could play the non-language version. So there's no communication, you just write down your moves, and you kind of have to communicate nonverbally through the actions that you take. We were doing research on this. DeepMind was doing research on this. And also University of Montreal and a couple other places as well. There was a lot of interest and progress, but we decided to take the risky bet of just jumping to the end point and instead of taking an incremental approach aiming for full natural language Diplomacy. And I'm glad that we aimed for that.

20. 被人类误认为人类的机器人 Bots that humans thought were human

Host

就是这样。你们所做的最令人惊叹的事情之一就是,你们基本上创造了让其他人——人类——以为是真人的机器人。

That's it. It seems like one of the pretty amazing things about what you all did is you basically created bots that other people that humans thought were other people.

21. 外交与合作 Diplomacy and Cooperation

Noam Brown

因此,它们必须学会如何相互协作,有时如何撒谎或欺骗,有时如何从博弈论的角度思考多步走法。所以这与下国际象棋或围棋截然不同,后者只是与对手对弈,然后几乎有一个概率性的走法树。你遇到了这些人类元素。你真的必须理解人类元素。而《外交》游戏除了自然语言成分之外,真正有趣的一点是,它是第一个在涉及合作的游戏中取得重大 AI 突破的游戏。这非常重要,因为归根结底,当我们制造这些下棋的 AI 时,我们并不是为了在游戏中击败人类。我们希望它们在现实世界中有用。如果你想让这些 AI 在现实世界中有用,那么它们也必须理解如何与人类合作。

And therefore they had to learn how to collaborate with each other, how to sometimes lie or deceive, how to sometimes think through multiple moves from a game theoretic perspective. And so it's a radically different thing than playing chess or Go against another person and then just having almost a probabilistic tree of moves or something. You run into these human elements. You really have to understand the human elements. And what's really interesting about Diplomacy aside from the natural language component is that it really is the first major game AI breakthrough in a game that involves cooperation. That's really important because at the end of the day when we make these AIs that play chess and Go, we're not developing them with the purpose of beating humans at games. We want them to be useful in the real world. And if you want these AIs to be useful in the real world, then they have to understand how to cooperate with humans as well.

Host

我和 Alot 曾讨论过人机协作玩法,以及考虑到我们已经接受 AI 会在游戏中获胜,这个概念是否还会存在。但我认为,AI 通过与人类合作来采取行动,这需要成为一项核心能力,这似乎是显而易见的。也许这只是我自我安慰的说法,但我希望与 AI 合作的能力仍然是一项非常重要的人类技能。

Alot and I were talking about centaur play and whether or not that would persist as an idea at all given that we've accepted that AIs are going to win games at this point. But I think the idea that AIs are going to take action by cooperating with humans, that needs to be a core capability seems obvious. And I am perhaps this is the making myself feel better story, but I am hopeful that that is a human skill that remains quite important, being able to cooperate with AIs.

Noam Brown

嗯,据我所知,人机协作玩法中,AI 在国际象棋等游戏中已经变得非常强大,以至于现在不太清楚人类是否真的增加了多少价值。我也是这么跟 Sara 说的。是的,我哭了。我理解。我接受。

Well, from what I hear, centaur play is like AIs have gotten so strong in games like chess that it's not clear if the human is really adding that much these days. That's what I told Sara, too. Yeah, I'm crying. I get it. I accept it.

Host

是的。我不认为人类在围棋这样的游戏中仍然有用,因为 AI 非常强大,但它们在每局游戏中也会偶尔犯一些非常奇怪的错误。而在《外交》游戏中,我认为除了 AI 之外,有一个经验丰富的人类是非常有帮助的。不过,最终我想象这些系统会变得如此强大,以至于就像国际象棋一样,人类只是在最后增加微小的差异。

Yeah. I don't think the humans are still useful in games like Go, because the AIs are super strong, but they will also sometimes make these really weird blunders a few times in each game. And in Diplomacy I think it's super helpful to have an experienced human in addition to the AI. Though eventually I'd imagine that these systems become so strong that it kind of goes the way of chess where the human is just adding a marginal difference at the end.

Noam Brown

是的,我其实在想,在人生这场游戏中,人类在人机协作中的窗口期能持续多久。没错。但没关系。我明白了。Alot 是对的。希望是永远。但谁知道呢?我们拭目以待。

Yeah, I'm actually just wondering how long that window is for humans in centaur play in the game of life. Right. But it's okay. I got it. Alot was right. Hopefully forever. But who knows? We'll see.

22. 扑克 AI 突破 Poker AI Breakthroughs

Host

那么,你介意解释一下你在扑克方面所做的工作以及你取得的一些突破吗?

So do you mind explaining the work that you've done in poker and some of the breakthroughs that you made there as well?

Noam Brown

是的,我的博士研究主要集中在如何让 AI 在无限注德州扑克中击败顶级人类玩家。特别是在博士期间,我专注于单挑无限注德州扑克,也就是两人扑克。这是一个长期存在的挑战问题。实际上,如果你回顾约翰·纳什关于博弈论的原始论文,文中讨论的唯一应用就是扑克。他实际上在论文中分析了一个简单的三人扑克游戏,并手动计算出了纳什均衡。然后在结尾他说:“哦,是的,用这种方法分析更复杂的扑克游戏会非常有趣。”所以我很高兴我们终于在 60 年后有机会做到这一点。有趣的是,我认为特别是在 AlphaGo 之后,这成了一个非常热门的问题,因为 AlphaGo 之后,有一个大问题:“好吧,AI 现在可以在国际象棋中击败人类,可以在围棋中击败人类。它们还有什么做不到的?”而它们做不到的一件大事就是推理隐藏信息,能够理解“好吧,这个对手知道一些我不知道的事情,而我知道一些他们不知道的事情。”并且能够在战略环境中克服这个问题,这是一个未解的重大问题。所以这基本上是我整个研究生阶段研究的重点。

Yeah, my PhD research was really focused on how to get an AI to beat top humans in the game of no limit Texas Hold'em poker. Specifically during my PhD it was on heads-up no limit Texas Hold'em poker. That's two-player poker. This was a long-standing challenge problem. Actually, if you go back to the original papers written on game theory by John Nash, the only application discussed in the paper is poker. He actually analyzes a simple three-player poker game in the paper and works out the Nash equilibrium by hand. And then at the end he says, 'Oh yeah, it'd be really interesting to analyze a much more complex poker game using this approach.' So I'm glad we finally got a chance to do that, 60 years later. It's interesting, I think especially after AlphaGo, this became a very popular problem because after AlphaGo, there was a big question: 'Okay, well, AIs can now beat humans at chess, they can beat humans at Go. What can't they do?' And the big thing that they couldn't do was reason about hidden information, be able to understand that 'Okay, this other player knows things that I don't know, and I know things that they don't know.' And being able to overcome that problem in a strategic setting was a big unanswered question. So that was the focus of my research from basically my whole grad school experience.

Noam Brown

有几个不同的研究实验室在从事这项工作。每年我们都会制作一个扑克机器人,并在一个名为“年度计算机扑克竞赛”的比赛中让它们相互对战。基本上,当我开始读博时,AI 在扑克方面已经取得了一些进展。所以这个比赛实际上变成了一场 Scaling(规模扩张)竞赛。在河牌圈,也就是德州扑克的最后一轮,大约有 25 亿种不同的手牌组合。我们做的是使用 K-means 聚类将这些手牌聚类,并将相似的手牌视为相同。这允许你计算一个扑克策略,因为现在你不必担心 25 亿种手牌并为每一种制定策略,你可以将它们分桶,现在你有大约 5000 个桶,然后你可以为这么多桶计算策略。这是在神经网络出现之前,所以我们使用 K-means 聚类而不是深度神经网络。但你可以把桶的数量想象成网络中的参数数量。所以在研究生阶段,这变成了一场 Scaling(规模扩张)竞赛。你的机器人能有多少个桶?第一年大约是 5000 个桶,然后我们增加到 30000 个,然后是 90000 个。每年我们都会拥有更大的模型,训练更长时间,并行化它们,它们总是能击败前一年的模型。2014 年,我们实际上赢得了年度计算机扑克竞赛,之后我们决定让我们的机器人与专家人类玩家对战。这是第一次“人脑对 AI 扑克竞赛”,我们邀请了四位顶级的单挑无限注扑克职业选手,让他们与我们的机器人打了 80000 手牌。结果机器人以相当大的差距输掉了。在这次比赛中,我意识到人类处理游戏的方式与我们的机器人非常不同。我们会在比赛前用 1000 个 CPU 训练机器人两个月。但到了实际比赛时,它会立即行动。而人类则不同。他们会提前练习,培养对游戏的直觉,但当他们与机器人对弈并遇到困难时,他们会坐在那里思考。

There were a few different research labs that were working on this. And what would happen is every year we would all make a poker bot, and we would play them against each other in this competition called the Annual Computer Poker Competition. Basically, when I started my PhD, there had already been some progress in AI for poker. And so the competition really turned into a competition of scaling. There are about 2.5 billion different hands that you could have on the river, the last round of poker in Texas Hold'em. And what we would do is cluster those hands together using K-means clustering and treat similar hands identically. That allows you to compute a policy for poker because now instead of having to worry about 2.5 billion hands and come up with a policy for each one, you can bucket them together, and now you have like 5,000 buckets, and you can actually compute a policy for that many buckets. This was before neural nets, that's why we were doing this K-means clustering instead of deep neural nets. But you can think of it as the number of buckets is like the number of parameters in your network. So in grad school, it turned into a competition of scaling. How many buckets could you have in your bot? First year it was like 5,000 buckets, then we got up to 30,000, then 90,000. Every year we would have bigger models, train them longer, parallelize them, and they would always beat the previous year's model. In 2014, we actually won the Annual Computer Poker Competition, and after that we decided to take our bot and play it against expert human players. This was the first Brains vs. AI Poker Competition, where we invited four top heads-up no-limit poker pros, and we had them play 80,000 hands of poker against our bot. And the bot actually lost by a pretty sizable margin. And it occurred to me during this competition that the way the humans were approaching the game was actually very different from how our bot was approaching it. We would train our bot for like 2 months leading up to this competition on 1,000 CPUs. But when it came time to actually play the game, it would act instantly. And the humans would do something different. They would practice ahead of time, develop an intuition for the game, but when they were playing the game against the bot and they were in a difficult spot, they would sit there and they would think.

23. 扑克 AI 中的搜索发现 Discovery of Search in Poker AI

Noam Brown

有时候思考 5 秒,有时候一分钟,但他们会思考,从而想出更好的解决方案。我突然意识到,这可能正是我们机器人所缺失的东西。所以比赛后我做了分析,想弄清楚:如果我们在实际对局中加入这种搜索、这种规划算法来制定更好的策略,能提升多少性能?答案是性能提升了大约 10 万倍。这相当于将模型规模、参数数量、训练量都扩展 10 万倍。而我博士前三年,只把规模提升了大约 100 倍。这已经很不错了,我为此自豪。但看到那个结果后,我意识到我博士期间所做的一切,与加入搜索并扩展搜索相比,都只是注脚。所以接下来一年,我几乎不停歇地工作,每周 100 小时,试图扩展搜索,在推理时投入尽可能多的算力。然后在 2017 年 1 月,我们又进行了一场比赛,对阵四位顶级扑克专家,奖金 20 万美元激励他们全力以赴。这次我们彻底击败了他们。扑克玩家们亲口说,他们没想到能以如此大的差距击败专家级扑克玩家。这就是我研究生阶段研究扑克 AI 的故事。

And sometimes it was like 5 seconds, sometimes it was like a minute, but they would think and that would allow them to come up with this better solution. And it occurred to me that this might be something that we're missing from our bot. So I did this analysis after the competition to figure out, if we were to add this search, this planning algorithm that would come up with a better strategy when it's actually in the hand, how much better could it do? And the answer was it improved the performance by about 100,000x. It was the equivalent of scaling the model, scaling the number of parameters, scaling the training by 100,000x. Now, the three years of my PhD at that point, I had managed to scale things by about 100x. And that's quite good. I was very proud of that. But when I saw that result, it made me appreciate that everything I had done in my PhD up until that point was just a footnote compared to adding search and scaling search. So for the next year, I just worked basically non-stop, like 100-hour weeks, trying to scale up search, make it as throw as much computation at the problem at inference time as possible. Then we did another competition in January 2017, where we played against four top expert poker players again. $200,000 in prize money to incentivize them to play their best. And this time we completely crushed them. People were literally telling us, poker players were literally telling us they did not think it was possible to beat expert poker players by that kind of margin. So that's the story of my grad school experience working on poker AI.

24. 多人扑克与可扩展搜索 Multiplayer Poker and Scalable Search

Noam Brown

那是针对两人扑克。之后我们转向多人扑克,六人扑克。同样,重大突破是我们开发了一种更可扩展的搜索技术。不再需要一直搜索到游戏结束,而只需向前搜索几步。非常有趣的是,我们进行了另一场比赛,机器人赢了,而训练这个机器人的成本如果在云计算服务上运行,不到 150 美元。我认为这表明这不仅仅是算力扩展的问题,而确实是算法上的突破。如果人们知道这种方法,这种结果 20 年前就能实现。

That was for two-player poker. We ended up after that working on multiplayer poker, on six-player poker. Again, the big breakthrough there was that we developed a more scalable search technique. So instead of always having to search to the end of the game, it could search just a couple moves ahead. And what was really interesting there is the bot we did another competition, the bot won, and that bot cost under $150 to train if you were to run it on a cloud computing service. I think that shows that this wasn't just a matter of scaling compute. It really was an algorithmic breakthrough. And this kind of result would have been doable 20 years ago if people knew the approach to take.

25. 对人类扑克策略的影响 Impact on Human Poker Strategy

Host

这在扑克方面是如何体现的?

How did that play out in terms of poker?

Noam Brown

好问题。作为最后时刻的添加,我们加入了这种能力。机器人的工作方式是,我们给它提供不同的下注大小。我们玩的游戏是 2 万筹码,盲注$100/$200,实际上是$50/$100。所以它可以下注从$100 到$20,000 的任何金额。能够同时下注$5,000 和$5,001 没有太大价值。所以我们离散化动作空间,限制它只考虑几个不同选项。问题在于:你给它哪些大小选择?在开发后期,我们有多余的算力,就加入了一些额外的大小,比如 4 倍底池、10 倍底池。这不会增加太多成本,为什么不给它这个选项呢?我原本以为它不会实际使用这些大小。但在比赛中,它大量使用了这些大小。有时它会向$100 的底池下注$20,000,这在职业扑克中闻所未闻。我一开始有点担心,以为这是个错误,对手玩家一开始也这么认为。但他们发现,自己总是陷入非常棘手的局面,在跟注或弃牌之间挣扎。这正是你打出好扑克的标志。如果你看到对手在决策中挣扎,说明你做对了。最后他们告诉我们:‘是的,这是我们想要融入自己打法的一点,加入这些所谓的超池下注。’以前典型策略是下注底池的四分之一到一倍之间,现在在职业扑克中,虽然不常见,但下注 5 倍、10 倍底池已成为策略的一部分。如果运用得当,这可以是非常强大的策略。很酷。我还想说,现在职业扑克玩家训练的方式,都使用机器人辅助。这很像国际象棋:你下棋,然后让机器人分析你的对局,看看是否犯了错误,哪里错了,下次如何改进。这个游戏已经被去神秘化,变得很像国际象棋。我把扑克描述为高维国际象棋。就像国际象棋,但你需要推理动作的概率分布,而不是离散动作。

That's a great question. So, as a last minute thing, we added this ability. The way the bot works, we give it different bet sizes that it can use. The game we were playing is 20,000 chips, $100/$200 blinds or $50/$100 blinds actually. So it can bet any amount from $100 up to $20,000. There's not much value in being able to bet both $5,000 and $5,001. So we would discretize that action space to constrain it to only considering a few different options. So there's a question of, 'What sizes do you give it the choice between?' Towards the end when we were developing this bot, we just had room for extra computation, so we threw in some extra sizes, like 4x the pot, 10x the pot. It doesn't cost that much more, so why not just give it the option? I didn't think it would actually use those sizes. Then during the competition, it ended up using those sizes a lot. It would sometimes bet $20,000 into a $100 pot, which was completely unheard of in professional poker play. I was a little worried about this because I thought it was a mistake at first, and I think the players we were playing against also thought it was a mistake at first. But then they found that they kept ending up in these really tricky situations, and they would just really struggle with whether to call or fold. That's how you know you're playing good poker. If you see the other person really struggling with the decision, that is a sign that you're doing something right. At the end they told us, 'Yeah, that's the one thing that we're going to try to incorporate into our own play, adding these what are called overbets into our strategy.' So instead of typically betting between a quarter of the pot and one times the pot, now in professional poker play, it's actually not common, but it is part of the strategy to bet sometimes 5x the pot, 10x the pot. If you can pull it off in the right way, it can be a very powerful strategy. It's pretty cool. I should also say the way professional poker players train now, they all use bots to assist them. It's a lot like chess where you play the game and then you have a bot analyze your play afterwards to see if you made mistakes, where you made mistakes, how you could do better next time. The game really has been demystified and become a lot like chess. I kind of describe poker as essentially high-dimensional chess. It's like chess where you have to reason about a probability distribution over actions instead of just discrete actions.

26. 最优玩法与纳什均衡 Optimal Play and Nash Equilibrium

Host

这很有趣,因为我认为以前人们并不真正相信扑克中存在完全最优玩法。他们理解概率分布,但如果你玩现场扑克,还有社交线索,对吧?社交玩法。这显然已经被扫除了。不是作为一种娱乐活动,而是作为真正获胜的策略。

It's really interesting because I don't think people really believed there was fully optimal play in poker before. They understood the probability distribution, but if you're playing live poker, there are social cues, right? And social play. That has clearly been swept out. Not as an activity of enjoyment, but in terms of a strategy that actually wins.

Noam Brown

是的,我认为很多人对此感到惊讶:存在一种最优的扑克玩法。有一种东西叫纳什均衡,如果你采用那种策略,你总是赢。嗯,它保证从长期来看,你在期望上不会输。原因是,如果你对阵另一个也采用纳什均衡的对手,显然你们不能都赢。要么一人输,要么平局。所以从期望上看,如果你们互相对阵,最终会打平。

Yeah, I think that's surprising to a lot of people, this idea that there is an optimal way to play poker. There's this thing called a Nash equilibrium, where if you're playing that strategy, you always win. Well, it guarantees that in the long run, you will not lose in expectation. The reason is that if you're playing against somebody else that's also playing the Nash equilibrium, obviously you can't both win. One of you is going to lose or you're going to tie. So in expectation, if you're playing against each other, you're going to end up tying.

27. 扑克中的纳什均衡 Nash Equilibrium in Poker

Host

但实际上,如果你在像扑克这样复杂的游戏中采用纳什均衡策略,对手会随着时间的推移犯一些小错误,而他们犯的每一个错误都会变成你口袋里的钱。所以你只需坚持纳什均衡,等待他们犯错,最终你就会赢。这现在已经成为扑克玩家的共识:从纳什均衡开始。如果你真的非常厉害,你可以观察其他玩家,看他们如何偏离纳什均衡、做出次优决策,然后你或许可以自己偏离均衡来利用这些错误。但真正稳妥的做法是坚持纳什均衡,让他们犯错,他们每犯一个错误就会损失金钱,而这些钱就进了你的口袋。我想我们的时间到了。非常感谢你参加我们的播客,Noam。

But in practice, what ends up happening is if you're playing the Nash equilibrium in a complicated game like poker, the other person is going to make these small mistakes over time, and every mistake that they make is money in your pocket. And so you just play the Nash equilibrium, wait for them to make mistakes, and you end up winning. And that is now the conventional wisdom among poker players, that you start by playing the Nash equilibrium. If you're really good, you can look at the other players, see how they're deviating from the Nash equilibrium, playing suboptimally, and maybe you can deviate yourself to capitalize on those mistakes, but really the safe thing to do is play the Nash equilibrium, let them make mistakes, and every mistake that they make costs them money, it puts money in your pocket. I think that's all we have time for. Thank you so much for joining us on the podcast, Noam.

Noam Brown

是的,非常感谢你的邀请。

Yeah, thank you very much for having me.

互动版:逐字朗读 + 针对本期提问 →