AlphaGo: The Match That Changed AI Forever
打开互动全文版(中英对照 + 朗读 + 问答)→十年前,AlphaGo 在围棋中击败李世石,标志着现代人工智能革命的开始。听突破背后的架构师讲述这个故事。
Ten years ago, AlphaGo defeated Lee Sedol in Go, marking the start of the modern AI revolution. Hear from the architects behind the breakthrough.
欢迎回到 Google DeepMind 播客。我是 Hannah Fry 教授。想象一下这个场景:2016 年 3 月,在韩国首尔的一家酒店套房里,两位选手正在对弈古老的围棋。这是一项复杂到难以想象的游戏,长期以来被认为机器无法掌握。一边是传奇的 18 次围棋世界冠军李世石,另一边是 AlphaGo——一个基于强化学习这一强大技术的神经网络 AI 系统。欢迎来到韩国首尔的 DeepMind 挑战赛现场。那是一个非常令人惊讶的落子。没有一位人类选手会选择第 37 手。经过 7 天数小时的激烈对弈,是的,那是一个激动人心的落子。李世石在棋盘上放下两颗棋子,示意最终认输。转眼之间,世界改变了。最终比分 4-1。祝贺 AlphaGo 和整个团队。那正好是十年前,此后 AI 领域发生了难以想象的变化。我们看到了大型语言模型的崛起、AI 智能体日益复杂,以及蛋白质折叠等科学重大挑战的解决。但在很多方面,现代 AI 革命可以说正是从韩国的那块木棋盘上开始的。因此,在本期节目中,我们想回顾过去并展望未来,看看这个教机器玩游戏的勇敢实验如何成为当今 AI 突破的基石。与我一起讲述这个故事的是最合适的嘉宾:Toby Graef,Google DeepMind 杰出研究科学家,他当时就在首尔,是 AlphaGo 项目的关键架构师;以及 Pushmeet Kohli,他领导 Google DeepMind 的科学工作,能告诉我们那些在围棋中开创的早期技术如今如何解决关键问题。欢迎两位来到播客。
Welcome back to Google DeepMind the podcast. I'm Professor Hannah Fry. Picture the scene. It's March 2016. Inside a hotel suite in Seoul, South Korea, two players are playing the ancient game of Go. A game of unimaginable complexity, long thought impossible for a machine to master. On one side is Lee Sedol, a legendary 18-time Go world champion. On the other, AlphaGo, a neural network-based AI system built on a powerful technique called reinforcement learning. Welcome to the DeepMind Challenge live in Seoul, Korea. That's a very surprising move. Not a single human player would have chosen move 37. After hours of intense gameplay spread over 7 days, Yeah, that's an exciting move. Lee Sedol placed two stones on the board to signal his final resignation. And in the blink of an eye, the world changed. Final result of 4-1. Congratulations to AlphaGo and to the entire team. That was exactly 1 decade ago, and the field of AI has changed unimaginably since then. We have seen the rise of large language models, the growing sophistication of AI agents, and the solving of scientific grand challenges like protein folding. But in many ways, the modern AI revolution arguably began right there, on that wooden board in South Korea. So in this episode, we wanted to look backwards and forwards to how a bold experiment in teaching machines to play games became the foundation stone for the AI breakthroughs of today. And with me are the perfect guests to tell that story. Toby Graef is a distinguished research scientist at Google DeepMind who was right there in Seoul as a key architect of the AlphaGo project. And Pushmeet Kohli who leads Google DeepMind science work and is the person to tell us how those early techniques pioneered in Go can tackle crucial problems today. Welcome to the podcast both of you.
Tory,我知道你本人就是一位出色的围棋选手。请给我们解释一下,为什么围棋被视为 AI 的一个好挑战?
Tory, I know you're an accomplished Go player yourself. Just explain to us why Go was seen as a good challenge for AI.
是的,围棋似乎是 AI 的完美挑战,因为游戏规则非常简单,却带来了如此复杂的对弈,包含战术、策略和复杂模式。一旦国际象棋被解决,或者说至少深蓝击败了世界冠军,围棋就成了一个开放的挑战。它比国际象棋复杂多个数量级,没有人预期它能很快被解决。然而,对于计算机科学家来说,它看起来如此优雅和简单,所以当时它是完美的攻克目标。
Yes, the game of Go seemed like the perfect challenge for AI because the game has such simple rules yet it leads to such complex gameplay with tactics and strategies and complex patterns. And once the game of chess had been solved as it were, or at least Deep Blue had won against the world champion, then Go was this open challenge. It's much more complex than chess by many orders of magnitude. And nobody was expecting it to be solved anytime soon. Yet it looks so elegant and simple for computer scientists, and so it was the perfect game to tackle at the time.
没人认为它能很快被解决,这个说法一针见血,对吧 Pushmeet?我知道你当时在微软工作,但这个问题被认为有多复杂?
I mean the idea of nobody thinking it would be solved anytime soon, that's sort of hits the nail on the head, right Pushmeet? I mean I know you were working at Microsoft at the time, but just how complex was this problem considered to be?
我认为它被认为极其复杂,这不仅是因为搜索空间的广度——你可走的步数,还有深度——你需要推理多长时间以及棋局有多长。在国际象棋中,你可能需要推理大约 60 到 70 步。在围棋中,这要长得多,这就带来了问题的挑战。
I think it was considered extremely complex and that is because not only because of the breadth of the search space of the number of moves you can make, but also the depth. How long you have to reason and how long the games are. In a game of chess, you might think about reasoning about 60 to 70 sort of moves. In a game of Go, it's much much longer and that leads to the challenge of the problem.
Tory,我知道你刚加入 DeepMind 时,作为一名围棋选手,你是不是第一天就和 AlphaGo 对弈了?
Tory, I know when you first started at DeepMind being a Go player, didn't you play against AlphaGo on your first day?
是的,没错。想象一下,我第一天到 DeepMind 上班。我认识几个人,包括 David Silver,他问我:‘Tory,你是围棋选手,对吧?能不能帮我们测试一下这个婴儿版的东西?’当时它甚至还不叫 AlphaGo。那是一个实习项目,他们刚刚从互联网上拿了数千局棋,或者可能是几十万局,训练了一个系统。我有机会成为最早和它对弈的人之一。但你可以想象,我既兴奋又紧张。那是我第一天上班,我被拉到中央的一张桌子对面,我想对面是 Aja Huang,他后来被称为 AlphaGo 之手,面无表情。然后我就和这个婴儿版 AlphaGo 对弈了。
Yeah, yeah, exactly. So, imagine I come first day at work at DeepMind. I know a couple of people including David Silver and he asks me 'Tory, you are a Go player, right? Couldn't you do us a favor and test our baby version of something that wasn't even called AlphaGo at the time, of course. You know, it was an internship project and they had just about taken a few thousand or games from the internet and had trained a system. Or a few hundred thousand games, maybe. And I had the opportunity to be one of the first people to play against it. But you can imagine I was excited, but I was also nervous. It was my first day at work and there I was being dragged to a centrally located table on the other side, I think it was Aja Huang who would later be known as the hand of AlphaGo with his poker face. And I got to play against this baby version of AlphaGo.
大概有很多人在看吧。
With people watching, presumably.
周围有很多人看着。你知道,无处可逃。后来 Demis 出现了,当然 David 一直在场。所以,你会怎么做?保守下棋,对吧?所以我就想,别犯错。这肯定没那么难。但当然,那正是那个版本的程序擅长的。它是在人类职业棋局上训练的,所以它完全知道如何应对常规下法。于是,随着这场小测试的进行,我的局面越来越差,最后以微弱差距输了。但我获得了‘第一个正式输给 AlphaGo 的人’的称号。那是一次相当难忘的经历。当然,之后所有人都认识了我。这是一种很好的自我介绍方式。一种让人谦卑的方式。
With a lot of people watching all around me. You know, there was no escape. Later Demis showed up and of course David was there the whole time. Yeah. And so, what does one do? Play conservatively, right? So, I just thought, just don't make a mistake. Surely this can't be so hard. But of course that was exactly what that version of the program was good at. It was trained on human professional games, so it knew exactly what to do against conventional play. And so, as this little test match proceeded, my position became worse and worse and I ended up losing by a small margin. But I took the crown of the first person who officially lost against AlphaGo. It was quite the experience. And of course, afterwards, everyone knew me. It was a wonderful way of introducing myself. A humbling way.
但请提醒我们,我知道算法从那个早期的实习项目阶段有了很大进步,但请大致解释一下它是如何工作的,特别是关于破解组合空间的想法。
But just remind us, I mean, okay, so I know that the algorithm advanced quite substantially from that early point where it was an internship, but just broadly, explain to us how it worked. And this idea about cracking the kind of combinatorial spaces in particular.
是的,我认为如果你看围棋,在任何给定时刻你可以走的步数是有限的。但如果你考虑并推理整个游戏状态,它是指数级的。你必须推理的状态数量呈指数增长,这就是游戏极其复杂的原因。
Yeah, so I think if you look at the game of Go, the number of moves that you can make at any given time, there are a finite number of moves. But if you look and reason about the overall game state, it's exponential. And that exponential growth in the number of states that you have to reason about is what makes the game extremely complicated.
那么,他们是如何破解的呢?他们发现的解决方案是什么?
So, how did they crack it then? What's the solution that they discovered?
AlphaGo 的美妙之处在于它包含了‘快思考’和‘慢思考’的元素。从某种意义上说,AlphaGo 是这两种思考过程的完美结合,共同应对这个极其庞大的搜索空间。我认为这与人类下棋的方式非常吻合。你知道,想象一下人类如何下国际象棋或围棋,我们也有能力看一眼局面,很快判断出这对黑方还是白方有利。我们也能看一眼局面,就看到一些有希望的着法。我们从不考虑所有可能的着法——国际象棋中可能有 20 或 30 种,围棋中可能有 200 或 300 种。我们会被某些着法立即吸引,这些着法可能甚至具有美学上的愉悦感,直觉告诉我们它们就是正确的。而这一元素由规划来补充,我们明确地推理各种可能性:如果我走这步,对手可能走那步,然后我必须用这步来应对。这两种不同的思维方式在人类下这些游戏时结合在一起,也在 AlphaGo 下棋时结合在一起。
The beauty of AlphaGo was there is this element of thinking fast and thinking slow. And AlphaGo in some sense was the perfect combination of those thinking fast and thinking slow processes coming together to take on this extremely large search space. And it matches quite well to how humans play the game, I think. You know, if you imagine how a human would play a game of chess or a game of Go, we also have the capacity to look at a position and pretty quickly appreciate if that's good for black or good for white. And we can also look at a position and already see moves that seem promising. We never look at all the possible moves, which would be maybe 20 or 30 in chess or 200 or 300 in Go. We are immediately drawn to certain, maybe even aesthetically pleasing moves that seemed like just the right ones, guided by our intuition. And that element is complemented by planning where we explicitly reason through the possibilities. If I make this move, my opponent might make that move, and then I have to counter with this move. And these two different ways of thinking come together in how humans play these games, and they also come together in how AlphaGo plays.
直觉和计算,可以这么说。
The intuition and the calculation as it were.
没错。那么,这是否就是灵感来源,让你思考自己和其他围棋选手是如何下棋的,并有效地从神经科学中直接汲取灵感?
Exactly. So, was that the inspiration that made you think about how you were playing the game, how other Go players were playing the game, and draw that direct inspiration from neuroscience effectively?
是的,我认为这绝对是一个方向,因为很多团队成员本身就是棋手,能够内省并观察我们如何应对棋局。然后这与深度学习结合,当时自 2012 年以来,深度学习已经发展成为一个方向,并首次为我们提供了工具来学习近似函数,比如价值函数,它接收棋盘并告诉我们黑白双方的优势;或者策略网络,它接收棋盘并根据职业选手选择某一步的可能性对可用着法进行排序。所以深度学习当时已经成熟,可以解决这个问题,并让我们有机会实现快速思考。慢速思考则与深蓝类似,是游戏树的搜索,这早已为人所知,我们现在可能称之为传统人工智能。
Yeah, I think that is definitely one direction because a lot of team members were actually game players who were able to introspect and see how we tackle the game. And then that came together with deep learning, which at the time, since 2012, had grown as a direction and for the first time gave us the tools to learn approximate functions for things like the value function that takes a board and tells us how good it is for either black or white, or the policy network that takes a board and effectively ranks the available moves according to how likely a professional player would take them. So deep learning was ripe at the time to tackle this problem and gave us the opportunity to implement the fast thinking. The slow thinking is not unlike what happened in Deep Blue. It's the search of the game tree that was already known and that we might now call good old-fashioned AI.
好吧,我的意思是,你早期输给了这个东西,但一旦它经过了团队中很多人的测试,我知道你让一位职业围棋选手来测试它,因为你邀请了樊麾来办公室。
Okay, I mean, you lost to this thing quite early, but once it had gone through a lot of the people on the team, let's say, I know that you tested it with a professional Go player because you had Fan Hui come into the office.
是的,没错。
Yeah. Exactly.
当时你有多大信心它能击败他?
How confident were you at that point that it was going to beat him?
我们的信心程度各不相同,这很有趣。我们很幸运找到了他。他是当时的欧洲围棋冠军,住在波尔多,然后过来了。我们引诱他和我们下棋。安排是他和当时的 AlphaGo 版本下 10 盘测试棋。我个人认为 AlphaGo 不可能已经达到击败欧洲冠军——一位职业选手的水平。所以我和大卫·西尔弗打了个赌。大卫·西尔弗很有信心,他说:“我觉得 AlphaGo 会 10-0 完胜。”我说:“不,我觉得 AlphaGo 至少会输一盘。”赌注是输的人要打扮成古代日本围棋大师的样子出现在办公室,并待上一整天。结果谁那样出现了?是我,因为实际上是 10-0。但这确实给了我们信心,让我们相信在不久的将来能够应对更强大的对手。
We had different levels of confidence, which was really interesting. We had been really lucky to find him. He was the European Go champion at the time. He lived in Bordeaux and came over. We lured him into playing this game with us. The setup was that he would play 10 test games against the version of AlphaGo at that point. I personally thought that AlphaGo could not possibly be at the point already that it beats the European champion, a professional player. So I had a bet with David Silver. David Silver was confident. He said, "I think AlphaGo is going to nail it 10-0." And I said, "No, I think AlphaGo will lose at least one game." The bet was that whoever lost would have to show up at the office dressed as an ancient Japanese Go master and be in the office for one day with that. Well, who showed up like that? It was me because it was in fact 10-0. But it did give us confidence and made us confident that we would be able to tackle even harder opponents in the near future.
当然,你们后来做到了。2016 年你们乘飞机去韩国首尔,与李世石对弈。跟我们说说,他到底是一位多么出色的棋手?
Which you did, of course. On a plane you got in 2016 to Seoul, in Korea, to play against Lee Sedol. I mean, just tell us, give us a sense of how phenomenal a player he actually is.
是的,李世石当时确实是顶尖棋手之一,甚至可能是最好的,他赢得锦标赛的记录令人难以置信。当时人们把他比作罗杰·费德勒,因为他的成功和智力 brilliance。所以我们非常荣幸他接受了我们的挑战。这也是一个巨大的挑战,因为我们必须设定一个日期,对吧?你不能说我们准备好了再通知你。日期定下来了,我们必须朝着那个日期努力,让 AlphaGo 足够强大。更增添紧张和兴奋的是,李世石坚信自己会赢。他认为 AlphaGo 获胜的可能性极低。当然,他的评估是基于他看到的 AlphaGo 与樊麾对弈的记录,他认为自己更强。但他没有意识到的是,AlphaGo 通过我们进行的训练和算法改进在不断提升。所以整个团队基本上都去了韩国,你无法相信那里人们的兴奋程度。在英国,围棋是一种小众活动,对吧?很少有人会下棋甚至了解它。但在韩国,人们非常兴奋。顶尖围棋选手都是名人。我们到了那里,成群的摄影师拍照。我们还有一个纪录片摄制组。想象一下,典型的计算机极客突然因为这场比赛成为世界瞩目的焦点。那真是一次冒险。
Yeah, so Lee Sedol was really one of the or maybe the best players at the time with an incredible track record of winning tournaments. He was compared to Roger Federer at the time for his success and intellectual brilliance. So for us it was a tremendous honor that he accepted our challenge to play against him. And it was a tremendous challenge because we had to set a date, right? You can't just say we'll tell you when we're ready. The date was set and we had to work towards that date to actually make AlphaGo strong enough. What added tension and excitement was that Lee Sedol was convinced that he would win. He thought it highly unlikely at the time that AlphaGo would win. Of course he was basing his assessment on the game records he had seen against Fan Hui and he assessed that he was better. But what he wasn't so aware of is that AlphaGo was constantly improving through training and algorithmic refinements that we made. So the entire team basically went to South Korea and you wouldn't believe the excitement of people there. In England, Go is a bit of a niche activity, right? Very few people would be able to play it or even know about it. But in South Korea people were so excited. The best Go players are celebrities. We came there and there were hordes of photographers taking pictures. We had a documentary film crew with us. So imagine typical computer geeks suddenly in the limelight of the world for this match. That was quite the adventure.
是的。我的意思是,你对 AlphaGo 的表现感到紧张吗?
Yeah. I mean, were you nervous about the performance of AlphaGo?
是的,我们确实很紧张。当然,我们有一个非常复杂的评估流程。你可以与你能接触到的棋手对弈,比如樊麾,这非常有帮助。你也可以与程序的早期版本对弈,并计算我们所说的系统的 Elo 评分,这基本上是根据你与其他版本(可能是早期版本)对弈的所有结果,来计算新版本的等级分。你可以很好地校准这些东西。但我们当然不知道李世石在那个尺度上处于什么位置。我们也希望有一定的缓冲,如果能强出不少,有一些确定性就好了。因为这是世界舞台,对吧?如果输了,对声誉会是一个打击。所以我们很紧张。我们一直工作到最后一刻。我们还需要确保系统非常稳定。你不想为了稍微改进而做最后一刻的改动,却冒系统变得不稳定的风险。但最终我们对它很满意,于是我们进入了那个现在很著名的酒店楼层,所有行动都在那里进行,所有媒体都在等待,然后开始了比赛。全世界的人都在观看,包括 Pushmeet。
Yes, we were definitely nervous. Of course we had a very sophisticated evaluation pipeline. You can test against players you have access to like Fan Hui, that was super helpful. You can also test against previous versions of the program and calculate what we call the Elo score of the system, which basically takes the outcomes of all the games you play against other versions, maybe earlier versions of your program, and calculates what the rating of the new version is. You can calibrate these things quite well. But of course we didn't know where on that scale Lee Sedol would be. And we wanted a cushion as well, it would be nice to be quite a bit better, to have some certainty. Because this is the world stage, right? If you lose this, that's a bit of a hit to the reputation. So we were nervous. We worked up to the last minute. We also needed to make sure that the system is really stable. You don't want to make last minute changes to make it that little bit better, but risk that it now becomes unstable. But in the end we were quite happy with it, and so we entered that now famous hotel floor where all the action happened, where all the press was waiting, and embarked on the match. And people were watching from around the world, including Pushmeet.
是的。那么,当时你在哪里?你在观看吗?
Yeah. So I mean, where were you at this point? Were you watching on?
是的,我在西雅图。我在第一局比赛的中途真正开始投入。很明显,AlphaGo 已经达到了那个特定的里程碑,你甚至可以看到媒体、解说员以及李世石本人的反应。
Yeah, I was in Seattle. I really started getting into it in the middle of the first game. It became so clear that AlphaGo had reached that specific milestone, and you could even see the reaction from the press and the commentators, and Lee Sedol himself.
嗯,你说在比赛中途很有趣,因为在比赛早期,是否清楚谁占上风?
Well, it's interesting you said in the middle of that game, because in the early stages of that game, was it clear who had the upper hand?
我认为从一个只是观看比赛的人的角度来看,我觉得在早期阶段,每个人都相当确信李世石会赢。事实上,只有当比赛进行下去,越来越接近最终结果时,他们才意识到,当你计算领地时,AlphaGo 有优势。事实上,这出乎人们的意料。
I think from a person who was just watching it, I felt that in the early stages, everyone felt quite confident that Lee Sedol would win. In fact, only as the game progressed and it became closer to the final outcome that they realized that as you count the territory, AlphaGo had an advantage. In fact, it came as a surprise to people.
你怎么看?嗯,我在现场有一次有趣的互动,和一位职业围棋选手,一位美国职业棋手,他坐在我旁边,我们一起观看比赛。当时角落里出现了一些棋局变化,他凑过来对我说:“你知道吗,我总告诉我的学生不要下那个蠢招,但 AlphaGo 刚刚就下了。所以,我觉得这太没希望了。”我回答说:“我不算专家,我们等等看吧。”这是我的反应。然后第一局结束后,这位先生走过来对我说:“这是我经历过的最非凡的事情。我非常感激能在这里见证一台机器能下出如此水平的围棋,我们能从中学到太多东西。”他已经开始接受这一点了。你要知道,这些人把一生都奉献给了围棋研究,他们往往从小就开始训练,直到现在这个年纪,只为精通这项游戏。所以,当一台机器可能匹敌甚至超越人类棋手时,对他们来说当然是一个冲击。
What did you think? Yeah, so I had this interesting interaction on site with a professional Go player, an American professional Go player, who was sitting next to me while we were watching and there was some sequence unfolding in a corner and he kind of approached me and said, "You know, I always tell my students not to play that stupid move that AlphaGo just played. So, I mean, it's pretty hopeless." And I was like, "I'm not as much of an expert. Let's just wait and see." was my reaction. And then after that first game, this gentleman came to me and said, "This is the most phenomenal thing I've ever experienced. I'm so grateful that I'm allowed to be here to witness that a machine can play Go at this level and there's going to be so much we can learn from it." And he was already embracing this. I mean, you have to imagine these people dedicate their lives to the study of this game and they often trained from being young children to their current age just to master this game. And so, of course, it comes as a shock to them that a machine might match or even exceed a human Go player.
因为如果说第一局 AlphaGo 赢了,那么在第二局中,AlphaGo 做了一件让所有人都大吃一惊的事。哇,那是一个非常令人惊讶的落子。职业解说几乎一致认为,没有哪个人类棋手会选择第 37 手。AlphaGo 表示,人类棋手下出第 37 手的概率只有万分之一。跟我们说说现在著名的第 37 手是怎么回事吧。
Because if that was the first game when AlphaGo won, in the second game, AlphaGo did something that really surprised everybody. Wow. That's a very surprising move. Professional commentators almost unanimously said that not a single human player would have chosen move 37. AlphaGo said there was a one in 10,000 probability that move 37 would have been played by a human player. Just explain to us what happened with the now famous move 37.
是的,那是一个非凡的场景。我当时坐在国际英语解说室里,我们的美国解说员迈克尔·雷德蒙面前有一块大演示板,他把所有棋子都摆在上面,向人们展示落子情况,并评论各种变化。他拿起对应第 37 手的棋子放在棋盘上,然后后退一步说:“不,这肯定错了。”他又把棋子拿了回来。然后他再看屏幕,说:“不,不,这确实是 AlphaGo 下的。”他又把棋子放了回去。他很困惑。你能看出来,这对人类棋手来说是一个非常反直觉的落子。那是一个五线的肩冲。这通常是人类棋手会避免的。在围棋中,经常有沿着边线的推挤,一方沿棋盘边线围空,另一方则向中央发展势力。如果这发生在三线和四线,通常被认为是大致均衡的,双方都有所得。但 AlphaGo 实际上是在暗示,即使你下在五线,给对方更多的实地,仍然是划算的。这就是让人们如此惊讶的地方——竟然存在这种情况,使得这样下是正确的。
Yeah, so this was a remarkable scene. I was sitting in the international English-speaking commentating room and Michael Redmond, our American commentator, he had this big demo board on the wall and he would put all the stones up there on the board to show people what was being played and comment on different variations. And so he took the stone corresponding to move 37 on the board and then he stepped back and said, "No, this must be wrong." And he took it back. And then he looked at the screen again and said, "No, no, that is actually what AlphaGo played." And he put it back. He was puzzled. You could see that it was such a counterintuitive move for a human player. It was a shoulder move on the fifth line. And this is typically something that human Go players avoid. So often in Go there is some kind of pushing going on along the edges and one of the players builds territory along the wall of the board and the other side builds influence towards the center of the board. And if that happens on the third and fourth line, this is considered to be roughly equitable, both sides get something out of it. But what AlphaGo was effectively suggesting is that it's still profitable if you do it on the fifth line and you give that much more territory to the other party. And that's what was so surprising to people, that there could be situations in which that would be correct.
所以,这不仅是一个非常特别的落子,而且在某种程度上代表了一种权衡即时实地与向中央发展势力这两个因素的新方式。我想,这超越了人类棋手通常的做法。
And so, not only was it a very special move, but it in a way represented a new way of weighing these two factors of immediate territory versus influence towards the center of the board against each other. Something that went beyond what a human Go player would normally do, I assume.
是的,绝对如此。我的意思是,有这样的时刻,你看到任何 AI 系统扩展人类知识的真正潜力,在这个具体案例中,人们多年来一直将围棋视为一个值得研究的领域。然后到了这个特定时刻,知识被扩展了。人们起初持怀疑态度,比赛中也是如此。当这步棋下出时,它被认为是幻觉,而不是错误,对吧?在它的影响在后续比赛中变得清晰之前,有相当长一段时间都是这样。
Yeah, absolutely. I mean, there are moments like this where you see the true potential of any AI system expanding human knowledge, where people have regarded, in this particular case, the game of Go as a thing to be studied for many many years. And there comes this particular point where that knowledge is expanded. And people were at first skeptical, which was the case in the game as well. When the move was played, it was considered a hallucination, not a mistake, right? For quite a bit of time before its implications became clear later on in the game.
没错,因为它被证明是第二场胜利的关键。
Exactly, because it proved to be pivotal to the second win.
这不仅仅是那场比赛中的一个时刻,我认为也是整个 AI 历史上的一个时刻。那个特定时刻向我们展示了,有时这些系统会产生洞察,我们甚至可能无法辨别它们是正确的还是惊人的突破,但它们将极大地影响我们以全新的视角看待整个研究领域。
It was not just a moment in that game, but it was also a moment, I think, in the whole sort of history of AI where that particular moment showed us that there will be times when these systems will produce insights which we might not even be able to discern whether they are the right things or amazing breakthroughs, but yet they will have a lot of influence in how we look at whole areas of study in a completely new light.
嗯,我还想谈谈第 78 手。这是李世石下的一步棋,它让 AlphaGo 困惑,导致它认输了。李世石在这里想干什么?他光这一步就花了七八分钟。嗯,看这步棋。这是一步激动人心的棋。哦。你知道,我其实不确定 AlphaGo 想在这里做什么。所以,他找到了它的弱点。那个挖的棋。世界冠军李世石在第四局中寻找 AlphaGo 的弱点,他找到了。此时 AlphaGo 已经连胜三局。现在李世石下了一步让系统困惑的棋。这么说公平吗?
Well, I also want to talk about move 78. This is a move that was played by Lee Sedol that confused AlphaGo, causing it to resign the game. What is Lee Sedol up to here? He's just burned like 7 or 8 minutes just on this move already. Hm. Look at that move. That's an exciting move. Ooh. You know, I'm not actually sure what AlphaGo is trying to do here. So, he's found his weakness. That wedge move. World champion Lee Sedol went looking for AlphaGo's weakness in game four and he found it. So, by this point AlphaGo has won three games in a row. And now Lee Sedol does a move that confuses the system. Is that fair to say?
是的,这么说完全公平。所以,第 78 手是李世石下的一步不寻常的挖。棋盘中央发生了一场非常有趣的战斗。李世石找到了这步棋,它也让人们感到惊讶,类似于第 37 手。从那时起,我们观察到 AlphaGo 不再能很好地把握局面。我们看到它下的棋对我们来说在不好的意义上说不通,你知道,第 37 手可能也说不通,但这些棋即使对我们这样的业余爱好者来说也显得奇怪,所以它被这步棋搞糊涂了。退一步说,让你明白为什么这仍然如此重要。你可能会说:“好吧,这是五局三胜的比赛,AlphaGo 已经赢了前三局。还有什么可证明的?”但我们当时在想,如果现在你会得出什么结论?他找到了破解之法,对吧?他发现了脆弱之处。
Yeah, that's absolutely fair to say. So, move 78 was an unusual wedge move that Lee Sedol played. There had been a very interesting battle as it were at the center of the board. And Lee Sedol found this move and it was also surprising to people, similar to move 37. And from then on we observed that AlphaGo didn't have a good grasp of the position anymore. We saw that the moves that it made didn't really make sense to us in a bad way, you know, 37 also didn't make sense to us maybe, but these moves even to amateurs like us seemed strange and so it had been confused by the move. And just to zoom out to give you a sense of why this still mattered so much. So, you might say, "Okay, it's a match of five games and AlphaGo has won the first three. What more is there to prove?" But then we were thinking, well, if now what would you conclude? He's got it figured out, right? He's found the fragility.
没错。所以,对人类来说,那将是一个人类的胜利。这就是为什么那场比赛和最后一场对我们来说仍然非常激动人心。但我们并非完全失望。我们当然失望,但我们也非常钦佩李世石。你知道,作为人类,能够找到这步棋。你只需想象这位大师,他一生致力于围棋,在这场战斗中一定非常艰难,对吧?看到这台机器下得如此完美,而他努力寻找出路。然后在第四局中,他找到了办法。正如他在新闻发布会上所说,我想他后来表示,他非常高兴和自豪,能够也许是最后一次代表人类找到战胜机器的方法。因为有些人称它为神之一手,不是吗?
Exactly. So, the human it would have been a human triumph. And so, that's why that game and the last one were still very exciting to us. But it wasn't entirely the case that we were disappointed. We were certainly disappointed, but also we had so much admiration for Lee Sedol too. You know, as a human to be able to find this move. You just have to imagine this master who has dedicated his life to playing this game in this battle that must have been so hard on him, right? To see this machine play so perfectly and him struggling to find a way. And then in game four he finds a way. And as he put it in the press conference, I think later he said that he was so happy and proud that he was able maybe for the last time on behalf of humanity to find a way to overcome the machine. Because some people called it the divine move, didn't they?
是的。
Yeah.
最终比分是 AlphaGo 以 4 比 1 获胜。围棋界的反应如何?
Well, the final score was 4-1 to AlphaGo in total. What was the reaction from the Go community?
围棋界非常关注这场比赛,结果当然很戏剧化,对很多人来说也出乎意料。所以人们的反应各不相同。有些人完全被结果震惊了,有些人不敢相信,还有些人认为某个时代结束了,因为现在最强的棋手可能不再是人类,而是机器。但总的来说,我们发现令人惊讶的是,人们对围棋的兴趣增加了。我认为现在下围棋的人比以前更多了,围棋界也真正接受了从 AlphaGo 中学到的东西。现在有很多程序基本上以与 AlphaGo 相同的方式运作,人们用它来教学,分析自己的棋局。总的来说,我认为它提升了整个围棋界。
Yeah, so the Go community followed the match very closely and of course the outcome was dramatic and for many people unexpected. So, people showed very different reactions. Some people were absolutely amazed and surprised about the outcome. Some people couldn't believe it. Others thought that some era had come to an end because now maybe the strongest Go player was no longer a human, but a machine. But overall, what we found amazing is that there was an uptick in interest in the game of Go. I think more people play Go now than did before, and the Go community really embraced the learning from AlphaGo. So, there are now many programs that work essentially the same way that AlphaGo does, and people use it for teaching purposes. They analyze their games through it, and overall, I think it has provided a lift to the whole Go community.
我想问问你,AI 界对这场比赛的反应如何?大家都在讨论什么?
Let me ask you about the reaction from the AI world to this match. What was the buzz? What was the conversation like?
李世石比赛,AlphaGo 对李世石的比赛,是一个关键的转折点。很多人,尤其是机器学习社区的人,他们一直把这些模型和技术当作数学和应用项目来研究,现在开始看到证据表明这些系统可以自我学习并超越人类知识。这是一个非常重要的点,因为在机器学习中,你用收集到的训练数据进行训练,你自然的期望是模型只会与那个分布一致。而证明你可以超越那个分布,并且这种洞察可以被世界利用,我认为这是整个经历中得出的一个惊人见解。它真正指出了人工智能的可能性,不仅是在围棋游戏中,而且在理解世界、化学、生物学、数学、计算机科学中。这些系统将能够发现并向我们揭示哪些类似于第 37 手的惊人类比?
The Lee Sedol match, the AlphaGo Lee Sedol match, was a key pivot point where a lot of people, especially in the machine learning community, who had been working on these models and techniques as a mathematical and applied project, started to see evidence that these systems can self-learn and go beyond human knowledge. And that is a very important point because in machine learning, you train with training data which has been collected, and your natural expectation is that the model is going to just be consistent with that distribution. To show that you can go beyond that distribution, and that insight then can be utilized by the world, I think is an amazing insight that comes out of this whole experience. And it really points to what is possible with artificial intelligence in not just the game of Go, but in the understanding of world, in chemistry, in biology, in mathematics, in computer science. What are these amazing analogs of move 37 that these systems will be able to discover and reveal to us?
我觉得你刚才提到的超越人类智能这一点非常迷人。但 AlphaGo 故事中最让我感兴趣的一点是,即使在 4 比 1 获胜之后,你们又构建了 AlphaZero,去掉了所有人类数据,所有它训练过的围棋棋谱,结果发现一旦去掉人类智能,它反而变得更好了,这让我很震惊。
I think that point that you made there about going beyond human intelligence is just so fascinating. But one of the things that I find most intriguing about the AlphaGo story, even after the victory of 4-1, is that you then built AlphaZero, where you took away all of the human data, all of the games of Go that it had been trained on, and discovered that once you take out the human intelligence, the thing actually improved, which is astonishing to me.
是的,从科学角度来看,可以说这比最初的 AlphaGo 是更大的一步。正如你所说,AlphaZero 系统无法访问任何人类棋谱,也不知道人类如何下棋,没有关于游戏的先验知识,只知道游戏规则,以及表示和学习我们讨论过的那些函数(策略网络和价值网络)的方法。所以基本上,它一开始完全随机下棋,因为它不知道什么是好棋或坏棋,但它通过下棋积累经验,学习哪些棋步更可能赢,哪些更可能输,哪些局面有希望,哪些没有希望,最终它开始下出越来越好的棋步。当然,它现在不受人类知识的限制。它发现的东西令人惊叹。首先,它重新发现了人类下棋的方式。这完全令人放心,你知道,围棋中有一些角部定式,我们称之为 joseki。或者在国际象棋中,有一些开局走法。这个系统现在更通用。它可以下围棋、国际象棋和将棋,如果我们以那种方式训练它,它还可以下任何其他棋盘游戏。所以,起初它重新发现了人类知识。我们想,‘哇,这太酷了。它找到了相同的开局等等。’然后我们观察其中一些开局,它不再下那些了。我们想,‘怎么回事?’它找到了一个反驳。所以,它发现并重新发现了人类知识,然后抛弃了它,因为它现在已经超越了人类知识,并找到了实际上更好的下法。它不再继续以这种人类的方式下棋。实际上是人类尚未发现的东西。
Yeah, from a scientific perspective, one could argue that that is even a bigger step than the original AlphaGo. Because as you were saying, the AlphaZero system doesn't have access to any human game records, or how humans play, didn't have access to prior knowledge about the game, how the game is played, but really only had access to the rules of the game, and means of representing and learning these functions that we talked about, the policy net and the value net. So basically, it starts playing entirely randomly at the beginning because it has no notion of what good or bad moves are, but it gathers experience from playing these games, and it learns what are moves that are more likely to lead to a win, what are moves that are more likely to lead to loss, what are positions that look promising, what are positions that are not promising, and eventually, it starts playing better and better moves. And now, of course, it's not limited by human knowledge. And what it discovered was amazing. So, first of all, it rediscovered ways of how humans play. And that was totally reassuring, you know, there's certain patterns in the corner in Go that we call a joseki. Or in chess, there are certain opening moves. The system was now more general. It could play chess, Go, and shogi and could have played any number of other board games if we trained it that way. And so, at first it rediscovers human knowledge. And we think, 'Wow, this is so cool. It finds the same openings and so on.' And then we look at some of these openings and it stops playing them. We think, 'What's going on?' It has found a refutation. So, it discovered rediscovered human knowledge and then it discards it because it has now gone beyond it and has found there's actually better ways of playing. I'm not going to continue playing in this human way. Stuff that humans hadn't found yet, effectively.
实际上是人类尚未发现的东西。
Stuff that humans hadn't found yet, effectively.
完全正确。对于 AlphaZero,当它下围棋时,它下棋的方式最终对我来说看起来很陌生。所以,这不是我从围棋老师那里学到的那种围棋,你知道,那种围棋的结构可能便于人类理解。这些棋步看起来非常自由,当时没什么意义。但 30 步之后,一切就都到位了。你会看到,‘哦,是的。哦,哇,现在说得通了。’等等,就好像它有某种远见,而它确实有,对吧?所以,从无到有达到那种水平的发现非常令人印象深刻。
Exactly. For AlphaZero, when it played Go, the way it played Go looked alien to me in the end. So, this wasn't the kind of Go that I had learned from my Go teacher, you know, which is structured maybe in a way that enables humans to understand it. These moves looked very free and didn't make much sense at the time. But 30 moves later everything would fall into place. You see, 'Oh, yeah. Oh, wow, that makes sense now.' And so on, as if it had the foresight, in a way, which it did, right? So, that discovery from nothing to that level of play was very impressive.
好的,我想给你看个东西,是你们在首尔时发生的事。因为正如你之前提到的,你们当时正在为 AlphaGo 的纪录片拍摄。有一些镜头没有剪进电影里,但被摄像机捕捉到了,当时他们正在收拾设备,但麦克风还开着。我不知道你是否听过这段小片段。让我放给你听。稍等。这是 Demis 和 David 在私下交谈。‘看到一个问题从被认为不可能变成几乎完成,真是令人惊叹。>> 可以解决蛋白质折叠问题。那简直是巨大的。我相信我们现在可以做到了。我以前就认为我们可以做到。但现在我们肯定可以做到。太美了。是不是很棒?>> 是的。’Pushmeet,你认为这捕捉到了当时的情绪吗?
Okay, so there's something I want to show you, something that happened actually when you guys were in Seoul. Because you as you mentioned before, you were being filmed for this documentary for AlphaGo. And there's some footage that didn't make it into the film, but it was captured by the cameras as they were sort of packing up but the microphones were still running. I don't know if you've heard this little clip. Let me play it for you. Hold on. This is Demis and David having a sort of private conversation. 'Well, it's just amazing seeing how quickly a problem that is seen as being impossible can change to being practically done. >> can solve protein folding. That's like I mean it's just huge. I'm sure we can do that now. I was I thought we could do that before. And now but now we definitely can do it. Beautiful. Isn't that great? >> Yeah.' Tore, do you think that that captured the mood at the time?
是的,那就是 AlphaGo 当时打开的那扇门。对吧?如果我们能做到这个,那我们还能做什么?因为这是一个有 10 的 170 次方种不同局面的游戏。这非常复杂。如果我们有原则性的方法来导航那种组合搜索空间,那么似乎我们也能处理其他大型组合搜索空间。当时,其中一个热门就是蛋白质折叠。绝对是的。而这就是你真正加入 DeepMind 团队的时刻,Pushmeet,因为当谈到 AlphaFold 时,你是那个故事中不可或缺的一部分。
Yeah, that was the kind of door that AlphaGo opened at the time. Right? If we can do this, then what else could we do? Because this is a game with 10 to the power of 170 different positions. This is super complex. And if we have principled ways of navigating that kind of combinatorial search space, then it seemed plausible that we would also be able to handle other large combinatorial search spaces. And at the time, one of the favorites was protein folding. Absolutely. And this is now the point really where you come aboard with the DeepMind team, Pushmeet, because when it came to AlphaFold, I mean you're an integral part of that story.
AlphaGo 是否直接影响了你们后来的工作,还是说这场胜利带来的信心让 Demis 说出了那些话?
Did AlphaGo directly influence what you guys went on to do, or was it the confidence of a victory that made Demis say things like that?
不,我认为 Demis 从一开始就对 AI 的发展目标有着非常清晰的认识。他确实将 AI 视为帮助我们更好地理解世界的工具。事实上,当 AlphaGo 比赛进行时,我在微软从事 AI 编程方面的工作。现在,AI 编程无处不在,但当时没有多少人从事程序合成和 AI 编程。Demis 希望我加入 DeepMind,我问他:我真正感兴趣的是让 AI 系统和机器学习系统来解决世界上最具挑战性的问题,并理解正在发生的事情。他的反应是:如果你想理解世界并解决最重要的问题,那么你必须加入 DeepMind,因为我们需要 AI 来深入理解世界并应对这些问题。所以,如果你对编程、网络安全、气候变化或难以治疗的疾病感兴趣,你就必须来领导如何将 AI 用于这些应用。
No, I think Demis from very early on had a very strong notion of what AI is being developed for. He really sees AI as a tool that will help us understand the world better. In fact, when the AlphaGo matches were happening, I was at Microsoft working on AI for programming. Now, AI for coding is everywhere, but at that time not many people were working on program synthesis and AI for coding. Demis wanted me to join DeepMind, and my question to him was: I am really interested in having AI systems, machine learning systems for solving the most challenging problems in the world and to make sense of what's happening. I think his reaction was: if you want to understand the world and solve the most important problems, then you have to join DeepMind because we will need AI to really understand the world deeply and to tackle these problems. So, if you are interested in learning to program, cybersecurity, climate change, or impossible-to-treat diseases, you have to come and lead the charge on how AI can be used for these applications.
我想问一下 AlphaGo 中的一些创新,以及它们如何最终应用到你们的科学项目中。AlphaGo 的一大成就是让巨大的搜索空间变得更容易处理。那么,搜索算法自那时起发生了怎样的变化?它们又是如何被用于科学的?
I want to ask about some of the innovations in AlphaGo and how they ended up in your science projects. One big thing AlphaGo did was make that gigantic search space more tractable. So, how have search algorithms changed since then, and how are they being used in science?
搜索是许多现实世界问题中不可或缺的一部分。我们刚刚谈到了蛋白质折叠,它可以被视为对所有可能结构空间的搜索。但举个更简单的例子,你也可以把搜索看作是对解决特定问题的算法的搜索。我们周围计算机所做的一切都涉及某种形式的矩阵乘法。即使是今天改变世界的机器学习系统和神经网络,也基于矩阵乘法。本质上就是取大的数字矩阵并将它们相乘。即使是最简单的矩阵乘法运算,也是你在学校和大学里学到的最简单的东西,但整个研究界却不知道两个矩阵相乘的最快方法是什么。如果你思考这个问题,可以将其视为一个搜索问题。存在一个可能的算法空间,现在在这个空间中搜索以找到最佳算法。问题是,这个问题的搜索空间甚至比围棋的搜索空间还要大。所以,我们需要做的第一件事就是提出一个名为 AlphaTensor 的智能体,它将矩阵乘法变成一个搜索问题,一个游戏。不是问你在围棋中赢了还是输了,而是问:你是否快速地将这两个矩阵相乘了?你是否用最少的步骤完全准确地相乘了这些矩阵?这就是游戏。1969 年 Strassen 提出了一种算法,此后 50 年没有进展。然后 AlphaTensor 找到了一种更好的方法来相乘这两个矩阵。这是证明相同技术可能实现的关键证据。
Search is such an integral part of many real-world problems. We just spoke about protein folding, which can be considered as a search over the space of all possible structures. But to give a simpler example, you can think of search as also the search for algorithms to solve a particular problem. Everything around us that computers do has some form of matrix multiplication underlying it. Even the fact that we have these machine learning systems and neural networks that are changing the world today, these neural networks are based on matrix multiplication. Essentially taking large matrices of numbers and multiplying them together. Even the very simplest operation of matrix multiplication is the simplest thing you learn in school and college, yet we don't know as a whole research community what is the fastest way of multiplying two matrices. If you think about that problem, you can reason about it as a search problem. There is a space of possible algorithms, and now search over that space to find the best algorithm. The issue is that the search space for that problem is even larger than the search space for Go. So, one of the first things we needed to do was come up with an agent called AlphaTensor, which made matrix multiplication a search problem, as a game. Instead of did you win or lose the game of Go, you're saying: did you multiply these two matrices together quickly or not? Did you multiply these matrices completely accurately in the smallest number of moves? That was the game. Strassen in 1969 had come up with an algorithm, and for 50 years there was no progress. Then AlphaTensor found a better way of multiplying these two matrices. That was a key proof point of what is possible with the same sort of techniques.
对于不熟悉矩阵乘法的人来说,每个大型语言模型本质上都是一个巨大的矩阵乘法问题。关于不同芯片的所有争论,都是因为有些芯片可以更快地相乘矩阵。即使速度上微小的提升,当扩展到全世界使用 AI 的规模时,也会产生巨大的差异。
For anyone not familiar with matrix multiplication, every large language model is essentially a massive matrix multiplication problem. All the fuss about different chips is because some can multiply matrices faster. Even small gains in speed, when scaled up to how much everyone is using AI, make gigantic differences.
是的,完全正确。从那以后,我们决定:不仅要解决矩阵乘法,还要解决你能想到的所有可能的算法。因此,我们的新智能体如 Alpha Evolve 在所有可能程序的空间中搜索,试图找到能够解决重要问题的最佳算法,无论是如何在数据中心调度作业(这对能源和算力利用率有影响),还是如何在网络中移动数据包来解决物流问题。解决这些搜索问题的相同基本方法现在已经在应用范围上得到了扩展。
Yeah, absolutely. And since then, we have said: let's not just tackle matrix multiplication, let's tackle all possible algorithms you can think of. So our new agents like Alpha Evolve search in the space of all possible programs, trying to find the best algorithm that can solve important problems, whether it's how to schedule jobs in a data center, which has implications for energy and compute utilization, or how to tackle logistics problems where you move packets around in a network. The same basic methodology of tackling these search problems has now been expanded in terms of what you can do with it.
但我在想策略网络,你描述的那种直觉,就像围棋选手看着棋盘说:“我认为这是一个有希望探索的方向。”如果面对的不是围棋棋盘,而是世界上所有可能的算法,你究竟如何在这种情境下创造直觉?你怎么知道如何缩小搜索空间?
But I'm thinking about the policy network, the intuition you described, where a Go player looks at the board and says, 'I think this is a fruitful direction to search.' If instead of a Go board, you have all possible algorithms of everything in the world, how on earth do you create intuition in that situation? How do you know how to narrow down the search space?
这是一个非常有趣的研究课题,我们现在开始思考如何应用像 Alpha Evolve 这样的智能体来发现新算法。有时这些算法对我们来说并不直观;事实上,它们可能违反直觉。有时你可以看到模式,问题中某些对称性,数学家和计算机科学家并不理解,但智能体却发现了这些对称性,并利用它们使解决方案更加高效。在某些情况下,我们就是不明白它是如何让事情变快的,但它们确实更快了。那么我们的挑战是:当你考虑人类和这些 AI 智能体协作时,我们如何确保产生的系统和算法对人类计算机科学家和工程师是可解释的?
This is a very interesting research topic we are now starting to think about when applying agents like Alpha Evolve to discover new algorithms. Sometimes those algorithms are not very intuitive to us; in fact, they could be counterintuitive. Sometimes you can see patterns, certain symmetries in the problem that mathematicians and computer scientists did not understand, but somehow the agent discovered those symmetries and exploited them to make the solution much more efficient. In some cases, we just don't understand how it made things faster, but they are faster. Then our challenge is: when you think about collaboration where humans and these AI agents work together, how do we make sure that the systems and algorithms produced are interpretable by human computer scientists and engineers?
这让我想起 AlphaGo 的情况,人们在终局观察 AlphaGo 时发现它并没有完全最优地下棋。他们非常惊讶地说:‘看,这步棋比 AlphaGo 下的更好。’它是不是下得不好?是不是在犯错?而答案是 AlphaGo 在优化我们给它的目标,即最大化获胜概率。人类倾向于使用一种启发式方法,他们希望比对手多占一些领地,并且认为差距越大越好。这通常是对的,但 AlphaGo 不在乎差距。对 AlphaGo 来说,赢半目就够了。所以在终局,它常常像是在戏弄对手,放弃分数,直到它确信能赢半目为止。你有时会看到这些反直觉的行为,但如果你深入探究,就能明白它们为什么会出现。因为算法和人类最终优化的目标略有不同。
It reminds me a little bit of the situation in AlphaGo where people in the end game were observing AlphaGo and found that it didn't quite play optimally. And they were really surprised to say, 'Look, this is a better move than what AlphaGo played.' You know, is it not playing well? Is it making mistakes? And the solution was that AlphaGo was optimizing the objective we had given it, which is to maximize the probability of winning the game. Humans tend to use a heuristic, which is they want to have more territory than the opponent by some margin, and they think the larger the margin is, the better it is for them. Which is often true, but AlphaGo doesn't care about the margin. For AlphaGo, it was enough to win by half a point. And so often in the end game, it seemed to be toying with the opponent and giving up points just up until the point where it was sure it could win by half a point. And you sometimes get these counterintuitive behaviors, but if you then drill deeper, you can see why they come about. Because the algorithm and the humans are ultimately optimizing for slightly different things.
没错。好吧,但这确实让我好奇。比如第 37 手,它超越了人类的能力。但同时,当第 37 手第一次出现时,人们认为它是一个错误,对吧?那么你怎么区分呢?我的意思是,如果算法想出了原创的东西,你能确定它不是幻觉吗?
Exactly. Yeah. Okay, but then that does make me wonder. So move 37 as an example of where it went beyond what humans were able to do. At the same time, when move 37 first came through, people thought it was a mistake, right? So how can you tell the difference? I mean, if the algorithm comes up with something that is original, can you be sure it's not a hallucination?
是的,我认为这是一个非常重要的点。对于大型语言模型,尤其是在最初版本开发时,它们会产生幻觉。它们会提出不正确的解决方案或完全无效的回应。这就是智能体框架的重要性所在,你将大型语言模型与验证器结合,验证器能够剔除哪些是幻觉,哪些是真正可能值得进一步研究的非凡结果。
Yeah, and I think this is a very important point. With large language models, especially when they were being developed initially in the first versions, they would hallucinate. They would come up with solutions which were not correct or responses which were completely invalid. This is where the importance of the agent harness comes into play, where you couple the large language model with a verifier, which is able to prune out when what is being hallucinated and what is actually something that might be remarkable that we need to investigate further.
但如果这些大型语言模型基于人类数据,你们会不会有局限在人类已发现之物的风险?我指的是教科书里已有的内容。
But then if these large language models are based on human data, is there a danger of you limiting yourselves to what humans have already discovered? I'm thinking of what's already in the textbook as it were.
当我们构建这些智能体时,我们故意增加了它们需要探索的内容。所以我们告诉模型,你必须超越你训练时的分布,你应该自由地探索更多。事实上,你可能会产生一些可能不合适或不正确的新东西,但我们有那个验证器和评估函数来剔除那些见解。我认为这正是卡尔·波普尔描述整个科学过程的方式。《猜想与反驳》是著名的文章,猜想可能就是幻觉。这是产生合理假设的能力,而反驳则是过滤掉错误、无效东西的步骤。我认为这也说明了为什么当前 AI 能力格局是这样的。也就是说,它在可验证领域非常出色。代码是可验证的领域。你定义目标,你可以为代码编写测试。第一个测试是它能编译,这已经是个好迹象。然后你在那些测试上测试它,但你有硬性标准来拒绝失败。这对这类任务非常重要。如果没有它,事情就变得棘手得多。例如,如果你研究开放的科学问题,你可能没有验证器告诉你这是对还是错。最终,所有东西都要通过实验,物理实验将是你需要的验证。
When we build these agents, we deliberately increase the amount of things that they have to explore. So, we tell the models that you have to go beyond the distribution that you were trained on and you should feel free to explore more. And in fact, you might produce new things which might not be appropriate or might not be correct, but we have that verifier and evaluation function to prune out those insights. I think this is really how Karl Popper would also characterize the whole scientific process. Conjecture and refutation is the famous essay, and conjecture is maybe hallucination. It's this production capability of producing plausible hypotheses, and then refutation is the step by which you filter out the things that are wrong, that don't work. And I think it also makes clear why the current AI capability landscape looks like it does. Namely, it is very good in verifiable domains. Code is a verifiable domain. You define the objective, you can write down tests for the code. The first test is that it compiles, which is already a good sign. Then you test it on those tests, but you have hard criteria to reject failure. Which is super important for these kinds of tasks. If you don't have it, things become much trickier. For example, if you work on open scientific problems, you might not have a verifier who can tell you that this is right or this is wrong. Ultimately, all from experiment, physical experiment will be the verification that you need.
没错,但这还有很长的路要走,不是吗?我指的是实验部分。因为我这里在想可解释性,回到你之前提到的点。考虑到风险比围棋棋盘上高得多,最终得到不易解释的结果有关系吗?
Right, but that's quite a long way down the road, isn't it? I guess the experimental part of it. Because I'm just wondering here about interpretability, coming back to the point that you made earlier. Does it matter that you might end up with results that are not easily interpretable here, given that the stakes are so much higher than they are on a board of a Go game?
是的,我认为确实有关系。科学也关乎沟通,对吧?如果你能提出新的见解,但如果你无法沟通,人们无法在此基础上发展,那么能产生的影响就有限。所以可解释性扮演着非常重要的角色。但它不是唯一的东西。以 AlphaFold 为例。AlphaFold 能够解决蛋白质结构预测这个惊人的问题。我们完全理解它所做的概念性操作吗?在机制层面,是的,但我们不完全了解可以用来重现人类级推理过程以做出相同预测的底层理论。我们需要以某种方式将它们转化为人类可消化的形式,让有限理性的人类心智能够理解。
Yeah, I think it does matter. Science is also about communication, right? If you can come up with this new insight, but if you are not able to communicate and people are not able to build on top of it, then there are limits to what impact will be achieved. So interpretability plays a very important role. But it's not the only thing. Take the example of AlphaFold. AlphaFold is able to solve this amazing problem of protein structure prediction. Do we understand completely the conceptual operations that it does? At the mechanistic level, yes, but we don't know completely the underlying theory that can be used to recreate a human-level reasoning process to make the same predictions. And we will somehow need to convert them to a human-digestible form that the bounded rational human mind will be able to comprehend.
我认为这里有一个非常有趣的点,即解释不仅需要说明你正在解释的现象,还需要考虑解释接收者的智力水平。所以,有时在 YouTube 上你可以看到这些东西以六岁、八岁、十岁、十二岁孩子的水平来解释。我不得不说,我很喜欢十二岁孩子的解释。这反映了这个事实。解释真的是现象与我们理解能力之间的桥梁。所以,很可能未来的 AI 系统会提出对它们来说可能显得简单的解释,但对我们来说正好能跟上 AI 系统,对吧?
I think there's a really interesting point there, which is that an explanation not only needs to account for the phenomenon that you're explaining, it also needs to account for the intellectual level of the recipient of the explanation. So, sometimes on YouTube you can see these things explained at the level of a six-year-old, an eight-year-old, a 10-year-old, a 12-year-old. I quite like the explanations for 12-year-olds, I have to say. And that reflects this fact. An explanation really is a bridge between the phenomenon and our capacity to understand it. So, it may very well be the case that future AI systems come up with explanations that might seem simplistic to them, but that are just about right for us to keep up with the AI system, right?
我的意思是,如果你看看我们的智能体,比如 AlphaProof,它们能做的事情是,你给它们开放的数学问题,它们会给你一个证明。那个证明是可验证的。
I mean, if you look at our agents like AlphaProof, what they are able to do is you give them open math problems, and they will give you a proof. And that proof is verifiable.
嗯。你可以判断它是否正确。
Mhm. You can tell whether it's correct or not.
正是。即使你不理解它。
Exactly. Even if you don't understand it.
是的,你可能不理解它,但你知道它是正确的,对吧?关于原始定理是否正确的不确定性现在已经解决了。但我们完全理解它吗?事实上,到目前为止,我们已有的结果,我们花了精力,然后将这些结果转化为数学家能够看到并说‘是的,这说得通。我实际上可以把它翻译成英文,而且一切正常’的形式。但由此产生了两个关键现象。一个是问题框架的重要性现在上升了。
Yeah, you might not understand it, but you know it's correct, right? The uncertainty about whether the original theorem was correct or not is now resolved. But do we completely understand it? In fact, till now the results that we have had, we have spent the effort and then converted those results into a form that mathematicians have been able to see and say, 'Yes, it makes sense. I can actually translate it into English, and it all works.' But there are two key phenomena that come out of it. One is that the importance of framing the problem now rises.
因为如果你不这样做,当我们试图解决这些非常难的数学问题,当我们给智能体这些难题时,挑战之一就是准确地指定问题,以便智能体能够理解它需要优化的奖励函数是什么。然后一旦它找到解决方案,还有一个挑战是如何将解决方案转换回人类可读的形式。不过,如果我们真的到了算法可以自己提出证明的地步,那么数学家在这其中还有什么角色呢?自私地说?
Because if you don't, one of the challenges when we are trying to solve these very hard math problems, when we are giving the agent these hard problems, is to specify the problem accurately so that the agent can understand what is the reward function that it needs to optimize for. And then once it finds a solution, there is a challenge of actually converting the solution back to a human-readable form. If we do get to a point though where an algorithm could just come up with its own proof, where's the role for mathematicians in all of this, speaking selfishly?
不,我认为数学家今天甚至更重要。因为这些智能体能够解决这些不可思议的问题。但哪些问题需要解决?你如何指定那个问题?这就是数学家和科学家发挥作用的地方。不过,我确实喜欢这样一个想法:有一天,也许黎曼假设会回来并说,‘是的,有一个证明。不幸的是,它超出了任何人类的理解能力。所以,你知道,抱歉了。’
No, I think mathematicians are even more important today. Because what these agents are able to do is they're able to solve these incredible problems. But what are the problems that need solving? How do you specify that problem? That's where mathematicians and scientists come in. I do like the idea though that one day there might be, I don't know, Riemann hypothesis and it comes back and says, 'Yes, there's a proof. Unfortunately, it's beyond any human's ability to understand it. So, you know, sorry about that.'
但实际上,我是说,我有点开玩笑,但如果我们在这里谈论推进科学知识和理解超越人类已经做到的,你认为你已经看到科学中的“第 37 手”例子了吗?
But actually, I mean, I'm joking slightly, but if we are talking here about advancing scientific knowledge and understanding beyond what humans have done, do you think you've seen examples of move 37 in science already?
是的,我绝对认为有。就拿矩阵乘法算法这个例子来说,这是人们研究了很多很多年的东西,但我们却能够提出一个新的算法。所以这确实是算法发现中的一个“第 37 手”时刻。我认为我们现在在科学的许多其他领域也看到了同样的事情,在数学、材料科学中,提出了我们认为现在稳定的新结构。所以有很多这样的事情,但最初的“第 37 手”时刻仍然非常相关,因为它在某种意义上是最早的,并且带来了超越人类理解的概念。
Yeah, I think absolutely. I think just the example of the matrix multiplication algorithm, it is something that people had studied for many, many years and yet we have been able to come up with a new algorithm. So that is genuinely a move 37 moment in algorithmic discovery. And I think we are now seeing the same thing in many other areas of science, in mathematics, in material science, coming up with new structures that we think now are stable. So there are a number of these things, but the original move 37 moment is still very relevant because it was in some sense the first and it brought about that concept of going beyond human understanding.
我在这里再次想到 AlphaZero,以及它如何真正摆脱了人类数据并展示了这些深刻的结果。另一方面,大型语言模型最终几乎成了智能的捷径,我想,这在很大程度上是基于人类数据的。这对你来说是一种令人惊讶的转折吗?
I am thinking here about AlphaZero again and how that really moved away from human data and showed these profound results. Large language models, on the other hand, ended up being almost a shortcut to intelligence, I guess, that was based very much on human data. Was that a sort of surprising turn of events for you?
是的,我认为这是我们观察到的一个有趣的事情。DeepMind 基于这样一个理念:我们将游戏作为现实世界的缩影。DeepMind 的理念是将智能体置于这些环境中,让它们学习如何掌握这些环境,从而增长它们的智能。然后大型语言模型发生的事情实际上是发现了一条捷径。不知何故,有大量结晶化的智能,如果你愿意这么说的话,以数据的形式存储在互联网上。首先是文本数据,也许还有图像、视频等等。而这条捷径实际上是首先挖掘所有这些数据,并基于此训练系统。这基本上就是第一代和第二代大型语言模型的基础。但当然,你到了一个点,首先,这不会带来新颖性。你现在处于现有人类知识的语料库中,我们知道这些模型在其中的能力有多强。但现在很难从中跳出来。我们如何超越我们已经知道的?我认为这就是过去几年社区再次探索 DeepMind 早期开创的方法的地方。当然还有环境中的强化学习。现在后训练的一部分常规上是强化学习的形式,要么基于人类生成的数据,要么基于环境中的问题,比如编码环境等等。所以现在我们正处于一个再次超越人类知识的时期。
Yes, I think that is an interesting thing that we observed. DeepMind was based on this idea that we use games as a microcosm of the real world. And the philosophy of DeepMind had been to place agents within these environments and let them learn how to master them and thereby grow their intelligence. And then what happened with large language models was really this discovery that there's a shortcut. That somehow there's this huge amount of crystallized intelligence, if you like, stored in the form of data on the internet. First text data, maybe images, maybe videos, and so on. And that the shortcut is really to first mine all of that data and train systems based on that. And that's basically the first and second generation of large language models that are based on that. But then of course, you come to the point where first of all, that doesn't lead you to novelty. You're now within this corpus of existing human knowledge and we know how competent these models are within that. But it's very difficult to get out of that now. How do we go beyond what we already know? And that's I think where now the community for the past few years is exploring the methods again that DeepMind pioneered early on. And there's of course reinforcement learning in environments. Part of the post training now is routinely forms of reinforcement learning and either on human generated data or also on problems on environments like coding environments and so on. And so now we're in a period where we're going again beyond human knowledge.
Pushmeet,你认为如果没有 AlphaGo,我们会处于人工智能革命的这个时刻吗?
Pushmeet, do you think that we would be here at this moment in the AI revolution if it hadn't been for AlphaGo?
我认为 AlphaGo 就是那个转折点,它变得非常清楚:我们在特定领域超越人类智能的转变时刻不是科幻小说,也不是几十年后。它正在发生。如果它能在围棋中发生,那么它没有理由不能在蛋白质结构预测、核聚变、材料科学中发生。那场比赛、第 37 手以及那次经历的遗产,就是我们今天所生活的。
I think AlphaGo was that transition point where it became very clear that the moment of transition where we go beyond human level intelligence in particular areas is not science fiction or many decades later. It is happening now. And if it could happen in the game of Go, there was no reason why it couldn't happen in protein structure prediction, in fusion, in material science. And the legacy of that match and move 37 and that experience is what we are all living in now.
我认为这实际上是结束这一集的好点。老实说,是 Demis 推了我一把。非常感谢你加入我。太棒了。是的,很荣幸。
I think that's a great point to end the episode actually. To be honest with you, Demis pushed me. Thank you so much for joining me. Amazing. Yeah, pleasure.
这些人类与机器故事中的重大范式转变时刻以前就发生过,但关于国际象棋,它始终只是一个计算问题。机器能通过蛮力取胜吗?AlphaGo 不同。这是机器第一次展示出更深层次的东西,一种真正的智能,将直觉与计算相结合,并将我们带到了超越人类能力的地步。现在,距离 AlphaGo 比赛已经过去了 10 年,这个领域以令人难以置信的速度发展。但当时困扰研究人员的许多问题现在比以往任何时候都更相关。你如何创建超越人类知识并能够产生新见解的 AI 系统?你如何将真正的新见解与幻觉区分开来?你一直在收听 Google DeepMind 播客,我是 Hannah Fry。今年我们还有更多集要发布。所以,请确保订阅我们的 YouTube 频道。我们很快再见。
These big paradigm-shifting moments in the story of humans and machines have happened before, but the thing about chess is that it was always just a question of calculation. Can a machine brute-force its way to a victory? AlphaGo was different. It was the first time that a machine had demonstrated something deeper, a genuine intelligence that combined intuition with calculation and took us beyond human capability. Now, 10 years on from the AlphaGo match, the field has moved at an incredible pace. But many of the questions that preoccupied researchers then are more relevant now than ever. How do you create AI systems that go beyond human knowledge and are capable of new insights? And how do you separate the genuinely new insights from hallucinations? You have been listening to Google DeepMind the podcast with me, Hannah Fry. We have got plenty more episodes to come out this year. So, please make sure you subscribe to our YouTube channel. We'll see you soon.