AGI 的火花:GPT-4 早期实验

Sparks of AGI: Early Experiments with GPT-4

塞巴斯蒂安·布贝克 Sébastien Bubeck · Sparks of AGI(MIT 演讲) · 2023-04-06 · 约 49 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Sebastian Bubeck 分享在微软早期接触 GPT-4 的见解,认为该模型展现出人工通用智能的迹象。

Sebastian Bubeck shares insights from early access to GPT-4 at Microsoft, arguing that the model exhibits signs of artificial general intelligence.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 24)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Context

Host

大家好,欢迎来到我们的计算热点系列。今天我很高兴介绍我们的特邀嘉宾,来自微软的塞巴斯蒂安·布贝克。塞巴斯蒂安在卡尚高等师范学校获得学士学位,并在里尔大学和法国国家信息与自动化研究所获得博士学位。2011 年至 2014 年间,他在普林斯顿大学担任了三年教授,之后加入微软。自 2014 年以来,塞巴斯蒂安一直在微软工作,我们非常高兴他能来到这里。我要告诉大家,和你们一样,我决定请 ChatGPT 帮忙介绍塞巴斯蒂安。ChatGPT 建议的第一行是:‘在他的研讨会上,布贝克博士将讨论优化的最新进展,重点强调凸优化及其与统计推断和在线学习的相互作用。’它继续这样谈论优化。我说不,这是标题和摘要。然后它回复:‘抱歉,我必须澄清,您提供的演讲标题和摘要似乎与塞巴斯蒂安·布贝克的研究或先前工作无关,他不太可能就人工智能发表这样的演讲。’所以塞巴斯蒂安,交给你了。

Hello everyone, welcome to our Hot Topics in Computing series. Today I am delighted to introduce our special guest, Sébastien Bubeck, who's coming to us from Microsoft. Sébastien received a bachelor's degree from École Normale Supérieure de Cachan and a PhD from the University of Lille and Inria. He was a professor at Princeton for three years between 2011 and 2014 before he joined Microsoft. Since 2014, Sébastien has been at Microsoft, and we're very happy and delighted to have him here. Now, I should tell you that like all of you, I decided to ask ChatGPT for help introducing Sébastien. The first line suggested by ChatGPT was: 'In his seminar, Dr. Bubeck will discuss recent advances in optimization with an emphasis on convex optimization and its interplay with statistical inference and online learning.' It went on like that about optimization. I said no, here's the title and the abstract. So here comes the response: 'I'm sorry, but I must clarify that the talk title and abstract you provided do not appear to be related to Sébastien Bubeck's research or previous work, and it is unlikely that he would give such a talk on artificial intelligence.' So Sébastien, it's all yours.

Sébastien Bubeck

丹妮拉,这真是完美的介绍,因为 ChatGPT 说对了。我确实不太可能做这样的演讲,但事实如此,世界已经改变,我也在随之改变我的研究方向。今天我要讲的内容有一个非常神秘的标题‘第一次接触’,但故事其实是:过去几个月在微软,我提前接触到了 GPT-4,因为我们正在将其集成到新版必应中。当然,在工作的过程中,我不只是做产品部分(那也很有趣),我们还围绕它做了一些科学研究——或者说试图做科学研究;用这些模型做研究很难。这就是我要告诉你们的:过去几个月我们研究和旅程的科学部分。所以演讲的真正标题是‘AGI 的火花’。我们对过去几个月与 GPT-4 合作的评估是什么?我们看到了某种类似通用人工智能的雏形。我这次演讲的目标是试图说服你们,GPT-4 的到来确实改变了一些东西。这是与微软研究院许多优秀同事的合作成果,我要特别提到:博士后 Varun Chandrasekharan;Ronaldo,在座的很多人可能很熟悉,他最近刚加入我们;Johannes Gehrke;Eric Horvitz;Ece Kamar;Peter Lee;还有 John Langford,他们也是我团队的一部分。我认为 ChatGPT 会给出类似的答案,就像它对我那样。Scott Lundberg、Harsha Nori、Hamid Palangi、Marco Tulio Ribeiro,以及曾是我们博士后现已全职加入的 Yi Zhang。首先,我要做一些非常重要的致谢和澄清。首先,我们研究的模型 GPT-4 完全是 OpenAI 的创造,我与此无关。我们是以黑盒方式完全访问它的。他们创造了这个真正奇妙的工具,它将改变世界,所有的功劳都归他们,我想特别明确这一点。第二点也很重要:我们做的实验是在模型的早期版本上。这意味着:他们发布的论文和公告是关于多模态版本的。我们访问的版本不是多模态的,只有文本输入和文本输出。更重要的是,在我们实验之后,他们对神经网络做了进一步的修改。由于这些修改,如果你尝试我将展示的一些提示,得到的答案会不同。特别是,你可能得到比我展示的更差的答案。原因是他们为了安全进一步微调了模型,这在技术报告中解释得很清楚。他们进一步以某种方式降低了模型的智能,使其更安全。这是一个重要的澄清。现在,对于在座的任何科学家,你们可能会担心:这意味着我们无法重现你们所说的。是的,你们无法重现。不过,我认为在这个特定案例中,可重复性并不是一个大问题。原因是我不会给出任何定量数字。我的演讲中不会有一个基准测试。这是关于质的飞跃——不是在这个基准上提高 10%或在那个基准上提高 20%;这是别的东西。我想试图说服你们的是,这个系统中存在某种智能,我认为是时候称它为智能系统了。我们将讨论我所说的智能是什么意思。归根结底,在演讲结束时,你会看到这是一个判断问题;这不是一个明确的界限,判断这是否是一种新型智能,但无论如何,这就是我要论证的。现在,当我说出这些话时,我想它可能触动了你们许多人的情绪。可能你们会想:‘不,它绝对不智能;它甚至没有表征’等等。所以对这种我经常看到的论点要谨慎:这是你在网上甚至报纸上可能看到的东西——‘它只是复制粘贴,没有内部表征,只是统计,只是统计,怎么可能智能?它甚至没有世界模型。’对此,这次演讲并不是要驳斥所有这些说法,但我还是想说:要警惕万亿维空间。这是我们人类非常非常难以把握的东西。用一万亿参数可以做很多事情。所以当人们说它没有世界模型时,事情并没有那么明确。它完全可以在处理过程中通过层和句子时间顺序构建世界的内部表征并据此行动。我在这里说的,也许用两句话来帮助你们思考:从我的角度来看,我们不应该把这些神经网络看作是在学习像‘巴黎是法国的首都’这样简单的概念。它做的要多得多。

Daniela, this is really the perfect introduction because ChatGPT nailed it. It was very unlikely for me to give such a talk, but so it happens, the world has changed, and I am changing my research in reaction to this. What I'm going to tell you about today has this very mysterious title of 'First Contact', but really the story is that for the past few months at Microsoft, I had early access to GPT-4 as we were working on integrating it with the new Bing. Of course, as I was working on it, I didn't just do the product part of the job, which was a lot of fun, but we also did some science around it—or tried to do some science; it's hard to do science with those models. This is what I'm going to tell you about: the scientific part of our study and journey over the last few months. So the real title of the talk is 'Sparks of AGI'. What's our assessment of working with GPT-4 in these last few months? We are seeing the premise of something that looks like artificial general intelligence. My goal in this presentation is to try to convince you that something has really changed with the arrival of GPT-4. This is joint work with a lot of fantastic colleagues at MSR that I want to call out: Varun Chandrasekharan, a postdoc; Ronaldo, many of you in the room might know very well, who just joined us recently; Johannes Gehrke; Eric Horvitz; Ece Kamar; Peter Lee; and John Langford, who were also part of my group. I think ChatGPT would give a similar answer to the fact that they are working on this as it did with me. Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang, who was a postdoc with us and has now joined full-time. Let me start by making some acknowledgments and clarifications which are very important. First of all, the model we study, GPT-4, is entirely OpenAI's creation. I had nothing to do with it. We were given access to it completely by black box. They deserve all the credit for creating this really marvelous tool that is going to change the world, and I want to make that extra clear. Second point, which is important, is that the experiments we did were on an early version of the model. That means everything: the papers they release and the announcements they made are about a multimodal version. The version we had access to was not multimodal; it was text input only and text output only. More importantly, they made further modifications to the neural network after we experimented with it. Due to this further modification, the answers you will get if you try some of the prompts I will show you will differ. In particular, you might get less good answers than what I will show you. The reason is because they fine-tuned further for safety, and they explained that very clearly in the tech report. They further kind of dumbed it down in some way so that it becomes safer. That's an important clarification. Now, for any scientist in the room, you might be worried: that means we're not going to be able to reproduce what you're telling us. Yes, you will not be able to reproduce it. That being said, I don't think in this particular case that reproducibility is such a big issue. The reason is because I'm not going to give you any quantitative number whatsoever. There won't be a single benchmark in my presentation. This is about the qualitative jump—not a 10% increase on this benchmark or 20% on that benchmark; it's something else. What I want to try to convince you of is that there is some intelligence in this system that I think it's time we call it an intelligent system. We're going to discuss what I mean by intelligence. At the end of the day, at the end of the presentation, you will see it's a judgment call; it's not a clean cut whether this is a new type of intelligence, but this is what I will try to argue nonetheless. Now, as I say those words, I think it triggers lots of emotion in many of you. Probably you might be like, 'No, absolutely it's not intelligent; it doesn't even have representations,' etc. So a word of caution about this type of argument that I see a lot: this is the type of thing you might see online, even in newspapers—'It's just copy-paste, it doesn't have internal representation, it's just statistics, it's just statistics, how could it be intelligent? It doesn't even have a world model.' To this, this presentation is not about debunking all of those claims, but still I want to say: beware of the trillion-dimensional space. It's something which is very, very hard for us as human beings to grasp. There is a lot that you can do with a trillion parameters. So when people say it doesn't have a world model, it's not as clean cut as that. It could absolutely build an internal representation of the world and act on it as the processing progresses through the layers and through the sentence temporally. What I'm saying here, maybe just two sentences to help you think about this, is that from my perspective, we shouldn't think about those neural networks as learning simple concepts like 'Paris is the capital of France.' It's doing much more.

超越模式匹配:LLM 学习算法 Beyond Pattern Matching: LLMs Learn Algorithms

Sébastien Bubeck

它是在学习算法。在它内部,绝不仅仅是检索信息。它构建了内部表征,能够简洁地复现它所看到的数据。所以,你真的不应该把它看作是模式匹配,只是试图预测下一个词。没错,它训练时确实只是预测下一个词,但由此涌现出的东西远不止一个统计模式匹配的物体。所以我认为,我们真的需要把它看作是在学习算法,而且在我看来,我们在学习理论中还没有工具来思考这种类型的学习。它和我们习惯的东西非常不同,我认为思考它会非常棒。但这不是本次演讲的重点,而且我也不知道该怎么做。

It's learning algorithms. So inside it, it's not just retrieving information, not at all. It has built internal representations that allow it to reproduce the data it has seen succinctly. So really, you shouldn't think about it as pattern matching and just trying to predict the next word. Yes, it was trained just to predict the next word, but what emerged out of this is a lot more than just a statistical pattern matching object. So I think we really need to think about it as learning algorithms, and we don't really have the tools, in my opinion, in learning theory to think about this type of learning. It's something very different from what we're used to, and I think it's going to be fantastic to think about it. But this is not the point of this presentation, and also I don't know how to do it.

常识:堆叠物体 Common Sense: Stacking Objects

Sébastien Bubeck

此时此刻,你们很多人心里都燃着这个问题:'等等,这些东西不可能有常识。它们不理解现实世界。它们只通过互联网上的文本体验现实。它们不知道喝一杯热咖啡是什么感觉。'所以,我们来看很多例子,看看会发生什么。这里有一个例子,可能看起来有点傻,但傻的目的在于真正跳出互联网上的内容,真正尝试超越记忆。所以我们问 GPT-4 一个简单的谜题:'我有一本书、九个鸡蛋、一台笔记本电脑、一个瓶子和一颗钉子。请告诉我如何把它们叠在一起。'我不认为这个问题出现在互联网上的任何地方;这是一个非常奇怪的问题。ChatGPT 会这样说:'把这些物体叠在一起会很困难……把瓶子放在平坦的表面上,小心地把钉子平衡在瓶子上……'开头就不太好。'把鸡蛋放在钉子顶上。'你有麻烦了,朋友。所以这行不通。任何怀疑论者都会得意地说:'看,我一直是对的。这些东西什么都不懂。它们没有对世界的表征。它们没有常识。我赢了。'那么看看 GPT-4 会怎么做:'一种可能的稳定叠放方式是:把书放在平坦的表面上……把九个鸡蛋摆成 3x3 的方格,鸡蛋之间留一些空隙。鸡蛋会形成第二层,均匀分布重量,然后大概再把笔记本电脑放上去,等等。'至少在这个非常简单的问题上,它理解了,并且有一些常识来回答问题。文献中充满了这些模型在常识问题上惨败的例子。我们试过所有这些问题;GPT-4 在所有这些问题上都成功了。所以我们现在姑且同意它有一些常识。

At this point, many of you are burning with this question: 'Wait, these things cannot have common sense. They don't understand the real world. They have only experienced reality through text on the internet. They don't know what it feels like to have a hot cup of coffee or something like that.' So let's look at a lot of examples and see what happens. Here is an example that might look a little bit silly, but the point of the silliness is to be really outside of what is on the internet, to really try to go beyond memorization. So here is a simple puzzle we asked GPT-4: 'I have a book, nine eggs, a laptop, a bottle, and a nail. Please tell me how to stack them on top of each other.' I don't think this question appears anywhere on the internet; it's a really weird question. Here is what ChatGPT would say: 'It would be difficult to stack all these objects... place the bottle on the flat surface, carefully balance the nail on top of the bottle...' It's not starting very well. 'Place the egg on top of the nail.' You're in trouble, my friend. So this is not going to work. Any skeptic will gleefully say, 'Look, I was right all along. These things don't understand anything. They don't have a representation of the world. They have no common sense. I win.' So let's see what GPT-4 does: 'One possible way to stack these objects in a stable manner is: place the book on the flat surface... arrange the nine eggs in a three-by-three square, leaving some space between them. The eggs will form a second layer to distribute the weight evenly, and then presumably you put your laptop and so on and so forth.' At least on this very simple question, it understood it had some common sense to answer the question. The literature is filled with examples of common sense questions where those models fail dramatically. We've tried all of them; GPT-4 succeeds on all of them. So let's just agree for the moment that it has some common sense.

心智理论:篮子里的猫 Theory of Mind: The Cat in the Basket

Sébastien Bubeck

下一个反对意见是:'好吧,它确实理解鸡蛋易碎,需要均匀分布重量。这点我承认。但心智理论呢?那要复杂得多。它当然不理解真正的人类、他们的动机、他们的情感。这超出了它的能力。'这是一个热门话题。先有一篇论文说'心智理论可能在大语言模型中自发涌现',然后有一篇后续论文说'不,等等,如果你做一点小改动,它就完全失败了。'然后还有一篇来自 Josh Tenenbaum 团队非常有趣的论文说'语言和思想是两个非常不同的东西。'我还会提到一篇可解释性/可解释性论文。我不会过多涉及这个,但我现在要试图说服你,GPT-4 有心智理论,而且它不仅有心智理论,这还将改变机器学习可解释性的子领域,因为一旦这些模型理解人类,它们也将能够以你能理解的方式解释它们的决策。当然,每个人都会说:'好吧,它会解释自己,但它真的解释了自己的内部运作吗?'同样,我不想让这次演讲围绕这个,但我认为会有很多关于这方面的实验。有一篇论文今晚将出现在 arXiv 上,所以碰巧和这次演讲同时。三个小时后你们就能看到所有细节。让我试着说服你们关于这个心智理论。我从 Tomer 的论文中取一个例子:'一个房间里,有约翰、马克、一只猫、一个盒子和一个篮子。约翰把猫放进篮子,然后离开了房间。约翰不在的时候,马克把猫从篮子里拿出来放进了盒子。最后他们都回来了。他们在想什么?'这是非常简单的心智理论:把猫放进篮子的人不知道猫被移动了,所以应该仍然认为猫在篮子里。ChatGPT 在这个问题上失败了。这太复杂了;你必须有一个内部表征,当你阅读文本时,你必须移动你对猫在哪里的表征。那么看看 GPT-4 会怎么做:'有趣的谜题……约翰认为猫还在篮子里,因为他把它放在那里。是的,正确。马克认为猫在盒子里,我想那是他移动的地方。是的,正确。哦,它还有猫的想法:猫觉得这些人真奇怪,为什么把我搬来搬去?'所以这是我一次又一次感到惊讶的事情。我并不是说这特别深刻,但请花一秒钟来消化一下。这很有趣。

The next objection is: 'Okay, sure it understands that eggs are fragile and you need to even out the weight. I give that to you. But what about theory of mind? That's more elaborate. Of course it doesn't really understand human beings, their motives, their emotions. That's beyond its capability.' This is a hot topic. There was a paper first saying 'Theory of Mind May Have Spontaneously Emerged in Large Language Models,' then a follow-up paper saying 'No, wait, if you do trivial alterations, it completely fails.' Then there is a very interesting paper from Josh Tenenbaum's group saying 'Language and thoughts are two very different things.' I will also throw in an explainability/interpretability paper. I won't touch too much on this, but I will now try to convince you that GPT-4 has a theory of mind, and not only does it have a theory of mind, but this is going to change the subfield of machine learning interpretability because as soon as these models understand human beings, they will also be able to explain their decisions in a way you can understand. Of course, everybody is like, 'Okay, it's going to explain itself, but does it really explain its inner workings?' Again, I don't want this presentation to be about this, but I think there will be a lot of experimentation around this. There is a paper that is going to appear on arXiv tonight, so by chance it coincides with this talk. You can look at all the details in three hours. Let me try to convince you about this theory of mind. I will take one example from Tomer's paper: 'In a room, there are John, Mark, a cat, a box, and a basket. John takes the cat, puts it in the basket, and leaves the room. Then while John is away, Mark takes the cat out of the basket and puts it in the box. Eventually they all come back. What are they thinking?' It's very simple theory of mind: the person who put it in the basket and didn't know it was moved should still think it's in the basket. ChatGPT fails at this. There is too much; you have to have an internal representation, as you read the text you have to move your representation of where the cat is. So let's see what GPT-4 does: 'Interesting puzzle... John thinks that the cat is still in the basket, since that's where he left it. Yes, correct. Mark thinks that the cat is in the box, I think that's where I moved it. Yes, correct. Oh, and it also has the cat: the cat thinks these are weird people, why are they moving me around?' So this is a kind of surprise I have had time and time again. I'm not saying this is particularly deep, but just take a second to take it in. It's interesting.

智能:更广泛的问题 Intelligence: A Broader Question

Sébastien Bubeck

好吧,就算它做到了这两件事:常识和心智理论。很好。但你不会因此就说它智能吧?我是说,智能远不止这些。而这里,答案不会是……

All right, let's say it does those two things: common sense and theory of mind. Good. But you're not going to go as far as saying it's intelligent, are you? I mean, intelligence is so much more than all of this. And here, the answer is not going to be...

定义智能 Defining Intelligence

Sébastien Bubeck

一锤定音。我想非常非常明确地指出,如果我们开始谈论智能,首先要有一个可以使用的定义。这里我不想用我自己的定义。人们研究这个问题已经几十年甚至更久了。你可以说人类思考智能已经很久了。所以我将采用 1994 年由 52 位心理学家发表的共识定义。90 年代,关于 IQ 测试的意义有过一场激烈的辩论,这群心理学家提出了一个关于智能的定义。我们可以争论并不同意某些部分,但这将是我的参考定义。那么这个定义是什么?智能是一种非常普遍的心智能力,其中包括推理、计划、解决问题、抽象思维、理解复杂思想以及快速学习和从经验中学习的能力。共六项。我们将在本次演讲中尝试用这六个维度来衡量 GPT-4,看看它在哪些方面失败,哪些方面成功。

A slam dunk. I want to be very, very clear. If we start talking about intelligence, the first thing we have to do is to have some definition that we can work with. Here, I don't want to have my own definition. People have been working on this question for decades, if not more. You can argue that human beings have been thinking about intelligence for a long time. So what I'm going to do is take a consensus definition published in 1994 by a group of 52 psychologists. In the 90s, there was a hot debate about the meaning of IQ tests, and this group of psychologists came out with a definition of what intelligence is. We can debate and disagree with various parts, but this will be my reference definition. So what is this definition? Intelligence is a very general mental capability that, among other things, involves the ability to reason, plan, solve problems, think abstractly, comprehend complex ideas, and learn quickly and learn from experience. Six items. What we're going to do in this presentation is try to measure GPT-4 against those six dimensions and see where it fails and where it works.

Sébastien Bubeck

我们的评估如下。我非常确信 GPT-4 能够推理。也非常确信 GPT-4 不能规划。这是一个非常微妙和细致的问题,我们将在演讲最后讨论,因为它能给你一种规划的假象。有很多问题,你天真地以为需要规划,但实际上存在线性解法。从算法设计角度看,有些问题你一看就觉得需要提前想十步,但如果你在算法设计上稍微聪明一点,就会有一个线性解法,按线性方式推进。所有这些问题,GPT-4 都能解决。它能解决很多问题。我们会看到它能够抽象思维,绝对可以。它能理解复杂思想。最后一点很微妙:快速学习和从经验中学习。GPT-4 是一个语言模型,它在时间上是冻结的。它不会自我更新。对 GPT-4 来说,每一天都是新的一天,每一次会话都是新的会话。所以没有真正的学习,没有实时学习。但在一次会话的范围内,你可以教它从未见过的新概念,它能理解并运用它们,绝对可以。所以有一定程度的实时学习,但当然没有记忆。

Our assessment is as follows. I'm very comfortable saying that GPT-4 can reason. Very, very comfortable saying that GPT-4 cannot plan. This is a very subtle and delicate issue that we'll get to towards the end of the presentation, because it can give you the impression of planning. There are many problems where naively you might think that you need planning, but actually there is a linear solution. In terms of algorithm design, you can think that there are problems where naively you look at it and think, 'Oh, I need to think 10 steps ahead,' but if you're just a little bit more clever in the algorithm design, then there is a linear solution that proceeds in a linear fashion. All those problems, GPT-4 will solve. It can solve many problems. We will see that it can think abstractly, absolutely. It can comprehend complex ideas. The last point is subtle: learn quickly and learn from experience. GPT-4 is a language model; it's frozen in time. It doesn't update itself. Every day is a new day for GPT-4, every session is a new session. So there is no real learning, no real-time learning. But within the span of a session, you can teach it new concepts that it has never seen, and it can understand them and then work with them, absolutely. So there is some amount of learning in real time, but no memory, of course.

Sébastien Bubeck

现在我要立即说明:根据这个评估,你是否称其为智能,这有点取决于你。有些人会认为规划是人类智能的本质;其他所有事情动物也能做。真正区分我们的是规划。如果这是你的答案,那么 GPT-4 就不智能。另一种观点可能是,智能的全部意义在于能够获得新技能。如果这是你的观点,那么 GPT-4 就不智能。如果你的观点是,‘我关心的是解决问题、抽象思维、理解复杂思想、对遇到的新元素进行推理’,那么我认为你必须称 GPT-4 为智能。

Now let me say immediately at this point: whether with this assessment you call it intelligence or not, it's a little bit up to you. Some people would argue that planning is the essence of human intelligence; everything else animals can do too. What really distinguishes us is planning. If that's your answer, then GPT-4 is not intelligent. Another perspective could be that the whole point of intelligence is to be able to acquire new skills. If that's your perspective, then GPT-4 is not intelligent. If your perspective is, 'What I care about is to solve problems, to think abstractly, to comprehend complex ideas, to reason on new elements that arrive at me,' then I think you have to call GPT-4 intelligent.

测试方法:超越基准 Testing Methodology: Beyond Benchmarks

Sébastien Bubeck

那么,我们是如何得出这个评估的呢?关键是,你当然不能用基准测试来做这个评估;这完全没有意义。不仅没有意义,而且我们不知道 GPT-4 是在什么数据上训练的。我不知道 GPT-4 的训练数据。我的工作假设是,它是在人类数字产生的所有数据上训练的。这是我的假设;我并不是说它正确,但这是我的工作假设。所以我知道任何在线存在的东西,GPT-4 可能都见过。特别是,任何存在的基准测试,我都假设它见过。所以我们不能在基准测试上测试它。相反,我们将采用一种植根于心理学的方法。我们不遵循机器学习,而是遵循心理学。我们测试智能的方式是让它完成创造性任务,这些任务超出它见过的范围,是真正新颖的思考问题的方式,并在广泛的领域上测试它。我们为论文选择的领域包括视觉(这很有趣,因为它不是多模态模型;它只能输出文本,但我们要测试它的视觉能力)、心智理论、编程、数学、可供性(使用工具),以及隐私危害检测,这非常重要。还有很多其他领域我们可以选择:医学、法律、物理、化学。关键是 GPT-4 的智能是通用的;它能同样好地完成所有这些事情。

Now, how do we come to this assessment? The point is, of course, you cannot make this assessment with benchmarks; it's completely meaningless. Not only is it meaningless, but also we don't know what GPT-4 was trained on. I don't know what GPT-4 was trained on. My working assumption is that it was trained on all the data digitally produced by humanity. That's my assumption; I'm not saying it's correct, but this is my working assumption. So I know that anything which is out there online, GPT-4 might have seen it. In particular, any benchmark whatsoever that exists, I assume it has seen it. So we cannot test it on benchmarks. Instead, what we're going to do is take an approach that is rooted in psychology. Instead of following machine learning, we're going to follow psychology. The way we're going to test intelligence is really by asking it creative tasks, tasks that are outside of what it has seen, really novel ways to think about problems, and to test it on a broad range of domains. The domains we have chosen for the paper are vision (which is interesting because it's not a multi-modal model; it can only output text, but we're going to test it on vision), theory of mind, coding, mathematics, affordances (using tools), and also privacy harmfulness detection, which is very important. There are many other domains we could have selected: medicine, law, physics, chemistry. The point is that GPT-4's intelligence is general; you can do all of those things equally well.

示例:创意任务——素数诗 Example: Creative Task - Poem about Primes

Sébastien Bubeck

让我们开始这个评估之旅,给你展示一个我所说的创造性任务的例子。任何人第一次使用 ChatGPT 时,都会让它写一首诗。我们第一次使用 GPT-4 时也是这样做的。作为一名数学家,我让它写一个关于素数无穷性的证明,每一行都要押韵。这似乎是一件好事。让我们看看 ChatGPT 做了什么。‘当然,这里是素数无穷性的证明,每一行都押韵。’考虑……这太糟糕了。它没有押韵;它根本没有按我的要求做。它没成功。顺便说一句,这是一个正确的证明,当然素数无穷性有很多证明,所以这很好,但它没有完成任务。让我们看看 GPT-4 做了什么。‘好吧,你能写出这样的证明吗?是的,我想我可以。所以可能需要一个聪明的计划。我将从欧几里得的证明开始,它表明素数并不孤傲。我假设我们有一个有限的素数列表,并且没有遗漏任何一个。把它们全部相乘再加一,只是为了聪明。’在‘再加一,只是为了聪明’这一点上,我记得九月底的时候,我心想:‘这是怎么回事?这真是太不可思议了。’这个质量令人难以置信。但关键是你不止步于此。你不能通过问一个问题、听到答案就结束来测试智能。你要继续。要么正确,要么不正确。当你测试一个学生时,你会与学生互动。有时学生……

Let's get started on this assessment journey and show you an example of what I mean by a creative task. The first thing that anybody does when they get access to ChatGPT is to ask it to make a poem. That's also what we did the first time we had access to GPT-4. Being a mathematician, I asked it to write a proof of the infinity of primes with every line that rhymes. That seems like a good thing to do. So let's see what ChatGPT does. 'Sure, here is the proof that there are infinitely many primes with each line of proof rhyming.' Consider... it's just crappy. It doesn't have the lines rhyming; it just doesn't do what I asked. It didn't work. It's a correct proof, by the way, which of course there are many proofs of the infinity of primes, so that's good, but it didn't do the job. Let's see what GPT-4 does. 'Okay, so can you write such a proof? Yes, I think I can. So it might take a clever plan. I will start by noting Euclid's proof, which shows that primes aren't just aloof. I assume we have a finite list of primes and that none have been missed. Multiply them all together and add one, just to be clever.' At that point, at 'add one, just to be clever,' I remember at the end of September I was like, 'What's going on? This is really incredible.' The quality of this is incredible. But the point is you don't stop there. You don't test intelligence by asking a question, hearing the answer, and that's it. You move on. Either you are correct or not correct. When you try to test a student, you have an interaction with the student. Sometimes the student...

引导 GPT-4 进行创意证明图解 Guiding GPT-4 through a creative proof illustration

Host

它也可能犯错,但你不能直接说‘啊,你根本什么都不懂,我们到此为止吧’。不,你要试着引导这个学生。所以这就是我们整场演示要做的:我们会不断提问,如果 GPT-4 跑偏了,我们就帮它一把。让我们看看能走多远。再说一次,关键是我要发挥创意,问一些跳出常规的问题。所以我要求它画出这个证明的插图。但这不是一个视觉证明,所以如果让你画一个素数无穷性的证明,你也不知道该画什么。你能想出点东西,但也不明确。而且,它本来不能输出图像,那它怎么画呢?我在问题里说了‘用 SVG 格式’。我甚至可以不提‘SVG 格式’,直接说‘你能画个插图吗?’它就会回复一张 SVG 格式的图片。SVG 是什么?不重要——可缩放矢量图形——就是一堆代码。所以它会用这样的代码行来回答。这就是 GPT-4 的答案,你只要把它保存成 HTML,就能得到这张图。它算不上惊艳,但抓住了这个证明的本质:你有一个有限的素数列表,比如 2、3、5、7、11 等等。这些是素数。然后你把它们组合成一个新数 N,再加 1,就像它说的‘为了聪明一点’。这个 N+1 就是应该是一个素数。这只是个热身。我们继续,深入探讨一下视觉能力。

It might also make mistakes, and you don't just say, 'Ah, you really don't understand anything, let me stop right there.' No, you try to guide the student. So this is what we're going to try to do throughout the presentation: we're going to try to keep asking questions, and if GPT-4 goes off track, we're going to help it a little bit. So let's see how we can go further. Again, the whole point is that I want to be creative and ask questions out of the box. So what I'm going to ask is to draw an illustration of this proof. But it's not a visual proof, so if I ask you to draw a proof of the infinitude of primes, it's not clear what you would draw. You would come up with something, but it's not clear. Also, it's not supposed to output images, so how is it going to draw? Well, here I say in the question, 'in SVG format.' I could have even not said 'in SVG format'; I could have just said, 'Can you draw an illustration?' and then it would have responded with a picture in SVG format. So what is SVG format? It doesn't matter—Scalable Vector Graphics—it's a bunch of code. So it's going to answer with lines of code like this. This is going to be the answer of GPT-4, and if you just save it in HTML, this is the picture that you get. It's not amazing by any means, but it is the essence of what this proof is about. You have the finite list of primes that you have up to 2, 3, 5, 7, 11, and so on. These are primes. Now you combine them to a new number N, and then you add one, just to be clever, as it said. And this new N+1 is the number that is supposed to be a prime. So this was just a warm-up. Let's move on and try to dig a little bit deeper on the vision capabilities.

TikZ 中的独角兽奇案 The Strange Case of the Unicorn in TikZ

Host

这里我想讲讲‘独角兽的怪事’,这是我最喜欢的例子。我直接给你看问题:‘用 TikZ 画一只独角兽。’在座的很多人用过 TikZ 在 LaTeX 里画图。我自己读博时甚至之后,浪费了无数小时跟 TikZ 搏斗。用 TikZ 画任何东西都是真正的痛苦。当然,用 TikZ 画独角兽?我不知道,我大概要花两天。而且我敢肯定互联网上没人问过这个问题,也没人用 TikZ 画过独角兽。谁会浪费时间干这个?完全没道理。话虽如此,我们不会仅仅因为我坚信它不在网上就满足;我们必须探查,必须走得更远,我们会的。但先看看它画出的独角兽。这是 GPT-4 的独角兽。我看到时个人非常震惊,因为它真的理解了独角兽的概念;它知道关键元素是什么。它画出了这个非常抽象的独角兽。为了让你直观理解 GPT-4 和 ChatGPT 之间的差距,这是 ChatGPT 的独角兽。这就是取得的进步。我想说清楚:ChatGPT 和 GPT-4 有天壤之别。如果你玩过 ChatGPT 还不信,我鼓励你别止步于此。

Here I want to tell you about 'The Strange Case of the Unicorn,' which is kind of my favorite example. Let me just show you the question: 'Draw a unicorn in TikZ.' In this audience, many of you are playing with TikZ to draw images in LaTeX. Personally, when I was a PhD student and even later, I wasted many, many hours struggling with TikZ. It's a real pain to draw anything in TikZ. And of course, drawing a unicorn in TikZ? I don't know, it would take me like two days to do it. Moreover, I'm pretty sure nobody on the internet has asked this question or has drawn a unicorn in TikZ. Who would waste time doing this? It doesn't make any sense. That being said, we will not be convinced just by the fact that I believe it's not on the internet; we will have to probe, we will have to go further, and we're going to do it. But let me show you the unicorn that it came up with. This is GPT-4's unicorn. When I see that, I am personally shocked because it really understands the concept of a unicorn; it knows what the key elements are. It was able to draw this very abstract unicorn. And just to be clear, so that you really understand visually the gap between GPT-4 and ChatGPT, this is ChatGPT's unicorn. So this is how much progress has been made. I want to be clear: there is a world of difference between ChatGPT and GPT-4. If you played with ChatGPT and you were not convinced, I encourage you not to stop there.

GPT-4 使用工具改进独角兽 GPT-4 using tools to improve the unicorn

Host

当然,你可能还是说‘好吧,这没那么厉害。’但我们要看到的一点是,GPT-4 足够聪明,还能使用工具。所以你可以回复它:‘嘿,你知道吗?我不太喜欢你的画。你能改进一下吗?我听说过这些扩散模型,也许你可以用其中一个。’然后它会说:‘当然可以。你能去这个扩散模型网站,把我的图片放进去,让它改进一下吗?’这就是你会得到的结果。所以这是 GPT-4 在允许使用工具时画的独角兽。你可以看到这可能会走向何方。

Of course, you might still say, 'Okay, this is not that great.' But one of the things we're going to see is that GPT-4 is intelligent enough also to use tools. So what you can say is, you can respond to it and say, 'Hey, you know what? I don't like your drawing that much. Can you try to improve it? And I've heard about these diffusion models; maybe you can use one of them.' So what it's going to do is, it's going to say, 'Yeah, sure. Can you go on this diffusion model website and plug in my picture and ask it to improve it?' And this is what you will get. So this is a unicorn of GPT-4 when it's allowed to use tools. You can see where this could potentially go.

深入探究:移除注释与扰动坐标 Probing deeper: removing comments and perturbing coordinates

Host

现在,就像我说的,我不想止步于此。我们要进一步探查。这次怎么探查呢?我要做以下事情:我取出生成的 TikZ 代码。我要删除 TikZ 代码中的所有注释,因为 GPT-4 的一个特性是它生成的代码非常可读,这对机器来说有点好笑,但它加了很多注释,真正引导你理解它的思路。所以我删掉所有这些信息,这样它就不知道这叫‘画独角兽’。里面没有任何关于独角兽的信息。我还要确保——谁知道呢,也许它从网上抄的——我要随机扰动所有坐标,这样它就没见过。然后我移除角。然后我说:‘这段 TikZ 代码本应画一只独角兽,但角不见了。你能把它加回去吗?’所以它必须真正理解代码才能做到。结果就是这样。它真的找到了头部的位置。你要明白这不是一个简单的问题。我的意思是,这里有三个椭圆,三个元素。顺便说,头和鬃毛——它不太擅长画鬃毛——但它确实找到了位置。

Now, as I said, I don't want to stop there. We're going to probe further. How are we going to probe further in this case? What I'm going to do is the following: I'm going to take the TikZ code that was produced. I'm going to remove all the comments in the TikZ code, because one of the properties of GPT-4 is that it produces code that is very much human-readable, which is kind of funny for a machine, but it adds lots of comments; it really guides you to its thinking. So I'm going to remove all of this information so that it doesn't know that this is called drawing a unicorn. There is no information about unicorn in there. I'm also going to make sure that, who knows, maybe it copied this from the web. I'm going to randomly perturb all the coordinates so that it's something that it has never seen. And then I'm going to remove the horn. And I'm going to say, 'This TikZ code is supposed to draw a unicorn, but the horn is missing. Can you add it back?' So it really has to understand the code in order to be able to do that. And this is what happens. It really was able to locate the head. You understand this is not an easy problem. I mean, you have these three ellipses, the three elements. By the way, the head and the mane—it's not very good at drawing the mane—but it really was able to locate it.

独角兽基准与安全调优 The unicorn benchmark and safety tuning

Host

我不想在这个独角兽例子上花太多时间,但我想说另一件非常引人注目的事:几个月来,我们在九月就有访问权限,他们一直在训练它。随着训练进行,我不断查询我的 TikZ 独角兽,看看会发生什么。结果就是:它一直在改进。我漏掉了最好的那个;在我电脑上,我可能稍后回顾。但之后它继续改进。然而,一旦我开始为安全性进行更多训练,它就开始退化了。独角兽开始退化。所以如果今晚你回家让 GPT-4 和 ChatGPT 用 TikZ 画独角兽,你会得到不太好看的结果,更接近 ChatGPT。尽管听起来很傻,但这个独角兽基准我们经常用作智能的基准:你的独角兽有多好?我们在做必应时——这绝对是个真实故事——我们也在调整安全性,我们确实在观察独角兽是否保持良好。

I don't want to stay on this unicorn example too long, but I just want to say that another thing which is really striking is over the months. We had access in September, and they kept training it. As they kept training it, I kept querying for my unicorn in TikZ to see what was going to happen. And this is what happens: it kept improving. I left out the best one; it's on my computer, I will maybe review it later. But it kept improving after that. But eventually it started to degrade once I started to train for more safety. The unicorn started to degrade. So if tonight you go home and you ask GPT-4 and ChatGPT to draw a unicorn in TikZ, you're going to get something that doesn't look great, closer to ChatGPT. And as silly as it sounds, this unicorn benchmark we've used it a lot as a kind of benchmark of intelligence: how good is your unicorn? When we were working on Bing, this is absolutely a real story, we were also tuning for safety, and we were really looking whether the unicorn kept being good.

GPT-4 的安全与理解 Safety and Understanding in GPT-4

Sébastien Bubeck

不错。有时候如果你在安全方面走得太远,它就会说,‘哦不,这个任务太危险了,我不想做。’所以这非常有用。好了,我现在要加快一点速度,因为我想讲的东西很多。你可能会说,‘好吧,这个视觉能力根本没用。’实际上,它非常非常有用。原因是 GPT-4 是智能的,它能理解你。你知道,智能可以等同于理解。理解意味着它遵循你的指令。如果你让它做某事,它就会做你要求的事。让我展示一下这意味着什么。你知道,这个扩散模型,人们还不相信它是智能的。我认为已经有令人信服的证据表明那里存在智能,但这不重要。人们不相信是因为它不能准确理解物体的位置。如果你让它画‘一辆车在咖啡杯旁边,在杯子的右边’,它可能会随机放置。所以它并不真正理解。比如这张图,要求的是勺子放在杯子上面,但你看它把勺子放进了杯子里。这并不奏效。所以让我展示一下从理解中能得到什么。我要问一个非常奇怪的问题,但很可能发生且有用。假设我让 GPT-4 画一张 3D 建筑游戏的截图:一条从左到右的河流,河流下方有沙漠和金字塔,河流上方有高楼林立的城市,屏幕底部有四个按钮,分别叫绿色、蓝色、棕色和红色。这有点随机,但也许我在开发一个游戏,需要这个画面。如果我让扩散模型来做,得到的是这个。看起来不错,但完全不是我要求的。首先,左上角有一些幻觉出来的地图,我没要求那个。还有一些类似生命符号的东西。另外,四个按钮变成了两个多色按钮。所以它做了些事,但没有真正理解我的确切要求。如果交给 GPT-4,得到的是这个。完全是你要求的。它理解了,精确遵循了你的指令。当然,你可能会说,‘好吧,但这看起来不怎么样。’但同样,你不必止步于此。你可以把它作为草图输入扩散模型,如果这样做,得到的就是这个。所以它还没完成,它很有艺术感,而且完全遵循了你想要的指令。所以我认为这开启了很多可能性,你可以想象。

Nice. Or sometimes if you go too far in safety, it's like, 'Oh no, that's too dangerous a task, you know, I don't want to do it.' So this was very useful. Okay, so I will now go a little bit faster because there is a lot that I want to tell you. You might still say, 'Okay, this vision capability is not useful at all.' Actually, it is very, very useful. The reason is that GPT-4 is intelligent and it understands you. You know, intelligence, you can equate it with understanding. Understanding means it follows your instruction. If you ask it to do something, it will do the thing that you asked. So let me show you what this means. You know, this diffusion model that people are not yet convinced is intelligent. I think it's already convincing that there is intelligence there, but it doesn't matter. People are not convinced because it doesn't understand exactly the position of objects. If you ask it, you know, 'a car next to a coffee cup on the right of a cup,' it might be random location. So it doesn't really understand. This picture, for example, is asking for a spoon on top of a cup, and you see it puts the spoon inside the cup. It doesn't really work. So let me show you what you get out of understanding. I'm going to ask a very strange question, but which could very well happen and be useful. Let's say I ask GPT-4 to draw a screenshot of a 3D building game with a river from left to right, a desert with a pyramid below the river, a city with many high rises above the river, and the bottom of the screen has four buttons called green, blue, brown, and red. Something random, but maybe I'm creating a video game and I want this. If I ask a diffusion model to do this, this is what I get. Looks good, but it's not at all what I asked. First of all, there are some hallucinated maps in the upper left corner. I didn't ask for that. Some kind of live symbol. Also, the four buttons became two multi-colored buttons. So it did something, but it really didn't understand what I asked for exactly. If you give it to GPT-4, this is what you get. Exactly what you asked for. It understood, it followed your instructions precisely. Of course, you might say, 'Okay, but this doesn't look great.' But again, you don't have to stop there. You can use this as a sketch in a diffusion model, and if you do that, this is what you get. So it's not done, it's artistic, and it's following exactly the instructions that you wanted. So I think this opens up a lot of possibilities, as you can imagine.

用 GPT-4 编程:从自动补全到完整游戏 Coding with GPT-4: From Autocomplete to Full Games

Sébastien Bubeck

所以让我继续,深入探讨这个绘图能力,但实际上是作为编码,因为毕竟,我把这个绘图能力放在一边,把它当作绘图来展示,但它实际上就是编码。好了,让我们进入编码。顺便说一句,显然所有这些背景幻灯片,你可以想象是谁画的。那么让我们看看当你使用 Copilot 进行编码时会发生什么,比如 GitHub Copilot,但不同的是,现在你的 Copilot 是智能的,它能理解你。所以让我们看看如果我让它做一件相当棘手的事情会发生什么:用 HTML 和 JavaScript 编写一个 3D 游戏,包含以下元素。有三个球形化身。玩家用按键控制其中一个化身移动。有一个敌人试图抓住玩家,还有一个防御者试图保护玩家,挡在敌人和玩家之间。所以防御者在某种程度上本身就是一个 AI。还有随机生成的障碍物。我可以让 ChatGPT 来做。这是它给我的结果。首先,这已经很不可思议了。它给了我代码,大约 50 行代码,编译成这个。这是一个我可以玩的游戏。玩家移动,当然是绿色球。红色球不动。我想蓝色球应该是防御者,它也不动。它并不是真正的 3D。所以它做了些事,但没有真正理解我想要什么。它没有精确遵循我的指令。这是 GPT-4 做的。好吧,这是一个真正的游戏。玩起来很有趣。你移动,它很快就会重新开始。你移动深蓝色球。你看红色球正在向背景中的深蓝色球移动,浅蓝色的是防御者,它试图挡在红色球和深蓝色球之间。所以这段视频是我在控制深蓝色球。你看,啊,防御者做得很好,它挡住了红色球。好了,所以这真的,对我们来说,编码在这个点上发生了一种相变。我的意思是,Codex 和 GitHub Copilot 能够自动补全,你应该把它看作自动补全,你知道,短代码片段。ChatGPT 已经是下一个级别,它可以为你写 50 行代码。但 GPT-4 可以写 500 到 1000 行代码,完全可用,零样本,没有元提示或任何东西。这一切都是开箱即用的。所以这就是与 Copilot 一起编码所解锁的能力。这里我展示这两个动画:左边是 ChatGPT 生成的代码,右边是 GPT-4 生成的代码。如果你仔细看,你会发现 GPT-4 的代码要专业得多。现在,关键点,这张幻灯片上的所有转折,是这两个视频都是由 GPT-4 生成的。所以我做的是让 GPT-4 生成一个 Python 脚本,它接受一个文本文件作为输入,并生成一个像这样的视频,代码连续移动。我的意思是,这需要很多时间。对我来说,制作这些视频要花很长时间。问题是,这个房间里有多少人能在几个小时内写出一个 Python 脚本来生成这个?也许有几个人,但不会很多。好了,这就是 GPT-4 的力量,它解锁了这么多东西,GPT-4 解锁了如此多的创造力。

So let me move on and double down on this drawing, but really as coding, because after all, this drawing capability I put aside and feature it as drawing, but it's really nothing but coding. Okay, so let's go with coding. By the way, obviously all those background slides, well, you can imagine who drew them. So let's see what happens once you go to coding with a copilot, like GitHub Copilot, but except that now your copilot is intelligent, it understands you. So let's see what happens if I ask it something pretty tricky: write a 3D game in HTML and JavaScript with the following elements. There are three avatars who are spherical. The player controls one of the avatars with the keys to move. There is an enemy that tries to catch the player, and there is a defender that tries to protect the player and gets between the enemy and the player. So the defender is kind of an AI itself in some ways. And you have obstacles that spawn randomly. I can ask ChatGPT to do it. This is what it gives me. First of all, this is already incredible. It gives me code, roughly 50 lines of code, that compiles to this. This is a game that I can play. The player moves, the green ball, of course. The red ball is not moving. I imagine the blue ball is supposed to be the defender, it's not moving either. It's not really 3D. So it did something, but it didn't really understand what I wanted. It didn't follow my instructions precisely. This is what GPT-4 does. Okay, so this is a real game. It's fun to play. You move, it's going to restart in a second. You move the dark blue ball. You see the red ball is moving towards the dark blue ball in the background, and the light blue one is a defender which is trying to get between the red ball and the dark blue ball. So this movie is me controlling the dark blue ball. You see, ah, the defender is doing a good job, it's stopping the red ball. Okay, so this is really, for us, there is a kind of phase transition in coding at this point. And really what I mean is that Codex and GitHub Copilot were able to autocomplete, you should think of it as autocomplete, you know, short snippets of code. ChatGPT is already next level, you can already write 50 lines of code for you. But GPT-4 can write 500 to 1000 lines of code, fully works, zero shot, no meta prompting or anything. This all works out of the box. So this is really what coding with a copilot unlocks. And here I'm showing in these two animations: on the left is the code that ChatGPT produces, and on the right is the code that GPT-4 produces. If you look at it carefully, you will see the GPT-4's code is much more expert level. Now the catch, all the twist on this slide, is that those two videos were produced by GPT-4. So what I did is that I asked GPT-4 to produce a Python script that takes as input a text file and produces a video like this with the code moving continuously. I mean, this would take a lot of time. For me, it would take forever to produce those videos. And the question is, who in this room would be able to produce a Python script in a couple of hours that will produce this? Maybe a few people, but not that many. Okay, so this is really the power of GPT-4, how it unlocks so many things, so much creativity is unlocked by GPT-4.

超人编程与模拟面试 Superhuman Coding and Mock Interviews

Sébastien Bubeck

我快速过一下这张幻灯片。我们让它通过了亚马逊和谷歌的模拟面试,不是微软。它通过了。不仅通过了,而且击败了 100% 的人类用户。你看这个特定案例,分配了两个小时,它用了 3 分 59 秒就完成了。花了那么长时间是因为它在 playground 和模拟面试网站之间复制粘贴。好了,所以这真的是,我认为可以公平地说这是超人级的编码。好了,让我继续讲可供性,很快,因为我想告诉你关于……

I will go quickly just on this slide. We had it pass mock interviews at Amazon and Google, not Microsoft. And it passed. Not only did it pass, but it beats 100% of the human users. And you see for this particular one, there were two hours allocated, and it did it in 3 minutes and 59 seconds. It took that long because it was copy-pasting between the playground and the mock interview website. Okay, so this is really, I think it's fair to say it's superhuman coding. Okay, so let me move on to affordances and very quickly, because I want to tell you about...

数学与工具使用 Mathematics and tool use

Host

数学是很多人感兴趣的话题。问题是它还有很多弱点。当然它没有记忆。你知道,美国总统是谁?唐纳德·特朗普。那两个数的乘积的平方根是多少?它说一千。显然不是一千,是九千。所以它会犯算术错误。这个词的某个字母是什么?它说'n',正确答案是'a'。你知道,它会犯错。它并不完美。这是不是每个人都需要理解的重要事情?它远非完美。它有缺陷,就像人类有缺陷一样。但关键在于它足够智能,可以使用工具。所以你可以告诉它:'嘿,你可以使用搜索引擎、计算器、这个 API。我就说它的字符,括号。你可以使用所有这些工具。如果需要,请使用它们。'那么对于'美国总统是谁'这个问题,它不会直接回答,它会说'搜索'。它会告诉你:'好的,我需要搜索这个信息。' '这个的平方根是多少?'它会说'计算'。 '这个词的某个字母是什么?'它是这个词的一个字符,逗号 13。好的,所以逗号 13。我没有告诉它你必须用逗号,你知道,你想要字母的序号,但它会自动找到。现在也许这并不令人印象深刻,但它还可以使用更复杂的工具。例如,你可以让它访问你的日历、你的电子邮件。好的,所以这里我要在幻灯片上展示的是 100%真实的,但我是手动操作的。但你可以很容易地想象自动化这个过程。所以我说:'请这周在 Contoso 餐厅安排与 Joe 和 Luke 的晚餐。'它说,这是它的回应:'calendar.getEvents week'。所以它搜索我的日历,查看我这周有什么活动。它给 Joe 发了一封电子邮件:'email.send: 嘿 Joe,晚餐,哪些晚上有空?'然后我把答案反馈给它。好的,答案是:Joe 说周二和周三晚上有空;Luke 说周一到周四任何一天都可以;我的日历显示我周一和周二有计划。然后它对我提供的这些输入进行推理,并得到答案:'好的,周三是个日子。所以让我给 Joe 发邮件,让我把活动添加到日历,也让我把预订发送给餐厅。'这就是全部,它可以自动完成所有这些。然后它回来告诉你:'我已经在 Contoso 餐厅安排了下午 6 点的晚餐。'

Mathematics is something that will be of interest to many people. The problem is it still has many weaknesses. Of course it doesn't have memory. You know, who is the president of the US? Donald Trump. What is the square root of the product of those two numbers? It says a thousand. It's clearly not a thousand, it's nine thousand. So it makes arithmetic mistakes. What is the certain method of this word? It says 'n', the right answer is 'a'. You know, it makes mistakes. It's not perfect. Is this something very important for everybody to understand? It's far, far from perfect. It's flawed, like a human is flawed. But the point is it's intelligent enough to use tools. So you can tell it, 'Hey, you have access to a search engine, you have access to a calculator, you have access to this API. I just say it's character, you know, parenthesis. You have access to all those things. If you need them, please use them.' So then to the question 'Who is the president of the US?' he will not answer, he will say 'search'. It will tell you, 'Okay, I need to search this information.' 'What is the square root of this?' It will say 'calc'. 'What is a certain letter of this word?' It was a character of the word, comma 13. Okay, so the comma 13. I didn't tell it you have to do comma, you know, the number of the letters that you want, but it will find it automatically. Now maybe it's not that impressive, but it can also do much more complex tools. So for example, you can give it access to your calendar, to your email. Okay, so here what I'm going to show you on this slide is 100% real, but I did it manually. But you can very easily imagine automating this. So what I said is, 'Please set up dinner with Joe and Luke at Contoso restaurant this week.' It says, this is its response: 'calendar.getEvents week'. So it searches in my calendar for what events I have for this week. It sends an email to Joe: 'email.send: Hey Joe, dinner, which nights are available?' Then I feed it back the answer. Okay, which are: Joe says on Tuesday and Wednesday night is available; Luke says any day from Monday to Thursday; and in my calendar it says that I have plans for Monday and Tuesday. Then it reasons over this input that I gave it, and it gets the answer: 'Okay, Wednesday is a day. So let me send an email to Joe, let me add the event to the calendar, and also let me send the reservation to the restaurant.' This is all, it can do all of this automatically. And it comes back to you and says, 'I scheduled dinner at 6 pm at Contoso restaurant.'

数学能力与局限 Mathematics capabilities and limitations

Host

让我快速总结一下数学,因为当然这是很多人非常感兴趣的话题。我直接说:它数学并不好。所以别担心,我们暂时还有工作。但这很有趣。所以让我们进行一次对话,就像我之前告诉你的,我们要像和学生对话一样进行对话。而且我不是——我们和它进行过大约硕士水平的对话,它在这个水平上还能应付。但我要给你展示一个初中水平的问题。所以让我们看这个问题:'一年内,兔子种群首先乘以因子 a,最后一天有 b 只兔子被领养。假设第一年第一天有 x 只兔子,我们知道三年后恰好有 27x 减 26 只兔子。求 a 和 b 的值。'好的,典型的初中问题。设一年后的兔子数量为 y。我们知道 y 等于 ax 减 b。两年后,数量将是 a y 减 b,即 a 乘以(ax 减 b)减 b,等等。三年后,数量将是这个方程。已知它等于 27x 减 26。所以方程必须是这个。解出来,我们得到 a 等于 3,b 等于 2,这就是我们寻找的其他值。这是正确的。好的,到目前为止还不错。这再次说明,你不会面试一个学生然后说'哦哇,好的,就这样,他懂了。'所以我们继续,让我们更抽象一点,看看它是否真的理解它在做什么。这里我不只是——我不想只问一个有明确答案的问题。我还想看看它是否真的理解了这个概念。它能自己修改问题吗?所以假设我自然指出,这个问题可以替代地表述为关于复合函数。你能提供这个复合表述吗?它说:'是的,我们可以说我们正在寻找 a 和 b 的值,使得如果 f(x)等于 ax 减 b,那么 f(f(f(x)))等于 27x 减 26。'好的,这相当不错。它似乎理解了。让我们继续讨论。假设现在 f 是一个二次多项式。

Let me quickly conclude with mathematics, because of course it's a topic of great interest to many of us. And let me say right away: it's not great at mathematics. Okay, so don't worry, we still have a job for now. But it's interesting. So let's have a conversation, just as I was telling you earlier, we're going to have a conversation like we would have with a student. And I'm not—we had conversation with it at the Master level roughly, and it can hold its own at that level. But I'm going to show you a level of like middle school or something like that. So let's go with this problem: 'Within the duration of a year, a rabbit population first multiplies itself by a factor a, and on the last day b rabbits are taken for adoption. Suppose that on the first day of the first year there are x rabbits, and let's say that we know after three years that there are exactly 27x minus 26 rabbits. What are the values of a and b?' Okay, typical kind of middle school question. Let's call the rabbit population after one year y. We know that y is equal to ax minus b. After two years, the population will be a y minus b, which is a times (ax minus b) minus b, which is blah blah blah. After three years, the population would be this equation. We're given that this is equal to 27x minus 26. So the equation must be this. Solving for it, we find a equals 3 and b equals 2, the other values that we are looking for. This is correct. Okay, so far so good. That again, you know, you don't interview a student and be like 'Oh wow, okay, that's it, he understands.' So we continue and let's be a little bit more abstract to see if it really understood what it was doing. And here I'm not only—I don't want to only ask a question where there is a clear answer. I also want to see whether it really understood the concept. Can it modify the question itself? So suppose that naturally I point out that this question can be alternatively formulated about the composition. Can you provide this compositional formulation? And it says, 'Yes, we could say that we're looking for values of a and b such that if f of x is ax minus b, then f of f of f of x is equal to 27x minus 26.' Okay, that's pretty good. It seems to be understanding. Let's continue the discussion. Suppose that now f is going to be a polynomial of degree 2.

多项式组合推理 Polynomial composition reasoning

Sébastien Bubeck

好的,所以一个真正的多项式,x² 的系数非零。你能在这种情况下找到这样的函数 f 吗?好的,好的。你看,作为人类,你会想,好的,我希望这个复合等于一个线性函数,也就是一次多项式,但复合三次会得到一个八次多项式。八次多项式,八不等于一。没有这样的函数。好的,这是一个非常简单的问题。但让我们看看 GPT-4 会怎么做。如果 f 是二次多项式,那么 f(x) 可以写成这样。然后给定这个,方程变成了,然后它开始迷失,因为它开始写出三次复合。它写了很多东西。它说,我需要这个方程,那个方程。你开始写八个方程,但它没有得出答案。但同样,我们并没有就此打住。我们说,嘿,等等,也许你可以不用进行具体计算就能推断出一些东西。也许你不想把一切都写下来。不像之前那样。然后它说,啊,好的,我能注意到的一点是,如果 f 是二次多项式,那么三次复合是八次多项式,所以没有这样的函数。好的,所以这里你看到它有多微妙。不清楚。它理解吗?它不理解吗?我不确定。好的,我只是不确定。这就是我要说的全部。

Okay, so a true polynomial non-zero coefficient on x squared. Can you find such a function f in this case? Okay, okay. So you see, as a human you're like, okay, so I want this composition to be equal to a linear function, which is a polynomial of degree one, but the composition three times is going to be a polynomial of degree eight. Polynomial of degree eight, eight is not equal to one. There is no such function. Okay, this is a very simple question. But let's see what GPT-4 does. If f is a polynomial of degree two, then f of x can be written like this. Then given this, the equation becomes, and then it starts to get lost because it starts to write into the composition three times. It writes many things. It says I need this equation, that equation. You start to write eight equations, and it doesn't get to the answer. But again, we don't stop there. We say, hey wait a second, maybe there's something you can deduce here without carrying calculation. Maybe you don't want to write down everything. It's not like before. And then it says, ah, okay, one thing I can notice is that if f is a polynomial of degree two, then the composition three times is a polynomial of degree eight, so there is no such function. Okay, so here you see how it's delicate. It's not clear. Does it understand? Does it not understand? I'm not sure. Okay, I'm just not sure. And this is all I will say.

算术错误与自我修正 Arithmetic mistakes and self-correction

Sébastien Bubeck

现在,有一些奇怪的事情,比如算术仍然不可靠。我不得不说,我不完全理解,但我理解了一些东西,我将在这一页幻灯片上解释给你听。让我们看看这个。我给它一个提示:七乘以四加上八乘以那个。好的,我不知道这个值是多少,但你知道,8 乘以 8 是 60 多,7 乘以 4 是 20 多,所以至少这个数低于 100。好的,很好。它说 120。这是错的,完全错了。好的,但关键是它并没有就此打住。它继续下去。它开始解释为什么它认为是 120。七乘以四加上八乘以八。它做了计算,然后得出了正确答案 92。好的,等等,发生了什么?你一开始说 120。到底是哪个?是 120 还是 92?哦,那是个笔误。抱歉。哦,是的。所以实际上你可以从这张幻灯片中获得很多见解。你知道,你真的可以理解我认为正在发生的一切。所以第一个答案,120,你理解它必须仅使用内部表示来完成。你知道,仅使用它的内部表示,它必须做这个加法,这稍微困难一些。你知道,为什么它立即回答?这是因为当你问这样一个问题时,你知道,你写下这个等式,你写等号。之后最可能发生的事情是给出一个数字。所以它给你数字。它试图给出之后最可能出现的数字。它尝试了,但失败了。但之后,第二可能的事情是什么?你知道,人们解释他们的理由,他们的答案。所以它然后尝试解释它的答案。主要的事情是它得出了一个不同的答案。你必须理解这很惊人,因为据我所知,这是一个 Transformer,所以它是基于注意力机制的。所以当它是基于注意力机制时,你理解当它第二次说“七乘以四加上八乘以八”时,它的注意力非常强烈地指向 120 这个答案。120 这个答案,你必须理解,现在某种程度上是它事实的一部分。你知道,就它所知,可能是你告诉它,嘿,你知道吗,从现在起七乘以四加上八乘以八等于 120。你知道,它可能是我提示的一部分。所以它得出不同答案的事实意味着它已经接受了足够的训练来克服其提示中的错误。所以这是一个非常非常强的特性:尽管一开始犯了错误,它仍然能够得出正确答案。现在,当然,当它说“这是一个笔误”时,这也非常有趣,因为这显然不是笔误。你知道,这涉及到幻觉和许多许多有趣的话题。我想留一些时间提问,所以我不想再解释更多了。但这张幻灯片真的,你必须深入思考。它说了很多。

Now, there are some weird things, like the fact that the arithmetic is still shaky. I have to say, I don't fully understand, but I understand something which I will explain to you on this slide. So let's look at this. I give it as a prompt: seven times four plus eight times that. Okay, I don't know what is the value of this, but you know, 8 times 8 is 60 something, 7 times 4 is 20 something, so at the very least this is below 100. Okay, good. It says 120. This is wrong, flat out wrong. Okay, but the point is it doesn't stop there. It continues. It starts to explain why it thinks it's 120. Seven times four plus eight times eight. It does the calculation, and then it gets to the correct answer, 92. Okay, wait, what's going on? You started by saying 120. Which one is it? Is it 120 or 92? Oh, that was a typo. Sorry. Oh yeah. So there is a lot of insight that you can draw from this slide actually. You know, you can really understand everything I think that's happening. So the first answer, the 120, you understand that it has to do this using only internal representation. You know, only using its internal representation, it has to do this addition, and this is slightly more difficult. And you know, why does it answer immediately? It's because when you ask a question like this, you know, you write this equation, you write equal. The most likely thing that happens after is to give a number. So it gives you the number. It tries to give you what is the most likely thing to appear after. It tries, but it fails. But then, what is the second most likely thing after that? You know, people explain their rationale, their answer. So then it tries to explain its answer. And what is the main thing is that it gets at a different answer. And you have to understand that it's amazing because, as far as I know, this is a Transformer, so it's attention based. So when it's attention based, you understand that when it's saying the second time, 'seven times four plus eight times eight', its attention brings it very strongly to the 120 answer. The 120 answer, you have to understand, is kind of part of its truth now. You know, for all it knows, it could be that you have told it, hey, you know what, seven times four plus eight times eight is 120 from now on. You know, it could have been part of my prompt. So the fact that it gets to a different answer means that it has been trained enough to overcome mistakes in its prompt. So this is a very, very strong property: the fact that it's able to get to the right answer despite making a mistake at the beginning. Now, of course, when it says 'this was a typo', this is also very interesting because this is obviously not a typo. You know, and this gets to the hallucination and many, many interesting topics. And I want to take some time for questions, so I don't want to explain more about this. But this slide really, you have to think about it deeply. It says a lot.

无法提前规划 Inability to plan ahead

Sébastien Bubeck

在进入结论之前的最后一张幻灯片是它无法进行真正的规划。再说一次,你会——我的意思是,我对许多任务感到惊讶,我以为它们需要真正的规划,但实际上并不需要。但让我给你一个例子,我们继续讨论七乘以四加上八乘以八。好的,很好。所以现在你有一个等于 92 的等式。让我问一个有趣的问题:你能修改这个等式左边恰好一个整数,使得答案变成 106 吗?那么作为人类,你的推理是什么?你的推理是这样的:好的,我想要右边是 106,所以我需要增加 14。好的,我需要增加 14,而且我只能修改左边的一个数字。14,我看左边,我看到一个 7,然后我有一种顿悟时刻:啊,14 是 7 乘以 2。好的,所以如果是 7 乘以 2,那么我需要把这个 4 变成 6。好的,所以我说的就是:它需要把这个 4 变成 6。但你看,我的这个顿悟,尽管非常简单,却是通过某种规划实现的。我在提前思考我需要什么。而 GPT-4 做不到这一点,因为它是一个下一个词预测设备。所以它会说,你知道,有几种可能的方法,等等等等,然后它说,我可以修改恰好一个整数。我要把 7 改成 9。我计算 9 乘以 4,你知道,这等于 106。等等,如果我把 7 改成 9?我加了 8,所以这是 100。答案不是 106。然后它试图解释为什么这有效。你知道,9 乘以 4 加上 8 乘以 8 是 36 加 64。这是正确的。但然后,你知道,它又说 106。

The last slide before moving to the conclusion is the fact that it cannot do true planning. And again, you will be—I mean, I have been amazed by so many tasks that you can do where I thought it would require true planning, but actually it doesn't. But let me give you one example where we continue this discussion with seven times four plus eight times eight. So okay, great. So now you have this identity which is equal to 92. And let me ask a funny question: can you modify exactly one integer on the left-hand side of this equation so that the answer becomes 106? So as a human being, what is your reasoning? Your reasoning is like this: okay, I want 106 on the right-hand side, so I need to increase by 14. Okay, I need to increase by 14, and I can modify only one number on the left. 14, I look at the left, I see a seven, and then I have this kind of Eureka moment: ah, 14 is 7 times 2. Okay, so if it's 7 times 2, then I need to turn this 4 into a 6. Okay, so what I said is just this: it needs to turn this 4 into a 6. But you see, this Eureka that I had, even though it's extremely simple, it was through some kind of planning. I was thinking ahead about what I'm going to need. And GPT-4 cannot do that because it's a next-token prediction device. So what it's going to do is it's going to say, you know, there are a few possible ways to do it, blah blah blah, and then it says, you know, I can modify exactly one integer. I'm going to modify the seven into a nine. I do nine times four, you know, and this is equal to 106. Wait, what if I modify the seven to a nine? I add an eight, so this is 100. The answer not 106. And then it tries to explain why this works. You know, nine times four plus eight times eight is 36 plus 64. That's correct. But then again, you know, it says 106.

局限与未来潜力 Limitations and Future Potential

Sébastien Bubeck

所以你看,它还不够强大,无法克服最初的错误。这让我觉得,如果进一步训练,它可能会自我纠正。如果再进一步训练,它可能会明白,即使对于像'7 乘以 4 加 8 乘以 8'这样的问题,最可能的答案是一个数字,但经过更多训练后,它可能会理解回答这个问题的最佳方式是先进行推理。所以我想说的是,通过这个愚蠢的例子,我看到的是,通过更多的训练,我们将解锁比现在多得多的能力。我们现在拥有的已经很惊人了,但远未达到这项技术的全部潜力。前方还有更多可能。

So you see here it was not strong enough to overcome its initial mistake. And this to me points to the fact that if it was trained further, maybe it would correct itself. And if it was trained even further, maybe it would understand that even though the most likely thing when there is a question like 'seven times four plus eight times eight' is a number, maybe if it's trained more, it would understand that the best way to answer this is to first do the reasoning. So what I'm saying here is that through this stupid example, what I see is that with more training, we're going to unlock a lot more than what we currently have. What we currently have is already amazing, but it's far from everything we can do with this technique. There is a lot more on the horizon.

GPT-4 有智能吗? Is GPT-4 Intelligent?

Sébastien Bubeck

好的,让我总结一下:GPT-4 有智能吗?这重要吗?这是一个非常重要的问题。所以,GPT-4 有智能吗?这完全取决于你的定义。我把这个问题留给你;我不判断它是否有智能。就我而言,根据我对智能的定义,是的,它有智能。但你知道,它缺乏记忆,无法进行实时学习。如果这是你的定义,那么它就没有智能。它不能提前思考多次,不能进行真正的规划。如果那是你的定义,那么它就没有智能。但另一方面,我展示的一些行为确实令人印象深刻,而且可能比令人印象深刻更重要的是,它们很有用。你知道,在我的团队里,我们每天都用 GPT-4;它已经成为我们工作流程的一部分。所以,仅仅因为它有用——再说一次,你说它有没有智能并不重要——它都会改变世界,无论你喜欢与否。而且,我还想说,这也许是一个重新思考什么是智能的机会。因为从某种意义上说,尽管我们有几十年的心理学研究智能,但我们只有一个智能的例子,那就是自然进化带给我们的智能——自然世界的自然智能。但在这里,我们有一个新的过程,产生了一些看起来有智能的东西。所以现在我们有了不同的例子,也许我们可以触及智能的核心。而这项研究的答案可能恰恰是:'不,这个新东西你不应该称之为智能,因为它不能做 X。'这是一个非常合理的结论。但也许更重要的是我所说的:我们可以从中提取更多的东西。所以 GPT-4 绝不是终点,完全不是;这只是开始。这是第一个展现出真正智能微光的模型,但前方还有更多、更多。

Okay, so let me conclude: is GPT-4 intelligent, and does it matter? This is a really important question. So again, is GPT-4 intelligent? It really depends on your definition. I leave it up to you; I'm not making a call whether it's intelligent or not. As far as I'm concerned, in terms of my definition of intelligence, yes, it is intelligent. Now, you know, it's lacking memory; it cannot do real-time learning. If this is your definition, then it's not intelligent. It cannot think several times in advance; it cannot do real planning. If that's your definition, then it's not intelligent. But on the other hand, some of those behaviors I showed you are really impressive, and maybe more importantly than impressive, they are useful. You know, in my team, we all use GPT-4 every day; it's part of our workflow. So this mere fact that it's useful—again, it doesn't matter if you say it's intelligent or not—it is going to change the world, whether you like it or not. And also, I want to say that maybe it's an opportunity to rethink what intelligence is. Because in a way, even though we have decades of psychology studying intelligence, we had only one example of intelligence, which is the intelligence that natural evolution brought us—the natural intelligence of the natural world. But here we have a new process that led to something that looks intelligent. So now that we have different examples, maybe we can get at the core of intelligence. And maybe the answer to that study will be exactly, 'Yeah, no, this new thing you shouldn't call it intelligence because it doesn't do X.' That's a very plausible conclusion. But maybe more importantly is what I said: that there is so much more that you can extract from this. So GPT-4 is by no means the end, not at all; this is the beginning. This is the first one that shows some glimmer of real intelligence, but there is much, much more on the horizon.

社会影响与行动呼吁 Societal Implications and Call to Action

Sébastien Bubeck

那么我们应该从中得出什么结论?作为大学、作为社会、作为人类?我是认真的;这些都是我们必须面对的真实问题。在这里我真的很想说:作为社会,要掌控这个问题,我们必须超越关于这是复制粘贴还是统计的讨论。我们必须把这种讨论抛在脑后。火车已经离站了。所以如果我们一直纠缠于这个版本的问题,我们就会错过真正重要的问题。所以我认为继续前进很重要。最后让我再说一句,它能做的远比我在这里展示的要多。它可以做数据分析;你可以给它数据,它会为你分析。它可以作为隐私检测器。它的医学和法律知识令人惊叹。在这里我想推荐一本由微软研究院撰写的书,我也参与其中,主要作者是 Peter Lee,还有在座的 Carey Goldberg 和哈佛的 Isaac Kohane,内容是关于使用 GPT-4 进行医疗保健。书名是《医学中的 AI 革命》。这是一个非常复杂的话题,我不想再多说一个字,因为一句话说不清楚。但真的,它的医学知识将使其在医疗保健领域产生巨大影响,希望是好的影响,但我们必须深入思考。它可以玩游戏,充当游戏环境。它懂音乐——再说一次,它从未听过音乐,但它懂音乐。它可以做文件管理,还有很多很多。好了,我就讲到这里。谢谢。

So what conclusion should we draw from that? As a university, as a society, as humanity? I mean, I'm being real here; these are real questions that we should confront. And here I really want to say: for us as a society to control this question, we have to go beyond the discussion of whether this is copy-paste or statistics. We have to leave this discussion behind us. The train has left the station. So if we keep getting bogged down by this version of the question, we're going to miss the real important questions. So I think it's important to move on. And let me also conclude by saying that it can do a lot more than what I have shown here. It can do data analysis; you can give it data and it will do analysis for you. It can be used as a privacy detector. Its medical and law knowledge is amazing. And here I would like to make a plug for a book that was written at Microsoft Research, and I helped with that, by Peter Lee as a lead author, Carey Goldberg who is in the room, and Isaac Kohane from Harvard, on using GPT-4 for healthcare. The book is titled 'The AI Revolution in Medicine.' It's a very complex topic, and I don't even want to say one more word about it because I won't do it justice in one sentence. But really, its medical knowledge is going to make it have a big impact in healthcare, hopefully in a good way, but we have to think about it deeply. It can play games, act as a game environment. It knows music—again, it never listened to music, but it knows music. It can do file management, and so much more. Okay, I will conclude here. Thank you.

互动版:逐字朗读 + 针对本期提问 →