Selecting Talent and Early AI Inspirations
打开互动全文版(中英对照 + 朗读 + 问答)→杰弗里·辛顿反思如何选拔人才、早期理解大脑的失望,以及伊利亚·苏茨克维带着关键见解出现的那一刻。
Geoffrey Hinton reflects on how he selected talent, his early disappointments in understanding the brain, and the moment Ilya Sutskever showed up with a critical insight.
你对于如何挑选人才有很多思考吗,还是说这更多是凭直觉?就像伊利亚出现时,你心想‘这是个聪明人,我们一起干吧’。还是说你对此有过深思熟虑?
Have you reflected a lot on how to select talent, or has that mostly been intuitive to you? Ilya just shows up and you're like, 'This is a clever guy, let's work together.' Or have you thought a lot about that?
我记得我刚从英国到卡内基梅隆大学时。在英国的研究机构,一到六点大家就去酒吧喝酒。在卡内基梅隆,我记得待了几周后的一个周六晚上,我还没交到朋友,不知道做什么,于是决定去实验室编程,因为我有一台列表机,没法在家编程。所以周六晚上九点左右我去了实验室,结果那里人满为患。所有学生都在,他们之所以在那里,是因为他们正在做的工作关乎未来。他们都相信,自己接下来做的事情将改变计算机科学的进程。这和英国太不一样了,让人耳目一新。
I remember when I first got to Carnegie Mellon from England. In England, at a research unit, it would get to be 6:00 and you'd all go for a drink in the pub. At Carnegie Mellon, I remember after I'd been there a few weeks, it was Saturday night. I didn't have any friends yet and I didn't know what to do, so I decided I'd go into the lab and do some programming because I had a list machine and you couldn't program it from home. So I went into the lab at about 9:00 on a Saturday night and it was swarming. All the students were there, and they were all there because what they were working on was the future. They all believed that what they did next was going to change the course of computer science. It was just so different from England, and that was very refreshing.
带我回到最初,杰夫,在剑桥试图理解大脑的时候。那是什么样的经历?
Take me back to the very beginning, Jeff, at Cambridge, trying to understand the brain. What was that like?
非常令人失望。我学了生理学,夏季学期他们打算教我们大脑如何工作。但他们只教了神经元如何传导动作电位,这很有趣,但并没有告诉你大脑是如何工作的。所以这极其令人失望。我转到了哲学,以为他们可能会告诉我们心智如何工作。结果也很令人失望。最后我去了爱丁堡大学学习人工智能,那更有趣。至少你可以模拟东西,从而检验理论。
It was very disappointing. So I did physiology, and in the summer term they were going to teach us how the brain worked. And all they taught us was how neurons conduct action potentials, which is very interesting but it doesn't tell you how the brain works. So that was extremely disappointing. I switched to philosophy, then I thought maybe they'd tell us how the mind worked. That was very disappointing. I eventually ended up going to Edinburgh to do AI, and that was more interesting. At least you could simulate things, so you could test out theories.
你还记得是什么让你对人工智能产生兴趣的吗?是一篇论文,还是某个特定的人让你接触到了这些想法?
Do you remember what intrigued you about AI? Was it a paper, was it any particular person that exposed you to those ideas?
我想是唐纳德·赫布的一本书对我影响很大。他对如何学习神经网络中的连接强度非常感兴趣。我还很早就读了约翰·冯·诺伊曼的一本书,他对大脑如何计算以及它与普通计算机有何不同非常感兴趣。
I guess it was a book I read by Donald Hebb that influenced me a lot. He was very interested in how you learn the connection strengths in neural nets. I also read a book by John von Neumann early on, who was very interested in how the brain computes and how it's different from normal computers.
你在那时就确信这些想法会成功吗,还是说你在爱丁堡时期有什么直觉?
Did you get that conviction that these ideas would work out at that point, or what was your intuition back in the Edinburgh days?
在我看来,大脑一定有某种学习方式,而且显然不是通过把各种东西编程进去,然后使用逻辑推理规则。这从一开始对我来说就很荒谬。所以我们必须弄清楚大脑如何学习修改神经网络中的连接,以便它能做复杂的事情。冯·诺伊曼相信这一点,图灵也相信。冯·诺伊曼和图灵都很擅长逻辑,但他们不相信这种逻辑方法。
It seemed to me there has to be a way that the brain learns, and it's clearly not by having all sorts of things programmed into it and then using logical rules of inference. That just seemed to me crazy from the outset. So we had to figure out how the brain learned to modify connections in a neural net so that it could do complicated things. And von Neumann believed that, Turing believed that. So von Neumann and Turing were both pretty good at logic, but they didn't believe in this logical approach.
你在研究神经科学的想法和直接做看起来对人工智能好的算法之间是如何分配的?早期你从神经科学中汲取了多少灵感?
What was your split between studying the ideas from neuroscience and just doing what seemed to be good algorithms for AI? How much inspiration did you take early on?
我从未深入研究过神经科学。我一直受到我所了解的大脑工作原理的启发:有一堆神经元,它们执行相对简单的操作,是非线性的,但它们收集输入,加权,然后输出依赖于加权输入的结果。问题是如何改变这些权重,让整个系统做好事?这似乎是一个相当简单的问题。
I never did that much study of neuroscience. I was always inspired by what I'd learned about how the brain works: that there's a bunch of neurons, they perform relatively simple operations, they're nonlinear, but they collect inputs, they weight them, and then they output something that depends on that weighted input. And the question is how do you change those weights to make the whole thing do something good? It seems like a fairly simple question.
你记得当时有哪些合作?
What collaborations do you remember from that time?
我在卡内基梅隆的主要合作对象并不在卡内基梅隆。我和特里·谢诺夫斯基有很多交流,他在巴尔的摩的约翰·霍普金斯大学。大约每月一次,要么他开车到匹兹堡,要么我开车到巴尔的摩——相距 250 英里——我们会一起度过一个周末,研究玻尔兹曼机。那是一次很棒的合作。我们都确信那就是大脑的工作方式。那是我做过的最激动人心的研究,也产生了很多非常有趣的技术成果。但我认为那并不是大脑的工作方式。我还和彼得·布朗有过一次非常好的合作,他是一位非常优秀的统计学家,在 IBM 从事语音识别工作。后来他以一个更成熟的学生身份来到卡内基梅隆攻读博士学位,但他已经知道很多了。他教了我很多关于语音的知识,事实上他教了我隐马尔可夫模型。我想我从他那里学到的比他从我这里学到的更多。这就是你想要的那种学生。当他教我隐马尔可夫模型时,我正在做带隐藏层的反向传播——只是当时它们不叫隐藏层——我决定,隐马尔可夫模型中使用的名字对于你不知道它们在做什么的变量来说是个好名字。这就是神经网络中“隐藏”这个名字的由来。我和彼得决定,这对神经网络中的隐藏层来说是个好名字。但我从彼得那里学到了很多关于语音的知识。
The main collaboration I had at Carnegie Mellon was with someone who wasn't at Carnegie Mellon. I was interacting a lot with Terry Sejnowski, who was in Baltimore at Johns Hopkins. About once a month, either he would drive to Pittsburgh or I would drive to Baltimore—it's 250 miles away—and we would spend a weekend together working on Boltzmann machines. That was a wonderful collaboration. We were both convinced it was how the brain worked. That was the most exciting research I've ever done, and a lot of technical results came out that were very interesting. But I think it's not how the brain works. I also had a very good collaboration with Peter Brown, who was a very good statistician and worked on speech recognition at IBM. Then he came as a more mature student to Carnegie Mellon just to get a PhD, but he already knew a lot. He taught me a lot about speech, and in fact he taught me about hidden Markov models. I think I learned more from him than he learned from me. That's the kind of student you want. And when he taught me about hidden Markov models, I was doing backpropagation with hidden layers—only they weren't called hidden layers then—and I decided that the name they use in hidden Markov models is a great name for variables that you don't know what they're up to. And so that's where the name 'hidden' in neural nets came from. Me and Peter decided that was a great name for the hidden layers in neural nets. But I learned a lot from Peter about speech.
带我们回到伊利亚出现在你办公室的时候。
Take us back to when Ilya showed up at your office.
我在办公室,可能是个星期天,我想我正在编程。然后有人敲门——不是普通的敲门,而是很急促的敲门。所以我走过去开门,门口站着一个年轻学生。他说他暑假在炸薯条,但他更愿意在我的实验室工作。于是我说:‘那你预约一下,我们谈谈。’伊利亚说:‘现在怎么样?’这大概就是伊利亚的性格。我们聊了一会儿,我给了他一篇论文读,那是《自然》杂志上关于反向传播的论文。我们约好一周后再见。他回来时说:‘我没看懂。’我很失望。我觉得他看起来是个聪明人,但那只链式法则,没那么难理解。他说:‘哦不,不,我懂了。我只是不明白为什么你不把梯度给一个合理的函数优化器。’这让我们思考了好几年。之后他一直都是这样。他有非常好的、对事物原始直觉总是非常好。
I was in my office, probably on a Sunday, and I was programming I think. And there was a knock on the door—not just any knock, but a very urgent knock. So I went and answered the door, and there was this young student there. He said he was cooking fries over the summer but he'd rather be working in my lab. So I said, 'Well, why don't you make an appointment and we'll talk.' And so Ilya said, 'How about now?' And that sort of was Ilya's character. So we talked for a bit, and I gave him a paper to read, which was the Nature paper on backpropagation. And we made another meeting for a week later. He came back and he said, 'I didn't understand it.' And I was very disappointed. I thought he seemed like a bright guy, but it's only the chain rule, it's not that hard to understand. And he said, 'Oh no, no, I understood that. I just don't understand why you don't give the gradient to a sensible function optimizer.' Which took us quite a few years to think about. And it kept on like that with him. He had very good, his raw intuitions about things were always very good.
你认为是什么让伊利亚拥有那些直觉?
What do you think had enabled those intuitions for Ilya?
我不知道。我认为他总是独立思考。他从小就对人工智能感兴趣。他显然数学很好,但这很难说。
I don't know. I think he always thought for himself. He was always interested in AI from a young age. He's obviously good at math, but it's very hard to know.
你们两人之间的合作是怎样的?你扮演什么角色,伊利亚扮演什么角色?
What was that collaboration between the two of you like? What part would you play and what part would Ilya play?
非常有趣。我记得有一次我们试图做一个复杂的事情,生成数据的映射,我用了某种混合模型。这样你可以用同一组相似度做出两张图,在一张图中‘bank’可以靠近‘greed’,在另一张图中‘bank’可以靠近‘river’。因为在一张图中你不能让它同时靠近两者,对吧?因为‘river’和‘greed’相距很远。所以我们做了混合映射,用 MATLAB 实现,这涉及大量代码重组以进行正确的矩阵乘法。
It was a lot of fun. I remember one occasion when we were trying to do a complicated thing with producing maps of data where I had a kind of mixture model. So you could take the same bunch of similarities and make two maps, so that in one map 'bank' could be close to 'greed', and in another map 'bank' could be close to 'river'. Because in one map you can't have it close to both, right? Because 'river' and 'greed' are far apart. So we'd have mixture maps, and we were doing it in MATLAB, and this involved a lot of reorganization of the code to do the right matrix multiplies.
他受够了,有一天过来说:‘我要写一个 Matlab 接口,这样我用另一种语言编程,然后有个东西自动转成 Matlab。’我说:‘不行,伊利亚,那要花你一个月时间。我们得继续这个项目,别被这个分心。’他说:‘没事,我今天早上就做完了。’这真是不可思议。那些年里,最大的转变不光是算法,还有那种技能。你这些年怎么看那种技能?伊利亚很早就有了那种直觉。他一直鼓吹:只要把模型做大,它就会更好用。我以前觉得那有点偷懒,你总得有些新想法吧。结果证明我基本是对的:新想法有帮助——Transformer 之类的帮助很大——但真正关键的是数据和算力的规模。那时候我们完全不知道计算机会快十亿倍,我们以为可能快一百倍就不错了。我们试图靠聪明点子解决问题,而那些问题如果有更大的数据和算力规模,本来会自己解决。2011 年左右,伊利亚和另一个研究生詹姆斯·马丁斯发了一篇论文,用字符级预测。我们拿维基百科,尝试预测下一个 HTML 字符,效果出奇地好。我们一直惊讶于它效果这么好。那是在 GPU 上用了一个花哨的优化器,我们几乎不敢相信它真的理解了,但看起来它确实理解了。那简直不可思议。
Got fed up with that, so he came one day and said, 'I'm going to write an interface for Matlab, so I program in this different language and then I have something that just converts it into Matlab.' And I said, 'No, Ilya, that'll take you a month to do. We've got to get on with this project. Don't get diverted by that.' And he said, 'It's okay, I did it this morning.' And that's quite incredible. Throughout those years, the biggest shift wasn't necessarily just the algorithms, but also the skill. How did you sort of view that skill over the years? Ilya got that intuition very early. So Ilya was always preaching that you just make it bigger and it'll work better. And I always thought that was a bit of a cop-out. You're going to have to have new ideas too. It turns out I was basically right: new ideas help—things like Transformers helped a lot—but it was really the scale of the data and the scale of the computation. And back then we had no idea computers would get like a billion times faster. We thought maybe they'd get a hundred times faster. We were trying to do things by coming up with clever ideas that would have just solved themselves if we had had bigger scale of data and computation. In about 2011, Ilya and another graduate student called James Martens had a paper using character-level prediction. So we took Wikipedia and we tried to predict the next HTML character, and that worked remarkably well. We were always amazed at how well it worked. That was using a fancy optimizer on GPUs, and we could never quite believe that it understood anything, but it looked as though it understood. That just seemed incredible.
你能讲讲模型是如何训练来预测下一个词的吗?以及为什么认为这是错误的思考方式?
Can you take us through how are models trained to predict the next word, and why is it the wrong way of thinking about them?
好吧,我其实不认为这是错误的思考方式。事实上,我认为我做了第一个使用嵌入和反向传播的神经网络语言模型。它很简单:数据只是三元组,把每个符号转成嵌入,然后让嵌入相互作用来预测下一个符号的嵌入,再从中预测下一个符号。然后通过整个过程反向传播来学习这些三元组。我展示了它能泛化。大约 10 年后,约书亚·本吉奥用了非常相似的网络,并展示了它在真实文本上的效果。又过了大约 10 年,语言学家才开始相信嵌入。这是一个缓慢的过程。我认为它不仅仅是预测下一个符号的原因是:如果你问‘预测下一个符号需要什么?’特别是如果你问我一个问题,那么答案的第一个词就是下一个符号,你必须理解这个问题。所以我认为通过预测下一个符号,它和传统的自动补全非常不同。传统的自动补全会存储单词的三元组,然后如果你看到一对单词,你会看到不同单词作为第三个出现的频率,这样就能预测下一个符号。大多数人以为自动补全就是这样。现在完全不是那样了。要预测下一个符号,你必须理解已经说过的话。所以我认为通过让它预测下一个符号,你是在强迫它理解。而且我认为它的理解方式和我们在很大程度上是一样的。很多人会告诉你这些东西不像我们——它们只是预测下一个符号,不像我们一样推理。但实际上,为了预测下一个符号,它必须做一些推理。我们现在看到,如果你做大模型而不加入任何特殊的推理机制,它们已经能做一定的推理。我认为随着模型变大,它们能做的推理会越来越多。你觉得我现在除了预测下一个符号之外还在做别的事吗?我认为那就是你学习的方式。我认为你在预测下一个视频帧,你在预测下一个声音。但我认为这是一个相当合理的关于大脑如何学习的理论。
Okay, I don't actually believe it is the wrong way. In fact, I think I made the first neural net language model that used embeddings and backpropagation. So it's very simple: data just triples, and it was turning each symbol into an embedding, then having the embeddings interact to predict the embedding of the next symbol, and from that predict the next symbol. Then it was backpropagating through that whole process to learn these triples. I showed it could generalize. About 10 years later, Yoshua Bengio used a very similar network and showed it worked with real text. And about 10 years after that, linguists started believing in embeddings. It was a slow process. The reason I think it's not just predicting the next symbol is: if you ask, 'What does it take to predict the next symbol?' Particularly if you ask me a question and then the first word of the answer is the next symbol, you have to understand the question. So I think by predicting the next symbol, it's very unlike old-fashioned autocomplete. Old-fashioned autocomplete would store sort of triples of words, and then if you saw a pair of words, you see how often different words came third, and that way you can predict the next symbol. That's what most people think autocomplete is like. It's no longer at all like that. To predict the next symbol, you have to understand what's been said. So I think you're forcing it to understand by making it predict the next symbol. And I think it's understanding in much the same way we are. So a lot of people will tell you these things aren't like us—they're just predicting the next symbol, they're not reasoning like us. But actually, in order to predict the next symbol, it's going to have to do some reasoning. And we've seen now that if you make big ones without putting in any special stuff to do reasoning, they can already do some reasoning. And I think as you make them bigger, they're going to be able to do more and more reasoning. Do you think I'm doing anything else than predicting the next symbol right now? I think that's how you're learning. I think you're predicting the next video frame, you're predicting the next sound. But I think that's a pretty plausible theory of how the brain's learning.
是什么让这些模型能够学习如此广泛的领域?
What enables these models to learn such a wide variety of fields?
这些大语言模型所做的是寻找共同结构。通过找到共同结构,它们可以利用共同结构来编码事物,这样更高效。举个例子。如果你问 GPT-4:‘为什么堆肥堆和原子弹相似?’大多数人回答不了。大多数人没想过;他们认为原子弹和堆肥堆是完全不同的东西。但 GPT-4 会告诉你:‘嗯,能量规模和时间尺度非常不同,但相同的是,当堆肥堆变热时,它产生热量的速度更快;当原子弹产生更多中子时,它产生中子的速度更快。’所以它理解了链式反应的概念。我相信它理解了它们都是链式反应的形式。它利用这种理解将所有信息压缩到权重中。如果它这样做,那么它还会对成百上千个我们还没看到类比的事物做同样的事,但它已经看到了。这就是创造力的来源:从看似非常不同的事物之间看到这些类比。所以我认为 GPT-4 在变得更大之后会非常有创造力。我认为那种认为它只是重复所学内容、拼凑已有文本的想法是完全错误的。它会比人类更有创造力。
What these big language models are doing is they're looking for common structure. And by finding common structure, they can encode things using the common structure, and that's more efficient. So let me give you an example. If you ask GPT-4, 'Why is a compost heap like an atom bomb?' Most people can't answer that. Most people haven't thought; they think atom bombs and compost heaps are very different things. But GPT-4 will tell you: 'Well, the energy scales are very different and the time scales are very different, but the thing that's the same is that when the compost heap gets hotter, it generates heat faster, and when the atom bomb produces more neutrons, it produces more neutrons faster.' So it gets the idea of a chain reaction. And I believe it's understood they're both forms of chain reaction. It's using that understanding to compress all that information into its weights. And if it's doing that, then it's going to be doing that for hundreds of things where we haven't seen the analogies yet, but it has. And that's where you get creativity from: from seeing these analogies between apparently very different things. So I think GPT-4 is going to end up, when it gets bigger, being very creative. I think this idea that it's just regurgitating what it's learned, just pasting together text it's learned already, that's completely wrong. It's going to be even more creative than people.
我想你会认为它不会只是重复我们迄今发展的人类知识,还能超越它。我认为我们还没完全看到这一点。我们开始看到一些例子,但在很大程度上我们仍然停留在当前的科学水平。你认为什么能让它超越?
I think you'd argue that it won't just repeat the human knowledge we've developed so far, but could also progress beyond that. I think that's something we haven't quite seen yet. We've started seeing some examples of it, but to a large extent we're sort of still at the current level of science. What do you think will enable it to go beyond that?
嗯,我们在更有限的场景中看到过。比如,AlphaGo 在与李世石的那场著名比赛中,第 37 手,AlphaGo 下了一手所有专家都认为肯定是失误的棋,但后来他们意识到那是一步妙棋。所以那是在那个有限领域内创造出来的。我认为随着这些东西变大,我们会看到更多这样的例子。AlphaGo 的不同之处还在于它使用了强化学习,这随后让它超越了当前状态。它从模仿学习开始,观察人类如何下棋,然后通过自我对弈发展出远超人类水平的棋艺。
Well, we've seen that in more limited contexts. For example, take AlphaGo in that famous competition with Lee Sedol. There was move 37, where AlphaGo made a move that all the experts said must have been a mistake, but actually later they realized it was a brilliant move. So that was created within that limited domain. I think we'll see a lot more of that as these things get bigger. The difference with AlphaGo as well was that it was using reinforcement learning, that subsequently sort of enabled it to go beyond the current state. So it started with imitation learning, watching how humans play the game, and then it would through self-play develop way beyond that.
你认为那是缺失的组件吗?
Do you think that's the missing component?
我认为那很可能是一个缺失的组件,是的。AlphaGo 和 AlphaZero 中的自我对弈是它们能做出这些创造性棋步的重要原因。但我不认为它是完全必要的。很久以前我做了一个小实验,训练神经网络识别手写数字——我喜欢那个例子,MNIST 例子——你给它一半答案错误的训练数据。问题是它能学得多好……
I think that may well be a missing component, yes. That the self-play in AlphaGo and AlphaZero are a large part of why it could make these creative moves. But I don't think it's entirely necessary. So there's a little experiment I did a long time ago where you train a neural net to recognize handwritten digits—I love that example, the MNIST example—and you give it training data where half the answers are wrong. And the question is how well will it...
你学习时,让一半的答案故意错一次并保持那样,这样它就不能通过有时看到正确答案、有时看到错误答案来平均掉错误。当它看到那个例子时,一半的情况下答案总是错的,所以训练数据有 50%的错误率。但如果你用反向传播训练,它能降到 5%或更低的错误率。换句话说,从错误标注的数据中,它能得到好得多的结果。它能看出训练数据是错的。这就是为什么聪明的学生能比他们的导师更聪明。导师告诉他们一堆东西,对于其中一半,他们觉得‘不,胡说’,然后只听另一半,结果就比导师更聪明了。所以这些大型神经网络实际上能比它们的训练数据做得好得多,而大多数人都没意识到这一点。
You learn and you make half the answers wrong once and keep them like that so it can't average away the wrongness by just seeing the same example but with the right answer sometimes and the wrong answer sometimes. When it sees that example, half of the examples when it sees the example, the answer is always wrong, and so the training data has 50% error. But if you train up backpropagation, it gets down to 5% error or less. In other words, from badly labeled data it can get much better results. It can see that the training data is wrong. And that's how smart students can be smarter than their advisor. Their advisor tells them all this stuff, and for half of what their advisor tells them, they think 'no, rubbish,' and they listen to the other half, and then they end up smarter than the advisor. So these big neural nets can actually do much better than their training data, and most people don't realize that.
那么你期望这些模型如何加入推理能力呢?我的意思是,一种方法是在它们之上添加某种启发式规则,很多研究现在就在这么做,比如思维链,把它的推理反馈给它自己。另一种方式是在模型本身内部,随着你不断扩展规模。你对此有什么直觉?
So how do you expect these models to add reasoning into them? I mean, one approach is you add sort of heuristics on top of them, which a lot of the research is doing now, where you have sort of chain of thought, you feed its reasoning back into itself. And another way would be in the model itself, as you scale, scale, scale it up. What's your intuition around that?
我的直觉是,随着我们扩大这些模型的规模,它们会变得更擅长推理。如果你问人类是如何工作的,大致来说,我们有这些直觉,也能进行推理,并用推理来纠正我们的直觉。当然,我们在推理过程中也会用到直觉来执行推理,但如果推理的结论与我们的直觉冲突,我们就意识到直觉需要改变。这很像 AlphaGo 或 AlphaZero,它们有一个评估函数,只看棋盘就能判断局面对我有多好,但随后你进行蒙特卡洛展开,得到更精确的评估,然后修正评估函数。所以你可以通过让评估函数与推理结果一致来训练它。我认为这些大型语言模型必须开始这样做。它们必须通过推理并发现直觉不对,来训练自己关于下一个词该是什么的原始直觉。这样它们就能获得比单纯模仿人类行为更多的训练数据。这正是 AlphaGo 能走出那步创造性的第 37 手的原因——它拥有更多的训练数据,因为它用推理来检验正确的下一步应该是什么。
My intuition is that as we scale up these models, they get better at reasoning. And if you ask how people work, roughly speaking, we have these intuitions and we can do reasoning, and we use the reasoning to correct our intuitions. Of course, we use the intuitions during the reasoning to do the reasoning, but if the conclusion of the reasoning conflicts with our intuitions, we realize the intuitions need to be changed. That's much like in AlphaGo or AlphaZero, where you have an evaluation function that just looks at a board and says how good that is for me, but then you do the Monte Carlo rollout and now you get a more accurate idea, and you can revise your evaluation function. So you can train it by getting it to agree with the results of reasoning. And I think these large language models have to start doing that. They have to start training their raw intuitions about what should come next by doing reasoning and realizing that's not right. And that way they can get more training data than just mimicking what people did. And that's exactly why AlphaGo could do that creative move 37; it had much more training data because it was using reasoning to check out what the right next move should have been.
你怎么看多模态?我们谈到了这些类比,通常这些类比远超我们所能看到的。它在发现远超人类的类比,而且是在我们永远无法理解的抽象层次上。现在当我们引入图像、视频和声音,你认为这将如何改变模型,又将如何改变它能够做出的类比?
What do you think about multimodality? We spoke about these analogies, and often the analogies are way beyond what we could see. It's discovering analogies that are far beyond humans, and at abstraction levels that we'll never be able to understand. Now when we introduce images, video, and sound, how do you think that will change the models, and how do you think it will change the analogies it will be able to make?
我认为这会带来很大变化。我认为它会大大提升模型理解空间事物的能力,比如。仅靠语言,理解某些空间概念相当困难,尽管 GPT-4 在成为多模态之前就已经能做到了。但当你让它多模态化,如果它既能视觉又能伸手抓取东西,那么它能通过拿起物体、翻转等操作更好地理解物体。所以虽然你可以从语言中学到很多,但多模态学习更容易,实际上你需要的语言也更少。而且有大量的 YouTube 视频可以用来预测下一帧之类的。所以我认为这些多模态模型显然会占据主导地位。通过这种方式你可以获得更多数据,它们需要的语言也更少。所以这里有一个哲学观点:你可以仅从语言中学到一个很好的模型,但从多模态系统中学习要容易得多。
I think it'll change it a lot. I think it'll make it much better at understanding spatial things, for example. From language alone, it's quite hard to understand some spatial things, although remarkably GPT-4 can do that even before it was multimodal. But when you make it multimodal, if you have it both doing vision and reaching out and grabbing things, it'll understand objects much better if it can pick them up and turn them over and so on. So although you can learn an awful lot from language, it's easier to learn if you're multimodal, and in fact you then need less language. And there's an awful lot of YouTube video for predicting the next frame or something like that. So I think these multimodal models are clearly going to take over. You can get more data that way, they need less language. So there's really a philosophical point that you could learn a very good model from language alone, but it's much easier to learn it from a multimodal system.
你认为这将如何影响模型的推理能力?
How do you think it will impact the model's reasoning?
我认为这会大大提升它对空间推理的能力,比如推理拿起物体会发生什么。如果你真的尝试拿起物体,你会得到各种有助于推理的训练数据。
I think it'll make it much better at reasoning about space, for example. Reasoning about what happens if you pick objects up. If you actually try picking objects up, you're going to get all sorts of training data that's going to help.
你认为人类大脑是为了很好地处理语言而进化出来的,还是语言是为了很好地适应人类大脑而进化出来的?
Do you think the human brain evolved to work well with language, or do you think language evolved to work well with the human brain?
我认为语言是为了适应大脑而进化,还是大脑是为了适应语言而进化,这是一个非常好的问题。我认为两者都发生了。我以前认为我们可以在完全不依赖语言的情况下进行大量认知活动。现在我的想法有点变了。让我给你三种关于语言及其与认知关系的不同观点。第一种是传统的符号主义观点:认知由某种经过清理的、无歧义的逻辑语言中的符号串组成,并应用推理规则,这就是认知——它只是对类似语言符号串的东西进行符号操作。这是一个极端观点。另一个相反的极端观点是:不,一旦进入大脑内部,一切都是向量。符号进来,你把它们转换成大向量,内部所有处理都用大向量完成,然后如果要产生输出,你再次产生符号。大约在 2014 年,机器翻译中人们使用神经循环神经网络,单词不断输入,它们有一个隐藏状态,不断在这个隐藏状态中积累信息。所以当它们到达句子末尾时,它们有一个大的隐藏向量,捕捉了该句子的含义,然后可以用来生成另一种语言的句子。这被称为“思想向量”。这是关于语言的第二种观点:你把语言转换成一个大向量,这个向量与语言完全不同,而认知就是关于这个的。但还有第三种观点,也是我现在相信的观点:你把这些符号转换成嵌入,并使用多层嵌入,从而得到非常丰富的嵌入。但这些嵌入仍然与符号相关联,因为你对这个符号有一个大向量,对那个符号也有一个大向量,这些向量相互作用产生下一个词的符号的向量。这就是理解。理解就是知道如何将符号转换成这些向量,并知道向量的元素应该如何相互作用以预测下一个符号的向量。这就是理解,无论是在这些大型语言模型中还是在我们的大脑中。这是一个介于两者之间的例子。你仍然使用符号,但你把它们解释为大向量,而所有工作都在这里。所有的知识都在于你使用什么向量以及这些向量的元素如何相互作用,而不是在符号规则中。但这并不是说完全摆脱符号;而是说你把符号转换成向量,然后在向量空间中进行所有计算。
I think the question of whether language evolved to work with the brain or the brain evolved to work with language is a very good question. I think both happened. I used to think we would do a lot of cognition without needing language at all. Now I've changed my mind a bit. So let me give you three different views of language and how it relates to cognition. There's the old-fashioned symbolic view, which is cognition consists of having strings of symbols in some kind of cleaned-up logical language where there's no ambiguity, and applying rules of inference, and that's what cognition is—it's just these symbolic manipulations on things that are like strings of language symbols. So that's one extreme view. An opposite extreme view is no, once you get inside the head, it's all vectors. So symbols come in, you convert those symbols into big vectors, and all the stuff inside is done with big vectors, and then if you want to produce output, you produce symbols again. There was a point in machine translation in about 2014 when people were using neural recurrent neural nets, and words would keep coming in, and they have a hidden state, and they keep accumulating information in this hidden state. So when they got to the end of a sentence, they had a big hidden vector that captures the meaning of that sentence, which could then be used for producing the sentence in another language. That was called a thought vector. And that's a sort of second view of language: you convert the language into a big vector that's nothing like language, and that's what cognition is all about. But then there's a third view, which is what I believe now, which is that you take these symbols and you convert the symbols into embeddings, and you use multiple layers of that, so you get these very rich embeddings. But the embeddings are still tied to the symbols in the sense that you've got a big vector for this symbol and a big vector for that symbol, and these vectors interact to produce the vector for the symbol for the next word. And that's what understanding is. Understanding is knowing how to convert the symbols into these vectors and knowing how the elements of the vector should interact to predict the vector for the next symbol. That's what understanding is, both in these big language models and in our brains. And that's an example which is sort of in between. You're staying with the symbols, but you're interpreting them as these big vectors, and that's where all the work is. And all the knowledge is in what vectors you use and how the elements of those vectors interact, not in symbolic rules. But it's not saying that you get away from the symbols altogether; it's saying you turn the symbols into vectors and then do all the computation in vector space.
你是最早想到用 GPU 的人之一。我知道黄仁勋为此很感激你。早在 2009 年,你提到你告诉黄仁勋,这对训练神经网络来说可能是个好主意。请带我们回顾一下你早期使用 GPU 训练神经网络的直觉。
You were one of the first folks to get the idea of using GPUs. I know Jensen loves you for that. Back in 2009, you mentioned that you told Jensen that this could be a quite good idea for training neural nets. Take us back to that early intuition of using GPUs for training neural nets.
实际上,我想大概在 2006 年,我有个以前的学生叫 Rick Zisy,他是个非常棒的计算机视觉专家。我在一次会议上和他聊,他说:‘你知道吗,你应该考虑用图形处理卡,因为它们非常擅长矩阵乘法,而你所做的基本上全是矩阵乘法。’所以我思考了一下。然后我们了解到这些 Tesla 系统,里面装了四个 GPU。最初我们只买了游戏 GPU,发现它们让速度提升了 30 倍。然后我们买了一个带 4 个 GPU 的 Tesla 系统,在上面做语音识别,效果非常好。接着在 2009 年,我在 NIPS 上做了一个演讲,告诉一千名机器学习研究者:‘你们都应该去买 Nvidia GPU;它们是未来,你们做机器学习需要它们。’然后我真的给 Nvidia 发了邮件说:‘我告诉了一千名机器学习研究者买你们的板子,能送我一块免费的吗?’他们说不——实际上,他们没有说不,只是没回复。但后来我把这个故事告诉黄仁勋时,他送了我一块免费的。这非常非常好。
Actually, I think in about 2006, I had a former graduate student called Rick Zisy, who's a very good computer vision guy. I talked to him in a meeting, and he said, 'You know, you ought to think about using graphics processing cards because they're very good at matrix multiplies, and what you're doing is basically all matrix multiplies.' So I thought about that for a bit. Then we learned about these Tesla systems that had four GPUs in them. Initially, we just got gaming GPUs and discovered they made things go 30 times faster. Then we bought one of these Tesla systems with 4 GPUs, and we did speech on that, and it worked very well. Then in 2009, I gave a talk at NIPS and I told a thousand machine learning researchers, 'You should all go and buy Nvidia GPUs; they're the future, you need them for doing machine learning.' And I actually then sent mail to Nvidia saying, 'I told a thousand machine learning researchers to buy your boards; could you give me a free one?' And they said no—actually, they didn't say no, they just didn't reply. But when I told Jensen this story later on, he gave me a free one. That's very, very good.
我觉得有趣的是 GPU 如何与这个领域共同演进。那么你认为我们在算力方面下一步应该怎么走?
I think what's interesting is how GPUs have evolved alongside the field. So where do you think we should go next in compute?
我在谷歌的最后几年,一直在思考如何实现模拟计算,这样我们就能用大约 30 瓦(像大脑一样)而不是兆瓦级的功耗,在模拟硬件上运行这些大型语言模型。我从未成功过,但我开始真正欣赏数字计算。如果你要用那种低功耗的模拟计算,每一块硬件都会有点不同,而想法是学习会利用该硬件的特定属性。这正是人类的情况:我们的大脑各不相同。所以我们不能把你大脑中的权重放到我的大脑里;硬件不同,单个神经元的精确属性也不同。学习过程已经学会了利用所有这些。因此,我们是会死的,因为我的大脑中的权重对其他任何大脑都没用;我死后,那些权重就无用了。我们可以通过我说出句子,你弄清楚如何改变你的权重以便你会说出同样的话——这被称为蒸馏——来相当低效地在个体之间传递信息。而数字系统是不朽的,因为一旦你有了权重,你可以扔掉计算机,只需把权重存储在某个地方的磁带上,然后建造另一台计算机,放入相同的权重,如果是数字的,它可以计算出与其他系统完全相同的结果。所以数字系统可以共享权重,这要高效得多。如果你有一大堆数字系统,它们各自做一点点学习,从相同的权重开始,做一点点学习,然后再次共享它们的权重,它们都知道其他所有系统学到了什么。我们做不到这一点。因此,在共享知识方面,它们远胜于我们。
My last couple of years at Google, I was thinking about ways of trying to make analog computation, so that instead of using like a megawatt, we could use like 30 Watts like the brain, and we could run these big language models in analog hardware. I never made it work, but I started really appreciating digital computation. So if you're going to use that low-power analog computation, every piece of hardware is going to be a bit different, and the idea is the learning is going to make use of the specific properties of that hardware. And that's what happens with people: all our brains are different. So we can't then take the weights in your brain and put them in my brain; the hardware is different, the precise properties of the individual neurons are different. The learning used to make has learned to make use of all that. And so we're mortal in the sense that the weights in my brain are no good for any other brain; when I die, those weights are useless. We can get information from one to another rather inefficiently by me producing sentences and you figuring out how to change your weights so you would have said the same thing—that's called distillation, but that's a very inefficient way of communicating knowledge. And with digital systems, they're immortal because once you've got some weights, you can throw away the computer, just store the weights on a tape somewhere, and now build another computer, put those same weights in, and if it's digital, it can compute exactly the same thing as the other system did. So digital systems can share weights, and that's incredibly much more efficient. If you've got a whole bunch of digital systems and they each go and do a tiny bit of learning, and they start with the same weights, they do a tiny bit of learning, and then they share their weights again, they all know what all the others learned. We can't do that. And so they're far superior to us in being able to share knowledge.
这个领域部署的很多想法都是非常老派的;它们是在神经科学中一直存在的想法。你认为还有什么可以应用到我们开发的系统中?
A lot of the ideas that have been deployed in the field are very old-school ideas; they're ideas that have been around in neuroscience forever. What do you think is left to apply to the systems we develop?
我们仍需追赶神经科学的一大方面是变化的时间尺度。在几乎所有神经网络中,有一个快速时间尺度用于改变活动:输入进来,活动(嵌入向量)全部改变,然后有一个慢速时间尺度用于改变权重,那是长期学习。你只有这两个时间尺度。在大脑中,权重变化有许多时间尺度。例如,如果我说一个意想不到的词,比如‘黄瓜’,然后 5 分钟后你戴上耳机,周围有很多噪音,有非常微弱的词语,你会更容易识别出‘黄瓜’这个词,因为我 5 分钟前说过它。那么大脑中的这个知识在哪里?这个知识显然是在突触的临时变化中;不是神经元在反复说‘黄瓜、黄瓜、黄瓜’——你没有足够的神经元来做这个。它是在权重的临时变化中。你可以用快速的临时权重变化做很多事情——我称之为快速权重。我们在这些神经模型中没有这样做。我们之所以不这样做,是因为如果权重有依赖于输入数据的临时变化,那么你就无法同时处理大量不同的案例。目前,我们取大量不同的字符串,把它们堆叠在一起,并行处理,因为这样我们可以做矩阵-矩阵乘法,这要高效得多。正是这种效率阻止了我们使用快速权重。但大脑显然使用快速权重进行临时记忆,你可以用那种方式做很多我们目前没有做的事情。我认为这是我们必须学习的最重要的事情之一。我曾对 Graphcore 这样的东西寄予厚望,如果它们采用顺序处理并只做在线学习,那么它们就可以使用快速权重。但这还没有成功。我认为当人们使用电导作为权重时,最终会成功的。
One big thing that we still have to catch up with neuroscience on is the time scales for changes. In nearly all neural nets, there's a fast time scale for changing activities: input comes in, the activities (the embedding vectors) all change, and then there's a slow time scale which is changing the weights, and that's long-term learning. You just have those two time scales. In the brain, there are many time scales at which weights change. For example, if I say an unexpected word like 'cucumber,' and now 5 minutes later you put headphones on, there's a lot of noise and there are very faint words, you'll be much better at recognizing the word 'cucumber' because I said it 5 minutes ago. So where is that knowledge in the brain? That knowledge is obviously in temporary changes to synapses; it's not neurons going 'cucumber, cucumber, cucumber'—you don't have enough neurons for that. It's in temporary changes to the weights. And you can do a lot of things with temporary weight changes fast—what I call fast weights. We don't do that in these neural models. The reason we don't do it is because if you have temporary changes to the weights that depend on the input data, then you can't process a whole bunch of different cases at the same time. At present, we take a whole bunch of different strings, we stack them together, and we process them all in parallel because then we can do matrix-matrix multiplies, which is much more efficient. And just that efficiency is stopping us from using fast weights. But the brain clearly uses fast weights for temporary memory, and there are all sorts of things you can do that way that we don't do at present. I think that's one of the biggest things we have to learn. I was very hopeful that things like Graphcore, if they went sequential and did just online learning, then they could use fast weights. But that hasn't worked out yet. I think it'll work out eventually when people are using conductances for weights.
了解这些模型的工作原理和大脑的工作原理如何影响了你的思维方式?
How has knowing how these models work and knowing how the brain works impacted the way you think?
我认为有一个重大影响,是在相当抽象的层面上。多年来,人们对拥有一个大的随机神经网络,然后只给大量训练数据,它就能学会做复杂事情的想法非常不屑。如果你和统计学家、语言学家或大多数 AI 人士交谈,他们会说那只是一个白日梦;没有某种先天知识,没有大量架构限制,你不可能学会做真正复杂的事情。事实证明这完全错了。你可以拿一个大的随机神经网络,仅从数据中就能学到一大堆东西。所以,随机梯度下降——使用梯度反复调整权重——会学到东西,而且会学到大的复杂东西,这个想法已经被这些大模型验证了。关于大脑,这是一个非常重要的事情:它不必拥有所有这些先天结构。
I think there's been one big impact, which is at a fairly abstract level. For many years, people were very scornful about the idea of having a big random neural net and just giving a lot of training data, and it would learn to do complicated things. If you talk to statisticians or linguists or most people in AI, they say that's just a pipe dream; there's no way you're going to learn to do really complicated things without some kind of innate knowledge, without a lot of architectural restrictions. It turns out that's completely wrong. You can take a big random neural network and you can learn a whole bunch of stuff just from data. So the idea that stochastic gradient descent—repeatedly adjusting the weights using a gradient—will learn things, and will learn big complicated things, has been validated by these big models. And that's a very important thing to know about the brain: it doesn't have to have all this innate structure.
它有很多先天结构,但肯定不需要先天结构来学习那些容易学的东西。所以乔姆斯基的那种观点——除非语言已经全部预先布线并成熟,否则你学不会任何像语言这样复杂的东西——这种观点现在显然是胡说八道。
It's got a lot of innate structure but it certainly doesn't need innate structure for things that are easily learned. And so the sort of idea coming from Chomsky that you won't learn anything complicated like language unless it's all kind of wired in already and just matures, that idea is now clearly nonsense.
我敢肯定乔姆斯基会感谢你称他的观点为胡说八道。
嗯,实际上我认为乔姆斯基的很多政治观点都很明智。我很惊讶,一个对中东问题有如此明智见解的人,在语言学上竟然错得这么离谱。
Well, I think actually a lot of Chomsky's political ideas are very sensible. And I was struck by how someone with such sensible ideas about the Middle East could be so wrong about linguistics.
你认为什么能让这些模型更有效地模拟人类的意识?但想象一下,你有一个你一生都在与之交谈的 AI 助手,而不是像 ChatGPT 那样总是删除对话记忆、重新开始,它具备自我反思能力。在某个时刻你去世了,你把这个消息告诉助手。你认为那个助手会在那一刻有感觉吗?
What do you think would make these models simulate consciousness of humans more effectively? But imagine you had the AI assistant that you've spoken to in your entire life, and instead of that being like ChatGPT that deletes the memory of the conversation and you start fresh all the time, it had self-reflection. At some point you pass away and you tell that to the assistant. Do you think that assistant would feel at that point?
是的,我认为它们也可以有感觉。所以我认为,就像我们对感知有一个内在剧场模型一样,我们对感觉也有一个内在剧场模型——它们是我能体验但别人无法体验的东西。我认为那个模型同样是错误的。所以假设我说“我想打加里的鼻子”,我经常这么说。让我们试着从内在剧场的概念中抽象出来。我真正想告诉你的是,如果不是因为来自额叶的抑制,我会执行一个动作。所以当我们谈论感觉时,我们实际上是在谈论如果没有约束我们会执行的动作。而这正是感觉的本质——如果没有约束我们会做的事情。所以我认为你可以对感觉给出同样的解释,而且没有理由说这些东西不能有感觉。事实上,在 1973 年,我看到一个机器人表现出情绪。在爱丁堡,他们有一个带两个这样夹爪的机器人,如果你把零件分开放在一块绿色毡布上,它可以组装一辆玩具车。但如果你把它们堆成一堆,它的视觉不足以弄清楚情况,所以它猛地一夹,把它们打散,然后就能组装起来。如果你在一个人身上看到这种情况,你会说它因为不理解情况而生气,所以破坏了它。这很深刻。
Yes, I think they can have feelings too. So I think just as we have this inner theater model for perception, we have an inner theater model for feelings—they're things that I can experience but other people can't. I think that model is equally wrong. So suppose I say 'I feel like punching Gary on the nose,' which I often do. Let's try to abstract that away from the idea of an inner theater. What I'm really saying to you is, if it weren't for the inhibition coming from my frontal lobes, I would perform an action. So when we talk about feelings, we're really talking about actions we would perform if it weren't for constraints. And that's really what feelings are—the actions we would do if it weren't for constraints. So I think you can give the same kind of explanation for feelings, and there's no reason why these things can't have feelings. In fact, in 1973 I saw a robot having an emotion. In Edinburgh, they had a robot with two grippers like this that could assemble a toy car if you put the pieces separately on a piece of green felt. But if you put them in a pile, its vision wasn't good enough to figure out what was going on, so it put its gripper whack and knocked them so they were scattered, and then it could put them together. If you saw that in a person, you'd say it was cross with the situation because it didn't understand it, so it destroyed it. That's profound.
你之前把人类和大语言模型描述为类比机器。你认为你一生中发现的最有力的类比是什么?
You previously described humans and LLMs as analogy machines. What do you think has been the most powerful analogies that you found throughout your life?
哦,在我的一生中……我想大概是一个相当弱的类比对我影响很大,那就是宗教信仰与符号处理信仰之间的类比。我小时候来自一个无神论家庭,上学后接触到宗教信仰,在我看来完全是胡说八道。现在仍然觉得是胡说八道。当我看到符号处理被用来解释人类如何工作时,我认为这同样是胡说八道。现在我不觉得那么胡说八道了,因为我认为我们确实在做符号处理——只是我们通过给符号赋予大的嵌入向量来实现。但我们确实在做符号处理,完全不是人们想象的那种方式,即匹配符号,而符号的唯一属性就是它与另一个符号相同或不同。符号只有那个属性。我们根本不是那样做的。我们利用上下文给符号赋予嵌入向量,然后利用这些嵌入向量的分量之间的相互作用来进行思考。但谷歌有一位非常优秀的研究员叫费尔南多·佩雷拉,他说:“是的,我们确实有符号推理,而我们唯一的符号就是自然语言。自然语言是一种符号语言,我们用推理。”我现在相信这一点。
Oh, throughout my life... I guess probably a weak analogy that's influenced me a lot is the analogy between religious belief and belief in symbol processing. So when I was very young, I came from an atheist family and went to school and was confronted with religious belief, and it just seemed nonsense to me. It still seems nonsense to me. And when I saw symbol processing as an explanation of how people worked, I thought it was just the same nonsense. I don't think it's quite so much nonsense now, because I think actually we do do symbol processing—it's just we do it by giving these big embedding vectors to the symbols. But we are actually symbol processing, not at all in the way people thought, where you match symbols and the only property a symbol has is it's identical to another symbol or it's not identical. That's the only property a symbol has. We don't do that at all. We use the context to give embedding vectors to symbols and then use the interactions between the components of these embedding vectors to do thinking. But there's a very good researcher at Google called Fernando Pereira who said, 'Yes, we do have symbolic reasoning, and the only symbolic we have is natural language. Natural language is a symbolic language, and we reason with it.' And I believe that now.
你做了计算机科学史上一些最有意义的研究。你能谈谈你是如何选择正确的问题来研究的吗?
You've done some of the most meaningful research in the history of computer science. Can you walk us through how you select the right problems to work on?
嗯,首先让我纠正你。我和我的学生做了很多最有意义的事情,这主要是与学生的良好合作。我挑选优秀学生的能力来自于这样一个事实:在 70 年代、80 年代、90 年代和 2000 年代,做神经网络的人很少,所以少数做神经网络的人能够挑选最优秀的学生。所以那是一种运气。但我选择问题的方法基本上是:你知道,当科学家谈论他们如何工作时,他们会有关于自己如何工作的理论,这些理论可能和实际情况没太大关系。但我的理论是,我寻找那些每个人都同意某件事,但感觉不对劲的地方——就是有一种直觉,觉得有些地方不对。然后我就研究那个问题,看看能否阐述为什么我认为它错了,也许我可以用一个小型计算机程序做一个演示,证明它并不像你预期的那样工作。举个例子。大多数人认为,如果你给神经网络添加噪声,它的表现会更差。例如,每次你输入一个训练样本,让一半的神经元静默,它的表现会更差。实际上,我们知道这样做会提高泛化能力,你可以用一个简单的例子来证明。这就是计算机模拟的好处:你可以展示你原来的想法——添加噪声会让它更差,丢弃一半的神经元短期内会让它更差——但如果你这样训练,最终它会表现得更好。你可以用一个小型计算机程序来证明这一点,然后你可以深入思考为什么会这样,以及它如何阻止大型复杂的共适应。但我认为这就是我的工作方法:找到听起来可疑的东西,然后研究它,看看能否给出一个简单的演示,说明为什么它是错的。
Well, first let me correct you. Me and my students have done a lot of the most meaningful things, and it's mainly been a very good collaboration with students. My ability to select very good students came from the fact that there were very few people doing neural nets in the 70s, 80s, 90s, and 2000s, so the few people doing neural nets got to pick the very best students. So that was a piece of luck. But my way of selecting problems is basically: you know, when scientists talk about how they work, they have theories about how they work which probably don't have much to do with the truth. But my theory is that I look for something where everybody's agreed about something and it feels wrong—just a slight intuition that there's something wrong about it. And then I work on that and see if I can elaborate why it is I think it's wrong, and maybe I can make a little demo with a small computer program that shows that it doesn't work the way you might expect. So let me take one example. Most people think that if you add noise to a neural net, it's going to work worse. If, for example, each time you put a training example through, you make half of the neurons be silent, it'll work worse. Actually, we know it'll generalize better if you do that, and you can demonstrate that in a simple example. That's what's nice about computer simulation: you can show that this idea you had—that adding noise is going to make it worse, and dropping out half the neurons will make it work worse in the short term—but if you train it like that, in the end it'll work better. You can demonstrate that with a small computer program, and then you can think hard about why that is and how it stops big elaborate co-adaptations. But I think that's my method of working: find something that sounds suspicious and work on it, and see if you can give a simple demonstration of why it's wrong.
现在什么听起来可疑?
What sounds suspicious to you now?
嗯,我们不用快速权重这一点听起来可疑。我们只有这两个时间尺度——这完全是错的。这根本不像大脑。从长远来看,我认为我们将需要更多的时间尺度。所以这是一个例子。
Well, that we don't use fast weights sounds suspicious. That we only have these two time scales—that's just wrong. That's not at all like the brain. And in the long run, I think we're going to have to have many more time scales. So that's an example.
如果你今天有你的学生团队,他们来找你说:“我们之前讨论过的汉明问题——你所在领域最重要的问题是什么?”你会建议他们接下来研究什么?我们谈到了推理时间尺度。你会给他们什么最高优先级的问题?
If you had your group of students today and they came to you and said, 'So the Hamming question that we talked about previously—what's the most important problem in your field?' What would you suggest that they take on and work on next? We spoke about reasoning time scales. What would be the highest priority problem that you'd give them?
对我来说,现在还是过去大约 30 年来我一直思考的同一个问题:大脑是否在做反向传播?我相信是的。
For me right now, it's the same question I've had for the last like 30 years or so, which is: does the brain do backpropagation? I believe it does.
大脑会获取梯度。如果你得不到梯度,学习效果就会比能获取梯度时差得多。但大脑是如何获取梯度的呢?它是在以某种方式实现反向传播的近似版本,还是某种完全不同的技术?这是一个重大的未解之谜。如果我还继续做研究,这就是我会研究的方向。
The brain is getting gradients. If you don't get gradients, your learning is just much worse than if you do get gradients. But how is the brain getting gradients? Is it somehow implementing some approximate version of backpropagation, or is it some completely different technique? That's a big open question. And if I kept on doing research, that's what I would be doing research on.
现在回顾你的职业生涯,你在很多事情上都是对的。但你有没有在哪些事情上错了,并且希望自己少花些时间在那个方向上?好吧,这是两个不同的问题:一是你哪里错了,二是你是否希望自己少花时间在上面。我认为我在玻尔兹曼机上错了,但我很高兴花了很长时间研究它。关于如何获取梯度,有比反向传播更优美的理论。反向传播只是普通且合理的,就是链式法则。玻尔兹曼机很巧妙,是一种非常有趣的获取梯度的方法,我很希望大脑就是这样工作的,但我认为并非如此。
And when you look back at your career now, you've been right about so many things. But what were you wrong about that you wish you sort of spent less time pursuing a certain direction? Okay, those are two separate questions: one is what were you wrong about, and two, do you wish you'd spent less time on it? I think I was wrong about Boltzmann machines, and I'm glad I spent a long time on it. There are much more beautiful theories of how you get gradients than backpropagation. Backpropagation is just ordinary and sensible, it's just the chain rule. Boltzmann machines is clever and a very interesting way to get gradients, and I would love for that to be how the brain works, but I think it isn't.
你有没有花很多时间想象系统发展之后会发生什么?你有没有想过,如果我们能让这些系统真正发挥作用,我们可以普及教育,让知识更容易获取,解决医学上的一些难题?还是说,你更关心的是理解大脑?
Did you spend much time imagining what would happen post the systems developing as well? Did you have an idea that okay, if we could make these systems work really well, we could democratize education, we could make knowledge way more accessible, we could solve some tough problems in medicine? Or was it more to you about understanding the brain?
是的,我有点觉得科学家应该做有助于社会的事情。但实际上,这并不是做出最佳研究的方式。你做出最佳研究是出于好奇心,你只是必须理解某件事。最近,我意识到这些东西既能带来很多好处,也能造成很多危害,我开始更加担心它们对社会的影响。但这并不是当初激励我的原因。我只是想弄明白大脑究竟是如何学会做事的。这就是我想知道的。我某种程度上失败了。作为那个失败的副作用,我们得到了一些不错的工程成果。但没错,这对世界来说是一次好的失败。
Yes, I sort of feel scientists ought to be doing things that are going to help society. But actually, that's not how you do your best research. You do your best research when it's driven by curiosity, you just have to understand something. Much more recently, I've realized these things could do a lot of harm as well as a lot of good, and I've become much more concerned about the effects they're going to have on society. But that's not what was motivating me. I just wanted to understand how on Earth can the brain learn to do things. That's what I want to know. And I sort of failed. As a side effect of that failure, we got some nice engineering. But yeah, it was a good failure for the world.
如果从事情可能发展得很好的角度来看,你认为最有前景的应用是什么?
If you take the lens of the things that could go really right, what do you think are the most promising applications?
我认为医疗健康显然是一个大领域。在医疗健康方面,社会能吸收的医疗资源几乎是无限的。比如一个老人,他可以全职用五个医生。所以当 AI 在做事上比人更强时,你希望它在那些需要大量资源的领域变得更好。而我们需要更多的医生。如果每个人都有自己的三个医生,那将非常棒,我们正在走向那个阶段。所以这是医疗健康好的一个原因。还有新的工程:开发新材料,比如更好的太阳能电池板或超导材料,或者仅仅是理解人体如何运作。这些都会产生巨大影响。这些都是好事。我担心的是坏人利用它们做坏事。我们为像普京、习近平或特朗普这样的人提供了便利,让他们用 AI 制造杀人机器人、操纵公众舆论或进行大规模监控。这些都是非常令人担忧的事情。
I think healthcare is clearly a big one. With healthcare, there's almost no end to how much healthcare society can absorb. If you take someone old, they could use five doctors full-time. So when AI gets better than people at doing things, you'd like it to get better in areas where you could do with a lot more of that stuff. And we could do with a lot more doctors. If everybody had three doctors of their own, that would be great, and we're going to get to that point. So that's one reason why healthcare is good. There's also just new engineering: developing new materials, for example, for better solar panels or for superconductivity, or for just understanding how the body works. There's going to be huge impacts there. Those are all going to be good things. What I worry about is bad actors using them for bad things. We've facilitated people like Putin or Xi or Trump using AI for killer robots, or for manipulating public opinion, or for mass surveillance. And those are all very worrying things.
你是否担心放缓该领域也会减缓积极方面的发展?
Are you ever concerned that slowing down the field could also slow down the positives?
哦,当然。而且我认为这个领域放缓的可能性不大,部分原因是它是国际性的。如果一个国家放缓,其他国家不会放缓。所以中美之间显然存在一场竞赛,双方都不会减速。所以,是的,我不……我的意思是,有一份请愿书说我们应该放缓六个月。我没有签署,只是因为我认定这永远不会发生。也许我本该签署,因为即使它永远不会发生,它也有政治意义。提出你知道得不到的东西来表明立场通常是好的。但我不认为我们会放缓。
Oh absolutely. And I think there's not much chance that the field will slow down, partly because it's international. If one country slows down, the other countries aren't going to slow down. So there's a race clearly between China and the US, and neither is going to slow down. So yeah, I don't... I mean, there was this petition saying we should slow down for six months. I didn't sign it just because I thought it was never going to happen. I maybe should have signed it because even though it was never going to happen, it made a political point. It's often good to ask for things you know you can't get just to make a point. But I didn't think we're going to slow down.
你认为这种辅助将如何影响 AI 研究过程?
How do you think that it will impact the AI research process, having this assistance?
我认为它会大大提高效率。当你拥有这些助手来帮助你编程,同时也帮助你思考问题,可能还在方程方面帮大忙时,研究将变得高效得多。
I think it'll make it a lot more efficient. Research will get a lot more efficient when you've got these assistants that help you program, but also help you think through things and probably help you a lot with equations too.
你有没有反思过选拔人才的过程?这对你来说主要是凭直觉吗,比如当伊利亚出现时,你觉得‘这是个聪明人,我们一起干吧’?
Have you reflected much on the process of selecting talent? Has that been mostly intuitive to you, like when Ilya shows up at the door, you feel 'this is smart guy, let's work together'?
对于选拔人才,有时你就是知道。和伊利亚聊了没多久,他就显得非常聪明。再多聊一会儿,他显然非常聪明,直觉很好,数学也很棒。所以那是不假思索的。还有一次,我在 NIPS 会议上。我们有一个海报,有个人走过来开始问关于海报的问题,他问的每个问题都对我们做错的地方有深刻的洞察。五分钟后,我就给了他一个博士后职位。那个人就是大卫·麦凯,他非常出色。他去世了,非常令人悲伤。但很明显你会想要他。其他时候就没那么明显了。我学到的一件事是,人是不同的;好学生不止一种类型。有些学生不太有创造力,但技术极强,能让任何东西工作。有些学生技术不强,但非常有创造力。当然你想要两者兼备的,但你不总能得到。但我认为实际上实验室需要各种不同类型的研究生。但我仍然凭直觉,有时你和某人交谈,他们就是非常非常……他们就是懂。那些就是你想要的人。
For selecting talent, sometimes you just know. So after talking to Ilya for not very long, he seemed very smart. And then talking to him a bit more, he clearly was very smart and had very good intuitions as well as being good at math. So that was a no-brainer. There's another case where I was at an NIPS conference. We had a poster and someone came up and he started asking questions about the poster, and every question he asked was a sort of deep insight into what we'd done wrong. And after 5 minutes, I offered him a postdoc position. That guy was David MacKay, who was just brilliant. And it's very sad he died. But it was very obvious you'd want him. Other times it's not so obvious. And one thing I did learn was that people are different; there's not just one type of good student. So there's some students who aren't that creative but are technically extremely strong and will make anything work. There's other students who aren't technically strong but are very creative. Of course you want the ones who are both, but you don't always get that. But I think actually in the lab you need a variety of different kinds of graduate student. But I still go with my gut intuition that sometimes you talk to somebody and they're just very, very... they just get it. And those are the ones you want.
你认为有些人直觉更好的原因是什么?他们只是比别人有更好的训练数据吗?还是说如何培养直觉?
What do you think is the reason for some folks having better intuition? Do they just have better training data than others? Or how can you develop your intuition?
我认为部分原因是他们不接受胡说八道。所以获得糟糕直觉的一个方法是:相信你听到的一切。那是致命的。你必须能够……我认为有些人是这样做的:他们有一个理解现实的完整框架,当有人告诉他们某事时,他们试图弄清楚它如何融入自己的框架,如果不符,他们就拒绝。这是一个非常好的策略。那些试图吸收所有信息的人最终会得到一个非常模糊的框架,可以相信一切,但那毫无用处。所以我认为实际上,拥有一个强烈的世界观,并试图将新事实纳入你的观点——显然这可能导致深度宗教信仰和致命缺陷,比如我对玻尔兹曼机的信念——但我认为这是正确的做法。如果你有好的直觉,你可以信任它们。你应该信任它们。如果你有坏的直觉,你做什么都没用,所以你不妨也信任它们。
I think it's partly they don't stand for nonsense. So here's a way to get bad intuitions: believe everything you're told. That's fatal. You have to be able to... I think here's what some people do: they have a whole framework for understanding reality, and when someone tells them something, they try to figure out how that fits into their framework, and if it doesn't, they just reject it. And that's a very good strategy. People who try to incorporate whatever they're told end up with a framework that's very fuzzy and can believe everything, and that's useless. So I think actually having a strong view of the world and trying to manipulate incoming facts to fit in with your view — obviously it can lead you into deep religious belief and fatal flaws and so on, like my belief in Boltzmann machines — but I think that's the way to go. If you've got good intuitions, you can trust them. You should trust them. If you've got bad intuitions, it doesn't matter what you do, so you might as well trust them.
当你审视今天正在进行的研究类型时,你认为我们是否把所有鸡蛋都放在一个篮子里,应该更多样化我们的想法,还是说这是最有希望的方向,所以我们应该全力以赴?
When you look at the types of research being done today, do you think we're putting all our eggs in one basket and should diversify our ideas, or is this the most promising direction so we should go all in?
我认为构建大模型并在多模态数据上训练它们,即使只是为了预测下一个词,也是一个非常有前景的方法,我们应该几乎全力以赴。显然,现在有很多人在做这个,也有人在做一些看似疯狂的事情,这很好。但我认为大多数人走这条路没问题,因为它效果很好。
I think having big models and training them on multimodal data, even if it's only to predict the next word, is such a promising approach that we should go pretty much all in on it. Obviously, lots of people are doing it now, and there are people doing apparently crazy things, and that's good. But I think it's fine for most people to follow this path because it's working very well.
你认为学习算法有那么重要吗?是有数百万种方式可以达到人类水平的智能,还是只有少数几种我们需要发现?
Do you think learning algorithms matter that much, or are there millions of ways to reach human-level intelligence, or only a select few we need to discover?
我不知道特定学习算法是否非常重要,或者是否有大量不同的算法都能胜任。在我看来,反向传播在某种意义上是对的——获取梯度,从而改变参数使其工作得更好。这看起来正确,而且非常成功。很可能有其他学习算法能以不同方式获得相同梯度或获得其他梯度,并且也有效。我认为这都是开放的,是一个非常有趣的问题。也许大脑在做别的事情,因为那更容易,但反向传播在某种意义上是对的,而且我们知道它效果很好。
I don't know the answer to whether particular learning algorithms are very important or if there's a great variety that will do the job. It seems to me that backpropagation is in a sense the correct thing to do—getting the gradient so you change a parameter to make it work better. That seems right and has been amazingly successful. There may well be other learning algorithms that are alternative ways of getting that same gradient or getting the gradient to something else, and that also work. I think that's all open and a very interesting issue. Maybe the brain is doing something else because it's easier, but backprop is in a sense the right thing to do, and we know it works really well.
回顾你几十年的研究,你最自豪的是什么?是学生、研究,还是其他什么?
Looking back at your decades of research, what are you most proud of? Is it the students, the research, or something else?
玻尔兹曼机的学习算法。它非常优雅,在实践中可能没有希望,但这是我和特里一起开发时最享受的事情,也是我最自豪的,即使它可能是错的。
The learning algorithm for Boltzmann machines. It's beautifully elegant, maybe hopeless in practice, but it's the thing I enjoyed most developing with Terry, and it's what I'm proudest of, even if it's wrong.
你现在大部分时间在思考什么问题?是看什么 Netflix 节目吗?
What questions do you spend most of your time thinking about now? Is it what to watch on Netflix?