Ilya Sutskever 谈深度学习革命与 AlexNet 论文

Ilya Sutskever on the Deep Learning Revolution and the AlexNet Paper

伊利亚·苏茨克维尔 Ilya Sutskever · Lex Fridman 播客 · 2020-05-08 · 约 97 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 联合创始人兼首席科学家 Ilya Sutskever 讨论深度学习革命的催化时刻、他对神经网络的直觉以及来自人脑的灵感。

Ilya Sutskever, co-founder and chief scientist of OpenAI, discusses the catalytic moment of the deep learning revolution, his intuition about neural networks, and the inspiration from the human brain.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 40)

全文 · Full transcript(中英对照)

引言 Introduction

Host

以下是与伊利亚·苏茨克弗的对话,他是 OpenAI 的联合创始人兼首席科学家,历史上被引用最多的计算机科学家之一,引用次数超过 16.5 万次,在我看来,他是深度学习领域有史以来最杰出、最有洞察力的人之一。在这个世界上,很少有人能像伊利亚那样,让我愿意在深度学习、智能乃至生活等话题上与之交谈和头脑风暴。无论是在麦克风前还是幕后,这都是我的荣幸和乐事。这次对话是在疫情爆发前录制的。对于所有承受着这场危机带来的医疗、心理和经济负担的人们,我向你们送去爱意。保持坚强,我们同在,我们会战胜这一切。这里是人工智能播客。如果你喜欢,请在 YouTube 上订阅,给出五星评价,并在 Patreon 上支持播客,或者直接在 Twitter 上与我联系:@lexfridman。像往常一样,我会先播几分钟广告,中间绝不会插入广告打断对话。希望这不会影响你的收听体验。本期节目由 Cash App 呈现,它是 App Store 排名第一的金融应用。下载时请使用代码 LexPodcast。Cash App 让你可以给朋友汇款、购买比特币、以低至 1 美元投资股市。由于 Cash App 允许你购买比特币,我想提一下,加密货币在货币历史背景下非常迷人。我推荐《货币崛起》这本书,它和有声书都很棒。分类账上的借贷记录始于大约 3 万年前,美元诞生于 200 多年前,而比特币作为第一种去中心化加密货币,仅在 10 多年前发布。因此,从历史来看,加密货币仍处于发展的早期阶段,但它仍在努力并可能重新定义货币的本质。所以,如果你从 App Store 或 Google Play 下载 Cash App 并使用代码 LexPodcast,你将获得 10 美元,Cash App 还会向 FIRST 组织捐赠 10 美元,该组织致力于推动全球青少年的机器人和 STEM 教育。现在,以下是我与伊利亚的对话。

The following is a conversation with Ilya Sutskever, co-founder and chief scientist of OpenAI, one of the most cited computer scientists in history with over 165,000 citations, and to me one of the most brilliant and insightful minds ever in the field of deep learning. There are very few people in this world who I would rather talk to and brainstorm with about deep learning, intelligence, and life in general than Ilya. On and off the mic, this was an honor and a pleasure. This conversation was recorded before the outbreak of the pandemic. For everyone feeling the medical, psychological, and financial burden of this crisis, I'm sending love your way. Stay strong, we're in this together, we'll beat this thing. This is the Artificial Intelligence Podcast. If you enjoy it, subscribe on YouTube, review it with five stars, and have a podcast support it on Patreon, or simply connect with me on Twitter at @lexfridman. As usual, I'll do a few minutes of ads now and never any ads in the middle that can break the flow of the conversation. I hope that works for you and doesn't hurt the listening experience. This show is presented by Cash App, the number one finance app in the App Store. When you get it, use code LexPodcast. Cash App lets you send money to friends, buy Bitcoin, invest in the stock market with as little as one dollar. Since Cash App allows you to buy Bitcoin, let me mention that cryptocurrency in the context of the history of money is fascinating. I recommend Ascent of Money as a great book on this history. Both the book and audiobook are great. Debits and credits on ledgers started around 30,000 years ago, the US dollar created over 200 years ago, and Bitcoin, the first decentralized cryptocurrency, released just over 10 years ago. So given that history, cryptocurrency is still very much in its early days of development, but it's still aiming to and just might redefine the nature of money. So again, if you get Cash App from the App Store or Google Play and use the code LexPodcast, you get $10, and Cash App will also donate $10 to FIRST, an organization that is helping advance robotics and STEM education for young people around the world. And now here's my conversation with Ilya.

AlexNet 论文与早期直觉 The AlexNet Paper and Early Intuitions

Host

你是与亚历克斯·克里热夫斯基和杰夫·辛顿共同撰写著名的 AlexNet 论文的三位作者之一,这篇论文可以说是标志着启动深度学习革命的重大催化时刻。当时,带我们回到那个时代。你对神经网络、对神经网络的表征能力有什么直觉?也许你可以提一下,在接下来的几年里,直到今天,这 10 年间,这种直觉是如何演变的?

You were one of the three authors, with Alex Krizhevsky and Geoff Hinton, of the famed AlexNet paper, which is arguably the paper that marked the big catalytic moment that launched the deep learning revolution. At that time, take us back to that time. What was your intuition about neural networks, about the representational power of neural networks? And maybe you could mention how that evolved over the next few years up to today, over the 10 years.

Ilya

是的,我可以回答这个问题。大约在 2010 年或 2011 年的某个时候,我在脑海中连接了两个事实。基本上,我的认识是这样的:在某个时刻,我们意识到我们可以用反向传播端到端地训练非常大的——我不应该说非常大,你知道,按今天的标准它们很小——但大型和深层的神经网络。在某个时刻,不同的人得到了这个结果。我也得到了这个结果。我第一次意识到深度神经网络很强大的时刻是 2010 年詹姆斯·马丁斯发明了无海森优化器,他端到端地训练了一个 10 层神经网络,没有预训练,从头开始。当那发生时,我想,就是它了。因为如果你能训练一个大型神经网络,一个大型神经网络可以表示非常复杂的函数。因为如果你有一个 10 层的神经网络,就好像允许人脑运行若干毫秒。神经元放电很慢,所以在大概 100 毫秒内,你的神经元只放电 10 次,所以这也类似于 10 层。而在 100 毫秒内,你可以完美地识别任何物体。所以我想,我当时已经有了这个想法:我们需要在大量监督数据上训练一个非常大的神经网络,然后它一定会成功,因为我们可以找到最好的神经网络。而且还有理论说,如果你有比参数更多的数据,你就不会过拟合。今天我们知道这个理论其实非常不完整,当数据少于参数时你会想要过拟合,但肯定的是,如果数据多于参数,你就不会过拟合。所以神经网络严重过参数化的事实并没有让你气馁。

Yeah, I can answer that question. At some point in about 2010 or 2011, I connected two facts in my mind. Basically, the realization was this: at some point we realized that we can train very large—I shouldn't say very, you know, they're tiny by today's standards—but large and deep neural networks end to end with backpropagation. At some point, different people obtained this result. I obtained this result. The first moment in which I realized that deep neural networks are powerful was when James Martens invented the Hessian-free optimizer in 2010, and he trained a 10-layer neural network end-to-end without pre-training, from scratch. And when that happened, I thought, this is it. Because if you can train a big neural network, a big neural network can represent very complicated functions. Because if you have a neural network with 10 layers, it's as though you allow the human brain to run for some number of milliseconds. Neuron firings are slow, and so in maybe 100 milliseconds, your neurons only fire 10 times, so it's also kind of like 10 layers. And in 100 milliseconds, you can perfectly recognize any object. So I thought, so I already had the idea then that we need to train a very big neural network on lots of supervised data, and then it must succeed because we can find the best neural network. And then there's also theory that if you have more data than parameters, you won't overfit. Today we know that actually this theory is very incomplete, and you want overfitting when you have less data than parameters, but definitely if you have more data than parameters, you won't overfit. So the fact that neural networks were heavily over-parameterized wasn't discouraging to you.

Host

所以你当时在想,参数数量巨大这个事实是没问题的。一切都会好的。

So you were thinking about the theory that the number of parameters, the fact there's a huge number of parameters, is okay. It's going to be okay.

Ilya

我的意思是,之前有一些证据表明它还算可以,但理论主要是——理论是如果你有一个大数据集和一个大型神经网络,它就会工作。过参数化并不算什么大问题。我想,对于图像,你只需要添加一些数据增强,就没问题了。

I mean, there was some evidence before that it was okay-ish, but the theory was most—the theory was that if you had a big data set and a big neural net, it was going to work. The over-parameterization just didn't really figure much as a problem. I thought, well, with images you're just going to add some data augmentation, it's going to be okay.

Host

那么怀疑来自哪里呢?

So where was any doubt coming from?

Ilya

主要的怀疑是:我们能训练更大的——我们是否有足够的算力用反向传播训练足够大的神经网络?我认为反向传播会起作用。这一点并不明确。是否有足够的算力来获得一个非常有说服力的结果?然后某个时候,亚历克斯·克里热夫斯基编写了这些极快的 CUDA 内核来训练卷积神经网络,然后就是砰的一声。我们开始吧。我们拿 ImageNet 来训练,这将成为最伟大的事情。

The main doubt was: can we train a bigger—will we have enough compute to train a big enough neural net with backpropagation? Backpropagation I thought would work. This image wasn't clear. Would there be enough compute to get a very convincing result? And then at some point, Alex Krizhevsky wrote these insanely fast CUDA kernels for training convolutional neural nets, and that was bam. Let's do this. Let's get ImageNet and it's going to be the greatest thing.

Host

你的直觉主要来自你自己和他人的实证结果吗,比如实际证明一个程序可以训练一个 10 层神经网络?还是有一些纸笔或白板思考的直觉,比如因为你刚刚把一个 10 层大型神经网络与大脑联系起来?你刚才提到了大脑。那么在你对神经网络的直觉中,人脑是否起到了构建直觉的作用?

Was your intuition most of your intuition from empirical results by you and by others, so like just actually demonstrating that a piece of program can train a 10-layer neural network? Or was there some pen and paper or marker and whiteboard thinking intuition, like because you just connected a 10-layer large neural network to the brain? So you just mentioned the brain. So in your intuition about neural networks, does the human brain come into play as an intuition builder?

Ilya

当然。我的意思是,你知道,在人工神经网络和大脑之间的类比必须精确。但毫无疑问,大脑是深度学习研究人员直觉和灵感的巨大来源,从 60 年代的罗森布拉特开始。比如,如果你看看神经网络的整体概念,它直接受到大脑的启发。像麦卡洛克和皮茨这样的人说,嘿,大脑里有这些神经元,嘿,我们最近了解了计算机和自动机。我们能不能用计算机和自动机的一些想法来设计某种计算对象,它要简单、可计算,并且有点像大脑?然后他们发明了神经元。所以当时他们受到了启发。然后有福岛邦彦的卷积神经网络,后来还有杨立昆,他说,嘿,如果你限制神经网络的感受野,它会特别适合图像,结果证明这是真的。所以只有极少数例子表明与大脑的类比是成功的。而我想,嗯,如果训练得足够努力,人工神经元可能和大脑没有太大区别。

Definitely. I mean, you know, you've got to be precise with these analogies between artificial neural networks and the brain. But there is no question that the brain is a huge source of intuition and inspiration for deep learning researchers, all the way from Rosenblatt in the 60s. Like, if you look at the whole idea of a neural network, it is directly inspired by the brain. You had people like McCulloch and Pitts who were saying, hey, you've got these neurons in the brain, and hey, we recently learned about the computer and automata. Can we use some ideas from the computer and automata to design some kind of computational object that's going to be simple, computational, and kind of like the brain? And they invented the neuron. So they were inspired by it back then. Then you had the convolutional neural network from Fukushima, and then later Yann LeCun, who said, hey, if you limit the receptive fields of a neural network, it's going to be especially suitable for images, as it turned out to be true. So there were a very small number of examples where analogies to the brain were successful. And I thought, well, probably an artificial neuron is not that different from the brain if it's trained hard enough.

大脑与人工神经网络差异 Differences between brain and artificial neural networks

Host

我们就假设是这样,然后继续。现在我们处于深度学习非常成功的时代,所以让我们少眯着眼,睁大眼睛说,人脑和人工神经网络之间有什么有趣的差异?我知道你可能不是专家,既不是神经科学家也不是生物学家,但粗略来说,未来一二十年里,人脑和人工神经网络之间有什么让你感兴趣的差异?

Let's just assume it is and roll with it. So now we're at a time where deep learning is very successful, so let us squint less and say, let's open our eyes and say, what is an interesting difference between the human brain? Now I know you're probably not an expert, neither a neuroscientist nor a biologist, but loosely speaking, what's the difference between the human brain and artificial neural networks that's interesting to you for the next decade or two?

Ilya

这是个好问题。大脑中的神经元和人工神经网络之间有什么有趣的差异?我觉得今天,人工神经网络——我们都同意在某些方面人脑远远胜过我们的模型,但我也认为人工神经网络在某些方面相比大脑有许多非常重要的优势。看优势和劣势是找出重要差异的好方法。所以大脑使用脉冲,这可能重要也可能不重要。

That's a good question. What is an interesting difference between the neurons in the brain and our artificial neural networks? I feel like today, artificial neural networks—we all agree that there are certain dimensions in which the human brain vastly outperforms our models, but I also think that there are some ways in which artificial neural networks have a number of very important advantages over the brain. Looking at the advantages versus disadvantages is a good way to figure out what is the important difference. So the brain uses spikes, which may or may not be important.

Host

这真是个有趣的问题。你觉得它重要吗?这是人工神经网络和大脑之间一个大的架构差异。很难说,但我的先验概率不高,我可以解释原因。有些人对脉冲神经网络感兴趣,他们基本上发现需要把非脉冲神经网络模拟成脉冲形式,这样才能让它们工作。如果你不把非脉冲神经网络模拟成脉冲,它就行不通,因为问题是:为什么它应该行得通?这联系到反向传播和深度学习的问题。你有一个巨大的神经网络——它为什么能工作?学习规则为什么能工作?这不是一个不言自明的问题,尤其是如果你刚进入这个领域,读了早期论文,你会说:‘嘿,人们说让我们构建神经网络。这是个好主意,因为大脑是神经网络,所以构建神经网络会有用。现在让我们弄清楚如何训练它们。应该有可能正确训练它们,但怎么做?’所以核心思想是代价函数。这就是核心思想。代价函数是一种根据某种度量来衡量系统性能的方法。

That's a really interesting question. Do you think it's important or not? That's one big architectural difference between artificial neural networks and the brain. It's hard to tell, but my prior is not very high, and I can say why. There are people who are interested in spiking neural networks, and basically what they've figured out is that they need to simulate the non-spiking neural networks in spikes, and that's how they're going to make them work. If you don't simulate the non-spiking neural networks in spikes, it's not going to work, because the question is: why should it work? And that connects to questions around backpropagation and questions around deep learning. You've got this giant neural network—why should it work at all? Why should the learning rule work at all? It's not a self-evident question, especially if you were just starting in the field and you read the very early papers. You can say, 'Hey, people are saying let's build neural networks. That's a great idea because the brain is a neural network, so it would be useful to build neural networks. Now let's figure out how to train them. It should be possible to train them properly, but how?' And so the big idea is the cost function. That's the big idea. The cost function is a way of measuring the performance of the system according to some measure.

Host

顺便说一句,那是个大想法。实际上,让我想想:想到这个想法难吗?这个想法有多大——即存在一个单一的代价函数?让我停一下。监督学习是一个难以想到的概念吗?我不知道。所有概念事后看来都很简单。

By the way, that is a big idea. Actually, let me think: is that a difficult idea to arrive at, and how big of an idea is that—that there's a single cost function? Let me pause. Is supervised learning a difficult concept to come to? I don't know. All concepts are very easy in retrospect.

Ilya

是的,现在看起来确实很简单。但我问这个问题的原因,我们也会讨论,是因为还有其他东西。有没有一些东西不一定有代价函数?也许有很多代价函数,或者有动态代价函数,或者完全不同的架构?因为我们必须这样思考才能得到新东西,对吧?所以唯一没有明确代价函数的好例子是生成对抗网络(GANs)。同样,你有一个博弈。所以不是考虑一个你想优化的代价函数——你知道你有一个算法,梯度下降,它会优化代价函数,然后你可以根据它优化什么来推理系统的行为——对于 GANs,你说:‘我有一个博弈,我会根据博弈的均衡来推理系统的行为。’但这一切都是为了提出这些数学对象,帮助我们推理系统的行为,对吧?

Yes, that's what it seems—trivial now. But the reason I asked that, and we'll talk about it, is because there are other things. Are there things that don't necessarily have a cost function? Maybe have many cost functions, or maybe have dynamic cost functions, or maybe a totally different kind of architecture? Because we have to think like that in order to arrive at something new, right? So the only good example of things which don't have clear cost functions is GANs. Again, you have a game. So instead of thinking of a cost function where you want to optimize—where you know that you have an algorithm, gradient descent, which will optimize the cost function, and then you can reason about the behavior of your system in terms of what it optimizes—with GANs, you say: 'I have a game, and I'll reason about the behavior of the system in terms of the equilibrium of the game.' But it's all about coming up with these mathematical objects that help us reason about the behavior of our system, right?

Host

这真的很有趣。GANs 是唯一的例子。这有点像代价函数是从比较中涌现出来的。我不知道它是否有代价函数。我不知道谈论 GANs 的代价函数是否有意义。这有点像生物进化的代价函数或经济的代价函数。你可以谈论它会趋向的区域,但我不认为代价函数的类比是最有用的。所以如果进化并没有真正的代价函数——类似于我们数学概念中的代价函数——那么你认为深度学习中的代价函数在阻碍我们吗?

That's really interesting. GANs are the only one. It's kind of like the cost function is emergent from the comparison. I don't know if it has a cost function. I don't know if it's meaningful to talk about the cost function of GANs. It's kind of like the cost function of biological evolution or the cost function of the economy. You can talk about regions to which it will go towards, but I don't think the cost function analogy is the most useful. So if evolution doesn't really have a cost function—like a cost function based on something akin to our mathematical conception of a cost function—then do you think cost functions in deep learning are holding us back?

Ilya

是的,你刚才提到代价函数是一个很好的第一个深刻想法。你认为这是个好主意吗?你认为我们会超越这个想法吗?自我对弈在强化学习系统中开始触及这一点。没错,自我对弈以及探索相关的想法,比如试图采取让预测器惊讶的行动。我是代价函数的忠实粉丝。我认为代价函数很棒,它们对我们非常有用,而且我认为只要能用代价函数做事,我们就应该用。而且,也许我们有可能提出另一种深刻的看待事物的方式,其中代价函数不那么核心,但我不知道。我认为代价函数——我是说,我不会打赌代价函数会被淘汰。

Yeah, so you just kind of mentioned that cost function is a nice first profound idea. Do you think that's a good idea? Do you think it's an idea we will go past? Self-play starts to touch on that a little bit in reinforcement learning systems. That's right, self-play and also ideas around exploration where you're trying to take actions that surprise a predictor. I'm a big fan of cost functions. I think cost functions are great and they serve us really well, and I think that whenever we can do things with cost functions, we should. And you know, maybe there is a chance that we will come up with some yet another profound way of looking at things that will involve cost functions in a less central way, but I don't know. I think cost functions are—I mean, I would not bet against cost functions.

Host

关于大脑,你还能想到其他可能不同且有趣的东西,供我们在设计人工神经网络时考虑吗?我们谈了一点脉冲。我的意思是,有一件事可能有用:我认为神经科学家已经弄清楚了大脑学习规则的一些东西,或者说我在说脉冲时间依赖可塑性(STDP)。

Is there anything else about the brain that pops into your mind that might be different and interesting for us to consider in designing artificial neural networks? So we talked about spiking a little bit. I mean, one thing which may potentially be useful: I think neuroscientists have figured out something about the learning rule of the brain, or I'm talking about spike-timing-dependent plasticity.

Ilya

等等,抱歉,脉冲时间依赖可塑性?是的,那是什么?STDP。这是一种特定的学习规则,利用脉冲时间来决定如何更新突触。所以有点像:如果突触在神经元放电之前向神经元放电,那么它会增强突触;如果突触在神经元放电之后不久向神经元放电,那么它会削弱突触。大致如此。我 90%确定是对的,所以如果我说错了什么,别太生气。但你说的时候听起来很聪明。但时间——那是缺失的一点。时间动态没有被捕捉。我认为这就像大脑的一个基本属性:信号的时间。

Wait, sorry, spike-timing-dependent plasticity? Yeah, what's that? STDP. It's a particular learning rule that uses spike timing to figure out how to update the synapses. So it's kind of like: if the synapse fires into the neuron before the neuron fires, then it strengthens the synapse; and if the synapse fires into the neuron shortly after the neuron fires, then it weakens the synapse. Something along this line. I'm 90% sure it's right, so if I said something wrong here, don't get too angry. But you sounded brilliant while saying it. But the timing—that's one thing that's missing. The temporal dynamics is not captured. I think that's like a fundamental property of the brain: the timing of the signals.

Host

嗯,循环神经网络——但你认为那是一个非常粗略的简化版本。我想循环神经网络有一个时钟。大脑似乎是它的连续版本,是允许所有可能时间的泛化,在这些时间中包含了一些信息。你认为循环神经网络中的循环能捕捉到与时间相同的现象吗?这似乎是……

Well, recurrent neural networks—but you think of that as a very crude simplified version. There's a clock, I guess, for recurrent neural networks. It seems like the brain is the continuous version of that, the generalization where all possible timings are possible, and within those timings, this contains some information. You think recurrent neural networks, the recurrence in recurrent neural networks, can capture the same kind of phenomena as the timing? That seems to be...

递归与神经网络 Recurrence and Neural Networks

Host

最近自然语言处理和语言建模的许多突破都来自不强调循环的 Transformer。你认为循环会卷土重来吗?

So much of the recent breakthroughs in natural language processing and language modeling have been with transformers that don't emphasize recurrence. Do you think recurrence will make a comeback?

Ilya

嗯,某种形式的循环,我认为很有可能。通常意义上的用于处理序列的循环神经网络,我觉得也有可能。对你来说什么是循环神经网络?一般来说,循环神经网络是什么?你有一个神经网络,它维护一个高维隐藏状态,当观测到达时,它通过某种连接方式更新其高维隐藏状态。

Well, some kind of recurrence, I think very likely. Recurrent neural networks for processing sequences as they're typically thought of, I think it's also possible. What is to you a recurrent neural network? Generally speaking, what is a recurrent neural network? You have a neural network which maintains a high-dimensional hidden state, and when an observation arrives, it updates its high-dimensional hidden state through its connections in some way.

Host

你更广义地理解它,比如将知识库视为隐藏状态的专家系统?还是更受限制的形式,比如带有门控单元的 LSTM?

Do you think of it more generally, like expert systems with a knowledge base that is a hidden state? Or is it the more constrained form with gating units like LSTMs?

Ilya

隐藏状态在技术上就是你描述的那样,LSTM 或 RNN 内部的隐藏状态。但应该包含什么?如果你想用专家系统类比,可以说知识存储在连接中,短期处理在隐藏状态中完成。是的,可以这么说。

The hidden state is technically what you described, the hidden state that goes inside the LSTM or the RNN. But then what should be contained? If you want to make the expert system analogy, you could say that the knowledge is stored in the connections and the short-term processing is done in the hidden state. Yes, you could say that.

Host

你认为在神经网络内部构建大规模知识库有未来吗?

Do you think there's a future for building large-scale knowledge bases within neural networks?

Ilya

当然。

Definitely.

深度学习成功的关键思想 Key Ideas Behind Deep Learning Success

Host

让我再退一步看。神经网络已经存在了几十年。你认为哪些关键想法导致了它们在过去的 10 年里取得成功,从 ImageNet 开始?

Let me zoom back out. Neural networks have been around for many decades. What do you think were the key ideas that led to their success in the past 10 years, starting with ImageNet?

Ilya

深度学习在成功之前的关键事实是它被低估了。从事机器学习的人根本不认为神经网络能做多少事。人们不相信大型神经网络可以被训练。有很多争论,因为没有硬事实——没有真正困难的基准,如果你做得很好,你可以说‘看,这是我的系统’。那时这个领域变得更像工程领域。想法都在那里。缺少的是大量的监督数据和大量的算力。一旦有了这些,还需要第三样东西:信念——相信如果你把已有的正确东西与大量数据和大量算力混合,它就会起作用。所以缺失的部分是数据、算力(以 GPU 的形式出现),以及意识到需要将它们结合起来的信念。

The key fact about deep learning before it started to be successful is that it was underestimated. People who worked in machine learning simply didn't think that neural networks could do much. People didn't believe that large neural networks could be trained. There was a lot of debate because there were no hard facts—no benchmarks that were truly hard, where if you do really well, you can say 'look here is my system.' That's when the field becomes more of an engineering field. The ideas were all there. The thing that was missing was a lot of supervised data and a lot of compute. Once you have that, there is a third thing needed: conviction—the conviction that if you take the right stuff which already exists and mix it with a lot of data and a lot of compute, it will work. So the missing piece was the data, the compute (which showed up in terms of GPUs), and the conviction to realize that you need to mix them together.

Host

所以算力和监督数据的存在让经验证据说服了计算机科学界的大多数人。有一个关键时刻,Jitendra Malik 和 Alyosha Efros 持怀疑态度,而 Geoffrey Hinton 则相反。ImageNet 就是那个时刻。

So the presence of compute and supervised data allowed empirical evidence to convince the majority of the computer science community. There was a key moment with Jitendra Malik and Alyosha Efros, who were skeptical, and Geoffrey Hinton, who was the opposite. ImageNet served as that moment.

Ilya

没错。他们代表了计算机视觉界的支柱。大师们聚在一起,突然之间就发生了转变。光有想法和算力是不够的,还需要说服存在的怀疑论。有趣的是,人们几十年都不相信。但实际上,情况很混乱,因为神经网络确实在任何事情上都不起作用,而且几乎在任何事情上都不是最好的方法。说这些东西没有吸引力是理性的。这就是为什么你需要非常困难的任务来产生无可辩驳的证据。这就是我们取得进步的方式,也是今天这个领域取得进步的原因——因为我们有这些代表真正进步的硬基准,并且我们能够避免无休止的争论。

That's right. They represented the big pillars of the computer vision community. The wizards got together, and all of a sudden there was a shift. It's not enough for the ideas and compute to be there; it's for it to convince the cynicism that existed. It's interesting that people just didn't believe for a couple of decades. But in reality, things were confusing because neural networks really did not work on anything, and they were not the best method on pretty much anything. It was rational to say this stuff doesn't have any traction. That's why you need very hard tasks that produce undeniable evidence. That's how we make progress, and that's why the field is making progress today—because we have these hard benchmarks that represent true progress, and we are able to avoid endless debate.

机器学习的统一 Unity of Machine Learning

Host

你对视觉、语言、强化学习以及深度学习的基础科学都有贡献。作为学习问题,视觉、语言和强化学习之间有什么区别?有什么共同点?你认为它们是相互关联的,还是根本不同的领域,需要不同的方法?

You've contributed to vision, language, reinforcement learning, and the fundamental science of deep learning. What is the difference between vision, language, and reinforcement learning as learning problems, and what are the commonalities? Do you see them as interconnected, or fundamentally different domains requiring different approaches?

Ilya

机器学习是一个具有巨大统一性的领域——思想的重叠,原理的重叠。事实上,只有一两个或三个非常简单的原理,它们几乎以相同的方式适用于不同的模态和不同的问题。这就是为什么今天,当有人写一篇关于改进视觉中深度学习优化的论文时,它也会改进 NLP 应用和强化学习应用。我会说计算机视觉和 NLP 今天非常相似。它们的区别在于输入模态不同,但底层原理是相同的。

Machine learning is a field with a huge amount of unity—overlap of ideas, overlap of principles. In fact, there are only one or two or three principles that are very simple, and they apply in almost the same way to different modalities and different problems. That's why today, when someone writes a paper on improving optimization of deep learning in vision, it improves NLP applications and reinforcement learning applications as well. I would say that computer vision and NLP are very similar to each other today. They differ in that they have different input modalities, but the underlying principles are the same.

架构的统一 Unification of architectures

Ilya

架构上略有不同:我们在自然语言处理中使用 Transformer,在视觉中使用卷积神经网络。但有一天这种情况也可能会改变,所有东西都统一到单一架构下。如果你回到几年前,自然语言处理中有大量的架构,每个不同的小问题都有自己的架构。今天,所有不同任务都只用一种 Transformer。如果再往前推,碎片化更严重。AI 中的每个小问题都有自己的子专业化和一套技能,人们知道如何手工设计特征。现在这一切都被深度学习吸收了。我们有了这种统一。所以我期望视觉也能与自然语言统一。或者更准确地说,我不该说期望,我认为这是可能的。我不想太确定,因为我认为卷积神经网络在计算上非常高效。

Slightly different architectures: we use transformers in NLP and convolutional neural networks in vision. But it's also possible that one day this will change and everything will be unified with a single architecture. If you go back a few years ago in natural language processing, there were a huge number of architectures for every different tiny problem. Today, it's just one transformer for all those different tasks. If you go back in time even more, you had even more fragmentation. Every little problem in AI had its own sub-specialization and set of skills, people who would know how to engineer the features. Now it's all been subsumed by deep learning. We have this unification. So I expect vision to become unified with natural language as well. Or rather, I shouldn't say expect; I think it's possible. I don't want to be too sure because I think the convolutional neural net is very computationally efficient.

强化学习统一语言与视觉 Unification of RL and supervised learning

Ilya

强化学习不同;强化学习确实需要稍微不同的技术,因为你确实需要采取行动,确实需要处理探索问题,方差要高得多。但我认为即使在那里也有很多统一性。我预计在某个时候,强化学习和监督学习之间会有更广泛的统一,强化学习会做出决策让监督学习进行得更好。我想象它会是一个大黑箱:你把东西扔进去,它自己就知道怎么处理你扔进去的任何东西。

RL is different; RL does require slightly different techniques because you really do need to take action, you really do need to do something about exploration, your variance is much higher. But I think there is a lot of unity even there. I would expect that at some point there will be some broader unification between RL and supervised learning, where somehow the RL will be making decisions to make the supervised learning go better. It will be, I imagine, one big black box: you just shovel things into it, and it just figures out what to do with whatever you shovel it.

学习行动的独特之处 RL integrates language and vision

Host

强化学习几乎结合了语言和视觉的某些方面。有长期记忆的元素需要利用,也有非常丰富的感官空间的元素。所以它似乎是两者的结合。

Reinforcement learning has some aspects of language and vision combined almost. There are elements of long-term memory that you should be utilizing, and elements of a really rich sensory space. So it seems like it's the union of the two or something like that.

Ilya

我会说得稍微不同。我会说强化学习两者都不是,但它自然地与这两者接口和集成。

I'd say something slightly differently. I'd say that reinforcement learning is neither, but it naturally interfaces and integrates with the two of them.

语言与视觉哪个更难 What is unique about learning to act

Host

你认为行动有根本不同吗?策略、学习良好行动有什么独特之处?一个例子是,当你学习行动时,你基本上处在一个非平稳的世界中,因为随着你的行动改变,你看到的东西开始变化。你以不同的方式体验世界。而对于更传统的静态问题,你至少有一个分布,你只是将模型应用于该分布,情况并非如此。你认为这是一个根本不同的问题,还是只是理解问题的一个更困难的泛化?

Do you think action is fundamentally different? What is unique about policy, learning to act well? One example is that when you learn to act, you are fundamentally in a non-stationary world because as your actions change, the things you see start changing. You experience the world in a different way. This is not the case for the more traditional static problem where you have at least some distribution and you just apply a model to that distribution. Do you think it's a fundamentally different problem, or is it just a more difficult generalization of the problem of understanding?

Ilya

这几乎是一个定义问题。肯定有大量的共同点。你计算梯度,你在两种情况下都试图近似梯度。在强化学习的情况下,你有一些工具来减少梯度的方差。你这样做。有很多共同点:在两种情况下使用相同的神经网络,你计算梯度,你应用它。所以肯定有很多共同点。但也有一些小的差异,并非完全不重要。这真的只是你的观点问题,你的参考系是什么,你在看这些问题时想放大或缩小多少。

It's a question of definitions almost. There is a huge amount of commonality for sure. You take gradients, you try to approximate gradients in both cases. In the case of reinforcement learning, you have some tools to reduce the variance of the gradients. You do that. There's lots of commonality: use the same neural net in both cases, you compute the gradient, you apply it. So there's lots in common for sure. But there are some small differences which are not completely insignificant. It's really just a matter of your point of view, what frame of reference, how much do you want to zoom in or out as you look at these problems.

AI 中的惊喜与幽默 Which problem is harder: language or vision

Host

你认为哪个问题更难?像诺姆·乔姆斯基这样的人认为语言是一切的基础,所以它是一切的基础。你认为语言理解比视觉场景理解更难,还是相反?

Which problem do you think is harder? People like Noam Chomsky believe that language is fundamental to everything, so it underlies everything. Do you think language understanding is harder than visual scene understanding, or vice versa?

Ilya

我认为问一个问题是否困难有点不对。我认为这个问题有点不对,我想解释为什么。一个问题困难是什么意思?无趣的愚蠢答案是有一个基准,基准上有人类水平的表现,还有达到人类水平所需的努力。所以从我们在一个非常好的基准上达到人类水平还需要多少努力的角度来看……

I think that asking if a problem is hard is slightly wrong. I think the question is a little bit wrong, and I want to explain why. What does it mean for a problem to be hard? The non-interesting dumb answer is there's a benchmark and there's human level performance on that benchmark, and there's the effort required to reach the human level. So from the perspective of how much until we get to human level on a very good benchmark...

Host

是的,比如一些……我明白你的意思。

Yeah, like some... I understand what you mean by that.

Ilya

我想说的是,这在很大程度上取决于……一旦你解决了一个问题,它就不再困难了,这总是正确的。所以某件事是否困难取决于我们今天的工具能做什么。今天,真正的人类水平语言理解和视觉感知是困难的,因为没有办法在接下来的三个月内完全解决这个问题。所以我同意这个说法。除此之外,我的猜测和你的一样好。我不知道。

What I was going to say is that a lot of it depends on... once you solve a problem, it stops being hard, and that's always true. So whether something is hard or not depends on what our tools can do today. Today, true human level language understanding and visual perception are hard in the sense that there is no way of solving the problem completely in the next three months. So I agree with that statement. Beyond that, my guess would be as good as yours. I don't know.

Host

哦,好吧。所以你对语言理解有多难没有基本的直觉?

Oh, okay. So you don't have a fundamental intuition about how hard language understanding is?

Ilya

我想我知道……我改变主意了。假设语言可能更难。我的意思是,这取决于你如何定义它。如果你指的是绝对顶尖的 100%语言理解,我选语言。但如果我给你看一张有字母的纸,那是视觉吗?你明白我的意思吗?你有一个视觉系统。你说它是最好的类人视觉系统。我给你看一本书,我给你看字母。它能理解这些字母如何形成单词、句子和意义吗?这是视觉问题的一部分吗?视觉在哪里结束,语言在哪里开始?

I think I know... I changed my mind. Let's say language is probably going to be harder. I mean, it depends on how you define it. If you mean absolute top-notch 100% language understanding, I'll go with language. But then if I show you a piece of paper with letters on it, is that vision? You see what I mean? You have a vision system. You say it's the best human level vision system. I show you a book and I show you letters. Will it understand how these letters form into words and sentences and meaning? Is this part of the vision problem? Where does vision end and language begin?

Host

是的,所以乔姆斯基会说它从语言开始。所以视觉只是某种结构和思想基本层次的一个小例子,这些思想已经以某种方式在我们的大脑中表示,通过语言表示。但视觉在哪里停止,语言在哪里开始?这是一个非常有趣的问题。一种可能性是,如果不使用同一种系统,就不可能对图像或语言实现真正深刻的理解。所以你会免费得到另一个。

Yeah, so Chomsky would say it starts at language. So vision is just a little example of the kind of structure and fundamental hierarchy of ideas that's already represented in our brain somehow, that's represented through language. But where does vision stop and language begin? That's a really interesting question. One possibility is that it's impossible to achieve really deep understanding in either images or language without basically using the same kind of system. So you're going to get the other for free.

Ilya

我认为很有可能,如果我们能获得一个,我们的机器学习可能足够好,我们也能获得另一个。但这不是 100%。我不是 100%确定。而且我认为这在很大程度上确实取决于你对完美视觉的定义,因为实际上,阅读是视觉,但它应该算吗?

I think it's pretty likely that yes, if we can get one, our machine learning is probably that good that we can get the other. But it's not 100%. I'm not 100% sure. And also I think a lot of it really does depend on your definitions of perfect vision, because really, reading is vision, but should it count?

Host

对我来说,我的定义是,如果一个系统看了一张图像,然后系统看了一段文本,然后告诉我一些关于那件事的东西,并且我真的很印象深刻,那是相对的。你会印象深刻半小时,然后你会说,‘嗯,我的意思是,所有系统都这样做。’但问题是:它们不这样做。是的,但我和人类在一起没有这种情况。人类持续让我印象深刻。

To me, my definition is if a system looked at an image and then the system looked at a piece of text and then told me something about that, and I was really impressed, that's relative. You'll be impressed for half an hour, and then you're gonna say, 'Well, I mean, all the systems do that.' But here's the thing: they don't do that. Yeah, but I don't have that with humans. Humans continue to impress me.

Ilya

真的吗?嗯,那些……好吧,我是一夫一妻制的粉丝,所以我喜欢和某人结婚,和他们在一起几十年的想法。所以我相信,是的,有可能有人持续给你带来愉快、有趣、机智、新的想法。朋友,是的。我想是的。他们持续让你惊喜。

Is that true? Well, the ones... Okay, I'm a fan of monogamy, so I like the idea of marrying somebody, being with them for several decades. So I believe in the fact that yes, it's possible to have somebody continuously giving you pleasurable, interesting, witty, new ideas. Friends, yeah. I think so. They continue to surprise you.

深度学习中最美的思想 Surprise and humor in AI

Host

这种惊喜,你知道,那种随机性的注入似乎是持续灵感的良好来源,比如机智、幽默。我觉得那会是一个非常主观的测试,但如果你有足够多的人在房间里……是的,我明白你的意思。我感觉我误解了你说的“打动你”的意思。我以为你指的是用它的智能打动你,比如它多好地理解一张图片。我以为你的意思是,我要给它看一张非常复杂的图片,它能正确理解,然后你会说‘哇,真酷’。2020 年 1 月的系统还做不到这一点。

The surprise, it's, you know, that injection of randomness seems to be a nice source of continued inspiration, like the wit, the humor. I think that would be a very subjective test, but I think if you have enough humans in the room... Yeah, I understand what you mean. I feel like I misunderstood what you meant by impressing you. I thought you meant to impress you with its intelligence, with how well it understands an image. I thought you meant something like I'm going to show it a really complicated image and it's going to get it right, and you're going to say, 'Wow, that's really cool.' Systems of January 2020 have not been doing that.

Ilya

不,我认为这归根结底是人们点赞互联网内容的原因——它让他们发笑,所以是幽默、机智或洞见。我相信我们也会得到这些。

No, I think it all boils down to the reason people click like on stuff on the internet, which is it makes them laugh, so it's humor or wit, or insight. I'm sure we'll get that as well.

深度学习为何有效的直觉 Most beautiful idea in deep learning

Host

请原谅我这个浪漫化的问题,但回顾过去,你遇到过的最美丽或最令人惊讶的深度学习或 AI 想法是什么?

So forgive the romanticized question, but looking back, what is the most beautiful or surprising idea in deep learning or AI in general you've come across?

Ilya

我认为深度学习最美妙的地方在于它真的管用。我是认真的,因为你有这些想法:你有一个小小的神经网络,你有反向传播算法,然后你有一些理论说,这有点像大脑,所以如果你把它做大,如果你把神经网络做大并在大量数据上训练它,那么它就会做大脑所做的同样功能。结果证明这是真的。这太疯狂了。现在我们只是训练这些神经网络,让它们更大,它们就不断变得更好。我觉得整个神经网络 AI 能工作简直难以置信。

I think the most beautiful thing about deep learning is that it actually works. And I mean it, because you got these ideas: you got the little neural network, you got the backpropagation algorithm, and then you got some theories as to, you know, this is kind of like the brain, so maybe if you make it large, if you make the neural network large and you train it on a lot of data, then it will do the same function the brain does. And it turns out to be true. That's crazy. And now we just train these neural networks and you make them larger and they keep getting better. And I find it unbelievable that this whole AI stuff with neural networks works.

神经网络的未发现特性 Intuitions on why deep learning works

Host

你是否已经建立起一种直觉,知道为什么?有没有一些零碎的直觉或洞见,解释为什么这一切能行?

Have you built up an intuition of why? Are there little bits and pieces of intuitions, of insights of why this whole thing works?

Ilya

我的意思是,肯定有一些。我们知道优化,我们现在有很好的——我们有大量的经验证据,大量的经验理由相信优化应该能解决我们关心的几乎所有问题。

I mean, some definitely. We know that optimization, we now have good—we've had lots of empirical, huge amounts of empirical reasons to believe that optimization should work on almost all problems we care about.

Host

你有没有什么洞见——你刚才说经验证据是你主要的——经验证据说服了你。这就像进化论:它是经验性的,它告诉你这个进化过程似乎是设计能在环境中生存的生物体的好方法,但它并没有真正让你了解整个系统内部是如何运作的。

Did you have insights of what—so you just said empirical evidence is most of your sort of—empirical evidence kind of convinces you. It's like evolution: it's empirical, it shows you that this evolutionary process seems to be a good way to design organisms that survive in their environment, but it doesn't really get you to the insides of how the whole thing works.

Ilya

我认为这是一个很好的类比:物理学。你知道,你说‘嘿,让我们做一些物理计算,提出一些新的物理理论并做出一些预测’,但你必须进行实验。你必须做实验,这很重要。所以这里也差不多,只不过有时实验先于理论。但情况仍然是:你有一些数据,你提出一些预测,你说‘是的,让我们做一个大的神经网络,训练它,它会比之前的任何东西都好得多,而且随着你让它变大,它实际上会继续变得更好。’结果证明这是真的。当一个理论像这样被验证时,真是太神奇了。这不是一个数学理论,它更像是一个生物学理论。所以我认为深度学习和生物学之间有一些不错的类比。我会说它像是生物学和物理学的几何平均。这就是深度学习:生物学和物理学的几何平均。

I think it's a good analogy: physics. You know how you say, 'Hey, let's do some physics calculation and come up with some new physics theory and make some prediction,' but then you gotta run the experiment. You gotta run the experiment, it's important. So it's a bit the same here, except that maybe sometimes the experiment came before the theory. But it still is the case: you have some data and you come up with some prediction, you say, 'Yeah, let's make a big neural network, let's train it, and it's going to work much better than anything before it, and it will in fact continue to get better as you make it larger.' And it turns out to be true. That's amazing when a theory is validated like this. It's not a mathematical theory, it's more of a biological theory almost. So I think there are not terrible analogies between deep learning and biology. I would say it's like the geometric mean of biology and physics. That's deep learning: the geometric mean of biology and physics.

Host

我想我需要几个小时才能理解这一点,因为要找到生物学所代表的几何平均——嗯,生物学:在生物学中,事物非常复杂,理论很难有好的预测性。而在物理学中,理论又太好了。在物理学中,人们做出这些超精确的理论,做出惊人的预测。而机器学习机制介于两者之间,有点中间。但如果机器学习能帮助我们发现两者的统一,而不是某种中间状态,那就太好了。但你说得对,你是在试图同时兼顾两者。

I think I'm going to need a few hours to wrap my head around that, because just to find the geometric mean of what biology represents—well, biology: in biology things are really complicated, theories are really hard to have good predictive theory. And in physics, the theories are too good. In physics, people make these super precise theories which make amazing predictions. And machine learning mechanics is in between, kind of in between. But it'd be nice if machine learning somehow helped us discover the unification of the two, as opposed to some of the in-between. But you're right, you're kind of trying to juggle both.

取得进展的困难 Undiscovered properties of neural networks

Host

那么你认为网络中还有美丽而神秘的属性尚未被发现吗?

So do you think there are still beautiful and mysterious properties in your networks that are yet to be discovered?

Ilya

当然。我认为我们仍然严重低估了深度学习。

Definitely. I think that we are still massively underestimating deep learning.

Host

你认为它会是什么样子?如果我知道,我早就做了。

What do you think it will look like? Like, if I knew, I would have done it.

Ilya

是的,如果你看看过去 10 年的所有进展,我会说大部分——有少数情况出现了一些感觉像是真正新想法的东西,但总的来说,每年我们都认为‘好吧,深度学习就到这了。’不,它实际上走得更远。然后第二年,‘好了,现在是深度学习的顶峰,我们真的结束了。’不,它走得更远。它每年都在继续前进。这意味着我们一直在低估,我们一直不理解它,总是有令人惊讶的属性。

Yeah, so if you look at all the progress from the past 10 years, I would say most of it—there have been a few cases where some things that felt like really new ideas showed up, but by and large, it was every year we thought, 'Okay, deep learning goes this far.' Nope, it actually goes further. And then the next year, 'Okay, now this is peak deep learning, we are really done.' Nope, goes further. It just keeps going further each year. So that means that we keep underestimating, we keep not understanding it, as surprising properties all the time.

未来突破与算力需求 Difficulty of making progress

Host

你认为取得进展越来越难了吗?

Do you think it's getting harder and harder to make progress?

Ilya

这取决于我们指的是什么。我认为这个领域将在相当长一段时间内继续取得非常稳健的进展。我认为对于个别研究人员,尤其是做研究的人来说,可能会更难,因为现在研究人员数量非常多。我认为如果你有大量算力,那么你可以做出很多非常有趣的发现,但随后你必须应对管理一个巨大的算力集群、尝试实验的挑战。所以有点难。

It depends on what we mean. I think the field will continue to make very robust progress for quite a while. I think for individual researchers, especially people who are doing research, it can be harder because there is a very large number of researchers right now. I think that if you have a lot of compute, then you can make a lot of very interesting discoveries, but then you have to deal with the challenge of managing a huge compute cluster, trying to experiment. So it's a little bit harder.

突破与算力需求 Future breakthroughs and compute requirements

Host

所以我在问这些没人知道答案的问题,但你是我认识的最聪明的人之一,所以我会继续问。那么让我们想象未来 30 年深度学习的所有突破。你认为这些突破中的大多数可以由一个人用一台计算机完成吗?在突破的空间里,你认为算力和大规模努力会是必要的吗?

So I'm asking all these questions that nobody knows the answer to, but you're one of the smartest people I know, so I'm going to keep asking. So let's imagine all the breakthroughs that happen in the next 30 years in deep learning. Do you think most of those breakthroughs can be done by one person with one computer? Sort of in the space of breakthroughs, do you think compute will be necessary and large efforts will be necessary?

Ilya

我的意思是,我不能确定。你说一台计算机,你指的是多大?你很聪明。我是指一个 GPU?我明白了。我认为这不太可能。我认为这非常不可能。我认为深度学习的栈开始变得相当深。如果你看它,从想法、构建数据集的系统、分布式编程、构建实际集群、GPU 编程,到把它们整合在一起。所以现在这个栈变得非常深,我认为一个人要在栈的每一层都成为世界级水平是相当困难的。

I mean, I can't be sure. When you say one computer, you mean how large? You're clever. I mean one GPU? I see. I think it's pretty unlikely. I think it's pretty unlikely. I think that the stack of deep learning is starting to be quite deep. If you look at it, you've got all the way from the ideas, the systems to build the datasets, the distributed programming, the building the actual cluster, the GPU programming, putting it all together. So now the stack is getting really deep, and I think it becomes quite hard for a single person to become world-class in every single layer of the stack.

Host

那像 Vladimir Vapnik 真正坚持的,用 MNIST 并从很少的例子中学习,从而能够更高效地学习。你认为在那个领域会有可能不需要巨大算力的突破吗?

What about the—like Vladimir Vapnik really insists on taking MNIST and trying to learn from very few examples, so being able to learn more efficiently. Do you think there'll be breakthroughs in that space that may not need the huge compute?

Ilya

我认为这将是一个非常——我认为会有大量的……

I think it will be a very—I think there will be a large number of...

深度双重下降论文 Breakthroughs and Compute Requirements

Ilya

总的来说,有些突破并不需要巨大的算力。也许我应该澄清一下,我认为有些突破确实需要大量算力,而且构建真正能做事儿的系统也需要巨大算力。这很明显:如果你想做某件事,而这件事需要巨大的神经网络,那你就得搞一个巨大的神经网络。但我认为,还有很多非常重要的研究工作可以由小团队和个人来完成。

Breakthroughs in general that will not need a huge amount of compute. So maybe I should clarify that I think that some breakthroughs will require a lot of compute, and I think building systems which actually do things will require a huge amount of compute. That one is pretty obvious: if you want to do X, right, and X requires a huge neural net, you got to get a huge neural net. But I think there will be lots of room for very important work being done by small groups and individuals.

反向传播与替代方案 Deep Double Descent Paper

Host

你也许可以谈谈深度学习科学的话题。说说你最近发表的一篇论文:关于深度双重下降,即更大的模型和更多的数据反而有害。我觉得这篇论文很有意思。你能描述一下主要思想吗?

You may be on the topic of the science of deep learning. Talk about one of the recent papers that you released: that deep double descent where bigger models and more data hurt. I think it's a really interesting paper. Can you describe the main idea?

Ilya

当然。事情是这样的:多年来,少数研究人员注意到,当你把神经网络做得更大时,它反而工作得更好,这似乎与统计学观点相矛盾。然后有人做了分析,表明实际上存在一个双重下降的凸起。而我们做的工作是证明双重下降几乎发生在所有实际的深度学习系统中。

Definitely. So what happened is that over the years, some small number of researchers noticed that it is kind of weird that when you make the neural network larger, it works better, and it seems to go in contradiction with statistical ideas. And then some people made an analysis showing that actually you got this double descent bump. And what we've done was to show that double descent occurs for pretty much all practical deep learning systems.

Host

你能退一步解释一下吗?双重下降图的 x 轴和 y 轴是什么?

Can you step back? What's the x-axis and the y-axis of a double descent plot?

Ilya

好的。你可以做这样的事情:拿一个神经网络,在保持数据集固定的情况下,慢慢增加它的规模。如果你慢慢增加神经网络的规模,并且不做早停——这是一个非常重要的细节——那么当神经网络非常小时,你把它变大,性能会迅速提升。然后你继续让它变大,在某个点上性能会变差,最差的时候恰好是在它达到零训练误差、精确的零训练损失的时候。然后当你让它继续变大时,性能又开始变好。这有点反直觉,因为你期望深度学习现象是单调的。很难确定这意味着什么,但它在线性分类器的情况下也会发生。直觉基本上归结为:当你有一个大数据集和一个小的模型时,那么小的随机... 基本上什么是过拟合?过拟合就是你的模型对数据集中那些小的、随机的、不重要的东西非常敏感,在训练数据集中。没错。所以如果你有一个小模型和一个大数据集,可能会有一些随机的东西——你知道,一些训练样本随机出现在数据集中,而其他样本可能不在——但小模型对这种随机性不太敏感,因为它是相同的;当模型很大时,几乎不存在不确定性。所以,在最基本的层面上,对我来说最令人惊讶的是,神经网络在能够学到任何东西之前,并不会每次都很快地过拟合,尽管参数数量巨大。

Okay, great. So you can do things like take a neural network and start increasing its size slowly while keeping your data set fixed. So if you increase the size of the neural network slowly, and if you don't do early stopping — that's a pretty important detail — then when the neural network is really small, you make it larger, you get a very rapid increase in performance. Then you continue to make it large, and at some point performance will get worse, and it gets the worst exactly at the point at which it achieves zero training error, precisely zero training loss. And then as you make it large, it starts to get better again. It's kind of counter-intuitive because you'd expect deep learning phenomena to be monotonic. And it's hard to be sure what it means, but it also occurs in the case of linear classifiers. And the intuition basically boils down to the following: when you have a large data set and a small model, then small tiny random... so basically what is overfitting? Overfitting is when your model is somehow very sensitive to the small random unimportant stuff in your data set, in the training data set. Precisely. So if you have a small model and you have a big data set, and there may be some random thing — you know, some training cases are randomly in the data set and others may not be there — but the small model is kind of insensitive to this randomness because it's the same; there is pretty much no uncertainty about the model when it is large. So okay, so at the very basic level, to me it is the most surprising thing that neural networks don't overfit every time very quickly before ever being able to learn anything, given the huge number of parameters.

Host

所以这里有一种方式……

So here there is one way...

Ilya

那么让我试着给出一个可行的解释。假设你有一个巨大的神经网络,有大量的参数。现在我们假装一切都是线性的,虽然实际上不是,但我们就假装一下。那么存在一个大的子空间,神经网络在其中达到零误差,而 SGD 会大致找到那个子空间中范数最小的点。可以证明,当维度很高时,这个点对数据中的小随机性不敏感。但是当数据的维度等于模型的维度时,数据集和模型之间存在一一对应关系,所以数据集的微小变化会导致模型的巨大变化,这就是性能变差的原因。

So maybe let me try to give the explanation that will work. So you got a huge neural network, let's suppose you have a huge number of parameters. And now let's pretend everything is linear, which is not, but let's just pretend. Then there is this big subspace where a neural network achieves zero error, and SGD is going to find approximately the point that's right, approximately the point with the smallest norm in that subspace. Okay, and that can also be proven to be insensitive to the small randomness in the data when the dimensionality is high. But when the dimensionality of the data is equal to the dimensionality of the model, then there is a one-to-one correspondence between all the data sets and the models, so small changes in the data set actually lead to large changes in the model, and that's why performance gets worse.

Host

那么模型有更多参数、比数据更大是好的?

So then it would be good for the model to have more parameters, to be bigger than the data?

Ilya

没错,但前提是你不要真的停下来。如果你引入早停或正则化,几乎可以让双重下降的凸起完全消失。

That's right, but only if you don't really stop. If you introduce early stopping or regularization, you can make the double descent bump almost completely disappear.

Host

什么是早停?

What is early stopping?

Ilya

早停就是在训练模型时监控测试或验证性能,如果某个时刻验证性能开始变差,你就说‘好了,停止训练’。如果你表现好,那就足够好了。所以神奇的事情发生在那个时刻之后。所以你不想要早停。如果你不做早停,就会得到一个非常明显的双重下降。

Early stopping is when you train your model and you monitor your test or validation performance, and then if at some point validation performance starts to get worse, you say, 'Okay, let's stop training.' If you're good, you're good enough. So the magic happens after that moment. So you don't want to do the early stopping. Well, if you don't do the early stop, you get a very pronounced double descent.

Host

你对为什么会发生双重下降有什么直觉吗?

Do you have any intuition why this happens? Double descent?

Ilya

哦,抱歉。直觉是这样的:当数据集和模型具有相同数量的自由度时,它们之间存在一一对应关系,因此数据集的微小变化会导致模型的显著变化。所以你的模型对所有随机性都非常敏感,无法丢弃它们。而事实证明,当数据远多于参数,或者参数远多于数据时,得到的解对数据集的微小变化不敏感。所以它能够丢弃那些小的变化、随机性以及你不想要的虚假相关性。

Oh, sorry. The intuition is basically this: when the data set has as many degrees of freedom as the model, then there is a one-to-one correspondence between them, and so small changes to the data set lead to noticeable changes in the model. So your model is very sensitive to all the randomness; it is unable to discard it. Whereas it turns out that when you have a lot more data than parameters, or a lot more parameters than data, the resulting solution will be insensitive to small changes in the data set. So it's able to discard the small changes, the randomness, the spurious correlation which you don't want.

神经网络能否推理 Backpropagation and Alternatives

Host

杰夫·辛顿建议我们需要抛弃反向传播。我们之前已经稍微谈过这个,但他建议我们干脆扔掉反向传播,重新开始。当然,这有点幽默,但你怎么看?训练神经网络可能有什么替代方法?

Jeff Hinton suggested we need to throw backpropagation away. We already kind of talked about this a little bit, but he suggested that we just throw away backpropagation and start over. I mean, of course some of that is a little bit humorous, but what do you think? What could be an alternative method of training neural networks?

Ilya

嗯,他确切的说法是,既然在大脑中找不到反向传播,那么看看能否从大脑的学习方式中学到东西是值得的。但反向传播非常有用,我们应该继续使用它。你的意思是,一旦我们发现了大脑中的学习机制,或者该机制的某些方面,如果事实证明在大脑中找不到反向传播,我们也应该尝试在神经网络中实现它。那么,我想你的回答是反向传播非常有用,所以我们为什么要抱怨呢?我个人是反向传播的忠实粉丝。我认为这是一个伟大的算法,因为它解决了一个非常根本的问题,即在某些约束下找到神经回路。而且我不认为这个问题会消失,所以我认为不太可能出现什么截然不同的东西。有可能发生,但我现在不会押注于此。

Well, the thing that he said precisely is that to the extent you can't find backpropagation in the brain, it's worth seeing if we can learn something from how the brain learns. But backpropagation is very useful and we should keep using it. You're saying that once we discover the mechanism of learning in the brain, or any aspects of that mechanism, we should also try to implement that in neural networks if it turns out that we can't find backpropagation in the brain. Well, I guess your answer to that is backpropagation is pretty damn useful, so why are we complaining? I personally am a big fan of backpropagation. I think it's a great algorithm because it solves an extremely fundamental problem, which is finding a neural circuit subject to some constraints. And I don't see that problem going away, so that's why I think it's pretty unlikely that we'll have anything which is going to be dramatically different. It could happen, but I wouldn't bet on it right now.

神经网络搜索小电路 Can Neural Networks Reason?

Host

你认为神经网络能被训练出推理能力吗?为什么不能?

Do you think neural networks can be made to reason? Why not?

Ilya

嗯,如果你看看 AlphaGo 或 AlphaZero,AlphaZero 的神经网络下围棋——我们都同意围棋需要推理——比 99.9% 的人类都强。仅凭神经网络本身,没有搜索,这难道不是神经网络能推理的存在性证明吗?

Well, if you look at AlphaGo or AlphaZero, the neural network of AlphaZero plays Go, which we all agree is a game that requires reasoning, better than 99.9% of all humans. Just the neural network without search—doesn't that give us an existence proof that neural networks can reason?

Host

稍微反驳一下,我们都同意围棋是推理。我认为它有一些相同的元素:推理类似于搜索,有顺序地逐步考虑可能性,并在此基础上构建,直到得出某个洞见。所以下围棋就是这样。当单个神经网络在没有搜索的情况下做到这一点时,这是在受限环境中存在类似推理过程的存在性证明。但更通用的推理,棋盘之外——还有另一个存在性证明:我们人类。

To push back a little, we all agree that Go is reasoning. I think it has some of the same elements: reasoning is akin to search, with a sequential stepwise consideration of possibilities, building on them until you arrive at some insight. So playing Go is like that. When you have a single neural network doing that without search, that's an existence proof in a constrained environment that a process akin to reasoning exists. But more general reasoning, off the board—there is another existence proof: us humans.

Ilya

好吧。那么你认为能让神经网络推理的架构会与今天的架构相似吗?

Okay, all right. So do you think the architecture that will allow neural networks to reason will look similar to today's architectures?

Host

我认为会的。我不想说得太绝对,但产生推理突破的神经网络很可能与当前架构非常相似——可能更深一些。这些新路线如此强大,为什么它们不能学会推理?人类能推理,为什么神经网络不能?

I think it will. I don't want to make overly definitive statements, but it's possible that the neural networks producing reasoning breakthroughs will be very similar to current architectures—maybe a bit deeper. These new lines are so insanely powerful, why wouldn't they be able to learn to reason? Humans can reason, so why can't neural networks?

Ilya

那么你认为我们看到的神经网络做的事情只是弱推理,而不是根本不同的过程?再说一次,没人知道答案。我会说神经网络能够推理,但如果你在不需要推理的任务上训练它,它就不会推理。这是一个众所周知的效果:神经网络会以最简单的方式解决你提出的问题。

So do you think the kind of stuff we've seen neural networks do is just weak reasoning, not a fundamentally different process? Again, nobody knows the answer. I would say neural networks are capable of reasoning, but if you train a neural network on a task that doesn't require reasoning, it won't reason. This is a well-known effect: the neural network solves the problem you pose in the easiest way possible.

神经网络中的长期记忆 Neural Networks as Search for Small Circuits

Host

这引出了你描述神经网络的一个精彩方式:你把它们称为对小电路的搜索,而通用智能是对小程序的搜索。你能详细说明这个区别吗?

That takes us to one of the brilliant ways you describe neural networks: you refer to them as the search for small circuits, and general intelligence as the search for small programs. Can you elaborate on that difference?

Ilya

是的。我当时的准确说法是:如果你能找到输出你手中数据的最短程序,那么你就可以用它做出最佳预测。这是一个可以数学证明的理论陈述。你还可以证明,找到最短程序是不可计算的——没有有限的计算量能做到。所以神经网络是实际可行的次优选择。我们找不到最短程序,但我们可以找到一个拟合数据的小电路。现在我会修正这一点:甚至是一个大电路,其权重包含少量信息。如果你把训练想象成将熵从数据集缓慢传输到参数,那么权重中的信息量最终不会很大,这解释了为什么它们泛化得这么好。大电路可能有助于正则化和泛化。

Yes. What I said precisely was: if you can find the shortest program that outputs the data at your disposal, then you can use it to make the best prediction possible. That's a theoretical statement that can be proven mathematically. You can also prove that finding the shortest program is not computable—no finite amount of compute can do it. So neural networks are the next best thing that actually works in practice. We can't find the shortest program, but we can find a small circuit that fits our data. Now I would amend that: even a large circuit whose weights contain a small amount of information. If you imagine training as slowly transmitting entropy from the dataset to the parameters, the amount of information in the weights ends up not being very large, which explains why they generalize so well. The large circuit might be helpful for regularization and generalization.

Host

你认为尝试学习类似程序的东西重要吗?

Do you see it as important to try to learn something like programs?

Ilya

当然。如果我们能做到,就应该去做。我们推动深度学习的原因是我们可以训练它们。训练是第一位的——这是一个我们不能违反的支柱。可训练意味着从零开始,并收敛到知道很多。我们不能离开那个支柱。如果我们说“让我们找到最短程序”,我们做不到,所以它再有用也没用。

Definitely. If we can do it, we should. The reason we push on deep learning is that we are able to train them. Training comes first—it's a pillar we cannot violate. Being trainable means starting from scratch and converging towards knowing a lot. We can't move away from that pillar. If we say 'let's find the shortest program,' we can't do that, so it doesn't matter how useful it would be.

Host

你认为找到小程序的问题仅仅在于数据吗?

Do you think the matter of finding small programs is just about the data?

Ilya

不。目前,还没有人成功找到程序的好先例。找到程序的方法是训练深度神经网络来做,这是正确的方向。但还没有实现。原则上,应该是可能的。

No. Right now, there are no good precedents of people successfully finding programs really well. The way you'd find programs is to train a deep neural network to do it, which is the right approach. But it hasn't been done yet. In principle, it should be possible.

Host

你能详细说明你“原则上可能”的见解吗?你不认为为什么不可能?

Can you elaborate on your insight that in principle it's possible? You don't see why it's not possible?

Ilya

这更像是说,押注深度学习失败是不明智的。如果这是一种人类似乎能做的认知功能,那么用不了多久就会出现一个能做的深度神经网络。我已经不再押注神经网络不行了,因为它们不断给我们惊喜。

It's more a statement that it's unwise to bet against deep learning. If it's a cognitive function that humans seem to be able to do, then it doesn't take too long for some deep neural net to pop up that can do it too. I've stopped betting against neural networks because they continue to surprise us.

长期记忆与知识 Long-Term Memory in Neural Networks

Host

那长期记忆呢?神经网络能有长期记忆吗,比如知识库,能够长时间聚合重要信息,然后作为有用的上下文?

What about long-term memory? Can neural networks have long-term memory, like knowledge bases, being able to aggregate important information over long periods of time that would then serve as useful context?

推理基准与惊人成就 Long-term memory and knowledge in neural networks

Host

你可以根据状态表征来做决策,因此拥有基于决策的长期上下文。从某种意义上说,参数已经做到了这一点。参数是神经网络全部经验的聚合,因此它们算作长期知识。人们已经训练了各种神经网络作为知识库,并研究了语言模型作为知识库。所以这方面是有工作的。但从某种意义上说,你认为这只是一个想出更好的机制来忘记无用信息、记住有用信息的问题吗?因为目前还没有真正记住长期信息的机制。

Representations of state that you can make decisions by, so have a long-term context based on what you make in the decision. So in some sense, the parameters already do that. The parameters are an aggregation of the entirety of the neural net's experience, and so they count as long-term knowledge. People have trained various neural nets to act as knowledge bases and investigated language models as knowledge bases. So there is work there. But in some sense, do you think it's all just a matter of coming up with a better mechanism of forgetting the useless stuff and remembering the useful stuff? Because right now, there are not mechanisms that do remember really long-term information.

Ilya

你具体指什么?我喜欢“具体”这个词。

What do you mean by that precisely? I like the word precisely.

Host

我在想知识库所代表的那种信息压缩,有点像创建……我为我以人类为中心的关于知识是什么的思考道歉,因为神经网络不一定能解释它们发现的知识。但对我来说,一个好的例子是知识库能够随着时间的推移构建起类似维基百科所代表的知识。这是一个非常压缩的结构化知识库,显然不是实际的维基百科或语言,而是像语义网,语义网所代表的梦想。所以这是一个非常好的压缩知识库,或者神经网络在不可解释意义上类似的东西。

I'm thinking of the kind of compression of information the knowledge bases represent, sort of creating a... I apologize for my human-centric thinking about what knowledge is, because neural networks aren't interpretable necessarily with the kind of knowledge they have discovered. But a good example for me is knowledge bases being able to build up over time something like the knowledge that Wikipedia represents. It's a really compressed structured knowledge base, obviously not the actual Wikipedia or the language, but like a semantic web, the dream that semantic web represented. So it's a really nice compressed knowledge base, or something akin to that in the non-interpretable sense as neural networks would have.

Ilya

嗯,如果你看神经网络的权重,它们是不可解释的,但它们的输出应该是非常可解释的。那么如何让像语言模型这样非常智能的神经网络变得可解释呢?你让它们生成一些文本,那么文本通常就是可解释的。

Well, the neural networks would be non-interpretable if you look at their weights, but their outputs should be very interpretable. So how do you make very smart neural networks like language models interpretable? Well, you ask them to generate some text, then the text will generally be interpretable.

Host

你认为这是可解释性的典范吗?比如,你能做得更好吗?因为你不能……我想知道它知道什么和不知道什么。我希望神经网络能举出它完全愚蠢的例子和完全聪明的例子。我现在知道的唯一方法就是生成大量例子并用我的人类判断力。但如果神经网络对此有一些自我意识就好了。

Do you find that the epitome of interpretability? Like, can you do better? Because you can't... I'd like to know what it knows and what it doesn't know. I would like the neural network to come up with examples where it's completely dumb and examples where it's completely brilliant. And the only way I know how to do that now is to generate a lot of examples and use my human judgment. But it would be nice if a neural network had some self-awareness about it.

Ilya

是的,100%。我非常相信自我意识。我认为神经网络的自我意识将允许实现你描述的那些能力,比如让它们知道自己知道什么和不知道什么,以及知道在哪里投入以最优化地提升技能。至于你的可解释性问题,实际上有两个答案。一个答案是,我们有神经网络,所以我们可以分析神经元,并尝试理解不同神经元和不同层的含义。你确实可以这样做,OpenAI 在这方面做了一些工作。但还有另一个答案,即以人类为中心的答案。你看一个人,你无法读懂他们的思想。你怎么知道一个人在思考什么?你问他们:‘嘿,你对这个怎么看?你对那个怎么看?’然后你得到一些答案。你得到的答案具有粘性,因为你已经对那个人有了一个心智模型。你已经对那个人是什么、他们如何思考、他们知道什么、他们如何看待世界有了一个大致的理解。然后你问的每件事,你都在添加到那个模型上。这种粘性似乎是人类一个非常有趣的特性:信息是有粘性的。你似乎能记住有用的东西,很好地聚合它们,并忘记大部分无用的信息。这个过程与神经网络所做的也非常相似,只是目前神经网络在这方面要糟糕得多。这似乎并没有根本性的不同。

Yeah, 100%. I'm a big believer in self-awareness. I think neural net self-awareness will allow for things like the capabilities you describe, like for them to know what they know and what they don't know, and for them to know where to invest to increase their skills most optimally. And to your question of interpretability, there are actually two answers. One answer is we have the neural net, so we can analyze the neurons and try to understand what the different neurons and different layers mean. You can actually do that, and OpenAI has done some work on that. But there is a different answer, which is the human-centric answer. You look at a human being, you can't read their mind. How do you know what a human being is thinking? You ask them: 'Hey, what do you think about this? What do you think about that?' And you get some answers. The answers you get are sticky in the sense that you already have a mental model of that human being. You already have an understanding of a big conception of what that human being is, how they think, what they know, how they see the world. And then everything you ask, you're adding on to that. That stickiness seems to be one of the really interesting qualities of human beings: information is sticky. You seem to remember the useful stuff, aggregate it well, and forget most of the information that's not useful. That process is also pretty similar to what neural networks do, it's just that neural networks are so much crappier at it at this time. It doesn't seem to be fundamentally that different.

语言中神经网络的历史 Reasoning benchmarks and impressive feats

Host

但为了在推理上再停留一会儿,你说为什么我不能推理?对你来说,什么样的推理成就基准会让你印象深刻?如果你不知道我们能做什么,你心里已经有想法了吗?

But just to stick on reasoning for a little longer, you said why can't I reason? What's a good impressive feat benchmark to you of reasoning that you'll be impressed by? If you don't know what we're able to do, is that something you already have in mind?

Ilya

嗯,我认为写出非常好的代码、证明非常难的定理、用创造性解决方案解决开放性问题,以及定理类数学问题。我认为这些也是非常自然的例子。你知道,如果你能证明一个未证明的定理,那么很难说你不推理。顺便说一句,这又回到了关于硬结果的观点。机器学习和深度学习作为一个领域非常幸运,因为我们有时能够产生这些明确的结果。当它们发生时,辩论会改变,对话会改变。我们有能力产生改变对话的结果。然后当然,就像你说的,人们会认为这是理所当然的,并说那其实不是一个难题。嗯,在某个时候,我们可能会用完难题。是的,整个死亡问题是一个我们还没有完全解决的棘手问题。也许我们会解决它。

Well, I think writing really good code, proving really hard theorems, solving open-ended problems with out-of-the-box solutions, and sort of theorem-type mathematical problems. I think those ones are a very natural example as well. You know, if you can prove an unproven theorem, then it's hard to argue you don't reason. And by the way, this comes back to the point about hard results. Machine learning and deep learning as a field is very fortunate because we have the ability to sometimes produce these unambiguous results. And when they happen, the debate changes, the conversation changes. We have the ability to produce conversation-changing results. And then of course, just like you said, people kind of take that for granted and say that wasn't actually a hard problem. Well, at some point we'll probably run out of hard problems. Yeah, that whole mortality thing is kind of a sticky problem that we haven't quite figured out. Maybe we'll solve that one.

扩展与语义理解 History of neural networks in language

Host

我认为在你整个工作以及最近 OpenAI 的工作中,一个引人入胜的事情是语言模型领域的一个对话改变者。你能简要描述一下在语言和文本领域使用神经网络的最新历史吗?

I think one of the fascinating things in your entire body of work, but also the work at OpenAI recently, one of the conversation changers has been in the world of language models. Can you briefly try to describe the recent history of using neural networks in the domain of language and text?

Ilya

嗯,有很多历史。我认为 Elman 网络是 80 年代应用于语言的一个小型循环神经网络。所以历史确实相当长。而改变神经网络和语言轨迹的东西,就是改变深度学习轨迹的东西:数据和算力。所以突然之间,你从学习一点点的小语言模型转变过来。特别是对于语言模型,有一个非常清晰的解释,说明为什么它们需要大才能好,因为它们试图预测下一个词。当你什么都不知道时,你会注意到非常粗略的表面模式,比如有时有字符,字符之间有空格。你会注意到这个模式,你会注意到有时有逗号,然后下一个字符是大写字母。你会注意到那个模式。最终你可能会开始注意到某些词经常出现。你可能会注意到拼写是一个东西。你可能会注意到语法。当你所有这些都做得很好时,你开始注意到语义。你开始注意到……

Well, there's been lots of history. I think the Elman network was a small tiny recurrent neural network applied to language back in the 80s. So the history is really fairly long. And the thing that changed the trajectory of neural networks and language is the thing that changed the trajectory of deep learning: data and compute. So suddenly you move from small language models which learn a little bit. And with language models in particular, there is a very clear explanation for why they need to be large to be good, because they're trying to predict the next word. When you don't know anything, you'll notice very broad stroke surface-level patterns, like sometimes there are characters and there is a space between those characters. You'll notice this pattern, and you'll notice that sometimes there is a comma and then the next character is a capital letter. You'll notice that pattern. Eventually you may start to notice that there are certain words that occur often. You may notice that spellings are a thing. You may notice syntax. And when you get really good at all these, you start to notice the semantics. You start to notice the...

GPT-2 与 Transformer 架构 Scaling and Semantic Understanding

Host

事实如此,但要实现这一点,语言模型需要更大。所以我们来深入探讨一下,因为这是你和诺姆·乔姆斯基可能意见相左的地方。你认为我们实际上是在采取渐进步骤:更大的网络、更多的算力将能够触及语义,能够理解语言,而无需乔姆斯基所认为的那种对语言结构的基本理解,比如将你的语言理论强加给学习机制。所以你是说,学习可以从原始数据中学习语言背后的机制?

Facts, but for that to happen the language model needs to be larger. So let's linger on that, because that's where you and Noam Chomsky could disagree. So you think we're actually taking incremental steps: a sort of larger network, larger compute will be able to get to the semantics, to be able to understand language without what Chomsky likes to think of as a fundamental understanding of the structure of language, like imposing your theory of language onto the learning mechanism. So you're saying the learning can learn from raw data the mechanism that underlies language?

Ilya

嗯,我认为这很有可能,但我也想说我并不确切知道乔姆斯基所说的“将你的结构强加给语言”是什么意思。我不太确定他的意思。但从经验上看,当你检查那些更大的语言模型时,它们表现出理解语义的迹象,而较小的语言模型则没有。几年前我们在做情感神经元的工作时就看到了这一点。我们训练了一个小型 LSTM 来预测亚马逊评论中的下一个字符,我们注意到当 LSTM 的规模从 500 个 LSTM 单元增加到 4000 个时,其中一个神经元开始表示评论的情感。这是为什么呢?情感是一个非常语义的属性,而不是句法属性。对于那些可能不知道的人来说,情感就是评论是正面的还是负面的。所以我们有非常清晰的证据表明,小型神经网络无法捕捉情感,而大型神经网络可以。为什么会这样?我们的理论是,在某个点上,你耗尽了要建模的句法,于是开始关注其他东西。随着规模的扩大,你很快耗尽了要建模的句法,然后你真的开始关注语义。就是这个想法。

Well, I think it's pretty likely, but I also want to say that I don't really know precisely what Chomsky means when he talks about imposing your structure on language. I'm not 100% sure what he means. But empirically, it seems that when you inspect those larger language models, they exhibit signs of understanding the semantics, whereas the smaller language models do not. We've seen that a few years ago when we did work on the sentiment neuron. We trained a small LSTM to predict the next character in Amazon reviews, and we noticed that when you increase the size of the LSTM from 500 LSTM cells to 4000 LSTM cells, then one of the neurons starts to represent the sentiment of the review. Now, why is that? Sentiment is a pretty semantic attribute, it's not a syntactic attribute. And for people who might not know, sentiment is whether it's a positive or negative review. So here we had very clear evidence that a small neural net does not capture sentiment while a large neural net does. And why is that? Well, our theory is that at some point you run out of syntax to model, so you start to focus on something else. With size, you quickly run out of syntax to model, and then you really start to focus on the semantics. That's the idea.

Host

所以我不想暗示我们的模型有完全的语义理解,因为那不是真的,但它们确实显示出语义理解的迹象,部分语义理解,而较小的模型没有显示出这些迹象。

And so I don't want to imply that our models have complete semantic understanding, because that's not true, but they definitely are showing signs of semantic understanding, partial semantic understanding, but the smaller models do not show those signs.

对 Transformer 性能的惊讶 GPT-2 and Transformer Architecture

Host

你能退一步说说什么是 GPT-2 吗?它是过去几年改变对话的大型语言模型之一。

Can you take a step back and say what is GPT-2, which is one of the big language models that was the conversation changer in the past couple of years?

Ilya

是的。GPT-2 是一个拥有 15 亿参数的 Transformer,它在大约 400 亿个文本 token 上进行了训练,这些文本来自 Reddit 上获得超过三个赞的文章所链接的网页。

Yes. So GPT-2 is a Transformer with one and a half billion parameters that was trained on about 40 billion tokens of text, which were obtained from web pages that were linked to from Reddit articles with more than three upvotes.

Host

什么是 Transformer?Transformer 是近期历史上神经网络架构最重要的进步。什么是注意力机制?也许这也是个问题。因为我认为那是个有趣的想法,不一定是技术上的,而是注意力机制的概念与循环神经网络所代表的概念的对比。

And what's the Transformer? The Transformer is the most important advance in neural network architectures in recent history. What is attention, maybe too? Because I think that's the interesting idea, not necessarily technically speaking, but the idea of attention versus maybe what recurrent neural networks represent.

Ilya

是的,问题是,Transformer 是多个想法的同时结合,注意力机制是其中之一。你认为注意力机制是关键吗?不,它是一个关键,但不是唯一的关键。Transformer 之所以成功,是因为它同时结合了多个想法,如果你去掉其中任何一个,它都会逊色很多。所以 Transformer 使用了大量的注意力机制,但注意力机制已经存在了好几年,所以那不可能是主要的创新。Transformer 的设计方式使其在 GPU 上运行得非常快,这带来了巨大的不同。这是一点。第二点是 Transformer 不是循环的,这也很重要,因为它更浅,因此更容易优化。换句话说,它使用了注意力机制,非常适合 GPU,并且不是循环的,因此更浅且更容易优化。这些因素的结合使其成功。所以现在它充分利用了你的 GPU,让你在相同的算力下获得更好的结果,这就是它成功的原因。

Yeah, so the thing is, the Transformer is a combination of multiple ideas simultaneously, of which attention is one. Do you think attention is the key? No, it's a key, but it's not the key. The Transformer is successful because it is the simultaneous combination of multiple ideas, and if you were to remove either idea, it would be much less successful. So the Transformer uses a lot of attention, but attention existed for a few years, so that can't be the main innovation. The Transformer is designed in such a way that it runs really fast on the GPU, and that makes a huge amount of difference. This is one thing. The second thing is the Transformer is not recurrent, and that is really important too, because it is more shallow and therefore much easier to optimize. So in other words, it uses attention, it is a really great fit to the GPU, and it is not recurrent, therefore less deep and easier to optimize. The combination of those factors make it successful. So now it makes great use of your GPU, it allows you to achieve better results for the same amount of compute, and that's why it's successful.

翻译与经济影响 Surprise at Transformer Performance

Host

你对 Transformer 和 GPT-2 的效果感到惊讶吗?你从事语言研究,在 Transformer 出现之前就有很多好想法,所以你看到了前后的整个革命过程。你感到惊讶吗?

Were you surprised how well Transformers worked and GPT-2 worked? So you worked on language, you've had a lot of great ideas before Transformers came about in language, so you got to see the whole set of revolutions before and after. Were you surprised?

Ilya

是的,有一点。我的意思是,很难记住,因为你适应得很快,但这确实令人惊讶。事实上,我要收回我的话:这相当惊人。看到它生成这种质量的文本真是令人惊叹。而且你要知道,当时我们已经看到了 GAN 的所有进展,GAN 生成的样本非常惊人,有逼真的面孔,但文本并没有太大的进步。突然之间,我们从 2015 年的 GAN 一步跨越到最好、最惊人的 GAN,对吧?这真的很震撼,尽管理论预测如果你训练一个大型语言模型,当然应该得到这样的结果。但亲眼看到又是另一回事。然而我们适应得很快,现在有一些认知科学家写文章说 GPT-2 模型并不真正理解语言。所以我们很快就适应了它们能够如此好地建模语言这一事实的惊人之处。

Yeah, a little. I mean, it's hard to remember because you adapt really quickly, but it definitely was surprising. In fact, I'll retract my statement: it was pretty amazing. It was just amazing to see it generate text of this quality. And you've got to keep in mind that at that time, we had seen all this progress in GANs, improving the samples produced by GANs were just amazing, you had these realistic faces, but text hadn't really moved that much. And suddenly we moved from whatever GANs were in 2015 to the best, most amazing GANs in one step, right? And it was really stunning, even though theory predicted that if you train a big language model, of course you should get this. But then to see it with your own eyes is something else. And yet we adapt really quickly, and now there are cognitive scientists writing articles saying that GPT-2 models don't truly understand language. So we adapt quickly to how amazing the fact that they're able to model the language so well is.

Host

那么你认为让我们印象深刻的门槛是什么?我不知道。你认为这个门槛会不断被提高吗?

So what do you think is the bar for impressing us? That it... I don't know. Do you think that bar will continuously be moved?

Ilya

当然。我认为当你开始看到真正巨大的经济影响时,那在某种意义上就是下一个门槛。因为现在,如果你思考正在运行的人工智能,它真的很令人困惑,很难理解所有这些进步意味着什么。有点像,好吧,你取得了一个进步,现在你可以做更多事情,然后你又得到了另一个改进,又有了一个很酷的演示。在某个时候,我认为人工智能领域之外的人再也无法区分这种进步了。

Definitely. I think when you start to see really dramatic economic impact, that's when I think that's in some sense the next barrier. Because right now, if you think about the working AI, it's really confusing, it's really hard to know what to make of all these advances. It's kind of like, okay, you got an advance and now you can do more things, and you got another improvement, and you got another cool demo. At some point, I think people who are outside of AI can no longer distinguish this progress anymore.

语言与视觉模型的联系 Translation and Economic Impact

Host

我们离线时谈到过将俄语翻译成英语,以及有很多优秀的俄语作品不为世界其他地方所知。这对中文来说也是如此,对许多科学家和一般的艺术作品也是如此。你认为翻译是我们将看到巨大经济影响的地方吗?

So we were talking offline about translating Russian to English and how there's a lot of brilliant work in Russian that the rest of the world doesn't know about. That's true for Chinese, that's true for a lot of scientists and just artistic work in general. Do you think translation is the place where we're going to see sort of economic big impact?

Ilya

我不知道。我认为有大量的……我的意思是,首先,我想指出翻译在今天已经非常庞大了。我认为数十亿人主要通过翻译与互联网的大部分内容互动。所以翻译已经非常庞大,而且也非常积极。我认为自动驾驶将产生巨大影响,虽然不知道具体何时发生,但我不会押注深度学习会失败。

I don't know. I think there is a huge number of... I mean, first of all, I want to point out that translation already today is huge. I think billions of people interact with big chunks of the internet primarily through translation. So translation is already huge, and it's hugely positive too. I think self-driving is going to be hugely impactful, and it's unknown exactly when it happens, but again, I would not bet against deep learning.

Host

所以这是深度学习整体而言,但你一直……

So that's deep learning in general, but you keep...

扩展与 GPT-2 Connection between language and vision models

Host

自动驾驶的学习,是的,深度学习用于自动驾驶,但我刚才说的是语言模型。我们稍微岔开一下,确认一下:你不觉得驾驶和语言之间有联系吗?

Learning for self-driving, yes, deep learning for self-driving, but I was talking about language models. Let's see, just to spear it off a little bit, just to check: you're not seeing a connection between driving and language?

Ilya

不,不。好吧,好吧。它们都用神经网络。会有一种诗意的联系。我想可能像你说的,会有某种统一,走向一种能同时处理语言和视觉任务的多任务 Transformer。那会是一个有趣的统一。

No, no. Okay, all right. They both use neural nets. There'll be a poetic connection. I think there might be some, like you said, there might be some kind of unification towards a kind of multi-task transformer that can take on both language and vision tasks. That'd be an interesting unification.

AI 模型的负责任发布 Scaling and GPT-2

Host

现在看看,关于 GPT-2 我还能问什么?很简单,没什么好问的。就是这样:你拿一个 Transformer,把它做大,给它更多数据,突然它就做出了所有那些惊人的事情。是的,美妙之处在于 GPT-2,Transformer 从根本上来说很容易解释和训练。你认为更大的模型会继续在语言上表现出更好的结果吗?可能吧。有点像,你觉得 GPT-2 的下一步是什么?

Now let's see, what can I ask about GPT-2 more? It's simple, it's not much to ask. It's so: you take a transformer, you make it bigger, you give it more data, and suddenly it does all those amazing things. Yeah, one of the beautiful things is that GPT-2, the transformers are fundamentally simple to explain and to train. Do you think bigger will continue to show better results in language? Probably. Sort of like, what are the next steps with GPT-2, do you think?

Ilya

我肯定,看看更大版本能做什么是一个方向。另外,还有很多问题。有一个问题我很好奇,就是下面这个。现在,GPT-2 我们喂给它来自互联网的所有数据,这意味着它需要记住互联网上所有那些随机事实。如果模型能以某种方式用自己的智能决定它想学习、接受哪些数据,拒绝哪些数据,那就好了,就像人一样。人们不会不加选择地学习所有数据;我们对自己学什么非常挑剔。我认为这种主动学习会非常好。

I think for sure, seeing what larger versions can do is one direction. Also, there are many questions. There's one question which I'm curious about, and that's the following. So right now, GPT-2, we feed all this data from the internet, which means that it needs to memorize all those random facts about everything on the internet. And it would be nice if the model could somehow use its own intelligence to decide what data it wants to study, accept, and what data it wants to reject, just like people. People don't learn all data indiscriminately; we are super selective about what we learn. And I think this kind of active learning would be very nice to have.

Host

是的,听着,我喜欢主动学习。所以让我问:数据的选择——你能再详细说明一下吗?你认为数据的选择是,我有这种感觉,如何选择数据的优化,也就是主动学习过程,即使在不久的将来也会成为很多突破的地方,因为那里公开的突破不多。我觉得可能有公司保密的私下突破,因为如果你想解决自动驾驶,如果你想解决一个特定任务,这个根本问题必须解决。但你对这个领域总体怎么看?

Yeah, listen, I love active learning. So let me ask: does the selection of data—can you just elaborate that a little bit more? Do you think the selection of data is, like, I have this kind of sense that the optimization of how you select data, so the active learning process, is going to be a place for a lot of breakthroughs even in the near future, because there hasn't been many breakthroughs there that are public. I feel like there might be private breakthroughs that companies keep to themselves, because the fundamental problem has to be solved if you want to solve self-driving, if you want to solve a particular task. But what do you think about the space in general?

Ilya

是的,所以我认为对于像主动学习这样的东西,或者实际上对于任何像主动学习这样的能力,它真正需要的是一个问题。它需要一个需要它的问题。如果你没有任务,就很难研究这种能力,因为那样的话,你会想出一个人工任务,得到好结果,但并不能真正说服任何人,对吧?就像,我们现在已经过了在 MNIST 上用一些巧妙的公式得到结果就能说服人的阶段了。没错。事实上,你可以很容易地在 MNIST 上提出一个简单的主动学习方案,获得 10 倍加速,但那又怎样?我认为主动学习需要——主动学习会随着需要它的问题的出现而自然产生。这就是我的看法。

Yeah, so I think that for something like active learning, or in fact for any kind of capability like active learning, the thing that it really needs is a problem. It needs a problem that requires it. It's very hard to do research about the capability if you don't have a task, because then what's going to happen is you will come up with an artificial task, get good results, but not really convince anyone, right? Like, we're now past the stage where getting a result on MNIST with some clever formulation will convince people. That's right. In fact, you could quite easily come up with a simple active learning scheme on MNIST and get a 10x speedup, but then so what? And I think that with active learning, it needs—active learning will naturally arise as there are problems that require it pop up. That's how I would, that's my take on it.

AI 开发竞赛 Responsible release of AI models

Host

OpenAI 在 GPT-2 上还提出了另一个有趣的事情:当你创建一个强大的人工智能系统时,它会产生什么样的有害影响并不清楚。因为如果你有一个能生成相当逼真文本的 AI 模型,你可以开始想象它会被机器人以某种我们甚至无法想象的方式使用。所以对于它能做什么,存在这种紧张感。所以你做了一件非常勇敢、我认为意义深远的事情:你发起了一场关于这个的对话——比如,我们如何向公众发布强大的人工智能模型?如果我们真的发布,我们如何私下与其他公司,甚至是竞争对手,讨论我们如何管理系统的使用等等?所以从整个经历中,你发布了一份报告。但总的来说,你从思考如何发布这样的模型中得到了什么见解吗?

There's another interesting thing that OpenAI has brought up with GPT-2, which is: when you create a powerful artificial intelligence system, and it was unclear what kind of detrimental effect it will have. Because if you have an AI model that can generate pretty realistic text, you can start to imagine that it would be used by bots in some way that we can't even imagine. So there's this nervousness about what it's possible to do. So you did a really kind of brave and, I think, profound thing: you started a conversation about this—like, how do we release powerful artificial intelligence models to the public? If we do it at all, how do we privately discuss with other, even competitors, about how we manage the use of the systems and so on? So from that whole experience, you released a report on it. But in general, are there any insights that you've gathered from just thinking about this about how you release models like this?

Ilya

我认为我的看法是,AI 领域一直处于童年状态,现在它正在走出那个状态,进入成熟状态。这意味着 AI 非常成功,也很有影响力,而且它的影响不仅大,还在增长。因此,在发布系统之前开始思考其影响似乎是明智的,也许早一点而不是晚一点。以 GPT-2 为例,就像我之前提到的,结果确实令人震惊,而且似乎合理——不是确定,而是似乎合理——像 GPT-2 这样的东西很容易被用来降低虚假信息的成本。所以问题是什么是最好的发布方式,分阶段发布似乎合乎逻辑。一个小模型被发布了,有时间看看有多少人以各种很酷的方式使用这些模型。有很多非常酷的应用;我们不知道有任何负面应用。所以最终它被发布了。但其他人也复制了类似的模型。不过,这是一个有趣的问题:就我们所知。

I think that my take on this is that the field of AI has been in a state of childhood, and now it's exiting that state and entering a state of maturity. What that means is that AI is very successful and also very impactful, and its impact is not only large but it's also growing. So for that reason, it seems wise to start thinking about the impact of our systems before releasing them, maybe a little bit too soon rather than a little bit too late. And with the case of GPT-2, like I mentioned earlier, the results really were stunning, and it seemed plausible—it didn't seem certain, it seemed plausible—that something like GPT-2 could easily be used to reduce the cost of disinformation. And so there was a question of what's the best way to release it, and staged release seemed logical. A small model was released, and there was time to see how many people use these models in lots of cool ways. There have been lots of really cool applications; there haven't been any negative applications we know of. And so eventually it was released. But also, other people replicated similar models. That's an interesting question, though: that we know of.

Host

所以在你看来,分阶段发布至少是问题答案的一部分:我们一旦创建了这样的系统该怎么做?这是答案的一部分。是的。还有其他见解吗?比如,说你根本不想发布模型,因为它对你的业务有用。嗯,已经有很多人不发布模型了,对吧?当然。但是当你有一个非常强大的模型时,是否有某种道德伦理责任去沟通,就像你说的?当你拥有 GPT-2 时,不清楚它可能被用于多少虚假信息。这是一个悬而未决的问题,要得到答案可能需要你和外部其他非常聪明的人交谈。请告诉我,有没有某种乐观的途径让世界各地的人们在这些案例上合作,还是说一家公司与另一家公司交谈仍然非常困难?

So in your view, staged release is at least part of the answer to the question of how do we—what do we do once we create a system like this? It's part of the answer. Yes. Is there any other insights? Like, say you don't want to release the model at all because it's useful to you for whatever the business is. Well, there are plenty of people who don't release models already, right? Of course. But is there some moral ethical responsibility when you have a very powerful model to sort of communicate, just as you said? When you had GPT-2, it was unclear how much it could be used for misinformation. It's an open question, and getting an answer to that might require that you talk to other really smart people that are outside your particular group. Have you—please tell me there's some optimistic pathway for people across the world to collaborate on these kinds of cases, or is it still really difficult for one company to talk to another company?

Ilya

这绝对是可能的。与其他地方的同事讨论这类模型,听取他们对如何做的看法,绝对是可能的。不过有多难呢?我是说,你看到这种情况发生吗?我认为这是一个需要逐步建立公司间信任的地方,因为最终所有 AI 开发者都在构建注定会越来越强大的技术。所以思考方式是,最终我们都在同一条船上。是的,我倾向于相信我们本性中更好的一面,但我确实希望,当你在某个特定领域构建一个非常强大的 AI 系统时,你也会考虑其影响。

It's definitely possible. It's definitely possible to discuss these kinds of models with colleagues elsewhere and to get their take on what to do. How hard is it, though? I mean, do you see that happening? I think that's a place where it's important to gradually build trust between companies, because ultimately all the AI developers are building technology which is bound to be increasingly more powerful. And so the way to think about it is that ultimately we're all in this together. Yeah, it's—I tend to believe in the better angels of our nature, but I do hope that when you build a really powerful AI system in a particular domain, you also think about the implications.

构建 AGI:深度学习加思想 Race for AI development

Host

潜在负面后果是,一个有趣又可怕的可能性是,AI 开发会变成一场竞赛,迫使人们封闭开发,不与他人分享想法。我不喜欢这样。我当了十年纯粹的学者,真的很喜欢分享想法,这很有趣,也很激动人心。

Potential negative consequences of it's an interesting and scary possibility that it'll be a race for AI development that would push people to close that development and not share ideas with others. I don't love this. I've been like a pure academic for 10 years. I really like sharing ideas and it's fun, it's exciting.

AGI 的模拟与现实世界 Building AGI: deep learning plus ideas

Host

你认为构建一个人类级别的智能系统需要什么?我们谈到了推理、长期记忆,但总的来说,需要什么?

What do you think it takes to build a system of human-level intelligence? We talked about reasoning, we talked about long-term memory, but in general, what does it take?

Ilya

嗯,我不确定,但我觉得是深度学习加上也许另一个小想法。

Well, I can't be sure, but I think deep learning plus maybe another small idea.

Host

你认为自我对弈会参与其中吗?你谈到过自我对弈的强大机制,系统通过在竞争环境中与技能相近的其他实体探索世界来学习,并逐步改进。你认为自我对弈会是构建 AGI 系统的一个组成部分吗?

Do you think self-play will be involved? You've spoken about the powerful mechanism of self-play where systems learn by exploring the world in a competitive setting against other entities that are similarly skilled, and so incrementally improve. Do you think self-play will be a component of building an AGI system?

Ilya

是的,所以我要说的是:构建 AGI,我认为将是深度学习加上一些想法,而自我对弈将是其中之一。自我对弈有一个惊人的特性:它能以真正新颖的方式给我们惊喜。例如,几乎每个自我对弈系统——无论是 DotaBot,还是 OpenAI 发布的多智能体捉迷藏,当然还有 AlphaZero——都产生了令人惊讶的行为,创造性地解决问题。这似乎是 AGI 的一个重要部分,而我们的系统目前还不能常规地展现出来。这就是我喜欢这个方向的原因,因为它能给我们惊喜。一个 AGI 系统会从根本上让我们惊讶。

Yeah, so what I would say: to build AGI, I think is going to be deep learning plus some ideas, and I think self-play will be one of those ideas. Self-play has this amazing property that it can surprise us in truly novel ways. For example, pretty much every self-play system — both DotaBot, I don't know if OpenAI had a release about multi-agent where you had two little agents playing hide and seek, and of course AlphaZero — they all produced surprising behaviors, creative solutions to problems. That seems like an important part of AGI that our systems don't exhibit routinely right now. That's why I like this direction, because of its ability to surprise us. An AGI system would surprise us fundamentally.

Host

但准确地说,不是随机的惊喜,而是找到问题的惊喜解决方案,同时还要有用。

But to be precise, not just a random surprise, but to find a surprising solution to a problem that's also useful.

Ilya

对。

Right.

AGI 的具身与意识 Simulation vs real world for AGI

Host

很多自我对弈机制都用在游戏或至少模拟环境中。你认为在通往 AGI 的道路上,有多少会在模拟中完成?你对模拟有多大信心,相比之下,让系统在真实世界(无论是数字真实世界数据还是实际的物理机器人世界)中运行呢?

A lot of self-play mechanisms have been used in game contexts or at least in simulation contexts. How far along the path to AGI do you think will be done in simulation? How much faith do you have in simulation versus having a system that operates in the real world, whether digital real-world data or actual physical world of robotics?

Ilya

我不认为这是非此即彼。模拟是一种工具,有特定的优势和劣势,我们应该使用它。

I don't think it's an either-or. Simulation is a tool, it has certain strengths and weaknesses, and we should use it.

Host

对自我对弈和强化学习的一个批评是,它们目前的结果虽然惊人,但都是在模拟或非常受限的物理环境中展示的。你认为有可能脱离模拟环境,在非模拟环境中学习吗?或者你认为有可能以照片级真实和物理真实的方式模拟真实世界,从而通过模拟中的自我对弈解决真实问题吗?

One of the criticisms of self-play and reinforcement learning is that its current results, while amazing, have been demonstrated in simulated or very constrained physical environments. Do you think it's possible to escape simulated environments and learn in non-simulated ones? Or do you think it's possible to simulate the real world in a photorealistic and physics-realistic way so we can solve real problems with self-play in simulation?

Ilya

我认为从模拟到真实世界的迁移绝对是可能的,并且已经被许多不同团队多次展示。在视觉领域尤其成功。此外,OpenAI 在夏天展示了一个机器人手,它完全在模拟中训练,以某种方式实现了从模拟到真实的迁移。

I think transfer from simulation to the real world is definitely possible and has been exhibited many times by many different groups. It's been especially successful in vision. Also, OpenAI in the summer demonstrated a robot hand which was trained entirely in simulation in a certain way that allowed for sim-to-real transfer to occur.

Host

这是为了魔方吗?

Is this for the Rubik's cube?

Ilya

没错。

That's right.

Host

我不知道那是在模拟中训练的。完全在模拟中训练?

I wasn't aware that was trained in simulation. It was trained in simulation entirely?

Ilya

真的吗?所以,它不在物理环境中?手没有训练?不,100% 的训练都在模拟中完成,模拟中学到的策略被训练得非常有适应性,以至于当你迁移时,它能很快适应物理世界。

Really? So what, it wasn't in the physical? The hand wasn't trained? No, 100% of the training was done in simulation, and the policy that was learned in simulation was trained to be very adaptive, so adaptive that when you transfer it, it could very quickly adapt to the physical world.

Host

那些用长颈鹿之类的干扰,不是模拟的一部分吗?

The kind of perturbations with the giraffe or whatever, those weren't part of the simulation?

Ilya

嗯,模拟通常被训练得对许多不同事物具有鲁棒性,但不是视频中那种干扰。所以它从未用手套训练过,也从未用毛绒长颈鹿训练过。所以理论上,这些是新颖的干扰。

Well, the simulation was generally trained to be robust to many different things, but not the kind of perturbations we had in the video. So it's never been trained with a glove, it's never been trained with a stuffed giraffe. So in theory, these are novel perturbations.

Host

不是理论上,实际上那些就是新颖的干扰。

It's not in theory, in practice those are novel perturbations.

Ilya

嗯,这是一个干净的小规模但清晰的从模拟世界到物理世界迁移的例子。我还要说,我预计深度学习的迁移能力总体上会增强。迁移能力越强,模拟就越有用,因为你可以先在模拟中体验某事,然后学到教训,再将其带到现实世界,就像人类玩电脑游戏时一直做的那样。

Well, that's a clean small-scale but clean example of transfer from the simulated world to the physical world. I will also say that I expect the transfer capabilities of deep learning to increase in general. The better the transfer capabilities, the more useful simulation will become, because then you can experience something in simulation and learn a moral of the story which you can carry with you to the real world, as humans do all the time when they play computer games.

智能测试 Embodiment and consciousness for AGI

Host

让我问一个具身的问题,继续谈 AGI。你认为 AGI 需要拥有身体、一些人类元素如自我意识、意识、对死亡的恐惧或物理空间中的自我保存吗?

Let me ask an embodied question, staying on AGI for a sec. Do you think AGI requires having a body, some human elements like self-awareness, consciousness, fear of mortality, or self-preservation in physical space?

Ilya

我认为拥有身体会有用。我不认为它是必要的,但肯定非常有用,因为你可以学到没有身体学不到的东西。但同时,我认为如果没有身体,你也可以补偿并仍然成功。

I think having a body will be useful. I don't think it's necessary, but I think it's very useful to have a body for sure, because you can learn things which cannot be learned without a body. But at the same time, I think that if you don't have a body, you could compensate for it and still succeed.

Host

你这么认为?

You think so?

Ilya

是的。嗯,有证据支持这一点。例如,有很多人天生聋盲,他们能够补偿缺失的模态。我特别想到海伦·凯勒。所以即使你无法与物理世界互动……

Yes. Well, there is evidence for this. For example, there are many people who were born deaf and blind, and they were able to compensate for the lack of modalities. I'm thinking about Helen Keller specifically. So even if you're not able to physically interact with the world...

Host

让我问一个更具体的问题——我不确定是否与拥有身体有关——但意识的概念,以及更受限的版本是自我意识。你认为 AGI 系统应该拥有意识吗?不管你认为意识是什么。

Let me ask on the more particular — I'm not sure if it's connected to having a body or not — but the idea of consciousness and a more constrained version of that is self-awareness. Do you think an AGI system should have consciousness? Whatever the heck you think consciousness is.

Ilya

很难回答,因为定义起来太难了。这绝对有趣、迷人。我认为我们的助手很可能会拥有意识。

Hard question to answer, given how hard it is to define. It's definitely interesting, fascinating. I think it's definitely possible that our assistants will be conscious.

Host

你认为这是从网络中存储的表征中涌现出来的吗?当你越来越能表征世界时,它自然就出现了?

Do you think that's an emergent thing that just comes from the representation stored within your networks? That it naturally emerges when you become more and more able to represent more of the world?

Ilya

我会提出以下论点:人类是有意识的,如果你相信人工神经网络与大脑足够相似,那么至少应该存在有意识的人工神经网络。

I'd make the following argument: humans are conscious, and if you believe that artificial neural nets are sufficiently similar to the brain, then there should at least exist artificial neural nets that are conscious too.

放弃对 AGI 的控制 Test of intelligence

Host

你非常依赖那个存在性证明。好吧,但这就是我能给出的最好答案。不,我知道,我知道。仍然有一个悬而未决的问题:大脑中是否有一些我们尚未发现的魔法——我不是指非唯物主义的魔法,而是大脑可能比我们想象的更复杂、更有趣。如果是这样,那它应该会显现出来,在某个时刻我们会发现无法继续取得进展。但我认为这不太可能。所以我们谈论意识,但让我谈谈另一个定义不清的概念:智能。我们讨论过推理,讨论过记忆。你认为什么是好的智能测试?图灵通过模仿游戏用自然语言提出的测试让你印象深刻吗?你心中有没有什么系统能做到的事情会让你深感震撼?

You're leaning on that existence proof pretty heavily. Okay, but it's just that that's the best answer I can give. No, I know, I know. There's still an open question if there's not some magic in the brain that we're not—I mean, I don't mean a non-materialistic magic, but that the brain might be a lot more complicated and interesting than we give it credit for. If that's the case, then it should show up, and at some point we will find out that we can't continue to make progress. But I think it's unlikely. So we talk about consciousness, but let me talk about another poorly defined concept: intelligence. Again, we've talked about reasoning, we've talked about memory. What do you think is a good test of intelligence for you? Are you impressed by the test that Alan Turing formulated with the imitation game, with natural language? Is there something in your mind that you will be deeply impressed by if a system was able to do?

Ilya

我的意思是,很多事情。今天的能力存在某些前沿,而前沿之外还有事物。任何这样的事物都会让我印象深刻。例如,一个深度学习系统能解决非常普通的任务,比如机器翻译或计算机视觉,并且在任何情况下都不会犯人类不会犯的错误。我认为这还没有被证明,我会觉得这非常了不起。

I mean, lots of things. There are certain frontiers of capabilities today, and there exist things outside of that frontier. I would be impressed by any such thing. For example, I would be impressed by a deep learning system which solves a very pedestrian task like machine translation or computer vision, and never makes a mistake a human wouldn't make under any circumstances. I think that is something which has not yet been demonstrated, and I would find it very impressive.

Host

所以现在它们会犯错,而且错误不同。它们可能比人类更准确,但仍然会犯一系列不同的错误。所以我猜测,有些人对深度学习的怀疑在于,当他们看到这些错误时会说:‘这些错误毫无道理。如果你理解了概念,就不会犯那种错误。’我认为改变这一点会激励我。那将是进步。

So right now they make mistakes, and they differ. They might be more accurate than human beings, but they still make a different set of mistakes. So my guess is that a lot of the skepticism some people have about deep learning is when they look at their mistakes and say, 'Well, those mistakes make no sense. If you understood the concept, you wouldn't make that mistake.' I think changing that would inspire me. That would be progress.

Ilya

是的,说得好。但我也很不喜欢人类那种批评模型不智能的本能。这和我们在批评任何其他生物群体时将其视为‘异类’的本能是一样的。因为很有可能 GPT-2 在很多方面比人类聪明得多。这绝对是事实。它有更广博的知识,甚至在某些话题上可能更有深度。很难判断深度意味着什么,但确实在某种意义上,人类不会犯这些模型犯的错误。是的,自动驾驶汽车也是如此。同样的情况可能还会继续发生在许多 AI 系统上。我们觉得这很烦人。这是 21 世纪的过程:分析 AI 的进步就是寻找一个系统以人类不会的方式大失败的情况,然后很多人写文章,公众普遍相信这个系统不智能。我们通过认为它不智能来安慰自己,只因为这一个轶事案例,而且这种情况似乎还在继续。

Yeah, that's a really nice way to put it. But I also just don't like that human instinct to criticize a model as not intelligent. That's the same instinct as when we criticize any group of creatures as 'the other.' Because it's very possible that GPT-2 is much smarter than human beings in many things. That's definitely true. It has a lot more breadth of knowledge, and even perhaps depth on certain topics. It's kind of hard to judge what depth means, but there's definitely a sense in which humans don't make mistakes that these models do. Yes, the same applies to autonomous vehicles. The same is probably going to continue being applied to a lot of AI systems. We find this annoying. This is the process of the 21st century: analyzing the progress of AI is the search for one case where the system fails in a big way where humans would not, and then many people writing articles about it, and then the public generally gets convinced that the system is not intelligent. We pacify ourselves by thinking it's not intelligent because of this one anecdotal case, and this seems to continue happening.

Host

是的,不过这话也有道理。我确信也有很多人对当今存在的系统印象深刻。但我认为这和我们之前讨论的观点有关:判断 AI 的进步本身就令人困惑。你知道,有一个新机器人演示了某样东西;你应该有多印象深刻?我认为一旦 AI 开始真正推动 GDP 增长,人们就会开始印象深刻。

Yeah, I mean there is truth to that though. There are people also, I'm sure, that plenty of people are also extremely impressed by the system that exists today. But I think this connects to the earlier point we discussed: that it's just confusing to judge progress in AI. And you know, you have a new robot demonstrating something; how impressed should you be? I think that people will start to be impressed once AI starts to really move the needle on the GDP.

Host

所以你是可能创造 AGI 系统的人之一。不是你,而是你和 OpenAI。如果你真的创造了一个 AGI 系统,并且能和它——他,她——共度一个晚上,你会聊什么?第一次。

So you're one of the people that might be able to create an AGI system here. Not you, but you and OpenAI. If you do create an AGI system and you get to spend the evening with it—him, her—what would you talk about? The very first time.

Ilya

嗯,第一次我会问各种各样的问题,试图让它犯错,然后我会惊讶于它不犯错,并继续问广泛的问题。

Well, the first time I would just ask all kinds of questions and try to get it to make a mistake, and I would be amazed that it doesn't make mistakes, and just keep asking broad questions.

Host

你觉得会是什么样的问题?会是事实性的,还是个人的、情感的、心理的?你怎么看?

What kind of questions do you think? Would they be factual, or would they be personal, emotional, psychological? What do you think?

Ilya

所有方面。你会寻求建议吗?当然。我的意思是,我为什么要限制自己和这样一个系统交谈呢?

All of that. Would you ask for advice? Definitely. I mean, why would I limit myself talking to a system like this?

Host

现在,再次强调,你确实是可能身临其境的人之一。所以让我问一个深刻的问题。我刚刚和一位斯大林历史学家聊过;我和很多研究权力的人聊过。亚伯拉罕·林肯说过:‘几乎所有的人都能忍受逆境,但如果你想测试一个人的品格,就给他权力。’我想说,21 世纪的力量,也许是 22 世纪,但希望是 21 世纪,将是创造 AGI 系统,以及直接拥有和控制 AGI 系统的人。那么你怎么看?在花了一个晚上与 AGI 系统讨论之后,你觉得你会做什么?

Now, again, let me emphasize the fact that you truly are one of the people that might be in the room where this happens. So let me ask a sort of profound question. I've just talked to a Stalin historian; I've been talking to a lot of people who are studying power. Abraham Lincoln said, 'Nearly all men can stand adversity, but if you want to test a man's character, give him power.' I would say the power of the 21st century, maybe the 22nd, but hopefully the 21st, would be the creation of an AGI system, and the people who have direct possession and control of the AGI system. So what do you think? After spending that evening having a discussion with the AGI system, what do you think you would do?

Ilya

嗯,我想象的理想世界是,人类就像公司的董事会成员,而 AGI 是 CEO。所以我喜欢这样的图景:有各种不同的实体——不同的国家或城市——居住在那里的人们投票决定代表他们的 AGI 应该做什么,然后代表他们的 AGI 去执行。我觉得这样的图景非常有吸引力。而且可以有多个 AGI:一个城市一个,一个国家一个。它实际上会试图将民主进程提升到一个新的水平。而董事会总是可以解雇 CEO,本质上按下重置按钮,说:‘在这里重新随机化参数。’

Well, the ideal world I would like to imagine is one where humanity are like the board members of a company where the AGI is the CEO. So I would like the picture where you have some kind of different entities—different countries or cities—and the people that live there vote for what the AGI that represents them should do, and then the AGI representing them goes and does it. I think a picture like that I find very appealing. And you could have multiple AGIs: one for a city, one for a country. And it would be trying to, in effect, take the democratic process to the next level. And the board can always fire the CEO, essentially press the reset button and say, 'Re-randomize the parameters here.'

Host

嗯,让我——这其实很好。这是一个美丽的愿景。我认为只要有可能按下重置按钮。你认为总是有可能按下重置按钮吗?

Well, let me sort of—that's actually okay. That's a beautiful vision. I think as long as it's possible to press the reset button. Do you think it will always be possible to press the reset button?

Ilya

所以我认为这绝对可以构建——我从你那里真正理解的问题是:人类或人们能否控制他们构建的 AI 系统?是的。我的答案是:绝对可以构建出愿意被人类控制的 AI 系统。哇,这是它们的一部分——所以它们不是不得不被控制,而是它们的存在——它们存在的目标之一就是被控制。就像人类父母通常想帮助他们的孩子,希望他们的孩子成功一样。这对他们来说不是负担;他们乐于帮助孩子,喂他们,给他们穿衣,照顾他们。我坚信对于 AGI 来说同样可能。可以编程一个 AGI,以这样的方式设计它。

So I think that it's definitely possible to build—so the question I really understand from you is: will humans or people have control over the AI systems that they built? Yes. And my answer is: it's definitely possible to build AI systems which will want to be controlled by their humans. Wow, that's part of their—so it's not that they can't help but be controlled, but that their existence—one of the objectives of their existence is to be controlled. In the same way that human parents generally want to help their children, they want their children to succeed. It's not a burden for them; they are excited to help the children, to feed them, to dress them, and to take care of them. And I believe with highest conviction that the same will be possible for an AGI. It will be possible to program an AGI, to design it in such a way.

艾伦·图灵的结束语 Relinquishing Power Over AGI

Host

它会有一种深刻的驱动力,并且乐于实现它,这个驱动力就是帮助人类繁荣。但让我退一步。在你创造 AGI 系统的那一刻,我认为这是一个非常关键的时刻。在那个时刻和由 AGI 领导的民主董事会成员之间,必须有一个权力的放弃。乔治·华盛顿说,尽管他做了很多坏事,但他做的一件大事就是放弃了权力。他首先不想当总统,即使当了总统,他也没有像大多数独裁者那样无限期地任职。你认为自己能够放弃对 AGI 系统的控制吗?考虑到你可以对世界拥有多大的权力?首先是财务上,赚很多钱,对吧?然后通过拥有 AGI 系统来控制。

Way that it will have a similar deep drive that it will be delighted to fulfill and the drive will be to help humans flourish. But let me take a step back. To that moment where you create the AGI system, I think this is a really crucial moment. And between that moment and the democratic board members with the AGI at the head, there has to be a relinquishing of power. Says George Washington, despite all the bad things he did, one of the big things he did is he relinquished power. He first of all didn't want to be president, and even when he became president, he didn't keep serving as most dictators do for indefinitely. Do you see yourself being able to relinquish control over an AGI system, given how much power you can have over the world? At first financial, just make a lot of money, right? And then control by having possession of an AGI system.

Ilya

我觉得这很容易做到。我觉得放弃这种……我是说,你知道,你描述的那种场景对我来说听起来很可怕,仅此而已。我绝对不想处于那种位置。

I'd find it trivial to do that. I'd find it trivial to relinquish this kind of... I mean, you know, the kind of scenario you are describing sounds terrifying to me, that's all. I would absolutely not want to be in that position.

Host

你认为你代表 AI 社区中的多数还是少数?

Do you think you represent the majority or the minority of people in the AI community?

Ilya

嗯,我的意思是,这是一个开放的问题,一个很重要的问题。大多数人是否善良?这是另一种问法。所以我不知道大多数人是否善良,但我认为在关键时刻,人们可以比我们想象的更好。

Well, I mean, open question, an important one. Are most people good? Is another way to ask it. So I don't know if most people are good, but I think that when it really counts, people can be better than we think.

Host

说得好。是的。你能想到将 AI 的价值观与人类价值观对齐的具体机制吗?这是……你在开发 AI 系统时考虑这些持续对齐的问题吗?

That's beautifully put. Yeah. Are there specific mechanisms you can think of of aligning AI's values to human values? Is that... do you think about these problems of continued alignment as we develop the AI systems?

Ilya

是的,当然。在某种意义上,你问的这类问题是……所以如果你必须把这个问题翻译成今天的术语,是的,这将是一个关于如何让一个强化学习智能体优化一个本身是学习得到的价值函数的问题。如果你看看人类,人类就是这样,因为人类的奖励函数、价值函数不是外部的,而是内部的。没错。有一些明确的想法关于如何训练一个价值函数,基本上是一个目标,你知道,一个尽可能客观的感知系统,它将单独训练以识别、内化人类对不同情境的判断,然后该组件将被整合为价值,作为更有能力的强化学习系统的基础价值函数。你可以想象这样一个过程。我不是说这就是过程,我是说这是一个你可以做的事情的例子。

Yeah, definitely. In some sense, the kind of question which you are asking is... so if you have to translate that question to today's terms, yes, it would be a question about how to get an RL agent that's optimizing a value function which itself is learned. And if you look at humans, humans are like that because the reward function, the value function of humans is not external, it is internal. That's right. And there are definite ideas of how to train a value function, basically an objective, you know, and as objective as possible perception system that will be trained separately to recognize, to internalize human judgments on different situations, and then that component would then be integrated as the value, as the base value function for some more capable RL system. You could imagine a process like this. I'm not saying this is the process, I'm saying this is an example of the kind of thing you could do.

Host

那么关于人类存在的目标函数,你认为人类存在中隐含的目标函数是什么?生命的意义是什么?

So on that topic of the objective functions of human existence, what do you think is the objective function that is implicit in human existence? What's the meaning of life?

Ilya

哦,我认为这个问题在某种程度上是错误的。我认为这个问题暗示存在一个客观答案,一个外部答案,你知道,你生命的意义是 X,对吧?我认为实际情况是,我们存在,这很神奇,我们应该努力充分利用它,并努力在我们存在的短暂时间里最大化我们自己的价值和享受。

Oh, I think the question is wrong in some way. I think that the question implies that the reason there is an objective answer, which is an external answer, you know, your meaning of life is X, right? I think what's going on is that we exist and that's amazing, and we should try to make the most of it and try to maximize our own value and enjoyment of a very short time while we do exist.

Host

有趣的是,行动确实需要一个目标函数。它肯定以某种形式存在,但很难明确表达,也许不可能明确表达,我想这就是你的意思。这是强化学习环境的一个有趣事实。

It's funny because action does require an objective function. It's definitely there in some form, but it's difficult to make it explicit and maybe impossible to make it explicit, I guess is what you're getting at. And that's an interesting fact of an RL environment.

Ilya

嗯,但我的观点略有不同,人类有欲望,他们的欲望创造了驱动力,导致他们……你知道,我们的欲望就是我们的目标函数,我们个人的目标函数。我们后来可以决定改变它,我们之前想要的不再好了,我们想要别的东西。是的,但它们非常动态。必须有某种潜在的弗洛伊德式的东西,比如性,有人认为是对死亡的恐惧,还有对知识的渴望,你知道所有这些事情。繁衍,所有那些进化论的观点。似乎可能存在某种基本的目标函数,其他一切由此产生。但这似乎……因为这非常重要。我认为那可能是一个进化的目标函数,即生存、繁衍并确保你的孩子成功。这是我的猜测。但这并没有回答生命的意义是什么的问题。我认为你可以看到人类是这个大过程、这个古老过程的一部分。我们……我们存在于一个小星球上,仅此而已。所以既然我们存在,就努力充分利用它,并尽可能多地享受,尽可能少地受苦。

Well, but I was making a slightly different point, is that humans want things, and their wants create the drives that cause them to... you know, our wants are our objective functions, our individual objective functions. We can later decide that we want to change that, what we wanted before is no longer good, and we want something else. Yeah, but they're so dynamic. There's got to be some underlying sort of Freud, there's things like sexual stuff, there's people who think it's the fear of death, and there's also the desire for knowledge, and you know all these kinds of things. Procreation, the sort of all the evolutionary arguments. It seems to be there might be some kind of fundamental objective function from which everything else emerges. But it seems... because that's very important. I think that probably is an evolutionary objective function, which is to survive and procreate and make sure you make your children succeed. That would be my guess. But it doesn't give an answer to the question what's the meaning of life. I think you can see how humans are part of this big process, this ancient process. We are... we exist on a small planet and that's it. So given that we exist, try to make the most of it and try to enjoy more and suffer less as much as we can.

Host

让我问两个关于生活的傻问题。第一:你有遗憾吗,那些如果能重来你会做得不同的时刻?第二:有没有你特别自豪、让你真正快乐的时刻?

Let me ask two silly questions about life. One: do you have regrets, moments that if you went back you would do differently? And two: are there moments that you're especially proud of that made you truly happy?

Ilya

所以我可以回答这个问题。两个问题我都可以回答。当然有。我做了大量的选择和决定,事后看来我不会那样做,我确实感到一些遗憾。但你知道,我试图安慰自己,当时我尽了最大努力。至于我自豪的事情,有……我很幸运有一些我自豪做过的事情,我为之自豪,它们让我快乐了一段时间。但我不认为那是幸福的来源。

So I can answer that. I can answer both questions. Of course there are. There's a huge number of choices and decisions that I've made that with the benefit of hindsight I wouldn't have made them, and I do experience some regret. But you know, I try to take solace in the knowledge that at the time I did the best I could. And in terms of things that I'm proud of, there are... I'm very fortunate to have things I'm proud to have done, things I'm proud of, and they made me happy for some time. But I don't think that that is the source of happiness.

Host

那么你的学术成就,所有的论文,你是世界上被引用最多的人之一,我提到的所有在计算机视觉和语言等方面的突破。那是你幸福和自豪的来源吗?

So your academic accomplishments, all the papers, you're one of the most cited people in the world, all the breakthroughs I mentioned in computer vision and language and so on. Is that the source of happiness and pride for you?

Ilya

我的意思是,所有这些事情当然是自豪的来源。我非常感激做了所有这些事情,做这些事情也很有趣。但幸福来自……你知道,你可以……幸福,嗯,我目前的观点是,幸福在很大程度上来自于我们看待事物的方式。你知道,你可以吃一顿简单的饭而因此感到快乐,或者你可以和某人交谈而因此感到快乐。或者相反,你可以吃一顿饭,却因为饭菜不够好而失望。所以我认为很多幸福来自于此。但我不确定。我不想太自信。在不确定性面前保持谦逊似乎也是这整个幸福事情的一部分。

I mean, all those things are a source of pride for sure. I'm very grateful for having done all those things, and it was very fun to do them. But happiness comes from... you know, you can... happiness, well, my current view is that happiness comes from, to a very large degree, from the way we look at things. You know, you can have a simple meal and be quite happy as a result, or you can talk to someone and be happy as a result as well. Or conversely, you can have a meal and be disappointed that the meal wasn't a better meal. So I think a lot of happiness comes from that. But I'm not sure. I don't want to be too confident. Being humble in the face of the uncertainty seems to be also a part of this whole happiness thing.

Host

嗯,我认为没有比生命意义和幸福讨论更好的结束方式了。伊利亚,非常感谢你。你给了我一些不可思议的想法。你给了世界很多不可思议的想法。我真的很感激,谢谢今天的谈话。

Well, I don't think there's a better way to end it than meaning of life and discussions of happiness. Ilya, thank you so much. You've given me a few incredible ideas. You've given the world many incredible ideas. I really appreciate it, and thanks for talking today.

Ilya

是的,谢谢你来。我真的很享受。

Yeah, thanks for stopping by. I really enjoyed it.

Host

感谢收听与伊利亚·苏茨克维的对话。感谢我们的赞助商 Cash App。请考虑通过下载 Cash App 并使用代码 Lex Podcast 来支持本播客。如果你喜欢这个播客,请在 YouTube 上订阅,在 Apple Podcast 上给予五星评价,在 Patreon 上支持,或者直接在 Twitter 上联系我@lexfridman。现在让我……

Thanks for listening to this conversation with Ilya Sutskever. And thank you to our presenting sponsor, Cash App. Please consider supporting the podcast by downloading Cash App and using code Lex Podcast. If you enjoy this podcast, subscribe on YouTube, review it with 5 stars on Apple Podcast, support on Patreon, or simply connect with me on Twitter at @lexfridman. And now let me...

Closing Quote from Alan Turing Closing Quote from Alan Turing

Host

最后,用艾伦·图灵关于机器学习的一段话作为结束:“与其试图编写一个模拟成人思维的程序,为什么不尝试编写一个模拟儿童思维的程序呢?如果随后对它进行适当的教育,就能得到成人的大脑。”感谢收听,期待下次再见。

I'll leave you with some words from Alan Turing on machine learning: "Instead of trying to produce a program to simulate the adult mind, why not rather try to produce one which simulates the child's? If this were then subjected to an appropriate course of education, one would obtain the adult brain." Thank you for listening, and hope to see you next time.

互动版:逐字朗读 + 针对本期提问 →