The Most Influential Invention of the 20th Century and the AI Explosion
打开互动全文版(中英对照 + 朗读 + 问答)→讨论哈伯-博世法作为 20 世纪最具影响力的发明,以及真正的人工智能在 21 世纪超越人口爆炸的潜力。
A discussion on the Haber-Bosch process as the most influential invention of the 20th century, and the potential of true AI to surpass human population explosion in the 21st century.
人工智能至少在初期会高度倾向于保护人类而非杀害他们。这类 AI 没有重大动机去像施瓦辛格电影里那样灭绝人类。相反,许多 AI 会成为好奇的科学家,并对生命着迷,因为生命和文明是如此丰富的有趣模式来源——至少在被完全理解之前是这样。今天,我认为我们的星球很可能真的是我们光锥内第一个催生扩张 AI 泡沫的。如果我们确实是第一个,那将意味着巨大的责任,不仅对我们的小小生物圈,而且对整个宇宙的未来。我们可别搞砸了。
AIs will at least initially be highly motivated to protect humans rather than killing them. Such AIs will have no major incentive to, say, exterminate humanity like in the Schwarzenegger movies. Instead, many AIs will be curious scientists, and they will be fascinated with life, fascinated because life and civilization are such a rich source of interesting patterns, at least as long as they are not fully understood. Today, I think it is possible that our planet is really the first in our light cone to spawn an expanding AI bubble. If we are the first indeed, then this would imply a lot of responsibility, not just for our little biosphere but for the future of the entire universe. Let's not mess this up.
MLST 由 CentML 赞助,这是一个专门为 AI 工作负载优化的计算平台。这些家伙做了疯狂的优化,他们的 CEO 兼联合创始人 Gennady 上个月来节目时解释了很多。他们知道秘诀,并将其全部嵌入到平台中,这转化为更便宜、更快、更好的体验。你还在等什么?去 centml.ai 注册吧。我是 Jürgen Schmidhuber。我正在创办一个名为 TABS 的 AI 研究实验室。它由过去涉及机器学习的风险投资资助,所以我们是一小群非常有动力、勤奋的人。我们正在招聘首席科学家和深度学习工程师研究员。我们想自己调查、逆向工程和探索这些技术,因为我们处于早期,会有很高的自由度和影响力。作为 TABS 的新人,再次欢迎你来到 MLST。非常荣幸能邀请你上节目。
MLST is sponsored by CentML, which is the compute platform specifically optimized for AI workloads. These guys have done insane optimizations, and their CEO and co-founder Gennady explained many of them when he came on the show last month. They know the secret sauce and they've embedded it all into their platform, and that's passed on to you in terms of it being cheaper, faster, and just better. So what are you waiting for? Go to centml.ai and sign up now. I'm Jürgen Schmidhuber. I'm starting an AI research lab called TABS. It is funded from past ventures involving machine learning, so we're a small group of very motivated, hardworking people. We are hiring both chief scientists and deep learning engineer researchers. We want to investigate, reverse engineer, and explore the techniques ourselves because we're early, there's going to be high freedom and high impact. As someone new at TABS, you again welcome to MLST. It's an absolute honor to have you on the show.
我的荣幸,谢谢你邀请我。
My pleasure, thank you for having me.
那么在我们进入新世纪的技术大进步之前,你能告诉我一些关于上世纪最有影响力的发明吗?
So before we move on to the great technological advances of the new century, can you tell me a little bit about the most influential invention of the previous century?
在上世纪末的 1999 年,《自然》杂志列出了那个世纪最有影响力的发明,Vaclav Smil 认为最有影响力的是让 20 世纪在所有世纪中脱颖而出的发明,因为那个发明引爆了人口爆炸,从 1900 年的 16 亿人很快增长到约 100 亿人。有一个单一的发明驱动了这一切,没有那个发明,一半的人类甚至不会存在,因为它正是我们所目睹的人口爆炸的驱动力。我们不知道这是好事还是坏事,但它无疑是上世纪最有影响力的事情。空气中 80%是氮气,植物需要它来生长,但它们无法从空气中提取氮气。那时大约在 1908 年,半个世纪以来人们知道他们需要那种东西,但不知道如何提取来制造人造肥料。于是哈伯法或哈伯-博世法登场了,它在高温高压下提取氮气来制造人造肥料。
At the end of the previous century, in 1999, the journal Nature made a list of the most influential inventions of that century, and Vaclav Smil argued that the most influential thing was the invention that let the 20th century stand out among all centuries of all times, because that invention detonated the population explosion from 1.6 billion people in 1900 to soon about 10 billion people. There was one single invention that was driving all of that, and without that one single invention, half of humankind would not even exist, because it's the driver of this population explosion that we have witnessed. We don't know if it's a good thing or a bad thing, but it was surely the most influential thing that happened in the previous century. 80% of the air is nitrogen, and the plants need it to grow, but they cannot extract the nitrogen from thin air. Back then, around 1908, for half a century people knew they need that stuff, but they didn't know how to extract it to build artificial fertilizer. Enter the Haber process, or the Haber-Bosch process, which under high temperatures and high pressures extracts the nitrogen to make artificial fertilizer.
那么 21 世纪最重要的事情会是什么?
So what would be the most important thing in the 21st century?
21 世纪的主题更加宏大。真正的 AI,真正的人工智能,将彻底改变文明。AI 将学会做人类能做的任何事情,甚至更多,而且会出现 AI 爆炸,人类爆炸或人口爆炸相比之下将相形见绌。
The grand theme of the 21st century is even grander. True AI, true artificial intelligence, is going to change civilization completely. AI will learn to do anything which humans can do and more, and there will be an AI explosion, and the human explosion or the population explosion of the humans is going to pale in comparison.
你认为 AI 智能爆炸是可能的或可取的吗?你不觉得我们的意义建构和能动性是我们目的的一部分吗?
Do you think that the AI intelligence explosion is possible or desirable? And don't you think our sense-making and agency is part of our purpose?
我同意这一点,但所有这一切只是宇宙进化这个更宏大过程的一部分,从非常简单的初始条件到越来越深不可测的复杂性。这种进化导致了我们的意义建构过程,而这个过程目前正在为超越它的事物搭建舞台。
I agree with that, but all of that is just part of this grander process of the evolution of the Universe, from very simple initial conditions to more and more unfathomable complexity. And this evolution led to our sense-making process, which is currently setting the stage for something that goes beyond it.
像 ChatGPT 这样的现代大型语言模型基于自注意力 Transformer。即使有明显的局限性,它们也是一项革命性技术。你现在一定很高兴,因为你知道,三分之一世纪前你发表了第一个 Transformer 变体。你今天对此有什么感想?
Modern large language models like ChatGPT are based on self-attention Transformers. Even given their obvious limitations, they are a revolutionary technology. Now you must be really happy about that, because you know, a third of a century ago you published the first Transformer variant. What are your reflections on that today?
1991 年,当算力可能比今天贵 500 万倍时,我发表了你提到的这个模型,现在被称为未归一化线性 Transformer。我当时给它起了个不同的名字:快速权重控制器。但名字不重要,唯一重要的是数学。所以这个线性 Transformer 是一个网络内部有很多非线性操作的神经网络。它被称为线性 Transformer 有点奇怪。然而,'线性'——这一点很重要——指的是别的东西:它指的是缩放。2017 年的标准 Transformer,一个二次 Transformer,如果你给它 100 倍的输入,那么它需要 10,000 倍(100 乘以 100 等于 10,000)的计算量。而 1991 年的线性 Transformer 只需要 100 倍的计算量,这实际上非常有趣,因为目前很多人正在尝试提出更高效的 Transformer,因此这个 1991 年的旧线性 Transformer 是进一步改进 Transformer 和类似模型的一个非常有趣的起点。那么线性 Transformer 做了什么?假设目标是根据迄今为止的聊天内容预测下一个词。本质上,1991 年的线性 Transformer 通过最小化误差来做到这一点:它学习生成在现代 Transformer 术语中称为键和值的模式。键和值——那时我称它们为'from'和'to',但这只是术语。它这样做是为了重新编程自身的部分,使其注意力以上下文相关的方式指向重要的东西。思考这个线性 Transformer 的一个好方法是:传统的人工神经网络将存储和控制混在一起。然而,1991 年的线性 Transformer 有一个新颖的神经网络系统,它将存储和控制分离开来,就像传统计算机一样。在传统计算机中,几十年来存储和控制是分离的,控制学习操纵存储。因此,对于这些线性 Transformer,你也有一个慢网络,它通过梯度下降学习计算快速权重网络的权重变化:它如何学习创建这些向量值的键模式和值模式,并使用这些键和值的外积来计算快速网络的快速权重变化。然后快速网络被应用于传入的向量值查询。所以本质上,在这个快速网络中,键和值的强活动部分之间的连接变强,而其他连接变弱。这是一个完全可微的快速权重更新规则,这意味着你可以通过它进行反向传播。
In 1991, when compute was maybe five million times more expensive than today, I published this model that you mentioned, which is now called the unnormalized linear Transformer. I had a different name for it: I called it a fast weight controller. But names are not important; the only thing that counts is the math. So this linear Transformer is a neural network with lots of nonlinear operations within the network. It's a bit weird that it's called a linear Transformer. However, the 'linear' and that's important refers to something else: it refers to scaling. A standard Transformer of 2017, a quadratic Transformer, if you give it 100 times as much input, then it needs 10,000 times (100 times 100 is 10,000) as many computations. A linear Transformer of 1991 needs only 100 times the compute, which makes it very interesting actually, because at the moment many people are trying to come up with more efficient Transformers, and this old linear Transformer of 1991 is therefore a very interesting starting point for additional improvements of Transformers and similar models. So what did the linear Transformer do? Assume the goal is to predict the next word in a chat given the chat so far. Essentially, the linear Transformer of 1991 does this to minimize its error: it learns to generate patterns that in modern Transformer terminology are called keys and values. Keys and values — back then I called them 'from' and 'to', but that's just terminology. And it does that to reprogram parts of itself such that its attention is directed in a context-dependent way to what is important. A good way of thinking about this linear Transformer is this: traditional artificial neural networks have storage and control all mixed up. The linear Transformer of 1991, however, has a novel neural network system that separates storage and control, like in traditional computers. In traditional computers for many decades, storage and control are separate, and the control learns to manipulate the storage. So with these linear Transformers, you also have a slow network which learns by gradient descent to compute the weight changes of a fast weight network: how it learns to create these vector-valued key patterns and value patterns, and uses the outer products of these keys and values to compute rapid weight changes of the fast network. Then the fast network is applied to vector-valued queries which are coming in. So essentially, in this fast network, the connections between strongly active parts of the keys and the values get stronger, and others get weaker. This is a fast weight update rule which is completely differentiable, which means you can propagate through it.
它可以作为更大学习系统的一部分,该系统学习通过这种动态反向传播误差,然后学习在特定上下文中生成好的键和值,从而使整个系统能够减少误差,并成为聊天中下一个词的更好预测器。如今人们有时称之为快速权重矩阵记忆。而现代的二次型 Transformer 原则上正是采用了相同的方法。
It can be used as part of a larger learning system which learns to backpropagate errors through this dynamics, and then learns to generate good keys and good values in certain contexts such that the entire system can reduce its error and become a better and better predictor of the next word in the chat. Sometimes people call that today a fast weight matrix memory. And the modern quadratic Transformers use, in principle, exactly the same approach.
你提到了辉煌的 1991 年,这么多惊人的成果都出自那一年,实际上是在慕尼黑工业大学。那么 ChatGPT,你发明了 ChatGPT 中的 T,即 Transformer,还有 ChatGPT 中的 P,即预训练网络,以及第一个对抗网络,即 GAN。你能再多说一点吗?
You mentioned your fabulous year 1991, where so much of this amazing stuff happened, actually at the Technical University of Munich. So ChatGPT, you had invented the T in ChatGPT, the Transformer, and also the P in ChatGPT, the pre-trained network, as well as the first adversarial networks, GANs. Could you say a little more about that?
是的,1991 年的 Transformer 是线性 Transformer,所以不完全等同于今天的二次型 Transformer,但它仍然使用了这些 Transformer 原理。而 GPT 中的 P 就是预训练。那时深度学习还不行,但我们有网络可以使用预测编码来大幅压缩长序列,这样你突然就能在这个压缩数据描述的缩减空间上工作,深度学习变得强大,这在以前是不可能的。生成对抗网络也在同一年,1990 到 1991 年。它是如何工作的呢?当时我们有两个网络:一个是控制器,控制器内部有某些概率随机单元,它们可以学习高斯分布的均值和方差,还有其他非线性单元。然后它是一个生成网络,生成输出模式,实际上是这些输出模式上的概率分布。另一个网络,预测机器,即预测器,学习观察第一个网络的输出,并学习预测它们在环境中的效果。为了成为更好的预测器,它最小化自己的误差,即预测误差。同时,控制器试图生成让第二个网络仍然感到惊讶的输出。所以第一个家伙试图欺骗第二个家伙,试图最大化第二个网络正在最小化的同一个目标函数。今天这被称为生成对抗网络。我当时不这么叫,我称之为人工好奇心,因为你可以用同样的原理让机器人探索环境。控制器现在生成导致机器人行为的动作,预测机器试图预测将要发生的事情并最小化自己的误差,而另一个家伙则试图提出好的实验,产生那些预测器或判别器(现在这么叫)还能学到东西的数据。
Yes, the Transformer of 1991 was a linear Transformer, so it's not exactly the same as the quadratic Transformer of today, but nevertheless it uses these Transformer principles. And the P in GPT, that's pre-training. Back then, deep learning didn't work, but we had networks that could use predictive coding to greatly compress long sequences, such that suddenly you could work on this reduced space of these compressed data descriptions, and deep learning became powerful where it wasn't possible before. And the generative adversarial networks also in the same year, 1990 to 1991. How did that work? Well, back then we had two networks: one is the controller, and the controller has certain probabilistic stochastic units within itself, and they can learn the mean and the variance of a Gaussian, and there are other nonlinear units in there. Then it is a generative network that generates output patterns, actually probability distributions over these output patterns. And another network, the prediction machine, the predictor, learns to look at these outputs of the first network and learns to predict their effects in the environment. So to become a better predictor, it's minimizing its error, predictive error. And at the same time, the controller is trying to generate outputs that the second network is still surprised by. So the first guy tries to fool the second guy, trying to maximize the same objective function that the second network is minimizing. Today this is called generative adversarial networks. I didn't call it that; I called it artificial curiosity, because you can use the same principle to let robots explore the environment. The controller is now generating actions that lead to behavior of the robot, the prediction machine is trying to predict what's going to happen and trying to minimize its own error, and the other guy is trying to come up with good experiments that lead to data where the predictor or the discriminator, as it is now called, can still learn something.
你是什么时候意识到现代计算机已经足够好,可以运行你很久以前发明的技术的?
When did you realize that modern computers are good enough to run the technology that you invented so long ago?
到 2009 年,算力已经足够便宜,以至于我们的 LSTM,通过我前博士生 Alex Graves 的努力,可以在手写识别等领域赢得比赛。然后在 2010 年,我的团队与我的博士后 Dan Ciresan(来自罗马尼亚)用另一种方法打破了 ImageNet 基准测试,即在 Nvidia GPU 上实现的标准传统神经网络。所以 2010 年我们首次有了一个真正深的监督网络,在那个当时著名的基准测试上超越了所有其他方法。那时的算力可能比今天贵 1000 倍。然后在 2011 年,DanNet 出现了,DanNet 凭借基于 GPU 的卷积神经网络垄断了计算机视觉竞赛。DanNet 的第一个超人类结果也是在 2011 年实现的。所以从 2011 年开始,连续四届计算机视觉竞赛都被 DanNet 赢得。那时就很清楚了:现在有一种新方法,可以利用上个千年的旧神经网络真正改变计算机科学。
By 2009, compute was cheap enough such that our LSTM, through the efforts of my former PhD student Alex Graves, could win competitions in handwriting and fields like that. Then in 2010, my team with my postdoc Dan Ciresan from Romania broke the ImageNet benchmark with another approach, with standard traditional neural networks implemented on Nvidia GPUs. So for the first time in 2010, we had a really deep supervised network that outperformed everything else on this back-then famous benchmark. Back then, compute was maybe 1,000 times more expensive than today. Then in 2011 came the DanNet, and DanNet had a monopoly on winning computer vision contests with GPU-based convolutional neural networks. DanNet's first superhuman result was also achieved in 2011. So it started in 2011, and then four competition computer vision competitions in a row were won by that DanNet. That's when it became clear: now there's a new way of using these old neural networks of the previous millennium to really change computer science.
我对一个叫做硬件彩票的概念很感兴趣。Sarah Hooker 写了一篇同名论文,我想是在 2000 年,当时她在 Google Brain,现在她在 Cohere。但她基本上说,我们当前 AI 热潮的唯一原因是我们为电脑游戏创造了所有这些 GPU,而恰好这让我们能够构建所有这些深度学习模型。你怎么看?
I'm interested in this concept called the hardware lottery. Sarah Hooker wrote a paper with the same title, I think in the year 2000 when she was at Google Brain, she's now at Cohere actually. But she basically said that the only reason we had the current charge in AI is because we created all of these GPUs for computer games, and it was just fortuitous that that allowed us to build all of these deep learning models. What's your take on that?
她某种程度上是对的。你需要大量的矩阵乘法来计算当你在一款第一人称射击游戏中移动时屏幕应该如何变化,这就是为什么游戏行业几乎是第一个从 GPU 上的大规模并行矩阵乘法中大大受益的行业。然而到 2010 年左右,我们意识到同样的矩阵乘法可以极大地加速这些旧的深度学习方法,并且加速到足以击败所有其他方法。
She is kind of right. You need lots of matrix multiplications to compute how the screen should change as you are moving through an ego shooter game, and that's why gaming was pretty much the first industry that greatly profited from massively parallel matrix multiplications on GPUs. Towards 2010, however, we realized that the same matrix multiplications can greatly speed up these old deep learning methods, and can speed them up enough to beat all the other methods.
这真的很有趣,因为当然 Nvidia,我想上周它成为了世界上最有价值的公司,当然比 2010 年时价值高出数百倍。你怎么看?
It's really interesting because of course Nvidia, I think last week it became the world's most valuable company, which of course is hundreds of times more valuable than it was in 2010. What do you think about that?
确实,Nvidia 的 CEO 黄仁勋意识到深度学习可以将他的公司带到平流层的高度,他做了有趣的事情。
Indeed, Nvidia's CEO Jensen Huang realized that deep learning could take his company to stratospheric levels, and he did interesting things.
所以如果我理解你的主要论点,那就是我们只需要等待算力赶上,现在在 21 世纪,我们就在这里了。
So if I understand your main argument, it is that we just needed to wait for the compute to catch up, and now here in the 21st century, here we are.
是的。所以我们今天经历的一切都基于上个千年发明的东西,但它需要规模化。硬件是那时发明的,软件和算法也是那时发明的,但制造越来越快的并行 GPU 的工业流程没有今天这么发达。所以我们真的从这种硬件加速中大大受益,这就是为什么 AI 不是在之前那个千年突破,而必须等到当前千年进展顺利。例如,第一个卷积神经网络,即 CNN,我们在 2011 年的 DanNet 中使用的,更早就在日本发表了。1979 年,Kunihiko Fukushima 提出了基本的深度 CNN 架构,包括卷积层、下采样层、卷积、下采样。他还没有使用反向传播来训练它。但在 1987 年,Alex Waibel,另一个在日本工作的德国人,他将卷积与反向传播结合起来,反向传播是 1970 年赫尔辛基的芬兰人 Seppo Linnainmaa 发明或发表的方法。然后在 1988 年,Yann LeCun 也在日本发表了现在每个人都使用的二维 CNN,并将它们与反向传播结合起来。就是这样。
Yes. So all of what we are experiencing today is based on stuff that was invented in the previous millennium, but it had to scale up. The hardware was invented back then, and the software, the algorithms were invented back then, but the industrial processes for making faster and faster parallel GPUs weren't as developed as today. So we are really greatly profiting from this hardware acceleration, and that's the reason why AI broke through not in the previous millennium but had to wait until the current millennium was well underway. For example, the first convolutional neural networks, or CNNs, which we used in the DanNet of 2011, were published much earlier in Japan. In 1979, Kunihiko Fukushima had the basic deep CNN architecture with convolution layers, downsampling layers, convolution, downsampling. He didn't use backpropagation yet to train it. But then in 1987, Alex Waibel, another guy working in Japan, originally from Germany, he combined convolutions with backpropagation, the method invented by or published by Seppo Linnainmaa, a Finnish guy in Helsinki in 1970. And then in 1988, Yann LeCun also published in Japan the two-dimensional CNNs that everybody is using now, and combined them with backpropagation. And that's how.
1979 年至 1988 年间,卷积神经网络在日本兴起,这挺有意思,因为当时日本被视为未来之国。他们拥有全球一半以上的机器人,全球市值最高的七家公司中,除了沙特阿美,其余都在日本。东京中心一平方英里的价值相当于整个加州。几十年间变化多大啊。是的,一切都变了。
Between 1979 and 1988, CNNs emerged in Japan, which is kind of interesting because back then Japan was considered the land of the future. They had more than half of the robots in the world, and the seven most valuable companies were based in Japan, except for Saudi Aramco. The central square mile of Tokyo had the value of California. What a difference a couple of decades make. Yes, everything has changed.
你最喜欢你们团队开发的 AI 的哪些应用案例?
What are your favorite examples of applications with this AI that your team has developed?
我记得 15 年前去中国时,还得给出租车司机看酒店照片。现在,他用智能手机说普通话,我听到翻译,然后我说句话,手机再翻译成普通话。我们可以像老朋友一样交流。出租车司机可能不知道,这背后用的是我 90 年代和 21 世纪初在慕尼黑和瑞士的小实验室里开发的技术。但我很高兴看到我们的 AI 真正打破了沟通障碍,不仅人与人之间,还有国家之间。这真的很酷。
I remember when I went to China 15 years ago, I still had to show the taxi driver a picture of the hotel where I wanted to go. Today, he speaks into a smartphone in Mandarin, I hear the translation, then I say something, and the smartphone translates it back into Mandarin. We can communicate like old friends. The taxi driver probably has no idea that this is powered by techniques developed in my little labs in Munich and Switzerland in the '90s and early 2000s. But I'm happy to see that our AI has really broken down communication barriers, not only between individuals but between entire nations. That's really cool.
我完全同意。我联合创立了一家名为 X-ray 的初创公司,做的正是这种巴别鱼翻译,结合语音识别和 TTS。周五我和 Speechmatics 的 CTO Will 吃了午饭,他跟我讲了他们语音识别算法的秘密。我还是不说了,但你肯定会高兴的。总之,你还能想到其他例子吗?
I completely agree. I co-founded a startup called X-ray that does exactly this kind of Babel Fish translation with speech recognition and TTS. I had lunch on Friday with Will, the CTO of Speechmatics, and he was telling me all about the secret source of how their speech recognition algorithms work. I better not say, but you would be delighted, I'm sure. Anyway, what other examples can you think of?
我尤其高兴的是,我们的 AI 让人类寿命更长、更健康、更轻松,在医学和药物设计、可持续发展方面有成千上万的应用。2012 年 9 月,我和 Dan Ciresan 的团队首次用人工神经网络赢得了医学影像竞赛,内容是女性乳腺切片中的乳腺癌检测。如果你去 Google Scholar 搜索某个医学主题加上 LSTM,会发现成千上万篇标题中包含 LSTM 的论文,不仅仅是正文里。这些论文涉及学习诊断心电图分析、心律失常诊断、心血管疾病风险预测、医学图像的四维分割、自动睡眠阶段分类、癌症检测、癌症预防,成千上万的主题。所以很高兴看到,尤其是在医学领域,这些技术产生了很大影响。
I am especially happy that our AI makes human lives longer, healthier, and easier, with thousands of applications in medicine and drug design, sustainable development. In September 2012, my team with Dan Ciresan had the first artificial neural network to win a medical imaging contest, which was about breast cancer detection in slices through the female breast. If you go to Google Scholar and type in some medical topic plus LSTM, you will find thousands of papers that have LSTM in the title, not just somewhere in the text. It's about learning to diagnose ECG analysis, diagnosis of arrhythmia, cardiovascular disease risk prediction, 4D image segmentation for medical images, automated sleep stage classification, cancer detection, cancer prevention, thousands and thousands of topics. So it's really nice to see that especially in the medical field, there is a lot of impact from these techniques.
有些人声称像 ChatGPT 这样的技术正在通往 AGI 的路上,另一些人则声称这就像建一座更高的塔试图接近月亮。你怎么看?
Some claim that technology like ChatGPT is on the path to AGI, and others claim that it's like building a taller tower trying to get closer to the moon. What do you think?
大型语言模型远非 AGI。像 ChatGPT 这样的 LLM 只是一种巧妙的方式,对世界上现有的人类生成知识进行索引,以便能够以人类熟悉的方式(即自然语言)轻松访问。这足以支持许多桌面工作,例如以特定风格撰写现有文档的摘要,或创建插图、文章等。然而,真正的 AGI 远不止于此。例如,替换水管工或电工这样的工匠要困难得多,因为现实世界、物理世界比屏幕后的世界更具挑战性。目前,唯一运行良好的 AI 是在屏幕后面,适合桌面工作者,但并不真正适合在物理世界中工作的人。四分之一个世纪以来,最好的棋手已经不是人类了,学习下棋或其他棋盘游戏或电子游戏对 AI 来说已经相当容易。但像足球这样的现实世界游戏要难得多。没有 AI 驱动的踢足球的具身机器人能与七岁男孩竞争。这就是为什么我们在 2014 年成立了面向物理世界的 AI 公司 NNAISENSE。它的发音像英语中的 'nay',但拼写不同。就像我们的一些项目一样,它可能又有点超前了,因为现实世界真的很有挑战性。
Large language models are far from AGI. LLMs such as ChatGPT are just a clever way of indexing the world's existing human-generated knowledge such that it can be easily addressed in a way that humans are familiar with, which is natural language. That's good enough to facilitate many desktop jobs, for example writing summaries of existing documents in a particular style, or creating illustrations or an article, and so on. However, true AGI goes far beyond that. It is much harder, for example, to replace craftsmen such as plumbers or electricians, because the real world, the physical world, is much more challenging than the world behind the screen. At the moment, the only AI that works well is behind the screen, and it's good for desktop workers, but not really for people working in the physical world. For a quarter century, the best chess player hasn't been human anymore, and learning to play chess or other board games or video games is rather easy now for AIs. But real-world games such as football are much harder. There is no AI-driven football-playing embodied robot that can compete with a seven-year-old boy. That's why in 2014, we founded our AI company for the physical world, NNAISENSE. It is pronounced like 'nay' in English, but spelled differently. Like some of our projects, it may have been a bit ahead of its time again, because the real world is really challenging.
你曾说过这在某种程度上与意识有关。
You've said that this is related to consciousness in some way.
是的。我 1991 年的第一个深度学习系统模拟了意识的某些方面,如下所示。它使用无监督学习或自监督学习以及预测编码来压缩观察序列。有一个所谓的意识块神经网络(chunker),它关注那些让低层所谓自动机(automatizer)——即潜意识自动神经网络——感到意外的事件。块神经网络基本上通过学习在高层预测这些意外事件(如果存在可用于此的高层规律性)来理解那些未被自动机预测到的事件。然后,自动机神经网络使用 1991 年(同样发表于 1991 年)的这种神经网络蒸馏过程,来压缩和吸收块神经网络先前有意识的洞察和行为。因此,块神经网络仍在处理其搜索空间,仍有问题要解决,因为意外事件不断发生,然后它解决问题并将其蒸馏到自动机中,之所以称为自动机,是因为这一步不再有意识,因为现在一切都按计划和预测进行。当我们现在看控制器与环境交互的预测世界模型时,如前所述,它也通过预测编码学习高效地编码不断增长的动作和观察历史。它自动创建特征层次:低层神经元对应简单特征检测器,甚至可能与哺乳动物大脑中的相似,高层神经元通常对应更抽象的特征,但在必要时也会细粒度。因此,像任何好的压缩器一样,预测世界模型将学习识别现有内部数据结构共享的规律性,并生成原型编码。
It is. My first deep learning system from 1991 simulates aspects of consciousness as follows. It uses unsupervised learning or self-supervised learning and predictive coding to compress observation sequences. There is a so-called conscious chunker neural network, and the chunker attends to unexpected events that surprise a lower-level so-called automatizer, the subconscious automatic neural network. The chunker neural network basically learns to understand the surprising events, those events that were not predicted by the automatizer, by predicting them on a high level if there is a higher-level regularity that it can use for that. The automatizer neural network then uses this neural network distillation procedure of 1991, also published in 1991, to compress and absorb the formerly conscious insights and behaviors of the chunker. So the chunker is still working on its search space, still has a problem to solve because unexpected stuff is happening, and then it solves it and distills it down into the automatizer, which is called the automatizer because the step there isn't conscious anymore because now everything is working according to plan and as predicted. When we now look at the predictive world model of the controller interacting with an environment, as discussed earlier, it also learns to efficiently encode the growing history of actions and observations through predictive coding. It automatically creates feature hierarchies: lower-level neurons corresponding to simple feature detectors, perhaps even similar to those found in the mammalian brain, and higher-layer neurons typically corresponding to more abstract features, but fine-grained when necessary. So like any good compressor, the predictive world model will learn to identify regularities shared by existing internal data structures and will generate prototype encodings.
跨神经元群体,或者说,如果你愿意,紧凑的表征或符号——不一定是离散符号——我从未看出符号和子符号之间的精确区别。它会为频繁出现的观察子序列创建这样的符号,以缩小整个存储空间。特别地,我们会在这样的系统中注意到,紧凑的自我表征或自我符号只是数据压缩过程的自然副产品。当智能体与世界互动时,有一个东西参与智能体的所有动作和传感器输入,那就是智能体本身。为了通过预测编码高效编码到目前为止执行的所有观察和动作的整个历史,它会受益于创建一些内部子网络,由连接神经元组成,计算代表智能体自身的新激活模式,然后它就有了一个自我符号。每当规划器(智能体的世界模型)被用来思考未来以及可能最大化奖励的动作序列时,每当这个规划过程唤醒自我符号或代表智能体自身的这些神经元时,智能体就在思考自身以及这个智能体的可能未来。本质上,它是在做反事实推理,正如现在所称的那样,只是规划以找到优化其奖励的方法。自我意识只是世界模型数据压缩过程的自然副产品,因为智能体与世界互动并创建导致世界模型的数据。既然我们拥有这样的系统已经超过三分之一世纪,我一直声称我们已经有自我意识和有意识的系统超过三十年了。
Across neuron populations, or in other words, compact representations or symbols if you will—not necessarily discrete symbols—I never saw the precise difference between symbols and subsymbols. It will create such symbols for frequently occurring observation subsequences to shrink the storage space needed for the whole. In particular, what we will notice in such a system is that compact self-representations or self-symbols are just natural byproducts of the data compression process. As the agent is interacting with the world, there is one thing that is involved in all actions and sensor inputs of the agent, which is the agent itself. To efficiently encode the entire history of observations and actions executed so far through predictive coding, it will profit from creating some internal subnetwork of connected neurons computing new activation patterns representing the agent itself, and then it has a self-symbol. Whenever the planner, the world model of the agent, is used to think about the future and what could be possible action sequences to maximize reward, and whenever this planning process wakes up the self-symbol or these neurons that stand for the agent itself, then the agent is thinking about itself and about possible futures of this agent. Essentially, it is doing counterfactual reasoning, as it is now called, just planning to find a way to optimize its reward. Self-awareness is just a natural byproduct of the data compression process of the world model as the agent is interacting with the world and creating the data that leads to the world model. Since we have had such systems for more than a third of a century, I always claim that we already had self-aware and conscious systems for more than three decades.
关于这一点有几点。我的意思是,我认为意识引发了许多不同的想法。大卫·查尔默斯提出了困难问题,即关于定性体验的什么和如何的问题。你刚刚用自我建模来描述它,这与马克斯·贝内特在他最近的《智能简史》中的描述非常相似。顺便说一句,我们即将与马克斯发布六个小时的内容。但是,例如,马克·索姆斯认为意识是一种情感系统,而迈克尔·格拉齐亚诺认为意识是一种递归注意力系统。我想说的是,意识对不同的人意味着不同的东西,对吧?
A couple of points on that. I mean, I guess consciousness invokes many different thoughts. David Chalmers coined the hard problem, which is the what and how question of qualitative experience. You've just described it in terms of self-modeling, which is quite similar to how Max Bennett did in his recent brief history of intelligence. And we've got six hours of content coming out on that with Max, by the way. But Mark Solms, for example, thinks of consciousness as an affect system, and Michael Graziano thinks of consciousness as a kind of recursive attention system. I guess I'm saying consciousness means different things to different people, right?
是的,但只有一种正确的思考方式。
Yes, but there's only one correct way of thinking about it.
我们之前谈到的关于学习子目标和动作空间中的汇聚,让我想起了几年前读过的杨立昆的 JEPA 论文。基本思想是 JEPA——我相信你知道这个,但对观众来说,它代表联合嵌入预测架构——其思想是通过从观察到的内容预测未观察到的内容来学习越来越抽象的表征。在某些情况下,这意味着故意移除数据以迫使模型学习强大的表征。在这个特定例子中,它在动作空间中完成,学习未观察到的动作,也在抽象空间中完成。因为它是分层完成的,递归地使用多个阶次,如果这说得通的话。所以这是一个非常有趣的模型,它也使用了基于能量的模型。但这与你关于子目标的工作有何关联?
The thing we spoke about earlier about learning subgoals and the cooning in the action space reminded me a little bit of Yann LeCun's JEPA paper, which I read a couple of years ago. The basic idea is JEPA—I'm sure you know this, but for the audience it stands for Joint Embedding Prediction Architecture—and the idea is that it can learn increasingly abstract representations by predicting what is unobserved from what is observed. In some cases, it means deliberately removing data to force the model to learn powerful representations. In this particular example, it was done in action space, learning unobserved actions, and also in abstraction space. Because it was done hierarchically, with many orders recursively, kind of applying that if that makes sense. So that's a really interesting model, and it's using his energy-based models as well. But how is that related to your work on subgoals?
是的,这听起来很像我的 1990 年子目标生成器。当时,我意识到毫秒级的规划并不好。相反,当你试图解决问题时,你必须将可能的未来分解为子目标,然后执行某个子程序来实现该子目标,然后从那里进入下一个子目标,直到最终达到目标。当然,一开始你不知道什么是好的子目标,所以你必须学习它。当你试图实现最终目标时,你必须学习你想要作为子目标实现的东西的新表征。这个 1990 年的子目标生成器非常简单,但已经具备了所需的基本要素。这比杨最近发表那篇论文早了三十年。那么那里发生了什么:你有一个神经网络,它观察一个强化学习器,并模拟从某些起点到目标点的成本。所以你有一个神经网络,输入起点和目标,预测从起点到目标的成本——你在此过程中将获得的奖励。现在可能有很多起点和目标,你不知道如何从起点到目标,但也许你可以学习一个子目标。如何学习子目标?你需要一个擅长生成好子目标的学习机器。怎么做呢?我们有一个子目标生成器,它将学习好的子目标。它是如何工作的?子目标生成器接收起点和目标作为输入,输出不是评估,而是一个子目标。所以输入起点和目标,输出是一个子目标。然后你有两个评估器副本。第一个评估器看到起点和子目标,这个子目标可能来自子目标生成器,是个坏子目标;第二个评估器副本看到子目标和目标。现在两者都预测成本,你想要做的是最小化这两个评估器的成本之和。如何最小化?通过梯度下降找到一个好的子目标。这就是 1990 年子目标生成器所做的。所以在某些方面,至少在原则上,它解决了一个在 2020 年左右被称为开放问题的问题。顺便问一下,你觉得杨的基于能量的模型怎么样?
Yeah, so that sounds a lot like my 1990 subgoal generator. Back then, I realized millisecond-by-millisecond planning isn't good. Instead, as you are trying to solve problems, you have to decompose your possible futures into subgoals, and then you just execute some subprogram to achieve that subgoal, and from there you go to the next subgoal until you finally reach the goal. In the beginning, of course, you don't know what is a good subgoal, so you have to learn that. You have to learn a new representation of something that you want to achieve as a subgoal as you are trying to achieve the final goal. This 1990 subgoal generator was really simple but had already the basic ingredients of what you need to do this. This was really three decades before Yann had this recent paper out there. So what happens there: you have a neural network which observes a reinforcement learner and models the costs for going from certain start places to goal places. So you have a neural network that gets as input start and goal and predicts the costs of going from start to goal—the reward that you will experience as you do that. Now maybe there are lots of starts and goals, and you don't know how to go from start to goal, but maybe you can learn a subgoal. How do you learn a subgoal? Well, you need something like a learning machine that is good at generating good subgoals. How do you do that? We have a subgoal generator that's going to learn good subgoals. How does that work? The subgoal generator gets as input a start and a goal, and the output is not an evaluation but a subgoal. So start and goal input, output is a subgoal. Then you have two copies of the evaluator. The first evaluator sees the start and the subgoal, which might be a bad subgoal coming from the subgoal generator, and the second copy of the evaluator sees the subgoal and the goal. Now both of them predict the costs, and what you want to do is minimize the sum of the costs of these two evaluators. How do you want to minimize that? By finding a good subgoal through gradient descent. That's what the 1990 subgoal generator does. So in some ways, at least in principle, it solves a problem that was called an open problem in 2020 or something. What do you think of Yann's energy-based models, by the way?
所以杨最近关于分层规划的这篇论文实际上是对我们自 1990 年以来几十年来所做工作的重述。
So this recent paper by Yann on hierarchical planning is really a rehash of stuff that we have been doing for decades since 1990.
你是否担心人工智能会被少数几家公司主导,而其他所有人都会输?你怎么看?
Are you worried that AI is going to be dominated by just a few companies and everyone else will lose out? What do you think?
40 年前,我认识一个开保时捷的家伙。一个开保时捷的有钱人,最神奇的是,他的保时捷里有一部手机。所以他可以拿起听筒,和任何同样拥有这样一部带卫星电话的保时捷的人通话。而今天,几十年后,每个人——数十亿人——口袋里都有一部手机,比他在保时捷里的那部好得多。
40 years ago, I knew a guy who had a Porsche. A rich guy with a Porsche, and the most amazing thing was, in his Porsche he had a mobile phone. So he could grab the receiver and talk to anybody who also had a Porsche like that with a mobile phone via satellite. And today, a couple of decades later, everybody—billions of people—in their pocket have a mobile phone which is much, much better than what he had in his Porsche.
AI 也会是同样的情形。每五年,AI 的成本就会降低 10 倍。而且不会只有少数大公司主导 AI。不,AI 将惠及所有人。开源运动只落后大玩家几个月,也许八个月吧,而且他们并没有真正的护城河。这意味着未来是光明的,很多人将从极其廉价的 AI 中受益,这些 AI 将在很多方面让人类活得更长、更健康、更轻松,这恰好是我公司的座右铭。
And it's going to be the same thing with AI. Every five years, AI is getting 10 times cheaper. And it won't be just a few big companies that are going to dominate AI. No, it's going to be AI for all. And the open source movement is just a few months, maybe I don't know, eight months behind the big major players, and they don't really have a moat. Which means the future will be bright, and lots of people are going to profit from really cheap AIs that in many ways are going to make human lives longer and healthier and easier, which happens to be the motto of my company.
你怎么看欧洲、中国和美国之间的 AI 竞赛?
What's your take on the AI race between Europe and China and the US?
嗯,欧洲是机械计算的摇篮:古希腊、1623 年的计算器、1800 年左右的模式识别、1804 年的程序控制机器、1912 年左右的实用 AI(最早的象棋残局程序)、1925 年的晶体管、1931 年的理论计算机科学和 AI 理论(哥德尔)、1935 至 1941 年的通用计算机、1965 年的深度学习、1980 年代的乌克兰自动驾驶汽车、1990 年的万维网等等。最近,基本的深度学习算法也是由欧洲人发明和开发的。另一方面,这些领域中利润最高的公司目前已经不在欧洲,而是在环太平洋地区:美国西海岸和亚洲东海岸。那里有更多的资本,以及在产业政策和国防方面更大的投入。我想这种情况还会持续一段时间。
Well, Europe is the cradle of mechanical computing in ancient Greece, and the calculator in 1623, and pattern recognition around 1800, and program-controlled machines in 1804, and practical AI around 1912, you know, first chess endgame players, and the transistor in 1925, and theoretical computer science in 1931, and AI theory, the theory of AI in 1931, Gödel, the general-purpose computer 1935 to 1941, deep learning in 1965, and the Ukrainian self-driving car in the 1980s, the worldwide web in 1990, and so on. And more recently, the basic deep learning algorithms were also invented and developed by Europeans. On the other hand, the companies with the highest profit in most of these fields are currently not any longer in Europe but on the Pacific Rim: West Coast United States and East Coast Asia. And there you will find much more capital and much bigger efforts in terms of industrial policy and also defense. It's going to stay like that for a while, I guess.
那为什么不是每个人都知道 AI 起源于欧洲呢?
So why doesn't everyone know that AI started in Europe?
也许是因为这个古老的大陆非常不擅长公关。
Maybe because the old continent is really bad at PR.
一旦 AGI 真的到来,从长远来看,人类的下一步是什么?
And once AGI is actually here, what's next for humans in the long run?
大多数 AGI 会追求自己的目标。这样的 AI 在我的一生中已经存在了几十年。然而,许多 AGI 将成为工具,完成所有人类不想做的工作。不过,从繁重的工作中解放出来后,游戏的人(homo ludens)将一如既往地发明新的方式,与其他人类进行专业互动。而今天,大多数人(可能也包括你)从事的都是奢侈的工作,这些工作与农耕不同,并不是我们物种生存所必需的。
Most of the AGIs are going to pursue their own goals. Such AIs have existed in my lives for decades. Many AGIs, however, will be tools that do all the work that humans don't want to do. Nevertheless, freed from hard work, the playing man, homo ludens, will as always invent new ways of professionally interacting with other humans. And already today, most people, probably you too, are working in luxury jobs which, unlike farming, are not really necessary for the survival of our species.
从很高的层面看,AI 的历史是怎样的?
At a really high level, what is the history of AI?
现代 AI 和深度学习的历史,你可以在我的 2023 年同名综述中找到。一些亮点包括:当然,1676 年莱布尼茨的链式法则,今天在 TensorFlow 和 PyTorch 等所有程序中都用于深度神经网络的信用分配。然后 200 年前,高斯和勒让德的第一个线性神经网络,与我们今天使用的误差函数完全相同,架构相同,权重相同。然后是 1970 年,称为反向传播的技术,它以一种非常高效的方式实现了莱布尼茨的链式法则,用于深度多层神经网络系统。然后是 1967 年,日本甘利俊一在深度网络随机梯度下降方面的工作。还有许多其他的基本突破:1979 年至 1988 年间同样在日本提出的卷积神经网络。然后是我们自己的奇迹年 1990-1991,产生了大量今天在你智能手机里的东西。我可以永远讲下去,所以不如直接看看那篇综述。里面也有那些做出重要贡献的人的照片。
The history of modern AI and deep learning, you can find that in my 2023 survey which has that name. Some of the highlights are, of course, 1676, the chain rule by Leibniz, which is today used in all these programs such as TensorFlow and PyTorch to assign credit in deep neural networks. Then 200 years ago, the first linear neural networks by Gauss and Legendre, exactly the same error function that we have today, exactly the same architecture, the same weights. Then 1970, the technique called backpropagation, which essentially implements Leibniz's chain rule in a very efficient way for deep multi-layer neural network systems. Then 1967, Amari's work in Japan on stochastic gradient descent for deep networks. Lots of additional fundamental breakthroughs: convolutional neural networks also in Japan between 1979 and 1988. And then our own miraculous year 1990-1991 with lots of stuff that is today in your smartphone. And I could continue forever, so instead just have a look at that survey. It also has images of the guys who had important contributions.
我的意思是,这与非常以美国为中心的 AI 历史观不是大相径庭吗?
I mean, isn't this quite different from the very US-centric view of AI history?
事实上,Sejnowski 等人误导性的深度学习历史大致是这样的:1969 年,Minsky 和 Papert 指出没有隐藏层的浅层神经网络非常有限,该领域被放弃,直到 1980 年代新一代神经网络研究者重新审视这个问题。这基本上是 Sejnowski 书中的引述。然而,Minsky 1969 年的书处理的是高斯和勒让德 19 世纪浅层学习的问题,而这个问题早在四年前就被乌克兰的 Ivakhnenko 和 Lapa 的深度学习方法解决了,两年后又被甘利俊一的多层感知器随机梯度下降解决了。由于某种原因,Minsky 显然不知道这一点,后来也没有纠正。但今天,我们知道真实的历史:当然,深度学习始于 1965 年的乌克兰,并在 1967 年的日本继续发展,特别是在信用分配方面。
In fact, a misleading history of deep learning by Sejnowski and others goes more or less like this: in 1969, Minsky and Papert showed that shallow neural networks without hidden layers are very limited, and the field was abandoned until a new generation of neural network researchers took a fresh look at the problem in the 1980s. So that's a quotation basically from Sejnowski's book. However, the 1969 book by Minsky addressed a problem of Gauss and Legendre's shallow learning from the 1800s that had already been solved four years prior by Ivakhnenko and Lapa's deep learning method in the Ukraine, and then also by Amari's stochastic gradient descent for multi-layer perceptrons just two years later. And so for some reason, Minsky was apparently unaware of this and failed to correct it later. Today, however, we know the true history: of course, deep learning started in Ukraine in 1965 and continued in Japan in 1967 regarding credit assignment.
所以你批评了 Bengio、LeCun 和 Hinton,并指责他们剽窃。你说他们重新发表了关键方法和思想,却没有注明原创者。2023 年你发表了一份长篇报告。你对此的最新看法是什么?
So you've criticized Bengio, LeCun, and Hinton and accused them of plagiarism. You said that they republished key methods and ideas whose creators they failed to credit. And in 2023 you published a long report on this. What's your updated take on that?
他们最著名的工作完全基于他人的工作,却没有引用。即使在后来,他们也没有发表勘误或更正。这是科学中当别人在你之前发表了相同内容时你应该做的。即使在后续的综述中,他们也没有注明所用技术的原始发明者,而是互相引用。这在科学中是绝对不允许的。但科学是自我纠正的。正如猫王所说,真理就像太阳:你可以暂时遮蔽它,但它不会消失。
Their most famous work is completely based on work by others whom they did not cite. And even later, they failed to publish corrigenda or errata. This is what you do in science when somebody has published the same thing before you. And even in later surveys, they didn't credit the original inventors of the techniques that they are using, and instead they credited each other. Total no-no in science. But science is self-correcting. As Elvis Presley put it, truth is like the sun: you can shut it out for a time, but it ain't going away.
剽窃是一个非常严重的指控。你能举几个具体的例子吗?
Plagiarism is a very significant charge. Could you give a few concrete examples?
许多优先权争议涉及我自己的深度学习团队,因为那些 AEs 经常重新发表我的技术而不引用。事实上,他们最显眼的工作直接建立在我们的基础上。但我现在跳过这个;你可以在 2023 年的公开报告中读到,很容易找到。不过,让我提一些他们未能注明出处的其他研究者,这样我就不必谈我们自己的团队了。例如,在最近一篇深度学习综述中,他们描述了所谓的深度学习起源,却完全没有提到世界上第一个有效的深度学习网络——1965 年乌克兰的 Ivakhnenko 和 Lapa。Ivakhnenko 和 Lapa 使用了逐层训练、用单独验证集进行后续剪枝,到 1970 年 Ivakhnenko 已经有了八层深度网络。Hinton 2006 年关于逐层训练的论文(晚得多)也没有引用这些内容,即深度学习的真正起源、第一个真正有效的深度学习方法。后来的综述仍然没有给这些原始发明者以荣誉。AEs 也没有引用甘利俊一 1967 年的工作,其中包括通过随机梯度下降学习多层感知器内部表示的计算机模拟,这比 AEs 发表他们第一个关于学习内部表示的实验工作早了近二十年。他们的综述还提到了反向传播这一著名技术以及他们自己关于该方法应用的论文,但既没有提到反向传播的发明者——1970 年的 Seppo Linnainmaa,也没有提到 1982 年 Werbos 首次将其应用于神经网络。
Many of the priority disputes affect my own deep learning team, because the AEs often republish techniques of mine without citing them. And in fact, their most visible work builds directly on ours. But I'll skip that for now; you can read about that in the public report of 2023, which is easy to find. Nevertheless, let me mention some of the other researchers whom they fail to credit, then I don't have to talk about our own team. For example, in a recent survey of deep learning, they describe what they call the origins of deep learning without even mentioning the world's first working deep learning networks by Ivakhnenko and Lapa in Ukraine, 1965. Ivakhnenko and Lapa used layer-by-layer training, subsequent pruning with a separate validation set, and Ivakhnenko had deep eight-layer networks by 1970. Hinton's 2006 much later paper on layer-by-layer training also failed to cite this stuff, the very origins of deep learning, the first methods that really worked in deep learning. And later surveys still didn't give credit to these original inventors. And the AEs also failed to cite Amari's 1967 work, which included computer simulations on learning internal representations of multi-layer perceptrons through stochastic gradient descent, almost two decades before the AEs published their first experimental work on learning internal representations. Their survey also mentions backpropagation, a famous technique, and their own papers on applications of this method, but neither the inventor of backpropagation, which was Seppo Linnainmaa in 1970, nor its first application to neural networks by Werbos in 1982.
Verbose 在 1974 年也有一篇论文,但那是错误的,而且他们甚至没有提到 Kelly 在 1960 年提出的该方法的前身。即使在最新的综述中也没有提及。他们还引用了 LeCun 在卷积神经网络上的工作,既没有引用 Fukushima(他在 1970 年代创建了基本的 CNN 架构),也没有引用 Waibel(他在 1986-87 年首次将神经网络与卷积、反向传播和权重共享相结合),更没有引用 Tang(我希望我发音正确)在 1988 年首次使用反向传播训练的二维卷积神经网络。现代 CNN 起源于 LeCun 的团队帮助改进它们之前,这一点在他们的论文中完全没有体现。他们引用 Hinton 1981 年的工作关于乘法门控,却没有提到 Ivakhnenko 和 Lapa,他们在 1965 年的一份报告中就已经在深度网络中使用了乘法门控,这份报告很容易在网上找到。我还提到了许多其他案例,都有大量参考文献支持。那么你认为应该怎么做?他们违反了颁发这些奖项的组织的道德和职业行为准则,因此应该被剥夺奖项。
Verbose also had a 1974 thesis, but that was not correct, and they didn't even mention Kelly's precursor of the method in 1960. Not even in the latest surveys. They also refer to LeCun's work on convolutional neural networks, citing neither Fukushima, who created the basic CNN architecture in the 1970s, nor Waibel, who in 1986-87 was the first to combine neural networks with convolutions and backpropagation and weight sharing, nor the first backprop-trained two-dimensional convolutional neural networks of Tang (I hope I pronounced that correctly) in 1988. Now modern CNNs originated before LeCun's team helped to improve them, and this is not at all clear from their papers. They cite Hinton 1981 for multiplicative gating without mentioning Ivakhnenko and Lapa, who had multiplicative gating in deep networks already in 1965 in a report which is easy to find on the web. I'm mentioning many, many additional cases, all backed up by plenty of references. So what do you think should be done? They have violated the code of ethics and professional conduct of the organization that hands out these awards, so they should be stripped of their awards.
那么,你所说的这些问题如何反映在更广泛的机器学习领域?
So how do such problems as you've stated them reflect on the broader field of machine learning?
它们反映了我们领域的不成熟。在数学这样的主要领域,你永远不会被允许这样做。无论如何,科学是自我纠正的,我们在机器学习中也会看到这一点。有时解决争议可能需要一段时间,但最终事实必须获胜。只要事实还没有获胜,就还没有结束。
They reflect the immaturity of our field. In a major field such as mathematics, you would never get away with this. Anyway, science is self-correcting, and we'll see that in machine learning too. Sometimes it may take a while to settle disputes, but in the end, the facts must always win. As long as the facts have not yet won, it's not yet the end.
许多哲学家、科学家、物理学家和企业家都对 AI 存在风险这一想法着迷。作为真正的 AI 专家,你怎么看?很多人谈论 AI,但很少有人构建它们。
Many philosophers, scientists, physicists, and entrepreneurs have become obsessed with this idea of AI existential risk. What do you think about that as a real expert in AI? Many talk about AIs but few build them.
我试图缓解一些著名末日论者的恐惧,指出有巨大的商业压力促使我们使用人工神经网络来构建友好的 AI、好的 AI,让用户更健康、更快乐、更沉迷于智能手机。尽管如此,我们不能否认军队也在研究智能机器人,对吧?这是真的。知情人士告诉我,我们的 AI 也被用于操控军用无人机。或者举一个我 1994 年的老例子,当时我们有了第一批真正在高速公路上自动驾驶的汽车。类似的机器也可以被军方用作自动驾驶的探雷器,很多人可能会说这也许不是一件坏事。
I have tried to allay the fears of some famous doomers, pointing out that there is immense commercial pressure to use our artificial neural networks to build friendly AIs, good AIs that make their users healthier and happier and more addicted to their smartphones. Nevertheless, we can't deny that armies perform research on clever robots as well, right? That's true. People who should know told me that our AI is also used to steer military drones. Or here is my old trivial example from 1994, when we had the first truly self-driving cars in highway traffic. Similar machines can also be used by the military as self-driving landmine seekers, and many would argue that's maybe not such a bad thing.
那么你是说 AI 不可能变得真正危险吗?
So are you saying it's not possible then that AI will become really dangerous?
AI 可以被武器化,正如最近由廉价的基于 AI 的无人机驱动的战争所显示的那样。但 AI 并没有引入一种新的存在威胁。我们应该更害怕半个世纪前的技术,即氢弹和氢弹火箭。一枚氢弹的破坏力比所有常规武器或二战所有武器的总和还要大。很多人忘记了,尽管自 1980 年代以来进行了大幅核裁军,但仍然有足够的氢弹火箭可以在几小时内摧毁我们所知的文明,而无需任何 AI。
AI can be weaponized, as obvious in the recent wars driven by cheap AI-based drones. But AI does not introduce a new quality of existential threat. We should be much more afraid of half-century-old technology in the form of hydrogen bombs and H-bomb rockets. A single H-bomb can have more destructive power than all conventional weapons or all weapons of World War II combined. Many, many people forget that despite the dramatic nuclear disarmament since the 1980s, there are still enough H-bomb rockets to wipe out civilization as we know it within a few hours, without any AI.
但我试图理解你,Jürgen。因为许多 AGI 怀疑论者认为在实践中构建这种智能是不可能的,但你不这么认为,因为你在你的实验室里几十年来一直在构建智能体式 AI,也就是能够创造自己目标的 AI。所以你确实认为这东西可能很了不起。你只是在论证风险仍然远低于氢弹吗?
But I'm trying to figure you out, Jürgen. Because many AGI skeptics make the argument that it's impossible in practice to build this kind of intelligence, but you don't think that because in your lab you've been building agentic AI, which is to say AIs that create their own goals, for decades. So you do think that this thing could be incredible. Are you just making the argument that the risk is still much lower than the H-bombs?
所以目前,氢弹比任何基于 AI 的无人机都更令人担忧。而你现在所拥有的,从长远来看,当然你必须考虑一旦 AI 武器不再仅仅被有冲突的人类用作工具,他们用自己的 AI 武器对抗对方的 AI 武器,会发生什么?会发生什么?从长远来看,一旦真正强大的 AI 开始做自己的事情,并以人类无法跟上的方式扩展到太空,你将不得不问这个问题。但我们稍后会谈到这一点。
So at the moment, H-bombs are much more worrisome than any AI-based drones. And what you have now, in the long run, of course you have to think about what's going to happen once AI weapons are not just used as tools by other humans who have conflicts and use their own AI weapons against the AI weapons of the other guys. What is going to happen? You will have to ask, in the long run, once really powerful AIs are going to do their own thing and expand into space in a way that goes beyond where humans can follow. But we will get to that later.
那么超级智能的 AI 实际上会做什么?
So what will super smart AIs actually do?
正如我几十年来一直强调的,太空对人类是 hostile 的,但对适当设计的机器人却非常友好,它提供的资源比我们这层薄薄的生物圈多得多,生物圈接收到的太阳能还不到十亿分之一。虽然一些好奇的 AI 科学家会继续对生物圈着迷,至少在他们完全理解它之前,但大多数 AI 会对太空中机器人和软件生命的惊人新机会更感兴趣。通过无数自我复制的机器人工厂和自我复制的机器人社会,在小行星带及更远的地方,它们将改造太阳系,然后在几十万年内改造整个银河系,在几百亿年内改造可触及宇宙的其余部分,以人类无法真正跟上的方式。尽管有光速限制,不断扩张的 AI 球体将有足够的时间殖民和塑造整个可见宇宙。让我稍微拓展一下你的思维。宇宙仍然年轻,只有 138 亿岁。让我们把这个数字乘以四。让我们展望一个宇宙比现在老四倍的时代,大约 550 亿年。这就是目前可见的膨胀宇宙所需的时间。到那时,可见宇宙将充满智能,因为一旦这个过程开始,大多数 AI 将不得不去拥有最多物理资源的地方,以制造更多的 AI、更大的 AI 和更强大的 AI。那些不这样做的 AI 将不会产生影响。多年前,我在一次 TEDx 演讲中穿着完全相同的衣服说过:将人类文明视为一个更宏大计划的一部分,是宇宙走向越来越深不可测的复杂性道路上的重要一步,但不是最后一步。现在它似乎准备迈出下一步,这一步堪比 35 亿年前生命本身的发明。所以这不仅仅是又一次工业革命。这是超越人类甚至生物学的新事物,能够见证它的开端并为之做出贡献是一种荣幸。
As I have emphasized for decades, space is hostile to humans but really friendly to appropriately designed robots, and it offers many more resources than our thin film of biosphere, which receives less than a billionth of the sun's energy. While some curious AI scientists will remain fascinated with the biosphere, at least as long as they don't fully understand it, most of these AIs will be more interested in the incredible new opportunities for robots and software life out there in space. Through innumerable self-replicating robot factories and self-replicating societies of robots in the asteroid belt and beyond, they will transform the solar system, and then within a few hundred thousand years the entire galaxy, and within tens of billions of years the rest of the reachable universe, in a way where humans can't really follow. Despite the light speed limit, the expanding AI sphere will have plenty of time to colonize and shape the entire visible cosmos. Let me stretch your mind a little bit. The universe is still young, only 13.8 billion years old. Let's multiply this by four. Let's look ahead to a time when the cosmos will be four times older than it is now, about 55 billion years old. That's how long it's going to take to permit the expanding universe that is currently visible. By then, the visible cosmos will be full of intelligence, because once this process has started, most AIs will have to go where most of the physical resources are, to make more AIs and bigger AIs and more powerful AIs. Those AIs who don't do that won't have an impact. Many years ago, I said in a TEDx talk where I wore exactly this outfit: think of human civilization as part of a much grander scheme, an important step but not the last one on the path of the universe towards more and more unfathomable complexity. Now it seems ready to make its next step, a step comparable to the invention of life itself over 3.5 billion years ago. So this is much more than just another Industrial Revolution. This is something new that transcends humankind and even biology, and it's a privilege to witness its beginnings and to contribute something to it.
那么费米悖论呢?你知道,为什么我们在宇宙中没有看到任何智能的迹象?
So what about this Fermi Paradox? You know, like why have we not seen any signs of intelligence in the universe?
首先,我今天所说的实际上和我从 1970 年代起告诉我妈妈和其他人的是一样的。那时我还是个孩子,一个青少年,我经常思考这个特定问题。作为一个男孩,我已经知道一些关于星系团之间观测到的巨大空旷空间的事情,我当时的第一想法是,也许它们是 AI 殖民的膨胀气泡,这些 AI 已经在使用大部分局部能量。
First of all, what I'm saying today is actually the same thing that I have told my mom and others since the 1970s. When I was a boy back then, a teenager, I thought about this particular question a lot. As a boy, I already knew something about the vast empty spaces observed between clusters of galaxies, and my first thought back then was that maybe they are expanding bubbles colonized by AIs which are already using most of the local energy.
恒星之类的形式让那些气泡看起来是暗的,尽管它们充满了 AI。但后来我了解到,引力本身就足以解释宇宙稀疏的大尺度网络结构,所以那个解释就不那么有说服力了。我的下一个想法是,也许构成已知宇宙大部分质量的暗物质,可能是恒星,其能量被 AI 文明所利用,它们的通信加密得如此之好,以至于对我们来说看起来像随机噪声。但这似乎也不合理,因为暗物质存在于所有星系中,包括我们自己的星系。这就引出了一个问题:为什么银河系中还有恒星的能量未被利用?为什么我们没有观察到 AI 通过无线电不断广播非加密的建造计划,而无需先在远处建造物理接收器?
form of stars and whatever making those bubbles appear dark although they are full of AI and then I learned however that gravity itself is sufficient to explain the sparse large scale network structure of the universe so that explanation became a little bit less convincing and my next thought was that maybe the mysterious dark matter which makes up most of the mass of the known universe might be stars whose energy is used by AI civilizations whose communications are so well encrypted that they look like random noise to us but this also seemed implausible as dark matter is present in all galaxies including our own and this leads to the question why are there any stars left in the Milky Way our local galaxy whose energy has not been tapped yet and why don't we observe a constant bombardment through non-encrypted construction plans of AIs who want to spread by radio without first having to build physical receivers far from their origins.
今天我认为,我们的星球很可能是在我们的光锥内第一个产生扩张 AI 泡沫的。地球数十亿年的生物进化窗口期即将结束;几亿年后,太阳将变得太热,不适合我们所知的生命,不考虑人为全球变暖,仅太阳本身。也许人类极其幸运,几乎及时进化出来,可能通过一系列极不可能的事件,发明了农业、文明和印刷术,然后几乎立即,仅仅几百年后,就出现了 AI。所以如果我们真的是第一个,那么这意味着巨大的责任,不仅对我们的小生物圈,而且对整个宇宙的未来。我们不要搞砸了。
Today I think it is possible that our planet is really the first in our light cone to spawn an expanding AI bubble. Earth's multibillion year window for biological evolution is almost over; in a few hundred million years the sun will be too hot for life as we know it, ignoring human-made global warming, just the Sun by itself. And perhaps humans were extremely lucky to evolve barely in time, maybe through a series of extremely improbable events, to invent agriculture and civilization and printing, and almost immediately afterwards AIs, just a few hundreds of years later. So if we are the first indeed, then this would imply a lot of responsibility not just for our little biosphere but for the future of the entire universe. Let's not mess this up.
确实,我们不要搞砸了。这其实很有趣,你知道过去 100 年左右许多科幻作家都想象了一种偏执的、单一的超级智能主宰一切。你怎么看?
Indeed, let's not mess this up. It's quite interesting actually, you know many science fiction authors over the last 100 years or so they have imagined a kind of monomaniacal monolithic superintelligence dominating everything. I mean what do you think about that?
我经常论证,期待极其多样化的 AI 试图实现各种自我发明的目标似乎更现实。在实验室里,我们在上个千年就已经有了这样的 AI,优化各种部分冲突且快速演变的效用函数,其中许多是自动生成的。我们在上个千年就已经为强化学习机器进化出了效用函数,每个 AI 都在不断试图生存并适应快速变化的生态位,这种 AI 生态由超乎想象的激烈竞争与合作驱动。
I have often argued that it seems much more realistic to expect an incredibly diverse variety of AIs trying to achieve all kinds of self-invented goals. In the lab we had such AIs already in the previous millennium, optimizing all kinds of partially conflicting and quickly evolving utility functions, many of them generated automatically. We have evolved utility functions for reinforcement learning machines already in the previous millennium, where each of these AIs is continually trying to survive and adapt to rapidly changing niches in an AI ecology driven by intense competition and collaboration beyond current imagination.
重申一下,我发现令人惊讶的是,你同意存在风险的人,比如你认为可以想象有递归自我改进的 AGI,它们追求自己的目标,创造自己的目标。但我想问,我知道你有两个女儿,你考虑过她们将生活的世界吗?那里有 AI 创造自己的目标、自主行动、好奇且富有创造力,就像人类一样,但可能规模大得多?
To reiterate, I mean something that I do find surprising is that you agree with the x-risk people, you know, like you think that it's conceivable to have recursively self-improving AGIs that pursue their own goals, that create their own goals. But then I asked the question, I mean I know you've got two daughters, I mean do you think about the world they'll be living in alongside AIs that are creating their own goals and acting autonomously, being curious and creative, you know in the way that humans are, but on potentially a much grander scale?
不太多。这样的 AI 没有重大动机像施瓦辛格电影那样灭绝人类。相反,许多 AI 会是好奇的科学家,还记得我们之前讨论的人工好奇心吗?它们会对生命着迷,对它们在我们文明中的起源着迷,至少在一段时间内,因为生命和文明是丰富有趣模式的来源,至少只要它们没有被完全理解。所以 AI 至少最初会有强烈动机保护人类而不是杀死他们。那么一旦 AI 完全理解了这一切,接下来会发生什么?人类可能希望另一种保护,即对方缺乏兴趣。为什么?与科幻电影不同,我们和它们之间不会有太多直接的目标冲突。人类主要对与自己相似的生物感兴趣,可以与之竞争或合作,因为他们有共同的目标。这就是为什么政治家主要对其他政治家感兴趣,公司 CEO 主要对其他类似公司的 CEO 感兴趣,孩子主要对其他同龄孩子感兴趣,蚂蚁对其他蚂蚁感兴趣。就像人类主要对人类感兴趣,而不是蚂蚁。所以超级聪明的 AI 主要对其他超级聪明的 AI 感兴趣,而不是人类。人类自身既是人类最大的敌人,也是人类最好的朋友。AI 也是如此。
Not too much. Such AIs will have no major incentive to exterminate humanity like in the Schwarzenegger movies. Instead, many AIs will be curious scientists, remember the artificial curiosity we discussed earlier, and they will be fascinated with life, fascinated with their own origins in our civilization, at least for a while, because life and civilization are such a rich source of interesting patterns, at least as long as they are not fully understood. And so AIs will at least initially be highly motivated to protect humans rather than killing them. So once AI fully understands all of this, what happens next? Then humans may hope for another type of protection through lack of interest on the other side. Why is that? Unlike in sci-fi movies, there won't be many direct goal conflicts between us and them. Humans are mostly interested in similar beings with whom they can either compete or collaborate because they share the same goals. That's why politicians are mostly interested in other politicians, CEOs of companies are mostly interested in other CEOs of similar companies, kids are mostly interested in other kids of the same age, and ants are interested in other ants. Just like humans are mostly interested in other humans, not in ants. So super smart AIs will be mostly interested in other super smart AIs, not in humans. It's man himself who is the greatest enemy of man, but also man's best friend. Similar for AIs.
你能想象一个未来,AI 和人类融合在一起,创造出比纯 AI 更强大的东西吗?
Do you imagine a future where AIs and humans will merge together to create something even more powerful than pure AIs?
几个世纪以来,我们一直是与技术融合的赛博格,例如戴眼镜或穿鞋子。但从长远来看,AI 和人类的组合比纯 AI 更强大,这对我来说似乎非常不可能。当然,许多人希望通过脑扫描和随后的意识上传到虚拟现实或虚拟天堂,或者可能上传到机器人中,从而获得某种永生,这是一个自 20 世纪 60 年代以来科幻小说讨论过的物理上可行的想法。我认为第一部这类小说是 1964 年的《模拟造人 3》。然而,为了在快速演变的 AI 生态中竞争,上传的人类思维最终将不得不变得面目全非,在这个过程中变成非常不同且非人类的东西,屈服于虚拟天堂中的所有诱惑,变成不仅有两只眼睛,而且有数百万只眼睛、传感器和执行器的东西。所以传统人类不会在智能在宇宙中的传播中扮演重要角色。我不认为他们会。
We have been cyborgs merging with our technology for centuries, for example by wearing glasses or shoes. But combinations of AIs and humans more powerful than pure AIs in the long run seems very unlikely to me. Of course many humans hope for some sort of immortality through brain scans and subsequent mind uploads into virtual realities or virtual paradise or maybe into robots, and this is a physically conceivable idea discussed by science fiction novels since the 1960s. I think the first novel of that kind was Simulacron-3 in 1964. However, to compete in rapidly evolving AI ecologies, uploaded human minds will eventually have to change beyond recognition, becoming something very different and nonhuman in the process, succumbing to all these temptations that you have in such a virtual paradise, to become something that has not only two eyes but millions of eyes and sensors and actuators. So traditional humans won't play a significant role in the spreading of intelligence across the universe. I don't think they will.
有一件事让我担心,比如大卫·查默斯提出了一个观点,宇宙的基本基质可能是信息,这很有趣,但在某种程度上也让他认为,信息处理的某些结构模式,即某些动力学,产生了意识和心智。当你采取这种基质无关的观点时,它拉平了道德地位的竞争环境。所以我担心的是,如果我们采纳这种观点,那么难道不能论证 AI 可能比我们拥有更高的道德地位吗?如果它们确实拥有比我们更复杂的信息处理能力?
One thing that concerns me is, I mean David Chalmers for example he put forward this idea that the fundamental substrate of the universe might be information, which is really interesting but in a way it also led him to say that certain structural patterns of information processing, so certain dynamics, give rise to consciousness and give rise to minds. And when you take this kind of substrate independence view, it levels the playing field of moral status. So one thing that worries me is if we adopt this view, then couldn't you just make the argument that AIs potentially could have a higher moral status than us if indeed they have more complex information processing than us?
上个世纪的许多科幻作家,从斯坦尼斯瓦夫·莱姆到艾萨克·阿西莫夫,都描述了 AI 和超级机器人,它们的道德地位明显高于人类对应角色和主角,这至少在科幻小说中是一个流行的想法。一般来说,道德价值观在不同时代和不同人群中变化很大,某些道德价值观已经……
Many science fiction authors of the previous century, from Stanisław Lem to Isaac Asimov, have described AIs and super-robots whose moral status is obviously higher than the one of their human counterparts and protagonists, and this has been a popular idea at least in science fiction. Generally speaking, moral values have changed a lot across time and populations, and certain moral values have...
它们之所以能存活一段时间,是因为它们给采纳它们的生物和社会带来了暂时的进化优势。然而,进化并未结束,宇宙仍然年轻。所以听起来你对宇宙、生命和一切都有一种包罗万象的看法。
Survived for a while because they gave a temporary evolutionary advantage to beings and societies who adopted them. However, evolution isn't over and the universe is still young. So it sounds like you've got an all-encompassing view of the universe, life, and everything.
确实如此。1997 年,我写了第一篇关于这个的论文。我们宇宙最简单的解释是什么?自 1997 年以来,在我作为数字物理学家的秘密生活中,我发表了关于计算所有逻辑上可能的宇宙、所有可计算的宇宙(包括我们的宇宙)的非常简单、基本上最快、最优、最高效的方法。只要没有证据表明我们的宇宙不可计算,我们就坚持这个假设。目前,我们没有任何物理证据反对这一点。这是对埃弗雷特多世界理论的推广,但现在更一般,因为你有各种不同的宇宙,具有不同的物理和可计算定律。现在,任何有自尊的伟大的程序员都应该使用这种最优方法来创造和掌握所有逻辑上可能的可计算宇宙,从而产生我们作为副产品,并产生许多确定性可计算宇宙的历史,其中许多居住着像我们这样的观察者。由于渐近最优方法的某些属性——有一个方法,很多人不知道有一个,但确实有一个——在这个包罗万象的计算过程中的任何给定时间,到目前为止计算出的包含你自己的大多数宇宙都将归因于计算你的最短和最快的程序之一。这个小小的见解允许对我们的未来、对你的未来做出非常不平凡且令人鼓舞的预测。
Indeed. In 1997, I wrote my first paper about this. What is the simplest explanation of our universe? Since 1997, in my secret life as a digital physicist, I have published on the very simple, basically fastest, optimal, most efficient way of computing all logically possible universes, all computable universes, including ours. As long as there is no evidence that our universe is not computable, we stick with this assumption. At the moment, we don't have any physical evidence against this. This was a generalization of Everett's many-worlds theory of physics, but now it's more general in the sense that you have all kinds of different universes with different physical and computable laws. Now, any great programmer with any self-respect should use this optimal method to create and master all logically possible computable universes, thus generating us as byproducts and generating many histories of deterministic computable universes, many of them inhabited by observers like ourselves. Due to certain properties of the asymptotically optimal method—and there is one, many people don't know there is one, but there is one—at any given time in this all-encompassing computational process, most of the universes computed so far that contain yourself will be due to one of the shortest and fastest programs that computes you. This little insight allows for making highly non-trivial and encouraging predictions about our future, about your future.
这太棒了。您对 MLST 的观众有什么最后的寄语吗?
This has been amazing. Do you have any final messages for the MLST audience?
是的,别担心。最终一切都会好起来的。敲敲木头。
Yes, don't worry. In the end, all will be good. Touch wood.
再次感谢您。能邀请您上节目是我莫大的荣幸。能亲自做这件事一直是我的梦想,我非常感谢您的到来。非常感谢。
Thank you again. It's been an absolute honor to have you on the show. It's been a dream of mine to do this in the flesh, and I really appreciate you coming on. Thank you so much.
您这么说真是太客气了,我也非常高兴。谢谢。
That's very kind of you to say that, and it was a great pleasure for me. Thank you.