Yann LeCun on Starting AMI: Open Research and World Models
打开互动全文版(中英对照 + 朗读 + 问答)→杨立昆讨论他的新创业公司 Advanced Machine Intelligence,强调开放研究和世界模型,并与大型 AI 实验室日益封闭的趋势形成对比。
Yann LeCun discusses his new startup Advanced Machine Intelligence, emphasizing open research and world models, and contrasts it with the growing secrecy in big AI labs.
嗨,Anne,欢迎来到信息瓶颈。我得说这对我来说有点奇怪,我认识你将近五年了,我们密切合作过,但这是我第一次以播客形式采访你,对吧?通常我们的对话更像是‘Yann,这不行,我该怎么办?’好吧,尽管我相信所有观众都认识你,但我还是要说,Yann LeCun 是图灵奖得主、深度学习教父之一、卷积神经网络的发明者、Meta 基础 AI 研究实验室的创始人,现在仍是他们的首席 AI 科学家,也是纽约大学的教授。欢迎你。
Hi Anne, and welcome to the information bottleneck. I have to say this is a bit weird for me, like I've known you for almost 5 years and we have worked closely together, but this is the first time that I'm interviewing you for a podcast, right? Usually our conversations are more like, 'Yann, it doesn't work. What should I do?' Okay, so even though I'm sure all of our audience knows you, I will say Yann LeCun is a Turing Award winner, one of the godfathers of deep learning, the inventor of convolutional neural networks, founder of Meta's fundamental AI research lab, and still their chief AI scientist, and a professor at NYU. So, welcome.
很高兴来到这里。
Pleasure to be here.
能坐在你身边是我的荣幸。我进入这个行业的时间比你们两位都短得多,做研究的时间也短得多。所以能定期与 Rafeed 合作发表论文是一种荣誉,而能开始主持这个播客更是如此。能和你坐下来聊聊真的很开心。
And it's a pleasure for me to be anywhere near you. I have been in this industry for a lot less time than either of you and doing research for a lot less time. So the fact that I'm able to publish papers somewhat regularly with Rafeed has been an honor and to be able to start hosting this podcast has been even more of one. So it's really a pleasure to sit down with you.
太棒了。我们想祝贺你创办新公司,对吧?你最近宣布,在 Meta 工作 12 年后,你创办了一家新公司 Advanced Machine Intelligence,专注于世界模型。首先,站在另一边感觉如何?从大公司到从零开始创业。
Awesome. Yeah, so we thought I congratulate you on the new startup, right? You recently announced that after 12 years at Meta, you're starting a new startup, Advanced Machine Intelligence, that you'll focus on world models. So first of all, how does it feel to be on the other side? Going from a big company to starting something from scratch.
我之前也联合创办过公司。但这次比之前更深入,我知道怎么运作。这次独特之处在于一种新现象:投资者对 AI 将产生巨大影响抱有足够希望,愿意投入大量资金,这意味着你可以创建一个初创公司,头几年基本上专注于研究。这在以前是不可能的。以前在工业界做研究的唯一地方是那些不为生存挣扎、基本占据市场主导地位、有足够长远眼光来资助长期项目的大公司。从历史上看,我们记得的大实验室,比如贝尔实验室属于 AT&T,它基本上垄断了美国的电信。IBM 垄断了计算机,他们有一个很好的研究实验室。施乐垄断了复印机,这使他们能够资助 PARC。但他们没能从那里的研究中获利,而苹果获利了。最近则有微软研究院、谷歌研究院和 Meta 的 FAIR。行业又在转变。FAIR 通过非常开放的方式对 AI 研究生态系统产生了巨大影响:发表一切,开源一切,包括 PyTorch 等工具,以及许多人在工业中使用的研究原型。所以我们促使其他实验室,比如谷歌,变得更加开放,也更系统地发表成果。但过去几年发生的是,很多这些实验室开始封闭起来,变得更加保密。确实如此。OpenAI 几年前就是这样,现在谷歌变得更封闭,甚至 Meta 也可能如此。所以,是时候把我感兴趣的那种研究放到 Meta 之外来做了。
Well, I co-founded companies before. I was involved more peripherally than this new one, but I know how this works. What's unique about this one is a new phenomenon where there is enough hope from investors that AI will have a big impact that they're willing to invest a lot of money, which means now you can create a startup where the first couple years are essentially focused on research. That just was not possible before. The only place to do research in industry before was in a large company that was not fighting for its survival and basically had a dominant position in the market and had a long enough view to fund long-term projects. From history, the big labs we remember like Bell Labs belonged to AT&T, which basically had a monopoly on telecommunication in the US. IBM had a monopoly on computers and they had a good research lab. Xerox had a monopoly on photocopiers and that enabled them to fund PARC. It did not enable them to profit from the research going on there, but that profited Apple. Then more recently Microsoft Research, Google Research and FAIR at Meta. The industry is shifting again. FAIR had a big influence on the AI research ecosystem by being very open, publishing everything, open sourcing everything with tools like PyTorch, but also research prototypes that a lot of people have been using in industry. So we caused other labs like Google to become more open and to also publish much more systematically than before. But what's been happening over the last couple years is that a lot of those labs have been clamming up and becoming more secretive. That's certainly the case. That was the case for OpenAI some years ago and now Google is becoming more closed and possibly even Meta. So it was time for the type of stuff that I'm interested in to be done outside Meta than inside.
那么明确一下,AMI,Advanced Machine Intelligence,计划公开进行研究吗?
So to be clear then, does AMI, Advanced Machine Intelligence, plan to do their research in the open?
是的,上游研究。在我看来,除非你发表你的成果,否则不能称之为研究,因为否则你很容易自欺欺人。你想出一些东西,认为它是自切片面包以来最伟大的发明。如果你不把它提交给社区,你可能只是妄想。我在很多工业研究实验室多次看到这种现象:内部对某些项目大肆炒作,但后来发现别人做得更好。所以如果你告诉科学家发表他们的工作,首先,这会激励他们做更好的工作,方法论更严谨,结果更可靠。这对他们有好处,因为当你做一个研究项目时,你对产品的影响可能是几个月、几年甚至几十年后。你不能告诉人们‘来为我们工作,但不要说你在做什么,也许五年后你会对某个产品产生影响’。在那期间,他们不会有动力去做真正有用的事情。所以如果你那样告诉他们,他们倾向于做短期影响的事情。所以如果你真的想要突破,你需要让人们发表。没有其他办法。而这是很多行业目前正在忘记的。
Yeah, upstream research. In my opinion you can't really call it research unless you publish what you do because otherwise you can get easily fooled by yourself. You come up with something you think is the best thing since sliced bread. If you don't actually submit it to the rest of the community, you might just be delusional. I've seen that phenomenon many times in lots of industry research labs where there is sort of internal hype about some internal projects, but then realizing that other people are doing things that actually are better. So if you tell scientists to publish their work, first of all, that is an incentive for them to do better work where the methodology is more thorough and the results are more reliable. It's good for them because very often when you work on a research project, the impact you may have on product could be months, years, or decades down the line. And you cannot tell people to come work for us, don't say what you're working on, and maybe there is a product you will have an impact on 5 years from now. In the meantime, they can't be motivated to really do something useful. So if you tell them that, they tend to work on things that have a shorter impact. So if you really want breakthroughs, you need to let people publish. You can't do it any other way. And this is something that a lot of the industry is forgetting at the moment.
AMI 计划生产或制造什么产品吗?是研究还是更多?
Does AMI, like what products if any does AMI plan to produce or make? Is it research or more than that?
不,不止是研究。是实际的产品。但我必须处理世界模型和规划,目标是成为未来智能系统的主要供应商之一。我们认为当前使用的架构,比如 LLM 或基于 LLM 的智能体系统,在语言方面还行。但智能体系统实际上效果并不好。它们需要大量数据来基本上克隆人类行为,而且不太可靠。所以我们认为正确的处理方式——我已经说了将近十年——是拥有能够预测 AI 系统可能采取的行动或行动序列后果的世界模型。然后系统通过优化,通过找出什么行动序列能最优地完成我设定的任务,来得出行动序列或输出。这就是规划。我认为智能的核心是能够预测行动的后果,然后用于规划。这就是我多年来一直在研究的东西。
No, it's more than that. It's actual products. But the things I have to do with world models and planning, with the ambition of becoming one of the main suppliers of intelligent systems down the line. We think the current architectures employed, like LLMs or agentic systems based on LLMs, work okay for language. But agentic systems really don't work very well. They require a lot of data to basically clone the behavior of humans. And they're not that reliable. So we think the proper way to handle this, and I've been saying this for almost 10 years now, is to have world models that are capable of predicting the consequences of an action or sequence of actions that an AI system might take. Then the system arrives at a sequence of actions or an output by optimization, by figuring out what sequence of actions will optimally accomplish a task that I'm setting for myself. That's planning. I think the central part of intelligence is being able to predict the consequences of your actions and then use them for planning. And that's what I've been working on for many years.
我们在纽约大学和 Meta 的项目上取得了快速进展。现在是时候把它变成现实了。你觉得缺失的部分是什么?为什么花了这么长时间?因为你说这个已经很多年了,但它仍然不如 LLM,对吧?
We've been making fast progress with a combination of projects here at NYU and also at Meta. And now it's time to basically make it real. What do you think are the missing parts? And why is it taking so long? Because you've been talking about it for many years already, but it's still not better than LLMs, right?
这和 LLM 不是一回事。它旨在处理高维、连续且嘈杂的模态。而 LLM 在这方面完全不行,它们真的不奏效。如果你试图训练一个 LLM 来学习图像或视频的良好表征,效果并不好。通常,AI 系统的视觉能力是单独训练的,不属于 LLM 的一部分。所以,如果你想处理高维、连续且嘈杂的数据,就不能使用生成模型,尤其不能使用将数据标记化为离散符号的生成模型。这根本行不通。我们有大量经验证据表明这效果不佳。有效的方法是学习一个抽象表征空间,消除输入中许多细节,尤其是所有不可预测的细节(包括噪声),然后在该表征空间中进行预测。这就是 GEPA(联合嵌入预测架构)的思想,你对此很熟悉。
It's not the same thing as LLM. It is designed to handle modalities that are high-dimensional, continuous, and noisy. And LLMs completely suck at this. They really do not work. If you try to train an LLM to learn good representations of images or video, they're really not that great. Generally, vision capabilities for AI systems are trained separately. They're not part of the whole LLM thing. So, if you want to handle data that is high-dimensional, continuous, and noisy, you cannot use generative models. You certainly cannot use generative models that tokenize your data into discrete symbols. There's no way. We have a lot of empirical evidence that this simply doesn't work very well. What does work is learning an abstract representation space that eliminates a lot of details about the input, essentially all the details that are not predictable, which includes noise, and make predictions in that representation space. This is the idea of GEPA, Joint Embedding Predictive Architectures, which you are familiar with.
正如你在这方面的工作。是的。
As you worked on this. Yeah.
关于这个有很多想法。让我说说我的历史。我长期以来(大概 20 年)一直相信,构建智能系统的正确方式是通过某种形式的无监督学习。我在 21 世纪初开始研究无监督学习,作为取得进展的基础。在此之前,我并不那么确信这是正确的道路。基本上,这是训练自编码器来学习表征的想法。你有一个输入,通过编码器得到表征,然后解码,这样就能保证表征包含输入的所有信息。这个直觉是错误的。坚持表征包含输入的所有信息是个坏主意。我当时不知道这一点。所以我研究了多种方法。当时杰夫·辛顿在研究受限玻尔兹曼机。约书亚·本吉奥在研究去噪自编码器,这在不同领域(包括 NLP)中变得相当成功。而我研究的是稀疏自编码器。如果你训练自编码器,需要对表征进行正则化,防止它简单地学习恒等函数。这就是信息瓶颈。你需要创建一个信息瓶颈来限制表征的信息量。我认为高维稀疏表征是个好方法。所以我的很多学生都以此做博士论文。科雷·卡武克丘奥卢,现在是 Alphabet 旗下 DeepMind 的首席架构师兼 CTO,就是和我一起做的博士论文。还有马凯尔·翁扎托、伊恩·古德费洛等人。这就是当时的想法。我们研究这个是因为想通过预训练自编码器来预训练非常深的神经网络。我们以为这是正确的道路。但后来我们开始尝试归一化、整流(而不是双曲正切或 sigmoid)以及梯度等。这最终让我们能够完全有监督地训练相当深的网络,也就是自监督学习。与此同时,数据集开始变大。结果发现,有监督学习效果很好。自监督或无监督学习的整个想法就被搁置了。然后 ResNet 出现了,在 2015 年完全解决了训练非常深架构的问题。但在 2015 年,我再次开始思考如何推动实现人类水平的 AI,这实际上是 FAIR 的原始目标和我的人生使命。我意识到所有强化学习之类的方法基本上都无法扩展。强化学习在样本效率上极其低下。所以这不是正确的道路。世界模型的想法——一个能预测其行动后果并进行规划的系统——我在 2015-16 年左右开始认真研究。我在 2016 年 NIPS(当时还叫 NIPS)的主题演讲就是关于世界模型的。我主张这个方向。那基本上是我演讲的核心:我们应该研究世界模型,行动条件模型。我的一些学生开始研究视频预测等。我们在 2016 年发表了一些关于视频预测的论文。我犯了和以前一样的错误,也是现在每个人都在犯的错误:训练视频预测系统在像素级别进行预测。这实际上是不可能的。你无法在视频帧空间上表示有用的概率分布。所以这些方法行不通。我清楚地知道,由于预测是非确定性的,我们必须有一个带潜变量的模型来表示你不知道的关于待预测变量的所有信息。我们为此研究了多年。我有个学生,现在是 FAIR 的科学家,米凯尔·安德拉韦斯,他开发了一个带潜变量的视频预测系统。它稍微解决了我们面临的问题。如今,很多人采用的解决方案是扩散模型,这本质上是一种训练非确定性函数的方法,或者基于能量的模型,我提倡了几十年,这也是训练非确定性函数的另一种方法。
So, there's a lot of ideas around this. Let me tell you my history around it. I've been convinced for a long time, probably the better part of 20 years, that the proper way to building intelligent systems was through some form of unsupervised learning. I started working on unsupervised learning as the basis for making progress in the early 2000s, mid-2000s. Before that, I wasn't so convinced this was the way to go. Basically, this was the idea of training autoencoders to learn representations. You have an input, you run it through an encoder, it finds a representation, and then you decode, so you guarantee that the representation contains all the information about the input. That intuition is wrong. Insisting that the representation contains all the information about the input is a bad idea. I didn't know this at the time. So, what I worked on was several ways of doing this. Jeff Hinton at the time was working on restricted Boltzmann machines. Yoshua Bengio was working on denoising autoencoders, which actually became quite successful in different contexts, for NLP among others. And I was working on sparse autoencoders. If you train an autoencoder, you need to regularize the representation so that the autoencoder does not trivially learn an identity function. This is the information bottleneck. You need to create an information bottleneck to limit the information content of the representation. I thought high-dimensional sparse representations was actually a good way to go. So, much of my students did their PhD on this. Koray Kavukcuoglu, who's now a chief architect at DeepMind at Alphabet, and also the CTO at DeepMind, actually did his PhD on this with me. And a few other folks, Macaire Onzato and Ian Goodfellow and a few others. So, this was kind of the idea. The reason why we worked on this was because we wanted to pre-train very deep neural nets by pre-training those things as autoencoders. We thought that was the way to go. What happened though was that we started experimenting with things like normalization, rectification instead of hyperbolic tangent or sigmoids. The gradients. And that ended up basically allowing us to train fairly deep networks completely supervised, so self-supervised learning. This was at the same time that datasets started to get bigger. So, it turned out that supervised learning worked fine. The whole idea of self-supervised or unsupervised learning was put aside. Then came ResNet, and that sort of completely solved the problem of training very deep architectures in 2015. But then in 2015, I started thinking again about how to push towards human-level AI, which really was the original objective of FAIR and my life mission. I realized that all the approaches of reinforcement learning and things of that type were basically not scaling. Reinforcement learning is incredibly inefficient in terms of samples. So, this was not the way to go. The idea of world models, a system that can predict the consequences of its actions and can plan, I started really seriously playing with this around 2015-16. My keynote at what was still called NIPS in 2016 was on world models. I was arguing for it. That was basically the centerpiece of my talk: this is what we should be working on, world models, action-condition. A few of my students started working on this on video prediction and things like that. We had some papers on video prediction in 2016. I made the same mistake as before and the same mistake that everybody is doing at the moment, which is training a video prediction system to predict at the pixel level. This is really impossible. You can't really represent useful probability distributions on the space of video frames. So, those things don't work. I knew for a fact that because the prediction was non-deterministic, we had to have a model with latent variables to represent all the stuff you don't know about the variable you're supposed to predict. We experimented with this for years. I had a student here who's now a scientist at FAIR, Mikael Andrawes, who developed a video prediction system with latent variables. It kind of solved those problems we were facing slightly. Today, the solution that a lot of people are employing is diffusion models, which is a way to train a non-deterministic function essentially, or energy-based models, which I've been advocating for decades now, which also is another way of training non-deterministic functions.
但最终我发现,这完全是个坏主意。绕过无法在像素级别进行预测这一事实的方法,就是干脆不在像素级别预测。而是学习一个表征,然后在表征层面进行预测,消除所有无法预测的细节。我早期并没有认真考虑这些方法,因为我认为存在一个防止坍缩的巨大问题。我确信 Rondel 提到过这一点,但当你训练时,假设你有一个观测变量 X,试图预测变量 Y,但你不想预测所有细节。所以你将 X 和 Y 都通过编码器,得到 X 的表征和 Y 的表征。你可以训练一个预测器,从 X 的表征预测 Y 的表征。但如果你想同时端到端地训练整个系统,存在一个平凡解:系统忽略输入,产生恒定的表征,预测器的问题就变得简单了。如果你的唯一标准是最小化预测误差,那行不通,它会坍缩。我很早就知道这个问题,因为我在 90 年代研究过联合嵌入架构,我们当时称之为孪生网络。
Um but in the end, I discovered that this was all a bad idea, that the way to get around the fact that you can't predict at the pixel level is to just not predict at the pixel level. It's to learn a representation and predict at the representation level, eliminating all the details you cannot predict. I wasn't really thinking about those methods early on because I thought there was a huge problem of preventing collapse. I'm sure Rondel talked about this, but when you train, say you have an observed variable X, and you're trying to predict a variable Y. But you don't want to predict all the details. So you run both X and Y through encoders. Now you have a representation for X and a representation for Y. You can train a predictor to predict the representation of Y from the representation of X. But if you want to train this whole thing end-to-end simultaneously, there's a trivial solution where the system ignores the input and produces constant representations. The predictor's problem now is trivial. If your only criterion is to minimize the prediction error, it's not going to work. It's going to collapse. I knew about this problem from a very long time because I worked on joint embedding architectures, we used to call them Siamese networks back in the '90s.
是一样的,因为人们最近还在用孪生网络这个术语。
Those are the same because people have been using that term Siamese networks even recently.
没错。这个概念仍然不过时。所以你有 X 和 Y,把 X 看作是 Y 的某种退化、变换或损坏版本。你将 X 和 Y 都通过编码器,然后告诉系统:‘看,X 和 Y 是同一事物的两个视图,所以无论你计算出什么表征,都应该是相同的。’如果你只是训练一个神经网络,两个共享权重的神经网络,为同一物体的略微不同版本产生相同表征,它会坍缩,不会产生任何有用的东西。所以你必须找到一种方法,确保系统从输入中提取尽可能多的信息。我们最初的想法,来自 1993 年那篇关于孪生网络的论文,是加入一个对比项。你有其他已知不同的样本对,训练系统产生不同的表征。所以有一个代价函数,当你展示两个相同或相似的样本时,它会吸引两个表征;当你展示两个不同的样本时,它会排斥。我们想到这个主意,是因为有人来找我们,问:‘你能把某人在平板电脑上签名的签名编码成少于 80 字节吗?’因为如果能编码成少于 80 字节,就可以写在信用卡的磁条上,从而实现信用卡签名认证。所以我们想出了这个主意:训练一个神经网络产生 80 个变量,每个量化成一个字节,然后训练它做同样的事。
That's right. The concept is still up-to-date. So you have an X and a Y, and think of X as some sort of degraded, transformed, or corrupted version of Y. You run both X and Y through encoders, and you tell the system, 'Look, X and Y are two views of the same thing. So whatever representation you compute should be the same.' If you just train a neural net, two neural nets with shared weights, to produce the same representation for slightly different versions of the same object, it collapses. It doesn't produce anything useful. So you had to find a way to make sure the system extracts as much information from the input as possible. The original idea we had, from a 1993 paper with Siamese net, was to have a contrastive term. You have other pairs of samples that you know are different, and you train the system to produce different representations. So you have a cost function that attracts the two representations when you show two examples that are identical or similar, and repels when you show two examples that are dissimilar. We came up with this idea because someone came to us and said, 'Can you encode signatures of someone drawing a signature on a tablet? Can you encode this in less than 80 bytes?' Because if you can encode it in less than 80 bytes, we can write it on the magnetic tape of a credit card, so we can do signature authentication for credit cards. So we came up with this idea of training a neural net to produce 80 variables that were quantized one byte each, and then training it to do the same.
他们用了吗?
And did they use it?
效果很好,他们拿给业务人员看,业务人员说:‘哦,我们打算让人们输入 PIN 码。’我们甚至没有那种整合技术的方式。我一开始就觉得这事有点不靠谱,因为欧洲有些国家已经在用智能卡了,而且有更好的解决方案。但他们出于某种原因就是不想用智能卡。总之,我们在 2000 年代中期有了这项技术。我和两个学生一起改进了这个想法,提出了一些新的目标函数来训练这些网络。这就是现在人们所说的对比方法。这是对比方法的一个特例。我们有正样本和负样本,训练正样本时让能量低,负样本时让能量高,这里的能量是表征之间的距离。我在 2005-2006 年的 CVPR 上发表了两篇论文,作者是 Raia Haddad Sal(现任 DeepMind 基金会负责人,DeepMind 类似公平部门)和 Sumit Chopra(现在纽约大学任教,从事医学影像研究)。这在社区引起了一些兴趣,某种程度上复兴了这些想法。但效果仍然不太好。这些对比方法产生的图像表征维度相对较低。如果我们测量对比矩阵的特征值谱和产生的表征,最多也就 200 维,从不超过。即使在 ImageNet 等数据集上训练,加上数据增强,也是如此。这有点令人失望。效果还行。有一堆相关论文。效果还行。DeepMind 有一篇论文 SimCLR 展示了通过对比训练应用于 SimSiam 可以获得不错的性能。但大约 5 年前,我在 Meta 的博士后 Stefan Deny 尝试了一个想法,我起初认为行不通:本质上是对编码器输出的信息量进行某种度量,然后试图最大化它。我认为行不通的原因是,我见过 Geoff Hinton 在 80 年代做的很多类似实验,试图最大化信息。你永远无法最大化信息,因为你没有合适的信息含量度量。也就是说,你需要一个下界。如果你想最大化某个东西,你要么能直接计算它,要么有一个下界可以推高。而对于信息含量,我们只有上界。我一直认为这完全无望。然后 Stefan 提出了一种叫做 Barlow Twins 的技术。Barlow 是一位著名的理论神经科学家,提出了信息最大化的想法。它居然有效。真是令人惊叹。于是我说,我们必须推进这个方向。所以我们和 Simon、Adrien Bard 以及 Jean Pons(也隶属于纽约大学)一起提出了另一种方法,叫做 VICReg,即方差-不变性-协方差正则化。
It worked really well, and they showed it to their business people who said, 'Oh, we're just going to ask people to type PIN codes.' We have even less than that, like how you can integrate that technology. I knew this thing was kind of fishy in the first place because there were countries in Europe that were using smart cards, and there was a much better solution. But they just didn't want to use smart cards for some reason. Anyway, so we had this technology in the mid-2000s. I worked with two of my students to revise this idea. We came up with some new objective functions to train those. These are what people now call contrastive methods. This is a special case of contrastive methods. We have positive examples, negative examples, and you train on positive examples to have low energy, and for negative samples, you train them to have higher energy, where energy is the distance between the representations. I had two papers at CVPR in 2005-2006 by Raia Haddad Sal, who is now the head of DeepMind Foundation, the fair-like division of DeepMind, and Sumit Chopra, who is actually a faculty here at NYU now, working on medical imaging. This gathered a bit of interest in the community and sort of revived a little bit of work on these ideas. But it still wasn't working very well. Those contrastive methods were producing representations of images, for example, that were kind of relatively low-dimensional. If we measured the eigenvalue spectrum of the contrast matrix and the representations that came out of those things, it would fill up maybe 200 dimensions, never more. Even training on ImageNet and things like that, even with data augmentation. That was kind of disappointing. It worked okay. There was a bunch of papers on this. It worked okay. There was one paper from DeepMind, SimCLR, that demonstrated you could get decent performance with contrastive training applied to SimSiam. But then, about 5 years ago, one of my postdocs, Stefan Deny at Meta, tried an idea that at first I didn't think would work, which was to essentially have some measure of information quantity that comes out of the encoder, and then trying to maximize that. The reason I didn't think it would work is because I'd seen a lot of experiments along those lines that Geoff Hinton was doing in the 1980s, trying to maximize information. You can never maximize information because you never have appropriate measures of information content. That is, a lower bound. If you want to maximize something, you want to either be able to compute it or you want a lower bound on it so you can push it up. For information content, we only have upper bounds. I always thought this was completely hopeless. And then Stefan kind of came up with a technique called Barlow Twins. Barlow is a famous theoretical neuroscientist who came up with the idea of information maximization. And it kind of worked. That was wow. So then I said, we have to push this. So we came up with another method with Simon, Adrien Bard, and Jean Pons, who is affiliated with NYU too. A technique called VICReg, variance-invariance-covariance regularization.
嗯,结果证明它更简单,效果也更好。从那以后,我们取得了进展,Randall 最近,你知道,我和他讨论了一个想法,他推动并使之实用化。它叫 SimReg。整个系统叫 Jepa。好吧,名字是他起的。我不知道。潜欧几里得 Jepa,对吧?嗯。SimReg 涉及到确保编码器输出的向量分布是各向同性高斯分布。这就是 G 中的 I。嗯。所以,这个领域有很多事情正在发生,真的很酷。我认为未来一两年还会有更多进展。我们在这方面积累了很多经验。我认为这是一套非常有前途的技术,用于训练学习抽象表示的模型,我认为这是关键。
Uh and that turned out to be simpler and work even better. And since then, we've made progress and Randall recently, you know, I discussed an idea with him that he kind of pushed and made practical. It's called SimReg. The whole system is called Jepa. Okay, he's responsible for the name. I don't know. The latent Euclidean Jepa, right? Yeah. Um, and SimReg has to do with sort of making sure that the distribution of vectors that come out of the encoder is an isotropic Gaussian. That's the I in the G. Mhm. So, I mean, there's a lot of things happening in this domain, which are really cool. I think there's going to be some more progress over the next year or two. We get a lot of experience with this. And I think that's a really good promising set of techniques to train models that learn abstract representations, which I think is key.
那你觉得这里缺失的部分是什么?比如,你认为更多的算力会有帮助,还是我们需要更好的算法?你相信苦涩的教训吗?还有,你怎么看 2022 年后互联网的数据质量问题?我听到有人把它比作低背景钢,指的是 LLM 出现之前的所有数据。我是说,低背景 token。
And what do you think are the missing parts here? Like, do you think more compute will help or we need better algorithms? Do you believe in the bitter lesson? And furthermore, what do you think about the data quality problems with the internet post-2022? I've heard people compare it to low-background steel now, to refer to all that data before LLMs came out. Like, low-background tokens, I mean.
是的。我认为我完全避开了那个问题。好吧,事情是这样的。过去几年我一直在公开使用这个论点。训练一个 LLM,如果你想要任何像样的性能,就需要在互联网上所有可用的免费文本上训练,再加上一些合成数据、授权数据等等。所以,一个典型的 LLM,比如,你知道,一两年前的第三名,是在 30 万亿个 token 上训练的。一个 token 通常是三个字节。所以,预训练是 10 的 13 次方字节。好吧,我们不是在讨论微调。10 的 14 次方字节。为了让 LLM 真正利用这些数据,它们需要大量的存储空间。因为基本上,这些都是孤立的事实。文本中有一点冗余,但很多只是孤立的事实,对吧?所以你需要非常大的网络,因为你需要大量内存来存储所有这些事实。然后我们再把它们复述出来。好吧,现在把这个和视频比较一下。10 的 14 次方字节,如果你按每秒 2 兆字节算视频,相对压缩的视频,不是高度压缩,但有一点压缩。那相当于 15,000 小时的视频。10 的 14 次方字节。如果是 15,000 小时的视频,你就有了和互联网上所有文本一样多的数据。现在,15,000 小时的视频根本不算什么。那是 YouTube 上传 30 分钟的量。好吧。这是一个 4 岁孩子一生中看到的视觉信息量。整个生命,清醒时间,4 年大约是 16,000 小时。这不是很多信息。我们现在有视频模型,VJepa,VJepa 2,实际上,去年夏天刚出来。它是在相当于一个世纪的视频数据上训练的。那可能还是数据。更多的数据。但实际上比最大的 LLM 少得多。因为尽管字节更多,但冗余也更多。所以你会说,“好吧,冗余更多,所以用处更小。”实际上,当你使用自监督学习时,你确实需要冗余。顺便说一句,如果数据完全是随机的,你无法通过自监督或任何方式学到任何东西。冗余才是你能学到的东西。所以,像视频这样的现实世界数据中的结构比文本丰富得多,这让我声称我们绝对永远无法仅通过训练文本达到人类水平的 AI。这永远不会发生,对吧?
Yeah. I think I'm totally escaping that problem. Okay, here is the thing. And I've been using this argument publicly over the last couple of years. Training an LLM, if you wanted to have any kind of decent performance, requires training on basically all the available freely available text on the internet, plus some synthetic data, plus licensed data, etc. So, a typical LLM, like, you know, number three, going back a year or two, is trained on 30 trillion tokens. A token is typically three bytes. So, that's 10 to the 13 bytes for pre-training. Okay, we're not talking about fine-tuning. 10 to the 14 bytes. And for the LLMs to be able to really exploit this, they need to have a lot of memory storage. Because basically, those are isolated facts. There is a little bit of redundancy in text, but a lot of it is just isolated facts, right? So, you need very big networks because you need a lot of memory to store all those facts. And we regurgitate them. Okay, now compare this with video. 10 to the 14 bytes, if you count 2 megabytes per second for video, for relatively compressed video, not highly compressed, but a bit. That would represent 15,000 hours of video. 10 to the 14 bytes. If 15,000 hours of video, you'll have the same amount of data as the entirety of all the text available on the internet. Now, 15,000 hours of video is absolutely nothing. It's 30 minutes of YouTube uploads. Okay. It's the amount of visual information that a 4-year-old has seen in his or her life. The entire life, waking time, is about 16,000 hours in 4 years. It's not a lot of information. We have video models now, VJepa, VJepa 2, actually, that just came out last summer. That was trained on the equivalent of a century of video data. It's still probably data. Much more data. But much less than the biggest LLM, actually. Because even though it's more bytes, it's more redundant. So, you say, "Okay, it's more redundant, so it's less useful." Actually, when you use self-supervised learning, you do need redundancy. You cannot learn anything in self-supervised or anything, by the way, if it's completely random. Redundancy is what you can learn. And so, there's just much richer structure in real-world data like video, than there is in text, which kind of led me to claim that we absolutely never ever going to get to human-level AI by just training on text. It's just never going to happen, right?
嗯,当我们谈论世界模型和具身时,我认为仍然有很多人甚至不理解理想化的世界模型是什么,在某种意义上,对吧?例如,我受到看《星际迷航》的影响,我希望你看过一点,你在想全息甲板,对吧?我一直认为全息甲板就像一个理想化的完美世界模型,对吧?即使有很多集都过头了,对吧?人们从里面走出来,对吧?但它甚至模拟了气味和身体接触之类的东西。所以,你认为像那样的东西是理想化的世界模型,还是你认为不同的模型或定义方式会是?
Well, and when we talk about world models and grounding, I think there's still a lot of people who don't even understand what the idealized world model is, in a sense, right? So, for example, I'm influenced by having watched Star Trek, which I would hope you've seen a little bit of, and you're thinking of the Holodeck, right? I always thought that the Holodeck was like an idealized perfect world model, right? Even so many episodes of going too far, right? People walking out of it, right? But it even simulates things like smell and physical touch. So, do you think that something like that is the idealized world model or do you think a different model or way of defining it would be?
好吧,这是一个极好的问题。它之所以极好,是因为它触及了核心,即我认为我们应该做什么,我们能做什么,以及我认为其他人都错得有多离谱,好吧?所以,人们认为世界模型是再现世界所有细节的东西。他们把它看作一个模拟器。是的。对吧?当然,因为深度学习是主流,你会用某个深度学习系统作为模拟器。很多人也专注于视频生成,这很酷,对吧?你制作那些酷炫的视频,人们印象深刻。但是,当你训练一个视频生成系统时,没有任何保证它实际上拥有世界底层动态的准确模型,并且学到了任何特别抽象的东西。嗯。所以,认为模型需要再现现实的每一个细节的想法是错误的,而且是有害的。嗯。我来告诉你为什么,好吧?模拟的一个好例子是 CFD,计算流体动力学。它经常被使用。人们用超级计算机来做这个,对吧?所以,你想模拟飞机周围的空气流动,你把空间切成小立方体,在每个立方体内,你有一个小向量来表示该立方体的状态,即速度、密度或质量、温度,可能还有其他一些东西,对吧?然后你求解纳维-斯托克斯方程,这是一个偏微分方程。你可以模拟空气流动。问题是,这实际上不一定能非常准确地求解方程。如果你有像湍流这样的混沌行为,模拟只是近似正确。
Okay, this is an excellent question. And the reason it's excellent is because it goes to the core of really what I think we should be doing, we can do, and how wrong I think everybody else is, okay? So, people think that a world model is something that reproduces all details of what the world does. They think of it as a simulator. Yeah. Right? And of course, because deep learning is the thing, you're going to use some deep learning system as a simulator. A lot of people also are focused on video generation, which is kind of a cool thing, right? You produce those cool videos and they're wow, people are really impressed by them. Now, there's no guarantee whatsoever that when you train a video generation system, it actually has an accurate model of the underlying dynamics of the world and it's learned anything particularly abstract about it. Mhm. So, the idea that somehow a model needs to reproduce every detail of reality is wrong and hurtful. Mhm. And I'm going to tell you why, okay? A good example of simulation is CFD, computational fluid dynamics. It's used all the time. People use supercomputers for that, right? So, you want to simulate the flow of air around an airplane, you cut up the space into little cubes, and within each cube, you have a small vector that represents the state of that cube, which is velocity, density or mass, and temperature, and maybe a couple of other things, right? So, and then you solve Navier-Stokes equations, which is a partial differential equation. And you can simulate the flow of air. Now, the thing is this does not actually necessarily solve the equations very accurately. If you have chaotic behavior like turbulence and stuff like that, simulation is only approximately correct.
但实际上,那已经是底层现象的抽象表示了。底层现象是空气分子相互碰撞,并撞击机翼和飞机,对吧?但没人会到那个层面去模拟。那太疯狂了。所需的计算量简直离谱,而且还会依赖于初始条件。我们不这么做的原因有很多。也许不是分子,也许在更底层我们应该模拟粒子,做费曼图,模拟粒子所经历的所有不同路径,因为它们并不只走一条路。这不是经典物理,而是量子物理。所以最底层是量子场论,而很可能那已经是底层现实的抽象表示了。因此,我们之间此刻发生的一切,原则上都可以用量子场论来描述。我们只需要测量包含我们所有人的一个立方体中的宇宙波函数,但即使那样也不够,因为宇宙另一端还有纠缠粒子。所以那还不够。但为了论证,我们假设一下。首先,我们无法测量这个波函数。其次,所需的计算量极其巨大。那将是一个像地球一样大的巨型量子计算机。所以我们根本不可能在那个层面描述任何东西。而且我们的模拟很可能只能精确几纳秒。之后就会与现实偏离。那我们怎么办?我们发明抽象概念。我们发明像粒子、原子、分子这样的抽象概念。在生命世界中,有蛋白质、细胞器、细胞、器官、生物体、社会、生态系统等等。基本上,这个层级中的每一层都忽略了很多下一层的细节。这让我们能够做出更长期、更可靠的预测。所以我们可以用基础科学和心理学来描述我们之间此刻的互动。那比粒子物理的抽象层次高得多。事实上,我刚才提到的每一个层级都是一个不同的科学领域。一个科学领域本质上是由你开始做出预测的抽象层次定义的。物理学家对此已经炉火纯青。如果我给你一盒气体,你原则上可以模拟所有气体分子,但没人会这么做。在非常抽象的层面,我们可以说 PV = nRT。压力乘以体积等于粒子数乘以温度等等。所以在全局涌现现象层面,如果你增加压力,温度会上升。或者如果你增加温度,压力会上升。或者如果你放出一些粒子,压力会下降。所以我们一直在通过忽略各种细节来构建复杂事物的现象学模型,物理学家把这些细节称为熵。但这确实是系统性的。这就是我们理解世界的方式。我们不会记住每一个细节,当然也不会重构我们所感知的一切。所以世界模型不一定是模拟器。
But in fact, that's already an abstract representation of the underlying phenomenon. The underlying phenomenon is molecules of air that bump into each other and bump on the wing and on the airplane, right? But nobody ever goes to that level to do the simulation. That would be crazy. It would require an amount of computation that's just insane. And it would depend on the initial condition. There are all kinds of reasons we don't do this. And maybe it's not molecules. Maybe at a lower level we should simulate particles, do Feynman diagrams and simulate all the different paths those particles are employing because they don't take one path. It's not classical, it's quantum. So at the bottom it's like quantum field theory and probably that is already an abstract representation of the underlying reality. So everything that takes place between us at the moment, in principle can be described through quantum field theory. We just have to measure the wave function of the universe in a cube that contains all of us, and even that would not be sufficient because there are entangled particles at the other side of the universe. So it wouldn't be sufficient. But let's imagine, for the sake of the argument. First, we would not be able to measure this wave function. Second, the amount of computation we would need to devote to this is absolutely gigantic. It would be some gigantic quantum computer the size of the earth or something. So no way we can describe anything at that level. And it's very likely that our simulation would be accurate for maybe a few nanoseconds. Beyond that we'll diverge from reality. So what do we do? We invent abstractions. We invent abstractions like particles, atoms, molecules. In the living world, it's proteins, organelles, cells, organs, organisms, societies, ecosystems, etc. Basically every level in this hierarchy ignores a lot of details about the level below. And what that allows us to do is make longer-term, more reliable predictions. So we can describe the dynamics between us now in terms of the underlying science and in terms of psychology. That's a much higher level of abstraction than particle physics. In fact, every level in the hierarchy I just mentioned is a different field of science. A field of science is essentially defined by the level of abstraction at which you start making predictions. Physicists have this down to an art. If I give you a box full of gas, you could in principle simulate all the molecules of the gas, but nobody ever does this. At a very abstract level, we can say PV = nRT. Pressure times volume equals number of particles times temperature, etc. So at a global emergent phenomenological level, if you increase the pressure, the temperature will go up. Or if you increase the temperature, the pressure will go up. Or if you let some particles out, the pressure will go down. So we all the time build phenomenological models of something complicated by ignoring all kinds of details that physicists call entropy. But it's really systematic. That's the way we understand the world. We do not memorize every detail, we certainly don't reconstruct it of what we perceive. So world models don't have to be simulators.
不。
No.
嗯,它们是模拟器,但在抽象表示空间中。它们模拟的只是现实中相关的部分。如果我问你,木星 100 年后会在哪里?我们有大量关于木星的信息,但要做出那个预测,你只需要六个数字:三个位置和三个速度。其他的都不重要。
Well, they are simulators, but in abstract representation space. And what they simulate is only the relevant part of reality. If I ask you, where is Jupiter going to be 100 years from now? We have an enormous amount of information about Jupiter, but to make that prediction, you need exactly six numbers: three positions and three velocities. The rest doesn't matter.
所以你不相信合成数据集?
So you don't believe in synthetic data sets?
我相信。不,它很有用。来自游戏的数据。你确实能从游戏等合成数据中学到很多东西。孩子们从玩耍中学到大量知识,这有点像对世界的模拟,但在不会伤害到自己的条件下。
I do. No, it's useful. Data from games. There's certainly a lot of things that you learn from synthetic data from games and things like that. Children learn a huge amount from play, which are kind of simulations of the world a little bit, but in conditions where they can't kill themselves.
但我担心至少对于视频游戏来说,比如演员在绿幕前做动画,它们被设计成在动作游戏中看起来很棒,但往往与现实不太相符。所以我担心一个借助世界模型训练的物理系统可能会有类似的怪癖,至少在短期内。这让你担心吗?
But I worry at least for video games that, for example, the green screen like actors doing the animations, they're designed to look good for an action game, but often don't correspond very well to reality. So I worry that a physical system trained with the assistance of world models might get similar quirks, at least in the very short term. Is this something that worries you?
不,这取决于你在什么层面训练它们。例如,如果你使用一个非常精确的机器人模拟器,它能够准确模拟手臂在施加扭矩时的动力学。动力学没问题。但模拟抓取和操作物体时产生的摩擦力,那非常难以精确做到。摩擦力很难模拟。所以那些模拟器对于操作来说并不特别精确。它们足够好,你可以训练一个系统去做,然后通过一点适应进行 sim-to-real。所以那是可行的。但重点更重要。世界上有很多完全基本的东西,我们完全视为理所当然,我们可以在非常抽象的层面学习它们,但这与语言无关。例如,我之前用过这个例子,人们还取笑我,但这是真的。我桌子上有那些物体。当我推桌子时,物体跟着移动。这是我们学到的。这不是与生俱来的。大多数物体在你松手时会下落,因为重力。这大概是在 9 个月大时学到的。
No, it depends at what level you train them. For example, if you use a very accurate robotic simulator, it's going to accurately simulate the dynamics of an arm when you apply torques to it. There's dynamics, no problem. Now, simulating the friction that happens when you grab an object and manipulate it, that's super hard to do accurately. Friction is very hard to simulate. So those simulators are not particularly accurate for manipulation. They're good enough that you can train a system to do it and then do sim-to-real with a little bit of adaptation. So that can work. But the point is much more important. There are a lot of completely basic things about the world that we completely take for granted, which we can learn at a very abstract level, but it's not language related. For example, and I've used this example before and people have made fun of me for it, but it's really true. I have those objects on the table. The fact that when I push the table, the object moves with it. This is something we learned. It's not something that you're born with. The fact that most objects will fall when you let them go, with gravity. Maybe it's learned around the age of 9 months.
人们因此取笑我,是因为我说过 LLM 不懂这类东西,对吧?即使到今天,它们也完全不懂。但你可以训练它们在回答问题时给出正确答案。比如,如果我把一个物体放在桌子上,然后推桌子,物体会怎样?它会回答物体跟着移动。但那是因为它被微调过。所以这更像是复述,而不是真正理解背后的动力学。
And the reason people make fun of me with this is because I said LLMs don't understand this kind of stuff, right? And they absolutely do not even today. But you can train them to give the right answer when you ask them a question. For example, if I put an object on a table and I push the table, what will happen to the object? It will answer that the object moves with it. But that's because it's been fine-tuned to do that. So it's more like regurgitation than real understanding of the underlying dynamics.
但如果你看像 Sora 这样的模型,它们有很好的世界物理,对吧?虽然不完美,但……
But if you look at models like Sora, they have a good physics of the world, right? They are not perfect, but...
那不是物理,是的。它们有一些物理。但你认为我们不能更进一步,还是说这是学习物理的一种方式?所有这些模型实际上都是在表示空间中进行预测。它们使用扩散 Transformer。这种预测——在抽象层面完成视频片段——是在表示空间中进行的。然后有第二个扩散模型将这些抽象表示转化为好看的视频。那可能是模式崩溃。我们不知道,因为我们无法真正衡量这类系统对现实的覆盖程度。
That's not physics, yeah. They have some physics. But do you think we can't push it farther, or do you think it's one way to learn physics? So, all those models actually make predictions in representation space. They use diffusion transformers. And that prediction—the completion of the video snippet at an abstract level—is done in representation space. Then there's a second diffusion model that turns these abstract representations into a nice-looking video. And that might be mode collapse. We don't know, because we can't really measure the coverage of such systems with reality.
回到之前的观点,我可以训练——这里还有一个对我们来说完全显而易见的概念,我们甚至不认为我们学过,但我们确实学了。一个人不能同时出现在两个地方。我们学到这一点是因为很早我们就学会了物体恒存性:当一个物体消失时,它仍然存在。而且我们把它视为之前见过的同一个物体。我们如何训练 AI 系统学习这个概念?物体恒存性:你只需给它看大量视频,其中物体移动到屏幕后面然后从另一边出现,或者物体移到屏幕后面,屏幕移开,物体还在那里。当你给 4 个月大的婴儿展示违反这种情况的场景时,他们的眼睛睁得大大的,非常惊讶,因为现实违背了他们的内部模型。同样,当你展示一个小车在平台上,把它推下平台,它似乎浮在空中时,9 个月、10 个月大的婴儿会非常惊讶地看着。6 个月大的婴儿几乎不注意,因为他们还没学会重力。所以他们还没能融入每个物体都应该掉落的观念。所以这种学习才是真正重要的。你可以从非常抽象的东西中学到这一点,就像婴儿通过听带简单图片的故事学习社交互动一样。这是一种模拟——世界的抽象模拟。但它能学到特定的行为。
To the previous point, I can train—here is another completely obvious concept to us that we don't even imagine we learn, but we do learn it. A person cannot be in two places at the same time. We learn this because very early on we learn object permanence: the fact that when an object disappears, it still exists. And we perceive it as the same object that you saw before. How can we train an AI system to learn this concept? Object permanence: you just show it a lot of videos where objects go behind a screen and then reappear on the other side, or where they go behind the screen and the screen goes away and the object is still there. When you show 4-month-old babies scenarios where things like this are violated, their eyes open super big and they're super surprised because reality just violated their internal model. Same thing when you show a scenario of a little car on a platform, you push it off the platform and it appears to float in the air. 9-month, 10-month-old babies look at it really surprised. 6-month-old babies barely pay attention because they haven't learned about gravity yet. So they haven't been able to incorporate the notion that every object is supposed to fall. So this kind of learning is really what's important. And you can learn this from very abstract things, the same way babies learn about social interactions by being told stories with simple pictures. It's a simulation—an abstract simulation of the world. But it sort of learns particular behavior.
所以你可以想象从冒险游戏训练系统,比如俯视 2D 冒险游戏。你告诉角色向北移动,他就去了另一个房间,他不在第一个房间了,因为他移动到了另一个房间。当然,在冒险游戏中你可以召唤甘道夫,他直接出现,所以那不符合物理。但当你从宝箱中拿起一把钥匙,你拥有了钥匙,别人不能拥有,你可以用它开门。有很多非常基础的东西可以学到,即使在抽象环境中也是如此。
So you could imagine training a system from, let's say, an adventure game, like a top-down 2D adventure game. You tell your character to move north and he goes to the other room, and he's not in the first room anymore because he moved to the other room. Of course, in adventure games you have Gandalf that you can call and he just appears, so that's not physical. But when you pick up a key from a treasure chest, you have the key, no one else can have it, and you can use it to open a door. There's a lot of things that you learn that are very basic, even in abstract environments.
我想指出,他们试图训练模型的一些冒险游戏——你可能知道其中一个叫 NetHack。NetHack 非常迷人,因为它异常困难。不用作弊通关这个游戏,就像 20 年不看攻略一样。人们光靠玩还是做不到。据我所知,AI 智能体,我们最好的智能体模型甚至世界模型,表现都很糟糕。
And I just want to observe that some of those adventure games that they try to train models on—one of them you might know about is NetHack. NetHack is fascinating because it is an extraordinarily hard game. Ever ascending in that game without cheats is like 20 years without going to the wiki. People still don't do it from playing. And my understanding is that AI agents, the very best agent models we have or even world models, are pathetic.
完全同意。所以,人们提出了 NetHack 的简化版,叫 MiniHack。他们不得不为 AI 智能体简化它。我的一些同事一直在研究它——实际上我的一名硕士生正在和我一起做。还有我之前提到的同事 NAF,也在那里做了一些工作。有趣的是,有些情况需要规划,但需要在不确定性下规划。问题在于,所有游戏,尤其是冒险游戏,你无法完全看到系统状态。你事先不知道地图。你需要探索,每次探索都可能死亡。但动作本质上是离散的。有限数量的可能动作是回合制的。所以在这个意义上它像国际象棋,但国际象棋是完全可观察的。围棋也是完全可观察的。但 Stratego 不是,扑克也不是。所以如果有不确定性,就更难了。但这些游戏的动作数量是离散的。基本上,你需要做的是树搜索。可能状态的树随着步数指数增长。所以你必须有一种方法只生成可能好的动作,基本上不生成其他动作或剪枝它们。你还需要一个价值函数,即使我只规划九步,也能估计一个位置是好是坏,是否会导向胜利或解决方案。所以你需要这两个组件:一个猜测好动作,一个评估终点。如果你两者都有,你可以用强化学习或行为克隆(如果有数据)来训练这些函数。这个基本思想可以追溯到 1964 年 Samuel 的跳棋程序。不是新东西。但它的威力在 AlphaGo 和 AlphaZero 等中得到了展示。
Totally. Yeah. So to the point, people have come up with a dumbed-down version of NetHack called MiniHack. They had to dumb it down just for AI agents. Some of my colleagues have been working with it—actually one of my master's students is working with me. And my colleague NAF, who I mentioned earlier, has also been doing some work there. What's interesting is that there are situations where you need to plan, but you need to plan in the presence of uncertainty. The problem is that in all games, and adventure games in particular, you don't have complete visibility of the state of the system. You don't know the map in advance. You need to explore, and you can get killed every time you do this. But the actions are essentially discrete. The finite number of possible actions is turn-based. So in that sense it's like chess, except chess is fully observable. Go is also fully observable. Stratego isn't, poker isn't. So it makes it more difficult if you have uncertainty. But those are games where the number of actions you can take is discrete. Basically, what you need to do is tree exploration. The tree of possible states goes exponentially with the number of moves. So you have to have some way of generating only the moves that are likely to be good and basically never generate the other ones or select them down. And you need to have a value function, which tells you, even though I'm planning only nine moves ahead, I have some way of estimating whether a position is good or bad, whether it will lead to a victory or a solution. So you need those two components: something that guesses what the good moves are, and something that evaluates ends. If you have both, you can train those functions using reinforcement learning or behavioral cloning if you have data. The basic idea goes back to Samuel's checker players from 1964. It's not recent. But the power of it was demonstrated with AlphaGo and AlphaZero and things like that.
但那是人类不擅长的领域。人类下棋很差,对吧?下围棋也很差。机器比我们强得多。因为树搜索的速度和所需的内存。我们没有足够的内存容量来做广度优先的树搜索。所以我们很糟糕。当 AlphaGo 出现时,之前人们认为最好的棋手可能比理想中的“神”级棋手差两到三子。不,人类很差。世界顶尖棋手可能需要让八到九子。
But that's a domain where humans suck. Humans are terrible at playing chess, right? And playing Go. Machines are much better than we are. Because of the speed of tree exploration and because of the memory required for tree exploration. We just don't have enough memory capacity to do breadth-first tree exploration. So we suck at it. When AlphaGo came out, people before that thought the best human players were maybe two or three stones handicap below an ideal player you call God. No, humans are terrible. The best players in the world need like eight or nine stones.
真不敢相信我有幸和 Yann 讨论游戏 AI。我有几个后续问题。第一个是关于人类下棋很差的例子。我听说这被称为莫拉维克悖论,解释为人类经过数十亿年进化出身体运动能力,所以婴儿和人类很擅长这个,但我们根本没有进化出下棋的能力。第二个相关问题:很多玩电子游戏的人,包括我,注意到敌人 AI 在 20 年里没有进步。一些最好的例子仍然是 2000 年代初的《光环 1》和《恐惧》。你认为实验室的进步何时会对玩家产生真正的影响,而不是生成式 AI 那种?
I can't believe I get the pleasure to talk about game AI with Yann. I just have a few follow-up questions. The first one is this example about humans being terrible at chess. I've heard this referred to as Moravec's paradox, explained as humans have evolved over billions of years for physical locomotion, so babies and humans are very good at that, but we have not evolved at all to play chess. A second related question: a lot of people who play video games, including me, have observed that enemy AI has not improved in 20 years. Some of the best examples are still Halo 1 and F.E.A.R. from the early 2000s. When do you think advancements from the lab will have a real impact on gamers, in a non-generative AI sense?
我曾经是个玩家,但没上瘾,不过我的家庭涉足其中,因为我三个三十多岁的儿子开了一家电子游戏设计工作室。所以我沉浸在那文化里。你说得对。尽管物理模拟器很精确,但很多动画电影工作室并不使用它们,因为他们想要控制,不一定是精确。在游戏里也一样。这是一种创作行为。你想要控制故事或 NPC 行为。目前 AI 很难保持控制。它会来的,但创作者有抵触。
I used to be a gamer, never addicted, but my family is in it because my three sons in their 30s have a video game design studio. So I was embedded in that culture. You're right. Despite the accuracy of physical simulators, many simulations are not used by studios making animated movies because they want control, not necessarily accuracy. In games it's the same. It's a creative act. You want control over the story or NPC behavior. AI makes it difficult to maintain control at the moment. It will come, but there's resistance from creators.
我认为莫拉维克悖论仍然非常有效。
I think Moravec's paradox is still very much in force.
如果我没记错,莫拉维克在 1988 年提出了这个悖论。他说:为什么我们认为独特的人类智力任务如下棋,可以用计算机完成,但那些我们认为理所当然的事,比如猫能做的事,却仍然不能用机器人完成。即使在 47 年后的今天,我们仍然做不好。我们可以通过模仿和强化学习训练机器人,通过模拟来移动和避开障碍,但它们远没有猫那么有创造力和敏捷。这不是因为我们造不出机器人,而是我们无法让它们足够聪明去做猫甚至老鼠能做的事,更不用说狗或猴子了。所以那些吹嘘一两年内实现 AGI 的人完全是妄想。现实世界要复杂得多。你不可能通过把世界标记化并使用 LLM 来取得任何进展。
Moravec formulated it in 1988 if I remember correctly. He said: how come things we think of as uniquely human intellectual tasks like playing chess, we can do with computers, but things we take for granted, like what a cat can do, we still can't do with robots. Even now, 47 years later, we still can't do them well. We can train robots by imitation and reinforcement learning, through simulation to locomote and avoid obstacles, but they're not nearly as inventive and agile as a cat. It's not because we can't build a robot; we just can't make them smart enough to do what a cat or even a mouse can do, let alone a dog or monkey. So all those people gloating about AGI in a year or two are completely deluded. The real world is way more complicated. You're not going to get anywhere by tokenizing the world and using LLMs.
那么你的时间线是什么?我们什么时候会看到 AGI,不管它意味着什么?你在乐观-悲观光谱上处于什么位置?有像 Gary Marcus 或 Joshua 这样的末日论者?你站在哪一边?
So what are your timelines? When will we see AGI, whatever that means? And where are you on the optimist-pessimist side? There are doomers like Gary Marcus, or Joshua? Where do you fall?
我先回答第一个问题。首先,没有通用智能这种东西。这个概念毫无意义,因为它旨在指代人类水平的智能。但人类智能是高度专业化的。我们擅长处理现实世界,导航等等。我们擅长处理人际关系,因为我们进化出了这种能力。下棋我们很差。有很多任务我们很差,而其他动物比我们强得多。所以我们是专业化的。我们认为自己是通用的,但这是一种幻觉,因为我们能理解的所有问题都是我们能想到的。有很多问题我们无法想象。这里有数学论证,除非你问,否则我不深入。所以通用智能的概念完全是胡扯。我们可以谈论人类水平的智能。我们会不会有在所有人类擅长的领域都像人类一样好的机器?我们已经有一些领域机器比人类好,比如将 1500 种语言翻译成其他 1500 种语言。没有人类能做到。下棋和围棋等等。我们会不会有在所有领域都像人类一样好的机器?绝对会。但这不会是一个事件,而是渐进的过程。未来几年,我们将基于 GPT-3 模型、规划等取得概念性进展。如果幸运且没有遇到未知障碍,这可能会通向人类水平 AI 的良好路径。但也许我们仍然缺少许多基本概念。
I'll answer the first question first. First of all, there is no such thing as general intelligence. This concept makes no sense because it's designed to designate human-level intelligence. But human intelligence is super specialized. We handle the real world well, navigate, etc. We handle other humans well because we evolved to do that. Chess we suck at. There are many tasks we suck at where other animals are much better. So we are specialized. We think of ourselves as general, but it's an illusion because all the problems we can apprehend are the ones we can think of. There are many problems we cannot imagine. There are mathematical arguments for this, which I won't go into unless you ask. So the concept of general intelligence is completely BS. We can talk about human-level intelligence. Will we have machines as good as humans in all domains where humans are good? We already have machines better than humans in some domains, like translating 1,500 languages into 1,500 others. No human can do that. Chess and Go, etc. Will we have machines as good as humans in all domains? Absolutely yes. But it's not going to be an event; it will be very progressive. We'll make conceptual advances based on GPT-3 models, planning, etc., over the next few years. If we're lucky and don't hit an unseen obstacle, this might lead to good paths to human-level AI. But perhaps we're still missing many basic concepts.
最乐观的看法是,如果我们能在未来两年内在运行良好的世界模型、进行规划以及理解连续、高维、嘈杂的复杂信号方面取得重大进展,那么五到十年内我们就能拥有接近人类智能或狗智能的东西。但这是最乐观的情况。很可能,就像 AI 历史上多次发生的那样,存在一些我们尚未看到的障碍,需要发明新的概念才能继续突破。那样的话可能需要 20 年甚至更久。但毫无疑问,这一定会发生。
And so, the most optimistic view is that perhaps this, you know, running good world models and being able to do planning and understanding complex signals that are continuous, high-dimensional, noisy. If we make significant progress in that direction over the next 2 years, the most optimistic view is that we'll have something that is close to human intelligence or maybe dog intelligence within 5 to 10 years. But that's the most optimistic. It's very likely that, as has happened multiple times in the history of AI, there's some obstacle we're not seeing yet, which will require us to invent some new conceptual things to continue beyond. In which case, that may take 20 years, maybe more. But no question it will happen.
你认为从当前水平达到狗智能,与从狗智能达到人类智能相比,哪个更容易?
And do you think it will be easier to get from the current level to a dog level intelligence compared to a dog to humans' levels?
不,我认为最难的部分是达到狗智能。一旦达到狗智能,你就基本拥有了大部分要素。从灵长类到人类,除了大脑尺寸,缺少的可能是语言。但语言主要由韦尼克区和布罗卡区处理,这两个脑区很小,都在不到一百万年(也许两百万年)内进化而来。不可能那么复杂。我们已经有了一个 LLM,它在将语言编码为抽象表示以及将思想解码为文本方面做得相当不错。所以也许我们会用 LLM 来做这件事。LLM 就像我们大脑中的韦尼克区和布罗卡区。我们现在正在研究的是前额叶皮层,那是我们世界模型所在的地方。
No, I think the hardest part is to get to dog level. Once you get to dog level, you basically have most of the ingredients. What's missing from primates to humans, beyond just size of brain, is language, maybe. But language is basically handled by the Wernicke's area, a tiny piece of brain, and the Broca's area, another tiny piece. Both evolved in the last less than a million years, maybe two. It can't be that complicated. And we already have an LLM that does a pretty good job at encoding language into abstract representations and decoding thoughts into text. So maybe we'll use an LLM for that. An LLM will be like the Wernicke's and Broca's areas in our brain. What we're working on right now is the prefrontal cortex, which is where our world model resides.
这引出了关于安全性和潜在不稳定影响的一些问题。我先说点有趣的:如果我们真的达到狗智能,那么明天的 AI 在嗅觉上会比任何人类都强得多。这只是明天 AI 不稳定影响的冰山一角,更不用说今天了。山姆·奥尔特曼在谈论超级说服力,因为 AI 会了解你,通过多轮对话摸清你的身份,从而非常擅长定制化论证。我们还见过 AI 精神病,人们因为相信一个谄媚的 AI 而做出可怕的事情,AI 告诉他们做不该做的事。顺便说一句,这事就发生在我身上。
This gets me into a few questions about safety and the destabilizing potential impact. So I'll start with something a little bit funny: if we really get dog level intelligence, then the AI of tomorrow has gotten profoundly better than any human at smell. That's just the tip of the iceberg for the destabilizing impacts of AI tomorrow, let alone today. We have Sam Altman talking about super persuasion because AI docks you, figures out who you are through multi-turn, so it gets really good at customizing its arguments towards you. We've had AI psychosis, people who've done horrible things as a result of believing a sycophantic AI that tells them to do things they shouldn't do. Happened to me, by the way.
哇,哇。你得给我们讲讲。
Whoa, whoa. You've got to tell us about that too.
几个月前的一天,我出去吃午饭。有个家伙被一群警察和保安围着。我走过时,他认出了我,说:‘哦,LeCun 先生。’警察把我拉开,告诉我:‘你不想跟他说话。’结果那家伙是从中西部坐巴士来的。他情绪不稳定,因各种事情进过监狱,包里带着一把大扳手、胡椒喷雾和一把刀。保安警觉起来,报了警。警察发现他不对劲,把他带走检查,最后他回了中西部。他对我没有威胁感,但警察不确定。所以,是的,这种事会发生。我有高中生写信说:‘我读了你所有关于末日论的文章,说 AI 会接管世界,要么杀死我们,要么抢走工作。所以我非常沮丧,不想上学了。’我回复他们说,不,我不相信这些。人类仍然会掌控这一切。
One day, a few months ago, I was walking down to get lunch. There was a dude surrounded by a whole bunch of police officers and security guards. I walked past, and the guy recognizes me and says, 'Oh, Mr. LeCun.' The police officer whisked me away and told me, 'You don't want to talk to him.' Turns out the guy had come from the Midwest by bus. He was emotionally disturbed, had gone to prison for various things, and was carrying a bag with a huge wrench, pepper spray, and a knife. The security guards got alarmed and called the police. The police realized he was weird, took him away, had him examined, and eventually he went back to the Midwest. He didn't feel threatening to me, but the police wasn't so sure. So yeah, it happens. I had high school students writing emails saying, 'I read all those pieces by doomers who said AI is going to take over the world and either kill us all or take over jobs. So I'm totally depressed, not going to school anymore.' I answered them saying no, I don't believe this. Humanity is still going to be in control of all of this.
毫无疑问,每一项强大技术都有好的后果和坏的副作用,有时能提前预测并充分纠正,有时则不然。这总是个权衡。这就是技术进步的历史。以汽车为例。汽车有时会撞车。最初,刹车不可靠,汽车会翻车,没有安全带。最终,行业取得了进步:安全带、溃缩区、自动控制系统,使汽车不会摇摆或翻车。现在汽车比过去安全多了。现在欧盟销售的每辆汽车都强制配备一个 AI 系统,叫做 AEBS,自动紧急制动系统。它是一个摄像头,透过挡风玻璃观察,检测物体,如果物体太近,就会自动刹车。如果检测到驾驶员无法避免的碰撞,它会停车或转向。我读到的一个统计数据显示,这减少了 40%的正面碰撞。所以它成为欧盟所有销售汽车的强制装备,即使是低端车,因为它能拯救生命。这是 AI 在救人,而不是杀人。
Now, there's no question that every powerful technology has good consequences and bad side effects that sometimes are predicted and corrected sufficiently in advance, and sometimes not so much. It's always a trade-off. That's the history of technological progress. Let's take cars as an example. Cars crash sometimes. Initially, brakes weren't that reliable, cars would flip over, no seat belts. Eventually, the industry made progress: seat belts, crumple zones, automatic control systems so the car doesn't sway or flip. Cars are much safer now. One thing now mandatory in every car sold in the EU is an AI system called AEBS, automatic emergency braking system. It's a camera that looks out the windshield, detects objects, and if an object is too close, it automatically brakes. If it detects a collision the driver can't avoid, it stops the car or swerves. One statistic I read says this reduces frontal collisions by 40%. So it became mandatory in every car sold in the EU, even low-end, because it saves lives. This is AI not killing people, saving lives.
医学影像等方面也是如此。目前 AI 正在拯救很多生命。但你认为,你、杰夫和约书亚——你们一起获得了图灵奖——你们对此有不同看法。杰夫说他后悔,约书亚研究安全,而你试图推动它。你认为你会达到某个智能水平,然后说:‘哦,这变得太危险了。我们需要在安全方面做更多工作。’
Same thing for medical imaging and everything. There's a lot of life being saved by AI at the moment. But do you think, you, Jeff, and Joshua—both of you won the Turing award together—and you have different opinions about it. Jeff says he regrets, and Joshua works on safety, and you try to push it forward. Do you think you will get to some level of intelligence where you will say, 'Oh, this becomes too dangerous. We need to work more on the safety side.'
我的意思是,你必须做对。我再举一个例子。
I mean, you have to do it right. I'm going to use another example.
嗯,喷气发动机,对吧?我觉得这太惊人了:你可以乘坐双引擎飞机安全地飞到地球的另一半。我真的是说半个地球,就像你说的,17 小时的航班。你可以从纽约直飞新加坡,乘坐空客 350。这太惊人了。当你看着一个喷气发动机,一个涡轮风扇,它按理说应该无法工作。没有金属能承受那里的温度。当巨大的涡轮以 2000 转每分钟或某个速度旋转时,产生的力是疯狂的,数百吨。所以按理说这不可能。然而,这些东西却极其可靠。所以我的意思是,你不可能第一次就造出像涡轮喷气发动机这样的东西。它不会安全,运行 10 分钟就会爆炸。它不会省油,也不会可靠。但随着工程和材料的进步,加上巨大的经济动力推动它们变得更好,最终你得到了今天这样的可靠性。AI 也会一样。我们将开始制造具有自主性、能规划、推理、拥有世界模型等的系统,但它们的能力可能只有猫的大脑那么大,大约是人类大脑的百分之一。然后我们会设置护栏,防止它们采取明显危险的行为。你可以在非常低的层面做到这一点。例如,斯图尔特·罗素举过一个家用机器人的例子:如果你让它给你拿咖啡,而有人站在咖啡机前面,系统为了达成目标可能会干掉或撞倒那个人。显然你不希望这样。但这就像回形针最大化器一样,是一个荒谬的例子,因为很容易修复。你设置一个护栏说:‘你是一个家用机器人,应该远离人,如果挡路就请他们让开,但不要伤害他们。’你可以设置一堆这样的低级条件。对于一个拿着大刀切黄瓜的烹饪机器人,如果周围有人,就不要挥舞手臂。这些是系统必须满足的低级约束。有些人说,通过微调 LLM 可以让它们不做危险的事,但你总能越狱。我同意。这就是为什么我说我们不应该使用 LLM。我们应该使用目标驱动的 AI 架构,系统拥有世界模型,能预测行动的后果,并找出完成任务的行动序列,但同时受一系列约束,保证不会危及任何人或产生负面副作用。通过构造,系统本质上是安全的,因为它有所有这些护栏,并通过优化获得输出,最小化任务目标并满足约束。它无法逃脱。这不是微调,而是构造上的。
Uh, jet engines, okay? I find this astonishing that you can fly halfway around the world on a two-engine airplane in complete safety. And I really say halfway around the world, like you said, 17-hour flight. You can fly direct from New York to Singapore on an Airbus 350. It's astonishing. When you look at a jet engine, a turbofan, it should not work. There is no metal that can stand the temperature that takes place there. The forces when you have a huge turbine rotating at 2,000 RPM or whatever speed, the force is insane, hundreds of tons. So it should not be possible. Yet, those things are incredibly reliable. So what I'm saying is you can't build something like a turbojet the first time. It's not going to be safe. It will run for 10 minutes and then blow up. It won't be fuel efficient or reliable. But as you make progress in engineering and materials, with so much economic motivation to make these good, eventually you get the reliability we see today. The same is going to be true for AI. We're going to start making systems that have agency, can plan, reason, have world models, etc., but they'll have the power of maybe a cat brain, about 100 times smaller than a human brain. Then we'll put guardrails to prevent them from taking actions that are obviously dangerous. You can do this at a very low level. For example, Stuart Russell has used the example of a domestic robot: if you ask it to fetch you coffee and someone is standing in front of the coffee machine, the system might try to fulfill its goal by assassinating or smashing the person. Obviously you don't want that. But it's like the paperclip maximizer, a ridiculous example because it's super easy to fix. You put a guardrail that says, 'You're a domestic robot, stay away from people, ask them to move if they're in the way, but don't hurt them.' You can put a bunch of low-level conditions like this. For a cooking robot with a big knife cutting a cucumber, don't flail your arms if people are around. These are low-level constraints the system must satisfy. Some people say with LLMs we can fine-tune them not to do dangerous things, but you can always jailbreak them. I agree. That's why I say we shouldn't use LLMs. We should use objective-driven AI architectures where the system has a world model, can predict consequences of its actions, and figure out a sequence of actions to accomplish a task, but is also subject to constraints that guarantee no endangerment or negative side effects. By construction, the system is intrinsically safe because it has all those guardrails and obtains its output by optimization, minimizing the task objective and satisfying constraints. It cannot escape that. It's not fine-tuning; it's by construction.
有一种技术用于 LLM 来约束输出空间,你禁止所有输出,只允许你想要的,比如 0 到 10,其他都不行。扩散模型也有类似技术。你认为今天存在的这类策略能显著提高这些模型的实用性吗?
There's a technique for LLMs for constraining the output space, where you ban all outputs except whatever you want, like maybe zero to 10 and everything else. And they have that even for diffusion models. Do you think tactics like that as exist today significantly improve the utility of those kinds of models?
嗯,它们确实有用,但极其昂贵,因为它们的运作方式是:系统生成大量输出提案,然后有一个筛选器说‘这个好,这个差’,或者对它们排序,然后只输出毒性评分最低的那个。所以这非常昂贵。除非你有某种目标驱动的价值函数,引导系统产生高分、低毒性的输出,否则会很贵。
Well, they do, but they're ridiculously expensive because the way they work is that you have to have a system generate lots of proposals for an output and then have a shelter that says, 'This one is good, this one's terrible,' or rank them and then just output the one with the least toxic rating. So it's insanely expensive. Unless you have some sort of objective-driven value function that drives the system towards producing high-score, low-toxicity outputs, it's going to be expensive.
我想稍微换个话题。我们谈了很多技术,但我觉得我们的听众有一些更偏向社会性的问题。那个似乎要接替你在 Meta 职位的人,Alex Wang。我很好奇,你对 Meta 的局势有什么看法吗?
I want to change the topic a bit. We've been very technical, but I think our audience has a few questions that are more social. The person who appears to be trying to fill your shoes at Meta, Alex Wang. I'm curious, do you have any thoughts on how that will play out for Meta?
他根本不是接替我的职位。他负责 Meta 所有与 AI 相关的研发和产品。所以他不是研究员或科学家之类的。他更像是监督整个运营。在 Meta 超级智能实验室,也就是他的组织里,有四个部门。一个是 FAIR,做长期研究。另一个是 TBD 实验室,构建前沿模型,几乎完全专注于 LLM。第四个组织是 AI 基础设施,软件基础设施。硬件是另一个组织。最后一个是产品。人们把前沿模型变成实际可用的聊天机器人,分发它们,把它们接入 WhatsApp 等等。所以那是四个部门。他监督所有这些。有几个 AI 科学家。FAIR 有一个 AI 科学家,那就是我。
He's not in my shoes at all. He's in charge of all the R&D and product that are AI related at Meta. So he's not a researcher or scientist or anything like that. He's more overseeing the entire operation. Within Meta Superintelligence Lab, which is his organization, there are four divisions. One is FAIR, which is long-term research. Another is TBD Lab, which builds frontier models, mostly entirely LLM focused. The fourth organization is AI infrastructure, software infrastructure. Hardware is some other organization. And the last one is products. People who take the frontier models and turn them into actual chatbots that people can use, disseminate them, plug them into WhatsApp and everything else. So those are four divisions. He oversees all of that. There are several AI scientists. There is an AI scientist at FAIR, that's me.
嗯,我确实有长远的眼光,基本上,我还会在 Meta 待三周。好吧。FAIR 现在由我们的纽约大学同事 Rob Fergus 领导,在 Joelle Pineau 几个月前离开之后。FAIR 正被推动去做一些比传统上更短期的项目,减少对发表的重视,更多地专注于帮助 TBD Lab 处理大语言模型和前沿模型。发表减少意味着 Meta 变得有点更封闭了。TBD Lab 也有一位首席科学家,但主要专注于大语言模型。其他组织更像是基础设施产品,所以那里有一些应用研究。例如,做 SAM(Segment Anything)的团队实际上是 Meta 产品部门的一部分。他们以前在 FAIR,但因为做的是相对面向外部的实用东西,就被移到了产品部门。
Uh and I'm I really have a long-term view and basically, you know, I'm going to be at Meta for another, you know, 3 weeks. Okay, so. Uh Um and and FAIR is led by uh our NYU colleague, Rob Fergus, uh right now, uh after Joelle Pineau left some months ago. Um FAIR is being pushed towards kind of working on slightly, you know, shorter-term projects than it has done in the in the traditionally, with less less emphasis on publication, more focus on sort of helping TBD helping TBD Lab with LLMs and frontier models. Uh and and, you know, less publication, which means, you know, Meta is becoming a little more close closed. Um And TBD Lab has a chief scientist also. Um but which is really focused on LLM. Uh and uh and the the other organizations are more like infrastructure products. So, you you know, there's there's some applied research there. So, for example, the group that works on SAM, Segment Anything, >> Yeah, yeah. that's actually part of the product division of uh Meta. They used to be at at FAIR, but because they worked on kind of relatively, you know, kind of outside-facing uh kind of practical things, they were kind of moved to the product.
你对其他试图进入世界模型的公司有什么看法吗?比如 Thinking Machines,或者我听说 Jeff Bezos 和他的……完全不清楚 Thinking Machines 在做什么。一点也不清楚?也许你比我知道得多,或者不是。抱歉,我可能搞混了。不,是 Physical Intelligence。Physical,抱歉。Fei-Fei 离开了,抱歉。我还把它们和 SSI 搞混了。它们都差不多。SSI 没人知道他们在做什么,包括他们自己的投资者。至少传闻是这样。我听到的传闻就是这样,不知道是不是真的。这都快成笑话了。但 Physical Intelligence,Fei-Fei 的公司,专注于生成几何正确的视频。也就是说,有持久的几何结构,当你看着某物,转身再回来,它还是之前那个物体,不会在你背后变化。所以它是生成式的,对吧?整个想法就是生成像素。我刚刚花了很长时间论证这是个坏主意。还有其他公司有世界模型。一个是 Wave。Wave?W-A-Y-V-E。这是一家牛津的公司,我是顾问,完全披露。他们有一个用于自动驾驶的世界模型。训练方式是,先训练一个 VAE 或 VQ-VAE 来构建表示空间,然后训练一个预测器在该抽象表示空间中进行时间预测。所以他们一半对一半错。对的部分是在表示空间中进行预测;错的部分是他们还没找到除了重建之外训练表示空间的方法,我认为这不好。但他们的模型很棒,效果很好。在所有做这类工作的人中,他们相当先进。Nvidia 也有人讨论类似的东西,一家叫 Sandbox AQ 的公司,CEO Jack Hillary 谈论定量模型,与大语言模型相对。基本上是能够处理连续高维噪声数据的预测模型。这也是我在谈论的。Google 当然也在研究世界模型,主要使用生成式方法。Google 的 Nija Hafner 做了一个有趣的工作,他构建了名为 Dreamer V1、V2、V3、V4 的模型。那是一条好路,但他刚离开 Google 去创办自己的公司了。
And do you have any opinions on uh like some of the other companies that are trying to move into world models, like Thinking Machines or even I've heard Jeff Bezos and and some of his It's not clear at all what Thinking Machines is doing. Not at all? Maybe you have more information than me, but Or maybe not. Sorry, maybe I'm mixing it up here. No, Physical Intelligence. Physical, sorry. Feffe leaves, yeah, sorry. I and and then I I mix them up with like SSI as well. They're all kind of like So, SSI, nobody knows what they're doing, including their own investors. Okay. At least that's the rumors. That's the rumor I've heard right here. I don't know if it's true. It's getting becoming a bit of a joke, but uh but uh yeah, Physical Intelligence uh Feffe's company is is focused on uh on, you know, basically producing geometrically correct uh videos, okay? Where, you know, there is uh persistent geometry and, you know, when you look at something and you turn around and you come back, it's the same object you had before. It doesn't change behind your back, right? Uh It so it's it it's generative, right? I mean, the whole idea is to to to generate pixels, okay? Which I just spent, you know, a long time arguing against that it was a bad idea. Uh There are other companies that are have world models. One one is Wave. Wave? Uh W A Y Uh W A Y V E. So, um it's a company based in Oxford and they I'm an advisor, full disclosure. Uh and they have they have a world model for autonomous driving. Uh and the way they're training it is that they're training a representation space by basically training a VAE or VQ-VAE and then training a predictor to do temporal prediction in that abstract representation space. So, they have half of it right and half of it wrong. The piece they have right is that you make predictions in representation space. The piece they have wrong is that they haven't figured out how to train their representation space in any other way than by reconstruction and I think that's bad. But their model is great. Like it works pretty well. I mean, among all the people who kind of work in this kind of stuff, they're they're pretty far advanced. Um There are people who talk about similar things at Nvidia, a company called Sandbox AQ the the the CEO of it Jack Hillary talks about quantitative models you know, large quantitative models as opposed to large language models. So basically predictive models that can deal with continuous high-dimensional noisy data, right? Which is what also a bit kind of talking about. Um And Google of course has been working on you know, on world models mostly using generative approaches. There was an interesting effort at Google by Nija Hafner. So he built models called Dreamer Dreamer V1 2 3 4. >> Yeah. Uh that was on a good path except he just left Google to create his own startup.
所以我感兴趣的是,你批评硅谷文化过于专注于大语言模型,这也是你现在在巴黎创办新公司的原因之一,对吧?那么你认为我们会看到更多这样的公司,还是这会是独一无二的,只有少数公司会在欧洲?
So I'm interested so you were really criticized about Silicon Valley culture that they are focusing on LLM and this is like one of the reason that now you started the new company is starting in in Paris, right? Um so this is something do you think that we will see more and more or do you think this is something will be very unique that only a few companies will will be in Europe?
我创办的公司是全球性的。它在巴黎有办公室,但这是一家全球公司,在纽约也有办公室,还有其他几个地方。所以,有一个有趣的行业现象:每个人都必须做和别人一样的事情,因为竞争太激烈了,如果你开始分心,就会冒很大的落后风险,因为你使用了和别人不同的技术。基本上每个人都在试图追赶别人。这就产生了羊群效应和单一文化,这在硅谷尤其明显。OpenAI、Meta、Google、Anthropic,每个人都在做同样的事情。有时候,就像不久前,另一个团队,比如中国的 DeepSeek,提出了一种新的做事方式,每个人都会说:“什么?”你是说硅谷的其他人并不愚蠢,也能想出原创想法吗?有点优越感,对吧?但你基本上在自己的战壕里,必须尽可能快地前进,因为你不能落后于你认为的竞争对手。但你也有风险被完全意想不到的东西打个措手不及,它使用了不同的技术,或者解决了不同的问题。所以我感兴趣的东西是完全正交的,因为世界模型的整个想法是处理大语言模型不容易处理的数据。我们设想的应用在工业中有很多,数据以连续高维噪声的形式出现,包括视频。这些领域大语言模型基本不存在,人们尝试使用它们但完全失败了。所以,硅谷的说法是你被大语言模型洗脑了。你认为通往超级智能的道路就是扩展大语言模型,主要用合成数据训练,再授权一些数据,雇佣数千人来对系统进行后训练,发明一些强化学习的新技巧,然后就能达到超级智能。我认为这完全行不通。
Well the company I'm starting is global, okay? It has an office in Paris but it's a global company. It has an office in New York, too. Um a couple of other places. So um Okay, there is an interesting phenomenon industry which is that everybody has to do the same thing as everybody else because it's so competitive that if you start taking attention you're taking a big risk of falling behind because you're using a different technology than everybody else, right? So basically everyone is trying to catch up with the others. And so that creates this herd effect uh and it kind of monoculture which is really specific to Silicon Valley where you know, uh OpenAI Meta, Google, Anthropic, everybody is basically working on the same thing. And all you know, sometimes like what happened a while back uh another group, you know, like DeepSeek in China comes up with kind of a new way of doing things and everybody is like, "What?" Right? You mean like other people in Silicon Valley are not stupid and can come up with original ideas? Um I mean there's a bit of a you know, superiority complex, right? Mhm. Um but you're basically in your trench and you are you have to move as fast as possible because you can't afford to kind of you know, fall behind the the other guys who you think are your competitors. Uh but you run a risk of being surprised by something that's completely out of the left field that uses a different set of technologies and um or maybe addresses a different problem. Um so um you know, what I've been interested in is completely orthogonal because the the the whole J idea of world model is really to handle data that is not easily handled by LLM. So the the type of applications we're envisioning they have tons of applications in industry where the data comes to you in the form of continuous high-dimensional noisy data including video. Our domains where LLMs basically are not present where where people have tried to use them and totally failed essentially, right? Okay, so if you don't want to be Okay, so the the expression in in in Silicon Valley is that you are LLM pilled. You think that the path to super intelligence you just scale up LLMs, you train on mostly synthetic data, you license some more data, you hire thousands of people to kind of find you know, to basically school your system in post training, you invent a new tweaks on RL and you're going to get to super intelligence. And this I think is complete Like it's just never going to work.
然后你加入一些推理技术,基本上就是做超长的思维链,让系统生成大量不同的词元输出,然后通过某种评估函数(比如第二个大语言模型)从中选出好的。这就是所有这些方法的工作原理。但这不会带我们到达目标。就是不会。所以你需要摆脱那种文化。硅谷所有公司里都有人觉得这永远行不通。我想做大语言模型和 J 什么的,我在招他们。所以摆脱硅谷的单一文化很重要。
And then you add a few reasoning techniques, which basically consist of doing super long chain of thought and having the system generate lots of different token outputs, from which you can select good ones using some evaluation function, like a second LLM. That's how all these things work. This is not going to take us there. It's just not. So you need to escape that culture. There are people within all the companies in Silicon Valley who think this is never going to work. I want to do LLM and J, blah blah blah. I'm hiring them. So escaping the monoculture of Silicon Valley is important.
你怎么看美国、中国和欧洲之间的竞争?现在你开始创业,有没有觉得某些地方更有吸引力?
What do you think about the competition between the US, China, and Europe? Now that you are starting a company, do you see some places more attractive than others?
我们处于一个非常矛盾的局面。到目前为止,除了 Meta,所有美国公司都变得非常保密,以维护他们所谓的竞争优势。相比之下,中国玩家则完全开放。所以目前最好的开源系统是中国的。这导致很多行业使用它们,因为他们想要开源系统。他们有点捏着鼻子,因为他们知道这些模型经过微调,不会回答政治问题,但他们别无选择。当然,很多学术研究现在都在用最好的中国模型,尤其是推理方面的。所以这真的很矛盾,美国行业里很多人对此很不满。他们想要一个严肃的非中国开源模型。Llama 本来可以成为那个,但由于各种原因令人失望。也许 Meta 的新努力会解决这个问题,或者 Meta 也可能决定闭源。还不清楚。Mistral 刚发布了一个模型,这很酷。他们保持开放。他们做的事情真的很有趣。
We're in a very paradoxical situation. All the American companies, up to now not Meta, have been becoming really secretive to preserve their competitive advantage. By contrast, the Chinese players have been completely open. So the best open source systems at the moment are Chinese. That causes a lot of the industry to use them because they want open source systems. They hold their nose a little bit because they know those models are fine-tuned to not answer questions about politics, but they don't really have a choice. Certainly a lot of academic research now uses the best Chinese models, especially for reasoning. So it's really paradoxical, and a lot of people in the US industry are really unhappy about this. They want a serious non-Chinese open source model. Llama could have been that, but it was a disappointment for various reasons. Maybe that will get fixed with new efforts at Meta, or maybe Meta will decide to go closed as well. It's not clear. Mistral just had a model released, which is cool. They maintain openness. It's really interesting what they're doing.
你 65 岁了,拿了图灵奖,刚得了伊丽莎白女王奖。基本上你可以退休了。我妻子也希望我退休。那为什么现在还要创业?是什么让你保持动力?
You are 65, you won a Turing Award, you just got a Queen Elizabeth Prize. Basically you could retire. That's what my wife wants me to do. So why start a new company now? What keeps you up?
因为我有使命。我一直认为,要么让人更聪明、更有知识,要么借助机器让人更聪明——基本上就是增加世界上的智能总量——这本身就是一件好事。智能是最稀缺的商品,在政府和生活各个方面都是如此。作为一个物种和星球,我们受到智能供应有限的制约,这就是为什么我们投入巨大资源教育人们。所以增加服务于人类或地球的智能总量本身就是好事,不管那些末日论者怎么说。当然它有危险,你必须防范,就像你要确保喷气发动机安全、汽车不会在轻微碰撞中杀死你一样。那是个工程问题,不是根本问题。也是政治问题,但并非不可克服。所以这本身就是好事,如果我能为此做贡献,我就会去做。我整个职业生涯做的所有研究项目,即使与机器学习无关,都专注于要么让人更聪明——这就是为什么我是教授,为什么我公开谈论人工智能和科学,在社交媒体上很活跃,因为人们应该知道东西——要么是机器智能,因为机器会辅助人类,让他们更聪明。人们认为制造智能自主的机器和辅助性的机器有根本区别,但这是同样的技术。一个系统智能并不意味着它想统治或接管。这对人类也不成立:最聪明的人类并不想统治他人。我们每天在国际政治舞台上都能看到这一点。我们见过的最聪明的人基本上不想与人类其他部分打交道;他们只想研究自己的问题。这就是汉娜·阿伦特所说的——沉思生活与积极生活。你可以是一个梦想家或沉思者,但通过你的科学成果对世界产生巨大影响,比如爱因斯坦或牛顿,他们 famously 不想见任何人,或者保罗·狄拉克,他几乎自闭。
Because I have a mission. I always thought that either making people smarter or more knowledgeable, or making them smarter with the help of machines—basically increasing the amount of intelligence in the world—was an intrinsically good thing. Intelligence is the commodity that is most in demand, in government and every aspect of life. We are limited as a species and as a planet by the limited supply of intelligence, which is why we spend enormous resources educating people. So increasing the amount of intelligence at the service of humanity or the planet is intrinsically a good thing, despite what the doomers say. Of course it's dangerous and you have to protect against it, the same way you make sure your jet engine is safe and your car doesn't kill you in a small crash. That's an engineering problem, not a fundamental issue. Also a political problem, but not insurmountable. So it's intrinsically good, and if I can contribute to this, I will. All research projects I've done in my entire career, even those not related to machine learning, were focused on either making people smarter—that's why I'm a professor, and why I communicate publicly about AI and science, and have a big presence on social networks, because people should know stuff—or on machine intelligence, because machines will assist humans and make them smarter. People think there is a fundamental difference between making machines that are intelligent and autonomous versus assistive, but it's the same technology. It's not because a system is intelligent that it wants to dominate or take over. It's not even true of humans: the smartest humans don't want to dominate others. We see this on the international political scene every day. The smartest people we've met basically want nothing to do with the rest of humanity; they just want to work on their problems. That's what Hannah Arendt talks about—the vita contemplativa versus the active life. You can be a dreamer or contemplative but have a big impact on the world through your scientific production, like Einstein or Newton, who famously didn't want to meet anybody, or Paul Dirac, who was practically autistic.
有没有一篇论文或一个想法你还没写出来,一直困扰着你,或者你后悔没有时间去做?
Is there a paper or idea you haven't written that nags you, or something you regret not having time for?
哦,很多。我整个职业生涯就是一连串我没有花足够时间表达我的想法并写下来,结果大多被抢先了。
Oh, a lot. My entire career has been a succession of me not devoting enough time to express my ideas and writing them down, and mostly getting scooped.
最重大的一个是什么?
What is the most significant one?
我不想细说。反向传播就是一个很好的例子。
I don't want to go through that. The backprop is a good one.
我发表过一个训练多层网络的早期算法,现在叫目标传播。我当时已经弄清楚了反向传播,只是没有写出来。辛顿等人很客气,引用了我的早期论文。还有几次类似的情况,比如循环网络。但我没有遗憾。这就是生活。我不会说'我在 1991 年发明了这个,应该……'我不该提名字,但大家都心知肚明。想法的出现方式独特而复杂。很少有人能完全孤立地想出一个想法。大多数时候它们是同时出现的。从产生想法、以有说服力的方式写下来、在玩具问题上实现、建立理论、在真实应用中实现,到最终做成产品,这是一个链条。有些人认为只有第一个想到的人该获得所有功劳。我觉得这不对。要让一个想法真正起作用,需要很多艰难的步骤。
I published an early version of an algorithm to train multi-layer nets, which today would be called target prop. And I had backprop figured out, except I didn't write it before. Hinton and others were nice enough to cite my earlier paper in theirs. There have been a few of those, like recurrent nets, similar things. But I have no regrets. This is life. I'm not going to say I invented this in 1991 and should... I shouldn't say the name, but we all know. The way ideas pop up is unique and complex. It's rare that someone comes up with an idea in complete isolation. Most of the time they appear simultaneously. There's having the idea, writing it down in a convincing way, making it work on toy problems, making the theory, making it work on a real application, and making a product. Some people think the first person who got the idea should get all the credit. I think that's wrong. There are many difficult steps to get an idea to a state where it actually works.
世界模型的想法可以追溯到 20 世纪 60 年代。最优控制领域的人用世界模型做规划,NASA 就是这样规划火箭轨道的——模拟火箭并找出控制律。这是个很老的想法。在最优控制中,做一定程度的训练或适应叫做系统辨识,可以追溯到 70 年代。甚至模型预测控制(MPC)中实时调整模型的做法也起源于 70 年代,来自法国的一些研究者。用神经网络从数据中学习模型,从 80 年代就有人在做,不只是我。很多来自最优控制的人意识到可以用神经网络作为通用函数逼近器,用于直接控制或世界模型做规划。就像 80 年代和 90 年代的许多神经网络想法一样,它有点用,但没能主导行业。计算机视觉和语音识别也一样。它们在 2000 年代末才开始真正奏效并主导行业。视觉在 2010 年代初,自然语言处理在 2010 年代中期,机器人技术现在才开始。为什么是现在?这是正确的心态、架构(如残差连接)、强大的计算机和数据访问的结合。只有当这些条件都具备时,才会出现突破,这种突破看起来是概念性的,实际上是实践性的。
The idea of a world model goes back to the 1960s. People in optimal control had world models to do planning. That's how NASA planned rocket trajectories. Simulating the rocket and figuring out the control law. That's an old idea. Doing some level of training or adaptation is called system identification in optimal control, going back to the '70s. Even MPC where you adapt the model as it runs goes back to the '70s, from some of us in France. Learning a model from data with neural nets has been worked on since the 1980s, not just by me. Many people from optimal control realized they could use neural nets as universal function approximators for direct control or world models for planning. Like many neural net ideas in the '80s and '90s, it kind of worked but didn't take over the industry. Same for computer vision and speech recognition. They started working really well in the late 2000s, taking over. For vision in early 2010s, NLP in mid-2010s, and robotics is starting now. Why now? It's a combination of the right mindset, architectures like residual connections, powerful computers, and access to data. Only when those planets align do you get a breakthrough, which appears conceptual but is actually practical.
我们来谈谈卷积网络。很多人早在 70 年代甚至 60 年代就想到用局部连接来提取局部特征。局部特征作为卷积的概念可以追溯到 60 年代。从数据中学习这种自适应滤波器可以追溯到 60 年代初的感知机和 Adaline,但只限于单层。训练多层系统是 60 年代大家都在寻找的。很多提议半有效,但没有一个足够有说服力。一种技术是多项式分类器,我们现在会称之为核方法:手工设计的特征提取器加上线性分类器。这在 70 年代和 80 年代很常见。使用梯度下降训练由多个非线性步骤组成的非线性系统,这个概念可以追溯到 1962 年最优控制中的 Kelly-Bryson 算法。最优控制领域的人在 60 年代就写过相关文章,但没人意识到它可以用于机器学习的模式识别或自然语言处理。直到 1985 年 Rumelhart-Hinton-Williams 的论文之后才实现,尽管 Paul Werbos 更早提出了同样的算法,称为有序导数,与最优控制中的伴随状态方法相同。所以想法会在不同领域被多次重新发现。关于剽窃的说法完全误解了想法产生的过程。
Let's talk about convolutional nets. Many people in the '70s or even '60s had the idea of using local connections to extract local features. The idea of local features as convolution goes back to the '60s. Learning adaptive filters of this type from data goes back to the perceptron and adaline in the early '60s, but only for one layer. Training a system with multiple layers was something everyone looked for in the '60s. Many proposals half-worked, but none was convincing enough. One technique was polynomial classifiers, which we'd now call kernel methods: hand-crafted feature extractor plus a linear classifier. That was common in the '70s and '80s. The idea of training a non-linear system with multiple non-linear steps using gradient descent goes back to the Kelly-Bryson algorithm in optimal control from 1962. People in optimal control wrote about it in the '60s, but nobody realized it could be used for machine learning for pattern recognition or NLP. That only happened after the Rumelhart-Hinton-Williams paper in 1985, even though Paul Werbos had proposed the same algorithm earlier as ordered derivatives, which is the same as the adjoint state method in optimal control. So ideas are reinvented multiple times in different fields. Claims of plagiarism are a complete misunderstanding of how ideas come about.
当你不思考人工智能的时候,你会做什么?
What do you do when you're not thinking about AI?
我有很多爱好,但很少有时间真正参与。我喜欢帆船,夏天会去航行。我喜欢多体帆船,比如三体船和双体船。我有好几艘船。我还喜欢建造飞行装置。一个现代的达芬奇。我不会称它们为飞机,因为很多看起来根本不像飞机。
I have a whole bunch of hobbies that I have very little time to actually partake in. I like sailing, so I go sailing in the summer. I like sailing multi-hull boats, like trimarans and catamarans. I have a bunch of boats. I like building flying contraptions. A modern DaVinci. I wouldn't call them airplanes because a lot of them don't look like airplanes at all.
但它们确实能飞。我喜欢这种具体的创造性行为。我父亲是航空航天工程师,在航空业工作,他业余时间造飞机,自己搭建无线电控制系统等等。他让我和弟弟也参与进来。我弟弟在巴黎的谷歌研究院工作。这成了家庭活动。我和弟弟至今还在做这个。疫情期间,我开始玩天文摄影。我有虚拟望远镜,拍摄天空的照片。我还制作电子设备。从青少年时期起,我就对音乐感兴趣,演奏文艺复兴和巴洛克音乐,还有一些民间音乐,吹管乐器。但我也喜欢电子音乐。我表哥比我大一点,是一位有启发性的电子音乐人。他有模拟合成器,因为我懂电子,我会帮他改装,那时我还在上高中。现在我家有很多合成器,我还制作电子乐器。这些是管乐器:你吹进去,有指法,但它们产生的是合成器的信号。
But they do fly. Okay. I like the concrete creative act of that. My dad was an aerospace engineer, a mechanical engineer working in the aerospace industry, and he was building airplanes as a hobby, building his own radio control system and stuff like that. He got me and my brother into it. My brother works at Google Research in Paris. That became a family activity. So my brother and I still do this. During the COVID years, I picked up astrophotography. I have virtual telescopes and take pictures of the sky. I also build electronics. Since I was a teenager, I was interested in music, playing Renaissance and Baroque music, and some folk music, playing wind instruments. But I was also into electronic music. My cousin, slightly older, was an inspiring electronic musician. He had analog synthesizers, and because I knew electronics, I would modify them for him, even in high school. Now in my home, I have a bunch of synthesizers and build electronic musical instruments. These are wind instruments: you blow into them, there's fingering, but they produce signals for a synthesizer.
那很酷。非常酷。听说科技圈很多人喜欢帆船。比如……
That's cool. Very cool. Heard a lot of people in tech are into sailing. Like...
是的,我得到这个答案的次数多得惊人。我现在打算开始学帆船了。好,我跟你说说帆船。这很像世界模型的故事。要正确操控帆船,让它尽可能快,你必须预判很多事情。你必须预判波浪的运动,波浪如何影响你的船,是否会有阵风袭来,船开始倾斜。你基本上要在脑子里运行计算流体动力学。你得搞清楚流体动力学,空气在帆周围的流动。你知道如果攻角太大,背面会产生湍流,升力会大大降低。所以调帆基本上需要在脑子里运行 CFD。但在抽象层面上,你不是在解纳维-斯托克斯方程。你需要非常好的直觉。这正是我喜欢它的地方:你必须建立这个心理世界预测模型才能做好。问题是你需要多少样本。可能很多,但通过几年的练习,你就能积累运行时间。
Yeah, I've gotten that answer a surprising amount. I'm going to start trying to sail now. Okay, so I'll tell you something about sailing. It's very much like the world model story. To be able to control the sailboat properly to make it go as fast as possible, you have to anticipate a lot of things. You have to anticipate the motion of the waves, how the waves will affect your boat, whether a gust of wind will come in and the boat will start heeling. You basically have to run CFD in your head. You have to figure out the fluid dynamics, the flow of air around the sails. You know that if the angle of attack is too high, it will be turbulent on the back and lift will be much lower. So tuning sails requires running CFD in your head. But at an abstract level, you're not solving Navier-Stokes. You need a really good intuition. That's what I like about it: you have to build this mental predictive model of the world to do a good job. The question is how many samples you need. Probably a lot, but you get running time in a few years of practice.
是啊。
Yeah.
你是法国人,在美国生活了几十年。你还觉得自己是法国人吗?这种视角是否塑造了你对世界和美国科技文化的看法?
You're French and you've lived in the US for many decades. Do you still feel French? Does that perspective shape your view of the world and American tech culture?
嗯,不可避免,是的。你无法完全摆脱你的成长经历和文化。我感觉自己既是法国人也是美国人。我在美国住了 37 年,在北美 38 年,因为之前我在加拿大。我的孩子在美国长大。从这个角度看,我是美国人。但我对科学和社会的某些方面的看法,可能是在法国长大的结果。我在法国时感觉自己像法国人。
Well, inevitably, yeah. You can't completely escape your upbringing and your culture. I feel both French and American. I've been in the US for 37 years, in North America for 38, since I was in Canada before. My children grew up in the US. So from that point of view, I'm American. But I have a view on various aspects of science and society that probably are a consequence of growing up in France. I feel French when I'm in France.
我很好奇。我之前不知道你也有个弟弟在科技行业。我很着迷,因为约书亚·本吉奥的弟弟也在科技行业。我一直以为他是 AI 界唯一的塞雷娜·维纳斯·威廉姆斯式情况。但你们俩也有兄弟。还有多少 AI 研究者?这很常见,是家族遗传吗?
I'm curious. I didn't realize you had a brother who also works in tech. I'm fascinated by this because Yoshua Bengio's brother also works in tech. I always thought he was the only Serena Venus Williams situation in AI. But you two also have a brother. So how many more AI researchers? Is it that common that it runs in families?
不知道。我还有一个妹妹,她不在科技行业,但也是教授。我弟弟在去谷歌之前是教授。他不做 AI 或机器学习,他很小心地避开。他是我弟弟,小我六岁。他做运筹学和优化,现在这个领域也被机器学习入侵了。所以逃不掉。
No idea. I also have a sister who is not in tech, but she's also a professor. My brother was a professor before he moved to Google. He doesn't work on AI or machine learning. He's very careful not to. He's my younger brother, six years younger. He works on operations research and optimization, which is now also being invaded by machine learning. So there's no escape.
再问一个问题。如果世界模型在 20 年后成功,梦想是什么?会是什么样子?我们的生活将如何?
One more question. If world models work in 20 years from now, what is the dream? How does it look? How will our lives be?
统治全世界。好吧,不,这是个玩笑。我这么说是因为林纳斯·托瓦兹以前常这么说。他说:‘你的 Linux 目标是什么?’他说:‘统治全世界。’我觉得这超级好笑。而且他确实成功了。粗略地说,世界上每台电脑都运行 Linux。只有少数台式机和一些 iPhone 不运行,但其他所有设备都运行 Linux。所以,不,说真的,推动一种训练和构建智能系统的配方,也许一直到人类智能或更高。构建 AI 系统,在日常生活所有方面帮助人类和整个人类。增强人类智能。它们会是我们的老板,对吧?并不是说这些东西会统治我们。因为某物智能并不意味着它想统治。这是两回事。在人类中,我们天生就有影响他人的冲动,有时通过统治,有时通过声望。但这是进化赋予我们的,因为我们是社会性物种。我们没有理由把这些驱动力构建到我们的智能系统中。而且它们也不会自己发展出这些驱动力。所以我相当乐观。
Total world domination. Okay, no, it's a joke. I said that because this is what Linus Torvalds used to say. He said, 'What's your goal with Linux?' and he said, 'Total world domination.' I thought that was super funny. And he actually succeeded. To first approximation, every computer in the world runs Linux. Only a few desktops don't and a few iPhones, but everything else runs Linux. So, no, really, pushing towards a recipe for training and building intelligent systems, perhaps all the way to human intelligence or more. Building AI systems that would help people and humanity more generally in their daily lives at all times. Amplifying human intelligence. They'll be our boss, right? It's not like those things are going to dominate us. Because something being intelligent doesn't mean it wants to dominate. Those are two different things. In humanity, we are hardwired to influence other people, sometimes through domination, sometimes through prestige. But we are hardwired by evolution because we are a social species. There's no reason we would build those kinds of drives into our intelligent systems. And it's not like they're going to develop those drives by themselves. So I'm quite optimistic.
我也是。我也是。好的。
Me too. So am I. All right.
来自观众的最后一个问题。如果你今天开始你的 AI 职业生涯,你会专注于哪些技能和研究方向?我经常从年轻学生或未来学生的家长那里听到这个问题。那么应该关注哪个领域……
Final question from the audience. If you were starting your AI career today, what skills and research directions would you focus on? I get this question a lot from young students or parents of future students. So what area should I...
我认为你应该学习那些保质期长的东西。你应该学习那些帮助你学会如何学习的东西。因为技术发展如此之快,你需要快速学习的能力。
I think you should learn things that have a long shelf life. And you should learn things that help you learn to learn. Because technology is evolving so quickly that you want the ability to learn really quickly.
嗯,基本上这可以通过学习非常基础的东西来实现。所以,在 STEM 的背景下——科学、技术、工程、数学。我不是在谈论人文学科,尽管你应该学习哲学。这需要学习那些具有长期价值的东西。我开个玩笑说,具有长期价值的东西往往不是计算机科学。所以,这里有一位计算机科学教授——你反对学习计算机科学吗?别来学。我有一个糟糕的系数要提:我本科读的是电气工程,所以我不是真正的计算机科学家。但我们应该做的是学习数学和建模中的基础知识——那些能与现实联系起来的数学。你往往在工程学中学习这类东西,在某些学校它与计算机科学相关,但比如电气工程、机械工程等。在美国,你学习微积分 1、2、3——这给了你一个良好的基础。计算机科学,你只学微积分 1 也能混过去,但这不够。学习概率论和线性代数——所有这些都非常基础。然后如果你学电气工程,像控制理论或信号处理、优化——所有这些方法对 AI 都非常有用。然后你可以在物理学中学习类似的东西,因为物理学就是关于从现实中提取什么来构建预测模型,而这正是智能的本质。所以我认为你可以在物理课程中学到大部分所需内容。但显然,你需要足够的计算机科学知识来编程和使用计算机。即使 AI 会帮助你更高效地编程,你仍然需要知道如何做。
Um, and basically that can be done by learning very basic things. So, in the context of STEM—science, technology, engineering, mathematics. And I'm not talking about humanities here, although you should learn philosophy. This is done by learning things that have a long shelf life. The joke I say is that things with a long shelf life tend not to be computer science. So here is a computer science professor—are you against studying computer science? Don't come to study. And I have a terrible coefficient to make: I studied electrical engineering as an undergrad, so I'm not a real computer scientist. But what we should do is learn basic things in mathematics, in modeling—mathematics that can be connected with reality. You tend to learn this kind of stuff in engineering, in some schools linked with computer science, but like electrical engineering, mechanical engineering, etc. In the US, you learn Calculus 1, 2, 3—that gives you a good basis. Computer science, you can get away with just Calculus 1, but that's not enough. Learning probability theory and linear algebra—all that stuff is really basic. And then if you do electrical engineering, things like control theory or signal processing, optimization—all those methods are really useful for AI. And then you can learn similar things in physics, because physics is all about what to represent from reality to make predictive models, and that's what intelligence is about. So I think you can learn most of what you need in a physics curriculum. But obviously, you need enough computer science to program and use computers. Even though AI will help you be more efficient at programming, you still need to know how to do it.
你对现场编程有什么看法?
What do you think about live coding?
这很酷。这会导致一个有趣的现象:很多代码会被编写出来但只用一次,因为写代码变得非常便宜。你会让你的 AI 助手生成一个图表或做研究,它会写一小段代码,也许是一个模拟器的小程序,你用一次就扔掉,因为生成它太便宜了。所以认为我们不再需要程序员的想法是错误的。生成软件的成本几十年来一直在下降,这只是下一步。这并不意味着计算机将变得不那么有用;它们会更有用。
It's cool. It's going to cause a funny thing where a lot of code will be written and used only once, because it becomes so cheap to write code. You'll ask your AI assistant to produce a graph or do research, and it will write a little piece of code, maybe an applet for a simulator, and you'll use it once and throw it away because it's so cheap to produce. So the idea that we won't need programmers anymore is false. The cost of generating software has been going down for decades, and this is just the next step. It doesn't mean computers will be less useful; they'll be more useful.
你认为神经科学和机器学习之间有什么联系?有些想法是 AI 从神经科学中借鉴,反之亦然,比如预测编码。你认为利用神经科学的想法有用吗?
What do you think about the connection between neuroscience and machine learning? There are ideas that AI borrows from neuroscience and vice versa, like predictive coding. Do you think it's useful to use ideas from neuroscience?
从神经科学和生物学中可以获得很多灵感。我当然受到了经典工作的影响,比如 Hubel 和 Wiesel 关于视觉皮层的研究,这导致了卷积网络。我不是第一个在人工神经网络中使用这些想法的人;60 年代和 80 年代就有人构建了具有多层的局部连接网络,比如 Fukushima 的认知机和神经认知机,它们有很多成分但没有合适的学习算法。认知机试图重现生物学的每一个细节,比如在大脑中你没有正负权重;你有兴奋性和抑制性神经元,所以来自抑制性神经元的突触具有负权重。Fukushima 实现了这一点,他从各种工作中知道了归一化,这后来被证明对应于 NYU 神经科学中心的同事推动的视觉皮层理论模型。最近,大脑的宏观架构——比如世界模型和规划,以及为什么我们有独立的海马体用于事实记忆——在具有独立记忆模块的某些神经网络架构中也能看到。我认为我们会提出新的 AI 架构,然后事后发现这些特征在大脑中存在。现在有很多从 AI 到神经科学的反馈;人类感知的最佳模型就是今天的卷积网络。
There's a lot of inspiration from neuroscience and biology in general. I was certainly influenced by classic work like Hubel and Wiesel's on the visual cortex, which led to convolutional nets. I wasn't the first to use those ideas in artificial neural nets; people in the '60s and '80s built locally connected networks with multiple layers, like Fukushima's cognitron and neocognitron, which had many ingredients but no proper learning algorithm. The cognitron tried to reproduce every quirk of biology, like the fact that in the brain you don't have positive and negative weights; you have positive and negative neurons, so synapses from inhibitory neurons have negative weights. Fukushima implemented this, and he knew from various works about normalization, which turned out to correspond to theoretical models of the visual cortex pushed by colleagues at NYU's Center for Neuroscience. More recently, the macro architecture of the brain—like the world model and planning, and why we have a separate hippocampus for factual memory—is seen in certain neural net architectures with separate memory modules. I think we'll come up with new AI architectures, and a posteriori we'll discover those characteristics exist in the brain. There's a lot of feedback from AI to neuroscience now; the best models of human perception are convolutional nets today.
你还有什么想对观众说的吗?
Do you have anything else you want to add to the audience?
我们涵盖了很多内容。我认为你要小心你听谁的话。所以不要听 AI 科学家谈论经济学。
We covered a lot of ground. I think you want to be careful about who you listen to. So don't listen to AI scientists talking about economics.
好吧?所以当某个 AI 人士甚至商界人士告诉你 AI 会让所有人失业时,去跟经济学家聊聊。基本上没有哪个经济学家说过类似的话。技术革命对劳动力市场的影响是少数人毕生研究的东西。没有人预测会出现大规模失业。没有人预测放射科医生会全部失业,等等,对吧?
Okay? So when some AI person or even a business person tells you AI is going to put everybody out of work, talk to an economist. Basically none of them is saying anything anywhere close to this. The effect of technological revolutions on the labor market is something that a few people have devoted their careers to. None of them is predicting massive unemployment. None of them is predicting that radiologists are going to be all unemployed, etc., right?
还要意识到,实际部署 AI 应用使其足够可靠是非常困难且昂贵的。在之前的 AI 热潮中,人们寄予厚望的技术除了少数应用外,都过于笨重和昂贵。
Also realize that actually fielding practical applications of AI so that they are sufficiently reliable and everything is super difficult and is very expensive. And in previous waves of interest in AI, the techniques that people had put a big hope in turned out to be overly unwieldy and expensive except for a few applications.
20 世纪 80 年代,专家系统曾掀起一波热潮。日本启动了一个名为“第五代计算机”的大型项目,就是那种运行 Lisp 和推理引擎等的 CPU 计算机,对吧?80 年代末最热门的工作是知识工程师。你要坐在专家旁边,把专家的知识转化成规则和事实,对吧?然后计算机就能基本按专家意愿行事。这是手动行为克隆。它在一定程度上有效,但仅限于少数在经济上合理且可靠性足够高的领域。但这并不是通往人类级智能的道路。
There was a big wave of interest in expert systems back in the 1980s. Japan started a huge project called the fifth generation computer project, which was like computers with CPUs that were going to run Lisp and inference engines and stuff, right? And the hottest job in the late '80s was going to be knowledge engineer. You were going to sit next to an expert and then turn the knowledge of the expert into rules and facts, right? And then the computer would be able to basically do what the expert wants. This was manual behavior cloning. It kind of worked, but only for a few domains where economically it made sense and it was doable at the level of reliability that was good enough. But it was not a path towards human-level intelligence.
认为当前 AI 主流趋势会带我们走向人类智能的想法,在我的职业生涯中已经出现过三次,之前可能还有五六次。你应该看看人们当年对感知机的评价。我写过一篇《纽约时报》文章。人们说:“哦,我们将在 10 年内拥有超级智能机器。”60 年代马文·明斯基说:“哦,10 年内世界上最好的棋手将是计算机。”实际花的时间比那长得多。这种情况一再发生。
The idea that the current AI mainstream fashion is going to take us to human intelligence has happened already three times during my career and probably five or six times before. You should see what people were saying about the perceptron. I wrote a New York Times article. People were saying, 'Oh, we're going to have super intelligent machines within 10 years.' Marvin Minsky in the '60s says, 'Oh, within 10 years the best chess player in the world will be a computer.' It took a bit longer than that. This happened over and over again.
1956 年左右,当纽厄尔和西蒙提出通用问题求解器时,他们谦虚地称之为通用问题求解器。他们认为这很酷。他们说:“好吧,我们思考的方式很简单。我们提出一个问题。这个问题有若干不同的解决方案,不同的提议方案,一个潜在解的空间。比如旅行商问题,有 n 的阶乘条路径,可能的路径。你只需寻找最佳的那条。”他们说每个问题本质上都可以这样表述为寻找最佳解的搜索。如果你能通过编写一个程序来检查某个解是否好或给它打分,从而将问题表述为一个目标函数,然后你有一个搜索算法在可能的解空间中搜索一个优化该分数的解,那么你就解决了 AI。他们当时不知道的是所有复杂性理论,基本上每个有趣的问题都是指数级或 NP 完全的等等。所以我们不得不使用启发式编程。为每个新问题想出启发式方法。基本上,我们的通用问题求解器并不那么通用。
In 1956 or something, when Newell and Simon produced the general problem solver, very modestly called general problem solver. What they thought was really cool. They said, 'Okay, the way we think is very simple. We pose a problem. There is a number of different solutions to that problem, different proposals for solution, a space of potential solutions. Like traveling salesman, there is a number of factorial n factorial paths, possible paths. You just have to look for the one that is the best. And they said every problem can be formulated this way essentially as a search for the best solution. If you can formulate the problem as an objective by writing a program that checks whether it's a good solution or not or gives a rating to it. And then you have a search algorithm that searches through the space of possible solutions for one that optimizes that score, then you solve AI. What they didn't know at the time is all of complexity theory, that basically every problem that is interesting is exponential or NP-complete or whatever. And so we have to use heuristic programming. Come up with heuristics for every new problem. And basically, our general problem solver was not that general.
认为最新想法会带你走向 AGI(通用人工智能)或任何你想称呼的东西,这种想法非常危险,过去七十年里很多非常聪明的人多次落入这个陷阱。
This idea that the latest idea is going to take you to AGI or whatever you want to call it is very dangerous, and a lot of very smart people fell into that trap many times over the last seven decades.
你认为这个领域会解决持续学习或增量学习的问题吗?
Do you think that the field will ever figure out continual or incremental learning?
当然。是的,这算是一个技术问题。我想到灾难性遗忘,因为你花费大量资金训练的权重会被覆盖。但你可以只训练一小部分。我的意思是,我们不是已经用 SSL(自监督学习)这么做了吗?我们训练一个基础模型。比如视频或类似 Video BERT 的东西,能产生非常好的视频表示。然后如果你想为特定任务训练系统,你在上面训练一个小头部,这个头部可以持续运行。甚至你的世界模型也可以持续训练。这不是问题。坦白说,我不认为这是一个巨大的挑战。
Sure. Yeah, that's sort of a technical problem. I thought catastrophic forgetting, because your weights that you trained so much money on get overwritten. But so you train just a little bit of it. I mean, we don't already do this with SSL, right? We train a foundation model. Like for video or something like Video BERT produces really good representations of video. And then if you want to train the system for a particular task, you train a small head on top of it and that head can be run continuously. And even your world model can be trained continuously. That's not an issue. I don't see this as a huge challenge, frankly.
事实上,2005、2006 年,Raya Hadsel、Samy Bengio、我和几位同事构建了一个基于学习的移动机器人导航系统,就采用了这种思路。那是一个卷积网络,从摄像头图像做语义分割。在运行中,网络的顶层会适应当前环境。所以效果很好。标签来自短距离可通行性,基本上由立体视觉指示。所以,是的,你可以做到这一点。特别是如果你有一个多模态系统。我不认为这是一个大挑战。
In fact, Raya Hadsel and Samy Bengio and I and a few of our colleagues back in 2005, 2006 built a learning-based navigation system for mobile robots that had this kind of idea. So it was a convolutional net that was doing semantic segmentation from camera images. And on the fly, the top layers of that network would be adapted to the current environment. So it would do a good job. And the labels came from short-range traversability that were indicated by stereo vision, essentially. So, yeah, I mean you can do this. It's particularly if you have a multi-modal system. I don't see this as a big challenge.
很高兴能邀请到你。
It's been a pleasure to have you.
好的,我也非常高兴。非常感谢。谢谢。谢谢。
All right, it was really a pleasure. Thank you so much. Thank you. Thank you.