世界模型:下一场 AR 革命的推动力

World Models: Enabler for the Next AR Revolution

杨立昆 Yann LeCun · 苏黎世联邦理工学院计算机视觉与几何组 · 2026-06-09 · 约 59 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

探讨世界模型如何弥合机器学习与人类智能之间的差距,实现在复杂真实环境中的快速适应和零样本学习。

Exploring how world models can bridge the gap between machine learning and human-like intelligence, enabling rapid adaptation and zero-shot learning in complex real-world environments.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 28)

全文 · Full transcript(中英对照)

世界模型与莫拉维克悖论 World Models and the Moravec Paradox

Yann

是的,我要谈谈世界模型。它可能是下一场 AR 革命的推动力。房间里可能有很多机器学习的人。我有坏消息要告诉你们:机器学习很糟糕。当我们把机器的学习能力与人类和动物进行比较时,显然存在巨大差距。人和动物可以非常快速地学习新任务,只需要很少的尝试和样本。人有常识,动物也有,物理常识。有很多任务我们可以零样本完成,即使以前从未遇到过。我们如何用机器做到这一点?我们有非常强大的 AI 技术,每个人都在用,但它们并不能真正处理现实世界。它们无法处理连续的、高维的、有噪声的数据。相比之下,语言很容易。现实世界是混乱的,语言是简单的。这与 Vladlen 和 Jitendra 之前提到的 Moravec 悖论有关:对计算机来说简单的事情对人类来说困难,而对人类来说复杂的事情对计算机来说并不那么困难,比如下棋、符号积分、解方程、证明数学定理等。那么,一个 10 岁的孩子怎么能做到你想让家用机器人做的事情,而且大多数任务在第一次被要求时就能完成,而无需专门训练?他们可能不想做,但能做到。为什么一个青少年可以在几小时的练习中学会开车,而自动驾驶汽车公司拥有数百万小时的训练数据,却无法用这些数据让机器仅仅通过模仿人类就能达到同样水平的可靠性?否则,我们就会有 L5 级自动驾驶汽车,但我们没有。在消费级汽车领域,最多只有 L2 或 L3 级,而机器人出租车是经过大量工程设计的,配备了各种传感器和其他东西。所以我们不断遇到这个 Moravec 悖论,如果你相信智能需要具身,我们就必须超越这一点。当然,一些哲学家,当然还有一些语言学家,认为这并非必要,但我认为是必要的。

Yeah, I'll talk about world models. Possibly the enabler for the next AR revolution. So, there's a lot of machine learning people I think in the room, perhaps. I have bad news for you. Machine learning sucks. You know, basically when we compare the learning abilities of machines with humans and animals, clearly there is a big gap. People and animals can learn new tasks extremely quickly and with very few trials, very few samples. People have common sense, animals too, physical common sense. There's a lot of tasks that we can accomplish zero shot even if we've never faced them before. And how do we do this with machines? We have very powerful AI techniques that everybody is using, but they don't really handle the real world. They don't handle continuous, high-dimensional, noisy data. Language is easy by comparison. The real world is messy. Language is simple. This connects with what Vladlen said earlier and Jitendra as well, the Moravec paradox: things that are simple are difficult for computers and things that are complicated for humans turn out not to be that difficult for computers, like playing chess, computing integrals symbolically, solving equations, proving math theorems, etc. So, how is it that a 10-year-old can do what you would like a domestic robot to do and do most of those tasks without actually being trained to do them the first time you ask them? They may not want to do it, but they can. How come any teenager can learn to drive a car in a few hours of practice, yet the self-driving car companies have literally millions of hours of training data, and despite that, they can't use those millions of hours to get a machine to just imitate humans to drive at the same level of reliability? Otherwise, we'd have level five self-driving cars, and we don't. At best, in the consumer car business, we have level two or three, and the robo-taxis are very heavily engineered with various sensors and other things. So, we keep bumping into this Moravec paradox, and we really have to go beyond this if you believe that intelligence requires grounding. Of course, some philosophers, and certainly some language people, don't believe that's necessary, but I think it is.

皮亚杰与智能本质 Piaget and the Nature of Intelligence

Yann

像 Vladlen 一样,我们在瑞士,靠近让·皮亚杰的家乡。他对我影响很大。他在 1970 年代末与诺姆·乔姆斯基在法国进行了一场辩论,讨论语言是先天的还是习得的。那场辩论有记录,参与者中有一位曾与皮亚杰共事,是麻省理工学院的教授,他谈到了感知机,说那些简单的机器学习模型能够学习令人惊讶的复杂任务,这可能证明学习是可能的,与乔姆斯基的观点相反。这个人就是西摩·帕珀特。他是麻省理工学院的教授,而在那之前的 10 年,他写了一本书,基本上扼杀了整个神经网络领域,指出了感知机的局限性。但 10 年后,他却在这里论证这些东西实际上值得研究。不管怎样,让·皮亚杰说:‘智能不是你知道什么,而是你不知道时做什么。’事实上,他从未说过这句话。这是杜撰的。但有一些心理学家将他的思想提炼成了这句话,他从未说过。所以他被引用了这句话。所以智能不是陈述性知识的积累。LLM 是陈述性知识的积累,它们有用的主要原因就是能积累大量陈述性知识。智能不是技能的集合。你也许可以建造一台机器来完成任何任务,只要投入足够的资源,包括自动驾驶之类的事情。但这并不是真正的智能。智能是在大约 20 小时内学会开车的能力。或者用很少的训练学习任何新任务,或者零样本完成新任务。这才是真正的智能,也是皮亚杰的意思。这意味着我们不会有简单的智能衡量标准,因为任何特定任务,你总是可以投入足够的时间和精力来攻克它。所以更重要的是你的适应能力。这与 Vladlen 所说的有关,比如 AGI 的概念完全是胡说八道。人类智能是专门的。人类智能的特点是快速适应,我们能学习新任务。我们每个人都知道不同的知识,拥有不同的技能。这是因为我们接触了不同的环境,解决了不同的问题。我们是适应性的。这才是真正的智能。

Like Vladlen, we're in Switzerland outside Jean Piaget. He was a big influence on me. He had a debate with Noam Chomsky in France in the late 1970s, where they were debating whether language was innate or learned. There were transcriptions of that debate with people participating, and one of them was a guy who had worked with Jean Piaget, who was a professor at MIT, and was talking about the perceptron, saying there are those simple machine learning models that are capable of learning surprisingly complex tasks, and that may be evidence that learning is possible, contrary to what Chomsky was saying. This guy was Seymour Papert. He was professor at MIT, and 10 years before that, he had written a book that basically killed the entire field of neural nets, pointing out the limitation of the perceptron. But here he was, 10 years later, arguing that those things were actually interesting to study. Anyway, Jean Piaget says, 'Intelligence is not what you know, it's what you do when you don't know.' In fact, he never actually said this. This is apocryphal. But there are other psychologists who basically kind of distilled his thinking into this sentence, which he never said. So he's kind of quoted as saying that. So intelligence is not an accumulation of declarative knowledge. LLMs are an accumulation of declarative knowledge, not just, but the main reason they're useful is because they can accumulate a lot of declarative knowledge. Intelligence is not a collection of skills. You can probably build a machine to accomplish any task if you spend enough resources on it, including things like self-driving. But that's not really what intelligence is. Intelligence is the ability to learn to drive in about 20 hours. Or to learn any new task with very little training, or accomplish new tasks in a zero-shot way. That's really what intelligence is, and that's really what Jean Piaget means. That means we're not going to have any simple measures of intelligence because any particular task, you can always, if you spend enough effort and time, you can always kind of crack it. So it's more about how adaptive you are. This connects to something that Vladlen said, like the notion of AGI is complete nonsense. Human intelligence is specialized. The characterization of human intelligence is that it's very quickly adaptive and we can learn new tasks. We all know different sets of knowledge and have different skills. It's because we've been exposed to different environments and we've had to solve different problems. We're adaptive. That's really what intelligence is.

人类如何学习:观察与直觉物理 How Humans Learn: Observation and Intuitive Physics

Yann

那么,人类和动物是如何学习的呢?生命早期有很多学习是通过观察进行的。所以,两个月大的婴儿可以发展出自己肢体的动力学模型,但基本上无法影响世界。它不能移动物体或其他东西。但它可以学到很多关于世界的东西。婴儿能很快学会的一件事是世界是三维的。为什么?因为物体与我们之间的距离是解释当我们移动头部时视野如何变化的最佳方式。当然,婴儿不一定自己移动头部,但他们被移动,所以他们看到视差,并由此推导出世界是三维的。我们今天可以用学习机器做到这一点。它们仅仅通过被动观看视频就能学会世界是三维的。所以这很有趣。像物体恒存这样的基本概念学得很快。稳定性、刚性等概念也是如此。但是,我们所谓的直觉物理,比如惯性、重力,实际上人类婴儿需要 9 个月才能学会。大多数动物则更快。

Okay, so how do humans learn and animals? There's a lot of learning that takes place in the early months of life mostly by observation. So, a two-month-old baby can develop a dynamical model of its own limbs but basically cannot affect the world. It can't move an object or anything. But it can learn a lot of things about the world. One thing a baby can learn really quickly is that the world is three-dimensional. Why? Because the fact that an object has a distance from us is the best way to explain how our view of the world changes when we move our head. And of course babies don't necessarily move their head but they are being moved, so they see parallax and derive from this the fact that the world is three-dimensional. We can do this with learning machines today. They learn that the world is three-dimensional only by being exposed passively to videos. So that's an interesting thing. Basic concepts like object permanence are learned really quickly. Notions of stability, rigidity, and things like that. But then, what we would consider intuitive physics, things like inertia, gravity, that actually takes 9 months for human infants. Shorter for most animals.

婴儿与机器对重力的学习 Learning about gravity in infants vs. machines

Yann

如果你把一个 8 个月大的婴儿放在高脚椅上,旁边放一堆玩具,孩子很可能会系统性地把所有玩具扔到地上,观察结果。他们是在做实验,验证重力适用于一切。这需要很长时间。这是怎么发生的?是什么样的学习在发生?他们在做实验,但也可以通过观察来了解重力。如果你展示一个场景:一辆车在平台上,你把它推下去,它似乎浮在空中,6 个月大的婴儿几乎不会注意——他们还没学会重力。10 个月大的婴儿会非常惊讶,就像那个小女孩。这就是心理学家衡量婴儿是否学会了某个世界概念的方式:违背预期。我们可以用这些技术来测试机器学习系统是否获得了某种常识。关于这一点有很多可以说的。Jitendra 和我合作了一篇论文,主要由 Emmanuel Dupoux 撰写,Jitendra 和我的贡献很小。

If you put an 8-month-old on a high chair with a bunch of toys, the child will likely systematically take all the toys and throw them on the floor, watching the result. They're doing an experiment that gravity applies to everything. That takes a long time. How does that happen? What type of learning is taking place? They're doing the experiment, but they can also learn about gravity just by observation. If you show a scenario where a car is on a platform and you push it off, and it appears to float in the air, a 6-month-old will barely pay attention—they haven't learned about gravity yet. A 10-month-old will look very surprised, like the little girl. That's how psychologists measure whether a baby has learned a particular concept about the world: violation of expectation. We can use those techniques to test whether machine learning systems have acquired some notion of common sense. There's a lot that can be said about this. Jitendra and I collaborated on a paper, mostly written by Emmanuel Dupoux, with Jitendra and I having very little contribution to this set of questions.

智能定义与 AGI 怀疑论 Definition of intelligence and AGI skepticism

Yann

如果智能不是技能或陈述性知识的积累,那它是什么?它是无需事先训练就能完成新任务、解决新问题的能力。AGI 这个短语毫无意义。人类智能是专门化的。问题不在于‘你是否知道如何做所有事?’,而在于‘你能快速学会如何做任何事吗?’或者广泛的事情。底部有一篇有点哲学性的论文,由我的一些年轻同事撰写。

What is intelligence, if not an accumulation of skills or declarative knowledge? It's the ability to accomplish new tasks, solve new problems without prior training. AGI makes no sense as a phrase. Human intelligence is specialized. The question is not 'Do you know how to do everything?' but 'Can you learn quickly how to do anything?' or a wide spectrum of things. There's a somewhat philosophical paper at the bottom written by some of my young colleagues.

扩展 LLM 不会产生类人智能 Scaling LLMs won't lead to human-like intelligence

Yann

仍然有很多人,尤其是在美国西海岸,相信通过扩大 LLM 规模,也许在合成数据上训练,使用后训练和强化学习中的一些技巧,我们就能达到他们所谓的 AGI。我认为这是不可能的。我相信具身智能。一个简单的计算:今天典型的 LLM 在约 20 万亿个单词上训练,对应约 30 万亿个 token。每个 token 是 3 字节,所以数据量约为 10^14 字节。这需要任何人花大约 40 万年才能读完。相比之下,一个 4 岁孩子一生中看到的东西:大约 16 小时的清醒时间,相当于约 30 分钟的 YouTube 上传内容。我们有 200 万根视神经纤维,每根每秒传输约 1 字节。所以通过视觉获得的数据量约为 10^14 字节——与 40 万年的文本数据量相同。有了互联网上所有人类产生的文本,仅靠文本训练,我们无法获得任何接近人类智能的东西。这是不可能的。你可能会说视频比文本冗余得多,但这是特性,不是缺陷。如果你想训练一个系统,特别是使用自监督学习,你需要冗余。没有冗余,你什么也学不到。冗余是好事,但你不能有太多。

There are still many people, particularly on the US West Coast, who believe we'll reach what they call AGI by scaling up LLMs, maybe training on synthetic data, using a few tricks in post-training and reinforcement learning. I think that's impossible. I believe in grounded intelligence. A simple calculation: a typical LLM today is trained on about 20 trillion words, corresponding to about 30 trillion tokens. Each token is 3 bytes, so the data volume is about 10^14 bytes. That would take about 400,000 years for any human to read. Compare that with what a 4-year-old has seen during their life: about 16 hours of wake time, which is about 30 minutes of YouTube uploads. We have 2 million optic nerve fibers carrying about 1 byte per second each. So the data volume through vision is about 10^14 bytes—the same amount as 400,000 years of text. With all human-produced text on the public internet, we're not going to get anything like human-like intelligence by just training on text. It's just not going to happen. You might say video is much more redundant than text, but that's a feature, not a bug. If you want to train a system, particularly using self-supervised learning, you need redundancy. Without redundancy, you can't learn anything. Redundancy is good, but you don't want too much of it.

智能系统的属性:推理模式 Properties of intelligent systems: inference mode

Yann

另一个问题是关于智能系统的正确属性。一个重要属性是推理模式。它是通过固定层数的神经网络传播来计算输出吗?还是考虑另一种方式:通过搜索与输入最兼容的输出来计算输出。你观察一个场景,通过感知模块产生当前世界状态的表示。你可以直接产生一个动作——那是反应式系统。或者你可以想象一个动作,让智能系统判断这个动作对于这个观察是否合适,是否能完成任务。这里的客观函数表征任务是否完成。把它看作一个成本函数,不用于学习,而用于推理。就像概率推理模型中的负对数似然,或者我更喜欢的能量函数。推理是在推理时搜索最小化某个能量函数的输出的过程。这在计算上本质上比固定层数的传播更强大。

Another question is about the right properties of intelligent systems. An important property is the mode of inference. Does it compute its output by propagating through a fixed number of layers of some neural net? Or consider the alternative: computing the output by searching for an output that is most compatible with the input. You observe a situation, run it through some perception module to produce a representation of the current state of the world. You can directly produce an action—that's a reactive system. Or you could imagine an action and have the intelligent system figure out if it's a good action for this observation, whether it will accomplish the task. The objective here characterizes whether the task has been accomplished. Think of it as a cost function, not used for learning but for inference. Like negative log likelihood in a probabilistic inference model, or as I prefer, an energy function. Inference is a process of searching for an output that minimizes some energy function at inference time. That's intrinsically more powerful computationally than just propagation through a fixed number of layers.

LLM 推理与世界模型推理对比 Contrasting LLM inference with world model inference

Yann

对比左边的模型,它类似于 LLM:取一个输入窗口,通过一个具有几千亿参数的固定层数的大神经网络,产生一个 token。然后将该 token 移入输入,产生下一个 token,以此类推。这是自回归预测,每个 token 涉及通过固定层数的固定计算量。这不是一个好的推理模型。你强迫 LLM 进行推理的方式是诱使它生成更多 token。但这不是我们推理的方式。我们在内部推理,不是在 token 空间,甚至不是语言。对比右边的模型,它是前一个模型的轻微特化。你感知世界,了解当前状态,然后想象一系列动作,一个动作提议。将其输入内部世界模型,该模型预测结果,然后将结果输入一个衡量任务完成程度的客观函数。然后通过优化,你搜索一个最小化这个能量的动作序列。在推理时——我还没谈到学习。在我看来,这是一个更强大的模型。但你需要一个世界模型。

Contrast the model on the left, which is LLM-like: take a window of inputs, run through a fixed number of layers of a big neural net with a few hundred billion parameters, produce one token. Then shift that token into the input and produce the next token, etc. That's auto-regressive prediction, and every token involves a fixed amount of computation through a fixed number of layers. This is not a good model of reasoning. The way you coerce an LLM to do reasoning is by tricking it into generating more tokens. But that's not how we reason. We reason internally, not in token space, not even in language. Compare this with the model on the right, a slight specialization of the previous one. You perceive the world, get an idea of the current state, then imagine a sequence of actions, a proposal for an action. Feed it to an internal world model, which predicts the outcome, and then feed that to an objective that measures how well the task has been accomplished. Then by optimization, you search for an action sequence that minimizes this energy. At inference time—I haven't talked about learning yet. In my opinion, that's a much more powerful model. But you need a world model.

能量最小化的推理与规划架构 Architecture for Reasoning and Planning via Energy Minimization

Yann

现在,如果你有……我大约五年前确定了这种想法或架构。我写了一篇长论文,2022 年放到了网上,包含一些通用架构等。如果你想拍照,这里有二维码,可以访问。它相对易读,但有点长。它基于这样一个理念:推理和规划是必不可少的,它们基本上通过能量最小化而非前向传播进行。而且,要做到这一点,你需要一个世界模型。所以,和我之前描述的过程一样,有一些额外的技巧。你观察环境,感知模块产生世界初始状态的表示,但只表示你当前感知到的内容。因此,你可能需要将其与记忆中的内容结合起来,以获得对世界状态的完整认识,至少是你所知道的。然后,你将其连同提议的动作序列一起输入世界模型,世界模型预测该动作序列的结果。你将其输入一个目标函数,即能量函数,它衡量特定任务完成的程度。因此,如果任务完成,该函数输出零;如果任务未完成,则输出某个正数,也许还衡量与任务完成之间的距离。直观地说,你可以有另一组目标作为护栏,确保系统将要经历的任何状态序列不会杀死任何人、伤害任何人或产生任何有害影响。因此,以这种方式构建的系统可以具有内在安全性,因为它必须服从并优化每个输出产生的护栏目标。这对于大语言模型来说并非如此。大语言模型,使其安全或无毒的唯二方法是通过微调。而且总有办法打破条件化,即越狱系统。在这里,你无法越狱这样的系统。它除了优化护栏目标和任务目标之外什么也做不了。当然,如果你有一个世界模型,在座的很多机器人系统最优控制专家都知道,你可以将这个世界模型应用于多个时间步,每个动作序列可以分解为一个序列。护栏可以应用于序列中的所有步骤。这就是你使用世界模型的方式。而通过优化进行规划的方式类似于模型预测控制(MPC),这是最优控制中非常经典的东西,可以追溯到 20 世纪 60 年代。

Now, if you do have... I've settled on this kind of idea or architecture about 5 years ago. I wrote a long paper about it that I put online in 2022 with some general architecture, etc. If you want to take pictures, here are QR codes, you can get to it. It's relatively easy to read, but kind of long. And it's really based on this idea that reasoning and planning are essential, and they basically proceed by energy minimization rather than forward propagation. And that for this to work, you need some world model. So, same process that I described before, there are a few additional tricks. You observe the environment, perception module produces a representation of the initial state of the world, but only a representation of what you currently perceive. So, you may have to combine this with the content of the memory to get a complete idea of the state of the world, what you know about it at least. Then you feed this to your world model together with a proposal for an action sequence, and your world model predicts the outcome of that action sequence. You feed this to an objective, an energy function that measures to what extent a particular task has been accomplished. So, this function outputs zero if the task is accomplished and some positive number if the task is not accomplished, and perhaps measures some distance to the task being accomplished. So, intuitively, you can have another set of objectives that are guardrails that would ensure that whatever state sequences the system is going to take the world through is not going to kill anyone or hurt anyone or have any kind of deleterious effect. And so, a system constructed this way can be made intrinsically safe because it has to obey and optimize the guardrail objective with every output it produces. This is not the case for an LLM. An LLM, the only way it can be made safe or non-toxic is by fine-tuning it. And there is always a way to break the conditioning, to jailbreak the system. Here, you can't jailbreak a system like this. It can do nothing but optimize the guardrail objectives and the task objective. Of course, if you have a world model, certainly a lot of robotics system optimal control people in the room, you can apply this world model multiple time steps and each action sequence can be decomposed into a sequence. The guardrails can be applied to all the steps in the sequence. That's the way you would use a world model. And the way you plan by optimization there is akin to model predictive control, MPC, very classical stuff in optimal control going back to the 1960s.

分层规划与训练世界模型 Hierarchical Planning and Training World Models

Yann

但最终,你想要的是能够进行分层规划的系统。我们所有人都会分层规划。动物也会分层规划。什么是分层规划?假设我坐在纽约大学的办公室里,明天想去巴黎。我无法以 10 毫秒为单位的肌肉动作来规划整个巴黎之旅,那是人类的基本动作。我做不到,因为首先,时间太长了。其次,我没有信息。我不知道当我走到街上时,要等多久才能打到出租车,对吧?所以,我无法规划整个事情。我必须进行分层规划。所以,我要做的是在高层说:‘嗯,我不知道去机场要花多长时间,但大概一个到一个半小时。所以,我需要去机场赶飞机。好了,这是一个两步的高层计划。我不需要知道很多细节就能制定这个计划。现在,我有一个子目标,即到达机场。我是说,在纽约,去机场包括走到街上、打车、去机场。现在我需要走到街上。我在纽约大学的一栋楼里,这包括走到电梯、按按钮、下楼、走出门。现在,我有一个子目标是到达电梯,对吧?等等。所以,你可以沿着整个层级往下走,到了某个点,你需要采取的动作非常简单。这是你熟悉的事情。你可能不需要动用全部脑力来规划这个动作。你大概可以不用思考就从椅子上站起来。那可能只是策略。但本质上,我们最终希望系统能够进行分层规划。我们如何解决这个问题?这是一个未解决的问题。如果你是一个机器人专家,或者从事人工智能机器人领域,或者智能体人工智能领域,如果你正在攻读这个主题的博士学位,这是一个很好的主题。它完全是开放的。没有人知道如何做到这一点。没有人证明他们知道如何做到这一点。好了,那么现在的大问题是我们如何训练这些世界模型?分层的还是不分层的?假设先不分层。首先,我们必须弄清楚给它们什么架构。在当今时代,自然的直觉是训练一个生成模型。事实上,我一直在尝试训练类似世界模型的东西,大约 15 年了,前 10 年基本都失败了。因为我试图训练生成模型。什么是生成模型?自监督学习在语言领域取得了令人难以置信的成功,对吧?你取一串单词,去掉一些单词,破坏输入,然后通过一个大型神经网络运行破坏后的输入,训练它恢复缺失的部分。这对文本非常有效。所以有像 BERT 这样的原始模型做过这个。大语言模型是这种情况的一个特例,其中你只去掉最后一个单词。所以整个系统只是试图生成序列中的下一个词。但如果做得对,它非常有效。但如果你把它应用到视频上,它就不起作用了。所以,如果你拿一个视频,向系统展示视频的初始片段,并要求它在像素级别预测接下来会发生什么,它实际上并不奏效。你从系统中得到的视频表示并不特别好。原因是你根本无法预测视频中发生的一切。有无数种可能的事情。在文本中,这很容易,因为只有有限数量的单词,所以你可以让系统产生一个概率分布,覆盖字典中所有可能的单词或词元。

Ultimately, what you want though is something that can do hierarchical planning. All of us do hierarchical planning. Animals do hierarchical planning. What is hierarchical planning? Let's say that I'm sitting in my office at NYU and I want to be in Paris tomorrow. There's no way I can plan my entire trip to Paris in terms of muscle actions 10 ms by 10 ms, which are the elementary actions that humans can do. I can't do that because first of all, it's too long. But second of all, I don't have the information. I don't know if when I'm going down on the street how long I'm going to have to wait before a taxi stops. Right? So, there's no way I can plan the entire thing. I have to do hierarchical planning. So, what I have to do is at a high level, I have to say, 'Well, I don't know how long it's going to take me to go to the airport, but maybe roughly an hour and an hour and a half. So, I need to get to the airport and catch a plane. Okay, that's a two-step high-level plan. I don't need to know many details to make that plan. And now I have a sub-goal, which is being at the airport. I mean, New York, so going to the airport involves going down on the street and hailing a taxi and go to the airport. Now I need to go down on the street. I'm in a NYU building, that involves walking to the elevator, pushing the button, getting down, and walking out the door. Now I have a sub-goal of getting to the elevator, right? Etcetera. So, you can sort of go down this entire hierarchy and at some point you get to a point where the action you need to take is very simple. It's something that you are familiar with. You may not have to use your full mental power to plan the action. You can probably stand up from your chair without having to think about it. That could be just policy. But essentially, ultimately, we want systems to do hierarchical planning. How do we solve that? This is an unsolved problem. If you're a roboticist or an AI for robotics kind of person or agentic AI kind of person, if you're studying a PhD on this topic, this is a great topic. It's completely open. Nobody knows how to do this. Nobody has proved that they know how to do this. Okay, so now the big question is how are we going to train those world models? Hierarchical or not? Let's say non-hierarchical to start. So first of all, we have to figure out what architecture to give them. And a natural instinct in these days and age is to train a generative model. And in fact, I've been working on trying to train world model like things for about 15 years, mostly failing for the first 10. Because I was trying to train generative models. What's a generative model? So self-supervised learning has been incredibly successful, astonishingly successful in the context of language, right? You take a string of words, you remove some of the words, you corrupt the input, and then you run the corrupted input through some big neural net, and you train it to recover the missing parts. That works amazingly well for text. So there are original models like BERT that used to do this. An LLM is a special case of this where the only word you remove is the last one. So the entire system is trying to just produce the next word in a sequence. But it works amazingly well if you do it right. It doesn't work if you apply it to video. So if you take a video and then you show the initial segment of the video to the system, and you ask it to predict what's going to happen next at a pixel level, it doesn't really work. The representations you get out of the system for your video are not particularly good. And the reason is you simply cannot predict everything that takes place in a video. There's an infinite number of plausible things. In text, it's easy because there is only a finite number of words, and so you can get the system to produce a probability distribution over all possible words or tokens in your dictionary.

视频预测的挑战 Challenges of Video Prediction

Host

但视频不行,对吧?可能的视频帧数实在太庞大了。

But, you can't do this with video, right? There's just an incredibly large number of possible video frames.

Yann

举个例子。如果我拍一段这个房间的视频,我从这里开始,慢慢旋转相机,停在这里,然后让系统继续生成视频。它可能会预测我们是在一个空教室或礼堂里,房间大小有限,这边可能有窗户等等。但系统绝对无法预测你们每个人的样子,或者哪些椅子是空的。这根本不可能,因为你没有足够的信息。所以,当你训练系统做这种预测时,你就毁了它。

Let me take an example. If I take a video of this room, right? I start here, and I kind of slowly rotate the camera, I stop here, and I ask the system to continue the video. You know, it's probably going to predict, you know, we are in some sort of empty classroom, auditorium, and you know, the room has a finite size, there might be windows on this side, and things like that. There's absolutely no way the system can predict what all of you look like. Or which, you know, chairs are unoccupied. It's just impossible. You just don't have the information. So, when you train a system to make this kind of prediction, you kill it.

Host

当然,你会说:“哦,但我们可以训练系统生成好看的视频,对吧?视频生成。”

Now, of course, you're going to tell me, "Oh, but we can train system to produce cute videos, right? Video generation."

Yann

是的,但这种预测通常是在表示空间进行的,而不是像素空间。只有在第二阶段才将预测转化为高分辨率、高帧率的视频。而且系统只需要生成一个好看的视频,不需要真正表示所有可能的视频。所以这是一个简单得多的问题。

Yes, but this prediction usually is done in representation space, not in pixel space. It's only a second stage that actually turns the predictions into high-resolution, high-frame-rate videos. And the system only needs to produce one cute-looking video. It doesn't need to actually represent all plausible videos. So, which is a much simpler problem.

Yann

正如我所说,过去 15 年的大部分时间里我都在尝试解决这个问题。这是一篇 10 年前的论文,我们试图训练一个神经网络,从四帧上下文预测两帧短视频片段。结果得到模糊的预测。为什么?因为系统预测的是所有可能情况的平均值。当然,你可以用潜在变量模型(比如扩散模型)来修正,但我们当时不知道。我们尝试了 GAN 等方法,不太成功。但也许使用潜在变量模型会有帮助,尤其是扩散模型,它们当然能生成好看的视频。但它们真的理解世界吗?证据显示没有。

As I said, I've been kind of attempting to work on this for the better part of the last 15 years. So, this is a 10-year-old paper where we tried to train some neural net to predict short video clips, you know, two frames from four frames of context. You get blurry predictions. Why? Because the system predicts the average of everything that can happen. Of course, you can correct that with latent variable models, like diffusion models, but we didn't know that at the time. We tried to use GANs and stuff like that. Wasn't too successful. But perhaps using latent variable models would help, diffusion models in particular, which of course produce cute videos. Do they actually understand the world? The evidence is no.

联合嵌入预测架构(JEPA) Joint Embedding Predictive Architecture (JEPA)

Yann

所以,这是我的解决方案。我称之为联合嵌入架构,更准确地说,是联合嵌入预测架构(JEPA),如右图所示。左边是生成式架构。你观察到 X,可能还观察到动作 A,以及结果 Y,系统试图以最精细的细节重建 Y。而在 JEPA 中,你观察到 X、Y 和 A,但你对 X 和 Y 都进行编码,预测在表示空间中进行。这是主要区别。系统通过构建 Y 的表示,可以消除输入中那些不可预测的信息。这使得预测更抽象、细节更少,但在某种程度上更准确。

So, here's my solution. My solution is an architecture I called joint embedding, or more precisely joint embedding predictive architecture, JEPA, which is shown on the right. On the left you have generative architecture. You observe X, maybe you observe A, an action that is taking place, and you observe the result Y, and the system is trying to reconstruct Y in its most minute details. With JEPA, you observe X and Y and A, but you encode both X and Y, and the prediction takes place in that representation space. Major difference. What the system can do is essentially eliminate from the input, by constructing a representation of Y, it can eliminate all the information about Y that is simply not predictable. Right? And that makes the prediction more abstract, with fewer details, but more accurate in a way.

Yann

如何训练生成模型?训练生成模型很容易,因为代价函数只是重建损失。你只需训练它重建。你可以把它训练成自编码器,但需要限制编码中的信息量,或者训练成去噪自编码器,这是很多技术(如掩码自编码器)尝试的方法。这意味着对输入进行某种破坏,然后训练自编码器恢复原始输入。顺便说一句,扩散模型是这种去噪方法的一个特例。

How do you train a generative model? It's easy to train a generative model because the cost is just a reconstruction cost. It's just going to, you know, you're just training it to reconstruct. You can train it as an autoencoder, but then you need to restrict the information content in the code, or as denoising autoencoder, which is what a lot of techniques have attempted to do like masked autoencoders and things like that. So that means taking an input, corrupting it in some ways, and then training an autoencoder to recover the initial one. And by the way, diffusion models are a bit of a special case of this sort of general thing of denoising.

Yann

关键在于,当你用这类系统学习图像表示时,你得不到好的表示。如果你用这种方式获得的图像表示,输入到有监督的下游任务中,结果并不好。要获得好结果,你必须使用联合嵌入架构。所有使用自监督学习训练图像或视频表示的最佳系统,都使用联合嵌入,没有一个使用重建。所有最好的都是这样。

The value is when you train systems of this type to learn representations of images, you don't get good representations. If you use the representation of images obtained this way, you feed it to a downstream task that you train supervised, you train a head supervised. The results you get are not great. To get good results, you have to use joint-embedding architectures. All the best systems that use self-supervised learning to train an image or video representation system, all use joint-embedding. None of them uses reconstruction. All the best ones.

Yann

要么你将其应用于图像:你有同一场景的两个视图,训练神经网络生成表示,并告诉系统“我希望这两个表示相同”。要么你使用破坏技术:取一个输入,以某种方式破坏或变换它,然后训练 JEPA 架构从破坏版本的表示预测原始图像的表示。

Either you apply this to images, either you have two views of the same scene, and you train a neural net to produce representations, and you tell the system, "I want those two representations to be identical." Or you use this corruption technique. You take an input, you corrupt it, or transform it in some ways, and then you train this JEPA architecture to predict the representation of the original image from the representation of the corrupted version.

Yann

这有一个大问题:系统可能会崩溃。生成模型在某种程度上也会崩溃。比如,如果你尝试训练一个自编码器而不限制编码的信息量,它只会学到恒等函数,这就是崩溃,学不到任何有用的东西。同样,这类系统也会崩溃。如何崩溃?它基本上完全忽略输入,产生恒定的表示,预测问题变得微不足道。所以,如果你只是试图最小化预测误差,系统就会崩溃,不会做任何有用的事。因此,自监督学习用于联合嵌入系统的全部技巧就在于如何防止崩溃。

There's a big issue with this, which is that the system can collapse. The generative models can actually collapse to some extent. Like, if you try to train an autoencoder without a restriction on the information content of the code, your autoencoder is just going to learn the identity function, and that's a collapse. It's not going to learn anything useful. Similarly, a system like this can collapse, and how can it collapse? It can essentially completely ignore the inputs, produce constant representations and the prediction problem is trivial. So, if you're just trying a system of this type to minimize the prediction error, it's going to collapse. It's not going to do anything useful for you. So, the whole trick of how you do self-supervised learning for joint embedding system is how you prevent collapse.

Yann

我有个最喜欢的概念。我会讨论其他方法,但我最喜欢的是信息最大化。你基本上设计一个目标函数,衡量编码器输出的表示的信息量,然后尝试最大化这个信息量。所以代价函数是负信息或类似的东西。

There is my favorite concept for this. I'll talk about other ways to do this but my favorite concept to prevent collapse is information maximization. So, you basically come up with some objective function that measures some sort of information content of the representation that comes out of your encoders. And you try to maximize that information content. So, your cost function is minus the information or whatever.

Yann

过去六七年里,有很多技术,比如 MNCR、NCR 平方、WMSE、Seegrid、VICReg 和 Barlow Twins。Barlow Twins、VICReg、Seegrid 来自和我一起工作的人,其他来自其他团队。MNCR 来自伯克利,NCR 平方来自纽约大学神经科学的同事 Sam Chelleli。所以 JEPA 这个想法越来越流行。Google Scholar 上大约有 1700 篇论文明确提到了联合嵌入预测架构。

There's a bunch of techniques for this since like the last 6 or 7 years with names like MNCR, NCR squared, WMSE, Seegrid, VICReg, and Barlow Twins. The Barlow Twins, VICReg, Seegrid come from people working with me. The other ones from other groups. MNCR comes from Berkeley and NCR squared from a colleague at NYU Neuroscience. It was Sam Chelleli. So, this idea of JEPA is getting popularity. There's about 1,700 papers that mention joint embedding predictive architecture spelled out on Google Scholar.

Yann

这类方法有一个问题:如何衡量信息量?我们需要一个可微的信息量度量作为代价函数,这样我们才能反向传播梯度并最大化它。

There's an issue with this type of method, which is how do you measure information content? We need to have a cost function that is a differentiable measure of information content, so we can backpropagate gradients and maximize it.

测量信息内容的挑战 Challenges in Measuring Information Content

Yann

坏消息是,首先,我们实际上没有客观的信息含量度量,因为所有恰当的定义都基于知道你要测量信息含量的向量或任何东西的分布。但我们不知道分布。我们只有来自编码器的样本。那么,如何从有限数量的样本计算信息含量?这是第一个问题。第二个问题是,要最大化某个东西,你需要信息含量的下界,这样当你最大化时,你才能推高实际的信息含量。问题是,我们拥有的每个经验度量都是上界。那么,我们该怎么办?我们想出一个好的上界,然后祈祷。我们展示一些定理之类的。

And the bad news is, first of all, we don't actually have objective measures of information content because all the proper definitions are based on knowing the distribution of the vectors or whatever that you want to measure the information content of. And we don't know the distribution. We only have samples coming out of an encoder. So, how do you compute information content from a finite number of samples? That's the first problem. Second problem is to maximize something you would need a lower bound on information content so that when you maximize, you push the actual information content up. Problem is, every empirical measure that we have are all upper bounds. So, what do we do? We come up with a good upper bound and we cross our fingers. And we show some theorems and whatever.

基于能量的模型框架 Energy-Based Models as a Framework

Yann

所以,这种技术以及许多其他技术,比如正确解释如何训练自监督学习系统以及实际上每个学习系统的方法,是我称之为基于能量的模型的框架,我已经倡导了大约 20 年。基本思想是这样的。如果你想捕捉两个变量 X 和 Y 之间的依赖关系,X 和 Y 之间并没有真正的函数关系。所以,你不能对给定的 X 有一个单一的 Y。对吧?它只是一种依赖关系,而不是一个函数。就像是一种关系或某种映射,但不是函数。所以,如右边图表所示,你有一堆数据点,那些是黑点。它们指示了 X 和 Y 之间的某种依赖关系。既然你不能运行一个从 X 计算 Y 的函数,你如何捕捉这种依赖关系?一种方法是学习或构建一个对比函数,即能量函数,它告诉你这个 XY 空间中的点是否靠近训练数据。所以,我们把它想象成某种景观,黑点在山谷里。在瑞士,那会是一个湖。然后,你知道,你会得到等高线,对吧?当你向外移动,离开那些区域,海拔升高。能量升高。现在,如果我给你一个 X 值,你可以推断出一堆与 X 兼容的 Y 值。它们是使能量最小化的 Y 值。所以,这就是我之前谈到的推理,通过优化进行推理,而不是前向传播。但你也可能反过来做。如果我给你一个 Y,你可以从 Y 推断 X。而且你可以给我多个答案。所以,在像视频预测这样的情况下,基本上有无限数量的可能答案,训练这类系统的正确方法是基于能量的模型来思考。顺便说一句,概率模型是一个特例,其中你的能量有特定形式,训练方式有特定损失函数。所以,如果你愿意,它是一个比概率推理和学习稍微更通用的框架。

So, this technique and many others, like the way to properly explain how you can train self-supervised learning systems and every learning system really, is a framework I call energy-based models that I've been advocating for 20 years or so. It's basically the basic idea is like this. If you want to capture the dependency between two variables, X and Y, there is no real functional relationship between X and Y. So, you cannot have a single Y for a given X. Right? It's just a dependency, but it's not a function. Like it's a relation or some kind of mapping, but not a function. So, indicated by the diagram on the right here, you have a bunch of data points, so those are the black dots. And so, they indicate some sort of dependency between X and Y. How do you capture this dependency given that you cannot run a function that computes Y from X? So, one way to do this is to learn or build a contrast function, energy function, that tells you whether a point in this XY space is near the training data or not. So, we think of it as some sort of landscape where the black dots are in the valley. In Switzerland, it would be a lake. And then, you know, you get level curves, right? As you move out, outside of those regions, the altitude goes up. The energy goes up. Now, if I give you a value for X, you can infer a bunch of values for Y that are compatible with X. They are values of Y that minimize the energy. So, it's the kind of inference I was talking about earlier, inference by optimization, not by forward propagation. But you can also possibly do it the other way around. If I give you a Y, you can infer X from Y. And you can give me multiple answers. So, in situations like video prediction, where there is basically an infinite number of possible answers, the proper way to train the system of this type is to think of it in terms of energy-based models. And by the way, probabilistic models are a special case, where your energy has a particular form and the way you train it has a particular loss function. So, it's a slightly more general framework, if you want, than probabilistic inference and learning.

防止能量模型崩溃 Preventing Collapse in Energy-Based Models

Host

好的,那么要训练一个基于能量的模型,你必须防止崩溃。

Okay, so what you to train an energy-based model, you have to prevent collapse.

Yann

我之前告诉你的崩溃问题会表现为能量函数处处平坦。你训练系统最小化一堆训练样本的能量,系统给你的能量函数处处为零。这就是学习恒等函数的自编码器所做的。一个忽略输入并为所有东西产生恒定表示且预测误差为零的喷射包。所以,这是一种崩溃。为了防止崩溃,你需要做两件事之一。一种是对比方法。你在数据区域之外生成点,然后推高能量。你提出一些成本函数,确保数据点的能量降低,而其他点的能量更高。有很多这样的方法。还有另一组方法,我逐渐更喜欢,即正则化方法,它们通过最小化可以具有低能量的空间体积来工作。所以,如果你推低某些区域的能量,其余部分必须上升,因为只有少量体积的低能量可用。那么,在实践中,你如何将其付诸实践?就是这两种方法之一。

The collapse problem I was telling you about before will be manifested by the energy function being flat everywhere. You train the system to minimize the energy for a bunch of training samples, and what the system gives you is an energy function that is zero everywhere. That's what an autoencoder that learns the identity function does to you. A jetpack that ignores the input and produces constant representation at zero prediction error for everything. So, it's a collapse. To prevent collapse, you need to do one of two things. One is contrastive methods. You generate points outside the region of data, and you push the energy up. You come up with some cost function that makes sure the energy of the data points come down, and the energy of other points is higher. And there's a whole bunch of them. And there is another set of methods which I've come to prefer, regularized methods, which work by minimizing the volume of space that can take low energy. So, if you push down the energy of certain regions, the rest has to go up because there is only a small volume of low energy to go around. So, in practice, how do you sort of reduce this to practice? One of those two methods.

信息最大化与表征学习 Information Maximization and Representation Learning

Yann

那么,让我们回到信息最大化这个想法。我想训练这个编码器来最大化某种信息度量。假设我通过一个编码器运行一批样本。我得到一个矩阵,其中每一行是一个样本的表示。每一列是表示中一个变量在所有样本中的值。有两种方法可以使这个矩阵信息丰富。一种是确保所有行都不同。另一种是确保所有列都不同。你需要确保列不同,因为如果所有列都相同,那就意味着表示中的每个变量都携带相同的信息。当然,这信息量不大。所以,你希望表示中的每个变量与其他变量最大程度地解耦,从而提供独立于其他变量的信息。所以,这可以称为维度对比方法的一个例子,这是一种正则化方法。然后,在底部,使所有行都不同的那种准则,是对比方法或样本对比方法。样本对比方法在某些应用中非常流行。许多感知流水线都是用称为 CLIP 的技术训练的,它基本上是一种对比方法,在图像和文本之间进行联合嵌入。但我更喜欢另一种。

So, let's go back to this idea of information maximization. I want to train this encoder to maximize some measure of information. Let's say I run a batch of samples through one of the encoders. I get a matrix where each row is the representation for one sample. Each column is the value of one variable in the representation for all samples. There's two ways to make that matrix informative. One way is to make sure all the rows are different. And the other way is to make sure all the columns are different. You want to make sure the columns are different because if all the columns are the same, that means every variable in the representation carries the same information. And of course, that's not very informative. So, you want each variable in the representation to be maximally disentangled from the other ones to give you independent information from the other variables. So, that would be an example of what we can call dimension contrastive methods, which is a form of regularized method. And then at the bottom, the type of criterion that makes the rows all different, those are contrastive methods or sample contrastive methods. Sample contrastive methods are very popular for certain applications. A lot of the perceptual pipelines in lens are trained with a technique called CLIP, which basically is a contrastive method that does joint embedding between images and text. But I prefer the other one.

抽象表征的必要性 The Need for Abstract Representations

Yann

所以,你需要找到输入的抽象表示才能进行预测,这个想法实际上非常自然。作为人类,我们一直在这样做。作为科学家和工程师,我们一直在这样做。动物也这样做。让我解释为什么。原则上,我可以解释或模拟此刻这个房间里发生的一切,用量子场论或粒子物理学的水平,对吧?可以模拟这个房间里每个粒子的轨迹。那将实际上模拟我们所有的大脑过程等等。所以原则上,运行模拟我可以判断你们中是否有人真的理解我在说什么。或者你是否在睡觉。或者你是否真的感到无聊。但当然,这完全不可行。

So this idea that you need to find an abstract representation of an input to be able to make prediction is actually very natural. We do this all the time as humans. We do this all the time as scientists and engineers. Animals do it too. Let me explain why. In principle, I could explain or simulate everything that takes place in this room at the moment at the level of quantum field theory or particle physics, right? Could simulate the trajectory of every particle in this room. And that would go down to actually simulating all of our brain processes and everything. So in principle, running the simulation I could figure out if any of you actually understands the word I'm saying or not. Or if you are sleeping right now. Or if you are actually bored. But of course that's completely impractical.

科学中的抽象 Abstractions in Science

Yann

在科学中,我们发明抽象概念来进行预测,这些抽象忽略了系统底层状态的许多细节。我们从量子场到粒子、原子、分子、蛋白质、细胞器、细胞、生物体、个体、社会、生态系统,每一层都是描述世界的一个特定抽象层次。通过忽略下层的大量细节,我们能够做出比下层更长期的预测。这就是为什么理解此刻这个房间里发生的事情,更合适的层面是心理学而非粒子物理学。当然,物理学家总是嘲笑别人说,你只要应用物理学就行了,甚至心理学在某种程度上也是应用物理学。但事实上,化学中有一些特定知识并非直接源自物理学。所以,这种抽象实际上包含了新的知识或信息或结构,这些在底层并不明显。

And then you know what we do in science is that we invent abstractions to allow us to make predictions and those abstractions ignore a lot of details about the underlying state of the system. So we invent those abstractions, you know, from quantum field to particles, atoms, molecules, proteins, organelles, cells, organisms, individuals, societies, ecosystem. Every level in this hierarchy is a particular level of abstraction with which we describe the world. Which allows us to make longer range predictions, if you want, than the levels below by ignoring a lot of details about the level below. Which is why the way to understand what goes on in this room at the moment is more at the level of psychology than at the level of particle physics, right? Now, of course, physicists always make fun of everyone saying like, you know, you just apply physics, right? Even psychology is applied physics to some extent. But in fact, you know, there is specific knowledge about chemistry that does not derive directly from physics, right? So this abstraction actually contains new knowledge or information or structure, if you want, that was not apparent at the level below.

世界模型与模拟器 World Models vs Simulators

Yann

所以 Jetpack 的想法正是建立在这个概念上:你需要找到一种抽象来进行预测。假设你想设计一架飞机,你需要设计飞机的气流。你进行计算流体动力学模拟,模拟机翼周围的气流。你通过速度和密度等参数,对机翼周围每个小立方体中的空气状态进行建模,然后求解纳维-斯托克斯偏微分方程,从而模拟气流。但实际上,这忽略了底层机制的大量细节。底层机制是空气分子相互碰撞并撞击飞机,但你永远不会在那个层面模拟流体。那太复杂了,而且由于细节过多,它会很快偏离现实。所以,你必须忽略细节才能做出准确的长期预测。我们在科学中一直这样做。因此,世界模型不应该是模拟器。它们应该在抽象空间中工作,不应该是数字孪生(那是个流行词),也绝对不应该是生成模型,正如我刚才解释的。它们也不应该是视频生成。很多人正在研究视频生成,并称之为世界模型,但它们不是世界模型,它们是视频生成系统。所以,我演讲的一个重要信息是:如果你想使用世界模型,就不要研究视频生成。这是不同的问题。如果你想制作可爱的视频,那就研究视频生成;但如果你想控制机器人或工业过程,或者理解世界,就不要研究生成。

So this idea of Jetpack really kind of constructs on this concept that you need to find an abstraction to be able to make predictions. Let's say you want to design an airplane. You need to design the airflow for the airplane. You do computational fluid dynamics, right? You simulate the flow of air around the wing. You model the state of the air in every little cube around the wing by basically the velocity and the density and things like that. And then you solve Navier-Stokes partial differential equations. And that simulates the flow of air. But in fact, it's ignoring a huge amount of details in the underlying mechanism. The underlying mechanism is molecules of air bumping into each other and bumping on the plane. But you never simulate fluids at that level. It's just too complicated. And also it would diverge from reality really quickly because it has too many details. So, you have to ignore details to be able to make accurate long-term predictions. And so, we do this in science all the time. And so, world models should not be simulators. Right? They should work in abstract space. They should not be digital twins, you know, that's a buzzword. They should definitely not be generative models, as I just explained. And they should not be video generation. So, a lot of people are working on video generation and they call this world models. They're not world models. They're video generation systems. So, one big message from my talk is that if you want to use world models, do not work on video generation. This is a different problem, okay? If you want to produce cute videos, work on video generation. But if you want to control robots or industrial processes or understand the world, do not work on generation.

复杂系统的世界模型 World Models for Complex Systems

Yann

你需要世界模型来控制那些无法通过写一组方程来建模动力学的复杂系统。如果你有一个类人机器人或任何类型的机器人,你可以直接写下动力学方程,然后模拟机器人的动力学,让类人机器人做空翻、功夫等等,这很简单。但一旦机器人开始与现实世界交互,情况就复杂得多,实际上更难简化为简单的方程。但想想一个复杂系统,比如涡轮喷气发动机、化工厂、病人,或者以复杂方式与现实世界交互的机器人。你无法将其简化为少数几个方程。你需要做的是学习整个系统(你控制的系统及其与环境的交互)的能量模型,以便进行预测并规划一系列动作以达到特定结果。这就是一个模型。这个概念很古老,可以追溯到 20 世纪 60 年代,是最优控制的根源。

You want world models to control complex systems where you cannot model the dynamics of the system by writing a bunch of equations. Okay, if you have a humanoid robot or any kind of robot, you can just write down the dynamical equations and then simulate the dynamics of the robot and you can get your humanoid robot to do somersaults and kung fu and whatever, right? That's simple. As soon as the robot starts to interact with the real world, that's a lot more complicated. And that is actually more difficult to reduce to simple equations. But then, you know, think about a complex system, like I said, a turbojet or a chemical plant or a patient or a robot, but a robot that interacts with the real world in complex ways. You cannot reduce this to a small number of equations. What you have to do is basically learn an energy model of the whole system, the system you control and its interaction with the environment, so that you can make predictions and you can plan a sequence of actions to arrive at a particular outcome. So that's one model. I mean, the concept is very old. It goes back to the 1960s. It's the root of optimal control.

Sigreg:各向同性高斯正则化 Sigreg: Isotropic Gaussian Regularization

Yann

好的,现在我要介绍一个我非常喜欢的具体技术,我认为我们将在未来几个月和几年内扩展它,以实现我之前提到的信息最大化。它叫做 Sigreg,即 sketch isotropic Gaussian regularization。技巧如下:你将一批样本通过编码器,得到表示空间中的一组点,表示空间的维度任意。我们试图让这些点的分布成为各向同性高斯分布,即所有维度具有相同的方差。为什么?因为各向同性高斯分布中所有变量都是独立的,所以每个变量都携带最大信息量。它也是给定方差下熵最大的分布,但我们并不关心这个。有趣的是它使变量相互独立。那么如何做到呢?当然,我们没有分布,只有空间中的一组点。这可能是高维空间,比如 2000 维,而我们可能只有几百或几千个点。如何确保它是高斯分布?技巧是:将每个点投影到单个方向上,得到边缘分布。当然,你仍然有离散点,没有连续密度。一个技巧是计算这些点的累积分布函数,它是一条阶梯函数,因为在一维上有离散点。然后你可以问:我的点的经验累积分布与高斯分布的累积分布之间的距离是多少?你可以这样做,因为你知道高斯分布的样子。对于每个点,在阶梯函数上,你可以判断它是在理想高斯分布的左边还是右边。这给出了一个梯度:我应该把这个点往这个方向还是那个方向移动?在这个投影上。这给出了一个梯度,对于批次中的每个训练样本。如果你通过梯度下降优化这个代价函数,它会使分布沿着这个投影的边缘分布成为高斯分布。

Okay, so now I come down to a particular technique that I'm very fond of, which I think we're going to expand over the next few months and years to do this information maximization that I was telling you about earlier. And it's called Sigreg. That means sketch isotropic Gaussian regularization. Okay, the trick here is the following. You run a batch of samples through your encoders, and what you get is a bunch of points in the vector space of dimension whatever the dimension of your representation space is. We're going to try to make the distribution of those points as isotropic Gaussian with the same variance in all dimensions. Why? Because an isotropic Gaussian is a distribution where all the variables are independent. Okay? So they're maximally informative individually. And it's also the distribution that has maximum entropy for a given variance, but we don't really care about that. What's interesting is that it makes the variables independent of each other. Okay, so how do we do this? Now, of course, we don't have the distribution. We just have a bunch of points in that space. And it may be a high-dimensional space like 2,000 dimensions. And we may have, you know, a few hundred or a few thousand points. Like how do we make sure this is a Gaussian? So here's the trick. The trick is you project the individual points along a single direction. And what you get is a marginal distribution. Okay? Now, of course, you still have discrete points. You don't have a continuous density. You have discrete points. Okay, so one trick you can do is compute the cumulative distribution that those points give you, right? So, it's a staircase, right? Because you have discrete points in one dimension. And then what you can do is you can ask, "What is the distance between the staircase, the cumulative empirical cumulative distribution of my points, and the cumulative distribution of, let's say, a Gaussian?" You can do that, because you know what the Gaussian looks like. And for every point, you can tell on the staircase, you can tell if it's to the left or to the right of the ideal Gaussian. And so, that gives you a gradient. Like, do I move the point this way or that way? In that projection. Okay? It gives you a gradient. Now, for every training sample in your batch. Okay. Now, if you make the distribution by gradient descent by optimizing this cost function, it's going to make the distribution Gaussian along the marginal of that distribution along this projection.

高斯化与世界模型 Gaussianization and World Models

Yann

但现在有一个定理:如果你沿着很多很多方向这么做,极限情况下,你的联合分布实际上是一个各向同性高斯分布。所以我们现在需要做很多投影。对于所有这些投影,计算那些梯度。移动点或者通过网络反向传播,改变权重,使得点移动,从而整体分布变得更像高斯分布。如果你把这个方法应用到像左上角那样的分布上,比如一个 X 形,这实际上是 1024 维中的二维。然后你进行梯度下降。你只是在这里移动点。你不训练神经网络。我提倡的技术在左边。你会得到某种类似高斯分布的东西。这在实际中确实有效。我们实际上把它应用到了训练动作条件的世界模型上,并用于规划,效果还不错。源代码是公开的。非常简单。你可以在一个 GPU 上训练它。我们需要做的基本上就是把这个技术规模化。还有其他一些事情需要做,但这是主要的。在简单情况下,你可以训练这个世界模型,并用它来规划简单的动作,比如在 push T 或模拟环境中的简单机器人场景。这需要规模化,但这是一项不错的工作。几天前我们发表了一篇理论论文,其中假设你的数据底层分布实际上是各向同性高斯分布,假设你从世界获得的观测是这些点的某种复杂非线性变换,比如螺旋变换,你训练一个带有 Sigreg 的神经网络,它会在表示空间中恢复原始高斯分布。这并不能证明它在所有情况下都有效,但它证明了如果你的原始解释变量是高斯分布,系统会恢复这些变量,最多差一个旋转。

But now, there's a theorem that says, if you do this along lots and lots of directions, in the limit, your joint distribution is actually an isotropic Gaussian. So, what we need to do now is do many projections. For all of those projections, compute those gradients. Move the points or backpropagate through the network, change the weights so that the points move so that the overall distribution gets more Gaussian. If you apply this to a distribution like the one on the top left here, like an X, these are actually two dimensions among 1,024. Then you do gradient descent. You just move the points here. You don't train the neural net. The technique I'm advocating for is on the left. You get something that's sort of Gaussian-ish. This really works in practice. We actually applied it to training world models that are action conditioned and we've used them for planning and it works decently. The source code is available. It's very simple. You can train it on one GPU. What we need to do with this technique is scale it up basically. There are a few other things we need to do, but that's the main one. In simple cases you can train this world model and use it to plan simple actions in a push T or simple robotic situation in simulated environments. That needs to be scaled up, but it's a good work. There is a theoretical paper we put out just a few days ago where if you make the hypothesis that the underlying distribution of your data is actually an isotropic Gaussian, if you assume that the observations you get from the world are some sort of complicated non-linear transformation of those points, like a spiral transformation, you train a neural net with Sigreg on it, it will recover the original Gaussian in the representation space. It's not a proof that it works in every case, but it's a proof that if your original explanatory variables are Gaussian, the system will recover those variables up to a rotation.

自监督学习:C-Reg 与蒸馏方法 Self-Supervised Learning: C-Reg and Distillation Methods

Yann

我们可以用这些技术进行自监督学习,训练图像识别系统。还有另一组技术值得一提,因为它们效果非常好,而且到目前为止已经被规模化。C-reg 概念上是我最喜欢的方法,但它非常新,我们还没有规模化它。而那些基于蒸馏的方法,我们进行了规模化,在图像和视频上都取得了非常好的结果,比如 I-jeppa 和 V-jeppa 技术。那么这些蒸馏方法的基本思想是什么?你仍然有两个编码器。这是一个 jeppa 架构。你取一个输入,对其进行变换、破坏或掩码等操作,然后训练系统在表示空间中进行预测,但不通过右侧编码器传播梯度。这两个编码器架构相同,并且某种程度上共享权重,但有趣的是,右侧编码器使用左侧编码器权重的指数移动平均。左侧编码器接收梯度并持续更新。右侧编码器更新得更慢,并且共享权重。这源于一些直观的想法。Google DeepMind 的一些人使用类似技术来稳定强化学习中的方差,他们意识到可以将其应用于图像的自监督学习。他们称之为 BYOL,即 bootstrap your own latent。来自 Meta 的许多方法,特别是 SimSiam、MoCo 等,都使用了这种指数移动平均的思想。我在这里展示的一个特定方法叫 I-jeppa,它产生了非常好的结果。我们能够用 I-jeppa 将其结果与一种生成式方法 MAE(掩码自编码器)进行比较。它不仅更好,而且训练速度也快得多。

We can use those techniques in the context of self-supervised learning to train an image recognition system. There is another set of techniques which I should mention because they work really well and they are the ones that have been scaled up so far. C-reg is conceptually my favorite method but it's very recent and we haven't scaled it up. Whereas those other methods that are based on distillation we scaled them up and we got really good results both for images and video with techniques like I-jeppa and V-jeppa. So what's the basic idea of those distillation methods? You still have those two encoders. This is a jeppa architecture. You take an input, you transform it or corrupt it or mask it or something, and then you train the system to predict in representation space but you don't propagate gradient through the encoder on the right. Those are two encoders with identical architectures and they kind of share the weights but the funny thing is that the encoder on the right uses an exponential moving average over time of the weights of the encoder on the left. The encoder on the left gets gradient and gets updated all the time. The encoder on the right gets updated slower essentially and shares the weights. This is derived from some intuitive ideas. Some people at Google DeepMind who were using techniques like this to stabilize the variance in reinforcement learning realized you could apply this to self-supervised learning from images. They call this BYOL, bootstrap your own latent. There is a whole bunch of methods coming from Meta in particular SimSiam, MoCo etc. that use this exponential moving average idea. A particular method called I-jeppa which I show here produced really really good results. What we were able to do with I-jeppa is compare the results of I-jeppa with a generative approach called MAE, masked autoencoder. It is not only better but it is much faster to train.

DINO 及其应用 DINO and Its Applications

Yann

另一种技术叫 DINO。我相信你们很多人都听说过。我知道你们有些人用过它,因为机器人演示中有项目实际使用了 DINO。这是我以前在 Meta 巴黎的同事做的,完全是自监督的。它是一个联合嵌入架构,使用蒸馏,但带有各种技巧,我就不解释了。背后有很多工程工作。这些系统目前基本上能产生最好的通用图像表示。如果你有任何类型的视觉任务要做,那可能是最好的图像编码器。但我们所做的,其中一项就是使用 DINO 作为编码器,然后训练世界模型并进行规划。让我给你们看一个可爱的视频。这里有一个模拟环境的初始状态,具有相当复杂的动力学。顶部有目标,底部你看到的是规划器使用这个训练好的世界模型,在不到 25 步内将世界配置到尽可能接近原始状态的动作序列。这已经应用于许多不同的场景,比如双摆和 push T 等等。效果非常好。

Another technique is called DINO. Many of you I'm sure have heard of it. I know some of you have used it because there were projects in the robot demos that actually used DINO. This is done by some of my former colleagues at Meta in Paris and it's completely self-supervised. It's a joint embedding architecture. It's using distillation, but with various tricks, which I'm not going to explain. There's a lot of engineering that goes behind it. Those systems basically at this time produce the best generic representations of images. If you have any type of vision task that you want to do, that's probably the best encoder for images. But what we've done is among other things use DINO as an encoder and then train a world model and do planning. Let me show you just a cute video on this. You have an initial state here of a kind of simulated environment that has pretty complex dynamics. And you have goals at the top and at the bottom what you see is the sequence of actions of a planner that uses this trained world model to get the world to a configuration as close as possible to the original one in less than 25 steps. This has been applied to a number of different scenarios like double pendulum and push T and whatever. It works really well.

V-Jeppa 与常识 V-Jeppa and Common Sense

Yann

我们最近把它应用到了视频上。所以,你取一段视频,掩码掉一大块,然后训练 JAPA 再次生成好的表示,这样你就可以从部分掩码视频的表示预测完整视频的表示。一旦系统训练好,你就用编码器从视频中提取特征,并在其顶部训练一个头部来完成某些任务,效果非常好。它在许多传统的视觉任务上达到了最先进水平,特别是来自视频的任务,比如动作识别、动作预测等等。我想提一件有趣的事,而不是用结果表格让你们厌烦,那就是这些系统,特别是 V-Jeppa,已经学到了一定程度的常识。我们可以用 V-Jeppa 做的一件事是,因为我们训练它预测视频中接下来会发生什么。我们可以尝试预测,我们可以测量它的内部预测误差。我们可以给它看一段视频,并监控每个时间步的内部预测误差。系统会取一个 16 帧的窗口。所以我们就在视频上滑动这些帧,并测量接下来 16 帧的预测误差。酷的是,如果你给它看一段视频,其中发生了不可能的事情,一些不物理的事情,预测误差会飙升到最高。

We more recently applied it to video. So, you take a video, you mask a big chunk of it, and you train the JAPA to again produce good representations so that you can predict the representation of full video from the representation of partially masked one. Once the system is trained, you use the encoder as a way to extract features from the video and you train a head on top of it to accomplish some task and it works really well. It's state of the art for a lot of traditional vision tasks particularly from video like action recognition, action prediction and stuff like that. The one interesting thing that I want to mention instead of boring you with the table of results is that those systems, V-Jeppa in particular, has learned some level of common sense. One thing we can do with V-Jeppa because we train it to predict what's going to happen next in the video. We can try to predict, we can measure its internal prediction error. We can show it a video and monitor the internal prediction error at every time step. The system takes a window of 16 frames. So we just slide those frames on the video and measure the prediction error for the next 16 frames. The cool thing is that if you show it a video where something impossible occurs, something unphysical, the prediction error shoots to the roof.

自监督学习与常识 Self-supervised learning and common sense

Yann

就像早期幻灯片里那个小女孩看着汽车没有掉下来的场景一样。同样,如果你有一段球被扔出去然后消失的视频,预测误差会飙升。这很有趣,因为至少在我看来,这是第一次看到一个完全自监督的系统获得了一定程度的常识。它能告诉你什么是可能的,什么是不可能的。

So it's like the little girl in one of the early slides who looks at the scene of the car not falling. Same thing. You have a video of a ball being thrown and the ball disappears. Prediction error will shoot through the roof. So that's interesting because it's the first time, at least from my point of view, that I've seen a completely self-supervised system acquire some level of common sense. It'll tell you what's possible, what's not possible.

Yann

我跳过这个。这很可爱,但它只是说 V-JEPA 可以用于规划,这些新版本在规划等方面做得更好。但这里有一个有趣的事情。记得我告诉过你,婴儿学习世界是三维的方式,是因为这是解释当你移动头部时视野变化的最佳方式。所以我们采用了某个版本的 V-JEPA 或 V-JEPA 2.1 学到的表征,然后在上面训练了一个头部来从单张图像预测深度。它做得非常好,结果甚至比 V3 还好。这表明,这个系统仅仅通过训练在表征层面预测或填补视频中的空白,就基本上理解了世界是三维的。我是说,带引号的理解。它理解了物体的概念。如果你把这个表征作为输入给一个分割系统,它也能工作得不错。还有其他各种应用。

Let me skip this. It's cute but it just says V-JEPA can be used for planning, and these new versions do a better job at planning and everything. But here is an interesting thing. Remember I told you the way babies learn that the world is three-dimensional is because it's the best way to explain how your view changes when you move your head. Okay, so we took the representation learned by some version of V-JEPA or V-JEPA 2.1, and then we trained a head on top of it to predict depth from a single image. And it does a really good job. It produces really good results. In fact, better than V3. And what that shows is that this system, by just being trained to predict or fill in the blanks in videos at a representation level, basically understands that the world is three-dimensional. I mean, understands with double quotes. Understands the notion of object. If you use the representation as input to a segmentation system, it works decently well. And for various other things.

放弃生成模型与 LLM Abandon generative models and LLMs

Yann

好了,我来总结一下。这很有趣,对吧?所以,放弃生成模型。我是说,如果你研究 LLM,当然要放弃。但你不应该研究 LLM。至少如果你在学术界,你绝对不应该研究 LLM。你没有什么可以贡献的。所以,放弃生成模型,转而采用联合嵌入架构。如果你对智能感兴趣,放弃概率模型,转而采用那些基于能量的模型。我没有时间真正解释为什么。我提出了一个论点,支持那些正则化方法或通过变量而非样本进行信息最大化的方法。所以,放弃对比方法,尽管它们有很多实际应用。我一直在说放弃强化学习。我不是真的说放弃,我是说最小化它的使用,因为它在样本效率方面极其低下。我知道这里有研究这个的人,但 RL 就像你绝望时别无选择才做的事情。你必须通过观察进行大部分学习,学习世界模型等等。一旦你有了好的表征,你可以在上面使用 RL,因为你已经有了好的表征。你不需要太多样本。有时你无法避免。当然,如果你对在 AI 上取得真正进展感兴趣,在某种接地气的、面向真实世界的 AI,如果你想要物理 AI,不要研究 LLM。也不要研究生成模型。

Okay, let me conclude. So, it's funny, huh? So, abandon generative models. I mean, if you work on LLM, of course. But you should not work on LLM. At least if you're in academia, you should absolutely not work on LLM. There is nothing you can bring to the table. So, abandon generative model in favor of joint embedding architectures. If you are interested in intelligence, abandon probabilistic models in favor of those energy-based models. I didn't have time to really explain why. I made an argument in favor of those regularized methods or information maximization through variables instead of samples. So, abandon contrastive methods, which again have a lot of practical applications. I've been saying forever to abandon reinforcement learning. I don't really mean abandon. I mean, minimize its use because it's so horribly inefficient in terms of sample efficiency. And I know there are people here who work on this, but RL is like what you do when you're desperate and there is nothing else you can do. You have to do most of the learning by observation, learning world models, etc. And once you have good representations, you can use RL on top of it because you already have the good representations. You won't require too many samples. Sometimes you can't avoid it. And certainly, if you're interested in making real progress in AI, in sort of grounded AI for the real world, if you want physical AI, don't work on LLMs. Don't work on generative models, either.

Yann

所以,你可能猜到了,这让我在硅谷不太受欢迎。是的。所以,我离开了 Meta,很多人可能知道,去年年底我成立了一家新公司叫 Ami Labs。Ami Labs 的宗旨是面向真实世界的 AI,比如物理 AI。机器人是一个用例,但不仅如此。还有工业过程的控制。任何高维、连续、嘈杂的问题,LLM 完全无能为力。这就是我们正在解决的问题。就这些。非常感谢。

So, as you can probably guess, this does not make me very popular in Silicon Valley. Yes. And so, I left Meta, as many of you probably know, at the end of last year and formed a new company called Ami Labs. And the purpose of Ami Labs is sort of AI for the real world, like physical AI. Robotics is a use case, but it's not just that. It's control of industrial processes. Like, anything that is high-dimensional, continuous, and noisy, for which LLMs are completely helpless. This is the kind of problems we're working on. And that's it. Thank you very much.

问答:表征空间中的护栏 Q&A: Guardrails in representation space

Host

好的,我知道有很多问题。也许我们问一两个,然后就得结束了。所以,请快速提问和回答。

Okay, so I know there are many questions. Maybe we'll take one or two, but then we have to wrap up. So, please quick questions and quick answers.

Audience

谢谢你的演讲。我想问一下你在早期幻灯片中提到的护栏,你也谈到了 MPC。工程师喜欢 MPC,因为他们可以加入约束,在状态空间(如 3D 空间)中描述它们。但根据我的理解,在你的系统中,一切都在表征空间中工作。我如何将像“不要撞墙”这样的约束放入这个表征空间?你设想系统自己学习约束,还是工程师真的可以加入它们?

Thanks for the talk. I wanted to ask about the guardrails that you mentioned on one of the earlier slides where you also talked about MPC. Engineers love MPC because they can put in their constraints, describe them in state space, like 3D space. But from what I understand, in your system, everything works in representation space. How do I even get a constraint like 'don't bump into the wall' into this representation space? Do you envision the system learning the constraints by itself, or can engineers really put them in?

Yann

不,你需要在表征之上学习一个非常小的头部,将其映射到你感兴趣的约束。所以那部分必须训练,但你可以用非常少的样本训练它,因为它基本上是一个很小的投影。

No, you would have to learn a very small head on top of your representation that maps that to the constraint you're interested in. So that part has to be trained, but you can train it with a very small number of samples because it's a tiny projection basically.

Audience

但你需要为每种可能加入的约束使用不同的编码器,不是吗?

But you need a different encoder for each kind of constraint that you might want to put in, no?

Yann

嗯,你需要为每个约束使用不同的投影器,对吧。所以如果你的任务是开门,我不是在说约束,我是在说任务目标。你需要一个成本函数来告诉你门是否打开了,对吧?所以那可能需要在训练完成任务时训练,但基本上只需要两个样本。好了。

Well, you need a different projector for each constraint, right. So if your task is to open a door, I'm not talking about constraint, I'm talking about a task objective. You need some cost function to tell you whether the door is open or not, right? And so that might have to be trained when you're trained to accomplish the task, but basically that requires two samples. All right.

Host

好的,我想我们得在这里结束了。非常感谢,Yann。

Okay, I think we'll have to leave it here. Thank you, Yann, very much.

Yann

好的。谢谢。

All right. Thank you.

互动版:逐字朗读 + 针对本期提问 →