智能的物理学:从沙堆到潜在空间

Physics of Intelligence: From Sand to Latent Space

马蒂厄·维亚尔 Matthieu Wyart · ML Street Talk · 2026-08-10 · 约 79 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

一位物理学家探讨统计物理原理如何解释大脑的高效性,以及为什么在潜在空间预测优于在词元空间预测。

A physicist explores how principles from statistical physics explain the brain's efficiency and why predicting in latent space beats token space.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

引言与研究动机 Introduction and Research Motivation

Matthieu

我是 Matthew Wyatt,美国约翰霍普金斯大学和瑞士洛桑联邦理工学院的正教授。你知道那些能生成我们从未见过的新图像、或说出从未听过的新句子的机器吧。我们的大脑似乎能用比机器少 10 万倍的词汇量来学习语言。为什么会这样?我们是不是做错了什么?我非常感兴趣的是,我们应该在非常底层的词元空间进行预测,还是更应该训练机器去预测抽象概念。所以这些年来我们一直在做的,就是试图建立一个基于物理学的框架,真正以统一的方式回答这些不同的问题。

I am Matthew Wyatt. I'm a full professor at Johns Hopkins University in the US and at EPFL in Switzerland. You know those machines that can build new images that we've never seen before or say new sentences that were never heard before. Our brain seems to learn languages with 100,000 times less words than machines. Why is it so? Are we doing the wrong thing? And I'm very interested in, you know, should we predict in token space at a very low level or more should we train machines to predict abstractions. And so what we've been doing over the years is trying to build a framework based on physics that's really tried to answer those different questions in a unified manner.

Matthieu

乔姆斯基提出了“刺激贫乏”论证,认为从例子中学会创造性实际上是不可能的。但如果你有一个深层架构,就存在一种巨大的隐式偏向,去构建那些粗粒度的变量。所以如果你想想大语言模型或扩散模型,它们孕育概念的方式,它们仅从统计中涌现。那些抽象,它们涌现出来,它们就在数据中,它们涌现出来,当你把那些预测周围相似上下文的配置归为一组时,这些概念就涌现了。

Chomsky gave this poverty of stimulus argument, you know, arguing that it was actually impossible to learn to become creative from example, but if you have a deep architecture there's a huge implicit bias to build those coarse-grained variables. And so if you think about LLMs or diffusion models, the way they breed concepts, they emerge from statistics alone. Those abstractions, they emerge, they are there in the data, they emerge, and those concepts emerge if you group together configurations that predict similar context around them.

Host

这非常相关,因为你有一篇论文,基本上是说我们应该在潜在空间而不是词元空间进行预测。

And this is very pertinent because you've got a paper out basically saying that we should predict in the latent space, not the token space.

Matthieu

所以再次强调,在这些模型中,我们发现那些内省的、从自身潜在表征中学习的算法,在样本复杂度方面要强大得多。它们最终会学到同样的抽象,但速度要快得多。就像如果你从不犯错,也许这标志着你在科学上有点固守常规。而我们有些人想去探索丛林,在丛林里你可能会犯错。我的意思是,是的。

So again, in those models, what we found is that those algorithms that are introspective, that learn from their own latent, are much more powerful in terms of sample complexity. They will eventually learn the same abstraction but much faster. Like if you never do mistakes, maybe it's a sign that you're staying a bit on the beaten path in science. And some of us want to explore the jungle, and in the jungle you can be wrong. I mean, yes.

赞助商插播 Sponsor Break

Host

快速插播。智能体每天都在变得更聪明,但即使是最聪明的智能体,如果没有正确的上下文和正确的工具,也会卡住。这就是 Notion 发挥作用的地方。随着最近自定义智能体的推出,Notion 成为了团队和智能体并肩工作的协作式 AI 工作空间。现在,他们的新开发平台正在把这个工作空间变成开发者可以构建的基础设施。这正是我运营 MLST 的方式。整个节目都放在 Notion 里。我的嘉宾、发布日历、商业方面,一切都在那里。但变化在于它现在是智能体式的。我只需和我的智能体对话。它可以是 Claude 或任何智能体式框架。然后它通过 MCP 或 CLI 与 Notion 对话,事情就完成了。然后我可以在手机上访问它。这绝对是一个颠覆性的改变。

Quick pause. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That is where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now their new development platform is turning that workspace into infrastructure developers can build on. Now, this is exactly how I run MLST. The whole show lives in Notion. My guests, the publishing calendar, the commercial side, everything is in there. But what's changed is that it's now agentic. I just talk to my agent. It can be Claude or any agentic harness. And then it then talks to Notion via the MCP or the CLI and it's just done. And then I can access it on my phone. It's an absolute game changer.

从物理到机器学习 From Physics to Machine Learning

Matthieu

所以,是的,我实际上是个物理学家。我真的很喜欢学物理,因为你必须处理所有可能尺度上的自然现象,而我专注于物理学中的一个特定领域,叫做统计物理。统计物理本质上是一个领域,你试图理解有多少实体、粒子相互作用产生集体现象。所以一个经典的例子是,你拿水,冷却系统,在某个时刻砰的一声,它冻结了,完全改变了它的组织。所以我最初开始研究股票市场,那里有相互作用的智能体影响价格的演变,这是一种非常有趣的随机游走。然后我去研究复杂系统。复杂系统是具有粗糙能量景观的物理系统。这意味着,如果你,有点像你刚才飞越阿尔卑斯山。如果你在这些山上扔一个球,它可能会停在许多不同的点,所以能量景观有许多亚稳态。这些系统真的很有趣,是有记忆的物理系统。我研究过几个这样的系统,比如沙子。沙子的美妙之处在于它是一个复杂系统。如果你准备 1 万堆沙子,每一堆都不同,但它又有有趣的相变。所以如你所知,如果你倾斜一层沙子,在某个时刻它会流动。这意味着能量景观是粗糙的,你处于亚稳态,但你倾斜了这个能量景观,你经历了一个相变,然后整个系统流动,尽管它是非常密集的粒子,却设法避开了彼此。所以我一直对从几何上理解这些问题非常感兴趣。

So, yeah, I'm a physicist actually. I really liked learning physics because you have to deal with nature at all possible scales, and I focused on one specific field in physics which is called statistical physics. Statistical physics is essentially the field where you try to understand how many entities, particles, interact together to do collective phenomena. So a classical example is you take water, you cool down the system, and at some point boom, it freezes, completely changing its organization. So I started to work on that initially on the stock market where you have interacting agents that influence the evolution of the price, which is a very interesting sort of random walk. And then I went to study complex systems. So complex systems are physical systems with a rough energy landscape. It means that if you are, it's a bit like if you're flying above the Alps like you just did. If you are throwing a ball in those mountains, it could stop at many different points, so the energy landscape has many metastable states. And those systems are really intriguing, the physical systems with memory. I worked on several of those, for example sand. So what's beautiful about sand is that it is a complex system. If you prepare 10,000 piles of sand, each of them is different, but it has intriguing again phase transitions. So as you know, if you tilt a layer of sand, at some point it's going to flow. It means that the energy landscape was rough and you are in a metastable state, but you tilted this energy landscape, you had a phase transition, and then the entire system flows, although it's very dense particles managed to avoid each other. So I've been very interested in understanding geometrically those questions.

Matthieu

但后来,大约九年前,我是一个围棋棋手,一个水平很差的围棋棋手,但我喜欢下棋,我被 AlphaGo 等迷住了。所以我开始思考机器学习,我开始把它看作一个复杂系统。确实如此,因为当你训练一台机器时,你构建了一个函数,如果你很好地拟合数据,这个函数就低,它被称为损失函数或成本函数。所以我们非常好奇这个景观的几何形状。我们发现的是,这个景观实际上具有与沙子完全相同的相变。这意味着当你参数不足时,当你没有足够的参数时,你有一个粗糙的景观,有许多亚稳态。如果你训练你的机器,并且你多次训练它,它会最终到达不同的位置,实际上卡在那里。但如果你有足够的参数,那么突然系统可以流动。景观有许多平坦的山谷,基本上能量为零。所以确实有密切的类比。我们在大约九年前发现了这一点,同时其他人也发现了非常相似的,我的意思是相同的现象,并称之为“双重下降”。所以现在这个名字已经固定下来,但这个双重下降,这个双重下降的峰值,对物理学家来说实际上是一个堵塞转变。

But then, like nine years ago, I'm a Go player, a poor Go player, but I enjoy playing, and I was mesmerized by AlphaGo and so on. So I started to think about machine learning, and I started to think of it as a complex system. And it is, because when you train a machine, you build a function that is low if you fit well your data, it's called the loss function or cost function. So we are very intrigued by what is the geometry of this landscape. And what we discovered is actually that this landscape has exactly the same phase transition as sand. It means that when you are underparameterized, when you don't have enough parameters, you have a rough landscape with many metastable states. And if you train your machine and you train it many times, it will end up in different positions where it's actually stuck. But if you have enough parameters, then suddenly the system can flow. The landscape has many flat valleys which have essentially zero energy. So there is really a close analogy. We discovered that like nine years ago, at the same time others found a very similar, I mean the same phenomenon, and called it double descent. So now this name has stuck, but this double descent, this peak of the double descent, is really for physicists a jamming transition.

Matthieu

所以这把我带到了机器学习。也许就此结束,我的意思是,在过去的四年里,我们一直对另一个我认为更有趣的景观非常感兴趣。那是数据的景观。所以如果你想到一张图像,让我们称 X 为一张图像,它是一个向量。你可以问这些图像的密度是多少,即 X 的密度。这个问题关系到世界的结构是什么,我们认为这是真正理解机器如何工作的关键。

So that brought me to machine learning. And just maybe to finish with that, I mean, in the last four years we've been very much interested in another landscape that I think is even more interesting. It's a landscape of data. So if you think about an image, let's call X an image, it's a vector. You could ask what is the density of those images, rho of X. And this question relates to what is the structure of the world, and we think it's key to actually understand how a machine works.

Host

这么说有意义吗?因为显然你是一位物理学家,你正在将这种分析视角应用于大语言模型,而天真地看,我会说,嗯,它感觉不像物质基底,感觉它没有与现实世界中的事物相同的动力学类型。但确实,当我们观察大语言模型的训练动态,当我们观察它们学习到的表征类型时,我们可以采用物理学的视角,说存在粗粒化、相变等等。

Does it make sense to talk about, because obviously you're a physicist and you're applying this lens of analysis to large language models, and like naively I'm looking at this and saying, well, it doesn't feel like a material substrate, it doesn't feel like it has the same type of dynamics as things do in the real world. But indeed, when we look at the training dynamics of LLMs and when we look at the types of representations they learn, we could adopt a physics lens and say there are coarse grainings and there are phase changes and whatnot.

沙丘与损失景观的类比 Analogy between sand and loss landscape

Host

你觉得做这个类比是合理的吗?

Do you think it's coherent to make that analogy?

Matthieu

是的,我认为这是科学中一个令人着迷的事实:某个概念可以应用于如此多样的现象。我认为科学本质上就是建立在这种类比之上的。比如,惠更斯是最早提出光是一种波的人之一。他是怎么提出的?他注意到海洋上的波浪可以相互交叉而不相互作用,他注意到光也是如此,于是他做了这个类比。我觉得这甚至很难谈论,因为我认为它是如此根本,我们总是在用类比来构建理解。所以,对于我给你的关于沙子和机器损失景观的具体例子,我认为在这种情况下类比非常直接,因为在这两种情况下,你所拥有的本质上是自由度。一种情况下是沙粒,另一种情况下是大模型的参数,而在这两种情况下,系统都在试图满足约束。对于沙子,本质上粒子只是试图避免彼此接触,但对于参数,它们集体试图做的是拟合数据。所以数据越多,约束就越多,最终,我们从物理学中论证的普适性在那里适用:如果你有一个约束可满足性问题,并且你有连续的自由度可以连续变化,那么砰,你就有了一个普适类。所以在这个意义上,是的,这类问题有某种普遍性。但这是一个非常具体的例子;我不想说所有事情总是相同的,但沙子的堵塞问题和机器的损失景观问题非常相似。

Yes, I think it's a sort of mesmerizing fact of science that some concept can be applied in so diverse phenomena. I think science is essentially built on those kinds of analogies. I mean if you think about Huygens, who was one of the first to propose that light was a wave. How did he propose that? He noticed that waves on the ocean could cross each other without interacting, and he noticed it was the same for light, and so he made this analogy. I think it's even hard for me to talk about because I think it's so fundamental that we are always building our understanding in terms of analogies. So for the specific example I gave you about the sand and the loss landscape of machines, I think the analogy is very direct in this case because in both cases what you have are essentially degrees of freedom. In one case those are the particles of sand, in the other case they are the parameters of your large model, and in both cases the systems are trying to satisfy constraints. For sand, essentially the particles are just trying to avoid each other, but for the parameters, what they are trying to collectively do is to fit data. So the more data you have, the more constraints you have, and at the end, the universality that we've argued from physics applies there: if you have a problem of satisfiability of constraints and you have continuous degrees of freedom that can change continuously, then boom, you have a universality class. So in this sense, yes, there's something universal about those kinds of problems. But that's a very specific example; I don't want to say that everything is always the same, but this specific problem of jamming of sand and the one of loss landscape of machines is very much the same.

Host

这是一个非常诱人的想法,因为我认为约束无处不在,在进化中我们有自然趋同的模式,反复出现的模式,比如蟹化。我想对此唯一的批评是,感觉在神经网络或计算机中,我们没有同样的物理约束。你知道,我们没有两个物体不能同时接触这样的物理定律等等。所以约束是通过数据中的统计模式存在的,但它们仍然对训练过程施加压力。这些仍然是有效的约束吗?

It's such a tantalizing idea because I think it is constraints all the way down, and in evolution we have naturally convergent patterns, reoccurring patterns like carcinization. And I guess the only critique to this is it feels like in neural networks or just in computers, we don't have the same kinds of physical constraints. You know, we don't have two objects that can't touch each other at the same time and the laws of physics and so on. So the constraints are there by dint of statistical patterns in the data, but they still apply pressure on the training process. Are those still valid constraints?

Matthieu

是的。这里我并不是在谈论计算机本身的任何约束。我真的是以某种抽象的方式思考算法。算法在做的是一种梯度下降,在两种情况下都沿着能量景观流动,这就是类比所在。类比与物质方面无关。诚然,一种情况下是物质,另一种情况下是算法,但如果你正确思考,在某种程度上它们是相同的。

Yes. So here I was really not talking about any sort of constraint of the computer itself. I was really thinking in some abstract way about the algorithm. What the algorithm is doing is some sort of gradient descent flowing down an energy landscape in both cases, and that's where the analogy is. The analogy is not related to the material aspect. It's true that in one case it's a material, in the other case it's an algorithm, but it is at some level the same if you think about it correctly.

Host

有趣的是,你把约束看作是算法本身,而不是能量景观。我明白了。我觉得你倾向于存在某种通用学习算法。但我的直觉是,数据和世界作为约束更有意义。这说得通吗?

It's so interesting that you're thinking of the constraints as being the algorithm rather than the energy landscape itself. And I get it. I think you're leaning towards there being some kind of a universal learning algorithm. But the way I intuit it is it's almost like the data and the world are more meaningful as constraints. Is that legible?

Matthieu

好的。我想我们会讨论创造力,然后我们也会大量讨论约束,但在我思考中,那将是另一个空间。所以我开始告诉你我们在讨论损失景观,这里的约束只是提供数据。稍后,我非常感兴趣讨论的是,如果你想想世界本身,句子,数据本身,忘掉实际学习它的算法,数据本身是非常受约束的。所有可能的句子在句法上并不都是有效的。所以我认为在这两种情况下思考约束都非常有用,但我认为它们是截然不同的约束。

Okay. So I think we will be talking about creativity, and then we will also be discussing a lot about constraints, but there will be, in my way of thinking, in another space. So I started to tell you we were discussing loss landscape, and the constraint here was just to feed data. Later on, something I'm really interested to discuss is, if you think about the world itself, the sentences, the data itself, forget about the algorithm that's actually learning it, the data itself is very constrained. All possible sentences are not valid in terms of syntax. So I think thinking of constraints is very useful in both cases, but I think of them as very different kinds of constraints.

Host

你还认为自己首先是一个物理学家吗?因为现在我们谈论的是学习的物理学,我们谈论的是机器学习。你作为物理学家的直觉是如何融入这个世界的?

And do you still think of yourself as a physicist first? I mean because now we're talking about the physics of learning, we're talking about machine learning. How does your instinct as a physicist come across into this world?

Matthieu

这是一个非常有趣的问题。我们在物理系里经常要辩论这个问题,因为我们需要招聘人员,什么是物理学,我们总是问这个问题。确实,我最初开始思考机器学习时,它与复杂系统和无序固体等密切相关。现在我正在重新思考,现在我正在与语言学家和神经科学家等组织会议。但我仍然认为内心深处它是物理学。我认为我们需要把物理学带到这个领域,做语言学的物理学。这是一个很长的讨论,但也许物理学非常特别的地方,其他领域的自然科学家也这样做,但我们真正试图这样做。首先是建立理论与实验之间的对话。例如,当有一项新技术时,它提出了大量新颖的问题,我们可以开始思考它,建立理论,然后理论我们建立简单模型,然后理论必须对正在发生的事情有预测性。所以首先,我们非常擅长建立这种对话,建立经验科学。第二个方面是建模。世界超级复杂。如果你试图制作一个一英里等于一英里的地图,它永远不会帮助你。所以你需要构建世界的漫画。这是物理学家所做的艺术。用爱因斯坦的话说,模型应该尽可能简单,但不能更简单。所以这意味着存在张力。要在一个良好的复杂性水平上描述问题真的很难,而且这也取决于你具体问什么问题。所以我认为物理学家也擅长开发这类模型。举个例子,我告诉过你相变。一个世纪前,皮埃尔·居里在思考磁性,以及当你改变温度时,突然这些材料变得有磁性,它们粘在你的冰箱上,但在更高的温度下它们不会。那么发生了什么?如果你在微观层面思考,这是极其复杂的量子力学。

That's a very interesting question. We have to debate it all the time in physics departments because we need to hire people and what is physics, and we always ask this question. And it's true, I started initially when I started to think about machine learning, it was closely related to complex systems and disordered solids and things like that. And now I'm rethinking, now I'm organizing conferences with linguists and neuroscientists and so on. But I still think deep down it's physics. I think we need to bring physics to this field and do the physics of linguistics. And that's a long discussion, but maybe what is very special about physics, other fields natural scientists do it too, but we're really trying to do that. First is to build a dialogue between theory and experiments. So for example, when there is a new technology, it's asking a huge number of novel questions, and we can start thinking about it, making theory, and then the theories we make simple models, and then the theories have to be predictive on what's going on. So first of all, we are very good at building this dialogue, building some empirical science. The second aspect is modeling. The world is super complicated. If you try to make a map where one mile is one mile, it will never help you. So you need to build a caricature of the world. And it's an art that physicists have done. To paraphrase Einstein, a model should be as simple as possible, but not simpler. So it means there's tension. It's really difficult to describe the problem at a good level of complexity, and it depends also specifically on what question you're asking. So I think physicists are also good at developing those sorts of models. Just to give you an example, I told you about phase transition. One century ago, Pierre Curie was thinking about magnetism and the fact that when you change the temperature, suddenly those materials become magnetic and they stick to your fridge, but at higher temperature they don't. So what's going on? If you think of it at the microscopic level, it's awfully complicated quantum mechanics.

物理与简单模型 Physics and Simple Models

Matthieu

但你知道,那个被广泛接受、在相变领域取得巨大进展、并催生了诺贝尔奖和数学等领域成果的描述,是一个非常简单的模型,即伊辛模型。在这个模型中,晶格上的箭头与相邻箭头相互作用,试图对齐。所以,对现象的一个非常粗略的描述,对于思考这个问题来说,是一个好的描述。所以我认为这也是我们可以为这些问题带来的东西。最后,我想你提到的类比,比如重整化,我们物理学家试图在所有尺度上思考问题,为了快速做到这一点,他们不得不在不同领域之间建立类比。所以,这就是我们能带来的。

But you know the description that stuck and that made huge headways in terms of phase transition and led to Nobel Prizes, in fields in math and so on, is a very simple model, the Ising model, where you have essentially arrows on a lattice that are interacting with their neighbors to try to align. So a very crude description of the phenomenon was a good one to essentially think about this problem. So I think that's also what we can try to bring to those questions. And lastly, I think what you were describing, analogies like renormalization, I mean we are physicists that try to think about problems at all scales and to do it fast they had to build analogies between different fields. So that's what we can bring.

Host

我是说,当我们和乔姆斯基交谈时,他对物理学这个事业相当不屑。他谈到,你知道,运动这个最初的难题,他说牛顿运用了机器,也就是机械宇宙观,但他保留了幽灵。所以,你知道,我们仍然不知道心智和意识是如何运作的,诸如此类。但他更是指出,很多物理学是理想化。然后有一个有趣的问题:我们的理论是真的用来理解宇宙如何运作,还是更多是为了帮助我们以我们能理解的方式去理解?所以我们是不是故意遗漏了一些东西?

I mean when we spoke with Chomsky, he was quite disparaging about the enterprise of physics. He was talking about, you know, the original hard problem of motion and he said that Newton exercised the machine, you know, the mechanical universe view, but he left the ghost intact. So, you know, we still don't know how mind and consciousness works and all of that kind of stuff. But he was rather pointing to this notion that a lot of physics is idealization. And then there's an interesting question about whether our theories are really to actually understand how the universe works or are they more to help us understand in terms that we can understand. So are we intentionally leaving something out?

Matthieu

是的,我认为两者都有。我是说,它们确实与世界互动,因为物理学中的这些理论对于构建技术非常有用。想想激光。我是说,一半的技术,很大一部分,都来自理论。但另一大部分则相反:技术提出了巨大的问题。所以是的,我认为我们也需要理论来构建一条高速公路,一条快速思考问题的高速公路。然后困难的问题在于,你需要你的理论达到什么精度,这取决于你提出的问题。

Yes, I think both. I mean, certainly they are interacting with the world because those theories in physics are super useful to build technology. Think about the laser. I mean, half of the technologies, a big fraction of them, are coming from theory. But another big portion of it is coming from the other way around: technology is asking immense questions. So yes, I think we need theory also to have a sort of highway, to build your highway of thinking super fast about problems. And then the difficult question is at which level of precision do you need your theory to be, and that depends on the question you're asking.

导师与轨迹 Mentors and Trajectory

Host

那么谁是你的导师?谁激励了你?你读了什么书?你是如何走上现在这条物理学家之路的?

And who were your mentors? Who inspired you? What books did you read? How did you kind of land on your current trajectory as a physicist?

Matthieu

嗯,这对我来说是个复杂的问题,因为我父母其实都是物理学家。当我开始读博士时,我试图通过转向经济物理学来逃离他们,做更多金融和经济方面的事。然后在博士期间,我再次被物理学以及沙子如何流动等事物所吸引。后来我做博士后时,总是被我们的大脑所迷住,我们如何思考。那时我甚至花了一年时间在类似珍妮莉亚农场这样的神经科学研究所,遇到了很多很棒的人,学到了很多。但我觉得在那个阶段,很多理论,比如你有很多连接的神经元以及它们的动力学,有点偏向应用数学,与真正的功能脱节。而关于如何学习智能、规则、约束、语言等问题,还没有达到我真正想操作的水平。所以我放弃了,回到了物理学。然后我认为,本质上是技术的发展摆在我们面前。所以我喜欢的一个类比是工业革命,当时热机出现,这又是一个技术先行的案例,然后你必须理解它们背后的原理,它们能有多高效,效率的极限是什么,等等。然后卡诺,一位法国物理学家,写了一篇优美的文本,读起来像哲学,几乎没有数学,引入了熵等概念,那是热力学的开端。非常深刻的思想来自某个技术事实。在这里我认为也是一样的。这些机器令人惊叹。你看,它们有创造力。你给它们一堆图像,那些扩散模型突然像画家一样,它们构建新面孔,组合新面孔。怎么会这样?或者它们创造出从未听过的句子。而乔姆斯基和其他人说这将会极其困难。它们做到了。所以为什么?所以是的,是被问题所吸引。这就是驱动力。

Well, that's a complex question for me because my two parents are physicists actually, and when I started to do my PhD, I tried to escape them by going to econophysics, doing more finance and economy. And then already during my PhD, I started to be fascinated again by physics and how sand flows and things like that. And then when I was a postdoc, I was always mesmerized by our brain, how we think. And I tried at that time to, you know, I even spent one year in a place like Janelia Farm, a neuroscience institute, and I met lots of fantastic people and I learned a lot. But I felt at that stage that a lot of the theory, you know, how you have many connected neurons and their dynamics, was a bit applied math and detached from real function. And then the question of how do you learn intelligence or rules, constraints, language, and so on, it was not at the level that I really wanted to operate. So I gave up, I went back to physics. And then I think it's essentially the development of the technology that's facing us. So I think the analogy I like is the industrial revolution, when heat engines emerged, again one case where technology was first, and then you had to understand what's behind them, how efficient can they be, the limit to their efficiency, and so on. Then Carnot, a French physicist, came up and wrote a beautiful text that reads like philosophy, essentially no math, introducing concepts like entropy, and it was the beginning of thermodynamics. Very deep ideas coming from some technological fact. And here I think it's the same. With those machines, it's amazing. You look, they are creative. You give them a bunch of images and those diffusion models suddenly like a painter, they build new faces, they compose new faces. How can it be? Or they create sentences that they have never heard before. And Chomsky and others said that it would be extremely hard to do. They do it. So why? So yeah, it's being fascinated by questions. That's what drives me.

乔姆斯基的批评 Chomsky's Criticism

Host

是的,你知道,抱歉又提起乔姆斯基,但他说大语言模型就像推土机。他说,‘我爱推土机。它们清雪很棒,但它们不是对科学的贡献。’他还说,‘我有一个理论。什么都行,对吧?它探索所有自然法则,任何可能的东西。’他说,当你有一个科学理论时,你必须解释为什么事情是这样,为什么不是那样。但你刚才说,当我们发现蒸汽机时,我认为你相信那实际上是构建理论的垫脚石。但对乔姆斯基来说,能力和表现之间有着巨大的区别。他有一个关于深蓝的精彩表述,那个下棋的东西,他说那有点像推土机赢得举重比赛。所以这几乎无关紧要。这是不连贯的。

Yeah, I mean, you know, sorry to bring Chomsky back again, but he said that LLMs are like bulldozers. You know, he says, 'I love bulldozers. They're great for clearing the snow, but they're not a contribution to science.' And he says, you know, 'I've got a theory. Anything goes, right? It explores all the laws of nature, anything that can be.' And he says that when you've got a scientific theory, you have to explain why are things this way, why are things not that way. But you were just saying, you know, when we discovered the steam engine, I think you believe that actually was the stepping stone to building theory. But for Chomsky, there's a huge difference between competence and performance. He had this wonderful expression about Deep Blue, the chess thing, and he said that's a little bit like a bulldozer winning a weightlifting competition. So it's almost inconsequential. It's incoherent.

Matthieu

是的。所以我非常尊重他。我们很多工作实际上受到他的启发。我们建模数据的方式是他构建的分类的一部分。但我认为他说的没有错。这不是一个理论,但显然没有错。但这不是重点。对我来说,重点是这是一个惊人的观察,它引发了一系列问题。你面对一个能学会创造力的机器。你可以打开它。你可以看它的神经元,人工神经元,你可以问它如何编码句法、语义等等。它是怎么做到的?你知道,即使最终大脑不是那样运作的。我认为在某种意义上,我很想思考大脑,并对大脑有所见解,但我发现机器中的智能本身也极其有趣。所以是的,就像卡诺的时代。好吧,我们面临新技术,它提出了许多问题。我认为我们应该,我希望物理学家,你问的是物理学,我们正在推动物理系也投资于这些方向。所以是的,不幸的是,有可能让一个系统以许多不同的方式执行功能,对吧?例如,大语言模型,它们显然有语言能力,但这是不同的。

Yeah. So I respect him a lot. A lot of our work is actually inspired by him. The way we model data is sort of part of the classification he built. But I think what he's saying is not wrong. It is not a theory, but it's obviously not wrong. But it's not the point. To me, the point is that it's an amazing observation and this is raising a bunch of questions. You have a machine facing you that can learn to be creative. You can open it up. You can look at its neurons, artificial neurons, and you can ask how is it encoding syntax, semantics, and so on. How is it doing it? You know, even if ultimately the brain doesn't work like that. I would argue in some sense, I would love to think about the brain and to have something to say about the brain, but I find intelligence in machine also extremely interesting in itself. So yeah, it's just like at the time of Carnot. Okay, we're facing new technology, it's asking many questions. I think we should, I hope physicists, you're asking about physics, we're pushing physics departments to also invest in those directions. So yeah, unfortunately it is possible to make a system perform a function in many many different ways, right? So for example, LLMs, they apparently have linguistic competence, but it's different.

抽象与学习 Abstraction and Learning

Host

然后你可能会说,哦,只要它做了我想让它做的事,那就不重要了。但我觉得我们可以更严格一点。有一件事似乎缺失了,那就是抽象概念的获取。你在这方面做过一些很棒的工作,但当然,当我使用语言模型时,有一点对我来说非常明显,即使我们在做爬山算法并解决这些数学问题时,它们也会穿越“意大利面怪物”,得到正确答案,但却是出于错误的原因,而且它们似乎处于抽象之山的低处。而我们人类有能力做这种粗粒化处理,或者使用隐喻,在抽象之山的更高处工作。我只是觉得,当我们使用随机梯度下降并仅从数据中学习时,模型并没有获得这些高层次的抽象。

And then you might just say, oh, it doesn't matter if it does the thing I want it to do, it doesn't matter. But I think we can be a little bit more rigid here. One thing that seems to be missing is the acquisition of abstractions. Now, you've done some amazing work on this, but certainly when I use language models, what's abundantly clear to me, even when we do hill climbing and we solve these mathematical problems, they traverse the spaghetti monster and they get the right answer, but for the wrong reasons, and they seem to be low down the abstraction mountain. And what we do is we have the ability to do this coarse graining or to use metaphor and work higher up the abstraction mountain. And it just feels to me that when we use stochastic gradient descent and we just learn from data, the models aren't acquiring these high-level abstractions.

Matthieu

恰恰相反。我们……所以基本上我认为这些问题非常有趣,非常深刻,我们正以物理学家的方式来处理它们。所以是的,我们思考这个问题的方式是,是的,世界非常复杂。所以也许我们坚持语言。我认为要理解机器如何工作,我们首先需要很好地理解需要学习的数据是什么。在这一点上,我认为语言学家在描述他们处理的数据方面是最令人印象深刻的领域之一。特别是乔姆斯基等人,几个世纪以来一直认为文本的底层是树,你知道,不同层次的抽象。对于图像也有类似的论证,称为模式理论。你知道,我们在物理学中对此非常熟悉。所以在物理学中,假设你取一种液体。我之前告诉过你液体。你可以在原子层面描述它。但如果你杯子里有数十亿数十亿的原子,那对你描述杯子并没有太大帮助。所以我们物理学家所做的是构建粗粒化变量,比如压力、场、速度、密度等等。对于这个系统,这有点简单,因为如果你愿意,基本上有两个描述层次,本质上是微观和宏观。真实数据有分层的多尺度描述层次。如果你想到一张图像,你可以在像素级别思考,非常低层次,而在非常高的层次,你可以把它想成描述图像内容的标题。而且你有很多中间步骤。在低层次,你可以从像素开始形成边缘和小几何图形,在某个时刻你可以形成眼睛、鼻子和耳朵,并理解它构成一个头。所以你有许多不同的描述层次。所以你问的问题正是我们想理解的。所以第一件事是我们如何建模?因为如果你看看最复杂的上下文无关文法,这些是基于乔姆斯基引入的树的模型。本质上,这些模型背后的想法是,如果你想描述像语言文本这样的线性对象,你可以通过某种底层树来描述它,并且你有隐藏变量生活在这棵树上,你描述这些隐藏变量如何产生隐藏变量字符串的方式。所以本质上你是在描述一种递归地生成句子的生成方式。但同样,如果你想把这些模型用于英语,那非常复杂。那么什么是好的模型?同样,这取决于你问的问题。没有绝对的答案。但对于我们正在问的那种问题,就像你问的,我们在什么意义上构建这些粗粒化变量?我们构建数据的上下文无关文法模型,其中底层有那些树。但像你作为物理学家,我们构建一个合成世界。一旦我们构建了它,我们玩的游戏是我们必须相信它足够丰富。所以在我们的例子中,我们希望它捕捉到世界存在某种层次隐藏结构的事实。所以我们捕捉到了这一点,但我们希望使它易于处理。所以也许这有点技术性。在这种情况下,我们开始第一个模型,基本上是一棵在几何上冻结的树,而生产规则,即树如何产生字符串,是随机选择的。最后,随机性,虽然违反直觉,但在物理学中常常使事情更简单,允许我们计算这个模型中的任何相关性,从那里我们实际上可以理解机器如何学习它们。确实,如果你有一个糟糕的机器,比如一个非常浅的网络,首先,即使是那些模型也基本上无法学习。所以有很多话要说,也许我可以回到高维的情况,那非常困难。但如果你有一个深度架构,我们发现的情况恰恰与你说相反。深度架构能够解决这个任务的原因,恰恰是因为它们理解,就像物理学家理解压力、速度、场一样,它们从数据的统计中理解这个隐藏的层次结构。否则它们永远无法完成这项工作。所以它们理解存在这个隐藏的层次结构,并且从中它们可以执行你想要的任务。

That's the opposite. We will... So essentially I think those questions are super interesting, super deep, and we are approaching them as physicists. So yes, so the way we're thinking about this question is that yes, the world is very complicated. So let's stick to language maybe. And I think to understand how machines work, we first need to understand very well what is the data that needs to be learned. And on that, I would say the linguists have been one of the most impressive fields in terms of characterizing the data they are dealing with. In particular, Chomsky and others have argued for centuries that underlying texts are trees, you know, a different level of abstraction. It's also been argued for images. It's called pattern theory. You know, we're very used to that in physics. So in physics, let's say you take a liquid. I told you about liquid before. You can describe it at the level of atoms. But if you have billions of billions of billions of atoms in your glass, it's not going to help you describe the glass so well. So what we did as physicists is to build coarse-grained variables, like pressure, field, velocity, density, things like that. For this system, it's sort of simpler because there is a single... I mean, there are two levels of description if you want, essentially microscopic and macroscopic. Real data have layered multiscale levels of description. If you think about an image, you can think of it at a pixel level, very low level, and at a very high level you could think of it as the caption that describes what's on the image. And you have very intermediary steps. At low level, you could start to make from pixels edges and little geometrical figures, and at some point you could make eyes and nose and ears and understand that it makes a head. So you have many different levels of description. So the question you are asking is the one we wanted to understand. So the first thing was how do we model this? Because if you look at the most complicated context-free grammar, so those are the sort of models which are based on trees that Chomsky introduced. Essentially the idea behind those models is that if you want to describe linear objects like language text, you can describe it by some underlying tree, and you have hidden variables living on those trees, and you describe the way whereby those hidden variables can give rise to strings of hidden variables. So essentially you're describing a generative way to make sentences in a recursive fashion. But again, if you want to feed those models to English, it's very complicated. So what is a good model? And again, it depends on the question you're asking. There's no absolute answer to that. But for the sort of question we're asking, like the one you're asking, which is in which sense do we build those coarse-grained variables? Well, we build context-free grammar models of data where you have those trees underlying it. But like you as a physicist, we build a synthetic world. And once we build it, the game we're playing is that we have to believe that it's rich enough. So in our case, we want it to capture the fact that there is some hierarchical hidden structure to the world. So this we capture, but we want to make it tractable. So maybe it's a bit technical. In this case, we started the first models essentially at a tree that was frozen in geometry, and production rules, which is how the tree gives rise to a string, were randomly chosen. Finally, randomness, although it's counterintuitive, in physics often makes things simpler, allowed us to compute any correlation in this model, and from that we could actually understand how they are learned by machines. And indeed, if you have a poor machine like a very shallow network, first of all, even those models would be essentially unlearnable. So there would be a lot to say, and maybe I can come back to that in high dimension, it's very hard. But if you have a deep architecture, what we find is precisely the opposite of what you're saying. The reason why deep architectures can solve this task is precisely because they understand, just like the physicist understood about pressure, velocity, field, they understand from the statistics of data this hidden hierarchy. Otherwise they would never be able to do this job. So they understand that there is this hidden hierarchy, and from it they can perform the task that you want.

Host

只是稍微复述一下,以便大家理解,这个想法是,我们在这里谈论的是语法,但更广泛地说,我们认为世界上存在结构化的生成过程。

Just to kind of play that back just so that everyone understands that the idea is that there is... I mean, we're talking about grammar here, but more broadly we think that there are structured generative processes in the world.

Matthieu

是的。

Yes.

Host

所以我们可以把这些看作是某种受约束的生成模型。我们在这里谈论的是句法,当我们做机器学习时,我们观察那个生成模型的输出,而学习过程理想上不应该是记忆原始输出,而应该是抽象地理解生成它的模型,因为创造力在于尊重深层结构和约束。如果你有了结构,你可以继续生成更多更多的东西,并且你遵守规则,你就有创造力等等。所以你说你做了实验。你创建了一个参数化的数学生成模型。你可以有任意多的深度,你发现在浅层网络上,基本上它并没有真正学习任何抽象结构。但是当你有深层网络时,它……

So we can think of those as being some kind of a constrained generative model. So we're talking about syntax here, and when we do machine learning, we look at the output of that generative model, and the learning process ideally should be not to memorize the raw output, but it should be to understand abstractly the model which generated it, because creativity is about respecting the deep structure and the constraints. If you have the structure, you can go on and generate many many more things, and you obey the rules, you're creative and all the rest of it. So you're saying that you've done experiments. So you've created a mathematical generative model which is parameterized. You can have as much depth as you want, and you found on shallow networks that basically it wasn't really learning any of the abstract structure. But when you have deep networks, it was...

Matthieu

是的,没错。所以也许我可以把它放到创造力的背景下,以及我们刚刚与乔姆斯基的讨论中。

Yeah, exactly. So maybe I can put it into the context of creativity and this discussion we just had with Chomsky.

创造力与乔姆斯基论点 Creativity and Chomsky's Argument

Matthieu

所以有一个关于创造力的问题。我将在这个非常狭义的意义上使用这个词:能够生成满足句法规则硬约束的新句子,而这些句子是孩子从未听过的。乔姆斯基提出了刺激贫乏论证,认为从例子中学习变得有创造力实际上是不可能的。这基本上是对他论证的一个非常粗略的总结,但我描述了这样一个事实:你有这种生成性的树状丰富上下文无关语法,假设这一点并真正捕捉到世界是抽象概念层级结构的事实。但还有其他可能的生成语法。有些要简单得多。一种叫做正则语法。这只是一个粗略的类比,但本质上这个想法是,一组词可能会固定下一个词的概率。而乔姆斯基的论点是说,即使你给我一百万句话,我可以用上下文无关语法来拟合这些句子,但我也可以用更简单的正则语法来拟合,它在分类上更简单,但要拟合这些句子,它必须非常复杂,有很多很多规则。而先天论和经验论在这个问题上有很大的争论。再次,我们觉得我们想以物理学家的身份来解决这些问题。所以在我们理想化的世界里,真实世界是,机器能否学会有创造力?我们发现,如果你有一个浅层网络,乔姆斯基所担心的完全正确。你学到一些东西,但你学不到这种有趣的生成语法。你基本上是在记忆,什么也做不了。但如果你有一个深层架构,就会有巨大的隐式偏向来构建那些粗粒度变量。层级架构很容易导致某种迭代计算。所以我们发现,确实,你可以通过接触非常少量的句子来学会有创造力。所以让我们说,在我们的模型中,如果 D 是句子的大小,句子的数量是巨大的,随 D 呈指数增长,但你需要看到的有创造力的句子数量只是随 D 呈多项式增长。所以这些模型确实是他论证的一个反例。最终这来自于机器有很强的隐式偏见。它们不是以平等的方式比较所有假设。如果你是深层的,你就能学到。所以这是那个论证的一个反例。所以在某种意义上,我认为就什么需要是天生的而言,如果你有一个深层架构,那就能做很多。这并不是说我在反对乔姆斯基的论证。这并不意味着他推断的是不正确的,对吧?不是因为我认为一个论证不正确,就说明陈述不正确。我不想暗示我们的大脑只是一个深层网络,而且没有更聪明的机制来学得更好。实际上,大脑可以用比那些机器少十万倍的语言接触来学习。所以我认为关于大脑如何工作有很多问题,而且它们很迷人。

So there is a question of creativity. I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraints of syntactic rules that the child would never have heard before. And Chomsky gave this poverty of stimulus argument, arguing that it was actually impossible to learn to become creative from example. Essentially, this would be a very crude way of summarizing his argument, but I described the fact that you have this sort of generative tree-like rich context-free grammars, assuming that and really capturing the fact that the world is a hierarchy of abstract concepts. But you have other possible generative grammars. Some are much simpler. One is called regular grammar. So this will be a caricature, but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chomsky's argument is to say that even if you give me one million sentences, I can fit those sentences by a context-free grammar, but I can also fit them by a much simpler regular grammar, simpler in its classification, but to fit those sentences it would have to be awfully complicated, many, many rules. And that nativism and empiricism, big debates on this question. And again, we felt like we want to address those questions as physicists. So in our idealized world where the true world is, can a machine learn to be creative or not? And what we find is that if you have a shallow network, what Chomsky worried about is completely true. You learn some, you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those coarse-grained variables. Hierarchical architecture leads very easily to some iterative calculation. And so what we found is that indeed you can learn to be creative by being exposed to a very small number of sentences. So let's say in our model, if D is the size of the sentence, the number of sentences is huge, exponential in D, but the number of sentences you need to see to be creative is only polynomial in D. So these models are really a counterexample to his argument. And ultimately it comes from the fact that machines have strong implicit bias. They are not comparing all hypotheses in an equal fashion. And if you're deep, you learn. So that's a counterexample to that. So in some sense, I would argue that in terms of what needs to be innate, if you have a deep architecture, that does lots of that. This is not to say that I'm arguing against Chomsky's argument. It doesn't mean that what he inferred is incorrect, right? It's not because I think an argument is incorrect that the statement is incorrect. I don't want to imply that our brain is just a deep net and that there are not much smarter mechanisms to learn much better. And actually the brain can learn with 100,000 times less exposition to words than those machines. So I think there are lots of questions about how the brain works and they are fascinating.

Host

我认为你们可能都是对的。所以乔姆斯基,即使在 50 年代,他也不是第一个这样做的人,但他提出了这些非常基本的转换规则,可以组合在一起,这持续了相当长一段时间,但他们意识到有很多问题和边缘情况,然后最终极简主义纲领出现了,它更加简洁,就像移动和合并,而且非常抽象。因为我同意你的观点,这些网络显然具有句法能力,这意味着它们绝对在创造,因为它们绝对能创造新颖的语法句子。但在广泛的环境中,它们并不具有创造力,因为它们不理解我们理解的世界上的许多其他抽象概念,这就是为什么我们需要提示它们去创造。你知道,它们可以渲染一只狗的图像,但它们不够有创造力,不知道什么是有趣且与世界连贯的狗图像。但我想谈的另一件事是,这触及了一个观点,即随着我们训练网络的时间更长,它们变得更大更深,它们开始分解。因为我们有这样一个概念,它们有这些破碎纠缠的表征,它们在浅层理解事物。它们理解一些东西而不理解其他东西。我和 Goodfire 的 Tom McGrath 谈过,那是一家大型机械可解释性公司,他多年来一直在研究网络,他说随着网络变得越来越大,它们变得更加分解,他实际上相信它们正在趋向某种自然的分解,但目前有点奇怪,它们在某些方面有分解,然后在其他方面有分裂。但你有点明白我的意思,因为我们知道神经网络是一个有限状态自动机,而乔姆斯基意义上的语言介于上下文无关和上下文敏感之间。他对此不是特别具体,但我们知道作为一个数学事实,它不是乔姆斯基所描述的抽象意义上的生成语法,但在那些模型中,它仍然是一个在某种较低意义上的连贯生成语法。

I think it's possible that you're both correct. So Chomsky, even back in the 50s, and he wasn't the first to do this, but he came up with these very basic transformative rules that could be composed together, and that went on for quite a while, but they realized there were lots of problems and edge cases, and then eventually the minimalist program came out, and it was even more parsimonious, it was like move and merge, and that's very abstract. Because I agree with you, these networks clearly have syntactic competence, which means they absolutely are creating, because they absolutely can create novel grammatical sentences. But in a broad setting though, they're not creative because they don't understand many other abstractions in the world that we do, which is why we need to prompt them to be creative. You know, they can render an image of a dog, but they're not creative enough to know what an interesting and worldly coherent image of a dog is. But another thing I wanted to get to is this is touching on the idea that networks as we train them for longer and they get bigger and deeper, they start to factorize. So because we have this notion that they have these fractured entangled representations, they understand things at a shallow level. They understand some things and not other things. I spoke with Tom McGrath at Goodfire, it's a big mechanistic interpretability company, and he's been studying networks for years, and he says as they get bigger and bigger they become more factorized, and he actually believes they're converging towards some kind of natural factorization, but at the moment it's a bit weird that they have some factorization and then they have some fractionation in other areas. But you kind of see where I'm going with this, because we know that a neural network is a finite state automaton, and language in Chomsky's sense is somewhere between context-free and context-sensitive. He wasn't super specific about that, but we know as a mathematical fact that it's not a generative grammar in the abstract way Chomsky was describing, but it is still a coherent generative grammar in some lower sense in those models.

Matthieu

首先,我们发现,随着你越来越多地训练机器,那些分解或抽象会逐步创建,如果你有一台巨大的机器,你给我越来越多的数据,那么你开始处理越来越抽象的概念。从这种观点来看,这些是最难学的。所以这是我对你问题的回答。第二个问题是,是的,我想你指的是 Transformer 有有限深度的事实,所以如果你考虑那些有 50 个补语的句子,并且像那样循环,可能很难重现,等等。但我认为那些更像是学术问题,你在实践中永远不会遇到,因为循环 50 次的句子极其罕见。所以我不确定它是否真的,我知道有些人对此投入了很多关注,但我更,作为一个喜欢做实证研究的人,我不知道这些担忧在实践中是否真的相关。

First of all, what we find is that as you train the machine more and more, those factorization or abstraction are created progressively, and if you have an immense machine and you give me more and more data, then you start to play with more and more abstract concepts. Those are the hardest to learn in this viewpoint. So that's my take on the being of your question. The second question was, yeah, I think you're referring to the fact that transformers have finite depth, and so if you are thinking about sentences where you have 50 complements of sentences and that are looping like that, it may be very hard to reproduce, and so on. But I think those are more like academic problems that you never encounter in practice, because sentences that loop for 50 times are extremely rare. So I'm not sure it's really, I know that some people put a lot of attention on that, but I'm more, as someone who likes to do empirical studies and so on, I don't know if those worries are actually relevant in practice.

Host

但你同意存在一个抽象谱系吗?所以,是的,这是一种不同类型的句法能力,但也许这并不重要。有表现能力这种东西。

Would you agree though that there is a spectrum of abstraction? So, yes, it's a different type of syntactic competence, but maybe it doesn't matter. There's the performance competence type thing.

编码代理与意图 Coding agents and intention

Matthieu

一个有趣的例子是,我把整个代码库喂给 Claude Code,有意思的是它并不真正理解我的意图。如果我把它放进一个循环里,放进一个智能体里,说“只管修 bug,让这个软件不断进化”,它不会尊重那些深层约束,也就是我的心理约束,比如我到底想通过这个实现什么,我会怎么做。所以它具备句法能力,知道怎么写代码。

An interesting example is I feed my entire codebase into Claude Code and isn't it interesting that it doesn't really understand what my intentions were. So if I put it in a loop, I put it in an agent and I say just fix the bugs and just keep evolving this software, it doesn't respect the deep constraints, now are my mental constraints like what was I trying to achieve with this, what would I have done. So it has the syntactic competence, it knows how to write the code.

Host

你是说这只是网络还不够好的问题吗?当它们真正理解,比如有了心智理论,能更抽象地理解世界如何运作,最终我们就能自主创造出能直接做出 Microsoft Word 之类的编码智能体,而且存在通往那种能力的路径?

Are you saying this is just a matter of the networks aren't good enough yet, when they do understand, you know, when they have a theory of mind and they understand how the world works even more abstractly, eventually we could just autonomously create coding agents that will just make Microsoft Word or something, and there is a path to that level of competence?

组合与创造力 Composition and creativity

Matthieu

所以本质上我们在说的是,想象你有一个扩散模型,它学会组合新面孔。它所做的就是,当它看过足够多的低级特征,比如鼻子、眼睛和嘴巴,它就理解了游戏规则,并把它们组合起来。我们还喜欢这个描述的一点是,我们可以做出非平凡的预测,用真实图像或真实文本来检验,也许我们稍后会回到这一点,因为我认为这是物理学中非常重要的一部分。

So essentially what we're saying is that imagine that you have this diffusion model that learns to compose new faces. Essentially what it's doing is that when it has seen enough low-level features like nose, eyes and mouth, it understands the rules of the game and it composes them together. And what we also liked about this description is that we can make non-trivial predictions that we test with real images or with real text, and maybe we'll come back to that because I think it's a really important part of physics.

Host

创造力不只是把满足约束的碎片拼在一起。虽然当你有一个新想法时,往往是把现有想法组合成一个新整体。但我认为创造力可以远超于此。我的意思是,如果我们想想我们讨论过的作为物理学家意味着什么,科学如何推进,那就是创造力的例子。如果你想想牛顿对行星运动等的理解。我们谈过实验与理论之间的对话,在好的描述层次上构建模型,类比。而且我不同意你所说的所有这些都在机器里。我认为没有理由我们有一天不能造出能做那事的机器。我不确定仅仅 Scaling 是否会导致那结果。我认为也许我们需要更多反思我们作为科学家如何运作,以提出好的数据集和好的程序来教机器成为好的科学家。

It's not creativity is not just putting pieces together that satisfy constraints. Although when you have a new idea often it's putting existing ideas together into a new whole. But I think creativity can be much more than that. I mean if we think about what we discussed about what it means to be a physicist and how science proceeds, it's an example of creativity. If you think about creativity like Newton's understanding of motion of planets and things like that. We talked about dialogue between experiments and theory, building models at a good level of description, analogies. And I don't think I agree with you that all that is in the machine. I think there's no reason why we would not be able one day to build machines that can do that. I'm not sure if just scaling up things will lead to that. I think maybe we need to do more introspection of how we function as scientists to come up with a good dataset and good procedures to teach machines to be good scientists.

Matthieu

它们只是一个例子,比如科学中的创造力需要很多能力来创造真正新的东西,以及如何与我们周围的世界互动,我认为机器没有这些。然而我同意你,仅仅 Scaling 我认为不会带来完全成功;我们需要在这些机器中发展其他能力。

They're just an example like creativity in science requires a lot of abilities to create something really new and how to interact with the world around us that I don't think machines have. And yet I think I agree with you that just scaling up I don't think will lead to total success; we need to develop other abilities in those machines.

Host

是的,我想我只是想理解差距在哪里,因为那会与你的论点一致。如果你说我们能学习世界的抽象结构,并在一个领域具备生成能力,那我们为什么不能?因为对我来说,创造力不只是关于连贯性和尊重约束。变革性创造力在我心中是关于发现有趣的新子空间。所以我们可以集体地穿越这些约束,有时偶然遇到这些迷人的新子空间,然后继续探索它们,当我们回顾发现它们之后,我们会想,哦,那是一个非常变革性的创造性垫脚石。

Yeah, I think I'm just trying to understand what the gap is because it would be consistent with your argument. If you're saying that we can learn the abstract structure of the world and be generatively competent in one domain, why would we not? Because for me, creativity is not just about coherence and respecting the constraints. Transformative creativity in my mind is about discovering interesting new subspaces. So we can traverse these constraints collectively and serendipitously sometimes we happen upon these fascinating new subspaces and we go on to explore them, and when we look back after discovering them we think oh that was a very transformatively creative stepping stone.

Matthieu

是的,我认为如果你看看科学史,它非常……我的意思是想想数学家发明虚数,或牛顿描述行星运动,或惠更斯关于波和衍射的东西。我的意思是人类一直看到波进入港口并被衍射,所以开始形成更圆形的图案。但如果你只是向机器展示那些图案,它会愚蠢地预测下一帧,因为速度会传播。但作为物理学家我们做什么?我们必须首先——有些人非常擅长观察某些东西很有趣。然后你必须简化几何。所以也许你把它放在一个非常简单的几何中,然后你必须做我描述的建模的一个方面:我们如何建模这个,等等。所以我认为所有这些都非常需要,而且我不明白你怎么能——这是你与世界的互动和简化世界,这需要互动。所以我不认为你可以仅仅通过查看所有曾经写过的东西而不强制这些互动来学习它。

Yes, I think if you look at the history of science, it's very... I mean think about mathematician inventing imaginary numbers or Newton describing the motion of planets or Huygens and something about waves and diffraction. I mean humans forever have seen waves entering a port and being diffracted, so starting to make more circle-like patterns. But if you're just showing those patterns to a machine, it would sort of stupidly predict the next frame because velocity will propagate. But what do we do as physicists? We have to first—some people are really good at observing that something is intriguing. Then you have to simplify the geometry. So maybe you put it in a very simple geometry, and then you have to do an aspect of modeling that I was describing: how do we model this, and so on. So I think all that is very much needed, and I don't see how you could—it's your interaction with the world and simplifying the world, and that requires an interaction. So I don't think you can learn it just by looking at everything that was ever written without enforcing those interactions.

刺激贫乏与学习不变性 Poverty of stimulus and learning invariances

Host

是的,回顾一下,我们之前对比过乔姆斯基的“刺激贫乏”论证。他基本上是说,从现实角度看,孩子们拥有的感觉数据量不足以让他们学会这种语法。而你的论文证明了实际上是可以的,因为你创造了这个生成函数,深度网络能学会它。但我想理解是如何做到的。所以你说例如网络会遇到歧义,当有足够多的数据且其中存在偏差时,网络就能突然顿悟并学会这种不变性。给我讲讲那个。

Yes, and to recap, we were contrasting before that Chomsky has this poverty of stimulus argument. He was essentially saying that it's not really possible, realistically, with the amount of sense data that children have, for them to learn this grammar. And your paper demonstrated that actually it is, because you created this generative function and deep networks could learn it. But I want to understand how. So you said for example that the networks encounter ambiguity, and when there's a sufficient amount of data which has a bias in it, then the network can suddenly grok it and it can learn this invariance. Tell me about that.

Matthieu

好的。所以这是关于机器如何真正构建那些粗粒度变量或抽象。我们在各种情况下研究过:监督学习,比如试图分类猫和狗;然后我们转向生成模型,比如下一个词预测或扩散模型;最近我们转向可能更聪明的算法,试图在更抽象的空间中预测。所以也许我可以从中间开始这个讨论。想想扩散模型或大语言模型,它们试图预测非常低级的 token 或像素或低级特征。所以本质上我们论证的是,也许再次与一个更简单的算法类比,即 10 年前引入的 word2vec,一个美丽的想法。那里的想法是如何构建一个有趣的词向量表示:对每个词我想关联一个向量。有一个美丽的想法:你可以取这个词,做一个小机器,一个隐藏层,所以你有一层神经元,训练这个机器预测附近的词。所以本质上它是一个基于共现训练的机器,即两个词在同一句子中出现的频率,简单来说。你在这里意识到的是,如果两个词是同义词,它们会有相似的上下文,所以这个机器会用相同的向量表示这两个词。

All right. So this is about how does the machine actually build those coarse-grained variables or abstractions. And we looked at it in various cases: supervised learning where you're trying to classify cats and dogs, and then we went to generative models like next-token prediction or diffusion models, and very recently we went to maybe smarter algorithms that are trying to predict in more abstract spaces. So maybe I can start this discussion in the middle. So think about models like diffusion models or LLMs that are trying to predict very low-level tokens or pixels or low-level features. So essentially what we argue is that maybe there's an analogy again with a simpler algorithm which is word2vec introduced 10 years ago, a beautiful idea. And so the idea there was how can we build an interesting vectoral representation of words: to each word I want to associate a vector. And there was a beautiful idea: what you can do is take this word, make a little machine, a one hidden layer, so you have one layer of neurons, and train this machine to predict the words nearby. So essentially it's a machine that's trained on co-occurrence, how often two words co-occur in the same sentence, let's say to say it simply. And what you realize here is that if two words are synonyms, they will have a similar context, and so this machine will represent those two guys with the same vectors.

粗粒度变量与抽象 Coarse-grained variables and abstraction

Matthieu

所以这就是,你将不再有那些不同的具体形态,而只有意义。对我来说,这就是粗粒化变量的一个例子。这很关键,因为我们之前谈到,在高维空间中的学习应该极其困难。所以本质上,这台机器能摆脱大量它不关心的东西,这一点非常重要,而要做到这一点,它必须桥接这些粗粒化变量。Word2vec 在一个层面上做到了这一点,在一个抽象层次上,一个低层次的抽象。所以本质上,我们说的是深度架构、扩散模型或大语言模型,它们做的正是这件事,但是以递归的方式。一旦它们理解了意义,它们就会把这些意义组合成更高层次的意义。也许让我举个例子。想想街道,街道这个概念。有行人、有汽车、有人行道,有无数种可能的街道,但有一个概念能把所有这些不同的配置组合在一起,会非常有用。这就是街道的概念。所以如果你想想大语言模型或扩散模型,它们构建概念的方式,是从统计中涌现出来的。那些抽象是涌现出来的,它们就在数据中,它们涌现出来,而如果你把那些能预测相似上下文的配置组合在一起,这些概念就会涌现出来。所以也许如果你有一条街道,通常附近有房子,也许房子有颜色或边缘,所以你会预测颜色和边缘。这样一来,至少在这些模型中,你会发现如果你有足够的数据,你能学到所有的抽象。但随着你变得越来越抽象,你会遇到一个问题,因为你总是试图通过说它们如何具有预测性来构建这些抽象,但这是在非常低的层次上。而当你非常抽象时,你如何预测像素或颜色等,是一个超级嘈杂的信号。所以本质上,这就是为什么在这些模型中我们发现——我们有经验证据,我很乐意谈谈经验证据——越抽象的概念越难学,因为本质上你的信号,随着你变得越来越抽象,会被稀释。

So this is, and so you will have, instead of the incarnation of those different, you will just have the meaning. So this is for me an example of coarse-grained variable. And essentially it's key because we talked about the fact that learning in large dimension should be extremely hard. So essentially it's really important that this machine manages to, in some sense, get rid of a lot of things they don't care about, and to do that they have to bridge those coarse-grained variables. So word2vec is doing that at one level, at one level of abstraction, at a low level of abstraction. So essentially what we're saying is that deep architectures, diffusion models or large language models, they do exactly that but in a recursive fashion. So once they have understood the meaning, they will group this meaning into supra-meanings. So maybe let me give an example. Think about streets, the concept of streets. You have passersby, you have cars, you have sidewalks, you have an immense number of possible streets, but it would be very useful to have a concept that groups all those different configurations together. And that's the concept of street. And so if you think about LLMs or diffusion models, the way they build concepts, they emerge from statistics alone. Those abstractions emerge, they are there in the data, they emerge, and those concepts emerge if you group together configurations that predict similar contexts around them. So maybe if you have a street, typically you have houses nearby, and maybe the houses have colors or edges, and so you would predict color and edges. And with that, in those models at least, you find that if you have enough data, you can learn all the abstractions. But as you get more and more abstract, you have a problem because you're always trying to build those abstractions by saying how they are predictive, but at a very low level. And when you're very abstract, how you predict pixels or colors and so on is a super noisy signal. So essentially that's why in those models we find—and we have empirical evidence, and I'm happy to talk about empirical evidence—that the more abstract concepts are the toughest to learn, because essentially your signal, as you get more and more abstract, gets diluted.

Host

嗯。

Yeah.

Matthieu

所以这就是,是的。这就是我们认为你构建那些潜在变量或抽象的机制。

So this is yes. So this would be the mechanism whereby we think you build those latent variables or abstractions.

Host

是的,这个想法太诱人了,我们一会儿就会讲到潜在变量,因为那也是值得一谈的好话题。但你是不是在暗示存在某种自然的分解?那么你认为不同的网络,也许架构不同,在给定相同数据的情况下,会几乎收敛到相同的逻辑分解吗?这种层次化的分解。

Yeah, it's such a tantalizing idea, and we'll get to the latent stuff just in a minute because that's also a great thing to talk about. But are you suggesting that there is some kind of natural factorization? So do you think that different networks, perhaps with different architectures, given the same data, would almost converge towards the same logical factorization of the data, this hierarchical factorization?

Matthieu

所以在我们从数学上发明的理想世界里,只要网络是深的,这是成立的。浅层网络什么也做不了,但如果你有 Transformer 或 CNN 等等,它们本质上会构建非常相似的抽象。实际上,在这些论证中,也预测了你需要多少数据来构建它们。所以这是一种证据——我们做预测,这就是我们玩物理的方式:我们做预测,然后检验它们。所以我们预测你需要多少数据来学习多少个不同层次的抽象,我们发现,在 Scaling 方面,不同架构所需的数据量是相似的。所以我仍然认为,不同的架构会导致细微的差异,不同的电路等等,但总体而言,我认为是的,那些抽象确实存在。再说一遍,它们是一组能预测相似周围环境的配置,如果它们有很强的信号,如果它们有很强的预测能力,它们会更早形成。

So in our dream world that we invented mathematically, it's true as long as the networks are deep. So shallow networks they don't do anything, but if you have a transformer or CNN and so on, they build essentially very similar abstractions. Actually, in this, those arguments also predict how many data you need to build them. So this is a sort of evidence—we make predictions and that's what we play as physics: we make predictions and we test them. So we predict how many data you need to learn how many different levels of abstraction, and we find similar in terms of scaling, similar number of data independently of the architecture. So I do still think that different architectures are going to lead to slight differences and so on, different circuits and so on, but the big picture I think is yes, those abstractions actually really exist. And again, they are a set of configurations that predict a similar surrounding, and if they have a strong signal, if they have a strong predictive power, they are formed earlier.

Host

把它们想象成唯一真实的抽象,这太诱人了。但正如我们之前所说,我们知道它不像乔姆斯基所说的合并操作。当我们做机制可解释性研究,比如观察网络如何做加法时,很奇怪它们会把三角函数组合在一起;它们不是用我们能理解的方式做的。也许那只是架构的局限。也许如果我们有合适的可学习图灵机,它们会在抽象树的更高层收敛。但我想我们需要讨论的一个边缘话题是维度灾难。所以一直存在这个统计规律,基本上在高维空间中,要使问题可处理所需的数据量呈指数增长。关于为什么不是这样,有各种理论。有流形假设,即内在维度更低。我们和 Randall Balestriero 聊过这个。他有神经网络样条理论,他说在高维空间中,所有数据都是外推。没有流形。它实际上是以输入敏感的方式进行样条分解。很多人对此有不同的看法。

It's so tantalizing to think of them as being the one true abstractions. But we know, as we said earlier, it's not like the merge operator that Chomsky was talking about. When we do mechanistic interpretability and look at how networks do addition, for example, it's super weird that they're composing trigonometric functions together; they're not doing it the way we can. And maybe that's just a limitation of the architecture. Maybe if we had proper learnable Turing machines, they would converge higher up the abstraction tree. But I suppose a tangential thing that we need to talk about is this curse of dimensionality. So there's always been this statistical law essentially that when we have high dimensions, the number of data that you need to make it tractable increases exponentially. And there are all of these theories about why that's not the case. There's the manifold hypothesis. So the intrinsic dimension is lower. We spoke with Randall Balestriero about this. He's got this spline theory of neural networks, and he said that in high dimensions all data is extrapolation. There's no manifold. It's actually doing this spline decomposition in an input-sensitive way. Lots of people have different ideas about this.

Matthieu

但你说这种涌现行为实际上是它可处理的方式。

But you're saying that this kind of emerging behavior is actually how it is tractable.

Matthieu

完全正确。所以实际上,在我们开始思考创造力之前,我们的第一个工作就是试图理解什么样的数据结构能让深度网络真正发挥作用。所以正如你所说。也许我可以再说一遍。在物理学中,我们知道体积等于长度的维度次方。所以在三维中,L 的立方;在二维中,L 的平方;L 是长度。所以想想高维。如果你想想一张图像,D 至少天真地说是像素的数量。如果你想想文本,它可能天真地说是句子中的单词数。所以那些体积是巨大的。它们是指数级的。它们在维度上是指数级的。所以这意味着,即使你给我一万亿个点,因为体积如此巨大,它们彼此之间也极其遥远。极其遥远。所以如果你有一个只是插值的机器,现在你问一个关于新测试点的问题,你可以从数学上证明,如果数据几乎没有结构,比如你试图学习回归某个平滑函数,那是没有希望的。我的意思是,你唯一能外推并拥有泛化能力的方法,是把这些点聚在一起。这意味着你需要指数级数量的数据。你拥有的数据比宇宙中的原子还多。所以那根本不可能。所以对我来说,这是一个完全根本的问题,而且确实,在文献中有时会把它抛在一边,说,好吧,说维度是图像上的像素数量是超级天真的。

Exactly. So actually, that's before we started thinking about creativity, our first work was really trying to understand what data structure allows deep nets to actually perform. So it's exactly as you said. Maybe I can say it again. So in physics, we know that a volume goes like a length to the exponent of the dimension. So in 3D, L cubed; in 2D, L squared; L is a length. So think about a large dimension. So if you think about an image, D may be the number of pixels, at least naively. If you think about text, it may be the number of words in your sentence, again naively. So those volumes are huge. They're exponential. They're exponential at large in the dimension. So what it means is that even if you give me one trillion points, because the volume is so huge, they're extremely far away from each other. Extremely far away. And so if you have a machine that's just interpolating, and now you ask a question about a new test point, you can prove mathematically that if the data has little structure, like you're trying to learn to regress some function that's smooth, it's hopeless. I mean, the only way you will extrapolate and have power to generalize is if you bring those points together. It means you have an exponentially large number of data. You have more data than atoms in the universe. So it's just impossible. So to me, this is a completely fundamental question, and it's true that sometimes in the literature it's tossed aside by saying, okay, it's super naive to say that the dimension is the number of pixels on an image.

维度灾难与深度架构 Curse of Dimensionality and Deep Architectures

Matthieu

事实上,数据确实应该位于一个低维流形上。如果你去测量,它确实位于低维流形上,但这个维度仍然很大。对我来说,最大的问题是,如果这就是答案,那就意味着像核方法这样非常简单的算法是深度网络甚至浅层网络的祖先。我是说,它们做得完美。如果你给它们一个低维流形,你不需要任何有趣的架构。但如果你在文本上使用它们,我可以告诉你它惨败。我是说,它完全不起作用。所以问题是,为什么需要深度架构?你说的一些东西并没有回答这个问题。所以这才是我们真正在寻找的问题。本质上,答案是,如果世界是分层的,如果它有那些隐藏的粗粒度变量,那些机器非常擅长发现它们,而且它们可以用不是很大的数据量来发现它们,再次强调,是维度上的多项式。一旦它们发现了它们,这就是对数据的一种总结,而不是逐像素描述。哦,有鼻子、耳朵等等。所以你本质上是在降低问题的维度,你可以解决维度灾难。所以我认为这个解释的优势在于,无论你提出什么解释,它都必须解释为什么需要深度网络。

In fact it should really be that the data lie in a lower dimensional manifold. And if you try to measure it, it's true that it lies in a lower dimensional manifold, but this dimension is still large in dimension. And to me the big problem is that if this was the answer to this question, it would mean that very simple algorithms like kernel methods are ancestors of deep nets or even shallow networks. I mean they do it perfectly. And if you give them a low dimensional manifold you don't need to have any interesting architecture. But if you use them on text, I can tell you it fails lamentably. I mean it completely does nothing. So the question is why do you need deep architectures? And some of the things you said do not answer that question. So that's really the question we're looking after. And essentially the answer to that is that if the world is hierarchical, if it has those hidden coarse-grained variables, those machines are super good at discovering them, and they can discover them with a number of data that is not huge, polynomial in the dimension once again. And once they discover them, it's a sort of summary of what the data is, you know, instead of describing pixel by pixel. Oh, is there a nose, ears, and so on. So you're reducing the dimension of the problem essentially and you can solve the curse of dimensionality. So I think this explanation has the advantage that whatever explanation you come up with, it has to explain why you need deep networks.

Host

是的。是的。你知道,前几天我和 Goodfire 的 Tom 交谈时,他说可解释性很大程度上是从神经表示到文本,试图内省它们。他认为我们可以有一种新的训练方法,从文本回到神经表示。所以我们发现这些新兴的模块化结构,在训练过程中我们鼓励它们变得更加纯粹、更加进化。但也有其他人谈论类似的想法。例如,Yann LeCun 有一个想法叫做联合嵌入预测架构。这非常相关,因为你有一篇论文基本上说我们应该在潜在空间预测,而不是在词元空间。他的想法本质上是,如果我们真的在潜在空间预测,那么我们可以比在环境空间预测更高效地利用样本。给我讲讲这个。

Yes. Yes. You know when I was speaking with Tom from Goodfire the other day, he was saying that so much of interpretability is going from essentially neural representations to text, to try and introspect about them. And he thinks we could have a new type of training method where we go from text back to neural representation. So we discover these emerging modular structures and during training we encourage them to be even more pristine, even more evolved. But there are other folks talking about similar ideas as well. Yan LeCun, for example, he's got this idea called a joint embedding prediction architecture. And this is very pertinent because you've got a paper out basically saying that we should predict in the latent space, not the token space. And his idea essentially is that if we actually predict in latent space, then we can be significantly more sample efficient than if we predict in the ambient space. Tell me about that.

Matthieu

是的。所以我……好的。所以这是我们过去一两年一直着迷的问题。正如我们刚才讨论的,大脑学习语言的数据比机器少得多。所以机器很神奇。它们说英语肯定比我说得好。但在某种智能定义下,它们做这些任务需要的数据比我们多得多。那为什么我们如此不同?有很多假设,但领域内讨论的一件事是,那些大型语言模型最终做的事情看起来有点琐碎。就像你遮住一个词元,然后试图发现它。即使在我们的模型中做到这一点,你也会发现你需要理解世界的完整分层抽象。即使要做好这一点。实际上,我们开始研究下一个词预测,因为我一直有,至少十年了,这个与维度灾难相关的问题:当我们产生语音时,想想句子的结尾,也许我之前说了 30 个词,可能的句子数量是巨大的。我怎么能记住那 30 个词来做这件事?这怎么可能?实际上,那些模型给出了一个优雅的答案,因为当你试图预测下一个词元时,如果你说了一个长句子,也许你会有一个粗粒度变量来描述句子前半部分的粗略含义,而当你接近你要说的内容时,你会有越来越精细的描述。所以至少对我来说,这种思维方式为我的悖论提供了一个可能的解决方案。但无论如何,即使你试图学习下一个词元,你也需要构建那些抽象。但我告诉过你,这样做的一个问题是,如果你非常抽象,它需要大量数据,因为你通过将数据中预测相似周围环境的配置组合在一起来构建这些抽象,但在低层次上,比如相似的像素。所以回到正题,文献中提出的建议,实际上在神经科学中也很有趣,有一种观点认为大脑可能在做某种非常有趣的自监督学习,而不是仅仅预测眼睛的下一个画面,它试图预测其皮层下一个活动。所以在某种抽象空间中预测。这些想法也出现在机器学习中,你提到了 Yann LeCun,还有其他模型,它们非常有趣。再次,这个想法是,与其在词元层面预测,我能否在更抽象的空间预测?他们开发了非常有趣的机器来做这件事。你可以想到孪生网络:你有一个机器,你复制它,一个机器看到完整数据。它是老师,另一个机器看到数据的一些遮挡版本,你的学生必须预测的不是被遮挡的词元,而是能看到它们的老师如何表示那些词元。这很美,对吧?所以那些网络在做某种内省。而且一直有很大的争论:它是否更好?因为毕竟,那些 LLM 在做奇妙的事情。所以我们觉得在样本复杂度方面基本上没有理论,所以我们觉得我们需要定量地思考这个问题。所以再次,用同样的模型,我们玩的游戏是开发一个框架,用单一视角来应对许多不同的问题。所以维度灾难、创造力,现在是从自己的潜在空间学习。还有缩放定律,也许我们会谈到那些。所以再次,在这些模型中,我们发现那些内省的、从自己的潜在空间学习的算法在样本复杂度方面要强大得多,它们最终会学到相同的抽象,但更快。要构建抽象,你需要组合配置。再想想街道。

Yes. So I... Okay. So that's a question we've been fascinated by in the last year or two. As we just discussed, the brain learns languages with much less data than machines. So machines are amazing. They speak better English than me, for sure. But in some definition of intelligence, they need many more data than us to do those tasks. So why are we so different? There are many hypotheses, but one thing that's discussed in the field is the fact that those large language models at the end do something that seems a bit trivial. It's like you mask a token and you try to discover it. Even to do that in our models, you find that you need to understand the full hierarchical abstraction of the world. Even to do that well. Actually, we started to work on next token prediction because I always had, for at least one decade, this sort of question related to the curse of dimensionality, which was: how come when we produce speech, think about the end of a sentence, maybe I said 30 words before, the number of possible sentences is huge. How do I need to memorize those 30 words to do that? How is it possible? And actually those models gave a sort of elegant answer to that, because what happens when you try to break the next token is that if you said a long sentence, maybe you would have a coarse-grained variable that describes a coarse meaning of the first half of the sentence, and as you approach what you're going to say, you have a finer and finer more precise description. So at least to me, this sort of way of thinking led to a possible solution for my paradox. But in any event, even if you try to learn the next token, you need to build those abstractions. But I told you that one problem with doing this is that if you're very abstract, it needs a lot of data, because you build those abstractions by bringing together configurations in the data that predict a similar surrounding, but at a low level like similar pixels around. So going back, what has been proposed in the literature, actually it's interesting also in neuroscience, there is this notion that maybe the brain is doing some sort of very interesting self-supervised learning where instead of just predicting what's going to be the next frame on its eyes, it's trying to predict the next activity of its cortex. So predicting in some sort of abstract space. And these ideas also emerged in machine learning, and you talked about Yann LeCun, and there are also other models, and they're extremely interesting. And again, the idea is instead of predicting at the level of token, can I predict in more abstract space? And they developed very interesting machines to do that. You can think about twins: you have one machine, you duplicate it, and one machine is shown the entire data. It's a teacher, and one machine is shown some occluded version of the data, and your student has to predict not the tokens that were occluded but how those tokens were represented by the teacher that could see them. It's beautiful, right? So those networks are doing some kind of introspection. And there's been a big debate: is it better or not? Because after all, those LLMs are doing fantastic things. And so we felt that there was essentially no theory on that on sample complexity, and so we felt we needed to think quantitatively about this question. And so again, with the same kind of model, the game we're playing is to develop a framework where with a single viewpoint you try to engage with many different problems. So curse of dimensionality, creativity, and now learning from your own latent. Also scaling laws, maybe we'll talk about those. And so again, in those models, what we found is that those algorithms that are introspective, that learn from their own latent, are much more powerful in terms of sample complexity, and they will eventually learn the same abstraction but much faster. To build abstraction, you need to bring configuration. Think again about the street.

层级抽象的效能 Efficiency of hierarchical abstractions

Matthieu

所有那些配置,你需要理解它是一个实体,一条街道。而扩散模型或下一个词预测所做的是,它们要整合这些配置的信号,必须与像素周围非常低层次的特征相关。我告诉过你,抽象事物与非抽象事物之间的这种相关性是存在的,但非常嘈杂。但想象一下,当你开始理解汽车和行人的概念,也理解了房屋的概念时,那么本质上这些方法可以通过预测那些配置不是房屋的像素而是附近房屋的概念来构建街道的概念,这样信号就大得多,因此你需要更少的数据来从噪声中提取信号。所以是的,我们确实发现,要理解世界的层级结构,至少在那些简单模型中,这样效率要高得多。

All those configurations you need to understand it's one entity, a street. And what diffusion or next token prediction does is that the signal they have to bring those together has to do with pixels around very low-level features. And I told you that this correlation between abstract things and things that are not abstract, it exists, but it's very noisy. But imagine instead that when you started, when you understood the concept of cars and passersby, and you also understood the concept of houses, then essentially what those methods do is they can build the concept of street by predicting that those configurations have not the pixel of the painting of the house but just the concept houses nearby, and then the signal is much larger, and so you need much less data to extract the signal from noise. So yes, we do find that to understand the hierarchical structure of the world, at least in those simple models, it's much more efficient.

Host

是的。所以我的意思是,很多人会知道 Yann LeCun 在视觉领域的工作。比如 Barlow Twins 和所有那些联合嵌入预测架构,大致来说,你有一个类似孪生网络的东西,然后你可能做某种掩码预测。所以你可能会遮挡一侧的图块,然后你在嵌入空间(即潜在空间)而不是环境空间上学习这个预测函数。但这也涉及到他关于基于能量的模型的更广泛哲学。所以大致想法是,你可以将领域特定知识注入预测架构,而且能量是可组合的。所以你可以极其具体,实际上有代表领域中事物的变量。但我们这里讨论的是相当通用的东西。这有点像归纳偏置,并不是真正的领域特定。所以它可以适用于任何类型的视觉,也可以适用于任何类型的语言模型,而且正如你所说,它在样本效率上显著更高。但我们是否仍然存在这个问题:它是否学到了真正好的通用抽象?它更高效,但是否仍然缺少什么?你知道,我们有那些“银河大脑”式的抽象。我们可以从看似无限的可能关系中选出事物之间的元关系,这只是朝那个方向迈出的一步,但还没有完全到位。

Yeah. So I mean, many folks will know Yann LeCun's work in the vision space. So Barlow Twins and all of these joint embedding prediction architectures, roughly speaking, where you have something like a Siamese network and then you might do some kind of mask prediction. So you might occlude tiles from one side, and you're learning this prediction function over the embeddings, the latents, rather than the ambient space. But this also goes into his broader philosophy about energy-based models as well. So the rough idea is that you can imbue domain-specific knowledge into a prediction architecture, and energies are composable. So you could be ridiculously specific and actually have variables that represent things in the domain. But what we're talking about here is something which is quite generic. It's a little bit like an inductive prior which is not really domain-specific. So it could work for any type of vision or it could work for any type of language model, and it's significantly more sample-efficient as you just said. But do we still have this issue that it is learning really good general abstractions? It's more efficient, but is there still something missing? You know, like we have these galaxy brain abstractions. We can just select these meta-relations between things from a seemingly infinite set of possible relations, and is this just one step in that direction, but not all the way?

Matthieu

也许我应该作为物理学家说一句谨慎的话。到目前为止我讨论的是同一种理论方法,但我们可以实证检验它,做出自然预测并测试它们。最后这部分,也就是一个月前的那篇论文,是我们现在正在测试的理论。所以当我谈论它时,我带着谨慎。你知道,我认为作为理论家,我们想要的是有预测性的理论,它们做出预测。对我们来说,严谨并不意味着有定理,而是我们回去测试那些预测。所以我们正在做这件事。

Maybe I should indicate a word of caution as physicists. What I've been discussing so far was the same sort of theoretical approach, but we could test it empirically and make natural predictions and test them. This last part, which is a paper that's one month old, is a theory that we are now testing. So when I talk about it, I talk about it with caution. You know, I think it's nice that's what we want to do as theorists: to have theories that are predictive, they make predictions. And being rigorous for us doesn't mean having a theorem. It's us going back and testing those predictions. So we are in the process of doing that.

Matthieu

所以是的,我的意思是,那里有一个非常深刻的问题。我们在那些简单模型中发现,确实你用更少的数据学到了这种抽象。但现在有一个问题:那些抽象在你的机器中是否被表示,以及你用它们做什么?而且有像从那些表示中你可以做图像分割或分类等任务,现在你可以在分类上与监督方法竞争。所以有证据表明它做得非常好。但例如,如果你想将它们与大型语言模型进行比较,我的意思是,大型语言模型,我们也喜欢它们,因为它们具有生成性。我们可以与它们交谈,然后它们回应。所以例如,这是我们正在研究的一个问题,我仍然不知道答案:一旦我发现了那些变量,本质上我创建了某种世界的编码器。我能否用不那么多的数据创建一个解码器,并从它们构建生成模型?所以我能否真的回去说我可以与那些下一个词预测竞争,并构建生成性的东西?我不知道。所以这对我来说完全开放。所以是的,你可以构建那种非常有趣的世界表示。现在,在我们的模型中,我们知道应该有什么,我们可以检查它在那里。但如果你不知道东西在哪里,你如何利用这些信息最有效地执行特定任务?所以我认为对我来说,这是未来几年一个迷人的研究领域。

So yes, I mean, there's a very deep question there. What we find in those simple models is that indeed you learn this abstraction with much less data. But now there is a question of those abstractions: are they represented in your machine, and what do you do with them? And there are things like from those representations you can do tasks like segmentation in images or classification, and now you can be competitive with supervised methods on classification. So there is evidence that it's doing a very very good job. But for example, if you want to compare them to large language models, I mean, large language models, we like them also because they are generative. We can talk to them and then they respond. So for example, that's a question we're working on and I still don't know the answer: once I have discovered those variables, essentially I created some sort of encoder of the world. Can I create with not so many data a decoder and build a generative model from them? So can I really go back and say I can compete with those next token prediction and build something generative? I don't know. So this is completely open to me. So yes, you can build those sort of very interesting representations of the world. Now, in our model, we know what should be there and we can check that it's there. But if you don't know where things are, how do you use this information to do specific tasks most efficiently? So I think to me it's a fascinating field of study for the years to come.

Host

而在你大约一个月前刚发布的这篇论文中,预测潜在空间而不是词元,你应该解释一下图一。我们现在把它放到屏幕上,但你实际上可视化了,并且对为什么使用潜在空间而不是词元更高效有某种分析性的解释。

And in this recent paper that you just released about a month ago, the predict latents not tokens, you should explain figure one. We'll put it on the screen now, but you actually visualize and have a kind of analytical explanation for why it's more efficient using latents and not tokens.

Matthieu

完全正确。所以是的,作为物理学家,我们喜欢做的是拥有可以改变参数的模型,然后我们做出 Scaling 预测,然后很容易测试你的预测。你把它画在双对数坐标上,然后你就能看到。所以我们喜欢有那种参数来测试我们的想法。但就概念图而言,这个图有三个网络。第一个是监督学习。所以在我们的模型中,数据是树状的。有一个顶层根。也许如果你考虑图像,这可能是说如果你的图像是猫或狗或其他什么,你看不到任何东西,那些是隐藏的纤维,然后你只看到数据,即输入。所以一个问题是你需要多少数据才能从输入中对树的根进行分类?那是监督学习。然后也许我会因为时间原因跳过那个。我的意思是,中心图更像是扩散模型或下一个词预测,而这个图展示的正是我试图告诉你的概念。这是街道的概念,在街道下面有行人、汽车、人行道。所以下面的节点把它们想象成行人,上面的节点是街道的概念。而在那些模型中真正重要的是你如何将这个概念与数据的非常低层次方面相关联。而在那些模型中相当清楚的是,当你沿着这棵树走远时,相关性会下降,因为你每次必须做出几个选择,这导致相关性下降。

Exactly. So yeah, what we like to do as physicists is also to have models where we can vary parameters and then we make scaling predictions, and then it's very easy to test that to test your prediction. You plot it in a log-log and you see. So we like to have those kind of parameters to test our ideas. But in terms of the conceptual picture, this figure has, I think, three networks. The first is supervised learning. So in our models, the data are tree-like. There's a top root. Maybe if you think about images, this is maybe saying if your image is a cat or a dog or whatever, and you don't see anything, those are hidden fibers, and then you see just what's the data, which is the input. So one question would be how many data do you need to from the input be able to classify the root of your tree? That's supervised learning. And then maybe I will skip that for reasons of time. I mean, the central figure is more like diffusion models or next token prediction, and what this figure is showing is really the concept I was trying to tell you. It's this concept of streets, and below streets you would have passersby, cars, sidewalks. So the nodes below think of them as passersby, and the node above is a concept of street. And really what matters in those models is how do you correlate this concept with very low-level aspects of your data. And what's in those models something that's quite clear is that as you go away along this tree, the correlation decreases because you have to make several choices every time, and that leads to decreasing correlation.

层级学习与样本复杂度 Hierarchical Learning and Sample Complexity

Matthieu

所以中间的面板展示的是,当你试图接近顶层路线,也就是试图构建抽象概念时,你是在与非常低层的东西相关联,这在树上是一段很长的距离。我们知道,每沿着这棵树移动一次,我们就得在所需的数据量上付出一个乘法代价。所以这就是为什么学习隐藏层级所需的数据量是树深度的指数级。这仍然不错,因为你想,输入的维度也是树深度的指数级。这意味着你可以用问题维度的多项式级来学习,这比指数级好得多,指数级就意味着不可能。所以下一个词预测是有效的,但它仍然是树深度的指数级。现在看最后一个面板,你真正看到的是你可以做一些非常不同的事情。同样,当你试图构建“街道”这个概念时,你只需要预测附近有什么,比如房子,而你已经理解了“房子”这个概念。现在在你的图上,相关性更近了,所以相关性更强。相关性越强,你总是有信号噪声比。你知道,你需要足够的数据来测量相关性,但如果信号强,你测量它所需的数据就少得多。一旦你测量到它,砰,你就能构建这些抽象。所以这基本上就是这个图展示的。它总结了我们在街道和房子这个例子中讨论的内容。

So what the middle panel would show is that when you try to approach the top route, so when you're trying to build an abstract concept, you're correlating with very low level, it's a long distance along the tree. And so we know that every time we move along this tree, we have to pay a multiplicative cost in the number of data we need. And so that's why the number of data you need to learn your hidden hierarchy is exponential in the depth of the tree. Which is still good because if you think about it, the dimension of the input is also exponential in the depth of the tree. It means you can learn that polynomially in the dimension of your problem, which is much better than exponential, which means impossible. So next-token prediction works, but it's still exponential in the depth of the tree. Now if you think about the last panel, what you would really see is that you can do something very different. Again, when you try to build the concept of street, you will just predict what's nearby, you know, houses, and you've already understood this concept of house. And now the correlation is much closer on your graph, so it's much more correlated. The correlation being much stronger, you always have a signal to noise. You know, you need enough data to measure correlations, but if the signal is strong, you need much less data to measure it. And once you measure it, boom, you can build those abstractions. So this is essentially what this figure shows. It's a summary of what we've been discussing in this example of street and houses.

Host

令人沮丧的是,我们知道很多能推进前沿的方法,但 OpenAI 和 Anthropic 还在训练老式的 Transformer。我和 Sakana 的 Clément Jones 谈过这个,他是 Transformer 的发明者之一。他说任何新方法都必须好得惊人,因为我们在硬件、优化器和编译器上投入了太多时间。这有一个完整的生态系统,实际上很难转向。Llama 确实有几家新创业公司,但我得到的印象是他专注于垂直领域。我们还没有做过大规模尝试这些新模型的登月计划。

It's so frustrating that we know so many things that could advance the frontier, but OpenAI and Anthropic, they're still training old-school transformers. And I spoke to Clément Jones at Sakana about this, who was one of the inventors of the transformer. And he said any new method has to be crushingly better because we've invested so much time in hardware and optimizers and compilers. There's an entire ecosystem about this, and it's actually very difficult to steer the ship. Llama does have a couple of new startups, but the read I'm getting is that he's focusing on vertical domains. We haven't yet done the moonshot where we try these new models on mass.

Matthieu

是的,我同意。但我还是要提醒一句,这些大语言模型是生成模型,这一点很重要,因为你可以与它们互动,它们会产生推理等等。而且我仍然不知道,即使在理论上,即使我理解了世界上隐藏的所有层级,我能否有效地利用它回到词元级别的预测。但如果我们可以,那至少意味着在概念上我们可以做一个更好的生成模型。但你看,这很微妙,但理解世界结构(就像构建一个编码器)和将其解码为数据非常低层的方面是有区别的。

Yes. So I agree. I will still say a word of caution that those LLMs are generative models. That's very important for them to be, because you can interact with them and they produce reasoning and so on. And I still don't know even theoretically if, even though I understood all this hierarchy hidden in the world, I can use it efficiently to go back to a prediction at a token level. But if we could, then it would mean at least conceptually that we could do a much better generative model. But you see, it's subtle, but there's a distinction between understanding the structure of the world, which is like building an encoder, and then decoding it for a very low-level aspect of the data.

Host

你提到了扩散模型,你在这方面有一篇很棒的文章。但只是概念上,你认为它们与 Transformer 有什么不同?

You mentioned diffusion models and you had a great paper out about that. But just conceptually, how do you think they are different from something like a transformer?

Matthieu

嗯,通常它们基于 Transformer 架构。所以更多的是指目标。一种情况是掩蔽未来,比如预测下一个词元,而扩散模型实际上是随机掩蔽随机位置。我认为这非常相似,在我们的理论中,两者的样本复杂度相同。所以并不是一个比另一个有巨大优势。从某种意义上说,唯一的区别是你填充被掩蔽内容的顺序。

Well, some often they are based on transformer architecture. So it's more the objective you mean. In one case you are masking the future, like you're predicting the next token, whereas in diffusion models you're actually masking randomly at random positions. I think it's very similar, and in our theory it's the same sample complexity for both. So it's not like one has a huge advantage over the other. In some sense, the only difference is the order in which you're filling up what is being masked.

Host

哦,这很有趣。我的意思是,直觉上我认为你首先有任意数量的扩散步骤,也许你会说这类似于在正常网络上训练时做更多的反向传播。而且扩散模型似乎先学习全局关系,然后收紧,而 Transformer 和 CNN 似乎有局部性偏差,这可能会有一些不同。

Oh that's interesting. I mean because intuitively I think of it as you first of all you have an arbitrary number of diffusion steps and maybe you would say that's analogous to just doing more back passes during training on a normal network. And there might be some difference in terms of the diffusion model seems to learn global relationships first and then tightens up, whereas transformers and CNN seem to have a locality bias.

Matthieu

我喜欢通过关注样本复杂度来简化讨论,也就是你需要多少数据来学习。这就是我们找到类似量的地方,因为你进入了这种前向和后向过程的计算。所以我真正谈论的是样本复杂度。然后我们在两种情况下发现,随着数据量的增加,你从下到上学习这些约束、这些语法规则。所以先是低层,然后是高层,两种情况都是。我们实际上对这些说法有信心,因为它们做出了非平凡的预测。例如,你可以预测,随着你训练扩散模型生成文本越来越多,最初它会是随机的。如果你没有数据,它生成的是垃圾。随着数据量增加,它应该开始形成连贯的单词。然后随着更多数据,形成连贯的词组,然后是连贯的完整句子。这是这些模型的预测,即上下文的连贯性应该随着数据量的增加而稳步提高,我们实际上可以检查扩散模型,也可以检查下一个词预测的缩放定律理论。

And I like to simplify the discussion by focusing on sample complexity, how many data do you need to learn that. That's where we find analogous quantities, because you are going into sort of compute of going back with this forward and backward process. So I'm really talking about sample complexity. And then what we find in both cases is that as you increase the number of data, you learn those constraints, those grammatical rules, bottom up. So first low level, then higher level in both cases. And we could actually have some confidence on those statements also because they make non-trivial predictions. So for example, you would predict that as you train a diffusion model more and more to generate text, initially it would be random. If you don't have data, it's regenerating crap. As you increase the number of data, it should start to form coherent words. Then later on with more data, coherent group of words, and then coherent full sentences. And this is a prediction of those models that the coherence of the context should steadily increase as you increase the number of data, that we could actually check for diffusion models and also check the theory of scaling laws of next-token prediction.

Host

大约一年前,你也有一篇关于缩放定律的论文。

About a year ago you had a paper about scaling laws as well.

Matthieu

是的。

Yes.

Host

给我讲讲。

Tell me about that.

Matthieu

随着数据量的增加,或者你投入的算力增加,或者参数数量增加,你的性能会稳步提高。Kaplan 等人的这一观察对我们所有人产生了巨大影响,因为它促使科技公司投入更多,甚至可能建造核电站。所以它产生了巨大的技术影响。但对我们理论家来说,有点尴尬的是,这基本上完全不被理解。你知道,这些缩放定律在定量上有指数。例如,描述如果你将数据量乘以 10,你的表现会好多少。而在这个问题上,理解非常有限。

As you increase the number of data or you increase how much compute you put or you increase the number of parameters, your performance steadily improves. And this observation by Kaplan and others had a huge impact for all of us, because it drove the tech companies to just invest more and build, you know, maybe nuclear plants. So it has a huge technological impact. But it's a bit embarrassing for us theorists that essentially it's not understood at all. You know, those scaling laws quantitatively, they have exponents in them. For example, describing how well you perform better if you multiply the number of data by 10. And there was very limited understanding on that question.

自然语言缩放定律理论 Theory for scaling laws in natural language

Matthieu

几个月前,我和 Franchesco Cagneta、Alan Ravventos 以及 Soya Ganguli 一起,针对这个问题提出了一个理论。这个理论受到我之前提到的那些合成世界的启发,但剥离了我们从那些模型中学到的核心经验,以便对自然语言做出定量预测。本质上,这个理论预测存在一个简单的配方来提取这些指数。这再次说明,其根本在于:如果你给我更多数据,我就能学习更抽象的概念,而更长的范围会导致更长的相关性。但归根结底,你需要测量的两个量是:第一,词或词元是相关的,而且这种相关性(这一点之前就众所周知)随着两个词之间距离的幂次而衰减。由此你可以测量出指数,它们取决于你作为数据集的语言。你可以测量它们。然后我们论证了另一个关键量,它与文本的熵有关。文本的熵在 50 年代就已经被香农讨论过。这是一个美妙的问题。本质上,熵是平均在一个位置上可能出现的词数的对数。我们认为非常重要、并且我们最终能用大语言模型或其他架构测量到、且结果一致的是:在看到一个包含 n 个词元的句子后,剩余的熵是多少。如果你看到 n 个词元,看到的词元越多,可能性就越少。那么熵是多少?在玩具模型中,它是幂律,在现实生活中,也发现了幂律。所以本质上,我们论证的是,用这两个指数,你可以按照我们指定的方式组合它们,得到大语言模型在自然语言上的训练曲线指数,而且效果非常好。我们对此非常兴奋。另外,我还得说,它还做出了非平凡的预测,即损失应该如何依赖于你给出的上下文,以及数据的数量。所以它是两个变量的函数,我们预测它应该以非常特定的方式弯曲,我们能够测试它,并且我们也观察到了这一点。

So a few months back with Franchesco Cagneta, Alan Ravventos, and Soya Ganguli, we proposed a theory for this problem, inspired by those synthetic worlds I told you about, but detaching the essence of the lesson we learned from those models to make quantitative predictions for natural languages. Essentially, the theory predicts that there is a simple recipe to extract those exponents. And essentially, this is saying again that what's underlying it is the fact that if you give me more data, I can learn more abstract concepts, and that longer range leads to longer range correlations. But at the end of the day, the two quantities you need to measure are: one, the fact that words or tokens are correlated, and that this correlation, it was well known before, decreases as a power of the distance between those two words. From that, you can measure exponents, and they depend on the language you look at as your dataset. You can measure them. Then there is another key quantity we argue, which is related to the entropy of text. So entropy of text has been discussed already by Shannon in the 50s. It's a beautiful question. Essentially, it's the log of the number of possible words that you would have at one location on average. What we argue is very important to look at, and that we could finally measure with LLMs or other architectures, and we find consistent results, is what is the entropy left after a sentence of n tokens. If you see n tokens, the more tokens you see, the fewer possibilities you have. What is the entropy of that? And in the toy model, it's a power law, and in real life, it's also a power law that is found. So essentially, what we argue is that with those two exponents, you can combine them in a way that we specify to get the training curve exponent of LLMs acting on those natural languages, and it works very well. So we got very excited with that. Also, I have to say that in addition, it's making a non-trivial prediction in terms of how the loss should depend on the context that you give it, and also the number of data. So it's a function of two variables, and we predict that it should bend in a very specific way, and we could test it, and we also observed it.

Host

你能给我一些直觉上的理解吗?我的意思是,当我们谈到香农时,人们经常看到一张图,在一个句子中,每个词都会降低熵。而现在我们几乎是在总体规模上讨论,当我们有海量文本语料库时,熵只会不断下降。这意味着什么?这是否意味着问题会随着时间变得更容易?这是否意味着模型会继续变得更好,或者可能会有某种相变?会发生什么?

And can you give me some intuition on that? I mean, it's often, you know, when we speak about Shannon, there's a graph that people often see that during a sentence every single word reduces entropy. And we're now talking almost at the population scale, that when we have a huge corpus of text, entropy is just going down and down. What does that mean? Does that mean that the problem is getting easier over time? Does it mean that the models will just continue to get better, or maybe there'll be some phase change? What's going to happen?

Matthieu

首先,我必须非常谨慎,因为我们能在学术范围内测试我们的理论,也就是 10 亿参数、10 亿词元。好吗?我告诉过你,随着词元数量的增加,这个机器开始使用越来越大的上下文。我们可以可视化这个上下文。所以我可以告诉你我们的理论在什么上下文规模下得到了测试,大约是 50 个词元,或者两三个句子。我认为这很好,因为那里有所有的句法等等。所以有很多东西。但这就是我确信的地方,我的意思是,我认为我们有一个非常稳健的故事,我认为它会成立。我的意思是,这个领域必须进一步研究并做出决定。但我们真正没有做的,因为以我们的手段不可能做到,是测试这个理论是否适用于远远超过三四个句子的情况。所以我不知道我们提出的机制在那里是否仍然适用,或者是否是完全不同的东西。

So, I should be extremely careful first of all, because we could test our theory at the academic range, so it means 1 billion parameters, 1 billion tokens. Okay? And I told you that as you increase the number of tokens, this machine starts to use a context that's larger and larger. And we can visualize this context. So I can tell you for which context scale our theory was tested, and it's about 50 tokens, or two or three sentences. I think it's great because there's all syntax and so on there. So there's a lot of stuff. But that's where I'm confident that, I mean, I think we have a very robust story that I think will hold true. I mean, the field has to investigate further and decide. But what we really have not done, because it's not possible with our means, is to test this theory for much beyond three or four sentences. And so I don't know if the mechanism we put forward still applies there, or if it's something completely different.

向科学家解释深度学习 Explaining deep learning to a scientist

Host

如果你能把深度学习解释给任何一位科学家,无论健在还是已故,你会选谁?

If you could explain deep learning to any scientist dead or alive, who would it be?

Matthieu

哦,这个问题我从未想过。我和我父亲关系非常紧密,他是一位物理学家,我们经常讨论科学。在他生命的最后阶段,他实际上对神经科学非常感兴趣。他思考各种派别之类的东西。所以我会选他。是的。

Oh, that's a question I never thought about. I had a very strong bond with my dad, who was a physicist, and with whom we discussed a lot about science. And at the end of his life, he was actually very interested in neuroscience. He thought about all factions and things like that. So it would be him. Yes.

科学中的错误 Being wrong in science

Host

在你的职业生涯中,有没有什么事情是你完全搞错了,然后改变了想法的?

Are there any things in your career that you've been completely wrong about and you've changed your mind?

Matthieu

我认为在科学中犯错误是完全正常的,但我认为非常重要的一点是,一旦你确信自己犯了错误,你就应该说出来,让所有人都清楚,即使你不再相信它,你也不应该坚持不放。我的意思是,所以是的,这肯定会发生。你知道,我们作为物理学家玩的游戏是提出世界模型,然后做出预测。我们已经觉得这就是我们的工作。然后它需要由我们或其他人来测试。有时你的预测不成立,因为那不是好的模型,但这就是我们建立理解、假设等等的方式。所以从这个意义上说,是的,这经常发生,但应该如此。我的意思是,在某种意义上,科学就应该这样运作。你提出假设,然后真正去测试它们。所以是的,我认为重要的是,如果你有创造力,如果你冒险,你就应该犯错误。如果你从不犯错误,也许这标志着你在科学上有点走在老路上。而我们中的一些人想要探索丛林。在丛林中,你可能会犯错。我的意思是,是的。但是……

I think it's completely fine to make mistakes in science, but I think what's very important is that once you're convinced that you made a mistake, you state it, and it's obvious for everybody, and you don't sort of encroach on it, even if you don't believe in it anymore. I mean, so yes, it happens for sure. You know, the game we're playing as physicists is to propose models of the world and then make predictions. Already that, we feel, is doing our job. Then it needs to be tested by us or by others. Sometimes your prediction does not hold because it was not the good model, but that's how we build understanding, hypotheses, and so on. So in this sense, yes, it often happens, but it should. I mean, that's how science should work in some sense. You're making hypotheses, and then you really test them. So yeah, I think it's important that if you're creative, if you take risks, you should make mistakes. If you never make mistakes, maybe it's a sign that you're staying a bit on the beaten path in science. And some of us want to explore the jungle. In the jungle, you can be wrong. I mean, yes. But...

结束语 Closing remarks

Host

非常荣幸能邀请你上节目。非常感谢你的到来。

It's been an honor having you on the show. Thank you so much for joining us.

Matthieu

非常感谢。很有趣。谢谢。

Thank you so much. It was fun. Thanks.

互动版:逐字朗读 + 针对本期提问 →