加速理解:物理领域的基础模型

Accelerated Understanding: A Foundation Model for Physics

阿尼玛·阿南德库马尔 Anima Anandkumar · Latent Space · 2026-09-04 · 约 27 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anima Anandkumar 和 Benedict 讨论他们的新公司 Accelerated Understanding,旨在构建跨行业的统一物理模拟 AI 模型。

Anima Anandkumar and Benedict discuss their new company Accelerated Understanding, aiming to build a unified AI model for physical simulations across industries.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 16)

全文 · Full transcript(中英对照)

开场 Introduction

Host

大家好!我是 RJ。今天和我搭档的是 Brandon,这里是 Latent Space 播客,聊的是科学领域的 Lightning AI。一周前,我们和 Anima Anandkumar 一起发布了一期非常棒的节目。反响很好,大概是因为就在我们发布前几天,她把自己的公司 Accelerated Understanding 从幕后带到了台前。但一个月前我们录制那次对话时,她还不能谈论自己的公司。所以现在 Anima 和她的联合创始人 Benedict 一起回到播客,来聊聊 Accelerated Understanding,以及上次录制时她必须保密的内容。顺便说个小细节:《时代》杂志刚刚把 Anima 评为年度百大最具影响力人物之一,我们也想聊聊这个。但首先,Benedict,恭喜。Anima,欢迎回来。请给我们讲讲 Accelerated Understanding 吧。

Hello everyone! I am RJ. I'm here with my co-host Brandon on the Latent Space podcast about Lightning AI for science. A week ago we released a fantastic episode with Anima Anandkumar. It received a great response, probably due to the fact that she brought her company Accelerated Understanding out of the shadows just days before our release. But when we recorded the conversation a month ago, she couldn't talk about her company. So now Anima is back on the podcast with her co-founder Benedict to talk about Accelerated Understanding and what she had to keep secret during our previous recording. By the way, a small detail: Time magazine just named Anima one of the 100 most influential people of the year, and we want to talk about that too. But first, Benedict, congratulations. Anima, welcome back. Please tell us about Accelerated Understanding.

加速理解愿景 The Vision Behind Accelerated Understanding

Anima

所以,本质上,我们在反思从 Anima 所有成功项目中观察到的现象:导管、天气模型,以及她上次播客里谈到的所有东西。我们意识到语言技术领域发生了什么:如果你回头看,在最近这次 AI 革命之前的“旧时代”,人们会针对特定任务训练模型。比如,我们有拼写检查模型、翻译模型、自动补全模型。然后 OpenAI 出现了,说:我们把所有东西合并到一个模型里吧。结果证明这效果惊人,几乎超越了所有专用模型。所以现在我们问自己:我们能不能对物理建模和理解物理过程做同样的事情?

So, essentially, we were reflecting on what we had observed from all of Anima's successful projects: catheters, weather models, all the things she talked about in the last podcast. And we realized what happened in language technology: if you look back, to those "old days" before the recent AI revolution, people trained models for specific tasks. For example, we had spell-checking models, translation models, auto-completion models. And then OpenAI came along and said: let's just combine everything into one model. It turned out that it works phenomenally well and outperforms almost all specialized models. So now we ask ourselves: can we do the same for physical modeling and understanding physical processes?

Host

是的。你知道,我们在语言模型中看到的这种通用性和可扩展性——它们在物理世界中的对应物是什么?这正是 Accelerated Understanding 押注的方向。

Yes. You know, that versatility and scalability that we've seen in language models—what's their counterpart for the physical world? This is what Accelerated Understanding is betting on.

Host

我想到的第一个问题是:所以你们实际上是想把来自不同领域的一堆数据应用到一个模型上,创造一种“上帝版本”,一个物理版本,或者说物理世界的版本。我担心的问题是:虽然语言有很多形式,但它仍然是语言,但对于天气模拟、导管和核反应堆来说,物理可能非常不同,所以那里可能没有太多知识迁移。那么,有什么证据,或者你为什么认为这仍然是一个适合迁移的任务?

The first question that comes to mind is: so you're actually trying to apply a bunch of data from different domains to one model, to create a kind of "version of God," a physical version, or a version of the physical world. The question that concerns me is: although language has many forms, it is still language, but for weather simulations, catheters, and nuclear reactors, the physics may be very different, and so there may not be much transfer of knowledge there. So what is the evidence or why do you think this is still a task that is amenable to transfer?

Anima

是的,嗯,首先,当然,在物理世界方面会有更多的挑战,对吧?没有现成的互联网包含所有这些类型的物理数据。所以,根本没有真实数据。此外,当涉及科学发现时,它总是新的东西。因此,根据定义,这不可能在训练数据中。因此,仅仅依赖纯数据驱动的 AI 是不够的。而正是在这里,你知道,加入物理定律真的很有帮助。考虑一下,我们可以用现有的模拟器来训练我们的模型,但真正的“亮点”将是利用物理定律本身进行自我改进。所以物理将是我们整个过程中的主要参考点。这也是不同行业的一个共同特征。无论是导管中的流体还是火箭,流体动力学的原理都是一样的。这些是不同的状态,由雷诺数和类似参数决定,但它们有共同点。我们看到它在不同领域如何运作:能源、半导体、航空航天——它们看起来非常不同,但它们有共同的基本原理。即使我们说不同的系统中有不同的数学方程,比如偏微分方程,它们之间仍然有共同点,比如随时间的变化。这可能是能量守恒定律、物体恒常性等方面,或者因果关系。所以即使在由不同方程描述的物理中,也有许多隐含的共性。这就是为什么在不同行业和数学模型中有许多共同特征,神经网络可以识别这些特征。

Yeah, well, first of all, of course, there are a lot more challenges when it comes to the physical world, right? There is no ready-made internet with all these types of physical data. So, there is simply no real data. Besides, when it comes to scientific discoveries, it's always something new. Therefore, by definition, this cannot be in the training data. Therefore, relying solely on purely data-driven AI will not be enough. And this is where, you know, adding the laws of physics really helps. Consider that we can use existing simulators to train our model, but the real "highlight" will be self-improvement using the laws of physics themselves. So physics will be our main reference point throughout the process. And this is also a common feature for different industries. Whether it's a fluid in a catheter or a rocket, the principles of hydrodynamics are the same. These are different regimes, determined by Reynolds numbers and similar parameters, but they have something in common. And we see how it works in different areas: energy, semiconductors, aerospace - they look very different, but they have common fundamental principles. Even if we say that different mathematical equations exist in different systems, such as partial differential equations, there is still something in common between them, such as change over time. This could be the law of conservation of energy, aspects such as the constancy of objects, or cause-and-effect relationships. So there are many implicit commonalities even in physics, which is described by different equations. That is why there are many common characteristics across different industries and mathematical models that neural networks can recognize.

Host

当然。那么,最有趣的时刻是什么?我们现在在创建这个模型方面处于什么阶段?你们进展了多远?

Of course. So, what are the most interesting moments? What stage are we at now in creating this model? How far have you progressed?

Anima

我们已经训练她一年多了。我们进行了所有必要的架构实验。有趣的是,我们必须重新发明很多东西,因为我们的数据结构完全不同。比如,如果你看语言,有在一维上增长的 token。先进模型可以容纳数百万个。对我们来说,数据在多维上增长,特别之处在于我们不想妥协。所以我们决定,“让我们按照真实世界的方式来做一切。”所以我们的空间在三维中扩展。它不像视频模型那样被压缩。它实际上保持在这三个维度中。然后你还有时间上的展开。而在这里,人们又选择了自回归的简单路径,也就是一步一步预测,但这样误差会累积,你会失去连续性和对输入数据中到底是什么导致了这个结果的理解,这对设计来说极其重要。因此,即使对于时间维度,我们也倾向于模拟完整的展开,我们实际上就是这么做的。但然后你会发现自己处于一个突然有四个维度的情况。这四个维度都独立增长。当你把这些数字相乘,你会很快达到数十亿甚至数万亿的上下文。这就是我们模型实际达到的。我们可以训练多达一万亿的输入上下文。我们能够训练一万亿上下文长度的输入和输出数据。我们能够用 5 万亿的上下文进行推理。与语言模型相比,这些只是毫无意义的数字。而且我们甚至没有使用这些方法中的大多数。如果你看视频模型,它们使用很多技巧来提高效率。例如,它们平均像素。它们把像素分成 patch。我们不这样做。我们仍然能够处理这个令人难以置信的上下文。但现在让我们回到我们做了什么?我们如何实现这一点?把所有这些东西放进去是相当困难的。教所有这些是相当困难的。你可能听说过像 FSDP 这样的标准技术,你把数据分成段,然后在 GPU 中组装一层进行训练。但我们的层太大了,你无法在 GPU 内部组装它们。我们的数据样本太大了,无法放在一个加速器甚至一个完整节点上。所以我们不得不重新发明整个分片基础设施,整个分段策略,来大规模训练这些模型。我们做到了。我们进行了数百次训练启动。我们训练了多达一万亿参数的模型。所以我们真的证明了它是有效的。我们确保的另一件事是我们从多物理中受益。我们不仅仅是在扩展以前小规模做的事情,我们实际上是在把这些东西结合起来。

We have been training her for a little over a year. We conducted all the necessary architectural experiments. The interesting thing is that we have to reinvent a lot of things because the structure of our data is completely different. If you look at language, for example, there are tokens that grow in one dimension. Advanced models can accommodate millions of them. For us, data is growing in many dimensions, and the special thing is that we don't want to compromise. So we decided, "Let's do everything the way it is in the real world." So we have space expanding in three dimensions. It is not compressed like in video models. It actually remains in these three dimensions. And then you still have the unfolding in time. And here again, people choose the easy path of autoregression, that is, predicting one step at a time, but then errors accumulate, and you lose continuity and understanding of what exactly in the input data led to the result, which is extremely important for design. Therefore, even for the time dimension, we prefer to simulate full deployment, which is what we actually do. But then you find yourself in a situation where there are suddenly four dimensions. And all four of these dimensions grow independently. And when you multiply these numbers, you very quickly get to billions or even trillions in context. And that's what we actually achieved with our models. We can train with up to a trillion input contexts. We are capable of training with input and output data that are a trillion contexts long. We are able to perform inference with a context of 5 trillion. These are just meaningless numbers when compared to language models. And we don't even use most of these methods. If you look at video models, they use a lot of tricks for greater efficiency. They, for example, average pixels. They break them into patches. We don't do that. We are still able to work with this incredible context. But now let's get back to what we did? How do we implement this? It's quite difficult to place all this. It's quite difficult to teach all this. You may have heard of standard techniques like FSDP, where you break the data into segments and then assemble a layer in the GPU for training. But our layers are so large that you can't assemble them inside a GPU. Our data samples are so large that they don't fit on an accelerator or even a full node. So we had to reinvent the entire sharding infrastructure, the entire segmentation strategy, to train these models at scale. And we did it. We conducted hundreds of training launches. We trained models with up to a trillion parameters. So we really proved that it works. Another thing we've made sure of is that we benefit from multiphysics. We're not just scaling up what was previously done on a small scale, we're actually combining these things together.

物理模型协作学习 Collaborative Learning in Physics Models

Anima

所以我们特意挑选了一批我们认为在特性和挑战上非常多样化的物理分支,把它们放进同一个模型里,对它们进行了完整的 4D 部署,并且我们能够教会它们。事实上,我想补充一点:一个同样规模的模型,如果覆盖多个物理分支,效果会比把这些参数单独隔离给每个分支更好。所以,即使你为每个分支单独制作同样大的模型,结果仍然会更差。这意味着系统受益于协作学习。这不仅仅是说我们扩大模型规模,它就会变得更好——这确实是真的——但添加新的物理分支能帮助它更好地掌握每一个分支。这和我们在语言模型中观察到的涌现式学习是同一类型。

So we deliberately chose a number of branches of physics that we think are very diverse in their characteristics and challenges, put them in one model, did a full 4D deployment for them, and we're able to teach them. In fact, I wanted to add that a model of the same size that covers several branches of physics works better than isolating all these parameters for each individual branch. So, if you had separate models made as large as the original, the result would still be worse. This means that the system benefits from collaborative learning. It's not just that we scale up the model and it gets better. This is true, but adding new branches of physics helps her better master each of them. This is the same type of emergent learning that we observe in language models.

Host

在你们的那期节目里,我们聊了很多关于神经算子的内容,特别是你们放入模型中的某些归纳偏置,比如傅里叶神经算子——你进入傅里叶空间,在球面上工作,并且对几何有归纳偏置。当你们试图处理许多不同的几何时,迁移学习是如何工作的?你们目前是只局限于矩形几何,还是有某些技术能让你们整合不同类别的偏微分方程,也许带有不同的时间导数之类的?

In your issue, we talked a lot about neural operators, particularly certain inductive biases that you put into the model, such as Fourier neural operators, where you go into Fourier space, work on a sphere, and have an inductive bias with respect to the geometry. How does transfer learning work when you are trying to work with many different geometries? Are you currently limiting yourself to just rectangular geometries, or are there any techniques that allow you to integrate different classes of partial differential equations, perhaps with different time derivatives or something like that?

神经算子与分辨率 Neural Operators and Resolution Flexibility

Anima

我们不能透露我们架构的所有细节,因为那是知识产权,但我可以说神经算子是基础,因为这是我们确保分辨率不变性的方式。Benedict 提到推理时上下文长度为 5 万亿,我们实现了这一点,训练时是 1 万亿。但这意味着我们每次都会这样做吗?不,不是每个应用都需要这么高的分辨率。

We can't reveal all the details of our architecture as it is intellectual property, but I can say that neural operators are the foundation, as this is how we ensure resolution invariance. Benedict talked about a context length of 5 trillion during inference, which we achieved, and a trillion during training. But does that mean we will do it every time? No, not every application requires such high resolution.

Host

不,对吧?

No, right?

Anima

这适用于最复杂的物理任务,其中每个细节都至关重要。我们可以提供很长的上下文。而对于那些不需要这么高分辨率的任务,或者在设计研究阶段、任务刚开始时,你不需要每个细节。你不需要花费那么多计算资源。因此,这种灵活性至关重要。这立刻让我们区别于其他所谓的世界模型——视频模型或视觉模型——因为它们都在训练和输出时假设固定的分辨率。这对我们的视觉特性来说是可以接受的,对吧?因为我们不一定需要扩大规模,尤其是在游戏或娱乐中。在那里我们承认物理是近似处理的。但在工程设计或科学发现中,我们不能承受这样的简化。因此,神经算子是关键,它提供了处理不同上下文长度(相当于不同分辨率)时的灵活性。

This applies to the most complex physical tasks, where every detail really matters. We can provide a great length of context. And for those where it's not necessary, or at the design research stage, at the beginning of the tasks, you don't need every detail. You don't need to spend so much computing resources. Therefore, such flexibility is critically important. This immediately sets us apart from other so-called world models—video models or vision models—because they all assume a fixed resolution during training and output. And that's acceptable for our visual characteristics, right? Because we don't necessarily need to scale up, especially in games or entertainment. There we will admit that the physics is somewhat approximate. But in engineering design or scientific discovery, we cannot afford such simplifications. Therefore, neural operators are the key to providing flexibility when dealing with different context lengths, which is equivalent to different resolutions.

Anima

另一方面,如果你想想那些在语言上表现如此出色的 Transformer 架构,它们根本无法支持 5 万亿的上下文长度,无论你投入多少算力。所以,这种二次复杂度是不可行的,也是不必要的,因为物理世界比“万物与万物”的任意关联更有结构。我们必须重新思考什么最适合物理世界,而我们已经以一种非常系统的方式做到了这一点,类似于先进的语言实验室进行大量系统实验和架构演进,从早期模型到专家混合等等。我们在加速理解方面已经经历并继续经历同样的演进,以确保最佳的架构和最佳的硬件利用,考虑到通信带宽和所有其他要求。因此,这种协同设计至关重要。

On the other hand, if you think about using transformer architectures that work so well with language, they simply wouldn't be able to support a context length of 5 trillion, no matter how much computing power you put in. So, such quadratic complexity is unfeasible, as well as unnecessary, since the physical world has more structure than just the arbitrary correlation of "everything with everything." We have to rethink what works best for the physical world, and that's what we've done in a very systematic way, similar to how advanced language labs do a lot of system experimentation and architecture evolution, from early models to a mix of experts, and so on. We have gone through and continue to go through the same evolution in our accelerated understanding to ensure the best architecture and best use of hardware, considering communication bandwidth and all other requirements. Therefore, such collaborative design is critically important.

推荐原集 Recommendation to Watch Original Episode

Host

快速说明:对于那些还没看过原集的人,我们建议你们去看。如果想了解神经算子,我认为你们应该从大约第 20 分钟开始看。我们会加一个直接链接,这样你们可以直接跳到我们开始讨论其他内容的部分。不过,是的,我想确保我们快速提到了这一点。

Quick note: for those who haven't watched the original episode, we recommend you do so. If you want to learn about neural operators, I think you should start around the 20th minute. We'll add a direct link so you can jump straight to the part where we start discussing other things. But yeah, I wanted to make sure we mentioned that quickly.

万亿级上下文 Achieving Trillion-Token Contexts

Host

嗯。那么,可能对你们每个人来说——我知道很多可能是专有信息,但你们能解释一下如何达到 1 万亿或 5 万亿 token 的直觉吗?以及你们如何从算法和基础设施两方面解决这个问题?

Ahem. So, probably to each of you—I know a lot of this is probably proprietary information, but can you explain the intuition of how you get to a trillion or five trillion tokens? And how do you, let's say, solve this issue from both the algorithmic side and the infrastructure side?

Anima

是的。如何得到这么大的数字其实很简单。如果你说每个空间维度有千分之一的分辨率,并且你想精确解析一千个时间步,那么你得到的就是 1 万亿。这是一个非常巨大的数字,如果你稍微想想数学,这个数据量有多大?例如,对于我们 5 万亿的运行,结果是 22 TB,而你希望看到这样的数据出现在加速器的内存中。就是这样,否则一切都会运行得很慢。这里有个思考方式:显然,我们可以处理任何规模,因为我们已经在大型集群上得出了结论。我们一直在做推理,例如,我们可以在 MacBook 或 Mac Studio 上运行上下文较短的较小模型,甚至中等偏大的模型。所以你有整个范围。因此,要实现这一点,你需要开发一系列扩展技巧,但主要的是——你需要一个工作数据集,它必须放在某个地方,而这决定了所有事情。再说一次,我不能具体说你们在做什么,但最终可以安全地假设,你想通过模型的带宽以某种方式让数据片段相互作用。这意味着在很大程度上,你还必须依赖具有非常可靠和高质量连接的大型集群。所以我们已经经历了很多 LLM 日益面临的挑战,现代方法和老式 HPC 方法相结合——我们必须弄清楚如何容纳一个不适合单个节点的样本。我们必须弄清楚如何在一个层内实现交互,在那里我们可以将状态存储在加速器中。所以所有这些技巧都需要大量的扩展和大量的基础设施工作,而我们做到了。

Yes. How you get such large numbers is quite simple. If you say you have a thousandth resolution in each spatial dimension, and you want to accurately resolve a thousand time steps, then you get a trillion. This is a really huge number, and if you think about the math a little, how big is this amount of data? For example, for our 5 trillion run, the results were 22 terabytes, and you want to see things like that in the accelerator's memory. That's how it is, otherwise everything will work slowly. Here's how to think about it; Obviously, we can work with any size, as we have already concluded on huge clusters. We've been doing inference, for example, we can run small models with less context or even medium-large models with less context on a MacBook or Mac Studio. So you have this whole range. So, to implement this, you need to develop a bunch of scaling tricks, but the main thing is—you need a working dataset that has to fit somewhere, and that's what determines everything. And again, I can't speak to what exactly you're doing, but ultimately it's safe to assume that you want to interact pieces of data with each other in some way through the bandwidth of the model. And this means that to a large extent you will also have to rely on large clusters with very reliable and high-quality connections. So we've already gone through a lot of the challenges that LLMs are increasingly facing, where modern approaches and good old HPC methods are combined—we had to figure out how to accommodate a sample that doesn't fit into a node. We had to figure out how to enable interaction within a layer where we could store state in the accelerator. So all these tricks require a lot of scaling and significant infrastructure work, which we managed to do.

数据生成与整合 Data Generation and Integration

Host

我们刚才在谈数据。这实际上,在我看来,是一个非常有趣的问题,我们还没有详细探讨过。我们能不能稍微详细说明一下?例如,这些数据长什么样?你们主要是大规模求解微分方程来生成训练数据吗?你们从各种来源获取物理数据吗?你们如何整合物理数据和建模、计算数据?

We're just talking about data. This is actually, in my opinion, a very interesting question that we have not yet explored in detail. Can we, can we elaborate on this a little bit? For example, what does this data look like? Are you mainly solving differential equations on a large scale to generate training data? Do you obtain physical data from a variety of sources? How do you integrate both physical data and, you know, modeling and computational data?

Anima

嗯。这实际上是一个相当有趣的点,因为这曾经是一个“瓶颈”,尤其是对于许多这类狭窄的替代模型。你可能见过这种情况:公司拥有出色的数据集,你有很棒的演示,但他们真正感兴趣的领域超出了样本的范围,结果什么也没实现。所以我们想摆脱这个限制。

M-hm. And this is actually quite an interesting point, as this used to be a "bottleneck", especially for many such narrow surrogate models. And you may have seen it. The company has an excellent data set. You have a great demo, but the area they are really interested in is beyond the scope of the sample, and nothing is implemented. So we wanted to get rid of this limitation.

PDE桥接模拟现实 Bridging simulation and reality with PDEs

Anima

与此同时,我们都知道,在物理过程的模拟中,核心主题是模拟与现实之间的差距,这个差距必须以某种方式弥合。那么我们如何同时实现这两个目标呢?有趣的是,对于偏微分方程(PDE),我们相当有信心,因为数学是已知的:当你正确求解 PDE 时,你就正确地对物理进行了建模。当然,你仍然需要确保你反映了你试图解决的任务,但你可以通过确保你的技能领域足够广泛,以保证保持在分布范围内,从而弥合模拟与现实之间的差距。那么,这为学习本身带来了什么呢?我们可以使用数值模拟器生成我们需要的尽可能多的训练数据。更进一步,我们可以做所谓的课程工程,即从简单的方程开始,从较低的分辨率开始,用更简单的数据。这与语言模型的训练方式形成鲜明对比。例如,对于语言,你下载整个互联网。这是一个巨大的混乱。一切都混在一起。有些是正确的,有些不是。这根本不是我们人类学习的方式。持续学习是一个更好的选择。所以我们有机会这样做。我们可以使用数值模拟器构建我们自己的训练程序,但这并不止于此。如果你停在那里,你会像机器学习中的其他一切一样,最终得到训练分布的平均质量。我们想要超越这些限制。这就是自我改进发挥作用的地方。所以,最有趣的是,你可以将这些偏微分方程(PDE)既用于数值模拟器生成数据,又如果你明智地行动,用作训练信号。你可以测试你的模型在 PDE 上的表现,并将其作为额外的训练信号,这样你突然发现自己处于一种可以在模型质量上超越数据质量的情况。与语言相比,这里的区别在于,语言模型需要来自人类或其他奖励信号的反馈来自我改进,而这些反馈非常稀缺。它们只是告诉你“是”或“否”,竖起或放下手指。我们有密集的反馈,因为有许多物理定律,你可以再次为它们的位置定义程序,并从它们那里获得丰富的反馈。因此,这种自我改进可能更有效,这正是我们在实验中观察到的。

At the same time, we all know that for the simulation of physical processes, the main theme is the gap between simulation and reality, which must somehow be bridged. So how do we achieve both goals? Interestingly, with partial differential equations (PDEs), we are pretty confident because the math is known: when you solve a PDE correctly, you model the physics correctly. Of course, you still need to make sure you are reflecting the task you are trying to solve, but you can bridge the gap between simulation and reality by making sure that your skill domain is broad enough to be guaranteed to stay within the distribution. Now, what does this give us for the learning itself? We can use numerical simulators to generate as much training data as we need. And even more, we can do what's called curriculum engineering, which is to start with simple equations, start with lower resolution, with simpler data. This is in stark contrast to how language models are trained. For example, for a language, you download the entire internet. This is a huge chaos. Everything is mixed up. Some of this is correct, some of it is not. And this is not at all the way we humans would learn, for example. Learning consistently is a much better option. So we have the opportunity to do that. We can build our own training program using numerical simulators, but it doesn't end there. If you stopped there, you would, like everything else in machine learning, eventually end up with an average quality of your training distribution. And we want to go beyond these limits. This is where self-improvement comes into play. So, the most interesting thing is that you can use these partial differential equations (PDEs) both for numerical simulators to generate data and, if you act wisely, as a training signal. You can test how well your model performs on the PDEs themselves and use that as an additional training signal, so that you suddenly find yourself in a situation where you can outperform the data quality in the quality of your model. And the difference here compared to language is that language models need feedback from people or other reward signals to self-improve, which are very rare. They simply tell you "yes" or "no", a raised or lowered finger. We have dense feedback because there are many physical laws, and you can again define a program for their location, and also receive rich feedback from them. Therefore, such self-improvement may be even more effective, and this is exactly what we observe in our experiments.

多尺度物理处理 Handling multiscale physics

Host

也许这已经隐含在你的回答或之前的陈述中,但物理学中让我感兴趣的一点是多尺度的概念,你知道,你需要同时表示在大尺度和小尺度上发生的过程。通常,高质量物理建模的困难恰恰在于你需要跨越巨大数量级的不同尺度,问题是如何同时表示它们?你究竟如何解决这个问题?

Maybe this was already implicit in your answer or previous statements, but one thing that interests me in physics is the idea of multiscale, where, you know, you need to represent processes that happen on both large and small scales simultaneously. And often the difficulty of high-quality physical modeling is precisely that you need different scales that span huge orders of magnitude, and the question is how to represent that simultaneously? How exactly do you solve this problem?

Anima

所以如果你看看我们发表的许多关于神经算子的工作,你会发现我们正是在解决这个问题,对吧?也就是说,如果存在多个尺度,你还需要不同的分辨率。当然,你可以用最大的细节做所有事情,但这极其昂贵。因此,最好先以较低分辨率收集更多数据,也许甚至使用更简单的求解器,这些求解器简化模型并忽略小效应。它们不精确,但它们是理解一般效应和主要趋势的良好起点,然后用考虑这些更精细细节的更专门的求解器对它们进行微调。神经算子具有这种灵活性,因为它们允许你在模型内部混合不同的尺度,而不是像物理学中许多其他混合机器学习方法那样在外部嵌入它们。这使我们能够更高效地处理数据,更好地学习。

So if you look at a lot of our published work with neural operators, you'll see that we're addressing exactly that, right? That is, if there are multiple scales, you also need different resolutions. Of course, you can do everything with maximum detail, but it is extremely expensive. Therefore, it is better to collect more data at lower resolution first, perhaps even using simpler solvers that simplify the model and ignore small effects. They are imprecise, but they are a good starting point to understand the general effects, the main trends, and then fine-tune them with more specialized solvers that take into account these finer details. And neural operators have this flexibility because they allow you to mix different scales within the model, rather than embedding them externally, as many other hybrid machine learning methods in physics do. And that, you know, allows us to be much more efficient at working with data and better at learning.

应用与商业化 Application areas and commercialization

Host

我们时间不多了。我想听你谈谈这里的应用领域是什么?你首先想实现什么?你已经走出隐身模式,所以你可能在商业化方面做了一些更公开的事情。你能多告诉我一些吗?

We don't have much time. I want to hear a little bit about what are the application areas here? What are you trying to implement first? You've come out of stealth mode, so you're probably doing something more public in terms of commercialization. Can you tell me a little more about this?

Anima

我们从两个角度来看待这个问题。有物理学的分支,也有它们可以服务的领域。因此,如果我们考虑当前非常有趣的领域,因为每个人都需要改进,对每个人来说主要的领域是半导体和能源——这正是我们的专长。我们不能完全透露我们与谁沟通以及我们的客户是谁,但在半导体领域,有许多问题物理学变得越来越相关。有趣的是,如果你看看那里许多事情是如何运作的,尤其是在设计芯片本身时,方法更像是:“让我们从数字开始,固定数字设计,让我们通过一次物理验证,因为 PDK 要求这样那样的参数,否则台积电不会为你生产,”就这样。比如,“哦,我们有足够的库存来做这个吗?”“但我们今天没有利用这些库存来提高生产力。”所以,当你能够理解物理学允许什么来推动边界,并且我能否稍微重新安排一些东西来进一步推动这些边界时,就有潜力开辟更多的可能性。这是其中一个领域。显然,半导体制造过程本身也有所有这些有趣的物理挑战,我们正在解决。所以有很多很多工作。我们很高兴。

We look at this from two perspectives. There are branches of physics, and there are areas they can serve. So, if we consider the areas that are currently very interesting, because everyone needs improvements there, the main ones for everyone are semiconductors and energy—and this is exactly our specialization. We can't fully disclose who we communicate with and who our customers are, but within semiconductors there are many problems where physics is becoming increasingly relevant. And what's interesting is that if you look at how a lot of things worked there, especially when designing the chips themselves, the approach was more like: "let's start with the numbers, fix the digital design, let's run it through physical verification once, because PDK requires such and such parameters, otherwise TSMC won't produce it for you," and that's it. Like, "Oh, do we have enough inventory to do this?" "But we're not using that inventory to improve productivity today." So there's the potential to open up a lot more possibilities when you can understand what physics allows to push the boundaries, and can I rearrange things a little bit to push those boundaries even further? This is one of the areas. Obviously, the semiconductor manufacturing process itself also has all these interesting physical challenges that we're solving. So there is a lot, a lot of work. We are delighted.

数模半导体设计 Joint design of digital and analog semiconductors

Host

那么,我理解得对吗,这是数字和模拟半导体的联合设计?

So, do I understand correctly, this is a joint design of digital and analog semiconductors?

Anima

显然,你必须从某个地方开始,而且现在整个流程也在改进。我们正在逐步提供这个物理基础,这当然是我们前进的方向。有不同的公司制造不同的部件。我们从这个方向开始,因为我们相信它开辟了巨大的机会,而以前那里有许多简化。你处理导线长度,进行某些优化。你想,“哦,我某个地方有热击吗?”我某个地方有电磁浪涌在困扰我吗?但不是,“哦,我有一堆其他东西,实际上很酷,几乎没什么问题,我能利用那个吗?”陡峭地。

Obviously, you have to start somewhere, and there's this whole pipeline that's also being improved now. We are taking care to gradually provide this physical foundation, and that is, of course, where our path is headed. There are different companies that make different parts. We are starting in this direction because we believe it opens up great opportunities where there were previously many simplifications. You worked with wire lengths, made certain optimizations. You're thinking, "Oh, do I have a heatstroke somewhere?" Do I have an electromagnetic surge somewhere that is bothering me? But not, "Oh, I have a bunch of other stuff where it's actually pretty cool and there's hardly anything going on, can I use that?" Steeply.

能源应用 Energy applications

Host

至于能源,我想我们之前已经谈过核能。我想那是方向之一。你也涉足太阳能和其他类似的东西吗?

And energy, I guess we've already talked about nuclear energy before. I suppose that's one of the directions. Are you also into solar energy and other things like that?

Anima

作为设计的一部分,那就是:“我想要一个具有特定能力的物理对象。”这是我们可以做得非常好的事情。

There is as part of the design, that is: "I want a physical object with certain capabilities." This is something we can do very well.

物理通用模型 Universal Models in Physics

Host

还有其他领域,问题是:“我对世界有一个观察,这能告诉我关于世界的什么?”以地热能为例,你需要知道地下的热量在哪里,以及可以泵送的水的合适混合物。或者,我们所需的所有电子设备的关键矿物在哪里?所有这些都属于我们的范畴,你通过扎实的物理学来实现这一点,但同时,在这些物理学中,事物之间并没有太大不同。所以,当你的模型变得足够普及时,突然间很多门就打开了。

There are other areas where the question is, 'I have an observation of the world, what does that tell me about the world?' Take geothermal energy, where you need to know where the heat is down there, with the right mixture of water that can be pumped. Or where are the critical minerals for all the electronic devices we need? All of this falls within our purview, and you achieve this through solid physics, but at the same time physics where things are not that different from each other. So when you reach a point where your model becomes universal enough, suddenly a lot of doors open.

Guest

稍微扩展一下这个想法,你会看到这种迁移发生在你与真实商业客户进行的应用中。

Just to expand on that idea a little bit, you see this transfer happening in the applications that you do with real commercial customers.

Host

没错。

That's right.

Guest

我们看到了这种物理上的提升。我们之前提到过。我们训练模型时想:让我们确保我们的模拟器运行良好,设置正确,我们确实可以训练它。然后我们想,哦,让我们把这个加到通用模型里。

We see this physical rise happening. We mentioned this before. We trained models where we thought: let's just make sure our simulator works well, is set up correctly, and we can actually train it. And then we thought, oh, let's add this to the universal model.

Host

我总是想检查这类事情。有一大堆流程工作需要完成,以确保一切正确。

I always want to check things like this. There is a whole bunch of pipeline work that needs to be done to make sure everything is correct.

Guest

我们注意到,就我们在窄模型中看到的性能特征而言,与我们在宽模型中突然看到的相比,宽模型在各方面都更好。

And we noticed that in terms of the performance characteristics that we see in the narrow model, compared to what we suddenly saw in the wide model, the wide model was better across the board.

未来计划与招聘 Future Plans and Hiring

Host

我们时间不多了,但也许这是最后一个问题。加速理解项目下一步是什么?你想让人们知道什么?你们在招人还是开设办公室?发生了什么?

We're short on time, but maybe this is the last question. What's next for Accelerated Understanding? What do you want people to know? And are you hiring people or opening offices? What is happening?

Guest

是的,当然,所有这些。我的意思是,我们在扩展我们的模型,并继续扩大我们的公司。所以是的,自上周以来我们收到了很多请求。所以我正在处理这些。如果我还没有回复任何人,抱歉,但这确实是一个激动人心的时刻。你知道,对我们来说,这只是扩展模型之旅的开始,对吧?如果我们要走向科学发现和发明中面临的最大挑战,那确实是复杂的物理学、复杂的现象。但与此同时,我们需要忠于我们的商业方面,确保我们在这个过程中为客户创造大量价值。正是这种结合真正吸引着我们。

Yes, of course, all of that. I mean, we're scaling our models and continuing to scale our company. So yes, we've been getting a lot of requests since last week. So I'm just dealing with this right now. Sorry if I haven't replied to anyone yet, but this is truly an exciting time. And you know, for us this is just the beginning of the journey of scaling models, right? If we are to move towards the greatest challenges we face in scientific discovery and invention, it's really complex physics, complex phenomena. But at the same time, we need to be true to our commercial side and ensure that we create a lot of value for our customers along the way. It is this combination that really fascinates us.

时代百大认可 Time 100 Recognition

Host

嗯,实际上,我只想请你用 30 秒谈谈《时代》百大人物。这是怎么发生的?这一切是怎么发生的?我从未获得过这样的奖项,所以我很好奇。

Well, actually, I just want to ask you for 30 seconds to talk about the Time 100. How did it come about? How did it all happen? I've never received such an award, so I'm just curious.

Guest

你知道,当我入选《时代》百大 AI 人物时,我感到非常惊讶和意外。去年我也很幸运获得了《时代》百大影响力奖。所以我为此去了迪拜。而且,你知道,我会见了这个领域的许多杰出人物。现在,你知道,今年,我认为他们试图汇集对 AI 有不同观点的人,这很好。所以是的,这是一件好事。

You know, I was just amazed and very surprised when I made the Time 100 AI list. And I was also very lucky to receive the Time 100 Impact Award last year. So I was in Dubai for this. And, you know, I met with many prominent figures in this field. And now, you know, this year, I think it's good that they're trying to bring together people with different perspectives on AI. So yes, it's a good thing.

结束语 Closing Remarks

Host

是的。嗯,恭喜你。你当之无愧。非常感谢你再次加入我们;我们期待看到加速理解项目如何发展。

Yes. Well, congratulations. You deserve it. And thank you very much for joining us again; We look forward to seeing how the Accelerated Understanding project develops.

Guest

谢谢你再次邀请我们。

Thank you for inviting us again.

Host

谢谢。保重。再见。

Thank you. Take care of yourself. See you later.

互动版:逐字朗读 + 针对本期提问 →