AI 科学家已至:自主实验室与合成超级智能——Periodic Labs

AI Scientists Are Here: Autonomous Labs & Synthesis Superintelligence — Periodic Labs

利亚姆·费杜斯 Liam Fedus · Latent Space · 2026-10-08 · 约 85 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Periodic Labs 的 Liam 阐释为何仅有智能还不够,以及自主物理实验室如何闭合假设与现实之间的循环。

Liam of Periodic Labs explains why intelligence alone is insufficient and how autonomous physical laboratories close the loop between hypothesis and reality.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 49)

全文 · Full transcript(中英对照)

为何需要物理实验室 Why Physical Labs Are Needed

Liam

但我们也认为,即使 Fable 7 变得非常完美,甚至比今天更好,它也不在乎——你仍然必须花实验来获得结果。原因在于,机器学习在它被训练过的内容上非常有效。但科学发现,按定义,是关于你没有被训练过的内容。这就是为什么我们建造这些实验室,让开放或封闭的模型都可以用它们来与宇宙互动,因为我们相信,不做测试就不可能做出伟大的发现。没有人能就这样,从零开始,创造出室温超导体。

But we also think that even if Fable 7 becomes very perfect, even better than today, it doesn't care — you'll still have to spend experiments to receive results. And the reason for this is that machine learning is very effective on what it was trained on. But scientific discovery, by definition, is about what you weren't trained on. And that's why we build these laboratories, so that either open or closed models could use them for interaction with the universe, because we believe that it is impossible to do great discovery without testing. No one can just like that, from scratch, create a room-temperature superconductor.

赞助商信息 Sponsor Message

Host

在我们进入今天的节目之前,我有一小段话要对听众说。谢谢。如果没有你们决定点击我们的视频并观看,我们就无法提供你们想看的关于 AI 工程、科学和娱乐的内容。几乎每天来看我们——赞助商在申请,但幸运的是,你们中有足够多的人订阅我们,支持这一切;它是无广告的,我们希望将来也保持这样。但我对你们所有人只有一个请求。最重要且完全免费的事情,就是按下“订阅”按钮。这是我只想告诉你们的一件事;我有时会请求你们。这对我和我那每周努力为你们创作 In Space 的团队来说意义极其重大。如果你们这样做,我保证我们永远不会停止努力,让这个节目变得更好。现在让我们进入正题。

Before we move on to today's episode, I have a small message for listeners. Thank you. We would be unable to provide you content about AI engineering, science, and entertainment that you want to see if you have not decided to click on our videos and watch them. Visit us almost every day — sponsors are applying, but fortunately, enough of you subscribe to us to support everything; it's ad-free, and we want it to remain so in the future. But I only have one request to all of you. The most important and completely free thing you can do is press the button 'sign up'. It's the only thing I want to tell you about; I will ask you sometime. And this means extremely much for me and my team, which works hard to create In Space for you every week. If you do it, I promise that we will never stop working to make this show even better. Now let's move on to affairs.

介绍Liam和Doga Introducing Liam and Doga

Host

好的。我们在 Periodic。今天我们请到了来自 Periodic 的 Liam 和 Doga。恭喜你们,也感谢你们接待我们。

OK. We are here, in Periodic. We are today here with Liam and Doga from Periodic. Congratulations and thank you for accepting us in yourselves.

Liam

是的,我们很高兴来到这里,也感谢你们到来。

Yes, we are happy to be here, and thank you for coming.

智能必要但不充分 Intelligence Is Necessary but Insufficient

Host

我想从你们网站上的一句引言开始,我非常喜欢它。上面说:“智能是必要的,但不充分。当想法与现实相符时,新知识就被创造出来。”这看起来很明显,但为什么它不是不言自明的呢?

I want to start with one quote from your site, which I really like. There it is said: 'Intelligence is necessary, but insufficient. New knowledge is created when ideas are found to agree with reality.' It seems like this is obvious, but why is this not self-evident?

Liam

我认为这是我们正在谈论的 Periodic 的论点和前提。他们在创建这个时就想到了:你不能仅仅“想通”通往决策的道路。宇宙如此复杂,为了真正扩展知识的边界并取得进展,必须创建假设,然后检查它们是否真实。仅仅重读教科书或文章,就像房间里的一次思考或短路,不会让你想通实验的所有可能结果,预期的和意外的。你确实需要这个迭代过程。这是我们的主要信念,为什么我们创建这个,也是为什么我们从一开始就决定我们需要什么:人工智能系统、智能、物理世界的模拟,但也需要开发物理上的高效实验。我们相信,从这一点中会产生另一种类型的智能。

I think it's kind of the thesis and premise of Periodic, which we are talking about. They thought about it when they created this: you cannot just 'think through' the way to a decision. The universe is so complicated that in order to really expand the boundaries of knowledge and achieve progress, it is necessary to create hypotheses, and then check whether they are true. Just rereading a textbook or article, just like a thought or short circuit in a room, won't let you think through all possible results of experiments, expected and unexpected. You really need this iterative process. And this is our main belief, why we created this, and that's why from the very beginning we decided what we need: artificial intelligence systems, intelligence, simulation of the physical world, but also development of physical highly productive experiments. And we believe that from this another type of intelligence arises.

从零搭建实验室 Building a Lab from Scratch

Host

我认为这也很了不起,就像 OpenAI 和 Google 的情况一样,你们在那些大型实验室里没有任何资源。也就是说,你必须去创造你自己的东西,因为你仿佛一边走一边制定自己的游戏规则。

I think this is also remarkable, as in the case of OpenAI and Google, you don't have any resources in those large laboratories. That is, you had to go and make your own thing, because you are as if you are making up your own rules of the game as you walk.

Liam

是的,我的意思是,类似这样的指令——我认为这永远不会发生。它不存在。我们试图联合化学家——固态化学家、固态物理学家、实验家、理论家、硬件和软件工程师、LLM 专家、计算机科学家。其中一些技术非常新;例如,高性能实验、机器人操纵器仅在最近 3 年左右才开始应用。使用力场的经验——其中一些力场也非常新。但是的,我们觉得母实验室非常重要,要有这种专注,并把所有这些人为非常协作的工作联合起来。我认为历史上最接近的例子是像贝尔实验室这样的地方,那里不可思议的理论家、实验家和化学家一起工作,取得了不可思议的成就。所以我们试图在这里做同样的事情。

Yes, I mean that a command similar to this — I think this one will never happen. It did not exist. We are trying to unite chemists — solid-state chemists, solid-state physicists, experimentalists, theorists, engineers from hardware and software, LLM experts, computer scientists. And some of these technologies are very new; for example, high-performance experiments, robotic manipulators began to be applied only in the last 3 years or about that. Experience working with force fields — some of these force fields are also very new. But yes, we felt that the mother laboratory is very important to have this focus and merge all these people for very collaborative work. And I think the closest examples from history are places like Bell Labs, where incredible theorists, experimentalists, and chemists worked together and achieved incredible things. So we try to do the same here.

与大型AI实验室的差异 Differences from Large AI Labs

Host

我想谈谈力场,但首先,在我们进入这个之前,它是什么?嗯,但在我们讨论它之前,我认为,为了提供背景,最大的区别之一——你知道,我们的观众一部分是 AI 工程师,一部分是科学家。因此,让我们简短一点,让我们对工程师说,我们问:如果你在 OpenAI 这样的大型实验室工作,有什么区别?你在那里做的事情和这里发生的事情有什么区别?我想你已经暗示了这一点,但也许更清楚地说:你必须如何重新思考你的思路?

I would like to talk about force fields, but first, before we move on to this, so what is it? Um, but before we get to it, I think, for context, one of the biggest differences between — you know, part of our audience — this is AI engineers, and some — scientists. Therefore, let's be brief, let's talk to engineers and we ask: in what difference if you worked at OpenAI as a large laboratory? What is the difference between what you do there, and because, you know, what is happening here? I think you already hinted at this, but perhaps more clearly: how you must rethink your train of thought?

Liam

嗯,我的意思是,一件有趣的事情是,现在我们的强化学习环境实际上来自我们的物理实验室。我们的数据来自我们的物理实验室,这是我们的主要真理。仅仅基于文章或教科书中已知的答案进行优化是不够的,因为我们超越了这些限制。我认为这是最大的区别之一。但随之而来的是问题,比如说,在不确定性条件下接受解决方案。因此,当你在数学中进行优化时,它具有高精度。你实际上不必处理变异性、不确定性或错误测量,而对我们的过程来说,这是关键。例如,当我们研究材料发现的循环时,东西从烤箱里出来时并没有标记,对吧?甚至标记过程也可能是随机的和不准确的。有时不同机器之间会出现偏差。也许我们缺乏遥测,所以我们以为我们在这个温度下做这件事,实际上它差了某个量。而智能能够感知这些嘈杂的数据并接受合理的决策,就像科学家会做的那样——这是另一套推理策略。所以我认为,在这之间,有大量的共同点。标准的强化学习训练过程、创建使用工具的智能体、减少离散度、检查结论是否与训练数据不一致。但我们必须更进一步,认真思考:当存在高度不确定性时,如何执行准确的工作?如何极其有效地使用有限的数据集?我们——我们扩展流程,但它仍然与数字环境非常不同,在数字环境中,你可以任意添加更多环境或迭代。我们没有这样的机会。因此,样本效率——另一个关键点。在我看来,这有点抽象,这是一个在科学播客中已经讨论了一段时间的话题。

Well, I mean that one interesting thing is that now our learning environments with reinforcement literally come from our midst — physical laboratories. Our data is coming from our physical laboratories, and this is our main truth. It's not simple enough to perform optimization based on the answers that are already known from articles or textbooks, because we go beyond these limits. And I think this is one of the largest differences. But along with this, questions arise, let's say, acceptance of solutions in conditions of uncertainty. Therefore, when you do optimization in mathematics, it has high accuracy. You don't really have to deal with variability, uncertainty, or false measurements, whereas for our process this is key. For example, when we work on the cycle of opening materials, things don't come out of ovens already marked, right? Even the labeling process may be stochastic and inaccurate. And sometimes deviations occur between different machines. Maybe we are lacking telemetry, so we thought we were doing this at this temperature, and actually it differed by a certain amount. And intelligence capable to perceive these noisy data and to accept reasonable decisions, how would a scientist do it — this is another set of reasoning strategies. So I think that between this there is a huge amount of commonality. Standard training with reinforcement in process, creation of agents that use tools, dispersion reduction, checking that there was no conclusion disagreements with training data. But we have to go further and seriously think about: how to perform accurate work when there is a high level of uncertainty? How to extremely effectively use a limited set of data? We are much — we scale processes, but it's still very different from a digital environment, where you can arbitrarily add more environments or iterations. We don't have such an opportunity. Therefore, sample efficiency — another key point. It seems to me a little abstract, and this is a topic that has already been discussed for some time in science podcasts.

科学中的实验不确定性 Experimental Uncertainty in Science

Host

但也许在你谈论实验不确定性时,解释一下会很有用:你正在进行实验的具体过程是什么样的,在引入 AI 之前,一个人会如何处理这个问题?你能否举一个具体的例子,说明你创造的材料类型,以及在这个过程中测量不确定性究竟出现在哪里?

But perhaps it would be useful when you are talking about experimental uncertainty, explain: what did it look like a specific process, which you are conducting experimentally, and how would a person approach this problem, before we move on before you are involving AI? Does the you specific example type the material you create, and which one exactly uncertainty in measurements occur during this process?

Liam

我认为这里有很多测量。其中之一是,在我看来,每个科学实验都必须进行某种降维。这与编程或数学非常不同。因此,当你在做数学或编程时,整个上下文可以为人或语言模型所用。这是一组公理,一些人们推导出的结果。这是一组函数和 API 调用。所以,你思考所需的一切都在手边,你只需要非常聪明并弄清楚这一点。

I think there are many here measurements. One of them is that, in my opinion, everyone scientific experiment must perform a certain reduction dimensions. And this very different from programming or mathematics. Therefore, when are you working out mathematics or programming, the whole context can to be available person or language models. This is a set axiom, some consequences that people with They were taken out. This is a set functions and calls APIs. So, everything you need needed for considerations, do you have under hand, and you just must be very smart and figure this out.

Liam

在物理学中,如你所知,我们一开始要处理的原子数量比我们能在计算机中保存的要多得多。所以,显然,我们必须从描述系统的初始测量数量和比特数,转变为我们能塞进计算机的比特数。你知道,这可能是热力学的初始前提。起初,他们对蒸汽机的工作方式非常困惑,但后来发现,你可以找到五六个热力学变量,它们能很好地解释正在发生的事情,这太不可思议了。这就是物理学的美丽之处。例如,你有一个充满原子的房间,有 10 的 23 次方或 10 的 27 次方个原子,但你只需要温度和压力,以及一些简单的变量,这就能告诉你一切。也就是说,在大多数情况下,你只需要知道这些,对吧?

In physics, as you know, we start with more quantity atoms than ever before we can save on computer. So, obviously, we will have to go from the initial number of measurements and bits describing system, to the number bits that we can fit into a computer. And you know, this is probably was the initial prerequisite thermodynamics. At first they were very confused the way they worked steam engines, but then it turned out that can you find five or six thermodynamic variables that are quite explain well that is happening, and this incredible. This is it. the beauty of physics. For example, you have a room full of atoms, there are 10 of them in the 23rd or 10 in to the 27th power, but you only needed temperature and pressure, and also simple variables, and this tells you about all. That is, he says you prefer that need to know in in most cases, isn't that right?

Host

假设你创造了一种材料,你能做类似 X 射线结构分析的事情,但它是非常有损的投影,或者说是错误的?这实际上不是结构,或者它没有向你展示真实的结构。它只显示了结构的主要组成部分的一部分。那么,一个人如何利用这些简单的主要组成部分来思考它们,以及如何将其扩展到人工智能?

Let's say you create material and can you do something like x-ray structure analysis, but it is very lossy projection, or wrong? This is actually not a structure, or she doesn't show you the real structure. It only shows part of the main components of this structures. So how a person can take these simple main components, to ponder over them, and how to extend this to artificial intelligence?

Liam

所以,从理论物理的角度来看,是的?我们认为能量、能量的平均值及其涨落非常重要。由此引出了许多热力学规定,然后我们明白,能垒反应也很重要。动力学非常重要。所以,一个人可以尝试理解这个极其困难的系统,通过降维发明某些描述符。人们也非常喜欢原子图像。例如,你有没有想过局部环境看起来如何,或者它由八面体、四面体组成,化学尤其受益于此。所以,这是从理论方面,而从实验方面,你试图收集尽可能多的数据。所以你有实验室日志,你写下你看到的一切,等等。但要明白,当然,有很多你会错过的东西。

So, from the side theoretical physics, Yes? We decided that energy, average the value of energy and its fluctuations are very important. From this a lot was brought out thermodynamic provisions, and then we understood that energy barriers reactions are also very important. Kinetics very important. So a person can try to understand this extremely difficult system, inventing certain descriptors from reduced dimensionality. People are also very like atomistic picture. For example, whether Have you wondered how? looks local environment, or it consists of octahedra, tetrahedra, and chemistry especially benefits from this. So, this is from the theoretical side, and from the experimental side you are trying to collect as much as possible more data. So in you have a laboratory journal, you write down everything you see, etc. But with the understanding that, which, of course, is a lot What will you miss?

Liam

所以,正如 Liam 提到的,所有这些都会产生问题。例如,你的炉子随着时间的推移会磨损,因为在使用过程中,你处理的部分材料会蒸发并覆盖加热元件。所以每次你使用炉子时,情况都会变得越来越糟。另一个问题是,炉子内的温度并不理想均匀。因此,你把材料放入炉子的位置会影响结果。我们到目前为止还没有这个问题,但也许某天会出现。光学工具非常容易受到振动的影响。我记得,当我攻读博士学位时,在一个实验室里有一个严重的问题:结果有时非常不同。最终,他们发现有人在晚上走动。在楼上,这种振动影响了激光装置。所以是的,科学非常困难,但这正是它特别之处,不是吗?

So, as Liam mentioned, all these arise problems. For example, your oven over time will wear out, because during use part materials that you process, will evaporate and will cover heating element. So every time, when you use oven, it becomes everything worse and worse. Other the problem is because the temperature in the oven is not ideal uniform. Ago the place where you put the material in furnace, will affect the result. We have so far that there is no this problems, but, maybe someday will appear. Optical tools very suffer from vibrations. I remember, when I was working on doctoral, in one the laboratory was serious problem: results sometimes were very different. Eventually, they found out that Someone was walking at night. floor above, and this vibration affected laser installation. So yes, science is very difficult, but that's what makes her special, isn't it?

Liam

还有一件事,我们与 Liam 谈过:当前许多改进集中在数学、编程和理论计算机科学上,因为这对 LLM 来说更容易,但在现实生活中,大多数需要智能的任务更类似于科学。存在坚实的不确定性、噪声和缺乏上下文,但你必须保持智能并决定下一步做什么。在我看来,值得补充的重要一点是:数学或理论计算机科学家的最优策略考虑可能不是科学的最优策略。是的。当你的输出是嘈杂的晶体结构,或者,我不知道,某种集合变量时,如何定义 RL 奖励函数?

One more thing about We spoke with Liam: many current improvements focused on mathematics, programming and theoretical computer science, because it easier for LLM, but in real life most of the tasks that require intelligence, more similar to science. There is a solid uncertainty, noise and the absence context, but you must remain intelligent being and to make a decision that to do next. It seems to me, important point that it is worth adding: optimal strategies considerations for mathematics or theoretical computer scientists can not be optimal for science. Yes. How to define a function RL rewards when your way out is noisy crystalline structure or, not I know, some kind of set collective variables?

Liam

例如,也许我会给出循环的一部分,并更深入地探讨它。所以,作为发现新材料循环的一部分,我们必须首先弄清楚到底创造了什么。这涉及原子在一起的稳定性,以及我们是否期望材料具有我们需要的性质。我们不仅需要新的原子构型,我们希望它们组合起来具有必要的性质。接下来——如何合成它?实际上如何制造这个物体?这也不是一件容易的事。所以,即使你知道结构会保持,确定其创建的处理条件的过程也极其困难。但话说回来,当你完成这两个步骤并创造出某些东西时,它没有标签,所以你需要对其进行表征,并弄清楚你到底做了什么。

For example. Maybe I'll give one part of the cycle and I'll dig a little deeper into her. So, as part of opening cycle new materials, we must first to find out what exactly create. It concerns stability of atoms together, and also that do we expect that the material will have we need properties. To us not just needed new configuration atoms, we want to in combination they had necessary properties. Next—how synthesize it? As actually make this object? This too not an easy task. So, even if you you know what the structure is will hold, process of determination processing conditions for it creation is extremely difficult. But then again, when you have completed these two steps and something created, it does not have labels, so you need it to characterize and find out what exactly you are done.

Liam

所以,在这种情况下,而不是把强化学习环境看作我们只是启动实验、等待几天、试图找出是否发现了室温超导体?这绝对不可能。我只是想问这个。是的,你不能。你可以让你的 GPU 工作一周或类似的事情。这绝对不可能,因为你没有足够的智能体来足够减少分散。时间部署花在不同工具之间的切换上。你只能等待。实验有物理限制,例如直接合成的时间或炉子的工作。所以这太慢且太嘈杂。因此,我们正在考虑将 AI 程序作为围绕我们收到的数据集创建智能体。在表征的情况下,你正在寻找一个从强化学习环境中学习的环境,该环境因识别真正存在的相而获得奖励。你能够拥有原始实验数据并有效地处理吗?我们这样做:我们将 X 射线照射在材料上。X 射线具有频率或波长,大约与原子间距离成正比。所以你得到良好的衍射图案,这是一种晶体结构的印记。而识别实际存在的东西,是一项非常困难的任务。

So, in in that case, instead of to think about learning environment with reinforcement as about something where we just we launch experiment, we are waiting a few days and we are trying to find out if they found we are a superconductor at room temperature or not? It absolutely impossible. I just wanted to ask. about this. Yes, you are not. you can leave yours GPU to work on a week or something like that. This is absolutely impossible, because in you are not enough agents to enough to reduce dispersion. Time deployment is spent on switching between different tools. You Just wait. There are physical limitation experiments, for example, time direct synthesis or work ovens. So this is too much slowly and too much noisily. Therefore we we are considering the program AI as creation agents around that the dataset that we received. And in case with the characteristic you looking for an environment learning from reinforcements, which rewarded for identification really present phases. Are you able to having raw experimental data, effectively to work out? We do it's like this: we direct X-rays rays on the material. X-rays the rays have frequency or length wave, which is approximately proportional to distance between atoms. So you you get good diffraction patterns, which is of a kind imprint crystalline structures. AND identification of that what actually present, there is very a difficult task.

AI用于实验科学 AI for experimental science

Liam

因此,我们早期的一些 AI 工作就是专门为这个创建系统的。它让我们能快很多,因为如果你在实验上花那么多时间,你很快就会在理解到底得到了什么这件事上遇到瓶颈。但话说回来,是的,现在用强化学习来做环境已经相当简单了:识别出存在的相,就给奖励。识别出错误的相,或者没有真正化学有效性的东西,就惩罚。它算是一个小的辅助工具,可以用在流程的不同环节。当你把它们全都拼起来,就得到一个完整的发现循环。

Therefore, some of our early AI work was dedicated to creating systems for this. It allows us to move a lot faster, because if you spend so much on experiments, you very quickly are faced with a bottleneck in understanding what exactly was received. But again, yes, environment learning from reinforcement is now quite simple: you reward for identification of phases present. You punish for incorrect phases or things that don't have real chemical validity. It's kind of a small supporting thing that can be used for different parts of the process. And when you put it all together, you get a full opening cycle.

Liam

好。也许还有一个要素是,当它积累了大量实验数据之后,你可以给世界状态加上一个时间标记函数。这样你就可以说:在这个日期,我们有如下实验证据。然后你可以创建带强化学习的环境,你说:“好,给定这组实验数据,以及一个选择——科学家的下一步是什么,或者那个实验的结果是什么?”这也非常有意思,因为它让我们能够实现一些从外部更难做的程序,因为如果之前的模型已经记住了这部分数据,而你又试图应用强化学习,那么如果它已经知道答案,而你又创建依赖这个答案的任务,它就可以模拟工作。

OK. Maybe one more element would be that, as it accumulates a large array of experimental data, you can add a marking function in time for the state of the world. Thus, you can say: on this date we had the following experimental evidence. And you can create learning environments with reinforcement, where you say, "Okay, having this array of experimental data and, you know, a choice, which was the next scientist's step or what was the result of that experiment?" And this is also very interesting, because it allows us to implement types of programs that it would be more difficult to do from the outside, because if the previous model already remembered part of this data, and you're trying to apply learning with reinforcement, then if it already knows the answer, and you create tasks that depend on this answer, it can simulate work.

Liam

所以,因为这看起来像是:“哦,她已经知道答案了,所以她不必进行复杂的物理推理。”她其实不需要做计算和模拟,它就能得到正确答案,然后你固定这个策略,并提高这些推理方法的权重。这些策略将没有论据可以泛化到新系统。因此,它们对我们没用。但当我们积累大量实验数据,并把它们全部关联起来,让整个过程可追溯时,它就能让我们创建在其他条件下根本不可能的新型环境。

So, since this looks like: "Oh, well, she already knows the answer, so she doesn't have to perform complex physical reasoning." She does not really need to perform calculations and simulations, it gets the correct answer, and then you fix this strategy and increase the weight of these methods of reasoning. These strategies will have no arguments to generalize to new systems. Therefore, they are not useful for us. But when we accumulate arrays of experimental data and link it all together, having traceability of the whole process, it allows us to create new types of environments that are simply impossible in other conditions.

物理定律已穷尽了吗? Have we finished the laws of physics?

Host

我想问我的问题,但我怕我跑题了。我还是问吧,因为我不是科学家,而是工程师,我熟悉机器学习,但不是物理这边。好,这个问题有几个版本,但我认为主要的问题是:我们是否已经发展出了大部分物理定律?如果是这样,那怎么解释我们仍然没有完美的模拟器?

I want to put my question, but I'm afraid that I'm getting off topic. I still ask, because I am not again a scientist, but an engineer, I'm familiar with ML, but not from the physics side. Okay, there are a few versions of this, but I think the main thing is the question: whether or not we have already developed most laws of physics? And how so? It turns out that we have there are still no perfect simulators?

Liam

这是个好问题。这是个非常基础的问题,它就像这样:“哈哈,是啊,挺可爱的。”但我说:“老兄,我有这么多物理教科书,你什么意思?我们还有什么没完成?”是的,确实如此。有几件事,对吧?是的?因此,是的。我认为,首先,我们绝对没有完成任何物理定律的工作。即使你在量子层面——好吧,也许在非常大的尺度上——也是,但在人类层面,在材料层面,我们就像……还剩下什么?所以,这真的是个好问题。而且非常好奇为什么我们还没完成,对吧?

That's a great question. Like, this is a very basic question, and it is like this: "Haha, yeah, that's cute." But I say, "Dude, I have all these physics textbooks, so what do you mean? What haven't we finished yet?" Yes, this is happening. A few things, right? Yes? Therefore, yes. I think, first of all, we definitely haven't finished work with none of the laws of physics. Even if you at the quantum level— okay, maybe on very large on a scale—also, but on a human level, at the level of materials, we're like... what else is left? So, this is really a good question. And very curious why we haven't finished yet, have you? Yes?

Liam

嗯,好吧,有一个非常常见的故事,大家都在谈论,他们说。当狄拉克发展量子力学时,大约在 1920 年代末,他写了一本量子力学教科书,他似乎是这样呈现的,好像一切都已经完成了。有一句话人们很喜欢引用。我其实不确定它有多准确,但狄拉克大概说过:“剩下的就是化学了”。他的想法是,他能解决关于氢原子的问题。全部。是的,也许他能解决像一维氢原子链这样的系统,但它解决不了,比如说,氮和氧的相互作用,但这没什么,这只是简单的化学。是的。

Um, well, there is one very common story that everyone is talking about, they say. When Dirac developed quantum mechanics, somewhere in the late 1920s, he wrote a textbook on quantum mechanics, and he seemed to present it this way, as if everything is already done. And there is one phrase that people love to quote. I actually am not sure how accurate it is, but probably Dirac said: "The rest is chemistry". The idea was that he could solve the problem about the hydrogen atom. All. Yes, maybe he could solve the system, such as a 1D chain of hydrogen atoms, but it could not solve, let's say, interaction of nitrogen with oxygen, but this nothing, it's simple chemistry. Yes.

Liam

物理学家真的很喜欢单个的东西,比如原子,非常简单的系统,而这就是所有能解的东西。球形奶牛。是的,是的,球形奶牛。但事实证明,我们在过去 100 年学到的东西,首先,它不是真的。它不是我们在氢上所做工作的简单扩展。而且它不容易……剩下的——它不只是化学。事实上,有很多物理。我的意思是,我们仍然没有弄明白的事情之一就是高温超导。但还有更多。

Physicists really love singles, like atoms, very simple systems, and that's all that can be solved. Spherical cow. Yes, yes, spherical cow. But it turns out that what we learned over the last 100 years, firstly, it is not true. It's not easy expansion of what we did with hydrogen. And it's not easy... The rest— it's not just chemistry. In fact, there are many physics. I mean that one of the things that we still haven't figured out is high temperature superconductivity. But there is many more.

Host

对你来说,“高温”是多少——175,200?

And "high" for you— is it 175, 200?

Liam

说“温度超导”,他们指的是非常规超导。因此,有传统超导,它主要由电子-声子相互作用引起。在这种情况下,你可以看到同位素效应。如果你取同样的系统,但其中一个元素的重量不同,你可以看到超导温度下降,是的,正如预期的那样。所以,这是传统的。但更高温度的超导体,比如铜氧化物,那些高于 77 K 的,例如 93 K,事实证明不遵守这个物理。但我们不知道它们遵守什么物理。它们简直是不可思议的超导体。但这不仅仅是关于这个。所以这绝对是那些非常热门的话题之一。我们仍然不知道这方面的理论。但还有太多我们不理解的事情。

Saying "temperature superconductivity", they mean unconventional superconductivity. Therefore, there is traditional superconductivity, which is mostly caused by electron-phonon interaction. And in this case you can see the isotope effect. If you take the same system, but with different weights of one of the elements, you can see that the temperature of superconductivity drops, yes, as expected. So, this is traditional. But higher temperature superconductors, such as cuprates, those above 77 K, for example 93 K, it turns out not to obey this physics. But we don't know what physics they obey. They're simply incredible superconductors. But it's not just about this. So this is definitely one of those very popular topics. We still don't know theories for this. But there are so many more things which we do not understand.

Liam

例如,我们仍然没有——我们甚至无法模拟简单量子力学系统中的强关联。密度泛函理论——非常强大的工具。它在某些事情上极其准确。但一旦出现强电子关联,它就无法捕捉某些效应。我的意思是,有太多东西需要发现。这就是为什么我对 AI 在这个行业的应用感到兴奋,因为理论上、计算上和实验上,我们需要打开的还有很多。我们在这里所做的每一项改进,都改善了人们的生活,因为,你知道,我们越理解材料和固体物理,我们就能为它们创造最好的设备。

For example, we still haven't— we can't simulate even strong correlation in simple quantum mechanical systems. Density functional theory— very powerful tool. It's incredibly accurate in some things. But as soon as strong electronic correlation arises, it is not able to catch some effects. I mean that there are so many things to find out. Here's why I am excited about application of AI in this industry, because theoretically, computationally and experimentally, to us much more is needed to open. And each improvement we here we do, has improved life of people, because, you know, what we understand better materials and physics of solid body, the best devices we have we can for them create.

Liam

菲尔·安德森有一句名言“多者异也”,其本质是,理解某事物的所有基本定律是可能的,但当你加入许多元素时,它们的行为在质上不同于简单物理所应暗示的。所以理解许多物体如何一起行为,看起来像是简单的事情。好像基本定律很简单,但集体行为如此复杂,以至于真的很难模拟。这看起来像是涌现和多样性,简直不可思议。我们在深度模型学习中也看到了这一点,对吧?我们在深度学习中到处看到幂律。我们不理解这个,但它非常让人想起物理,那里有涌现和多样性。

There is a famous quote from Phil Anderson "more is different", the essence of which is that it is possible to understand all the basic laws of something, but when you add many elements, they behave qualitatively different than simple physics should have suggested. So understanding how many objects behave together seems like something simple. As if the basic laws are simple, but collective behavior is so complicated that it's really hard to simulate. This looks like emergence and versatility, which is simply incredible. And we see this also in deep models learning, right? We see the power laws everywhere in deep learning. We don't understand this, but it is very reminiscent of physics, where there is emergence and versatility.

Host

在物理世界里,有很多关于……涌现的理论工作。你知道在哪。当然。

In the physical world there was a lot of theoretical work about the emergence of... You know where. Certainly.

过程与最终结果 Process versus final result

Host

是的。呃,我另一个额外的问题是关于过程本身,我们姑且这么叫,是的,过程和最终结果。假设你想得到你追求的某些材料属性,然后你必须找出像以前一样达到的过程。

Yes. Uh, and my other additional question concerned the process itself, let's call it, yes, the process and the final result. Let's say you want to get certain properties of materials to which you strive, and then you have to find out the process as before to reach.

可逆工程与代理模型 Reversible Engineering and Surrogate Models

Host

分享这些东西、训练模型,理想情况下能逆向工程任何过程,以任何理论最终状态为目标,这值得吗?这说得通吗?

Is it worth it to share these things, to train the model, which is ideally reversible engineer any process, with the aim of any theoretical final state? Does this make sense?

Liam

也许吧,回到我们之前关于原子数量超过阿伏伽德罗常数的讨论,我们永远无法确切知道发生了什么。我认为这就是为什么我们永远无法理想地逆向工程一切。但我们只是试图做到足够好,以改善材料的特性。

Perhaps, returning to our conversation about that since atoms more than a number Avogadro, we never we won't know for sure the context of what happened. I think there is. the reason why we we will never be able to ideally reverse-engineer everything. But we just we are trying to do that's good enough, to improve characteristics material.

Host

是的。你问是否存在一个直接模型,输入传入数据就能预测周末?你想要……是的,你想创建一个代理模型,能够理解……输入数据——这些原子构型的最终状态,而预测是……分工,可以这么说。

Yes. You ask whether there is a direct model, where you enter the incoming data and you can predict the weekend? And you want... Yes, you want create a surrogate a model that can understand... input data—this the final state of this atomic configurations, and the forecast is... to divide the work, yes so to speak.

Liam

这是可能的,例如,一个团队负责过程,另一个负责性质,然后在他们之间安排竞争。是的,我认为我们可以在将这些代理工作过程划分到不同领域方面取得重大进展。我认为我们还可以在预测合成方面投入大量实证工作。

It is possible, for example, so that one the team was engaged process, and the other— properties, and just arrange competition between them. Yes, I think we we can achieve significant progress in in terms of dividing these agency workers processes into different spheres. I think that we we can also to spend a lot empirical work regarding forecasting synthesis.

Host

是的。嗯,你知道,利用这些原子构型,工具可以使用吗?比如,有了这些实验数据和之前的出版物或文章,你真的达到这一步了吗?毕竟你检查了,你知道,你能做到吗?玩还是不玩?我这样做的原因,我说,也许甚至是为了把这看作一家公司。你知道台积电是一种半导体工厂,但他们不开发这些半导体。你可以在材料领域成为台积电,人们来了,说他们想要什么,你定义过程,然后……

Yes. Um, well, you know, taking this configuration atoms that tools can be use? As, having these experimental data and previous publications or articles, you really Are you getting to this point? AND after all you you check, you know, Are you able to do this? to play or not? The reason why I do this I say, maybe even in in order to think about this as about company. You know how TSMC is a kind of factory for semiconductors, but they are not develop these semiconductors. You could become TSMC in in the field of materials, where people come and they say what they want, and you define the process and...

Liam

是的,这就像编译器的问题。是的。也就是说,给定这些要求……就像《星际迷航》里那样。是的,是的,是的。但是,考虑到这些要求,这是真实有效的构型集吗?它能存在吗?目标函数?不,我知道。这太多了,规模太大了。是的。我的意思是,我们谈论的是我们作为超级智能合成的一项任务。所以,我认为这与我们的愿景一致。

Yes, it's like a compiler matter. Yes. That is, given these requirements... Like in "Star Trek." Yes, yes, yes. But, with considering these requirements, is this real valid set configurations? Can Does it exist? Objective function? Not I know. This is too much, too large-scale. Yes. I mean that we are talking about one from our missions as a superintelligence synthesis. So, I think, this coincides with ours vision.

相变与学习 Phase Transitions and Learning

Host

是的。我想更详细地探讨你之前提到的一点,即关于相变。我认为这涉及到 Doji 关于涌现的评论。相变的概念,我想对许多听众来说可能很陌生。所以你能解释一下什么是相变,以及为什么它对强化学习环境之类的来说是有用的信号吗?

Yes. I want more details to dwell on, what did you say a little bit earlier, namely about phases. And I think this regarding the comment Doji about what it is Emergence. The concept of phase transition, I think, maybe to be a stranger to many in the audience. So can you explain what it is phase transition and why is this useful signal for something like learning environment with reinforcement?

Liam

相变非常迷人。我认为最容易理解的例子是冰融化或水沸腾。你提高冰的温度,它看起来仍然像冰一样,但在某个温度下它停止加热并变成液体。所以,它突然从固态相变为液态。另一个例子,与我们非常相关——这是伊辛模型。我认为计算机科学和数学也在研究这个,也许用其他名字,但核心是——自旋向上和向下,它们有不同的条件相互作用,如果都向上或一个向上一个向下。在相当高的温度下,它们通常简单地随机向上或向下。熵获胜。你降低温度,一切看起来都一样,但在某个点它们突然理想地组织起来。是的,相变。相变非常方便,因为它们使物理学家更容易研究某些现象。它们还具有某些空间相关性,对我们帮助很大。所以,在我们的相变实验室中,相变通常发生在我们混合前驱体——实际上是不同的晶体,然后提高温度或做其他事情来促使它们反应。然后原子开始反应并形成新的晶体。它在射线照片上显示出来,因为通常,如果原子几何形状急剧变化,那么射线照片会显著变化。但实际上,在典型的实际发现工作中,当你第一次尝试某样东西时,它不起作用,结果不仅仅是“是”或“否”,通常是混合相。其中通常有一点前驱体,可能有一点非晶相,还有一堆你可能不打算创建的相。这成为一个挑战,这就是为什么我们很难将模拟和人工智能结合起来进行特性表征。所以模拟总是可能说:“哦,这个相不像我们预测的那样,但它是它的一个小变化”。允许我现在花时间。稳定 2D 晶体的功率计算场。或者人工智能可以说:“考虑到之前的实验和这个,这可能不是我们想要得到的正确相,所以让我们改变条件。”

Phase transitions are fascinating. I think that most understandable for people an example is melting ice or boiling water. You increase ice temperature, he still looks how ice behaves like ice, but at a certain temperature it stops heating and turns into liquid. So, he suddenly passes from the solid phase into liquid. Another example, very relevant for us—this is the Ising model . I think computer science and mathematics too studying this, perhaps, under other names, but at the core—the back up and down, and they have different conditions interactions, if both directed up or one up, and the other down. At quite high temperature they usually simple randomly directed up or down. Entropy wins. You are lowering temperature, and everything looks the same, but at some point they suddenly ideally are being organized. Yes, phase transition. Phase transitions are very convenient because they make it easier for physicists the study of certain phenomena. They also have certain spatial correlations that we They help a lot. So, in our phase laboratories transitions usually occur when we mix precursors —actually different crystals, and then increase temperature or we do something else to encourage them to reactions. Then the atoms are starting to react and form a new one crystal. It displayed on radiographs, because, as a rule, if geometry atoms sharply changes, then radiograph changes significantly. But in fact during typical practical discovery work materials when you trying something for the first time, it doesn't work, and the result is not just "yes" or "no", and usually very mixed phase. In it there is usually a little precursors, maybe to be a little amorphous phases, and also a bunch of phases which you may not were going to create . This becomes a challenge, and that's why we had very difficult to combine simulations and artificial intelligence for carrying out characteristics. So simulations always they might say, "Oh, this phase is not like what we predicted, but it's her little one variation". Allow I need to spend time now. power calculation fields for stable 2D crystal. Or artificial intelligence can say: " Considering previous experiment and this, this, probably not the right phase, which one we want get, so let's "Let's change the conditions."

Host

是的,对听众来说:XRD——这是 X 射线衍射,你已经描述过了。

Yes, and for the listeners: XRD —this is an X-ray diffraction, which you already described.

Liam

是的,在某些情况下会出现一种相,它以前从未被固定过。它不在任何文章或数据库中,因此人工智能必须使用这些工具来说:“那么,什么原子构型可以解释这些相?”但这对控制科学过程变得极其重要,因为,假设我们有目标相,我们似乎需要爬到山顶才能获得更好的测量。所以如果它是 1%,我们想提高相纯度,我们想要大量精确的测量这些实验数据来说什么存在什么不存在。所以缺点之一是,你只能清楚一个你可以观察的变量。

Yes, some cases appears a phase that is simply nowhere never been before fixed. She is not there. in any articles or databases, therefore artificial intelligence has use these tools to to say: "Well, what atomic configuration could explain these phases?" But it becomes incredibly important for control scientific process, because , let's say we have target phase, and we it seems necessary to climb to the top mountains to get better measurements. So if it's 1%, and we we want to increase phase purity, we we want to have a lot precise measurements these experimental data to say that what is present and what is not . So one of disadvantages are that in you are only clear a variable that you can watch.

不确定性与化学直觉 Uncertainty and Chemical Intuition

Host

回到关于你如何应对不确定性的陈述,似乎你答案的一部分是你确实观察到某种明确的特征。这里没有解释的余地。对吗?这就是背后的逻辑,你只是需要发现一个新相,因此这让你担心什么?我的意思是,在许多这些情况下仍然存在不确定性。

Returning to statement about how are you coping with uncertainty, seems to be part of your answer is that you do observed signature somewhat unambiguous. Here there is no room for interpretations. That's right. ? This is the logic that stands behind this, do you just needed to be found a new phase, and therefore this What worries you? I mean, in many of these cases still remains uncertainty.

Liam

嗯,你知道,很简单的事情是复制。当然,我们正在进行复制。但不,我认为这是一个复杂的任务,因为它不是完全确定性的,而且存在某些模糊性和其他事情。两个不同的相实际上不能对应同一个图案。所以我认为对于系统来说,使用化学直觉非常重要,说:“好吧,考虑到合成条件和这些先验知识,这极不可能。”这可以用来帮助消除模糊性。

Well, you know, pretty simple The thing is to replicate. Of course, we are conducting replicate. But no, I I think it's complicated. task, because it is not complete deterministic, and there are, well, certain ones ambiguities and others things. Two different phases actually can't to answer one and the same pattern. And so I think that for systems are very important use chemical intuition to say: "Okay, this is it would be extremely unlikely, considering the conditions synthesis and these prior knowledge." AND this can be used, to help eliminate ambiguity.

Host

所以,你在某种程度上输入了先验数据。基于此,你期望什么?

So, you are to some extent enter here a priori data. Based on that, What do you expect?

Liam

是的,我认为……热力学是最重要的先验因素,还是错的?然后物理学也是重要的先验因素。然而我们这里有一件事帮助很大,那就是材料特性的多模态性。

Yes, I think... Thermodynamics is most important a priori factor, or wrong? And then physics is also important a priori factor. Yet one thing we have here It helps a lot, it is multimodality characteristics materials.

多模态表征 Multimodal Characterization

Liam

我们可以进行 X 射线衍射(XRD),有些相在衍射图上可能看起来相似,但我们也可以测量电学和磁学性质。我们可以借助电子显微镜测量它们的形貌。然后,一些在 XRD 上看起来相似的相,在其他测量中会显得不同。所以多模态确实很有帮助。而且,AI 在这里非常有用,因为人的上下文和算力有限。如果你同时给一个人 10 种不同的模态,并让他们按顺序分析所有模态,这有点困难。但对 AI 来说,这并不真的需要超级智能。这很棒,是的。它就像超级智能,只是能执行许多计算。它能查看所有。这引导我们讨论在不确定性下更好的决策,系统长时间整合来自许多不同工具的信号,通过各种实验和重复,创建更深入、更好的模型,而不是过于依赖一个设备的一次测量。

We can conduct X-ray diffraction (XRD), and some phases may look similar in the diffractogram, but we can also measure electrical and magnetic properties. We can measure their morphology with the help of some electron microscopy. And then some phases that look similar on XRD will look different in other measurements. So multimodality really helps. And again, AI is very useful here, because people have limited context and computational power. So if you give a person 10 different modalities simultaneously and ask them to analyze them all in sequence, that's a little difficult. But for AI, it doesn't really require superintelligence. And that's great, yes. It's like a superintelligence, it just can perform many calculations. It can view all. It leads us to phrases about better decision-making under uncertainty, where the system integrates signals from many various tools over a long time, through various experiments and replicates, and creates a deeper, better model, instead of relying too much on one measurement with one device.

数据过载与选择 Data Overload and Selection

Host

让我补充一下,你一段时间以来一直在说,数据量简直超过了任何智能计算机或处理器的容纳能力。所以在某个时刻,你必须拒绝部分数据。即使你有所有这些重复,即使你可能有 10 种不同的方法来精确测量同一件事。因为如果你的一个传感器温度错了呢?所以当你丢弃这些数据时,你如何决定怎么做?

Let me add, you have been saying for some time that the data is simply more than any smart computer or processor can accommodate. So at a certain moment you have to reject part of the data. Even if you have all these replicates, even if you have probably 10 different ways to measure the same thing exactly. Because what if one of your sensor temperatures is wrong? So when you are discarding this data, how do you decide what to do?

Liam

作为一名来自机器学习的专家,渴望数据,我想要一切,对吧?然后我就,你知道,我建模,我研究一切。我不知道我是什么。我不知道,我只是把车扔向它。是的,是的。所以我们不会丢弃数据。嗯,我想也许我应该意思是,如果上帝在观察实验,数据会比人们能放进计算机的更多。不幸的是,我们凡人无法观察所有原子。所以即使数据超过我们能保存的,我们并不真的必须访问。例如,我无法追踪所有原子的位置。嗯,所以不幸的是,我们仍然处于任何收集的数据都非常宝贵的情况,我们不会删除任何东西。也许一种思考方式,与 AI 方面联系,是思考探测,就像你在探测基础模型时,对吧?这个模型有所有这些数据。因此,上帝有完整的图景,他做什么材料和实验——这个小线性探针,在输出上有八位或类似的东西。这就是为什么你拿这个巨大的东西,粗略地构建她的碎屑信息。现在这要么是传统的人类工作,要么现在是你的工作代理——从那些八位信息中重建实际上正在发生的事情。正如我所说,类比,你无法完全看到每个参数,每个激活对于这个入口。嗯,是的,你得到一些数据子集。或者一些,嗯,你知道,概括所有这一切的迹象。是的。

As a specialist from machine learning, thirsty for data, I want everything, truth? And then I just, you know, I model, I study everything. I don't know what I am. I don't know, and I just I'm throwing a car at it. Yes, yes. So we don't throw away the data. Um, I think maybe I should have meaning that if God was watching the experiment, it would be more data than people can fit in a computer. Unfortunately, we mortals cannot watch all atoms. So even if the data is more than we can save, we don't really have to access. For example, I can't track the position of all atoms. Um, so unfortunately, we are still in situations when any data collected is very valuable, and we don't delete anything. Maybe a way to think about it, to connect with the AI side, is to think about probing, as when you are probing the base model, right? This model has all this data. Therefore, God has the complete picture of what he does material and experiment—this small linear probe, which has eight bits or something like that on exit. And that's why you take this huge thing and roughly build her to the crumb information. And now this is either traditional human work, or now your job agent—to reconstruct from those eight bits information that actually is happening within. As I say, analogy, where you can't have full visibility each parameter, each activation for this entrance. Hmm, yes, you get some subset of data. Or some, um, you know, generalized signs all of this. Yes.

麦克斯韦妖类比 Maxwell's Demon Analogy

Host

我猜你可能觉得麦克斯韦妖的论点很有趣,因为你能讲讲吗,因为我……他之前提到过。我实际上忘了。

I assume that you probably think Maxwell's demon argument is very interesting because can you just tell it, because I no... He mentioned it earlier. I actually forget him.

Liam

嗯,你知道,在 19 世纪,人们意识到系统的熵必须增长,或者,好吧,对于封闭系统,保持不变,但它不能减少。这是热力学第二定律。有一个这样的想象实验:如果存在一个微小的全知恶魔,他能观察原子,每当一个原子能量低时,就打开一扇门,让它进入另一个房间,并且一直这样做,以减少系统的熵,这将违反热力学第二定律。很长一段时间,我认为没有非常令人信服的答案来解释为什么会这样。但后来出现了更现代的科学工作,德兰道尔提出了一个论点,即要删除信息需要浪费能量,而麦克斯韦妖通过你的行动将获得这么多信息,他必须部分删除,在信息删除过程中,他应该花费能量,这将保证熵的增加。所以虽然我们还没有达到麦克斯韦妖的水平,我认为你的问题非常有远见,关于未来,我们将保留如此多的数据,以至于我们将不得不部分删除。

Well, you know, in the 1800s people realized that the entropy of a system has to grow or, okay, for closed systems, stay unchanged, but it cannot decrease. This is the second law of thermodynamics. There was such an imaginary experiment: what would happen if there existed a tiny omniscient demon who could watch atoms, and every time when an atom had low energy, would open a door, letting it into another room, and would do this all the time, to reduce the entropy of the system, which would violate the second law of thermodynamics. And for a long time I don't think there was a very convincing answer why this is happening. But then more modern scientific work appeared, de Landauer put forward the argument that to remove information you need to waste energy, and Maxwell's demon through your actions would have access to this amount of information that he would have to partly delete, and during information deletion he should spend energy that would guarantee an increase in entropy. So although we are not yet at so equal to be by Maxwell's demon, I think that your very question is far-sighted regarding the future, where we will keep so much data that we will have to partly remove.

传感器饱和与可见性 Sensor Saturation and Visibility

Host

是的,不,当然不是。我指的是类似这样的规模。这让我想起我有一份工作。空调,因为那正是它的工作方式。但我也认为这个“为什么”问题就是不要买下世界上所有的传感器,不要让所有测量设备过度饱和,不要连续复制一切。我们怎么做?对。好的,是的。好的,是的。是的,是的。对。好的。是的,我认为是这样。如何的问题,是的,没错。

Yes, no, of course not. I meant something like this scale. This reminds me of I have a job. air conditioner, because That's exactly how it works. But I also think this "Why" questions just don't buy it all the sensors in the world and don't oversaturate everything measuring devices, not duplicate everything in a row ". How do we do this? Right. Okay, yes. Okay, yes. Yes, yes. Right. OK. Yes, I think so. the question of how Yes, that's right.

Liam

你获得单个原子的可见性,那是不可能的,但我们添加了大量的遥测数据,这实际上减少了隐藏变量,如果你愿意的话。是的。是的,只是出于好奇:你是否担心你只在加利福尼亚?你不想吗?例如,在尼泊尔或澳大利亚做这个?是的,我们可以。那就是我们有计划开设更多实验室。因为,当然,重力和你所在的位置,以及你在哪里,都有价值。专业知识也可能不同,是吗?比如这里,我们有一些经验,例如,斯坦福的物理和工程系有某些知识,我们真的很喜欢帮助,因为我们可以从那里雇佣人,我们的资源可以合作。正如你所说,地理可以有价值,但在不同的城市,物理、化学领域的经验可能会有差异。是的,嗯,我们有,这是 Zoom。我心中有数,但仍然,你的实验会受到你所在位置的影响。是的。是的,我有。意思是未来的周期表不是一个实验室。我认为我们需要联系你的朋友杰夫飞入太空和其他东西。是的,我认为是这样。有必要有这样的实验室网络,并且,正如刚才提到的杰夫,对于每个实验室,根据经验或,比如说,材料限制,有最佳位置。有许多不同的限制。嗯,好的,其中之一……

You gain visibility individual atoms that impossible, but we add a huge amount of telemetry, and this is actually reduces the number hidden variables, if you want. Yes. Yes, and just with curiosities: whether What worries you is that you are only in California? Don't you want to? to do this, for example, in Nepal or Australia? Yes, we could. That is we have plans open more laboratories. Because, of course, gravity and where you are directed, and where you are, have value. Also expertise may differ, Yes? Like here, we have some experience, for example, physical and engineering departments of Stanford have certain knowledge, which we really like help because we can hire people from there, our resources can cooperate. As you said, geography can have value, but in different cities can there may be differences in experience in the field physics, chemistry. Yes, well, we have for This is Zoom. I have on mind, but still, for your experiments will be affected by where you are are you. Yes. Yes, I have. meaning that the future Periodic is not one laboratory. I think we need to contact your friend Jeff fly into space and something else. Yes, I think so. it is necessary to have such a network of laboratories, and, as just noted Jeff, for each of them is optimal place depending on experience or, let's say, restrictions in materials. There is many different restrictions. Um, okay, one of them...

集成计算工具 Integrating Computational Tools

Host

比如说,我感兴趣的是,你如何使用计算工具?或者你如何将计算工具整合到工作流决策中?特别是当,比如说,计算和实验不一定一致时。嗯,啊,还有从理解角度的高层理论。那么,这一切如何协同工作?

Let's say I'm interested, how do you work with computational tools? Or how you integrate computational tools in workflow decision making? Especially when, let's say, calculations and the experiment is not necessarily coincide. Hmm, ah also theory from understanding view high level. And so, How does it all work? together?

Liam

是的,看这里。这很有帮助,有些事情在模拟中更容易、更精确地执行,但有些事情在实验中更容易、更准确。一般来说,当然,实验更准确,但某些测量很难通过实验完成。这就是为什么你做一个失败的版本,你得到不准确的数据。嗯,让我。举几个例子。

Yes, look here. It helps a lot that there are things that are easier and more precisely, to perform in simulations, but there are things, which are easier and more accurate to perform in experiment. In general, of course, experiments are more accurate, but certain measurement is very hard to do experimentally. That's why you do a failed version, and you get inaccurate data. Um, let me. to give a few examples.

DFT简介 Introduction to DFT

Host

那么,其中之一——我们使用密度泛函理论来估算材料的生成焓。抱歉,稍等。你能解释一下密度泛函理论和信息熵吗?

So, one of them—we use density functional theory for estimation of enthalpy of formation of your materials. Sorry, small pause. Can you explain the theory density functional and information entropy?

Liam

是的,当然。抱歉,生成焓,是的。是的,密度泛函理论——这可能是材料领域最常用的模拟方法。它源自 Kohn 和 Hohenberg 的理论,Kohn 因此获得了诺贝尔奖。他明白量子力学非常昂贵,因为你要处理每个波函数的指数级希尔伯特空间。因此,必须解决一个非常困难的任务,尤其是随着系统尺寸的增加。Hohenberg 和 Kohn 明白,你其实不需要处理这个指数空间。你需要考虑的一切就是电荷密度,至少对于基态量子力学系统而言。所以,这个定理其实相当简单。我想连物理学生都能理解。嗯,但想法是,基态量子力学系统的所有性质都是电荷密度的泛函,这简直不可思议。因为电荷密度只是一个三维对象,而希尔伯特空间和波函数是指数级的。所以,这是一个非常有趣的观察,然后 Kohn 与他的博士后 Sham 发表了另一篇文章,关于 Kohn-Sham 波函数,他们提出了一种实用的方法来求解材料量子力学性质的近似。今天我们在实验室以及实验室之外非常积极地使用这个方法。生成焓——这是我们可以做的一种测量,用来理解材料的能量。也就是说,本质上,我们想问:这种材料的能量是多少?如果这个能量高于这些原子可以转变成的其他材料,那么这种材料就不太可能被创造出来。所以稳定性意味着它位于其他材料及其能量的凸包上。嗯,就这些吗?已经足够清楚了,是否值得我们……我是说,也许更简单,一般结论。因此,密度泛函理论——这是一种方法,把随电子数或系统尺寸指数级复杂的东西,以一定的误差,近似地降低到 n 阶近似。嗯,所以现在你可以计算以前根本不可能计算的值,但要付出一定的代价。当然,如果我们能完美地做到,就不需要实验室了,但是,你知道,我们做不到。这可能不错。参考我们自己的,历史上的第二集。Heather Kulik 说过一句著名的话:对于没有材料,你的 AlphaFold,不仅在计算上,而且因为真实数据根本不存在。我们并不真正知道晶体材料结构是什么样子,而 DFT 通常是我们最好的方法……是的,我想是这样。我是说,我要补充两点:DFT 按定义不一定是近似。Kohn-Sham 定理表明它可以是精确的,但是,正如你指出的,我们还没有获得交换关联能泛函。而对于基于电荷密度的模型,我们不知道动能泛函是什么样子。即使我们有完美的 DFT,我仍然认为我们需要实验室,因为即使有理想的 DFT,我们可以在计算机中拟合 10 的 23 次方个原子。你现在非常执着于这一点。你知道,好吧,好吧。是的,但是随着指数……随着计算能力的增长,人们可以声称,当 n 立方缩放时,最终你基本上可以计算一切。只要等到……哦是的,我们需要足够强大的计算机,是的。是的,是的。但在 n 立方缩放之前,你可以……作为一个非科学家,我会做一个观察。我不知道你是否想……这是讨论还是不是?我来自金融界。我们有高斯 copula,这是一种评估成本的方法,例如通过相关性评估信用违约互换,并将一切简化为单一核心。这在精神上与 VAE 非常相似,你计划将所有这些事情简化为唯一一个参数。我想知道这是否不是到处都一样的把戏。你的方法听起来比我的更多维。但 VAE 只是,你知道,一个 sigma,一个 sigma。但听起来一样吗?是的,是的,密度泛函理论仍然有足够的电荷,很好。是的。没错。你有点过于概括了,但这很好。实际上,这是最重要的时刻。有点。就个人而言,我不知道预测新材料稳定性的最佳方法。它肯定不理想,但肯定比我记得的其他方法更好。所以,回到你的问题,我们使用它们的原因是一些事情模拟比实验更容易,而且它们很好地互补。然后我们可以循环地做这件事。我们相信仅靠模拟永远不够,但在模拟循环、AI 和实验中,我认为我们可以比以前快得多。允许模拟的一件事——是可扩展性快得多。我的意思是,现在启动一台计算机比开设新实验室容易得多。你如何避免,或者,假设,你正在扩大 DFT,你如何避免你的模型过于集中于此,同时保持合理性,因为现实世界最终是真理?你对哪个感兴趣?如果你,你知道你如何……嗯,我们也有,也有一个实验室。是的,你在扩大实验室,但我的意思是,你可以想象将 DFT 扩展到数百万、数亿次现代计算能力,而在你的实验室里,我假设你每天做数百、数千次?我不知道。嗯,即使在非常高通量的情况下,你的能力仍然相当有限,对吧?是的,我的意思是,人们以及实验室都清楚 DFT 的局限性。例如,即使我们可以如你所说花费 1 亿次 DFT 测试,我们也永远无法检查它们的微观结构。微观结构是指 DFT 通常模拟完美晶体,但晶体实际上并不理想,它们的微观结构,即中程有序结构,可以显著影响性质。其他问题当然在于,一些性质,如高温超导,不可能用 DFT 轻松建模。再次是的,即使你花费 1 亿次 DFT 测试,这也不够。因此,我们必须依赖启发式方法、实验等。所以,在实验室实际执行之前,所有这些计算有一个巨大的过滤器。这是由人、LLM 完成的。这是混合的吗?它可以是混合的。

Yes, of course. Sorry, enthalpy, yes. Yes, the density functional theory — this is probably the most used simulation method for materials. It comes from the theory of Kohn and Hohenberg, for which Kohn received the Nobel Prize. He understood that quantum mechanics is very expensive because you are dealing with exponential Hilbert space for every wave function. Therefore, it is necessary to solve a very difficult task, especially with increasing size of systems. Hohenberg and Kohn understood that you don't really need to work with this exponential space. Everything that you need to think about is the charge density, at least for ground state quantum mechanical systems. So, this theorem is actually quite simple. I think even a physics student will understand. Um, but the idea is that all properties of ground state quantum mechanical systems are functionals of charge density, which is simply incredible. Because charge density is only a three-dimensional object, while Hilbert space and wave functions are exponential. So, it was a very interesting observation, and then Kohn posted another article with his postdoc Sham, about the Kohn-Sham wave functions, who proposed a practical way to solve approximations of quantum mechanical properties of materials. And today we very actively use this in our laboratory, and beyond its borders. Enthalpy of formation — this is one of the measurements that we can do to understand the energy of a material. That is, in essence, we want to ask: what is the energy of this material? If this energy is higher than other materials into which these atoms can turn, it is unlikely such material will be able to be created. So stability means that it lies on the convex hull of other materials and their energies. Hmm, is that it? It was enough to make it clear whether it is worth it to us... I mean, maybe simpler, general conclusion. Therefore, density functional theory — this is the way to take something that is exponentially complicated depending on the number of electrons or system size, and reduce it to approximately, with a certain error, to the nth approximation degree. Um, so now you can calculate values that were previously simply impossible to calculate, but for a certain price. Of course, if we could do that perfectly, we wouldn't need a laboratory, but, you know, we can't. This could be good. By reference to our own, the second episode in history. Heather Kulik said she is famous for saying: for no materials is your AlphaFold, not only computationally, but also because the true data simply does not exist. We don't really know what crystalline material structures look like, and DFT is generally our best way to... Yes, I think so. I mean, I'll add two things: DFT by definition does not have to be an approximation. The Kohn-Sham theorem shows that it can be accurate, but, as you noted, we don't have access to the exchange-correlation energy functional yet. And for models based on charge density, we do not know what the kinetic energy functional is like. And even if we had perfect DFT, I still think that we need a laboratory, because even with ideal DFT we can fit 10 to the 23rd power of atoms in a computer. You are very fixated on this at the moment. You know, okay, okay. Yes, but with exponentially... With growth of computational capacities, one can claim that when n-cube scaling, eventually, you will be able to basically calculate everything. Simply wait until... Oh yes, we need big enough powerful computers, yes. Yes, yes. But until n-cube is good scaling, you can... As a non-scientist, I will make an observation. I don't know if you want... Is this a discussion or not? I came from finances. We had Gaussian copula, which was a way of assessing cost, for example, credit default swaps through correlations and reducing everything to a single core. This is very similar in spirit to VAE, you plan to reduce all these things to the only one parameter. I wonder if this is not the same trick everywhere. Your method sounds a little more multidimensional than mine. But VAE is just, you know, one sigma, one sigma. But does it sound the same? Yes, yes, density functional theory still has enough charge, great. Yes. True. You are a little generalizing, but this is pretty good. It actually is the most important moment. Sort of. Personally, I don't know the best method for prognostication of stability of a new material. It is definitely not ideal, but definitely better than other methods that I can remember. And so, returning to your question, the reason why we use them is that some things are easier to simulate than to experiment, and they are good complements to one another. And then we can do this cyclically. We believe that simulations alone will never be sufficient, but in the simulation loop, AI and experiments, I think we can move significantly faster than before. One thing that allows simulations — is scalable much faster. I mean that now it is much easier to start a computer than to open new laboratories. How do you avoid, or, let's say, you are scaling up DFT, how do you avoid your models being too concentrated on this, while remaining justified by the fact that the real world is ultimately the truth? Which one are you interested in? If you, you know how you... Well, we also have everything, there is also a laboratory. Yes, you are scaling the laboratory, but I mean, you can imagine scaling DFT to millions, hundreds of millions for modern computational capacities, while in your laboratory, I assume you make hundreds, thousands per day? I don't know. Well, even in the case of very high throughput, you still have the ability quite limited, truth? Yes, I mean that people, as well as laboratories, are well aware of DFT constraints. For example, even if we could spend as you said 100 million tests with DFT, we could never check their microstructure. Microstructure is the idea that DFT usually models perfect crystal, but crystals are actually not ideal, and their microstructure, that is, the structure of medium order, can significantly influence properties. Other problems, of course, consist in the fact that some properties, such as high-temperature superconductivity, are impossible to easily model using DFT. Again yes, even if you spend 100 million DFT tests, this will not be enough. Therefore, we have to rely on heuristics, experiments, etc. So, there is a huge filter from all these calculations before what actually is performed in laboratories. It's done by people, LLM. Is this a mixture? It could be a mixture.

材料科学中的实验校准 Experimental Calibration in Materials Science

Liam

例如,很长一段时间里,Materials Project——最好的开源 DFT 数据库——用实验数据来校准 DFT。所以他们实际上用实验数据来校准 DFT。我们公司内部也这么做。当我们对某个化学体系感兴趣时,我们会做实验和模拟,然后进行校准。

For example, for a long time the Materials Project, the best DFT database with open source, used experimental calibration for DFT. So they actually used experimental data for DFT calibration. We do it here too, within the company. When we are interested in a certain chemical system, we conduct experiments and simulations and we calibrate them.

Liam

这是另一点。我跟很多生物学家朋友聊过,他们总是很惊讶,当你说在材料科学里你不能只做 XRD 结构分析就弄清楚它到底是什么。因为在生物学里你可以看晶体蛋白结构,通常能精确到埃米级重建。材料科学的 XRD——X 射线衍射——抱歉,衍射——实验 XRD 与生物学相比,差别在哪里?我们丢失了哪些信息,为什么那是一种从损失中来的投影?

This is another point. I talked to many biologist friends, and they are always very surprised when you say that in materials science you cannot just do XRD structure and find out what it actually is. Because in biology you can look at crystalline protein structure, and usually it is being reconstructed from accurate to angstrom. Where is the difference between XRD for materials—X-ray diffraction—diffraction, sorry—experimental XRD for materials compared, say, to biology? Which information are we losing and why is that a kind of projection from losses?

Host

很难不放弃,疯狂地回应这个问题。但难道不是吗?在你看来,有机化学和生物学几乎来自一个 VC 维更低的模型,也就是按柯尔莫哥洛夫复杂度更低?我们似乎能把它表示成一维序列,几乎可以编译。所以,关键在于按柯尔莫哥洛夫复杂度更容易看出复杂性。一维?我是说 DNA 或 RNA——它只是序列。大多数人会说这更像是二维的,但我不是生物学家,所以我不确定他们是对的。

It's hard not to give up, crazy, responding to this question. But isn't it? Does it seem to you that organic chemistry and biology almost come from a model with lower dimension VC, that is, lower complexity by Kolmogorov? We seem to be able to present it as a 1D sequence, which is almost possible to compile. So, the point is that it's easier to see complexity by Kolmogorov. 1D? I mean, DNA or RNA—it's just sequence. Most people would say that this is rather 2D, but I'm not a biologist, so I'm not sure that they are right.

Liam

但在无机化学里,一切都是化学,否则就只是疯狂。这显然不是来自低柯尔莫哥洛夫复杂度的模型,因为这些三维无机晶体可能包含金属、绝缘体、超导体、金刚石。这都是真的吗?差异大得难以置信。人们尝试过,但我们从未找到一维表示。例如,很难为三维无机晶体找到 SMILES 字符串。原子之间相互作用的方式很奇怪。例如,共价键、离子键、金属键。

But with inorganic everything is chemistry otherwise, it's just madness. This clearly does not come from a model of low Kolmogorov complexity, because these 3D inorganic crystals may contain metals, insulators, superconductors, diamond. Is this all true? Incredibly different. People tried, but we could never find a one-dimensional representation. For example, it is very difficult to find a SMILES string for three-dimensional inorganic crystals. And the way atoms interact with each other is strange. For example, covalent bonds, ionic bonds, metallic bonds.

Liam

但是的,这非常有趣。我有时会想——你研究过固态电解质的问题吗?为什么人们不能用固态电解质替代电池中的液态电解质?那里形成的都是这些枝晶。本质上,锂形成结构,穿透固态电解质并破坏它。你看这个,你觉得与生物学相比它有多原始?毕竟,在生物学中这些系统能够实现如此复杂的纳米技术进展。而我们甚至无法让两个界面协同工作,因为一切都会断裂。尽管无机材料可能非常耐高温。用它们可以建造航天飞机。可以用硅进行计算,遵循摩尔定律。所以是的,这是一种妥协。

But yes, that's very interesting. I sometimes wonder—have you ever studied problems with solid electrolytes? Why can't people replace liquid electrolytes in batteries with solid? All are formed there—these dendrites. Essentially, lithium forms structures that penetrate into the solid electrolyte and destroy it. You look at this and how much do you think it is primitive compared with biology? After all, in biology these systems are capable of such complex nanotechnological progress. While we cannot even force two interfaces to work together, because everything breaks. Although, probably, inorganic materials can be very resistant to high temperatures. With them they can build space shuttles. Can use silicon for calculations according to Moore's law. So yes, this is a compromise.

Host

是的,我们和其他几位嘉宾讨论过这个。我认为对我来说主要区别之一是,在生物学中我们有一套可以从数百万年进化中借用的工具,进化已经给了我们一切,我们只需拿来用。但进化也限制了 VC 维,使其对蛋白质等来说足够低。像细胞这样的大规模系统中的涌现行为可能复杂得多。

Yes, we discussed this with a few other guests. I think one of the main differences for me is that in biology we have a set of tools that can be borrowed from millions of years of evolution, which has already given us everything, and we can just play with it. But also evolution limits VC-dimension, making it low enough for proteins, etc. Emergent behavior in large-scale systems such as cells may be much more complicated.

Liam

是的,还有意识。

Yes, and consciousness too.

Host

是的,但可能更具体地说:如果你看 X 射线,拍一张完美晶体的照片,照射它并得到光谱。它仍然不能提供关于晶体结构的唯一信息,对吧?

Yes, but possibly even more specifically: if you look at X-ray, take a picture of a perfect crystal, enlighten it and get a spectrum. It still doesn't work—unique information about what the structure of the crystal is, right?

Liam

不能。而这——什么?这只是平均行为。但例如,歧化——这是非常复杂的事情,对吧?也许你有一个完美晶体,平均来看 X 射线显示完美,但实际上有另一种模式,原子稍微向右或向左偏移,尽管平均来看它们在中间。

No. And this—what? This is just average behavior. But, for example, disproportionation—this is a very complicated thing, right? Maybe you have a perfect crystal, which on average looks perfect by X-ray, but actually there is another pattern, where atoms are a little shifted to the right or to the left, although on average they are in the middle.

Host

嗯,是的,这是艰难的生活,是的。那么蛋白质呢?你本质上知道某个蛋白质是什么,它服从某些规则。所以你可以将直接模型与此结合。

Hmm, yes, this is a hard life, yes. Well, what about proteins? You essentially know what a certain protein is, that it subordinates certain rules. So you can combine a direct model with this.

Liam

当然。然后你可以转向,比如说,单晶 X 射线衍射分析,这会有帮助。

Of course. And then you can move to the side, let's say, monocrystalline X-ray diffraction analysis, and this can help.

材料科学的复杂性 Complexity of Materials Science

Host

当我们推出科学播客时,有那么多元素:科学中的 AI、AI 与数学、AI 与物理等等。然后那里有点像一场竞赛,哪个更难?我认为有可能科学地证明材料科学是最难的。

When we launched the science podcast, there were so many elements: AI in science, AI and mathematics, AI and physics, and all that. And then there was something there like a competition, which one is more difficult? And I think it's possible to scientifically substantiate that materials science is the most difficult.

Liam

嗯,我们选择它不一定是因为复杂性。我的意思是,有很多领域确实更容易。我们有很棒的大型材料类模拟器,而在生物学中,模拟细胞、器官或整个生物体极其困难。所以,这是一个巨大的优势。

Well, we chose it not necessarily because of complexity. I mean, there are a lot of areas where it really is easier. We have wonderful simulators for large classes of materials, while in biology it is incredibly difficult to model a cell, organ or whole organism. So, this is a huge advantage.

Host

是的,这不是很有趣吗?至少所有这些都是等同一维或二维的?但然后出现了许多数量级和结构,这正是创造复杂性的原因。

Yes, isn't that interesting? Which is at least all of this is equal one-dimensional or two-dimensional? But then many more are appearing orders of magnitude and structures, and that's exactly what creates complexity.

Liam

这很奇怪。而材料中的复杂性,也许更多在于微观结构。

This is strange. While complexity in materials, perhaps more is in microstructure.

Host

是的,我是说材料在不同长度尺度上也包含不同的复杂性。所以是的,微观结构,而且不止于此。

Yes, I mean that the materials also contain different complexity on different scales of length. So yes, microstructure and not only.

梦想材料与科幻 Dream Materials and Science Fiction

Host

好吧,稍微换个话题,作为一种放松。如果我们谈论科幻,假设你可以创造任何更有价值的材料?你知道,我这里有个清单。显然,室温超导体,但还有,你提到了电池。我一直认为如果电池比锂离子更好,那你就走运了。你知道,它们已经存在一百年了。最后一个——碳纳米管,主要用于太空电梯。不知道这是否是想到的三个。或者在你的圈子里谈论的,真正的梦想是什么?

Well, to change it up a bit, as a kind of discharge. If we talk about science fiction, suppose you can create any materials that are more valuable? You know, I have here a list. Obviously, this superconductors at room temperature, but also, you mentioned batteries. I always thought that if the batteries will be better than lithium-ion, then you are in chocolate. You know, they already exist a hundred years. And the last one—carbon nanotubes, mainly for space elevators. Not I know if these are the three that come to mind. Or there are those about which speak in your circle, which is a real dream?

Liam

当然,超导体、磁体,特别是如果可能减少对稀土或难以找到的过渡金属(如钴)的依赖。减少电池中钴的数量。在我看来,根本上重要的一件事是,我们能否接近计算能效的兰道尔极限。

Certainly, superconductors, magnets, especially if possible to reduce dependence on rare earth or transition metals that are difficult to find, such as cobalt. Reducing the number of cobalt in batteries. One of the things that, in my opinion, is fundamentally important, is whether we can approach the Landauer limit for energy efficiency of calculations.

Host

你能解释一下什么是兰道尔极限吗?

Can you explain what the Landauer limit is?

Liam

是的,如果你看我们在一次浮点运算上花费多少能量,那么可以看到类似摩尔定律的行为,一切都变得指数级更高效。因此,我们每次浮点运算消耗的能量呈指数级减少。我的意思是,我最近没查过。是的,那是 10 年前。是的,那很棒。我记得当我第一次发现并研究它时。我们当时,怎么了?距离它大约 10 到 20 个数量级左右。而上次我看时,我们奇怪地接近了。我们在某个地方,我不知道,距离它三到四个数量级,或者实际上更近。很简单,多年来呈指数级发展。

Yes, if you look at how much energy we spend on one flop of calculations, then you can see behavior by type of Moore's law, where everything becomes exponentially more effective. Therefore, we spend exponentially less energy per unit flop of calculations. I mean, I don't checked it lately. Yes, it was 10 years ago. Yes, that's great. I remember when I first found out about it and investigated. And we were, what's up? Somewhere between 10 and 20 orders of magnitude further away from it or something like that. And the last time I looked, we are strangely close. We are somewhere, I am not I know, in three or four orders of magnitude from it, or actually even closer. It's simple, exponentially developed during for many years.

可逆计算与Landauer极限 Reversible Computing and Landauer Limits

Liam

是的,相当令人印象深刻。从我打开它那一刻起,大约 10 或 15 年前?我们必须做的一件事就是散热。因为当这些计算被执行时,热量会释放——不可避免——但我们必须为此留出余量。所以,如果我们在计算效率上接近兰道尔极限,对人类来说将是惊人的,对吧?因为这意味着我们以最高能效进行计算——这可能是人类最重要的事情之一——因为能量是宇宙真正的货币。

Yes, quite impressive. From the moment I opened it, around 10 or 15 years ago? One of the things we must do is dissipate heat. Because when these calculations are performed, heat is released—inevitably—but we must allot for it. So if we get closer to Landauer limits in efficiency of calculations, it will be amazing for humanity, right? Because it means we perform calculations—probably one of the most important things humanity does—at maximum energy efficiency, because energy is the real currency of the universe.

Host

这会影响量子计算吗?只是随便抛个想法。

Does this affect quantum computing? Just throwing an idea in the air.

Liam

理论上,这是大规模、极度并行的计算。所以正是如此——它们受到翻转计算的影响。但说实话,我不是专家。我不理解。如果你能做可逆计算,你就不需要在过程本身上花费能量,因为你不删除任何信息。它们是可逆的。但我不知道这是否可行。我不知道。量子计算,我想,应该以某种方式遵守这些定律,因为它们必须遵守热力学。所以这是一个可逆过程。但你想到的是什么?

Theoretically, this is massive, extremely parallel calculation. So that's exactly it—they are affected by turnover calculation. But honestly, I'm not an expert. I don't understand. If you can do reversible calculations, you don't need to spend energy on the process itself, because you don't delete any information. They're reversible. But I don't know if it works. I don't know. Quantum calculations, I suppose, should somehow obey these laws, because they must obey thermodynamics. So it's a reversible process. But what do you have in mind?

Host

与 Elad Gil 最难忘的对话之一是关于他实际上对量子计算非常怀疑。他说,即使它们今天就在这里,也没有应用。我就像,‘哦豁。’

One of the most memorable conversations with Elad Gil was about how he was actually very skeptical of quantum calculations. He said that even if they were here today, there are no applications. I'm like, 'Oho.'

Liam

我会说,这违反了安全性。就这样。我同意这一点。例如,人们说如果你今天有一台量子计算机,你可以更好地建模事物。我问他:‘好吧,假设如此。你会建模什么?’‘你都在建模同样的完美晶体。你不在乎 DFT 存在的问题,即使 DFT 在预测能力上是完美的。’所以是的,我真的这么认为——人们有点忽略了一些事情,因为量子计算很迷人。如果我们能在量子逻辑空间而不是经典逻辑中计算,这太令人兴奋了,以至于人们没有注意到,当它被创造出来时,这实际上会带来什么。也许这是件好事,因为也许当它被创造出来时,它会做出我们甚至不知道我们能想象的惊人事情。

This violates safety, I would say. That's all. I agree with that. For example, people say that if you had a quantum computer today, you could model things much better. And I asked him: 'Okay, let's assume this. What would you model?' 'You are all modeling the same perfect crystal. You don't care about problems that remain, which DFT has, even if DFT were perfect in its ability to prognosticate.' So yes, I really think so—that people are a little ignoring some things, because quantum calculations are fascinating. If we can calculate in quantum logical space instead of classical logic, this is so exciting that people don't pay attention to the fact that this will actually give when it will be created. Maybe that's a good thing, because perhaps when it will be created, it will do amazing things we don't even know we can imagine.

实验室自动化 Automation in Laboratories

Host

好的。我本来要转到自动化话题——你们在实验室里闭合了实验循环。你谈到了机器人操纵器。他们谈到了……但基本上,这就是你在这里的原因。我最喜欢你的引述是,你实验室里的每件设备都将拥有 140 的智商。这是什么意思?

Good. OK. I was going to go to automation topics—laboratories where you closed the cycle of their experiments. You talked about robotic manipulators. They talked about... But basically, that's all you're here for. My favorite quote of yours is that every piece of equipment in your laboratory will have an IQ of 140. What does this mean?

Liam

在初始阶段,当我们扩大流程规模时,我们明白我们有巨大的瓶颈,仅仅是由于与设备打交道。因此,对于其中一个物种,我们有设备技术员和科学家,他们正在寻找某种形态,并试图理解:‘好的,我们到底在制造什么?看看这是否符合我们的意图。’我们很快,在扩大实验室规模时,遇到了这些狭窄的地方,我们只是没有时间。所以我们开始使用软件方法在扫描电子显微镜(SEM)上收集数据。而由于这个非常简单的程序——选择视野、拍照、靠近、拍照——出现的数据并不是很有用。它们对于我们真正想要计算的东西来说太原始了。因此,在那一刻,直接在机器中实施 AI 系统以开始控制这些过程变得非常合适。现在它们拥有我们试图实现的目标的完整上下文。所以,实验的目标是什么?我们试图合成什么?其他实验证据是什么?它真的在机器内部查看并收集这些数据。这现在非常有价值,因为正如我们已经谈到的这些隐藏变量,不能考虑一切,但如果你能在实验期间进行更智能的数据收集,你为未来 AI 系统和预测提供的数据将变得更好。你将拥有更丰富的数据集。这就是我们所说的实验室中的一切必须具有令人难以置信的智能,以便我们使数据最大化有用的意思。所以这是这背后的一种动机和灵感。

In the initial stages, when we scaled processes, we understood that we had huge bottlenecks, simply due to working with equipment. Therefore, for one of the species we had equipment technicians and scientists, who were looking for a certain morphology and tried to understand: 'Okay, what exactly are we making? To see if this is consistent with our intentions.' And we very quickly, scaling the laboratory, encountered these narrow places where we just didn't have time. So we started using software approaches to data collection on scanning electron microscope (SEM). And the data that appeared thanks to this very simple program that chose the field of view, took a picture, approached, made the picture, were not very useful. They were too primitive for what we really wanted to count. Therefore, at that moment it became really appropriate to implement AI systems directly in machines to start controlling these processes. And now they have the full context of what we are trying to achieve. So, what was the goal of the experiment? What are we trying to synthesize? What were other experimental evidence? And it really looks inside machines and collects this data. And this is very valuable now, because as we already talked about these hidden variables that cannot take into account everything, but if you can do more intelligent data collection during the experiment, your data for future AI systems and forecasts will become much better. You will just have much richer data set. And that's what we have in mind when we say that everything in the laboratory has to be incredibly intellectual, so that we make data maximum useful. So this was a kind of motivation and inspiration for this.

AI在设备中的现状 Current State of AI in Devices

Host

关于在每个设备中实施智能,目前的状况如何?

What is the current state of cases regarding implementation of intelligence in every device?

Liam

嗯,我认为一个有趣的方面是:延迟是什么?没错。在管理这个工具时,云延迟,但也有设备上的计算。这是真的。还有与各种物理过程相关的时间线。如果 API 调用或类似的东西太慢,即推理 token 或工具调用的量不与过程本身的延迟相对应,那么就不可能实现。当然。没错。

Well, I think one interesting aspect is: what is the delay? That's right. In managing this tool, cloud delay, but there is also calculations on devices. This is true. There is also a timeline associated with various physical processes. And if the API call or something like that is just too slow, i.e., volume of reasoning tokens or tool calls does not correspond to the delay of the process itself, then it's impossible to realize. Of course. That's right.

Host

这几乎就像,你知道,无人驾驶汽车,对吧?因此,我,快速地说,我真的很惊讶听到这个,因为我想象的考虑延迟与许多这些事情执行数小时相比似乎很小,对吧?或者也许你的实验快得多,或者有高带宽或其他什么。我只是惊讶地听到这个。

It's almost like, you know, unmanned cars, right? Therefore, I, speaking quickly, I'm really surprised to hear this because the delay that I imagined for consideration seems small compared to what many of these things are performed for hours, right? Or maybe your experiments are much faster or have high bandwidth or something else. I'm just surprised to hear that.

Liam

最终,实验确实需要很多小时、数天,但也许有单独的阶段,延迟会很重要。所以,例如,你可能想要更精确的控制。另一件事——它有多贵,因为同样的原因,一些模型分析数据可能非常慢,也使它们非常昂贵。最后,人类的耐心。如果一个人需要分析特定的 XRD 图案,他们必须等待 2 小时才能得到好的结果,在我看来,这与 2 分钟得到结果完全不同。

Eventually, experiments really take up a lot of hours, days, but maybe separate stages where delay will matter. So, for example, you may want to have more accurate control. Another thing—how expensive is it, because it's the same reason why some of these models can be very slow to analyze data, also makes them very expensive. And, finally, human patience. If a person needs analysis of a specific XRD pattern and they have to wait 2 hours before getting a good result, this, in my opinion, is not at all the same as getting it in 2 minutes.

Host

哦,当然。是的。然后推理,你向这个模型提出不止一个挑战,对吧?是的?你可能执行许多许多 token 来达到这些模式。就像在循环中启动模拟。

Oh, of course. Yes. And then reasoning, you do more than one challenge to this model, right? Yes? You are performing potentially many, many tokens to reach these patterns. Like launching simulations in a cycle.

Liam

没错。深度研究在这里不仅仅是一个挑战。工具。这不是一个结论流。所以延迟可能会显著增长。

That's right. Deep research in this is not just one challenge. Tool. This is not one stream of conclusions. So the delay may grow significantly.

Host

我明白你的意思。等等,是什么?你的部分延迟是由于工具调用?我假设工具调用——这主要是 DFT 或计算方法,因此如此。所以,你的延迟中有多少是由于它们,而哪些实际上是推理,呃……

I understand that you mean. Wait, what is it? Part of your delays are due to tool calls? I assume that tool calls—this is mostly DFT or computational methods, and therefore so. So, what is your share of delays due to them, and which one is actually reasoning, uh...

Liam

这真的取决于过程。

It really depends on the process.

Host

好的,是的。还有一件事我在想,这可能是一个从微小到宏大的创造,我假设一般的诱惑或典型的发展是渐进的,当一切由人管理,然后你找到方法自动化并进一步发展。困扰我的是,有时人们就是这样开发东西的,但这是一个局部最小值,不是最优的。

Okay, yes. And one more thing I'm thinking about, this is probably a creation big from small, where I assume that general temptation or typical development is gradual, when everything is managed by people, and then you find ways to automate this and develop further. It bothers me that sometimes that's how people develop things, but this is a local minimum, not optimal.

Liam

没错。

That's right.

人形机器人与实验室自动化 Humanoid Robots and Automation in Labs

Host

什么时候才真正值得把人形机器人的工作直接放进去?

When is it actually worth it to take humanoid work and just put it there?

Liam

我们似乎坚持这样的观点:创造人形机器人实际上会拖慢我们实现某些目标的进度。

We seem to adhere to the opinion that creating humanoid robots would actually slow us down in achieving some of our goals.

Host

只是澄清一下。

Just clarifying.

Liam

再说,嗯,很多这些——只是你每天做的事。我们看到了,但我不知道真实情况。

Again, well, a lot of this — it's just what you do every day. We see this, but I don't know the real situation.

Host

许多人都有这样的机器人。

Many people have such robots.

Liam

是的。我认为其中一个过程是,将人与自动化连接起来,你可以很快地确定实验过程中的瓶颈,并开始逐一消除它们。因此,例如,如果有一个非常复杂的灵巧性任务,人类非常擅长,而自动化需要数月时间,也许不值得花大量时间在上面。也许更有前景的是那些非常常规、耗费科学家和技术人员大量时间、且容易自动化的事情。所以,我认为这有一定的实用性。但最终,我们从实验室想要什么?这是大量的数据、高质量的数据、各种数据,这些是我们的目标,而完全自主不是目标。我们使用自动化作为实现数据目的的手段。

Yes. I think that one of the processes is that, connecting people and automation, you can very quickly define narrow places in the experimental process and start eliminating them one by one. And therefore, for example, if there is something very complicated task on dexterity, in which people are wonderful, and the automation of which would take months, maybe it's not worth it to spend a lot of time on it. And perhaps more promising is something very routine that takes a huge amount of time for scientists and technicians, and what is easy to automate. So, I think this has a certain pragmatism. But in the end, what do we want from laboratories? This is a huge number of data, high quality data, various data, and these are our goals, and full autonomy is not the goal. We use automation as a means of achieving data purposes.

Host

是的。

Yes.

Liam

但我认为自动化的另一个方面是,工作可能在实验室中减少错误数量,拥有能够全面审查我们所有数据的人工智能系统非常有用,因为有时,如果数据中存在某种排列错误,它们可以检测到。例如,在我们某个阶段出现了一个循环错误,因为有人从机器上加载错了,模式有争议。所以 AI 分析了这些数据并说:“嗯,考虑到启动的内容,这不是预期的结果,”并且它动态地考虑了这个数据数组。然后它明白了:如果我们执行循环排列并将一切转回,数据就变得一致了;然后我们能够回到物理基础设施,并理解在下载过程中犯了一个错误。我相信数据质量管理是 AI 在物理世界中应用的基础,这正是我们实施的事情,以改进工作流程。但在这个时刻你明白:好的,这是自动化的另一个绝佳机会。我们如何使这个过程如此可靠和稳定,以至于类似的错误永远不会重复?是的,我认为如此。实验数据中背景噪声水平的降低非常重要。文献中的大多数实验数据噪声水平如此之高,以至于它们通常比理论方法(密度泛函理论 DFT)的准确性更差。因为,你知道,当不同的人在世界各地不同的实验室、不同的时间进行实验时,这确实在结果中增加了很大的变异性。所以我们希望,由于这些工作流程的标准化,最大的优势之一将是噪声降低。

But I think that one more aspect of automation is that work potentially will assume a smaller quantity of errors in laboratories, and it was very useful to have artificial intelligence systems with full review of all our data, because sometimes, if there is some permutation in the data, they can detect it. So, for example, on one of our stages a cyclical error arose, because one loaded from the machines wrong, and the patterns were controversial. So AI analyzed these data and said: 'Well, considering what was launched, it is not the expected result,' and it considers this data array in dynamics. And then it understood: if we perform a cyclic permutation and turn everything back, the data becomes agreed; and then we were able to return to the physical infrastructure and understand that during the download an error was made. And I believe that management of data quality is fundamental for application of AI in the physical world, and this is exactly the thing that we implement so that we improve the working process. But at this moment you understand: okay, this is another great opportunity for automation. How do we make this process so reliable and stable, so that similar mistakes are never repeated? Yes, I think so. A decrease in the level of background noise in experimental data is very important. The majority of experimental data in the literature have such a high noise level that they usually do worse than the accuracy of methods of theory, density functional theory (DFT). Because, you know, when different people spend experiments in various laboratories, in different parts of the world and at different times, this really adds up to very much variability in results. So we hope that thanks to standardization of these workflows, one of the largest advantages will be noise reduction.

模型栈与数据优势 Model Stack and Data Advantage

Host

你相当开放。你谈到了如何使用开源模型代码并重新训练它们用于类似任务。或者这样更好?我可以想象一种情况,你有几个模型执行一个任务,你为这个任务优化它们,然后有一个通用的前沿模型控制一切。这是一个好的心智模型吗?还有更多我不知道的阶段吗,我猜?

You are quite open. You talked about how you use open models source code and retrain them for similar tasks. Or is this better? I can imagine a situation where you have several models that perform one task, you optimize them for this task, and then there is some universal frontier model that controls everything. Is this a good mental model? Are there any more stages that I don't know about, I guess?

Liam

我们主要使用开源和闭源模型的组合。我们在建模、构建机器学习堆栈和模拟堆栈中从中获得了巨大的好处。我们不需要推动这个方向。但有许多领域延迟太高,成本太大。但在某些情况下,我们可以真正地,通过访问这些数据,超越这些系统所能达到的,即使是最高的努力。所以这不仅仅是在价值或效率上的帕累托改进,但在某些情况下,当你有访问别人没有的数据时,你可以超越这些限制。描述它的一种方式——从计算效率的角度。因此,通过访问这些数据,你可以比一些先进的模型在计算中更有效。它成为我们项目中真正决定性的因素。

We mainly use a combination of models with open and closed source code. We get a huge benefit from this in our modeling, building the machine learning stack and a simulation stack. We don't need to press this direction. But there are many areas where the delay is too high, and the cost too big. But in some cases we can really, having access to this data, go beyond what these systems are capable of even with the highest efforts to consideration. So it is not necessary only in value or efficiency by Pareto, but in some cases you can get out beyond these limits, when you have access to data which no one else has. One of the ways to describe it — from the point of view of efficiency calculations. Therefore, having access to these data, you can be a lot more effective in calculations compared to, you know, some advanced models. And it became a truly decisive factor in our program.

Host

是的,我正在考虑。数据质量或访问它们作为计算上的妥协。我认为存在某种汇率,美元换计算或美元换数据,在我看来,天平可能倾向于数据。我的意思是,当然,你会同意这一点。这尤其适用于当我们谈论实验数据的噪声水平时。我们需要高质量的数据。如果你为大量噪声花费巨大的计算资源,不会有什么好结果。对你来说更难,因为你积极尝试节俭。你也实际上试图获得零结果并从中学习,而且……

Yes, I am considering it. Data quality or access to them as a compromise on calculations. I think there is a certain exchange rate dollars for calculation or dollars for data, and it seems to me that the pendulum can lean towards data. I mean, of course, you will agree with this. This is especially true when we say, for example, about the level of noise in experimental data. We need high-quality data. If you spend a huge amount of computational resources for a lot of noise, nothing good from this will work. And it's harder for you because you are actively trying to be thrifty. You too are actually trying to get zero results and learn from this, and...

Liam

这是真的。

This is true.

Host

是的。

Yes.

零结果与负数据 Zero Results and Negative Data

Host

你提到过在其他几个平台上的零结果。对于材料来说,零结果看起来如何,你如何有效地使用它,特别是当我认为,你通常的成功信号相当罕见时?

You mentioned about zero results on several other platforms. How does a zero result look for materials and how do you effectively use it, especially when, as I think, your common signal of success is usually quite rare?

Liam

嗯。零结果可能意味着我们有意创建某种结构,但所有证据表明我们没有创建它。

Mhm. Well, zero result may mean that we had the intention to create a certain structure, but all the evidence indicates that we did not create it.

Host

你有没有曾经不打算创建结构?

Have you ever not planned to create a structure?

Liam

当然。是的。

Of course. Yes.

Host

好的。你花费阴性对照。

OK. You spend negative control.

Liam

当然。好的。当然。有杂质相,会破坏你的性质。

Certainly. OK. Of course. There are impurity phases, which will kill your properties.

Host

嗯,是的。

Well, yes.

Liam

好的。当然。这可能是有毒的,你知道。

OK. Of course. This could be toxic, you know.

Host

没错。100%。

Exactly. 100%.

Liam

好的。当然。嗯,在某些情况下,负面结果——这是当我们打算做一件事,但实际上能够识别出一个以前未知的新结构。所以从某种意义上说它是负面的。我们试图做一件事,但从数据中出现了其他东西。

OK. Of course. Well, and negative result in some cases — this is when we were going to do one thing, but in fact were able to identify a new structure that was previously not known. So in a sense it's negative. We sought to do one, but from the data something else emerged.

Host

是的。

Yes.

Liam

而且一般来说,当你训练机器学习算法时,特别是在基础层面,如果你使用分类算法。如果你没有负样本,如果一切都是正样本,你就无法很好地训练。这在材料科学中尤其是一个大问题,因为人们通常只发表他们成功合成的晶体,但几乎从不发表合成过程中失败的消息。有时他们可以。嗯,而且,我们从来不确定晶体是否通常未合成,对吧?这可能只是规模问题。

And also in general, when you teach a machine algorithm training, especially at the basic level, if you use a classification algorithm. If you do not have negative samples, you will not be fine to train, if everything is only positive. And this is especially a large problem in materials science, because people usually publish only those crystals that they managed to synthesize, but almost never publish messages about failures during synthesis. Sometimes they can. Um, and also, we never definitely know whether the crystal is in general unsynthesized, right? It could be simply a matter of scale.

Host

没错。是的,规模问题。例如,合成方法、技术的问题。是的。

That's right. Yes, the question of scale. For example, the question of method of synthesis, technology. Yes.

Liam

所以,嗯,我认为,当我们进行自己的实验并在我们尝试做的背景下得到负面结果时,这真的很有帮助。然后我们甚至可以教分类算法。我认为还有另一个重要时刻,当你有这样一系列负面结果,最终——由于过程的迭代——你取得了积极的结果。

So, um, I think, it really helps when we conduct our own experiments and we get negative results in the context of what we tried to do. Then we can even teach a classification algorithm. I think there is another important moment when you have such a series of negative results, and finally — thanks to iterations of the process — you achieve a positive result.

数据作为过程工程 Data as process engineering

Liam

这材料真的很有意思。我会把它称为数据类型的过程工程。因为迭代——你究竟如何得到正确的结果?在材料科学中,关于某个东西究竟是如何被创造出来的,存在太多模糊性。所以是的,即使材料是已知的,它的复现也可能是一项困难的任务。有些东西在学校教科书里,但另一些东西真的处在机会的边缘,所以我们积累这些经验。然后我们创建一个系统,它拥有若干负面结果,会告诉你如何真正达到那个正面的案例。

It's really interesting material. I would call it data-type process engineering. Because iteration—how do you actually come to the right result? In materials science there's so much ambiguity regarding how exactly something was created. So yes, even if the material is known, its reproduction may be a difficult task. There are some things in school textbooks, but other things really are on the verge of opportunities, and so we accumulate this experience. And then we create a system which, having a number of negative results, will tell you how to actually get to that positive case.

Host

回到结果分类器,我几乎会假设,如果你完全靠自己去做,你的大多数结果实际上会是负面的,而不是正面的。相当讽刺的是,因为文献只给你正面的结果。所以如果你从零开始,问题似乎恰恰相反:你有过剩的负面结果,而没有足够的正面结果。正如你所说,如果你为每个实验打标签,它们主要是负面的。但如果你为每个 campaign 这样做,并假设 campaign 在我们取得成功时结束,那么它可能会更平衡一些。

Going back to the result classifier, I would almost assume that if you do everything on your own, most of your results will actually be negative, not positive. Quite ironically, because literature gives you only positive results. So if you started it from scratch, it seems the problem is actually the opposite: you have a surplus of negative results and not enough positive. And as you said, if you made labels for each experiment, they would be mainly negative. But if you did this for each campaign, and assumed that campaigns end when we achieve success, then it could be a little more balanced.

Liam

我明白。所以你的推理链可以覆盖 campaign 的目标,而不是……好的。这是一个关键时刻。这样的数据几乎在其他任何地方都不存在,我们花了大量时间把完整的科学谱系过程植入模型。所以追踪所有这些数据,比如对话、直觉猜测、实验室里做了什么、在哪里进行计算、在哪里写代码——所有这一切的组合——这是极其宝贵的。我认为总体目标不是学习科学的最终结果,而是学习你自己的科学过程活动。这非常类似于你试图从递归自我改进(RSI)中实现的东西,即教模型训练最好的模型。但现在你是在教模型进行更好的实验。然而,区别在于它不会回到基础模型,所以这是……除非你开发的物理模型改进芯片,进而改进人工智能。所以这是……

I understand. So your chains of reasoning could cover campaign goals, not... OK. This is a key moment. Such data virtually nowhere else exists, and we spend a lot of time to lay the full scientific pedigree process into the model. So tracking all these data, such as conversations, intuitive guesses, what was done in laboratories, where calculations were held, where code was written—a combination of everything—this is extremely valuable. And I think the overall goal is not about learning on the final results of science, but learning on your own scientific process activities. This is very similar to what you would try to make from recursive self-improvement (RSI), where the model is taught to train the best models. But now you are teaching the model to conduct better experiments. However, the difference is that it is not coming back to the base model, so this is... unless physical models that you develop improve chips that then improve artificial intelligence. So this is...

Host

所以它更广泛,更……是的,RSI 更高。这是一个特殊的防火墙。但这是否意味着,比如说,在 Fable 6 或者 GPT 8 中,为这种自我反思设置的高阶痕迹逻辑可以让它无需任何逻辑就能做到,因为这是同一个思维过程,对吧?

So it's broader, more... Yes, the RSI is higher. This is a special firewall. But does this mean that it is possible that, let's say, in Fable 6 or maybe GPT 8, which is set up for such self-reflection, higher-order traces logic could let her do it without any logic, because it's the same thought process, right?

Liam

是的,我的意思是,在机器学习中,存在在不确定性条件下接受解决方案的情况。显然,在启动 AI 循环时会有噪声。但我们相信,当你真正与物理世界互动时,还有其他问题。此外,我认为还有另一个方面,即总权重的压缩——可能性仍然存在,非常有价值。如果在结论期间的考虑就足够了,所有领先的实验室都会停留在 GPT-4 的水平,并会说:‘好吧,从现在起我们只在推理期间改进。’因此,我们相信,由于物理科学和工程训练之间的差异,在我们自己的权重中进行压缩将导致不同类型系统和能力的出现。我们还相信,即使第七代模型会非常好,甚至比现在更好,它仍然必须花费实验来获得结果。原因是机器学习非常擅长处理它所训练的内容。而科学发现几乎按定义就是你没有训练过的东西。这就是为什么我们建造这些实验室,以便开放和封闭的模型都可以用它们来对宇宙进行实验,因为我们不相信不做测试就能做出伟大的发现。

Yes, I mean that in machine learning there is acceptance of solutions under conditions of uncertainty. Obviously, there is noise when launching AI cycles. But we believe that there are other problems when you actually interact with the physical world. In addition, I think there is another aspect, namely that compression of total weight—the odds are still there, very valuable. If considerations during the conclusion were enough, everyone leading laboratories would stop at the level of GPT-4 and would say: 'Okay, from now on we will improve only during inference.' Therefore, we believe that because of differences between physical sciences and engineering training, compression in our own weight will lead to the appearance of different types of systems and capabilities. We also believe that even if the model of the seventh generation will be very good, even better than now, it will still have to spend experiments for receiving results. And the reason is that machine learning very good copes with the fact that what it is trained on. And scientific discovery is almost by definition that which you weren't trained on. That's why we build these laboratories, so that both open and closed models could use them for experiments with the universe, because we do not believe that it is possible to do great discovery without testing.

Host

是的,没有人能在 AI 的帮助下从零发明室温超导体。

Yes, no one can invent a room-temperature superconductor from zero with the help of AI.

Liam

是的,是的。作为曾与小规模湿实验室密切合作的人,我当然比某些人更怀疑从零获得科学结果。但差不多就是这样。

Yes, yes. I have in mind, as a person who worked closely together with small wet laboratories, I am certainly more skeptical about receiving scientific results from scratch than some others. But that's about it.

Host

这一点,我想,有些人可能会问,所以。

That, I think, some can ask, so.

Liam

是的,我认为非常重要的是要把我们在数学和理论物理中看到的结果与物理世界区分开来。是的,这是两件不同的事情。这是两件非常不同的事情。我的意思是,即使在理论物理中,到目前为止这是理论计算机科学、编程和数学。也许理论物理是下一个。它会来的,是的。

Yes, I think that's very important to distinguish the results we see in mathematics and theoretical physics from the physical world. Yes, these are two different things. These are two very different things. I mean that even in theoretical physics, so far this was theoretical computer science, programming and mathematics. Perhaps theoretical physics next. It will come, yes.

Host

是的,是的。哦,是的。我能理解吗?你能达到的模型抽象的某个心智层次?所以,例如,关于自动化,我回到这个话题,我们谈到了 campaign,关于个别分析和进行测试,以及人们不擅长应对这个,因此我们必须把人从中移除。还有其他层次吗?例如,在 campaign 之上可以是物理理论,你正在检验的,而下面一个层次——是什么?

Yes, yes. Oh, yes. Can I get it? A certain mental level of model abstractions to which can you reach? So, for example, regarding automation, I'm going back to this topic, where we talked about campaign, about individual analyses and conducting tests, and about people being bad at coping with this, therefore we must remove people from this. Are there others? For example, at the level above the campaign can be a physical theory, which you are checking, and one level below—what?

Liam

你知道,我喜欢从这个角度思考,然后想‘好吧,现在这个 API 调用。也就是说,你再也不用碰这个了。而这个——是的,不,不,它仍然是 90% 的人类工作。’你知道,也许我们可以以某种方式绘制一张领域地图。有没有,或者还有其他层次吗?我可以给你命名一些层次。我觉得你提出了一个好问题。我没有非常系统的答案,但让我们谈谈一些层次。一个层次——这是原子结构。它是一个抽象,对吧?在我们所做的中,没有理想的原子结构,这只是一个近似。另一个是连续介质模型。所以它不再是原子结构,而是一个呈现材料的连续网格。当然。然后,在另一个测量中,有热力学。假设你永远做这个实验,最终状态会是什么?还有另一个抽象层次——动力学。认识到我们不会永远做这个实验,所以时间有价值。那么,反应会多快发生或不发生,即使它能量更低或不低。因此,热力学、动力学、原子、连续,还有什么?还有其他抽象层次吗?嗯,我的意思是,你关于可能存在共同的新理论突破,影响许多 campaign 的观点,有道理。也许。

You know, I like to think about it from this point of view, and then think that 'okay, now this API call. That is, you never again will have to touch this again. And this one—yes, no, no, it's still 90% human work.' You know, and maybe we could somehow draw a map of territories in such a way. Is there, or are there other levels? I can name you some levels. I feel that you put a good question. I have not very systematic answers, but let's talk about some levels. One level—this is atomic structure. It's an abstraction, right? There is no ideal atomic structure in what we do, this is only an approximation. The other is continuum model. So it's not anymore atomic structure, a continuous mesh that presents the material. Of course. And then, in another measurement, there is thermodynamics. Let's say you spend this experiment forever, what will be the final state? And there is another level of abstraction—kinetics. Recognizing that we don't do this experiment forever, so time has value. So, how soon will the reaction happen or not, even if it has a lower energy or not. Therefore, thermodynamics, kinetics, atomic, continuous, what else? Are there other levels of abstractions? Well, I mean that your opinion on that may exist common new theoretical breakthrough, which affects many campaigns, has sense. Perhaps.

Host

是的,作为投资者,我想说:这就是瓶颈,伙计们,当我们做到时,我们就会克服,我们就会得到其他一切。我不知道。

Yes, as an investor, I want to say: here it is, bottleneck, guys, and when we do it we will overcome, we will get everything else. I don't know.

自动表征与模拟 Automated Characterization and Simulation

Liam

从一开始,我们的假设是,因为大型狭窄环节之一具有自动化特征,所以混合粉末进行尝试并不那么困难。但如果你无法表征和分析这一点,然后明智地决定下一步该做什么,你真的不会从随机混合粉末中获得很大好处。因此,我们专注于自动化表征,通过模拟闭环,因为那正是你可以测试许多事情的方式,以采取平衡的观点决策并决定第二天做什么。

From the very beginning, our hypothesis was that because one of the large narrow places has automated characteristics, therefore it's not that difficult to mix powders for attempts. But if you can't characterize and analyze this, and then wisely decide which one should be the next step, you really won't get a big benefit from random mixing powders. Therefore, we focused on automated characteristics, closing the loop through simulations, because that's exactly how you can test many things, to take a balanced view decision and decide what to do next day.

循环中的人类参与 Human Involvement in the Cycle

Host

也许还有一个额外的问题。人们如何存在于这个循环中?也就是说,你在哪些时刻吸引一个人?我的意思是,我假设你可能有……是否有可能偶然加入任何阶段?

Maybe there is an additional question. How do people exist in this cycle? That is, at what moments are you attracting a person? I mean, I assume that you probably have... Is it possible by chance to join in at any stage?

Liam

不,当然不是。并增加价值。但你如何决定把时间和金钱花在什么上?看起来这可能是一起的。你有实验室里的人在物理上移动材料,然后你有科学家管理关于选择哪个活动的决策,你还有人们……我想我只是重复同样的事情。问题问你。

No, of course not. And add value. But how do you decide what to spend your money on time? It seems that this may be all together. You have people in the laboratory, which are physically moving materials, then you have scientists who manage decisions about which campaign to choose, and you still have people who... I guess I just repeat the same thing. Question for you.

Host

是的。我认为所有这些也在不断变化。

Yes. I think that all of this also constantly is changing.

Liam

是的。我的意思是,即使你采取领导活动,在初始阶段它完全由人管理。现在它越来越是人工智能控制和人类管理的结合。Darsh 说,一开始我们确定特征是大狭窄环节。在开始阶段,它是人驱动的,科学家非常超负荷,平衡显著转向人工智能一侧。这允许科学家,以前把所有时间花在这些澄清上,现在将你的工作提升到另一个层次。但我认为我们……它实际上是时间的函数。所有这些层面都在不断变化。

Yes. I mean that even if you take leadership campaigns, in the initial stages it was completely managed by people. Now it is increasingly more combination controlled by artificial intellect and managed by people. And Darsh said that at the very beginning we determined that the characteristic was large narrow place. At the beginning stages it was people-driven, and scientists were very overloaded, and balance significantly shifted to the side artificial intelligence. And this allowed scientists, who previously spent all my time on these clarification, now to elevate your work to another level. But I think we... it actually kind of function of time. All constantly changing at these levels.

不同领域的搜索方法 Search Methods in Different Domains

Host

从一般搜索的角度来看,有没有一些事情在物理世界中工作得很好,但在其他世界中不工作,或者反之?例如,进化搜索在大型语言模型中相当流行。我不认为它在这里工作得很好。所以,特别是在模拟中,人们使用进化搜索,你知道,足够长。嗯,够了。将很好地应对预测结构,例如搜索低能量结构。好的,所以它有效。然后,半导体行业的工程师-技术专家 DOE 参与,规划实验,他们经常使用一些零阶,或贝叶斯优化,或进化搜索,或者,我认为,他们使用强化学习(RL),但想法类似,是的。你需要有一些算法来找出。零阶优化它仍然像那样工作,对吧?因此,好的,我想回到我之前的问题,即——Scaling(规模扩张)。所以我想知道你的愿景是 Scaling(规模扩张)实验室?

From the point of view of general search, are there things that work well, say, in the physical world, but do not work in other worlds, or vice versa? For example, evolutionary search is quite popular in major language models. I don't think that it works well here. So, especially in simulations, people use evolutionary search, you know, enough is enough long. Hmm, that's enough. Will cope well with forecasting structures, for example, searching for structures with low energy. Okay, so it works. And then, the engineers-technologists in semiconductor industry DOE is involved, planning experiments, and they often used some zero order, or Bayesian optimization, or evolutionary search, or, I think, they use learning from reinforcement (RL), but the idea is similar, yes. You need to have some algorithm to find out. Optimization zero order it still works like that same, right? Therefore, okay, I want to return to the question I had earlier, namely—scaling. So I wonder what yours is. Vision scaling laboratories?

生物学与材料中的扩展策略 Scaling Strategies in Biology and Materials

Liam

有不同的方法,我的意思是,在生物学中你可以做很多酷技巧。你可以标记东西,使用 DNA 测序来接收各种有趣的数据。如果你能将你的复杂分析转移到测序,你可以将其扩展到数百万或数亿或类似的东西。你能使用类似的技巧吗?最终,每个分析都有自己的时间实现。我在博客上有一篇最喜欢的出版物,我想在发布说明中引用。但时间看起来像什么?这些实验室的执行?你对材料的 Scaling(规模扩张)从根本上说是线性的,在某种意义上,你只需要更多资源来做更多事情,你能以智能方式并行化过程,结合它们,这里有什么技巧吗?

There are different ways, I have meaning that in biology you can do a lot cool tricks. You can mark things, use DNA sequencing for receiving various interesting data. And if you can move your complex analysis for sequencing, you can scale its to millions or hundreds of millions or something like that. Can you use similar tricks? Ultimately, every analysis has its own time implementation. I have favorite publication on the blog I want quote in release notes. But what does time look like? Execution for these laboratories? Your scaling for materials fundamentally linear in the sense that what do you just need more resources to do more things, do you can you parallelize processes intelligent ways, combine them, is there are there any tricks here?

RL环境与数据扩展 Reinforcement Learning Environments and Data Expansion

Liam

嗯,我认为思考这个问题的一种方式——强化学习环境——它不是简单的启动实验并等待,直到智能体完全完成工作,进行计算并通过所有工具到最终结果。我们本质上,我们开始使用各种工具和所有这些不同的活动生成数据。而开始为机器学习扩展这一点的一种方式——这是创建基于强化的学习环境,再次基于当时实验或计算活动的数据。或者你也可以取工具的子集说:“好的,让我们创建一个基于这个单独工具的强化学习环境”。或者,也许它将是不同工具的组合,以达到某种奖励状态。所以,这就是方式,通过这种方式,这个有限的数据集如果以不同方式考虑,你可以为教学目的扩展。所以,这不适用于实际进行实验,但从机器视角教学,可以给你一些与你带来的生物学例子类似的类比。

Well, I think one of ways to think about this—the RL environment—it's not quite simple launch experiment and wait, until the agent is completely will do the job, making calculations and passing through all tools to final result. We, in essence, we start generating data using various tools and all these different campaigns. And one of ways to start expand this for machine learning—this creation learning environments with reinforcement on based, again, data about experimental or computational campaigns for that moment. Or you too you can take subsets tools and say: "Okay, we let's create an environment learning from reinforcement on based on this alone tool". Or, maybe it will be combination of different tools to to achieve a certain reward status. So, this is the way, by which this limited set data if to consider it in different, you can expand for purposes teaching. So, this does not apply actual carrying out experiment, but with machine view teaching, can give you some analogues to biological examples that you brought.

组合喷涂与超导测量 Combinatorial Spraying and Superconductivity Measurement

Liam

所以,你知道,人们尝试过的方法之一——这是组合喷涂。我不知道你是否听说过这个?也就是说,你取溅射靶材并创建梯度组成。所以,当你查看最终产品时,通过来自每个前驱体的梯度,在不同地方形成不同的晶体。所以,一次你可以尝试数百或数千种晶体。这是一个例子,本质上,与你的类似,不是吗?

So, you know, one of the methods that people tried—this is combinatorial spraying. I don't know whether you have heard of this? That is, you take sputtering targets and create a gradient composition. So, when you look at the end product, through gradient from each predecessor different are formed crystals in different places. So, at one time you can try hundreds or thousands of crystals. It an example that, in essence, similar to yours, isn't it?

Host

是吗?是的。而且,据我所知,众所周知,到目前为止这工作得不太好,因为结果证明存在扩散,所以一切都混合了。

Yes? Yes. And, as far as I know it is known that this is so far doesn't work very well, because it turns out there is diffusion, so everything mixed.

Liam

然而,我们想到但尚未实施的一种方法,许多人说——这是测量超导性,这是一个瓶颈,因为与 X 射线衍射(XRD)不同,不存在高通量的超导性测量方法,一次测量可能占用大约数小时。因此,人们正在考虑取不同的候选者,将它们混合成一个样品并进行测试。如果它显示出超导性,你就知道其中一个成分起作用了,然后你可以像他们做 COVID 测试之前那样行动,所以这种想法存在。有些有效,好的,有些无效。

Yet one method that we thought, but not yet implemented, and about which many says,—this measurement superconductivity, which is a bottleneck, because, unlike from the X-ray diffraction (XRD), not exists highly productive measurement methods superconductivity, and one measurement can to occupy about hours. Therefore, people were thinking about to take different candidates, mix them into one sample and protest. And if he shows superconductivity, you you know that one of components worked, and then you can act like this before they do COVID tests, so ideas this kind exist. Some of them work Okay, some don't.

新资金与实验室扩展 Scaling with New Funding and Laboratory Expansion

Host

是的。所以,我认为你将宣布关于大量参与资金。你计划如何扩展?

Yes. So, I think you are going to announce about great involvement funds. How are you planning to scale?

Liam

我们有一段时间一直在设计我们的实验室,我认为资源将允许我们显著扩展它们,并增加算力,既从 AI 方面,也从一般计算方面。但我认为极其重要的是我们还能建造什么新类型的实验室。我认为我们已经在门洛帕克这里的第一个实验室中检查了这个循环,我们将继续扩展到新类型的实验室,重复同样的过程。是的,我们试图建造的那种类型的实验室,我认为,以前没有以这种规模或这种方法创建过。我们在自己的建设中学习了许多原因,这些教训帮助我们准备后续版本。

We have been for some time we are engaged in design our laboratories, and I think the resources will allow us to scale significantly them, as well as increase computational power both from the side AI, and from the side general calculations. But I think that extremely important also what new types of laboratories we can build. I think we already checked this cycle in our first laboratories here in Menlo Park, and we will continue expand to new types laboratories, repeating the same process. Yes, laboratories of the type that we are trying to build, I think, not before were created on this scale or with with this approach. We many reasons learned during own construction, and these lessons help us prepare the following versions.

扩展工具与硬件 Scaling Tools and Hardware

Liam

然后,就像你说的,我们可以扩大规模、雄心和工具的质量。有些工具非常昂贵。因为我们开始更好地理解什么能带来最大回报,我们就可以在它们上投入更多。这使得硬件工程成为流程的核心部分。例如,在初始阶段,只是为了速度,我们购买了现成的设备。而随着我们扩大实验室规模,你就得自己创造。我们被迫自己创造。所以,我们明白:好吧,这个设备相当快。这些组件确实相当快,但在这些规模下它们成了“瓶颈”。或者它们可能是与质量相关的数据,例如,这个机器人机械臂的静止位置在已经混合了物质的板子上方。所以有污染的风险。因此,在我们设计设备时,我们这样重新设计,使静止位置远离这些东西。这些是细微的细节,允许我们降低实验活动中的噪声水平。所以,我们获取这些知识,然后扩大规模。

And then we can, just like you said, increase scale, ambition and quality of tools. Some tools are very expensive. Because we start to better understand what brings us the greatest return, we can invest more in them. And this makes hardware engineering a central part of the process. For example, in the initial stages, just for speed, we bought ready-made appliances. And since we are expanding laboratory scale, you create your own. We are forced to create your own. So, we understand: okay, this device is pretty fast. These components are really quite fast, but at these scales they become a "narrow" place. Or they can be things that are quality-related data, for example, the rest position of this robot-manipulator is above a plate with already mixed substances. So there is a risk of pollution. Therefore, in our design of equipment we rework it like this, so that the resting position is away from these things. These are the thin details that allow us to reduce noise levels in our experimental campaigns. So, we take this knowledge and we scale them.

公司结构与招聘 Company Structure and Hiring

Liam

你知道,你根据任务来构建公司结构。你有很多命令。我想当我们查看招聘页面时,我们感到惊讶,并想:我们不完全理解 Periodic 在做什么。也许值得在职位页面上添加组织原则。所以我们在寻找 AI、基础设施、计算、实验角色、硬件供应和杂货方向的专家。我们想要实现合成超级想法,对吧?我们觉得这需要化学、物理、建模、理论以及处理薄膜和粉末的相关经验。工程设备——看起来是非常不同的角色,但实际上它们的目标是整体的:创建一个超级穹顶。这也适用于 LLM 研究人员、基础设施,甚至杂货工程师;因为后者保证研究成果成为产品,可以被实验室的实验者使用。

You know, you build company structure according to their tasks. You have a lot of commands. I think we were surprised when we looked at the job page, and thought: we are not completely understand what is engaged in Periodic. Maybe it's worth adding principles of organization to deal with vacancies page. So we are looking for specialists in the fields of AI, infrastructure, calculations, experimental roles, hardware provision and grocery directions. We want to achieve synthetic super idea, right? And we feel that for this is needed relevant experience in chemistry, physics, modeling, theory, and also in working with thin films and powders. Engineering equipment - it would seem very different roles, but actually they are holistic in their purpose: creating a superdome. This also applies to LLM researchers, and infrastructure, and even groceries engineers; because the latter guarantee that the results of research are becoming a product that can be used by experimenters in laboratories.

多学科方法 Multidisciplinary Approach

Liam

是的,要闭合与物理世界互动的循环,你需要对这个世界上有影响力,建立实验室,进行有效的活动,创建 AI 系统并有效使用计算工具。这是复杂性,但同时也是 Periodic 的机会。这样一群人以前从未聚集在一起。这不可避免地是一个多学科问题。

Yes, to close the cycle of interaction with the physical world, you need to have influence into this world, to build laboratories, to conduct effective campaigns, create AI systems and effectively use computational tools. It is complexity, but at the same time an opportunity for Periodic. Such a group of people never before were not collected together. It inevitably multidisciplinary problem.

动手工作原则 Hands-On Work Principle

Liam

是的,对我们来说还有一个重要的指导原则:我们希望非常熟练和有经验的人直接亲自动手工作。你知道,现代生活导致了一个事实,特别是在学术环境中,当某人在研究方面非常熟练时,我们给他加上了太多关于写资助申请和教学的责任,以至于在这个行业中没有剩下时间进行实际研究。但如果你看看贝尔实验室、IBM,那些取得重大进展的机构,那么这些都是非常有经验的人,他们亲自工作。例如,巴丁每天都在实验室,尽管他是理论家。亚历克斯·穆勒在实验室进行实验,即使他是领导者。所以我们在这里尝试做同样的事情。我们有员工是不同行业的世界领袖,但他们执行实际工作。

Yes, and one more important guiding principle for us: we want very skilled and experienced people to do the work directly with their own hands. You know, modern life has led to the fact that especially in academic environment when someone is really skilled in research, we put on him so many responsibilities regarding writing grants and teaching, that there is none left time for practical research in this industry. But if you look at Bell Labs, IBM, institutions that made significant progress, then these were very experienced people, who worked personally. For example, Bardeen was in laboratories every day, although he is a theorist. Alex Muller conducted experiments in laboratories, even being its leader. So we try to do the same here. We have employees who are world leaders in different industries, but they perform practical work.

关键团队成员 Key Team Members

Host

你想让我数数他们,你知道,炫耀一下吗?

Do you want me to just count them, you know, show off?

Liam

哦,我很乐意。是的。所以,你知道,当我写论文时,我最喜欢的计算材料科学专家,大约和我同龄,是 Murat Aykol。他从我们的顾问之一 Chris Wolverton 那里获得了博士学位。他可以说是我们的计算材料科学专家。他非常了不起。他有时,你知道,真的是很好的实验家,尽管他从事建模。他非常了解系统。我们的实验室负责人——Joe Chekaluski,麻省理工学院教授,他休假来全职与我们合作。Joe 是一位了不起的物理学家,非常了解超导性,但他也在日本当过教授,所以他很好地掌握了合成和化学,尤其是物理学方面。我们有 Daniel Chica,他曾在 Mercury 合成小组——也许是世界上最好的固体化学小组——读研究生,Daniel 是其中的合成专家之一。所以他就像个魔法师。我的意思是,有很多这样的例子。

Oh, I would be happy to. Yes. So, you know, when I was writing my dissertation, my favorite specialist in computational materials science, who is approximately my age, was Murat Aykol. He defended his dissertation from one of our advisors, Chris Wolverton. And he, we can say, our expert in computational materials science. He is incredible. He sometimes, you know, really good experimentalist, although he is engaged in modeling. He understands systems very well. Our leader of laboratories—Joe Chekaluski, Professor at MIT, who took vacation to work with us full-time. Joe is an incredible physicist who understands superconductivity very well, but he was also a professor in Japan, so he mastered synthesis and chemistry well, especially for physics. We have Daniel Chica, who was a postgraduate student in the group Mercury synthesis—perhaps the best chemistry group in solid body in the world—and Daniel was one of its experts on synthesis. So he is like a wizard with his hands. I mean, there are so many like this examples.

更多团队成员 More Team Members

Liam

不,我知道。我想到的是,例如,Dima Babakhin,他领导我们的许多 AI 倡议和主要语言模型。他是神经网络中注意力机制的发明者。呃。哦,Babakhina。注意 Babakhina。是的,没错。是的。他非常参与实际工作,深入研究追踪、训练数据分布,以及如何将其与实验室联系起来。化学、对话——我的意思是,他深入钻研,是的。他非常投入。深入。他在蒙特利尔有实验室,甚至有固体合成书籍。Ray Nakano,他曾在 OpenAI 担任相机工作的技术经理。他在实验室工作。因此,当我们参观实验室时,我们起初想:“哦,这位新科学家,这位新技术专家看起来像 Ray。”但不,那确实是 Ray,他实际上和科学家一起在实验室里,自动化机器。这就像同样的 DNA 和我们在这方面取得成功所需的人才类型。

No, I know. I have in mind, for example, Dima Babakhin, he leads many of our AI initiatives and major language models. He was the inventor of the attention mechanism in neural networks. Ahem. Oh, Babakhina. WARNING Babakhina. Yes, that's right. Yes. He is very involved in practical work, delving into tracing, distribution data for training, and how to relate this to laboratory. Chemistry, conversations—I mean, he dives deep, yes. He is very immersed. Deep. He has laboratories in Montreal even has synthesis books solids. Ray Nakano, he was technical manager of camera work in OpenAI. He works in laboratories. Therefore, when we spent excursions in laboratories, we at first thought: "Oh, this new scientist, this new technical truly a specialist looks like Ray." But no, it was literally Ray, who was actually in laboratories together with scientists, automating machines. And it's like the same DNA and the type of people we needed to achieve success in this matter.

部署工程师角色 Deployment Engineers Role

Host

所以,我们还看到你们有部署工程师的职位。我们正要推出一个关于部署工程师工作的播客,因为有很多工程师可以承担这个角色。他们甚至可能不知道在 PeriodX 对他们来说有这样的空缺。这个角色是什么,他们如何与客户合作?

So, we also saw that you have positions of engineers from deployment. We are just going to launch a podcast about work engineers from deployment, because there are so many engineers who could would take on this role. They may even don't know what's in PeriodX for them there is such vacancy. What is this role and how they working with customers?

Liam

我们为自己创造了我们自己的产品、自己的工具、自己的 AI 和计算能力。所以我们是我们自己的第一个客户。现在我们把这些相同的工具用于不同的行业,现在我们专注于半导体行业。然而,要做到这一点,因为这是非常封闭的公司。我们必须能够在最受保护的环境之一中工作。因此,我们的部署工程师和研究人员将直接在对象上工作,整合我们的 AI 系统或计算系统,帮助我们的合作伙伴更快地实现目标。因此,本质上,我们在自己的研究和材料工程中创造的工具和专有技术,我们现在提供给行业。我们相信这是帮助合作伙伴实现最终结果和目标的最有效方式,而不仅仅是把技术传递给他们说:“自己想办法吧。”此外,这些部署工程师将在本地得出结论,并且能够根据数据训练模型。再次,当我们采用我们的系统,它理解这些领域,我们展开它,我们可以用数据专家客户或合作伙伴来做,这样他们就可以用自己的智慧和理解这些的系统,这比仅仅参考 API 或未训练的模型要有效得多。

We created our own products, own tools, own AI and computing power for ourselves. So we were the first client for ourselves. Now we take these same tools for different industries, and now we are focused on semiconductor industry. However, to do this, because it is very closed companies. We must be able to work in one of the most protected environments. Therefore, our engineers from deployment and researchers will work directly on objects, integrating our AI systems or computational systems to help our partners faster achieve goals. Therefore, essentially, tools and know-how that we created during own research materials and engineering, we are now offer industry. And we believe that this most effective a way to help partners to achieve final result and goals, not just to pass it on to them technology and say: "figure it out for yourself." In addition, these engineers from deployments will be to draw conclusions locally, as well as will be able to teach models on data. Again, when we take our system, who understands these areas, and we unfold it, we can do it data expert customers or partners so that they could own with one's own intellect and systems that understand this very much more effective than simply referring to API or untrained models.

AI for Science新角色 New Roles in AI for Science

Host

所以这很重要——相比标准岗位,从部署转过来的工程师角色被扩展了,因为它需要机器学习的深厚专业知识。你必须在数据学习、基础设施,以及物理学的视角上都极其精准。我们与高科技行业合作。他们必须理解化学、材料科学以及某些类型器件等多个领域。所以这些就是我们正在创造的角色。这部分也意味着愿意在台湾的某个设施长期投入。

So this is significant — an expanded role for engineers from deployment compared to standard, because it requires deep expertise in machine learning. You need to be incredibly accurate with a view of learning on data, infrastructure, and also from the point of view of physics. We work with high-tech industries. They must understand various areas such as chemistry, materials science, and some types of devices. So these are the roles that we are creating now. This is partly a willingness to spend for a long time at a facility in Taiwan.

Host

好吧,我不了解这类客户。我很好奇你们的是什么样的。变现——你们现在主要的收入方式是什么?未来可能会变。但我想说的是,在我看来,这是一种人们对于研究型 AI 实验室并没有完全理解的模式。比如,如果你把自己比作贝尔实验室,这正是人们谈论的,我想,对吧?比如,什么——我觉得一个很好的类比是软件开发。

Okay, well, I don't know this type of customer. I wonder what yours is. Monetization — what is the main way in which you are earning now? But in the future it may change. But I guess I mean that this is a model that, in my opinion, people don't completely understand regarding research AI labs. For example, if you compare yourself to Bell Labs, this is what people are talking about, I think, don't you? For example, what is — I think a very good analogue there would be software development.

从Copilot到自主研究 From Copilots to Autonomous Research

Liam

是的。所以在软件开发初期,我们有安全——你知道,GitHub Copilot,然后是早期版本的 ChatGPT。人们把这些东西当作副驾驶来用,帮助自己找到解决方案。而因为自动化改进了,现在我们有了像 Codex 这样的东西,我们很少有工程师还像以前那样写代码。我们相信类似的情况也会发生在物理学上,你可以有系统来加速研究人员、科学家——材料科学家、材料工程师、工艺工程师——的工作。但随着它在自主性、智能和机会上的增长,你可以开始评估结果。所以,帮助——我想达到这个目标状态。我认为这对我们来说是一个非常有趣的领域。

Yes. So, at the beginning of the development of software we had security — you know, GitHub Copilot, then early versions of ChatGPT. People used these things like co-pilots, to help them find a solution. And because automation improved, now we have things like Codex, and very few of our engineers write code just like it was before. And we believe that a similar situation can also happen with physics, where you can have systems for speeding up the work of researchers, scientists — materials scientists, engineers from materials, process engineers. But as it grows in autonomy, intelligence, and opportunities, you can start to evaluate results. So, help — I want to achieve this target state. And I think that's a very interesting area for us.

Host

是的。我的意思是,我们觉得我们对固态物理和科学最好的贡献,可能就在于让这些工具和主题变得有利可图。就像 ChatGPT 让计算机科学和 LLM 专业在大学里前后变得热门得多。我们想展示,物理学、固态和材料科学领域的研究可以产生巨大的商业影响。然后它会吸引更多关注,年轻人会想学物理,那将是一个梦想。而且,你知道,贝尔实验室确实产生了巨大的商业影响,对吧?所以,他们没能把自己一些不可思议的成就商业化。他们当然做过一些根本无法商业化的研究,比如——宇宙背景。但他们从真空管中获得了大量收益,真空管连接了东海岸到西海岸的电话线。

Yes. I mean that we feel that our best contribution to solid state physics and science could consist in the fact that to do these tools and themes profitable. Just like how ChatGPT made CS and LLM majors a lot more popular in colleges before and after. We would like to show that research in the field of physics, solid state and materials science can have a big commercial influence. And then it will attract more attention, young people will want to study physics, which would be a dream. And, you know, Bell Labs really made a huge commercial influence, or wrong? So, they were not able to commercialize some of their own incredible achievements. They, of course, conducted some research, which was simply impossible to commercialize, as — here is the cosmic background. But they got a lot of benefits, for example, from a vacuum tube, which connected the eastern coast from western telephone numbers lines.

技术与资本交织 Technology and Capital Intertwined

Host

是的,我觉得思考技术和资本如何不可思议地交织在一起也非常有趣。因此,如果你看聊天机器人多年来的进展,想象你在数,好吧,聊天机器人的进展,比如说,从 2010 年到 2015 年。我们可能很难说它们在那五年里改进了多少。而如果你看 2021 年到 2026 年这段时期,那就是天壤之别。发生的是,ChatGPT 和其他类似的系统能够实现合规、产品市场,这完全改变了资本的格局。它改变了招聘、算力、数据等方面的情况。良性循环。

Yes, and I think that's also very interesting to think about — what technology and capital is incredibly intertwined. Therefore, if you look at what the progress was in chatbots during many years, imagine that you're counting, okay, what is the progress of chatbots, say, since 2010 until 2015. We would probably hesitate to say how much they improved over that five-year period. And if you look at the period from 2021 to 2026, so this is, you know, the sky and Earth. What happened is that ChatGPT and others like it — systems were able to achieve compliance, product market, and this completely changed the landscape of capital. It changes the situation with hiring, computational capacities, data, etc. Virtuous circle.

Liam

没错。所以成功会催生成功。

That's right. And therefore success generates success.

Host

正是如此。所以,技术与资本和资源紧密相关,我们想在物理世界实现同样的事情。

Exactly. So, technology is incredibly tightly related to capital and resources, and we want to achieve the same in the physical world.

开源与资金 Open Source and Funding

Host

是的。我观察到的一个趋势是,从科学资助——政府拨款等——转向风险投资和私人资本。你们是否计划开源你们自己的任何成果?比如,发布数据集或实际模型?或者也许在未来——即使这些模型不是最先进的,它们仍然可以是有用的发布。

Yes. One of the trends that I am observing is the transition from science funding — government grants, etc. — to venture and private capital. Are you planning to open the weekend code of any of your own achievements? For example, to release datasets or the actual model? Or maybe it's in the future — even if these models are not state-of-the-art, they can still be useful releases.

Liam

当然。也就是说,我们有几个方向。我们对开源软件做出贡献,比如 pymatgen、代码库 Materials Project、Custodian、它们的 DFT 运行器,而 Torch Sim 是由我们的一位研究员 Abhijit Gangan 创建的,至今仍由他支持。Jackson MD,Abhijit 也在支持。所以,我们对开源做了很多贡献,包括你提到的 Megatron。是的,是的。我们是开源软件的非常积极的贡献者。我们也有学术项目资助,通过它我们为大学团队提供资金,在我们看来,这些团队真的在合成超级穹顶的方向上取得进展。这也非常值得关注。看起来这项资助的第一个成果文章很快就要出来了,所以看到更多这样的出版物会很有趣。

Certainly. That is, there are several of our directions. We do contribution to the open software provision, such as pymatgen, code base Materials Project, Custodian, their DFT runner, and Torch Sim was created by one of our researchers, Abhijit Gangan, which is still his supports. Jackson MD, Abhijit supports it. So, we do a lot of contributions to open source, including as you said, Megatron. Yes, yes. We are very active contributors to open source software. We also have academic program grants, within which we provide financing to university groups that, in our opinion, really are advancing in the direction of synthetic superdome. And this is also very nice to watch. It seems the first article by the results of this funding has to come out soon, so it will be interesting to see more like this publications.

超导之路 The Path to Superconductivity

Host

好。让我们想象你成功了——实现了 100% 的实验室自动化,它做所有事情并正确表征每个方面。你拥有一切必要的——能够执行任何任务的智能体。我还是不太明白。你如何达到超导是很清楚的。有许多复杂问题需要解决。哪条路?在这之前,什么时候,之后?本质上,对于大多数强关联或高温系统,没有理论或模型?

Good. Let's imagine that you are successful — implemented 100% laboratory automation, which does everything and right characterizes each aspect. You have everything necessary — agents able to perform any tasks. I still don't quite get it. It's clear how you will reach superconductivity. There are many complex problems that need to be solved. Which way? Before this, when, after? In essence, there is no theory or models for most highly correlated or high-temperature systems?

Liam

是的,我们认为不难想象能够提供超导性的化学空间。实际上难的是合成。这就是为什么我们如此强调合成中的超级智能。也许,一个历史例子是 Alex Muller,当时他认为过渡金属氧化物可能在通往超导的道路上表现良好。事实上,他尝试研究镍酸盐。镍、氧和一些阳离子——没有成功,然后他尝试铜酸盐,就是铜、氧和一些阳离子,这次成功了,之后他获得了诺贝尔奖。但后来我们了解到,镍酸盐也是一个好主意。我的意思是,镍酸盐和铜酸盐在元素周期表中并排。

Yes, our opinion is that it's not hard to imagine chemical spaces that could provide superconductivity. It's actually hard to synthesize. That's why we do this great emphasis on superintelligence in synthesis. Perhaps, one historical example is Alex Muller, when he thought that oxides of transition metals can be good on the way to superconductivity. In fact, he tried to work with nickelates. Nickel, oxygen and some cations — it did not work, and then he tried cuprates, simply copper, oxygen and some cations, and this it worked, after which he got the Nobel Prize. But later we learned that there were also nickelates — a good idea. I have in mind that nickelates and cuprates stand side by side in the periodic table.

Host

是的。是的。Haidong Wang,我们在斯坦福的一位教授,一个不可思议的人,意识到他可以以薄膜形式创建镍酸盐并展示超导性。因此,如果我们拥有真正强大的合成超级智能实验室,并且我们能够扩展这些东西,我认为我们不会用完想法,或者 LLM 不会终结关于哪些方向值得尝试的想法。比如 3D 过渡金属、氧——这里有很多相关的想法。许多人尝试研究钴酸盐,即用钴代替铜或镍。所以我认为就是这样。极其迷人。也许,值得学习另一个历史例子——一个发现二硼化镁的日本团队。如你所知,这是常压下最高温度的常规超导体。他们发现它,只是尝试了一堆材料。看起来他们测试了 30,000 种不同的物质。

Yes. Yes. Haidong Wang, one of our professors at Stanford, an incredible person, realized that he could create nickelates in the form of thin films and demonstrate superconductivity. Therefore, if we have laboratories that will be really powerful superintelligence for synthesis and we will be able to scale these things, I think we don't have run out of ideas, or in LLM will not end ideas about what directions worth trying. Like 3D transition metals, oxygen — there's plenty of it here, related ideas. Many people tried to work with cobaltates, i.e. instead of copper or nickel to cobalt. So I think that's it. Extremely fascinating. Perhaps, it's worth learning another historical example — a Japanese group that discovered magnesium diboride. As you know, this is a normal superconductor with the highest temperature at atmospheric pressure. And they opened it, just trying it out a bunch of materials. It seems they tested 30,000 various substances.

自动化与材料发现 Automation and Materials Discovery

Host

其中 30 种表现出有趣超导性,MgB2 就是其中之一。但二硼化镁作为前驱体在架子上放了几十年。BCS 理论早在 1957 年就提出了。因此,人们发现这些材料的方式并不是理论的直接结果。并不是因为他们早先无法合成它们。只是因为他们尝试了很多东西。所以想象一下,如果我们有你描述的那种完美自动化,听起来像做梦一样——我们会在一个月内尝试 Taketo 团队整个职业生涯研究的 30,000 种选项。我们就会大大扩展成功的机会。

30 of them demonstrated interesting superconductivity, and MgB2 was one of them. But magnesium diboride lay on the shelves for decades like a precursor. And the BCS theory was invented back in 1957. Therefore, the way in which people found these materials was not a direct result of theories. It was not because they couldn't synthesize them earlier. It was just because they tried a lot of everything. So imagine, if we had that perfect automation that you describe, it sounds like a dream—and we would try those 30,000 options that Taketo's group investigated throughout their careers, for one month. We would just significantly expand opportunities for success.

Liam

不,但是,是的,你所做的一切非常鼓舞人心。祝贺你所有的成功。对我来说,你似乎创造了这样一个现代的研究乐园。我想象招聘你一定极其容易。所以我有点嫉妒,但你付出了很多才实现这一点,因此……是的。

No, but yes, what you've done is very inspiring. Congratulations on all your successes. To me, it seems that you create such a modern playground for research. I imagine hiring you must be extremely easy. So I'm a little jealous, but you have worked many to achieve this, therefore... Yes.

Host

是的。嗯,非常感谢。谢谢。和你们两位聊天很愉快。

Yes. Well, thank you very much. Thank you. It was nice to chat with both of you.

Liam

是的,非常有趣。是的。

Yes, very interesting. Yes.

互动版:逐字朗读 + 针对本期提问 →