From ChatGPT to Atoms: Liam Fedus on AI for the Physical World
打开互动全文版(中英对照 + 朗读 + 问答)→ChatGPT 联合创始人、前 OpenAI 副总裁 Liam Fedus 分享他从物理学到 AI 的历程,以及他正在构建的面向原子的 AI 基础实验室。
Liam Fedus, co-creator of ChatGPT and former VP at OpenAI, discusses his journey from physics to AI and his new venture building an AI foundation lab for atoms.
今天我们与 Liam Fedus 对话。Liam 是 ChatGPT 的联合创始人之一,我想现在几乎每个人都在用。他曾任 OpenAI 的后训练副总裁,之前还在 Google Brain 工作,参与了许多早期 AI 创新。Liam 将向我们介绍他的公司 Periodic Labs,这家公司专注于为原子世界构建 AI 基础实验室。换句话说,就是如何利用 AI 影响物理世界、材料科学、化学等领域。非常激动人心的话题,很高兴今天能与他交流。Liam,非常感谢你今天来到 Priors。
The Dana Prize we're talking with Liam Fedus. Liam is one of the co-creators of ChatGPT, which I think almost everybody uses at this point. He was the VP of post-training at OpenAI and before that was at Google Brain, where he worked on a variety of really early AI innovations. Liam will be telling us a bit about Periodic Labs, his company, which is focused on building an AI foundation lab for atoms. In other words, how do we impact the physical world, material sciences, chemistry, etc. using AI. Very exciting topic and excited to be talking with him today. Liam, thank you so much for joining us today on the Priors.
是的,非常感谢你的邀请。很高兴见到你。
Yeah, thank you so much for having me. It's great to see you.
嗯。那么,也许我们可以这样开始——我认为你在替代模型类型方面做着非常有趣的事情,特别是针对材料科学和物理世界。实际上,你正在构建的是一个面向原子的 AI 基础实验室,我觉得这很迷人。
Yeah. So, maybe what we can do I I think you're doing incredibly interesting things in terms of alternative types of models, specifically for material sciences, for the physical world. Effectively, what you're building is an AI foundation lab for atoms, which I think is fascinating.
没错。
That's right.
但也许我们可以先多聊聊你的背景。我知道你曾是 OpenAI 的副总裁,参与过最早万亿参数模型的工作等等。能跟我们多讲讲是什么让你走到今天这一步的吗?
But maybe we can start with just a little bit more of your background. You know, I think you were VP at OpenAI, you worked on one of the first trillion parameter models ever, etc. Could you tell us a little bit more about just like what got you here and
嗯,再早之前,我本科是物理专业。花了一些时间做暗物质研究。我们有一个装置,对暗物质的方向敏感。所以非常有趣。
Yeah, so even further back I was a physics major in undergrad. Spent some time doing dark matter research. We had a apparatus that was directionally sensitive to dark matter's direction. So, it was very interesting.
为什么——我很想回到这个话题——为什么现在 AI 领域有这么多物理学家?你看 Dario Amodei 领导 Anthropic,Adam Brown 在 Google,还有很多人,他们都有物理背景。
Why why are there so many I'd love to come back to this but why are there so many physicists in AI right now? So, you look at Dario Amodei who runs Anthropic. You look at Adam Brown at Google, you look at a variety of people and they all kind of have these physics backgrounds.
是的,我以前的经理 Josha 在 Anthropic 也是物理背景。
Yeah, my old manager Josha also physics at Anthropic.
嗯,你觉得这是为什么?
Yeah, why why do you think that is?
我认为这是一种很好的思考世界的方式。非常讲原则,非常严谨的科学家,非常仔细。而且,我不知道,我觉得这只是一个非常不可思议的领域。你在计算机科学和 AI 中有很高的杠杆效应。所以,我认为很多物理学家看到了这一点。特别是在高能物理领域,希格斯玻色子的发现之后,我想很多高能物理学家都在寻找下一步。最终,它受限于新的设备,你知道,去推动下一个能量前沿。我认为很多物理学家在审视自己的技能组合,看到其他领域的进展,然后说:‘嘿,我觉得我在其他地方也能做出巨大贡献。’
I think it's a great way to think about the world. It's like very principled, very like hard-nosed scientists, very careful. And I don't know, I think it's just it's such an incredible field. You have such high leverage in computer science, in AI. And so, I think a lot of physicists were seeing that. Particularly in like high energy physics, the discovery of the Higgs, I think a lot of high energy physicists were sort of looking for what's next. Ultimately it becomes bottlenecked on the new um apparatus for, you know, pushing the next energy frontier. And I think a lot of physicists were looking at their skill set and looking at the progress elsewhere and and saying like, 'Hey, I think I could be a huge contributor elsewhere.'
看到弦理论学家、研究黑洞和各种效应的人纷纷转向 AI,这非常迷人。几乎感觉我们像是在重造曼哈顿计划,只不过现在追求的是不同形式的智能。
And it's been fascinating to see like string theorists and people working on black holes and all sorts of effects like kind of moving into AI. It's almost It almost feels like we're recreating a Manhattan project or something except now what we're seeking is, you know, different forms of intelligence. So
是的,没错。
Yeah, that's right.
这个视角挺有意思。抱歉打断一下。那么,你学过物理,研究过暗物质。
Kind of interesting that perspective. Sorry to interrupt. So, you know, you studied physics, you worked on dark matter.
没错。然后基本上,在物理研究生阶段,我总是被机器学习问题吸引。我在研究粒子重建,这本质上就是机器学习问题。但我觉得,如果我真的想推动机器学习的前沿,我应该进入计算机科学领域。所以最终我去了 Google Brain,与第一年的 resident 们重叠。那绝对是一群了不起的人,也是 Google Brain 的一个了不起的时期。我的意思是,那个时代创造了分布式训练策略、混合专家模型、Transformer。那是历史上非常丰富的时期,也是一个有趣的寒武纪时代,人们只用少量 GPU 和很小的合作就在推动前沿。这个领域要早得多,我认为研究中存在很多多样性和熵,非常有趣。
That's right. And then I was basically and then in grad school in physics I was always gravitating towards the machine learning problems. I was looking at particle reconstruction and it's effectively machine learning problems. But it felt if I really wanted to push frontier of machine learning, I should be in, you know, computer science. So ended up at Google Brain, was overlapping with the first year residents there. Absolutely remarkable group of people, remarkable period for Google Brain. I mean this era of when there's the creation of like distributed training strategies, mixture of experts, the transformer. It was a really rich period in that history and it was a fun kind of like Cambrian era where you people were really pushing the frontier with just like a handful of GPUs, really small collaborations. The field was a much much earlier and I think there was a lot of diversity and entropy in the research and it was very fun.
所以大概是 2010 年代末左右?
So it was kind of late 2010s or so, something like that.
那是 2016、2017 年。当时 Google Brain 还很小,最终被 DeepMind 吸收或合并了。所以我在 Google 待了很多年?主要是做架构工作。所以一直在推动稀疏性,这让你能够更高效地大规模服务模型,并真正推动我们所能做到的规模。到 2022 年底,我对产品创造变得非常兴奋。技术变得非常引人注目。所以,我和其他一些 Google 员工一起去了 OpenAI。
This was 2016, 2017. So Google Brain at that point was still very small and eventually was subsumed by DeepMind or combined with DeepMind. So is that Google for many years? Mostly just doing architecture work. So was really pushing um sparsity. That allows for uh you more efficient serving of models at scale and just really pushing the scale of what we could do. Towards late 2022, um really became excited about the creation of products. The technology was getting very compelling. And so, I ended up at OpenAI with um some other Googlers as well.
嗯。你在 OpenAI 具体做什么工作?
Mhm. And what did you work on specifically at OpenAI?
嗯,目标是实现 GPT-4 的产品化。我们 OpenAI 有 GPT-4,它已经预训练好了,还有一些粗略的后训练。问题在于:‘我们如何把这个极其强大的模型变成产品?’我们 brainstorm 了很多想法,比如写作机器人、编程机器人,你知道,当时很自然。嗯,我们最不感兴趣的想法之一是会议机器人。它只是坐在 Google Meet 里,做笔记,然后发送待办事项。但 John Schulman 很有主见。他说:‘我们认为应该保持非常通用。我们做聊天机器人吧。’这成了那几个月努力的一大部分。
Well, so the goal was we need to come up with some productionization of GPT-4. So, we OpenAI had GPT-4. It was pre-trained and there were some like um the rough post-trains on it. And there was questions about like, 'Okay, how do we turn this incredibly powerful model into products?' And we're all spitballing ideas like writing bot, uh coding bot, you know, very natural at the time. Yeah. Some of our least interesting ideas were a meeting bot. So, it would just sit in a Google Meet, take notes, and then send out like to-do's after. But John Schulman was very opinionated. He's like, 'We think we should keep it very general. Let's do chatbot.' And that became a large part of the effort um for those few months.
嗯,没错。好的,所以你做的是 ChatGPT。
Mhm. That's right. Okay, yeah. So, you worked on ChatGPT.
没错。
That's right.
嗯,显然,我觉得那像是整个 AI 革命的发令枪,或者至少是公众意识的起点。我之前就开始投资这个领域了。
Um and obviously, I felt like that was kind of the starting gun of this whole AI revolution or at least in in terms of people's awareness. Like I'd started investing in the area beforehand.
对。
Right.
但在 ChatGPT 出现之前,这几乎是个秘密,突然之间每个人都意识到有这种强大的技术可用。
But it it seemed like almost as a secret up until ChatGPT came out of and suddenly everybody realized that there's this powerful technology available.
是的。
Yes.
那怎么又把你带回到材料、原子和物理世界了呢?我知道那是你学术上的起点,但考虑到现在语言领域正在发生如此大的变革,是什么让你回归?
How did that lead you to materials and atoms and you know, the physical world again? I know that was sort of your starting point in terms of academics, but what brought you back given how much is being transformed right now through language?
我认为只是将这些系统连接到物理世界的必然性。我和其他人在 Sprat Periodic 持有的观点是,除非你开始将这些事物连接到物理世界,否则你不会看到科学和技术有同样的加速。科学最终不是坐在房间里使劲思考。嗯,你必须进行实验,从中学习,必须与现实互动。而 2022 年底 ChatGPT 的创造是一项重要的技术,但它仍然太弱了。比如,我们无法用那个时代的技术做 Periodic。我认为在那之后的几年里,我们看到了不断改进的模型。嗯,我们看到了推理。
I think just the inevitability of connecting these systems to the physical world. The opinion that I and others held at Sprat Periodic was you're not going to see the same kind of acceleration in science and technology unless you start connecting these things to the physical world. Science ultimately isn't sitting in a room thinking really hard. Um you have to conduct experiments, you have to learn from them, you have to interface with reality. And the creation of ChatGPT in late 2022 um was a, you know, important technology, but it was still far too weak. Like we couldn't have done periodic on technology of that era. I think over the next few years past that, we saw ever improving models. Um we saw reasoning.
我认为测试时推理变得非常重要。这带来了更可靠的错误修正和工具使用。我们看到编码智能体和其他智能体的兴起。我认为这些是后来将这些系统连接到物理世界所需的基础技术。这在 2022 年的 AI 技术下根本不可能实现。
I think test-time inference became really important. That led to more reliable error correction, more reliable tool use. And we see the rise of coding agents and other agents. And I think those were foundational technologies necessary to then connect these systems to the physical world. It was just not possible with the AI technology of 2022.
我想物理世界还缺少的另一件事是数据,或者至少是容易获取的数据。你看语言方面的大型基础模型,它们基本上是以互联网为主要语料库训练的,并通过各种方式用其他数据源增强。对于你正在做的事情——试图对物理世界中的原子进行建模——你是如何看待这一点的?
I guess the other thing that's missing from the physical world is data, or at least data that's easily accessible. So you look at something like the big foundation models on the language side, and they're basically trained on the internet as a major corpus. It's augmented in all sorts of ways with other data sources. How do you think about that for what you're doing where you're trying to model atoms in the physical world and how all that stuff kind of works?
所以是实验。我们有物理模拟,也有实验。正如你指出的,机器学习系统在训练数据和训练任务上表现良好。我认为有时存在关于 AGI、ASI、RSI 的神话。我们看到越来越强大的系统,但如果它们无法获取原始数据来做出明智的决策,它们就会受到限制。
So experiment. I mean we have simulation physics simulations, and we have experiment. And, you know, I think exactly as you're pointing out, ML systems are good on the data you've trained them on, on the tasks you've trained them to do. I think sometimes there's this mythology of AGI, ASI, RSI. And I think what we see increasingly powerful systems, but they do become limited if they don't have access to the raw data to actually make informed decisions.
你需要多少数据?我知道有一些关于数据规模的研究,关于如何爬山式地得到一个好模型。你需要运行多少次实验?需要多少数据点?你如何考虑需要生成的数据点的多样性?我有点好奇这实际上是什么样的。
How much data do you need? And I know that there's some data scale related research and other things in terms of how you kind of hill climb towards a really good model. How many experiments do you need to run? Or how many data points do you need? Or how do you think about the diversity of data points you need to generate? I'm a little bit curious what that actually looks like tangibly.
现有模型有一些泛化能力。所以我们不需要重新构建一个能理解和编写英语或代码的系统。我们是在利用……
There is some generalization from the existing models. So we don't need to reproduce a system that can understand and write English or write code. So we're kind of leveraging...
你们是用开源模型还是闭源模型,还是混合使用?
Are you using open source for that or closed source models or some mixture?
混合使用。例如,Periodic 在改进编码模型上投入了零精力。我们对 Codex、Claude Code 印象深刻,这对公司来说是巨大的加速器。但我们把机器学习工作集中在现有前沿不够好的地方。回到数据问题,我们利用了开源模型中约数十万亿的 token。这给了我们非常基础的理解。但一旦我们进入特定的发现领域、化学空间,我们就能看到非常高的样本效率。所以系统不是从随机初始化的神经网络开始的,它对世界有很强的先验知识。
A combination. For example, Periodic spends zero effort on improving coding models. We're incredibly impressed by Codex, Claude Code, and that's been a huge accelerator for the company. But we focus our machine learning efforts where the existing frontiers are not sufficiently good for us. Going back to the data question, we're leveraging on the order of tens of trillions of tokens that went into open source models. And that's given us a very foundational understanding. But once we start moving into specific discovery areas, chemical spaces, we can see a very high level of sample efficiency. So the system isn't starting as a randomly initialized neural net. It has a strong prior on the world.
那么这个先验来自哪里?是什么数据提供了这些信息?只是通用的……
And so where does that prior come from? What data is it that informs that? Just general...
就是论文,是的,正如你指出的互联网。然而,这还不够。我们团队的一名工程师在查看一个报道的材料属性时,发现从文献中提取的值跨越了好几个数量级。所以如果你用这些数据训练一个机器学习系统,你最多只能建模这个分布,但离真实值还很远。这就是实验数据发挥作用的地方,它提供了基础。但非常重要的一点是,它不仅仅是一个数据池,而是一个交互式的闭环系统,非常强大。一旦你有了实验数据,你可以查看它,寻找异常、模式,以及与模拟数据、文献的一致性,然后这有助于驱动下一组实验。所以它不仅仅是一个数据池,而是一个非常活跃的循环。
Just papers, yeah, the internet as you're pointing out. However, that's insufficient. One of the engineers on our team was looking at a reported material property and it was just sort of extracted values from literature and it was really interesting to see the reported value spanned many orders of magnitude. And so you train an ML system on that and it's like well the best you can do is model this distribution but you're no closer to a ground truth. And that's where experimental data comes in where you now have a grounding in this. But really important, it's not just a pool of data. It's this interactive closed loop system that is so powerful. Once you have the experimental data you can look through it. You can look for aberrations. You can look for patterns. You can look for consistency with simulation data, with literature, and then that helps drive the next set of experiments. So it's not just a pool of data. It's this very active loop.
我明白了。那么你如何考虑数据的多样性?我看看像 AlphaFold 或一些蛋白质折叠相关的模型,它们很了不起。我以前是生物学家,一个晶体结构如果成功的话需要数年时间,因为你不能确定是否能在特定试剂条件下结晶出特定蛋白质。然后 AlphaFold 出现了,你可以任意建模蛋白质世界中的任何东西,这是一个惊人的突破。但那是基于一个已经存在的非常特定的数据集,有大量结构,经过数十年的工作。对于每个材料领域,你需要多努力才能启动这样的过程?还是你选择你认为有前景的特定领域然后进行泛化?
I see. And then how do you think about diversity of data? So I look at something like AlphaFold or some of the protein folding related models, which are amazing. I used to work as a biologist, and a crystal structure would take years if it happened at all, because you wouldn't necessarily be certain if you could crystallize a specific protein under certain reagent conditions. And then AlphaFold comes out, and you can just arbitrarily model anything on the protein world, which was an amazing breakthrough. But it was a very specific data set that already existed that had lots and lots of structures. Over decades of work. How hard do you have to bootstrap that for every single material's domain, or do you choose specific ones that you think and then generalize?
我们在内部看到,在数据丰富的领域取得了最大的进步,这带来了最高的加速率。但我认为你可以考虑不同层次的泛化。对于受量子力学效应强烈支配的系统,存在一些泛化。但如果你构建了一个非常精确地建模量子力学对象的系统,它对流体动力学或其他抽象层次帮助不大。所以我们看到的泛化相当不错,但几乎就像是从第一原理出发……
We have seen internally the greatest advances where we have an abundance of data in some space. And that has led to the highest rate of acceleration internally. But I think you can think of different levels of generalization. For systems that are strongly governed by quantum mechanical effects, there is some generalization there. But if you produce a system that has modeled quantum mechanical objects really accurately, it's not really helping much on fluid dynamics or another kind of level of abstraction. And so the generalization we're seeing is quite good, but there's almost like a first principles you can...
哦,这很有趣。所以你可以做类似的事情:这里是化学合成的基本步骤,这里是量子力学,这里是原子一般如何相互作用的不同方面,比如范德华力等等。
Oh, that's so interesting. So you could do like, here are the basic steps of chemical synthesis. Here's quantum mechanics. Here's different aspects of how atoms interact in general or van der Waals forces or things like that.
完全正确。
Absolutely.
哦,这太有趣了。很酷。那么从架构角度来看,你们有什么独特或有趣的做法吗?或者你能谈谈你们实际上是如何构建这些模型的吗?
Oh, that's so interesting. Yeah, that's cool. And then from an architecture perspective, is there anything unique that you're doing or interesting? Or can you talk a little bit about how you're actually constructing some of these models on top?
语言模型非常强大,是一种非常自然的界面。所以我们继续使用它们。但我们几乎把它们看作一个编排层。所以它有点像副驾驶助手,也是一个可以指导实验的系统。它还在编排其他专门的模型。所以我们确实构建了专门为原子系统设计的神经网络,其中包含一些对称性感知。
Language models are incredibly powerful. It's a very natural interface. And so we continue to use these. But we think about them almost as an orchestration layer. So that's sort of a co-pilot assistant, but also like a system that can direct experiments. And it's orchestrating other specialized models as well. So we do construct neural nets that are specially designed for atomic systems, where there's some symmetry awareness.
嗯,这些模型延迟更低,而且已经针对这一点进行了微调。所以基本上,你可以把它看作一个编排层,能够吸收文献、处理我们的实验数据、处理不同的模态,同时还能使用专门的神经网络作为工具和奖励函数。所以这是一个整体系统。
Um and those have much lower latency, and they've been like fine-tuned for that. And so, basically, you can kind of think of this like orchestrating layer that can ingest literature, it can go through our experimental data, it can go through different modalities, but they can also use specialized neural nets as tools, as reward functions. So, it's like an overall system.
好的,这很有道理。我看到很多人也在为客服或其他领域设计这类方法。看起来这是随着这些模型在不同用例中应用而涌现出的常见架构。
Okay. Yeah, that makes a lot of sense. Yeah, I've seen a lot of people architect those sorts of approaches even for things like customer support or other areas. Like, it seems like it's the common architecture that's emerging as you're doing these different use cases of these models. Yeah.
是的,但 Transformer 一直非常强大。
Yeah. But, transformers have been very powerful.
是的,这真的很酷。所以,如果看语言领域,它有一个非常独特的地方,也是我认为 OpenAI、Anthropic 等公司发展如此迅速的原因,就是它直接嵌入了一个人类存在的巨大领域——所有语言。而所有语言意味着企业软件和企业交互,也意味着消费者行为。这基本上就是我们与世界互动的方式。
Yeah, yeah. And that's really cool. So, if I look at the language world, one of the things that was pretty unique about it, and it's the reason that I think these companies like OpenAI, Anthropic, and others are growing so fast, is it just plugged into a very big domain of human existence, which is all language. And all language means enterprise software and enterprise interactions, and it means consumer behavior. It's basically how we interact with the world.
是的。
Yes.
嗯,其他领域似乎需要更大的跨越。例如,在机器人领域,世界上存在各种有趣的机器人,但它们的足迹相对于语言来说非常有限。材料科学似乎也是如此。那么,你如何看待首先在哪里商业化,或者与谁合作,或者你们是否在优先开发特定的产品领域?
Um it seems like there's a little bit more of a leap for other areas. So, for example, in robotics, there's really interesting things different types of robots that exist in the world, but the footprint of that is quite limited relative to language. And the same seems to be true for material sciences. So, how do you think about where you're going to commercialize this first, or who you're going to work with, or are there specific domains of products that you're working on first?
我们开始与科学家密切合作。我们把 Periodic 视为我们的零号客户,看看如何改变这个科学领域的研究方式。但在所有与物理世界打交道的行业和企业中,都存在巨大的机会。那些受限于材料工程、工艺工程的人。同样,这些也是工程师们询问数据、寻找异常、调试机器、寻求更好配方的自然界面。这实际上也是一个相当普遍的事情。所以我们在内部创建了一个小测试场,现在我们对正在构建的技术感到非常兴奋,并希望看到它更广泛地加速先进制造业。
So, we've begun working very closely with scientists. We've treated periodic as our customer zero and seeing how can we transform how this field of science is done. But there's huge opportunities across all of these industries, all these enterprises that are interfacing with the physical world. People who are bottlenecked by materials engineering, process engineering. And again, those are kind of the same natural interfaces where engineers are asking questions about their data, they're trying to find aberrations, they're trying to debug machinery, they're trying to get to a better formulation. It's actually a quite universal thing as well. And so we've kind of created our little testing ground internally and now we're sufficiently excited about the tech we've been building and to see this acceleration for advanced manufacturing more broadly.
你们的模型是为第三方开发材料,还是开发自己的材料然后在市场上销售?因为这让我有点想起生物技术模式。在生物技术领域,你可以与大药企合作,帮助他们开发药物并收取版税,或者自己开发药物。在你们所做的背景下,你们如何看待这一点?
And is your model going to be developing materials for other third parties? Is it developing your own materials that you then sell in the market? Like because it almost reminds me a little bit of a biotech model. Yeah, where in biotech you can either partner with a big pharma and then effectively help them create a drug and take a royalty on it or you can build your own drugs. How do you think about that in the context of what you're doing?
我们把自己视为这些公司的智能层。所以你可以把它看作记录系统、不同实验的控制平面以及达成解决方案。但就像你说的,这里的一些突破可能具有非常高的价值,可能更类似于我们在生物技术等领域看到的发现模式。但一开始,我们只是把它当作一个软件业务来考虑。
We're thinking about us ourselves as an intelligence layer for these companies. So you can think about system of record, control plane for different experiments and getting to solutions. Um but like you're saying, there is a very interesting aspect of some breakthroughs here could have you know really high value and it might be more akin to a discovery model like we've seen in biotech and elsewhere. But starting thinking about our just as a software business.
你听说过《钻石时代》吗?你听说过《钻石时代》吗?
Have you ever heard of the Diamond Age? Yeah. Have you heard of the Diamond Age?
不,实际上我没听说过。
No, I haven't actually.
尼尔·斯蒂芬森的书。这本书写于 90 年代,里面有两个关键概念。一个关键概念是有一个 AI 导师被释放到世界上,它教会了大量年轻女孩各种技能,这是一个关于 AI 教育的非常有趣的事情。与此同时,
Neil Stephenson book. It's basically this book about It was written in the 90s okay and there's two key concepts in it. One key concept is there's effectively an AI tutor that's unleashed on the world and it kind of teaches huge numbers of young girls all sorts of skills and it's a very interesting thing about AI education and then in parallel,
为什么特别针对年轻女孩?
Why young girls in particular?
嗯,基本上,一位 AI 研究科学家为他的女儿创建了一本入门书,然后中国人偷走了它,克隆并分发到全国,因为他是为年轻女孩构建的,所以突然之间中国每个年轻女孩都有了它。
Uh, basically this, AI research scientist creates a primer for his daughter and the Chinese steal it and clone it and distribute it across the country and because he built it for young girls, it suddenly every young girl in China has it.
对,对。
Right, right.
原因。这是一个非常典型的中国知识产权盗窃事件。
The reason. It's this very, China theft of IP kind of thing.
是的,没错。
Yes, right.
书的另一部分是关于物质管道进入每个人的家中,他们都有 3D 打印机,你下载蓝图,它就能在物理世界中创造你需要的任何东西,有些人开始进化出不同的纳米机器人来做不同的事情。这是一个非常先进的 AI 加材料的未来世界。
The other part of the book is about matter pipes into everybody's homes and they all have 3D printers and you download blueprints and it just creates whatever you need in the physical world and some people start evolving different nanobots to do different things. It's this very advanced kind of AI plus materials kind of future world.
是的。
Yes.
嗯,假设 PeriodX 成功,你对 10 年后世界的愿景或概念是什么?
Um, what is your vision or conception of what our world looks like in 10 years assuming PeriodX is successful?
嗯,我的意思是,正如你指出的,我们正在从不仅仅是写文章、写软件的系统,转向真正生成物质的系统。我认为这对半导体、航空航天、能源有着深远的影响,而且对于能否加快世界物理发展的步伐至关重要。我们看到数字领域变化有多快。软件工程现在看起来甚至与 6 个月前截然不同。但我认为我们在物理世界也看到了类似的机会。当然,原子是硬的,所以会有一些物理限制,但仅仅因为原子是硬的,并不意味着不能加速一两个数量级。只是理解大量数据并更快地找到解决方案。嗯,是的,所以我认为我们正在努力做的是赋予人类这种原子重排和合成的能力,我们认为这将是一个巨大的加速器。所以,如果我们的物理世界能以数字世界的几分之一速度跟上,我认为生活将会变得截然不同。
Well, I mean, I think, as you're pointing out, you're going from systems that aren't just writing essays, not just writing software, but to literally generating matter. And I think it's has pretty profound implications to semiconductors, aerospace, energy and I think it's incredibly important for can we increase like the pace of just like the physical development of the world. I mean, we see how quickly the digital realm is changing. Software engineering now looks wildly different than even 6 months ago. Um, but I think we see like, you know, similar opportunities in the physical world. Of course, like atoms are hard and so you will have, some limits of physics, but just because atoms are hard doesn't mean there's not an order of magnitude or two to speed up. Um, just making sense of huge amounts of data and getting to solutions more quickly. Um, yeah, so I think what we're trying to do is give humanity this agency for atomic rearrangement, synthesis and we think it's going to just be a huge accelerator. So, I mean if our physical world could keep up at some fraction to our digital world, I think life will just feel dramatically different.
这确实可能是一场革命。是的,这让我想起材料领域的农业革命。在那里,产出生产力突然大幅飙升。
Kind of the revolution that could really come. Yeah, it kind of reminds me of almost the materials equivalent of the agricultural revolution. Yeah, where you suddenly had a massive spike in productivity of output.
完全正确。似乎一直有各种瓶颈在限制我们,而你们正在努力解决这些问题。没错。
Exactly. And it seems like there's been all sorts of bottlenecks that have constrained us until now that you folks are trying to address. That's right.
是的。你对自己正在做的工作中最兴奋的是哪方面?
Yeah. What aspect of the work that you're doing are you most excited about?
我们这些群体之间的迭代。我的意思是,这本质上就是一个多学科问题。我们有物理学家和化学家与世界上一些顶尖的 AI 研究人员密切合作,再与一些最优秀的工程师密切合作。
The iteration with our between these groups of people. I mean, it's like this is just irreducibly a multi-disciplinary problem. We have physicists and chemists working really closely with some of the top AI researchers in the world, working closely with some of the best engineers in the world.
这种多学科、真正紧密的合作简直不可思议,因为亲眼目睹一个领域如何发生根本性变化。有些人在某个领域做了几十年的研究,现在看到,哦,在这些系统下,在智能系统下,它可能看起来非常不同。我经常用机器学习来类比。回到早期的 Google Brain 时代,前沿是由几个 GPU 和几个人推动的。现在你看这个时代,它真的工业化了,有几十、几百名研究人员与几十万、数百万个 GPU 一起工作,由缩放定律决定和驱动。一切都是关于 Scaling(规模扩张)。它带来了可预测性。它让我们能够向这个领域投入大量资本。我认为物理科学、物理工程也会有非常相似的特性,我们会建立这些缩放属性并带入这种思维。所以,这个领域的周期性真正思考的是如何将更大规模的实验集应用于此。智能系统实现了这一点,自动化实现了这一点,而且你确实需要两者。自动化的改进可能会很快在智能方面造成瓶颈。科学家们非常感受到这一点,他们不习惯在那种吞吐量水平下工作。他们根本无法理解这么多数据。
And this multi-disciplinary, really close collaboration is just absolutely incredible because seeing firsthand how a field can fundamentally change. People who have been doing research for in some cases decades in a field and now seeing, oh, under these systems, under intelligent systems it could look very different. And I use an analogy to machine learning a lot. Going back to the early Google Brain days where the frontier is pushed forward by a few GPUs and a few people. Now you look at this era where it's really industrialized and there are dozens, hundreds of researchers working together with hundreds of thousands, millions of GPUs dictated and driven by scaling laws. Everything is about scaling. It's given that predictability. It's allowed us to put huge amounts of capital into this field. And I think the physical sciences, physical engineering will have a very similar property where we establish these scaling properties and bring that mindset. So, Periodic in this field is really thinking about how do we bring much larger scale sets of experiments to bear on this. And intelligent systems have enabled this, automation has enabled this, and you really need both. An improvement to automation where you can soon become create bottlenecks in intelligence. And the scientists very much feel this where they're not used to working at that level of throughput. And they just can't simply make sense of so much data.
有意思。是的,所以我想在规模方面,真正让前沿实验室在 LLM 方面受益的一点就是资本规模,以及由此带来的 GPU 规模和数据规模。
It's interesting. Yeah, so I guess in terms of scale here, one of the things that's really benefited the frontier labs on the LLM side is just scale of capital and therefore scale of GPU and scale of data.
当然。
Of course.
在你看来,这同样是一个资本密集型领域吗?
Is this similarly a capital intensive area in your mind?
是的,我们需要更多资本。GPU 非常昂贵。有趣的是,算力成本相对于物理基础设施实际上令人惊讶,你知道,那么多钱花在算力上,物理基础设施有时反而更低,但它的前置时间很长,而且要让这些校准良好、运行正常的物理系统存在内在困难。但从资本角度来看,主要是算力成本。
Yeah, we will require more capital. GPUs are so extraordinarily expensive. What was interesting is just the compute cost relative to physical infrastructure is actually surprising where you know, just so much money is spent on the compute that the physical infrastructure sometimes is actually lower, but it has very large lead times and there's intrinsic difficulty of having these well-calibrated, well-functioning physical systems. But from a capital perspective, it's primarily a compute cost.
是的,这真的很有趣。如果你看看斯坦福博士后相对于机器学习工程师的成本,差别很大。我的结论是,许多从事科学工作的人,尤其是在学术中心环境中,相对于他们的社会价值,报酬过低。
Yeah, it's really interesting. If you look up the cost of a Stanford postdoc, for example, relative to a machine learning engineer, it's like such a big difference. And my take away is that many people working in science, particularly in academic center setting, are very under-compensated relative to their societal value.
绝对如此。
Absolutely.
是的,所以我总是喜欢公司帮助人们融入,既在人类影响方面,也在真正大规模做事和以不同方式做事的能力方面。所以,这对你团队的人来说一定非常令人兴奋。
Yeah, and so I always like it when companies kind of help bring people into the fold in terms of both human impact, but also that ability to do things at real scale and really do things a different way. So, it must be very exciting for the people on your team.
是的,我的意思是,一些加入我们的科学家是世界上最好的,与他们合作绝对不可思议。
Yeah, I mean it's like some of the scientists who joined us are among the best in the world and it's been absolutely incredible working with them.
是的,听起来你建立了一个如此惊人的跨学科团队。你现在有没有在积极寻找的特定角色,或者你真正想招聘的关键岗位?
Yeah, I mean, it sounds like you've built such an amazing interdisciplinary team. Are there specific roles you're actively looking for right now or key things that you really want to hire up?
当然。在我们的网站上,我们把世界分解成了比特和原子。这是一个松散的分类,但在比特方面,我们真正考虑的是 AI 方面的中训练、预训练角色,总是有更多基础设施角色,而在原子方面,比如控制工程、系统工程,但现在也在考虑与产品工程结合。所以,有很多活跃的职位等等。
Absolutely. So, on our site, we have decomposed the world into bits and atoms. It's a loose taxonomy, but on the bit side, we're really thinking about mid-training, pre-training roles from the AI side, always more infrastructure roles, and on the atom side, like control engineering, system engineering, but also now thinking about spanning that with product engineering. So, a lot of active roles, etc.
全面开花。
Across the board.
是的,那真的很酷。所以,我认为现在每个人都在深入思考或感到兴奋的一件事是 AGI、ASI,这些先进系统在不同事情上与人类一样好或更好,或者在广泛事物上的能力非常通用。你如何看待这在整体基础模型曲线背景下发生的事情?因为你显然在开发其中一些系统方面发挥了重要作用,然后你如何看待将其具体应用于你正在工作的某些领域?
Yeah, that's really cool. So, I think one of the things that everybody's really thinking deeply about or is excited about right now is AGI, ASI, sort of these advanced systems that are as good as humans or better than humans at different things or are very generalizable in terms of their abilities to do a broad swath of things. How do you think about that within the context of what's happening over the overall foundation model curve? Cuz obviously you were very integral in terms of the development of some of these systems, and then how do you think about that applied specifically to some of the areas you're working in?
我认为一个谬误是把智能看作一个标量。我们一直看到这些系统有非常奇怪的尖峰性,实际上可以构建一个在某些数学领域是世界级的系统,但如果你对问题做一些扰动,它就会大幅退化。所以,就像一个差劲的高中生。因此,这些系统存在这种奇怪的尖峰性。
I think one fallacy is thinking about intelligence as a scalar. We've consistently seen these systems have a very odd spikiness, and it's actually possible to architect a system that is world-class on some math domain, but then you could do some perturbations to the questions and actually degrade it substantially. So, it's like a bad high school student. And so, there's this odd spikiness to these systems.
所以,基本上你可以制造一个系统,它在某件事上像天才,但在其他很多事情上不太擅长。
So, basically you can make a system that's like a genius at one thing and not very good at a bunch of other stuff.
我想我要说的是,这些领域实际上可能非常接近。所以,有时泛化可能不直观。但我认为递归自我改进的一种方式类似于大约 10 年前的神经架构搜索。我认为软件工程有一条非常清晰的路径。因此,这些系统在这个领域变得如此令人难以置信地令人印象深刻,这得益于大量数据、非常便宜的可验证环境,比如你可以用几个 CPU 检查单元测试从失败到通过。基本上是瞬间的。AI 研究员和软件工程师之间没有领域专业知识差距。显然,这将成为并且正在成为下一代系统的更大贡献者。
And I guess the point I was making is those fields can actually be quite adjacent. So, sometimes the generalization can be non-intuitive. But one way I think about recursive self-improvement is really kind of akin to neural architecture search from roughly 10 years ago. And I think there's a very clear path for software engineering. So, these systems have become so incredibly impressive on this domain as a result of huge amounts of data, really cheap verifiable environments, like you can check unit tests go from failing to passing with just a few CPUs. It's basically instantaneous. There's no domain expertise gap between an AI researcher and a software engineer. And obviously this will become and is becoming a larger contributor to the next generation of the system.
你认为什么时候会变成一切都是机器自我改进,而不是人类指导或需要大量人类干预?所以,你认为那是 2 年后?5 年后?10 年后?
When do you think it just flips into we just everything is machine self-improvement versus human directed or needs a lot of human intervention. So, do you think that's 2 years away? Do you think it's 5 years away? Do you think it's 10 years away?
嗯,我想基于我所说的,我认为有一个领域上的注意事项。所以,推进软件工程自我改进,我认为你会有一个系统可以编写完整的代码库、识别错误、重构代码,但它不会突然理解生物学。对吧?就像那里存在知识上的领域差距。但除此之外,软件工程中有一套策略与科学或工程策略不同。
Well, I guess building on what I was saying is I think there's a domain caveat to that. So, rolling forward that software engineering self-improvement, I think you're going to have a system that can write complete repositories, identify bugs, refactor code, but it doesn't suddenly understand biology. Right? It's just like there's a domain gap there in knowledge. But even beyond that, there are sets of strategies done in software engineering that differ from scientific or engineering strategies.
所以,它不像是在不确定性下做决策。它是非常可验证的,这推动了我们的很多工作。
So, it's not like decision-making under uncertainty to the same degree. It's very verifiable and that's driven so much of our work.
嗯。
Mhm.
在那个领域,我认为现在差不多正在发生。
In that domain, I think it's happening now-ish.
嗯。
Mhm.
而且我认为我们在 AI 研究中也会看到同样的情况。
And I think we'll see the same thing too for AI research.
嗯哼。
Uh-huh.
那是一个更慢的外循环,因为现在的实验不只是检查一些单元测试是否通过,而是检查缩放属性是什么?这个模型收敛了吗?系统的泛化能力如何?这需要 GPU,需要很多小时的实验,但我认为这也会发生。
That's a slower outer loop because now the experiment isn't just checking some unit tests passing, but it's checking what was the scaling property? Did this model converge? What's the generalization of the system? That requires GPUs, many hours of experiments, but I think that will also happen.
这些都是人们在评估现有模型时使用的评估方法,所以它们确实有那个效用函数,那个可以由自我学习驱动的反馈循环。
And those are all evals that people use today as they're looking at existing models and so they do have that utility function, that feedback loop that can be just driven by self-learning.
没错,没错。但同样,这些东西与物理世界的连接将非常关键,因为这两个系统都是在针对那个领域的闭环中训练的。所以这是一个用于软件工程的闭环,一个用于 AI 研究的闭环。这就是 Periodic 的前提。我们需要这些实际做科学、实际做工程的闭环。而这两个领域,我认为世界其他领域会以一定的延迟跟进。这又是我们正在构建的基础技术。
That's right. That's right. But again, the connection of these things to the physical world is going to be so critical because both of these systems are being trained in a closed loop against that domain. So it's a closed loop for doing software engineering, a closed loop for doing AI research. And that's the premise of Periodic. Like we need to have these closed loops of actually doing science, of actually doing engineering. And these two domains are how I think the rest of the world will go with some delay. And this is again the foundational technology that we're building.
非常有趣。你认为你需要足够好的机器人系统才能为你正在做的事情建立那个闭环吗?
Super interesting. Do you think you need sufficiently good robotic systems in order to have that closed loop for what you're doing?
不,但它是一个巨大的加速器。
No, but it's a huge accelerator.
换句话说,你需要像 pie 或 skulls 这样的东西才能让 Periodic 在闭环系统方面达到逃逸速度吗?
In other words, do you need something like pie or skulls or something else to work in order for Periodic to hit that escape velocity in terms of a closed loop system?
不,但它是一个巨大的加速器。Periodic 的目标是生成大量高质量、多样化的数据。自动化对此是辅助。所以现在我们也在雇佣人员,并且我们有非常可靠的自主部分。如果你有一个灵巧的人形机器人,能够走进一个非结构化的实验室,理解并可靠地遵循指令,那将是一个巨大的加速器。目前物理系统的自动化需要非常仔细的设计,而且很慢,但我认为随着机器人技术的改进,这将会加速。但这些混合系统的可靠性已经足以产生大量可靠的数据,只是会进一步加速我们。
No, but it's a huge accelerator. The goal for Periodic is to generate high quantity, high quality data, diverse data. And automation is assistance to that. So right now we employ people as well and we have autonomous parts that are very reliable. If you had a dextrous humanoid who could wander into an unstructured lab and make sense and follow instructions reliably, that would be a huge accelerator. Right now the automation of physical systems requires very careful design and it's slow, but I think with improvements in robotics, this is going to accelerate. But already the reliability of these hybrid systems is sufficient to produce huge amounts of reliable data, but it's just going to accelerate us further.
是的,我问的原因之一是这家公司 Color 的用户。我们构建了自己的液体处理机器人系统。我们购买液体处理机器人,但必须大幅调整它们。我们有摄像头使用机器学习来监控系统并进行调整。我们必须 3D 打印零件以减少平台上的振动,因为我们处理的是非常小体积的液体。所以有大量的定制化,而固件中的功能很糟糕,编写代码也很痛苦。相比之下,如果有一个像现代系统那样工作的机器人系统,那就好了。这就是我问的原因:如果你真的想做高通量实验,你需要这些底层系统能够完成所有的液体处理、滴定等工作。
Yeah, one of the reasons I ask is a user in this company Color. We built our own liquid handling robotic systems. We'd buy liquid handling robots but then we have to adjust them dramatically. We had cameras that would use ML to monitor the system and make adjustments. We had to 3D print parts to decrease vibrations on the platform because we were dealing with such small volumes of liquid. So there was an enormous amount of customization versus just having in the firmware for it was awful and writing against that was painful. Versus just having a robotic system that would work like a modern system in all the ways that you'd conceive. So that's the reason I was asking is if you really want to do high throughput experiments, you need these underlying systems to be able to do all the liquid handling and titration and stuff.
是的,没错。我认为现在我们几乎更多是使用现成的机器人。它们非常简单,非常商品化。在这方面没有进行大量的创新。但同样,随着这些更通用的机器人系统出现并达到这个可靠性阈值,它也将成为启动新实验室的巨大加速器。
Yeah, that's right. I think right now we're using almost more like off-the-shelf robotics. It's very simple, very commoditized. Not doing a huge amount of innovation on that front. But again, as these more general robotic systems come to be and hit this reliability threshold, it's going to be a massive accelerator for spinning up new labs as well.
自从大约十年前你在谷歌工作以来,你见证了 AI 世界发生的各种不同的事情。你在 Transformer 模型诞生时就在那里。你在 ChatGPT 诞生时也在那里。除了 Periodic 之外,未来几年在 AI 方面你最兴奋的是什么?
You've seen such a wide range of different things happen in the AI world since you were working at Google about a decade ago. You were there during the birth of the transformer model. You were there for the birth of ChatGPT. What are you most excited about outside of Periodic over the next few years in terms of what's happening with AI?
我的意思是,当然是机器人技术。再次,我对 AI 系统与物理世界的接口感到非常兴奋,我们正在接近其中的一个角度,那就是科学工程,我们需要这些数据来取得进展。但仅仅是通过机器人技术对物理世界的代理和控制就将具有变革性。所以我对这些接口层非常兴奋。我认为这将是一个巨大的机会。因为,你知道,世界上有多少软件工程师,相比之下,从事物理世界工作的人有多少?而且到处都有劳动力短缺。所以,是的,我认为这将是一个非常有趣的十年。
I mean, of course robotics. Again, I'm just so excited about the interface of AI systems with the physical world and we're approaching one angle of that, which is science engineering and we need that data in order to make those advances. But simply just agency and control of the physical world via robotics is going to be transformative. So I'm very excited about these interface layers. I think that's going to be such a massive opportunity. Because, you know, how many software engineers are there in the world versus people who are like the physical world? And there's just labor shortages everywhere. So, yeah, I think it's going to be a very interesting decade.
哦,太棒了。非常感谢你今天加入我们。
Oh, amazing. Well, thank you so much for joining us today.
是的,非常感谢。今天聊得很愉快。
Yeah, well, thank you so much. That was very good chatting today.