The Next Frontier: Scaling AI with Scientific Data
打开互动全文版(中英对照 + 朗读 + 问答)→来自 LILA Science 的 Rafa Gomez-Bombarelli 和 Andy Beam 讨论未来实验室应如何像数据中心一样,以科学数据作为新前沿来扩展 AI 规模。
Rafa Gomez-Bombarelli and Andy Beam from LILA Science discuss how the lab of the future should feel like a data center, scaling AI with scientific data as the new frontier.
欢迎收听 Latent Space Science。我是 Brandon,今天和我的联合主持人 RJ 一起。我们有幸请到了 LILA Science 的 Rafa Gomez Bambarelli 和 Andy Beam。我们就直接开始,请你们自我介绍一下?
Welcome to Latent Space Science. I'm Brandon, I'm here with my co-host RJ. Today we have Rafa Gomez Bambarelli and Andy Beam from LILA Science. We'll just start off and will you introduce yourself?
谢谢邀请我们上播客。老听众,第一次打电话。很兴奋能来这里。我是 Andy,LILA 的首席技术官。我从事 AI 研究已经大约 20 年了,可以追溯到深度学习之前的日子,支持向量机、随机森林之类的。我在 2010 到 2014 年左右读了神经网络方向的博士,正好是深度学习兴起的时候。很明显神经网络是值得押注的方向,但自动求导库当时还没开发出来。所以我手动做反向传播,就像当年上学时翻山越岭一样。我对 AI 在医疗和生命科学领域的应用非常感兴趣。我妻子是医生,所以我看着她应对各种困难,觉得 AI 显然是很多问题的自然解决方案。我在哈佛医学院做了博士后,从事早期医疗 AI 工作。我真正热爱的是 AI,很想知道 AI 能解决哪些问题。但我也一直对创业充满好奇。所以我离开学术界一年,帮助创办了一家名为 Generate Biomedicines 的公司,这是一家早期的生成式生物学公司。我是那里的机器学习创始负责人,在接下来的五六年里,我享受了那种教授兼创业者的混合角色。我在哈佛有一个实验室,介于公共卫生学院和医学院之间,既做方法研究,也做大量应用工作。那是一份很棒的工作,但我感觉到 AI 的浪潮正在发生重大变化。我想参与其中。所以我开始思考在哪里可以从事 AI 前沿的工作,解决真正令人兴奋的问题。学术界有很多优点,但获取大规模算力或大规模资源不是它的强项。我早期是 LILA 的顾问,当他们的核心理念成型后,我非常兴奋。基本上,科学是一个无限的 token 生成器,可以用来大规模训练模型。除了创建一个能解决科学问题的新前沿模型,我还有什么理由去做其他事呢?所以我开玩笑说,两年前我挂起了花呢夹克,离开了学术界,全职加入 LILA 担任首任 CTO。
Yeah, thanks for having us on the podcast. Long time listener, first time caller. Excited to be here. I'm Andy, I'm the chief technology officer at LILA. I've been an AI researcher now for something like 20 years going back to the pre-deep learning days, SVMs, random forest, things like that. Did a neural net PhD around 2010 to 2014 right as deep learning was taking off. It was clear neural nets were the thing to back, but auto grad libraries really hadn't been developed yet. So I did the back prop by hand, walking uphill both ways kind of thing. Got very interested in AI for healthcare and life sciences. My wife's a physician, so I watched her struggle through different things and thought that AI was obviously a natural solution for a lot of those problems. Did a post doc at Harvard in the medical school doing early work on medical AI. I'm in it for the AI. I was really interested in what problems could AI solve. But I've also always been startup curious. So I took a break from academia for a year and helped start a company called Generate Biomedicines, which was an early generative biology company. I was the founding head of machine learning there and got to do the fun kind of hybrid professor startup founder thing for the next five or six years. So I had a lab at Harvard, between the School of Public Health and the Medical School, doing methods research but also a lot of applied work. Those were a great set of jobs, but I got a sense that the AI moment was changing in a very significant way. And I wanted to be a part of it. So I started to think about where could I work at the frontier of AI and on really exciting problems. Academia has a lot going for it. Access to scaled compute is not one of the things that it has going for it or scaled resources. So I'd been an early advisor for LILA and got very excited once the thesis crystallized. But basically, science is an infinite token generator to train models at scale. Why would I want to work on anything other than creating a new frontier model that can solve scientific problems? So I kind of joke that I hung up the tweed jacket two years ago, left my position in academia and joined LILA full-time as the inaugural CTO.
我叫 Rafa,是 LILA 物理科学首席科学官兼联合创始人。我过去是一名计算化学家。我们使用一种商品化资源——算力。很明显,我们可以扩大算力来做分子模拟,这产生了足够多的数据,以至于在 2010 年代初我们意识到自己遇到了数据问题,我的方向也由此转变。我和 David Duvenaud 以及 Ryan Adams 合作,融合了当时感觉像是深度学习在科学中的首批实例。我是最早将生成式 AI 用于化学的人之一,使用自编码器处理 token 化的分子。我深深爱上了潜在空间。实际上,我们有一个和你们非常相似的标志,不过是针对分子的,那个图已经自成一派了。这会是刻在你墓碑上的图。没错。我的学生们有一个 Slack 频道,专门在它出现时发出来。我的故事没有 Andy 那么丰富,但在 2015-2016 年时期也经历了同样的转变。我在哈佛博士后期间创办了一家公司,是一家计算材料平台公司,然后去了麻省理工学院,在那里建立了我的材料科学与工程研究组。该组的工作处于分子模拟和 AI 的交界处,比如材料结构的生成模型、用于分子模拟中我们想看到的非常酷的梯度的自动求导。到 2022-2023 年,事情正如 Andy 刚才提到的方向转变。我们看到了苦涩的教训在计算生成数据上的体现。这就是为什么 Meta、DeepMind 和微软都有团队在做 AI 计算材料科学。但很明显,我们需要弥合差距,让 AI 服务于真正的材料科学,而不仅仅是计算版本。这正好与再次创业的机会吻合。我现在非常兴奋,正在推动这个整合的愿景,即跨越所有我们可以在实验室验证的科学模态进行科学推理。
I go by Rafa. I'm the chief scientific officer for physical sciences at LILA and a co-founder. I was a computational chemist back in the day. We used a commodity resource that is compute. So it was clear that we could scale up the compute to do molecular simulations and that produced enough data that in the early 2010s we realized we had a data problem and things switched gear for me right around then. I worked with David Duvenaud and Ryan Adams in blending what felt like the first instances of deep learning for science. I was one of the first people to do generative AI for chemistry with an auto encoder on tokenized molecules. I fell deeply in love with latent spaces. We actually have a very similar logo to yours, but for molecules, and that figure has taken its own life. This is the one that will be on your tombstone. Exactly. And my students have a Slack channel just to post it when it shows up in the wild. Not as much of a story as Andy, but the same converting in the 2015-2016 era. I spun out a company out of my postdoc at Harvard, a computational materials platform company, and then went to MIT, where I started my group in material science and engineering. The group there was working at the interface of molecular simulations and AI with things like generative models for material structure, autograd for really cool gradients that we want to see in molecular simulations. By 2022-2023, things were taking the turn that Andy just mentioned. We had seen the bitter lesson come to computationally generated data. That's the reason why Meta, DeepMind, and Microsoft have teams doing AI for computational material science. But it was clear that we needed to bridge a gap and make AI for actual material science, not just the computational version. That lined up with the opportunity to start spinning out something again. I'm very excited now to be pushing this integrated vision of scientific reasoning across all the modalities of science we can validate in the lab.
好的。这让我想到 LILA 的核心理念是什么?你们似乎有一个非常宏大的目标。
All right. That brings me to what is LILA's thesis? It seems like you have a very ambitious goal here.
好问题。我尽量给你一个 TLDR,然后我们可以深入几层。所以,就像 Rafa 说的,我们全力押注苦涩的教训和 Scaling。我们认为那些能扩展且通用的方法会击败那些不能的。这听起来直白无误,但实际上与 AI 研究 70 年历史中的大部分直觉相悖。我们的认识是,过去 4、5、6 年大型语言模型的兴起,源于大规模算力和大规模数据的结合。那些数据来自互联网,是人类生成的,而且我们已经用完了。正如 Ilya 去年在 NeurIPS 上所说,我们只有一个互联网。它就像我们开采的化石燃料。我们从中榨取了每一盎司的数据,但它已经用完了。
Yeah, it's a great question. I'll try and give you the TLDR and then we can go a couple levels deeper. So, like Rafa said, we are all in on the bitter lesson and scale. We think that methods that scale and that are general beat those that are not. That actually sounds straightforwardly true, but is actually counterintuitive and contrary to much of the 70-year history of AI research. But the realization that we had is that what gave rise to large language models over the last 4, 5, 6 years is the combination of scale compute and scale data. That data came from the internet. It was human generated and we have used it all. As Ilya said at NeurIPS last year, we have but one internet. It's the fossil fuel we fracked. We got every ounce of data that we could out of the internet, but it's gone.
你知道,人们通常谈论不同的 Scaling 维度。有算力,有数据,而对于科学来说,数据不一定是无限的资源。你的意思是,我们现在想为数据增加一个新的 Scaling 维度。
You know, people normally talk about different scaling axes. You have compute, you have data, and for science, data is not necessarily an infinite resource. And your point is that we now want to add a new scaling axis for data.
我们认为未来的实验室应该像一个数据中心。一排排服务器机架尽可能密集,同时也尽可能节能,诸如此类。
We think that the lab of the future should feel like a data center. Rows of server racks as densely packed as possible and also as energy efficient as possible and things like that.
所以 AI 的问题是:下一个互联网规模的数据集从哪里来?在后预训练时代,我们进入了带可验证奖励的强化学习。人们经常谈论 RL,但 RL 本质上是一种让模型自己生成数据的方式,奖励信号会强化好的数据、惩罚坏的数据。这在数学和编码问题上已经是一个非常有效的框架。但在 Lyra,我们相信,科学本身——运行科学方法、用自然和实验作为验证器——才是这个框架的终极版本。所以我们正在构建的东西,我们称之为 AI 科学工厂。它们是面向科学的规模化验证器,让我们能够大规模进行后训练,并推动推理模型的能力边界。这就是我们的核心论点。
And so the question AI is like, where is the next internet scale data set coming from? Post the pre-training era, we moved into reinforcement learning with verifiable rewards. So people talk about RL a lot. But really what RL is is a way for a model to generate its own data and the reward signal reinforces good data and penalizes bad data. So that has been a very productive framework for problems in math and coding. But what at Lyra we believe is that actually science — running the scientific method and using nature and experiments as verifier — is like the ultimate version of that. And so what we're building, we'll talk about these things that we call AI science factories. They are scaled verifiers for science so that we can do post-training at scale and push out the frontier of what reasoning models are capable of. So that's like the thesis in a nutshell.
你的提议基本上是,人们通常谈论不同的 Scaling 轴:算力、数据、参数。而对于科学来说,数据不一定是无限的资源。你的观点是,我们现在想要有一个新的数据 Scaling 轴。
Your proposal is basically, you know, people normally talk about different scaling axes. You have compute, you have data, you know, you have parameters. And for science, data is not necessarily an infinite resource. And your point is that we now want to have a new scaling axis for data.
没错。
Correct.
所以我引用一下我在 Escalate Bio 的一些朋友写的一篇博客,非常好的博客,我推荐你读一读。它说:“你的实验有一个运行时间。”那么,你的数据收集的运行时间是多少?
So, I quote some of my friends at the Escalate Bio have a blog post, really good blog post, I recommend you read it. It says, "Your experiment has a run time." So, what is the run time of your data collection?
这是个很棒的问题。显然,它因实验而异。比如,你不能让核糖体跑得更快,至少据我所知是这样。生物学设定了速度的上限。在材料科学和化学中,时间尺度更小,长度尺度更大。但你真正问的是一个技术问题:如何训练一个模型,使其面对反馈量级相差几个数量级的反馈机制?所以我们把整个 Lyra 视为能够在不同长度尺度上生成不同类型数据的系统。一旦数据生成,我们就可以同步训练模型。对于我们做的一些实验,长度尺度是几天或几周。那么问题就是:我们能否多路复用,能否在单位时间内获得更多数据?但无限的 token 生成器仍然存在。我们只需要解决另一边的技术问题,把所有这些拼图对齐并训练到模型中。
I mean, so that is an awesome question. It obviously varies by experiment. So, like you can't make the ribosome go faster, at least to my knowledge. Biology sets a limit for how fast you can go. In material sciences and chemistry, there are smaller time scales, there are bigger length scales. What you're actually kind of asking is a technical question though. So, how do you train a model against feedback mechanisms that vary by orders of magnitude in terms of feedback? So, we think about all of Lyra as being able to generate different kinds of data on different length scales. We can then synchronize how we train the model once that data has been generated. Again, for some of the experiments we do, the length scales are on the order of days or weeks. And then the question is can we multiplex, can we get more data per unit of time? But, the infinite token generator is still there. We just have to solve the technical problem on the other side of that to be able to line all these pieces up and train it into the model.
那么,当你说无限的 token 生成器仍然存在时,你是什么意思?因为你可以想象很多不同的科学 token,有些 token 提供的信息比其他多得多。某些东西你可以大规模收集,比如喜欢 NGS 的人,基本上可以收集无限量的 NGS 数据。
So, when you say the infinite token generator is still there, what do you mean by that? Because there are many different scientific tokens you can imagine, and some tokens provide much more information than others. And certain things you can collect maybe at scale, like people who love NGS, you can basically collect an infinite amount of NGS data.
是的。
Yeah.
然而,在某些情况下,另一个人类基因组可能只是一个增量更新,而不是……
And yet, there are certain cases where, you know, another human genome is probably going to be an incremental update versus...
是的,比如我的基因组相对于参考基因组,只有几千字节的信息量。那里没有太多信息。所以你说得完全正确。我们不想一遍又一遍地生成同一种数据。因此,我们正在构建的平台与传统自动化框架有本质区别。实际上,我们正在构建的实验平台优先考虑通用性和灵活性,而不是原始吞吐量。我们希望模型能够设计新的实验协议、运行协议并接收反馈,即使这个实验我们自己都没想过要做。所以下一个增量 token 必须是对模型有价值的东西,而不是又一个 NGS 样本——那种已经收益递减的东西。
Yeah, like my genome relative to a reference genome is like a couple kilobytes worth of information. There's not a lot of information there. So, you're exactly right. So, we don't want to generate the same kind of data over and over again. And so, the platform that we're building is qualitatively different than traditional automation framework. So, actually the experimental platform that we're building prioritizes generalizability and flexibility over raw throughput. We want the model to be able to design a new experimental protocol, run the protocol, and receive the feedback even if that's not an experiment we have thought about doing ourselves. So it's the next incremental token has to be something that is valuable to the model versus yet another NGS sample to teach it something where it's already hit diminishing returns.
所以当你说 Lyra 的下一个实验时,我想到的是传统上你会进入实验室,以某种方式重新配置实验室,然后手动运行一些实验,可能需要几周时间。在 Lyra,实验室是如何为新的实验重新配置的?
So when you say next experiment at Lyra what I think of is traditionally you would go into the lab and reconfigure the lab in whatever way and then run some experiments by hand maybe over the course of weeks or whatever. How does the lab get reconfigured for the new experiment at Lyra?
可以把实验室想象成一个图。每个仪器都是图中的一个节点,节点之间的边表示这两个仪器之间有物理传输层。
The way to think about the lab is it's almost like a graph. And so each instrument is a node in this graph and an edge between the node indicates that there's a physical transport layer between those two instruments.
嗯。
Yeah.
我们稍后可能会展示一个视频,但我们有一个物理传输层,几乎连接了我们在 Lyra 购买的所有仪器。目前这些是平面电机系统,有一个 96 孔板在轨道上磁悬浮。你可以对板的位置进行毫米级控制,所以你可以——我几乎把它想象成 PCI 总线,每个仪器——我不确定这个类比是否合适——
And so we'll probably have a video that we'll show in a bit but we have a physical transport layer that connects almost every instrument that we have bought at Lyra to each other. These are currently planar motor systems where there's a 96-well plate that magnetically levitates over a track. You have sort of millimeter control over where that plate goes and so you can — I think of it almost as like a PCI bus where each instrument — I'm not sure on analogies for like —
这个类比不错。
It's a good one.
我觉得一半的观众可能不知道 PCI 总线是什么,所以——
I think half of the audience might not know what a PCI bus is so —
是的。它是主板上的通用串行总线,允许你连接新设备。所以如果你插上新的显卡或新的硬盘,有一条总线让那个设备与计算机的其他部分通信。
Yeah. So it's a universal serial bus that on your motherboard allows you to connect a new device. So if you plug a new graphics card, if you plug a new hard drive in, there's a bus that allows that device to speak to the rest of your computer.
这适用于生物系统、材料系统等等。
And this works for like bio systems and material systems and etc.
越来越是这样,但还不完全。所以关于自动化,另一件要记住的事是,有一个很长的尾巴
Increasingly but not totally yet. So the other thing to keep in mind about automation is there's a very long tail
是的。
Yes.
你必须解决很多问题才能实现自动化。到目前为止,人们还没有以这种灵活的方式思考端到端自动化。所以,有些仪器现在还没有连接进来。例如,材料科学中没有很多高通量自动化,我们为此构建了定制仪器,然后将其接入。但这里有一个 80/20 法则:容易接入和自动化的东西直接插入 PCI 总线;而那些不容易的,仍然由人来操作。人们仍然会把样品移动到那里。或者,拧开试管盖是一件很难自动化的事情。很多实验室工作都假设你有对生拇指并且很擅长使用它们。我们看到一些关于 Lyra 的讨论把它描述成一家自动化公司,但这其实是错误的视角。我们不是自动化最大化主义者。我们实际上是 token 生成最大化主义者和灵活性最大化主义者。所以,随着时间的推移,我们会自动化那些有意义的环节,同时在当前使用有意义的解决方案。系统设计实验,给出指令。比如,哦,人们需要实际做这件事。那么你就招募一些员工去做。一切都是一次 API 调用。
of things that you have to solve to be able to automate. And to date, people have not been thinking about end-to-end automation in this flexible kind of way. And so, there are instruments that are not connected to this now. There's not a lot of high throughput automation in material sciences, for example, and we've been building custom instruments for that that then are brought on board. But there's like an 80/20 rule I play here where things that are easy to onboard and automate are plugged directly into the PCI bus. And then things that are not, people still move. People will still move a sample to that. Or turns out that removing a cap from a test tube is a very hard thing to automate. It's like a lot of the lab assumes that you have opposable thumbs and you're good with them. And some of the things that we've seen discussed about Lyra frames this as an automation company and that's like kind of the wrong perspective to think about what we're doing. We're not automation maximalists. We are actually sort of like token generation maximalists and flexibility maximalists. So, we will over time automate things that make sense to automate and then again, use solutions now where they make sense. So, the system designs the experiments. It gives instructions. There's like, oh, people need to actually do this thing. So, you recruit some of the staff to go and do that thing. Everything's an API call.
嗯。
Yeah.
所以,有时候你调用 API,背后是一个机械臂;有时候,是一个人的手臂在操作。
And so, sometimes when you call an API, there's a robot arm. Sometimes there's a human arm that does something.
就在 API 层之下。
Literally below the API line.
嗯,有意思。我觉得我们想把资源花在刀刃上,做出理性的决策。有时候,一个人十分之一秒就能完成的事,硬要去自动化反而没意义。关键在于模型要有能力给出指令来检验假设,并且所有数据都可见、透明、可存储,这样那些 token 才能流回模型。
Yeah, well, funny thing. I think that again, we want to spend resources where it makes sense and make rational decisions. And sometimes it just doesn't make sense to try and automate a step when a person can do it in a tenth of a second. But what matters is that the model has the ability to give instructions to test hypotheses and that all of that data is visible, transparent, stored so that those tokens flow back into the model.
你的 AI 模型会做完整的实验设计吗——不只是调整比例或来源,而是超出已有方案的那种?
Do you have your AI models doing entire experimental designs which are beyond just a pre-existing protocol where you tweak relative ratios or sources, or like what goes into a pipe or pipette?
这取决于你对“新颖”的阈值。在表达方案和一些基因编辑工作中,我们测试了平台与人类的能力对比。模型零样本能达到 80%,人类零样本是 0%。我们现在做完全开放式的自由实验了吗?还没有。那是目标,我们正在朝那个方向努力,那是我们想要的终局。但我们确实看到了在极短时间内完成大量人类智力劳动的能力。
I mean, it depends on your threshold for novelty. Certainly for expression protocols, for some gene editing work we've done, we have tested the platform's ability to do that versus humans. The model gets about 80% of that zero shot. Humans get 0% of that zero shot. Are we doing fully open-ended free-form experimentation now? No, not yet. That is the goal. But we're building towards that. That is the end state we want to be in. But we have seen the ability to do what would be an enormous amount of human intellectual labor over a very short time horizon.
那么,当你让 AI 模型自由设计新实验时,如何确保这些是值得测量的东西?如何验证这是一个好策略,而不是白白浪费钱?
So when you're giving your AI models free reign to start designing new experiments, how do you make sure that these are things that should be measured, or that you validate that this is a good strategy, or that you didn't just waste a bunch of money?
首先,这背后可能有一个安全问题。我们从一开始就非常重视,对吧?既包括安全也包括安保:数据安全和模型建议的安全性。我们有一个非常强大的团队,在强有力的领导下不断壮大。这是第一层。我们有强大的 AI 安全协议,类似于人们在大型语言模型中研究的那些提升考虑,只不过这次是动真格的。
The first one is there is maybe an underlying safety question there. And I think that we've been taking it very seriously from the beginning, right? Both security and safety: security of the data and the safety of the model suggestions. We have a very strong team. It's growing under very strong leadership. That's the first layer. We have strong AI safety protocols that look similar to the sort of uplift considerations that people have been looking into for large language models. Only it's absolutely for real.
在实验室自动化场景中,处理生物物理材料科学这类问题时,实际需要担心哪些危险?我通常想到的是恶意行为者,或者系统足够复杂以至于可能真的输出危险结果。根据我对 Lila 的理解——我们还没聊到,可能一会儿会提到——安全似乎还不是当前的主要担忧。
In a lab automation setting where you're working on a biophysical material science type problem, what are actually the dangers you have to worry about? I generally think of malicious actors, or situations where you have a sufficiently complicated system that it could genuinely output something dangerous. It seems like from the scope of Lila as I understand it, which we haven't talked about yet, maybe it'll come in a minute, but it doesn't seem like safety is actually going to be a major concern at this point.
这是我们从一开始就必须认真对待的事,我们承担不起出错的代价。我同意你的看法。目前,操作权在 Lila 员工手中,他们的利益与我们的使命一致,对平台的理解也一致。所以,我们确实不用担心恶意行为者。但我们仍需要在一定程度上担心模型给出的建议。我认为这更多是触及实验室安全的问题,而非恶意。我不认为会出现模型建议使用剧毒化学品的涌现行为,更多是让仪器溢出、混合了不该混合的化学品。所以,从一开始就需要有化学 EHS 安全层,因为我们做的是开放式工作。
I mean, it's something we need to take seriously from the beginning. It's something where we cannot afford to not get it right. I agree with you. Right now, it's in the hands of Lila employees whose interests are aligned and whose understanding of the platform is aligned with our mission. So, I agree we don't have to worry about malicious actors. We still need to worry to some degree about the model giving a suggestion. I think it's more about things that start touching into lab safety than malicious. I don't think we're going to have emergent behavior where the model suggests an extremely toxic chemical. It's more about pushing an instrument such that it overflows, or combines chemicals it shouldn't have. So, I think there's a chemical EHS safety layer that needs to be there from the beginning because we're doing open-ended work.
我确实认为 Rafa 说得对,安全是不能拖延的,因为能力曲线往往是 S 形的,可能看起来一切正常,然后突然出现你没想到模型能做到的事。所以我们在这方面非常主动。我们有 AI 安全团队,就像 Rafa 说的。但我也认为他们说得对,我们可以用有意义的方式约束问题,这是面向公众的通用 AI 系统做不到的。我们还可以依靠生物安全等级之类的东西,以及传统的实验室安全来帮忙。
I mean, I do think Rafa is right in that safety is not something you can procrastinate on because capability curves tend to be sigmoid shaped and it can look like everything's fine, and then all of a sudden there's something that you didn't anticipate the model being able to do. So, we are definitely proactive on that side. We have an AI safety team like Rafa said, but I think they're also right in that we can constrain the problem in meaningful ways in the way that a broad-based AI system that interacts with the general public cannot. We can also lean on biosafety levels and things like that, good old-fashioned lab safety to help in the meantime.
当然,模型并非暴露在所有信息中。我们不一定需要把所有实验能力暴露给所有科学问题。对于抗体设计问题,我们可能甚至不需要让模型知道我们有装气体的气罐,因为它用不到。所以我们仍然可以在特定科学领域的问题范围内保持创造力。
And of course, the models are not exposed to everything. We don't necessarily need to expose all the experimental capabilities to all the scientific questions. For an antibody design question, we probably don't even need to expose the model to the fact that we have gas canisters that contain gases, because it's not going to need them. So we can still be creative within questions that relate to one particular area of science.
嗯。
Yeah.
不过你的问题很有意思:如何判断某件事是否危险?这其实很难。或者说,如何判断它是否浪费?我们在电催化剂方面做了一些工作,Lila 内部有人在这个领域发表了 40 篇论文。模型最初的一些建议很无聊,然后从无聊变成了他认为的愚蠢。这些是非铂族电催化剂,用于从水中分离氢和氧来制氢,结果它们成了我们制造过的最好的非铂族电催化剂。所以,明显错误和令人惊讶的新颖之间的界限,即使对人类专家来说也很难判断。因此我们会做一些浪费的事,因为我们想知道两者之间的区别。
Your question though is interesting: how do you know if something is dangerous? That's actually kind of hard to do. Or actually, how do you know if it's wasteful? Some of the work we've been doing in electrocatalysts, we have someone inside of Lila who's published 40 papers on the topic. Some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turn out to be our best non-platinum group electrocatalysts that we've made. So the line between obviously wrong and surprisingly novel, even to a human expert, is hard to know. And so we will do wasteful things because we kind of want to know the difference between the two.
这就引出了一个问题:像你刚才描述的实验,结果是否有效是显而易见的,对吧?但你可以想象,之前伯克利实验室的一些工作曾因测量被误解而引发争议。你怎么知道你的有效性测量或你正在优化的指标实际上是正确的?
That brings up the question: for an experiment like what you're describing now, it is obvious whether it works or not, right? But you can imagine, and there was some controversy in previous work at Berkeley Lab around measurements that were misinterpreted. How do you know that your measurements of effectiveness or whatever you're optimizing are actually correct?
是的,我对这方面非常熟悉。我想说,我们不能因为它是 AI 就放松科学严谨性的标准。不像 5 年前,我们刚开始做 X 和 Y 的遗传模型时,大家会说:“嗯,挺可爱的,有点用。”就像对待孩子一样。但现在我们已经过了那个阶段,我们需要让 AI 科学达到与人类主导的科学相同的标准。
Yeah, I'm very familiar with that part of the landscape. I would say we cannot relax our standards of scientific rigor because it's AI. It's not like maybe 5 years ago when we started doing genetic models for X and Y, they were like, 'Yeah, it's cute. It kind of works.' Like you would do with a kid. But now we're past that, and we need to hold AI science to the same standard we hold regular human-led science.
我认为 2023 年的那篇论文是社区的一个转折点。AI 社区中,AI for science 的人一直对渐进式进步感到兴奋。那时,我们开始触及更广泛社区的意识,他们说:‘太棒了,但现在我们要以最高标准来做事。’所以,我们有很多实验科学家。我想回到 API 这个点。我们有幸从零开始建立一家公司,员工对 AI 有认知且充满热情。通过我合作过的人和网络,我们组建了一支由实验科学家和自动化工程师组成的团队,他们真正相信这个使命并希望实现它。他们很包容。每当 AI 给出非常错误的结果,他们就会按下红色按钮,说‘这是个坏主意。’但他们也很包容地尝试假阳性。假阳性对人类科学家来说很糟糕,因为你去尝试某个东西,它不工作,但模型很棒——它大大降低了不确定性。对操作员来说,这有点扫兴,因为你以为会得到很酷的东西。我们做到了,回到 Andy 提到的点,直到 3 个月前,人们还在审批 AI 的决定。大约 3 个月前,我们开始看到模型的疯狂想法开始让人惊讶,但出奇地好。就像‘我不知道,我想我们需要试试。’我们看到人们实验性地通过包容来挑战 AI 的能力。人与计算机的界面在过去几个月里非常有益。
I think that 2023 paper was a switch over from the community. Part of the AI community, AI for science people, were always excited to see incremental progress. At that point, we started collectively touching upon the rest of the community's awareness, and they were like, 'Fantastic, but now we're going to talk about the way we do things to our highest standard.' So, I think we have lots of experimentalists. I want to go back to the API point. We've had the fortune by starting from zero to build a company where people are AI aware and AI excited. Across all the people and networks I collaborated with, we've managed to build a team of experimentalists and automation engineers who really believe in the mission and want to make it happen. They're taking this graciously. Whenever AI gives something very wrong, they push the red button and say, 'This is a bad idea.' But they're also gracious in trying false positives. False positives are terrible for human scientists because you go to try something and it doesn't work, but the model is fantastic—it reduces uncertainty a lot. For the operator, it's a bummer because you thought you were going to get something cool. We've managed, going back to the point Andy made, until 3 months ago people would be approving AI decisions. About 3 months ago, we started seeing that the model's crazy ideas started being surprising to people, but surprisingly good. It's like, 'I don't know, I guess we need to try.' We see the switch over to the ability of people experimentally to challenge the AI by being gracious. That interface of human and computers has been very rewarding over the last few months.
另外,让模型控制实验室迫使你构建基础设施,以暴露你通常想要或关心的数据。也许没有实验是错的,但你希望有能力解释结果。如果你有一个实验,你拟合一个统计模型;你试图做的是用输入的变化来解释输出的变化。我们有能力解释输出的变化,因为我们测量了很多不同的东西,因为我们必须向模型暴露这些数据。所以我们可以说:‘好的,那天实验室湿度不对。也许这正好解释了问题。’然后我们可以按一个按钮,重新运行实验来验证。我们不相信,就像 Rafa 说的,在某种意义上我们必须对任何结果更加怀疑,但我们可以快速重新运行那个实验,因为实际上一切都是软件。
Also, giving the model control of the lab forces you to build infrastructure to expose pieces of data that you would normally want to or care about. Maybe no experiment is wrong, but you want the ability to explain the outcome. If you have an experiment, you fit a statistical model; what you're trying to do is use variation in inputs to explain variation in outputs. We have the ability to explain variation in outputs because we measure so many different things, since we have to expose that to the model. So we can say, 'Okay, the humidity was off in the lab that day. Maybe that explains exactly the thing.' Then we can push a button and rerun the experiment to verify. We do not believe, like Rafa said, in some sense we have to be more skeptical of any outcome, but we can quickly go and rerun that experiment because it's all software effectively.
你觉得团队花了很多时间在验证上吗?或者说,时间分配是怎样的?
And you find that the team spends a lot of time on verification, or how's the breakdown?
越来越少。一开始,验证并不是很多。有一个执行过程:我们有了假设,有了一套指令,这些指令会发送到 API。我们为其担保,然后 API 的一部分是人们在做事。我认为这一点会保留——这是劳动密集型部分。但检查直觉是否正确,我认为我们开始看到在超级智能的局部峰值中,我们更多地是在支持这种超级智能行为的涌现,而不是在把关,确保想法不是浪费。所以我认为这在各个领域都在发生。详细说一下 Andy 提到的例子,我们非常关心能源和可持续性。我们不仅仅是生物技术公司;我们真正关心能源、可持续性和材料。我们正在尝试制造绿色氢气。要制造绿色氢气,你需要用光来分解水分子。你需要支付一部分能量,因为那是储存在化学键中的能量,你从阳光和电力中获得。然后还有一些开销,称为过电位,这与世界不完美、事物有损耗有关。损耗来自一种叫做催化剂的东西。今天的催化剂还可以,但昂贵且稀有——它们由钌和铱制成。所以我们建立了一个模型来探索如何不使用这两种元素。人们发表论文,几周前有篇文章说‘钌,低钌合金用于 XYCs。’当然,你可以掺杂,稀释,但仍然是同样的根本问题——你只是少用了 50%。所以我们让模型自由探索这个问题,我们有能力制造材料、测量性质、测量催化作用、测量稳定性。在第二代或第三代的序列学习中,这种信息和模型所知之间的相互作用,我们开始看到一些建议——措辞没问题,它使用了我们使用的概念——它说:‘我不会把这个想法应用到那个元素上。我不会那样把它们组合在一起。’结果证明,这些是我们迄今为止性能最好的化学品。
Less and less. At the beginning, it wasn't so much. There's an execution: okay, we got a hypothesis, we got a set of instructions that's going to go off to the API. We vouch for it, and then from part of the API are people doing things. I think that will stay—that's the labor of it. But the double-checking that the intuitions were right, I think we're starting to see in this superintelligence local spikes of places where we are supporting this emergence of superintelligent behavior more than we are gatekeeping and that the ideas are not just wasteful. So, I think that's happening for domains. To elaborate on the example Andy mentioned, we care a lot about energy and sustainability. We're not just a biotech; we really care about energy, sustainability, and materials. We're trying to make green hydrogen. To make green hydrogen, you need to use light to split the water molecule. A bunch of that energy you need to pay for because it's the energy stored in the chemical bond, and you get that from sunlight and electricity. Then there is some overhead called the overpotential, which has to do with the fact that the world is not perfect and things are lossy. The loss comes from something called the catalyst. Today, the catalysts out there are okay, but they're expensive and rare—they're made out of ruthenium and iridium. So we set up a model to explore what we can do to not use these two elements. People do papers, and there was something a couple of weeks ago that said, 'Ruthenium, low ruthenium alloys for XYCs.' Sure, you can dope it down, water it down, but still the same fundamental problem—you're just using 50% less. So we set the model loose on this problem, and we have the ability to make the material, measure the properties, measure the catalysis, measure the stability. On the second or third generation of sequential learning, this interplay between information and what the model knows, we started seeing suggestions that were like—the words were fine, it was using the concepts we use—that said, 'I wouldn't apply that idea to that element. I wouldn't have put them together in that way.' And it turns out those have been our best performing chemicals so far.
我确实想谈谈为什么我们不是生物技术公司,但在此之前,沿着这个思路问最后一个问题。强化学习以奖励破解而闻名。你刚才——我忘了你说什么了。我不知道你是否说过你在用强化学习,或者学习迭代。我非常担心,你投入一些奖励,你真的可以破解物理科学,这是用算力做不到的。
I do want to get to why we're not a biotech, but before we do, one last question along this train of thought. RL is famous for reward hacking. You just—I forget what you said. I don't know if you said you were using RL, or learning iterations. I'd be very concerned that you throw some rewards and you can really hack the physical sciences in a way that you can't do with compute.
是的,我不反对这一点。
Yeah, I'm not going to disagree with that.
100% 同意——
100% agree with—
你见过的最有趣的奖励破解例子是什么?
What's the funniest example of reward hacking you've seen?
我们有很多有趣的强化学习失败案例,不完全是奖励破解。一个是当我们训练早期的一个任务时,比如,你能做一个板图吗?你能把实验条件布置在板上吗?当人问它时,它会生气——你让模型做一个板图,它做了,然后你说‘实际上,你能换这些试剂吗?’它就会骂人。它会说:‘这是一个 96 孔板。拜托,老兄。没那么难。’在思维链中,我不知道那是从哪里来的,但它会……
We have lots of funny RL fails that are not explicitly reward hacking. Well, one is when we trained one of the early things we did was like, can you just make a plate map? Like, can you lay the experimental conditions out on a plate? It got annoyed when the person would ask—you would ask the model to do a plate map and it would do it, and then it would be like, 'Actually, could you change these reagents?' and it would swear. It would be like, 'It's a 96-well plate. Come on, man. It's not that hard.' In the chain of thought, I don't know where that came from, but it would...
对,互联网的某个角落里有这些。我不在乎那个。所以我们看到很多这种有趣的人格特质,都是强化学习的结果。有一些明显的强化学习病态现象,比如重复。思维链会崩溃,然后一遍又一遍地重复它的最终答案。出于某种原因,这有时确实能可靠地带来更高的奖励。我们不太确定为什么病态或不可读的思维链在某些情况下会导致更高的奖励。
Yeah, somewhere in the internet's in there. I didn't care about that. So we've seen lots of funny personality quirks like that as a function of RL. There's obvious RL pathologies like repetition. The chain of thought will collapse and just repeat its final answer over and over again. For some reason, that reliably sometimes leads to higher rewards. We're not sure exactly why pathological or non-legible chain of thoughts in some cases lead to higher rewards.
抱歉,我想打断一下。我们在讨论强化学习。你描述的方式听起来就像大家都在做的思维链上的强化学习。但你们的强化学习实际上有一个实验步骤。
Sorry, I want to interrupt. So we're talking about RL. The way you're talking about it sounds like just RL on chain of thought, like everybody's doing. But your RL actually has a lab step.
对。
Yeah.
如果陷入病态循环,那是不是意味着实验室就在一遍又一遍地做同一个实验?
If you're in a pathological loop, does that mean the lab is just doing the same experiment over and over?
不是这样的。思维链,退一步说,是模型用来解决问题的词元。如果你在解一道数学题,你会做定理一、定理二、推论、引理——你把问题分解。在科学中,思维链也有类似的部分。有推理过程:我正在为这个靶点制造一种抗体。我对这个靶点了解多少?已知的表位有哪些?我的攻击计划是什么?在思维链中,还有工具调用。我可能会使用结构预测模型来了解序列如何在三维空间中折叠。所以工具调用是思维链的一部分。在 Lyra,有趣的是,实验室仪器也是工具调用,或者如果是一个工作流,那就是一系列工具调用。
It's not. So a chain of thought, to step back, is tokens that the model uses to solve a problem. If you were solving a math problem, you would do theorem one, theorem two, corollary, lemma—you decompose the problem. In science, the chain of thought has some of that too. There's reasoning: I'm trying to make an antibody for this target. What do I know about this target? What are the known epitopes? What's my plan of attack? In the chain of thought, there are also tool calls. I might use a structure prediction model to get a read of how the sequence folds in three-dimensional space. So tool calls are part of the chain of thought. At Lyra, the fun thing is that the lab instruments are also tool calls, or a series of tool calls if a workflow.
这一切都是人类可读的,都是英文。
It's all human legible. It's all in English.
对。所以我们看到的一些病态现象是,它直接跳过了所有中间部分——我们认为对解决问题很重要的部分——直接给出答案。它说:“这种情况下我不需要做实验。我不需要调用工具。”在某些情况下,因为我们已经做过实验,所以我们可以判断,但出于某种原因,这实际上在某些情况下并不是一个坏策略。这里有些神秘之处。
Right. And so, some of the pathologies we've seen is it just skips all the middle part, which we would think is important for solving a problem, and goes right to the answer. It says, 'I don't need to do an experiment in this case. I don't need to call a tool.' In some cases where we can judge something because we've already done the experiment, for some reason, it is actually not a bad strategy in some cases. There's some mystery there.
它是个理论家。
It's a theorist.
对。它已经完成了计算。这可能有点跑题,但它实际上是在潜在空间中思考。它输出词元。所以思维链通常是模型实际计算过程的不可靠叙述者。我们正在思考的一个大问题是,当我们开始处理一个问题时——比如 Rafa 提到的电催化剂——我们实际上不知道什么是对什么是错。我们应该在多大程度上依赖思维链,而不是仅仅相信实验、相信验证器、相信模拟器作为最终的真相来源。
Yeah. It's done the calculation. And this is probably too much of a tangent, but it actually thinks in latent space. It emits tokens. So the chain of thought is often an unreliable narrator for what computation the model is actually doing. One of the big things we're trying to think about is when we're moving into working on a problem, like Rafa said for electrocatalyst, where we actually don't know what right and wrong looks like. How much should we rely on the chain of thought versus just trusting the experiment, trusting the verifier, trusting the simulator as the ultimate ground truth.
所以,Lyra 不是一家生物技术公司。我认为 Lyra 在这方面其实相当独特。
So, Lyra is not a biotech company. Lyra is actually fairly unique, I think, in this way.
我参与过生物技术公司,也帮助创办过生物技术公司。通常,目标是冲刺到临床试验。你想开发一个资产。你开发平台是为了在进入哪个领域时有选择权。但一旦你有了临床资产,你就把所有东西都置于医学诱导的昏迷状态,然后通过临床试验。如果顺利,其他事情才能进行。所以我们把那个选项拿掉了。在 Lyra,模型本身就是有价值的东西。从这个意义上说,我们更像是一个新实验室,试图思考一种新的方式来推进基于语言模型的核心推理能力。
I've been involved with biotechs. I've helped start biotechs. Often, the goal is to sprint to a clinical trial. You want to develop an asset. You develop the platform in service of having optionality of what space you move into. But once you have the clinical asset, you put everything into a medically induced coma, and you get through the clinical trial. If it goes well, then other things get to. So we are taking that option off the table. The model itself is the thing of value at Lyra. In that sense, we're much more like a neo lab, trying to think of a new way to push forward capabilities of a core reasoning LM-based model.
连实验室平台都不是?
Not even the lab platform?
嗯,实验室平台是词元生成器。那是数据生成机制,最终是 Lyra 的护城河。一旦它持续扩展,我们每单位时间、每单位平方英尺能生成的数据量就会增加,这反馈给模型使其更智能,然后模型建议下一个实验。所以我们真正专注于让这个核心模型尽可能高效和智能。我们可以讨论这如何适用于不同的商业策略,但最终,我们有兴趣创造这种新型的 AI 模型。
Well, the lab platform is the token generator. That is the data generation mechanism that ultimately is the moat for Lyra. Once that continues to scale, the amount of data we can generate per unit time, but per unit square foot, will go up, and that feeds back into the model to make it smarter, which then suggests the next experiment to do. So we really are focused on making this core model as performant and smart as possible. We can talk about how that lends itself to different commercial strategies, but ultimately, we're interested in creating this new type of AI model.
我想引用 Octant Bio 的 Sridhar Kota 几周前的一条精彩推文:“机器学习在药物发现中的商业模式是什么?因为如果你需要数据来训练模型,但如果你已经有了数据,你还需要模型做什么?”
I want to quote Sridhar Kota from Octant Bio who had a great tweet a few weeks ago: 'What is the business model in ML for drug discovery? Because if you need the data to train the model, but if you have the data, what do you need the model for?'
当范围很窄时,这是对的。在任何特定的科学垂直领域都是如此。我用的类比是,如果你回到 10 年前,试图创建一个编程助手模型,你只会得到编程数据。你不会同时得到莎士比亚诗歌、 carnitas 食谱。事实证明,当模型能够在更广泛、更深入的数据上训练时,会产生溢出效应。我们打的核心赌注是,这在科学领域也是成立的。如果模型在越来越广泛的数据集上训练,你在特定领域所需的数据量就会减少。在某些情况下,如果该领域与模型已经见过的领域相邻,数据需求会降到零。所以有一个数据效率的论点,表明拥有一个能够生成广泛科学数据的通用平台是有益的。我还要提到,我们也在使用已经是商品的东西:公共数据集、模拟器,而实验平台是对这些现有商品资源的补充。
That is true when you are narrowly scoped. That is true within any given vertical of science. The analogy I would use is if you went back 10 years and tried to create a coding assistant model, you would just get coding data. You wouldn't also get Shakespeare poetry, carnitas recipes. It turns out that there is spillover as the model is able to train on a broader swath of data and a deeper cut of data. The core bet we're making is that is true for science. If the model is trained on an increasingly broad set of data, the amount of data you need in a given domain is reduced. In some cases, it will be reduced to zero if it's adjacent to what the model has already seen. So there's a data efficiency argument that would suggest having a general platform that can create a broad swath of scientific data. I'll also mention that we are using things that are already commodities: public datasets, simulators, and the experimental platform is a complement to these existing commodity resources.
这引出了关于适用域概念的问题。你有不同的尺度——量子、化学、不同的生物领域——它们产生完全不同类型的信息和实体之间的关系。我对跨领域方法的一个担忧是,这些领域之间到底有没有迁移?而 carnitas 食谱和象棋问题有一个共同点,即它们都是用语言写的,但在这里,这些领域之间几乎有完全独立的模型。
This brings up a question about the concept of applicability domain. You have different scales—quantum, chemical, different bio realms—that result in completely different types of information and relationships between entities. One concern I would have with a cross-cutting approach is, is there domain transfer between these at all? Whereas carnitas recipes and chess problems have the commonality that they're written in language, you almost have a completely separate model between these domains.
人类科学家在所有那些领域工作。
Human scientists work on all of those domains.
对。而且他们大多用书面语言互相交流……
Correct. And they mostly communicate with each other in written language using...
好的。
Okay.
使用工具。我认为存在一个共同的推理过程,让人能够解决每个领域中的问题。所以我认为这种逻辑也适用于我们正在训练的推理模型,它同样使用工具,能做数学、能写代码,但所有知识都存储在一个地方。
Using tools. I would say there's a common reasoning process that allows someone to solve problems in each one of those domains. And so I think that logic carries over to a reasoning model that we're training that again uses tools, can do math, can do code, but it's having all of that knowledge stored in one place.
对我来说,领域迁移的一个经典例子是复杂性理论和量子引力之间的关系,对吧?现在很多量子引力理论基本上都认识到两者背后有相同的数学。你有没有这种“哦,这个领域居然适用于那个领域”的例子?
One classic example for me of domain transfer is between complexity theory and quantum gravity, right? Where now a lot of the quantum gravity theories are basically recognizing the identical math behind the two. Do you have examples of this kind of 'oh man, this domain actually applies to this domain'?
是的。所以我们整理了这个包含 10 万亿科学 token 的推理数据集,这些推理轨迹在生命科学、化学和材料科学中经过了实验验证。我们发现这个通用模型往往能击败领域专用模型。所以很难指出模型里到底是什么让它做到这一点,它发现了什么联系。但显然,在科学领域看到更多数据,以样本对样本的方式,击败了领域专用的推理模型。
Yep. So, we have assembled this reasoning dataset of 10 trillion scientific tokens, reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences. And we have seen that this general model often beats the domain-specific models. And so, it's hard to point to what's in the model that is making it that way, what connections it has realized. But clearly having seen more data across all science beats, in a sample-for-sample kind of way, domain-specific reasoning models.
科学的未来是语言,对吧?
And the future of science is language, right?
嗯,是的,所以我……
Well, so yeah, so I...
是的,化学的未来是语言。
Yeah, future of chemistry is language.
是的,是的。
Yeah, yeah.
也许吧。
Maybe.
我不认为这是必要的。我不认为这是科学超级智能的必要条件。我的意思是,上周 Demis Hassabis 有句话,可能不值得提炼所有那些存在于……你知道,有些数据模态与语言截然不同。他们总是告诉我,‘好吧,Rafa,英语是图灵完备的。所以你可以用英语表达一切。’
I don't think it's necessary. I don't think that's a necessary condition for a scientific superintelligence. I mean, there was this quote from Demis Hassabis last week that it might not be worth distilling all the ways that live in sort of... you know, there are data modalities that are so different from language. And they always tell me, 'Well, Rafa, English is Turing complete. So you could express everything in English.'
图灵完备的语言,但是的,是的,是的。
Turing complete languages, but yes, yes, yes.
所以,我同意这一点。可能有些地方因为问题的本质而更高效……我的意思是,你做过几何深度学习,对吧?我认为对于几何,也许……你知道,我跟我同事 Tess Mead 聊过。我认为几何是那种人们觉得可能有些……问题的本质更适合其他架构的领域。所以,如果我们需要调用蛋白质折叠模型,或者需要调用等效的扩散模型来生成晶体结构,那完全没问题。所以,我认为未来科学与我们交流的方式肯定是通过语言。模型需要一直用英语思考化学。也许是的,也许不是。不要用英语思考化学。
So, and I agree with that. There might be places where it's more efficient because of the nature of the... I mean, you've done geometric deep learning, right? I think for geometry and maybe... you know, I call my colleague Tess Mead. I think geometry is one of those places where people feel that there might be some... just the nature of the problem is more amenable to other architectures. So, if we need to call a protein folding model or we need to call an equivalent diffusion model to make crystal structures, that's fair game. So, I would say the future of the way science talks with us for sure is through language. That the model needs to be thinking in English about chemistry all the time. Maybe yes, maybe not. Don't think about chemistry in English.
他们用英语讨论它。
They talk about it in English.
完全正确。我同意 Rafa 说的一切。基于 token 的推理结合工具使用非常强大,我认为我们提出的主张是,在科学领域我们才刚刚触及皮毛。我们并不是要把领域专用模型蒸馏成一个推理模型。它可以高效地使用这些工具。所以,推理(通常用英语,但也用 Python 等)与工具使用的结合非常强大,我们在科学领域还处于非常早期的阶段,去理解我们能把它推进到什么程度。
Exactly. And I agree with everything Rafa said. Like token-based reasoning with tool use is very powerful and I think the claim that we're making is we have barely scratched the surface for that in science. We're not trying to distill domain-specific models into a reasoning model. It can use those tools productively. And so it's just the combination of reasoning often in English but also in Python and things like that combined with tool use is very powerful and we're very early in science in understanding how far we can push that forward.
我明白了。那么你能举一些你们正在进行的、有代表性的项目例子吗?
I see. So can you give some examples of campaigns that you are running that are representative?
实际上,在你回答之前,我们先退一步。我意识到我们还没解释清楚,你不只做生物。不只是任何技术生物。所以我认为这正好引出这个话题。那么,不只是技术生物,你在科学方面还做什么?
Actually, before you do that, let's take back. I realized we still haven't explained that you don't just do bio. Not just any tech bio. So I think this is a great lead into this. So not just tech bio, what do you do in terms of science?
科学?不。
Science? No.
嗯,所以也许是的。我们训练模型的方式涵盖生命科学:DNA、RNA、蛋白质、细胞、小分子、不同种类的化学物质和不同类型的材料。这就是我们现在的范围,诚然……
Um, so maybe yeah. So the way that we train the model is across life sciences: DNA, RNA, proteins, cells, small molecules, different kinds of chemistries and different types of materials. So that is where we are scoped now, which is admittedly...
但材料本身也不仅仅是那样,它的范围也和生物那边一样广泛。
But materials itself is also not just that, it's also as largely scoped as everything that's in the bio side.
我们来举一些例子。今天我们可以制造薄膜、粉末、量子点。我们有一个可爱的量子点。你们熟悉吗?量子点是某些电视中的发光技术,你需要控制它们,使它们具有完全相同的纳米尺寸,而制造的纳米尺寸决定了它们会是什么颜色。你需要最纯的红、最纯的蓝和最纯的绿,才能为你的电视打造出非常锐利和丰富的色彩,而且你需要让它们高度均匀。它们都必须相同,否则颜色就会混合。所以我们有一个可爱的演示:在我们的自动驾驶实验室里,我们让来访者挑选一个波长,也就是他们希望量子点呈现的颜色,然后我们启动机器。模型进行推理,有时我们甚至会加入模型从未见过的新化学物质,看看它如何应对。机器运行着,在大约一个到一个半小时的参观结束时,机器已经制造出了一代或更多代量子点,这些量子点往往能达到——否则我们不会这么做,对吧?——达到人们建议的颜色。所以,我们有能力制造大量材料。我们可以配制液体、聚合物和软物质。我们非常关心能源和可持续性。因此,我们在电化学方面有很强的能力,涉及化学转化与作为可再生能源的电力的相互作用。我们关心传统催化,也关心材料的机械性能。所有这些都汇聚在项目中,比如我们制造催化剂、用于腐蚀或航空航天高性能机械应用的高性能涂层。在过去的几周里,我们与一个外部合作伙伴一起,启动了多个我们以前没做过的冲刺项目,涉及从粘合剂到冷却液等各种领域。所以,我们能够越来越多地启动令人兴奋的发现,在开放式的化学和材料科学空间中。
We're going to give some examples. So today we can make thin films, we can make powders, we can make quantum dots. We have a cute quantum dot. Are you folks familiar? Quantum dots are the luminescent technology in some TVs and you need to control to make them of exactly the same nanometer size, and the nanometer size you make them controls what color they're going to be. And you need the purest red and the purest blue and the purest green to make really sharp and rich color palette for your TV, and you need to make them as homogeneous. They all need to be the same. Otherwise the color gets blended. So we have a cute demo where our self-driving lab, we ask our visitors to pick a wavelength, what color you want your quantum dot to be when they come into the office, and then we fire off the machine. The model reasons, even sometimes we throw in new chemicals that the model had never seen just to see how it moves. The machine is running and by the end of this sort of hour and hour and a half tour, the machine has made maybe one, maybe more generations of quantum dots that tend to hit—otherwise we wouldn't do it, right? Then to hit the color that people suggested. So, we have the ability to make lots of materials. We can formulate liquids and polymers and soft matter. We care about energy and sustainability a lot. So, we have a good chunk of electrochemistry capabilities about the interplay of chemical transformations and electricity as a renewable energy source. We care about traditional catalysis and we care about mechanical properties of materials. And all this comes together in programs where, you know, we make catalysts, we make high-performance coatings for corrosion or aerospace high-performance mechanical applications. And over the last few weeks with an external partner, we started multiple sprints of things we weren't doing before that touch, you know, from adhesives to cooling fluids. So, we've been able to sort of more and more spin up just exciting discoveries in open-ended chemistry and material science spaces.
你们有没有在量子点和蛋白质设计之间建立联系?
Do you have any connection between quantum dots and, let's say, protein design?
做这件事的是同一个平台。是同一套能力。所以,有一个共享的基础设施让我们能在同一个屋檐下做所有这些事情。如果没有这种连接组织,那么我们的能力……我们就无法做所有这些事情。我们有没有做过类似机制可解释性的事情,比如查看模型内部,看看来自电催化剂的见解是否启发了……我们还没有深入进行机制可解释性的研究。我们看到,随着平台越来越成熟,我们执行这些项目的能力也变得越来越快。
It's the same platform that does that. It's the same set of capabilities. And so, there's a shared infrastructure that lets us do all of those things under the same roof. If there were no connective tissue, then our ability to do... we just would not have the ability to do all those things. Have we done like the mechanistic interpretability thing where we look inside the model and see does this insight from electrocatalyst inform... We haven't done a deep dive on the mechanistic interpretability thing. We have seen that our ability to do these programs has gotten faster as the platform has become more mature.
那么,LMP 在你的工具包中是不是很常见?因为它,你刚才提到的这些想法有很大一部分得以实现?这在生物方面当然说得通,但我对生物的了解远多于材料。这是你们实验室工具包中的一个常见主题吗?
So, is LMP just like a common thing in your toolkit? Because of this, it enables a large fraction of these ideas you've mentioned? It certainly makes sense on the bio side, but I know bio much more than materials. Is that a common theme in your lab toolkit?
AI 科学工厂的能力越强,我们就能越快地去追求新的目标产品概况和激动人心的新机遇。因为模型能做的事情更多了,实验室能做的也更多了,我们的科学家在整合新能力时也更灵活、更快速。我们拥有的仪器越多,添加新仪器的速度就越快。这与我们公司的类型相符。这里有些超大规模扩展的影子:软件的 Scaling 由硬件的 Scaling 支撑。数万平方英尺的实验室空间上线,配备数十到数百台仪器,这给了我们快速行动的广度。例如,我的同事 Heather Kulik 最近来过这个播客。我们正在看到颠覆的一个领域是由分子与金属相互作用制成的材料。我们的模型是在小分子药物发现上训练的,它们学到的所有化学知识都迁移到了推理金属有机框架材料上,这些材料可以用来从空气中捕获二氧化碳或过滤氨气。
The AI science factory: the more capabilities it has, the faster we've been able to go after new target product profiles and new exciting opportunities. Because the model is prepared to do more things, the lab can do more things, our scientists are more flexible and faster in incorporating new capabilities. Adding new instruments has become faster the more instruments we have. So it goes with what type of company we are. There are echoes of hyper scaling here: scaling in software is backed by scaling in hardware. Having tens of thousands of square feet of lab coming online with dozens to hundreds of instruments gives us the breadth to move fast. For example, my colleague Heather Kulik was on this podcast recently. One area we're seeing disruption is in materials made from the interaction of a molecule with a metal. Our models had been trained on small molecule drug discovery, and all the chemistry they learned carried over to reasoning about metal-organic framework materials that can take CO2 out of the air or filter ammonia.
我觉得这很迷人。很多时候我看到人们做机器学习,他们在一个大数据集上训练,然后迁移到新领域,但迁移量通常很小。
I find that fascinating. Many times I've seen people work on machine learning where they train on some big dataset and then move to a new domain, and often the amount of transfer is small.
是的。
Yes.
那么,是不是有一组“原色”你经常组合在一起,从而产生你的实验?
So, are there a group of primary colors that you combine together that often result in your experiments?
在生物学方面,明显的候选是核酸能力、自我再表达以及我们关心的下游资产。这些核心能力可以衍生出大量不同的事情。在材料方面,我认为是配方。它如此常见和重要,以至于一开始我们并没有把它当作那些更炫酷的方向之一。很多工业界和世界各地的人都关心配方:混合液体和粘稠的东西来制造其他粘稠的东西。比如润滑剂、睡眠纳米颗粒、除臭剂、消费品、工业产品、药品、用于皮肤移植的凝胶。所有这些都源于混合粘稠材料。这是我们正在打造的一项能力,一个非常常见的平台,随着我们与更多人交流,它不断出现。
On biology, the obvious candidates are nucleic acid competence, self-re-expression, and downstream assets we care about. These core competencies give rise to a factorial number of different things you can do. On the material side, I think formulation. It's so common and so important that it wasn't one of the first flashier places we thought of. A lot of people in industry and the rest of the world care about formulation: mixing liquids and gooey things to make other gooey things. That's lubricants, sleeping nanoparticles, deodorant, consumer products, industrial products, medicine, gels for skin grafts. All these things emerge from mixing gooey materials. That's a muscle we're building, a very common platform that shows up all the time the more we talk with people.
有意思。
Interesting.
我在思考 Scaling 时有一个经验法则:每次系统规模扩大一个数量级,你面临的问题就会完全改变。你们选了最难的两个问题:材料和生物。它们以难以推向市场而闻名,时间跨度 10-15 年,尤其是材料的 Scaling。你们是怎么考虑的?你们只是说我们做发现,还是说我们会做到那一步?
I have a rule of thumb when thinking about scaling: every time you scale an order of magnitude in a system, your set of problems completely changes. You guys picked the two hardest problems: materials and bio. They are notoriously difficult to get to market, with 10-15 year time horizons, especially for materials scaling. How are you thinking about that? Are you just saying we're discovery, or that we'll get to it?
我停止做学术讲座前准备的最后一个讲座叫做《材料与化学中 Scaling 的苦乐参半的教训》。在 AI 中,Scaling 是好事,因为它给了你路线图。但在化学和材料中,Scaling 很可怕,因为只有那些能规模化的东西才重要。我们非常清楚这一点。我们的产品团队和实验室团队都知道。例如,在量子点案例中,我们用同一个配方从个位数毫升做到了近一升。在某些地方,我们当前的能力已经触及 Scaling 和技术成熟度。我们正在构建系统,使其能够在进行实验时就推理出以后什么会重要。这将体现在我们的无稀土或无铂族催化剂中。问题的本质恰恰在于我们需要能够规模化这些。所以从第一个实验开始,我们就考虑供应链。我们已经读了每一篇论文。我们有一个技术经济分析智能体,随时准备对我们做的任何事情进行技术经济分析。归根结底,我们不会去做临床试验,也不会为某个特定工艺建造中试工厂。我们会与客户合作,或者如果我们发现某个东西非常棒,就直接去卖。但通常我们会交接,支持客户在治疗发现和材料创新上的痛点。这些也都与 Scaling 有关。
The last academic lecture I prepared before I stopped giving them was called 'The Bittersweet Lesson of Scaling in Materials and Chemistry.' In AI, scaling is a good thing because it gives you a roadmap. In chemistry and materials, scaling is spooky because only the things that can scale matter. We're extremely cognizant. Our product team and lab team all know. For instance, in the quantum dot example, we used the same recipe from single-digit milliliters to almost a liter. There are places where our capability today takes bites into scaling and technology readiness level. We are making the system such that it can reason about what will matter later as experiments are done now. This will be in our rare-earth-free or platinum-group-free catalysts. Precisely, the nature of the question is that we need to be able to scale these. So it's supply-chain conscious as we fire off the first experiment. We've already read every paper. We have a techno-economic analysis agent ready to do the techno-economics of anything we do. At the end of the day, we're not going to do clinical trials or make pilot plants for one particular process. We will work with our customers, or if we find something amazing, we just go sell it. But typically, we hand off, supporting therapeutic discoveries and materials innovations at the pain points our customers have. Those also have to do with scaling.
你们目前进展到什么程度了?
How far have you gotten so far?
在生命科学方面,我同意 Rafa 说的一切。人们会如何使用这个平台,就像科学领域的云代码之类的东西。吸引早期客户的一点是,我们不是一家体内 CAR-T 公司。有很多体内 CAR-T 公司,现在非常热门。6 个月前我们看到了 Capsid 被收购,金额超过 20 亿美元。体内 CAR-T 是一种非常新的治疗模式,以前用于血癌,现在越来越多地用于自身免疫疾病。我们内部确实拥有做体内 CAR-T 所需的三项能力:结合剂设计、LMP 配方和 mRNA 设计。
On the life sciences, I agree with everything Rafa said. The way to think about how people would use the platform is like a cloud-code-ish kind of thing for science. One thing that has drawn early customers to us is that we're not an in vivo CAR-T company. There are lots of in vivo CAR-T companies; it's super hot right now. We saw the Capsid acquisition for over 2 billion dollars 6 months ago. In vivo CAR-T is a very new therapeutic modality, previously for blood cancers but now increasingly for autoimmune disease. We did have the internal triumvirate of capabilities needed to do in vivo CAR-T: binder design, LMP formulation, and mRNA design.
为了让人们知道 CAR-T 是什么,因为它真的非常酷。
Just so people know what CAR-T is, because it's really freaking cool.
它非常酷。
It's so cool.
我也可以聊聊 CAR-T 吗?
Can I talk about CAR-T as well?
可以可以可以。聊聊 CAR-T。
Yeah yeah yeah. Talk about CAR-T.
CAR-T 从 80 年代末或 90 年代就开始研究了,大约在 2010 年左右在癌症领域真正火起来。以前的做法是提取患者的 T 细胞,然后在上面改造一种叫做嵌合抗原受体的结构,告诉 T 细胞去杀死哪种细胞。
So CAR-T has been worked on since the late 80s or 90s. It really caught fire around 2010 or so for cancers. The way it used to work is you'd extract someone's T cells, engineer what's called a chimeric antigen receptor that goes on top of that, which tells the T cell what kind of cell to go and kill.
所以你基本上是在改造人的 T 细胞。把它们取出来,改造一下,让它们表面有这个奇怪的抗原受体。
So you're basically modifying people's T cells. You take them out, modify them so they have this weird antigen receptor on their surface.
就是“寻找并摧毁”标签。通常他们使用一种叫做 CD19 的蛋白质,它优先在 B 细胞上表达。当 B 细胞恶性转化时,会导致血癌和自身免疫疾病。这样做会几乎清除患者所有的 B 细胞库,附带损伤很大。但本质上你是告诉 T 细胞去杀死什么。
Seek and destroy tag. Usually they use a protein called CD19, which is preferentially expressed on B cells. When B cells become malignant, they create blood cancers and autoimmune diseases. You wipe out almost the entire B cell repertoire when you do this. There's a lot of collateral damage. But essentially you're telling the T cell what to go and kill.
这真正开始火起来是在 2015 年左右。它既昂贵又缓慢。你必须提取患者的 T 细胞,进行改造,每次输注大约 40 万美元。它仍然是许多癌症的奇迹疗法。有点跑题,但有一个叫 Emily Whitehead 的孩子在费城儿童医院接受治疗。她是 CAR-T 治愈儿童癌症的首批病例之一。她本来要被转去临终关怀,结果接受了 CAR-T。再稍微跑题一下:她差点因为最初的 CAR-T 治疗发烧而死。她能活下来的唯一原因是治疗她的医生有个患儿童关节炎的女儿,知道某种特定抗体可以抑制她对 CAR-T 的 IL-6 反应。这里面有很多关于 AI 用于科学的内容,所有那些必须发生的机缘巧合。如果你把骰子再掷一千次,很可能不会在那个时刻遇到那个医生,恰好知道该给她用什么抗体,让治疗变得治愈而非致命。所以,这些就是我们想要自动化的那种机缘巧合。
This really started to catch fire around 2015. It was expensive and slow. You have to extract someone's T cells, engineer them. It's like $400,000 per infusion. It's still a miracle cure for many types of cancer. Too much of a tangent, but there's this child named Emily Whitehead who was treated at CHOP. She was one of the first cures in pediatric cancer by CAR-T. She was going to be referred to hospice care, got CAR-T. Another slight tangent: she almost died of a fever from the initial CAR-T treatment. The only reason she survived is because the doctor treating her had a daughter with pediatric arthritis and knew that a specific antibody would blunt her IL-6 response to CAR-T. There's a lot to unpack there in terms of AI for science, all the serendipity that had to happen. If you roll that dice a thousand more times, you probably don't get that doctor at that moment who knew exactly what antibody to give her to make the treatment curative instead of lethal. Those are the types of serendipity we'd like to automate.
所以它又慢又贵。后来人们意识到,只需一次输注,如果你取一段编码嵌合抗原受体的 mRNA,把它包在脂质纳米颗粒里,外面加上 CD8 靶向部分,然后输注进去,它就会结合到 T 细胞上,被摄入,脂质纳米颗粒溶解,mRNA 出来,嵌合抗原受体表达并呈现在 T 细胞表面。
So it's slow and expensive. People then realized that through just an infusion, if you take an mRNA that encodes for the chimeric antigen receptor, put it in a lipid nanoparticle, put a CD8 targeting moiety on the outside, then give it as an infusion, it will bind to the T cell, get ingested, the lipid nanoparticle dissolves, mRNA comes out, chimeric antigen receptor gets expressed and presents on the T cell.
所以你是在重新编程 T 细胞,让它们表达这些奇怪的抗原。
So you're reprogramming the T cells to express these weird antigens.
字面意义上的编程生物学。
Literally programming biology.
是的。
Yeah.
然后 T 细胞就去执行任务,清除任何有 CD19 的东西。
And then the T cell goes and does its thing and wipes out whatever has CD19 in this case.
恶性 B 细胞导致了许多血癌,也导致了许多自身免疫疾病。B 细胞经常针对自身抗原产生抗体。最近,六个月前,作为一位诺贝尔奖得主实验室大约六年工作的成果,加上大约一亿美元的研发投入,我们看到了体内 CAR-T 治疗自身免疫疾病的一些最令人信服的临床前数据。那是一家叫 Capstan 的公司。他们被 AbbVie 以大约 21 亿美元收购。
Malignant B cells explain a lot of blood cancers. They also explain a lot of autoimmune diseases. B cells often make antibodies in response to autoantigens. Recently, six months ago, as the result of about six years of work spun out of a Nobel Prize winner's lab, and about a hundred million dollars of R&D, we saw some of the most compelling preclinical data for in vivo CAR-T treatment of autoimmune diseases. It was by a company called Capstan. They were bought by AbbVie for about 2.1 billion dollars.
在 Lyell,我们一直在孤立地研究这三件事。大约六个月前,Lyell 内部一个两到三人的团队尝试看看我们在体内 CAR-T 方面能做什么。我们一直在研究的是 mRNA 设计。对于大多数 RNA 药物来说,你可以调节的最大旋钮是表达峰值和表达持久性:当你给人接种疫苗或其他 mRNA 药物时,每单位 mRNA 能产生多少蛋白质?我们开发了一些超强的 UTR(非翻译区),它们位于蛋白质编码区两侧,决定了这些表达特性。大约是 Moderna 和 Pfizer 参考值的 10 倍。在六个月的时间里,我们在非人灵长类动物中获得了体内数据,B 细胞清除效果明显优于 Capstan 的数据,持久性也显著更好。拥有更多的 CAR 表达可能是改善 CAR-T 疗法最有效的方法之一。受体的数量决定了 T 细胞在找到坏细胞后与之结合的可能性。T 细胞是连环杀手:它们杀死一个细胞,然后去下一个。它们能持续多久取决于 CAR 表达的持久性。
At Lyell we had been working on all three of those things in isolation. About six months ago, a team of two or three people inside Lyell tried to see what we could do in in vivo CAR-T. What we had been working on was mRNA design. For most RNA medicines, the biggest knob you can turn is expression peak and expression durability: how many proteins do you get per unit of mRNA when you give someone a vaccine or other mRNA medicine? We developed some monster UTRs, untranslated regions which flank the protein coding region, which dictate those expression properties. Something like 10x the references from Moderna and Pfizer. Over the course of six months we got to in vivo data in non-human primates where B cell depletion was significantly better than what was shown in the Capstan data, and the durability was also significantly better. Having more CAR expression is probably one of the most potent ways to improve a CAR-T therapy. The number of receptors expressed dictates how likely that T cell is to bind to the bad cell once it finds it. T cells are serial killers: they kill a cell, then go to the next one. How long they can do that is dictated by the durability of CAR expression.
我们不是一家 CAR-T 公司,我们是科学极客,喜欢做酷的事情。我们在大约六个月内达到了那个验证点,一直到你可能考虑为新的临床资产提交 IND 的阶段。我们不会那样做,我们不会进行临床试验。那会占据全部精力。但一些在 Lyell 待了很久的人认为,这本质上可以做成一个两到三人的全职等效初创公司。几个有领域知识的科学家,结合模型和平台,可以在六个月内完成五年的生物技术工作,只花总投资的 10%。我们现在考虑的很多商业关系本质上就是零全职等效初创公司模式。有人带着一个想法来:如果市场上有一个 CAR-T 可以结合两种东西,如果是双特异性的,或者有其他特性,我知道那个东西能填补的市场空白。我们的商业合作实际上就是现在运行在 Lyra 上的虚拟初创公司。有人带着一个明确的问题来。他们不知道如何实现。可能涉及靶点识别之类的事情。但他们可以在更短的时间内、以更低的成本有效地运行整个项目。
We're not a CAR-T company, we're science nerds, we like to do cool stuff. We got to that proof point in about six months, all the way up to where you might think about filing an IND for a new clinical asset. We're not going to do that, we're not going to do a clinical trial. That would be all-encompassing. But some folks who had been around Lyell for a long time saw that as a way to do essentially a two to three person FTE startup. A couple scientists with domain knowledge, combined with the model plus platform, can do five years worth of biotech work over a six-month period for 10% of the total investment. A lot of the commercial relationships we're thinking about now are essentially the zero FTE startup model. Someone comes with an idea: if there was a CAR-T in the market that could bind to two things, if it was a bispecific, or if it had these other properties, I know the hole in the market that thing would plug into. Our commercial engagements are effectively virtual startups running on Lyra now. Someone comes with a well-specified problem. They don't know how to get there. There may be things related to target identification. But they can effectively run that entire program over a much shorter amount of time at a fraction of the cost.
所以合作伙伴来找你,说,我有这个想法。我不想建实验室。我不想招团队。我只想把它做成。
So a partner comes to you and says, I have this idea. I don't want to build a lab. I don't want to hire a team. I just want to get it done.
没错。
Yep.
所以可能是大学里的某个学者,说,我有这个想法。我做过一点验证。
So it could be some academic at a university that's like, I have this idea. I kind of did a little bit of validation.
我觉得这能行。我能参与进来吗?跟你们团队一起干六个月,把它做成?
I think it'll work. Can I do it? Can I sit with you guys for six months and make it work?
对,这么想就对了。合同上大致是这样:有一笔平台使用费,我们要支付试剂、系统运行和一些管理费用,然后还有利润分成。这是一个可扩展的模式。随着平台越来越好,我们不用只做几十个,而是可以同时做几百个、几千个虚拟初创公司,在平台上开发。这样我们既有短期收入来维持运营,又能与决定跟我们合作的人建立利润共享的伙伴关系。
Yeah, I mean that's the right way to think about it. The way it contractually plays out is there's a platform access fee, we have to pay for reagents and running the system and some overhead, and then there's some upside sharing. That is a scalable model. As the platform gets better, instead of doing dozens of those, we can do hundreds and then thousands of simultaneous virtual startups being developed on the platform, where we have revenue that helps pay the bills in the near term, but then we also have this upside partnership with folks who decide to build with us.
这太棒了,因为我们看到的是:人们越来越倾向于摆脱所有多余的基础设施,利用自动化,专注于创意本身。
It's amazing because this is what we're seeing: people are more and more pushing towards getting rid of all the extraneous infrastructure and using automation and focusing on the idea.
我的想法是:我们大多数人投身科学是因为好奇、想解答问题。我本身是计算机科学家,喜欢通过软件来回答问题。但如果我必须用二进制编程,那乐趣就大打折扣了。
The way I think about it is like most of us got into science because we're curious and want to answer questions. I'm a computer scientist by training and I like to answer questions through software. However, if I had to program in binary, I would enjoy that significantly less.
这些都是高层抽象,越来越高层。以前只有 Python 和 Java,现在有云代码帮我更快地回答问题。打个比方:科学家们现在还在用二进制编程。他们有一个问题想回答,必须把它编译成实验方案,然后去做体力活,移液移到得关节炎。这就是科学界的二进制编程。所以我们想帮科学家提升抽象层次。也许你的想法行不通,大多数临床试验都会失败,但如果你不用同时做体力活和部分脑力活来把各个环节拼凑好,至少可以快速失败。
They're high-level abstractions, increasingly high-level abstractions. It used to just be Python and Java. Now it's like Claude Code that helps me answer questions faster. The analogy is that scientists are still programming in binary. They have a question they want to answer, they have to compile that down to an experimental protocol, then go do the manual labor and get arthritis moving liquids from one place. That's the equivalent of scientific programming in binary. So we're trying to help scientists move up the abstraction ladder. Maybe your idea isn't going to work. Most clinical trials fail, but you can at least get to failing fast if you don't have to do both the physical labor and some of the intellectual labor to get all the pieces in the right place.
大多数临床试验都会失败。只有大约 5% 到 8% 的临床试验能从 IND 走到获批。所以发现本身并不是瓶颈。
Most clinical trials fail. Somewhere between 5 and 8% of clinical trials actually get from IND to approval. So the discovery is not actually the constraint.
我感兴趣的是——你刚才提到了一种经济建模智能体。我记不清你具体怎么称呼它了,但这似乎就是需要解决的问题。你怎么看?
I was interested—you were talking about the sort of economic modeling agent. I can't remember what exactly you called it, but that seems like the problem to solve. How do you think about this?
你是指临床试验的成功率。
You mean the success rate of clinical trials.
嗯,是生物和材料领域 Scaling 背后的经济模型。通常有一个巨大的流程。一旦你有了你认为最终的东西,比如 IND 或材料领域的开发候选物,通常还需要大约 10 年的临床试验或材料科学领域的认证才能把它变成产品。而瓶颈往往是规模化生产、监管、安全性等问题,这些通常很难提前回答。所以每次你都得掷骰子。人们处理这个问题的典型方式基本上是投资组合模型,从融资角度看,你唯一能赚钱的方式就是在一定风险校准下进行 Scaling。
Well, the economic model underlying scaling in general for both bio and materials. Oftentimes there's this huge process. Once you have something which you consider final, like an IND or development candidate for material, there's still usually like 10 years of clinical trials or qualification in the material science world to just get that into a product. And often the bottlenecks there are things about scale manufacturing, about regulatory, about safety, things that are often just very hard to answer up front. So every time you do this you just have to roll the die. The typical way people deal with this is essentially a portfolio model, and financing-wise it's very much the only way you can make money is if you scale with some level of risk calibration.
对。
Yeah.
听到你能具体做这些事情真的很令人兴奋,但这对更大的问题有什么帮助呢?即使你立刻解决了这些问题,那也只是整个问题的 10%。
It's really exciting to hear that you can do these things specifically, but how does it feed into the larger thing where even if you solve these problems immediately, it's still only 10% of the problem?
美国生物技术输给中国生物技术的原因不是创新问题。还有一个监管框架需要支持快速临床试验。FDA 最近在这方面有所动作,既涉及某些情况下需要提交的临床前数据,也涉及我们如何运行和监控试验。所以认为一家公司能独自改变这一点是疯狂的。必须与监管机构协同推进。然而,临床前成功概率的微小提升非常重要。从投资组合理论的角度看,这会使投资更具吸引力。这意味着预期中药物能更快到达患者手中,失败更少。所以我要说,这正是我们目前关注的领域:一个受益于一百万个独特 mRNA 设计、最大化已知能转化为治疗益处的因素的系统所创造的药物,将显著提升那些临床前成功率。我想说,掷一个灌铅的骰子总比掷一个公平的骰子好,所以我们只是尽量让骰子灌铅。
The reason why US biotech is losing to Chinese biotech is not because of an innovation problem. There's a regulatory framework too that has to go to enabling fast clinical trials. The FDA has made motions towards that recently, both for the preclinical data you have to submit in some cases, but also how we will run and monitor trials. So it would be crazy to think that one company could change that on its own. It has to be done in tandem with the regulators. However, the minor moves in preclinical probability of success matter a lot. From a portfolio theory perspective, it makes the investment much more attractive. It means in expectation medicines get to patients faster, fewer of them fail. So I would say that is the area we're focusing on now: a medicine created by a system that has had the benefit of a million unique mRNA designs to maximize things known to translate to therapeutic benefits will meaningfully move those preclinical success rates. I guess I would say it's better to throw a loaded die than a fair die, so we're just trying to make the die as loaded as possible.
我想我的想法——这是我一直在思考的问题——是如何实现转化,在材料领域里对应的概念是什么,我们怎么称呼它。我真的很想看到一个模型,它能思考这些因素,进行推理,并且非常擅长说:我把我的设计筛选到那些我认为能通过三期临床的。
I guess my thinking—this is the thing I think about constantly—is how do you bring translation, right, in whatever the equivalent is in materials, so how we name that. I really want to see a model that thinks about these are the factors and reasons about and is very good at saying I'm filtering my designs to the ones that I think are going to make it through phase three.
我妻子是生物技术领域的转化科学家,所以我经常看到,确实,你们应该用 AI 做转化科学。从某种意义上说,我认为这是一种呼应,尤其是在材料和化学领域。我们的工具可以调用过程工程模拟器,计算出应该使用多大的管道直径和换热器,才能让工艺规模化在经济上划算。所以,也许——我不知道我们是否会在测量之前就设置一个过滤器——但能够提前推理下游会发生什么,这正是转化 AI 要做的事:提前推理什么才是重要的。因为我想说,一旦你有了 IND,你就被锁定了,在化学领域也是如此。
My wife is a translational scientist in biotech, so I see it reminds me very often: yeah, you guys should be doing AI for translational science. In a sense, I think that's sort of the echo, especially in materials and chemistry. Our tools can call process engineering simulators and go figure out what pipe diameters and what heat exchangers you should be using in order to scale up the process for the economics to be worth it. So still, maybe there is—I don't know if we're going to gang up a filter until we go measure it—but the ability to reason now about the things that will come downstream, which is sort of what translational AI would do, is to reason now about what's going to matter. Because I would just say earlier, when you have the IND you're locked in, and it's true on the chemistry side.
你选择的序列对应的分子、你要给哪个群体用药、以及如何衡量成功——这些都是你之后要做的选择。我要说明一下,我不做临床前的工作,我做的是化学和材料。而这正是我们现在试图通过验证器和数据源来灌输的行为类型——这些数据源要么是有人已经考虑过的,要么是物理条件允许的,要么是我们能测量出足够好的代理指标来预示后续结果。
The molecule that the sequence you've chosen, of course which population you're going to give it to and how you're going to measure success, those choices you make afterwards. And I would say, to be clear, I don't work on the preclinical stuff but on the chemistry and materials. That is precisely the type of behaviors we're trying to instill now with the verifiers and the data sources that we can access, either because somebody has thought about them, either because the physics allows it, or because we can measure good enough proxies that tell us what's going to happen later.
而且明确一下,所有这些事情都是我们内部经常讨论的。我们已经快要病态地过度限定范围了。
And to be clear, all those things are things that we talk about internally a lot. We're already on the verge of being pathologically over scoped.
但我绝对相信,随着模型变得更聪明,随着它们消化 clinicaltrials.gov 的数据,随着我们与制药公司合作并接触到那个“饼干罐”——即试验前的概率——这些概率将发生有意义的改变。在生物制造方面,能够接触到良好的生物制造工艺和放大工艺,我们认为模型也能在那里做出贡献。我们只是选择将大量商业和合作活动集中在我们认为目前可以解决的科学前沿上,但目标是超越它。
But I'm absolutely like the belief that we have is that as models get smarter, as they ingest clinicaltrials.gov, as we partner with pharma companies and get access to that cookie jar of pre-trial probabilities, these probabilities will meaningfully change. On the biomanufacturing side, having access to biogood manufacturing processes and scale-up processes too, we think the models will be able to contribute there. We've just chosen to focus a lot of our commercial and collaborative activity on the sort of frontier of science that we think we can address now, but the goal is to push past that.
如果让我总结,这有点像工具调用。
If I may summarize, it's kind of like a tool call.
是的,全是 token,全是工具调用。Token 和工具调用就是一切。
Yeah, it's all tokens, it's all tool call. Tokens and tool calls are all you need.
但还有你提到的那个治疗 IL-6 抗体的医生的推理机制。那个人是从生活经验和阅读文献的结合中学到的。既然我们相信我们的论点——广度能给我们带来优势——我们就会通过做更多我们正在做的事情来变得更好。
But also the reasoning mechanisms that maybe you mentioned for that doctor that treated the IL-6 antibody that you mentioned. That person had learned that from a combination of lived experience and reading the literature. And since we believe our thesis that the breadth gives us there, we will get better at those things by doing more of the things we do.
在那么多反事实的世界里,那位医生并不是治疗 Emily Whitehead 的人,而 CAR-T 可能看起来只是 Eroom 定律上的又一块墓碑——又一个失败的药物。所以我确实认为,成功率从 2% 变成了 98%,仅仅因为那个人恰好就在现场。所以如果我们能把这一点操作化,你就能大幅改变概率。
So many counterfactual worlds that doctor was not the one treating Emily Whitehead in that case, and CAR-T may have looked like it might have been yet another gravestone in Eroom's Law, yet another failed drug. So I do think that went from a 2% success probability to a 98% just because that person happened to be in the room. And so if we could just operationalize that, you're going to move a lot of probabilities.
这是一个很好的例子,说明仅仅拥有非常广泛的科学信息知识就能起作用。所以这几乎就像谷歌,因为那个药当时只用于小儿关节炎——另一个非常小众的医学领域。
That's a good example of where just having really broad knowledge of scientific information. So that's almost like Google, because that was only being used in pediatric arthritis, another very niche area of medicine.
我明白了。是的。
I see. Yeah.
你的团队里有 Ken Stanley,他著有《为什么伟大不能被计划》一书,非常推崇研究中的开放性和偶然性。那么开放性在 Lila 扮演什么角色?
So you have Ken Stanley on your team, famously wrote the book 'Why Greatness Cannot Be Planned' and is very big on open-endedness and serendipity in research. So what is the role of open-endedness at Lila?
哦,是的,Ken 非常棒。对于那些不了解的人,Ken 开创了机器学习和 AI 的一个领域,叫做开放性,我认为这是机器创造力。比如,我们如何让模型进行开放式探索,并且对什么有趣、我们应该探索什么有品味。如果你只是一个擅长考试的人,你不可能拥有科学超级智能。如果你想想强化学习即使在大规模下在做什么,它是以一种冷酷的瓦肯人式、Spock 式的方式回答问题,但你很可能只在有限的程度上认为那个模型是极其有创造力的。所以 Ken 在 Lila 创建了一个开放性团队,来承担推理挑战的外循环或元部分。我们如何让我们的模型不仅能回答难题,而且首先能提出有趣的问题?这确实是 Ken 的任务。过去几个月他一直在组建一个世界级的团队,他们现在正在厨房里烹饪。我想到今年年底,我们会看到 Ken 团队的一些很酷的东西。
Oh yeah, Ken is awesome. For those of you who don't know, Ken pioneered an area of machine learning and AI called open-endedness, which I think of as machine creativity. Like how do we get models to do open-ended exploration and also have a sense of taste about what's interesting, what things we should go down. So you can't have scientific superintelligence if you're just a good test taker. If you think about what reinforcement learning is doing even at scale, it's answering questions in a kind of ruthlessly Vulcan-esque, Spock kind of way, but you probably only in limited ways would think of that model as being supremely creative. And so Ken has created or built an open-endedness team at Lila to sort of take on the outer loop or the meta part of that reasoning challenge. So how can we get our models to not only be able to answer tough questions but ask interesting questions in the first place? That's really Ken's mandate. He's been building a world-class team over the last several months, and they're in the kitchen cooking now. I think by the end of this year we'll have some cool stuff from Ken's group to share.
我们来看一段实验室的视频,它会展示几个不同的东西。好的,那可能是一个剥离器或封膜机。当你把板从一个仪器移动到另一个维度时,显然里面有液体,大多数生物学都是湿的。所以你要把这些贴纸贴在上面。那是一个板正在被封膜。好了,开始了。它正在拿起一个……是的,这是在液体处理器的内部。我们等它切换到更广的镜头,这样你就能看到 PCI 总线和一些机器人。液体处理器……
We're going to hop into a video here of the lab that's going to show a couple of different things. Okay, so that's probably a peeler or sealer. When you move plates from instrument to dimension, obviously there's liquid in it, most of biology is wet. So you put these stickers on it. So that was a plate being sealed. All right, so here we go. It's picking up a... Yeah, so this is inside of a liquid handler. Let me go, we'll wait till it gets to a wider shot here, so that you can see the PCI bus and see some of the robotics. So the liquid handlers...
是磁性的……是的,这是平面电机系统,板在这里磁悬浮。这是 PCI 总线,传输层连接所有仪器。你可以看到那里有放置所有仪器的台子。机械臂把它拿起来,现在要把它转移到另一个板,以便进行下一步。这里有一点交通控制要做,实际上它们会停下来一会儿,等交通拥堵缓解。
Is the magnetic... Yeah, so this is the planar motor system here where the plate magnetically levitates. This is the PCI bus where the transport layer connects all the instruments. You can see benches there where all the instruments sit. Robot arm picks it up, is now going to transfer it to a different plate to go on to the next. There's a little bit of a traffic control thing that you have to do here, like they actually will go and park for a while while traffic congestion clears.
嗯。
Yeah.
这是 PCI 总线的远景。同样,这一切都是完全受控的,完全自动化的。这是一个材料科学的例子。
And here's a long shot of the PCI bus. And again, all that's fully controlled, all that's fully automatic. And this is a material science example.
这是一个物理科学的例子,是的,它把我们带回了 Scaling 的观点。这里,它正在以放大的形式制造我们的氢催化剂——那是一种含有材料纳米颗粒的墨水。那是一个旋涂机,你可以从它旋转板这一点猜出来。然后,这是一个机器人处理器在移动小块催化剂进行测试。还有这种好看的紫色 90 年代霓虹风格……
This is a physical science example, yeah, where it takes us back to the scaling point. Here, it's making our hydrogen catalysts in a scaled-up form factor by... that's an ink that contains nanoparticles of the material. That's a spin coater as you can guess from the fact that it spins the plates. And then, this is a robotic handler moving around little piece of catalyst to test. And this nice-looking purple 90s neon vibe...
这叫做磁控溅射机,我们让原子从源飞出,沉积在腔室的另一侧,形成非常薄的原子薄膜,我们可以根据三四个源上的元素任意混合。我们只是将它们汽化,让它们飞过腔室,制成这些非常节省材料的薄膜。我们可以用非常少的材料做到这一点,它是我们在催化、腐蚀、机械性能等许多应用中快速设计、制造、测试的主力设备之一。很多东西都可以在这种非常方便的形式下进行测试。
This is called a magnetron sputtering machine where we make atoms fly from a source and deposit on the other side of the chamber in a very thin atomic film, where we can make arbitrary mixes of elements based on what's on the three or four sources. We just vaporize them and make them fly over the chamber and make these nice thin films that are very material efficient. We can do this with very little material, and it's one of the workhorses for us to design, make, test fast in many applications in catalysis, in corrosion, in mechanical properties. Many things you can test in this sort of very convenient form factor.
液体处理器和一些那些机器,大部分是现成的。然后你们想出了这种形式,适用于很多那些机器,无论是材料还是生物领域。
The liquid handlers and some of those machines, those are kind of off-the-shelf mostly. And then you've come up with this sort of form factor that works for lots of those machines both for material and for bio.
量子点就是两者结合的一个很好的例子。
Quantum dot is a good example of the combination of the two.
这其实是一个我们改造用于量子点合成和酶的液体处理器。我认为这说明用 20、30、40、50 台仪器就能走得很远。关键是它们必须放在模型能够使用的平台上。这对我来说是加入 Laila 后的一个重大启示,因为我之前并没有真正接触过实验室自动化。这不是我期望的那种自动化。很多自动化都是点自动化:液体处理器旁边挂着一个平板,你可以输入数据,但那个设备本来就不打算——有时甚至是故意设计成——不与其他设备通信。所以我们做了很多工作,我开玩笑说我们拥有生物学领域最大的保修失效集合,因为我们编写了自己的定制驱动和固件,以获得对这些仪器的底层精细控制,并让它们互相通信。视频很酷,因为你看到了磁悬浮板。但你看不到的是将所有东西缝合在一起的定制软件封装。这很大程度上归结于非常棘手的软硬件接口挑战。有些机器仍然运行着 Windows 95。想想你怎么自动化它。我们实际上用了一个视觉语言模型来控制一台 Windows 95 机器,因为那是唯一的自动化方式。
It's actually a liquid handler that we've repurposed for quantum dot synthesis and enzyme. I think that speaks to how far you can get with 20, 30, 40, 50 instruments. The key is they have to be on platforms that the model can use. This was a big eye-opener for me coming into Laila, because I wasn't in lab automation meaningfully before. It's not the automation I was hoping for. A lot of automation is point automation: there's a tablet attached to the side of a liquid handler where you can enter data, but that device is not meant—and sometimes purposely designed—not to talk to other things. So a lot of what we've done, I joke that we have the world's largest collection of voided warranties in biology, because we wrote our own custom drivers and firmware to get low-level granular control over these instruments and make them talk to each other. The video is cool because you see magnetically levitating plates. What you don't see is the custom software wrapper that stitches it all together. A lot of this comes down to really hard software-hardware interface challenges. Some machines still run Windows 95. Think about how you automate that. We actually have a vision language model controlling a Windows 95 machine, because that's the only way to automate it.
嗯。
Yeah.
因为那是唯一的自动化方式。
Because that's the only way to automate it.
开玩笑说用机械手指按按钮,但……
To joke about a mechanical finger pressing buttons, but...
开玩笑,但没开玩笑,我们真那么做了。我们真的用了一个机器人去推设备侧面的 iPad。
Joke, but no, we did that. We actually used a robot to push the iPad on the side of the thing.
另外要指出的是,这仍然是为人设计的自动化。仪器放在大约齐胸高的台面上,因为假设需要有人伸手进去维护或加试剂。这是我们认为实验室自动化会变成的样子的 V0 或 V0.5 版本。因为我们决定垂直整合并拥有整个软硬件栈,V2 看起来会非常不同,我们将能够集成各种东西。这在材料科学领域已经发生了,因为那些能力根本不存在。我们通常用 XY 坐标来思考实验室。随着集成,我们还会加入 Z 维度,因为我们可以堆叠东西。所以单位体积的 token 数将是我们考虑的重点。但我们认为未来的实验室不应该让人轻易走进去。它应该感觉像一个数据中心,你看到一排排服务器机架,后面有空间让维修车通过来维护节点。它应该尽可能密集,同时也尽可能节能。所以回答你的问题,我们现在使用现成的商品化设备,因为这样起步合理,但随着时间的推移,外形因素肯定会发生很大变化。
The other thing to call out here is that this is still automation made for people. The instruments sit on benches which are approximately chest high because there's the assumption that someone needs to reach in to service it or fill reagents. This is V0 or V0.5 of what we think lab automation will look like. Because we've decided to vertically integrate and own the hardware-software stack, the V2 will look very different, where we'll be able to integrate things. This is already happening in material sciences because those capabilities just don't exist. We often think about labs in terms of XY coordinates. As we integrate, we'll have a Z component too because we'll be able to stack things. So tokens per unit volume is what we'll be thinking about. But we think the lab of the future should not be made for people to easily walk into. It should feel like a data center where you see rows of server racks, with room for a crash cart behind to service the nodes. It should be as densely packed as possible and as energy efficient as possible. So to answer your question, we're using commodity things now because it makes sense to get started, but over time the form factors will change quite a bit.
我明白了。我只是有点惊讶,你们能想出这种通用尺寸的托盘,能满足你们大部分问题的需求。
I see. I'm just a little surprised that you can come up with this common size of tray that kind of matches your needs for a good percentage of your problems.
嗯。
Mhm.
这种尺寸能满足你们大部分问题的需求。
That kind of matches your needs for a good percentage of your problems.
嗯,这只是逆向推导。96 孔板是实验室自动化的实验原子单位。所以我们现在在材料科学中也采用 96 孔板的外形。不是所有东西都适合那种外形,但采用 96 孔或 384 孔板格式所能覆盖的范围……
Well, it's just working backwards. 96-well plates are the atomic unit of experimentation in lab automation. So we now do 96-well form factors for material sciences as a result. Not everything fits into that form factor, but the coverage you get from adopting a 96-well or 384-well plate format...
80/20。
80/20.
80/20,没错。是的。
80/20, exactly. Yeah.
是的。我想你可以看到有些沉积材料的碎片更大。所以我们仍然使用板状来搬运它们,但样本数量更少。我想有些可能是 12 个,4x3 排列。
Yeah. I think that you can see some of those where the pieces of deposited material were bigger. So we still use the plate shape to carry them over, but then the number of samples, right? That you have them are smaller. They I think some of them are like maybe 12 4 x 3.
嗯。
Yeah.
这让我想到你们之前问到的 Scaling 问题,以及当你扩大规模时,问题会变得不同。我们公司共同期待的一个问题是数据中心规模的 AI 科学工厂的编排和调度。当你有所有可以并发运行的实验时,你如何考虑移动所有这些样本和连接所有这些仪器的物流与编排,从而为我们的客户创造最大价值,为我们的模型提供最大信息?那才是令人兴奋的部分。那个问题将与我们目前思考的其他问题非常不同。
This takes me to a point you folks asked earlier about scaling and how when you scale, your problems are different. A problem we're looking forward to collectively at the company is the orchestration and scheduling of a data center-sized AI science factory. When you have all the experiments you could run concurrently, how do you think about the logistics and orchestration of moving all these samples and interfacing all these instruments to create maximum value for our customers and maximum information for our model? That's the exciting part. That problem will look very different from some of the other problems we're thinking about now.
我们考虑的——Rafa 说过——在此基础上进行编排就像 Slurm 队列之类的东西,让你全局最大化系统吞吐量。但随着系统变大,使用同样的抽象来思考吞吐量、调度、编排,最大化吞吐量的复杂性也会增加。所以如果你是一个约束满足问题爱好者,我们有一个最酷的问题可以思考。你是把 Scaling 看作一个可以复制的集群,还是把你的所有液体处理器和旋涂仪分布在实验室的不同区域?
What we think about—Rafa said—orchestration on top of that is like a Slurm queue or something that lets you globally maximize throughput of the system. But using those same abstractions to think about throughput, scheduling, orchestration as the system gets larger, the complexity in maximizing throughput. So if you're a constraint satisfaction problem nerd, we have one of the coolest ones to think about. Are you thinking about scaling as one cluster that you just cookie-cutter, or do you have all your liquid handlers and spin coaters in different parts of the lab?
目前,我们拥有的基本上是一个大的全连接图。这无法无限扩展。有些材料会释放有害烟雾,所以出于安全原因要隔离。我不知道未来科学集群的具体配置,但我想它可能比你猜的要少——几百台,也许几千台。但我们确实以与扩展数据中心相同的方式来思考扩展:一栋多层建筑,数百万平方英尺,一个无人值守设施,24/7 运行,实时生成数据,具有与数据中心相同的正常运行时间。这非常难做到。但那是我们试图逆向推导的终点。
Currently, what we have is essentially one big fully connected graph. That won't scale indefinitely. Some material stuff throws off hazardous fumes, so that's isolated for safety. I don't know the exact configuration of the science cluster of the future, but I think it will probably have fewer instruments than you might guess—hundreds, maybe thousands. But we do think about scaling it the same way you'd scale a data center: a multi-level building, millions of square feet, a lights-out facility running 24/7, generating data in real time, with the same uptime you'd expect from a data center. That's very hard to do. But that's the endpoint we're trying to work backwards from.
在通往那个终点的路上,你需要解决哪些问题?
What problems do you need to solve on the way to that endpoint?
不,这又回到了我之前关于实验运行时间的问题,因为 Scaling(规模扩张)有不同的含义。一种是实验设计,它本身具有可扩展性,但可能以信噪比或其他概念为代价,但能以某种成本快速高效地获取广泛数据。另一种 Scaling 是低吞吐量但大规模并行化。所以一般来说,我会把它们视为两类不同的问题。我不认为同一种策略能普遍适用于它们。那么作为科学家,哪种类型的 Scaling 对你更重要?
No, this goes back to my previous question though about the runtime of your experiments too, because scaling means different things. One of them is experimental design, which intrinsically scaled but maybe at the cost of signal-to-noise ratio or some other idea, but getting broad data quickly and efficiently at some cost. Or scaling is lower throughput but just parallelizing wildly. So in general, I would approach those as two different sets of problems. I don't think the same strategy really works for them in general. So what types of scaling is more important for you as a scientist?
我认为逐轮迭代比那种广泛、高度复用、噪声很大的方式更重要。
I would say round-over-round iteration is more important than a broad, hugely multiplexed, highly noisy kind of thing.
所以迭代时间才是真正的关键。
So iteration time is really the single thing.
是的。
Yeah.
好的。那么这是否限制了你想要关注的领域?比如,现在如果我们尝试解决一个新问题,我们是否会问:我们能否用更快的迭代来解决这个问题,而不是用那种需要一个月周转期的大规模复用方式?
Okay. So does that limit the domains that you want to focus on? Like, now do you think if we're going to try to tackle a new problem, we ask can we just solve this problem with faster iteration versus something where the answer is we scale up by massively multiplexing something but with month-long turnaround?
并行化和复用有些不同,对吧?所以有时候……
Parallelizing and multiplexing are somewhat different, right? So sometimes...
没错。
That's right.
我会说混合样本,我们喜欢混合样本。
I would say pooled we love pooled.
是的。
Yeah.
所以我们喜欢混合样本,因为它既快又广。
So pooled we love because it gives you fast and broad.
混合样本是由什么组成的?
What is pooled made from?
混合样本就像 DNA 编码库,里面有一堆乱七八糟的东西,实验结束后你可以把那些垃圾挑出来。
Pooled are things like DNA-encoded libraries, where you have a bunch of crap in it and you can sort out the crap after you do the experiment.
某种形式上,实验设计允许你同时进行一千、一百万或十亿次实验。实验的设置方式使得读出结果能选出优胜者。所以你在一个板上尝试一百万个东西,然后得到一个读出结果,或者一千个优胜者的一千个读出结果。整个复用……
Somehow the form of the assay allows you to throw a thousand or a million or a billion experiments at the same time. And the way the assay is set up, the readout picks the winner. So you try a million things in one plate and you get one readout or a thousand readouts of the thousand winners. Multiplex the whole...
所有生物技术都只是把你想要的任何读出结果映射到 GSC 上,是的,你可以 PS。对,目标,复用。就是这样。是的,你可以获得大量数据。
All biotech is just mapping whatever readout you want onto the GSC and yeah, you can PS. Yes, target, multiplex. There you go. Yeah, you can get lots of data.
一个优雅的论点是:如果这个领域的标准是一个月,而我们只需要四天,那么四天的学习周期就很棒,因为它真的能推动该领域的进展。这就是我们的自动化工程师和团队思考其他测量方式的地方。在冷却剂和催化领域,有些地方我们制造了不同的仪器,测量另一种属性,结果发现响应速度快了一千倍。例如,在吸附方面,我可以简单说一下。在气体吸附中,人们通常测量他们加压一定量的气体。对于我提到的从空气中吸走 CO2 的 MOF 和 COF 材料,根据理想气体定律,你知道你把多少气体放进小盒子,然后等待气体被材料吸附,检查压力,从压力差知道有多少气体进去了。然后你再次加压,看看又进去了多少。如果这听起来很慢,那是因为它确实很慢。这叫做 BET。每个样品需要大约一天,而且很难并行化,因为需要另一条气体管线、另一个罐子。或者你可以从其他可并行化的仪器中获取其他类型的代理测量。这就是我们现在在实验室里构建的东西:我们不测量压力,而是测量另一种我们关心的属性,这种属性是压力能告诉我们的读出结果,但我们可以在大约一小时内对 96 个金属有机框架进行 96 孔板测量。所以速度可能快了 2500 倍。这是一个需要一点独创性的地方。与其他读出结果相比,一小时仍然很慢,对吧?电化学中的其他东西,也许我们可以在几分钟内完成。但现在我们比之前的方法快了一千倍。
The elegant argument would be that if the standard for this field is a month and it's going to take us four days, a four-day learning cycle is amazing because it's really going to move the needle for that part of the field. This is where our automation engineers and our teams are thinking about other ways of measuring things. In coolants and in catalysis, there are places where we just made different instruments that measure a different property that turns out responds a thousand times faster. For instance, in sorption, I can tell you a little bit. In gas sorption, people typically measure they pressurize an amount of gas. For the MOF and COF materials I was talking about sucking CO2 out of the air, you know how much from the ideal gas law, you know how much gas you put in the little box and then you wait for the gas to be adsorbed in the material, you check the pressure and from the difference in pressure you know how much went into the thing. Then you up the pressure again and you see how much extra went. If this sounds slow, it's because it's very slow. It's called BET. This takes like a day per sample and it's very tough to parallelize because it's another gas line, another canister. Or you can take other types of proxy measurements from other instruments that are parallelizable. That's something we built in the lab now where instead of measuring pressure, we're measuring another property we care about that is a readout for what pressure would tell us, but we can do 96 well plates for 96 metal-organic frameworks in like an hour. So it's maybe 2,500 times faster. This is a place where there's a little bit of room for ingenuity. An hour is still slow compared to other readouts, right? Other things in electrochemistry, maybe we can do in a minute. But now we're a thousand times faster than the way we were doing it.
我认为你问题的答案还取决于我们认为模型是从完全空白起步,还是已经会走或会跑。如果我们关心的某个领域或问题,很明显我们使用的基础模型的权重中知识为零,那么我们可能更喜欢用一个大而慢的东西来引入知识。如果我们认为模型在该领域已经相对胜任,那么我们会非常倾向于快速串行迭代循环。所以我们会两者都做。赌注在于:随着模型性能的提升,样本效率也会提高,因此从逐轮实验中获得的复利将超过从一个大而嘈杂但广泛的数据集中获得的收益。
I think the answer to your question also depends on how much we think the model is starting from a dead start versus a walk versus a jog. If there's some area we care about, some question, it's clear there's zero knowledge in the weights of the base model that we're using, then we may prefer a big slow thing to move it in. If we think it's already relatively competent in that, then we would vastly prefer the rapid serial fast iteration cycle. So we will do both. The bet is that as the model performance improves, the sample efficiency goes up, and therefore the compound interest you get from round-over-round experimentation will outweigh that from a big noisy but broad data set.
那么你有没有担心?这只是我随口说说,但你是否担心你会很快饱和那些可以用……解决的问题?
So do you have any concern? This is just thinking out loud, but do you have concern that you're going to quickly sort of saturate the problems that you can solve using...
担心还是希望?
Concern or hope?
或者两者都有。也许是担心和希望。但也许你正在部署这些系统,现在因为它们很新,所以有很多空白领域你可以去攻克所有适合高通量实验的问题。你可能会这样做几年,然后突然一切都变了,你不得不彻底改造你价值数百万美元的基础设施。
Or either. Concern and hope, maybe. But maybe you have these systems that you're putting in place, and right now because they're new, there's a lot of green field you can go and tackle all these problems that are amenable to high-throughput experimentation. You're going to do that for a couple years, maybe, and then all of a sudden everything is different, and you have to completely retool your like multi-million dollar infrastructure.
明确地说,我希望这是真的。我希望两年后我们不必再测量结合 KD。如果不用再做那个,我会非常兴奋,因为模型已经基本掌握了结合动力学。
I hope that that is true, to be clear. I hope that we don't have to measure a binding KD again in two years. If we didn't have to do that, I'm very pumped about that because the model has essentially mastered binding kinetics.
所以你认为最终会达到模型知道如何做那一步,你不需要……
So you would think that eventually you get to the point where the model knows how to do that, you don't...
让我们再回到 PCI 总线。我们真正想做的是减少将新仪器引入平台所需的时间。你希望那感觉很像 USB。我不知道你们多大年纪,但我小时候,你买一个新设备,驱动在软盘上。你得撞墙才能把驱动装上,两天后你的打印机还只是勉强能用。
Let's go back to the PCI bus again. What we actually want to do is to reduce the time it takes to bring a new instrument on platform. You want that to feel a lot like a USB. I don't know how old you guys are, but when I was a kid, you got a new device, you got the drivers on a floppy disk. You had to beat your head against the wall to get the driver to install, and two days later your printer only kind of works.
是的。
Yeah.
所以那有点像……
So that's kind of what...
对对对,完全正确。如果你是个 Linux 硬核用户,你今天仍然可以体验那种感觉。
Yeah, yeah, exactly. And if you're a Linux hardcore person, you can still live that experience today.
你的音频驱动还是不能用。
Your audio driver still doesn't work.
这就是目前在生物学和物理科学领域将新仪器引入平台的现状——我们就像在用软盘上的驱动程序和说明书来让它工作。所以我们希望统一平台能实现的是,仪器接入时间最终趋近于零:厂商提供规格,模型读取它,正确的 API 被抽象出来。我们正在与一些仪器厂商合作,让这个过程更容易,但我认为我们思考如何调节系统的方式很大程度上受限于当前的做法。所以我们希望,两年后统一平台能让仪器接入从 30 天的任务变成 30 分钟的任务。这很难,我可能错了,我们也许做不到,但那就是我们指向的未来。目前,我们实际上已经可以非常快速地更换现有仪器。比如,如果我们需要把 Hamilton 换成另一种液体处理器,这个更换已经可以很快完成。所以我们有合理的信心,仪器接入会随着时间变得更快、更好、更可靠。我们不想在 2036 年做 2026 年的科学。所以我们希望一些仪器会被淘汰,或者我们的测量方式会改变。否则,我们和其他人对未来十年进步速度的很多假设就都错了。
So that is what it's like to bring a new instrument onto a platform in biology and physical sciences now — we're in the driver on a floppy disk and the manual to try to get it to work. So again, one of the things we hope a unified platform enables is instrument onboarding time eventually goes to zero, where you have the spec from the manufacturer, the model reads it, and the right APIs get abstracted. We're working with some instrument vendors to make this process easier, but I think a lot of the way we think about modulating a system is conditioned on how we do it now. So we're hoping that a unified platform makes onboarding an instrument two years from now a 30-minute exercise versus a 30-day exercise. It's a hard thing to do; I could be wrong, we might not be able to do it, but that is the future we're pointing to. Currently, we actually can swap out existing instruments very quickly. So if we need to replace a Hamilton with a different liquid handler, that swap happens very quickly already. So we do have some reasonable belief that onboarding instruments will get faster, better, more reliable over time. We don't want to be doing 2026 science in 2036. So we hope that some of these instruments get deprecated or the way we measure things changes. Otherwise, lots of assumptions we and everyone else made about the rate of progress in the next decade will have been wrong.
而且我们已经从仪器厂商那里受益了,对吧?我希望我们遇到的问题是你描述的那种——我们做完了所有能用这些仪器做的科学。那会提高对仪器厂商的要求。我们现在的仪器和 10 年前的光束线一样强大。我们今天做的测量,10 年前需要你向联邦政府申请凌晨 2 点的时段,在某个地方浪费几个晚上的睡眠,在非常亮的中子或 X 射线源上测量。而今天,厂商制造的仪器可以放在量子点旁边或蛋白质表达旁边。
And we're already benefiting from the instrument vendors, right? I wish the problem we have is what you're describing — that we'd run out of science to do with the instruments. That would up the ante for the instrument vendors. The instruments we have now are as powerful as a beamline would have been 10 years ago. We're taking measurements today that 10 years ago would have required you to ask the federal government for a time slot at 2:00 in the morning somewhere out there, to waste a couple of nights of sleep taking measurements at a really bright neutron or x-ray source. And today the vendors make instruments like those that we can put next to the quantum dot or next to the protein expression.
是的。
Yeah.
我希望那是一个理想但非常非常不可能的终局。而且我确信会有新的科学问题需要我们用现有仪器来回答。
I wish that's an end state that is desirable but very, very unlikely. And I'm sure there's going to be new science to be asking of the instruments we have.
我们有过嘉宾谈到这两个主题。首先,你买的任何设备都不是为高通量 AI 科学设计的。其次,每天都有新的科学设备出现,它们打开了 5 年、10 年前不可能的事情。比如在线核磁共振,有很多表征、小型化,还有更高的分辨率、更亮的光源,这些都具有变革性,并且与我们正在做的自动化高通量科学非常契合。
We've had guests who have had both of these themes. First of all, none of the devices you buy are set up to do high-throughput AI science. And also, there are new scientific devices which come up every day that just open up something which was impossible like 5, 10 years ago. Like inline NMR, there's lots of characterization, miniaturization, and also more resolution, more bright sources that are just transformational and they marry really well with the kind of automated high-throughput science we're doing.
所以我们正在搬进马萨诸塞州剑桥的这个设施,这只是一个 3D 渲染图。这是一个 10 万平方英尺的空间,我们将采用 AMR(自主移动机器人)作为部分运输方式。你可以在那里看到一些。我们会把它放在节目笔记里。好了,稍微换个话题。你之前提到了你的 10 万亿 token 的科学数据堆。
So we're moving into this facility in Cambridge, Massachusetts and it's just a 3D rendering. It's a 100,000 square foot space and we will move towards AMRs, autonomous mobile robots as some of the transport. So you can see some of that there. We'll put it in the show notes. Okay, so kind of switching topics a little bit. So you were talking about your scientific pile of 10 trillion tokens.
嗯。
Mhm.
当我听到 10 万亿时,我的第一反应是:‘哇,听起来很多。’这大约是 3000 个人类基因组,测序成本大约 300 万美元。这大约是 Evo 和核苷酸 Transformer 等大型基础模型的 1/2000。所以从某种意义上说,这是很多数据。从另一种意义上说,这不算多。而且并非所有 token 都相同。所以我很好奇:创建这个数据堆的过程是怎样的?你的思考过程是什么?然后,10 万亿 token 中实际有多少有用信息?
When I hear 10 trillion, my first thought was, 'Man, that sounds like a lot.' Things like this is 3,000 human genomes, which would cost roughly $3 million to sequence. It is roughly 1/2000 of the several large foundation models like Evo and nucleotide transformer and so on. So in some sense, it is a lot of data. In another sense, it's not a lot of data. And not all tokens are the same. So I'm curious: what went into creating this? What were your thought processes? And then how much actual useful information is in 10 trillion tokens?
是的。所以这些 token 和我们从互联网或后训练运行中计数后训练 token 的方式一样。所以这些再次——RL 是思考 RL 的最佳方式,它是一种数据生成机制。它是一种引导模型产生越来越有价值的 token(比如更好的 token)的方法。所以这些是在 Lyra 的许多不同科学 RL 环境中运行该过程的结果,其中 token 混合了英语、工具调用和实验反馈。所以它们是我们一直在讨论的准英语 token,由 tokenizer 分词。这就是它们的来源。
Yeah. So it's tokens in the same way that we think about counting post-training tokens from the internet or from post-training runs. So these are again — RL is the best way to think about RL is a data generation mechanism. It's a way to steer the model towards more and more valuable tokens, like better tokens. And so these are the result of running that process across many different scientific RL environments at Lyra, where the tokens are a mix of English, tool calls, and experimental feedback. So they're quasi-English tokens as we've been talking about, tokenized by the tokenizer. So that's where they came from.
所以你不是在一般地对序列进行分词。
So you're not tokenizing sequences in general.
我们不是一般地对序列进行分词。隐含地,因为如果模型被问到关于 DNA 的问题,里面会有 DNA token。我们不是下载了 dbGaP 或 PDB 或 Swiss-Prot 之类的东西并在序列级别进行分词。这些是推理 token,由模型生成,并经过实验验证。
We're not tokenizing sequences in general. Implicitly, because if the model is asked a question about DNA, there are DNA tokens in there. It's not like we downloaded dbGaP or the PDB or Swiss-Prot or something like that and tokenized at the sequence level. These are reasoning tokens, model-generated, that are experimentally verified.
除此之外,你还有 AlphaFold、核苷酸 Transformer,以及所有进入这个数据堆的测序数据。所以 10 万亿 token 是……
On top of this, you also still have your AlphaFold, your nucleotide transformer, you have all your sequencing data which goes into this. So 10 trillion tokens is...
10 万亿。
10 trillion.
所以是 10 万亿。我在笔记本电脑上看到 10 T……
So 10 trillion. I'm looking at 10 T on my laptop and...
我们认为这个数据量很重要的原因是:预训练语料库通常在 15 到 30 万亿 token 之间。所以,正是在这个规模上,你才会看到这些涌现现象发生。所以,一旦你进入万亿 token 量级,我们有信心这足以让模型开始掌握并展现出涌现能力。
The reason why we think that level of data is important: pre-training corpuses are usually somewhere between 15 and 30 trillion tokens. And so, that's the scale at which you see these emergent things happen. And so, once you're in the trillion token regime, we feel confident that that's enough for the model to start to master and see emergent capabilities.
那么,你是从头开始训练模型,还是使用一些开源模型……
So, are you starting from scratch with your model or you have some open source...
是的,嗯,再次,为了保持雄心勃勃但不过分病态,我们决定不承担预训练,因为你需要做的黑魔法太疯狂了,而且我们已经被赠予了价值约 10 亿美元的算力,以开放权重模型的形式。是的。所以我们从一个已经预训练的开放权重模型开始,我们做出的假设是,该模型已经在互联网和大部分科学文献上进行了预训练。因此,它是对已知知识的良好科学先验,因此是一个很好的大本营,可以在此基础上构建。
Yeah, well, again, in the interest of being ambitiously over-scoped but not pathologically so. We have not decided to take on pre-training as well, just because the black magic that you have to do is insane and we've been gifted something like a billion dollars worth of compute in the form of open-weight models. Yes. So, we start with an open-weight model that has been pre-trained, and the assumption we are making is that the model has been pre-trained on the internet and a large fraction of the scientific literature. Therefore, it's a good scientific prior over what is known and therefore a good base camp to build upon.
所以,这是在已有的数万亿 token 之上又加了 10 万亿……
So, it's 10 trillion on top of the trillions that...
是的。我们大量使用 Nemotron,因为我们与 Nvidia 有合作。我认为该模型的预训练和后训练大约用了 30 万亿 token。
Yeah. And we use Nemotron quite a bit because we have a partnership with Nvidia. And I think there's like 30 trillion tokens that go into the pre- and post-training for that model.
那么,在这些推理 token 的生成过程中,你们实际上也在创建一些相当有用的数据集。你们有没有考虑过独立开源其中一些数据集,即使不包含推理模型本身?这对社区可能仍然很有价值,而且完全不会削弱你们的护城河。
So, in the process of these reasoning tokens, you are also creating what are arguably probably just rather useful datasets themselves. Have you thought about independently releasing some of those datasets open source, even in the absence of the reasoning model? That may still be quite valuable to the community but doesn't actually deteriorate your moat at all.
我们在这个过程中开发的一个东西是一个测试套件,包含大约一千个独特的科学强化学习环境,你可以放入前沿模型、你自己的模型,或者我们的模型。所以,我们几乎肯定会开源其中的一部分。有些基于我们生成的数据,有些是我们为社区整理的数据。因此,会有一个我们整理的基准测试的开源版本,可能还会附带一些训练数据。
So, one of the things that we've developed along the way is a test suite of something like a thousand unique scientific RL environments where you can drop in a frontier model, you can drop in your own model, we drop in our models. So, almost surely we're going to open source a subset of that. Some of it based on data that we've generated, some of it that we have curated for the community to use. So, there will be some open source version of the benchmark that we've assembled as part of that. And there will be, probably some training data that goes along with that.
酷。
Cool.
你们内部有没有实际运行实验室的基准测试?比如衡量……也许不是这个说法。你们有实验性的自动化实验控制吗?
Do you have benchmarks internally that actually operate the lab? Like a benchmark for how well does... Maybe not the way of saying it. Do you have experimental automated experimental controls?
是的,我认为对于每一个……我们从公司成立之初就组建了多学科团队,专注于解决具体的封闭式问题。我们的工作方式一直是:从零开始训练模型,调用每个人都会直接使用的前沿模型,以及我们自己的内部模型,然后进行基准测试。所以,我们做的每件事都有内部基准。不过,这些领域非常具体,对吧?它们不像 Andy 描述的那个基准那么通用和全面,因为它们是我们真正关心的东西、我们想要交付的产品以及我们希望产生影响的地方。但在所有这些领域,我们通常看到,Andy 描述的那种经过科学预训练、能够调用工具的模型,通常会碾压我们与之比较的任何其他模型。
Yes, I mean we have, I think for every one of the... we've put together from the beginning of the company these multidisciplinary teams to work on specific closed-ended problems. And the modus operandi has always been to benchmark training something naively from zero, calling frontier models that everybody would use right out of the box, and our own internal models. So with everything we've done, we do have an internal benchmark. Now, the domains are very specific, right? They're not as general and all-encompassing as the benchmark that Andy was describing because they are the things we really care about and the products that we want to deliver and the places where we want to make a difference. But in all those places, we've typically seen that the scientifically pre-trained model that Andy is describing with access to tool calling typically demolishes anything else that we compare it to.
值得思考的是,我们试图做的事情如何与 LLM 形成互补。如果你考虑一个经过实验验证的推理轨迹,你认为互联网或预训练语料库中有多少这样的轨迹?
It's worth thinking about what we're trying to do, how that is additive with LLMs. If you think about an experimentally verified reasoning trace, how many of those do you think exist on the internet or in the pre-training corpus?
数量级为零?
Order of zero?
数量级为零,没错。与下一个数量级相比,它肯定可以四舍五入为零。所以,我们刚刚看到,向模型展示经过实验验证的推理轨迹带来了不可思议的提升。即使我们在参数上相对于前沿模型处于劣势,仅仅展示一个经过实验验证的推理轨迹,我们就能看到立竿见影的提升。
Order of zero, yeah. It certainly rounds down to zero versus the next order of magnitude. So, we have just seen an incredible lift from showing the model that... Even if we're at a parameter disadvantage relative to the frontier models, just showing it an experimentally verified reasoning trace. You see immediate lift when we do that.
Lila 是一家 Flagship 公司。Flagship 基本上是全球顶尖的生物技术孵化器之一。他们大概有 30 次成功的 IPO。你们是 Generate Biomedicines 的一部分,最近刚刚成功上市。所以,Lila 在生物技术方面非常出色。从历史上看,它非常偏向单一资产的传统生物技术。
Lila is a Flagship company. Flagship is basically one of the biotech incubators in the world. They've had something like 30 successful IPOs, I think. You're part of Generate Biomedicines, which just had a successful IPO very recently. So, Lila is very good at biotech. I would say from history it's very much single asset, traditional biotech.
Flagship 在生物技术方面非常出色。
Flagship is very good at biotech.
Flagship。对对对。我刚才说了什么?Lila。对对对。嗯,是的,Flagship 在生物技术方面非常出色,因为历史上它一直非常专注于单一资产。我想在过去几年里,随着 Generate、Expedition、Velo 的出现,它们开始向更多平台型业务拓展。
Flagship. Yes, yes. What did I just say? Lila. Yes, yes. Well, yeah, Flagship is very good at biotech because historically it's been very focused on single assets. I guess in the last few years with Generate, with Expedition, Velo, there are some branching out into more platform things.
我很好奇一件事:Lila 如何融入更广泛的 Flagship 生态系统?Lila 现在这样有什么具体原因吗?为什么从单一资产转向科学推理?更广泛的互动是什么?特别是,你提到你们有一种 CAR-T 药物已经达到了 IND 阶段。所以,你们显然有生态系统将其转化为实际成果。我很好奇这将会走向何方。
I'm curious about one thing: how does Lila fit into the broader Flagship ecosystem? Was there a specific reason why Lila is now? Why the sort of pivot from single asset into scientific reasoning? And what is the broader interaction? In particular, you mentioned that you had a CAR-T drug which was at the level of IND. So, you clearly have the ecosystem to make that into something. So, I'm curious where this is going.
好问题。我先简单介绍一下 Flagship 的背景,然后谈谈……我们最初都像同一个多能干细胞,但后来发生了分化。Flagship 公司的传统路径是……Generate 的历史是这样的:我在 2018 年早期担任 Generate 的顾问。当时有一个想法,用机器学习进行蛋白质工程。我和 Flagship 的其他几位同事,以及一些外部人员,包括达特茅斯学院的教授 Gabor Gregorian,一起参与了这个项目。我们从 Flagship 获得了种子资金,然后将其剥离出来。我们致力于构建技术,通常的协议是 Flagship 在 A 轮融资中是唯一投资者。然后 B 轮通常是外部资本首次进入的节点。正如你所说,它们最终往往成为资产型公司。Generate 有一种治疗哮喘的单克隆抗体正在进行三期试验,后面还有治疗 COPD 的一期试验。我认为 Flagship 的一些人,尤其是我们的 CEO Jeff Bultzen,他参与创建了许多这样的公司,他注意到需要反复招聘相同的团队。你需要机器学习团队,需要平台团队。所以我认为他看到了这些公司之间的共同 DNA,于是想,不如让一家公司来支持所有这些不同的事情。Lyra 成立的第一年基本上就是 O1 发布的时候。所以我们所有要素都已就位,很明显我们可以创建一个平台来支持一种新的科学模型。早期我们不知道如何变现,商业策略是什么。经过几年,我们对此有了更清晰的认识,但两年前我们的核心信念是“苦涩的教训”是正确的。科学可以成为一个无限的 token 生成器,如果……在运营上,我们与普通 Flagship 公司的不同之处在于,外部投资在 A 轮之前就进来了。同样,A 轮的领投方不是 Flagship。所以我们确实有那个血统,我们来自波士顿。我们确实从一家创建了 110 家初创公司的公司那里学到了很多共享经验。我想是的。他们通常……Generate 是 FL 56 57。实际上是合并的。早期 Lyra 是 96 97。所以 Flagship 有着极其悠久的公司创建历史。
Yeah, great question. Let me do a little Flagship framing and then I'll talk about... So, we all started as the same pluripotent stem cell but there's differentiation that we all take. The traditional path for a Flagship company is... The history of Generate is: I was an early advisor to Generate, a consultant over 2018. There was this idea to use machine learning for protein engineering. Me and a couple other folks at Flagship and some other external folks who came in. Dartmouth professor named Gabor Gregorian was part of this. Got seed money from Flagship to then go and spin that out. We worked on building the technology and then usually the deal is that Flagship is the sole investor during a Series A. And then the Series B is normally the first point at which external capital comes into that. To your point, they often end up being asset-based companies. Generate has a Phase 3 trial for a monoclonal antibody to treat asthma, and a Phase 1 behind that to treat COPD. I think the recognition from some folks at Flagship, especially our CEO Jeff Bultzen, had created, been involved in creating a lot of these companies and he saw, he's like hiring the same team over and over again. You need the ML team, you need the platform team. And so I think he saw shared DNA between all these companies and thought, let's have one company that can essentially support all these different things. Year one of Lyra was essentially when O1 dropped. And so we had all these pieces in place and it just became clear that we could create a platform to support a new kind of scientific model. In the early days, we didn't know how to monetize that, what the commercial strategy was. We've gotten a lot of clarity over that over the years, but the core conviction that we had two years ago was the bitter lesson is correct. Science could be an infinite token generator if... Operationally, the way that we're different from a normal Flagship is outside investment came in before the Series A. Again, the lead of the Series A was not Flagship. So we do have that lineage, we do come from Boston. We do have a lot of the shared learning from a company that has created 110 startups. I think so. They normally... So Generate was FL 56 57. It was actually a merge. In the early days, Lyra was 96 97. And so Flagship has this enormous, long history of creating companies.
那么为什么我听说你们有一种非常有前景的 CAR-T 疗法,比如你说的那种,为什么不直接把它授权出去呢?或者这还在计划中?
So why is it that when I hear you know you have a very promising CAR-T therapy like what you said you had a dizzy like why not just you know partner that out or maybe this is on the horizon or something?
简短的回答是,我们确实在围绕 CAR-T 疗法开展商业合作。其中一些是为了进一步开发,以增强或改变某些特性,比如针对新的适应症。但我们用那一个 CAR-T 项目基本上启动了好几个合作项目。
The short answer is that we are engaging in commercial partnerships around CAR-T therapies for sure. Some of them are further development to increase some of the or change some of the properties, like going after novel indications. But we've used that one CAR-T to essentially launch several partnership programs.
好的,所以这有点像原理验证,但它本身并不完全是药物所需的东西?
Okay, so it's just sort of like the proof of principle but it itself was not you know quite exactly what a drug needed to be or something.
嗯,澄清一下,我们本可以去尝试授权或合作那个具体项目。但我们发现,用它来围绕进一步开发建立多个合作关系会更好。
Well, just to be clear, we could go and try to license or partner that specific thing. We found that it was better to take that and secure several partnerships around further development of it.
你们基本上是在做某种代码开发的事情,其中……
You're basically doing some sort of code development thing where...
这就是虚拟初创公司的理念:一家公司围绕其中一个适应症成立一个虚拟初创公司,他们基本上向我们支付收入以进行进一步开发,我们还有里程碑之类的。是的。
This is the virtual startup idea where a company starts a virtual startup around one of these indications and they essentially pay us revenue to further development, and again we have these milestones and things around it. Yeah.
那么长期来看,既然 Flagship 专门做生物,从未真正涉足材料领域,这对 Flagship 或 Lila 的战略有什么影响?这有关系吗,还是说在这一点上你们已经启动并利用了这一点?
So like long term, since Flagship is specifically bio and is never really branched into materials, how does that sort of weigh on Flagship or Lila's strategy? Does that play into it at all, or is it like at this point you've kind of launched and sort of used that?
从资源角度看,如果你这么快就分化到不同的领域,对吧?我认为使命的广度显然从一开始就超越了生物技术。我们需要招聘的人来自不同的网络。我们不得不购买的仪器来自不同的供应商,而不是 Flagship 通常用的那些。所以我认为这是为什么感觉有些不同的原因之一。但这也是使命的核心,对吧?我们无法在狭窄的领域让这些东西发挥作用。根据定义,我们希望尽可能广泛,因为新兴行为将从那里产生。
A lot of the resource-wise, if you differentiated into something different so quickly, right? I think part of the breadth of the mission clearly was beyond biotech from day one. The people we needed to hire came from different networks. The instruments we had to buy came from different vendors than the Flagship vendors would have usually been. So I think that was part of the reasons why it feels somewhat different. But it's also core to the mission, right? We cannot get these to work on a narrow field. By definition, we want to be as broad as we can possibly be because that's where the emerging behaviors are going to come from.
而且我认为,如果你看看现在在 Lila 工作的人员构成,它会与普通生物技术公司截然不同。我们招聘或竞争的对象,有时甚至能赢过那些考虑 Frontier Lab offer 的人。我们有大量的软件工程和技术人员。我们在 GPU 上的花费对于一家生物技术公司来说是非典型的,我敢说。我认为,如果我们称自己为生物制药公司,我们可能拥有全球前三的 GPU 集群。这确实是我们的 DNA 的一部分,但我们一直有意做出决策,让我们走上我们认为最有前途的轨道。所以这不仅仅是孩子反抗父母之类的事情。我们认为这个论点是正确的,它指向一家不仅对生物技术,而且对材料和化学都非常有价值且重要的公司。
And I think if you looked at the composition of people who work at Lila now, it would look categorically different than what you would expect like a median biotech company to look like. So we hire out of or compete for and sometimes win against people who are considering Frontier Lab offers. We have a heavy software engineering and tech presence. The amount that we spend on GPUs would be atypical for a biotech, I will say. I think that if we called ourselves a biopharma, we probably would have a top three GPU cluster in the world. It's true that that's part of our DNA, but we've been intentional about trying to make decisions that put us on what we think is the most promising trajectory for us. So this isn't just like kids rebelling against their parents or something. We think that the thesis is right and it points towards a very valuable, but also important company for not just biotech, but for materials and chemistry.
好的,这引出了我最后一个问题。哪个更难,材料还是生物学?
Okay, that brings me to what I think is my last question. What's harder, materials or biology?
它们实际上都非常难。有趣的是,对吧?它们……
They're actually very difficult. It's funny, right? They're...
我感觉我们就要像蜘蛛侠 meme 那样了……
I feel like we're about to do the Spider-Man meme in like...
当我参与第一波 AI 用于小分子药物发现时,我的意思是,原子层面的步幅和生成,你知道,计算机模拟。所以我认为最难的是小分子。它拥有化学的所有难点,比如知识推理或合成。然后它还有关于生物学、不良反应和免疫反应推理的所有难点。
When I was around for the first merry-go-round of AI for the small molecule drug discovery, I mean, the atom-wise stride and generate, you know, the in silico. So I think the hardest is the thing that is a small molecule. It has all the difficulties of chemistry, of knowledge reasoning or synthesis. And then it has all the difficulties of the reasoning about biology and adverse effects and immune response.
是的,但反过来说,我们的工具箱里有很多可以从生物学借鉴的技巧,对吧?所以它更难,但你也有……
Yes, but the counterpoint being that we have so many tricks in our toolkit which you can borrow from biology, right? So it's harder, but you also have...
嗯,我认为材料更难。它们有很好的模拟器,这是我们在生物学中没有的。
Well, I think materials are harder. So they have the benefit of great simulators that we don't have in bio.
在材料科学中,你没有生物学中那种成熟的高通量自动化。对我来说,材料作为一个学科很有趣,因为没有像中心法则那样的统一原则。材料意味着很多不同的东西。我实际上仍然不太理解当我们说材料科学时,统一原则到底是什么。然后商业动态完全不同。就像 CAR-T 一样,我们知道如果想直接变现该怎么做。对于材料,有供应链、设备,你在实验室做的测试只能部分预测材料在使用寿命中的表现。而且数学更难。
In material science you don't have the mature high-throughput automation that you have in biology. For me, materials as a subject is interesting because there's not a unifying principle like the central dogma. Materials means lots of different things. I actually still don't quite understand the unifying principle when we say material science, what exactly that means. And then the commercial dynamics are completely different. Like again with CAR-T, we know if we wanted to how to monetize that directly. With materials, there's a supply chain, there are devices, the testing that you do in the lab is only partially predictive of the lifetime of how that material will be used. And the math is harder.
我的意思是,就供应链而言,两者都很重要。也许你可以用一些产品验证和确认来替代临床试验。比如“认证”这个术语。所以两者有直接的类比,也都有难点。
I mean, in terms of supply chains, they still matter for both. Maybe you replace clinical trials with some product validation and verification. Like qualification is the term. So there are direct analogies and there are hard parts for both of them.
嗯,经济学非常不同。如果你通过了临床试验,你就赚钱了。那个东西很有价值。而制造成本很少成为障碍。
Well, the economics are very different. If you pass a clinical trial, you make money. That thing is valuable. And how much it costs to make it is very rarely the blocking element.
在生物学中承保资产比在材料中容易得多。你看。
Much easier to underwrite an asset in biology than it is in materials. See.
你们知道一家制造超导体的公司名字吗?这个问题总是出现。你们在做……是的,我们关心磁铁,我们关心超导体。它们是非常酷的科学。你们知道一家制造超导体的公司名字吗?没人知道。这些东西超级重要。
Do you guys know the name of a company that makes a superconductor? This always comes up. Are you guys doing... Yeah, we care about magnets, we care about superconductors. They're really cool science. Do you folks know the name of a company that makes super... No one knows. Like these things are super important.
它们用于 MRI。
That they're used in MRIs.
正是。
Exactly.
那是我知道的唯一商业应用。
That's the only commercial application I know.
但事实证明,对吧?这些东西当你成功时,你有点……当你制造出一种很酷的材料,它很有用,你是一家默默无闻的公司,制造这个东西,很成功,现金流很好,但你无法突破,你知道的。
But it turns out, right? Like these things when you succeed, you kind of are... when you make a cool material that does something, you're kind of a nameless company that makes this thing and is successful and has good cash flows, but you don't get to break sort of you know.
人人都知道大型制药公司,但除此之外,你还有像 3M 这样的公司,对吧?这种……
Everybody knows a big pharma, but other than you know, you've got your 3M's, right? This sort of...
是的,而且大多数大型材料公司都是闭门造车。大多数商业合作看起来就像是让他们告诉你什么才是重要的问题。而且开放创新的生态系统相对较少。材料领域有一些东西被公认为有价值,但我认为这与生命科学非常不同。
Yeah, and like most of the big material companies are behind closed doors. Like most of commercial engagements look like getting them to tell you what the important problem is. And there's like less of an open innovation ecosystem. There's a couple things in materials that are obviously recognized to be valuable, but like it's just I think very different than than life science.
还有一件事我们没怎么谈到,我想提一下。我认为在化学,尤其是材料领域,政府资助的研究是一个重要驱动力。就像政府认为除了资助 NIH 进行早期开放科学、假设驱动的科学之外,他们不需要做药物发现一样。政府和国家安防以独特的方式推动材料创新。从我们与英国政府的合作中就能看到这一点。我们有合作伙伴关系,与美国政府合作,获得奖项,参与开发材料和技术,这是推动创新的生态系统中一个不同的部分。这也不同。
And maybe one of the last things that we haven't touched upon a lot and I want to flag out. I think in chemistry and especially materials, government sponsored research is a big driver. So in the same way that you know, the government doesn't feel they need to do drug discovery other than funding NIH for early stage open science, hypothesis-driven science. You know, the government and national security drive materials innovations in ways that are unique. And you see this in the way we engage with the British government. We have partnerships. We work with the US government. We have awards. We participate in sort of developing materials and technologies, which is a different part of the ecosystem that drives innovation. That's also different.
是的,确实如此。你们在参与 Mission Genesis 吗?
Yeah, definitely. Are you guys working in Mission Genesis?
我们是列名的合作伙伴之一。我们与许多国家实验室一直有合作关系,所以我们一直在做这方面的工作。
We were one of the named partners. We've had an ongoing relationship with a lot of the national labs, and so we have been working on that.
通过这个发送 25?
With this send 25?
是的。
Yeah.
上周的 Genesis 灯塔提案。
Genesis lighthouse proposals last week.
我在开玩笑。我在开玩笑。我在开玩笑。他离开学术界时以为 Great Writing 已经过去了,结果还得写……
I was joking. I was joking. I was joking. He left academia thinking Great Writing was behind him only to have to write to...
是的。
Yeah.
他们谈到两万。
They speak about 20,000.
是的。
Yeah.
我们喜欢问所有嘉宾的问题是:如果你能凭一己之力消除你领域中的一个瓶颈,你可以自行定义领域,那会是什么瓶颈?
The question that we like to ask all of our guests is if you could remove a bottleneck in your domain by fiat, and you can define domain by fiat, what would that bottleneck be?
嗯。
Mhm.
对我来说,我要回到过去做基于物理模拟的 Rafa。我会说是从模拟到现实。我是说,对于来自基于物理世界的人来说,模拟到现实有一个……
To me I'm going to go to old timing Rafa that was doing physics based simulation. I would say the sim to real. I mean sim to real for the people that come from sort of the physics based world. I mean the sim to real having like an...
你能解释一下这是什么意思吗?
Can you explain what that means?
我的意思是?这些人通常是在机器人学的语境下说的,你的 3D 空间虚拟模拟可以训练机器人在物理空间中移动,但存在一个差距,他们称之为模拟到现实的差距。对我们做基于物理模拟的人来说,我们做粘性物质的分子模拟,做硬物质的电子结构模拟,它们还行,但不够具有预测性。这就是为什么,如果不是因为这个,我们可能就不必建造材料自驱动实验室了,因为我们就能直接预测了。所以我认为我们知道物理规律,但它并不完全具有预测性,也就是说,我们在物理上训练的模型也无法缩小差距,因为它们仍然缺失这些,或者说它们是在不够好的近似上训练的。所以我认为我们在材料 AI 领域追求了十年的事情是:如果我们用计算数据训练,能否回答现实世界的实验问题?如果我能回到过去并消除瓶颈,那将是底层模拟的准确性,我们一直依赖它进行训练。
What I mean? So these people have typically meant it in the context of robotics where your virtual simulations in 3D spaces kind of allow you to train robots that will move in physical spaces, but there's a gap and they call it the sim to real gap. For us in physics based simulations is that you know, we do molecular simulations of gooey stuff. We do electronic structure simulations of hard stuff and they're okay, but they're not predictive enough. So and this is the reason why you know, if I if it wasn't for that, maybe we wouldn't have had to make a self-driving lab for materials because we would have been able to just predict. So I think we know there's physics, but it doesn't quite go the way to being predictive, meaning that the models that we train on physics cannot possibly close the gap either because they're still missing these rather they're trained on on approximations that are just not good enough. So I think the thing we've been chasing for a decade in AI for materials has been sort of if we train on computational data, can we answer real world experimental questions? And that would have been the place where if I if I get to go also back in time in addition to taking the bottleneck out, it would be the the accuracy of the underlying simulation that we've been training on all the time.
所以这有点像 Heather Kulik 说的,材料领域没有 AlphaFold。
So this is sort of like Heather Kulik said, there is no AlphaFold for materials.
有趣的是,AlphaFold 是在实验数据上训练的,所以这是不同的,我的意思是,这有点意思。
Well, the funny thing is AlphaFold was trained on experiments, so it's a different I mean, that's kind of funny.
是的。
Yeah.
她和我都来自做基于物理模拟的背景,她指出一个完全没有模拟的东西,这恰恰反映了同一个根本问题:Meta 已经产生了数千万、数亿的训练数据点,但它们都是虚拟模拟,对我们实际想做的事情来说不够有力。
She and I we both come from doing physics-based simulations, and the fact that she called out something that had no simulations in it whatsoever is kind of a meeting the same underlying issue, which is like all these, you know, Meta has produced tens of millions hundreds of millions of training data points, but they're all virtual simulations that just don't carry enough water for the thing we actually want to do.
这可能会是一个无聊且显而易见的问题,但有一个指标用来衡量训练运行的效率,叫做平均浮点运算利用率(MFU)。GPU 有一个标称的峰值浮点吞吐量,即在最佳情况下,做你不关心的计算时,单位时间内能完成多少浮点运算。MFU 总是峰值理论浮点运算的一个很小比例。对于强化学习,它总是在 5% 到 6% 左右。换句话说,我们只得到了我们付费的 GPU 计算能力的 5%。所以,如果我能凭一己之力挥动魔杖,让我们的堆栈达到 100% 的平均浮点运算利用率,我会这么做,因为我们不仅能更快得到答案,还能购买更少的 GPU,并将资金重新部署到实验室或其他地方。
This is going to be like a boring and obvious one, but like there's a metric that you use to track how efficient your training runs are. It's called mean flop utilization or MFU. So, the GPU comes with an advertised like peak flop throughput, which is under the best situation, doing a calculation that you don't actually care about, how many floating-point operations can you do per unit of time. MFU is always a very small fraction of peak theoretical flops. And for reinforcement learning, it's always somewhere like around 5 to like 6%. So, said differently, that means that we're getting like 5% of the actual GPU computing power that we're paying for. So, if I could by fiat wave a wand and make our stack perform at like 100% mean flop utilization, I would do that because we would one get to the answer faster, but then also be able to buy fewer GPUs and redeploy that capital to the to the lab or something like that.
有趣的是,因为你的 rollout 不是受实验室限制吗?
Interesting though because your rollouts, aren't they constrained by the lab?
是的,但当我们训练一个大模型时,所有那些数据。强化学习训练流程非常复杂。所以,一种规模化做法是让模型不断进行 rollout,等待足够多的轨迹堆积起来,然后反向传播到模型中。另一种做法是将其分解,并行训练一批专家模型,它们要么生成数据,要么自己训练,然后将其蒸馏回中心模型。第二种方式更高效。因为所有这些事情发生在不同的时间尺度上,当你拥有 10 万亿个 token 并希望尽可能高效地通过模型处理它们时,你仍然需要在其上进行一些强化学习。所以,如果我们能获得我们支付的所有浮点运算,我会宣布这一点。
They are, but when we train a big model, like all of that data. So, there's RL training pipelines are very complicated. So, like one way to think about how you would do this at scale is just to have the model doing rollouts left and right waiting for enough trajectories to pile up and then back propagating that into the model. A different way to do that would be to factorize that, have a bunch of expert models that are trained in parallel that are either generating data or being trained themselves, and then you distill that back into the central model. And second way is the most efficient the more efficient way to do that. So, cuz all those things are happening at different time scale and so it's that big when you have the 10 trillion tokens and you want to push them through the model as efficiently as possible, you're still going to be doing some reinforcement learning on top of that. So, like if you could get all the flops that we're paying for, I would I would buy fiat declare that.
酷。好吧,在我们结束之前,你有什么想对观众说的吗?
Cool. Well, yeah, before we end is there anything you want to leave the audience with?
让我说说我们为什么在这里。我们现在在旧金山有一个办公室,位于旧金山市中心的 Fremont 街 181 号。目前大约有 20 到 30 人在那里工作,但我们正在积极扩张。我们正在从技术栈的各个领域招聘人才。所以,后训练方面显然在积极招聘。在领域 AI 方面,比如生命科学和材料科学,我们也在招聘。
Let me say like why we're here. So, we have an office in San Francisco now. It's 181 Fremont Street in downtown San Francisco. There's currently 20-ish 30-ish people who sit there, but we are looking to expand that aggressively. We're looking to pull from sort of all areas of the stack. So, both like post trading obviously aggressively hiring for that. Folks who've been working in like domain AI like life sciences and material sciences we're also hiring for that.
目前这里没有湿实验室,所以全是计算工作。如果今天大家听到的这些内容中有任何感兴趣的,欢迎给我或 Raphael 发消息。
No wet lab here currently, so it's all computational work. If any of this stuff that people have heard about today sounds interesting, feel free to shoot either me or Raphael a message.
感谢你来做客。这真是一场非常非常精彩的对话。非常感谢。
Thank you for being here. It's been a really, really fascinating conversation. Appreciate it.
感谢你的邀请。
Thank you for having us.
嗯。
Yeah.