ESMC: World Modeling for Protein Design
打开互动全文版(中英对照 + 朗读 + 问答)→Alex Rives 讨论 ESMC 的可编程生物学世界建模方法,设计蛋白质结合剂和抗体,以及蛋白质语言模型背后的缩放定律。
Alex Rives discusses ESMC's world modeling approach to programmable biology, designing protein binders and antibodies, and the scaling laws behind protein language models.
所以 ESMC 也在探索可编程生物学,但我会说方式非常不同。它从一种世界模型的角度出发,基本思路是你有一个预测模型,然后你在这个世界模型中搜索,找到满足你设计标准的蛋白质分子。我们利用这一点已经成功设计了许多蛋白质结合物。但最令人兴奋的是,我们还能用它来设计抗体,比如单链可变区片段(scFvs)。
So ESMC is also approaching programmable biology, but I would say in a very different way. It's approaching it from this kind of world modeling perspective where the idea is basically you have a predictive model and you know you're going to search the world model to find protein molecules that satisfy kind of whatever design criteria that you have. So we've been able to use this to actually now go and design many protein binders. But I think sort of most excitingly, we've been able to use this to actually design antibodies, SCFVS.
你好,欢迎收听 Latent Space AI for Science 播客。我是 R.J. Haneki,Muromix 的 CTO。
Hello, welcome to the latent space AI for science podcast. I'm R.J. Haneki, CTO of Muromix.
是的。今天我是 Brandon。很高兴邀请到 Biohub 的科学主管 Alex Rives。你能快速介绍一下自己吗?
Yeah. And, uh, I'm Brandon today. It's a pleasure to have Alex Rives, uh, head of science at Biohub. Yeah. Would you like to introduce yourself real quick?
是的。谢谢邀请,很高兴来到这里。我是 Biohub 的科学主管,是一名计算机科学家,研究 AI 在生物学中的应用,我的很多工作都集中在生物学语言模型上。
Yeah. Yeah. Thank you for having me here. It's great to be here. Um, I'm head of science at Biohub. I'm a computer scientist uh and I work on AI for biology and a lot of my work has been on language models for biology.
在这期播客发布时,你应该已经推出了几个令人兴奋的新模型。回顾这些模型,我不禁觉得你可能是目前蛋白质生物学领域最信奉“苦涩教训”的人。你能解释一下这对生物学意味着什么,以及你为什么如此坚定并热衷于这条道路吗?
By the time this podcast is released, you will have put out several new exciting interesting models. Going over them, I couldn't help but have the kind of thought that you might be the most bitter lesson person in protein biology right now. Can you give a little context about what that means for biology and you know why you're so committed and excited to this route?
好吧,我接受这个说法。我相信缩放定律。我从 2018 年夏天就开始研究这个了。我的团队在 Metaphair 时训练了第一个用于蛋白质生物学的 Transformer 语言模型。我一直认为,当你训练模型预测进化产生的下一个词元时,生物信息会涌现。我们团队多年来一直在探索这个想法,我们看到了缩放曲线,随着每一代模型规模提升一个数量级,新的能力不断涌现。
Well, I'll take that. Um, I believe in scaling laws. So, you know, I guess I've been working on this for, you know, since since the summer of 2018. Um, and so my team when we were at Metaphair trained uh really the first transformer language model for protein biology. And so I guess you know I I've always thought that there would be kind of emergence of biological information as you train a model to predict the next token that evolution creates. So our team has really explored that idea over a number of different years and we've really kind of I think seen the scaling curve and really seen as we have have increased models by an order of magnitude kind of in each generation that you know there's this emergence of new capabilities.
是的。你说能力随着代际缩放而涌现。你研究这个大概有 8 年了。但并非一开始就顺利,对吧?有一些迹象表明缩放可能有效。我们即将看到一些新结果,我认为你以前所未有的方式清晰地证明了这一假设。但你对这个方向有很强的信念,而我未必会如此确信它能以同样的方式奏效。蛋白质语言和自然语言不同。有相似之处,但如果你用正常温度采样一个自然语言 Transformer,你会得到胡言乱语;而用无限温度采样蛋白质语言模型,你会得到有效的蛋白质,尽管可能不是有趣的蛋白质。这毕竟是不同的领域。那么,你认为蛋白质有什么特别之处,使得这种方法同样有效?
Yeah. So you've been you say emergence of capabilities scaling over generations. You've been working at this as you said for I guess it would be 8 years now or something like that. It didn't always work that way right like there was signs that scaling might work. You know we'll be getting to some new results where I think really you've kind of clearly demonstrated this hypothesis in a way that hasn't happened before. But you seem to have like a strong commitment to this in a way that I'm not necessarily sure I would have been so convicted that it would work in the same way. I mean proteins are not the protein language is not the same thing as natural language. There are similarities but if you start sampling a transformer at you know a normal language transformer at temperature you're going to get gibberish. you sample a protein language model at infinite temperature, you're going to get something which is a valid protein if not a not interesting protein despite the fact that is a different domain for a different reason. I'm not necessarily sure that I would I primarily assume the natural language model insight would transfer over. So what is specifically about proteins that you thought was special or you you know that would make this also valid?
这是一个非常有趣的问题,也是当前 AI 领域更深层次的问题。AI 是一门经验科学,我们不一定有理论指导,但有很强的经验证据支持缩放。我受到启发的点是:想想进化,想想我们拥有的蛋白质数据——数据库中有数十亿条蛋白质序列。这些序列包含模式,早在几十年前人们就知道,蛋白质家族序列中存在模式,这是因为进化在约束下运作。比如,一个蛋白质序列折叠成三维结构,序列中的两个残基可能在折叠结构中接触。进化不能独立地选择它们:在一个位置做出选择后,必须在另一个位置做出兼容的选择。从基因测序开始,人们就在不同生物的同源蛋白质中看到这些反映基础生物学的模式。ESM 背后的想法是:如果我们将这一原理应用于整个进化过程,跨越生命产生的所有蛋白质多样性,让语言模型预测进化在所有生物背景下会选择哪些氨基酸,那么整个图景中蕴含着关于蛋白质基础生物学的巨大信息。这就是最初的灵感:模型需要预测下一个词元,实际上我们使用掩码语言建模训练这些模型,它们预测序列中被掩码的部分,从而学习那些约束进化选择的潜在规则。
Yeah, I mean it's a really interesting question. I think kind of a deep question across AI right now more broadly and you know I think you know what's what's so interesting is AI right now is is such an empirical science and so we don't have you know theory that can always guide us in these things but we have this really strong empirical evidence of scaling the thing that I was motivated by is you know if you think about evolution and you know you think about the data that we we have around proteins we have databases that have billions of protein sequences. And you know, those those sequences contain patterns and you know it had had been long been known so that you know this is going back you know decades kind of before you know we started working on this with language models but that there are patterns the sequences of protein families that come there because of the constraints that evolution is operating under. So you can think about, you know, like a um a protein sequence that folds into a three-dimensional structure in space. And you can, you know, imagine that there are two residues or amino acids that are in this sequence that might be in contact in that folded structure. And so evolution isn't free to choose those independently from each other. If it makes a choice at at one position, it kind of has to make another choice that's going to be compatible at the next position. So going back, you know, all the way to the beginning of gene sequencing when people first began to be able to to look at this and kind of look at different related, you know, the same protein and related organisms, you could start to see these kind of patterns that are reflecting the fundamental underlying biology. So the idea behind ESM, kind of the thinking behind ESM was, okay, what if you were to apply this principle of across all of evolution, across kind of the vast diversity of proteins that have been generated across all of life and, you know, basically have a language model kind of predict the amino acids that evolution will choose to place in proteins across all of those biological contexts. So you can think that there's just this this kind of like incredible amount of information in that total picture about the underlying biology of proteins. And so that was really the idea that sparked this is is you know as as a model is having to predict the next token and actually we train these models with mass language modeling. So they're predicting kind of tokens that are masked out of various parts of the sequence that it would have to learn something about those kind of underlying constraints that are shaping which tokens evolution can choose.
是的。那么回顾一下历史,你们刚刚发布了 Evolutionary Scale Modeling Cambrian,对吧?是这个名称吗?是的。这大概是系列中的第四或第五个模型。如果追溯到它们被称为 ESM 之前,可能更多。
Yeah. So maybe for a bit of history um so you know you have you you just released um evolutionary scale modeling Cambrian, right? Is that what it's called? Yeah. And this is like the maybe fourth or fifth in a series of models. I think maybe even more if you go back before they were called ESM.
它们从一开始就叫 ESM。我们有过不同模型的分支。这个可以说是第四代模型。实际上,我们一年前就训练了这个模型。现在我们在 Biohub,首次以 MIT 许可证完全开源这个模型,我们非常兴奋。但最大的新进展是,我们真正构建了一个蛋白质生物学的世界模型。其基础就是 ESMC。
Well, they were called ESM from the start. Yeah. We had sort of various branches of the different models. Yeah. So, so this one I would say is is kind of a a fourth generation model. Um it's actually a model that we trained a little over a year ago. Now that we're at Biohub, we're um we're we're open sourcing this this model fully under MIT license for the first time. So, we're really excited to do that. But kind of the the big thing that is new here is that we've really kind of built a world model of protein biology. So the foundation of that is ESMC.
但你知道,利用 ESM 的表示,我们现在构建了一个结构预测模型。这是下一代 ESM Fold 模型。我们还使用了机械可解释性和稀疏编码技术,真正深入审视语言模型的表示空间,并提取出模型实际用于表示蛋白质生物学的底层特征。综合所有这些,我们能够预测蛋白质结构,预测蛋白质构成的基本特征,从而建立跨进化关联。我们可以将这个模型反转来设计蛋白质。我们利用它创建了一幅全面的蛋白质生物学图景。我们整合了全球最大的蛋白质序列数据库,总计 68 亿条非冗余蛋白质。其中,我们已经解析了 11 亿条蛋白质的预测结构。我们还计算了所有这些蛋白质的特征,以便在整个进化和蛋白质生物学中建立这些关联。
But you know, using the representations of ESM, we've now built a structure prediction model. And this is the next generation ESM Fold model. And then we've also used the techniques of mechanistic interpretability and sparse coding to really start to look deeply into the representation space of the language model and be able to pull out the underlying features that the model actually uses to represent protein biology. So bringing all of this together, we're able to make predictions for protein structure, predictions about the underlying features that proteins are made out of that allows us to build linkages across evolution. We're able to take this model and invert it to design proteins. And we've used this to create a comprehensive picture of protein biology. So we put together all the world's largest protein sequence databases. That amounts to 6.8 billion non-redundant proteins. And then we've resolved predicted structures for 1.1 billion of those. And we've also computed features across all of those so that we can make these linkages basically all across evolution and protein biology.
68 亿,其中你解析了 12 亿的结构?是……
6.8 billion, of which you've resolved structure for 1.2? Is that...
11 亿。
1.1.
11 亿。那其他的呢?
1.1. So what about the others?
基本上,我们对该数据库进行了聚类,以 70% 的序列一致性为标准。所以实际上我们为所有蛋白质都解析了结构,因为每个聚类都有一个聚类中心。我们预测那里的结构,然后可以预期其他蛋白质会有相似的模板结构。会有微小变化,但它们具有相同的折叠。
So basically what we did is we took that database and we clustered it at 70% sequence identity. So it's really resolving structures for everything in the sense that for each cluster we have a cluster center. We're predicting the structure there and then we can expect that the other proteins are going to have a similar template structure. There will be small variations but they have the same fold.
大约 12 亿个聚类覆盖了 68 亿条序列。
1.2 billion or so clusters that are covering the 6.8 billion.
是的。
Yeah.
好的,有意思。既然我们在讨论 Scaling,你怎么知道这个数字是合适的?比如,你怎么知道专注于这 11 亿个聚类,并且这个分辨率对这个模型来说是合适的?
Okay. Interesting. And maybe since we're talking about scaling, how do you know that this is the right number? Like how do you know that focusing on these 1.1 billion and that's the right resolution for this model?
我们选择它们是为了真正覆盖整个空间。所以我认为关于这个数据库,可以说它是有史以来最全面的蛋白质结构和功能图景。它为我们对蛋白质结构多样性的认知增加了数亿个结构,同时也创建了这个特征空间,使我们能够发现蛋白质在进化中的这些关联。因此,我们可以看到进化中涌现出非常有趣的主题,例如连接基因编辑系统,这些系统在序列上相距甚远,但它们共享一些底层功能模式和结构同源性,模型能够将它们汇聚起来并找到这些联系。
Well, we've chosen them so that they really cover that entire space. So I think what I can say about this database is it's really the most comprehensive picture of protein structure and function that's been created. It's adding hundreds of millions of structures to our knowledge of the diversity of protein structure, and it's also creating this feature space that allows us to find these linkages between proteins across evolution. So we can see really interesting themes emerging across evolution, linking for example gene editing systems which are very far apart in sequence but they share some underlying functional patterns, structural homology that the model's able to bring together and find those connections.
现在我们讨论的是机械可解释性部分。所以,如果我理解正确,你们使用稀疏自编码器和其他技术来理解:当我用蛋白质激活网络时,我看到的输出模式是什么,它们之间如何关联?如果我理解正确,你们有一些序列,它们基于实际序列不相关或仅部分相关,但在行为上它们有相似的行为,因此它们激活了相似的网络。这是对你刚才所说的总结吗?
Now we're talking about the mechanistic interpretability part. So you have, if I understand correctly, you use sparse autoencoders and other techniques to understand: when I activate the network using a protein, what are the patterns of outputs that I'm seeing and how do they relate to each other? If I understand correctly, you have these sequences that are unrelated or only partly related based on the actual sequence, but in terms of behavior they have similar behavior and therefore they are activating similar networks. Is that kind of the summary of what you just said?
是的。基本上,我们在 ESM 模型家族的所有不同层上训练了稀疏自编码器。这个家族实际上有三个模型:一个 3 亿参数模型、一个 6 亿参数模型和一个 60 亿参数模型。然后我们对那个 60 亿参数模型(真正最先进的蛋白质语言模型)的特征空间进行了非常深入的分析。我们发现,非常有趣的是,出现了一个特征层次结构。它的有趣之处在于,它确实对应于几十年来、一个世纪的生物学实验所发展出的还原论生物学图景。但很酷的是,这是在没有任何先验知识的情况下涌现的。它是语言模型自己学到的。所以 SAE 的有趣之处在于,它们只是揭示了表示空间的内在结构。这个模型是在蛋白质序列上训练的,它只是被训练来预测进化会选择哪些氨基酸。然后不知何故,这导致了这种非常有序的特征空间的涌现,它具有层次结构,你可以在其中看到从基本的生化特性和蛋白质的基本结构构件,到非常大的功能主题,这些抽象概念与人类对蛋白质功能的理解相联系。
Yeah. So basically what we've done is we've trained sparse autoencoders across all the different layers of the ESM model family. So there are actually three models in that family: a 300 million parameter model, a 600 million parameter model, and a 6 billion parameter model. And then we've done a very deep analysis of the feature space of that 6 billion parameter model, which is really the state-of-the-art protein language model. So what we find, what's really interesting, is there's this hierarchy of features that emerges. What's really interesting about it is it really corresponds to the reductive picture of biology that has been developed over many decades, a century of biological experiments. But what's so cool is this is emerging without any prior knowledge. It's been learned by the language model. So the interesting thing about SAEs is they're really just revealing the intrinsic structure of the representation space. So this model's been trained on protein sequences. It's been trained just to predict the amino acids that evolution will choose. And then somehow this is leading to the emergence of this very ordered feature space that has this hierarchical structure where you can really see everything from the basic biochemical properties and the basic structural building blocks of proteins to these very large functional themes, these abstract concepts that connect to the human picture of protein function.
你是否有一个假设或感觉,为什么序列本身之间存在关系,即使它们被移位、切割和以不同方式重组?我可以想象这可能是有效的,因为蛋白质本身也是分层的。所以也许层次结构在移动,但相同的序列,但功能单元,我想那些有相关的结构。这里的假设是什么?
And do you have a hypothesis or feel for why if there are relationships between the sequences themselves even if they're like shifted and cut up and recombined in different ways? I can imagine that might work because proteins are kind of hierarchical in their nature as well. So maybe the hierarchy moves around but the same sequence but the functional units, I guess those have related structures. What is the hypothesis here?
这是一个非常有趣的问题。我想我可以推测一下。我不认为我们完全理解这一点。但让我举一个具体的例子。亲核肘部是一个核心功能基序,人们认为它可能在进化中独立出现,在不同时间、不同蛋白质家族中。但它有这个非常清晰的结构基序,你可以在晶体结构中看到。我们基本上发现,模型对这个亲核肘部有一个单一的特征,并且它在这些进化上多样化的家族中被激活。真正完全不同的结构拓扑,可能完全独立进化的蛋白质。但模型使用这一个特征来表示它。那么它为什么会这样做?我认为这是一个非常有趣的问题。
I mean it's a really interesting question. I think I can speculate about it. I don't think we completely understand this. But let me give a concrete example. So the nucleophilic elbow is this core functional motif that people have thought maybe this has emerged independently in evolution at different times in different protein families. But it has this very clear structural motif that you can see in a crystal structure. What we found basically is that the model has a single feature for this nucleophilic elbow, and it's activating across these evolutionarily diverse families. Really completely different structural topologies, proteins that probably evolved entirely independently from each other. But the model is using this one feature to represent that. So why does it do that? I think it's a really interesting question.
我认为一个答案是压缩的概念,以及模型需要发展某种潜在的隐变量来帮助解决这个序列预测任务。因为任何一个氨基酸的选择都与序列中所有其他氨基酸的选择完全纠缠在一起。所以预测蛋白质中氨基酸应该放在哪里是一个非常复杂的任务。但要真正做好这件事,模型必须开始拥有这些代表生物学的隐变量,让它能够观察一个蛋白质并判断在不同上下文中应该出现哪些氨基酸。这就是直觉。我会把它类比于语言建模。我深受 Zellig Harris 1954 年的论文《分布结构》影响。我认为那篇论文也影响了许多语言建模领域的人。它聚焦于语言,并阐述了这样一个观点:一个词出现的上下文集合是由该词的意义决定的。所以 Zellig Harris 设想,当你观察词在哪些上下文集合中出现的统计模式时,你就能推导出语言的意义。你会拥有一种反映语言潜在意义的统计结构。对我来说,这是解释为什么在互联网文本上训练的语言模型会学到关于意义的东西的最有说服力的理由之一。它会学到更深层、更根本的东西。所以你可以把同样的道理应用到生物学上:一个氨基酸可能出现的上下文是由蛋白质的结构、功能及其生物学角色决定的——这是非常复杂的现象,包括蛋白质的内在生物学特性以及它与其他蛋白质的关系、功能和进化等。但这些决定了上下文集合,因此你可以想象,氨基酸使用中的那些统计模式直接反映了这些潜在的隐变量,而模型将会学到关于这些隐变量的东西。
I mean, I think one answer is sort of the idea of compression and the idea that the model needs to have some kind of underlying latent variables that it develops to help solve this sequence prediction task. Because the choice of any amino acid is completely entangled with the choice of all the other amino acids in the sequence. So this is a very complex task to try to predict what amino acids should be where in a protein. But to really do this well, the model would start to have to have these hidden variables that are representing the biology that allow it to look at a protein and say, okay, what amino acids should be there in all these different contexts? So that's sort of the intuition. I would draw the parallel to language modeling. I was very influenced by a paper by Zellig Harris called "Distributional Structure" from 1954. I think that paper influenced a lot of people in the language modeling field as well. It focuses on language and articulates this idea that the set of contexts in which a word appears are determined by the meaning of that word. So what Zellig Harris imagined is that as you looked at the statistical patterns of what words appear in what context sets, you would be able to derive the meaning of language. You would have this statistical structure that would mirror the underlying meaning of language. For me, that's one of the most convincing explanations for why a language model trained on the text of the internet is going to learn something about meaning. It's going to learn something deeper and more fundamental. So you can think about the same thing in biology, where the contexts in which an amino acid can occur are determined by the structure, the function of the protein, its biological roles—very complex phenomena, both the intrinsic biology of the protein and its relation to all of the other proteins, function, evolution, and so on. But those are what determine the context sets, so you would imagine that those statistical patterns in the use of amino acids directly reflect those underlying hidden variables, and the model is going to learn something about those hidden variables.
我完全相信这个观点,它看起来很有道理。实际上我确实相信这个方向,但有很多方面让我觉得它可能行不通。其中之一是数据可用性。我们通常有什么类型的序列数据?我认为 ESMC 相比之前的模型有一些新的数据源,这可能会有帮助。但通常我们可用的序列类型对医学、人类生物学或疾病生物学等特定需求有很强的偏向性。所以并不是说随便拿一个朴素的数据集就能得到有趣的缩放定律。我很好奇 ESMC 的突破具体是什么。也许我们可以回顾一下 ESMC 之前的一些 ESM 前辈,它们有哪些优势,以及 ESMC 克服了哪些局限性,还有那些发展。
I definitely buy that; it seems plausible. I actually do really believe in this direction, but there are a lot of ways I think about this where maybe it wouldn't work. One of them is data availability. What type of sequence data do we normally have? I think that ESMC in particular has some new data sources compared to previous models, which might be helpful. But often the type of sequences we have available have a very strong bias towards certain specific needs for medicine or human biology or disease biology. So it's not necessarily that if you take a naive dataset you're going to get an interesting scaling law. So I'm curious about what in particular was the breakthrough in ESMC. Maybe we can go back a bit and talk about some of the other ESM predecessors which got here before ESMC, and how they had their strengths but also limitations that ESMC overcame, and what the developments were.
是的。我承认我更擅长 Scaling(规模扩张)。我是 Scaling 的粉丝。我确实认为增加数据、增加参数并进行压缩会带来更强大的模型。但同样正确的是,数据的底层结构和分布非常关键,你完全正确。所以有些数据集在学习这些通用原理方面会比其他的更有价值。但我认为这与许多关于数据收集的生物学直觉相悖。通常当你考虑需要什么数据时,你是在试图回答一个非常具体的科学假设。你想要一个控制得很好的实验,你确实需要多个重复,这是非常聚焦的。所以我认为思维方式的转变是:如果你想学习蛋白质的通用表示,你真正想要的是让氨基酸在尽可能多的进化上下文中出现。这才是你真正想要的。这就是我对数据的看法。我认为 ESM2(上一代模型)和 ESMC(新一代模型)之间的变化在于,它们规模大致相同——相同的算力规模,相同的参数规模。ESM2 用了很多算力,但 ESMC 用了更多。但不仅仅是算力,数据才是关键。当我们训练 ESM2 时,我们观察到两件事。第一,随着参数和算力的增加,我们看到了改进。所以我们有一个十亿参数规模的模型,一个百亿参数规模的模型,更大的模型比小的更好。但如果你看参数规模与能力(能力指表示保真度,即捕捉蛋白质结构的程度)的图表,你会发现 ESM2 存在收益递减。ESM2 是在 UniRef 上训练的。而对于 ESMC,我们加入了宏基因组学,所以我们在训练数据中增加了数十亿条序列。
Yeah. Well, I'll admit that I am better at scaling. I am a scaling fan. I do think that just increasing the data and increasing the parameters and having that compression is going to lead to more powerful models. But it is also true, and I think you're absolutely right, that the underlying structure and distribution of the data is really critical. So some datasets will be far more valuable for learning these general principles than others. But I think it goes against a lot of biological intuitions about collecting data. Normally when you think about what data you want, you're trying to answer a very specific scientific hypothesis. You want a very well-controlled experiment, you really want multiple replicates, it's something very focused. So I think the change in the way of thinking is to think: what you really want if you want to learn a general representation of proteins is to see amino acids in as many evolutionary contexts as possible. That's really what you want. That's how I think about data. And I think what changed between ESM2, which was the previous generation model, and ESMC, which is this new generation model, because they're both at approximately the same scale. Same scale of compute. Same scale parameters. ESM2 got a lot of compute, but ESMC got even more compute. But it's not just the compute. The data was really the critical thing here. When we trained ESM2 we observed two things. The first was that as we increased the number of parameters and compute, we saw improvements. So we had a model at the billion parameter scale, a model at the 10 billion parameter scale, and the larger scale model is better than the smaller scale model. But if you look at a plot of parameter scale versus capability—for capability we're looking at representational fidelity, how well it captures protein structure—you could see that there are diminishing returns in ESM2. ESM2 is trained on UniRef. And for ESMC, we added metagenomics. So we added billions more sequences to the training data.
你能解释一下 UniRef 和宏基因组学是什么意思吗?
Could you explain what UniRef and metagenomics mean?
好的。UniRef 是序列生物学的黄金标准数据集。它从各种不同的测序资源中获取序列,进行聚类以去除你提到的冗余,并创建了蛋白质生物学的确定性覆盖。
Yeah. So UniRef is the gold standard dataset of sequence biology. It takes sequences from across a wide variety of different sequencing resources, clusters them to remove some of the redundancy that you mentioned, and creates a definitive coverage of protein biology.
与经典基因测序并行发展的是宏基因组测序,人们前往各种不同的生物群落和环境,从世界各地采集样本,对那里存在的自然多样性进行测序。比如来自热液喷口的蛋白质,来自南极附近寒冷环境的蛋白质,或者深海、土壤、人类肠道——各种不同的环境。这是一种非常不同的数据收集方式。你不是试图理解某个特定生物体的特定基因组或某个特定蛋白质,而是收集一堆东西,混在一起,提取序列。你完全不知道这些序列来自什么生物,甚至不一定知道某个序列是不是蛋白质,但你可以根据某些背景猜测,然后说:“好吧,我们把它们放在一起。这些可能是我们找到的蛋白质序列。我们不把它们分配给某个生物体,也不分配给更大的背景。我们只是说这可能是蛋白质,然后训练它。”
What has happened in parallel to classical gene sequencing is this idea of metagenomic sequencing, where people go out into all kinds of different biomes and environments and collect samples from the world and just sequence the natural diversity that's present there. So, proteins from a hydrothermal vent or proteins from a frigid environment near the South Pole, or the deep ocean, or soil, or the human gut — all kinds of different environments. This is a very different way of collecting data. Instead of trying to understand a specific genome of a specific organism or trying to understand a specific protein, you just collect a bunch of stuff, mix it up in a pot, get the sequences out. You have no idea what organisms these are from. You don't necessarily even know if a given sequence is a protein, but you can guess based upon certain contexts and you say, "Okay, we throw these together. These are likely protein sequences we find. We're not assigning them to an organism. We're not assigning them to a larger context. We're just saying this is probably a protein. Let's train on it."
没错。而且你甚至得不到完整的基因组,只能得到一些重叠群,这些重叠群常常是断裂的,甚至包含部分蛋白质。所以数据非常嘈杂。还有一个技术性问题:如果我理解正确的话,你实际上并不是用设备直接测序蛋白质,而是对制造这些蛋白质的 DNA 进行测序。也就是说,你找到 DNA,然后寻找指示蛋白质序列起始和结束的标记。是这样吗?
That is right. Yeah. And you don't even get the full genomes. You just get these kind of contigs that often are broken and have even partial proteins. So, the data is really noisy. One more nerdy question: if I understand correctly, you're not actually using a device that sequences proteins. You're sequencing the DNA that would manufacture those proteins. So, you're finding DNA and then looking for markers that indicate the beginning and end of a protein sequence. Is that kind of...
对,完全正确。基本上就是测序基因序列,然后从这些序列翻译出蛋白质。所以你会去挖下水道——不是说你本人去挖,但确实有下水道,可能还有纽约地铁,各种地方。
Yeah, that's exactly right. Basically sequencing genetic sequences and then you translate the proteins from those sequences. So you're digging up sewers — not you personally, but there are sewers, probably New York City subway, all kinds of things.
那么我自然想问:你构建了这个模型,并且认为已经做了去重,得到了一个没有太多冗余的良好表示集。还有多少蛋白质有待发现?如果我们有十倍多的资源,你认为是否还有十倍多的蛋白质可以发掘?
So the natural question to me is: you built this model and you think that you've kind of deduplicated it so that you have a good representational set without a lot of redundancy. How much more is there? If we had an order of magnitude more resources, do you think there is an order of magnitude more proteins to discover?
我觉得是的。我不完全确定,但蛋白质数量巨大,我认为我们在测量地球生物多样性方面只触及了皮毛。有一些核心蛋白质在所有生命中都保守存在,这我们知道。但随着你进入这些不同的环境,进化不断创造新的基因和新的蛋白质。据我理解,其中很多来自病毒、细菌和其他微生物。它们之间长期处于冲突状态,这促使它们以有助于在极端或任何环境中生存的方式重组 DNA。正是这一点导致了蛋白质的惊人多样性。
I think so. I'm not entirely sure, but there are a lot of proteins and I think we've barely scratched the surface of measuring Earth's biodiversity. There are core proteins that are conserved across all of life, we know that. But as you go into these different environments, there are new genes and new proteins constantly being created by evolution. A lot of this, from my understanding, is viruses and bacteria and other microorganisms. These guys are in this basically long-running conflict with each other that causes them to recombine their DNA in ways that help them survive in extreme or whatever environment. And that's what's causing this incredible diversity of proteins.
没错。四十亿年来,生命在地球上各种不同的生态位中并行进行实验。我们看到了所有这些结果。所以这种组合效应——这就是为什么你认为从宏观角度看,多样性可能远不及微观尺度,因为微观尺度上有这种惊人的组合效应。
That's right. And just 4 billion years of life running experiments in parallel all across the earth in all kinds of different ecological niches. We see the outcome of all that. So the combinatorial effect — that's why you believe that from a macroscopic perspective there's maybe not nearly as much diversity as at the microscopic scale, because you have this incredible combinatorial effect.
是的。那里有巨大的多样性。那么回到之前的话题……
Yeah. I mean there's just tremendous diversity there. So going back then...
是的。我知道这很棒,对吧?我们还可以讨论数据和构建细胞模型,从分子层面走向更高的生物复杂性。但为了完整描述 ESMC:最大的变化是加入了这些宏基因组序列。我们看到的是,规模不再有收益递减。这实际上说明,对于 ESMC 来说,ESM2 是数据受限而非算力受限。我们可以绘制一个漂亮的缩放定律:在小规模上训练模型,观察它们在给定算力预算下能达到的最佳表示保真度,然后画一条外推线,完美预测更大规模模型能达到的效果。所以这是一个漂亮的缩放。对 ESMC 的唯一改动是让它训练时更高效,但我认为数据才是真正的驱动力。
Yeah. I know it's great, right? I think we could also talk about data and building models of the cell and going from the molecular level to higher levels of biological complexity. But to complete the description of ESMC: the big change was adding these metagenomic sequences. What we saw is that there are no longer diminishing returns to scale. That's really saying that ESM2 was data limited rather than compute limited for ESMC. There's a beautiful scaling law we can plot: we train models at smaller scales, look at the best representational fidelity they can achieve for a given compute budget, and draw a line of extrapolation that beautifully predicts what larger scale models will achieve. So there's this beautiful scaling. The only changes to ESMC were to make it a more efficient model for training, but I think the data is really the big thing driving that.
所以它基本上还是一个标准的普通 Transformer,有一些技巧——到这个阶段大家都有一些技巧——语言模型,加上大量数据。
So it still is basically just a standard vanilla transformer, a few tricks — everyone has a few tricks at this point — language model and just a lot of data.
这与 AlphaFold 之类的方法形成鲜明对比,对吧?AlphaFold 在模型中内置了大量归纳偏置来预测蛋白质结构。
This is very much in contrast to something like AlphaFold, right? Where you have a lot of inductive bias built into the model to predict protein structure.
没错。这里的想法是:我们能否直接学习正确的结构?不提供任何先验,让机器学习自己找出结构是什么。
That's right. And the idea here is: can we just learn the right structure? Don't give any priors, just allow machine learning to figure out what that structure is.
那么你在 ESM3 中也曾绕道先验,对吧?或者不完全是先验,而是使用了更多直觉或更多人工设计。你认为 ESM3 是一次绕道吗?你是不是最终说“让 C 更大”,然后突然就奏效了,现在你发现我们不再需要先验了?这是一个关键洞见,还是你仍然认为先验有空间?
So you also had your own detour into priors with ESM3, right? Or maybe not priors, but using more intuition or more human design. Do you think ESM3 was a detour? Did you just end up saying "let's make C bigger" and suddenly it worked, and now you learned that we don't need priors anymore? Is that a key insight, or do you still think there's room for priors?
我认为两者都需要。两者都有其位置。ESM3 的目标是真正让生物学变得可编程。我们试图思考:编程语言是什么?你如何让生物学家能够提示模型并设计结构和功能?所以我们认为它需要正确的轨道。但我要说,ESM3 与 ESM 的哲学非常一致,因为我们所做的基本上就是为这一系列进化上多样的蛋白质预测结构,并将其用作训练数据。
I think we need both. There's a place for both. The goal for ESM3 was to really make biology programmable. We were trying to think: what is the programming language? How are you going to allow biologists to prompt a model and design structure and function? So we really thought it needed the right tracks. But I would say that ESM3 was very consistent with the philosophy of ESM, because what we did was we basically predicted structures for this vast array of evolutionarily diverse proteins and used that as training data.
所以模型现在从序列模式、结构模式、功能模式中学习。但我认为,模型在序列上学习的这种综合能力,你可以想象引入更多多维信息会构建出更好的表示空间。如果你是一个程序员,或者你在构建语言模型然后构建编码智能体,你会从预训练所有数据开始,然后通过某种后训练(可能是强化学习)来进入编程部分。你有没有想过对 ESMC 进行后训练,以赋予它同样的可编程能力?你认为能否在不使用所有归纳偏置(比如结构图谱,它们大多是对某种有趣知识的蒸馏)的情况下获得可编程性?但我想也许那只是另一种模型的预训练。
So the model's now learning from sequence patterns, learning from structural patterns, learning from functional patterns. But I think that same kind of synthesis that the model is learning on sequences, you could imagine bringing in more multi-dimensional information would build an even better representation space. If you are a coder or you're building language models and then building coding agents, you start with pre-training on everything and then you go to doing the programming part by some sort of post-training, probably RL. I mean, have you thought about post-training ESMC to give you the same abilities for programmability? Do you think you could get programmability without doing all of the inductive biases which involve like atlas of structures and which mostly distill some sort of interesting distillation but I guess maybe that isn't some kind of pre-training of a different model.
是的。我认为这是一个非常有趣的问题,关于这些模型能在多大程度上相互转换。我觉得这一点还没有完全被理解,但思考如何去做以及什么才是正确的方法,这是一个非常有前景的方向。
Yeah. I mean I think it's a really interesting question kind of to what degree can you interconvert these models. I don't think that's fully understood yet, but I think it's a very promising direction to think about doing that and what are the right ways to do that.
所以,ESMC 也在接近可编程生物学,但我要说方式非常不同。它从世界模型的角度出发,基本思路是你有一个预测模型,然后你搜索这个世界模型,找到满足你任何设计标准的蛋白质分子。因此,我们已经能够用它来设计微型蛋白结合剂,但最令人兴奋的是,我们还能用它来设计抗体、单链抗体(scFv),并且在少量试验中看到了非常令人振奋的成功率。
So, ESMC is also approaching programmable biology, but I would say in a very different way. It's approaching it from this kind of world modeling perspective where the idea is basically you have a predictive model and you're going to search the world model to find protein molecules that satisfy whatever design criteria that you have. So, we've been able to use this to actually now go and design mini protein binders, but I think most excitingly, we've been able to use this to actually design antibodies, scFvs, and we're seeing really exciting success rates in a small number of trials now.
是的。那么,你能解释一下什么是 scFv 吗?
Yeah. So, can you explain what those scFvs are?
是的。scFv 基本上是一种单链抗体。它是一种治疗形式,包含一条重链和一条轻链。抗体通常有重链和轻链,然后由一对重链和轻链、另一对重链和轻链组合在一起识别靶点。因此,治疗中使用的这些形式有不同变体。scFv 的有趣之处在于它只有一条重链和一条轻链,所以它能形成非常复杂的结合界面,两个不同的亚基结合在一起与靶点结合。这是一种重要的治疗形式;大约四分之一的新药是抗体。所以它确实是医学中的关键形式之一。基本上,我们能够看到的是,你可以搜索 ESMC,实际上找到达到治疗功能和活性所需亲和力水平的抗体。
Yeah. So, an scFv is basically a single chain antibody. So it's a kind of therapeutic modality that basically has a heavy chain and a light chain. An antibody has a heavy chain and a light chain, and then it basically has a pair of one heavy chain, one light chain, another heavy chain, and one light chain that come together to recognize a target. So there are different variations of these modalities that are used therapeutically. And so what's interesting about the scFv is it has one heavy chain and one light chain. So it is able to form these very complex binding interfaces where you can have two different subunits coming together to engage a target. These are an important therapeutic modality; something like a quarter of new drugs are antibodies. So it's really one of the critical modalities for medicine. And basically what we're able to see is that you can search ESMC and you can actually find antibodies that are reaching the level of affinity that is needed for therapeutic function and activity.
蛋白质设计领域在过去 5 年爆炸式发展。每个人都在做蛋白质设计;很多人对此充满热情。我高层次的天真理解是,像微型结合剂这样的东西相当可行。人们已经相当常规地成功做到了。到了纳米抗体和 scFv,设计起来就有点难了,而抗体通常仍然遥不可及。其中一个常见原因是,如果你处于 AlphaFold 范式,你没有多序列比对(MSA),对吧?抗体的进化压力在很多方面与其他所有东西的进化压力相反;它们追求多样性,而不是沿着非常受限的路径进化。所以我很好奇,你尝试过更大的结构吗?你在那方面成功了吗?还是说你觉得由于某种原因这仍然很难做到?
The protein design space has kind of exploded in the last 5 years. Everyone is doing protein design; many people are excited about protein design. My high level naive understanding of the field is that things like mini binders are quite doable. People have done that quite routinely successfully. By the time you get to nanobodies and scFvs, they're a little bit harder to design, and antibodies are still actually quite out of reach often times. One of the common reasons for this is if you're in the AlphaFold paradigm, you don't have MSAs, right? The evolutionary pressure for antibodies is actually the opposite in many ways of what the evolutionary pressure is for everything else; they go for diversity rather than trying to evolve along a very constrained path. So I'm curious, did you try larger structures and is that something that you've seen success on, or is this something that you still think for some reason it might be hard to do?
我们实际上可以把 scFv 重新格式化为抗体。所以我认为这是最快的方法。我们还没有尝试过完整的 IgG。我看不出有什么理由行不通。实际上,我们还没有做。我们决定现在发布这个,是因为我们觉得它已经达到了一个点,我们看到了比过去可能做到的显著进步。所以我们只是想把它推出去,但我认为还有更多进步的可能。因此,我们有很多合作来研究这里的一些其他应用。关键在于它是一个通用模型。所以对我来说,最令人兴奋的是:一个用于蛋白质序列、结构和功能的通用模型。你可以搜索它,治疗设计基本上就从搜索中涌现出来。
We can actually take the scFvs and reformat them as antibodies. So I think that would be the quickest approach to do that. We've not tried full IgGs. I don't see any reason why that wouldn't work. Actually, it's something we haven't yet. We've decided we're basically releasing this now because we feel like it's reached a point where we're seeing a really significant step above what's been possible in the past. And so, we just wanted to get it out there, but I think there's a lot more progress that's possible. So, we have a lot of collaborations to look at some of the other applications here. The thing about it is it's a general model. So to me, that's the most exciting thing about it: a general model for protein sequence, structure, and function. You can search it and therapeutic design basically emerges from that search.
是的。对我来说,你提到你没有使用多序列比对(MSA),这是让 AlphaFold 工作得非常好的关键见解之一。而你不需要它就能让模型基本上和 AlphaFold 3 一样好,这让我非常兴奋。因为这意味着你的论点——尽可能覆盖可能的蛋白质空间,看看涌现行为是什么——如果这是一种涌现行为,我们能够复现多序列比对的效果,那么其他那些我们可能没有数据但也能以涌现方式做到的事情是什么呢?
Yeah. I mean to me, you mentioned that you're not using MSAs, multi-sequence alignments, which was one of the critical insights that allowed AlphaFold to work really well. And the fact that you didn't need that in order to make it work basically as well as AlphaFold 3 is really exciting to me. Because that means that your thesis of let's cover the space of possible proteins as well as we can and see what the emergent behaviors are, so that if this is an emergent behavior that we're able to replicate what happens with multi-sequence alignment, what are the other things that maybe we don't have data for but that we are able to also do in an emergent way?
我可以说实际上我们在抗体上做得明显更好。所以我认为这是非常酷的一点。这是我们当初的一个论点:抗体可能不会像预测分子结构拓扑那样从进化信息中受益。所以我认为你现在可以看到,表示空间中包含了一些关于抗体的非常有趣的东西。
I would say actually we're doing significantly better on antibodies. So I think that's one of the things that's really cool. That's one of the theses that we had: antibodies are not going to benefit from evolutionary information probably in the same way that predicting the structural topology of a molecule will. So I think you kind of see that now where the representation space is containing something that's really interesting about antibodies here.
我想谈谈,因为你提到了一个对我来说非常有趣的东西,那就是虚拟细胞以及它可能如何与此接口或在此工作。我真的很想知道:你在机制可解释性中找到了其他东西吗?有哪些有趣的东西不仅仅是验证生物学,而是出现了意想不到的模式?你发现过类似的东西吗?
I want to talk about because you mentioned something very interesting to me which was talking about virtual cell and how this maybe interfaces or did this work here. I'm really interested to know: were you able to find other things in your mechanistic interpretability? What were some interesting things that weren't just validating biology, but there's a pattern that was unexpected. Did you find anything like that?
这很复杂。
It's complicated.
所以,我们实际上必须去验证其中的一些东西。我认为我们看到了一些有趣的关联。例如,进化上关系遥远的基因编辑系统在这个空间中聚集在一起,其方式与我们对该基因编辑系统起源的了解一致并反映了这一点。这非常令人兴奋。但地图中有许多蛋白质以不同的方式聚集在一起,我们目前还不知道它们是什么,也不知道它们的功能。一个假设是,这些是新型基因编辑系统。我认为在这个图谱中,会有一些非常有趣的科学发现基础。如果你想想人们如何寻找新的基因编辑系统,他们通常是在大型基因序列数据库中挖掘,寻找与之相关的不同序列模式或结构模式。实际上,ESM 图谱的第一个版本就被 Funang 的团队用来发现了一个新的基因编辑系统。我认为有很多我们不了解的生物学正在等待被发现。能够连接蛋白质之间的点,使我们从已知出发对未知进行推断——这正是让我兴奋的地方。我认为自然界可能已经发明了用于许多应用的蛋白质。想想 Taq 聚合酶,它使 PCR 成为可能,这种酶来自生活在热温泉中的细菌。也许解决气候变化的方案就在蛋白质生物学的某个地方。可能存在着构建完全绿色化学基础设施的各种基础模块。可能还有新的药物和疗法。但问题是如何找到它们?能够连接这些点确实是开始打开蛋白质生物学发现空间的一种方式。
So, because we have to actually go and validate some of these things. I think what we saw are interesting connections. For example, distantly evolutionarily related gene editing systems cluster together in this space in ways that are consistent with and reflect our knowledge of the origin of those gene editing systems. That's really exciting. But there are a number of proteins in that map that are brought together in different ways where we just don't know what they are right now. We don't know what they do. One hypothesis is that these are novel gene editing systems. I think in this atlas there's going to be some really interesting basis for scientific discovery. If you think about how people go out and look for new gene editing systems, they're typically mining large genetic sequence databases and looking for different sequence patterns or structural patterns linked to that. Actually, the first version of the ESM atlas was used by Funang's group to find a new gene editing system. I think there's a lot of biology out there that we don't understand that's waiting to be discovered. Being able to connect the dots between proteins so that we can go from what we know today to make inferences about the unknown—that's what I'm excited about. I think there are proteins for so many applications that nature has probably invented. You think about the Taq polymerase which enables PCR, which came from a bacteria living in a thermal hot pool. There may be the solution to climate change somewhere in protein biology. There are probably all kinds of building blocks for completely green chemistry infrastructure out there. There's probably new medicines and therapies. But the question is how do you find those? Being able to connect the dots is really one way to start to open up that space of protein biology to discovery.
我很好奇,ESMC 的进步之一是多聚体预测的改进,基本上是蛋白质-蛋白质相互作用,即预测两种蛋白质相互作用方式的能力。我认为你现在声称比任何人都做得好,对吗?如果我错了请纠正。
I'm curious, one of the advancements of ESMC is an improvement in multimer, basically protein-protein interactions, the ability to predict the way two proteins interact. I think you now claim to do better than anyone else, right? Correct me if I'm wrong.
是的,我认为我们在开放模型中是最先进的。
Yeah, I think we're state of the art for open models.
我知道有些人认为对虚拟细胞非常有用的一个东西是整个人类转录组中每一对蛋白质的完整映射。你有没有考虑过这样做,作为虚拟细胞的开始,比如创建那个地图?
One thing which I know some people would find very useful for virtual cell is just an entire mapping of every single pair of proteins inside the human transcriptome. Have you thought about doing this as a beginning to a virtual cell, like create that map?
我认为类似的东西会非常有价值。ESMFold 2 的另一个特点是它非常快,因为它不需要多序列比对。你可以直接从序列进行推理。只需几秒钟就能获得原子分辨率的预测。所以这是 Biohub 一个非常有趣的应用。我们在想的另一件事是,我们能否在实验上解决这个问题?我们正在构建的一个东西是冷冻电子断层扫描,我们正在构建能够极大提高在原子水平观察细胞时对比度的系统。我希望在未来某个时候能看到结构上经验解析的相互作用组。有一些相当大的技术障碍和需要开发的技术来克服这一点,但我认为这是可能的。我们可以用计算方法开始获得一个代理,这将会非常强大。但我认为结构预测的未来很大程度上将转变为结构测定,将这些用于蛋白质建模的工具与实验数据结合起来,以便我们能够构建一个由经验生物学和我们能观察到的东西所启发的图景。
I think something like that would be really valuable. The other thing about ESMFold 2 is that it's a really fast model because it doesn't require the multiple sequence alignment. You can do inference directly from the sequence. It takes seconds to get an atomic resolution prediction. So that's one really interesting application at Biohub. The other thing we're thinking about is can we actually experimentally resolve this? One of the things we are building is cryo-electron tomography, and we're building systems that can greatly increase the contrast when looking at the cell at the atomic level. I hope to see structurally empirically resolved interactomes at some point in the future. There are some pretty big technical hurdles and technologies that have to be developed to overcome that, but I think that's going to be possible. We can use computational methods to start to get a proxy of that, and that's going to be really powerful. But I think a lot of the future of structure prediction is going to turn into structure determination, bringing together these tools for modeling proteins with experimental data so that we can develop a picture informed by empirical biology and what we can observe.
那么,如果我没理解错的话,这是不是就是这里的愿景:你有一个实验室在回路中的东西,其中有一个智能体与你的冷冻电镜等设备对话,然后它预测你感兴趣的属性。它测序基因组或创建基因组。它从基因组中创建蛋白质。然后它用某种版本的显微镜观察它。你刚才叫那个显微镜什么来着?
So is that the vision here, if I'm understanding correctly, that you have a lab-in-the-loop kind of thing where you have an agent that is talking to your cryo-ET and whatever, and then it predicts a property that you're interested in. It sequences the genome or creates the genome. It creates the protein from the genome. It then observes it with some version of this microscope. What did you call the microscope again?
冷冻电子断层扫描。
Cryo-electron tomography.
好的。然后你做任何实验或观察它,然后你把它用作实验室在回路中,说:“哦,好的,它这样折叠。因此,我想检查的下一个实际上是另一个”,并使用主动学习系统。这是你在这里阐述的那种愿景吗?
Okay. And then you do whatever experiments or you observe it and then you use this as a lab-in-the-loop to say, "Oh okay, this folds this way. Therefore, I want to check the next one that I want to check is actually a different one" and use an active learning system. Is that sort of the vision you're articulating here?
嗯,我认为下一个生物学时代会有几个基本原则。现在是一个非常有趣的时期,因为我们正处于一个新科学范式的开端。是什么定义了那个范式?我认为有几个原则:数据生成,这将非常关键;计算性的预测性数字生物学表示——你可以把 ESM 视为第一代,AlphaFold 也是这些方法的第一代,我们可以开始思考随着我们对更多生物复杂性进行建模,那会是什么样子;然后你有反馈原则;还有我们拥有可扩展的智能,可以应用于生物问题的每一个单元。所有这些结合在一起意味着什么?我认为我们将拥有越来越强大和准确的分子、基因组、细胞,最终是生理学的数字表示。那是你想要达到的目标。我们将不得不沿着复杂性尺度上升,生物复杂性的各个层次需要跨越一个数据障碍。有些数据不存在,需要生成才能达到那种预测保真度。然后我们将拥有推理。
Well, I think there are going to be a few fundamental principles for the next era of biology. It's such an interesting time right now because we're at the beginning of a new scientific paradigm. What is defining that paradigm? I think there are a few principles: data generation, that's going to be really critical; computational predictive digital representations of biology—you can think of ESM as first generation, AlphaFold as first generation of those approaches, and we can start to think about what that looks like as we model more biological complexity; then you have the principle of feedback; and you have the principle that we have intelligence now that's scalable and can be applied to every unit of a biological problem. What would it mean for all of that to come together? I think we're going to have increasingly capable and accurate digital representations of molecules, genomes, cells, ultimately physiology. That's where you want to get. We're going to have to go up that complexity scale, the levels of biological complexity that require traversing a data barrier. There's data that does not exist that needs to be generated to achieve that level of predictive fidelity. And then we're going to have reasoning.
我认为这意味着,我们可以利用预测性预言机,以数字方式并行推理成千上万、数百万乃至数亿个科学假设,这些预言机实际上能够预测实验的结果。因此,我们提出问题的规模以及能够提出的问题类型,将通过这种反馈发生根本性改变,这一点至关重要。模型需要有一个 Scaling 维度,即构建数据以获得准确的表示;还需要一个反馈维度,即模型能够从生物学中学习,进行数字推理,将问题缩减为少量实验假设,检查每个实验的结果,更新理解,并以此方式构建知识。我认为这就是未来的样子,我们必须构建每一个组件。Biohub 真正想要做的,是将实验层和技术层结合起来,使这些 AI 模型能够与生物学互动并进行实验。我们看到,在能够通过计算获得反馈的封闭领域中,取得了令人难以置信的进展,但实验生物学当然是完全开放的。因此,那里的反馈原则将非常不同。但会有类似 RLVR 的方法应用于实验,让模型真正构建知识、从知识中学习,并能够开发出越来越准确的表示。
And I think you know what that will mean is that we can reason over thousands, millions, hundreds of millions of scientific hypotheses in parallel digitally using predictive oracles which can actually predict the outcome of an experiment. So the scale that we can ask questions and the kinds of questions that we can ask will just fundamentally change through that feedback is going to be critical. The models are going to need to have a scaling dimension of building the data to have those accurate representations and then a feedback dimension where the models can learn from biology, can reason digitally, can reduce that to a small number of experimental hypotheses, examine the outcome of each of those experiments, update their understanding, and build knowledge in that way. So I think that's what it's going to look like and we kind of have to build each of those components. What Biohub is really trying to do is to bring together the experimental and the technology layer that will actually allow us to have these AI models interact with the biology and do experiments. And I think we see incredible advances in areas where we can get feedback computationally, in enclosed domains, but of course experimental biology is completely open-ended. And so the feedback principle there is going to be very different. But there's going to be something like RLVR with experiments where we can have models that are just really building knowledge and learning from that knowledge and being able to develop more and more accurate representations.
你是 Biohub 的科学负责人。也许有个趣闻,对于不知道的人来说,Ledin Space 的科学板块基本上是在大约 6 个月前,马克·扎克伯格和普莉希拉·陈做客本播客之后或作为回应而推出的。实际上,能邀请你来到这里,有种回到起点的感觉,非常令人兴奋。马克为 Biohub 想要实现的目标描绘了相当宏大的愿景,我认为你刚刚阐述的正是那个愿景非常自然的延续。我记得你当时刚加入,好像才两周。
You're the head of science at Biohub. Maybe fun fact for those who don't know, the science section of Ledin Space was basically launched after or in response to Mark Zuckerberg and Priscilla Chan on this podcast about 6 months ago. It's actually very exciting to have you here and kind of come full circle. Mark laid out quite an ambitious vision for what Biohub wants to accomplish and I think you just laid out a very natural successor to that. I think you had just joined at like you were there two weeks.
我是在十月底加入的,十一月初就启动了。
I joined at the very end of October and launched at the beginning of November.
是的。我好奇的一点是,在你看来,Biohub 现在处于什么阶段?你想实现什么?你的目标是什么?对于那些还没看过马克和普莉希拉那期节目的听众,我们当然推荐他们去看。链接在描述里。另外,即使你才来了短短六个月,你学到了什么吗?愿景有没有演变?你认为这会走向何方?ESMC 如何融入其中?你最近宣布的虚拟生物学计划又如何融入?然后我觉得你还在做其他几件事,我们还没谈到。
Yeah. One thing I'm curious about is in your eyes, where is Biohub now? Like what do you want to accomplish? What are your goals? Big picture goals for listeners who haven't watched the episode with Mark and Priscilla. We recommend of course that they go watch it. Link in the description. And then have you learned anything even in just the short time of six months you've been here? And like has the vision evolved and where do you see this going? How does ESMC fit into this? How does the virtual biology initiative that you recently announced fit into this? And then I think there's like several other things that you're working on that we haven't even touched on.
是的。我每天都在学习新东西。但我的想法是,我们正在为这个新范式建立一个科学机构。要做到这一点,这个机构将由前沿实验生物学、用于测量和观察的前沿技术,以及前沿人工智能驱动。
Yeah. Well, I'm learning things every single day. But the way I think about it, we're building a scientific institution for this new paradigm. And to do that, it's an institution that's going to be powered by frontier experimental biology, frontier technology for measurement, for observation, and it's going to be powered by frontier artificial intelligence.
而且这一切都是开源的,对吧?
And this is all open source, right?
这是一个慈善项目。所以我们的目标是加速科学进步。我们的使命是治愈或预防疾病。要做到这一点,我们相信我们的理解存在根本性的差距,我们需要加速科学来跨越这个差距。因此,我们真正在思考生物理解的每一个层面,从最基本的层面,比如细胞中蛋白质的原子,一直到生理学和疾病中的细胞系统,以及我们如何创建能够捕捉这种复杂性、让我们理解这种复杂性的模型。我认为,如果你思考治愈疾病是什么样子?它不是一颗药丸,不是传统意义上的药物。它必须是一个能够建模和理解疾病潜在生理学的系统,并且对每一个人类个体、每一个不同的基因组都有所区分。它还必须能够将分子尺度的事件与疾病在生理学中的表现联系起来。所以这是一个极其复杂、极其困难的问题。对我们来说,我们正在努力搭建这些复杂性的层级,并构建科学家可以用来回答基本问题的基础工具。因此,我们正在创建原子级成像。我们正在创建光片显微镜,让我们能够观察所有细胞在发育生物体中如何移动和发育。我们正在创建具有空间和时间分辨率的炎症图谱。我们正在创建细胞编程和免疫细胞重编程,以便能够设计完全可编程的疗法。我们还在每个层级创建这些数字表示,以便加速科学、模拟正在发生的事情,使生物物质、蛋白质、细胞和基因组变得可编程。我认为所有这些都必须结合在一起。我认为,如果你有专注力,并且将生物学和计算层紧密集成在一起构建,这就是我们取得最快进展的方式。在过去的 10 年里,我认为我们一直是开放科学的主要倡导者之一。我们是一个既资助又建设的组织。在我们的资助中,我们一直支持开放科学。在我们的建设中,我们也一直践行开放科学。所以这将继续下去。这是非常根本的。我们不是一家药物开发公司。我们不是要制造疗法。我们是在努力构建推动科学进步的技术。
It's a philanthropy. So our goal is to accelerate science. Our mission is to cure or prevent disease. To do that, our belief is that there's a fundamental gap in our understanding and we need to accelerate science to traverse that gap. So we're really thinking about every layer of biological understanding that goes from the most basic level like the atoms of a protein in a cell all the way to systems of cells in physiology and disease and how can we create models that can capture that complexity, can allow us to understand that complexity. And I think if you think what does the cure to disease look like? It's not a pill, it's not a medicine in the conventional sense. It's going to have to be a system that is capable of modeling and understanding the underlying physiology of disease in a way that's differentiated for every single human being, for every single different genome. And it's going to have to be able to link events all the way at the molecular scale to the manifestation of disease in physiology. So it's an incredibly complex, incredibly hard problem. And for us, we're trying to ladder up those layers of complexity and we're trying to build the foundational tools that scientists can use to answer the fundamental questions there. So we're creating atomic level imaging. We're creating light sheet microscopy that allows us to observe how all the cells move and develop in a developing organism. We're creating spatially and temporally resolved maps of inflammation. We're creating cellular programming and immune cell reprogramming to be able to actually design completely programmable therapies. And we're creating these digital representations at each of these layers so that we can accelerate the science, simulate what's happening, make biological matter and make proteins and cells and genomes programmable. And I think all of that has to come together. And I think if you have the focus and you build the biology and the computational layers together so that they're tightly integrated, that's how we're going to make the fastest progress. For the last 10 years, I think we've been one of the big champions of open science. We're an organization that both funds and builds. And in our funding, we've always supported open science. And in our building, we've always done open science. So that's something that's going to continue. It's just really fundamental. We're not a drug development company. We're not trying to generate therapies. We're trying to build the technology that moves science forward.
所以我认为马克有一个概念,如果你提供正确的工具,那么整个科学界都可以利用它们。是的。显然你非常相信蛋白质语言建模作为一种工具。那么,对于提升我们应对人类疾病能力的整体进步,下一个最重要的工具是什么?
So I think Mark had this concept about if you provide the right tools, then the entire scientific community can leverage them. Yeah. So obviously you believe strongly in protein language modeling as a tool. What is the next most important tool for advancing a general improvement in our ability to tackle human disease?
所以我认为,我们接下来要应对的复杂性是细胞的复杂性。这将是极其困难的——数十亿种蛋白质。
So I think the next level of complexity that we have to address is the complexity of the cell. And I mean this is going to be tremendously hard — billions of proteins.
所以你说这极其困难。是的。如果你来说这会轻而易举,我们就会……
So you say it's tremendously hard. Yeah. If you come and say it's going to be easy peasy, we would just...
嗯,我认为这是一个值得挑战的问题。但它需要今天还不存在的技术,需要新的建模方法,以及可能还不存在的机器学习的架构和思想。所以有一些深刻而根本的问题需要解决。但同样,我们一步一步来。所以我们从分子层面开始,我们知道那确实是根本性的,我们可以开始将其与细胞生物学中的可观测现象联系起来。
Well, I think it's a worthy challenge. But it requires technology that doesn't exist today, requires new modeling approaches and probably architectures and ideas in machine learning that probably don't yet exist. So there are deep and fundamental problems to solve. But again, we take it step by step. So we kind of start at the molecular layer and we know that that is really fundamental, and we can begin to link that to observables in cellular biology.
我很好奇,因为这个问题在我脑海里已经很久了:我们有虚拟细胞模型,有分子尺度模型,也有几篇论文试图将它们联系起来。但你们在做什么?因为这听起来像是你正在思考的重点。
I'm really curious because this has been the question that's been on my mind for a long time: we have virtual cell models, we have molecular scale models, and there have been a few papers about trying to link them. But what are you guys doing? Because it sounds like this is becoming top of mind for you.
那么让我们用蛋白质生物学来类比。我认为我们的蛋白质数字表示之所以强大且有用,是因为它们能够泛化。它们能够对与训练数据中完全不同的蛋白质进行预测。它们能够泛化,以至于你可以设计全新的折叠、新的结合界面、新的结构。所以这就有了一定程度的泛化能力或通用性。简而言之,它们可以预测我们尚未进行的实验的结果,这些实验它们尚未经过训练。因此,数字表示要有价值,就必须能够用来回答新的问题。我认为这是关键。所以我们在细胞方面还没有达到那个水平。我认为当前一代被称为虚拟细胞的模型,它们对底层数据有很好的表示,但当你在一个新的未观察到的背景下进行新的干预时,它们预测会发生什么的能力非常有限。但要能够回答关于细胞生物学的基本科学问题,我们需要一个能够做到这一点的模型。所以我们对这个问题的思考就从那个想法开始:要到达那里需要什么?
So let's make the analogy with protein biology. What I think makes our digital representations of proteins powerful and useful is that they generalize. They're able to make predictions for proteins that are entirely unlike the proteins in their training data. They're able to generalize so that you can design fundamentally new folds, new binding interfaces, new structures. So there's this degree of what we call generalization or generality. In short, they can predict the outcome of an experiment that we haven't already made, that they haven't already been trained on. So for digital representations to be valuable, they've got to be able to be used to answer a new question. I think that's the critical thing. So we're not there with cells. I think with the current generation of models that are being called virtual cells, they are good representations of the underlying data, but they have a very limited ability to predict what will happen when you make a novel intervention in a novel unobserved context. But to be able to answer the fundamental scientific questions about cellular biology, we need a model that can do that. So our thinking about this starts with that idea: what's it going to take to get there?
回到蛋白质-蛋白质相互作用,人类相互作用组。如果你有了那个,仅仅预测静态结构——静态结构在某种意义上不足以理解很多生物学。动力学对大多数人来说是一个更有用的工具。你可以从静态开始,它能给你一些洞见,但它很少是完整的答案。所以你有一个能够预测许多不同蛋白质的模型。我们可能在 PDB 中解析了许多这样的蛋白质,有些还没有。考虑到动力学和相互作用更重要,你如何弥合这个差距?对我来说,这似乎是从一个非常微观的模型走向更接近虚拟细胞的关键步骤之一。你实际上必须能够模拟局部蛋白质、RNA、DNA、脂质或细胞中漂浮的任何其他东西的局部相互作用。这是你们会试图弥合的目标吗,还是我误解了?有没有其他方式你想象中能弥合这两者?我的意思是,总有一天可能有一台计算机能够从第一原理模拟细胞,但我们离那还很远,对吧?我认为那远远超出了当前计算技术的范围。我的意思是,即使是模拟单个蛋白质分子折叠的物理过程——基本上我们只能对少数快速折叠的蛋白质做到这一点,但仅此而已。
Going back to protein-protein interaction, the human interactome. If you had that, just predicting static structures — static structures are in some sense not enough for a lot of understanding biology. Dynamics are probably for most people a much more useful tool to have. You can start with static, it can give you some insight, but it's very rarely the full answer. So you have a model which is capable of predicting a lot of different proteins. We probably have many of these resolved in the PDB, some of them we don't. Given that dynamics and interactions are more important, how do you bridge that gap? To me, that seems like maybe one of the key steps in going from a really microscopic model of things to something which is closer to a virtual cell. You actually have to be able to model local interactions of local proteins or RNA or DNA or lipids or whatever else is floating in the cell. Is that sort of a goal that you would try to bridge, or maybe I'm misunderstanding? Is there another way you would imagine bridging these two? I mean, one day it'll probably be possible to have a computer that can simulate the cell from first principles, but we're very far from that, right? I think that's far beyond reach of current computational technology. I mean, even simulating the physics of the folding of a single protein molecule — basically we could do it for a fast folding few, but that's really about it.
是的。所以存在这种双重视角,这种互补的生物学视角。一种视角是第一性原理还原论,所有生物学都可以用更基本的术语解释,用基本的物理、化学、生化术语。我认为历史上有一条很长的研究线,一直试图以这种方式理解和模拟生物现象。而且我认为历史上该领域曾相信,蛋白质折叠问题或蛋白质结构预测问题的解决方案将来自这种第一性原理模拟。而它真的出乎意料地出现了,这个问题可以通过模式识别或这种机器学习方法来解决。所以我认为历史上通过信息论、通过信息来理解生物学一直是富有成效的。你可以把细胞看作一台计算机,一个信息处理机器。在信息术语中,有一些非常基本的原则将基因组中编码的信息与转录的基因以及最终产生的细胞表型联系起来。所以如果我们能够在底层程序层面建模和理解细胞,我认为那给出了正确的抽象。我说的正确抽象是什么意思?嗯,我认为我的意思是今天可能的抽象,因为我们处于大规模信息论的时代。克劳德·香农有一个关于下一个字符的理想预测器的想法,他有一篇非常漂亮的论文,他试图计算英语的熵,并想象基本上取一个无限上下文,然后下一个字符的熵是多少。在当时,那是无法想象的——需要巨大的想象力才能想象那个理想预测器。但今天我们越来越接近能够构建它,而且我们可以对文本做到这一点。那么对于生物学,那个预测器会是什么?这就是 ESM 的想法:它将学习所有生物现象的底层结构。所以如果你从细胞的角度来思考,如果我们能收集足够多的可观察的细胞生物学输出,以揭示底层程序、模式和结构,那么我们就可以创建细胞的信息论描述。我认为那将足以理解疾病。
Yeah. So there's this dual view of biology, this dual complementary view of biology. One view is that kind of first-principles reduction, where all of biology is explainable in more basic terms, in basic physical, chemical, biochemical terms. And I think historically there's a long research line that has really sought to understand biological phenomena and simulate biological phenomena in that way. And I think historically the field had believed that the solution to the protein folding problem or the protein structure prediction problem would come from this kind of first-principles simulation. And it really came out of nowhere that this could be solved using essentially pattern recognition or this type of machine learning approach. So I think historically it has been productive to understand biology through information theory, through information. You can think about the cell as a computer, as an information processing machine. In informational terms, there are these very basic principles that link the information coded in the genome to the genes that are transcribed to the phenotypes of the cell that will result. So if we could model and understand the cell at the level of its underlying programs, that sort of gives, I think, the right abstraction. What do I mean by the right abstraction? Well, I think I mean the abstraction that is possible today because we're in the era of information theory at scale. Claude Shannon had this idea of the ideal predictor of the next character, and he had this really beautiful paper where he tried to compute the entropy of the English language and imagine basically taking an infinite context and then what is the entropy of the next character. At that time, it was unimaginable — it took a great leap of imagination to imagine that ideal predictor. But today we get closer and closer to being able to build that, and we can do that for text. So what would that predictor be for biology? And that's kind of the idea of ESM: it would learn the underlying structure of all biological phenomena. So if you think about that from the standpoint of the cell, if we can collect enough outputs of cellular biology that we can observe to reveal the underlying programs, patterns, and structure, then we could create the information-theoretic description of the cell. And I think that would be sufficient to understanding disease.
这让我想起现在信号通路研究中的很多工作,其中蛋白质在一系列不同的蛋白质-蛋白质相互作用级联中,最终以某种方式导致细胞表型变化。你如何将其转化为可以规模化的事情?或者也许是别的东西,但例如,你如何做到这一点?
This reminds me of a lot of the work that happens in signaling pathways right now, where you have a protein in a cascade of different protein-protein interactions that eventually cause a phenotypic change in the cell in some way. How do you translate that into something that can be scaled? Or maybe it's something else, but how do you, for example, that?
回到苦涩的教训,我们需要数据。我的意思是,我认为你知道为什么蛋白质生物学的这些进步成为可能。它们之所以成为可能,是因为几十年来——对于蛋白质结构来说,半个世纪的实验工作来确定蛋白质的结构——以及整个科学界在测序基因组和宏基因组方面的努力。这创建了这个数据集,你可以真正地在其上训练规模,并真正学习这些更深层的原理。
Going back to the bitter lesson, we need data. I mean, I think you know why these advances in protein biology have been possible. They've been possible because of decades—for protein structure, half a century of work to experimentally determine the structure of proteins—and the effort across the scientific world to sequence genomes and metagenomes. That's created this dataset that you can really train a scale on and really learn these deeper principles.
但这两个数据集实际上在很多方面都相当不同。比如 PDB,一堆非常费力构建的蛋白质结构,其中许多是单个博士论文,然后可能后来有类似的后续,一个博士论文可能有 10 个。然后这个——人们估计创建 GDB 需要大约 130 亿,一个非常大的数字。人们创建 PDB 是因为每个单独的蛋白质都是独立有用的。人们创建它并不是为了“我们在解决蛋白质结构”。他们看到,“哦,这个我们相信参与这个通路的蛋白质,让我们了解这个蛋白质,这样我们就可以靶向它,”等等。当然,这里有一些注意事项,但总体而言,很多基因组数据——特别是对于人类、病毒或细菌——也是出于非常具体的原因而测序的,对吧?嗯,这些事后有用是很好的,但我想知道从现在开始,特别是随着虚拟生物学计划,Biohub 的虚拟生物学计划大约是 5 亿美元,我想。而且我相信未来 Biohub 还会有更多大型计划。你有机会非常具体、有意识地收集数据,以解决 ML 问题,而不是依赖于为其他目的而策划创建的数据集。所以考虑到这个新机会,你会如何不同地做事?当你基本上可以从第一原理做任何事情时,你如何考虑数据收集以广泛地推动科学?
But those two different datasets are actually in many ways quite different. Like PDB, a bunch of very painstakingly constructed protein structures, many of which were an individual PhD thesis, and then maybe the follow-up similar ones came later, which might have been 10 of them for a PhD thesis. Then this—people estimate it's like 13 billion to create the GDB, some very large number. The reason people created the PDB was because each individual protein was independently useful. People didn't create it with the sake of 'we're solving protein structure.' They saw that, 'Oh, this protein we believe is involved in this pathway, let's understand this protein so we can target it,' and so on. Of course, there are some caveats here, but at a high level, a lot of this genomic data was—especially for humans or viruses or bacteria—sequenced for a very specific reason as well, right? Well, it's great that these are useful after the fact, but I wonder if now going forward, especially since with the Virtual Biology Initiative, Biohub's Virtual Biology Initiative is like half a billion dollars, I think. And I'm sure there will be more large initiatives coming from Biohub in the future. You have the chance to be very specific, deliberate, and now collect data for the sake of solving a problem with ML rather than depending on a dataset which was curated, created for some other purpose. So given that new opportunity, how do you do things differently? How do you think about data collection to enable science broadly when you have the option of doing basically anything from first principles?
稍微介绍一下背景。几周前我们宣布了虚拟生物学计划。我们基本上说,我们将内部投资 4 亿美元用于数据创建和技术开发,以规模化数据生成,从而能够增加我们可以同时测量的模态数量。我们还宣布,我们将投入 1 亿美元来推动 Biohub 外部的数据生成工作。所以,我们认为这只是实际所需的一小部分,对吧?但希望基本上是通过做出这一初步承诺,为一些真正在思考这个问题的团队提供启动资金,努力构建所需数据的不同核心领域,这将成为一个催化剂,激励其他团队加入并为此做出贡献。这就是我们真正希望看到的。这个想法是,这是一项基础广泛的努力。所以不仅仅是我们。所以我可以谈谈我对这里需要生成什么数据或可以生成什么数据的看法。但我们也希望与科学界真正合作地推进这件事。所以这部分也是听取科学家们想要什么。所以从我的角度来看,这里有几个关键原则。首先是速度。好吧。所以建立蛋白质数据花了几十年,我们不能等几十年。我们需要弄清楚如何在几年内做到这一点。你看看通用 AI 的发展速度,生物学的限制将从根本上受限于实验科学和数据。所以我们真的需要尽快解决这个差距。所以我认为这是一件关键的事情:看看我们今天可以规模化哪些技术,以开始描绘细胞的信息架构。所以有速度,然后还有泛化的概念。所以回到我之前所说的,我们希望模型能够充当生物学的预言机。它们可以预测你尚未进行的实验。那么我们如何能够做到这一点?我们需要在多种不同的背景下观察多种不同的干预措施。所以这类似于在互联网上训练语言模型或在所有进化多样性上训练蛋白质语言模型的原理。对于细胞生物学来说,这看起来像什么?所以我们必须规模化干预生物学。
A little bit of context. We announced a few weeks ago the Virtual Biology Initiative. We basically said we're going to invest $400 million internally in data creation and development of technology to scale data generation to be able to increase the number of modalities that we can measure simultaneously. We also announced that we're going to commit $100 million to catalyzing efforts outside of Biohub to generate data. And so, we think that's a fraction of what's actually needed to do this, right? But the hope basically is that by making this initial commitment, giving starting funds to some of the groups that are really thinking about this, working to build different core areas of the data that's going to be needed, that that's going to be a catalyst that's going to galvanize other groups to come in and contribute to this. So that's what we really hope to see. The idea is that this is a broad-based effort. So it's not just us. So I can say kind of what my perspective is on what data needs to be generated here or what can be generated. But we also want to approach this really collaboratively with the scientific community. And so part of this is also hearing from scientists what they want. So from my view, there are a few key principles here. The first is speed. Okay. So it took decades to build the data for proteins, and we can't wait decades. We need to figure out how to do this in a couple of years. You look at the rate that general AI is developing, and the limitation in biology is going to be fundamentally limited by experimental science and data. So we really need to work to address that gap as quickly as possible. So I think that's one key thing: looking at what are the technologies that we can scale up today to begin to give this picture of the information architecture of the cell. So there's speed, and then there's also the idea of generalization. So going back to what I was saying before, we want models that can serve as oracles for the biology. They can predict an experiment that you haven't done. And so how are we going to be able to do that? We're going to need to look at a multitude of different interventions in a multitude of different contexts. And so it's kind of similar to the principle of training a language model on the internet or training a protein language model across all of evolutionary diversity. What does that look like for cellular biology? And so we have to scale interventional biology.
这包括像扰动生物学、perturb-seq 测量之类的方法,我们可以结合转录组、成像以及细胞信息层次的其他层面。科学界有很多团队正在研究这类问题,并且已经准备好进行规模化。第二个是空间生物学。我认为这将非常重要,它能帮助我们真正理解细胞在上下文中的状态。孤立地理解细胞并不是我们需要的,也不是目标。细胞是人体中极其复杂系统的一部分。要理解疾病,我们必须理解细胞如何相互作用、它们形成的系统和回路。所以我们需要看到这些。空间生物学正在快速发展,是一个真正准备好规模化的领域。这就是目前可以规模化的方向。Biohub 在过去 10 年里实际上在这些领域做出了开创性的资金承诺。我们资助了像人类细胞图谱这样的项目,构建了 Tabula Sapiens(大型细胞图谱),以及 Cell by Gene(单细胞转录组学数据库)。我们真的希望在此基础上继续发展。我不知道目前最大规模的项目中有多少细胞;大概在 10 亿个细胞左右。所以我们需要再提升几个数量级。这既需要扩展现有技术,也需要下一代技术。因此我们也在资助和支持该领域的努力。我们非常希望更多地关注跨模态。能否同时观察表型、转录层、蛋白质组学层面发生的情况,并将其与基因组和表观遗传状态联系起来?我们希望看到所有这些。所以真正推动技术更快发展,以揭示更多这些联系和生物学信息,并以更可扩展的方式实现。
So that looks like things like perturbation biology, perturb-seq measurements, where we can look at combined transcription, imaging, other layers of the cellular information hierarchy. There are a number of groups across the scientific world that are working on problems like this that are ready to scale. The second is spatial biology. I think that's going to be really important, and that's going to help us to really understand the cell in context. Understanding the cell in isolation is really not what we need; it's not the goal. The cell is part of an incredibly complex system in the body. To understand disease, we have to understand how cells interact, the systems they form, the circuits they form. So we need to see that. Spatial biology is undergoing rapid progress and is an area that's really ready to scale up. That's kind of what can scale now. Biohub has actually over the last 10 years made pioneering funding commitments in those areas. We've funded efforts like the Human Cell Atlas, built Tabula Sapiens which built large cell atlases, and built Cell by Gene, which is a database of single-cell transcriptomics. We're really looking to build on that. I don't know how many cells there are in the largest efforts; we're probably around like a billion cells or something like that today. So we've got to go multiple orders of magnitude from that. That involves scaling the technologies that we have now, but it also involves the next generation of technology. So we're also funding and supporting efforts in that area. We really want to look more at cross-modality. Can you simultaneously see the phenotype, observe the transcriptional layer, understand what's happening proteomically, link that to the genome, and the epigenetic state? We'd like to be able to see all of that. So really pushing technology to be developed faster that can reveal more of those connections and more of that biology, and do that in a more scalable way.
这很有意思,因为我听到的这些想法,很多都是人们已经在考虑的生物学规模化方向。那么,下一项能够实现数据收集的技术是什么?回到生物学苦涩教训的主题,不仅算力和参数有缩放定律,现在数据收集在某种意义上可能也有缩放定律。下一个重大机遇在哪里?所以你提到开发新技术是这项计划的一部分。
It's interesting because when I hear most of those ideas, they're often the things that people already think about in terms of scaling biology. What is the next technology that is going to allow for enabling data collection technology? Going back to the theme of bitter lesson for biology, you don't have just scaling laws on compute and parameters, but now the scaling laws probably in data collection in some meaningful sense. Where are the next big opportunities there? So you're talking about developing new technology as part of this initiative.
是的。所以我认为基本上就是扩展现有技术,能够增加我们可观察的干预措施数量,增加我们可测量的参数数量。所以真正是越来越多维度的测量,并降低成本。更好的基因测序,更好的细胞封装方法,能够同时测量转录组和其他层面。这里有一个有趣的帕累托前沿:如果预算固定,你在改进检测方法上花多少时间,在规模化上花多少时间?你如何权衡?
Yeah. So I think it's basically scaling what we have now, being able to expand the number of interventions that we can look at, expand the number of parameters that we can measure. So really more and more multi-dimensional measurement, and drive down the cost. Better gene sequencing, better ways of encapsulating cells and being able to measure what's happening not just the transcriptome but other layers simultaneously. There's an interesting Pareto frontier there about if you have a fixed budget, how much time do you spend on improving your assay versus how much do you spend on actually scaling it? Where do you weigh in there?
我们必须两者都做,对吧?因为我认为,利用现有技术,通过相对合理的投资,我们肯定能将数据量提升到现在的 10 到 100 倍。但要再提升 10 倍或更多,就需要更多的技术开发。但另一个非常重要的原则是反馈。我认为这将非常关键。你可以将其视为需要发生的一个技术开发层面。我认为现在有很多好的事情正在发生,自动化、灵活机器人技术将加速这一进程,实验设计也是如此。
We have to do both of those things, right? Because I think with current technology, we can definitely get data 10x to 100x where it is today with relatively reasonable investments. But then to get another 10x or more, that's going to require a lot more technology development. But the other really big principle is going to be feedback. And I think that's going to be really critical. And I think you can see that as a layer of technology development that's going to need to occur. And I think there's a lot of great things happening right now, automation, flexible robotics that's going to accelerate where that can go, and experimental design as well.
我们通常会问嘉宾,你会移除哪个瓶颈来解锁一些东西,但我们刚刚花了很长时间讨论这个。我想问这个问题,但换个角度:也许稍微超出你的领域,比如语言建模或供应链,某个瓶颈可能不那么明显,也不是你直接从事的工作,但可能对生物学或 Biohub 的工作有影响。
So we typically ask our guests what is a bottleneck that you would remove that would sort of unlock things, but we just spent a long time talking about that. I want to ask it but I'm going to give it a spin: maybe a little bit outside of your domain, like language modeling or supply chain, something that is a bottleneck that is maybe nonobvious and not directly something that you are working on, but that maybe has impact on the work of biology or Biohub in particular.
这个问题很难回答,因为瓶颈太多了。我经常想到的一个是算力,但我觉得这很明显。目前它在很多方面是整个 AI 的瓶颈。特别是因为我们正在训练这些大规模模型,我们总是关注算力。我认为我们同时受到数据和算力的限制。作为一个从事生物学的团队,我们拥有令人难以置信的算力资源,但就像所有从事 AI 的团队一样,限制因素就是你拥有多少算力。
It's a hard question to answer because there are just so many bottlenecks. The one that I always think about is compute, but I think that's a pretty obvious one. It's the bottleneck for all of AI in many ways right now. Especially because we're training these large scale models, we're always focused on compute. And I think we're limited both by the data and compute. We're in a position where we have incredible compute resources for a team working in biology, but like all teams working in AI right now, the limit is just how much compute power you have.
所以如果你能把算力提升 100 倍,你认为 ESMC 会好得多吗?
So if you could 100x your compute, you think that ESMC would be way better?
它肯定会好得多。我们还需要扩展数据。所以这两件事必须同时进行。
It would definitely be way better. We also need to scale data. So both of those things would have to happen in tandem.
你基本上已经用尽了目前可用的数据吗?
Have you basically exhausted what's available right now for data?
我不这么认为。不,我不这么认为。
I don't think so. No, I don't think so.
现有的那些大型数据集,或者你可以……
The large data sets out there or you could...
嗯,更多参数。所以我们把 ESM C 训练到了 60 亿参数。
Well, more parameters, you know. So we trained ESM C up to six billion parameters.
但我是说在可用数据方面,你是否已经用尽了大部分公开数据?
But I'm saying in terms of data available, like have you exhausted most of what's publicly available in terms of like...
不,还没有。我们刚刚构建的图谱实际上比 ESMC 训练所用的序列和结构更多。所以肯定还有一点空间。
No, not yet. And the atlas that we just built actually has more sequences and structures than ESMC was trained on. So definitely have a little room to go.
那么这是一个数量级的跃升还是两倍?具体是怎样的?
So is that an order of magnitude jump or twice as much? How does that work?
是的,我认为 ESMC 是在大约 10 亿条序列上训练的。所以很可能有大约 1000 亿条序列。其中很多是高度冗余的。
Yeah, I think ESMC is trained on say order a billion sequences. So there's definitely probably order 100 billion sequences. A lot of them are largely redundant.
1000 亿。
100 billion.
是的。
Yeah.
哦,好的。为了得到那十亿,你从 68 亿中筛选出来。那么,如果对那 1000 亿进行类似的聚类并找出独特的序列,你认为结果会是多少?
Oh, okay. To get that billion, you whittle down from six billion 6.8 billion. Right. So of those 100 billion, if you were to similarly cluster and find unique ones, what do you think? Where do you think it would land?
序列实际上并不冗余,对吧?这取决于你对冗余的定义,因为我认为从微小的遗传变异中可以学到大量信息。这些变异在非常精细的层面上揭示了蛋白质结构和功能的基本决定因素。因此,当我们思考蛋白质空间时,拥有跨越广泛蛋白质家族的巨大序列多样性对于结构预测能力的出现至关重要。我认为,大的多样性训练模型去理解并发展结构的表征,但要发展功能的表征,这些微小的变异才是关键。模型还没有被训练到那种深度,去真正理解这些微小但关键的序列模式。一个单点突变就足以破坏蛋白质的功能。
The sequences aren't actually redundant, right? It really depends on what you mean by redundancy because there's I think a tremendous amount that you can learn from small genetic variations, right? Cuz these are these are really revealing of you know kind of the the very basic determinants of protein structure and function at a very fine level. So I think that you know as we think about protein space you know having a vast diversity of sequences across a wide range of protein families is you know really critical for the emergence of this kind of structure prediction capability because I I think kind of large diversity is what trains the model to understand to develop a representation of structure but I actually think that to develop a representation of function it's these very small variations that are important and so I I do think that there is probably a lot more you know it's like the models haven't yet been trained at that level of kind of just like really deep understanding of these very small but critical patterns in and sequence I mean a single a single mutation is enough to destroy the the function of a protein
所以你可以设想,实际上用全部 68 亿序列重新训练,其他一切保持不变,但……
So you could conceivably actually take all 6.8 billion of those retrain everything's the same but
是的,你可以训练比那更多的数据,我的意思是那已经是聚类后的结果了。
Yeah you could train on more than that there even I mean that's kind of clustered down so
也许问题在于,在多大程度上你不会遇到收益递减定律?听起来你有计划开发 ESM4 或 EMD 之类的模型,但我只是想知道,是否在某个点上你会耗尽预训练数据?人们常谈论耗尽预训练数据。是的,在某个点上。
Yeah maybe the question is how far wouldn't you hit the law of diminishing returns here I mean it sounds like you have plans for an ESM4 or an EMD or whatever you want to developing yeah but I'm just wondering it's you know at some point is this actually something that you could exhaust you know people talk about exhausting the pre-training data in on Yeah. At some point. Yeah. At some point.
是的。但这并不是你在未来几年内可以想象做到的事情。即使你没有耗尽数据,对于你想要预测的应用,你也会遇到很多收益递减,也许你的资源更适合花在其他地方。
Yeah. But it's not actually something you could conceivably imagine imagine doing in the next few years or even if you don't exhaust it, you hit a lot of diminishing returns for you know the applications that you're trying to predict here where maybe your resources are better spent somewhere else.
这基本上是一个经验问题,对吧?这确实是一个经验问题。
I mean it's basically it's an empirical question, right? It's truly an empirical question.
所以我们就是不知道。我的意思是,对于 ESM2,我们不确定,因为存在一些收益递减。
And so we just we just don't know, you I mean with the SM2 we weren't sure because there were some diminishing
对于 ESMC,现在没有收益递减了,所以你可以看看那个,从缩放定律外推,有足够的数据来训练下一个模型。
Returns with the SMC you know now now there aren't right so you can kind of look look at that extrapolate from from the scaling law there and you know there is enough data to train that next next model so
另一个我们常问的问题是,有什么行动号召吗?如果听众想参与、被雇佣或构建东西,你希望他们做什么?
And the other question that we usually ask is any call to action or what what do you want people to go take action on if the listeners want to get involved, get hired, get build things. What would you ask people to do?
我们刚刚宣布——或者说,在这期播客发布时,我们已经宣布了 ESMC 和这个蛋白质生物学世界模型。它将开源,采用 MIT 许可证,我们希望人们使用它。我们希望这能成为解锁科学的工具。我们很乐意合作。我们有一个团队在从事这方面的工作,我们希望听到人们的声音,了解我们可以构建什么来加速他们的科学研究。
Well, we just announced or or I should say we are we are going at the time that this this podcast comes out, we will have announced ESMC and this this world model for protein biology. Um, it's going to be open source. It's it's going to be MIT licensed and we want people to use it. You know, we want we want this to be a tool that can unlock science. We're excited to collaborate. We have a team that that that works on that and we want to hear from people and understand, you know, what what we can build that can help to accelerate their science.
是的,我们可能会在这个频道上举办一个演示或论文俱乐部之类的活动。敬请期待。
Yeah, we might have a um a demo slashpaper club of some sort on this channel. So stay tuned.
是的,敬请期待。我们会邀请你和你的团队,谁有空都可以来。一旦论文最终预印本发布,我们会在 Way in Space 论文俱乐部花一个小时重点讨论它。
Yeah, stay tuned for that. We'll we'll invite you and your team um whoever can make it. We'll feature this paper once it's in final preprint form and spend some time on it for an hour on the way in space paper club.
好的。感谢你和我们聊天。
Yeah. Uh, thanks for chatting with us.
太棒了。很高兴见到你们。
Awesome. Yeah, great to meet you guys.