Accelerating Biology with AI: Mark Zuckerberg, Priscilla Chan, and Alex Rives on Biohub
打开互动全文版(中英对照 + 朗读 + 问答)→马克·扎克伯格、普莉希拉·陈和亚历克斯·里夫斯讨论 Biohub 如何利用开源工具和 AI 对蛋白质、细胞及整个生物系统进行建模,以加速科学进步。
Mark Zuckerberg, Priscilla Chan, and Alex Rives discuss how Biohub uses open-source tools and AI to model proteins, cells, and whole biological systems, aiming to accelerate scientific progress.
今天在 No Priors,我们请到了马克·扎克伯格、普莉希拉·陈和亚历克斯·里夫斯。我们将讨论 Biohub 以及他们现在开始大规模应用 AI 来构建细胞世界模型和生物学中不同层面交互的各种努力。马克、普莉希拉,谢谢你们来做客。
Today on No Priors, we're joined by Mark Zuckerberg, Priscilla Chan, and Alex Reeves. We'll be talking about Biohub and all their various efforts to now start applying AI at scale to do world models of cells and different levels of interactions across biology. Mark, Priscilla, thank you for doing this.
谢谢邀请我们。这很有趣。
Yeah, thanks for having us. This was fun.
亚历克斯,恭喜你有了新使命。
Alex, congratulations on new missions.
谢谢。
Thank you.
你们把 Biohub 作为主要的慈善事业,并承诺投入 5 亿美元用于这个虚拟生物学计划。能跟我们说说为什么这么做,以及你们是如何从“我们应该资助这个”变成“这就是我们”的?
You guys made Biohub your primary philanthropic effort and then committed $500 million to this virtual biology initiative. Can you tell us a little bit about why do that and how did you go from we should fund this to this is like who we are?
所以,对于目前形式的 Biohub,我们感到非常兴奋。我们认为它非常适合我们是谁、我们能带来什么以及我们能共同实现什么。但这项工作始于 10 年前,当时我们在思考如何回馈社会,马克想建立一个能在本世纪末治愈、预防和管理所有疾病的组织。我们和科学家们开了一系列搞笑的会议,那些著名的诺贝尔奖得主科学家们都在嘲笑我们。那是你的起点吗?我们就要治愈所有疾病。
So, Biohub in its current form, we're super excited about. We feel like it's a really good fit for who we are and what we bring to the table and what we can achieve together. But this work started 10 years ago when we were thinking about how can we give back and Mark had wanted to build an organization that could cure, prevent and manage all disease by the end of the century. And we had a series of hilarious meetings with scientists that like famous Nobel Prize winning scientists were just laughing at us. Is that was that your starting line? We're just going to cure all disease.
不,不。要说明的是,我们不认为我们会是治愈疾病的人。我们的目标一直是构建能够加速整个科学领域的工具。这样,整个科学界才能共同治愈所有疾病。但我仍然认为本世纪末这个目标有点牵强。现在,我觉得它太保守了。所以我们一直说,“好吧,我们有过一系列有趣、尴尬、有教育意义的对话,我们问,好吧,但为什么?你为什么认为这不可能?”然后,你知道,作为房间里的人,我们只是说,“嗯,我不知道为什么。你告诉我。”最后,人们说,“好吧,如果你真想知道的话。”我们说,“不,我们确实想知道。这似乎很重要。”你知道,他们说我们各自为政,发表的信息得不到共享,被长期封锁,我们也没有工具。他们举例说,一个实验室的博士后构建了一个很好的工具,但它只存在于他们的电脑上,当他们毕业时,工具就没了。我们听到的是,很难构建共享工具来加速科学,构建共享知识库来快速推动科学发展。这就是我们开始思考的地方:如果这些是问题,我们能贡献什么?
No. No. And to be clear, we don't think that we're going to be the ones curing the diseases. Our goal was always to build tools that could accelerate the whole scientific field. That way, the scientific field collectively could cure all the diseases. But still, I thought that by the end of the century was a stretch. Now, I think it's like too conservative. And so we kept being like, "Okay, well, we had these series of funny, awkward, educational conversations where we're like, okay, but like why? Like why do you think it's impossible?" And like, you know, just being the person in the room is just like, "Well, I don't know why. You tell me." Finally, we got people to like they're like, "Fine, if you really must know." And we're like, "No, we do. It seems important." It's you know they were like well we work in silos and when you publish information doesn't get shared it gets locked up for long periods of time and we don't have tooling you know they gave the example of like we build a great tool by one postdoc in a lab and it lives on their computer and when they graduate the tool is gone and they just it was what we heard was very hard to build shared tools to move science faster build the shared knowledge base to quickly move science faster and that's sort of where we began in thinking about okay like if those are the problems like what can we contribute.
是的,我的意思是,最初的 Biohub 模式基本上是通过将工程师和科学家聚集在一起,专注于长期工具开发,它确实奏效了。我们从 CZI 开始做了很多不同的事情,随着时间的推移,我们觉得科学部分确实在起作用,于是我们不断投入更多,直到现在它基本上是我们做的主要事情。我们已经将最初的旧金山 Biohub 扩展到几个地方,现在有纽约、芝加哥。真正的重点和统一主题是虚拟生物学计划,围绕利用能够生成的独特数据集来有效建模,从最小的蛋白质开始,最终到细胞和整个生物系统。这就是我们演变的方式:这在一定程度上是一个 AI 问题,你想建立一个前沿 AI 实验室,但你需要将其与前沿生物学工作结合起来,这样才能理解和获取构建这些模型所需的数据。因为与语言模型不同,互联网上有大量数据,但生物学并非如此。我的意思是,显然存在许多不同的数据集,是学术界和科学家几十年来生成的,但我想放入其中的很多东西并不存在,对吧?比如你想可视化人们以前无法看到的东西,这就是我们做成像工作的原因。你想记录体内发生的事情,这就是我们做细胞工程工作的原因,对吧?你想以以前不可能的方式测量炎症等,这就是芝加哥 Biohub 专注于构建这类设备的原因。这将从根本上创造新型数据集,从而允许新型模型。我认为这非常令人兴奋。回到你所说的,如果科学领域主要需要工具开发,那么这将赋能整个领域的科学家更快地工作,我们认为通过这种对工具开发的长期关注,我们可以提供这一点。
Yeah, I mean so the original Biohub model was basically focus on long-term tool development by bringing together engineers and scientists across multiple universities to focus on long-term tool development and it basically worked and you know we started off with CZI doing a number of different things and I think over time we just felt like okay the science piece is really working and we just kept on investing more and more and more in it until now it is basically the primary and main thing that we're doing and we've expanded the original San Francisco Biohub to a handful now at this point there's New York there's Chicago the real focus and the unifying theme at this point is the virtual biology initiative around taking the unique data sets that are able to be generated in order to model effectively starting with the smallest pieces of proteins but then eventually cells and whole biological systems. But that's kind of how we've evolved is, you know, this idea that some of this is an AI problem and you want to build a frontier AI lab, but you need to couple that with a frontier biology effort that can do the work of being able to understand and get the data that you need to actually be able to build these models. Because unlike language models, well, there's just like a lot of data out there on the internet. That's not really the case with biology. I mean, there are obviously a bunch of different data sets that exist that academia and scientists have generated over the decades, but a lot of the stuff that I think we want to put into this, it doesn't exist, right? It's like you want to be able to visualize things that people haven't been able to see before, which is why we're doing the imaging work. You want to be able to record things that are going on inside the body, which is why we're doing the kind of cellular engineering work, right? You want to be able to measure things like inflammation in ways that haven't been possible which is why the Chicago Biohub is focused on building those kind of devices and being able to do that and that will fundamentally create new types of data sets that will allow new types of models and I think is just a very exciting thing that going back to what you're saying if the scientific field primarily needs kind of tool development that now is going to empower scientists across the field to be able to do their work faster that's what we think we can provide through this kind of long-term focus on tool development.
但我认为,从我们起步到与 Alex 合作的工作之间有一条有趣的线索。我们的第一个 RFA 就是围绕单细胞测序。我们想观察单个细胞中转录的 RNA。这虽然可行,但当时对细胞如何表达 DNA 的理解仍处于早期阶段。最初我们只是资助方法,让人们描述如何操作,以便其他人可以共享这些方法。后来这演变为我们资助人类细胞图谱,它现在是最大的单细胞转录组数据库之一。科学家们越来越难以注释这些数据,所以我们构建了 Cell by Gene,一个简单的注释工具。一个社区围绕它形成,并开始贡献更多数据,而这些数据并非我们创建或资助的。现在 Cell by Gene 是一个知识库,许多基于转录组的模型都依赖它,科学界经常使用。但总有人批评说:‘这不过是集邮,收集数据碎片,我们无法从中提取科学知识和智慧。’我们一度无法反驳。后来,当大语言模型成为热门话题时,你可以想象我们有多高兴——它们能理解海量数据。对我来说,问题是:我们能否真正理解生物学的工作原理?将其从发现驱动的科学转变为工程驱动的科学,从而系统性地理解活细胞如何工作以及为何出错。当我们看到那个时刻时,我们认为:‘就是它了。这里可能发生真正重大的事情。’Alex,你从 Metaphair 起步,但你已经走上了这条路——你在 Evolutionary Scale 组建了团队,筹集了风险投资,并在模型上取得了进展。Mark 和 Priscilla 是怎么说服你,说这是实现使命的正确道路的?
But I think there's a fun throughline from where we started to our work with Alex now. Our very first RFA here was around single-cell sequencing. We wanted to look at the RNA transcribed in individual cells. That was possible but still early in understanding how different cells express their DNA. At the beginning, we were just funding methods, getting people to describe how to do it so others could share that methodology. Then that became us funding the Human Cell Atlas, which is now one of the largest databases of single-cell transcriptomes. It was getting hard for scientists to annotate the data, so we built Cell by Gene, a simple annotation tool. A community formed around it and started contributing more data that we had nothing to do with creating or funding. Now Cell by Gene is a corpus of knowledge that many transcriptomic-based models are based on, used regularly by the scientific community. But there were always critiques: 'This is just stamp collecting, gathering bits of data, and we won't be able to pull scientific knowledge and wisdom out of it.' We didn't have an answer for a while. Then imagine our delight when large language models became a huge topic of conversation—they could make sense of large amounts of data. For me, it was: What if we could actually understand how biology works? Move it from a discovery-based science to an engineering-based science, where we could systematically understand how living cells work and why things go wrong. When we saw that moment, we thought, 'This is it. Something really big could happen here.' Alex, you started at Metaphair, but you were on the path—you assembled a team at Evolutionary Scale, raised venture, and made progress in your models. What was the pitch from Mark and Priscilla that said that's the right way to go after the mission?
嗯,对我来说,真正关键的时刻是我意识到他们确实将这件事视为前沿 AI 与前沿生物学的融合。我早已确信,这确实是一个刚刚开始的科学新时代——人工智能将带来无限可能。我们身处大规模信息理论的时代,拥有能够预测下一个词并从中学习世界模型的系统。它们能从数据中学习生物学。所以很明显,要打造下一个时代的下一所机构,你需要前沿人工智能和前沿生物学。你需要让这两者形成反馈,让模型从生物学中学习。你还需要合适的规模和合适的人才。这感觉就是正确的道路。
Well, I think for me it was really the moment when I understood that they really saw this as an integration of frontier AI and frontier biology. I had developed conviction that this is really a new era of science that's just beginning—what's going to be possible with artificial intelligence. We're in the age of information theory at scale, and we have these systems that can predict the next token and learn world models from that. They can learn biology from the data. So it was really clear that to build the next institution for the next era, you would need frontier artificial intelligence and frontier biology. You would need to put those things in feedback, having models that learn from the biology. And you need the right scale and the right people. This felt like the way to do that.
你们一直在研究各种不同的模型。生物学早期的一些突破,比如 AlphaFold——一个 Google 模型,它展示了大规模蛋白质折叠的可能性,而人们之前并未意识到这是可行的,这发生在大型 Transformer 浪潮之前。你们在不同尺度上研究不同的事物:增量分子建模、蛋白质折叠、基于细胞的研究,以及思考生物学中更大尺度的系统。你认为这如何从微观延伸到宏观?你提到从构建模块开始,逐步向上,但建模细胞行为与建模蛋白质折叠非常不同。数据不同,建模也不同。你认为这一切都是相似的——只是数据和训练——还是说处理这些系统的方式存在差异?
There are a variety of different models you've been working on. Some of the earliest breakthroughs in biology were things like AlphaFold—a Google model that showed protein folding at scale in a way people didn't realize was tractable, pre-dating the big transformer wave. You're working on different things at different scales: incremental molecular modeling, protein folding, cell-based stuff, and thinking about larger-scale systems in biology. How well do you think that extends from micro to macro? You mentioned starting with building blocks and building up, but modeling cellular behavior is very different from modeling protein folding. The data is different, the modeling is different. Do you think it's all similar—just data and training—or are there differences in how you deal with these systems?
我的意思是,可能确实存在一些差异。每一层最终都会在性质上有所不同。但你需要理解蛋白质相互作用才能理解细胞如何工作。不先理解蛋白质建模,你就无法直接研究细胞。如果你想理解免疫系统或一群细胞如何相互作用,不先理解细胞就很难。你可以在非常高的抽象层次上模拟一个系统,但如果你真想理解它如何工作,你就需要分层构建模拟。这基本上就是我们的方法,从构建模块和蛋白质开始。每种建模技术都需要收集不同类型的数据。我认为我们会在各个层面看到进步。但战略的一个重要部分是这种分层构建的观点。我们独特的一点是,我们非常刻意地将 AI 工作和湿实验室工作视为一个整体,并做了大量工作将它们融合在一起。巧妙之处在于,我们可以收集数据,帮助我们连接各个层级。你可以观察细胞内的转录组学及其定位。我们可以观察半透明的斑马鱼,看大脑发育过程中不同细胞的变化。我们有传感器可以观察细胞间通讯和不同分子。因此,我们可以策略性地选择实验和数据类型,以跨越这些层级,提供连接组织,驱动建模的魔力。
I mean, there are probably some differences. Each layer is going to end up being somewhat qualitatively different. But you need to understand protein interactions to understand how cells work. You can't go straight to cells without understanding protein modeling. And if you're trying to understand something like the immune system or how a bunch of cells interact, it's tough without first understanding cells. You might be able to simulate a system at a very high level of abstraction, but if you really want to understand how it works, you want to build simulations at each level hierarchically. That's basically the approach we're going through, starting with the building blocks and proteins. There will be different types of data to collect for each modeling technique. I think we'll see advances across the board. But a big part of the strategy is this view that you need to build it up hierarchically. One unique thing about us is we were very intentional that the AI efforts and the wet lab efforts were a single effort, and we've done a lot to bring them together. The neat thing is we can pull and gather data that helps us connect across the hierarchy. You can look at transcriptomics with space within a cell and where it's localizing. We can look at translucent zebrafish and see development across different cells when the brain develops. We have sensors that allow us to look at cell-cell communication and different molecules. So we can be strategic about the types of experiments and data we collect to bridge across these levels, providing connective tissue that drives the modeling magic.
我之所以问这个问题,是因为我以前是生物学家。我有生物学博士学位,在湿实验室工作了近十年。一直缺少的是生物学不同层次之间的整合性。发育生物学家各自为政,分子生物学家做不同的实验。所以我对此很好奇。
The reason I asked the question by the way is I used to be a biologist. I have a PhD in biology and I worked in wet labs for almost a decade. One of the things that was always lacking was this integrative nature across the different layers of biology. The developmental biologists would work on their own, the molecular biologists would be doing different experiments. So that's what I was curious about.
通常生物学有还原论观点和系统观点,这两类人并没有真正深入合作。所以你们所做的一个令人兴奋之处就是如何弥合这一鸿沟。这也是我提问的基础。
Typically there's a reductionist view of biology and a systems view, and those people didn't really work together deeply. So one of the exciting things about what you're doing is how you're bridging that. That was the basis for the question.
是的。我想补充一点,我们正处于生物学信息理论的时代。生物学有不同层次的复杂性和层级,每个层次都由更低层次构成。要获得更完整的描述,以及能够泛化并数字化回答实验问题的系统,就需要在每个层次上建立正确的建模基础。我们能做到的独特之处在于,在每个不同层次构建信息,收集这些信息以及连接点,并且以能够揭示底层信息架构的规模去做。这对于构建能够回答新实验问题的数字表征至关重要。
Yeah. And if I could add something there, I think we're in the age of information theory in biology. There are levels of complexity and hierarchy in biology, and each level is constituted by the lower levels. As you want a more complete description and systems that can generalize and answer experimental questions digitally, you need the right basis for modeling at every level. What's unique about what we can do is to build information at each of these different layers, collect them, collect those connection points, but also do it at the scale that will reveal the underlying information architecture. That's going to be critical to build digital representations that can answer new experimental questions.
这项努力最激励我的事情之一是 Priscilla 所说的:我们对生物学有太多不了解。如果我们能了解呢?这与其他 AI 问题非常不同,那些问题中我们复制人类行为,而且很多数据都在互联网上。即使不假装理解所有人类行为,也能预测很多。我认为你们发布中最有趣的一点是你提到的机制可解释性——我们能否从模型认为正在发生的事情中提取新知识?你能谈谈这个吗?
One of the things that inspires me most about this effort is what Priscilla said: there's so much we don't understand about biology. What if we could? That's very different from other AI problems where we replicate human behavior, and a lot of that data is on the internet. Without pretending to understand all human behavior, you can predict a lot of it. I thought one of the most interesting things in your release was the mechanistic interpretability stuff you alluded to—can we extract new knowledge from what the model believes is happening? Can you talk a little bit about that?
是的,我对此非常兴奋。在机制可解释性方面,传统上它被应用于大型语言模型,目的是理解表征空间、如何计算事物,以及是否与我们直观的世界理解相联系。已经开发出了丰富的工具来提出这些问题。对于生物学,我们训练的一类模型是蛋白质语言模型——在蛋白质的代码上训练。它们学到的任何关于生物学的东西都是涌现出来的。我们已经看到它们能学习生物结构和功能,这是从词元预测训练任务中涌现出来的。当我们考虑这些模型中的机制可解释性时,我们实际上是在看到未知,因为模型已经在数十亿个蛋白质序列上训练过,包括已知和未知的生物学。然而,它们发展出的表征对应于几个世纪以来建立的还原论生物学图景。所以你可以将我们完全不了解的蛋白质与了解的蛋白质联系起来,因为模型表征空间中存在一个底层的结构语法将它们连接起来。
Yeah, I'm really excited about that. In mechanistic interpretability, traditionally it's been applied to large language models with the goal of understanding the representation space, how it computes things, and whether that connects to our intuitive understanding of the world. There's a rich toolkit developed to ask those questions. For biology, one class of models we train are protein language models—trained on the codes of proteins. Anything they learn about biology is emergent. We've seen they can learn biological structure and function, emergent from the token prediction training task. When we think about mechanistic interpretability in those models, we're really seeing the unknown because the models have been trained on billions of protein sequences, both known and unknown biology. Yet they develop representations that correspond to the reductive picture of biology built over centuries. So you can connect the dots between proteins we know nothing about and those we do, because there's an underlying structural grammar linking them in the model's representation space.
极端情况下,我们可能会理解以前不了解的身体系统,或者新疗法的作用机制,因为我们可以询问模型并探究那个表征。
At the extreme, it could be we're going to understand systems in the body that we didn't before, or the mechanism of action for a new treatment, because we can ask the model and interrogate that representation.
没错。希望在于你真正了解它做出预测的底层基础,从而打开黑箱,理解模型所代表的生物学。
That's right. The hope is that you really learn the underlying basis for how it's making predictions, so you open up the black box and can understand the biology the model is representing.
所以帮朋友问一下——你们都相信风险投资支持的公司是影响世界的一种方式。收集斑马鱼数据、数据范围、湿实验室工作或规模是怎样的?是什么让这项工作更适合这个大型非营利生态系统努力,而不是风险投资支持的公司?
So asking for a friend—you all believe in venture-backed companies as a way to have impact on the world. What was it like collecting data on zebra fish, the span of the data, the wet lab work, or just the scale? What makes this a better fit for this big nonprofit ecosystem effort versus a venture-backed company?
嗯,我认为我们只是想给整个科学界提供工具。为了产生最大的影响,其实并不清楚如果我们想的话,是否不能把它当作一门生意来运营。我只是认为,通过将其作为开源项目来做,更快地让更多科学家使用,我们会产生更大的影响。这就是我们的方法。
Well, I think we just want to give tools to the whole scientific community. In order to have the biggest impact, it's not actually clear that we couldn't run it as a business if we wanted to. I just think we'll have a bigger impact by getting this in more scientists' hands quicker by doing it as open source projects instead. That's kind of the approach.
我不确定……我的意思是,显然你们是以非营利公司的形式在做。在那之前做模型时会遇到一些问题。你需要筹集大量资金来构建算力集群。我认为在很多方面,数据实际上更是一个制约因素。因为如果你看看这些模型与语言模型的规模相比,它们更小,但更小是因为数据量更少。要获取数据,并不是说哪里有工厂你可以付钱生产数据。你实际上需要发明新的科学方法才能做到,比如我们在纽约做的那种细胞工程,或者芝加哥的那些设备。这就是为什么当我们谈论前沿生物学和前沿 AI 这个概念时,前沿生物学意味着你需要做真正的科学来推进不同的生物学方法,以便能够观察到那些产生数据、进而输入模型的东西。所以这不像现成的东西你可以直接创造。那是一个相当大的工程。我不知道有多少类似的事情是以生物技术公司的形式完成的。我认为这只是我们正在做的事情的雄心规模,以及我们承诺做这件事的时间跨度。我觉得部分理论是,如果你要构建如此复杂的工具,你大概需要 10 到 15 年的时间跨度来推进这些工作。还有所需的资本规模。我想没有规定说你不能作为一个资金极其充裕的初创公司来做,但我认为这样更合理。而且这在战略上也简化了,不需要考虑如何用不同的东西赚钱。我们只是想把模型交到人们手中。我们以开源形式发布它们。我认为这是一件非常有价值的事情。再说一次,理论不是我们要治愈疾病。我们不会。而是我们想帮助加速整个科学领域的进步步伐。
I'm not sure that I mean obviously you were doing it as a nonprofit company. A bunch of the modeling before then you run into certain issues. I mean you have to raise a large amount of money in order to build compute clusters. I think in a lot of ways the data is actually even more of a constraint. Because if you look at the scale of these models compared to language models, they're smaller, but they're smaller because the amount of data is less. In order to get the data, it's not just like there's some factory somewhere that you can pay to produce the data. You actually need to invent new novel scientific approaches to be able to do, for example, the type of cellular engineering we're doing in New York or the types of devices in Chicago. Which is why when we're talking about this concept of frontier biology and frontier AI, the frontier biology is you need to do real science to advance different biological methods in order to be able to observe the things that create the data that go to the model. So it's not just like an off-the-shelf thing that you can create. Now, that's a pretty big effort. I don't know that there are that many things like that that are done as biotechs. I think it's just the scale of the ambition of what we're doing, the horizon over which we're committed to doing it. I think part of the theory is like if you're building tools that are this complicated, you kind of want to have a 10 to 15 year time horizon on building out these efforts. And then the scale of capital required. I guess there's no rule that said that you couldn't do it as an incredibly well-funded startup, but I think that this just made more sense. And then it also is simplifying strategically to not have to think about how you're going to make money with the different things. We just want to get the models in people's hands. We release them as open source. I think that that's a very valuable thing to do. And again, the theory isn't that we're going to cure the diseases. We're not. It's that we want to help accelerate the pace of progress for the whole scientific field.
作为这里最没有赚钱经验的人,我想说,我们工作的这种中立非营利性质实际上有助于吸引更多人加入这项事业。而要真正实现理解人类生物学全貌并治愈、预防、管理所有疾病的使命,你确实需要整个学术生物技术产业团结起来,以某种统一的方式共同努力。部分原因是外面有很多人才,将任何人才排除在努力之外都是无益的。而且疾病有非常长的长尾。有常见病,即使是常见病,我认为如果你拆解心脏病、癌症、神经退行性疾病,甚至拆解痴呆症或抑郁症,都有很多很多很多子类别,变得越来越小众。这还不算罕见病的漫长长尾。这些疾病常常被忽视,在我们寻找最有效的方式影响多数人生活时,它们不会被带动。但如果你分散努力,把工具交到很多人手中,你就会开始遇到这样的人:你知道吗,我对脊髓性肌萎缩症超级感兴趣,这是我非常关心的事情。如果你把工具交到那个人手中,他们就能取得进展。在某种程度上,如果你必须集中精力并下大赌注,你可能不会这么做,因为那只是一个 niche 的、个体的小群体疾病,但如果我们能理解那个疾病过程,它反过来会帮助我们解锁更多关于人体如何运作的知识。
As the person least experienced with making money here, I would say that the sort of neutral nonprofit nature of our work actually helps harness more people to enter this effort. And to actually achieve the mission of understanding the totality of human biology and to cure, prevent, manage all disease, you actually do need the entire academic biotech industry to come together and to work on this in a sort of unified way. In part because there's a lot of talent out there and it's not helpful to exclude any talent from the effort. And there's a super long tail of diseases. There are the common ones, and even the common ones I think if you unbundle heart disease, cancer, neurodegenerative diseases, even if you unbundle like dementia or depression, there are many, many, many subcategories that become more and more niche. And that's not even looking at the long long tail of rare diseases. Those often get orphaned and don't get brought along when we're looking at what the most efficient way to impact the lives of many. But if you sort of decentralize the effort and put the tools in many people's hands, you start getting people who are like, you know what, I am super interested in spinal muscular atrophy and that's something I care deeply about. And if you put the tools in that person's hands, they're going to be able to make progress. In a way, if you had to focus your efforts and make big bets, you probably wouldn't because it's just a niche individual small group disease that actually will in turn, if we can understand that disease process, helps us unlock knowledge about a lot more about how the human body works.
关于这项工作会首先影响哪些疾病领域,你有什么想法或预测吗?我知道这些事情很难预测,但仅就工作的性质和模型的性质而言,短期到中期内,你最乐观的领域有哪些?
Do you have any thoughts or predictions in terms of what disease areas this work will impact first? I know it's very hard to be predictive about these things, but just given the nature of the work and the nature of the models, are there areas you're most optimistic about in the short to medium term?
至少我不是这么想的。我的想法是,我们想理解生物学是如何运作的。理想的世界是,你会说,我理解这个人的遗传学。所以我想从个体层面思考人。我想理解这个人的遗传学。我想理解他们患不同疾病的风险。我想理解比如基因变异、蛋白质和疾病过程之间的机制联系。因为如果你理解了那个链条,你就可以设计一个蛋白质,设计一种为他们量身定制的药物,并实际进行干预。而现在,我相信我们都有过生病的经历,如果你有某种哪怕稍微不标准的病症,你会去 PubMed,查一篇论文,看补充材料,然后开始看方法,你会想,这篇论文里有我吗?我们只是在猜测。我们真的没有机制上的理解。我们说,好吧,你有点像我们研究过的那些人,这种药有点影响我们认为涉及的通路。试试看会发生什么。时间过去了,有时有效,有时无效。所以我的目标是能够将个体作为个体来治疗,理解机制并能够干预。不同的疾病在填补整个链条的不同阶段。所以对于某些疾病,你只想理解哪些基因变异实际上导致疾病,哪些不导致。这本身就可以给患者带来巨大的力量。如果除此之外,有些疾病我们理解了链条,但就是无法干预并改变特定的蛋白质功能,那也非常令人兴奋。比如,如果我们能设计一种蛋白质来实际改变生理机能,那么我们就能真正治愈某人。但对我来说,这与理解人们最初如何生病同样令人兴奋。
That's actually not how I think about it at least. The way I think about it is like we want to understand how biology works. The ideal world is you would say I understand the genetics of this person. So I want to think about people at the individual level. I want to understand the genetics of this person. I want to understand the risks they have to different illnesses. I want to understand the mechanistic connection between say a gene variant, a protein, and a disease process. Because if you understand that through chain, then you can design a protein, design a drug bespoke to them, and actually make an intervention. And right now, and I'm sure we've all had experiences being sick, and if you have something that's even remotely non-standard, you go into PubMed, you look up a paper, you look up the supplement, and then you start going through the methods, and you're like, am I represented in this paper? And we're just making guesses. We really have no mechanistic understanding. We're saying like, okay, you're kind of like these people that we studied, and this drug kind of impacts the pathway that we think is implicated. Let's try and see if anything happens. And time passes and sometimes it works and sometimes it doesn't. So my goal is to be able to treat the individual as an individual, understand the mechanisms and be able to intervene. And there are different diseases that are at different stages of filling out that whole through line. And so for some diseases, you just want to understand which gene variants actually cause disease and which don't. And that in itself can be super empowering to patients. And if beyond that there are some diseases where we understand the chain we just can't intervene and change a specific protein function. That's super exciting too. Like if we could design a protein to actually change the physiology, then we can actually cure someone. But to me, that is just as exciting as understanding contributing to our understanding of how someone gets sick in the first place.
是的。这是一个非常令人兴奋的愿景,因为你基本上是在说,你可以带来可推广的工具,为每个人提供非常个性化的东西。是的。这就是这种方法的力量——你构建的这些大模型可以应用于任何地方。我知道你之前提到,你打算在一百年内尝试治愈、预防所有疾病,而且你说,考虑到 AI 的所有进步,实际上可能更快。
Yeah. That's a very exciting vision because you're basically saying you can bring generalizable tools to provide very personalized things for each individual person. Yes. And that's the power of the approach is you have these big models that you build that can then apply anywhere. I know that you mentioned earlier that you were going to try and cure prevent all diseases within a hundred years and you mentioned hey it could actually be sooner now given all the advances in AI.
你对我们何时会更接近那个目标有什么想法吗?
Do you have some thought of when we think we'll be closer to that goal or some?
我的意思是,我乐观地认为它会更快。我认为复杂之处在于这是一个动态系统。所以如果你解决了一个问题,显然还会有未来的问题需要处理。我不认为目前我们意识到的这些问题是唯一需要解决的。但我认为 AI 的进展确实非常令人兴奋。另一件我想说的是,我们更关注系统而非特定疾病。例如,一个似乎非常重要的领域是炎症。我们讨论过很多次。这是芝加哥生物中心的一个重点。关于炎症有很多数据。很明显它与许多不同疾病相关。但与其研究特定疾病,我们认为通过更广泛地理解炎症,可以让其他公司利用这些工具来开发特定疗法。另一个例子是免疫系统。我认为它非常适合研究我们在细胞工程方面的一些工作,当我们从蛋白质到细胞再到体内整个动态系统层层递进时。它有些特殊。细胞可以在体内四处移动。显然这在应对不同疾病中扮演重要角色。如何让免疫系统更好地运作?但如何连接最后一英里,我认为这更适合生物技术公司或其他单独研究特定领域的学者来做。这就是我们如何思考构建工具集来帮助加速所有这些其他人的工作。至于时间线是 10 年,希望现在不到 100 年,我认为对于普通医生或患者来说,思考进展中哪些是外部可见的很有用。你在 UCSF 与患者合作了很长时间。如果确实在加速进展,医生应该关注什么?人们应该关注什么?
I mean, I'm optimistic it'll be sooner. I think the thing that's complicated is that it's a dynamic system. So if you fix something, there will obviously be future things that you need to work on. I don't think that the current set of things that we're aware of are going to be the only things that need to get worked out. But I think that the progress with AI is really very exciting. The other thing I'd say is we really look at more kind of systems than specific diseases. For example, one area that seems really important to understand is inflammation. We talked about this a bunch. This is a big focus of the Chicago Biohub. There's a lot of data on that. It seems quite clear that it's connected to a bunch of different diseases. But rather than studying the specific diseases, we think that by trying to understand inflammation more broadly, that will make it so that other companies can then use these tools to work on specific therapies. Another example is the immune system. I think it's a very good case to study for some of the work we're doing in cellular engineering, when we're kind of layering up from proteins to cells to whole dynamic systems within the body. It's sort of privileged. The cells can travel around through the body. Obviously that has a big part in addressing different diseases. How do you make the immune system function better? But exactly how you connect that last mile, I think, is going to be more something that biotech or other academics individually studying things will be better suited to do. So this is how we think about building out the tool set that just helps accelerate all these other folks. Whether the timeline is 10 years, hopefully less than 100 now, I think it's useful for maybe your average doctor or patient to think about what's externally visible in the progress here. You worked with patients for a long time at UCSF. What should doctors look out for? What should people look out for if you're actually accelerating progress?
这部分让我对进展感到非常兴奋,尤其是 Alex 和他的团队推出的这个项目。我认为很明显科学将开始快速发展。但对我来说不太清楚的是我们如何将其转化到临床以及具体会是什么样子。我认为必须改变的是我们进行临床研究的方式。我的希望是我们真正缩短了实验室研究到患者影响之间的距离。但那里有很多步骤,我们需要真正照顾患者的人创造性地思考,并思考如何安全地部署。这是一个我们还有工作要做的差距。我们与 Jennifer Doudna 在 UCSF 的 Cures 项目合作。所以我们正在初步了解研究的部署需要如何改变,因为研究进展会很快,但这仍在形成中。
This is the part I'm super excited about the progress for us especially with this launch that Alex and his team have put forth. I think it's very clear that science is going to start moving pretty quickly. And I think the thing that's less clear to me is exactly how we translate to the clinic and what that looks like. I think what has to change is actually the way we do clinical research. My hope is that we're really shortening the distance between bench research and patient impact. But there's a lot of steps there that we need people who actually take care of patients to think creatively and think about how to deploy safely. That's a gap that we have some work in. We partner with Jennifer Doudna's Cures program at UCSF. So we're dipping our toe in understanding how the deployment of research needs to change given how quickly research will be progressing, but that one is still shaping up.
也许我可以谈谈我们最近的发布,因为我认为它也...
Maybe I could say something about our most recent launch because I think it also...
请,我们应该明确地谈谈。
Please, we should explicitly.
是的。大约一周前,我们发布了新的 ESM Fold。这基本上是一个用于蛋白质生物学科学发现的开放系统。它是一个经过训练的蛋白质生物学世界模型。它基于语言模型。所以它经过了数十亿蛋白质序列的训练,学习了这些蛋白质生物学的涌现表征,然后我们可以用它来预测原子分辨率的蛋白质结构。它非常快,极快。所以它展示了结构预测中速度和准确性的帕累托最优前沿。这使我们能够表征蛋白质宇宙中非常广阔的区域。我们折叠了超过 11 亿个蛋白质,预测了它们的结构,并通过机制可解释性识别了连接所有蛋白质的特征。但我认为这个模型最令人兴奋的是它是一个非常通用的蛋白质生物学模型。你可以把它当作一个世界模型。你可以真正开始搜索世界模型的空间来设计新的蛋白质。它在几乎每个结构预测基准上都达到了最先进水平,尤其是在蛋白质-蛋白质相互作用和蛋白质-抗体相互作用上,这对治疗设计非常关键。我们发现你现在可以用这个模型来设计蛋白质和设计单链抗体。你可以在数字上完成所有这些,然后在少量实验试验中,基本上是一个 96 孔板,你从数字上的数十万条轨迹中选择,实际合成 96 个蛋白质,在实验室中用一个非常简短简单的实验周期进行测试。我们发现了纳摩尔级别的结合物。这确实是治疗活性的水平。所以它真正展示了你可以拥有这些通用模型。我们没有为抗体设计模型。我们没有设计一个能够结合特定靶点的模型。我们只是设计了一个能够理解蛋白质的模型,然后蛋白质设计作为涌现属性出现。我还认为它展示了开放科学和开源的力量,因为我们将其作为开放发现引擎发布,所以任何人都可以在此基础上构建。它将那些需要筛选数十万或数百万抗体的高强度实验室实验,变成了你可以直接启动一个实例进行计算,然后就能生成抗体。
Yeah. So about a week ago, we announced the new ESM Fold. This is basically an open system for scientific discovery in protein biology. It's a world model of protein biology that's been trained. It's a language model based. So it's been trained on billions of protein sequences, learns these emergent representations of protein biology, and then we can use it to make predictions of atomic resolution protein structure. It's really fast, blazing fast. So it's illustrating this Pareto optimal frontier of speed and accuracy in structure prediction. This allows us to characterize really vast stretches of the protein universe. We folded over 1.1 billion proteins and predicted their structures and identified features connecting all of them through mechanistic interpretability. But I think the thing that was most exciting about this model is it's this really general model of protein biology. You can use it as a world model. You can actually start to search the space of the world model to design new proteins. It's hitting state-of-the-art across pretty much every structure prediction benchmark, especially on protein-protein interactions and protein-antibody interactions, which is really critical for therapeutic design. What we found is you can actually now use the model to design proteins and to design single-chain antibodies. You can do all of this digitally, and then in a small number of experimental trials, basically a 96-well plate, you select from hundreds of thousands of trajectories digitally, actually synthesize 96 proteins, test them in the lab in a really short easy experimental cycle. We found nanomolar binders there. That's really the level for therapeutic activity. So it's really showing that you can have these general purpose models. We didn't design a model for antibodies. We didn't design a model to be able to bind one particular target. We just designed a model that could understand proteins, and you get protein design as an emergent property. I also think it illustrates the power of open science and open source, because we release this as an open discovery engine, so really anyone can build on it. It takes what are really intensive laboratory experiments where you have to screen through hundreds of thousands or millions of antibodies in high throughput screens, and you can really just spin up an instance and compute, and now be able to generate antibodies.
你应该多谈谈我们如何在进行抗体筛选时获取那些数据,然后我们进行了验证,我们观察了细胞中的 PDL,然后在冷冻电镜下观察,以及所有这些如何补充和验证了你在模型中看到的结果。
You should say more about how we took that data when we did an antibody screen, and then we validated, we looked at PDL in cells, and then we looked at it under the cryo, and how all that complemented and validated what you were seeing in the models.
没错。是的。
That's right. Yeah.
所以我认为,真正去实验室里表征这些分子至关重要。我们这里有一个结构生物学中心,配备了极其强大的冷冻电镜,因此我们能够从生物物理和功能角度观察这些蛋白质。我们为几个治疗相关靶点设计了蛋白质,并能在其按预期工作时确认其功能。
So I mean I think it's really critical to actually go and characterize these molecules in the lab. We have a structural biology center here with incredibly powerful cryo-EM microscopes, so we're really able to look at these proteins biophysically and functionally. We designed proteins for several therapeutically relevant targets and we're able to confirm their function when it works the way it's supposed to.
是的,这非常了不起。
Yeah, it's very amazing.
我们还能观察结构,从而在结合界面看到原子级别的分辨率。
We're able to look at the structure also, so you can see atomic resolution at the binding interfaces.
没错。我知道你的很多工作都集中在基础研究和构建基本原理上。如果看实际转化为药物或药物开发,临床试验通常需要 15 年,花费 15 亿美元。其中大约 5000 万是分子和临床前工作,耗时几年。剩下的 14.5 亿和十多年实际上是药物开发部分。很多瓶颈在于监管问题、患者招募等各种因素,但也与药物在吸收或毒性等方面的试验失败有关。你是否考虑过解决分子设计的另一链条,还是主要关注基础生物学和初始分子?
Correct. I know a lot of your work is really focused on basic research and building out the fundamentals. If I look at actual translation into drugs or drug development, often a clinical trial will be 15 years, cost $1.5 billion. About $50 million of that often is the molecule and pre-clinical work, a few years of work. The other $1.45 billion and decade plus is actually the drug development side. A lot of that seems gated on regulatory issues, recruitment, a variety of things, but also has to do with failure of drugs in trials around absorption or toxicity. Have you considered tackling that other chain of molecular design, or is the primary focus more on basic biology and initial molecules?
我的希望是,在构建这个细胞如何工作的综合模型时,也能预测脱靶效应。我认为你可以用生物模型做到这一点,因为目前一些脱靶效应只是我们不知道你的肾细胞也表达了这种受体,当我们在人体中测试时,就看到了肾毒性。如果你有一个单细胞图谱,涵盖所有不同的细胞类型,其中一些在我们建模之前并未被预测到,你就可以开始查看哪些细胞实际上拥有你本以为只针对的靶点的受体,从而在进入人体试验前预测一些下游效应。我认为这实际上是转录组学模型最令人兴奋的应用之一,可以理解当你干预时不同细胞会如何反应。但当你考虑递送机制和患者护理时,就必须创造性地思考首先想治愈哪种疾病。有些疾病更容易递送治疗药物,或者风险收益比更合理。我想去年我们都受到婴儿 KJ 的启发,当时 CHOP 的团队成功递送了 CRISPR 疗法,编辑了一个本会导致严重神经毒性并改变他一生的突变。那个疾病是经过精心选择的,因为我们需要靶向他的肝细胞,而且我们可以轻松递送一种能在肝脏起效的产品。我认为,正是这种选择正确应用的创造力和能力,能帮助我们解锁首批应用。
I mean, at least my hope in building this comprehensive model of how cells work is actually also being able to predict off-target effects. I think you can do some of that with biological models because right now some off-target effects are just that we didn't know your kidney cell also expressed this receptor, and then when we test it in humans, we see it happening and see renal toxicity. If you have a single cell atlas that looks at all the different cell types, some of which were not predicted before we modeled them, you can start looking at which cells actually have receptors for the target you thought you were exclusively targeting and be able to predict some of these downstream effects before we get into human trials. I think that's actually one of the more exciting applications of a transcriptomic model to understand how the different cells will react when you intervene. But when you think about delivery mechanisms and patient care, you start having to be creative about what disease you want to cure first. There are certain diseases that will be easier to deliver a therapeutic to, or the risk-reward makes more sense. I think we were all inspired by baby KJ last year when the team at CHOP was able to deliver a CRISPR therapeutic to edit a mutation that would have inevitably led to significant neurotoxicity and altered his life. That disease was very carefully chosen because we needed to target his liver cells, and we could easily deliver a product that would work in his liver. I think that's when the creativity and wherewithal to choose the right applications can help us unlock the first applications.
也许再补充一点,因为你描述了传统的药物开发过程。我认为这些工具有潜力对该过程产生很大影响。但有趣的是,真正开始思考可能开启的新范式。如果开发药物、设计分子、通过所有阶段的障碍大大降低,这意味着什么?你有了可编程生物学,可以真正开始为每个患者创造药物。我认为这对我们如何进行药物开发以及医学的未来有着巨大的影响。
Maybe something to add to that also, because you described the conventional drug development process. I think these tools have the potential to have a lot of impact on that process. But what's interesting is to really start to think about the new paradigms that can open up. What does it mean if the barrier to develop a drug, to design a molecule, to get through all of those stages is so much lower? You have programmable biology, and you can really start to create a medicine for every individual patient. I think that has enormous implications for how we do drug development and what the future of medicine looks like.
是的。当 FDA 接受基于个人视角的虚拟一期临床试验时,那将是激动人心的一天。
Yeah. It'll be an exciting day when the FDA accepts a virtual clinical trial for phase one or something based on some personal view of that person.
是的。即使达不到那一步,想想你看到加速的具体机制。我想如果人们觉得他们可以预测对肾细胞的影响,或者因为有了更广泛的理解而对毒性有更强的洞察,他们会愿意尝试更多的项目。
Yeah. Or even short of that, thinking about the specific mechanisms where you see this acceleration. I imagine if people feel like they can predict impact in kidney cells or have a stronger perspective on toxicity because they have this broader understanding, they'll be willing to try many more programs.
是的。患者招募也可能改变。我们有一个项目叫“罕见即唯一”,基本想法是很多人关注最常见疾病,但还有长尾,公司关注这些疾病的经济效益并不好。但如果你能让患者群体聚集起来组织起来,说“嘿,我们愿意尝试这种实验性药物”,那么由于你提到的成本,以及它在总成本中占比巨大,如果你能扭转这一点,那么如果你能更容易地生成某种东西,并与一群人配对,经济上就会合理得多。我认为科学和工程中有趣的一点是,你常常会在常见问题上碰壁,但很多时候,你从发现边缘案例中发生的某种罕见或奇怪的事情中,能学到更多关于系统的知识。所以我认为这始终是其中有趣的一部分,而且与这里的情况联系得很好,因为现在你将能够支持长尾的新想法被尝试,并可能更容易地得到测试。
Yeah. The recruitment could also change. We have this program Rare is One, and the basic idea is that a lot of people focus on the most common diseases, but there's this long tail, and the economics don't quite work out for companies to focus on those diseases. But if you can make it so that groups of patients can come together and organize and say, 'Hey, we would take an experimental drug on this,' then because of the cost you're talking about and how that's a huge amount of the overall cost, if you can flip that, then it actually makes the economics make a lot more sense if you can generate something more easily and pair it with a group of people. I think one of the interesting things from science and engineering is that often you can hit your head against the wall on common problems, and in this case diseases, but a lot of times you learn a lot more about a system from finding some kind of rare or weird side thing happening in an edge case. So I think that's always been an interesting part of this that actually connects pretty well because now you're going to be able to enable a long tail of new ideas to get tried and potentially tested more easily.
是的,关于罕见病这一点说得很好。在我们的罕见病队列中,首先,它们非常鼓舞人心且强大,但患者群体正在自我组织患者登记、自然史登记、生物样本库,他们正在组织自己的临床试验。有一个疾病小组的基因疗法在 3 到 5 年内就取得了进展,而不是几十年。速度如此之快,是因为患者自己组织起了科学家或临床医生可能需要的资源。这太不可思议了。
Yeah, that's a really good point on rare. In our rare disease cohorts, first of all, they're incredibly inspiring and powerful, but patient groups are self-organizing patient registries, natural history registries, biobanks, they're organizing their own clinical trials. There's gene therapy that one disease group has moved forward over the course of like 3 to 5 years rather than decades. The speed is so fast because the patients themselves have organized the resources that a scientist or clinician might need. It's incredible.
你们三位在不同场合都提到过开放生态系统在如此广阔领域中的力量。我认为你们描述的关于开源和数据收集的广度或多样性的逻辑,也应该适用于语言模型世界和多模态 AI 世界。你觉得对吗?你在这里所做的工作是否改变了你对 AI 和 Meta 的看法?
All three of you have mentioned at different points the power of open ecosystems in such a large space. Like I think some of that logic around open-source and the breadth or diversity of data collection that you were describing, it should also apply in the language model world and the multimodal AI world. Do you think that's right? Does any of the work you're doing here change how you think about AI and Meta?
我认为总体上是类似的理念。你知道,Priscilla 也谈到过这一点,我们很多工作重点都是构建工具来赋能个人。这是我许多工作的共同主题:把技术交到个人手中。我们不相信那种高度集中的未来,即少数机构主导一切进步。我们的愿景不是出现一个中央超级智能来解决所有科学问题。我认为人非常重要,而且未来会更加重要。给人们更多工具来提高生产力,将是任何积极未来的关键部分。历史上进步从来不是通过集中化实现的,而是通过赋能个人去尝试那些有点非主流、别人认为不是好主意的事情——因为他们觉得那些好主意已经被做过了。所以我认为这非常核心。从某种程度上说,这就是为什么你会创造像社交媒体这样的东西,对吧?为了给人们发声。我认为在通过个人 AI 赋能人们方面,开源是其中一种体现,但不是唯一方式。它确实是一种方式,基本上就是说我们要把这项技术交到每个人手中。在科学方面,我认为这非常有意义,我们深度致力于开源。当然,也有一些重要的考量需要平衡,比如生物安全等问题,我们需要思考如何处理。但总体而言,这深深植根于我们在 Biohub 所做工作的理念中,也可能是我所做很多事情的共同主题:我们相信一个积极的未来是,你把技术作为工具,交到个人手中,社会就是这样进步的。
I mean, I think it's sort of a similar philosophy overall. And you know, Priscilla was talking about this, that a lot of our focus is building tools that empower individuals to do things. That's a common theme across a lot of the things that I work on: just putting the technology in individuals' hands. We don't believe in this very centralized future where there should be a small number of institutions that basically are advancing all this stuff. Our vision is not that there's going to be some central superintelligence that solves all of science. I think people are really important, and I think we'll be more important in the future. Giving people more tools to be more productive is going to be a critical part of any kind of positive future. That's how progress has always been made historically, right? It's not through centralization. It's through empowering individuals to try things that are somewhat out of the mainstream, that other people didn't think were good ideas because they thought they were good ideas that already have been done. So I think that's very central to the whole ethos. To some degree, it's like why you create something like social media, right? To give people a voice. I think a lot of the stuff that I care about in terms of empowering people with individual AI, open source is one instantiation of it. It's not the only way to do it. It certainly is one way that you are basically saying we're going to take this technology and put it in everyone's hands. In terms of science, I think it really makes sense, and we're deeply committed to open source. There are obviously interesting considerations on this that are important too, because there's a lot of considerations around biosafety and things like that that we're going to need to balance and think through how to handle. But I think overall, this is very deep in the ethos of the work that we're doing both at Biohub and probably a theme for a lot of the stuff that I do: we believe that a positive future is one where you build a technology as a tool, you put it in individuals' hands, and that's kind of how society makes progress.
你在 Biohub 有着极其雄心勃勃的使命,然而在这里工作的 AI 科学家也可以去商业企业工作。你如何看待人才问题,以及如何吸引人们来到 Biohub?
You have this incredibly ambitious mission at Biohub, and yet the AI scientists that work here could also go work in commercial enterprises. How do you think about the talent and how to bring people to Biohub?
你想从哪里开始?我认为 AI 研究人员市场非常火热,但这也意味着他们需求旺盛,可以从事他们想做的事情。我认为这又回到了关于前沿 AI 和前沿生物学的观点,对吧?在这里工作的 AI 研究人员可以去任何主要实验室研究语言模型之类的东西,但那些实验室没有前沿生物学部分。所以我认为这里还有一个非常大的使命成分:在这里你可以做独特的工作,在其他地方真的做不到。如果你的关注点就是这个,那么我认为世界上没有任何其他组织同时在做前沿生物学和前沿 AI。
I mean, where do you want to start? I think it's a very hot market for AI researchers, but I think that part of what that means is that there's a lot of demand and they are very in demand and can work on the things that they want to work on. And I think this gets back to this point again about frontier AI and frontier biology, right? So the AI researchers who work here could go work on language models or things at any of the main labs. But those labs don't have the frontier biology part attached to it. So I think that there's also a very large mission component to this: there's an ability to do this unique work here that you just can't really do at other places. If that's what your focus is, then I don't actually think that there's any other organization in the world that's doing both the frontier biology and the frontier AI.
是的。Alex,你为什么在这里?
Yeah. Why are you here, Alex?
我认为这很简单。我们的使命是照护和预防疾病。我认为这真是太……
I mean, I think it's really simple. Our mission is to take care of and prevent disease. And I think it's just such a...
你面不改色地说在不到 100 年的时间里。
You say with a straight face in a less than 100-year timeline.
现在非常严肃。不再有……
It's very serious now. There's no more...
那是……是的。是的。这是一个非常强大的使命,我认为……
That's... Yeah. Yeah. It's a really powerful mission and I think...
是的,我的意思是,这只是……我认为科学家们对此非常有动力。这是人们深受激励的事情。而且我认为我们正处于这样一个时刻,这似乎确实是可以实现的。我认为我们正在建设一个非常独特的地方,在那里我们解决这个问题,我们拥有资源和正确的东西来真正去追求并实现它。
Yeah, I mean, it's just... scientists, I think, are very motivated by that. It's something people are deeply motivated by. And I think we're at this moment in time where that actually seems like something that can be achieved. And I think we're building a really unique place where we're tackling that problem, and we have the resources and the right things to actually really go after that and do that.
是的。我作为一个与许多研究科学家交谈并招聘他们的人,对此深有共鸣。他们想知道你是否拥有数据、工具、算力和人才,然后使命是什么。所以我实际上认为这非常有竞争力。
Yeah. I mean, that resonates with me as somebody who talks to and hires a lot of research scientists. They want to know if you have the data, if you have the tools, if you have the compute, if you have the talent, and then what the mission is. So I actually think that's super competitive.
另一件事是你不需要一个非常大的团队,对吧?所以我认为世界上有趣的一点是:人们关心不同的使命,这很好。我认为这正是为什么构建这些工具并让人们有能力探索他们关心的事物——无论是科学还是其他一切——是推动社会进步的如此强大的方式。人们关心不同的事情。要在 AI 上取得进展,你不需要成百上千的 AI 研究人员。我认为你可以通过一个由十几人或几十人组成的非常强大的团队取得进展。找到关心这个使命的人并不是特别困难。我的意思是,这是世界上非常重要的事情。所以我认为这只是世界上很酷的一点:人们显然被不同的使命所吸引。
The other thing is that you don't need a very large team, right? So I think it's an interesting thing about the world: people care about different missions, and that's good. I think that's part of why building these tools and giving people the ability to explore what they care about—whether it's across science or just across everything—is such a powerful way to make progress in society. People care about different things. In order to make progress in AI, you don't need many hundreds of AI researchers or thousands or anything like that. I think you can really make progress with a very strong group of a dozen or a couple dozen people. And finding people who care about this mission is not a particularly hard thing. I mean, this is a super important thing in the world. So I think that's just kind of a cool thing about the world: people obviously are drawn to different missions.
所以,我认为即使关注这个领域的人,最简单的思维模型基本上是:蛋白质结构预测模型和蛋白质-蛋白质相互作用模型。然后有一个基础理解的层面,还有一个理论是,有一天我们能够零样本地将东西直接推向临床,并且命中率大大提高。要从 ESMFold 2 走到这一步,需要发生什么?这可行吗?
So, I think the simplest mental models that folks have, even if they're paying attention to the space, are essentially: okay, structure prediction models for proteins and protein-protein interaction models. And then there's this one piece which is fundamental understanding, and then there's this theory of someday we're just going to be able to zero-shot things into the clinic with a much better hit rate. What needs to happen for us to go from ESMFold 2 to this other piece? Is that feasible?
我觉得这是个很好的问题。我对此非常乐观。一方面,历史上人们可能花整个职业生涯来解决这些问题:如何有效优化药物?如何通过临床前阶段?如何进行早期安全性测试?我认为,当有了新的科学范式时,曾经困难的问题会通过新范式变得简化。所以我非常乐观,许多核心问题会通过这些模型以涌现的方式得到解决。一个很好的例子是毒性。如果你能真正数字化地模拟一切,并预测药物在人体内的分布和结合位置,你就有了解决这类问题的开端。所以我认为,一旦我们在分子层面有了这些准确的表征,就会看到许多核心问题取得非常快速的进展。
I think that's a great question. I mean, I would say that I'm really optimistic on that. So, I think on the one hand, these are problems that historically people could spend an entire career working on: how do you figure out how to effectively optimize a drug? How do you get it through pre-clinical? How do you do the early safety? I think that when you have a new scientific paradigm, questions that were once hard become simplified through the new paradigm. And so I'm very optimistic that many of these core problems will be solved in an emergent way through these models. And I think one great example of that is toxicity. If you can really digitally simulate everything and be able to predict where a drug is going to distribute and bind across the human body, you have the beginning of a solution to that kind of problem. So I think that once you have these accurate representations at the molecular level, we're going to start to see really rapid progress on a lot of these core problems.
自发布以来,过去一周你看到的最令人兴奋的模型使用或实验是什么?
What is the most exciting use or experimentation with the models you've seen in the last week since release?
是的,看到它被集成到各种事物中真是太棒了。我认为一个非常有趣的现象是,人们将它连接到智能体系统,进行自动化设计,并自动化整个流程。所以这又是一个例子,说明如何将智能体和前沿 AI 与生物学世界模型结合起来,真正推理生物学,并开始自动化整个设计过程。
Yeah, it's just been great to see it get integrated in all kinds of things. I think one of the really interesting things that we've been seeing is people connecting it with agentic systems to do automated design and automate that whole process. So it's really another example of how you can bring together agentic and frontier AI with the ability to have a world model for biology and actually reason about biology and start to automate the entire design process.
你如何决定研究议程的下一步?比如生物学世界模型,然后我可以扩大规模,添加更多数据。添加数据在新方法和领域方面并非易事。你是否从更大的生态系统中获取关于人们如何使用它的反馈?什么会让它更有用?还是说我们确实清楚下一步要寻找的结构或覆盖范围?
How do you decide what the next step in the research agenda is? It's like world model for biology and then I could scale it up, add more data. Adding data is non-trivial in terms of new methods and domains. Do you take input from the larger ecosystem about how people are using it? What would make it more useful? Or is it really that we understand the next step of structures or coverage that we're looking for?
我认为有两件事。我们对下一个大挑战有看法,那就是虚拟细胞,真正能够沿着生物复杂性的层级上升到细胞层面。
I think there's two things. So we have a view on the next big challenge, which I think is the virtual cell, and really being able to ladder up the hierarchy of biological complexity to the cell.
抱歉,一个非常基本的问题:这个虚拟细胞模型,我应该期望的输入和输出是什么?
Sorry, very basic question: this virtual cell model, what is the input and output I should expect?
是的,我认为对此有不同的看法,但最终你想要的是一个能够真正建模每个复杂性层级的系统。所以是蛋白质组层、遗传层、转录组层,并将这些与表型连接起来。你需要足够的通用性,以便能够向模型询问关于一个它未训练过的上下文中的新干预措施的问题,并从中得到答案。我们作为领域需要弥合的差距是能够真正做出那些可以泛化的预测。所以这需要巨大的努力来生成数据。
Yeah, I think there are different views on that, but what you ultimately want is a system that can really model each of the levels of complexity. So the proteomic layer, the genetic layer, the transcriptomic layer, and connect that to the phenotype. And you need enough generality so that you can ask the model questions about a new intervention in a context that it hasn't been trained on and get an answer from it. The gap that we need to close as a field is being able to really make those predictions that can generalize. So that's going to require an enormous effort to generate data.
然后关于你决定下一步做什么,我认为这是一个相当正常的约束管理过程,对吧?我认为世界上每个领域的每个实验室可能都感到算力受限。我认为这里可能也是如此。所以总是有问题:我们是应该更专注于推进蛋白质部分?还是应该做更多细胞方面的工作?这些是关于如何排序的持续辩论。然后在那之内,存在一个帕累托前沿,关于你想训练不同模型到什么程度,而模型的大小也取决于你拥有的数据规模。所以我认为其中一部分只是你想在曲线上处于什么位置以及正常的约束,但我认为这可能是任何研究组织都会经历的过程:你想朝所有这些不同的方向前进,你只是试图进行约束优化,在某个时间点在一件事上取得足够进展以做出世界级的工作,同时播下一些种子,这些种子在未来几年也能开花结果。
And then in terms of what you decide to do next, I think this is a pretty normal process of constraint management, right? I think every lab in every field across the world probably feels compute constrained. I think that's probably true here too. So there are always questions: should we double down more on advancing the protein piece? Should we do more of the cellular stuff? Those are ongoing debates in terms of how you sequence that. And then within that, there's being at the Pareto frontier about how much you want to train the different models, and the size of the models is also dependent on the scale of the data that you have. So I think it's some of that is just where you want to be on the curves and normal constraints, but I think this is probably the same process that any research organization goes through: you want to go in all these different directions and you're just trying to constraint-optimize and make enough progress to do world-class work at one thing at a time while planting some seeds that can blossom over the next couple years as well.
这是我职业生涯中至少见过的最具活力的技术时期。我的意思是,AI 领域发生的一切都令人兴奋,每周都有新变化。你是感到疲惫还是振奋?
This has been the most dynamic period of technology, at least I've seen over my career. I mean, it's so exciting in terms of everything that's happening with AI and every week there's something new that's changed. Are you tired or invigorated?
两者都有。我觉得每个人都处于狂躁阶段。是的,是振奋和疲惫的结合。
I'm both. I feel like everybody's in a manic phase. Yes, it's a combination of invigorated and exhausted.
现在事情非常不可预测。真的很难知道接下来会发生什么。我们在模型方面看到了指数增长的早期迹象,智能体流程开始以非常有趣的方式出现,模型越来越多地帮助模型,但这仍然非常早期。如果从现在回看 5 年,你要定义相对于你的努力什么是成功,我知道事情非常动态且变化很大,但你有为生物中心提供工具这一共同主线,你有大规模赋能科学家这一共同主线。回看 5 年,有没有一件你特别想确保完成或实现的事情,或者一个主要目标?
Things are very unpredictable right now. It's really hard to know what's coming. We have these early signs of exponentiation on the model side with agentic flows that we're starting to see in really interesting ways, models starting to help more and more with models, but that's still very early days. If you're thinking back 5 years from now and you were to define what success was relative to your efforts, and I know things are very dynamic and changed a lot, but you have this common thread of tooling for the biohub, you have a common thread of empowering scientists at scale. Looking back 5 years from now, is there a specific thing that you really want to make sure that you've accomplished or achieved, or a primary goal?
嗯,我认为我们对想要围绕生物学构建的这一套层级化的世界模型有相当清晰的认识。
Well, I think we have a pretty clear view of this hierarchical set of world models that we want to build around biology.
另一部分是我们希望做世界上最高质量的工作。我认为我们基本上已经为此做好了准备,拥有一支世界级的 AI 研究团队和一系列世界级的生命科学研究机构——Biohub。这是世界上任何其他组织都不具备的基础条件。但拥有很多优秀要素并不能保证成功。所以,对我来说,5 年后回头看,我相信其他实验室或项目也会尝试做出与我们目标相近的东西。而我认为我们应该能够做出更有意义、更独特的知识贡献。这正是做任何研究时都应该追求的目标。如果我们做到了,我想大家都会感觉非常好。我也预计,在某个时候,我们会开始看到使用模型的人产生更多创意。但我有足够的信心,这部分会自然实现,对我来说更重要的是确保我们做出世界级的工作,如果做到了,其他事情几乎会水到渠成。
And the other part of that is that we want to do the highest quality work in the world. I think we're basically set up to do that between having a world-class AI research team and this collection of biohubs, which are world-class life sciences research organizations. That's fundamentally a setup that no other organization in the world has. But you can have a lot of great ingredients and that doesn't guarantee that you succeed. So, to me, 5 years from now, looking back, I'm sure other labs or efforts will try to produce things that approximate what we're trying to do. And I just think that we should be able to do something that is meaningfully better and a unique intellectual contribution to the world. That's what you're trying to do whenever you do any kind of research. So if we do that, I think we'll all feel very good. I would also expect that at some point we'll just start seeing a lot more idea generation from the people using the models. But I have enough faith that that part will materialize that for me it's more just about making sure that we do world-class work and I think if we do, the rest almost will take care of itself.
最后一个问题。假设现在是 2026 年中,过去一年你对 Biohub 或这个领域最大的认知更新是什么?
Very last question for you. Snapshot of it's mid 2026. What's the biggest update in your own thinking about Biohub or the domain from the last year?
从过去一年来看。你也是去年加入的。我认为最大的变化是我们基本上调整并正式确定了 Biohub 是我们慈善事业的主要焦点。所以这是一个非常大的转变。但 Alex 和团队的加入很有意思,不仅因为这是一个世界级的团队——你们已经合作了一段时间——而且因为这个领域变化非常快。我认为有一件事被低估了:这是一个极其有才华的团队,他们彼此了解,合作良好,稳定且优秀。人们在一个稳定的环境中长期良好协作所带来的复利效应被低估了。所以这是非常重要的一点。但我们想做的部分事情是,在 Alex 领导之前,Biohub 的先前领导者基本上是主要对技术感兴趣的生物学家。而现在我认为我们真正扭转了这一点。显然你也有生物学背景,但你首先是一位 AI 研究者,在 AI 和生物学方面都有背景。我认为这深刻反映了我们期望未来如何驱动更多价值。所以这些可能是过去一年中我们工作方面最大的更新。而且这是一个新的领导者,不仅仅是一个领导者,而是一个团队,我认为这非常好。至于行业其他方面,它正按轨道发展。这有点疯狂,因为当你有一条指数增长曲线时,指数曲线的感觉是它增长得太快,以至于情感上觉得它不可能继续下去。但指数曲线的本质是它不仅会继续,还会加速。所以我认为这伴随着各种情绪和心理,但根本上,当你看到行业的曲线时,它一直保持在那条曲线上,这对所有这些领域都有非常深远的影响。当然,这验证了并让人对在那些如果保持轨道就会实现的事情上做出巨大投资感到非常良好。而且我们似乎就在轨道上。所以这是非常好的消息。
Well, from the last year. I mean, you joined in the last year. I think the biggest thing is that we basically rotated and formalized that Biohub is the main focus of our philanthropy. So, this has been a very big shift. But Alex and the team coming in has been interesting not only because it's a world-class group—you guys have worked together for a while—but also because stuff is changing so much in the field. I think one thing that's underrated is that this is an extremely talented group of people who also know each other and work well together and are stable and good. That is underestimated in terms of the compounding benefit of people being able to work well in a stable environment over time. So that's a really important piece. But part of what we wanted to do was, prior to Alex leading the effort, the previous leaders of the Biohub were basically primarily biologists who were interested in technology. Now I think we've really flipped that. Obviously you have a background in biology as well, but you are primarily an AI researcher who has a background in AI in biology. I think that's a deep reflection of the way we expect this to drive more value in the future. So those are probably the biggest updates in the last year in terms of the work we're doing. And it's a new leader, not just a leader, but a team, that I think has been really good. And then on the rest of the industry, it's on track. It's kind of a crazy thing because when you have an exponentially growing curve, the way an exponential curve feels is it's growing so quickly that the emotional feeling is it can't possibly keep going. But the nature of an exponential curve is it doesn't just keep going, it keeps accelerating. So I think that has all these emotions and psychology attached to it, but fundamentally when you look at the curve in the industry, it has remained on that curve, which has all these very profound implications for all of these domains. Certainly it validates and makes one feel very good about making a very big investment in the things that will play out if you stay on that track. And it seems like we are. So that is very good news.
我认为你们在那里做的最重要的事情是实际上正在与真实的生物学形成闭环,因为代码和研究是闭环系统,迭代非常快。而这是一个开环系统,所以你们正在闭合这个循环,这对进步至关重要。
I think the most important aspect of what you're doing there is you're actually closing the loop with the actual biology, because with code and research it's closed loop systems and so they're very fast to iterate. This is an open loop system so you're closing a loop and that's really crucial to progress.
是的。
Yeah.
对我来说,现在我们推动的战略以及 Alex 掌舵后最大的变化之一是,以前我们有出色的团队大致朝着同一个方向前进,并理解我们工作的潜在合作和相互联系。但现在我们手挽手一起前进。方向非常明确,非常令人兴奋。有点吓人,但这是一个真正的团队,相互配合,努力朝着这个目标前进。这需要很多努力,也需要我们的团队工作成熟到一定程度,使得相互锁定确实有意义。
For me, one of the biggest changes with the strategy we're driving now and Alex at the helm is, you know, before we had amazing teams moving generally in the same direction and understanding the potential collaborations and interconnectedness of our work. But now we are arms linked moving together. It's very directed and it's very exciting. It's a little bit scary, but it's like truly a team playing off each other in trying to make progress towards this goal. And that has taken a lot of work, but also the maturity, our teams being able to have their work at a level of maturation where it actually does make sense to interlock.
太棒了。祝团队保持在曲线上,感谢你们做这些。
Amazing. Well, to teams being on the curve, thank you guys for doing this.
感谢你加入我们。
Thank you for joining us.
谢谢。
Thank you.
在 Twitter 上关注我们 @no prior pod。订阅我们的 YouTube 频道,如果你想看我们的脸。在 Apple Podcasts、Spotify 或你收听的地方关注节目。这样你每周都能收到新一期。在 no-briers.com 注册邮件或查找每期节目的文字稿。
Find us on Twitter at no prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.