Xaira Therapeutics:AI 原生药物发现与虚拟细胞模型

Xaira Therapeutics: AI-Native Drug Discovery with Virtual Cell Models

王波 Bo Wang · Latent Space · 2026-07-21 · 约 90 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Xaira Therapeutics 正在构建一个 AI 原生的药物发现平台,利用虚拟细胞模型预测细胞反应,加速药物开发。

Xaira Therapeutics is building an AI-native drug discovery platform with virtual cell models to predict cellular responses and accelerate drug development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 32)

全文 · Full transcript(中英对照)

开场与嘉宾介绍 Introduction and Guest Introductions

Host

真正让我大开眼界的是,当我看到模型做出预测,直接打印出世代变化的热图,看着实际原始数据,把线性基线预测、真实值和 X-Cell 预测排在一起,视觉上非常清楚地看到 X-Cell 预测比线性基线更接近真实值。这就是我开头提到的“哇”时刻。这是第一次有人能把不是一项,而是七项基因组扰动实验整合在一起。我们生物学家立刻注意到的一点是,其中一些扰动是上下文通用的。

And what really blew my mind away is when I saw the model make prediction just print out the heat map of the generation changes look at the actual raw data and line up the linear baseline prediction the ground truth and X-Cell prediction all together it's visually very clear to see that X-Cell prediction is much more similar to ground truth than the linear baseline this is a wow moment I was talking about in the beginning this is the first time that someone can put together not just one prosy but seven genome web perturbsy campaigns together. Something that uh jumped out to us biologists right away is that some of the partations are context universal.

Host

大家好,我是 Mirror OMIX 的首席技术官 R.J. Haniki。这位是 Brandon Anderson,他在 Atomic AI 开发 RNA 疗法。这里是“潜空间:AI 用于科学”播客。本播客贯穿的主题之一是,实验室、实验和现实世界可能对“AI 用于科学”还是“B2B SaaS”这类事情影响最大、关联最密切。今天我们非常高兴邀请到 Xaira Therapeutics 的 Bo Wang 和 Ci Chu 来到演播室。在 Xaira,他们和其他人一起构建 AI 药物发现平台。他们利用高通量实验系统收集非常大的数据集,然后训练 AI 模型来预测你体内细胞对药物和疗法的反应。非常高兴你们来,我是你们工作的忠实粉丝。你们两位向听众介绍一下自己吧?

Hi, I'm R.J. Haniki, CTO of Mirror OMIX. This is Brandon Anderson who builds RNA therapeutics at Atomic AI and this is the latent space AI for science podcast. One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether uh something is AI for science or something like B2B SAS. Uh we're really happy to have in the studio with us today Bo Wang and Ci Chu from Xaira Therapeutics. At Xaira, they're building with a bunch of other people a AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Um why don't you two introduce yourselves to the listeners?

Bo

大家好,我叫 Bo Wang。我是 Xaira Therapeutics 的高级副总裁兼生物医学 AI 负责人。大约八个月前加入 Xaira,之前是加拿大多伦多大学的副教授。

Hello everyone, my name is Bo Wang. I'm SVP and head of biomedical AI at Xaira Therapeutics. Joined Xaira about eight months ago and before that I was associate professor at the University of Toronto in Canada.

Chu

嗯,我的名字非常难发音,除非你会说普通话。所以我通常叫 Chu,就像 Chewbacca 或 Pikachu 那样,你最喜欢的虚构角色。我是 Xaira 的 AI 驱动发现高级副总裁。大约两年多前加入,当时公司还在隐身模式。在这里我领导高通量生物学团队,生成用于训练 AI 模型的数据,并思考它们的应用。在此之前,我在 AI、大数据和生物学的交叉领域工作了大约十年。之前我在 Citro 领导虚拟发现平台,再之前我在从 Google X 分拆出来的 Verily 工作。

And um I'm such uh my first name is incredibly difficult to pronounce unless you speak Mandarin. So I go by Chu uh as in Chewbacca or Pikachu. I think your favorite fictional character. I'm the SVP of AI enabled discovery at Xaira. I joined about more than two years ago when it was still in stealth mode. Uh and here I lead the high throughput uh biology group generating the kind of data that will feed our AI models and also think about their applications. Um before this I uh spent about a decade in uh at the intersection of AI and uh uh big data and biology. uh as I previously I worked at in Citro leading uh the invitual discovery platform there and before that I was at Verily which spun out of Google X.

Xaira的使命与AI平台 Xaira's Mission and AI Platforms

Host

好的。你们在 Xaira 这家公司,它正处于名字容易混淆和巨额融资的前沿。嗯,Xaira 大概几年前从隐身模式出来,然后突然就变得非常庞大。所以,我很好奇你们能否解释一下 Xaira 的使命是什么,他们的核心论点是什么,比如 Xaira 有什么特别之处,以及你们未来的发展方向。

Okay. So you you are at Xaira the company which is on the prao frontier of confusing names and mega rounds. Um so uh Xaira is I think kind of came out of stealth like a few years ago and just really big org kind of out of nothing. So, um I'm curious if you can explain a little bit about what is Xaira's mission, what is their thesis statement, like what is, you know, special about Xaira and, um, you know, kind of where you're going in the future.

Bo

是的,Xaira 是一家 AI 驱动的药物发现公司。我们使命的核心是利用 AI 平台生成更好的疗法,以推进患者护理。因此,我们将利用不同的 AI 能力来制造药物。我们正在构建三个主要的 AI 平台。第一个是蛋白质设计工作,源自我们的联合创始人 David Baker 博士在华盛顿大学的团队。当前这一代蛋白质设计者中有很多都在我们公司。那里的思路是利用先进的 AI 技术来开发针对以前无法成药靶点的分子。第二个 AI 平台,我想我们今天会花很多时间讨论,就是我和 B 已经研究了一段时间并刚刚发布预印本的那个。那就是虚拟细胞或生物学基础模型。那里的希望是构建一个 AI 模型来预测生物学,正如你所说,预测哪些基因和药物分子会影响细胞生物学。第三个部分我们刚开始构建,是患者表征模型,目标是让 AI 模型能够理解哪些患者会对哪些疗法产生反应。所以希望这些平台技术加在一起,能帮助我们更快地制造更好的药物,并且成功率比以往的技术更高,把过去那种手工艺式的试错时代,越来越转变成一门工程学科。

Yeah, Xaira is a uh AI uh enabled drug discovery company and at the core of our mission uh we're using AI platforms to generate better therapeutics uh to advance patient care and so we will be making drugs uh using different AI capabilities. There are three main AI platforms that we're building here. Uh the first one is protein design work that spun out of uh our co-founder uh Dr. David Baker's group from Udub. Uh a lot of the current generation of protein designers are here in the company. So there the thinking is to use advanced AI technology to develop um uh molecules against previously unable targets. The second AI platform I guess we'll spend a lot of time talking about today is the one that uh B and I have been working on for quite some time and just released a pre-print on. That's the virtual cell or foundation model of biology work. there the hope is to build a AI model to predict biology exactly like you said uh and predict what genes and drug molecules will affect cell biology and the third piece which we're beginning to build now is patient representation models and the goal there is to have AI models that can understand which patients will respond to which therapeutics so hopefully together these platform technology will help us make uh better drugs faster and uh with a higher success rate than previous technologies to transform what is used to be artisal tri and era in the past into more and more into an engineering discipline.

Chu

是的,我认为 Xaira 的不同之处不仅在于那十亿美元左右的融资,还在于 Xaira 是为数不多的 AI 原生的药物发现公司之一,它覆盖药物发现的各个环节,从早期的靶点识别、蛋白质设计、小分子,一直到一期到三期临床试验。我们的目标是用 AI 加速药物发现的每一个部分,这样不仅提高药物开发的成功率,而且大大缩短周期时间,这样我们就能有新药,而不是每 10 到 20 年才有一次。所以希望我们能缩短周期,为患者提供更多有用的药物。

Yeah, I think what sets Xaira different is not just the one billion uh around but also um I think Xaira is one of the very few AI native companies for drug discovery that works from end to end of all sections of drug discovery from as early as you know target ID and the protein design small molecules and two phase one to three clinical trials. We aim to use AI to accelerate every part of the drug discovery so that not only we increase the success rate of uh developing drugs but also greatly you know reduce the cycle time so that we can have you know new drugs instead of every 10 20 years so hopefully we can have the cycle time so we have more useful drugs for patients.

模型瓶颈与整合 Bottlenecks and Integration of Models

Host

这真的很有趣。我知道现在很多人对第三件事感兴趣,有些人可能称之为从实验室到临床的转化。嗯,瓶颈在哪里?你们有这三个模型。你们正在解决的瓶颈是什么?你们是怎么做的?为什么这样做?

That's really interesting. I know there's a lot of interest right now in that third thing some maybe called translation from the lab to the clinic. Um where are the bottlenecks? So you have these three models. What are the bottlenecks that you're addressing and and sort of like how are you doing that? Why are you doing it that way?

Bo

我们是一家 AI 原生的公司。在药物发现的几乎所有环节,我们都试图用 AI 来彻底改变我们开发药物的方式。早期部分,我们构建因果基础模型,有时我们称之为虚拟细胞;蛋白质方面,我们有最先进的蛋白质工程模型;我们还有患者表征学习模型。我认为 Xaira 要做的不仅是开发 AI 模型,还要创建正确的数据集来赋能这些模型。让我在 Xaira 工作感到兴奋的是,我们总是致力于将三个 AI 模型连接起来,而不是让它们各自独立工作。所以当我们设计虚拟细胞模型时,我们会寻找连接:我们能否找到更容易应用蛋白质工程模型的靶点?甚至当我们设计细胞因果模型时,能否连接到患者表征?什么样的患者数据能连接细胞模型,以便我们展示临床效用?所以让我真正兴奋的是,在我加入之前,我是计算生物学系或计算机科学系的教授,主要在计算机上工作,看数据、看阵列等等。但一旦来到 Xaira,让我兴奋的是我能和像 Chu 这样的人交谈,还有很多药物猎手,你知道,经验极其丰富的药物猎手,真正理解他们的痛点。所以当我们设计 AI 模型时,我们会思考真正让生物学家兴奋的问题。

There is a AI native company. Uh almost every parts of the sections of drug discovery we trying to use AI to revolutionize how we develop drugs. So the early part we build causal foundation models uh or sometimes we call virtual cell uh proteins we have uh you know state-of-the-arts protein engineering models and we have also patient representation learning uh models and I think what Xaira is trying to do is not only we develop AI models but also we create the right data sets to empower these models and I think what's really make me excited to work at Xaira is we always aim to connect three AI models together instead of letting them work individually by their own. So when we design virtual cell models, we look for connections to that can we find targets that is easier to to apply the protein engineering models and then even when we design the cellular causal models, can we connect to patient representations? what are the right patient data to connect the cellular models so that we have something to show uh clinical utilities. So I think what really make me excited is um before I joined there I'm kind of a professor in computational uh biology department or computer science department where we mostly working on computers we look at the data look at arrays uh etc. But once coming to Xaira, what really excites me is that I get to talk to people like Chu, lots of uh drug hunters, you know, extremely experienced drug hunters to really understand their pinpoint. So when we design AI models, we think about questions that really excites biologists.

X-Cell影响与药物发现挑战 X-Cell's Impact and Drug Discovery Challenges

Bo

稍后我们可以聊聊,在开发出 X-Cell 之后,我收到的一个很有成就感的信号是,生物学家们惊叹,这是第一次,生物学家发现模型能准确预测这些未见过的细胞系对不同扰动的反应。所以真正让我兴奋的是干实验(或 AI 模型)与湿实验(或生物学,甚至最终到临床)的整合。

So later maybe we can talk about how one of the rewarding signals I received after we developed X-Cell is that like the wow moments from biologists that this is the first time biologist actually find the model can predict exactly how these unseen cell lines respond to different perturbations. So that's kind of the part really excites me is the integration of kind of dry lab or AI models to wet lab or the biology or even eventually to the clinical side.

Host

说到这个临床模型。我知道你们的目标是把一种药物一直推进到 FDA 批准,甚至更远。我们现在处于什么阶段?我不知道你是否能谈这个,但你们能从临床试验中收集数据并反馈回来吗?

With this clinical model. Um I know you guys are aiming to you know take a drug all the way to to FDA approval um and beyond. Where do we stand now? I don't know if you're able to talk about this, but like are you able to collect data from clinical trials and tie that back yet?

Bo

正如 B 所说,我觉得如果你想想药物发现过程,其实很简单,对吧?你只需要找到正确的靶点,制造正确的分子,然后找到正确的患者给他们用药。当然,每一步都极其难以做对。到目前为止,就像我刚才说的,这很大程度上依赖试错和猜测。我认为主要问题是我们没有真正合适的生物数据来驱动预测模型的训练。而在蛋白质设计领域,我认为那是我们目前进展最快的领域。

As B said, I think if you think about drug discovery process, it's easy, right? You just need to find the right target, make the right molecule, and find the right patients to give them to. Of course, each of those steps are incredibly difficult to get right. And so far, like I said just now, it relies a lot of on trial and error and guess work. And the main issue I think is that we don't have the right biological data really to power the training of a predictive model. And in protein design space I think that's where we have seen the most rapid progress so far.

Bo

这部分是因为我们有大量高质量数据,是过去 70 多年由整个社区整理出来的。

That's partially because we have a lot of data high quality data over 70 years curated by the entire community.

Bo

人们把蛋白质结构存入一个叫 PDB 的数据库。我们也有很多年从不同基因组收集的序列数据,也能帮助训练模型。正是这些收集和积累的高质量数据,带来了蛋白质设计和 AlphaFold 及其他折叠模型的革命。在其他领域,比如临床模型预测或虚拟细胞,我们远没有同样规模的高质量数据,我认为这主要是数据限制问题。所以回答你的问题,我们正大力投入生成这些数据,特别是实验室中细胞生物学的因果数据,我认为这让我们也能在算法方面创新,从而带来虚拟细胞模型。在患者方面,这是一个非常有趣的问题。也许这是最难获取的数据之一,因为获取高质量的患者样本本身就很难。将其与正确的临床注释匹配,以便你能真正学习分子数据和临床反应之间的桥梁,那就更难了。你也许能跨不同疾病严重程度做到这一点,但要收集正确的数据来预测哪种药物治疗在特定患者身上会或不会起效,就更难了。所以这需要很多思考和仔细的整理来生成数据。我们刚开始进入这个领域,但希望很快能分享更多。

People deposit protein structures into a database right called PDB. We also have a lot of sequence data right collected over the years from different genomes that can help inform the model as well. And it's these high quality data that are collected and accumulated that ushered in this revolution in protein design and AlphaFold and other folding models. In the other domains such as clinical model prediction such as virtual cell we are nowhere near the same kind of massive data that are high quality and I think it's mainly a data limitation issue so to your question that's where we're in very invested in generating these data particularly causal data in cell biology in a lab and that's I think what made it possible to innovate on the algorithm side as well to usher in virtual cell models on the patient side it's a very interesting question. Perhaps that's one of the hardest data to get because getting access to high-quality patient samples is difficult in itself. Getting it matched to the right clinical annotation so that you can actually learn the difference the bridge between molecular data and clinical response. That's even harder. And you might be able to do that across different disease severities, but it will be harder to collect the right data to predict which drug treatment will or will not respond in a particular patient or not. And so that takes a lot of thought and a lot of careful curation to generate data out of. And so we're beginning to go into that area, but hopefully we'll be able to share more soon.

Host

太棒了。也许我们现在该换个话题。你们刚发布了 X-Cell。你们来描述一下吧,我可能会说错。

Awesome. Maybe we should switch gear now. You just released X-Cell. Why don't you guys describe I I'll butcher it.

Bo

X-Cell 是 Xaira 的第一个虚拟细胞模型。它是一个 AI 模型,可以预测对基因扰动的反应。当然我们可以将其扩展到其他类型的干预,比如药物扰动、化学扰动等。那么你能为正在收听的非生物学家描述一下什么是扰动吗?你指的是什么?

X-Cell is Xaira's first virtual cell model. It is an AI model that can predict the response to genetic perturbations. Certainly we can extend it to other type of interventions such as drug perturbations, chemical perturbations, etc. So can you just describe for the nonbiologists in the that are listening what is a perturbation? What do you mean by that?

Bo

在我们的细胞中,当 Bo 谈到基因扰动时,我们的人类细胞通常有 2 万个基因。并非所有细胞都同等地表达每个基因。这就是为什么你的眼细胞、皮肤细胞、心脏细胞,即使它们共享相同的基因组,功能却非常不同。这在很大程度上取决于选择性基因表达,它决定了细胞的类型和状态。所以我们做的是构建一个模型,你可以在计算机中模拟从细胞中删除某些基因,这是一种计算机模拟扰动,也就是说,如果我降低细胞中这个基因的表达,对细胞的其他部分有什么影响,生物学后果是什么。

In our cells when Bo talk about genetic perturbations our cell human cell typically have 20,000 genes. Not all cells express every gene equally. That's why your eye cell, your skin cell, your heart cell, even though they share the same genome, they function very differently. A lot of that's determined by selective gene expression that determine the type and the state of the cell. So what we do is to build a model that you can in silico ablate certain genes from the cell that is a in silico perturbation that's to say if I reduce the expression of this gene in the cell what is the implication for the rest of the cells what's the biological consequence.

Host

你基本上是把一个基因的旋钮调低,没错。

You basically turn the knob down on one gene that's right.

Bo

然后看看那个细胞里所有其他基因会发生什么。

And then that what happens to all the other genes in that cell.

Host

正确,当然希望是预测对所有其他基因的影响,但可能比基因表达更多的东西,比如细胞的功能。

Correct and the hope is of course to predict the effect on all the other genes but maybe even more things than gene expression such as the function of the cell.

Bo

好的。

Okay.

Bo

这很重要,因为这与治疗相关,因为很多药物是抑制剂,它们正是通过降低蛋白质或基因的活性来起作用的。所以如果我们能从基因扰动预测开始,希望我们也能进入通路抑制预测等等。通路就是一组基因,它们之间相互交流,这个基因表达一种蛋白质,那个蛋白质对另一个基因有影响,等等。这是一个基因和蛋白质的长链反应。这就是所谓的通路。如果你中断它或以某种方式改变它,那就会对细胞的更大表型产生影响,即细胞的样子、行为等。

And that's important that's therapeutically relevant because a lot of drugs are inhibitors and they function through exactly that turning down the activity of a protein or gene. And so if we can start with gene perturbation prediction the hope is that we can also go to pathway inhibition prediction so on and so forth. So a pathway is just a set of genes that all kind of talk to each other by this this gene expresses a protein. That protein has some impact on another gene and so forth and so on. There's this long-chain reaction of genes and proteins. And then so that's called a pathway. And so if you interrupt that or somehow change it then that has an impact on the larger phenotype of the cell, what the cell looks like, does etc.

Host

完全正确。

That's exactly right.

Host

是的。所以你们有所谓的虚拟细胞,或者你们正在创建虚拟细胞,而虚拟细胞现在非常流行。很多人对这个概念感兴趣,但我认为你们的方法有些独特,或者与其他人的做法不同。你能解释一下人们说虚拟细胞时通常是什么意思吗?有哪些不同的其他策略,然后你们的具体策略是什么?

Yeah. So you have what you called a virtual cell or you're creating a virtual cell and virtual cells are very popular these days. A lot of people are interested in this concept but I think your approach is somewhat unique or separate from other people are doing. Can you explain what do broadly people mean when they say virtual cells? What are some of the distinct other strategies and then like what is your specific strategy that you're going for?

Bo

当然,虚拟细胞是一个非常高层次的概念,用来描述一个 AI 模型,它能够预测或描述细胞的样子,或预测细胞在特定干预后的表达或功能。这是一个非常高层次的概念。首先,这不是一个新想法。我们大约 20 年前就有虚拟细胞项目,但那时我们有时称之为虚拟细胞 1.0。哦,那是人们试图推导微分方程,用数学来描述某些通路干预的反应,正如你刚才提到的,通过将这些方程拟合到不同的观察结果,但很大程度上这是一次失败的尝试,因为生物学太复杂了,无法用几个预定义的微分方程来写。

Certainly virtual cell is a very high level term to describe an AI model that is able to predict or describe what cell looks like or predict the cell expressions or cell functions after certain interventions. It's a very high level concept. It was first of all it was not a novel idea. We had virtual cell project almost 20 years ago but back then sometime we call it virtual cell 1.0. Oh, is that people trying to derive differential equations to try to use mathematics to describe what's the response for certain pathway interventions as you just mentioned and by fitting these equations to different observations and largely speaking that was a failed attempt in the sense that the biology is just way too complicated to write in a few predefined set of differential equations.

单细胞基因组学基础模型 Foundation Models for Single-Cell Genomics

Bo

随着语言模型的兴起,用数据驱动的方法让 AI 模型模拟细胞对不同干预的反应,这一想法开始流行起来。大约三年前,也就是 ChatGPT 发布后差不多四个月,我们在多伦多大学的实验室发表了单细胞基因组学领域最早的基座模型之一,名为 scGPT。你可以把它理解为一个类似 GPT 的单细胞模型。它很快就流行起来,因为这是我们第一次拥有一个基座模型,能够用同一个模型处理不同的下游任务。例如,我们可以用同一个模型整合不同批次的单细胞 RNA 测序数据,也可以用同一个模型预测多组学整合。

Moving forward with the rise of language models, the idea of using AI models to mimic how cells respond to different interventions through a data-driven approach started to become popular. About three years ago, almost four months after ChatGPT was released, our lab at the University of Toronto published one of the early foundation models of single-cell genomics, called scGPT. You can think of it as a GPT-like model for single cells. It quickly became popular because, for the first time, we had a foundation model that could tackle different downstream tasks using the same model. For example, we could use the same model to integrate different batches of single-cell RNA-seq data, and we could use the same model to predict multiomic integrations.

Host

我们来定义一下这些概念。批次,整合不同批次的 RNA 测序数据。也就是说,你有不同的设备,可能是不同的实验室在收集数据。

Let's define those things. Batches, integrating different batches of RNA-seq. So you have different equipment, maybe different labs collecting data.

Bo

是的,不同的实验室、一天中不同的时间、甚至月相不同,这些因素都会对收集到的数据产生很大影响。所以有一个大问题:当存在所有这些与基因表达无关、仅仅源于测量方式的差异时,我该如何比较这个数据集和那个数据集?

Yeah, different labs, different times of day, different phases of the moon, whatever. Those factors actually have a big impact on the data you collect. So there's a big problem: how do I even compare this dataset to that dataset when there are all these other differences that have nothing to do with gene expression, just how I measured it?

Host

我们称之为批次效应。我们当然希望在去除批次效应的同时保留细胞类型,这是我们要保留的更重要的生物学信息。

We call that batch effect. We certainly want to remove the batch effect while preserving the cell types, which are the more important biology we want to preserve.

Bo

这类似于图像分类器中的坦克问题,模型会捕捉到与真正关心的底层生物学无关的虚假特征。整合不同批次的核心思想是在去除批次效应的同时保留生物信号。在这些基座模型出现之前,在单细胞领域,生物学家必须为每个任务选择所谓的专家级最先进方法。有了像 scGPT 或 Geneformer 这样的基座模型,我们希望带来的是一个能解决单细胞所有任务的统一模型。

This is analogous to the tank problem in image classifiers, where models pick up on spurious features that have nothing to do with the underlying biology you actually care about. The core idea of integrating different batches is to keep the biological signals while removing the batch effect. Before these foundation models, in the single-cell domain, for every task biologists had to choose the so-called specialist state-of-the-art approaches. With foundation models like scGPT or Geneformer, what we hope to bring is one model that solves all tasks in single cells.

Bo

随着基座模型的普及,许多研究人员在 CZI(陈-扎克伯格倡议)下聚集起来,我们在《细胞》杂志上发表了一篇观点文章,首次提出了“虚拟细胞”这一术语——实际上是“虚拟细胞 2.0”。其理念是采用数据驱动的方法:如果我们无法描述它,那就去学习它。这就是虚拟细胞的想法。我们能否构建一个语言模型或类语言模型,来预测细胞类型的样子、细胞对不同干预的反应,并最终通过简单的计算机模拟取代所有细胞实验,甚至无需进行实际实验?

With the popularity of foundation models, many researchers came together under the CZI (Chan Zuckerberg Initiative) and we published a perspective paper in the journal Cell, coining for the first time the term 'virtual cell' — actually 'virtual cell 2.0'. The idea is to use data-driven approaches: if we cannot describe it, let's learn it. That's the idea of virtual cells. Can we build a language model or language-type model to predict what cell types look like, how cells respond to different interventions, and eventually replace all cellular experiments by simply running simulations on a computer without even running the actual experiments?

Host

也许为了提供更多背景,你可以这样想:虚拟细胞只是一个通用概念。细胞内有 2 万个基因,在大多数人类细胞中,大约有 4000 到 5000 个基因在任意时刻通常处于活跃状态,或以合理水平表达。所以你看一个正常细胞,可能有 4000 到 5000 个基因在活动。在许多情况下,药物的作用方式是靶向某个蛋白质,或某种使蛋白质增多或减少的物质,或阻止蛋白质发挥某种功能。你的目标是,给定细胞中的一些基因——每个细胞都有不同的基因组成——会发生什么变化?某些通路会消失吗?某些通路会增强吗?由此,你可以通过理解改变一个特定基因或一组基因如何改变一切,来预测药物将如何起作用。这样理解对吗?

Maybe for a bit more context, you can think about this: a virtual cell is just a general concept. Cells have 20,000 genes in them, and in most human cells, roughly 4,000 to 5,000 are usually active at any given time or expressed at reasonable levels. So you look at a normal cell, you might have 4,000 to 5,000 genes doing things. In many cases, the way medicine works is you target a protein or something that makes proteins more common or less common, or stops the protein from doing something. Your goal is, given some number of genes in a cell — every cell has a different composition of genes — what is going to change? Will some pathway die off? Will some pathway grow? From this, you could predict how a medicine is going to work by just understanding how changing one specific gene or some cluster of genes could change everything. Is that a correct understanding?

Bo

是的,这是对虚拟细胞的一个正确的高层理解。这个领域目前的情况是,我们缺乏对虚拟细胞的具体定义,人们几乎将基座模型等同于虚拟细胞。但在我看来,虚拟细胞可能是一个比基座模型更广泛的概念。基座模型主要提供可靠的、具有语义意义的细胞表征。但我认为虚拟细胞更具动态性:我们能否构建 AI 模型来预测不同细胞状态随时间的发展,甚至描述不同细胞分辨率下的空间变化?在我看来,我们确实还处于开发这种全面虚拟细胞模型的早期阶段,基座模型仅仅是一个起点。

Yeah, that's a correct high-level understanding of virtual cells. What's happening in this field is that we lack a concrete definition of virtual cells, and people almost equate foundation models with virtual cells. But in my view, virtual cells are probably a much broader concept than just foundation models. Foundation models mostly provide reliable, semantically meaningful representations of cells. But I think virtual cells are more dynamic: can we build AI models that even predict the development of different cell states across time, or even describe the spatial changes of cells at different cellular resolutions? In my understanding, we are really at the early stage of developing such comprehensive virtual cell models, and the foundation model is really just the starting point.

Host

AI 模型总是从数据开始。你正在构建一个高通量实验系统,或者已经构建并正在继续开发。这听起来非常酷也非常复杂。你能告诉我们这涉及什么吗?你在做什么?你在运行哪些实验?这如何为构建 AI 模型提供信息?为什么这样做,而不是使用 CLX 基因数据库——一个汇集了公共数据集的基因表达数据集合?

AI models always begin with data. You are building a high-throughput experiment system, or have built and are continuing to develop one. That sounds really cool and really complicated. Can you tell us what that entails? What are you doing? What are the experiments you're running? How does that inform the building of an AI model? And why do this rather than pick up the CLX gene database, which is a collection of gene expression data aggregated over public datasets?

Bo

是的,好问题。我想接着 Bo 的话说。我认为 Bo 说了一些非常深刻的东西:从表征模型——生物学的基座模型——到虚拟细胞。关键区别在于扰动预测或生物学中的动态过程。这是一个因果概念。为此,我认为我们需要因果数据。如果你看看 CellxGene,那是一个很棒的数据集,最初整理了超过 3300 万个细胞,现在远不止这个数。当 scGPT 在 Bo 多伦多实验室的这个数据集上训练时,它主要是一个观察性分析数据集。这是描述性数据,不是因果性的,而且主要分析健康人类捐赠者。因此,在这个数据集上训练的模型非常擅长描述性任务,比如协调批次效应、去除不同实验室和技术的影响。但我认为,我们和该领域的许多其他人都发现,这些在描述性数据上训练的模型在因果任务、扰动任务(我们称之为反事实任务)上,尚未超越线性模型。反事实任务就是:“如果我对细胞做了这个,会发生什么?”这对生物学家来说很直观,因为描述性数据集中的相关性数据可以用许多可能的因果结构来拟合。在一个非常简单的例子中,假设你在描述性数据集中观察到基因 A、B、C 一起上升和下降。你可以推断 A 调控 B 和 C,这就是为什么当 A 上升时,B 和 C 也上升。

Yeah, great question. I want to pick up where Bo left off. I think Bo said something pretty profound: going from a representation model, the foundation model of biology, to a virtual cell. The key difference there is perturbation prediction or dynamic processes in biology. That's a causal concept. For that, I think we need causal data. If you look at CellxGene, that's a fantastic dataset that was curated at the beginning with more than 33 million cells, now a lot more than that. At the time when scGPT was trained on that dataset coming out of Bo's lab in Toronto, it was mostly an observational profiling dataset. It's descriptive data, not causal, and mostly profiling healthy human donors. So the model trained on this dataset is very good at doing descriptive tasks such as harmonizing across batch effects, removing effects from different labs and technologies. But I think both us and many others in the field have found that these models trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks: 'If I did this to the cell, then what would happen?' That makes intuitive sense to a biologist because the correlation data in the descriptive dataset can be fit with many possible causal structures. In a very simplistic case, let's say you observe genes A, B, C all go up and down together in your descriptive dataset. You can infer that A regulates B and C, that's why when A goes up, B and C also go up.

从观测数据学习因果的挑战 Challenges in Learning Causality from Observational Data

Bo

你也可以说 B 调控 A 和 C,这同样完全合理。你还可以说 A 调控 B 和 C,而这一切又由某种不同的东西调控。你看,问题就在这里。有无数种方式可以将一个因果调控网络拟合到这组数据中。从根本上说,我们认为观测数据在真正学习因果关系方面能力不足。这就是为什么我们很早就意识到,我们需要真正开始训练并构建因果数据集,来训练一个因果模型。

You might also say that B regulates A and C, and that would be perfectly reasonable as well. You could also say that A regulates B and C is completely regulated by something different. You see the problem there. And there's n number of ways to fit a causal regulatory network into this group of data. Fundamentally, we believe observational data are underpowered to learn causality truly. And this is why we realized pretty early on that we need to really start training and building causal datasets to train a causal model.

Host

那么有哪些方法可以做到这一点呢?

So what are the ways to do that?

大规模生成因果数据:Perturb-seq Generating Causal Data at Scale: Perturb-seq

Bo

我认为这个领域已经成熟,可以大规模开展这些工作。我们称之为“超生物学”的技术,有很多方法可以大规模生成这些因果数据。我们专注的技术叫做 Perturb-seq。对于不熟悉这项技术的听众来说,它将高通量混合 CRISPR 扰动与单细胞 RNA 技术相结合,构建二维数据集。让我详细解释一下。

I think the field has come of age to do these at scale. A technique that we call hyperbiology, and there are many ways to generate these causal data at scale. The technique that we have focused on is something called Perturb-seq. So for the listeners who are not familiar with that technology, it combines high-throughput pooled CRISPR perturbation together with single-cell RNA technology to build 2D datasets. Let me break that down.

Host

好的,请讲。

Yeah, go please.

理解扰动与基因表达 Understanding Perturbations and Gene Expression

Bo

我们刚才谈到,一个细胞中至少有 2 万个待测量的分析物。这些就是基因。它们既是待测量的特征,也是用来扰动细胞的杠杆。为了清晰起见,我们称它们为扰动和基因表达。扰动在一个轴上,描述细胞的特征在另一个轴上。

So we just talked about in a cell there are at least 20,000 analytes to measure. These are the genes. These are both the features to measure. These are also the levers to perturb the cells with. So for clarity, let's call them perturbations and gene expressions. For perturbation on one axis and the features that you measure that describe the cell on the other axis.

Perturb-seq原理:CRISPR与混合实验 How Perturb-seq Works: CRISPR and Pooled Experiments

Bo

Perturb-seq 是一项利用实验室生物学最新突破的技术,即 CRISPR-Cas9。这些是源自细菌的酶,可以破坏哺乳动物细胞(例如人类细胞)中的基因表达,而且我们可以逐个进行。所以我可以一次敲除一个基因。当然,如果我想在单个实验中完成全部 2 万个基因的敲除,那将极其难以扩展。可能需要一个巨大的工厂和很多机器人才能做到。或者你可以用混合的方式。我喜欢混合实验。这些实验具有超强的可扩展性。我们有实验室技巧,可以做到每个细胞破坏一个基因,但在一个单一的混合实验中,跨越许多许多细胞完成全部 2 万个基因的扰动,并且完全随机化。这样就没有批次效应,没有板间差异。基本上,我们利用某种组合技巧,首先以不同的组合扰动所有不同的基因,然后你可以读出结果并进行一些数学处理,基本上你就在一个实验中获得了大量不同的实验。

Perturb-seq is a technique that leverages the latest breakthrough in lab biology, CRISPR-Cas9. These are bacterially derived enzymes that allow you to disrupt gene expression in mammalian cells, in human cells for example, and we can do so in a one-at-a-time fashion. So I can take out one gene at a time. Of course, that would be incredibly difficult to scale if I want to do all 20,000 gene expression knockouts in one single experiment. Probably need a huge factory, a lot of robots to do that. Or you can do them in a pooled fashion. And I love pooled experiments. These are hyperscalable. So we have lab tricks that allow us to disrupt one gene per cell, but do all 20,000 genes across many many cells in one single pooled experiment, perfectly scrambled. So there's no batch effects. There's no plate-to-plate variation. Basically, we use some sort of combinatorial trick to first perturb all the different genes in different combinations, and then you can read them out and do some math on it, and you basically pull out a whole bunch of different experiments in one experiment.

Host

正确。这需要条形码技术,而条形码实际上是通过直接读出每个细胞中存在哪种 CRISPR 引导 RNA 来实现的。对于 CRISPR-Cas9,这种源自细菌的机制要在数百万个细胞中工作,你只需要向每个细胞递送两样东西。你需要递送蛋白质,即执行工作的 Cas9 蛋白,以及一个由一段称为引导 RNA 的短 RNA 编码的地址条形码。引导 RNA 通过 Watson-Crick 碱基配对(ATCG)纯粹地告诉蛋白质在细胞中去哪里。因此,它与基因的一部分匹配。它足够长,可以说这将匹配正确的基因,然后引导它连接到正确的基因并减少该特定基因在细胞中的表达。

Correct. It requires barcoding technology, and that barcode is actually achieved by directly reading out what kind of CRISPR guide RNA is present in which cell. So for CRISPR-Cas9, this bacterially derived machinery to work in millions of cells, you just have to deliver two things to each cell. You have to deliver the protein, the Cas9 protein that does the job, and you have to deliver an address barcode encoded by a short piece of RNA called guide RNA. And the guide RNA tells the protein where to go in the cell purely via Watson-Crick base pairing, ATCG. So it matches a part of the gene. It's sufficiently long to say this will match the correct gene, and then that guides it to connect to the right gene and reduce the expression of that particular gene in the cell.

设计引导RNA与基因沉默 Designing Guide RNAs and Silencing Genes

Bo

正确。我们设计这种引导 RNA 去往基因的启动子区域。那是每个基因在转录开始前的一段起始序列。如果我们带着正确的引导 RNA(即沉默器)将 Cas9 蛋白带到那里,启动子就会被关闭,该基因将永远不会再被转录。因此,我们有效地调低了该基因的表达水平。所以,你只需要知道哪个引导 RNA 在哪个细胞中,这可以通过基因组读出完成。这就是条形码,然后你可以推断哪个基因在哪个细胞中被沉默了。这就是在扰动方面扩展通量的方式。

Correct. We designed this guide to go to the promoter part of a gene. That's the beginning stretch of every gene before the transcription starts. And if we bring the Cas9 protein in there armed with the right guide, the silencer, that promoter will get shut off and that gene will never be transcribed again. So effectively we tune down the expression level of that gene. And so all you have to know is figure out which guide RNA is in which cell, and that can be done using genomic readouts. That's the barcode, and you can then infer which gene is being silenced in which cell. So that's the way you scale throughput on the perturbation side.

单细胞RNA-seq扩展读数 Scaling Readout with Single-Cell RNA-seq

Bo

在读出方面,这是一个二维数据集,对吧?我们刚才谈到了一个维度。在读出方面,我们利用单细胞 RNA-seq 技术。这些也是过去十年中发展起来并扩展的技术,可以让你同时读出每个细胞中全部 2 万个基因的表达水平。因此,有了高通量 CRISPR 扰动和高通量单细胞 RNA-seq 技术,我们突然就能生成这些二维数据集,系统地扰动或敲除/敲低细胞类型中人类基因组的每一个基因,并读出对同一细胞中其他所有基因的影响。因此,我们生成这些丰富的二维数据集,与训练 AlphaFold 模型的 PDB 数据在规模和类型上没有太大区别,对吧?想想看,那是数十万个蛋白质条目。如果那些是行,列就是每个氨基酸的 XYZ 坐标。那也是二维数据集。我认为正是这些丰富的二维数据集为生物学基础模型的训练提供了动力。

On the readout side, it's a 2D dataset, right? So we just talked about one of the dimensions. On the readout side, we leverage single-cell RNA-seq technologies. So these are also recent technologies in the last decade that have been scaled that can let you read out expression levels of all 20,000 genes simultaneously from each cell. So armed with both high-throughput CRISPR perturbation and high-throughput single-cell RNA-seq technologies, all of a sudden we can generate these 2D datasets where we systematically perturb or knock out/knock down every single gene in the human genome in the cell type, and we read out the impact on every other gene in the same cells. So we generate these 2D rich datasets, not that different than the size and type of PDB data that trained AlphaFold models, right? If you think about that, that's hundreds of thousands of protein entries. If those are the rows, columns are the XYZ coordinates of every single amino acid. That's also a 2D dataset. And I think it's these types of rich 2D datasets that power the training of foundation models of biology.

扩展至百万细胞:工程挑战 Scaling to Millions of Cells: Engineering Challenges

Host

我觉得很有趣的是,你们把一个相当简单的检测方法,我理解这是 NGS 测序,对吧?下一代测序,非常高通量。你们用这个来扩展一个简单的扰动响应,单独来看可能并不那么有趣,但扩展到了基本上任意数量细胞的巨大规模。我想你们做了 2500 万左右。

I find it really fun how you have turned a fairly straightforward assay in using, I get this is NGS sequencing, right? Next-generation sequencing, very high throughput. You've used this to scale a simple perturbation response, which is individually maybe not all that interesting, to this massive scale of over basically an arbitrary number of cells. I think you did 25 million or something.

Bo

实际上远不止这个数。2500 万是经过最严格质量过滤后得到的结果。

So it's actually a lot more than that. So 25 million is what came out of the most stringent quality filtering.

Host

实际上,弄清楚如何进行 CRISPR 和单细胞 RNA-seq 既是一个科学挑战,也是实验第一部分的一个工程挑战。我们经常需要收获数千万甚至数亿个细胞,它们经过各种质量漏斗,最终为 Bo 和团队提供最高质量的数据。这非常困难,因为你可以想象,所有这些技术之前都由学术界发表过,它们在小规模实验中效果很好。但当你考虑将它们扩展到全基因组扰动时,我们说的是处理数亿个细胞。学术界发表的技术通常都是处理新鲜细胞的。细胞仍然是活的,如果你的整个实验只需要一两个小时,那可能没问题。但要处理跨越 14 小时工作日的细胞就不太容易了。那是数亿个细胞。所以一天结束时,我常和团队开玩笑说,你可以很容易地从细胞和实验室里的科学家身上检测到压力信号。

It's actually as much of a scientific challenge to figure out how to do CRISPR and single-cell RNA-seq as it is an engineering challenge in the first part of the experiment. Often times we have to harvest tens if not hundreds of millions of cells, and they go through various quality funnels to arrive to give Bo and team the highest quality data at the end. That's incredibly difficult to do because, as you can imagine, all of these techniques have been published by academia before, and they work very well in small-scale experiments. But when you think about scaling them to a genome-wide perturbation, we're talking about handling hundreds of millions of cells. Techniques that are published in academia used to be all about handling fresh cells. Cells are still alive, and that may be okay if your entire experiment takes only an hour or two. It's not quite easy to handle cells across a 14-hour day. That's hundreds of millions of cells. And so by the end of the day, I used to joke with my team, you can easily detect stress signals from the cells and from your scientists in the lab.

Host

是的。

Yeah.

Bo

很快我们意识到这不是进行数据生成的方式。机器学习对质量非常敏感,我们希望为 AI 团队提供最高质量的数据。因此,我们投入了大量的工程思考,并逐步将整个工作流程工业化。

And quickly we realized that's not the way to do these data generation. Machine learning is very quality-dependent, and we want to give the highest quality data to our AI teams. So we're putting a lot of engineering thought and industrializing the whole workflow step by step.

化学固定与干细胞 Chemical Fixation and Stem Cells

Host

引入化学固定,以便在实验开始时锁定细胞状态,但要找到不影响后续所有生物学、分子生物学步骤的方法。它不影响数据质量,这样我们就能以时间错开的方式生成所有数据,且不易受批次效应影响。你有一点没提到,就是你用了某种干细胞。而且显然,你没有脑细胞或血细胞,如果有的话,那会产生巨大的组合效应。那么你为何确信在干细胞上工作——据我理解,实际上有些血细胞被解锁了干细胞行为,这也会给细胞带来某种压力——你为何确信这是研究脑细胞或其他细胞的好代理?

Introduce chemical fixations so that we lock the state of the cells in at the beginning of this experiment but figure out ways that it doesn't disrupt all of the biology molecular biology steps afterwards. It doesn't impact data quality so that we can do all of these data generation in a timeshifted operational manner that's very not prone to batch effects. One thing that you didn't mention is that you're using some sort of stem cells. And so obviously like you don't have brain cells or blood cells or if you did then you would have a big combinatorial effect on that. So how are you convinced that working on stem cells, which are, my understanding is that there are actually blood cells that have been sort of the stem cell behavior has been unlocked on them and that causes some sort of stress on the cell as well. So you have these like sort of not quite blood cells that are stressed and then why are we convinced that that is a good proxy for a brain cell or whatever you're studying?

Bo

嗯,不完全是。我们其实不是从干细胞开始的,那是后来的发展。我们开始生成数据时,发布了我们讨论的方法以及前两个数据集,那是去年六月预印本中当时全球最大的扰动数据发布。我们称之为 X-Atlas/Orion 数据集,它实际上来自两个细胞系。这个领域的很多早期工作都是从细胞系开始的。这些是癌细胞系。其中一个是癌细胞系,另一个只是细胞系。这些是永生化细胞。有些来自癌症,因此称为癌细胞系;其他则来自原代细胞,但已被永生化,并在各个实验室中培养了多年。你可以想象,人们一开始用这些细胞系,因为它们容易操作,容易规模化,容易培养出大量细胞。事实证明,能培养数百万个细胞对进行这些大型实验至关重要。所以我们先从这些开始,它们实际上仍然保留了源自结直肠癌和造血细胞的细胞类型特征。但在最近的预印本中,我们扩展到了更多细胞类型。现在有些仍然是细胞系,比如我们选择用细胞系做 T 细胞,但有些已经进入原代细胞。所以我们在 iPSC(诱导多能干细胞)中做了一个实验,另一个实验是——我们认为这是迄今最雄心勃勃、最酷的筛选——这是一个多细胞类型干细胞分化项目。实际上,我们在一个实验中无限制地将 iPSC 分化为 10 种不同细胞类型,并对它们进行了全基因组规模的扰动。你可以想象,我们不是只做了 1 万个不同的生物学实验,而是做了 1 万乘以 10 种细胞类型,几乎是文库对文库的实验。我们为什么这么做?我们认为在数据收集的初期阶段,正如我们所说,我们仍处于虚拟细胞构建的早期,数据的背景、多样性和丰富性很重要。这不仅仅是细胞总数或测序读段总数的问题,而是每美元比特数和信息含量。所以我们不仅要在遗传扰动图谱上扩展,还要在生物学背景上扩展,这样我们才能给 AI 团队提供最丰富的数据集,以构建可泛化的模型。

Yeah, not quite. So we didn't actually start with stem cells. That was more of a later development. When we started data generation, we put out the method that I talk about as well as the first two data sets, which is the world's largest perturbic data release at the time last June in the preprint. We call a data set X-Atlas/Orion that was actually generated from two cell lines. A lot of this field's early work started with cell lines. These are cancer cell lines. One of them is a cancer cell line, the other is just a cell line. These are immortalized cells. Some of them are derived from cancers, hence cancer cell lines. Others are just derived from primary cells but have been immortalized many times grown for many years in various labs. People start with these cell lines in the beginning, as you can imagine, because those are easy to do. It's easy to scale, easy to grow a lot of cells out of. Turns out the ability to grow millions of cells is actually critical for doing these large experiments. So we started there first and they actually still capture the characteristics of the cell types that are derived from colorectal cancer as well as hematopoietic cells. But later on, in the most recent preprint, we actually expanded to many more cell types. Now some of these are still cell lines, our T-cells we chose to use cell lines, but some of these have now gone into primary cells. So we did one experiment in iPSC, these are induced pluripotent stem cells, and another experiment in, and we think this is the most ambitious and coolest screen that we've done to date. This is a pen differentiation multi-cell type stem cell project. So effectively we differentiate iPSC into 10 different cell types in one single experiment without restriction and we did a genome scale perturbation across them. So you can imagine instead of just generating 10,000 different biological experiments we did 10,000 by 10 cell types, was almost a library on library experiment. Why are we doing this? We think that in the beginning phase of data collection, as both said, I think we're just in the early days of virtual cell building, context and diversity and richness of the data matters. It's not just the total number of cells or total number of sequencing reads. It's about bits per dollar and information content. So we want to scale not only in the genetic perturbation landscape but we also want to scale across biological context so that we can give our AI teams the best rich data set to build a generalizable model on.

空间背景与虚拟细胞模型 Spatial Context and Virtual Cell Models

Host

你有没有考虑过——你说背景,但显然这些实验中的细胞已经被某种方式——我忘了术语——从它们的同伴中分离出来了,对吧?有没有考虑使用空间转录组学或其他基于成像的技术来构建模型,在扰动的同时考虑细胞邻近的背景?

Is there any thinking about, so you say context, but obviously these cells in these experiments have been sort of de—I forget the term—but they've been separated from their cohorts, right? Is there thinking about using spatial transcriptomics or other imaging-based technologies to build models with perturbations but in the context of the cells that it lives near?

Bo

好问题。我们正从几个方面考虑这个问题。第一,这实际上正是我们首先想要构建虚拟细胞模型的原因。你可能会想,既然你可以在这些细胞系中进行全面筛选,为什么还需要模型?你可以直接做实验并生成数据。当然,如果你的问题只是关于细胞系中的细胞生物学,你说得对,我们不需要模型,对吧?至少对于遗传筛选,我们可以直接做实验。但你也说得对,很多时候好的靶点、生物学洞见并不来自细胞系,而是来自原代细胞、来自器官中天然生理环境中的细胞,甚至多器官共同作用产生的涌现特性。很多神经疾病就是这样。你无法在动物系统、器官或所有这些复杂的转化模型中进行详尽的高通量实验。你可以做一些实验,但这些实验昂贵且高风险。构建一个模型,可以在可扩展的大规模数据上训练,并通过微调和迁移,在这些复杂模型中做出高质量的因果预测,这样我们就能进入实验室,验证最高质量的假设。我认为这就是构建虚拟细胞模型的意义所在。但从 AI 的角度,我认为你完全正确,我相信未来的虚拟细胞模型应该能够整合多种模态,而不仅仅是 RNA 表达。空间单细胞 RNA 测序已经是热门技术,甚至对 SGBT 也是如此。我们实际上有一个扩展版本,称为 SGBT spatial,专门为空间单细胞设计。我们还有关于早期尝试用 H&E 图像预测基因表达的论文,已经能发现一些信号。所以最终我预测,虚拟细胞模型将不仅能整合 RNA,还能整合更多功能相关的数据,例如蛋白质组学或其他调控组学(如表观遗传学),将所有描述性组学数据集结合起来,预测细胞功能的未来状态。我认为这可能是虚拟细胞模型的未来。

Great question. So we're thinking about that in a couple ways. Number one, that's actually exactly why we want to build a virtual cell model in the first place. You might think that well you can already do exhaustive screening in these cell lines, why do you still need a model? You can just do the experiment and generate the data. Certainly, if your query is just about cell biology in cell lines, you're right. We don't need a model, right? At least for genetic screening, we can just do this experiment. But you're also correct that often times good targets, biological insights are not about cell lines. These are about primary cells, about cells in their native physiological context in organs or even multiorgan coming together and having some emerging properties. A lot of neurological disease are that way. You cannot do exhaustive high throughput experimentation in animal systems or in organs or in all of these complex translational models. You can do some experiments and these are expensive and high stake. The ability to build a model that can be trained on massive data where it is possible to scale and be trained in a way that can be fine-tuned and transferred to make high quality causal predictions in these complex models so that we can go into the lab and have the highest quality hypothesis possible to validate. I think that's the whole point about building a virtual cell model. But from AI side, I think you're absolutely right that I believe the future virtual cell model should be able to incorporate multiple modalities not just RNA expressions. Spatial single cell RNA-seq is already a popular technology even for SGBT. We actually have an extended version. We call it SGBT spatial that is specifically designed for spatial single cells. And we also have papers on early attempts to try to take the H&E images to predict the gene expressions. There's already some signals you can find. So eventually what I predict is that virtual cell model will be able to integrate not only RNA, can integrate more functionally related, for example proteomics or other regulatory side of omics such as epigenetic, to overall combine all your descriptive omics data sets to predict the future states of the cellular functions. I think that's probably the future for virtual cell model.

空间转录组学与蛋白质组学解析 Spatial Transcriptomics and Proteomics Explained

Host

我们最近邀请了 Noetic 的 Ron Alpha 和 Dan Bear 作为嘉宾,想多了解一点的观众,我想我们在那里讨论得很深入。如果你想了解背景,可以去看看。但你能解释一下什么是空间转录组学和空间蛋白质组学吗?

We had Ron Alpha and Dan Bear from Noetic as guests recently, and viewers who want to hear a little bit more about that, I think we go quite in depth there. So if you want background you can go to that. But can you explain a little bit about what spatial transcriptomics and spatial proteomics are?

Bo

那么这里可能先讲一点历史。在我们有单细胞 RNA 测序之前,我们有 RNA 测序,在那之前我们有微阵列技术。

So maybe a bit of a history lesson here. Before we had single-cell RNA-seq, we had RNA-seq, and before that we have microarray technologies.

从批量到单细胞再到空间组学 From bulk to single-cell to spatial omics

Bo

RNA seeker 微阵列过去做的事情是取一块我的组织,把它全部磨碎,放进搅拌机,想象一下做成一杯冰沙,然后从那块组织中的不同细胞里取出所有 RNA,测量它们的表达水平。第一次能同时测量全部 2 万个基因的表达,这很棒。我们以前不得不一个一个地做,但不好的是我们不知道哪个 RNA 来自哪个细胞。如果你处理的是多细胞组织,这尤其是个问题。你想把 RNA 归到免疫细胞、皮肤细胞、成纤维细胞、角质形成细胞上,但你做不到,因为你把一切都磨成了冰沙。单细胞技术让你能做的是逐个细胞分析。所以现在我可以把 RNA 基因表达归到它们来源的细胞上。但还有一个问题:我不知道信号在空间上来自哪里。对许多疾病来说这很重要。比如在免疫肿瘤学中,你想知道 T 细胞何时靠近肿瘤细胞,或者 T 细胞何时无法穿透实体瘤,它们之间有什么区别;或者 T 细胞攻击肿瘤细胞与不攻击时,它们之间有什么区别。为此你需要空间信息,你需要原位观察细胞在其环境中的状态。所以现在有不同的技术来解决这个问题。本质上,取那块组织,我不需要再磨碎它。我只需做一个横截面,把它放在一张玻片上,我可以用标准技术如 H&E 染色来测量它的形态。我还可以用多重免疫荧光分析来测量许多蛋白质表达。最终,我还可以通过一些最新的空间组学分析,在所有细胞的天然空间坐标上观察全基因组范围的基因表达。所以你有了每个细胞的 XY 坐标,还有我们之前谈到的所有分子分析物。这对基因组学领域来说是一个令人兴奋的新方向。

What RNA seeker microarray used to do is take a chunk of my tissue, grind it all up, put it in a blender, imagine making a smoothie out of it, and take all of the RNA from different cells in that piece of tissue and measure all of their expression levels. It is great for the first time you can measure gene expression all 20,000 at a time. And we used to have to do them one at a time, but it is not great in that we don't know which RNA came from which cell. And this is particularly a problem if you're dealing with a multicellular piece of tissue. You want to attribute RNA to the immune cell, to the skin cell, to the fibroblast, to the keratinocytes, but you can't because you grind everything up in a smoothie. What single cell technology allows you to do is analyze them cell by cell. So now I can attribute RNA gene expression to the cell that they originate from. But there's still a problem. I don't know spatially where the signal comes from. And for many diseases it matters. In immuno-oncology, for example, you want to know when T cells are close to a tumor cell, or when a T cell is not able to penetrate the solid tumor, what is the difference between them, or when a T cell is attacking the tumor cell versus when it is not, what is the difference about them. And for that you need spatial information. You need to observe cells in situ in their context. And so now there are different technologies that solve that problem. Essentially, take that chunk of tissue. I don't have to grind it up anymore. I just make a cross-section, lay it down on a piece of slide, and I can measure its morphology using standard techniques like H&E staining. I can then also measure many protein expressions using multiplex immunofluorescence assays. Ultimately, I can also look at the gene expression up to genome-wide in all of these cells in their native spatial coordinates by using some of the latest spatial omics assays. So you have the XY coordinates of every cell, but also all of the molecular analytes that we talked about earlier. And that's an exciting new direction for the genomics field.

Host

就我个人而言,你可以想象空间组学给 AI 建模增加了更多难度,因为你不是看单个细胞,而是必须看邻近的微环境细胞,才能更好地学习具有空间一致性的表征。这是当前空间基础模型面临的挑战。但那个环境对于正确理解癌症等疾病至关重要,比如免疫细胞、癌细胞和非免疫细胞之间的相互作用,对于预测某些临床反应的空间感知生物标志物至关重要。我认为构建这样的模型极其重要。

Personally, you can imagine that spatial omics adds more difficulty to AI modeling because instead of looking at individual cells, you have to look at the neighboring niche cells to better learn the representation that is spatially cohesive. That is the challenge the current spatial foundation models are facing. But that context is going to be crucial for correct understanding, let's say, cancer, where the interaction of immune cells and cancer cells and non-immune cells is crucial for spatially aware biomarkers to predict some of the clinical response. I think that would be extremely important to build such models.

Bo

回到 X-Cell,这你知道大概也能为空间模型提供信息,对吧?你可以让一个细胞在一个位置。你可以想象,好吧,我可以直接扔掉坐标,一次只对一个细胞做推理。现在我可以创建一个更复杂的模型来做这件事,但它也知道它的邻居是谁。

Getting back to X-Cell, this you know presumably can inform a spatial model as well, right? You can have one cell in one place. If you can imagine, okay, I can just throw away the coordinates and just do inferences on one cell at a time. And now I can create a more complicated model that does that but it also knows who its neighbors are.

Host

你说得完全正确。但我们当前发布的版本并不处理空间组学。不过,我们正在进行的工作以及下一版 X-Cell 肯定能够推断不同细胞的空间感知表征。

You're absolutely right. But the current version we are releasing, we are not dealing with spatial omics. However, definitely our ongoing work and the next version of X-Cell will be able to infer the spatially aware representations for different cells.

Bo

我明白了。

I see.

Host

我们稍微谈了一下数据收集。让我们谈谈架构。给正在听的 AI 工程师们来点干货。

We've talked about the data collection a bit. Let's talk about the architecture. Get some red meat for the AI engineers listening in.

从自回归到扩散语言模型 From autoregressive to diffusion language models

Bo

当然。让我们回顾一下虚拟细胞建模的历史,特别是虚拟细胞 2.0。我认为我们的 scGPT 为大多数单细胞基础模型奠定了基础。我们采用了自回归训练,与 ChatGPT 在语言上的训练方式极其相似,对吧?我们使用下一个词预测。所以我们模仿 ChatGPT 在语言上的训练方式来训练一个基于所有细胞的单细胞基础模型。这样做,他们必须假设基因有一个内在顺序,对吧?我们假设基因顺序的方式是通过注意力机制。还有许多其他方法使用不同的基因顺序。有些简单到只是根据表达值对基因进行排序。也有更复杂的方法来对不同的基因排序。但本质上你必须假设一个基因顺序。

Sure. Let's get to the history of virtual cell modeling, particularly virtual cell 2.0. I think our scGPT kind of sets the foundation for most of the foundation models of single cells. We adopted autoregressive training, extremely similar to how ChatGPT is trained on languages, right? We use next-token prediction. So we mimic the way ChatGPT is trained on languages to train a single-cell foundation model on all cells. By doing that, they have to assume an inherent order of genes, right? The way we assume the order of genes is by attention mechanism. There are many other methods that use different orders of genes. Some are as simple as just ranking the genes based on the expression values. There are also more complicated methods to rank different genes. But inherently you have to assume an order of genes.

Host

而且为了明确起见,当你说基因时,它们本质上是有序的,对吧?它们是用 ATGC 拼写出来的句子,对吧?所以基因本身有核苷酸,还有这条长链,对它们有一个顺序是很有意义的。但我们谈论的是不同的东西。那是表达数据,表达水平。所以表达水平意味着我在测量时看到了这些基因中的多少个。对于 DNA 序列,ATG 的顺序对我们来说完全合理,对吧?但对于表达数据,它们实际上只是矩阵。所以很难假设基因有一个内在顺序。即使你打乱基因的顺序,我认为生物学不会改变太多。然而,由于语言模型的训练方式,每个人都有某种预设的技巧来训练这样的模型,所以很容易采用。这就是所有单细胞基础模型的起步方式。

And just to be clear, so when you talk about genes, those are intrinsically ordered, right? They're a sentence spelled out in ATGC, right? So genes themselves have the nucleotides and there's this long chain and that makes a lot of sense to have an order to them. But what we're talking about is something different. That's the expression data, the expression levels. So the expression level means how many of these genes did I see when I was measuring. For DNA sequences the order of ATG makes total sense to us, right? But for expression data they're literally just matrices. So it's really hard to assume an inherent order of genes. Even if you shuffle the order of genes, I think the biology doesn't change much. However, because of the way language models are trained, everybody has kind of preset tricks to train such a model, so it's easy to adopt. That's how all the foundation models are started for single cells.

Bo

然后我很快意识到,使用扩散语言模型,我们实际上不需要假设基因的顺序。相反,我们可以有一个双向扩散过程来生成这种长的高维基因表达数据集。所以想一想,自回归训练和扩散语言模型之间的区别是什么?你可以把自回归训练想象成打字,例如“我喜欢咖啡”——你必须先打“我”,然后“喜欢”,然后“咖啡”,有一个内在顺序。但扩散语言模型,你可以把它当作编辑。你从一个非常模糊、粗糙的句子迭代生成一个句子,然后你可以迭代地完善它。基因表达也是如此,你可以生成一个非常粗糙的基因表达表征,然后从嘈杂的表征迭代到更精细的表征。所以你迭代地编辑基因表达预测,直到它最小化损失。所以这是一种非常不同的生成式预测扰动后响应的哲学。事实证明它实际上更适合单细胞组学。这就是为什么我们从类似 scGPT 的模型转向当前的 X-Cell 模型,它使用扩散语言模型。

And then I quickly realized that with diffusion language models we actually don't need to assume the order of genes. Instead, we can have a bidirectional diffusion process to generate such long high-dimensional gene expression data sets. So just to think about it, what's the difference between autoregressive training versus diffusion language models? You can think of autoregressive training as typing, for example, 'I like coffee' — you have to type 'I' and then 'like' and then 'coffee', there's an inherent order. But diffusion language models, you can treat it as editing. You iteratively generate a sentence from a very vague, rough sentence, and then you can iteratively refine it. So same thing with gene expression, you can generate a very rough representation of the gene expressions and then iteratively from noisy representation to more refined representations. So you kind of iteratively edit the gene expression predictions until it minimizes the losses. So this is a very different philosophy to generatively predict the response after perturbation. And it turns out it actually fits more to single-cell omics. So that's why we switched from scGPT-like models to the current X-Cell model, which uses diffusion language models.

Host

当我想到 Transformer 时,它们本质上是作用于集合的对象。社区花了很多时间试图让它们具有某种因果顺序。

When I think of transformers, they're fundamentally objects which operate on sets. The community spends a lot of time trying to make them things which have some sort of causal ordering to them.

架构选择:扩散vs自回归 Architecture Choice: Diffusion vs Autoregressive

Host

嗯,但如果你只是天真地拿一个 Transformer,它本质上是一种集合操作,对吧?既然如此,为什么要从扩散模型或自回归大语言模型的角度来思考呢?为什么不让你的初始预测策略变成类似“取一组基因,每个基因都有自己独热编码的身份”这样的方式,然后把它用作一种预测?这在我看来是更自然的架构。我知道你和很多人在做类似的事情,我一直有点困惑,为什么社区里会有这种偏见。

Uh, but if you just naively take a transformer, it's a set operation, right? So given that, why think about this in terms of diffusion or autoregressive LLMs? Why not have your initial prediction strategy be something like taking just a set of genes, each of which has its own one-hot encoded identity, and then use that as a sort of prediction? That seems like a much more natural architecture to me. And I know you and a lot of people work on things like this, and I have been somewhat confused why there's this bias in the community about this.

Bo

这是个好问题。我觉得你提到的更接近表示学习,也就是取一组基因,然后尝试把它们投影到低维潜在空间。但我们构建虚拟细胞生成式建模时,关心的是预测细胞的动态变化,所以我们想要生成式模型。这就是为什么我们主要使用仅解码器架构来生成完整的转录组,而不是只预测一个预先定义的小基因集,因为你要建模整个基因调控网络,而那是极其高维的,对吧?

That's a good point. I think what you're referring to is more related to representation learning, where you can take sets of genes and try to project them to a low-dimensional latent space. But what we care about for building generative modeling for virtual cells is predicting the dynamics of cells, so we want to have generative models. That's why we mostly use decoder-only architectures in order to generate the full transcriptomics instead of just predicting a predefined small set of genes, because you want to model the whole gene regulatory networks, which are extremely high-dimensional, right?

Host

所以确认一下,输入是基因加扰动,输出是新的基因表达水平。是基因表达水平加扰动作为输入,输出是……

So just to be clear, input is genes plus a perturbation, output is new gene expression levels. Is that gene expression levels plus perturbation is input, output is...

Bo

细胞,比如对每个细胞。

Cells, like for each cell.

Host

对,对,就是这样。嗯。

Correct. Correct. That is correct. Yeah.

Host

好的。那么,我对扩散语言模型的理解是——你可以纠正我,因为我对它们了解不多——我觉得它们就像 BERT,但你会反复做同样的事情。这个理解大致对吗?

Right. Okay. And so the way that I think about diffusion language models, and you can correct me here because I don't know a lot about them, but the way I think about them is they're like BERT but you do it over and over again. Is that kind of a good...

Bo

嗯,这是对扩散语言模型工作原理的大致理解。

Yeah, that is a rough understanding of how diffusion language models work.

Host

嗯。所以你只是反复应用扩散过程,基本上就是反复去掩码或编辑。类似于图像扩散模型反复细化图像。在这种情况下,它基本上是一个 Transformer。它是一个 Transformer,但它会反复更新这个“句子”——在这里是一堆表达水平——一遍又一遍。

Yeah. So you just apply the diffusion process, which is basically unmasking or editing, correct, once over and over again. Similar to how an image diffusion model kind of refines the image over and over again. In this case, it's a transformer basically. It is a transformer, but it's like repeatedly updating the sentence, in this case a bunch of expression levels, over and over again.

Bo

没错。实际上,在我们的论文中,我们展示了随着扩散步数的增加,损失函数持续下降,预测与真实值的拟合度持续上升。这意味着模型开始理解如何迭代地细化预测。

That is correct. Actually, in our paper we show that as the number of diffusion steps goes on, the loss function keeps decreasing, and the fitness of the prediction to the ground truth keeps increasing. This means the model starts to understand how to iteratively refine the predictions.

Host

我明白了。我们刚才在讨论扩散模型和自回归模型的对比。我注意到论文里有很多关于使用各种东西进行预条件处理的讨论。你能稍微谈谈这个吗?

I see. We're talking about diffusion versus autoregressive. I noticed in the paper there's a bunch of discussion of preconditioning using a whole bunch of stuff. Can you talk a little bit about that?

Bo

我们在 X-Cell 中做出的另一个重大创新,是将先验知识融入模型的方式。在生物学中,融入生物先验通常是个好主意,因为生物学家已经花了几十年去理解一些生物学规律。我们如何告诉模型一些先验知识、一些关于细胞的元数据呢?在 X-Cell 之前,人们通常只尝试融入单一类型的先验,比如用基因调控网络作为先验来预测扰动,有时也会尝试融入 PPI 作为先验。据我所知,X-Cell 是最早尝试融入极其多样化生物先验的模型之一。在我们的预印本中,我们融入了五种类型的先验,包括文献——简单到直接问 ChatGPT“告诉我关于这个基因的一切”,然后我们把它嵌入为 GenePT 的嵌入。没错,就是 GenePT。我们还融入了 PPI,即蛋白质-蛋白质相互作用网络;我们还融入了 DepMap,这是与癌症相关的必需基因信息、形态学信息;我们甚至尝试融入 scGPT 嵌入,这基本上就是细胞类型。通过将一组先验知识作为模型的条件,模型在上下文特定预测方面变得更加准确。更有趣的是,通过观察不同先验的权重,我们实际上可以理解哪些先验知识对特定细胞类型更重要。这增加了模型的可解释性。我们发现,将扩散语言模型与非常多样化的先验集结合,X-Cell 在泛化到未见过的上下文方面表现更好。这就是我们为 X-Cell 做的一些 AI 创新。

Another major innovation we made in X-Cell is the way we incorporate prior knowledge into the model. Incorporating biological priors has always been a good idea in biology in general, because biologists spend decades to understand some of the biology already. How do we tell the model some of the prior knowledge, some metadata about the cells? Before X-Cell, what people do is they try to incorporate a single type of prior, for example using gene regulatory network as a prior to predict the perturbations. Sometimes trying to incorporate PPI as a prior as well. X-Cell, to my knowledge, is one of the first models that tries to incorporate an extremely diverse set of biological priors. So in our preprint, we incorporate five types of priors, including literature as simple as just ask ChatGPT to tell me everything about this gene, and then we embed that as the embedding to GenePT. Exactly, that's GenePT. And we also incorporate PPI, protein-protein interaction networks. We also incorporate DepMap, which is cancer-related essential gene information, morphology information. We even try to incorporate scGPT embeddings, which is basically cell types. So with a set of prior knowledge as conditions to the model, the model starts to have more accuracy in terms of context-specific predictions. And what's more interesting to us is that by looking at the weights of different priors, we can actually understand which prior knowledge is more important to these particular cell types. So it adds more interpretability to the models. So we find that combining diffusion language model plus a very diverse set of priors, X-Cell does much better in generalizing to unseen contexts. So this is some of the AI innovations we made for X-Cell.

Host

现在模型工作时,你需要提供所有这些上下文吗?还是说这些就像预条件处理,如果你愿意,模型也可以在没有它们的情况下工作?

Do you now need to provide all of that context in order for the model to work, or are those like preconditioning that it can also do without if you want?

Bo

我们不再需要融入这些先验知识,因为它们已经是模型内部可学习的参数了。不过,你提到的更像是虚拟细胞的“可提示”或“上下文学习”。我们也可以做到这一点,基本上就是通过向先验知识中添加更多条件,来提示模型朝特定方向预测。

We don't need to incorporate these prior knowledge anymore because these are already learnable parameters inside the models. However, what you suggest is more promptable or in-context learning for virtual cells. We can do that as well, basically by adding more conditions into the prior knowledge so as to prompt the model to predict towards certain directions.

Host

换句话说,模型现在利用了你在训练时提供的先验知识进行学习,所以它不需要它们,但因为你在训练时提供了它们,所以它有一些优势。但如果你能在推理时提供这些先验,你甚至可以获得更多优势。

In other words, the model now takes advantage of the learning using the priors that you provided during training and doesn't need them, but it has some advantage because you provide them during training. But you can even get more advantage if you are able to provide those priors during inference.

Bo

嗯。

Yeah.

Host

好的。哇,不错。嗯,那这有多重要呢?我的意思是,每当我看到大型机器学习论文里塞进一大堆东西时,我总在想,大的阿尔法在哪里,小的阿尔法在哪里?这些东西到底增加了多少?它们只是带来一点点的增量性能提升,还是所有这些对泛化真的至关重要?

Okay. Wow. Nice. Yeah, how much does that matter? I mean, whenever I see big machine learning papers with tons of things thrown in, I'm always wondering where's the big alpha and where's the little alpha? How much are these adding? Is this just a little bit of incremental performance boost, or are all these actually crucial to generalization?

Bo

所以我们必须考虑多个因素。数据贡献了多少?AI 架构贡献了多少?即使是架构,从自回归训练切换到扩散模型的增量是多少?先验知识的增量是多少?当然,所有这些都需要非常具体的消融研究。从经验来看,我们发现数据集的质量和数量最重要。这就是为什么我们对发布 Pisces 数据集感到非常兴奋,它包含 16 种不同的细胞类型和 2500 万个细胞,而且是全基因组范围的。如果你从计算机的角度来看,这是一个巨大的张量:全基因组扰动、全基因组转录组学,加上细胞数量,乘以条件数量。所以这是一个巨大的张量。而且由于 Perturb-seq 技术,我们没有批次效应。所以你不需要模型去克服批次效应。这已经是一个优势了。

So there are multiple factors we have to consider. How much contribution did the data contribute? How much contribution did the AI architectures contribute? Even for the architecture, what's the delta from switching to autoregressive training to diffusion? What's the delta from the prior knowledge? Certainly all of these need very specific ablation studies. From empirical experience, we find that the quality and the amount of the datasets matter the most. This is why we were extremely excited to publish the Pisces dataset, which has 16 different cell types and 25 million cells, and it's genome-wide. You have a huge tensor if you really think about it from a computer perspective: genome-wide perturbation, genome-wide transcriptomics, plus number of cells, times number of conditions. So it's a massive tensor. And because of the Perturb-seq technology, we don't have batch effects. So you don't need the model to climb the hill of batch effect. So that's already an advantage.

数据质量、架构与先验知识 Data Quality, Architecture, and Prior Knowledge

Bo

我们发现,在高质量的扰动数据集上训练已经能给模型带来很大的提升。我们还做了消融实验:如果用相同的数据集训练所有现有的虚拟细胞模型,包括最先进的 scGPT 和原始的 SGV,我们观察到的差异是什么?我们也在论文中报告了这些结果。我们发现,从自回归训练转向扩散语言模型,在一些较难的任务上,特别是对未见任务的泛化能力,带来了显著的提升。而先验知识或多或少是条件特定的。对于某些细胞类型,一些先验知识产生了巨大的差异,但对于某些细胞类型,差异似乎微乎其微。我们正在思考如何更好地整合先验知识。我们仍然相信,让模型了解大量现有的生物学知识应该是有帮助的,但也许是我们通过交叉注意力整合先验知识的方式限制了元数据的范围。但我认为这确实是一个研究课题。但总的来说,如果非要排个序,我的顺序是:数据集的质量(在规模之上),然后是架构,最后是先验知识。但这当然只适用于我们的模型。我相信不同的架构选择会有不同的贡献排名。

So we find that training on high-quality perturbation data sets already gives a big boost to the models. We also did an ablation: if we train all the virtual cell models out there, including state-of-the-art scGPT and the original SGV, on the same data sets, what's the delta we are observing? We report the results there as well. We find that switching from autoregressive training to diffusion language models gives significant improvements on some of the harder tasks, particularly generalization to unseen tasks. And the prior knowledge more or less is condition-specific. For certain cell types, some of the prior knowledge makes a huge difference, but for certain cell types, the delta seems to be marginal. We are thinking about how to better incorporate the prior knowledge. We still believe that letting the model know a big chunk of existing biology should be helpful, but maybe it's the way we incorporate the prior knowledge through cross-attention that limited the scope of the metadata. But I think it's certainly a research topic. But overall, if we have to give an order, my order would be: the quality among scale of the data sets, then the architecture, and then the prior knowledge. But certainly this only applies to our model. I'm sure different choices of architecture have different ranks of contributions.

模型规模与投资 Model Scale and Investment

Host

首先,这真的非常吸引人,很酷的模型。我希望大家都有机会看看这篇论文。显然,投入了大量资源来做这件事。我不知道你们能否透露具体有多少。这是一大笔钱。不管怎样,运营湿实验室,可能是非常复杂的训练运行。我想这是一个大约 40 亿参数的模型,对吗?

First of all, this is really fascinating, very cool model. I hope everyone has a chance to look at the paper. There's obviously a lot of resources that were put into doing this. I don't know if you guys can disclose how much. It's a lot of money. Whatever it was, operating a wet lab, probably very complicated training runs. I think there's a couple four billion parameter model. Is that right?

Bo

49 亿。

4.9 billion.

Host

是的。50 亿参数的模型。所以模型大得多。训练可能用了很多 GPU。与把资金投入到湿实验室工作和传统流程(到目前为止基本上是现状)相比,你从这项努力中获得的提升是什么?

Yeah. Five billion parameter model. So much larger model. Probably took a lot of GPUs to train. What's the lift that you get from this effort versus let's just put the money into wet lab work and the sort of traditional pipeline that basically has been the status quo up until now.

虚拟细胞愿景与泛化 Vision of Virtual Cells and Generalization

Bo

生物学是一门多尺度的学科。在最基本的层面上有 DNA 序列。有细胞,有多细胞的组织片段,共培养,你有组织,有动物系统,最后是人类。我认为我们希望能够朝着光谱的右端进行因果预测。最终在人类身上进行因果预测,知道哪些药物对哪些患者有效,但这很难收集更高级的数据。因此,虚拟细胞的整个愿景是在可能的地方生成数据,以便我们能够将因果预测转移到更转化、更复杂的系统。当然,你已经可以挖掘数据了。我们生成了大量数据,正如 Bo 所说,7 个筛选,16 种不同的生物学背景,基因组规模的扰动。其中已经有很多好的想法。我们在预印本中放了一张图,专门研究了 T 细胞失活。我们已经看到了一些已知的生物学,TCR 复合物。我们还看到了一些推定的新生物学,我们非常兴奋地在实验室验证。其中一些实际上也在去年 12 月 Alex Marson 实验室(也在湾区)发表的一个非常近期的筛选中被捕获。所以看到这个非常兴奋。但希望不仅仅是挖掘现有数据。希望是模型能够泛化,我们将能够在未来进行计算机模拟实验。以前没有人知道需要多少数据、什么样的数据才能做到这一点。整个领域都在等待模型能够击败线性基线在扰动预测中,并且能够超越上下文泛化,不仅仅是在你有训练数据的细胞系内,而是超出那个上下文。这就是为什么你需要一个模型。所以对我们来说非常兴奋的是,在这篇预印本中,我们看到了这种泛化能力的几个演示。我们首先在 T 细胞中做了。我们专门为此目的生成了数据。我们生成了一个静息 T 细胞扰动筛选。所以这些是处于基线状态、未激活的 T 细胞。现在我们有一个激活的 T 细胞扰动。

Biology is a multiscale discipline. There are DNA sequences on the most fundamental level. There are cells, there are multicellular pieces of tissues, co-cultures, you have tissues, you have animal systems and finally you have human. I think we would like to be able to do causal prediction towards the right of the spectrum. Ultimately do causal prediction in human, know what drugs will work in which patients, but that's very difficult to collect higher data on. And so the whole vision of virtual cell is to generate data where it is possible so that we can transfer the causality prediction towards the right, towards the more translational, the more complex systems. Certainly you can mine the data already. We generated a lot of data, as Bo said, seven screens, 16 different biological contexts, genome-scale perturbation. There's a lot of good ideas in that already. There's a figure that we put out in the preprint that just looks into inactivation of T-cells. We already saw some known biology, TCR complex. We also saw some putative new biology which we're very excited to validate in the lab. Some of that were actually also caught out in a very recent screen last December published from Alex Marson's lab, also in the Bay Area. So very excited to see that. But the hope is to not just mine the existing data. The hope is that the model can generalize and we will be able to do in silico experiments into the future. Nobody knows before how much data and what kind of data are needed to do that. The whole field is waiting for the demonstration that the model can beat linear baseline in perturbation prediction and it can generalize out of context, not just within a cell line you have training data on, but out of that context. That's why you need a model. So what's very exciting for us is that in this preprint we saw that generalization capability in a few demonstrations. We first did in T-cells. We actually generated the data expressly for this purpose. We generated a resting T-cell perturbation screen. So these are T-cells in their baseline condition, not activated. And now we have an activated T-cell perturb.

Host

所以 T 细胞激活意味着我要去杀死什么东西。

So just T-cell activation means I'm going, I'm trying to kill something.

Bo

不,这些是调节性 T 细胞。是的,我们激活它们的受体,使它们开始增殖。

No, these are regulatory T-cells. Yes, we activate their receptors so that they're starting to proliferate.

Host

嗯,它们变得更活跃,它们可以做它们的生理工作。

Um, they become more active, they can do their physical job.

Bo

而且我们只,关键的是,我们只在静息 T 细胞上训练模型。

And we only, critically, we only train the model on the resting T-cell.

Host

然后我们告诉模型,嘿,这就是激活 T 细胞的样子,现在去预测所有扰动在这个激活 T 细胞中会做什么。

And we told the model, hey, this is how the active T-cell looks like, now go and predict what all of the perturbations are going to do in this active T-cell.

Bo

而且模型没有见过扰动在激活 T 细胞中如何起作用。我们设置了几个严格的测试。一个线性基线采用了静息情况下的扰动增量。只是将其线性转置到激活 T 细胞上,这就是我们的线性基线。基本上把这看作一个组合扰动预测问题。扰动之一是细胞的激活。另一个是所有基因组扰动。我能不能简单地将两个效应线性相加,那就是线性基线。第二,我们应用了该领域的其他模型。最后但同样重要的是,我们应用了 X-Cell。关键的是,X-Cell 没有见过激活 T 细胞,它不仅能准确预测已知的生物学,TCR 复合物,准确预测它们的效果,即这些会失活 T 细胞,这正是我们预期看到的,而且它还正确预测了我们在筛选中发现的推定 T 细胞失活剂。所以这对我们来说非常兴奋,这表明我们可能能够完全在上下文之外,在未见过的上下文中使用这些虚拟细胞模型,并预测新的生物学。所以我们非常兴奋地跟进这些命中并在实验室验证。简单提一下,我们看到的这个模型的其他几个令人兴奋的泛化案例。记得我们做了一个多细胞类型分化的 iPSC 实验。在那里,我们特意从训练中留出一个细胞类型。所以模型没有见过那个细胞类型,用其他细胞类型和其余数据集训练。模型在那个未见过的细胞类型中,对数千个基因、数千个扰动做出了非常好的预测。这再次表明模型能够跨细胞类型泛化。最后一个实验,我认为我们非常兴奋的是,我们在 T 细胞系上训练,但就在最近,Alex Marson 实验室发表了一个原代 T 细胞筛选。那是一项令人印象深刻的工作。在原代细胞中做这种规模的筛选并不容易。很少有实验室有这种能力。

And the model has not seen how perturbations work in active T-cells. And we set up a couple of rigorous tests. One linear baseline took the perturbational delta in the resting case. Just transpose that linearly onto the active T-cell, and that's our linear baseline. Essentially think about this as a combinatorial perturbation prediction problem. One of the perturbations is activation of the cell. The other is all of the genomic perturbations. Can I just linearly add the two effects together and that would be a linear baseline. And second, we apply other models from the field. And last but not least, we applied X-Cell. Critically, X-Cell has not seen active T-cells, and it's able to make accurate predictions not only on the known biology, the TCR complex, predicting their effect accurately, that these are going to inactivate T-cells, which is exactly what we would expect to see, but also it predicted the putative T-cell inactivators that we found in the screen correctly as well. So that's very exciting to us, and that suggests the possibility that we might be able to use these virtual cell models completely out of context, in an unseen context, and predict new biology. And so we're very excited to follow up on those hits and validate them in the lab. Just very briefly, a couple other cases that we saw exciting generalization capability of this model. Remember we did a multi-cell-type differentiated iPSC experiment. There we specifically held out one cell type from training. So the model has not seen that cell type, trained on the other cell types as well as the rest of the data sets. The model made very good predictions across thousands of genes, thousands of perturbations in that unseen cell type. So again suggesting the model's ability to generalize out of cell type. And the last experiment I think we're very excited about is that we trained on a T-cell cell line, but there was just very recently a primary T-cell screen published from Alex Marson's lab. That's an impressive amount of work. It's not easy to do this scale screening in primary cells. Very few labs have that kind of capabilities.

对原代细胞的泛化 Generalization to primary cells

Bo

在 T 细胞系中做这件事要容易得多。同样,模型能够从细胞系泛化到原代细胞,并在那里做出准确的预测。

Much easier to do that in T-cell lines. Again, the model is able to generalize out of cell lines into primary cells and make accurate predictions there.

Host

所以,他们实际上扰动的是原代细胞,而不是细胞系?

So, they actually perturbed primary cells, not cell lines?

Bo

从供体采集的原代 T 细胞。我们跨越多个供体,而仅在一种 T 细胞系上训练的 X-Cell 能够对来自原代 T 细胞实验的多个供体做出预测。

Primary T-cells harvested from donors. We were across multiple donors, and X-Cell trained on just one T-cell line is able to make predictions across multiple donors from primary T-cell experiments.

Host

这是对整个理论的验证,对吧?你可以在这些有点奇怪的细胞上训练,而且效果会很好,因为你很好地覆盖了领域,或者不管怎样,你能够真正预测来自真实人体的真实细胞。

This is a validation of the whole theory, right? That you can train on these slightly weird cells and it will be good because you're covering the domain well enough, or whatever it is, that you're able to actually predict in real cells that come directly from real people.

Bo

没错。是的,我认为构建虚拟细胞并不是要取代生物实验,正如你提到的。我们真正想做的——虚拟细胞的圣杯——是让模型能够泛化到那些更难甚至不可能进行生物实验的未见情境。对吧?到目前为止,X-Cell 专注于细胞系,最终我们希望扩展到更复杂的生物系统,比如动物、类器官,以及我们之前提到的,最终到患者,到人类生物学。请记住几个数字:90% 的疾病没有治愈方法,大多数药物在患者身上的三期临床试验中失败。三期试验的成功率低至 5% 到 10%。

That's right. Yeah, I think building a virtual cell is not to replace biological experiments, as you mentioned. What we're trying to do—the holy grail of virtual cell—is to have a model that generalizes to unseen contexts that are harder or even impossible to conduct biological experiments on. Right? So far, X-Cell is focusing on cell lines, and eventually we want to extend to more complicated biological systems such as animals, organoids, and eventually, as we mentioned before, to patients, to human biology. And bear in mind a few numbers: 90% of diseases have no cure, and most drugs fail at phase three clinical trials on patients. The success rate of phase three trials is as low as 5 to 10%.

Host

三期是指?

And phase three means?

Bo

患者试验的最后阶段。

The final stage on the patient trials.

Host

所以那是从二期的毒性泛化到三期的疗效,对吧?从小群体到大得多的群体。

So that's when you generalize from toxicity in phase two to efficacy in phase three, right? From a small cohort into a much larger cohort.

Bo

哦,抱歉。毒性是一方面,对吧?小群体和大群体。

Oh, sorry. Toxicity is one, right? Small cohort and so a large cohort.

Host

所以泛化问题:好吧,这种药——我非常仔细地挑选了患者,效果很好,但现在我有了更多患者,突然效果就不太好了。这就是你要解决的大问题。

So the generalization problem: okay, this drug—I've very carefully selected my patients and it works pretty well, and now I get a bunch more patients and suddenly it doesn't work very well. And that's the big problem that you're addressing.

Bo

虚拟细胞的承诺当然是:我们能否构建这样一个模型,学习所有因果生物学,从而能够基于此最终预测患者的反应,这样对于某些药物,我们可以选择合适的患者进行临床试验?对吧?所以这是一个长期愿景,但我们已经看到一些早期希望,即 X-Cell 在多样化的因果数据集上训练后,已经能够泛化到一些未见过的细胞类型。所以当然还有很多实验要做来验证这个模型,甚至不断微调这个模型,但我认为我们确实看到了一些早期希望。

Certainly the promise of virtual cell is: can we build such a model that learns all the causal biology so that it can be grounded to predict the response eventually on patients, so that for certain drugs we can select the right patients to conduct the clinical trials on? Right? So this is a long-term vision, but we already see some early hopes that X-Cell, trained on diverse sets of causal datasets, can already generalize to some unseen cell types. So certainly there's a lot of experiments to be done to validate this model, and even continuously fine-tune this model, but I think we certainly see some early hopes.

基础模型vs线性基线 Foundation models vs linear baselines

Host

你刚才谈到线性模型,这引出了那个著名或臭名昭著的 ARC 挑战,关于扰动。一直有这样一个主题:复杂的基础模型往往无法击败线性基线。我想听听你的看法。首先,这有什么不同吗?我的意思是,我认为你们自己的一些模型过去可能也难以击败线性基线。你们当前的数据策略或未来的方向有什么不同吗?领域的发展方向是什么?基础模型、虚拟细胞模型与这些简单基线相比,各自扮演什么角色?

So you were talking about linear models, and this brings up this famous or infamous ARC challenge about perturbation. And there's been this theme about complicated foundation models oftentimes not beating linear baselines. I'd like to get your take about that. First of all, is this different? I mean, I think some of your own models might also have had trouble beating linear baselines in the past. Is there something different about your current data strategy or where you're going? And where's the field going? And what is the role of foundation models versus virtual cell models versus these simple baselines?

Bo

是的,有几点。首先,那些基准测试,正如你提到的,是在非常小的扰动数据集上进行的,人们报告的指标大多是 MAE。当然你可以想象,因为单细胞数据集非常稀疏,细胞的平均谱——你可以想象它是最小化 MAE 的一个很好的最小值,一种局部最优。这就是为什么有时细胞的平均谱的 MAE 甚至低于技术重复,而技术重复被认为是扰动实验的黄金标准。所以这本身就表明那个指标不可靠。然而,大多数这些基准测试仍然在比较那些在静态表达数据集上训练的基础模型,比如 scGPT 或 Geneformer。在内部,我们也发现,当涉及 MAE 时,有时 scGPT 无法超越线性模型,原因就是我刚才说的。而 X-Cell 与这些静态表达模型(如 scGPT 或 Geneformer)的不同之处在于,我们实际上不是训练基因表达数据集,而是训练因果数据集。我们训练大量的全基因组扰动数据集,以便更好地学习干预的动态。在我们的预印本中,我们也广泛地与线性模型进行了比较。正如 Chu 提到的,线性模型完全无法扩展到未见过的细胞类型。你可以很快想象为什么。我相信,在更困难的任务中,特别是在泛化任务中,那些在正确数据上训练的基础模型或其他更复杂的 AI 模型将超越这些线性模型。这就是为什么我一直提到,正确的数据集加上正确的 AI 模型将带来巨大的改进。但我认为该领域仍然需要更多的生物学验证,才能更确信虚拟细胞方向是正确的。

Yeah, a few things. First of all, those benchmarks, as you mentioned, are conducted on perturbation datasets which are very small datasets, and the metrics people report are mostly MAE. Certainly you can imagine, because single-cell datasets are so sparse, the average profiles of the cells—you can imagine it's a great minimum, kind of local optimum, to minimize the MAE. This is why sometimes the average profile of cells has lower MAE even than technical replicates, which are considered ground truth for perturbation experiments. So that itself shows that that metric is not reliable. However, most of these benchmarks are still comparing foundation models that train on static expression datasets, such as scGPT or Geneformer. Internally, we also find that when it comes to MAE, sometimes scGPT fails to outperform linear models just because of the reasons I just stated. And what sets X-Cell apart from these static expression models such as scGPT or Geneformer is that we actually, instead of training on gene expression datasets, we train on causal datasets. We train on massive amounts of genome-wide perturbation datasets so that it learns better about the dynamics of the interventions. And in our preprint, we extensively compare with linear models as well. And as Chu mentioned, linear models totally failed to extend to unseen cell types. You can quickly imagine why. And I believe that foundation models or other more complicated AI models that train on the right data will outperform these linear models in harder tasks, particularly in generalization tasks. And that's why I keep mentioning that the right dataset with the right AI model will lead to huge improvements. But I think the field still needs to see more biological validations to be more convinced that the virtual cell direction is the right one.

Host

是的,我认为该领域缺乏一致且统一接受的基准。被衡量的东西才会被改进。在我们的论文中,我们衡量了——我认为我们投入了很多思考并看到模型真正大放异彩的指标之一是围绕基因表达变化的指标。所以,你知道,Pearson delta——预测变化与扰动后真实变化之间的相似性。这很难作弊;你必须真正把变化弄对。而真正让我震惊的是,当我看到模型做出预测,打印出基因表达变化的热图,查看实际原始数据,并将线性基线预测、真实值和 X-Cell 预测排在一起时,视觉上非常清楚地看到 X-Cell 的预测比线性基线更接近真实值。这就是我一开始提到的“哇”时刻。

Yeah, I think the field suffers from a lack of consistent and uniformly accepted benchmarks. What gets measured will get improved. And in our paper, we measured—I think one of the metrics that we put a lot of thought into and saw the model really shine is metrics around gene expression changes. So, you know, Pearson delta—the similarity between predicted changes and ground truth changes upon perturbation. That's very hard to cheat on; you have to really get the changes right. And what really blew my mind away is when I saw the model make predictions, just print out the heatmap of the gene expression changes, look at the actual raw data, and line up the linear baseline prediction, the ground truth, and X-Cell prediction all together. It's visually very clear to see that X-Cell prediction is very much more similar to ground truth than the linear baseline. This is the wow moment I was talking about in the beginning.

Bo

没错。而且不难理解为什么。当我们把——这是第一次有人能把不仅仅是单个扰动,而是七个全基因组扰动活动放在一起。让我们这些生物学家立刻注意到的是,有些扰动是情境通用的,意味着无论你在什么细胞类型中实验,基因都做同样的事情。你可能不会感到惊讶,这些是你的管家基因,对吧?当然,它们在每个细胞里都做同样的事情。然后还有其他所有的基因簇,它们具有非常情境特异的功能。它们在不同细胞中做不同的事情。同样,不难想象为什么。

That's right. And it's not hard to understand why. When we put—this is the first time that someone can put together not just one perturbation but seven genome-wide perturbation campaigns together. Something that jumped out to us biologists right away is that some of the perturbations are context universal, meaning that the genes do the same thing regardless of the cell types you experiment in. Might not be surprising to you that these are your housekeeping genes, right? Of course, they do the same thing in every cell. And then there are all of these other clusters of genes that have very context-specific functions. They do different things in different cells. Again, not hard to imagine why.

单一vs组合扰动 Single vs combinatorial perturbations

Host

你们的扰动总是单基因扰动,还是会有更多?因为我对调控网络的理解是,有时候一个基因就能做很多事情。比如,我认为男性之所以分化,只是因为胚胎发育第七天左右某个基因被激活,然后这个基因就决定了一切。但有时候你又会有大片的基因网络,它们都非常冗余,从而允许更微妙的反馈机制等等。所以我可以想象很多单基因扰动可能没什么意义。你可能会想开始采用更具组合性的策略。

Are your perturbations always single gene perturbations or do you have more? Because my understanding of regulatory networks is often times sometimes it can be a single gene does a ton of things. For example, I think males are just differentiated due to one gene being enabled at like day seven of embryo development or something and that differentiates everything is this one gene. But then sometimes you have large networks of genes which all are very redundant which allows for more subtle feedback mechanism and so on. So I could imagine a lot of single gene perturbations as being kind of irrelevant. And that you might want to start having a more combinatorial strategy here.

Bo

是的,这是个很好的问题,Bo 和我对此思考了很多。实际上,你提到生殖生物学很有意思。我研究了很多雌性细胞中的补偿机制,即剂量补偿机制。雌性细胞有两条 X 染色体,雄性细胞只有一条 X 染色体,为了匹配 X 染色体的剂量输出,哺乳动物细胞采用的策略是:一个基因产生一种不编码任何蛋白质的 RNA,只是非编码 RNA。这个 RNA 包裹住雌性的一条 X 染色体,并下调该染色体上大部分基因的表达,把它推到细胞核的一个角落里,这被称为巴氏小体,之后它就不再被使用了。所以我完全同意你的观点,一个基因可以发挥很大作用,但在生物学中,你还有冗余、补偿以及各种机制,敲除一个基因并不总能观察到表型。如果四个基因冗余执行同一功能呢?对吧?只去掉一个是不够的。所以我们最初是从单一细胞类型、功能缺失、单基因扰动,并且只以 RNA 表达作为输出来开始的,现在我们正沿着所有这些维度扩展平台。这就是我们今天为训练像 X-Cell 这样的模型构建数据脚手架的方式。我们现在开始在这三个维度上扩展平台:超越单纯的转录组学,看向多模态数据;超越单一基因扰动,也看向通路激活和失活,开启或关闭整个基因级联反应;并且超越细胞系和单一培养,进入更复杂的、与转化医学相关的系统,比如原代细胞、类器官,甚至直接进行体内扰动筛选。我们相信,有了所有这些扩展,用于训练模型的数据将更加令人兴奋。这也是为什么我们将 PPI 网络作为先验知识整合到模型中。虽然目前的模型是在单基因扰动上训练的,但一旦模型训练完成,你实际上可以在模型上仅通过计算机模拟来预测组合扰动,对吧?也就是说,你可以同时扰动两个基因的 token,看看响应是什么。当然,如果没有在真实的组合扰动数据集上训练,准确性可能不够,但至少有了这样的模型,我们可以开始利用计算机模拟扰动来生成假设。

Yeah, that's a great question and Bo and I have thought about this a lot. Actually, it's interesting that you brought up reproductive biology. I study a lot in female cells the compensation, the dosage compensation mechanisms there. Female cells have two X chromosomes. Male cells have one X chromosome to match the dosage output from the X chromosomes. The strategy that the mammalian cells employ is one gene that produces an RNA that does not encode for any protein, just a non-coding RNA. That RNA wraps around one of the female X chromosomes and turns down most of the gene expression from that chromosome, shoves it away in a corner of the nucleus, and it's called a Barr body. It is never heard from again. So absolutely agree with you, one gene can do a lot, but in biology you also have redundancy, you have compensation, you have all kinds of mechanisms where knocking down one gene is not sufficient to always see a phenotype. What if four genes redundantly do the same thing, right? Taking out one is not going to be sufficient. So where we started with one cell type at a time, loss of function, single gene perturbation and look at only RNA expression as the output, we're expanding the platform along all of those axes. So that's what we do today to build a scaffold of the data for training models like X-Cell. We are now beginning to grow in all three axes of the platform. Going beyond transcriptomics alone to look at multimodal data. Going beyond just one gene perturbation alone to look at also pathway activation and inactivations, turning on and off entire cascade of gene chain reactions. And also going beyond just cell lines, monocultures, into more and more complex translationally relevant systems, into primary cells, into organoids, and doing even direct in vivo perturbation screens. So we believe with all of that expansion, the data will be all the more exciting to train models on. This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now is trained on single gene perturbations, once the model is trained, you can actually predict combinatorial perturbations just on the model in silico, right? So in the sense that you can just perturb the tokens of two genes at the same time and see what's the response. Certainly without training on the actual combinatorial perturbation dataset, the accuracy may not be there, but at least with the existence of such models, we can start to generate hypotheses using in silico perturbations.

AI时代科学家的角色 Role of scientists in the AI age

Host

你如何看待科学家在 AI 时代角色的变化?而且你可能,因为你不是在使用,你的重点不是语言模型本身、智能体式科学之类的东西,那么你的观点可能会略有不同。你在构建这些非常具体的模型,但还有一点,你和你的学生如何能保持如此高的节奏?我怀疑这可能部分与生成式 AI 有关,但还有,你如何看待科学家,特别是学术界科学家的角色变化?

How do you see the role of the scientists changing in the age of AI? And you may, because you're not using, your focus is not language models themselves and agentic science and things like that, then you may have a different slightly different take. You're building these very specific models, but still, one, how are you and your students able to maintain such a high pace? And I suspect it may have something to do with generative AI partly, but also, and how do you see the role of the scientist, the academic, changing?

Bo

是的,这是个很好的问题。我官方的时间分配是 80% 在 Xaira,20% 在大学任职,但现实是我 100% 在 Xaira,100% 在这件事上。

Yeah, that's a great question. So my official split of time is 80% on Xaira, 20% on my university affiliations, but turns out the reality is 100% on Xaira, 100% on this.

Host

你发明了时间机器。这就是答案。

You invented a time machine. That's the answer.

Bo

所以我保持跟进的方式是,我的实验室使用大量智能体式 AI 来每天监控所有 AI 论文。每周我们开组会,讨论 AI 用于生物学、AI 用于医疗保健等不同主题。而且说实话,即使作为教授,我也觉得跟上进度极其困难。AI 的节奏快得令人难以置信,以至于有时我醒来会感到焦虑,比如“天哪,这篇论文已经这么多人发表了,我们未发表的工作怎么办?”你可以想象,学生们可能面临 10 倍的焦虑。所以有时我试图鼓励学生真正使用不同的工具,努力保持专注,找到一个我们成为专家的细分领域,对吧?但在生成式 AI、智能体式 AI 的时代,我发现人们,至少是学术界人士,做科学的方式现在非常不同。总的来说,大多数教授或学生现在在资金、发表速度方面都非常挣扎。这就是为什么你会看到很多重大突破来自工业界,对吧,比如 AlphaFold。所以学术界如何在这样的智能体式 AI 时代生存甚至繁荣,肯定是每个人都在思考的问题。我们看到很多教师离开大学加入工业界,仅仅是因为资源的原因,对吧?如果你在做 AI 研究,你的学校有足够的 GPU 吗?当学生加入教授实验室时,他们经常问的第一个问题是你有多少 GPU,对吧?所以当然,在这方面,工业界比学术实验室有巨大优势。但我认为学术实验室的优势在于创新速度,以及特定学术实验室可以极其专业的细分领域。此外,拥有思考的自由有时也让你更容易在工业界人士可能没想到的想法上创新。总的来说,我认为整个领域需要更多创新才能跟上,希望借助不同工具,也希望政府开始更多地投资学术界,因为我仍然深信学术界是整个领域创新的主要来源,尤其是在生物技术方面。所以希望我们能看到更多对学术界的投资,这样我们才能维持下去。

So the way I try to keep up is, I mean, my lab uses lots of agentic AI to try to monitor all the AI papers every day. And every week we have lab meetings, we try to discuss different topics in AI for biology, AI for healthcare, etc. And it's, to be very frank, even as a professor, I find it's extremely hard to catch up. The pace of AI is just so incredibly fast, and to the point that sometimes I feel anxiety waking up, like, oh my god, this paper already, so many people published, and what happened to our existing unpublished work? And suddenly you can imagine students probably face 10x anxiety. So sometimes I try to encourage students to really use different tools, try to stay focused, find a niche area that we become experts on, right? But with the era of generative AI, agentic AI, I find the way people, at least academics, do science is extremely different now. Overall, most of professors or students in academics start to be very struggling in terms of fundings, in terms of the pace of publications. And that's why you can see lots of major breakthroughs come from industry, right, like AlphaFold, for example. So how academics survive or even thrive in such an era of agentic AI is certainly something everybody is thinking about. We see lots of faculty leave university and join industry simply for the reasons of resources, right? If you are doing research on AI, do you have enough GPUs at your school? The first question you should ask when a student joins a professor's lab, the first question they often ask is how many GPUs do you have, right? So certainly in that sense, industry has a major advantage over academic labs. But I think what academic labs have advantages on is really the kind of the pace of innovations and also the niche areas this specific academic lab can be extremely expert on. So also having the freedom of thinking sometimes also makes you easier to innovate on ideas that maybe industry people didn't even think about. Overall, I think the whole field needs to be a lot more innovative to catch up, and hopefully with the help of different tools, and hopefully the government starts to invest more into academia, because I still deeply believe that academia is the main source of innovation for the whole field, and particularly when it comes to biotech. So hopefully we see more investment into academia so that we stay afloat.

政府为何应资助学术界 Why Government Should Fund Academia

Host

我同意你的观点,但为什么你会这么想?为什么资金不应该直接流向产业界?政府为什么要把钱投入学术界?

I agree with you, but why do you think that? Why shouldn't the money just go to industry? Why should the government put any money into academia?

Bo

我仍然相信学术自由的力量。这实际上是学术教授存在的初衷,他们不仅教学,还做研究。同时教学和研究也有好处,因为当你教授一门学科时,你实际上必须成为该领域的专家,这迫使你不断更新知识库,并找到简单的方式将知识传授给学生。通过这样做,你实际上开始在不同想法上创新。例如,我在多伦多大学教授一门关于深度学习和神经网络的大课,每年有 600 名学生。通过教学,我被迫每年更新幻灯片和讲义,我阅读大量材料试图更新自己,以便找到方法将所有知识传授给学生。这就是我不断更新自己的方式。每当有转变——我清楚地记得当我们有 GPT-2 时,我们更新了讲义,然后是不同的大语言模型,多线程 GPU 通信如何用于训练更大规模的神经网络等等。所以我强迫自己更新模型,通过与不同学生交流,我们确实产生了许多新颖的想法,将前沿 AI 模型应用于生物学或医疗保健中非常特定的细分领域。也许这是加拿大学术体系独有的。作为学术界教授,我们还能接触到许多医疗数据集,这些数据由于法律或监管原因,产业界很难获取。这就是为什么你会看到我们通过加拿大学术医院发表的一些论文,我们开发了一些最先进的超声图像基础模型。所以我的意思是,一定程度的学术思维自由确实推动了许多创新。我仍然相信——也许这有偏见——但我仍然相信,拥有一定程度的学术思维自由将带来许多在产业界无法实现的创新。

I still believe in the power of academic freedom. This is actually the original motivation for the existence of academic professors, who not only teach but also do research. And there are benefits of teaching and research at the same time, in the sense that when you teach a subject, you actually have to become an expert on it, and that forces you to keep updating your knowledge base and find easy ways to convey your knowledge to students. By doing that, you actually start to innovate on different ideas. For example, I teach a big class at the University of Toronto about deep learning and neural networks. It's a gigantic class with 600 students every year. By teaching that, it forces me to update the slides and lectures every year, and I read lots of materials trying to update myself so that I can find ways to convey all the knowledge to students. That's how I keep updating myself. Whenever there's a transition—I remember vividly when we had GPT-2, we updated the lecture, then different language models, how multi-thread GPU communication is used in training larger-scale neural networks, etc. So I force myself to update the models, and by talking to different students, we really generate lots of novel ideas to apply cutting-edge AI models to very specific niche areas in biology or healthcare. Maybe it's unique to the Canadian academic system. By being a professor in academia, we also have access to lots of healthcare datasets that are very hard for industry to access due to many legal or regulatory reasons. That's why you see some of the papers we publish through academic hospitals in Canada, where we developed some of the state-of-the-art foundation models for ultrasound images. So that's what I mean: there's a certain level of freedom of academic thinking that really drives lots of innovations. And I still believe—maybe it's biased—but I still believe that having a certain level of freedom of academic thinking will lead to lots of innovation that is unsinkable in industries.

Host

我 100% 同意 Bo 的观点。想想我们做的实验室工作流程,很多都是建立在最初由学术界开创的创新之上。CRISPR 当然是在学术界发现的。CRISPR 应用于高通量筛选也首先在学术界得到验证。单细胞 RNA 测序,这种液滴包裹的单细胞 RNA 测序,首先在学术界得到验证,然后不同的公司试图将其商业化。将这些结合起来做 Perturb-seq 也首先在学术界开创,对吧?Chris Box 实验室、Jonathan Weissman 实验室、Aviv Regev 实验室——许多先驱都在学术界。然后我认为,尤其是在实验室方面,这种创新需要很长时间,发现过程可能非常偶然,因此可能不适合纯产业界承担。但一旦它们显示出早期前景,进行 Scaling(规模扩张)、增强稳健性,并生成不仅海量而且高质量的数据,尤其是针对 AI 规模的数据,我认为这可以在产业界做得很好,无论是思维方式还是我们能够支持的资源,有时学术实验室难以匹敌。

I 100% agree with Bo there. Just thinking about the lab workflow that we do, a lot of these are building upon innovations that were first pioneered in academia as well. CRISPR, of course, was discovered in academia. CRISPR applied to high-throughput screening was also demonstrated in academia first. Single-cell RNA-seq, this kind of droplet-encapsulated single-cell RNA-seq, was first demonstrated in academia, then different companies tried to build it up into commercial offerings. Putting all of these together to do Perturb-seq was also first pioneered in academia, right? Chris Box lab, Jonathan Weissman's lab, Aviv Regev's lab—many of the pioneers in academia. And then I think, especially on the lab side, this innovation takes so long and the discovery process can be so accidental that it's perhaps not ideal for pure industry to take on. But once they show early promise, scaling them and robustifying them and generating data that's not only massive but high quality, especially for AI scale, I think that's something that can be very well done in industry, both the mindset as well as the kind of resources that we can support, which sometimes can be hard for academic labs to match.

开放科学与数据共享 Open Science and Data Sharing

Host

Xaira 在发布数据集和模型方面非常慷慨。鉴于你刚才谈到的一些事情,你似乎非常致力于开放科学。而且,关于虚拟细胞或理解人类生物学的最佳、最重要的数据策略的讨论。你认为学术界下一步应该怎么走?既然 Xaira 现在可能拥有相当于几十个或几百个生物实验室的预算,如果你是一位学者、教授,尤其是在湿实验室,你会想专注于什么?

Xaira has been very generous with releasing your datasets and your models. You seem to be very committed to open science, given some of the things you were just talking about. And you know, the discussion about what is the best most important data strategy for virtual cells or understanding human biology. Where do you think academia should go next? And since Xaira has probably a budget comparable to probably dozens or hundreds of biolabs right now, what do you think if you are an academic, a professor, especially in a wet lab, what would you want to be focusing on?

Bo

首先,我是开放科学的坚定信徒。这就是为什么我们在这里讨论的所有模型和数据都是开源的。你可以从我实验室的 GitHub 上找到所有的数据权重。

First of all, I am a deep believer in open science. That's why all the models we talk about here and the data are open-sourced. You can find all the data weights from my lab's GitHub.

Host

非常全面。

It's very thorough.

Bo

是的,谢谢。我相信开放科学的原因是,正如 Chentian 提到的,大多数时候学术实验室从一个想法开始并制作原型。它不可扩展,甚至不是一个好产品,而产业界可以将其规模化。这也是为什么 scGPT 迅速成为前公司中最广泛使用的单细胞基础模型之一。这对我们来说非常鼓舞人心,这也是 Xaira 开始开源一些数据集和模型的部分原因。部分原因是我们认为虚拟细胞领域还处于早期阶段,保留某些数据集或模型没有帮助,因为太早了。一个更好的双赢局面是,这个领域的每个人开始共同贡献数据,共同贡献模型,交流想法,这样这个领域就能以更快的速度前进。我们在蛋白质领域看到了成功的例子,对吧?由于 PDB 等开源数据的可用性,我们有 AlphaFold、RoseTTAFold 和 AlphaFold2 等模型,它们也是开源的,所以我们可以快速迭代不同的模型。这就是为什么你看到蛋白质领域蓬勃发展的原因。我们想为虚拟细胞做同样的事情:让我们把所有数据集放在一起,让我们有相同的标准协议来生成高质量的数据集,让我们把所有资源放在一起生成下一代虚拟细胞模型。至于学术实验室,Chentian 可以评论湿实验室学者如何生存。但从干实验室的角度来看,我确实鼓励大学里所有的干实验室 AI 研究人员开始与产业界合作,这样他们就能获得更多资源来发展自己的想法。随着智能体式 AI 时代的到来,现在每个人都能编程。所以更重要的是对你的项目有正确的品味,这样你就不会漫无目的地消耗 token,对吧?所以我们希望更多的学术学生和教授有更高的研究品味,这样我们就能正确利用 AI 智能体。

Yeah, thank you. And the reason I believe in open science is that, as Chentian mentioned, most of the time academic labs start with an idea and prototype it. It's not scalable, it's not even a good product, and industry can take it to scale things up. This is also why scGPT quickly became one of the most widely used single-cell foundation models in former companies. So that's very encouraging to us, and this is why Xaira is also starting to open-source some of the datasets and some of the models. Part of the reason is that we believe virtual cells are such an early field, and it doesn't help to withhold certain datasets or models because it's so early. A better win-win situation is everybody in this field starts to contribute data together, contribute models together, exchange ideas, so that this field can move forward at a much faster pace. We see successful examples in protein space, right? Because of the availability of open-source data such as PDB, we have models such as AlphaFold, RoseTTAFold, and AlphaFold2, which are also open-source, so we can quickly iterate different models. That's why you see a booming situation in protein space. We want to do the same thing for virtual cells: let's put all the datasets together, let's have the same standard protocols to generate high-quality datasets, let's put all the resources together to generate the next generation of virtual cell models. When it comes to academic labs, Chentian can comment on how wet lab academics can survive. But from a dry lab perspective, I do encourage all the dry lab AI researchers in universities to start collaborating with industries so that they can get more resources to develop their own ideas. And with the era of agentic AI, now everybody can code. So it's more important to have the right taste about your projects so that you don't just burn tokens without purpose, right? So we want more academic students and professors to have higher taste in research so that we make the right utility of AI agents.

AI时代培养品味 Developing Taste in the Age of AI

Host

作为教授,显然你的工作就是要有品味。但作为学生,在一个如此多的思考和科学过程基本上可以外包给大语言模型的世界里,你如何培养品味?这不是一个你被迫绞尽脑汁、艰难学习品味的世界。

As a professor, it's your job to have taste, obviously. But as a student, how do you develop taste in a world where so much of the thinking and scientific process could be essentially outsourced to an LLM? There's not a world where you're forced to bang your head against something and learn taste the hard way.

智能体AI时代的训练 Training in the Era of Agentic AI

Bo

这就是为什么我们需要学术训练,让你进入一个你完全不了解的领域,希望毕业后你能成为这个特定领域的专家。这就是为什么你必须经历不同的项目,与同行交流,与教授交流,以了解什么是好的研究品味。但更重要的是,我总是教学生,学习某事的最好方法就是直接去编程。通过编程,你能发现论文中数学方程里隐藏的细节,而这些细节常常被省略。但在智能体式 AI 时代,情况略有不同:过去我们花大量时间编码,少量时间调试,但现在我们让智能体做大部分编码,而我们花大部分时间调试,这对我来说非常有趣。我们和实验室的学生讨论了很多关于如何发现 AI 产生的错误。那么,如何找到 AI 特别擅长的领域,以及如何找到 AI 仍然受限的地方?这同样需要大量的领域专业知识。最终,你仍然需要用真实世界的证据来验证你的模型,对吧?所以,与湿实验室合作,与临床团队合作,验证你的模型,为你的品味提供反馈信号,正如我们讨论的,这无疑是一个非常重要的训练项目。

This is why we need academic training where you get into a field you know nothing about, and hopefully after you graduate, you become the expert about this particular topic in the world. This is why you have to go through different programs, talk to your peers, talk to your professors, to get an idea about what's good research taste to begin with. But more importantly, I always teach students that the best way to learn something is to just program it. By programming, you kind of know what's the details hidden in all the mathematical equations in the paper, which often you omit. But in the era of agentic AI, it's slightly different in the sense that we used to spend lots of time coding and a little bit of time debugging, but now we let the agent do most of the coding, but we spend most of the time debugging, which seems to be infinitely interesting to me. And we had lots of discussions with the students in the lab about what's the best way to spot bugs by AIs. So how do you find places where AI is particularly good at, and also how do you find places AI are still limited at? You need lots of domain expertise as well. And in the end, you still need to validate your model using real-world evidence, right? So that's why collaboration with wet labs, collaboration with clinical teams to validate your model, provide feedback signals to your taste, as we discussed, is certainly a very important training program.

学术界与产业界的共生创新 Symbiotic Innovation in Academia and Industry

Bo

在湿实验室方面,学术界扮演着极其重要的角色。我认为我们将进入一个共生创新和思想交叉融合的时代,就像 AI 领域一样。我认为学术界仍然不断涌现出伟大的想法,但工业界现在也在架构等方面贡献新的想法。在湿实验室方面,工业界似乎能够很好地扩展这类数据生成。但生物学远不止基于细胞的扰动,也不止 RNA 测序。我们想测量许多其他分析物,对吧?蛋白质、代谢物、脂质、蛋白质-蛋白质相互作用。我们如何在单个细胞之外大规模地做到这一点?我们希望能够在原生环境中测量细胞相互作用、空间分布,甚至整个动物水平的体内扰动。同样,我们如何在学术界的大量创新中实现规模化?实际上,上周就有一篇伟大的论文。那么,我们如何将所有这些连接起来?我认为我们还有很多年的工作要做,才能完全破解所有生物学的数据生成,我们需要工业界的规模、工业化和创新,我们也需要学术界的这些。我认为我们将共同把这个领域推向新的高度。

On the wet lab side, academia has extremely important roles to play. I think we'll enter a field of an era of symbiotic innovation and cross-pollination of ideas, just like in AI field. I think we have great ideas coming out of academia still all the time, but industry now increasingly are contributing new ideas on architecture on all of that as well. In the wet lab side, certainly industry seems to be able to scale this type of data generation quite well. But biology is so much more than just cell-based perturbations, beyond RNA-seq. We would like to measure many other analytes, right? Proteins, metabolites, lipids, protein-protein interactions. How do we do that at scale beyond individual cells? We would like to be able to measure cell interaction, spatial, cell in their native context, or even whole animal level in vivo perturbations. Again, how do we do that at scale with lots of great innovation coming out of academia? Actually, just one great paper last week. And so, how do we connect all of these together? I think we have many years of work ahead of us to fully crack data generation for all biology, and I think we need the scale, the industrialization, the innovation from industry, and we also need that from academia. I think we will together move this field to the next level.

魔法棒:蛋白质测序与时间动态 Magic Wand: Protein Sequencing and Temporal Dynamics

Host

我们一直想问每个人一个问题:在你的领域,你可以说是 AI 和高通量实验,或者随便你怎么定义,如果你能挥动魔杖,消除一个瓶颈或解决一个关键问题,那会是什么?

One question that we've been trying to ask everyone is: in your field, which you could say maybe is AI and high-throughput experimentation, or however you want to define it, if you could wave your magic wand and have a bottleneck removed for you or a key problem solved, what would that be?

Bo

蛋白质。

Protein.

Host

好的。

Okay.

Bo

如果有办法进行蛋白质测序或高通量蛋白质测量,达到与基因组学相同的规模,那将是惊人的。我接受的是基因组学训练,但如果我能做到这一点,我会立刻采用这项技术。RNA 很神奇,它预示着哪些蛋白质将被制造出来,但蛋白质基本上是细胞中的功能单位。不仅它们的丰度重要,它们的翻译后修饰、它们在细胞中的定位也很重要。如果我们能大规模地测量所有这些,包括它们的构象状态、修饰、丰度和定位,无论是单细胞还是空间层面,我认为这样的数据集将对训练下一代前沿模型极其有用。我知道在这个方向有很多创新,我迫不及待想看到它成熟。

If there is a way to do protein sequencing or high-throughput protein measurement, the same scale that we can do genomics, that will be amazing. I'm training in genomics field, but if I can do that, I would incorporate that technology in a heartbeat. RNA is amazing. It foreshadows which proteins are going to get made, but proteins by and large are the functional units in a cell. Not only does their abundance matter, their post-translational modification matter, their localization in a cell matter. If we can measure all of those things, their conformational states, their modifications, their abundances, their localizations at scale, single cell or even spatially, I think such datasets will be incredibly useful to train the next generation of frontier models. I know there's a lot of innovations in that direction. Can't wait to see that become of age.

Bo

我的希望是看到测序技术的突破。不仅仅是降低成本,而是能够在不同时间点对同一细胞进行测序的技术。我认为目前这方面非常缺乏,因为要对细胞进行测序,你必须杀死细胞,对吧?那么,我们能否有一种技术,可以在不同时间点测量同一组细胞的细胞状态?我认为这将为数据集带来一个非常不同的维度,使我们能够开始测量细胞的时间动态。到目前为止,我们测量和建模的一切都是极其静态的。那么,我们能否有一种技术,可以在不同时间点测量同一组细胞的不同反应?这将为建模细胞动态释放巨大的机会。对我来说,这才是真正的虚拟细胞。

My hope is I hope to see a breakthrough in sequencing technology. Not just the reduced cost, but sequencing technology that can sequence the same cells at different time points. I think this is much lacking right now because in order to sequence the cell, you have to kill the cell, right? So can we have a technology that can measure the cell states at different time points for the same set of cells? I think that will bring a very different dimension to the dataset so that we can start to measure the temporal dynamics of cells. So far everything we measure, everything we model is extremely static. So can we have a technology that measures different cell responses at different time points for the same set of cells? That will unlock massive opportunity to model the dynamics of cells. To me, that is the real virtual cell.

Host

这真是一个有趣的想法。我从未想过这一点。哇。你能接受哪怕是部分的,比如基因的小片段,或者少量转录本的 3' 端区域吗?

That's a really interesting idea. I never would have thought about that. Wow. Would you be okay with even just partial like small snippets of genes or maybe three prime regions of a small number of transcripts?

Bo

从一小部分基因组合开始,对吧?但最终,如果我们说的是那个神奇的目标,最终如果我们能有一个系统,在不同时间点观察细胞如何演化,并且我们有足够的数据来实际建模这种发展,我认为那将是真正的虚拟细胞建模。

Start with a small set of gene panels to begin with, right? But eventually, if we're talking about the magic one here, eventually if we can have a system that observes how cells evolve at different time points and we have enough data to actually model such development, I think that would be real virtual cell modeling.

Bo

有一些小的尝试,哦,抱歉,早期的尝试,比如从细胞中吸出一部分,几乎像从细胞中取样活检,进行一部分细胞质测量,这可能与你提到的想法类似。还有一个公共实验室的工作,让细胞分泌小囊泡,他们在细胞培养基中收集这些囊泡,以纵向测量细胞产生的东西。但还没有技术能让你在保持细胞存活的同时测量整个细胞转录组。鱼与熊掌不可兼得。

There are small attempts in, oh sorry, earlier attempts in just, for example, sucking out portions of the cells, taking almost biopsies from the cell to do a fraction of the cytoplasm measurements, so that might be similar to the idea you talk about. There's also work from a public lab to have the cells secrete little vesicles and they harvest that in the cell culture media to measure what the cells are producing longitudinally. But there hasn't been technology that can let you measure the entire cell transcriptome while still keeping a cell. You can't have the cake and eat it.

闭幕词与行动号召 Closing Remarks and Call to Action

Host

酷。是的,感谢你抽出时间与我们聊天。非常棒,学到了很多。有很多非常有趣的讨论,我特别欣赏你对开放科学的承诺,以及你发布的所有酷炫模型和数据。你还有什么最后的想法,或者有什么想让观众知道或关注的?

Cool. Yeah, thank you for taking the time to chat with us. It's been great. Learned a lot. It was a lot of really interesting discussions, and I think I especially really appreciate your commitment to open science and all of the cool models and data you've released. Is there any last thoughts you have or anything you'd like the audience to know or follow up with?

Bo

总的来说,我认为虚拟细胞是一个非常新且快速发展的领域。我们希望越来越多的人加入我们。我们的 X-Cell 论文已经发表,我们期待收到你的评论和反馈。另外,我们正在招聘。是的,我们一直在寻找有才华的工程师、技术专家、生物学家、药物研发人员、AI 科学家、计算生物学家。所以请查看 zera.com,寻找开放的职位。我们很乐意与你交流。

Overall, I think virtual cell is such a new and fast-moving field. We hope to have more and more people join us. And our X-Cell paper is out, and we look forward to receiving your comments and feedback. And also, we're hiring. Yeah, we are always looking for talented engineers, technologists, biologists, drug hunters, AI scientists, computational biologists. So look on zera.com, look for the open roles. We'll be happy to chat with you.

Host

谢谢。非常感谢。

Thank you. Thank you very much.

Bo

太好了。谢谢。

Great. Thank you.

互动版:逐字朗读 + 针对本期提问 →