Xaira Therapeutics 正在构建一个 AI 原生的药物发现平台,利用虚拟细胞模型预测细胞反应,加速药物开发。
Xaira Therapeutics is building an AI-native drug discovery platform with virtual cell models to predict cellular responses and accelerate drug development.
要点 · TL;DR
虚拟细胞模型需要因果扰动数据,而不仅仅是观察数据,才能预测药物反应。 Virtual cell models need causal perturbation data, not just observational data, to predict drug responses.
高质量、大规模的扰动数据集是模型性能最关键的因素。 High-quality, large-scale perturbation datasets are the most critical factor for model performance.
扩散语言模型在预测扰动反应方面优于自回归模型。 Diffusion language models outperform autoregressive models for predicting perturbation responses.
核心观点 · Key points
虚拟细胞模型需要因果扰动数据,而不仅仅是观察性数据,才能预测药物反应。 Virtual cell models require causal perturbation data, not just observational data, to predict drug responses.
高质量、大规模的扰动数据集是模型性能最关键的因素。 High-quality, large-scale perturbation datasets are the most critical factor for model performance.
扩散语言模型在预测扰动反应方面优于自回归模型。 Diffusion language models outperform autoregressive models for predicting perturbation responses.
整合多样化的生物学先验知识提高了对未见细胞类型的泛化能力。 Incorporating diverse biological priors improves generalization to unseen cell types.
开放科学和数据共享对于推动虚拟细胞领域的发展至关重要。 Open science and data sharing are essential for advancing the virtual cell field.
虚拟细胞模型可以泛化到未见情境,从而在原代细胞和复杂系统中进行预测。 Virtual cell models can generalize to unseen contexts, enabling predictions in primary cells and complex systems.
反共识 · Contrarian takes
在观察性数据上训练的基础模型在扰动任务上往往无法超越线性基线。 Foundation models trained on observational data often fail to beat linear baselines on perturbation tasks.
该领域缺乏统一的基准;当前的指标如MAE在评估扰动模型时不可靠。 The field lacks unified benchmarks; current metrics like MAE are unreliable for evaluating perturbation models.
学术实验室难以跟上步伐;工业界在资源和数据生成方面具有主要优势。 Academic labs are struggling to keep pace; industry has major advantages in resources and data generation.
虚拟细胞的瓶颈不是模型架构,而是缺乏因果性、时间序列数据。 The bottleneck for virtual cells is not model architecture but the lack of causal, temporal data.
测序技术会杀死细胞,限制了建模;需要非破坏性的时间测量。 Sequencing technology that kills cells limits modeling; non-destructive temporal measurement is needed.
大规模蛋白质测量比RNA更能理解细胞功能。 Protein measurement at scale is more important than RNA for understanding cellular function.
本期章节 · Chapters(共 32)
开场与嘉宾介绍Introduction and Guest Introductions
Xaira的使命与AI平台Xaira's Mission and AI Platforms
模型瓶颈与整合Bottlenecks and Integration of Models
X-Cell影响与药物发现挑战X-Cell's Impact and Drug Discovery Challenges
单细胞基因组学基础模型Foundation Models for Single-Cell Genomics
从观测数据学习因果的挑战Challenges in Learning Causality from Observational Data
大规模生成因果数据:Perturb-seqGenerating Causal Data at Scale: Perturb-seq
理解扰动与基因表达Understanding Perturbations and Gene Expression
Perturb-seq原理:CRISPR与混合实验How Perturb-seq Works: CRISPR and Pooled Experiments
设计引导RNA与基因沉默Designing Guide RNAs and Silencing Genes
单细胞RNA-seq扩展读数Scaling Readout with Single-Cell RNA-seq
扩展至百万细胞:工程挑战Scaling to Millions of Cells: Engineering Challenges
化学固定与干细胞Chemical Fixation and Stem Cells
空间背景与虚拟细胞模型Spatial Context and Virtual Cell Models
空间转录组学与蛋白质组学解析Spatial Transcriptomics and Proteomics Explained
从批量到单细胞再到空间组学From bulk to single-cell to spatial omics
从自回归到扩散语言模型From autoregressive to diffusion language models
架构选择:扩散vs自回归Architecture Choice: Diffusion vs Autoregressive
数据质量、架构与先验知识Data Quality, Architecture, and Prior Knowledge
模型规模与投资Model Scale and Investment
虚拟细胞愿景与泛化Vision of Virtual Cells and Generalization
对原代细胞的泛化Generalization to primary cells
基础模型vs线性基线Foundation models vs linear baselines
单一vs组合扰动Single vs combinatorial perturbations
AI时代科学家的角色Role of scientists in the AI age
政府为何应资助学术界Why Government Should Fund Academia
开放科学与数据共享Open Science and Data Sharing
AI时代培养品味Developing Taste in the Age of AI
智能体AI时代的训练Training in the Era of Agentic AI
学术界与产业界的共生创新Symbiotic Innovation in Academia and Industry
魔法棒:蛋白质测序与时间动态Magic Wand: Protein Sequencing and Temporal Dynamics