Alex Reeves 讨论 ESMC 的可编程生物学世界建模方法,设计蛋白质结合剂和抗体,以及蛋白质语言模型背后的缩放定律。
Alex Reeves discusses ESMC's world modeling approach to programmable biology, designing protein binders and antibodies, and the scaling laws behind protein language models.
要点 · TL;DR
用宏基因组数据扩展蛋白质语言模型,可产生涌现能力并达到最先进的结构预测水平。 Scaling protein language models with metagenomic data yields emergent capabilities and state-of-the-art structure prediction.
ESMC 作为蛋白质生物学的世界模型,无需多序列比对即可进行设计。 ESMC acts as a world model for protein biology, enabling design without multiple sequence alignments.
通过稀疏自编码器进行机制可解释性,揭示了蛋白质模型中的层级生物学特征。 Mechanistic interpretability via sparse autoencoders reveals hierarchical biological features in protein models.
核心观点 · Key points
缩放定律适用于蛋白质语言模型,更多数据和算力会带来涌现能力。 Scaling laws apply to protein language models, with more data and compute leading to emergent capabilities.
加入宏基因组数据消除了收益递减,比仅用 UniRef 实现了更好的 Scaling。 Adding metagenomic data removes diminishing returns, enabling better scaling than with UniRef alone.
ESMC 构建了蛋白质生物学的世界模型,无需多序列比对即可进行结构预测和设计。 ESMC builds a world model of protein biology, enabling structure prediction and design without MSAs.
通过稀疏自编码器的机制可解释性揭示了与生物学对应的层次化特征。 Mechanistic interpretability via sparse autoencoders reveals hierarchical features mirroring biology.
下一个前沿是细胞建模,需要新的数据和技术来构建数字表示。 The next frontier is modeling the cell, requiring new data and technologies for digital representations.
反共识 · Contrarian takes
蛋白质语言模型能以高成功率设计抗体和单链抗体,优于先前方法。 Protein language models can design antibodies and SCFvs with high success rates, outperforming prior methods.
不需要多序列比对等归纳偏置;仅靠 Scaling 数据和算力就能达到最先进的结构预测。 No inductive biases like MSAs are needed; scaling data and compute alone yields state-of-the-art structure prediction.
抗体(为多样性而进化)从语言模型中获益比从进化多序列比对中更多。 Antibodies, which evolve for diversity, benefit more from language models than from evolutionary MSAs.
细胞最好通过信息论而非第一性原理物理模拟来理解。 The cell is best understood through information theory, not first-principles physics simulation.
蛋白质中的微小遗传变异对学习功能至关重要,而不仅仅是大的多样性。 Small genetic variations in proteins are crucial for learning function, not just large diversity.
本期章节 · Chapters(共 24)
ESMC与世界模型简介Introduction to ESMC and World Modeling
扩展承诺与苦涩教训Commitment to Scaling and the Bitter Lesson
蛋白质扩展为何有效Why Scaling Works for Proteins
ESM模型历史与ESMC发布History of ESM Models and ESMC Release
ESM Fold与蛋白质结构预测ESM Fold and Protein Structure Prediction
稀疏自编码器的机制可解释性Mechanistic Interpretability with Sparse Autoencoders
蛋白质建模中的压缩与隐变量Compression and hidden variables in protein modeling
ESM模型的数据可用性与扩展Data availability and scaling in ESM models
宏基因组测序与蛋白质发现Metagenomic sequencing and protein discovery
模型学习与后训练实现可编程性Model learning and post-training for programmability
ESMC:可编程生物学的世界模型ESMC as a world model for programmable biology
蛋白质设计空间与抗体Protein design space and antibodies
无需MSA与涌现行为No MSA needed and emergent behaviors
机制可解释性与意外模式Mechanistic interpretability and unexpected patterns
ESM Atlas与生物学发现ESM Atlas and Biological Discovery
ESMC Multimer与虚拟细胞ESMC Multimer and Virtual Cell
实验室闭环愿景Lab-in-the-Loop Vision
AI驱动生物学愿景Vision for AI-driven biology
细胞的复杂性Complexity of the cell
连接静态结构与动态Bridging static structures and dynamics
机器学习驱动生物学的数据收集Data collection for ML-driven biology
扩展生物学:扰动与空间生物学Scaling biology: perturbation biology and spatial biology