Chai Discovery:AI 原生的药物发现软件工厂

Chai Discovery: The AI-Native Software Factory for Drug Discovery

马修·麦克帕特隆 Matthew McPartlon · Latent Space · 2026-08-11 · 约 95 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Chai Discovery 的创始人讨论他们的 AI 模型和设计套件如何将药物发现从瀑布流程转变为敏捷循环,并与大型制药公司建立合作。

Chai Discovery's founders discuss how their AI models and design suite are transforming drug discovery from a waterfall process into an agile loop, with partnerships with major pharma companies.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 36)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Backgrounds

Host

欢迎收听《Len Space AI for Science》。我是 Brandon,在 Atomic AI 从事 RA 疗法研发。今天和我一起主持的是联合主持人 R.J. Honiki,Mirror OMIX 的 CTO 兼联合创始人。今天我们非常荣幸地邀请到了 Chai Discovery 的 Magma Partland 和 Neil Patil 来到演播室。Chai 是一家蛋白质设计初创公司,成立大约两年半,在这几年里引起了不小的轰动。他们有几项非常令人兴奋的公告,我想他们今天会告诉我们。那么,首先,你们两位能介绍一下你们的背景以及在 Chai 的职责吗?

Welcome to Len Space AI for science. I'm Brandon. I build RA therapeutics at A Atomic AI. I'm joined by my co-host R.J. Honiki, CTO and co-founder of Mirror OMIX. It's a pleasure to have with us in the studio today Magma Partland and Neil Patil of Chai Discovery. Chai is a protein design startup which is about 2 and a half years old and has made quite a splash in those few years. They have several very exciting announcements that I think they'll tell us about today. But yeah, to get started, could you two give us a bit about your background and what you do at Chai?

Matt

是的,非常感谢你们邀请我们。我们今天非常兴奋能来聊聊 Chai。我是 Matt McPartland,Chai 的联合创始人之一。我的背景是博士期间做 AI 生物学相关的研究。实际上,我博士一开始是学理论计算机科学的,后来才转到这个方向。我做这行大概 8 年了,进入这个领域的时候正好是蛋白质结构预测刚刚开始显现出生命力的时期,也就是 AlphaFold 1 的时代。后来在 AlphaFold 2 期间我也在这个领域,见证了很多有趣的发展。我一直对把这类技术应用到现实世界很感兴趣,而 Chai 正好是一个绝佳的机会。

Yeah, thank you very much for having us. We're super excited to talk about Chai today. I'm Matt McPartland. I'm one of the co-founders of Chai. My background is in AI biology related stuff during my PhD. I actually started my PhD in theoretical computer science and then transitioned to this later. I've been doing this stuff now for about 8 years and I kind of came into the field at an interesting time where protein structure prediction was just starting to see signs of life. So this is like AlphaFold 1 days. And was in the field during AlphaFold 2 and got to see a lot of the interesting developments at that time. So yeah, I'd always been pretty interested in applying this stuff in the real world and Chai was just a perfect opportunity to do that.

Neil

我是 Neil Patiel,在 Chai 负责平台和产品。主要是训练模型的基础设施、模型服务,以及产品化部分,也就是让你能使用模型的设计套件。我的路径比较曲折。大概 15 年前我开始编程,在 App Store 里做应用,对那种多巴胺冲击上瘾了。后来被机器人技术吸引,做了一段时间,2018、2019 年搞自动驾驶汽车。后来真的厌倦了,心想暂时不想碰硬件了。于是转行加入了 SaaS 公司 Vanta,成为最早的一批员工之一,跟着公司一起成长。之后自己创办了一家安全公司。做了几年后,我想,你知道吗,原子其实挺酷的。我想做点更有意义的事情,所以大约一年前,就在 Chai 2 发布后不久,我加入了 Chai,帮助处理平台和商业化方面的事情。

And I'm Neil Patiel. I help lead platform and product here at Chai. So a lot of the stuff around infrastructure to train models, serve them, and then the productization piece, the design suite that lets you use the models. I kind of have a more meandering path. I got into programming like 15 years ago, making apps in the app store. Got really addicted to the dopamine hits you get from that. And then actually got nerd sniped by robotics and worked on that for a bit. Self-driving cars in like 2018, 2019. Got really jaded and was like, I don't want to touch hardware for a while. Ended up switching and joining a SAS company called Vanta as one of the first employees there and kind of grew with it. Started my own security company afterwards. Got a few years into that and I was like, you know what, atoms are kind of cool. I want to work on something a little more meaningful and so I joined Chai about a year ago, right after Chai 2 was announced, to help with a lot of the platform and commercialization pieces.

Host

太棒了。这就像悲伤的五个阶段之类的。

Awesome. It's like the five stages of grief or something.

Matt

是的。我们已经到了接受阶段。

Yeah. We're at acceptance.

合作与商业模式 Partnerships and Business Model

Host

太棒了。你们现在有四个大合作伙伴,还筹集了一大笔资金。能跟我们聊聊这些合作吗?我真正想知道的是,你们跟投资人和客户说了什么,让他们愿意达成这些大交易?

Awesome. You have these, I think, four now big partnerships and raised a whole bunch of money. Can you tell us a little bit about those partnerships and then what I really want to know is what are you telling investors and customers that is so compelling that they're willing to do these big deals?

Matt

是的。我们非常幸运,首先与礼来合作,然后与辉瑞、诺和诺德以及我们的 Gen X 合作。我觉得这是一段非常有趣的旅程,而且我认为我们的商业模式对很多人来说也很有吸引力。我们非常关心合作伙伴的成功。Chai 作为一家公司,真的取决于合作伙伴的成功。我想 Neil 可能对我们实际提供的东西以及为什么这么有吸引力有一些有趣的见解。所以我把话交给他。

Yeah. So we've been very fortunate to partner first with Eli Lilly and then with Pfizer, Novo Nordisk, and our Gen X. Yeah, I think it's been a really interesting ride and I think our business model is also very compelling to a lot of people. We really care about the partners succeeding. Chai as a company really depends on how the partners succeed. I think Neil probably has some interesting takes on what we actually offer and what makes that so compelling. So I'll hand it over to you.

Neil

是的,你知道,药物发现是一个非常漫长的过程,对吧?很多制药公司花费大量时间,年复一年,投入数十亿美元,试图找到最初的治疗候选药物。所以在 Chai,我们训练模型来加速这个过程,找到那些初始的结合剂等等。你知道,有很多生物公司,AI 生物公司,他们在自己研发药物。我们不这么看自己,对吧?我们几乎把自己看作一个中立的药物制造软件工厂。因此,这让我们能够去与所有这些其他制药公司合作,支持他们的药物发现之旅。所以,是的,很多资本只是另一个证明点,证明我们可以真正加速这个软件工厂,对吧?去攻克更难的模态,训练更大的模型,最终构建我们的合作伙伴和客户要求我们构建的东西。

Yeah, I mean, as you all know, drug discovery is a very lengthy process, right? And a lot of these pharma companies are spending lots of time, you know, years and years and billions of dollars trying to find initial therapeutic candidates. And so at Chai, we train models that can help accelerate that process and kind of find those initial binders and then some. And you know, there's a lot of bio companies, AI for bio companies that are like making their own drugs. We really don't see ourselves that way, right? We see ourselves as almost a neutral software factory for making medicines. And so that's what lets us go then work with and support all of these other pharmas in their kind of drug discovery journey. And so yeah, I mean a lot of this capital is just another proof point that we can sort of start to really accelerate that software factory, right? Go after harder modalities, train bigger models and ultimately just build what our partners and customers ask us for.

Host

但为什么是你们,而不是其他结构生物学公司?为什么他们愿意向你们购买?

But what is it that why you and not other structural companies? Why are they compelled to buy from you?

Matt

Chai 的核心理念一直是成为软件和建模层,我认为这在当时非常有争议。就像每个人,你知道,这个玩法绝对是……

The thesis of Chai has always been to be the software and modeling layer, which I think was very controversial at the time. Like everyone, you know, this play is definitely...

Host

两年前,现在已经是完全不同的世界了。

Two years ago and it's already like a completely different world.

Matt

是的。这太疯狂了。人们尝试这种玩法有一段时间了。我认为模型当时真的还没准备好。即使对我们来说,一开始也是在冒险。我们基本上押注模型会发展起来。我在工作中看到了早期的生命力迹象,我们的 CEO Josh,他在 Meta 的那个团队里参与了最初的 ESM 论文,他看到了相当早期的迹象,表明这里可能存在缩放定律。我认为我们实际上将能够开始设计东西。

Yeah. It's pretty crazy. Like people tried this play for a while. And I think like the models just really weren't there yet. And even for us, we were taking a risk in the very beginning. We were kind of banking on the models getting there. And I had seen early signs of life in my work and our CEO Josh, he was on the original ESM papers on that team at Meta and he was seeing pretty early signs of life that there might be scaling laws here. I think we'll actually be able to start designing things.

产品愿景与未来 Product Vision and Future

Matt

它看起来不太像 ChatGPT,而更像 Autodesk、SolidWorks 或 Figma,你知道,如果你用过那些工具,你可以加载你的分子。这几乎就像一个类似 Photoshop 的设计套件。你有相当于画笔工具的东西来绘制你的表位。你有相当于内容感知填充工具的东西,从 Chai 生成你的结合剂。我想补充一点,是的,目标发现、命中发现和优化这个概念,每一步都有一个门槛,需要几个月到几年的时间,这是一种非常瀑布式的模型,对吧?早期尝试和获取东西的成本非常昂贵。但我认为,正如 Matt 所说,如果你开始进入一个模型能给你真正有前景的候选物的阶段,你就可以开始让它看起来更像一个循环,对吧?这类似于在软件开发中变得更加敏捷。但现在下一个问题是激动剂,对吧?比如如何可靠地一次性击中细胞上的开关,对吧?或者购买特异性或 ADC,对吧?我认为随着模型变得更好,我们将不得不与产品一起攀登这些抽象层次。如果你有这些非常好的结构预测、结合和设计的基本单元,并且你可以组合它们,那么你就可以开始进入科学的外循环。

It looks a lot less like a, you know, a ChatGPT and a lot more like Autodesk or SolidWorks or Figma, you know, if you've used those things where you can kind of load up your molecule. There's this almost like Photoshop-esque design suite. You have this equivalent of a paint tool to kind of paint your epitope. You have this equivalent of a content-aware fill tool to kind of get your binders generated from Chai. And I think to add to that, right, yeah, this notion of target discovery and hit discovery and optimization where each of these has a gate and takes a few months to a few years is this very like waterfall model, right? Where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, right, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop, right? It's akin to becoming more agile in software development. But now the next problem is like agonists, right? Like how do you reliably one-shot hitting a switch like on a cell, right? Or buy specifics or ADCs, right? And I think this levels of abstraction that we're going to have to climb with the product as the models get better. If you have like these really good primitives for structure prediction and binding and design and you can kind of compose them, then you can start to just like grow into like the outer loop of science.

结构预测进展与通用性赌注 Structure prediction progress and the bet on generality

Matt

结构预测现在变得非常出色。一个疯狂的想法是,直到 2021 年我们才有了多聚体结构预测模型。那是 5 年前,我们才开始用深度学习真正预测两个蛋白质的形状。AlphaFold 1 就是这样,AlphaFold 2 是巨大的突破。但一年后 AlphaFold 2 multimer 就出来了。所以你真的需要那个才能解锁设计。我们当时甚至没想同时预测多个蛋白质。大约在那时,逆折叠开始奏效,我们意识到,哦,蛋白质设计在实验室里真的管用。这要归功于 Baker 实验室,他们对所有模型做了出色的实验室验证,但我觉得我们开始看到他们做有趣的事情,真正在现实世界的实验上工作。现在可能是开始押注这个的时候了。在那之前,你可能会从针对某个目标的活动中获取一些实验数据,取得一些进展,在某个特定案例上持续爬山。那时候通用模型还不存在。所以我们认真对待这个赌注,决定尽可能努力,真正追求方法的通用性。当 Chai 2 出来时,也就是我们继 Chai 1 之后的第二篇论文,我们向世界展示了这实际上是可能的,而且可以规模化。我们不是只针对一两个目标展示;我们全力以赴。Josh 喜欢说我们设定了一个大胆的全公司挑战,设计针对 50 个目标的抗体,我们看到了生命的迹象。我们说,让我们用真实的统计数据来做,看看这是否真的有效。

Structure prediction is getting really good. Like one crazy thought is that we didn't have a multimer structure prediction model until like 2021. That was 5 years ago when we could start using deep learning to actually predict the shape of two proteins at once. AlphaFold 1 was like that, and AlphaFold 2 was a huge breakthrough. But then AlphaFold 2 multimer came out like a year later. So you really needed that to unlock design in the first place. We weren't even trying to predict multiple proteins at once. And around that time, inverse folding started working, and it was like, oh, protein design actually works in the lab. Credit to the Baker lab for doing excellent lab validation on all their models, but I think we're starting to see them do interesting things and actually work on real-world experiments. Now is probably the time to start betting on this. Before then, you might take some experimental data from a campaign on one target you cared about and make some progress, keep hill climbing in that one specific case. General models weren't really a thing back then. So we took that bet pretty seriously and decided to push as hard as possible and really shoot for generality in our approach. When Chai 2 came out, our second paper after Chai 1, we showed the world that this is actually possible and possible at scale. We didn't show it for one or two targets; we went all in. Josh likes to say we set a bold company-wide challenge to design antibodies to 50 targets, and we saw some signs of life. We said, let's do this with real statistics and see if it actually works.

目标选择与50目标挑战 Choosing targets and the story behind the 50-target challenge

Matt

我们选择这些目标的过程很有意思。我们说,我们要选什么目标?应该选一些有趣的目标。那时我们正在和 CRO 合作,摸索湿实验室流程。我们决定,在尝试了一些蛋白质之后,这些是有趣的目标,我们应该关注这些。但有一半时间目标就是不行。我们还在学习。我们说,也许我们应该选择 CRO 已经验证过的目标。所以我们拿到 CRO 的目录,看看他们做过什么,然后限制在一个有趣的集合里。从中我们选了 50 个目标,设计抗体,对一半有了命中。那时,制药公司开始意识到这里有生命的迹象,这可能真的适用于他们的一些项目。

It's an interesting story of how we chose these targets. We said, what targets are we going to choose? We should choose some interesting targets. At that point we were ramping up with CROs and figuring out what our wet lab process looks like. We decided, after trying some stuff with many proteins, here are the interesting targets, this is what we should look at. Half the time the targets just didn't work. We were still learning. We said, maybe we should go with targets that the CROs have actually validated. So we got the CRO catalog, saw what they'd already worked on, restricted that to an interesting set. From that we chose 50 targets, designed antibodies against them, got hits to half. At that point, pharma starts to realize that there are signs of life here, and this might actually work in some of our programs.

抗体:挑战与吸引力 Why antibodies are a challenging but attractive domain

Host

所以抗体可能比其他结构预测问题更具挑战性。那为什么要攻克抗体?也许先退一步,什么是抗体?你用它做什么,为什么它是一个有吸引力的目标?

So antibodies are maybe a more challenging domain than other structural prediction problems. So why tackle antibodies? Maybe back up, what is an antibody? And what do you do with it, and why is it an attractive target?

Matt

大家常说的类比是锁和钥匙的问题,你的目标是你试图结合的蛋白质,可能是疾病蛋白,那是你的锁,你想设计一把能插进去的钥匙,在我们的情况下就是卡在那里。抗体的有趣之处在于它们是真正灵活通用的蛋白质。在很多方面它们非常通用,在很多方面它们实际上相当统一。但至少它们如何结合目标是非常通用的。所以你在设计结合界面时有很多选择。抗体的结构预测问题,预测抗体如何结合目标,钥匙如何插入锁,一直是一个出了名的难题。好的一面是我们在结构预测上取得了很大进展,整个领域已经走了很长的路。但在设计场景中,你可以对想要的设计类型和关注的结构类型更加挑剔。在某些情况下,设计一个作为抗体的蛋白质结合剂实际上可能比预测它如何结合目标更容易。如果你有选择的自由,你可以只挑容易的案例。

The analogy everyone gives is this lock and key problem, where your target is a protein you're trying to bind to, maybe a disease protein, that's your lock, and you want to design a key that fits into it, and in our case just sticks there. The interesting thing with antibodies is that they're really flexible general proteins. In many ways they're very general, in many ways they're actually pretty uniform. But at least how they bind to a target is very general. So you have a lot of optionality in how you design this binding interface. The structure prediction problem for antibodies, predicting how the antibody binds to the target, how the key fits into the lock, has been a notoriously difficult problem. The nice thing is that we've made a lot of progress on structure prediction, the field as a whole has come a long way. But in the design setting, you can be a lot more selective about the types of designs you want to make and the types of structures you focus on. In some cases, it might actually be easier to design a protein binder that is an antibody than to predict how it might bind that target in general. If you have the freedom to choose, you can just pick the easy cases.

抗体结构与治疗应用 Antibody structure and therapeutic applications

Host

所以抗体,身体里有一套完整的机制与抗体协同工作。身体自然用它做什么,你能用它做什么不是自然的但对治疗有用的?

So antibodies, there's a whole machinery in the body that works with antibodies. What does the body do with it naturally, and what can you do with it that is not natural but useful for therapeutics?

Matt

这是从一个非生物学家的角度来说的,但我把抗体想象成 Y 形蛋白质,就像手指比出的和平手势。每个手指就像抗体的一条臂。实际上只有指尖,也就是抗体的尖端,参与结合。这使得它们成为该特定区域非常好的治疗设计目标。好的一面是,除了尖端之外的部分相对恒定。这被称为抗体的框架区。在设计问题中,你通常只设计指尖部分,而且你实际上可以在很大程度上选择你的免疫系统已经识别的框架区。所以抗体是这些 Y 形蛋白质,你的免疫系统非常熟悉。它是你身体抵御病原体和其他类型疾病的防线之一。一端通常连接细胞表面的蛋白质,另一端帮助免疫系统识别病原体。但你也可以做像 ADC(抗体药物偶联物)这样的事。那意味着在另一端放一个药物,当你结合到某个东西时,它会将药物释放到细胞中。

This is coming from a non-biologist, but I think of antibodies as these Y-shaped proteins, like a peace sign with your fingers. Each of these fingers is like an arm of the antibody. It's really only the tips of your fingers, the tips of the antibody, that engage in binding. This makes them really nice therapeutic design targets for that particular region. The nice part is that the rest apart from the tips is relatively constant. This is called the framework region of an antibody. In the design problem, you're typically just designing the very fingertips, and you can actually choose, for the most part, these framework regions that your immune system already recognizes. So antibodies are these Y-shaped proteins that your immune system recognizes really well. It's one of your body's lines of defense against pathogens and other types of diseases. On one end, they connect to proteins on the surface of a cell typically, and the other end helps the immune system identify a pathogen. But you can also do things like ADCs, antibody-drug conjugates. That means putting a drug on the other side, and when you bind to something, it releases the drug into the cell.

Host

对吧?它们就像这个非常通用的框架,对吧?在末端你有这些 CDR 环,你可以设计它们来结合任意东西,也许一端结合癌细胞,另一端结合有毒分子。

Right? They're like this very general framework, right? Where on the ends you have these CDR loops, and you can design them to bind to arbitrary things, where maybe one end you bind to a cancer cell, the other end you bind to a toxic molecule.

精准药物设计 Precision drug design

Matt

你现在是在把有毒分子精准递送到癌细胞上,对吧?或者你只是让两端结合到某些东西上,强制诱导邻近效应,在体内产生某种效果。或者,你知道,历史上很多药物其实真的只是关于阻断,对吧?比如拮抗剂行为。但也许你可以有激动剂行为。我们实际上可以非常精准地按下开关。比如有一种 GPCR 蛋白,它们是细胞膜上的门铃蛋白。你有一个抗体,经过非常精准的工程改造,以某种方式戳它,引发下游连锁反应。我认为,随着这些模型的发展,我们开始能够达到那种精准度,这真的很令人兴奋,对吧?我们可以真正靶向一个非常特定的表位,也就是结合位点,一组非常特定的原子,让抗体去攻击。历史上,对于很多药物,你基本上是在蛮力尝试,你知道,很多抗体,就是试图弄出一堆东西,看看哪个有效。但那可能只让你得到一个结合到靶分子某个位点的结合物,并不能让你精准设计你要戳的位置。

You're now precision delivering that toxic molecule to a cancer cell, right? Or you just have two ends bind to things and kind of force induced proximity to have some effect in the body. Or, you know, a lot of drugs historically are really just about blocking things, right? Like antagonist behavior. But maybe you can have agonist behavior. We can actually really precisely press a switch. Like there's a GPCR protein, which are these doorbell proteins that sit in your cell membrane. You have an antibody very precisely engineered to poke it in a certain way that causes a downstream chain reaction. And I think one of the things that's really exciting about where we're getting to with some of these models is we can start to get that precise, right? We can really target a very specific epitope, meaning binding spot, a very specific set of atoms to have the antibody go after. Historically, with a lot of drugs, you're just kind of brute forcing, you know, a lot of antibodies and just trying to come up with a bunch of things and see what sticks. But maybe that gets you a binder to some spot of your target molecule, but that doesn't let you precisely engineer where you're poking after.

传统药物设计流程 Traditional drug design process

Host

我知道你不是生物学家,但你知道在这些模型出现之前,他们过去是怎么设计这些的吗?比如,你要经历什么艰苦的过程才能找到……

I know you're not biologists, but do you have any idea about how they used to design these before these models came up? Like what was the grueling process you would do to find...

Matt

或者什么是,实际上仍然是……是的,对于已经进入临床的药物,什么仍然是当前最先进的方法?

Or what is, which is actually still... Yeah, what still is the state-of-the-art in terms of drugs which have made it to the clinic?

Josh

是的。我们的 CEO 乔什喜欢说,我们最大的竞争对手是老鼠,或者在某些方面是大自然。所以传统上,这类药物样分子要么是在免疫接种活动中发现的。比如,你真的就是让老鼠感染一种疾病,然后看它产生什么抗体来对抗。其他方法包括超大规模的酵母展示等等。所以你可能从这样的问题开始:嘿,我真的很喜欢这个框架,我该怎么找出正确的环来设计以结合这个靶点?我就是要尽可能多地尝试,真的就是在干草堆里找针。这通常意味着你要针对这一个靶点筛选至少数十亿个潜在分子。在这种情况下,你最终可能得到一两个,或者十几个潜在的命中分子。你对这些命中分子其实了解不多。你只知道它们能粘到靶点上。你并不一定知道它们结合在哪里,或者它们是否具有药物样性。

Yeah. Josh, our CEO, likes to say that our biggest competitor is the mouse, or nature in certain ways. So traditionally, these types of drug-like molecules were either discovered in immunization campaigns. So you literally just infect a mouse with a disease and see what antibodies it makes to try to combat that. Other ways of doing this is like super large yeast display and so on. So you might start with, hey, I really like this framework and how am I going to figure out the right loops to design to bind this target? I'm just going to try as much as I possibly can and just literally search for a needle in a haystack. And this would be on the order of at least billions of potential molecules that you're screening against this one target. In that case, you might end up with one, two, maybe a dozen potential hits to this target. You actually don't know much about those hits. All you know is that they kind of stick to the target. You don't know necessarily where or if they're even drug-like.

AI驱动设计的优势 Advantages of AI-driven design

Josh

我认为 Chai 的一个重大区别,也是我们的合作伙伴肯定喜欢看到的,就是你在这个设计过程中可以非常有意图性。你可以说,我想在这个特定区域结合这个靶点。你甚至可以在我们验证了设计之后,回头去看这些设计。所以你可以回头检查,说,这个抗体是否以我预期的方式与靶点结合?我认为这真的会产生我追求的治疗效果吗?知道你有正确的结合姿态的一个酷之处在于,你现在还可以把选择性设计进去。

I think one big separator of Chai, and a thing that definitely our partners like to see, is you can be really intentional with how you want to do this design process. You can say, I want to bind this target in this particular area. You can even go back and look at the designs after we've validated our designs. So you can go back and look and say, is this antibody engaging the target in the way that I expect? Do I think this will actually have the therapeutic effect that I'm going after? One of the cool things about knowing that you have the right binding pose is that you can now also design selectivity into that.

选择性与交叉反应 Selectivity and cross-reactivity

Host

你们的平台有某种技术来实现选择性吗?

Does your platform have some technique for doing selectivity?

Josh

是的,在建模方面,尤其是在产品方面,我们融合了很多想法来处理选择性和交叉反应性。所以在某些情况下,你希望你的分子结合一个靶点,同时避免另一个。所以你可能会有蛋白质的健康变体和疾病变体。你想避免疾病变体,或者你可能有一些其他类似的蛋白质,它们在你的体内实际上并不有害,你不想人为地阻断它们。所以在建模方面,是的,我们已经想出了办法,但我认为在产品方面更有趣。那么,你如何让客户或合作伙伴能够实际有意地设计这些东西呢?

Yeah, there's a nice mix of ideas that went both into the modeling side and especially on the product side for dealing with selectivity and cross-reactivity. So in some cases, you want your molecule to bind one target and avoid another one. So you might have healthy variants of a protein and disease variants of a protein. You want to avoid the disease variant, or you might have some other similar protein that's not actually harmful in your body that you don't want to artificially block. So on the modeling side, yeah, we've come up with ways of doing that, but I think it's even more interesting on the product side. So how do you enable customers or partners to go through and actually intentionally design for these things?

交叉反应定义与设计策略 Defining cross-reactivity and design strategy

Host

是的。也许让我们退一步,定义一下交叉反应性,对吧?事实证明,当你开发一种药物时,你不一定会直接把它注射到人体内,对吧?比如,你可能想先在猴子身上测试。而猴子可能有大致相似但略有不同的变体。所以你的药物不仅需要结合人类变体,还需要结合猴子变体,对吧?所以我们试图在模型和产品中建模的方式,是让你能够考虑这些非常普遍的情况,比如你说,嘿,我正试图设计一种能同时结合这两者的东西,这样我才能真正去开发药物。让我识别出保守的区域,意思是两者之间变化不大的区域,然后靶向那个精确的区域。然后类似地,对于选择性,对吧?也许你可能想要,在人体内有一个非常相似的蛋白质,如果你不小心结合了那个,那就非常糟糕,你只想结合靶蛋白。这就是为什么很多药物会失败、有毒或有非常严重的副作用。所以这有点像组合问题:只结合这些,避免那些。我认为最近一些进展真正令人兴奋的是,我们在这些模型能获得的特异性水平上取得了许多改进。

Yeah. And maybe to back up and define cross-reactivity, right? Like it turns out when you're developing a drug, you're not necessarily going straight to injecting that into a human, right? Like you might want to put it in monkeys first, for example. And the monkey might have a mostly similar but slightly different variant of it. And so your drug not only needs to bind to the human variant, but also the monkey variant, right? And so the way we've tried to model the models and the product is to let you account for those very general cases where you say, hey, I'm trying to design something that can bind to both of these things so that I can actually go and develop the drug. Let me actually identify maybe the region that's conserved, meaning doesn't change much between the two, and target that exact region. And then similar with selectivity, right? Maybe you might want to, there's a very similar protein in the human that if you accidentally bind that one, that's very bad, and you only want to bind the target protein. And that's why a lot of drugs fail or are toxic or have really bad side effects. So it's kind of a combinatorial problem of bind only these things and avoid only these. And I think what's been really exciting with some of the progress recently has been a lot of the improvements we've been able to make on the level of specificity we can get with those models.

替代方法与反向筛选 Alternative approaches and counter-screening

Host

所以你不仅在设计结合物,还在确保它不会结合到其他东西上。所以其他公司尝试解决这个问题的方法,是拥有某种分子或信号通路,说,只有当我结合这个而不结合那个时,我才触发。但你说的是,你直接设计一个抗体,它实际上只会结合你关心的那个东西。

So you're not only designing the binder, but you're also making sure that it doesn't bind to another thing. So other ways that companies have tried to tackle this is by having some molecular or signaling pathway that says, I only fire if I bind this one and this one doesn't bind. But you're saying you just design an antibody that actually will only bind to the thing that you care about.

Josh

我们正在达到这样的程度,在某些情况下,我的意思是,这很微妙,对吧?但在某些情况下,你实际上可以尝试这样做。

We're getting to the point where in some cases, I mean, it's nuanced, right? But in some cases, you can actually try that.

Host

好的。太棒了。所以你是说,你基本上称之为反筛选,或者在你的平台中,你可以可靠地针对大量多样化的蛋白质进行反筛选,这些蛋白质可能会在下游造成问题。

Okay. That's amazing. So you're saying you essentially call it counter-screen, or you have in part of your platform you can reliably counter-screen against a large diverse set of proteins which might be issues for downstream.

Josh

是的。我会说,更准确的框架是,你可以非常具体地说明你关心结合什么,以及你关心避免什么。

Yeah. I would say the framing is more you can be very specific about what you care about binding versus what you care about avoiding.

资金与模型扩展 Funding and Model Scaling

Matt

但我觉得,比如说,我们现在筹到的很多钱,会让我们能训练更大的模型,这些模型可能会更通用,能同时处理更多的事情。

But I think you know, for example, a lot of the money that we're raising now will let us train bigger models that can maybe be even more general and start to account for even more things at the same time.

Host

对。也许我们应该往回讲。我们来聊聊 Chai 系列模型的历史吧。

Right. Maybe we should back up. Let's talk about the history of the Chai series of models.

Matt

嗯,要不你来讲讲这个故事?

Well, why don't you tell the story?

Host

我们大约两年半前创办了 Chai。最初几个月,我们想:“好,我们要做蛋白质设计。” 我们一直在做这个,也取得了一些进展。我们觉得:“哦,这挺有意思的。” 我们有一些想法和模型。然后正好 AlphaFold 3 发布了。我们当时在讨论,天哪,我们真的需要一个 MSA 流程。我们需要所有这些基础设施。

We started Chai around two and a half years ago. At first couple months we were like, "All right, we're gonna work on protein design." And we were working on this. We were making some progress. We were like, "Oh, this is pretty interesting." We had some ideas and models. And then that was right when AlphaFold 3 came out. And we were talking about like, man, we really need an MSA pipeline. We need all this infrastructure.

Matt

MSA 就是多序列比对。

MSA is multiple sequence alignment.

Host

为什么这……我们之前讲过,但用两句话说说什么是 MSA,为什么它很重要?

Why is this... We've covered this before, but what is an MSA in like two sentences and why is it important?

Matt

所以如果你想预测蛋白质的结构,看一堆非常相似的蛋白质序列可能会非常有用。这些非常相似的序列告诉你的是,哪些位置,也就是哪些氨基酸,在这个蛋白质的许多变体中最终是保守的。如果你看到高水平的保守性,或者高水平的关联突变,这通常表明这些氨基酸在三维空间中是接近的。这样你就有了蛋白质的二维视图,可以用来帮助预测三维结构。

So if you want to predict the structure of a protein, it might be really useful to see a bunch of very similar protein sequences. And what those protein sequences that are really similar tell you is kind of what positions, like which amino acids end up being conserved across many variants of this protein. And if you see high levels of conservation or high levels of correlated mutations, it typically gives you some indication that these amino acids are close in 3D space. So you kind of have this 2D view of a protein which can then be used to help you predict the 3D structure.

Host

所以你是从进化中学习哪些是保守的,因为那些不保守的东西可能破坏了蛋白质,导致某些生物死亡或没能存活下来。

So you're learning from evolution what was conserved because the things that weren't conserved probably broke the protein and something died or didn't make it.

Matt

完全正确。对,是的。说实话,这能行得通真是了不起。这是我最喜欢的生物学事实之一。所以我们当时想,天哪,要是能有大量基础设施就好了。所以 AlphaFold 3 发布了。我们想,嘿,我们应该开源这个模型。我们应该埋头苦干,把我们需要的所有基础设施都建起来。我认为从长远来看这肯定会有回报,既是推动我们走到现在的动力,也是对整个社区的贡献。

Exactly. Right. Yeah. It's pretty remarkable that this works honestly. One of my favorite bio facts here. So we were thinking like, oh man, it'd be nice to have a lot of infra and whatever. So AlphaFold 3 came out. We're like, hey, we should open source this model. We should just bunker down, build all the infrastructure that we need. I think this will pay back in the long term for sure, as a forcing function to be where we are and also to contribute to the community as a whole.

Host

所以有趣的是,你选择了,好吧,我们实际上在做的是,我们在构建一个模型,但我们真正在做的是学习如何构建基础设施。你是这个意思吗?

So it's interesting that you chose, okay, this is what we're actually doing here, we're building a model, but what we're really doing is learning how to build the infrastructure. Is that kind of what you're saying?

Matt

是的,完全正确。我在博士期间也建过很多类似的基础设施,但不是为公司达到生产级别的。所以那时我们大概有五个人。我们五个人在 Chai,然后我们说,好吧,这就是我们的推动力。我们有一个明确的目标要去实现。非常直接。让我们把这件事做起来,看看我们能多快完成。

Yeah, that's exactly right. And I had built a lot of similar infrastructure in my PhD but not at a production level for a company. So at that point I think we were five people. So there were five of us at Chai and we're like, all right, this is our forcing function. We have a clear goal to work towards. It's very direct. Let's get this thing going and see how fast we can do it.

Host

你们当时是坐在 OpenAI 的办公室里。是吗……

You guys were at this time sitting in the OpenAI office. Is that...

Matt

我们坐在 OpenAI 的办公室里?是的。在使命区。

We were sitting in the OpenAI office? Yeah. In the mission.

Host

对吧?那背后的故事是什么?这真的很有意思。

Right? So what's the backstory on that? That's really interesting.

Matt

我们的另外两位联合创始人 Josh 和 Jack 实际上和 OpenAI 的一些人有关系。OpenAI 共同领投了我们的种子轮。所以我们当时想,好吧,我们只有五个人,要不要找个办公室?结果那个办公室大部分是空着的。所以我们就在 OpenAI 的办公室里坐了一段时间。我们把它做成了开源,学习了基础设施。

Two of our other co-founders, Josh and Jack, had a relationship with some of the OpenAI people actually. OpenAI co-led our seed round. So we were kind of thinking, all right, should we get an office while we're only five people? And it turned out that office was mostly vacant. So we got to sit in the OpenAI offices for a while. We built it open source, learned about infrastructure.

Host

是的。所以在那之后,我们真正把目光投向了蛋白质设计。值得指出的是,Chai 1 是一个结构预测模型,对吧?所以你有序列,它折叠成的结构是什么?然后那就是 Chai 2。

Yeah. So then after that, we really set our sights down on protein design. And worth pointing out, Chai 1 was a structure prediction model, right? So you have the sequence, what is the structure that it folds to? And then that was Chai 2.

Matt

是的,Chai 1。Chai 1 完成了。还有一个疯狂的故事。看看我们能不能分享。但这是个很搞笑的故事。我们当时想:“天哪,我们真的很想第一个发布这个。” 然后我们说:“好,我们还有一周时间。模型快训练完了。” 我们想:“我们要不要建一个网页服务器?” 然后我们想:“哦,也许不用。” 但最后我们还是搭了一个完整的网页服务器,这样人们就可以直接使用,而不是下载 Git 仓库。这对生物学家来说尤其麻烦。而且我们确实希望人们使用这个。所以让我们搭一个网页服务器。让我们把技术报告发出去,所有这些。所以我们连续 48 小时没睡,就为了把论文搞定,把网页服务器上所有最后的事情做完。然后 Josh 那天早上要接受彭博电视台的采访。我们已经连续 48 小时没睡了。所以 Josh 跑进一个房间去做彭博电视台的采访。我记得大概是早上 7 点。大家都在办公室里。我们不想被看到,无所谓。采访者说:“哦,有趣的公司。看起来这里好像没有员工。” 但那是段非常快乐的时光。我觉得早期创业的日子就是超级有趣。所以在那之后,我们把目光投向了设计。我们真正想的是,我们一直把抗体放在心上。我们认为这是最容易处理的问题。蛋白质的好处是,你有这种漂亮的序列表示。而且已经有很多研究是关于如何自回归生成序列的。你怎么……这个序列生成问题已经被充分研究了。所以我们想,在生物领域,哪里是应用序列生成的好地方?很自然的就是做氨基酸的线性序列。所以我们开始做设计。Chai 的一个独特之处是,我们不是……我们在设计抗体,我们是一家抗体公司,但我们不会把自己局限在一个治疗领域。所以我们试图非常普遍地解决这个问题。所以我们想,我们能设计许多蛋白质吗?我们能设计抗体吗?我们能搭建常规复合物吗?所以真正从整体上看如何设计蛋白质。这最终导致了 Chai 2 模型。那是我们的第一个旗舰设计模型,Chai 2 论文和我们的“大胆靶点发现”项目就是在那里产生的。所以我们为那篇论文设计了针对 50 个靶点的抗体。大约一半获得了结合物,我认为平均结合命中率约为 20%。

Yeah, Chai 1. Chai 1's finished. One other crazy story there. Let's see if we can actually share this. But this is a hilarious one. So we were like, "Oh man, we really want to be the first to put this out." And we were like, "Okay, we're one week out. The model's almost done training." We're like, "Should we build a web server?" And then we're like, "Oh yeah, maybe not." And then we ended up spinning up this whole web server so people could use it rather than just download the git repo. It's kind of annoying, especially for biologists. And we actually wanted people to use this. So let's spin up a web server. Let's get the technical report out, all this stuff. So we ended up being up for 48 hours straight, just getting the paper over the line, getting all the last things done on the web server. And then Josh was interviewing with Bloomberg TV or something that morning. And we'd been up for 48 hours straight. So Josh runs into a room to do this interview on Bloomberg TV. And I think it was like 7:00 in the morning. Everyone's in the office. We didn't want to be seen, whatever. The interviewer is like, "Oh, interesting company. Doesn't look like there are any employees here." But yeah, it was a really fun time. I think the early startup days were just super fun. So yeah, after that we set our sights on design. And really what we were thinking is we always had antibodies in mind. We thought of this as the most tractable problem. The nice thing with proteins is you have this beautiful sequence representation. And there's already a lot of research been done on how do you autoregressively generate sequences? How do you... this sequence generation problem is well studied. So we were thinking, what's a nice area to apply sequence generation to in the bio space? And it's pretty natural to do linear sequences of amino acids. So we started working on design. A unique thing about Chai is we're not... we're designing antibodies, we're an antibody company, but we don't really pigeonhole ourselves into one therapeutic area. So we tried to really tackle this problem very generally. So we were thinking, can we design many proteins? Can we design antibodies? Can we scaffold regular complexes? So really take a holistic view on how do you design proteins in general. And that eventually led to the Chai 2 model. So that was our first flagship design model, and that's where the Chai 2 paper and our bold target discovery project came in. So we designed antibodies to 50 targets for that paper. Got binders to about half of them, with I think on average around a 20% hit rate for binding.

Chai 1 简介 Introduction to Chai 1

Matt

之后就开始研究 Chai 3。那是我们最新的模型系列,不过先停一下。

And then afterwards started working on Chai 3. So that's our latest series of models, but a break there.

Host

但在我们讨论 Chai 3 之前,你能给我们讲讲,特别是对那些可能不熟悉结构预测模型的听众,这个模型是什么样的?它大体上是怎么工作的?

But before we talk about Chai 3, can you tell us, especially for listeners that may not be familiar with structure prediction models, what does the model look like? How does it work in general?

Matt

我们来看看 Chai 1。Chai 1 大致有一个分词器、一个 Transformer、一个看起来像语言模型的东西,然后还有一个看起来像图像扩散模型的东西,它们就像拼接在一起。分词器不是那种典型的词级分词器。这就像我有一个分子里的一堆原子,现在我想把它们拉进我称为“token”的东西,供我那个像语言模型的主干使用。然后它条件化这个大的扩散模型,扩散模型会输出图像,也就是某个 3D 结构。

Let's take a look at Chai 1. Chai 1 has this like roughly a tokenizer, a transformer, something that looks like a language model, and then something that kind of looks like an image diffusion model, and they're all just like stick stitched together. The tokenizer is not your typical kind of words-of-X style tokenizer. This is like I have a bunch of atoms in a molecule, and now I want to pull those into what I would call tokens for my LM-looking trunk. And then that conditions this kind of big diffusion model which will then emit the image which is some 3D structure.

Host

那么输入是原子还是氨基酸?

So is it atoms or is it amino acids that are the input?

Matt

这也是个有趣的问题。我们有所有这些不同的输入轨道。生物学的一个特点是数据在某种意义上天生是多模态的。你有这种 token 序列表示。每个 token 都有一组原子悬挂在上面。然后你还有不同原子的某些属性。原子可能有不同的电荷,可能有不同的元素类型,就像元素周期表。然后这些都被捆绑成 token。一旦分词,你就可以用非常标准的方式处理。但最终你必须回到这些 3D 坐标。所以为了预测结构,这只是一个 3D 对象,而这个对象要经过,或者说要输出这个对象,你要经过一个看起来像图像扩散模型的东西,从 token 回到原子表示。

It's an interesting question as well. So we have all these different input tracks. One thing about biology is the data is inherently multimodal in a sense. You have this token sequence representation. Each of these tokens has a set of atoms that kind of dangles off. And then you also have some properties of the different atoms. An atom might have a different charge. It might have a different element type. So like periodic table of atoms. And then these all get bunched together into tokens. Once tokenized, you can process this in very standard ways. But ultimately you have to get back to these 3D coordinates. So in order to predict the structure, this is just some 3D object, and that object goes through, or to emit that object, you go through what looks like an image diffusion model where you kind of go back from tokens back to the atom representation.

Host

我明白了。所以 token 输入,Transformer 建立不同 token 之间的关系,然后扩散模型把那个潜在表示变成 3D 结构。

I see. So the tokens go in, the transformer establishes the relationship between the different tokens, and then the diffusion model turns that latent representation into a 3D structure.

Matt

完全正确。是的。

That's exactly right. Yeah.

Host

好的,太好了。所以那是 Chai 2。那是 Chai 1。

Okay, great. So that's Chai 2. That was Chai 1.

Matt

好的。所以 Chai 1 遵循了模型。是的。就像,所有这些生物的东西听起来有点吓人,比如原子、token、氨基酸。归根结底,我个人的背景是理论计算机科学。我早年都在做这个,到博士后期才转过来。但我认为你需要的背景和你在任何其他机器学习领域需要的背景非常相似。有很多领域特定的东西要学。但我想说的一个类比或轶事是,人们认为除非你是生物学家,否则你不能做 AI 生物。但这有点像,除非你是导演,否则你不能做视频模型。有很多超级领域特定的东西,比如,哦,是的,要理解视频中的光照等等,但归根结底这些只是机器学习问题,它们都以同样的方式解决。

Okay. So Chai 1 followed the model. Yeah. It's like, and like all this bio stuff sounds kind of scary, like atoms, tokens, amino acids. At the end of the day, my background personally is theoretical computer science. That's what I spent all my earlier years doing, transitioning to this pretty late in my PhD. But I think the background that you need is really similar to the background that you need for any other field of machine learning. There are all these domain-specific things that you learn about. But one analogy or anecdote I like to say is people think you can't work on AI bio unless you're a biologist. But it's kind of like you can't work on video models unless you're a director or something. There are all these super domain-specific things like, oh yeah, to understand lighting in video and things like that, but at the end of the day these are just machine learning problems and they're all solved the same way.

Chai 2:设计能力 Chai 2: Design Capabilities

Host

好的。那么 Chai 2 在能力上有了飞跃,架构也变了,对吧?

Okay. So then Chai 2 there's a jump in capability as well as an architectural change, right?

Matt

是的。我们披露的关于 Chai 2 的是它是一个全原子扩散模型。所以我们仍然试图预测 3D 空间中的原子,但我们这样做的方式是模型实际上有能力设计原子、放置它们、决定哪些原子真的在那里。所以表示一个氨基酸(如蛋白质 token)的一种方式是通过哪些原子存在。所以在 Chai 2 的情况下,我们只是预测,好吧,让模型选择它想保留哪些原子,然后映射回有哪些氨基酸。

Yeah. What we've disclosed about Chai 2 is it is an all-atom diffusion model. So we're trying to predict atoms in 3D space still, but we're doing it in such a way that the model actually has the ability to design atoms, place them, decide which atoms actually are there. So one way to represent an amino acid like a protein token is by which atoms are present. So in the Chai 2 case, we were just predicting, all right, let the model just kind of pick what atoms it wants to keep, and then map that back to what amino acids there are.

Host

你能用 Chai 2 做哪些 Chai 1 做不到的事?只是更好,还是带来了新的能力?

What are you able to do with Chai 2 that you can't do with Chai 1? Is it just better or are there new capabilities it brings?

Matt

是设计,对吧?所以 Chai 1 让你说,嘿,我知道氨基酸序列,那个文本字符串,我知道结构。

It's design, right? So Chai 1 lets you say, hey, I know the sequence of amino acids, that text string, and I know the structure.

Host

你会从基因组得到,或者确切地说。

That you would get from the genome, or exactly.

Matt

Chai 2 说,好的,我有一个目标结构,我想设计一个结合剂。Chai 2 会生成候选分子、候选药物,它们能结合到那个目标。所以这是一种设计模型或设计模型家族。我认为那才是真正跨越有用性门槛的地方。我的意思是,Chai 1 非常有用,因为你至少可以凭直觉理解它,推理结构,看到你在看什么。但最终目标是设计药物和设计新分子。我认为 Chai 2 在一年前真的跨越了用抗体做这件事的性能门槛。

Chai 2 says, okay, I have a target structure that I want to design a binder to. Chai 2 will then generate candidate molecules, candidate medicines that bind to that target. And so this is kind of a design model or design family of models. And I think that's where you really cross the threshold of usefulness. I mean, Chai 1 is very useful because you can at least intuit it and reason about the structure and see what you're looking at. But the ultimate goal here is to design medicines and design new molecules. And I think Chai 2 really crossed the threshold of performance for doing that with antibodies a year ago.

Host

这里的一个类比可以回到图像领域。所以 Chai 1 就像,你知道,这张图像里有一只猫,谢谢 Chai 1。而 Chai 2 就像,我给你看一个背景,也许我用一些图像信息提示你,比如,嘿,把一只猫放在田野里,Chai 2 实际上会给你一张猫在田野里的图像,你会说,那图像看起来不错或不好。你可能还有其他模型来对图像排序。但根本上这是生成问题。

One analogy here would be back to the image domain. So Chai 1 would be like, you know, there is a cat in this image, thanks Chai 1. And Chai 2 is like, I'll show you a background, maybe I'll prompt you with some image information like, hey, put a cat in a field, and Chai 2 actually gives you back an image of a cat in a field, and you're like, that's a good-looking image or it's not. You might have some other model which kind of ranks the image. But fundamentally it's the generative problem.

Matt

是的。所以把这个类比再推进一步,也许更像你给它看一个背景,然后它生成,有一只猫,然后它同时生成猫的图像,并且在这个领域里有一只猫是有意义的,而且猫在图像中也合理。所以这是一个有趣的问题,因为你必须同时生成两样东西,序列和结构。

Yeah. So taking that analogy a step further, it's maybe more like you show it a background and then it generates, there is a cat, and then it generates an image of the cat at the same time, and it makes sense that there is a cat in this field and also that the cat works in the image. So it's an interesting problem because you have to generate two things at the same time, both the sequence and the structure.

Host

你能不能,如果你不知道你能不能,但你能谈谈这是怎么工作的吗?你如何共同设计序列,使结构也合适且合理?

Can you, if you don't know if you can, but could you talk a bit about how that works? How do you co-design the sequence in a way that the structure also fits and makes sense?

Matt

一种思考方式是经典的做法。让我们在结构预测中讨论两者。比如,好吧,我知道序列,我可以从中大致算出 3D 形状。然后还有逆折叠问题,即给定一个 3D 形状,给我一个能折叠成这个形状的序列。现在你需要同时做这两件事。

One way to think about it is the classic way of doing this. Let's talk about both in structure prediction. Like, all right, I know the sequence and I can from that roughly figure out the 3D shape. And then there's the inverse folding problem, which is given a 3D shape, give me back a sequence that would fold into this. And now you kind of need to do both things at the same time.

迭代设计与EM算法 Iterative Design and EM Algorithm

Matt

但我认为类似的原理也适用。你可以让模型稍微思考一下这个结构应该是什么样的,然后让模型的另一部分思考什么样的序列可能支持这个结构。扩散模型的一个好处是,你可以相当缓慢且迭代地做这件事,给模型时间去思考如果结构这样改变,序列应该如何变化,然后来回调整,直到它收敛到某个自洽的结果。

But I think similar principles apply. You can have the model think a little bit about what this structure should look like, then have another part of the model think about what sequence might support this. A nice thing with diffusion is you can do this pretty slowly and iteratively, giving the model time to consider how the sequence should change if the structure changes, and play this back and forth until it converges on something self-consistent.

Host

这几乎就像 EM 算法。

It's almost like an EM algorithm.

Matt

是的,完全正确。所以你现在有了这个模型,Chi 2,它能够预测或采样一个结构以及生成该结构的序列。但仅仅因为你能生成一个结构,并不一定意味着它足够准确,可以用来做点什么。那么,你在这个基础上还有其他脚手架吗?还有额外的问题吗?你是单次生成这些,还是需要生成数千个然后进行排序或打分?有一个候选可能还不够。你采样出一个结构之后会怎么做?

Yeah, exactly. So you have this model now, Chi 2, which is able to predict or sample a structure and a sequence that generates that structure. But just because you can generate a structure doesn't necessarily mean it's accurate enough to do something. So do you have other scaffolding on top of that? Are there additional problems? Are you one-shotting these things, or do you need to generate thousands and then rank or score them? Having a candidate might not be enough. What do you do once you sample a structure?

Host

传统上,当共设计和蛋白质结构设计开始成为一件事时,我们在指标上很迷茫。你怎么知道你的蛋白质——你设计了一些序列和结构——我怎么知道这是否靠谱?我可以随便说。

Traditionally, when co-design and protein structure design started becoming a thing, we were at a loss for metrics. How do you know your protein—you design some sequence and structure—how do I know this is legit or not? I can tell you anything.

Matt

现在这完全超出领域了,对吧?从定义上讲。

It's totally out of domain now, right? By definition.

Host

是的。作为人类,你可以看着这个,心想,我不知道,但看起来没问题。甚至生物学家也会说,我不知道这东西是否真的能折叠。也许有些部分看起来是对的。甚至我们的生物学家也对我们的一些最终有效的设计感到惊讶。当时我们做的是,我们提出了一堆指标,而 AlphaFold 确实促成了这一点。你取你预测的序列,通过一个完全不同的结构预测方法——完全独立于你的模型——然后说,如果一个独立模型认为这个序列折叠成类似的结构,那么它正确的可能性就比先验可能性更高。所以你可以取你的序列,测量这个结构预测方法与实际预测的结构有多一致。你可以将你的设计与独立模型的结构预测进行比较。这成为验证你的设计模型正确性的一个非常好的方法。人们在一段时间内有点玩弄这些基准,并不断推进。事实证明,如果你的所有蛋白质看起来都一样,那么很容易获得自洽的结构设计。这会产生很多问题,但后来人们开始在此基础上不断增加更多内容。

Yeah. And as a human, you can look at this and think, I don't know, it checks out. Even biologists are like, I have no idea if this thing actually folds. Maybe some of it looks right. Even our biologists are surprised by some of our designs that do end up working. What was done at the time is we came up with a bunch of metrics, and AlphaFold really enabled this. You'd take the sequence you predicted, run it through a completely distinct structure prediction method—completely independent of your model—and say, if an independent model thinks this sequence folds to a similar structure, then it has a higher likelihood of being correct than whatever the prior likelihood would be. So you can take your sequence and measure how consistent this structure prediction method is with the structure you actually predicted for that sequence. You can compare your design to an independent model's structure prediction. That became a really good way of gaining conviction that your design model was correct. People kind of gamed these benchmarks for a while and kept pushing. It turns out it's easy to get self-consistent design of structures if all your proteins look identical. There are a lot of problems this creates, but then people started adding more and more on top of this.

Matt

是的,这是社区中一些人已经承认的一个有趣观点。那么你是怎么解决这个问题的?你可以看到,如果你同时使用你的预言机和采样器,你最终会收敛。你如何阻止这种情况,或者如何说服自己你在做有价值的事情?

Yeah, that's an interesting point that some people in the community have acknowledged. So how did you solve that? You can see that if you use your oracle and your sampler at the same time, you eventually converge. What do you do to stop that or to convince yourselves that you're doing something valuable?

Host

结构预测方法的一个好处是,通常你对模型预测的置信度有一定的校准。事实证明,这些模型可以给你一个校准良好的置信度预测。它不只是说“我认为结构是这样的”,而是会说“我认为结构是这样的,但这里有些部分我不太确定”。你可以将其聚合为一个标量。通常人们会看不仅是我有多自洽,而且这个独立模型有多喜欢它输出的结构。这是早期获得信心的一种方式。人们经常做的另一件事是看生成的多样性,因为你可能有一个完全一致的模型,给你很好的置信度预测,但可能每次都是相同的结构,每次都是相同的序列。所以你也想看看解决方案有多多样。从某种意义上说,我能解决多少这些新问题?如果我有很多钱来验证,你会怎么做?我能去做冷冻电镜或类似的事情来尝试找出结构,获得一些真实数据吗?

One of the nice things about structure prediction methods is that usually you have some calibration in how confident the model is in its prediction. It turns out these models can give you a well-calibrated confidence prediction. Rather than just saying this is what I think the structure looks like, it'll say this is what I think the structure looks like, and here are the parts I'm not really certain about. You can aggregate this down to a single scalar. Typically what people do is look at not only how self-consistent I am, but how much does this independent model even like the structure it output. That was one way early on to gain confidence. Another thing people often do is look at the diversity of their generations, because you could have a model that's perfectly consistent, gives you great confidence predictions, but might be the same structure every time, same sequence every time. So you also want to see how diverse the solutions are. How many of these new problems can I solve in a sense? If I had a lot of money to validate, how would you do that? Can I go and do cryo-EM or something like that to try to figure out the structure, get some ground truth?

Matt

更多的是反馈循环非常慢。你可以这样验证几个结构,但这可能需要几个月,而且这不是一个非常可扩展的方向。我认为这是整个领域的问题。人们花了很多时间,尤其是在 Chai,思考如何更大规模地验证这些问题。我们如何提高验证的通量或缩短周期时间?因为如果你要等几个月才能知道你的模型是否正确,那么在研究环境中很难这样迭代。好消息是情况正在好转。现在有整个湿实验室网络可以合作,他们会运行这些测定和实验,告诉你诸如你的蛋白质是否与其靶标结合等信息。幸运的是,我们不是以年为单位——我们已经缩短到几周。虽然没有 LLM 领域那么快,在那里你可以扩大评估规模,投入更多算力,几小时内就能得到结果,但已经足够快,可以开始递归地自我改进。我们也花了很多时间思考我们可以计算哪些硅基指标,这些指标可能预测实验室的成功。你问到了冷冻电镜——你必须测量结构,这非常昂贵,因为你必须冷冻蛋白质,然后向它发射电子束,看它们如何反弹。我记得有一个非常有趣的小故事。看看我能不能分享。

It's more that the feedback loop is really slow. You can validate a few structures like this, but it might take months, and it's not a very scalable direction. I think that's a problem for the field as a whole. People are spending a lot of time, especially at Chai, thinking about how to validate these problems at a bigger scale. How do we increase the throughput of our validation or increase the cycle time? Because if you're waiting months to figure out if your model was correct, it's hard to iterate in a research environment that way. The good news is this is getting a lot better. There's a whole network of wet labs you can work with that will run these assays and experiments and tell you things like whether your protein binds to its target. Thankfully we're not at years—we're down to weeks. Not as fast as LLM land where you can scale up an eval and throw more compute and get results back in hours, but fast enough to start recursively self-improving. We also spend a lot of time figuring out what metrics we can compute in silico that are predictive of lab success. You asked about cryo-EM—you have to measure the structure, and that's so expensive because you have to freeze the protein and shoot electron beams at it and see how they bounce off. I remember there's this really funny anecdote. We'll see if I can share it.

验证结构预测 Validating structure predictions

Matt

但就像,你知道,Chai 的论文里,我们确实这么做了。我们拿模型预测的一些蛋白质去跑冷冻电镜,拿到结果后我们想,等等,结果看起来不对,因为我们已经把预测叠在电子云、点云上,却看不出任何差别。重点是,我们现在已经到了这样的阶段:这些结构预测模型与验证出的真实原子位置只差几个埃甚至更少。

But like, you know, the paper in Chai, we actually did that. We took some of the proteins that the model predicted and ran Cryo-EM, and we got the results back and we're like, wait, the results look wrong because we had overlaid the prediction over the electron cloud, the point cloud, and we didn't see any difference. The point is, we're getting to the point now where these structure prediction models are within a few angstroms or less of the actual atomic positions that you validate.

Host

而这次误差是 0.33 埃,相当于原子宽度的三分之一。

And in this case it was a 0.33 angstrom error, which is 1/3 the width of an atom.

Matt

我们当时想,这根本不可能是对的。显然他们只是把错误的设计发回给我们了。

And we were like, this can't even be right. Like clearly they just sent us back the wrong design.

Host

他们只是把我们的设计发回来了。

They just sent us back our design.

Matt

对,就是这样。

Yeah. Exactly. Yeah.

Host

你们检查过数据泄漏吗?

Did you check for data leakage?

Matt

呃,是的,这次没有。我们特意选了这些靶点,确保没有已知的抗体结合剂。所以如果我们真的命中了,那绝对是这个靶点的第一个抗体命中。对。

Uh yeah, in this case there were no, so we actually chose these targets specifically to have no known antibody binder. So if we did get a hit, it was definitely the first antibody hit to this target. Yeah.

Host

我觉得生物学有一点是我以前没意识到的,就是有多少工作真的像是在黑暗中摸索,这甚至不是比喻。你根本看不到这些东西长什么样,对吧?所以结构模型如此重要,因为现在你可以在原子级别预测这些东西的样子,这让你能进一步做像 Chai 2 那样的设计模型。

I think that's one of the things I didn't realize about biology was like just how much of it is literally feeling around in the dark, and that's not even a metaphor. You literally can't see how these things look, right? So structure models are so huge because now you can actually predict within an atom how these things look, and that enables you to then do things like Chai 2 with the design models.

Matt

在我看来,这就是 AI 用于科学的核心问题之一,对吧?在很多情况下,你根本不知道如何衡量你的问题。所以很难验证。

This to me is AI for science, one of the cornerstone problems, right? You don't know—you fundamentally don't even know how to measure your problem in a lot of cases. So it's very difficult to validate.

Host

对。所以你们用 Chai 2 得到了亚埃级预测。Chai 3,为什么是 Chai 3?有什么改进?

Yeah. So you're getting these sub-angstrom predictions with Chai 2. Chai 3, why Chai 3? What's better or what?

为何推出Chai 3 Why Chai 3

Matt

对,我觉得对于 Chai 3,说实话,有 Chai 2、Chai 2.5、Chai 2.7,最后还有 Chai 3,每次我们都看到性能越来越好。Chai 3 的主要问题是,我们看 Chai 2 和它能解决的靶点。Chai 2 之后有很多内部讨论:嘿,我们成功为这 50 个靶点中的一半做出了分子结合剂。另外 25 个呢?我们能做些什么让它们更好?然后我们产生了分歧:我们是该研究这些错过的靶点,弄清楚它们有没有我们可以关注的特性,还是干脆押注模型——如果我们投入更多时间,减少经验驱动,真正押注模型会变得更好,模型会不会自己达到目标?我们明确选择了后者。我们押注模型会变得更好,并在这条路上全力以赴。

Yeah, I think with Chai 3, honestly, there was a Chai 2, there was a Chai 2 and a half, there was a Chai 2.7, there was eventually a Chai 3, and each time we saw better and better performance. The main thing with Chai 3 is we looked at Chai 2 and the targets it could solve. There was a lot of internal discussion after Chai 2: hey, we made successful molecule binders to half of these 50 targets. What about the other 25? What can we do to make those better? And then we were split: should we study these targets that we missed and figure out exactly if there are properties of these that we can look at, or should we just bet on the models—will the models just get there if we put more time into being less empirically driven and really bet on the models getting better? And we definitely took the latter approach. We bet on the models getting better and just pushed as hard as we could on that front.

Host

扩大模型、数据等,来构建更准确的模型。

Scaling up the model, the data, whatever, to build more accurate models.

Matt

对。

Yeah.

Host

是准确性吗?这是主要的吗?是结合亲和力吗?我们——

Is it accuracy? Is that the main thing? Is it binding affinity? What do we—

Matt

所以,我认为结合亲和力是一个大问题。你不能只是弱结合。为了让这成为一个有用的工具,尤其是对我们的合作伙伴,我们需要开始生产达到或非常接近治疗级别的分子,这意味着它们必须结合得非常紧密。它们还必须具有可开发性。它们必须具有所有这些良好的治疗特性。

So, I think binding affinity is a big one. You can't just bind weakly. In order for this to be a useful tool, especially for our partners, we need to start producing molecules that are at or very close to therapeutic grade, which means they have to bind really tight. They also have to be developable. They have to have all these nice therapeutic properties.

Host

还有可开发性——我想我们谈过,他提到了 Chai 2.5,对吧?我们在 Chai 2 发布几个月后发布了它。我们对分子的可开发性做了一项研究。对于听众来说,显然分子必须结合得好且紧密,但还有其他你关心的特性。用非生物学术语来说:它安全吗?稳定吗?容易制造吗?会自聚集吗?我们惊喜地发现,在这些领域我们能提升和推动性能的程度。

And developability—I think we talked about, he mentioned Chai 2.5, right? Which we released a few months after Chai 2. There was a study we did on the developability of the molecule. For the audience, obviously the molecule has to stick well and stick tightly, but there are these other properties you care about. To use non-biological terms: is it safe, is it stable, is it easy to manufacture, does it self-aggregate? And we've been pleasantly surprised at how much we've been able to climb and push the performance in those areas.

Host

看起来你想做抗体的原因之一就是可开发性。

It seems like one of the reasons that you want to do antibodies is the developability.

Matt

对,用抗体框架,你免费获得了很多东西,对吧?

Yeah, you get a lot for free there, right, with that antibody framework.

Host

对,这很有意思。我的意思是,对我来说,市面上有很多结构预测模型。我觉得是这些其他辅助因素,可能对产品的实用性影响最大。

Yeah, it's interesting. I mean, to me, there are many structure prediction models out there. I feel like it's these other ancillary factors that are probably going to be the most impactful in the usefulness of a product, let's say.

Matt

对,没错。

Yeah. Right.

Host

对,绝对。结构预测的好处是有一个可以对比的基准真相。对于设计,你并没有真正拥有这个。你会说,‘这里有个新的疾病分子。给我一个结合剂。’如果你想知道这东西是否真的结合,你必须把它送到实验室等一段时间。对于结构预测,你可以说,‘好吧,模型以前没见过这个序列。它从没见过任何接近的东西。它真的能折叠成正确的形状吗?’我们可以把它从数据集中拿出来检查。所以我一直认为结构预测是一个很好的速通式基准,用来验证想法。

Yeah. Absolutely. The nice thing about structure prediction is there is a ground truth that you can compare against. For design, you don't really have that. You're like, 'Here's some new disease molecule. Give me a binder for that.' And if you want to know if this thing really binds, you have to send it off to the lab and wait a while. For structure prediction, you can be like, 'All right, the model hasn't seen this sequence before. It's never seen anything close. Does it actually fold up into the correct shape?' And we can just kind of hold that out of the data set and check. So I think I've always thought of structure prediction as this really nice speedrun kind of benchmark to validate ideas on.

Matt

对。抱歉,我不是那个意思——我指的是,你知道,一般的结构模型。对。但没错,正是如此。所以也许我们可以多谈谈进入产品方面的事情。感谢你的到来。我实际上——我的意思是,就像我说的,我真的认为这不仅适用于这类结构模型,也适用于虚拟细胞等等。真正影响最大的是药物开发过程周围的所有其他东西。所以你能谈谈这个吗?

Right. Sorry, I didn't mean to say—I meant, you know, sort of structural models in general. Yeah. But yes, exactly. So maybe we can talk a little bit more about getting into the product side of things. Thank you for coming. I actually—I mean, like I said, I really think this goes throughout not only for, you know, sort of structural models like this but also virtual cell and whatever. It's really all the other stuff around the drug development process that is going to have the biggest impact. So can you talk a little bit about that?

Chai 2后的产品侧 Product side after Chai 2

Matt

对,我觉得在 Chai 2 之后谈这个其实是件好事,因为我认为从产品角度来看,Chai 2 是开始变得真正有趣的地方,对吧?有了 Chai 2,我们跨越了实用性的门槛。在我们发表那篇论文后,很多制药公司和生物技术公司来找我们,说,‘嘿,这个模型可能能为我们做些事情。我们能使用它吗?’然后我们想,‘哦,天哪,我们应该构建一个产品,对吧?我们应该构建一些东西让你能使用那个模型。’那正好是我加入的时候,然后就有了一场疯狂的构建,既要构建产品——我们可以谈谈它的形态——也要去确保算力,这样我们才能把那些模型提供给我们的合作伙伴。

Yeah, I think that's actually a good thing to talk about after Chai 2, because I think Chai 2 is where it started to get really fun from a product perspective, right? With Chai 2, we crossed the threshold of usefulness. After we released that paper, we had a lot of pharmas and biotechs approach us and say, 'Hey, this model might be able to do some stuff for us. Can we use it?' And then we're like, 'Oh man, we should build a product, right? We should build something to let you use that model.' And that's right around when I joined, and there was this mad buildout to both build the product—which we can talk about the shape of—and also go and secure the compute, actually, so we can go and serve those models to our partners.

制药平台的安全与IP Security and IP in pharma platform

Matt

我认为另一个非常有趣的第三点是关于安全和知识产权。我们希望成为一个非常中立的平台,任何人都可以在上面设计药物,但正如你们所知,制药行业是出了名的对知识产权敏感的行业。我加入时,很多人告诉我这不可能——他们不会把数据放在一个平台上,然后让所有新药都从里面生成。有一点安全背景帮了忙。如果你在数据分段和单租户设置上非常激进,几乎是为每个客户部署一个单独版本或单独账户,你实际上可以构建一个平台,然后交付给他们。去年夏天,我们开始这么做。我们一直在和礼来谈,他们是最早与我们密切合作的合作伙伴之一,共同打造了那个设计套件的 V1 版本,你可以用它来设计一些分子。也许值得谈谈那个设计套件。

And I think another third piece there that was really interesting is around security and IP. I think we want to be a very neutral platform that anyone can design medicines on, but as you guys know, pharma is this notoriously IP-sensitive industry. When I joined, a lot of people told me this can't be done—they're not going to put their data in a platform and have all their new medicines generated out of it. Having a bit of a background in security helped. If you're really aggressive about how you segment data and set up single tenancy, where you're almost deploying a separate version or separate account in the product per customer, you can actually build a platform and then go and ship it to them. Through the summer of last year, we started doing that. We'd been talking to Eli Lilly, and they were one of the first partners to really work with us closely on that, making the V1 of that design suite that you can use to engineer some of those molecules on. Maybe it's worth talking a bit about that design suite.

设计套件:分子工程可视化 Design suite: visual tool for molecule engineering

Matt

我认为我们现在拥有非常强大的模型,只要以正确的方式设定条件,它们就能做各种疯狂的事情。如果你给它们正确的上下文,关于你追求的结构或模型的约束——比如,嘿,我想设计一个抗体,能命中这个 GPCR 蛋白,但不与细胞膜碰撞,同时也针对那个特定的表位。我们看了看,想也许我们可以给它加个聊天机器人界面,那样对话起来很容易,但我们想构建的是某种非常视觉化的东西。而借助一些结构预测模型,你终于可以构建出真正视觉化的东西。如果你看看 Chai 的产品,它看起来更像 Autodesk、SolidWorks 或 Figma,而不是 ChatGPT。你可以加载你的分子。这几乎是一个类似 Photoshop 的设计套件。你有相当于画笔工具的东西来绘制你的表位。有一个相当于内容感知填充工具的东西,可以从 Chai 生成你的结合剂。你还有很多科学分析和绘图功能来理解模型的结果。但我们惊讶的是,做这件事本身就有很多复杂性,这样你在提示这些模型给你建议时就不会搬起石头砸自己的脚。

I think we have these really powerful models now that can do all of these crazy things if you condition them in the right way. If you give them the right context about the structure you're going after or the constraints around the model—like, hey, I want to design an antibody that hits this GPCR protein, but doesn't collide with the cell membrane and also targets the specific epitope on that as well. We looked at it and thought, I guess we could put a chatbot around it, which would be really easy to talk to, but we're trying to build something almost very visual. And you can finally build something really visual with some of these structure prediction models. If you look at the Chai product, it looks a lot less like a ChatGPT and a lot more like Autodesk, SolidWorks, or Figma. You can load up your molecule. There's this almost Photoshop-esque design suite. You have the equivalent of a paint tool to paint your epitope. There's an equivalent of a content-aware fill tool to get your binders generated from Chai. You also have a lot of the scientific analysis and plotting to understand the results of the models. But we've been surprised at how much complexity is actually just in doing that, so you don't shoot yourself in the foot when you're prompting these models to give you advice.

说服药物化学家使用AI工具 Convincing med chemists to use AI tools

Host

那么,你是和那些设计抗体的人坐在一起,然后他们向你抱怨之类的吗?你如何说服药物化学家使用你们的工具?因为药物化学家讨厌 AI 工具——出了名的,比如,我不想碰这个东西,或者我不理解它,他们不会碰他们不理解的东西。

So, are you sitting with people who are designing these antibodies, and then they're complaining to you or whatever? How do you convince med chemists to use your tools? Because med chemists hate AI tools—notoriously, like, I don't want to touch this thing, or I don't understand it, and they will not touch things which they do not understand.

Matt

嗯,模型运行得非常好,这很有帮助。当我们有了 Chai 2 和 Chai 2.5 的结果时,我认为这已经足够激活能量,让制药公司和这些公司里的科学家们说,哦,让我们试试吧。实际上,Chai 能不能试着针对几个目标运行模型,让我们看看结果?然后我们做了,结果很好,他们就说,好吧,让我试试那个产品,让我用用看。我认为制药行业实际上非常务实。到目前为止,我对我们合作过的每个人都印象深刻。他们对此非常务实,而且他们愿意被证明是错的。我实际上并不怪他们不信任模型。我用过这些模型,他们怀疑是对的。我看到新版本时总是相当怀疑——我一直如此。所以你确实只需要向他们展示证据。他们可以给你一个他们感兴趣的目标,或者可能是他们过去研究过的东西。他们可能不想一开始就分享知识产权,但他们可以说,嘿,我过去在这个特定目标上遇到过麻烦。让我们看看你们在这个上能做得怎么样。一旦你向他们展示了证据,他们几乎压倒性地愿意接受。

Well, it helps a lot to have the models working really well. When we had the results of Chai 2 and Chai 2.5, I think that's enough activation energy where pharma companies and the scientists within these companies are like, oh, let's try it. Actually, can Chai just try running the model against a few of these targets and let's look at the results? Then we do that, and the results are good, and they're like, okay, let me try to get on that product and use it. I think pharma is incredibly pragmatic, actually. I've been very impressed with everyone we've worked with so far. They're very pragmatic about this, and they're willing to be proven wrong. I actually don't blame them for not trusting the models. I have used these models, and they're right to be skeptical. I am pretty skeptical when I see a new release—I always have been. So you really just need to show them the proof. They can give you a target they're interested in, or maybe something they've worked on in the past. They probably don't want to share IP right out of the gate, but they can be like, hey, I've had trouble with this particular target in the past. Let's see how you guys can do on this. Once you show them the proof, they almost overwhelmingly are willing to accept that.

Matt

我来自网络安全背景,之前也做过安全产品。那些年是黑暗的岁月,因为你花了很多时间向那些出奇地不技术的人销售。你以为网络安全人员非常技术化,但在很多情况下他们不是,而且这是向非常不成熟的客户进行艰难的企业销售。我一直很惊喜地发现,我是多么喜欢与我们的合作伙伴和客户合作。这些科学家一生中花了 5 年、10 年、20 年研究一个目标,通常在某些情况下,他们研究了关于它的一切。他们非常老练,非常聪明。与他们合作是一座金矿。我们学到了很多关于如何改进产品的知识。有一个轶事:几个月前,我们展示了一个与制药合作伙伴合作的目标的一些结果,房间里的一位科学家开始流泪哭泣。

I come from a cybersecurity background and have worked on security products before. Those were dark years, because you spend a lot of your time selling to people who are surprisingly not that technical. You think cybersecurity people are very technical, but in many cases they're not, and it's this uphill enterprise slog to a very unsophisticated customer. I've been pleasantly surprised by how much I enjoy working with our partners and customers. These are scientists who have been spending 5, 10, 20 years of their life working on one target, often in some cases, and they've studied everything about it. They're very sophisticated and very smart. Getting to collaborate with them is a gold mine. We learn a lot about how to make the product better. There's this anecdote: a few months ago, we were showing some results from a target with a pharma partnership, and one of the scientists in the room started tearing up and crying.

Host

哦,哇。

Oh wow.

Matt

你真是说到点子上了。她当时,我们问,怎么了?她说,没什么,我只是真的花了 10 年试图为这个东西找到一个初始结合剂,而你们帮助我做到了。

You really hit the nail on the head with that one. She was like, we were like, what's wrong? She's like, no, I've just literally spent 10 years trying to get an initial binder to this thing, and you guys were able to help me do it.

Host

哦,那真是……

Oh, that's...

Matt

回答你的问题,那感觉真的很特别。

And that feels really special, to answer your question.

内部跨学科团队与严谨性 Internal cross-disciplinary team and rigor

Matt

你知道,在每个合作项目里,我们当然有科学家和计算生物学家团队。我们楼里也有自己人。我觉得我非常欣赏 Chai 的一点就是它的跨学科性。我们有工程专家、像我这样生物背景不那么强的人,我们有很棒的 AI 科学家或机器学习科学家,但我们还有一批科学家,他们加入进来帮我们测试模型的极限,看看 Chai 2 到底能做什么、哪些靶点能做、哪些不能,并为研究方向提供信息。

You know, there are, of course, the teams of scientists and computational biologists that we're working with within each of our partnerships. There's also the people we have within the building. I think one of the things that I really appreciate about Chai is how cross-disciplinary it is. We have people who are engineering experts and less bio experts like myself, we have great AI scientists or ML scientists, but we also have a bunch of scientists that we work with and that have joined to help us test the limits of the models, see what Chai 2 is actually capable of, what targets it can do, what it can't, and inform some of the research direction.

Host

我想补充一点。在 Chai 2 时代,我们一开始是一群工程师和有 AI 生物背景的人,没有硬核的实验室科学家。我们在那个领域最早招的人之一是 Nathan Rollins,我记得他 14 岁就在 Baker 实验室工作,18 岁从哈佛毕业,21 岁左右在 Marks 实验室拿到博士学位。他一开始对 Chai 非常怀疑,后来结果出来了,他说,这有点意思,可能行得通。等 Chai 2 的结果出来后,他说,我得让它无懈可击,先别庆祝。所以我觉得能有这种严谨性真的很好,有那些在实验室待过、自己设计过蛋白质的人,还有像 Andy 那样领导过多个治疗项目、亲自把药物推进到临床的人。我们内部有这些人,他们就在用产品,真正在实战中检验它。

I want to add to that. In the Chai 2 days, we kind of started with a bunch of engineers and people who had AI bio experience. We didn't have a hardcore lab scientist. One of our first hires in that realm was Nathan Rollins, who I think started working in the Baker lab at 14, graduated from Harvard at like 18, and got his PhD by like 21 or something, in the Marks lab. He was super skeptical about Chai at first, and then the results started to come in, and he was like, okay, this is kind of interesting, this could work. Then once the Chai 2 results came back, he was like, I need to bulletproof this, nobody celebrate yet. So I think it's been really nice to have that level of rigor, to have people who have spent time in the lab, designed proteins themselves, and in the case of Andy, led several therapeutic programs and brought drugs to the clinic themselves. We have all these people internally at Chai just using the product and really battle testing it.

Host

如果你没有自己的平台,对吧?我的意思是,你没有自己的项目,对吧?你是纯粹的平台或合作模式,对吧?

If you don't have your own platforms, right? I mean, so you don't have your own programs, right? You're pure platform or partnership model, right?

Matt

对。

Yeah.

Host

如果你基本上没有一个需要不断推进的用例,你怎么去实战检验?或者如果你只是往前推,成功的话最终会得到自己的候选分子,那你怎么处理?

How do you battle test something if you basically don't have a use case where you have to continuously push it forward? Or if you are just pushing things forward, when you just end up with your own candidates if you're successful, then what do you do about that?

Matt

我们有内部案例的基准。有一组靶点已经有已知的治疗药物,还有一组靶点是我们选来挑战自己的。我们不断优化这个集合,往里加东西。内部科学团队做的就是扩展它,运行实验来获得初始结合剂。我们并不在乎去开发那些药物,我们做这些只是为了验证和让模型更好。

We have benchmarks of our own internal cases. There's a set of targets that have known therapeutics against them, and there's a set of targets that we pick to push ourselves. We're constantly refining that set and adding to it. That's what the internal science team helps with: expanding that and running experiments to try to get initial binders. We don't care about going and developing those drugs; we just do that in service of validating and making our models better.

Host

当然,和合作伙伴之间也有一个循环。

And then of course there's a loop with our partners too.

Host

你会把自己定位为命中发现,还是,用行话说,命中到先导优化?你处在哪个环节?命中发现可能是你能做的一部分,但后面的部分往往更定制化、更特殊。我的意思是,你怎么平衡?对我来说,解决通用的先导优化似乎比解决发现难得多。

Would you consider yourself hit discovery, or are you, I guess using some jargon, hit to lead optimization? Where do you live in this? Hit discovery might be one part of it, which you can do, but the later parts are often much more bespoke and special. I mean, how do you balance that? It seems much more difficult to me to solve general lead optimization than to solve discovery.

Matt

我认为理想情况下,我们真的希望能够不把这看作一系列阶段。我们之所以这么想,部分原因是初始分子通常不足以成为药物。我们现在正处在拐点。我们在 Chai 内部真的看到,模型已经非常接近能产生最终可能成为药物、或者非常接近药物的分子。所以我们尽量不在命中发现、先导优化、临床前管线的各个部分之间做太多区分。我们的北极星是直接从模型中产出类药分子。当然这很难,会有很多障碍。你需要能够提示模型做到这一点,你需要整个强化学习栈来学习不同的性质,诸如此类。但我认为这是完全可以实现的。

I think ideally we really want to be able to, rather than think of this as a bunch of stages. Part of the reason why we think of it that way is because the initial molecules are usually not good enough to be drugs. We're kind of at the inflection point now. We're really seeing this internally at Chai where the models are getting pretty close to producing molecules that could eventually be, or are very close to, drugs. So we try not to make too much of a distinction between hit discovery, lead optimization, all the different parts of this pre-clinical pipeline. Our north star is to really produce drug-like molecules straight out of the models. Of course this is going to be hard, and there are going to be tons of roadblocks. You need to be able to prompt the model to do this, you need the whole RL stack to learn different properties, things along those lines. But I think it's very achievable.

Host

我想补充一点,靶点发现、命中发现和优化,每个阶段都有门槛,需要几个月到几年,这是非常瀑布式的模型,早期尝试和获得东西的成本非常高。但我想马特的意思是,如果你开始进入一个模型能给你真正有希望的候选分子的状态,你就可以让它看起来更像一个循环。这类似于软件开发中变得更敏捷。在内部,我们有两个北极星。乍一看,它们几乎听起来矛盾。研究方面的北极星是开始从头单次生成越来越好、尽可能接近下一阶段就绪的候选药物。但在产品方面,我们也确实想扩展到这些迭代工作流,比如我得到一个结合剂,从实验室得到一些结果,用它来调节我的下一轮模型。我认为它们听起来矛盾,但实际上并不矛盾,因为会发生的是,研究将更好地为特定类别的药物识别从头候选,比如拮抗剂,阻断东西,可能更容易一点。我们可以达到一个状态,在那里我们可以单次生成相当好的药物。但下一个问题是激动剂,如何可靠地单次命中细胞上的开关?或者特异性或 ADC。随着模型变得更好,我们将不得不和产品一起攀登这些抽象层次。几个月前我变得非常存在主义的一件事是,天哪,我们在产品中构建的所有这些可视化分子之类的东西,也许当马特发布 Chai 4 时,我不得不把它们全部扔掉。但我认为这就是现在构建产品的现实。你实际上越来越少把它们当作目的本身。

I think to add to that, this notion of target discovery, hit discovery, and optimization, where each of these has a gate and takes a few months to a few years, is this very waterfall model where the cost of trying things and getting things early is very expensive. But I think to what Matt's saying, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop. It's akin to becoming more agile in software development. Internally, we kind of have two north stars. At first pass, they almost sound contradictory. The north star in research is to start to de novo one-shot better and better medicinal candidates that are as close to being ready for the next phase as possible. But also within product, we do want to sort of expand into whatever these iterative workflows look like, where maybe I get a binder, I get some results from the lab, I'm using that to condition my next run of the model. I think they sound contradictory, but they're actually not, because what's going to happen is research is going to get better at identifying a de novo candidate for a specific class of drugs, say like antagonists, blocking things, a little bit easier maybe. We can get to a state where we can one-shot pretty good drugs there. But now the next problem is like agonists, how do you reliably one-shot hitting a switch on a cell? Or by specifics or ADCs. There's kind of this levels of abstraction that we're going to have to climb with the product as the models get better. One of the things I got very existential about a few months ago was, man, all this stuff we're building in the product to visualize molecules and do this, maybe I'm just going to have to throw it all away when Matt ships Chai 4. But I think that's kind of the reality of building products now. You're actually using them less as an end in and of itself.

产品演进与抽象化 Product Evolution and Abstraction

Matt

就像你可能曾经构建过预期能持续 20 年的软件,现在它可能只预期持续一年,但它是传递价值、赋能研究、进而带你走向下一步的桥梁。所以我猜想我们可能会在越来越高的抽象层级上重写我们的产品,对吧?就像现在,我们可能有一些更接近 Cursor 的东西,你以检查代码的方式检查分子,因为你确实需要验证正在形成的键和你得到的东西的属性。但然后,你知道,你会到达一个点,那些东西已经解决得足够好,产品实际上只是在帮你编排这些假设的战役,对吧?或者你可能有一个靶点,你在编排一系列不同的表位选择来对抗它。然后你可能再上升一个抽象层级,现在你在对整个通路中的所有靶点进行一场完整的战役,对吧?嗯,我认为真正令人兴奋的是,如果你有这些非常好的结构预测、结合和设计的原语,并且你能组合它们,那么你就可以开始进入科学的外循环,对吧?然后,你知道,也许事情会自行运转,最终你会得到一些非常非常非常酷的药物。

Like maybe you'd have built software that was supposed to last like 20 years. Now it's supposed to last maybe one year, but it is the bridge to deliver value and kind of enable the research that then gets you to the next thing. And so I'd imagine we're probably going to rewrite our product at higher and higher levels of abstraction, right? Like maybe like right now we have something a little bit more akin to cursor where you're, you know, inspecting the molecule in the same way you're inspecting the code because you really need to verify like the bonds that are forming and the properties of the things that you're getting. But then, you know, you get to a point where that stuff is solved enough where now the product is actually just helping you orchestrate these like campaigns of hypotheses, right? Or maybe you have like one target and you're like orchestrating a bunch of different epitope choices or whatever against that. And then maybe you're going up one level of abstraction where you're now doing a whole campaign against all of the targets within a pathway, right? Um, and uh, I think what's really exciting about that is if you have like these really good primitives for structure prediction and binding and design and you can kind of compose them, then you can start to just like grow into like the outer loop of science, right? And then, you know, maybe the thing runs itself and uh, you start to really get to some really, really, really cool drugs at the end of it.

Host

我其实想深入探讨你刚才说的表位预测,因为我认为领域内很多人会认为这可能比寻找抗体和结合剂难得多。你认为整体上,以及就 Chai 而言,表位预测的最先进水平在哪里?这个问题是否有合理可解决的时间范围?哦,还有,也许你能定义一下表位预测吗?

I actually want to push on what you just said about epitope prediction because I think a lot of people in the field would argue this might be the much harder problem than finding antibodies and binders. Where do you think that the state-of-the-art is in general and also with regards to chai in terms of epitope prediction and like is this a problem which has a reasonable solvable time horizon? Oh and also maybe can you define epitope prediction?

Matt

我会从几个不同的层面来思考。最基本的层面是,好吧,我有一个想靶向的疾病,哪些蛋白质真正负责。就像真正从生物学上弄清楚发生了什么,比如我应该首先用药物靶向什么。我猜一旦你弄清楚了,那在某种程度上就成了结构生物学问题。你会说:“好吧,这组蛋白质是负责的,那里发生了什么?”嗯,这个蛋白质正在与某个它不应该相互作用的蛋白质相互作用。传统上,你只想阻断那个相互作用或抗体中的某些东西。但那正是这些蛋白质相互作用的地方,以及你想要破坏的相互作用类型。那通常就是表位。它就像你想要的蛋白质上实际阻断的位点。这是一个极其困难的问题。嗯,我同意你的看法。这是更难的问题。嗯,就像你需要的上下文量,以及你需要获得的全局理解,才能真正弄清楚什么在相互作用以及如何相互作用。

I'll think of this at like some different levels. So the most basic level is okay, I have some disease that I want to target and what proteins are actually responsible there. Like actually figuring out biologically what's going on, like what should I be targeting in the first place with the drug? I guess once you figure that out, um it's kind of like a structural biology problem at that point. You're like, "All right, this like set of proteins is responsible and like what's going on there?" Well, this is interacting with some other protein that it shouldn't be interacting with. And conventionally, you'd just like want to block that interaction or something within anybody. Uh, but that's kind of where these proteins interact and like the type of interactions that you want to disrupt. That's typically like the epitope. It's like the actual site on the protein that you want to block. This is a ridiculously hard problem. Uh, I'm with you on this. This is like the harder problem. Uh, like just the amount of context that you need and like the global understanding that you need in order to like actually figure out what's interacting and how.

Host

但也许让我们看几个具体案例。让我们想想,比如 SARS 病毒 3 出现,或者新型流感之类的。你会怎么做?我的意思是,你认为在这种情况下你真的能合理应对吗?

But maybe let's take a few specific cases. Let's think about what about um SARS KV3 comes around or the new flu or whatever. What would you do there? I mean is that something that you think you could actually reasonably tackle in that case?

Matt

就像是的,你可以运行一个结构预测模型,看看模型认为这个东西会结合在哪里。如果它对此非常自信,你可能会说,好吧,这就是我们想要阻断的位点。我认为总体上仍然非常困难,即使结构预测已经变得非常好,很多人认为 AlphaFold 2 解决了结构预测。其实不然。就像 AlphaFold 2 得到了,我认为 11%,多聚体版本在抗体抗原预测案例中只有 11% 的正确率。这意味着 90% 的时候它是错的。

Like yeah, you could just run a structure prediction model maybe and like see where the model thinks this thing will bind. Uh if it's highly confident in that, you might say okay here is like the site that we want to block. I think in general still very hard and even like structure prediction it's getting really good and like a lot of people think AlphaFold 2 like solves structure prediction. Not really. Like AlphaFold 2 got like I think 11% the multimer version of this got like 11% of antibody antigen prediction cases correct. That means 90% of the time it's wrong.

Host

是的。我的意思是,但 AlphaFold 2 用 MSA 解决了一类单体蛋白质。是的。是的。所以我的意思是,我认为 MSA 可能是这里的关键点,因为 MSA 在某种程度上是让一切运转的魔法。它就像一个模板,在某种意义上关于结构应该是什么,而抗体在进化上几乎不可能有模板,对吧?每个人都必须有独特的抗体,适应他们一生中经历过的事物。

Yeah. I mean but but uh AlphaFold 2 solved a certain class of monomeric proteins with MSA. Yeah. Yeah. So I mean the and that's the MSA I think might be the key point here because MSAs are sort of the the the magic which makes it all work. It's like a it's a template in some sense about like what the structure should be and antibodies almost evolutionarily can't have a template, right? Everyone has to have unique antibodies accustomed to the things that they've experienced over the course of their life.

Matt

是的。所以,对。而且为了澄清,我自己也得理解这一点,所以也许我能帮助那些不熟悉的听众。抗体,抗体的全部意义在于它能识别身体以前没有遇到过的新事物。所以抗体的设计,与其他类型的蛋白质不同,系统被设计成你可以快速重组它的不同组件,以匹配来自未知病原体的蛋白质。所以这就是为什么它在进化中不像其他蛋白质那样保守。

Yeah. So, right. And just to clarify, I had to understand this myself so maybe I can help the listeners who aren't familiar. An antibody the whole point of an antibody is it can identify new things that it hasn't the body hasn't encountered before. So the design of antibodies as opposed to other types of proteins is to the system is designed so that you can quickly recombine different components of it in order to match proteins that are from unknown pathogens more or less. And so this is why it's not conserved in evolution the way that other parts other proteins are.

Matt

是的。所以回到表位预测问题。嗯,我认为它仍然很难。我认为有很多案例可能是可处理的,但我认为总体上,如果你想为一个新靶点发现这个,仍然是一个非常困难的问题。也许虚拟细胞会是那里最接近最先进水平的东西,但那仍然还有一段路要走。

Yeah. So so like back to the epitope prediction problem. Um I I think it's still hard. I think like there there are a lot of cases that that are maybe tractable, but I think in general like if you want to discover this for a new target, uh still still a really difficult problem. Maybe virtual cell would be like the closest thing to state-of-the-art there, but that's still still a ways out.

Host

我想稍微深入一下产品,因为我对目前所有结构相关事情的经济学有些不明白。显然很多人认为它非常非常有价值。所以,你知道,我可能没完全理解,但当你看到开发一个抗体的成本,你知道,可能就几百万美元,对吧?当你从你以某种方式识别了一个靶点,然后你说,好吧,我需要一个抗体来匹配这个,然后我必须用各种方式优化它,然后也许我尝试它,我的意思是对于抗体,你通常更快地进入动物实验。如果你看看,如果你有先见之明,选对了靶点和技术,把药物一直带到上市要花多少钱,可能是 5 亿美元,通常那个 26 亿美元的数字是把所有失败也算进去的。所以如果你只看那一次成功的成本,取决于疾病可能更少,但你知道,5 亿美元可能是一个不错的中位数。所以你在一个 5 亿美元的战役中节省了大约几百万美元。

I wanted to dig in a little bit on the product cuz I there's something I don't understand about the economics of basically all all the structural stuff that's happening right now. And obviously a lot of people think it's very very valuable. So there's, you know, I'm not grokking something, but when you look at the cost of developing an antibody, you know, it maybe is a couple million dollars, right? When you go from you you you've identified a target somehow and then you say, okay, I need an antibody to match this and then I have to sort of optimize it in various ways and then maybe I try it in I mean with antibodies, you go to animal typically faster. If you look at how much does it cost to bring if you are prescient and pick the right target and the right technology to get all the way to drug it might be half a billion typically that $2.6 billion number is advertised over all the failures as well. So if you look at just the cost of that one success depending on the the disease maybe less but you know half a billion might be a good median number or something. So you're you're saving like a couple million dollars in a half billion dollar campaign.

抗体能力与新模态 Antibody Capabilities and New Modalities

Host

那为什么这这么有吸引力?

So why is this so attractive?

Matt

我可能会稍微挑战一下这个前提,从几个方面来说,对吧?比如,当然,如果你只是想为一种非常简单的靶点获取抗体,那也许可以,对吧?但我认为最让我们兴奋的是,我们的合作伙伴以更复杂的方式使用抗体,对吧?比如,在 Chai 2 中,我们展示了 GPCR 激动剂活性,对吧?你可以非常精确地按响细胞门铃蛋白上的开关,可以这么说。

I would maybe challenge the premise a bit, in a few ways, right? Like, okay, sure, if you're trying to get an antibody for a very simple kind of target, maybe, right? But I think what we've been most excited by is our partners using antibodies in more sophisticated ways, right? Like, for example, in Chai 2, we showed GPCR agonist activity, right? Where you can really hit the switch on a cell doorbell protein, so to speak, in a very precise way.

Host

如果你做不到那么精确,用抗体做到这一点非常非常难,对吧?

Very, very hard to do that with antibodies if you can't be that precise, right?

Matt

所以你是在解锁一种新能力。我更倾向于认为,这不像是在说,哦,我把现有能做的药物做得更快。我是说,确实也有那部分,对吧?但更像是,嘿,你怎么去追求更好的靶点,对吧?那些可能更精确、更有效的靶点,对吧?

So you're unlocking a new capability. I would think about it as less like, oh, I'm taking the existing drugs that I can do and making them faster. I mean, there is some of that too, right? But it's like, no, there are just like, hey, how do you go after better targets, right? That are, you know, maybe more precise, more effective, right?

Matt

我觉得除此之外,还有一些药物模式是你通过免疫方法根本无法发现的。比如,你不会去设计那种疯狂的多特异性、弹头式、超强效的格式。这些东西真的需要你从第一性原理出发去设计。即使仅仅以双特异性抗体为例,两条臂现在需要结合不同的靶点。而且你在结合率上会有这种乘法效应。所以,如果你在第一条臂上找到结合物的概率是十亿分之一,在第二条臂上也是十亿分之一。

I think also, on top of that, there are drug modalities that you just can't discover with immunization. Like, you're not going to design your crazy multi-specific, warheaded, super intense formats. These are really things where you kind of have to design these from first principles. Even just with bispecifics in particular, like both arms need to now bind different targets. And you've kind of got this multiplicative effect on your binding rate. So like, if you have a one in a billion chance of finding a binder in arm one and a one in a billion chance in arm two.

Host

嗯。

Yeah.

Matt

你不可能——用传统方法根本行不通。

You're not—this just isn't going to work with the traditional approach.

Host

确实。

Exactly.

Matt

我想考虑的另一点是,对吧,你不只是帮助合作伙伴开发一种药物,对吧?可能有一组靶点,或者他们正在尝试开发一组药物。平台方法的好处在于,而不是我们开发单个药物,我们可以随着他们追求更多靶点、更雄心勃勃的靶点而扩展。

I think the other thing I'd think about is, right, you're not just helping your partner with maybe one drug, right? There might be a portfolio of targets, or they're going after a portfolio of drugs that they're trying to make. And the nice thing about the platform approach, rather than we are developing individual drugs, is we can sort of scale with them as they pursue more targets, in addition to more ambitious targets.

Host

对。所以它让你把学习集中在那个子领域,这样每个人都能从中受益。

Right. So it lets you concentrate your learning in a subdomain of that, so that everybody benefits from that.

Matt

没错。这就是——但我——好吧。所以我——那这些能力有哪些?你提到了一些。还有没有你们正在追求的、真正有趣的能力?

Exactly. That's the—but I—okay. So I—so what are some of these capabilities? You mentioned a few. Are there more that are really interesting that you guys are chasing?

Matt

是的。我是说,我们谈到了交叉反应性。我们谈到了选择性。我们谈到了双特异性抗体的一些非常有趣的额外模式,对吧?有一系列我们的合作伙伴一直在要求我们做的事情,我们一直在努力,但我不能深入太多,因为那会开始暴露他们正在追求的一些靶点。但我认为重点是,一旦你变得精确,你就能开始制造一些非常非常酷的药物。

Yes. I mean, we talked about cross-reactivity. We talked about selectivity. We talked about some of these really interesting additional modalities with bispecifics, right? There's a set of things that our partners have been asking us for that we've been working on that I can't get too into, because then that starts to reveal some of the targets that they're going after. But I think the point being, once you get precise, you can start to do some really, really cool drugs.

Host

这是一种新技术,对吧?在制药行业,技术意味着你如何递送你的治疗药物,所以这可能有点像 CAR-T 是一种技术,对吧?所以这可能是一种新技术,因为你可以拥有这些高度工程化的……

It's a new technology, right? So like technology in pharma means how do you deliver your therapeutic, and so this is maybe a kind of thinking about like CAR-T is a technology, right? And so this is maybe a new technology in the sense that you can have these highly, highly engineered...

Matt

对,这源于公司的使命,即真正将药物发现从科学实验转变为工程学科,对吧?你如何进入生物学的精密工程阶段,在这个阶段,你可以几乎以声明式的方式定义你想要的东西,然后让模型填补空白,为你实现它?

Right, and that comes from the mission of the company, which is to really turn drug discovery from a scientific experiment to an engineering discipline, right? How do you get to the precision engineering phase for biology, where you can start with almost declaratively defining the thing you're trying to get, and have the model fill in the gaps and get you that?

Host

那么从科学到工程的最大障碍是什么?

So what is the biggest blocker from going from science to engineering?

Matt

天哪,有太多事情了。这就是关于——你知道——

Oh man, there's so many things. That's the thing about—you know—

Host

嗯。嗯。

Yeah. Yeah.

Matt

我甚至不想谈这个,比如头疼的事情有多少——

I don't even want to talk about this, like the amount of headaches—

Host

太晚了。你已经知道了。

Too late. You already know.

Matt

好吧。好吧。所以,就像,当你实际解析——首先,生物学家的文件格式。就像,我——他们就是不在乎。没有标准化的——有标准化的文件格式。它们是最好的吗?我真的不知道。但也有很多你想打包进去的信息。我有这个结构。这是解决它的人。这是我用来解决它的方法。有很多事情在进行。然后,根据你用来确定这个 3D 结构的方法,你可能会有该结构的多个副本。其中一部分可能没有被真正解析,或者你会说,它可能在这里,也可能在那里。我就给你两个选项。所以,在工程方面处理这类数据的实际解析问题真的很难。

Okay. Okay. So, like, just when you're actually parsing—first of all, file formats for biologists. Like, I just—they just don't care. There's no standardized—there are standardized file formats. Are they the best? I don't really know. But there's also just a lot of information that you want to pack in. I have this structure. Here are the people who solved it. This is the method I used to solve it. There's a lot of stuff going on. And then, depending on the method that you use to actually figure out what this 3D structure is, you might have multiple copies of that structure. Part of it might not have really been resolved, or you're like, it could be here, it could be there. I'm just going to give you both options. So the actual parsing problem on the engineering side of working with this type of data is really difficult.

Host

不过,这似乎是 LLM 可以擅长的事情。

This seems like something that LLMs can excel at though.

Matt

它们通常不知道所有的边缘情况,对吧?

They don't know all the edge cases often, right?

Host

这更多是回到简单性的方法。比如 LLM 非常好。我完全承认这一点。

This is more back to just a simplicity approach. Like LLMs are very good. I will absolutely give you that.

Matt

然后你会想,我真的想——这个函数应该有 20 个特例,还是我们应该在如何处理这个问题上非常有原则,我们应该,我想,更……

Then you're thinking about, do I really want to—should this function have 20 special cases, or should we be really principled in how we approach this, and should we be, I guess, more...

Host

有主见。

Opinionated.

Matt

有主见,是的。比如,我们在做这件事时应该多有主见?我们想要一个对人类来说足够容易理解的策略,当我们阅读代码库时,我们真的需要知道这里发生了什么,潜在的问题是什么,有时这只需要看例子。但然后,我想,好吧,一旦你解决了所有的基础设施工作以及如何将数据输入模型,接下来就是扩展模型,然后是扩展模型周围的基础设施,以训练越来越大版本的模型,这是 Neil 和产品团队实际做的很多工作……

Opinionated, yes. Like, how opinionated should we be in how we do this? We want a strategy that's easy enough for humans to understand, and when we're reading through the codebase, we really need to know what's going on here, what are the potential problems, and sometimes that just comes down to looking at examples. But then, I think, okay, once you've kind of figured out all the infra work and how you get data into the models, there's then scaling the model, there's then scaling the infrastructure around the model to train bigger and bigger versions of this, and that's a lot of work that Neil and the product team actually...

Host

是的,我的意思是,那本来是我的答案,就是基础设施部分。我的意思是,你知道,不是老调重弹,但算力——获得算力并以正确的方式使用它,是一个巨大的挑战,你知道,尤其是对初创公司来说。

Yeah, I mean, that would have been my answer, is the infrastructure part. I mean, you know, not to beat a dead horse, but compute—getting the compute and using it in the right way is such a challenge, you know, especially for startups.

Matt

这真是——是的。Anthropic 单枪匹马地阻碍了科学。Anthropic 开放……

This has been such a—yeah. Anthropic is single-handedly holding back science. Anthropic opening...

Host

不,我的意思是,关于这一点,就像我们——

No, I mean, and to that point, like we—

Matt

我的意思是,他们也在加速科学,但这就像一种奇怪的——

I mean, they're also accelerating science, but it's like this weird—

Host

完全。比如,我在 Chai 经常帮忙的事情之一就是为公司购买算力。

Totally. Like, one of the things that I help a lot with at Chai is buying compute for the company.

Matt

最糟糕的工作,兄弟。我不会推荐。压力非常大。

Worst job, man. I would not recommend it. It is very stressful.

Host

但你知道,即使是去年九月,对吧?

But you know, even September of last year, right?

Matt

回到你,硬件工作。

Back to you, the hardware job.

Host

嗯。嗯。我知道。没错。用错误的方式。

Yeah. Yeah. I know. Exactly. In the wrong way.

算力紧张与市场动态 Compute crunch and market dynamics

Matt

但是,你知道,去年九月,我们开始真正注意到事情变得紧张了,对吧?我们很多推理都在现货市场和按需市场上进行,有时候会遇到算力紧张的日子,我们就想,好吧,我们可能应该开始提前为自己购买一些算力。而且,我觉得每个人可能都会这么说,但是,天哪,这太难了。我没有意识到这有多大的幂律效应,对吧?比如,有 10000 台 B300 设备在各地出货,超大规模云厂商和最大的 AI 实验室买走了 95% 以上,然后初创公司就只能争抢剩下的残羹冷炙。

But, you know, September of last year, we started to really notice things were getting tight, right? We were doing a lot of our inference on spot and on-demand markets, and we'd have these days where you'd just get these capacity crunches, and we're like, okay, we should probably start to get ahead of buying some compute for ourselves. And, I mean, I think everyone probably says this, but man, it was hard. I didn't realize how much of a power law this is, right? There's, say, 10,000 B300 units shipping everywhere, and the hyperscalers and the biggest AI labs are buying 95+% of it, and then you kind of have the startups fighting over the scraps.

Matt

而且我觉得另一件非常有趣的事情是,特别是如果你看看这些较新的算力版本,对吧?Vera Rubin 或 B300,很多这些东西都是为 LLM 优先构建的,对吧?你有这些带巨大 KV 缓存的系统,有 72 个 GPU 都需要互相通信,对吧?显然那里的一些性能提升对我们有帮助,但有趣的是,算力市场在多大程度上被 LLM 主导了。

And I think the other thing that's really interesting, especially if you look at these later compute versions, right? The Vera Rubins or the B300s, a lot of this stuff has been built very LLM-first, right? You have these systems with huge KV caches where you have 72 GPUs that are all required to talk to each other, right? And obviously some performance gains there help us, but it's kind of interesting just how much the compute market has gotten LLM-pilled.

Matt

我觉得可能有一整套的算力栈和推理优化等等,需要为这类模型做。而且我认为这类模型会像 LLM 一样大、一样有影响力,但几乎就像算力还没有意识到这一点,无论是在容量方面还是在软件栈方面。所以我们实际上花了很多时间,甚至只是做基本的算力优化,让它们更好地为我们拥有的模型类型工作。

I think there's probably a whole set of compute stack and inference optimizations and things that need to be made for this class of models. And I think this class of models is going to be just as big, just as impactful as LLMs, but it's almost like the compute kind of doesn't realize that yet, both in the capacity sense but also in the software stack sense. So we actually spend a lot of our time, even just doing basic optimizations of compute, to get them to work better for the types of models that we have.

Host

是的,我知道一些结构化模型比 LLM 更具递归性,例如,所以这改变了你需要的算力与内存的比率等等。根据模型类型,你们做了哪些酷炫或有趣的优化?

Yeah, I know that some structured models are more recursive than LLMs, for example, and so that changes sort of the compute-to-memory ratio that you need and things like that. What are some of the cool or interesting optimizations that you've done there, depending on the type of model?

Matt

比如我们可以回到一种 try-one 类型的模型。在这种情况下,我们遵循 Fold 2/3 架构,在那里你不是在正常的序列表示上做注意力,而是在某种意义上松散地在成对表示上做注意力。所以你可以把它想象成一个长度为 L 平方的序列,而不是通常的长度 L。如果你在那上面做注意力,你实际批处理的方式最终会是 L 立方。现在你处于一个相当重的算力状态。你投入到每个 token 的 FLOPs 数量保持相当高。从 SRAM 传输到其他任何地方的内存带宽开销,在这些架构中是一个真正的瓶颈。所以即使像层归一化这样简单的事情也可能需要很长时间,实际上,这可能是你使用的算力的很大一部分。

So like we can go back to a try-one-type model. In that case, we're following the Fold 2/3 architecture, and there you're rather than doing attention over a normal sequence representation, you're in a sense loosely doing attention over a pair representation. So you can think of this as a sequence of length L squared rather than typically length L. If you're doing attention over that, the way that you actually batch this up, it ends up being L cubed. Now you're in a pretty heavy compute regime. The amount of FLOPs that you're putting into every token stays pretty high. The amount of memory bandwidth overhead of just transferring that from SRAM to whatever, that's a real bottleneck in these architectures. So even something as simple as a layer norm can take a long time, actually, and that can be a significant amount of the compute that you're using.

Matt

所以我认为在我们这边,我们花了很多时间优化它,并且非常认真地对待工程,所以这些操作至少更好。我们一直在关注新芯片与旧版本相比的表现。有时训练和推理甚至不同,当然 Neil 对此非常了解。

So I think on our side, we've spent a lot of time just optimizing it and taking engineering very seriously, so that these operations are at least better. We're always looking at how new chips perform compared to the older versions. Sometimes that's even different for training versus inference, and of course Neil knows this really well.

持久执行与Temporal Durable execution and Temporal

Host

嗯,所以有你在单个 GPU 上做的事情,然后还有你如何编排 GPU 集群,对吧?你基本上对计算进行分片,对吧?所以当你在 Chai 上设计一个分子时,不一定是一次调用,对吧?而是很多 GPU 被投入到这个问题上,跨越很多算力。而且,实际上,我想说软件工程中最难做对的事情之一是持久执行。你们熟悉这个术语吗?我能继续讲一点吗?

Well, so there's what you're doing on the individual GPU, and then there's how do you orchestrate fleets of GPUs, right? You basically shard your computation, right? And so when you're designing a molecule on Chai, it's not necessarily one call, right? It's a lot of GPUs being thrown at the problem across a lot of compute. And actually, I would say that one of the hardest things to get right in software engineering is durable execution. Are you all familiar with that term? Can I go on a little?

Matt

最终,如果你在非常广泛的基础设施上计算大量数据、模型调用,你总会遇到这些问题,基础设施的某些部分不稳定,对吧?也许你从中获取数据的存储桶宕机了,或者你的数据库因为事务太多而出现故障,或者你的 GPU 出错,对吧?我以前在几家公司工作过,你花了很多时间处理这个问题。你基本上把所有这些队列和所有重试都放在一起,用胶带粘在一起,结果变成了一团糟,现在原本应该是一个理想情况下相当简单的分布式计算,你却把 95% 以上的时间花在了所有这些排队和重试的事情上。

Ultimately, if you're computing a lot of data, model calls across a very wide set of infrastructure, you always run into these problems where some part of the infrastructure is flaky, right? Maybe the bucket you're grabbing your data from goes down, or your database has a blip because there are too many transactions against it, or your GPU errors out, right? I've been at companies before where you spend so much of your time just dealing with this. You're basically putting all of these queues and all of these retries, and you're duct-taping things together, and it becomes this mess where now what used to be ideally a pretty simple computation that's just distributed, you're ending up spending 95+% of your time on all of this queuing and retry stuff.

Matt

我们是 Temporal 这家公司的忠实粉丝。基本上,有这样一个想法:看,如果你最终只是想让一个长时间运行的任务运行起来,你需要什么?你需要一个队列。你需要你的不稳定组件从队列中拉取。你需要一些重试逻辑,如果失败就把东西放回队列,对吧?然后你需要一个完整的编排系统来把所有队列连接起来并监控它们。Temporal 真正酷的地方在于,这是一家发明了用于做这件事的框架的公司。

We're huge fans of this company called Temporal. Basically, there's this idea: look, if you're just trying to get a really long-running job to run at the end of the day, what do you need? You need a queue. You need your flaky thing pulling off of the queue. You need some retry logic to put things back on the queue if they fail, right? And then you need some whole orchestration system to just tie all the queues together and monitor them. What's really cool about Temporal is that this is a company that's kind of invented a framework for doing this.

Matt

我们早期做出的一个非常有帮助的技术决定是尽可能多地在 Temporal 上运行东西,对吧?无论是从应用程序调用数据库以确保数据库事务成功完成,好吧,让副作用放在 Temporal 上,这样它们就能被智能地重试,而不需要我们编写自己的队列逻辑,对吧?或者与模型调用相关的事情,或者与编排非常长的数据管道相关的事情。重点是,其中一个原语就是,嘿,你需要正确地进行持久执行,这样你就不会陷入重试地狱。这是一个非常深入的工程问题,除非像我和 Jack 一样,你以前被这个问题烧过很多很多次,否则你不会意识到。

One of the technical decisions we made early on that was very helpful was to run as much stuff as we can on Temporal, right? Whether those are calls out to the database from the app to make sure the database transaction goes through without failing, okay, let's have side effects sit on Temporal so that they get retried smartly without us having to write our own queue logic, right? Or things related to model calls, or things related to orchestrating really long data pipelines. Point being, one of those primitives is just like, hey, you need to get durable execution right so that you're not stuck in retry hell. It's a really deep engineering thing that you wouldn't realize unless, like for me and Jack, you've been burned by this many, many times before.

Matt

而且我认为我们现在处于这种状态,对吧,我们又筹集了 4 亿美元。我得再去买一个算力集群。我们将会有非常非常非常大的运行以及推理和训练集。所以把这些基础打好,实际上才能让我们做更有雄心的事情。

And I think we're at this state now, right, where we've raised another $400 million. I have to go buy another compute cluster. We're going to have really, really, really large runs and inference and training sets. And so getting those foundations right is what's actually going to let us do more ambitious things.

生物学中的工程原语 Engineering Primitives in Biology

Matt

为了回答你的问题,我其实认为,让生物学更像工程学的瓶颈,很大程度上在于拥有正确的工程原语。

And to kind of answer your question, I actually think that's a lot of the bottleneck to making biology more like engineering is just having the right engineering primitives.

Host

我在模型这边有一个类似的题外话。实际上,这些问题的好处之一是它们非常可见。所以至少你知道,比如,这个崩溃了,这个失败了,我们只是看到损失曲线没有下降,或者我们看到奇怪的梯度行为之类的。我认为很多同样的原则,比如工程优先,也适用于研究团队。我喜欢说的一件事是,复杂性和吃苦头药丸(bitter lesson)从根本上是对立的。例如,我认为 AlphaFold 3,我可能记错数字,但我觉得它有大约 23 个子模块,在这一点上,这是一个非常难以优化和研究的系统。你会想:“好吧,如果我调整子模块 30 或 21 里的这个东西,会发生什么?整个系统会怎样?”你总是可以想:“嘿,我们可以通过添加模块 24 来改进,但你应该这样做,还是应该考虑删除一些东西,降低复杂性?”但我认为在 Chai,这是一个非常根本的东西,就是工程文化,以及非常偏向简单。你们见过 SpaceX 发动机的图片吗?就像 Raptor 1 有很多管道,而 Raptor 2 没有。我们办公室墙上就挂着那张图,因为我觉得这真的很对,对吧?你怎么才能删除、删除、再删除更多东西?

I have an analogous tangent on the model side. Actually, one of the things that's kind of nice about those problems is they're super visible. So at least you know, like, hey, this crashed, this failed for us, we just see loss curve didn't go down or we see weird gradient behavior or whatever. I think a lot of these same principles, like engineering first, that also applies on the research team. One thing that I like to say is that complexity and being bitter lesson pill are fundamentally at odds. For example, I think AlphaFold 3, I might get this number wrong, but I think it was like 23 submodules, and at that point that's a really difficult system to optimize and study. You're like, "All right, what happens if I tweak this thing in submodule 30 or 21? What happens to the whole system?" And you can always think, "Hey, we can make this better by adding module 24, but should you, or should you think about just removing things and lowering that complexity down?" But I think that's a pretty fundamental thing at Chai, just the engineering culture and being very simplicity-biased. Have you all seen the picture of the SpaceX engines? It's like Raptor 1 has a bunch of pipes, and Raptor 2. We have a picture of that on our office wall, because I mean it's just true, right? Like how do you delete, delete, delete more things?

Matt

是的。

Yeah.

Host

但你能做到这一点的唯一方式是,我的意思是,AlphaFold 2 和 AlphaFold 3 之所以有效,是因为它们相对来说是小型模型。它们非常消耗算力,但数据效率非常高。

But the only way you can accomplish that is, I mean, the reason AlphaFold 2 and AlphaFold 3 worked is they were small models, relatively speaking. They were very compute-intensive, but they were very data-efficient.

Matt

是的。

Yes.

Host

而且有一个又一个的归纳偏置,是由人类直觉和可能是艰苦斗争得来的经验引入的。

And there was inductive bias after inductive bias brought in by human intuition and probably hard-fought experience.

Matt

嗯。

Mhm.

Host

它们非常高效。如果你试图推翻那些东西,你知道,它们不像纸牌屋。一切都是在其之上的增量改进。为了超越这一点,在我看来,你确实需要新的数据来源。你至少需要以一种更高效的方式从根本上不同地对待数据。

They're incredibly efficient. If you try to knock down those things, you know, they're not like a house of cards. Everything is an incremental improvement on top of it. In order to get beyond that, it seems to me like you really just need new sources of data. You need to at least treat data fundamentally different in a way that is much more efficient.

Host

我其实有点惊讶听到你们已经扩展到了那个程度,因为我觉得你们在做的事情和社区的想法非常不同。我不知道你是否能对此发表评论。

I'm actually kind of surprised to hear that you have scale to that degree, because I suggest that you're doing something very different from what the community is thinking. I don't know if you can comment about that.

Matt

我们是相当第一性原理的人,整个 Chai 的研究团队都是。除了我和 Kevin,真的,我们是仅有的有生物学背景的人,而且即便如此,我们也已经相当远离了。所以我认为我们试图把每个问题都看作一个核心机器学习问题。我们试图思考在其他领域中的类比是什么。所以即使对于图像模型,比如 CNN 是为了处理图像而构建的,所以图像应该以补丁的方式看待,那是那里的良好归纳偏置。然后人们会说,好吧,你可以把这个东西分词化,扔进 Transformer,它就会工作,而且它确实最终工作了,即使是在相对较小的数据集上。但我认为特别是对于蛋白质,这真的很难。没有那么多结构数据。有大量的序列数据,这是 ESM 能够工作的解锁因素之一。你可以让它直接在 Transformer 上运行。如果你尝试用实验结构数据做同样的事情,祝你好运。

We're pretty first-principles people, the whole research team at Chai. Except for me and Kevin, really, we're the only people with a bio background, and even still, we're pretty far removed. So I think we try to look at every problem as a core ML problem. We try to think of what's the analog in other spaces. So even for image models, like CNNs were built to process images, so images should be looked at in patches, that was the nice inductive bias there. Then people are like, well, you can just tokenize this thing, throw it into a transformer, and it's going to work, and it did end up working, even on a relatively small dataset. But I think for proteins in particular, it is really hard. There's not as much structural data. There's a ton of sequence data, and that's one of the unlocks for ESM working. You can get that to just run on a transformer. If you try to do the same thing with experimental structure data, good luck.

Host

我的意思是,有一篇 Apple 的论文,他们在 disfold 上进行了蒸馏,这真的很酷,你可以在非常大的数据集上进行蒸馏,并获得良好的信号,但它完全没有泛化,因为它不是推理。它真的是模式匹配。比如你提到的这些三角层,例如,它们确实有一个非常好的归纳偏置。也许它不是论文最初提出的三角不等式,但它是一个干净的归纳偏置,而且它明确地是让它起作用的东西之一,但它带来了巨大的成本。

I mean, there was that Apple paper where they distilled on the disfold, which was actually really cool that you could distill on a very large dataset and you could get good signal, but it didn't generalize at all because it wasn't reasoning. It was really pattern matching. Like one of the things, these triangle layers you were talking about, for example, they do have a very nice inductive bias. Maybe it's not the triangle inequality like the paper originally proposed, but it's a clean inductive bias, and it unambiguously is one of the things which made it work, and it just comes at a huge cost.

Matt

是的。

Yeah.

Host

是的。不,我认为这绝对是正确的。这些层相当昂贵。这限制了你能用这些架构做什么。它们不仅在算力方面昂贵,而且在现代 GPU 上也不高效。你有小的隐藏维度,大的序列维度,这恰恰与 GPU 设计用来处理的东西相反。从三角层中得到的一个启示是,在某种意义上,你只是在用参数换取算力。这是思考这个问题的一个心智模型。我可能想向问题投入更多算力,然后用它来换取参数,因为我无法容纳那么多,我实际上无法存储这些大的成对表示,同时仍然进行正常的注意力机制。所以我认为你可以从 AlphaFold 这样的想法中抽象出一些基本的东西,但你可以调整这些,并以你自己的方式开始在此基础上构建。

Yeah. No, I think that's definitely true. These layers are pretty costly. And that kind of limits what you can do with the architectures. Not only are they costly in terms of compute, they're just not efficient on modern GPUs either. You have small hidden dimensions, large sequence dimensions, like it's exactly the opposite of what GPUs are designed to process. One takeaway from triangle layers is you're kind of just trading off parameters for compute in that sense. That's one mental model for thinking about this. I might want to throw more compute at the problem and just trade that off for parameters, because I won't be able to hold as many, I can't literally store these large pair representations and still do normal attention. So I think there are fundamental things you can abstract from the ideas like AlphaFold, but you can kind of just tweak these and start building off of them in your own way.

Host

听起来你们在这个方向上有相当多的研究,基础研究。我想对于寻找“书呆子狙击”的听众来说,以及针对新问题的机器学习工程,这可能是一个非常不同于许多社区正在走的研究方向。

It sounds like you have quite a bit of research, fundamental research, going into this direction. I guess for audience looking for a nerd snipe, and ML engineering for new problems, probably something very different research direction than a lot of the community is going in.

Matt

是的。是的。我认为我们在 Chai 构建的东西在很多方面都非常独特,但也与核心机器学习擅长的事情紧密相连,就像我之前说的那样。我们试图把每个问题都映射到一个核心机器学习问题。我们思考,如果这是一个大语言模型或类似的东西,你会如何处理?但归根结底,我们真的很重视简单性。我们真的鼓励没有生物学背景的人不要害怕这些东西。

Yeah. Yeah. I think what we built at Chai is very unique in a lot of ways, but also very tied to what core ML is good at, kind of what I was saying before. We try to map every problem into a core ML problem. We think, how would you approach this if it were an LLM or something like that? But yeah, at the end of the day, we really value simplicity. And we really encourage people who don't have a bio background to not be scared of this stuff.

产品通用性与设计哲学 Product generality and design philosophy

Matt

而且我认为这也延伸到产品层面,你知道,这里需要取得平衡,对吧,比如你把产品做得有多通用?你是构建一个交叉反应工作流、一个选择性工作流和一个特异性工作流,还是你们都说,不,让我们把模型做得足够通用,让它能说“我要任意结合或避免某些东西”,然后你在 CAD 套件里就有一个非常通用的筛选器,你可以说“嘿,我只想避免或结合这些不同结构的这些部分”,对吧?而且我认为,你知道,就像机器学习团队——我没有正式的生物学背景,产品和平台团队的大多数人也没有正式背景。现在,我可能会后悔我说的话,对吧?因为我确信有无数细微差别,我不想显得太鲁莽或天真。但是,你知道,我认为有时候不被所有这些“哦,这些细微差别,这个那个”所拖累是有帮助的,你可以大胆地做到最大程度的通用,因为,呃,你知道,这就是我们在研究中看到的。你可以——模型非常通用,这让产品也能非常通用。

And I think that extends into the product too, where you know there's a balance to be had here, right, between like how general do you make the product? Like do you build a cross-reactivity workflow and a selectivity workflow and a by-specifics workflow, or do you all say no, like let's make the model general enough to say I'm going to like condition on arbitrarily binding or avoiding something, and then you just have a very general like screen in your CAD suite where you can say hey I just want to avoid or bind to these parts of these different structures, right? And I think, you know, kind of like the ML team—like I don't have a formal bio background, most of the product and platform team doesn't have a formal background either. Now, there's some amount of like maybe regretting my words that I'm gonna have, right? Because I'm sure there are, you know, a million nuances and, you know, I don't want to come off as too brash or naive there. But, you know, I think sometimes it's helpful to not be burdened by all of the, oh, these nuances and this and that, and you get to kind of bet and be maximally general because, uh, you know, that's kind of what we're seeing in the research. You can—the models are very general, that lets the product be very general.

Matt

我回想起我计算机科学理论的日子——我的第一位导师说,我们在研究某个问题,需要某个东西的多项式时间算法,他总是告诉我,永远不要低估多项式时间的力量——这基本上就是你可以选择任何你想要的指数。我的第一篇论文就是针对这个问题的 n 的 20 次方时间算法,我说,安迪,我完全照你说的做了。

I'm thinking back to my CS theory days—my first adviser was like, we're working on some problem and we needed a polynomial-time algorithm for something, and he would always tell me, never underestimate the power of polynomial time—like this is basically you're allowed to choose whatever exponent you want. And my first paper was an n to the 20th time algorithm for this problem, and I was like, Andy, I did exactly what you said.

Host

他说,等等,我不是那个意思。

He's like, wait a minute, I didn't mean it like that.

Matt

是的,但我认为这有点——你真的可以帮助自己,当你觉得“好吧,我可以做任何我想做的事,然后再简化它”时,你会解放很多。我认为这确实是一种非常基本的思考方式,我们在 Chai 经常利用这一点。

Yeah, but I think it kind of—you can really help yourself, you can free yourself a lot when you're like, all right, I can kind of do whatever I want and then kind of simplify it later. And I think that's really like a pretty fundamental way of thinking about things that we leverage a lot at Chai.

行业前景与竞争 Industry outlook and competition

Host

蛋白质设计和结合剂的空间实际上是一个相当拥挤的领域。我很好奇你对这个领域、这个行业的总体看法是什么。

The space of binders of protein design and binders in general is actually a fairly crowded space. I'm curious about what your general outlook of the field, the industry is.

Matt

我的意思是,我可以讲个轶事。我大概三四年前在 NeurIPS,对吧?就是 RF diffusion 刚出来后的那届。我和 Baker 实验室的某个人聊天,他们说,天哪,我一次就搞定了。我不认为他们当时甚至用“one-shot”这个词。当时“one-shot”还不是一个术语,但他说,我刚刚从 RF diffusion 得到了皮摩尔级别的结合剂,然后直接扔进冷冻电镜里。太好了。对吧。但这似乎并没有解决问题。就像,不是“哦,天哪,现在每个——”。

I mean, I can go back to some anecdote. I was at maybe NeurIPS three or four years ago, right? The one right after RF diffusion came out. I was talking to someone in the Baker lab and they're like, man, I just one-shotted. I don't think they even use one-shot. One-shot wasn't even a term back then, but they was like, I just got picomolar binders out of RF diffusion and just like threw in the cryo. Great. Right. It didn't seem like that just solved the problem. Like, it's not like, oh man, now every—

Host

是的。

Yeah.

Matt

但我认为有很多人已经看到,你实际上可以在至少某些类别中做得相当好的蛋白质设计。我想说,是迷你蛋白还是迷你结合剂?嗯,讽刺的是,纳米结合剂实际上比迷你蛋白更小或更大,或者可能更难一点。抗体通常被认为更难。但问题是,这是否是某种可以并且将会在某种程度上商品化的东西。你如何竞争?这个领域从这里走向何方?

But there are lots of people who I think have seen that you can actually do protein design at least in some categories quite well. I'd say like is it mini proteins or mini binders? Um, ironically nanobinders are actually smaller than or larger than mini proteins or maybe like a little bit harder. Antibodies are typically considered even harder. But there's this like is this something which can and will be commoditized at least in some part. How do you compete? Like where does this—where do you—where does the field go from here?

Matt

我的意思是,我认为答案是以上所有情况都有。我认为对于某些类型的药物或疗法,可能会有某种商品化层,对吧?同时,我们将能够做越来越雄心勃勃的药物,而且你会——这就像 LLM 领域发生的事情,对吧?你有开源模型,它们可能在某些方面是通用的、有用的,但人们仍然在购买前沿模型,对吧?实际上,如果你看看捕获的价值量,实际上是闭源前沿模型——你知道,整个蛋糕在增长,但增长如此之快,即使开源模型的份额在扩大,前沿模型仍然能够捕获大部分价值。

I mean, I think the answer is it's kind of all of the above. Like I think there probably will be some commodity layer for certain types of modalities or drugs, right? I think at the same time we're going to be able to do even more and more and more ambitious drugs, and you're going to—it's just like what's happened in LLM land, right? Like you have your open-source models that are maybe general and helpful for some things, but people are still buying frontier models, right? And actually if you look at the amount of value captured, it's actually the closed-source frontier models—you know, the whole pie is growing, but it's growing so fast that even as the open-source models' share expands, the frontier models are still able to capture the majority of the value.

Host

对吧?原因是什么?

Right? And what are the reasons for that?

Matt

第一,如果你有更多的智能,你会去处理更困难的任务,对吧?我认为如果我们有更智能的生物模型,我们会去处理更疯狂的任务,对吧?但还有一点,我的意思是,我不使用开源模型的原因很多,比如你知道,我没有得到像 Claude Code 这样的东西,对吧?我没有得到像 cloud——你知道,我认为有一个产品层需要构建,它和模型层一样重要。

One, like if you have more intelligence, you're going to go after harder tasks, right? I think if we have more intelligent biomodels, we're going to go after more crazy biotasks, right? But then also too, like I mean a lot of the reason I don't use the open-source model is 'cause like you know, I don't get like Claude Code, right? I don't get like cloud—you know, I think there's a product layer to be built that is just as important as the model layer.

Matt

我们从合作伙伴和楼里的人那里学到了很多,比如他们在使用模型时真正遇到的困难是什么,对吧?其中一些是最愚蠢的事情,对吧?比如,你知道,我想更好地可视化这个部分,并专注于那个。还有一些实际上是非常复杂的事情,我们必须为它们构建一些相当垂直的产品。而且,看,也许在时间的尽头,AGI 会一次性解决所有问题,无关紧要,但我认为距离那一步还有相当长的时间,对吧?我认为产品在这方面会产生巨大的差异。这就是我的答案。我的意思是,你可能有一个更以模型为导向的答案。

We learn a lot from our partners and the people in the building as well, just like what are the really tough things that they get stuck on using the models, right? And some of them are like the dumbest things, right? Like, you know, I want to be able to better visualize this piece and like focus on that. And some of them are actually very sophisticated things that we then have to build some pretty vertical product for. And look, maybe in the fullness of time, like AGI one-shots everything and doesn't matter, but I think there's quite a bit of time until we get there, right? And I think the product makes a huge difference for that. That'd be my answer. I mean, you probably have a more model-forward answer.

Host

不,我认为生物学是缓慢的,这是一件好事,而且没有那么多标记数据。所以你可以利用所有公开可用的序列信息,那可能给你一个很好的基础模型,但你仍然需要对这些数据进行一些测量,这仍然相当耗时,然后你需要在此基础上迭代。所以我认为即使在解锁方面也存在数据障碍——比如如果我们真的想做到零样本设计候选,开始生成几乎可以进入临床的分子。我认为这不仅仅是——你知道,AGI 可能不会马上解决这个问题。我认为那里肯定有一些技术障碍,对吧?

No, like I think biology is slow, which is one kind of nice thing, and there's not that much labeled data. So like you could take all the publicly available sequence information out there, that might give you a good base model, but you still need some measurements on that data that's still pretty time-consuming, and then you need to iterate on that. So I think there are even just data blockers there into unlocking—like if we really want to do this zero-shot design candidate, start generating molecules that are almost ready to go into the clinic. I think there's more to that than just—you know, AGI might not solve that right away. I think there are definitely some technical blockers there, right?

差异化与战略 Differentiation and strategy

Host

但即使在专业公司领域——我的意思是,我不打算点名,但我认为大概有 10 到 15 家蛋白质设计初创公司。我认为 Chai 似乎走了两条路,一是全合一产品,二是如果你没有自己的数据模式,你就不试图做自己的平台。你知道,这最终会帮助你获胜,还是会成为障碍?我只是好奇。

But even in the space of specialist companies—I mean, I'm not going to start naming them, but there are, I think, probably 10 or 15 protein design startups. I think the two things which it sounds like Chai has gone on is one, all-in-one product, and two, you are not trying to do your own platform if you don't have your own data mode. You know, is that going to help you win out in the end, or is that going to be a blocker? I'm just curious about that.

Matt

是的,这是个好问题。

Yeah, that's a great question.

合作模式与数据策略 Partnership Model and Data Strategy

Matt

是的,我们绝对没有启动自研管线的计划。我们非常认真地对待合作模式。从个人立场来说,我很喜欢这种激励对齐:我们让模型变得更好,合作伙伴取得更大成功,然后这又会自我迭代。所以我认为 Chai 的一个独特之处就在于,既能与很多人合作,又能获得产品反馈,知道这是非常真实的,掌握在大型药企手中,他们真的在用这些模型开展项目。所以我们真的必须做到模型优先、以模型为中心,我们需要持续交付价值。这给研究团队和产品团队带来了很大压力,首先要服务好这些需求。研究团队总是追求更好的版本。我的想法是:如果你是一个不那么短视、更有远见的公司,那么模型和数据两边都有很多工作要做,但我认为两者都没有被耗尽。说我们不需要更多数据是愚蠢的,但说模型停滞不前、只能用数据解决问题也是愚蠢的。所以我认为两边都有巨大的增长空间,我们都很重视。我还要反驳一下“数据护城河不存在”的前提。这就好比说,所有与 Anthropic 合作的企业都不让 Anthropic 用他们的数据训练,所以你就无法构建擅长企业工作流的模型。第一,我们正在这方面投资——有办法把算力转化为数据,获得更多数据,我们正在这么做。但第二,你想获取的是什么样的数据?与这么多合作伙伴紧密合作并支持他们的一个酷炫之处在于,我们能真正了解哪些东西对研究有帮助。所以,我们不是在真空中做研究,基于假设性的酷炫想法,而是能根据合作伙伴自然提出的需求,进行有依据的研究。

Yeah, so definitely no plans of starting a pipeline. We take the partnership model pretty seriously. And personally, I love the incentive alignment: we make the models better, the partners succeed more, and it iterates on itself. So I think that's a pretty unique part of Chai—being able to partner with a lot of people and get feedback on the product, knowing it's very real, in legit big pharma hands, and they're actually running campaigns on this stuff. So we really have to be model-forward and model-focused; we need to keep delivering value. That puts a lot of pressure on the research team and the product team to serve these things. The research team always aims for better and better versions. The way I think about this is: if you're a bit less myopic and forward-thinking as a company, there comes a point where there's a lot to do on both the model and data side, but I don't think either is exhausted. It would be stupid to say we don't need any more data, but also stupid to say the models are stuck and we can only use data to solve these problems. So there's tons of room to grow on both sides, and we're taking both very seriously. I would also push back on the 'no data moat' premise. That would be like saying all the enterprises that work with Anthropic aren't letting Anthropic train on their data, so you can't build models that are good at enterprise workflows. One, we are investing in this—there are ways to turn compute into data and get more, and we're doing those. But then two, what is the kind of data you're trying to get? What's cool about working closely and supporting so many partners is we get to learn what would be helpful in research. So rather than doing research in a vacuum based on what would hypothetically be cool, we can do informed research based on what our partners have organically asked us for help with.

Host

我明白了。我猜你们不能基于合作伙伴的数据训练通用模型吧?你们会训练专门的模型吗?比如有没有一个 Artist 模型和一个辉瑞模型?

I see. Do you assume that you aren't allowed to train general models based upon your partner's data? Do you train special specialized models for like—is there an artist model and a Pfizer model?

Matt

是的,我的意思是,很多这类合作——而且这些都是公开的——我们正在与他们合作,为他们训练或微调我们模型的一个版本。我认为随着时间的推移,我们在这方面还能做更多。我哥哥创办了一家公司叫 Applied Compute,很不错的公司。他们正在为 LLM 做类似的事情,帮助企业真正理解其语言数据的价值,并将其用于专门任务。我认为在生物数据方面,我们也有可能开辟一个全新的领域。

Yeah, I mean, a lot of these deals—and this is all public—we are working with them to train or fine-tune a version of our model for them. And I think there's probably so much more we can do there over time. My brother started a company called Applied Compute. Great company. They're kind of doing this for design for LLMs, helping enterprises really understand the value of their language data and do that for specialized tasks. I think there's a whole world where we could potentially do that for biological data.

Host

那里的价值是什么?使用他们的数据能带来什么提升?我的意思是,仅仅是因为数据更多,还是因为它针对特定问题进行了专门化?

What is the value there? Like, what is the lift that you get from using their data? I mean, is it just that it's more data or is it more that it's specialized to a problem?

Matt

你知道,他们从实验中获得大量科学数据,这些数据可能有助于我们的模型在他们关心的特定候选物或靶点上表现得更好。

You know, they have a lot of scientific data that they're getting from experiments that can maybe help our models do better in particular classes of candidates or targets that they care about.

Host

嗯。

Yeah.

Matt

我的意思是,甚至像这样简单的事情:他们可能只是有一些偏好的做事方式,而这些方式并非 Chai 模型原生支持的,他们可以问产品团队,从某种意义上说,‘嘿,我们喜欢我们的设计具有属性 X,你能确保它们具有这些属性吗?’所以即使是这样简单的事情,对他们来说也有相当大的影响。

I mean, even something as simple as they might just have some preferred way of doing things that might not be native to the Chai model, and they can ask the product team, in a sense, 'Hey, we like our designs to have property X, can you make sure they have those?' So even things as simple as that actually have a pretty big impact for them.

Host

是的。这符合我个人的一个假设:所有 AI 公司,尤其是生物和科学领域的公司,实际上都是咨询公司。我认为制药行业尤其如此,因为你正在开发一种新药,对吧?这几乎从定义上就是新的,所以在很多情况下,现有的东西必须定制,除非你只是在重复旧的东西。但很多大型药企正在推动科学的边界。

Yeah. So this goes along with a pet hypothesis I have that all AI companies, and especially bio and scientific ones, are actually consulting companies. Pharma I think is particularly the case because you're developing a new drug, right? It's almost by definition new, so the existing stuff has to be customized in many cases, unless you're doing something that's just reiteration of old stuff. But a lot of the big pharma are pushing the boundaries of science.

Matt

是的,当然,我们的目标是让模型非常通用,让产品非常通用,让它强大。但是,是的,与每个客户都有集成工作。回答你的问题,这样做确实会获得一定的防御性,对吧?与这些合作伙伴建立信任关系的好处在于,希望如果我们第一年执行得非常好,那么他们就会继续与 Chai 合作,做更雄心勃勃的药物。

Yeah, I mean, certainly we aim to make the models very general, we aim to make the product very general, we aim to make it powerful. But yeah, there is integration work with every customer. And to answer your question, you do get some defensibility just by doing that, right? And what is nice about building trusted relationships with these partners is hopefully, if we execute really well over the first year, then they'll continue working with Chai to do more ambitious drugs past that.

Host

我的意思是,他们很难换供应商,对吧?

I mean, it's going to be hard to switch, right?

Matt

我希望如此。是的。

I hope so. Yeah.

Host

仅仅是为了获得安全性。

Just getting the security.

Matt

是的。也许另一个有趣的点是:如果你从每个词元的角度来看,我不知道还有哪个领域的下游价值能像制药业这样高。你想想最终出来的药物——这些可能是数十亿美元的资产。就 GLP-1 而言,我认为这两种 GLP-1 药物加起来可能是一个万亿美元的资产。

Yeah. Maybe one other interesting point is: if you think of this on a per-token basis, I don't know if there's another domain where the downstream value of a token is as valuable as it is for pharma. You're thinking about the actual drugs that come out—these can be multi-billion dollar assets. In the case of GLP-1s, I think the two GLP-1 drugs combined are maybe a trillion dollar asset.

Host

是的。我的意思是,直到大约 3 个月前,GLP-1 的总收入超过了所有 AI 实验室的总和。

Yeah. I mean, up until I think 3 months ago, GLP-1's total revenue was more than all of the AI labs put together.

Matt

是的,我觉得人们没有意识到这一点。我自己之前也没意识到。太疯狂了,对吧?

Yeah, I don't think people realize that. Like I didn't realize that. It's crazy, right?

Host

但市场却低得多。相对来说,市场如此,真是疯狂。

But yet the market is way lower. It's crazy how relatively speaking the market is.

Matt

而且,你知道,我之前没意识到制药业在多大程度上像风险投资业务,对吧?从某种意义上说,他们是在下非常大胆的赌注。

And you know, I didn't realize how much of a VC business pharma is in, right? They're in some sense taking really ambitious bets.

硅谷历史与生物技术融资 Silicon Valley History and Biotech Funding

Matt

我觉得最酷的事情之一就是研究硅谷的历史。人们想到硅谷就会想到软件,但在 80 年代,最大的风险投资成果之一是基因泰克。因为这是一个典型的 VC 模式,你得到一串令牌,可以在下游产生巨大的价值。

I think one of the things that was really cool was studying the history of Silicon Valley. People think of Silicon Valley with software, but in the '80s, one of the biggest venture outcomes was Genentech. Because it's such a VC model, you get a string of tokens that can give you so much value downstream.

Host

顺便提一下 Outposting 关于金融和融资的博客系列,真的很棒。

Just a general shout out to Outposting's blog series about finance and funding. Really fantastic.

Matt

在那之前,我知道很多这些观点,但我没有意识到这个兔子洞有多深。

Before that, I knew a lot of those points, but I didn't realize how deep that rabbit hole went.

Host

是的,生物制药领域最大的问题可能其实就是融资模式。

Yeah, it's maybe the single biggest problem in biopharma is actually just the funding model.

Host

你听说过阿拉姆定律吗?

Have you heard of Aram's law?

Matt

嗯,哦,是的。

Yeah, oh yeah.

Host

更反过来了。摩尔定律反过来。在算力领域,它有一个漂亮的指数对数线性缩放,但在制药领域恰恰相反。制造药物的成本在指数级增长。每种药物投入的资金以指数级速度增长,这很有趣。

More backwards. Moore's law backwards. In compute, it scales with a nice exponential log-line scaling, but in pharma it's the exact opposite. The cost of making a drug is increasing exponentially. The amount of money put in per drug is growing at an exponential rate, which is pretty interesting.

Matt

是的,这保证了在某个时候新药开发的边际回报将是负的。

Yeah, which guarantees at some point the marginal return on new drug development will be negative.

Host

完全正确。所以除非有人想出办法解决这个问题。

Exactly. So unless someone figures out how to fix this.

Matt

我认为我们可能正处于翻转这些曲线的边缘。

I think we might be on the verge of flipping some of these curves.

Host

弯曲 S 曲线。

Bending the S-curve.

Matt

是的。为了强调这一点,制药和 VC 从根本上都是在优化投资组合。这就是它们之间的联系。

Yeah. Just to belabor the point, pharma and VC fundamentally both are optimizing a portfolio. That's the connection there.

Host

是的,把制药公司看作老练的资本配置者,拥有一系列目标并在它们之间分配。这对我来说是一个很大的重新框架。希望制药公司未来能承担更大的风险,追求真正酷的药物靶点。

Yeah, thinking of pharma as sophisticated capital allocators, with a portfolio of targets and allocating between them. That was a big reframe for me. Hopefully pharma can take riskier bets and pursue really cool drug targets in the future.

Matt

这个类比实际上也是我们在 Chai 思考研究的方式。我们的研究团队相对较小,大约 10 人,与 Isomorphic 或 DeepMind 相比。我们几乎把它看作一份投资工作,将想法投资于算力。从这个意义上说,我们真的只是资本配置者。

That analogy is actually how we think about research at Chai as well. Our research team is relatively small, around 10 people, compared to Isomorphic or DeepMind. We think of it almost like an investing job, investing ideas towards compute. In that sense, we're really just capital allocators.

Host

是的,我觉得这可能有点太巧妙了,但我想提出一个更广泛的观点,我们认为 Chai 的每个人都是某种资本配置者。我们只有 30 人,这是因为我们雇佣的每个人,尤其是现在有了 AI 的赋能,都在将注意力分配到正确的想法上,并分配他们的算力。

Yeah, I think maybe this is too cute, but I would make the broader point that we think of everyone at Chai as a bit of a capital allocator. We're only 30 people, and that's because everyone we hire, especially now empowered with AI, is allocating their attention into the right ideas and allocating their compute.

Matt

这是机器学习 AI 项目和科学的一个特征。而如果你在为 B2B SaaS 公司构建 API,你的限制主要是人。但如果你在构建硬件、AI 模型或科学相关的东西,瓶颈是实验室、算力和其他东西。所以你必须处于这种有限分配的心态,有一些射门机会。我如何分配这些射门?

This is a characteristic of machine learning AI projects and also science. Whereas if you're building an API for a B2B SaaS company, your limit is mostly people. But if you're building hardware, AI models, or something scientific, the bottleneck is the lab, the compute, other things. So you have to be in that mentality of limited allocation, some shots on goal. How do I allocate those shots?

Host

嗯,我会说是也不是。我同意它更像那样,但回到为 B2B 公司构建 API 的例子,那个 API 有增量成本。你必须支持它,它增加了复杂性,它是另一个需要营销和销售的东西。也许你应该把它分配到产品路线图上的另一个赌注。在一个构建东西变得非常便宜的世界里,稀缺的是注意力,既包括你保持产品简单的注意力,也包括客户理解如何使用它的注意力。我不认为这是二元的,而是我们都会更像注意力的分配者。

Well, I would say yes and no. I agree it's a bit more like that, but going back to the example of building an API for a B2B company, that API has incremental cost. You have to support it, it adds complexity, it's another thing to market and sell. Maybe you should allocate that into a different bet on your product roadmap. In a world where building things gets really cheap, the scarce thing is attention, both yours to keep the product simple and your customer's to understand how to use it. I see it less as binary and more as we're all going to be a little more like allocators of attention.

Matt

是的,这就是高管们所做的。我们都正在成为高管。

Yeah, which is what executives are. We're all just becoming executives.

Host

嗯,我听了一个 Satya Nadella 的播客,他说微软想让每个人都成为无限思维的管理者。如果你把这个推向极端,每个人都会成为高管。我当然觉得自己像个高管,我每天都和 Claude 交谈。

Well, I was listening to a podcast with Satya Nadella, and he says Microsoft wants to make everyone a manager of infinite minds. If you take that to the extreme, everyone's going to be an executive. I certainly feel like an executive, and I talk to Claude every day.

Matt

一群实习生,他们都出去热切地解决你可能想要也可能不想要的问题,但他们正在解决问题。

A little suite of interns who are all going out and eagerly solving problems you may or may not have actually wanted, but they're solving the problems.

瓶颈移除问题 Bottleneck Removal Question

Host

所以,我们有两个典型的问题要问。我们已经问了一个,但我要更直接地问一次。如果你能通过法令消除问题空间中的一个瓶颈,那会是什么?

So, we have two typical questions that we ask. We've already kind of asked one, but I'm going to ask it again more directly. If you could remove a bottleneck from your problem space by fiat, what would that be?

Matt

这是一个有趣的问题。我觉得有一件事会非常好,我总是处于研究状态,很难关掉。对我来说,可能是一般蛋白质设计的验证循环。能够立即说,嘿,这个有效,这个无效。仍然有一些在黑暗中摸索的感觉。在 Chai,我们认真对待这一点,但这涉及到验证假设并确定事情是否有效。

That's an interesting question. I think one thing that would be really nice, I'm always in research land, very hard to turn off. For me, it's probably the validation loop of protein design in general. Being able to instantly say, hey, this works, this doesn't. There's still a bit of walking around in the dark. At Chai, we've taken this seriously, but it's along the lines of validating hypotheses and knowing for certain that things work.

Host

是的,这肯定是一个未解决的问题。

Yeah, that's an unsolved problem for sure.

Matt

未解决的问题。

Unsolved problem.

Host

是的,而且会非常有价值。

Yeah, and would be hugely valuable.

Matt

非常有价值。是的。我要给出一个更抽象的回答,实际上是人才默默无闻。

Hugely valuable. Yeah. I'm going to take a much more abstract answer to that, which is actually talent obscurity.

人才流动与可及性 Talent flow and accessibility

Matt

我觉得你知道有很多聪明人去做大语言模型(LLM)了。有很多人正在成为 SaaS 的软件工程师,但我觉得没有那么多聪明人去做生物。我高中时没做生物,因为我想我可以拿起电脑编程做应用,但如果我想做生物,我得去学习、在学校拿好成绩,也许还要读个博士之类的,对吧?也许这是原因之一。我觉得另一个原因是,很多这些东西真的很晦涩。我们在播客里抛出了很多大词。你没法真正可视化这些东西。这是我们在 CH 非常关心的一件事:如何让整个东西在我们的网站和产品中变得可视化。我们来这里的部分原因是,我觉得更多人应该意识到,你不需要有超级专业的生物背景就能在计算层面做出贡献。我经常思考人才流动,以及人才在经济中的去向。90 年代大家都涌向人才,2000 年代以来人们涌向科技,但大科技公司吞噬了大量人才,直到几年前,现在也许大语言模型和大型 AI 实验室正在吞噬很多优秀人才。但在元层面上,如何更好地分配人才?自私地说,我希望更多人才进入生物领域。我们可能希望更多人才进入制造业和物理世界的事情,以及美国面临的其他问题。但是的,我觉得更好地传达这一点,如果我有扩音器对所有人说话,我会尝试这样做。

I think you know there's a lot of smart people going and working on LLMs. There's a lot of people that are working and becoming software engineers for SaaS, but I think just not that many smart people go and work on bio. I didn't work on bio in high school because I thought I could pick up my computer and program apps, but if I want to work on bio, I have to go study and get good grades in school and maybe get a PhD or whatever. Maybe that's one reason for it. I think another reason is a lot of this stuff is really obscure. We threw around a lot of big words during this podcast. You can't really visualize the things. It's one of the things we care a lot about at CH: how do we make the whole thing feel visual on our website and in the product? Part of the reason we're here is I think more people should realize you don't need to have a super specialist bio background to contribute to this computationally. I think a lot about talent flows and where talent goes in the economy. In the '90s everyone was flowing to talent, and since the 2000s people have been flowing to tech, but big tech ate up a lot of the talent until a few years ago, and now maybe LLMs and the big AI labs are eating up a lot of the good talent. But at the meta level, how do you allocate talent better? Selfishly, I want more talent going into bio. We probably want more talent going into manufacturing and physical world things and these other problems that the US has. But yeah, I think communicating that better would be the thing that if I had a megaphone to talk to everyone, I would try to do that.

Host

好的。那么这就引出了第二个问题,也许答案是一样的,但你希望人们从这期节目中得到的一个要点是什么?

Okay. So then that leads to the second question, which is maybe the answer is the same, but what is the takeaway, one takeaway that you would like people to have from the episode?

Matt

是的,我觉得生物学一直是一个有点晦涩的领域,你在黑暗中摸索。你不知道你在看什么。你在实验中面对非确定性。你必须在很长一段时间内进行非常漫长且迭代的试错循环。在某个时刻,你跨越了计算能力的门槛,当你能把折叠模型做到埃米级别以内,当你能让设计模型给出超过 50% 的命中率,现在你可以把它们放进 96 孔板,实际上得到 48 个有趣的结合剂。你开始达到这样的点:现在你可以声明式地精确工程化你想要的东西,而不是赌自然或试错来达到目标。我们在软件领域也发生过同样的事情,你可以写代码并确定性地得到结果,或者在电气工程中,你不用手绘原理图,而是可以把它放进 Cadence 设计系统,在软件中完成,或者机械工程中的 CAD,你可以精确工程化你的零件并打印或制造出来。同样的事情正在生物领域发生,而且发生得非常快。

Yeah, I think biology has been this somewhat obscure feeling field where you're stumbling around in the dark. You don't know what you're looking at. You're dealing with non-determinism in your experiments. You're having to do a very long and iterative trial and error loop across a very long amount of time. And at some point you're crossing that threshold of what you can do computationally when you can get folding models down to being within an angstrom, where you can get design models to give you hit rates north of 50%, where now you can put them in a 96 well plate and actually have 48 interesting binders. You start to get to the point where now you can declaratively precision engineer what you want rather than betting on nature or trial and error to get you there. We had the same thing happen in software where you can write code and deterministically get an outcome, or in electrical engineering where instead of your schematic being drawn out, you can put it in Cadence design systems and get it made in software, or CAD for mechanical engineering where you can precision engineer your part and get it printed or manufactured. The same thing is happening in bio, and it's happening very quickly.

Host

是的。这真的为很多有趣的人打开了大门,也许以前它不那么清晰或可及,对吧?像我这样的软件工程师,像 Matt 这样的研究人员,显然我们仍然需要专家,但通才往往能真正加速该领域的精密工程。

Yeah. And that really opens the door for a lot of really interesting people, or maybe it wasn't as scrutable or accessible before, right? Like software engineers like myself, researchers like Matt, you know, obviously we're still going to want the specialists, but the generalists can often really accelerate the precision engineering happening in the domain.

Matt

是的,对我来说最大的收获是这个领域真的在起作用。它不仅具有商业吸引力,而且研究实际上显示出生命的迹象。它甚至不仅仅是显示出生命的迹象,生命的迹象已经被展示了。我们实际上处于模型有效的阶段。它们正在交付价值,而且还有大量非常有趣的研究问题需要解决。所以我认为这个领域的低垂果实比其他领域多得多。而且我认为你能产生的影响力,尤其是作为研究人员,在这个领域是无与伦比的。对我们来说,我们都是使命驱动的。但即使你不是,也有很多有趣的谜题要解决。有这种 3D 几何的角度。如果你喜欢扩散模型,在这方面有无数问题要解决。我们在 Chai 1 中有这个类似 LLM 的主干。核心机器学习被这些问题触及的太多了。虽然我们取得了巨大进展,但仍有很多工作要做。而且我认为这是最值得从事的领域之一,同时对人类产生一些最大的影响。

Yeah, I think for me the biggest takeaway is that the field is actually working. Not only does it have commercial traction, but the research is actually showing signs of life. It's not even just showing signs of life, the signs of life have been showed. We're actually in a place where the models work. They're delivering value, and there's still tons of really interesting research problems to solve. So I think there's a lot more low-hanging fruit in this field than there would be in other fields. And I think the amount of impact that you can have, especially as a researcher, is just unmatched in this field. For us, we're all very mission-driven. But even if you're not, it's a lot of fun puzzles to solve. There's this kind of 3D geometry angle. If you like diffusion models, there's a million problems to solve in that regard. We have this LLM-looking trunk in Chai 1. There's just so much of core machine learning that is touched by these problems. We've made a ton of progress, but there's still a lot to be done. And I think it's just one of the most interesting fields to be working in, while also having some of the largest impact on humanity.

Host

非常感谢你长途跋涉。这是一次很棒的 22 分钟步行。是的。我们期待追踪 Chai 的进展。太棒了。谢谢你们。非常感谢。

Thank you so much for making a long journey. It's been a great 22-minute walk. Yeah. And you know, we look forward to tracking Chai's progress. Awesome. Thank you guys. Thank you very much.

互动版:逐字朗读 + 针对本期提问 →