从偶然发现到工程设计:用 AI 改造生物学

From Serendipity to Design: Engineering Biology with AI

乔什·迈尔 Josh Meier · Training Data · 2026-08-04 · 约 47 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Chai Discovery 联合创始人探讨 AI 如何将药物发现从试错转变为以设计为导向的工程过程。

Chai Discovery co-founders discuss how AI is turning drug discovery from trial-and-error into a design-oriented engineering process.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 24)

全文 · Full transcript(中英对照)

指导原则:简洁 Guiding Principle: Simplicity

Josh

我们的一大指导原则就是简单。比如你看 Chai 1 这样的模型,我认为 Chai 1 里有 23 个不同的子模块。当你想在这样的系统上迭代时,会变得非常困难,因为你需要独立理解每一个子模块,理解它们所有的行为和动态,而这实际上不太能 Scaling。所以你可以开始思考:我该如何简化它?我该如何识别什么才是真正重要的?一旦你做到了这一点,整个研究过程以及识别这类 Scaling 方向就会变得简单得多。

One of our big guiding principles is just simplicity. So when you look at a model like Chai 1, I think there are 23 distinct submodules in Chai 1. And when you're trying to iterate on something like that, it gets really hard because you need to understand each of these submodules independently. You need to understand all of their behaviors and dynamics, and that doesn't actually scale that well. So you can start to think, okay, how do I simplify this? How do I identify what's really important? And once you have that, the whole research process and identifying these types of scaling directions becomes a lot simpler.

Chai Discovery 的大构想 The Big Idea of Chai Discovery

Host

今天我们请到了 Josh 和 Matt,Chai Discovery 的两位联合创始人。Chai 正在用 AI 工程化分子,可以说是一个面向生物学的 Foundation Model 实验室。这听起来是个很大的想法,但让我先从这里开始:这个大想法是什么?

We're here with Josh and Matt, two of the co-founders of Chai Discovery. Chai is engineering molecules with AI. So it's something of a foundation model lab for biology. Sounds like a big idea, but let me just start there. What is the big idea?

Josh

感谢邀请我们上节目。我们在 Chai 想做的令人兴奋的事情之一,就是让药物发现过程更像工程学。比如我们看到了所有用 LLM 做代码生成的工作,对吧?那之所以效果很好,是因为代码是一种非常简单的抽象,可以表达你想做的事情。而生物学今天不是这样的,它有很多试错。我想如果我们回到现代生物技术的早期,我们今天拥有的很多药物都是相当随机地被发现的,有点偶然。我们想做的,是让这个过程更工业化一些,并尝试开发出我们需要的工具,以便拥有像代码和现代工程中那样的抽象层,让我们能快速迭代,并把它带入生物学和药物发现。

Thanks for having us on the show. One of the exciting things we're trying to do at Chai is to make the drug discovery process look a little bit more like engineering. We've seen all the work happening with LLMs for code generation, for instance, right? And that works really well because code is a very simple abstraction to get across what you want to do. Biology doesn't look like that today. It's a lot of trial and error. I think if we go back to the early days of modern biotech, a lot of the medicines that we have today were discovered quite randomly. It was a bit serendipitous. What we're trying to do is allow us to industrialize that process a little bit more and try to come up with the tools that we'd need in order to have abstraction layers the same way we have that in code and modern engineering, which allows us to iterate really quickly, and try to bring that into biology and into drug discovery.

工程与试错边界 Engineering vs. Trial and Error Boundary

Host

那么让我就此问你一个问题。在可以被工程化的东西和需要在现实世界中测试的东西之间,几乎存在一条边界。而且感觉这条边界一直在移动——药物开发过程中可以被工程化而非试错的比例似乎在增加。这样想合理吗?如果合理,你能谈谈具体是哪些发展推动了这条边界的移动吗?

So let me ask you a question on that. There's almost this boundary between that which can be engineered and that which needs to be tested in the real world. And it feels like that boundary has been moving over time—the proportion of the drug development process that can be engineered as opposed to trial and errored seems to be increasing. Is that a reasonable way to think about it? If so, can you talk about what specific developments have pushed that boundary over time?

Josh

我认为这是一种思考方式。我们经常这样看:实验室是验证你所做事情是否正确的重要部分。实际上,验证在 AI 中也是一个非常大的主题。如果你能评估你的模型有效,你能验证它,那么你就可以开始爬山式改进并取得进展。所以我认为实验室是非常重要的部分。我认为问题在于我们如何让药物发现看起来更像药物设计。我们称这个领域为药物发现的原因之一,是我们经常在干草堆里找针。我们会筛选数百万或数十亿个分子,试图找到一个有效的。如果我们能输入你想要的分子的理想状态,然后模型能将其具体化,那将非常强大。所以这甚至不是减少实验室测试的问题。我的意思是,也许这会作为结果发生。我实际上可能持相反观点——也许我们会做更多的实验室测试,因为投资回报率会增加,就像现在软件工程师变得更有生产力,所以对软件工程师的需求更大一样。对实验室的需求可能会更大。但我认为关键是我们如何改变这里的范式,使其更以设计为导向。

I think that's one way to think about it. The way we often look at this is that the lab is an important part to verify that what you're doing is correct. And actually verification is a very big theme in AI as well. If you can evaluate that your model works, you can verify it, then you can start to hill climb and make progress on it. So I think the lab is a very important part of this. I think the question is how do we take drug discovery and make it look a lot more like drug design. One of the reasons why we call this field drug discovery is we're often looking for a needle in a haystack. We'll screen millions or billions of molecules, try to find one that works. If we can instead put in the dream state of the molecule you want and then the model can materialize that, that's going to be really powerful. So it's not even a question of reducing lab testing. I mean maybe that happens as a result. I actually might even take the opposite side of the coin—maybe we'd actually do even more lab testing because the ROI will increase, the same way there are more software engineers or there's more demand for software engineers now that they become more productive. There might be more demand for the lab. But I think the key part is how do we change the paradigm here and make it more design-oriented.

前沿技术演进 State-of-the-Art Evolution

Host

我们能谈谈今天的最先进水平是什么吗?然后也许让我们回忆一下——五年前、三年前、一年前。有哪些重大突破?最先进水平在近年来是如何变化的?

Can we talk about what is state-of-the-art today and then maybe let's take a little trip down memory lane—five years ago, three years ago, one year ago. What have been some of the major breakthroughs and how has state-of-the-art changed over the recent years?

Josh

是的,我认为这个领域真的在进化,尤其是在过去十年。直到——你可以回顾一下,有一个两年一度的蛋白质折叠竞赛。每两年,通常一群学术团队会参加这个蛋白质折叠竞赛。会发生的是,你会保留一些从未有人见过的蛋白质靶标,它们从未被公开存放,然后所有这些团队竞争,看谁能最好地预测蛋白质。直到 2018 年,你才开始看到性能的巨大飞跃,然后最终在 2020 年又有了 AlphaFold 2。但真正开启整个领域的,是深度学习的出现。所以一开始,你做的事情——这让我回到我在导师实验室的最初日子。我们研究了最早用深度学习做这件事的系统之一。我们试图预测蛋白质中氨基酸之间的距离,然后我们有一个渲染器,它接收这些嘈杂、看起来不完整的距离,然后从中生成一个蛋白质。相当了不起的是,用相对最少的信息,你可以构建一个机器学习系统,它实际上能产生看起来真的像蛋白质的东西,而且在大多数情况下是正确的。但仍然存在很大的差距,它没有达到我们今天的分辨率。几乎就像早期的图像生成模型,有点颗粒感,有点模糊,而现在就像,天哪,几秒钟就能生成超高分辨率的图像。所以我们看到了同样的进化随时间发生。所以从蛋白质折叠开始,然后一旦这变得更可实现,人们就想解决其他子问题。所以现在,给定一个蛋白质结构,我能否设计一个可能折叠成该结构的序列?这越来越接近药物设计问题。所以你开始思考,我如何开始工程化匹配特定形状或可能执行特定功能的蛋白质?人们独立研究了这个问题,那是在 2021、2022 年左右。然后这些想法开始融合。真正是扩散模型的出现,我们开始能够同时生成蛋白质结构和序列。我可以开始让这些我们称之为提示的东西变得越来越现实。

Yeah, I think the field has really evolved, especially over the last decade. It really wasn't until—you can actually look back, there's this biannual protein folding competition. Every two years, usually a bunch of academic groups would compete in this protein folding competition. What would happen is you kind of hold out these protein targets that no one's ever seen before. They never get deposited publicly, and then all these groups compete to predict the proteins the best. And really it wasn't until 2018 where you started to see this big step change in performance, and then finally again in 2020 with AlphaFold 2. But really it was kind of the advent of deep learning that kicked off this whole field. So at first, what you do—these are kind of like bringing me back to my original days in my advisor's lab. We worked on one of the first systems to do this with deep learning. What we would try to do is predict the distances between amino acids in a protein, and then we'd have some renderer which took in these kind of noisy, incomplete-looking distances and then emitted some protein from that. What's pretty remarkable is that with relatively minimal information, you could build a machine learning system that could actually produce something that really looked like a protein and for the most part was correct. There were still pretty big gaps. It didn't quite have the resolution of what we have today. Almost like with the early image generation models where it was a little bit grainy, a little bit blurry, and now it's like, oh my gosh, this is crazy high-definition images in seconds. So we kind of saw that same evolution happen over time. So it started with protein folding, and then once that became more realizable, there were these other subproblems that people wanted to solve. So now, given a protein structure, can I design a sequence which might fold to that? And that's getting closer and closer to the drug design problem. So you're kind of now thinking, okay, how do I start engineering proteins that match a certain shape or might perform a specific function? And so people studied that problem independently, and that was kind of 2021, 2022 time. And then these ideas began to merge together. It was really the advent of diffusion models where we started to be able to generate a protein structure and a sequence kind of simultaneously. And I can start making these what we would call prompts more and more realistic.

公司创立介绍 Introduction and Company Founding

Host

所以我现在可以针对某个特定形状的靶点来提示这些模型,说“嘿,我想在这里结合”,就像给问题加上更多现实世界的约束。你们在 2024 年看到了什么,让你们觉得那是创办公司的正确时机?

So I can now prompt these models on a target that might have a certain shape and I can say hey I want to bind over here kind of like add more real world constraints on the problem. What did you guys see in 2024 that made you think that was the right time to start the company?

Josh

是的,我们进行了很多讨论。我记得其中一次早期讨论是在 Matt 家里,当时 Matt 给我看了一些关于抗体-抗原结构预测的结果。Matt 当时在谈蛋白质折叠。但长期以来,人们认为抗体的蛋白质折叠太难了。比如,人们觉得蛋白质数据库里没有足够的抗体数据来解决这个问题。因此,人们认为抗体设计遥不可及。早期人们用深度学习做蛋白质设计的工作,很多都不是针对抗体的,而是针对一类叫做迷你蛋白的蛋白质。迷你蛋白本身确实很有趣,但它们不是大多数制药行业关注的对象。所以圣杯问题是,我们能否在计算机上设计抗体蛋白,尤其是那些具有全部治疗功能的抗体。我们的想法是,如果你无法预测抗体的结构,你怎么可能设计出抗体呢?

Yeah, there were a lot of discussions that went into it. I remember one of these early discussions actually back in Matt's house when we were, you know, Matt was showing me some of these results on antibody-antigen structure prediction. Matt was just talking about protein folding. But for a long time people thought that protein folding for antibodies was just too hard of a problem. People were like there was not enough data for antibodies in the protein data bank, for instance, to solve this problem. And people as a result thought that antibody design was going to be out of reach. A lot of the work that people even did on protein design in the early days with deep learning, it wasn't antibodies. It was these different class of proteins called mini proteins, which are actually really interesting in their own right, but they're not what most of the drug industry is looking at. So the holy grail was whether we could design antibody proteins on the computer, especially ones that had all that therapeutic function. And our thinking was that if you couldn't predict what an antibody looks like, how are you ever going to design one?

Host

传统上,人们认为生物学很可怕,要了解的东西太多了。

Traditionally, people have thought like biology is scary. There's so much to know.

Josh

我对生物学感到恐惧。

I'm terrified of biology.

Host

说实话,我也是。我的背景从来不是生物学。我学的是纯数学,博士开始读的是理论计算机科学,直到第三年我才转向深度学习、蛋白质结构预测这些领域。所以对我来说这完全是一个新领域,看起来很疯狂。但归根结底,它比人们想象的要简单得多,问题之间的关联性也更强。所以我觉得人们刚进入这个领域时,会想:“天哪,什么是抗体?什么是迷你蛋白?什么?”这些其实都只是氨基酸序列。归根结底,这些只是模型的不同类型的提示。就像你给 ChatGPT 一个数学问题,它既能回答你的数学问题,也能帮你做英语作业。所以我们的模型也是同样的情况。我们只是有某种方式来表示这些氨基酸序列,然后我们也有办法设计和预测它们。我认为从这个角度来看,事情就变得清晰多了。

I am as well honestly. Like my background was never biology. I studied pure math and started my PhD in theoretical computer science, and it was only after my third year that I ended up switching into deep learning, protein structure prediction, all of this stuff. So it was like a totally new field to me. Seemed insane. But at the end of the day, it's much simpler and the problems are much more interconnected than one might think. So I think people when they start going into the field, they're like, "Oh man, what's an antibody? What's a mini protein? What?" These are all just sequences of amino acids. At the end of the day, these are just different types of prompts for the model. But in the same way where you might have a math problem that goes into ChatGPT, ChatGPT can both answer your math problem and help you with your English homework. So really we have the same type of thing going on with our models. We just have some way of representing these sequences of amino acids, then we have a way of designing and predicting those as well. And I think in that lens, things become a lot more clear.

Host

Matt,你稍微提到了你的背景。Josh,谈谈你的背景,然后更广泛地说,为了实现这个目标,你们汇集了多种不同的学科。所以你能谈谈你的背景,以及团队其他核心成员的背景,以及这些东西是如何结合在一起的?

And Matt, you mentioned your background a little bit. Josh, talk a bit about your background and then more broadly, in order to pull this off, you have a bunch of different disciplines that kind of come together. So can you just talk a bit about your background and that of some of the other core members of the team and how these things all fit together?

Josh

是的,我从小就对人工智能和生物学充满热情。所以,我想和 Matt 不同,我不是从理论计算机科学开始的。我上高中时有一个干细胞实验室,所以我从小就对生物学感到兴奋,而且我是作为程序员长大的。我的职业生涯真正始于 OpenAI,我是那里的早期团队成员。那时它还是一个非营利组织,所以在那里工作是一段很好的时光。我们做了 GPT1、GPT2 和缩放定律。当时的问题是,如果模型能学会说英语、德语、法语,为什么它们不能学会说 DNA 和蛋白质的语言?从那以后,这基本上就是我的研究议程。我认为这与 Matt 进入这个领域的时间大致吻合。那也是一个非常重要的时期,对吧?因为如果你看看我们带来的方法,就会发现有很多边缘上的变化。Matt 谈到了这个领域的一些历史。但关键在于我们如何找到正确的深度学习架构、正确的算力配置和正确的模型架构来实现这一点。应该应用哪些正确的任务?我们在讨论我们怎么知道 2024 年是创办公司的正确时机。正如 Matt 所说,这些都是不同类型的氨基酸序列,人们认为抗体这类问题太难了。而我们开始看到第一批迹象表明这实际上开始奏效了。我认为这在很大程度上得益于我们为这些问题带来的一些新架构,当时我认为是扩散模型。

Yeah, I've been excited about AI and biology since I was a kid. So, I guess unlike Matt, I didn't start with theoretical CS and get into that way. But I went to a high school with a stem cell lab. So I was just always excited about biology as a kid and I grew up as a programmer. I really started my career at OpenAI, so I was on the early team there. It was a nonprofit back then, so it was a pretty good time to be there. We did GPT1, GPT2 scaling laws. And the question was like if the models can learn to speak English, German, French, why can't they learn to speak DNA and protein? And that was kind of my research agenda since then. And I think that sort of intersects with around the time like Matt got into the field as well. And it was a pretty important time as well, right? Because if you look at the kind of methods that we are bringing in, like there's been a lot of these changes on the edges if you will. And Matt talked a bit about the history of what's happened in the field here. But it's all about how do we find the right deep learning architectures with the right compute configuration and the right model architectures to make this happen. What are the right tasks to apply it to? We're talking about how we even knew that 2024 is the right time to start the company. As Matt was saying, these are all different kinds of amino acid sequences and people thought that the antibody class of problem was going to be too hard. And we started to see the first signs of life that actually this was starting to work. I think a lot of it fueled by some of the new architectures we're bringing to the problems, where I think it was diffusion models back then.

Host

第一次有人能够生成看起来合理的蛋白质是在扩散模型出现的时候。所以那真是疯狂。当时有很多生成建模方法,如果你有大量数据,它们就能奏效。人们让这些方法在图像上奏效了。一路上有很多技巧让它们越来越好。我们在自己的领域也肯定借鉴了很多这些想法。但真正起作用的是扩散模型的出现。

The first time anyone was able to generate reasonable looking proteins was the advent of diffusion models. So it was pretty crazy. There were a bunch of generative modeling approaches that would kind of work if you had a bunch of data. So people got these working for images. There are a bunch of tricks to make this better and better along the way. We've definitely borrowed a lot of those ideas in our domain as well. But it was really once diffusion models came around.

Host

那么,对于扩散模型为什么有效,有没有什么直觉上的解释?

And is there an intuition for why diffusion models work?

Josh

是的。所以扩散不是一个一步到位的过程。我认为在那之前,主要的生成设计范式叫做变分自编码器。在那种情况下,你是说我只想压缩我的输入分布。你有一些蛋白质,你想让它们看起来像模糊的高斯向量。这个任务真的很难。也许今天如果我们非常努力地尝试,我想我们能解决它。但在当时,我认为它还没有发展到足以解决我们的问题。真正效果很好的是给模型更多的时间思考,并向它展示更多的例子。比如,这里有一个看起来有点破损的蛋白质,你如何让它变得更好?你可以把那个蛋白质破坏得越来越厉害,让它看起来越来越嘈杂,越来越破损。然后教模型一些快速的小捷径。好的,我这样让它稍微好一点。然后你可以一遍又一遍地问,让它稍微好一点,让它稍微好一点。把问题分解到那个尺度,对生物学来说效果非常好。

Yeah. So diffusion is not a one-step process. So I think kind of up until this point, the main generative design paradigm was called variational autoencoders. And in that case you're saying I want to just compress my input distribution. You have some proteins, you want to make these look like fuzzy Gaussian vectors. That task is just really hard. And maybe today if we tried super hard, I think we could crack it. But at the time, I think it wasn't really developed enough to work for our problems. What turned out working really well was just kind of giving the model more time to think and showing it more examples. Like here's a slightly broken looking protein. How do you make it better? And you can kind of break that protein more and more and more. And you can make it look more and more noisy, more and more broken. And teach the model just quick little shortcuts. All right, here's how I make it slightly better. And you can just keep asking over and over again, make it slightly better, make it slightly better. And kind of like breaking the problem down to that scale worked really, really well for biology.

独角兽类比介绍 Introduction and unicorn analogy

Host

是啊,这个“让它变得更好”让我想起我和女儿们喜欢玩的一个游戏:我们让 ChatGPT 给我们一个独角兽,然后我们让它变得更强,我们一直告诉它要变得更强,等我们结束时,我们就拥有了世界上最强的独角兽。所以差不多,对吧?

Yeah, the make it slightly better reminds me of a game that I like to play with my daughters where we have ChatGPT give us a unicorn and then we make it stronger and we just keep telling it to make it stronger and by the time we're done we have the strongest unicorn in the world. So about the same, right?

Josh

是啊,呃,就这么简单。

Yeah, that's uh it's that easy.

Host

是啊。好吧。也许不一样。

Yeah. Okay. Maybe not the same.

Josh

嗯。

Yeah.

组建四语团队 Building a quadrilingual team

Host

所以我觉得在这个领域,你需要组建一个像“四语”团队那样的群体,就像化学家、生物学家、AI 人员的复仇者联盟,这确实是个挑战。嗯,你们是怎么找到人、说服他们加入团队的?你们的超级明星是谁?

So I feel like in this domain you need to assemble kind of like a quadrilingual group of people like an Avengers squad of chemistry people biologists AI folks and so that's a challenge. Um how have you guys gone about finding people convincing people to join the team and and who are your superstars?

Josh

我们在公司里对此非常务实。如果你看看创始团队,主要是 AI 研究人员。嗯,就是那些研究 Scaling(规模扩张)模型或让模型在这个领域工作的人。但随着每一代模型的出现,我们为下一个里程碑所需的人才类型也在变化,或者说可能扩大了,对吧?比如看 Chai 2,那是我们开始引进世界上一些最不可思议的抗体工程师和科学家的时刻。嗯,比如 Sonic Stream 的 Andy Young,实际上当我们雇用他时,人们问我们是不是转向了全栈药物管线,因为他们说,如果 Andy 加入你的团队,你不这么做就太疯狂了。嗯,但我认为 Andy 在过去几个月里进行的抗体活动比他在整个职业生涯中可能进行的还要多,这非常酷。团队里还有像 Nathan Roland 这样的人。Nathan 实际上是在家上学,然后很早就上了大学。所以他 14 岁就加入了 David Baker 的实验室,David Baker 因蛋白质设计获得了诺贝尔奖,18 岁开始读博士,他有很多创意。随着模型变得更好,我们需要建立产品团队,因为虽然研究人员可能让模型变得非常强大,但你需要构建正确的产品界面,让模型真正有用。嗯,那时我们开始引进那些构建了我们今天所知的一些最激动人心产品的人。比如我们的联合创始人 Jack 曾在 Stripe 工作,Munaz 是 Stripe 的前 10 名代码贡献者之一,Neil 在安全变得重要之前经营着自己的网络安全公司。随着我们向大合作伙伴部署并开始扩展,我们引进了那些在 GPU 黑客方面做了很多工作的人,以便扩展我们的系统,使它们在规模化运行时不会崩溃。前几天我们收到一封来自超大规模服务商的电子邮件或 Slack 消息,我们有一个集群出了问题,他们说“哦,GPU 太热了,我觉得你们运行太多了”,我们说“这不是重点吗?”我心想“很好,我们在做我们的工作”。至少你问怎么说服那些人加入?我认为很多还是归结于结果和明确的需求。我们以前没有在抗体设计模型中雇用抗体工程师,那些人能做什么?我们甚至担心开始这个趋势,因为对于某些下一代格式,比如复杂的抗体,如多特异性抗体,它们甚至不适用于 Chai 2。所以实际上当那些人出现时,花了几周时间模型才能达到可以处理这些有趣案例研究的程度。但幸运的是,进展足够快,能够把它带上线。所以我认为我们总是在发展团队,追求下一个里程碑。结果我们的团队也非常小。这样每个人都有点超负荷,这意味着我们必须优先排序。这迫使我们专注于最重要的事情。

So we've been really pragmatic about this at shop. If you look at the founding team, it was mostly AI researchers. Um so people who had uh worked on either scaling models or getting them to work in this domain. But really with each generation of model, uh the kind of people we've needed for the next milestone has has changed uh or I'd say probably has expanded, right? Uh so if you look at Chai 2, right? That's the point when we started to bring in um some of the most incredible like you know antibody engineers and scientists in the world. Um, uh, one of the Sonic stream, Andy, Andy Young, uh, actually when we hired him, uh, people asked us if we had pivoted into building a full stack drug pipeline because they're like, you'd be crazy not to do that if like Andy joined your team. Um, but, uh, I think Andy has has enjoyed running more antibbody campaigns in the past couple of months than he's probably run in his whole whole career. Uh, which is very cool to see. You have folks like Nathan Roland's on the team. Like Nathan was actually homeschooled and then started college very early on. So he joined like David Baker's lab who won the Nobel Prize for protein design like when he was 14. uh start his PhD when he was 18 and has so many creative ideas on our uh as the models started to get better, we needed to build up a product team uh because while the researchers might get the models to point that they're uh they're very powerful, you need to build the right product interfaces that the models are actually useful. Um and uh that's where we started to bring on uh people who have you know built some of the most exciting products we know about today. Uh like our co-founder Jack worked at Stripe, Munaz who was one of the top 10 code contributors at Stripe. uh Neil who ran his own cyber security company before security started to become very important as we deployed this to our big partners as as we started to scale up we've brought in people uh who've uh uh really done a lot of like the GPU hacking if you will in order to like scale up our systems that they they they don't break when we're running them at scale we had an email from uh or a Slack message from one of our hyperscalers the other day uh where you know we had a cluster I think that had some issue to it and they're like oh like the GPUs got too hot I think you guys are running too many and we're like isn't that the point right that's probably I was like good we're doing our job At least you ask like how do you convince those people to join? I think again a lot of it comes back to uh the results and like a clear need. We didn't hire antibody engineers before in an antibody design model like what are those folks going to do? We were even I think worried when we started that trend because for some of the um next generation formats the complicated antibodies like multisspecifics like they didn't even work with chai 2. Um so it actually took a couple of of uh weeks when some of those people showed up before the models could work at a point that they could work on some of these interesting case studies. Uh but fortunately the progress was fast enough to kind of bring that online. So I think we're always like evolving that team and going for you know that next milestone. We've got the team very small as a result too. Um so this way you know everyone is a little bit like slightly over capacity I think which means we have to prioritize. It forces us to work on the things that really matter most.

Host

我认为在研究方面,我们的创始工程师之一 Kevin Woo 拥有第一个,我认为是有史以来第一个蛋白质扩散模型。嗯,这说明了 Kevin 的执行速度。他是一位了不起的工程师,我认为从一开始工程就一直很重要。嗯,所以实际上我们的研究人员都是优秀的工程师,我们非常关心构建高质量的代码库。归根结底,我们技术上是一家软件公司。我们是 AI 研究人员,我们是蛋白质设计师,我们有很多身份,但我们的交付物是某种软件。所以我们一直保持这个高标准,同时也努力用伟大的研究人才和能够推动可能性前沿的人来平衡。

I think on the research side as well one of the founding engineers Kevin Woo he had the first uh I think it was the first protein diffusion model like ever. Uh and that speaks to Kevin's speed of execution. like he is a heck of an engineer and like I think engineering has just always been like important since day one. Um so like really even our researchers are all excellent engineers and like we really care about building a high quality code base. Like at the end of the day we are technically like a software company. We're we're AI researchers. We're we're protein designers. We're we're a lot of things but uh our deliverable is some piece of software. So we've kept that like kept that bar really high while also trying to like you know level that with great research talent and people that can actually like push the frontiers of what's possible.

顿悟时刻与分子质量 Aha moments and molecule quality

Host

你们在这个领域似乎有过很多“啊哈”时刻,比如 AlphaFold 是一个“啊哈”,我们能弄清楚蛋白质如何折叠,然后扩散模型是“啊哈”,我们能生成蛋白质。嗯,似乎你们最新的模型是另一种“啊哈”时刻,你们不仅生成看起来像蛋白质的分子,而且它们还具有治疗特性,比如能非常紧密地结合。嗯,它们有很高的命中率。你能谈谈你们的模型现在产生的分子质量,以及为了做到这一点,这些模型背后投入了什么吗?

And you've had a number of it seems like aha moments in the field like alpha fold was an aha we can figure out how a protein folds and then the diffusion models aha like we can generate proteins. Um, and it seems like your latest models are a kind of another aha moment where you're not only generating molecules that look like proteins, but they also have therapeutic properties like they can bind really tightly. Um, they have really high hit rates. Can you talk sort of about the quality of the molecules that your models are producing now and everything that sort of went in to those models to make them able to do that?

Josh

是的,如果我们看看产生的分子质量,这要追溯到我们创办公司时的论点之一,我们真的想专注于分子的从头生成。如果你看看当时很多药物发现 AI 的工作,都是关于如何拿一个分子,让它变得更好一点,对吧?我们刚才在某种意义上谈过这个,但那是通过实验室在环的方式,对吧?我拿一些数据,试着用那种方式改进。我们开始的问题是,我们能不能真的在计算机上做所有这些,有没有办法,也许有足够的数据,或者我们可以收集足够的数据,这样我们就可以零样本生成一个具有很多这些特性的分子。所以要做到这一点,我们需要做的第一件事是设计具有非常高成功率的分子。当我们创办公司时,抗体设计的最先进水平大约是 0.1% 的结合率。也就是说,你设计的分子中只有千分之一能在实验室中实际结合。嗯,所以首先这意味着你必须筛选大量分子才能找到好的。

Yeah, if we look at the the quality of the molecules that's coming out, if it goes back to, you know, one of the thesis when we started the company that uh we really wanted to focus in on this like denovo generation of the molecules. If you look at what a lot of the drug discovery AI work was at the time, it was about how do I take a molecule and just make it a little bit better, right? Which we just talked about in a sense, but it was doing it with a lab in the loop uh style, right? Where I take some data, try to make it better that way. And the the question uh we started with is well can we actually just like do all that on the computer right is there a way that the there maybe there would be enough data or we could collect enough data um so that we could just zero shot a molecule that has a lot of these properties. So the first thing we needed to do to get there um was to design molecules with really high success rates. When we started the company the state-of-the-art for antibbody design was about like a 0.1% binding rate. So one in a thousand of the molecules you design would actually bind in the lab. Um so first of all that means you have to screen a lot of molecules to find some good ones.

提升命中率与扩展 Improving hit rate and scaling

Josh

这也意味着你在这个过程中的梯度其实很弱。所以对很多靶点,你不会得到任何命中。对于你确实得到一些命中的靶点,你也没有足够的数据来判断是否具备类药性质。所以我们真正专注于如何让这个过程更准确。用我们的 Chai 2 模型,我们达到了大约 15% 的成功率。所以现在如果你筛选一千个分子,你能得到 150 个回来。现在你可以开始获得一些关于分子性质的有趣统计数据。这让我们能够在此基础上迭代,在实验室里建立正确的评估,并真正尝试在这方面爬山。我们现在到了一个可以真正把很多这些不同性质从一开始就融入引擎的阶段。然后也许你可以想想这是如何发生的,或者我们是如何处理这个问题的,以及为什么我们认为它会继续改进。

It also means that the gradient you get on your process is actually quite weak. So for many targets you won't get any hits. For the ones where you do get some hits, you won't have enough to actually see whether you're getting drug-like properties. So we really focused on how to make this process more accurate. With our Chai 2 model, we got about a 15% success rate. So now if you screen a thousand molecules, you're getting 150 back. Now you can start to get some interesting statistics on the properties of the molecules. And that allowed us to iterate on that, build the right evaluations around that in the lab, and really try to hill climb that as well. And we're getting to a point now where we can actually bake in a lot of these different properties from the start into this engine. And then maybe if you think about how this happens or how we're approaching this, and why we think it's going to continue improving.

Host

是的。说实话,我觉得 Josh 和我都对这工作进展之快感到惊讶。我们最初做预算时,以为三四年内能有 20% 的命中率就不错了。我们当时的目标是 1%——我们觉得百分之一就很了不起了。我们觉得这会是开创性的。我们有一套理念和方法,结果效果非常好。这种方法与机器学习其他领域成功的方法非常相似。人们把生物学当作一个定制问题或定制领域,但实际上它与自动驾驶或语言模型的原则相同。我们之前讨论过的一件事是你如何找到缩放定律。将一段文本 token 化是一回事,将生物学 token 化又是另一回事。你能简单谈谈这个挑战以及你们是如何解决的吗?

Yeah. So honestly, I think Josh and I were both surprised at how quickly this worked. When we were originally budgeting this, we thought maybe 20% hit rate in three or four years. We were really shooting for 1%—we thought one in a hundred would be amazing. We thought this would be a groundbreaking thing. We had a philosophy and an approach we wanted to take, and it just ended up working really well. That approach is very similar to what's worked in the rest of machine learning. People treat biology as this bespoke problem or bespoke field, but really it's the same principles as self-driving or LM principles. One of the things we've talked about before is how you manage to find scaling laws. It's one thing to tokenize a string of text, it's another thing to tokenize biology. Can you say a couple of words about that challenge and how you guys solve it?

Josh

是的。我非常喜欢 Chai 的一点是,我们是一家非常信奉“苦涩的教训”和经验主义的公司。所以我们坚信要 Scaling 数据、Scaling 模型、Scaling 算力。要做到这一点,显然你需要找到缩放定律,否则你只是在浪费时间和资源。而且我认为,在不透露太多的情况下,我们的一大指导原则是简单。所以当你看到一个像 Chai 1 这样的模型时,Chai 1 中有 23 个不同的子模块。当你试图在这样的模型上迭代时,会变得非常困难,因为你需要独立理解每个子模块,理解它们所有的行为和动态,这不太容易扩展。所以你可以开始思考,好吧,我如何简化它?我如何确定什么才是真正重要的?一旦你有了这些,整个研究过程和识别这些 Scaling 方向就会变得简单得多。

Yeah. So one thing that I really like about Chai is that we're a very bitter lesson and empirical company. So we really believe in scaling data, scaling models, scaling compute. To do that, obviously you need to identify scaling laws, otherwise you're just wasting time and resources. And I think, without giving away too much, one of our big guiding principles is simplicity. So when you look at a model like Chai 1, there are 23 distinct submodules in Chai 1. When you're trying to iterate on something like that, it gets really hard because you need to understand each of these submodules independently, understand all their behaviors and dynamics, and that doesn't scale well. So you can start to think, okay, how do I simplify this? How do I identify what's really important? Once you have that, the whole research process and identifying these scaling directions becomes a lot simpler.

Host

我认为另一件有趣的事情是,我们从深度学习其他领域成功的经验中汲取了很多教训。但正如你指出的,数据本身是不同的。Chai 的模型完全是从零构建的。我们不是在一些蛋白质数据上微调 GLM 之类的模型。我们是从头开始构建一切。不过我认为,公司建设过程中的很多工作,是采取一种理念并真正坚持下去,在此基础上迭代,并拥有一些指导原则。当你建立一家蓝天研究公司时,Chai 几乎就像这些新实验室之一,对吧?我们正在攻克一个重大的人工智能问题,随着我们取得进展,这为我们的客户带来了机会。但如果你要从事如此开放的工作,你需要一些原则来指导你。而且我认为我们在公司中贯彻这些原则方面做得非常好。所以 Matt 提到的很多事情……

One of the other things I think is interesting about this is we take a lot of these lessons from what's worked in the rest of the deep learning space. But as you're pointing out, the data itself is different. The models at Chai are completely built from scratch. We're not fine-tuning GLM or something like that on some protein data. We built everything from the ground up. I think a lot of the company building process though is taking a philosophy and actually sticking with it and iterating on that, and just having some guiding principles. When you're building a blue sky research company, if you will, Chai is almost like one of these Neo Labs, right? We have this big AI problem we're going after, and as we make progress on it, that opens up opportunity for our customers. But if you're going to work on something so open-ended, you need some principles to guide you. And I think we've done a very good job on tracking those principles in the company, working on it. So a lot of the things that Matt...

Host

你能分享一下吗?指导原则是什么?

Can you share them? What are the guiding principles?

Josh

所以我认为简单是 Matt 提到的原则之一。还有这种“苦涩的教训”式的 Scaling 算力、数据和模型。还有就是要非常严谨。这在领域内非常重要。在生物学中你很容易自欺欺人——湿实验室的误差线实际上也很大。所以这与代码生成等领域有些不同。比如看 SWE-bench,大约一年前,人们会说“哦,我达到了 16%,然后 17%,然后 18%”。但在生物学中,如果你的实验室误差是正负 5%,那可能都是一样的。所以这实际上意味着,就你希望看到的模型阶跃变化而言,门槛非常高,但你也需要对自己是否在进步非常诚实。所以你可以提出一个花哨的模型,看起来在一两个新任务上表现不错,但如果你真的想构建一个能推动领域前进的产品,那么展示它在更广泛范围内有效是非常重要的。

So I think simplicity was one of the ones that Matt mentioned. It's this bitter lesson-ness of scaling compute and data and models. It's being really rigorous. That's something that's so important in this space. You can fool yourself so easily in biology—the error bars in the wet lab are actually quite large as well. So it's actually a little bit different than if you look at code generation, for instance. If you look at SWE-bench, people were like, maybe a year ago, people were like, "Oh, I got 16%, then 17%, then 18%." I mean, in biology, if you're plus or minus 5% in your lab, that might all be the same. So it actually means the bar is really high in terms of the step changes you want to see with the models, but you also need to be really honest with yourself about whether you're making progress or not. So you could come up with some fancy model that looks like it works well on one or two new tasks, but it's very important to show that it works more generally if you're actually trying to build a product that can bring the field forward.

Host

是的。我觉得这很有意思,因为生物学是那些本质上不太可验证的领域之一,而你们在向客户和像我们这样对生物学知之甚少的人展示进展方面做得非常好。那么我们能谈谈评估和模型进展中可验证的部分吗?你们怎么知道你们的模型在变得更好?

Yeah. And that's I think pretty interesting because biology is one of those inherently not-so-verifiable domains, and you guys have been really good at sort of showing your progress to customers and to people like us who know very little about biology. So could we talk a little bit about the evals and the verifiable part of the model progress? How do you guys know that your models are getting better?

Josh

嗯,首先我想说,我实际上认为这是更可验证的领域之一。它实际上是一个非常客观的读数。即使看代码生成,对吧?也许它是一个可验证的任务:我的代码编译了吗?它解决了这些单元测试吗?但你如何考虑品味呢?我是否写了一些无法维护的草率代码?那是什么样的?当我们在实验室设计分子时,我们实际上可以非常具体地说明许多这些性质。所以也许我们得到了一个与靶点结合的分子,但我能制造它吗?那可能是你版本的技术债务,但你可以衡量它。而且我认为这些评估实际上使这个领域更可验证。也许验证需要更长一点时间——不像 5 秒钟就能得到读数并运行单元测试。

Well, I would say first of all that I actually think this is one of the more verifiable domains. It's actually a very objective readout. If you look at something even like codegen, right? Maybe it's a verifiable task: did my code compile? Did it solve these unit tests? But how do you think about the taste, right? Did I write some really sloppy code that can't be maintained? What does that look like? When we think about designing a molecule in the lab, we can actually be quite specific about many of these properties. So maybe we got a molecule that binds to the target, but can I manufacture it? That might be your version of tech debt, but you can measure that. And I think those evals actually make this domain more verifiable. Maybe it takes a little bit longer to validate it—it's not like 5 seconds to get a readout and run a unit test.

模型进展与硬目标 Model Progress and Hard Targets

Host

是的。当我们谈论进展和模型有多好时,生物学中有一类靶点,正如你提到的,可以通过传统筛选方法解决。这些方法耗时很长,非常缓慢且原始。然后还有一些靶点用现有方法根本无法触及,无论出于什么原因,它们都是不可成药的。那么,就模型进展而言,我们在哪里?一方面是在现有靶点上更快地生成分子;另一方面是解锁新的生物学、新靶点,以及我们以前无法成药的东西。

Yeah. When we talk about progress and how good the models are, there's like a domain of targets in biology that you can, as you mentioned, sort of address with traditional screening methods. They take a long time. They're very slow and rudimentary. And then there are targets that just aren't addressable with existing methods. They're not druggable for whatever reason. So where are we in terms of model progress, in terms of working on existing targets and generating molecules faster? That's one end of the spectrum. And then the other end of the spectrum is unlocking novel biology, new targets and things that we couldn't drug before.

Josh

这其实回到了我们的核心建模理念。实际上有很多次我们都在想,天哪,如果我们有第 24 个模块,我们就能解锁那个新靶点。但我们又会问,这真的是我们想要长期维持的吗?这是渐进式的改进,还是真正的复利式进步?这真的能帮助我们泛化到那一整类我们目前无法触及的靶点吗?所以我们的理念一直是,好吧,让我们继续专注,找到你的缩放定律。总有一个临界点,如果模型能把损失压得更低,它就必须理解它所处理的靶点的一些非常本质的东西。举个例子,前几天我和 Paul 在讨论,也许我们看待蛋白质上某些糖基化位点的方式,我们想,哦,也许我们该用不同的方式来表示它们。但我们又想,即使我们不这样表示,也会有某些序列基序告诉模型这里应该有一个糖基化位点,以便进一步降低损失。模型应该自己学会这一点。所以靶点有所有这些隐藏的特征,如果你真的相信缩放定律,你相信模型会达到那个水平,那么这些类型的靶点就应该随着更好的模型而解锁。当然,你仍然必须非常认真地对待这件事,仍然需要所有适当的验证。你需要真正挑战自己,确保这确实有效。但我认为我们的方法一直是,有了更好的模型,我们应该能够解锁很多这样的靶点。

This kind of goes back to our core modeling philosophy. There have actually been plenty of times where we're like, man, if we had a 24th module, we can actually unlock that new target. And we're like, is that really something that we want to maintain long term? Is this incremental or is this actually a compounding improvement? Will this actually help us generalize to the broader class of targets that we really can't hit? And so our philosophy has been like, all right, let's just continue to focus, identify your scaling laws. There comes a point where if the model is able to push loss down even further, it has to understand something very intrinsic about the target that it's operating on. So one example we were talking about the other day, Paul and I, maybe the way that we're looking at certain glycosylation sites on proteins, we're like, oh, we might want to represent them differently or something like this. And we're like, well, even if we didn't represent them, there are certain sequence motifs that will tell the model there should be a glycosylation site here in order to drive loss down further. The model should just have to learn that. So there are all these hidden features of targets where if you really believe in scaling laws, you believe the models will get there, these types of targets should just unlock with better models. And of course, you still have to take this very seriously and you still need all the proper validation. You need to really challenge yourself and make sure that this is truly working. But I think our approach has always been that with better models, we should be able to unlock a lot of these targets.

硬目标的跨学科性 Interdisciplinary Nature of Hard Targets

Josh

另一个有趣的事情是,如果我们看看 CHI 的不同团队,每个团队所说的“硬靶点”实际上都不一样。在科学团队,可能是一个不可成药的 GPCR 靶点,还没有人能在功能上调节它。在机器学习研究团队,可能是模型就是无法折叠这个东西,不知道它长什么样。然后在产品团队,可能是,哦,我有所有这些修饰,比如糖基化,而且它是一个膜蛋白,我该怎么向用户展示它?实际上,每个团队对硬靶点的定义不同,我认为这是一个特性,而不是一个缺陷。这意味着如果我们想要取得广泛的进展,每个人都在以不同的方式并行推进,这意味着我们总能取得平稳的进展。在 CHI 通常不会只有一个瓶颈。不是那种,哦,如果我们多一个模块就能解决问题,或者如果我们以某种方式把它推向产品就能解锁。我们正在努力构建一个统一的解决方案,因为归根结底,公司的目标是构建一个分子计算机辅助设计套件。不是要制造一两个分子,也不是要推出五条有趣的治疗管线。而是要改变药物发现的方式。如果我们要做到这一点,我们就必须攻克所有硬靶点,无论你如何定义“硬”。

One of the other interesting things is if we look at the different teams at CHI, what people will call a hard target is actually different in literally every team. So on the science team it might be an undruggable GPCR target. No one's gotten something that has modulated that in a functional way. On the ML research team it'll be something like, oh, the model just can't fold this thing up. It doesn't know what it looks like. And then on the product team it might look like, oh, I've got all these modifications like my glycosylations and it's a membrane protein, how do I represent that to the user? And actually the fact that it's different for each of these groups I think is a feature rather than a bug. It means that if we want to make broad progress, everyone is kind of pushing in parallel on these different ways, and that means that there's very smooth progress that we can make all the time. There's usually not one bottleneck at CHI. It's not like, oh, if we only had that one extra module things would work, or if we only tried to push this into the product in some way we could unlock it. We're trying to build this unified solution because at the end of the day the goal of the company is to build a computer-aided design suite for molecules. It's not to make one or two molecules. It's not to get a pipeline of five interesting therapies that we bring to market. It's to change the way that medicines are discovered. And if we're going to do that, we need to work on all the hard targets regardless of how you define hard.

分子计算机辅助设计套件 Computer-Aided Design Suite for Molecules

Host

是的。请多谈谈这个分子计算机辅助设计套件的想法。那是什么意思?

Yeah. Say more about this idea of computer-aided design suite for molecules. What does that mean?

Josh

所以,从核心上讲,这回到了让生物学更像一门工程学科的观点。我们不是去大型文库中捞出一个分子,也不是做大量的试错。你希望能够预先指定设计分子的原则,然后有一个引擎能够将其实现为我们将在实验室中创造的实际分子。你看,你可能仍然需要在实验室和模型上进行一些迭代,因为也许你的假设是错的。但我们想要加速的是,再次拥有那个计算机辅助设计套件,这样你就可以很快地从想法变成可测试的假设。如果这个循环现在需要大约 9 个月来发现你的分子,而如果它需要 9 周或 9 天,那么每一个数量级的提升都会极大地扩展你能真正筛选的想法数量。我认为这最终是领域走向更好药物的方式。

So, at its core, it goes back to this point about making biology look more like an engineering discipline. So, we're not going and fishing something out of a large library or doing a ton of trial and error. You want to be able to specify upfront the principles that go into designing your molecule and then have an engine that can actually realize that into some molecule that we're going to go and create in the lab. And look, you still might do some iteration on the lab and on your model because maybe your hypothesis was wrong. But what we want to speed up is actually again have that computer-aided design suite so that you can go from idea to testable hypothesis very quickly. And if that loop right now takes something like 9 months, I don't know, to go and discover your molecule versus if it takes 9 weeks or it takes 9 days, each order of magnitude just scales in a very big way the number of ideas you can really sort through. And I think that's ultimately how the field is going to converge on better medicines.

行业未来 Future of the Industry

Host

这确实回到了人们有时会讨论的问题:我们关心的是速度还是靶点的难度?在某个时刻,它们也会以这种方式汇聚。因为一个硬靶点,如果我们能更快地迭代假设,那么也许就更容易攻克它。所以,如果我们能稍微脱离当前现实,走到足够远的未来,我们不是在从现在向前思考,而是从未来向回思考,2035 年、2040 年、2100 年,随你选。这个行业会是什么样子?让我们想象一下,分子计算机辅助设计套件已经成为标准。让我们想象一下,大量的创新已经向下游流动到流程中一些湿实验室的部分。你能描绘一下 10 年后这个行业可能的样子吗?

It really comes back to like people sometimes talk about do we care about speed or do we care about the difficulty of the targets. At some point they converge in this way as well. Because a hard target, if we can iterate through hypotheses a lot faster, then maybe it'll be easier to crack it. So if we can maybe detach ourselves from the present reality and go far enough into the future that we're not thinking present forward, we're actually thinking future back, 2035, 2040, 2100, whatever you want. What does the industry look like? Let's imagine that computer-aided design suite for molecules has become a standard. Let's imagine a lot of innovation has flowed downstream to some of the wet lab parts of the process. Can you paint the picture of what the industry might look like 10 years from now?

Josh

是的,我认为那将是一个真正激动人心的时刻。我们可以从几个角度来看。首先,我们开发的药物质量有望提高。今天我们进入临床的很多分子都很难发现。你找到一个完成了 80% 的分子,也许我还是会推进它,以满足我的时间表。它可能会让一些患者受益。

Yeah, I think it's going to be a really exciting time. And we can look at this on a couple of angles. So, first of all, the quality of the medicines that we develop will hopefully go up. There's a lot of molecules we put into the clinic today that are really hard to discover. You find something that's like 80% of the way there, maybe I'm just going to advance it anyways to hit my timelines. It's probably going to benefit some patients.

药物发现未来 The Future of Drug Discovery

Josh

但然后你会被击败,你知道,一年后又被别人超越,这真的不是行业里最有效的资源利用方式。有些疾病今天还是太难攻克。人们一直在尝试治疗阿尔茨海默症,但不幸的是,进展不如我们期望的那样。有些东西今天去攻克并不经济。想想个性化药物、罕见病,这些患者群体可能较小的领域。但同样,如果我们能更快地迭代这些想法,如果我们能更快地推出产品,如果我们能以更低的成本做到,那么这些也可能变得可行。所以我认为我们可以在很多不同的方向上发力。这意味着未来真的非常光明。

But then you get beat, you know, a year later by someone else, and it's really just not the most efficient spend of resources in the industry. You've got the kind of diseases that are just too hard to go after today. People have been trying to drug Alzheimer's forever, and unfortunately, you know, haven't made as much progress as we'd like. You have things that just aren't economical to go after today. Think about personalized medicines, rare diseases, things where maybe the patient population is going to be smaller. But again, if we can iterate through those ideas faster, if we can launch something faster, if we can do it in a less expensive way, then those probably come into reach as well. So I think there are just so many different axes that we're able to push on. And I think that means that the future is really bright.

Host

过去你要么想成为同类首创,要么想成为同类最佳。是的。现在你更想成为同类最后,因为你实际上只想成为最终答案。所以我认为你会看到很多的是,在设计药物类型时会有更多的针对性。比如这种药会非常针对目标疾病。它不会有今天很多药物都有的某些相互作用,比如某些负面相互作用。很多这些东西实际上你都能在计算上建模,也许今天还不行,但肯定有通往那里的路径,我认为这是最让我兴奋的事情之一。

It used to be like you either want to be first in class or best-in-class. Yeah. Now it's like you want to be last in class, because you actually just want to be the final answer. So I think a lot of what you'll see is just way more intentionality in the types of drugs that you're designing. Like this drug will be super specific to the disease of interest. It won't have certain interactions, like certain negative interactions that a lot of drugs today do. A lot of this stuff is actually like you're able to model most of this computationally, like maybe not today, but there's definitely a path towards getting there, and I think that's one of the most exciting things for me.

Host

非常酷。我们能谈谈你们做的商业模式决策吗?因为我认为很多时候人们认为 Chai 和 Isomorphic 是同一类公司。Isomorphic 在开发药物。你们是在让现有行业更高效、更好、更快、更便宜地开发药物。为什么你们决定走那条路,而不是 Isomorphic 那条路?

Very cool. Can we talk about a business model decision that you guys made? Because I think a lot of times folks think about Chai and Isomorphic in the same neighborhood. Isomorphic is developing drugs. You guys are enabling the existing industry to develop drugs more efficiently, better, faster, cheaper than they have before. Why did you decide to go down that path versus, you know, the Isomorphic path?

Josh

是的,首先,我认为这两条路都很棒,它们都能创造巨大的价值。我们一直对为行业构建基础设施感到非常兴奋。我们开始时的赌注,回到几年前在 Matt 家看到的那些结果,就是未来大多数药物将以这种方式被发现。如果是这样,就需要有人去构建那个基础设施来实现它。我认为部分原因也是我们很多创始团队成员,包括 Matt 和我,之前都在公司里构建过全栈药物管线,对吧?我们当时在构建 AI 模型,把那些药物推进临床。同样,我们也得到了惨痛的教训。我们的想法是,随着模型变得更好,我们想把更多的钱花在制造更好的模型上,而不是把那些资源转移到临床试验之类的事情上。我们商业模式的一个有趣之处在于,随着模型变得更好,它实际上为我们赢得了继续投资于它们的权利,对吧?你与这些制药公司的合作伙伴关系今天正在产生回报,并允许我们继续投资。所以从这个角度看,这是一个更具可扩展性的业务。我只是觉得我们能为世界创造的最终影响要大得多。我们一直想与生态系统广泛合作。这又回到了在生物学中自欺欺人的观点。如果你只研究少数几个药物靶点,你可能会设计出最精妙的分子,为世界创造大量价值。但你可能会只见树木不见森林,因为也许你的模型不能泛化到那一年人们将要制造的其他 500 个分子。如果你去与人合作,你就不能自欺欺人。比如看看我们的合作伙伴礼来、诺华、Organics、辉瑞,这些公司不会把东西视为理所当然。你必须真正在这些合作中交付成果,他们才会认真对待你。这意味着它不能只是有时有效。比如当我们在 Chai 发布模型时,它们必须真正有效。它们必须为我们的合作伙伴带来价值。而不是说,哦,我们做了一个分子,它不完全有效,我们让化学家稍微清理一下。所以这几乎是一个更难实现的业务。我认为这就是为什么你很少看到有人追求这个模式的原因之一。如果你再次致力于药物管线,比如以后可能有办法修复问题,当然会有临床结果的证明,但当你采用这种基于合作的模式时,你必须对自己的模型非常严谨,你的模型必须真正有效,否则那些合作伙伴不会轻易来。所以我认为这在很多方面让我们的生活更艰难,但我认为如果我们能成功,这也是更有回报的道路。

Yeah, first of all, I think both of these paths are great, and they can create tremendous value. We've always been really excited about building infrastructure for the industry. Our bet when we started, going back to those results in Matt's house, you know, a few years ago, was that this is how most future drugs are going to be discovered. And if that's the case, somebody needs to go and build that infrastructure to make it happen. I think part of this too is a lot of our founding team, including Matt and I, had worked in companies before where we had built these full-stack drug pipelines, right? And we were building AI models, we were putting those drugs into the clinic. Again, we're pretty bitter lesson as well. And our thinking was, as the models get better, we want to be spending more of the money making better models as opposed to diverting those resources into clinical trials and things like that. And one of the interesting things about our business model is, as the models get better, it actually wins us the right to continue investing more in them, right? And you have partnerships with these pharma companies that are paying off today and allow us to continue investing in it. So it's a much more scalable business in that way. And I just think about the ultimate impact that we can create for the world is a lot larger. We've always wanted to just partner broadly with the ecosystem. It goes back to this point about fooling yourself in biology. If you work on a small number of drug targets, you might come up with the most exquisite molecules, creating a ton of value for the world by doing that. But you might miss the forest for the trees, because maybe your model doesn't generalize to the other 500 molecules that people are going to make that year. And if you go and partner with people, you just can't fool yourself. Like you look at our partners Eli Lilly, Novartis, Organics, Pfizer, like these are not companies that take the stuff for granted. You have to really deliver on these partnerships for them to take you seriously. And that means it can't just work some of the time. Like when we ship models at Chai, they really have to work. They have to deliver value to our partners. And it's not like, oh, we made some molecule, it doesn't fully work, we're gonna have our chemist clean it up a little bit. So it's almost a harder business to pull off. I think that's one of the reasons why you haven't seen it pursued many times. If you work on a drug pipeline again, like there might be ways to fix things up later, there will be the proof of like what happens in the clinic of course, but when you have this partnering-based model, you have to be really rigorous about your models, your models have to work really well, because otherwise those partners are not going to come easily. So it's made our life harder, I think, in many ways, but I think it's also the more rewarding path if we can get it to work.

Host

那么,你学到了什么?我的意思是,你知道,你算是走出了实验室,在现实世界中,为真正的制药公司交付实际价值。当你开始与这些合作伙伴合作时,你学到了什么?有没有什么惊喜,比如他们的需求与你预期的不同,或者他们在这方面的专业水平与你预期的不同?与这些合作伙伴合作有没有什么惊喜或收获?

Well, and what have you learned? I mean, you know, you're out of the lab, so to speak, and in the real world, you know, delivering real value for actual pharma companies. What have you learned as you start to work with these partners in terms of any surprises in terms of how their needs might have differed from what you expected, or how their level of sophistication around this might be different than what you expected? Have any surprises or any learnings from working with these partners?

Josh

是的,我们进入这些合作时,我想很多人告诉我们制药公司不知道如何使用 AI。这不是一个技术前沿的行业之类的。但说实话,我们的经历并非如此。我认为这些合作伙伴非常严谨,非常严格的客户,对吧?在他们开始部署这些东西之前,他们会测试我们的每一个说法,对吧?但当他们看到数据时,他们会全力以赴,对吧?因为制药是一个创新行业。我觉得仅仅思考制药行业的整体经济学就很有趣。如果你在制药行业开发一个产品,比如一种药,你只能在特定时期内拥有该药的独家经营权,对吧?你必须继续再投资。礼来现在是一家万亿美元的制药公司。如果他们得不到更多重磅药物,他们不会永远是万亿美元的制药公司。我认为这迫使这些公司真正在采用新技术、部署新技术并努力保持领先方面全力以赴。制药是一个竞争非常激烈的领域。有很多人试图将这些药物带给患者,顺便说一句,这对患者来说很棒,但这意味着你必须在这个领域保持领先。

Yeah, so we went into these partnerships, I think a lot of people told us that pharma doesn't know how to use AI. These are not, it's not like a tech-forward industry and things like that. And to be honest, that hasn't really been our experience. I think that these are again very rigorous partners, very rigorous customers, right? And they're going to test every one of our claims, right, before they start to deploy these things. But when they see the data, they go all in, right? Because pharma is an innovation industry. I think it's interesting just to think about even the whole economics of the pharma industry. If you build a product in pharma, right, like a drug, you only have exclusivity on that drug for a certain period of time, right? And you have to continue to reinvest. Eli Lilly is a trillion-dollar pharma company right now. If they don't get more blockbuster drugs, they will not be a trillion-dollar pharma company forever. And I think that forces these companies to really be on their game of adopting new technologies and deploying them and trying to stay ahead. Pharma is a very competitive arena. There's a lot of people trying to bring these drugs to patients, which by the way is great, great for patients, but it means that you have to be on top of your game here.

采用与激励机制 Adoption and Incentive Structure

Josh

而且我认为这意味着,一旦你跨过那道门槛,模型开始运转,我们就看到采用率迅速上升,人们开始以极具创造性的方式思考如何使用这些模型。

And I think that means that once you're through that door and your models are working, we've seen an upscale in adoption very quickly, and people thinking about how to use the models in incredibly creative ways.

Host

我也非常喜欢由此产生的激励机制。更多地采用合作伙伴模式,随着我们的模型变得更好,我们也能为客户带来更好的结果,等等。所以我认为这对我们来说也是一个很好的附带效应。

I really like the incentive structure that it creates as well. Taking more of the partnership model, as our models get better, we get better results to our customers, and so on. So I think that's just a nice side effect for us as well.

Josh

Chai 一直非常专注。而且确实,随着模型变得更好,你可以用更好的数据迭代出更好的模型,等等。我们看到很多我们的合作伙伴用模型来生成结合剂、抗体等等,我们也在内部大量进行“吃自己的狗粮”式的测试。我们有一个完整的科学团队在使用这些模型,试图理解它们能做什么。很多反馈都归结为:既然我们解锁了用例 X,那我们可以生成什么类型的新数据?我们如何以此让模型变得更好?而这正是 Chai 的长期愿景。我们从不认为我们会构建一个能解决所有问题的单一模型。我们知道,就像在其他领域一样,这个模型会有多次迭代,随着你得到越来越好的模型,你会产生这种飞轮效应,我想。

Chai has been incredibly focused. And really, as the models get better, you can iterate on a better model with better data, and so on. A lot of what we see and what our partners are using the model for—like generating binders, antibodies, and so on—we're also doing a lot of dogfooding in-house. We have a whole science team that's using the models and just trying to understand what they can do. And a lot of that comes back to: now that we've unlocked use case X, what type of new data can we generate? How can we make the model better that way? And really, that's been the long-term vision of Chai. We never thought we're going to build this one model that's going to solve all of this. We know that, like in other fields, there are going to be multiple iterations of this model, and as you get better and better models, you get this flywheel effect, I guess.

Host

嗯,你提到了我们能生成什么数据。我们能稍微谈谈数据吗?你们显然不能直接去网上抓取所有构建模型所需的数据。数据从哪里来?它如何随时间累积?你能简单谈谈数据作为模型输入的情况吗?

Well, you mentioned what data can we generate. Can we touch on data for a minute? You guys can't exactly just go scrape the internet and have all the data you need to build your models. Where does the data come from? How does it compound over time? Can you just say a word about data as an input to your models?

数据来源与方法 Data Sources and Approaches

Josh

是的,所以我认为数据的主要来源,或者说黄金标准的数据来源,是蛋白质数据库。实际上,自 1970 年代以来,真正的实验室科学家们一直在向其中提交蛋白质和其他分子的晶体结构。如果没有这些,结构预测和设计就不会存在。所以这是结构数据的一个来源。当我进入这个领域时,我确实采用了基于结构的方法。我对预测结构非常感兴趣——你如何考虑一个能输出 3D 坐标的机器学习模型?那看起来不像 LLM,也不像图像模型。这真的是一个独立的类别。有趣的是,Josh 采取了完全相反的方法。他是 ESM 的原始作者之一,那是在理解如何将语言模型应用于蛋白质序列方面的一项开创性工作。这项工作的真正酷之处在于,如果你能训练一个语言模型来理解蛋白质序列,最终它会内部表示 3D 结构。这发生的原因非常有趣。为了预测缺失的氨基酸——就像你预测句子中的下一个词一样,我想预测蛋白质中的下一个氨基酸——要有效地做到这一点,你真的需要了解该氨基酸的即时微环境是什么样的,因为这告诉你与周围环境兼容的氨基酸是什么。要做好这一点,你需要了解蛋白质的 3D 形状。所以我采取了这种真正基于结构的方法,而 Josh 比我更极端——他说,我们只需要看序列,这就会自然涌现。所以,数据的两个主要来源是:用于结构的蛋白质数据库,以及这些庞大的、可能达到数万亿 token 规模的序列数据库。一旦你有了真正好的模型,你可以做的就是将它们运行在序列数据库上,以获得新的结构。所以,你再次看到这种复合效应:随着你的模型变得更好,它们在预测这些结构方面越来越准确,然后你就有越来越多的训练数据用于下一代模型。

Yeah, so I think the primary source of data, or the gold standard source of data, is the protein database. It's actually just legitimate lab scientists who, since 1970, have been depositing crystal structures of proteins and other molecules over time. And really, without that, structure prediction and design wouldn't have been a thing. So that's one source of structural data. When I started in the field, I really took a structure-based approach. I was super interested in predicting structure—how do you think about a machine learning model that can output 3D coordinates? That doesn't look like an LLM. It doesn't look like an image model. This is really in its own class. Josh, interestingly, was taking the exact opposite approach. He was one of the original authors on ESM, and that was a really seminal work in understanding how to apply language models to protein sequences. What's really cool about that work is if you can train a language model to understand protein sequences, what ends up happening is it ends up representing the 3D structure internally. And there's a really interesting reason why this happens. In order to predict missing amino acids—the same way you'd predict the next word in a sentence, I want to predict the next amino acid in a protein—to do that effectively, you really need to understand what that amino acid's immediate microenvironment looks like, because that tells you what compatible amino acids are with everything surrounding it. And to do that well, you need to understand the protein's 3D shape. So I was taking this really structure-based approach, and Josh was even more bitter than me—he's like, we're just going to look at the sequences and this is just going to emerge. So yeah, the two main sources of this data are the protein database for structures and then these massive, maybe even on the order of trillions of tokens, sequence databases. And what you can do once you have really good models is just run them on the sequence database to get new structures out. So again, you have this compounding effect: as your models get better, they get more and more accurate at predicting these structures, and then you have more and more training data for the next series of models.

范式与数据生成 Paradigm and Data Generation

Host

有时候人们会问我们 Chai 现在到底在追求哪种范式?我喜欢我们团队的一点是,实际上我们两者都不是。我们非常务实;我们想解决问题。我们并不真正关心——是序列方法还是结构方法?在实践中,最终当然会是两者兼而有之。我想你会感到惊讶,但互联网上的生物序列 token 可能比英语语言 token 还要多。现在,很多这些数据并不是那么有用。它们可能不是冗余的,可能非常嘈杂,但确实有大量数据存在,而将它们整合起来需要很多技巧。而且我认为令人兴奋的是,随着模型已经发展到我们可以在实验室中设计东西的程度,比如针对新靶点,我们实际上也可以使用模型来生成数据。所以我们在 Chai 做的所有实验产生了大量“废气”,这些也有助于让模型变得更好。所以这与我们在 LLM 中看到的起飞类似。比如当我在 OpenAI 时,我们研究 GPT-1 架构的强化学习,但没有成功,因为模型不够好。但一旦基础模型足够好,你就可以开始做这类实验。我认为在我们的领域现在也开始出现类似的类比,模型已经达到了一个点,人们对数据重新产生了兴趣,以及我们如何真正让模型参与到实现这一目标的过程中。我认为这创造了另一个关于模型复合改进的非常有趣的循环。

Sometimes people ask us at Chai about these days—which paradigm are you actually going after? I think one of the things I love about our team is that it's actually just neither. We're very pragmatic; we want to solve the problem. We don't really care—is it a sequence approach? Is it a structure approach? In practice, it's going to end up being both, of course. I think you'd be surprised, but there might even be more biological sequence tokens on the internet than English language tokens. Now, a lot of that data is not that useful. It might not be redundant. It might be very noisy, but there's a lot of data out there, and there's a lot of art in bringing it together. And I think the exciting thing is that as the models have gotten to a point now where we can design things in the lab, right, to new targets for instance, we can actually use the models to generate data as well. So there's a lot of exhaust from all the experiments that we're doing at Chai, which also help to make the models even better. So it's a similar kind of takeoff that we saw with LLMs. Like when I was at OpenAI, we worked on reinforcement learning of a GPT-1 architecture, and it did not work because the models weren't good enough. But once the base model got good enough, then you could start to do those kinds of experiments. And I think there's a similar analogy that's starting to happen in our world now as well, where the models have reached a point where there's actually a renewed interest in data and how do we actually bring the models into the loop on making that happen. I think that creates another really interesting cycle on compounding improvement of the models.

竞争格局 Competitive Landscape

Host

是的,我们刚才谈到了模型的复合改进,你提到制药行业是一个竞争极其激烈的领域。这很有趣,因为我认为人们对使用 AI 和 ML 生成蛋白质和分子的兴趣重新燃起,所以你们的领域实际上已经变得相当竞争激烈。而你们显然在保持前沿方面做得非常出色。对你们来说,这是部署的一年。你们已经锁定了一些制药合作伙伴关系,这些关系正在让你们的模型变得更好。但你们如何看待竞争格局以及如何保持在前沿?

Yeah, we're talking a bit about compounding improvement in the models, and you said something about how pharma as an industry is an incredibly competitive landscape. And it's interesting because I think there's been a renewed interest in using AI and ML to generate proteins and molecules, so your arena has actually become quite competitive. And you guys have obviously done an amazing job at staying at the frontier. It's been the year of deployment for you. You've locked up a number of pharma partnerships that are making your models better. But how do you guys think about the competitive landscape and staying at the front of the frontier?

评估与模型能力提升 Evaluation and pushing model capabilities

Josh

嗯,首先,我认为这要回到看待结果、不自欺欺人、保持严谨的态度上。所以,我们做大量模型评估的原因之一,主要是为了在模型本身进行爬山优化。我认为,我们带来的许多能力,在领域内都是首创的。看看我们的 Chai 2 模型,对吧?达到了一个成功率,你不再需要做大规模文库筛选就能看到结果。几个月后,展示了我们如何引入之前讨论过的许多可开发性属性,比如分子的可制造性。所以在很多情况下,我们是在推动模型前进,试图看到这些能力涌现,然后我们迅速锁定它们。我们如何让这些能力更加突出,使它们达到生产就绪,并能交付给客户?

Well, first of all, I think it goes back to looking at the results and not fooling yourself and being rigorous. So, one of the reasons why we do a lot of evaluation of the models, we mainly do it just to hill climb in the models itself. I think if you look at many of the capabilities that we've brought in, they've kind of been like first in the field. You look at our Chai 2 model, right? Getting to a success rate of the models that you didn't have to do large library screening anymore to see results. A couple months later, showing how we could bring in a lot of these developability properties we've talked about before, like the manufacturability of the molecules. So in many cases, we are pushing the model forward and trying to see these capabilities emerge, and then we try to quickly lock those in. How do we make those capabilities even more pronounced that they become production ready and we can ship them to our customers?

Josh

我想我喜欢告诉团队的一件事是,你知道,我们并不是与其他模型提供商正面竞争。我认为我们所有人都在与自然对抗。自然其实是一个相当不错的基线。人们长期以来一直以某种方式进行药物发现,你知道 Matt 谈到过也许你不想在 Chai 模型上加第 24 个模块,但人们已经在现有的湿实验方案上加了第 240 个模块,而且这些方案已经被调优得相当不错了。

I think one of the things I like to tell the team is that, you know, it's not like we are head-to-head with other model providers or something like that. All of us, I think, are working against nature. Nature's actually been a pretty good baseline. People have done drug discovery a certain way for a long time, and you know Matt talked about how maybe you don't want to add module 24 to the Chai model, but people have added like module 240 to the existing wet lab protocols, and they have been tuned quite considerably.

Josh

我们团队的一位科学家 Andy Young,实际上是二十年前在 MIT 最早研究酵母展示的人之一。他在辉瑞和基因泰克有 20 年的经验,真正磨练了这些方法,名下有一种药物获批,还有抗体。我认为你看看这样的人,Andy 知道如何用现有工具制造好的抗体,而这实际上是我们需要达到的标准。当然,我认为 AI 的上限会比我们以前做到的更高,否则做这件事的意义何在?我们创办公司不是为了制造一个快 10 倍的鼠标,对吧?我们创办公司是为了制造以前不可能实现的突破性药物。但最终,为了获得采用,我们需要达到那个标准。我认为几个月前我们达到了那个拐点。这就是为什么你看到很多大型制药公司的公告。但现在我们只需要继续磨练,让这些东西变得更好。

One of the scientists on our team, Andy Young, was one of the first people working on yeast display at MIT actually, like two decades ago. And he's got 20 years experience at Pfizer and Genentech, really honing in these methods, has a drug approval to his name and antibodies. And I think you look at someone like that, and Andy knows how to make a good antibody with existing tools, and that is actually the bar that we need to clear. Now, of course, I think the ceiling on AI is going to be a lot higher than what we've managed to do before, otherwise what would be the point of doing this? We didn't start the company just to make a 10 times faster mouse, right? We started this to make breakthrough medicines that weren't possible before. But ultimately, that is the bar that we needed to clear in order to get adoption. I think we hit that inflection point a couple months ago. That's why you've seen a lot of these big pharma announcements. But now we just need to continue to hone in on making these things even better.

在 Chai 工作的优劣 Best and worst things about working at Chai

Host

对于正在收听、觉得“哇,这听起来很酷。我想知道在 Chai 工作会是什么样子”的人来说,在 Chai 工作最好的事情是什么?最糟糕的事情是什么?

For somebody who's listening who thinks, "Wow, this sounds pretty cool. I wonder what it would be like to work at Chai." What is the best thing about working at Chai? And what is the worst thing about working at Chai?

Josh

是的,我可以谈谈其中的一些。我会即兴思考。最糟糕的事情……所以我认为最好的事情就是每个人都真正以使命为导向。每个人都如此专注于他们正在做的事情。我在其他公司工作过。我见过最接近的可能是博士实验室里的那些人。但每个人都难以置信地有动力。我们都非常努力。有一个明显的共同目标。我认为这真的很罕见。而且我认为这要追溯到我们从一开始就有的专注。我们一直有一个清晰的理念,一个清晰的计划,关于事情如何实现,如何变得更好。而且 Chai 的每个人都非常认同这一点。看到每个人投入的奉献精神,真是太棒了。

Yeah, I can speak to some of this. I'll think of this on the fly. The worst thing... So I think the best thing is just how actually mission-driven everybody is. Everyone is so dedicated to what they're doing. I've worked at other companies. The closest I've ever seen to this is maybe some of the guys in my PhD lab. But everyone is just incredibly motivated. We all work really hard. There's an obvious shared goal. And I think that's really rare to see. And I think this goes back to just the focus that we've had since the beginning. We've always had a clear philosophy, a clear plan on how things are going to get there, how things are going to get better. And really everyone at Chai is very bought into this. It's pretty amazing just to see the amount of dedication that everyone's putting in.

Josh

关于 Chai 最不喜欢的事情。不是直接关于蒲公英巧克力。也许。不,

Least favorite thing about Chai. Not directly on top of dandelion chocolate. Maybe. No,

Host

也许是下一个办公室。

Maybe the next office.

Josh

是的。实际上,现在最不喜欢的部分可能只是让事情在规模上运转。而且,我真的不知道那意味着什么。实际上,当我们创办 Chai 时,我们有 128 个 GPU,我当时想,这是实验室可能的最大规模。这太疯狂了。我来自博士小组,我们四个人共享 8 个 GPU,我当时觉得,我真是 GPU 富翁。太疯狂了。所以现在在 Chai,我们有更多的基础设施要维护,我们有更多的算力资源。幸运的是 GPU 是可并行的。但这也带来了它自己的一系列问题。比如如何让集群长期保持健康?如何进行大规模训练运行?如何让它连续运行数月?即使它真的挂了,如何自动恢复这些任务?如何减少通信开销?如何优化模型并充分利用你拥有的资源?所以我认为这些是不断出现的问题。它们是很好的问题,但我也认为它们真的很难解决。是的,我只是很兴奋能解决这些问题。

Yeah. Actually, so probably the least favorite part now is just getting things to work at scale. And really, I didn't even know what that meant. Actually, when we started Chai, we had 128 GPUs and I was like, this is the most scale possible for a lab. Like this is crazy. I came from my PhD group where there were four of us sharing eight GPUs, and I was like, I just felt GPU rich. It was crazy. So now at Chai, we have a lot more infrastructure to maintain, we have a lot more compute resources. Luckily GPUs are parallelizable. But that also brings up its own set of problems. So just like how do you keep a cluster healthy over time? How do you get that large training run? So how do you keep that running for months on end? And even when it does die, how do you automatically resume these things? How do you keep all of your communication down? How do you optimize the models and make the best use of the resources that you have? So I think these are a lot of problems that continuously pop up. They're good problems to have, but I think they're also really difficult to solve. Yeah, and I'm just excited to work on this.

Host

是的,我记得我们合作过几次的一位 CEO,Frank Slugman,说过这样一句话:你要么承受失败的痛苦,要么承受成长的痛苦。你宁愿承受成长的痛苦。

Yeah, I remember one of the CEOs that we've worked with a couple times, Frank Slugman, had this line about, you either have the pain of failure or the pain of growth. You'd much rather have the pain of growth.

Josh

是的,完全正确。

Yeah, that's exactly right.

Josh

我想我最喜欢的部分可能是结果。这听起来有点老套,但没有什么比它有效更让人兴奋的了,对吧?而且知道你在,你知道,我认为我们公司里的许多人,对吧?我们在 AI 领域工作了很长时间,对吧?你积累了所有这些经验,然后知道你把它应用到了真正重要的事情上。甚至我认为,就像 Matt 和我,我们研究这个问题已经大约 10 年了,对吧?几年前,我看着,是的,我们在写一些很酷的论文,每个人都在庆祝。我们真的在让世界变得更好吗?这真的会影响一些患者吗?我认为现在答案是肯定的。我们已经达到了一个点,这将给世界带来巨大的改变。每次你得到这些突破性的结果,每当产品上有新功能让客户的生活更轻松,每当科学团队传来新的实验结果,说实话,总是令人振奋地意识到,这真的会以一种相当深刻的方式改变世界。

I think my favorite part is probably the results. And that sounds a bit cliche, but there's nothing like it's working, right? And just knowing that you're, you know, I think many of us in the company, right? We've been working in AI for a long time, right? There's all this experience you've built up, and to know that you're applying it to something that really matters. Just even I think just Matt and I, we've been working on this problem for like 10 years, right? And a couple of years ago, I'm looking at, yeah, we're writing some cool papers, everyone is celebrating this. Are we actually making the world better? Is this actually going to impact some patients? And I think now the answer is actually yes. We have reached the point where this is going to make a big difference in the world. And every time you get one of these breakthrough results, anytime there's a new feature on the product that makes the lives of our customers easier, whenever there's new lab results coming back from the science team, it's just always so, honestly, exhilarating to realize this is actually going to change the world in a pretty profound way.

生产挑战与代码库优先级 Production challenges and codebase priorities

Josh

我觉得这也许就说到最不喜欢的部分了,对吧?你知道,我们在经营公司,有真正的合作伙伴依赖我们,事情必须得运转起来,对吧?你发布新一代模型,怎么确保里面没有 bug?怎么确保没有回归问题?对吧?这不再是那种蓝天研究问题——哦,我们有了些酷炫结果,然后就继续往前走了。随着团队壮大,我们不得不把生产级代码库作为高度优先事项。我们怎么确保代码库处于一种能让更多人参与贡献的状态?所以,我们的另一位联合创始人 Jack 喜欢说:如果你想长期快跑,有时短期就得稍微慢一点,对吧?确保你在构建真正的东西。这又回到复利那个想法,回到我们不会为了凑合下一个功能而临时加第 25 个模块。所以有时候,你特别兴奋想赶紧拿到下一个结果,就想直接跳进去。但我们有真正的合作伙伴,世界上一些最大的公司现在都依赖我们。重要的是我们要意识到,要把这份责任放在心上,确保我们构建的系统能持续运转。

I think that also then maybe comes to the least favorite side, right? Like you know, we're running the company and it's like we have real partners that are relying on this and like things have to work, right? And you ship a new model generation. How do you make sure that there's no bugs in that? How do you make sure you don't have regressions? Right? This is no longer just like a blue sky research problem of like oh we got some cool results and we move on. We've had to have really high priorities on like you know having production level code bases as the team grows. How do we make sure that the codebase is in a state that more people can contribute to this? So, something that one of our other co-founders, Jack, likes to say is that if you want to move fast in the long term, you sometimes have to just move a little bit slower in the short term, right? And make sure that you are building something. Again, that goes back to that compounding idea. It goes back to we don't add module 25 to make the next thing work. So sometimes, you know, you're so excited to get to the next result and you just want to jump into it. But we have real partners, some of the biggest companies in the world that are now relying on us. And it's important that we realize that we take that responsibility to heart and we make sure we're building systems that continue to work.

Chai Discovery 名称由来 Origin of the name Chai Discovery

Host

我有个特别想问的问题。你们为什么叫 Chai Discovery?

I have a burning question. Why are you guys called Chai Discovery?

Josh

化学(Chemistry)和 AI?但我们也喜欢 Chai(茶)。

Chemistry and AI? But we love Chai as well.

Host

哦。

Oh.

Josh

办公室里有很多茶主题的东西。

It's a lot of chai theme stuff in the office.

Host

这是个好问题。我也不知道这个。

That's a good question. I didn't know that either.

Josh

Josh 是远见者。那都是他想的。

Josh is the visionary. That was all him.

Host

这名字确实很用户友好。

It is a very user friendly name.

Josh

是啊。我们也想要一个名字——你看生物技术公司的名字都特别复杂。我们想要一个简单得多的名字。我们想把这整个事情变简单,对吧?所以需要一个简单的名字来搭配。

Yeah. We also wanted a name that like biotech companies have such complicated names. We wanted something that's going to be much simpler. We're trying to make this whole thing simpler, right? So, we need a simple name to go along with it.

未来6-12个月的期待 Excitement for the next 6-12 months

Host

太棒了。接下来 6 到 12 个月,你们最期待什么?

Awesome. What are you guys most excited about in the next 6 to 12 months?

Josh

对我来说,就是正在发生的部署。我们已经宣布了几个这样的合作,我真的很期待听到合作伙伴们上线带来的结果。这又回到真正产生影响这一点,也是为什么我很高兴看到这些合作在协议签署后依然进展顺利。我们和这些人合作时,模型不是躺在某个架子上,而是真正用在真实项目里。人们试图攻克毁灭性疾病,如果 Chai 能给他们一个有效的分子,真的能改变患者的生活。所以我真的很期待看到那会怎样。这里的进展速度令人难以置信,模型的速度也一样。比如一年前,你不可能零样本生成一个分子,还能对项目会成功有把握。现在这已经变了。有人可能零样本生成一个分子,然后说:我觉得我们现在要把这个项目推进到临床。一年后,他们甚至可能让第一批分子进入患者体内。这种速度简直不可思议。而且,你知道,有时候想到这些会起鸡皮疙瘩——比如,好吧,我的模型将要……Matt 有我们上一家公司的专利,关于在电脑上生成分子,而现在这些东西已经用在患者身上了。再想想规模,我不知道,几年后,我们会有几十个?还是几百个 Chai 分子进入人体?想到这对患者可能意味着什么,有点让人震撼。

I think for me it's just the deployments that are happening. So, we've announced a couple of these partnerships and I'm really excited just to hear about the results that our partners are bringing online. It goes back to this point of making a real difference and also why I'm so happy to see how these partnerships are going even post into the agreement and as we're working with these folks that the models are not just sitting on a shelf somewhere like they're actually being used on real programs. People trying to approach devastating diseases where if Chai could give them a molecule that works it could really change the lives of patients. So I'm really excited to see how that goes. Just the pace of progress here is incredible but also just like the pace of the models. Like a year ago you could not zero shot a molecule and have a good sense that your program was going to work. Like now that's changed. Someone might zero shot a molecule and be like, I think we're going to bring this program to the clinic now. And then a year later, you know, they might even have some of those first molecules going into patients. And just the speed of that is just incredible. And it's, you know, sometimes you get some shivers thinking about this that like, okay, my model is going to... Matt has some patents from our last company about just generating the molecule on the computer and these things are now in patients. And just to think about the scale, like I don't know, a few years from now, do we have dozens? Do we have hundreds of Chai molecules going into people? It's a bit mind-blowing to think about what that might look like for patients.

Matt

我觉得让 Chai 独特的一点就是这种投入。人们会说:你工作真多。你不会倦怠吗之类的。但实际上真的很容易,而且当你取得进展、看到我们正在取得的进步时,会非常有动力,就会想:天哪,下一个模型会比上一个更好。我们发现了这个新东西,等等。这对我来说太有激励性了。更多的是:我们接下来还能解锁什么?怎么让这些东西更可控?比如有人带着某个靶点来找我们,而不是只希望我们碰巧得到好的亲和力之类的,你知道,我们真的能控制这个吗?我们能说我们想要一个恰好 10 纳摩尔的结合剂之类的吗?我觉得有很多技术问题我们正处在解决的边缘。对我来说,把这些东西确定下来,把这一切都搞定,然后看看接下来会通向哪里,真的很有动力。

I think one of the things that again like what makes Chai unique like the dedication. Like people are like man you work a lot. And like you like aren't you burnout or whatever. It's like it's actually really easy and like it's very motivating when you're making progress seeing the progress that we're making and just like thinking man the next model is going to be even better than the last. We identified this new thing so on. Like that's so incredibly motivating for me. It's more of just like what can we unlock next and like how do we make these things more controllable. Like when someone comes to us with a certain target instead of just like hoping we get good affinity or something like that, you know, can we actually control this? Can we say like we want exactly a 10 nanome or binder or things like that. Like there are a lot of technical things that I think we're kind of right on the brink of solving. And for me it's like it's really motivating just to pin those things down and just like get all of this over the line and see kind of where that leads to next.

结束致谢 Closing thanks

Host

太棒了。Matt、Josh,感谢你们为工程生物学所做的一切,也感谢你们今天和我们分享你们的故事。

Awesome. Matt, Josh, thank you for engineering biology and thank you for sharing your story with us today.

Josh

谢谢你们。感谢邀请我们。

Thank you guys. Thanks for having us.

互动版:逐字朗读 + 针对本期提问 →