From GANs to Diffusion: AI's Evolution in Protein Modeling
打开互动全文版(中英对照 + 朗读 + 问答)→Evan Fineberg 和 Sergey Udov 讨论从 GAN 到扩散模型在蛋白质结构预测中的转变,他们的物理学背景,以及 Genesis Molecular AI 改进 AI 用于药物发现的方法。
Evan Fineberg and Sergey Udov discuss the shift from GANs to diffusion models for protein structure prediction, their physics backgrounds, and Genesis Molecular AI's approach to improving AI for drug discovery.
我清楚地记得,在 2017、2018 年左右,大家都在讨论 GAN,认为生成对抗网络显然是图像生成的未来。显然,它们在蛋白质或蛋白质折叠系统上效果并不好,我们不得不等待正确的原语被创造出来,结果就是扩散模型,它被证明是更适合这个领域的原语。很酷的一点是,现在对于那些对真正核心的基础 AI 研究感兴趣的人来说,实际上一些最具创新性的扩散研究正发生在我们领域,也就是 3D 结构预测。当时没人能预料到这一点,但现在它可以说是扩散模型的一个支柱。大家好,欢迎收听 Latent Space AI for Science 播客。我是 Mirmix 的 CTO R.J. Haniki。和我一起的是联合主持人 Brandon Anderson,我们很荣幸邀请到 Genesis Molecular AI 的创始人兼 CEO Evan Fineberg,以及在加入 Genesis 担任 CTO 之前领导了 Llama 2 和 Llama 3 预训练的 Sergey Edunov。
I remember very clearly in like 2017, 2018 talking about GANs and how generative adversarial networks were clearly the future of image generation. Obviously, they didn't work very well for proteins or protein folding systems, and we sort of had to wait for the right primitive to get created, and that turned out to be diffusion, which turned out to be a much more useful primitive for the space. What's kind of cool is right now for people that are interested in really core fundamental AI research, actually some of the most innovative diffusion research is happening in our field, in 3D structure prediction right now. No one would have predicted that then, but now that's a pillar of diffusion, I'd say. Hi there. Welcome to the Latent Space AI for Science podcast. I'm R.J. Haniki, CTO of Mirmix. I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Fineberg, founder and CEO of Genesis Molecular AI, and Sergey Edunov, who led Llama 2 and Llama 3 pre-training before he joined Genesis as CTO.
大家好,我是 Sergey。我上学时学的是物理,那是很久以前的事了。毕业后,我碰巧从事软件工程工作,当时以为再也不会用到物理了。当机器学习兴起并成为热门时,我发现机器学习中的许多事情实际上和物理中做的事情非常相似。于是我就搭上了机器学习的顺风车,做了很多 AI 研究,在 Facebook 的 FAIR 团队待了相当长一段时间。后来我领导了 Llama 团队,负责 Llama 2 和 Llama 3 模型,最近我决定再次调整职业方向,重新找回一点物理的根基,于是加入了 Genesis 担任 CTO。
Hi, I'm Sergey. I studied physics at school. This is a long time ago. After graduation, I happened to work in software engineering. I thought I would never need physics again. When ML came up and became a thing, it turns out that a lot of things you're doing in ML were actually very similar to what you would do in physics. So I jumped in on the ML bandwagon and did a lot of AI research, being a part of the FAIR team at Facebook for quite a while. I later led the Llama team for Llama 2 and Llama 3 models, and then recently I decided to pivot my career again and recover my roots in physics a little bit, and I joined Genesis as CTO.
大家好,我是 Evan,Genesis Molecular AI 的创始人兼 CEO。和 Sergey 一样,我也是物理专业出身。我成长过程中和家人有点不同,家里几乎所有人都是某种医疗专业人士,我姐姐则成了一位成功的电视编剧、剧作家和小说家。我的热爱是物理和计算机科学。但我对成年人的认知是,你应该帮助别人,最好是帮助患者。所以我一直在寻找正确的方式去做这件事。来到斯坦福后,我在 Vijay Pande 的实验室攻读博士,2010 年代中期,我们对机器学习在图像和语言方面取得的惊人进展感到兴奋,虽然这个说法现在有点过时了。差不多在 Sergey 在 FAIR 研究大型图的同时,我在斯坦福研究许多小型图。分子其实就是由原子、键和空间相互作用构成的网络。如果你在正确的时间出现在正确的地点,利用我们的物理背景来改进用于分子分析的 AI 算法,我们就在图机器学习领域发表了几篇论文。和 Sergey 一样,我以为有了机器学习就不再需要物理了。但事实证明,随着 Genesis 的发展,我们在 GPU 上运行的大量蛋白质模拟也派上了用场。过去几年,我们非常兴奋地探索如何为一个全新的领域构建基础模型,并让它们对患者有用。
Hey, I'm Evan. I'm the founder and CEO of Genesis Molecular AI. Like Sergey, I was also a physics major. I was a bit different from everyone in my family growing up. Almost everyone was a medical professional of some kind, and my sister became an accomplished TV writer, playwright, novelist. My love was physics and computer science. But my mental model of an adult was that you should help people and help patients ideally. So I was always searching for the right way to do that. After arriving at Stanford, doing my PhD in Vijay Pande's lab, we were excited in the mid-2010s about everything amazing going on in machine learning, to use a dated term, for images and for language. Around the same time that Sergey was at FAIR working on a lot of big graphs, I was at Stanford working on many small graphs. Molecules are really networks of atoms and bonds and spatial interactions. If you're at the right place at the right time to bring to bear our backgrounds in physics to improve AI algorithms for looking at molecules, we published a few papers in the area of graph machine learning. Like Sergey, I thought I wouldn't need this physics again because there's machine learning. But it turns out the massive amounts of simulations on GPUs of proteins that we ran also came in quite handy as Genesis evolved. We've been really excited the past few years to figure out how to build foundation models for a totally new domain and make them useful for patients.
是的,这很好地引出了我的第一个问题。你们两位在这个分子和生物学的机器学习领域已经工作了大约 10 年。可以说,自那以后,整整一代技术生物公司来了又去。虽然很多分子机器学习相当有效,但有一个领域历来对机器学习建模有抵抗力,那就是蛋白质-小分子相互作用的世界。随着 Genesis 最近取得的一些进展,似乎你们可能真的开始在这个领域取得了我们很久没见过的实质性改进。能谈谈你们做了什么,Genesis 做了什么,导致这些改进的发展,以及为什么你认为这比那些模棱两可的传统机器学习策略是真正的改进吗?
Yeah, that's a really nice lead into my first question. You've both been in this domain of machine learning for molecules and bio for roughly 10 years. It's kind of an entire generation of tech bio has come and gone since then. While a lot of machine learning for molecules has been quite effective, one domain where it's been historically resistant to machine learning modeling has been the world of protein-small molecule interactions. With some of the recent advances that Genesis has put out, it seems like you might have actually started to make real improvement in this area in a way that we haven't seen for a long time. Can you talk about what you have done, what Genesis has done, the developments which led to improvement, and why you think this is actually a real improvement over some of the traditional machine learning strategies which were ambiguously helpful?
我完全同意,Brandon。非常了不起的一点是,当我们最初作为斯坦福 AI 研究的衍生公司创立时,我们担心自己可能太晚了。你还记得吗,我们第一次见面差不多就是那个时候,七年前,正如你所说,几乎是该领域的不同世代。当时我们只是一家小型种子公司,而该领域已经有一些现有公司,他们筹集的资金比我们多几个数量级。已经有一些非常令人兴奋的学术论文发表,我们不得不不断回答一个问题:是否还有空间让另一家公司做 AI 药物发现?现在回头看,这听起来很疯狂,事后诸葛亮嘛。我们当时的主张和现在一样:有 20,000 个蛋白质编码基因,每个都可能导致疾病。就像 AI 的其他领域一样,并不是从零到一,然后问题就突然解决了。现实是,从来没有一个 iPhone 时刻。即使是手机,也需要随着时间的推移进行迭代开发,才变成今天的样子。人们喜欢回顾性地谈论 ChatGPT 时刻,但最早广泛使用的 ChatGPT 模型远没有现在有用。我认为我们应该以同样的方式看待 AI 背景下的药物发现:随着发展,我们应该期待解决越来越多的问题,会有大的飞跃,也会有迭代改进,但它们会随着时间的推移而累积。这就是为什么我认为,就像 10 年前还很早,但从 AI 角度看,价值已经开始为药物开发创造一样,快进 10 年,我们的技术——我们将在本次讨论中谈到——比那时有用得多。未来 10 年,我们还会看到另一个指数级的改进。就像你的汽车在自动驾驶方面今天能做到的事情——我们在旧金山录制,你看向窗外看到 Waymo,几年前我们会被这些技术震撼到。所以我认为我们都应该这样看待:每隔几个月、每年,技术都会变得越来越有用,用于用人工智能发现新药物,而且这种情况在未来几年将继续如此。
I totally agree, Brandon. One of the really remarkable things is that when we were founding the company initially as a spin-out of the research we were doing in AI at Stanford, we were afraid that we could be too late. I mean, you remember that you and I first met around that time, right? Seven plus years ago, as you rightly point out, almost a different generation in the area. At the time when we were a little seed company, there were already incumbents in the space who had raised orders of magnitude more money than us. There were already some really exciting academic papers that were out, and we had to constantly answer the question: is there room for another company doing AI in drug discovery? Which is a crazy thing to say looking with a backward-facing light, and hindsight is 2020, I suppose. The claim that we made then is the same that we're making now: there are 20,000 protein-coding genes, and each of them can cause a disease. In the same way that in other domains of AI, it's not like they were zero to one and then suddenly that problem was solved. The reality has been there's been no single iPhone moment. Even the phone required iterative development over time to become what it is today. People love to talk retrospectively about the ChatGPT moment, but the first ChatGPT models that became widely used were vastly less useful than they are now. I think we should look at drug discovery from an AI context in the same way: as we go, we should expect to solve more and more problems, and there will be large leaps, there will be iterative improvements, but they'll compound over time. That's why I think, in the same way that 10 years ago was very early but clearly value was starting to be created for drug development from an AI perspective, fast forward 10 years, our technologies, which we'll talk about in this discussion, are vastly more useful than they were then. In the next 10 years, we'll see another exponential improvement as well. In the same way that your car in terms of autonomy can do things today—we're recording this in San Francisco, you look outside and see Waymos, we'd be blown away a few years ago by those technologies. So I think we should both be looking at it as every few months, every year the technology gets more and more useful for discovering new medicines with artificial intelligence, and that will continue to be true for years to come.
也许我们可以稍微回溯一下时间,解释一下大约十年前蛋白质小分子药物发现的状况。我们当时实际上能用机器学习模型做什么?有哪些失败模式是人们当时从实践角度没有预料到的?
Maybe we can go back in time a little bit and explain like what was the state of protein small molecule drug discovery let's say about a decade ago. What could we actually expect to do with machine learning models and what were some of the failure modes that people I think didn't really see coming from a practical standpoint back then?
有趣的是,我稍微打破第四面墙。我进来时并不真正了解这些问题,但 Brandon 真的很懂我的语言,因为就像在座的其他人一样,尽管我的背景更偏向物理和计算机科学,十多年来我一直非常专注且充满激情地研究如何用 AI 制造药物。我见证了太多地壳运动般的巨变。给你一个具体例子,我为那些生物学背景不多的人提供一点背景:药物发现类似于为锁找钥匙,锁通常是蛋白质(有时是核酸),药物就是钥匙。它通常是小分子、肽、抗体或其他形式,想要结合到那个锁、那个蛋白质、受体上,并改变它的某些特性。如果是酶,你通常想阻止它发挥作用。有时你想激活那个内源蛋白,但这是引入一个外源分子来改变生物通路。这个过程中一个必要但不充分的部分是找到一种能很好结合受体或蛋白质的分子。当然,这还不够。我们还需要确保它不结合某些抗靶点,即你想避免的蛋白质,这通常会导致毒性。你需要 ADMETox,通常大约 30 个特性,以确保你的分子除了有效之外还安全,并能到达正确的组织。所有这些都至关重要,都是必要的,但没有一个是充分的。但回答你关于蛋白质-配体相互作用的问题,一个具体例子是:我们长期以来有一个假设,如果我们能预测药物和蛋白质的 3D 结构(对计算机科学的人来说想象成一个点云),并且能高精度建模,那自然会带来更准确的结合亲和力测量或效力预测。这个假设根本没法检验,因为预测这些 3D 姿态、复合物 3D 结构的模型太差了。或者,如果你能预测,它需要巨大的计算资源,以至于你不如直接解决问题,对吧?你不如直接解出 3D 晶体结构或冷冻电镜结构,这可能需要数万美元、数月或数年。整个博士论文、博士后有时可以 24/7 工作,试图解出一个 3D 结构。假设是,如果我们能把这个速度提高几个数量级,有时更准确——这可以通过人工智能或之前的分子动力学实现——那么你不仅能加速药物发现,还能为以前被认为不可成药的靶点发现药物。那只是一个假设。直到最近几年才证明,如果我们系统性地提高这些预测的准确性,如果我们能证明蛋白质-配体 3D 复合物的准确性,我们也能使效力预测更准确。这是我个人感到非常欣慰的,因为我们一直在展示这一点,因为几年前我们创办公司时那只是一个假设,现在由于许多不同想法的融合,它已成为现实,这些想法使得 Pearl 和其他类似的基础模型成为可能。
It's funny, I mean breaking the fourth wall a little bit. I came in not really knowing any of these questions, but Brandon really knows how to speak my love language because, like the other people in this room, though my background is much more on the physics and CS side, I've been very focused in a deeply passionate way about the area of how can we make medicines with AI for over a decade. I've seen so many tectonic shifts. To give you one concrete example, and I'll give a little background for those who don't have as much biological background: drug discovery is akin to finding a key for a lock, where the lock is usually a protein (in some cases a nucleic acid), and the drug is the key. It's usually a small molecule, a peptide, an antibody, or some other modality that wants to bind to that lock, that protein, the receptor, and change something about it. If it's an enzyme, you usually want to stop it from functioning. Sometimes you want to activate that endogenous protein, but it's to introduce an external molecule, an exogenous molecule, to change something about that biological pathway. A necessary but not sufficient part of that process is finding a molecule that binds well to that receptor or that protein. Of course, that's not sufficient. We also need to make sure it does not bind to certain anti-targets, certain proteins you want to avoid. That often mediates toxicity. You need ADMETox, which is typically like 30-ish properties to make sure your molecule is safe in addition to being effective and gets to the right tissue. These are all critical, and all are necessary, and none are sufficient. But one concrete example to answer your question about protein-ligand interactions is that we had a hypothesis for a long time that if we can predict the 3D structure, the 3D coordinates (imagine for the CS people like a point cloud) of your drug and the 3D positions of the protein, and if you could model that with high accuracy, that would lead naturally to measuring binding affinity or predicting potency more accurately. That was a hypothesis that fundamentally could not be tested because models for predicting those 3D poses, the 3D structures of complexes, were so bad. Or if you could predict this, it requires so much computational resources that you might as well just solve the problem, right? You might as well solve the 3D crystal structure or cryo-EM structure, which can cost tens of thousands of dollars, months or years. Entire PhD theses, postdocs can sometimes work literally 24/7 trying to solve one 3D structure. The hypothesis was that if we make this orders of magnitude faster, sometimes more accurate, which we can talk about with artificial intelligence or before that with molecular dynamics, you would thereby not only accelerate drug discovery but enable the discovery of medicines for targets that were previously thought to be undruggable. That was just a hypothesis. It took until the last few years to show that if we systematically improve the accuracy of those predictions, if we can prove the accuracy of a protein-ligand 3D complex, we can make potency prediction more accurate as well. That's something I am personally extremely gratified to feel like we've been showing, because that was just a hypothesis when we started a company a few years ago, and now it's become a reality thanks to the convergence of a lot of different ideas that were able to enable Pearl and other foundation models like it.
你能谈谈 Pearl 吗?我认为 Pearl 的一个有趣之处在于,它试图扩展在我看来似乎不可扩展的东西。RNA 小分子,你知道,通常如果你扩展它们,模型就会变成纯粹的模式匹配器。这些东西喜欢告诉你你已经知道的东西,有时它们可能不太擅长告诉你你不知道的东西。所以似乎有迹象表明 Pearl 实际上已经超越了这一点。其中的关键见解是什么?
Can you talk about Pearl a bit? I think one of the interesting things about Pearl in my mind is it was an attempt to kind of scale what seems unscalable to me. RNA small molecules, you know, normally if you scale them, models just come out as just like pure pattern matchers. These things love to tell you what you already know and sometimes they're maybe not so great at telling you something you don't know. So maybe it seems like there's hints that Pearl has actually moved beyond this. What were the key insights which went into this?
当然。顺便说一句,Brandon,你描述的情况不应该让在座的任何纯机器学习从业者感到惊讶,因为机器学习最初就是为模式识别而构建的。当你启动 Claude 或 ChatGPT 或其他类似工具时,也是一样:它在最接近训练数据时表现最好,并且会进行模式匹配。所以我认为,在 AI 与物理世界相遇时,紧迫感在于如何外推、如何构建可泛化的模型,这正是我们一直努力的方向。也许 Sergey 想对此做更多说明。也许退一步讨论一下 Pearl 是什么。它是一个结构预测模型,基本上意味着它输入一个蛋白质序列和一个你试图附着到该蛋白质上的配体表示,然后预测该蛋白质和配体在 3D 空间中作为结构看起来如何。
Sure. And by the way, Brandon, what you describe, that shouldn't surprise any pure ML practitioners in the audience, just because machine learning is initially constructed for pattern recognition. It's the same with when you fire up Claude or ChatGPT or what have you, it's going to do best when it's closest to the train data and it's going to pattern match off it. So I think that has been the big sense of urgency in AI meeting the physical world: how to extrapolate, how to make generalizable models, and that's what we've been really hard to work on. Maybe Sergey wants to give some more color on that. Maybe to step back a little bit and discuss what Pearl is. It's a structure prediction model which basically means it takes as an input a protein sequence and a ligand representation that you try to attach to this protein and it predicts how this protein and ligand are going to look together as a structure in the 3D space.
所以人们有时用的一个术语是共折叠。它是多重折叠。AlphaFold 3,然后一些,还有 Boltz 和 OpenFold 3 等也以各自的方式实现了其中一些想法。
So like a term for this people sometimes I've used is like co-folding. It's multiple folding. AlphaFold 3 and then some of the and then Boltz and OpenFold 3 and so on have also implemented some of these ideas in their own way.
是的,有几个模型,有些是开源的,有些是闭源的,都在追求类似的方向。我们的根本区别在于我们专注于小分子空间。很多模型在做蛋白质-蛋白质相互作用,这对小分子空间也有效。乍一看,这听起来像是一个更简单的任务,因为,嗯,它是小分子,对吧?那为什么会具有挑战性呢?但现实是,小分子的搜索空间非常巨大。宇宙中有 10 的 60 次方个类药小分子。所以祝你好运搜索那个空间。而且还有各种变化,比如你可以旋转它,这些分子可以有不同构象。
Yes, there are several models, some are open source, some are closed source, that are pursuing similar direction. Our fundamental difference is that we're focusing on small molecule space. A lot of models are doing protein-protein interactions, and it works well for small molecule space. On the first glance, it may sound like it's an easier task because, well, it's a small molecule, right? So why would it be challenging? But the reality is the search space is so vast for small molecules. There are 10 to the 60 drug-like small molecules in the universe. So good luck searching that space. And there are also all variations like you can rotate it, you can have different conformations of those molecules.
也许直觉是这样的:当你在谷歌搜索中输入时,如果你只输入几个词,你会得到一大堆匹配结果。但如果你有一个非常具体的查询,那么你会得到少数匹配,而且它们更可能匹配。
Maybe the intuition there is that you have like when you type in a Google search, if you just do a few words you're going to get a huge list of matches. But if you have a very specific query then you're going to have a few matches and they're more likely to match.
类似地,对于小分子,如果你的查询分子很小,那么你很可能需要搜索一个巨大的可能匹配空间。而如果你有一个非常复杂的分子,那么很容易排除不匹配的情况。这样理解对吗?
So similarly with a small molecule, if you have a very small query which is your molecule, then it's likely that you're going to have to search a huge space of possible matches. Whereas if you have something very complex, then it's easy to rule out a match. Is that a good way to think about it?
这确实是一种理解方式。对我来说,这纯粹是计算复杂度的问题——要找出哪个小分子能附着到这个蛋白质上。这个问题太庞大了,没有好的模型根本不可能解决。
That's one way to think about it. Definitely. For me, it's just the computational complexity of figuring out which small molecule will attach to this protein. It's such a vast problem to solve that it's impossible to do without good models.
就像大海捞针。
It's like finding a needle in a haystack.
是的。
Yeah.
除了你的针之外,其他所有东西都非常危险,或者根本不结合。
Where everything except your needle is very dangerous or just doesn't bind.
或者它们都是针。
Or they are needles.
对,没错。在针堆里找干草可能更贴切。
Yeah. Right. Right. Finding hay in a needle stack might be a better analogy.
是的。
Yeah.
那么现在我们可以考虑如何训练这些模型。人们常用的训练数据是所谓的 PDB,这是一个公开的数据库,包含所有历史上的晶体结构,但规模并不大,大约有 20 万个晶体结构,而且很难扩展。创建新的晶体结构需要大量的时间、精力和金钱。虽然有一些项目在推进,但扩展速度非常缓慢。所以扩展这个数据库相当困难。但我们发现,在小分子领域,实际上可以用物理学来建模小分子,模拟它们的行为,从而生成更多数据。这在蛋白质-蛋白质相互作用中不一定可行。
So now we can think about how we train those models. A lot of training data that people use is the so-called PDB, which is a public database of all historical crystal structures, and it's not that big. It's like 200,000 crystal structures, and it's very hard to expand. It takes a lot of time, energy, and money to create new crystal structures. Although there are some projects pursuing this, it's still expanding at a glacial pace. So expanding this database is pretty hard. But what we figured out is possible and very relevant to our small molecule space: in small molecule space, you can actually model your small molecules with physics. You can model their behavior, and that allows you to create more data. But you can train a model on something which is not necessarily possible in protein-to-protein interactions.
因为蛋白质太复杂了。
Because they're too complex for a protein.
是的,那些是非常大的分子,用物理学建模非常困难。不是不可能,只是计算上非常困难。
Yeah, those are very big molecules. So it's very hard to model them with physics. Not impossible, it's just computationally very hard.
所以你的意思是,你们用分子动力学之类的方法很好地建模了这些小分子,从而能以低得多的成本将这些结构纳入训练集,作为合成训练数据使用。
So I guess what you're saying is you do a really good job of modeling these small molecules using MD or something like that. And so you can put together those structures in your training set at a much lower cost, and use that as your synthetic training set.
或许我们可以退一步,谈谈我们的路线图。这与 LLM 的 Scaling 路线图没有本质区别。回想一下,在 LLM 中,我们有所有阶段:首先是预训练 Scaling,然后是后训练 Scaling(微调或 RL,现在大家都在做 RL),最后是推理时 Scaling。这三个概念相互关联,正是它们带来了当今最先进的 LLM。在我们这边,情况也基本类似。我们也有预训练 Scaling,通过生成大量合成数据来训练更好的模型。我们确实这么做了。然后我们开始做的第二件事,实际上是 LLM Scaling 的第三步:推理时 Scaling。这本质上非常相似。在 LLM 中,推理时 Scaling 指的是思考令牌——LLM 不会立即给出答案,而是先思考一段时间,然后给出回应。我们的模型也做了类似的事情:模型被迫思考,只不过它不是用语言令牌思考,而是用晶体结构——不是完全实体化的晶体结构,而是某种内存中的晶体结构表示。模型会反复处理这些表示。在这个过程中,我们使用基于物理的引导来将模型输出导向正确的方向。我们发现这能大幅提升模型性能。
Maybe to step back a little and talk about our roadmap. It's not fundamentally different from the roadmap for LLM scaling. Remember, in LLMs we have all the stages: you have pre-training scaling, then you have post-training scaling where you do either fine-tuning or RL—now everybody's doing RL—and then you do inference time scaling. All three concepts connect, and that's what led us to state-of-the-art LLMs these days. It's not fundamentally different on our side. We also have pre-training scaling where we create a lot of synthetic data to train better models. We do that. Then the second thing we started doing is actually the third step in LLM scaling: we started doing inference time scaling. And it's fundamentally very similar. In LLMs, when we talk about inference time scaling, we're talking about thinking tokens, where the LLM, instead of giving you the answer right away, goes and thinks for a while and then comes up with a response. Well, we are doing a very similar thing with our models, where a model is forced to think, except it's not thinking in language tokens; it's thinking in terms of crystal structures—not fully materialized crystal structures, but some sort of crystal structure representation in memory. The model kind of goes back and forth with those. And we use physics-based guidance during this process to steer the model output in the right direction. What we found is that it improves model performance by a lot.
思考令牌的工作原理很容易理解,因为它基本上就是你看不到的转录部分。但对于你们的模型,你们是否有某种循环,比如循环 Transformer?还是说不止于此?我知道你提到了基于物理的验证作为循环的一部分。
It's easy to understand how thinking tokens work because it's just basically parts of the transcript that you don't see. But for your models, do you have some sort of loop, like a loop transformer, or is there more to it? I know you mentioned physics-based verification as part of that loop.
我们模型的一个基本模块是基于扩散的头部。这和人们用于图像和视频生成的扩散模型是一样的。我们用它来生成晶体结构。扩散头部本质上是迭代的,需要多个步骤。你在不断优化预测的结构。在这个过程中,你可以将模型引导到正确的方向。
One fundamental block of our models is a diffusion-based head. It's the same diffusion models that people use for image and video generation. We use them for crystal structure generation. A diffusion head is by nature iterative; it's multiple steps. You're refining your predicted structure. As you do this process, you can steer the model in the right direction.
所以你在循环中有一个引导模拟。你进行被动扩散,观察输出结果,运行某个验证器,然后据此说“我喜欢这个”或“我不喜欢这个”,或者给出一个方向向量。是这个意思吗?
So you have a steering simulation in the loop. You go through a passive diffusion, look at what comes out, run some verifier, and use that to say "I like this" or "I don't like this" or give a vector to go towards. Is that the idea?
是的。
Yeah.
所以你的扩散头部基本上在学习某种力场,你在基于扩散的力场和基于物理的力场之间进行平衡?可以这样理解吗?
So it's like your diffusion head is basically learning something like a force field, and you're balancing between a diffusion-based force field and a physics-based force field? Is that a way to think about it?
我很难理解这些模型到底在底层学习什么。我认为可解释性本身就是一个巨大的研究挑战。你可以识别 Transformer 中某个神经元可能在想什么,但这非常困难。我们有一些内部表示,但输出是晶体结构。很难确切知道内部发生了什么。
I struggle to understand what those models are really learning underneath. I think the whole idea of explainability is a big research challenge. You can identify what a neuron in a transformer is potentially thinking about, but it's a really difficult problem. We have some internal representation, but at the output it comes out as a crystal structure. It's really hard to tell what exactly is happening inside.
有没有可能在这个背景下做类似标准机制可解释性的事情,比如稀疏自编码器或更复杂的方法?还是说做这类事情有问题?
Is it possible to do something like standard mechanistic interpretability, like sparse autoencoders and more complicated things in this context? Or is there a problem with doing that kind of thing?
我认为作为一个研究方向会很有趣。但我们目前没有在追求这个。未来可能会探索。
I think it would be interesting as a research direction. That's not something we are pursuing. Potentially we could explore this.
好的,明白了。
Okay. Yeah.
Sergey 比我更有资格谈论 LLM 在我们领域的可迁移性。Sergey 很谦虚,但他曾在 Meta 领导 Llama 2 研究团队。他训练了地球上使用最广泛的语言模型之一。不过,在物理和可解释性方面,我想补充的是:我喜欢从输入、模型本身和输出来看待 AI 模型。我们发现,在尽可能多地包含物理先验(当然不能偏向我们渺小人类的信念)的情况下,结果和可解释性都会得到改善。但我一直认为,AI 从根本上说就是表示学习。
Sergey is a lot more credible than I am talking about LLM translatability to our space. Sergey is being humble, but Sergey led the Llama 2 research team at Meta when he was still there. So he's trained one of the most widely used language models on the planet. But I guess what I can add on the more physical and interpretability side is: I like to look at AI models in terms of the inputs, the model themselves, and the output. We find it really improves outcomes as well as interpretability—to your question—when you can include as many physical priors as possible, without biasing the model to what we puny humans believe. But I've always had the view that AI fundamentally is representation learning.
对我来说,这其实和语言与视觉领域发生的情况没有太大不同。当我们开始使用卷积神经网络时,我们是在对图像施加一个真正的人类先验——我们说,图片就是像素网格,这有其内在特性,那是人类的构造,我们在此基础上构建了卷积网络。而语言则是一维的 token 序列,我们先构建了 RNN,然后是 Transformer,但本质上仍然是在对如何查看这些数据施加一个相当强烈的先验。我们在这里的看法也差不多。从输入的角度看,我们领域相比传统 AI 领域的劣势在于,我们没有互联网可用。我们不能直接下载 Reddit 帖子,或者订阅《华尔街日报》,然后训练一个模型,预训练模型就工作得相当好。最接近的等价物是 Sergey 提到的 RCSB 蛋白质数据库,它只有大约几十万个晶体结构或其他结构,不过每个结构的 token 很多,所以这个数字描述的潜在信息要多得多。因此,我们必须在输入侧更聪明,就像 RJ 提到的,生成更多的预训练数据,在模型架构本身中尽可能多地利用物理知识,目的是让模型不必重新学习太多物理知识,从而减少过拟合的可能,然后在输出阶段强制施加物理性。
And to me, that's actually not that different than what's happened in the language and the vision domains where we're enforcing a real human prior on images when we started using convolutional neural nets is we're saying well pictures are grids of pixels. There's something inherent about that. That's a human construct and we constructed convnets on top of that and language as sequences of tokens in one dimension. We first built RNNs, then build transformers, but it's still that's baking in a pretty serious prior on how to look at that data. And we don't view it very differently here. In our case, from the input perspective, the disadvantage that we have in our field over the more traditional domains of AI is that we don't have the internet to work with. Like we can't just download Reddit posts and buy some subscriptions to the Wall Street Journal and train a model and voila the pre-trained model works fairly well. The closest equivalent that Sergey mentioned is the RCSB protein data bank which has more like a couple hundred thousand crystal structures or structures and others although there's a lot more there's a lot of tokens per structure so there's a lot more latent information that that figure would describe. So we have to be clever on the input side which is generating more pre-training data as was RJ's mentioning using physics as much as possible in terms of the model architecture itself the idea is to not have the model have to relearn as little physics as possible so it's less likely to overfit and then at the output stage enforcing physicality.
我认为这也关系到我们公司的一个主要关注点,那就是从一开始我们就专注于小分子和中等大小分子的发现。这是什么意思呢?小分子通常是指可以制成药片服用的药物,也就是口服生物可利用的,这是大多数人对药物的理解。而中等大小分子,比如大环化合物、多肽,这些突破了传统“五规则”但仍在医学中不断发展的药物形式。尽管如此,小分子仍然占 FDA 批准药物的 65%,所以我们谈论的是最大的一块蛋糕。我们从第一天起就明智地聚焦于药物发现者、药物化学家、CAD 科学家需要什么,才能比以往更快、更好地发现药物。因此,我们希望从一开始就构建不仅具有人类可解释性,而且具有可用性,特别是物理可用性。我们希望这些模型的输出能够被需要 3D 坐标才能实际起作用的物理方法所使用。用术语来说就是力场。我们希望模型的输出能够作为力场的输入。有几篇论文发表了,我记得去年有一篇在《细胞》上,它表明尽管有很多关于 AlphaFold 解决药物发现的声称,但人们尝试将 AlphaFold 生成的蛋白质结构用于传统对接,却发现没有价值。那些结合口袋的分辨率不够高,无法用于物理筛选方法。所以我们希望构建的系统能够从一开始就被人类使用,被来自物理学和计算化学社区的众多相邻且强大的工具使用,实现互操作性。
I think that also goes to one of our main focuses as a company which is we've been focused from the beginning on small and medium-sized molecule discovery. What does that mean? So small molecules tend to be drugs that you can take as a pill. So it's orally bioavailable. It's how most people think of medicine and medium-sized molecules. So think about macro cycles, peptides, modalities that break the traditional rule of five but are still growing modalities in medicine. That said, of course, small molecules are still 65% of FDA approved drugs. So we're talking about the biggest part of the pie. We're building this with a judicious focus from day one on what do drug hunters, what do medicinal chemists, what do CAD scientists, what do they need to discover drugs faster and better than they could before. And so we wanted to build that sort of human not only interpretability but usability from the beginning and also physical usability. We want outputs from these models to be used by physical methods that require 3D coordinates to actually make sense. Use the term force field. We want the outputs of our model to be useful as inputs to a force field. There was a few papers that came out. One was in Cell I think last year which showed that for all the claims about AlphaFold solving drug discovery people try to take AlphaFold produced protein structures use them for traditional docking and found no value in it. Those pockets just weren't high resolution enough. They couldn't be used by physical screening method. And so we wanted to be able to build our systems they'd be useful by humans, useful by all the many adjacent and powerful tools from the physics and computational chemistry communities from day one to have that interoperability.
是的。所以现在可能是个好时机,稍微退一步看看。当你开发一种药物时,有很多步骤,从发现、毒性到可用性。这些都有各自的名称,我让你来陈述,然后是临床部分,最后希望是批准。你们的模型处于哪个位置?或者说它们涉及了过程的哪些部分?步骤有哪些?简单来说,我读到过大概有 12 个步骤,也许我不知道,药物发现中的“接受”是最后一步,除非你的三期试验失败了。没错。那么,你们的模型在发现或开发过程中如何帮助人们完成他们的工作?
Yeah. So this might be a good time to just step back for a minute. When you develop a drug, there are many steps to that from discovery and toxicity and availability. And so there's names for all these things which I'll let you state and then there's the clinical stuff and finally hopefully approval. Where do your models sit in that or which parts of that process do they say and what are the steps? Just like briefly I've read that there's 12 steps maybe I don't know 12 steps in drug discovery acceptance is the last one except that your phase three trial failed. Exactly. Yeah. So where do your models help people to do their jobs in the discovery or in the development process?
正如你正确指出的,如果我们从头到尾来看。首先,我们需要确定是什么靶点导致了疾病,是什么信号级联出了问题,是什么导致了患者的表型。这并非易事,但有很多疾病——当然不是全部——其病因其实非常明确。最明确的是单基因疾病,但即使在其他疾病中,通过各种检测也能清楚看出是某个过表达、过度活跃或缺失的基因或蛋白质导致了疾病。所以我们知道很多(但不是全部)疾病的原因。然后,一旦确定了靶点,我们就需要为这个靶点寻找药物,理想情况下药物要有选择性、效力强,并且能到达正确的组织。有些疾病是全身性的,有些则非常特定于某个器官或组织。你需要确保药物能到达那个组织。接下来是 GLP 讨论和 IND 申报的过程。IND 就是研究性新药。这里的 GLP 不是“良好实验室规范”,谢谢,我们现在不是在讨论代谢空间。然后是临床试验,通常规模递增:一期主要关注安全性,二期和三期,最后如你所说,希望获得 FDA 或 EMA 等监管机构的批准。我们的观点是,人工智能最高杠杆的应用是药物发现和药物设计过程。原因在于,尽管所有环节都有价值,我也为所有同行喝彩,其中很多人我认识,他们在不同层面工作,但首先要明确的是,就像视觉模型和编码模型非常不同一样,它们可能在底层有一些相似之处,但要正确做好任何一个领域都需要真正的专注。从靶点识别(即生物学)到药物发现,再到准备监管文件,再到帮助临床试验的患者分层,这些模型之间的差异即使不是更大,也是类似的。这些都是非常重要、互补但最终截然不同的问题。我想到了很多这样的情况:患者去看医生,得到了非常明确的诊断。此外,医生会告诉他们,我们已经取得了突破,理解了这种疾病发生的原因。我们甚至对你的基因组进行了测序。我们知道或者测序了你的肿瘤。我们知道它导致了你的病情,但我们没有针对性的疗法。
As you rightly point out if we go from the very beginning to the very end. First we need to identify what target is causing the disease what signaling cascade what's going wrong what's causing the phenotype of the patient. And that's not a trivial exercise, but there are many certainly not all but many diseases where the cause is actually very clear. The most clear are the monogenetic diseases but even in others it's from various panels it's clear that this overexpressed or overactive or deleted gene or protein is causing the disease. So we know many but not all of the universe of causes of diseases. Then once a target is identified, we have the process of finding a drug for that target that's ideally selective, very potent. It gets the right tissue. Some diseases are multi-system. Some are very specific to a certain organ or tissue. You want to make sure that drug can get that tissue. And then there is the process of GLP talks and IND enablement. IND investigational new drug you have. GLP in this case is not good not not we're not good lab practices thank you we're not talking about the metabolic space right now. There's clinical trials usually ascending size phase one starts typically more with safety then phase two and phase three and as you point out ideally approval at whatever regulatory body the FDA or EMA or what have you. Our contention is that the highest leveraged application of artificial intelligence is the drug discovery and drug design process. And the reasons are that even though they're all valuable and I cheer on all the peers, many of whom I know who are working at all the different parts of the stack, but I think the first thing to make clear is just like a vision model is going to be very different than a coding model. They might share some similarities under the hood but requires real focus to do any one area correctly. There's a similar if not greater difference from the sort of models you need for target identification i.e. biology to drug discovery to preparing for regulatory filings to helping segment patients for clinical trials. These are all very important complementary but ultimately distinct problems. I think about all of the times where a patient goes to a doctor, gets told that they have a very clear diagnosis. And in addition, the physician will tell them there's been breakthroughs where we understand why this disease happens. We've sequenced your genome even. We know or sequenced your tumor. We know it's causing your condition, but we do not have a selective therapy.
针对你的病情,目前没有精准医疗,但至少他们知道问题所在,却不知道如何治疗。我们的目标是尽可能让更多患者被告知他们患有非常具体的疾病,并且有特定的治疗方法。这需要降低发现和开发新药的成本曲线,但更重要的是解决那些所谓的“不可成药靶点”——那些对传统方法耐药甚至顽固的蛋白质,我们需要找到成药的途径。为了澄清我的理解,有发现阶段和临床阶段。在发现阶段,你可以努力找出哪个靶点对这种疾病影响最大,或者为了解决某个医学问题,我应该追求什么。然后临床阶段回答的是这个东西是否真的有效?两者之间甚至可能存在反馈循环。但如果你缺少中间环节,比如“我们如何实际构建这种药物,使其能够击中靶点”,它不够选择性,所以会杀死细胞或对细胞做我不想它做的事情,那就不重要了,对吧?我可以知道答案,这种药应该针对这个靶点,但如果我不知道如何有效、强力地做到这一点,那么答案就无关紧要。你是这个意思吗?
There is no precision medicine for your condition and it has the hope that at least they know what's wrong with me but they don't really know how to treat it. And our aim is to have as many moments as possible where patients are told they have a very specific condition and there's a specific treatment for them. That's going to require bending the cost curve of discovering and developing new medicine, but also importantly solving those cases where there are certain quote undruggable targets, undruggable proteins that we need to figure out how to drug, where those targets have proven resistant to traditional methods or even intractable in some cases. Just to clarify my understanding, there's the discovery part and then there's the clinical part. And so there can be effort there where okay, I want to find what's the target that is impacting this disease the most or how do I maybe what's the thing that I want to go after in order to solve a medical problem. And then there's the clinic where you're answering does this thing actually work? And maybe there's a feedback loop between them even. But if you don't have the thing in the middle that's like, okay, here's how we actually build this drug so that I can hit my target. It's not selective enough so it kills cells or does things to cells that I don't want it to do. So then it doesn't matter, right? Like I can have the answer, this drug should go after this target, but then if I don't know how to do that effectively, potently, etc., then the answer doesn't matter. Is that kind of what you're saying?
是的。人们喜欢争论临床试验的成功率,但现实是,如果你关注那些针对与疾病有密切遗传关联的蛋白质的候选药物,或者至少生物学机制明确、动物模型能很好地转化到疾病、分子预测具有良好的药代动力学且患者血清中的药物浓度足够高,这些分子的 FDA 批准率相当高。从 1 期到 3 期结束的成功率远高于平均水平。人们喜欢引用 10% 的成功率,但那实际上低估了。因为我们通常知道哪些基因、蛋白质、靶点导致了疾病。它们只是很难成药,或者我们给患者的分子不如它们应有的目标产品特征。如果你只看临床前数据,专注于具有良好生物学特性、预测分布良好且安全性好的靶点,这些分子很可能会获批,并为患者创造巨大价值。因此,我们的观点是,我想到那些进行试验、渴望这类分子的医生,以及那些寻求更具选择性疗法的患者。所以,我们认为这是 AI 在更广泛的医疗保健领域中最具杠杆效应的应用。
Yeah. R.J. if you because people love to debate about success rates in clinical trials, but the reality is if one focuses on those drug candidates that are aimed at proteins that have a close genetic linkage to a disease and/or at least where the biology is well understood, where the animal models translate well to the disease, where the molecule is predicted to have good pharmacokinetics and the levels in the blood in patients are in serum are going to be high enough. Those molecules have fairly high FDA approval rates. The success rate from phase 1 to end of phase 3 is substantially higher than average. People love to cite the 10% success rate, but that's really a low ball. Because we often know what genes, what proteins, what targets are causing the disease. They're just really hard to drug or the molecules we put into patients are inferior to the target product profile that they deserve. In the cases where if you just look at the pre-clinical data, you focus on good targets with good biology that are predicted to be distributed well and have good safety profiles, those molecules are very likely to get approved and create tremendous value for patients. And so our view is I think about the physicians that run trials that are clamoring for those kinds of molecules. I think about the patients that are looking for more selective therapies. And so our view is that is the highest leverage application of AI in healthcare medicine more broadly.
这里有个商业问题。那些靶点生物学清楚、结构清楚,你唯一需要做的就是找到正确的锁,这类靶点似乎可能已经被挑选过了,从某种意义上说,它们是容易的靶点,对吧?那么这类机会有多大?
Just a business question here. What is the space of things which are that the target is the biology is understood, the structure is understood and that you know the only thing you really needed to do is find the right lock that seems like you probably would have already picked over that class of targets in some sense those are easy targets right so how much opportunity is there for that?
所以有两个正交的概念:已知的生物学与靶点成药的难易程度是正交的。不幸的是,有时它们似乎呈负相关,因为从验证角度看最有吸引力的靶点往往很难成药。
So there's two orthogonal concepts: known biology is orthogonal to the ease with which one can drug that target. Sometimes unfortunately it seems they anti-correlate in that often it seems that the most appealing targets from a validation perspective seem to be really hard to drug.
不,但它们是负相关,还是我们只是摘完了所有同时解决这两个问题的低垂果实?
No, but do they anti-correlate or is it just that we have picked all the low hanging fruit which do solve both those?
我认为很可能的情况是,许多所谓的“更容易”的靶点,由于各种原因,其中很多已经被成药了,但并非全部。即使在这些情况下,人们也喜欢制造一种错误的二分法,即容易与困难靶点,或者这个已被成药、那个已被摘取,因此就没有了。在 Genesis,我们大部分工作是通过大型制药公司进行的。我们主要做的事情就是为吉利德等大型制药公司提供 AI 服务,我们最近宣布了与 Insights 合作的扩展,我们对此感到兴奋。我不能透露他们正在研究的具体靶点细节,但我可以说,从首创化学物质开始,范围很广。这意味着我们认为这个靶点会导致疾病,但还没有已知的分子能结合它。我们需要找到第一个结合剂,真正的从零到一案例。但我们研究的靶点种类繁多,更多的是我所说的 1 到 10 案例,有时有已获批的药物,有时只有临床前分子但不理想。如果你能改进这些临床前分子,使其成为开发候选药物,准备进入患者试验,或者改进现有的临床或已获批药物,就能为患者创造巨大价值。我举一个公开的例子。我们完全不研究这个靶点,但看看 ALK 抑制剂的发展历程,具体来说,看患者生存曲线图。你可能会说,当第一个 ALK 抑制剂问世时,为什么还需要另一个?我们已经把 ALK 成药了,完成了。但这是质的差异,而不仅仅是量的差异。当你比较第一代 ALK 抑制剂和后来几代 ALK 抑制剂的患者生存曲线时,患者活得更久,受益人数更多。所以我认为,不仅首创药物有价值,同类最佳药物对患者也有巨大价值。我认为这两个领域都是我们关注的,这样说你能理解吗?
I think it is likely the case that many of the quote easier targets for a variety of reasons have been sort of many of them have been drugged not all but even in those cases people love to have this false dichotomy of easy versus hard targets and well this one's drugged versus it's picked over therefore it's not. We at Genesis do most of our work through large pharma. Like most of what we do is providing AI services basically to major pharma companies like Gilead and we recently announced our expansion of our collaboration with insights which we're excited about and I can't give details of specific targets that they're working on but what I can say is there's a real range from first-in-class chemical matter. What that means is we believe this target causes a disease. There's no known molecules that bind to that target. We need to find the first binders ever, the true zero to one cases. But there's a wide variety of targets we work on that are more I'd say 1 to 10 cases where sometimes there is an approved agent or sometimes there are molecules that are only pre-clinical but they're not optimal where if you can improve upon those preclinical molecules and get them to development candidates they're ready to get into patients or if you can improve upon the existing clinical or approved agents you can create a lot of value for patients. I'll give a public example. We're not working on this target at all, but you just look at the progression of ALK inhibitors, ALK inhibitors, and you look at, to make it very concrete, charts of patient survival. You might have said, well, when the first ALK inhibitor came out, why do we need another? We've drugged ALK. We're done. But it is a qualitative, not just a quantitative difference. When you compare patient survival curves from first generation ALK inhibitors to later generation ALK inhibitors of how much longer those patients live, how many more get benefit from it. And so I think there's not only value in first-in-class, there's enormous value for patients in best-in-class too. And I think both are areas that we focus on if that makes sense.
好的。到目前为止,我们讨论了 Genesis 建模在蛋白质配体结合这个特定问题上的应用。你们在开发这个方面整体上是如何运作的?我的意思是,我假设你们肯定——我确实知道你们还研究过解决更广泛的早期药物发现问题的其他方面。我一直开玩笑说,当诺贝尔奖颁给 AlphaFold 3 时,很多人认为药物发现已经被解决了。
Okay. Yeah. So we so far have talked about Genesis modeling in terms of let's say this one specific problem of protein ligand binding. How does your overall world work in terms of developing this? I mean I assume that you must I mean I do know that you have other things that you've worked on in solving the broader early phase drug discovery problem. I keep joking about that when Nobel Prize was given for AlphaFold 3, a lot of people thought that like drug discovery is solved.
而事实远非如此,差得很远。
And it's very far from the truth, very far.
是的,你也许可以预测晶体结构,对吧?
Yeah, you can maybe predict crystal structure, right?
在低分辨率下。
At low resolution.
在低分辨率下。嗯,我认为我们在进步。
At low resolution. Well, we're doing better, I think.
是的。
Yes.
假设你能预测晶体结构。
Assume you can predict crystal structure.
这是否意味着药物发现已被解决?不,显然没有。
Does it mean the drug discovery is solved? No, obviously not.
是的,更广泛的社区,比如蛋白质折叠社区和药物发现社区,立即指出还有其他需要考虑的因素,比如动力学,仅凭一个静态结构不足以理解相互作用的机制。所以除了静态结构之外还有很多东西。我认为很多人认为这些非常有用,确实加速了科学进展,但这显然是一个起点,而不是终点。
Yeah, the broader community like protein folding community and drug discovery community immediately said there's other things you care about like dynamics, like just a single static structure is not enough to understand what's going on interactions. So there's lots of things there in addition to just having a static structure. I think many people thought those were very useful, like they've really accelerated science, but it was very clearly a starting point and not an end condition.
没错,但如果你看看领域之外的人,比如普通大众,普遍认为这个问题已经解决了。
Right, but if you look outside of people working in the field, like general popular community, it was a pretty widely held belief that the problem is solved.
实际上,除此之外——我必须澄清这一点,否则整个机器学习结构生物学社区都会跳出来说:“不,不,这不是一个已解决的问题。”必须加上这些提醒。
Now in reality, besides—I just have to clarify that otherwise the entire ML structure biology community will jump on us and say, 'No, no, it's not a solved problem.' Got to throw those caveats in.
绝对不是一个已解决的问题。
Absolutely not a solved problem.
是的。实际上,你需要预测许多其他性质,比如我还没提到的 ADME 性质。基本上,你是在设计一把钥匙来开一把锁——那个在体内与蛋白质结合的小分子。但你不希望这个小分子与体内所有东西都结合,因为这可能会引起问题。你还得确保这个小分子是可溶的,这样你才能以药片形式服用。这些也是重要的性质。你还要确保它没有其他副作用。预测所有这些性质与预测晶体结构本身同样重要,甚至更重要。当然,作为一家公司,我们对此非常自豪。我们不仅仅在构建预测晶体结构的模型,我们还在构建预测所有这些性质的模型,基本上让药物研发人员的工作效率提升 100 倍。
Yeah. In reality, you need to predict so many other properties, like ADME properties that I haven't mentioned. Basically, you're designing a key for a lock—the small molecule that sticks to your protein in the body. But you don't want this small molecule to stick to everything in your body, because that's probably going to cause issues. You also want to make sure that small molecule is soluble so that you can ingest it as a pill. Those are important properties too. You want to make sure it doesn't have any other side effects. Predicting all of those properties is just as important as, or maybe even more important than, predicting the crystal structure itself. And of course, as a company, we're very proud of it. We're not just building models for predicting crystal structure; we're building models for predicting all of those properties and basically enable drug hunters to become 100x more effective in their daily job.
我对你们的管线很好奇。你提到你们有不同的制药合作伙伴,比如 GSK 或其他公司。他们是否有由你们设计的靶点药物已经进入临床试验或获批?我们在这方面进展如何?
I'm curious about the pipeline that you have. You mentioned you have different pharma partners, GSK or whomever you're working with. Do they have targets that were sort of drugs that were designed by you that are in clinical trials approved? Like where do we stand with all that?
关于制药公司我们能说的很有限;他们以保密著称,这可以理解。我可以举一个最近公开披露的例子:我们刚刚扩大了与 Insilico 的合作,大约一年多前开始初步合作。这在一定程度上涵盖了药物发现过程的两端。我们合作的一个项目是一个极具挑战性的靶点,已有化学物质存在,但关键一步是获得 DC(开发候选化合物),即一个分子或理想情况下的一组分子,都有可能成为一期临床试验的药物。我们合作时,将我们的基础模型在 Insilico 的数据上进行微调,并紧密协作以更快地获得 DC。在那个案例中,我们取得了实质性进展,这也是我们很高兴扩大合作的原因之一。另一方面,我们合作的另一个领域是一个与严重疾病密切相关的蛋白质,但没有任何已知的化学物质——没有专利、没有论文,也没有与该蛋白质结合的小分子的共晶结构。因此,我们必须找到第一个已知的与之结合的化学物质,然后推进这些所谓的“命中”化合物。命中化合物是第一个与你的蛋白质结合的分子。然后将它们推进到在生化测定(更基于酶)和细胞测定(即在活细胞模型中分子也有活性)中具有活性的抑制剂。正是基于这些具体成果,我们希望扩大合作。除此之外,我们能分享的信息非常有限。但 Genesis 的目标是创造患者渴望的药物。实现这一目标的方式是与尽可能多的制药和生物技术合作伙伴合作,他们的比较优势在于发现、临床开发和商业化,而我们的比较优势在于 AI。我们将这两种专长以协同而非简单叠加的方式结合起来,共同制造出原本不可能的药物。这就是我们追求的目标——你提到的那些临床成果——我非常期待在未来几年中分享这些成果。
We're limited in what we can say about pharma companies; they are famously secretive, which makes sense. What I can say, which was recently publicly disclosed as an example, is we just expanded our partnership with Insilico, and that started approximately a little over a year ago with initial work together. That spanned the two sort of bookends of the drug discovery process in some way. One of the programs we worked on together was a case where a very challenging target where chemical matter existed, but there's a key binary event which is getting to a DC, or development candidate, and that is the molecule, or ideally a set of molecules, all of which are possible to be the agent for a Phase I clinical trial. We had to do some work together where we take our foundational models, our base models, we fine-tune them on Insilico's data and use them in close collaboration to work together to get to a DC faster. And we're getting substantially closer in that case, which is one of the reasons we're excited to expand our work together. On the flip side, one of the other areas we worked together was on a protein with very nice linkage to a very severe disease where there was no known chemical matter—no patents, no papers in this case, no co-crystal structure of a ligand (another synonym for a small molecule) that bound to that protein. So we had to find the first ever known chemical matter to bind to it, and then we progress those what are called hits. Hits are the first molecules that bind to your protein. Progress those into inhibitors that are active in biochemical assays (which are more enzyme-based) and cellular assays (meaning in an actual living cell model, your molecule is active as well). So it was really based on a concrete set of accomplishments together that we want to expand the collaboration. Unfortunately, outside of that, there's really limited information we can share. But the objective of Genesis is to create medicines that patients wish they had. The way we'll be able to do that is by working with as many pharma and biotech partners as possible, for whom their comparative advantage is discovery, clinical development, commercialization, and our comparative advantage is in AI. We can put those two expertises together in a synergistic, not just additive, way and make medicines together that otherwise would not be possible. That's the name of the game we're in for—those clinical outcomes that you're pointing out—and I'm very excited to be able to share those as they arise in the coming years.
你们在网站上提到的一个点是,一埃(angstrom)阈值下,蛋白质结构预测以及它与小分子的结合变得有用。你能谈谈为什么是这样吗?当其他人失败时,你们是如何做到的?你提到过模型中的偏差、合成数据,可能还有更多。这如何影响下游的 ADMET 等性质?
One thing that you guys talk about on your website is this one angstrom threshold at which a protein structure prediction and the binding between that and a small molecule becomes useful. Can you talk a little bit about why that is the case? How did you do it when others have failed? You've spoken a little bit about that with the biases in the model, the synthetic data, maybe there's more. How does that impact downstream things like ADMET and all that stuff?
当我们谈论分子相互作用时,人们通常测量的两埃尺度太大了。你可以想象成图像生成模型:你生成一张图片,它是模糊的。所以是的,大体上就是这样,对吧?但你无法分辨细节,而细节在这里至关重要。在两埃的精度下,你的整个芳香环可能被翻转,但它仍然是一个有效的输出。最糟糕的是,与模糊的图像不同,你甚至不知道它是模糊的,对吧?你翻转一个芳香环,它看起来没问题。更糟的是,如果是一个杂环芳香环,那你就真的麻烦了。这真的很重要,因为单个原子需要在这里建立连接。所以我们追求的精度——一埃——非常重要。这有点像很多生成式图像模型,它们会填充一个根本不存在的细节。
When we're talking about molecular interaction, the scale of two angstroms that people typically measure is just too big. You can think about it like in an image generation model: you would generate a picture and it's fuzzy. So yeah, it's like in general, right? But you can't discern the details, and the details really matter here. With two angstrom accuracy, your entire aromatic ring can be flipped and it will still be a valid output. The worst part is that unlike an image which is blurry, you don't even know it's blurry, right? You flip around an aromatic ring and it looks just fine. And yeah, or even worse, a heterocyclic aromatic ring, and then you're really in trouble. And it really matters in this case because individual atoms need to establish connections here. So that level of accuracy we're pushing for—one angstrom—is really important. So maybe it's almost like with a lot of generative image models, they will fill in a detail which just doesn't exist.
所以,如果你尝试做法医分析,让你的视觉 Transformer 增强一张图像,突然它弹出一张脸,但那实际上不是那个人的脸,对吧?但你不知道,它看起来还挺像那么回事。所以,如果你试图把它作为结构假设来用,作为药物化学家,你就完蛋了。如果你作为调查员看错了脸,那是不行的。
So if you were to try to do forensics and you tell your vision transformer to enhance an image and suddenly it just pops up a face, but that's actually not the person's face, right? But you don't know that and it feels just fine. And so if you're trying to use this as a structural hypothesis, as a med chemist, you're kind of screwed. If you're looking at the wrong face as an investigator, it's not going to take.
是的,我真的很喜欢你的这个比喻。就像 Evan 说的,在下游,你运行一堆其他模型,比如基于物理的模型,所有这些小问题都会累积。然后显然,如果你做出了错误的预测,根本性的错误,那么下游的预测也会出错。我只想说,对于你的问题,我非常欣赏你提到的 2000 年代中期电视剧里增强法医图像的梗。所以我想说,我很感激。
Yeah. I really like your framing. And then as Evan said, at the downstream, you run a bunch of other models like physics-based models and all of those little issues, they compound. And then obviously if you made the wrong prediction, fundamentally wrong, then the downstream predictions are also going to be wrong. I just wanted to say for your question, I really appreciate your mid-2000s TV reference about enhancing forensic images. So I just want to say I appreciate that.
在我们回到细节之前,我们谈到了构象。所以也许逆向观点是,构象甚至不是一个定义明确的概念,思考这些事情的最佳方式是概率分布。有一个小分子可能以这样的方式存在于结合口袋中。我的意思是,这个构象是不是人类使用的一种抽象,而在某种意义上并非真实?所以我其实有点想听听,我们有这个警告阈值,这对药物化学家非常重要。我也会和一些化学家聊,他们会说这个构象甚至不是真实的,而只是一个最可能的构型。你通常如何看待这种抽象?你们是否探索构象空间并将其作为工具提供?你们如何处理单一结构 vs. 集合?我不知道这个问题是否有意义。
Before we get back into details here, we talked about poses. So maybe the contrarian take is that a pose isn't even a well-defined concept and that the best way of thinking about these things is this probability distribution over things. And there's a small molecule which probably lives in a binding pocket something like this. I mean, is this pose like an abstraction that humans use and not really in some sense ground truth? So I'm actually almost kind of interested to hear that we have this warning threshold. This is really important for med chemists. I would also talk to some chemists who would say this pose isn't even real, versus just like a most probable configuration. How do you think about that abstraction in general? And do you explore conformational space and provide that as tools? And how do you interact with just a single structure versus an ensemble? And I don't know if that question even makes sense.
是的,这是一种抽象,但它是非常有用的抽象。它帮助我们建立信心,确信特定的模型输出实际上是有效的,而不是直接凭空捏造了什么。因为最终重要的是结合亲和力或效力,你可以直接用模型预测它,跳过整个构象生成步骤。但那样你只有一个数字,而这个数字可能完全是幻觉,你没有办法验证这个数字是否有意义。所以,尽管构象并不完美,它们仍然是整个过程中非常有用的工具。
Yes, it's an abstraction, but it's a very useful abstraction. It helps us to build up confidence that a particular model output is actually valid. It did not just straight up hallucinate something. Because ultimately what matters is binding affinity or potency, and you can straight up predict that with your model and just skip the entire pose generation step. But then you only have a single number, and that number might as well be completely hallucinated, and you have no means to validate whether that number even makes any sense. So as much as poses are not perfect, they're still a very useful tool for the entire process.
抱歉,纠正一下。我担心我跳进了一个技术兔子洞,但仅仅因为你有一个构象,还有熵和焓的贡献,预测结合亲和力不仅仅是把能量算对,而是这个分子是否真的有可能进入结合口袋,并且长期待在那里?所以这远不止是亲和力效力预测。它不仅仅是“这是正确的构象吗?”,而是“这个分子是否具备在这种状态下长期存在的更广泛特性?”你们又是如何处理这些问题的呢?
Sorry, going into that correction. I'm afraid I'm jumping into a technical rabbit hole here, but just because you have a pose also, there's entropic and enthalpic contributions, and predicting binding affinity is not just about getting the energy right, but actually is this even likely to make it into the binding pocket and is it likely to live there long term? So it's much more than just affinity potency prediction. It's more than just 'is this the right pose?' but 'does this have the broader properties it needs to be a long-lived molecule in this state?' And how do you kind of deal with those things too?
给更广泛的 AI 听众一个提示,这可能会引起从业者和用户的共鸣:显然,过去六个月蓬勃发展的一个东西是智能体。我们热爱智能体。谁不喜欢呢?有很多必要但不充分的条件。我想说我对一个深刻问题的回应是:我们都记得智能体是什么样子,比如说去年年中。可以说有正面价值,也有负面价值。两者都可以被智能体放大。为什么?智能体的有用程度取决于它们所编排的底层模型。想想编码。如果你的编码模型即使产生微小但真实的 bug,你的智能体只会放大这些问题,你最终得到的不仅是垃圾,而且可能是反有用的东西,可能会给用户错误的信息。我们都记得去年年中那个时代,让很多人对 LLM 用于智能体工程的宣称失去了耐心。有些东西变了;显然达到了一个阈值。尽管这些模型仍然不完美,但 LLM 在软件工程中的效用现在已经非常明显。这对我们来说是一个巨大的推动力,非常有用,取代了大量繁琐的编码工作,并使其更专注于一些更重要的战略问题。你可以从中直接类比到我们在这里讨论的内容。我们正在开发一个用于 24/7 药物发现的智能体平台。你可以想象数百名药物化学家和 CAD 科学家夜以继日地为你的药物靶点工作。这个项目的代号是 Sapphire。前提是我们需要底层模型在构象、3D 复合物预测、效力、ADME 方面都足够好,以便智能体 24/7 使用这些模型来创造药物化学家真正想要制造、而不是嘲笑的分子。如果你的模型 RMSD 在 1.8、1.9,那很可能是垃圾。让我直截了当地说:你听到的是人们对 3D 构象的效用及其真实性的双峰分布。现实是,对于一个高效力的配体,几乎可以肯定分子的大部分具有非常明确的 3D 位置,甚至达到半埃精度。如果你不相信我,你可以打开 PDB 中的电子密度图。这些都在线可用,你实际上可以看到在某些情况下,芳香环中间缺失密度。所以这不仅仅是你的有机化学教授在黑板上展示的构造。你可以实际看到一个甜甜圈,一个芳香环的电子密度环面。然而,会有溶剂暴露的区域,这些区域对结合亲和力不太重要,但可能对溶解度或分子的其他性质很重要。这些有时定义不那么明确,因为有时在现实中它们更具动态性,在溶剂中摆动。但作为预测结合自由能的上游指标,关键部分是让与蛋白质特异性相互作用的分子核心正确到亚埃分辨率。
A note for the wider AI audience that probably resonates with both practitioners as well as users: obviously something that's blossomed in the past six months is agents. We love agents. Who doesn't? There's a lot of necessary but not sufficient conditions. I'd say my response to a deep question there: we all remember what agents were like, let's say mid last year. And let's just say there is positive value and there's negative value. And both can be amplified by agents. Why is that? Agents are only as useful as the underlying models that they're orchestrating. Let's think about coding. If your coding model even makes subtle but real bugs, your agents are just going to amplify those issues, and you're going to end up with not only slop but something that may be anti-useful, that might give the user incorrect information. We all remember what that age was like mid last year that made a lot of people lose patience with claims about LLMs for agent engineering. Something changed; clearly a threshold was met. And even though these models are still not perfect, the utility of LLMs for software engineering are so obvious now. It's been a huge tailwind and very useful for us for replacing a lot of the drudgery of coding and getting it focused a bit more on some of the more strategic issues that matter. You could draw a direct analogy from that to what we're talking about here. We are working on an agentic platform for 24/7 drug discovery. You can imagine fleets of hundreds of med chemists and CAD scientists working nights and weekends all the time for your drug targets. The code name for that gem is Sapphire. The prerequisite for that was we needed the underlying models for pose, 3D complex prediction, potency, ADME to all be good enough for an agent using these models 24/7 to create molecules that medicinal chemists would actually want to make and not laugh at. If your model is sitting at 1.8, 1.9 RMSD, that's slop most likely. Let me be really direct about it: what you're hearing is a bimodal distribution from people of the utility of a 3D pose and how real it is. The reality is that for a highly potent ligand, almost certainly there is a large portion of that molecule with a very well-defined 3D position down to even half an angstrom. If you don't believe me, you can open up the electron density diagram in the PDB. It's all available online, and you can literally see in some cases aromatic rings with missing density in the middle. So it's not just a construct on a blackboard that your organic chemistry professor showed you. You can literally see a donut, a torus of electron density for an aromatic ring. However, there are going to be solvent-exposed areas often times that will be less important for binding affinity but maybe are important for solubility or other properties of your molecule. Those sometimes are less defined because sometimes in reality they're more dynamic, flopping around in solvent. But the critical piece as an upstream indicator that's valuable for predicting the free energy of binding is to get the core of your molecule that specifically interacts with the protein to be correct to sub-angstrom resolution.
直观来说,为什么这很重要?
And why does this matter just to use it intuitively?
氢键,在座各位可能都听说过。它是最关键的非共价相互作用形式,核酸靠它连接,多数配体与蛋白质靠它相互作用,蛋白质形成二级结构也靠它。氢键有非常特定的角度和距离,距离是从供体到受体的重原子,范围是 2.7 到 3.3 埃。如果我没算错,这个区间只有 0.6 埃。超出这个范围就不是氢键了:小于这个距离是碰撞,大于这个距离相互作用会迅速减弱很多。埃,对不了解的人来说,是纳米的十分之一。所以药物发现本质上是一门分辨率的科学。如果你的精度不够,就无法用于下游关心的事情,包括效力预测和前瞻性设计——下一步该合成什么分子?因此我们认为,初创公司的创新历史表明,做得最好的那些公司都专注于一个定义明确但非常重要的问题。我们能够获得更高分辨率的预测,源于我们明智地专注于小分子设计,而不是试图包罗万象。
So a hydrogen bond everyone in the audience has probably heard about a hydrogen bond. It's the most critical form of non-covalent interactions. It's how nucleic acids are held together. It's how most ligands and proteins interact. It's how proteins form secondary structures. Hydrogen bonds have a very specific angle and distance. And the distance is from the donor to the acceptor heavy atom. It's 2.7 angstroms to 3.3 angstroms. And if I do my math right, that's a 0.6 angstrom gap. And outside of that, it's not a hydrogen bond. If it's less than that, it's a clash. If it's more than that, the interaction is much much weaker pretty quickly. An angstrom for those who don't know is one tenth of a nanometer. So drug discovery really is a science of resolution. And if your accuracy is not sufficient, it will therefore not be useful for the downstream things that you care about, which is both potency prediction but also prospective design: what molecule do I make next? So that's why our view is that the history of innovation in startups is that the ones that do best are ones that focus on one well-defined but very important problem. And we think our ability to get higher resolution predictions stems from our judicious focus on small molecule design rather than boiling the ocean.
对,这就是“是什么”和“为什么”。那么“怎么做”呢?你们是如何达到 1 埃甚至亚埃精度的?
Yeah. So that's the what and the why. So what about the how? How did you get to one angstrom and sub one angstrom?
我要给你一个极其无聊的答案。
I'm going to give you an extremely boring answer.
好。
Okay.
我认为这对整个 AI 领域都成立:AI 中三件事最重要:数据、基础设施和演化。
Which I think is actually true for the entire AI field: three things matter in AI: data, infrastructure, and evolves.
好。
Okay.
对。你只能改进你衡量的东西。一旦你非常谨慎地衡量重要的事情,并且团队里有真正有才华的人,我们就会想办法去优化那个指标。所以从一开始,我们就专注于小于 1 埃的精度,这导致过程中一系列小决策不断累积。对吧?如果你的团队根本不看这个指标,你就永远训练不出擅长它的模型。如果你一直关注它,你就会实现它。所以正确的目标加上好的科学,它会渗透到整个技术栈的方方面面。它会渗透到你如何看待数据:比如我们如何过滤数据?有些数据噪声更大,也许我们不需要看到它,或者训练后期不需要看到它。它会渗透到你的模型架构和损失函数中。
Right. So you can only improve what you measure. Once you are very careful about measuring what matters and you have really talented people on the team, we're going to figure out how to hill climb that measure. So from the start we actually focused on less than one angstrom precision, and that led to a bunch of small decisions in the process that compound. Right? If your team never looks into this metric at all, then you will never train a model that is good at it. And if you're constantly looking into that, then you're going to achieve that. So it's the right objective plus good science, and it propagates everywhere through the whole stack. It propagates to how you look at the data: maybe how do we filter out data? Some data is more noisy, so maybe we don't need to see it, or maybe you don't need to see it later in the training. It propagates through your modeling architecture and propagates through your loss.
你们朝着这个具体目标努力了多久?我很好奇。这不是你在社区里广泛听到有人谈论的目标。人们确实说 RMSD 小于 2 埃是标准基准,我从未听人说过 1 埃是截止值,直到读到 Genesis 的成果。你们决定这是需要优化的数字有多久了?这种专注在公司发展过程中有多直接?我只是想知道这是怎么来的。
How long have you been working towards this specific goal? I'm curious. This is not something that you broadly hear in the community talk about wanting. People do say RMSD less than two angstroms is sort of the canonical benchmark, and I've never heard someone say one angstrom is the cutoff until reading things coming out of Genesis. How long have you decided this is the number we need to hill climb on? How direct has this focus been in the evolution of the company? I'm just wondering how this came about.
我认为我们最重要的秘密之一是我们正在从事真实的药物项目,无论是与合作伙伴还是内部进行。当你真正做真实的药物项目时,你会看到失败模式,看到什么有效、什么无效。当你看到真实项目的输出时,哪种失败模式正在发生就非常明显了。而且很明显,2 埃对于这些场景根本行不通。
I think one of our most important secrets is that we are working on real drug programs, either with partners or in house. And when you actually work on real drug programs, you see the failure modes, you see what works and what doesn't work. And it's pretty obvious when you see the outputs of real programs, what kind of failure modes are happening. And it becomes very obvious that two angstroms is just not working out for those setups.
对。但我的意思是,有很多非常聪明的药物化学家,他们也非常仔细地思考基准测试,是我非常尊敬的人,他们也实际执行过药物发现项目。所以我只是好奇为什么这没有成为社区的一部分?是不是社区从未能在某个获胜基准上取得成功,于是他们就把 2 埃当作一个可以追求的目标?或者我不知道。
Right. But I mean there's a lot of really smart med chemists who also think very carefully about benchmarks, people I respect very much, and who also have prosecuted actual drug discovery programs. So I'm just kind of curious why hasn't this become part of the community? Is this literally just that the community has never been able to succeed at a winning benchmark that they kind of settled on two angstroms as something we can aspire to? Or I don't know.
但我的意思是,根据我在制药行业的经验,有很多问题在子领域的技术专家中是已知的,但这些知识没有流出制药行业,或者没有得到应有的关注,因为部分信息是专有的,在公司之间流传但从未真正公开发布。这是你对现状的估计吗?
But I mean to your point, it sounds like in my experience with pharma, there are plenty of problems for which there are these known things amongst the technical experts in a subdomain, but it doesn't get out of pharma or it doesn't get the attention it deserves because some of the information is proprietary and gets passed from company to company but never really released into the public. Is this your estimation of what's going on?
你听说过 SWE-bench 吗?
Have you ever heard of SWE-bench?
所以呢?
So?
Gemini 在 SWE-bench 上表现相当不错。有时 Gemini 发布模型,在那些软件基准测试上获胜。现在请举手,如果你正在用 Gemini 写代码,而不是其他明显的竞争对手。
Gemini does pretty well on SWE-bench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors.
对,不是我。
Yeah. Not me.
没有人。谁会这么做?它在实践中明显更差。我想说,实际上在这个案例中,如果你看看 RMSD 小于 2 是怎么来的——我长期在这个领域——它最初来自对接研究,远在 AI 模型用于姿态预测之前,来自学术机构的基于物理的对接研究。因为通常专有软件制造商不愿意将自己的方法与其他方法进行基准比较,所以学者们必须获得许可并尝试,然后他们引入了 RMSD 小于 2,这并不奇怪。他们是学者,不是用这些东西来制造药物,而是用来写论文。所以这种来源被 AI 社区重新利用了。但实际趋势更符合我们所说的方向。第一个重大创新是 PoseBusters,由牛津大学的一个实验室提出,它指出 RMSD 本身是不够的,我们还需要关注物理有效性。所以我说的 PoseBusters 不是基准测试,而是指标。牛津改进了这一点。在 OpenBind 的最新版本中,我们刚刚在 Genesis 上发布了基于 OpenBind 数据集的基准测试。我们稍后会讨论。但如果你看他们的原始论文,他们的默认指标使用了 RMSD、PoseBusters 有效性和 LDDT。所以很明显,已经有人承认 RMSD 小于 2 是不够的,这个领域正在迅速演变以承认这一点。
No one. Like why would you do that? It's obviously worse in practice. And I'd say actually that in this case, if you look at the provenance of how that happened, RMSD less than two came originally from again, I've been in the space forever, from docking studies long before AI models for pose prediction, when physics-based docking studies from academic institutions. Because usually the proprietary software makers didn't want to benchmark their methods against other methods, so academics had to get licenses and try it, and then they'd introduce RMSD less than two, which is not surprising. They're academics. They're not using these things to make drugs. They're using it to write papers. And so that sort of provenance got repurposed by the AI community. But the actual trend is much more in the direction that we're saying. So the first big innovation there was PoseBusters, put out by a lab at University of Oxford, which pointed out that RMSD itself is insufficient and we need to look at physical validity as well. So that's PoseBusters I'm talking about, not the benchmark but the metric. So Oxford improved that. And in the latest release from OpenBind, we just published our benchmarks at Genesis on OpenBind set. We'll talk about that in a bit. But if you look in their original publication, their default metrics are using RMSD, PoseBusters validity, and LDDT. So there's clearly an acknowledgement that's already been made that RMSD less than two is insufficient, and the field is now rapidly evolving to acknowledge that.
你认为学术文献会建立某种基准,也许是 instro,也许是某个 LDD 或其他指标,这些指标会逐渐收敛,然后希望作为社区推动更准确的建模策略。我想说,我们领域曾存在评估危机,现在正处于转型期。以前这个领域相对安静,但如果你每年去 NeurIPS 和 ICML,研讨会一年比一年拥挤。这些评估正在转型,因为我们意识到了之前方法的缺陷。
You think that the academic literature is going to establish some benchmark, maybe it's one instro, maybe it's some LDD or some other metric that will kind of converge and then hopefully as a community drive forward these more accurate modeling strategies. I would say there was an eval crisis in our field that is now in transition right now as our field which was previously, let's say, a lot more quiet, but if you go to NeurIPS and ICML every year, year after year, the workshops get a lot more crowded. Those evals are in transition now that we realize the flaws of what came before it.
你提到选择正确的评估,然后像平常一样做好机器学习,这很有意思。但我注意到 Evan 之前说过,你们一开始用的是所谓的 potential net——我想你之前没定义过——就是那种基于图的网络。之后你们做了大量计算模拟,现在生成式建模兴起后,你们又回到了这个方向。我很好奇这个演变过程是怎样的?计算技术在一段时间内对构建 ML 有多关键?是计算反过来推动了生成式建模,还是你们开始收集数据来构建生成式模型?计算数据就像是生成模型的数据记录方式。这段历史是怎样的,是什么导致了这些决策,你能谈谈吗?
That was a really interesting point about picking the right evals and then just doing good machine learning like you normally would do. But I kind of picked up on something Evan said earlier, which is that you started out with, I guess, what we call potential net. I don't think you defined that earlier, but this graph-based network. And then you did a lot of computational simulations kind of after that, and now you're sort of going back into generative modeling once it took off. I'm curious about how that evolution worked and how did you find computational techniques to be really crucial to building on the ML for a while? Did that computation lead back into generative, or did you start collecting data to build a generative model? Computational data is like a way of data documentation for generative models. What's the history of that and what led to those decisions, to the extent you can talk about?
我得说,我写的最后一行 PyTorch 代码,比 Sergey 最近提交的那行要早得多。所以我想让他也发表一下看法。
I will say that the last line of PyTorch I've written is much further back in history than the most recent line of PyTorch that Sergey has committed. So I want to make sure that he gives his opinion as well.
我想确保我直接回答你的问题。我还想回应你之前提到的一个问题,我觉得它可能被忽略了,就是你问我们如何做其他事情,比如 AD 预测。所以我想就此说几句。我可以花一整天时间,也很乐意深入讨论 Pearl 和它的演变,我们会的。但我想强调一点:它是药物发现过程中的一个重要支柱,但不是唯一的支柱。你提到了 potential net 论文,我们很高兴看到它变得有影响力。当年你和我在这个领域工作时,它还非常未来,但现在它已经是现实了,这很棒。大约同一时间,我们还发表了一篇关于 ADME 预测的神经网络论文,实际上我们发表了两篇。
I want to make sure that I'm going to directly answer your question. I want to make sure that I address something you said a little bit of time ago that I think kind of got lost, which is you asked about how we do other things including AD prediction. So I just want to make a comment about that. I can spend all day and I'm happy to dive into details about Pearl and the evolution, and we will. I just want to give a shout out to the fact that it is one important pillar but not the only pillar that matters in the trajectory of a drug discovery campaign. And you mentioned that potential net paper, which we're happy to see had become influential since you and I were working in the space when it was very much in the future, let's say, but now it is the future, so it's great. Another paper we published around the same time was on neural networks for ADME prediction, and we published two actually.
哦,抱歉,谢谢。吸收、分布、代谢、排泄和毒性,这五个方面构成了 ADMET。
Oh sorry, thank you. Yeah, absorption, distribution, metabolism, elimination, and toxicity are the five prongs that comprise ADMET.
这些就是你之前提到的那些性质,要想让药物成功,这些性质都必须达标。
These are all the properties you talked about before that you just have to get right in order to make a drug successful.
没错。所以大概有 30 多种检测,你可以想象,如果你是做神经网络的人,就像一个多任务神经网络或多头网络,它需要预测 30 多种性质,如果其中任何一个不在正确范围内,你的分子就不是药物,而只是一个工具。这些性质包括溶解度,Sergey 提到过,这对制剂很重要,比如能否做成口服药片。还有口服生物利用度,是否抑制某些酶,比如细胞色素 P450 及其不同变体。hERG 是一个关键通道,抑制过度会导致心脏毒性。这些就像字母汤一样,大多数人没听说过。
Correct. So I would say there's over 30 or so assays, each of which you can imagine if you're a neural net person, like a multitask neural network or multi-head, it's got to predict over 30, you know, three dozenish properties, each of which if it's in the wrong range means your molecule is not a drug, it's just a tool. So these are things like solubility, which Sergey mentioned, important for formulation, as in can it be made into a pill that you can take orally. Your oral bioavailability, whether or not you're inhibiting certain enzymes called the cytochrome P450s and its different variants. hERG, a critical channel which if you inhibit it too much can cause cardiotoxicity. So these are an alphabet soup of things that most people here haven't heard of.
而且很多性质极其困难,因为这不是单一的因果效应。通常有很多过程、很多通路共同决定了这一个终点。
And a lot of these are extremely hard specifically because it's not like a single causal effect. There are often many processes, many pathways which are involved in defining this one endpoint.
没错。
Correct.
而且数据集往往小得可笑。
And data sets are often comically small.
除非有制药公司帮忙,但至少在开源领域,这是个难题。
Unless you have, I guess, pharma to help you out, but at least open source it's a hard problem.
公共领域的数据很稀疏。实际上,有些性质可以直接预测,而有些,正如你指出的,是其他信号事件的混合体。比如是否抑制 CYP 3A4,这非常具体,是抑制某个特定的蛋白质。
It's sparse in the public domain. I'd say there's actually a range from really directly predictable properties to ones that, as you point out, are actually amalgams of other signaling events. So like whether or not you're inhibiting CYP 3A4, that's really specific. It's a certain protein that you're inhibiting.
所以像 Pearl 这样的模型,实际上可以建模并期望看到一些性能。是的,有用的预测。
So that's something that Pearl, for example, could actually model and expect to see some performance here. Yeah. Useful prediction.
我们在其他人之前就开始了 AD 预测的 3D 工作。有一系列工作最终促成了公司的成立。当年 AI 还能发表在同行评审期刊上,而不只是 arXiv 上的随机白皮书,我们发表了 MoleculeNet,它被引用了数千次,我们还发表了一篇论文,证明多任务图神经网络在当时是大型制药数据集上做 ADMET 预测的最佳方法。那篇论文是 AI 用于 ADMET 领域被引用最多的论文之一。很多论文,你知道,写了就过去了,但这两项工作都变得很有影响力。所以从一开始,我们的历史就是专注于药物发现,而不是只解决一个问题。这样我们就有精力专注于构建药物发现所需的所有 ML 模型,而不被生物发现过程、靶点识别或临床试验所分心,而是真正专注于药物发现所需的所有任务,包括分子生成,对吧?这些都很重要。我想在开头就指出这一点,但正如你正确指出的,Brandon,Pearl 是一个 3D 结构预测模型,许多 ADMET 性质都可以用这个框架来建模。因此它是有用的,不仅仅用于靶点效力预测。所有这些都很重要,我们必须全部解决。
We've been doing 3D work on AD prediction before anyone else was. And there was a bunch of work that ended up launching the company. Back in the day when AI could still be in peer-reviewed journals, not just sort of random white papers on arXiv, we published MoleculeNet, which has been cited a few thousand times at some point, and we also published this paper showing that multitask graph neural networks were the best at the time for doing ADMET prediction on large pharma data sets. And that paper is one of the most cited papers on AI for ADMET ever. Now a lot of papers, as you know, it's like you write them and you kind of move on, but that work, both those works, have become quite influential as well. So our history from the beginning has been to work on not just one problem but to focus on drug discovery, and in so doing, being able to have the bandwidth to focus on building all of the ML models that are needed for drug discovery and not get distracted by the biological discovery processes that are needed, the target ID side, or the clinical trial side, but really focused on all those tasks that are required for drug discovery and also molecular generation, right? These are all important. And I just wanted to make the point at the beginning, but as you rightly point out, Brandon, Pearl is a 3D structure prediction model and many of these ADMET properties can be posed as those. So it is therefore useful, and not just on-target potency prediction as well. All of them matter, and we've had to tackle all of them.
你们对所有这些测试都用 Pearl 吗,还是只针对其中一些?
Do you use Pearl for all these tests or is it focused on some of them?
我们会在未来一段时间公开分享一些东西。但我觉得我们最近已经有足够多的结果了。到目前为止,如果这说得通的话。每次发表东西,团队都要付出很多努力。所以我们会在未来一段时间内做这件事,但最近发表的是 OpenBind 的结果,这是一个标准的 3D 预测任务。
We'll be sharing some things publicly in the coming period. But I think we've had enough news of results lately, I think. So far, if that makes sense. Every time you publish something, it's actually a lot of work for the team to put together. So we'll do that in the coming period, but the most recent thing we published was obviously the OpenBind results, which was for a standard 3D prediction task.
你刚才提到有很多合成数据和机器学习建模,带有生成式模型的特点。我想谈谈在开发这个模型的过程中,计算、机器学习和湿实验数据之间的反馈。我们讨论过先验知识。通常,蛋白质-配体模型很难在真正泛化的意义上进行扩展,但你最近的 Pearl 论文和其他结果展示了真正的泛化。我很好奇这些方面是如何协同作用,产生这种能力的。计算很有趣,但必须小心:分子动力学有偏差,可能给出糟糕的结果。这一切是如何协同工作的,尤其是在公司的发展历程中?
So you talked about there being a lot of synthetic data and ML modeling with a generative model flavor. I want to talk about the feedback between computation, ML, and wet lab data in developing this model. We discussed priors. It's been typically hard to scale protein-ligand models in a way that meaningfully generalizes, but your recent Pearl paper and other results have shown true generalization. I'm curious how these aspects feed together to give this power. Computation is interesting, but you have to be careful: MD has biases and can give poor results. How did all this work together, especially in the history of the company?
有一个经典概念叫叙事谬误:事后看来,一切似乎都显而易见,你可以画出一条从 A 点到 Z 点的线性路径,但现实总是更有趣、更混乱、更非线性。对于 Sergey 和我来说,我们训练神经网络已经超过十年,记得很多不同的时代。如果能告诉十年前的自己我们现在的技术能力,我们会非常兴奋。如果告诉自己是怎样走到这一步的,可能看起来有些明显,但有些事确实很难预见。例如,将生成式 AI 用于分子领域并不是新概念,但十年前将其付诸实践非常困难。我曾共同运营斯坦福 AI 沙龙。2017-2018 年,我们讨论 GANs 作为图像生成的未来,也有工作尝试将 GANs 用于蛋白质构象或配体姿态。但出于同样的原因——模式崩溃——它们在蛋白质上效果不佳。我们不得不等待正确的基元:扩散模型,它被证明更有用。有趣的是,图像和视频模型有些用扩散,有些则转向了自回归。现在,一些最具创新性的扩散研究发生在 3D 结构预测领域,这在十年前没人能预料。与此同时,早在用于化学的 3D 扩散模型出现之前,我们就构建了各种药物发现工具,包括基于物理的方法预测效力和分子生成。正如 Sergey 指出的,有 10 的 60 次方种类药分子;高效搜索这个空间很难。我们一直在研究这个问题,这些工具在我们想要将新兴的共折叠领域——令人兴奋但尚未成熟——提升到有用甚至不可替代的水平时,正好可用。我们构建了其他基元,所以能够将它们组合起来用于合成数据管道或推理时技术。这有一定的规划性,但就像任何发现——我不是在比作青霉素——如果你足够长时间专注于一个问题,并在实验室或电脑前埋头苦干,你就能创造运气,促成一些进展。
There's the classic concept of the narrative fallacy: in retrospect, everything seems obvious and you can draw a linear process from point A to point Z, but reality is always more interesting, messier, and more nonlinear. For Sergey and me, who've been training neural nets for over a decade, we remember many eras. We would have been so excited if we could tell ourselves the capabilities of our technologies 10 years ago. If we told ourselves how we got there, it might seem obvious, but some things were hard to foresee. For example, using generative AI for molecular space is not new, but reducing it to practice was difficult a decade ago. I used to co-run the Stanford AI salon. In 2017-2018, we talked about GANs as the future of image generation, and there was work applying GANs to protein conformations or ligand poses. But for the same reasons GANs were tricky for images—mode collapse—they didn't work well for proteins. We had to wait for the right primitive: diffusion, which turned out much more useful. Interestingly, image and video models sometimes use diffusion, but some have gone autoregressive. Now, some of the most innovative diffusion research is happening in 3D structure prediction, which no one would have predicted 10 years ago. In parallel, long before 3D diffusion models for chemistry, we built tools for drug discovery using physics-based methods for potency prediction and molecular generation. As Sergey pointed out, there are 10^60 drug-like molecules; searching that space is hard. We worked on that, and those tools were available when we wanted to take the nascent area of co-folding—exciting but not ready for prime time—and make it useful and sometimes irreplaceable. We had built other primitives, so we could put them together for synthetic data pipelines or inference-time techniques. There was some planning, but like any discovery—I'm not comparing to penicillin—if you focus on a problem long enough and spend time banging your head against the wall, you can create the luck that enables developments.
实验室如何与开发过程互动?
How does the lab interact with the development process?
正如我提到的,除了结构预测,我们还训练其他模型。对于这些模型,实验室输出立即可用——比如效力预测,这些输出可以直接用于模型训练。但我最期待的是强化学习。我认为它即将进入我们的领域,我们已经看到早期迹象表明它在我们模型上有效。最初,你可能通过基于物理的反馈,利用强化学习循环来改进模型,但最终你可以实现实验室在环的设置:你的模型生成预测,你基于这些预测进行合成,测量下游属性,然后将这些反馈回模型。
As I mentioned, we train other models besides structure prediction. For those, lab outputs are extremely useful right away—like potency prediction, those outputs can be used directly for model training. But what I'm most excited about going forward is reinforcement learning. I think that's coming in our field, and we've seen early signs of it working with our models. Initially, you might use physics-based feedback to improve models through an RL loop, but eventually you can go all the way to lab-in-the-loop setups: your model produces predictions, you synthesize based on those, measure downstream properties, and feed those back into the model.
你们如何获得足够的量?是有自动化,还是通过大规模活动生成大量数据?
How do you get enough volume? Do you have automation, or large campaigns that generate a lot of data?
我非常自豪的一点是我们与 Insight 的合作。他们是一家了不起的公司,非常擅长生成数据:获取化合物、生成它们、创建它们、测量下游属性,并将结果反馈回来。
One thing I'm super proud of is our partnership with Insight. They are an amazing company, extremely good at producing data: taking compounds, generating them, creating them, measuring downstream properties, and sending results back.
对我们来说,Genesis 和 Insight 之间的这种合作简直是天作之合,我们能够快速训练模型、给出预测并迅速获得结果反馈。
This is such a partnership kind of born in heaven for us between Genesis and Insight where we are able to train models, give predictions and have results back from inside super quickly.
所以你们的部署本质上包括了实验室迭代。
So your rollouts include a lab iteration essentially.
是的,太棒了。而且扩散过程很慢,有时速度差不多。给它一周时间。
Yeah, that's amazing. And diffusion is slow, so it's about the same speed of that sometimes. Give it a week.
确实如此。我坚信公司通常只擅长一两件事,专注时才能做到最好。同样,我们一直坚定不移地专注于开发用于药物发现的最佳 AI 模型,而 Insight 也以同样狂热的专注优化药物发现和开发。另一个人们喜欢谈论的趋势当然是中国在生物技术领域的崛起。这是房间里的大象,没必要回避。它现在非常符合时代精神,很多西方公司越来越依赖 CRO 来做大量湿实验工作。Insight 的内部实验能力变得如此先进,生产力极高,正如 Saki 所说,这对我们来说是完美的匹配。因为我们擅长的是模型的持续学习。所以我们希望设计、制造、测试、分析的循环尽可能快,并根据实验室看到的情况持续微调甚至重新训练模型。这种合作是第一个实现真正联合基础模型训练(基于历史数据和前瞻数据)的合作之一。作为机器学习从业者,我们都知道数据是关键输入,是核心要素。所以这让我们非常兴奋,从我们目前的对话来看,这涵盖了结构效力以及各种 ADMET 属性。我认为这会立即提升我们模型的能力,并极大地加速药物发现。
It is true. It's like I'm a big believer that companies are typically really good at one or two things and do best when they can focus on it. And in the same way we've been just really doggedly focused on developing the best AI models for drug discovery, like Insight has that level of maniacal focus on optimizing drug discovery and development. And another sort of trend people like to talk about is of course the rise of China in biotech. It's the elephant in the room that there's no point avoiding it. It's so in the zeitgeist right now and so many Western companies have become to rely more on CROs to do a lot of their wet lab work. Insight became so state-of-the-art in terms of the experimental capabilities they had in house, their productivity is extremely high, and as Saki said, it's a match made in heaven for us. Because what we thrive on is continuous learning of the models. So we want to have design, make, test, analyze cycles that are as rapid as possible and continuously fine-tune in some cases depending, retrain the models based on what we see in the lab. And so that partnership is one of, if not the first ever, that enables that sort of true joint foundation model training on historical and also prospective data. And as ML practitioners we all know here the data is such a critical input, the whole critical ingredient. So it's just extremely exciting for us, and that's a range you know from reflecting our conversation so far, from structure potency and a variety of ADMET properties as well. So I think it immediately will improve the strength of our models and be really powerful for generally accelerating drug discovery.
我们之前开玩笑说扩散所需的时间,但不知道你是否能透露,你给他们发邮件,比如 10 个或 100 个化合物,然后他们返回这些实际化合物的测量结果,需要多久?一天、一周、一个月还是一小时?
And we were joking around about the time that it takes to do diffusion, but I don't know if you can disclose this, but how long does it take to you know you email them with here's some, I don't know how many 10, 100 compounds, and then they get back to you with measurements of those realized compounds in what, a day, a week, a month, an hour?
这取决于情况。有些化合物很容易合成,比如反应条件很成熟,常规条件就能成功;但有时耦合反应不像文献说的那样顺利,我们需要尝试不同条件。所以要看情况。我这么说,是为了让观众明白,尽管外界有很多关于机器人实验室自动化合成的说法,但现实要复杂得多。
It depends. Like some compounds are really easy to knock out, really known like a very well-characterized reaction where the usual conditions just work, and sometimes it's oh this coupling actually didn't work like the literature said it would and we have to try different conditions. So it depends. I'm saying this so the audience understands that for all the claims out there about just like robotic labs automating synthesis, the reality is just a lot more complicated than that.
抱歉。那里有哪些复杂情况?你彻底……
Sorry. What are some of the complications there? You radically...
它们都做某种形式的自动化。我不是要你跟他们争论,但关于自动化实验室出问题的地方,有什么相反的观点?这也是 Sergey 提到的,我们对强化学习感到兴奋的原因之一,因为它绕过了必须快速进行设计-制造-测试-分析循环并将实验室结果反馈到模型的问题。理想情况下,我们可以投入越来越多的 GPU,让模型自我训练,就像我们在 2010 年代末对棋盘游戏或最近对编程所做的那样。那是我们的关键北极星。然而,在有人告诉你数据室里会有一群天才解决药物发现之前,我们领域有些问题并不天然适合强化学习,需要实验干预。为了具体说明原因,举个例子:编程一直是我人生的北极星。但我在湿实验室里作为平庸的实验科学家度过了相当长的时间,直到研究生院才找到真正的快乐,可以整天编程。现实很混乱。要制造一个新分子,你必须合成它。这意味着不同的化学试剂、催化剂、温度和溶剂以正确的方式和正确的方案组合,试图制造新分子。之后你需要纯化它。如果不纯化,阳性或阴性检测结果可能是假象。你还需要表征它,使用核磁共振和质谱等技术来证明你小瓶里制造的东西正是你最初想要制造的,这并不像听起来那么简单。这不仅仅对我们如此,对任何做 AI 材料科学的人也一样。同样的挑战:我制造的是我实际想要的东西吗?这根本不是一个小问题。在小分子药物发现中比在材料中容易,但仍然不简单且耗时。我们见过很多案例,大分子筛选在论文中听起来很棒,因为你可以一次性测试数百万甚至数十亿个化合物,有更多机会,而且还能为模型提供数据,获得数百万数据点。但现实是,高通量筛选(无论是 DEL DNA 编码库还是传统筛选)的预测结果与实际重新合成分子并进行低通量高保真实验之间的 R 平方低得惊人。假阳性率因各种原因非常高。这就是为什么人们对 AI 进行预测如此兴奋的原因之一,因为理论上 AI 预测可以比许多湿实验(尤其是高通量工作)更干净,产生更高的富集率和更高的真阳性率。
So they all do some form of automation. What like and again not trying to get you to get in some fight with them but like what was the contrary take here on okay what goes wrong with automated lab. Okay. So, this is also one of the reasons as Sergey mentioned, we're really excited about reinforcement learning because it circumvents the problems of having to do very fast design make test analyze cycles from the lab to feed back into the model. Ideally, we can throw more and more GPUs at the problem and have the model self-trained in the same way that we once did for board games in the late 2010s or most recently for coding. Like that, that's a key northstar for us. However, before anyone tells you that there will be a country of geniuses in a data room just solving drug discovery, there are some problems in our space that don't lend themselves to RL as naturally and that will require experimental intervention. And so to drive in the specifics of why to give some examples, coding was always my northstar what I wanted to do with my life. But I spent a fair bit of time myself in the wet lab as a very mediocre experimental scientist before I found my true joy back in you know when I got to graduate school and I could just focus on coding all day. The reality is quite messy. So in order to make a new molecule you have to synthesize it. And what that means is different chemical reagents and catalysts and temperatures and solvents coming together in the right way in the right protocol to try to make your new molecule. And once you do that you need to purify it. If you don't do that your positive or negative assay read could be an artifact. You need to characterize it. So you need to use NMR and other techniques like mass spec to indicate that what you made in your vial is what you actually sought to make in the first place, which is not as trivial as it sounds either. That's not true only for us by the way, for anyone doing material science for AI. It's the same sort of challenge: did I make what I actually made is actually not a trivial question at all. It's easier in small molecular drug discovery than it is in materials but it's still non-trivial and takes time. We've seen so many cases where screenings of large molecules sounds great in paper because you can test millions billions of compounds in one go so you have more shots on goal and as a bonus you have data for your model, right, you just got millions of data points. While okay the reality is that the translation of high throughput screens whether it's DEL DNA encoded libraries or more traditional screens, the R squared of those predictions to like the actual business of resynthesizing a molecule de novo and doing a low throughput high fidelity experiment is a shockingly low R squared. That false positive rate is enormous for a variety of reasons. And so that's one of the many reasons why people are so excited about AI making those predictions because in theory they can be even cleaner than a lot of the wet lab work, especially the high throughput work, and yield to higher enrichment and higher true positive rates.
那么为什么 AI 可以更好?仅仅是因为实验经常出错吗?即使有非常精确的机器人来做,仍然有太多变量,系统中有太多误差吗?
So why can it be more? Is it just because the experiments go wrong a lot and even does it matter that you have some sort of very precise robot doing it? It's still there's too many variables. There's too much slop in the system.
有几个原因。一个是药物发现经常涉及寻找异常值。
It's a few things. So one is that drug discovery frequently involves finding the outliers.
是的。
Yeah.
我们这个领域最疯狂的一点是,分子中追求的特性常常相互矛盾。结合力会随着化合物越疏水、越油腻而增强,因为蛋白质结合口袋是油腻的。但猜怎么着?这会让你的溶解度变差。
One of the many aspects that's so wild about our space is that the properties that one is seeking in a molecule very often anti-correlate. Binding tends to improve the more hydrophobic, the more greasy your compound is because protein binding pockets are greasy. Guess what? That makes your solubility worse.
没错。
Yeah.
然后你通常通过增加极性来改善溶解度。但突然,对于那些还记得高中生物的人来说,细胞被脂质双分子层包围。你的分子可能因为太极性而无法穿过细胞膜。所以你想要的特性常常相互矛盾,这就导致分子的多参数优化感觉像打地鼠一样。你经常在寻找真正的异常值,这通常需要制造根本新颖的分子。而今天可以自动化的化学种类相当有限。因此,你在广阔的化学空间中搜索那些真正顶尖的帕累托最优化合物的能力实际上非常有限。速度上的好处与你能制造的分子实际质量和新颖性之间存在着严峻的权衡。
And then you want to improve your solubility often by adding polarity to your compound. And suddenly, for those that remember high school biology, cells are surrounded by lipid bilayers. Your molecule might not get through the cell membrane because you made your molecule too polar. So the properties you want often anti-correlate and that's where you get where multiparameter optimization of molecules ends up feeling like playing whack-a-mole. So you're often searching for real outliers and that often requires making molecules that are fundamentally novel. And the kinds of chemistries that today can be automated are fairly constrained. So your ability to search chemical space broadly for those really top top Pareto optimal compounds is actually very limited. So the benefits you get in speed have a very harsh trade-off with the actual quality and novelty of the molecules that you can make.
所以你可以很快地做很多无聊的事情。是这样吗……
So you can do a lot of boring stuff very quickly. Is that...
如果你幸运的话?是的,那会是……是的,如果你幸运的话。
If you're lucky? Yeah, that would be... Yeah, if you're lucky.
但要真正回答你那些前沿问题,很难让机器人系统做那些事情。
But to get to the actual answer to your sort of cutting-edge questions, it's hard to get a robotic system to do those kinds of things.
是的,我们很希望那能实现。那会是我们所做事情的巨大推动力,但今天还做不到。
Yes, we would love for that to happen. It'd be a huge tailwind for what we do, but that's not available today.
不是我们今天所处的状态,对吧?我的意思是,我知道再次不要求你透露任何你不能说的,但 Insight 的方法是什么让它们如此快速有效?
Not where we're at today, right? And so I mean, I know again not asking you to disclose anything that you can't, but what is Insight's approach that makes them so fast and effective?
我认为如果一个人愿意专注于一个问题并对分心说“不”,那么奇迹就会发生。我认为其中一点是,他们的人才密度非常高,并且他们愿意做现在大型或中型制药公司并不普遍做的事情,那就是在内部建立如此多的实验能力。
I think if one is willing to focus on one problem and say no to distractions, it's amazing what can happen. And I think one is that their talent density is very very high and they're willing to do what is not universally done now in large or medium-sized pharma, which is build so many experimental capabilities in house.
嗯。
Mhm.
而且我实际上认为,考虑到全球制药开发动态已经发生了巨大变化,更多公司有可能尝试这样做。没错,CRO 接管了一切,然后突然变得商品化,然后你做前沿实验的能力就下降了。关于垂直整合的好处已经争论了几百年。它伴随着所有领域的风险、挑战和前期成本,但拥有端到端控制流程的所有好处。
And I actually think that it's possible more companies to try to do that given the global dynamic of pharmaceutical development that's changed so dramatically. And right, so CRO is taking over everything, then now suddenly has become commoditized, and then your capability to do the cutting-edge experiments goes down. It's been a debate for hundreds of years about the benefits of vertical integration. It comes with all the areas of risk and challenges and upfront cost, but has all the benefits of controlling the process end to end.
当它起作用时,当你端到端地做事时,你可以做出惊人的事情。
When it works, you can do amazing things when you do things end to end.
是的。
Yeah.
当然,它伴随着更大的挑战,我认为需要一种意愿去做。这又引出了我的另一个问题,再次更偏向行业。那么你们,也许现在是时候谈谈你们的转型了,比如从 Genesis Therapeutics 更名为 Genesis Molecular AI,我认为这可能与公司战略的转型相辅相成,或者从我的理解来看,是不是我们要专注于 AI 部分,然后将其出售给制药公司,而不是成为一家也销售工具的制药公司。这大致准确吗?
Of course it comes with greater challenges and I think an appetite needed to do it. And this kind of brings me to another question I have, again more industry. But so you guys, so maybe now is a good time to talk about your pivot from, like, so there's name change from Genesis Therapeutics to Genesis Molecular AI, and that I think goes hand in hand with maybe a pivot in the strategy of the company, or is that in terms of okay we're going to like what it sounds like to me is we're going to focus on the AI portion and then sell that to pharma rather than be a drug company that also sells the tool. Is that sort of accurate?
当我们创办公司时,公司的起源是我们在斯坦福开发的基础深度学习研究方法,对吧?我们试图解决的目标是如何为患者带来最大的影响?当时,公司的创立纯粹是 AI 研究。那时还没有真正的先例,一家纯模型公司仅仅通过为制药或大多数传统行业的合作伙伴创造价值而成功。另一个组成部分是公司建立在 AI 研究之上。我认为我们真的想向人们展示我们对生物技术领域是认真的。所以我认为这是 Genesis Therapeutics 这个名字起源的一部分。我们在制药界那些久经考验的人、根深蒂固的药物化学家、我们真正想合作的药物开发者看来更认真。所以我认为那里有一点试图通过命名决定论的成分。但也承认,我认为当时还没有人真正证明你可以成为一家纯 AI 模型公司并在该领域取得成功。那显然是 2019 年。所以我们有幸成为该领域的真正先驱、开拓者。我们也有过早的劣势。当时以为晚了,但结果发现还是早期。随着公司的发展和演变,我们的结构有点像双螺旋,因为我们拥有极其强大、专注的 AI 研究,并且我们很幸运能够招募到一些最有经验和成就的药物猎手,比如我们的管理团队和董事会成员,他们发现、共同发明、开发了许多 FDA 批准的药物。你与大多数药物化学家交谈,大多数药物化学家如果能在整个职业生涯中参与一个获得 FDA 批准的分子,就会感到非常幸运。而我们真的很幸运能够说服这些有成就的药物猎手从一开始就加入我们的管理团队和董事会,并指导方向。所以结果总是一样的。我们谈论的一切,只有当我们最终帮助到患者及其家人时才有意义。我们从一开始就认为,大规模做到这一点的方式主要是与大型中型制药公司和生物技术公司合作。这样他们可以做他们最擅长的事情,即新颖的生物学、临床开发。我们可以做我们最擅长的事情,即 AI,并将两者结合起来。除此之外,正如 Sergey 所说,拥有内部湿实验室数据生成有许多优势,明确地吃自己的狗粮,从真正的药物化学家那里获得非常直接的人类反馈,以极其坦诚的方式构建我们需要的东西。当然,第三点是宇宙中最有价值的东西,我喜欢开玩笑说是错误之镜测试。如果你问人们你最想要的一件东西是什么,并且你真的思考过,大多数人会说是为自己或家人准备的药物。因此,拥有自己的管线项目具有巨大的内在价值。
When we started the company, the origin of the company was fundamental deep learning research methods that we had developed at Stanford, right? And the objective we were trying to solve for was how can we have the greatest impact for patients? And at the time, the founding of the company was purely an AI research. There was no real precedent at that time for a pure model company really making it in the realm of just creating value for partners in pharma or really most legacy industries. And the other component was the company being founded on AI research. I think we really wanted to show people that we were serious about the biotech domain. So I think that's part of the origin of let's call it Genesis Therapeutics. We seemed a lot more serious to the tried-and-true folks in pharma, the dyed-in-the-wool med chemists, drug developers who we really wanted to work with. So I think there was an element of an attempt at nominative determinism there a bit. But also an acknowledgement that I don't think anyone had really shown at that time that you could be a pure AI model company and really make it in the space. And that was obviously this is 2019. So we had the benefit of being real forebears, pioneers in the space. We had the downside of being pretty early. Thought to be late, but turns out it was still early innings. As the company grew and evolved, we were really structured kind of like a double helix in that we had extremely strong, focused AI research and we were fortunate to be able to recruit some of the most experienced and accomplished drug hunters, like people on our management team and board who have discovered, co-invented, developed many FDA approved drugs. You talk to most med chemists, most med chemists feel very fortunate if they're able to work on one molecule that gets FDA approval in their whole career. And we're really fortunate to convince these sort of accomplished drug hunters to come join our management team and board at the beginning and steer the show. So the outcome has always been the same. Everything we're talking about, it only matters if we end up helping patients and their families at the end of the day. And we have felt from the beginning that the way to do that at scale is to primarily partner with large medium-sized pharma companies and biotech companies. So they can do what they're best at, which is novel biology, clinical development. We can do what we're best at, which is AI, and bring those two together. In addition to that, as Sergey has said, there's numerous advantages to having in-house wet lab data generation, clearly dogfooding the models, getting really direct human feedback from real med chemists, what we need to build in an extremely candid way. And of course, the third thing is that the single most valuable thing in the universe, what I like to joke is the mirror of error said test. If you ask people what is the one thing you most want and you really thought about it, most people would say a medicine for themselves or a family member. And so there is just immense intrinsic value to having one's own pipeline programs.
当我们与制药合作伙伴互动时,他们普遍喜欢我们自己也在做分子研发,因为这表明我们不是一群只会敲键盘的人,而是亲身实践我们推销的愿景,并且我们自己也需要使用这些模型。因此,这些模型在额外的方式下经过了实战检验,我们也被“驯化”了,深知这个领域的真正难点。我们不是在兜售空中楼阁。
And when we interact with pharmaceutical partners, on average, they love that we're working on molecules ourselves because it means that we're not just a bunch of keyboard jockeys, but we live the dream that we are pitching and we need to use these models and practice for ourselves, too. So, they're battle tested in an additional way and we're sort of domesticated in that we know what's really difficult about the space. We're not just selling some pipe dream.
在我看来,过去一两年发生了一个转变:制药公司突然对购买现成工具产生了兴趣,而之前很多公司尝试构建 AI 但遇到困难,于是它们变成了制药公司,因为那是它们唯一能赚钱的方式,对吧?你筹集一大笔钱,希望十年后能有一种药物,如果成功就大获全胜,否则就倒闭或出售资产。现在似乎正在改变。最近在药物发现、病理学等许多相关领域出现了大量 AI 收购。你看到这个转变了吗?这是否是你们转向销售模型的部分原因,还是仅仅因为那个转变而重新品牌?到底发生了什么?
It seems to me like there's been a shift in the last, I don't know, year to two years where suddenly pharma has become interested in buying tools off the shelf whereas previously there were a lot of companies that started trying to build AI and then had trouble and so they became pharma companies because that's the only way they could make money, right? You raise a bunch of money and hope that in 10 years you have a drug and if that drug is successful then you've made it and otherwise then you fold or sell off the assets or whatever. And it seems like now that's changing. You have a lot of recent AI acquisitions in drug discovery and pathology and lots of different related areas. So have you seen this shift and is that part of the reason you are shifting towards selling models or is this just rebranding because of that shift? What's going on?
为了完成那个叙事弧线,更名只是为了反映我们实际在做的事情:我们是一家 AI 公司,从第一天起就是如此,当时我们还在用 PyTorch 编码。当 Sergey 在 PyTorch 上扩展 Transformer,第一个做到这一点时,我们也在 PyTorch 上扩展图神经网络,也是最早这样做的之一。所以这反映了我们从第一天起就有的身份。同时也暗示了我们自己是非常认真的分子制造者。我们在那个意义上是全栈的。另外,明确一点,我们的使命是创造尽可能多的、否则不会被创造的药物,我认为实现这一目标的方法是将我们的技术直接交到尽可能多的药物开发者手中,不仅是我们自己,还包括大型制药、中型制药和生物技术公司。这也是我们一直在构建智能体的部分原因,这样我们那些变得更加强大的模型,可以被那些可能一生从未写过一行代码的药物化学家和 CAD 科学家轻松使用和编排。
To finish that narrative arc, the name change was just to reflect what we were really doing in practice, which is we are an AI company as we have been from day one when we were coding in PyTorch. Building the first, when Sergey was scaling transformers on PyTorch, the first one to do that, we were scaling graph neural nets in PyTorch, one of the first ones to do that. So it's reflecting what our identity has been from day one. But also alluding to the fact that we are very serious molecule makers ourselves. And we are full stack in that way. We are also, to be clear, our mission is to create as many medicines as possible that would otherwise not have been created, and I think the way to get there is to put our technology directly into the hands of as many drug developers as possible, not just our own but in big pharma, midcap pharma, and biotech. And that's partially why we've been building our agents so that our models, which have become so much more powerful, can be easily used and orchestrated in the hands of med chemists and CAD scientists who might have never ever written a line of code in their life.
我很有兴趣听到你们现在开始将智能体引入开发流程。听起来从一开始,你们的理念就更多是:这些是给科学家用的工具,你们制造最好的工具,让他们优化任何他们使用的流程,但最终这是给药物化学家的工具。而现在你们在转变,你们说我们实际上将拥有自动化的系统,它们做出决策并推进。首先,你确实暗示你们可能已经达到了某个阈值,比如去年 11 月的那个阈值,它突然变得神奇地有用。所以也许这是部分原因,但还有其他原因促使你们朝这个方向走吗?为什么是现在?
I find it interesting to hear that you have now started bringing in agents into your development process. It sounds like from the very get-go your philosophy was much more these are tools for scientists and you make the best possible tool and this lets them optimize whatever pipeline they're going to use, but ultimately this is a tool for med chem. And now you are switching and you were saying we are actually going to have automated systems which make decisions and then advance. First of all, you did allude to you've kind of hit maybe the threshold, like the last November threshold where it just magically became useful. So maybe that's part of it, but are there other reasons for going in this direction and why now?
好吧,让我反过来问你一个问题。你认为一个人类药物化学家或 CAD 科学家实际上能学会并精通多少种工具?
Well, let me ask you a question in return. How many tools do you think a human med chemist or CAD scientist can realistically learn and be proficient in?
哦,我会说他们能学会多少以及实际使用多少?因为我看到他们喜欢用很多工具,但可能并不一定擅长。
Oh, I would say how many can they learn and how many do they use? Because I've seen that they like to use lots of them, but maybe they're not necessarily good with them.
但这真的很难,对吧?很多工具极其复杂。
But it's really hard, right? A lot of the tools are extremely complex.
它们有很多参数需要正确设置和配置,对吧?
They have a lot of parameters that you need to set properly and configure them, right?
而这正是智能体极其擅长的地方。它们实际上知道如何使用所有工具以及如何编排它们。
And this is where agents are extremely good at. They actually know how to use all of the tools and how to orchestrate them. Well,
所以我不想打断你,当你说智能体时,有些人用这个词仅仅指自动化的东西,但另一些人用智能体特指有一个 LLM 负责编排决策。我的意思是,在生物技术之外,几乎普遍是后者,但有时在生物领域,人们仍然用智能体指代任何基于决策的、自动化脱离人类的系统,也许其中一些甚至更像是传统的基于规则的系统,对吧?
So I wouldn't interrupt you when you say agent that some people like using the term agent to just mean something automated, but other people use agent to specifically mean you have an LLM which is responsible for orchestrating decisions. I mean outside of biotech it's almost universally the latter, but sometimes in the space of bio people are still using agents to refer to any sort of decision-based system which is automated out of humans and maybe some of them are even more like traditional rules-based systems or something, right?
所以在我们的案例中,我指的是 AI 对智能体的定义:一个 LLM,能够使用我们多年来辛苦构建的所有工具。这是一个非常强大的概念,因为现在作为 CAD 科学家或药物化学家,你不需要深入了解所有细节和需要设置的所有超参数,你可以让这个东西去运行、找出并解决问题。这就是我们之前讨论的所有问题变得非常重要的地方:智能体需要能够理解这些模型产生的内容。为什么我们关心晶体结构?因为智能体实际上可以查看你预测的晶体结构并基于此做出决策。我们认为这对迭代和改进超级重要。
So in our case, I mean the AI definition of agents: an LLM which is able to use all of the tools that we have been painstakingly building over the many years of existence of this company. And that's a very powerful concept because now as a CAD scientist or med chemist, you don't need to deeply understand all of the details and all of the hyperparameters that you need to set and you can just let this thing go and figure out and solve the problems. And this is where all of the previous questions that we discussed come back as very important: agents need to be able to understand what these models are producing. Why do we care about crystal structure? Well, the agent can actually look at your predicted crystal structure and make decisions based on that. Something that we believe is super important to iterate and improve.
你是说你可以拿一个晶体结构——我猜在这种情况下它被框架化为一个图或三维图像,Molestar 以某种方式解释它——然后它实际上可以就什么与什么结合做出决策并提出假设。你们已经达到了看到它做出有效决策的地步了吗?
You're saying that you can take a crystal structure which I guess in this case is framed as a graph or like a three-dimensional image that Molestar is interpreting in some way and then it actually can make decisions upon what is binding to what and make hypotheses. You've gotten to that point where you have seen it making effective decisions?
对。所以这个问题肯定有多个层面。你可以从基本的图像理解开始,或者使用额外的工具来理解晶体结构中发生了什么,甚至你可以训练一个能够原生理解晶体结构表示的模型。就像图像模态是 LLM 的一种模态一样。如今,晶体结构也可以成为 LLM 的一种模态。
Right. So there are definitely multiple layers to this question. You can start with like basic image understanding right here, or you can use additional tools to understand what is happening in the crystal structure, or you can even train a model which is able to natively understand crystal structure representation. Like image modality is a modality for LLMs. These days, crystal structure can be a modality for LLM as well.
我可以在那些以某种方式捕获 3D 晶体结构 token 的序列上进行微调,对吧?
I can fine-tune it on sequences that look like that are somehow capturing tokens that are 3D crystal structures, right?
是的。
Yeah.
那么,如果历史上你们把人类视为你们为其制造工具的对象,现在你们有了智能体,你们如何重新构想人类与智能体的互动以及决策过程?这是希望本质上将人类自动化掉吗?你们想扩大规模并做更多假设吗?
So, if you historically have treated humans like you were making tools for humans, so now you have agents, how do you reimagine the interaction of humans and agents and the decision-making process? Is this you hope to essentially automate humans out? You want to scale up and do more hypotheses?
是的。
Yeah.
此外,怀疑论者还会问:你怎么确保你的智能体不会做出奇怪的决定,比如调错了超参数,然后给你一堆……
And then also the skeptic would ask: how do you make sure that your agents are not making weird decisions where they tweaked the wrong hyperparameter and now they give you a bunch of...
对。你并不足够了解那些参数,所以不知道发生了这种情况。
Yeah. You don't understand the parameters well enough to know that that happened.
这是个好问题。我认为答案可以从编码工具的演变中看到。还记得 Cursor 刚出现的时候吗?它只是顶级的自动补全工具。它当时就有用,但不像现在这样,我打开 Cursor,它就能帮我搞定一切。所以我认为我们会看到非常相似的轨迹。那些智能体已经很有用了,但人类需要提供持续的反馈,并在过程中引导和指导智能体。随着时间的推移,它们会变得更加独立。不过,我不相信完全自动化会取代人类。我相信人类会利用这些工具在工作中变得高效得多。人类将为智能体需要做的事情提供战略方向,而智能体将代表人类去执行。
It's a great question. I think the answer can be seen in how coding tools have evolved. You remember Cursor when it first appeared? It was just top autocomplete. It was still useful, but not to the extent it is now, where I can fire up Cursor and it just goes and does everything for me. So I think we'll see a very similar trajectory here. Those agents were already useful, but humans need to provide consistent feedback and basically steer and guide the agent along the way. Over time, they'll become more independent. Still, I don't believe in full automation replacing humans. I believe humans will become way more efficient in their jobs using these tools. A human will provide strategic direction for what the agent needs to do, and the agent will execute on behalf of the human.
我认为这是有史以来成为有创造力的人、或者说任何人的最好时代。历史的漫长弧线,从我们决定不再狩猎采集……我不确定。我想我要安定下来,搞搞农业。我认为那开启了数千年的苦差事。有时你会想我们为什么要搞农业,因为我看到的一切似乎表明农业在短期内可能是净亏损。
I think this is the best time ever to be a creative human being, or really any human. The long arc of history, starting from when we decided that hunting and gathering... I don't know about that. I think I'm going to settle down and do this whole agriculture thing. I think that started thousands of years of drudgery. Sometimes you wonder why we did agriculture, because everything I've seen seems like agriculture was probably a net loss in the short term.
从长远来看,我看到的唯一好处是,如果你是一个狩猎采集者,生病了、天生有遗传病或受伤了,你就麻烦了。我们帮不了你多少。医学根本不存在。但最终,我认为文明的最高成就是,社会中的病人不再被视为负担,而是我们想要帮助的人。这在动物界(包括人类)的历史上并非如此。这是一个可悲的现实,但却是事实。现在,因为医学,我们把病人视为我们想要帮助、能够帮助的人,他们甚至可以成为更有生产力的社会成员,如果你在乎这个的话。快进到现在,我认为很多创造性工作和日常工作,如果你把它看作一个饼图,时间是我们最宝贵的资源。其中很大一部分被不需要太多思考或创造力的重复性任务占据。现在我们创造了额外的深度思考空间。对我们来说,解决一些技术上最具挑战性但重要的问题,我们可以利用这些工具变得极其高效和富有创造力。这适用于我们的直接用户:药物化学家、药物发现者、CAD 科学家。他们可以成为药物发现活动的总战略家。当你有数百名药物发现科学家在大型数据中心全天候工作时,发现方面的产出只会更大,尤其是在由懂行的专家科学家使用时。
The only benefit I see in the long run is that if you were a hunter-gatherer and you got sick, or were born with a genetic illness, or got injured, you were in trouble. There wasn't much we could do for you. Medicine just didn't really exist. But ultimately, I think the crowning achievement of civilization is that the sick people among us are not seen as burdens, but as people we want to help. That hasn't been true in the history of the animal kingdom, including humans. It's a sad reality, but it's true. Now, because of medicine, we see sick people as people we want to help, who we can help, and they can become even more productive members of society if that's what you care about. Fast forward to now, I think so much of creative work and day-to-day work, if you look at it as a pie chart, time is our most precious resource. So much of it was taken up by repetitive tasks that didn't require much thought or creativity. Now we've created additional headspace for deep thinking. For us, solving some of the most technically challenging but important problems, we can use these tools to become wildly more productive and creative. That goes to the med chemists, drug hunters, CAD scientists who are our direct users. They can be grand strategists for a drug discovery campaign. When you have hundreds of drug discovery scientists working 24/7 on large data centers, the output in terms of discovery is just going to be greater, especially if it's used by expert scientists who know what they're doing.
我们可能早该聊到这个了,但你的 Pearl 模型基于最近一个你之前提到的开放数据集挑战,取得了一些非常令人兴奋的结果。抱歉我们之前没聊到,但我想给你一个机会谈谈这个。有一个开放绑定,是 EV A721A 蛋白酶吗?这是一个非常难的目标,有很多因素使得社区或共折叠方法难以内部运作。你们内部运行了这个,得到了一些有趣的结果。
We probably should have talked about this earlier, but you have some really exciting results from your Pearl model based on a recent open dataset challenge you alluded to earlier. Sorry we didn't talk about it before, but I'd like to give you a chance to talk about this. There's this open bind, was it EV A721A protease? It was a very hard target, had a lot of things which made it hard for the community or for co-olding methods to work internally. You ran this internally and had some fun results.
是的,我们也可以回到关于评估以及评估普遍存在问题的对话。如果你看看开源模型及其在公共基准上的表现,基本上都差不多。当 OpenBind 提出这个新基准时,令人惊讶的是,模型在该基准上的表现差异很大。部分原因是因为它只是一个单一目标,我理解。但也因为这个目标是那些模型在训练中从未见过的。所以以前没有人针对那个特定目标进行过优化,没有人爬过那个坑。我们非常好奇我们的模型在这个我们训练或开发模型时也从未见过的特定目标上,开箱即用表现如何。所以当它出现时,我们显然想研究一下。我们在上面运行了 Pearl 系统,结果让我们既惊讶又兴奋。我们的数字远高于其他公开可用的模型。对我来说,这反映了 Pearl 在实际药物发现项目中的表现,因为我们之前在内部项目或合作项目中见过这种数字差异,但我们无法谈论它们,因为很多数据要么是合作数据,要么是内部数据,我们不能披露。这里有一个完美的例子,它是一个外部目标,我们事先不知道,只是运行了一下。我们看到了这种惊人的差异。这就是让我兴奋的地方。这个目标有一些特点,比如它有一个灵活的环,当你将配体插入正确位置时需要移动。很多方法都无法处理。Pearl 在找出如何移动这个环方面异常出色,而且我们基本上每个姿态都是正确的。我们很好地预测了这种移动。所以这是令人兴奋的部分。
Yeah, we can also go back to our conversation on evals and what's wrong with evals in general. If you look at open source models and their performance on public benchmarks, it's kind of very similar across the board. What was surprising when OpenBind came up with this new benchmark is that you see a wide variety of performance of the models on that. Part of it is because it's just a single target, I get it. But it's also because this is a target that those models haven't seen during training. So nobody has optimized for that specific target before, nobody has climbed on it. It was very interesting for us to see how our model would perform out of the box on this particular target that we also have never seen before during training or when we developed the model. So when it came up, we obviously wanted to look into it. We ran the Pearl system on it, and we were genuinely surprised and excited about the results. Our numbers are way higher than other published openly available models. To me, this reflects how Pearl actually behaves in real drug discovery programs, because we have seen this kind of difference in numbers in internal programs or partnership programs before, but we were unable to talk about them because a lot of this data is either partnership data or internal data we cannot disclose. Here you have a perfect example where it's an external target, we didn't know, and we just ran on it. We see this staggering difference. That's what makes me excited. There are some specifics of this target, like it has a flexible loop that needs to move when you insert your ligand into the right location. A lot of methods are just unable to handle it. Pearl was exceptionally good at figuring out how to move the loop, and we are basically correct for every single pose. We are predicting this movement really well. So that's the exciting part.
我对此很好奇。Pearl 有什么特点让它能够模拟这样一个动态目标?
I'm curious about this. What is it about Pearl that allows it to model a dynamic target like that?
Pearl 的训练方式,是训练它同时预测蛋白质结构和配体。蛋白质结构和配体一起,不一定与孤立蛋白质的形状或形式相同。所以整个训练过程将 Pearl 推向了那个方向。
Pearl, the way it's trained, it's trained to predict together protein structure and ligand. Protein structure and ligand together are not necessarily in the same shape or form as the protein in isolation. So the whole training process kind of nudges Pearl in that direction.
那么训练集中有很多这样的例子吗,具有这种动态特性,即在结合过程中构象发生变化?整个训练集中都有诱导契合吗?
So are there lots of examples in the training set like that, where there's this dynamic nature, where once during binding the conformation changes? There is induced fit throughout the training set?
是的。
Yeah.
这可能是从你们使用的基于物理的模拟中继承而来的。
This is probably inherited from the physics-based simulations that you're using.
这当然有帮助。我认为 Sergey 提到的另一个因素是,我们不仅仅是在同一个基准上最大化。我们显然会发布和运行姿态。
It certainly helps. I think the other component that Sergey alluded to was that we haven't just been benchmark maxing on the same. We obviously publish and run poses.
去年年底,我们在 GTC 上与 Nvidia 联合发布了 Pearl 技术报告,展示了在 runs 和 poses 上的最先进结果。我们使用 runs、poses 和 pose busters,Pearl 在这些基准测试中表现最佳。我们做到了。
We published the state-of-the-art numbers on runs and poses last year when we published the Pearl technical report with Nvidia at GTC at the end of last year. We use runs and poses, we use pose busters. Pearl does the best on those benchmarks. We did that.
不过,正如 Siri 所暗示的,在我们合作的药物靶点上,Pearl 与次优模型之间的差距甚至更大。这在一定程度上反映了我们公司的目标不是最大化基准测试成绩,而是为那些通常研究 PDB 中不存在的困难药物靶点的生物技术和制药公司创造价值。因此,我们必须构建能够外推的模型。开放购买挑战就是最新的例子。
However, as Siri is alluding to, the gap from Pearl to the next best model is even larger on the partner drug targets that we work on. That in part reflects that the objective of our company is not maxing benchmarks. The objective is creating value for biotechs and pharma companies that usually work on hard drug targets different from what's in the PDB. So we've had to build models that are able to extrapolate. The open buy challenge was just the latest example of that.
在结束之前,我们总是会问两个问题。第一个是,如果你能通过行政命令消除行业中的一个瓶颈,那会是什么?
Before we wrap up, we always ask two questions. One is if you could remove a bottleneck from your industry by fiat, what would that be?
我可以说,我目前的主要瓶颈是 GPU。
I can say my main bottleneck right now is GPUs.
好吧。每个人都这么说。
Okay. Everyone says that.
是的,你看,GPU 价格不断上涨,LLM 公司正在吸走所有的 GPU 算力。我坚信我们所做的事情对人类真正重要——为自己、为所爱的人发现新药,是我们每个人一生中某个时刻都需要的东西。所以我认为为药物发现解除 GPU 瓶颈非常重要。
Yeah, look, GPU prices are going up and up, and LLM companies are sucking up all the GPU capacity out there. I firmly believe that what we're doing is genuinely important for humanity—discovering new medicines for yourself, for your loved ones is something we all need at some point in our life. So I think it's very important to unblock that GPU bottleneck for drug discovery.
Anthropic 实际上是否因为 GPU 采购而阻碍了科学进步?你是不是要搞事情?我不知道能不能问这个,但你们在考虑非 Nvidia 的 GPU 吗?我是说,我猜你们是 Nvidia 的合作伙伴,随便吧。但我看到一些公司正在转向多架构模型,以便利用其他供应商来应对短缺。
Is Anthropic actually throttling science because of GPU purchases? Are you getting spicy? I don't know if you can ask this, but are you looking at non-Nvidia? I mean, I guess you guys are Nvidia partners, whatever. But I've seen that some companies are pivoting towards a multi-architecture model to take advantage of other suppliers because of the shortage.
Nvidia 一直是我们的大力支持者。他们两次投资 Genesis,我们密切合作——不仅在资本方面,我们还共同优化内核。Pearl 是由 Genesis 和 Nvidia 共同撰写的。他们与我们一起公开了开放思维的结果。我们对此非常感激。他们是一家非常非常大的公司,而我们还没有达到像 Meta 那样的客户规模。尽管如此,我们仍然感谢他们给予我们时间和关注。但我认为这在一定程度上是出于自身利益:显然没有人能把握市场时机,但传统 LLM 领域(聊天机器人、编码等)的炒作和投资,在某个时候,其阿尔法收益将无法与投资相匹配。我认为,无论经济如何变迁,药物总是有高需求。无论是出于帮助人类的愿望,还是寻找下一个大事件的自身利益,我确实认为包括 Nvidia 在内的芯片制造商将希望更多地投资于生命科学,因为需求始终很高,而纯 LLM 领域的阿尔法收益正变得有些可疑。我的意思是,所有 LLM 公司现在都开始涉足生命科学,对吧?GTC Roslin,有云计划,还有很多初创公司等等。
Nvidia has been a great supporter of us. They've invested twice in Genesis, and we collaborate closely—not just in terms of capital, but we optimize kernels together. Pearl was co-authored by Genesis and Nvidia. They publicized the open mind results with us. We've been very grateful for that. They're a very, very big company, and we're not at the scale of a customer like Meta or something. So we're grateful that they've given us time and attention nonetheless. But I think part of that is a self-serving aspect: obviously no one can time the market, but the amount of hype and investment in the traditional LLM space for chatbots, coding, etc., at some point the amount of alpha will not be what's desired compared to the investment. I think regardless of the vicissitudes of the economy, medicines will always be high in demand. Whether it's a desire to help humanity or a self-serving interest in looking for the next big thing, I do think that chipmakers including Nvidia are going to want to get a lot more invested in life sciences because it will always be high in demand, and the amount of alpha left in pure LLM space is just getting a little questionable. I mean, all the LLM companies are starting to look into life sciences right now anyway, right? GTC Roslin, there's a cloud initiative, there's lots of startups and whatever.
是的。我们认为这些都是对我们有利的顺风。
Yep. We think those are all great tailwinds for us.
那你呢?除了 GPU,还有什么瓶颈?
And then what about you? What's the bottleneck besides GPUs?
哦,我当然和 Sergey 的回答一样。如果我能挥动魔杖,我希望拥有一个巨大的 H100 集群。那将太棒了。
Oh, I had the same answer as Sergey for sure. If I could wave a magic wand, I'd love to have just an enormous H100 cluster. That would be amazing.
然后第二个问题是,你对我们的观众有什么行动号召吗?也就是说,AI 工程师和科学家们,参与进来,申请工作,你的行动号召是什么?
And then the second question we ask is just do you have a call to action for our audience? So meaning AI engineers plus scientists, get involved with something, come apply for jobs, like what's the call to action for them?
我们确实在招聘。我们正在寻找这个领域的顶尖人才。我个人从 LLM 领域转到了药物发现。在 Genesis 之前,我从未做过药物发现。我发现这个领域非常迷人且极其有趣。在 LLM 公司,你经常看到研究科学家对研究架构感到兴奋,但却被迫做无聊的事情。老实说,LLM 架构相对无聊。我不知道,我可能疏远了你一半的听众。
We are definitely hiring. We're looking for top talent in the space. I personally transitioned from the LLM space to drug discovery. I had never done drug discovery before Genesis. I find this field fascinating and extremely interesting. In LLM companies, you often find research scientists excited about working on architectures but kind of forced to work on boring stuff. Honestly, LLM architectures are relatively boring. I don't know, I probably alienate half of your audience.
我想他们可能也这么想。但说到底,它就是一个 Transformer 层——论文发表于 2017 年,现在你去任何 LLM 实验室,都会看到他们的模型中非常相似的组件。
I think they probably think the same thing. But it's like it's a transformer layer in the end—the paper was published in 2017, and you go to any LLM lab today, you will see very similar pieces in their models.
是的,有一点 MoE 的风格,但除此之外,架构基本上与 2017 年发现的非常相似。对吧?嗯,我们的架构和模型实际上非常不同,并且非常有趣。所以如果有人对研究架构感到兴奋,这个领域实际上是一个非常有趣的工作场所。
Yes, there is a little bit of flavor of MoEs, but otherwise architectures are fundamentally very similar to what was discovered back in 2017. Right? Well, our architectures and our models are actually very different and very interesting to work with. So if somebody is excited about working on architectures, this space is actually a very interesting place to work.
是的,我想评论一下。不仅你们的架构与 LLM 领域非常不同,而且你们也与许多其他生物 ML 领域正在做的事情非常不同。它在子领域和更广泛的范围内都是独特的。
Yeah, I'd like to comment. Not only are your architectures very different from the LLM space, but you're also very different from what a lot of the other bio ML space is doing as well. It's sort of unique both in its subdomain and more broadly.
我们看到,最近的一位嘉宾是 ESM fold 的创始人之一。是的。他们引用了我们几个月前发布并开源的模型架构论文。因此,作为 Genesis 的 AI 研究员,你有很多机会发表论文、开源和影响该领域,除了通过与真正制造改变人们生活的药物的顶级客户合作来影响该领域的明显方式之外。所以这既非常具有智力吸引力,也有很多机会影响该领域并让你出名。但如果你愿意,你也可以继续做 LLM 公司机器中的一颗螺丝钉。没有冒犯,没有冒犯。
We saw as a recent guest you had one of the progenitors of ESM fold. Yeah. They cite our model architecture paper that we published and open-sourced a few months ago. So the work that you do here at Genesis as an AI researcher, there are lots of opportunities for publication, open source, and influencing the field beyond the obvious ways to influence the field by working with top customers that are actually making medicines that change people's lives. So it's both incredibly intellectually engaging and a lot of opportunities to influence the field and get your name out there. But if you want, you could also just stay being a cog in the machine of an LLM company. And no shade, no shade.
好了。非常感谢你们来做客并和我们交谈。我个人非常享受这次对话。希望观众也是。Evan,Sergey,谢谢你们。我们期待看到事情的发展。希望你们能从中收到一些好的工作申请。
All right. Well, thank you so much for coming and speaking with us. I've personally really enjoyed this. Hopefully the audience as well. Evan, Sergey, thank you. And we look forward to seeing how things develop. Hopefully you get some good job applications out of this.
是的,我们喜欢这个节目,所以非常感谢你们抽出时间聊天。
Yeah, we love the show, so we really appreciate you taking the time to chat.
是的,谢谢你们邀请我们。
Yeah, thank you for having us.