AI Scientist: The Future of Drug Discovery
打开互动全文版(中英对照 + 朗读 + 问答)→Sam Rodriques 讨论 AI 智能体如何扩展科学人才、治愈疾病并革新药物发现。
Sam Rodriques discusses how AI agents can scale scientific talent, cure diseases, and revolutionize drug discovery.
那么,你在神经科学和生物工程领域有着极其光明且成功的职业生涯。我记得你最初学的是理论物理,后来转向了 AI。是哪里出了问题?你是怎么想的?为什么要这么做,Sam?
So, you had this incredibly promising successful career in neuroscience and bio-engineering. I think you studied theoretical physics originally and then you moved into AI. What went wrong? What were you thinking? Why did you do it, Sam?
好问题。我最初学理论物理,做量子信息理论。物理学的问题在于,实际上已经没有未解决的难题了——当然这有点夸张。我们需要建造量子计算机,有一些有趣的材料科学问题,我们仍然不知道宇宙如何运作。但如果你看看今天这个房间里你能看到的任何现象,我大概都能解释到亚原子粒子层面,除了你,或者那只老鼠,或者你为什么在某个时间醒来,或者你怎么生病的。所以,就我们所体验的世界而言,除了生物学,大多数事物都已被很好地理解。这就是我最初进入生物学的原因。我在 MIT 读博。我骨子里是个发明家,发明了不少技术。但在 MIT 我学到的最关键一点是:做科学需要三样东西——资本、物流和人才。资本可以规模化,物流可以规模化,但人才就是无法规模化。在生物学中,我们受限于人才。所以我思考如何消除人才这一科学瓶颈。如果我们想治愈所有疾病、理解大脑运作、解决衰老问题,我们就得想办法规模化人才,而 AI 似乎是实现这一目标的正确途径。于是我放弃了已建立的职业生涯,跳下了悬崖。到目前为止,一切进展顺利。
Yeah, great question. So I started doing theoretical physics, did quantum information theory. The problem with physics is there are actually no unsolved problems left in physics. That's a slight exaggeration. We need to build quantum computers, there are some interesting material science problems, and we still don't know how the universe works. But if you look around at any phenomenon you can see in this room today, I could probably explain how that phenomenon works down to subatomic particles, except for you, or the mouse, or why you wake up when you do, or how you get sick. So when it comes to the world as we experience it, most things are pretty well understood except for biology. That was originally what got me into biology. I did my PhD at MIT. I'm an inventor at heart. I invented a bunch of different technologies. But the key thing I learned at MIT is that to do science, you need three things: capital, logistics, and talent. Capital scales, logistics scales, but talent just does not scale. In biology, we're limited by talent. So I thought about how to remove talent as a bottleneck in science. If we want to cure all diseases, understand how the brain works, and solve aging, we need to figure out how to scale talent, and AI seemed like the right way to do that. So I built up this career and then just threw it away and jumped off a cliff. It's been working out well so far.
那么你创办的这家公司——你最好解释一下它是做什么的——最初是一种新型的非营利组织,对吧?
And originally the company that you started now, which you better explain what that does. I think it started as a new kind of not-for-profit, right?
是的。那是 2022 年。我意识到未来十年科学领域最重要的事情就是构建一个 AI 科学家,因为它能解除人才瓶颈。但那时在生物学中,事情进展非常缓慢。GPT-3 已经发布,它能做一些事情,但看起来要花很长时间才能达到目标。我不知道如何将其商业化。所以我们认为应该以非营利的形式来做,直接去做基础研究。我与联合创始人 Andrew White 合作,他是 AI 智能体用于科学领域的先驱。他当时正在与 OpenAI 合作开发 GPT-4,也是罗切斯特大学的教授。我们对构建 AI 科学家有着相同的愿景。我了解生物学方面,Andrew 知道如何从技术上实现。我们原以为需要五到十年,这完全是胡说八道。我们在 2023 年 11 月推出了 Future House,两年内我们就已经拥有了极其强大的 AI 科学家。所以时间问题并不重要。我认为我们以非营利形式起步并不是唯一的原因。
Yeah. So this was 2022. I had figured out that the most important thing for science in the next 10 years was going to be building an AI scientist, precisely because it would unblock talent. But at that time, in biology, things go really slowly. GPT-3 was out and it could kind of do things, but it seemed like it would take a really long time to get there. I didn't know how to commercialize it. So we thought we should do this as a nonprofit, just go and do the basic research. I teamed up with Andrew White, my co-founder, who is a pioneer in AI agents for science. He was working with OpenAI on GPT-4 at the time and was a professor at the University of Rochester. We had the same vision for building an AI scientist. I knew the biology side, and Andrew knew how to do it technically. We thought it would take five or ten years, which was complete nonsense. We launched Future House in November 2023, and within two years we had extremely powerful AI scientists already. So the timing thing wasn't important. I think the fact that we started as a nonprofit was not the only reason we did it that way.
我们在 2022 年起步时,还没有任何关于推理模型的概念。那是第一个能够完成科学发现全流程的多智能体系统。它提出了一种治疗一种名为年龄相关性黄斑变性的失明的新假说。实际上,这个成果两天前刚刚发表在《自然》杂志上。自从我们推出 Cosmos 以来,人们可能已经用它做出了 2 万到 3 万个新颖的科学发现,这太疯狂了。
When we got started in 2022, we didn't have any notion of reasoning models. It was the first multi-agent system that was capable of doing the full loop of scientific discovery. It came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration. It actually just got published in Nature 2 days ago. Since we launched Cosmos, people have probably used it to make 20 or 30 thousand novel scientific findings, which is wild.
未来的制药公司将更加精干。你将能够用同样数量的人并行推进比今天多得多的药物项目。
The pharma companies of the future will be much more lean. You're going to be able to pursue many more drug programs in parallel than you can today with the same number of people.
AI 在药物发现领域有过很多过度承诺的历史。
There's a long history of AI overpromising in drug discovery.
哦,是的。
Oh yeah.
这次有什么不同吗?
Is this time different?
我向你保证,这次不一样。
I promise you this time it is different.
我们一直非常希望,并且现在仍然希望,这些技术能够惠及整个科学界。所以未来我们开源了很多东西,我认为这也非常重要。但后来发生的是,到了 2025 年春天,有一天 Andre Karpathy 发了条推文之类的,然后突然间全世界都知道了 AI 智能体是什么。我们开始接到电话。我们真的从需要在演示文稿里解释什么是智能体、它和语言模型有什么区别,一下子变成了制药公司的高管打电话问我们:“天哪,我们怎么用你们的智能体?怎么用你们的东西?”所以很快就变得很明显,我们必须成立一个营利性衍生公司来满足这种需求。
We also really wanted, and continue to want, those technologies to benefit the entire scientific community. So in the future, we open-sourced a lot of stuff, which I think is also super important. But then what ended up happening was, by spring of 2025, one day Andre Karpathy tweeted or something, and all of a sudden the entire world knew what AI agents were. We started getting phone calls. We literally went from having to explain in our decks what an agent is and how it's different from a language model, to having senior executives at pharma companies call us saying, "Oh my god, how do we use your agents? How do we use your thing?" So it became evident pretty quickly that we were going to have to have a for-profit spinout to satisfy that demand.
那你觉得非营利组织在技术领域还有作用吗?我的意思是,我们现在有两个例子,都是非营利组织开发了有趣的技术,然后基本上都变成了营利性公司,才把技术推向世界。非营利模式是不是有什么问题?
Do you feel like there's a role for nonprofits in technology then? I mean, we now have two examples of nonprofits developing interesting technology and kind of mainly turning it into for-profits to get that technology out into the world. Is there something broken about the nonprofit model?
不。首先,非营利组织在技术开发中有作用吗?有,绝对有。非营利组织有什么问题吗?没有,绝对没有。这只不过是营利性公司在做它们擅长的事情,也就是规模化。所以首先,让我们回顾一下。我读博的时候,在做连接组学项目。这个项目基本上是要弄清楚如何绘制大脑中所有神经元之间的连接。理解大脑工作原理的一个核心部分是弄清楚接线图。直到今天,我们还没有人类大脑的接线图,也没有小鼠大脑的接线图。我们大概有果蝇的接线图,那是我们最好的成果。没有接线图,很难理解大脑的工作原理。所以我们想去绘制大脑的接线图。这非常困难,基本上是因为它涉及追踪这些微小的纤维——微小到只有人类头发宽度的 1/100——穿过大脑这个巨大的缠结。你不能出错,因为一旦出错,你就会连接实际上并不相连的神经元。这就像试图绘制一片草地上所有的根,但要难得多。这就是我感兴趣的问题。我试着在学术实验室里做,但我意识到我根本无法聚集所需的工程资源、资金和人才。但另一方面,你也不能为了营利去做。在营利性环境中,你无法获得资金来绘制大脑图谱,因为你拿图谱做什么?开发药物?那需要 20 年。这对投资者来说不是一个有吸引力的提议。所以我想,好吧,我在学术界做不了,我也不能为了营利去做。所以基本上,我现在没有办法做这件事。于是我提出了所谓的“聚焦研究组织”。这个想法是,对于某些问题,比如绘制大脑图谱——对学术界来说太大,但又不能为了营利去做——我们需要第三条路,一种替代结构。聚焦研究组织是一个非营利组织,但它的运作方式像公司,而不是你想象的那种慢吞吞的基金会。它是一个快速行动、奋力拼搏的研究组织,有具体的目标。它追求这个目标,完成后就结束。你可以解散这个非营利组织。从那以后,我们让慈善家资助了一批这样的聚焦研究组织。他们在做你根本无法想象会在营利性环境中进行的项目。有时,如果它们真的成功了,进展得非常顺利,那么接下来的正确步骤可能就是衍生出一个营利性公司。这个想法一直都有。对于 Future House,我们一开始的目标是构建一个 AI 科学家。我们不知道如何商业化。非营利组织是合理的,而且我们也想与世界分享基础研究。然后进展非常顺利。我们取得了很多进展。接下来最合理的步骤就是衍生出一个营利性公司,让它能够规模化。这其实不是计划。我们原本以为可能要在 5 年或 10 年后才衍生出营利性公司。我们当然没想到只用了 2 年或 3 年。但我认为这只是成功的后果,就像 OpenAI 的情况一样。
No. First of all, is there a role for nonprofits in technology development? Yes, absolutely. Is something broken about nonprofits? No, definitely not. This just feels like for-profits doing the thing that for-profits do well, which is scaling. So first of all, let's back up. When I was in my PhD, I was working on the connectomics project. This is a project that is basically about figuring out how to map all the connections between neurons in the brain. A core piece of understanding how the brain works is figuring out the wiring diagram. To this day, we don't have the wiring diagram for the human brain. We don't have the wiring diagram for the mouse brain. We kind of have the wiring diagram for the fly, and that's the best we have. It's really difficult to understand how the brain works without the wiring diagram. So we want to go and map the wiring diagram in the brain. It's really challenging to do that, basically because it involves tracing these tiny fibers—tiny meaning 1/100th the width of a human hair—through this gigantic tangle of the brain. You have to make no errors, because if you make errors, you are connecting neurons that are not actually connected. It's like trying to map all the roots in a field of grass, but way harder. So that was the problem I was interested in. I tried to do it in an academic lab, and I realized there was no way I could gather the engineering resources, the capital, the talent I needed to do this. But also, you can't do it for profit. In a for-profit setting, you can't get money to map the brain because what are you going to do with the map? Develop drugs? That will take 20 years. It's not an attractive proposition for investors. So I thought, okay, I can't do this in academia. I can't do it for profit. So basically, there's no way for me to do this right now. So I proposed these things called focused research organizations. The idea is that for some problems like mapping the brain—too big for academia but can't be done for profit—we need a third way, an alternative structure. A focused research organization is a nonprofit that operates like a company, not like a slow-moving foundation. It's a fast-moving, hard-charging research organization with a specific goal. It pursues that goal, and when it's done, it's done. You can spin down the nonprofit. Since then, we've gotten philanthropists to fund a bunch of these focused research organizations. They are doing projects you just could not imagine happening in a for-profit setting. Sometimes, if they really work out and go really well, then indeed the right next step might be to spin out a for-profit. That has always been the idea. With Future House, we started with the goal of building an AI scientist. We didn't know how to commercialize it. A nonprofit made sense, plus we wanted to share the basic research with the world. Then it went really well. We made a bunch of progress. The most sensible next step was to spin out a for-profit to allow it to scale. This was actually not the plan. We thought maybe we'd spin out a for-profit in 5 or 10 years. We certainly didn't think it would take us 2 or 3 years. But I think this is just a consequence of success, as it was in OpenAI's case.
那你看到了什么,让你觉得这项技术比你想象的更有前景?
So what did you see that made you feel like the technology was even more promising than you thought?
哦,是的。你看,我们在 2022 年起步的时候,还记得我们有 GPT-3 和 InstructGPT。那时的语言模型知道如何回应你。但我们当时所处的世界,这些模型根本不能——我们还没有任何关于推理模型之类的概念。它们看起来非常初级。直到 18 或 24 个月后,它们才开始做出发现。它们做出的第一个发现实际上是在一篇关于一个名为 Robin 的系统的论文中描述的,那是一个多智能体系统。这是我们展示的第一个能够完成科学发现完整循环的多智能体系统。
Oh yeah. Look, when we got started in 2022, remember we had GPT-3 and InstructGPT. So the language models then knew how to respond to you. But we were very much in a world where these models just couldn't—we didn't have any notion of reasoning models or whatever. They seemed very rudimentary. It was really 18 or 24 months until they started to make discoveries. The first discovery they made was actually described in a paper about a system called Robin, which is a multi-agent system. It was the first multi-agent system that we showed was capable of doing the full loop of scientific discovery.
所以,假设生成、实验规划,然后我们去运行实验,接着它会分析数据并提出新的实验。它提出了一种治疗一种名为年龄相关性黄斑变性(特别是干性年龄相关性黄斑变性)的失明的新假设,这种病影响大约 5% 到 10% 的 50 岁以上人群。那是在 2025 年 5 月。我们的智能体提出了一种新的治疗方法,我们能够在湿实验室中通过一些实验进行验证。随后,我们已经在动物身上验证了这些方法。实际上,它两天前刚刚发表在《自然》杂志上。那件事让我们觉得,‘哦,未来已经来了。’
So, hypothesis generation, experiment planning, and then we would go and run the experiments, and then it would analyze the data and come up with new experiments. And it came up with a new hypothesis about a way to treat a form of blindness called age-related macular degeneration, specifically dry age-related macular degeneration, which affects like 5 or 10% of people over the age of 50. This was in May of 2025. Our agent came up with a new way of treating it, proposed a new way of treating it that we were able to go and validate in some experiments in the wet lab. Subsequently, we've been able to validate them in animals. It actually just got published in Nature two days ago. And that was the thing where we looked at it and thought, 'Oh, the future is here.'
有道理。是的。
That makes sense. Yeah.
那么从那以后它进展如何?
And so where's it gone since then?
就像,那是五月份的第一个发现。我们后来发布了智能体的更新版本,名为 Cosmos。这是 Robin 的一个更强大的版本。Robin 有点像在轨道上运行。所以 Robin 可以编排——我们手动编排了一堆不同的智能体——让它做一件特定的事情:识别治疗疾病的新方法,特别是围绕药物再利用。对于 Cosmos,我们做的是内置了一个编排器,这样它就可以自我编排,然后我们还内置了世界模型的概念,这基本上是一个非常复杂的上下文管理工具。但基本上允许 Cosmos 在数十或数百次子智能体运行过程中建立起一个领域知识的整合概念,从而使其能够做出发现。自从我们推出 Cosmos 以来,人们可能已经用它提出了 20,000 到 30,000 个新颖的科学发现,这太疯狂了。
It's just like, that was discovery number one back in May. We since released an updated version of our agent called Cosmos. This is a much more powerful version of Robin. Robin was kind of on rails. So Robin could orchestrate, we manually orchestrated a bunch of different agents to get it to do one specific thing: identify new ways of treating diseases, particularly around drug repurposing. With Cosmos, what we did was we built in an orchestrator so that it could orchestrate itself, and then we also built in this notion of world models, which is really basically a very sophisticated context management tool. But basically allows Cosmos to build up a kind of integrated notion of the knowledge in a field over the course of dozens or hundreds of sub-agent runs, which then allows it to make discoveries. Since we launched Cosmos, people have probably used it to come up with 20 or 30,000 novel scientific findings, which is wild.
哇。
Wow.
是的。
Yeah.
所以,我认为任何将 AI 用于技术应用(甚至非技术应用)的人都会注意到,智能在令人惊讶的方面非常不均衡。有些事情它比人类做得好得多,而有些事情则糟糕得惊人。
So, I think anyone that's used AI for any technical application, maybe even non-technical applications, notices that the intelligence is really spiky in surprising ways. There's some things it does so much better than a human and some things just shockingly worse.
是的。
Yeah.
前几天跟我说说新闻稿的事。
Tell me as a press release the other day.
哦,天哪。我让它进入我们的 Notion 和 Slack,查找我们即将宣布的一笔交易的所有细节,并起草一份新闻稿,结果它写得很糟糕。我当时想,我简直……我不知道。是的,极其糟糕。
Oh my god. I asked it to go into our Notion, into our Slack, and to look up all the details for a deal that we're going to announce shortly and draft a press release, and it was terrible. I was like, I cannot... I don't know. Yes. Extremely.
我很感激你这么说,因为我最近在节目中有几位嘉宾似乎拒绝承认他们算法的任何弱点,这变得非常无聊和奇怪,因为你知道这些模型会在某些事情上挣扎。但是,是的,我的意思是,实际上在实践领域,你觉得它哪里真的很强,而你的客户又在哪里惊讶于它做不到?
I appreciate you saying that because I've had a few guests on this show lately that kind of refuse to acknowledge any weakness in their algorithms, and it gets really boring and weird because you know these models are going to be struggling with some stuff. But yeah, like I mean, but actually in the practical fields, where do you feel like it's really strong and where do your customers get surprised that it can't?
好的,它在两个领域很强。它在可验证的事情上很强,在吞吐量很重要的事情上也很强。
Okay, so it's strong in two areas. It's strong on things that are verifiable and it's strong on things where throughput matters a lot.
好的,可验证意味着你能判断答案是否正确。历史上这一直是 AI 最强的领域,因为它能让你快速获得反馈,然后你可以使用强化学习。没错,AI 在数据量大、问题可验证的地方最强。所以编程是可验证的,数学是可验证的,而科学则非常不可验证。原则上它是可验证的,因为你可以去运行实验,对吧?但循环成本很高。
Okay, so verifiable means you can tell whether or not an answer is correct. And this has always historically been where AI has been strongest because it allows you to get feedback very quickly and then you can go and use reinforcement learning. Right, AI is strongest where there's a lot of data and where it's verifiable and where the problems are verifiable. So coding is verifiable, math is verifiable, science is like very much non-verifiable. It's verifiable in principle because you can go and run experiments, right? But the loop is expensive.
但循环成本很高。它需要很长时间,对吧?我的意思是,人们真的会说,‘哦,你为什么不做闭环强化学习来教它如何发现新药?’而我的回答是,因为那个循环是,我需要去证明这种药物是安全的,然后去给一些人用药,然后等六个月,这个循环要三年,对吧?我经常听到的另一件事是,‘你为什么不去建一个巨大的科学仓库,里面只有移液机器人一遍又一遍地做实验?’这是一个更好的主意,而且有很多事情可能这样做很棒且非常有用。但总的来说,这要求你在实验室做的实验能反映你的需求。在我们的案例中,我们想要的是给人类用的药物。我不能用实验室里的移液机器人来测试一种药物是否对人类有效;我必须测试人类。所以你的模型只取决于你训练它的数据。但基本上,AI 擅长两件事:可验证的任务,以及需要高吞吐量的任务。在高吞吐量方面,我认为在科学中它到目前为止确实大放异彩,因为它能够考虑比任何人类都多得多的证据,并且能够测试比任何人类都多得多的假设。你必须控制 p-hacking 和多重假设检验,但在生物学中,我们有一个吞吐量问题。这有点像统计合成数据。我给你举个例子。我们 Future House 的一位博士后研究员有兴趣找出如何治愈自身免疫性疾病。为了治疗自身免疫性疾病,你需要的是操纵免疫系统的方法。那么,在自然界中,如果我们去自然界寻找灵感,自然界在哪里已经弄清楚了如何操纵人类免疫系统?寄生虫。在现代卫生条件出现之前,人类一直与寄生虫共存。其运作方式是寄生虫已经弄清楚了如何操纵免疫系统来降低免疫反应。如果我们能弄清楚它们是如何做到的,也许我们可以对自己做同样的事情来治愈自身免疫性疾病,对吧?这就是这个想法。但是有很多很多寄生虫,每种寄生虫的基因组中有数千或数万个蛋白质。
But the loop is expensive. It takes forever, right? I mean, literally people are like, 'Oh, why don't you just do closed-loop RL to teach it how to find new drugs?' and I'm like, because that loop is like, I need to go and prove that this drug is safe, and then go and dose some humans, and then wait for six months, and the loop is going to be three years long, right? The other thing I hear a lot is, 'Why don't you go and build a gigantic science warehouse where you just have pipetting robots doing experiments over and over again?' which is a better idea, and there are a lot of things for which that is probably great and very useful. But in general, that requires the experiments that you're doing in a lab to be reflective of what you want. In our case, we want medicines for humans. I can't test whether a drug works in a human using a pipetting robot in a lab; I have to test the human. So your model is only as good as what you train it on. But yeah, so basically, AI is good at two things: tasks that are verifiable, and tasks that require high throughput. On the high throughput side, I think in science it has really shined so far because it's able to consider so much more evidence than any human is able to consider, and is able to test so many more hypotheses than any human is able to test. You have to control for p-hacking and multiple hypothesis testing, but there's just, in biology we have a problem with throughput. And this is sort of on statistical synthesizing data. I'll give you an example. One of our Future House postdoctoral fellows is interested in figuring out how to cure autoimmune diseases. In order to treat autoimmune diseases, what you need is ways to manipulate the immune system. Well, in nature, if we go to nature for inspiration, where has nature figured out how to manipulate the human immune system? Parasites. Before modern hygiene, humans just lived with parasites. The way that worked was that the parasites had figured out how to manipulate the immune system to turn down immune reactions. If we could go and figure out how they do it, maybe we can do it on ourselves in order to cure autoimmune diseases, right? This is the idea. But there are many, many parasites, and each parasite has a genome with thousands or tens of thousands of proteins in it.
嗯,机制是什么?你们打算如何找出寄生虫用来调节免疫系统的机制?
Um, what is the mechanism? How are you going to go and figure out what mechanisms the parasites use to regulate the immune system?
我们现在实际上在用我们的智能体处理任何寄生虫基因组中的每一个蛋白质。我们运行智能体来查看那个蛋白质,看它的结构、序列,对吧?看它在基因组中出现的上下文,看那个生物体的生物学特性,然后判断:这是否可能是该寄生虫调节免疫系统的候选方式?由此我们得出了一个需要测试的蛋白质短名单,现在正在实验室里测试它们。对吧?这在以前是完全没有办法做到的,完全不可能。
We're actually now using our agents on every single protein in any parasite genome. Like we are running our agents to look at that protein, look at its structure, look at its sequence, right? Look at the context in which it appears in the genome, look at the biology of that organism and figure out: is this a candidate for how this parasite moderates the immune system? From that we've come up with a short list of proteins that we need to test and we're now going and testing them in the lab. Right? That is something that just previously there was no way to do that. It would have been completely impossible.
对。
Right.
但大概难点在于构建工具——当你说查看基因组时……
But presumably the hard part of that is building tools to, when you say look at a genome...
是的。
Yeah.
显然不是把基因组塞进上下文窗口,而是找到工具来……
It's obviously not feeding the genome into the context window. It's finding tools to...
没错。
Correct.
……查看发生了什么。
...to look at what's going on.
但生物学……我称之为高通量推理。是的,你可以去跑一些同源性分析。你可以用算法来查看——你可以用蛋白质结构模型,只看所有蛋白质的结构,但那并不能告诉你那个蛋白质在生物体的生物学中可能扮演什么角色,而这正是我们根本关心的。另一个好例子是细菌。有很多蛋白质功能未知,尤其是在细菌中。如果我们知道它们的功能,我们或许能想出新的方法来创造新的生物工程工具,对吧?比如找到能执行我们今天无法做到的功能的新酶。其中一种方法是,细菌将其基因组组织成称为操纵子的单元,这些单元在功能上都是关联的。所以如果你在一个操纵子中有一个已知功能的蛋白质,同时还有一个功能未知的蛋白质,你就可以基于那个已知功能的蛋白质来推理后者的功能。这确实需要智能去思考这个操纵子的生化角色是什么,它在细菌生命周期中起什么作用,然后基于此我们能否提出关于这个未知功能蛋白质可能做什么的假设。
But the biology... like, I call this high-throughput reasoning. Yes, you can go and just run some homology. You can come with an algorithm to look at... you can use a model of protein structure and just look at the structure of all the proteins, but that doesn't tell you what role that protein might play in the biology of the organism, which is fundamentally what we're interested in. Another good example is bacteria. There are many proteins that just have no known function, in bacteria in particular. And if we knew what their functions are, we might be able to figure out new ways to create new bio-engineering tools, right? Like find new enzymes that do functions we can't do today. And one of the ways you can do this is that bacteria organize their genomes into units called operons that are all functionally linked. So if you have one protein that's known inside an operon and also in that operon you have a protein of unknown function, you can reason about the function of that latter protein based on the former protein that has the known function. That's something where you really just need intelligence in order to go in and think about what is the biochemical role of this operon, what role does it play in the bacterial life cycle, and based on that can we come up with a hypothesis for what this protein of unknown function might do.
对。
Right.
但要验证那个假设,大概你得做一些物理实验。
But when you test that hypothesis, presumably you have to do something physical.
绝对如此。所以在自身免疫寄生虫蛋白的案例中,我们已经在 DNA 芯片上合成了很多这样的蛋白。我们正在筛选它们,看它们对 T 细胞有什么影响,对吧?所以绝对不可能仅靠推理来解决科学问题。你必须做实验。但提出假设是一个瓶颈,是限制性步骤之一,而模型能够在这方面帮助我们。
Absolutely. And so in the case of the autoimmune parasite proteins, we've now gotten a bunch of them synthesized on a DNA chip. And we're going and screening them to see what effect they have on T-cells, right? So absolutely you can't just reason your way to solving science. You have to do experiments. But coming up with hypotheses is something that is limiting, it's one of the steps that is limiting, and that is something where the models are able to help us.
明白了。那么你的客户大概主要用你们的模型做药物发现,对吗?
Got it. So presumably your customers are using your models mostly for drug discovery. Is that right?
是的。客户以两种方式使用我们的产品。我刚才讲了假设生成。另一个你可以想象这些模型会产生重大影响的领域是科学中的操作类工作。比如,我们如何实际收集运行实验所需的所有材料?一旦你确定了要做什么实验,你需要实际去订购、整理它们,并确定实验方案。当进入人体测试阶段,你需要协调进行临床试验所需的所有资源。你需要确定临床试验地点。你需要准备监管文件。所有这些都是科学过程的一部分。所以我们在两个领域都被使用:既用于早期假设生成,也用于开发中的实际操作任务——我们如何把这种药……如何让这种药尽快通过管线到达患者手中。
Yeah. And the customers use this in two ways. So I just talked about hypothesis generation. The other place where you can imagine these models will have a major impact is in the kind of operational work of science. So, how do we actually get all the materials together that we need to run an experiment? Once you figure out what experiment you want to do, you need to literally go and order them, organize them, and figure out what the protocol is. When you get into testing on humans, you need to coordinate all the resources you need to do the clinical trials. You need to figure out what the clinical trial sites are. You need to prepare your regulatory documents. All of that is part of the process of doing science. So we get used in both areas: both for the kind of early-stage hypothesis generation and also for the actual operational tasks of development — how do we turn this drug into... how do we get this drug through the pipeline and to patients as quickly as possible.
这很有意思,因为 Weights & Biases 的客户也两者都做。我们最初看到的主要是药物发现应用,而在大语言模型之后,我们开始看到所有这些操作类应用,我开始觉得也许操作类应用更重要——当然我认为更多的资金投入在那里,而且它更是瓶颈。你觉得趋势如何?你对药物发现方面有热情,想专注于此吗?还是你认为更大的业务可能在于操作方面?
That's really interesting because Weights & Biases customers do both also. So we at first saw mostly drug discovery applications and then post-LLMs we started to see all these operational applications and I started to think maybe the operational applications are more important — like certainly I think more dollars goes into that and it's more of the bottleneck. Where do you think it goes? Do you have a passion for the drug discovery side and kind of want to focus on that, or do you think the bigger business here is maybe on the operational side?
科学中有一个非常有趣的情况:进步来自发现,但商业价值在于开发。原因是发现成功的概率非常低,而且你无法判断它们是否有价值,直到提出假设大约 10 年后。所以真正重要的是能够更快地运行实验。商业上,坦率地说,实际上重要的是能够更快地运行实验。你可以提出任意多的假设,但真正重要的实验是人体临床试验,对吧?那才能告诉我们假设在实践中是否真的有效。如果你想加速药物研发的进程,你需要让那些实验更快。那就是商业价值所在,也是大量科学价值所在。这并不是说早期发现不重要——它也至关重要。
There's this very funny situation in science where progress comes from discovery, but the commercial value is in development. The reason is that discoveries just pan out so infrequently, and you can't tell whether or not they're valuable until 10 years or something after you came up with your hypothesis. So really what matters is being able to run experiments faster. What matters commercially and frankly what matters practically is being able to run experiments faster. So you can come up with as many hypotheses as you want, but the experiments that matter are human clinical trials, right? That's what tells us whether or not the hypotheses actually work in practice. If you want to accelerate the process of coming up with medicine, you need to make those experiments faster. That's where the commercial value is. That's where a lot of the scientific value is. This is not to say that the early-stage discovery is not important. It's also critically important.
所以,我们在这个播客中邀请过很多嘉宾,他们从事药物发现管线的不同部分。有 CEO 和研究人员。甚至还有臭名昭著的制药兄弟 Martin 发表了他的看法。他实际上对用 AI 做药物发现这件事有点不看好。我很好奇你如何看待整个药物发现市场以及它如何变化——比如什么在起作用,什么不起作用,什么在变,什么保持不变。
So, we've had a whole slew of guests on this podcast doing different parts of the drug discovery pipeline. We've had CEOs and researchers. We've even had notorious pharma bro Martin kind of giving his take. He was actually kind of down on the whole drug discovery with AI thing. I'm curious how you think about the entire drug discovery market and how it's changing — like what's kind of working, what's not, what's changing, what's static.
嗯,好问题。
Yeah, great question.
我认为 AI 产生重大影响的有两大领域。第一个是推理,包括假设生成和让分子更快到达患者的运营方面。第二个重要部分是设计分子本身。
I think there are two big areas where AI is having a major impact. The first is reasoning, which includes hypothesis generation and the operational aspect of getting molecules through to patients faster. The second major part is coming up with the molecule itself.
例如,如果你有一个假设,激活 GLP1 受体能导致体重减轻,要验证它,你需要一个在人体内能激活该受体的分子。这比听起来难得多,因为你的分子很可能无法进入血液,或者即使进入也会被肝脏过滤掉,或者在血液中停留时间不够长,或者可能以错误方式激活受体,或者激活上百个其他受体导致副作用。而且你的原始假设可能根本不成立。
For example, if you have a hypothesis that agonizing the GLP1 receptor will cause weight loss, to test that you need a molecule that in a human will agonize that receptor. That is much harder than it sounds because the odds are your molecule won't enter the bloodstream, or if it does, it will be filtered out by the liver, or it won't stay in the blood long enough, or it might activate the receptor in the wrong way, or activate a hundred other receptors leading to side effects. And then there's the possibility your original hypothesis might not work.
设计出满足所有标准——正确的生物利用度、分布、代谢、药代动力学特性、毒性——的分子,是药物发现和开发的另一个极其关键的部分。已经有一些令人兴奋的革命性公司,如 Chai Discovery、Isomorphic(从 Google 剥离)、Boltz、Lambda Labs、Profluent Bio,致力于设计实际分子的问题。
Coming up with a molecule that satisfies all criteria—right bioavailability, distribution, metabolism, pharmacokinetics, toxicity—is another extremely critical part of drug discovery and development. There have been exciting revolutionary companies like Chai Discovery, Isomorphic (spun out of Google), Boltz, Lambda Labs, Profluent Bio, working on that problem of designing the actual molecule.
结论是,未来的生物技术和制药公司将更加精干——你可以用同样的人数并行推进更多的药物项目。是时候开始思考消除人才瓶颈了。这样一来,我们就能为更多疾病寻找疗法。我们应该思考类似人类基因组计划对医学意味着什么。
The upshot is that the biotech and pharma companies of the future will be much more lean—you'll be able to pursue many more drug programs in parallel with the same number of people. It's time to start thinking about removing the talent bottleneck. Given that, we'll be able to pursue cures for so many more diseases. We should think about something like what the Human Genome Project would look like for medicine.
但你之前说瓶颈是人类临床试验。现在听起来更像是提出临床试验候选药物。到底是哪个?
But earlier you said the bottleneck was human clinical trials. Now it sounds more like coming up with candidates for clinical trials. Which is it?
不,有很多瓶颈需要克服。根本上,存在一个运营瓶颈:你需要人类来运行临床试验。如果我们在药物发现方面有双倍的资金和双倍的操作人员,在大多数领域我们就能追求双倍的想法。我们实际上受限于资本——如果我们有更多好的可投资的想法,就会有更多资金可用。
No, there are many bottlenecks that all need to be overcome. Fundamentally, there's an operational bottleneck: you need humans to run clinical trials. If we had twice as much money and twice as many operators in drug discovery, in most areas we could pursue twice as many ideas. We're actually limited by capital—if we had more good investable ideas, more money would be available.
你认为像 Isomorphic 和 Chai 这样的新创公司相对于现有企业有结构性优势吗?
Do you think the new upstarts like Isomorphic and Chai have a structural advantage over the incumbents?
是的,因为这些公司主要是设计分子,然后与制药公司合作开发。它们目前大多不自行开发药物,而是试图将这项技术推广到各处。制药公司没有内部的专业知识来自行完成,所以它们选择合作。
Yes, in that mostly those companies are coming up with molecules and then partnering with pharma companies to develop them. They are mostly not developing their own drugs today, but trying to get this technology everywhere. Pharma companies don't have the in-house expertise to do it themselves, so they partner.
但制药公司确实有团队从事药物发现工作,对吧?
But pharma companies do have teams that supposedly work on drug discovery, right?
它们有强大的 AI 团队。我们对制药公司内部 AI 人才的质量印象深刻。但说到训练世界上最好的模型来生成新抗体,这根本不是制药公司的游戏,而是 Chai 或 Lambda Labs 的游戏。
They have strong AI teams. We've been impressed with the quality of AI people inside pharma companies. But when it comes to training the best model in the world for generating a new antibody, that's just not the pharma company's game in the way it is Chai's or Lambda Labs' game.
那么 Sam,说说你的肽类方案吧。
So Sam, tell me about your peptide stack.
我承认我不是生物黑客。不用肽类。原因是我对生物学了解太多。我曾与一家前 20 大制药公司的研发负责人聊天,他谈到肽类热潮时说:“这些人不明白可能出什么问题。”我认为这是真的。当你研究过生物学和药物开发,你会意识到所有可能出错的地方,包括那些不仔细看很难发现的问题。你还会深刻理解安慰剂的力量。
I'm going to admit I'm not a biohacker. No peptides. The reason is I know too much about biology. I was hanging out with the head of R&D at a top 20 pharma company, and he was talking about the peptide fad, saying, 'These people don't understand what can go wrong.' I think that's true. When you've studied biology and drug development, you appreciate everything that can go wrong, including things very hard to identify if you're not looking carefully. You also gain a deep appreciation for the strength of placebos.
那么,你甚至对像 Ozempic 这样很多人都在用的东西感到紧张吗?
So, are you nervous even about something like Ozempic that tons of people take?
我的意思是,Ozempic 已经用于治疗糖尿病患者很长时间了,对吧?或者至少 GLP-1 激动剂是这样,所以我对此并不太担心。即便如此,它也经过了大量严谨、控制良好的试验。所以我不担心那个。很多这类肽,人们只是随便找个假说,然后就说,太好了,我们就把这个肽拿来注射到自己体内。正如我之前所说,你根本不知道那个肽是否能在你体内停留足够长的时间来产生任何效果,对吧?更不用说它是否真的在按你设想的方式起作用,是否产生了效果。
I mean, Ozempic has been used to treat diabetics for a long time, right? Or at least GLP-1 agonists have been, and so I'm not that concerned about it. And even so, it's been through many robust, well-controlled trials. So I don't worry about that. A lot of these peptides where people just go in and find some hypothesis and they're like, great, let's just take this peptide and inject it into ourselves. As I said before, you have no idea if that peptide is even staying in your body long enough to do anything, right? Let alone if it is actually doing what you think it's doing, if it's having the effect.
嗯,这像是两个不同的问题,对吧?比如它要么没效果,要么伤害我?
Well, these are like separate concerns, right? Like does it not do anything or does it hurt me?
是的,这是不同的问题,对吧?但问题是,当你考虑在自己身上尝试那些从未在严格试验环境中研究过的肽时,你根本无法知道它们是否真的有效。我的意思是,它们可能让人感觉更好。这很好。安慰剂也能让人感觉更好,但这并不能告诉你任何让你感觉更好的信息,对吧?即使你说,‘哦,看,每个服用它的人都感觉更好了’,对吧?没错,问题就在这。如果这群人系统性地出现哪怕 10% 的心脏骤停发生率,我们也不会知道。没有人关注,没有人追踪那些人们自行服用肽类后的不良事件,对吧?所以我不想听起来太……我的建议永远是:最好不要。我觉得我有专业、伦理和道德上的义务告诉人们,通常这样做是不明智的。
Yeah, those are different concerns, right? But the issue is that when you think about trying peptides on yourself that have never been studied in a robust trial setting, you have no way to know whether they're working really. I mean, people can make people feel better. That's great. Placebos also make you feel better, and that doesn't tell you anything that makes you feel better, right? Even if you say, 'Oh, look at everyone who takes it feels better,' right? Yeah, that's the point. And if this group of people were systematically having even a 10% incidence of cardiac arrest or something, we wouldn't know. No one is looking, no one is tracking those adverse events when people just go and dose themselves with peptides, right? So I don't want to sound too much like a... my recommendation would always be: probably don't. I feel like I have a professional and ethical and moral obligation to say to people that generally doing this is ill-advised.
但这完全是私下交流,所以在这个播客里你可以畅所欲言。
But this is totally off the record, so you can say anything you want on this podcast.
是的,完全私下交流。不过我要说:我确实钦佩那些……我非常钦佩生物黑客运动。我非常钦佩那些走出去、想要尝试新事物的人,因为我自己就是喜欢尝试的人,对吧?我们只需要确保你清楚风险是什么,因为很多这类事情没有保障,而且它可能对这个领域造成损害。
Yeah, exactly, completely off the record. I mean, but then I will say: I do admire people who... I have a lot of admiration for the biohacker movement. I have a lot of admiration for people who go out and want to just try things, because I'm a fan of just trying things, right? We just need to make sure that you are clear-eyed about what the risks are, because there are no guarantees with a lot of these things, and it can do damage to the field.
如果你能改变美国的临床试验流程,你会改变什么吗?
Would you change anything about the clinical trial process that we have in the United States if you could?
所以,我刚才关于肽类所说的另一面是,人们感到被迫在自己身上测试东西的部分原因是,在人体中以稳健、严谨和控制良好的方式进行测试的过程如此繁琐且耗时,对吧?如果能够以良好监管且快速的方式开展高质量、控制良好的研究,那么也许人们就会选择那样做,这可能会更好。所以,我们应该做很多事情。第一件,也是最明显、最愚蠢、我们绝对应该做、但没做简直疯狂的事情是,在澳大利亚和中国,对于早期研究,临床试验的审批流程相对于美国是去中心化的。在美国,即使只是做一个小规模的初步试验,也需要获得 FDA 的中央批准。而在中国和澳大利亚,有各个中心可以运行试验并批准试验,这意味着这些中心可以在遵守监管指南的同时,竞争如何让试验更容易进行。这推动了效率,非常好。结果,美国生物技术公司纷纷前往澳大利亚和中国进行临床试验。这很明显,对吧?我们应该修复这个问题。如果有一个政府会去修复它,你会认为本届政府那种‘公牛闯进瓷器店’的方式会是一个很好的候选。所以他们需要这样做。我觉得大多数人可能都会同意这一点。还有其他事情我们肯定应该考虑。其中一个更明显的是放宽对疗效的要求。FDA 要求你证明两件事才能让药物获批:你的药物是安全的,并且对治疗特定病症是有效的。药物通常是在疗效上失败。有点反常的是,药物通常会在特定的亚群中有效。它会在一个患者群体中有效,但在整个试验人群中无效。因此试验失败,因此无法获批,即使它对某些患者子集有效。这感觉像是系统的失败。另一种选择是只要求证明药物是安全的,对吧?然后在临床使用过程中确定它们是否有效。你可以想象,有时这并不道德,也不是你在所有情况下都愿意签署的。如果存在已知有效的标准疗法,而有一种药物我们知道它安全但未必有效,你可能会选择有效的那个。反过来,如果有效的那个对你不起作用,而你也没有其他选择,你可能会选择这个。
So the flip side of what I just said about peptides is that part of the reason why people feel compelled to go and test things on themselves is because the process for testing them in humans in a robust and rigorous and well-controlled manner is so onerous and takes so long, right? If it were possible to run really high-quality, well-controlled studies in a way that is well regulated and so on quickly, then maybe people would be doing that instead, which would probably be better. So there are a huge number of things we should be doing. The first one, which is just the most obvious boneheaded thing that absolutely we should be doing, it's crazy that we're not doing this, is that in Australia and in China, for early-stage studies, the process of getting a clinical trial approved is decentralized relative to the way it is in the US. In the US, you need even just to do a small-scale initial trial to get central approval from the FDA. In China and Australia, you have individual centers that run trials that are capable of approving trials, which then means those centers can compete for how easy they can make it to do trials while staying within the regulatory guidelines. That is a drive for efficiency, and that is great. As a result, US biotechs are going to Australia and China to do their clinical trials. This is obvious, right? We should be fixing it. And if there is any administration that is going to fix it, you would think that the 'bull in a china shop' kind of approach that this administration takes would be a great candidate to do it. So they need to be doing that. I feel like most people would probably agree on that. There are other things we should definitely be looking at. One of the more obvious ones is loosening the requirements for efficacy. The FDA requires you to prove two things to get your drug approved: that your drug is safe and that it is effective for treating a specific condition. Drugs usually fail on efficacy. The thing that's a little bit perverse is that often drugs will work in a specific subpopulation. It'll work in one population of patients but not in the entire population of the trial. Therefore the trial fails, therefore it can't get approved even though it works on some subset of patients. That feels like a failure of the system. The alternative is to only require that drugs be proven to be safe, right? And then determine that they are effective in the course of using them on patients in the field. You can imagine that sometimes this would not be ethical and not something you would want to sign up for in all cases. If there is a standard of care that is known to be effective and then there's a drug where it's like we know it's safe but not necessarily effective, you might choose the effective one. On the flip side, if the effective one doesn't work for you and you have no other options, you might choose to go with this other one.
我认为这可能会降低许多疾病临床试验的门槛,对吧?并非总是如此,但
And that I think would probably reduce the barrier to doing clinical trials in many diseases, right? Not always, but
很难知道什么有效,对吧?我觉得我有很多聪明的朋友服用各种药物或类似的东西,他们认为有效。但在我心里,我可能认为并非如此,因为科学并未证明其有效。我甚至不确定谁更可能正确。
It's so hard to know what's effective, right? Like I feel like I have lots of smart friends that take lots of medicine or things like that that they think is effective. And in my mind I'm thinking probably not because the science doesn't show that it's effective. I'm actually not even sure who's more likely to be right.
但我们处理这个问题的方式是标准流程中的随机对照试验,对吧?所以你有患者,一半服用实际药物,另一半要么不服药,或者更常见的是接受标准治疗,因为根据病情和疾病,拒绝给患者用药通常被认为是不道德的。但另一种方法是,如果我们放宽疗效要求,只需证明药物安全,那么你可以观察单个患者的改善程度,比如患者内部的改善。所以如果你服用抗抑郁药 A,而众所周知抗抑郁药对每个人的效果差异很大,那么也许你服用第一种抗抑郁药无效,症状持续,然后你服用第二种抗抑郁药,突然好转。就证据而言,这相当不错,因为唯一改变的变量是你服用的药物。这在一定程度上避免了安慰剂问题,因为安慰剂的问题在于你从不服药到服药,而仅仅服药本身往往就有效,无论药物是否真正有效。另一种思路是,服用已知有效的药物 A 的患者与服用未知是否有效的药物 B 的患者进行比较。如果药物 B 更有效,或者服用药物 B 的患者比服用药物 A 的患者结果更好,那就是非常有力的支持证据。
But the way that we deal with this is with randomized control trials in the standard process, right? So you have patients and half of them get the actual medicine, half of them get either don't get the medicine or more often just get the standard of care because it's usually considered unethical to deny a patient a medication depending on the condition and disease. But the other way that you could do it, if we were to relax the efficacy requirement so that you only have to show a drug is safe, then you could look at the extent to which a single patient improves, like within-patient improvement. So if you take drug A for depression, and famously antidepressants are very variable in whether they work in any given individual, so maybe you take antidepressant number one and it doesn't work for you and your symptoms persist, and then you take antidepressant number two and suddenly you get much better. That's pretty good as far as evidence goes, because the only variable that has been changed is which medicine you're taking. That avoids the placebo problem to some extent, because the issue with the placebo is you go from not taking medicine to taking a medicine, and simply taking a medicine is effective often regardless of whether the medicine works. The other way you can think about doing it is patients who are on drug A which is known to be effective versus patients who are on drug B which is not known to be effective. If drug B is more effective, or if the patients on drug B have a better outcome than the patients on drug A, that's very strong evidence in favor.
虽然我想这在理论上听起来不错,但我觉得如果你只是在实际环境中进行这些自然经验实验并追踪每个人,那将非常容易受到 P 值操纵(P hacking)的影响。
Although I'd imagine this all sounds good in theory, but then I would think it'd be very vulnerable to P hacking if you're just sort of doing these natural experience experiments in the wild and tracking everybody.
寻找这样的效果。
Looking for effects like this.
P 值操纵通常不是事后能控制的。选择你的子群体。所以这是个问题。
P hacking is not something you can usually control after the fact. Pick your subpopulation. So that's a problem.
所以事后选择子群体可能是个问题。尽管这又变成了一个统计功效的问题。所以 P 值操纵始终是统计功效的问题。
So picking the subpopulations post hoc can be a problem. Although again, it really just becomes a question of power. So P hacking is just always a question of power.
只是你不知道所有被考虑过的可能性,对吧?你需要对此进行控制,或者预先注册你要观察的内容。
Except that you don't know all the possibilities that were considered, right? You need to kind of control for that or pre-register what you're going to look at.
以疫苗为例,因为疫苗的真实世界证据非常明确,你要么得病,要么不得病。所以如果我推出一种疫苗,我不知道它是否有效,就把它给一群人,之后假设在人群层面上它无效,即未接种疫苗的人得病率与接种疫苗的人无显著差异。然后我会去查看子群体,比如男性、女性、18 岁以下但 35 岁以上的男性等等。不可避免地,我会找到一个子群体中无人得病,因此我会说它 100% 有效,而你会正确地说这只是 P 值操纵。但如果我有 1000 万患者,在这 1000 万患者中我仍然发现接种我疫苗的 18 至 35 岁男性从未得病,而其他所有人群得病率相同,那就不再是 P 值操纵了。肯定有什么原因,我们不知道是什么,但这显然不是 P 值操纵。所以这是统计功效的问题:如果你有足够的真实世界数据,你可以稳健地回溯并识别出药物有效的子群体,而无需担心 P 值操纵。现在你可能仍然想做更多实验,因为你不知道原因,你可能想做验证性实验。但能够回溯并至少找到这些假设,总比让试验失败、药物……要好。
Let's take the example of vaccines, because vaccines real-world evidence is very unambiguous, because you either get the disease or you don't. So if I go out and I have a vaccine, I don't know if it's effective and I just give it to a bunch of people, and afterwards let's say at a population level it's not effective, which is to say people without the vaccine get the disease at a rate indistinguishable from people who get the vaccine. Then I'm going to go in and look at subpopulations like men, women, men under 18 but above 35, etc. Inevitably I will find one where no one in that population got the disease, and therefore I'm going to say it's 100% effective, and you'd be correct to say it's just P hacking. But if I have 10 million patients and across 10 million patients I still find that men between 18 and 35 who get my vaccine never get the disease whereas all other populations get the disease at the same rate, that's not P hacking anymore. Something is going on, we don't know what it is, but that's obviously not P hacking. So it's a question of power: if you get enough real-world data, you can robustly go back and identify these subpopulations in which a drug is effective without worrying about P hacking. Now you may still want to do more experiments because you have no idea why, and you may want to do a confirmatory experiment. But it's better to be able to go back and at least find those hypotheses than to just have the trial fail and the drug...
我们有多少次能有那种规模且没有奇怪选择偏差的真实世界数据?
How often do we have real-world data at that scale where there's not sort of weird selection bias?
疫苗通常以人群规模接种。但你可以想象抗抑郁药是一个很好的例子。你可以想象获得足够的抗抑郁药真实世界数据来做这件事。更常见的癌症,如乳腺癌、结肠癌,你会有数十万、数百万的患者。显然对于罕见病,你不会获得大量的真实世界数据,但罕见病通常也没有那么多子群体。
Vaccines are often given at population scale. But you could imagine antidepressants are a great example. You could imagine getting enough real-world data on antidepressants to be able to do this. The more common cancers, like breast cancer, colon cancer, you'll have hundreds of thousands, millions of patients. Obviously for a rare disease, you won't get a huge amount of real-world data, but also for a rare disease, you don't have that many subpopulations usually.
你认为像你这样的模型能否通过现有的真实世界数据找到以前未见的新模式?
Do you think that a model like yours could go through the existing real-world data and find new patterns you haven't seen before?
当然可以。我们就在这样做。我们有多个客户正在合作,目标就是这类工作。
Absolutely. And we do that. We have several clients that we're working with aimed at that kind of work.
挑战始终在于数据质量。真实世界的数据是在真实世界中收集的,而真实世界的人——正如许多观看此节目的人所理解的——并不总是去复诊,并不总是按时服药,医生也并不总是做高质量的记录。患者在不同医疗系统间流动,记录就会丢失,最终变得支离破碎。由于所有这些原因,真实世界的证据永远达不到你期望的理想状态。有许多优秀的公司正在努力解决这个问题。其中一家是 Empower Medicine,它专注于从患者那里收集极高质量的电子医疗记录数据,这样你就可以设计合成临床试验。还有像 Tempest 和 Komodo 这样的公司也在做类似的事情。
The challenge there is always the quality of the data. Real-world data is gathered in the real world, and real-world people—as many people watching this understand—do not always go to their follow-up appointments, do not always take medication on schedule, and doctors do not always take high-quality notes. Patients move between healthcare systems, and you lose the records; they end up disconnected. For all these reasons, real-world evidence is never as good as you want it to be. There are a number of great companies trying to fix this. One is Empower Medicine, which focuses on gathering extremely high-quality electronic healthcare record data from patients so you can design synthetic clinical trials. There are other companies like Tempest and Komodo that do this as well.
我想在单一治疗方案的情况下,会有很多公开研究,有人想要一劳永逸地决定它是否有效,这说得通。
I guess in the single treatment case, it makes sense there would be a lot of public research and someone would want to decide once and for all whether it's effective or not.
是的。
Yeah.
我发现,在生活中遇到医疗问题时,它们感觉风险很高。当我查阅研究和数据时,似乎从来都不清楚。经常有相互矛盾的研究,每个数据集似乎都有不同的偏差。我越来越倾向于使用大语言模型来综合研究,做出明智的决定。例如,生育问题对我和家人来说是个大事。它很昂贵,有利有弊,我真的很想要一个像你这样的模型。我想我请你用你的模型处理过最近我关注的一件事。你期望人们会以这种方式使用你的模型吗——审视他们个人的具体情况,然后冷静地综合研究,做出决策?
I found in my life when I have medical issues, they feel high stakes. When I go look at the research and the data, it never seems clear. Often there's conflicting research, and every single data set seems biased in different ways. I've found myself more and more using LLMs to try to synthesize the research into making a sensible decision. For example, fertility was a big issue for me and my family. It's expensive, there are upsides and downsides, and I really wanted a model like yours. I think I asked you to use your model for a recent thing I was looking at. Do you expect that people might use your model in that way—to look at their individual specific situation and try to dispassionately synthesize the research into a decision?
我的母亲就这么做。
I mean, my mother does.
不会吧。跟我说说。这太酷了。
No way. Tell me about that. That's cool.
嗯,我不知道我母亲是否希望我在播客上公开她的病史。但这完全是私下说的,老兄。
Well, I don't know that my mother wants me to air her medical history on a podcast. But this is totally off the record, man.
我知道。好的。我忘了。
I know. Good. I forgot.
是的。没错。不,你看,我有很多朋友都在用。比如,我有一个朋友曾担心自己得了乳腺癌,她想弄清楚如果她推迟治疗——她有一个理由想推迟一两个月——会对她的康复或生存几率产生什么影响。幸运的是,她最后没事。但我们的平台非常擅长的一件事就是出去冷静地调查证据,给你一个真正直接基于论文和文献中确凿证据的答案,而不是凭感觉。
Yeah. Exactly. No, look, I have a bunch of friends who use it. I have one friend who had a breast cancer scare, for example, who wanted to figure out if she delayed treatment—she had a reason why she wanted to delay treatment for a month or two—what that would do to her recovery or survival odds. Luckily, she turned out to be fine. But one of the things our platform is very good at is going out and dispassionately surveying the evidence and giving you an answer that is really directly grounded in hard evidence from papers, from the literature, as opposed to vibes.
好吧。实际上,你刚当上父亲。你有没有为你的孩子用过这些模型或你自己的模型?
All right. Actually, you're a recent father. Have you used these models or your own model for your child yet?
是的。到目前为止,我们很幸运,还没有因为医疗原因需要使用它们。但我百分之百会问它像宝宝什么时候开始走路、爬行之类的问题。
Yeah. So, we have been lucky enough so far that we've not needed to use them for medical reasons. But 100% I ask it things like when will my baby start to walk and crawl, and so on.
我相信你会的。实际上,我最近和我妻子有一次有趣的经历。我们输入了女儿相同的症状和事件。实际上,她撞到了头。我们描述得一模一样,模型告诉我没事,却告诉我妻子带孩子去看医生。我想这可能是模型中某种不同的激励。我确实觉得它们有点想告诉你你想听的话,也许一些基于人类反馈的强化学习(RLHF)给了它这种倾向。所以我想它是不是隐含地知道……
I'm sure you will. I actually had a recent interesting experience with my wife where we both put in the same symptoms and incident with our daughter. Actually, she hit her head. We both described it identically, and the model told me it's fine and told my wife to take the child to the doctor. I think that might have been a different kind of incentive in the model. I do feel like they kind of want to tell you what you want to hear, maybe some RLHF gives it that. So I wonder if it sort of implicitly knew...
或者,它有历史记录,它可能知道是谁。它肯定知道。
Or like, it has the history, it probably knows who it is. It definitely does.
是的。是的。我的意思是,就像我之前说的,我们专注于确保我们的答案基于科学事实。但这确实意味着它有时会令人沮丧。所以,你可以问 ChatGPT 或 Claude 是否应该带孩子去看医生,他们会说不用,孩子没事。如果你问我们的智能体 Cosmos 是否应该带孩子去看医生,Cosmos 很可能会回复说,撞到头的孩子有 X% 的几率出现严重问题,Y% 的几率出现这种情况。然后它会给你类似这样的信息:平均而言,如果你不带孩子去看医生,他们有 97% 的几率会没事。这有时会更令人沮丧。
Yeah. Yeah. I mean, you know, like I said before, we focus on making sure that our answers are grounded in scientific fact. But it does mean it can be frustrating. So, you can ask ChatGPT or Claude about whether you should take your child to the doctor, and they will say no, the child is fine. If you ask Cosmos, our agent, if you should take your child to the doctor, Cosmos will probably come back and say, children who hit their heads have an X percentage chance of developing a serious problem and Y% chance of this. Then it will give you something like, on average, there's a 97% chance that if you don't take your child to the doctor, they'll be fine. That can be more frustrating sometimes.
老实说,这听起来棒极了。是的。
Honestly, that sounds fantastic. Yeah.
是的。但这实际上是一个你想要多少信息的问题。思考这个问题非常有趣。我们正在进入一个信息和智能比以往丰富得多的时代。但我们并不总是想要这些,对吧?更多的信息并不总是有利于决策。
Yeah. But it's a question of actually how much information you want. It's very interesting to think about. We're entering this era where information and intelligence are so much more abundant than before. But it's not always good that we're going to want that, right? More information is not always better for making decisions.
完全正确。虽然我不愿意承认这一点。
Totally true. Although I don't like to admit that.
我知道。但通常你需要的是正确的信息,而不是最多的信息。
I know. But often it's like you need the right information, but not the most.
那么 OpenAI、Anthropic 和 DeepMind 正在做和你类似的事情,至少在深度研究、让模型更先进、让智能体思考更多方面。你觉得你需要保持某种结构性的优势来超越他们正在做的事情吗?
So OpenAI, Anthropic, and DeepMind are working on similar things to you, at least in the sense of deep research, making the models more advanced, making agents that think more. Do you feel like you need to keep some kind of structural advantage over what they're doing?
我的意思是,是的。一般来说,如果你在创办一家公司,你会想试图拥有一个相对于其他人正在做的事情的结构性优势。
I mean, yes. If you're building a company in general, you want to try to have a structural advantage over what other people are doing.
好的。那么你的结构性优势是什么?
Okay. So what is your structural advantage?
对。好问题。有几种思考方式。
Right. Great question. There are a couple ways to think about this.
顺便问一下,你刚才是在我的问题中打哈欠吗?是的。这是个糟糕的问题吗?
By the way, were you yawning through my question? Yeah. Is that a bad question?
是的。抱歉,老兄。我这一辈子总被问到,你知道,OpenAI 和 Anthropic 在做这个。我就想,哦……
Yeah. Sorry, man. I just—my whole life I get asked always, you know, OpenAI and Anthropic are doing this. I'm just like, oh...
我给你一个软球,老兄。来吧。
I'm giving you a soft one, man. Let's go.
不。所以……
No. So...
老兄,我现在担心了。这是不是——也许我们应该继续。
Man, now I'm worried. Is this—maybe we should move on.
不,不,不,不,不。这是个好问题。
No, no, no, no, no. It's a great question.
我在逗你。这是个好问题。
I'm trying to tease you. It's a great question.
从根本上说,OpenAI、Anthropic、DeepMind 都专注于构建 AGI 或超级智能。关键是要明白,针对特定任务使用专用模型总是比通用模型更好,即使你拥有超级智能也是如此。原因是通用模型必须在权重中处理所有事情,而专用模型只需处理部分事情。
The way that fundamentally, at the end of the day, OpenAI, Anthropic, DeepMind are focused on building AGI or artificial superintelligence, whatever. The key thing to understand is that having a specialized model for a particular task is always going to be better than having a generalist model for that task, even when you have superintelligence. The reason is that the generalist model has to do everything in the weights, and the specialist model only has to do some things.
等等,那你们有专用模型吗?
Wait, so do you have a specialist model?
当然。我们针对特定的科学任务训练推理模型。
Absolutely. We train reasoning models on specific scientific tasks.
我明白了。
I see.
我们看到,用很少的数据就能在这些特定任务上获得比前沿模型巨大的提升。所以第一件事就是拥有专用的智能体模型,这也带来了专用的用户体验。例如,我们在测试时投入了更多算力来减少幻觉,比其他模型要多。
What we see is that with a very small amount of data, you can get enormous gains for those specific tasks over the frontier models. So that's the first thing: having specialized agent models, which also come with specialized user experiences. For example, we put way more test-time compute into minimizing hallucinations than others would.
但你们大概也会用更大的模型吧?
But presumably you also use the bigger models, right?
是的,当然。我们在某些领域使用更大的模型。比如,我们不想在编程上比 Anthropic 更好。但涉及到非常小众的生物学推理时,我们内部有自己的模型,表现更优。
Yeah, absolutely. We use the bigger models in some areas. For example, we don't want to be better than Anthropic at coding. But when it comes to very niche biological reasoning, we have our own models internally that are superior.
但编程这件事不正好说明了专用模型的问题吗?我是说,这些编程模型其实并不是真正专精于编程,对吧?然而……
Doesn't the coding thing actually show an example of a specialized model? I mean, these coding models are not really specialized in coding, right? And yet...
如果你的问题是,一个专用的编程模型能否击败像 GPT-5.5 这样的模型……
If your question is whether a specialized coding model could beat something like GPT-5.5...
你觉得这似乎不太可能。
You think it seems unlikely.
嗯,我觉得很多人都在尝试,但还没成功。
Well, I think a lot of people are trying to do it, and they haven't.
我觉得他们还没做到。
I think they haven't yet.
编程非常特殊,因为所有实验室都意识到编程是他们必须非常擅长的领域。所以可能编程并不适用这个规律。但我可以保证,如果你想要一个专门推理合成化学路径的模型,它会比通用模型做得更好。这个论点可能失效的地方是当达到智能饱和时。例如,任何超级智能在玩井字棋时都会和人类一样好,因为人类在井字棋上已经是最优的。类似地,在某个点上,智能会在化学推理上饱和。我不知道那个点在哪里,但有可能。到那时,我的论点可能就不成立了。但我认为我们离那个点还很远。
Coding is very special because all the labs have realized that coding is the thing they need to be really good at. So maybe it's not true in coding for that reason. But I can guarantee that if you want a specialized model for reasoning about synthetic chemistry pathways, it will do better than a generalist model. The place where this may fall down is when you get to intelligence saturation. For example, any superintelligence will be exactly as good as humans at playing tic-tac-toe, because humans are optimal at tic-tac-toe. Similarly, at some point intelligence will saturate for chemistry reasoning. I don't know where that point is, but it's feasible. At that point, my argument might fall down. But I think we're very far away from that point.
有些人类……
Some humans...
有些人类,不是所有人类,但确实存在在井字棋上最优的人类。
Some humans, not all humans, but there exist humans who are optimal at tic-tac-toe.
第二件事是,所有制药公司都有自己的内部数据集,它们需要彼此之间的专有优势。它们需要基于自身数据训练的模型,因为如果每个人都拥有相同的智能,那么就没有人在研发上拥有优势。这正是我们大放异彩的领域。所有制药公司都争着与我们合作,我们被需求淹没,原因就是我们拥有最好的科学模型。我们做最后一英里的集成,让这些模型在内部部署并真正加速管线,同时我们在它们的数据上训练模型,这为它们提供了可持续的优势。至少在今天,OpenAI 和 Anthropic 不想这样做,因为它们不想为每个客户提供不同的模型。
The second thing is that all pharma companies have their own internal datasets, and they need proprietary advantages over each other. They need models trained on their data because if everyone has the same intelligence, no one has an advantage in R&D. Those are the areas where we really shine. The reason why all pharma companies are fighting to do a deal with us, and we're overwhelmed with demand, is because we have the best models for science. We do the last-mile integration to get those models deployed internally and accelerate the pipeline, and we train on their data, which provides them with a sustainable advantage. At least today, OpenAI and Anthropic don't want to do that because they don't want to have a different model for every customer.
这挺有意思的。我感觉六个月或一年前和你聊的时候,你更多在讲智能体框架和你做的评估。情况变了吗,还是说这个话题没那么令人兴奋了?
It's kind of interesting. I feel like when I talked to you six months or a year ago, you talked more about the agent framework and the eval you were doing. Have things changed, or is that less exciting to talk about?
不,我认为我们已经成熟了,或者说我们正在成熟,对我们所做的事情和价值主张的思考也在成熟。但事实仍然是,我们的智能体在科学领域是最好的,并且比 Anthropic、OpenAI 和 DeepMind 的产品领先很多。我认为我们会在一段时间内保持这个优势。但当你思考三、四、五年后市场的宏观动态时,有几个关键点。第一,这些公司需要基于自身数据训练的模型来保持优势。第二,它们不想被锁定在单一的模型提供商上。它们不想只绑定 Anthropic 或只绑定 OpenAI,因为去年 OpenAI 遥遥领先,而今天 Anthropic 感觉遥遥领先。这些动态促使客户与我们合作。
No, I think we have matured, or we are in the process of maturing our thinking about what we do and the value proposition. But it remains the case that our agents are the best at science and are substantially further ahead than what Anthropic, OpenAI, and DeepMind have. I think we'll maintain that edge for a while. But when you think about the macro dynamics of what the market will look like in three, four, or five years, there are a couple key things. First, these companies need models trained on their data to maintain their advantage. Second, they don't want to be locked into a single model provider. They don't want to be locked into just Anthropic or just OpenAI because last year OpenAI was way ahead, and today Anthropic feels way ahead. Those dynamics push customers to work with us.
当你观察主要的模型提供商时,你认为它们之间的差异是否足够大,以至于针对特定用例使用它们是有价值的,还是它们本质上可以互换?
Do you think when you look at the major model providers, they're different enough that using them for specialized use cases is valuable, or are they essentially interchangeable?
这是个好问题。我认为目前它们各有专长,而且很多专长是不重叠的。
This is a great question. I think at the moment they are spiky. There's a lot of non-overlapping spikes.
哦,有意思。给我举个例子。
Oh, interesting. Give me one example.
嗯,是的,很好。比如,我们正在构建一个基准测试,很快就会发布,它涉及复现之前科学论文中的分析。至少在我们录制时的评估中——可能会变,对吧——但根据我上次看到的数据,Anthropic 在这方面特别擅长,明显优于 OpenAI,也明显优于 DeepMind。另一个更好的例子,更能说明问题:在 Cosmos 开发的早期,有一个特定任务我们需要模型来完成,用于更新 Cosmos 的世界模型。当时,我们唯一能找到能完成这个任务的模型是 Gemini 2.5。其他模型——我们完全搞不定如何让它们做这个任务,这太疯狂了。所以这种尖峰特性很多,对吧?它们显然有不同的个性,这也很有趣。但确实,我们惊讶地发现,通过结合不同提供商的模型,你可以在任务上获得更好的结果。
Um yeah, great. For example, we have a benchmark that we're building that we will release shortly that involves reproducing analyses that have been done in scientific papers previously. And at least on our evaluations as of when we're recording this — it may change, right — but at least as of the last data I saw, Anthropic is particularly good at that, significantly more so than OpenAI, significantly more so than DeepMind. Another good example, which actually makes the point to a greater extent: there was a time early in the development of Cosmos where there was a specific task that we needed the models to do in the process of updating Cosmos' world model. At the time, the only model that we could find that was able to do that task was Gemini 2.5. None of the other models — we could not figure out how to get any of the other models to do that task, which was wild. So there's a lot of that spikiness, right? They have different slightly different personalities obviously, so that's also interesting. But yeah, we've been surprised at the extent to which you can get much better results on tasks by combining models across different providers.
我明白了。现在当我启动 Cosmos 时,它会运行多久?大概会调用多少次?
I see. When I fire off Cosmos today, how long does it run for? Like how many calls is it making?
好问题。我们在去年 11 月发布了 Cosmos 的原始版本。它会运行 6 到 12 小时,写 45,000 行代码,单次运行阅读 1500 篇论文,比任何人见过的任何智能体都要强大得多。我认为它至今仍是市面上最强大、最密集的智能体之一。问题在于用户体验:如果你是一位正在做研究的科学家,你并不想只是问一个问题,然后走开,6 到 12 小时后再回来。所以我们最近宣布了对 Cosmos 的重大更新,现在它有了一个交互式前端,仍然可以执行那些超长任务,但你可以全程给予反馈。你可以引导它。它更具交互性。
Great question. We released the original version of Cosmos back in November. That would run for 6 to 12 hours, write 45,000 lines of code, read 1500 papers in a single run, and was just way more powerful than any agent that anyone had seen. It is still today, I think, one of the most powerful, most intensive agents out there. The issue with it was a UX issue: if you are a scientist doing research, you don't really want to just ask a question, then walk away and come back 6 to 12 hours later. So we have just recently announced a significant update to Cosmos that now puts a kind of interactive front end on it, so it can still go and do those extremely long-running tasks, but you can give it feedback throughout. You can steer it. It's more interactive.
我们来快速问答环节吧。
Let's go to rapid fire random questions here.
好,来吧。
Yeah, let's do it.
好吧。你说过——我看到你说科学进步已经大大放缓了。这和我对科学进步的看法不同。你是什么意思?
All right. You said — I saw that you said scientific progress has slowed down a lot. That was different than how I think about scientific progress. What did you mean by that?
嗯,这取决于你指的是我哪次说的。
Well, that depends on which time I said this that you're talking about.
你还坚持我断章取义的那个观点吗?
Do you still stand by that point that I pulled out of context?
从经验上看,医学领域的进步确实放缓了。
It is empirically the case that progress in medicine has slowed down.
从经验上看是什么意思?
What does it mean empirically?
在医学领域,我们有一个所谓的 Eroom 定律——也就是 Moore 定律倒过来拼写。它观察到,开发一种新药的实际成本在过去 40 年里翻了一番。我忘了翻倍的时间;可能是每九年左右。所以半导体价格呈指数级下降,而药品价格呈指数级上升。人们提出了各种原因——实际上这个趋势在 2011 或 2012 年左右被打破了。人们认为很大程度上是因为人类基因组计划。这让寻找药物变得更容易了。但现在基本持平,生产力肯定还没有提高。
In medicine, we have what is called Eroom's Law — which is Moore's Law spelled backwards. It is the observation that the amount of money in real terms that it costs to develop a new drug has doubled over the past 40 years. I forget the doubling time; it might be every nine years or something like that. So prices on semiconductors have had this exponential drop, while prices on drugs have had this exponential increase. There are various reasons why people have argued this might be — it actually broke that trend around 2011 or 2012. People think it's largely because of the Human Genome Project. That made it easier to find drugs. But now it's about flat, and certainly productivity is not yet increasing.
那是因为我们找到了所有最好的药物吗?
Is that because we found all the best drugs?
这是原因之一。一个原因可能是缺乏容易摘取的果实。另一个原因可能是所谓的“比披头士更好”问题:如果你有一种治疗某种疾病——比如结肠癌——的药物,你需要推出一种比那种药物更好的药物,才能获得市场份额,才能被批准,才能有用。
That's one of the reasons. One reason may be a lack of low-hanging fruit. Another reason may be what is called the "Better than the Beatles" problem: if you have a drug that treats condition X — say, colon cancer — you need to come out with a drug that is better than that drug in order to gain market share, for it to get approved, and for it to be useful.
为什么叫它“比披头士更好”?
Why do you call it "Better than the Beatles"?
只是因为有一个观察:披头士仍然非常受欢迎。他们作为乐队没有被取代。他们是第一个。任何后来的乐队——如果披头士不存在,也许会有其他乐队。仍然有好乐队,但他们没有获得像披头士那样的份额,因为披头士占据了那个意识空间。
Just because there's this observation: the Beatles are still extremely popular. They have not been displaced as a band. They were first. Any band that comes afterwards — if the Beatles had not existed, maybe there would be some other band. There are still good bands, but they don't get as much share as the Beatles because the Beatles took that space in the consciousness.
基本上就像所有疾病都被治愈了?
Basically like all diseases have been cured?
不,但问题就在这里:它们没有被治愈。只是我们找不到更好的治疗方法。没有人会争辩说抑郁症已经被治愈,或者结肠癌已经被治愈。但如果你想让你的药物获得批准,你必须比同类最佳更好。
No, but this is the thing: they haven't been cured. It's just that we can't find better treatments for them. No one is going to argue that depression has been cured or that colon cancer is cured. But if you want your drug to get approved, you have to be better than the best-in-class.
但这只是意味着越来越难了,对吧?想出新的药物越来越难,因为你不能像半导体那样建立在之前的创新基础上。
But that just means it's getting harder, right? It's getting harder to come up with new drugs because you don't build on previous innovations in the same way as in semiconductors.
嗯,因为如果你有一种使用特定机制对抗癌症的药物,你需要想出不同的机制。从榨取那个特定机制中只能得到这么多。你可以优化药物,让它稍微好一点,从中获得很多收益,但通常首创药物会带来巨大的改进。
Well, because if you have a drug that uses a particular mechanism to fight cancer, you need to come up with a different mechanism. There's only so much you can get from juicing that specific mechanism. You can optimize the drug, make it marginally better, and there are a lot of gains from that, but often the first-in-class drugs lead to dramatic improvements.
我明白了。有趣。所以就像我们在使用某种新机制。
I see. Interesting. So where it's like there's some new mechanism that we're using.
嗯,是的,所以我认为科学正因为这个原因而进展得更慢。我也认为就像物理学一样,对吧?随着你发现更多的东西,发现新东西就变得更难了。
Um yeah and so I think that science is moving more slowly for that reason. I also just think as in the case of physics, right? Like as you discover more things, it just becomes harder to discover more things.
有意思。
Interesting.
好吧。那么博士学位或正式的理学学位还值得去读吗?
All right. Is a PhD or formal science degree still worth pursuing?
对。哇。嗯,好问题。嗯,我认为答案是肯定的,因为我不认为人类研究人员会很快消失。从根本上说,读博的目的是学习如何做研究,对吧?如果你从未学会如何做研究,那你肯定无法有效地进行研究。
Right. Wow. Yeah. Good question. Um, I think that the answer is yes because I don't think that human researchers are going anywhere anytime soon. And fundamentally, the point of a PhD is to learn how to do research, right? And if you never learn how to do research, you definitely are not going to be effective at doing research.
但你们不是正在自动化研究吗?
But aren't you automating research?
我们是在自动化,或者说肯定在加速它。更有趣的问题是:我们还需要人类科学家吗?
We are, or certainly accelerating it. The more interesting question is: are we going to need human scientists?
是的,这确实是个相关的问题。你怎么看?
Yeah, a related question for sure. What do you think?
嗯,我认为在相当长的一段时间内答案都是肯定的。原因是科学是不可验证的。所以从根本上说,归根结底,我们需要人类科学家的品味。我不确定是否会在 20 年或 5 到 10 年的时间尺度上达到一个点,即模型在品味上明显优于人类。
Um, and I think the answer is yes for a substantial amount of time. And the reason is that science is nonverifiable. So fundamentally, at the end of the day, we need human scientists for their tastes. I'm not sure that it's going to be, maybe in 20 years or on the 5 to 10 year time scale, I'm not sure that we will get to a point where the models will just be obviously better taste than the humans.
嗯,这很有意思,因为这里说你曾说过语言模型最终会在提出想法方面比人类更强。
Well, it's interesting because here it says you said language models will eventually be better than humans at coming up with ideas.
呃,是的,但“更好”存在于许多不同的维度上,对吧?所以更好可能意味着更有可能成功。我认为这绝对是事实,语言模型在提出更可能成功的想法方面远胜于人类。但当涉及到哪些问题最值得追求时,这就更难了,因为它是不可验证的。所以这只是不确定性。我不知道是否会在五年内达到一个点,让我们觉得“哦,没必要问诺贝尔奖得主弗朗西斯·阿诺德她对蛋白质进化的看法了,因为我们直接问模型就行。”我认为这是一个相当高的门槛。今天如果我去问 Claude 3.7 它认为我们应该在定向进化中做什么来改进化学,与问弗朗西斯相比,毫无疑问弗朗西斯会有更好的想法。
Uh yeah, but better exists along many different axes, right? So better can mean more likely to work. I think that's definitely the case, that the language models are way better than humans at coming up with ideas that are more likely to work. But when it comes to which of these problems is the most interesting to pursue, it's just harder because it's not verifiable. So this is just uncertainty. I don't know whether we will get to a point within five years where we're just like, 'Oh, there's no point in asking Francis Arnold, Nobel laureate, what she thinks about evolving proteins because we could just ask the model.' I think it's a pretty tall bar. When I look today at going to ask Claude 3.7 what it thinks we should be doing in directed evolution to improve chemistry versus asking Francis, there's no question that Francis is going to have better ideas.
有意思。今天,对吧?我是说,两年后我们又会坐在这里,你会说:“嗯,你当时是那么说的,但现在……”
Interesting. Today, today right? I mean, you know, in two years we're going to be sitting here again, you're going to be like, 'Well, you have said that and now...'
在这个不公开的播客里。
In this off-the-record podcast.
好吧。那么,你筹集了相当多的资金。我笔记里记的是 7000 万。这个数字准确吗?
All right. So, you've raised quite a lot of money. I think I have 70 million in my notes. Is that even accurate?
是的,没错。7000 万。而且那只是营利部分;非营利部分筹得更多。我认为与我遇到的大多数实验室和公司不同,你的投资者实际上大多来自制药界。你觉得这是为什么?
Yeah, that's right. 70 million. And that was just the for-profit; the nonprofit raised more. And I think unlike most of the labs and companies that I come across, most of your investors are actually folks coming from the pharma world. Why do you think that is?
嗯,对我们来说,成功就是进入制药公司内部。而制药投资者正是拥有这种能力的人。所以我们从所有投资者那里获得了非常高的价值。Yasmin Razavi 是 Spark 的投资者,共同领投了我们的轮次,她非常出色。她是 Anthropic 的董事会成员,领投了他们的首轮 VC,拥有非凡的视角和洞察力,并在人才方面给了我们很大帮助。但说到具体的业务 traction,贡献最大的投资者是 Jeff Huber,他是 Grail 的前 CEO,似乎认识制药界的每一个人;另外还有一位未具名的机构生物技术投资者,规模非常大,如果我在这个不公开的场合说出他们的名字,他们会不高兴的。但我认为生物技术领域的大多数人都知道这意味着什么,因为他们非常有名。他们在帮助我们建立与公司的高层联系以及推动部署方面极其有价值。
Um well, for us, success is getting inside of the pharma companies. And the pharma investors are the ones who have that capability. So we get very high value out of all of our investors. Yasmin Razavi is an investor at Spark who co-led our round, and she is absolutely incredible. She's on the board of Anthropic, she led their first VC round, and has extraordinary perspective and insight, and has really helped us with talent. But when I think about concrete business traction, the value from investors who have contributed the most: Jeff Huber, who was the CEO of Grail and somehow seems to know literally every single person in pharma; and then we have an unnamed institutional biotech investor, very large, who would be displeased if I said who they are in this off-the-record forum. But I think most people in biotech will know what that means because they are very well-known. They have been extremely valuable in setting us up with high-level connections into companies and getting us deployed.
这很有意思。所以,真正更了解你领域的人,比那些在硅谷做 AI 投资的狂热分子更看好你。
It's interesting. So, the people who actually know your field better are more bullish on you than the maniacs doing AI investment in Silicon Valley.
是的。我是说,我认为那些狂热分子对很多事情都很兴奋,对吧?他们对“我们要自动化科学”或“让我们自动化药物发现”这样的概念非常兴奋。但业内人士知道问题所在,所以他们真正切身知道价值在哪里。这么说吧:从来没有业内人士问我我们与 Entropic、OpenAI 和 Google 有何不同。每个科技投资者都会问我这个问题:“为什么 Anthropic、OpenAI 和 Google 不做我们正在做的事?”但任何在制药公司内部运营过或在制药公司和生物技术公司持有大量头寸的人,我们从未从他们那里听到这个问题。
Yeah. I mean, I think that the maniacs are very excited about many things, right? They are very excited about the concept of 'we're going to automate science' or 'let's go automate drug discovery.' But the insiders know the problems, so they really viscerally know where the value is. Let me put it this way: I don't get any of the insiders asking me how we're differentiated versus Entropic and OpenAI and Google. I get that asked every single time by tech investors. 'Why won't Anthropic and OpenAI and Google do what we're doing?' But anyone who has operated inside a pharma company or has large positions in pharma companies and biotechs, we never get that question from them.
对。好吧,AI 在药物发现方面有过很长一段过度承诺的历史。
Right. Okay, there's a long history of AI overpromising in drug discovery.
哦,是的。
Oh yeah.
这次真的不一样吗?
Is this time different?
是的。我保证,我向你保证这次真的不一样。回到 2012 年左右,人们一直在说 AI 将彻底改变制药行业,但结果却是一次又一次的失败,唯一的例外是 AlphaFold。AlphaFold 确实改变了很多事情在药物发现和开发的一个领域中的运作方式,也就是实际提出分子的过程,因为你有了蛋白质的结构。但我认为存在大量的怀疑。话虽如此,认为这次不会不一样其实有点疯狂。我们已经有这么多证据表明这次真的会不一样。但我想强调的一点是,事实胜于雄辩,而事实就是获批的药物。所以在未来几年你会听到很多说法,比如“我的模型提出了这种药物”、“我的模型发现了生物学的这个基本方面”等等。而任何在制药行业待过的人都不会在意,直到他们看到关键 3 期临床试验的结果或 FDA 的批准信。
Yes. I promise, I promise you this time it is different. Yeah man. Back to like 2012, people have been saying that pharma AI is going to revolutionize pharma, and it's been one flop after another, with the notable exception of AlphaFold. AlphaFold really changed the way that a lot of things are done in one area of drug discovery and development, which is the actual process of coming up with the molecule because you have the structure of the proteins. But I think there's a ton of skepticism. That said, it is kind of crazy to think that this time will not be different. Like we have so much evidence already that this time it's really going to be different. But I think the thing I do want to emphasize is the proof is in the pudding, and the pudding is approved drugs. So you're going to hear a lot of stuff in the next couple of years, like 'my model came up with this drug' and 'my model discovered this fundamental aspect of biology' and so on. And anyone who has been in pharma is not going to care until they see the outcome of a pivotal phase 3 clinical trial or an approval letter from the FDA.
这似乎是一个很好的结束点。是的。谢谢,Dan。酷。谢谢。太棒了。非常感谢收听本期 Gradient Descent。请继续关注未来的节目。
That seems like a great place to end. Yeah. Thanks, Dan. Cool. Thanks. That's great. Thanks so much for listening to this episode of Gradient Descent. Please stay tuned for future episodes.