AlphaFold:从世纪难题到诺贝尔奖

AlphaFold: From Grand Challenge to Nobel Prize

约翰·江珀 John Jumper · Google DeepMind · 2025-11-28 · 约 48 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

约翰·詹珀回忆因 AlphaFold 获得诺贝尔奖的经历,从焦虑等待到庆祝的起泡酒,并反思他从物理博士辍学到领导生物学最伟大突破之一的非传统道路。

John Jumper recounts winning the Nobel Prize for AlphaFold, from the anxious wait to the celebratory sparkling wine, and reflects on his unconventional path from dropping out of a physics PhD to leading one of biology's greatest breakthroughs.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 15)

全文 · Full transcript(中英对照)

AlphaFold 的影响与诺奖 AlphaFold's impact and Nobel Prize announcement

Host

我在推特上看到有人评论说:‘他们怎么拿到我的结构的?DeepMind 怎么拿到我做了但还没发表的东西?’他们简直不敢相信,这居然是一台机器瞬间完成了多年艰苦的工作。所以,你从一个宽泛的假设——精子表面有某种蛋白质起这个作用——开始,AlphaFold 说‘我觉得是这一个’,然后你去做详细的实验来确认。现在你可以思考不孕不育之类的问题了。如果你在那个蛋白质中看到突变,那可能就是不孕的原因。也许我们可以考虑治疗它。欢迎收听 Google DeepMind 播客。我是 Hannah Fry 教授。今天我们谈论的是 AlphaFold,现代科学中最非凡的技术突破之一。这个工具被描述为 AI 做过的最有用的事情。说实话,这可能是低估了。这是一个 Google DeepMind 的 AI 系统,它解决了生物学最大的挑战之一——预测蛋白质(生命的基本构建块)的 3D 结构。其最新版本 AlphaFold 3 现在可以以前所未有的精度模拟所有生命分子的结构和相互作用。其影响是巨大的。AlphaFold 已经绘制了数亿个蛋白质结构,来自 190 个国家的超过 300 万研究人员现在使用它的数据库。它正在改变药物发现。2024 年,诺贝尔化学奖授予了 Google DeepMind 的 Demis Hassabis 和 John Jumper,他是我们今天的播客嘉宾。这个故事我们从第一季就开始关注了,那大约是八年前,远在它成为头条新闻之前。所以,如果你是第一次接触 AlphaFold,想知道它为什么这么受关注,你可以在描述中找到我们之前的解释性剧集链接。欢迎来到播客,John。

I saw this comment from someone on Twitter saying, 'How did they get a copy of my structure? How did DeepMind get this thing that I had done and not yet published?' Like they couldn't believe that this was literally a machine doing years of painstaking work all at once in a flash. So you go from this broad hypothesis that there's some protein on the surface of sperm that does this. AlphaFold says I think it's this one. And then you go do your detailed experiments to confirm. And now you can think about questions like infertility. If you see mutations in that protein, maybe that's a cause of infertility. Maybe we can think about treating that. Welcome to Google DeepMind the podcast. I'm Professor Hannah Fry. Today we are talking about AlphaFold, one of the most extraordinary technological breakthroughs in modern science. A tool that has been described as the most useful thing that AI has ever done. And in truth, that might be an understatement. This is a Google DeepMind AI system that solves one of biology's grandest challenges, predicting the 3D structures of proteins, the fundamental building blocks of life. Its latest version, AlphaFold 3, can now model the structure and interactions of all of life's molecules with unprecedented accuracy. And the impact has been seismic. AlphaFold has mapped hundreds of millions of protein structures and more than 3 million researchers across 190 countries now use its database. It is transforming drug discovery. And in 2024, the Nobel Prize in Chemistry was awarded to Google DeepMind's Demis Hassabis and John Jumper, who is our guest on today's podcast. This is a story that we have been following on this podcast since season 1, which was nearly eight years ago, long before it hit the headlines. So if you are coming to AlphaFold for the first time and wondering what all the fuss is about, you can find our previous explainer episodes linked in the description. Welcome to the podcast, John.

John

哦,真令人兴奋。

Oh, it's exciting.

Host

我想自从你获得诺贝尔奖以来,我还没采访过你。告诉我,你得知这个消息时在哪里?

I don't think I've interviewed you since you won your Nobel Prize. Tell me, where were you when you found out about it?

John

我待在家里,因为我已经够紧张了。我觉得有十分之一的机会。所以我想,如果在家失望的话,我就坐在床上。我原本计划是睡过去,然后一个电话叫醒我,我就得了诺贝尔奖,但我睡不着,因为你知道电话可能打来的那天,对吧?

I stayed home because I was nervous enough. I thought there was a chance, like one in ten a chance. And so I figured I would be disappointed at home and I was kind of just sitting in the bed. My original plan was I'll sleep through it and a phone call wakes me up then I've got the Nobel but I couldn't sleep because you knew the day that the phone call might happen, right?

Host

你知道那天。事实上,我知道它计划在 11 点公布。我知道获奖者大约提前一小时接到电话。

You knew the day. I knew in fact the kind of time that it was scheduled to be announced at 11:00. I knew that winners were called about an hour beforehand.

John

所以到了大约 10:30,我说:‘哦,好吧,我想今年不是了。’我告诉我妻子,她说:‘不,不,等等。’就在她让我等的时候,我的手机亮了,显示来自瑞典的电话。幸好这不是世界上最恶毒的恶作剧电话。是的,这真是一件非凡的事情。你接起电话,他们说:‘John Jumper 博士在吗?’‘是的。’‘我有一些好消息。’‘太好了。’‘请稍等。’然后他们,我想你等着。嗯,我想他们当时在尝试,部分问题在于他们一开始没有 Demis 或我的电话号码。所以无论如何,他们最终很晚才给我们打电话,但后来他们终于安排好了,他们让那个人接电话,他说:‘我有一些改变人生的消息。’他们在 60 到 90 秒内没有说‘诺贝尔’这个词,那是我生命中最长的一分钟,因为这次通话没有其他解释。我记得我做的第一件事就是跑去洗澡,因为我知道接下来一整天都没时间了。但之后,你知道,消息公布了,你进来,看到团队,我们进行了精彩的庆祝。我们把当地女招待的所有起泡酒都买光了。

So by about 10:30 I said, 'Oh well, I guess not this year.' And I told my wife and she goes, 'No, no, wait.' And as she's telling me to wait, my phone lights up with a phone call from Sweden. And thankfully it was not the world's meanest prank call. And yeah, it was just kind of this extraordinary thing. And you answer and they say, 'Is Dr. John Jumper available?' Yes, I have some wonderful news. Great. Can you please hold? And right, so they get, I think you hold. Well, I think they were trying to, part of the problem was they didn't have either Demis or my phone number initially. So anyway, they ended up calling us very late, but then they were finally arranging and they pulled the person on and he says, you know, I have something like I have some life-changing news and they don't say the word Nobel for like 60 to 90 seconds, which was the longest minute of my life as there's no other explanation for this call. And I remember the very first thing I did is run to get a shower because I knew I was going to get no time for the rest of the day. But after that, you know, it was announced, you come in, you see the team, you have this amazing kind of celebration. We bought the local waitress out of sparkling wine. Um,

Host

只有最好的才行。

Only the best will do.

John

我不是品酒专家,我们和朋友一起庆祝,整栋楼都充满了不可思议的派对气氛。太棒了。

It was, I'm not a connoisseur and we were celebrating with friends and there was just this incredible kind of party just across the floors of our building. It was amazing.

Host

问题是,这是一个关于你个人的非凡故事,对吧?因为你的第一个博士,你的物理学博士,你退学了,对吧?

The thing is, it's an extraordinary story of you as an individual, right? Because your first PhD, your physics PhD, you dropped out, right?

John

是的。是的。

Yeah. Yeah.

Host

所以从那个(我想一定是很艰难的经历)到成为诺贝尔奖得主,你的工具被用于成千上万的学术论文。

And so going from that, which I think must have been quite a hard experience to live through, to being a Nobel Prize winner and having your tool being used in tens of thousands of academic papers.

John

我的意思是,我要说退学对我来说是一件非常幸运的事。我当时在做错误的事情。我并不真的想做,所以我就离开了。因为我离开了,我实际上进入了一个计算生物学小组,他们在用定制计算机芯片模拟蛋白质方面做了很棒的工作。然后我回去,现在通过另一系列意外读了一个化学博士。我没有那些强大的计算机。所以为什么不进入 AI 呢?为什么不尝试使用复杂的算法来弥补算力的不足?我一定是第一个因为缺乏计算能力而不是因为丰富计算能力而进入 AI 的人。然后我很幸运地找到了一份工作,与我过去尝试做的一切都有关系,然后它成功了,我获得了诺贝尔奖。

I mean, I will say dropping out was a very lucky thing for me. I was doing the wrong thing. I didn't really want to and so I just left. And because I left, I actually fell into this computational biology group that was doing amazing work on custom computer chips to simulate proteins. And then I go back and I do my PhD now in chemistry by another set of accidents. And I didn't have those great computers. So why not get into AI? Why not try and use sophisticated algorithms to make up for a lack of compute? I have to be the first person to get into AI because of a lack of computational capability rather than an abundance. And then I got lucky enough to kind of find a job that had something to do with everything I'd ever tried to do in my past and it then it worked out and I get a Nobel.

Host

现在人们对你的反应和以前不同吗?

Do people react differently to you now then?

John

哦,我的意思是,嗯,我认为有各种各样的人。有和我一起读化学博士的人,他们知道我是一个相当好的物理学家和一个糟糕的化学家。有每天和我一起工作的人,我想我仍然是 John,但现在是有诺贝尔奖的 John,所以他很忙。但还有我遇到的所有其他人。我的意思是,我会接到电话,令人惊讶的是,很多电话都以‘很荣幸能与您交谈’开头,我有时想‘我也是’。有一种差异,或者至少是兴奋,它是这个巨大 AI 世界的象征,以及它对于应用 AI 解决现实世界问题意味着什么。然后你是诺贝尔奖得主。所以你可以对任何事情发表意见,而且即使不准确,也被认为有点道理。所以人们希望你出现在各种场合,只是为了象征诺贝尔奖被赢得了,然后你就完成了,这作为一个科学家来说并不是一件非常令人满意的事情,但你拥有这个平台,也许你可以用它来影响公众对科学的看法,以及科学资金的投入。所以所有这些事情都混合在一起,形成了一种疯狂的组合。我想,就本科毕业到大致退休的时间而言,我大约处于职业生涯的中点。

Oh, I mean, well, I think there's all sorts of people. There are the people that I did my chemistry PhD with who knew me as a pretty good physicist and a lousy chemist. There are the people that I work with every day and I'm still, I think, just John but now John with a Nobel so he's busy. But then there are all the people I meet. I mean, I would get on phone calls and a surprising number of my phone calls start with 'It's such an honor to speak with you' and I sometimes think 'and also with you.' There's a certain type of difference or at least excitement and it's a symbol of this giant AI world and what it can mean in terms of applying AI to solve real world problems. And then you're a Nobel Prize winner. So you're allowed to have an opinion on anything and it's supposed to be a bit valid even if it's not. And so people want you to show up to things just so that you can symbolize that a Nobel Prize was won and then you're done, which is not a very satisfying thing to do as a scientist, but you have this platform that maybe you can use to affect how the public thinks about science, how it funds science. So all of these things kind of roll together in this wild combination. I'm, I would say, roughly at the midpoint of my career in terms of time since undergrad and time until rough retirement.

AlphaFold 的影响与意义 AlphaFold's impact and significance

Host

所以我得想想下半场该做什么,这确实是个不小的挑战。

And so I've got to figure out what to do in the second half and that's a fair amount to live up to.

Host

是的,绝对如此。那件事压力很大。我的意思是,距离 AlphaFold 2 在 CASP 上突破并击败预测挑战才过去 5 年。你当时意识到这项工作的潜在重要性了吗?

Yeah, absolutely. A lot of pressure going on in that thing. I mean we're still only 5 years on from that CASP breakthrough really when AlphaFold 2 smashed the prediction challenge. Did you realize at the time the potential significance of the work that you were doing?

John

我们确信两件事,对其他一些事则完全不确定。我认为我们相当确信的两件事是:我们非常确定它有效,甚至在进入 CASP 之前。我们测量得很好,知道我们在 CASP 中的表现,我们理解并很谨慎。我们知道我们已经解决了这个重大挑战。但科学中重大挑战的常规想法是,你解决了它,会有盛大的庆祝,然后你会去构建利用这些想法解决重大挑战的有效实用系统,这标志着一个时代的开始。我认为真正让我震惊的是,我们训练的那些权重、那个系统、那个计算机软件,至今对从事该领域的科学家来说具有如此巨大的实际重要性。那个软件本身被用于所有这些不同的应用领域,所有这些不同类型的科学都基于这个黑箱计算机程序发表。它进入科学实践的程度,我认为真的超出了我的想象。

We were sure of two things and totally unsure of some others. I think the two things we were pretty sure of: we were very sure it worked even before we entered CASP. We had measured very well, we knew about how we would do in CASP, that we understood and we were careful. We knew that we had solved this grand challenge. But the normal thought in a grand challenge in science is that you'll solve it and there'll be a great celebration and then you will go build effective useful systems that use the ideas that enabled you to solve the grand challenge, and that this was kind of the beginning of an era. I think the real shock to me is, you know, those weights that we train, that system, that piece of computer software has been so incredibly practically important to scientists working in this field to this day. That the actual bit of software is used that makes this difference in all these different application areas, all this different type of science published on top of this as a blackbox computer program. And the extent to which that has entered into scientific practice has been really, I think, beyond my imagination.

Host

是的,我的意思是,很难夸大这带来的真正意义。我看到有说法称 AlphaFold 是 AI 做过的最有用的事情,对吧?公众似乎还没有完全意识到这一点,是吗?

Yeah, I mean it's really difficult to overstate the genuine significance that this has had. I saw one thing where AlphaFold was described as the most useful thing that AI has ever done, right? That hasn't sort of landed with the public yet, has it?

John

我认为人们很难理解,而你从事科学传播工作,知道科学有多难,治愈疾病有多难。我们必须非常努力才能获得关于细胞如何工作、身体如何运作的零碎知识。对于蛋白质结构预测或蛋白质结构,也就是 AlphaFold 所做的,我认为人们很难理解这个过程,它需要在实验室里花一年时间。我见过一些博士论文,它们只是朝着确定 X 的结构取得进展,并不意味着完成了,只是他们觉得自己更接近了一点,需要毕业。我指的是一个蛋白质,一个单独的片段。而我们将这项工作变成一台机器,在 5 分钟内给出非常好的答案,然后推动下游更多的工作。我记得大概有,我没看最新数字,但 3 万到 3.5 万篇科学论文引用了 AlphaFold。3.5 万项对我们理解生物学的贡献都建立在这一进步之上。我认为思考 AlphaFold 的正确方式不是我们解决了生物学中的所有问题,我们远未做到。我认为对于关心细胞结构的那部分生物学,即结构生物学,也许我们整体上使其加快了 10%。我们放大了这一巨大的社会投入,最终将带来变革性的科学。在某些狭窄的领域,比如蛋白质设计,正被这种理解所改变。

I think it's hard for people to appreciate, and you work in science communication, how very hard science is, how very hard curing disease is. We have to work extremely hard to get smaller bits of knowledge about how say the cell works, how the body works. For protein structure prediction or protein structure, what AlphaFold does. I think it's hard for people to appreciate this process, you know, takes a year in the lab. I've seen PhD theses that are progress toward determining the structure of X. And that doesn't mean they finished it, just they feel like they're a little bit closer and they need to graduate. I think that of one protein, of one individual piece. And the notion that we'll turn that work into a machine that gives you a really good answer in 5 minutes and then that enables so much more work downstream of it. I think there's something like, I haven't looked at the recent number but 30, 35,000 different scientific papers that cite AlphaFold. 35,000 different contributions to our understanding of biology that build on top of this advance. And I think the right kind of way to think of AlphaFold is not certainly that we've solved all problems in biology, we very much haven't. I think for this slice of biology that cares about what structures in the cell look like, structural biology, maybe we've made it 10% faster overall across the whole thing. We've amplified this enormous effort in societal expense and then ultimately we will have transformative science. And there's certain narrow areas, say protein design, that are just being transformed by this understanding.

Host

我认为对我来说,真正证明这一突破重要性的方式之一是,当你们发布 2 亿个蛋白质结构时,生物学家的反应。请告诉我一点,因为你们就是直接发布了,对吧?

I think one of the ways for me that really demonstrates just how important this was as a breakthrough was the way that biologists reacted when you published the 200 million protein structures. Just tell me a little bit about that because you just put it out there, right?

John

哦,是的。最初的发布规模小一些,但我想也有 40 万个。我记得在我们发布代码后大约一周,真正的专家们开始使用它,他们说,‘这真的能解决难题。’但其他生物学家说,‘不,这些不可能是我研究的那些真正的难题。’然后当然我们发布了那个巨大的数据库,AlphaFold 数据库,人们说,‘好吧,让我看看这个 AI 引擎有多蠢,点开他们感兴趣的蛋白质,’我想他们期待嘲笑它。然后他们坐在那里,惊讶不已。我看到推特上有人评论说,‘他们怎么拿到我的结构的?他们怎么,DeepMind 怎么拿到我做了但还没发表的东西?’他们简直不敢相信这真的是一台机器瞬间完成了多年艰苦的工作。同样令人惊叹的是,社区迅速形成了对 AlphaFold 能做什么和不能做什么的理解,例如对单个氨基酸变化不敏感,以及如何将其融入工作流程并在此基础上开展工作。我原以为人们需要几年才能真正弄清楚如何正确融入。如何确保我查看置信度指标,AlphaFold 说这个看起来可靠,那个不可靠?而这在几个月内就发生了。科学社区发展出了这种极其迅速、虽然不是完全完美但实际上非常好的理解,非常非常快。所以人们在 AlphaFold 的基础上做出了优秀的科学成果。我们大概在 7 月左右发布了它,人们几乎立即就在上面做出了非常优秀的科学。到那年年底,非常酷的工作就出现了。我认为这证明了科学家们多么渴望真正有效的工具来推动知识进步,一旦找到,他们就会使用。

Oh yeah. So the original release was a bit smaller but was still I think it was like 400,000. What I remember was there was maybe a week in between when we had put out our code and the real experts were playing with it and they were like, 'This really works on hard problems.' But all the other biologists were like, 'No, these can't be real hard problems like I work on.' And then of course we put out this huge database, AlphaFold database, and people are like, 'Well, let me just see how dumb the AI engine was and click on their protein of interest,' I think expecting to make fun of it. And then they sat there and they were amazed. I saw this comment from someone on Twitter saying, 'How did they get a copy of my structure? How did they get, how did DeepMind get my this thing that I had done and not yet published?' Like they couldn't believe that this was literally a machine doing years of painstaking work all at once in a flash. What was also amazing about it is how rapidly this turned into a community understanding of what AlphaFold does and what it doesn't do, for example not sensitive to single amino acid changes, and how to build this into their workflow and do work on top of it. I thought it would take years for people to really figure out what's the right way to build it in. How do I make sure that I look at the confidence measures where AlphaFold is saying this looks like a reliable answer, this doesn't? And this happened within a matter of months. That science community developed this incredibly rapid, not totally perfect but actually really good understanding very very fast. And so people were doing excellent science on top of AlphaFold. We released this in I think July or something like that and people were almost immediately doing really excellent science on top of it. That really cool work was coming out by the end of that year. And I think that's just a testament to how much scientists are looking for really effective tools that help them push knowledge forward and then when they find it, they use it.

Host

那么,你现在如何跟进人们用 AlphaFold 所做的工作?我的意思是,因为它已经深深嵌入人们做生物学的方式,大概他们不会每次都给你发邮件吧。

Well, how do you stay on top of the work that people are doing with AlphaFold now? I mean, because it's so embedded in the way that people do biology now that presumably they're not sending you an email every time.

John

谢天谢地。我实际上还是经常在 X 或其他平台上搜索‘AlphaFold’这个词,看看随机的成果。我喜欢那些突然冒出来的东西,比如‘哦,我们用这个做了件奇怪的事。’另一种方式,在公司的一个好处是,如果有很酷的事情发生,会有人注意到并发布到我们的聊天室里,收集各种人们觉得酷的东西。所以收集这些经验非常有价值,也很有趣,对吧?你会对那部分工作有一种间接的拥有感。

Thank goodness. I actually end up still pretty often just put the word AlphaFold into searching on whatever X or and just see the random work. What I love is the random things that pop up that say, 'Oh, we use this to do this weird thing.' I think the other way, one of the nice things about being at a company is if really cool things happen, someone will notice and they'll post it to the chat room that we have to collect various things that people find cool. And so it's so valuable to collect that experience and it's so much fun, right? You feel this kind of vicarious ownership of a little bit of that work.

AlphaFold 的非常规用途 Unusual uses of AlphaFold

Host

那么告诉我,你见过的最随机、最不寻常的 AlphaFold 用途是什么?

Tell me then, what is the most random unusual use of AlphaFold that you've come across?

John

我非常喜欢的一个是关于大黄蜂中的一种蛋白质,他们试图了解大黄蜂的种群、繁殖,以及它们的生物学特性,以便促进授粉并理解蜂群崩溃等问题。事实上,他们用 AlphaFold 研究了一些参与蜜蜂生命周期的重要蛋白质。最终你可以看到这如何导向保护工作。我觉得非常有趣的是,结构生物学如何回响到我们关心的所有事情,从食物到工业生产再到其他一切。它们都是相连的,因为都是同样的生物学,对吧?植物、动物,它们基本上有相同的蛋白质。我们最初主要考虑的是人类健康,并没有想如何帮助蜜蜂种群,但结果就是这样。

One that I really love is this protein in bumblebees and they're trying to understand bumblebee populations, you know, reproduction, but their biology to try and, you know, enable pollination and understand things like colony collapse. And so, in fact, there were some important proteins involved in the honeybee life cycle that they were studying with AlphaFold. And you can see ultimately how this kind of leads on to be conservation. And I think it's so interesting to see the structural biology of this echoing into all these things that we care about from food to industrial production to everything else. They're all connected because it's all the same biology, right? You know plants, animals, they have basically the same proteins. We were definitely thinking most about human health. We weren't thinking how am I going to help you know honeybee populations and but here we are.

John

还有一个很好的故事,人们试图理解人类受精过程,即卵子和精子相遇、结合并最终融合。

There was another really nice story that people were trying to understand human fertilization when an egg and a sperm meet and come together and eventually fuse.

Host

他们想找到参与精子附着到卵子的确切蛋白质,对吧?

They want to find the exact proteins that were involved in sperm sticking to eggs, right?

John

精子附着到卵子。我认为他们当时已经掌握了所有卵子蛋白质的全貌,但并非所有精子蛋白质。事实上,有两个独立的研究小组做了这件事。他们说:“嗯,我们只知道精子表面有大约 2000 种蛋白质。为什么不直接尝试所有蛋白质,看看哪些能附着到我们已知的卵子蛋白质上?”如果你考虑用实验方法来做,那就像是要花上两千年,每年 10 万美元,然后才能发一篇《自然》论文——这不是可行的方法。但 AlphaFold 非常快,而且他们有一些可用的计算机。所以他们尝试了所有蛋白质,两个小组都得出了同一种蛋白质,TM 什么的。我记不清数字了,但这是他们之前不知道功能的蛋白质,现在发现 AlphaFold 说它能附着到卵子上,这就是受精的第一步。当然,他们不会仅仅相信 AlphaFold,对吧?它是一个计算系统。

And sperm sticking to egg. And they I think had the full picture of all the egg proteins, but not of all the sperm proteins. And in fact, there were two independent groups that did this. And they said, "Well, there are only 2,000 proteins that we know that are on the outside of sperm. Why don't we just try all of them and see which ones stick to the proteins that we know are on egg?" And if you think about doing this experimentally, that's like, well, I'll spend the next two millennia, you know, a year of time $100,000 each of the next millennium and I'll get a nice paper in nature is not that's not a feasible approach. But AlphaFold is pretty fast and they had some computers available. So they tried all of them and they both came out with this one protein TM something. I can't remember the number, but this was the one that they didn't know what it did before and now they find out it's AlphaFold says it sticks to the egg and that this is how kind of the first step of fertilization and so of course they don't just trust AlphaFold right it's a computational system.

John

所以他们去做了实验,说:如果我移除那个蛋白质或改变它,会发生什么?他们发现,如果改变或移除那个蛋白质,精子和卵子会靠近但不会受精。所以从一个大致的假设——精子表面有某种蛋白质起这个作用,AlphaFold 说我认为是这一个——然后你去做详细的实验来确认,现在你可以考虑诸如不孕不育的问题。如果在那蛋白质中看到突变,那可能是不孕的原因,也许我们可以考虑治疗。我认为从粗略的假设,到中间的 AlphaFold,再到实验确认,最终我们或许可以在此基础上考虑药物设计,但我们必须先获得这种生物学理解,为细胞中的所有部分赋予意义。我认为 AlphaFold 在早期阶段真正帮助的是为细胞各部分赋予意义,然后像 Isomorphic 这样的公司用它来构建具有靶向效应的小分子。

So they went and they said well what happens if I remove that protein or if I change that protein they find if they change that protein or remove that protein then sperm and egg will get close but they won't fertilize so you go from this broad hypothesis there's some protein on the surface of sperm that does this AlphaFold says I think it's this one and then you go do your detailed experiments to confirm and now you can think about questions like infertility now if you see mutations in that protein maybe that's a cause of infertility maybe we can think about treating that and we go I think from kind of rough hypothesis AlphaFold in the middle confirm with experiment and now maybe we can think ultimately about something like drug design on top of this but we have to get this biological understanding first that we bring meaning to all those pieces in the cell and that's what AlphaFold I think really helps with in the early stages is bringing meaning to the parts of the cell and then later companies like Isomorphic use it in order to build small molecules that have targeted effects.

Host

我们应该谈谈 AlphaFold 2 和 AlphaFold 3 的区别,对吧?因为 AlphaFold 2 是从氨基酸序列预测蛋白质结构。但生物学当然不仅仅涉及蛋白质,对吧?还有所有其他生物分子:DNA、RNA、小药物分子、离子(带电粒子)等等。所有这些都在相互作用。那么,在过程中多早你就意识到需要改变 AlphaFold 2 的基本模型以纳入这些额外分子?

We should talk about the difference between AlphaFold 2 and AlphaFold 3 though, right? Because I mean AlphaFold 2 was like predicting the structure of proteins from these you know strings of amino acids. But of course biology isn't only about proteins, right? You've got all of these other biomolecules. You've got DNA. You've got RNA. You've got like small drug molecules for example. You've got ions you know charged particles etc. Like all of these are interacting with each other. Um, so how early in the process did you know that you needed to change the fundamental model of AlphaFold 2 in order to incorporate those additional molecules?

John

所以,甚至在全世界知道 AlphaFold 2 之前,我们就坐在那里做梦,做梦有两个原因。一是我们有很多蛋白质天然存在于所谓的复合物中,多个蛋白质粘在一起。有时不一起预测它们的结构就没有真正的方法。所以我们已经在考虑这个多蛋白质问题。例如,正如你所说,结合药物、小分子,可能只有 20 个原子。你可以想想阿司匹林,对吧?它粘在蛋白质上。我们知道这非常重要。我们说,但以后再说。我们开始谈论这个梦想,这个目标。我们称之为全 PDB,对吧?蛋白质数据库(PDB)是我们使用的数据源,但我们拿过来后扔掉了许多东西。哦,这个有 RNA 或 DNA 附着在蛋白质上。好吧,我们扔掉 RNA 或 DNA,只让 AlphaFold 预测蛋白质。因为无法处理那些额外分子,我们无法处理那种复杂性,尽管我们非常想处理。我们有 20 种氨基酸产生 20 种结构,然后我们进行预测,我们所有的代码都基于此。我们想,这是以后的挑战,但最终我们会开始做。我们几乎立刻意识到,我们在 AlphaFold 中做的许多决定非常好、非常有帮助,但扩展到更复杂的事情时非常烦人。另一项工作是,我们试图找出如何简化 AlphaFold,我们想,好吧,AlphaFold 很复杂,但也许有些东西可以去掉。

So even before the world knew about AlphaFold 2, we were sitting there and dreaming and we were dreaming for two reasons. One is that we had a lot of, for example, proteins that exist naturally in what are called complexes, multiple proteins stuck together. And sometimes there's no real way to predict their structure without predicting them all together. So we're already thinking about this multi-protein problems. For example, as you say, bind drugs, small molecules, maybe 20 atoms. You can think of aspirin, right? Sticks to a protein. And we knew that this was really important. We said, but later. And we started to talk about this kind of dream about this goal. We would call whole PDB, right? So the Protein Data Bank, the PDB is the data source we use, but we take it in and we threw away a lot of the things. Oh, this has RNA or DNA attached to the protein. Well, let's throw away the RNA or DNA and just have AlphaFold predict the protein. And because you can't handle those extra molecules, we couldn't handle that complexity that we were very like driven by. We have 20 amino acids that produces 20 types of structures and then we will predict and all our code was kind of based around that. And we're like, it's a challenge for later, but eventually we'll start doing it. And one of the things we almost immediately realized is a lot of the decisions we made in AlphaFold were very good and very helpful and very annoying to extend to more complicated things. The other bit of work is that we were trying to figure out how to simplify AlphaFold and we thought okay well AlphaFold is complicated but maybe there are some things we can remove.

Host

那么架构从 AlphaFold 2 到 AlphaFold 3 是如何转变的?

How has the architecture shifted from AlphaFold 2 to AlphaFold 3 then?

John

有很多变化,但我会说有两个大的变化主题,当我们试图处理更多像 DNA、小分子等时。我们采用了一种称为扩散架构的东西,一种处理不确定性的不同方式,我可以告诉你更多,但另一个是真正思考了很多关于进化和进化数据的作用。

So there's a lot of changes but I would say there's two big themes of changes when we were trying to handle much more of the kind of the DNA small molecules etc. We adopted this thing called a diffusion architecture a different way in which we handle our uncertainty and I can tell you more about that but then I think the other one was really thinking a lot about the role of evolution and evolutionary data.

Host

那么让我问你这个问题,因为这是我记得对 AlphaFold 2 相当关键的一点。这个想法是,蛋白质实际上在许多不同生物中多次进化,实际上有一些关于它们进化历史的线索,会指示氨基酸在最终折叠形状中可能的位置。

Well, let me ask you about that then because this is one of the things that I remember being quite key to AlphaFold 2. This idea that actually proteins have evolved in lots of different creatures numerous times and actually there is like some clues about their evolutionary history that will indicate where amino acids are likely to end up in the final folded shape.

AF2 与 AF3 的进化信息差异 Evolutionary information in AlphaFold 2 vs AlphaFold 3

Host

所以即使你从一串氨基酸开始,你也不是完全盲目,因为这种事情以前发生过,对吧?这最终成为了模型的一个关键部分,但我觉得也可能是让模型对其他分子不太灵活的部分之一。这样说公平吗?

So that even if you're starting with a string of amino acids, you're not going in totally blind because this sort of stuff has happened before, right? That ended up being quite a key part of the model, but it was also, I think, potentially one of the parts that made the model quite inflexible to other molecules. Is that fair?

John

所以,AlphaFold 2 以这种丰富的方式在几乎每个模块的每个部分都使用了进化信息,说‘这里是进化信息,以防你需要’。但我们在 AlphaFold 3 中研究的很多内容,我们知道我们正在走向的方向,并没有进化信息。所以我们对着它大喊,却什么也没有。我们有点担心这既会减慢网络速度,也可能导致其工作方式出现一些不良动态。所以我们决定将其从大部分网络中移除,转而强调几何信息,这才是真正始终存在的东西。结果效果非常好,实际上比我们预期的还要好。

So, AlphaFold 2 used evolutionary information in this exuberant way at kind of every part of almost every block, saying, 'Here's the evolutionary information in case you need it.' But a lot of what we studied in AlphaFold 3 that we knew we were moving toward didn't have evolutionary information. And so we were shouting at it with nothing. And we were kind of worried that this was both slowing down the network and also possibly leading to some bad dynamics in how it works. And so we decided to just take that out of most of the network and otherwise emphasize the geometric information, the thing that really is always there. And that turned out to work exceedingly well, actually better than we expected.

Host

我想知道我们是否可以用一个类比来解释这种架构上的差异。假设你在策划一场婚礼,需要安排整个婚礼的座位表。你有所有这些客人——这些就是你的氨基酸——你必须弄清楚每个人该坐在哪里。有几种不同的思考方式。你可以考虑成对互动,比如这个人坐在那个人旁边,这是好的互动吗?但这对这边这个人或那边那桌意味着什么?但你也可以考虑你对这些人的历史了解。我这个类比怎么样?

I wonder if there's an analogy we can use here for the difference in this architecture. Let's imagine that you're planning a wedding and you have to do the seating plan for the entire wedding. You have all these guests—those are your amino acids—and you have to work out where each one sits. There are a couple of different ways you can think about this. You could think about pairwise interactions, like this person sitting next to this person, is that a good interaction? But then what does that mean for this person over here or that table over there? But you could also potentially think about the history of what you know about those people. How am I doing so far as an analogy?

John

我可以接受这个。我觉得,你知道,有些人一起上过学。有些人曾经约会过,然后分手得很糟糕,对吧?那些可能是……

I can go for this. I think, you know, some people went to school together. Some people used to date and had a terrible breakup, right? Those might be...

Host

坐在一起。

Sit there together.

John

你可能不想让他们紧挨着坐,除非你真的想找点火花。但之前我们只讨论了婚礼客人坐哪里,现在我们要考虑花艺布置在哪里。

You probably don't want to sit them right next to each other unless you're really looking for sparks. But before, we just talked about where the wedding guests sit, and now we think about where the flower arrangements are.

Host

这很好。

That's nice.

John

对吧?我们考虑所有这些其他东西,它们共同构成了婚宴晚餐。

Right? We think about all these other things that come together to become the reception dinner.

Host

在这个类比中,AlphaFold 2 非常关注客人的历史,对吧?它不断地根据过去检查他们最适合的位置。这对蛋白质来说很好,但一旦你试图加入接待处的其他元素,就变得困难了。一旦你开始引入其他生物分子,你就不想那么关注历史了。

In this analogy, AlphaFold 2 was very focused on the history of the guests, right? It was sort of continually checking where they might best fit based on their past. And that's great for proteins, but it's difficult once you try to include other elements of the reception. And once you start bringing in other biomolecules, you don't want to focus so much on history.

John

我觉得这完全正确。不过我想说的一点是,我们总是让这些历史信息可用,也就是进化历史,在类比中就是我们对他们过去的了解。我们发现,AlphaFold 我们认为并没有太多依赖它,除了在最开始的时候,它就像在说‘哦,这些人应该在一起。这些人应该分开。我知道一些事情。’但后来它训练自己忽略了它。所以通过检查并发现我们可能没有使用这些信息,也许我们应该停止不断地将其附加到处理过程中。

I think that's all true. I think one thing though I would say is that we kind of always made this history available, this evolutionary history, and in the analogy, what we know about their past. And what we would find is that AlphaFold, we think, was not relying on it much other than at the very beginning, where it was like saying, 'Oh, well these people should probably be together. These people should probably be apart. I know a couple of things.' But then it kind of trained itself to ignore it. And so by inspecting and seeing that we're probably not using that information, maybe we should stop kind of attaching it constantly into the processing.

Host

但结果,你成功地大幅简化了模型。

But then as a result, you managed to massively simplify the model.

John

我不会说我们大幅简化了,但我们让它大幅更准确,突然我们就能解决新问题了。事实上,我们做了轻微调整,然后得到了一个更好的模型,结果发现即使是蛋白质-蛋白质问题,与配体或核酸等无关的问题,也因这种科学和改进而大幅改善。

I wouldn't say we massively simplified, but we made it massively more accurate and suddenly we were doing new problems. And in fact, we made light adjustments and then we made a much better model, and then it turned out that even that protein-protein problem, something that has nothing to do with ligands or nucleic acids or anything else, even that protein-protein problem got massively better from this kind of science and improvement.

AF3 中的扩散与幻觉问题 Diffusion in AlphaFold 3 and hallucination concerns

Host

所以扩散是一种不同的训练神经网络的想法。AlphaFold 2 系统非常侧重于蛋白质,侧重于蛋白质主链的形状。在 AlphaFold 3 中,我们转向了扩散,你基本上说,‘这是蛋白质的模糊图像。我把整个蛋白质加上一些噪声、一些误差,就像你用错误的处方眼镜看它,然后它猜出正确答案。’然后你不断优化它。这给了我们非常好的局部几何理解,知道如何让事物极其精确,因为这是它在小尺度上所做的,以及处理大系统的方法。这给了我们一种新方法,我们不必过于纠结蛋白质具体看起来如何的细节,因为它们不同于 DNA、RNA 和小分子。好处是它让我们非常容易地处理我们研究的这个广阔宇宙。坏处是它导致了更高的幻觉率,出现奇怪的东西。所以我们需要用不同的方式处理这个问题。

So diffusion is this different idea in how you train a neural network. The AlphaFold 2 system was really heavily based around proteins, around the kind of shape of a protein backbone. In AlphaFold 3, we went to diffusion where you basically say, 'Here's a blurry image of the protein. I kind of took all of the protein and added some noise, some error, like you looked at it with the wrong prescription glasses, and then it guessed the right answer.' And you have it constantly refined. And so what this gave us was a really great understanding of local geometry, of how to make things extremely precise because that's what it does at small scales, and this way of tackling big systems. And that gave us a kind of new approach that we didn't have to get so involved in the details of exactly how proteins look because they're different than DNA. They're different than RNA and small molecules. And the upside is that it made it really, really easy to handle this wide universe of things that we study. The downside is that it led to a higher rate of hallucination of weird stuff appearing. And so then we needed to handle that in different ways.

Host

嗯,这是 AlphaFold 2 和 AlphaFold 3 的一大区别,对吧?你引入了随机性,幻觉的可能性。人们在使用时应该对此有多担心?有没有危险,因为他们觉得 AlphaFold 2 非常准确,就把 AlphaFold 3 当作某种神谕?

Well, this is one of the big differences between AlphaFold 2 and AlphaFold 3, right? That you have this introduction of stochasticity, the potential for hallucinations. How much should people be concerned about that when they're using it? Is there a danger that they think because AlphaFold 2 was so on the money, they think of AlphaFold 3 as though it's some kind of oracle?

John

我认为生物学家很棒的一点是,作为科学家,他们对工具深怀疑虑。AlphaFold 2 确实有一个优势,错误的答案通常看起来很愚蠢。没有人看到那个会说‘这绝对是一个蛋白质’。

I think one of the wonderful things about biologists is that as scientists, they're deeply skeptical of their tools. AlphaFold 2 did have an advantage that wrong answers often looked stupid. They didn't—no one looked at that and said, 'That's definitely a protein.'

Host

而 AlphaFold 3 中的错误答案有时更可信。

Whereas wrong answers in AlphaFold 3 are sometimes more plausible.

John

但我认为人们已经变得非常擅长,虽然不是完美无缺,但非常擅长说,‘嗯,AlphaFold 2 也告诉我它认为自己有多准确,在置信度度量中。我也应该利用这一点。’所以这是一种社会知识。科学家使用的任何实验或工具都有局限性。

But I think that people have gotten really good, not uniformly perfect, but really good at saying, 'Well, AlphaFold 2 is also telling me how accurate it thinks it is in the confidence measure. I should also use that.' And so it's this kind of social knowledge. There's no experiment or tool that scientists use that is without limitation.

Host

对。即使是实验结构测定也有所有这些已知的缺陷。所以我认为科学家使用得相对较好。我还没看到有人使用 AlphaFold 3 出问题。我认为这是因为现在它已经成为科学家教育和社区的一部分,当你使用计算方法时,你要看这些置信度度量。我们根据置信度给蛋白质上色,最终我们也把它看作一个工具,用来提出假设,然后通过实验检验这些假设。AlphaFold 2 也绝非完美,它只是非常非常有用。

Right. Even experimental structure determination has all these known faults. And so scientists, I think, use it relatively well. I haven't seen any from people using AlphaFold 3. I think because it's now such a part of the education and community of scientists that when you use computational methods, here are the confidence measures you look at. We color our proteins by confidence and ultimately we also think of it as a tool where we'll induce hypotheses and we'll test those hypotheses experimentally. And AlphaFold 2 is not perfect by any means. It's just very, very useful.

AI 与科学的可解释性 Interpretability in AI and Science

Host

可解释性在这其中有多重要?我的意思是,人类想要理解为什么 AlphaFold 以特定方式折叠蛋白质。很多人对此感兴趣,你会听到有人非常自信地宣称,我们只有在完全理解 AI 系统的情况下才能使用它们。他们几乎意味着,如果我能写出一个算法来代替那个 AI 系统就好了。我认为这种渴望——那是个讨厌的黑箱,如果它不是黑箱就好了——这种要求我们必须完美理解的狭隘要求,老实说有点奇怪。我想到了我们在科学中完全满意于没有这种理解的例子。比如,实验科学本身。看看早期晶体学中人们如何结晶蛋白质,当时并不清楚这些结构是否与液体中自由漂浮的蛋白质一样。更多的实验表明大多数时候是差不多的。所以我们在科学中一直以这种部分可解释的方式工作。我认为可解释性有很好的应用,比如我想理解网络以便改进它,做出更好的 AlphaFold 版本。我之前讲过一些我们如何做这类工作的故事。有些人说他们想要可解释性以便信任它。我认为更重要的是,如果你真的想知道是否信任一个答案,我们有很好的表征表明我们的置信度指标是可靠的指南,人们在实践中用它来判断答案是否可能正确。

How important is interpretability in all of this? I mean this idea that humans want to understand why AlphaFold is folding a protein in a particular way. There's a lot of interest in it and you will hear people sometimes make very confident pronouncements that you know we can only use AI systems if we perfectly understand what they do. And they almost mean if I can write down an algorithm that I could use instead of that AI system. And I think it's this desire of that's an annoying black box. What if it just wasn't a black box? And I kind of feel like that kind of narrow demand for we must understand it perfectly is honestly kind of a weird demand. And I think about cases in which we've been perfectly happy not having that in science. One, for example, is just experimental science in general. If you look at say how someone crystallizes a protein early in crystallography, it wasn't clear if those structures were going to look just like a free protein floating in liquid. And more experiments kind of said most of the time it's about right. So we always in science totally worked in this kind of partial interpretability way. I think there's really good applications of interpretability when we think about okay I want to understand the network so that I can change it and make a better version of say AlphaFold. I described some stories earlier about how we do that kind of work. Some people will say well I want interpretability so I can trust it. I think more important than that is if you really want to know whether you trust an answer, well, we have pretty good characterization that our confidence metrics are a reliable guide and people use them in practice to decide when an answer is probably true.

John

事实上,我很希望看到更多关于 AlphaFold 的可解释性工作。到底是什么让它如此广泛地泛化?我认为可以做得更多,但这不一定会给人们带来他们以为会得到的东西。

In fact, I'd love to see a lot more interpretability work go on for AlphaFold. What exactly leads it to generalize so widely? I think there could be more done, but that won't necessarily give people what they think it will give them.

Host

我想有时候,比如罗马人建造桥梁和水渠,并没有完全理解重力,他们没有牛顿方程。AlphaFold 在这里是不是有点像罗马建筑、罗马工程,但用于生物学?我们现在能够用你和你的团队创造的工具建造东西,尽管我们不一定完全理解它们为什么有效。

I think sometimes if like okay the Romans for example were building bridges and aqueducts without having you know a full understanding of gravity right they didn't have like Newton's equations is there a way in which AlphaFold here is like Roman building Roman engineering but for biology like we are able to build stuff now with the tools that you and your team are creating even though we don't necessarily have full insight into why they're working.

John

我的意思是,罗马人是很久以前就这样做了。

I mean Romans are way back in order to do this.

Host

我觉得这很棒。

I mean this feels great.

John

想想现代喷气式飞机。

Think about a modern jet airplane.

Host

是的。

Yeah.

John

或者现代汽车。

Or a modern car.

Host

是的。

Yeah.

John

对。我们理解流体动力学的纳维-斯托克斯方程等等,对湍流也有一点了解。

Right. We understand you know the Navier Stokes equations of fluid dynamics etc. A bit about turbulence.

Host

尽管有这种底层理解,

Despite that low-level understanding,

John

我们既建造风洞来测量气流,也构建模拟,针对这种精确的机翼几何形状展示空气如何流过。

we both build wind tunnels to measure flow and we build simulations that for this precise wing geometry show you how the air goes over it.

Host

嗯。

Mhm.

John

我认为罗马建桥的比喻可能更贴切地描述了我们的 AI 开发方式:我们有一些直觉,就像罗马人有直觉,他们建造了美丽的桥梁,但他们没有所有方程和完全理解,却建造了他们需要的东西,并且能够驾驶马车过桥。所以在这个意义上,我们在 AI 中部分凭直觉操作,但下游使用 AlphaFold 等工具的人,我认为更像是拥有一个你可能不完全理解的伟大计算包。你的专长并不在于气流如何导致湍流,但你学会了如何改变和适应,并用这个工具进行更大规模的科学。我认为这确实是 AI 用户所做的,稍微不那么罗马的版本。

I think the Roman bridge building is maybe a better analogy of how we do AI development where we have some intuitions just like the Romans had intuitions and they built some beautiful bridges but they didn't have all the equations and full understanding and yet they built things they needed and yet they were able to drive carts across the bridge they built right so we are in that sense operating partially intuitionally in AI but downstream the people using tools like AlphaFold I think It's more like having a great computation package that maybe you don't exactly understand. Your expertise is not exactly in how this airflow results in turbulence, but then you figure out how to change it and adapt and you work with this tool to do your larger scale science. And I think that's that's really the slightly less Roman version of what uh AI users are doing.

Host

嘿,听着,罗马人,没有贬低罗马人的意思。

Hey, look, the Romans there's no shade on the Romans.

John

罗马人给了我们什么?是的。

What have the Romans given us? Yes.

Host

完全正确。好吧,我们来谈谈一些下游应用。今年早些时候,我采访了 Isomorphic Labs 的 Max Jaderberg 和 Rebecca Paul,他们正在将你的 AlphaFold 工具用于药物设计。看到你们建造的东西以这种方式实际应用于药物设计,感觉如何?

Exactly. Exactly right. Okay. Okay. Well, let's talk about some of those downstream applications because earlier this year I got to speak to Max Jaderberg and Rebecca Paul from Isomorphic Labs about how they're using your AlphaFold tool in drug design. What has that been like to see this thing that you guys built actually being implemented in drug design in that way?

John

我认为看到它走这么远真的非常了不起,并且成为其中的一部分。关于药物设计,有一点是它不仅仅是蛋白质结构预测。我喜欢提醒人们,一个蛋白质结构大约花费 10 万美元,而一种药物大约花费 10 亿美元,对吧?所以这告诉你,它不能全是蛋白质结构测定。看到人们试图在此基础上进一步发展这些想法,并真正找到将其整合到应用中的方法,这真的很特别。我们在整个制药行业都看到这一点:我们如何围绕这个构建流程,最终得到给患者服用的分子,通过所有这些不同的测试。你知道,有些会帮助分子如何结合,或者生物学是什么?这个蛋白质到底是不是靶点?而有些我们几乎无关,比如这种药物是否会在肝脏中代谢?也许有一种蛋白质-小分子相互作用可以用来帮助这一点。但大多数情况下,AlphaFold 可能不是那个工具。我认为我们既需要理解生物学的工作,也需要专门制造靶向药物的分子,这两者都很重要。这是一个令人兴奋的组合。

I think it's just really extraordinary to see it carried so far and to be a part of, you know, one of the things about drug design is it's not just protein structure prediction. I like to remind people that a protein structure costs about $100,000 and a drug costs about a billion, right? So that can tell you that it can't all be protein structure determination. I think it's really exceptional to see people trying to build on and take these ideas further and really find also a way in order to integrate it into application. We see this across the farm industry like how are we going to build processes around this that enable us to ultimately end up with molecules that are dosed in patients that pass all these different things. You know, some will help with how does the molecule stick or what is the biology? Is this protein a target at all? And some that we have, you know, very little to do with, you know, will this drug be metabolized in the liver, right? Maybe there's a protein small molecule interaction that you can use to help that. But for the most part, AlphaFold is probably not the tool. And I think it's really important that we have, you know, both the work in how do we understand biology and then the work in specifically how do we make molecules to drug targets. It's an exciting combination.

Host

我想,在和他们的谈话之前,我没有完全意识到的是,实际上,找到一个分子与特定蛋白质靶点结合只是治愈疾病的一小部分,对吧?比如阿尔茨海默症,我们知道蛋白质参与其中,但甚至不一定有靶点可寻。

I think one of the things that I hadn't quite appreciated until having that conversation with them is that actually, you know, finding a molecule to bind to a particular protein target is such a small subset of curing disease, right? Like I mean Alzheimer's is an example where we know that proteins are involved but there isn't even necessarily a place to target yet.

John

嗯,我们甚至不知道。我们仍然不知道淀粉样蛋白β积累是否在因果链中。它是症状吗?

Well, we don't even know. We still don't know if amyloid beta accumulation is in the causal chain. Is it a symptom?

Host

对吧?那是一种蛋白质。但分解它,我的意思是,已经有一些开始看到效果了。我认为这是一个最重要的例子:如果你考虑找到一种与蛋白质结合的药物,找到一种至少在动物模型中大部分无毒的药物,完成所有这些早期药物设计的艰难阶段——我们正确地认为这些非常非常困难,需要多年时间。但对我来说,更大的问题是,即使你做了所有这些,90%的药物在临床试验中失败。所以即使你做对了所有事情,它们仍然无效或不安全,对吧?我们通过实验确定这一点。而这很大程度上是我们对生物学的巨大无知,对吧?我们不知道阿尔茨海默症、自闭症的原因。

Right? That's a protein. But breaking that up, I mean, there's been some, you know, starting to see a bit of effect. I think it's an example of one of the most important things to say if you think about finding a drug that sticks to a protein, finding a drug that's non toxic, at least for the most part in animal models, finding doing all these hard stages of early drug design that we rightly say are very very difficult, take people off in years. But the bigger problem to me is like even when you do all that, 90% of drugs fail in clinical trials. So even though you do all those things right, they still don't work or they're still not safe, right? We determine this experimentally. And a lot of this is our grand ignorance of biology, right? That we don't know the causes of Alzheimer's of autism.

AlphaFold 在疾病理解中的作用 AlphaFold's role in understanding disease

Host

即使我们有了病因的线索,比如亨廷顿病有非常明确的遗传关联,制造一种真正改善患者生活的分子仍然非常困难。还有很多巨大的问题有待解决。我们才刚刚开始涉足计算生物学这个领域,或者说在继续推进,但仍有大量工作要做。

Even when we have ideas of cause, for example Huntington's with very clear genetic correlates, still making a molecule that actually makes those patients' lives better is so very difficult. There are so many giant problems left. We're only starting this world of computational biology, or continuing let's say, but still there's so much left.

Host

那么,让我理解一下。请给我一个例子,说明 AlphaFold 如何帮助理解疾病。

Well, let me understand that then. So give me an example of a way that AlphaFold can be used to help understand disease.

John

我喜欢 AlphaFold 的一个案例,是最近的一个:人们试图理解胆固醇如何在体内运输。有一种蛋白质参与将脂肪分子从一个地方运输到另一个地方。我相信它也存在于一些与心脏病相关的斑块中。当我们开始理解这种生物学时,AlphaFold 贡献了一个很好的部分:这个分子的详细结构,而他们只能用一种叫做冷冻电镜的方法拍出非常模糊的图像,这对该技术来说并不罕见。但那个模糊的图像与 AlphaFold 的结构非常匹配。所以现在你可以说,好吧,这就是那个运输胆固醇的东西。也许我可以干扰或改变它运输胆固醇的方式。也许我可以添加一个小分子。但当然,你可能会首先说,哇,那是蛋白质。为什么不直接加一种阻断它的药物呢?我认为你会立刻发现那会很糟糕。你的身体不是偶然拥有这种蛋白质的。这种蛋白质的目的不是导致心脏病,对吧?它的目的是将脂肪分子运送到细胞中需要的地方。

One case study I like from AlphaFold, one that is somewhat recent: people were trying to understand how cholesterol is moved around the body. There is this protein that is involved in the transport of fatty molecules from one location to another. I believe it is also found in some of the plaques that build up that are correlated to heart disease. Even as we start to understand that biology, we have this nice piece that AlphaFold contributed: the detailed structure of this molecule that they could only take an extraordinarily fuzzy picture with a method called cryo-electron microscopy, which is not an uncommon outcome for that technique. But then that fuzzy piece actually really well matched with the AlphaFold structure. So now you can say, okay, well this is that thing that is moving cholesterol. Maybe I can interfere with or change how it moves cholesterol. Maybe I can add a small molecule. But of course, your first thing you might say, whoa, that's the protein. Why don't you just add a drug that blocks it? And I think you would immediately find out that would be really bad. Your body didn't have this protein by accident. The purpose of this protein is not to cause heart disease, right? The purpose of this is to move fatty molecules where they need to be in the cell.

Host

你确实需要它,对吧?你确实需要它。你将会需要它。所以实际上你需要弄清楚的是,这如何给你一些新想法,来改变它在细胞中的行为,而不杀死患者,并改善他们的生活。我认为 AlphaFold 是那个故事的一部分,但不是终点。

You sort of need that, right? You sort of need that. You're going to sort of need that. So what you're actually going to need to figure out is how this might give you some new ideas to change how this behaves in the cell without killing the patient and making their life better. And I think AlphaFold is a part of that story. It's not the end.

从预测到设计:AlphaFold 的意外影响 From prediction to design: AlphaFold's unexpected impact

Host

不过,这一切似乎有一个自然的下一步。如果你预测蛋白质的形状,然后用这些模型来解释蛋白质在人体中的功能,那么接下来是不是要设计新的蛋白质?

It feels like there's a natural next step to all of this though. If you are predicting the shape of proteins and then using those models to interpret the function of proteins in the human body, does it then go on to designing new proteins?

John

哦,是的,人们一直想要那样。他们看着这些美丽的蛋白质说,我希望人类也能做到,对吧?所以有很多出色的工作,事实上很多是在 David Baker 的实验室完成的,他和 Demis 以及我一起获得了诺贝尔奖。AlphaFold 在这方面已经惊人地具有变革性:我们如何从已经构建了理解它的计算系统,到设计我们自己的蛋白质?事实上,很大一部分新批准的药物是蛋白质,通常是抗体,最初通过非常有趣的方式发现:向小鼠或羊驼注射你想要针对的蛋白质,利用它们的自然免疫系统来找到它。但我们开始非常认真地讨论如何设计具有我们想要效果的蛋白质。事实证明,其中最重要的部分是你可以设计很多你认为可能有效的东西。在实验室测试非常耗时、困难且昂贵。所以关键之处在于使用 AlphaFold 作为自然的代理:试图说我们如何整合 AlphaFold 对蛋白质如何结合的理解,如何利用它来最大化蛋白质设计的信号?人们已经取得了非凡的成功;他们变得非常非常擅长让蛋白质在想要的地方结合。

Oh yeah, people have wanted that. They've looked at these beautiful proteins and said, I wish humans could do that, right? And so there's been all this exceptional work, and in fact a lot of it done at David Baker's lab who, with Demis and I, won the Nobel. AlphaFold has actually been shockingly transformative at this: saying how do we go from now we've built these computational systems that understand it, how are we going to design our own proteins? In fact, a large portion of new approved drugs are proteins, normally antibodies discovered initially in very interesting ways: injecting mice or llamas with something that you want to build a protein against and using their natural immune system to find it. But we are starting to talk very seriously about how we are going to design proteins to have the effect we want. And it turns out that the most important part of that is that you can design many things you think might work. It's extraordinarily time-consuming, difficult, and expensive to test in the lab. So what's been so important there is using AlphaFold as a proxy for nature: trying to say how do we integrate AlphaFold's understanding of how proteins stick together when they do, how do we use that to maximally make a signal for protein design? And people have gotten extraordinarily successful; they've gotten really, really good at getting proteins to stick just where they want.

Host

哇。好吧,但等等。因为 AlphaFold 的初衷并不是看蛋白质如何结合。

Wow. Okay, but hang on. So because that wasn't the original intention of AlphaFold to see how proteins stick together.

John

它们的初衷并不是看它们如何结合。事实上,这是来自 Twitter 的一个早期惊喜,两个不同的人说,你知道吗,如果你想知道两个蛋白质是否结合,我们正忙于构建一个多蛋白质的、正确完成的系统。他们说,好吧,只需把这两个蛋白质放在一起,中间加一些随机氨基酸,看看它们是否以这种方式结合。那是世界上最好的判断蛋白质是否结合的系统。哇。我们没想到我们在构建一个能够以非常深入的方式帮助人们设计蛋白质的系统。我们以为我们会利用这个根本性突破,然后继续做下去。然后人们说实际上它已经有效了。我认为这是一个在 AI 中反复出现的宏大故事。所以也许我们应该预料到:如果你训练一个模型在某个任务上非常非常出色,它必须学习很多深层事实。比如说,如果你想在结构预测上非常出色,它会学习关于蛋白质如何相互作用的深层事实。如果你做对了实验,你就能获取这些知识。但有一个完整的领域,我认为人们开始称之为 alpholdology,人们会找出哪些方法有效。他们只是把它当作一个非常酷的黑盒子,可以开始实验并尝试自己的想法。我认为有很多非常伟大的科学已经并且正在沿着这个方向进行。我们仍在探索,并且开始有关于如何制造酶、进行化学反应的蛋白质的工作?如何做真正复杂精密的事情?我们仍然,你知道,大自然仍然会嘲笑我们设计蛋白质的能力。但我们开始开发这些非常有趣的工具,可能是治疗性的,也可能是探究细胞的方法:你可以将两个蛋白质放在一起,看看细胞如何因为你的操作而改变。我认为我们不仅会在治疗工具中看到这种相互作用,而且现在我们能够以令人兴奋的方式探究细胞,每次我们开发出这样的能力,我们都会对细胞有更干预性的理解,并将其应用于医学和合成生物学。

Wasn't the intention of them to see how they stick together. In fact, that was an early surprise from Twitter where two different people said, you know, if you want to know if two proteins stick together, we were busy making a multi-protein like properly done system. They said, well, just take those two proteins and put some random amino acids in the middle and see if they stick together that way. And that was the best system in the world for seeing if proteins stick together. Wow. We didn't think we were making a system that could help people design proteins in a really deep way. We thought we would go use this fundamental breakthrough and go on and do it. And then people said actually it already works. And I think it was this grand story that does show up again and again in AI. So maybe we should have expected it: that if you train a model to be really, really good at a task, it has to learn a lot of deep facts. Say if you want to be really good at structure prediction, it learns some deep facts about how proteins interact. And if you do just the right experiments, you can kind of access that knowledge. But there was this whole field of what I think people started to call alpholdology where people would find out which things worked. They would just treat it as this really cool black box that they could start experimenting with and try their own ideas on. I think there was a lot of really great science that has been and continues to be done in that vein. We're still figuring out and there's starting to be work on like how do we make enzymes, proteins that do chemistry? How do we do really complicated sophisticated stuff? We're still, you know, nature would still laugh at us on our ability to design proteins. But we are starting to develop these really interesting tools that are maybe therapeutics that are also maybe ways to interrogate the cell: that you can bring two proteins together and see how the cell changes because you do that. I think we'll get this interplay in not only the tools we use for therapeutics but now our ability to poke the cell in exciting ways to interrogate it, and every time we develop that, we'll develop a more interventional understanding of the cell that we will bring forward to medicine and synthetic biology.

AlphaProteio 与蛋白质设计 AlphaProteio and Protein Design

Host

你这里描述的是 AlphaProteio 吗?

Is this AlphaProteio that you're describing here?

John

AlphaProteio 是 Google DeepMind 内部的一个项目,致力于蛋白质设计,思考结合和酶的问题,并真正尝试找出如何获得可靠的系统,尤其是针对这些超级困难的问题。我认为我们在设计领域仍然看到了很多成功,但当你实际设计蛋白质时,你必须去实验室测试它们;没有其他方法可以知道。在找出预测它们是否有效的正确方法方面,AlphaProteio 的工作表明我们可以在这方面取得越来越大的进展。

So AlphaProteio is Google DeepMind's internal effort to do protein design and think about problems in binding and enzymes, and really trying to figure out how we get reliable systems, especially for these super hard problems. I think what we're seeing is in the design space still a lot of success, but when you're actually designing proteins, you have to go to the lab and test them; there's no other way to find out. And in finding out the right ways to predict if they're going to work, the AlphaProteio work has shown that we can get further and further in doing this.

Host

那么给我举个例子,有哪些类型的蛋白质是目标,能够设计出来会很棒?

Give me an example then of some of the type of proteins that are the target, right, that would be nice to be able to design.

John

你知道吗,我觉得如果你问任何一位蛋白质设计师,他们都会有一个最爱,而他们的最爱真的是:我们能否让蛋白质做像碳捕获这样的事情?我们能否真正构建出对应对气候变化有意义的酶?我认为其他你真正看到的例子,比如降解微塑料或环境塑料。不过,我想说的一点警示是,对于所有这些,当你谈论实际应用时,就像人们对药物设计的理解是‘让分子结合,药物设计就完成了’——但事实并非如此,对吧?你需要更多特性:你需要它在各方面都可耐受,你需要它能制成药丸,你需要所有这些其他东西。同样在酶方面,你可能会想,‘哦,你只需要让这个反应发生。’酶是一种催化化学反应的蛋白质。但不对,实际上你需要它能多次进行这个反应,对吧?足够多次,这样你就不必为每个反应都制造新的蛋白质。你需要它足够快。你需要它不进行某些其他反应。你还有所有这些其他特性。我认为当我们思考从‘哦,这可能有点意思’到‘这真的非常有效’时,还有很多工作要做。不过公平地说,有趣的是,在合成进化酶方面,人们已经在使用它们了。你知道,很多洗衣粉中含有设计的蛋白质,我觉得这很迷人。我认为这是人们能认识到的少数几个设计蛋白质的应用之一。

You know what, I think if you ask any protein designer, they will have a favorite, and their favorite is really: can we make proteins do things like carbon capture? Can we actually build enzymes that meaningfully contribute to addressing climate change? I think other ones you really see, for example, degrading microplastics or environmental plastics. I think one of the things also though I'll say as a caution is that for all of these, when you talk about doing a real application, just like people's conception of drug design is 'get molecule to stick, drug design done right' — and that's not the case, right? There are so many more properties you need: you need it to be tolerable in all these ways, you need it to be pill-formulatable, you need all these other things. Similarly in enzymes, you might think, 'Oh, you just need to make this reaction happen.' An enzyme, right, is a protein that catalyzes a chemical reaction. But no, actually you need it to be able to do this many times, right? Enough that you're not constantly having to make new proteins for each reaction. You need it to be fast enough. You need it to not do certain other reactions. You have all these other properties. And I think there's a lot more to be done as we think about going from 'oh, maybe this is kind of interesting' to 'this really, really works.' Although in fairness, interestingly, on kind of synthetically evolved enzymes, people are already using them. You know, there's a lot of washing powder that has designed proteins, which I find fascinating. I think it's one of the few applications of designed proteins that people would recognize.

工程生物学 vs. 预测 Engineering Biology vs. Prediction

Host

是的,绝对如此。但工程化生物学比仅仅预测要难多少呢?

Yeah, absolutely. How much harder is it though to sort of engineer biology than it is to just predict?

John

我非常经验主义。你应该三年后再问我,到时候我们就知道了。它既更容易也更难。我喜欢用一个类比:如果你想弄清楚一个物体是什么,你可能会说,‘这是一辆自行车吗?’我会看到两个轮子、一条链条、一些车把,然后说,‘是的,那是一辆自行车。’但有两个轮子、一个车把和一条链条并不能让某物成为一辆能工作的自行车。所以当你设计某物时,你必须把所有的细节都做到足够正确,以至于它实际上能工作。我认为我们在蛋白质方面还在摸索这一点。目前,蛋白质结构预测可以说是‘已解决星号’——它是一个非常非常有用的系统,但并不完美。设计还没有解决,但我认为它正在快速进步,而且我不认为 15 年后我们还会说蛋白质设计极其困难。

I'm very empirical. You should ask me in 3 years and we'll know. It's easier and it's harder. One analogy I like to say is: if you were trying to figure out what an object is, you might say, 'Is this a bicycle?' And I would see two wheels, a chain, some handlebars, and I would say, 'Yeah, that's a bicycle.' But having two wheels, a handlebar, and a chain doesn't make something a working bicycle. So when you're designing something, you have to get all the details right enough that it actually works. And I think we're still figuring this out in proteins. And right now, protein structure prediction is, let's say, solved star — it's a very, very useful system, it's not perfect. Design is not yet solved, but I think it is advancing rapidly, and I don't think we'll be still talking about protein design being incredibly difficult in 15 years.

AI 与生物学:思考 vs. 实用 AI and Biology: Thinking vs. Utility

Host

好吧,让我们稍微放大视野,更广泛地谈谈人工智能和生物学,因为整个对话让我想起了几年前我们上次采访你时你说过的话,我有一段小片段可以放给你听。

Well, okay, let's zoom out a little bit on AI and biology more generally, because this whole conversation has reminded me of something that you said when we last interviewed you a few years ago, and I've got a little clip that I can play you what you said.

John (from clip)

我认为记住这一点非常重要:我们开发的这些非常强大的技术仍然远未达到真正的人工智能,那种可以谈论思考、做决定等等的人工智能。

I think it's really important to remember that these are really powerful techniques that we've developed that are still far short of a real artificial intelligence that you can talk about thinking and making decisions and everything else.

Host

我觉得这很有趣。那是 2022 年,对吧?我想知道你现在对此有何反思。你认为机器现在开始以智能的方式理解生物学了吗?你改变想法了吗?

I think that's so interesting. So that was 2022, right? I wonder how you reflect now on that. Do you think that machines are beginning to sort of understand biology in an intelligent way now? Have you changed your mind?

John

我认为无论它们是否能思考,它们在解决问题方面都非常有用。它们离人工智能或 AGI 有多远,我认为这几乎无关紧要。我认为真正有趣的地方在于我们能否将这些系统描述为足够可靠。我们能否为它们找到有用的事情做?我认为我们需要更加功利地看待它。当然,像 AlphaFold 这样的机器,我不一定会用‘思考’这个词。我不知道我们是否处于那种情况,对吧,我们过去常说,‘好吧,智能就是下棋。’我们应该研究下棋,因为一旦我们有能下棋的机器,我们就基本上有了智能。当然,我们确实有了能下棋的机器,在超人水平上,那是在 1994 年的卡斯帕罗夫比赛。但那并不是通向能读写的机器的道路。所以我认为我们总是抓住这些问题说,‘好吧,这就是问题所在,’或者人们相当乐观地将某物命名为‘人类最后的考试’——一个如此困难的问题,如果你解决了它,就没有必要再向机器提出问题了。而我对如何找到那些在某种意义上如此简单的问题非常感兴趣,以至于我们可以在它们上面做得非常好,并在构建 AGI 之前构建非常有用的系统。这些就是那种科学问题。当然,你想使用与试图构建 AGI 的人相关的技术。它们是强大的技术,但我们不必陷入哲学争论。我们可以直接构建有用的系统。事实上,我认为整个行业都在思考如何构建对软件开发人员有用的系统,对写作人员有用的系统,从而扩展我们解决问题的性质。然后我们会看到我们是否最终得到 AGI,但我们肯定会得到有用的系统。

I think that whether or not they can think, they're extraordinarily useful for solving problems. How far they are from AI or AGI, I think that's almost beside the point. I think the really, really interesting point is where we can characterize these systems as reliable enough. Do we find useful things for them to do? I think we need to be much more kind of utilitarian about it. And certainly machines like AlphaFold, I wouldn't necessarily apply the word 'think.' And I don't know if we're in the situation, right, that we used to say, 'Okay, intelligence was playing chess.' And we should work on chess because once we have machines that play chess, we've basically got intelligence. And of course, we got machines that played chess really well at a superhuman level in, what, 1994 was the Kasparov match. And that wasn't the path that led us to machines that can read and write. And so I think we always reach for these problems and say, 'Well, this is the problem,' or people rather optimistically name something 'humanity's last exam' — a problem so hard that if you solve it, there's no point in posing problems to machines anymore. And I'm very interested in how do we find those problems that turn out to be so easy in a certain sense that we can do incredibly on them and build very useful systems before we build AGI. Those are the kind of science problems. And of course, you want to use related techniques to the people trying to build AGI. They're powerful techniques, but we don't have to get tied up in the philosophy. We can just build useful systems. In fact, I think the whole kind of industry is thinking a lot about how do we build useful systems that matter for people doing software development, that matter for people doing writing, that expand the nature of the problems we solve. And then we'll see if we end up with AGI, but we will certainly end up with useful systems.

迈向模拟细胞? Towards a Simulated Cell?

Host

那么,生物学中最有用的系统呢?我的意思是,你有 DeepMind,你可能拥有所有这些针对人类生物学不同方面的系统,比如 AlphaFold、AlphaGenome、AlphaProteio 等等。你能把它们整合到一个系统中吗?我的意思是,这里是否有目标要构建一个模拟细胞?你知道,我以前从事模拟工作,模拟就是我会写下所有小部件局部做它们小事的规则,然后我把它们全部混合在一起,转动一个大曲柄,然后我就得到了结果。但我们甚至没有细胞的零件清单。我认为所有这些效应不会给我们一个经典模拟的模拟细胞。

So how about the most useful system of all in biology? I mean, you have DeepMind, you might have all of these different systems for lots of aspects of human biology, like AlphaFold, AlphaGenome, AlphaProteio, and so on. Can you bring those together in a single system? I mean, is there a goal here to sort of build a simulated cell? You know, I used to work in simulation, and simulation is that I will write down the rules for how all the little pieces do their little thing locally, and then I'll put it all mash it together and turn a big crank, and then I will get it. But we don't even have a parts list for the cell. You have all these effects that I think are not going to give us like a classical simulation simulated cell.

John

我认为没错。我认为模拟细胞的想法是一个非常引人注目的愿景,但我觉得我们离它还很远。我认为我们实现它的方式不是试图从第一原理模拟一切,而是构建能够预测关键属性和行为的模型,然后整合它们。这更像是构建一个相互对话的模型系统,而不是一个单一的庞大模拟。我认为这就是 AlphaFold 和其他工具的发展方向:它们提供了拼图的碎片,我们需要弄清楚如何有效地组合它们。

I think that's right. I think the idea of a simulated cell is a very compelling vision, but I think we're a long way from it. And I think the way we'll get there is not by trying to simulate everything from first principles, but by building models that can predict key properties and behaviors, and then integrating them. It's more like building a system of models that talk to each other, rather than a single monolithic simulation. And I think that's where AlphaFold and other tools are heading: they provide pieces of the puzzle, and we need to figure out how to combine them effectively.

将 AlphaFold 与 LLM 集成 Integrating AlphaFold with LLMs

John

我认为我们要做的是构建真正有用的系统,从 AlphaFold、文献和基因组中提取信息,并用它来讲述关于生物学的重要且有用的事情。我认为其核心技术之一可能是找到我们在狭义 AI 系统中的理解与我们在大型语言模型方面对广义机器学习的理解之间的正确融合。

I think what we're going to do is build really useful systems that draw information from AlphaFold, from the literature, from the genome, and use that to say really useful things about biology that matter. And I think quite possibly one of the core technologies of that will be finding the right fusion of what we understand in narrow AI systems and what we're understanding about broad machine learning in terms of large language models.

Host

那么,你是如何将这些系统结合起来的?是否有来自大型语言模型的想法可以应用?

So is that how do you bring those systems together? Are there ideas from large language models that can be applied?

John

很容易说我们只需让大型语言模型把 AlphaFold.exe 当作工具调用。但还有很多其他问题:如果 AlphaFold 生成一个结构,这些大型语言模型能真正很好地理解结构吗?它们能在多大程度上像人类一样理解这些 3D 坐标,甚至比人类更好?它们如何从 DNA 测序和其他来源引入信息?我认为要实现深度整合,让模型既像 AlphaFold 一样了解蛋白质和蛋白质结构,又理解整个生物学文献,绝非易事。我有点希望我们能实现,但我们必须去构建它。

It's very easy to say we'll just have your large language model call AlphaFold.exe as a tool. But there are all these other problems: if AlphaFold produces a structure, can these large language models actually understand structure really well? To what extent can they understand these 3D coordinates as well as a human, or better than a human? How do they bring in information from DNA sequencing and all these others? I think it's far from trivial how we get these deep integrations so that a model can understand as much about proteins and protein structure as AlphaFold, but also understand the entirety of the biology literature. I'm kind of hopeful we'll get there, but we have to build it.

生物学中难以计算的部分 Aspects of biology resistant to computation

Host

你认为生物学中是否有某些方面会抗拒计算预测?

Do you think there are aspects of biology that will resist computational prediction?

John

肯定会有一些方面。如果你问关于进化或生命起源的深刻问题,你从什么数据中学习?做什么实验?你必须从很远的数据中提取信息来回答这个问题。你可能了解一些化学知识,也许能更快地做这些实验,但你肯定不是直接从数据中学习。或者我们谈论进化,绘制系统发育树,但最终我们只有现存物种的 DNA 和一点点过去的信息。我认为这类事情会非常困难。然而,随着我们构建这些 AI 工具,合理假设的空间将会缩小。它会说‘可能不是那样,因为这个原因,可能也不是那样’,我们的实验会更好。在某种贝叶斯意义上,我们对合理生物学答案的先验会因计算工具而缩小,实验将帮助解决它们。这种相互作用会变得更紧密。随着我们做更多实验,或使用 AI 进行蛋白质设计等,获得更多工具来探测细胞,我们将学到更多,做得更多。但有些事情会更难,有些更容易,容易的事情会先发生。

There will certainly be aspects. If you ask deep questions about evolution, or the origin of life, what data are you learning from? What experiments? You're going to have to draw data very far away to answer that question. You might know something about chemistry. You might be able to do these experiments a bit faster, but you're certainly not directly learning from data. Or we talk about evolution and we draw phylogenetic trees, but ultimately we just have the DNA of the species that exist right now and a little bit into the past. These kinds of things I think will be very hard. However, as we build these AI tools, the space of reasonable hypotheses will narrow. It will say 'probably not that for this reason, probably not that' and our experiments will be better. In a certain Bayesian sense, our prior over what are reasonable biological answers will narrow because of our computational tools, and experiments will help resolve them. This interplay will get tighter. As we do more experiments or use AI for things like protein design that give us more tools to poke the cell, we will learn more and do more. But some things will be harder and some easier, and the easier things will happen first.

结束语 Closing remarks

Host

容易的事情会先发生。约翰,非常感谢你加入我们。

The easier things will happen first. John, thank you so much for joining us.

John

期待所有这些容易的事情实现。

Looking forward to all those easy things falling.

Host

我认为人们很容易对科学抱有非常浪漫的想法,认为科学是关于揭示宇宙隐藏的真理。作为研究者,你的目标是逐步构建这幅图景,以理解生命的机制。这正是让约翰关于可解释性的想法如此迷人的原因,因为它彻底颠覆了传统。AlphaFold 毫不掩饰地不关心‘为什么’。相反,它是一个可以可靠加速科学家工作的工具。当你想起约翰·詹珀的科学生涯才过半程,却已获得诺贝尔奖时,你会意识到他并非在捍卫旧范式;他正在构建下一个范式。如果约翰的焦点是实用性而非理解,那么当建造了 AI 有史以来最有用东西的人告诉你这才是重要的时,你不得不怀疑他是否在向我们展示科学未来的方向。您正在收听的是我与汉娜·弗莱教授共同主持的 Google DeepMind 播客。如果您喜欢本期节目,请留下评论或评分。接下来,我们将采访 Google DeepMind 的两位联合创始人德米斯·哈萨比斯和沙恩·莱格。相信我,您绝对不想错过。请订阅我们的 YouTube 频道。下次见。

I think it's really easy to have a very romantic idea of science, that it's about uncovering the hidden truths of the universe. Your aim as a researcher is to build this picture piece by piece to understand the mechanisms of life. That is what makes John's ideas about interpretability completely fascinating, because that is turning things on their head. AlphaFold is unashamedly not about the 'why'. Instead, it's a tool that can reliably accelerate the work scientists can do. When you remember that John Jumper is only halfway through his career as a scientist and already has a Nobel Prize, you realize he isn't defending an old paradigm; he is building the next one. If John's focus is on utility rather than understanding, when the person who built the most useful thing AI has ever done tells you that is what matters, you have to wonder if he's showing us where science is headed next. You have been listening to the Google DeepMind podcast with me, Professor Hannah Fry. If you enjoyed this episode, please leave a comment or review. Coming up, we have interviews with two of Google DeepMind's co-founders, Demis Hassabis and Shane Legg. Trust me, you will not want to miss them. Subscribe to our YouTube channel. See you soon.

互动版:逐字朗读 + 针对本期提问 →