AlphaFold, Nobel Prize, and the Future of AI in Biology
打开互动全文版(中英对照 + 朗读 + 问答)→诺贝尔奖得主 John Jumper 探讨 AlphaFold 的影响、开放数据的作用,以及 AI 如何从预测生物学转向设计生物学。
Nobel laureate John Jumper discusses AlphaFold's impact, the role of open-access data, and how AI is shifting from predicting biology to designing it.
约翰·江珀博士领导了构建 AlphaFold 的团队,这个 AI 模型通过解决一个长达数十年的难题重塑了分子生物学。他于 2024 年获得诺贝尔化学奖。现在,作为 Google DeepMind 的杰出科学家,他专注于下一个前沿:利用 AI 不仅预测生物学,还要设计生物学。我们在利物浦的 ISMB ECCB 大会上,这里计算生物学的前沿与 AI 的未来交汇。这里的讨论不仅仅是解决问题,而是重新构想我们提出的问题。我是 Leila Risby,和我一起的是我的联合主持人 Steven Horn,以及播客嘉宾 Luke Yates。各位,很高兴再次见到你们。
Dr. John Jumper led the team that built AlphaFold, an AI model that reshaped molecular biology by solving a decades old challenge. He was awarded the Nobel Prize in chemistry in 2024. Now, as a distinguished scientist at Google DeepMind, he's focused on the next frontier, using AI not just to predict biology, but to design it. We're in Liverpool at ISMB ECCB where the cutting edge of computational biology meets the future of AI. Here the talk isn't just about solving problems. It's about reimagining the questions we ask. I'm Leila Risby and I'm joined by Steven Horn, my co-host, and Luke Yates, a guest of the podcast. Guys, it's great to have you here again.
很高兴来到这里。很高兴来到利物浦。所以,AlphaFold,约翰·江珀,大讨论,了不起的嘉宾,你知道,显然是一位真正的开拓者,当然在 AlphaFold 项目中发挥了关键作用,昨晚的主题演讲非常精彩,真的相当惊人。
It's great to be here. It's great to be here in Liverpool. So, AlphaFold, John Jumper, big discussion, amazing guest, you know, obviously a real trailblazer, very instrumental in the AlphaFold project, of course, and really fascinating keynote lecture last night. It was really rather amazing.
所以,这是你很了解的事情,对吧?那么,你想问他哪一个问题?你昨晚可能一直在想这个问题,对吧?你在想,我要问他什么?那么,你想问他哪一个问题?
So, this is something you know a lot about, right? So, what's the one question you want to put to him? You probably been thinking of this last night, weren't you? You were thinking, what am I going to hit him with? So, what's the one question you want to ask him?
嗯,作为一名结构生物学家,我特别为这个社区感到自豪,我们建立了一个非常大的开放获取的蛋白质结构库,这些结构是通过实验确定的。我们严格守护着它,确保提交的结构达到一定标准,当然,这对 AI 的训练模型至关重要,特别是对 AlphaFold,用于从蛋白质序列生成这些预测。我想知道这类数据对构建大型语言模型有多重要。
Well, I think as a structural biologist, I'm particularly proud of the community, the fact that we generated a very large open access repository of protein structures which are experimentally determined. We've jealously guarded that in terms of making sure the structures deposited reach certain criteria and of course that was instrumental as a training model for AI for AlphaFold in particular for generating these predictions from protein sequences. And I guess I kind of want to know how instrumental is this type of data for building large language models.
Luke,你能简单介绍一下 AlphaFold 是什么以及为什么它如此重要吗?
Luke, can you just give us a little bit of an overview of what AlphaFold is and why it was so monumental?
所以,在我看来,解析蛋白质结构对于理解生命如何运作至关重要。蛋白质是生命的基石。它们在我们的细胞和身体中工作,帮助我们维持生命。我们拥有大量的遗传数据,知道所有生命王国中许多蛋白质的序列,一直到海底的微生物。但挑战之一是生成结构需要在实验室中花费数月甚至数年的专门工作。因此,有一个古老的问题,AlphaFold 解决的挑战是:我们能否仅根据氨基酸序列准确预测三维结构?答案是肯定的,相当准确。而 AlphaFold 就是那个突破,使我们能够以极短的时间揭示并完成这类工作。所以我们不再需要在实验室中进行艰苦的工作。我们可以在计算机上生成结构。
So, in my view, solving protein structures is really essential to understand how life works. Proteins are the building blocks of life. They do the work in our cells, in our bodies to help keep us alive. And we have a huge amount of genetic data and we know all the sequences of lots of proteins across all kingdoms of life right through to microbes at the bottom of the ocean. But one of the challenges is that generating that structure requires months and years of dedicated work in a laboratory. And so there is an age-old problem that the challenge that AlphaFold solved was: can we predict accurately the three-dimensional structure based only on the amino acid sequence? And the answer is yes, quite accurately. And AlphaFold was that breakthrough that allowed us to uncover and to be able to do that kind of work in a fraction of the time. And so we don't need to do the laborious work in the laboratory. We can generate the structures on a computer.
Stephen,你对这个领域最感兴趣的是什么?我知道你不像 Luke 那样是科学家,但这是 Google DeepMind。他们在甚至没有涉足的领域取得了巨大进步,并且彻底击败了对手。那么,你怎么看?
And Stephen, what's so interesting to you about this area? I mean, I know you're not a scientist like Luke, but this is Google DeepMind. They made huge strides in a field that they weren't even involved in and they wiped the floor clean. So, what do you think of that?
我怎么看?嗯,我认为这完全是关于创新,不是吗?我认为是关于那种专注和驱动力,它带到了学术领域,如果你愿意这么说的话,令人着迷的是,你如何用那种驱动力,用那个糟糕的美国词,'货币化',对吧,科学和技术。所以我真的很着迷。我也很着迷看到一位杰出的科学家在工业界。你知道,我们做这行 20 年了,大多数公司很难让杰出的科学家为他们工作,因为他们想在所谓的公共科学领域工作。所以,我很想知道诺贝尔奖得主约翰如何处理这个问题。
What do I think of that? Well, I think it's all about innovation, isn't it? And I think it's about that sort of focus and the sort of drive that brings to fields which are academic fields if you like, which is just fascinating how you can, in that awful American word, monetize, right, sort of science and technology with that drive. So, I'm really fascinated to see. I'm also fascinated to see a prominent scientist in industry. You know, we've been doing this for 20 years and most companies struggle to get prominent scientists to work for them because they want to work in what's known as public science. So, be really interested to see how John Nobel Prize winner deals with that.
当然。是的。我也对稍微宏观一点的视角感兴趣,显然 AlphaFold 出现的时间比过去两三年我们看到的 ChatGPT 革命要早得多,所以我认为它有点像我们预期在其他领域、其他行业将会发生的事情的缩影。AlphaFold 刚出现时,被认为会取代结构生物学家的职位,他们想,我们到底要做什么?这东西能为我们做一切。但他们并没有失业。我们在 ISCB,还有很多结构生物学家、分子生物学家仍在工作。所以我认为这表明当 AI 革命来到其他领域时,我们还有很多事情要做。
Sure. Yeah. And I'm also interested in zooming out a little bit and obviously AlphaFold came about quite some time ago compared to the chat GPT revolution we've seen in the last 2-3 years and so I think it's kind of like a microcosm of what we are to expect to come in other fields in other industries where AlphaFold when it first came about was thought to take the jobs of structural biologists they were thinking what on earth are we going to do? This thing can do everything for us. But they're not out of a job. We're here at ISCB and there's lots of structural biologists, molecular biologists still working away. So I think it shows that there's going to be a lot for us to still do when the AI revolution comes to other areas.
嗯,期待见到约翰。我认为这将是一次非凡的讨论。
Well, looking forward to meeting John. I think it will be an extraordinary discussion.
我非常兴奋。我想他就在那边。
I'm very excited. I think he's just over there.
他就在那里。现在我们很高兴请到了 DeepMind 的约翰·江珀博士。约翰,很高兴你来到这里。
He's just there. We're now delighted to be joined by Dr. John Jumper from DeepMind. John, it's great to have you here.
哦,很高兴来到这里。
Oh, it's wonderful to be here.
那么,我想从利物浦的这次会议开始。是什么吸引你来到 ISMB ECCB?
So, I'd like to start off with this meeting at Liverpool. What drew you to ISMB ECCB?
我的意思是,这是一个大型会议。令我惊讶的是这里有这么多我认识的真正伟大的科学家。另一个原因当然是珍妮特·桑顿女爵士,她在诺贝尔奖之后立即找到我,说你必须来 ISMB,我不会对她说不,但这是一个很棒的计算社区会议,看到它真的很令人兴奋,老实说,对我来说也很近。我从伦敦过来,所以这是一个很棒的会议,也是一个美丽的地方。
I mean, this is a huge conference. What's amazing to me is how many people here, really great scientists that I know. I mean the other is of course Dame Janet Thornton who comes to me right after the Nobel and says you have to come to ISMB and I'm not going to tell her no but it's a wonderful meeting of the computational community and I it's really exciting to see it's honestly also local for me. I came up from London so it's just a wonderful meeting and a beautiful place.
我想这里有很多科学家使用过你的工具 AlphaFold。那么你能告诉我们 AlphaFold 是如何工作的吗?
And I imagine there's plenty of scientists here who have used your tool AlphaFold. So can you tell us a bit about how AlphaFold works?
所以 AlphaFold 是一个 AI 工具,它预测一个非常具体实验的结果。这就是确定蛋白质的结构。这是实验室生物学家经常做的实验。每年大约确定 11,000 到 12,000 个新结构。但每个结构代表科学家大约一年的工作来确定。这真的是巨大的工作量。这些数据在过去大约 50 年里被收集到蛋白质数据库(Protein Data Bank)中,这是一个实验结果的存储库。我们能够开发新的 AI 方法,这些方法在预测这个实验的结果、从序列预测蛋白质结构方面效果要好得多。这是生物学家一直用来规划下一步实验、提出疾病假说、尝试开发针对特定蛋白质的药物。所以我们提供的系统已经被全球大量科学家使用,这非常令人振奋。我无法告诉你,即使在这个会议上,有多少次有人走过来对我说,我们每天都用 AlphaFold,并且对此非常兴奋。
So AlphaFold is this AI tool that predicts the results of a very specific experiment. This is determining the structure of a protein. And this is an experiment that lab biologists do commonly. About 11-12,000 new structures are determined every year. But each one represents about a year of work for a scientist to determine. It's really an enormous amount of work. And these data have been collected kind of over the last about 50 years into the Protein Data Bank, this repository of these experimental results. And we were able to develop new AI methods that work tremendously better at predicting the results of this experiment, predicting the structure of a protein from its sequence. This is something that biologists use all the time in order to plan their next experiments, develop hypotheses about disease, try and develop drugs that target a specific protein. And so the system that we made available has been used really worldwide by huge numbers of scientists and it's so heartening. I can't tell you how many times someone will come up to me even at this conference and say we use AlphaFold every day and just so excited about it.
这也非常重要,因为科学工作非常困难。实验科学需要投入大量工作,能够构建一个帮助科学家更快完成工作的工具真是太棒了。
It's so important also because the work of science is so difficult. There's so much work that goes into experimental science and it's really wonderful to be able to build a tool that helps scientists do it faster.
不,我想接着这个话题说,我认为它彻底改变了许多研究人员的生活。我的意思是,我每天都在实验室里使用它。但我想问的是,你有没有注意到一些例子,他们以非常有趣的方式使用了 AlphaFold,引起了你的注意?
No, I think just to carry on from that, I think it has revolutionized a lot of researchers' lives. I mean, I use it in my lab every day. But I guess what I wanted to ask really, have there been any examples that you've paid attention to where they've used AlphaFold in a really interesting way that caught your attention?
这样的例子太多了。我觉得有一个特别有趣:我们原本知道 AlphaFold,就像我刚才讲的故事一样,AlphaFold 提供对某种实验的预测。但有趣的是,人们开始用我们不知道它能做的方式做事。我发现一个非常引人注目的例子是,人们使用 AlphaFold 时会说:‘好吧,我不知道哪些蛋白质会结合。我应该预测什么?为什么不预测所有东西?为什么不运行 AlphaFold 2000 次,看看会得到什么?’最近确实有一个相当高调的结果,关于精子和卵子如何结合以及参与该相互作用复合体的蛋白质。人们运行了 2000 次 AlphaFold 预测,将精子表面的每种蛋白质与卵子上已知的蛋白质进行比对。他们发现了一种特定的蛋白质,AlphaFold 可以预测其结构。然后他们通过实验证实了这一点。事实上,这个结果是由两个计算小组独立发现的,其中一个小组通过实验验证,这是一个足够高调的结果。我记得我在《纽约时报》上读到过,文章提到了 AlphaFold,而我们除了构建这个工具之外与此无关,但人们却能在《纽约时报》上发表文章,不是关于 AI,而是关于使用我们的工具在生物学上的新发现,这真的让我感到非常温暖。
There have been so many. I think one that's been really interesting for me is we kind of knew AlphaFold, you know, the story I just told, right? AlphaFold provides a prediction of a certain type of experiment. But what's interesting is people have started to find ways to do stuff that we didn't know it would do. One that I found really compelling is that people will use AlphaFold. They'll say, 'Well, I don't know which proteins are coming together. What should I predict? Why don't I predict everything? Why don't I run AlphaFold 2,000 times and see what comes out?' And there was actually a pretty high-profile result recently how egg and sperm come together and the proteins involved in that interaction complex. And people ran 2,000 AlphaFold predictions, every protein on the surface of sperm against these known proteins on the egg. And they found one particular protein that AlphaFold could say this was the structure. And then they confirmed it experimentally. In fact, it was discovered independently by two groups computationally and one showed experimentally and it was a high enough profile result. I remember I read about it in the New York Times and it mentioned AlphaFold and we had nothing to do with this other than building the tool and here are people getting articles in the New York Times that isn't about AI it's about this new discovery in biology using our tools and that just absolutely kind of warms my heart.
那么我想问一个关于蛋白质数据库的问题。你多次提到它的重要作用,作为我们社区的一部分,我们一直小心翼翼地守护和把关它。我想问的是,蛋白质数据库和那个实验数据存储库对于 AlphaFold 的开发和突破有多重要?
So I just wanted to come on to a question about the Protein Data Bank. You've mentioned a few times how instrumental it's been and as a part of our community we've guarded it and gatekept it rather jealously I guess. And I wanted to ask how important was the Protein Data Bank and that repository of experimental data for the development and the breakthrough that came with AlphaFold?
PDB 非常重要。AI 当然是一个将两件事结合在一起的工具。你编写一个计算机程序或数学公式,但它实际上变成了一个计算机程序,这给了你学习的骨架。然后你引入数据,这实际上把它填充成有用的东西。你应该把机器学习或 AI 系统看作是代码加数据产生有用的东西。因此,拥有这些数据是绝对必要的。实验界的远见卓识在于他们收集了这些数据,你不能在没有让 PDB 接受你的结构的情况下发表期刊文章,这种高质量的把关和策展确保了数据的可用性和可靠性。当然,很有趣的是,如果你和计算生物学家交谈,他们会说 PDB 是我们拥有的最干净的数据源。PDB 之后我们该怎么办?而你和实验学家交谈,他们会说:‘哦,但这是错的,那是错的,那是建立在希望之上的。’看到这两种观点很有趣,但两者都是对的。但我认为这确实证明了实验界的远见,他们以一种我们在许多领域都没有做到的方式团结在一起。也许蛋白质序列和 UniProt 等是另一个很好的例子。但我认为展望未来,我们的数据也必须跟上科学的变化。我认为结构生物学中最令人兴奋的一些问题是原位测量的问题。我们如何收集断层扫描?我们如何收集——对于不知道的人来说,这些是快速冷冻的整个细胞切片,这样你就能看到蛋白质在其天然环境中的样子。我认为一个非常有趣的问题是,什么将成为细胞切片的 PDB 等效物,可能用于不同形式的显微镜。我认为这些是一些非常有趣的问题,如果我们处理得当,将产生让我们的生活更美好的 AI 模型。如果我们必须从一千个略有不同的来源中提取数据,也许我们就不会成功。我认为数据策展有很大的机会在未来影响我们拥有的 AI 系统。
The PDB was so important. The AI is of course a tool that really brings together two things. You write a computer program or a mathematical formula, but really it becomes a computer program and this gives you the skeleton of learning. Then you bring data and that actually fills it out into something that works. You should think of machine learning or AI systems as code plus data yields something hopefully useful. So it's absolutely essential to have these data. It's really the foresight of the experimental community that they collected these data, that you can't publish a journal article without getting the PDB to accept your structure, that there's very high quality gatekeeping and curation that ensures that the data are available, that are reliable. Of course, it's very funny if you talk to computational biologists, they'll say the PDB is the greatest cleanest source of data we have. What are we going to do after the PDB? And you talk to experimentalists and they say, 'Oh, but this is wrong and that's wrong and that was built on hope.' It's very funny to see these two, but they're both true. But I think it's really a testament to the foresight of the experimental community that they came together in a way that we haven't in a lot of areas. Maybe protein sequences and UniProt and others are another great example. I think going forward though, our data has to also keep track with how our science is changing. I think some of the most exciting questions in structural biology are the questions of in situ measurement. How are we going to collect tomograms? How are we going to collect, for people who don't know, these are flash-frozen slices of whole cells so you get to see the proteins in their native environment. I think one really interesting question is what is going to be the equivalent of the PDB for slices of cells, possibly for different forms of microscopy. I think these are some really interesting questions that if we get them right will yield AI models that make our lives better. And if we have to pull it from a thousand sources that are all somewhat different, maybe we won't. I think there's a real opportunity for data curation to make a difference in the AI systems we have in the future.
我能问你几个问题,更多是关于你的职业生涯吗?作为科学家,我们长期以来对物理学界进行了很多采访,我一直感到惋惜。工业界总是抱怨他们很难招募到最优秀的科学家,因为最优秀的科学家想在公共科学领域工作。你认为你的奖项会以某种方式帮助招募科学家进入工业界吗?
Can I ask you a couple of questions pulling more towards your career? As a scientist, we've done a lot of interviews over a long period with the physics community and it's always been bemoaning to me. Industry is always bemoaning that they really struggle to recruit the best scientists because the best scientists want to work in public science. Do you think your award is going to help with that in a way of recruiting scientists into industry?
我认为这个奖项的一个很好的地方是,它承认了顶尖科学可以来自工业界,并且这些科学对科学家的工作产生了直接的影响。我认为情况也是如此,尤其是在像 AI 这样的计算领域,算力很重要。这是做实验的能力,对吧?GPU 在某种意义上就是我们的试剂。所以我认为工业实验室在组织算力和组织团队方面发挥了重要作用。我认为工业研究中一个很大的机会是我们可以组建一个团队。AlphaFold 是一个大约 15 人的团队,对吧?我们可以把这些人放在一起,他们一起成功或失败,没有第一作者的问题,不需要分配功劳。你确实会在内部分配,说这个人做了这个,那个人做了那个,但作为一个团队,你们一起成功或失败,这支持了专业化,支持了工作。
I think one thing that is very nice about the award is it's an acknowledgement that top-tier science can come out of industry and science that makes a really direct impact on the work of scientists. I think it's also the case, especially when you look in a field like AI computationally, computational power is important. It's the power to do the experiments, right? GPUs are our reagents in a certain sense. So I think industrial labs have had this big role to play in organizing the compute and organizing the teams. I think one thing that is really a great opportunity in industrial research is that we could put together a team. AlphaFold was a team of about 15, right? We can put that group of people that are succeeding or failing together that don't have the first author problem, that aren't trying to parcel out. You do parcel out internally and you say this person made this, that person made that, but as a group you succeed or fail and that supports specialization, that supports work.
你提到数据极其重要,计算架构是 AI 方程的另一部分。每个人都能访问蛋白质数据库。那么你的方法中是什么放大了这一点,让你从 AlphaFold 中获得那样的价值?
You mentioned data being incredibly important and computing architectures being another part of the equation when it comes to AI. Everyone had access to the Protein Data Bank. So what was it about your approach that amplified that for you to get that value out of AlphaFold?
好问题。我喜欢说 AI 有三个要素:数据、算力和研究——也就是这个计算机程序是什么的问题。对于 PDB,每个人都能访问相同的数据。实际上,我们在 2018 年左右构建了 AlphaFold 1,在 2020 年左右构建了 AlphaFold 2,两者之间数据没有变化。我们使用了完全相同的训练数据,但学到的却多得多。Alcareshi 实验室的一项研究用 1%的可用数据重新训练了 AlphaFold 2 模型,发现其准确度与 AlphaFold 1 相当甚至更高。所以我们的研究价值大约是数据的 100 倍。研究让你从每个珍贵的实验数据点中获取更多信息。我们有全新的方法和极高速度的团队文化——有人提出新想法,同周内就有人在此基础上发展。相比之下,学术界的分布式周期是论文发表后 3-6 个月才在其他系统中测试,而我们则是下周。我可以讲具体的架构创新,但有一种迷思认为 AI 是现成的东西,只要有足够算力扔给数据就行。这不是真的。研究是极其重要的一部分,尤其是在蛋白质机器学习领域,这个交叉学科之前并未被认真对待。我们认真对待科学:如何提出新想法、测试想法、在开发架构变化时运用科学方法。我们以极快的速度做到了这一点。现代 AI 中有大量的研究和突破,不仅仅是顶层想法,还有所有中间的小想法,它们累积起来构成了变革性的系统。
That's a great question. I like to say there are three ingredients of AI: data, compute, and research—the question of what is this computer program? With the PDB, everyone had access to the same data. In fact, we built AlphaFold 1 around 2018 and AlphaFold 2 around 2020, and we did not change the data between them. We kept the exact same training data, yet we learned tremendously more. There was a study by the Alcareshi lab that retrained an AlphaFold 2 model on 1% of the available data and found it was as accurate or more accurate than AlphaFold 1. So the research we did was worth about 100 times the data. Research lets you get more information out of each precious experimental point. We had very new approaches and a team culture of incredibly high velocity—someone would have a new idea and someone would build on it within the same week. Compared to academia's distributed cycle where a paper comes out and is tested 3-6 months later, for us it was next week. I can tell you specific architectural innovations, but it's a meme that AI is something you take off the shelf and throw at data with enough computers. That's not true. Research is an enormous part, especially in protein ML where the intersection hadn't been taken seriously. We took seriously the science: how to come up with new ideas, test them, use the scientific method as we developed architectural changes. We did this incredibly quickly. There's a tremendous amount of research and breakthroughs in modern AI, not just top-level ideas but all the little ideas in between that add up to a transformative system.
你从那种在专利局提出疯狂理论的孤独天才,发展到了现在的协作团队,由于跨学科的需求,这是科学的未来。
You've gone from this idea of lone geniuses coming up with crazy theories in patent offices to now collaborative teams, which is the future of science because of this interdisciplinary need.
我认为两者都有空间。当你有明确的目标且大家都能达成一致时,团队合作尤其有效。但一个变化是,你需要根据当前最先进水平来测试你的想法。在 AlphaFold 1 和 AlphaFold 2 之间,我们在 GDT 尺度上可能提高了 30 分。单个最大增益的想法价值 2.5 分;大多数是 1 分或更少。所以我们是在一步步攀登。我不认为这是增量式的或工程性的——有些想法很出色,但即使出色的想法也只值一点,你需要很多这样的想法。我们喜欢讲述孤独想法带来成功的故事,但即使在历史上,也是更广泛的科学家群体在彼此想法的基础上构建,才达到今天的高度。
I think there's room for both. Teams work especially well when you have a clear objective that everyone can agree on. But one thing that changes is you need to test your ideas against the state-of-the-art. Between AlphaFold 1 and AlphaFold 2, we were maybe 30 points better on the GDT scale. The single idea that gave the largest gain was worth two and a half points; most were one or less. So we were climbing our way forward. I don't think it's incremental or engineering—some ideas are brilliant, but even brilliant ideas are worth a bit, and you need many of them. We love to tell stories of lone ideas leading to success, but even in history, a wider group of scientists built on each other's ideas to bring us where we are today.
你谈到成为公共知识分子,并且这要求做正确的事。能详细说说吗?
You've talked about being a public intellectual and that it has a demand to do the right thing. Can you tell me more about that?
诺贝尔奖很棒。我告诉年轻学生我会推荐它。它是科学的象征,也是计算生物学前景的象征——我们可以在计算机上做对实验科学家重要的工作。随之而来的是许多请求:与年轻科学家交谈、与权威科学机构交谈、与政府交谈。总会有邀请你以伟大知识分子的身份发言。
The Nobel Prize is wonderful. I tell young students I would recommend it. It's a symbol of science and the promise of computational biology—that we can do work on computers that matters for experimental scientists. It comes with many requests: talking to young scientists, august bodies of science, governments. There's always an invitation to speak as a grand intellectual.
嗯,当然你获得诺贝尔奖是因为在科学、理解计算机架构和蛋白质机器学习方面的贡献,而不是因为政府政策或其他什么。所以我认为总有一种责任感,你想利用这个平台在力所能及的时候做好事。当有机会参加一个精彩的活动时,你通常应该尽量去做。同时,你应该记住你也是人,这个奖项是对一项成就的认可。我在结构预测算法的设计上有很多话要说,我可能对科学组织有所贡献,但这需要带着更多的谦逊,而不是谈论蛋白质结构预测架构的细节。所以我认为这真的是在机遇与谦逊之间取得平衡:你真正知道什么?你能为这些重大的问题真正贡献什么?
Well, of course you received the Nobel Prize for a bit of science and understanding computer architecture and machine learning of proteins, not for government policy or anything else. So I think there's always this responsibility that you want to use that platform to do good when you can. When there's an opportunity to come to a wonderful event, you should generally try and do it. At the same time, you should remember that you too are human, that this prize recognizes an accomplishment. I have a lot to say in the design of structure prediction algorithms, I may have something to contribute to the organization of science, but it needs to be done with a lot more humility than I talk about details of protein structure prediction architecture. So I think it's really balancing the opportunity with the humility of what do you really know, what can you really contribute to these big important questions.
正如你所说,你最重要的工作是当父亲,对吧?
And as you said, your most important job being the father, right?
哦,非常是。这是另一回事。我有三个年幼的孩子,6 岁、9 岁和 11 岁。我想确保他们不会因为我在外面忙于做重要的事情而缺少一个积极、投入的父亲。我认为要把握好所有这些平衡,而且身处这个仍然发展迅速、仍然非常重要的研究领域,你想做出贡献。我很大一部分——我妻子说如果由我决定,我会躲起来只做科学,这确实很真实。但我想所有这些,只是试图找到正确的平衡。诺贝尔奖是一个美妙的认可,表明我的职业生涯为世界做了一些好事,我希望我职业生涯的剩余部分也能为世界做一些好事。
Oh, very much. And that's the other thing. I have three young kids, you know, 6, 9, and 11. And I want to be incredibly sure that they are not shortchanged of having an active, engaged father because I'm off busy doing important things. I think getting all that balance right, and being in this research area that's still incredibly quickly developing, still incredibly important, and you want to contribute. There's a good portion of me—my wife says if it was just up to me, I would just hide away and do science, and it's pretty true. But I think all of those and just trying to find the right balance. The Nobel is a wonderful acknowledgement that my career has done some good for the world, and I would like the remainder of my career to do some good for the world too.
那么继续谈到你职业生涯的剩余部分,接下来是什么?你现在希望回答的下一个大问题是什么?AlphaFold 在那个飞跃中如此重要。
So just moving on to the remainder of your career, what is next? What's the next big question that you're hoping to answer now? AlphaFold has been so substantive in that leap forward.
我认为 AlphaFold 和类似的技术将继续成长和发展。我实际上非常感兴趣的是,我们能否找到下一个问题?我能否找到下一个看起来不太可能但会发生的事情?老实说,大约一年前我开始思考这个问题:我们正在开发语言模型、聊天机器人,它们显然能够阅读和写作科学内容。它们显然有点理解。我记得我第一次感到惊讶的经历之一:我给它一篇论文的上下文,说‘列出这篇论文中所有实际进行的实验’,它做得很好。我心想,这并不简单,要阅读一篇密集的科学论文并提取出来。但真正的问题是,如果你问它们一个关于这个实验会发生什么或其他你非常了解的详细问题,它们往往做得不太好。我对这个问题非常感兴趣:什么能让我们实现生物推理?我们能在多大程度上让语言模型深入理解科学,像科学家那样与之互动?我不会过多谈论我们如何尝试做到这一点,因为我还没有成功。我得保留一些秘密。但我对如何发展推理这个问题非常感兴趣,因为有很多问题没有 PDB。事实上,即使你收集了文献中的所有数据,也不会很多。所以我们真的必须学习。机器学习的伟大技巧是,把你感兴趣的问题和一些看起来大部分不相关的数据,让它们变得足够相关。拿一个你想要结构的蛋白质,找到另一个完全不同家族的蛋白质,从中学习一点,然后应用到这里。这真的很神奇。我们如何在科学推理中做到这一点?我认为这是一个重大问题,一年前我们开始尝试时,人们对此的怀疑比今天多得多,当然是由于这些技术的快速发展,但前方还有巨大的挑战。
I think AlphaFold and technologies like it will continue to grow and develop. I'm actually very interested in, can we find the next question? Can I find the next thing that will happen that doesn't look as plausible now? Honestly, about a year ago I started to get into this question of we're developing language models, chatbots, things that can clearly read and write science. They clearly kind of understand. I remember one of my first surprising experiences: I just gave it a paper in the context and said, 'List me all the actual experiments performed in this paper,' and it did a fine job. I thought to myself, that's non-trivial, to read a dense scientific paper and pull that out. But the real question is, if you ask them a detailed question about what will happen in this experiment or something else that you really know well, they tend not to do so well. I'm very interested in this question: what gets us to biological reasoning? To what extent can we make language models understand science in a deep way, engage with it as a scientist perhaps does? I won't say too much about how we're trying to do it because I haven't succeeded yet. I got to keep some secrets. But I'm very interested in this question of how we are going to develop reasoning, because there are so many problems that don't have a PDB. In fact, even if you collected all the data in the literature, it wouldn't be much. So we're really going to have to learn. The grand trick of machine learning is to take a question you're interested in and take some data that looks mostly unrelated and make it just related enough. Take a protein that you want the structure of and find another protein in a completely different family and learn a bit from that that you can apply here. That's really what is amazing. How do we do this in scientific reasoning? I think it's one of these grand problems that people expressed much more doubt when we started playing with this a year ago than they do today, of course due to the rapid development of these technologies, but there are enormous challenges ahead.
是的,当然。我可以想象,如果你得到一个足够好的模型,那么它可能能够突出我们知识中的空白和需要探索的地方。我想那会是最终目标。
Yeah, of course. And I can imagine if you get a model good enough, then I guess it could highlight the gaps in our knowledge and places to go looking. I guess that would be the ultimate aim.
我的意思是,如果你真的擅长理解这一点,你可以做很多很多事情。我们会有知识上的空白。我们也会对实验所暗示的内容有所理解,能够找到真正相关的文献。我们可以做很多事情,但我们需要坐下来,不只是梦想,而是找到从我们目前的位置到这些工具真正有效和值得信赖的未来之间的局部路径。
I mean, there are many, many things that you can do if you actually get good at understanding this. We will have gaps in our knowledge. We will have some understanding also of what's implied by their experiments, ability to find really relevant literature. There's much that we can do, but we need to sit down and not just dream, but find the local part of how do we get from where we are to this future when these tools are really effective and trustworthy.
所以,在你做过的一些采访中,我发现非常有趣的一点是,你谈到了 AI 的速度和——我不会说 AI 的进化,因为可能甚至不是进化——这引出了两个问题。一个是,你认为中期内事情会如何发展,但还有公众信任。公众信任已经不足。大多数人今天不理解我们在谈论什么,而且它越来越快,越来越多的事情被完成。我们需要做什么才能让人们信任 AI?
So one of the things that I found very interesting again in some of the interviews that you've done, you talk about the speed of AI and the speed of—I wouldn't say evolution of AI because potentially not even evolution—and that brings me to two questions. One is, where you see things going in the medium term, but also public trust. There's already a shortage of public trust. Most people don't understand what we're talking about here today and it goes faster and faster and faster and more gets done. How—what do we need to do to get people to trust AI?
就信任而言,我个人在 AlphaFold 上的经历非常有趣。我们在很多方面都不确定。即使当我们的 CASP 结果出来时,那是一个公开评估,表明在盲测中我们可以预测结构。我仍然认为实验界有很多怀疑:‘嗯,也许那些更容易,也许……’直到大约六个月后,我们的论文发表,代码可用,预测可用,人们才真正开始——预测在网站上可用,所以他们说,‘我就查一下。’很多人已经解析了实验结构但还没有存入 PDB,他们说,‘为了好玩,我去看看。’我记得有人在推特上说,‘他们怎么得到我的结构副本的?’对吧?‘你是通过 Gmail 发送的吗?’社区内完全不敢相信这些方法突然之间就真的相关了。
Really in terms of trust, it's been a very interesting experience for me personally with AlphaFold. We weren't sure in a lot of ways. Even when our CASP results came out, that was this public assessment that said that blinded we could predict structures. I still think there was a lot of doubt in the experimental community: 'Well, maybe those are easier, maybe...' It wasn't until about 6 months later when our paper came out and our code was available and predictions were available that people really started—the predictions were available on a website, so they said, 'I'll just check.' And a lot of people had solved an experimental structure and hadn't yet deposited in PDB, and they said, 'Just for fun, I'll go see.' I remember someone on Twitter saying, 'How did they get a copy of my structure?' Right? 'Did you send it by Gmail or something?' It was absolute disbelief within the community that suddenly these methods all at once were really relevant.
我记得当我们发布这个时,我们正在与 Imble EBI 和 Samir 等人讨论如何呈现这些信息。我们担心的不是人们不信任我们。我们觉得一旦他们真正接触到预测,他们会花一些时间——我以为会比实际花的时间更长——但人们会学会它。我们也担心他们会过度信任。如何让他们以恰到好处的方式怀疑?我们投入了大量精力开发所谓的置信度度量,让模型预测它在各个维度上离答案有多远。我认为我们做的最好的决定之一是采用了我们开发的数值置信度度量 pLDDT,范围是 0 到 100。我们决定定义四个阈值:高于 90,我们将结构涂成深蓝色;低于 50,涂成红色表示危险。这四个阈值非常清晰地对应了人们应该如何处理:相当信任、可能信任、可能不信任、绝对不信任。实验社区对此接受得很好。我觉得他们真的说:‘嗯,这些是不完美的。我们知道它们有错误,部分是因为人们看到了,但我们也知道错误往往与模型表示不信任时高度相关。’然后他们会查看、思考、设计下一个实验、测试预测的各个方面,并很快将其整合到科学实践中,以大致正确的方式信任它。我认为这之所以发生得这么快,是因为科学界总是处理不确定性。他们从不完全信任另一个实验,也从不完全信任自己的实验。他们总是在思考各种可能性。我的意思是,你是一名实验科学家,你知道——
And I remember when we were putting this out and we were actually talking with Imble EBI and Samir and others like how do we present this information? We were worried about—we weren't as worried that people wouldn't trust us. We kind of figured once they really had access to the predictions it would take some time. I thought it would take more time than it did, but people would learn it. We also worried they would trust it too much. How would we get them to doubt it just the right amount? And we had put a lot of effort by this point into what we would call confidence measures where the model predicts how far it is from the answer in various dimensions. And I think one of the better decisions we made was we took this kind of numeric confidence measure called pLDDT that we had developed. It was 0 to 100. And we decided to just define four cutoffs. Above 90, we would color the structure dark blue. Less than 50 would be red for danger. And we had these four things that were very clearly aligned with what people should do with this: probably trust it quite a bit, probably trust it, probably don't trust it, definitely don't trust it. And the experimental community took to this and others really well. I felt like they really said, 'Well, these are imperfect. We know they have errors, in part because people have seen them, but also we know that the errors tend to be well correlated with when the model says it doesn't trust it.' And then they'll look at it, think about it, design their next experiment, test aspects of this prediction, and they very quickly integrate it into scientific practice, trusting it in about the right way. And I think the reason this happens so quickly is the scientific community always deals with uncertainty. They never fully trust another experiment. They never fully trust their own experiments. They're always thinking about all the ways. I mean, you're an experimentalist and you know—
是的,绝对如此。我们有时连自己的手都不信任。
Yeah. Absolutely. We don't trust our own hands sometimes.
正是如此。所以他们习惯了这种策略性的怀疑。以仍然能让他们取得进展的方式怀疑。足够的怀疑使他们不被误导,同时又有足够的信任使他们能继续前进。这与科学家的工作方式很好地融合了。我认为 AI 的更大问题是,人们会逐渐理解大型语言模型的输出意味着什么,它们有什么用处,什么没有用处。我个人看到了一些相似之处。我出生于 1985 年,所以我的变革性经历是在 90 年代,当时互联网是新的。有‘你怎么能信任互联网?任何人都可以在互联网上发布东西’这样的说法。当然这是真的,互联网上有很多愚蠢的话,但我们也学会了如何使用和驾驭这个工具,如何思考。我记得有人警告过关于维基百科的事。‘任何人都可以编辑维基百科’是 90 年代每个学生被告知的。我们学会了驾驭它,并从中获得了很多用处。我认为我们会看到类似的情况。AI 迅速成为一股重要的力量,社会将适应它的优点和缺点。这不会总是容易的,但我认为这会发生。希望我们能够开发更多的度量,特别是在科学领域,置信度度量。我认为我们在 AlphaFold 中的这个想法——了解模型置信度的方法就是让它预测自己会有多错——本身就是一个有趣的事实。这是一件奇怪的事:你可以让一个模型告诉你它错了,而不告诉你正确答案是什么。我可以给出一个人类类比:你经常在考试结束后知道自己考得好不好,而不知道所有答案。但这非常有趣。我认为我们会看到它随着时间的推移而发展,但世界和科学系统也会调整。我当然看到过 AlphaFold 在科学中的一些糟糕使用,但绝大多数情况下,我认为人们使用得很好,开发了正确的技术,并在社区中分享。实际上是社区学会了如何使用这些工具来更好地完成工作。
Exactly. And so they're used to this kind of strategic doubt. Doubting in just the way that still lets them make progress. Just enough skepticism that they aren't misled and yet enough trust that they can still make progress. And that integrated really well into how scientists work. And I think the larger question of AI is people developing over time this understanding of what these outputs of a large language model mean, what they are useful for and what they're not. I personally see some parallels. I was born in 1985, so my transformative experiences were in the '90s when the internet was new. There was this 'how can you trust the internet? Anyone can put something on the internet.' And of course that's true, and there's a lot of really dumb things said on the internet, but also we learned how to use and navigate this tool, how to think. I remember being warned about Wikipedia. 'Anyone can edit Wikipedia' was what they would tell every school student in the '90s. And we learned to navigate that and still obtain a lot of use out of it. And I think we'll see something similar. AI became a really relevant force very rapidly, and society will adjust to its strengths and weaknesses. It won't always be easy, but I think this will come. And hopefully we'll be able to develop more measures, certainly in the scientific space, confidence measures. And I think this notion that we had in AlphaFold, that the way to know a model's confidence is just to ask it to predict how wrong it will be, was itself a fun fact. It's a weird thing: you can get a model to tell you it's wrong without telling you what the right answer is. I can give a human analogy: you often come out of a test knowing if you did well or poorly without knowing all the answers. But it is very interesting. I think we'll see it develop over time, but also the kind of world and scientific systems will adjust. And I certainly have seen some bad uses of AlphaFold in science, but overwhelmingly I think people are using it really well and developing the right techniques and sharing in the community. It is really the community that learns how to use these tools to do their work better.
我想就科学过程本身追问一下。传统上,科学是关于解释和从机制上理解发生了什么,然后找到解决方案。你认为如果我们甚至不理解问题就解决了它,这可以吗?我的意思是,如果目标只是解决问题,我们解决了它,但不知道中间发生了什么,特别是在 AI 是一个难以理解的黑箱的背景下,这真的重要吗?这如何改变科学?
I just want to follow up with regard to the scientific process itself. Traditionally it's all been about explanations and understanding mechanistically what's happening and then getting to the solution. Do you think it's okay if we solve the problem without even understanding it? I mean if the goal is just solving the problem and we solve it and we don't really know what happens in between, particularly in the context of AI being a black box which is difficult to understand, does that really matter? And how does that kind of change science?
嗯,我认为这是一个有趣的问题,但我对此有两种看法。我甚至要说,科学不是关于解释,而是关于检验。我们有一个假设,从中推导出一个结论,然后我们看看——特别是一个令人惊讶的结论——这个令人惊讶的结论本身是否为真。从这个意义上说,这很干净地过渡到了像黑箱这样的东西,也许不是解释,而是做出预测的系统。事实上,这并不是什么新鲜事。我们进行计算机模拟已经很长时间了。你可以解释你是如何进行这个计算机模拟的,但然后你运行这个计算机模拟的百万、十亿步。你得到预测并检验它们,但在实践中,你并没有对分子动力学模拟中我们如何达到这个结果有一个解释。当然,这并没有成为科学的巨大障碍。我们有关于如何构建模拟器的理论。我们有运行模拟的结果,然后我们将它们整合到更大的实践中,最终我们检验结果。所以,当我们用 AlphaFold 做这件事时,我们并不以蛋白质结构为最终目标。我们需要它们是因为它们帮助我们理解生物学,帮助我们提出新的假设,然后我们检验这些假设。当我们用计算输入制造药物时,我们检验的是药物,而不是计算。我认为这不是一个根本性的挑战。而且我认为这里也有科学,但科学不是——你不能在中间切断它。你不能说外星人给了你一个黑箱,然后就是科学了。科学是我们对机器学习如何发展理解有了一定的理解。
Well, I think it's an interesting question but I have kind of two perspectives on it. I would even say that science isn't about explanation; it's about testing. We have a hypothesis, we deduce a conclusion from it, and we see—especially a surprising conclusion—if that surprising conclusion itself is true. So in that sense, that transfers quite cleanly over to having something like a black box, maybe not explanations, but systems that make predictions. And in fact, this is not a new thing. We've been doing computer simulations for a long time. You can explain how you came to this computer simulation, but then you run a million, a billion steps of this computer simulation. You get predictions and you test those, but you don't in practice have an explanation of, in a molecular dynamics simulation, how we got to this. And of course that hasn't been a huge stumbling block for science. We have theories of how you build simulators. We have the result of running simulations, and then we integrate them into our larger practice, and ultimately we test the consequences. So right, when we do this with AlphaFold, we don't want protein structures as the final goal. We want them because they help us understand biology, they help us make new hypotheses, and then we test those. When we make drugs with computational inputs, we test the drugs, not the computation. I think that's not a fundamental challenge. And I think also there is a science here, but the science isn't—you can't cut it quite in the middle. You can't say aliens handed you a black box and now science. The science is we have an understanding of how machine learning kind of develops understanding.
我们正在发展一种局部化的科学,关于如何构建蛋白质结构预测系统?我们能改变什么?什么才是重要的?然后这个数据加机器学习的整体循环产生了一个预测器,它在科学中有清晰的解释。但我觉得部分原因是人们不太看到第一部分,或者做这部分的人较少,所以他们只看到了产物,认为这是个挑战。但我认为最终我们会继续做预测,继续判断这些预测是否普遍可靠,就像我们一直以来对待理论和计算一样。我认为这不会对范式构成重大挑战,也许以后如果系统能自己提出假设并验证,但就目前而言,它很好地融入了科学范式。
We're developing kind of localized science in terms of how do we make protein structure prediction systems? What can we change? What matters? And then this overall loop of data plus machine learning yields a predictor that you then use that has a clean explanation in a science. But it's a little bit I think we're I think it's in part because people kind of don't see the first part as much or it's done by fewer people and so but they see the artifact that that they have this challenge. But I think ultimately we'll just continue doing we'll continue making predictions. will continue deciding whether those predictions are generally generally reliable the same way we've done with theories and with computation forever and I think it will it will not show up as a major kind of challenge to the paradigm maybe later if we have systems that are hypothesizing and testing them but right now I think it fits nicely within the scientific paradigm
我只想非常感谢你抽出时间与我们交谈,我们真的很感激,这对我们来说是一次非常精彩的讨论,所以非常感谢。
I just want to say thanks ever so much indeed for taking time out of your schedule to talk to us we've really appreciate it it's been such a fascinating discussion for us so thank you very much
非常感谢。所以,各位,John Jumper,多么棒的采访,多么亲切的人。
Thank you very much. So guys, John Jumper, what an interview and what a gracious guy.
哦,了不起的嘉宾,对吧?我当时没问他是否带着诺贝尔奖章。我现在有点希望我问了。我想看看它。
Oh, phenomenal guest, right? I didn't quite ask him if he had his Nobel medal on him at the time. I kind of wish I did now. I wanted to see it.
是的,我觉得与 John 的讨论非常有趣,因为现在我们听到的都是 AI 靠蛮力,谁算力最强谁就能成功。但他强调了研究的重要性,没错,数据和算力很重要,谷歌当然有很多算力,但研究才是 AlphaFold 成功的关键。
Yeah, I thought the discussion with John was really interesting because I think all we hear at the moment is that AI is about brute force. whoever has the most computational power is, you know, going to make it. And I think him emphasizing the importance of research and yes, data and compute are important and of course Google has a lot of compute, but the research really was the key to the success of AlphaFold.
是的,我认为科学方法在研究中的主导地位仍然不可动摇。你可以生成一个可能产生结果的黑箱,但你仍然需要验证那个结果,才能真正知道它有多真实。他做到了,并且谈到了他如何通过 AlphaFold 的迭代使其更准确、更强大。对于这些预测产生的蛋白质结构也是如此,你仍然需要验证。他在这里时明确确认了这一点。所以那很好。
Yeah, I think the idea of the scientific method still reigns supreme in research. The idea that you can generate this black box that might produce a result, but you still have to validate that result to really find out how truthful it is. And he did that and he talked about that with his own AlphaFold increments to make it more accurate and more powerful. And the same is true for the protein structures that come out of those predictions. You're still left with the need to validate. And he clearly confirmed that when he was here. So that was nice.
他对你和你的同事在那个领域所做的工作给予了高度赞扬,不是吗?是的,我认为他非常重视蛋白质数据库,他评论说这可能是训练大型语言模型最干净的数据集之一,这说明了我们需要为其他生物学领域开发同样好的数据集。他在思考如何让 AI 用生物学思维进行下一轮迭代时暗示了这一点,为了准确,数据从哪里来?我认为我们现在必须思考,对于其他生物学问题,下一个版本的蛋白质数据库是什么。
And he had a lot of compliments to pay about the work that you and your colleagues have done in that whole area, isn't he? Yeah, I think he really clearly values the Protein Data Bank and the fact that he kind of commented that it was possibly one of the cleanest data sets you could train a large language model on, I think speaks to the need to produce and develop other equally good data sets for questioning other areas of biology. And he hinted at that with the next iteration of thinking about ways to get AI to think in biological terms, where's it going to get the data from in order to be accurate, and I think we have to think about now what is the next version of the Protein Data Bank but for other biological questions.
但甚至超越生物学问题,材料科学、物理学、化学,甚至销售、商业金融文档。但我们如何将其纳入我们的系统和流程,确保我们收集了正确的数据、干净的数据,并且以对系统、组织和团队容易的方式做到这一点?
But even beyond biological questions, material science, physics, chemistry, in actual sales, yeah, business financial documents. But how do we build that into our systems, our processes that we are collecting the right data, we're collecting the clean data as well, and how we doing that in a way that's easy for our systems, our organizations, our groups.
精彩的讨论。请来了一位很棒的嘉宾,不是吗?我真的很喜欢。多么棒的节目,对吧?
Fascinating discussion. Great guest to have on, wasn't it? I really really loved it. And what a show, right?
多么棒的节目。我会说这是最好的之一。
What a show. One of the best, I'd say.
我也这么认为。它会在我脑海里停留很久。总之,感谢你们今天加入我们。Luke,Steven。
I think so. It's going to live in my head for a long time. Anyway, well, thank you guys for joining us today. Luke, Steven.
谢谢邀请。
Thanks for having me.
今天我们做了一期多么精彩的节目。感谢你的收看。如果你喜欢更多这样的内容,请点赞和订阅。
What an amazing show we had today. Thank you for joining us. Please like and subscribe if you'd like to see more content like this.