谷歌的 AI 科学家:从自动化 Kaggle 到科学发现——对话谷歌院士 John Platt

Google's AI Scientist: From Kaggle Automation to Scientific Discovery — John Platt, Google Fellow

约翰·普拉特 John Platt · Latent Space · 2026-09-22 · 约 121 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

谷歌院士 John Platt 讲述如何从自动化 Kaggle 的尝试,发展出将科学问题映射为可评分任务的 AI 科学家。

Google Fellow John Platt discusses how an attempt to automate Kaggle evolved into an AI scientist that maps scientific problems into scorable tasks.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 42)

全文 · Full transcript(中英对照)

预测模型与描述模型 Predictive vs. descriptive models

Host

你是在说引入一些显式的先验,这些先验基于人类的直觉,或者在这个情况下,可能是 LLM 的直觉?

Are you talking about introducing explicit priors that you know based upon some human intuition or maybe in this case LLM intuition?

John

当你谈到多重假设检验时,对吧,有预测模型和描述模型。预测模型就像,假设你有一些输入和一些输出,你只是想构建一段代码,试图在某些数据集上达到最低的错误率——一个统计模型。描述模型实际上是科学试图达到的目标,即,它应该能够外推,因为它内部捕捉了某种物理规律或对现实的某种描述,然后你可以用它来外推。是的,牛顿想到了苹果和引力,但引力实际上并不是关于苹果的,对吧?如果你采用 17 世纪的机器学习模型,比如“哦,苹果会掉下来”,但行星呢?你知道,我不知道。我没有关于行星的数据,所以谁知道它们会怎样,对吧?这两者之间的区别有点模糊,对吧?因为当物理学家或科学家出现时,他们使用自己的直觉,或者甚至比直觉更多。本质上,可能有一堆他们知道的关于世界的坚实事实,然后他们确保他们构建的任何模型都与已知的事实一致。

When you talk about multiple hypothesis testing, right, there's predictive models and there's descriptive models. A predictive model is like, let's say you just have some inputs and some outputs, and you just want to build a piece of code that tries to have the lowest error rate on some data set—a statistical model. A descriptive model is actually what science is trying to get to, which is, okay, it should be able to extrapolate because it has sort of the physics or some description of reality captured within it, and then you can use it to extrapolate. Yes, Newton thought of apples and gravity, but gravity isn't actually about apples, right? If you take the 17th century machine learning model like, oh, apples will fall, but how about planets? You know, I don't know. I have no data about planets, so who knows what they do, right? The distinction between those is a little bit blurry, right? Because when a physicist or scientist comes, they use their intuition or maybe even more than intuition. Like essentially there's maybe a solid pile of facts that they know about the world, and then they make sure that whatever model they build is sort of consistent with what's known.

介绍John Platt Introducing John Platt

Host

我的联合主持人 R.J.。今天非常高兴能邀请到 John Platt。John 是 Google 院士兼 Google Research 应用科学负责人。他的背景真的很有趣。我猜几分钟前我们聊天时,你把自己描述成一个超级书呆子。

My co-host R.J. It's a pleasure to have John Platt with us today. John is a Google fellow and head of applied science at Google Research. He has really a fun background. I guess you described yourself when we were talking a few minutes ago as a mega nerd.

John

哦,是吉咖书呆子。

Oh, giga nerd.

Host

吉咖书呆子。吉咖书呆子。他对一切都充满热情。这真的很明显。是的。如果我说错了什么,请纠正我,但你 14 岁开始上大学,18 岁在加州理工学院开始攻读博士学位。你的导师或联合导师是 John Hopfield,对吧?

Giga nerd. Giga nerd. He's excited in absolutely everything. And it really shows. Yeah. You was correct me if I'm wrong about any of this stuff, but so you started college at 14 and started your PhD at 18 at Caltech. You were advised or co-advised by John Hopfield, right?

John

哦,是的。是的。

Oh, yeah. Yeah.

Host

是的。是的。他两三年前刚获得了诺贝尔奖。

Yeah. Yeah. Who just won a Nobel Prize in you know two or three years ago.

John

是的。

Yes.

Host

所以 John 创造了几个教科书级的算法。一个被称为 Platt scaling,另一个是序列最小优化,这是训练 SVM 的教科书算法。即使在今天,如果你使用 sklearn,它仍然存在。John 发现并命名了两颗小行星,还因 2006 年的技术开发获得了奥斯卡奖。所以如果你看过皮克斯电影,你就见过 John 的算法和工作。John 的 Erdős–Bacon 数是六或三,两边都是三。我将跳过你大约 20 年的职业生涯,但然后跳到 Google,在 Google 科学部门工作,你从事过聚变、量子计算、气候建模和许多其他主题。这大致正确吗?

So John created several responsible for several textbook algorithms. One known as Platt scaling, another one sequential minimal optimization which is the textbook algorithm for training SVMs. Even today it's still if you use sklearn it's there. John has discovered and named two asteroids, has an Oscar for technical developments from 2006. So if you've ever watched a Pixar movie you've seen John's algorithms and work. John has an Erdős–Bacon number of six or three, three and three from either side. And I'm going to skip over like 20 years of your career but then jumping to Google at you know working at Google sciences you worked on fusion quantum computing climate modeling and many other topics. Is that more or less right?

John

没错,是的。

That's right, yeah.

Host

好的,好的,酷。我有没有遗漏今天重要的内容?

Okay, okay, cool. Did I miss anything important for today?

John

没有,我的意思是,我还做过很多应用数学和信号处理以及各种有趣的事情。

No, I mean, I've also done lots of applied math and signal processing and all sorts of fun things.

Host

是的。是的。我想你的维基百科上还有一个关于专利和 iPhone 的有趣故事。

Yeah. Yeah. I think you also your Wikipedia has a fun story about patents and the iPhone too.

John

是 iPod。

The iPod.

Host

iPod。是的。是的。

iPod. Yeah. Yeah.

John

是的。欢迎。

Yeah. Welcome.

Host

谢谢。谢谢邀请我。

Thank you. Thank you for having me.

ERA与可评分任务 ERA and scorable tasks

Host

你能告诉我们关于 ERA 的事情吗?我想这个缩写是这样发音的。我知道在 Google 内外有很多不同的半相关的东西。所以你能告诉我们一些关于 ERA 的细节以及它特别之处吗?

Can you tell us about the ERA, is the I think the way that the acronym is pronounced. And I know that there's a lot of different semi-related stuff out there both within and outside of Google. So what can you tell us a little bit about the details of ERA and what makes it special?

John

嗯,我们在 Google Research 做 AI for science 已经超过 10 年了。大约 10 年前,主要是使用,我不知道现在你们怎么称呼它,也许是经典机器学习模型,比如卷积网络之类的,它们是用来解决特定科学问题的特定模型。但大约两年前,我们对过去几年出现的这些更通用的 LLM 感到非常兴奋,我们想知道能用它们做什么。当然,很多人一直在尝试并试图弄清楚正确的做法是什么。我们偶然发现了这种映射。换句话说,我们发现许多不同的科学问题可以映射到我们称之为可评分任务的东西。所以你经常可以把一个科学问题表述为,天哪,我真的想要一段代码,它能最大化某个分数。令人惊讶的是,通过映射到这个框架,你可以在许多不同的科学问题上取得很大进展。嗯,一件事是很多科学家花大量时间构建模型。它们可能是统计模型,也可能是基于物理的模型。如果是统计模型,就像机器学习中那样,你的评分函数是,嗯,我有一些数据集,我希望模型和数据集的拟合度提高。我们稍后可以讨论过拟合,但

Well, we've been doing sort of AI for science in Google Research for more than 10 years now. And around 10 years ago it was very much using, I don't know what you call it now, maybe classical machine learning models, you know, things like convolutional nets or whatever, and they were specific models to build to solve specific science problems. But about 2 years ago we got very excited about these more general LLMs that have popped up in the last few years and we were wondering what can be done with them. And of course a lot of people have been playing and trying to figure out what the right thing to do is. And we kind of stumbled into this mapping. In other words, we found that many different scientific problems can be mapped into something we call scorable tasks. So you can often phrase a scientific problem as a gosh, I really would like to have a piece of code that maximizes some score. And it's surprising the number of different sort of scientific problems you can make a lot of progress on by mapping into that framework. Well, one thing is a lot of scientists spend a lot of time sort of building models. They might be statistical models or they might be physically based models. And if it's a statistical model like in machine learning, your scoring function is, well, I have some data set and I'd like to have the fit the model and the data set go up. And we can talk about overfitting in a minute, but

Host

那是我们的问题之一。

That was one of our questions.

John

然后这实际上非常有趣。所以机器学习是这种可评分任务的一个子集,对吧?但你可以做其他类型的事情,特别是 Michael Brenner,他是 ERA 论文的主要作者。他非常非常熟练,因为他喜欢现在用一个工具在一个晚上搞定一篇科学论文。所以应用数学中有个东西叫渐近展开,就是你在问一个常微分或偏微分方程,但比如常微分方程,行为如何,某个参数中有一个 epsilon,你试图说明当 epsilon 趋于零时它如何表现。结果你可以把它变成一个经验任务,本质上要求它提出一些渐近正确的解,然后你检查渐近解对于 epsilon 等于 1e-4 之类的值是否正确,然后你检查那个拟合。但然后你要求 Gemini,也就是其核心 AI,进行数学推理,尝试解决问题,同时最大化对数据的拟合。所以你实际上可以做很多技巧,因为底层改变代码的东西或底层做决策的东西不是一个随机过程。它本身就是一个 AI,很聪明,了解事物,对世界知道很多。你可以解决很多有趣的问题,因为那个核心内循环是一个拥有大量先验知识的 AI。所以这就是诀窍。

And then that's actually very interesting. So machine learning is kind of a subset of this sort of scorable task, right? But you could do other kind of things like especially Michael Brenner who's the lead author on the ERA paper. He's very very skilled because he likes to sort of knock out a scientific paper in an evening now with a tool. So there's something in applied math called asymptotic expansions, and which is you're asking how does an ordinary differential or partial differential equation but say ordinary differential equation behave, there's some parameter has an epsilon in it and you're trying to say how does it behave as epsilon goes to zero. And it turns out you can turn that into an empirical task by essentially asking that it proposes some solutions that are asymptotically correct and you check to see if the asymptotic solution is correct for like epsilon equals 1e-4 or something and then you check that fit. But then you ask Gemini, which is the core AI underneath it, to do the mathematical reasoning, try to solve the problem while also maximizing the fit to the data. So you can actually, there's a lot of sort of tricks you can do because it's not that the underlying thing that's altering the code or the underlying thing that's sort of making the decisions is not a random process. It's an AI itself that is smart and knows about things and knows a lot about the world. You can solve a lot of interesting problems because that sort of core inner loop is an AI that has huge amounts of prior knowledge. So that's sort of the trick.

科学问题映射为可评分任务 Mapping Scientific Problems to Scorable Tasks

Host

我们一直在四处奔走,试图把许多科学问题映射成可评分的任务,并尝试解决它们。这真的挺有意思的,我很乐意聊聊至少我参与过的那些。

So we've been running around trying to map lots of scientific problems into scorable tasks and trying to solve them. It's really been kind of fun, and I'm happy to talk about the ones that I've been involved in at least.

Host

好啊,我很想听听一些更……那个统计学的,听众们大概都知道。你刚才提到的很有道理。还有哪些其他有意思的?

Yeah, I would love to hear about some of the more... So, it's a statistical one is what everyone listening will probably know about. What you just mentioned makes sense. What are some of the other interesting ones?

John

让我想想,非统计学的。让我想想,因为我们有一些有趣的,比如我们刚在 arXiv 上贴了一篇论文,或者其实我觉得可能在 GitHub 上。就是……你在遥感领域经常遇到这种情况,因为总是有权衡。有卫星在地球上方飞行,它们能多频繁地重访地球上的某个地点、空间分辨率是多少、像素有多大、光谱分辨率是多少,也就是有多少个波段,这些之间有权衡。理想情况下,你希望有持续的地球监测,每五分钟一帧,高光谱分辨率,10 厘米级别。你得不到那样的数据。但比如说监测 CO2,大气中的 CO2 浓度,你可以从一颗卫星获取数据,比如 OCO-2 或 OCO-3。OCO-3 其实附着在国际空间站上,但它能给你一小条 CO2 测量数据,精度很高,分辨率也相当高。你其实可以尝试……因为很多都在红外波段,像 GOES 这样的气象卫星有一些红外波段,它基本上每 5 分钟拍一张照片,但像素非常大,光谱分辨率也不太好,因为它不是为探测 CO2 设计的。所以你就让一个去估计另一个,并加入其他数据,比如当前天气、长期反照率之类的,于是你做出了一个非常好的模型,几乎能做超分辨率,一种有信息的超分辨率,从一颗卫星到另一颗。所以这算一个例子。

Let's see, that are not statistical. Let's see, because we have interesting ones like one that we just put a paper up on arXiv is... or actually I think it might be on GitHub. It's... you often run into this in remote sensing because there's always a trade-off. There's satellites flying above the Earth and there's a trade-off between how frequently they can revisit a spot on the Earth, what their spatial resolution is, how big the pixels are, and their spectral resolution, so how many bands they have. And ideally you'd like to have monitoring of the Earth that's constant and, you know, a frame every five minutes at hyperspectral resolution at, at whatever, 10 centimeters. You can't get that. But for example to monitor CO2, the atmospheric concentration of CO2, you can take data from one satellite that's, for example, it's OCO-2 or OCO-3. OCO-3 is actually attached to the International Space Station, but so it gets you like a little strip of CO2 measurements that are highly accurate and pretty high resolution. You can actually try to do... there you could, because a lot of it is in the infrared, weather satellites like GOES has some infrared bands and it takes a picture every essentially 5 minutes but the pixels are very large and it doesn't have such great spectral resolution in terms of... it wasn't designed to find CO2. So you just ask one to estimate the other and you shovel other data in like what's the current weather, what's sort of the long-term albedo, and so you came up with this very nice model that can do almost like super resolution, an informed super resolution of one satellite to another. So that's like one example.

Host

嗯。嗯。好。所以任何科学问题你都能映射到这个框架里。输入到误差是这种映射,输出是代码,是吗……

Yeah. Yeah. Okay. So any scientific problem that you can map into this framework. So the input is to error is sort of this mapping and the output is code, is that...

John

嗯,差不多吧。我是说,输入是……我们在产品里设置的方式是你直接开始说话,对吧?因为很多时候怎么做这种映射并不显然,尽管像 Michael Brenner 这样的专家知道怎么做。所以我们实际上写了一个智能体来帮你,它会和你对话,帮你定义你的可评分任务应该是什么。所以已经有一个 Gemini 实例坐在那里试图帮你写代码。所以这其实几乎像一个中间结果。你开始和它谈你的问题,它试图在下面生成一个 Python notebook,里面有一个评分,本质上是一个可评分的函数……它会产生一个分数,然后它开始以巧妙的方式变异那个 notebook,因为还是 Gemini,它会不断提出试图最大化分数的代码。那么这和一般的、能优化 notebook 的通用智能体系统有什么不同?

Well, sort of. I mean, the input is... the way we've got it set up in the product is you just start talking, right? And so because a lot of times it's non-obvious how to do this mapping, although experts like Michael Brenner know how to do it. So we actually wrote an agent that actually helps you, it sort of talks to you to try to help you define what your scorable task should be. So there's already sort of an instance of Gemini sitting there trying to help you write code. So that's actually sort of almost like an intermediate result. You start talking to it about your problem and it tries to produce essentially a Python notebook underneath that has a score, essentially a function with a scorable... which essentially produces a score, and then it starts to mutate that notebook in a clever way because again it's Gemini and it will try to sort of keep proposing code that tries to maximize the score. So what is different about this than just a generally a general agentic system that can sort of optimize notebooks?

Host

现在它本质上是它自己的……用 2026 年的现代说法,我们其实在 24 和 25 年就做过这个,但用现代说法,它是一种专门的 harness,运行一个算法,在早期叫蒙特卡洛研究。所以本质上它保留成百上千个可能的 notebook 实例,然后选一个,我可以解释它怎么选,它决定好,我能做什么?Gemini 问自己我能做什么让那个 notebook 更好,然后它会做一个新的并测试,再放回候选池。所以你可以想象候选池其实是树结构的,因为每个候选可能有子节点,你要做的是基于一种叫……其实是一种相当标准的强化学习算法,叫上置信界,UCB。所以你本质上选……它是一种乐观算法。所以它试图估计比如变异的第 95 百分位结果,它试图估计那个,然后选上界最高的,最乐观的上界。换句话说,它不总是选表现最好的 notebook。它试图预测当前表现加上它猜测的两个标准差。所以它总是试图……所以它会四处搜寻。

Right now it's essentially its own... in modern 2026 parlance, we actually worked on this in '24 and '25, but in the modern parlance it's kind of a specialized harness that runs an algorithm which in former era was Monte Carlo research. So essentially it's keeping hundreds or thousands of possible instances of notebooks and then it selects one, and I can explain how it selects one, and it decides well, what can I do? Gemini asked itself what can I do to make that notebook be better, and then it will make a new one and test it and then put it back into the candidate pool. So you can imagine the candidate pool is actually tree structured because every candidate possibly has some children, and what you do is you pick based on something called the... it's actually a fairly standard algorithm from reinforcement learning called upper confidence bound, UCB. So you essentially pick... it's an optimistic algorithm. So it tries to estimate say what's the 95th percentile outcome of mutation, and it tries to estimate that and it picks the one with the highest bound, the highest optimistic bound. So in other words, it doesn't always pick the best performing notebook. It tries to predict like what's the current performance plus two sigma of its guess. And so it's always trying... so it hunts around.

Host

所以基本上是高召回。

So high recall basically.

John

高召回,它试图下注,以便最有效地取得进展,这并不总是贪婪地做最好的候选。有时是第五好的。我们还……我们还玩过它有点重组,从两个候选里取想法把它们砸在一起,试图做出第三个候选。

High recall, it's trying to make its bet so that it most efficiently tries to make progress, which isn't always greedily doing the best candidate. Sometimes it's the fifth best. We're also... we've also played around where it kind of recombines, it sort of takes ideas from two candidates and smashes them together and tries to make a third candidate out of that.

Host

它如何播种初始候选池?

How does it seed the initial candidate pool?

John

嗯,神奇之处在于 Gemini 底层其实很擅长写代码。我是说,你只要让它给我写个东西,因为你有问题的文本描述。不是说你给个评分函数就开始。你说函数的文本描述,你可能给它……事实上我们在某些东西下面有,比如这里有五篇人们试图解决这个问题的论文,它挺聪明的。它实际上会去读论文,然后真的会先尝试写代码。可能不太好,也可能好,你知道,有时有 bug,它基本上返回负无穷,但然后它会尝试变异代码,说“哦,它会试着让它更好。”所以挺酷的。你其实不必给它……我是说,如果你想可以给一些起始代码,但不必。

Well, that's the amazing thing is that underneath Gemini is actually good at writing code. I mean you just ask it write me a thing because you have a textual description of the problem. It isn't just oh here's your scoring function start. You say a textual description of the function and you might give it... in fact we have under some things like here are five papers that people tried to solve this problem with and it's kind of smart. It actually goes and reads the papers and will actually take a first stab at code. It might not be great or it might, you know, sometimes it has bugs and it returns essentially minus infinity, but it will then try to mutate the code and say, "Oh, it'll try to make it be better." So it's pretty cool. You don't actually have to give it... I mean, you can if you want give it some starter code, but you don't have to.

Host

你启动了多少个智能体?我猜也许不是智能体,或者每次迭代你启动多少个不同的,你知道,树分支?

How many agents are you spinning up? I guess maybe not agents or how many different, you know, tree branches are you spinning up at each iteration?

John

哦,每次迭代。嗯,有权衡。你想做很多并行工作,但如果并行太多,你就不能从之前的东西学习。所以,现在,我们用大约……默认是 10 并行。所以你一次尝试长出 10 个叶子。

Oh, at every iteration. Well, there's a trade-off. You'd like to do a lot of parallel work, but if you do too much parallel work, you can't learn from previous things. So, right now, we use about... the default is 10 par. So you try to grow 10 leaves at a time.

Host

好。

Okay.

John

这似乎是正确的权衡。

That seems to be about the right trade-off.

Host

当你说不能从之前的迭代学习,那意味着编排器或者有某种……是的。

When you say you can't learn from previous iterations, that means that the orchestrator or is there some sort of... Yeah.

智能体内省与并行搜索 Agent introspection and parallel search

Host

你刚才说有些步骤能够重组或做出超越分数的决策,那是什么?我想问的一个问题是:作为人类,当你做某种机器学习项目时,你不会只看我们试图优化的那一个指标。很多时候还有正交的指标,有时甚至是一些洞察,比如仅仅观察训练曲线就能给你关于正在发生什么的直觉,或者查看具体的例子。它会做这种内省吗?

What's that you said about some step that's able to recombine or make decisions beyond just the score? I guess one of the questions is: as a human, when you're doing some sort of ML project, you don't just look at one metric that we're trying to optimize. Often times there are orthogonal metrics, sometimes even insights such as just watching training curves can give you intuition about what's going on, or looking into specific examples. Does it do any sort of introspection like this?

John

嗯,它有历史记录。但我的意思是,为什么你不能同时做太多事情:如果你同时有 10 个并行搜索,那么第一个实际上看不到第二到第十个在做什么。所以如果你一次做一千个,那么你使用了大量算力,却没有太多交叉学习。而一旦你完成一小批,你就能获得历史记录。显然你必须修剪它,以免上下文窗口爆炸,但你能获得它在写代码时思考的内容以及代码结果的历史记录。所以它可以从之前的尝试中学习。

Well, it has the history. But I meant why you can't do too many things in parallel is if you have 10 parallel searches at once, then number one can't actually see what numbers two through 10 are doing. So if you do a thousand at once, then you're using a huge amount of computation without a lot of cross learning. Whereas once you finish a little batch, you get the history. Obviously you have to prune it so it doesn't blow up the context, but you get the history of what it was thinking about as it was writing the code and the results of the code. So it can learn from its previous attempts.

Host

好的。那它会跨批次学习吗?

Okay. And does it learn across?

John

哦,是的。是的。本质上它就像一个共享的上下文窗口。是的。所以它是一边进行一边思考。它不像是一千个完全独立的、互不相关的分支。

Oh yes. Yes. Essentially it's like one shared context. Yes. So it is thinking as it goes along. It's not like it's a thousand different completely independent branches.

Host

你确实在推动 Gemini 的长上下文能力。

You're really pushing Gemini's long context abilities.

John

没错。而且你必须做好正确的管理等等。

That's right. And you have to do the right management and stuff.

Host

是的。是的。是的。好的。哦,那很酷。我的意思是,回到 R.J. 的问题,这里的关键点是,第一个关键目标,我猜是确定你试图优化的具体分数,对吧?我同意,有时那确实是问题中最难的部分。所以我觉得有趣的是,我不确定我总是会信任我的智能体去做那部分。那部分似乎更像是循环中更人类的任务。

Yeah. Yeah. Yeah. Okay. Oh, that's cool. I mean, going back to R.J.'s question, the key point here being the first key goal is to, I guess, identify what specific score that you are trying to optimize, right? Sometimes that is, I agree, kind of the hardest part of the problem. And so I find it interesting that I'm not sure I'd always trust my agent to do that part. That part seems like the more human task in the loop.

John

是的。而且你经常必须小心。事实上,你做的很多事情都非常元。我想我们做的一切都非常高层。你必须确保——一个常见的情况是,你提出一个评分函数,或者智能体提出,或者你们一起做,然后迭代会找到一种作弊或钻空子的方式,就像,哦不,我不是那个意思。所以你必须经历,经常玩,有一个循环,你迭代说,不不,我不是那个意思,或者你必须在指令中告诉它,好吧,不要这样做。所以是的,经常有迭代,即使有智能体帮助,你也不一定能从第一天就得到正确的评分函数。事实上,这真的很棒,因为——我的意思是,在过去,也就是 2024 年左右——很多研究生会花大量时间做科学软件,写代码本身就非常费力,所以你可能会尝试几件事,或者几件非常相关的事,然后你就停了,因为你必须写论文,你必须做下一个实验。这个东西有点在下面,有点无情,因为它不断尝试,不断尝试,不断尝试。所以使用它的人现在几乎把所有时间都花在正确的层面上,几乎在科学创造力的层面上。拥有一个成本函数意味着什么?你明白我的意思吗?所以这几乎就像是科学问题的本质。你现在不那么关注细节了,哦,我必须导入这个 CSV 文件,或者我必须让这个数据库工作,或者别的什么。你现在几乎是在哲学上深入思考你实际的科学问题,而不是在担心数据库的肮脏泥潭里。事实上,智能体可以做的一件很酷的事情是实际上向你建议数据集,比如,哦,你有没有想过也许拉入这个数据集并做一个连接?所以它会就你可以连接的数据集提出建议,这很酷。

It is. And often you have to be careful. In fact, a lot of what you do is very meta. I guess everything we do is very high level. You have to make sure that—one common thing is you come up with a scoring function, or the agent does, or you do it together, and then the iteration finds a way to cheat or hole, like, oh no, I didn't mean that. And so you have to go through and often play and have a loop around it where you iterate like, no no, I didn't mean that, or you have to tell it in its instructions, okay, don't do this. So yeah, there's often iterations, and even with agent help you don't necessarily get the right scoring function from day one. And in fact, it's really neat because—I mean, in the old days, i.e., 2024 or something—a lot of grad students would spend a lot of time doing scientific software, and it's just so much effort to write code at all that you kind of try maybe a few things, or a few things that are very related, and then you sort of stop because you have to write your paper, you have to do your next experiment. This thing is kind of underneath, kind of relentless because it keeps trying and keeps trying, keeps trying. And so the people who use it are now spending all their time almost at the right level, almost at the scientific creativity level. What does it mean to have a cost function? You know what I mean? And so that's almost like the essence of the scientific problem. You're not so much now in the details of, oh, I have to import this CSV file, or I have to get this database to work, or whatever. It's you're now sort of thinking almost like deeply philosophically about your actual scientific problem, not down in the grungy goop of worrying about databases. In fact, one cool thing the agent can do is actually suggest data sets to you, like, oh, have you thought about maybe pulling in this data set and doing a join? And so it'll make suggestions about data sets you can join with, which is kind of cool.

Host

回到你刚才说的,关于智能体喜欢破解东西和奖励破解,有没有什么有趣的故事或有趣的故事,关于事情喜剧性地失控?

Going back to what you said a second ago, in terms of agents love to hack things and the reward hack, are there any fun stories or interesting stories about where things were comedically run off the rails?

John

天哪,我一时想不起来了。我知道其他人遇到过。我不确定我是否有足够的细节来表达那种喜剧性,但它确实——你会感到惊讶。

Boy, I'm blanking. I know other folks have run into it. I don't know if I have enough details to sort of express the comedy of, but it does—you kind of get surprised.

Host

是的。嗯,我不确定我有没有任何真正具体的——抱歉,我一时想不起来了。

Yeah. Um, I don't know if I have any really concrete—sorry, I'm blanking.

John

不,没关系。是的。我总是喜欢把机器学习想象成古老的精灵故事,在猴爪之前,就像小心你的愿望,因为你会得到它。

No, it's fine. Yeah. I always like to think of machine learning as like the old genie stories before monkey paw, like careful what you wish for because you're going to get it.

Host

是的,没错。这在这件事上非常明显。所以你必须小心。但另一方面,它有一些知识。好消息是 Gemini 对很多事情了解很多,比任何一个人能做到的都多。所以它至少知道——特别是如果你指向论文,比如这里有五篇论文以某种方式尝试做这个。所以某种程度上它确实有那种精灵的感觉,但某种程度上它也做理智的事情。这就是为什么记住进化编码的整个想法。它从 70 年代就存在了。每个人都喜欢这样做。比如,哦,让我们变异列表代码或其他东西来做事情。但它没有起飞的原因是代码空间中的随机变异几乎毫无价值。我的意思是,就像 DNA 一样,大多数事情是有害的。所以在这里,哦不,不,我们实际上可以找到——它有点在底层知道,并且知道有趣的方向去尝试,这就是为什么这个东西有效,底层循环本身就是一个 AI。所以是的,它可能过拟合,有你暗示的有趣的精灵问题,但它也有一定程度的理智,因为它了解世界,并且有世界知识在里面。

Yeah, that's right. And that happens very much with this. So you have to be careful. But on the other hand, it has some knowledge. The nice thing is that Gemini knows a lot about many things, more than any one person can do. So it at least knows—especially if you point papers, like here are five papers that tried to do this in some way. So to some extent it does have that genie feel, but to some extent it also does sane things. This is why remember the whole idea of evolutionary coding. It's been around since the 70s. Everyone's loved to do that. Like, oh, let's mutate list code or whatever to do things. But the reason why it just hasn't taken off is that random mutation in code space is pretty much worthless. I mean, just like DNA, it's like most things are harmful. So here it's like, oh no, no, we can actually find—it sort of knows underneath and knows interesting gradients to try, which is why the thing works, that the underlying loop itself is an AI. So yes, it can maybe overfit and have funny genie problems like you allude to, but it also has some amount of sanity because it knows about the world and it has world knowledge in it.

Host

不过那篇论文,你在做 Gemini 2.5,我想你知道 Gemini 已经进步了不少。你有指标吗,或者你——你知道这是一个你持续使用的工具,听起来你在改进,我想知道你们内部是否看到了这个工具的有效性几乎像相变一样?它——在过去一两年里,改进有多剧烈?

The paper though you were doing Gemini 2.5 and I think you know Gemini has advanced quite a bit. Do you have metrics or have you—you know this is a tool that you're continuously using and it sounds like you're improving and I'm wondering like do you internally have you seen like almost like a phase transition in how effective this tooling has been? Has it—how dramatic has the improvement been over the last like I guess year or two?

John

哦,嗯,我是说一两年。是的。惊人。

Oh well I mean year or two. Yeah. Amazing.

引言与Gemini进展 Introduction and Gemini Progress

Host

换句话说,每一个甚至每一个半版本,我的意思是,我认为在 Gemini 2.0 下这是不可能的。

In other words, every even every half version of, I mean essentially I think it would have been impossible under Gemini 2.0.

John

是的,我认为是这样。换句话说,它不会成功。

Yeah, I think so. In other words, it wouldn't have worked.

Host

所以,你从 2.5 开始,那就像是只是

So, you started at 2.5 and that was like just the

John

哦,不。实际上我们已经尝试实验这些东西有一段时间了。好吧。然后事情就是不起作用,然后它们开始起作用,现在它们简直太棒了。所以,Gemini 主要版本的进展简直令人惊叹。

Oh, no. We've been trying to experiment with these things actually for a while. Okay. And things just weren't working and then they started to work and then now they're just amazing. So, the progress on Gemini major versions has just been stunningly amazing.

Host

是的。是的。我认为这是很多人都有过的经历,那些看起来不可能的事情突然变得神奇地有用,而且速度非常快。

Yeah. Yeah. I think this is an experience a lot of people have been having where things were just seemed impossible whatever are suddenly becoming magically useful like really quickly.

John

所以如果人们,我甚至对科学家说这个,因为有些人说哦我试过 2.0 什么的,我不喜欢,哦是的,那是很久以前,那是一年前,那就像是很久,那是永恒以前,对吧,是的,事实上所有我们甚至有一篇预印本,我们结合了

And so if people are, I even say this to scientists because there's some people like oh I tried whatever 2.0 go and I didn't like oh yeah that was a long time that was a year ago that was like a long that was eternity ago right yes and in fact all the we even have one of the preprints where we've sort of combined

Host

呃呃 era 和 anti-gravity,你知道整个 anti-gravity 的整个框架也相当惊人,就是那个你可以拉入很多论文,它可以为你写很多代码,并且

uh uh era with anti-gravity and that you know the whole that whole harness of anti-gravity is pretty amazing too and it's that that's the one where you can sort of pull in lots of papers and it can write lots of code for you and

John

所以是的

so yeah

Host

那是公开可用的还是

is that publicly available or is that

John

呃 anti-gravity 是的。是的。

uh the the anti-gravity Yeah. Yeah.

Host

嗯,anti-gravity 肯定是公开的。

Well, anti-gravity is certainly publicly.

John

哦,抱歉。抱歉。是的。呃 era 加上 anti-gravity。

Oh, sorry. Sorry. Yeah. Uh the the era plus anti-gravity.

Host

呃 还没有。

Uh not yet.

John

好的。好的。还没有。好的。

Okay. Okay. Not yet. Okay.

代码变异与过拟合 Code Mutation and Overfitting

Host

我觉得这个领域真的很有趣,因为就像你说的,基本上从计算机科学诞生以来就有某种形式的代码变异。嗯,典型问题大概是过拟合或多重假设检验问题,我认为这可能更匹配这个问题,即你基本上现在我的假设是这个算法有效,现在我的假设,所以你面临的风险是它呈指数级爆炸,对吧,因为现在突然我有这些,就像我优化的超超参数,所以你探索的状态空间爆炸,因此似乎更容易过拟合到一个问题。你对此有什么想法?因为另一方面,根据我的经验,我甚至试过嗯你知道那种开源的 ERA 版本。嗯,我把它接入 Claude,它现在正在运行,所以我不能告诉你它运行得怎么样。

I find this area really fascinating cuz like you said, it's been there's been some form of code mutation out there since the dawn of computer science basically. Um the canonical problem is sort of the overfitting or multiple hypothesis testing problem I think which is maybe a little bit better match to the problem where you're basically my hypothesis now that this algorithm work now my hypothesis and so that um you run the risk that sort of it has exponentially exploded right because now suddenly I have like these it's like hyper hyperparameters that I'm optimizing and so that you have this explosion of state space that you're exploring and so that it seems much easier to sort of overfit to a problem. What are your thoughts about that? Because on the other hand, empirically my experience, I even tried um the you know sort of open-source version of ERA. Um I kind of strapped it into Claude and it's running right now so I can't tell you how well it's working.

John

好的,我很好奇。

Okay, I'm curious.

Host

是的,我会告诉你的。嗯,但我只是好奇想知道,这是一个一直在我脑海中的问题,关于一般的科学 AI,那么你在这方面的经验是什么,那种实地经验?

Yeah, I'll let you know. Um, but I'm just curious to know this is a question that's been on my mind about just general AI for science and what so what are your experiences with this sort of on the the sort of on the ground?

John

我想你的问题中嵌入了两个问题,对吧,因为当你谈论多重假设检验时,有预测模型和描述模型,对吧,呃,你知道我能看出我很久以来一直在做机器学习和统计吗

I guess there's two questions sort of embedded in your question I think right because when you talk about multiple hypothesis testing right there's predictive models and there's descriptive models right uh I you know can you tell I've been doing machine learning for a long time in statistics

Host

这实际上是一个非常好的观点,不过你介意解释一下,展开一下吗?我不确定这是每个人都会,观众都会熟悉的东西,就像

that's actually a really good point though do you mind explaining that expanding that I'm not sure that's something that everyone would audience would be familiar with like

John

尤其是在现代,我认为人们试图模糊这两者

especially in in in modern days I think people are trying to sort of obscure the two

Host

如果你从 LLM 开始,我不确定这种区分会有意义

if you started with LLM I'm not sure that distinction would be meaningful

John

没错。所以,预测模型就像,假设你有一些输入和一些输出,你只是想构建一段代码,试图在某些数据集上具有最低的错误率。那只是一个统计模型,对吧?描述模型实际上是科学试图达到的,即,它应该能够外推,因为它有某种物理学或实际现实的某种描述被捕获在其中。然后你可以用它来外推。嗯,因为就像你知道的,是的,牛顿想到了苹果和重力,但重力实际上不是关于苹果的,对吧,如果他只是拟合,如果他把它当作机器,你知道 17 世纪的机器学习模型,就像哦苹果会掉落,但行星呢,你知道我不知道,我没有关于行星的数据,所以谁知道它们会做什么,对吧

That's right. So, a a predictive model is like let's say you just have a you have some inputs and you have some outputs and you just I just want to build a piece of code that tries to just have the lowest error rate on some data set. That's just a statistical model, right? A descriptive model is actually what science is trying to get to, which is okay, it should be able to extrapolate because it has sort of the physics or the actual some description of reality that's captured within it. And then you can use it to extrapolate. Um because it's sort of like you know yes Newton thought apples and gravity but gravity isn't actually about apple right if he had just fit if he had taken as a machine you know the 17th century machine learning model like oh apples will fall but how about planets you know I don't know I have no data about planets so who knows what they do right

外推与模型类型 Extrapolation and Model Types

Host

所以,所以当你说外推时,我很好奇,好吧,所以当你说外推时,所以有不同的方式我可以思考这个,其中之一是你说的物理模型或世界模型,你是在谈论引入显式先验,你知道基于某些人类直觉,或者也许在这种情况下是 LLM 直觉,或者你是在谈论这是物理实际上是由模型学习的,或者

so so when you say extrapol I'm curious okay so when you say extrapolative so there's different ways I could think about this one of them is you said a model of physics or model of the world are you talking about introducing explicit priors that you know based upon some human intuition or maybe in this case LLM intuition or are you talking about this is the the physics is actually learned by the model or the

John

更多的是

it's more

Host

世界的底层过程在模型之下

the underlying process of the world is under the model

John

这些之间的区别有点模糊,对吧,因为当物理学家或科学家来时,他们使用他们的直觉,他们或大约或甚至可能不仅仅是直觉,就像本质上可能有一堆坚实的事实,他们知道关于世界,然后他们确保他们构建的任何模型都与已知的一致。嗯,目前在 era 中,它有点是 LM 直觉,本质上这就是我试图说的,在下面有一个好的梯度,特别是如果你把它指向现有的论文,它会尝试构建那种在下面合理的模型,因为如果再次特别是如果你给它指导,比如哦确保纳入这个和这个,或者看看这些论文,你会得到这些东西,所以它会,你可以引入对某些模型选择的偏见,它会有一个偏见,因为它只是它自己的小世界知识在自身内部积累,在预训练中。

the distinction between those is a little bit blurry right because when a physicist or scientist comes they use their intuition and they or about or or maybe even more than intuition like essentially there's maybe a a solid pile of facts that they know about the world and then they make sure that whatever model they build is sort of consistent with what's known. um currently in era it is it's sort of it is LM intuition essentially that's what I was trying to say about having a good gradient underneath that that especially if you point it at existing papers it will try to build models that are kind of sane underneath because if again especially if you give it guidance like oh be sure to incorporate this and this or look at these papers you get these things so it will there is a you can introduce a bias towards um certain model choices and it will have a bias because it it just its own little sort of world knowledge is accumulated inside of inside of itself in in uh in sort of pre-training.

物理信息建模 Physics-Informed Modeling

Host

你能给我们举些例子说明那可能是什么样子吗?它是用某种方式建模,而这种方式实际上有所不同吗?比如,如果你在处理偏微分方程,人们有这些形式化方法,比如神经算子,或者你可以在某种意义上编码一个微分方程。我想那叫做物理信息神经网络之类的。它就像用于流体和偏微分方程类型建模系统的一种方法,或者对于分子系统,经常有这种关于等变性的想法。模型是在利用文献中开发的这些技巧,还是为某种物理先验添加一些权重?当它引入某种知识时,那看起来是什么样子?

Can you give us examples of what that might look like? Is it modeling something in a way where it's actually different? For example, if you're doing something with partial differential equations, there are these formalisms people have, like neural operators, or where you can encode a differential equation in some sense. I think that's called physics-informed neural networks or something. It's like one for fluid and kind of PDE-type modeling systems, or for molecular systems, there's often this idea about equivariance. Is the model picking up on these tricks which have been developed in the literature, or is it adding some weights for some physical prior? What does that look like when it introduces some sort of knowledge like that?

John

我不确定我是否有足够的数据来说,哦,你知道,73% 的情况下它会这样做。但特别是当你指向现有论文时,它会尝试——事实上,它会非常擅长将论文中描述的方法适应于问题。事实上,它会做得非常出色。你实际上经常可以重新创建或逆向工程一篇论文。这又是 Michael Burner 实际上喜欢做的事情。他会说,哦,那听起来像一篇有趣的论文。哦,我们实际上为这个做过——我们有这个东西。这实际上有点像个 hack。我向 Michael 建议过这个,有一位 MIT 教授想出了些代码来做本质上——如果你有一个固定面积的屋顶,你想最大化一天内捕获的太阳能,太阳能,你可以建造当然能捕获更多阳光的东西,你可以让它设计一个包含镜子或支柱或太阳能板的部件,以你喜欢的任何角度或尺寸堆叠,通常有最大高度,然后试着弄清楚——让它探索那个设计空间。我相信 Michael——我们可以问他——我相信它实际上只是——我不认为他真的安装了模拟器。我认为代码只是复制了代码,因为它内部有像编码智能体一样的东西。所以只是从论文中复制了代码并弄清楚了这一切。所以是的,它非常好,特别是如果给它一个指向别人做过什么的指针。它非常擅长那种——哦,我还没见过它尝试等变建模;如果你了解 Clebsch-Gordan 系数,那可能会变得非常复杂。挺有趣的。是的。所以我不知道它是否会做真正的等变的东西,但它实际上知道很多。我记得当 Gemini 2.5 出来时——我知道这不完全是关于 ERA 的,但我记得当时想,'哦,这是一个新世界。'当 2.5 出来的那天,因为我说,'嘿,Gemini 2.5,你能给我写一些提升决策树代码吗?'它做到了。而且它有效。它就是做到了。是的。我说,是的,这是一个新世界。所以是的,我想回到你的问题,我认为是的,如果你给它一些指导,比如,哦,你知道,重要的是放入这种东西,它会的。所以它不一定会,至少我们没有看到,发现完全新的——就像如果你不知道 Clebsch-Gordan,去钓鱼吧,我不知道,你明白我的意思吗?如果它不知道某件事,它不会完全从头发现一种新的物理模型,但如果你告诉它关于世界上已知的有趣约束,它肯定会进化。我不知道我是否回答了你的问题。

I don't know if I have enough data to sort of say, oh, you know, 73% of the time it does this. But especially when you point it at existing papers, it will try to—in fact, it will do very well at adapting the methods that are described in the papers for the problem. In fact, we'll do an amazing job. You can actually often just recreate or reverse engineer a paper. That's again what Michael Burner actually likes to do. He'll say, oh, that sounds like an interesting paper. Oh, we did this actually for the—we had this thing. It was actually kind of a hack. I suggested this to Michael where there's this one MIT professor who came up with some code to do essentially—if you have a rooftop with a fixed area and you want to maximize the amount of solar power you capture over a day, sort of solar energy, you can build up which of course captures more sunlight and you can have it design a widget involving mirrors or struts or solar panels to sort of stack at whatever angles or sizes you like, and usually with a maximum height, and then try to figure—let it sort of explore that design space. And I believe Michael—we can ask him—I believe it actually just—I don't think he actually installed the simulator. I think the code just reproduced the code because it has like a coding agent inside of it. So just reproduced the code from the paper and sort of figured it all out. So yes, it's very good, especially if given a pointer to what other people have done. It's very good at kind of like—oh, I haven't seen it try equivariant modeling; that can get very hairy if you know about Clebsch-Gordan coefficients. It's pretty fun. Yes. So I don't know if it'll do the true equivariant stuff, but it actually knows a lot. I remember actually when Gemini 2.5 came out—I know this is not exactly about ERA, but I remember sort of thinking, 'Oh, this is a new world.' When like the day 2.5 came out, because I said, 'Hey, Gemini 2.5, can you write me some boosted decision tree code?' And it did. And it worked. It just did. Yes. And I said, yeah, this is a new world. So yes, I think to loop back to your question, I think yes, if you give it sort of guidance about, oh, you know, it's important to put this kind of thing in, it will. And so it won't necessarily, at least not that we've seen, discover completely new—like if you didn't know about Clebsch-Gordan, go fish, I don't know, you know what I mean? If it didn't know about something, it won't completely discover a new kind of physical model from scratch, but it will certainly, if you tell it about interesting constraints about the world that are known, it will certainly evolve. I don't know if I answer your question.

人类创造力与AI Human Creativity and AI

Host

所以,所以未来一两年仍然有空间留给人类。

So, so there's still room for humans for the next year or two.

John

哦,事实上,回顾过去,我认为完全有空间留给人类,因为我不知道。我的意思是,我们有 co-scientist 试图帮助你提出假设生成,但真的,我仍然没有看到那种创造力、哲学和那种仔细的严谨性。你完全需要人类。我不认为人类会消失。它们可以提出奇怪的 sug,我实际上用 co-scientist 解决了一个地球化学中的有趣问题,我了解到一种我不知道在岩浆中发生的新离子,但——所以它会告诉你有趣的事情,你会学到东西,但我不认为它能替代人类的创造力。

Oh, in fact, there's going back I think there's totally room for humans because I don't know. I mean, we have co-scientist that tries to help you come up with sort of hypothesis generation, but really I still haven't seen sort of the creativity and the philosophy and the sort of the careful rigor. You totally need the humans. I don't see humans going away. They can make strange suggestions and I've used co-scientist actually for an interesting problem in geochemistry and I learned about a new kind of ion I didn't realize happened in magma but I—and so it'll tell you interesting things and you'll learn stuff but I don't think it sort of substitutes for human creativity.

锯齿前沿与进展 Jagged Frontier and Progress

Host

你知道,回顾我们刚才谈到的两年,2 点,你知道,2,你有 Gemini 20 到 2.5,这就像你已经看到了一个飞跃,现在又过了一两年,现在我们到了 3.5,你知道,你说这工作得好多了。我的意思是,每当你看着图表,你知道,你可以——如果某件事看起来像指数,它要么——你要么处于 S 形曲线中,要么处于起飞的开端,对吧?我猜每个指数最终都会变成 S 形曲线。

You know going back we were just talking about two years 2 point you know two you had Gemini 20 to 2.5 and this was like you already saw a leap and now it's been another year or two and now we're at 3.5 and you know there's you're saying this is working much better. I mean whenever you look at a graph you know you can if something looks like an exponential it can either you can either be in a sigmoid or you can be at the beginning of a takeoff right I guess every exponential turns into a sigmoid eventually.

John

每个指数最终都会变成 S 形曲线。

Every exponential turns into a sigmoid.

Host

但问题是我们处在那个曲线的哪个位置。

But the question is like where are we on that.

John

我的意思是,我猜我非常相信那种整个锯齿状的——是的,锯齿状前沿,所以至少我所看到的,我的意思是,我不知道几年后会发生什么,但是的,在编码能力、收集知识和寻找相关事物方面,有一些大的尖峰,这巨大而美妙,我认为这对科学家来说很棒。到目前为止,在严谨性方面,它有点少,我们可以谈论像国际数学之类的,但总的来说数学,但但但就哲学和创造力而言,我认为它仍然有点不行,我——也许它会——也许一切都会——有些人说一切都会膨胀并通过,但我仍然看到很多非常强烈的锯齿状。所以我可以看到,好吧,也许它会变得特别擅长编码,特别擅长拟合模型,特别擅长提出建议等等,但我不知道。到目前为止,没有,到目前为止你需要人类。

I mean I I guess I'm a big believer in sort of the whole jagged—yeah the jagged frontier and so certainly at least what I see I mean I don't know what's going to happen in a couple years but yes there there's some big spikes out in jaggedness in terms of coding ability and just gathering knowledge and and finding related things and and that's huge and wonderful which I think is great for scientists. So far, it's kind of less in terms of uh rigor and um we can talk about things like the international methad, but and math in general, but but but in terms of sort of philosophy and creativity, I think it's still kind of not and I I maybe it'll maybe everything will people some people are saying everything's going to inflate and pass, but I'm still seeing a lot of very strong jaggedness. So I I can see well maybe it'll get to be extra good at coding and extra good at at fitting models and extra good at making suggestions and things but I don't know. So far no so far you need the humans.

多重假设检验 Multiple Hypothesis Testing

Host

是的,我想回到关于多重假设检验的问题。

Yeah I want to get back to the question about the multiple hypothesis testing.

John

抱歉,我完全喜欢这个题外话。嗯,多重假设检验是当你有一个描述性模型,你说这就是世界运作的方式,你有一堆数据,你拿十亿支飞镖扔出去,所以你必须小心。有一种叫做错误发现率的东西,对吧,所以问题是。

Sorry totally love the the tangent. Um multiple hypothesis testing is when you make have a descriptive model and you're saying this is the way the world works and you have a bunch of data and you take a billion darts and you throw and so you have to be careful. There's some something called false discovery rate right and so the question is.

Host

这是在寻找描述性模型,还是在寻找预测性模型?从根本上说,科学家在那里是为了确保所说的任何东西都是描述性的。

Is this finding descriptive models or is this finding sort of predictive models and fundamentally the scientist is there to make sure that whatever is saying is descriptive.

AI作为科学利器 AI as a Power Tool for Science

John

到目前为止,我们还无法用这些组件构建出一个能真正发现全新物理学或全新科学的系统,但这是一个强大的工具,可以帮助你发现全新的科学。所以也许我是在试图回避你的问题。我认为这触及了核心。

We haven't been able to make a system so far out of these pieces that can really discover completely new physics or completely new science, but this is a power tool to help you discover completely new science. So maybe I'm trying to unask your question. I think this gets at the heart of it.

Host

是的。

Yes.

John

是的。但然后你说,那纯粹的过拟合呢?好吧,我们先放一边。它并不是试图找出一个描述性的世界模型。那仍然取决于科学家。但单纯的过拟合呢?

Yes. But then you're saying what about just pure overfitting? Okay, let's set aside. It's not trying to figure out a descriptive model of the world. That's still up to the scientist. But what about just plain old overfitting?

Host

嗯。

Yeah.

John

是的。你必须非常小心,因为它是一个强大的工具。它可能会切掉你的手指。你明白我的意思吗?你必须非常小心,必须非常严谨。事实上,现在你必须更加小心、更加严谨,以免自欺欺人。你真的需要极其小心地设置非常隐蔽的保留集,你绝对不能看。你必须超级超级严谨,确保你不会完全……因为它是一个完全强大的工具。

Yes. You have to be very careful because it's a power tool. It can slice your fingers off. You know what I mean? You have to be very careful and you have to be very rigorous. In fact, now you have to be more careful, more rigorous to not fool yourself. You really really need to be just excruciatingly careful about having very hidden hold out sets you don't look at. You have to be just super super rigorous to make sure that you don't completely... because it is a total power tool.

Host

所以,如何不切掉手指的问题,答案是:你需要使用相同的技术,但要非常小心。

So the question to how do I not slice my fingers off is you need to use the same techniques but be very careful with them.

John

是的。

Yes.

Host

这是一个非常清晰的答案。

That's a very clear answer.

John

是的。好吧。是的。其实我觉得我没听过任何嘉宾说过这个。是的。这就像……我认为这是一项非常重要的技能,可能是新时代最重要的技能之一……人们经常谈论品味。

Yeah. Okay. Yeah. I actually don't think I've heard any guests say that. Yeah. It's like it is I think a very important skill, maybe one of the most important skills in the new... people talk a lot about taste.

Host

嗯。

Yeah.

John

但也许这是品味的一种变体,但品味现在就像是另一面。就像严谨性。就像……是的。事实上,如果有的话……

But and maybe this is variation of taste but the taste is like the other side right now. It's like the rigor. It's like the... Yes. Uh yes. In fact, if anything...

Host

嗯。品味还是严谨,哪个更重要?

Yeah. taste or rigor, which one's more important?

John

嗯,我不知道。我认为人们至少在我看来……我的意思是,研究研究人员也是软件开发人员,这在很多领域有很多重叠。我看到软件工程师……几乎就像显然有很多担忧,比如哦不,我该怎么办,你知道编码似乎变得自动化了,所以我认为两者都有,我认为很多人被吸引到,好吧,我将成为创意来源,所以我会尝试发现新科学,我会尝试发现新产品,我会尝试真正非常非常有创意,我再次坚信我认为这不会消失,也有一些人被吸引到严谨性,比如,哦,我想确保这不会崩溃。我想确保这能扩展。我想确保这没有错。我认为你两者都需要。而且我认为你需要真正擅长两者的人。但他们不一定是同一个人。但是的,我认为你需要……我认为这甚至比科学更广泛,就像软件工程的发展一样。是的,将会是,你知道,带来创造力的人和带来严谨性的人,我认为这些将是支柱。

Well, I don't know. I think people at least the way I am viewing it is... I mean the aspect of that research researchers are software developers which there's a lot of overlap in a lot of fields. I'm seeing that software engineers are... it's almost like obviously there's a lot of concern like oh no what am I going to do you know coding seems to be getting automatic so I think there sort of both I think there's a lot of people get pulled into well I'll be the creative source so I'll try to figure out new science I'll try to figure out new products I'll try to sort of really be very very creative and I again I'm strong believer that I don't think that's going to go away there's also people sort of pull towards rigor like, oh, I want to make sure this doesn't crash. I want to make sure this scales. I want to make sure this isn't wrong. I think you need both. And I think you need people who are really good at both. But they don't necessarily have to be the same people. But yes, I think you need... I think this is even broader than science just as sort of software engineering evolves. Yeah, it'll be, you know, the people who will bring the creativity and the people who will bring the rigor and I think those will be sort of anchors.

Kaggle竞赛与过拟合 Kaggle Competitions and Overfitting

Host

我还看到一些其他相关工作。有一个非常酷的排行榜,你知道,比如来自斯坦福的 claw 排行榜、科学问题的智能体排行榜。我不知道你是否熟悉。对我来说,让不同的智能体在排行榜上竞争似乎是一个非常有趣的想法。所以,如果你稍微眯眼看,ERA 所做的有点像排行榜,但它是内部的,并且重新组合想法。而你对这个有什么看法,这是你们正在做的事情吗?这有什么问题或优势吗?

There's some other things that I've seen are related work out there. There's a really cool leaderboard for, you know, like claw leaderboard, agent leaderboard for scientific problems from Stanford. I don't know if you're familiar with it. It seems like a really interesting idea to me to have, you know, sort of different agents kind of competing on the leaderboard. So, it seems like if you squint a little bit, what ERA is doing is kind of a leaderboard, but it's internal and it's recombining ideas. Whereas what are your thoughts about this and do you is that like a thing that you guys are working on and is there problems with that or advantages to that?

John

讽刺的是,你知道,整个 ERA 项目实际上是因为人们可能没有意识到 Kaggle 实际上是 Google 的一部分。哦,是的。所以它被称为自动 Kaggle 问题。所以实际上就是这样,让我们尝试有一个系统可以在 Kaggle 竞赛中获胜。所以这就是为什么它有这样的形状。这就是项目开始的方式。这又回到了过拟合,对吧?如果你曾经真正参加过 Kaggle 竞赛。

Ironically, you know, the whole era project actually started because people may not realize Kaggle is actually part of Google. Oh, yeah. And so it was called the auto kaggle problem. So it was actually like that's what it was is let's try to have a system that can sort of you know win at Kaggle competitions. So that's sort of why it sort of has the shape. That's sort of how the project started. And it goes back to sort of overfitting, right? If you've ever actually competed in a Kaggle competition.

Host

我参加过 Kaggle 竞赛,或者我参加过一次。这是一个非常有趣的现象,因为过拟合非常猖獗。是的。是的。人们能够以某种方式对某些数据集进行过拟合,这真的令人印象深刻。

I have done Kaggle competitions and or I've done one. It is a really interesting phenomenon because there's this overfitting is like rampant. Yeah. Yeah. And it's really impressive how people can overfit to certain data sets in a way that is Yeah.

John

或者我们甚至有一个有趣的项目,如果你愿意,我可以多谈谈,试图减少凝结尾迹。如果你愿意,我很乐意谈谈。我们有一个凝结尾迹竞赛,人们实际上击败了我们,但他们发现我们的标签中有半像素误差,他们……因为这与中心还是左下角有关,比如 00 在哪里?是在像素的左下角还是中心?你明白我的意思吗?所以他们发现了这一点并利用了它,并挤压出一些额外的东西。因为事实证明,当你制作人工数据并旋转它时,你必须确保考虑到那个半像素偏移。所以他们……是的,人们或人们自己会像这些 LLM 一样行事,并试图在这些事情上进行奖励黑客。哦,这又回到了……那是什么?古德哈特定律。你知道,古德哈特定律。任何成为目标的指标都不再是好的指标。是的。所以这就是……我的意思是这很好,只是你必须非常非常小心,你必须再次拥有像严谨性层次,比如好吧,但我们会这样做,我们会为此优化,但你必须意识到好吧,现在古德哈特定律适用了,你必须小心,所以这就是为什么整个人工智能领域一直在不断耗尽这些东西的很多原因,因为再次古德哈特定律适用于你制作的每个排行榜。所以再次,你只需要退后一步,非常非常小心。

Or we even we had a we have a fun project I can talk about more if you like that tries to mitigate contrails gen contrails. I'm happy to talk about that if you want. And we had a contrail cattle competition and people actually beat us, but they found that we had a half pixel error in our labels and they and they cuz it had to do with the center versus the lower left like where is 00? Is it in the lower left of the pixel or is it in the center? You know what I mean? So, they found that and exploited that and and squeezed whatever a little bit extra stuff. Because it turns out when you make artificial data and you rotate it, you have to make sure that you take into account that half pixel offset. So they so yes people or people themselves will be act like these LLMs and and try to sort of reward hack upon these things. Oh, it sort of goes back to um what is it? Goodart's law. You know, good hearts law. Any maybe quote, let's say any metric that becomes a target is no longer good as a metric. Yeah. And so that's the I mean it's good and it's just you have to be you have to be very very careful and you have to again you have to have like layers of rigor like okay but we'll do this and we'll optimize for this but you have to realize okay that's just now goodart's law applies and you have to be careful and so that's a lot of reasons why the whole AI field has been kind of constantly exhausting these these things because again goodart's law applies individually to every to every leaderboard you make. So again, it's sort of you just have to step back and be very very careful.

Host

我可能这不是……我不知道,不,这真的对我的思考很有用。你知道,随着我们邀请嘉宾,这是一个反复出现的主题:如何管理 LLMs 和科学带来的所有复杂性。

I maybe that's not I don't no that that's really useful to my thinking. This is you know as we've had guests on it's been a recurrent theme of how do you manage all this the complexity that's introduced by LLMs and atte science.

Host

我想我的后续问题是关于 Kaggle 中的过拟合。嗯,是的。如果你有一个自动 Kaggle 问题,那么问题是,给定自动 Kaggle,它有多经常成功?我的意思是,我假设你可能只是在所有 Kaggle 竞赛上运行了这个。

I think my followup question was about overfitting in Kaggle. Um, yeah, it is. If you you had an auto autocaggle problem and then the question is given autocaggle, how often was it successful? I mean, I assume you probably just ran this on like all of your Kaggle competitions or something.

John

嗯,我们在各种所谓的游乐场竞赛上尝试过,它在游乐场竞赛中做得非常非常好。嗯,我们进入了不同的竞赛。

Well, we tried it on various um like what they call playground competitions and it did very very well in the playground competitions. Um uh we've entered into different competitions.

Kaggle之外的竞赛 Competitions Beyond Kaggle

John

结果发现,其中一些——过去几年里,各种排行榜和竞赛之类的数量已经远远超出了 Kaggle。所以我们做得很好,其中有一件事我们特别自豪,就是 CDC 设立了这个竞赛,让你预测下周美国每个州和地区将出现的 COVID 和流感病例数。你要提前一周预测,ERA 在这上面做得非常好。

Some of them, it turns out, in the last few years the number of leaderboards and competitions and whatnot has just exploded far beyond Kaggle. So we've done very well, and some of them—one thing we're super proud of is the CDC set up this competition where you try to predict next week's number of COVID and flu cases that will happen in every state and territory in the US. You try to predict a week in advance, and ERA did super well on that.

Host

这挺有意思,因为某种意义上 Google 用 Google Flu 发明了用数据追踪疾病进展这个概念。所以有点好笑,你差不多是 20 年后又绕回了原点之类的。

It's funny because in some sense Google invented the concept of using data to track disease progression with Google Flu. So it's kind of funny that you were sort of going full circle 20 years later or something like that.

John

那个做得非常好。我们参加的其他一些竞赛,我们就不那么出色了,往往是因为——你知道,有时你在这类竞赛里表现多好,取决于你投入了多少心血,以及你愿意不愿意去抠最后那 0.001。所以 ERA 做得不错,你相当接近了,但我们没能在最后那 30 名左右的位置上把差距补上,因为没人去把最后那 0.1 削掉。

So that did very well. Other ones where we've entered, we weren't quite as good, often because—again, you know, sometimes how well you do in these competitions is a measure of how much TLC you put into it and how much you're willing to squeeze the last 0.001. So ERA did well. You got pretty close, but we didn't close the gap in the last whatever, 30 places or whatever, because no one was there to shave the last 0.1 off the thing.

人在回路中 Human in the Loop

Host

它是非常迭代式的吗?就像你做到——我在论文里看到了那些图表,你会看到它发现某个东西时出现这些阶梯式的变化,然后平掉,然后——所以它非常依赖人在回路中吗,比如,好,你在这个问题上卡住了,试试这类东西?

Is it very iterative? Like you get to it—I saw the charts in the paper, and you sort of get these step changes as it discovers something, and then flat, and then—so is it very much human in the loop, like, okay, you've stalled on the problem, try this kind of thing?

John

是的。在外层循环,我觉得那几乎是更有趣、更有创造性的部分。所以是的。哦,来看看这篇论文。哦,你在做坏事,或者你懂我意思吧?所以这几乎就像有一个极度热切、不睡觉的研究生之类的。你告诉它一些东西,你带着它到处走。

Yes. At the outer loop, which I think is almost like the more fun, creative part. So yes. Oh here, have a look at this paper. Oh, you're doing something bad, or you know what I mean? So it's almost like having a hyper-eager grad student or something who doesn't sleep. And you sort of tell it things and you sort of guide it around.

Host

有人介入的频率有多高,对比——外层循环是什么样子的?

How often does someone intervene versus—like what does the outer loop look like?

John

它可能会跑几个小时,然后回来给你一些例子,然后你可以,随你喜欢,你可以继续尝试,继续戳它。

It might run for a few hours and come back and give you some examples, and then you can, as far as you like, you can sort of keep trying and keep poking at it.

Host

所以它很大程度上就是被设计成人在回路中的。

So it's very much designed to be human in the loop, then.

John

是的。是的。

Yes. Yes.

Host

有意思,因为我试过的很多其他工具往往都是一次性的。

Interesting, because a lot of the other tools I've tried tend to be very one-shot.

John

嗯,我想这取决于你的定义,对吧?我是说,显然——你跟它对话,你启动它,它会跑几个小时然后回来。但当然然后你说——但那时就是人类创造力发挥作用的地方,然后你就在做外层循环,你差不多每隔——你知道,取决于你想不想睡觉——但每隔几个小时你就去再试一次。

Well, I guess it depends on your definition, right? I mean, it's obviously—you talk to it and you start it and it'll go for some number of hours and then come back. But of course then you say—but then that's where the human creativity kicks in, and then you're sort of doing the outer loop where you sort of every—you know, depends if you want to sleep—but every few hours you go and you give it another try.

Host

你给这东西多少预算?就像你不小心花掉了一百万美元那种?

What kind of budget are you giving this thing? Like you blew through a million dollars accidentally kind of thing?

John

我其实不知道,因为我们用的是对 Gemini 的内部调用。所以我其实不知道。

I don't actually know, because we're using internal calls to Gemini. So I don't actually know.

Host

但是你看,有 token 预算,但也有——我在解决一个计算上很昂贵的问题。

So but look, there's token budget, but then there's also like I'm solving a problem that is computationally expensive.

John

哦是的,那也是。它本质上——在底层,因为评分函数本身可能有蒙特卡洛估计之类的。是的。所以你实际上最终会为了做一次模拟就用掉大量算力,比如你内部有个模拟器,它就得跑一次模拟。所以是的,你可以花掉相当多的 CPU 或 GPU。

Oh yes, that also. It essentially—underneath it, because the scoring function itself might have Monte Carlo estimation or whatever. Yes. So you actually end up using a lot of compute just even to do a simulation, like if you have a simulator inside, it has to run a simulation. So yeah, you can spend a fair amount of just CPU or GPU.

预算与缩放定律 Budget and Scaling Laws

John

所以我在 ERA 和 Claude 上的小实验,是为一些分类问题构建一个神经网络。显然——如果你有足够的数据,那么更大的网络效果更好,但训练成本更高,你就开始遇到一个问题:如果我的预算固定,我该如何管理预算,好把我的钱花在最有效的方案上。我认为这仍然是我们需要弄清楚的事情。但这当然和一个研究生试图训练一个非常非常大的神经网络或非常非常大的数据集没什么不同。他们自己也得——有某种东西,比如,哦,有没有缩放定律?我能外推吗?所以这是同一个问题,但也许更紧迫,因为它就是会撞上这个问题,因为它是如此不知疲倦。它撞上这个问题的速度比研究生快得多。

So my little experiment with ERA and Claude is to build a neural network for some classification problems. And so they obviously—like if you have enough data, then larger networks work better, but they're more expensive to train, and you start to run into a question of how do I manage my budget if I have a fixed budget, so that I'm spending my dollars on the most effective solutions. And I think that's still something we need to figure out. But it's of course no different than if you have a grad student and they're trying to train a very very large neural network or a very very large data set. They themselves have to—there's some thing like, oh, is there a scaling law? Can I extrapolate? So it's the same problem, but maybe more urgent because it just runs into this problem because it's so relentless. It runs into the problem much quicker than a grad student could.

凝结尾迹问题 The Contrails Problem

Host

你优化的东西之一是凝结尾迹。你能稍微谈谈那个吗?

One of the things that you optimized was contrails. Can you talk a little bit about that?

John

嗯,让我也许花一两分钟谈谈凝结尾迹这个问题。

Well, let me maybe spend a minute or two talking about the contrails problem.

Host

好的,请讲。

Yes, please.

John

作为背景,是凝结尾迹,不是化学尾迹,那是阴谋论。

For context, contrails, not chemtrails, which is a conspiracy theory.

Host

嗯是的。不过你也应该讨厌凝结尾迹,但也许不是因为同样的理由。

Uh yes. Although you should also dislike contrails, but maybe not for the same reason.

John

凝结尾迹就是——如果你见过喷气式飞机后面形成的那些白云,那些叫凝结尾迹。结果发现它们增加了——至少根据人们的估计——大约 1% 的人为全球变暖是由凝结尾迹造成的。为什么会这样?我就谈谈其中的物理原理吧。结果发现其实有两种相互抵消的效应。凝结尾迹——嗯,有时如果你见过它们,它们会拉成条然后消散。那些其实没什么影响。但有时它们会持续很长时间。你会在天空中看到几乎像华夫饼一样的所谓持久性凝结尾迹。它们有两种效应。那些是薄薄的白云。所以它们反射阳光,但那当然只在白天发生。结果发现所有物体都会发出所谓黑体辐射。地球以它当时的温度发出,大约 300 开尔文。它在远红外,大约 10 微米。而在那些波长上,凝结尾迹的反照率非常低。它们几乎本质上是黑的。所以它们会吸收一点向外发出的红外辐射,然后向两个方向重新发射。所以本质上它们会反射一部分向外散失的热量。所以它会像毯子一样把热量困住。而因为这在一天 24 小时都发生,它们往往是增温的。

So contrails are—if you've ever seen those white clouds formed behind jets, those are called condensation trails or contrails. And it turns out they add—at least according to the estimates that people have—about 1% of all anthropogenic global warming is caused by contrails. Why is that? I can just talk about maybe the physics of that. So it turns out that there's actually two countervailing effects. Contrails are—well, sometimes if you've ever seen them, they streak and then they kind of go away. Those don't really do anything. But sometimes they last for a long time. You'll just see in the sky almost like a waffle of just persistent contrails, they're called. And there's two effects that they have. Those are thin white clouds. So they reflect sunlight, but that only of course happens during the day. It turns out all objects emit something called black body radiation. And the Earth does at whatever the temperature is, about 300 Kelvin. It's in the far infrared, around 10 microns. And at those wavelengths, contrails have very low albedo. They're almost essentially black. And so they'll absorb a little bit of the outgoing infrared radiation and then re-emit it both directions. So essentially they'll reflect some of the outgoing heat. So it'll trap heat like a blanket. And so because that happens 24 hours a day, they tend to be warming.

凝结尾迹与气候影响 Contrails and Climate Impact

John

事实证明,这又令人惊讶,存在一些不确定性,但航迹云,即由凝结尾迹产生的卷云,可能覆盖,尤其是在欧洲等航空交通繁忙的地区,实际天空的百分之几被额外的凝结尾迹覆盖。因此,在欧洲这样的地方,它至少局部增加了约每平方米一瓦的辐射强迫。这意味着,为了让你有个概念,全球平均的人为变暖约为每平方米三瓦。所以在飞机交通繁忙的地方,局部变暖可能很大。那么你能做什么?嗯,事实证明,凝结尾迹是由大气中冰过饱和的区域引起的。它们有点像冰糖。就像你制作冰糖时,你得到一种糖分过多的水溶液。任何一点糖都会使所有的糖结晶出来。就像在这凝结尾迹中,这些区域往往呈薄饼状,只有几百米高。如果你飞过它,喷气发动机的尾气中含有少量水分,这些水分会变成液滴然后冻结。然后每克,如果你在这个糟糕的区域,每克水、冰或烟灰你排放出去,大约有 10 公斤的水被吸出。所以有巨大的 10,000 比 1 的固化比例。所以这是一个大问题。所以你能做的是,你可以找出这些区域,它们当然是不可见的,这些冰饱和区域在哪里,然后告诉飞机飞到下面,你只需要下降所谓的两个飞行高度层。所以实际上它消耗一点燃料,但不多,以避免这些糟糕的区域。所以我们建立了一个系统,查看卫星图像并尝试检测凝结尾迹在哪里。所以我们基本上有一个连续的监测系统,然后尝试建立一个模型,因为事实证明天气模型不够准确,无法找到这些冰过饱和的地方。所以我们建立了一个自定义模型,再次像卷积网络或 UNET 之类的,来预测它们将在哪里发生,然后我们将地图提供给飞行规划软件,以便他们可以避开它,并以低成本大幅减少航空对气候的影响。

And it turns out it's surprising again, there's some uncertainty about it, but contrail cirrus, cirrus that sort of comes from contrails, might cover, especially in places like Europe which has a lot of airline traffic, a few percent of the actual sky is covered by additional contrails. So it adds in those places like Europe about one watt per square meter of forcing, locally at least. Which means that, just to give you a sense, all of anthropogenic warming averaged across the whole globe is about three watts per square meter. So in places of high airplane traffic, it can be a lot of warming locally. So what can you do? Well, it turns out contrails are caused by areas in the atmosphere that are ice supersaturated. They're a little bit like rock candy. So like when you have rock candy, you get a water solution that has too much sugar in it. And any little bit of sugar in it will just crystallize all the sugar out. Just like in this contrail, these regions tend to be kind of pancake shaped, only a few hundred meters tall. And if you fly through it, the jet exhaust has a little bit of moisture in it which will turn into droplets and then freeze. And then for every gram, if you're in this bad region, for every gram of water, ice or soot you put out, about 10 kg of water gets sucked out. So there's this enormous 10,000 to 1 curing ratio. So it's a big problem. So what you can do is you can figure out where these regions, they're invisible, of course, these regions of ice saturation are, and then tell the plane to go underneath, and you only have to drop essentially what they call two flight levels. So it actually costs a little bit of fuel but not very much to avoid these bad regions. So we built a system that looks at satellite images and tries to detect where contrails are. So we have essentially a continuous monitoring system and then try to build a model of, because it turns out the weather models are not quite accurate enough to find these places of ice supersaturation. So we build a custom model, again like a convolutional net or a UNET or something, to essentially predict where they're going to happen so that then we give maps to a flight planning software so they can dodge it and inexpensively reduce the climate impact of aviation by a lot.

Host

你能预测的物理原理是什么?对吧。是不是我卫星上看到了,然后明天我觉得它还会在那里,因为飞机去同一个地方,或者

What's the physics behind why you can predict that? Right. Is it just I see it in the satellite and then tomorrow I think it'll be there because planes go to the same place or

John

哦,不。是因为你试图检测这些冰过饱和区域,因为它们非常非常持久。基本上它们

Oh, no. It's because you're trying to detect these regions of ice supersaturation because they're very very persistent. Essentially they're

Host

哦,它们很持久。

Oh, they're persistent.

John

哦,是的。是的。我的意思是,没人确切知道,但它们可以持续数天。基本上,它们是由他们认为的暖湿空气注入到对流层顶边界,即平流层底部引起的。然后当湿度到达那里时,它会长时间停留,然后逐渐消散。

Oh, yeah. Yeah. I mean, no one knows exactly, but they could last for days. Essentially, they're caused by, they think, warm moist air being injected just at the boundary of the tropopause, just at the bottom of the stratosphere. And then when humidity gets up there, it sticks there for a long time and then gradually dissipates.

Host

明白了。所以,所以这只是一些

Got it. So, so it's there's just a sort of

John

它们就像大气中的坏点,你不想飞过。

they're like bad spots in the atmosphere you don't want to fly through.

Host

对。好的。所以一旦你确定了,它至少可能持续几天。

Right. Okay. And so once you've establish that, it's probably good for a couple days at least.

John

呃,好吧,你必须不断预测它在哪里。

Uh well, you have to keep predicting where that is.

Host

是的。你使用的模型,你提到像 CNN 之类的。

Yeah. And the the models you're using are you mentioned like CNN's or something like that.

John

没错。我们还没有用 era 级别的模型替换那些。但出现了一个非常有趣的问题,那就是你想知道这个凝结尾迹造成了多少变暖,以及它增加了多少全球变暖。因为例如,你可能想找到最大的那些,因为有一些燃料成本,也许飞机避开它要花一点钱。所以你说,哎呀,我想知道它造成了多少。但这实际上就是他们所说的反事实问题,就像好吧,你制造了一个凝结尾迹,一定量的红外辐射发生了。所以如果你小心的话,我们可以测量它,但如果没有凝结尾迹,会发生什么。这是一个非常难以估计的事情,因为你无法访问那个

That's right. And we haven't replaced those with era level models yet. But there was a very interesting problem that came up which is you sort of want to know well just how much warming did this contrail make and how much did it add to global warming. Because for example you might want to find the biggest ones because there's some fuel cost and maybe it costs a bit of money for the airplanes to avoid it. So you say well gee I'd like to kind of know how much it did. But that's actually what they call a counterfactual problem like okay you made a contrail and certain amount of infrared radiation happened. So we can measure that if you're careful but would have happened if there hadn't been a contrail there. That's a very difficult thing to estimate because you can't access the universe where the

Host

没有发生的宇宙。

didn't happen.

John

所以你必须制作这些叫做反事实模型的东西。而且那些是,我不知道你的听众是否知道反事实模型实际上很难拟合和制作。呃,我们当时,记得有两个,有两个,一个是反射阳光,然后是红外。事实证明,测量反射阳光的影响实际上更困难,我们实际上被卡住了两年。我们有一个模型,对于 outgoing longwave radiation 的反事实模型效果还行,但对于反射阳光就不行。era 实际上帮助我们找到了一个模型,它搜索了所有的混杂因素,并弄清楚了如何估计它,因为我们再次,我们甚至有测试代码,在人工数据上,因为你可以注入人工制造的人工数据集,其中注入了凝结尾迹,然后弄清楚,哦,我们知道它有多少,因为我们注入了它,所以再次,我们自己的尝试甚至没有通过我们自己的测试,但 era 这个东西实际上通过了,并解决了这个问题。所以,是的,我们正在写一篇论文,我们有一篇关于 outgoing longwave radiation 的论文,但我们有一篇论文,呃,还没有提交,但我们在 EGU 上讨论过,我想,呃,我们实际上解决了这个问题。

So you have to make these things called counterfactual models. And there's those are if I don't know if your listeners know counterfactual models are actually pretty tricky to fit and make. uh and we were there was remember there were two there were two um the reflecting the the sunlight and then there's the infrared. It turns out the uh measuring what the effect of reflecting sunlight is actually more difficult and we were actually stuck on it for 2 years. We we had a model that worked okay a counterfactual model for the outgoing longwave radiation but not for the reflected sunlight. era actually helped us find a uh a model that sort of searched all the confounders and sort of figured out like oh how how can we estimate it because we again we even had like test test code on on sort of artificial because you can kind of inject artificial make artificial data sets where there sort of injected um contrails and sort of figure out oh well we know how much it was because we we injected it and so again our own attempts didn't even pass our own um tests but era this thing actually did and and sort of unstuck this problem. So, yeah, we're in the middle of writing up a paper uh we have a paper about the outgoing long range radiation, but we have a paper um that uh it's not submitted yet, but we've talked about it at uh at EGU, I think, um where um we actually solve this problem.

Host

那么,era 提出的模型是像一大团代码,还是相当基础,只是你需要直觉来开发?

So and and the the models that error comes up with are they just like a big monstrosity of of of code or are they like pretty basic and is just you needed the intuition to develop that?

John

是的,实际上在这个特定情况下,它更像是后者,它基本上帮助识别了,它是一个非常简单的模型,有一些混杂因素,我们只是没有尝试过那种组合,它效果非常非常好。所以是的,它实际上提出了,而且事后看来很明显。所以我认为这是一个巨大的胜利。

Yes, it's actually in this particular case it was actually more of the latter that it essentially sort of helped identify what the it was a very simple model with some just a some number of confounders that we just hadn't we just hadn't tried that combination before and it worked very very well. So yeah, it actually sort of came up with the the and it was seen in in retrospect. So that was I think a a big win.

Host

是的,这很有趣。我知道你在气候方面做了很多工作。你还做过哪些其他事情?

Yeah, that's interesting. I know you've done a lot of work in climate. What other other um stuff have you done?

John

我想我谈过这个,对吧。我谈过 CO2 的事情。那很有趣,因为它仍然是,呃,估计大气中 CO2 是一个有趣问题的原因是我们实际上不知道进出生物圈的碳通量是多少。

I think I talked about this right. The the I talked about the CO2 thing. That was pretty fun because it's this it's still quite a uh the reason why estimating CO2 in the atmosphere is an interesting problem is we actually don't know what the carbon flux is into the bio in and out of the biosphere.

生物圈碳吸收的不确定性 Uncertainty in the Biosphere's Carbon Uptake

John

我们确实知道生物圈会吸收碳。我们向大气中排放大量二氧化碳,其中一部分被海洋吸收,主要通过无机化学过程和一些浮游植物,还有很大一部分被陆地吸收。但关于具体会发生什么的误差棒相当大,而 50 年后的误差棒更是非常大。就像 2100 年的模型——我们不知道生物圈将如何应对不断升高的温度和二氧化碳。所以我们实际上不知道会有多少二氧化碳被吸收,仅仅因为吸收过程的不确定性,误差棒就高达 300 ppm 的二氧化碳。顺便指出,现在大约是 440、450 ppm。所以这个数字是巨大的。我的意思是,情况可能极其糟糕,或者只是不太好,但 300 ppm 是一个巨大的不确定性。所以,如果能弄清楚我们能否减少这个不确定性,那将是非常好的。因此,二氧化碳浓度这项工作就像是朝着解决那个问题迈出的一步。

We do know that the biosphere captures carbon. We emit a bunch of CO2 into the atmosphere, and some of it gets absorbed into the ocean, mostly through inorganic chemistry and some phytoplankton, and a lot of it gets absorbed on land. But the error bars about what happens are moderately large, and the error bars 50 years from now are very large. Like the models for 2100 — we don't know how the biosphere will react to ever-increasing temperatures and CO2. So we don't actually know how much of the CO2 will be absorbed, and the error bars are 300 ppm of CO2 just from the uncertainty of what gets absorbed. And just to point out, right now there's about 440, 450 ppm. So it's huge. I mean, it could be seriously amazingly awful, or not great, but the 300 ppm is an enormous uncertainty. So it would be really nice to figure out, can we reduce that? So this CO2 concentration work is like one step towards solving that.

谷歌气候与天气建模突破 Google's Breakthroughs in Climate and Weather Modeling

Host

我知道谷歌在气候建模和天气预报方面也取得了非常大的改进,对吧?我今年参加了 NeurIPS,就是最近这届 NeurIPS,我去了气候分会场。我可能只听了两场报告之类的,但它真的让我大开眼界——过去大概十年左右,气候建模发生了某种阶跃式变化。我知道其中很多工作是在谷歌完成的。你能谈谈谷歌和其他地方发生了什么,使得气候和天气建模出现了如此巨大的转变吗?

And I know that Google has made some really big improvements in climate modeling and weather prediction as well, right? I was at NeurIPS this year, this last NeurIPS, and I stopped by the climate track. I maybe only had a chance to listen to two talks or something, but it really blew my mind — the sort of step change that's happened in the past, I don't know, maybe 10 years or whatever, in terms of climate modeling. And I know a lot of that happened at Google. Can you talk a little bit about what has happened at Google and other places that has allowed that really big transition in climate and weather modeling?

John

好的,让我来说。人们经常把气候和天气混为一谈,因为它们本质上是相同的物理过程。

Okay, so let me start. People often collapse climate and weather together, because they're fundamentally the same physics.

Host

是的。

Yeah.

John

尽管,至少对于大气物理而言——显然当你开始考虑冰和陆地时,你知道,气候是长期的天气。所以一个完整的地球系统模型,也就是气候模型,其复杂性要比大气模型大得多。比如你必须实际测量进出陆地的水通量和二氧化碳通量是多少,或者冰会发生什么。所以天气模型已经出现了阶跃式变化。抱歉,我想把这点说清楚:天气大约最多能预测 15 天,因为天气本身,或者说大气,似乎是混沌的。

Although, at least for the atmospheric physics — obviously when you start having ice and land, you know, climate is long-term weather. And so the complexity of a full Earth system model, which is a climate model, is much, much bigger than an atmospheric model. Like you have to actually measure what's the water flux and the CO2 flux in and out of the land, or what will happen with ice. So there has been a step change with weather models. Sorry, I want to make this clear: weather is up to 15 days, approximately, because weather itself, or the atmosphere, appears to be chaotic.

Host

我希望——我不知道我是否应该解释一下混沌。

I'm hoping — I don't know if I should explain chaos.

John

本质上就是蝴蝶效应,对吧?微小的扰动——比如蝴蝶扇动翅膀——天气在 2 到 3 周后就会完全不同。所以天气预测是试图预测大气在比如 2 周内的实际轨迹。这已经是一个巨大的阶跃式变化,而且这是因为——这甚至不是新的 LLM 技术,而是基于 2018 年代的机器学习技术,加上大量数据和大量算力。所以有很多非常聪明的工作,其中很多来自谷歌,做出了新的天气模型,这非常棒。事实上,我们有一个非常巧妙的突破,因为现在我们显然可以更准确地提前很多天预测气旋、热带气旋的路径。像牙买加这样的地方曾遭到一场可怕飓风的重创。很多经典模型实际上没有预测到,部分原因是——尤其是强度增强——它完全由海面温度驱动,因为飓风,人们可能没有意识到,本质上是热机。它们将海洋中的热量转化为巨大的大气运动。所以天气方面很棒。气候则困难得多,因为你实际上不关心——你不是试图预测 2070 年西雅图是否会下雨。你试图得到的是平均值。而使其困难的是所谓的非平稳性。所以实际上,字面意义上,底层物理,或者底层的——比如植物行为不同,冰的行为也不同。所以对真正的气候模型使用经典机器学习非常非常困难。这就是为什么你们讨论的描述性模型和多重假设检验在气候领域极其严重,因为我们没有 30 年后的数据,我们也不想等 30 或 50 年才知道我们是否正确,或者我们是否过拟合。所以无论我们做什么,都必须小心,并尝试剥离出子问题。而如何注入——如何构建一个能够预测未来但仍受我们所知约束的大模型——这是一个迷人的问题。我认为它仍未解决,但这是一个很好的问题,因为再次强调,这些不确定性。我们真的很想知道 60 年后气候会发生什么。所以这仍然是个问题。这是一个非常非常有趣的研究问题,但到目前为止,AI 还没有彻底改变它,因为它非常非常抗拒,再次因为数据问题。这是一个低数据问题。

Essentially it's the butterfly effect, right? That small perturbations — like a butterfly flaps its wings — and the weather will be completely different in, you know, 2 or 3 weeks. So weather is trying to predict the actual trajectory of the atmosphere over, say, 2 weeks. And that has been a huge step change, and that's because — that's been a lot of not even the new LLM stuff, that was based on the 2018-era machine learning stuff, and just a large amount of data and a large amount of compute. So there's been a lot of very clever work, and a lot of it from Google, making new weather models, and it's been great. In fact, we had a really neat breakthrough, because now we can apparently predict tracks of cyclones, tropical cyclones, much more accurately many days in advance. And so places like Jamaica had got hammered by a terrible hurricane. And a lot of the classic models didn't actually predict it, partially because — especially the intensification — it's all being driven by the surface temperature, because hurricanes, people might not realize, are heat engines essentially. They convert heat in the ocean to big atmospheric motions. So weather has been great. Climate is much more difficult, because you don't actually care about — you're not trying to predict whether it's going to rain in Seattle in 2070. You're trying to get averages. And what makes it difficult is that it's what they call non-stationary. So in fact, literally the underlying physics, or the underlying — like, plants are behaving differently, and ice behaves differently. And so it's very, very difficult to use classical ML on true climate models. And so that's why the whole discussion you guys had about descriptive models and multiple hypothesis testing is incredibly severe in climate, because we have no data from 30 years out, and we don't want to wait 30 or 50 years to find out whether we were right or that we overfit. So whatever things we do, you have to be careful and try to peel off sub-problems. And the problem of exactly how do you inject — how do you build a big model that can predict into the future but is still constrained by what we know — it's a fascinating problem. I think it's still unsolved, but it's a great problem to have, because of again these uncertainties. We really would like to know what will happen in 60 years to the climate. So it's still a thing. It's a very, very interesting problem to work on, but so far AI has not revolutionized it, because it's very, very resistant, again because of this data problem. It's a low-data problem.

混沌、吸引子与气候预测 Chaos, Attractors, and Climate Prediction

Host

蝴蝶效应,天气的混沌本质,是否也影响气候?还是说时间尺度如此之大,以至于你有一个封闭系统,你知道,也许它在两极之间振荡之类的,但当你从那个时间尺度看时,它更平稳?

Does the butterfly effect, the chaotic nature of weather, also impact climate? Or is the time scale so large that you have a closed system for which, you know, maybe it's oscillating between poles or whatever, but it's sort of — when you look at it at that time scale, it's more stationary?

John

不幸的是,它以另一种方式是非平稳的。但最初整个混沌理论——也许有人提出了它,但在气象学中要追溯到一位叫洛伦兹的人,他有一个非常简单的模型,常微分方程。所以气候和天气的区别是:天气是你在吸引子上的位置,而气候是关于吸引子本身的统计特性。气候的问题在于我们正在改变它。所以吸引子本身在变化,在移动。而且可能存在——大家都谈论临界点——那意味着吸引子突然改变。麻烦在于这非常非常非常难以预测。所以甚至吸引子的形状也会快速变化,或者可能变化。而麻烦在于,当你运行模拟器时,你不知道:它变得不稳定是因为我的模型不好,还是实际的物理不稳定性?而且极其难以区分。

It's unfortunately non-stationary in a different way. But the original whole chaos thing was — well, maybe people came up with it, but in meteorology it was back to a person named Lorenz, who had this very simple model, ODE. So the difference between climate and weather is: weather is where are you on the attractor, and climate is about the statistics itself of the attractor. The problem with climate is that we're altering it. So the attractor itself is changing, is moving. And there could be — everyone talks about tipping points — that means that the attractor suddenly changes. And the trouble is that's very, very, very difficult to predict. So even the attractor's shape changes quickly, or could. And the trouble is when you run a simulator, you don't know: did it go unstable because my model's not great, or is it an actual physical instability? And it's extremely difficult to tell the difference.

稀疏历史数据的挑战 The Challenge of Sparse Historical Data

Host

你该怎么办?尤其是当——对我来说,不仅你没有未来数据,你真的也没有多少过去数据。你可以做一些测量、冰芯和很多东西来尝试做到这一点。但 100 年前没有人拿着仪器。

What do you do? Especially when — I mean, to me it strikes me that not only do you not have future data, you really don't have much past data. You can do some measurements and ice cores and lots of stuff to try to do that. But there was nobody with an instrument 100 years ago.

气候模型与过程模型 Climate Models and Process Models

John

所以如果你有年度数据之类的,也许运气好的话,在任何一个地点能有 50 个数据点,对吧?所以我们做任何事情都必须受到已知信息的严格约束。但这真的很难。我只是在告诉你人们所处的两难困境。人们会做这些——事实上,整个应用科学领域的人,我觉得气候是最极端的——会做这些叫做过程模型的东西。你做的事情是——我见过那些代码——好吧,你知道,我要做一个还原论者,我要把那个可怕而复杂的气候问题拆解成一千个碎片,然后我要找到那个写了关于第 763 号碎片论文的专家,他把某个数据拟合成了一个三次函数。比如说,有一件非常神秘的事情,与凝结尾迹有关,就是冰在云中如何表现?你深入研究会发现,一切都很复杂。但结果是,当你制造一条凝结尾迹时,它能持续多久?这取决于——因为凝结尾迹蒸发的方式是冰开始积累,就像我说的,然后冰晶变大,然后它们下落。但当然,它们下落的速度取决于它们的形状,而形状是未知的。凝结尾迹从内部潮湿区域向外混合的程度有多大?同样,人们有近似值,但他们不知道。所以不确定性非常混杂。而且这不仅仅是,哦,约翰关心凝结尾迹。结果是,冰的微物理学的实际物理特性对气候模型的行为有非常强的影响。而我们就是不知道。所以我想说的是,这非常棘手,不是一个已解决的问题。

So if you have annual data or whatever, maybe you have, if you're lucky, 50 data points in any one location, right? So whatever we do has to be very constrained by what we know. But it's just very difficult. I'm just telling you the horns of the dilemma people are on. People make these—in fact, people in applied science in general, I would say climate is the most extreme—make these things called process models, where what you do—and I've seen the code—oh well, you know, I'm going to be reductionist and I'm going to take the horrible complicated climate thing and boil it down to a thousand pieces, and then I'm going to find the expert who wrote a paper about piece number 763, and he fit a cubic to some data. For example, one thing that's very mysterious, which is related to contrails, is how does ice behave in clouds? It turns out you might again get—everything is complicated once you dig into it. But it turns out that when you make a contrail, how long does it last? Well, it depends on—because the way contrails can evaporate is ice starts to accumulate, as I said, and then the ice crystals get big and then they fall. But of course, how quickly they fall depends on their shape, which is not known. And how much does the contrail mix from the moist inside the contrail out? Again, people have approximations, but they don't know. And so the uncertainty is very much confounded. And it's not just, oh, John who cares about contrails. It turns out that the actual physics of the microphysics of ice has very strong implications about what climate models do. And it's sort of—we just don't know. So I'm trying to say it's very gnarly and it's not a solved problem.

John

我的希望是,有了工具——也许不是今天这个时代,而是明天这个时代,因为记住,正如我之前说的,它不仅能够拟合数据,还能阅读论文,对吧?问题是,它能读的论文比我们多得多。所以也许我们可以整合所有数据,或者所有人们仔细评估过的知识,比任何一个人写代码和拟合数据所能做到的要多得多。我的意思是,那将是无比辉煌的。我们今天没有这个,但这是我对未来更神奇工具的希望之一——真正能够以合理的方式写代码,甚至更加合理,因为它会受到我们迄今为止积累的所有科学知识的约束。那将是惊人的。我们今天没有这个。

My hope is that with tools—maybe not like today's era but maybe tomorrow's era, because remember, as I was saying before, it not only can fit data, it can read papers, right? And the question is, it can read a lot more papers than we can. So maybe we can integrate all the data or all the knowledge that people have carefully evaluated, much more than any one person writing a piece of code and fitting data. I mean, that would be utterly glorious. We don't have that today, but that's sort of one of the hopes that I have even for a more amazing tool in the future—something that really can write code in a sane way, even much more sane because it'll be constrained by all the scientific knowledge that we've accumulated so far. That would be amazing. We don't have that today.

Host

我的意思是,我觉得这跟生物学非常相似。

I mean, that strikes me as being very similar to biology.

John

哦,是的。天哪。对。如果你曾经接触过生物学,甚至只是看看生物学,生物系统中有太多的例外和太多的权宜之计。是的。天哪。所以,是的。如果我们能有一个东西,真正将所有的已知科学知识与数据整合起来,并尝试合成新的模型和新的东西,那将是惊人的。

Oh, yes. Oh, boy. Right. If you've ever played with biology or even looked at biology, there are so many exceptions and so many hacks in the biological systems. Yes. Oh boy. So, yeah. It would be amazing if we could have a thing that could really integrate all known scientific knowledge with data and try to synthesize new models and new things.

Host

我想你说的是,AI 在某种程度上可以成为这里的解锁钥匙,因为模型是如此零碎。它们必然是零碎的。所以能够拼装这个拼图——不是要混用比喻,而是拼装这个拼图——仅仅拥有规模和容量实际上就有很大帮助。

I think what you're saying is that AI can be an unlock here to some extent because the models are so piecemeal. They necessarily piecemeal. And so being able to assemble the jigsaw puzzle—not to mix metaphors, but to assemble like a really this jigsaw puzzle—having just scale and capacity actually helps a lot.

John

没错。这些 AI 拥有的一件事是,不知何故——你知道,人类,即使我读了很多书,我觉得,但对我来说,要整合 n 平方个不同的——我的 n 是我一生中读过的论文数量。它相当大,对我来说,即使做那个 n 平方的事情也很困难。但不知何故,在那些数十亿、数十亿的参数中有如此多的数据,而且你还可以让它访问 PDF,它就能以某种方式开始把人们不会整合的东西整合起来。所以这又是——我开始在 ERA 内部看到一些小小的迹象。我不是说 ERA 今天就能做到这一点。但是的,这就是我对未来走向的希望。

That's right. The one thing that these AIs have is somehow—you know, humans, even I'm pretty well-read, I think, but it's just difficult for me to kind of integrate across the n squared different—my n is the number of papers I've read in my life. It's pretty big, and it's just difficult for me to even do that n squared thing. But somehow there's just so much data in those billions and billions of parameters, and you can also give it access to read PDFs, that it can somehow start to pull things together that people wouldn't do. So that's again—I'm starting to see little indications of that inside of ERA. I'm not claiming that's what ERA does today. But yeah, that's sort of my hope of where this is going to go.

Host

我想我听过很多人提出,通往智能的路径是将 LLM 与某种形式的搜索结合起来。所以实际上,令人惊讶的是,ERA 就在这样做,对吧?某种非常强大的数据库查询加上好的搜索算法是一种方式。当然,我的意思是,还有整个——我想人们仍然在做整个 RAG 的事情,当然。如果你想想,谷歌本身,你知道,那 10 个蓝色链接的东西,

I think I've heard a lot of people suggest something like the route to intelligence is to combine LLMs with some form of search. So it's actually amazingly like ERA doing that, right? Something which is maybe a very strong database lookup with a good search algorithm is one way. And of course, I mean, there was the whole—I mean, people still do, I guess, the whole RAG thing, of course. And if you think about it, Google itself, you know, the 10 blue links thing,

Host

它曾经是,或者说现在也是一种 AI 形式,甚至在我们有 LLM 之前,对吧?因为它就像——你可以把自己带回 2010 年或 2015 年。你可以问谷歌任何东西,它都会告诉你一些东西,对吧?所以,

it was or is a form of AI even before we had LLMs, right? Because it was like—you could cast yourself back to whatever 2010 or 2015. You could ask Google about literally anything and it would tell you stuff, right? And so,

John

实际上,令人惊讶地好。

Surprisingly well, actually.

Host

是的,令人惊讶地好,因为互联网上可能有人写过关于它的东西。所以如果你能匹配那个——

Yeah, surprisingly well, because somebody on the internet has written about it probably. So if you can match that—

John

事实上,这就是我想来谷歌的原因之一——那真是一件了不起的事情,对吧?

In fact, that was one of the reasons why I wanted to come to Google—it was just that was such an amazing thing, right?

Host

我想知道我们有多少观众在谷歌之前做过搜索,知道那种体验有多糟糕。是的,我记得 1998 年,我想——我想当谷歌——就像我在用 AltaVista,然后我——我不知道,谷歌刚刚发布,我用了它,我就——抱歉,数字——我就像扔掉烫手山芋一样扔掉了它,立即开始使用谷歌。不,它是如此——那是一种 AI 形式。所以是的,它可能是——是的,可能只是能够访问所有这些,并同时将其记在脑中——

I wonder how much of our audience did a search pre-Google and just know how bad that experience was. Yeah, I remember 1998, I think—I think when Google—like I was using AltaVista and like I—I don't know, Google had just gotten released and I used it and I just—sorry, a digital—I just dropped it just like a hot potato or something and started immediately using Google. No, it's so—that is sort of a form of AI. And so yes, it might be—yeah, it could be that just having access to all of that and sort of keeping it in mind at the same time—

Host

那种科学发现的模型,只要它能成功,也让人感到安慰,因为它是还原论的,所以你可以查看各个部分并理解它们。所以它找到了要组装的确切东西,但它们实际上可能从根本上都是人们发明或迭代过的东西,所以所有那些小碎片都是可以单独理解的,然后你也可以把它们组合成一个连贯的图景。我认为对于生物学或气候科学这样的许多问题,除非答案以那种形式呈现,否则我们不会信任它。

that model of scientific discovery, in as much as it pans out, is kind of comforting too, because it is reductionist, so that you can look at the individual parts and understand them. So it found the exact things to assemble, but they're all actually maybe fundamentally things that people have invented or done iterations on, and so all those little pieces are individually understandable, and then you can also put them together into a coherent picture. I think for a lot of problems like biology or climate science, I don't think we would trust the answer unless it was in that shape.

John

是的。

Yeah.

Host

因为如果有一个巨大的黑箱模型说:“哦,这就是细胞的工作原理。”那就像,我相信它吗?我的意思是,我不知道我是否相信它,因为我无法检查它。所以——

Because if there was some giant black-box model that said, "Oh, this is how a cell works." It's like, do I believe it? I mean, I don't know if I believe it because I can't examine it. So—

John

但我的意思是,要反驳它的话,如果它真的运行得很好——

But I mean, to argue against it though, like if it works really well—

Host

但你必须收集——你必须——我的意思是,你显然必须测试它,但一个统计模型,你又得——它必须外推。

But you'd have to gather—you'd have to—I mean, you have to test it obviously, but a statistical model, you have again—it has to extrapolate.

John

是的。

Yeah.

Host

是的。而且它必须外推到极端情况或那种黑天鹅事件。

Yeah. And it has to extrapolate to the extreme or the sort of the black swan events.

John

没错。所以这非常困难。这就是为什么像自动驾驶汽车这样的事情非常非常——这是一个非常困难的问题。

That's right. And so it's very hard. This is why things like self-driving cars are very, very—it's a very difficult problem.

数据驱动模型与可解释模型 Data-driven models vs. interpretable models

Host

全都是边角案例。它们做得这么好,真是令人惊讶。这让我想到一个有意思的点:从物理学的世界来看,模型通常是一个方程或少数几个方程,它们唯一地定义一个系统及其一切,你只要去解这个系统,就能知道你需要知道的一切。而像 AlphaFold 这样的东西对很多人来说是一种转变——以前他们认为,蛋白质折叠这个问题,只要找到正确的力场、有正确的计算引擎,就能解决蛋白质折叠。而用数据驱动的方式真正解决它,这个想法在 AlphaFold 1 出现前几年才出现。有意思的是,它迫使人们几乎重新评估什么是科学,因为 AlphaFold 和类似的模型极其强大。它们作为工具打开了很多东西,但在核心上,它们往往不能像大多数物理学家历史上所希望的那种物理学那样给出直觉。所以我想,有句老话说:所有模型都是错的,但有些是有用的。

It's all corner cases. It's kind of amazing how well they've done. That's an interesting point, thinking about coming from the world of physics, where a model was usually a single equation or a small number of equations which uniquely define a system and everything about it, and you just crank, you just find a solution to the system, and you now know everything you need to know. And I think something like AlphaFold was kind of a shift for a lot of people, where before they thought, oh, protein folding is a problem where if we find the right force field and we have the right computational engine, we will solve protein folding. And the thought of even really solving it in a data-driven way only appeared a few years before AlphaFold 1 came out. And it's interesting that it's sort of forced people to re-evaluate almost what is science, because AlphaFold and similar models are incredibly powerful. There's a lot of things that they've opened up as tools, but at their core they often don't give intuition in nearly the same way that, let's say, the physics that most physicists historically would have wanted. And so I guess there's this old saying: all models are wrong, some are useful.

John

这是 Box 说的。

Box said that.

Host

对。你觉得什么时候数据驱动的模型就足够了,什么时候你需要某种可解释的、人类真正能理解的东西?

Yeah. When do you find the data-driven models to be sufficient, and when do you want something which is interpretable that humans can actually understand?

John

我觉得这几乎可以归结为天气和气候的区别。如果你处在数据丰富的领域,比如天气,甚至因为 PDB 而数据丰富的蛋白质,你就能感觉到:是的,换句话说,我有足够的数据来覆盖它,所以像 AlphaFold 这样的统计模型就能做到。事实上很多人很乐意用 AlphaFold。我觉得它真的革新了我的理解——我不是生物化学家,但人们似乎很喜欢它。他们做的一件了不起的事是,他们穷尽式地在所有 PDB 上运行了它并发表了结果,这真的非常非常酷。

I think it boils down to almost like the difference between weather and climate. If you're in a data-rich regime, like weather, or even proteins because of the PDB, you can feel, oh yes, in other words, I've got enough data to kind of cover it, and so a statistical model like AlphaFold can do it. And in fact a lot of people happily use AlphaFold. I think it's really revolutionized my understanding—I'm not a biochemist, but people seem to love it. And one amazing thing they did is they exhaustively just ran it on all PDB and published it, which is just really, really cool.

Host

大概有六十亿个蛋白质结构。所以绝大多数其实相当准确。

It's like six billion protein structures. So the vast majority are actually quite accurate.

John

对。所以这真的很了不起。但它感觉是封闭的,如果你明白我的意思。但当它像气候那样是开放的、非平稳的,或者你必须做这些大的外推时,你就必须谨慎得多。或者也许又是生物学——生物学里可能有某些部分,比如 Arc Institute 有个虚拟细胞挑战赛,结果挺有意思。

Yeah. So that's just amazing. So but it feels closed, if you know what I mean. But when it's like climate and it's open and it's non-stationary, or you have to make these big extrapolations, you have to be much more cautious. Or maybe biology again—there may be parts of biology, like there was this virtual cell challenge from the Arc Institute that had a funny result.

Host

我知道人们——可能有一些过拟合,至少那会说明更多,抱歉。

I know that people were—there may have been some overfitting, at least that would say more, sorry.

John

或者我想也许在高层次上,我觉得简单的基线效果非常非常好。

Or I guess maybe at a high level, I think the simple baselines work very, very well.

Host

对,对,就像你做生物学时的一个经典做法,就是永远从简单的基线开始。也许这大概就是好的机器学习的一般原则——理解你最简单的情况。而在生物学里有很多问题,即使你有很多数据,它们也极其抗拒任何超出简单基线的东西。

Yeah, yeah, just like one of the classic things in whenever you do biology is just always start with a simple baseline. Maybe this is probably just good ML in general—understand your simplest case. And in biology there are many problems where they're extremely resistant to anything beyond the simple baseline, even if you have a lot of data.

John

没错。事实上我也这么告诉别人。我说永远先拟合线性回归,就拟合——

That's right. And in fact I tell people the same thing. I say always just fit linear regression, just fit—

Host

就去做。

Just do it.

John

就去做。就做线性——

Just do it. Just do linear—

Host

或者 SVM。我是说 SVM 只是线性回归的另一种形式。

Or SVMs. I mean SVMs are just a different form of linear regression.

John

对。那什么时候你需要更偏过程模型的东西?我觉得就是当你——我是说,气候在一端,我不知道,天气也许在另一端,也许那太极端了——但我觉得就是你在数据丰富度上处于什么位置。什么时候你能感觉到,哦不,我确实有一个封闭的问题,而且我觉得我实际上能覆盖它。

Yes. So when do you need the more process-model-y thing? I think it's just when you have—I mean, sort of climate is on one end, and I don't know, weather maybe on the other end, maybe that may be too extreme—but I think just where are you on the data-richness thing. When can you feel like, oh no, I really have a closed problem, and I think I can actually cover it.

Host

对,一个你的数据完全覆盖的封闭问题。对,我觉得这在我所见到的里面也很说得通。

Yeah, a closed problem that your data fully covers. Yeah, I think that makes a lot of sense in what I've seen as well.

气候建模的目标 Goals of climate modeling

Host

对,我有点好奇的一件事是,当你在做气候建模时,你谈到了凝结尾迹,你谈到了,我猜,二氧化碳预测——你试图达成的广泛目标是什么?其中一个我猜是做干预,另一个可能是为保险之类的事情做预测,或者你如何帮助针对某种气候变化做调整?我想,对你个人或者对整个领域来说,主要目标是什么?

Yeah, I think one thing I'm kind of curious about is, when you're working on climate modeling, what are the—you talked about contrails, you've talked about, I guess, CO2 predictions—what are the broad things you're trying to accomplish? So one of them is, I guess, making interventions, and the other one might be making predictions for things like insurance, or how do you help adjust for some sort of climate change? What are the principal goals, I guess, for you specifically or the community at large?

John

我觉得,你知道,就像任何领域一样,大概有很多不同的目标。对我、对我的团队来说,我们对干预非常非常感兴趣。比如哪些干预在相对成本下是可行的。我是说,凝结尾迹挺了不起的,因为事实证明这个干预成本相当低。而且凝结尾迹的一个了不起之处是它们是局部的,不像二氧化碳之类的东西,因为如果一个国家决定修复自己上空的凝结尾迹,它实际上改善了自己的——我是说,它有全球影响,但它主要是在自己上空稍微改善气候。所以他们喜欢这一点。

I think, you know, just like any community there's probably many different goals. For me, and my team, we're very, very interested in interventions. So like which ones are possible at relative cost. I mean, contrails was kind of amazing because it turns out that the intervention is quite low cost. And also one amazing thing about contrails is they are local, unlike things like CO2, because if a country decides to fix contrails over itself, it actually improves its—I mean, it has global effects, but it mostly improves the climate a little bit over themselves. So they like that.

Host

不过我猜,如果你处在寒冷气候里,想把它变暖,这就成了你自己的——你可以自己掌控它。

I guess though, if you are in a cold climate and you want to warm it up, this is now your own—you could own it.

John

对。事实证明这有点不对称。变暖本质上是恒定的,而且是全球性的。降温只在太阳以好的角度照在你上方时才发生。所以很少——变暖方面有不确定性界限。有很多很多凝结尾迹,当然大多在夜间,那里主要是变暖,而且我们非常确定,大概两个标准差。没有多少凝结尾迹是你能说,哦,我确定它在降温,我想要更多。只有在极地、极地夏季,你才知道凝结尾迹在降温,因此如果消除它们就会变暖。但南极洲上空基本没有航班,夏季极地上空也没那么多。

Yeah. It turns out that it's a little bit asymmetric. The warming is constant essentially and global. The cooling only happens when the sun is at a good angle over you. So it's very rare that—there are uncertainty bounds in terms of the warming. There are many, many contrails, mostly at night of course, where it's largely warming, and we're very sure in terms of like two sigma. There are not very many contrails where you say, oh, I know for sure that it's cooling and I want more of them. So only over the poles in polar summer do you know that the contrails are cooling, and therefore if you got rid of them they would warm up. But there are essentially no flights over Antarctica and not that many over the poles in the summer.

Host

所以没人会——住在寒冷气候里的人不会恶意利用这个。

So no one's gonna—no one who lives in a cold climate is going to use this maliciously.

John

嗯,是的,在那个——嗯,他们不会确定它是在变暖还是降温,所以他们会做一些事。所以大多数情况下我们只是忽略——我们不建议人们飞那些航线。

Well, yes, in that—well, they wouldn't know for sure whether it was warming or cooling, and so they would do stuff. And so mostly we just sort of ignore—we don't recommend that people fly those.

Host

你还提到了一个关于经济学的有意思的点。

You also brought up an interesting point about the economics.

气候解决方案与市场 Climate Solutions and the Market

John

我是说,历史上人们对某些气候变化干预措施有很多抵触,但从某种意义上说,市场已经接管了。到现在,毫无疑问,可再生能源和电池几乎普遍优于替代方案。

I mean, there was a lot of resistance historically about certain climate change interventions, which in some sense the market has just taken over. At this point, unambiguously, renewables and batteries are just almost universally better than alternatives.

Host

对于非移动的,我是说,你指的是移动的?

For non-mobile, I mean, you mean mobile?

John

是的,这是个很好的观点。比如飞机,我们还没有解决方案。

Yeah, that's a really good point. Like planes, we do not have a solution.

Host

没错。我是说,有一些电池驱动的飞机,但它们非常小,航程也非常有限。

Correct. I mean, there are some battery-powered planes, but they're very small and have very limited range.

John

它们可能永远不会真正……很难想象物理上会非常非常……除非我们发明出类似核电池的东西,那会很棒,但我们不知道怎么做。

They probably will never actually be... It's hard to imagine the physics would be very, very... unless we came up with something like nuclear batteries, which would be kind of amazing, but we don't know how to do that.

Host

或者即使我们做到了,我认为风险是人们会太害怕核电池出问题之类的。

Or even if we did, I think the risk is that people would be too afraid of a nuclear battery going wrong or something.

John

哦,是的。是的。既然我们不知道它们是什么,我们就不知道风险是什么。我想我们没有风险。是的。

Oh yeah. Yeah. Since we don't know what they are, we don't know what the risk is. I guess we don't have the risk. Yeah.

Host

是的。所以我们不知道。所以是的,这就是问题所在……我谈论过……我做过关于气候变化的演讲,我谈到“糟糕的饼图”、“悲伤的饼图”,

Yeah. So we don't know. So yeah, that's the problem with... I talk about... I have given talks about climate change and I talk about the pie chart of badness, pie chart of sadness,

John

也就是说,气候变化没有单一的银弹,对吧?我们的经济中产生温室气体的来源太多了。所以它们都必须被解决,或者很多很多都必须被解决。所以没有单一的东西。我是说,我研究过核聚变。核聚变很酷,如果它足够便宜,可能真的能解决很多问题,但我们不知道,因为我们还不知道它是否可行。

which is there's no one silver bullet for climate change, right? There's so many different things that contribute greenhouse gases just from across our economy. So they sort of all have to be fixed or many, many of them have to be fixed. So there's no one single thing. I mean, I've worked on fusion. Fusion is cool and it might actually knock a lot of them out if it's cheap enough, which we don't know because we don't know if it'll work yet.

Host

我是说,核聚变是那种有趣的东西,笑话总是说核聚变永远还要 30 年,但我认为现在实际上可能不到 30 年了。

I mean, fusion is one of those interesting things where the joke was always fusion is 30 years away, but I think it's actually now less than 30 years away, maybe.

John

是的。不,我认为很有可能有人甚至在这个十年结束前就能实现商业相关的核聚变。所以,我认为它可能还有三年,而不是 30 年。

Yeah. No, I think there's a definite probability that someone will make commercially relevant fusion even by the end of this decade. So, I think it's like three years away, not 30 years away.

Host

这非常真实。有趣的是,我认为很多……我现在要宣传一下我们的其他几集,但很多实际上归结为材料科学。有趣的是……

This is very real. Interestingly enough, I think a lot of that... I'm going to now just stump or advertise some of our other episodes, but a lot of it actually comes down to material science. Interestingly enough in that...

John

嗯,我……哦,太好了。当然。当然。当然。嗯,是的。抱歉,如果你愿意,我们可以谈谈核聚变……

Well, I'm... Oh, super. Sure. Sure. Sure. Well, yes. Sorry, we could talk about fusion if you...

Host

是的。是的。实际上有两件事。一是核聚变。另一个是更好的控制系统,我认为这实际上……

Yeah. Yeah. There was actually two things. One is fusion. The other is better control systems, which I think is actually...

John

是的。事实上,Google DeepMind 一直在研究托卡马克的控制系统,以确保它们不会本质上变得不稳定并发生破裂。是的,没错。

Yes. And in fact, Google DeepMind has been working on control systems for tokamaks to make sure they don't essentially go unstable and go disrupt. Yeah, that's right.

Host

破裂本身也相当有趣。

Disruptions are quite interesting in themselves.

John

是的。

Yes.

Host

是的。基本上整个……托卡马克中的所有能量汇聚成一个小束,然后它击中你的真空室,你非常非常难过。

Yeah. It's basically the entire... all energy in the tokamak columnates into one little beam and then it hits your vacuum chamber and you're very, very sad.

John

非常难过。

Very sad.

Host

是的。是的。我认为人们相信 ITER 在花费 300 亿美元后可以开启,然后破裂,基本上就变成一个 300 亿美元的砖块之类的。

Yeah. Yeah. I think people believe that ITER could be turned on after $30 billion, disrupt, and then basically have a $30 billion brick or something.

John

哦,是的,我想你可以试着修补它。我记得工作……我们再次,在 LM 之前,我们与一家名为 TAE 的核聚变公司合作,我在他们的控制室里,是的,有点难过。你必须非常小心。我们在做系统来推荐新实验,他们非常非常怀疑和偏见,不知道该往哪个方向走,因为即使在人类控制下,就像……他们在做某个实验,然后你听到一声巨响,就像“哦不”,然后你知道,然后设备停机两周,他们修补一些……

Oh yeah, I guess you could try to patch it. I remember working... we again, before LM, we worked with a fusion company called TAE and I was in their control room and yes, it was kind of sad. You have to be very careful. We were making systems to recommend new experiments and they were very, very skeptical and jaundiced about which way they should go because even under human control it's like... they were doing some experiment, then you hear this big bang and it was like oh no, and then it's like you know, then the apparatus is down for two weeks as they patch some...

Host

你在破裂期间在场?

You were there during a disruption?

John

哦不,这是……抱歉,他们有场反位形。

Oh no, this is... sorry, they have field reversed configuration.

Host

哦,好的,那有它自己的……我是说,你知道,有一些电弧。那么,那是什么?抱歉,我不熟悉。

Oh okay, which has its own... I mean, things you know, there's some arc. So, so what is that? Sorry, I'm not familiar.

John

哦,哦,什么是场反位形?嗯,事实证明托卡马克并不是……尽管它们可能是研究最多的等离子体形式。有许多不同种类的架构,本质上试图稳定和压缩等离子体的方法。有一种形状,本质上是一种自包含的橄榄球状等离子体,称为场反位形,其中内部和外部的磁场是相反的。所以,它们被一种称为分界面的东西分开。这在理论上是不稳定的,但在实践中是稳定的。例如,当你运行磁流体动力学 MHD 代码时,在那个假设下它是不稳定的,但那个假设并不是现实世界的工作方式。所以是的,它多年来不太受青睐,但 TAE 和其他人,我认为 Helion,有 FRC,因为它们实际上相对稳健。你实际上可以把它们撞到墙上,它们仍然保持稳定。是的,但你仍然可能得到放电和东西,在你的真空室上打洞,这有点不幸。

Oh, oh, what's the field reversed configuration? Well, it turns out tokamaks are not... although they're perhaps the most studied form of plasma. There's many different kinds of architectures, essentially ways to try to stabilize and compress plasmas. There was a shape, essentially, it's essentially a self-contained football plasma called a field reversed configuration where essentially the magnetic field inside and outside are opposite. So, they're separated by something called a separatrix. And that is sort of in theory unstable but in practice stable. Like for example, when you run magnetohydrodynamics, MHD code, it's unstable under that assumption, but that's an assumption that's not the way the real world works. And so yeah, it was kind of disfavored for many years, but TAE and other people, I think Helion, have FRCs because they are actually relatively robust. You can actually knock them against walls and they'll still stay stable. And yes, but you can still get discharges and things that punch holes in your vacuum chamber, which is kind of unfortunate.

Host

为了澄清,所以你有这些核聚变反应堆。它们是或试图成为反应堆,也许是装置。装置,你创造等离子体。等离子体是磁性的……

For clarification, so you have these fusion reactors. They are or trying to be reactors maybe and apparatuses. Apparatuses and you create a plasma. The plasma is magnetically charged...

John

或者被约束。是的。

Or confined. Yes.

Host

或者被约束。所以它被磁场约束。所以你有某种磁系统,可以由计算机调节,然后计算机试图维持约束。

Or confined. So it's confined by a magnetic field. So you have some sort of magnetic system that is tunable by a computer and then the computer tries to kind of maintain the confinement.

John

嗯,FRC 一旦你制造出来,它们就有点持续。有不同的方法试图确保你……好吧,所有核聚变都归结为所谓的劳森判据。本质上……它解释了为什么核聚变很难。本质上,你可以很容易地在信封背面展示密度、温度,以及本质上能量损失,称为约束时间。它是能量在等离子体中衰减到 1/e 所需时间的倒数。所以这三个数字的乘积必须大于某个常数,然后你才能得到核聚变,如果不满足,就不能。这解释了它是三个数字的乘积这一事实,解释了为什么核聚变如此困难,因为每种方法都有一个阿喀琉斯之踵,其中一个数字不是很大,然后他们拼命试图让它更高。

Well, FRCs kind of once you make them they're sort of sustained. There's different ways of trying to make sure you... Okay, so all of fusion boils down to something called the Lawson criteria. There's essentially... and it explains why fusion is hard. Essentially, you can just very easily on the back of an envelope just show that the density, the temperature, and essentially the energy loss, it's called the confinement time. It's one over the amount of time it takes for the energy to decay away, 1 over e in a plasma. So the product of those three numbers has to be bigger than some constant and then you can get fusion and if you don't then you don't. And that explains the fact that it's a product of three numbers explains why fusion is so hard because every approach has an Achilles heel where one of those numbers is not very big and then they try to desperately make that be higher.

Host

而且每种方法都不同。

And every approach is different.

聚变进展与警示 Fusion Progress and Caution

John

每种聚变方法都各不相同,当出现那些令人喘不过气的聚变新闻时,你们很多人得保持一点怀疑,因为它会说,现在约束时间稳定了 X 分钟之类的,它只谈了三个数字中的一个,但你必须三个数字都达标才能实现聚变。我认为整个领域正在取得很大进展,非常令人兴奋,但你对那些只谈一个数字的耸动新闻文章确实得谨慎一点。

Every approach to fusion is kind of different, and a lot of you have to be a bit skeptical when there are all these breathless news things about fusion, because it'll say, you know, now confinement time is stable for X minutes or whatever, and it's talking about one of the three numbers, but you have to have all three numbers before you can get fusion. I think the whole field is making a lot of progress and it's very exciting, but you do have to be a little bit cautious about the breathless news articles that only talk about one number.

Host

那么其中的计算部分是什么?

So what is the computational part of that?

John

哦,不幸的是,不管好坏,这取决于具体方法。对于托卡马克,正如布伦丹所说,它基本上是稳定的,除了偶尔会出现这种不稳定性,把所有能量都集中到一个地方猛击,所以你必须让一切处于控制之下。所以它是一个控制系统。而 FRC 本身有非常简单的不稳定性。例如,它们有一种叫做 Z 不稳定性的东西。所以没关系。它是稳定的。它只会晃动。它真的会来回晃动,但你只要做一个所谓的 PID 控制器,让那个橄榄球保持在反应堆中心,一切就都没问题了。

Oh, unfortunately, for better or worse, it depends on the approach. So for tokamaks, as Brendan said, it's mostly stable except that there's occasionally this instability that takes all the energy and smacks it into one place, and so you have to keep everything under control. So it's a control system. FRCs themselves have very simple instabilities. So, for example, they have what they call a Z instability. So it's fine. It's stable. It'll just wobble. It'll literally wobble back and forth, but you just make what they call a PID controller that keeps the football in the center of the reactor and things are fine.

Host

它是通过调整磁场来做到这一点的。

And it does that by adjusting the magnetic field.

John

是的。我认为它实际上调整的是电场。它把它来回推。人们遇到的问题真的取决于他们决定采用哪种等离子体架构。

Yeah. It actually adjusts, I think, the electric field. It sort of knocks it back and forth. The issue that people have really depends on which plasma architecture they're deciding to use.

气候经济学与干预 Climate Economics and Interventions

John

气候问题在某种程度上是政治性的,基本上是因为经济,可能主要是,也许还有其他因素,但它的经济性,你知道,你必须说服人们以某种方式多花钱,或者你必须有一个解决方案,恰好既在经济上更好,又对气候更好。这很难。

Climate is sort of political because of economics, basically, probably mostly, maybe other stuff, but the economics of it, you know, you have to persuade people to somehow spend more, or you have to have a solution that has this happy coincidence where it's both economically better and better for the climate. That's hard.

Host

是的。在很多情况下,它并不是。我的意思是,很多情况下它并没有得到回报,但……

Yeah. In many cases, it's not. I mean, many cases it hasn't been earned, but...

Host

嗯,你想想,比如,甚至预测天气,对吧?你可以做准备,你能看出那在经济上可能有益。那么你在干预方面做了什么工作,那又如何与经济相互作用?听起来像凝结尾迹那个——我在一项分析中做了,并说实际上这很棒,因为它的经济影响很低但价值很高。

Well, you think about like, okay, predicting even weather, right? You can prep and you could see how that could be economically beneficial. So what kind of work are you doing with interventions and how does that kind of interact with economics? Like it sounds like the contrails one—I did it in an analysis and said actually this is great because it's very low economic impact but high value.

John

没错。所以如果你尝试——我认为有一种能源干预。所以你必须与现有的能源形式竞争。这并非易事,除非有共同利益,或者有一些巧妙的共同利益。就像,再次强调,这非常推测性。这不是我们的工作。有一家初创公司——我不知道你是否看到了新闻。我想是去年,有人发现如果你把汞注入聚变反应堆,中子通量实际上可以把汞嬗变成金,然后你可以卖掉黄金。我觉得这非常聪明。它可能行不通,但……

That's right. So if you try—I think there's sort of energy intervention. So you have to compete with existing forms of energy. And that's not trivial unless there's a co-benefit or there's some sort of clever co-benefit. Like, again, this is highly speculative. It wasn't our work. There was a startup that was—I don't know if you saw the news. It was last year, I think, where someone figured out if you inject mercury into a fusion reactor, the neutron flux can actually transmute the mercury into gold, and then you can sell the gold. Which I thought was very clever. It might not work, but...

Host

你知道,作为一名物理学家,我想从聚变反应堆中得到的一样东西是氦,但那是另一回事了。抱歉。

You know, as a physicist, the one thing I want out of a fusion reactor is helium, but that's a different story. Sorry.

John

哦,你——哦,氦-3。嗯,我的意思是,氦-4 有点无聊,尽管它正在变得——因为战略储备已经关闭。它更少了。是的。

Oh, you—oh, helium-3. Well, I mean, helium-4 is kind of boring, although it is getting—because the strategic reserve has been shut down. There's less of it. Yes.

Host

当然我想要氦-3。嗯,甚至不只是为了聚变,只是为了制造稀释制冷机用于……

And of course I want helium-3. Well, not even just a fuse, just to make dilution refrigerators for...

John

Quantis 之类的。是的。如果我们用完了氦,我们想到的那么多技术实际上就都泡汤了。

Quantis or so. Yeah. So much technology we think about actually just goes out the window if we run out of helium.

Host

确实如此。

That's true.

John

没人在想——抱歉,这完全是题外话。

No one's thinking about—sorry, that's like a complete aside.

Host

呃,是的。美国有一个氦战略储备,是为了我们非常重要的飞艇舰队。

Uh yes. The fact that the US had a helium strategic helium reserve was for our very important blimp fleet.

John

是的。

Yeah.

Host

但不管怎样,他们保留了几十年。所以那很好。但后来我们停止了。我们把它全处理掉了。它全都升到空中了。

But they kept it for decades anyway. So that was nice. But then we stopped. We got rid of it all. It all went up in the air.

John

是的。在气球之类的东西里。

Yes. In balloons and stuff.

Host

或者从天然气井里出来。

Or out of natural gas wells.

John

是的。

Yeah.

Host

嗯,抱歉,我们现在在谈氦。

Um sorry, now we're talking about helium.

John

是的。不。那么,那么干预措施。有哪些最令人兴奋、最有趣的?

Yeah. No. So, so interventions. What are some of the most exciting interesting ones?

Host

嗯,我对聚变非常兴奋。我的意思是,我不知道那算不算干预。那是一种能源,因为如果我们能让它工作,并且能让它的资本成本足够低,那实际上会大有帮助,因为至少目前的模式是——可再生能源很棒。理想情况下,你想让一切都电气化,对吧?这有问题,因为比如你不能让航班电气化,但你可以尝试让很多东西电气化,你知道,电动汽车。你必须弄清楚如何让水泥或钢铁之类的东西电气化。那些很难,尤其是像炼钢——无论如何都需要还原能力,本质上——你要加碳,还要还原铁矿石。所以让一切都电气化有很多困难。但如果你能让一切都电气化,那么所需的电量将增长五倍,你可以尝试发展可再生能源加电池。而试图把最后一点都挤出来,成本会越来越高,因为你只需要越来越多——你需要大量电池来覆盖最后几个百分点,甚至 10% 或 20%。所以我们确实需要某种能覆盖最后 20% 的电力,某种基载电力。所以聚变可能适合这个。所以那非常令人兴奋。再次强调,没有一种银弹能覆盖所有情况。所以我很乐意谈论任何具体情况,但这有点像……

Well, I'm very excited by fusion. I mean, I don't know if that's intervention. That's sort of a source of energy, because if we can make it work and we can make it be sort of low enough capital cost, that will actually help a lot, because at least the current models are—renewables are great. Ideally you'd like to electrify everything, right? Which has problems because like you can't electrify flights, but you can try to electrify a lot of stuff, you know, EVs. You'd have to figure out how to electrify things like cement or steel. Those are hard, especially things like making steel—want reduction power anyway to essentially—you're adding carbon and you're reducing iron ore. So there's a lot of things that are difficult about electrifying everything. But if you could electrify everything, then the amount of electricity required would grow by a factor of five, and you could try to grow renewables plus batteries. And trying to squeeze all of it out, it starts getting ever more expensive because you just need ever more—you need like a huge number of batteries to cover the last few percent or even 10 or 20%. So we do need some sort of power that can cover the last 20%, something that's base load. So fusion might be a thing for that. So that's super exciting. Again, there's no one silver bullet that can cover all the cases. So I'm happy to talk about any specific case, but it's sort of like...

Host

世界是一个非常复杂的地方,全球经济是一个非常复杂的地方。所以一般来说,谈论干预措施非常困难。

The world is a very complicated place and the global economy is a very complicated place. So it's super hard to talk about interventions in general.

John

是的。

Yeah.

决策与洛杉矶大火 Decision-Making and LA Fires

Host

也许不是干预措施,我好奇的一件事是这如何影响决策,例如,我们建造什么?我们如何建造?我想你来自洛杉矶,对吧?或者至少你……

Maybe instead of interventions, one thing I'm curious about is how does this make affect decisions and to, for example, like what do we build? How do we build? I think you're from LA, right? Or at least you...

John

嗯,我待了 11 年。

Well, I spent 11 years.

Host

好的。所以是的,你人生中很多时间都在洛杉矶度过。我的意思是,洛杉矶基本上大部分地区都烧毁了。可能离你以前住的地方很近。所以,你知道,这是我认为很多人有点预见到的事情。也许部分是由于监管问题,但部分是由于其他问题,你知道,我们完全没有准备,似乎缺乏对下一步该做什么或如何调整的准备。

Okay. So yeah, you spent a lot of your life in LA. I mean, LA just basically large parts of it just burned down. And maybe probably close to where you used to live. So the, you know, this is something that I think a lot of people kind of saw coming. Maybe partially due to regulatory issues, but partially due to other issues, you know, and we were completely unprepared and it seems like there is a lack of preparation about what to do next or to sort of adjust for this.

野火检测与气候韧性 Wildfire Detection and Climate Resilience

Host

你是否研究过预测新的风险评估,或者建议我们实际改变什么来让社会为即将到来的事情做好准备,无论我们是否真的采取行动解决根本问题?

Have you worked on basically predicting new risk assessments or suggestions of what we actually change to maybe harden society for what's coming, regardless of whether or not we actually do something to solve the underlying problem?

John

没错。事实上,谷歌在气候危机韧性方面投入了大量努力。我们有一个非常有趣的项目叫 Fire Fires。不知道你是否了解。我们发现,对于野火,很多野火如果你及早发现,扑灭一个这个房间大小的野火非常容易。但即使是一英亩大小,也会变得非常非常困难。当然,在某些情况下,它们可以从这个房间大小指数级增长到一英亩。所以可能很难捕捉,但它们通常从小规模开始并持续一段时间。我们发现,如果你有一个低地球轨道卫星的全球星座,可以在中波红外探测,这本质上回到黑体,即火的温度。它们在中波红外中很突出。所以我们设计了一个传感器,如果你建造,这取决于它们的轨道,但大约 50 到 80 个,你实际上可以找到这个房间大小的火灾,大约 5 米见方,可能比这个房间稍大,5 米见方,在地球上任何地方,再次取决于你有多少卫星,在 15 到 20 分钟内。你必须发射相当多的卫星,比如 80 个,才能在 15 分钟内发现,然后你实际上可以干预。如果你想让火燃烧燃料并且你认为安全,你可以决定不干预,但如果它会爆炸成不安全的东西。所以你们与一个名为 Earth Fire Alliance 的非营利组织合作,我们是其中一部分,所以他们开始,我们已经发射了一颗原型卫星。我们与一家名为 Muon Space 的公司合作实际制造卫星。所以这很酷。我们有野火边界检测,并通过谷歌传播信息。所以我们可以从现有卫星和现有数据源中找出火灾的边界,然后通过安卓手机或搜索告诉人们关于火灾的信息。我们与美国林务局合作,制作新的火灾传播模型,因为这又回到了 70 年代由一位名叫 Rothermel 的人提出的基于过程的模型。所以我们实际上制作了一个小的神经网络代理模型,基于一个新的本质,以便能够非常非常快速地运行它。所以我们与林务局合作。所以是的。是的,我们非常非常感兴趣于尽量减少,因为人们可能没有意识到,世界卫生组织估计全球每年有 30 万例因野火烟雾导致的超额死亡。

That's right. In fact, there's a big effort at Google into something called climate crisis resilience. And so we had a very fun project called Fire Fires. I don't know if you know about this. So it turns out that for wildfires, a lot of these wildfires you could see if you only caught them early enough, it's very easy to put out a wildfire the size of this room. But even if it's like an acre, it gets much much harder. And of course under certain circumstances they can grow exponentially from the size of this room up to an acre. So that might be hard to catch, but they often start small and spend a while. So we figured out that if you had a global constellation of low earth orbit satellites that could detect in the midwave IR, which goes back to the black body essentially, that's the temperature of fire. They stand out in the midwave IR. And so we designed a sensor that if you built, it depends on exactly what their orbits, but roughly 50 to 80 of them, you could actually find fires about the size of this room, about 5 m on a side, maybe it's a bit larger than this room, 5 m on a side, anywhere on the planet and again depending on how many satellites you had within like 15 to 20 minutes. You'd have to put a fair number up like 80 to get them within 15 minutes and then you could actually intervene. You could decide not to if you wanted to have the fire burn fuel and you thought it was safe, but if it was going to blow up to something unsafe. So you worked with a now a nonprofit called Earth Fire Alliance that we're part of and so they're starting to we've launched one satellite which is a prototype. We've worked with a company named Muon Space to actually make the satellite. So that's cool. We have wildfire boundary detection and we propagate that information out through Google. So we can actually figure out from existing satellites and existing data feeds where the boundaries of fires are and then we sort of tell people through their Android phones or through search about fires. We've worked with the US Forest Service on making new models for how fires propagate because again that goes back to these process-based models from the 70s by a person named Rothermel. So we've actually made a little neural network proxy model based on a new essentially to be able to run it very very quickly. So we worked with the forest service on that. So yeah. Yeah, we're very very interested in trying to minimize because it turns out people might not realize the World Health Organization estimates that there are 300,000 excess deaths a year across the world from wildfire smoke.

Host

是的。我记得,是不是,我们已经几年没有经历过真正糟糕的火灾季节了。也许四五年前,有一片烟雾云穿过美国北部和加拿大,我认为造成了很多呼吸问题。

Yeah. I mean I remember, was it, it's been a few years since we had a really bad fire season. Maybe what four or five years ago there was this cloud of smoke which crossed all of northern US and Canada and caused a lot of respiratory issues I think.

John

是的。

Yeah.

Host

是的。而且很难追踪。我的意思是,你必须通过这些,你必须从统计手段估计这些超额死亡。

Yeah. And it's very hard to track. I mean, you have to get these, you have to estimate these excess deaths from statistical means.

John

但是的,这是一个非常严重的公共卫生问题,也非常可怕,烧毁人们的房屋,很糟糕。

But yeah, it's a very serious public health problem and also just very scary and burns people's houses down and it's terrible.

Host

我的观点是,气候变化有点像一种严重的疾病。你是治疗症状?即适应,还是试图攻击根本问题?答案是,如果足够严重,两者都要,对吧?

My view is that climate change is sort of like a serious disease. Do you treat the symptoms? I.e. do you adapt or do you try to attack the underlying thing? And the answer is well if it's serious enough both, right?

John

所以是的,我们在谷歌非常重视适应,尤其是气候韧性,我们试图给人们提供信息工具来帮助。这就是为什么我们在研究天气和旋风预测等。所以这一切实际上是一致的。所以不仅仅是,你说得对。不仅仅是干预。也是气候韧性。

And so yes, so we take adaptation especially around climate resilience very seriously at Google and we try to give people informational tools to help. That's part of the reason why we're working on weather and then cyclone prediction and things. So it all actually hangs together. So it's more than just you're right. It's more than just interventions. It's climate resilience too.

Host

是的。在西海岸经历了四五个火灾季节,它们可能非常非常糟糕。以前不是这样,我的意思是,我在内华达山脉有一个小木屋,是的,以前是,哦,你知道夏天很好,然后现在就像,不是每年,但是的,就像你知道冬天、春天、夏天和烟雾。

Yeah. Having lived through four or five fire seasons on the West Coast, they can be quite quite nasty. And it used to not be, I mean I have a cabin up in the Sierra Nevada mountains and yeah it used to be oh you know summertime it's nice and then now it's like well not every year but yeah there's like you know winter, spring, summer and smoke.

John

是的。火灾季节你只是

Yeah. Fire season you're just

Host

是的。我想待在火线的西边,然后

Yeah. I want to stay to the west of the fire line and

John

是的。是的。所以这是我感兴趣的另一个事情,谷歌也非常感兴趣,就是气候韧性。

Yes. And Yes. So that is another thing that I'm interested in and that Google's also very interested in is climate resilience.

Host

这些红外传感器是否足够小,可以搭便车在微型卫星网格中?比如,是否有意义,几乎更便宜只是雇佣,我的意思是付钱给某人来监视一个星座?

Are these IR sensors small enough that they could hitch a ride in like a micro satellite grid? Like would it make sense to, would it be almost cheaper just to hire, I mean to pay someone who's watching a constellation?

John

哦,它们没那么小。问题是你需要制冷,因为

Oh, they're not that small. The thing is you need refrigeration because

Host

哦,好的。好的。

Oh, okay. Okay.

John

因为它是中波红外。所以,你必须保持冷却。

because it's midwave IR. So, you have to keep it cool.

Host

所以,这些必须是它们自己的特殊

So, these would have to be their own special

John

卫星。它们不是超级大。

satellites. They're not super large.

Host

好的。它们不像,你知道,地球同步轨道上的卫星是这些巨大的怪物,因为所有的光学和谁知道什么。但嗯是的,

Okay. They're not like the, you know, the satellites in geosynchronous orbit are these giant monsters because of all the optics and who knows what. But um yeah,

John

而且它们有,你知道,它们基本上只是眼睛传感器,分辨率为 5x5。呃,不,不,另一个可爱的事情是分辨率大约是 50 乘 50 米,但你可以使用超分辨率,因为它本质上是多光谱的,你有点知道火在哪里,所以是的,上面撒了不少 AI 来达到 5x5 米。

and they have, you know, they're basically just eye sensors with a resolution of 5x5. Uh, no, no, that's the other cute thing is the resolution is about 50 by 50 m, but you can use super resolution because it's essentially multispectral and you sort of know where fires are and you have so yeah, there's a fair sprinkling of AI on top of them to reach that 5x5 meter.

Host

是的。是的。

Yeah. Yeah.

John

当你也有,你对你读出的东西进行卷积,对吧?也

When you also have you have a convolution over the what you're reading out, right? As well

Host

我忘了帧率。嗯,卫星在移动。我不记得点扩散函数是什么。对不起。我不,但你说得对。它们确实移动,但我不记得它们有多快,我不记得它们有多快,呃,这也有一种叫做扫帚传感器的东西。所以,有一个有趣的事情,试图你以一种方式展开光谱,也有,这是一个有点复杂的东西。它不仅仅像宝丽来。它是一个复杂的传感器。

the I forget the frame rate. Um, the satellite is moving. I don't remember what the point spread function is. I'm sorry. I don't uh but you're right. They do move, but I don't remember how fast they I don't remember how fast they uh this also has something called a broom sensor. So, there's this funny thing of trying to you kind of spread out the spectrum one way and there's also it's it's a somewhat complicated thing. It isn't just a like a Polaroid. It's it's a complicated sensor.

AI在科学中的演进 Evolution of AI in Science

Host

稍微换个话题。你知道,你在 AI 和科学的交叉领域已经有一段时间了。我觉得你有点来回穿梭。你怎么看这个领域的发展?因为我觉得现在发展得非常快。

Sort of switching gears a little bit. You know, you've been at the intersection of AI and science for quite some time. I think you've sort of wound your way into and out of it back and forth. How do you see the field has evolved? Because I feel like it's evolving very quickly now.

给年轻科学家的建议 Lessons for Young Scientists

Host

那么,你学到的、你认为社区已经学到的一些教训是什么?如果你是一位年轻科学家或年轻从业者,你认为这应该如何改变?这应该如何改变你应对未来的方式?

And like what are the sort of lessons that you've learned that you think the community has learned and how do you think this should change if you are a young scientist or young practitioner? How should this change how you should approach the future?

John

嗯,我认为在过去 12 到 18 个月里发生了一次相变。所以,我的意思是,正如我所说,我们过去做的很多事是构建这些专用模型来解决单个问题。如果你认为那是你的工作,那还挺有趣的。你找到一个 problem,解决它。再找到一个 problem,解决它。

Well, I think there has been a phase change in the last 12 to 18 months. So, I mean a lot of what we used to do as I said was build these specialized models to solve individual problems. And if you think that's your job, it's kind of fun. You find a problem, you solve it. You find another problem, you solve it.

John

但现在我们有了这些更通用的 AI 事物。我认为整个 AI for science 社区还在摸索。事实上,他们所做的工作如此之新,以至于我们集体都不确定什么是最好的做法,或者也许没有唯一最好的,也许有一个工具链,我认为我们都在试图弄清楚我们应该做什么。

But now we have these much more general AI things. And I think the whole AI for science community is kind of still feeling around. There's such the fact that they're working is so new that collectively we're not sure like what's the best thing to do or maybe there's no one best maybe there's a tool chain and I think I think we're all trying to figure out like what should we do

Host

所以有一个问题:年轻科学家应该做什么?

and so there's a question of uh what should uh young scientists do?

John

我认为会是,你知道,我有一个儿子,他刚满 21 岁,他非常喜欢 AI、编程和化学,我看着他,我认为他做得对,因为他既在学习很多,又试图成为 RNA 领域的专家,但他也在使用 vibe coding 和使用所有工具。我认为这是正确的答案,因为每个人都在摸索,仍然要成为一个深度领域。我不认为领域专业知识会消失,因为这又回到了很多人说的,这又回到了品味,以及试图弄清楚人们如何在不做所有苦工的情况下获得品味。这是一个有趣的开放问题,但要发展领域专业知识,同时也要尝试使用所有可用的不同工具,因为并不是说哦是的我们知道会发生什么,聪明的老人知道会怎样,不,我们也在实验,所以嗯是的,所以我会说一定要发展领域专业知识,尝试使用这些工具,并尽你所能解决大而难的科学问题,仍然有一个巨大的开放问题,关于你实际上如何处理物理实验室工作,这不会消失,因为你知道实验是 ground truth,而且

I think it would be I you know I I I have a son who's uh who just turned 21 and he's really into both sort of AI and coding and chemistry and I look at him I think he's doing the right thing because he's both learning a lot and trying to be a domain expert about RNA but he's also sort of using vibe coding and using all the tools. I think that's I think that's the right answer is is is because everyone's figuring it out still be a deep domain. I don't I don't think domain expertise is going away because it goes back to a lot of people who said it goes back to taste and trying to figure out how people get taste without doing all the the grunt work. That's an interesting open question but develop domain expertise but also try and play I would say with all the different tools that are available because it's not like oh yes we know what's going to happen and the smart old people are knowing what's like no we're experimenting too and so uh yeah so I would say uh definitely develop domain expertise and try to use these tools and try to solve uh big hard scientific problems as best you can there's still the huge open issue about what do you actually do about physical lab work that is not going away because you know experiments are the the ground truth and

Host

而且是一个瓶颈

and a bottleneck

John

而且是一个瓶颈。我的意思是人们在谈论 lab in the loop,但这仍然非常非常非常开放,因为你怎么,据我所知,没有人有一个能做所有事情的通用实验室。有很多非常具体的、你知道的、可控的实验室。

and a bottleneck. I mean people are talking about a lab in the loop but that's still very very very open because how do you how you no one has I think as far as I know a general lab that does everything. There's a lot of very specific you know labs that are controllable.

John

嗯,所以我认为只是我们经历了这次相变。嗯,这看起来超级令人兴奋。再次,我会建议人们玩任何可用的工具,并发展某种深度的领域专业知识和品味,尽你所能。我也会建议人们不要害怕,尝试东西,你知道,嗯,你知道我总是很高兴,我们在 Google 有学生研究员,他们来了,做各种疯狂的事情,那总是令人愉快。所以是的,人们应该尝试各种疯狂的事情,看看会发生什么。

Um so I think it's just we've gone through this phase change. uh it seems super exciting. Again, I would advise people to play with whatever tools are available and to develop sort of deep domain expertise and and taste to the extent you can. And I would advise people also not not to be scared and like try stuff you know uh you know I'm always happy uh we have student researchers at Google and they come and they do sort of wild and crazy things and that's always just delightful. So yeah, people should be trying sort of wild and crazy things and see what happens.

AI时代培养专长 Developing Expertise in the AI Era

Host

是的,这可能是一个没有答案的问题,但当我想到我如何发展专业知识以及很多人如何发展专业知识时,是从一个简单的定义问题开始,然后反复敲打它,然后在那个探索过程中,你学到更多,你知道有些方式你变得更广,有些方式你变得更深,但你仍然只是用头撞问题,现在这个问题可能瞬间可解,这个过程教会你解决更难问题所需的技能。你会给你儿子什么建议?

Yeah, this may be a question without an answer, but when I think about how I developed expertise and how a lot of people developed expertise, it was by starting with a, you know, a a simple defined problem and then hammering it and then in that process of exploration, you learn more and you know some ways you go broader, some ways you go deeper, but you still the process of just banging your head against the problem which now would be instantly solvable teaches you the skills you need to solve harder problems. What advice would you give to your son for that?

John

你知道,我不知道。也许这有点像徒步,那就是你,是的。我的意思是,你显然不能开车到所有地方。或者你可以开车上山。

You know, I don't know. Maybe it's a bit like hiking, which is you Yes. I mean, you obviously can't drive everywhere. Or you could you could drive up the mountain.

Host

是的。

Yeah.

John

或者你可以徒步上山。也许偶尔徒步上山是可以的,甚至有趣,即使你可以开车上山。

Or you could hike up the mountain. And maybe it's okay, even fun to occasionally hike up the mountain even if you can drive up the mountain.

Host

嗯是的。我的意思是,

Uh yeah. I mean,

John

在过去,我会听起来像一个真正的老人。在过去,

in the old days, I'm going to sound like a real old man. In the old days,

Host

在六个月前的过去,

in the old days of six months ago,

John

嗯,不,不,我甚至在想 80 年代和 90 年代的日子,你知道很多人认为哦有开源包,有学习,有各种我们没有的东西,我不得不写自己的数值库,我不得不写自己的机器学习,我写过 boosting,可能重写了四次,用四种不同的语言,所以现在我知道 boosting,你知道就像这样,所以也许不采取完全简单的方式,我的意思是显然有一个权衡,比如哦但我想尽可能高效和富有成效。是的。但你也必须发展肌肉。所以这也许有点像当运动员。就像有些地方有些时候你实际上在做 exploit,当你试图尽可能快地跑。然后还有训练时间,所以也许人们只需要训练。

well, no, no, I was even thinking of in the old days of the 80s and 90s like you know a lot of people take it like oh there's open source packages there's there's learn there's there's all sorts of things we didn't have that I had to write my own numerical library I had to write my own machine learning I've written boosting probably rewritten it four times four different languages and and so now I know boosting you know it's like and so maybe uh not taking the totally easy I mean obviously I mean there's this trade-off like oh but I want to be as efficient and and productive as possible. Yes. But you also have to develop the muscles. So it's a little bit maybe like being an athlete. Like there are places there are times when you're actually doing exploit when you're trying to run as fast as you can. And then there's also training time and so maybe people just have to train.

Host

而且完全有可能,如果你花时间真正反复敲打并做艰苦的工作,即使在那里进展较慢,这也会为你的更大的,嗯,你知道,长期生产力带来红利。就像即使局部那一刻你没有通过利用 LLM 来最大化生产力,这也会反馈到某些东西中。

And it's entirely plausible that if you spend time actually hammering away and doing the hard work, even if it goes slower there, that pays dividends into your larger um you know productivity long term. Like even if locally that one moment you are not being maximally productive by not exploiting an LLM that feeds into something.

John

我希望如此。我希望,我不知道的是,我希望人们在他们的职业生涯中,这很难,因为对,整个世界似乎都想优化一切,而你的错。不,我没有。这是不,我没有。这是我的错,但只是那种,你知道,嗯,有时你必须留出时间。就像在 Google,尤其是在我的团队,我们有 20% 时间的概念,我在自己的团队中仍然非常非常强烈地试图保护它。就像你可以做任何事,如果你想学习东西,如果你想尝试东西,你甚至不必告诉我。你甚至不必告诉我。事实上,我可能不应该告诉。嗯,你知道,只是做事情为了,为了,确切地说。为了学习,也因为那是创造力的源泉。我不想占用人们的时间,让他们没有,就像他们不能感觉可以玩或学习或尝试新的疯狂的事情。所以我知道 20% 时间是不寻常的,世界上似乎有这种强烈的动力,就像我说的优化和挤出一切,但当你过度优化时,你确实会失去一些东西,你有点过拟合。

I hope so. I hope the thing I don't know is I hope people in their careers it's hard because right the the the whole world seems to want to optimize everything and your fault. No, I don't. It's No, I don't. It's my fault, but it's just sort of the the the the you know, uh and sometimes you have to set aside time. Like at Google, especially in my group, we have this concept of 20% time, which I still I very very strongly in my own group try to protect. It's like you can do whatever you if you want to learn stuff, if you want to try stuff, you don't even have to tell me. You don't even have to tell me. In fact, I probably shouldn't tell. Uh you know, just do stuff for for Exactly. for learning and also because that's where the sort of creative juices are. I don't want to so occupy people's time where they have nothing like where they they can't feel like they can play or learn or try new crazy things. So I know 20% time is unusual and there just seems to be this strong impetus in the world to just like like I said optimize and squeeze everything out but you do lose something when you hyperop you sort of overfit.

Host

你过拟合到生产力,因为你如此

You overit overfit to productivity as you're so

John

没错。所以我知道我的建议可能是逆流而上,对抗,嗯,也许文化规范。

that's right. So I I know that might be my advice might be swimming upstream against uh perhaps cultural norms.

Host

这里总是出现在我脑海中的是,我,有一个,问题不是静止的。

The thing that always comes up for me here is that I there's a the problem is not stationary.

AI时代的持久技能 Enduring Skills in the Age of AI

Host

有一套新的技能组合将成为未来的正确技能组合,对吧?我脑子里总是在纠结这个问题:这项技能是不是一项持久的技能?或者也许昨天它还不持久,但今天它会变得持久,因为——就像我看到的,你知道,举个明显的例子,当你使用 LL 技能迁移时,你有点变成了一个管理者。

There's a new skill set that will be the right skill set for the future, right? And the question in my mind is always just tangling: is this a skill that is an enduring skill? Or maybe it wasn't enduring yesterday, but today it will be enduring because—like, I've seen that, you know, for just an obvious one, you become sort of like a manager when you're using LL skills transfer.

John

嗯,有些技能不是这样,但很多技能确实如此。所以作为管理者,你会失去对正在发生的事情的细节的跟踪,你信任你的下属或智能体或其他什么来管理这些,以便他们能向你汇报并回答,你知道,那些高层次的问题,并正确判断小事情。那么,这就是我们现在的处境吗,我们就像中层管理者一样?不,我的意思是,我不知道。再说一次,我有一点管理工作,但你不想成为一个空壳。

Well, some of them don't, but a lot of them do. And so as a manager, you lose track of the details of what's going on and you trust your people or agents or whatever to have that managed so that they can report up to you and answer, you know, sort of the high-level questions and get the judgment about the little things correctly. And so, is that what we've come to, that we're just like middle managers now? No, I mean, I don't know. Again, I have a little management work, but you don't want to be an empty suit.

Host

换句话说,因为这些东西可能会出错,尤其是那些非常奇怪的大语言模型。它们不会犯和人类一样的错误。

In other words, because the things might get it wrong, especially LLMs that are sort of really weird. They don't make the same kind of mistakes that humans make.

John

所以你不能完全信任它们。你必须严谨,就像,你知道,戳一戳它,确保没问题。不过,你也应该戳一戳你自己写的软件。我的意思是,你不应该信任自己。这是我学到的一件事。费曼怎么说来着?你知道,你绝对不能欺骗自己,而你是最容易被欺骗的人。

And so you can't, you know, fully trust them. You have to be rigorous and like, you know, poke at it and make sure. Although, you should be poking at software that you write yourself, too. I mean, you shouldn't trust yourself. That's one thing I've learned. What did Feynman say? You know, you absolutely can't fool yourself and you're the easiest person to fool.

Host

所以,那可能是——我不知道我能否量化什么是持久的,但不知怎的,是根本性。我的意思是,真正学习一个关于世界的领域,例如。所以这就是为什么我喜欢生物学或物理科学。我认为那些是持久的、根本的东西。数学非常持久,但即使是严谨和检查之类的事情,这可能又回到了管理,你想真正确保大语言模型产生正确的东西,或者它们没有以某种方式作弊,但再次,你也应该对自己这样做。

So, that might be—I don't know if I can quantify what's enduring, but somehow fundamentalness. I mean really learning a domain that is about the world, for example. So this is why I like well biology or physical sciences. I think those are enduring sort of fundamental things. Math is very enduring, but even things like rigor and checking and that sort of thing, which goes back to maybe management, that you want to really make sure that the LLMs are producing the right things or they haven't cheated in some way, but again, you should be doing that to yourself, too.

John

是的。

Yeah.

Host

所以,我认为有一些持久的东西,还有创造力和跳出框框思考的持久价值。所以,我认为那些是持久的。我不知道。所有这些都有一些非常根本的东西。

So, I think there's some enduring and also just the enduring value of sort of creativity and thinking out of the box. And so, I think those are enduring. I don't know. There's something very fundamental about all of those.

费曼的计算课程 Feynman's Class on Computation

Host

所以,你提到了费曼。嗯,如果你不介意我换个话题。

So, you mentioned Feynman. Um, if you don't mind me changing gears.

John

我上过费曼的一门课。是的。不是随便什么课。

I took a class from Feynman. Yeah. Not just any class.

Host

是的。他那种计算物理学课程。他实际上和霍普菲尔德以及卡弗·米德一起教的。嗯,那很有趣。当时我认为也许费曼知道。我觉得我们甚至不知道问题是什么。我的意思是,我猜费曼试图说哦让我们做量子模拟,我猜这是嗯结果证明是正确的答案,但是的,当时它嗯我猜它很酷,首先它是由 DARPA 资助的,所以你应该每周去一次获得上等肋排晚餐,嗯但我只是每周都去,嗯无论如何为了学生。

Yes. The his sort of physics of computation class. He did it with in fact Hopfield and and and uh uh Carver Mead. Uh that was fun. At the time I don't think any maybe Feynman knew. I felt like none of us knew what even the problem was. I mean I guess Feynman was trying to say oh let's do quantum simulation which I guess is uh it turned out to be the right answer but yes at the time it was well I guess uh it was cool first of all it was DARPA funded it so you were supposed to go once a week to get a prime rib dinner uh but I just went every week uh anyway for the students.

Host

是的。所以只是为了多一点背景,这门课基本上是在费曼和我不记得还有谁提出了量子计算机的概念之后,没有真正知道它是什么,但知道有某种——

Yeah. So just for a little bit more context, this class was basically the class right after Feynman and I forget who else proposed the concept of a quantum computer without really knowing what it was, but knowing that there was some sort of—

John

嗯,有《底部还有很大空间》那篇文章,我认为那是在很早的时候——我是在 82 年上的。

Well, there was plenty of room at the bottom essay which I think was in the very early—this I took it in '82.

Host

哦,那是他教的时候。

Oh, that was when he taught it.

John

他确实说了底部还有很大空间。但课程的结构是,嗯,周二有客座讲座,然后费曼会在周四站起来解释为什么那全是错的。嗯,这很有趣。然后我们遇到了其他人也遇到过的东西,叫做费曼效应。也许他太有魅力了什么的。他会解释事情,然后说:“是的,是的,我明白了。”然后你走出去想:“不,不,我没有。”

And he did say there was plenty of room at the bottom. But the way the class was structured, it was um uh it was like a guest lecture on Tuesday and then Feynman would stand up on Thursday and explain why that was all wrong. Uh which was pretty fun. And then we encountered something which other people have also encountered something which called the Feynman effect. Maybe he was so charismatic or something. He would explain things and he would say, "Yes, yes, I understand." And then you walk out to think, "No, no, I did not."

Host

是的。所以,这有点有趣。但很多客座的人,我的意思是,也许展示了混乱。我的意思是,很多人,我想是丹尼·希利斯来了,嗯有所有这些有趣的客人。所以,我会说在他们所有人的联合中,我认为展示了当时发生的大规模混乱,因为有很多像,哦,我们应该让计算机可逆吗,因为我们必须确保,你知道,它们甚至可以是可逆的吗?每次操作的热量下限可以为零吗,或者有某种热力学极限,那曾经是个大问题。我不认为现在是个大问题。

Yeah. So, it was kind of fun. But a lot of the guest people, I mean, maybe the sort of showed the chaos. I mean, a lot of people, I think it was um Danny Hillis came, uh there was all these interesting guests. So, I would say in the union of all of them, I think showed the sort of mass confusion of what was going on because there was a lot of like, oh, should we make computers reversible because we have to make sure, you know, can they even be reversible? Can the sort of the bottom limit of the heat per operation be zero or is there some sort of thermodynamic limit that was like a big deal. I don't think that's a big deal now.

Host

是的。但等等。所以你在谈论兰道尔极限,对吧?或者不是它们上面的东西。

Yeah. But wait. So you're talking about the Landauer limit, right? Or not what's on them.

John

嗯,当时所有这些问题都是关于你能有可逆的像台球计算机吗,它们可以是可逆的吗等等,还有类似什么样的计算,这就是为什么丹尼·希利斯来了,你知道,我认为那是在连接机的时代,以及最初的思维机器,不是疯狂,而是 80 年代最初的那个。

Uh well, there at the time it was all this question about like can you have reversible like billiard ball computers and can they be reversible and things so and also sort of like what kind of computing that's why Danny Hillis came like you know I think that was in the era of the connection machine and what the original thinking machines not mad but the original one in the 80s.

Host

嗯,所以他来了。

Uh so he came.

John

1982 年通用计算的状态是什么样的,我的意思是,我的意思是,在这一点上——

What was the state of like general computation in 1982 like I mean I mean at this point—

Host

四舍五入误差,我们为零。我的意思是,我们——我记得当我,抱歉,我到了那里,嗯我是卡弗·米德的嗯 CIS 管理员,我们有一台 VAX 11/750,可能做 1 MIP,100 万次操作,整个研究组共享一个 80 兆字节的磁盘驱动器,有洗碗机那么大。那非常令人兴奋。嗯,那么复杂性理论呢?它是什么?因为我知道现在对量子复杂性以及它与引力的关系有很多兴趣。我想知道,我脑子里没有关于复杂性何时——

To rounding error we had zero. I mean we had—I remember when I sorry I got there uh I was Carver Mead's um uh CIS admin and we had a VAX 11/750 that maybe did a MIP, 1 million operations and the whole uh research group shared an 80 megabyte disc drive that was the size of a dishwasher. It was it was very exciting. Uh so what about complexity theory? What what was it? Because I know that there's a lot of interest in quantum complexity and how it relates to gravity right now. And I wondered I don't have a mind in my mind about when complexity—

John

我们可以试着查一下。我不认为有,当然整个有一个完整的层次结构,你说得对,量子复杂性类。

We could try to look it up. I don't think there's of course the whole there's a whole hierarchy of you're right quantum complexity classes.

Host

我认为那是在那之后发展起来的,因为这是在 1982 年。事实上,我甚至不确定有没有门。我们从未谈论过门模型。没有量子门模型——

I think that was developed after that because this was in 1982. In fact I I'm not even sure there was a gate. We never talked about the gate model. There was no gate model of quantum—

John

我是说量子门。是的。是的。是的。

I mean quantum gate. Yeah. Yeah. Yeah.

Host

是的。是的。我忘了。抱歉。我认为很多那些——

Yeah. Yeah. I forget. I'm sorry. I think a lot of those—

John

是的。我认为那些直到 90 年代左右才出现。是的。

Yeah. I don't think those came out until like the what 90s or something. Yeah.

量子信息教科书 Quantum Information Textbook

John

我记得读过 Nielsen 和 Chuang 的经典量子信息教科书,现在回想起来,那本书提到了很多这些点。

I remember reading Nielsen and Chuang, the classic quantum information textbook, which brings up a lot of those points now that I think about it.

Host

是的,量子信息教科书。

Yeah, quantum information textbook.

John

那本教科书是在 90 年代末或 2000 年代初写的。

That textbook was written in the late '90s or early 2000s.

Host

没错。

That's right.

John

我认为那是第一本奠定该领域通用知识的教科书,但我可能记错。现在我觉得,年轻时我并不认为这是人人都在用的书。

And I think that was the first textbook that put down the general knowledge of the field, but I could be wrong. I think now that when I was young, I didn't take it for granted that this is the book that everyone uses.

Host

是的,但这是……嗯,你记得在 80 年代初,神经网络引起了大量兴奋,这真的很有趣,因为当时我们真的不知道自己在做什么。没人知道自己在做什么。

Yeah, but this is a... Well, you remember in the early '80s, there was a whole bunch of excitement around neural networks, and it's really interesting because again, we really did not know what we were doing. No one knew what they were doing.

John

有趣的是,先是神经网络,然后是支持向量机,再是神经网络。

There was like an interesting like neural networks then SVMs then neural networks.

Host

哦不,不,不,这之前,因为在 80 年代,每个人都说这有能力彻底改变计算。

Oh no no no, this is before that because in the '80s everyone said this has the capability of revolutionizing computing.

John

但那又怎样……嗯,我的意思是,Hopfield 网络引起了巨大的兴奋。事实上,NeurIPS 就源于 Snowbird 的一个研讨会,名义上是私人的,但每个人都想挤进去,所以他们创办了 NeurIPS。

But what does... well, I mean there was tremendous excitement around Hopfield networks. In fact, NeurIPS came out of a workshop at Snowbird that was nominally private but everyone tried to crash, and so they spun up NeurIPS.

Host

Snowbird 是一个附带计算会议的滑雪之旅。

Snowbird is a skiing trip with a computation conference attached.

John

没错。我因为不会滑雪而学会了滑雪,然后就一直去。总之,那个研讨会源于 1985 年的圣巴巴拉研讨会,而后者又源于加州理工学院的一些本地活动,叫做 Hopfests。

That's right. I learned to ski because I didn't know how to ski and then I kept going. Anyway, that workshop came out of the Santa Barbara workshop in 1985, which came out of some local things at Caltech called Hopfests.

Host

所以,所以是的,人们想,哇,有些东西涉及……事实上讽刺的是,有关于联想记忆的东西,如果你深入探究 Transformer 是什么,它们就是联想记忆。所以事实上甚至有一篇论文叫做你知道的,Hopfield 网络就是你所需要的。所以

So, so yeah people thought oh wow something involved in fact ironically there was like yeah something about associative memory and if you actually dig down into what transformers are they are associative memory. So in fact there was even a paper called you know uh hopfield networks are all you need. Uh so

John

是的,我忘了那个标题是引用那篇论文。

Yeah I forgot that title was a reference to that paper.

Host

呃,对那个时代,所以实际上就像整个事情又回到了原点,事实上我认为我们确实,呃,我们作为一个领域集体地彻底改变了计算机科学,但希望和梦想完全超出了能力,因为再次我们实际上没有算力。

Uh to the era to the era so it's actually it's like it's like the whole thing has sort of come full circle and in fact I think we did uh we have uh um collectively as a as a field uh revolutionized computer science but it was the the sort of the hopes and dreams completely outstripped uh the capabilities because again we effectively had zero compute

John

是的,我认为机器学习的历史就是不同范式,比如算力与内存,Scaling(规模扩张)与数据在不同层面变得可用。

Yeah I think there's the the history of machine learning is different paradigms as like compute versus memory uh scaling versus data become available in different levels.

Host

没错。事实上,我认为人们没有意识到神经网络胜出的原因是它们是唯一受算力限制的东西,尽管现在我们有了 Transformer,事情又变得受内存限制了。而且它们建立在 BLAS 之上,BLAS 被 GPU 等优化。我的意思是,我不确定,事实上我相当确定大脑不是通过矩阵乘法工作的,但只是算法与硬件共同进化了。

That's right. And the fact that I think people don't realize that the reason why neural networks won is they're the one compute limited thing although again with we with now that we have transformers things are getting memory limited again but but and the fact that they ride on top of blaz and so the fact that blaz was being optimized by things like GPUs I mean I don't actually know if in fact I'm pretty sure the brains don't work by matrix multiply but but it was just that that did that the algorithms co-evolved with the with the hardware

John

这就是我们走到今天的原因。谁知道呢?我的意思是,如果我们走了另一条路,人们真正关心其他计算方式,谁知道我们会最终得到什么架构。我不知道。我的意思是,很多这些不是从……我想人们黑入 PlayStation 来训练神经网络之类的,或者做……我猜也许更早,为了超级……

and And that's that's why we're here. Who knows? I mean, if we'd gone down some other path where people really cared about some other compute making who knows what what architecture we have ended up with. I don't know. I mean, didn't a lot of this start with like I think people were hacking PlayStations to trade to train neural networks or something or to do I guess maybe it was even before that for like super

Host

有一个有趣的故事,来自我研究生院的朋友 Brian Kotenzaro 在 Nvidia,他……Brian,抱歉如果我记错了,但他来到 Nvidia,很难获得支持,基本上他们加倍投入游戏,有一天他和 Jensen 开会,说服他,让我们……你知道,我们有很多人使用 CUDA 做 BLAS 和深度学习,说服了 Jensen,他说那是一次 15 分钟的对话,他们第二天就整个公司转型了之类的。好吧,我从未在……我的意思是,我有朋友 Dave Kirk,他是第一任首席科学家,还有 Bill Deli,都是我研究生院的朋友。所以是的,但我不知道具体细节。

There's a there's a funny story by um a friend of mine from grad school uh Brian Kotenzaro at video where he Brian sorry if I get this wrong but um he came to Nvidia and um and was having a lot of trouble getting traction and and and he basically they were they were doubled and tripled down on on gaming and he basically one day had a meeting with Jensen and convinced him let's you know like this is we have all these people using CUDA for blaz basically and and for deep learning in particular and convinced Jensen and they he said it was like a 15-minute conversation and they pivoted the whole company the next day or whatever. Okay, this I've never worked in I mean I have friends uh uh uh Dave Kirk uh who's the first chief scientist and and Bill Deli were both friends of mine from grad school. So yes, but I don't know the details of of way.

John

现在记住,人们,比如 Hinton 的团队,即使在 2010 年左右,他们使用 GPU 进行深度学习,但即使在 2007、2008 年,它们也不比 CPU 快多少。再次,它们真正开始超越……事实上我认为这不是巧合,在原始 ImageNet 和语音识别等时代。所以我认为这不是巧合,它们真正起飞是在 GPU 超过 CPU 的时候。所以那不是巧合。

Now remember that people um uh like in in um Hinton's group even in the around 2010 they were using GPUs to do the deep learning but but even in as even in I would say 2007208 they weren't that much faster than CPUs. Again they they they only started really exceeding in fact I think this not a coincidence uh in the era of you know the original imageet and some of the speech some of the speech recognition stuff. So that's I think that's not a coincidence that they really researched when when when GPUs passed CPUs. So there's that that's not a coincidence either.

Host

是的。我真的很想知道一个人怎么有机会命名小行星。

Yeah. I really want to know how does one get the opportunity to name an asteroid.

John

哦,好吧,又是在加州理工学院,有一些很棒的人,Jean 和 Carolyn Shoemaker,他们在教一门行星科学课。我喜欢行星科学,所以是的,作为那门课的一部分,他们带我们经历了小行星发现的过程。现在是 80 年代,所以现在有各种惊人的系统。有一段时间有个叫 LINEAR 的东西自动化了这个过程,但这是在自动化之前。所以是的,这只是课程的一部分,你要做的步骤,尽管我们在课上顺序乱了,但 40 年前我们用来做这个的步骤是:你去帕洛马,有一个快速望远镜,现在 literally 是博物馆展品,在他们的游客中心,但当时是真实的东西,你放一张胶片,拍一张照片,然后等几分钟,再拍同一片天空的相同照片,你把它放进立体镜,就像回到加州理工学院,你会看看有没有东西漂浮,因为它会移动一点点。所以它会弹出来,你可以看到它,因为它在不同的眼睛中不同。然后你能看到它投影在不同的位置。

Oh well again is Caltech there was there's were some wonderful people uh Jean and Carolyn Shoemaker and they were teaching a class in um in planetary science. I like planetary science and so yes as part of that class uh they took us they they took us through the sort of asteroid discovery thing. Now this is in the 80s so now there's all sorts of amazing amazing um systems. Well there was for a while there something called linear uh which automated the thing but this was before anything was automated. So yeah, it was just part of a class and what you would do is you go the steps even though we did it out of order in the class but the steps we used to do this is 40 years ago is you would go to Palomar there would be a fast telescope which literally now a museum piece that's in their visitor center but at the time it was a real thing you would put a piece of film in it you would take a picture and then you'd wait for a few more minutes and take a same picture of the same point of sky you'd put in a stereoscope like back at Caltech you would see if anything you would look around you would see if anything floated Because it's this it would move by a tiny bit. So it would pop up and you can see it was because it's different in different eyes. Then you would be able to see it as a something that was projected in a different place.

Host

没错。所以它 literally 会弹出来。然后你会去一个测量显微镜,测量已知恒星参考和这个漂浮物的位置。所以如果你得到足够准确的测量,然后你会把它发送给一个人,我不认为现在还是他,他的名字是 Brian Marsen,在亚利桑那州的小行星中心。然后他有一个大的软件系统来拼凑这些,这些叫做 apparitions。

That's right. So you would pop out at you literally. And then you would go to a measuring microscope and you would take measurements of known star references and where this floater is. And so if you got accurate enough measurement, then you would send it off to a person, I don't think it's him anymore, his name Brian Marsen at the Minor Planet Center in in Arizona. And then he had a big um uh software system to do um to sort of piece together these are called apparitions.

小行星发现与命名 Asteroid Discovery and Naming

John

如果你发现了,并且碰巧做了一次观测,那是最后一次出现,让他的软件能把它连成一个大轨道,那么你就能获得发现权,可以给这颗小行星命名。但现在情况很惊人。智利现在有一个天文台叫薇拉·鲁宾天文台,还有一台了不起的望远镜叫西蒙尼巡天望远镜。它基本上把这个过程自动化了。基本上就是拍摄大量天空的帧。这是一台极其惊艳的仪器。所以他们在 6 周内发现了 11,000 颗小行星。

And if you found and if you happen to have done an observation, which was the final apparition, which allowed his software to connect it into one big orbit, then you would get discovery rights and you could name the asteroid. But now it's like amazing. There's this observatory now in Chile called the Vera Rubin Observatory and there's this amazing telescope called the Simony Survey Telescope. Essentially automates this process. Essentially can just take many frames of the sky. It's an utterly stunning instrument. And so they discovered 11,000 asteroids in 6 weeks.

Host

哇哦。

Oh wow.

John

是啊。所以现在

Yeah. So it's now

Host

那到底有多少颗?我是说至少

How many of them are there? I mean at least

John

嗯,我想这取决于截止标准,但没错。不过可能有数百万颗小行星。

Well, it depends I guess on the cut off but yeah. So, but there probably millions of asteroids.

Host

所以,仍然有很多机会。

So, there's still a lot of opportunity.

John

确实。但我甚至不知道他们是否还费心去命名。

It's true. But I don't even know if they bother name.

Host

所以,我不知道。也许你不需要,我不知道。

So, I don't know. Maybe you don't need I don't know.

John

但人们在做掩星。抱歉,我可以一直聊这个。有一种东西叫掩星。比如我的一颗小行星,我当时在给一位业余爱好者发邮件。如果一颗小行星碰巧从一颗恒星前面经过,它就会变暗。就像他们发现系外行星那样。但如果你超级幸运,它会变暗,然后再次变暗,因为有一颗卫星。所以我的一颗小行星,这位业余爱好者发现它周围有一颗小卫星。所以那是我想是去年

But people are doing occult. Sorry, I can I'll talk about this forever. There's there's some there's stuff called occultation. Like one of my asteroids I I was uh talking I was sending email to a amateur. Uh I if an asteroid happens to pass in front of a star, it uh dips. You like when they discover exoplanets. But if you're super lucky, it'll dip and then it will dip again because there's a moon. And so one of my asteroids uh this amateur uh found a little moon uh around it. So So that was I think last year

Host

围绕那颗小行星。

around the asteroid.

John

是的。

Yeah.

Host

有意思。

Interesting.

John

所以那挺酷的。所以仍然有空间

So that's pretty cool. So there's still room for for

Host

你的小行星有一颗卫星。

your asteroid has a moon to it.

John

是的。事实上,显然其实这里有点故事,因为那其实不算是我的小行星,我找到了两颗。我把其中一颗以我父亲命名。我犹豫了。所以 Carolyn 把它以华盛顿大学的一位教授命名,我说:“哦,但我想命名它。”于是我犹豫了,她很好心,把她的一颗给了我。然后他们把那颗以我母亲命名,就是那颗有卫星的。

Yes. In fact, apparently it's actually actually there's a little bit of a story there because it's not really my asteroid that I found two of them. I named one of them after my dad. I waffled. So Carolyn named it after a professor at University of Washington and I said, "Oh, but I wanted to name it." And so I whed it and she was nice and so she gave me one of hers. And so then they named that one after my mom and that's the one that has the moon.

Host

哦,好的。

Oh, okay.

John

所以以我母亲命名的那颗小行星有一颗卫星。

So So the asteroid named after my mom has a moon.

Host

它有多大?

How big is it?

John

主体他们认为大约四公里宽。我尽量把单位说对。卫星大约 1 公里。所以实际上这是一个相当大的双星系统。它们不完全是双胞胎,但没错。

It's like the main body they think is about four kilometers across. I'm trying to get the units right. And the moon is about 1 km. So it's actually it's not it's a pretty big binary. They're not quite twins, but yeah.

Host

那个班上其他人

Did other people in that class

John

发现了吗?我想还有一个人。这有点碰运气。是的。而我找到了两颗,这相当不寻常。

discover? I think there was one other person. It's a little bit of a crapshoot. Yeah. And the fact that I found two was quite unusual.

Host

有什么原因吗?是纯粹运气好,还是整天待在那里?我试过吗?我知道。我尽量小心,但

Was there a reason why you was it just pure luck or just staying there all day? Do I try? I know. I I tried to be careful, but

John

是的。

Yeah.

Host

一个人怎么才能获得奥斯卡奖?

How does one become get an Oscar?

John

是的。学院奖。

Yeah. Academy Award.

Host

呃,好吧,也许我又是在正确的地方。呃,我的论文导师叫 Albar,嗯,我 1986 年实际上在一个叫 Shumber 的地方做学生实习生,那是个做石油发现的地方,但他们有一个 AI 实验室,所以我们都在做计算机图形研究,呃,我们都在想,你知道,当时这是革命性的,我意识到现在这被认为极其无聊,比如你可以实际用物理模拟器来制作计算机图形电影,当时我觉得哇那真的很酷,所以我说你知道你可以用弹性理论来制作柔软的东西,所以我说有点像这是弹性理论,我写了一个弹性模拟器,我做了你知道布料和你知道有弹性的东西之类的。所以他们说哦哇那很酷,所以那类东西的后代成了很多人们在你知道 Pixar 和他们的各种电影中使用的物理模拟器。所以是的,我告诉实习生,半开玩笑地说,如果你作为实习生做得非常好,你就能获得奖励。那么,从你做那项工作到那时隔了多久?

Uh well, again, maybe I was at the right place. Uh I uh my thesis advisor is named Albar and um I was a student intern actually in 1986 at a place called Shumber which is a boil discovery place but they had an AI lab and so they were we were all doing sort of computer graphics research and uh we were all thinking you know there at the time this was revolutionary I realize it's now considered incredibly boring like you could actually use like physics simulators to to make computer graphics movies and at the time I was like wow that's really cool so I said you know you could use the theory of elast elicity to make floppy things and and so I said I sort of like here's the theory of elasticity and I wrote it elastic simulator and I made you know fabric and you know stretchy things and and stuff. So they said oh wow that's cool and so the sort of descendants of that became a lot of the physics simulators that people used in in um you know Pixar and their various movies. So yes I tell I tell interns sort of half jokingly well if you do a really good job as an intern you can get you get awarded. So, so how like how long was that between the time you did that work and then

Host

20 年?

20 years?

John

是 20 年,

It was 20 years,

Host

这其实并不罕见,对吧?因为当他们给你颁发学院奖时,他们想确保,哦是的,它被广泛使用,每个人都在用之类的。所以,是的,但当然显然在那 20 年里,嗯,每个人都在做。这很明显,但你知道,

which is actually not atypical, right? Cuz they want to when they give you an Academy Award, they want to make sure like, oh yeah, it's sort of well used and everyone uses it and stuff. So, yeah, but but of course the obviously like in that 20 years like, well, everyone does it. It's obvious, but you know,

Host

这是为某部特定电影或其他什么做的具体工作吗?

this was specific work done for a specific movie or something.

John

不,它就像一篇论文,发表在《Cigarette》上。

No, it was like a it was like a paper in cigarette.

Host

它是一篇论文,然后 Pixar 就成了我认为某种程度上重要的核心,就像他们很多

It was it was it was a paper and then Pixar just became like I think somewhat important core to like a lot of their

John

是的。事实上我的一些朋友,事实上很多朋友做了很多这些模拟器。所以是的

Yes. And and in fact some of my friends in fact a lot of my friends did did a lot of these simulators. So yeah

Host

我很好奇,既然你一直与量子计算相关,并直接在 Google 应用科学部门工作过一段时间。我想你现在不在那了,但我很好奇想看看

I'm curious about since you have been quantum computing adjacent and worked on quantum computing directly at Google applied sciences for a while. I think you're not currently on that but I'm curious to see what

John

我还在涉猎,你还在涉猎。是的。那么你认为量子计算这些年的发展轨迹会怎样?因为这是那种东西,我猜有点像核聚变,一开始有种感觉,非常令人兴奋,然后似乎很长时间没有进展,也许现在我们又看到它令人兴奋的迹象,我,或者也许那只是我这种量子相关的看法。

I'm still dabbling in you're still dab. Yeah. So where do you see the trajectory of quantum computing over the over the years going? Because this is one of those things which I guess kind of like fusion was sort of had a a sense at first like being very exciting and then seeming like it wasn't going anywhere for a long time and then maybe now we're seeing hints again of it being exciting and I I'm or maybe that's like my sort of quantum adjacent.

Host

我想很多人都经历过那种。我倾向于平均,我倾向于把东西平均化,我想我现在够老了,会把东西按几十年来平均,而它就是进步。

I think a lot of people have gone through that. I tend to average I tend to average things out over the I guess I'm old enough now where I sort of average things out over the decades and it's just progress.

John

但我的意思是,比如从你知道,从费曼开始,甚至试图弄清楚这作为一个概念意味着什么,一直到现在的你知道,我,最近有 Willow 的量子纠错结果,这实际上似乎真正

But I mean like see going from you know starting from fman just trying to figure out even what this means as a concept all the way up to now where you know I there's this recent Willow result of quantum error correction which actually seems genuinely

Host

用正确的缩放定律之类的可以实现。

the right achievable with like the right scaling laws and stuff.

量子计算时间线 Quantum Computing Timeline

Host

你还认为量子计算要 20 年才能实现吗?或者要等到某个简单但实用的算法真正成功?你觉得这正在加速,还是仍然会……因为我们在很多领域看到它进展非常快,而且非常出乎意料。我想知道你认为这件事会很快实现,还是不会。

Do you still think that quantum is 20 years away, or until some simple but practical algorithm actually succeeds? Do you see this accelerating, or do you think this is still going to be something... because there are a lot of areas where we saw this advance very quickly and it was very unexpected. I'm wondering if this is a thing that you think will be quick or will not be.

John

我觉得介于两者之间。这不纯粹是软件的事,因为我们需要构建足够大、足够稳定的系统,才能成为量子计算机,而不是量子装置。我们正处于所谓的 NISQ 时代——中等规模量子——这个词是 John Preskill 创造的。抱歉,这不是个很好的术语,但我想这就是我们有的术语。量子团队写了一篇很棒的论文——我想我是众多作者之一——关于一种叫量子回声算法的东西,它现在就可能适用。它本质上是费曼提议的风格。技术术语是,你可以尝试将哈密顿量拟合到观测数据,比如 NMR。这意味着你有一个参数化的物理模型,你用量子计算机来调整参数,并在一个循环中试图找出正确的参数。所以你可以,例如,解码 NMR 参数。这是一个实际使用的东西。现在的问题是,它是否会大到足以取得突破——这仍然有待确定。Google 的量子团队在 Neven 几年前制定的路线图上执行得非常好,他们继续沿着这条路前进,不断扩大规模。当他们达到上一个里程碑时,应该能够拥有一台做惊人事情的量子计算机,他们继续推进。所以是的,我认为是几年到十几年的尺度。我没有跟上他们说的确切日期,所以你应该问 Hartmut 具体什么时候会发生。但是的,他们正在前进。有趣的问题是,超导计算会成为赢家,还是其他替代技术之一——这仍然有待确定。我仍然认为超导是非常有前途的,因为它非常可扩展。所以不,我不认为它要 30 年。我不认为它是明天。我不认为会——即使有突然的硬件突破,这些东西非常挑剔。有人可能会有绝妙的想法,但仍然需要一段时间,因为从根本上说,至少在 NISQ 时代或一段时间内,这些本质上是模拟计算机,所以它们往往非常非常挑剔。所以我不期望突然发生某种相变。没有人弄清楚如何将量子比特扩展到任意大小并相互作用——我认为仍然在 100、200 个量子比特的数量级。

I think sort of in between. It's not a purely software thing because we need to build systems that are large enough and stable enough to be a quantum computer instead of a quantum apparatus. We're in the what they call the NISQ era — intermediate scale quantum — which was coined by John Preskill. It's not a very good term, sorry, but I guess that's the term we have. The quantum team made this wonderful paper — I think I'm co-author on one of many — for something called the quantum echoes algorithm, which could be applicable now. It's in the style of Feynman's proposal. The technical term is you could try to fit a Hamiltonian to observed data like NMR. That means you have a physical model that's parameterized, and you use the quantum computer to adjust the parameters and try to figure out inside of a loop what the right parameters are. So you could, for example, decode NMR parameters. That is a practical thing that's used. Now the question is whether it will be big enough to make breakthroughs — that's still TBD. The quantum team at Google has been executing amazingly well against a roadmap that Neven laid out a few years ago, and they're continuing to march down this thing, scaling up. When they hit their last milestone, they should be able to have a quantum computer that does amazing things, and they're continuing to march it along. So yeah, I think on the scale of a few to several years. I haven't kept up on exactly what date they're saying, so you should ask Hartmut exactly when that's going to happen. But yeah, they're marching along. The interesting question is whether superconducting computing will be the winner or one of the other alternate technologies — that's still TBD. I still think superconducting is a very promising thing because it is very scalable. So no, I don't think it's 30 years away. I don't think it's tomorrow. I don't think there's going to be — even if there's a sudden hardware breakthrough, these are very finicky things. Someone might have a brilliant idea, but it will still be a while because fundamentally, at least in the NISQ era or for a while, these are fundamentally analog computers, and so they tend to be very, very finicky. So I wouldn't expect all of a sudden some phase change happens. No one has figured out how to scale up qubits in a way that they interact to arbitrary size — still on the order of 100, 200 qubits I think.

Host

你是指在……好的,那么……

You mean in terms of... okay so...

John

在……方面,嗯,有点像……是的,没错。所以对于超导量子比特,尽管人们正在努力,但物理量子比特与逻辑量子比特的比例仍然相对较大,对于你需要的错误率来说。所以也许会有突破,我不知道。或者有其他东西可能——你不需要那个比例那么高,因为很多与你在 2D 芯片上布局量子比特的 2D 连通性有关。对于像中性原子这样的东西,理论上它们可以连接任何东西到任何其他东西,但当然在实践中我们并不真正知道,我们实际上不知道中性原子的局限性是什么。至少我不知道中性原子的局限性——也许超级专家知道。

In terms of... well it's a little like... yeah yeah exactly. So for superconducting qubits, although people are working on it, the ratio of the number of physical qubits to logical qubits is still relatively large for the error rates that you need. So maybe there'll be a breakthrough there, I don't know. Or there are other things that might have — you don't need that ratio to be so high because a lot of it has to do with the 2D connectivity of the chips that you lay out your qubits on — a 2D chip. For things like neutral atoms, they in theory can connect anything to anything else, but of course in practice we don't really know, and we don't actually know what the limitations of neutral atoms are. At least I don't know the limitations of what neutral atoms are — maybe the super experts know.

Host

特别是超导量子比特有这种张力,你想要有干净清洁的谐振器——你通过与环境解耦来获得干净的谐振器——然后你通过耦合谐振器来获得良好的相互作用,这涉及与环境耦合。所以看起来……

Superconducting qubits in particular have this sort of tension where you want to have nice clean resonators — you get a clean resonator by decoupling from the environment — and then you get good interactions by coupling resonators, which involves coupling to the environment. So it seems like...

John

没错,但至少……但是是的,但就像整个环境,然后还有像小孔,你想要穿过它们到你的邻居。

True, but at least... but yes, but there's like the whole environment, and then there's like little tiny holes that you want to go through to your neighbors.

Host

这不像基本的不确定性关系之类的。这像是技术限制或困难之类的。

This is not like a fundamental uncertainty relationship or something. It's like this is a technological limitation or difficulty or something.

John

是的,Google Quantum 的硬件团队非常非常熟练。他们非常非常熟练。所以是的,他们真的很擅长做这些设计并让这些东西真正工作。所以我发现他们令人印象深刻。

Yeah, the hardware team in Google Quantum is very, very skilled. They're very, very skilled. So yeah, they're really good at making these designs and making these things actually work. So I find them impressive.

Host

是的。

Yeah.

John

看到这种进步会很令人兴奋。

It'd be exciting to see that advance.

Host

是的。是的。是的。

Yeah. Yeah. Yeah.

John

是的。

Yeah.

Host

我想我们只能屏住呼吸等待。

I guess we'll just have to hold our breath and wait.

John

就等待。是的。我想我只是非常耐心,所以我就开始工作。所以提前 20 年……

Just wait. Yeah. I guess I'm just very patient, so I just start working. So 20 years ahead of...

Host

这就是 Dave Bacon,他负责 Google Quantum 的软件团队,总是取笑我,'哦不,我不能在量子计算领域工作,那是 10 年前的事了,因为它会……你总是提前 20 年,所以我必须等 20 年,就像好吧,你知道 10 年过去了……'

That's what Dave Bacon who runs the software team in Google Quantum always teases me like, 'Oh no, I can't work in quantum computing, it was 10 years ago because it's going to be... you're always 20 years ahead of time and so I have to wait for 20 years like okay well you know 10 years have gone by like...'

John

一半了。

Halfway there.

Host

一半了。我不认为那是绝对的,是的,但我也在 10 年前开始研究聚变。

Halfway there. I don't take that as an absolute, yeah, but I also started working on fusion 10 years ago.

John

也许你实际上是因果的——就像你开始研究某件事,现实就赶上了。

Maybe you're actually causal — like you start working on something and reality just catches up.

Host

我想是的。也许。谁知道呢?我不知道。

I guess so. Maybe. Who knows? I don't know.

John

我开始工作——我在 90 年代初就喜欢卷积网络。

I started working — I loved convolutional nets in the early '90s.

Host

Yann LeCun 声称……我查过了。

Yann LeCun claimed... I checks out.

John

我创造了卷积网络这个词。据我所知,这可能是真的,因为每个人都叫它 LeNet,因为它是一个非常具体的东西。他们都为 Yann 工作。我说好吧,我不为 Yann 工作。我不想叫它 LeNet。所以我叫它卷积网络。所以我不知道。那是一个更通用的术语。这是个好术语。

I coined the term convolutional net. As far as I can tell that may be true, because everyone called it LeNet because it was a very specific thing. They all worked for Yann. I said well I don't work for Yann. I don't want to call it LeNet. So I called it a convolutional net. So I don't know. That was a more generic term. It's a good term.

Host

是的。

Yeah.

John

在你走之前,有什么你想让观众记住的吗?有什么信息你想传达的吗?

Before you go, is there anything you want the audience to take away? Any messages you want to deliver?

Host

我认为错误就是一个例子。但我对 AI 在科学中的潜力感到非常兴奋。我认为它将成为科学家的惊人强大工具。而且我认为科学家不会被取代。

I think error is an example of it. But I'm really amazingly excited about the potential of AI for science. I think it's going to be an amazing power tool for scientists. And I think scientists won't be replaced.

对AI研究的期望 Hopes for AI Research

Host

我希望他们会把所有时间都花在创造性的事情上、严谨的事情上、哲学的事情上。所以我觉得这会非常酷。

I'm hoping that they'll spend all their time on the creative stuff, on the rigorous stuff, on the philosophy stuff. So I think it's going to be way cool.

John

我也希望如此。

I hope so, too.

Host

是的。我的直觉是,很多人——我听到很多人这么说。我希望他们是对的。我希望这不仅仅是在应对一个令人不适的现实。

Yeah. My intuition is that a lot of — I've heard a lot of people say that. I hope they're right. I hope it's not just a sort of coping with a reality that's uncomfortable.

消除瓶颈 Removing a Bottleneck

Host

我们差点忘了问的另一个问题:如果你能通过命令,像变魔术一样消除你所在行业的一个瓶颈——无论你怎么定义——那会是什么?

The other question we almost forgot to ask: If you could remove a bottleneck in your industry, by fiat, then like magic — however you want to define that — what would that be?

John

如果我能许一个魔法愿望,我会说:请有人造出那个万能实验室,你只要发送一个 JSON 数据块,它就能做任何实验。

If I could get a magic wish, I would say someone please make the everything lab that you could send a JSON blob to and it will do any experiment at all.

Host

好的。自动化的,

Okay. Automated,

John

但它必须能做任何事。所以本质上,我想我们必须解决那种 AI 完备的机器人学问题。但如果我们做到了,那会非常惊人,因为现在像 ERA 这样的东西,全是计算性的。所以必须有人去收集数据。所以是的,如果我们能突破这一点。哦,那会太棒了。那会绝对太棒了。

but it has to be anything. So essentially, I guess we have to solve the sort of AI-complete robotics problem. But if we did, then that would be stunning because right now things like ERA, it's all computational. So someone has to gather the data. So yeah, if we could just break that. Oh, that would be so amazing. That would be so utterly amazing.

结束语 Closing Remarks

Host

是的。太好了。我真的很感谢你抽出时间来见我们。而且我觉得你是飞过来的,还稍微调整了你的日程来

Yeah. Great. So I really appreciate you taking the time to see us. And I think you kind of flew in and adjusted your schedule a little bit to

John

是的,我本来是要飞越旧金山回家,所以我在旧金山降落了。

Yeah, I was sort of flying over San Francisco to get home and so I landed in San Francisco.

Host

所以你费了很大劲才来到这里。我们真的很感激。和你聊天真的很有趣。

So you made a big effort to be here. We really appreciate that. It was really fun to talk to you.

John

嗯。

Mhm.

Host

是的。听着,这非常愉快。

Yeah. Listen, it's been a blast.

John

好的,酷。谢谢你邀请我。

Okay, cool. Thank you for having me.

Host

是的,不客气。

Yeah, you're welcome.

互动版:逐字朗读 + 针对本期提问 →