AI in Math: From IMO to Millennium Problems
打开互动全文版(中英对照 + 朗读 + 问答)→Grant Sanderson 讨论为何 AI 在数学上的成功并不自动意味着通用人工智能,以及解决黎曼猜想是否意味着白领工作的自动化。
Grant Sanderson discusses why AI's success in math doesn't automatically lead to AGI, and whether solving the Riemann Hypothesis would imply automation of white-collar work.
今天我和 Grant Sanderson 聊天,他运营着 Blue and Brown,现在正在做一个新项目,记录 AI 在数学领域取得的进展。我想和你聊聊这个,因为 AI 在数学上的进步是所有领域中最快的。所以,这里发生的一切,以及我们看到 AI 进步的方式,都会告诉我们,随着 AI 变得越来越好,世界其他地方会发生什么。我想从三年前我第一次采访你时问的那个问题开始。我问你,一旦 AI 能在国际数学奥林匹克竞赛中获得金牌,那不就是 AGI 吗?考虑到这些题目有多难,它难道不能做任何人类能做的事吗?你的回答事后看来非常明智且正确:它只会像所有其他被攻克的基准一样,成为又一个基准。显然,从那以后 AI 在整体上变得更好了,但不会有一个顿悟时刻。首先,我很好奇你的直觉,为什么这被证明是对的。其次,我好奇这种狭隘性还能持续多久。所以,当 AI 解决了千禧年大奖难题时,你认为那时是否仍然可能有很多人类在做的工作,AI 还无法在经济中自动化?
Today I'm chatting with Grant Sanderson who runs through Blue and Brown and is now working on a new project documenting the progress AI is making in math. I wanted to talk to you about this because AI has been making the fastest progress in mathematics as of any other field. So whatever is happening here and whatever way we're seeing AI progress happen or not happen would tell us about what will happen to the rest of the world as AI gets better and better. So, I wanted to start with this question I asked you when I first interviewed you three years ago. And I asked you once we have AIs that can get gold in the International Math Olympiad, wouldn't that just be AGI? Wouldn't this just be able to do anything any human can do given how hard these problems are? And you had an answer which in retrospect turned out to be very wise and correct which is like it'll be another benchmark like all these other benchmarks that are passing. Obviously, it has gotten better in general ways since then, but there won't be some aha moment when this happens. First, I'd be curious to get your heuristics on why that turned out to be true. And second, I'm curious how long you think this narrowness can continue to be true. So, by the point that AI has solved the millennium prize problem, do you think it's still possible that at that point there's lots of tasks that humans are doing that AI still can't automate in the economy?
这是个有趣的问题,因为如果不提前知道解决方案长什么样,就很难回答。我的意思是,就拿 IMO 来说,我认为你三年前那个问题的精神在于,这些问题的解决方案似乎确实需要创造力,而问题的设计者会试图让你无法轻易通过训练来应对。但 IMO 的一个肮脏秘密是,很多题目其实真的可以通过训练来掌握。所以,随着整个 AI 和数学项目的进行,我认为你指出的一个有趣之处在于,AI 有一个尖峰前沿,数学正好就在其中一个尖峰上。但这种尖峰有某种分形性质,因为当你放大数学内部的特定进展时,有些东西比其他东西容易得多。所以,如果我们只考虑 IMO——这已经是旧闻了——大约两年前它们就表现得相当好了。如果不是因为以下原因,它们在 2024 年本应获得金牌。它们非常擅长,基本上直接解决了几何题。IMO 有四类问题:几何、数论、代数和组合数学。几何题在 2024 年只需 19 秒就能解决,因为它有点像暴力求解器。而肮脏秘密是,对学生来说,也有某种暴力方法可以应对。组合数学是百搭题,更像是需要玩味的谜题。那年的考试中有两道组合数学题。并非总是如此。有四类,六道不同的题。所以哪类会有两道题是随机的。如果几何题更多,它们那年就能拿金牌。但它在组合数学题上表现不佳。你知道,那些试图守住数学最后阵地的人可能会说,嗯,那些才是更需要创造力的题目。即便如此,我认为你问题的精神是,如果它们解决了千禧年大奖难题,那是否也能服务于大量白领工作?这暗示着,从现在到那个目标之间的速率限制因素,与让白领工作变得更好的速率限制因素是相同的。我们可以设想几种不同的方式,比如专注于黎曼假设——解决它会是什么样子?一种可能性是,这些东西在某个特定知识领域极其擅长,并且非常深入地了解它,然后了解另一个领域,再了解另一个领域。你指出过这一点。拥有这种超人类广度、对所有领域都如此了解的东西,却不仅仅是找到连接它们的闪电,这很奇怪。我认为我们开始看到一些火花,即它实际上找到了它擅长的东西之间的联系。我相信我们会谈到这个。如果黎曼假设的解决方案具有这种性质,那对我来说,与做好白领工作所需的能力截然不同。而且有理由相信,这实际上可能就是解决方案的性质。不知道你是否知道 Hugh Montgomery 和 Freeman Dyson 在 IAS 谈话的故事。这是个题外话,但这是个有趣的故事,我不知道是不是在午餐时发生的。基本上,这位数论学家试图理解黎曼 zeta 函数零点对之间的统计相关性。黎曼假设就是关于所有这些零点是否都在一条直线上,他找到了一个可以问的定量问题,并写下一个公式,看起来像 1/sin^2 之类的。物理学家 Freeman Dyson 说,我知道那个表达式。那个表达式出现在研究随机 Hermitian 矩阵的特征值中,而这又出现在研究原子核能级时。这两个看似不同的事物的统计规律相同,这引发了一个潜在的探索:嘿,随机矩阵理论是否有某些方面可能与黎曼 zeta 函数相关?我认为这还是个开放问题,比如那里是否有成果可挖?这种连接两个不同领域的方式,如果黎曼假设的解决方案是进一步探索这样的想法,那它具有你期望 LLM 擅长数学的那种特征:它们是量子物理的专家,是解析数论的专家,它们应该能够看到这种相似性,而不需要 Montgomery 和 Dyson 一起吃午餐并恰好聊到那个。这与白领工作完全不同,对吧?就比如你很难用 AI 做编辑,并不是因为它什么都知道,你只需要它找到中间的闪电。
It's an interesting question because it's hard to answer without knowing what the solution looks like ahead of time. I mean, if we take the IMO, that's something where I think the spirit of your question three years ago was in looking at how some of the solutions to these problems really seem to require creativity and the designers of these problems, they'll try to have them come up with things that you can't train for as easily. I think the dirty secret with the IMO is that you really can train for a lot of them. And so with the whole AI and math project undergoing, I think as you point out one of the reasons it's interesting at all is that there's a spiky frontier to AI, math is just right there in one of the spikes. But there's kind of a fractal nature to that spikiness because when you zoom into the specific progress within math, you have some things that are a lot easier than others. So if we just think about IMO which is old news at this point, it's kind of like two years ago that they're really doing quite well. They would have gotten a gold in 2024 if not for the following reason. They were very good. They just cold solved geometry basically. And the IMO has these four categories of problems: geometry, number theory, algebra, and combinatorics. So geometry just solves in like 19 seconds in 2024 because it's kind of a brute force solver. And the dirty secret is for students there's also sort of a brute force way that you kind of can go at it. Combinatorics is the one that's the wild card of much more like playful puzzly seeming problems. And there were two combinatorics problems on that year's test. There's not always. There's four categories, six different problems. So it's kind of a tossup which one is going to have two questions. Had it been more geometry questions, they would have gotten a gold that year. But it struggles on those combinatorics ones. And you know, someone who's trying to keep that torch of the last holdout of math for humanity might say, well, you know, those are the ones that require the more creativity. Even then though, I think the spirit of your question on like if they're solving a millennium prize problem, does that also service a lot of white collar work? It suggests that whatever the rate limiter is between where we are now and that is the same as the rate limiter for making things better at white collar work. We could maybe paint a couple different ways that we focus on, I don't know, Riemann hypothesis, like what would it look like to solve that? One possibility would be these things are extremely good at a specific domain of knowledge and just knowing it very deeply and then knowing another domain and knowing another domain. And you've pointed this out. It's like bizarre to have something with this superhuman breadth that knows all the fields so well that's not just finding those lightning bolts that connect them. I think we're starting to see sparks of that of actually finding connection between the things that it's an expert at. I'm sure we'll talk about it. If the nature of the solution to the Riemann hypothesis was something like that, that feels pretty distinct to me than what's necessary to get good at white collar work. And there's a reason to believe actually that that might be the nature of the solution. I don't know if you know the story of Hugh Montgomery and Freeman Dyson at the IAS like talking. This is a side tangent, but it's just kind of a fun story on how, I don't know if it was over lunch or something like that. Basically, you have this number theorist who is pointing out just trying to understand the statistical correlation between pairs of zeros of the Riemann zeta function. So the Riemann hypothesis is all about do all these zeros sit on a straight line and he's finding this quantitative question you could ask about and he writes down a formula. It looks like 1 over sin^2 or something like that. Freeman Dyson, a physicist, is like, I know that expression. That expression comes up in studying the eigenvalues for random Hermitian matrices, which was something that comes up in studying the energy levels of a nucleus. And the idea that the statistics of those two seemingly different things were the same, sort of prompted a potential exploration on, hey, are there aspects of random matrix theory that might be relevant to Riemann zeta function? And I think it's a little bit of an open question, like is there fruit to be had there? That kind of bridging together from two different fields, if it turned out that the solution to the Riemann hypothesis was exploring an idea like that even further, that has this character of kind of how you expect LLMs to be good at math: they are an expert at the quantum physics, they're an expert at the analytic number theory, they should be able to see that similarity in a way that doesn't require like Montgomery and Dyson to be having lunch and happening to talk about that. That's totally different from white collar work, right, in terms of the extent to which you maybe have a hard time using an AI as an editor. It's not because they know everything and you just need them to find that lightning bolt in between.
另一种可能性是……什么类比合适呢?也许可以想想费马大定理,从费马提出这个问题到最终解决方案的样子,解决方案涉及极其复杂的数学工具,对吧?这个问题的美妙之处在于,你可以用非常简单的方式表述它。你问的是 x 的 n 次方加 y 的 n 次方等于 z 的 n 次方,当 n 大于 3 时是否存在整数解?你可能会期待有一种初等数论方法能解决它,但就我们所知,并没有。而实际的解决方案,也许有更简单的办法,但这可能就是必须走的路。它建立在一套极其复杂的思想之上,这些思想基于几个世纪的工作,围绕椭圆曲线展开,还有另一座围绕模形式的思想大山。这两座大山都必须先建好,你才能提出连接它们的问题。所以,如果黎曼猜想的解决方案需要建造一座新的大山,那是一种技能——提出正确新思想的能力——这感觉与它们目前智能的特性截然不同,并不是你从雇来的视频剪辑师那里需要的东西。但如果它有能力建造大山,即正确的新理论,能结晶化我们该如何思考一个学科,那这种智能水平高到令人惊讶,如果它不渗透到经济其他领域,仅仅停留在数学本身的大山建造上,那才奇怪。
Different possibility would be... what's the right analogy? Maybe like if we think of Fermat's Last Theorem, between the moment of Fermat phrasing the question and then what the solution itself looks like, where ultimately the solution involves such heavy machinery in math, right? So the beauty of that problem is you can phrase it so simply. You ask about x to the n plus y to the n equals z to the n. Do you have integer solutions for this when n is bigger than 3? And it's something you might expect there to be an elementary number theory approach to it, but just as far as we can tell, there's just not. Whereas the actual solution, you know, maybe there is something simpler, but this might be what it has to be. There's such a complicated set of ideas that build on centuries of work, centered around elliptic curves and then this other mountain of ideas centered around these things called modular forms, and both of those mountains have to be built before you can ask the right question that connects it. So if the solution to the Riemann hypothesis involved building a new mountain, that's a kind of skill — the ability to come up with the right new ideas — that feels sufficiently different from the character of how they're intelligent right now that it's not like that's what you need from your hired video editor per se. But if it's capable of building mountains that are the correct new theory that crystallizes how we should be thinking about a subject, that's just such a level of intelligence that then it starts to feel like it would be surprising if that didn't permeate into other aspects of the economy besides just the mountain building for math itself.
是的。或者至少,即使它不能真的做到白领人类能做的每一件事,它也会产生变革性的影响,就像在 IMO 拿金牌并没有对世界产生变革性影响一样。首先,我想指出我完全是在移动球门柱,因为当我大约两三年前采访 Dario 时,我问过这个问题:为什么他们没能利用广博的知识将想法联系起来,从而做出新发现?这似乎是那种即使一个中等智力的人拥有这么多信息,也能从某种药物引起偏头痛而另一种药能治这个那个,也许同一种药能同时治疗这两种病,从而得出医学诊断的事情。嗯,我不知道。从外行人的角度看,数学显然是一个领域,其中找到单位距离问题猜想的反例就是这类事情的一个例子。所以完全是移动球门柱。但接下来我们可以问:既然 AI 能做这件我们本以为它们应该能做的事,下一个基准是什么?下一件令人印象深刻的事是什么?这里有几个候选想法。一个是首先提出有趣的问题,另一个是提出新的对象或概念化,从而创造或统一领域。关于第一个。现在我们只是训练这些模型……我们有这些千禧年大奖难题,因为,我不知道,数学家们注意到黎曼提出了黎曼 zeta 函数这个概念,因为他认为它可能与素数的密度有关,或者这个函数的零点可能与素数有关。所以弄清楚为什么我们一开始认为这是一个有趣的研究对象?为什么我们要构建这个对象并试图回答关于它的问题,特别是这个问题?这似乎是下一个基准。
Yeah. Or at the very least, even if it couldn't literally do every single thing white collar humans can do, it would just have transformative effects in the way that getting gold in the IMO did not have transformative effect on the world. First of all, I do want to point out that I'm totally moving the goalpost here because when I interviewed Dario about two, three years ago, I asked this question about why haven't they been able to use their vast knowledge to connect ideas together and come up with a new discovery that way. That seems like the kind of thing even if a moderately intelligent person had knew this much information, they'd be able to come up with a medical diagnosis from the fact that this drug causes migraines and this other thing, you know, whatever does this and maybe that it's the same drug that can cure both things. And yeah, I don't know. From an outsider's perspective, mathematics seems clearly like a field where finding this counterexample to the unit distance problem conjecture was like an example of this kind of thing. And so total goalpost moving. But then we can ask okay what is the next benchmark now that AIs can do this thing that we should have thought they should be able to do? What is the next thing that would be quite impressive? And there's a couple of candidate ideas here. So one could be coming up with interesting problems in the first place and the other is coming up with new kinds of objects or conceptualizations that create or unify fields. On the first one. Right now we just train these models to... we have these Millennium Prize Problems because, I don't know, mathematicians have noted like Riemann came up with this idea of this Riemann zeta function, and because he thought that it would have some connection with the density of prime numbers, or if the zeros on this function would have some connection to prime numbers. So figuring out that why do we think this is an interesting thing to study in the first place? Why were we building this object and trying to answer questions about it and answer this particular question about it? Seems like the kind of thing that would be the next benchmark.
我的意思是,你举了两个很好的例子。对于好奇单位距离猜想的人来说,有一个叫 Polylog 的数学频道制作了一个很棒的视频,他们讨论了这个问题。其中一个人,因为所有这些讨论都促使人们反思做数学的过程,对吧?他们说,“啊,这个东西能做这么令人印象深刻的事,这对我们意味着什么?”他引用了这样一句话:好的数学家证明定理,伟大的数学家提出猜想,最伟大的数学家提出定义。这几乎完全就是你这里的框架——我们需要猜想生成器,然后是定义生成器。那是最高级别的数学家。我不太明白如何将其具体化为一个基准,因为通常当我想到“基准”这个词时,我想到的是类似球门柱的东西,球是否穿过球门,你可以明确地说“是的,这完成了”。部分原因是为了能够做像 RLVR 这样的事情,但部分原因也是为了能够知道你没有移动球门柱。在回答中,你知道,OpenAI 可以有一个标题说“否证了单位距离猜想”,因为这是一个清晰明确的事情——它做对了。而想象一下试图有一个标题说“D5.4 提出了一个非常好的猜想”,对吧?就像我们保证每个人都认为这是一个好猜想。但它就是没有同样的效果。
I mean, you highlight two pretty good examples there. For anyone curious about the unit distance conjecture, there's this really nice video by math channel called Polylog where they talk about it. And one of the people in that, because all of these discussions causes people to reflect on the process of doing math, right? They're like, "Ah, this thing can do this impressive stuff, what does that mean for us?" And he highlights this quote: how good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest mathematicians come up with definitions. And that's more or less exactly your framing here on like those two — we need the conjecture generator and then the definition generator. That's the premium tier mathematician. I don't understand how exactly you would make that a benchmark in the sense that usually when I think of the word benchmark I'm thinking something that you have like it's a goalpost that the ball is through the goal or it's not, like you can clearly say yes this is done. Partly to be able to do things like RLVR but also partly just to be able to know that you haven't moved the goalpost. In answering, you know, OpenAI can have their headline on disproving the unit distance conjecture because it's a clear distinct thing — it's like it did it right. Whereas imagine trying to have a headline on like "D5.4 came up with a really good conjecture," right? Like we promise everyone thinks it's a good conjecture. It just doesn't land the same way.
但也许这并不否定这是值得思考的正确方向。所以,如果它最终以基准的形式出现,并且我们有一个分数说它通过了这个基准,因为我们可以量化一个猜想有多好,我会感到惊讶。但可能真正需要的是,你会感觉到与数学家对话时的语气转变,关于它与他们合作的有用方式,对吧?你提到的那个系列,目前还没有制作,可能几个月内也不会,形式是我们采访了很多数学家。有趣的是,我们大约一年前就开始做这个了。看到他们在 2025 年中期和现在 2026 年之间谈论 AI 的方式有了一点语气转变,这很有趣。你知道,在现实世界中,那是非常短的时间。在 AI 世界中,那是永恒,对吧?我们能够看到在这永恒中语气的转变。
But maybe that doesn't negate the fact that that's the right thing to be thinking about. So I would be surprised if it ever took the form of looking like a benchmark and we have a score saying that it's passed this benchmark because we can quantify how good a conjecture it is. But probably the nature of what it would take is that you would feel a tone shift in conversations with mathematicians about the way that it's useful to work with, right? And this series that you referenced that is not at all produced yet and probably won't be for a couple months takes the form of us interviewing a lot of mathematicians. And what's interesting is we started doing this like over a year ago. And it's fun to see a little bit of a tone shift in the way that they talk about AI between mid 2025 and where we are now in 2026. You know, in the real world that's a very short amount of time. In the AI world that's eons, right? And we're able to see over those eons this tone shift.
我认为衡量猜想生成能力的方式会更主观,比如那种语气转变——数学家们说他们不只是用它来解决问题,而是退后一步决定他们的研究领域应该是什么,与某个模型的对话对此确实有帮助。我不太可能看到它以头条新闻的形式出现,说这又是另一个被攻克的基准测试。
I think the way that you'd measure conjecture generating ability is going to be more subjective, like that tone shift where mathematicians say they're not just using it to solve their problems, but as they step back and decide what their research field should even be, a conversation with such and such model was genuinely helpful for that. I don't think it's likely that you'd see it in the form of a headline saying that this was yet another benchmark knocked down.
对吧?所以很有趣的是,那些你无法为其制定基准测试的东西,至少在当前范式下,也是你无法轻易训练的东西,对吧?因为基准测试和训练环境之间其实没有根本区别。
Right? And so it's very interesting the kinds of things you can't make benchmarks for are also the kinds of things, at least in the current paradigm, you can't easily train for, right? Because there's really no fundamental difference between a benchmark and a training environment.
是的。
Yes.
我认为很容易提出某种二分法,比如“这是 AI 不能做某事的深层原因”,然后结果证明,你只是思考方式不对,实际上它很快就能做到。但我还是会提出几个。而且我认为,很可能在相对短期内,我们会有办法训练 AI 做这类事情,但这似乎必须不同于当前的强化学习训练。
I think it's very easy to come up with some dichotomy of like "here's a deep reason why AI can't do a certain thing" and then it turns out, well, you're just thinking about it the wrong way and actually it can do it pretty soon thereafter. But I'm going to come up with a couple anyway. And I think that this will probably turn out that there are ways in which we can train AI to do these kinds of things in the relatively near term, but it seems like it would have to be different from current RL training.
所以我好奇的是,在我看来驱动数学和科学领域许多重大进展的东西,是提出一种思考问题的新方式或理解世界的新方式,从而统一不同领域、催生全新领域、解决我们最初甚至没想过要解决的问题。比如爱因斯坦思考广义相对论,并不是因为他想解释为什么光会弯曲或黑洞为什么存在。这些现象他最初甚至不知道需要解释。但在数学中,从外部来看,似乎常常有这样一种方式:证明一个特定问题可以激发一种新的概念化,进而产生一个全新的领域、一种全新的思维方式,这种思维方式极具生产力,而另一种则不然。我很想听听你谈谈,伽罗瓦提出群论并区分出他对五次方程根式无解的解,而阿贝尔几年前提出了一个不同的证明却没有提出群论。但如果你想做一个验证循环:“群论是一个有趣的概念吗?这里有什么有用的东西吗?为什么这个证明更好?”这个验证循环可能长达 100 年,涉及密码学的出现、物理学的进步、群论思想的相关性以及理解物理学中的对称性等等。就像一个长达 100 年的验证循环:“为什么这首先是一个富有成效的概念?”
So the thing I'm curious about, and the thing it seems to me that drives a lot of the big progress in mathematics and in science generally, is coming up with a new way to think about a problem or a new way to understand the world that then unifies different fields, spawns entire new fields, solves problems we weren't even thinking we were trying to solve in the first place. Like the reason Einstein was thinking about general relativity is not because he wanted to explain why light bends or why black holes exist. These are phenomena he didn't even know needed to be explained in the first place. But in mathematics, it often seems, okay, a total outsider, I don't even know the details of what I'm talking about here, from the outside it seems like there's often ways to say prove a specific problem that can motivate a new conceptualization, one which results in a whole new field, a whole new way of thinking, which is immensely productive, and one which doesn't. I'd be curious to hear you talk about whether Galois coming up with group theory and distinguishing his solution to the quintic having no formula for the roots, and Abel coming up with a different proof a few years earlier that didn't come up with group theory. But then if you wanted to do a verification loop on "is group theory an interesting concept? Was something useful done here? Why is this proof better?" Potentially that verification loop is 100 years long and it involves cryptography coming around, physics making progress, the ideas in group theory being relevant and understanding symmetries in physics and all those kinds of things. Like a 100-year verification loop of "why is this a productive concept in the first place?"
是啊。哎呀,等等,你戳到我的痛处了,因为我有一个关于伽罗瓦的项目,本来打算 2022 年做,但搁置了,不过我花了大约一年的时间思考他的工作。所以我有可能会不小心在细节上说得太长,你拉住我。这对你的情况来说是一个完美的例子,因为描述为什么它是一个有价值的见解,并非来自直接的实用性。所以当然,如果你考虑强化学习验证环境,这将会非常难做。但有趣的是,即使是当时的人类验证者,也花了很长时间才认识到它的用处。我认为爱因斯坦的广义相对论,人们几乎立刻就觉得这是一个好理论。而伽罗瓦理论之所以是一个如此有趣的例子,是因为你有一个长达 100 年的思想片段,它流经许多不同人的头脑,才最终成为数学界公认的好东西。
Yeah. Boy, wait, yeah, you struck a nerve because I had this project about Galois that I was going to do in 2022 that I put on the shelf, but I spent like a year of my life thinking a lot about what he did. So there's a risk of me accidentally talking too long on the specifics; hold me back on it. It's a perfect example for your case because describing why it was a valuable insight does not come from immediate utility. And so certainly if you're thinking about RLVR environments, it's like, okay, this is going to be really hard to do. But it's interesting to note how even with human verifiers at the time, it took a really long time to recognize it as being useful. Like I think Einstein with GR, people sort of felt, you can feel this feels like a good theory right away. What makes the Galois theory such an interesting example is you have literally this 100-year segment of an idea that flows through many different people's heads before it settles into something that the math community agrees is good.
所以稍微回溯一下,你想了解这个问题的背景吗?
So to back up a little bit, do you want the background on the problem at all?
好吧。嗯,我们都在学校学过二次公式。
All right. Well, we all learn about the quadratic formula in school.
我以为你会说我们都在学校学过群论。
I thought you were gonna say we all learn about group theory in school.
我们都知道我错过了那节课。我们都在学校学过关于二次公式的群论。所以这在某种意义上已经为人所知;比如希腊人能解二次方程,但他们并没有真正用代数来写。所以更像是阿拉伯人写下了那个公式。有一个有趣的故事,关于一些决斗的意大利数学家,不是真正的决斗,而是智力挑战,他们秘密地找到了三次方程的公式,然后很快又找到了四次方程的公式。所以数学家的自然开放问题是:你能找到一个解五次方程的公式吗?四次公式是怪物。把它写下来很疯狂。你通常不会完整地写下来;你把它分解成一个过程性的东西。所以你可能认为这些东西有指数增长的复杂性。所以几百年来没有人真正回答这个问题。通常我们说阿贝尔是第一个证明的人。他是挪威一位早熟的年轻数学家,他证明了这根本不可能。不是你能找到五次公式。他以为自己找到了一个,但他证明了不可能。不过我认为真正的功劳,你得稍微回溯一下,谈谈拉格朗日,他找到了关于这个问题的正确提问方式。如果你愿意,我可以深入细节,但我只讲一个非常概括的层面。他在研究这个问题,并认识到能够解这些多项式实际上与理解某些代数表达式的对称方式密切相关,或多或少。比如如果我写下 a + b + c + d,只是把四个变量加起来,如果我置换它们,表达式的值不变。而如果我写 (a + b) * (c + d),有些置换不改变它,但有些会。他有一个非常精彩的见解:如果你能找到这样的表达式,有四个自由变量,但所有置换只取三个不同的值,这与能够将四次方程降为三次方程有出人意料的关系。所以他开始通过问“嗯,我想知道我能不能扩展这个”来探讨“我们能不能找到五次多项式”。
We all know I missed that class. We all learned about group theory about quadratic formula. So this was known in some sense; like Greeks could solve quadratics, but they didn't really write things in algebra. And so it's really more like the Arabs that wrote down that formula. There's this delightful story around some dueling Italian mathematicians, not real duels, just intellectual challenges, who secretively found a formula for the cubic, and then very shortly thereafter found a formula for degree 4 polynomials. So natural open question for mathematicians is: can you find a formula that solves degree 5 equations? Now the degree 4 formula is monsters. It's wild to write it down. You usually don't really write it down in full; you break it up as a procedural thing. So you might believe these things have exponentially increasing complexity. So many hundreds of years nobody is really answering that question. Usually we say Abel was the first to prove it. He was this young precocious Norwegian mathematician and he showed it's simply impossible. It's not that you can find a quintic formula. He thought he found one but he showed it's impossible. I think the real credit though, you have to back up a little bit and talk about Lagrange, where Lagrange found the right kind of question to ask about this. I can go into the details if you want, but I'll give it a very high level. He was studying the question and he recognized being able to solve these polynomials is actually very related to understanding the way that certain algebraic expressions are symmetric, more or less. So like if I write down a plus b plus c plus d, just adding four variables, if I permute those it doesn't change the value of the expression. Whereas if I write like a plus b multiplied by c plus d, some of the permutations don't change it but some of them do. And he had this really nice insight about how if you can find expressions like this that have four free variables but all the permutations take on three distinct values, that had this unexpected relationship with being able to reduce degree 4 into degree 3. So he started approaching the "can we find a quintic polynomial" by saying, "Hmm, I wonder if I can extend that."
要扩展那个方法,你需要一个带有五个自由变量的表达式,使得当你用所有 5! 种排列去置换它们时,它只取四个或更少的值。这就像你可以把它放进谜题书里,放进一个 12 岁孩子也能参与的脑筋急转弯里。而且你很容易就会觉得那是一项不可能完成的任务。所以拉格朗日坐在这里想:‘嗯,这是我试图解决这个问题的一个策略。我能找到一个五次多项式,这个策略对它不适用吗?看起来可能是不可能的,至少从这个策略来看。’但那是历史上第一次,人们本能地意识到某种关于对称性的问题才是研究这些多项式的正确方式。在他心里,那只是一个方法;当时尚未发现的是,实际上存在更紧密的联系,而且也许我们不应该寻找公式,而应该问相反的问题:你能证明它不可能吗?所以他算是播下了那颗种子。
And to extend that method you would have to have an expression that has five free variables such that as you permute them over all the five factorial permutations it takes on only four values or fewer. So that's like you could put that in a puzzle book. You could put that in a brain teaser that like a 12-year-old can engage with. And it's not too hard to find yourself feeling like that's an impossible task. And so Lagrange is sitting here saying, 'Hmm, here's a strategy that I'm trying to solve this problem. Can I find a quintic polynomial this strategy doesn't... it seems like it might be impossible, at least from this strategy.' But that was the first time in history that people had the instinct that some kind of question about symmetry was the right way to be studying these polynomials. In his mind it was just a way; it had yet to be discovered that actually there's a tighter connection and also like maybe rather than searching for the formula we should be asking the opposite question: can you prove that it's impossible? So he sort of planted that seed.
大约 50 年后。阿贝尔肯定读过拉格朗日并受其影响。伽罗瓦,我们知道他在爱上数学时非常喜欢拉格朗日。所以很难想象,这两位年轻的天才,他们都在那个问题上产生了非常相似的见解,这并非源于拉格朗日。但回到你的问题:你能验证这是一个好主意吗?拉格朗日并没有得出任何结果。他从未解决这个问题,因此我们也就无从知道那是否是正确的问题。他提出了它。它本身有一些内在有趣的东西。而且当时它对数学并不那么重要;大多数人更感兴趣的是它在物理学中的应用。这几乎属于那种边缘的、近乎娱乐消遣的东西。比如阿贝尔,你知道,他开始研究五次方程,但后来有人建议他把更多精力放在研究椭圆函数上,所以他在英年早逝前的主要工作都在那方面。他 26 岁死于肺结核。然后伽罗瓦,他把这两个想法都推向了正确的方向,真正理解了抽象的本质。他实际上在监狱里写了一篇非常精彩的文章。他就像……我们可以聊聊他的人生故事,相当疯狂。但他只是个十几岁的少年。他在监狱里。他曾试图提交他的数学论文,但被拒绝了。所以这又回到了可验证的奖励。当时的验证者——学院——拒绝了他写的东西,因为坦率地说,它不太连贯;不是一个完整的证明。他没有清晰地阐述这个理论到底是什么。他只是一个初出茅庐的年轻数学家,还在摸索方向。所以那里的验证奖励是‘嗯,不行。’但他有种直觉,觉得里面有些东西。所以他写了一篇长篇大论,论述数学的本质是随时间推移而变化的,他谈到了代数本身的出现,从仅仅用数字思考到熟练运用纯粹的代数表达式,而不受这些表达式解释的束缚。他有一种直觉,认为还有另一层抽象,那才是我们应该做的:不是思考公式本身,而是思考这些公式背后的对称性。但这仍然是一个相当模糊的理论。所以如果你想问:‘好吧,可验证的奖励是他解决了一个别人没解决的问题吗?’嗯,阿贝尔证明了五次方程不可解。你会问:‘那伽罗瓦在做什么?’原则上,伽罗瓦理论能让你针对一个具体的多项式,给出规则来判断这个多项式是否有你能写出来的根。例如,x^5 - 1,你知道一个解是 1;或者 x^5 - 2,你可以写出 2 的五次方根。所以并不是每个五次多项式你都写不出解,但你能找到一个具体的多项式,证明你无法用根式写出解吗?他也没有完全解决这个问题;他的理论要抽象得多……他没有针对一个具体的例子证明不可能。所以甚至描述他解决了什么问题都非常棘手。然后他死了。这是一个非常浪漫的故事:他进行了一场决斗。我们可以深入聊聊。有很多传说,比如据说他在决斗前夜写下了所有想法。
Like around 50 years later. Abel definitely read Lagrange and was influenced by it. Galois, we know that he loved Lagrange when he was like falling in love with math. And so it's very hard to imagine that like these two young geniuses, the fact that they both come up with like pretty similar insights around that problem, it's not like born from Lagrange. But to your question on like are you able to verify that this was a good idea? There wasn't any like result that Lagrange came to. There's never like he solved the problem and therefore we know that that was like the right question to ask. He asked it. There's some like intrinsically interesting thing. It also wasn't very important for math at the time; like most people were more interested in like the applications to physics. This is almost in that like side, almost recreational hobbyist type thing. Like Abel, you know, he started working on quintic stuff but then he was advised to spend more of his effort studying elliptic functions, and so more of his work was on that before he died young. He died at 26 from tuberculosis. And then Galois, he pushed both of those ideas like to the right direction where he really understood the nature of abstraction. And so he had this really nice piece that he wrote while he was in prison actually. He was like... we could talk all about his life story. It's pretty wild. But he's like this teenager. He's in prison. He had tried to submit his math papers and they had been rejected. So again it's like verifiable reward. The verifier function that is the academy at that time is rejecting what he wrote because frankly it was not very coherent; like it wasn't a complete proof. He wasn't giving like a clear thought of like what the theory actually was. He was just like a young fledgling mathematician getting his bearings. So it's like the verified reward there is like 'eh, no good.' But he has some instinct that there's something there. So he's writing this diatribe on like the nature of math being something which undergoes these shifts over time, and he talks about like the advent of just algebra itself and going from just thinking in terms of numbers to having a certain fluency just with like pure algebraic expressions where you're not tied to interpreting those expressions. And he has this instinct that there is another layer of abstraction that seems like what we should be doing, where rather than thinking about the formulas themselves, thinking about like what symmetries underlie those formulas. But it was still a pretty ill-defined theory. So if you're trying to say, 'Okay, is the verified reward that he has solved a problem that other people haven't?' It's like, well, Abel proved that quintics are unsolvable. And you say, 'What was Galois doing?' Well, in principle, the thing that Galois theory will let you do is take a specific polynomial and it gives you the rules to say, does that specific polynomial have roots that you could write down? For example, like x^5 - 1, you know that a solution is 1, or x^5 - 2, you can write down fifth root of 2. So it's not that every quintic polynomial you can't write down the solution, but could you find a specific one where you prove you can't write the solution using radicals? He also didn't even solve that exactly; like he has a much more abstract... he didn't show for a specific example that he couldn't. So even describing like what problem did he solve is very tricky. So then he dies. It's this very romantic story of he has this duel. We can get more into it. There's a lot of myth around, like supposedly he writes up all his ideas the night before the duel.
真的吗?他试图让它们发表。
Really? He tried to get them published.
五次方程似乎对你的健康不利。
The quintic doesn't seem to be good for your health.
非常糟糕。对。对。对。对。如果你是个年轻的天才,别研究五次方程。所以他让他的兄弟和密友把这些笔记交给高斯。交给当时重要的数学家,因为我觉得这里面有东西。即使那样,它也没有真正被接受。他的兄弟和朋友试图把它们传出去。又过了 20 年,刘维尔才看到这些笔记,觉得里面可能有些东西,并试图整理它们,理解伽罗瓦到底想表达什么。然后又是大约 20 年,直到若尔当真正整理出类似现代群论的东西,并将其归功于伽罗瓦。你可以很容易地想象历史会不同,如果这些想法是从数学的其他地方产生的,而伽罗瓦如果是一个不那么生动的人物,他可能就被历史遗忘了。
It's very bad. Yeah. Yeah. Yeah. Yeah. If you're a young genius, don't work on the quintic. And so he asks his brother and his close friend like get these notes to Gauss. Get these notes to like the important mathematicians of the day cuz I think there's something here. Even then it didn't really take. So his brother and his friend like tried to get them out. It wasn't another 20 years until Liouville like sees these notes, sees that maybe there's something in them and tries to like clean it up and understand like what was Galois getting at? And then even then it was another 20 years or so until Jordan actually puts together something like a modern treatment of group theory that they attribute it to Galois. You could easily imagine history turning differently where like these ideas were kind of coming about from other points in math and like Galois could have been forgotten in history if he was a less like vivid character.
但从拉格朗日隐约觉得“根的对称性可能是正确方向”,到它看起来像现代群论,这中间跨度很长。很多时候,它甚至通不过人类审稿人的验证奖励,对吧?因为到了某个人桌上,他们说“我不太确定这里面有什么东西”。到了另一个人桌上,他们也不确定。你必须有那么一个人能识别出它的价值。而且即便如此,那时它并没有解决实际问题。正如你指出的,密码学和物理学之类的东西,要到 20 世纪才出现伽罗瓦思考“嗯,也许理解某些群如何分解,与粒子由什么构成有关系”。他基于一个纯粹的群论问题预言了夸克的存在。那是群论最有趣的应用之一:甚至预言夸克的存在都是一个群论问题。那是在拉格朗日之后很久才有的东西。所以你必须问,衡量进步的方式是什么,不是基于解决一个问题,对吧?那在某种程度上捕捉了伽罗瓦脑子里说“我觉得这里有东西”时的直觉。拉格朗日脑子里说“我觉得这是正确的思考方式”时的直觉是什么?刘维尔脑子里说“嗯,这个早已去世的年轻人散乱的笔记可能有点东西”时的直觉是什么?很难说清楚。但我正在制作的另一个系列视频是关于“压缩即智能”这个想法的。虽然这不是我真正要讲的角度,但确实有点道理:更小、更具预测性的表达感觉更智能。所以我想知道,你能在多大程度上给出某种可验证的奖励,不仅仅是围绕“你解决了吗”或“它解决了什么”,而是围绕解决问题所需概念的小巧性。
But between the time of Lagrange having this inkling that maybe symmetries of roots is the right way to go, to where it at all looks like modern group theory, you've got this long span. A lot of the time it's like not even passing the verified reward of human reviewers, right? Because it gets on someone's desk, they say 'I don't really know if there's anything here.' Gets on someone's desk, they don't. You have to have this one person sort of recognizes it. And even then, it's not really solving practical problems at that point. As you point out, cryptography and physics and things like that, you have to get into the 20th century before you have Galois thinking 'Hmm, maybe understanding the nature of how certain groups break down has this relationship with what particles are made out of.' And he anticipates quarks based on a purely group theoretic question. That's one of the more interesting applications of group theory: to even predict the existence of quarks is a group theoretic question. That's so long after Lagrange before you have anything like that. And so you have to ask, what is the way of measuring progress that's not based on solving a problem, right? And that's somehow capturing what is the instinct that's inside Galois's mind when he says 'I think there's something here.' What's the instinct that's inside Lagrange's mind when he says 'I think this is the right way to think about it.' What's the instinct inside Liouville's mind when he says 'Hmm, these scattered notes from this long dead youngster might have something to them.' So hard to put a finger on that. But I mean, a different series of videos I'm making right now is about the whole 'compression is intelligence' idea. And even though this isn't really the angle I'm taking, there is something to the idea that the smaller expression that's more predictive feels more intelligent. And so I wondered the extent to which you can give some kind of verifiable reward around not just 'did you solve it' or 'what is it solving', but around the smallness of the concepts required to do it.
我的意思是,回到黎曼猜想的解,如果 AI 解决了它会是什么样子?我认为可能发生的第三种方式是它直接更努力地工作,对吧?就像你可能有一个费马大定理的初等证明,洋洋洒洒几千页,但难以理解,而更清晰的看法是用椭圆曲线之类的东西。也许有一个千页的黎曼猜想证明,但没人能从中得到什么。而你真正想要的是那些思想的简洁、压缩版本,它们能促进人类理解。我不知道,柯尔莫哥洛夫复杂度,也许你可以把它纳入你试图量化“优雅”的尝试中。但我不认为这很容易,但我确实认为这是为了奖励伽罗瓦式的直觉而不是仅仅奖励“你解决了一个问题”所必须做的事情。为科学提出一个核心启发式方法非常困难。但显然人类不知何故一直在这样做,而且显然 AI 在某个时候也会做到。
I mean, going back to Riemann Hypothesis solutions, what would that look like if an AI solves it? I think a third way that it could happen is it just straight up works harder, right? In the same way that you could maybe have an elementary proof of Fermat's Last Theorem that's just spelled out over thousands of pages that would be incoherent, but the cleaner way to view it is with elliptic curves and all that. Maybe there's some thousand-page proof of the hypothesis that's like no one's really getting anything out of it. And what you actually want is what are the succinct, compressed versions of those ideas that would then lend themselves to human understanding. I don't know, Kolmogorov complexity, maybe you throw that into your attempt to quantify what you mean by elegance. But I don't think it's easy, but I do think it's something you would have to do in order to reward the Galois-like instinct rather than just rewarding 'have you solved a problem.' It's very hard to come up with a core heuristic for science. But clearly humans have been doing this somehow, and obviously AI will do it at some point.
嗯,这不仅与可验证奖励相关,而且最终目标大概是理解,即人类理解。所以即使你有一个千页的数学证明或某个宏大的新物理理论,目标也是理解,对吧?如果目标是预测性,你可以让自动化工程师去造火箭芯片之类的东西。我们不知道它们怎么工作,但我们可以星际旅行。但会有很多人想要理解。你仍然需要某种简洁函数,将“这里有一种复杂的思考方式”提炼成正确的那个,就像牛顿的万有引力定律那样。你仍然需要训练 AI 能够做到这一点,并找到压缩的表示。
Well, it's relevant also not just in terms of verified reward, but presumably the end goal is understanding, like human understanding. And so even if you do have some thousand-page proof of some math thing or some grand new physical theory, the goal is understanding, right? Maybe if the goal is predictiveness, you can just have automated engineers go off and build rocket chips or something. We're like, we have no idea how these work, but we can get between stars. But there's going to be a lot of people who want to understand. You're still going to want whatever the concision function is that distills down 'here's this complicated way of thinking' into the right one, like the equivalent of the universal law of gravitation for Newton. You would still want to train AIs to be able to do that and find the compressed representation.
我在印度长到 8 岁,所以除了英语,我还会说古吉拉特语。既然 Google 刚刚发布了 Gemini 3.5 Live Translate,我想在这个中插里测试一下会很有趣。3.5 Live Translate 能自动检测超过 70 种语言,并几乎实时翻译成目标语言。在你说话时,它保持你原有的语速和格式进行实时翻译,就像现在这样。我 2024 年去过中国,当时我想,如果我能实时翻译与研究人员和街上遇到的人的对话,那次旅行会高效得多。现在我们有了这项技术。所以如果你在构建一个需要实时翻译的应用,你 100% 应该看看 Gemini 3.5 Live Translate。它现在通过 Gemini Live API 和 AI Studio 提供。访问 a.studio/live 开始使用。
I grew up in India till I was 8 and so in addition to English I also speak Gujarati. And since Google just released Gemini 3.5 Live Translate, I thought it'd be fun to put it to the test in this midroll. 3.5 Live Translate automatically detects more than 70 different languages and translates them in almost real time into the target language. Live Translate your original speed and format while speaking, just like it's doing right now. I visited China back in 2024 and I remember thinking at the time that this trip would have been so much more productive if I could have been able to live translate the conversations I'm having with researchers and random people I meet on the street. Now we have that technology. So if you're building an app that needs live translation, you should 100% check out Gemini 3.5 Live Translate. It's available now via the Gemini Live API and in AI Studio. Go to a.studio/live to get started.
所以人们特别担心数学方面,AI 会证明黎曼猜想,而我们对数学的理解并不会因此变得更好。我对此有几个问题。第一个是,这是否是你应该预料到的事情?人类在解决大问题时提出一般的自然对象和子目标等等,难道不是因为这在处理复杂重要问题时很有用吗?所以我们可以从理论上思考,这会不会是解决黎曼猜想的更简单方式,而不是提出与思考问题相关的自然抽象?然后第二,从经验上看,当 AI 今天在问题上取得进展时,我们观察到的是这样吗?当 AI 通过单位距离问题猜想提出那个反例时,你可以阅读它的思维链,对我来说它不可理解,因为我对数学一无所知,但对其他数学家来说似乎是可以理解的,它利用了已知的数学概念,证明了它们之间的关系,并且使用了自然语言,结果加速了我们对这个对象和这个猜想之间联系的理解。
So people have this worry about mathematics in particular that the AIs will prove the Riemann Hypothesis and our understanding of mathematics won't be any the better for it. I have a couple of questions about this. The first one is whether this is like a thing you should expect. Like, isn't the reason humans come up with general natural objects and subgoals and whatever when we're working on a big problem is that it's just useful when you're trying to work on a complicated important problem? And so we can just think about theoretically, would this even be a simpler way to solve the Riemann Hypothesis as opposed to just coming up with the natural abstractions that are relevant to thinking about the problem? And then two, empirically, is this what we observe when AIs do make progress on problems today? When the AI came up with that counterexample through the unit distance problem conjecture, you can just read its chain of thought and it seems it's not understandable to me because I don't know anything about mathematics, but it seems to other mathematicians it was understandable and it made use of known concepts of mathematics and proved relationships between them and all the natural language, and as a result accelerated our understanding of the connection between this object and this conjecture.
那么,从经验上看,这真的是我们应该担心的事情吗?
So is this even empirically a thing we should be worried about?
我认为这取决于解决方案的性质。如果我们分解解决黎曼假设和今年另一个大问题——一个关于原始集的厄尔多斯问题,编号 11,196——的三种可能方式,它带有那种从看似不同领域引入想法的特征。一旦你把基本想法呈现给数学家,你说:“如果我们用马尔可夫链过程,从下往上概率地证明这个东西,而不是从上往下,并用冯·曼戈尔特函数呢?”如果你对懂行的人这么说,他们大概知道怎么继续。所以我们有这种非常小的想法,形式是一个领域的专长,另一个领域的专长,在它们之间画一道闪电。这些会非常容易被人类理解,因为你只需要展示那些连接的起点和终点。如果它的特征是“造山”,你就得花更多时间去理解那座新建的山,因为它像一条新线索,而不仅仅是它们之间的闪电。如果进步的本质只是纯粹的苦干,就是超级长的东西,没有新理论,只有长链条推理,那么你就会有那个消化过程。所以我不认为有一个明确的答案。我认为这取决于解决方案会是什么样子。在造山方面,看看它是否默认非常人类可理解,就像我们看到伟大数学家提出新理论的方式,还是它是一种外星式的不同山脉,我们甚至需要重新处理我们所使用的抽象类型,这会非常有趣。
I think it depends on the nature of the solution. If we break down the three possible ways of solving the Riemann hypothesis and another big one from this year, a certain Erdos problem numbered 11,196, about primitive sets, it had that character of bringing an idea from a seemingly different field. As soon as you present the basic idea to a mathematician, you say, "What if we use a Markov chain process where we show that this thing is one from the bottom up probabilistically rather than the top down, and use the von Mangoldt function?" If you say that to someone in the know, they'd kind of know how to run with it. So we have this very small idea that has the form of expertise in one field, expertise in another, draw a little lightning bolt between them. Those are going to be very human-parsible, because all you have to do is show the start and end point of what those connections are. If the character of it is mountain building, you do have to put in a lot more time to understand that new mountain that was built, because it's like a new thread, not just a lightning bolt between them. And if the nature of the progress was just raw hustle, just a super long thing with no new theories but a long chain of reasoning, then you would have that digestion process. So I don't think there's one clear answer. I think it depends on what the solution would look like. On the mountain building side, it would be really interesting to see if it is by default very human-understandable, like the way we see new theories from great mathematicians, or if it is an alien different kind of mountain being built where we even have to reprocess the kinds of abstractions we engage with.
对。
Right.
嗯,这里最接近的例子是 ABC 猜想的尝试性证明。我们也许不该深入讨论那个,但它很可能不是一个正确的证明。基本上,这是一个日本知名数学家提出的全新思维方式,数学家们花了很长时间才理解他在说什么。它给人一种外星数学的感觉,是理论构建,而不仅仅是长链条推理。他称之为“跨宇宙几何”之类的。恐惧在于它也会那样,然后就像 ABC 猜想一样,人们花多年时间爬山,然后发现:“该死,这不对。”如果它被证明是错的,但看起来很像真的。即使它是对的,爬一座新山也需要大量努力。
Well, the closest example here would be the attempted solution of the ABC conjecture. We maybe shouldn't get into that one, but it's probably not a correct solution. Basically, it's this whole new way of thinking that an otherwise reputable mathematician in Japan came up with, and it just took mathematicians a long time to even parse what he was saying. It had the feeling of an alien bit of mathematics that's theory building, not just a long chain of reasoning. He called it interuniversal geometry or something. The fear would be that it does that, and then much like the ABC conjecture, people work for years to go up the mountain and they're like, "Dang it, this just isn't right." If it turns out to be wrong but it really looked right. Even if it was right, there's a lot of effort to hike up a new mountain.
对。如果我们陷入那种情况,David Bessis 有一篇很棒的博文叫《定理经济的衰落》,他谈到了这个。历史上,就像你说的,数学是提出这些定义和问题,然后证明关于它们的定理。定理证明的东西得到了所有功劳,但它实际上是寄生在提出定义上的。历史上,在功劳分配上这不是问题,因为如果你提出了定义,你很可能就是提出定理的人。但现在我们处于一种情况,如果有价值的工作是提出洞见,而 AI 只是自动化了后一部分,想象一个场景:AI 提出了关于世界上许多重要猜想的直接论证,然后我们只有这些证明。现在由人类或未来的 AI 来整合。我确信如果你能接触到它,会让你更容易思考这里发生了什么。有没有更深层的方式理解这个证明为何成立,从而更容易提出群论背后的想法?
Yeah. If we end up in that situation, David Bessis had a really great blog post called "The Fall of the Theorem Economy" where he talks about this. Historically, as you were saying, mathematics is coming up with these definitions and problems, and it's about proving theorems about them. The theorem proving stuff is what gets all the credit, but it's really a parasite on the coming up with the definition stuff. Historically, it's not been a problem in terms of credit allocation because if you come up with the definition, you're probably going to be the guy who comes up with the theorem. But now we're in a situation where if the valuable work is the coming up with the insight and AI just automates the latter part, imagine a scenario where AI comes up with like the abble direct arguments about a bunch of important conjectures in the world, and then we just have these proofs. Now it's up to humans or future AIs to consolidate. I'm sure if you had access to it, it would make it easier for you to then think about what is going on here. Is there some deeper way to understand how this proof works that would make it easier to come up with the ideas behind group theory?
是的,我认为这会非常有帮助。因为尝试发现新数学很大程度上是不断犯错。你试图解决一个问题,感觉不像是在不断迈出正确的登山步伐。大多数时候感觉像随机醉汉走路,你做一件事,然后错了,不断发现。所以至少如果你知道消化你所知道的东西最终会导向一个正确的解,那感觉就是进步,仅仅因为它提供了知道它导向解的感觉。在最近的数学史中有很多例子,感觉触及超过了掌握,有些东西在理解之前很久就被证明了。我最喜欢的论文开头之一,甚至不是研究论文,而是说明性论文,来自一位名叫 Timothy Chow 的数学家,他试图理解一个叫“力迫”的概念。有一个问题叫连续统假设,大致问的是:自然数有一个无穷大,实数有一个无穷大,中间有没有东西?答案既是肯定的也是否定的。这取决于你的公理。它有点超出我们通常公理系统的范围,这是一个有趣的答案,但描述它的方法真的很难理解。它叫力迫。在这篇论文的开头,他写道:“我想提出一个未解决的说明性问题的想法,我们确实证明了它,但我们并不真正知道为什么它是真的。”突然他提出了那个说明性问题的部分解。你可以想象为什么我喜欢这个框架,因为这就是我的全部生活。我不做研究数学。它完全关于什么是最清晰的理解方式,即使它已经被证明。证明和解释之间是有区别的。
Yeah, I think it would be hugely helpful. So much of trying to discover new math is mostly being wrong. You're trying to solve a problem, it doesn't feel like constantly taking the correct step up the mountain. Mostly it feels like a random drunken walk where you're doing a thing and then you're wrong and constantly discovering. So if at the very least you know that trying to digest what you know is ultimately leading to a correct solution, that feels like progress simply because it provides a sense of knowing that it leads to a solution. There are plenty of instances in the recent history of math where it feels like the reach has sort of exceeded the grasp, where there are things that are proven long before they're understood. One of my favorite openings to a paper, not even a research paper, more an expository one, is from a mathematician named Timothy Chow who was trying to understand a concept called forcing. There's this problem called the continuum hypothesis that more or less asks: you have a size of infinity for the natural numbers, you have a size of infinity for the real numbers, is there something in between? The answer is both yes and no. It depends on your axioms. It's sort of outside the scope of our usual axiom systems, which is an interesting answer, but the method to describe it is just really hard to understand. It's called forcing. In the beginning of this paper, he writes: "I want to propose the idea of an unsolved expository problem, where sure we've proven it but we don't really know why it's true." Suddenly he proposes a partial solution to that expository problem. You can imagine why I loved that framing, because this is my whole life. I don't do research math. It's just wholly about what's the most clear way to understand this, even if it's proven. There is a difference between proof and explanation.
所以从这个角度来说,我认为你基本上触及了那个区别的重要性。
And so on that side I think that you are basically getting to the importance of that distinction.
是的。这将是主要的动机——或者说动机必须改变,不仅在数学领域,也在其他科学领域——从证明关于世界的事实,转向将证明整合成问题或更高层次的洞见。但我们午餐时讨论过你最近的一个演讲,关于设计如何帮助我们理解事物,然后极限情况下,概念化一个想法和想法本身之间真的有区别吗?所以如果你想想狭义相对论和时空图以及闵可夫斯基时空,这是一种我们用来阐释为什么存在长度收缩和时间膨胀的方式,但这是现实吗?所以阐述似乎在某种意义上就是解释本身。
Yeah. And that will be the main incentive for—or the incentive would have to change in not just mathematics but in other areas of science—from proving things about the world to consolidating proofs into problems or higher level insights. But we had a discussion earlier at lunch about a recent talk you were giving about design and how it helps us understand things, and then in the limit, is there really a difference between the conceptualization of an idea and the idea itself? So if you think about special relativity and spacetime diagrams and Minkowski spacetime, it's like a way in which we illustrate this idea of why there's length contraction and time dilation, but is that the reality? So the exposition does seem to be the explanation in some sense here.
是的。那里有几件有趣的事。第一,似乎那些提出真正新颖洞见的人,与那些在沟通中非常清晰的人之间存在很强的相关性。你可能会想,鉴于大学生的经历往往是教他们的专家不一定是该主题的最佳解释者,因为他们被自己的专业知识宠坏了。但至少在某些情况下,似乎那些真正提出新颖东西的人——比如爱因斯坦或克劳德·香农——你读他们的论文,它们真的非常清晰,对吧?你不会觉得“哦,这只是给专家的,你得用砍刀劈开它”。他们是非常好的阐述者,费曼也有这个特点,非常好的阐述者。所以也许大脑中在研究层面提出正确新思维方式的同一部分,也拥有这种善于解释的诀窍。我认为这与 AI 相关,我以前认为 AI 会成为自动定理证明器,而数学家的角色将转向像我的工作——解释这些东西。我有点怀疑实际上它们也会非常擅长这一点,并且可能比大多数人类更擅长解释和提炼的部分,而这实际上并不是留给数学家的——更像是消化和解释发生了什么。可能这就是事情发展的本质。我可以想象——我们可以讨论可能不是这样的方式——但很可能提出真正好的新想法来解决新问题的同一个东西,也擅长解释它。这是我的信念改变的方式。
Yeah. I mean there are a couple interesting things there. One is it seems like there's a really strong correlation between the people who come up with genuinely novel insights and also who are actually quite clear in their communication of it. Like you might imagine, given that the experience of a university student is often that the expert teaching them is not necessarily the best explainer of that topic because they are so spoiled by their expertise. But what seems at least in some cases to be the case is that the people who are really coming up with something quite novel—so you've got like Einstein or Claude Shannon or something—you read their papers, they're really lucid papers, right? It doesn't feel like 'oh this is just for the experts and you have to chop through it with a machete.' They're very good expositors, like Feynman has this characteristic too, very good expositor. And so maybe the same part of the brain that comes up with the correct new way of thinking about it at a research level also has this knack for good explanation. And I think this is pertinent to the AI one where I kind of used to think that AIs will become these automated theorem provers, but the role of the mathematicians is going to shift towards like my job—explain these things. I kind of suspect that actually they'll also be quite good at doing that and probably just better than most humans are at doing the explanation half and distilling half, and that's actually not what's left for the mathematicians—it's like digesting and explaining what was going on. Probably the nature of how these things are going. I could envision—we can talk about ways this might not be it—but probably the same thing that is coming up with the really good new idea that solves some new problem is just also good at explaining it. That's a way my beliefs have changed.
你认为你最后会做什么——或者你和人类数学界会做什么?
What's the last thing you think you will be doing—or both you and also what the mathematical community, the human mathematical community, will be doing?
我可能会一直做我现在做的事情直到我死。即便如此——
I will probably be doing something like what I am until I die. Even so—
如果末日论者是对的,也许那完全一样——出于同样的原因。
And if the doomers are right, maybe that'll be the same exactly—it'll be for the same reason.
是的。是的。你知道,就像“给一个人生火,他暖和一夜;但点燃一个人,他暖和一辈子。”所以这就是我对 AI 的看法。不,因为解释者或教师的部分功能是为某人好奇的事物增加清晰度。这是一回事。但另一部分则更具关系性,更像是提供动力、提供策展感。我听过一个有趣的观点,认为数学家最终会变得更像艺术博物馆的策展人,而不是其他什么。AI 甚至知道如何很好地解释,但你还是希望有人帮你在这个近乎无限的空间中导航,哪些想法值得投入。即使 AI 在某种意义上更擅长这一点,我认为我们仍然会更喜欢一个我们与之有关系的真人,因为我们被激励去对事物感兴趣的方式是一种社会现象。如果你有特定的技术要构建,那可能不同。但我认为收听这个播客的人,他们首先信任你对什么是有趣话题的策展。他们并不是因为你的下一个话题正是他们事先想理解的而来到这里。他们信任你作为策展人。所以我的角色,以及可以说其他数学家的角色,可能实际上会微妙地转向那个策展方向,即哪些想法值得展示。而这正是我现在工作的很大一部分,即使现在也是如此——我认为人们常常认为制作视频的大部分时间花在视觉效果上。当然,有一点。它不是立竿见影的,但实际上很多工作只是决定首先说什么值得说,或者什么值得放在那里。因为那正是——我想参与其中,而且我认为我与某些人建立了信任,他们好奇我会选择提出什么,即使 AI 在这方面更好,就像人类音乐家总是会有一个角色,因为他们背后的故事具有社会功能,即使某个模型输出的 MP3 文件的客观质量更好。这就是我看到的我的工作未来的样子。
Yeah. Yeah. You know, it's like 'build a man a fire and he's warm for one night, but set a man on fire and he's warm for the rest of his life.' So that's where I am with AI. No, because some of the function of an explainer or a teacher is to add clarity to a thing that someone's curious about. That's one thing. But some of it is a little bit more relational and a little bit more like providing motivation, providing a sense of curation. Like one interesting take that I've heard about what mathematicians will end up being is actually more analogous to art museum curators than anything else, where the AIs—they even know how to explain it really well, you know, out there. But you still want someone to help you navigate in this nearly infinite space of what ideas are worth engaging with. Someone kind of doing that, and even if AIs were in some sense better at that, I think we would always still prefer a human that we had a relationship with, because the way that we get motivated to be interested in things is a social phenomenon. If you have some specific technology you're trying to build, you know that might be different. But I think the people listening to this podcast, they sort of trust your curation on what's an interesting topic in the first place. It's not that they're landing on here because whatever your next topic is, that's what they in a prior sense wanted to understand. They're trusting you as a curator. So my role, and arguably that of other mathematicians, might actually just shift subtly into that curation direction of what ideas are worth displaying. And that's a lot of my job right now, even now, is basically—I think people think a lot of the time for a video goes into the visuals. Like sure, a little it is. It's not like immediate, but actually a lot of it is just deciding what's worth saying in the first place or what's worth putting there. And because that is just—I want to engage with that and I think I have a trust with certain people and they are curious what I would choose to put forward, even if the AIs are better than that, in the same way that human musicians are always going to have a role because of that social function of the story behind them, even if the objective quality of the MP3 file coming out is better from some model. That's kind of what I see happening to my job.
是的。我想回到之前的问题——就在 AI 跨越了这个门槛,这个重要的基准,能够连接现有想法以提出新发现或证明或反驳某些东西。就在它跨越这个门槛时,我们想“好吧,但下一步是什么?”我想——
Yeah. I want to go back to this question earlier—we were sort of just as AI has crossed this threshold, this important benchmark of being able to connect existing ideas to come up with a new discovery or prove or disprove something. Just as it's crossed this threshold, we're like 'okay, but what's the next thing?' I want to just—
这方面还有很多工作要做。仅仅因为几道闪电已经——我仍然认为未来几年会有一个蓬勃发展的连接时期。所以在极限情况下,你甚至可以说——我不知道这样说是否准确——但可能很多最大的突破在某种程度上看起来就是这样。就像广义相对论。哦,你只是把黎曼几何和狭义相对论连接起来,对吧?所以随着 AI 在这个连接事情上变得越来越好,也许很多重大突破在本质上并不是不同的。我不知道你对此有什么看法。
There's a lot more to do on that one. Just because a couple lightning bolts have been—I still think there's this flourishing future over the next couple years of really connecting. And so in the limit, you could even say—I don't know if this is accurate to say it—but potentially a lot of the biggest breakthroughs look like this at some level. It's just general relativity. Oh, you just connect together Riemannian geometry and special relativity, right? And so as AI keep getting better and better at this connection thing, maybe a lot of big breakthroughs are not really of a different qualitative nature. I don't know if you have a take on that.
嗯,我的意思是,很多讨论都集中在解题和数学的那种性质上,比如解决一些新鲜的问题。我得说,大多数数学家可能并不会把自己的工作描述成专注于攻克下一个难题。你了解朗兰兹纲领吗?这与其说是一个数学领域,不如说是一种研究精神。费马大定理就是其中一个体现:你有两个看似毫不相干的东西,它们之间的联系导致了一个解法。朗兰兹是一位数学家,他有一封著名的信,基本上阐述了很可能存在更多这样的联系,并且还更具体地说明了这些联系的性质,以至于你可以想象一张大地图:这边有一个山谷,那边有一座山,还有一片平原。很多数学家把自己的工作描述成试图理解这张地图上的线索,以及其中的进展。这不像“这里有一个具体问题,我们知道可以通过那个联系解决”。更多的是,已经有足够多的案例表明,大问题是通过发现联系而被攻克的,以至于这几乎成了预先发现联系。所以,任何时候你遇到一位数学家,问问他们自己的工作更像是朗兰兹纲领,还是更像针对某个特定问题,你会得到一个明确的分化。但 AI 成为超级连接器的可能性,感觉可能会是这项追求中的一个放大工具。不过这很难衡量,因为这又回到了我们之前说的:你怎么给“你做到了”打分?如果是攻克一个问题,你有一个明确的方式说“你做到了”。你可以写头条新闻,作为 AI 公司你可以做公关宣传说“我们做到了”。而如果感觉那是画对了联系,你可以围绕它写定理,这就是那个领域论文的样子。但我认为这将需要更多的人参与其中,来基本判断“我们追求的是哪种联系”。我猜测,未来五年这些模型最有用的进展,就是真正填补那个如果你是多领域专家就能画出的联系图景。就像你指出的,我们居然还没做到这一点,这有点令人惊讶,对吧?
Well, I mean, a lot of the conversation focus has been on problem solving and that nature of math, like taking on fresh problems. I would say it's not even a majority of mathematicians who would characterize their work as really targeting the next problem to dig down. Are you familiar with the Langlands program? So this is not even a field of math so much as it is a research ethos where the last theorem is one inkling of this: you had these two different seemingly disparate things and a connection between them led to a solution. Langlands was a mathematician. He has this famous letter essentially spelling out how it seems likely that there are a lot more connections like that, and even got a little more specific about the nature of the connections such that you might imagine this large map: you've got this valley over here, this mountain over here, and this set of planes over there. There are a lot of mathematicians who would characterize their work as being part of trying to understand the threads on this map, and the progress there. It's not like here's this one specific problem that we know will be solved by that connection. It's more that there have been enough cases where big problems were knocked down by finding connections that it's almost preemptively finding the connections. So anytime you run into a mathematician, ask them whether the character of their work is more akin to Langlands program or more akin to targeting one particular problem, and you get a certain bifurcated split there. But the possibility of AIs being supercharged connectors feels like it might be an amplifying tool in that pursuit. It's hard to measure though, because this cuts to what we were saying earlier: how do you assign a score to say yes, you've done it? If it's knocking down a problem, you have a clear way of saying yes, you've done it. You can write the headline, you can have your PR move as the AI company to say we did it. Whereas if it feels like that was the right connection drawn, you can write theorems around it, and this is the nature of what the papers in that field look like. But I think it will require a lot more human-in-the-loop to basically say what was the kind of connection that we're going for. That's my guess on what most of the useful progress from these models will look like in the next five years: just really filling in that landscape of connections that you can draw if you're an expert in multiple fields. As you've pointed out, it's kind of surprising we haven't already had this, right?
我很好奇的是,在技术层面上是什么导致了这种解锁。因为一方面,你可以在脑子里描绘一个解释,说明为什么你可能是所有这些领域的专家却没有画出那些联系:当推理发生时,推理的方法是这种自回归的思维链现象。自回归实际上是一种非常奇怪的生成方式。想象你是一个聪明人。我把你锁在一个盒子里。你与外界互动的唯一方式是你收到一张纸条,有人说:“你能预测接下来会发生什么吗?”然后你预测下一个是什么,然后你的记忆被清空。然后你又得到另一张纸条,你继续……想象这个过程重复了很多次。然后另一端出来的是什么?他们会说:“看看你写的这篇文章。”你可能会看着它说:“这太糟糕了。这不是我会写的文章。”因为反复预测某件事的过程,与你作为作家构思和深思熟虑的方式截然不同。特别是,很可能发生的是你成了上下文的奴隶。你可能在回答某个特定领域的问题,所以你调用了所有相关的上下文,然后朝着那个方向走。而真正产生实质内容的联系,本质上是非常不可能的。你可以做所有你想要的强化学习来试图在某些方面变得更好,但到底是什么在专门提升和激励做出这些不太可能的联系,而绝大多数联系都不是可预测的下一个词?所以情况可能是,你拥有这种被锁在盒子里的智能,但与之互动的方式很奇怪。所以我好奇的是:你是否通过质疑 token 生成的前提而获得过任何成果?我不认为这就像调整温度那么简单,但有没有什么方法可以利用现有的智能水平,找到正确的方式来激发那些联系,从而解锁我们见过的这些东西?还是说你需要再多一点智能,使得在预测的层面上,它能够预测自己应该做出那种跨领域的闪电连接?
What I'd be curious to know at a technical level is what causes the unlock there. Because on the one end, you can kind of paint an explanation in your head for why you could be an expert in all of these things and not be drawing those connections: when the thing is reasoning, the method of reasoning is this autoregressive chain of thought phenomenon. Autoregression is actually a really weird way to produce stuff. Imagine you're an intelligent person. I've locked you in a box. The only way you have of interacting with the world is that you receive a slip of paper and someone says, "Can you predict what will come next?" Then you predict what will come next, and then your memory is wiped. Then you get another slip of paper, and you go, "..." Imagine that was done a whole bunch. Then what comes out on the other end? They're like, "Look at this essay that you wrote." You might look at that and be like, "This is awful. That's not the essay that I would have written." Because the process of repeatedly predicting something is just pretty different from how you would think as a writer to compose it and think it through. In particular, what would probably happen is you're sort of a slave to your context. You might be answering some question about some particular field, so you draw in all the context around that and you go there. The connection that actually is where all the substance is going to come from is by its nature a very unlikely one. You can do all the RL that you want to try to get better in some way, but what's the thing that's specifically upweighting and incentivizing making these unlikely connections when the vast majority of them aren't the predictable next token that would come in there? So it might be the case that you have this intelligence that's sort of locked in there inside that box, but it's just a weird way of interacting with it. So the thing I'm curious about is: do you ever get any fruit by just questioning the premise of how tokens are generated every now and then in some way? I don't think it would be as simple as manipulating the temperature or something like that, but are there any things that you can do that take the existing level of intelligence but find the right ways of sparking those connections that unlock these sorts of things that we've seen? Or do you need just a little bit more intelligence such that at the level of prediction it's kind of predicting that it should be making that lightning bolt to another field?
我认为更有成效的是从数据角度推理,而不是架构甚至损失函数。我们有做文本的扩散模型,它们生成的东西并没有完全不同的特征;只是它们还没有被充分探索。我认为更相关的是,无论你采用什么架构、什么损失函数,数据在激励你产生什么。而且看起来它们确实在变得更好,先别管数学。我的意思是,我们确实有过几个这样的例子。但如果你只看为什么它们在成为自主智能体方面变得更好?我只是不知道,它们好像处于一个环境中,自回归地生成“让我们退一步,在整个代码库中搜索”这样的步骤,然后“让我们退一步,评估我的错误”是有效的。
I think it's more productive to reason about data, rather than architecture or even loss function. We have diffusion models that do text, and the kinds of things they produce are not of a whole different character; they just haven't been explored as much. I think the more relevant thing is what is the data on which whatever architecture, whatever loss function you have, is incentivizing you to produce. And it does seem like they're getting better at, forget about math. I mean we did have a couple of examples of this kind of thing. But if you just look at why are they getting better at being autonomous agents? It just, I don't know, they have like they're in an environment where autoregressively producing the step that says "let's step back and do a search over the whole codebase," and then "let's step back and assess my mistake" is the thing that works.
我猜想,在科学或数学进步的情况下,你会遇到前沿数学问题,这些问题需要数学家专门设计,因为它们需要连接两个不同的领域。有各种巧妙、部分合成的方法来制造越来越难的问题,需要这种连接,例如通过消除假设,仍然要求 AI 继续得出答案。损失函数是什么其实并不重要。真正的问题在于,你是否能创造一个激励这种能力的环境。
I assume what happened in the case of progress in science or maybe in math is you have frontier math-like problems which require mathematicians specifically designed them because they require connecting together two different fields. There's all kinds of clever, partially synthetic ways to make harder and harder problems that require these kinds of connections, for example by eliminating assumptions and still requiring the AI to continue to get to the answer. It doesn't really end up mattering what the loss function is. It's really about can you come up with an environment which incentivizes this ability.
是的,感觉你应该能做到。
Yeah, it feels like you should be able to.
是的,我无法说出解锁这一切的正确方法,但如果未来三年内没有出现更多这样的“闪电”,那会相当令人惊讶。所以这是一个值得思考的重要问题:我们通常考虑单个系统有多聪明,却没有考虑 AI 拥有源于它们其他事实的优势。在这种情况下,关键事实是我们可以并行化并任意扩展它们。所以无论它们的能力水平如何,这不仅仅是数学史上一个古怪的天才做出一些联系然后死于决斗。而是将那个水平线普遍应用于所有在该能力水平下可及的问题。我觉得这是数字思维天生拥有的众多优势之一,我们对此思考不足。其他优势包括它们可以融合所有知识——至少会有技术实现这一点——以及你可以生成具有相同知识水平的副本。但这种并行化是一个非常重要的特性。我很好奇你的预测:即使它们不如人类数学家聪明,但它们是数十亿个——因为出于公关原因,AI 公司正在投入数十亿美元——数量本身就有其质量。
Yeah, I can't speak to the correct ways of doing that that unlock all this, but it would be pretty surprising if over the next 3 years there's not just a lot more of those lightning bolts. So this is an important thing to think about: we often think about how smart a single system is, and we don't think about AI having advantages that are more the result of other facts about them. In this context, the key fact is that we can parallelize and arbitrarily scale them. So whatever level of capability they have, it's not just one idiosyncratic genius in the history of mathematics who makes a few connections and then dies in a duel. It's universally applying that waterline across all problems that are accessible at that level of capability. I feel like this is among the many advantages that digital minds inherently have that we don't think enough about. The other ones include the fact that they can merge all the knowledge together, at least that there will be techniques that allow this to happen, and that you can spawn off copies with identical levels of knowledge. But this parallelization is quite an important property. I'd be curious about your predictions: even if they're not as smart as human mathematicians, the fact that they are billions of—because for PR reasons the AI companies are dumping billions and billions of dollars at this—quantity has a quality all of its own.
这似乎方向正确。我认为,如果我们以蒙哥马利和戴森在 IAS 的对话为例,那暗示了黎曼假设或黎曼 ζ 函数零点与随机矩阵之间的联系,这感觉像是你可以尝试自动化的那种事情。你有代表所有这些领域专业知识的智能体,基本上——好吧,我们都知道一个研究所比一个人聪明,而让人们都在同一地理位置的原因是你想要那些偶然的对话发生。那么,在智能体之间设计这种对话会是什么样子?这很有趣,因为你指出你可以某种程度上汇聚所有知识。所以我真的想知道,其中一个优势是否在于你可以做相反的事情:有时当 AI 失败时,是因为它陷入了一个糟糕的思维链,很难摆脱出来。所以你就重新开始。人类也一样:有时你以某种方式思考,而需要做的就是退一步。也许有时那种形式——有故事说人们长时间试图证明某事,然后某个时刻他们说,“等等,如果我试图证明它是不可能的,证明相反的情况呢?”然后解开自己的上下文,以全新的思维重新开始。你可以想象系统化这一点,或者让多个不同的智能体被故意赋予不同的上下文,然后尝试比较和对比。我们对自己的上下文没有同样程度的操控能力。
That seems in the right direction. I think if we take that conversation between Montgomery and Dyson at the IAS that suggests some connection between Riemann hypothesis or Riemann zeta function zeros and random matrices, that feels like the kind of thing you could try to automate. You have agents representing expertise in all these, and basically having—okay, we all know that an institute is smarter than an individual, and the reason for having people all in the same geographic location is because you want those serendipitous conversations to happen. What does it look like to sort of engineer those between agents? It's interesting because you point out you can sort of pull all your knowledge. So I actually really wonder if one of the advantages is that you can do the opposite of that, where sometimes when an AI is failing, it's because it gets into a bad chain of thought and it's really hard to get it out of it. So you just start again. Same deal with humans: sometimes you start thinking about it in a certain way, and what's required is to just back up. Maybe sometimes the form of that—there are stories about people trying to prove something for a long time, and then at some point they say, 'Hang on a second, what if I tried to prove that it's impossible, prove the opposite?' and that unwinding your own context and going at it with a fresh mind. You could imagine systematizing that, or having multiple different agents deliberately given different pieces of context and trying to compare and contrast. We don't have the same level of manipulation on our own context.
在这个 AI 与数学系列中,第一集将关于他们解决 IMO 问题的时候,我想聚焦于一个他们失败的具体 IMO 问题,这个问题很多非常聪明的学生也失败了。陶哲轩也失败了。它的本质基本上是人们对这个问题非常生气,因为他们称之为“恶搞题”。我几乎不想剧透,因为我想围绕引导某人进入情境来构建这一集,让他们不知道结果会有一个简单的解法,因为你可以真正共情一个学生解决这个问题时的感受。基本上,有一种非常优雅的方式,基于作为国际数学奥林匹克竞赛问题的背景,你会觉得那将是解法。这个解法的特点非常诱人,但很难证明它是最好的。原因在于它不是。有一个几乎“脑死亡”的解法才是最好的。所以这与整个 AI 故事的相关性在于,对于人类来说,回答那个问题需要的是逃离你的上下文。逃离你身处 IMO 的上下文。逃离你被训练来解决这些竞赛数学问题的方式。如果你只是把它当作一个扔给街边路人的脑筋急转弯,他们很可能回答得很好。有时在其他背景下的人类研究中也希望如此:有时只是能够说,刷新你的思维,以完全不同的方式来处理。所以在数字思维拥有的所有优势中,这实际上可能是其中之一:更系统化地,刷新你的思维是什么样子,尝试回答两个独立的问题,分出两个智能体,一个试图证明它,一个试图反驳它,一个这样尝试,一个那样尝试,它们故意拥有不同的上下文。我很好奇,如果我们三年后再次进行这次对话,有多少重大成果具有那种基本上擦除先前上下文、尝试一堆不同东西的特征,而不是融合一堆不同结果的特征。
In this AI and math series, the first episode will be about when they solved the IMO, and I want to focus on one specific IMO problem that they failed on, which is one that a lot of very smart students failed on. Terry Tao also failed on it. The nature of it is basically that people were very mad at the problem because they called it a troll problem. I almost don't want to spoil it because I want to construct the episode around leading someone in without knowing that it turns out to have a simple solution, because you can really empathize with what it's like to be a student solving this. Basically, there's a really elegant way of going down what you really feel is going to be the solution based on the context of being an International Math Olympiad problem positioned as it is. The character of the solution is really enticing, but it's kind of hard to prove that it's the best. The reason is that it's not. There's this almost brain-dead solution that is the best. So the relevance of that to the whole AI story is that for a human, what's required to answer that question is to escape your context. Escape the context that you're in the IMO. Escape the context of the way you've been trained to solve these contest math problems. If you just approached it like a brain teaser that I throw to someone off the street, they'd probably answer it well. You sort of want the same sometimes for human research in other contexts, where sometimes just being able to say, refresh your thinking, come at it completely differently. So of all the advantages that digital minds have, that might actually be one of them: a little more systematic, what does it look like to refresh your thinking, try answering two separate questions, spin off two agents, one who's trying to prove it, one who's trying to disprove it, one who tries it this way, one who tries it that way, and they deliberately have different contexts. I would be curious to see if we're having this conversation three years from now, how many of the significant results that make headlines have that character of basically erasing the context previously, trying a bunch of different things, as opposed to merging the results of a bunch of different...
这非常有趣,因为人们对 AI 的一个常见担忧是这种熵坍缩,即它们都以相同的方式思考,因为它们的训练方式相似。这就是为什么它们不擅长写作。它们基本上只是走同一条路,有相似的说话模式等等。
This is incredibly interesting because a common concern people have about AIs is this entropy collapse where they all think the same way because they're trained in similar ways. This is why they're bad at writing. They kind of just go down the same path and have similar patterns of speaking and so forth.
但也许 AI 真正的关键优势在于,你可以系统地——听起来单位距离问题猜想之所以花了那么久才被证伪,部分原因是人们假设该猜想是正确的。所以他们大多在想办法证明它。因此,AI 的一个关键优势或许是,通过系统地尝试否定某个命题并尝试证明其正面,或者系统地给不同智能体赋予不同的偏见,来增加熵。
But maybe actually the key advantage AIs have is that you can systematically — it sounded like one of the reasons the unit distance problem conjecture took so long to be disproven was because people assumed the conjecture was actually true. So they were mostly trying to figure out ways to prove it. And so maybe one of the key advantages the AIs will have is actually to increase the entropy by systematically trying out both the negation and trying to prove the positive of any given statement, or being able to systematically give different agents different biases.
说得好。人类科学史上很重要的一点是,爱因斯坦正是被“事物在不同参考系中应看起来相同”这一偏见所驱动,他还有其他类似的偏见,但这些对他的思想形成了关键影响。而你可以系统地审视一系列启发式方法,看看哪些在解决特定问题时富有成效。
That's a good point. It seems like an important thing in the history of human science is that Einstein is just really motivated by this bias that things should look the same in different reference frames, and then he had multiple other biases like these, but that is just very formative in his thinking. And you can just systematically survey a bunch of heuristics and see which ones are being productive at a given problem.
是的。所以你的建议基本上是,在提示层面系统地增加熵,尽管在自回归层面不可避免地会坍缩。
Yeah. And so you would suggest basically like systematically increasing entropy at the prompt level, even though you have this inevitable collapse at the autoregression level.
是的。爱因斯坦是个有趣的例子,因为他有“事物应是相对的”这一偏见,他还有“上帝不掷骰子”的偏见,对吧?你几乎要确保不会不小心让所有大语言模型都变成爱因斯坦,因为那样可能会阻碍量子力学的进展,对吧?这恰恰说明,科学中不存在唯一正确的启发式方法。你实际上需要多个独立的研究项目,各自拥有自己的启发式方法。
Yeah. And I mean Einstein would be an interesting example because he's got this bias towards things should be relative. He also has a bias towards like God should not play dice, right? And it's almost like you want to make sure that you don't accidentally have all of your LLMs be Einstein, because you might halt on quantum mechanics progress, right? Which actually goes to show you that there is not a correct heuristic for science. You actually just need multiple independent research programs with their own heuristics.
完全正确。这感觉像是老派软件,对吧?只要你能以某种方式描述它,你就有老派软件以某种方式放大那种熵。如果你能为想要提示的不同思维方式建立一个清晰的分类体系,你探索整个分类体系,然后每个个体各自运行。但我认为这里有一个设计问题:如何准确描述不同的方法。简单的方法是:你是想证明它还是证伪它?更难的是,要列出所有可能用来证明它的策略,并确保你在探索时具有足够的广度。
Exactly. Yeah. And that feels like old school software, right? As long as you're able to describe that in some way, you have old school software that amplifies that entropy in some way. And if you're able to put a clear ontology to the distinct ways of thinking that you want to prompt, you explore that full ontology and then each individual one runs off doing what it is. But I think there's a certain design question there on how exactly you describe the different approaches. The easy one is: are you trying to prove it or disprove it? The harder one would be to say what are all the tactics that you could take to prove this and make sure that you're applying sufficient breadth to exploring that.
我觉得人们没有充分认识到,当你给这些模型配备像 Cursor 这样的好工具时,它们能为你处理多少事情。例如,我开始在 Bilibili 上发布我的节目,希望能吸引越来越多的中国观众,但上传的所有内容都需要剪掉赞助片段。通常这意味着我得让我的编辑回去翻看所有旧节目,剪掉广告,然后重新导出。但只需花我发那条 Slack 消息的时间,我就能直接告诉 Cursor 去做,省得麻烦他们。至于播客的研究,我建立了一个完整的仓库,里面放了所有与近期节目准备相关的书籍和论文。我能够把一切整合起来,因为 Cursor 工具非常擅长帮助模型准确找出需要提取的信息,无论是来自我的仓库还是来自网络,以便回答我在研究时遇到的问题。所以,无论你现在在做什么,试试用 Cursor 来处理。访问 cursor.com/locash 开始吧。
I don't think people appreciate the kinds of things that these models can just go handle for you when you equip them with a good harness like Cursor. For example, I started publishing my episodes on Bilibili for a hopefully burgeoning Chinese audience, but everything I upload there needs the sponsored segments cut out. Normally, that would have meant that I would have to ask my editors to go back through all the old episodes, cut out the ads, and re-export everything. But in about as much time as it would have taken me to send them that Slack message, I can just tell Cursor to do it instead and spare them. And for research for the podcast, I have a whole repo that I've set up where I've just put every single book and paper that's been relevant to prepping for any of the recent episodes. And I've been able to hodgepodge everything because the Cursor harness is just extremely good at helping the model figure out exactly what information to pull, whether that's from my repo or from the web, in order to answer the questions I have while I'm doing research. So, whatever you happen to be working on right now, just try pointing Cursor at it. Go to cursor.com/locash to get started.
显然,AI 在数学领域的进展比其他领域快得多,人们指出领域的可验证性是关键原因。我认为这是两个重要原因之一,但我觉得人们确实忽略了另一个原因。我不在实验室内部,不知道实际情况,但这完全是一个天真的理论。好,一个与“为什么 AI 在数学上进展如此之快”相关的问题:为什么在计算机使用方面进展如此缓慢?
Obviously, AI for math is making a lot faster progress than everything else, and people point to verifiability of the domain as the key reason this is happening. I think that's one of the two important reasons, but I don't think people really neglect the other one. And I'm outside the labs. I don't know what's actually going on, but this is a totally naive theory. Okay, a tangential question to why AI is making so much progress in math. Why has it been so slow at computer use?
计算机使用实际上非常可验证——比如,我的 Etsy 包裹到了吗?或者我的活动预订成功了吗?这些都是极其可验证的事情。计算机使用缺乏的是可重复性。因为网站有机器人检测器,而且运行并行 rollout 需要巨大的算力。很难在亚马逊上对同一个结账流程运行一千次并行 rollout,因为你会被 Andy Jassy 封杀,对吧?所以你可以尝试构建每个网站的克隆。这非常耗费人力,而且拖慢速度。顺便说一句,目前用深度学习来学习一项技能需要这么多并行 rollout 的原因,是我们还没有解决样本效率问题。你就像用吸管吸取监督信号。当然,人们正在研究许多不同的技术,但根本上,我们训练 AI 的方式存在这个大问题和大限制。对于代码,你可以将某个进展水平容器化到一个仓库中,然后启动数千个并行容器,说“尝试实现这个功能”,这完全是确定性的。因为它是确定性的,你可以解决信用分配问题,因为你知道,导致这个 rollout 成功而另一个失败的原因,差异就是那个起作用的改动。这样你就解决了信用分配问题。如果情况从不同的起点开始,信用分配问题就变得难得多。但现实世界中的大多数事情都很难用同样的方式容器化。编码和数学是例外。但如果你只是想弄清楚如何建立一个成功的新业务,如何在市场上交易一天并赚钱?你做不到——你必须与现实世界互动,事情每天都在变化,这意味着你不能反复重放、重复和模拟。但数学当然是例外。我觉得这实际上是该领域以及编码领域进展的重要驱动力。不仅仅是可验证性,还必须具有可重复性。人们指出的 AI 进展迅速的第三个原因,是他们非常关注 Lean 和形式化。
Computer use is actually very verifiable — like, is my Etsy package coming? Or is my event booked? These are extremely verifiable things. What computer use lacks is grindability. Because websites have bot detectors, and also it takes a tremendous amount of compute to run parallel rollouts. It's very hard to just run a thousand parallel rollouts at the same checkout flow on Amazon, because you'll get shut down by Andy Jassy, right? And so you could try to build clones of every single website. This is very labor intensive and slows you down. And the reason, by the way, you need to do so many parallel rollouts in order to learn a skill currently with deep learning is that we haven't solved sample efficiency. You're sucking supervision through a straw. Of course people are working on many different techniques, but fundamentally there's this big problem and this big constraint in the way we train AI. With code, you can containerize a given level of progress in a repository and then spin out thousands of parallel containers and say, "Try to implement this feature," and it's totally deterministic. Because it's deterministic, you can solve the credit assignment problem because you know that whatever caused this rollout to succeed and this one to fail, the diff is the thing that worked. And this way you solve the credit assignment problem. If you have situations that are starting off at different starting points, this credit assignment problem becomes much harder to solve. But most things in the real world are just very hard to containerize in the same way. Coding and math are exceptions to this rule. But if you're just trying to figure out how do I build a new business that succeeds, how do I go trade in the markets for a day and make money? You can't — the fact that you had to interact with the real world and things change day after day means that you can't keep replaying and grinding and farming the simulator. But math, of course, is the exception. And I feel like this is actually an important driver of progress in this domain and also in coding. It's not just verifiability. It has to be grindable. The third reason that people point out that AI is making fast progress is they focus a lot on Lean and formalization.
再说一次,我真的完全不知道实验室里在发生什么。我觉得 Lean 对于当前 AI 的进展来说并不那么重要,或者说,为什么 AI 能解决单位距离问题?他们推翻了单位距离猜想的反例。他们发布的思维链,或者说思维链的重写版本,里面根本没有用到 Lean。我认为 Lean 提供的基于过程的监督——每一步都正确——似乎不如拥有一个可验证的、可反复打磨的结果来得重要。
Again, I have literally no idea what's going on in the lab. I feel like Lean just doesn't matter that much for the current level of progress in AI, or like why is AI able to solve the unit distance problem? Well, they disproved the conjecture by the unit distance problem. They released the chain of thought, or at least a rewrite of the chain of thought didn't have any Lean in it. I think it just like the process-based supervision that Lean provides, where you know each step is correct, seems less relevant than just having this grindable outcome that is verifiable.
有趣的观点,可打磨性更重要。我想我会说……是的。好吧。所以天真地看,你可能会认为 Lean 为数学提供了独特的东西,因为你可以看到它是否能证明。你有老式软件可以告诉你对或错。你把它当作你的验证器。我的意思是,能佐证你观点的是,最初的尝试——我再回到 IMO——一开始 DeepMind 基本上就是这么做的:全部用 Lean,然后第二年就全用自然语言了。所以按你的说法,并不需要。
It's an interesting point, grindability mattering more. I guess I will say on the... Yeah. Okay. So naively you might think Lean provides something unique for math because you're able to see if it can prove it. You have old-school software that can tell you yes or no. You use that as your VR. I mean, what would corroborate your point is the idea that the initial attempts—again I'll just circle back to IMO—it's like initially DeepMind basically does that: everything in Lean, and then the next year it's all in natural language. So to your point, not needed.
我确实认为那个形式化领域还有一个尚未被探索的好处,那就是目前你仍然需要最终由人类来审查单位距离猜想的反例,说它看起来没问题,这给事物的无限可探索性设定了某种边界。如果你考虑像 AlphaGo 或 AlphaZero 那样的东西,它们在自己的宇宙里,只是下围棋并自我探索,完全可能偏离任何人类需要检查的轨道,但它们仍然有自动化的可验证奖励。这不仅仅是你可以在上面做强化学习;还在于你基本上永远不需要检查,你可以直接把算力倾注进去,就像探索围棋的宇宙一样。有趣的是——也许这不会成功,但我认为关于这是否会有所产出,结论仍悬而未决——有了 Lean,你可以想象一个基本上永远运行的程序,不断尝试扩展 Mathlib。Mathlib 是一个 GitHub 仓库,基本上就是把所有数学写成代码。它离所有数学还很远,但他们希望它成为所有数学的代码化版本,你可以问:这个证明正确吗?写这些证明非常费力。围绕它有一个完整的子社区。但你可以想象,如果你有一个 AI,你只是说“尝试扩展 Mathlib”。也许它是一个分支,这样里面就没有垃圾,因为人们对想放进去的内容有特定品味。所以你有了纯 AI 数学库的分支,它就一直运行,永不停止。它不需要任何人检查,对吧?它可以一直进行下去。它可能会提出自己的猜想,提出自己的理论和不同的定义。也许很多都是无用的,但它有这棵可以无限生长的树。这是数学独有的、其他领域没有的东西:你可以按下开始,然后倾注算力,十年后回来看,说“你有什么成果”,而那里一定会有东西,对吧?然后有一个问题:它有用吗?你怎么判断出来?这是一件很有趣的事情。如果这没有产生某种有趣的数学洞见,那才奇怪呢,对吧?所以我认为这才是真正的理由:好吧,Lean 在这个故事中有两种不同的重要性。这是第一种:你可以放手,甚至不检查,进展就会发生。你可以对围棋这么做。我不认为你能对自然语言数学这么做。
I do think there is a yet-to-be-explored benefit of that formalization domain, which is that at the moment you still need ultimately a human is reviewing that counterexample to the unit distance conjecture to say it looks good, and that provides a certain bound on how endlessly explorable things are. If you consider like AlphaGo or AlphaZero style stuff where they're just off in their own universe, just playing a bunch of Go and exploring themselves, completely going potentially off the rails of what any human needs to look at, but they still have this automated verifiable reward. It's not just that hey you can do RL on that; it's also you basically never have to check in and you can just pour compute at them like exploring the universe of Go. What stands to be interesting—maybe this won't pan out, but I think the jury should still be out on whether this will yield anything—with Lean you could imagine having a basically endlessly running program that's constantly trying to extend Mathlib. So Mathlib is this GitHub repository that's basically like all of math written in code. It's very far from all of math, but they want it to be all of math written in code that you can ask: is this proof correct? It's very labor-intensive to write these proofs. There's a whole sub-community around it. But you could imagine what if you just had an AI where you say simply try to extend Mathlib. Maybe it's a fork of it so that it doesn't have trash in it because people have certain taste for what they want to be in there. So you have your fork of the pure AI math lib and it just goes and it just doesn't stop. It doesn't need anybody to check in on it, right? It could just keep going. It might come up with its own conjectures, it might come up with its own theories and different definitions. Maybe many of them are useless, but it just has this infinite tree that it can grow out. That's a very unique thing that math has that nothing else has, where you could press go and then just pour compute at it and look away for 10 years and then come back and say what do you have, and there's going to be something, right? And then there's a question: is it useful or not? How do you sus that out? That's just an interesting thing to be able to do. It would be very surprising if that didn't yield some sort of interesting mathematical insight from it, right? So I think that's the real case for okay there are two different ways that Lean is important in this story. That's the first one of them basically: how it's like you could let go, not even check in, and progress will be made. You can do that with Go. I don't think you can do that with natural language math.
这非常有趣。你看到 Karpathy 的自动研究想法了吗?他写了一个基本上只有一个 Python 文件的程序,用来做基本的 LLM 训练,然后有一个仓库,LM 智能体尝试对文件进行修改。如果它加速了速通,修改就被保留。Eric Jang,他曾来讲解 AlphaGo 的工作原理,在尝试构建一个非常强的围棋机器人时也做了类似的事情。他对这类事情有一些有趣的观察:它非常擅长运行实验并沿着那条路走下去,但不擅长在死胡同处停下来,也不擅长做极度并行的事情。不管怎样,这将来可能会改变。思考它在极限情况下是什么样子非常有趣。
This is very interesting. Did you see Karpathy's auto research idea? He wrote this basically one Python file that does basic LLM training and then just had a repo where LM agents would try to make modifications to the file. If it sped up the speedrun, the modification stays. Eric Jang, who came on to explain how AlphaGo works, did a similar thing when he was trying to build a very strong Go bot. And he had interesting observations about the kinds of... it's really good at just go running an experiment and going down that path, but it's bad at stopping at dead ends and just doing extremely parallel things. Anyways, this will probably change in the future. It's very interesting to think about what it looks like in the limit.
我的意思是,这本质上就像人类数学研究的机构,对吧?就像一个以有趣且有用的方式扩展的库,这样你没有任何基于结果的监督。没有你想要激励的结果,但你有步骤正确的过程。你只是不知道它是否在朝有趣的方向前进。
I mean, this is fundamentally like what the human institution of mathematical research is, right? Just like this is a library extended in interesting and useful ways, and this way you don't have any outcome-based supervision. There's no outcome that you're trying to incentivize, but you have a process that the steps are correct. You just don't know if it's going in an interesting direction.
但是,是的,如果你这样做,你不想完全脱轨,在逻辑空间里随机游走。你可能想要某种监督模型,试图提供关于它是否有用的启发式方法。但没错,就是那种性质的东西。我的意思是,你知道,人们正在研究它,那是五年后的事情之一。我很想知道未来的我们会不会谈论它可能毫无进展,但陶哲轩曾谈到一个研究项目:基本上就是穷举搜索可能的代数空间。比如你可以想象应用于代数系统的不同公理。当我们提出群论时,有一个特定的公理系统,它们看起来有点像任意规则,除非你知道动机。但基本上,如果你尝试所有公理呢?其中任何一个会产生有用的东西吗?绝大多数在某种程度上只是垃圾,比如全部坍缩成没有有趣的结果。但偶尔会有这么一个小岛,一个完全不同的公理系统,至少从它能产生的定理数量来看似乎很丰富。而这正是你想象中自动证明器所擅长的核心内容。
But yeah, you would like if you were doing that you don't want to completely go off the rails and do a random walk through the space of logic. You'd probably want some supervisor model that's trying to provide heuristics on whether it's useful or not. But yeah, something of that character. I mean, you know, people are working on it and that's one of those five years from now. I'd be curious to be able to get the future version of us talking about whether maybe that goes nowhere, but Terry Tao was talking about one research project: it's basically try to exhaustively search the space of possible algebras. Like you could imagine different axioms that you apply to algebraic systems. And so when we come up with group theory, there's a certain axiom system that has this flavor of they kind of look like arbitrary rules unless you know the motivation. But it's basically what if you tried all of them? Do any of these yield useful things? And the vast majority of them is just trash in some way. Like it all collapses to no interesting results. But every now and then there would be this little island of a completely different type of axiom system that at the very least seems rich in terms of the number of theorems that can come out of it. And that's like bread and butter for what you would imagine automated provers being good for.
就像探索那个空间,看看其中哪一个最终能变成某种东西,也许其中某个岛屿实际上能变成某种你可以事后赋予动机的东西,说这就是那种试图达到的结构,就像你可以想象看着群的公理,不知道它关于对称性,但事后意识到“哇,这与研究对称性非常相关”。所以你可以想象那种味道的结果,但不仅仅是探索可能的代数系统,而是所有可能的任何公理的逻辑推论。
like exploring that space and seeing which one of them turns out to be something and like maybe one of those islands actually turns out to be something you can retroactively put motivation on to say this is the kind of structure that's trying to get at in the same way that you could imagine looking at the axioms for a group not knowing that it's about symmetry but retroactively realizing like wow this is very relevant to studying symmetry. So you could imagine results of that flavor but instead of just exploring possible algebra systems it's like all possible like logical consequences of any kind of axiom
关于是否可以在没有 Lean 的情况下提供基于过程的监督这一点。所以 DeepSeek 有他们的 DeepSeek Math 模型,他们发表了一篇关于如何训练它的论文。
on the point about whether you can provide process based supervision without lean. So deepseek had uh their uh deepseek math model that and they released a paper on how they trained it
而且这相当有趣。所以他们有一个问题,自然语言证明的问题是你不知道它是否正确。所以他们有一个验证器,然后这个验证器由一个元验证器训练,确保他们训练这个模型解决的所有问题(比如在解题艺术中)验证器都能给出良好的反馈,而且它确实有效。所以这很有趣。
and it was quite interesting. So they have um the problem with having natural language proofs is you don't know if it's correct or not. And so they have a verifier and then the verifier is trained by a metaverifier that makes sure that any all the problems that they're training this model to solve in like the art of problem solving that the verifier is giving good feedback on that and it like it works. And so it's just interesting.
是的,带有某种元验证的自然语言验证至少在已发表的文献中似乎有效,而且在我们使用的已发布产品中似乎也有效,比如如果你看看编码智能体。
Yeah, natural language verification with some sort of metaverification kind of work at least seems to work so far in the published literature and also it seems to work in the published products that we're using like if you look at coding agents
嗯。
Mhm.
他们在编写干净代码和重构代码等方面变得越来越好。而且我确信存在基于过程的、比如 LLM 作为评判者的东西,它们试图提供品味,说“嘿,这是写这个函数的干净方式吗?我们是否有相同模块化形式的重复?”等等。我觉得这应该也适用于数学,对吧?即使你只在自然语言中工作,似乎数学比其他任何领域都更可信赖验证器。
they're getting better and better at like writing clean code and refactoring code and stuff like that. And I'm sure that that there's process based like llm as judge kinds of things which are saying trying to provide taste and say hey is this like a clean way to write this function are we like are there are there duplicates of the same kind of modular forms and so forth. Um I feel like that should also work for mathematics right it's like it doesn't seem it seems more plausible for math than anything else even if you're only working in natural language that you could trust a verifier.
我的意思是,我们之前讨论过为什么它们不擅长写作,你知道,我在问为什么你不能直接让它们成为好的评判者。如果我给它们两篇学生写的文章,它们能说出哪一篇更准确、更有洞察力。那么,为什么你不能直接有一个验证器说“这是一篇好文章还是不是”?也许最终的失败在于,即使它们擅长区分 B 级文章和 A 级文章,它们实际上并不擅长区分 A 级文章和那种你真正想读的、在 Substack 上可追踪且有洞察力的东西。它们实际上最终更喜欢没有洞察力的文章。是的。
I mean, you and I were talking earlier about why they're bad at writing, and you know, I was asking like why you can't just have like they seem to be good judges. If I give them two essays that like students write, they'd be able to say which one's more like accurate and insightful. Um, so why can't you just have like a verifier saying like is this a good piece of writing or not? And like maybe the ultimate failure there is like even if they're good at discriminating between like a a B essay and an A essay, they're not actually good at discriminating between like an A essay and like a thing you actually want to read that would be, you know, followable on Substack and insightful and all of that. Like they actually end up preferring just uninsightful pieces of writing. Yeah.
所以在数学方面,我想问题会是:仅仅知道“这是一个正确的证明吗?”这一步,本身就适合自动验证器,即使是在自然语言中。你可能仍然能取得大量进展。但我仍然喜欢那种脱离 Lean 的逻辑树,因为你可以真正地偏离轨道,对吧?就像没有对先前表述方式的约束,就像每个人谈论的 AlphaGo 中的第 37 步那样,什么东西能让你跳出先前的启发式?似乎在这种探索中与世界其他部分脱节是有成效的,作为自然语言数学前沿的补充研究追求。
And so on the math front, I guess the the question would be like that step to simply know like is this a correct proof or not? that lends itself to like an automated verifier even in natural language. Uh you could probably still make a ton of the progress. It still doesn't like I still like the sort of tree of logic out of lean front just in that you can really go off the rails, right? like there's just no constraint on like the previous way that things had been phrased before in the same way that you know everyone talks about like move 37 um in like AlphaGo and such like what is the thing that lends itself to just going outside the prior heristics and you it seems productive to have uh uh a disconnection from the rest of the world in that exploration as like a complimentary research pursuit to the natural language math front.
嗯,我的意思是,Lean 的另一个相关点是:假设你有纯自然语言的强化学习环境,和纯自然语言的证明集,人们说过像“先行的 AI 数学家”那样,它们每天生成 10 篇论文,产生一堆东西。如果错误率哪怕有一点,Alex Conurvich 就谈过这个。作为数学家,这会变得难以忍受,因为每次我看到其中一篇,我都不确定是否值得花时间,即使 100 篇里有 99 篇是正确的。我不知道是否值得花时间去读,因为找出那个错误非常费力。而且如果你发现你把所有时间都花在一篇垃圾论文上,那会非常令人沮丧。所以,任何能给你一个绿色勾号的东西,说即使这很难理解,即使会很痛苦,你至少知道它是正确的。其他每个领域都会为此拼命,对吧?而数学拥有这一点,如果模型也能将它们的自然语言证明形式化的话。所以这似乎很巨大,对吧?拥有这种能力,每个领域都会喜欢。所以我认为你是对的,Lean 作为 VR 环境对于一般数学进步的重要性可能被高估了,但我绝对不会把它排除在故事之外。
Um I mean the other the other relevance of lean there would be like okay let's say you have your um pure natural language um RL environments and you have a pure natural language uh set of proofs and people have said like precede AI mathematicians and they go and they generate like 10 papers a day that produce a bunch of stuff um if the error rate if there's like any error rate to that at all. So, Alex Conurvich has talked about this. It becomes insufferable like as a mathematician because you you would basically be like I'm every single time I see one of these I kind of don't know if it's worth my time even if 99 out of 100 of them are right. Um I don't know if it's worth my time to even go through it because it's really labor intensive to find what that error would be. And it's like really frustrating if it turns out you spent all your time on a paper that was trash. And so having anything that's able to give you that green check mark that says even if this is going to be complicated to understand, even if it's going to be a pain, you at the very least know it is correct. Like every other field would kill for that, right? And like math has that um if if the models are also able to take their natural language proofs and formalize them. And so that seems huge, right? The ability to have that like every field would love to have something like that. And so I think you are right that lean is maybe overrated on the side of the importance of it being used as a VR environment for any kind of like just progress in math generally. But I I I definitely wouldn't write it out of the story.
是的。
Yeah.
我也喜欢将数学的这种扩展作为我们文明即将发生的事情的隐喻。
I I also love this extension of math as a metaphor for like what's going to happen to our civilization pretty soon.
当然。
Sure.
对。就像几千年来,人类一直在建立这个知识、理解和一切事物的语料库,我们现在已经将其提炼到这些模型中,在某个时刻,模型将任意地扩展它。顺便说一下,在写作方面,我实际上有一个理论,为什么写作的进展比其他领域更差。但我认为其中一个原因就是你所说的,它们不仅不擅长判断 A 与 B,而且会被 B* 完全带偏,
Right. just like for millennia humanity is building this like corpus of knowledge and understanding and everything that we have now distilled into these models and at some point the models will just like extend that arbitrarily. Um by the way on the writing front I actually have I I have a theory of why writing is making worse progress than these other domains. But I think one of one of them is what you said that they're bad at judging not only A versus B, but they get like just totally derailed by B star,
好的,
okay,
就是那种糟糕的文章,它恰好击中了所有
which is this like shitty essay that just hits all the um
所有 A 应该击中的花哨功能,然后奖励黑客行为就完全脱轨了。但我认为另一个重要的事情是,写作不像代码和数学那样模块化。你知道,你可以用许多不同的方式写一个函数,它们大致做同样的事情,当然你希望它非常干净等等,但归根结底,它工作就行。同样,引理和数学也是如此,然后你可以有一个最终产品,它不同于生产它的方式,所以代码是产生某个最终产品的东西,而你希望有一个功能性的最终产品。而在写作中,最终产品直接就是 AI 正在生产的东西,每个段落、句子、单词都很重要,因为那就是实质。
all the bells and whistles that like A is supposed to hit and then so the reward hack thing just like totally goes off the rails. But I think the other important thing is that writing is not modular in the same way that code and math are. like you know you can write a function many different ways and they kind of do the same thing and of course you want it to be very clean and stuff but like at the end of the day it works it works same with like lemas and mathematics and then you know you can like have some end product that is different from the way it is produced so the code is the thing that produces some end product and you are you want a functional end product um whereas in writing the end product is directly the thing the AI is producing and each paragraph sentence to word matters because that is a thing that is like like that is the substance.
它不像写作中产生的某种独立的东西,不能像代码那样,即使写得乱七八糟也能产生你想要的结果。
It's not like some separate thing that is produced out of the writing. It can't just be slop in the way that code can be slop and still produce some outcome that you want.
但你刚才指出,实际上我们在让智能体编写不仅功能正确而且整洁的代码方面已经进步了很多。为什么同样的进步——让你从仅仅功能正确到写出整洁、可合并的 PR——没有同样带来更清晰的写作呢?
But you were just pointing out how actually we've gotten much better at agents writing not just functional code but clean code. Why is it not the case that the same progress that allows you to go from merely functional to clean and mergeable PR doesn't also result in clearer writing?
是的,说得好。我的意思是,难道不是吗?我同意他们在很多方面是糟糕的写作者,但对于我消费的很多写作内容,我发现最好直接复制粘贴到 LLM 里,然后说“给我解释一下”。解释会比人类写的东西更好。所以很有趣,我们说他们是如此糟糕的写作者,而我的显示偏好却是“能不能让另一个来给我解释?”即使我在电话里和人类专家实时交谈。如果是他们独有的、未编码在分布中的知识,我希望他们给我解释。但如果为了理解那个,我需要先理解一个更基础的概念,我宁愿在社交上可以接受地说“我们先停一下,我想问 LLM 那是怎么工作的,然后再回到你那个特殊的知识点。”
Yeah, that's a good point. I mean, has it not? Like I agree there are many ways in which they're terrible writers, but for a lot of writing I consume, I find it's better to just copy-paste it into an LLM and say 'explain this to me.' The explanation will be better than the thing produced by the human. So it's funny that we say these are such terrible writers, and also my revealed preference is just 'can I have another one explain it?' even when I'm talking to a human expert live on a call. If it's a piece of knowledge they have that only they have, that's not encoded in the distribution, I want them to explain it to me. But then if in order to understand that I need to understand a more basic concept, I would prefer if it was socially acceptable for me to just say 'let's pause there, I was going to ask an LLM how that works, and then we can come back to your special piece of knowledge.'
嗯,那是蒸馏,对吧?一种解释。所以如果我在考虑作为散文作者的质量视角,如果我给你一本书读,我想要一份读书报告,那么我可能相信 LLM 会给我一份更好的读书报告。但我认为人们说它写作不好时真正想说的是——写作是什么?它不只是对已有思想的蒸馏。不只是“如何清晰解释”,因为他们是好的解释者。它关乎洞察力。这就是自动回归生成内容变得非常奇怪的地方,因为当你写作时,你隐约知道为了写好,你必须有一个不可预测的元素。而且这不只是像在脑海中提高温度之类的;而是精确知道在什么时候你想做出一个不可预测的举动。那才是更有洞察力的。所以即使它更擅长解释已有的事物,最初生成那本你想蒸馏的书的是什么?不是 LLM 生成了它然后你只需要它。而是一位作者,通过大量探索世界中的想法,然后决定哪些方面有趣,哪些呈现方式构成连贯、动机良好的叙事。他们以某种方式把这一切组合起来。如果他们是好作者,你可能会倾向于读他们的书而不是蒸馏版。但尽管如此,是什么让一开始就值得探索,并且你还要上传它?我认为正是那一面——当人们说它们写作不好时——那种不可预测性,故意选择新颖的东西,直接与事物产生的方式相矛盾。
Well, that's distillation, right? An explanation. So if I'm thinking about quality of view as an essay writer, if I give you a book to read and I want a book report, then I might believe that okay, the LLM maybe gives me a better book report. But I think what people are really getting at when they say it's bad at writing—what is writing? It's not just distillation of pre-existing ideas. It's not just 'how do you explain clearly' because they are good explainers. It's about the insight. And this is where it gets like auto-regression is a very weird way to generate stuff because when you're writing, you sort of know in order for it to be good, you have to have an element of the unpredictable. And it's not just like increasing temperature in your mind or something; it's knowing exactly the correct point when you want to make an unpredictable move. And that's going to be what's more insightful. So even if it's better at explaining a pre-existing thing, what generated that book that you wanted distilled in the first place? It wasn't an LLM that generated it and you just needed it. It's some author who through a lot of exploration of ideas in the world and then deciding what aspects of it were interesting and which ways of presenting it were the coherent, well-motivated narrative. They put that all together in some way. And if they're a good author, it's probably one that you would err on the side of reading their book instead of the distillation. But still, what makes it worthwhile to explore at all in the first place and you're uploading it at all? I think it's all of that side of it that's when people cite them being bad at writing—that element of unpredictability, of being deliberately choosing something that's novel, that's very directly contradictory to the way that things are being produced.
是的,说得好。我认为他们也非常不擅长构建关于人的良好心智模型,我认为这是写作中非常重要的技能。所以 Annie Matushak 和另一位我现在忘记名字的合作者做了一份有趣的报告,他们试图教 LLM 编写好的间隔重复提示。我真的很喜欢这个,因为尽管这看起来完全是一个随机的技能,但人们正在谈论一年内的递归自我改进,而我们却无法让这些东西写出好的闪卡。这是怎么回事?
Yeah, that's a good point. I think they're also really bad at building really good mental models of people, which I think is a very important skill in writing. So Annie Matushak and another collaborator whose name I'm forgetting right now did an interesting report where they tried to teach LLMs to write good spaced repetition prompts. I really like this because even though it seems like a totally random skill, it's like people are talking about recursive self-improvement in a year and we can't get these things to write good flashcards. What's going on there?
对。
Right.
他们尝试了许多不同的技术。聪明的人们,他们尝试对开源模型进行强化学习,尝试了各种方法,包括思维链以及他们发送给最佳闭源模型的大提示等。在我看来,关键约束是:写好一张卡片是关于预测某人在三个月后的心智状态:他们会如何关联问题,那时他们会想到什么样的答案,以及那个启发是否激发了你实际上想从你要制作卡片的段落中提取的细节。我认为写作也与此类似。当你写东西时,之所以这是一个如此耗费精力、耗时漫长的过程,是因为每个词或每个句子你都应该思考:此刻我读者的脑海里正在发生什么?即使我把措辞颠倒过来,把结尾短语放到开头,比如这是你在阅读句子其余部分之前脑海中出现的第一个图像。那种自动回归可能不擅长那种事情。也许有一种更像扩散模型的属性,考虑整体而不是逐句进行。但我也认为这需要大量的心智化,而这些模型奇怪地在这方面挣扎。
And they tried many different kinds of techniques. Sophisticated people, they tried to RL open-source models, they tried all kinds of things including chain of thought and the big prompt they sent to the best closed-source model, etc. And the key constraint it seemed to me was that writing a good card is about projecting somebody's mind in 3 months: what is the way in which they will associate the question, what kind of answer will they be thinking by that moment, and is that the elicitation that inspires the detail you actually want to take away from the passage you're trying to make cards about. I think writing also is similar to this. When you're writing something, the reason it's such an enervating process that takes so long is each word or each sentence you should be thinking: what is happening in my reader's mind right now? Even if I flip the phrasing around where the end phrase goes to the beginning, like this is the first image that comes to your mind before you read the rest of the sentence. That kind of maybe auto-regression is bad at that kind of thing. There's maybe a more diffusion-like property of considering the whole rather than going sentence by sentence. But also I think that requires a lot of mentalizing, which these models weirdly struggle at.
嗯,有趣的问题:它们在这方面挣扎很奇怪吗?我可能会搞砸这个——你知道当你引用你曾经读过的研究时,可能那个研究并不真实之类的。有一个非常令人难忘的研究。好吧,假设你想测试人们的情商。你展示一张某人面部表情的闪卡,有人试图描述那是什么情绪。网上有一些很好的测试,会有一张脸和四种可能的情绪,要准确描述正确的情绪出奇地难,但你也感觉到确实有一个正确答案。如果你在生活中对人们尝试这个,你会注意到那些社交能力强的人做得很好,而那些更左脑思维的人则不行。好的,这是一种你可以做的测试。我模糊记得一个实验,他们让刚打过肉毒杆菌的人做前测和后测,后测中他们解读他人表情的能力差了很多。这感觉有点奇怪。
Well, interesting question: is it weird that they struggle at that? So I might butcher this—you know how when you cite studies that you once read and it's like maybe the study wasn't real or something. There's one very memorable one. Okay, so let's say you want to quiz people's EQ. You show a flashcard of someone's facial expression and someone's trying to describe what's that emotion. There are really good tests online that'll have a face and then four possible emotions, and it's surprisingly hard to describe exactly the correct emotion, but you also get the sense there really is a correct answer. And if you try this with people in your life, you'll notice that the ones who are pretty plugged in socially do really well on it, and the ones who are a little more left-brained don't. Okay, so that is a kind of test you can do. I vaguely remember an experiment to this effect where they took people who had freshly gotten Botox in some way, and they did a pre-test and a post-test, and post-test they were just much worse at reading people's expressions. That feels kind of weird.
等等,他们打了肉毒杆菌?
Wait, they got Botox?
所以,做测试的人,就像你做完测试,然后去打肉毒杆菌,脸就僵了,现在你理解所见情绪的能力变差了,对吧?想法是,理解你所看到的情绪,部分原因在于你自己在面部层面上的模仿,比如动动你的面部肌肉。你看到那个表情,你模仿它,然后在某种非常潜意识的层面上,你会想,“哦,那是焦虑”。所以从这个意义上说,如果模型的心智理论很差,当然,它们知道一切,因为它们读了所有人写的东西。但在真正能设身处地理解你的层面上,就像我的面部肌肉模仿你的面部肌肉那样,这才是帮助我理解你感受的方式。这完全不奇怪。它们没有面部肌肉。它们的大脑运作方式完全不同。就像一个外星人在试图共情。它怎么可能有心智理论?心智理论会是一个非常涌现的东西,而我们则可以直接把它接入我们自己的大脑,我们有现成的硬件来直接放置它。
So the person taking the test, it's like you do the test and then you go and get Botox and your face is all frozen and now you are worse at understanding the emotions of what you see, right? And the thought is that part of understanding the emotion you're looking at is doing it yourself at a facial level, like moving your face muscles. You see that, you mimic that, and you're like, "Oh yeah, that's anxiety" at some very subconscious level. So in that sense, if it is the case that models have bad theory of mind, sure, they know everything because they read what everyone wrote. But at a level of actually being able to put themselves in your shoes in the same way that my face muscles are mimicking your face muscles, that's what helps me understand how you feel. Not surprising at all. They don't have face muscles. Their brain works completely differently. It's like an alien trying to empathize. How could it have theory of mind? It would be this very emergent thing to have theory of mind, whereas we can just plug it into our own minds, and we've got the ready-made hardware to just place it in.
这很有意思。从这个角度看,这并不那么令人惊讶。
That's very interesting. From that lens, it's not that surprising.
好的,Grant,我们都是 James Street 的合作伙伴。我相信这些年来你和很多 James Street 的人打过交道。你觉得他们或他们的文化有什么独特之处?
Okay, Grant, we are both partners with James Street. I'm sure over the years you've interacted with a lot of James Streeters. What have you found that's unique about them or their culture?
我今年和他们做了一次访谈,这很有意思,因为他们通常不对外露面。在业内,他们以极高的员工留存率而闻名,人们就是愿意留在那里。我觉得深入了解这一点,我记得有人评论说,尽管人们有研究员、交易员或工程师这样的职位头衔,但他们常常不知道同事的实际角色是什么,因为每个人都在做一点其他事情。即使你名义上是交易员,你也在做很多研究;即使你名义上是研究员,你也在做很多编程。我猜想,这可能就是他们留存率如此之高的部分原因,因为任何想要成长的人都有机会做很多不同的事情。
I did this interview with them this year that partly was interesting because they don't usually have anything outward facing. In the industry, they're known for having a pretty wild retention rate, like people just stay there. I think getting an inside view of that, I remember one of the comments someone was saying even though the people have role titles like researcher or trader or engineer, they often don't know what their colleague's actual role is because everyone's doing a little bit of everything else. Even if you're officially a trader, you're doing a lot of research; even if you're officially a researcher, you're doing a lot of coding. I suspect maybe that's part of why they have the insane retention that they do, because anyone who wants to be growing just has the chance to do a lot of different kinds of things.
好的,Grant,这次我来帮你宣传一下。如果你想看 Grant 和那里一些人做的完整访谈,请访问 3b1b.co/jainstreet。
All right, Grant, I'll do the plug for you this time. If you want to watch this full sitdown interview that Grant did with some of the folks there, go to 3b1b.co/jainstreet.
好的,Grant,让我们多聊聊 AI 和数学。关于使用 LLM 学习,你有什么建议?
All right, Grant, let's talk more about AI and math. What advice do you have about using LLMs to learn?
正如我所说的,对于很多众所周知的概念,我觉得它们非常有帮助。但往往再往下问几条消息,我试图理解某个东西时,它们自己就很困惑,或者把我搞糊涂了,而且解释得不对。然后我知道,和合适的人聊三分钟就能解决我的困惑。我不知道。我觉得我们会越来越想用这些东西来学习。所以,你有没有注意到更有效地使用它们来理解概念的方法?
As I was describing, for a lot of well-known concepts I find them very helpful. But often just a couple of further messages down and I'm trying to understand something, and they're so confused themselves or confusing me, and they don't explain the right way. Then I know that talking to the right human could clear up my confusion in three minutes. I don't know. I feel like more and more we're going to want to use these things to learn things. So, have you noticed ways to use them more productively to understand concepts?
我很想听听你的看法。我先说说我的。我觉得学习中的一个相关洞见是认识到“谁”比“什么”更重要。所以,给任何大学生的选课建议:少关心你已有的兴趣,因为它们现在有点随意;多关心教课的人是不是好老师,是不是你能产生共鸣的人。我认为在选择读什么,比如读什么书时,作者是谁可能比是否是你之前的兴趣更重要。如果你之前喜欢过一本书,就去读那个作者的其他作品,而不是读同一主题的另一本书。我这就联系到 LLM 了。所以,尝试学习某样东西时,看维基百科页面和看,比如说,哲学主题的斯坦福哲学百科全书,或者数学主题的普林斯顿数学指南,感觉是不同的。区别在于,那些文章是由一个人精心撰写的,他试图真正围绕它构建动机和一切,而维基百科是一个达到的局部最小值,基本上每个句子都必须正确。我认为好的阐述在过程中不太关心正确性,但你可以故意构建一些有点错误的东西,然后在过程中纠正,这在众包环境中会被编辑掉。所以,LLM 的解释目前对我来说很像维基百科,也就是说,很了不起。想象一下维基百科之前的世界,要花多长时间才能找到并弄清楚东西。但尽管如此,维基百科页面最有用的部分是什么?通常只是底部的参考文献。你查看关键参考文献,然后去读它们。有时这能提供更好的概述。所以我经常只是问 LLM:“我应该读谁?”也许我还可以给出一些具体的学习方式。我实际上被这个误导过一次,我记得我试图学习半导体之类的东西。我想:“这感觉很视觉化。但全是文字。有没有什么很好的可视化数学视频——不,不是数学,抱歉——一个很好的可视化视频来解释你提到的概念?”然后 Claude 说:“有,这里有几个。”第一个是:“这是来自 Three Blue One Brown 的一个。”我想:“我保证没有。”结果那是一个真实的视频,一个真实的链接,但只是把别人的视频错误归因了。而且那个视频很好。我点击观看那个视频来学习那个东西,体验比继续在那里提问好得多。所以从这个意义上说,基本上把它当作一个超级增强版的谷歌,用来精准定位到正确的人类撰写的资源。
I'm curious to hear your take on this. I'll give mine. I feel like a relevant insight in learning was recognizing that who matters more than what. So, advice to any college student when choosing courses: care a little bit less about your pre-existing interests because they're kind of arbitrary right now, and care a little bit more about whether the person teaching it is a good educator and someone you resonate with. I think in choosing what to read, like what books to read, who the author is maybe matters more than if it's a prior interest. If there's a book you've liked before, read what else that author has written rather than reading another thing on that subject. And I'm getting to LLMs on this. So there's a difference in feel for trying to learn something if you look at a Wikipedia page versus if you look at, say, a philosophy topic on the Stanford Encyclopedia of Philosophy, or a math topic on the Princeton Companion to Mathematics. The difference is that the articles are deliberately written by one individual who tries to actually craft a motivation around it and everything, whereas Wikipedia is this local minimum that's reached where basically every sentence has to be correct. I think a good exposition cares a little bit less about correctness on the way, but you can deliberately craft things that are a little bit wrong that you correct along the way, which gets edited out in a crowd source environment. So LLM explanations feel to me at the moment a lot like Wikipedia, which is to say amazing. Imagine the world before Wikipedia, how long it would take to find and sus things out. But nevertheless, what's the most useful part of a Wikipedia page? It's often just the references at the bottom. You look at the key references and go to them and read them. Sometimes that gives a much better overview. So often I like to just ask an LLM, "Who should I read?" and maybe I can even give some specifics on ways I want to learn. I actually got gaslit by this once where I remember trying to learn about semiconductors or something. I was like, "This feels very visual. This is all text. Is there any really good well-visualized math video or, not math, sorry, a well-visualized video explaining the concepts you're getting at?" And Claude was like, "Yeah, here's a couple." And the top one, it was like, "Here's one from Three Blue One Brown." I'm like, "I can guarantee that there's not going." And it was an actual video, an actual link, but it just had misattributed someone else's. And it was good. I had a much better experience clicking over and watching that video to learn about the thing rather than trying to proceed forward with questions there. So in that sense, basically using it like a very souped-up version of Google to zero in on the right human-written resource.
那你呢?你经常使用这些。最好的方式是……
What about you? You engage with these a lot. What's the best way to...
我觉得你说到点子上了。
I think you put your finger on it.
我最有效率的学习经历是,当有人类制作的某种作品——无论是文章、书籍还是视频——以正确的方式组织相关概念,并逐步建立动机,说明为什么下一个想法有助于解决接下来会遇到的问题,然后利用 LLM 对书籍指出的这个分支做一点修剪。
The most productive learning sessions I've had is when there's some artifact that a human has produced, whether it's an article, a book, a video that organizes the relevant concepts in the correct way and builds up the motivation of why building up the next idea would be relevant to solving the next problem you'd encounter and the next idea and the next idea and then using the LLMs to just do a little bit pruning around this branch that the book has identified.
所以我当时在读 Steven Strogatz 关于混沌和非线性动力学的教材。我很喜欢那本书。
So, I was going through Steven Strogatz's textbook on chaos and nonlinear dynamics. I love that book.
对,就是那本混沌的。我也很喜欢。
Yeah, the chaos one. I love that book.
我读这本书时感觉非常愉悦,就像你的视频以书籍形式呈现一样。他太棒了,超级有趣。我的学习方式是:屏幕三分之一放他的大学讲座视频,三分之一放教材对应部分,三分之一放 LLM。我就在想,如果我回到大学现场听这堂课,肯定完全听不懂。那些学生一定很聪明,因为我需要暂停、读教材、和 LLM 对话,然后再重新开始。但有了他策划的正确概念理解顺序和激发理解概念的合适问题,就完全不同了。
And so I was going through it and it was like bliss. It was like your videos in book form. He's so good. It was super fun. And the way I was learning it is like I'd have on one third of the screen his lecture from university. On one third of the screen I'd have that part of the textbook and on one third of the screen I have an LLM. And I was actually thinking if I was back in college and watching this lecture live, it would just totally go over my head. Like these kids must be really smart because I'm like pausing and reading the textbook and talking to and then just restarting again. But with him curating what is the right order to understand concepts, what is the right problem to motivate understanding a concept.
哦,还有一件事 LLM 非常不擅长,而优秀的人类能做到:当你提问时,他们会说“实际上,你思考这个话题的方式不对。你应该问的问题是……组织这些概念的正确方式是 X。”而 LLM 真的做不到这一点。
Oh, also another thing LLMs are really bad at is a thing a really good human can do: when you ask a question, they say like 'Actually, you're just not really thinking about this topic the correct way. The question you want to be asking is... The correct way to organize these concepts is X.' And the LLM just can't really do that.
对,它有点太迎合了。我的意思是,它们最终非常顺从,你知道,总是“哦,多么有洞察力的问题”。你希望去掉这些。
Yeah. It's a little too placated. I mean this is ultimately like they're very supplicant, you know, it's very 'Oh, what an insightful question.' You want to strip that down.
说得好,我觉得这触及了心智理论。比如,意识到提出某种问题表明你的心智结构与讲解者不同。有时人们会过度这样做,对吧?我认为一个真正好的老师,比如在初中数学课堂上,如果学生问了一个表明他们用不同方式思考的问题,你很难当场认真对待:“等等,用那个方法能得出正确答案吗?”然后再说“哦,别那样,我们这样做”。真正好的老师能够柔道般地将学生创造性的思考方式融入进来。LLM 做不到这一点,对吧?它们不会重新框定你的问题,而是直接跑偏。
That's a good point and I think that cuts to theory of mind a little bit. Like recognizing that to ask a certain kind of question reveals that the mental structures are not the same as what the explainer has. And sometimes people do this to a fault, right? Like I think a really good teacher, let's say you have a middle school math classroom, if a student asks a question that suggests they're thinking about it in a different way, it's actually really hard to take seriously in the moment, 'Hang on, could you get to a right answer with that?' Before you say, 'Oh, instead of that, let's do this.' And the really good teachers are able to jiu-jitsu the creative way that the student was thinking about it and bring it in. LLMs aren't doing that, right? When they are not reframing your question. Instead, they kind of run off.
但至少,感觉这里有三个层次。LLM 是一个层次,好的讲解者是另一个,而 A+ 级的讲解者能够柔道式地利用你的思维方式,并说“哦,那在那种情况下有用”。所以也许存在一个循环,五年后,LLM 仍然在做同样的事,但方式更好。
But at the very least, it feels like there's three levels here. LLM is at one, good explainer is at another, but then the A+ explainer is the one who can jiu-jitsu your way of thinking and say, 'Oh, that's where that's useful.' And so maybe there is a certain cycle all the way around where again, five years from now, the LLM will still be doing that, but in a better way.
对于那些经常给你发邮件问这个问题的学生,你有什么建议?他们说:“我想学数学,我对这个学科充满热情,但看到 AI 取得的进展,我不知道是否还值得把它作为职业。”这不仅适用于数学领域的人,也适用于那些注意到自己领域因 AI 而获得生产力提升的人。编程与此非常接近。你对这些人有什么建议?
What is your recommendation to students who email you this question all the time: 'I want to pursue mathematics, I'm really passionate about the subject, but seeing all the progress AIs are making, I don't know if it makes sense for me to pursue this as a career.' And this is relevant not only to people in mathematics but to people who are noticing that their field is getting productivity gains from AI. So coding is very adjacent to this. What advice do you have for people?
我不会相信我给出的任何建议——这大概是我要说的前提。但即使在 AI 出现之前,对于任何你将要从事的工作,真正理解钱从哪里来、你实际增加什么价值、以及这两者之间的联系,都感觉非常重要。而且我认为人们对此的思考少得惊人。尤其是学生,他们处于这样的环境中:他们可能想进入数学领域,因为他们一直擅长数学,并且因为正确完成下一个环节而得到奖励。当他们想成为数学家时,是因为这是继续从事这件事的一种方式。他们想的是“我要去人们能做这个的地方”,而不是思考“我为别人增加了什么价值,以及这在多大程度上是薪水流向我的原因?”因为不同情况差异很大。在某些情况下,一位非常有声望的数学家,他们的存在为大学带来了品牌价值,这就是大学想要他们的原因。在某些情况下,NSF 拨款是因为你相信基础科学具有公共产品属性,并且围绕它有一个机构,然后有一整套官僚体系试图作为我们所谓公共产品的代理,还有一套复杂的流程来正确预测你的进展是否符合资助精神。有时就是直接教学,对吧?人们喜欢把孩子送到有专家教学的机构,这就是你做的,你通过成为专家提供品牌价值,然后通过教学提供直接价值。所以无论 AI 是否能证明定理,无论我们是在 2016 年还是 2026 年,那些想成为数学家的学生思考得不够多,但我认为值得思考。对我来说,我当时并没有刻意思考,而是偶然进入了这条职业道路——基本上数学探索可以货币化为娱乐,我偶然发现了这一点。我很感激,但那是意外,不是刻意的。如果我能批判性地思考,也许可以避免依赖偶然性,更有意地做到这一点。
I wouldn't trust any advice that I give. That would be how I'd couch it. But even pre-AI, it feels very important for any job that you're going to go into to really understand where the money's coming from and what value you're actually adding, and the connection between those two. And I think often a surprisingly small amount of thought is put towards that. Especially students, they are in this environment where they probably want to go into math because they've always been good at it and they've just been rewarded in life for proceeding through the next hoop correctly. And when they think they want to be a mathematician, it's because it's a version of getting to continue to engage with that. It's like, 'Well, I'll go where people get to do this,' rather than thinking 'What value am I adding to other people and to what extent is that the reason that salary is flowing in my direction?' Because it's actually quite different in different cases. In some cases, it's a very prestigious mathematician and their presence at a university lends a certain brand value, and that's why the university wants them. In some cases, it's the NSF grant given because you've got this public good belief that basic science has, and you've got this institution around that, and there's going to be this whole bureaucracy around trying to act as a proxy for what we think that public good is, and a whole song and dance around how to correctly make them predict that your progress will be in the spirit of that funding. Sometimes it's just straight up teaching, right? It's like people like to send their kids to an institute that has experts teaching them, and that's what you're doing, and you are providing the brand value by being an expert and then the direct value by being a teacher. So regardless of whether AIs are proving theorems or not, or whether we're talking in 2016 or 2026, that is a thing that not enough students thinking 'I want to be a mathematician' think about, but I think it's worth thinking about. For me, I just wasn't necessarily thinking about it and kind of stumbled into this career path where basically math exploration can be monetized as entertainment, and I stumbled into that. I'm very grateful that I did, but it was an accident. It wasn't a deliberate thing, and I think I could have avoided relying on serendipity and maybe done that a little bit more by design had I been thinking critically about it.
所以回到你的问题,如果自动定理证明几乎实现了,而且它们还是很好的解释者,那么即使是为了让人类理解,我认为数学家所承担的许多社会角色实际上并没有太大变化,对吧?公众仍然觉得基础科学有价值,我们信任数学家的判断,让他们决定时间花在哪里最值得。而声望来自那个群体内部——是其他成员说这是一个很好的结果,而不是像拨款申请者那样真正理解代数数论才能判断结果好坏。所以会存在一种内部文化,决定什么构成有价值的贡献。也许它会从定理证明转向好的定义写作,也许是博物馆策展人的想法。但你会拥有同样的群体,只要整个社会仍然重视基础科学的前提。而且如果我们处于 AI 带来的富足世界,可能这方面的资金反而会更多,对吧?在机构声望方面,比如他们的讲师是谁。
So to your question, if it's the case that you have almost automated theorem proving and then let's say it's the case they're also really good explainers, so it's like even to get the human understanding, I think a lot of the social role that mathematicians serve actually doesn't change that much, right? You still have a sense of as a public we sort of feel like there's value to basic science, and we're trusting in the judgment of mathematicians to determine where their time is best spent. And the prestige comes from within that community. It's like other members saying that this was a really good result, more than it is like the grant writer who really understands algebraic number theory to understand that it's a good result. And so there's going to be some inner culture of what constitutes valuable contributions. Maybe it shifts away from theorem proving and maybe it shifts towards good definition writing. Maybe it's that museum curator idea. But you're going to have that same community, and as long as society as a whole is still valuing the premise of basic science. And if we're in the abundance world of what AI brings, probably there's more funding in that direction in some sense, right? On the side of prestige to institutions for who their lecturers are.
我其实认为教学是后 AGI 时代最稳定的工作之一,因为它非常注重关系。如果父母有富足的财富,他们愿意把钱花在好的教学和教育上,而这远不止是解释。即使大语言模型是很好的解释者,老师所做的是一种社交、辅导、导师式的事情,这可能是未来 50 年里最稳定的职业之一。所以,既然很多数学家的角色与之重叠,你作为即将进入这个领域的学生,可以朝这个方向努力。我实际上认为更多的学生应该考虑并重视成为一名数学教育者的想法,以及这对下一代所能带来的价值。我要再次强调,我不认为自己有资格对未来的年轻数学家说“你应该如何看待未来”,因为我是一个 YouTuber,对吧?我不是他们想要进入的机构里的人。所以我是作为旁观者发言。但这感觉像是普遍的好建议:知道钱从哪里来,知道你在哪里能融入其中。如果你只是问这些问题,你实际上已经领先于其他所有初出茅庐的准数学家了。
I mean, I actually think teaching is one of the most stable post-AGI jobs that there is, because it's so relational. It's so like this is where parents want to spend their money if they have an abundance of wealth, is on good teaching and good educating, and it goes so far beyond explanations. It's like even if LLMs are good explainers, the thing that a teacher is doing is such a social, coaching, mentor type thing that that's probably one of the most stable careers that's going to exist over the next 50 years. And so in so far as what a lot of mathematicians' role overlaps with that, you as the prospective student going into it, you could lean into that. I actually think a lot more students should think about and give credence to the idea of being just a math educator and the value that that can serve towards the next generation. So I'll couch again on I don't think I'm the one to say here, prospective young mathematician, here's how you should think about the future, because I'm like a YouTuber, right? I'm someone who is not in the institution that they are thinking of going into. And so I'm speaking as an outsider looking in. But it feels like generally good universal advice: know where the money is coming from, know where you plug into that. And if you're just asking those questions, you're actually already steps ahead of all the other fledgling prospective mathematicians.
是的。事实上,我认为在疯狂的世界里,在 5 到 10 年内,AI 不仅会提出千禧年大奖问题的解决方案,还会首先提出全新的问题、全新的数学领域和对象等等。在那个世界里,首先有大量的富足,其次,AI 思维走得最远的地方,也就是它们超越我们视野最远的地方,将是数学。
Yeah. And in fact, I think in the crazy world, in the world where within 5 to 10 years, the AIs are coming up with not only solutions to the Millennium Prize problems, but coming up with totally novel problems to be solving in the first place, novel mathematical fields and objects and stuff. It is in that world where first of all there's a ton of abundance, and two, the things that AI minds will have gone furthest in, where they will have seen furthest beyond our horizons, will be mathematics.
而且会有巨大的需求,比如“AI 看到了什么?你能解释给我们听吗?”
And there will be so much demand of like, what have the AI seen? Can you explain it to us?
是的,我觉得在那个世界里,如果还有任何工作的话,提炼 AI 所学到的东西肯定是其中之一。
Yeah, I feel like in that world, if there's any jobs whatsoever, surely distilling what the AIs have learned will be one of them.
而且有趣的是,这一切都假设它没用,对吧?我们并没有谈论所从事的数学的实际应用。所以只要有任何经济效用,你可以想象,那些理解它并能决定它应该指向哪里的人,通过作为策展人做出判断,将数学巨兽指向有用的方向,实际上具有更大的经济价值。这突然之间就比以往更有杠杆作用了。我能问你一个问题吗?显然,AI 用于数学的一个问题不仅是它能不能做,还有它做得好吗?
Also it's funny because all of this sort of presumes that it's useless, right? Like we're not talking about the actual practical applications of what math is being done. So in so far as there's any economic utility to it, you would imagine that the people who understand it and are able to make the decision of where it should point, they actually have a lot more economic value by being able to make that judgment as curator and point this behemoth of new math in a useful direction. Like suddenly that's a much more levered move to make than it had been previously. Can I actually ask you about that? So the obviously one question for AI for math is not only can it do it, but is it any good?
是的。
Yeah.
或者它对任何事情有用吗?你描述了任务群论中的所有方式。我们试图解决这个问题,试图找出关于不同类型函数根的各种随机事实,现在有各种实际应用跨越许多不同领域。你有没有感觉,如果我们完全达到人类数学领域加速 10 倍或 100 倍的地步,会发生一些疯狂的事情,还是我们实际上会被其他领域瓶颈所限制?
Or is it any good for anything? You were describing all the ways in mission group theory. We're trying to solve this, we're trying to figure out random facts about the roots of different kinds of functions, and now there's all these different applications that are practical across many different fields. Do you have some sense of if we just totally get to a place where mathematics, the field of human mathematics, is accelerated 10x or 100x, that something crazy happens, or are we just actually going to be bottlenecked by other fields?
我认为有些领域可能会。我的意思是,这非常尖峰,对吧?我觉得像代数数论的进展,感觉不太可能解锁什么……但我不知道。我记得和一位做动力学和偏微分方程求解的数学家聊过,他提到他的团队有一些想法……让我看看我总结得对不对。就像波音制造飞机的方式:他们制造飞机,然后做大量测试,然后必须根据测试拆解并重新组装。他们基本上有一些见解,关于如何在模拟中做更多事情,这样就不必拆解和重建,这为波音节省了数十亿美元之类的。然后他们就开始资助那个团队。这非常明显与应用相邻,因为偏微分方程就是这样。所以你可以想象,那个领域的进展确实会解锁一些东西。我不知道是不是这些阶跃变化,但也许更多是在发动机设计变得更流畅,或者你提出正确的机翼形状,而不是运行一大堆复杂的计算流体动力学模拟,或者你能够加速你的计算流体动力学模拟,因为某些纯数学见解使它们更高效。我打赌你会看到很多很好的渐进式改进。数学的重大突破立即转化为巨大的经济突破似乎不太可能,比如你解决了纳维-斯托克斯问题,然后解锁了模拟更多东西的能力。但你可能会在这些边缘看到一些有意义的溢出,从纯数学见解泄漏到其他领域。
I think there's some fields that probably will. I mean it's super spiky, right? I think like progress in algebraic number theory, it feels unlikely that that then unlocks some... but I don't know. I remember talking to this mathematician who does more like dynamics and PDE solving type stuff, and he was referencing basically like his group had some ideas that... let me see if I summarized this right. It's like the way that Boeing would make planes is they would make it and then they would do a bunch of tests and they had to disassemble it and reassemble it based on those tests, and they essentially had some insights on how to do more things in simulations such that you don't have to deconstruct and rebuild, and it saved Boeing just like billions of dollars or something. And then they just started funding that group. Which is so... it's much more obviously application adjacent, because PDEs just sort of are that. So progress in that domain you would imagine actually does unlock some things. And I don't know if it's these step changes, but maybe it's more on the side of engine design becomes just a little bit more fluid, or you coming up with the right wing shape instead of running a whole bunch of complicated CFD, or maybe you're able to speed up your CFD simulations because of certain pure math insights that make those more efficient. I bet you'd just see a lot of great incremental improvement there. It seems less likely that the massive breakthroughs in math immediately turn into this massive economic breakthrough, like you solve the Navier-Stokes problems and then that unlocks an ability to simulate more things. But you probably will see at those fringes some meaningful leakage outside of the pure math insights into other things.
而且,有很多人在从事 AI 工程师、物理工程师、材料科学等领域的工作,你得想象他们应该有能力审视 AI 在数学上的洞见,并判断这些洞见是否在某种程度上相关。所以,这又是那种我不会在这里断言一定会发生的事,但如果未来五年内没有出现直接归因于 AI 数学进展的经济上有价值的改进,那会有点令人失望,也有点令人惊讶。如果 AI 只是解决了一堆空泛的问题,却没有触及任何直接与物理世界相关的数学,那确实会让人失望。
But also I mean there's a ton of people working on things like AI engineers, physical engineers, material science, and things like that that would be... you have to imagine that they would be in a good position to look at the AI math insights and decide if they're relevant in some way or not. And so, it's another one of these things where I'm not going to sit here and predict that there will be, but it would be a little bit disappointing and a little bit surprising if there weren't over the next 5 years economically valuable improvements that were directly referable to the AI progress in math. Like that just would be kind of disappointing if it was just taking down a bunch of airish problems and none of them actually... it wasn't doing any of the math that actually directly touches the physical world.
是的。我同意你的观点,数学史上很多内容就是建立这些概念和联系的堆叠。有时这些堆叠会相互连接,或者你在别处发现了应用。至少,你建立了巨大的堆叠,然后当社会在奇点期间取得更广泛的进展,当我们进入奇点的工业部分时,你就有所有这些不同的想法,希望它们能在世界的其他地方有用。我是说,就像我说的,正在发生的事情有趣的一点是,它让人们退后一步问,数学是什么?也许其中一个尴尬的结论会揭示,天哪,在过去,数学已经完全变得无用了。
Yeah. I mean to your point about a lot of history in mathematics is about building up these piles of concepts and connections and whatever. Yeah. And sometimes the piles connect with each other or you discover an application somewhere else. At the very least you just build up this huge pile and then as broader progress in society happens during the singularity when we get to the industrial part of the singularity, you just have all these different ideas that you can hopefully are useful in other parts of the world. I mean, yeah, it like I said, one of the interesting things about what's happening is it causes people to step back and ask like, what is math? And maybe one of the awkward conclusions of it will be revealing like, oh man, over the last it's just become wholly useless.
就像所提出的问题已经变得与物理上可应用的东西如此脱节,这是数学家们必须面对的事情之一。每个人都会看着说,等一下,你们应该在哪里……如果那里有那么多 10 倍的进步,为什么我们在这里看不到?然后,每次我们写那些资助提案,说相信我们,椭圆曲线的进展将有助于密码学,这反而凸显了一个事实,也许它并没有帮助。所以这是一种可能性。
Like the kind of questions being asked have become so divorced from things that are physically applicable that that's one of the things mathematicians have to come to terms with where everyone will look and be like, a second, where are you guys supposed to... if there's so much that's like 10x progress there, why aren't we seeing it over here? And then it's like every time we wrote those grant proposals that said trust us, the elliptic curve progress is going to help with cryptography, it shines a light on the fact that maybe it doesn't. So that's one possibility.
嗯,Grant,这非常有趣。非常感谢你参加。
Um Grant, this is super fun. Thanks so much for doing it.
当然。我的荣幸。
Absolutely. My pleasure.