AI Surpasses Human Understanding in Math: OpenAI's Noam Brown on the Erdős Conjecture Disproof
打开互动全文版(中英对照 + 朗读 + 问答)→Noam Brown 讨论 AI 模型如何在数学推理上超越人类能力,以 OpenAI 最近推翻 Erdős 单位距离猜想为例。
Noam Brown discusses how AI models have surpassed human ability in mathematical reasoning, exemplified by OpenAI's recent disproof of the Erdős unit distance conjecture.
欢迎来到《博弈论》。我是贝恩资本风投的 Slater,今天我们的嘉宾是 Noam Brown。Noam 在 OpenAI 专注于多智能体推理和测试时算力扩展。他是 O1(OpenAI 首个公开推理模型)的关键人物之一。去年,他的团队横扫了数学和计算机竞赛圈,他们的模型在 IMO、IOI 和 ICPC 中获得了金牌。今天我们将讨论 OpenAI 最近对埃尔德什单位距离猜想的证伪,以及 AI 在数学研究中的一般应用。Noam,感谢你今天来做客。
Welcome to the Game Theory. I'm Slater with Bain Capital Ventures, and today we're talking with Noam Brown. At OpenAI, Noam focuses on multi-agent reasoning and test time compute scaling. He was one of the key people behind O1, OpenAI's first public reasoning model. Last year his team swept the math and CS contest circuit. Their model won gold medals at the IMO, IOI, and ICPC. Today we'll talk about OpenAI's recent disproof of the Erdős unit distance conjecture and about AI for math research in general. Noam, thanks for coming today.
当然,我很高兴来到这里。
Of course. I'm happy to be here.
我们今天坐在这里,大约是在 OpenAI 宣布证伪埃尔德什平面单位距离猜想一周后。这是我想谈的第一件事。我的第一个问题是,为什么是现在?IMO 的结果大约在 10 个月前,已经非常令人印象深刻。为什么我们今天才看到数学研究的第一个重大成果,而不是 10 个月前?模型或系统有什么不同?
Well, we're sitting here today about a week after OpenAI announced the disproof of the Erdős conjecture on the planar unit distance problem. And it's the first thing I want to talk about. Maybe my first question is why is this happening now? The IMO results, which were incredibly impressive, happened about 10 months ago. Why are we seeing the first major results in math research today as opposed to 10 months ago? What's different about the models or the systems?
从很多方面来说,这并不令人惊讶。大约一年前我们得到 IMO 结果时,就预计进展会继续,我们会看到数学上更重要的成果。最终模型会达到证明人类无法证明的事物的程度。在过去一年里,模型开始证明一些未证明的问题时,就有一些初步迹象。我认为有人质疑那是否只是因为人类没有给予足够关注,或者模型做的事情并不那么有趣。这个结果的酷之处在于,它可能是第一个模型证明了对数学家来说真正有趣且非常重要的问题的案例。我们觉得这迟早会发生,模型能力越来越强,最终就发生了。我要说的是,这个结果不是我或我的团队促成的;它是模型能力增强的自然结果。OpenAI 的一些研究人员决定看看模型是否到了能证明这类问题的水平,他们用那个问题测试了模型,答案是肯定的。
In many ways it's not that surprising. When we got the IMO result about a year ago, we expected progress to continue, that we'd see more significant results in math. Eventually we'd reach a point where the models are proving things that humans can't prove. There were some initial signs over the course of the year when the model started proving things that were unproven. I think there were questions about whether that was just because humans hadn't paid much attention, or if it was doing things that weren't that interesting. What's cool about this result is it's perhaps the first case of a model proving something that was really interesting to mathematicians and also a very significant problem. We felt this was going to happen eventually, and the models kept becoming more capable, and it finally happened. I should say it wasn't me or my team that enabled this result; it's a natural consequence of the models becoming more capable. Some researchers at OpenAI decided to see if it was at the level where it could prove a problem like this, and they ran it on that problem, and the answer was yes.
你对组合学特别感兴趣吗?你觉得组合学是 AI 可能取得快速进展的领域,还是说在幕后你们尝试了很多开放问题,而这个问题恰好是第一个被解决的?
Were you especially interested in combinatorics? Do you feel like combinatorics is an area where you felt rapid progress with AI was more possible, or is it that behind the scenes you were trying lots of open problems and this just happened to be the one that fell first?
我不认为组合学有什么特别之处。实际上我对组合学还有点看跌。但碰巧我们尝试了很多不同的问题,这个问题的答案是正确的。
I don't think there's anything special about combinatorics. I was actually a little bit bearish on combinatorics. But it just so happened that we tried a bunch of different problems and this one the answer came back as correct.
这个构造本身非常有趣,因为它是人类数学家很难想出来的那种东西。它非常复杂,与埃尔德什的原始论证有些相似之处,但复杂得多。你们发布了一篇配套文章,附有数学家的评论。Jacob Zimmerman 提到他之前尝试过类似的东西,但太复杂了——用他的话说,水域太危险了。这引发了一个问题:AI 系统是否会在与人类截然不同的方面更擅长数学?它们可能在不同维度上达到极限。也许即使人类数学家被要求寻找反例,也永远不会想到这个。你觉得模型在数学上比人类更强的方式是否与人类不同,而不是像人类但只是好一些?
The construction itself is very interesting because it's the kind of thing a human mathematician would have really struggled to come up with. It's very complicated, with some analogies to the original Erdős argument, but quite a lot more complex. You released a companion piece with commentary from mathematicians. Jacob Zimmerman remarked that he tried something like this a while ago, but it was too complicated—the waters were too treacherous. This raises the question of whether AI systems are going to be better at math in ways that are really different from humans. They might be maxed out in different dimensions. Maybe even a human mathematician given the prompt to find a counterexample might never come across this. Do you feel the models are better at math in a way that's different from humans, as opposed to being like a human but just some factor better?
这很有趣。模型已经达到了超越我理解它们能力范围的程度。理解这类问题的证明超出了我的能力,也超出了大多数人的能力。所以我现在很难说模型在数学上的弱点和强点是什么。但从我与更熟悉这个领域的人(包括 OpenAI 内部和外部数学家)的交流中,我感觉模型非常擅长结合数学的许多不同领域。在这方面,我认为它们目前已经超人类了。然后还有一个单独的问题,即通过复杂问题进行推理。可以说它们还没有超人类,但我们也看到了快速进展。所以我认为我们会看到这种尖峰智能,模型在某些方面超人类,在其他方面则不然。但随着时间的推移,随着模型能力越来越强,它们将变得全面更强,而这正是我们正在看到的。
It's interesting. The models have reached a point where they have surpassed my ability to really understand what they're capable of and what they excel at. Understanding the proofs of these kinds of problems is beyond my ability, and beyond most people's ability. So it's hard for me to say what the weak and strong points of the models are in math. But the sense I get from talking to people more familiar with the space, both at OpenAI and external mathematicians, is that the models are extremely good at combining a lot of different areas of mathematics. In that respect, I would say they are superhuman at this point. Then there's the separate question of reasoning through complex problems. It's arguable that they're not superhuman yet, but we're also seeing rapid progress. So I think we're going to see this spiky intelligence where the models are superhuman in some respects and not quite in others. But over time, as the models become more capable, they will become more capable across the board, and that's what we're seeing.
这不仅仅是它们越来越擅长吸收大量数学知识并将其组合起来。而是它们在这方面进步的同时,也在提升基本的推理能力。所以我认为短期内我们会看到更多这样的成果——以新颖的方式结合不同数学领域,做出以前没人想到的、非常酷的证明。但随着时间的推移,它将在几乎所有方面变得超人类。我不知道这个过渡期需要多久,可能是几年,也可能是十年,但我认为我们最终会达到那一步。
It's that it's not just like oh they're getting better and better at ingesting a lot of mathematics and combining it. It's that they're getting better at that but they're also getting better at just like the fundamental reasoning capabilities. So I think we will in the short term see more of these results where it's like combining different areas of mathematics in novel ways that nobody thought to do before and then having it be like really cool proofs in this kind of way but over time it will just become superhuman in basically every respect. And I don't know how long that transition period is going to take. It could be a couple of years, it could be 10 years, but I think we'll get there eventually.
对。我好奇的一点是,你是否推测存在一个半人马时期,即人类辅助数学的阶段?还是你认为这个时期太短,长远来看没什么意思?想听听你的看法。
Yeah. What do you... That is one thing I was curious about is if you speculate that there is this kind of centaur period for human-aided math or if you think that period is too short to be interesting in the long run. Curious what you think about that.
是的,我认为肯定会有半人马时期。经典的类比是国际象棋和围棋。1997 年加里·卡斯帕罗夫输给深蓝,之后有大约十年时间,国际象棋 AI 在某些方面更强,但人类加 AI 比单独的人类或单独的 AI 都强。而现在,AI 在国际象棋上太强了,人类几乎没什么可贡献的。所以很明显,已经超出了人类能增添价值的范围。我认为数学也会类似:AI 在某些方面极其有效,我们可能会达到一个点,如果你必须在人类数学家和 AI 数学家之间选择,你会选 AI 数学家。但即使到了那个点,你也可以说人类加 AI 比任何一方单独都更有效。
Yeah, I think there definitely will be a centaur period. The classic analogy, which I think is appropriate, is to things like chess and Go. In chess, you have Garry Kasparov losing to Deep Blue in 1997, and then there is this 10-year period where chess AIs were better in some ways, but human plus AI was actually better than human alone or AI alone. At this point, the AIs are so good at chess that the human doesn't really add anything. So it's clear that it is beyond the human's ability to add anything to the table. I think we'll see something similar with mathematics: the AIs are extremely effective in some ways, and we might reach a point where if you had a choice between a human mathematician or an AI mathematician, you would choose the AI mathematician. But even when we reach that point, you could argue that the human plus the AI would be more effective than either alone.
对。
Yeah.
这个时期有多长很难说,可能很短,也可能很长。我认为很大程度上取决于数学进步的轨迹。而且这也是与象棋和围棋类比失效的地方,因为象棋和围棋变得超人类是因为自我对弈,AI 能收集无限数据,与自己对抗并无限提升。但数学中没有类似的自我对弈,因为它不是两人零和博弈。所以能力提升可能需要更长时间。话虽如此,能力提升一直非常迅速,所以我不确定。这是一种权衡。我试图平衡两方面:一方面 AI 进步到目前为止极快,另一方面随着达到超人类能力,进步可能会放缓。
Now, how long that period is, it's hard to say. It could be very short, it could be long. I think a lot of it depends on the trajectory of mathematical progress. I think also this is where the analogy to things like chess and Go breaks down because chess and Go became superhuman because of self-play and the ability of the AI to collect infinite data and play against itself and get arbitrarily better. And we don't really have a similar version of self-play in mathematics because it's not a two-player zero-sum game. So it could take longer for the capabilities to improve. Now, that said, the capabilities have been improving very rapidly, so I don't know. It's a trade-off. I'm trying to balance the fact that AI progress has been extremely fast so far versus the fact that there is the possibility that progress slows down as you reach superhuman capabilities.
对,我听说过关于自我对弈的猜测,比如一个模型出题,另一个尝试解答。你可以对抗性地让问题更难,但可能很难奏效。这似乎比下围棋这种封闭环境要模糊得多。
Yeah, I've heard speculations around the self-play idea of, you know, one model proposes problems, the other tries to solve them. You can adversarially make the problems more difficult, but maybe it's very hard to get to work. It seems less clear how you would do that than in a very boxed setting like playing Go.
是的,我认为人们低估了两人零和博弈在自我对弈中的独特性。在两人零和博弈中,目标非常明确:收敛到最小最大策略。你只需要让 AI 互相博弈,不需要任何外部数据或人类数据。它们互相博弈,找出如何击败对方,然后从错误中学习,最终收敛到最小最大策略。一旦超出两人零和博弈,这个性质就不复存在了。我喜欢举的例子是最后通牒博弈。你了解最后通牒博弈吗?
Yeah, I think people underappreciate how unique two-player zero-sum games are when it comes to self-play. In these two-player zero-sum games, you have a very clear objective, which is to converge to a minimax policy. You just have the AIs play against each other. You don't need any external data. You don't need any human data. Just have the AIs play against each other and figure out how to beat each other and then learn from those mistakes and they will converge to the minimax policy. When you go outside of a two-player zero-sum game, you don't have this property anymore. An example I like to give is the ultimatum game. Are you familiar with the ultimatum game?
不了解。
No.
好的,最后通牒博弈是一个非两人零和博弈。玩家 A 有 100 美元,她必须选择给玩家 B 多少钱,可以给 0 到 100 美元之间。B 选择接受或拒绝。如果 B 拒绝,两人都得 0 美元。如果 B 接受,则按 A 的提议分钱。比如 A 出 20 美元,B 接受,则 A 得 80,B 得 20。如果 B 拒绝,两人都得 0。你可以想象自我对弈会收敛到什么结果。你猜会是什么?
Okay, so the ultimatum game is a non-two-player zero-sum game. One player, Alice, has $100 and they have to choose how much to give to Bob. They can offer anywhere between $0 and $100. And Bob chooses to reject or accept. If Bob rejects, then both players receive $0. If Bob accepts, then the money is split according to what Alice proposed. So, let's say Alice offers $20. Bob accepts, then Alice gets $80, Bob gets $20. If Bob rejects, they both get zero. So, you could imagine what self-play would converge to in this game. If you had to guess, what do you think it would converge to?
比如总是接受。
Like always accepting.
对,总是接受,然后 A 基本只给 1 美元或 1 美分。
Yeah, always accepting and then Alice offering basically like a dollar or a penny.
对。
Yeah.
如果换成人类,结果就不会这么好,对吧?如果我对你说,我有 100 美元,我给你 1 美分,你选择接受或拒绝。你会怎么反应?
If you were to do this with humans, it would not go very well, right? If I were to offer you, okay, I have 100 bucks, I'm going to offer you a penny. You have a choice of accepting or rejecting. How would you react?
对,我会……
Yeah, I would...
所以,一旦超出两人零和博弈,自我对弈收敛到优美、完美、客观正确策略的想法就不再成立。数学也会遇到类似问题。你可以说,我们在数学中做自我对弈,一个智能体出任意难的问题,另一个尝试解答。但问题在于,出题者可以出不可能的问题,或者出很难但对人类没意思的问题。比如 50 位数乘法。对没有计算器的大语言模型来说很难,但有意思吗?没意思。
So, once you go outside of two-player zero-sum games, this idea of self-play converging to a beautiful, elegant, perfect strategy that is objectively correct no longer holds. And you run into a similar problem with things like math. You could say, okay, we'll do self-play in mathematics where one agent is trying to propose arbitrarily difficult problems and the other agent is trying to solve them. The problem is the proposer agent could propose impossible problems, or they could propose problems that are difficult but not interesting to humans. For example, 50-digit multiplication. That's going to be difficult for an LLM to do if it doesn't have a calculator, but is it interesting? Not really.
对,很难规范化,对吧?即使有简短解法或已知解法的问题,也很难想出那种能推动训练朝正确方向前进的问题。
Yeah, it's very hard to regularize, right? Even things where there's a short solution or a known solution, it's hard to come up with problems that would push you in the right direction for training.
对。所以,这并不是说自我对弈行不通。
Yeah. So, it's not to say that self-play can't work.
我认为自我对弈将变得极其重要,它很可能会推动这些模型超越超人表现,达到难以想象的智能水平。但这并不像围棋和国际象棋这些两人零和游戏那样容易。这就是为什么我不清楚半人马时期会持续多久的主要原因,因为这取决于真正解锁自我对弈需要多长时间。
I think self-play is going to be extremely important, and it will probably be the thing that propels us past superhuman performance to unimaginable levels of intelligence for these models. But it's not as easy as it was for Go and chess, which are two-player zero-sum games. That's the main reason why it's unclear to me how long the Centaur period will be, because it depends on how long it takes to really unlock self-play in a proper way.
非常有趣。回到埃尔德什问题一会儿,你之前说了一件很有趣的事:模型已经变得如此之好,以至于非专家人类很难判断一个解决方案是否正确。我很好奇这个具体结果的微观细节。你们有自动化测试和内部数学家团队。但当你认为有了解决方案,到真正相信已经完全解决,这中间的过程是怎样的?
Super interesting. Coming back to the Erdős problem for a minute, you said something earlier that was really interesting: the models have gotten so good that it's hard for non-expert humans to tell whether a solution is correct. I was curious about the micro-level details of this specific result. You have automated testing and a team of in-house mathematicians. But what's the play-by-play of thinking you have a solution and then actually believing you've gotten all the way there?
基本上,我们训练了一个新模型。它不是为数学设计的,而是一个通用推理模型。OpenAI 有数学博士,有些是前数学教授,对数学感兴趣。他们作为副业,想看看这个新模型的能力。他们在许多未解决问题上运行模型,看它是否认为有解。一个挑战是验证这些结果。模型在大量问题上运行后,对于这个问题,它返回说已经解决了。然后有一个艰难的验证过程,因为需要多个数学领域的专业知识。有一个有趣的为期一周的时期,我们认为它正确但不确定。我们不得不去找外部数学家讨论这个候选证明是否有效。一旦确认正确,就面临该怎么做的问题。老实说,当他们告诉我模型证明了这一结果时,我不知道这是否了不起。我以前听说过几十个埃尔德什问题被解决。真正说服我和许多人的是,与外部数学家交谈后,他们表示这是一个非常重要的结果。这就是为什么我们强调获取外部数学家的反馈——我们说它重要是一回事,外部数学家说它重要是另一回事。
Basically, we trained a new model. It wasn't designed for math; it's a general-purpose reasoning model. We have people with math PhDs at OpenAI, some former math professors interested in math. They decided as a side thing to see where the capabilities of this new model were. They ran it on a bunch of unsolved problems to see if it thought any were correct. One challenge is verifying these things. So the model was run through many problems, and for this one it came back saying it was solved. Then there was a difficult process of verifying that it's correct, because you need expertise in several areas of math. There was an interesting week-long period where we thought it was correct but weren't sure. We had to go to external mathematicians to discuss if the candidate proof was valid. Once it became clear it was correct, there was a question of what to do. Honestly, when they told me the model proved this result, I didn't know if it was a big deal. I'd heard of dozens of Erdős problems being solved before. What convinced me and many others was talking to external mathematicians who said it was a very significant result. That's why we emphasized getting feedback from external mathematicians—it's one thing for us to say it's significant, another for them to say it.
是的,我同意。我记得本科时就听说过这个结果。这是数学家花大量时间研究的那种问题,不是一个缺乏关注的领域。然而,这种方法可能受到的关注很少。AI 模型的一个有趣之处在于,你可以在这些问题上投入远超人类社会能力的注意力和努力。数学的许多领域可能只有少数人在最前沿工作。很容易想象在不久的将来,按在这些问题上花费的时间衡量,大部分努力来自 AI 系统而非人类数学家。我想知道仅凭这一点是否就能推动大量发现。
Yeah, I agree. I remember hearing about this result as an undergrad. It's the kind of thing mathematicians spend a lot of time on, not an area that's been attention-starved. Yet this approach has probably received very little attention. One interesting thing about AI models is you could potentially pour a level of attention and effort into these problems that far exceeds human society's capacity. Many areas in math might only have a couple of people working at the absolute forefront. It's easy to imagine a near-term future where the vast majority of effort, measured by hours spent on these problems, comes from AI systems rather than human mathematicians. I wonder if that alone should propel a lot of discovery.
我认为可能是这样。但我想强调一点,我们并没有在解决这个问题上投入大量精力。最终使用的算力并不多。我们后来展示了一些图,x 轴是测试时算力,y 轴是解决埃尔德什问题的概率。为了说明这一点,为了制作那张图,我们必须在每个数据点上对问题运行模型 100 次。所以如果每个数据点都运行 100 次,你可以感觉到这并非最昂贵的事情。那张图的一个很酷的地方是,随着你投入更多测试时算力,解决问题的概率显著上升。我的结论是,可能还有很多其他问题这个模型可以解决,只是我们还没有投入必要的算力。
I think that could be the case. But one thing I want to emphasize is that we did not put a lot of effort into solving this problem. It was not a lot of compute at the end of the day. We later showed plots where the x-axis is the amount of test-time compute and the y-axis is the probability of solving the Erdős problem. To put that in perspective, to make that plot, we had to run the model on the problem 100 times for every single data point. So if we're doing 100 runs for every data point, you can get a sense it's not the most expensive thing. One cool thing from that plot was that as you put more test-time compute into solving the problem, the probability of solving it goes up significantly. My takeaway is that there are probably a bunch of other problems that this model can solve, and we just haven't put the necessary amount of compute into them.
鉴于这一结果,你对哪些领域感到乐观?你应该更新信念,认为真正的、数学年鉴级别的全新发现是可能的。你有了这个新工具。你想把它指向哪里?
Are there any areas you're optimistic about, given this result? You should have a belief update that true, annals-of-mathematics-level novel discoveries are possible. You have this new instrument. Where do you want to point it?
这是个好问题。事实是,在 OpenAI,我们有巨大的杠杆来使模型更有效。我们处于一个奇怪的位置,比其他人早几个月就能接触到这些前沿模型。诱惑在于,遍历所有未解决的数学问题,在其上运行模型,看看哪些能被证明,把所有时间都花在这上面,因为我们可以获得令人难以置信的结果。
It's a good question. The truth is, at OpenAI we have a huge amount of leverage to make the models more effective. We're in this weird spot where we have access to these frontier models a few months before anyone else. It's tempting to go through all unsolved mathematical problems, run the model on them, and see what can be proven, spending all our time doing that because we can get incredible results.
但我不认为那实际上是我们的最高杠杆。我认为我们的最高杠杆是让模型更好、更强大、更安全,尽快将它们推向世界,并让数学家能够使用这些模型来解决所有未解难题。所以,对我们来说,尝试解决所有这些问题确实很诱人,我实际上会鼓励我在 OpenAI 合作的研究人员不要走那条路,因为我认为最重要的是专注于改进模型本身。而这正是我最看好的。这次成果并非数学专用,模型或算法也不是特制的,它只是一个通用模型。它的能力不限于数学。所以,我认为它在做机器学习研究、提高模型效率以及让我们实现递归自我改进方面也会非常有效。
But I don't think that is actually the highest leverage for us. I think the highest leverage for us is making the models better, more powerful, safer, getting them out to the world as quickly as possible, and enabling mathematicians to use these models to solve all of these unsolved problems. So, for us, it's really tempting to try to solve all these problems, and I actually try to discourage researchers that I work with at OpenAI from going down that route because I think that the most important thing is for us to focus on improving the models themselves. And that's actually what I'm most bullish about. The fact that this was not a math-specific result, that this was not a math-specific model or algorithm or anything, it was just a general-purpose model. Its capabilities are not limited to mathematics. So, I think it's also going to be very effective at doing things like machine learning research and making the models more effective and enabling us to basically do recursive self-improvement.
是的,我记得你是我最早交谈过的人之一,真正相信用通用推理模型来做这类事情的想法。比如竞赛数学,大概一年前有一场大辩论,你知道,你是否必须走那种 Lean 形式化的路径?你能用通用推理模型吗?如果我没记错的话,你们用于所有竞赛的模型,比如 IMO、IOI、ICPC,实际上是同一个模型。
Yeah, I remember you were one of the first people that I talked to that really believed in this idea of using the general purpose reasoning models for these types of things. For contest math, for example, I think like a year ago there was this big debate of, you know, do you have to go through the sort of Lean formalization path? Can you use a general purpose reasoning model? If I remember correctly, the model that you guys used for all of the contests, like IMO, IOI, ICPC, was effectively the same model.
基本上是同一个模型。我认为对于 ICPC 和 IOI,它是由几个模型搭建的,但最重要的模型与 IMO 中使用的模型相同。
Basically the same model. I think for ICPC and IOI it was like a scaffold of a few models, but the most important model was like the same model as was used in the IMO.
那么为什么这对你如此重要?是这种希望这些改进能普遍用于递归自我改进之类的事情的感觉吗?是因为你关心从这些竞赛领域到每天数亿人使用的模型的迁移吗?你为什么这么早就做出了那个决定?
And so why was that so important to you? Is this feeling of like I want these improvements to be generally useful for things like recursive self-improvement? Is it that you cared about transfer from these contest domains to the model that hundreds of millions of people use every day? Why were you so early to make that decision?
对我来说,始终是那些极具影响力的事情:真正通用的系统,并试图将问题分解到实际所需的最低限度,然后找出什么可以规模化。尤其是在这个领域进展如此之快的时候,很容易被“技术宅狙击”而去做那些非常擅长解决 IOI 等问题的定制模型,但最终在六个月后就没用了。我认为我们必须认识到有一个目标,那就是 AGI 或超级智能,随便你怎么称呼,我们不想在这条路上绕弯子。所以,我认为做像 IMO 这样的事情作为通往 AGI 路上的里程碑很好,但我不认为我们想被它过度分心。
For me, it was always about what was extremely impactful: the really general purpose systems, and trying to break things down to the actual bare minimum that's needed, and then figuring out what could be scaled. Especially at a time when the field is progressing so quickly, it's really easy to get nerd sniped into making custom-built models that are very, very good at solving the IOI or whatever, but ultimately will not be useful 6 months down the road. I think it's important for us to recognize that there is a goal, and that goal is AGI or superintelligence or whatever you want to call it, and we don't want to take detours from that route. So, I think it's great to do things like the IMO as a milestone along the way to AGI, but I don't think we want to get sidetracked by it too much.
对。它应该像是攀登曲线过程中的副产品,而不是投入资源只是为了占领那个位置。
Right. It should be like a byproduct of ascending the curve as opposed to putting resources just to claim the post.
是的。这是一种持续的张力,因为获得这些令人印象深刻的结果并向世界展示、获得认可,这种诱惑太大了,我认为我们必须抵制这种诱惑。
Yeah. And it is a constant tension because there's such a temptation to get these really impressive results to present to the world and get recognition, and we have to resist that temptation, I think.
那么,如果路径是将模型发布给人类数学家,想法是他们将花时间摆弄这些模型并产生结果。我好奇的一件事是,谁会非常擅长这个,以及这是否与当前的手动数学家群体有所不同。这里的类比是理查德·汉明,他是物理学中模拟的早期倡导者。他对贝尔实验室的总裁说,今天 90%的实验在物理世界中进行,10%在计算机上。这很快就会转变为 90%在计算机上,也许 10%在物理现实中。然后他描述了在模拟领域非常擅长的人不一定是在实验领域非常擅长的人。那么,擅长用这些模型证明新数学的人,是否与擅长手动做数学的人不同?我们应该期待令人印象深刻的新结果来自菲尔兹奖得主,还是来自非常擅长使用这些模型的人,后者可能是不同的一群人?
So, if the path is releasing the models to human mathematicians, and the idea being that they'd be the ones that spend time playing with them and producing results. One thing I wondered about is who's going to be really good at that and if it's different in any ways than the current set of manual mathematicians. The analogy here is Richard Hamming, who was an early proponent of simulations in physics. He talked to the president of Bell Labs and said, today 90% of experiments are done in the physical world and 10% on the computer. That's going to switch immediately to 90% on the computer and maybe 10% in physical reality. Then he described how the people who are really good in the simulation regime are not necessarily the same people who are really good in the experimental regime. Are the people who will be really good at proving new math with these models different than the people who are really good at doing math manually? Should we expect impressive new results to come from Fields Medalists or from people who are really good at using these models, which might be a different set of people?
我认为可能是不同的一群人。我认为新技术总是如此,需要适应并以非常不同的方式处理问题。还有一个问题是,如果模型在某些方面极其出色,而在其他方面低于人类水平,那么最成功的人将是那些与模型互补的人。所以这种特征可能与今天成为伟大数学家的条件非常不同。这有点难说,因为这取决于 AI 模型将在哪些方面强、哪些方面弱,以及今天什么造就了伟大的数学家,但我认为最成功的人将是那些与 AI 模型互补得很好的人。
I think it's possible that it is a different set of people. I think this is always the case with new technology where it requires adapting and approaching things very differently. And it's also a question of if the models are extremely good in some respects and below human performance in other respects, then the people that are going to be most successful are the people who are the complements to the models. So that profile could be very different from what makes a great mathematician today. It's a little hard to say because it depends on where the AI models are going to be strong and where they're weak and what makes a great mathematician today, but I think the people that will be most successful will be the people that complement the AI models really well.
你看到任何早期迹象了吗?比如我能想象的一件事是数学已经变得非常专业化。我记得我本科时,我问过一个关于椭圆曲线的非常重要的结果,我说如果我花时间试图理解它,我能做到吗?一位著名的数学教授说不行,因为他作为五年级研究生花了六个月试图理解,但感觉并没有真正弄懂。它太专业了。
And do you see any early hints of what that might be? Like one thing I can imagine is that math has become very specialized. I remember when I was an undergrad, I asked about this really important result about elliptic curves and I said if I really spend time trying to understand it could I? And this well-known math professor said no, because he spent like six months trying to do it as a fifth year grad student and he didn't feel like he understood it. It's just so specialized.
我能想象的一种情况是,通才可能再次迎来高光时刻,连接许多不同领域很有价值。另一种情况则完全不同,更像是凭直觉判断模型擅长什么,并指出正确的问题。我很好奇,你是否有早期迹象,知道谁会是新的数学家,这些互补技能在哪里?
One thing I can imagine is that maybe the generalists have another day in the sun where connecting a lot of these different areas is valuable. Another thing I can imagine is totally different, more about having intuition for where the models are good and pointing at the right problems. I'm just curious if you have any early hints of who the new mathematicians are, where those complementary skills are.
根据我个人使用这些模型进行 AI 研究的经验,它们的研究品味往往不太好。所以对我来说,这其实很棒。我在 IC 类工作上还算可以,但从不突出。现在我觉得自己与这些模型形成了很好的互补,效率高了很多。如果数学领域也类似,我不会感到惊讶——提出正确问题、知道该研究什么的能力会变得最有价值,因为模型目前可能不太擅长这个。我可以想象,非常擅长这一点的人会与数学 AI 模型形成良好互补。
In my personal experience using these models to do AI research, they tend not to have very good research taste. So for me, this has actually been amazing. I was okay at the IC kind of work, but never exceptional. Now I feel like I'm such a great complement to these models, able to be much more productive. I wouldn't be surprised if it's similar in mathematics, where the ability to ask the right questions and know what to investigate becomes the most valuable thing, because the models probably aren't very good at that right now. I could see people extremely good at that being a good complement to AI models in math.
所以也许有人研究品味好,擅长构建纲领,比如研究朗兰兹纲领,或者像格罗滕迪克那样,有一个更大的纲领,知道该走哪条路。
So maybe somebody who has good research taste and is good at program construction, like working on Langlands or being Grothendieck style, having a larger program that is the right path to go down.
嗯。
Mhm.
有意思。我好奇的是,这种构造与针对埃尔德什问题的另一种论证之间的区别。在你的合作文章中,菲尔兹奖得主蒂莫西·高尔斯说,起初他以为 OpenAI 证明了一个更紧的上界,之后他失眠了,觉得如果那是真的,我们就全完了。后来他得知那是一个反例构造,推翻了埃尔德什的上界,他觉得这更合理,像是你能想象 AI 模型想出来的东西。我好奇你是否区分这两者,是否认为这是一个真正的哲学区别。
Interesting. One thing I was curious about is this kind of construction versus a different kind of argument for the Erdős problem specifically. In your companion piece, the Fields Medalist Timothy Gowers said that at first he thought OpenAI had proved a much tighter upper bound, and he lost a lot of sleep after that, thinking if that's true then we're all completely cooked. Then when he learned it was a construction of a counterexample that disproved Erdős's upper bound, that felt more plausible, like the kind of thing you could imagine an AI model coming up with. I'm curious if you draw that distinction, if you view that as a real philosophical distinction.
我认为,撇开埃尔德什问题不谈,如果让我预测 AI 在数学领域的进展,我想会是这样的:先拿到 IMO 金牌,然后开始证明一些不太重要或研究者不太关注的未解决问题。接着证明一个重要的结果,但方式不会让整个数学领域失眠,不会让人觉得在各方面都超越人类能力。然后进展继续,最终达到数学家们确实失眠的地步。我认为这种进展是自然的,我们只是看到了进展中的又一个点。这就是我对当前能力水平的预期。一年后,能力会高得多,如果它做出今天数学家无法想象的事情,我也不会惊讶。
I think, separate from the Erdős problem, if I were to predict what the progress of AI in mathematics would look like, I think it would be: you get something like IMO gold, then you start proving some unsolved problems that aren't that significant or that researchers haven't paid much attention to. Then you get to something like proving a significant result, but in a way that doesn't cause a field of mathematicians to lose sleep, saying this is beyond human ability in every respect. Then progress continues, and eventually you get to that point where mathematicians do lose sleep. I think it's natural to expect this progression, and we're just seeing another point along that progression. That's what I would expect from this level of capability. A year from now, the capability will be much higher, and I wouldn't be surprised if it's doing things unimaginable to mathematicians today.
嗯。
Mhm.
是的,很多数学家的评论已经表明,这是一种改进原始构造的深刻方式。原始的埃尔德什构造相对初等,改进它需要来自代数数论的许多不同思想。能把所有这些都记在脑子里并花大量时间研究这个问题的人非常少。所以 AI 系统可能比人类有另一个重大优势:现有研究文献浩瀚,大多数专业数学家只知道其中一小部分。可能有很多问题,如果你把五篇论文放在他们面前,说读完这些,不离开这个房间直到证明出结果,他们能取得进展,但你无法对许多不同领域都这样做。我想知道我们可能非常接近实现的一次性红利,就是把来自非常不同研究领域的想法结合起来,产生新东西。也许这是更近期的领域之一。我很好奇你怎么看。
Yeah, it already feels like a lot of the mathematician commentary is that this is a deep way of improving the original construction. The original Erdős construction is relatively elementary. To improve it requires many different ideas from algebraic number theory. The number of people who have all those things in their head and can spend serious time on the problem is very small. So another way AI systems might have a major advantage over humans is that the existing research literature is vast, most professional mathematicians only know a small fraction. There are probably lots of problems where if you put five papers in front of them and said read these and don't leave this room until you've proved a result, you could make progress, but you can't do that for many different areas. I wonder about this one-time dividend we might be very close to achieving, just putting together ideas from very different research areas to produce something new. Maybe that's one of the things that's more near field. I'd be curious what you think.
我确实认为短期内很多这类突破会来自这里。模型在如此多的不同数据源、如此多的不同论文上训练,它能够连接那些不可能在每个领域都深入跟进的想法。所以这已经是模型在一段时间内明显超越人类的地方了。
I do think that's going to be short-term where a lot of these breakthroughs come from. The fact that the model is trained on so many different data sources, so many different papers, it's able to connect ideas that are impossible to keep up with in an in-depth way across every single field. So this is already something where it's pretty clear the models have been superhuman for a while.
嗯。
Mhm.
所以我认为这是起点,并且进展会从此继续。
And so I do think that this is where it begins, and I think it just progresses from there.
也许你是在提升一个层次。你认为数学为什么是一个特别好的领域?我会把它放在 AI for Science 的大框架下,用 AI 模型做出新的科学发现。我很好奇你对数学是否是一个特别肥沃的领域的看法,以及你对哪些其他领域感到乐观。
Maybe you're kind of bumping one level up. What are the reasons why you think math is a particularly good domain? I'd put it inside of AI for science writ large, where making new scientific discovery with AI models. I'd be curious for your take on if math is a particularly fertile domain for this, and what other domains you're optimistic about.
我认为数学是一个特别好的领域,因为它非常自洽。你可以严格验证结果,而且有明确的目标。我乐观的其他领域包括蛋白质折叠,比如 AlphaFold,以及材料科学,AI 可以发现具有所需特性的新材料。此外,药物发现和化学等领域也很有前景。
I think math is a particularly good domain because it's very self-contained. You can verify results rigorously, and there's a clear objective. Other domains I'm optimistic about include protein folding, like AlphaFold, and materials science, where AI can discover new materials with desired properties. Also, areas like drug discovery and chemistry are promising.
是在物理学内部,还是物理学的某个特定分支?也许模型在粒子物理问题上特别擅长。我很好奇你最乐观的领域是哪里。
Is it inside of physics or a particular branch of physics? Maybe the models are very good at particle physics problems or something. Curious where you're most optimistic.
我认为数学很特别,因为它纯粹受限于推理能力。在物理学中,我的印象——我不是物理学家,所以不确定,但和物理学家交流时,感觉这个领域很大程度上受限于实验结果。你可以提出各种疯狂的理论,但最终需要投入资金进行实际物理实验来验证或获取更多数据。数学没有这个问题。你只需坐在房间里长时间深入思考,就能想出惊人的东西。所以我认为短期内,最大的收益将来自那些受限于纯推理能力而非物理实验或实验数据的领域。
I think math is special because it is purely bottlenecked by reasoning. In physics, my impression—I'm not a physicist, so I don't know for sure, but talking to physicists, it sounds like a lot of the field is bottlenecked by experimental results at this point. You could make all sorts of crazy theories, but ultimately you need to put money into actual physical experiments to validate them or get more data. You don't have that problem with mathematics. You can just sit in a room and think hard for a long time and come up with something amazing. So I think we'll see the biggest benefits in the short term from domains bottlenecked by pure reasoning ability, not by physical experiments or experimental data.
嗯。
Mhm.
所以,例如湿实验室会是一个瓶颈。在物理学中,收集实验数据可能是一个瓶颈。数学不会有这个问题。问题在于还有哪些领域也没有这个瓶颈。我实际上认为,在某些方面,比如 AI 研究——这还有争议。一方面,很多研究受限于拥有大量 GPU 和运行大规模实验。另一方面,用小型实验和资源也能做很多事。所以总体而言,我对 AI 模型在 AI 研究本身取得进展持乐观态度。
So wet labs, for example, are going to be a bottleneck. In physics, collecting experimental data could be a bottleneck. Mathematics won't have that. The question is what other fields also don't have that bottleneck. I actually think in some ways, things like AI research—it's debatable. On one hand, a lot of research is bottlenecked by having a ton of GPUs and running large-scale experiments. On the other hand, there is a lot you can do with small-scale experiments and resources. So I am overall pretty bullish on the ability of AI models to make progress on actual AI research.
你觉得我们距离下一届 IMO 大约还有两个月?你认为今年会是饱和的一年吗?去年分数已经是 42 分中的 34 分,明显是金牌水平。我可以想象很快 IMO 就会变得不那么有趣,因为你预期每个问题都能拿满分。我很好奇你认为竞赛数学中有趣的前沿是什么,或者你是否认为我们现在必须转向数学研究。
Do you think we're about two months out from the next IMO? Do you view this as the year that saturates? Last year scores were already 34 out of 42, clearly gold medal level. I could imagine soon getting to the point where the IMO becomes less interesting because you expect to max out on every problem. I'm curious what you view as the interesting frontier in contest math, or if you think we have to go to math research now.
我记得去年我们获得 IMO 金牌时,另一个实验室的人发短信问我:‘你觉得明年能拿满分吗?’我告诉他,如果明年我们还没有一个任何人都能用来在 IMO 上拿满分的模型,我会很失望。如果我们最新的内部模型拿不到满分,我会感到惊讶。如果现在发布的模型拿不到满分,我也会失望。所以我认为,现在或很快,竞赛数学或竞赛编程将不再有趣。真正的前沿是实际未解决的问题——在现实世界中做真正的研究。
I remember when we got the IMO gold last year, someone from another lab texted me and asked, 'So, do you think you'll get a perfect score next year?' I told him I'd be disappointed if we don't have a model out by next year that anybody can use to get a perfect score on the IMO. I would be surprised if we don't get a perfect score with our latest internal models. I would also be disappointed if we don't get a perfect score with the released models at this point. So I don't think competition math or competition coding will be interesting anymore either now or pretty soon. The frontier really is actual unsolved problems—doing real research in the real world.
是啊,也许正确的评估就是能否把问题 PDF 上传到公开的 GPT-5.5 然后拿满分。也许这就是今年有趣的事情。
Yeah, maybe the right evaluation is just whether you can upload a PDF of problems to the publicly available GPT-5.5 and max out. Maybe that's the interesting thing this year.
是的。我还想说,我们从未真正提过,但我们的模型已经能解决去年 IMO 的第六题有一段时间了。我们决定不宣传,因为我们觉得这个领域已经超越了竞赛数学,但我们已经能在第六题上拿到满分好几个月了。
Yeah. I should also say, we never really mention it, but our models have been able to solve problem six from the IMO last year for a while now. We decided not to advertise it because we felt the field has already progressed beyond competition math, but we've been able to get perfect scores on P6 for several months.
我正想问这个,因为历史上第六题是最难的题目之一。去年在人类参赛者中,可能只有六个人得了满分。我很好奇第六题是否有特别之处让 AI 系统更难解,还是说它对人类和 AI 都同样难。我记得那是一个组合铺砌问题。我听说过一种观点,认为几何或组合几何问题特别难,因为人类有视觉直觉。我很好奇你是否认为第六题有什么特别之处。
I was going to ask about that because historically problem six is one of the harder problems. Last year among human participants, maybe only six people got full marks. I was curious if there was something specific about problem six that made it harder for AI systems, or if it's just uniformly harder for humans and AI. I remember it was a combinatorial tiling problem. I've heard the argument that geometry or combinatorial geometry problems are particularly hard because humans have visual intuition. I'm curious if you think there's anything special about problem six.
这是个好问题。我觉得有趣的是,对人类和 AI 来说,难度之间存在很大的相关性。所以这是一个主要因素——对人类来说很难的问题,对 AI 模型来说也很难,这并不太令人惊讶。我也认为问题的性质很重要;模型在几何和几何理解方面稍显落后。但模型全面进步,所以最终它们突然能解决这个问题并不令人意外。
It's a good question. I think it's interesting how much correlation there is between difficulty for humans and difficulty for AIs. So that is a major factor—it was a really hard problem for humans, so it's not too surprising it's hard for AI models. I also think the nature of the problem matters; models lag a bit when it comes to geometry and geometric understanding. But models just get better across the board, so it wasn't surprising that eventually they were suddenly able to solve it.
你是否觉得竞赛数学和研究数学之间存在本质上的鸿沟?竞赛问题必须是自包含的,人类大约一小时内能解出,而且往往偏向某些已知技巧,比如鸽巢原理。我很好奇,在 AI 研究方面,让模型处理竞赛数学和开放研究问题之间是否有本质区别。
Do you feel there's a different-in-kind chasm between contest math and research math? Contest problems have to be self-contained, solvable by a human in about an hour, and they often favor certain known techniques like the pigeonhole principle. I'm curious if you felt any difference in kind on the AI research side between pointing the model at contest math versus open research problems.
我认为存在时间跨度的差异。IMO 给你一个半小时解决一个问题。研究问题的时间跨度要长得多。所以关键区别在于时间尺度和更深入探索的需求。
I think there is a difference in horizon. The IMO gives you an hour and a half to solve a problem. Research problems have a much longer horizon. So the key difference is the time scale and the need for deeper exploration.
我的同事 Alex Way 指出的一个事情,我当时没意识到但事后觉得确实如此,就是如果你看 AI 模型在数学上的进展,它大致遵循这样的轨迹:模型越来越擅长解决需要人类数学家花更长时间的问题。所以,想想 2023 年我们在做 GSM8K,而 GSM8K 人类数学家大概 5 秒就能解出来。
And one of the things that my colleague Alex Way pointed out, which I didn't realize at the time but I think is really true in retrospect, is that if you look at the progression of AI models when it comes to math, it kind of follows this trajectory of the model getting better at doing problems that would take a human mathematician longer and longer. So, if you think in 2023 we were doing GSM8K. And GSM8K takes a human mathematician like 5 seconds.
嗯。
Mhm.
然后 2024 年我们在做 MATH。那些问题是高中水平,可能大学低年级水平,人类数学家大概需要一分钟。然后到了 AIME,人类数学家可能需要 10 分钟。然后到了 IMO,人类数学家需要大约 100 分钟,一个半小时。基本上每年你都在这个方面进步一个数量级,就像人类数学家解题所需的时间跨度。有理由认为这个轨迹可能不会持续,但到目前为止它一直相当稳定地持续着。所以竞赛数学和研究数学的区别在于,研究数学的时间跨度要长得多。
And then in 2024 we're doing MATH. And those problems are high school level, maybe early college level, and that takes a human mathematician maybe a minute. And then you get to the AIME that could take a human mathematician maybe 10 minutes. And then you get to the IMO and it takes a human mathematician about 100 minutes, about an hour and a half. And you can just basically every year you progress like an order of magnitude in this respect. Like the horizon how long it would take a human mathematician to do it. And there's reasons to think maybe that trajectory doesn't continue but it's been continuing pretty regularly so far. And so the difference between competition math and research math is that research math is over a much longer horizon.
嗯。
Mhm.
但同样,模型最终能够完成人类数学家需要 15 小时或一周才能完成的事情,这并不令人惊讶,只是模型需要变得更强大。还有一个区别是,IMO 在很多方面对 AI 模型来说是 adversarial 的困难,因为我们讨论过这些 AI 模型有不同的优势和劣势。比如它们非常擅长了解数学的许多不同领域,并能结合不同领域的结果。而另一种能力是坐下来长时间思考一个难题并推理出来。在这方面它们擅长,也在变得更好,但我还不会说它们在这方面超人类。而 IMO 特别强调第二部分。它真的不是关于了解很多不同领域的数学,而是关于坐下来长时间推理一个非常困难的问题。在这方面,模型在 IMO 获得金牌其实相当有趣。因此,我认为不到一年后我们开始看到模型证明数学家无法证明的结果并不令人惊讶,因为我们已经知道,一年前它在推理难题的能力上已经达到 IMO 金牌水平,而且它可能已经更擅长了解许多不同领域的数学并综合它们的结果。所以在这方面,能够做研究数学这么快发生可能并不太令人惊讶。
But it's also, I guess, not surprising that the models eventually are able to do things that would take a human mathematician 15 hours or a week, and it's just the models have to become more capable. There is also a difference in that in many ways the IMO was adversarially hard for AI models because, you know, we discussed there are different strengths and weaknesses of these AI models. Like they're really good at knowing a bunch of different areas of mathematics and being able to combine results from different fields. And then there's the being able to sit and work on a difficult problem for a long period of time and just reason through it. Where they're good, they're getting better, but I wouldn't say that they're superhuman in this respect yet. And the IMO really emphasizes the second part. It's really about not knowing a bunch of different areas of mathematics, it's about sitting down and reasoning through a very difficult problem for a long time. And in that respect it's actually pretty interesting that the model was getting a gold medal at the IMO. And for that reason I think it's not surprising that less than a year later we're getting the starts of proving results that mathematicians cannot prove because we already know, okay, it was at IMO gold level in the ability to reason through hard problems a year ago, and it was probably already better at knowing a bunch of different areas of mathematics and combining results between them. So in that respect, being able to do research math probably isn't too surprising that it happened so quickly.
是的,就像在竞赛环境中,你无法利用模型相对于人类最具比较优势的那些方面。好吧,也许稍微推测一下,五年后人类数学家会做什么?工作如何变化?你想象人们会如何花时间?是像提出新的研究计划,然后让模型在后台运行六小时看看是否有结果?还是别的什么?我很好奇你会推测人类数学家在你认为可预测的时间范围内会做什么。
Yeah, it's like you're sort of not able to use the things the models are best where they have the most comparative advantage relative to humans in the contest setting. And well, okay, maybe speculating a little bit, what are human mathematicians going to be doing in five years? Like how does the job change? How do you imagine people will be spending their time? Is it things like coming up with new research programs and then just letting the model crank in the background to see if there's a there there over the next six hours? Or is it something different? I'm curious what you would speculate human mathematicians are doing in whatever time horizon you think is predictable.
是的,五年现在太遥远了,我不知道五年后的世界是什么样子。我无法推测五年后任何事情的样子,因为进展太快了。
Yeah, five years is so far down the road at this point that it's like I don't know what the world looks like in five years. I can't speculate what anything is like in five years. Just because progress is so fast.
嗯。
Yeah.
我认为一两年后,很多数学会像今天的编程一样,大量工作是与模型合作。模型做大部分工作,你有点像在引导它们,指向正确的方向,挑战它们,并与它们一起得到这些结果。在很多方面,我有点惊讶这没有首先发生在数学领域。实际上我本来预期会这样。但我认为它离编程的现状并不远。
I think a year or two from now, I think a lot of mathematics looks like the way coding is today, where it's a lot of working with the models. The models are doing a lot of the work and you're kind of steering them and pointing them in the right direction and challenging them and working with them to do these kinds of results. In many ways, I'm kind of surprised that it didn't happen in mathematics first. That's what I would have expected actually. But I don't think it's far behind where coding is today.
接下来会发生什么?你认为会不会有限度地向人类数学家发布模型,以便在这些领域取得进展?我很好奇你怎么看。对我来说,OpenAI 专注于让模型变得更好完全合理。但如果你只看数学这个领域,作为研究社区,我们明年如何快速进步?接下来会发生什么?你认为下一步是什么?
What happens next? Do you think maybe there's some limited release of the model to human mathematicians to make progress in these areas? I'm curious how you think about it. It makes total sense to me that OpenAI's focus would be on making the models better. But if you were looking just at math as a domain and asking, as a research community, how are we going to make fast progress here over the next year? What's about to happen? What do you think are the next steps?
在 IMO 结果出来后,我们确实讨论过是否应该提前向数学家提供这个模型。最终,这只是一个带宽和进行那种定制部署的难度问题。而且,进展太快了。我们希望尽快把模型发布给所有人。
We after the IMO result, we actually did discuss should we make this model available in advance to mathematicians. And ultimately, it was just a matter of the bandwidth and the difficulty involved in doing that kind of custom deployment. And also, I mean, progress is so fast. We want to get the model out to everybody.
嗯。
Yeah.
尽快。所以,为特定群体加速会引入很多复杂性,我认为不值得。我认为最好把所有这些努力都集中在尽快把模型发布给所有人上。所以,我不认为我们会做特殊案例,但大部分情况下我们真的专注于尽快把模型发布给所有人,让所有人都能用它来证明各种疯狂的结果。
As quickly as possible. And so, fast-tracking it for a particular group introduces a lot of complexity that I just don't think is really worth the overhead. I think it's better for us to just focus all that effort into getting the model out for everybody as fast as possible. So, I don't think maybe we'll do some special cases, but I think for the most part we're just really focused on getting the model out to everybody as quickly as possible and allowing everybody to use this thing to prove all sorts of crazy results.
太棒了。好吧,我们时间快到了。有没有什么我们没谈到但你想谈的?你觉得有什么重要的需要理解或者想引起注意的吗?
Awesome. Well, we're almost at time. Is there anything that we didn't talk about that you wanted to talk about? Anything that you think it's important to understand or that you just want to call attention to?
你知道,自从这个结果出来后我一直在想的一件事是,证明的水平我达不到,我的意思是,也许如果我花很多时间去理解,我能理解,但我没花那么多时间。而且我肯定有很多人无法理解这类结果的证明,更不用说自己去证明了。
You know, one of the things I've been thinking about ever since this result is that the proof is at a level that I'm not, I mean, maybe if I really spent a lot of time trying to understand it, I would be able to understand it, but I haven't really spent that much time. And I'm sure there's a lot of people that can't understand the proofs for these kinds of results, let alone proving it themselves.
而且我认为,随着模型在许多方面变得超人类,我们将面临这样一个问题:证明本身对人类数学家来说将变得过于难以验证。
And I think as models become superhuman in many respects, we're going to have this problem where the proofs themselves are going to be too difficult for human mathematicians to verify.
嗯。
Mhm.
我认为这本质上是一个经典的可扩展监督问题。当证明本身超出了人类真正能理解的范围时,你如何让模型向人类数学家证明这个证明是正确的?
And I think this is a classic problem that's basically scalable oversight. How do you get the models to prove that the proof is correct to human mathematicians when the proof itself is beyond what humans can really comprehend?
嗯。
Mhm.
我很兴奋数学能成为这类可扩展监督问题的先锋。我认为这将是一个有趣的问题,我希望我们能解决它,因为我觉得它对许多其他领域也很重要。
I think I'm excited for mathematics to be like the vanguard for this kind of question of how do you do scalable oversight. And I think it's going to be an interesting problem, and one that I hope we can figure out, because I think it's going to be important for a lot of other domains as well.
是的。不,谢谢你做这个。这太棒了。
Yeah. No, thanks for doing this. This is great.
非常棒。
It's been great.