OpenAI's IMO Gold: The Trio Behind the Breakthrough
打开互动全文版(中英对照 + 朗读 + 问答)→Alex、Cheryl 和 Noam 讨论他们如何以小型团队和新颖的测试时计算扩展技术在国际数学奥林匹克竞赛中获得金牌。
Alex, Cheryl, and Noam discuss how they achieved gold at the International Math Olympiad with a small team and a novel technique for scaling test-time compute.
进步的速度真的令人惊叹。我认为在数学上这一点尤其明显。Alex 发过推文:就在几年前,这些模型还在为小学数学题挣扎。我记得即使在 2024 年,GSM8K 还是每个人发布模型时的标准评估,然后数学评估只流行了一小段时间,接着变成了 AMC,再后来是 USA(J)MO。这个速度已经横扫了所有这些数学基准。今天,我们邀请到了 Alex Wei、Cheryl Sue 和 Noam Brown,他们是 OpenAI 模型背后的三人组,该模型刚刚在国际数学奥林匹克竞赛中获得了金牌。IMO 金牌是通往超级智能竞赛中最重要的里程碑之一。这项突破之所以特别引人入胜,不仅在于数学能力,更在于底层架构:用于扩展测试时算力和处理难以验证任务的通用技术,其应用范围远超竞赛数学。就在一年前,模型还只能对数学问题进行十分之一分钟的推理,而现在,我们已经有了能够推理和集中注意力长达 100 分钟的系统。对超级智能的希望是,随着我们将推理扩展到数千或数十万小时,我们可以开始解决人类在数学、科学等领域最伟大的未解难题。Alex、Cheryl 和 Noam 加入了我们的《训练数据》节目,讨论他们的方法,并分享这一历史性成果背后的一些幕后趣事和心得。请享受节目。Alex、Cheryl、Noam,非常感谢你们今天加入我们。我们请到了 OpenAI 首个 IMO 金牌背后的团队。祝贺你们所有人。这是一项里程碑式的成就。
The pace of progress is really astonishing. I think you see it so clearly in math. Alex tweeted about this: even a few years ago, these models were struggling with grade school math. I remember even in 2024, GSM8K was used as the standard eval when everyone would release a model, then it was math for a short period, then it became AMC, then USA(J)MO. The pace has just blown through all of these math benchmarks. Today we're joined by Alex Wei, Cheryl Sue, and Noam Brown, the trio behind the OpenAI model that just achieved gold medal performance at the International Math Olympiad. The IMO gold is one of the most important milestones in the race to artificial superintelligence. What makes this breakthrough particularly fascinating isn't just the mathematical chops, but the underlying architecture: general-purpose techniques for scaling test-time compute and handling hard-to-verify tasks that extend far beyond competition math. We've now gone from models that can reason about math for a tenth of a minute just a year ago, to systems that can reason and concentrate on the order of 100 minutes. The hope for superintelligence is that as we scale reasoning to thousands or hundreds of thousands of hours, we can begin to solve humanity's greatest unsolved problems in math, the sciences, and more. Alex, Cheryl, and Noam joined us on Training Data to talk about their approach and share some behind-the-scenes fun and learnings behind this historic result. Enjoy the show. Alex, Cheryl, Noam, thank you so much for joining us today. We have with us the team behind OpenAI's first gold medal at the IMO. Congratulations to you all. It's a momentous achievement.
谢谢。
Thanks.
是的,谢谢。
Yeah, thank you.
我想深入了解一下这背后的起源故事。我知道 IMO 金牌一直是一个难以企及的目标,AI 领域的每个人都在追逐它很久了。我记得 2021 年 Sam 向我们推销时,它就在幻灯片上,我当时想,‘啊,那看起来真的很遥远。’我想了解这个具体努力的更直接的起源故事。你们什么时候开始考虑这个的?它是怎么发生的?
I'd love to get into a little bit of the origin story behind this. I know that the IMO gold has been this elusive thing that everyone in AI has been chasing for a long time. I remember back when Sam pitched us in 2021, it was on the slides and I remember thinking, 'Ah, that seems really far away.' I'd love to understand the more immediate origin story for this specific effort. When did you guys start thinking about this and how did it come about?
是的,我认为这是我们长期以来一直在思考的事情。我记得在 OpenAI 的第一周,Noam 就问我:‘你觉得模型什么时候能拿到 IMO 金牌?’我当时认为 2025 年非常不可能。但这始终萦绕在我们心头。就像你说的,Sam 很多年前也提过。但具体到这个努力,我认为真的只有几个月,从最后的冲刺到为今年的 IMO 做好准备。当然,我们一直在改进我们的强化学习算法。这些想法大概六个月前开始成形,但为今年 IMO 做点什么的最后冲刺只有几个月。
Yeah, I think it's something we've been thinking about for a long time. I remember in my first week at OpenAI, Noam asked me, 'When do you think the model will get IMO gold?' I thought it was really unlikely in 2025. But it's always been on our minds. As you said, Sam many years ago as well. But this specific effort, I think it was really only maybe a couple months since the last sprint to get everything ready for this year's IMO. Of course, we've been working on improving our RL algorithms. The ideas for this started coming together maybe six months ago, but the last push to try to do something for this year's IMO was only a couple months long.
太神奇了。参与的团队有多大?
It's amazing. And how big is the team involved?
我们当然是在 OpenAI 很多人的工作基础上建立的。没有来自推理团队、扩展组织、预训练和强化学习训练人员的帮助,这是不可能的。但就核心团队而言,我认为只有我们三个人。这是一个非常小而精干的努力。
We're definitely building on a lot of folks' work at OpenAI. This is not possible without a lot of help from people working on inference, the scaling org, the people who train the pre-training and the RL training. But in terms of the core team, I would say it's just the three of us. It was a super small, scrappy effort here.
太疯狂了,就你们三个人。
That's crazy, just the three of you.
主要是 Alex。Alex 研究这项技术已经有一段时间了,Cheryl 和我在临近 IMO 时很高兴能帮忙,让它成为现实。
It was mostly Alex. Alex has been working on this technique for a while, and Cheryl and I were happy to help out as we were getting closer to the IMO to make it a reality.
太酷了。这到底是怎么发生的?你是自我指导、自我选择‘我想研究 IMO 金牌,我要带我们实现它’吗?你是如何主动提出要研究这样的项目的?
That's so cool. And how does this even come about? Do you self-direct and self-choose 'I want to work on IMO gold and I'm going to get us there'? How do you even raise your hand to work on something like this?
我认为这只是一个感觉,也许有可能。也许我们努力几个月,就能实现。OpenAI 的一个好处是,研究人员确实有权进行他们认为有影响力的研究。Alex 提出了一个想法,说有一种新技术可能很有帮助。老实说,当时有不少怀疑。有些人支持,但每个人都觉得我们应该给他们自由去探索和追求。然后它开始显示出一些强有力的证据,人们仍然有点怀疑,但更多人开始兴奋起来。最终它变成了更实质性的东西,现在人们显然对此非常兴奋。
I think it was something where it just felt like maybe it's possible. Maybe if we push a bit for a couple months, we can just get there. One of the nice things about OpenAI is that the researchers are really empowered to do the kinds of research they think is impactful. Alex had this pitch that there's this new technique that could help a lot. Honestly, there was a decent amount of skepticism. Some people were supportive, but everybody felt we should give them the freedom to explore this and pursue it. Then it started showing some strong evidence, and people still were a little skeptical, but more people were getting excited. Eventually it turns into something more substantial, and now people are obviously very excited about it.
你能再多说一点关于那些强有力的证据吗?你们看到的早期迹象是什么,让你们真正投入进去?
Can you say a little bit more about the strong evidence? What were some of the early signs that you all were seeing that made you really lean in?
我认为就是在难以验证的任务上取得了进展。以前,我们的很多工作更侧重于可验证的奖励。我们只是看到在这些更难验证的任务上有了更多改进,这让我们很兴奋。
I think it's just progress on hard-to-verify tasks. Previously, a lot of our work was more focused on verifiable rewards. We were just seeing more improvement on these harder-to-verify tasks, which is what made us excited.
也许在这方面,你们是如何验证你们得到的结果是正确的?我看到你们在 GitHub 上发布了证明,但你能再多说一点你们是如何知道你们已经找到了答案的吗?我的理解是,它们的完成方式与人类回答的方式有点不同。
Maybe on that front, how did you even verify that the results you had were right? I saw that you published the proofs on GitHub, but can you just say a little bit more about how you even know that you've discovered the answers? My understanding is that they're done a bit differently from how a human might answer them.
是的,我确实认为模型输出的风格有点糟糕。
Yeah, I do think the style of the model outputs is a little atrocious.
糟糕不是我要用的词。说创意吧,像外星语言。
Atrocious isn't the word I was going to use. Say creative, like an alien language.
是的,这是一个非常小而精干的努力,所以我们没有为了人类可读性而大力优化。但这是我们知道如何做的事情。我们可以用 ChatGPT 那样非常易读的方式做同样的事情;我们在这里也可以做同样的事情。
Yeah, so it was a very small, scrappy effort, so we didn't optimize as hard for human readability. But that's something we know how to do. We can do the same stuff in the same way that ChatGPT is very readable; we can do the same things here.
你们还需要优化人类可读性吗?这重要吗?
Do you even need to optimize for human readability? Like is that even important?
我认为如果你展示给人类看,他们更喜欢可读性。我们实际上讨论过,你知道,我们得到了证明,比如,好吧,因为你其实可以把它们扔给 ChatGPT,让它重写成更易读的形式。证明仍然是正确的,只是更易读一些。然后我们想,哦,我们在网上发布这些时,是发布经过 ChatGPT 处理的更易读版本,还是直接发布原始版本?我们决定,为了完全透明,我们会发布原始版本,让人们自己判断。
I think if you're showing this to humans, they prefer readability. We were actually discussing, you know, we got the proofs like, okay, because you could actually just run them through ChatGPT and ask ChatGPT to rewrite them in a more readable way. And the proofs are still correct, they're just a little more readable. And we were like, oh, should we when we post these online, post the more readable version that's run through ChatGPT or just post the raw version? And we decided, I think for full transparency, we'll post the originals and people will figure it out.
你们 OpenAI 员工里有很多 IMO 奖牌得主和参赛者,对吧?你们业余时间会兼职给模型生成的答案打分吗?
You guys have a bunch of IMO medalists and participants in the staff at OpenAI, right? Do you guys moonlight in your spare time grading the answers that the model produces?
比如在测试期间,是的。我们看了很多样本。但具体给这些打分,我们聘请了外部的前 IMO 奖牌得主。所以每个证明由三位奖牌得主评分,每个都达成了一致共识,确认正确性。
Like during the testing, yeah. We read a lot of samples. But for grading these specifically, we hired external former IMO medalists. So each proof was graded by three medalists and for each one they reached unanimous consensus on the correctness.
我还想说,对我来说,不知道 Cheryl 怎么样,但这些证明超出了我的理解能力。我虽然是数学专业,但从未真正做过竞赛数学,这个模型写的东西已经超出了我的评分能力。
I should also say that for me, I don't know about Cheryl, but for me the proofs are beyond my ability to comprehend. I was a math major and I never really did competition math, and the stuff that this model is writing about is beyond my ability to grade.
是的,我也是。我觉得这更令人惊叹,模型有多聪明。
Yeah, same. I think that's what makes it even more amazing, just how smart the model is.
完全同意。那第六题呢?为什么今年 IMO 的所有模型都没有解出来?你们的模型甚至没有尝试第六题。能多说说这道题的特点吗?传统上第六题总是 IMO 中最难的,对吗?
Totally. What about problem six? How come none of the models at this year's IMO had a solution? And your model didn't even attempt problem six. Can you say more about what makes that problem? And traditionally problem six is always the hardest at the IMO, is that right?
是的,通常是第三题或第六题。
Yeah, I think problem three or problem six usually.
好的。为什么?再多说说第六题有什么不同,以及你们从中学到了什么。我记得你发推说,你的模型知道自己解不出第六题,这给了你希望。也请多谈谈这一点。
Okay. Why? Just say a bit more about what made problem six different and what you learned from it. I think you tweeted that the fact that your model knew that it couldn't solve problem six was one of the things that gave you hope. So just say a bit more about that as well.
对于第六题,它确实是一个非常难的问题。我觉得就算给我几个月时间思考,甚至给我一个大提示关于解题的主要思路,我也不认为自己能解出来。这是一个极其困难的问题,有太多可以尝试的方向,但找到证明的路径非常狭窄。我认为这就是数学的难处。是的。我们投入了大量算力在第六题上。但我觉得很好的是,模型没有试图幻觉或编造一个答案,而是会说没有答案。我的意思是,当你觉得它做了那么多工作却只说没有答案时,确实有点失望,但我认为它能够承认这一点是好的。
For problem six, it's just a really tough problem. I think if you gave me months to think about it, even if you gave me a big hint about the main idea to solve problem six, I don't think I'd be able to get there. It's just a crazy tough problem where there are so many things you can do and there's a very narrow path to finding the proof. I think it's one of those things, math is just hard. Yeah. And we threw a lot of compute at problem six. But I think it was good to see that the model doesn't try to hallucinate or just make up some solution, but instead will say no answer. I mean, it is kind of disappointing when you feel like it's done so much work just to say no answer, but I think it's good that it actually acknowledges that.
是的,这是一种惊人的自我意识,认识到自己的天花板。因为我记得至少几年前,这些模型总是试图提供帮助,编造一个答案,对吧?所以看到这一点,我认为这些模型有了惊人的自我意识。当我们发布推理模型时,我和一些教授、数学家、计算机科学家聊过,我问他们是否觉得这些模型有价值。答案通常是肯定的,但他们抱怨的一件事是,如果问模型一个它不知道答案的问题,它就会输出一个非常有说服力但错误的答案。他们必须非常仔细地检查,才能确定是否完全正确,或者模型是否偷偷翻转了某个不等式之类的。很高兴看到这个模型,如果它不知道,它就会承认自己不知道,至少更频繁地这样做。
Yeah, that's an amazing level of self-awareness of your own kind of ceiling. Because I remember at least a couple years ago with these models, they'd always try to be helpful and make up an answer, right? And so to see this is just an amazing level of self-awareness from these models. When we released the reasoning models, I talked to some professors, mathematicians, computer scientists, and I was asking them if they were finding value in these models. The answer was frequently yes, but the one thing they would complain about is if they ever asked the model a question that it didn't know the answer to, it would just output a very convincing but wrong answer. And they would have to go through it very carefully to figure out if it was exactly correct or if there was some flip of an inequality or something that the model snuck in there. It's nice to see that this model, if it doesn't know, it will just acknowledge that it doesn't know, at least more frequently.
我猜你们内部有没有像赌池之类的东西,赌你们今年能不能拿到 IMO 金牌?内部氛围怎么样?
I guess internally did you guys have like a betting pool or something going on whether you guys were going to win IMO gold this year? And what was the internal vibe?
我觉得我们感觉很有机会,但我们也觉得不是稳赢。肯定有一类问题模型可能比人类更吃力,但另一类问题模型会非常非常强。我认为今年介于两者之间,第六题对当今最先进的模型来说还是遥不可及。而且我认为总的来说,这些困难的组合数学问题,比如第六题,更具挑战性,模型仍然难以应对。
I think we felt like we had a strong shot, but I think we also felt that it wasn't like a lock. There's definitely a distribution of questions where the models would probably struggle more than the humans, but then there's another distribution of questions where the models would be really, really strong. And I think this year was somewhere in the middle, where problem six is just out of reach of state-of-the-art models today. And I think maybe in general, these hard combinatorics problems, which problem six was, are more challenging, and that's still something that the models struggle with.
组合数学有什么特点让它具有挑战性,而相比之下几何你们似乎做得很好?
What is it about combinatorics that makes it challenging versus geometry, for example, which seems like you guys do well at?
我认为对于组合数学,可能是因为它更抽象、更高维。而且通常组合数学问题需要一些信念的飞跃或洞察力的飞跃,模型不太擅长。我认为模型更擅长需要一系列小步骤的问题,比如。
I think for combinatorics, it's probably because it's a little more abstract, a little more high-dimensional. And often times combinatorics problems sort of require leaps of faith or leaps of insight that the models aren't as good at. I think the models are more good at problems that require a bunch of smaller steps, for example.
从你们的角度看,内部氛围是乐观还是不乐观,觉得你们能拿金牌?
What about from your guys' perspective? Was the internal vibe optimistic or not that you all were going to get gold?
我觉得不是特别乐观。我认为他们肯定知道有可能,但我觉得即使一两个月前,感觉还需要改进很多,我想我们确实改进了。我记得比赛前大概两个月,我和 OpenAI 的另一位研究员聊天,我们说,‘好吧,如果要打赌,我是个爱打赌的人。我很乐意赌。’
I feel like it wasn't super optimistic. I think they definitely knew that it could happen, but I think even like a month or two months back, it definitely felt like it'd have to improve quite a bit, which I guess we did. I remember I was talking to another researcher at OpenAI maybe two months before the competition and we were like, 'Okay, if we were to bet, I'm a betting man. I'm happy to bet.'
是的,你是。
Yes, you are.
然后我说,‘你愿意接受什么赔率?’因为我愿意赌。我说,‘我们会拿金牌。’他说,‘真的没机会。’他说他很乐意接受比如二比一的赔率赌模型输。所以概率不到三分之一,但他不想和我们对着赌。他觉得赌团队输会带来不好的氛围,所以没接这个赌。
And I was saying, 'What odds would you take?' Because I was willing to bet. I'm like, 'We were going to get gold here.' And he was like, 'There's really no chance.' And he said that he would gladly take like two to one odds against the model winning. So like less than one-third chance, but he didn't want to bet against us. So he thought it would be bad vibes to bet against the team winning, so he didn't go for the bet.
你从他那赚了点零花钱吗?我真希望我赚了。我是说,你早就知道这一点,因为你们当时——我记得你大概 15 个月前发推说 AIME 上只有 12%的正确率,对吧?所以即使你永远不想跟 Scaling 和 OpenAI 对着干,你们取得的成就斜率也实在太惊人了。进步的速度在数学上体现得特别明显。我记得 Alex 也发过推,说几年前这些模型还在小学算术上挣扎。甚至 2024 年的时候,GSM8K 还是大家发布模型时的标准评估,然后很快变成了 MATH,接着是 AIME,再然后是 USAMO。它就这么一路横扫所有数学基准,真的令人震惊。
Did you make some pocket change to him? I wish I did. I mean you knew that so because I mean you guys were at I think you tweeted 12% on AIME like 15 months ago, right? So even though you never want to bet against scale and OpenAI, it's just an astounding slope of what you all have accomplished here. The pace of progress is really I think you see it so clearly in math. And I think Alex tweeted about this where even a few years ago these models were struggling with grade school math. And I remember even in 2024 that GSM8K was used as the standard eval when everybody would release a model, and then it was MATH for a short period of time, and then it became AIME, and then it became USAMO. The pace that it's just gone blown through all of these math benchmarks is really astonishing.
嘿,我记得两年前还在 GSM8K 上训练模型呢。
Hey, I remember training a model on GSM8K two years ago.
是啊,那些日子已经过去了,对吧?评估都饱和了。接下来呢?你觉得明年这个时候,我们能解决千禧年大奖难题吗?
Yeah, we're past those days, huh? Saturated the evals. What's next? Do you think at this point next year, do you think we'll be solving Millennium prizes?
我觉得那些还很遥远。一方面,想想从 GSM8K 以来数学进步有多大——两年前它还是大家努力攻克的标准。那真是惊人的进步。但另一方面,也要想想人类需要花多少时间。GSM8K 是小学算术,数学好的人几秒钟就能解一道。现在我们从几秒钟进步到优秀学生平均每题一个半小时。IMO 是三个题四个半小时。而研究数学,同样的优秀学生成长为研究员,可能需要 1500 小时。所以思考时间差了上千倍。千禧年大奖难题则耗费了整个领域、人们一生的思考,而且大多数还没什么进展。所以一方面,我们取得了这么多进步,非常令人兴奋;另一方面,看到从 1.5 小时到数万、数十万小时的人类思考还有多远,也让人感到谦卑。
I think those are still very far away. On one hand, you think about how much math progress has been made since GSM8K, which was just two years ago the standard people were trying to push on. That's an astounding level of progress. But also, you think about how much time it takes for people. GSM8K problems are grade school math; someone good at math takes a couple seconds. Now we've gone from a couple seconds to something that takes brilliant students an hour and a half per problem on average. The IMO is three problems in four and a half hours. Then research math is going to take those same brilliant students, now grown up as researchers, maybe 1500 hours. So there's a thousand times more thinking time. And the Millennium Prize problems have taken entire fields, people's lifetimes of thinking, and we still don't have much progress on most of them. So on one hand, it's super exciting that we've made so much progress. On the other hand, it's humbling to see how much further progress has to go from an hour and a half to tens of thousands, hundreds of thousands of hours of human thinking.
完全同意。Noam,我觉得你在这方面预见未来功不可没。我记得你还没加入 OpenAI 之前就来拜访过我们,谈到游戏中的结果,以及如果让模型思考几个小时、几十个小时会发生什么。你确实看到了未来。
Totally. Noam, I think you deserve a lot of credit for seeing the future on this. I remember you visited us before you even joined OpenAI, talking about the results from gameplay and what happens if you let a model think for hours and tens of hours. Credit to you, you've really seen the future on this.
谢谢。是啊,看到它真的实现很令人兴奋。当你把算力时间、推理时间从 0.1 分钟量级扩展到 100 分钟量级时,会出现哪些难题?从宏观层面讲,因为不是所有人——我们大多数听众不是 AI 研究员——但要让模型保持在正轨上,难点是什么?一个明显的挑战是:如果你让模型思考 1500 小时,那么为了评估它,你也得让它思考 1500 小时。所以最终,模型评估会成为进展的一个重大瓶颈。我们还没到那一步。如果让模型思考一个半小时,没什么大不了的,我们可以跑那些测试。但要跑一个模型思考一个月的测试,就得花一个月才能完成。所以如果你要等那样的结果,进展就只能那么快。
Thank you. Yeah, it's exciting to see it actually happen. What are the hard things that happen as you scale compute time, inference time, from the order of 0.1 minutes to the order of 100 minutes? At a high level, because not everyone, most of our listeners are not AI researchers, but what are the hard things that happen to keep the model on the rails so to speak? One thing we can point to is clearly a challenge: if you have the model thinking for like 1500 hours, then in order to eval, you have to have it think for 1500 hours. So eventually the evaluation of the models becomes a significant speed bump on progress. We're not really at that point yet. If we have the model think for an hour and a half, it's no big deal. We can run those tests. But to run a test where the model is thinking for a month, it takes a month to finish that test. So progress can only advance so fast if you want to wait for those kinds of results.
我觉得你们俩都在多智能体团队。帮我理解一下多智能体系统在这方面扮演什么角色。
I think both of you are on the multi-agent team. Help me understand what role multi-agent systems play in this.
是的,所以除了让模型长时间思考并在难以验证的任务上取得进展外,这还涉及扩展并行算力。所以其中有多智能体的成分。我们可能无法深入讨论具体技术细节。但这确实是我们为 IMO 扩展测试时算力的一种方式。顺便说一句,关于多智能体扩展并行算力,我们做的方式是真正优先考虑技术的通用性。比如,我研究过扑克 AI。Alex 和我都研究过外交 AI。Alex 是 Cicero 团队的成员。那些项目我真的很自豪,但也是我们花了多年时间才取得成果的项目。随着 AI 进步速度如此之快,感觉花时间开发一个只能做那一件事的定制系统并不是最好的时间利用。所以我们所有人都真正优先考虑了通用技术。我们用于扩展思考时间、处理难以验证的任务以及并行算力的技术,都是通用技术,我们计划或已经用于其他系统。
Yeah, so in addition to having the model think for a very long time and make a lot of progress on hard-to-verify tasks, this also involved scaling up parallel compute. So there's a multi-agent component to that. We're probably not going to be able to go into too much detail about the exact techniques. But that was certainly one way we were able to scale up test-time compute for the IMO. By the way, one thing I'll add for the multi-agent scaling parallel compute thing is that the way we did it, we really tried to prioritize generality in our techniques. For example, I worked on AI for poker. Alex and I actually both worked on AI for diplomacy. Alex was on the team that worked on Cicero. Those were projects I'm really proud of, but they were also projects we spent years working on to achieve that result. With the pace of AI progress being so fast, it felt like that wasn't the best use of time to develop a very bespoke system that could only do that one task. So we all really prioritized general-purpose techniques in all this. The techniques we used for everything—scaling up thinking time, working on hard-to-verify tasks, and parallel compute—are all general-purpose techniques that we're either planning or have used for other systems as well.
这就是你们选择不用 Lean 的原因吗?我的理解是,今年官方 IMO AI 赛道是 Lean 版本。所以你们才不选 Lean 吗?
And is that the reason you all chose not to do this in Lean? My understanding is the official IMO AI track was a Lean interpretation this year. Is that why you guys chose not to go with Lean?
是的,没错。我是说,Lean 作为工具当然有很多价值。比如数学家觉得它有用。但我们的优先事项是通用推理能力,而 Lean 有其局限性。所以这就是为什么我们优先选择自然语言。
Yeah, that's right. I mean, there is certainly a lot of value in Lean as a tool. Mathematicians find it useful, for example. But the priority for us is really general-purpose reasoning capabilities, and Lean has its limitations. So that's why we wanted to prioritize natural language.
我外行的理解是,Lean 是一个形式化验证工具。你们的结果是否基本上说明,带有规模的非形式化验证可以达到甚至超越形式化验证的水平?这是正确的结论吗?
My layman's understanding is Lean is a formal verification tool. Does your result here basically say that informal verification with scale can perform at the same level or even surpass formal verification? Is that the right takeaway?
我不认为那是正确的结论。我不知道。
I would not say that's the right takeaway. I don't know.
Alex,你有什么想法?
Alex, you have thoughts?
我觉得这就像是两个正交的组成部分。我认为非形式化数学是一个有趣的问题,因为它代表了在难以验证的任务上扩展测试时算力的一个困难核心。这反映了我们从通用角度感兴趣的一个非常广泛的任务集合中的困难。我认为 Lean 更狭窄一些,因为世界上很多事情可以用非形式化推理来处理,而不是形式化。
I say that these are just like two orthogonal components here. I think we found the informal math an interesting problem because it represents a kernel of difficulty around scaling up test-time compute for hard-to-verify tasks. That represented difficulties from a very broad set of tasks that we were interested in from a general-purpose standpoint. I think Lean is a little more narrow, where a lot more of the world can be approached with informal reasoning than is formalizable.
我不认为狭义 AI 有什么问题。狭义 AI 可以非常有效,在某些领域显然远超通用 AI。我认为正确的思考方式是,就像人类数学家从 Lean 中发现很多价值一样。通用 AI 可以与专注于形式化数学的更狭窄系统兼容,并且这种组合可以因此变得更好。
I don't think there's anything wrong with narrow AI. Narrow AI can be very effective and obviously far surpass general-purpose AI in certain domains. I think the right way to think about it is in the same way that human mathematicians find a lot of value in Lean. General AI can be compatible with a more narrow system focused on formal mathematics, and the combination can be better because of it.
我在推特上看到 OpenAI 的几个人,我想你们也提到过,这个系统是用与 OpenAI 最近许多发布非常相似的方法和基础设施构建的。上周我们邀请了聊天智能体发布的 Issa 来播客。你能多说一点关于这个相似的基础和方法是什么吗?
I saw on Twitter from multiple folks at OpenAI, and I think you guys have mentioned this as well, that this system was built with a very similar approach and infrastructure to many of the recent launches from OpenAI. We had Issa from the chat agent launch on the podcast last week. Can you say a little more about what the similar foundation and approach is?
从基础设施的角度来看,我们都使用相同的基础设施。但就这个问题的核心而言,就像 Alex 说的,这里没有什么专门针对 IMO 的东西。真正的希望是,我们可以使用 Alex 研究的那些用于非可验证任务和扩展测试时算力的技术,并将其应用于推理的其他领域或模型能力的其他方面,从而构建更强的模型,不断改进智能体、ChatGPT 以及其他一切。
Infrastructure-wise, we all kind of just use the same infrastructure. But as far as the core of this question, like Alex said, there's nothing very bespoke to IMO here. The hope is really that we can use the techniques that Alex worked on for non-verifiable tasks and scaling up test-time compute, and be able to apply this to other areas of reasoning or other areas of model capabilities in general, and just build stronger models, keep improving agent, keep improving ChatGPT and everything else.
告诉我 IMO 当天的实际体验。感觉怎么样?
Tell me about the actual experience of IMO day. What was it like?
我们当时在等题目出来,因为一旦参赛者完成考试,题目就会公布。所以我们大概在深夜,可能是凌晨 1 点左右把题目输入模型。老实说,我去睡觉了,因为凌晨 1 点,我不会熬夜四个半小时。我早上起来再看。但我想这两位实际上熬夜了,看着模型实时运行。
We were waiting for the problems to come through because once the participants finish the exam, they get posted. So we plugged the problems into our model around pretty late at night, maybe like 1 a.m. or something. Honestly, I went to sleep because it's 1 a.m., I'm not going to stay up for four and a half hours. I'll just wake up in the morning and see. But I think these two actually stayed up and got to watch the model and see it come in in real time.
是的,非常有趣。
Yeah, it was a lot of fun.
有没有人想打电话给我说,‘醒醒,醒醒。我们成功了。’
Did anybody want to call me like, 'Wake up, wake up. We got this.'
有几次 Alex 太累了,决定小睡一会儿,但我们告诉他,‘好吧,确保你的手机静音,这样如果我们需要叫醒你,我们可以打电话。’有一次我们确实不得不打电话给他,但我觉得他没醒。
There were a couple moments where Alex was so exhausted that he decided to take a nap, but we told him, 'Okay, just make sure your phone is on silence so that if we need to wake you up, we can call you.' At one point we did actually have to call him, but I don't think he woke up.
太棒了。那一定非常激动人心,特别是当结果出来的时候……所以你凌晨 1 点开始,那么大概早上 9 点就知道结果了?
That's awesome. It must have been such a thrill and such a high, especially for that to come through at like... So you started at 1 a.m., so you must have known at like 9 a.m. then?
哦,是四个半小时。
Oh, it's four and a half hours.
第一个花了四个半小时?
Four and a half hours for the first?
是的。我不知道。我的意思是,我们能看到题目进来。所以我只是确保系统保持稳定,然后 Alex 在那里阅读,看看模型做得怎么样。
Yeah. I don't know. I mean, we can kind of see the problems come in. So I'd just be making sure the systems are staying stable, and then Alex is over there reading and seeing whether or not how the model's doing.
所以你是在做实时的人工验证,看它是否真的……我自然对结果非常焦虑,所以我只是看着模型取得的进展。你可以观察到一些东西。然后我也手动检查了,因为我们本来要发给评分员,但我实在太好奇了,所以也手动检查了。
So you were doing the live human proof checking to see if it was actually... I was naturally very anxious about the results, so I was just looking at the partial progress the model was making. You can sort of observe that. And then I also hand-checked things because we were going to send these out to the graders, but I was just also hand-checking them because I was so curious.
好吧,下次叫我。我想去那里凑热闹。我会去睡觉。听起来太棒了。
Okay, well call me next time. I want to come hang out there for that. I'll go to sleep. That sounds awesome.
这些模型的一个很酷的地方是,我看不懂证明。但当你看到模型在思考时,它会在这个过程中用自然语言表达它的不确定性或信心。它会说一些暗示其信心的词语。如果它非常有信心解决了问题,它会说很多‘好’。如果不确定,它会加入很多问号。所以很酷的是,我可以大致跟上,看看模型对自己的进展感觉如何,尽管我无法判断它是否正确。
One of the cool things about these models is that I can't understand the proofs. But when you see the model thinking about it, it will express its uncertainty or its confidence in natural language throughout the process. It will just kind of say words that hint at its confidence. If it's really confident that it figured it out, it'll say 'good' a lot. And if it's unsure, it'll throw in a lot of question marks. So it's cool that I can kind of follow along and see how the model is feeling about its progress, even though I can't really tell if it's got it correct or not.
是的,你会得到可怕的‘看起来很难’。
Yeah, you get the dreaded 'seems hard'.
你在第六题上得到了那个。
You got that on problem six.
经常得到那个。‘没有进展,很难。’
Got that a lot. 'No progress hard.'
‘看起来很难。继续。太糟糕了。’
'Seems hard. Keep going. Too bad.'
太棒了。展望未来,你们已经在竞赛数学中取得了巅峰成果。我想你们明年可以参加普特南竞赛,但基本上已经登顶了,对吧?那么下一步是什么?
Wonderful. Looking ahead, you've gotten the pinnacle results in competition math. I guess you can go do Putnam next year, but you're basically at the top, right? So what's next?
实际上对于普特南竞赛,由于每道题的时间比 IMO 少,而且更注重知识储备,我们在评估中发现模型非常擅长普特南题目,比 IMO 题目还要好。所以我认为这里的边界不再是这些时间非常有限的竞赛题,而是那些需要更长时间和更深入思考才能解决的问题。
Actually for Putnam, the problems, since the exam is less time per problem than the IMO and it's a little more knowledge-heavy, we actually found in our eval that the model was really, really good at Putnam problems, better than it was at IMO problems. So I think the frontiers here are really not about these very time-boxed competition problems anymore, but about problems that really take longer periods of time and more deep thinking to solve.
这真的很酷。好吧,所以你们现在要开始证明新的定理了。但我认为,在这些时间有限的竞赛题和真正的研究突破之间,存在一个非常令人生畏的差距,后者需要一年的工作量,大约 1500 小时而不是 1.5 小时。
It's really cool. Okay, so you're going to start proving novel theorems now. Again, I think there's this very intimidating gap between these very time-boxed competition problems and a real research breakthrough, which takes a year's worth of work, like on the order of 1500 hours instead of 1.5.
是的,完全同意。
Yeah, totally.
我想相关的是,我昨晚听了 Demis 的播客,他提到最难的事情实际上是提出有趣的问题来解决。我很好奇你们是否同意这一点。
I guess relatedly, I was listening to the Demis podcast last night and he mentions that the hardest thing is actually coming up with the interesting problems to solve. I'm curious if you all agree with that.
我认为这有一定道理。这些模型现在确实很擅长解决这些问题。但提出这些问题仍然是一个挑战,不过我想指出的是,我们正在见证令人难以置信的进步速度。总会有下一个难关。最初大语言模型出现时,问题是:如何让它们推理?然后我们让它们学会了推理,但接下来是如何让它们在难以验证的任务上推理?现在它们已经能在难以验证的任务上推理了。我认为下一个难关将是:如何让它们提出新颖的问题?即使是出一道国际数学奥林匹克竞赛题也是一个挑战,需要很多数学家和大量工作。但我不认为有任何根本性的障碍阻止我们实现这一目标。
I think there's some truth to that. These models are really good now at solving these problems. Coming up with them is still a challenge, but I think it's also worth noting the incredible pace of progress. There's always a next hurdle. Originally when LLMs came out, it was like, how do we get them to reason? Then we got them to reason, but then how do we get them to reason on hard-to-verify tasks? And now they can reason on hard-to-verify tasks. I think the next hurdle is going to be: how do we get them to come up with novel questions? Even creating an IMO question is a challenge, and it takes a lot of mathematicians and work. But I don't see any fundamental barriers that block us from getting there.
说得好。你们在数学上的成果能完全泛化吗?你们会变得更擅长科学推理、通用推理吗?擅长竞赛数学是否意味着在其他方面也很出色?
I love that. Do your results in math fully generalize? You're just going to be better at scientific reasoning, general reasoning? Does being great at competition math make you great at everything else?
我认为我们的方法并不是要擅长竞赛数学。我们专注于开发通用技术来改进强化学习。我们非常兴奋能在数学以外的其他领域改进模型,并希望让模型在日常使用中更有用。这是一个相当新的成果,甚至让 OpenAI 内部的人都感到惊讶。下一步是将这一成果更广泛地整合到我们的模型中,全面提升推理能力。但这需要一些时间来完成这个过程并部署到全世界。所以我认为它会到来,只是需要更多时间。
I think how we approached this was not like we should be great at competition math. We were focused on developing general-purpose techniques to make our reinforcement learning better. We are very excited to improve our models in other domains beyond math, and hopefully make models more useful for everyday usage. This is a pretty late-breaking result. It was a surprise even to people internally at OpenAI. The next step is to incorporate this more broadly into our models and improve reasoning capabilities across the board. But it's going to take some time to go through that process and deploy it to the world. So I think it's going to come, but it'll take a little more time.
这些模型做国际数学奥林匹克竞赛题更难,还是物理奥林匹克竞赛题更难?
Is it harder for these models to do the IMO or the Physics Olympiad?
我认为肯定是物理奥林匹克竞赛更难,因为它有实验部分。
I think definitely the Physics Olympiad, because it has an experimental section.
哦。或者你必须……我们得先解决机器人问题。我之前不知道,我以为只是在纸上答题。
Oh. Or you have to... we need to solve robotics first. I didn't realize that. I thought it was just done on a piece of paper.
是的。所以我认为模型可能在纸面部分表现不错,但还需要一段时间才能做实验,而不是用世界模型。
Yeah. So I think the model will probably be good at the paper part, but it will be a bit of time before it can do the experiments, not with a world model.
好的。你们会发布这个模型让客户试用吗?Ruof 的儿子是个数学奥林匹克选手,他说‘我想用这个数学奥林匹克模型’。人们能用到它吗?
Okay cool. Are you going to release this model for customers to play with? Ruof's son is a math olympiad kid and he's like 'I want access to the math olympiad model.' Will people be able to play with this?
我们希望让数学家能够使用它。我们仍在研究具体的实现细节。但我认为我们开发出这个非常擅长数学的系统真的很酷,而且我们想看看数学家能用它做什么,这很合理。实际上我已经和一位斯坦福数学教授通过邮件联系了。大约一年前,在我们宣布 o1 之前,他给我发邮件说:‘嘿,你想合作解决一些困难的数学问题吗?’我基本上告诉他,我认为我们只需要推进通用推理能力,最终它们就能帮助你解决困难的数学问题。我认为那实际上是最有希望的途径。他有点怀疑,但每次我们发布推理模型,他都会发邮件跟进,问‘现在它能解决这个问题了吗?’我一直在测试,虽然我不知道输出是什么,但我把结果发回给他,他说‘嗯,那是错的。’这次他又发邮件跟进,问同样的问题:‘现在它能解决了吗?’它仍然解决不了,但至少这次它意识到自己解决不了。所以我认为这是一大步。我们很好奇是否还有很多其他问题,数学家们想用这个模型来挑战,看看它能否应对。
We want to make this accessible to mathematicians to use. We're still trying to figure out the exact details of how we make that happen. But I think it's really cool that we've developed this system that is incredibly good at math, and it makes sense that we want to see what mathematicians can do with it. I've actually already been emailing with a Stanford mathematics professor. He emailed me about a year ago before we announced o1, and he was like, 'Hey, do you want to do a collaboration on solving hard math problems?' Basically what I told him is that I think we just need to advance general reasoning capabilities, and eventually they're going to be able to help you with your hard math problems. I think that's actually the most promising route. He was a little skeptical, but every reasoning model release, he's emailed me with a follow-up, like 'Can it solve this problem now?' I've been plugging them in, and I don't know what the output is, but I email it back to him, and he says, 'Yeah, that's wrong.' He emailed me a follow-up this time with the same problem, asking 'Can it solve it now?' It still can't solve it, but at least this time it recognizes that it can't solve it. So I think that's a big step. We're curious to see if there are a lot of other problems out there that mathematicians want to challenge this model with and see if it can take them on.
太棒了。祝贺你们所有人。我认为这是一个重大的成果,整个领域已经等待了很长时间,而且它是由一个三人团队在两个月内完成的,这非常了不起。祝贺你们,感谢你们参加《训练数据》节目。
Amazing. Congratulations to you all. I think this is a momentous result that the entire field has been waiting for for a very long time, and the fact that it was accomplished by a team of three people in a span of two months is extraordinary. Congratulations and thanks for joining us on Training Data.
谢谢。
Thank you.
感谢邀请我们。
Thanks for having us.