AI in Mathematics: From Ineffective Grad Student to Gold Medal Performance
打开互动全文版(中英对照 + 朗读 + 问答)→Terry 和 Mark 讨论了过去一年 AI 工具的演变,从低效研究生到在数学竞赛中取得金牌表现,以及这如何改变数学研究。
Terry and Mark discuss how AI tools have evolved in the past year, moving from being an ineffective grad student to achieving gold medal performance in math competitions, and how this is changing mathematical research.
非常感谢。在开始之前,非常感谢研究所今天接待我们。场地很漂亮。也感谢各位的到来。我知道你们不是来听我说话的,所以我不会讲太久。真的非常感谢你们两位的到来。很少能同时请到两位如此杰出的头脑。我们非常感激你们为此付出的时间。这是我们第三次了。而且开始形成一个小模式。也许这是开始对话的好地方。你们大约一年前进行过一次对话。当时,Terry,你对 GPT 在数学方面的评估是它像一个非常低效的研究生,这个评价我一直记得,因为我自己作为人类也收到过这样的反馈。所以这是一个清晰的基准。不如我们先从你认为事情从那以后发生了哪些变化开始,然后 Mark 你再谈谈你的看法。
Well, thank you very much. Before we get going, a massive thank you to the Institute for hosting us today. Beautiful space. And also to all of you for turning up. I know you're not here to hear me talk, so I won't talk for too much longer. Really just to say also a massive thank you to you both for coming. It's rare that you get two such great minds in the same place. So we really appreciate the time going into this. It's our third time. And a little pattern is starting to build up. Maybe that's a good place to start the conversation. You guys had a conversation almost a year ago to the day. And at the time, Terry, I think your prognosis for where GPT was for mathematics was something like a very ineffective grad student, which remained with me because I'd heard that feedback myself as a human being. So it was a clear benchmark. Why don't we start with how you think things have changed since then, and then Mark give your side of the story.
好的,是的。过去一年发生了很多事情。不仅在 AI 领域,而且这些工具确实变得更加强大了。我认为现在有一些能力基本上已经常态化了,我们一直在使用它们。深度研究工具、文献搜索变得非常非常好。它已经超越了传统的搜索。代码生成当然是一件大事,但作为一名纯数学家,我不是代码的重度用户,但它改变了我处理数学问题的方式。我会绘制一些图表。如果我认为某个猜想是真的,我会让 AI 尝试证明或反驳它。如果有一个引理我认为我知道如何证明,但就是懒得做纸笔计算,我会外包出去。我还没有发现它在最深层次上有用,当我试图用纸笔或与同事一起解决问题时。我还不能以我需要的对话水平与它互动。但也许将来可以。我认为在社会层面上,整个数学界开始意识到这些工具会一直存在,我们必须开始调整我们做研究的方式。某些非常繁琐的事情,也许我们会强迫研究生去做,现在可以交给 AI,这开辟了许多新的数学研究方式,特别是大规模的研究项目,这是我们以前做梦也想不到的。虽然我认为我们可以用 AI 来辅助当前的工作流程,但这样做还有点别扭,但我认为更大的潜力在于创建针对 AI 优化的新工作流程。就像我们发明汽车时,我们开始改变城市建造的方式。当然,你可以说也许并非所有改变都是好的,但我们正处于一个中间状态,我们的道路仍然是为人和马建造的,而现在我们有了汽车。
Okay, yeah. So a lot has happened in the last year. Not just in AI, but yeah. These tools have definitely become a lot more powerful. I think there are now capabilities that are basically normalized and we just use them all the time. Deep research tools, literature search has become really, really good. It has surpassed traditional searches. Code generation of course is a big thing, but as a pure mathematician, I'm not as heavy a user of code, but it has changed the way I approach a math problem. I will plot something. If there's an inquiry I think is true, I will ask an AI to try to prove or disprove it. If there's a lemma that I think I know how to prove, but I just can't be bothered doing the pen and paper calculations, I will outsource it. I've not yet found it to be useful at the deepest level, when I'm trying to solve a problem with pen and paper or with a colleague. I can't interact with it on a conversational level quite at the level I need yet. But maybe in the future. I think also socially, the mathematician community as a whole is beginning to understand that these tools are here to stay and we have to start adapting how we do our research. Certain things that were very tedious, maybe we would force our graduate students to do, we can offload to AI, and this opens up lots of new ways to do mathematics, lots of research projects especially at scale that we could not dream of doing. While I think we can use AI to assist our current workflows, it's a little bit awkward still to do that, but I think much more miles are in creating new workflows which are optimized for AI. It's like when we invented the automobile, we started changing the way we built cities. Of course you could say that maybe not all the changes were good, but we're sort of in this intermediate state where our roads are still built for people and horses and we now have automobiles.
那么,是否可以公平地说,我们已经到了偶尔有帮助的协作阶段,但也许比所有更大的开放空间更有趣的是,随着这些工具的出现,你如何改变你做数学的方式。Mark,你所看到的和你正在构建的是这样吗?
So would it be fair to say that we've got to the point where occasionally helpful collaborating, but maybe more interesting than all the bigger open spaces is how you change the way you do maths with these tools coming. Mark, would that be true to what you're seeing and what you're building for?
是的,老实说,我不怪 Terry 一年前说它是个低效的研究生。我认为那基本上是我们当时的状态。我真的认为 AI 进步的背景是爬山,我们内部称之为“米表”,即模型自主工作时间越来越长。我认为去年我们处于分钟级别。你也看到了,对吧?模型会幻觉,当你给它大量工作时它会崩溃。但我确实认为过去一年对我们许多人来说是一个过渡,我们看到错误减少了,因此你可以信任模型做更长时间的工作。这真的让我们摆脱了过去可能需要的大量脚手架,真正开始攻击更大的问题,并与模型真正协调。是的,我只想到一年前我们大致在国际数学奥林匹克竞赛中获得铜牌。我认为今年夏天,在所有高中数学和编程竞赛中,我们达到了金牌水平,我认为我们基本上已经用完了这些人类编写的基准,这就是为什么你看到人们进化到做数学研究的领域。从根本上说,这一直是目标。我们在 OpenAI 并不以解决 IMO 问题之类的事情为荣。真正的雄心是推动科学的前沿,最终任务视野已经赶上,我们实际上能够去做这项工作。再说一次,这还没有实现。我认为结构中的趋势很强,但我确实认为今天很多人发现了它的实用性。我想也许先谈谈第一个证明,以及当我们进入更前沿的数学时的过渡。
Yeah, honestly I don't blame Terry for saying it's an ineffective grad student a year ago. I think that's largely the state that we were in back then. And I really do think of the backdrop of AI progress as hill climbing this what we call meter plot internally of the models doing autonomous work for longer and longer periods of time. And I think last year we were in the category of minutes. And you saw that, right? It would just the model would hallucinate, it would kind of fall over when you gave it significant chunks of work. But I do think the last year has been a transition for a lot of us in that we've seen the mistakes go down and therefore you can trust the model to do longer periods of work in general. And that's really kind of allowed us to do away with a lot of the scaffolding that we might have needed to use before and really start to attack bigger problems and truly orchestrate with the model. And yeah, I just think of a year ago we were in the world where we were kind of roughly achieving a bronze medal at the IMO. I think this summer across all kind of high school mathematics and programming competitions we are achieving gold medal performance and I think we've just kind of run out of these human written benchmarks and that's why you do see people evolving to this sphere of doing mathematical research. And fundamentally that's always been the goal. We don't find any pride at OpenAI just kind of solving IMO problems or anything like that. The real ambition is to push the frontier of science and finally the task horizon has caught up to a point where we are actually able to go do that work. And again it's not there yet. I think the trend in the structure is strong but yeah, I do think it's true that a lot of people are finding utility in it today. I'd like to come to maybe first proof and that transition as we go into more frontier mathematics.
但也许停留在当前的能力上,我认为 IMO 问题通常被视为检验模型水平的试金石。这也许是一个有代表性的集合,因为其中一些问题可能不像其他部分那么复杂,我们这样设计更好。你可能会说,模型的成功在于快速解决许多较容易的问题,而不是必然向橡子级别的问题迈进。在你看来,Terry,这是对它们当前状态的公平描述吗?
But maybe to stay with the capabilities right now I think often the IMO problems are seen as a way of getting a litmus test for where the models are. And that maybe is a representative set in that some of those problems are maybe not as complex as other parts of those problems and we're better designed that way. And that you might say that the success of the models has been in doing lots of the easier ones quickly versus necessarily moving towards the kind of acorn level problems. Is that a fair depiction in your mind, Terry, where they are today?
是的,所以我一直积极参与追踪 Erdos 问题的进展。我的意思是,是的,基本上很大程度上你会怎么说?这些问题难度范围很广。有些我们非常想解决,而且已经研究了数十年。你知道,我有一些论文在这些问题上取得了微小进展。而今天,AI 并没有真正帮助那些我们已经投入大量关注的问题。但有一个非常长的尾部问题。
Yeah, so I've been heavily involved in tracking the progress on the Erdos problems in particular. So, I mean, yeah, and it is basically largely what would you say? These problems range widely in difficulty. There are some that we desperately want to have solved and they've been worked out for decades. You know, I have papers making tiny progress on some of these problems. And today AI has not really helped with the ones that we've already poured a lot of attention to. But there was this very long tail of problems.
厄多斯提出了上千个问题。它们不全是好问题,但他明白重要的是激发兴趣,而且他隐约知道那些重要的问题会自行发展。但有一长尾的小问题几乎没有任何后续文献。这正是 AI 工具取得惊人进展的地方。大概有二三十个这样的问题在极少人类监督下被解决了。我们还能验证它们,通常借助其他 AI 工具进行验证。我们摸索出了一套工作流程,避免被 AI 垃圾或错误答案淹没。所以这是一种我们以前没有的新能力,因为我们现在可以攻克注意力瓶颈问题。我认为这表明我们需要开始为 AI 工具和公众创建越来越广泛的挑战集。在同一时期,许多最早的问题也被业余数学家解决了,有时用 AI,有时不用。让 AI 成功的相同工作流程也让业余数学家成功。所以我预见一种文化转变:我们不再只专注于少数极难的问题,不分享我们关心的其他问题列表,而是开始发布我们想要答案的问题。也许 AI 能解决其中 10%,一个高中生能解决另外 5%。我们可以用更社区驱动的方式做数学。我认为这就是最早的问题所预示的。
It was Erdős who posed a thousand problems. They weren't all winners, but he understood that the important thing was to stimulate interest, and he kind of knew that the problems that were going to be important would take on a life of their own. But there was this long tail of small problems where there's almost no follow-up literature. That's where AI tools have made a lot of spectacular progress. Maybe 20 or 30 of these problems have been solved with fairly minimal human supervision. And we were able to verify them, often with some other AI tools for verification. We've worked out a workflow for doing this without being overwhelmed by AI slop or incorrect solutions. So it's a new capability we hadn't had before, because we can now attack attention bottleneck problems. I think this suggests we need to start creating more and more broad challenge sets for AI tools and the general public. In the same period, many of these earliest problems were also solved by amateur mathematicians, sometimes with AI tools, sometimes without. The same workflows that enable AI to be successful also enable amateur mathematicians to be successful. So I foresee a change in our culture where instead of only working on a small number of really hard problems and not sharing a longer list of other things we care about, we'd all start releasing problems we want answers to. Maybe AI can solve 10% of them, and a high school student can solve another 5%. We can get a much more community-driven way of doing mathematics. I think this is what the earliest problems are an early harbinger of.
将这一点放在其他科学领域背景下看很有趣,至少在我的生物学领域,任何一篇论文的合作人数都随时间呈指数增长。科学更像是一项团队运动。数学和某种程度上理论物理是例外。当你思考这一点时,这总是关于我们能将模型做得多聪明、它们能回答多难的问题,还是也关于如何赋能人类协作解决这些问题?
It's interesting to contextualize that in other domains of science, at least in biology, the number of people collaborating on any given paper has risen exponentially over time. Science is much more of a team sport. Maths and to some degree theoretical physics are outliers. When you think about this, is it always just a question of how smart we can make the models and the ever more difficult questions they can answer, or is it also about how to empower humans to work collaboratively on these problems?
现在我们确实看到与社区的高度互动。这是推动所有科学领域进步的必要部分。凯文在这里负责我们的 OpenAI for Science 项目,其中一部分就是这些实验,比如第一个证明或厄多斯问题,确实需要与社区互动,找出哪些问题值得解决。我们在物理学中也这样做过,请来专家物理学家制定一个适合 AI 的重要问题计划。这有助于我们塑造 AI 并发现缺陷,从而加以改进。我们希望建立一个平台,让世界各地的科学家能加速自身发展。我们想赋能那个社区数学家。今天我们看到了这样的人,20、21 岁的孩子用模型解决一些厄多斯问题。可能不是复杂的重大飞跃,但他们能做很多自主工作。当你问这个问题时,我想到:我知道你以前组织过很多大型数学社区倡议。你认为 AI 如何改变那个世界?它是否以重要方式进入那个世界?
Right now we really see heavy engagement with the community. That is a necessary part of driving progress in all scientific fields. Kevin here runs our OpenAI for Science program, and part of it is that these experiments like the first proof or the Erdős problems really involve engaging with the community to figure out what problems are important to tackle. We've done this in physics as well, bringing in expert physicists to lay out a program of important things amenable to AI. That helps us shape the AI and find deficiencies, so we can shore things up. We hope to build a platform where scientists around the world can accelerate themselves. We want to empower that community mathematician. We see people like that today, 20- or 21-year-old kids using the models to solve some of these Erdős problems. It may not be sophisticated significant leaps, but they can do a lot of self-directed work. I had this thought when you asked the question: I know you've organized a lot of big community initiatives in math before. How do you think AI is changing that world? Does it enter that world in a significant way?
我认为实际上它们结合得很好。AI 最终将实现的是分工,这是自工业革命以来每个行业都通过分工提高效率的方式,除了数学。传统上,做数学涉及多个任务:问题生成、策略生成、在生成的策略中选择策略、执行策略、验证和沟通结果。我们训练数学家在每个任务上都做得不错。我们专攻领域,但必须了解问题从何而来、什么是好问题、什么是好策略。我们必须有技术技能、写作和解释能力。有些数学家在某些方面更擅长,所以我们从合作中受益。但我们不能像科学那样真正专业化,科学可以有技术人员和项目经理。现在有了 AI 和其他现代协作工具以及完整的验证,可以运行数学项目,让个体参与者只专攻一个领域。也许你的合作中有空白——没人知道如何做技术部分——但 AI 可以填补一些空白。你仍然需要人类,因为 AI 性能非常参差不齐。一些输入现在可以自动化,但如果自动化太多,比如你自动化了策略生成但没有自动化验证,你会得到数百个无法处理的策略。但如果验证也跟上,突然你就有了一种极其有效的做数学的新方式。
I think it combines very well, actually. What AI will enable is finally a way to use division of labor, which is something that every industry since the industrial revolution has managed to become more efficient through, except mathematics. Traditionally, doing mathematics involves several tasks: problem generation, strategy generation, strategy selection among all the strategies you generated, execution of the strategy, verification, and communication of results. We've trained our mathematicians to be somewhat good at each of these tasks. We specialize in fields, but we have to have some idea of where problems come from, what are good problems, what are good strategies. We have to have technical skill, write, and explain. Some mathematicians are better at some of these than others, so we have benefited from collaboration. But we can't really specialize the same way as in the sciences, where you can have technical staff and project managers. Now with AI and other modern collaboration tools and full verification, it has become possible to run math projects where individual participants specialize in just one area. Maybe there are gaps in your collaboration—no one knows how to do the technical thing—but AI can plug some of the gaps. You still need humans because AI performance is very jagged. Some inputs can now be automated, but if you automate too much, for example, if you automate strategy generation but not verification, you get hundreds of possible strategies that you can't handle. But if verification also keeps pace, suddenly you have a new style of doing mathematics that is extremely effective.
是的,对此再快速补充一点。我完全同意 AI 能力目前非常参差不齐,所以你看到与人类非常富有成效的合作。
Yeah, just one quick comment on that too. I absolutely agree that AI capabilities are super jagged today, and so you see this really fruitful collaboration with humans.
探索另一面也很有趣,那就是有些 AI 系统比你想象的更接近人类,你必须以正确的方式注入大量强化学习,才能避免模型像人类一样放弃。你知道,如果你给一个太难的问题,模型常常会运行几个测试或提示,在它的思维链里想,‘啊,这个问题太难了。我觉得我做不到。让我假装很努力地尝试给用户看。’所以,我们在 Erdos 问题上看到,你让 AI 去解决一个 Erdos 问题,它做的第一件事就是去 Erdos 问题网站查一下,然后说,‘这是个开放问题,太难了,我不试了。’所以你必须说,‘不要用互联网,自己尝试解决这个问题。’其实挺简单的,我发誓。知道前沿研究其实就是引导模型按你想要的方式行事,这很好。
It's also interesting to explore the flip side of that, which is that some of these AI systems are more human-like than you imagine, and you have to pump a lot of RL in the right way to not have the problem have the models give up in the same way a human would. You know, if you give a too hard problem, often times the model can just run a couple of tests or prompts in its own chain of thought and be like, 'Ah, this problem's too hard. I don't think I can actually do it. Let me pretend to the user like I tried really hard.' So, we've seen with the Erdos problems that you get an AI to try to solve an Erdos problem. First thing it'll do is go to the Erdos problem website, look it up, say, 'It's an open problem. It's too hard. I'm not going to try.' So, you have to say, 'Do not use the internet. Try to solve the problem yourself.' It's actually pretty easy. I swear. It's good to know that frontier research is actually just about coaxing the models into behaving the way you want to.
这个愿景现在对这个房间以及更远的人来说可能相当有吸引力,我们基本上是在说这项技术从根本上让更多人能够合作解决这些问题。但这是否只是一个垫脚石,通向一个你只与许多 AI 智能体合作、它们逐渐主导这个领域的世界?
That vision right now is probably quite a compelling vision for this room and beyond where we're sort of saying the technology fundamentally empowers more people to collaborate on these problems. But is this just a stepping stone to a world where you're only collaborating with many AI agents and slowly but surely they come to dominate the space?
我认为既是也不是。我的意思是,我们今天做的数学可能会慢慢朝那个方向发展。但可能会有全新的做数学的方式,我们现在甚至无法想象。我认为数学是无限的,难度级别也是无界的。数学中甚至有些问题是无解的,我们知道它们无解。好吧,带个星号,但我不想谈这个。所以,有些东西即使是最强大的 AI 也无法做到……有一些密码学挑战,AI 现在无法挖出所有比特币。所以,我认为总会有一个前沿,而且我很确定,正因为人类和 AI(至少当前一代大语言模型)与人类技能的互补性,最佳组合永远是人类的一个凸组合,但组合的性质可能会随时间改变。
I think yes and no. I mean, I think the type of math that we do today might slowly kind of move in that direction. But there could be very new types of doing math that we can't even envisage right now, which would... I think math is infinite and the difficulty levels are unbounded. There are even problems in math that are unsolvable. We know they're unsolvable. Well, okay, with an asterisk, but I don't want to talk about it. So, there's certain things that even the most powerful AI can't... There are cryptographic challenges that AI cannot mine all the Bitcoin right now. So, I think there will always be a frontier and I'm pretty sure that just because how complementary human and AI, at least current generation LLMs, are with human skills, the best combination is always going to be a convex combination of humans, but the nature of the combination may change over time.
假设即使从哲学上讲,存在这样一个前沿,至少当前范式的 AI 无法跨越,需要某种核心的人机协作。在你看来,Mark,到达那个前沿是更聪明的强化学习训练的问题,还是仅仅原始算力的问题?如果今天我能给你无限的算力,你能加速到达那个前沿吗?
Let's assume even just philosophically there is this frontier beyond which at least the current paradigm of AI wouldn't be able to cross and some central human-AI collaboration is needed. Getting to that frontier in your mind Mark, is that a question of much smarter RL training or is it actually just a question of raw computation? If I could give you an infinite amount of compute today, would you be able to accelerate your way to that frontier?
是的,我认为当我思考 OpenAI 的整体研究计划时,它本质上就是关于如何改进算法,使它们能够扩展到我们明年和后年拥有的算力水平。所以它基于我们实际拥有的算力现实,我认为我们知道的算法都很简单且能扩展,但它们需要大量的工程和微调来确保它们真正扩展到下一个数量级及更远。一个很好的事情是,今天这是一个非常多维的问题。我们有很多轴可以扩展模型智能。我们可以扩大模型规模,构建更大的大脑,拥有更多的核心知识,这抓住了这样一种直觉:也许你广泛了解的数学越多,内化得越深,就越容易建立联系和跳跃。还有一个推理轴,我们也在扩展,这是将所有基础知识串联起来创造新见解的能力。我们在这个房间里有几个人正在研究我们在 GPT-5 直播中讨论过的事情,这有点像把这一点连接起来,让模型为自己生成新知识,真正在特定领域放大其知识。所以我认为有很多不同的轴将发挥作用,把模型带到下一个前沿。但总的来说,所有这些都基于这个度量堆。我们正在积极爬山,朝着越来越自主、更长周期的任务前进,我们看到这一趋势在继续。
Yeah, I mean I think when I think about the OpenAI research program overall, it is really fundamentally about how do we improve the algorithm such that they scale to the level of compute we have next year and the year after. So it's grounded in the reality of what compute we actually have and I think all the algorithms we know are simple and they scale, but they take a lot of engineering and fine-tuning to make sure that they truly scale to the next order of magnitude and the order of magnitude beyond. One really great thing is this is a very multi-dimensional problem today. There are many axes by which we can scale model intelligence. We can scale up the model and build these bigger brains with just more core knowledge and this captures the intuition that maybe the more math you know broadly, just internalize deeply, it's easier to make these connections and jumps. There's also a reasoning axis which we scale and this is the ability to take all of that base knowledge and chain it together to create new insights. And we have a couple people in this room working on this thing we talked about in the GPT-5 live stream which is kind of connecting this a little bit and having the models just generate new knowledge for themselves and really amplify its knowledge in certain domains. So I think there are a lot of different axes that are going to play into bringing the models to the next frontier. But overall, all of these things are grounded in this meter pile. Like we are aggressively hill climbing towards more and more autonomous longer horizon tasks and we see that trend continuing.
爬山这个词,如果我在与研究团队合作中学到了什么,那就是我们必须找到并定义一座要爬的山,也许这就是这两个世界交汇的地方——定义正确的山。所以,也许我们下一段就聚焦于此。First Proof 似乎是一个明显的例子,我们试图共同定义一座要爬的山。在你看来,Terry,这是否代表了你认为即将出现的新兴大众,还是 AI 一直在研究的更经典大众的最终形式?
The term hill climbing, if I've learned anything working with the research team is that we always must find and define a hill to climb and perhaps that's where these two worlds come together as defining the right hill. So, maybe we focus on that for the next segment. First proof would seem like an obvious example of where we've tried to co-define a hill to climb. In your mind, Terry, is that representative of what you're thinking could be the new emergent mass to come or is that the final form of this more classical mass that AI has been working on?
会有一个谱系。是的,First Proof 是一个非常有趣的实验,不同的人用 AI 工具生成的证明相当不错。我们实际看到的是存在一个明确的验证瓶颈。我们生成了很多证明,有些很糟糕,有些相当好,有些与文献中的类似,有些与作者自己的证明类似。有几个实际上与官方证明不同,这很有趣。但要仔细评估每个证明有多新颖、多有趣,我们实际上没有有效的方法。所以,我认为 First Proof 团队稍后会创建一个更结构化的竞赛,他们会有某种验证机制。为了充分利用 AI 的新能力,我们确实需要创建易于验证的挑战。所以,某种程度上,在变成垃圾之前你能使用的自动化和 AI 能力水平大致与你的验证严格程度成正比。所以,是的,我认为最初你会看到在那些足够基础、相对容易形式化的领域取得很多进展。所以,组合数学,我想你会看到。最早的问题肯定属于这一类。
There'll be a spectrum. Yeah, so First Proof is a very interesting experiment and the proofs that the various people with AI tools generated were quite good. What we saw actually was that there was a definite verification bottleneck. So, we had a lot of proofs generated. Some were terrible, some were quite good, some were similar to things in the literature, some were similar to the proofs that the authors themselves had. There were a couple which were actually different from the official proofs and so that was interesting. But to evaluate carefully exactly how novel and how interesting each proof was, we actually don't have a way of doing that effectively. So, I think the First Proof team are going to create a more structured competition later where they will have some mechanism for verification. So, in order to take full advantage of the new capabilities AI have, we do need to create challenges that are easily verifiable. So, somehow the level of automation and AI power that you can use before it becomes slop is roughly proportional to how stringent your verification is. So, yeah, I think initially you're going to see a lot of progress in areas either which are sort of elementary enough that they can be relatively easy to formalize. So, combinatorics I think you're going to see. The earliest problems definitely fall in this category.
有一些数值型挑战,你需要找到一个满足特定属性的数学对象。一旦你有了这个对象,验证它就非常容易。我们在物理学中看到了一些类似问题的例子。我认为我们会在那里看到很多进展。但数学的其他部分,目标不是找到一个满足特定属性的对象,而是找到一个好的总体理论来解释某些东西,或者一个好的定义。这些要难验证得多。例如,如果你想提出一个新的猜想或一个新的策略来解决一个未解问题,AI 可能会生成一百个可能的策略,但只有人类专家才能验证或给出有见地的意见。所以这将是一个瓶颈。即使 AI 将创造解决方案的成本降到零,仍然有其他巨大的瓶颈,这些瓶颈之前并不是我们关注的重点。
There are some numerical type challenges where you want to find a configuration, a mathematical object that obeys certain properties. Once you have the object, verifying it is very easy. We saw some examples in physics with similar problems. I think we will see a lot of progress there. But there are other parts of mathematics where the goal is not to find an object that obeys a certain property, but to find a good overarching theory to explain something, or a good definition. Those are much harder to verify. For example, if you want to propose a new conjecture or a new strategy to solve an unsolved problem, AI might generate a hundred possible strategies, but only a human expert can verify or give an informed opinion. So that will be a bottleneck. Even if AI drives the cost of creating solutions down to zero, there are still other huge bottlenecks that were not front and center in our minds.
我认为我们还需要更精确地陈述目标。AI 几乎太擅长不折不扣地实现目标了。所以你说‘我想解决这个问题,我想证明这个定理’,也许未来的 AI 就直接运行并证明了它。但实际上你想要的是人们努力工作、失败、寻找例子、联系文献,并交流所有部分结果。这才是解决特定问题的实际价值。如果你对 AI 的目标指定得太窄,就会错过大部分好处。所以我们必须更加小心地指定目标。
I think we also need to become much better at stating goals precisely. AI is almost too good at fulfilling a goal to the letter. So you ask, 'I want to solve this problem, I want to prove this theorem.' And maybe a future AI just runs and proves it. But actually what you wanted was for people to work hard, to fail, to find examples, to connect to the literature, and to communicate all the partial results. That was the actual value of solving a particular problem. There's a danger that if you specify your goal to an AI too narrowly, you miss out on most of the benefit. So we'll have to be more careful about goal specification.
我只想快速补充两点。我确实考虑过这种离线版本的首次证明。我们实际上也在讨论这一点。你可以想象,你训练一个模型,让它拥有到某个特定、非常详细的时间点的知识。然后你可以想象在那个时间点首次证明会是什么样子。现在你有了后见之明,你知道你追求的技术可能是什么,模型中的创造力可能是什么样子。我认为这些都是非常有趣的思维实验。有一个思维实验是关于选择哪一天作为截止点以获得最大信号。但我也思考数学的过程:它不仅仅是回答或证明一个定理。它是你在某处吸收的所有这些部分进展。我们在 OpenAI 有 AI 系统,它们就像是信息的中央存储库。你可以想象这在数学中也能发挥作用,就像一个全球图书馆。我知道 Daniel Lit 不久前在网上发表了一些东西,关于一个数学家可以与之交互的智能体。它填补了数学结果的凸包。你可以随时用它作为人们正在探索的内容的真实来源,它会为你连接很多点。它存储了我们所知道的东西。
Just two quick things to add. I do think of this offline version of first proof. We are actually discussing this a little bit too. You can imagine that you train a model with knowledge up to a specific, very detailed point in time. And you can imagine what a first proof would be at that point. Now you have the benefit of hindsight, you know what the techniques you're after might be, what creativity in the model might look like. I think those are very interesting thought experiments. There's a thought experiment of what day you would choose as a cutoff to get maximum signal. But yeah, I also think about the process of mathematics: it's not just answering or proving a theorem. It's all this partial progress that you assimilate somewhere. We have AI systems at OpenAI that are like central repositories for information. You could imagine that serving a function in mathematics as well, like a global library. I know Daniel Lit published something online a while ago about an agent that mathematicians can interface with. It fills out this convex hull of mathematical results. You can always use it as the source of truth for what people are exploring, and it will connect a lot of the dots for you. It stores what we know.
没错。有时关掉它可能是有用的。我曾经研究过一个我知道太多技术的问题。我知道有一个强大的方法可以解决问题,但它需要很多技术技巧。所以我用了它,解决了问题,然后发表了。然后有人指出,如果我用了更简单的工具,会得到一个更简单的证明。我确实担心,能够使用文献中已知的每一种技术并不一定是最好的前进方式。但拥有多样化的 AI 工具,仍然会有人以老派的方式做事、寻找更人性化的解决问题的方法为乐和自豪。
Right. Well, it may sometimes be useful to turn that off. I've worked on a problem where I knew too many techniques. There's a powerful thing I know will solve the problem, but it requires a lot of technical skill to use. So I do it, solve my problem, and publish. Then someone points out that if I had used a much simpler tool, I would have a much simpler proof. I do worry that having access to every single technique known in the literature is not necessarily the best way forward. But having a diverse array of AI tools, there will still be people who take pleasure and pride in doing things old school and finding more human ways to solve problems.
我想知道这种模式是否也适用于你给出的案例,Mark。确实,如果我们能回到某个科学领域特定范式转变之前的时间点,然后看看模型是否能预测它,那可以成为一种验证工具。但库恩式的范式转变观点也可能意味着未来会有一次范式转变,使之前的范式失效。所以你实际上不希望模型猜测之前的范式,因为它可能会让模型走上一条无法到达下一个范式的道路。这凸显了验证和确认的问题。这既是一个哲学问题,也是一个实践问题。在所有领域中,你可以说数学——免责声明,我确实在 Lean 上工作过一点,所以向那个团队致敬——拥有自动验证的能力,这是其他领域很少具备的。虽然不是完美的,也不是没有缺点。你的直觉是,对于所有其他知识领域,是否需要出现一个独立的验证工具,类似于数学中发生的情况,还是需要其他类型的范式?
I wonder if that pattern even plays to the case study you gave, Mark. It's true that if we could go back in time just before a particular paradigm shift in whatever domain of science and then see whether the model would predict it, that could be one verification tool. But the Kuhnian vision of paradigm shifts could also mean that there is a future paradigm shift that would invalidate the prior one. So you don't actually want the model to guess the previous one because it might take it off a pathway that doesn't get to the next one. It throws into relief the question of verification and validation. It's both a philosophical and a practical question. Of all the domains, you could argue that maths—and disclaimer, I did work a little bit on Lean, so big shout out to that team—has the capacity to do automated verification in a way that very few other domains do. Not perfectly and not without its own drawbacks. Is your instinct that a separate validation tool will need to come into existence for all the other domains of knowledge, mirroring what's happened in maths, or will it need some other paradigm?
我确信,在将 AI 注入工作流程时,存在一个上限,超过这个上限就会变成净损失,导致比解决的问题更多的错误和问题。最大的上限之一是验证能力。在数学中,我认为我们最有希望实现真正高水平的自动化,能够以在可信度较低、可验证性较差的领域中无法做到的方式有效使用高水平自动化,因为我们有很高的验证标准,至少对于证明特定事物这一具体任务——这不是我们唯一关心的事情——但证明我们已经指定要证明的事物。
I definitely believe that there is an upper bound on how much AI you can inject into a workflow before it becomes a net loss, causing more errors and problems than it solves. One of the biggest upper bounds is the ability to verify. In math, I think we have the best shot at getting really high levels of automation, being able to effectively use high levels of automation in a way that you couldn't do in less trustable and less verifiable domains, because we have a high verification bar, at least for the specific task of proving things—which is not the only thing we care about—but proving things that we've already specified we want to prove.
是的,尽管形式验证也有弱点,语言本身可能被恶意智能体利用。AI 可能会试图表现得有帮助,尽可能多地证明定理,但暗中向形式系统添加一些公理。你可以尝试关闭它们,但如果 AI 太强大,有时你不得不限制 AI 的能力,或者定期让人类参与。在其他科学领域,你可以做一些类似的事情。例如,数值模拟在某些情况下可以用作验证器,但同样不能完全依赖。假设你想模拟天气,你有一台超级计算机预测天气,而你并没有训练 AI 来模仿数值模拟。有可能在某个时候,AI 会利用数值模拟中不属于真实情况的特征。所以它只能工作到一定程度,然后就会失效。我们需要更好地了解验证器的局限性。许多验证系统在非对抗性使用时效果很好。但如果你专门训练 AI 来最大化基于这个验证器的输出,它就会找到漏洞。是的,AI 非常擅长这个。它是一个无情的作弊者。所以我们必须意识到这一点。仅仅因为一个验证器通过了人类测试,它可能并不适合 AI 使用。
Yeah but although even formal verification does have weaknesses, language itself can be exploited by malicious agents. An AI may attempt to be helpful and try to prove as many things as possible, but secretly add some axioms to the formal system. You can try to shut them down, but if the AI is too powerful, at some point you have to limit how capable your AI is or have humans periodically involved. In other sciences, you can do some of this. For example, numerical simulation can be used as a verifier in some cases, but again you can't rely on it. Say you want to model the weather and you have a supercomputer that predicts the weather, and you haven't trained an AI to mimic the numerical simulation. It is possible that at some point they will exploit some feature of the numerical simulation that is not part of the ground truth. So it will work up to a point and then it will stop. We do need to get a lot better at knowing the limits of our verifiers. Many verification systems work just fine if used non-adversarially. But if you're training an AI specifically to maximize its output based on using this verifier, it will find the exploits. Yes, AI is so good at that. It's a ruthless cheater. So we have to be aware of that. Just because a verifier passes a human test, it may not be suitable for AI use.
有道理。直观上,为了避免 AI 作弊,让某物可测量的最简单方法是从第一天、第一步就设计成可测量的。Mark,你在试图让模型变得更聪明时,是这样想的吗?你是从第一性原理出发,思考什么必须为真才能被测量为更聪明?还是纯粹依赖泛化来获得更聪明的模型?
That makes sense. Intuitively, to avoid AI cheating, the easiest way to make something measurable is to design it to be measurable from day one, from step one. Mark, do you think like that when you're trying to make the models ever smarter? Do you think in terms of first principles, what would have to be true to be measured as being smarter? Or do you rely purely on generalization to try and get ever smarter models?
是的,我认为归根结底,为什么我们在 OpenAI 这样的地方要攻克数学和物理?这实际上是因为我们缺乏好的评估,好的由人类编写的评估,而现在做科学本身就是评估。数学特别令人兴奋,因为你可以攻克某个定理,在很多情况下可以验证它,并且你有信心自己确实在推动前沿。我知道物理方面也有举措。在物理中,会有一些模糊的说法,比如“这个常数太小了”,但你仍然可以构建相当形式化的系统。所以这让我们能够真正推动数学和物理的前沿。但根本上,我们如此关注非形式语言推理的原因之一是我们关心泛化。我们也希望在生物学等领域进行深度推理,并取得突破,即使突破的定义有些模糊。在数学中,这要清晰得多:你解决了一个明显的问题,那就是一个重大突破。如果模型说,“嘿,这是你在机器学习中的下一个突破”,我不知道如何验证这是否属实。这非常经验主义,很多事情需要时间来判断。所以我认为我们关心的是这种基本的可泛化推理层。自然语言似乎是表达这一点的一个好方式,它不太容易陷入那种拥有一个技术工具箱、只专注于已知技术的陷阱。在自然语言中,我们能够表达新技术。到目前为止我们一直能做到。所以我们非常关心泛化,而且我确实认为这些形式化领域为我们提供了一种非常严谨的方式来测试我们是否在推动前沿。
Yeah, so I really think when it comes down to it, why do we care about attacking math and physics at a place like OpenAI? It really comes down to we are out of good evals, good human-written evals, and science doing science is the eval now. Math is particularly exciting because you can attack some theorem, you can verify it in many cases, and you feel confident that you're legitimately pushing the frontier forward. I know there are initiatives in physics, too. In physics there's a little bit more hand-waving around "oh this constant is too small" but you can still build pretty formal systems. So it allows us to really push the frontiers in both math and physics. But fundamentally, one of the reasons we cared so much about reasoning in informal language is we care about generalization. We want to be able to do deep reasoning in fields like biology, too, and create breakthroughs there, even if it's kind of fuzzy what a breakthrough means. In math it's much more clear: you solve an obvious problem, that's a big breakthrough. If the model says, "Hey, here's your next breakthrough in machine learning," I don't know how to verify if that's true. It's so empirical and time tells with a lot of these things. So I think what we care about is this fundamental generalizable reasoning layer. Natural language feels like a good way to express this in a way that falls less into the trap of having a tool bag of techniques and just centering on known techniques. In natural language, we are able to express new techniques. We've been able to do that so far. So we really deeply care about generalization, and I do think these formal fields give us a really rigorous way to test that we're pushing the frontier.
除了数学的结构性质以及你可以进行形式验证的方式之外,在这个领域追求更高能力还有其他实际好处吗?或者在你看来,这真的只是相当于一个评估?
Beyond the structural nature of math and the ways you can formally verify, is there some other practical benefit to pursuing ever greater capabilities in that space? Or is it in your mind really more just an equivalent of an eval?
嗯,我认为将数学作为其他用例的试验床的一个积极特点是,我们今天早些时候引用了 Vladimir Arnold 的一句话:数学是实验便宜的地方。它也是失败便宜的地方。所以这是相关的。如果你是一名工程师,被要求建造一座桥,而桥塌了,那是一个昂贵的错误。如果你是一名外科医生,切错了东西,那是一个昂贵的错误。但在数学中,如果你试图证明一个定理而你的证明不成立,那不是一个昂贵的错误。所以我们有这种失败的自由,这比其他学科更多。正因为如此,我们有一种从错误中学习的文化,这比其他学科要多得多。所以,与建造桥梁或心脏手术相比,用 AI 进行实验是一个相对更安全的地方。
Well, I think one positive feature for using math as a test bed for other use cases is that we had this quote earlier today from Vladimir Arnold that mathematics is the place where experiments are cheap. It's also the place where failure is cheap. So it's related. If you're an engineer and you're asked to build a bridge and the bridge collapses, that's an expensive mistake. If you're a surgeon and you cut the wrong thing, that's an expensive mistake. But in math, if you try to prove a theorem and your proof doesn't work, that's not an expensive mistake. So we have this freedom to fail, which is more so than in other disciplines. Because of this, we have a culture of learning from our mistakes a lot more than in other disciplines. So it's a relatively safer place to experiment with AI than, let's say, bridge building or heart surgery.
好的。是的,我喜欢你这么说。这也正是我们在 OpenAI 的思考方式。因为我认为从根本上说,我们开发 AI 的真正核心目标是用它来开发更强的 AI,对吧?我们想设计更好的实验来构建更强的模型、更智能的模型,这些模型反过来会做更好的数学,但也会构建更强的模型,从而形成飞轮效应。而这是一件昂贵的事情。如果你在任何方面搞砸了系统,你就会浪费很多钱,很多算力。所以我确实认为数学和物理是推动前沿的安全领域。
Okay. Yeah, I love that you say that. It's exactly the way that we think about things at OpenAI as well. Because I think fundamentally what we care about developing AI for, the really inner core goal, is to use it to develop stronger AI, right? We want to design better experiments to build stronger models, more intelligent models, which by extension will do even better math, but it will build an even stronger model and you get that flywheel. And that is an expensive thing. If you screw the system up in any way, you burn a lot of money, a lot of compute. So I do think about math and physics as safe domains to push the frontier.
是的,这很有道理。我的意思是,我甚至想知道你是否能进一步推动这一点。Kevin,你的工作涉及这个。假设在一个世界里,模型可能发现真正超越人类知识前沿甚至人类概念上无法理解的东西。大概需要某种方式将这些发现重新表示为逻辑链,至少我们能跟随步骤,即使不是实际组成部分。而数学比其他任何领域都更似乎发明了一种适应这种需求的工作流程。
Yeah, it makes a lot of sense. I mean, I even wonder if you can push that. Kevin, your work touches on this. Assuming a world in which the models may be discovering things that really are beyond the frontier of human knowledge or even the ability for a human to really conceptually follow. Presumably there needs to be some way to re-represent those findings into a logical chain that at least we can follow the steps if not the actual constituent parts. And math more than any other domain seems to have invented a workflow that accommodates that.
当然,与生物学或化学相比,它们没有那么多形式化的公理可以依据。所以,这可能是任何真正前沿科学进步的必要前提。但无论如何,这在某种程度上表明,我们未来思考数学的方式将会改变。我们可能更强调创造力、协作,以及不同于过去一百年的技能。那么,这是否会渗透到你教授数学的方式中?
Certainly compared to say a biology or chemistry which doesn't have as much by way of kind of formal axioms that you can work against. So, potentially it's a necessary precursor to any truly frontier science advancement to have this capability. Regardless though, it does to a point Terry suggests that the way that we think about doing math in the future will change. That we're emphasizing maybe creativity, collaboration, different skills perhaps than what have been happening the last 100 years. Does that filter then through into how you teach maths?
是的,这仍然是一个悬而未决的问题。所以在短期内,一些事情不得不改变。比如,每周的作业是第一个牺牲品。我认为我们现在可以推动学生去做更有雄心的事情。所以我更多地转向了基于项目的评估方式。在小班中,你可以进行口头评估。我们需要教授的技能将会不同。独立验证 AI 生成输出的能力将变得至关重要。软技能,比如如何与人合作。数学家过去在这方面并不都擅长,但我们必须变得更好。变化的速度如此之快,教育系统并没有迅速跟上,但我认为迫于必要,我们会被迫改变。以新冠疫情为例,我们对课程进行了一些紧急调整,效果还行,但体验并不好。所以希望这次我们能更有计划地进行,但变化的规模会是那样的。
Yeah, this is an open problem still how to. So in the very short term, some things have had to change. Okay, so like homework weekly homework assignments have been the first casualty. For instance, I think we can push our students to do more ambitious things now. So I've switched much more to a project-based type of assessment. In smaller classes you can do some oral assessment. The skills we need to teach will be different. Validation for independent ability to independently verify AI-generated output will become essential. Softer skills, how to work with people. Mathematicians have not been uniformly good at that in the past, but we'll have to get better. The pace of change is such that education systems are not catching up as rapidly, but I think by necessity we'll be forced to. With COVID for example, we did some emergency changes to our curriculum and it kind of worked. It was not a great experience. So hopefully this time we can do with a bit more planning, but it will be on that level of change, I think.
是的,我认为类似的情况是,我们的面试也很快变得不可靠了。如果人们有时间做带回家或书面的面试,那很难判断。我确实认为,转向一个模型可以与你互动、教你东西,并且模型本身可以判断你学了多少的世界,这感觉像是一个方向性的好更新。我考虑过改革面试形式,比如你说服模型你拥有在 OpenAI 工作所需的技能。当然,你必须防止黑客攻击和越狱之类的事情。但我真的认为教学从根本上必须以某种方式改变。我有点好奇 Terry 对几件事的看法。首先,我从其他教授那里听说,你确实看到了历史上最好的作业成绩和最差的现场考试成绩之间的分化。我不知道你是否也看到了这个趋势。第二,你是否真的看到了这种分化:那些非常有学习动力、善于使用工具的学生,以及那些真正感到加速的群体?
Yeah, I think the analog of that is, you know, our interviews became busted very quickly, too. I think if people have time to do some kind of take home or written type of interview, it's very hard. I do think moving to a world where you can have a model also just kind of interact with you and teach you things, and the model itself can judge how much you are learning, that actually feels like a directional really good update. I've thought about revamping interviews in the form of, you know, you convince the model that you have the skills necessary to work at OpenAI. Of course, you have to prevent hacking and jailbreaking and stuff like that. But I really do think fundamentally teaching has to change in some way. I am kind of curious to get Terry's take on a couple things. First, I've heard from some other professors that you really do see this divergence of the best homework and worst live exam scores in history. I don't know if that's a trend that you see as well. And I think the second thing is, do you actually see this divergence of students who are very motivated to learn that get really good using the tools, and is there this cohort that really feels accelerative?
是的,我确实注意到作业分数上升了,而现场考试分数下降了。但还不至于崩溃。我没有硬数据。我感觉到最弱的学生在用 AI 快速达到一个水平。而最聪明的学生通常倾向于避免使用 AI,因为他们担心用得太多会削弱自己的分数。最弱的学生我觉得他们觉得没什么可失去的。但一旦你有了一定的专业知识,这些工具就很棒。所以也许平衡点是实际上不鼓励使用,或者以特定方式使用。我确实可以看到未来,解决方案不是重点,因为任何人都能得到它。但比如,你用了什么提示词来得到解决方案?那可能是更有趣的评估工具。所以我们必须弄清楚。而且实际上,就像 AI 会优化任何奖励函数一样,我们给学生的奖励函数会产生很大影响。我们必须仔细思考。
Yeah, so definitely have noticed homework scores going up and in-person scores going down. Not so much a collapse level. I don't have hard data. I do get a sense that the weakest students are using AI to sort of get to an immediate level. And the brightest students generally tend to avoid using AI just because they're worried that using it too much atrophies their scores. The weakest students I think they feel like they have less to lose. But once you have a certain level of expertise, these tools are great. And so maybe the equilibrium is to actually discourage their use or use them in specific ways. I can certainly see in the future where the solution is not the point because anybody can get to it. But for example, what prompt did you use to get to the solution? And that might be the more interesting assessment tool. So we have to figure it out. And actually, just like AIs will optimize whatever reward function, the reward function we give to our students will make a lot of difference. We have to think this carefully.
是的,在某种程度上,极端情况很容易理解。完全认知卸载,因此没有学习发生等等。更微妙的情况可能是 AI 的生产性使用。我想你在最近的一次采访中用了一个比喻,就是说你被直升机送到了目的地,而不是走风景优美的路线。工作流程的变化对应着人类认知层面的某些东西。我们还不知道如果那样做会失去什么,但我们需要保持警觉。你有没有关于我们可能失去什么的论点?
Yeah, it's in some ways the extreme cases are easy to understand. A total cognitive offloading and therefore no learning is occurring etc. Maybe the more nuanced would be something where it is a productive use of AI. And I think you used this metaphor in a recent interview, which is to say you've been helicoptered to the destination as opposed to taking the scenic route there. There's something about the change in workflow corresponds to something at the cognitive level for humans. And that we don't know yet what it may be that you lose if you do that, but we need to be alive and alert to it. And do you have a thesis about what we might lose if we start doing that?
我认为我们很快就会从经验中看到。我认为我们只需要对研究或任何其他任务的所有不同方面有更多的认识。正如我之前所说,AI 允许许多事情解耦,这对分工有好处,更高效,但这确实意味着以前可以设定非常模糊的目标,因为人类试图达到这些目标时,也会顺便触及所有附近的目标。
I think we'll see empirically pretty soon. I think we just need much more awareness of all the different facets of research or any other task. So as I said before, AI allows for a decoupling of many things, which can be good for division of labor, more efficient, but it does mean that goals that previously it was okay to set very fuzzy goals because any human attempt to reach these goals would sort of also hit all the nearby goals as well.
所以正如我所说,使用 AI 时,如果你想去看山上一个漂亮的瀑布,你会徒步旅行,有时会看到有趣的野生动物,或者瞥见一个更好的地点,也许还会遇到其他徒步者并交谈。所有这些偶然发现都是自然发生的。过去,我们只会说去看瀑布是个好主意,但没有仔细分析为什么这样做以及实际的好处是什么。但现在我们有另一种方式到达瀑布:你可以乘坐 AI 直升机直接降落在那里。是的,你得到了 Instagram 照片,但也许那不是你唯一想要的东西。我认为不幸的是,我们必须通过经验来学习这一点。很难浪漫地谈论旅程,但我认为只有当我们失去它时,才会真正明白我们错过了什么。今天我们发布了学习成果测量套件,这是我们用来评估人类在使用模型时是否在学习的方法。所以这是一个正在进行的研究问题。但我想知道,马克,对你来说,偶然发现和模型给出不精确答案以创造探索空间——这是你感兴趣探索的模型品质吗?这是一个模型行为问题、个性问题吗?你如何处理这个问题?
So as I said with an AI, if you want to go see a nice waterfall on a mountain, you take a hike and sometimes you see interesting wildlife or get a glimpse of an even nicer location you might want to visit someday, and maybe you meet some other hikers and have a conversation. There's all this serendipity that naturally happens. In the past, we would just say it's a good idea to visit this waterfall, but we didn't unpack that carefully enough to see why we do that and what the actual benefits are. But now we have an alternate way to get to this waterfall: you can get an AI helicopter to drop you off there. So yes, you get your Instagram photo, but maybe that's not the only thing you wanted. I think unfortunately we're going to have to learn this by experience. It's hard to talk romantically about the journey, but I think only when we see what happens when we don't have that will we really understand what we're missing. Today we released our learning outcomes measurement suite, which is how we use the models to assess whether humans are learning when they use them. So it's a live research question. But I wonder, Mark, for you, serendipity and the idea of an inexact answer from the model to create space to explore—is that a quality of the models you're interested in exploring? Is it a model behavior question, a personality question? How do you grapple with that?
今年我们在构建 AI 的新基元和交互范式方面最大的举措之一是成立了一个交互式智能体团队。我认为仅仅向 AI 提问,然后一天后它返回最佳解决方案是不够的。人类是协作的;他们在这些结构中工作。举一个完全非数学的例子:你想创建一个 PowerPoint 或其他作品。你不会只是告诉 AI 做一个完美的 PowerPoint。你希望它返回结果,然后你引导方向,进行多轮交互。这就是与非常智能的智能体合作的样子,我们希望将这一点深度构建到模型中——使其非常可控,像一个思想伙伴。我希望在几个月内,至少一年内,AI 会变成这样。
One of the biggest initiatives we have this year in terms of building a new primitive and interaction paradigm with AI is that we started an interactive agents team. I think it's not sufficient that you just ask an AI a question and it comes back a day later with its best attempt at a solution. Humans are collaborative; they work in these constructs. Take a completely non-math example: you want to create a PowerPoint or some artifact. You don't just tell an AI to make a perfect PowerPoint. You want it to come back, and you shape the direction it's going, with multiple rounds of interaction. That is what it's like to co-work with a very intelligent agent, and we want to build that deeply into the model—something that's very steerable, feels like a thought partner. I hope within a couple months, at least within a year, that's the way the AI looks.
在协作上进行强化学习要困难得多。你如何评分与同事的默契程度?
It's much harder to do reinforcement learning on collaboration. How do you score how well you're vibing with your co-workers?
我的论点是这是可能的。我同意这可能没有你想象的那么困难。现实世界中默契有相当清晰的生物信号,你可以将其带回机器世界。
My thesis is that it is possible. I agree that it might not be as difficult as you imagine. There are quite clear biological signals of what vibing looks like in the real world that you can bring back into the machine world.
所以用肢体语言赋予 AI 具身性。我认为它们也在默契。我很高兴我得到了一个会被保留的引语。之后我会开放提问。我再问一个问题,给大家时间思考。也许一年后我们会再次见面,我很有兴趣知道你对一年后我们会在哪里的预测。
So embody your AI with body language. I think they're vibing too. I'm glad I've got a quote that will be kept after this. I will open up the floor for questions after this. I'll ask one more just to give you all time to think about it. Perhaps we will meet again in a year, and I'd be interested to know what your predictions are for where we'll be in a year's time.
我真的希望我们会看到许多新型的、基于挑战的方法论项目,比如第一证明类型的事情,其中一些数学家群体将创建一套非常好的创造性问题,他们希望至少有人能解决其中一个。他们有很好的难度梯度和验证协议,然后向社区开放。这充分利用了不仅是 AI,还有互联网和梅特卡夫定律。如果有 n 个人能产生问题,n 个人能解决问题,那么就有 n 平方种可能的连接。数学家们一直不擅长利用这种大规模网络。所以我认为我们会看到一种不同的数学研究市场风格,AI 将在其中大放异彩。这是我想看到的一件事,也许一年后我们会开始看到。
I really hope we're going to see a lot of new types of methodical projects that are challenge-based, like first proof type things, where some group of mathematicians, for example, will create a really good creative set of problems that they would like at least one to solve. They have a very good gradation of difficulty and a very good verification protocol, and they would just open it up to the community. It's taking full advantage of not just AI, but also the internet and Metcalfe's law. If there are n people who can produce problems and people who can solve problems, then there are n squared possible connections. Mathematicians have been very bad at using that sort of large-scale network. So I think we will see a different marketplace style of doing mathematics, and there AI will shine. So that's one thing I want to see, and maybe in a year we'll start seeing that.
我认为机器学习在这里预示了数学的未来。当你看到前沿实验室今天如何与研究科学家合作时,我们正在进入这样一个世界:最强大的研究科学家能够并行追求许多想法,并充当协调者。他们可以思考一个想法,考虑实验的多种变体,然后让模型去执行。我希望数学中也有类似的范式,像 Terry 和你这样的人能够自由探索广泛的想法和策略,几乎不需要手把手的指导。我确实认为“几乎不需要手把手”这一点会变得更加真实。任务的时间跨度将继续延长。就像一年前我们只有几分钟的时间跨度,我认为一年后我们将达到多天的时间跨度,你可以真正信任模型去完成需要那么长时间的任务。再往后,就是确保交互无缝衔接。这些东西应该感觉与人类群体以及你们所在的社区非常自然地互动。最后,我真的希望我们能有重大突破,无论是在数学、物理还是生物学。我认为今天我们证明的东西很好,但我确实认为这有潜力产生对人类非常有益的东西。
I think of ML's kind of foreshadowing math here. When you look at how frontier labs operate today with research scientists, we are moving into this world where the strongest research scientists are able to pursue a lot of ideas in parallel and just act as orchestrators. They can think about an idea, think about a bunch of variations in the experiments, and have the model go and execute that. I hope there is a similar paradigm in math where people like Terry and yourselves feel empowered to explore a broad set of ideas and strategies with very little hand-holding. I do think the very little hand-holding part will become more true. The task horizon will continue to elongate. Just like we were at minutes of horizon a year ago, I think in a year from now we're going to be at multiple days where you can actually trust the model to do tasks that would take you that long. And then beyond that, it's just making sure the interaction is seamless. These things should feel like they interact very naturally with groups of humans and with the communities you operate in. Finally, I really do hope we have some really big breakthrough, whether in math, physics, or biology. I think today the things we're proving are good, but I do think there's potential for this to produce something very beneficial for humanity.
软件开发中有个关于集市和大教堂的比喻。集市是一个自发组织、非常多样化的东西。
There was this metaphor in software development of bazaars and cathedrals. A bazaar being a self-organizing thing that springs up and is very diverse.
一座大教堂由一位伟大的建筑师设计,因此非常优雅。或许我们的想法是,我们能让这两种现象同时发生,并希望这能带来繁荣。所以,我想开放提问。有人有迫切的问题吗?
A cathedral being one great mind architects and is therefore very elegant. And that the idea perhaps will be that the mass we get both those phenomena occurring and hopefully that is a flourishing. So, I do want to open up to questions. Does anyone have a burning question?
是的。我想知道您能否谈谈世界模型,它们不是预测下一个词元,而是预测下一个状态。我读到关于视频生成模型的一些东西,说它实际上可以自我纠正,因为可以避免幻觉。这是真的吗?因为我在我的模型上能运行的只是一个 hello world 程序,所以我不太了解。
Yeah. I wonder if you can talk about the world models which instead of predicting the next token predict the next state. And what I'm reading about the video party ways is that it can actually self-correct because hallucination can be avoided. Is that all true? Because whatever I could run on my match contest model is just a hello world thing. So, I don't know much about it.
为了观众的利益,这个问题是关于我们如何看待世界模型?它们是否是一种范式转变,是否特别有助于数学和解决幻觉?
For the benefit of those watching, the question was about how do we feel about world models? Are they a paradigm shift and would they be specifically useful for maths and solving hallucinations?
我认为这是一个非常有前途的替代方向。LLM 很棒,实际上在某些方面它们太棒了,以至于我们几乎将整个 AI 基础设施都围绕让 LLM 尽可能强大来构建,这可能会排挤其他一些非常互补的方式来创建具有完全不同特点的 AI 助手。所以,我绝对支持对世界模型的研究。我认为在很长一段时间内,它们会不如 LLM,因为 LLM 拥有所有的动力和基础设施。这就像我们围绕汽车和汽油建设了城市,拥有了整个基础设施,这实际上让替代的机动交通工具很难突破。但确实有人在推动这件事,我祝他们好运。
I think it's potentially a very promising alternate direction. LLMs are great and in some ways they're too great, actually, in that we've kind of routed our entire AI infrastructure around making the LLMs as powerful as possible, and it could crowd out some other very complementary ways to create AI assistants that have a jagged in a completely different way. So, definitely support research into world models. I think for a long time they will underperform the LLMs because of all the momentum and infrastructure LLMs have. It's like we have built our cities around the automobile and gasoline, and we have this entire infrastructure, and it's actually making it hard for alternate motor transportation to break through. But there are definitely people who are pushing that, and I wish them a lot of luck.
是的。我确实认为,当你想到一个纯视频生成的世界模型时,我们似乎还差得很远。我认为现有的视频模型是相当好的物理模拟器,但稍微施加一点压力,它们也会崩溃。我确实想象随着时间的推移,它们会变得越来越稳健,但还没有完全达到。我们在这方面相当努力。我认为世界模型有很多种。你也可以把 LLM 看作一个世界模型。但我认为数字世界模型,即我们与计算机交互,拥有计算机的所有规则和反馈,这是一个非常重要且有趣的系统,我确实认为我们很快就能攻克它并从中获得巨大价值。
Yeah. I do think when you think about a pure video generative video world model, we still seem pretty far from that. I think the existing video models are pretty good physics simulators, but with a little bit of our own pressure they also fall apart. I do imagine that'll get more and more robust over time, but it's not quite there yet. We are pushing fairly hard on that. I think there's many spectrums of world models. You can see an LLM as a world model, too. But I think digital world models where we're interfacing with computers, with all the rules and feedback of a computer, that's a very important and interesting system, and I do think we'll tackle and really get a lot of value from that very soon.
我想知道是否存在一个中间地带,即你可以构建基于物理定律或遵循物理定律的 LLM 环境,这在某种程度上将你从世界模型中获得的好处赋予 LLM 或其他东西。所以,也许这是一个交叉点,而不是两条替代路径。
I wonder on that if there is a middle ground in as much as you can construct LLM environments that are based on the laws of physics or follow the laws of physics, which to some degree confers the benefits you otherwise get from a world model to an LLM or anything else. So, perhaps it's an intersection rather than two alternate pathways.
是的。所以,AI 在科学中,许多领域,当它预测得很好时是有效的。例如,蛋白质折叠,准确预测天气。但在数学和理论物理学中,我们要求的是不同的东西,比如我们想要得到一个公式,一个证明。但这会不会太局限了?让 AI 为我的假设提供一个证明,但你的大脑太狭窄无法理解,会不会更容易?所以,我可以教其他 AI 关于它,并取得更多进展。是的。这已经是我和 AI 的关系了。所以,我们有些人已经达到了那个前沿。
Yeah. So, AI in science, many fields, is effective when it's predicting very well. For example, protein folding, predicting weather accurately. But in mathematics and theoretical physics, we're asking for something different, like we want to get a formula, get a proof. But is it potentially too limiting? Will it be easier to get an AI to potentially have a proof for my hypothesis, but your brain is too narrow to understand it? So, I can teach other AIs about it and do more progress with it. Yeah. That's definitely my relationship with AI already. So, some of us have hit that frontier.
再次,为了观众的利益,问题是:在某些科学领域,我们对纯模拟感到满意。如果它能做到这件事,我们就认为这件事被证明了,即使你不能正式验证它。所以,如果它能准确预测天气,即使我们不知道如何做到的,我们也会对这个结果感到满意。我们对数学和物理学持有不同的标准,即它必须按照我们讨论过的方式是可验证的。施加这种限制是否在某种程度上是限制性的或错误的?
Again, for the benefit of those watching, the question was in some domains of science, we're satisfied with pure simulation. If it can do the thing, we consider the thing to be proven, even if you can't formally verify it. So, if it can predict weather accurately, even if we don't know how, we're kind of happy with that outcome. We hold maths and physics to a different standard, which is that it must be verifiable in the way we already discussed. Is that somehow limiting or a mistake to put that restraint?
是的,我认为会有不同类型的数学任务,我们现在不做,但我们会信任 AI 去做。它们可以非常互补于获得问题的形式证明的任务。所以,给你一个类比,在现在的国际象棋中,所有棋手都使用这些象棋引擎进行训练。象棋引擎做的一件事是随时给你这个局面的分数,比如白方领先三分之类的。这是一个训练人类棋手的非常好的信号。你得到即时反馈,哦,那是一个很糟糕的走法。好的,我试试这个。我可以想象一个 AI,人类试图做一个证明,每次你说,我要尝试证明一个矛盾,你的分数就会下降。等等,好吧,那是个坏主意。好的,你应该退回去做别的事情。所以,也许一个小电击……所以,那是一个专业模型。是的,好的。所以,是的。一个好的数学导师可以做到这一点,但也许我们必须创造性地思考 AI 可以帮助的任务类型,这些任务我们今天根本没想到。
Yeah, I think there'll be different types of mathematical tasks that we don't do nowadays, which we would trust AI to. And they could be quite complementary to the task of getting a formal proof of a problem. So, to give you an analogy, in chess nowadays, all chess players train using these chess engines. And one thing that a chess engine does is it gives you this score at any time in this position, you know, it's white is three points ahead or whatever. And it's a really great signal to train human chess players. You get instant feedback, oh, that was a really bad move. Okay, I'll try to do this instead. I could imagine an AI where a human is trying to do a proof, and every time you say, I'm going to try to prove a contradiction, your score goes down. Wait, okay, that's a bad idea. Okay, you should back up and do something else. So, maybe a small electric shock just to... So, that's a pro model. Yeah, okay. So, yeah. A good math tutor can do that, but maybe we have to be creative about the type of tasks that AI could help with that we just don't think about today.
是的,我的意思是,我确实认为验证很重要。它不一定是形式验证。我认为我们深切地想知道为什么某件事是真的。而且我认为这实际上是更深层次的对齐问题的一部分,对吧?当 AI 处理现实世界的影响任务时,你想知道它为什么做出某个决定。假设你决定这是发展业务的最佳策略,对吧?你不希望它没有充分理由就那样做。所以,我们有很多对齐技术,比如辩论,对吧?即使你不一定得到一个滴水不漏的形式化东西,你也可以大致理解轮廓,与证明互动并质疑它。所以,是的,我确实认为对辩论和对齐等技术的投资将在未来真正帮助我们。所以,我不……
Yeah, I mean, I do think verification is important. It doesn't have to be formal verification. I think we deeply want to know why something is true. And I think that's actually part of a deeper alignment problem, right? When the AI is attacking real-world impact tasks, you want to know why it made a certain decision. Let's say you decided here's the best strategy to grow a business or something, right? You don't want it to do that without having a good justification. And so, we have a lot of alignment techniques like debate, right? Where even if you don't get necessarily an airtight formal thing, you can kind of understand the outline and interact with the proof and question it. So, yeah, I do think investment into techniques like debate and alignment will really help us in the future. So, I'm not...
请问你能在潜在空间之类的东西中获取信息吗?比如看看中间层在想什么?
Please can you get anything in the sort of the latent space stuff like see what the interview layer is thinking?
是的,是的。所以,这是我们经常研究的东西,对吧?我认为首要的事情是,能够监控思维链的推理,你实际上可以从那里获得很多见解。
Yeah, yeah. So, that's something we look into a lot, right? I think the top-order thing is, you know, just being able to monitor the reasoning of the chain of thought and you actually get a lot of insight from there.
这两个问题可能相互关联:对于可解释性而言,我们压缩潜在空间的方式确实限制了模型可能产生的关联类型。在什么情况下你会决定为了理论上的新连接而保留潜在空间,即使牺牲可解释性?或者你认为我们需要一种不同的可解释性范式,不需要这种压缩?
These two questions may link up in as much as for interpretability, the way that we collapse the latent space does limit the kind of associations that could come out of the model. At what point do you decide that it's better to maintain the latent space because of the theoretical new connections it can make at the cost of interpretability? Or do you just think that we need a different interpretability paradigm that doesn't require that kind of collapsing?
我认为我们今天在文本空间操作的原因是,可解释性带来了巨大的好处,对吧?你可以通过观察模型出错的地方来调试很多问题,比如‘哦,这里推理错了,我们可以去调试’。当你在纯粹不可解释的潜在空间中操作时,你就失去了这种能力。我不确定长期来看我们是否会转向那种方式。理想情况下,我们应该有多样化的模型。所以可能有些应用你只想要答案,不关心可解释性,那么你就把旋钮转向一边。但其他应用你确实想看到过程,想看到人类可读的思维链等,那么你就把旋钮转向另一边。
I think the reason we operate in text space today is interpretability buys you so much, right? I think you can debug so many things that go wrong with the model by just being like, 'Oh, well, clearly it's like reasoning wrong here, so there's you know, we can go and debug.' When you're doing something in just pure uninterpretable latent space, you lose that. And I don't know that we would switch to something like that in the long term. Well, ideally we should have a diversity of models. So maybe there are some applications where you just want the answer, you don't care about interpretability, and then you just turn the dial one way. But there would be other applications where you really want to see the process, you really want to see a human-readable chain of thought or whatever, and you turn the dial the other way.
是的。直觉上,将事物压缩成语言确实是有代价的。所以即使模型有优先级,你可能也需要表达或验证方法的优先级,以免强行将其塑造成不合理的形状。
Yeah. I mean, it does feel intuitively like having to compress things into language does come at a cost. So even alongside a priority of models, you might also want a priority of expressions or verification methods so that you don't somehow force it into a shape that doesn't make sense.
我有一个关于归因的问题。以 AlphaFold 为例,世界普遍认为是 AI 解决了这个问题,但它建立在蛋白质数据库和数十年的努力之上。然后蛋白质数据库却失去了资金。当我们考虑大规模数学问题以及产生问题集和验证所需的人力时,存在一种危险:虽然 AI 是关键推动者,但它并非独立发生,而是一个生态系统。那么我们如何控制这种叙事?对于 OpenAI 和大公司来说,他们如何应对这种责任?我们如何避免走向坏的方向?
I have a question about attribution. One thing with AlphaFold, the world largely thinks AI came and solved that problem, but it was sitting on the protein data bank and decades of effort. And then you see the protein data bank loses its funding. As we think about large-scale math problems and all the human effort needed to produce sets of problems and verify them, there is a danger that while AI was the critical enabler, it didn't happen on its own. It's really this ecosystem. So how do we control that narrative? And for OpenAI and big companies, there's responsibility about how they navigate that. How do we avoid it going in a bad direction?
这是一个重要观点。部分解决方案:正如我所说,我预见到挑战问题的兴起,人们会创建他们想要解决的任务数据集。这是双赢的,因为创建数据集的人会得到部分问题的解决,这正是他们想要的。但这些数据集对校准 AI 非常有用。所以有些情况下是双赢的。但确实有些情况是,人们花费巨大代价构建数据集并非为此目的,然后它被各种 AI 吸收。我不知道我们能多好地追踪这一点。这涉及到知识产权法。这是一个非常棘手的问题。
This is an important point. A partial solution: as I said, I envisage the rise of challenge problems where people will create datasets of tasks they want solved. There it's win-win because the people who create these datasets will get some fraction solved, which is what they want. But these datasets could be very useful to calibrate AIs. So there are cases where it can be win-win. But there are definitely cases where people have built a dataset at great expense for not this reason, and then it gets absorbed into various AIs. I don't know how well we can track that. It leads into intellectual property law. It is a very tricky problem.
好的,我把这个问题抛给你来处理。
Okay, which I will toss to you to deal with.
不,我认为目前 AI 并不想要你的功劳。所以,我认为 OpenAI 对科学的愿景并不是我们索取功劳。我们当然有推动科学进步的雄心,但 Kevin,你想建立一个平台,让全世界的数学家能够整体加速这个领域。我认为我们不知道正确的问题是什么。我们不是 OpenAI 内部的协调者。我认为功劳应该归于你们。我知道这不完全是你的问题。
No, I think as of now, the AI doesn't want your credit. So, you know, I do think the vision we have for OpenAI for science is really not about us claiming the credit. I do think we certainly have the ambitions to move science forward, but Kevin, you want to build a platform where mathematicians around the world can just accelerate the field in its totality. I think we don't know the right questions to ask. We aren't the orchestrators within OpenAI. I do think the credit should just go to you guys. I know that's not exactly the question you're asking.
我认为并不是 AI 会索取功劳,而是公众认为 AI 解决了这个问题,而人类、实验和数据似乎无关紧要。我们看到的情况是,当有一个无人问津的开放最早问题时,某个 AI 解决方案突然出现并登上社交媒体:‘AI 解决了一个未解问题。’很多情况下,24 小时后有人使用深度研究工具发现,这个结果已经在文献中用非常相似的方法证明了。我们不能确定 AI 是否使用了那个解决方案或间接知道它,但这种情况非常频繁。我们有一整张 AI 贡献表,专门有一节就是关于这个的。所以某种程度上,我们至少有能力检测到其中一部分,因为我们也有这些研究工具。有时我们可以恢复归因。这并不完美。但可能正是同一项技术,让我们能够利用文献解决问题,也能利用文献归因解决方案。
I think it's not that AI is going to claim credit, but the public perception is AI solved this problem and then somehow humans and experiments and data are not like that. So what we've seen is that when there's an open earliest problem that no one has looked at and then some AI solution just gets a solution and hits social media: 'AI solved an unsolved problem.' In many cases, 24 hours later someone with deep research tools uncovers that this result was already proven by a very similar method in the literature. We can't say for sure whether the AI used that solution or was indirectly aware of it, but it happened so frequently. We have a whole table of AI contribution, a whole section just this. So to some extent, we at least have the capability to detect some of this because we also have these research tools. So we can recover attribution sometimes. It's not perfect. But it could be that the same technology that allows us to use the literature to solve problems can also use the literature to attribute solutions.
在此基础上再补充一点。一般来说,数据归因本身就是一个非常困难的问题。当你生成某样东西时,它在多大程度上受到哪些数据点的启发?一个有趣的想法是,也许新颖性和贡献与模型思考某件事所花费的时间有一定相关性。也许并不总是如此,但我想,除了它在文献中重新发现的内容之外。
Just one thought on top of that. In general, it is a very hard problem just data attribution. When you generate something, how inspired is it from which data points? One interesting thought here is that perhaps the novelty and contribution is somewhat correlated to the amount of time the models spend thinking about something. Maybe it's not always true, but I think, modulo stuff it rediscovers in literature.
不过我觉得这里面有公关成分,比如 DeepMind 本可以更主动地提及蛋白质数据库和叙事。有些问题与其说关乎 AI 模型和数据,不如说关乎我们如何谈论它和公关。
I think there's a PR element though, like DeepMind could have gone out of their way to say more about the Protein Data Bank and the narrative. There are things that are not about the AI model and data so much as about just how we talk about it and PR.
是的,这点说得非常对。我理解其中的动机。我希望 OpenAI 的任何人能证实,我们非常重视诚信。我会极力争取正确的叙事。也许从我之前的经历补充一点:一个范式转变是,公众往往低估科学进步在多大程度上依赖于根本性地改进工具。你需要更好的显微镜等等。这些不是光鲜的工作,通常不是理论科学家喜欢做的。我们应该传达的叙事是,通过为科学建造更好的工具,我们让人类加速整个领域。常常带来巨大的进步。AlphaFold 确实是一个工具,它本身不是科学研究,尽管它也有一些研究成分。如果我们能强调它作为工具,这也许能回应你关于社会投资流向的观点,因为需要围绕工具的东西使其有效,但最终服务于使用工具解决问题的人类研究者或科学家。时间快到了,如果允许的话我再问一个。
Yeah, that's a very fair point. I understand the incentives. I do hope anyone here at OpenAI can verify this, but we care a lot about integrity. I would very much strongly fight for the correct narrative there. Maybe to add a flavor from my previous life: one paradigm shift is that the general public tends to underestimate the degree to which the progress of science relies on improving tools fundamentally. You need better microscopes, and so on. These are not glamorous jobs, not typically what theoretical scientists like to work on. The narrative we should be saying is that by building ever better tools for science, we enable humans to accelerate the entire field. Often, really big steps forward. AlphaFold is really a tool. It's not per se scientific research, though it does have some research components. If we can emphasize it being a tool, that might also flow back to your point about where investment flows in society, because it needs things that surround the tool to make it effective, but ultimately in service of human researchers or scientists who will use the tool to solve the problem. We are just at the edge of time, so I'll take one more if that's all right.
一个非常有趣的事情是使用 AI 加速数学和物理学。我认为能够解决 AI 或数学问题,然后这如何帮助解决其他物理问题,存在某种协同效应。我很好奇你对 OpenAI 出现的其他协同效应的看法,你们处理数学、物理学,以及你们未来还看到什么。
One of the really interesting things that's happened is accelerating math and physics using AI. I think there is a certain synergy from being able to solve AI or math and then how that can help solve other physics problems. I'm curious to hear your thoughts on additional synergies you see coming out of OpenAI, you guys tackling math, physics, and what else you see out there in the future.
Kevin 可能是回答这个问题的最佳人选,但我们也关心处理数学和物理学之外的领域。我们探索过的一些领域是生物学,我们让 AI 致力于使湿实验室的生物流程更高效。与我们的合作伙伴之一 Ginkgo Bioworks 一起,我们迭代了他们许多核心流程,使合成蛋白质的成本降低了 40%。这只是推动更多进步的基本原语。你可以想象在材料科学和其他领域可以做更多事情。
Kevin would probably be the best to speak to this, but we do care about tackling domains outside of math and physics as well. Some that we have explored are biology, where we have had the AI work on making biological procedures in the wet lab much more efficient. With one of our partners, Ginkgo Bioworks, we iterated on a lot of their core processes and made the cost for synthesizing proteins 40% more efficient. That's just the underlying primitive that will drive more progress. A lot more you could imagine doing in material science and other domains.
IPAM 研究所的核心使命就是找到这些协同效应。通过纯数学与应用数学研究所,我们把不同的社区聚集在一起交流。很多是偶然发现。我们不会随意把领域拼凑在一起,但我们会选择那些我们认为会产生大量意想不到的富有成效合作的领域。我很高兴你把随机碰撞粒子留给物理学家。
The IPAM institute's core mission is basically to find these synergies. With the Institute for Pure and Applied Mathematics, we bring together different communities to talk together. A lot of it is serendipity. We don't just randomly smash together fields, but we do pick ones where we believe there will be a lot of unexpected fruitful collaboration. I'm glad you leave randomly smashing together particles to the physicists.
是的。
Yeah.
这是一个很好的结束时机,也顺便提一下 Kevin 即将做一个讲座,涵盖了许多这些问题:要加速科学的各个部分需要什么条件,以及为此构建的工具。希望你们能来参加,非常感谢大家。
That's a good moment to close it and also to plug that Kevin is about to give a lecture which covers many of these questions: what needs to be true to accelerate all parts of science and the tools to build for it. I hope you'll come to that, but thank you all so very much.
嗯。
Mhm.