Axiom Math:构建自我改进的推理引擎与 AI 数学家

Axiom Math: Building a Self-Improving Reasoning Engine with AI Mathematician

洪乐潼 Carina Hong · Gradient Dissent · 2026-02-05 · 约 51 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Axiom CEO Carina Hong 讨论他们的 AI 数学家如何结合生成与验证,在普特南考试中取得最高分,以及摇滚乐如何影响她的创业之旅。

Axiom's CEO Carina Hong discusses how their AI mathematician combines generation and verification, achieving top scores on the Putnam exam, and how rock and roll influences her startup journey.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

0. Axiom简介及其使命 Introduction to Axiom and its mission

Host

欢迎收听《梯度下降》,一档关于让机器学习在现实世界中落地的节目,我是主持人 Lucas Bewald。Carina Hong 是 Axiom Math 的 CEO 兼创始人,这家公司自称是一个自我推理系统,目前正处在构建 AI 数学家的前沿。他们在数学奥赛问题上取得了惊人的成绩,最近还在我们即将谈到的普特南测试中拿到了最高分。Carina 非常有趣,她曾是一位明星学术数学家,如今转型成为创始人。Axiom Math 在秘密运营后公开亮相,已融资 6400 万美元。她详细讲述了 Axiom 的工作原理、她认为这项技术如何随时间推移超越数学领域,以及中国硬核摇滚乐如何影响了她对管理和保持饥饿感的思考。希望你喜欢这次访谈。Carina,我们先从最基本的问题开始吧:什么是 Axiom,它在高层面是如何运作的?

You're listening to Gradient Descent, a show about making machine learning work in the real world, and I'm your host, Lucas Bewald. Carina Hong is the CEO and founder of Axiom Math, which is a company that calls itself a self-reasoning system and is currently on the cutting edge of building an AI mathematician. They've had astonishingly strong results on math olympiad problems and recently got the highest score on the Putnam test which we talk about. Carina is super interesting. She was a star academic mathematician now turned into a founder. Axiom math came out as stealth having raised $64 million. She talks in detail about how Axiom works, how she think it applies to more than math over time and also how rock and roll or hardcore rock and roll music in China influenced her thinking around management and staying hungry. I hope you enjoy this interview. Carina, why don't we start by the most basic question which is what is Axiom and how does it work at a high level?

Carina Hong

是的。Axiom 的使命是构建一个自我改进的推理引擎,它结合了生成与验证,我们认为这是当前 AI 领域被忽视的一个组成部分,我们想从 AI 数学家开始,因为数学是这种自我循环的绝佳试验场。我们使用像 Lean 这样的形式语言来锚定自然语言部分,正因为如此,我们可以解锁许多有趣的能力,且样本效率高得多。所以 Axiom 的总体愿景是:有一个证明器,一个能证明定理的系统;还有一个猜想器,一个能提出有趣猜想的系统;它们相互对话。它们相互对话,还有第三个组件叫知识库。所以如果我是证明器,我想知道哪些已经被证明、哪些我可以使用。如果我是猜想器,我想压力测试这个猜想是否合理。所以也许某个反例会帮助我,干脆不提出那个猜想。知识库是支持这两者的非常重要的组件,还有自动形式化,即能够自动将自然语言内容形式化为形式语言内容的系统。它把三者编织在一起。这就是 Axiom 的基本设置。

Yeah. So, Axiom's mission is to build a reasoning engine that is self-improving and that combines generation and verification, which we think is an overlooked component in the current AI landscape, and we want to start with an AI mathematician because math is a really good testing ground for this sort of self-loop. We use formal languages like Lean to ground the natural language counterpart, and because we are doing that, there are a lot of interesting capabilities we can unlock with much higher sample efficiency. So the general vision of Axiom is there is a prover, a system that can prove theorems, and also a conjecturer, a system that can propose interesting conjectures, and they talk to each other. They talk to each other, you have this third component called knowledge base. So if I am a prover, then I would like to know what are already proven and what I can use. If I'm a conjecturer, I would like to stress test whether this conjecture is reasonable. So perhaps some counterexample will help me out by just not suggesting that at all. And the knowledge base is a very important component supporting these two, and then auto-formalization, which is the system that can automatically formalize natural language stuff into formal language stuff. It's kind of weaving all three. So that's the basic setup of Axiom.

1. 简单示例演示Axiom运作 How Axiom works on a simple example

Host

那么,我们能不能在每个人都能理解的简单背景下分解一下,比如如果我们试图证明勾股定理,假设它不在你的知识库里,实际中会如何运作?

So can we break it down in the context of something simple that everyone could understand, like if we're going to try to prove the Pythagorean theorem, for example, assuming it's not in your knowledge database, how would that work practically?

Carina Hong

对,我觉得如果是勾股定理,它已经被证明过,希望是数百万次了。但如果我们做一个全新的数学问题,比如我们最近实际证明的,我们在 2020-2026 年普特南考试中取得了非常高的分数,因为他们在 12 月命名,按次年命名。所以我们在考试时限内答对了 12 题中的 8 题。我们会公布实际最终分数,但我们也公布了 12 题中的 9 题,因为第 9 题是在我发推两小时后解出的,这有点尴尬。所以你会想,哦等等,大家别转发那条推文了,转发这条吧。

Right, so I think if it's Pythagorean theorem, which is already proven, hopefully, millions of times. But if we do a de novo math problem, like we recently actually proved, we got a really strong score, winning score on Putnam exam 2020-2026 because they name it in December, they name it by the next year. So we within the exam time limit got eight out of 12. We are going to announce our actual final score, but we also announced nine out of 12 because I think the ninth problem came two hours after my tweet, which is a little bit awkward. So you're like, oh wait a second, everyone don't repost that tweet, repost this one.

Host

那不错。本播客的独家新闻。

That's nice. Breaking news here on this podcast.

Carina Hong

是的。12 题中答对 9 题是去年所有 4000 名人类参赛者中的第一名,这是普特南研究员,每年几乎稳定在前五名的最高荣誉,除了某些年份题目特别简单。今年的结果我们觉得非常鼓舞人心,尤其是最终证明发布时会确切说明解出了多少题。所以我认为这非常令人兴奋。我们见证了确定性工具和概率系统协同工作的力量,我因此热爱这种能力。

Yeah. Nine out of 12 was last year's number one among all 4,000 human participants, and it's Putnam Fellow, which is top honor, top five every year almost every year consistently, beyond some years where the problem somehow got really easy. This year's result we find very encouraging, especially given that the final proof release will say exactly how many solved. So I think that's very exciting. We witness the power of deterministic tooling and probabilistic system working together, and the capability I love because of that.

Host

好的。抱歉回到你的问题。也许我们先说一下,对于普特南考试,对于不是数学极客的人来说,这是一场数学本科生之间的竞赛,难度极高。中位数分数是零,而且参赛的是最优秀的数学本科生。所以答对任何一道题都是真正的成就,击败所有人则相当了不起。所有最优秀的数学家都在这个考试中表现出色。告诉我这个描述是否准确。

Okay. So sorry back to your question. Maybe we just say first of all, so for Putnam, for people who aren't math nerds, it's a contest among math undergraduates that's incredibly hard. The median score is zero, and it's the best math undergrads doing it. So getting any of the problems right is a real accomplishment, and beating everyone is kind of amazing. All the best mathematicians did super well in this. So tell me if that's a good description.

Carina Hong

是的,没错。我个人得分比我们的证明器系统差得多。所以这非常有趣。

Yeah, that's right. I personally got a much worse score than our prover system. So it's a very interesting thing.

Host

你得了多少?等等,让我们炫耀一下。你得了多少?

What did you get? Wait, let's brag a little bit. What did you get?

Carina Hong

嗯,我得了 12 题中的 4 题。你知道,这不算太好。我只进了前 250 名荣誉提名。所以不一定,但 Axiom 的最终分数可能会是我的三倍。所以我觉得我们应该办个“打败 Carina”派对,那总是很有趣。

Yeah. Well, so I got four out of 12. Which, you know, it's not that great. I got some top 250 honorable mention only. So not necessarily, but then that's like Axiom's final score will potentially triple me. So I think we should have a beat Carina kind of party, which is always fun.

2. 从学术之路到创业创始人 From academic path to startup founder

Host

你原本走在学术道路上,现在却经营着这家公司。是什么让你决定这样转变方向?

You were on an academic path and now you're running this company. What made you decide to shift gears like this?

Carina Hong

我五岁的时候就会听地下摇滚乐。它们显然非常政治叛逆,也有反叛的吸引力。我确实认为这家初创公司非常有趣,因为它是一次全新的冒险,你每天都有点像在过摇滚生活。

When I was five, I would listen to underground rock and roll music. They're obviously very politically rebellious and also have the contrarian appeal. I do think that this startup is really fun as in it's such a brand new adventure and you're a bit living the rock and roll day every day.

3. 人工智能与数学的未来 Future of mathematics with AI

Host

你对数学的未来有什么真正的预测?

What do you really predict for mathematics going forward?

Carina Hong

数学家将学会在不同于以往习惯的抽象层面上工作。所以如果你是一位伟大的数学家,就让你的直觉带你去任何想去的地方,让 AI 数学家成为勤奋的研究生,努力证明你的直觉。

Mathematicians will learn to work on a different abstraction than they are used to before. And so if you're a great mathematician, just let your intuition take you to wherever you want and have the AI mathematician be the diligent grad student trying to prove the intuitions you have.

4. 语言模型与人类的难度层级 Difficulty hierarchy for LMs vs humans

Host

你知道,我注意到直接使用 GPT 或 Gemini 这类语言模型(没有用你们的技术),它们在数学奥赛题上表现惊人地好——可能不是白金级题目,但分数已经很亮眼了。但我们有个共同的朋友 Robert Nishihara,他给了我一些脑筋急转弯题。其中一道很简单有趣的题,我放进 GPT 和 Gemini 里,它们居然解不出来。有意思。对我来说,数学奥赛比谷歌面试题难多了,但似乎语言模型没有同样的难度层级。你觉得什么对它们容易、什么难?你们的系统也一样吗?会不会有些对初级数学家来说很 trivial 的东西,对你们的系统反而很难?

You know, I've noticed using LMs like GPT or Gemini directly without your technology, they seem surprisingly good at math Olympiad questions, maybe not the platinum, but the scores are impressive. But then I have a mutual friend, Robert Nishihara, who gave me brain teaser questions. One simple fun question I put into GPT and Gemini actually couldn't solve it. Interesting. To me, math Olympiad is so much harder than Google interview questions, but it seems for LMs they don't have the same hierarchy of difficulty. Do you have an intuition for what's easy and what's hard for LMs, and is your system similar? Would you find some things trivial for a junior mathematician hard for your system?

Carina Hong

我觉得思考语言模型做数学这件事很有意思,尤其是在安全关键领域,比如代码——统计方法无法保证对边界情况有可证明的保证。这就像数学里,证明必须可靠,不能有致命缺陷。所以我们不相信把非正式模型扩展到数学 AI 的做法。对于某些数学问题,很容易论证那些需要精细检查的题目语言模型会很擅长;而依赖高层直觉、没有太多细节的题目,非正式模型可能贡献不错。对于计算,像数学插件这类工具很有帮助。比如分析学里一个简单的正性论证,就很有趣。我们计划发布一个改进版并附上评论,看看每个问题 AI 是怎么解的,以及是否符合人类思考的预期。在分析学里,AI 会做大量工作去处理人类学生或研究者直接忽略的东西。它会写大段代码来严格确保收敛和极限被仔细处理。这就是为什么早期 Mathlib 里,本科代数教材很快就被形式化了,但分析学花了很长时间。分析和代数通常是本科的两门课。我其实找到了一张图想给你看。

I think it's fascinating to think about LMs doing math and sometimes in safety-critical domains like code, where statistically just doesn't work in a way that you want provable guarantees on edge cases. This is similar to saying in math, a proof must be sound and have no critical flaw. That's why we don't believe in informal scaling of models to math AI. For certain math questions, it's easy to make a case that those requiring fine-grained checking will be very good for LMs, while those relying on high-level intuitions without detailed parts, the informal model will probably contribute well. For computations, certain tools like math plugins are helpful. It's interesting to see, for example, a simple positivity argument in analysis. We're planning an improved release with commentary, looking at each problem and how the AI solves it, and whether that's expected given how a human might think. In analysis, the AI does a lot of work for something a human student or researcher would just write off. It writes chunks of code to rigorously ensure convergence and limits are carefully handled. That's why in the early days of Mathlib, undergraduate algebra texts were formalized quickly, but analysis took a long time. Analysis and algebra are usually the two subjects at the undergraduate level. I actually found a graph I'd like to show.

Host

好的。

Okay.

Carina Hong

所以基本上,这不是一道普特南问题。这是 Erdős 问题 481。一个 Twitter 用户,一个本科数学系学生,自己解出了这道题,非常兴奋,就发帖了。我们想看看我们的系统能不能也解出来,作为 Twitter 日常互动。这就是系统做的事:它把问题分解成不同的子目标、结果或引理。从某个节点不是绿色开始,直到最后所有节点都变绿,这个节点才会变绿。

So basically, this is not a Putnam problem. This is Erdős problem 481. A Twitter user, an undergraduate math student, solved this problem by himself, got very excited, and posted about it. We're trying to see if our system can solve it too, as a Twitter daily engagement thing. This is what the system does: it breaks down into different subgoals, results, or lemmas. It starts from this node not being green, and then until everything is green at the very end, every other node turns green, this node will turn green.

Host

你能把它变成人类能理解的东西吗?

Can you turn that into something interpretable by a human?

Carina Hong

是的,我觉得可以。实际上,每个节点你都可以点进去,看到 Lean 代码。因为自动非形式化——把 Lean 转回自然语言——比反过来(自动形式化)要容易得多。模型见过大量英语,但没怎么见过 Lean,所以就像翻译:从英语到一种非常小众的语言,一个方向要容易得多。所以我认为可以,你可以点进去,这对想和系统交互、看它做证明的人来说会非常激动人心。

Yes, I think you can. In fact, the goal is for each node you can click into it and see the Lean code. Because auto-informalization, converting Lean back into natural language, is significantly easier than the other way, auto-formalization. Models have seen a lot of English, not a lot of Lean, so it's like translation: from English to a very niche language, one direction is significantly easier. So I think yes, you can poke into it, and that will be very exciting for someone to interact with the system when it's doing the proving work.

Host

酷。我想问:我注意到代码生成的一个现象——现在的语言模型更擅长生成能运行的代码,而不是人类容易理解的代码。这跟它们用的强化学习有关。我觉得你们可能也有类似问题:生成能证明为真的 Lean 代码,比生成人类能理解的代码更容易检查。我猜你们的模型生成的方式和人类不同。有些技术性数学证明已经很难理解了。你觉得我们会不会进入这样一个世界:模型证明了黎曼猜想,但没人能理解?每个人都同意每一步,但没人能真正获得直觉,知道发生了什么。

Cool. I wanted to ask: one thing I noticed with code generation is that LMs today seem better at generating code that works than code that is easily interpretable by a human. That makes sense with the reinforcement learning they use. I would think you might have a similar issue: generating Lean that proves true is probably easier to check than whether a human can understand it. I would think your models generate things in different ways than humans. I feel like some technical math proofs are already really hard to process. Do you think we're entering a world where models might prove the Riemann Hypothesis, but no one understands it? Everyone might agree with every step, but no one can actually get an intuition for what happened.

Carina Hong

对,对。所以我觉得这是个很有意思的问题。如果一个人证明了一个重大结果,我记得当年张益唐教授宣布他证明了一个数学结果时,他们很快召集了一个专家小组。我当时在牛津,有些教授告诉我们他们被召集去审查那个证明。因为有些数学家写得很晦涩,专家们也不能立刻明白是怎么回事。所以我想我会从这个角度开始。

Right, right. So I think it's a fascinating question. If a human proves a major result, I remember back when Professor Yitang Zhang announced he proved a math result, they quickly assembled an expert panel. I was at Oxford at the time, and some professors told us how they were gathered to examine the proof. Because certain mathematicians write in a very obscure way, the experts couldn't immediately know what was going on either. So I think I would start with that.

5. AI证明的可验证性 Verifiability of AI proofs

Carina Hong

我认为对于 Lean 证明这种具有可验证性特征的东西来说,这确实是个好消息。因为想象一下,假如某天 GPT 生成了一个一万行的 Lean 证明,声称它解决了一个开放猜想,尤其是它经过了思维链等数据的训练,我们根本无法确定真假。我记得 2024 年 12 月发生过这样的事,是关于 2025 年普特南数学竞赛的。Dan Hendrycks 在推特上发帖说,他把这些问题喂给了——我记得是 o1 Pro 还是 o3 Pro,我忘了——还有 o3 preview。然后人们给这些模型的输出打分,分数差异很大。推特评论区里没有人能确定那个解答是否正确。大家的共识是:不,它不对,因为所有这些问题都存在一个明显的漏洞,模型试图含糊其辞。而形式化系统的力量就在于你不能含糊其辞。事实上,它有点走向另一个极端——它几乎让人恼火地严格检查每一个细枝末节,比如收敛性和极限。我会想,你真的不需要花那么多篇幅去证明一个简单的正性引理,但它就是会这么做,而且这实际上成了我们生成的证明图中最重的节点之一。所以我觉得这很有趣。至少在优雅性和直觉方面,你能获得某种安心感。

I do think that for lean proofs which have this sort of verifiability feature, it's really good news. Because if you imagine, suppose GPT one day produces a 10,000-line Lean proof claiming that it has solved an open conjecture, especially because it's trained on chain-of-thought data among other things, we really can't be sure. So I think the scenario happened in 2024, December, about the 2025 Putnam. I remember Dan Hendrycks posted on Twitter about, hey, I feed these problems into I think it was o1 Pro or o3 Pro, I forget. And then o3 preview, and then people were grading those model outputs widely different grades. No one in the Twitter commentary area was able to say for sure whether that's a correct solution or not. The consensus was something like no, it's not correct, because for all these problems there is a significant hole that the model tries to handwave. The power of formal systems is you cannot handwave. In fact, it's doing a little bit too much of the opposite, which is that it's almost annoying that it's rigorously checking down to every fine-grain detail, like the convergence and limit. I'm like, you really don't need to spend that amount of space to prove the simple positivity lemma, but it does, and it's actually one of the heaviest nodes in the proof graph that we produce. So I think that's interesting. You have that sort of at least peace of mind when it comes to elegance and intuition.

6. AI作为数学家的协作者 AI as a mathematician collaborator

Carina Hong

对于很多问题来说,要给出两个证明是非常困难的。所以如果有人告诉我,嘿,我对黎曼猜想有两个证明,我觉得这比说只有一个证明要不可信得多。有些问题你本能地觉得只可能存在一个证明。所以,AI 找到的证明,去掉那些冗余的部分——比如反复重复相同的内容,你可以说它很丑陋——但如果你把这些去掉,精简后的版本很可能与人类可能找到的证明趋同。我认为这非常令人兴奋。所以,不要把 AI 当作二等公民,而是把它看作另一个数学家合作者,一个你可能从未听说过的合作者。这有点像拉马努金的故事:剑桥的哈代和李特尔伍德发现了这位来自印度的自学天才。我认为这是一种看待 AI 数学家的非常有趣的思维方式。

For a lot of the problems, it is very hard to produce two proofs. So if someone tells me that hey, I have two proofs for the Riemann hypothesis, I think that's a much more unlikely statement than say I have one proof. There are certain problems that you just feel like there is bound to be one proof. So that would probably mean that what the AI finds, minus the part where it might be doing redundant stuff like repeating the same thing again and again, you can call that ugly. But if you take those out, the streamlined version of it could, there's a good case to be made that it might be convergent to what a human would possibly find. And I think that's extremely exciting. So kind of treating AI not as some sort of second-class citizen, but as just another mathematician collaborator, maybe one that you have not heard of. It's similar to the Ramanujan case where Hardy and Littlewood at Cambridge just discovered this self-taught genius from India. I do think that's a very interesting mindset to think about AI mathematician.

7. 展示证明树 Demonstrating the proof tree

Carina Hong

说到这个,我还发现了另一种证明树。抱歉,希望这部分能剪掉。好,我又在共享主持人的屏幕了。这是系统在证明问题 A1 的过程。还有一些问题,这个显然很复杂。它证明得这么快。我们做了时间加速。我觉得如果我不刷新,它就会……嗯。这就是它在做的事情。

Speaking of that, I also found the other kind of proof tree. I'm sorry, I hope this can be cut out. Okay, so I'm sharing the host screen again. This is how the system is proving problem A1. And then there are problems, this is like a really complicated one obviously. It's proving it this fast. We kind of have like a time acceleration. And I think this is like, if I don't refresh it, it kind of... Yeah. So this is what it's doing.

8. 理解问题A3 Understanding problem A3

Carina Hong

A3 是一个大家都能理解的问题。这是一个博弈问题,标准的 Alice 和 Bob 玩一个游戏,涉及一个由 n 个数字组成的字符串,每个数字只能是 0、1 或 2。一开始所有数字都是 0,这是初始状态。然后每一步,你可以对其中一个数字加 1 或减 1,生成一个之前没有出现过的新字符串。如果你做不到,也就是说所有可能的字符串都用完了,那么对方获胜,Alice 总是先手。所以问题是哪个玩家有必胜策略?这是一个竞赛数学中标准的博弈论问题。很多学生看到它时,可能会用两三个竞赛数学中非常常用的结果来攻克这类问题。例如,在一个没有平局的游戏中,一定存在必胜策略。然后问题只是问哪个玩家。学生会尝试写下字符串,想象自己是 Alice 或 Bob,看看他们会如何进攻,并试图找到任何模式。这其实就是直觉发现的部分。如果我是 Alice 或 Bob,我会怎么做?真正解题的人会试图找出是否存在某些状态比其他状态糟糕得多。所以如果我面对一个字符串,我可能会想,如果某个数字是 0,那么我不能减,只能加。如果是 2,我不能加。所以数字为 1 的位置会给我更大的灵活性。然后尝试将这些很糟糕的状态或很好的状态泛化,从而确定序列的走向,或者当我将对手置于那种情况时,对手的合理行动是什么。我认为这些问题非常有趣。当然也有更繁琐的问题,比如 A2 和 B2,涉及很多微积分。因为普特南竞赛是最难的本科水平考试,不像 IMO 是高中水平,微积分是允许的。对于非常擅长积分的人来说,A2 或 B2 不会构成太大挑战。然而,对于 AI 证明器来说,我觉得它根本不觉得有挑战性,但它确实需要很多行 Lean 代码来解决一些被视为已知的问题。在某些情况下,AI 会证明出非常不同的东西,比如与人类数学家专家预期的解法截然不同,这是最有趣的案例。

A3 is a problem where people can definitely understand. It's a game problem where Alice and Bob, the standard Alice and Bob, play a game for a string of n numbers, and it can only be zero, one, or two. In the beginning, everything is zero, that's the start state. Then you can, at each move, either add or subtract one from one digit to create a new string that has not appeared before. If you cannot do that, as in you have run out of all the possible strings, then the other player wins, and Alice always goes first. So which player has a winning strategy? That's a standard sort of game theory problem in competition math. A lot of students, when they look at it, will probably have about two or three very widely used results in competition math to attack this kind of problem. For example, you're guaranteed, I think in a game of no ties, a winning strategy. And then it's just asking which player. Students will try to write down the string and imagine they are Alice or Bob to try to see how they would attack the game and try to see if there's any pattern they can find. This is really the part about intuition discovery. If I'm Alice or Bob, what would I do? People actually solving the problem will try to figure out, hey, are there certain states that are just a lot worse than others? So if I'm faced with a string, I will probably think that if I have zero on one number, then I cannot subtract, right? I can only add. And if I have two, I cannot add. So potentially one at a place will give me a lot more flexibility, and try to generalize these sort of really bad states or some really good states into determining the sequence of things or what the rational move for the opponent is when I put them in that situation. I think these are very interesting, very fun problems. There are obviously more tedious problems, like if you look at A2 and B2, there's a lot of calculus involved. Because the Putnam is the hardest undergraduate-level exam, unlike the IMO being high school level, calculus is seen. For people who are very proficient with integration, I don't think A2 or B2 will pose any significant challenge. However, for the AI prover, I don't think it's challenging at all, but it just does take a lot of lines of Lean code to solve some of the taken-as-given issues. For some cases, we have the AI prove very different things, like have a very different solution from what our human mathematician experts expect, which is the most fun to watch cases.

9. 主持人尝试解答A3 Host's attempt at A3

Host

那么我们来谈谈 A3 吧,这个问题很棒,对吧?很容易理解。我的思路是,好,让我试一下。看起来你只能加 1。所以 Alice 加 1,得到 1。显然 Bob 不能减 1,所以 Bob 加 1。好。那么玩家二赢了,对吧?是这样吗?实际上,首先,抱歉。你说玩家二可以加 1 得到 1,然后 Bob 加 1 得到 2。嗯。

Well, so let's talk about A3 because that's a great one, right? It's so easy to understand. Where my brain goes is like, okay, let me try one. It looks like all you can do is add one. So Alice adds one. Now you get one. Obviously Bob can't subtract one. So Bob adds one. Okay. So player two wins, I guess, right? Is that accurate? Actually, first of all, sorry. So you say player two person can add one to get one and then Bob goes and adds one and gets two. Yeah.

10. A3问题中Bob的获胜策略 Bob's winning strategy in A3 problem

Host

然后我觉得……我觉得没错。所以实际上 Bob 总是赢。

And then I think I think that's right. So actually Bob always wins.

Carina Hong

是的。Bob 有一个非常简洁的必胜策略,可以很精炼地描述出来。一旦你发现了它,Bob 根本不需要太在意 Alice 的走法,我觉得这恰恰让这道题比上一道 Pund 级别的题要简单一些。所以它只是 A3 难度。每场 session 有六道题,通常四、五、六会比一、二、三更难。基本上 Bob 只要执行那个策略,不用多想就行。这也让它在 Lean 中更容易追踪,因为需要跟踪的状态很少,也没有复杂的分支需要推理。

Yeah. There's a very clean winning strategy for Bob that can be described succinctly. So once you see it, Bob doesn't need to respond with much care to Alice's moves at all, which I think is what makes this problem slightly easier than a last pund problem level. So it's only A3. There are six problems for each session. Usually we'll expect four, five, six to be harder than one, two, three. And yeah, basically Bob can just execute that strategy without thinking much. So that makes it also much more trackable for a lean as well. There's very little state to track and no complicated branching to reason about.

Host

但我想,你知道,当我内省自己的思维时,对吧?我认为人类的做法是思考各种情况,寻找一个合理的策略,而不是形式化的数学证明。那么你觉得 Axiom 实际上采用了与人类不同的方法吗?

But I guess like I don't you know like when I introspect my own brain here, right? But I think the human approach here is to think about cases, look for a strategy that makes sense, not a formal mathematical proof. So do you think that Axiom is actually taking a different approach than a human approach?

Carina Hong

你是指这个具体的 A3 问题吗?

You mean for this specific A3 problem?

Host

对,就这个具体的 A3 问题。

For this specific A3 problem.

Carina Hong

是的,对于这个具体的 A3 问题,我认为 AI 的思考方式和人类是相当一致的。但在其他场景中,我们可以看到 AI 采取了与人类不同的路径,因为它知道自己可以暴力求解,而人类——如果你是一个真正的解题者,想要避免暴力求解——如果你知道暴力求解可能行不通,比如我在做一个欧几里得几何问题,完全不知道如何下手,但我知道暴力求解(复坐标法)肯定能行,我可能愿意花一个小时去暴力求解,但对于某些非几何问题就不一定了。我认为人类甚至会找到一个巧妙的一图证明。

Yeah, for this specific A3 problem, I would say that it's kind of aligned how the AI thinks about this problem and how a human thinks about it. There are other scenarios where we can see the AI taking a different path from humans, because it knows it can brute force through things where a human, if you are an actual task taker and you want to avoid brute force, if you know brute force might not work—like if I'm doing a Euclidean geometry problem and I have no idea how to solve it, and I know that brute force is the complex coordinate method that can definitely get me there, I am probably willing to sink one hour into brute forcing that, but not necessarily for some of the non-geometry cases. I think the human will find a clever one-diagram proof even.

11. 形式化验证的应用与里程碑 Applications and milestones of formal verification

Host

这是个很好的过渡。我想问问你这些应用的前景。沿途有哪些里程碑?我的前 CRO 是你们董事会成员,也是你们的投资者之一。我很好奇你是怎么向他们推销这个数学语言模型的。

So it's a good segue. I want to ask you about the applications of this. What are the milestones along the way? My former CRO is on your board and one of your investors. I'm curious how you pitched this math LM to them.

Carina Hong

是的。我觉得投资者自己也看到了这一点。我发现一件非常令人惊叹的事:当我和投资者交谈时,与某些基金存在真正的智力共鸣。我认为基金内部甚至有备忘录,标明这个类别甚至不是一个值得竞争的类别,而是他们认为未来不可避免的东西。所以如果你考虑硬件和软件中验证的滞后——设计团队的规模只有验证团队的三分之一或四分之一,验证在硬件场景中可能需要三年。在软件方面,例如,在我们拥有大量原生程序员使用各种编码模型(如 Cursor)之前,代码审查和安全关键场景的需求很大。我确实认为客户和企业正面临一个不幸的选择:要么“我只希望这东西不出错,所以可能让人类来检查或手动完成,而不是依赖 AI”,但这也非常痛苦,原因我们之前提到过:时间、资源、无法接触到某些专家。所以我认为,芯片领域高水平的正式验证专家与中等水平专家之间的差距——如果你能设法让 AI 处理那些非常棘手的案例,从而降低招聘门槛,那将非常有价值。例如,我看到 AWS 拥有最好的自动推理团队之一,但他们花了五年时间手动形式化 hypervisor 的内存隔离组件。这个数字令人震惊。这还不是整个 hypervisor,只是其中一部分。我认为随着更多 VIP 编码和更多智能体生成代码,验证边界情况是否覆盖也将非常重要。另一个用例是代码迁移。如果你有遗留代码,你想确保新代码与遗留代码完全等价——既不好也不差——因为这些遗留代码服务于重要的业务功能。这在数据库场景中也是一个非常有趣的用例。如果有恶意行为者,那么证明它实际上是一致的。目前数据库之间是分片的,在这个意义上它们是不一致的,因为存在恶意行为者的问题。所有这些都可以通过形式化验证来解决,提供一个非常有趣的解决方案。

Yeah. I think people saw it themselves as well. This is something I find really amazing: when I'm talking to investors, there are real intellectual alignments with certain funds. I think funds even have internal memos marking that this category is not even a category where it's good to battle; it's something they consider inevitable for the future. So if you think about the lag of verification in hardware and software, where design teams are one-third or one-fourth the size of verification teams, verification could take as I said three years in the hardware scenario. If you think about in software, for example, before we have a lot of the programmers who are native programmers using various coding models like Cursor, but there's a lot of need for code review and safety-critical cases. I do think that customers and enterprises are facing this unfortunate choice between 'I just want this to not be wrong, so I might have a human to eyeball or manually do it instead of relying on AI to do it,' but that's also very painful for the reasons we described: time, resources, the inability to access certain experts. So I think the difference between a high-expertise formal verification expert in chips and a medium-expertise one—if you can somehow make it the case that the AI can handle the very notorious cases and lower the bar for hiring, I think that would be incredibly valuable. For example, I see that AWS, which has one of the best automated reasoning teams, took them I think five years to formalize the memory isolation component of the hypervisor by hand. That's a shocking number. It's not the entire hypervisor, it's one part. I think also with now more VIP coding and more agents generating code, verifying that the edge cases are covered is going to be very important. Another use case is code migration. If you have legacy code and you want to make sure that the new code is completely equivalent to your legacy code—no better, no worse—because those legacy codes are serving important business functions. That's another very interesting use case in the database scenario. If there are bad actors, then proving that it is actually consistent. Currently databases are kind of sharded from each other and they are inconsistent in that sense because they have this problem where there might be bad actors. All these can be solved by formal verification, providing a very interesting solution.

Host

是的,我想在我更熟悉的软件领域,我觉得测试用例并不像形式化证明那样非黑即白,而且可能有点过度约束。我认为真正的挑战在于将“错误”的含义形式化到足够清晰,以至于你实际上可以……

Yeah, I mean I guess the domains that I'm more familiar with in software, I feel like the test cases are not quite as black and white as with formal proof, and it might be kind of over-constrained. I think the real challenge would be formalizing what it means to be wrong clearly enough that you could actually...

Carina Hong

是的,我觉得这非常有趣。特别是对于数据库场景,我认为拥有某种易于交互的抽象将是一个重大突破。所以我认为目前的情况是:有些用例存在迫切需求——我们真的需要验证一切,也许是因为法规要求,也许是因为我无法将其投入部署。所以目前我们让人类来做,以确保在最大程度上一切万无一失。然后还有一些“锦上添花”的用例,比如数据库场景,有它当然很好。我认为如果没有它,我们也许也能凑合,但显然拥有它需要权衡:使用起来有多痛苦,以及使用它能解锁多少能力。

Yeah, I do think it's very interesting. For the database case specifically, I think having some sort of abstraction that is easy to interface with will be a big unlock. So basically what I think is going on is that there are certain use cases where there is a dire need—we really need to verify everything, maybe because regulation mandates it, maybe because I just can't put that into deployment. So currently we are having humans do it to make sure that to the best extent possible everything is soundproof. And then there are 'good to have' cases, like in the database case, where it's really good to have. And I think if we don't have it, maybe we can live without it, but obviously having it is a weighing of how painful it is to use versus how much capability it will unlock when I use it.

12. AI验证与编码的三个层级 Three tiers of AI verification and coding

Carina Hong

我认为这些都是有趣的权衡。还有最后一类:比如编写一个可爱的网站可能不需要形式化验证。我不认为我们正在构建的 AI 数学家能写诗。这大概是一个三层结构。我们的目标显然是从第一层开始,然后最终拿下第二层。我认为当我们拿下第二层时,那将是 AI 更广阔领域中的一个重大时刻,远不止数学。例如,在第一种情况下,如果有一个由 Leslie Lamport 编写的 Paxos 协议一致性证明,你可以自动形式化该证明以及部分形式化过程。系统可以生成像 Rust 这样的强类型语言代码来执行这个协议。这种将实现与验证统一在生成中的想法非常强大。然后你可能会问:我能将这种方法从 Rust 扩展到一些弱类型语言到什么程度,慢慢缩小差距,同时知道不可能形式化验证 Python 中的所有代码,但可以验证大部分?人们会选择哪种编程语言本身也是一个不确定的问题。谈到验证用例,在 1980 年代,很多公司尝试用人类进行形式化验证,这些人非常擅长编写改进语言,比如 Lean 的古老对应物。当时它没有规模化,因为人类是一个瓶颈。也许 AI 形式化验证可以在 Axiom 这里卷土重来,从 AI 数学家开始。

I think all these are interesting trade-offs. Then there's this final group: maybe coding a lovable website wouldn't necessarily need formal verification. I don't think the AI mathematician we are building can write poetry. There is this sort of three-tier thing. Our goal is obviously to start from the first tier and then eventually take the second. I think by the time we take the second group, it will be a big moment in the broader landscape of AI, much beyond math. For example, in the first case, if there's a proof of the consistency of the Paxos protocol written by Leslie Lamport, you can auto-formalize both the proof and part of the formalization. The system can generate code in a strongly typed language like Rust to execute this protocol. This unifying of generation—implementation and verification in one—is a powerful idea. Then you might ask: how much can I extend that from Rust to some weakly typed languages, slowly closing the gap, knowing it's not possible to formally verify all code in Python, but you can do a majority of it? Which programming language will be people's choice is also an uncertain question. Talking about the verification use cases, in the 1980s, a lot of companies tried formal verification with humans who were very good at writing in improving languages, like the ancient counterpart of Lean. Back then, it didn't scale because humans are a bottleneck. Perhaps AI-form of verification can have a huge comeback here at Axiom, starting with an AI mathematician.

Host

是啊,挺有意思的。这在某种程度上感觉非常老派。我记得在 90 年代,人们对 AI 中的形式化证明有很多热情,但过去几十年很少听到。这有点意思。我想象从 80 年代和 90 年代中汲取东西,比如搜索定理并尝试构建证明。我认为也有很多现代研究。具体来说,推理模型、代码生成能力、在可验证领域应用强化学习的巨大改进,以及 Lean 成为广受赞誉的语言,有比以往更多的开发者对它感到兴奋,数学家也愿意接受它。我认为这三件事结合在一起,使得现在成为软件和硬件验证领域的一个绝佳时机。例如,在软件方面,有正在开发的 CSlib。斯坦福大学自动推理实验室由 Clark Barrett 教授领导,在 DeepMind 和 AWS 等行业合作伙伴的支持下进行。他们试图用 Lean 形式化所有本科 CS 文献。有人试图用 Lean 形式化机器学习的东西。我们正生活在一个非常有趣的时刻。回到数据库,我很好奇你的想法。例如,所有数据库都必须解决拜占庭将军问题。这些解决方案的一致性需要被形式化证明。这类事情是值得关注的,还是人们可以接受不这样做?我自己其实很好奇。

Yeah, it's funny. It feels very old-fashioned in a way. I remember in the 90s there was a lot of enthusiasm around formal proving in AI, and you don't hear about it much in the last few decades. It's kind of interesting. I imagine pulling things from the 80s and 90s in terms of searching over theorems and trying to build proofs. I think there is a lot more modern research as well. Specifically, the timing of reasoning models, code generation ability, great improvement applying reinforcement learning in a verifiable domain, and Lean becoming a widely celebrated language with more developers than ever excited about it, and mathematicians willing to embrace it. I think all these three things coming together make now a really good timing in the software and hardware verification landscape. In software, for example, there is CSlib under development. Stanford's Automated Reasoning Lab led by Professor Clark Barrett is doing it, supported by industry partners like DeepMind and AWS. They're trying to formalize all the CS literature undergrad in Lean. There are people trying to formalize machine learning things in Lean. It's a very interesting moment we are living in. Going back to the database, I'm curious about your thought. For example, all databases have to solve the Byzantine general problem. Consistency of those solutions needs to be formally proven. Is something like that of interest, or can people live with that? I'm actually quite curious about that myself.

Carina Hong

嗯,我一直很好奇。我也不是数据库专家,但这确实让我思考。曾几何时,数据库的主要应用更像是金融服务,人们非常关注确保事情完全一致,有真正的不变量。然后我一生中看到的是对 NoSQL 和放松一致性约束的数据库的兴趣激增,因为它们希望更高效地运行,保存所有用户数据并能够查询。还有其他这些相互竞争的东西,似乎不同。似乎大多数数据库应用可能不需要那种形式化一致性水平。我想知道那种……

Well, I always wonder. I'm also not a database expert, but it does make me think. There was a time when the big application of databases was more like financial services, and people really focused on making sure things were completely consistent, with real invariants. Then I think what I've seen in my life is this explosion of interest in NoSQL and databases that relax these consistency constraints because they wanted to operate more efficiently, save all user data, and be able to query it. There are these other competing things that seem different. It seems like the majority of applications for databases maybe didn't need the level of formal consistency. I wonder how that sort of...

Host

我觉得总会有一些应用需要你极其小心,但总是有权衡。在某些应用中,你……

I feel there will always be applications where you want to be incredibly careful, but there always are trade-offs. In certain applications you...

Carina Hong

我同意。对此有趣的反应是,你不必因为系统中的强推理组件而放弃效率。所以它会非常不同。再次强调,推理改进框架。例如,较早的 AlphaProof 系统,顺便说一句,它赢得了 IMO——老实说,那对我来说就是 IMO 时刻。28 分,银牌,金牌线是 29 分,但他们答对了除组合题外的所有题目。这和 2025 年 IMO 一样,很多公司得了 35 分,只错了一道组合题。两年的区别是 2024 年有两道组合题,2025 年有一道。在 AlphaProof 系统中,它是一个形式化系统。我们可能会看到这种权衡:如果你只有一个非形式化系统,它无法进行形式化证明。我认为将这两者结合起来将是革命性的,即使从客户的角度来看也是如此,因为有了自然语言部分。作为客户,你实际上不需要费力地精确说明你想要什么。我认为自动形式化能力的目标之一是,当面对歧义时,系统会怎么做?如果你只是……我真的很喜欢 Deep Research。我经常使用它。

I agree. I think the interesting response to that is that you don't have to give up efficiency because of the strong reasoning component in the system. So it will be very different. Again, the reasoning improver framework. For example, the older AlphaProof system, which by the way won the IMO—that to me was the IMO moment, to be honest with you. The 28 score, silver medal, with 29 being the gold cutoff, but they got every single question about the combinatorics one. That's the same as the 2025 IMO where a lot of the companies got 35, missing the one combinatorics. The difference between the two years is 2024 there are two combinatorics and 2025 there's one. In the AlphaProof system, it's a formal system. We might see that trade-off: if you have only an informal system, it cannot formally prove. I think combining these two is going to be game-changing, even from a customer perspective, because of the natural language part. You don't actually need to painstakingly spell out exactly what you want as a customer. I think part of the goals of the capability of auto-formalization is that when facing ambiguity, what does the system do? It will be not very fun if you just... I really like Deep Research. I use that a lot.

13. 交互反馈与自动形式化 Interactive Feedback and Autoformalization

Carina Hong

它会,或者说,有些智能体比如 Menace 会问我偏好这个还是那个。这是主动的,挺好的。我认为那是交互式反馈。但一旦列表到了 500 项,我显然就会失去耐心。面对模糊性时知道该怎么做,是自动形式化中非常重要的一部分。再说一次,自动形式化,大多数人认为它是一种数据生成方式,但对我们来说,这是一项核心技术。我们认为自动形式化可能比证明类似问题更难。这是一项核心技术,两者几乎不可分割。对我们来说,自动形式化的难点始终在于陈述本身,因为如果你想自动形式化一个证明,你有 Lean 检查器给出的信号,那没问题。但正确形式化陈述会很有挑战性,因为根据定义,你还没有要证明的解。在代码验证场景中尝试自动形式化非常有用,因为你有输入输出对的测试用例作为形式化规范的代理。我认为这种基于规范的编程,类似于 AWS 的 Carol 这样的产品,非常有价值且具有远见。

It will, or like, even some agents like Menace will ask me for my preference for this or that. That's active, that's good. I think that's interactive feedback. But once that list gets to like 500, then I will obviously lose patience. Knowing what to do when facing ambiguity is a very important part for autoformalization. And again, autoformalization, most people think about it as a data generation way, but for us it's a core technology. We think of autoformalization as harder potentially than proving a similar problem. It's a core technology and they can't be almost separated. Also for us, the difficulty of autoformalization is always the statement, because if you want to autoformalize a proof, you have the lean checker signal, that's fine. But formalizing the statement correctly is going to be challenging because by definition you don't have the solution to prove yet. It's very good in the code verification setting to try to autoformalize because you have the test cases of input-output pairs as a proxy to ground your formal spec. I think this sort of spec-based programming, similar to products like Carol from AWS, is just very valuable and visionary thinking.

Host

我有点力不从心,但这里是不是存在某种停机问题,限制了它可能的发展方向?仅仅因为有测试用例并不意味着你就能知道结果。

I mean, I'm a little bit out of my depth here, but isn't there a kind of halting problem that limits where this can possibly go? Just because you have cases doesn't mean you'll know.

Carina Hong

你不能像那样验证所有 Python 代码,那是不可能的。但你可以验证大部分代码。你可以用已有的能力做一些让人们生活更轻松的事情。我想我大约 10 分钟前说过,那是一个重要的限定条件。你不能处理所有代码,但可以处理相当多。

You cannot verify all code in Python like that, that's not possible. But you can do it for the majority of it. You can do things that can make people's life easier with the capability you have. I think I said before, about 10 minutes ago, that was an important qualification. You cannot do all code, but you can do quite a lot.

14. 商业化与研究阶段 Commercialization and Research Phase

Host

你们已经在进行技术的商业化了吗,还是仍处于研究阶段?

Are you already working on commercialization of your technology, or are you still in the research phase?

Carina Hong

实际上,我们现在是一家成立六个月的公司,假期过后。我们原以为研发阶段会花很长时间,但直到今天,我仍然认为我们只处在第零天。还有很多事情要做,但这并不意味着我们没有在与一些值得信赖的合作伙伴积极讨论这项技术能为他们解决什么问题。通过这些对话,首先,这非常拓展智力,而且感觉很像我的 C 轮融资过程——与你交谈的人有着真正的智力共鸣,每个人都觉得自己有点像秘密守护者。就好像我们看到了世界尚未看到的东西。当你和对的人交谈时,那种感觉令人振奋。我们在这方面正在进行一些早期的对话。

We are actually a six-month-old company now, after the holiday. We thought the R&D stage would take a long time, and I still today think we are only at day zero. There's so much to do, but that doesn't mean we aren't in active conversations with some trusted partners about what this technology could solve for them. Through these conversations, first of all, it's just very intellectually expanding, and it's really great to feel similar to my C round fundraising process—there's real intellectual alignment with the people you're talking to, and everyone feels a bit like a secret keeper. It's like we see something the world doesn't see yet. That feeling, when you are talking to the right people, is just exhilarating. We are doing some early conversations on that front.

15. 人工智能与数学的未来 Future of Mathematics with AI

Host

你对数学的未来有什么预测?当然,说这是一个有帮助的朋友和合作者很好,但它会变得越来越好。我想象在某个时候,所有新证明都会以这种方式生成。你认为学术数学还会继续吗?

What do you predict for mathematics going forward? Of course it's nice to say this is a helpful friend and collaborator, but it's going to get better and better. I imagine at some point all new proofs will be generated in this manner. Do you think academic mathematics continues?

Carina Hong

我认为数学会继续。我认为数学家将学会在他们不习惯的抽象层次上工作。抽象层次的改变遇到的阻力会比人们想象的小。举个例子:在 LaTeX 出现之前,你用打字机来排版数学论文。LaTeX 出现后,数学家们欢迎并使用了它。这可以是一个非常好的工具。此外,让计算机帮助数学研究的某些部分并不新鲜。计算机代数系统已经帮助了很多计算性工作。我们发现,这将解放那些已经培养出很好直觉的数学家,使他们免于辛苦检查每个技术引理的手工劳动。所以如果你是像陶哲轩这样伟大的数学家,就让你的直觉带你去任何地方,让 AI 数学家成为勤奋的研究生或博士后,努力证明你的直觉。一个非常重要的部分是,很多 AI 数学研究忽视了构造部分。证明和构造都是数学研究中非常重要的部分。构造有趣的例子作为猜想前的步骤——写下几个序列,看看它们是否给你灵感来制定下一个正确的引理,或者构造反例来说不要走那个方向。这将不依赖于我们讨论的形式化证明,而是依赖于数学发现,比如模式提升,能够生成有趣的模式。我们还没有触及那部分,我认为这将是一个非常有趣的循环:由数学家的直觉引导出一个猜想,然后行动证明者证明它,然后你以某种方式生成更多猜想。所以通过形式化猜想或通过端到端的翻译式代码库进行大规模发现,你发现更多猜想并证明它们,然后开始这个循环。数学家可以成为引导循环每一步的直觉的良好守护者。

I think mathematics will continue. I think mathematicians will learn to work on a different abstraction than they are used to. The change of abstraction is going to have less resistance than people think. Let's take an example: back then before LaTeX, you had a typewriter to typeset your math paper. After LaTeX came out, mathematicians welcomed it and used it. This can be a really good tool. Also, having computer help with some part of mathematics research is not new. Computer algebra systems have been aiding a lot of the more computational work. We find that this will free mathematicians who have already developed very good intuitions from the manual labor needed to painstakingly check every technical lemma. So if you're someone like Terence Tao, a really great mathematician, just let your intuition take you wherever you want and have the AI mathematician be the diligent grad student or postdoc trying to prove the intuitions you have. One very important part is that a lot of the AI for math efforts overlook the construction part. Proving and construction are both very important parts of mathematical research. Constructing interesting examples as a pre-conjecture step—write down a few sequences and see whether it gives you ideas for the right lemma to formulate next, or constructing counterexamples to say don't go that direction. That will rely not on the formal proving things we talked about but on mathematical discovery like pattern boost, being able to generate interesting patterns. We have that part we haven't quite touched on, and I think that will be a really interesting loop: a conjecture led by intuitions of a mathematician, then the action prover proves it, and then you generate more conjectures one way or another. So formal conjecturing or through the sort of end-to-end translative code base in mass discovery, you discover more conjectures and prove them, and then you start having this loop. Mathematicians can be good guardians of the intuitions guiding every step of the loop.

Host

我的意思是,不难想象你正在构建的系统可能比数学家更有直觉,知道该尝试什么,甚至知道什么对人类来说有趣,这有点有趣。

I mean, it doesn't take much imagination to also imagine that the system you're building could have better intuitions than mathematicians about what to try and even what's interesting to humans, which is kind of funny.

Carina Hong

不同的直觉。我认为有趣的是,如果我们认为机器觉得难和容易的事情与人类觉得难和容易的事情是一样的,那可能更符合“我们完蛋了”的叙事。

Different intuitions. I think the interesting thing is it will probably be more of a case where the 'we are doomed' narrative, if we think that what's difficult and easy for machines are the same as what's difficult and easy for humans.

16. AI与顶尖数学家对决 AI vs top mathematicians

Host

事实并非如此,对吧?我们认为推理者也会从人类那里学到很多。但我确实认为,有那前 0.00001% 的数学家,他们就像是智力巨人。我认为人工智能要赶上他们需要很长时间。

That's not the case, right? We think that the reasoner learns a lot from humans as well. But I do think that there is a sort of top 0.00001% of mathematicians, they are like intellectual powerhouses. I think for an AI to catch up to that takes a long time.

Carina Hong

有意思。因为你认为陶哲轩和普通学术数学家之间的差距,比数学本科毕业生和普通学术数学家之间的差距还要大。

Interesting. Because you think there's a bigger gap between a Terence Tao and an average academic mathematician than between someone with a math undergrad and the average academic mathematician.

Host

我也这么认为。但我也觉得,即使说到陶哲轩这样的数学家,也有多种风格。有些数学家是理论构建者,他们不是最擅长解题,但能通过连接理论文献中深层的点,开辟全新的领域。还有一类是解题型数学家。我认为,在不需要大量深层理论的领域——顺便说一句,这大大缩小了我们讨论的范围,因为即使在组合学中,也有很多问题目前与分析学、非常深层的分析学相关,所以这绝不是一个浅显的学科——确实有某些类型的数学家会比其他人更早将 AI 视为一种调研和等价函数工具。

I would think so. I would think so. But I also think that even when we talk about the Terence Taos, there are many styles of mathematicians. There are mathematicians who are theory builders. They are not the best problem solvers, but they are able to open up completely new fields by connecting the dots of things that are deep in the theory literature. And there are people who are problem solver type mathematicians. I do think that the problem solver type of mathematicians in domains that do not require a lot of deep theory — which by the way very much narrows the fields we're talking about, because even in combinatorics there are a lot of problems that are currently connected to analysis, very deep analysis, making it not a shallow subject at all — I do think there are certain types of mathematicians that will see AI as a survey and equivalent function more than the other types.

Carina Hong

有意思。你有没有试过把 P vs NP 或黎曼猜想这样的问题输入到你的系统中?

Interesting. Have you tried putting in something like P versus NP or the Riemann hypothesis into your system?

Host

哦,我觉得我们还没到那一步。看看结果如何。

Oh, I don't think we are there yet. See what comes out.

Carina Hong

我觉得我们还没到那一步,但我确实认为我们很快会发布一些有趣的结果,关于我们未曾预料到的惊人能力。老实说,我们也没想到 Putnam 竞赛的结果会那么好。不断看到技术能力压缩了甚至构建者自己的时间线预测,我认为这对我们这样早期阶段的初创公司来说是一个非常好的信号。

I don't think we are there yet, but I do think that there are interesting results we're going to put out soon about surprising capabilities that we didn't expect. Certainly, to be very frank with you, we did not expect Putnam to be that well — that the result to be that good as well. Constantly seeing that the capability of the technology compresses the timeline prediction of even the ones building it is a very good sign, I think, for a startup like us at this early stage.

17. 从学术界到创业的转型 Transition from academia to startup

Carina Hong

完全同意。我还有一个好奇的地方:你之前走的是学术道路,现在却在经营这家公司。能谈谈这个转变吗?是什么让你决定这样转换方向?

Totally. I guess one more thing I was curious about: you were on an academic path and now you're running this company. Could you talk a little bit about the transition? What made you decide to shift gears like this?

Host

我从小听硬核摇滚乐长大。我是说,我五岁时就会听那种有很多脏话的地下摇滚乐。我不知道为什么我妈妈没有做家长控制。你在说什么?

I grew up listening to hardcore rock and roll music. I'm talking about when I was five, I would listen to underground rock and roll music with a lot of curse words in it. I don't know why my mom did not do the parent control. What are you talking about?

Carina Hong

呃,就像中国有一些非常隐秘的乐队,摇滚乐队。它们显然在政治上很叛逆,也写各种话题。对于一个五岁的孩子来说,歌词相当有趣。我内心有一部分非常欣赏生活中发生的各种事情,也欣赏反主流的一面。所以我不认为自己在学术界时是最学术的,也绝对不是最创始人的创始人。我有点介于两者之间。我确实觉得这家初创公司非常有趣,因为它是一场全新的冒险,你每天都在过摇滚乐般的生活。熵很高,事情非常令人兴奋。你在赋能他人去做他们梦想的研究,去建立他们的科学遗产或实现他们的科学抱负。我也觉得确保事情运转、服务我们拥有的世界级研究人员是另一部分我非常享受的事情。我认为这个转变实际上并没有——我没有感觉到。我觉得我好像找到了我的故乡。我认为有一种非常特殊的雷电人格,以及所有那些关于“肩膀上的旅程”之类的说法,关于既非常自发又非常书呆子方法论的说法。我发现自己就是所有这些特质的集合。所以我真的很热爱我的工作。我以为会有某种冲击或转变。不,我感觉如鱼得水。真的很棒。我也觉得和这个团队一起工作——我不认为还有另一个团队同时拥有 AI、编程语言和数学这三大支柱,这非常有趣。编程语言更偏向系统侧,我们有很多人处于编程语言和 AI 的交叉点、AI 和数学的交叉点,以及所有这些应用领域(如芯片或代码验证)的交叉点。所以不知是运气还是什么,我们为这三大支柱招聘的人在他们各自的路径上也有这些应用领域的记录。所以感觉这是合适的团队,与他们共事是一生的荣幸。看到他们如何从各种角度思考问题——比如那些可能不懂机器学习的数学家给了我们很多灵感。编译器文献中有有趣的技术可以应用在这里。如果你把 AI for math 的 OG 们,比如 François Charton 和 Alberto,他们做数学发现,如果你让他们和形式化证明的人交谈,就会产生关于自我改进意味着什么的有趣想法。这些讨论每时每刻都在 ICM 办公室里发生,而我作为其中一员,我觉得非常谦卑,而且感觉这在某种意义上甚至比学术界更学术,因为它更多地扩展和延伸了我的智力边界。

Uh, like there are some really secret bands, rock and roll bands in China. They're obviously very politically rebellious and also write about all topics. The lyrics are quite interesting for a 5-year-old to listen to. There's this part of me that really appreciates a diverse range of things happening in my life and also the contrarian appeal. So I don't actually think of myself as the most academic academic when I was in academia, and I'm definitely also not the most founder founder. I'm somewhat in between. I do think that this startup is really fun, as in it's such a broad new adventure, and you're a bit living the rock and roll day every day. It's very high entropy and things are wildly exhilarating. You're empowering others to do their research of their dream, to build their legacy of scientific or achieve their scientific ambition. I also think this sort of making sure things run and serving the world-class researchers we have is another part that I really greatly enjoy. I think the transition has not actually — I don't feel it. I feel like somehow I found my native land. I think there's this very specific kind of thunder personality, and all the sayings of like trips on your shoulder, or all these sayings about being both very spontaneous and very nerdy methodological. I find myself to be everything in this portfolio. So I just really love my job. I thought there would be some sort of shock or transition. No, I feel like a fish in the water. It's really great. I think also working with this team of people coming from — I don't think there's another team with AI, programming languages, and mathematics, the three pillars together, which is very interesting if you think about it. Programming language is more system side, and we have many people in the intersection of programming language and AI, intersection of AI and math, and intersection of all these applied domains like chips or code verification with these. So by luck or something, the people we hire for the three pillars also have track records in their path in each of these application areas. So it felt like the right team, and working alongside them is a lifetime honor. Seeing how they think about things from all sorts of perspectives — like the mathematicians who might not know machine learning give us a great deal of inspirations. The compilers literature have interesting techniques that can be applied here. If you group the AI for math OGs like François Charton and Alberto in doing mathematical discovery, if you have them talk to the formal proving people, there are interesting ideas about what self-improvement means. These sorts of discussions happen every second inside the ICM office, and me being a part of that, I just think I'm very humbled by that, and I feel like it's even more academia in a sense, as in it expands and stretches my intellectual boundaries more.

18. 他为何享受创始人身份 Why he enjoys being a founder

Carina Hong

太棒了。这非常令人兴奋。我觉得这是一个很好的结束点。

Love it. That's super exciting. I think that's a great place to end.

Host

我不知道。如果你非要我追溯为什么我喜欢做创始人,我会说那个五岁时听硬核摇滚的我——和这有些关联。

I don't know. I think if you force me to trace why I enjoy being a founder, I would say that the 5-year-old hardcore rock and roll — there's some correlation to that.

Carina Hong

你能给我们推荐一首歌吗?我觉得我们应该把这首歌放在播客里。

Can you give us one song? I feel like we should put the song in the podcast.

Host

哦天哪。歌太多了,它们——我之后可以发给你一些。是的,我需要很长时间……

Oh Jesus Christ. There's so many songs and they are — I can send you something after this. Yeah, it takes me a long time to...

Carina Hong

发给我们一个播放列表吧。你有 Spotify 播放列表之类的吗?

Send us a playlist. Do you have a Spotify playlist or something?

Host

我有——我可能有一个 YouTube 播放列表。那些歌太——所有乐队现在都解散了。

I have — I can probably have a YouTube playlist. Those songs are so — all the bands are broke now.

19. 保持挑战者心态 Preserving the underdog mindset

Carina Hong

就像他们解散了,你会发现他们穷困时写的歌,那些成功的歌,比后来的好得多。我不认为是商业化的压力。我只是觉得,那种想要成功的强烈渴望和与之相伴的饥饿感起了作用。这正是我们在 Axiom 努力保持的东西。我们想要一直处于不舒服的状态。我们很清楚,有大型老牌公司存在,而我们入场稍晚,所以无论取得什么里程碑,我们都要保持这种后来者的心态。我认为保持这种小而强的文化,像孩子一样充满好奇,不把自己太当回事,这很重要。这就是为什么现在是我们建设的完美时机。我们正处于那个阶段。有点像研究生,他们有时会在读研期间做出最好的工作。

Like they are all disbanded and you find that the songs they write when they are poor, the ones that make it, are significantly better than the later ones. I don't think it's the pressure of commercialization. I just think there's something about the burning desire to make it and the hunger associated with that. That's something we try really hard to preserve here at Axiom. We want to be constantly uncomfortable. We are acutely aware that there are large incumbents and we are a little bit late to the game, and we want to have that underdog mindset regardless of what milestones we achieve. I think preserving that small and mighty culture, with curiosity like a child and not taking ourselves seriously, is important. That's why it's a perfect time for us to build right now. We're exactly at that phase. It's a bit like grad students who sometimes produce their best work.

Host

完全同意。

Totally.

Carina Hong

是的。

Yeah.

Host

很高兴见到你。这真的很有趣。谢谢。

Well, great to meet you. This is really fun. Thank you.

Carina Hong

很高兴见到你。太棒了。非常感谢。

Great to meet you. Awesome. Thank you so much.

Host

非常感谢收听本期 Gradient Descent。请继续关注未来的节目。

Thanks so much for listening to this episode of Gradient Descent. Please stay tuned for future episodes.

互动版:逐字朗读 + 针对本期提问 →