Noam Brown 讨论 AI 模型如何在数学推理上超越人类能力,以 OpenAI 最近推翻 Erdős 单位距离猜想为例。
Noam Brown discusses how AI models have surpassed human ability in mathematical reasoning, exemplified by OpenAI's recent disproof of the Erdős unit distance conjecture.
要点 · TL;DR
OpenAI 的通用模型推翻了一个埃尔德什猜想,展示了 AI 在结合数学领域方面的超人类能力。 OpenAI's general model disproved an Erdős conjecture, showing AI can combine math areas superhumanly.
数学中的自我对弈比游戏更难,因为没有明确目标,限制了 AI 的提升。 Self-play in math is harder than in games due to no clear objective, limiting AI improvement.
随着 AI 证明变得过于复杂,人类难以验证,可扩展监督成为关键。 Scalable oversight is key as AI proofs become too complex for humans to verify.
核心观点 · Key points
模型在结合不同数学领域方面已超人类,但在深度推理上尚未超越。 Models are superhuman at combining different math areas but not yet in deep reasoning.
数学中的自我对弈比零和游戏更难,因为缺乏明确目标。 Self-play in math is harder than in zero-sum games due to lack of clear objective.
AI 在数学上的进步遵循解决需要人类更长时间的问题的轨迹。 AI progress in math follows a trajectory of solving problems with longer human time horizons.
OpenAI 的最高杠杆是改进通用模型,而非解决特定数学问题。 The highest leverage for OpenAI is improving general models, not solving specific math problems.
可扩展监督是关键挑战,因为 AI 证明变得过于复杂,人类无法验证。 Scalable oversight is a key challenge as AI proofs become too complex for humans to verify.
反共识 · Contrarian takes
厄尔多斯反例并非由数学专用模型得出,而是通用推理模型。 The Erdős disproof was not a math-specific model but a general-purpose reasoning model.
竞赛数学可能很快饱和;真正的研究才是新前沿。 Competition math may soon be saturated; real research is the new frontier.
人类数学家的价值将转向提出正确问题和研究品味。 Human mathematicians' value will shift to asking right questions and research taste.
数学特殊之处在于它纯粹受推理瓶颈限制,不像物理或湿实验室。 Math is special because it's purely bottlenecked by reasoning, unlike physics or wet labs.
数学中的人机协作期可能比国际象棋短,因为缺乏自我对弈。 The centaur period in math may be shorter than in chess due to lack of self-play.
本期章节 · Chapters(共 15)
引言与埃尔德什猜想证伪Introduction and the Erdős Conjecture Disproof
AI 数学推理与半人马时期AI's mathematical reasoning and centaur period
自我对弈的重要性Importance of Self-Play
聚焦通用模型Focus on General-Purpose Models
谁将在数学领域善用 AIWho Will Excel with AI in Math
人机研究互补技能Complementary skills between humans and AI in research
证明界限与构造反例的区别Distinction between proving bounds and constructing counterexamples
AI 连接不同研究领域的优势AI's advantage in connecting diverse research areas
数学为何适合 AI 及其他领域Why math is a good domain for AI and other promising areas
数学与物理:推理瓶颈 vs 实验瓶颈Math vs Physics: Bottleneck by Reasoning vs Experiments
IMO 饱和与满分预期IMO Saturation and Perfect Score Expectations
第六题难度:几何与人机相关性Problem Six Difficulty: Geometry and Human-AI Correlation
竞赛数学与研究数学:视野与技巧Contest Math vs Research Math: Horizon and Techniques