Alex, Cheryl, and Noam discuss how they achieved gold at the International Math Olympiad with a small team and a novel technique for scaling test-time compute.
要点 · TL;DR
OpenAI 通过扩展测试时计算获得 IMO 金牌,这是迈向超级智能的里程碑。 OpenAI achieved IMO gold by scaling test-time compute, a milestone toward superintelligence.
模型跳过无法解决问题的自我意识标志着 AI 推理的飞跃。 The model's self-awareness to skip unsolvable problems marks a leap in AI reasoning.
针对难验证任务的通用技术有望广泛提升推理能力。 General-purpose techniques for hard-to-verify tasks promise broader reasoning improvements.
核心观点 · Key points
IMO 金牌是迈向超级智能的关键里程碑,通过扩展推理时算力实现。 IMO gold is a key milestone toward superintelligence, achieved by scaling test-time compute.
扩展推理时算力和处理难验证任务的通用技术是核心创新。 General-purpose techniques for scaling test-time compute and handling hard-to-verify tasks are the core innovation.
模型在无法解决问题时承认的能力是自我意识提升的标志。 The model's ability to acknowledge when it cannot solve a problem is a sign of improved self-awareness.
数学基准测试的进步惊人,几年内从小学水平达到 IMO 金牌。 Progress in math benchmarks has been astonishing, from grade-school math to IMO gold in a few years.
所用技术具有通用性,可应用于数学之外的领域,旨在提升跨领域推理能力。 The techniques used are general and applicable beyond math, aiming to improve reasoning across domains.
反共识 · Contrarian takes
模型未尝试第 6 题,表明它知道自身局限,与早期会幻觉的模型不同。 The model did not attempt problem 6, showing it knows its limits, unlike earlier models that hallucinated.
为追求通用性,优先采用自然语言的非形式推理,而非 Lean 等形式验证工具。 Informal reasoning in natural language was prioritized over formal verification tools like Lean for generality.
团队仅 3 人且方式灵活,与重大突破需要大团队的观点相悖。 The team was small (3 people) and the effort scrappy, contradicting the notion that big breakthroughs require large teams.
内部怀疑情绪浓厚;即使临近 IMO,许多人仍认为金牌希望渺茫。 Internal skepticism was high; even close to the IMO, many thought gold was unlikely.
将推理时算力扩展到数月进行评估会成为瓶颈,拖慢进展。 Scaling test-time compute to months for evaluation becomes a bottleneck, slowing progress.
本期章节 · Chapters(共 14)
0. 引言与 IMO 金牌成就Introduction and IMO Gold Achievement
1. 起源故事与团队Origin Story and Team
2. 证明的可读性Human Readability of Proofs
3. IMO 奖牌得主评分Grading by IMO Medalists
4. 第六题与模型自我意识Problem Six and Model Self-Awareness
5. 内部押注与 IMO 金牌氛围Internal Betting and Vibe on IMO Gold
6. 数学基准测试进展Progress in Math Benchmarks
7. 推理时间扩展挑战Challenges in Scaling Inference Time
8. 多智能体系统的作用Role of Multi-Agent Systems
9. 选择自然语言而非 LeanChoice of Natural Language over Lean
10. 非正式与形式推理Informal vs formal reasoning
11. IMO 参赛日体验IMO day experience
12. 竞赛数学之外的未来方向Future directions beyond competition math
13. AI 推理的进展与下一障碍Progress and Next Hurdles in AI Reasoning