AI Podcast › 诺姆·布朗 › 本期
大规模测试时计算与模型评估 Large-Scale Test Time Compute and Model Evaluation
诺姆·布朗 Noam Brown · No Priors 播客 · 2026-06-26 · 约 36 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview Sarah Goa 与 Nome Brown 探讨 AI 模型评估的缺陷、大规模测试时计算的影响,以及当前基准测试为何无法反映模型真实能力。
Sarah Goa and Nome Brown discuss the broken state of AI model evaluations, the impact of large-scale test time compute, and why current benchmarks fail to capture true model capabilities.
要点 · TL;DR 模型能力随测试时计算量增长,而不仅取决于权重。 Model capability scales with test-time compute, not just weights. 基准测试必须包含计算成本才能公平比较。 Benchmarks must include compute cost for fair comparison. 安全评估忽略了大规模测试时计算预算。 Safety evaluations ignore large-scale test-time compute budgets.
核心观点 · Key points 模型能力现在是测试时算力预算的函数,而不仅仅是模型权重。 Model capability is now a function of test-time compute budget, not just model weights. 基准测试必须用词元、成本或时间作为 x 轴来公平比较模型。 Benchmarks must be evaluated with an x-axis of tokens, cost, or time to compare models fairly. 当前的安全评估没有考虑大规模测试时算力预算。 Current safety evaluations don't account for large-scale test-time compute budgets. 在许多任务上,模型随着测试时算力增加而持续改进,不会饱和。 Models improve continuously with more test-time compute on many tasks, without plateauing. 模型发布周期比全面评估能力所需的时间更快。 The model release cycle is faster than the time needed to fully evaluate capabilities.
反共识 · Contrarian takes 当前模型可以高效思考数周,使得传统基准测试的饱和点无法达到。 Current models can think productively for weeks, making traditional benchmark plateaus unreachable. 行业处于发布基准测试网格但不控制算力的不良均衡中。 The industry is in a bad equilibrium of publishing benchmark grids without controlling for compute. 由于测试时算力需求,递归自我改进受时间瓶颈限制,而非智能。 Recursive self-improvement is bottlenecked by time, not intelligence, due to test-time compute needs. 模型缺乏研究品味,即使有巨大推理预算也无法完全取代研究人员。 Models lack research taste and cannot fully replace researchers even with huge inference budgets. 多智能体协调仍处于初期;模型不像人类文明那样积累知识。 Multi-agent coordination is still nascent; models don't accumulate knowledge like human civilization.
本期章节 · Chapters(共 19) 测试时计算评估模型 Evaluating Models with Test-Time Compute 思考时间与实际应用 Thinking time and practical use 基准测试与评估现状 Benchmark maxing and evaluation landscape 个人评估方法 Personal evaluation methods 扑克机器人的推理演进 Reasoning progression in poker bot creation 安全评估的影响 Implications for safety evaluations 测试时计算评估模型能力 Evaluating model capabilities with test-time compute GPT-5.5 推翻 Erdos 单位距离猜想 Disproof of Erdos Unit Distance Conjecture with GPT-5.5 大规模测试时计算对研究方向的影响 Impact of Large-Scale Test-Time Compute on Research Direction 模型在研究中的局限性示例 Examples of model limitations in research RSI 与渐进式起飞框架 Framing of RSI and gradual takeoff 多智能体与知识积累 Multi-agent and knowledge accumulation 无快速起飞的前沿竞争 Competition at the frontier without fast takeoff AI 风险的严重性 On the seriousness of AI risks 模型用于高风险决策 Using models for high-stakes decisions 与学术界在基准测试上的分歧 Disagreement with research community on benchmarks 路由层与测试时计算 Routing layers and test-time compute 结语 Closing remarks 引言 Introduction
阅读全文双语转录 →