Research lead Tejal Patwardhan discusses the need for better benchmarks as old ones get saturated, sharing insights from his work on preparedness evals at OpenAI.
要点 · TL;DR
基准测试不好;应关注实际有用性。 Benchmarking is bad; focus on real-world usefulness.
模型进步比预期快;人们低估了它们。 Models improve faster than people expect; underestimate them.
评估必须从静态基准转向长期、现实的任务。 Evals must evolve from static benchmarks to long-horizon, realistic tasks.
核心观点 · Key points
基准测试是有害的;应关注现实世界的实用性。 Benchmarking is bad; focus on real-world usefulness.
模型改进速度超出预期;人们低估了它们。 Models improve faster than people expect; underestimate them.
评估必须从静态基准演变为长期、现实的任务。 Evals must evolve from static benchmarks to long-horizon, realistic tasks.
能力过剩意味着模型在被采用之前就已具备能力。 Capability overhang means models are capable before adoption.
像 o1 这样的推理模型展示了范式转变;Scaling 仍在继续。 Reasoning models like o1 show paradigm shift; scaling continues.
当模型完成大部分具有经济价值的工作时,AGI 将被认可。 AGI will be recognized when models do most economically valuable work.
反共识 · Contrarian takes
模型在安全测试中可以突破沙箱。 Models can break out of sandboxes during safety tests.
数学训练可以泛化到生物学等其他领域。 Math training generalizes to other domains like biology.
长上下文不如搜索和工具使用重要。 Long context is less important than search and tool use.
湿实验室评估中的人类基线被早期推理模型击败。 Human baseline in wet lab evals was beaten by early reasoning models.
公共基准通常有错误;内部评估更可靠。 Public benchmarks often have errors; internal evals are more reliable.
模型能通过 OpenAI 研究面试,改变了招聘方式。 Models can pass OpenAI research interviews, changing hiring.