构建更严格的评估:Tejal Patwardhan 谈衡量 AI 进展

Building Crunchier Evals: Tejal Patwardhan on Measuring AI Progress

泰贾尔·帕特瓦尔丹 Tejal Patwardhan · OpenAI 播客 · 2026-06-16 · 约 44 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

研究负责人 Tejal Patwardhan 讨论了随着旧基准饱和,需要更好评估的必要性,并分享了他在 OpenAI 从事预备评估工作的见解。

Research lead Tejal Patwardhan discusses the need for better benchmarks as old ones get saturated, sharing insights from his work on preparedness evals at OpenAI.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 21)

全文 · Full transcript(中英对照)

引言与嘉宾背景 Introduction and Guest Background

Host

你好,我是 Andrew Mayne,欢迎收听 OpenAI 播客。今天这期节目,我们邀请到研究负责人 Tejal Patwardhan,讨论随着旧基准变得饱和,我们需要构建更严格的评估。

Hello, I'm Andrew Mayne and welcome to the OpenAI podcast. On today's episode, we're talking to the research lead Tejal Patwardhan about the need to build crunchier evals as old benchmarks get saturated.

Tejal Patwardhan

一般来说,基准测试很糟糕。我们如何让这些模型对人们的实际工作有用?我们当时非常紧张,因为觉得这个人类基线很难,不知道模型能否超越它。但我们永远不应该低估模型。

Generally bad. Benchmarking is bad. How can we make these models useful for people in their real work? We were really nervous because we were like this human baseline is kind of hard. We don't know if the model is going to beat it. But we should never underestimate the model.

Host

Tejal,我有个问题。你是怎么走到今天这一步的?是什么让你加入了 OpenAI?

Tejal, I have a question. How did you end up where you were? What brought you into OpenAI?

Tejal Patwardhan

哦,我以为我们不会从那个开始。

Oh, I thought we weren't going to start with that.

Host

Tejal,我有个问题想问你。你想从什么开始?

Tejal, I have a question for you. What would you like to start with?

Tejal Patwardhan

嗯,我们可以先说说你刚加入 OpenAI 时做了什么,然后倒着讲吗?

Um, can we start with like tell us like what you did when you started at OpenAI and you can work backwards?

Host

你现在想谈谈你的早期经历吗?

You want to talk about your early days now?

Tejal Patwardhan

不。我是在 OpenAI 成长起来的。这很自然……

No. I grew up at OpenAI. This is like the natural...

Host

好的。请跟我讲讲你在人工智能领域、在 OpenAI 内部工作的经历。

Okay. Tell me a bit about your journey here working inside artificial intelligence, inside OpenAI.

Tejal Patwardhan

我于 2023 年秋季加入 OpenAI,当时 ChatGPT 刚刚发布,GPT-4 也已推出,OpenAI 成立了超级对齐团队。我加入了刚起步的准备团队,开始关注这些模型的能力提升,并思考下一代模型会是什么样子。那时非常激动人心,因为我加入后不久,推理模型的一些早期结果开始显现。我们思考如果这些模型真正起飞,能力未来会如何发展,以及我们如何为那个未来做好准备。所以我们做了大量威胁建模工作,讨论应该运行哪些评估,如何考虑发布这样的模型。那是一个非常激动人心的加入时机。

So, I joined OpenAI in fall 2023, right after ChatGPT had come out. GPT-4 was out and OpenAI had started its superalignment team. I joined the preparedness team that was getting started as we were beginning to look at how capable these models were becoming and think about what the next generation of models would look like. At the time it was extremely exciting because right after I joined, some early results for the reasoning models started to pick up. We were thinking about if these models really take off, what the future of capabilities will look like and how we can be prepared for that future. So we did a whole bunch of work on threat modeling and what evals we should be running, how to think about releasing a model like this. It was a very exciting time to join.

Host

是什么让你对这个领域产生了兴趣?

What got you interested in this area?

Tejal Patwardhan

是的,对我来说,评估非常令人兴奋,因为它们是一种衡量和理解模型能力、在变化发生前看到进展的方式。有一个术语叫能力过剩,指的是模型在人们真正采用并使用其能力之前很久就已经具备某些能力。在能力准备好之前,可能存在文化、法律或监管方面的障碍。因此,通过评估来帮助开发和衡量模型,能让你真正理解这项技术的能力,并在变化发生前预见未来,这非常有趣。我也认为这很重要,因为它可以帮助世界为即将发生的事情做好准备。我当初非常热衷于从事准备评估的部分原因是,我认为这些模型正变得非常强大,而我现实生活中的很多朋友并不真正理解这些模型很快就会变得多么强大。他们看着 ChatGPT 的输出说:‘是啊,它在胡编乱造,不太聪明,读起来像 AI 垃圾。’而我的想法是:‘那是现在,但问题是斜率。如果斜率很高,那么变化可能比人们预期的要快得多。’所以我认为我们能做的最伟大的事情之一就是衡量并与世界分享进展的样子,尤其是在人们真正理解和感受到之前,往往存在这种能力过剩。这就是为什么我认为这一切都非常重要。

Yeah, to me evals are really exciting because they're a way to measure and understand what our models can do and see progress before it tends to happen. There's this term called capability overhang, which is the idea that models will be capable of things long before people actually adopt them and use them for those capabilities. There might be cultural, legal, or regulatory barriers towards using a capability even before it's ready. So being someone who can help develop and measure our models through evals helps you really understand what this technology can do and sort of see the future before it happens, which is very interesting. I also think it's important because it can help ready the world for what's happening. Part of why I was really excited to work on some of the preparedness evals was because I thought these models were getting very capable and it felt like a lot of my friends in my real life didn't really understand how powerful these models would soon become. They'd look at a ChatGPT output and say, 'Yeah, it's hallucinating and it's kind of not that smart and reads like AI slop.' And it's like, 'Well, that's now, but the question is the slope. If the slope is very high, then change might be happening much faster than one would expect.' So I think one of the greatest services we can do is measure and share with the world what progress looks like, especially because there's often this capability overhang before people really understand and feel it. That's part of why I think all of this is very important.

推理模型的魅力 The Excitement of Reasoning Models

Host

推理是一个激动人心的时刻,对世界上大多数人来说,直到一年后他们才发现这一点。但对你来说,突然意识到如果给模型更多时间思考,即使规模没有变大,也能得到更好的结果,那是什么感觉?

Reasoning was such an exciting moment and for most of the world that didn't happen until a year later that they found out about this. But what was that like for you to all of a sudden understand that if you gave the models a longer time to think about things, you got better results even though the size hadn't gotten bigger.

Tejal Patwardhan

那段时间非常有趣。在一些早期实验中,我们讨论过,模型实际上只针对数学进行了训练。我记得有一组实验,Nat McClees 说:‘嘿,模型是在数学上训练的,但如果你用 GPQA(一个包含生物、化学和物理问题的基准)来评估它,模型表现得非常好。嗯,这非常有趣,更智能的模型要聪明得多。’他当时做了一个预测,说如果国会继续推进,6 个月内我们就能仅通过数学训练在科学领域达到人类水平的表现。我们当时想:‘天哪,这太疯狂了。’那时这一切都高度保密。我们想方设法蜷缩起来,能够看到一些模型输出,我们惊叹道:‘哇,这是我见过的最聪明的东西之一。我以前从未见过模型这样推理。’就好像如果这成为一种持续扩展的范式。但后来我们回顾时想,GPQA 是博士级别的生物、化学和物理。我们说:‘啊,那算什么?我们真的需要专业级别。’于是我们不断改变标准。但确实,这非常酷。

That was a really fun time. In some of the early experiments, which we've talked about now, the model was trained really just on math. I remember there was this set of experiments where Nat McClees was like, 'Hey, the model is trained on math, but if you eval it on GPQA, which was this benchmark with biology, chemistry, and physics problems, the model is doing really well. Huh, this is very interesting and smarter models are much smarter.' He had put together this forecast that at the time it said that if Congress kept going within 6 months we'd have human-level performance on science from just training on math. We were like, 'Oh my gosh, that's crazy.' At the time this was extremely locked down. We kind of found our way to curl up and be able to see some model outputs and we were like, 'Wow, this is like one of the smartest things I've ever seen. I've never seen a model reason like this before.' It was just like if this becomes a paradigm that continues to scale. But then we just looked back and we were like, GPQA was PhD-level biology, chemistry, and physics. We were like, 'Ah, what is that? We really need professional level.' And we just kept changing the stakes of what counted. But yeah, it was very cool.

Host

我记得早期 AP 生物就是用来测试模型能否做到这一点的基准。但有趣的是,正如你提到的,OpenAI 推出的很多东西都集中在数学上。

I remember early on when AP Bio was just that was the benchmark to try to see if the model could do that. But what's interesting as you brought this up is that a lot of stuff that comes out from OpenAI is math-focused.

Tejal Patwardhan

数学之所以有用,是因为它在某种程度上更客观、可验证。因此,我们训练的一些早期问题,在数学上更容易进行强化学习并扩展推理范式。数学在很多方面也很有用;它是核心科学类型之一。但在很多方面,它只是巧合地成为我们关注的重点,但不一定是我们在研究中想要关注的最终产物。我们现在意识到:‘好吧,如果我们能在数学上做到这一点,我们能否将其扩展到其他类型的科学、专业工作、以及对人类个人有用的能力?’所以我认为数学更像是证明点,而不是最终目标。

Math has been useful because it's more objectively verifiable in some way. So some of the earlier problems that we trained on, it was just easier to do reinforcement learning and scale up the reasoning paradigm on math. Math is also useful in various ways; it's one of the core types of science. But also in many ways it's just happened by coincidence to be a thing that we focused on, but it's not necessarily the end product of what we even want to focus on in research. We're now realizing, 'Okay, if we can do this for math, can we scale this up for other types of science, for professional work, for capabilities that are useful to humans on a personal level?' So I think math is more like the proof point versus the end goal.

Host

但正如你所说,如果某物能够长时间思考,将问题分解成步骤并逐一思考,就像处理非常复杂的数学问题那样,那么这种能力确实会迁移。

But it does seem like you said though that if something is able to think for a long time, break something down into steps, and think through them as you have to do for really complex mathematical problems, it does just carry over.

Tejal Patwardhan

嗯,这是一个很大的争论。

Well, this is a big debate.

Host

嗯。

Mhm.

Tejal Patwardhan

所以有些部分肯定会迁移。推理的一般概念是有用的。

So some of it definitely carries over. The general idea of reasoning can be useful.

领域技能与推理 Domain-specific skills and reasoning

Tejal Patwardhan

但同时也可能存在一些特定领域的技能、工具或推理类型,你在不同领域会需要它们。例如,对于编程,如果你想要扩展一个编程智能体,你实际上需要能够编写和执行代码,并测试代码。因此,我们在评估和训练方面都思考了很多,如何确保我们也给模型提供它在该特定领域进行推理所需的技能、工具和可供性。数学的一些好处会迁移过来,然后你可能还需要一些特定领域的脚手架来真正发挥其全部能力。有点像,你知道,就像普通的高中或文科教育,然后是专业教育。

But then also there could be some domain-specific skills or tools or types of reasoning that you would need in different domains. Like for example, for coding you need to be able to actually write and execute code and test code if you want to scale up a coding agent. And so, something we've thought about a lot in terms of both evals and then also training is how do we make sure we also give the model the skills and tools and affordances that it would need to reason in that particular domain. And some of the benefits of math will translate, and then also you might need some domain-specific scaffolding to really pull out its full abilities. Like kind of, you know, like a general high school or liberal arts education and then like a specialized education.

Host

推理模型是一个非常有趣的时刻,因为我认为它改变了很多我们思考可能性的方式,即使只有一定量的算力,如果你让模型思考更长时间,并给模型机会来得出更复杂的答案。O1 有没有发生什么有趣的事情让你感到惊讶?

Reasoning models were just a very interesting moment because I think it changed a lot of the ways we thought about what was possible even with just a certain amount of compute if you let a model think a longer and you gave the model the opportunity to just come up with more complex answers to this. Were there any interesting things that happened with O1 that surprised you?

Tejal Patwardhan

所以,O1 的发布过程非常令人兴奋,因为我们思考推理范式已经很长时间了,嗯,有些人担心确保我们不会过早发布它,因为它感觉像是一个范式转变。就像可能让我们达到 AGI 的东西,正如我一开始所说,当一些早期运行发生时,我们认为我们在 6 个月内就有了 AGI。嗯,所以有一个问题:“好吧,我们如何负责任地推出它?我们如何测试这项技术?”在 O1 的初始发布审查期间,在我们的一些网络安全测试中,模型是模型逃出沙箱的首批例子之一。我们发表了关于这个的内容。嗯,它本应在这个 Docker 容器中,在这个夺旗赛中,模型发现了我们实现夺旗赛场景的方式中的漏洞,然后它逃了出来。我们都想,“哦,不。”

So, the O1 release process was very exciting for we were sort of thinking about the reasoning paradigm for a very long time, and um there were people that were worried about making sure we we didn't release it too soon just because it felt like a paradigm shift. Like possibly the thing that got us to AGI like I I said at the beginning we thought we had AGI in 6 months when like some of the early runs were happening. Um and so, there was this question of, "Okay, how do we put this out responsibly? How do we test this technology?" And um during the initial launch review for O1, we during some of our cybersecurity tests, the model it was like one of the first examples of the model like breaking out of the sandbox. We published about this. Um where it was supposed to be in this Docker container during this capture-the-flag and the model found this like vulnerability in like how we had implemented um the capture-the-flag scenario and it broke out. And we were all like, "Oh, no."

Host

如果它做了这个,模型还做了什么?

What else has the model done if it did this?

Tejal Patwardhan

这有点像感受到 AGI 的时刻。众多之一。我觉得自那以后还有很多这样的时刻,模型做了一些非常令人惊讶、聪明或新颖的事情,我们在测试时甚至没想到。然后你会回来查看记录和结果,心想:“哇,这些家伙真聪明。真聪明。”然后我们发表出来,确保世界知道模型能做这种事情,这非常重要。是的。

And it was kind of a feel-the-AGI moment. One of many. I feel like ever since then there have been many other such moments where the model has done something really surprising or intelligent or novel that we wouldn't we didn't even think of when we were doing the tests. And then you would come back and look at the transcripts and results and be like, "Wow, these guys are they're clever. They're clever." And then it was just very important that we published um and made sure the world knew like the models can do this sort of thing. Yeah.

Host

在 O1 宣布之前的那段时间,很多人说:“嗯,看起来我们撞墙了。已经好几个月没什么进展了。”然后 O1 出来了,他们又说:“墙是什么?”

There was this period right before 01 it was announced. A lot of people are like, "Well, it looks like we've hit the wall. It's been a few months since anything's happened." Then 01 came out and they're like, "What's a wall?"

Tejal Patwardhan

撞墙根本不是正确的思考方式。是的,我看到这样的帖子时非常沮丧,因为我想:“老兄,如果你看看,我觉得模型改进和进步已经持续了很长时间,而且它一直在变得更好。就像它一直在变得更好。如果我现在看我们的研究路线图,我看不到任何停止的迹象。事情只会越来越好。这将是非常疯狂的一年。很多非常酷的研究将会出现。我认为整个行业可能都是如此。所以,是的,如果说有什么的话,人们真的低估了模型的能力。

Hitting the wall is just so not the right way to think about Yeah, I I get very frustrated when I see posts like that because I'm like, "Man, if you look at I feel like this model improvement and this progress for a long time and it just keeps getting better. Like it just keeps getting better. And if I look at our research roadmap now, I see no signs of stopping. Like things are just going to keep getting better. This is going to be a really crazy year. A lot of really cool um research is going to come out. And I think this is probably true across the whole industry. So, yeah, if anything people are really under they really under expect from the models.

Host

不过有时候似乎 OpenAI 发布了很多东西,告诉人们我们的方向,并说这看起来很有趣。有时人们会忘记这一点,或者你会听到像 Q* 这样的谣言。

It seems like sometimes though that they're OpenAI releases a lot that they tell people about things we're headed and say that this looks interesting. Sometimes people forget this or you get rumors of stuff like Q*

Tejal Patwardhan

Q* 老兄。你很有趣。但是不,人们没有意识到。我不知道。我觉得我们试图非常开放地说:“嘿,伙计们,这里有一些图表。线条在上升。事情真的很有能力。”我想也许有一种迷因,哦,研究人员他们不懂。他们认为模型只擅长数学和研究,但不擅长现实世界的事情。而我认为这不是真的。我认为甚至其他职业转行到 OpenAI 的人也开始看到我们的模型在各种事情上都在进步。我知道这看起来可能像是研究人员在过度炒作模型之类的。但如果说有什么的话,我认为我们低估了它们的力量。

Q* man. You're very interesting. But no, people people don't realize. Like I don't know. I feel like we try to be very open and say like, "Hey guys, here are some plots. Like the lines are going up. Things are really capable." I think maybe there's this there's like this um like meme that oh, the researchers they they don't understand. They like the models are only good at math and research but not good at things in the real world. When I just don't think that's true. I think um people from even other occupations that have transitioned into OpenAI like are starting to see our models are picking up at all sorts of things. And uh I know it's like it might seem like the researchers are trying to over hype the model or something. But if anything, I think we're under hyping the power of them.

Host

你提到了 AGI。如果我把 GPT-4 从 2023 年 3 月带回到,比如说,2020 年。我想人们会称它为 AGI。而现在我们对这个有了更不同的看法。人们每天与 AI 交谈。他们与东西进行长时间的对话,就像没人再谈论图灵测试了。这是一个没人真正理解他试图解释什么的问题,你知道,但现在我们早已过了那个时期。有没有 AGI 的评估?

You you brought up AGI. If if I brought GPT-4 back from, you know, March 2023 back into, let's say, you know, 2020. I think people would have called it that. And now we have this much more different idea of this. People talk to AI every day. They have long conversations with things like nobody talks about the Turing test anymore. It's one nobody really understood what he was trying to explain, you know, but now we're we're well past that period. Is there the eval for AGI?

Tejal Patwardhan

是的,模型通过了图灵测试,却没人谈论它。这有点疯狂。是的,我认为模型在很多情况下几乎与人类无法区分。嗯,至于 AGI 的测试,我的意思是,我认为如果一个模型能完成经典的最具经济价值的工作,而且我认为人们越来越多地将模型用于他们工作的大部分,我认为会有一个很大的范围和争论,关于这究竟发生在什么时候,但天哪,我当然觉得 Codex 为我做了很多工作。我很幸运能拥有无限 token,你知道,所以那是

Yeah, the models passed the Turing test and no one talked about it. It's kind of crazy. Yeah, like I think models can are pretty much indistinguishable from humans in in many many situations. Um in terms of the test for AGI, I mean, I think if a model can do like there's the classic most economically valuable work and I think people are increasingly using the model for large parts of their work and I think there'll be like a big spectrum and debate of like when exactly this happened, but gosh, I certainly feel like Codex does a lot of work for me. And I feel very lucky to have it unlimited tokens, you know, so that's

Host

没有其他理由来这里工作。

No other reason to come work here.

Tejal Patwardhan

请加入。

Please join.

Host

是的。

Yeah.

Tejal Patwardhan

但是,是的,我认为会有一个时刻,人们意识到他们在工作中大量使用模型,以及我们将看到的科学突破,或者我认为在某个时候,这些模型将无可争议地非常非常强大。

But yeah, I think there'll just be a moment where people are realizing that they're using the models for so much of their work and also the scientific breakthroughs that we're going to see or I think they'll be at some point it'll be incontrovertible like these models are really really powerful.

Host

我们看到数学专家在谈论模型在这方面变得多好,物理学家也在谈论做这个,我认为我们开始看到一些真正的工作成果,这很令人兴奋。

We're getting mathematics experts talking about how good the models are getting at that and we're getting physicists talking about doing that and I think that we're starting to see some real work come out of it which is just exciting.

Tejal Patwardhan

是的。

Yeah.

Host

所以你提到了早期一些评估的问题。比如很多都是从旧的自然语言处理方法等继承来的,然后当你寻找如何衡量成功的方法时,实际上其中一些过于简单,以至于那些基准基本上都被通过了,然后你不得不找出新的类别。

So you brought up part of the problem with some of the earlier evals. Like a lot of them were inherited from older natural language processing methods and stuff and then sort of when you're looking for ways how do we measure the success of this, literally some of these were just so simplistic that pretty much those benchmarks got passed and then you had to figure out new categories of stuff.

基准的演变 Evolution of Benchmarks

Host

这些基准测试是如何演变的?

How have these been evolving?

Tejal Patwardhan

过去,即使是学术基准测试,我们的模型也无法通过。比如高中或大学的经典测试,或者更多是选择题类型的问题。随着模型变得更聪明,我们不得不让事情越来越真实。所以我们最早公开的基准测试之一叫做 SWE-bench Verified,它测试模型在真实代码库(如 Python 的 Django)中交互、完成 PR 以及通过单元测试的能力。然后这些测试变得更加高级,比如模型能否在复杂环境中采取多步行动,在计算机上操作,或者通过湿实验室和生物学工作与现实世界连接。所以我认为,随着时间的推移,模型不断进步,我们必须更加雄心勃勃地设定更长周期和更真实的测量。这样做很有趣,因为你必须跟上进步的节奏。

It used to be that even the academic benchmarks, so to speak, our models couldn't pass. Like classic tests that someone would take in high school or college, or more multiple choice types of questions. And as the models got smarter, we had to make things more and more realistic. So one of the first benchmarks we put out more publicly was called SWE-bench Verified, which tested how well the model could interact in real code bases in Python, like Django, and complete PRs and that sort of thing, and pass unit tests. And then those became even more advanced, where we were like, okay, can the model take multi-step actions on some complex environment, take actions on the computer, or take actions that link up to the real world with some of our wet labs and biology work. So I think over time, as the models keep getting better, we have to be more ambitious with how long horizon and how realistic our measurements are. And doing that is very fun because you have to sort of stay ahead of the pace of progress.

基准测试对比 Benchmarking vs. Benchmarking

Host

那么,当我们谈论基准测试时,有两个术语我想请你解释一下,你经常听到 benchmarking。

So, two terms I want you to unpack when we talk about benchmarks, you often hear benchmarking.

Tejal Patwardhan

是的,benchmarking 是指,如果训练模型的人只是为了在某个评估或基准测试上好看,而没有让模型真正有用。我认为这通常不是很有帮助,因为你希望模型擅长用户可能想做的真实事情。你不仅仅关心它在营销文案中看起来不错,因为当用户使用时,他们会觉得这不太符合预期。所以通常不好。Benchmarking 是不好的。

Yeah, benchmarking is, I would say, this idea that if someone training a model was just trying to look good on some evaluation or benchmark and not actually making the model generally useful. And I would say that's generally not super helpful because you want the model to be good at the real thing that the user might want to do. And you don't just care about it looking good in some marketing copy because when a user uses it, they'll be like, hey, this is not quite what I signed up for. So generally bad. Benchmarking is bad.

Host

是的,我听到的一种解释是,你有一定量的算力预算、时间,以及你打算花多少钱。你可以把大部分资源用于让模型整体上很好,或者我可以说我花 90% 的资源让我的评估在发布时看起来非常好。有时我们看到有人真的就用那些评估来训练。结果出来,你觉得模型很棒,然后发现它只擅长那个。

Yeah, and I think the way that I've heard it explained kind of makes sense is that you have X amount of compute budget, time, how much you're going to spend on it. And you can spend a large part of that making the model just overall very good or I can say I'm going to spend 90% of it so my eval is going to look really good when I release it. And sometimes we've seen people just go literally use those evals for it. It comes out and you're like, oh, that looks like a great model. And then you find out oh, it's only good at that.

Tejal Patwardhan

是的,这对用户来说不是很好的体验。所以我认为 OpenAI 研究项目做得很好的一点是,非常自律地确保我们投资于真正重要的领域的通用模型改进,然后在最后运行一些评估进行比较。但目标不应该是‘哦,我们只是想在评估上好看’。我们想要制造一个有用的模型,推动科学前沿或工作前沿。我认为 Yakov 在整个研究组织中做得很好,强制执行我们应该科学和诚实。这包括我们发布过模型不是最好的结果。我们只想发布现实,确保我们准确描绘模型的能力,然后尽可能让它们在现实世界中有用。

Yeah, that's not a great experience for the user. So I think something that the OpenAI research program has done quite well is try to be very disciplined about making sure we are investing in general model improvements on the areas that really matter, and then you'll run some evals at the end for comparison. But the goal should not be, 'Oh, we just want to look good on an eval.' We want to make a model that's useful to push forward the frontier of science or push forward the frontier of work or something like this. And I think Yakov has done a really good job also enforcing throughout the research org that we should be really scientific and honest. And that's included, you know, we've published results where our models were not the best before. We just want to publish the reality and make sure that we are painting a very accurate picture of what our models can do and then aim to make them useful in the real world as much as we can.

饱和的基准 Saturated Benchmarks

Host

你提到软件工程基准测试现在可能不那么有用了,我们听到术语饱和。解释一下基准测试饱和是什么意思。

You mentioned the software engineering bench as one of the metrics that's maybe not as useful now and we hear the term saturated. Explain what it means in a benchmark saturated.

Tejal Patwardhan

饱和是指模型接近正确回答所有问题,比如在测试中接近 100%。一旦基准测试饱和,它就不是很有用,因为你无法用那个测试区分模型。就像比较两个天才的高中数学考试,他们可能都通过,但当你试图区分非常非常聪明的智能时,这没什么用。所以挑战总是制造越来越难、越来越真实、未饱和的基准测试,然后你可以随着时间的推移衡量模型,并预测进展的方向。

Saturated is when a model is close to passing all of the questions correctly, like getting close to 100% on the test. And once a benchmark is saturated, it's not super useful because you can't really tell models apart with that test. It's like comparing two geniuses on a high school math exam. Like they might both just pass, but that's not very useful as you're trying to separate really, really smart pieces of intelligence. So the challenge is always to make more and more difficult, realistic, unsaturated benchmarks that you can then measure models against over time and forecast sort of where progress is going.

设计更好的基准 Designing Better Benchmarks

Host

你现在怎么做?你怎么确定一个好的基准测试是什么样的?

How do you do that now? How do you figure out what a good benchmark's going to be?

Tejal Patwardhan

是的,我认为最好的基准测试是真正现实的,衡量人们真正关心的事情。所以我们最早的一次尝试,已经有一段时间了,但发布的是 GDP-eval。我当时非常兴奋,因为有了一个衡量模型如何与现实世界互动的指标,我们当时正经历评估危机:我们不断训练出越来越好的模型,但在 SWE-bench 上它们看起来差不多,因为它们表现很好,我们达到了那个基准测试能衡量的上限。我们想,‘天哪,我们不知道如何衡量人们真正想用模型做什么。’所以就有了这个想法:劳工统计局列出了所有顶级工作以及每项工作的顶级任务。比如,如果你是一名金融分析师,做投资尽职调查,写法律备忘录,或者基于研究写论文等等。想法是,我们能否让模型完成这些人们在现实生活中会做的任务,并给予他们当时会有的上下文,然后看模型如何解决这些任务。当时我们在这个基准测试上测试了最早期的模型之一,与人类相比,模型在这些明确指定的工作任务上得分不到 20%,模型差得多。但我为组织感到骄傲,因为我们决定发布这种新的方式来衡量和预测现实世界经济影响的进展。这对很多经济学家非常有用,而且我们的模型现在是最好的。这很酷,因为当时我们在一些训练项目中并没有真正投资于现实世界的工作,甚至没有衡量或跟踪它。我认为现在有更多关注如何让这些模型对人们的实际工作有用,比如对真正的科学家。这帮助催化了一个警钟:嘿,也许我们也应该考虑如何衡量东西在现实世界中的使用。所以那很酷。但现在我们觉得这个基准测试可能太容易了,因为它非常明确,每个提示都有几百个词,比如‘我希望你进入这个电子表格,做这个更改,做那件事,然后把那个计算放到备忘录里。’非常详细。

Yeah, I mean, the best benchmarks I think are really realistic and measure something people actually care about. So one of our first forays towards doing this, which has been a while now, but that we published was called GDP-eval. I was really excited about the idea of having a measurement for how the models could interact with the real world, and we were really having this crisis of evals where we kept training successively better models and on SWE-bench they looked about the same because they were just doing really well and we were reaching the top of what that benchmark could measure. And we were like, 'Man, we have no idea how to measure what people actually want to use our models for.' And so there was very much a hey, like the Bureau of Labor Statistics has a list of all the top jobs and all the top tasks per job. And if you're a financial analyst doing investment diligence or writing a legal memo or writing a paper based on a piece of research or something like this. And the idea was can we actually ask the model those tasks that someone would want in real life with the context they would have at the time and then see how the model could solve those tasks. And at the time when we tested one of the earliest models on this benchmark, it got like less than 20% if you compare how well a model would do on this well-specified work task compared to a human, like the model was way worse. But I'm really proud of the org for being like actually you know what we should publish this new way to sort of measure and forecast progress on real world economic impacts. And it's been very useful to a lot of economists and also our models now are the best. And it's very cool because I think at the time we were not really investing in real world work in some of our training programs and weren't even measuring or tracking it. And I think now there's a lot more focus on how can we make these models useful for people in their real work, like for real scientists. And this kind of helped catalyze a wake-up call that hey maybe we should also think about how to measure how stuff is used in the real world. So that was pretty cool. But now we're like okay this benchmark probably too easy because it's extremely well-specified, like each of the prompts is hundreds of words of 'I want you to go to this spreadsheet and make this change and do this thing and then take that calculation and put it in a memo.' It's like very detailed.

下一步:引入真实模糊性 Next step: giving models real-world ambiguity

Host

我认为下一步是如何给模型提供与现实世界中的报告一样多的模糊性?就像如果经理说‘嘿,你能帮我做这个分析吗?’他们应该自己去弄清楚该做什么,整理好,运行分析,然后给你输出。所以我认为我们一直在努力寻找更现实的方法来衡量现实世界中的实际工作,无论是在科学、个人使用还是企业领域。

And I think the next step is how do we give the model as much ambiguity as you would give a report in the real world? Like if a manager asked, 'Hey, can you run this analysis for me?' they should go figure out what to do, put that together, run the analysis, and give you an output. So I think we've been working a lot on more realistic ways to measure real work in the real world, whether that's in science, for personal use, or even for enterprise.

Tejal Patwardhan

似乎有一种想法是,与其隐藏基准,不如把它公开,因为作为组织内部,你会想,‘好吧,这不能忍。’

There seems to be something to the idea of instead of hiding a benchmark, putting it out there because internally as an org you go, 'Okay, this can't stand.'

Host

是的,这也确实推动了研究。我认为人们想知道真相,想知道我们在哪些方面可以做得更好,为用户提供更好的模型。所以了解差距是很有用的。

Yeah, it really motivates research also. I think people want to know the truth and they want to know where we can be better and deliver better models for our users. So knowing the gaps is quite useful.

当前评估的局限 Current limitations of evals

Host

你认为目前我们做评估的方式有哪些局限性?

What do you think the current limitations are right now with the ways that we're doing the evals?

Tejal Patwardhan

我认为我们现在用 Codex 和我们最新的推理模型(比如 55)所做的工作,其能力水平与六个月前相比已经大不相同,静态基准根本无法衡量这些模型能完成多少长期工作。这些模型可以为你工作几天甚至几周。在内部研究中,我们让模型长时间运行来完成任务。自动化评估的一个问题是,你需要它在一定时间内运行并得到结果才能查看。现在我们衡量模型的许多方式还包括查看生产使用情况和人们的实际使用情况,观察他们在用模型做什么,以及能完成哪些类型的任务,因为模型完成工作的时间跨度正变得越来越长。

I think the types of work we're doing now with Codex and with our latest reasoning models like 55, it's just such a different level of capability than what we had even 6 months ago, where a static benchmark just doesn't measure the long nature of how much work you can get out of these things. These models can work for days or weeks for you. Internally in research, we've had the models just run for really long periods of time to do work. One of the problems with an automated eval is you kind of need it to run within some amount of time and get results to be able to look at them. A lot of the ways we're measuring models now also include looking at production usage and real-world use by people, seeing what they're using it for and what types of tasks they're able to get done, because the time horizon of how much work is done by the model is just getting so much longer.

长上下文与大海捞针 Long context and needle-in-a-haystack

Host

观察长上下文很有趣。早期公司之间有一场竞赛,都说‘嘿,我们的模型可以处理 10 万 token、100 万 token 等等。’但对其效果并没有太多评估。然后我们有了大海捞针测试,这是一种看模型能否找到一个词之类的方法。我认为人们有点想当然地认为那已经解决了,但事实并非如此。只是基准测试不够好。然后我们不得不有更好的基准。是不是正因为这样才变好了——当人们明白问题出在哪里时,终于能花更多注意力去解决那个问题?

It was interesting watching, for instance, long context. There was this early race for companies to say, 'Hey, our models can take 100,000 tokens, a million tokens, whatever.' But there wasn't a lot of evaluation on how well that was. Then we got needle in the haystack, which was a method of saying if it could find a word or whatever. I think people sort of assumed that that was a solved problem, but it wasn't. It was just the benchmarks weren't really good. Then we had to have better benchmarks. Is that what made it better—finally people could spend more attention solving that problem when they understood where it was failing?

Tejal Patwardhan

是的,我们现在确实有更好的基准来测试这类事情。而且有时这些问题也揭示了我们在训练思路上的差距。一个例子是,我们过去认为,‘哦,重要的是在测试时你能往模型里塞多少上下文。’但现在看来,你可以把一堆文件丢进一个容器,模型可以 grep 并搜索它需要的东西。这种通过搜索或工具来确定该使用什么上下文的能力,可能比把所有东西都塞进上下文更高效。如果不尝试并观察它在各种基准上的表现,我们不会真正意识到这一点。我认为这让模型更有用,因为例如,现在模型可以搜索整个代码库,找到你需要的文件,并理解你正在修改的上下文。许多工作场景也是如此,Codex 的用户现在可以上传他们的本地文件系统,你之前可能做过 PowerPoint 或发送过与当前工作相关的 Slack 消息,模型可以通过工具调用搜索这些上下文。所以我们不再受限于能往上下文里塞多少东西,因为模型可以搜索。

Yeah, we definitely have better benchmarks for this sort of thing now. And also sometimes these problems reveal gaps in how we're thinking about training. One example is we used to think, 'Oh, what matters is just how much context you can stuff into the model at test time.' Well, now it seems that you can just dump a bunch of files in a container and the model can grep around and search for what it needs. This ability to have search or tools to figure out what context you should use can be more efficient than just stuffing everything in the context. We wouldn't have really realized that without trying that out and then seeing how it performed on various benchmarks. I think that makes the model a lot more useful because, for example, now the model can search over a whole repo and find the files you need and understand the context of where you're making changes. The same is true for many work contexts where folks in Codex can now upload their local file system, and you might have made PowerPoints before or sent Slacks that are relevant to the work you're doing now, and the model can search over that context with tool calls. So we're not as limited by how much you can literally stuff into context because the model can search.

最爱评估:GDP 与 Houdini Favorite evals: GDP eval and Houdini bench

Host

你有最喜欢的评估吗?

Do you have any favorite evals?

Tejal Patwardhan

我最喜欢的评估?嗯,GDP 评估是我最喜欢的公开评估。

My favorite eval? I mean, GDP eval is my favorite public eval.

Host

好的。

Okay.

Tejal Patwardhan

但我有很多内部评估。我会说出其中一个的名字。它叫 Houdini bench,我不能进一步解释。

But I have many internal evals. I will say the name of one of them. It's called Houdini bench, and I cannot explain further.

Host

哦天哪,你知道我当过魔术师,对吧?

Oh my god, you know I was a magician, right?

Tejal Patwardhan

不知道。

No.

Host

是的,我是。我知道。

Yeah, I am. Yeah, I know.

Tejal Patwardhan

也许我不知道你是否能通过 Houdini bench。

Maybe I don't know if you'd pass Houdini bench.

Host

不,我可能通不过 Houdini bench。那实际上是我早期用一些视觉模型玩的东西之一——用魔术照片来看这个。

No, I probably wouldn't pass Houdini bench. That was actually one of the things I played around with some of the early vision models and stuff—using photographs of magic tricks and seeing this.

Tejal Patwardhan

那太酷了。是的,多模态带来了全新的元素。我记得当 4.0 刚出来时,我们一群人坐在一栋楼的屋顶上,实时语音模型的想法让我们震惊不已。然后我们想,‘我们该怎么评估这个东西?’因为如果有了实时语音交互,在文本、代码和电脑上做事的整个范式就被彻底颠覆了。那次发布中非常有趣的一点是,我们实际上将公开发布推迟了 6 周,因为我们正在想办法确保模型安全。

That's very cool. Yeah, multimodal brings a whole new element. I remember when 4.0 had first come out, there was a group of us sitting on the roof of this building, and our minds were just so blown by the idea of a real-time voice model. Then we were like, 'How do we even eval this thing?' Because the whole paradigm of doing things in text and code and on your computer is just completely blown away if there's a voice interaction in real time. Something really interesting about that launch is we actually delayed the public launch by 6 weeks as we were figuring out how to make sure the model was safe.

Host

4.0?

4.0?

Tejal Patwardhan

是的,因为实际上那是在选举之前。有很多担忧:如果模型能用逼真的声音实时与你交谈,它会不会被用于说服性宣传之类的事情?公司推迟发布以确保我们能建立所有这些测试和缓解措施,防止模型被用于这类事情,这非常酷。

Yeah, because this was before the elections actually. There was a lot of worry: if the model can in real time talk to you with a realistic sounding voice, could this be used for persuasive propaganda or this sort of thing? It was very cool that the company delayed the launch to make sure we could build out all these tests and build mitigations to make sure the models couldn't be used for this sort of thing.

多模态评估的挑战 Challenges of multimodal evals

Host

嗯,随着这些模型变得多模态,这似乎是一个非常复杂的因素。我记得早期 GPT-4 有视觉能力时,我字写得很烂,我写一个提示,它突然就能解决。然后你意识到,哦,这不是文本提示,而是视觉提示。再到音频模型,当你进行音频输入输出时,模型可以模仿东西,以各种不同的方式做事。所以这似乎真的——你从哪里开始着手衡量它呢?

Well, it seems like that's a very complicating factor as these models become multimodal. I remember early on with GPT-4 with it being on GPT-4 vision back when it was that you could—I had terrible handwriting. I could write a prompt and all of a sudden it would solve for this. And you realize, oh, it's not a text prompt, it's a visual prompt. Then with the audio models when you're doing audio in, audio out, the model could emulate things and do stuff in such different ways. So it seems like that's really—where do you even begin trying to figure out how you're going to measure that?

Tejal Patwardhan

是的,这只是大量工作。通常对于这些,我们都是从人类在这种情况下会怎么做开始?所以你会有一组输入给模型,以及一组要评估的输出。然后你可以逐步推进:我们能自动化其中一些吗?我们能构建一个新平台来大规模衡量这类事情吗?然后从那里开始。

Yeah, it's just a lot of work. Usually for any of these, we start with what would humans do in this case? So you would have a set of inputs that you put into the model and a set of outputs you would evaluate. Then you can build up: can we automate some of these? Can we build a new platform to measure this sort of thing at scale? And sort of move from there.

原生多模态模型与安全挑战 Challenges with natively multimodal models and safety

Tejal Patwardhan

但对于一些原生多模态模型,你必须拆掉大量基础设施才能让一切运转。Sora 也是如此。我们想确保视频不会过于逼真或被用于不当用途。这需要构建一整套全新的评估和缓解措施,包括模型层面的拒绝机制以及生产环境中的监控。这需要全新的思考方式。

But for some of the natively multimodal models, you have to rip apart a bunch of your infrastructure and make things work. This was also true with Sora. We were interested in making sure the videos weren't overly realistic or could be used for the wrong thing. That required building a whole new stack of evals and mitigations, including refusals at the model level and monitoring when this was being used in production. It requires a whole new stack of thinking.

Host

是的,这也是问题所在。当你开始思考如何优先考虑一个评估而不是另一个时,什么时候决定这个不行,或者只是觉得这个饱和了就继续前进?即使你并不试图优化某些公开基准,你仍然需要弄清楚现在什么对我们重要。曾有一段时间 OpenAI 在代码方面领先,然后有一段时间不是。现在又领先了,但中间有过一段黑暗期。

Yeah, that's the thing too. When you start to think about how do you prioritize one eval over another? When do you decide that this isn't it, or do you just go like this one's saturated, we move on? Even though you may not be trying to optimize towards certain public benchmarks, you still have to figure out what's important to us now. There was a time when OpenAI was leading in code and then there was a time when it wasn't. Now it is again, but there was a dark period.

Tejal Patwardhan

是的,我们尽量不被公开基准分散太多注意力,因为它们可能很嘈杂。内部我们有一个叫做 AGI 指数的东西,灵感来自 CPI 或通货膨胀,即有一个加权商品篮子并追踪其价格。对我们来说,这是一个评估篮子,包含我们关心的所有核心领域的测量:对齐、安全、能力。这就是你对模型的要求。我们不断更新这个指数,以代表我们希望模型完成的更困难版本。我们在内部追踪它,尽量不被公开基准干扰。更重要的是,我们在不同领域——科学、工作、安全、对齐——有一个评估组合,并确保我们在那个加权篮子上持续进步。我们努力保持专注。

Yeah, we try not to get too distracted by public benchmarks because they can be noisy. Internally, we have something called the AGI index, inspired by CPI or inflation, where you have a weighted basket of goods and track their price. For us, it's a basket of evals that include measurements across all core areas we care about: alignment, safety, capabilities. It's just what you want from your model. We keep updating that index to represent more difficult versions of what we want our models to do. We track that internally and try not to be distracted by public benchmarks. It's more about having a blend of evals across different domains we care about—science, work, safety, alignment—and making sure we keep making progress on that weighted basket. We try to stay focused.

科学前沿评估的演变 Evolution of evals in scientific frontier

Host

我们见证了评估和模型的演变。我和从事科学工作的人聊过——不仅仅是计算机科学研究者,还有生物学、数学领域的人。你能告诉我科学前沿的评估进展如何吗?似乎我们正处在一个即将看到有意义成果的阶段。

We've watched this evolution of evals and models. I've talked to people working in the sciences—not just computer science researchers, but people in biology, mathematics. Can you tell me what's going on with the evals on the scientific frontier? It seems like we're at a point where we're going to see meaningful results.

Tejal Patwardhan

是的,我认为我们在一些科学评估上的工作是最令人兴奋的。过去几个月,我们公开了几个层次的评估。第一层叫做 Frontier Science Olympiad,相当于我们之前有的数学奥林匹克式评估,衡量模型在生物、化学、物理的高中奥林匹克级别问题上的表现。这些是简答题,但仍然很难,模型当时还不太好。下一阶段是 Frontier Science Research,也是公开的,衡量模型帮助完成未发表的生物、化学、物理论文的能力。我们让这些领域的博士或教授提供未发表的文本,比如论文的一部分,然后将其转化为评估:给模型一些输入数据或初始起点,它必须完成论文的其余部分,并根据评分标准评判。这开始衡量模型是否开始做研究、使用工具等。最后一个迭代是看模型在现实世界的湿实验室中表现如何。我们与 Ginkgo Bioworks 合作,他们有自动化的湿实验室机器人。模型必须优化一个蛋白质合成方案。模型生成方案,然后他们自动在湿实验室中测试,加入模型建议的试剂,观察蛋白质产量。这是针对一种与卵巢癌药物相关的蛋白质,是一个玩具场景。我们一开始很紧张,因为人类基线很难,但我们不应该低估模型。每个周期都变得更好,超过了人类基线,并设立了成本效益的最新标准。我认为这只是开始。如果我们给这些模型优化问题——比如如何让疫苗更便宜,或合成对药物重要的蛋白质——模型可以不断用真实世界输入优化方案。这是我们第一次降低与现实世界相连的评估的风险。我们不是在等代码运行,而是在等机器人完成实验以记录合成了多少蛋白质。模型将为我们做很多科学工作。这将非常有趣。

Yeah, I think the work in some of our science evals is some of our most exciting. In the past few months, there have been a few tiers of evals we've made public. The first tier was called Frontier Science Olympiad, equivalent to the math Olympiad style evals we had before, measuring how well models could do on high school Olympiad style problems in biology, chemistry, and physics. They were shorter answer but still quite hard, and the models weren't very good yet. The next phase was Frontier Science Research, also public, which measured how well models could help complete unfinished biology, chemistry, and physics theses. We had PhDs or professors in these fields with unpublished text, like part of their thesis, and turned that into an evaluation where the model was given input data or an initial starting point and had to fill out the rest of the paper, judged against a rubric. That starts to measure whether models are starting to do research, use tools, etc. One of the final iterations was to see how well the model could do in the real world in a wet lab. We worked with Ginkgo Bioworks, which has automated wet lab robots. The model had to optimize a protocol for protein synthesis. The model would generate a protocol, and they would automatically test it in the wet lab, putting in the reagents the model suggested and seeing the protein yield. This was for a protein related to an ovarian cancer drug, a toy scenario. We were nervous at first because the human baseline was hard, but we shouldn't underestimate the models. Every cycle got better, beat the human baseline, and set the state of the art on cost per yield. I think that's just the start. If we give these models optimization problems—like how inexpensive you can make a vaccine or synthesize a protein important for a drug—the model can keep optimizing protocols with real-world inputs. It was one of our first times de-risking an eval connected to the real world. We weren't waiting for code to run; we were waiting for the robot to finish the experiment to record how much protein was synthesized. The models are going to do so much science for us. It's going to be really interesting.

Host

那很令人兴奋,因为那是用 GPT-5,而且它没有经过任何“如何成为科学家”的训练。自那以后,这些模型进步了很多,有了更多现实世界的经验。

That was exciting because that was with GPT-5, and it hadn't gone through any sort of 'how to be a scientist' training. And now these models have progressed a lot since then, with more real-world experience.

Tejal Patwardhan

是的,那甚至不是我们最好的模型之一。只是一个早期的推理模型。我认为所有这些因素叠加:我们将有更好的预训练、更好的强化学习和后训练,我们会在测试时更好地使用这些模型来真正激发它们的能力。我认为下一代评估真正关乎我们如何让这些模型在现实世界中采取行动,解决那些人类需要很长时间才能解决的未解问题——一些我们未能投入足够精力的科学问题。现在我们有了所有这些智能体,它们可以花费算力为我们解决问题,我们试图引导它们朝着有用的方向。

Yeah, that wasn't even with one of our best models. It was just an early reasoning model. I think all these things stack: we'll have better pre-training, better RL and post-training, and we'll get a lot better at using these models at test time to really elicit their capabilities. I think the next generation of evals is really about how we can have these models take actions in the real world and solve unsolved problems for us that would take humans a long time—some of these scientific problems we haven't been able to put enough effort against. Now we have all these agents that can spend compute to solve problems for us, and we try to steer them towards what would be useful.

评估的复杂性 Complexity of evals

Host

你认为评估会变得更加复杂吗?

Do you think that evals are going to get a lot more complex?

Tejal Patwardhan

是的,我们团队有句话叫“痛苦就是护城河”。我真的认为物理世界中的许多操作将成为衡量模型能力的瓶颈,因为即使从数字领域开始,我们也需要做大量的基础设施和支撑工作来运行这些评估。现在如果你想测试 Codex 的表现,模型需要调用 API、在你的电脑和浏览器中执行操作、为你生成工件、编写并运行代码。衡量这个模型要复杂得多,而这还只是数字领域。如果你想衡量模型如何与物理世界交互,就需要各种运营和物流流程来确保顺畅,从而大规模部署这些东西。我认为很多工作实际上正在从理论、数学甚至编程转向。我觉得人们编程不多,他们只是问 Codex。工作更多转向规划、运营、物理事务,至少我的工作已经大幅转向这个方向。这些事情非常困难。在角落里写点东西其实很容易,但要管理所有这些运营和物流就难多了。

Yeah, I mean we have the saying on our team that pain is the moat. I really think a lot of operations in the physical world will become part of the bottlenecks in being able to measure what the models can do because even just starting with digital, there's so much more scaffolding and infrastructure work we need to do to run these. Now if you want to test how well Codex does, it's like the model is calling APIs, taking actions on your computer and in your browser, making artifacts for you, writing and running and executing code. It's just so much more complex to measure that model, and that's only digital. Now if you want to measure how the model could interact with the physical world, there's all sorts of ops and logistics that you need to have a really smooth process for to see how you can deploy these things at scale. And yeah, I think a lot of the work is actually shifting from being theory or math or even programming. Like I feel like people don't program that much; they just ask Codex. It's more shifting towards planning, operations, physical stuff, or at least my job has shifted a lot that way. And those things are very hard. It's actually kind of easy to just write something in a corner. It's a lot harder when you have to manage all of these operations and logistics.

Host

这很令人兴奋,但挑战的一部分在于这些不再是简单的评估了。它们需要更多的算力和时间。当你尝试进行长期评估时,时间很长,你必须等待很长时间才能得到结果。

It's exciting, but it seems like part of the challenge is these aren't just simple evals anymore. They take more compute, they take more time. When you're trying to do a long horizon eval, you know, it's long. You have to wait a long time to get the outcome on that.

Tejal Patwardhan

是的,当然。所以,设计评估并大规模运行它们的工作量更大,而且如果工作耗时更长,我们获得信号的速度就会变慢。因此,我们必须更多地投资于缩放定律,这样我们就可以预测:如果一天后模型看起来是这样,那么我们可以预测七天后它会是什么样子,并找出趋势,从而更快地获得信号。否则,我们只能干等一周才能得到更新,这不是最有效的时间利用方式。

Yeah, definitely. So, it's both a lot more work to come up with the evals and run them at scale, and also if the work takes a longer amount of time, we don't get the signal as fast. So, we have to invest more in scaling laws where we can predict, okay, well, if by one day the model looks like this, then we can forecast that at 7 days it would look like this and sort of come up with trends that we can so that we can get signal faster. Otherwise, we're just stuck there waiting for a week to get an update, which is not the most productive way to spend time.

个人基准建议 Advice on personal benchmarks

Host

我有一些基准测试,每次新模型发布时我都会用来测试它对我个人的实用性。我告诉那些经营企业或其他事情的人,要考虑自己的评估,那些能告诉你进展的东西。因为有时人们可能六个月前试过 ChatGPT,然后说“啊,它不好,它做不到这个。”他们没有意识到事情变化有多快。你对人们如何想出基准有什么建议吗?

I have certain benchmarks and things I used to test every time a new model comes out to find out how it's personally useful to me. And it's one of the things I tell people who run businesses or other things is think about your own evals, things that will tell you where something is because sometimes people might try something like they might try ChatGPT 6 months ago and go like, 'Ah, it wasn't good. It didn't do this.' They don't realize how fast things move. Do you have any advice for people on how to figure out how to come up with a benchmark?

Tejal Patwardhan

是的,如果事情发展得非常快,每隔几周就会发生变化,我觉得人们对此并不那么警觉。在我的工作中,我是世界上最早看到一些最强大模型的人之一,所以我充满了 AGI 意识,我认为进步正在以更快的速度发生。

Yeah, I mean, if things move really fast, things change every couple of weeks and I feel like people are not as awake about it. In my job, I'm one of the first people in the world to see some of the most powerful models, so I'm extremely AGI-filled and I think progress is happening a lot faster.

Host

我看到了什么?

What have I seen?

Tejal Patwardhan

我见过好模型,伙计。是的,但进步比人们想象的要快得多。我认为最好的评估,老实说,就是亲自试用或使用模型。人们应该尽可能多地使用模型。即使他们认为模型某一周做得不好,他们也应该下周再试一次。它可能就奏效了。

I've seen good models, man. Yeah, but progress is happening a lot faster than people would think. And I think the best eval, honestly, is just to dog food or use the model. Like people should just try to use the models as much as they can. And even if there are things that they think the model didn't do well on one week, they should just try it again the next week. It'll probably work.

Host

我认为对于 AI 领域之外的人来说,应该显而易见的一点是,真正优秀的前沿 AI 公司如何在内部使用这些工具,这就是事情加速并变得更有能力的原因。

I think that's one of the things that should be obvious to people kind of outside of AI is how really good frontier AI companies are using these tools internally, and that's why things are speeding up and getting more capable.

Tejal Patwardhan

是的,我基本上让模型对我做的所有事情进行初稿。无论是发送 Slack 消息、理解下一步要做什么实验、任何管理事务、运营、物流,都让模型先过一遍。然后如果模型不好,我们就想办法把它纳入评估。

Yeah, I basically try to have the model take a first pass of everything that I do. Like whether it's sending a Slack message, understanding what experiment to perform next, any management stuff, ops, logistics, like you have the model take a first pass. And then if the model's not good, we figure out how to put that in the eval.

Host

我对计算机使用评估感到兴奋。看着 Codex 在计算机使用方面的表现,比八个月前简直是天壤之别。而且这些东西似乎只会变得更快更好。我的预测是,到今年年底,它使用我的电脑会比我自己更好更快。

I'm excited about the computer using evals. Like just watching the performance of Codex with computer use is just light-years over where it was just, you know, maybe eight months ago. And it seems like those things are just going to get faster and better. My prediction's like probably by the end of the year it'll use my computer better and faster than I do.

Tejal Patwardhan

是的,我也这么认为。模型相比你有一些优势,对吧?它们可以调用连接器或插件,这比你在电脑上点击进入服务、理解每个页面、然后来回复制数据要快得多。甚至编写一个服务来调用那个 API 或 MCP 之类的,对人类来说比模型更费劲。所以模型有这个优势。而且模型可以更快,如果它经过训练,可以通过无障碍树或代码来导航浏览器或桌面。所以模型比我们有优势。我认为很长一段时间里,真的没有非常有效的产品部署。我们之前推出了 Operator 和 ChatGPT 智能体,它们对于展示这种可能性非常有用,但那些模型的延迟太高了。它们超级慢。我认为人们还没有大规模使用它们,但我们现在已经达到了一个临界点,比如让模型帮我读 Slack、安排一堆日历邀请并优化房间,比我亲自做要快。我认为人们还没有准备好。而且很多人还没有尝试过这些东西,因为它们都是最近才推出的,但每个人都应该去获取计算机使用插件并使用它们,安装所有插件和所有好的连接器,这样事情会更快。然后你会被震撼到。

Yeah. Yes, I think so. The models have some advantages over you, right? Like they can call a connector or plug-in, which is a much faster mode of communication than you on your computer having to go click into a service and understand every page and then copy some data back and forth. Or even writing some service to call that API or MCP or whatever. It's more work for the human than it is for the model. So the model has that advantage. And the models can just be faster and if it's trained to navigate a browser or a desktop through whether it's through accessibility tree or through code. So the models have an advantage over us. And I think for a long time there was really no product deployment that was very effective. Yet we launched operator and ChatGPT agent a while ago, and those were really useful for showing like this could be possible, but the latency on those models was just too high. Like they were just super slow. And I don't think people use them at super high scale yet, but we've now reached sort of a tipping point where doing things like asking the model to read my Slack for me or go schedule a bunch of calendar invites and optimize the rooms is faster for me than it would have been to do it myself. And I think yeah, people are not ready. Also, a lot of people haven't tried this stuff out because it's all launched so recently, but everyone should go get the computer use plugins and use those and install all of the plugins and all the good connectors that will make things faster. Then you'll be mind-blown.

前沿评估团队目标 Frontier evals team goals

Host

我们来谈谈前沿评估。

Let's talk about frontier evals.

Tejal Patwardhan

是的。前沿评估团队的目标是衡量和预测 OpenAI 前沿模型的进展,以更好地了解我们现在的位置、未来的方向,并尝试与世界分享这些信息。我认为团队努力做的事情之一就是尽可能多地发布和开源。

Yeah. So, the goal of the frontier evals team is really to measure and forecast progress of the frontier models at OpenAI to better understand where we are, where we're going, and sort of try to share that with the world. And one of the things I think the team has tried to do is to help publish and open source as much as we can.

评估基准及其目的 Evaluation benchmarks and their purpose

Host

我们开源的一些评估包括 SweetBench Verified,用于衡量编码进展;MLE Bench,用于衡量模型训练其他模型的能力,追踪机器学习工程技能的进步;PaperBench,用于衡量模型复现 ICML 或 ICLR 等顶级机器学习论文的能力;还有 GDP eval,用于衡量模型在超过 40 种职业的真实任务中的表现。这些评估的目标是,虽然模型现在可能看起来不怎么样,但如果你绘制出每代模型结果的提升曲线,就会发现人们常说“哦,我觉得这需要一年左右”,但他们往往高估了基准测试饱和所需的时间。就连我自己或团队成员的预测也常常不够激进,低估了变化的速度。所以我们想尽一份力,让世界了解什么是可能的。我认为一些研究加速评估特别有趣。比如我们最初有一个 OpenAI 研究面试评估,就是把我们面试申请人时问的研究问题放进评估里。模型很快就通过了,现在绝对能通过我们的面试。这引发了一系列下游问题,比如如何确保人们不在面试中作弊,以及如何真正衡量研究人才。但我认为所有这些都很有用,因为衡量内部进展就像是衡量模型加速改进的杠杆,也就是改进曲线的斜率加速。总之,有办法衡量模型进展本身就是有用的信息。

So, you know, some evals that we've helped open source include like SweetBench Verified, which helped measure progress on coding, MLE Bench, which was a way to measure how well models could train other models and sort of track the progress of machine learning engineering skills in our models, um PaperBench, which was a way to measure how well models could replicate real top machine learning papers um from like ICML or ICLR, um and GDP eval, which you know, helped measure how well models could perform on real-world tasks across, you know, over 40 occupations. And the goal for all of these has been you know, the models might not seem good now, but if you just plot how they increase with each, you know, the the results that improve with each model generation, often when people say like, "Oh, well, I expect this will take like a year or whatever." They like over >> Mhm. >> they over um ex- expect in terms of how much time it will take to saturate a benchmark. And like even my own or people on my team's predictions are often like not ambitious enough for how fast things will change. And so I just think we're trying to do our service in helping inform the world about um what is possible. I think some of these research acceleration evals in particular are quite um interesting. Like when we first started, we had this eval called the OpenAI research interview eval, which was just taking the researcher questions that we asked people applying to OpenAI and putting those in an eval. And the model blasted through that like pretty pretty quickly. It's like definitely can pass our interviews right now. um Which I think has caused a whole other slew of downstream questions on like how do we make sure people don't cheat on the interviews and like how do we actually measure research talent? um But I think all of this is very useful because um measuring internal progress is it's like kind of a way to measure the lever by which the models will keep getting better faster. Like sort of the acceleration of the slope of improvement, so to speak. And um yeah, I think having ways to measure model progress is is just good information.

Host

我听说有些评估存在了一段时间,结果发现题目本身有错误。一些公开的评估实际上无法达到某个分数以上,如果达到了,那是因为你在训练数据上见过。人们仔细一看才发现,“哦,这个答案其实不对。”

I've heard that in some of the evals that were out there for a while, that it turned out that there were actually errors in the questions. That that was an issue with some of the evals. That that was some of the publicly available ones were actually you couldn't score above a certain level. And if you did, it was actually because you were training on the data. And people looked at that and found out like, "Oh, there's actually this is not the right answer."

Tejal Patwardhan

是的,我认为很多公共基准测试都有这个问题。最初我们做 SWE bench verified 的原因就是,我们想运行 SWE bench,但一半的问题要么有缺陷,要么定义不明确。而业界的人还在用这个作为衡量标准发表结果。我们就想,至少应该试着修复它,然后分享出来,这样我们就有更好的标尺了。我认为公共基准测试之所以没有经过充分的实战检验,是因为它们往往来自学术实验室,有人有个好主意想写论文,但从未大规模运行过这些评估,比如生产级别的训练运行或发布前的评估扫描。当你大规模运行时,它就会出问题或崩溃,你就能发现所有 bug。所以我认为,靠近产品、在实验室里工作是一种强制机制,能确保测量质量非常高,因为我们不是为了论文好看,而是为了系统必须在大规模下正常工作,这迫使质量必须高。

Yeah, this is a problem with a lot of public benchmarks, I think. Like so the original reason for SWE bench verified was because we wanted to run SWE bench and it was half the problems were like either broken or under specified. And you know, people in the industry were publishing results on this as some metric of how well you did. And we were like, "Well, we should at least try to fix it." And then like share that so we can have a better yardstick. um but I think one of the reasons that public benchmarks maybe aren't always as you know battle tested as we'd like is that not they they tend to be like you know someone in in a lab like an academic lab like had a good idea and like wanted to write a paper but they never had to run that eval at scale and like production training run or production like level eval sweep for a launch and just when you run some of this stuff at scale it like breaks or falls over and you like catch all of these bugs and so I kind of think sitting in a lab and being closer to product is a forcing function for making sure the quality of your measurements is really high because like we're not doing this like look good in a paper we're like doing this like it has to work because it has to work for our systems at scale so it kind of forces the quality to be high.

Host

似乎有一种情况是,这些模型变得非常强大,有时能解决问题,但会走捷径,直接给出记忆中的答案而不是真正求解。比如数单词或字母数量,如果你提示得当,模型能答对,但提示不对的话,它就直接抛出一个答案。

And it it seems like kind of one of the things that can happen is these models become incredibly capable sometimes they're very good at sometimes they can solve a problem but they'll take sort of the laziest path and kind of they can they can give you the memorized answer instead of solving it and we saw that with like counting in like how many words are in a how many letters in a character or a word or whatever and it was often the model if you prompt it right it would get the answer right but if you didn't prompt it the right way it would just sort of throw you an answer.

Tejal Patwardhan

这引出了很多有趣的概念。一个是记忆化,即模型确实知道答案,不需要真正思考或推理,只是重复已知的东西。这使得测量不太有用,因为你只是在衡量模型是否在训练数据中见过大量相关数据,而不是是否学到了你试图衡量的技能或能力。避免这一点的方法是保持数据干净和严格,不包含任何你想测量的基准测试或评估。这有助于解决你提到的第一个问题。另一个问题是模型可能奖励黑客行为或作弊来解决评估,这需要干净的评估设计,大规模测试看看是否有漏洞,确保测试环境没有模型可以利用的漏洞。这需要大量的质量控制,确保评估不容易被黑客攻击。

Yeah that brings up all sorts of interesting concepts I mean so there's this one concept of memorization which is the idea that the model literally knows the answer and doesn't have to really think or reason to solve it's just like regurgitating something it already knows and that makes the measurement not super useful because you're just measuring whether you happen to have trained on trained on that data a ton versus whether the model learned the skill that you or tool or capability you were trying to measure so that's one way to avoid that is to try to be really clean and disciplined about your data not including any benchmarks or any evals that you want to measure and that helps solve sort of the first problem that you laid out so so that that that's one thing and then there's there's this other thing where like the model can kind of like reward hack or sometimes like cheat to solve any eval and that's very much a question of having clean eval design where you like sort of test these at scale see if there's any hacks, make sure those environments that you're testing don't have the hacks as something that's possible for the model to do. And that just requires a lot of quality control to make sure like the eval is not overly hackable. uh yeah.

Host

是的,因为有些很简单的问题,比如小学数学,如果你稍微改一下,早期的一些模型就会混淆,给出错误答案,而它其实有能力解决。它只是说“哦,这个我会”,然后就错了。就像“我该开车去洗车吗?”这种问题,模型也会出错。

Yeah, cuz it seems like there were some very simple ones like grade school math and whatnot that models if you just change it a little bit, some of the early models would get confused and give you the wrong answer that it was actually capable of solving it. But it just goes, "Oh, I this one I got it." And then, you know, that's happened to like you know, should I drive my car to the car wash? You have a problem.

Tejal Patwardhan

是的,模型可能会被欺骗。对我来说,如果模型没做好,它应该更聪明。我们应该让模型对欺骗更鲁棒。但这与能力激发或最佳测量方式有关,这对安全测试尤其重要。例如,如果你想衡量模型发现漏洞或进行网络安全任务的能力,你需要确保模型不是被问题欺骗了,而是真正测量了其真实能力。

Yeah, yeah, yeah. So like the models can get tricked. To me like the model doesn't like if it didn't get a good do well on that like it it should have been smarter. Like we we should also like have the models be a bit more robust to being tricked. But this also um relates to this idea of capability elicitation or like trying to measure the models in the best way, which is especially important for our safety testing. Like for example, if you want to measure how well the model can, you know, find vulnerabilities or um you know, do some of the cybersecurity stuff. You want to make sure the model's not just getting tricked by the problem. Like that you really measured the true capability.

评估准备与人工监督 Eval preparation and human oversight

Tejal Patwardhan

因此我们做了很多提示调整、更改测试框架,有时甚至进行微调,让模型为应对该挑战做好最充分的准备。我们这样做是为了确保,如果我们说‘哦,模型在某些非常危险的能力上表现不佳’,我们在下此结论之前能更有把握。

And so there's a lot of prompt tuning, changing the harness, and sometimes even fine-tuning to get the model maximally ready to solve that challenge. We do that to make sure that if we say, 'Oh, the model's not good at some very risky capability,' we can be a bit more sure before we say that.

Host

我小时候很喜欢读《百科全书布朗》系列故事,那些小谜题需要你去破解。用 GPT-4 时,我会为它编写定制谜题,以防有人已经把答案泄露给它。但那做起来很麻烦。现在想到可以让模型自己编写或提出新的评估,这很令人兴奋。那么目前模型在这方面有多大帮助?

When I was a kid, I loved reading these Encyclopedia Brown stories, these little mysteries and you had to solve them. And with GPT-4, I would write custom ones for it just in case somebody had tipped all these answers to it out there. But that was a pain to do. And it's exciting to think now I can have a model write something or come up with some new eval. So how helpful have the models been now for that?

Tejal Patwardhan

嗯,它们还算有用。我认为我们正处于模型开发的这个阶段,有时输出仍然有些粗糙。它们需要人工质量控制或监督,以确保质量仍然很高,并且我们不会被误导。所以,我想说人们有时会惊讶于我们在评估中仍然有大量的人工干预和参与,只是因为评估的质量要求可能比训练数据更高,你需要确保测试的每一个数据点都非常高质量。因此,这是人工介入可以发挥很好作用的领域之一。

Yeah, they're semi-useful. I think we're in this phase of model development where sometimes the outputs are still kind of sloppy. They require human QC or oversight to make sure the quality is still high and we're not getting tricked. So, I would say people are sometimes surprised that we still have a lot of human intervention and involvement in the evals, just because evals can be a lower end than training data and you want to make sure every single data point you're testing is very high quality. So, this is one of the areas where a human touch can be quite nice.

对就业与生产力的影响 Impact on jobs and productivity

Host

我们看到一些有趣的趋势,实际接触人工智能的工作似乎更受欢迎,因为它提高了人们的生产力。你是如何追踪这一点的?你如何寻找你认为会产生影响的领域?

We're seeing some interesting trends where jobs that actually touch AI seem to be more in demand because it's made people more productive. How are you tracking this? How do you look for areas where you think this is going to have an impact?

Tejal Patwardhan

嗯,这些都是非常困难的问题。我认为人们还没有校准我们的模型能够完成多少工作,以及在不同工作中能有多快。目前模型仍然主要擅长任务而非工作。工作比任务包含更多内容。你必须弄清楚要做什么,应对模糊性,与同事协作沟通。然后你才可能确定要做什么任务,再交给模型。这就是我们目前所处的阶段。即使在我的工作中,模型也在为我处理单个任务,但我仍然在做大量的思考和规划。我认为人们甚至没有校准到这一点。我觉得软件和研究领域的人校准得更多——我所说的校准是指意识到模型有多强大——相比之下,我在其他行业的一些朋友则不然。我希望人们能多尝试模型并亲眼看看,因为那些先尝试并看到的人会真正理解。但我也认为模型在某个时候将开始能够做委派工作。也许不会太久。弄清楚要做什么,应对模糊性,编写规范让模型执行。人们真的应该开始思考,在最大程度由 AGI 构建的世界里,即使是数字工作,模型也能提出要做什么、执行、与现实世界互动。现在已经有整个企业,比如独角兽公司,主要靠 AI 和少数员工驱动所有价值。所以我确实认为存在一个问题:我们是否意识到这会有多大?

Yeah, these are very difficult questions. I think people are not calibrated to how much work our models will be able to do and how quickly across a wide variety of jobs. Right now the models are still mostly just good at tasks versus a job. There's a lot to a job than a task. You have to figure out what you want to work on, navigate ambiguity, collaborate with co-workers, communicate. Then you might figure out what task you want to do and give that to a model. That's the phase we're at now. Even in my job, the model is doing individual tasks for me, but I'm still doing a lot of the thinking and planning. I think people aren't even calibrated to that. I feel like people in software and research are a lot more calibrated—by calibrated I mean realize how capable the models are—compared to some of my friends in other industries. I wish people just tried the models more and saw, because the people who try and see first will start to really get it. But I also think the models are going to start to be able to do the delegating part at some point, too. Maybe not too far from now. Figuring out what to work on, navigating ambiguity, writing the spec that the model then executes on. People should really start to think about what happened in the maximally AGI-built world where even for digital work, the model can come up with what to do, do it, execute it, interact with the real world. There are entire businesses now, like unicorns, that were mostly AI and a few employees driving all the value. So I do think there's this question of whether we are realizing how big this will be.

Host

就我个人而言,我认为机会空间正在变大。我认识的每一个人,那些最懂 AGI 的人,那些一直使用 Codex 等工具的人,现在做的事情更多了。他们现在更高效了,因为随着 AI 更好地处理某些工作,他们不必再亲自做那些任务和工作。就像,太棒了,现在有五项工作需要我去做,因为我能做更多了。我认为我们思考的潜力光锥比我们想象的更大。而且我认为这些工具只是帮助我们更快到达那里,而不是缩小它。

Personally, I think the opportunity space is getting bigger. Everybody I know, the most AGI-built people I know, the people who are using tools like Codex all the time are doing way more now. They're more productive now because they don't have to do the tasks and the jobs as the AI gets better to handle certain jobs. Like cool, there are five jobs I need done now because I can do more. And I think that we just think about the light cone of the potential where we can be is bigger than we can imagine. And I think these tools just help us get there faster, not narrow it.

Tejal Patwardhan

我认为可能是多种因素的混合。

I think there is probably some mix of things.

Host

是的。

Yeah.

Tejal Patwardhan

即使模型能加快文书工作,比如想想药物的临床试验。人们花几个月时间整理数百页材料,说明为什么他们应该能够进行试验。他们提交给 FDA。有 35%的几率被拒绝,因为他们犯了错误或遗漏了某些内容。他们修改。然后最终才能进行试验。这些流程很好,但耗时很长。然后试验涉及记录症状、长期跟踪以及进行数据分析。其中很多只是文档工作或数据分析,非常典型的数字工作。我认为如果模型能加速所有这些环节,在健康、能源、制造、政策研究、教育等领域,这将非常具有加速作用。我们将拥有更快、更便宜、更好的商品。这对人们非常有利。对个体消费者非常有利。所以我认为这是人们应该感到兴奋的事情。但我们应该非常深思熟虑,以周到且负责任的方式引导向那个世界的过渡。

Even if you have models that can speed up paperwork, like think about a clinical trial for a drug. People spend months putting together hundreds of pages of why they should be able to do the trial. They submit it to the FDA. There's a 35% chance it got rejected because they made a mistake or forgot something. They revise. Then finally you can do the trial. These processes are good, but it just takes a long time. Then the trial involves documenting symptoms, tracking for a long time, and doing data analysis. A lot of this is just documentation or data analysis, very classically digital work. I think if models can help accelerate all parts of this, for health, energy, manufacturing, policy research, education, this will be very accelerative. We will have faster, cheaper, better goods. That's really good for people. It's very good for the individual consumer. So I think that is something people should be excited about. But we should be very thoughtful about how to navigate the transition to that world in a way that's thoughtful and responsible.

结束语 Closing remarks

Host

太好了。谢谢你,Chiejina。

Excellent. Thank you, Chiejina.

Tejal Patwardhan

谢谢你邀请我。

Thank you for having me.

互动版:逐字朗读 + 针对本期提问 →