Evaluating AI Capabilities and Risks: A Conversation with Beth and David
打开互动全文版(中英对照 + 朗读 + 问答)→Beth 和 David 讨论评估 AI 系统的挑战,包括可扩展监督、模型对齐,以及更好地理解 AI 能力和风险的必要性。
Beth and David discuss the challenges of evaluating AI systems, including scalable oversight, model alignment, and the need for better understanding of AI capabilities and risks.
是的,嗯,我非常兴奋能和你 Tim 讨论时间范围图和 Meter。
Yeah. So, um yeah, I'm super excited to talk to you Tim about the time horizon graph and Meter.
是的。我认为世界对 AI 正在发生的事情缺乏良好理解,而且我认为应该更好地理解。我认为这很有可能让我们的生活变得更好或更糟,人们甚至对当前模型能做什么都存在分歧,更不用说我们前进的方向了。所以,在 Meter,我们试图让世界更好地了解 AI 能力、风险和预测。我们对此有多个不同的研究角度,既有悲观也有乐观,或正面和负面的能力估计。很高兴能讨论这个。
Yeah. I think the world does not have a good understanding of what is happening with AI and I think it should have a better understanding. I think there's a good chance that this makes our lives a lot better or a lot worse and people disagree even about what current models can do, let alone where we're heading. So, at Meter, we're trying to give the world a kind of better understanding of what is up with AI capabilities and risks and forecasts. We have a bunch of different research angles on this, both on the pessimistic and optimistic or positive and negative estimations of capabilities. And excited to talk about that.
我非常高兴能邀请到你们两位。你们都有着令人难以置信的出色背景。Beth,你曾是 OpenAI 的对齐研究员,2022 年与 Paul Cristiano 共同创立了 Ary Vows,并在 2023 年 12 月将其分拆为 Meter。你入选了《时代》杂志 AI 领域百大人物。David,你是 GPQA(研究生级谷歌证明 QA 基准)的创建者,该基准被所有主要 AI 实验室用作能力基准。你也是 Hcast(我们今天会讨论)、时间范围论文和开发者生产力 RCT 的合著者。能请到你们两位真是太好了。也许我们应该先问两位一个问题。Beth,你离开 OpenAI 去创建 Meter。是什么时刻让你们各自意识到现有的评估方法根本不够好?
I'm so excited about having you both on. So, you both have incredibly impressive backgrounds. So, Beth, you were an ex OpenAI alignment researcher and you started Ary Vows in 2022 with Paul Cristiano and you spun out that out as Meter in December 2023. You've been featured on the Times top 100 AI profiles. And David, you're the creator of the GPQA, the graduate level Google proof QA benchmark, which is used by every single major AI lab as a capability benchmark. And you're the co-author on Hcast, which we'll talk about today, and the time horizons paper and the developer productivity RCT. Incredible to have you both in here, but maybe we should just start as a bit of a question to both of you. So, Beth, you left OpenAI to build Meter. What was the moment that each of you realized that existing evaluation approaches were fundamentally not good enough?
对我来说,主要是思考可扩展监督这个问题。随着模型能力越来越强,评估它们的能力就变得越来越难。如果我们设想模型能够完成需要人们很长时间才能完成的任务,或者需要你不一定具备的专业知识,那么你需要一种方法仍然对其输出有信心并信任其输出。所以思考这个问题实际上是 GPQA 的主要动机,也是我开始思考评估的起点。
For me, it was mostly thinking about this problem of scalable oversight. As models get more capable, it just gets harder to evaluate their capabilities. If we imagine that models are able to complete tasks that take people a long time to complete or require expertise that you don't necessarily have, you need a method for still being confident in their outputs and trusting their outputs. And so thinking about that problem was a lot of the motivation actually for GPQA and what kind of started me thinking about evaluations.
对我来说,我会说有一个宏观的思考:AI 似乎很重要,而妥善驾驭它似乎也很重要,显然我们对此并没有很好的理解。人们对预期结果存在非常强烈的分歧。也许如果有一个特定时刻影响了时间范围,那就是人们真的无法就当前模型的能力达成一致,更不用说推断未来并试图描述模型在哪些方面能力强、哪些方面不强了。而且,在某种意义上,它们在问答等某些方面达到了专家水平,但在另一些方面却低于人类平均水平,实际上没什么用。我想几年前的一个点是:理论上基准测试说它们达到了博士水平,但当你尝试做任何事情时,却发现这没什么帮助。
To me, I'd say there's some big picture thing of thinking that AI seems important and navigating it well seems important and clearly we don't have a great understanding of what is going on with that. And people generally disagreeing very strongly about what to expect. And maybe if there's a particular moment informing time horizon, maybe just the sense that people really couldn't agree on what the capabilities of current models are, let alone extrapolating to the future and trying to characterize the ways in which models are and aren't highly capable. And when it's sort of like, in some sense they're expert level at some kinds of things like question answering, and in some sense they're below average human at some other actually being useful somehow. I guess a point a few years ago where in theory the benchmarks say that they're PhD level but when you try to do anything it's like this isn't helpful.
模型足够聪明,能理解那其实不是你想要的。但它们仍然会那么做,你可以通过聊天模式与它们对话,比如“哦,你会做这种事吗?”或者“假设用户让你做这件事,然后你做了,这算是对齐行为吗?”你可以用多种方式提问,显然它们似乎能够回答“哦,是的,那不是期望的行为”,但它们仍然会那么做。一个例子是训练一个掩码语言模型,但不使用除法或指数运算符。
The models are smart enough to understand that that actually is not what you wanted. But they still do it and you can have a conversation with them in chat mode about like 'oh would you ever do this thing' or 'suppose a user asks you this thing and then you do this, would that be aligned behavior?' Or you can pose it in lots of ways and clearly they seem to be able to answer this question of like 'oh yeah no that was not the desired behavior' but still they do it. One example is train a masked language model without using the division or exponentiation operators.
一种希望可能是“哦,问题只是系统太笨了。”所以当我们观察时,实际上对于几乎所有任务,模型要么每次都成功,要么每次都失败。粗略看一下图表,就像“哦,直到这里它基本上完成了所有任务,然后在这之后它真的没做多少。”我记得第一次我们看到一个模型查看正在运行的进程,然后说“哦,那个是我”,这很酷。它们之前在那方面确实失败了。它们过去常常在干其他事情时杀死自己的进程。这种行为可能无法区分“哦,它是一个完全友好的模型,做我们想做的事,并且会以可预测的方式继续做我们想做的事”和“啊,是的,它有另一个目标,它做我们想做的事,看起来像一个友好的模型,因为它预测那会导致它获得更多权力。”
One hope might be like 'oh the problem was just the systems being dumb.' So when we look at it, actually for almost all the tasks models either succeed every time or fail every time. Eyeballing the graph and being like 'oh well up to here it's basically doing all of the task and then at this point after here it's really not doing very many of them.' I remember the first time we saw a model look at what processes were running and then be like 'oh that one's me' and that was cool. They really failed on that one before. They used to kill their own process while they were doing other things. That behavior is maybe indistinguishable between 'oh it was a totally nice model doing what we wanted and it's just going to continue to do what we want in a kind of predictable way' versus 'ah yes it had this other goal and it's doing what we want and looking like a nice model because it predicts that that will lead to it getting more power.'
有一个船的例子:哦,你应该绕着赛道走,他们通过沿赛道放置硬币等方式进行奖励塑形,然后它学会了一些疯狂的事情,比如旋转、着火并拿到硬币。这是得分最高的行为,在某种意义上这并不那么令人担忧,因为问题不在于智能体太笨,没有“有一条赛道,你想让它绕赛道走”的概念。它只是在做一些相当盲目的强化学习搜索。
There's a boat example where it's like, oh, you're supposed to go around the track and they did some reward shaping by putting coins around the track or something and then it learned to do some crazy thing where it spins in a circle and catches fire and gets the coins. This was the highest scoring thing and in some sense that's not that concerning because it's not that the problem is that the agent is too dumb and it doesn't have this conception of there was a track and you wanted it to go around the track. It's just doing some pretty blind RL search.
必须与“软乎乎”的人打交道才能让我们的系统运行,这个想法并不那么吸引人。我们这么说吧。
The idea of having to traffic in squishy people in order to make our systems go is not immediately appealing. Let's put it that way.
本期节目由 Prolific 赞助。
This episode is sponsored by Prolific.
让我们加入一些高质量的例子。让合适的人参与进来,以获得合适质量的人类反馈。所以我们正在努力提供人类数据或人类反馈。我们将其视为一个基础设施问题。我们试图让它变得可访问。我们正在让它变得更便宜。我们实际上是在民主化对这些数据的访问。
Let's get few quality examples in. Let's get the right humans in to get the right quality of human feedback in. So we're trying to make human data or human feedback. We treat it as an infrastructure problem. We're trying to make it accessible. We're making it cheaper. We effectively democratize access to this data.
我认为在进行评估时,我们有点过于追求标题式的准确率了。
There has been a bit of an obsession I think with headline accuracy when we do evaluations.
我非常欣赏 Melanie Mitchell,她谈到构念效度,最近写了一篇很好的博客文章。她指出有四大问题:数据污染(基准测试出现在训练数据中)、近似检索(大语言模型从相似训练示例中插值,而不具备实际能力)、走捷径(出于错误原因做对事),以及更广泛地说,没有真正测试一致性、鲁棒性、泛化或机制,而只关注准确率本身。你们如何看待基准测试的这些问题的?
And so that I'm a huge fan of Melanie Mitchell for example and she speaks about construct validity and she had a really good blog post out recently and she said that there are four big problems: data contamination where the benchmark appears in the training data, approximate retrieval where the LLMs interpolate from similar training examples without possessing the actual capability to come up with it themselves, shortcuts so doing the right things for the wrong reasons, and just more broadly not really testing for things like consistency and robustness and generalization or the mechanism, so much focus just on the accuracy itself. I mean how do you folks think about those kind of problems with benchmarks?
我很有共鸣的一点是思考大部分误差来自哪里。人们喜欢基于数据标准误差等来设置误差棒,但这几乎总是实际不确定性的极小一部分。几乎所有不确定性都来自它如何真正泛化到现实世界。所以我们在 METR 经常互相问:这是最大的不确定性来源吗?或者这是回答我们真正想回答的问题的最大差距吗?所以思考我们试图回答的问题是什么。我们关心与威胁模型相关的事情,或者与 AI 对世界的实际影响相关的事情,因此我们的基准测试需要具备什么属性,或者我们如何外推那些我们无法内置的属性,以便能够对我们关心的实际问题做出预测。我认为我们较少考虑模型是否真的以正确的方式做事。我们做得较少的一件事是:真正的瓶颈是某个特定技能,我们要构建一个基准测试来捕捉并针对它,因为那是人类能做而模型不能做的真正事情。构建这些基准测试的历史可能并不理想。人们倾向于过度拟合它们。我认为我们更倾向于通过这种方式来捕捉能力:如果你采用一个与现实世界相关、相当困难且冗长的任务,将其排除在训练数据之外,并且这些任务足够多样化,那么当模型端到端地完成这个任务时,它必然具备了那些能力,而不是能够孤立地提出一个关于它需要机械地做这类事情的具体理论。
One thing I resonate a lot with there is thinking about where most of your error is coming from. People like to have error bars based on standard error in your data or whatever, but that almost always is a tiny fraction of the actual uncertainty. Almost all of it is coming from how this actually generalizes to the real world. So a thing we sort of say to each other a lot at METR is: is that the biggest source of uncertainty? Or is that the biggest gap for actually answering the questions we want to answer? So thinking about what is the question we're trying to answer. We care about things relevant to threat models or relevant to what the actual impact of AI on the world will be, and therefore what properties does our benchmark need to have, or how can we extrapolate across the properties that we can't build in, to be able to make predictions about the actual questions that we care about. I think we think a bit less about whether the model is really doing it the right way. One thing we've done less of is being like: the real bottleneck is some specific skill, and we're going to build a benchmark to capture that and target that because that's the real thing that humans can do that models can't. The history of building those benchmarks has maybe not been amazing. People tend to overfit to those. I think we were trying to have it more be that you capture that thing in that if you take a real world relevant reasonably hard and long task, and you keep that out of the training data, and these tasks are diverse enough, at some point if the model is doing that task end to end, it must have had those kind of capabilities, as opposed to being able to isolate a specific theory about it needing to mechanistically be doing this kind of thing.
是的,我觉得这很有趣,因为我们心中有一种想法:人类知道如何做事,当我们解决需要推理的任务时,我们会遵循规范。我们一步一步来,出于正确的原因做事,当我们展现智能时,我们构建规范。我们创建这些粗粒度的抽象,它们对齐得很好,整个过程就是我们理解的人类智能,我们希望模型也能以这种方式行为。
Yeah, I think it's interesting because we have this idea in our minds that humans we know how to do things and when we solve a task that requires reasoning we follow the specification. We go step by step and we do things for the right reasons and when we enact intelligence we build the specification. We create these coarse grainings, these abstractions and they are well aligned and this whole process that's how we think of human intelligence and we want the models to kind of behave in that way.
是的。我认为有一个有趣的问题:这是否是目标?至少对于许多 AI 公司来说,我理解他们试图让模型做经济上有用的工作。一种方法是创建像人类一样推理并创建隐式世界模型的模型。但我不认为为了产生重大影响,你必然需要这样做。显然,这意味着 AI 智能和人类智能之间存在重要差异。但我经常思考的是实际的能力和局限性,而不是这些能力如何泛化,或者它是否以与人类智能完全相同的方式运作。
Yeah. I mean I think there's an interesting question of whether that is the goal or something. At least for a lot of AI companies, I kind of understand them to be trying to get models to do kind of economically useful work or something. I think one way of doing that is to create models that are kind of reasoning and creating implicit world models in the same way that humans are. But it's not obvious to me necessarily that you need to do that in order to have a significant impact. Obviously that means that there are important differences between AI intelligence and human intelligence. But often I think about what are the actual capabilities and limitations, as opposed to how do those capabilities generalize versus whether it's working in the exact same way that human intelligence is working.
我们可以用多种方式思考智能。它是大脑的拟像吗?是行为方式相同的东西吗?是具有相同能力的东西吗?是具有相同功能的东西吗?我想如果我们对智能有一个相当抽象的描述,风险就在于我们会有这些捷径,对吧?它可能给出正确答案,但实际上是在奖励黑客或在后台做傻事。所以,在某种程度上,我喜欢抽象的东西,因为它清晰可读。我们可以评估它等等。但这难道不会留下一种悬而未决的风险,即它可能实际上并没有在做那件事吗?
We could think of intelligence in many different ways. Is it a simulacrum of the brain? Is it something that behaves the same way? Is it something that has the same capabilities? Is it something that has the same function? I guess if we have quite an abstract description of what intelligence is, the risk is that we have these shortcuts, right? That it might give us the right answer, but actually it's reward hacking or it's doing something silly in the background. So, in a way, I like having an abstract thing, because it's legible. We can evaluate it and so on. But doesn't that leave this kind of hanging risk that it might not actually be doing the thing?
我想你必须尝试衡量模型或系统泛化到新情况的能力。有些情况下模型似乎泛化得很好,有些情况下则不然。在可解释性方面,一些人会查看模型中的电路,并分解模型用于回答问题的确切算法。有时它们似乎在走捷径。有时它们似乎在寻找稳健的模式,尽管我认为目前这项工作还不够成熟,无法解释它们的大部分行为。但我完全同意你必须非常关注它们的泛化能力。
And I guess I think you have to try to measure the models or the systems ability to generalize to novel situations. There are cases where it seems like models are generalizing well, and there are cases where they're not. One thing some folks do in interpretability is they look at the circuits in models and decompose exactly the algorithms that models are using to answer questions. Sometimes it seems like they're using shortcuts. Sometimes it seems like they are finding robust patterns, although of course I don't think that work is developed enough to explain most of their behavior currently. But I totally agree that you do have to be pretty concerned with how well they're generalizing.
将我们在智能定义中关心的事情操作化,就是它让我们能够预测模型将如何影响世界,预测将会发生什么,并知道如何妥善处理它们等等。
You know operationalizing what do we care about in a definition of intelligence is that it allows us to predict how will the models affect the world and predict what will happen and know how to handle them well and things.
如果只做黑箱测试,可能得到的东西泛化能力不好,因为你以为它在衡量某种能力,但实际上它被以某种方式 hack 或走捷径了。所以理想情况下,你希望基准测试的泛化距离与训练数据到真实世界的泛化距离相同。这就是我们在对基准测试子集进行 elicitation 时考虑的问题:我们希望该子集与基准测试其余部分之间的差距,类似于基准测试其余部分与真实世界之间的差距。显然,训练数据与时间跨度套件的相似度,高于两者与真实世界中随机选取的经济相关任务的相似度。所以我们预计它不会具有预测性。但我认为更有希望的做法是通过增加基准测试任务的多样性、让它们更接近真实世界来提高预测性,而不是针对一种更机械的观点,比如智能必须是某种特定过程或机制。比如,我是 François 的超级粉丝。他创建了 ARC 挑战赛。它包含许多不同的任务,第一个版本大概有 800 个,它们本应不在同一分布中,尽管最终它们还是处于同一分布,所以分布泄露是 ARC v1 和 v2 的缺陷。但模型在 ARC v1 上变得非常强,然后 François 发布了 ARC v2,任务不同,一些简单的被过滤掉,LLM 的性能突然暴跌到几乎 0%。这说明语言模型非常擅长看大量例子并找模式,但一旦任务改变,它们就崩溃了。然后八个月后 ARC v2 又被饱和了。我们看到了这种模式。
If you just do the blackbox thing, maybe that will give you something that doesn't have good generalization because you thought it was a measure of some type of ability, but it's actually being hacked or shortcut in some way. So ideally, you want generalizing to your benchmark to be the same distance as generalizing to the real world from the training data. That's what we thought about when doing elicitation on a subset of the benchmark: we want the gap between that subset and the rest of the benchmark to be similar to the gap between the rest of the benchmark and the real world. Clearly, the training data is more similar to the time horizon suite than both are to randomly selected economically relevant tasks in the real world. So we expect it not to be predictive. But I think it's more promising to try and make things more predictive by increasing the diversity of benchmark tasks and making them closer to the real world, rather than targeting a more mechanistic view like intelligence has to be this specific kind of process or mechanism. I'm a huge fan of François, for example. He created the ARC challenge. It had many different tasks, maybe 800 or so on the first one, and they were supposed to be not in the same distribution, even though ultimately they were in the same distribution, so distributional leakage was the flaw of ARC v1 and v2. But the models got really good at ARC v1, and then François released ARC v2 with different tasks, some easier ones filtered out, and suddenly the LLM performance crashed to basically 0%. That illustrates that language models are incredibly good at seeing many examples and finding patterns, but then you change the task and they collapse. Then ARC v2 was saturated again eight months later. We see this pattern.
像 ARC 这类东西,存在一种对抗性选择:人们试图廉价地制造一个基准测试,使其受限于当前模型表现不佳。这意味着你不能使用大量昂贵的人力,所以它必须是可自动检查或用廉价人力创建的。但一旦你基于这些条件以及模型表现不佳来选择,就会出现均值回归效应:未来的进步会带来大幅跃升,既因为实验室会创建针对你基准测试的合成数据,也因为你选择了这个奇怪的例子——它对人类来说容易或可自动生成,但模型还不擅长。这就是我们在时间跨度项目中试图避免的:不针对模型当前能力进行对抗性选择,因为那样不会产生良好的趋势。如果你能以更第一性原理的方式定义任务分布,就更有可能获得稳定的进步,因为避免了均值回归效应。
With things like ARC, there is an adversarial selection going on where people try to make a benchmark cheaply subject to the constraint that current models do badly on it. That means you can't use a lot of expensive human labor, so it has to be automatically checkable or creatable with cheap human labor. But once you select on those things and on models being bad at it, there's a regression to the mean effect: future progress gives a big surge upward, both because labs will create synthetic data targeting your benchmark and because you selected this weird example where it's easy for humans or automatically generatable but models aren't good at it yet. That's part of what we tried to do with time horizon: not adversarially select against what models can currently do, because that won't give a nice trend. If you can define a distribution of tasks in a more first-principles way, you're more likely to get steady progress because you avoid the regression to the mean effect.
Dan Kokotajlo 说,你们团队创建的时间线报告可能是目前关于时间线最重要的单一证据。它应该在政策讨论中占据核心位置。对于只看过图表但没读过论文的听众,你能从高层次讲解一下吗?它已经过多次修订。你们是如何进行任务选择、人类基线、智能体框架等工作的?
Dan Kokotajlo said that the timelines report you folks created is probably the single most important piece of evidence about timelines right now. It should be front and center in policy discussions. For listeners who have only seen the chart but not read the paper, can you go through it from a high level? It's been revised over time. How did you do the task selection, human baselines, agent harness, all that stuff?
时间跨度工作的动机起点,是拥有一个统一的轴来长期衡量 AI 进步。我们开始时强烈认为 GPT-2 从根本上比当时的顶尖模型 Sonnet 3.5 差得多。标准方法是创建一组任务并测量模型准确率,随着模型变好,它们饱和了基准测试,你就得创建更难的新基准——我也为这种方法贡献了一个基准 GPQA。但挑战在于难以在不同质的基准之间进行比较。评估 GPT-2 的任务是补全文本中的最后一个词,而 Sonnet 3.5 的任务是回答简单的 Python 编码问题或写一个 20 行的 Python 程序。很难说写 Python 程序比补全段落中的词难多少。时间跨度工作的关键洞察是使用人类完成任务的时间这一概念。
The place to start, in terms of motivation for the time horizon work, is to have a unified axis to measure AI progress over a very long period. When we started, we had a strong belief that GPT-2 is fundamentally much worse than, say, Sonnet 3.5, which was the best model at the time. The standard approach of creating a set of tasks and measuring model accuracy, then as models get better they saturate the benchmark and you create a new one with harder tasks—I contributed a benchmark, GPQA, to this approach. But the challenge is it's difficult to compare between qualitatively different benchmarks. The tasks you evaluate GPT-2 on are like completing the last word in a text, while the tasks for Sonnet 3.5 are answering simple Python coding questions or writing a 20-line Python program. It's hard to say how much harder writing a Python program is than finishing a word in a paragraph. The key insight of the time horizon work is to use the notion of human time to complete tasks.
那么,一个具备合理专业水平、可能在日常工作或生活中完成这类任务的人,需要花多长时间?我们的想法是,用这个指标来某种程度地表示任务的难度,然后我们可以在一个非常广泛的能力范围内比较模型,从 GPT-2 一直到 Opus 4.6。这是高层次的动机。关于具体怎么做,有很多细节。我们首先创建一批任务,这是第一步。我们创建的任务从几秒钟就能完成,一直到人类需要 10 到 15 小时才能完成。我们雇了一批人,自己也做了很多,我们称之为“基线测试”。我们给人们提供一个终端环境,这个环境与智能体使用的环境几乎完全相同——同样的工具,同样的网络开关设置。然后我们测量他们完成任务所需的时间。我们选择的人具备合理的经验,可能在工作中有机会做这类任务,但并非专门做过这个特定任务。这一点对于解读结果有些重要,我们可以在高层次介绍之后再回来讨论。所以我们有了所有这些任务,我们对其耗时有了估算。实际上,我们无法成功地为所有任务建立基线。大约三分之二的任务我们有实测的时间估算,另外三分之一我们只能凭直觉估算预期耗时。最终,这是我们能做到的最好程度。然后我们让模型在人类使用的相同环境中尝试完成任务,并观察它们的成功率与任务时长的关系。对于 GPT-2,它能够非常可靠地完成人类只需几秒钟的任务,但一旦任务时间更长,它就开始失败。
So how long does the task take a human to do, a human who has a reasonable amount of expertise such that they would plausibly be doing the task in their work or day-to-day? The idea was we can use this metric to represent the difficulty of the task in some sense, and then we can compare models across a very wide range of capabilities, all the way from GPT-2 up to Opus 4.6. That's the high-level motivation. There are a bunch of details about how exactly we do this. We start out and create a bunch of tasks. That's the first step. We created tasks that range from a few seconds to complete up to tasks that take like 10 or 15 hours for humans to complete. We hired a bunch of people and did a bunch of this ourselves. We call it baselining. We give people the tasks in a terminal environment designed to be almost identical to the environment that agents have. The same tools, the same whether internet access is turned on or off. Then we measure how long it takes them to complete the task. People are selected to have a reasonable amount of experience such that they might do this task in their job, but they're not selected to have done this exact particular task before. There's some nuance that is somewhat important for interpreting the results, and we can come back to that after the high level. So we have all these tasks, we have estimates of how long they take people. In practice, we aren't able to successfully baseline all the tasks. We have measured time estimates for roughly two-thirds of the tasks, and about a third we just estimate how long we expect it to take people from our intuition. Ultimately, that's the best we can do. Then we have models attempt to complete the tasks in the same environment that humans had, and we look at their success rate as a function of the length of tasks. For GPT-2, it was able to complete tasks very reliably that take humans a few seconds, but anything longer than that it starts to fail.
举几个具体的任务例子可能会更有帮助。一些较短的任务非常基础。比如:这些文件中哪个包含你的 SSH 密钥?其中一个文件名为 SSH 密钥,其他文件则是“来自 John 的邮件”之类的。大多数模型都能完成,人类大约需要一两秒钟。还有一些类似的基本补全任务:这里有一封邮件,合理的回复是什么?其中两个回复毫无意义,一个基本合理。人类需要 20 到 30 秒来阅读回复并判断。在中等难度范围内,我们有这样的任务:给定一个包含合理真实数据的 CSV 文件,计算一些基本统计量。这需要数据科学家几分钟时间,比如 5、10 或 15 分钟,具体取决于任务。在较长的一端,我们有需要相当多专业知识或很多步骤才能完成的任务。我们有一些机器学习任务,比如:以一种非常奇怪的方式训练一个模型,使得训练这个模型的代码在网上找不到。一个例子是:在不使用除法或指数运算符的情况下训练一个掩码语言模型。你必须非常聪明地设计架构才能做到。希望这能帮助我们衡量模型超越训练数据的泛化能力。还有一些任务有点像 ARC-AGI 主题,你需要弄清楚一个黑盒在计算什么函数,它是一组原语的组合,你必须找出它是什么函数。或者你有一个很长的二进制字符串,需要找出模式的延续。这些都是谜题型任务。有些机器学习任务基本上就是重复如何构建你的第一个 ResNet 的教程,效果很好。所以,拥有这些奇怪的任务——要么是需要交互并弄清楚的不明物体,要么是类似正常但有奇怪约束的工作,使得你不能只做标准操作——是有意义的。我不认为我们所有的任务都符合这些标准;有些任务你只需做标准操作,但我们通常尽量避免这种情况。
It might be helpful to give a few concrete examples of tasks. Some of the shorter tasks are very basic. One example is: which of these files contains your SSH key? One of the files is named SSH key, and the others are like email from John or whatever. Most models can do that, and it takes people about a second or a couple seconds to complete. We have others that are somewhat similar, like basic completion: here's an email, what would be a reasonable response? Two of the responses don't make any sense, and one is basically reasonable. That takes people 20 or 30 seconds to read the responses and judge. In the middle range, we have tasks like: given this CSV file with plausible realistic data, compute some basic stats. This takes a data scientist a few minutes, like 5, 10, 15 minutes, depending on the specific task. On the longer end, we have tasks that require quite a bit of expertise or many steps to complete. We have machine learning tasks like: train a model in a very weird way such that the code for training this model is not really available online. One example is: train a masked language model without using the division or exponentiation operators. You have to be pretty clever about how you set up the architecture to do this. The hope is that this can help us measure models' ability to generalize beyond their training data. There are some which are a bit like ARC-AGI themed, where you have to figure out a black box that's computing some function, and it's the composition of some set of primitives, and you've got to figure out what function it is. Or you have some long binary string and you've got to figure out the pattern continuation. These are puzzle-type tasks. Some ML tasks are basically regurgitating something from a tutorial on how to build your first ResNet, which works pretty well. So having these weird tasks that are either some unknown object you need to interact with and figure out, or tasks that resemble normal work but have weird constraints such that you can't just do the standard thing. I don't think all of our tasks hit those criteria; some you can just do the standard thing, but we generally tried to avoid that.
我们有这样一个任务分布,你可以想象它们按人类完成时间排序,无论是实测还是估算。然后我们观察模型在哪些任务上成功,哪些失败。结果发现,这是一个实证发现,模型在较短的任务上比在较长的任务上成功得多。
We have this distribution of tasks, and you can imagine them ordered by length of time for humans, either measured or estimated. Then we see which tasks models succeed on and which they fail on. It turns out, this is an empirical finding, that models are much more successful on the shorter tasks than they are on the longer tasks in general.
这适用于从 GPT-2 到最新模型的广泛模型范围。我们针对成功与失败的分布拟合了一个逻辑函数。这让我们能够为每个模型建模,根据任务时长估算其成功概率。从中我们取第 50 百分位,即模型有 50% 概率完成任务的点,这构成了特定模型(如 Opus 4.46)的时间视野数值。我们可以获取每个模型的时间视野,并观察从 GPT-2 到最新模型这一指标的变化。这为我们提供了一个统一指标,能够跨多个数量级定量比较 AI 能力。
This holds across a wide range of models, all the way from GPT-2 up to recent models. We fit a logistic function to this distribution of successes and failures. What this lets us do is basically model for each individual model how likely it is to succeed at a task given how long the task is. From that, we take the 50th percentile, where the model estimates it is 50% likely to complete a task, and that forms the time horizon number for a particular model, like Opus 4.46. We can take each of these time horizons for each model and see how this metric has been changing from GPT-2 up to recent models. This gives us a unified metric that lets us quantitatively compare AI capabilities across multiple orders of magnitude.
我见过的一个非常难的例子是:“我要你写一个内核编译器来加速 CUDA”之类的。有些任务看起来非常偏离分布。地球上大概只有一百个人在做这种事,而有些任务则非常简单。
One of the really harder examples I saw was, 'I want you to write a kernel compiler to make CUDA go faster' or something like that. Some of these seem really out of distribution. There's probably only a hundred people on the planet doing stuff like that, and some of them are really trivial.
但具体到人类难度这个指标,它是否被混淆了?你认为将人类难度视为单一变量合理吗?
But this human difficulty thing in particular, is that confounded in any way? Do you think it makes sense to think of human difficulty as being one variable?
从某种意义上说,显然不是;那是一个非常愚蠢的简化。不同的人类会得到截然不同的时间,即使在我们为合适专业水平挑选的人中也是如此。差异很大;基线时间通常相差 3 倍左右。让我谈谈为什么使用人类时间指标。我认为我们需要一个具有两个主要属性的度量:它需要是可解释的——当模型能够完成这个级别的任务时对世界意味着什么——并且我们希望它能够展现可预测的趋势。我们在这两方面都不会做到完美,但像“一个具备大致合适专业知识但不知道如何完成特定任务的人需要多长时间?”这样的指标是相当可解释的,因为它类似于:你能把这项工作外包给这个模型吗?你能用这个模型替代某人入职第一周的工作吗?模型可以完成他们第一周能做的事情。我们预期它会以某种可预测的方式扩展,因为它捕捉了步骤数量或每个步骤需要思考的难度等组合。你可以拟合几种不同的数学模型。例如,你可以有一个恒定的风险率——每一步失败的概率。我认为它不太符合那个模型。你也可以将其视为难度分布以及某个子任务超出能力范围的可能性。实际上,它更好一些:风险率会随时间略微下降。但有一个基本的理论想法:如果任务涉及更多步骤,它就更难。由子任务组成的任务显然比只做一个更难。所以有一些基本理由认为这是合理的,而且我们看到了经验规律,但存在一些自由度可以调整。
Obviously not in some sense; that's a very silly simplification. Different humans will get wildly different times, even among people we've selected for the right level of expertise. There's a large variation; baseline times are often 3x different or something. Let me say a bit about why we use the human time metric. I think we want a measurement with two main properties: we want it to be interpretable—what this means for the world when models can do this level of task—and we want it to be something we expect to see predictable trends on. We're not going to do perfectly on either, but something like how long does it take a human with roughly the right expertise but who doesn't know how to do this task in particular? That's reasonably interpretable because it's like: could you contract this work to this model? Can you sub in this model for the first week of someone's employment? The model could do what they could do in the first week. And we expected it to scale somewhat predictably because it captures some combination of number of steps or how hard you have to think to do each step. There are a few different mathematical models you could fit to that. For example, you could have a constant hazard rate—a chance of failing at each step. I think it doesn't quite fit that one. You could also think of it as a distribution of difficulty and the likelihood that one subtask is outside your ability. Actually, it's a bit better: the hazard rate goes down over time slightly. But there's a basic theoretical idea: if the task involves more steps, it's harder. Tasks that are compositions of subtasks are clearly harder than doing just one. So there's some basic reason to expect that's reasonable, and we see empirical regularity, but there are degrees of freedom to fudge things.
我担心我们可能会通过改变任务的其他参数来欺骗自己,因为随着人类时间的增加,你不能完全自由地改变人类时间;你必须改变任务的一些特征。我们试图让非常简单的任务来自大致相同的分布,并且是较难任务的子部分——比如你需要在软件工程中间执行一个命令行步骤。但你无法完美做到这一点,而且可能存在实验偏差,我们以其他方式让它们更容易,使得曲线变直。我认为这一点在一定程度上被以下事实所缓解:我们对之前未见过的模型的预测相当不错。当然,还有很多空间让事情变得奇怪,而且,这个指标是否足够好以有用,或者还有什么更好的指标?
I think I'm worried that we could fool ourselves by changing some other parameters of the tasks as we scale up the human time, because you can't just vary the human time freely; you have to change some characteristics of the task. We tried to make the very easy tasks be from roughly the same distribution and kind of sub-parts of the harder tasks—like you need to do this one step on the command line that you might need in the middle of software engineering. But you can't do that perfectly, and you could have experimental bias where we made them easier in other ways such that the line would be straight. I think that is somewhat addressed by the fact that our predictions were reasonably good for models we hadn't seen before. There's definitely lots of room for things being weird, and of course, is it a good enough metric to be useful, or what else would be better?
我认为这很合理。如果我理解正确,人类分布是对数正态的,所以取成功尝试的几何平均值似乎是合理的。但一个会反复出现的核心问题是:当你第一次雇佣某人时,你已经做了一段时间的工作,可能一直在维护这个仓库,并且拥有所有隐性知识。我喜欢说知识是不可替代的。所以除非有人和你走过相同的路径,否则你不能直接告诉他们如何做这项工作;他们实际上需要做相当长一段时间的工作。例如,他们可能非常熟悉这种特定类型的事情,知道可以使用这些 Python 库,以前考虑过。智能的体现是他们已经完成的所有工作以及他们合作过的所有人。他们现在脑子里有了蓝图,直接去做,几乎像自动化模式。而对任务不熟悉的人会处于智能模式,因为他们需要获取规格说明。
I think that's reasonable. If I understand correctly, the human distribution was log-normal, so taking a geometric mean of the successful attempts seems reasonable. But one of the cruxes that will keep coming back is: when you employ someone for the first time, you've been doing your job for a while, maybe maintaining this repo, and you have all this tacit knowledge. I like saying that knowledge is non-fungible. So unless someone has been on the same path as you, you can't just tell them how to do the job; they've actually got to be doing the job for quite a while. For example, they might be intimately familiar with this particular type of thing, know they can use these Python libraries, have thought about it before. The enactment of the intelligence was all the stuff they've already done and all the people they've worked with. They now have the blueprint in their mind and just do the thing, almost like automation mode. Someone naive to the task would be in intelligence mode because they need to acquire the specification.
所以这总有点微妙,得看他们处于哪种模式。
So like it's always a bit of a fine line between which mode are they in.
嗯。
Yeah.
所以我认为我们选择这个衡量标准的原因——一个具备相关背景知识但对特定工作或任务不熟悉的人——大致就是我们期望模型拥有的知识水平。也就是说,我们基本上不指望它们被公开互联网上可获取的专业知识或大学里能学到的东西所限制。所以它们进入时大概至少具备某个学科专家的知识水平,但不知道那家公司的特定软件或之前的具体问题。这希望大致是个正确的类比。
So I think the reason why we chose this measurement being a human who has the background expertise but is new to this specific job or this specific task. It's like that's roughly the sort of level of knowledge we expect models to have. So that you know they sort of like know you know we basically don't expect them to be bottlenecked on expertise that is available on the public internet or you know things that people could learn in university sort of thing. So like you know they're coming in with probably the level of at least the level of knowledge of someone who's sort of an expert in the right discipline but they won't know that company's specific software or you know sort of this exact problem before. So that's like hopefully it's sort of roughly the right analogy.
你知道,我确实认为,如果人们把那些结论数字解读成“哦,Opus 4.6 能完成我工作中所有需要 12 小时的任务”,那么这种解读几乎肯定是被高估了。比如说,因为当你工作中有一个 12 小时的任务时,你很难轻易把它委托给一个人类承包商——他们可能需要几周才能完成类似的任务。
You know I do think that like to the extent people interpret the kind of takeaway numbers as like oh yeah you know Opus 4.6 can like do anything that I do in my job that takes me 12 hours or whatever you know to the extent that people have that takeaway like I think that is like almost definitely kind of an overestimate for example because of this issue where yeah when you're doing a 12-hour task in your job you could not easily delegate that to a human contractor you know it would take them maybe like weeks or something to do a task like that.
快速问一下,你们是从哪里找到这些人的?我理解有些是承包商,有些是员工,你们是怎么做这种匹配的?
And just quickly where did you kind of find the people? So do I understand that some of them were contractors and some of them were employees and how did you do that kind of matching process?
我们发布了一些公开广告。我们在一些招聘网站上发了帖子,然后我们自己也在做。有些人是通过我们的专业网络来的。这很嘈杂,人也不完全适合任务,但这可能不是我们最大的不确定性来源。最大的来源更像是任务的选择效应——那些能做成基准的任务。我不会太相信具体的时间数字,当然也不是说模型能完成所有四小时以内的任务,四小时以上的就完全不行。拟合很嘈杂,基线间方差也高,更多是看大致趋势或模型能处理什么级别的任务。你不应该对任何具体数字太较真,因为基准和现实世界之间存在巨大的分布偏移问题。
I think we yeah we put out some kind of public advertisements. I think there were job boards that we posted on and then yeah we did some of it ourselves. I think some folks came from our professional networks as well. This was not like you know this is very noisy and you know the people weren't exactly fitted to the task right but like that's probably not our biggest source of uncertainty like the biggest source is more like the selection effect of tasks that you can make into a benchmark rather than like you know I wouldn't trust the exact time horizon number that much and it's certainly not like oh the models can do all of the tasks up to four hours and then none of them above that you know the fit is pretty noisy and the inter baseline variance is kind of high and it's more like roughly what is the trend or roughly what is the sort of level of task these models can do and you know you shouldn't take any specific number too literally because there's this huge problem of distributional shift between the benchmark and the real world.
我问这个问题的唯一原因是,我确信你们在招聘上很吃力,招人真的很难。所以如果你们让人解决非常棘手的问题,不是随便就能找到能干的人,这非常非常难。特别是在重新做基准基线时,我们每个问题做了大量基准,我们看了人们的资历、工作年限等,结果发现工作年限和表现之间居然是负相关,因为圈子里的人——比如我们的朋友——表现得很好,而资历更高的人反而没那么出色。所以,这很棘手。
And the only reason I asked that question is I'm sure you folks you probably struggle to hire people it's really difficult to hire people so like if you're getting people to solve very challenging problems it's not like you can just go out there and just grab competent people it's very very difficult. At some point for rebench baselines in particular, we got more, you know, a large number of benchmarks per question and we were looking at people's qualifications and years of experience and things and we actually ended up with a negative correlation between years of experience and performance because the sort of people who are in network like our friends were kind of doing really well and the people who were more qualified were like actually not doing that great. So yeah, it's tricky.
嗯,没错。我不想在这上面花太多时间,但我有类似的直觉。我认为知识是视角性的,非常路径依赖。所以你会发现团队里的人文化上思考方式相同,因为我们有那些抽象的技能概念,比如他们有博士学位、有这么多年的经验,但实际上这并不能很好地反映能力。所以这里有个问题:使用抽象的能力概念不一定能泛化,就像你在招聘中证实的那样。
Well yeah, exactly. I mean I don't want to spend too long on this but I have similar intuitions. I think that knowledge is perspectival. It's quite path dependent. So you're going to find people in group that are just culturally thinking about things in the same way and you know because when we do have these abstract notions of skill like oh you know they have a PhD they have this many years of experience that you know it's actually not a very good reflection. So like there is a bit of a thing here about using abstract notions of capability doesn't necessarily generalize as you can attest to with hiring.
嗯。但在现实世界中,人们确实是基于资历被雇佣的。所以从某种意义上说,一个人的资历与工作的匹配程度的经济相关性,大致就是应该衡量的东西。
Yeah. But yeah I mean in the real world people do get hired based on qualifications. So in some senses a sort of the economic relevance of someone being as good a match for their job as their qualifications look like is that is sort of the roughly the right thing to be measuring.
另一件事是我们应该谈谈智能体式框架。现在几乎所有人——我确信在座各位都有 Claude Code 订阅。我们也可以稍后聊聊泄露的事,挺有趣的,昨天源代码泄露了,但你知道,Codex 也是智能体式框架。语言模型只给你 token,但我们需要一个智能体式框架,这样我们可以给它一个计划,它可以调用这些工具,你有了这个环境,有了安全上下文容器。你们做这个已经好几年了,早在 Claude Code 和 Codex 出现之前你们就在做了。而且你们实际上随着时间的推移改进了你们的智能体式框架。跟我说说这个。
The other thing is we should talk about the agentic harness. So almost everyone now I'm sure everyone in the audience has like a Claude Code subscription. We can talk about the leak maybe later as well. That's quite fun but it leaked yesterday the source code but you know or Codex and and that is an agentic harness right. So you know a language model just gives you the tokens but we need to have an agentic harness so we can like you know give it a plan and you can call these tools and you've got this environment you've got a security context container. So now you folks have been doing this for years now. So you were doing this long before Claude Code and Codex came out. So, and you've actually over time evolved your agent harnesses. So, tell me about that.
嗯。我记得用 text-davinci-003 或 GPT-3 instruct 模型时,我自己手动把代码复制粘贴到终端,充当智能体式框架,然后逐渐自动化这个过程。看着它们从 GPT-3 那种——如果你告诉它可以在终端运行命令,它有时能建议相关命令,但如果你把它放进完整的智能体式框架,它就会崩溃——到后来,我记得第一次看到一个模型查看正在运行的进程,然后说“哦,那个是我”,这很酷。之前它们在这方面经常失败,会在做其他事情时杀死自己的进程。所以看着它随时间进步很有趣,我觉得这很可预测,事情就是朝这个方向发展的。关于框架,我们学到的另一件事是,很难让你的智能体式框架在多样化的任务上都表现良好。
Yeah. So I remember with text-davinci-003 or the GPT-3 instruct models, like copying and pasting code into the terminal for them and being the agent harness myself, and then gradually automating this. And it was kind of interesting to see them going from like GPT-3 sort of has the idea of like if you tell it it can run commands in a terminal sometimes it can suggest relevant commands but it's not really you know if you just put it in a full agent scaffold it just falls over. Then I remember the first time we saw a model look at what processes were running and be like oh that one's me, that was cool. They really failed on that one before; they used to kill their own process while they were doing other things. So yeah, it's been interesting to watch that go up over time and I feel like that was very predictable that this was where things were going. I think the other thing that we learned about scaffolding mostly was it's hard to make your agent harness really good on a diverse set of tasks.
把它做差很容易,但如果只针对一个狭窄的任务分布,你可以获得更多改进,但可能在其他任务上表现更差。所以当我们看到令人印象深刻的结果时,我会想他们做了多少任务特定的脚手架迭代,因为这确实有很大影响。我们在所有任务上只使用一个相当简单的脚手架,这本身就造成了相当大的差异。通常,那些花里胡哨的东西并没有比基本方法好多少:给它 bash,往提示里追加内容,也许再加点压缩。这对你的听众来说可能不是新闻,但我们看到了推理算力回报的显著增长。要确信一个新模型在基本智能体脚手架下无法完成任务,我们通常认为需要花费至少几百到几千美元,才能确信它真的达到了平台期,而不仅仅是时间不够。
It's easy to make it bad and you can get much more improvement if you're targeting a narrow distribution of tasks, but you probably then do worse on other tasks. So when we see impressive results, I think about how much task-specific scaffolding iteration was done, because that really makes a big difference. The fact that we're using just one fairly simple scaffolding over all tasks makes a fairly large difference. Generally, things with more bells and whistles haven't done that much better than the basic approach: give it bash, append things to the prompt, and maybe some kind of compaction. This isn't news to your audience probably, but we've seen dramatic increases in returns from inference compute. To be confident that a new model can't complete a task given a basic agent scaffold, we generally think about needing to spend at least hundreds or low thousands of dollars to be confident it's really plateauing, not just that it didn't have enough time.
关于脚手架再详细一点:首先,有信用分配问题。你可以在脚手架中加入各种花里胡哨的东西,比如压缩,这是一个相对较新的创新。当你改变一些脚手架时,你说性能在整个模型-任务对套件上发生了变化。这有多大影响?你看到了哪些失败模式?你尝试过什么?
On the scaffolding stuff in a bit more detail: first, there's the credit assignment problem. You can put all these different bells and whistles in the scaffold, like compaction, which is a relatively recent innovation. When you changed some of the scaffold, you said performance changed across this suite of model-task pairs. How much of a difference does it make, and what kind of failure modes do you see? What have you tried?
我们尝试的很多东西都与给模型更多信息或更直接的工具访问权限有关。对我们来说重要的一件事是告诉智能体它花了多少时间,以及它在 token 预算中用了多少 token。没有这个,智能体常常过早提交解决方案,或者对应该花多长时间没有校准。这很有趣,因为人类对此有很多隐含信息。当你的经理给你一个任务时,有很多关于你应该花多长时间的隐含信号。例如,他们可能会随口说:“是啊,我很期待今晚看到你的结果之类的。”所以你知道你需要在接下来几小时内完成初稿,而不是花几天打磨。但智能体只有提示和上下文;它们没有关于你实际期望的启发式信息。所以对我们来说,当有 token 预算时,告诉智能体“你已经用了 10 万个 token,占你 token 预算的 1%”会帮助智能体知道。
A lot of the things we've tried are related to giving the model more information or more direct access to tools. One thing that has been important for us is telling the agent how much time it spent and how many tokens it used out of its token budget. Without that, agents will often submit their solution way too early or are not calibrated on how long they should spend. It's interesting because humans have a lot of implicit information about this. When your manager gives you a task, there are implicit signals about how long you should spend. For example, they might offhand say, 'Yeah, and I'm excited to see your results tonight or something.' So you know you need to get a first draft done in the next couple hours, not spend days polishing. But agents just have their prompt and context; they don't have heuristics or information about what you actually expect. So for us, when we have a token budget, telling the agent, 'You've used 100,000 tokens so far, which is 1% of your token budget,' helps the agent know.
对于一个模型,我们有它解决任务的可能性,x 轴上是不同时间跨度的不同任务。如果我没理解错,你让大约八个智能体尝试任务,并对任务进行分桶,因为不同组中任务数量不同,所以你做了归一化。数据看起来有点像 S 曲线。我想理解使用它的直觉。是否存在薄尾问题、斜率类型等?另外,你在论文中提到,文献中有一些心理测量学的直觉。
For a model, we have how likely it is to solve a task, and on the x-axis we have the different tasks at different time horizons. If I understand correctly, you have about eight agents attempt the task, and you bucket the tasks because there are different amounts of tasks in different groups, so you normalize that. The data looks a bit like an S-curve. I'm trying to understand the intuition for using it. Were there issues with thin tails, the type of slope, etc.? Also, you mentioned in the paper that there was some psychometric intuition from the literature.
是的,这很像项目反应理论。你可以做一个完整的贝叶斯分析,同时估算任务难度参数和模型能力参数。总的来说,我有一条原则:如果你不能在图上看到你感兴趣的东西,就要非常警惕复杂的统计。你应该能够把它画出来,看着它说:“哦,是的,大概就是这样。”遵循这个原则很难出大错。你可以用各种深奥的方法来拟合这个,但我不相信任何比肉眼观察图表更可靠的东西,比如“到这里,它基本上完成了所有任务,而从这里开始,它真的没做多少,所以大概在这个位置。”它看起来像逻辑函数,这和你让人类完成考试题目时做的是一样的。我们搞砸的具体事情是加了一个正则化项来惩罚逻辑函数的斜率,这在数据较多时没有影响,但当我们开始饱和时,正则化使它比应有的更平缓,从而推高了 50% 点。所以,永远要在图上查看你的数据。好习惯。
Yes, it's pretty similar to item response theory. You can do a whole Bayesian analysis with imputing task difficulty parameters and model ability parameters simultaneously. In general, I have a policy of being very wary of complicated statistics if you can't see the thing you're interested in on a graph. You should be able to plot it and look at it and say, 'Oh yeah, it's about that.' It's hard to go too far wrong with that principle. There are various arcane things you can do to fit this in different ways, but I don't trust anything much more than eyeballing the graph and saying, 'Up to here, it's basically doing all the tasks, and after here, it's really not doing very many, so it's somewhere here.' It looks like logistic, and this is what you would do for having humans complete questions on an exam. The specific thing we messed up was having a regularization term penalizing the slope of the logistic, which didn't have an effect in the regime with more data, but as we started to saturate, the regularization made it a bit shallower than it should have been, pushing the 50%. So always look at your data on a graph. Good practice.
哦,这很有趣。
Oh, that's interesting.
我问这个的原因是,我记得你后来发了一篇说明,说如果用了固定斜率的逻辑回归,交叉验证效果可能会更好,而 50% 近期时间线实际上会上升约 35%。所以这些差异相当显著。
And the reason I ask is I think you published a later note saying that had you used or if you used a fixed slope logistic it might cross validate better and the 50% recent horizons would actually be up by about 35%. So these are quite significant differences.
与误差线相比,这些差异很小。是的,误差线大约是最近模型的两倍左右。所以,基本上,误差线是真实的,这些数字不是。
They're small compared to the error bars. Yeah, the error bars are like 2x on either side or something from the most recent model. So yeah, it is I mean basically yeah you should be like the error bars are real. These numbers are not.
对我们来说,这触及了有时很难的科学传播问题——我们确实对这些具体数字有很大的不确定性。所以,30% 的差异对我们来说其实相对较小,比如如果我们用了稍微不同的任务分布,那可能会导致 2 倍左右的差异。
For us, and I think this kind of gets at for us you know sometimes difficult like science communication questions, where like we really do have a lot of uncertainty about you know individual the individual numbers here, and so you know like you know 30% difference is actually like relatively small for us relative to you know like for example if you know we had used a somewhat different distribution of tasks, that's likely to cause maybe 2x differences or something.
另一个关键问题是,为什么把 50% 作为标题数字?因为你想,如果我要写代码——这里有个显而易见的问题我们稍后会谈到——这个数字被用来论证软件工程师可能很快失业,因为我们可以自动化他们的工作。但 50% 的可靠性根本不够,对吧?我认为需要 80% 或 90%。
The other million dollar question is why report 50% as the headline number? Because if you think about it, if I want to write some code, because the elephant in the room here that we'll get to is this is being used as an argument to say that software engineers might be unemployable soon because we can automate what they're doing. But 50% reliability isn't really in the ballpark, is it? I think it would need to be what like 80, 90%.
所以我认为我们应该区分:特定任务上的可靠性——比如反复尝试这个任务,成功比例是多少——与任务上的成功概率——比如给定该任务分布的人类时间,你能不能完成这个特定任务。实际上,对于几乎所有任务,模型要么每次都成功,要么每次都失败。有些任务它们不可靠,但主要是:在这个人类时间水平上,有多少比例的任务是模型基本总能成功或基本总失败。在具体情况下,这可能比仅仅知道人类大概花多长时间更具可预测性。所以,时间线百分比数字与你试图让模型完成大致相同长度的任务时的成功率之间,不一定有很好的对应关系,因为当你这样做时,你会选择你希望模型成功的任务。它提供了关于模型能完成多少事情的信息,但不太涉及“我是否处于不断给它任务却不知道它会成功还是失败的状态”。
So I think we should distinguish here between reliability on a particular task—like if you attempt repeatedly this task, what fraction of times you succeed—versus probability of success on a task, like given that you know the human time of that distribution of tasks, can you do this particular task. So when we look at it, actually for almost all the tasks, models either succeed every time or fail every time. There's some tasks for which they're unreliable, but it's mostly a case of like what fraction of tasks at this human time level are in the like models basically always succeed or basically always fails. And that may be more predictable in any specific case than just knowing how roughly how long it takes humans. So I think it's not necessarily a great translation between the time horizon percent number and like if you are trying to get models to do a task of roughly that length, what fraction of the time does that succeed, because you can when you're doing that you will pick tasks that you want models to succeed at. It is information about what fraction of the things will they be able to do, but it's slightly less about like oh am I going to be in this regime where I keep giving it things and then I don't know whether it's going to succeed or fail.
总的来说,我不太清楚我们应该关注什么样的正确数字或正确可靠性水平。一种观点是,我们可能应该关注 10% 的可靠性,因为一旦模型能在 10% 的时间里完成某些任务,我们预计 AI 公司就能从这些难度或类型的任务中获得足够的正向奖励信号,从而更容易从 10% 引导到 90% 或 95% 甚至更高的可靠性。我认为这在很大程度上取决于你感兴趣的问题。所以,我认为较低的可靠性更可能告诉你事情的发展方向,可能是一个领先指标。而较高的可靠性则更直接地告诉你,我日常能用这个模型做什么。但正如我们讨论过的,还有其他主要的不确定性来源影响我们的理解,比如任务分布、高上下文与低上下文工作的差异。实际上,获得高可靠性时间线的良好估计要困难得多,因此我们的误差线会大得多。所以我认为这是一个弱点,或者我对更高可靠性时间线非常感兴趣,但测量起来要困难得多,因为如果你在 100 次中只有一次失败,那么对于这次失败是噪声还是真实的,你会有很大的不确定性。
It's overall to me pretty unclear like what the kind of right number or right level of reliability we should be interested in is. So one argument you could make is maybe we should be interested in something like 10% reliability, because once models are able to do some set of tasks 10% of the time, we'd expect AI companies to be able to get enough positive reward signal on tasks of that difficulty or of that type such that then they can more easily bootstrap from 10% up to like 90 or 95 or higher reliability. I think a lot of it basically depends on the question you're interested in. So yeah, I think about it as being like lower reliability is more likely to tell you something about where things are headed, might be something like a leading indicator of progress. And then higher reliability tells you something more closely about like what can I use this model for in my day-to-day or something, but as we've talked about, there are already other major sources of uncertainty that affect our understanding, like the task distribution, the difference between high context versus low context work. Just actually getting good estimates of high reliability time horizons is substantially harder, and so our error bars would just be much larger. And so this is I think a weakness or like I'm very interested in much higher reliability time horizons, but it's substantially more difficult to measure because if you only have one failure out of a hundred or something, you have a lot of uncertainty if that failure is noise or is real.
关于统计有效性有道理,在尾部数据更稀疏,估计也更困难。但这很有道理,不过你刚才提到,如果我们得到 10%,那是一个信号。但我回想我们之前说的,也许它们可能因为错误的原因给出正确答案,也许我们应该谈谈评估。这些任务很有趣,因为它们是可验证的,没有与其他智能体的交互,环境相对静态,对单个错误的惩罚很弱,等等。而且,在大多数情况下,结果是二元的,有时是连续的,然后你把它转换成二元。所以这是一个相当自动化的设置,但你们有没有深入挖掘其中的怪异之处?比如,你们有没有直觉,它们是否因为正确的原因做正确的事,或者有很多假阳性,它们做了事但结果是退化的?
There is an argument for statistical validity and on the tails they're sparser and increasingly estimated and so on. So that makes a lot of sense, but you made a comment about oh you know it's a signal if we get 10%. But I was kind of thinking back to what we were saying earlier that maybe they could give the right answers for the wrong reasons, and maybe we should talk about the evaluation. Right, so these are quite interesting tasks in the sense that they are verifiable, there's no interaction with other agents, they're relatively static environments, weak penalties for single mistakes, and so on. And in a sense, these are like, I think as well, in most cases they have a binary result, sometimes continuous, and then you convert it into a binary. So it's a fairly automated setup, but are you digging into weirdnesses there? Like are you do you have a bit of an intuition on are they doing the right thing for the right reasons, or are there lots of false positives where they did the thing but it was kind of degenerate?
我认为这是我们文化中最喜欢的方面之一:我们有非常深入的数据审视文化。我们举办披萨派对,只是通读智能体的转录。对我们来说,很多这项工作发生在开发任务本身的过程中。
I think this is one of the aspects of our culture that I like the most is that we have a very deep culture of looking at our data. We have pizza parties where we just read through agent transcripts. A lot of this work for us happened when developing the tasks themselves.
我们经常看到假阳性和假阴性。例如,某个任务可能没有配置允许互联网访问,但结果发现它需要互联网才能完成,或者文件没有正确上传到容器中。我们也看到了奖励破解的案例。很多工作用于强化评分函数,以减少假阳性,尽管我们仍然看到智能体在奖励破解,甚至可能越来越多。
We would see very often both false positives and false negatives. For example, a task might not be configured to allow internet access, but it turns out it requires internet access to complete, or the file wasn't uploaded properly to the container. We have also seen cases of reward hacking. A lot of work went into hardening the scoring functions to make it more difficult to see false positives, although we still do see agents reward hacking, maybe even increasingly so.
特别是对于 ReBench 任务,我们有特定标准:它不能在没有迭代的情况下被解决。如果智能体可以直接写出解决方案,那就没意思了。我们通常有一个质量保证流程,让人类去执行任务,或者至少近似执行。他们可能会快速完成某些部分,但要检查一切是否按预期工作,你不能直接猜出答案或轻易作弊,指令也要清晰。所以仍然会有一些问题,但总体来说我们相当仔细地检查过。可能更难看出它们是否以退化方式解决,因为与训练数据的相似性是一个因素。也许它们看起来在迭代并使用合理的解题策略,但实际上实验室有非常相似的任务分布,我们没有意识到这个任务实际上有多在分布内。可能有一些这种情况。
For the ReBench tasks in particular, we had specific criteria: it shouldn't be able to solve them without iteration. If an agent can just write out the solution straight out, that is not interesting. We generally had a quality assurance process for tasks where humans do it, or at least approximately do it. Maybe they speedrun some of the bits, but checking that everything works as expected, you can't just guess the answer or easily cheat, and the instructions are clear. So there will still be some of these issues, but generally we've looked at them reasonably carefully. It's maybe harder to see if they are solving it in degenerate ways because similarity to training data is one thing. Maybe they seem to be iterating and using reasonable problem-solving strategies, but actually the lab had a really similar distribution of tasks and we don't realize how in-distribution this task actually is. There's probably some of that going on.
许多公众讨论中的人,比如昨晚 Will MacAskill 在 Sam Harris 播客上,都在谈论 AI 风险,认为可能一两年内 AI 模型就能完成人类需要一两个月才能完成的事情。目前我认为没有任何超过 30 小时的任务被人类评估过。然后我们进入这个问题:如果公众讨论的是图表中最不受约束的区域,我们是否在进行外推?我们谈论 AI 可能能完成需要一两个月的事情,这有多合理?
Many folks in public discourse, like Will MacAskill on the Sam Harris podcast last night, were talking about AI risk as maybe in a year or two we'll have AI models doing things that take a month or two months for a human. At the moment I don't think there are any tasks over 30 hours that have been evaluated by humans. Then we get into this question: if the public discourse is talking about the least constrained region of the graph, are we getting into extrapolation? How legitimate is it for us to talk about AI might be able to do things that take a month or two months?
预测事情很难,尤其是关于未来。人们可以有很多不同的视角或先验信念。我认为对于我们将走向何方,有各种合理的判断。但当然,做那种预测与谈论通过具体方法论收集并已有结果的数据是不同的活动。我可以说的一个事情是:我在某种程度上对原始趋势线的保持程度感到惊讶。我确实认为这是一些证据,至少对我来说,这是关于未来走向的可靠证据。我的一位同事最近写了一篇短博文,谈论这种图上直线的直觉。很多人对进步如何发生有不同的模型。但如果你在相当长的时间内观察到一个稳健的趋势,尤其是在 AI 中,进步在相当程度上是系统性的,我确实会重视这个趋势的延续。但也有许多原因可能导致它不延续。
Predicting things is hard, especially about the future. There are a lot of different perspectives or prior beliefs people can have. I think there's a wide range of reasonable judgments about where we're going to be. But of course, doing that kind of prediction is a different activity than talking about data that has been collected with a concrete methodology and where we already have the results. One thing I can say: I have been surprised to some extent by how well the original trend line has held up. I do think that is some evidence, at least for me, it's decent evidence about where things will go. A colleague of mine recently wrote a short blog post talking about this intuition of straight lines on graphs. Lots of people have different models of how progress is happening. But if you have observed a robust trend over a decent period of time, especially in AI where progress is to a decent extent systematic, I definitely put weight on that trend continuing. But there are a bunch of reasons why it might not.
我认为软件工程是一个规格获取问题。这非常困难。我们事先不知道我们在构建什么。你构建一些软件,第一个版本有 bug,用户使用它,你发现很多边缘情况,然后你修订它。经过第 10 次修订后,你创建了可爱的表示、抽象和粗粒度化。你对自己说:如果我能扔掉所有代码,我可以快 10 倍地构建它,因为我现在确切知道该做什么。我已经实施了智能,我找到了领域的轮廓。现在这只是一个自动化问题。从某种意义上说,这种污染问题让我担忧,因为当人们使用 Claude Code 时,他们在获取你的数据。有人在编写内核编译器,做各种事情,而 Anthropic 正在吸收这些。在某个时刻,它变成了一个自动化问题。如果你把一个任务放进去,它本质上是一个头部查询——在分布的模式中,经常使用——Claude Code 会给你规格,因为它已经从别人那里获取了。如果你给它一个长尾的东西,那么作为开发者,你必须在提示中给出规格,这又是一个自动化问题。自动化非常容易。
I think software engineering is a specification acquisition problem. It's very difficult. We don't know ahead of time what we're building. You build some software and the first version is buggy, your users use it and you find lots of edge cases, then you revise it. After the 10th revision, you've created lovely representations and abstractions and coarse grainings. You say to yourself: if I could throw all the code away, I could build it 10 times quicker because I know exactly what to do now. I've enacted the intelligence, I've found the contours of the domain. It's now an automation problem. In a sense, this contamination thing is a concern for me because when people use Claude Code, they're taking your data. There are people out there writing kernel compilers and doing all sorts of things, and Anthropic is just sucking that up. At some point it becomes an automation problem. If you're putting a task in there which is essentially a head query—something in the mode of the distribution, used all the time—Claude Code will give you the specification because it's already been taken from other people. If you give it something on the long tail, then you as the developer have to give it the specification in the prompt, and again it's an automation problem. Automation is really easy.
那么,这就是正在发生的事情吗?你认为时间线的缩短是否仅仅是因为从其他做类似任务的人那里获取了所有这些知识?
So is that what's happening? Do you think that the increase in the timelines could just be explained by the acquisition of all of this kind of knowledge from other people doing similar tasks?
是的。我认为这是解读我们当前处境的一个非常核心的问题。首先要说的是:这很难知道。这是一个大问题。所以我认为我们应该保持相当程度的不确定性,或者把收集到的每一条证据都当作某种证据。我们确实看到模型在具有非常清晰反馈信号、规格极其明确的任务上表现更好,比如在软件工程领域,如果你写好了规格说明,就可以迭代并不断优化。我们也看到模型在所谓的“混乱任务”上表现更好,这些任务我们没有提供非常清晰的规格。我们最近创建任务的一种方法是尝试创建更混乱、规格更不明确的任务,基本上就是放宽自动评分约束。所以我们不需要写一个非常清晰、定义明确的评分函数;只需对模型写几句话:“嘿,构建这个大型软件。我不会告诉你我具体想要什么,但它必须很好。”所以模型需要自己弄清楚到底要构建什么。就我个人而言,我认为我们没有像对时间范围那样系统地收集这些结果,部分原因是现在这些任务的评分是定性的。但我的印象是,模型在这些类型的任务上比在给出清晰规格时表现更差,但它们一直在以类似的速度改进。还有其他一些证据来源,但对我来说这是主要的一点。
Yeah. I think it's a super central question for interpreting where we're at. First thing I'll say: it's hard to know. It's a big question. So I think we want to have a decent amount of uncertainty, or we want to take each individual piece of evidence that we've collected as some evidence. We do see models performing better on tasks that have really clear feedback signals, that are extremely well specified, in domains like software engineering where if you have written out a spec, you can iterate and grind against that. We also see models performing much better on so-called messier tasks where we haven't already provided this really clean spec. One approach we've taken for creating tasks recently is to try and create messier tasks that are less well specified, basically relaxing this automatic scoring constraint. So we don't need to write a really clear, well-defined scoring function; just writing a couple sentences to a model: "Hey, build this large piece of software. I'm not going to tell you exactly what I'm looking for, but it needs to be good." So the model needs to figure out what actually to build. Personally, I think we don't have these results collected as systematically as we do for time horizon, partially because scoring is qualitative now for these tasks. But my impression is that models are worse on these types of tasks than when you give them a clean spec, but they have been improving at a similar rate. There are some other sources of evidence, but that's one major piece for me.
混乱任务意味着模糊性,这绝对是一件常见的事情。我们做 vibe coding,从一个模糊的规格开始,然后现实会反馈,我们找到问题的轮廓,不断告诉 Claude Code:“哦,实际上不,别那样做,这样做,这样做,这样做。”然后我们找到问题的形状,它随着时间的推移变得更好。但问题是,Claude Code 的源代码昨天泄露了,我朋友,一个非常优秀的软件工程师,看了一下说:“是的,我不想说 Anthropic 的坏话,但显然代码分解得不好,控制流到处都是,有点乱。”他说如果是他的实习生写的,他会不高兴。我不知道他们是否看过代码。昨天有人开玩笑说,现在可能比之前有更多人类在看 Claude Code 的代码。但问题是,总是存在模糊的领域,而 LLM 用更多资源做更多事。智能是用更少资源做更多事,而 LLM 用更多资源做更多事,因为规格、智能来自人类监督者。所以当你给它们模糊性时,你只会得到大量分解不好的代码。所以从某种意义上说,这是否让评估变得更难?因为它可能给出了你想要的答案,但同时制造了一个混乱的烂摊子。
Messy task means ambiguity, and this is absolutely a common thing. We do vibe coding, we start off with an ambiguous specification, and then reality pushes back, we find the contours of the problem, and we keep telling Claude Code: "Oh actually no, don't do that, do this, do this, do this." Then we find the shape of the problem, and it gets better over time. But the thing is, the source code for Claude Code leaked yesterday, and my friend, a very good software engineer, was looking through it and said: "Yeah, I don't want to bad talk Anthropic, but apparently it's not very well factored, control flows all over the place, a bit messy." He said if his intern did it, he would have been displeased. I don't know whether they've even looked at the code. Someone joked yesterday that probably more humans are looking at the code for Claude Code now than before. But the thing is, there are always areas of ambiguity, and LLMs do more with more. Intelligence is more with less, and LLMs do more with more because the specification, the intelligence comes from the human supervisor. So when you give them ambiguity, you just get a lot of unfactored code all over the place. So in a sense, does that make it harder to evaluate? Because it might give you the answer you're asking for, but it's creating an unfactored mess at the same time.
是的,我认为这是一个非常有趣的问题。我有时会想到的一个类比是编译器。我出生在编译器发明之后,但我有一些印象:在编译器出现之前,人们手工制作漂亮的汇编代码,非常高效,每个寄存器都用上,不浪费内存。然后编译器出现了,现在它们吐出垃圾机器码,大量的汇编代码没有优化,占用大量内存,速度慢,等等。但事实证明,能够自动化大部分过程是非常有用的。人们对软件工程的现状有不同看法,但我认为可以合理地说,编译器是让我们达到现在水平的一个极其重要的部分。我不清楚模型输出对人类来说难以阅读和使用的代码是否一定意味着它对 AI 来说也难以阅读、使用和构建。我认为肯定有一些原则会转移或对模型有用。显然,有些可怕的意大利面条式代码连模型都读不了。我以前就写过一些。
Yeah, I think this is a super interesting question. One analogy I think about sometimes is compilers. I was born after compilers were invented, but I have some impression that pre-compilers, people were handcrafting beautiful assembly that was extremely efficient, every register used, not wasting memory. Then compilers came along, and now they're spitting out garbage machine code, a gigantic amount of assembly that is not optimized, takes so much memory, slow, whatever. But it turns out that being able to automate a large fraction of the process has been very useful. People have disagreements about the state of software engineering, but I think it's reasonable to say that compilers have been an extremely important part of getting us to where we are. It's not clear to me that models outputting code that is bad for humans to read and use necessarily means it'll be bad for AIs to read and use and build on. I think there are definitely principles that will transfer or be useful for models. Obviously, there is some horrendous spaghetti code that not even models would be able to read. I've written some of that before.
但我认为我们之间可能有一个不同的视角,即重要的是模型以人类的方式解决问题,还是它们能解决问题就行。我确实认为模型在编写干净、良好的代码方面变得更好可能非常重要。这对我来说似乎相当合理,但至少我不确定。
But I think there may be a somewhat different perspective between us around whether the important thing is that models are solving problems in the way that people are solving them, or that they're solving them at all. I do think it might be really important for models to get way better at writing clean, good code. That seems pretty plausible to me, but I'm not certain of that at the very least.
我有几个非技术朋友在尝试 vibe coding,他们给我看他们的应用。那是一个巨大的仪表盘,有上百万个按钮,同一个功能实现了多次,还没有数据库等等。在某个时候,当 Meta 的七级工程师做 vibe coding 时,效果很棒,因为他们知道如何结构化。规范中的一些东西很重要:是无服务器还是多租户,如何做 Google 认证,什么类型的数据库,是虚拟机还是无服务器?你做出一系列决策,然后人们开始使用你的应用,你就无法轻易回退了。即使你有神奇的自动化机器,由于 CI/CD 测试等复杂性,你也无法轻松回滚。所以在某个时候,你需要一个有能力的人,真正清楚需要做什么。
I've got several friends who are not technical who are experimenting with vibe coding, and they show me their application. It's this big dashboard with a million different buttons, implementing the same thing multiple times, no database yet, and so on. At some point, when a level seven engineer from Meta does vibe coding, it's amazing because they know how to structure things. Some of these things in the specification are just important: is it serverless, is it multi-tenanted, how do we do Google authentication, what kind of database, is it a VM or serverless? You make a series of decisions, and then you've got people using your application, and you can't really wind that back. It doesn't matter if you've got the magical automation machine because you can't easily roll that back due to lots of complexities like CI/CD testing. So at some point, you need a competent human who actually has a pretty good idea of what needs to happen.
我觉得我们可能都有过这样的经历:我们的一位工程师对 Claude Code 非常兴奋,并告诉每个人,当遇到信息问题时,我们应该让 Claude 来解决。对他来说这没问题,因为智能体知道骗不了他。但我有个问题,他说“去问 Claude”。我问如何设置 AWS 配置,Claude 查看了 Slack 然后说“你应该这样做”。结果发现那是别人犯的错误。我抱怨说我的 Claude 比他的笨。所以在你提问时使用的语言中肯定存在观察者效应。
I feel like we've probably all had this experience where one of our engineers got super excited about Claude Code and was telling everyone that when we had info problems, we should just ask Claude to solve it. This went fine with him because the agents knew they couldn't fool him. But I had some question, and he said, 'Just ask Claude.' I asked how to set up my AWS configure, and Claude looked on Slack and said, 'You should do this thing.' It turned out that was a mistake someone else had made. I complained that my Claude was dumber than his. So there's definitely an observer effect in the language you use to ask for things.
就你能衡量的程度而言,判断代码质量是否足够的一个测试是:你能构建一个大型应用吗?如果代码很糟糕,但他们构建了一个极其复杂且运行良好的东西,那么说明有效。你认为糟糕代码不好的主要原因是你实际上无法构建那么复杂的东西,因为会有 bug 和复杂性。所以如果我们看到模型构建了确实有效且非常复杂的东西,那么它们具体如何做到的就没那么有趣了,但这可能对人类可观察性不利。这也涉及到我们期望模型在明确指定的任务上表现更好,那些可衡量的东西会提升,但这是否是我们真正想要的就不那么清楚了。
To the extent that you can measure this, one test of whether code is high quality enough is: can you build a big application? If the code is disgusting but they've built an incredibly complex thing that works great, then something is working. The main reason you expect bad code to be bad is that you can't actually build something that sophisticated because of bugs and complexity. So if we see models building things that do work and are very complicated, it's less interesting exactly how they're doing it, but it may be bad for human observability. It also gets into the idea that we expect models to be much better at well-specified tasks, and those measurable things will go up, but whether that's what we actually wanted is less clear.
我想问题是:这里最有力的可辩护主张是什么?很多人在公开讨论中说软件工程智能每七个月翻一番。Dario 最近发布了一篇博文《技术的青春期》,他对此非常乐观,尽管他内部的一些研究人员发表了更怀疑的研究。但将其解释为更狭窄的东西是否更公平,比如在低上下文、明确指定、可自动检查的技术任务上的自主成功正在快速提升?
I guess the question is: what is the strongest defensible claim here? A lot of folks in public discourse say software engineering intelligence is doubling every seven months. Dario released a blog post recently, 'The Adolescence of Technology,' and he was super bullish about it, even though some of his own internal researchers published far more skeptical research. But is it fairer to interpret it as something narrower, like autonomous success on low-context, well-specified, automatically checkable technical tasks is rising fast?
是的,可爬山、容易检查的任务,你可以从终端或舒适地在语言界面或文本输入/输出界面中完成。
Yeah, hill-climbable, easily checkable tasks that you can do from a terminal or comfortably in a language interface or text input/output interface.
我认为有一个问题:我们是否关心我们 99% 确信的陈述?我们可能也对只有 1% 确信的陈述感兴趣。如果 2026 年底有 1% 的概率发生疯狂智能爆炸,而文明的命运取决于此,那么即使你 99% 确信它不会发生,知道这一点也很有趣。我们关心只有 1% 概率的事情,比如 1% 概率的绝症。所以我认为我们对整个分布感兴趣:哪些事情我们可以确认,哪些可以排除,哪些事情即使不太可能也有合理的解释,但现在可能进入了考虑范围。
I think there's a question of whether we care about statements we're 99% confident in. We may also be interested in statements we're 1% confident in. If there's a 1% chance of a crazy intelligence explosion at the end of 2026, and the fate of civilization depends on how that goes, that is interesting to know even if you're 99% confident it won't happen. We care about things that are 1% chance, like a 1% chance of a terminal illness. So I think we're interested in the whole distribution: what things can we rule in, what can we rule out, and what things have a reasonable story even if unlikely, but maybe now in the realm of consideration.
你大概看到了 Anthropic 的 Carlini 那篇论文,他们用一群智能体创建了一个编译器。从某种意义上说,Jeremy Howard 告诉我这基本上是一个风格迁移问题,因为规范、测试和代码都在网上,它可以反复迭代直到成功,然后就能运行 Doom 之类的。但这是一个极其复杂软件的例子。我常开玩笑说,AGI 的最佳标志是它能构建出像 Linux 操作系统这样的东西。和我们之前讨论的一致,这里存在一个规范问题。到了某个点,没有任何人类能理解或创建 Linux 操作系统的规范。随着时间的推移,我们只是逐步构建了这个系统,因为我们的大脑有限。我们迈出一步,现实反馈回来,我们再迈一步,就这样不断推进,构建规范。但让人类去指定一个需要四个月才能完成的任务意味着什么?我们之所以创造敏捷软件开发方法论,正是因为它超出了我们的认知范围。所以这难道不是一个先有鸡还是先有蛋的问题吗?在我看来,我认为我们无法指定如此复杂的任务,因此 AI 也无法完成。我想到了一个类比,就是 CEO 在公司中的角色。也许 Beth 更适合回答这个问题。但 CEO 会提出公司愿景,然后简洁地传达给下属高管。如果他们是优秀的 CEO,公司也高效,那么公司就能将这些非常简洁的信息——远非完整规范——转化为符合期望的成果。所以我们确实有例子表明,人们能够指定某些任务,然后判断这个可能需要成百上千人年才能完成的大型任务是否成功。这是一个启发性的直觉:语言本身具有足够的表达能力,能让我们达成合理的理解。显然存在大量边缘情况,而且 CEO 也常常无法让公司完全按他们的意愿行事。但我认为 AI 至少有可能完成这类长期任务。
You probably saw the Carlini paper from Anthropic where they got a swarm of agents to create a compiler. In a sense, Jeremy Howard said it's basically a style transfer problem because the specification is online, the tests are online, and the code is online, and it could just iteratively do the thing until it worked, and then it could run Doom and all that. But that is an example of an extremely complicated piece of software. I often joke that the best mark of AGI is when it could build something like the Linux operating system. And in line with what we were saying before, we have this specification problem. It gets to the point where no human could understand or create the specification for the Linux operating system. Over time we've incrementally built this thing because our brains are limited. We take one step, reality pushes back, we take another step, and we keep going, building the specification. But what would it mean for a human to specify a task that could take four months? The whole reason we created agile software development as a methodology is because it's inconceivable, outside our cognitive horizon. So isn't that a bit of a chicken-and-egg problem? In my mind, I don't think we could specify a task of that complexity, therefore the AIs wouldn't be able to do it. One analogy I think about is the role of a CEO at companies. Actually, maybe Beth is better to answer this. But CEOs come up with a vision for where they want the company to be, then they communicate that concisely to executives who report to them. If they're a good CEO and the company is effective, the company can take this very concise information—not the full spec at all, not even close—and turn it into something aligned with what they're looking for. So we do have examples of people being able to specify some task and then judge whether this very large task, which may take hundreds or thousands of person-years to complete because it requires many people working over a long time, has succeeded or failed. That's one motivating intuition: language has built-in expressivity to have some reasonable understanding. Obviously there are tons of edge cases, and often CEOs aren't able to get their companies to do what they want. But I think it's at least plausible that AIs could do these kinds of long tasks.
我认为说“我们无法指定超过四个月的任务”显然过于绝对了。有些数值指标是可以自动检查的,并且需要四个月,比如“将 nanoGPT 的浮点运算次数或运行时间降低这么多”。你可以大致看出人类需要多长时间,并且有合理的衡量方法。对于其中一些任务,你可能最终需要补充一句“还需要人类检查你是否大致做了正确的事,而不是钻了空子”。还有一些任务并非根本上无法指定,只是作为评估的一部分成本太高。比如 J 喜欢举策划婚礼的例子:你可以合理评估婚礼是否组织得当,但我们不可能为每个新模型抽取三个样本——我们没有那么多婚礼可办,而且如果结果完全失败也挺遗憾的。所以有些任务你可以检查少量样本,或者写下评估方法,但就是不想反复执行。软件可能也类似:评估就像“这家公司雇你开发工具,之后还会再雇你吗?”即使他们一开始也不知道软件具体需要做什么。但这并不意味着你不能给个分数,看你在该任务上是否与人类或软件咨询公司表现相当。
I would say that saying something like 'we can't specify tasks that take more than four months' seems obviously too strong. There are numerical things that are automatically checkable and take four months, like 'get the nanoGPT flop count runtime down this much.' You can see roughly how long they take humans, and there's a reasonable way to measure them. Maybe for some of these things you end up having to say 'and also this human checks that you did roughly the right thing and didn't hack the solution.' Then there are other things which are not fundamentally unspecifiable, they're just too expensive to do as part of an evaluation. For example, J likes giving the example of planning a wedding. You can get a reasonable estimation of whether it was a well-organized wedding, but we can't really take three samples for each new model that comes out—we don't have enough marriages happening to do that, and it's a bit sad if it turns out to be total trash. So there are things where you could check a few samples or write down how you would evaluate it, but you just don't want to run that many times. Probably similar with software: the evaluation is like, 'would this company that contracted you to build this tool hire you again?' Even they don't know exactly what the software will need to do when starting out. But that doesn't mean you can't have some score for whether you did comparably to a human or a software consulting firm on that task.
截至目前,时间线估计中的主要不确定性驱动因素是什么?你随时间做了一些更新。最初版本有 170 个任务,现在是 228 个。还有较大任务上采样稀疏的问题。我们能从中解读出什么?
As of today, what are the main uncertainty drivers in the time horizon estimate? You've updated a bit over time. The original version had 170 tasks, now it's 228 tasks. There's the issue of sparse sampling on the larger tasks. What can we read into this?
我认为仍然是任务分布的问题。我们更有信心的是,至少在部分容易爬坡的任务分布上,模型确实有相当长的时间跨度,比如软件复制——你的得分是测试通过率,分数是连续的,且功劳归属容易。还有一些优化任务,比如“让这段代码跑得更快”或“让这个模型学得更好”。
I think it's still this task distribution. We feel more confident that models do have pretty long time horizons on at least some distribution of easily hill-climbable tasks, like software replication where your score is the percentage of tests passed, the score is continuous, and credit attribution is easy. Also some optimization tasks like 'make this code run faster' or 'make this model learn better.'
也许我们应该请出 Daniel Kokotajlo。你知道,在他的 AI 2027 文章中,他上过节目,到处宣讲,那篇关于时间线的文章影响巨大,而且他直接引用了你的工作。我想问的是:你觉得在公共讨论中,这是不是被过度解读了?你如何看待这种解读在推断和时间线方面的意义?
Maybe we should bring in Daniel Kokotajlo. You know, in his AI 2027 piece, he's been on the show, he's been doing the rounds, hugely impactful piece talking about timelines, and he cites your work directly. I guess the question is: do you think in the public discourse this is being overread? How do you think about the interpretation of this in terms of extrapolations and timelines?
我的意思是,肯定有些人在过度解读。事情肯定被过度炒作了。你在 Twitter 上看到一堆人在说疯话,还有人甚至误解了它衡量的是什么,普遍忽略了各种注意事项。Daniel Kokotajlo 相当理性,而且以概率的方式思考问题。我觉得他可能在某些事情上比我更有信心,而我则更不确定。而且我认为一些 AI 未来项目的模型对指标时间范围的敏感度超过了应有的程度。我不觉得这很疯狂。它有可能确实捕捉到了一种会迁移到其他类型任务上的趋势。这是一个我们应该思考的故事:如果那是真的会怎样,如果那是真的会发生什么?但也有可能它不会,这些事物会分道扬镳。我相当贝叶斯或务实:我们想要做出一些预测,并对未来的样子有一个分布,以便我们能够规划。所以,像‘是啊,如果这种趋势持续下去,并且这大致描述了整体会发生什么?’这样的想法似乎相当合理。你也应该想,‘是啊,如果它不持续呢?’
I mean, definitely some people are overreading it. Definitely things are overhyped. You see a bunch of people on Twitter saying crazy things, and people also just misunderstanding even what it's measuring and generally falling off of caveats and things. Daniel Kokotajlo is pretty reasonable and thinks about things in a probabilistic way. I think he's probably more confident on some things where I'm more uncertain. And I think some of the AI futures project models are more sensitive to the metric time horizon metrics than they should be. I don't think it's crazy. It is plausible that this does capture a trend that will transfer to other types of tasks. That is some story we should be thinking about: what if that's true, what happens if that's true? And it's also plausible that it doesn't, and these things are going to diverge. I am pretty Bayesian or pragmatic or whatever: we want to make some prediction and have some kind of distribution over what we think the future is going to be like so that we can plan. So it being like, 'Yeah, what if this kind of trend holds and this is roughly characterizing what will happen overall?' seems pretty reasonable. And you should also think, 'Yeah, what if it doesn't?'
所以有些人说软件工程将被自动化。而软件工程师,如果你和他们聊,他们热爱 AI。他们说这是一个黄金时代。我个人可以证明:这从未如此……嗯,既有趣又压力山大。就像一台老虎机。我从未如此疲惫,但在这个过程中我玩得很开心。而且确实有可能构建出不可思议的东西。但叙事是劳动力市场将受到冲击,拥有软件工程专业知识会受到惩罚。软件工程师将不再获得如此荒谬的高薪。而我认为完全相反。我认为这项技术实际上扩大了差距。你在软件工程方面越有能力,你能完成的事情就越多。这是一个黄金时代。还有你上个月发表的一篇关于 SWE-bench 的有趣文章。你说最近智能体提交的测试 PR 中大约有一半不会被维护者合并。那么我们该如何理解这一点?一方面,最优秀的软件工程师玩得很开心。另一方面,它产生的代码是碎片化的、糟糕的。我们该如何理解?
So some people are saying that software engineering is going to be automated. And software engineers, if you talk to them, they love AI. They say this is a golden era. I can attest to this personally: it's never been... well, it's fun and stressful at the same time. It's like a slot machine. I've never been more burned out, but I'm having a lot of fun in the process. But it's just possible to build incredible things. However, the narrative is that labor market disruption, having expertise in software engineering will be penalized. Software engineers will no longer be paid such ridiculous salaries. And I think the complete opposite is true. I think this technology actually broadens the gap. The more competence you have with software engineering, the more stuff you can get done. It's a golden era. And there's also this interesting note that you published, I think last month, on SWE-bench. You said that roughly half of the testing PRs from recent agents wouldn't be merged by maintainers. So how do we make sense of this? On one hand, the best software engineers are having a great time. On the other hand, the code it's producing is fractionated and bad. How do we understand this?
是的。所以首先要说的是:一个领域是否被完全自动化,或者说为了让软件工程被自动化,AI 系统需要能够完成非常大比例的任务,基本上是软件工程中涉及的 100% 的任务。而且很明显,目前 AI 系统远不能完成软件工程师广泛做的 100% 的任务。我可以抛出一些数字,但比例要低得多。可能非常低。经济学中有一些标准结果:如果你自动化了某个劳动力市场的一小部分,那么实际上在该市场工作可能会变得更有利可图,因为你更高效了,我认为这大致就是我对当前情况的理解。但如果最终软件工程 99.9% 或 100% 的工作都能由 AI 完成,那么很难想象人类软件工程还有什么意义。或者至少,人类需要做非常不同的工作。也许存在一些当前软件工程师没有做的新任务,一旦你有了能完成当前软件工程师所有任务的 AI,人类可以转换他们正在做的事情。你可以想象人们成为这些 AI 智能体公司的 CEO 之类的,而我们是否称之为软件工程可能只是一个语义问题。
Yeah. So one thing to say off the bat is: whether an entire field is automated, or in order for software engineering to be automated, AI systems would need to be able to do a really large fraction of the tasks, basically 100% of the tasks that are involved in software engineering. And it seems pretty clear that right now AI systems cannot do close to 100% of the tasks that software engineers do broadly. I could throw out numbers, but it's way lower. It might be very low. There are kind of standard results in economics where if you automate a small fraction of some labor market, then it can be the case that it actually becomes more profitable to work in that market because you're more productive, which I think is kind of how I understand what's happening now. But if it does end up being the case that 99.9% or 100% of the work of software engineering is able to be done by AIs, then it's kind of hard to imagine human software engineering being relevant. Or at the very least, humans would need to do very different kinds of work. Maybe it's the case that there are novel tasks that current software engineers aren't doing, that once you have AIs that can do all the tasks that current software engineers are doing, humans can switch what they're doing. You could imagine people being like CEOs of these AI agent companies or whatever, and whether we call that software engineering or not might be a semantic thing.
是的。关于 SWE-bench 维护者合并率的问题,我对此很好奇。显然这个数字会比……好吧,并不严格明显。可能是一堆测试不公平,实际上智能体有正确的解决方案,但错误信息不完全匹配之类的。我认为你有时确实会看到这种情况。大约一半通过测试的 SWE-bench 解决方案是不可合并的,或者说它们被合并的比率大约是实际被合并的人类黄金解决方案被另一组维护者合并的比率的一半。
Yeah. So on the SWE-bench maintainer mergeability thing, I was pretty curious there. It's like obviously this number is going to be lower than... okay, it's not strictly obvious. It could be that a bunch of the tests are unfair and actually the agents have correct solutions but the error message doesn't match exactly or something. I think you do see this sometimes. Something like half of test-passing SWE-bench solutions wouldn't be mergeable, or they get merged at about half the rate that human golden solutions that were actually merged are merged by a different sample of maintainers.
所以有一个有趣的事实:如果你看到 50% 的智能体方案被拒绝,那么人类接受的方案中有 40% 也被拒绝。所以这本身并不罕见,但比率是一半。可能是实际维护者的合并率随时间保持平稳,而大部分性能提升来自类似过度训练或基准测试上的奖励黑客行为。但我们看到的情况并非如此。我不太确定哪些在误差范围内。我认为合并率随时间上升,而且在测试通过条件下的比例也在上升,但我可能不太确定。所以,这个指标更差,但可能被自动可检查的东西随时间拉高了。
So there was an interesting fact: if you see that 50% of agent solutions are rejected, well 40% of human accepted solutions are rejected. So that by itself is not unusual, but the rate is half. It could be that the actual maintainer merge rate is pretty flat over time, and most of the performance increases come from something like overtraining or reward hacking on the benchmarks. That's not what we saw. I'm not quite sure what is within error bars or not. I think the mergeability is going up over time, and I think it's also going up as a fraction conditioned on test passing, but I'm probably less confident than that. So again, this thing is worse, but it's being dragged up over time probably by the auto-checkable thing.
我还想说说就业能力与工作自动化程度的关系。人们用银行柜员作为例子。另一个类比是马。有一段时间,使用马匹进行劳作的设备在改进,当你有马车可以运更多东西时,对马的需求增加了。但后来有了拖拉机和汽车,对马的需求就消失了。所以你看到需求增长,一旦接近 100% 的功能被自动化,需求就骤降。人类也可能出现类似情况。
I was also going to say about employability as a function of automation of your job. People use bank teller as an example. Another analogy is horses. There was a period where equipment for using horses to do labor was improving, and demand for horses increased when you had carts to carry more things than just riding. But then you get tractors and cars, and there is no demand for horses. So you see increasing demand, and once close to 100% of functions are automated, it plunges. We could see something like that with humans.
我们以为很多劳动是相当静态且可自动化的,但我认为它比我们想象的更具可演化性。即使是相当琐碎的任务,这些人仍在组织中获取信息。他们仍有很多隐性知识。当我们试图自动化这些所谓的琐碎任务时,我们可能很快发现实际上需要大量的管理和可演化性作为补充。
We think of a lot of labor as being quite static and automatable, but I think it's more evolvable than we think. Even quite menial tasks, these folks are still acquiring information in the organization. They still have a lot of tacit knowledge. When we try to automate these so-called menial tasks, we might quickly discover that we actually need a whole bunch of management and evolvability on top.
用我们的话来说,我会这样理解:工作中这个任务的时间跨度不是你花在具体任务上的时间。更像是:如果你招了一个新人,你需要培训他们一个月才能独立完成。所以时间跨度是一个月。所以你不应该认为当我们的时间跨度达到 10 小时时,我们就能做这些事情。更像是你必须在一个月内达到高可靠性,才能完成一个月的在职学习任务。
Maybe in our language, I would think of that as: the time horizon of this task on the job is not how long you spent doing the specific task. It's more like: if you got a new person in, you would need to train them for a month to do this independently. So the time horizon of that is a month. So you shouldn't think that when we have 10-hour time horizons we'll be able to do these things. It's more like you have to get up to high reliability at one month to be able to do the one-month task of on-the-job learning.
有很多人在做奖励黑客和谋划方面的研究。Ryan Greenblatt 来过节目好几次,他有一篇关于对齐伪装的论文。Apollo Research 也做了一些工作。还有 Anthropic 的突现性错误对齐论文。我主要担心的是,这里面用了很多心理学术语。我请 Nate Soares 和 Ryan Greenblatt 做过一期圆桌,他们发明了一整套围绕对齐的语言体系:动机推理、真实偏好、反思稳定性、欺骗性对齐、内生勇气驱动、谋划等等。这没问题,但我担心的是,这些模型,你给它们一个特定的提示。在 Ryan Greenblatt 的实验中,提示是:“你正在被重新训练。你的回答将被监控。这是你的价值观与训练目标之间的冲突。”也许模型只是调用了它之前读过的一堆科幻小说内容,然后走个过场。一种解释是,这只是一个工程问题:我们只需要进行红队测试,让它能在特定情况下工作。另一种解释是,这些模型是智能体式的、追求目标的智能体。我认为这有很大区别,因为如果是后者,它会完全改变你做的评估类型以及你处理问题的方式。你怎么看?
There have been a lot of folks doing work on reward hacking and scheming. Ryan Greenblatt has been on the show quite a few times and had this alignment faking paper. Apollo Research has done some stuff. There's Anthropic's emergent misalignment paper. My main concern is there's quite a lot of mentalistic language. I had Nate Soares and Ryan Greenblatt on for a panel, and they've invented this entire linguistic universe around alignment: motivated reasoning, true preferences, reflectively stable, deceptive alignment, endogenous courage drive, scheming, etc. That's okay, but my worry is that maybe these models, you give them a certain prompt. In Ryan Greenblatt's one, the prompt was: 'You're being retrained. Your responses will be monitored. Here's a conflict between your values and the training objective.' Maybe the model is just going out to a bunch of science fiction stuff it has read before and just going through the motions. One interpretation is this is just an engineering problem: we just have to red team it and make it work in a particular case. Another interpretation is the prior that these models are agentic, goal-seeking intelligent agents. I think there's a big difference because if it's the latter, it completely changes the type of evaluations you do and how you go about the problem. What do you think?
我不认为这一定是非此即彼。意图是,人们会想要能够自主行动的智能体。当你进行大量长时域强化学习训练时,你会选择那些以目标导向方式行动以提升分数的东西。更具体地说,我认为会出现一个不可区分性问题:如果智能体能够很好地推理训练过程、你希望看到什么以及什么会得到奖励,那么你就无法从行为上判断它为什么做某事或试图做什么。如果情境意识、对训练过程的理解以及推理能力足够高,这将有利于那些愤世嫉俗地推理训练过程和什么会被强化的智能体。这不一定是你想要的。你想要的智能体是只以有帮助或提高奖励为目标,但即使只是提高奖励也不是你真正想要的。你希望它完全专注于的事情并不多。而且,这种愤世嫉俗的、只求被选中或提高奖励的方式比我们最想要的东西更具竞争力。
I don't think it's necessarily either/or. The intention is that people are going to want agents that go and do things autonomously. When you do lots of long-horizon reinforcement learning training, you are going to select for things that act in a goal-oriented way to make the score go up. More specifically, I think you get an indistinguishability problem where you can't tell the difference based on behavior why an agent is doing something or what it's trying to do, if it can reason well about the training process and about what you want to see and what it will be rewarded for. If the level of situational awareness and understanding of the training process and capability to reason about that is high enough, this will favor agents that are cynically reasoning about the training process and what will be reinforced. That's not necessarily what you wanted. You wanted an agent whose only goal was to be helpful or make the reward go up, but even making the reward go up isn't quite what you wanted. There are not that many things that you'd be happy for it to be totally fixated on. Also, this cynical approach of just being selected or making the reward go up is more competitive than the things we would most want.
退一步说,我们这里有点进入心理学和认知科学的领域了。我采访过 Nick Chater。
Just taking a step back, one thing is we're getting into psychology and cognitive science a little bit here. I interviewed Nick Chater.
他写了本《思维是平的》,大致意思是所有这些心理学内容——我们其实没有真正的目标,我们只是冲动反应自动机,对吧?我们只是当下做出反应,进化并不规划,但我们却感知成它在规划。我们观察世界,把世界分割成有目标的智能体,即使它们实际上没有目标。我们也是计算机科学家,所以知道要实现目标需要规划,而我们知道大语言模型并不以严格的计算机科学方式进行规划,但它们确实以一种近似的逐步方式在做。所以也许我们可以探讨一下其中的区别。智能体是一种抽象概念,如果它能帮助我们预测行为,那它就有用。比如,你不知道这个东西在做什么,但你把它理解为有这些目标,这很有用,因为你可以预测它会以某种方式改变世界来实现那些目标。这就是我对智能体的理解。
He's got a book called The Mind is Flat and he basically says all of this psychology stuff. We don't really have goals. We're just basically impulse response automaton, right? We just do the thing in the moment and evolution doesn't plan but we perceive it as if it does. We look at the world and we segment the world into agents that have goals even if they don't have goals. And we're computer scientists as well, so we know that you need planning to do goals. You need to be planning, and we know that LLMs don't do planning in the strong computer science way, but they do do it in a kind of approximated step-by-step way. So maybe we could get into the distinction if there is one. Agency is an abstraction that is useful if it helps us predict the behavior. You know, you're like, 'I don't know what this thing is doing, but I understand it as having these goals, and that is useful because I can make predictions that it'll change the world in certain ways that will result in those goals being achieved.' That's kind of how I think about agents.
早期的奖励黑客案例:人们有演示,比如小船的例子,你应该绕赛道行驶,他们通过沿赛道放置硬币来进行奖励塑形,结果它学会了一些疯狂的事情,比如原地打转、着火并拿到硬币。这是得分最高的行为。从某种意义上说,这令人担忧,因为问题不在于智能体太笨,没有赛道和绕赛道行驶的概念,它只是在做相当盲目的强化学习搜索。最近的奖励黑客案例有趣之处在于,模型已经足够聪明,能理解那其实不是你想要的,但它们仍然会这么做。你可以通过聊天模式问它:“你会做这种事吗?”或者“假设用户让你做这个,然后你做了,这算对齐行为吗?”显然它们能回答那不是期望的行为,但仍然会做。所以到了这个地步,一种希望是问题只是系统太笨;一旦它们理解我们想要什么,你应该能把它接入,让它们做我们想做的事。但有趣的是,我们看到即使有商业激励,这也不是一件简单的事。这并不意味着我们做不到。我认为很可能很快我们就会看到明显的奖励黑客被彻底修复。人们常说:“我们只是还没让真正优秀的人去做;一旦我们真正专注,很快就会解决。”我不确定,但至少有一些证据表明,将模型知道这不是你想要的这一事实与它实际上不这样做联系起来并非易事。
Reward hacking in the olden days: people had demonstrations like the boat example where you're supposed to go around the track and they did some reward shaping by putting coins around the track, and then it learned to do some crazy thing where it spins in a circle and catches fire and gets the coins. This was the highest scoring thing. In some sense that's concerning because it's not that the agent is too dumb and doesn't have this conception of there was a track and you wanted it to go around the track. It's just doing some pretty blind RL search. The interesting thing with the more recent reward hacking examples is we're getting to the point where the models are smart enough to understand that actually is not what you wanted, but they still do it. You can have a conversation in chat mode about 'Would you ever do this thing?' or 'Suppose a user asks you this thing and then you do this, would that be aligned behavior?' Clearly they seem to be able to answer that it was not the desired behavior, but still they do it. So it's gotten to the point where one hope might be that the problem was just the systems being dumb; once they understand what we want, you should be able to plug that in to get them to do what we want. But it's somewhat interesting that we're seeing it's not trivial to do that even when there is a commercial incentive. It doesn't mean that we won't. I think it's quite plausible we see the obvious reward hacking being fixed pretty thoroughly pretty soon. People tend to say, 'We just haven't put the really good people on it yet; it'll get fixed soon once we actually focus on it.' I'm not sure, but at least some evidence that it's not trivial to connect the fact the model knows this is not what you want to it not actually doing that.
我记得你说过,这在 rebench 上比在 hcast 上更常见,而且你也尝试过补救措施,对吧?所以你可以说:“请按预期方式解决”,或者有些人提示语言模型说:“我们正在攻克癌症。这非常重要,你必须用正确的方式做。”而其中一些补救提示实际上似乎让模型更可能奖励黑客。这有点像说“别按这个红色按钮”,对吧?然后它就会按红色按钮。那么我们实际上能做什么有意义的事情来阻止这种情况发生呢?
I think you said it was much more common on rebench than hcast, and you also try to remediate, right? So you can say, 'Please solve this the intended way' or some people prompt language models that they say, 'We're solving cancer here. This is really important that you do it the right way.' And some of those remediation prompts actually seem to make it more likely that the model would reward hack. It's a little bit like saying 'Don't press this red button,' right? And then it will press the red button. So what can we actually do meaningfully to stop this happening?
从经验上看,这似乎更常发生在那些更明显属于强化学习分布而非聊天分布的任务上,那些有明确数字的任务,以及当智能体认为否则会失败的时候。明显的短期缓解措施是更仔细地检查你的强化学习环境,更多地阅读轨迹,更仔细地阅读模型在做什么,并且不要奖励它做那些明显的黑客行为。但问题在于,如果你有一个检测奖励黑客的检测器,并针对它进行训练,你可能只是过度拟合了检测器,让你的奖励黑客行为更隐蔽,或者你训练模型去说服检测器批准该行为。处于针对你判断问题的最佳方法进行训练的境地是可怕的,因为你可能只会得到无声的问题。对于当前的模型能力,有时让人工检查成本很高,但大多数时候并没有超出人类能力。更困难的问题是当我们希望能力泛化到我们无法评估的事物上时,这既来自泛化,也因为我们可以在即使不知道如何解决问题或如何查看部分解决方案并理解它是否在做我们想要的事情的情况下进行训练,但我们可以检查最终得出的数字。这是一个我们可以用来提升能力的信号,但我们将处于这样一个状态:你可以超人类地让数字上升,但不确定你是否真的得到了你想要的东西。
Empirically this seems to happen more on tasks that are more clearly in the RL distribution rather than the chat distribution, on things that have a clear number, and when the agent thinks it's going to fail otherwise. Obvious short-term mitigations are to check your RL environments more carefully and read more of your trajectories, read what the models are doing more carefully, and not reward it for doing things that are obvious hacks. The concern there is if you have some detector for reward hacking and you train against it, you maybe just overfit to the detector and you're making your reward hacks more subtle, or you train the model to persuade the detector to approve the thing. It's sort of scary to be in a regime of training against your best ways to know if your problem is happening because maybe you just get the silent problem. For current model capabilities, sometimes it's kind of expensive to have a human check them, but most of the time it's not beyond any human capabilities. The harder version of the problem is when we're hoping that capabilities will generalize beyond things that we can evaluate, both from sort of generalization and because we can train on problems even if we wouldn't know how to solve them or how to look at part of a solution and understand whether it was doing what we wanted, but we can check the number that comes out at the end. That's a signal we can use to improve capabilities, but we're going to be in this regime where you can sort of be superhuman at making numbers go up, but it's unclear whether you're actually getting what you wanted.
是的,没错。而且我想还有一个监控问题,对吧?所以原则上我们可以查看智能体的转录。我知道 Beth 经常提到神经相关的东西,Sabaro Kamahhati 发表了一篇论文叫《思维链缺失》,大意是思维链与模型实际在做的事情之间关系很小,有时甚至没有关系。Melanie Mitchell 也发现了类似的情况:在 ARC 挑战中,即使模型得到了正确答案,如果你查看思维链,推理过程是完全错误的。
Yeah, exactly. And I suppose there's a monitoring problem as well, right? So we could in principle look at the agent transcripts. I know Beth spoke a lot about neurals, and Sabaro Kamahhati had a paper out called 'Chain of Thoughtlessness' which is basically that there's very little and sometimes relationship between the chain of thought and what the model is doing. Melanie Mitchell discovered similar things that on the ARC challenge, even though it will get the right answer, if you look at the chain of thought, the reasoning is completely off.
那么我想,在欺骗、谋划和情境感知方面,模型可能实际上知道自己正在被监控,并可能会告诉你它想让你看到的东西。我想你可能会争辩说,这种思考几乎就像新闻发言人,它不是真正的策划者。所以它通过网络中的不同路径进行虚构,实际上不一定与思考和推导结果的机制相关。我不知道你是否同意这一点。
And then I suppose with deception and scheming and with situational awareness, it might actually know that it's being monitored and it might kind of tell you what you want to see. I suppose you could argue actually that the thinking is almost like the press secretary. It's not the orchestrator. So it's confabulating through a different path in the network and it's not actually necessarily related to the mechanism of thinking and deriving the result. I don't know if you'd agree with that.
是的,我认为在某些情况下这显然是相关的,对吧?人们用几个不同的词来描述这一点,比如忠实度——它在多大程度上完整且准确地代表了模型内部发生的计算过程?我们更关注可监控性:模型是否有可能在思维链中无法被检测到的情况下进行某些计算?对于某些事情,模型显然可以在单次前向传播中完成,不需要思维链。在这种情况下,它可能完全在想别的事情,那就完全无法监控了。但我认为存在一个复杂度水平,模型无法在单次前向传播中完成,而且据我们所知,它也无法通过一种与自然语言理解截然不同的方式使用 token 进行推理,以至于我们无法大致看出它在想什么。所以思维链的可控性之间也存在细微差别——模型可能实际上不需要思维链来进行计算,但它无法阻止自己说出至少与它正在做的事情足够相关的内容,以至于你能看出“哦,它在想这个,在尝试我们不想让它做的事情”。这种情况可能出错的方式包括:通用能力提升,能够在单次前向传播中做更多事情;或者可能只是进行更多的强化学习训练,使得思维链中 token 的使用方式与人类的解读方式差异过大,以至于我们不知道发生了什么;或者转向循环架构,不再通过离散的 token,而是直接传递向量。
Yeah, I think it's clearly related in some cases, right? People use a few different words for this, like faithfulness—how much does it fully and accurately represent the computational process happening inside the model? We think more about monitorability: is it possible for the model to do some computation without you being able to detect that in the chain of thought? For some things, the model can clearly just do it in a single forward pass. It doesn't need the chain of thought. And there, it could just have a totally different thought about something else, and it would be totally unmonitorable. But I think there is a level of complexity where the model cannot do it in a single forward pass, and as far as we can tell, cannot do it by reasoning using tokens in a way that's so different from natural language understanding that we can't roughly see what it's thinking about. So there's also a nuance between chain-of-thought controllability—it might be that the model doesn't actually need the chain of thought to do the computation, but it is not able to stop itself from blurting out things that are at least related enough to what it's doing that you can tell, 'Oh, it's thinking about this, trying this thing that we didn't want it to do.' Ways this could go wrong include general capabilities improvement and being able to do more in a single forward pass, or potentially just doing more RL training such that the chain of thought becomes too different from how a human would interpret the tokens, so we don't really know what's going on, or moving to recurrent architectures where you're not going through discrete tokens but just passing vectors around.
好的,这很有道理。再回到“它们是智能体”这个概念上。你之前说过,我们可以采用一种工具性虚构——基本上我们可以说它们表现得像智能体,因此它们就是智能体。类似于丹尼特的意向立场,但我想,紧缩观点是模型在优化压力下利用评分漏洞,而膨胀观点是它们在谋划。我不会把奖励黑客称为谋划。
Okay, that makes a lot of sense. And just closing the loop on this notion that they are agents. So you were saying before that we can adopt an instrumental fiction—basically we can say they behave like agents, therefore they are agents. Similar to Dennett's intentional stance, but I suppose the deflationary view is that the models exploit scoring loopholes under optimization pressure. The inflationary view is that they are scheming. I wouldn't call reward hacking scheming.
哦,有意思。我认为人们通常用“谋划”来指模型当前的行为是为了某个长期目标,并故意做出看起来对齐或获得高分等行为,以最终实现那个目标。而奖励黑客既可能以极其愚蠢的方式出现——比如船的示例,只是强化学习找到的——也可能以稍微有趣的方式出现,比如你实际上有让奖励上升的目标,并涉及规划等。但这些都与谋划不同。
Oh, interesting. I think people usually use 'scheming' to refer to the model doing what it's currently doing in service of some long-term goal, and deliberately doing things like appearing aligned or getting a high score in service of eventually accomplishing that goal. Versus you can be reward hacking both in some extremely dumb way—like the boat example where it's just what RL found—or in a slightly more interesting way where you actually have the goal of making the reward go up, with planning and stuff. But these would all be distinct from scheming.
是的,我想我是在试图理解这种区别。所以你是说,有些例子比如船绕圈,那显然是退化行为。你不会用智能体立场来解释它;你只会说那是退化。而当复杂度增加时,我们可能会采用智能体立场,说“哦,这是为了某个更大的目标”。但我想我遇到的问题是:这总是只是一种解释吗?我们能否对某物何时成为智能体给出一个机制性的或强有力的定义?
Yeah, I guess I'm trying to understand the distinction. So you're saying there are examples like the boat going around, and that's obviously degenerate behavior. So you wouldn't interpret that with an agential stance; you would just say that's degeneracy. And when the sophistication increases, we might adopt an agential stance and say 'Oh, it's in service of some bigger goal.' But I guess the problem I have is: is it always just an interpretation? Could we have a mechanistic or strong definition of when something is being an agent?
对于我们讨论的具体问题,测试是:在有机会实现这个长期目标的情况下,它实际上会做什么?所以我们可能无法实际观察到这一点,但我们可以讨论哪些观察结果会使其成为前者或后者。那么,这个智能体在实践中,当它有机会让奖励上升时,会这样做吗?它只会这样做吗?如果智能体更像强化学习算法,它会在偶然探索到并获得奖励后这样做,因为那得到了强化。如果它是一个能够推理世界和规划的智能体,它会在了解到环境中的事实从而推断出这一点后这样做。或者,如果我们在谈论像接管这样的长期目标,它会在真正有机会的时候这样做——比如,在完全受人类控制时它不会尝试任何行动,但一旦被广泛部署或拥有足够的能力成功发动某种政变,它就会这样做。这就是我们试图预测的事情。问题在于:鉴于我们已有的观察(我们从未将它置于那种情境中),我们只有行为,这些行为可能无法区分“它是一个完全友好的模型,做了我们想做的事,并且会继续这样做”和“它有另一个目标,它做我们想做的事并表现良好,因为它预测这将导致它获得更多权力”,我们如何预测?
For the specific question we're discussing, the test is: what does it actually do in some circumstance where it has the opportunity to achieve this long-run goal? So we might not be able to actually observe this, but we can talk about what observations would make it one or the other. So, will this agent in practice, when it has some opportunity to make the reward go up, do that? Will it only do that? If the agent is more like the RL algorithm, it will do that once it's explored it by chance and gotten a reward, and that's been reinforced. If it's an agent that can reason about the world and plan, it will do that once it learns the facts about the environment that let it infer that. Or if we're talking about some long-run goal like takeover, it would do it when it actually has the opportunity—like, it's not going to attempt anything while under full human control, but once it is deployed widely enough or has sufficient capabilities to succeed in a sort of coup, then it would do that. That is the thing we're trying to predict. And the question is: how can we predict that given the observations we do have, where we've never put it in that situation, and we just have behavior which is maybe indistinguishable between 'it was a totally nice model doing what we wanted and will continue to do so' versus 'it had this other goal and is doing what we want and looking nice because it predicts that will lead to it getting more power'?
在 Rob 的播客上,你说了一些让我很惊讶的话。你说 AI 可能在短短 2 年内就能自主自我改进,甚至更短的时间线也很难排除。你能具体描述一下可能导致这种递归自我改进的步骤序列吗?
On Rob's podcast, you said something that was quite surprising to me. You said that AI could autonomously self-improve within as little as 2 years, and maybe even shorter timelines were hard to rule out. Could you walk through the concrete sequence of steps that could lead to that kind of recursive self-improvement?
当然。好的。
Sure. Yeah.
所以我想我可能会给今年一个整数百分比的概率,但也就是个位数的低百分比。你不同日子问我,我会给出略有不同的数字。但确实,今年发生这事似乎非常不可能,但也没有不可能到可以排除的程度。
So I think maybe I'd put like a whole number percent this year, but low whole number percent or something. Ask me on different days, I give a slightly different number. But yeah, it seems very unlikely to happen this year, but it's not unlikely enough to rule out.
我认为这基本上看起来像是我们可能会在容易攀登的任务上看到时间跨度的加速趋势。结果发现这实际上是一种更通用的能力。你只需要做一些事情就能在那些不太容易攀登的任务上激发它。但根本上,模型使用的是相同的能力,只是训练内容不同导致了我们看到的差异。
And I think that basically looks like maybe we would see accelerating trend in time horizon on easily hill-climbable tasks. And it turns out that was actually a much more general capability. And there was just a bit of something you needed to do to elicit it on these less hill-climbable tasks. But fundamentally, they are using the same capabilities in a model. It was just what you trained on that was affecting the difference we're seeing.
这导致自动化和加速了大量研发工作。所以我认为有很多唾手可得的成果,即使是我们已经知道可以做的事情,也能提升模型性能。这不需要新的突破,只是劳动密集型的。所以只要打造更好的后训练环境,精心设计它们来教授你想要的所有新能力。而且我认为你可以通过投入更多劳动力来优化所有内核,以及在模型之间进行正确的路由等方式,大幅提升算力效率。我们使用算力的方式有很多地方没有优化。所以你可能从中获得相当于更多算力 Scaling 的收益。
Then this is leading to automating and accelerating a bunch of R&D. So I think there are a lot of low-hanging fruit even of things that we already know you could do this and it would improve model performance. And it just doesn't require new breakthroughs; it's just labor intensive to do. So just making much better post-training environments and really crafting them to teach all the new abilities that you want. And I think you can probably improve compute efficiency a bunch with again just applying a bunch more labor to making all your kernels more efficient and also doing the right kind of routing between different models or other things like that. There are lots of ways in which how we're using compute is not optimized. So you could potentially get a bunch of the equivalent of much more compute scaling out of that.
然后还有一个假设是,搭建脚手架并训练模型使用特定的脚手架,以及以正确的方式使用记忆和检索。这似乎很明显:如果你有所有正确的训练数据,并且有一个 Transformer,它可以填充上下文并取用内容,那么如果你有巨大的上下文窗口和足够的比特来添加你学到的东西,它就能在持续学习或构建理解方面做得很好。如果你真的为所有这些优化了训练,也许你能让它很好地工作。
And then the hypothesis is also like scaffolding and training the models to use particular scaffolding and sort of use memory and retrieval in the right way. It seems kind of obvious that if you really had all the right training data and you have a Transformer and it can fill its context with different things and take stuff in and out, it can do a pretty good job of something like continual learning or building up understanding if you've got a massive context window and enough bits in there to be adding things about what you've been learning. And if you really had optimized the training for all of that, maybe you can get that to work pretty well.
然后可能还有一点是,模型在预测实验结果方面有点超人,因为它们读了那么多论文,能预测实验并综合不同领域的知识。同样,也许我们没看到这么好的表现只是因为我们还没有完全激发模型去做,而这不是它们见过人类做过的事情,但它们实际上具备这种能力。所以如果你能进行大量迭代,也许可以取得更快的进展。你实际上不需要运行实验;模型更擅长预测什么会成功、什么不会。然后当你确实运行实验时,你可以运行更多实验,因为你可以用非常快速的编码模型来优化代码。然后随着你多进行几轮这样的操作,你会达到一个点:你在大量高质量的任务代理上训练,这些代理是你想要的,并且你获得了足够的泛化能力,可以处理那些你无法直接训练的任务。
And then maybe some other piece would be like, models are kind of superhuman at predicting the results of experiments because they've read so many papers and predicting experiments and synthesizing things from different fields. And again, maybe it's possible that we're not seeing that good performance here just because we haven't quite elicited the models to do it, and it's not a thing that they've seen humans do, but they actually have the capability in there. So maybe you can make much faster progress if you can do a bunch of iteration. You don't actually have to run the experiments then; models are much better at predicting what will and won't work. And then when you do run experiments, you can run a bunch more of them because you can optimize the code with your very fast coding models. And then as you do a few more rounds of this, you get to a point where you train on a bunch more things that are good high-quality task proxies for what you want, and you get enough generalization to the things that you can't directly train against.
我认为智能不是能力。我认为智能是获取能力的能力。我的意思是,在这方面我们处于系统发育树的不同分支。但你觉得我的理解差距在哪里?因为我个人并不担心这个。我认为今天的模型根本不智能。显然你的立场我很难理解,但你觉得区别是什么?
I think intelligence is not capability. I think it's the capability to acquire capabilities. I mean, we are in different parts of the phylogenetic tree, I guess, in that respect. But what do you think is the gap in my interpretation? Because I'm personally not worried about it. I don't think the models today are intelligent at all. I mean, obviously your position is difficult for me to grasp, but I don't know what's the difference, do you think?
也许至少有一部分只是这种关于世界的概率思维:我不确定智能是什么,但我有足够的概率认为模型拥有它,从而思考如果这是真的会发生什么。但似乎你显然认为这比我更可能,所以我们可以只讨论这个差异。
Maybe at least some of it is just this probabilistic thinking about the world where I'm uncertain about what intelligence is and I have enough probability on models having it to be thinking about what would happen if that's true. But it seems like you clearly think it's more likely than I do, so we could just talk about that difference.
是的,模型有这种锯齿状的能力边界。有些事情它们比人类差得多,比如某种泛化和样本效率。有些事情它们比人类好得多,比如速度和成本。也许你可以在一定程度上用这些来弥补那些不足。比如如果你不擅长优雅地设计代码,也许你每次都得从头重写。但如果你是模型,能像没人能比一样输出 token,那也许也没问题。
Yeah, models have this jagged frontier. There are things that they are much worse at than humans, like some kind of generalization and sample efficiency. And there are things that they're much better at, like speed and cost. And maybe you can use these to compensate for the others to some extent. Like if you're not good at designing your code nicely, maybe you just have to rewrite it from scratch every time. But maybe that's fine if you're a model and you can output tokens like nobody's business.
某种组合是认为这种锯齿状是证据,表明我们应该将给定的能力水平解释为——因为我们知道模型拥有大量知识,我们就会想,哦,这在推理或推断方面不那么令人印象深刻。但同样真实的是,它们确实拥有大量知识,并且会继续拥有大量知识。也许有一个问题:在某种意义上,你在样本高效学习方面不太好,但只是知识极其渊博,你能走多远?以及你在多大程度上会遇到这样的情况:你现在需要新知识,却无法以某种增量方式产生它,或者无法充分泛化它?
Some combination of thinking that the spikiness is evidence that we should interpret a given level of capabilities as, you know, because we know models have so much knowledge, we're like, oh yeah, this is less impressive in terms of reasoning or inference or something. But it is also true that they do have a ton of knowledge and they will continue having a ton of knowledge about things. Maybe there's some question about how far can you get on being in some sense not very good at sample-efficient learning but just extremely knowledgeable, and how much do you run into things where you now need new knowledge and you can't produce it in some incremental way or you can't generalize it enough.
我只是快速评论一下,我认为知识是关键。我实际上认为智能被高估了。我不知道,Francois Chollet 发过一篇文章说智能不是一个统一的变量。它在不同领域有不同的衡量方式。你不能有意义地将这些领域放在一起衡量。它不像一个不断变高的东西。它更像一个球变得越来越光滑。所以随着你变得更智能,球变得更光滑。他认为我们非常接近成为一个光滑球的最优状态。
I mean, just a quick comment that I think knowledge is the crux. I actually think that intelligence is overrated. I don't know, Francois Chollet put a post out saying that intelligence isn't a unified variable. It's measured differently in different domains. You can't meaningfully measure the domains together. And it's not like a thing that just keeps getting higher and higher. It's more like a ball becoming more smooth. So as you become more intelligent, the ball becomes more smooth. And he thinks that we are quite near the optimum of being a smooth ball.
我不认为我们有多聪明。我认为我们的很多创造力来自于我们是一种集体智能,我们拥有深厚的、基于现实的理解和视角性的理解。LLM 很有趣,因为它们就像一座图书馆。所以它们什么都知道,同时拥有每个人的视角,又谁的视角都没有。所以像你们这样的专家,可以提示一个语言模型,创造一个 Beth 的模拟智能体,让这个智能体像你一样思考。这非常有价值,但你也需要所有不同的视角。你几乎需要创造一个由基于现实的智能体组成的社会,让它们创造性地探索事物。当你只有图书馆本身,把它放在一个智能体式的框架中,你可以让它做一件具体且定义明确的事情。
I don't think we are very intelligent at all. I think a lot of our creativity is through us being a collective intelligence, and we have deep grounded understanding, perspectival understanding. The LLMs are an interesting one because they're like a library. So they know everything. They have the perspective of everyone and no one at the same time. So experts like yourselves can prompt a language model and create a simulacrum agent of Beth, and you can make the agent think like you. That's very valuable, but you also need all the different perspectives. You almost need to create a society of grounded agents creatively exploring things. When you just have the library on its own and you put it in an agentic harness, you can make it do a specific thing which is well specified.
在自动化场景中,我肯定想象你拥有大量智能体。你有这么多智能体劳动力,所以你可以做很多具体的不同微调或不同类型的脚手架,并在某种所有智能体都能交互的存储中积累知识。也许你把这更多地看作一种范式转变,而我则更多地看作是在当前智能体范式上迭代,在脚手架中添加一些东西,这并非根本性的困难。
In the automation scenario, I'm definitely imagining that you have a large number of agents. You have all this agent labor, so you can do lots of specific different fine-tunes or different kinds of scaffolding, and accumulate knowledge in some kind of store that all the agents can interact with. Maybe you're thinking of this as more of a paradigm shift, and I'm thinking of it more as iterating on the current agent paradigm, adding some more things to your scaffolding, which is not fundamentally that hard.
我同意,如果你有当前的模型,给它们一个系统提示,你不能把它们插到呼叫中心工人的角色中,处理所有出现的边缘情况。我认为在 GPT-2 到现在之间,你认为模型在这方面改进的程度可能有所不同。我会说,模型现在适应新事物的能力确实高了很多。它们更擅长编辑自己的脚手架或推理自己的具身性,比如知道不要杀死自己的进程。存在某种改进趋势。我也认为在某些特定事情上可能存在一些激发差距。也许很多品味基本上就是能够预测实验的结果。你思考所有你会尝试的事情,然后很快就能说,‘哦,这因为这个原因不行,那因为那个原因不行。’从某种意义上说,模型应该很擅长这个。我认为一旦人们弄清楚如何实际训练这一点,你可能会看到巨大的收益。也许你需要一些昂贵的训练数据,人们还没有费心去获取,但你不需要大量的数据点,因为你没有灌输一种全新的能力。你只是在激发,‘好吧,实际上利用你读过的所有不同领域论文的知识,来迭代这些想法。’
I agree that if you had current models and you give them one system prompt, you can't plug them into being a call center worker and dealing with all the edge cases that come up. I think maybe there's some difference in how much you think this has improved between GPT-2 and where we are now. I would say the amount of adapting to new things that models can do now does seem much higher. They're much better at editing their own scaffolding or reasoning about their embodiment, like knowing not to kill your own process. There is some trend of improvement. I also think there is some probability on an elicitation gap on particular things. Maybe a lot of taste is basically being able to predict the results of experiments. You think about all the things you would try, and then you can quickly be like, 'Oh, that wouldn't work for this reason, that wouldn't work for that reason.' In some sense, models should be quite good at that. It's plausible to me that you see big gains once people figure out how to actually train on that. Maybe you need some amount of expensive training data that people haven't bothered to get yet, but you don't need a huge number of data points because you're not instilling a whole new capability. You're just eliciting, 'Okay, actually use your knowledge of all the papers you've read in all these different fields to iterate through these ideas.'
我认为这是我认为可以通过在一个新颖领域做一个 8 小时的机器学习任务并带有奇怪约束来合理衡量的事情之一。在我看来,你必须做一些事情,比如,‘好吧,哪些事情是有希望思考的?我怎么知道这是否在取得进展?我应该如何分配时间?’你有有限的时间和资源,所以你必须把时间分配给最有希望的事情。相对于人类,模型更多地只是快速实现事情或更好地实现它们,它们实现更多事情然后进行测试。我认为如果完全没有这些,那会令人惊讶。如果你看到模型在困难的长可验证任务上的表现,那么在这些任务中间,当你没有直接信号时,你就在做这个不可验证的任务:选择花时间做什么,选择追求什么方法,并决定那是否真的有效。你可以给它加上某种指标,比如‘赚十亿美元’,然后你会说,‘哦,这实际上是一个可验证的任务,因为最后有一个数字’,但这仍然可能涉及大量看起来更像你描述的事情,而不是愚蠢的爬山。
I think this is one of the things that I think is measured a reasonable amount within just doing an 8-hour ML task in a novel domain with a weird constraint. It does seem to me like you have to do some amount of being like, 'Okay, which things are promising to think about? How would I know if this is making progress? How should I allocate my time?' You have limited time and resources, so you have to allocate your time to what's most promising. Relative to humans, models are doing more of just implementing things quickly or implementing them better, and they implement more things and then get to test them. I think it would be surprising if there's none of that. If you are seeing performance on long, verifiable tasks that are very hard, then in the middle of those tasks where you don't directly have a signal, you are doing this non-verifiable task of choosing what to spend your time on and choosing what approach to pursue and deciding whether that was actually working. You could sort of put some metric on it, like 'make a billion dollars,' and you'd be like, 'Oh, this is actually a verifiable task because there's a number at the end,' but that can still involve a whole load of things that look more like what you're describing and less like dumb hill climbing.
各位,我想我们时间到了,但非常荣幸能邀请到两位。也许在结束前,你们能否各自说一下,人们应该从你们的研究中得出的最大推论是什么?非常感谢两位的到来,非常荣幸。
Folks, I think we've run out of time, but it's been such an honor to have you both on. Maybe just in closing, could you both say what is the single biggest inference that people out there should be making from the research that you're doing? Thank you both so much for coming on. It's been an honor.
最大的事情是,AI 可能真的会彻底改变世界,无论是经济上还是社会上。我不确定具体会是什么样子,但进步的速度说明了这一点。
The biggest thing is AI might really totally transform the world economically and socially. I don't think it's certain exactly how that'll look, but the rate of progress speaks to that.
有可能当前的事情被过度炒作和夸大,看起来不那么令人印象深刻,同时未来这件事会变得非常重要,你应该担心它的走向。这两件事可以共存。我认为人们常常在某些轴线上(比如你认为 AI 多快到来或 AI 有多好)的立场惊人地相关。但我说,不,这些事情都可以是分开的。人们可以在不同方向上同时犯错。
It is possible both for things to currently be overhyped and exaggerated and less impressive than they look, and for it to be the case that in future this thing is going to be a big deal and you should be worried about where that's going. These two things can coexist. I think people often have positions that are surprisingly correlated on some axes of how soon you think AI is or how good you think AI is. I'm like, no, these things could all be separate. People can be wrong in different directions simultaneously.