智能商品化:2025 年 AI 的三个关键趋势

Intelligence as a Commodity: Three Key Trends for AI in 2025

杰森·魏 Jason Wei · 斯坦福 AI 俱乐部 · 2025-10-18 · 约 30 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

来自 Meta 超级智能实验室的 Jason 讨论了理解 2025 年 AI 的三个基本理念:智能商品化、验证者法则和智能的锯齿边缘。

Jason from Meta Super Intelligence Labs discusses three fundamental ideas to navigate AI in 2025: intelligence as a commodity, verifier's law, and the jagged edge of intelligence.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 14)

全文 · Full transcript(中英对照)

介绍与日程 Introduction and Schedule

Host

好了,大家好。我们马上开始,简单说一下日程。Jason 会先做个演讲,然后我们开放问答。介绍一下 Jason,他是 Meta 超级智能实验室的研究科学家。之前在 OpenAI 工作了两年,共同创建了 o1 和 deep research。再之前是 Google Brain 的研究科学家,他的工作帮助推广了思维链提示、指令微调和许多涌现现象。他的研究是现代 AI 领域最有影响力的之一,引用超过 9 万次。Jason,很高兴你今天能来,谢谢你和我们交流,请开始吧。

All right. Hello everyone. We're going to get started now just for a quick schedule. Jason's going to give a talk and then we're going to open it up to Q&A. And to introduce Jason, he's a research scientist at Meta Super Intelligence Labs. Previously he worked at OpenAI for two years where he co-created o1 and deep research. Before that he was a research scientist at Google Brain where his work helped popularize chain of thought prompting, instruction tuning, and a lot of emerging phenomena. His research is some of the most influential in the modern AI world with over 90,000 citations. And Jason, it's great to have you here today. Thanks for taking the time to speak with us and take it away.

Jason Wei

谢谢介绍,很高兴来这里。我今天会讲大概 25 到 30 分钟,然后问答环节。我会讲三个简单但基本的概念,帮助理解 2025 年的 AI。如果你问 AI 发展会如何改变世界,答案因人而异。我有个量化交易的朋友说 ChatGPT 很酷,但做不了他工作里的事。另一端,我问了一个顶级实验室的广告研究员,他说我们大概还有两到三年工作,之后 AI 就会取代我们。所以看法差异很大。我会讲三种思考方式。第一个趋势是智能将变成商品。获取知识或进行推理的成本和门槛将趋近于零。第二个我称之为验证者定律:训练 AI 完成特定任务的能力与验证该任务的容易程度成正比。最后一个是智能的锯齿边缘。AI 在特定任务上的能力和改进速度会因任务特性而异。首先讲智能作为商品。AI 进步有两个阶段。第一阶段是推动前沿,AI 还做不好,你在解锁新能力。比如看 MMLU 过去五年的进展,性能逐渐提升。第二阶段是能力商品化。这里有个例子,y 轴是时间,x 轴是达到特定 MMLU 性能的成本(美元)。趋势是每年使用特定智能水平模型的成本都在下降。为什么这个趋势会持续?我的论点是这是深度学习史上第一次自适应算力真正起作用。回顾到去年,深度学习都是固定算力,不管问题难易。但现在有了自适应算力,你可以根据任务调整算力。这首先在 o1 上展示,一年多前发布,显示增加测试时算力解决数学问题,性能更高。自适应算力能持续降低智能成本,因为你不必一直扩大模型规模。对于简单任务,你可以用最便宜的算力。第二部分是获取公共信息的时间。x 轴是前互联网时代、互联网时代、聊天机器人时代和智能体时代。问找一条信息要多久。比如想知道 1983 年釜山的人口,前互联网时代你可能开车去图书馆翻百科全书,几小时。互联网时代搜索浏览,几分钟。现在基本即时。更难的知识,比如 1983 年釜山有多少对夫妇结婚,前互联网时代你可能要飞韩国翻书。互联网时代容易些但仍困难。聊天机器人时代容易一点。智能体时代我认为几分钟内就能找到。

Yeah, thanks for the nice intro. Fun to be here. So I guess today I'll talk for I'll try to keep it not too long, maybe like 25 or 30 minutes and then we can do a Q&A session on whatever you guys want. Okay. So I'll talk today about three pretty simple but maybe fundamental ideas to understand to navigate AI in 2025. So if you ask this question of how our world is going to change with the development of AI, you get a broad spectrum of answers depending on who you ask. One of my quant trader buddies says ChatGPT is cool and all but it can't really do the stuff in his job. And on the other end, I asked an ad research at a top lab and he says we basically have two to three more years of working before AI takes our jobs. So there's a huge spectrum of how people think AI is going to play out. I'm going to talk about three ways of thinking about it. The first trend is that intelligence is going to become a commodity. The cost and accessibility of finding out knowledge or doing some reasoning is going to be driven towards zero. The second is something I call verifier's law: the ability to train AI to do a particular task is proportional to how easy it is to verify that task. And the final one is the jagged edge of intelligence. Both the capability and the rate of improvement of AI on a particular task will vary based on certain properties of those tasks. Okay, first let's talk about intelligence as a commodity. I would say there are two stages of AI progress. The first stage is when you're pushing the frontier. AI can't really do the thing well yet, and you're unlocking that new ability. You can see, for example, if you plot MMLU, a very common benchmark, over the past five years, you'll see gradual progress in performance. Then the second stage is once you have an ability, it becomes commoditized. Here's an example where the y-axis is time and the cost of getting a particular performance on MMLU in dollars. You can see the trend: every new year, the cost of using a model with a particular level of intelligence decreases. You might ask why this trend might continue. The argument I would make is that it's the first time in the history of deep learning that adaptive compute actually truly works. If you look at the entirety of deep learning up to last year, we were in a mode where the amount of compute used for a particular problem was fixed regardless of how hard the problem was, whether it was giving the capital of California or a very hard competition math problem. But now we're in an era of adaptive compute where you can vary the amount of compute used for your task. This was first shown with o1, which came out more than a year ago, showing that if you increase the amount of compute used at test time to solve math problems, the performance on that benchmark is higher. The reason adaptive compute means you can continue to decrease the cost of intelligence is that you don't have to keep scaling model size. For a very easy task, you can go to the limit of how cheap compute you'd have to spend. The second part is you can think about the time to retrieve certain pieces of public information. On the x-axis you have pre-internet era, internet era, chatbot era, and then agents era. You can ask how long it takes to find a certain piece of information. For example, if you wanted to know the population of Busan in 1983, before the internet you'd probably drive to the library and look in encyclopedias, taking a few hours. In the internet era, you'd search and browse websites, taking minutes. Now it's basically instant. For a progressively harder piece of knowledge, like how many couples got married in Busan in 1983, before the internet you might have to fly to Korea and dig through books. Internet era it might be easier but still challenging. In the chatbot era it's a bit easier. And with the agents era, I think this could be found within minutes.

智能作为商品 Intelligence as a Commodity

Jason Wei

第一个趋势是智能将变成商品。获取知识或进行推理的成本和门槛将趋近于零。AI 进步有两个阶段。第一阶段是推动前沿,AI 还做不好,你在解锁新能力。比如看 MMLU 过去五年的进展,性能逐渐提升。第二阶段是能力商品化。这里有个例子,y 轴是时间,x 轴是达到特定 MMLU 性能的成本(美元)。趋势是每年使用特定智能水平模型的成本都在下降。为什么这个趋势会持续?我的论点是这是深度学习史上第一次自适应算力真正起作用。回顾到去年,深度学习都是固定算力,不管问题难易。但现在有了自适应算力,你可以根据任务调整算力。这首先在 o1 上展示,一年多前发布,显示增加测试时算力解决数学问题,性能更高。自适应算力能持续降低智能成本,因为你不必一直扩大模型规模。对于简单任务,你可以用最便宜的算力。

So the first trend is that intelligence is going to become a commodity. The cost and accessibility of finding out knowledge or doing some reasoning is going to be driven towards zero. I would say there are two stages of AI progress. The first stage is when you're pushing the frontier. AI can't really do the thing well yet, and you're unlocking that new ability. You can see, for example, if you plot MMLU, a very common benchmark, over the past five years, you'll see gradual progress in performance. Then the second stage is once you have an ability, it becomes commoditized. Here's an example where the y-axis is time and the cost of getting a particular performance on MMLU in dollars. You can see the trend: every new year, the cost of using a model with a particular level of intelligence decreases. You might ask why this trend might continue. The argument I would make is that it's the first time in the history of deep learning that adaptive compute actually truly works. If you look at the entirety of deep learning up to last year, we were in a mode where the amount of compute used for a particular problem was fixed regardless of how hard the problem was, whether it was giving the capital of California or a very hard competition math problem. But now we're in an era of adaptive compute where you can vary the amount of compute used for your task. This was first shown with o1, which came out more than a year ago, showing that if you increase the amount of compute used at test time to solve math problems, the performance on that benchmark is higher. The reason adaptive compute means you can continue to decrease the cost of intelligence is that you don't have to keep scaling model size. For a very easy task, you can go to the limit of how cheap compute you'd have to spend.

Jason Wei

第二部分是获取公共信息的时间。x 轴是前互联网时代、互联网时代、聊天机器人时代和智能体时代。问找一条信息要多久。比如想知道 1983 年釜山的人口,前互联网时代你可能开车去图书馆翻百科全书,几小时。互联网时代搜索浏览,几分钟。现在基本即时。更难的知识,比如 1983 年釜山有多少对夫妇结婚,前互联网时代你可能要飞韩国翻书。互联网时代容易些但仍困难。聊天机器人时代容易一点。智能体时代我认为几分钟内就能找到。

The second part is you can think about the time to retrieve certain pieces of public information. On the x-axis you have pre-internet era, internet era, chatbot era, and then agents era. You can ask how long it takes to find a certain piece of information. For example, if you wanted to know the population of Busan in 1983, before the internet you'd probably drive to the library and look in encyclopedias, taking a few hours. In the internet era, you'd search and browse websites, taking minutes. Now it's basically instant. For a progressively harder piece of knowledge, like how many couples got married in Busan in 1983, before the internet you might have to fly to Korea and dig through books. Internet era it might be easier but still challenging. In the chatbot era it's a bit easier. And with the agents era, I think this could be found within minutes.

验证者定律 Verifier's Law

Jason Wei

第二个概念我称之为验证者定律:训练 AI 完成特定任务的能力与验证该任务的容易程度成正比。这是理解 AI 会在哪些领域成功、哪些领域挣扎的关键。对于容易验证的任务,比如有标准答案的数学题或可测试的代码,我们可以用强化学习等技术训练模型达到高性能。对于验证困难的任务,比如写小说或产生创意,训练 AI 就难得多,因为你很难判断输出好坏。这个定律解释了为什么我们在数学和编程等领域进展迅速,而在创意写作或科学发现等领域进展较慢。

The second idea is something I call verifier's law: the ability to train AI to do a particular task is proportional to how easy it is to verify that task. This is a key insight for understanding where AI will succeed and where it might struggle. For tasks where verification is easy, like math problems with known answers or code that can be tested, we can use reinforcement learning or other techniques to train models to high performance. For tasks where verification is hard, like writing a novel or generating creative ideas, it's much harder to train AI because you can't easily tell if the output is good. This law helps explain why we've seen rapid progress in areas like math and coding, but slower progress in areas like creative writing or scientific discovery.

智能的锯齿边缘 Jagged Edge of Intelligence

Jason Wei

最后一个概念是智能的锯齿边缘。AI 在特定任务上的能力和改进速度会因任务特性而异。有些任务对 AI 来说容易,比如下棋或翻译;有些则难,比如理解细微的人类情感或制定长期计划。锯齿边缘意味着 AI 的能力不是均匀的;它可以在某些领域超越人类,在另一些领域则不如人类。这对思考 AI 如何影响不同工作和行业很重要。比如,AI 可能擅长分析法律文件,但不擅长谈判合同。理解这个锯齿边缘有助于我们预测 AI 会在哪里产生最大影响,以及哪里仍然需要人类。

The final idea is the jagged edge of intelligence. Both the capability and the rate of improvement of AI on a particular task will vary based on certain properties of those tasks. Some tasks are easy for AI, like playing chess or translating languages, while others are hard, like understanding nuanced human emotions or making long-term plans. The jagged edge means that AI's abilities are not uniform; it can be superhuman in some areas and subhuman in others. This is important for thinking about how AI will impact different jobs and industries. For example, AI might be great at analyzing legal documents but terrible at negotiating a contract. Understanding this jagged edge helps us predict where AI will have the biggest impact and where humans will still be needed.

智能作为商品 Intelligence as a commodity

Host

然后你可以看到更困难任务的趋势。比如,如果你想问‘1983 年亚洲人口最多的 30 个城市,按当年结婚数量排序’,我认为现在可能几小时就能完成,但在互联网时代之前,回答这个问题需要几周。这里有个例子:这不是一个超级简单的问题,比如 1983 年釜山有多少人结婚。GPT-3 做不到,但 OpenAI Operator 可以,因为你需要访问一个叫 Kosis 的数据库,点击直到找到正确的数据库查询,然后才能找到答案。我们在 OpenAI 尝试衡量这一点的方法之一是通过一个叫 BrowseComp 的基准测试,代表浏览竞赛。它包含一系列问题,一旦你有了答案,验证很容易,但解决一个问题实际上需要很长时间。一个例子是:这里有一堆足球比赛的约束条件,然后找到符合所有约束条件的比赛。我们让很多人做这些问题。平均而言,有些问题人类需要两个多小时才能解决。许多问题,如果你看规模,人类在两小时内无法解决的比例远高于他们实际解决的。但你可以看到 OpenAI 的深度研究模型能解决大约一半的问题。所以进展相当不错。总结一下,你应该把智能看作一种商品。一旦我们通过 AI 实现了能力,成本将趋近于零。我认为这个趋势会继续。即时知识的概念:任何公开可用的信息,你都能立即获取。一些影响:一是基于知识的任意入门壁垒的领域民主化。编程绝对是其中之一——氛围编程就是一个很好的例子。个人健康是另一个。过去,如果你想做生物黑客实验,你会去看医生,他们会说‘就试试我告诉你的方法’,而不会帮你理解如何自己做实验。但现在 ChatGPT 几乎能给你任何好医生能提供的信息。另一个影响是私人内幕信息的相对价值略高。鉴于公共信息的成本现在更低,私人信息的相对价格就会更高。例如,知道不在市场上的可售房屋——那信息更有价值。最后,我们将实现无摩擦的信息访问。不再访问公共互联网,你会得到一个个性化的互联网,无论你想知道什么,都会有一个个性化的网站展示给你。

And then you can see the trend for even harder things. For example, if you want to ask, 'Of the 30 most populated cities in Asia in 1983, sort them by number of marriages in that year,' I think that's something that can be done now maybe in hours, but in the pre-internet era it would take weeks to answer. Here's an example: this is not a super easy question, like how many people got married in Busan in 1983. GPT-3 was not able to do this, but OpenAI Operator can, because you have to go to a database called Kosis and click around until you get the exact database query that's correct, and then you can find the answer. One way we've tried to measure this at OpenAI is via a benchmark called BrowseComp, which stands for browsing competition. It's a bunch of questions where once you have the answer, it's easy to verify, but it actually takes a long time to solve one of these problems. An example would be: here are a bunch of constraints for a soccer match, and then find the match that actually fits all those constraints. We asked a lot of humans to do these questions. On average, some questions took more than two hours for a human to solve. Many of the questions, if you look at the scale, humans were unable to solve in two hours versus the ones they actually solved. But you can see the deep research model from OpenAI can solve around half of them. So pretty good progress. To summarize, you should sort of see intelligence as a commodity. Once we achieve abilities with AI, the cost will be driven towards zero. I think the trend will continue. The idea of instant knowledge: anything that's publicly available information, you'll be able to access instantly. A few implications: one is democratization of fields that were previously gated by arbitrary barriers of entry based on knowledge. Coding is definitely one—vibe coding is a great example. Personal health is another. In the past, if you wanted to do biohacking experiments, you'd go to the doctor and they'd say, 'Just try the thing I told you,' and wouldn't help you understand what to do for your own experiment. But now ChatGPT can give you almost any information a good doctor could. Another implication is the slightly higher relative value of private insider information. Given that the cost of public information is now lower, the relative price of private information becomes higher. For example, knowing houses not on the market that can be sold—that information is more valuable. Finally, we'll have frictionless access to information. Instead of accessing a public internet, you'll get a personalized internet where whatever you want to know, there'll be a personalized website to show you.

验证不对称与验证者定律 Asymmetry of verification and verifiers law

Host

第二个想法:验证不对称性和验证者定律。验证不对称性是计算机科学中一个非常常见的概念:对于某些任务,验证解决方案比找到解决方案容易得多。例子:数独很难解,但如果有答案,验证很容易。另一个是编写运行 Twitter 的代码。显然需要一支由数千名工程师组成的团队,也许还有数百个埃隆在运营公司,才能生成网站,但验证它是否正常工作要容易得多——你只需渲染并点击。竞赛数学题:有些题解起来和验证一样容易,所以这是中间情况。数据处理代码:有时不同。如果你写一个脚本来处理数据,写起来很容易,但如果给我别人乱糟糟的代码,弄清楚他们的代码在做什么可能比我自己写或检查他们的代码花的时间更长。写一篇事实性文章:提出看似真实的说法很容易,但核实某个特定说法可能极其繁琐。这是相反不对称性的一个例子:容易生成一篇看似真实的文章,但验证它是否好需要更长时间。这甚至延伸到创造新饮食之类的事情。我可以断言最好的饮食是只吃野牛,这花了我 10 秒钟。但要验证这个说法是否真实,你需要大样本量,等待长期结果,而且可能充满噪声。所以这些是验证不对称性在谱系上的几个例子。你可以这样可视化:x 轴是生成的容易程度,y 轴是验证的容易程度。数独生成难度中等,但验证容易。Twitter 生成难,验证较易。最佳饮食生成容易,验证难。还有中间的东西。有趣的是,你可以通过提供特权信息来改善任务在这个平面上的位置。例如,在竞赛数学中,如果我提供答案键,检查就变得非常容易。或者如果你在写代码,我提供测试用例,就像我们在 SWE-bench 中做的那样,检查也变得非常容易。所以想法是:对于某些任务,你可以事先做一些工作,增加验证的不对称性。这就引出了我所说的验证者定律,或者如果你作为科学家对‘定律’这个词敏感,可以叫它验证者规则。

Second idea: asymmetry of verification and verifiers law. Asymmetry verification is a very common idea in computer science: for some tasks, it's much easier to verify a solution than to find it. Examples: Sudoku is very difficult to solve but easy to verify if you have the answer. Another is writing the code to run Twitter. Obviously it takes a team of thousands of engineers, maybe hundreds of Elons running the company, to generate the website, but to verify it's working is much easier—you just render it and click around. Competition math problems: some are just as easy to solve as to verify, so that's a middle case. Data processing code: sometimes it's different. If you write a script to process data, it's pretty easy to write, but if you give me someone else's messy code, it might take longer to figure out what their code is doing than to write my own or check theirs. Writing a factual essay: it's pretty easy to make feasibly true claims, but fact-checking a particular claim might be extremely tedious. That's an example of opposite asymmetry: easy to generate a feasibly true essay, but takes longer to verify it's good. This extends to things like creating a new diet. I can assert that the best diet is to only eat bison, which took me 10 seconds. But to verify whether this is true, you need a large sample size, wait for long-term outcomes, and it might be noisy. So these are examples of asymmetry of verification across a spectrum. You can visualize it: x-axis is how easy to generate, y-axis is how easy to verify. Sudoku is medium difficulty to generate but easy to verify. Twitter is hard to generate, easier to verify. Best diet is easy to generate, hard to verify. And things in the middle. The interesting point is you can improve where a task is on this plane by giving privileged information. For example, in competition math, if I provide an answer key, checking becomes very easy. Or if you're writing code and I give you test cases as we do in SWE-bench, checking also becomes very easy. So the idea is: for certain tasks, you can do some work beforehand and increase the asymmetry of verification. This brings us to what I'm calling verifiers law, or if the word 'law' triggers you as a scientist, you could call it verifiers rule.

验证者定律与验证不对称 Verifier's Law and Asymmetry of Verification

Jason Wei

但我基本想主张的观点是,训练 AI 解决任务的能力大致与该任务的可验证性成正比。这意味着任何可解且易于验证的任务最终都会被 AI 攻克。更具体地说,我认为可验证性取决于以下五个因素:第一,是否存在客观标准来判断回答的好坏?第二,验证速度有多快?第三,能否同时验证一百万个不同的候选回答?第四,噪声是否低?第五,是否获得连续奖励?也就是说,你只区分通过和不通过,还是给出完整的回答质量谱系?大多数 AI 基准测试本质上都易于验证,这正是验证者定律的一个很好的实例——我们可以看到,过去五年我们关心的所有基准测试都相对较快地被 AI 解决了。

But basically the claim I'd like to assert is that the ability to train AI to solve a task is basically proportional to how easily verifiable the task is. And the implication is any solvable easily verifiable task will eventually be conquered by AI. To be a little more concrete, verifiability I would say is a function of these five things. One, is there objective truth to what's a good response and what's a bad response? Two, how fast is it to verify? Three, can you verify like a million different proposed responses at once? Four, is there low noise? And five, do you get continuous reward? So do you differentiate only between passing and non-passing or do you give the entire spectrum of response quality? Most AI benchmarks by definition are easy to verify, and that's a nice instantiation of verifier's law where you can see that all the benchmarks we've cared about in the past five years have been solved by AI relatively quickly.

示例:AlphaEvolve Example: AlphaEvolve

Jason Wei

一个利用验证不对称性的绝佳例子是 DeepMind 的 AlphaEvolve,如果你们还没读过,我强烈推荐。他们通过投入大量算力进行采样,并结合智能算法,成功解决了符合这种验证不对称性的任务。这些任务包括数学问题和算力优化等。举个例子:有一个数学问题,要求放置 11 个六边形,使得能画出最小的外接六边形。这显然满足所有五个标准:客观性——直接画图就能验证;可扩展性——计算即可验证;低噪声——每次验证结果相同;连续奖励——六边形的大小直接衡量了答案的好坏。算法的工作原理是:使用一个大语言模型,采样一批候选解,根据任务定义进行评分,然后选取最好的解作为灵感输入模型进行下一轮采样。经过大量算力和迭代,性能随时间提升。他们聪明的地方在于绕过了泛化问题。在大多数深度学习中,我们关心从训练到测试的泛化——要么是同一任务的新样本,要么是新任务。但这里他们选择训练和测试相同的问题,你只需要找到单个特定问题的答案。这绕过了很多问题。因此,你必须选择那些有可能得到比已知答案更好的问题。

One great example of leveraging asymmetry of verification, I encourage you all to take a look if you haven't read this yet, would be AlphaEvolve from DeepMind. Basically they were able to solve tasks that fit this asymmetry of verification by just spending a lot of compute via sampling and a smart algorithm. It includes a bunch of tasks in math and optimizing usage of compute, etc. An example is this math problem: find the placement of these 11 hexagons where you can draw the smallest outer hexagon around it. This clearly satisfies all five criteria. It's objective—you can just plot it to check the answer. It's scalable to check because it's computational. It's low noise—you get the same result every time you check. And it's continuous reward—the size of the hexagon gives you a direct measure of which answer is better. The algorithm works by taking a large language model, sampling a bunch of candidate solutions, grading them because they have a way of grading by definition of the task, then taking the best one and feeding it into the model for the next round of sampling as inspiration. Once you spend a lot of compute and iterations, performance increases over time. The smart thing they did is they sidestepped generalization issues. In most of deep learning, we care about generalization from training to test—either same task unseen example or unseen task. But here they pick problems where train and test are the same, so you just want the answer to a single particular problem. That allows you to sidestep a lot of these issues. So you have to pick problems where you can possibly get a better answer than what you already know.

验证者定律总结 Summary of Verifier's Law

Jason Wei

总结一下,验证不对称性:每当你面对一个任务,我建议思考它在这个不对称平面上的位置。验证者定律指出,任何非常容易验证的事物最终都会被 AI 解决。评估基准和 AlphaEvolve 就是例子。一些启示:首先,最先被自动化的任务将是那些非常容易验证的任务。其次,如果你想创办公司,一个新兴领域就是提出衡量事物的方法,然后让 AI 来优化它们。

To summarize, asymmetry of verification: whenever you have a task, I would recommend thinking about where it is on this plane of asymmetry. Verifier's law or rule states that anything that's very easy to verify will eventually be solved by AI. Evaluation benchmarks and AlphaEvolve are some examples. Some implications: first, the first tasks that will be automated are those that are very trivial to verify. Second, one of the nascent areas, if you want to make a company, is coming up with ways to measure things that can then be optimized by AI.

智能的锯齿边缘 The Jagged Edge of Intelligence

Jason Wei

最后一件事:智能的锯齿边缘。如果你问 AI 将如何改变世界,人们的观点差异很大。我的一位前同事 Boaz 说,东海岸的人低估了变化的规模——他们认为当前模型做不到这个,不太考虑发展轨迹——而湾区的人可能低估了部署模型的一些摩擦和时间延迟。另一位我喜欢的 Run(你应该关注他)说,现在没有人应该给出或接受任何职业建议。每个人都普遍低估了变化的范围和规模以及未来的高方差。你在 Meta 的 L4 工程师朋友跟你说‘兄弟,CS 学位没用了’,他其实不知道。所以关于 AI 将如何影响不同行业,显然存在广泛的观点。一个长期存在的假设是快速起飞的想法:一旦你在某方面超越人类,你就会突然变得比人类强大得多,在短时间内获得巨大的智能。我认为这很可能不会发生。快速起飞可能是一个简化的版本:很多年你都无法用 AI 训练 GPT-n+1,然后在第二年突然就能做到。但我认为更像是渐进式的进步:每年你都在逐步推进 AI 自我改进的能力。也许第零年你甚至无法获取代码库;第零年中期你可以训练一些东西,但结果并不惊人。之后它可以自主训练,但效果不如交给 10 个最好的研究人员。有时你仍然需要人类干预才能让它持续良好运行。所以我认为自我改进能力更像是一个谱系,而不是二元的。另一个原因是,我相信自我改进的速度应该按任务来看待。存在一个不同任务的谱系。

Last thing: the jagged edge of intelligence. If you ask how AI will change the world, people have pretty different views. One of my ex-colleagues Boaz says that east coasters underestimate the magnitude of change—they think the current model can't do this, they don't think about the trajectory as much—whereas in the Bay Area maybe we underestimate some of the friction and time lag to deploy models. Another person, Run, who I like and you should follow if you don't, says that nobody should give or receive any career advice right now. Everyone has broadly underestimated the scope and scale of change and the high variance of your future. Your L4 engineer buddy at Meta telling you 'bro CS degrees are cooked' doesn't know. So there's clearly a wide range of opinion on how AI will affect different industries. One hypothesis that's been around for a long time is the idea of a fast takeoff: once you pass humans in a certain thing, you'll suddenly become much stronger than humans, with a short takeoff duration where you gain a huge amount of intelligence. I would say this is probably not going to happen. Fast takeoff might be a simplistic version where for many years you can't train GPT-n+1 with AI, and then in year two you can suddenly do it. But I think it would be more like gradual progress: every year you make gradual progress towards AI being able to self-improve. Maybe year zero you can't even get the codebase; halfway through year zero you can sort of train something but the result isn't amazing. Then after that it trains autonomously but it's not as good as if you gave it to the 10 best researchers. Sometimes you still need humans to intervene to keep it running well. So I think it's more of a spectrum of self-improvement ability rather than a binary thing. Another reason is that I believe the self-improvement rate should be looked at in a per-task fashion. There is a spectrum of different tasks.

智能的锯齿边缘 Jagged Edge of Intelligence

Jason Wei

你可以这样理解。存在一个锯齿状边缘,对吧?在峰值处,是我们目前特别擅长的问题,比如困难的数学题、某些竞赛编程。然后还有一些低谷,有点奇怪。例如,很长一段时间里,ChatGPT 会说 9.11 大于 9.9。还有像说弗林吉特语这种语言,我认为只有几百名美洲原住民会说。我不认为 ChatGPT 能做好。我不认为会出现这种情况:你有一个自我改进的模型,然后突然就能做好所有事情。我认为更可能的是右边这种情况:每个任务都有不同的改进速度。所以有些任务改进很多,因为它们可验证,并且你找到了能快速改进这些任务的算法;而其他任务,比如说话,可能受限于你去美洲原住民保留地记录语言。我认为这些任务不会改进得那么快。

Um and you can sort of think of it like this. So there's some jagged edge, right? So at the peaks you have problems that we can currently do especially well like hard math problems. Some types of competition coding. And then there are also these valleys that are a little bit weird. So for example, for a long time ChatGPT would say that 9.11 was greater than 9.9. And then you have things like speaking Flingit, which is a language I think only a few hundred Native Americans can speak. I do not think ChatGPT can do that well. And I don't think we'll be in this case where you have a self-improving model and then suddenly you can do everything well. I think you're more likely to be in this case on the right, which is that every task will have a different rate of improvement. So maybe some tasks improve a lot because they're verifiable and you come up with an algorithm that can improve those tasks quickly, and then other tasks like speaking, which is maybe bottlenecked on you going to a Native American reservation and documenting what the language is. I don't think those tasks will improve as quickly.

AI 改进速度启发法 Heuristics for AI Improvement Speed

Jason Wei

所以我会说几个启发式方法,用来思考 AI 在某些任务上的改进速度。一是 AI 擅长数字任务。我不知道,这个作业机器漫画实际上相当准确,考虑到它是在 1981 年创作的,描述了 AI 的工作原理。但你知道,《我,机器人》显然我们还没有,也许很快会有,但还没有。我认为 AI 在数字任务上发展更快的关键原因就是迭代速度。因为当你做数字任务时,你可以比用真实机器人做实验更容易地扩展算力。另一个显而易见的是,对人类更容易的任务往往对 AI 也更容易。所以你可以有一个难度谱系。我认为即将到来的一件事是,AI 能够完成人类可能因为生物大脑的局限性而无法完成的任务。例如,预测乳腺癌的发生是一个可能完成的任务,如果你读过 1000 万张乳腺癌图像,你就能找到预测它的模式,但作为人类,我们活不了那么久,也没有足够的注意力去做。另一个极其简单的启发式方法是,当数据丰富时,AI 往往表现更好。所以你可以看到一个非常清晰的例子:查看语言模型在不同语言上的数学表现。如果你绘制该语言的频率(即我们有多少数据)与表现的关系,趋势非常明显:数据越多,任务表现越好。然后,这个规则的一个特殊解锁或例外是,如果你有一个单一的目标指标,那么你可以采用 AlphaGo 或 AlphaZero 的策略,通过强化学习生成合成数据。所以我的前同事 Danny 发过一条很好的推文:只要任务提供清晰的评估指标,可以作为微调时的奖励信号,任何基准测试都可以快速解决。

So I will say a few heuristics for how to think about how fast AI will improve at certain tasks. One is I think AI is good at digital tasks. So I don't know, this homework machine comic is actually pretty accurate given that it was created in 1981 for how AI works. But you know, I robot obviously we don't have, maybe we'll have that soon but we don't have that yet. I think the core reason why development of AI has been so much faster on digital tasks is just iteration speed. Because when you're doing a digital task, you can scale up compute a lot more easily than you can scale up experiments using a real robot. Another kind of obvious one is that tasks that are easier for humans tend to be easier for AI. So you can have a spectrum of how hard things are for a human. One sort of thing that I think is coming is the ability to do tasks that maybe humans can't do because of fundamental limitations we have as humans with a biological brain. So maybe for example predicting the occurrence of breast cancer is a task that is possible if you can, let's say, if you've read 10 million images of breast cancer you could find the one pattern that allows you to predict it, but as humans we don't live long enough or have enough intention to do that. Another extremely simple heuristic is that AI tends to thrive when data is abundant. So you can think of a very clear example where you can look at math performance of language models in different languages. And if you plot the frequency of that language, in other words, how much data we have, versus performance, it's a pretty clear trend that the more data you have, the better you're going to do on that task. And then maybe a special unlock or exception to that rule is if you have a single objective metric then you can do the AlphaGo or AlphaZero tactic where you can generate essentially synthetic data via reinforcement learning. So my former colleague Danny do had this nice tweet: any benchmark can be rapidly solved as long as the task provides a clear evaluation metric that can be used as a reward signal during fine-tuning.

预测 AI 能力时间线 Predicting AI Capability Timelines

Jason Wei

我有一张之前展示过的表格,你可以用这三个启发式方法来预测 AI 何时能完成某些任务。比如,翻译前 50 种语言,简单,已经完成;调试基本代码,我认为在 2023 年完成,对人类中等难度,数字任务,数据容易获取。竞赛数学,对人类困难,但数字任务且数据容易获取,所以已完成。进行 AI 研究,对人类困难,数字任务,但数据或创建数据不太容易。所以我可能会说 2027 年。我只是在编数字。化学研究对人类也很困难,但不是数字任务。所以可能比 AI 研究晚。制作电影对人类非常困难,但数字任务且数据容易获取。所以可能 2029 年。股市预测对人类非常困难,数字任务且数据容易获取。所以这个我不太确定。翻译成它(某种语言),对懂的人类容易,数字任务,但数据不太容易获取。所以我认为 AI 能做的可能性很低。修理水管,对人类中等难度,但不是数字任务。数据容易获取?不太确定。理发是另一个我认为对 AI 很难的任务。然后如果你遇到真正困难的任务,比如传统手工地毯制作,对人类非常困难,需要一队人花一个月做一块地毯。不是数字任务,数据不容易获取。我认为他们不会很快做到。带女朋友去约会让她满意,不可能,不是数字任务,数据不容易获取。所以我认为我们不会有这个。我认为我们还能继续做一段时间。

And I have this table that I've shown before of you can just sort of use these three heuristics to predict when AI will be able to do certain things. So you say translation top 50 languages easy already done, debugging basic code did in 2023 I would say, and it's something that's medium difficulty for humans, digital, easy to get data. Competition math hard for humans but digital and easy to get data so done. Conducting AI research hard for humans and digital but not super easy to get or create data. So I would say maybe 2027. I'm just making up numbers here. Chemistry research is also hard for humans, but it's not digital. So I would say probably later than AI research. Making a movie very hard for humans, but digital and easy to get data. So maybe 2029. Stock market prediction very hard for humans, digital and easy to get data. So I'm not really sure about this one. Translation to it. Easy for humans that know how to do it, digital, but not very easy to get data. So I would say probably a pretty low chance that AI will be able to do this. Fix your plumbing. I would say medium difficulty for humans, but it's not digital. Easy at data, not really sure. Hair dressing is another one that I think will be pretty hard for AI. And then if you get the really hard ones, like for example, traditional spec carpet making, very hard for humans. It takes a team of people like a month to make a rug. Not digital, not easy to get data. I don't think they will do this anytime soon. Taking a girlfriend on a date that she's happy with. Impossible, not digital, not easy to get data. So I don't think we'll have this. I think we'll be in business for a while.

总结与启示 Summary and Implications

Jason Wei

好的。总结一下,智能的锯齿状边缘,我不认为会有快速的超级智能起飞,因为每种任务都有不同的能力和改进速度。AI 的影响将在满足某些属性的任务上最大,即数字任务、对人类容易、数据丰富。好的,那么对于影响,我认为某些领域将被 AI 极大加速。软件开发显然是其中之一。其他领域可能保持不变,比如理发。好的,很好。总结一下,智能和知识将变得快速而廉价。第二,验证者定律,测量是 AI 进步的驱动因素。最后,智能的边缘是锯齿状的。好的,很好。我就在这里结束。我有一个反馈表。如果你想给我的演讲提供反馈,我会阅读。还有,很高兴在 Twitter 上联系。

Okay. So to summarize, the jagged edge of intelligence, I don't think there will be a sort of fast superintelligence takeoff because every sort of task has a different capability and rate of improvement. The impact of AI will be largest on tasks that meet certain properties, namely they're digital, easy for humans, and data abundant. Okay so for implications, I think certain fields will be extremely heavily accelerated by AI. So software development is obviously one of those. And then other fields will probably remain untouched like hairdressing. Okay great. So in summary, intelligence and knowledge will become fast and cheap. Number two, verifiers law, measurement is a driving factor of AI progress. And then finally, the edge of intelligence is jagged. Okay, great. I will end here. I have a feedback form. If you want to give feedback on my talk, I will read it. And yeah, happy to connect over Twitter.

互动版:逐字朗读 + 针对本期提问 →