AI Models Commoditization and Public Market Test
打开互动全文版(中英对照 + 朗读 + 问答)→Benedict Evans 探讨前沿 AI 模型是否正在商品化、缺乏网络效应,以及随着 OpenAI 和 Anthropic 走向公开市场,投资者真正购买的是什么。
Benedict Evans discusses whether frontier AI models are becoming commoditized, the lack of network effects, and what public investors are actually buying as OpenAI and Anthropic head towards public markets.
当你自动化你的工作方式时,你总能看到那些即将消失的工作,因为它们就在眼前,但你不知道新工作会是什么。人类的需求是无限的。现在有多少人靠制作播客谋生?想象一下 10 年前预测这一点。在市场演化的某个阶段,如果你还在争论那个问题,你就是个傻瓜。但在初期阶段,你可能对其中一些问题有看法。你甚至可能没问对问题。我认为这就是我们今天面对这些东西时的处境。
When you automate your way work, you can always see the jobs that are going to go away because they're right there and you don't know what the new jobs are going to be. Human needs are infinite. How many people are earning a living from making podcasts now? Imagine predicting that 10 years ago. There's a stage in the evolution of the market where if you're still arguing about that, you're an idiot. But there's a stage at the beginning where you might have opinions about some of these questions. You're probably not even asking the right questions. That I think is where we are with this stuff today.
欢迎收听 Analyze 播客,这是一档专注于剖析全球商业、科技和媒体脉搏的优质播客。我是 Bernard Leung。AI 周期不再只是关于更好的模型,它正在成为对基础设施、分发和能源的资本市场考验。Anthropic 和 OpenAI 正走向公开市场,SpaceX 正在准备科技领域最具影响力的上市之一。Anthropic 今早发布的 Claude 3.5 表明,前沿能力现在与安全、接入和治理问题捆绑在一起。这些头条都指向同一个更深层的故事:巨额资本支出、日益商品化的模型、浅层但庞大的用户触达,以及一场关于谁控制无限智能策展层的新争夺。今天,我们检验模型层是否会变成一种公用事业,持久价值究竟在哪里积累,以及 AI 如何重塑全球的工作和竞争。今天与我一起的是常驻嘉宾 Benedict Evans,他是“AI is the World”背后的独立科技分析师。Benedict,欢迎回来。
Welcome to Analyze Podcast, the premium podcast dedicated to dissecting the pulse of business, technology, and media globally. I'm Bernard Leung and the AI cycle is no longer about just better models. It is becoming a capital markets test of infrastructure, distribution, and energy. Anthropic and OpenAI are moving towards public markets. SpaceX is preparing one of the most consequential listings in technology and Anthropic's launch this morning of Claude 3.5 shows how frontier capability now comes bundled with safety, access, and governance questions. These headlines point to the same deeper story: massive capex, increasing commoditized models, shallow but enormous user reach, and a new battle over who controls the curation layer over infinite intelligence. Today, we test whether the model layer becomes a utility, where durable value actually accrues, and how AI reshapes work and competition globally. With me today, recurring guest Benedict Evans, independent technology analyst behind 'AI is the World'. Benedict, welcome back.
谢谢邀请。
Thanks for having me.
那么,自从我们上次对话以来,你最近在忙什么?
So, since our last conversation, what have you been up to recently?
一直在坐飞机。人们想了解这个 AI 的东西。你可能听说过。
Been on lots of airplanes. People want to know about this AI thing. You may have heard about it.
是的,我想我们谈论 AI 的方式,对吧?我想直接切入问题。我的意思是,今年早些时候你开始每 6 个月更新一次你的演示文稿,对吧?而“AI is Eating the World”现在可能已经是 2026 版了。所以,我想先从一个最大的问题开始,那就是 OpenAI 和 Anthropic 以及公开市场测试。我订阅了你的通讯,其中一篇我非常喜欢的文章是关于 OpenAI 如何竞争的:前沿模型几乎接近我们所说的商品化,没有网络效应,现在每个人都在试图购买算力时间,而这并不是护城河。所以,OpenAI 和 Anthropic 都在朝着接近万亿美元市值的公开市场迈进。那么,如果买的不是护城河,公开投资者实际上在买什么?
Yeah, I guess the way we talk about AI, right? I just want to go straight to the question. I mean, I guess early this year you started updating your presentation every 6 months now, right? And 'AI is Eating the World' probably now is in the 2026 edition. So, I think I want to start off first with the biggest question, which is OpenAI and Anthropic and the public market test. So, one article that I really enjoy reading since I'm a subscriber of your newsletter is how OpenAI competes: is that the frontier models are pretty much near what we call commoditization with no network effects and now everybody is trying to buy compute time and this is not a moat. So, both OpenAI and Anthropic are now heading towards a public near trillions dollars capital market cap. So, what is the public investor actually buying if not a moat?
嗯,这里面包含几个不同的问题。第一个是,目前模型中没有明显的赢家通吃效应。有一个微妙但重要的区别:你是否在做一些别人无论多努力都做不到的事情,因为他们的业务性质,无论花多少钱都追不上?这就是谷歌在搜索领域的地位。微软花多少钱都追不上。YouTube、iOS、Windows 和 Instagram 也是如此——它们有内在的结构性原因,使得它们很难失去地位。而我们在大型语言模型中还没有看到类似的情况。当然,最大的挑战是我们不知道这个市场将如何演变。有些人也不认同这个说法,但你处于这个市场的非常早期阶段,你不知道它会如何演变。我们也不知道科学如何运作,也不知道科学将如何变化。所以,可能会发生一些事情导致网络效应,但目前还没有。因此,你面临的情况是,有大约三到六家公司投入大量资金,然后可能还有多达十几家公司愿意落后三到六个月。但没有什么机制意味着如果其中一家领先,它们就会进一步领先。所以,我们现在看到的是,Anthropic 让 Claude 运转起来,找到了产品市场契合点,编码方面也找到了产品市场契合点。并没有内在原因使得谷歌和 OpenAI 无法追赶并让自己的东西运转起来。它们可能做到,也可能执行失败,但这不是可以预测的结果。所以,这是一个基本观察。模型都差不多,有相同的基准分数和相同的评估结果。因为它们使用相同的基础设施、相同的训练数据和相同的算法。所以,这是意料之中的。那么,你就面临一个问题:首先,是否存在一些机制,比如资本的使用,能让其中一家拉开差距?其次,它们能向上走多远?模型就是整个体验吗?模型就是我们使用的吗?聊天机器人会成为通用用户界面吗?这是我们三年来一直在问的问题。
Well, there are several different questions embedded in that. The first of them is like at the moment there's no apparent winner-takes-all effect in models. There's a subtle but important distinction: are there things that you're doing that no one else can do no matter how hard they try, that they won't be able to do because of the nature of their business, that they won't be able to catch up no matter how much money you spend? Because that's where Google is in search. It doesn't matter how much money Microsoft spends, they can't catch up. And that's what happened with YouTube and with iOS and Windows and Instagram — there were inherent structural reasons why it's really difficult for them to lose their position. And we don't see equivalents of those in large language models yet. Now, the big challenge of course is we don't know how this market's going to evolve. Some people have a problem with that statement too, but you're at the very early stage of this market, you don't know how it's going to evolve. We also don't know how the science works and we don't know how the science will change. So, stuff may happen that means there are network effects, but right now there aren't. And so, you're in the situation where you've got, pick a number, sort of three to six companies spending a lot of money, and then maybe anything up to a dozen companies that are willing to be three to six months behind. But there's not some mechanic that means that if one of them gets ahead, they'll get further ahead. And so, what we see now is, you know, Anthropic got Claude working, and they've got product-market fit, and they got coding working, product-market fit. There's not some inherent reason why it's just impossible for Google and OpenAI to catch up with that, and to get their own things working. They may do, they may fail to execute, but there's not a kind of predictable outcome. And so, that's a sort of base observation. The models are all kind of the same with the same kind of benchmark scores, the same kinds of evals. Because they're using the same infrastructure, and the same training data, and the same algorithm. So, that's what you'd expect. So, then you get to a question: Firstly, would there be mechanisms, for example, use of capital that would allow one of them to pull ahead? Secondly, how far up the stack do they go? Is the model the whole experience? Is the model what we use? Does the chatbot become the universal user interface, which is a question we've been asking for three years?
是的。
Yes.
或者,主流的大众市场采用——不是每周用一次或每月几次,而是每天一直用——是否需要将模型封装在应用、用例、工具、上市策略、数据以及我们想到软件时的一切之中?你的手机每天使用 50 个不同的数据库。没错。但它们肯定都运行在 SQL 或其他东西上,但这并不是重点。那不是产品本身。那么,模型能向上走多远?智能体能向上走吗?还是模型必须作为 API 被其他人构建的东西使用?最终,如果有 500 个这样的东西,它们不可能都由 Anthropic 和 OpenAI 构建。不是因为无法让模型写代码,而是因为你需要构建业务,需要弄清楚它是什么,需要建立销售团队,需要做上市策略,需要处理写代码之后的所有其他事情。所以,我不知道为什么它们能比微软或苹果做得更好。这就引出了一个初步论点:好吧,模型是一种糟糕的用户体验。聊天是一种糟糕的用户体验,而模型是商品。
Or, does mainstream mass-market adoption — and adoption that's not using it once a week or a couple of times a month, but using it all the time every day — does that need this to be wrapped in apps and use cases and tooling and go-to-market and data and everything else that we think of when we see software, which is, you know, your phone has 50 different databases every day. That's right. But they will definitely run on SQL or whatever it is, but that's kind of not the point. That's not what the product is. And so, how far up the stack can the models go? Can an agent go up the stack? Or, do the models have to be APIs that are used by stuff built by other people? And in the end, if there are 500 of these things, they can't all be built by Anthropic and OpenAI. Not because you can't get the model to write the code, but because you've got to build the business, and you've got to work out what it would be, and you've got to build a sales force, and you've got to do go-to-market, and you've got all the other stuff that happens after you've written the code. And so, I don't know why they would be able to do that any more than Microsoft could or Apple could. And so that gets you to a kind of preliminary thesis: okay, models are kind of a crappy UX. Chats are crappy UX, and models are commodities.
需要大量的应用,模型公司不可能全部自己开发。因此,模型看起来最终会变成商品化的基础设施。
There needs to be loads of apps, the model company can't build all of those. And so the models kind of look like they're going to end up as commodity infrastructure.
嗯。
Mhm.
所以回到你关于定价的观点,显然我们现在正处于一个极端的价格紧缩时期。但我们应该继续推进。未来一年会有很多乘数效应四处涌现。你看,有万亿美元级别的资本支出正在投入。模型每年效率提升 50 倍、100 倍、200 倍。芯片也在不断进步。你不知道下一个模型会是什么样。下一个模型可能会高效得多。根据以往经验,我知道它们会消耗更多的 token。
So to your point about pricing, clearly right now we're in this moment of extreme pricing crunch. But we should just proceed. There's so many multipliers that are going to get flung all over the place in the next year. So, you've got a trillion dollars of capex coming down the pipe. The models get 50, 100, 200 times more efficient per year. The chips keep getting better. You don't know what the next model will be. The next model may be way more efficient. The next models, on past history, I know they use way more tokens.
是的。
Yes.
唯一真正找到产品市场契合点的,就是编程。想象一下,如果我们有一个真正有产品市场契合点、能让很多人愿意使用的东西,而不是编程——编程最终只有几千万用户,而不是几亿。所以所有这些杠杆都会被到处挥舞。
The only thing that's actually got product-market fit, really really got product-market fit, is coding. So imagine if we had something that had product-market fit that would actually have lots of people who wanted to do it, rather than coding, which is in the end like tens of millions of people, not hundreds of millions of people. So all those levers are going to get swung all over the place.
嗯。
Mhm.
那么最终可能会稳定在什么状态呢?我一直在想这个问题,你可以抛出一些类比。比如说,电信行业的情况是,移动网络有边际成本。科技界有很多人不理解这一点。移动网络有边际成本。如果有更多用户更频繁地使用,你就得建设更多网络,这需要花钱。结果就是,我们基本上每个月付 50 到 100 美元,持续了 20 年。虽然没那么简单,但我们付的钱大致相同。而我们使用的数据量增加了 1000 到 2000 倍。这是一个年收入万亿美元的行业,每年资本支出 2000 亿美元,但股票 20 年来毫无起色。
And so where might it settle out in the end? I was thinking about this, you can kind of throw analogies out there. So you can say, well, what happened in telecoms is that mobile has marginal cost. There's a bunch of people in tech that don't get this. Mobile networks have marginal cost. They have to, if you have more users using it more, you have to build more network and that costs money. So what happened is that basically we went to paying 50 or 100 dollars a month every month for 20 years. It's not as simple as that, but we pay roughly the same amount of money. We now use 1,000 to 2,000 times more data. It's a trillion dollar annual revenue industry, pays 200 billion dollars a year on capex, and the stocks have gone nowhere in 20 years.
不会吧,哇。
No, wow.
所以这是一个大行业,资本支出巨大,收入也很高,但几乎没什么利润。利润很少,作为投资者也不是什么好投资。因为所有价值都流向了上层。
So it's a big industry that spends a lot of money on capex and earns a lot of money, but there's not really any profit. There's not very much profit, and it's not a great investment as an investor. Because all the value went up the stack.
没错。
Correct.
所有酷的东西都是别人做的。
All the cool stuff was built by other people.
现在,你还可以做其他比较。比如云计算,回报相当不错,谷歌、微软和亚马逊之间也有差异化。但同样,没有网络效应。
Now, there's other comparisons you can make. So you can also point to cloud where there is a pretty good return and there is differentiation between Google, Microsoft, and Amazon. But there again, there's no network effect.
而且它们确实存在。所有企业都在使用这类服务,对吧?
And they are really there as well. All the businesses are using these type of services, right?
它们运行在 AWS 上,但 AWS 不会从你的 Uber 行程中抽成。所以它们实际上并没有从上层获得那么多价值或杠杆。它们只是一个关键组件。你还可以拿芯片来比较。芯片类比的价值在于,每一代都变得更贵。
They run on AWS, but AWS doesn't get a percentage of your Uber trip. So they don't actually get that much value or leverage further up the stack. They're just an essential component. You could point to chips. And the value of the chip comparison is that it gets more expensive with every generation.
英伟达。
Nvidia.
不,我不是说英伟达,我是说台积电之类的。
No, I don't mean Nvidia, I mean like TSMC and so on.
那也没错。
That is also correct.
半导体制造,由于洛克定律,每一代都变得更贵。所以现在基本上只有三家芯片公司,而三四十年前有几十家。
Semiconductor manufacturing, because of Rock's Law, gets more expensive every generation. And so now there's basically three chip companies, whereas 30 or 40 years ago there were dozens.
嗯。
Mhm.
光纤的类比很有意思,因为乍一看很蠢,因为 25 年前的光纤泡沫是远远超前于需求建设的。到处都是空光纤。而这次当然是在需求之后建设,比如你生产的每个 token 都能卖掉。然而,光纤类比的有趣之处在于大规模的去聚合。光纤发生的情况是,它被拆分开了。所以有管道、暗光纤、亮光纤、波长、IP 流,可能还有其他东西。然后有人在上面交易带宽。而且总有人拥有稍微新一点的设备,能提供稍微更便宜的东西。所以每个人都在以边际成本或更低的价格出售。
The fiber comparison is interesting because at first sight it's dumb, because the fiber bubble 25 years ago was built out way ahead of demand. There was all this empty fiber. Whereas this, of course, is being built out behind demand, like you can sell every token you can make. However, the interesting part of the fiber comparison is the massive disaggregation. So what happened with fiber was that you got split up. So you had ducts and dark fiber and light fiber and wavelengths and IP streams, and probably some other stuff. And then you have people trading bandwidth on top. And then there's always somebody who's got the slightly newer equipment with slightly more lower stuff. So everyone is selling it at marginal cost or less.
感觉 token 就像光纤一样。
Feels like tokens to fiber.
那么,问题就是:市场越去聚合,就越是这样。比如马克·扎克伯格说,‘如果我有闲置容量,我就卖掉。’人们试图交易它,还有人基于它构建 ETF 之类的东西。你越这么做,市场就越碎片化。与此同时,还有一个问题:你需要最新模型多久?你必须使用最新、最完整的模型多久?有多少用例可以使用 6 个月前的模型、3 个月前的模型、开源模型,或者运行得慢得多,或者在自己的设备上运行?有多少用例必须拥有完整的智能,在云端以高昂价格运行完整模型,消耗大量 token?
Well, so this is the question: the more the market gets disaggregated. So you've got Mark Zuckerberg saying, 'Well, if I've got spare capacity, I'll sell that.' And people trying to trade it and people building like ETFs on top of it. The more that you get, the more the market gets fragmented. And then in parallel, there's this other question of how long do you need the newest model? How long do you have to have the newest full-fat model? And for how many use cases can you use the 6-month-old model, or the 3-month-old model, or the open source model, or running it much more slowly, or running it on your own equipment? How many use cases have to have the full intelligence, the full model running in the cloud at the massive price, using lots of tokens?
嗯。对对对,没错。
Mhm. Yeah, yeah, yeah. That's right.
这就像智商钟形曲线的 meme。智商 50 的傻子和智商 150 的人都会说:‘嗯,商品通常不是高利润产品,算力每次都会变得更便宜。’然后你可以坐在中间试图分析,比如谷歌今年有多少 TPU?我这么说有点不公平,因为那很难算。但这实际上能告诉你什么呢?
And it's like the bell curve IQ meme. The dumb guy with 50 IQ and the person with 150 IQ both say, 'Well, commodities tend not to be high margin products and compute tends to get cheaper every time.' And then you can kind of sit in the middle trying to analyze, you know, how many TPUs does Google have this year? Which is, I'm being slightly unfair because that's a hard thing to do. But what does this actually tell you?
嗯。
Mhm.
我想起一个故事,抱歉我一直在独白,但我会……这个故事是我 1999 年刚当科技分析师时想到的。有一家公司销售电脑零件、电脑组件。所有这些。全部在线销售。
This is a story that occurred to me, and I'm sorry I'm monologuing, but I'll... The story that occurred to me when I was a baby technology analyst back in 1999. There was a company that sold computer parts, computer components. All this stuff. Sold all this online.
对对,组件,没错。早期的时候。
Yeah, yeah, components, yeah. The early days.
嗯。
Mhm.
嗯,这家公司最初是那种在电脑杂志后面有 20 页广告的公司,如果你还记得的话。很多价格信息。就是那种公司。你知道,90 年代电脑杂志的后面。比如一本 300 页的杂志,后面 150 页全是不同经销商插页,每份 20、30、40 页。他们基本上卖同样的东西,你翻来翻去,他们会说:‘嗯,他们有我想要的打印机,卖 260 美元,而他们卖 250 美元。’但这些——所以他们做了个网站。他们——我瞎编的——年营收 1 亿英镑。我们都去看了那家公司。回来的路上,我想另一家投行拿到了上市授权,估值大概 7 到 8 亿。回来的路上我们讨论这事,这位资深银行家说:‘嗯,这是个低利润的经销商,一次性销售。’所以没人明说。花了很多时间讨论亚马逊、互联网、消费者采用率、宽带、HTML、退货、定价和管理团队。这是一次性——低利润经销商,一次性销售。这就是一个论点。我想问的问题是:解释一下,为什么那些使用类似基础设施、类似模型、类似 token 来交付类似产品的人——而且会有三到六家,市场会碎片化,有更好更快更新更旧更便宜更贵的——为什么这里有定价权?
Um, and it had been converted from originally being one of those companies that had 20 pages in the back of a computer magazine, if you remember that. A lot with their prices. One of those companies. You know, the back of a computer magazine in the '90s. Like it was a 300-page magazine and the back 150 pages was one insert after another of 20, 30, 40 pages from different resellers. And they were basically selling the same stuff and you kind of page back and forth and they'd say, 'Well, they've got the printer I want for $260 and they're selling it for $250.' But these—so they make a website. They're doing—I'm making this up—they're doing 100 million quid a year in revenue. We all go up to see the company. On the way back, I think another investment bank had got a mandate to float it at like 7 or 800 million. And on the way back we're discussing it and this veteran banker says, 'Um, it's a low-margin reseller, one-time sales.' So nobody calls it obviously. A lot of time talking about Amazon and internet and consumer adoption and broadband and HTML and returns and pricing and the management team. It's a one-time—it's a low-margin reseller, one-time sales. And that's a thesis. And I suppose the question I'm asking is: explain to me why people who are basically using similar infrastructure, similar models, similar tokens to deliver a similar product—and there's going to be three to six of them and it's going to get fragmented with better and faster and newer and older and cheaper and more expensive—why is there pricing power here?
这就是关键。是的,我完全同意你的观点,而且你也很清楚地指出,唯一的用户界面——即使今年出现了智能体 AI——我们仍然在使用聊天机器人界面。
That's the point. Yeah, I really come to your point on that, and I think you're also very clear about the fact that the only UI—even with agent AI happening this year—we are still using the chatbot interface.
是的,仍然如此。
Yeah, it still does.
变化不大,对吧?
It's not much change, right?
它还是 DOS 系统。
It's still DOS.
是的。
Yes.
我的意思是,我认为这里还有另一条有趣的线索。那就是,如果你考虑经典的 S 曲线框架——技术如何演进。想想移动技术或互联网的发展方式。一开始有一段平坦期,它不起作用,没人关心,又笨又蠢。然后有一段时期,一切开始运转,曲线急剧上升,非常令人兴奋。
I mean, I think there's another strand that's interesting to pull on here. Which is that if you think about the kind of classic S-curve framing—how technology evolves. So, you think about the way that worked with mobile or the way that worked with the internet. There's a period at the beginning where it's kind of flat and it's not working and no one cares and it's dumb and stupid. And then there's a period where everything kind of starts working and the curve starts shooting up and it's really exciting.
没错。
That's right.
然后有一段时期,一切都完成了,已经做完了,曲线变平,有点无聊。
And then there's a period where it's all kind of done and it's been done and it flattens, kind of boring.
正确。
Correct.
这就是智能手机现在的状态。也是客机现在的状态。也是电动汽车出现之前汽车的状态。
Which is where smartphones are now. It's where airliners are now. It's where cars were until electric happened.
没错。
That's right.
你知道,今年的型号和五年前的型号没有太大区别。所有你想要的东西——一切都已做完,它变成了一种成熟、无聊的产品。我提到这个的原因是,当你处于那个阶段的开始时,非常不清楚这将如何运作,有很多潜在路径,因为大多数人们将要——需要人们——我该怎么说呢?你在 1985 年有辆车,比如你需要选择火花塞。你开车时车后面绑着一捆 50 条轮胎,因为每 3 英里就会爆胎。而且你通常带个机械师——如果你有钱买车,你也足够有钱雇个机械师——你需要他,因为车每 20 分钟就坏一次。然后什么都不管用。这非常令人兴奋,但什么都不管用。70 年代末、80 年代初的 PC 也一样。非常令人兴奋。如果你当时在用,你会觉得这太棒了。你说它不管用是什么意思?不,你知道,我们小时候的正常体验是,你在电脑上敲着敲着,抬头看屏幕,它死机了。然后你爬到桌子底下,拔掉电源再插上,希望只丢了不到 10 分钟的东西。所以一切都令人惊叹,但什么都不管用。让它工作所需的东西还不存在。你不太清楚这一切将如何演变。这就是 2005 年左右移动技术所处的阶段。
You know, there's not very much difference between this year's model and the model from five years ago. It's all the kind of stuff that you wanted—everything's kind of been done and it's become a kind of mature, boring product. The reason I mentioned this is when you're at the beginning of that phase, it's very unclear how this is going to work and there are many potential paths because most of the stuff that people are going to—that need people—how can I put this? You got a car in 1985, like you need to choose the spark plugs. And you drive along with a bundle of 50 tires on the back of the car because you get a puncture every 3 miles. And you generally have a mechanic with you—if you're rich enough to own a car, you're also rich enough to have a mechanic on staff—and you need it because the car breaks down every like 20 minutes. And then nothing works. It's incredibly exciting, but nothing works. Same thing with PCs in the late '70s, early '80s. Incredibly exciting. And if you were using it then, you thought this is great. And what do you mean it doesn't work? Like no, like you know, the normal experience when we were kids was you'd be tapping away in your computer and you look up at the screen and it had frozen. And you'd crawl under the desk and unplug it and plug it in again and hope that you hadn't lost more than like 10 minutes of stuff. And so like everything's amazing and nothing works. And the stuff that has to happen for it to work doesn't exist. And you don't quite know how all that's going to evolve. And so that was where mobile was in like 2005.
但你觉得这个周期会快得多。
But you feel that this cycle is going to be much faster.
我的观点是,是的,周期更快。是的。但这就是我们所处的位置——你知道,你完全有信心在 2005 年预测移动技术会如何发展吗?
My point is that yes, the cycle's faster. Yes. But that's where we are in the sense of—you know, are you completely confident that in 2005 you would have predicted how mobile was going to work?
不。
No.
甚至到 2010 年,还有很多人相信——科技界很多非常聪明的人都相信,这将会像 Windows 对苹果那样发展。苹果是封闭的,苹果会被消灭。
Or even in 2010, there were all these people who were convinced that—a lot of really clever people in tech were convinced that this was going to play out like Windows versus Apple. Apple was closed, Apple would get obliterated.
结果并非如此。
That didn't work out the same way.
事情并非如此,我们可以解释原因。我当时也不认为会发生,我是对的,但你知道,我在很多其他事情上错了。我以为 Windows Phone 可能有机会,但它没有,我也可以解释原因。我这么说是因为,在市场演变的某个阶段,你可以对事情如何发展有强烈而清晰的看法。而在市场的另一个阶段,如果你还在争论那个问题,你就是个白痴。比如如果你还在争论为什么会有人用 iPhone,或者你知道,你应该能在上面安装任何你想要的软件。那是 10 年前的争论了。注意点。但在开始阶段,你可能对其中一些问题有看法。你甚至可能没有问对问题。我认为这就是我们今天所处的境地。就像 80 年代的 PC,90 年代中期的移动互联网。我们随便选了 1997 年作为例子。
Not what happened, and we could explain why. And I didn't think that was going to happen and I was right, but you know, I've been wrong about plenty of other things. I thought Windows Phone might have a chance and it didn't, and I can explain why as well. The reason I'm saying this is like there's a stage in the evolution of the market where you can have strong clear opinions about how this is going to work. And there's a stage in the market where if you're still arguing about that, you're an idiot. Like if you're still arguing about why would anyone use an iPhone or you know, you should be able to install any software you want on it. Like that was an argument from 10 years ago. Pay attention. But there's a stage at the beginning where you might have opinions about some of these questions. You're probably not even asking the right questions. And that I think is where we are with this stuff today. Where the PC was in the '80s, where mobile the internet was in the mid-'90s. And we used like 1997 picking out that data random.
嗯。
Mhm.
我们处于移动技术 2006 年或 2010 年的阶段。嗯,有很多事情你根本不知道会如何发展,也不知道杠杆会是什么。移动技术有趣的地方在于,我们在 iPhone 之前就认为它已经起作用了。只是增长不快。而且看起来它似乎还行。
We're where mobile was in either 2006 or 2010. Um, where there's a lot where you just don't know how this is going to work and you don't know what the levers are going to be. The funny thing about mobile was like we thought it was working before the iPhone. It was just wasn't growing very fast. And it seemed like it was kind of working.
是的。
Yeah.
而且我们以为诺基亚会主宰世界,然后——
And we thought Nokia is going to dominate the world and then—
是的,诺基亚、Windows 和黑莓。当时并不清楚会有东西出现并彻底颠覆一切。我们都认为移动设备基本上就是 PC 的配件。
Yeah, Nokia and Windows and BlackBerry. And it wasn't clear that something was going to come along and just blow the roof off. And we all thought mobile was basically a PC accessory.
嗯。
Mhm.
当时大家问的是:你在移动互联网上能做什么?移动用例是什么?当然,实际发生的是,现在你说桌面互联网
And it was well, what do you do on the mobile internet? What are the mobile use cases? And of course, what actually happened is now you say desktop internet
嗯。
Mhm.
而移动设备让 PC 成了智能手机的配件。是智能手机。它就这样成了科技和移动使用的中心。2000 年或 2005 年没人这么说。
and the mobile made the PC a smartphone accessory. It's the smartphone. This is how it became the center of tech and the center of mobile usage. And no one in 2000 was saying that or 2005.
因为他们无法想象手机实际上就是个人电脑,而且能发展到那么大的市场规模,对吧?
Because they couldn't imagine that the mobile phone is actually the personal computer that actually can grow to that a large market size, right?
几乎就是字面意思。BlackBerry 的一位联合创始人说过一句很棒的话:他看着 iPhone 说,如果这东西成了,我们就是在跟 Mac 竞争。
Almost literally. I mean, this is a great quote from one of the co-founders of BlackBerry. He said he looked at the iPhone and if this thing works, we're competing with a Mac.
因为如果你不在科技圈,还是很难理解 iPhone 其实就是一台 Mac。
Because this is the thing it's still hard for if you're not in tech to understand that an iPhone is a Mac.
嗯。
Mhm.
它是一台 PC,广义上的 PC。而它当时竞争的是那些加了点算力的手机。
And it's a PC. PC in the general sense. Whereas it was competing with phones that had some compute added.
嗯。
Mhm.
然后突然之间,不,这完全就是一台 PC。它有 4 或 8 GB 的存储,这简直是大得离谱的存储量。而且它运行的是 PC 操作系统。所以,这就是为什么 BlackBerry 和 Nokia 彻底完蛋了。我以前打的比方是:你本来在做远洋邮轮生意,然后你看到喷气式客机,你觉得自己完全没有相关技能去造那玩意儿。
And then suddenly, no, this is just a whole PC. And it had like 4 or 8 gig of storage, which was just a ludicrously large amount of storage. And it was running a PC operating system. So, that was why BlackBerry and Nokia were completely screwed. The analogy I used to give was it was like you were in the ocean liner business and then you look at jet liners, and you think that you just don't have any relevant skills at all for making that thing.
让我换个方式提问。这是我作为一个东西方之间生活的人的观察。我发现这一批美国公司,你们 Anthropic 和 OpenAI,表现得更像中国公司。我来说说中国公司的特点。他们 996。根据 Ramit 过去两年的情况,我觉得所有公司都在周六报销,前提是大家疯狂工作。我知道有些 AI 研究人员就是这样。在中国,你看阿里巴巴不只是 eBay。它也是 PayPal。也是所有其他服务。
Let me point a question in another way. I take this from my observation being a person living between the East and the West. I find that this batch of US companies, your Anthropic and OpenAI, are behaving more like Chinese companies. And I'll give you some features of Chinese companies. They work 996. According to Ramit's last 2 years, I think all the companies are expensing on their Saturdays only if people are working like crazy. I know some of the AI researchers are. In China, you see Alibaba is not just eBay. It is also the PayPal. It is also all the other services.
更横向的竞争。
Much more horizontal competition.
对,横向的,而且他们试图更垂直整合。
Yeah. Horizontal, and they try to vertically integrate more.
你有一个更宽的矩阵,有很多行和列,很多公司填满了许多格子。
You got a much wider matrix with many rows and columns, and lots of companies have got lots of the cells filled in.
而且你看到这一批 AI 公司,他们不只是供应,比如 Anthropic 的 Claude Code 供应给 Cursor。现在它也开始跟 Cursor 竞争了。
And you see this batch of AI companies, they are not just supplying to like say Claude code Anthropic supplies to Cursor. Now it's also coming in to compete with Cursor as well.
嗯,我不这么认为。我不太确定。
Yeah, I don't think that's so. I'm not sure about that.
那你觉得呢?
So do you find that?
我觉得 996 这件事,对比在于:2015 年或 2020 年,当 Google 已经赢了,现金像消防水管一样流的时候,硅谷是什么样子?那里有很多人并不在工作
I think the 996 thing, the contrast there is just well, what was the Valley like in 2015 or 2020 when Google had won and had a fire hose of cash. And there were a lot of people there not working
免费水疗和宠物。
free spas and pets.
对,很多人工作不努力,因为反正钱就在那里。所以我不确定那是那回事。我认为这恰恰是我说的 S 曲线的本质:当一切都很棒很激动人心,你必须昨天就把所有东西建好。而且,在早期,你不知道它会变成什么样。
Yeah, a lot of people were not working very hard because, hey, the money's just there anyway. So I'm not sure that was that. I think that's just the nature of the S curve I was talking about, like when everything is amazing and exciting, and you have to build everything yesterday. Also, when it's early, you don't know what it's going to be.
对。
Correct.
这里有一个有趣的对比,但说得非常简单:Google 的回应是把它做成一个功能。现有玩家总是把它做成一个功能。
And there's an interesting contrast here, but put this very simplistically. Google's responses will make it a feature. Incumbents always make it a feature.
没错。深度研究是一个功能。
That's right. Deep research is a feature.
对。这也是 Apple 的做法:我们会把它加到搜索里,加到 Gmail 里。它会赋能新能力,让现有东西更好。Apple 现在也在这么做。
Yeah. This is also Apple, but we'll add it to search. We'll add it to Gmail. It will empower new capabilities and make the existing stuff better. This is also what Apple is doing now.
没错。
That's right.
而如果你不是现有玩家
Whereas for if you're not the incumbent
Anthropic 用的是核心工作。
Anthropic uses core work.
对。所以你没有地方把它作为功能添加。
Yeah. So, then you don't have places to add it as a feature.
没错。
That's right.
所以对于 Anthropic 和 OpenAI,问题是:“好吧,那这东西的用例是什么?你用它做什么?你怎么告诉人们他们能用它做什么?”
So, then for Anthropic and OpenAI, the question is, "Okay, well, what are the use cases for this? What do you do with it? How do you communicate to people what they would do with it?"
嗯。
Mhm.
而 OpenAI 去年年底的立场是:一切,无处不在,昨天就要。
And the position at OpenAI late last year was the line was everything everywhere yesterday.
对。
Yeah.
所以我们会试一个社交视频应用。我们会试一个应用商店。
So, we'll try a social video app. We'll try an app store.
我们再试一个。
We'll try another one.
我们再试另一个应用商店。
We'll try another app store.
对。
Yeah.
我们会试一个结账流程,电商结账流程。我们会试……我不知道。他们试了五六种不同的东西。就像一个应用
We'll try a checkout, an e-commerce checkout flow. We'll try I don't know. They were like half a dozen different things. Like an app
业务。
business.
各种东西,因为我们不知道它是什么,所以我们会都试一遍,然后在有效的那一个上加倍下注。结果发现有效的是软件开发。至少这是第一个有效的,而 Anthropic 先做到了。
Like all sorts of stuff because we don't know what it is and so we'll try all of them and double down on the one that works. It turned out that the one that works is software development. That's at least the first one that's working and Anthropic got that working first.
嗯。
Yep.
所以 OpenAI 有点在转向那个方向。
And so OpenAI is sort of pivoting to go into that one.
但 Anthropic 也间接开始扼杀 Cursor 的使用。所以你进入了一个有点像自相残杀的局面,因为这与美国互联网非常不同,在美国互联网,你总是在自己的车道上,比如 Google 是
But Anthropic also indirectly starts to kill off cursor usage. So you're getting into this place where it's like a cannibalization because this is very different from if I say the US internet, it's always you are in this lane like Google is
完全不是。Google 过去 20 年,人们过去 20 年一直在说 Google 在杀死这个创业公司,Apple 在杀死那个创业公司,Amazon 在杀死那个创业公司。如果你在做一个平台,这始终是过去五年人们争论的焦点:大科技公司杀死创业公司。如果你回头看看 Microsoft 和 Windows 的演变方式。
That's total. Like Google has spent the last 20 years people have spent the last 20 years saying Google's killing this startup or Apple's killing that startup or Amazon's killing that startup. If you're making a platform, this was always the whole 5 years of people arguing about competition like big tech companies killing startups. If you go back and look at the way Microsoft and Windows evolved.
对。
Yeah.
有一个时刻,我以前有一张幻灯片讲这个。有一个叫 Sideways 的程序
There's a moment I used to have a slide of this. There was a program called Sideways
对。
Yes.
你在 80 年代末可以买到。它大概 150 美元,按 80 年代末的美元算,所以现在大概 300 美元,不管通胀是多少。Sideways 的功能就是让你横向打印。就这些。
that you could buy in the late '80s. And it was like $150 in late '80s dollars, so like $300 now, whatever the inflation is. And what Sideways does is it lets you print out in landscape. That's it.
对,我记得。
Yep, I can remember that.
而那是 150 美元。
And that's $150.
嗯。
Mhm.
而今天那是 Windows 或 Mac OS 打印对话框里的一个复选框。所以重点是同样的事情,如果你买了
And today that's a checkbox in the print dialog box in Windows or in Mac OS. And so the point is the same thing if you bought
嗯。
Mhm.
电子表格拼写检查是另一个例子,一个不那么荒谬的例子。拼写检查曾是一个完整的软件类别。
Spreadsheet spell check another example, a less absurd example. Spell checking was a whole software category.
嗯。对,然后你得到了
Mhm. Yes, and then you got
你买 PC 杂志,它们会有 10 种不同拼写检查程序的集体评测。
You bought PC magazines and they would have a group test of 10 different spell checking programs.
它们都是两美元和 300 美元。它的工作方式是,你写好文档,然后保存,再打开拼写检查文档,指向你的文档。它会打开文档并运行拼写检查。所以我的意思是,有一整类东西会自然地集成在一起。那么问题来了,这种集成对语言模型能走多远?它能走多远?是基因编码吗?是事物固有的吗?你可以采取一种最大化的观点,认为这些模型能够做所有事情的方式是它们能自己写代码来做事情。
They were all two and $300. And the way it worked was that you would write your document and then you would save it and then you would open the spell checking document and point it at your document. It would open that and it would run spell check on that. So my point is there's a whole bunch of class of stuff that just naturally gets integrated. And so the question is how far does that integration go for an LM? How far does that go? Is it genetic coding? Is that inherent to the thing? You could take the maximalist view here and say that the way these models will be able to do everything is because they'll be able to write their own code to do stuff.
是的。
Yes.
所以你不会让模型从头开始用 token 自己做所有事情。相反,它会写一个 Perl 脚本来做。所以模型会写一段传统的常规软件来做那件事。因此,所有事情都将通过模型生成自己的中间代码层来完成。
So you won't have the model doing everything itself from first principles in tokens. Instead it'll make a Perl script to do it. So the model will write a piece of conventional traditional software to do that thing. And so everything will be done through the models generating their own intermediate layers of code.
是的。所以我回到你最喜欢的类比。你说在 1980 年代,人们拿到电子表格,但不知道用它做什么。但本质上,电子表格变成了 Tableau,以及人们在其上构建的各种软件,因为他们发现自己知道如何使用电子表格,对吧?对我来说,co-work 就相当于今天的电子表格。只不过它还能创建 Word 文档,创建代码脚本来帮我做你刚才说的事。也许是 UI 的变化,或者是因为你仍然需要通过告诉它你想要什么来编排,这让它变得困难。
Yeah. So I come back to your favorite analogy. When you say that in the 1980s people were given a spreadsheet, but they didn't know what to do with it. But essentially the spreadsheet became Tableau, all the different kinds of software that people built on top of because they discovered that they know how to use a spreadsheet, right? To me, co-work is the equivalent of a spreadsheet today. Except that it could create a word document, create a coding script to help me to do exactly what you were saying. Maybe that change in UI, or maybe it's because you still have to orchestrate by telling it what you want, which makes it difficult.
所以我认为这正在揭示正在发生的事情。
So I think this is realizing what's going on.
所以 co-work 在某种意义上就是我描述的,即你希望模型为你做一件事。它做那件事的方式可能是使用 Git,或者写一个 Excel 宏之类的。
So co-work in a sense is what I'm describing, which is that you want the model to do a thing for you. It may be that the way it does that thing is using Git, or writing an Excel macro or something.
然而它并没有做所有事情。
And yet it's not doing everything of that.
它并没有在 LLM 内部以 LLM 的方式做所有事情。LLM 仍然在使用工具。
It's not doing everything in the LLM as LLM. The LLM is still using a tool.
正确。
Correct.
然而,我不认为这解决了聊天机器人作为用户体验的根本问题。我认为有两个基本问题驱动着软件创建和公司创建:空白屏幕问题和锯齿前沿问题。空白屏幕问题是:你怎么知道你需要什么?你不会坐下来就知道应该有哪些按钮、哪些问题、哪些流程。大多数人无法坐下来精确地画出他们将要如何完成这项工作的流程图。当你购买软件时,你购买的是大量的机构知识和大量的思考:问题是什么、工作流应该是什么、它应该如何工作、每一步应该给你什么选项。你不是从第一性原理出发,自己坐下来闭上眼睛想,‘嗯,好吧,我该怎么解释我想要什么?’当你购买软件时,你得到的就是这些。他们已经为你解决了这些问题。而且大多数人不是那种意义上的工具创造者。一个非常优秀的平面设计师并不是设计 InDesign 的人。
However, I don't think that solves the underlying problem of the chatbot as UX. And I think there are two fundamental problems here that drive software creation, company creation. And that's the blank screen problem and the jagged frontier problem. The blank screen problem is how is it that you know what you need? You don't sit down and know what buttons there should be and what the questions are and what the flows are. Most people couldn't sit down and draw a flowchart of precisely how they're going to do this job. And when you buy software, what you're buying is a whole bunch of institutional knowledge and a whole bunch of thought as to what the problem is, what the workflow should be, how it should work, what options I should give you at each step. You're not working out from first principles and sitting down yourself, shutting your eyes and thinking, 'Hmm, okay, how am I going to explain what I want?' When you buy software, you're getting that. They've worked that out for you. And most people are not tool creators in that sense. The person who's a really good graphic designer is not also the person to design InDesign.
没错。
That's right.
所以我认为这是一大类问题。我认为第二个是锯齿前沿,我认为这比幻觉更好的表述。因为这些模型有些事做得很好,有些事做得不太好。即使是现在,每个新模型都一样。每个新模型的锯齿部分都在不同的地方。所以有些人说幻觉已经解决了,你会想,‘嗯,对你来说是的。’但对我来说不是。
So I think that's one broad set of problems. I think the second is the jagged frontier, which I think is a better phrasing than hallucination. In that these models do some things very well and other things not very well. Even now, each new model is the same. Each new model has the jagged bits in different places. So there are some people who say hallucinations are solved and you think, 'Well, for you they are.' But not for my use cases.
实际上,不。
Actually, no.
让我说完。显然我们还没有 AGI。因此,有些事这些模型能做,有些不能。但这以不可预测的方式出现。能做和不能做的事,我们不一定知道原因。你无法从命令行判断哪个能工作,哪个不能。作为一个普通用户,如果你不是每天读 AI 论文,你不会知道,‘哦,当然那个任务会做得很好。这个任务,你必须用完全不同的方式提问,而且可能仍然不对。’所以你有这个锯齿前沿,关于模型擅长什么。这不重要,因为我们还有一个锯齿前沿,关于我们凭直觉能理解它擅长什么或不擅长什么。然后还有一个锯齿前沿,关于我实际拥有的重要且有价值的用例。在软件开发中,这些大致对齐。但还有很多其他情况,它们不一定对齐。
Let me finish the point. Clearly we do not have AGI. Therefore, there's stuff these can do and stuff they can't. But that comes in unpredictable ways. The things that can and can't be done aren't something we necessarily know why. You can't tell from the command line which is going to work and which isn't. As a normal user, if you're not reading AI papers every day, you're not going to know, 'Oh, well, of course that task will work brilliantly. This task, you've got to ask in a completely different way and it may still not be right.' So you've got this jagged frontier of what the models are good at. That doesn't matter that there's a jagged frontier of what we can intuitively understand it's going to be good at or not. Then there's a jagged frontier of what actual use cases I have that are important and valuable. And with software development, those all kind of line up. But there's a bunch of others where they don't necessarily line up.
嗯。当然,还有 token 成本的问题,比如这个用例可能比那个用例用更多的 token,你可能没意识到为什么。或者我们甚至不知道为什么。而且那个用大量 token 的用例可能也不值多少钱。所以所有这些部分如何组合在一起,存在各种锯齿性。这就是为什么你雇佣软件开发者,你雇佣一家软件公司,他们坐下来解决了,‘好的,这就是它如何工作的。如果你这样做,如果你给它这些数据,如果你这样构建,如果你预加载这个,给它这个上下文,它就会工作,我们会确保如果你问那个,它会告诉你那行不通。’或者我们会告诉你,‘不,这部分你可能想检查那个数字。’你把它包装在工具中。但所有这一切就是说,你需要把它包装在工具、用例、市场推广和对实际问题的深入思考中,以及它如何工作。
Mhm. Never mind, of course, what the tokens cost as well, which is like this use case might use way more tokens than that use case, and you might not realize why. Or we might not even know why. And the one that uses loads of tokens might not be worth very much, either. So there are all these sorts of jaggedness in how all these bits fit together. And that's why you hire a software developer, you hire a software company who sat down and worked out, 'Okay, this is how it's going to work. It will work if you do it this, if you give it this data, if you frame it like that, if you pre-load it with this, you give it this context, and we'll make sure that we'll tell you if you ask it that that won't work.' Or we'll tell you, 'No, this bit you would probably want to check that number.' You wrap it in tooling. But all of that is to say, you need to kind of wrap it in tooling and use case and go to market and deep thought about what the actual problem is and how this is going to work.
嗯。我的意思是,这就是 SaaS 末日的问题,或者说 SaaS 末日的微妙之处在于,建立一家软件公司的难点不是写代码。
Mhm. I mean, this is the problem with the SaaS apocalypse, or the nuance of the SaaS apocalypse is that the hard part of building a software company is not writing the code.
是分发和网络。
It's the distribution and network.
是其他所有事情。是意识到那个问题的存在。
It's everything else. It's realizing that that problem exists.
因为有问题的人往往看不到问题。那些软件要解决的问题的拥有者,并没有意识到那是个问题。而且,你解决问题的方式——可能之前有三四个人尝试过但失败了。你找到了正确的方法,把问题转化成了另一个看起来不像原问题的问题。你还找到了正确的市场进入策略、正确的切入点和正确的数据获取方式,让整个事情运转起来。如果你在应付账款部门工作,或者在视频剪辑店做调色,那些都不是你的技能。你很擅长调色,但不擅长找出正确的切入点,甚至看不到问题的存在。
Because the people who have the problem don't often see it. The people who have the problem that the software is solving do not realize that's a problem. And also, the way you're solving it—maybe three or four other people have tried to solve it before and failed. And you've worked out the right way by turning it into a different problem that didn't look like that problem. And you worked out the right go-to-market, the right insertion point, and the right way to get the data to make the whole thing work. If you're sitting in accounts payable or doing color grading in a video editing shop, those aren't your skills. You're great at color grading, but not at working out the right insertion point or even seeing that the problem exists.
嗯。
Mhm.
所以,对我来说,这就是整个论点的挑战所在,即“模型会一路走到顶端”。除非你相信模型会 Scaling(规模扩张)到超级通用——不管你现在用什么术语——如果你真的相信模型能理解我刚才说的一切,知道所有内容,并在没有提示的情况下全部完成,那也行,但那样我们就有了比担心企业软件未来更大的问题。
And so, to me, that is the challenge of the whole thesis that says, "Well, the model will go all the way up to the top." Unless you believe that the models will just scale to be super general—whatever terminology you're using now—if you really believe that the models will be able to understand everything I've just said, know all of it, and do it all unprompted, then fine, but then we've got bigger problems than wondering about the future of enterprise software.
我认为更大的问题——这是作为一个每天都在实践的从业者说的——是我一直发现的一件事,就像你说的,这个 ChatGPT 的类比——我喜欢这个类比,实际上我也没法用别的方式表达。我觉得 AI 软件非常擅长用你以为想要的输出来欺骗你。演示很完美,但一旦进入生产环境,你就会开始看到缺陷在哪里。
I think the bigger problem—and this is speaking from being a practitioner doing this every day—is that one thing I always find is that, like you said, this ChatGPT analogy—I like that analogy and I actually couldn't articulate it differently. I feel that AI software is very good at deceiving you with the output you think you want. The demo is great, but once you get into a production situation, you start to see where the flaws come up.
它非常擅长。我是说,有各种各样的……
It's very good. I mean, there's all sorts of different...
如果你能描述一下。是的。
If you can describe this. Yeah.
它非常擅长给你一个看起来像是好答案的东西。
It's very good at giving you what a good answer would probably look like.
没错,但那不是好答案。
That's right, but it's not the good answer.
但那可能是也可能不是你想要的。这取决于用例、你是谁以及你为什么问。举个简单的例子:几年前,当这个技术刚开始起作用时,我去一个活动演讲,他们要我提供一份长传记,而我没有。我用 ChatGPT 生成了一个长传记,里面全是错误。但我意识到:如果我用 ChatGPT 来做这件事,我可以在 30 秒内修正它,因为我是自己的专家。我可以读传记然后说,“不,不是那家公司,不是这家公司,改这个,删掉那段。”完美。但如果你对我一无所知,那它就没多大用处。所以我的观点是:什么是好答案?有些问题不需要精确正确的答案,或者根本没有精确正确的答案。比如,“这是一张新快消品图片。为它建议 10 条广告语。”那没有正确答案。它有更好或更差的答案,但没有精确正确的答案。所以问题是:如何把问题转化为生成式 AI 问题?想想过去 10–15 年的机器学习,有一堆东西显然是机器学习问题。也有一大堆公司说,“我们意识到可以把它变成模式识别。机器学习就是模式识别。所以我们意识到可以把它变成模式识别,方法如下。”或者,“我们意识到可以把它变成图像识别。它看起来不像图像识别,但我们找到了把它变成图像识别的方法,然后你就可以用机器学习模型来自动化它。”所以这里有一个大问题:那些明显是 LLM 用例的东西,那些看起来真的不是 LLM 用例的东西——比如法律,或者也许这么说不对——比如它看起来应该是,但很难,因为你需要做对。然后还有很多关于如何管理所有这些问题的疑问。
But that may or may not be what you want. It depends on the use case, who you are, and why you asked. A trivial example: a couple of years ago, when this had just started working, I went to speak at an event and they asked for a long biography, and I didn't have one. I made a long biography in ChatGPT, and it was full of mistakes. But the thing that occurred to me is: if I'd used ChatGPT to do that, I could have fixed it in 30 seconds, because I'm an expert on myself. I can read the biography and say, "No, not that company, not this company, change this, delete that paragraph." Great. If you don't know anything about me, it wouldn't be very useful. So this is my point: what is a good answer? There are some questions that don't need a precisely correct answer, or where there is no precisely correct answer. Like, "Here's a picture of a new CPG product. Suggest 10 advertising slogans for it." That doesn't have a right answer. It has better and worse answers, but no precisely correct answer. So the question is: how do you turn things into generative AI questions? If you think about the last 10–15 years of machine learning, there was a bunch of stuff that was obviously a machine learning question. There were also a whole bunch of companies where they said, "We realized we could turn this into pattern recognition. Machine learning is pattern recognition. So we realized we could turn this into pattern recognition, and this is how." Or, "We realized we could turn it into image recognition. It didn't look like image recognition, but we worked out a way to turn it into image recognition, and then you can automate it with machine learning models." So there's this whole question of: the stuff that's obviously an LLM use case, the stuff that looks like it's really not an LLM use case—like law, or maybe that's the wrong way of putting it—like it looks like it should be, except that it's really hard because you need to get it right. And then there's a lot of questions about how you manage all those problems.
对。我能问你这个问题吗?当人们开始谈论工作替代时,我总是这样和他们争论。我说,问题是,律师工作中需要大量手动操作、重复性任务的部分被 AI 取代了。但作为律师的创造性方面——比如写出正确的条款,知道如何作为律师为客户辩护——那部分仍然存在,因为 AI 不会给你最好的答案。它给你一个看起来像正确答案的欺骗性答案,但不是正确答案。我这样说对吗?
Right. Can I ask you this question? When people start talking about job displacement, I always argue with them in the following manner. I say, look, the problem is that the lawyer's job that requires a lot of manual work, repeatable tasks, is taken away by AI. But the creative aspects of being a lawyer—as in putting the correct clause, knowing how to defend a client as a solicitor—that part still stays in the system, because AI is not going to give you the best answer. It's giving you an answer that looks deceptively like the real answer, but not the correct answer. Am I right to put it this way?
这里有两件事。首先,我认为有另一种讨论方式。我前不久刚写了点东西,那就是这很难预测。如果你试图回测……如果你回顾 20 世纪计算的目的,一半是为了自动化会计。所以我们有加法机、打孔卡、大型机、数据库、数据处理、电子表格、ERP。这一切都是为了自动化会计和簿记。你去看看美国人口普查数据,被称为会计师或簿记员的人数在整个 20 世纪基本呈直线上升。所以这就是为什么人们会去查维基百科上关于德国税收悖论的条目,这基本上是价格弹性。如果你让做某事更便宜,你是用更少的钱做同样的工作,还是用同样的钱做更多的工作,或者用更多的钱做更多的工作?你得到了新的投资回报率。但如果你停下来想一想,实际情况并非如此。实际发生的是,今天的会计师不做 50 年前他们会做的工作,但做了更多的工作。
So two things here. I think firstly, there's a different way of talking about this. I just wrote something about this a while ago, which is that it's very hard to predict this. And if you try to back-test it... So if you go back and look at the purpose of computing in the 20th century, half of it was to automate accounting. So we have adding machines, punch cards, mainframes, databases, data processing, spreadsheets, ERPs. It's all about trying to automate accounting and bookkeeping. You go look at the US census data, and the number of people who are called an accountant or a bookkeeper basically goes up in a straight line for the whole 20th century. So this is where people go and look up the Wikipedia entry for the German tax paradox, which is basically price elasticity. If you make it cheaper to do something, do you do the same work for less money, or do more work for the same money, or maybe more work for more money? You get a new ROI. But if you stop and think about it for a minute, that's not what happened. What actually happened is that an accountant today doesn't do the work they would have done 50 years ago, but way more of it.
他们正在做一大堆过去不可能做的事情,因为那时根本做不到。
They're doing a whole bunch of other stuff that they wouldn't have done then because it would have been impossible.
这让他们能做更多工作。
Makes them able to do more work.
而且它还催生了不同类型、不同性质的工作。坦白说,如果你 40 年前是个投行 banker,想做 DCF 模型,那得花一整周。现在你敲几下键盘,就能得到一个新的 DCF,只需一秒钟。
And but also it enables different work, different characters of work. Frankly, you know, if you were an investment banker 40 years ago and you wanted to do a DCF, well, that would take all week. And now you can type in a different whack and you've got a new DCF. It takes you a second.
现在更多是核心工作。
More core work now.
对,没错。我的意思是,现在会计师做的事情和 50 年前完全不同。
Yeah. Exactly. So my point is that the stuff that the accountants are doing now is not what they were doing 50 years ago.
嗯。
Mhm.
所以,能够自动化他们 50 年前的工作,反而解锁了各种其他可能性。反过来的例子最明显的是报纸行业。互联网并没有真正改变记者的工作本质,但记者的薪水却来自一个垄断本地广告的轻工业和运输业务。
And so the fact that you could automate what they were doing 50 years ago unlocked all sorts of other stuff. The inverse of this is what happened to most obviously newspapers. Where the internet didn't really change what it was to be a journalist, but journalists were being paid through a light manufacturing and trucking business that had a monopoly on local advertising.
但他们那是……
But they're that...
所以如果你说本地广告要消失了,对吧?
And so if you'd said local advertising is going away, right?
对,所以如果你坐下来审视文字编辑的工作,问互联网是否改变了它,答案是否定的,跟互联网毫无关系。但报社的薪水却是由完全不同的东西支付的。
Yeah, so if you'd sat and done looked at the job of a copy editor and said does the internet change the job of a copy editor? The answer is no, zero exposure to internet. Except that the newspaper the salary was being paid by something completely different.
所以,任何单一分析都无法判断某个工作是否暴露于风险。Uber 也一样。如果你问哪些工作暴露于智能手机,没人讨论。大家都在谈 GPS,没人谈打车税。所以从某种意义上说,这就像问为什么不直接买那些上涨的股票?我们无法预测这些事。我们无法完美预知未来。不过我认为还有第二个答案,更简单,比如杰文斯悖论和劳动总量谬误。
And so that wouldn't be any one analysis of is this job exposed here. Same thing with Uber. Like, if you'd been to say what jobs were exposed to smartphones, no one was talking. Everyone was talking about GPS, no one was talking about tap taxes. So in a sense what I'm saying here is like it's like if why don't we just buy any buy the stocks that go up? We should like you can't predict this stuff. You don't have perfect knowledge of the future. I think there's a second answer here though I think would be like yes, obvious you know, it's even a simpler like Jevons paradox lump of labor fallacy.
嗯。
Mhm.
劳动总量谬误在于,当你自动化掉一些工作时,你总能看到那些即将消失的工作,因为它们就在眼前,但你不知道新工作会是什么。
Lump of labor fallacy is that when you automate away work, you can always see the jobs that are going to go away because they're right there, and you don't know what the new jobs are going to be.
对。
Yeah.
但自动化掉那些工作,反而释放了新的经济需求,新的经济模式允许消费新事物,而人类的需求是无限的。现在有多少人靠做播客谋生?想象一下 10 年前预测这个。
But the fact that you've automated away that work has unlocked new economic demand, new economics that allow the consumption of new things, and human demand, human needs are infinite. How many people are earning a living from making podcasts now? Imagine predicting that 10 years ago.
嗯……
Um...
所以你不知道新工作会是什么,但你知道会有新工作,所以是的,过程会很痛苦,有些工作会消失。但过去 200 年一直如此。比如 200 年前,我们大多数人都是农民。
So you don't know what the new jobs could be, but you know there will be new jobs, and so yes, it will be painful and there will be jobs that go away. But that's what's happened continuously over the last 200 years. Like 200 years ago most of us were peasants.
嗯。
Mhm.
我们担心庄稼歉收。90% 的人都是农民。我们花了过去 200、250 年自动化掉那些工作,但工作岗位数量依然如故。所以你需要一个理论来解释为什么这次不同,而到了那一步,就只剩下很多空谈了,因为你可以说这次发生得更快。但真的更快吗?三年过去了,什么都没发生。
And we worried that the crops would fail. And 90% of us were peasants. And we spent the last 200, 250 years automating that away, and yet there are just as many jobs. So you kind of need to have a theory for why this is different, and at that point there's just kind of a lot of hand-waving because you can say well this is happening way quicker. I mean is it? It's worth three years in and nothing's happened yet.
但你看到招聘减少了。哦,不,如果你看就业数据,那并不成立,对吧?对,招聘。不,不,不,我同意你的观点。
But you see less hiring. Oh, but no, if you look at the job state days it's not true, right? Yeah. Hiring. No, no, no. I agree with you on this one.
最多也就是经济学家之间争论不休,到底发生了什么。我不认为经济学家们普遍认为我们明显看到了招聘下降。我想回到你之前说的一个更广泛的就业问题,你一开始提到的:实际的工作是什么?
At best you'll get a lot of argument about it amongst economists as to what's really going. I don't think there's any sense amongst economists that like this we're clearly seeing a decline in hiring yet. I think the broader kind of job question here is I want to come back to something else you said. Where you started, which is what is the actual job?
对。
Yeah.
你实际雇佣他们是为了什么?你为什么雇律师事务所?是为了让他们起草合同吗?有时答案是肯定的。有时答案是,我雇他们做一份标准合同,内容和所有合同一样。但不知怎的,当你真正雇律师时,总有一个理由说明为什么不是那样。
What are you actually hiring them for? Why did you hire the law firm? Did you hire them to make the contract? And sometimes the answer is yes. Sometimes the answer is I hire them to make a plain vanilla contract that does what all the contracts will say. And somehow or other when you actually hire a lawyer there's always a reason why it's not that.
但也许那就是你想要的。比如有时候你确实只想要那个东西。我想到的类比是亚马逊和零售商。因为亚马逊能给你那个东西。先抛开那些亚马逊上没有的东西,比如奢侈品。但为了论证,想想书是最好的例子。如果你确切知道想要哪本书,亚马逊能给你。但你怎么知道你想要那本书?
But maybe that's what you want. Like sometimes now you really just want the thing. And the analogy I was thinking about here is to think about Amazon and retailers. Because Amazon can get you the scoop. I mean set aside that there's a bunch of stuff that isn't on Amazon like luxury goods. But for the sake of argument Amazon can, if you think about like books would be the best example here. If you know exactly what book you want, Amazon can get you that book. But how do you know you want that book?
你得去搜索,对吧?
You have to search for it, right?
而当你去书店时,你通常会带着你原本不知道存在的书走出来。
And when you go to a bookshop generally you walk out with books you didn't know existed.
对。我仍然去书店。是的,我经常那样做。
Yeah. I still go to bookshops. Yes, I do that all the time.
这里的问题是,零售商的目的就是成为物流链最高效的终点吗?如果是这样,亚马逊可能更高效,但杂货店不一定。去街角商店或公寓楼下的店买牛奶,可能比从亚马逊订购更高效。
The question here is, is the purpose of the retailer to be the most efficient end point to a logistics chain? In which case Amazon is probably more efficient, not necessarily for groceries say. It's probably more efficient to go to the shop on the corner or the shop in the ground floor of your apartment building and buy milk than to order it from Amazon.
对。
Yeah.
但如果你不知道那个东西存在,那商店在做什么别的事?买书是一种休闲活动。它关乎服务、体验、建议、策展和推荐。亚马逊有 8 到 9 亿个 SKU。你不能去亚马逊说,“我想买点好看的东西装饰房子。”那不是一种顺序查询。
But if you don't know that that thing exists, what is it that the shop is doing something else? Book buying is a leisure activity. It's about service and experience and suggestion and curation and recommendation. Amazon has 800 or 900 million SKUs. You can't go to Amazon and say, 'I kind of feel like buying something nice for my house.' That's not a sequel query.
所以书店的工作变了。它变成了一个发现引擎。
So the bookstore has the job has changed. It has become a discovery engine.
关键是工作被分离了。
The point is the job is separated.
是的。
Yes.
功能只是以最有效的方式给我那个东西,还是别的什么?亚马逊能给你那个东西,但做不了别的。而大语言模型能给你那个东西,但它能做别的吗?另一种思考方式是,大语言模型本质上做的就是给你平均值。
Is the function just to get me the thing in the most efficient way possible or is it something else? And Amazon can get you the thing, but it can't really do the something else. And the LLM can make you the thing, but can it do the something else? Another way to think about this is that what LLMs do absolutely inherently is they give you the average.
对。
Yeah.
它们告诉你,这大概是大多数人会说的。那是你想要的吗?
They tell you this is what most people would probably say. Is that what you want?
不。如果你想……
No. If you want to...
嗯,也许要看情况,对吧?
Well, maybe it depends, it depends, right?
你去找律师,说我想给我的公司弄一份保密协议。
You go to your lawyer, I want an NDA for my company.
好,你是想要一份跟别人不一样的 NDA,还是就跟大家用一样的 NDA?
Okay, do you want a different NDA from what everybody else has or do you kind of just want the same NDA as everybody else?
大概就是一样的 NDA,对吧?
It's probably the same NDA, right?
你可能想让律师检查一下,但你不必付钱让律师写一份跟别人一模一样的文件。你现在可能就不该为这个花钱,但有了 AI 你肯定不会再这么做了。但问题是,你从他们那里得到的就是这个吗?还是得到了别的东西?
You maybe want a lawyer to check it, but you don't need to pay but you probably won't be paying the lawyer to write a document that's exactly the same as the document that everybody else uses. You probably shouldn't be now for that but certainly with AI you won't be doing that. But then the question is, is that what you were getting from them? Were you getting something else?
嗯。
Mhm.
而如果你想要——这就是我的观点——你想要的是平均值吗?你想要的是大家可能都会做的、都会说的那个平均数?答案往往是“是”。比如,怎么做意大利烩饭?米饭要煮多久?我不是在找什么原创答案,我就是要那个答案。
And if you want — and this is my point — do you want the average? Do you want the mean of what everybody would probably do? What everybody would probably say. And the answer very often is yes. Like, how do I make a risotto? How long should I cook the rice for? I'm not looking for an original answer here. I'm looking for the answer.
对。
Yeah.
但你可以把这个延伸一下。这是我冰箱的照片,晚饭该做什么?现在就能用。试试看,效果会很好,比你想的还好。
But then you can extend that. Here's a picture of my fridge. What should I cook for supper? That will work now. Try it. And it will work pretty well, better than you would expect.
是的,我相信你。
Yes. I believe you.
它很可能会给你一些你从没想过的、激进古怪的新点子。它会说:“嗯,我看到一些菠菜,一些瑞可塔奶酪。”这就是这件事的发展方向。
It may well give you some radical weird new thing you'd never thought of. It would say, 'Hmm, well, I can see some spinach. I can see some ricotta.' This is where this matter is going.
但你要的是平均值吗?这就是关键。你为什么去找律师?为什么去找贝恩、波士顿咨询或麦肯锡?就只是要一份演示文稿吗?这里有两部分。领英上全是些蠢货,觉得“克劳德给我做了这份演示文稿”。狭隘的批评是这份演示文稿很糟糕。如果麦肯锡给你这个,你会把他们赶出去。
But do you want the average? And that's the point. Why do you go to the lawyer? Why do you go to Bain or BCG or McKinsey? Is it just you want the deck? There's two parts to this. LinkedIn is full of idiots who think that Claude made me this deck. And the narrow criticism is the deck is terrible. If McKinsey gave you that, you'd throw them out.
但麦肯锡也在用这个,对吧?
But McKinsey is using that too, right?
是的。因为你要区分那些会被修复和改进的事情,而那些不同的事情是另一个问题。但这并不是你雇佣他们的原因。有时是。也许你在做普通的私募股权尽职调查,只想要一份演示文稿,说明市场结构和竞争对手。那么,是的,克劳德可以帮你做。或者至少,它可以快得多地完成,你也不用付那么多钱给他们,因为他们已经设置好了。但也许你付钱给他们不是为了这个。
Yes. Because you have to distinguish between the things that will get fixed and improve, and the things that are different are a different question. But that's not why you hired them. Sometimes it is. Maybe if you're doing a plain vanilla private equity due diligence and you just want a deck that says this is a market structure and these are the competitors. Well, then yes, Claude can do that for you. Or at any rate, it can do it vastly quicker, and you won't pay them as much for them to do it for you because they've got it set up right. But maybe that's not what you were paying them for.
我前几天跟一个朋友聊天,他在四大会计师事务所之一工作。他说客户的首席信息官想把 Oracle 从一个版本迁移到另一个版本,并付钱给四大中的一家做研究,问是否该这么做,答案是“是”。CEO 说这要花很多钱,这真的是最好的花钱方式吗?CEO 又雇了另一家四大来做第二意见,因为第一家说“是,你应该这么做,而且你应该雇佣我们 300 人来执行这个项目,这是一个为期两年的项目”。CEO 觉得有意思。所以我打算雇另一家公司,前提是他们不会拿到转换工作,我想让他们给我一个意见,我们到底该不该做。答案是“是,也许如果你把正常折旧年限翻倍,而且你要意识到:第一,那家公司的合伙人是 CIO 的好朋友;第二,CIO 真的很想要一个大项目;第三,CIO 快退休了,没什么政治资本。”所以 CIO 真正想要的是一份文件,告诉董事会他和董事会已经知道的事情。
I was talking to a friend the other day who was at one of the big four accounting firms. He said the client CIO wants to migrate from one version of Oracle to another and has paid one of the big four to do a study to say should we do this, and the answer was yes. The CEO says this is a lot of money, is this really the best way to spend it? CEO hires another of the big four for a second opinion because what the first firm has done is said yes you should do this and you should hire 300 of us to come in and do the project, it's a two-year project. CEO says interesting. So I'm going to hire another firm on the basis that they won't get the work to do the conversion and I would like them to give me an opinion on whether we should actually do this. And the answer is yes maybe if you double the normal depreciation life and also you do realize that number one the partner at that firm is best friends with the CIO. Number two the CIO really wants to have a big shiny project. Number three the CIO is about to retire and doesn't have much political capital. So what the CIO actually wants is a document to tell the board what he and the board already knows.
从来都不是关于问题本身。
It's never about the problem.
这算是一个极端案例。但你雇他们是为了什么?有时候你只是雇他们来获取数据,获取代码,获取内幕。你只是想要大家可能都会做的东西。我是说,极端情况下,你知道,那些对 AI 最不满的人。嗯,就像用 AI 写言情小说。
Now this is kind of an extreme case. But what is it that you're hiring them for? Sometimes you're just hiring them to get the data. You're just hiring them to get the code. You're just hiring them to get the scoop. You just want what everybody would probably make. I mean at the extreme, you know, the people who are most upset about AI. Well, it's like using AI to write romance fiction.
对。
Yeah.
因为它遵循一些非常直接、广为人知的套路。
Because it follows some very straightforward, well-understood conventions.
或者跟着一些好作家学写作。
Or follow some good writers for more writing.
对。或者就拿写作来说。挑战在于,你想要的是不是非平均值的东西?因为这是一个普遍有趣的理论问题。大语言模型或 AI 可以让你更前卫摇滚、更朋克、更新浪潮,或者更像不太好的涅槃乐队。但它不会知道大家都厌倦了迪斯科和前卫摇滚,经济很糟糕,我们担心这个那个,所以我们想要一种完全不同的声音,而朋克就能做到。也许那不是人们想要朋克的原因。但它不会知道人们想要不同的东西。或者即使知道,那也会是一个非常不同的幅度。但很容易看出,你可以用 AI 制造更多我们现在已有的东西。让它知道我们想要不同的东西以及具体是什么,这是一个不同的问题。它怎么会知道呢?因为默认情况下,那会是训练数据之外的东西。
Yeah. Or writing for example. The challenge is, do you want something that isn't the average? Because this is an interesting theoretical question in general. An LLM or AI can make you more prog rock or more punk or more new wave or more stuff that sounds like not very good Nirvana. But it won't know that everyone's really bored of disco and prog rock and the economy's terrible and we're worried about this and that, and so we want a completely different kind of sound and punk will do that. Maybe that's not why people wanted punk. But it won't know that people want something different. Or if it did, that would be a very different amplitude. But it's easy to see that you can use AI to make more of the stuff we've got now. It's a different problem for it to know that we would want something different and what it is we would want. How will it know? Because by default that would be something outside the training data.
是的。
Yes.
所以,我们绕了一大圈回到你的问题,但有多少次你雇人是因为你想要训练数据的平均值,又有多少次你想要训练数据之外的东西?你想要的是不是正常建议的东西。
And so, it's a long way we're kind of circling around your question, but how often is it that you hire something because you want the average of the training data, and how often is it that you want something that's out of the training data? You want something that's not what would be the normal suggestion.
没错。如果你是对冲基金经理,你在寻找阿尔法。所以你不会找大家都同意的东西,你会找大家都不一致的东西。
Correct. If you are a hedge fund manager, you'll be looking for the alpha. So you won't be looking for what everybody agrees on. You'll be looking for what everybody disagrees on.
对,但你不能直接去 ChatGPT 说“给我一些特别蠢的投资建议”,对吧?
Yeah, but you can't just go to ChatGPT and say give me some really stupid investment recommendation, right?
但你可以,不过——
But then you can, but like —
但我想你说的是,大语言模型解决了很多大多数人都同意是正确做法的事情。只是我们不知道新任务是什么。
But I think what you are saying is that the LLM is solving a lot of tasks that most people could agree is the right way to do it. Except we do not know what the new tasks are.
嗯,这又回到了自动化的问题,对吧?
Well, this comes back to automation, doesn't it?
是的。
Yes.
所以,你知道,从 19 世纪开始,自动化的核心就是拿一个任务,一遍又一遍地用完全相同的方式去做。
So, you know, the point of automation right back to the 19th century is take this one task and do it exactly the same way over and over and over again.
嗯。
Yeah.
而现在有了 AI,你可以把一堆以前需要人做的任务拿过来,说“做同样的任务”——可能不是完全一样的方式,因为它是概率系统——但还是一遍又一遍地做同样的事。
And so now with AI you can take a bunch of tasks that previously you needed people for and say do the same task maybe not exactly the same way cuz it's probabilistic system do the same thing over and over again.
而且用不同的方式重新设计,让它更高效。
And redesigned it in a different way so that it's efficient.
对,就像麦当劳故意把汉堡做得看起来不规则一样。
Yes, you well it's like the thing of like McDonald's making their hamburgers look artificially irregular.
没错。
That's right.
你想要那样吗?你想要每次都用完全一样的方式做吗?答案是有时候是的。但也不一定——那不一定是你找别人的原因。
Do you want that? Do you want it done exactly the same way every single time? And the answer is sometimes yes. But not necessarily that's not necessarily why you're going to somebody.
嗯,但我理解你的意思,对吧?那么我认为,不管有没有 AI 工具,创作者仍然会有工作,因为人类总是需要创造新事物、构建新事物、想出新的角度。我举个例子:有人说围棋已死,因为 AI 会下,但我们看到很多围棋大师现在在下以前被认为是禁忌的招法。
Yeah, but you I take your point to this, right? Then I think creators will still have a job whether we have AI tools or not because there's always a need human need to create new things, building new things, and coming up with angles that I mean I'll use the example of people say Go is dead because AI plays it, but people we've been seeing a lot of Go masters now playing things that previously were taboos.
嗯,我觉得这挺公平的——
Yeah, I mean I think that's a fair
但就像我说的,我并没有输给 H 案例。
But I'm well I'm well I'm I'm not losing to the H case as I was saying.
这个类比不好,因为就像说“没人会再跑步了,因为我们有车”。
That's a bad analogy cuz it's like saying you know, it's like saying no one will go running anymore cuz we have cars.
对,当然不是,对吧?人们还是会跑步,对吧?但我想说的是,这是否意味着实际上什么都没变,只是我们想法的重新配置?
Right, of course not, right? People still go for a run, right? But that's all all I'm saying is does that mean that actually nothing really changes but it's only the reconfiguration of what we think
所以我要反驳我刚才说的一切。有一种观点认为,过去 200 年我们实际上一直在自动化越来越高级的人类大脑功能。人类功能。你从把人当作役畜来自动化开始——比如拉、拖、搬东西。所以你先自动化腿。
So I would argue against everything I've just said. So there's a there's a view here that says that what we've actually been doing for the last 200 years is automating higher and higher level human brain functions. Human functions. So you start by automating human beings as beasts of burden. So like human beings like pulling, you know, hauling things and carrying things. And then you so you automate legs first.
好吧。
All right.
然后你自动化手臂和手指。
And then you automate arms and you automate fingers.
现在你开始走向大脑。
And now you're going towards the brain.
现在你要去大脑了。没错。然后你会到达顶端,什么都不剩了。这个框架很简洁,但我不确定它很有说服力。就像,你怎么知道那就是它呢?
And now you're going to go to the brain. Exactly. And so you'll reach the top and there'll be nothing left. Now, it's a neat framing. I'm not sure it's very convincing. It's like, well, how would you know that that's what it is that
而且机器人目前还不太好用。
And robots don't work very well yet.
嗯,因为你怎么知道它正在自动化所有人们想要和需要的东西呢?
Yeah, I mean, because how is it that you know that the things that were there that it's automating all the things that people want and all the things that people need?
嗯。我想这就是根本问题。
Mhm. And that is the underlying question, I suppose.
是的,而且这有点无法回答,因为和所有其他平台转变不同,我们不懂这背后的科学。
It is and it's a sort of unanswerable thing because unlike every other platform shift, we don't understand the science of this.
嗯。
Mhm.
所以,我说,你不知道会发生什么,但在深层意义上你知道。你不知道互联网会如何演变,但你知道 PC 要两三千美元一台,全世界只有一亿台,而且明年不可能有 50 亿人拥有 PC。
So, you know, I said, you know, you didn't know what was going to happen, but like at a deep level you did. You didn't know how the internet was going to evolve, but you knew that PCs cost two or three thousand dollars and there's only a hundred million of them in the world and they weren't going to be five billion people with a PC next year.
嗯。
Mhm.
而且电信公司不会在 1998 年底前给全世界所有人铺光纤。你大概知道一些基本约束。移动设备也一样。你不知道 iPhone 4 会是什么样,但你知道它不会有视网膜植入。
And that Telcos weren't going to give everybody in the world fiber in by the end of 1998. Like, you kind of knew some basic constraints of what could happen. Same with mobile. You didn't know what the iPhone 4 would look like, but you knew it wouldn't have a retinal implant.
嗯。
Yeah.
而且你知道,所以你知道它大致能走到哪里的物理约束。你大概知道成本。你大概知道每年会卖出多少台。而对于大语言模型,因为我们没有很好的理论描述来解释它们为什么这么好用,这就是 Scaling(规模扩张)的问题。我们从经验上知道它们已经 Scaling(规模扩张)了。但我们没有理论解释——一个好的理论解释——来说明这种 Scaling(规模扩张)是否会继续,以及会持续多久。
And you know, you know, so so you knew the basic physical constraints of where this could go. Um you knew what it roughly what it cost. You knew roughly how many people would be how many were getting sold every year and so on. Whereas with with LLMs, because we don't have a good theoretical description of why they work so well, this is the scaling thing. We know empirically that they have scaled. We don't have a theoretical explanation, a good like theoretical explanation for for whether that continues and what and how long.
嗯。
Mhm.
所以我们无法预测我们能用这些东西做什么。成本也一样。我的意思是,也许明天就有一篇论文说,你可以用 5% 的算力得到相同的结果。
Um so, we can't predict what it is that we'll be be able able to do with this stuff. The same thing about cost. I mean, you know, maybe that, you know, we have a paper tomorrow that says you can get the same results for 5% of the compute.
嗯。
Mhm.
我的意思是,也许不太可能,但我们没有办法确定地说出来。
I mean, I you know, maybe it's unlikely, but like we we don't have a way to state that definitively.
嗯。理论物理学家回答你那个关于大语言模型为什么这么好用的问题的方式是:恰好你放入十亿个 token,你就得到了所有东西的最佳梯度下降。这大概就是我看待它的方式。但我还有最后一个问题,对吧?嗯,听你说话,以及回顾你过去 10 年的所有作品,我一直觉得问你预测是错误的问题。这次我有个不同的问题。哪些指标能向你表明 AI 实际上已经“吃掉”了世界,或者还在“吃掉”世界——也就是说,我们如何思考 S 曲线?是我们必须看到一些尚未出现的非常不可预测的东西,还是这些东西总是渐进式变化的?
Mhm. The the the theoretical physicist's way of answering that question you say about why LLMs work so well is just happens that you put a billion tokens, you get the best gradient descent of everything. That that's that's probably the way I would look at it. But I'll just have one final question, right? Uh I think from listening to you and going through all your works over the last 10 years. I always I think asking you about predictions is the wrong question. I have a different question this time round. What are the indicators that would indicate to you where AI has actually eaten the world or still eating the world as in which part of the how do we think about the S curve? Is it have we have do we have to see something very unpredictable that hasn't shown up yet or is it that these things are always like incrementally changing?
嗯,所以我认为你可以让一群人同意或不同意:我们是在能力上有了阶跃变化,还是在原则上有了持续改进。我们是有了某种激进的加速,还是只是让编程变得好用了?
Well, so I mean I think you could you could get a bunch of people agreeing, disagreeing on whether we've had a step change in capabilities or whether we've got sort of continuous improvement in kind of the same thing in principle. Um, have we had some radical acceleration or is it just that we got coding to work?
是的。
Yes.
嗯,我认为可以肯定的是,那些每隔几个月才测试一次、从年初就没再看过的人,并不真正理解我们现在拥有的东西。嗯,但我不知道。嗯,做出这种量级的陈述很难。就像说“移动设备比互联网更重要吗?”嗯,
Um, I think it's certainly the case that people who only test this every couple of months and haven't looked at it since the beginning of the year do not really get what we have understand what we have now. Um, but I don't know. Um, it it is it's tough to make these kind of statements of magnitude. It's like saying was mobile a bigger deal than the internet? Well,
那是个无意义的问题,对吧?
That's a meaningless question, right?
我真的不知道该怎么回答那个问题。
How I don't I don't really know how one would answer that question.
嗯。但你的观点是,你只需要看看你现在做的事,也许 6 个月后再看看,看任务是否有戏剧性的改进。也许那是一个很好的指标,可以判断这个周期是否已经结束。
Mhm. But you your your point of view is that you're just going to look at maybe what you do now and maybe 6 months down the line and see if there's really a dramatic improvement towards the task. Maybe that's a good indicator just to see whether this cycle has played out all.
嗯,我在想怎么表达。你知道,我非常认同这个规律。我给你的类比是:想象你是一个会计,在 70 年代末看到第一个电子表格。那是改变人生的。它把一周的工作在一小时内完成。
Yeah, so I'm trying to think how to put this. You know, I'm very much the law you know, so I I've the analogy I I gave you like you know, imagine you're an accountant seeing the first spreadsheets in the late 70s. It's life-changing. Like you it doesn't it does it does a week of work in an hour.
嗯。
Yeah.
这就是现在一个软件开发者看代码的感觉。太震撼了。
And that's how what it's like to be a software developer looking at code code now. It's mind-blowing.
它完全改变了做这份工作的意义。
It completely changes what it is to do the job.
然而,现在想象你是 1978 年的一位律师,看着一张电子表格。好吧,这很聪明,我的会计应该看看这个,但这不是我的工作。同时,文字处理器也存在。
However, now imagine you're a lawyer looking at a spreadsheet in 1978. Okay, this is very clever. My accountant should see this, but that's not what I do. Now, word processors exist at the same time.
是的。
Yes.
但如果你是一位律师,看着电子表格,你会想:“嗯,这很聪明,我下周可能用它来做我的考勤表。”我的意思是,在 1978 年,你需要一台价值 1.5 万美元的电脑才能运行电子表格,但实际上,买一台 Apple II,配齐足够的内存和磁盘驱动器之类的东西,大概要花 1 万到 1.5 万美元。但现在想象你把这个放在一边。你是一位律师,看着这个觉得很好,我下周可以用它来做考勤表,但我不会每天都用,那不是我的日常工作。现在的情况和这个有点像。所以,有些人是会计,看着电子表格,大多数人都是周活跃用户和月活跃用户,而不是日活跃用户。即使你看 13 到 19 岁的年轻人,那里也是周活跃或月活跃用户多于日活跃用户。所以,大多数人仍然是看着电子表格的律师,他们说:“这很聪明,也许下周很有用。”我不确定解决方案是更好的模型。我认为解决方案是,你必须把它包装在不同的产品里。你必须思考如何创建一个使用场景,让它被设置好。现在,我意识到你有这样一件事,你每天都没注意到自己在做,而我做了一个东西来解决它。
But if you were a lawyer looking at a spreadsheet, you think, "Well, this is very clever. I might use it next week to do my timesheet." I mean, it says decide that you needed like 15 grand worth of computer to run a spreadsheet in 1978, but I mean, literally, it was like to buy an Apple II with all the enough memory and the disk drives and stuff with like 10 or 15,000 dollars. But now imagine you set that aside. You're a lawyer looking at this like that's great. I could use that to do my timesheet next week, but I'm not going to use it every day. That's not what I do all day. And it's kind of the same now with with with with with that at arms. So, some people who are the accountant looking at spreadsheet most people are weekly active users and monthly active users and not daily active users. Even like you look right in at like 13 to 19-year-olds, even there more people are daily or weekly or monthly active users than daily active users. And so, most people are still the lawyer looking at spreadsheet and they're saying This is very clever. It's quite useful maybe next week. And I'm not sure that the solution to that is a better model. I think the solution to that is now you have to like wrap that in different products. And you have to think about how you create a use case where it's set up. Now, I've realized that you have this thing that you didn't notice you were doing every day and I've made a thing that solves it for you.
是的。而且因为聊天界面并不是我做所有事情的界面。
Yeah. And because the chat interface is not that I do interface for everything.
确实不是。而且让它工作起来仍然很繁琐。你知道,举个例子,我经常旅行。
It isn't, no. And making it work is still fiddly. I mean, you know, take an example. So, I I I travel a lot.
嗯。
Mhm.
所以我需要整理我的费用。那么,正确的方法是什么?嗯,
So, I need to put my expenses together. So, what's the right way of doing that? Well,
我可以拿邮件,查看航班、出租车和酒店,然后把它们输入到电子表格里。选项一。
I can take the email, I can look at the flight and the taxi and the hotel and type and type them into a spreadsheet. Option one.
嗯。
Mhm.
因为我不需要把所有收据放在一起。选项二,我可以让 Gmail 查看所有进入邮件的收据,并自动把它们放入 Google 表格。我会去设置这个吗?这不应该是 Google 的工作吗?选项三,我可以使用金融科技或公司卡,它会自动为我做这件事。所以,我可以注册 XYZ 新卡,这是他们的功能之一,他们已经构建好了。
Cuz I don't need to put all the receipts together. Option two, I could get G get Gmail could look at all those receipts as they come into email and automatically drop them into a Google Sheet. Am I going to set that up? Shouldn't Isn't that Google's job? Option three, I could use a fintech or a corporate card that will automatically do that for me. So, I could sign up to XYZ new card and that's one of their features and they've built that.
而且他们使用机器学习,或者 Google 也会使用机器学习。
And they're using machine learning or Google just as Google will be using machine learning.
他们帮你处理费用。
And they help you to do the expenses.
选项——有一件事我不会做
Option The one thing I'm not going to do
嗯。
Mhm.
就是打开 Claude Code 说:“嘿,你能登录我的银行和 Gmail,这是我的 Uber 账户,你能计算哪些是……”我的意思是,你可以这样做,但那太疯狂了。
is load up Claude code and say, "Hey, can you log in to my bank and my Gmail and here's my Uber account and can you calculate which of them are like that's I mean, you could that would be insane.
嗯。
Mhm.
那会是一个非常非常愚蠢的做法。我的意思是,你可以做,但为什么?不,不,你到底为什么要那么做?
It would just be a really, really stupid way of doing it. I mean, you could do it you could do that. But why No, no, why on earth would you do that?
好观点。这是个不错的收尾点。嗯,Benedict,非常感谢你来做客节目。嗯,一个非常简单的问题:我的观众在哪里能找到你?我是你新闻通讯的订阅者,所以我就直接推荐了。
Good point. And that's a good place to wrap. Um Benedict, many thanks for coming on the show. Um just very simple question. Where do my audience find you? I'm a subscriber of your newsletter, so I'm just going to recommend it straight.
是的,我需要检查一下你是否是付费订阅者。
Yeah, I need to check if you're a paying subscriber.
哦,好的。好的。好的。
Oh, okay. Okay. Okay.
我可以把过去 10 年的所有收据都发给你。
I can send you all the receipts over the last 10 years.
嗯,我父母的 SEO 做得很好,所以如果你谷歌 Benedict Evans,就会找到我的网站。我每年会发布两次演讲内容。我会写一些关于我在想什么的东西,然后还有一份每周新闻通讯,是我每周的笔记,记录发生了什么、什么有趣以及它的意义。
So, um my parents have good SEO, so if you Google Benedict Evans, then there's my website and I do a presentation to a publish presentation twice a year. I write stuff about what I'm thinking about and then there's a weekly newsletter which is my notes for the week of what's going on and what was interesting and what it meant.
嗯。
Mhm.
谢谢你的电子表格类比。每次我思考产品时,总会想到它。
Thank you for the spreadsheet analogy. I always think about it every time I think about product.
好的。谢谢。
Right. Thank you.
谢谢。
Thank you.