Agentic Coding: From Useful to Game-Changing
打开互动全文版(中英对照 + 朗读 + 问答)→Benedict Evans 反思智能体编程如何成为 AI 的杀手级应用,从好奇转变为重塑科技行业的产品市场契合点。
Benedict Evans reflects on how agentic coding has become the killer app for AI, shifting from curiosity to a product-market fit that's reshaping the tech industry.
Benedict,欢迎回到 Acing Z 播客。
Benedict. Welcome back to the Acing Z podcast.
谢谢。
Thank you.
上次你来的时候,我们讨论了你演讲《AI 吞噬世界》的第一个版本。你写那篇演讲差不多是一年半以前了。现在,你总是以那些大问题开始你的演讲。但这次我很好奇,在进入未来的问题之前,我想让你反思一下,自从你最初做那个演讲以来,我们学到了什么?哪些事情已经发生了?让我们回顾一下过去一年发生了什么变化。
Last time you were here, we were discussing the first iteration of your presentation AI eats the world. You wrote it almost a year and a half ago. At this point, you always begin your presentation with the big questions. But I'm curious this time, before getting into the questions going forward, I want you to reflect on what have we learned since you originally made the presentation. What's played out? Let's reflect back on what's changed in the last year.
所以我认为我们对产品策略的分化有了更清晰的认识。我们对竞争紧张局势有了更深的感受,这已经超越了单纯的“用更多算力更快地做出更大的模型”。特别是 OpenAI 的策略经历了几次迭代,从“昨天一次性搞定所有事”到“哎呀,不,也许我们应该加倍押注编程”。显然,智能体式编程开始奏效了,因此科技领域的焦点大规模地集中到了这一点上,它拥有绝对的产品市场契合度,客户几乎是抢着要。当然,随之而来的是我们目前看到的供应紧张,以及供需、产能、资本支出、定价之间的不平衡。所以这是个大转变。我们从一个“这有点用、有点令人兴奋,但不太确定要拿它做什么”的时刻,变成了“好吧,它对编程有用。它对其他东西也有用吗?是的,几乎可以肯定,但这就是现在正在奏效的东西。”所以我们的焦点变得狭窄了很多。除此之外,图表上的数字不断攀升,模型越来越大,资本支出持续增长,使用量也在增长,人们用得越来越多。但两三年之前你可能有的那些基本问题,大部分都还没有答案。我们不知道模型领域会不会有赢家。我们不知道他们能否在技术栈上层捕获价值。我们不知道模型能做什么。我们看不到消费者会每天使用这个技术,而不是每周,以目前的技术水平来看。所以所有这些问题都还是开放的。
So I think we have much more of a sense of diverging product strategy. We have much more of a sense of competitive tension that goes beyond just 'make a bigger model faster with more compute'. We've had several iterations of OpenAI's strategy in particular, from sort of 'everything all at once yesterday' to 'oops, no, maybe we should double down on coding'. Clearly, agentic coding started working, and so all the focus in tech has narrowed in massively onto that as something that has absolute product-market fit, in the sense that the customers are pulling it out of your hands. And of course, that comes with the supply crunch around capacity and price imbalance of supply-demand, capacity, capex, pricing that we see at the moment. So that's the big shift. We had a moment of 'this is kind of sort of working and kind of exciting, but we're not quite sure what we're going to do with it' to 'right, it works for coding. Will it work for anything else? Yes, almost certainly, but that's what's working right now.' So we've got this much narrower focus. Otherwise, the chartman numbers keep coming up, the models keep getting bigger, the capex keeps growing, the usage keeps growing, people using this more. But most of the fundamental questions you might have had two or three years ago didn't really have answers. We don't know if there'll be a winner in the models. We don't know if they can capture value up the stack. We don't know how much the models can do. We don't see a way that consumers will use this daily rather than weekly with the technology we have right now. So all those questions are still open.
是啊。就编程而言,我们怎么能预见到它会成为真正起飞的应用场景呢?你对此有什么反思?
Yeah. And just on the coding, how could we have foreseen that would have been the use case that really took off? What's your reflection on that?
嗯,从确定性的角度你可以说:“看,谁在摆弄这些东西?软件开发者。软件开发者会试图让什么工作起来?软件开发。”所以从一个非常简单、天真的层面来说,是的,应该首先起作用的就是软件开发。我经常把这个时刻比作 97、98 年的互联网,但也像 70 年代末 80 年代初的个人电脑。它非常令人兴奋,但还不完全清楚它的用途,而且它还没有完全奏效。显然,人们用个人电脑做的第一件事就是制造电脑。而人们用大语言模型做的第一件事,从某种意义上说大语言模型就是电脑,就是制造更多的算力。所以这并不太令人惊讶。我认为转变发生在今年年初,很明显智能体式编程从“有点用”变成了“真正改变一切”。我不确定你是否能确定性地准确预测它何时会发生,以及编程会是第一个起作用的。
Well, deterministically you could have said, 'Look, who's messing about with this stuff? Software developers. What are software developers going to try and make work? Software development.' So at a very simplistic, naive level, well, yeah, the stuff that should work first is software development. I often compare this moment to the internet in '97, '98, but it's also like the PCs in the early '80s or the late '70s. It's incredibly exciting but it's not quite clear what it's for and it doesn't quite work yet. And clearly the first thing that people did with PCs was make computers. And the first thing that people are doing with LLMs, in a sense LLMs are computers, is to make more compute. So that's not terribly surprising. I think the shift has been at the beginning of this year clearly that agentic coding went from being kind of useful to really changing everything. And I'm not sure you could have deterministically predicted exactly when that was going to happen and that it was going to be coding that would work first.
那么关于这对工程师、初级工程师、高级工程师、就业讨论、团队组织方式等意味着什么,我们学到了什么?到目前为止我们学到了什么?
And what have we learned about what this means for engineers, junior engineers, senior engineers, the jobs discussion, how teams are organized, etc.? What have we learned so far?
我认为我们什么也没学到。我的意思是,六个月前这还不管用。每个人都在手忙脚乱地试图弄清楚这意味着什么。你很容易陷入噪音和细节中,以及某人在昨天的聚会上说了什么。“哦,天哪,事情就会这样发展。”这需要几年的时间才能稳定下来。别的先不说,就定价而言,我们面临着巨大的供需矛盾,进而影响定价。所以我们不知道一个团队会是什么样子。我认为人们在问一些新的问题,围绕那个显而易见的问题:你还会雇佣初级员工吗?如果会,他们做什么?你过去为什么雇佣初级员工?你雇佣他们是为了做他们实际做的事情,还是为了做别的事情?所以如果你自动化了一类过去由人完成的工作,那么会发生什么?这在软件开发中现在变得非常现实,因为你确实在自动化一堆过去由人完成的工作。所以这些问题现在是现实问题,而不是理论问题。但我不认为有人能说他们知道市场结构会是什么样子,或者三年后软件工程师的职业会是什么样子。我认为如果你认为现在就能知道,那简直是疯了。
I don't think we've learned anything. I mean, this didn't work six months ago. And everyone is scrambling around trying to work out what it means. You can get very into the noise and the detail and what somebody said at a party yesterday. 'Oh my god, that's how it's all going to work.' It's going to take a couple of years for this all to settle down. If nothing else, because of the pricing, we've got this enormous crunch between demand and supply and hence the pricing. So we don't know what a team is going to look like. I think people are asking new questions around the obvious one: do you hire junior people? And if so, what are they doing? And why were you hiring junior people in the past? Were you actually hiring to do the thing that they did, or were you hiring them to do something else? So if you automate away a class of stuff that used to get done by people, then what will happen? That becomes much more real now in software development because you actually are automating a bunch of stuff that used to be done by people. So those questions are now rather than theoretical. But I don't think anybody can possibly say they know what the market structure is going to look like or what the career of a software engineer is going to be in three years' time. I think it would be insane to think that you could know that yet.
是啊。谈谈 OpenAI。最让你惊讶的是什么,或者你是如何理解他们的战略发展以及他们未来面临的问题的?
Yeah. Talk about OpenAI. What's most surprised you, or how have you made sense of their strategy development and the questions they have going forward?
嗯,你知道,它一直是一个如此平静、没有戏剧性的环境。
Well, you know, it's always been such a tranquil, drama-free environment.
所以,显然他们遇到了 Fiji Simo 需要休病假的问题,这让事情有点混乱。看,很明显去年下半年,最后一个季度,他们的问题是对的:模型就是模型,但还有什么?我们如何让人们用这个做其他事情?你知道,让 ChatGPT 给出 15 个关于我们如何在基础设施之上构建价值的想法,然后我们全部实施。这几乎就是实际情况。然后 Anthropic 筹集的资金较少,他们说:“不,我们要专注于编程。”他们确实让编程成功了。不管这是刻意的策略还是他们偶然发现的,让别人去说吧,但显然这奏效了。但问题仍然存在:目前有效的是软件开发和一些其他领域的事情,然后有很多人对此感到兴奋,在边缘使用它,用于某些事情。而且很明显,硅谷那些买了一堆 Mac Studio 整天运行开源模型的人,与另外 40% 说“嗯,它有点用,我上周用它做了点事”的人之间存在巨大差距。我想,你怎么弥合这个差距?我不认为这个问题已经解决了。软件是一个已经跨越了那座桥梁的地方。然后还有很多其他地方,人们有点摸不着头脑,使用它到一定程度就停了;还有很多地方,公司用它来自动化某些特定的后台流程,你不需要用户自己去想怎么用这个新工具,而是说:“好,这里有一个我们可以解决的问题。”我去和美国以外、科技行业以外的公司以及顾问和投资者交谈。他们正在逐个地看这些点解决方案。所以,几天前我和一家大宗商品公司交谈,他们想用大语言模型来更好地预测现金流,因为他们与各种小生产商打交道,不一定知道发票什么时候能付清,而且这是一个利润率很低的业务。所以这很重要,他们想用大语言模型来更好地预测现金流。这与打开 ChatGPT 或 Claude 说“嘿,给我总结一下这周的会议”完全不同。
So, obviously they've had the issue with Fiji Simo having to take a medical leave, which kind of shuffled things up a bit. Look, clearly the second half, last quarter of last year, their question was right: the models are the models, but what else? And how do we get people to do other stuff with this? You know, ask ChatGPT for 15 ideas for what we could do to build value on top of infrastructure, and then we'll do all of them. It's almost literally what it looked like. And then Anthropic, with having less capital raised, said, "No, we're going to focus on coding." And they got coding working. Whether that was a deliberate strategy or they kind of stumbled into it is for others to say, but clearly that worked. But the question still remains: the stuff that's working right now is software development and some things in some other fields, and then there's a lot of people who are kind of excited about using this around the edges and using this for some things. And there's clearly this kind of wide gap between people in the Valley who bought a cluster of Mac Studios and are running open source models all day, versus those other 40% of people who say, "Yeah, it's kind of useful, I used it last week for something." And I'm like, how do you bridge that? I don't think that question is answered. Software is a place where that bridge has really been jumped over. And then there's a lot of other places where people are kind of scratching their heads and using it up to a point, and then there's a lot of places where corporations are using it to automate some specific back office process where you're not asking the user to work out what they do with the new tool; instead, you're saying, "Okay, here's a problem that we can solve." And I go and talk to companies outside America and outside of tech, and talk to consultants and investors. They're looking at those one-at-a-time point solutions. So, I spoke a couple of days ago to a commodities company, and they want to use LLMs to get better predictions on their cash flow because they deal with all sorts of small producers and they don't necessarily know when their invoices are going to get paid, and it's a very low margin business. So that's a big deal, and so they want to use LLMs to get better cash flow forecasting. That's very different from going to ChatGPT or Claude and saying, "Hey, give me a summary of my meetings this week."
是的。你能分享一下,在每周或每日用户方面,这与移动或其他平台的早期用户采用情况相比如何?
Yeah. Can you share how this compared with mobile or other sort of platform in terms of early user adoption on a weekly or daily user basis?
所以我认为有几种不同的方式来回答这个问题。其中之一是,我们总是站在巨人的肩膀上,增长总是在复合。所以移动不需要等待互联网或移动数据这样的蜂窝网络。移动互联网不需要等待它——它需要等待蜂窝数据,但它不需要等待互联网的出现,互联网不需要等待个人电脑,个人电脑不需要等待消费电子产品和半导体,等等。所以你总是有这种加速的采用。当我的老老板 Marc Andreessen 在开发 Netscape 时,整个地球上只有几千万台个人电脑。所以不,你不可能有 9 亿周活跃用户,因为没有 9 亿台个人电脑。所以总是有这种加速。这是第一点。我认为第二点是,在任何这些转变的早期阶段,都不清楚它会如何运作,而且什么都不好用。我刚好够老能记得这个。我不知道你多大,但任何 30 多岁的人都不太记得这样一个时代:你正在工作,然后屏幕上的所有东西都冻结了,你必须爬到桌子底下拔掉电脑插头,然后祈祷你过去一个小时做的一些东西还在。那现在已经不再发生了。回到 80 年代:你买了一张声卡。嗯,那是 300 美元。你的电脑不会有声音。好吧,那是 300 美元,而且那得花一个周末才能让它工作。我的意思是,我记得试图让这些东西工作。互联网也是一样:你得有一张装有 TCP/IP 的软盘,而且它很慢,你需要做的所有事情都不存在。移动也是一样。我们大致处于那个阶段。当然,不清楚这些东西中哪些会成功。现在也是一样:浏览器会成功吗?会是这个吗?会是那个吗?这一切如何组合在一起?在非常令人兴奋的东西和少数愿意投入工作让它工作的人之间,与把它变成一个只需按一下按钮就能用的东西之间,存在差距。我认为第三点是一个更具体的观察:我们之前提到的定价紧缩,在我看来很像 2009-2010 年移动数据发生的情况,一方面人们突然收到 5000 到 10000 美元的数据账单,另一方面,如果你有无限流量套餐——这有点像美国 iPhone 的情况,AT&T 推出了带有无限流量套餐的 iPhone——然后不幸的是每个人都买了 iPhone,开始使用 3G 和观看 YouTube,整个网络就瘫痪了,因为他们没有容量来支持。有趣的是,科技界仍然有人不明白蜂窝网络有边际成本,他们必须增加更多容量,而这需要更多钱。所以网络不得不争相让成本曲线与定价系统、基础成本和感知价值对齐,他们通过有上限的套餐、公平使用和限速等方式做到了这一点。
So I think there are a bunch of different ways to answer this. One of them is that we're always standing on the shoulders of giants, and the growth is always compounding. So mobile didn't need to wait for the internet or cellular networks like mobile data. Mobile internet didn't need to wait for it—it kind of needed to wait for cellular data, but it didn't need to wait for the internet to happen, and the internet didn't need to wait for PCs, and PCs didn't need to wait for consumer electronics and semiconductors, and so on. So you've always got this accelerating adoption. And when my old boss Marc Andreessen was working on Netscape, there were like double-digit millions of PCs on the entire planet. So no, you couldn't have 900 million weekly active users because there weren't 900 million PCs. So there's always that acceleration. That's one point. I think the second point is that at the early stage of any of these shifts, it's not really clear how it's going to work and nothing works. So I'm just about old enough to remember this. I'm not sure how old you are, but anyone in their 30s doesn't really remember a time when it was completely normal that you'd be working and then everything on the screen would just freeze, and you'd have to crawl under your desk and unplug the computer, and then pray that some of what you'd done in the last hour might still be there. That just doesn't happen anymore. And go back to the 80s: you bought a sound card. Well, that's $300. You won't have sound on your computer. Okay, that's $300 and it's like that's the weekend to make that work. I mean, I remember trying to get this stuff to work. And the same thing with the internet: you've got to get a floppy disc that has TCP/IP on it, and it's slow, and none of the stuff that you need to do existed. And the same with mobile. And we're kind of at that stage. And of course it's not clear which of these things are going to work. And that's the same thing now: is a browser going to work? Is it going to be this? Is it going to be that? How's this all going to fit together? And there's a gap between what's incredibly exciting and the small number of people who are willing to put the work in to get something to work, and just turning that into a thing where you can just press a button in real hands. I think the third point here is a much more tangible observation: the pricing crunch that we've already mentioned looks to me a lot like what happened with mobile data in 2009-2010, where suddenly people got bills for $5,000-$10,000 of data on one side, and on the other hand, if you had flat-rate data—which is kind of what happened in the US with the iPhone, AT&T launched the iPhone with flat-rate data—and then unfortunately everybody buys iPhones and starts using 3G and watching YouTube, and the whole network goes down because they just don't have capacity to do that. It's funny, there are still people in tech who don't understand that cellular networks have marginal cost, and they have to add more capacity and that costs more money. And so the networks kind of had to scramble to get the cost curve aligned with the pricing system aligned with the underlying cost and aligned with perceived value, which they kind of did with capped bundles, fair use, and throttling, and so on.
但另一面正是你现在看到的:一方面,你每月付 20 美元就能拿到价值 1 万美元的 token;另一方面,你瞎搞几天就收到一张 1 万美元的账单,你会想‘这他妈是什么?’这正是你现在从这些故事里看到的,也就是 2008-2010 年发生过的事,还有 2001-2003 年 GPS 的情况。但我认为这个类比中更有趣的部分是,自那以后,移动数据流量增长了大约 1500 到 2000 倍,移动网络的总收入约 1 万亿美元,每年资本支出约 2000 亿美元,而股价 20 年来一直持平,所有酷东西都是别人建的。他们都以为自己会建所有酷东西——比如我曾在拥有一家银行牌照的电话公司工作,因为他们以为自己会做移动银行,现在这听起来完全疯了。但关键就在这里:他们建了这套惊人、极其复杂、非常昂贵的全球基础设施,使用量一直在巨大增长,它改变了我们所有人的生活,我们都为此付费,而他们却没赚到钱,因为所有价值都向上游转移了。这绝对是 LLM 的核心问题:模型能完成所有事情吗,还是必须在上面建 300 个应用?你能直接对模型说‘帮我报税’,还是需要一个以 10 种不同方式在内部使用 AI 的报税工具?如果不能,那么成为基础模型提供商意味着什么?这仅仅是按边际成本出售的商品化基础设施吗?这似乎是目前人们很难理解的概念,因为你能卖出所有能生产的 token,所以你可以按 ROI 定价。但在未来几年,我们会有大约 1-2 万亿美元的资本支出投入,模型每年效率提升 100 到 200 倍,然后还有新模型。模型会用更多 token 还是更少?但无论我们达到什么新均衡,为什么这个均衡会是模型公司拥有定价权,而所有模型都差不多,用同样的芯片做差不多的事?它们为什么会有定价权?所以这是对你问题的长篇回答,但回顾历史:芯片公司没有捕获价值,ISP 没有捕获价值,移动网络运营商没有捕获价值。Windows 和 iOS 做到了,但它们在干别的事——它们有各种杠杆可以向上游走。当然,它们有网络效应,而模型没有。所以问题是:它们最终会像基础设施层,还是像操作系统层那样捕获价值并真正决定建什么?或者它们最终——讽刺的是 Netscape,马克·安德森曾著名地说要把 Windows 变成一堆调试得很烂的设备驱动程序,然后微软强行挤入市场,但结果证明网页浏览器不是关键,因为所有价值都在别处。所以我认为这是一堆关于这一切如何尘埃落定的问题,这又回到我对你问题的所有回答。有些事情,你不知道它会如何发展。
But the other side of that is exactly what you see now: on the one hand, you're paying $20 a month and you get $10,000 worth of tokens; on the other hand, you mess about for a couple of days and you get a bill for $10,000 and you're like, 'What the hell is this?' That's exactly what you see in these stories now, which is what happened in 2008-2010, and also in 2001-2003 with GPS. But I think the other interesting part of that analogy is that since then, mobile data traffic has risen by something like 1,500 to 2,000 times, and the mobile networks collectively have revenue of about a trillion dollars and spend about $200 billion a year on capex, and the stocks have been flat for 20 years, and all the cool stuff got built by somebody else. They all thought they were going to build all the cool stuff—like I worked for a phone company that had a banking license because they thought they would do mobile banking, which now seems absolutely insane. But that's the point: they built this amazing, incredibly sophisticated, very expensive global infrastructure with enormous growth in use all the time, and it changed all of our lives, and we all pay for it, and they didn't make any money from it because all the value moved up the stack. This is absolutely the central question for LLMs: can the model do the whole thing, or do you have to have 300 apps built on top of it? Can you just go to the model and say 'do my taxes for me,' or do you need a tax thing that uses AI in 10 different ways inside it? And if not, then what is it to be a foundation model provider? Is this just commodity infrastructure sold at marginal cost? That seems to be a very difficult concept for people to grasp right now because you can sell all the tokens you can make, so you can price it at ROI. But over the next couple of years, we've got like $1-2 trillion of capex coming down the pipe, and the models get 100x to 200x more efficient every year, and then there are new models. Will the models use more tokens or fewer tokens? But wherever we get to a different equilibrium, why would that equilibrium be one where the model companies have pricing power when the models are all kind of the same, doing kind of the same thing with the same chips? Why would they have pricing power? So it's a long answer to your question, but you go back and look over time: chip companies didn't capture the value, ISPs didn't capture the value, mobile network operators didn't capture the value. Windows and iOS did, but they were doing something else—they had all these levers to go up the stack. And of course, they have network effects, which models don't have. So the question is: do they end up like the infrastructure layers, or do they end up like the operating system layers and capture value and actually get to decide what gets built? Or do they end up—the irony there is Netscape, where Mark Andreessen famously said he was going to turn Windows into a set of badly debugged device drivers, and then Microsoft kind of crowbarred their way into the market, but it turned out that web browsers weren't the point because all the value was somewhere else. So I think that's a swirling mass of questions about how this settles out, which comes back to all my answers to your question. Some of this stuff, you don't know how it's going to work.
是的。目前还不清楚它更像互联网还是软件——大量价值或更高利润率发生在应用层,还是像云那样价值似乎存在于硬件层。目前看来,英伟达在上涨,利润率更高,并捕获了大量价值,但不确定这种情况是否会持续,或者是否会出现应用——我们是否会更像互联网。你甚至如何开始预测答案?
Yeah. It's unclear whether it looks more like the internet or sort of software where a lot of the value or just better margins happen at the application layer, or sort of the cloud where it seems the value existed at the hardware layer. And right now, so far it seems like Nvidia is going up and has better margins and is capturing a lot of the value, but it's unclear if that will remain the same or if there will be sort of applications—if we'll look more like the internet. How would you even begin to predict the answer to that?
嗯,对此有两个回答。有很多关于历史如何运作的说法,我最喜欢的一句是‘历史只教会我们一件事,那就是总会有事发生’。你总是可以事后说‘当然结果是那样’,但当时通常并不明显。特别是,我记得大约 15 年前,科技界很多非常聪明的人看着 iPhone 和 Android 说‘这又是开放对封闭,Android 会碾压 iPhone’,当然这并没有发生。我可以解释为什么,但所有这些比较都有用,却没有一个具有预测性。事后看来总是很明显。有趣的是——我最近做了几个播客并发布了这个演示,有一类评论说‘Benedict,你没做好你的工作;你应该告诉我们会发生什么,你应该做出预测,而你似乎只会说我们不知道’。这有两个问题。一是确实有一些地方我确实说了‘我认为这行不通;我认为它会那样发展’——比如我不认为基础模型是产品,我不认为聊天机器人是产品,我认为价值会更上游。但另一方面是,当你处于周期的这个阶段时,有很多路径,你不知道会是哪一条。试图说‘我认为会是那条’——你可能对,但你必须意识到这有多不确定,以及有多少不同的路径可能。这就是周期这一阶段的本质:所有赌注都是开放的。我们会到达 S 曲线向上弯曲并收窄的点。曾经有一个时刻 Windows Phone 可能成功——事后看来,不,它可能不会成功——但有一个时刻,移动将如何发展并不清楚。然后有一个时刻变得清晰,对,这就是正在发生的事。现在我们进入下一个问题。
Well, two answers to this. There are all these sorts of quotes about how history works, and my favorite one is 'history teaches us nothing except that something will happen.' You can always ex post facto say 'of course it worked out like that,' but it generally wasn't obvious at the time. In particular, I remember about 15 years ago, a lot of really clever people in tech looked at the iPhone and Android and said 'this is open versus closed again, and Android is going to crush the iPhone,' which of course isn't what happened. And I can go and explain why, but all of these comparisons are useful, none of them are predictive. It's always obvious in hindsight. It's funny—I've done a couple of podcasts recently and published this presentation, and there's a class of comment that says 'Benedict, you're not doing your job; you're supposed to tell us what's going to happen, you're supposed to make predictions, and all you seem to do is say we don't know.' There are two problems with that. One is that there are a bunch of places where I actually do say 'I don't think this is going to work; I think it's going to work like that'—like I don't think foundation models are a product, I don't think a chatbot is a product, I think the value will be further up. But the other side is that when you're at this stage in the cycle, there are many paths, and you don't know which one it's going to be. To try and say 'I think it's going to be that one'—you might be right, but you have to be conscious of how uncertain this is and how many different paths it could take. That's the nature of this part of the cycle: all bets are open. We get to the point where the S-curve kind of curves up and it narrows in. There was a moment when Windows Phone might have worked—in hindsight, no, it probably wasn't going to work—but there was a moment when it wasn't clear how mobile was going to work. And there's a moment when it was clear, right, this is what's happening. Now we move on to the next question.
科技的一个特点是,当你理解某件事、知道它如何运作以及会发生什么时,你就应该转向其他事情了。你应该始终寻找那些我们不知道答案的地方。因为,你知道,我已经大约五年没更新我的苹果电子表格了,因为我们知道结果了。他们想要——我不在乎明年的 iPhone 长什么样。我不关注他们在中国的市场份额。事情已经发生了。下一个问题。
One of the characteristics of tech is that the moment you understand something and you know how it works and what's going to happen is the moment you should move on to something else. You should always be looking for the places where we don't know what the answers are because, you know, I haven't updated my Apple spreadsheet in like five years because we know what happened. They want like I don't care what the next year's iPhone looks like. I don't pay attention to their market share in China. Like it happened. Next question.
你提到一个预测,认为基础模型不是产品;你认为它会向上移动。稍微解释一下其中的推理,以及这可能是什么样子。
You mentioned the prediction that you don't think foundation models are the product; you think it'll move up. Explain the reasoning there a bit and what that could look like.
所以我认为有三四个基本要素可以摆上台面。其中之一是,目前还不清楚如何以某种可持续的差异化方式构建一个从根本上优于其他所有人的模型。似乎没有网络效应。似乎没有像 Instagram、YouTube 或 Google 搜索那样的杠杆和策略。我们在大型语言模型(LLM)中没有看到类似的东西。现在有不同的侧重点。你知道,也许这个比那个好,也许你更喜欢这个而不是那个,但除了你愿意花钱之外,模型之间似乎没有根本性的差异化或竞争差异。第二个问题是,聊天机器人本身就像一个奇怪的、有限的 v1 用户界面。有些事、有些人和有些任务它做得很好,但对于大多数其他情况,你需要一堆其他东西:你需要工具,并且需要正确设置。它需要有正确的数据,需要配置和控制,需要有正确的用户界面,人们需要坐下来思考这应该如何工作,因为通常擅长使用工具并完成需要该工具的工作的人,与擅长决定工具应该是什么样的人不是同一群人。所以,你知道,非常擅长设计印刷出版物的人不应该去创建和设计工具。那是不同的技能。而且,你知道,非常擅长提供财务建议的人不是设计 TurboTax 的合适人选。那些是拥有不同技能的不同的人。所以,你在这个中间地带摸索。于是现在有了 Claude 做这个,Claude 做那个,还有技能等等。对我来说,这有点像,一个问题是谁来构建技能?另一个问题是,这有点像你在 Excel 中点击“文件-新建”得到的东西:这些是模板,它们能带你走一段路,但到了某个点,人们就会超越模板。我的演示文稿中有一张幻灯片,引用了几年前有人在 Twitter 上对我说的话:他们说他们是一名顾问,一半的工作是告诉使用 Excel 的人改用数据库,另一半工作是告诉使用数据库的人改用 Excel。所以这里有一个模糊混乱的地带:你需要专用软件吗?你需要横向软件吗?你需要纵向软件吗?但你不能只用 Excel 做所有事情。你知道,我们都见过靠一个 10 兆文件运行的部门。现在,我用 Numbers 运行我的业务,但用的是电子表格,但到了某个点你会超越它。那么,顺着这个思路,模型实验室能构建所有这些吗?当然不能。就像微软或苹果不能构建每一个 Windows 应用或每一个 iPhone 应用一样。那么,模型实验室有杠杆吗?它们是 Windows 吗?它们是 iOS 吗?同样,有网络效应吗?比如,如果你现在是一家律师事务所,购买了一个软件——你知道 A16Z 投资的所有企业软件——律师事务所、制造公司或银行多久会说“哦,这个用的是 Claude 还是 OpenAI,因为我们标准化在 Claude 上”?不,不是这样的,就像云服务那样。你不会说“我们公司标准化在 AWS 上”。你甚至不知道那个 SaaS 产品运行在哪个云上。这就是重点:它被抽象掉了,不是你的问题。所以基础模型看起来更像那样。从这个意义上说,它们看起来更像超大规模云服务商,因为它们可能拥有竞争优势,但在更上层的堆栈中,你没有杠杆,没有网络效应,没有控制权。这顺便促使我说,也许正确的比较是半导体行业,每一代都变得更贵,所以玩家越来越少。综上所述,模型是一种差异化商品,聊天机器人不是正确的用户界面或产品,公司也无法自己构建所有这些东西。因此,它们是底层基础设施。那么,它们有定价权吗?好吧,你会有,随便选一个数字,三到六家公司制造前沿模型,每年花费没人知道——没人诚实地知道——大约在 2000 亿到 2 万亿美元之间。此外,还会有一堆边缘模型和一堆开源模型。那么这会稳定在哪里?你会有大概六家公司相互竞争销售这些东西。那么价格纪律从何而来?尤其是当其中一些公司还有完全不同的商业模式时,比如 Google 卖广告,所以它们对定价的态度与 OpenAI 不同。所以我认为这里的挑战是,我们现在所处的位置和最终应该达到的位置之间存在差异,这有点像一年级经济学学生的讨论。现在我们处于供需、价格和容量的极度失衡期。但仅仅因为对 token 的需求是无限的,并不意味着你不能达到不同的价格均衡,因为移动数据就是这样发生的。对比特的需求是无限的。在过去 15 年里增长了 1500 倍。但你仍然有供需价格均衡,并且在世界大部分地区,电信运营商之间仍然有残酷的价格战,因为从根本上说,你是在向那些会来回切换的人销售一种商品,当然开发者也会来回切换。
So I think there's like three or four building blocks you can put on the table. One of them is that it's not clear how you could build a model that was fundamentally better than everybody else's model in some sort of sustainable differentiated way. There doesn't seem to be a network effect. There doesn't seem to be sort of levers you can pull and a strategy where Instagram is or YouTube is or Google searches. And we don't see an equivalent of that for LLMs. Now you have different emphasis. You know this maybe this one's better than that, maybe you like this one more than that, but there doesn't seem to be a sort of fundamental differentiation, fundamental competitive difference between the models except your willingness to spend money. Second problem is the chatbot itself is like a kind of a weird limited v1 UI and there's some things and some people and some kind of task where it works really well, but for most of the others you need a bunch of other stuff: you need tooling and it needs to be set up right. It needs to have the right data and it needs to be configured and controlled and have the right user interface and people need to have kind of sat down and thought about how this should work because generally people who are good at using the tool and doing the job that needs the tool are not the same people who are good at deciding what the tool should be. So you know people who are really really good at designing print publications are not the people who should create and design. That's a different set of skills. And you know people who are really really good at doing financial advice are not the right people to design TurboTax. Those are different people with different skills. So you have kind of groping around the middle of this. So you now have Claude for this, Claude for that and you have skills and so on. To me this is kind of like well one question is who builds the skill? Another question is like well that seems to be a bit like what you get if you do File New in Excel: like these are templates and they'll take you so far but at a certain point people outgrow the templates. There's a slide in my presentation which is a quote from somebody said to me on Twitter years ago: they said they were a consultant and half of the jobs were telling people who used Excel to use a database and the other half were telling people who used a database to use Excel. So there's this kind of fuzzy swirly place of like do you need dedicated software? Do you need horizontal software? Do you need vertical software? But you can't just do everything in Excel. There's always, you know, we've all seen the department that runs along on a 10 meg file. Now, I run my business in Numbers, but on a spreadsheet, but like there's a certain point where you outgrow that. And so following that on, well, can the model labs build all of that? Well, of course not. No more than Microsoft or Apple could build every Windows app or every iPhone app. So then, do the model labs have leverage? Are they Windows? Are they iOS? And again, well, is there a network effect? Like if you're a law firm right now and you buy a piece of software like you know do the all the pieces of enterprise software that A16Z is invested in, how often does the law firm or the manufacturing company or the bank say 'oh well does this use Claude or does it use OpenAI because we standardize on Claude'? Well no, that's not how it works anymore than it worked like that for cloud. You didn't say 'well our company standardized on AWS'. You don't even know what cloud what that SaaS product runs on. That's the whole point: it's abstracted away, it's not your problem. And so the foundation models seem to look more like that. They seem to look more like the hyperscalers in that sense in that they might have competitive advantages that further up the stack you don't have leverage, you don't have a network effect, you don't have control. That sort of prompts me incidentally to say well maybe the right comparison here is with semiconductors where with each generation it just gets more expensive and so you have fewer players. So all of that kind of taken together, the models are kind of differentiated commodities and the chatbot isn't the right UI or the right product and the companies aren't going to be able to build all of that stuff themselves. So therefore, they're low-level infrastructure. And so then, well, do they have pricing power? Well, you're going to have, pick a number, three to six companies making a frontier model, spending no one knows, no one honest knows, like something between $200 billion and $2 trillion a year on building these models. Plus, there'll be a bunch of edge and a bunch of open source. So where's this going to settle down? You're going to have as it might be half a dozen companies that are all competing to sell this stuff. And so where is the price discipline going to come from? Particularly when some of them have got like whole other business models as well like you know Google selling ads so they've got a different attitude to pricing to OpenAI. And so like I think the challenge here is there's a difference between where we are right now and where this should end up which is kind of a first year economics student kind of conversation. Right now we're in this period of extreme disequilibrium of supply and demand and price and capacity. But just because demand for tokens is infinite that doesn't mean that you can't get to a different price equilibrium because of course that's what happened with mobile data. Like demand for bits is infinite. It's grown 1500x in the last 15 years. But you still got your supply and demand price equilibrium and you still got a murderous price war between telcos in most parts of the world because fundamentally you're selling kind of a commodity to people who will swap back and forth and of course developers will also swap back and forth.
现在,我很高兴地说,这可能是完全错误的。我们可能会进入一个只有两家公司能造大语言模型、它们拥有定价权的世界,或者我们进入一个我们做的大部分事情都被模型吸收的世界,或者它们在上层拥有杠杆。这有点像我对 iOS 与 Android 的看法。仅仅因为你可以说,嗯,过去三次都是这样运作的,但这并不能证明这次也会如此。但这并不意味着你至少不应该提出这些问题。而且你当然应该,我只想说作为一个初步观察,目前这种状况是暂时的。我们处于极度稀缺的状态,然后我们有定价体系,有自由市场,还有资本支出激增,比如一万亿美元的资本支出。所以那些倍数会发生变化。
Now, this is, you know, I'm happy to say that this might be completely wrong. It may be that we get to a world in which there's only two companies that can make an LLM and they have pricing power, or we get to a world in which most of what we do gets subsumed into the model, or they have leverage further up the stack. And, you know, it's kind of my point about iOS versus Android. Just because you can say, well, it worked like that the last three times, that doesn't prove that it's what's going to happen this time. But it doesn't mean you shouldn't at least ask the questions. And you should certainly, I'll just say as a sort of a primary observation, like this situation right now is transitory. You know, we're in this extreme scarcity, and then we have a pricing system, and we have a free market, and we have a surge of capex, like a trillion dollars of capex. So those multiples are going to move around.
回到你之前提到的观点,比如,苹果就是苹果。作为过渡,下一个问题,你最关注哪些问题,或者我们应该重点关注哪些问题?
Going back to your point you made earlier, like, hey, you know, Apple's Apple. One next question, as a segue, what are some of the next questions that you're most focused on, or that we should be paying most attention to?
所以我认为回答这个问题的一种方式是,我们已经讨论过一些问题,比如模型能走多远,模型能否差异化等等。另一个问题显然是,在什么时候我们会看到越来越多的用例类别,其中模型已经足够好,我们不需要云端最昂贵、最快、最大、最重的模型,而是可以使用较旧的模型、开源模型,或者在设备上运行模型。显然,这就是苹果将在几周后讨论的内容。你能把多少计算推到设备上,那里的算力是免费的,或者至少对你来说是免费的,对开发者没有边际成本。
So I think one way to answer that is, some of the questions we've already talked about, like how far do the models go, can the models differentiate, and so on. I think another is obviously, at what point do we see more and more classes of use case where the models are good enough and we don't need the most expensive, fastest, biggest, heaviest model in the cloud, and you can use an older model, you can use an open source model, you can have a model running on device. Obviously, this is what Apple's going to be talking about in a couple of weeks. You know, how much can you push onto the device where the compute is free, or free to you anyway, doesn't have marginal cost for the developer.
另一个经典问题是,这几乎像是问题从技术领域转移出去了。所以如果你看一家律师事务所、咨询公司、投资银行,或者基本上任何专业服务领域,传统上都有金字塔结构,而你可以自动化金字塔底层人员做的大量工作,会发生什么?你只能这么说:如果你从未在律师事务所工作过,或者从未在贝恩、BCG、麦肯锡工作过,你可能不太清楚这是如何运作的,因为你可能并不真正知道那些初级员工在做什么,也不真正知道客户在为什么付费,以及这些东西如何重新配置。那么,AI 对金融意味着什么,既对内部招聘结构,也对你能创造的产品类型和利润率结构?对咨询公司意味着什么?对四大、三大、埃森哲、大型律师事务所和广告业意味着什么?你可能知道其中一些问题,但如果你不在那个行业,你并不真正知道答案。
Another classical question is, it's almost like the question moves out of technology. So if you're looking at a law firm or a consultancy or an investment bank, or basically anyone in professional services where you traditionally have this pyramid structure, and you can automate a great chunk of what the people at the bottom of the pyramid were doing, what happens? And the only thing you can say there is, if you have never worked at a law firm or never worked at Bain, BCG, McKinsey, you probably won't have a good idea of how this works, because you probably don't really know what it is that all those associates are doing, and you also don't really know what it is that the client is paying for, and how do those things get reconfigured. So, what does AI mean for finance, both for that internal hiring structure and the kind of products you can create and the margin structure? What does it mean for consultants? What does it mean for the big four, for the big three, for Accenture, for big law firms, and for advertising? And you can probably know some of those questions, but if you're not kind of in that industry, you don't really know the answers.
这让我想起了很多我在 A6C 时写的东西,我称之为“内容不是王”。我还写过“Netflix 不是科技公司”。我的观点是,如果你看 Netflix,整个事情是由科技行业构建的东西支撑的。但 Netflix 的所有问题都是电视/洛杉矶问题:什么节目、多少节目、什么类型的节目、该付给人才多少钱、是否该追求奖项、是否该做电影、是否该买体育版权、什么类型的体育?这些都是洛杉矶问题,不是旧金山问题。旧金山甚至没人知道正确的问题是什么。它们是媒体行业的问题。这就是我的观点:所有对 Netflix 重要的问题都变成了媒体行业的问题。这显然是关于特斯拉的巨大张力点。它是汽车公司还是科技公司?所以我的意思是,这些东西对法律意味着什么,既是律师的问题,也是那些非常了解律师事务所、它们实际如何运作、实际在做什么、客户实际在购买什么的人的问题。同样,生成式视频对好莱坞意味着什么?本·阿弗莱克可能比我更了解这个。他创办了一家公司,并以大约一亿美元的价格卖掉了。所以他显然懂。所以这是第二个问题:问题从 AI 领域转移出去,变成了半 AI、半其他领域的问题。
This reminds me a lot of something I wrote when I was at A6C, which I called 'Content Isn't King'. And I also wrote something that said 'Netflix Isn't a Tech Company'. The point I was getting at is that if you looked at Netflix, this whole thing is enabled by stuff the tech industry built. But all the questions for Netflix are TV/LA questions: what shows, how many shows, what kind of shows, what should you pay the talent, should you aim for awards, should you do movies, should you buy sports, what kind of sports? These are all Los Angeles questions. They are not San Francisco questions. No one in San Francisco even knows what the right questions are. They're media industry questions. And this was my point: that all the questions that matter to Netflix have become media industry questions. This is obviously the great tension point about Tesla. Is it a car company or a technology company? So what I'm getting at is, what does this stuff mean for law is kind of a question for lawyers as much as it is for people who understand a lot about law firms and how they actually work and what they're actually doing and what the clients are actually buying from them. Same thing for what does generative video mean for Hollywood? Ben Affleck probably knows a lot more about this than I do. He built a company and sold it for like a hundred million dollars. So obviously he does. So that's kind of a second question: the questions move outside of AI and they become sort of half AI questions, half something else questions.
然后是第三个层面,我想我可能应该早点说,那就是这一切与之前的平台转变有根本性的不同。对于 3G、iPhone 或网络等,你不知道接下来会发生什么,但你知道物理限制。比如 1995 年,你知道电信公司不会在下周给全世界每个人提供宽带,你也知道全世界的人不会都去买 PC,因为一台 PC 要 3000 美元。所以你大致知道什么可能发生、什么不可能发生的基本物理限制。而对于生成式 AI,显然我们不知道。我们可能在录完这个节目后看手机,看到一条推送通知说 OpenAI 的新模型发布了,价格只有原来的 2%,因为他们解决了某些问题。我不认为这在目前很可能会发生,但我们不知道这类问题的答案。所以,模型会变得多大?多好?多快?多便宜?模型的特性会以哪些方式改变?我们不知道。这与之前的平台转变不同,那时你确实知道基本约束。所以这会衍生出你的问题。从某种意义上说,这是我之前指出的:目前有产品-市场契合度的是编程。其他任何领域目前都没有同等的产品-市场契合度。我认为我这么说相当安全。当 Swapped 从去年年底的 90 亿美元年化收入增长到 470 亿美元年化收入时,那全是软件,不是吗?那么当其他领域的某个人让某个东西起作用时,会发生什么?
And then the third level, which I think I probably should have said earlier, is the way that all of this is sort of fundamentally different from previous platform shifts is that with 3G or the iPhone or the web or whatever it was, you didn't know what was going to happen next, but you knew the physical limits. Like in 1995, you knew that telcos weren't going to give everybody in the world broadband next week, and you knew that everyone in the world wasn't going to go out and buy a PC because a PC cost like $3,000. So you kind of knew the basic physical limits of what could and couldn't happen. And with generative AI, obviously, we don't. We might look at our phones when we get off this recording and there's a push notification that says that OpenAI's new model is out and it's like 2% of the price because they worked something out. I don't think it's very likely at this point, but we don't know those kinds of questions. So, how much bigger will the models get? How much better? How much faster? How much cheaper? In what ways will the characteristics of models change? We don't know. And that is different from previous platform shifts, where you did know the sort of fundamental constraints. And so that will spin off your questions. And in a sense, this is something I pointed to earlier: the place that's got product-market fit right now is coding. Nothing else has equivalent product-market fit right now. I think I'm pretty safe in saying that. When Swapped has gone from whatever it was, 9 billion run rate at the end of last year to $47 billion run rate, that's all software, isn't it? So what happens when someone else in some other field gets something working?
是的。如果你必须猜一下,编程之外有哪些用例可能产生日常活动?
Yeah. If you had to guess, what are the use cases outside of coding that could potentially yield daily activity?
所以,我几周前发布的那场演讲大致分为三个部分。第一部分是关于资本、资本支出、基础设施、基础模型和差异化,也就是我们刚才聊的那些。第二部分是:你如何用这些东西来构建软件?这对软件行业意味着什么?软件会变成什么样?利润率、公司以及其他一切又会发生什么变化?第三部分我称之为“变革”,也就是要讲到这个点了。
So there, I should say the sort of presentation that I published a couple of weeks ago has three sections. One of them is talking about capital and capex and infrastructure and foundation models and differentiation, which is the stuff we talked about. The second is: how would you build software with this, and what does this do for the software industry, and what would software look like, and what happens to the margins and the companies and everything else? And the third section I called 'Change', which is kind of getting to this point.
我以一句似乎会让某类人不爽的话开场,引用了尤吉·贝拉的名言:“预测很难,尤其是关于未来的预测。”我觉得这有一个回溯测试的点:想象一下在 1997 年问关于互联网的这类问题。你会得到什么?你又错过了什么?但我觉得看待这个问题的一种方式是:这就是自动化。它让一类过去人们做但无法自动化的事情,现在可以自动化了。然后问:那意味着什么?
I opened it with, again, what appears to upset a certain category of person, where I said the Yogi Berra quote: 'Predictions are hard, especially about the future.' And I think there's a sort of back-test point: imagine asking these kinds of questions about the internet in 1997. What would you have got? What would you have not got? But I think one way you can look at this is to say: this is automation. This makes a class of thing that people used to do that couldn't be automated, now you can automate that. And then say, well, what does that mean?
我提出了三四个可以按下的按钮。第一个是:这仅仅是价格弹性吗?这其实就是杰文斯悖论。比如,如果你让做事情变得更便宜,你是花更少的钱做同样多的事,还是花同样的钱做更多的事,还是花更多的钱做更多的事?因为它变得如此便宜。
I proposed three or four buttons to press. First one is: is this just price elasticity, which is really what the Jevons paradox is? Like, if you make it cheaper to do stuff, do you do the same amount of stuff for less money, or do you do more for the same money, or do you do more for more money? Because it becomes so much cheaper.
有没有以前做不到、现在变得便宜的事情?有没有以前很贵、是进入门槛的事情,比如像报社拥有印刷机那样?有没有某个成本上的进入门槛现在消失了?有没有因为这东西变便宜而在你的商业模式或竞争空间中被解锁的东西?最后一个问题是:什么东西以前完全不可能、成本高得令人望而却步,以至于根本没人想过,而现在却触手可及?
Was there something that you couldn't do before that now becomes cheap? Was there something that was expensive and was a barrier to entry, like owning a printing press as a newspaper? Is there something that was a barrier to entry in a cost space that now goes away? Is there something that gets unlocked in your business model or in your competitive space because this thing became cheap? And then the final question would be: what stuff was just completely impossible, cost-prohibitive, so that nobody even thought about it, and now that's within reach?
我在这里举的例子是:蒸汽机让火车成为可能,你买多少匹马都没用,你造不出火车或特快列车。一个更当代的例子是 YouTube 或者 Spotify。你知道,Spotify 的故事是:看看音乐行业过去 25 年。前半段是,如果你不用花 15 美元买一张 CD 来得到那首歌,会发生什么?但后半段是:如果每月 15 美元就能让你听到所有音乐呢?这在以前是完全不可能的。
The example I used to give here was: steam engines make trains possible, and it wouldn't matter how many horses you buy, you couldn't have a train or an express train. A much more contemporary example would be to point to something like YouTube, or indeed Spotify. Like, you know, Spotify says: step one, you look at the last 25 years of the music business. The first half is what happens if you don't have to buy a $15 CD to get that track. But then the second half is: what if $15 a month gets you all the music that there is? Which is something that was just completely impossible.
这就是做这类预测的问题所在。一方面,你会说一些听起来聪明又显而易见的话,但你实际上并不知道它对每个行业具体意味着什么。所以如果我们回到 90 年代末,说“我们知道互联网会摧毁实体分销的价值”,结果这对报纸和电影制片厂意味着完全不同的东西。报纸被彻底搞垮了,而电影制片厂基本上没怎么变。所以,还是得看情况。
This is all the problem with making predictions like this. On the one hand, you're going to say stuff that's kind of clever and obvious, but you don't actually know what it's going to mean industry by industry. So if we'd been back in the late '90s and we'd said 'we know the internet will destroy the value of physical distribution,' it turned out that meant completely different things for newspapers and movie studios. Newspapers got completely screwed by this, and movie studios are kind of not really changed very much. So again, it depends.
另一部分是:有些地方我觉得你可以问更有用的问题。让我感兴趣的一个是:这会如何改变广告、电商、品牌、营销以及我们购买的一切?因为广告是万亿美元级别的,零售是 25 万亿美元,所以这是一个规模可观的市场。我过去一直思考的是,Google、Meta 和 Amazon 并不真正知道那个产品是什么。它们知道它是一个 SKU,知道发布者在元数据字段里输入了什么,也知道买这个的人也买了那个。但它们不知道为什么,也不真正知道那些东西是什么。这就是为什么会有这样的笑话:“嘿 Amazon,我买了一个马桶座圈套,我不是在收集马桶座圈。”因为它并不真正知道马桶座圈是什么,也不知道人们不会买两个。实际上,它们应该知道,那应该是频率分析,但它们不知道。
The other part of this is: there are some places where I think you can ask more useful questions. The one that intrigues me is: how does this change advertising, e-commerce, brands, marketing, and everything we buy? Because advertising is a trillion dollars and retail is $25 trillion, so it's a reasonable-sized TAM. The thing I've always used to think about was that Google, Meta, and Amazon don't really know what that product is. They know it's an SKU, they know what the publisher typed in the metadata field, and they know that people who bought this also bought that. But they don't know why, and they don't really know what those things are. That's why you get jokes like: 'Hey Amazon, I bought a toilet seat cover, I'm not collecting toilet seats.' Because it doesn't really know what a toilet seat is, and doesn't know that people don't buy two. Actually, they should know that, that should be frequency analysis, but they don't.
有了 LLM,原则上你会知道那些东西是什么,人们为什么买它们,以及人们还会买其他什么东西。当然,“知道”是一个很难用的词。你说“知道”是什么意思?但至少,AI 系统能够达到的统计相关性水平会非常不同。这当然就是为什么你看到 Google 和 Facebook 的广告数据和转化率每个季度都在飙升,因为它们正在把所有这些整合到广告系统、推荐引擎和预测算法中。你会看到更多你喜欢的东西,你看到的广告也更可能是你想买的东西。所以它们的广告收入出现了突然的加速增长。
With an LLM, in principle you would kind of know what those things are and why people buy them and what other things people buy. And obviously 'know' is a difficult, tricky term to use. What do you mean when you say 'know'? But at a minimum, a much different level of statistical correlation of what an AI system would be able to do. Which is of course why you see the ad numbers and the conversion rates shooting up every quarter from Google and Facebook, because they're rolling all of this into their ad systems, recommendation engines, and prediction algorithms. You get shown more stuff that you would like, and the ads you're seeing are more likely to be things you'd like to buy. So they have this sudden acceleration in their ad revenue.
所有这一切都在说:你看看这些系统是如何工作的。现在它们说:“买那个的人可能会买这个。”你现在应该能够说:这里有一件外套的图片,它是什么,我在哪里可以买到?五年前,这真的行不通。十年前,肯定不行。五年前,可能也不行。现在应该可以了。然后你可以说:“好的,推荐 10 件类似但价格不同的外套,告诉我在哪里可以买到,并给出每件的优缺点。”你大概也能得到这个。然后你可以再进一步说:“看看我的 Instagram,推荐一件我应该买的冬装,能改变我的形象,但不要改变太多。”同样,三年前这完全是科幻小说。现在你会想,是的,你大概能做出一个能用的东西。
All of which is to say: you look at how these systems work. Right now they say: 'People who bought that could buy this.' You should now be able to say: here's a picture of a coat, what is it, where can I buy that? Five years ago, that really wouldn't work. Ten years ago, certainly wouldn't work. Five years ago, probably wouldn't work. Now that should work. And then you can say: 'Okay, suggest 10 other coats like that with different prices, tell me where I can buy them, and suggest the pros and cons of each one.' And you'll kind of get that too. And then you can push one step further and say: 'Look at my Instagram and suggest a winter coat I should buy that will change my look, but not too much.' Again, three years ago that would have been total science fiction. And now you think, yeah, you could probably build something like that that would kind of work.
这些关于计算机知道什么、能自动化什么、能提出什么建议的转变——回到最开始:每当你获得一项新技术,你都是从用旧方式做更多开始:更多的电子表格、更多的 PowerPoint、更多的电子邮件、更好的电子邮件。但重要的不是用旧方式做更多。而是做一些你用旧方式做不到的新事情。我的意思是,这是一个相当老生常谈的观察,但我们往往会忽略它。
Those kinds of shifts in what the computer knows, what it can automate, what suggestions it can make — going back right to the beginning: whenever you get a new technology, you start by doing the old thing but more: more spreadsheets, more PowerPoints, more email, better email. But the important stuff is not doing the old thing but more. It's doing something new that you couldn't have done with the old thing. I mean, it's a pretty banal observation, but we kind of lose sight of it.
那么,相比于自动化旧有工作,有哪些新事情是只有借助这个才能做到的?
And so what are the new things that you can only do with this as opposed to automating the old stuff?
我认为企业版的应用场景是:你拥有所有与客户 Zoom 通话的录音、Salesforce 中所有邮件的往来流程,以及用户如何使用你产品的所有遥测、指标和分析数据。那么,我们应该如何调整定价来改善客户流失率?这是大语言模型可能做到的,这跟“对呼叫中心的通话进行情感分析,然后告诉我哪些客户生气了”完全不同。你能进行的分析在抽象层面上发生了多重转变。当然,这就会催生新公司、摧毁旧公司,并创造新业务。但话说回来,我们现在就像在 1997 年,试图预测 Uber 和 Airbnb。如果我真能预测,那有个普遍观点:如果我们能预测未来,我们就活在平行宇宙了。风投的命中率就会是 100%,而不是十分之一。我们现在问的一个问题是:以前贵得离谱的事情,现在变得可能了?也许是像从头重建 YouTube 或重写 Linux 这样疯狂的事。
I think the enterprise version of this would be: you've got all Zoom calls with clients recorded, all the flows of emails in and out of Salesforce, and all the telemetry, metrics, and analytics of how people use our product. So how should we change our prices to improve our churn? That's something an LLM might be able to do, which is very different from saying, 'Do sentiment analysis on calls into the call center and tell me which customers are angry.' You get multiple shifts in the layer of abstraction around what analysis you can do. Of course, that then creates new companies and destroys old ones, and creates new businesses. But again, we're in 1997 and I'm trying to predict Uber and Airbnb. If I could actually do that, there's a general point: if we could predict what was going to happen, we'd live in a parallel universe. VCs would have a 10 out of 10 hit rate, not one in 10. One of the questions we're now asking is: what was unreasonably expensive to do before that is now possible? Maybe something crazy like rebuilding YouTube from scratch or rewriting Linux from scratch.
是的,这很有趣。一个常见的谬误是:新事物出现后,人们说“我们要用新东西来重建旧东西”。当然,我们要用开源来重建 Office,在网页上重建它。但结果呢?看看 Google Docs——它只占了 20% 的市场,因为那根本不是重点。有趣的是去做别的事情,做新的事情,去转移那个抽象层次,发现从未存在的问题。你在风投公司整天听路演时会有这种体验:有些东西你觉得“听起来有点用”,有些东西你觉得“我不确定那为什么会成功”。但有些东西填补了宇宙中的一个空洞。一旦有人解释给你听,你就会想:“哇,为什么之前没人做过?为什么没人看到那个东西存在?”这就是看初创公司的乐趣之一。人们会突然发现某个问题存在,而包括有这个问题的人在内,没人意识到它存在。然后他们就会去制造一个东西来解决它。这也是为什么我不认为模型能包揽一切。如果你回想一下你加入 S6C 以来看到的所有项目,有多少是业内人知道那是个问题的?通常答案是没有。业内没人觉得那是个问题,而且花了两年时间才向他们解释并说服他们那个问题确实存在。这就是为什么“金融业的中层管理者会用这个工具来解决一个巨大的全球行业问题”这种想法有问题。不,因为没人知道那个行业问题存在,更不用说想出正确的方法来构建一个工具去解决它了。
Yeah, it's funny. The paired fallacy is: the new thing comes along and says, 'We're going to build the old thing with the new thing.' Of course, we're going to build Office with open source, rebuild it on the web. And it turns out, look at Google Docs—it's got like 20% of the market because that's not the point. What's interesting is to do something else, something new, to shift that level of abstraction and spot problems that have never existed. It's the experience you get sitting in pitches all day at a venture firm: there's stuff where you think, 'That sounds kind of useful,' and stuff where you think, 'I'm not sure why that would work.' But there are some things that fill a hole in the universe. As soon as somebody explains it, you think, 'Wow, why did nobody do that before? Why did no one see that that thing existed?' That's part of the fun of looking at startups. People will suddenly work out a way that that problem existed, and no one—including the people who have that problem—realized it existed. Then they'll go out and make a thing to solve it. This is also why I don't see the model doing the whole thing. If you go back and think about all the pictures you've seen since you joined S6C, how many were things where people in the industry knew that was a problem? Often the answer is no. No one in the industry thought that was a problem, and it took two years to explain to them and persuade them that the problem existed at all. That's the problem with the idea that a middle manager in finance will use this tool to solve a big global industry problem. No, because no one knew that industry problem was there, let alone could work out the right way to build a tool to solve it.
这是否意味着 AI 之后的 SaaS 环境会比之前更不整合?也许更少捆绑,或者更少像微软企业那样的单一巨头。
Does this imply a less consolidated SaaS environment than before AI? Maybe less bundling or single behemoths like the Microsoft Enterprise.
天哪,你把我拉回现实了。SaaS 行业会变得更不整合吗,Benedict?这些都很好,但跟我们说说股票吧。我们能在这里打下什么样的基础?显然,构建软件会变得更便宜、更快速。显然,会有很多以前完全无法用软件做到的事情现在可以做了。所以竞争会更激烈。当然,这会带来新的利润率结构,但正如我们之前讨论的,我们并不真正知道那个利润率结构会是什么样子。你会转向基于结果的定价吗?要把企业软件中的每次按钮点击与损益表挂钩非常困难。有时在 Salesforce 之类的软件里可以,但有很多软件很难说:“我今天做的工作对 DPS 产生了这个影响,因此我们应该为此付这么多钱。”我认为从长远来看这说不通。无论如何,与现在相比,定价结构会如何演变?竞争会更激烈。构建东西会更容易、更快速。我对此的思考方式有两种有用的框架。一种是:如果你看看今天的企业软件版图,有三个类别。有大型横向系统:SAP、Workday、CRM、资本管理软件、薪资管理软件等等。然后是垂直软件。一个典型的美国大公司有大约 300 到 400 个 SaaS 应用,另外还有一千个他们自己构建或购买并在本地运行的应用。中间是 Excel、电子邮件和共享文件系统这个模糊的临时空间。东西在这些类别之间来回移动。原则上,每个 SaaS 应用都在做你本可以在 SAP 或 Excel 中完成的事情。例如,你可以在 Workday 中管理毕业生招聘,但在某个点上,如果你是普华永道,每年招聘数千名毕业生培训成会计师,你可能会有一款自己构建或请埃森哲构建的专用软件,而且你可能讨厌它。如果你是一家每年只招聘五名毕业生的公司,你会在电子邮件和共享的 Google 表格中完成这件事,因为你为什么要买软件呢?然后中间还有一个空间。
Gosh, way to bring me back down to earth. Is the SaaS industry going to be less consolidated, Benedict? That's all great, but tell us about the stocks. What are the building blocks we can put down here? Obviously, it's going to be way cheaper and quicker to build software. Obviously, there's going to be a bunch of stuff you could do with software that you just couldn't do before at all. So there will be more competition. Of course, this comes with a new margin structure, but as per our conversation earlier, we don't really know what that margin structure is going to look like. Are you going to go to outcome-based pricing? It's really hard to tie each button press in a piece of enterprise software to P&L. Sometimes you can in Salesforce or something, but there's an awful lot of software where it would be really hard to say, 'The work I did today did this to DPS, therefore this is what we should pay for it.' I don't think that makes sense in the long run. Anyway, what does the pricing structure look like over time versus now? There will be more competition. It will be easier to build stuff and quicker to build stuff. The way I think about this is there are maybe two framings that are useful. One is to say that if you think about the enterprise software fleet today, you've got three buckets. You've got your big iron horizontal systems: SAP, Workday, CRM, capital management software, payroll management software, and so on. Then you've got vertical software. A typical big US company has like 300 to 400 SaaS apps, and then another thousand apps that they built or bought internally running on prem. In the middle, you've got this fuzzy improvised space of Excel, email, and the shared file system. Stuff moves back and forth between those. In principle, every SaaS app is doing something you could have done in SAP or Excel. For example, you could have managed your graduate recruiting in Workday, but at a certain point, if you're PwC and you hire however many thousand graduates every year to train to be accountants, you probably have a piece of dedicated software that you built for yourself or hired Accenture to build, and you probably hate it. If you are a company that hires five graduates a year, you're doing that in email and a shared Google Sheet, because why would you buy software for that? Then there's a space in the middle.
你是用 Workday 做?用 Excel 做?用专用 App 做?现在又加上了聊天功能。你是用 LLM 做吗?有没有一个 LLM 工具让你能在 Salesforce 里做以前做不到的事?或者能在你的垂直应用里做以前做不到的事?你用 LLM 给自己搭个工具吗?就像某个公司部门可能还在用 15 年前有人搭的 10 兆 Excel 表格一样。没人知道它为什么能运行,没人知道它怎么工作的,但他们还在用。所以它就这样进入了一个广阔、碎片化、复杂的格局,成为完成那项任务的又一套选项。这是一种思考框架。另一种思考框架是:LLM 是放在栈顶还是栈底?一方面,栈底是 Salesforce 里的一个功能。你在 Salesforce 里,查看这个客户的历史记录、我们所有销售电话的上下文、我们的业务目标,然后建议一封邮件,或者建议我该怎么做、该在电话里对客户说什么。所以它是一个功能,一个受控的按钮,有工具和护栏,由那个特定用例驱动。另一种看法是我之前举的例子:去查看 Salesforce、Workday、我们所有的邮件和 Google Analytics,然后综合出一些你以前做不到的东西。这两种情况的张力在于:你把可能出错的概率性软件放在哪里,把无法回答这类问题的确定性系统软件放在哪里?你把数据库放在哪里,把 LLM 放在哪里?哪个在顶层,哪个在底层?答案可能是两者都有,取决于你在做什么以及它放在哪里。所有这些绕了一大圈,其实是在说:这对软件意味着什么?答案是更多的软件,多得多的软件。我的意思是,所有软件公司都是为了解决其他软件公司制造的问题而存在的。这是安全领域的一个笑话:所有安全软件都是为了解决其他安全软件制造的问题而存在的。显然,我们在 SaaS 时代就经历过这个。SaaS 给我们带来了一个数量级、两个数量级更多的软件。我们很可能也会看到这种情况。这导致 SaaS 末日论:所有投资者都在看着这些公司说,我们真不知道哪些公司会被这一切搞垮。肯定有一些,显然会有一些公司走到最后,x% 的 SaaS 公司会被消灭,但你不知道是哪些,所以你不应该把整个行业估值打五折。但显然你会想,在我搞清楚到底怎么回事之前,我不确定现在会长期持有软件股。
Do you do it in Workday? Do you do it in Excel? Do you do it in a dedicated app? And now you add chat to that. Do you do that in an LLM? Is there an LLM tool that means you can do that in Salesforce where you couldn't do it before? Or you can do it in your vertical app that you couldn't do before? Do you use the LLM to build yourself a tool for that? Just as you might have a company department that runs on a 10-meg Excel spreadsheet that someone built 15 years ago. No one knows it works. No one knows how it works, but they're still using that. So it arrives within this broad, fragmented, complicated landscape and it's another set of options for how you would do that task. So this is one framing to think about it. I think the other framing is: does the LLM go at the top of the stack or the bottom of the stack? On one hand, bottom of the stack is a feature inside Salesforce. So you're in Salesforce, look at the history with this customer, look at the context of every other sales call we've done, look at our business objectives, and suggest an email or suggest what I should do here and what I should say on the call to the customer. So it's a feature. It's a button that's controlled and has tooling and guardrails driven by that particular use case. The other way to look at it is the example I gave earlier: go look at Salesforce and Workday and all of our email and Google Analytics and then synthesize something you couldn't have done before. So the tension in both cases is: where do you put the probabilistic software that can make mistakes, and where do you put the deterministic system software that can't answer these kinds of questions? So where do you put the database and where do you put the LLM? Which is at the top and which is at the bottom? The answer is probably both, depending on what you're doing and where it goes. All of which is a long way of saying: what does this do to software? The answer is more software, like way more software. I mean, all software companies exist to solve problems created by other software companies. That was the joke in security: all security software exists to solve problems created by other security software. And clearly that's what we went through with SaaS. SaaS gave us an order of magnitude, two orders of magnitude more software. And we should probably expect that with this. What that gets to the SaaS apocalypse is all the investors are looking at these companies and saying, we don't really know which of these companies are going to get screwed by all of this. Some of them must be, obviously there must be some that go through the end of this and x% of all the SaaS companies out there are going to get wiped out by this, but you don't know which ones, so you probably shouldn't derate the whole thing by 50%. But clearly you're going to go, I'm not sure I'm going to be long software at the moment until I have some idea of what the hell's going on.
你在与 Ben Thompson 的对话中说,软件是有人坐下来设计了一个工作流,然后说从今以后这就是做这件事的正确方式。但你也说过,流程是从企业运行方式中生长出来的。这仅仅是需要时间,还是你认为我们需要更多来自这些垂直 AI 初创公司的实验和迭代,才能找到未来软件的正确形态?
You said in your talk with Ben Thompson that software is someone sat down and designed a workflow and said this is the right way of doing this from now on. But you also said that a process grows out of the way a business runs. Does that just take time, or do you think we need more experimentation, iteration from these vertical AI startups to get the right shape of software for the future?
嗯,从某种意义上说,也许一个有趣的转折是,这正是战略顾问和软件公司都在做的事情:他们观察公司内部的情况,然后说,嗯,这种做法很糟糕。有更好的做法,能更好地实现你的目标。软件公司将其编码到软件中,战略咨询公司则将其编码到工作流、图表、流程、培训和目标中,也许还会告诉他们购买一些软件来做这件事,或者现在越来越多地直接为他们构建软件。我认为这里要谈的另一件事是,组织内部有多少工作是隐性的、没有文档记录、不在训练数据中、也不是公司里任何人能坐下来给你画个流程图并解释清楚的。这正是 Bain、BCG、McKinsey 价值的一大部分:他们有权进入一家公司,与每个人交谈,与你在不同部门不能交谈的人交谈而不会被解雇,然后弄清楚事情实际是如何运作的,而不是应该怎样运作,以及为什么人们不执行战略——因为实际上他们的奖金目标取决于他们不执行战略——然后解决所有这些问题,成为一个从外部介入并给你答案的团队,然后你可以责怪他们,或者拥有那个预先准备好的解决方案。这些都是组织管理、人员运作方式以及人们如何解释自己工作的问题,非常难以写下来,非常难以融入一个 Claude 技能然后说,好了,做个 PPT 吧。所以这里有一个更广泛的挑战:如何让人们使用这些技术?人们如何采用新工具?如何找出帮助人们采用新工具的方法,并弄清楚你会用它们做什么新事情?这也正是云、网络、移动、互联网、PC、电子表格等发生的情况。
Well, in a sense, maybe an interesting turn on this is that this is both what strategy consultants and software companies do: they look at what's going on inside a company and say, well, this is a crap way of doing it. This would be a better way of doing it. It would achieve your objectives better. And a software company encodes that in software, and a strategy consultancy encodes that in workflows, charts, processes, training, and objectives, and maybe tells them to buy some software to do that thing, or maybe now increasingly builds them that software as well. I think another thing to talk about here is how much of what's done inside an organization is implicit and not documented, not in the training data, and not something that anybody in that company could sit down and draw you a flowchart of and explain to you. That's a big chunk of the value of Bain, BCG, McKinsey: they have license to come into a company, talk to everyone, talk to the people you're not allowed to talk to in a different org without getting fired, and go work out how this actually works as opposed to how it's supposed to work, and why people aren't doing the strategy because actually their bonus targets depend on them not doing the strategy, and work all of that out, and be a team ready to come in from the outside and give you the answer, and then you can blame them or have that pre-baked solution. These are problems in organizational management and how people function and how people can explain what they do that are very hard to write down and very hard to bake into a Claude skill and say, there you go, make a PowerPoint. And so there's a broader challenge of how does this always work: how do you get people to use these technologies? How do people adopt new tools? How do you work out how to help people adopt new tools and work out what new things you would do with them? Which is also what happened with cloud, the web, mobile, the internet, PCs, spreadsheets, and so on.
为此,你认为 AI 原生软件和新类型的界面之间是否存在某种共同进化?例如,新的客服 AI 平台可能没有那么多面向人类的 UI,或者记录系统软件根本没有前端,因为其主要用户将是直接查询它的 AI 智能体。所以我认为这些是很有趣的想法。我很难对此有强烈的看法,因为我对企业基础设施的采购方式没有深入了解。我想知道这些问题有多新。我记得 Chris Dixon 在 10-15 年前说过,API 是新的 BD,软件不需要 UI。软件公司可以只开放他们的 API。嗯,旧的东西又变新了。你不再需要 API;你只需要一个 MCP 服务器,人们就会直接接入,智能体就会直接接入。我不知道。
To that end, do you think there's some kind of co-evolution between AI-native software and new types of interfaces? For example, new customer service AI platforms that might not have had as much human-facing UI, or system of record software being built without a front end at all because its primary user will be AI agents querying it directly. So I think these are kind of interesting ideas. They're things I struggle to have a strong opinion on because I'm not deep in the weeds of how enterprise infrastructure gets bought. I wonder how new some of these questions are. I remember Chris Dixon saying 10-15 years ago that APIs are the new BD and software wouldn't need a UI. Software companies could just open up their APIs. Well, what's old is new. You don't need an API anymore; you just have an MCP server and people will just plug in, the agent will just plug into that. I don't know.
我认为这类事情的挑战在于,所有的决策实际上都是异常处理。问题总是:什么不能自动化?什么需要有人做决定、做出判断并持有观点,因为可能那件事没有被写下来,或者以前没发生过,或者看起来和以前不一样。我认为有各种方式可以思考如何区分哪些被自动化、哪些不被自动化。我在演示文稿中使用的方式是讨论什么是任务、什么是工作。通常,用于完成工作的任务可能会改变,但工作本身变化不大,或者工作向客户提供的价值变化不大。比如,想想 50 年前的会计师和今天的会计师。他们几乎不再花时间做同样的事情,但对客户来说,这差不多是同一件事。只是以完全不同的方式完成,涉及一系列不同的任务。其中一种更深刻或更抽象的方式来思考这个问题是:你在哪里想要平均值?你在哪里想要的是每个人都会这样做的方式?那是每个人都会做的,那是任何人都会说的,那是任何人都会制造的,那是任何助理都会制造的,那是任何人都会给我的,那是任何人都会给出的答案。相比之下,哪里不是你想要的?你在哪里想要一个新问题的答案,或者一个不同的答案,或者一个不同的想法?因为 LLM 非常擅长那些你可以描述人们如何做、并且你想要的是任何人都会那样做的方式的事情;而不太擅长那些你无法真正解释为什么那样做、并且你做得与人们通常做法不同的地方。
I think the challenge with a lot of this stuff is that all the decisions are really exception handling. Like the question is always what can you not automate? What requires someone to make a decision and some judgment and have an opinion about it because maybe that hasn't been written down or that didn't happen before or doesn't look quite the way it happened before. I think there are various ways of thinking about separating out what gets automated and what doesn't. The way that I used in the deck was to talk about what's a task versus what's a job. We often the tasks that are used to accomplish the job might change without the job itself changing very much or without the thing that the job is selling to the client changing very much. Like if you think about what accountants did 50 years ago and what accountants do today. They spend almost none of their time doing the same things but to the client it's kind of the same thing. It just gets done in a completely different way with a whole bunch of different tasks. And one of the more profound or abstract ways to think about this is where is it that you want the average? Where is it that what you want is the way that everybody will do this? That's the way everyone would do it. That's what anyone would say. That's what anyone would make. That's what any associate would make. That's what anybody would give me. That's the answer anyone would give. Versus where is that not what you want? Where is it that you want the answer to a new question or a different answer or a different idea? Because LLMs are going to be very good at anything where you can describe how people do it and where what you want is the way anybody would do that and not so good at where you can't really explain why you did it like that and where you're doing it differently to the way people would normally do it.
包括谷歌 CEO 在内的许多人都说,投资不足的风险比投资过度更大。是否存在某个资本支出水平,让这句话不再成立?我们现在是否正在接近那个水平?
Various people including the CEO of Google said that the risk of underinvesting is riskier than overinvesting. Is there any level of capex where that stops being true and are we getting there now?
嗯,这里有一个财务重力问题:微软、Meta 和谷歌今年都计划将超过 50% 的收入用于资本支出。我们认为电信行业是资本密集型的,但电信的资本支出只占收入的 15% 到 20%。所以今年四大公司的指引是 7000 亿美元。电信是 3000 亿,移动是 2000 亿,整个电信是 3000 亿。石油和天然气,根据定义和统计口径不同,从 7000 亿到 1 万亿美元不等。我记得这取决于你问谁。所以每年 7000 亿美元在全球大型基础设施成本中并不是一个不可能的数字。只是很多钱。显然,这些公司明年不可能花 1.5 万亿美元,如果花了,他们就得借钱,而且肯定无法长期维持这种支出水平。所以增长必须在某个点放缓,因为没有更多的钱了。当然,你可以谈论 ROI 和从投资中产生回报的能力。显然,资本市场愿意在一定程度上提供资金,但随便选一个数字,比如我们不能每年花 10 万亿美元在 AI 基础设施上,因为根本没有那么多钱。所以可用的资金有类似物理定律的上限。我目前不愿意给出比这更具体的说法。我的意思是,我几乎要回到我一开始说的:我们有很多倍数。所以需求远大于供给。另一方面,效率正在大幅提升。我们不知道下一个模型会是什么。我们不知道边缘计算和开源何时会介入。与此同时,你总是在追逐下一个模型。所以贯穿这一切的一条主线是:模型只有 3 到 6 个月、6 到 9 个月的相关性,随便你怎么说。而模型要花几十亿美元。你需要多少基础设施来做这件事?我认为这个规模还没有真正确定下来。我的意思是,显然有很多非常聪明的半导体分析师花大量时间试图给这些数字。这有点像在 90 年代末给带宽、互联网带宽定数字。你大致知道电子表格里有哪些行,但你真的不知道数值是多少。你只能说的是,嗯,这显然有物理限制。我认为回答这个问题的另一种方式是:如果你是谷歌、Meta 或微软,在某种程度上还有亚马逊和苹果,这有点像生存问题,你有一个 FOMO 问题。一方面,你目前投资的回报非常积极。另一方面,你不能让其他人在这方面领先而你却不参与,因为那样你的公司就完了,你不想像 2000 年代的微软、90 年代的 IBM 或者 2010 年代的 Meta 那样,被苹果不断打压。所以如果这是计算的未来,那么你必须参与其中。但显然,CFO 坐在那里说,嗯,这很好,但我们说的是多大程度的参与?我认为我们不知道。显然,在某个点,这条曲线必须趋平,因为它无处可去了。
Well, there's a financial gravity problem in that Microsoft, Meta and Google are all on line to spend over 50% of revenue on capex this year. And we think of telecoms as being capital intensive. Telecoms spend instead of 15 to 20% of revenue on capex. And so $700 billion is the guidance from the big four companies this year. Telecoms is 300, mobile is 200, total telecoms is 300. Oil and gas depending on which definition and which bits you're counting is anything from 700 billion to a trillion dollars. I think from memory depends on who you ask. So $700 billion a year is not an impossibly large amount of money in terms of what big global infrastructure costs. It's just a lot of money. Clearly, those companies could not spend $1.5 trillion next year or if they did, they'd have to borrow it and they certainly couldn't sustain that level of spending for any length of time. And so there's a certain point at which that growth has to slow down because there isn't any more money. Now clearly you can talk about ROI and your ability to produce returns from that investment. And clearly the capital markets are willing to fund that up to a point but pick a number at random like we can't spend $10 trillion a year on AI infrastructure because there isn't $10 trillion a year there to spend on it. So there are kind of laws of physics caps on the amount of money that's available. I'd hesitate to say something more tangible than that at the moment. I mean, I kind of almost go back to what I said at the beginning that we've got a bunch of multiples. So there's far more demand than supply. On the other hand, the efficiency is increasing massively. We don't know what the next model will be. We don't know where edge or open source come in yet, when edge and open source come in yet. And meanwhile, you're always chasing the next model. And so this is kind of the line that runs across all of it is the model is only relevant for 3 to 6 months, 6 to 9 months, whatever you want to say. And the model costs how many billion dollars. And how much infrastructure do you need to do that? I don't think that mass has really shaken out yet. I mean, obviously there are a bunch of very clever semiconductor analysts who spend lots of time trying to put numbers on this. It is kind of like trying to put numbers on bandwidth, internet bandwidth in the late 90s. You kind of know what the rows in the spreadsheet are, but you don't really know where the values are. All you can really say is, well, there are clearly physical limits on this. I think another way to answer the question is like if you're Google or Meta or Microsoft, to some extent Amazon, some extent Apple, this is sort of an existential problem and you have a FOMO problem in that on the one hand your returns on the investment at the moment are hugely positive. On the other, you can't let other people get away with this without you participating because then your company's gone and you don't want to end up like Microsoft in the 2000s or IBM in the '90s or indeed Meta in the 2010s where they are kind of continually getting shafted by Apple. So if this is the future of compute, then you need to be participating in it. But obviously at the same time the CFO is sitting there saying well yeah that's great but how much participation are we talking about here and I don't think we know. Clearly at a certain point that curve is going to have to taper off because there's nowhere else it can go.
你认为会出现关于“Token 最大化”的清算吗?公司是否可能过度使用 AI,当他们做适当的 ROI 研究时会撤回?
Do you think there's going to be a reckoning around token maxing? Is it possible that companies have been overshooting AI usage and when they do proper ROI studies they'll pull back?
嗯,显然,有人用最昂贵的模型在网上瞎折腾。这有点像 2010 年移动领域发生的情况。你收到一张 1 万美元的账单,然后你会说:“等等,我以为这是固定费率套餐。”发生了什么?所以有很多愚蠢/浪费的故事。我认为还有一个问题可能更有趣。显然,会有一个时刻——正如我多次说过的——我们正处于一个巨大的失衡时刻,定价必须与成本重新对齐,使用量必须与定价和 ROI 重新对齐。
Well, obviously, you've had people using the most expensive model to dick around on the internet. Which is kind of what happened with mobile in 2010. You got a $10,000 bill and you would have said, 'Wait, I thought this was a flat rate bundle.' What happened? So there are a bunch of silly/wasteful stories. I think there's also a point of what maybe is slightly more interesting as a question. And clearly there's going to be a point in which, as I've said several times, we're at a moment of massive disequilibrium and the pricing has got to get back into alignment with the cost and the usage has got to get into alignment with the pricing and the ROI.
挑战在于,在这么早的阶段,确实有点棘手。很难知道投资回报率是多少。这有点像在 90 年代末给每个人发互联网,然后说“好了,去提高生产力吧”。如果你去看德勤的调查,还有美联储的调查(都在我的演示里),你去问 CFO 们在哪里看到了收益,到目前为止大部分收益都是很难衡量的东西,比如更好的分析、更好的客户支持、更高的生产力。你可以更快地制作幻灯片,更快地做分析。这很难给它赋予一个财务价值。它确实有财务价值,但这不等于说“我们用 AI 做了这个新东西,它带来了这么多收入或省了这么多钱”。那些事情显然需要更长时间。建立一个新的收入来源比把这个工具给每个人、让他们更快地做电子表格要难得多。所以就会有点疑问:这需要多长时间?我认为这个问题的另一个答案当然是消费者剩余,这有点像 Excel 的情况。你知道,如果一个 DCF 模型需要你一周时间,那你可能只做一两个 DCF。如果 DCF 只需要 10 秒,那你就会做 50 个 DCF,但你可能没法为此多收钱。所以部分情况是,这些东西变成了竞争必需品,每个人都必须购买和使用。但你从中获得的成本节约或生产力提升就这样被竞争掉了。所以你没法为此多收费。我的意思是,如果你在麦肯锡、贝恩或波士顿咨询,以前做那个分析需要一周,现在只需要一天,你可能会做五倍的分析,向客户收同样的钱,你的成本基础也没有改变。这正是思考投资银行和财务分析时的方式。你只是用更少的人做了更多的分析,向客户收了同样的钱。
The challenge is it's a bit tricky at this early stage. It's quite hard to know what the ROI is. It's rather like giving everybody the internet in the late '90s and saying okay go off be more productive. And if you look at a survey from Deloitte, there's also a survey from the Fed that's in my presentation, where if you go and ask CFOs where have you seen the benefits, most of the benefits so far have been stuff that's pretty hard to measure, like better analytics, better customer support, more productivity. You can make more slides more quickly, you can do the analysis more quickly. It's kind of tough to put a financial value on that. It has a financial value, but it's not the same as saying, well, we made this new thing with AI and it had this revenue or it saved us this much money. Those things obviously take longer. It's harder to build a new revenue line than to give this to everybody and have them use it to make spreadsheets more quickly. So there's a little bit of like, well, how long does this take? I think the other answer to the problem here, of course, is consumer surplus, which is kind of what happened with Excel. You know, if a DCF takes you a week, then you probably only do one or two DCFs. And if a DCF takes you 10 seconds, then you do 50 DCFs, but you probably can't charge any more money for that. So some of what happens is that these things become competitive necessities and everybody has to buy it, use it. But the cost saving or the productivity gain that you get from it just gets competed away. So you don't get to charge more for it. I mean, if you're at McKinsey or Bain or BCG and doing that piece of analysis used to take a week and now it takes a day, you probably do five times more analysis and charge your customer the same, and your cost base hasn't changed either. Which is exactly the way to think about what happened with investment banks and financial analysis. You just did way more analysis with probably fewer people and charged customers the same amount of money.
你的论点中有一部分是模型最终会变成商品,然而历史上融资最快、最多的却是这些基础模型公司。那么,针对这一点,你对它们有什么建议吗?无论是集体建议还是针对某一家公司,以便它们适应?
Part of your thesis is this idea that models are going to end up as commodities, and yet the layer that's raising the most money in the fastest time in history is these foundation model companies. So given that, what advice might you have for them, either collectively or individually, in order to adapt?
我并不是说我知道它们会变成商品。我的立场更像是:这里有一连串的论证,表明从确定性角度看,这些东西看起来会是商品,请告诉我为什么它们不会。我就只承诺到这一步。我认为,这些公司融了这么多钱,我有点回到我之前关于移动行业的观点,这虽然不具有预测价值,但值得观察。移动行业非常庞大,在基础设施上投入了大量资金,但利润并不高,所有酷的东西都是别人做的。然后你问,资本回报率是多少?答案是,这取决于你在哪个市场,是在美国、欧洲、印度还是中国。但与此同时,这是一件值得做的事,它为某些人带来了回报,但它最终并没有控制整个行业,其他人从中获得了更多的价值。你知道,我脑子里没有具体数字。谷歌去年的净利润是多少?大概 500 亿美元吧。整个电信行业的净利润是多少?我真该订阅彭博终端,这样我就能立刻回答这些问题了。但可以肯定的是,谷歌、Meta、亚马逊、微软、苹果的利润总和超过了整个电信行业。所以这是一个谜题。你在推动前沿。你陷入了这样一个陷阱:你必须不断竞争,否则别人就会做,你就会落后。还有一件事我们完全没有谈到,那就是:我们不是在构建 AGI 吗?就像我们要把上帝装进盒子里,有些人确实相信这一点,尽管这很难分析。但也许吧。所以你会继续构建这些东西,但实际问题是:你如何做出人们想用的、不是软件开发的东西?我的意思是,软件开发是个好生意。但这是唯一的生意吗?假设让软件行业更高效值几千亿美元。很好,那可能值一万亿美元。但然后呢?你如何把这个扩展到经济的其他部分,扩展到其他所有人?这就是为什么会有这些关于与私募股权合作、与咨询公司合作的讨论。正如我们一直在讨论的,如果你真的在经营一家实体公司,实际上很难弄清楚怎么用这些东西。所以你会去找贝恩、波士顿咨询、麦肯锡,或者印孚瑟斯、高知特、IBM、埃森哲,或者私募股权公司。所以就有这样一种感觉:一方面,你在构建越来越大的模型,你觉得必须继续做下去;但另一方面,是的,但人们用它来做什么呢?
It's not that I know that they're going to become commodities. My position is more like, here is a chain of argument that says deterministically it looks like these things will be commodities, and explain to me why they won't be. That's as far as I would commit to that. I think the raising all this money, I kind of go back to my point about mobile, which again has no predictive value but it's a worthwhile observation. The mobile industry is very big and spends a lot of money on infrastructure and isn't very profitable, and all the cool stuff is done by somebody else. Then you ask, well, what's the return on capital? And the answer is, well, it depends which market, whether you're in America or Europe or India or China. But meanwhile, that was a worthwhile thing to do and it produced a return for somebody, but it didn't end up controlling the whole thing, and other people ended up getting more value from that than they did. You know, I don't have the number in my head. What is Google's net income last year? $50 billion or something. What was net income for the total telecoms industry? I should really subscribe to Bloomberg, then I could just answer these questions instantly. But it's a pretty safe bet that Google, Meta, Amazon, Microsoft, Apple produce more profits than the entire telecoms industry. So this is a puzzle. You're driving the frontier forward. You're kind of caught in this trap that you have to keep competing because otherwise they'll do it and you'll fall behind. You've also got this thing that we haven't talked about at all, which is, hey, aren't we just building AGI? Like we're going to build God in a box, which some people do believe, although it's kind of hard to analyze. But maybe. So you're going to carry on building this stuff, but the practical question is, how do you get things that people want to use that aren't software development? I mean, that's a good business. Is that the only business? Pick a number of how many hundreds of billions of dollars it is to make the software industry more productive. Great. That's worth a trillion dollars maybe. But then what? How do you expand this into the rest of the economy, into everybody else? Which is why you get these conversations about partnering with private equity, partnering with consultancies, where exactly as we've been discussing, guess what, it's actually quite hard to work out what to do with this stuff if you're actually running a real company. So you go to Bain, BCG, McKinsey, or Infosys and Cognizant and IBM and Accenture, or private equity shelters. So there is this sort of sense of, on the one hand, you're building these bigger and bigger models and you kind of feel like you've got to keep doing it, but on the other hand, yes, but what are people doing with it?
为什么大多数人看着 ChatGPT,却想不出今天能用它做什么?
Why do most people look at ChatGPT and not really think of anything to do with it today?
最后一个问题。你希望听众从你的演示中带走什么?我去年用过、今年又用了一次的东西,是我找到的一则 50 年代初的 IBM 广告,上面有一大群工程师举着计算尺,广告语是:“一台 IBM 电子计算器能给你 150 个额外的工程师。”这就像你在 A16Z 看到过多少张这样的宣传图。我们记得,每 10 年、15 年或 20 年,我们都会经历一波根本性的技术变革,它们都令人惊叹,改变一切,完全不同于以往。所以 AI 是令人惊叹的、变革性的,完全不同于以往。移动也是件大事,互联网也是,个人电脑也是,计算也是。那些也都是非常重大的事件,当时很难预测会发生什么。
Last question. Is there anything from the presentation that you want to make sure listeners leave with? The thing that I used last year and I used again is an IBM ad I found from the early '50s which has a picture of a sea of engineers all holding up slide rules, and it's an IBM ad that says, "An IBM electronic calculator gives you 150 extra engineers." That's like how many pictures have you seen at A16Z where that was the pitch. And we kind of remember that we go through these waves of fundamental technology changes every 10 or 15 or 20 years, and they're all amazing and change everything and are completely unlike anything that's happened before. And so AI is amazing and transformative and completely unlike anything that's happened before. Mobile was quite a big deal too, and so was the internet, and so were PCs, and so was computing. Those were all also very big deals where it was hard to tell what was going to happen.
所以我们应当默认一个基准情形:好吧,我们会再次经历这一切。你知道,这会产生一堆毁掉人们生活的东西,会让很多人失业。嗯,还会有一堆我们不太满意的东西。也会有一堆我们都认为很棒的东西。然后 20 年后,我们就会忘记曾经有一个世界,电脑做不了那些事。我是说,看看现在。我们通话已经一个小时了,电脑没崩溃,还在互相传输高清视频,就好像,嗯,当然能行。事实上,我还在用 iPhone 做这件事。我的 iPhone 通过 Wi-Fi 把视频流传输到我的 Mac 上,它就这么神奇地工作着,我们都不再注意它了。我认为这基本上就是我对于这一切最终会如何发展的概括。它会变得像魔法一样,20 年后我们只会说,嗯,当然就是这样。电脑一直都能做这些事。
And so we should sort of presume as a base case, okay, well, we're going to go through that again. And you know, that will produce a bunch of things that ruin people's lives and it will put a bunch of people out of work. And there'll be a bunch of stuff that we're not very happy about. And there'll be a bunch of stuff that we all think is great. And then in 20 years time, we'll kind of forget that there was a world when computers couldn't do that. I mean, here we are. We've been on this call for an hour and our computers didn't crash and we're streaming HD video to each other and it's like, well, of course that worked. In fact, I'm also doing it with my iPhone. So, my iPhone is streaming to my Mac over Wi-Fi streaming video here and it just works like it's magic and we don't notice it anymore. And I think that's really my kind of oneline description of how all of this is going to end up. It's going to be magic and in 20 years time we'll just say, well, of course that's how it is. Computers have always done that.
是的,这是一个很好的收尾。这个演讲叫做“AI 改变世界”,在 Benedict Evans 的网站上。非常精彩。还有很多我们没来得及聊的内容,所以一定要去看看。Benedict,这次对话非常棒。非常感谢你来做客播客。
Yeah, that's a great place to wrap. The presentation is called AI Essets the World. It is on Benedict Evans' website. It is excellent. There's a lot more that we didn't get to, so definitely go check it out. Benedict, this has been a great conversation. Thanks so much for coming to the podcast.
谢谢。聊得很开心。
Thanks. Great to chat.