Claude Code's Explosive Growth and the Future of AI Agents
打开互动全文版(中英对照 + 朗读 + 问答)→Boris Cherny 讨论 Claude Code 前所未有的增长、token 最大化以及产品轨迹是否可持续。
Boris Cherny discusses Claude Code's unprecedented growth, token maxing, and whether the product's trajectory is sustainable.
我从未见过如此深度的增长。然后它继续变得越来越指数级。Claude Code 100% 由 Claude Code 编写。Co-work 100% 由 Claude Code 编写。越来越多的功能完全由 Claude Code 编写,遍布 Anthropic 及其产品。所以我想听听你对 token maxing 的看法,以及你是否认为这构成了你们正在构建的产品使用量的很大一部分。
I've just never seen growth this deep. And then it just kept going more and more exponential. Claude Code is 100% written by Claude Code. Co-work is 100% written by Claude Code. An increasing number of features are fully written by Claude Code across Anthropic and products. So I want to hear your perspective on token maxing and whether you think that makes up a large portion of the usage of the products that you're building.
我不写代码。我提示 Claude。实际上现在,我主要做的是让一个 Claude 去提示其他 Claude。所以我甚至不和 Claude 对话。我有一个 Claude 在和我的其他 Claude 对话。
I don't write code. I prompt Claude. And actually nowadays, mostly what I'm doing is I have a Claude that prompts other Claudes. So I don't even talk to Claude. I have a Claude that's talking to my Claudes.
让我们与 Claude Code 负责人 Boris Cherny 聊聊该产品的爆炸性增长。路线图上的下一步是什么,以及这一切是否可持续?接下来就是这些内容。欢迎收听 Big Technology Podcast,一档关于科技世界及其他领域的冷静而细致对话的节目。今天我们为您准备了精彩的节目。Claude Code 负责人 Boris Cherny 来到我们的演播室。我们将全面讨论该产品、它的起飞方式、路线图上的下一步,当然还有它是否可持续。我们还将探讨 token maxing、token inefficiency 以及知识工作的未来。所以话题不少。Boris,很高兴见到你。欢迎来到节目。
Let's talk with Claude Code head Boris Cherny about the product's explosive growth. What's next on the roadmap and whether all this is sustainable? That's coming up right after this. Welcome to Big Technology Podcast, a show for coolheaded and nuanced conversation of the tech world and beyond. We have a great show for you today. Claude Code head Boris Cherny is here with us in studio. We're going to talk all about the product, the way it's taken off, what's next on the roadmap, and of course whether it's sustainable. Going to go into things like token maxing, token inefficiency, and then of course the future of knowledge work. So no lack of topics to cover. Boris, it's so great to see you. Welcome to the show.
谢谢邀请。
Yeah, thanks for having me.
那么,我们首先来谈谈 Claude Code 的增长。它非常巨大,对吧?我想在最近的一次活动中,Anthropic 的 CEO Dario Amodei 提到,对 Anthropic 产品的需求同比增长了约 80 倍。我记得去年这个时候和他交谈时,他对 Anthropic 达到 40 亿美元的年经常性收入感到兴奋。现在这看起来有点过时了。现在的数字说可能是 450 亿美元,对吧?所以,增长了 10 倍,需求增长了 80 倍。问题在于公司能以多快的速度满足这里的需求。但请谈谈 Claude Code 在这部分需求中所占的比例,以及你在需求增长和使用人数方面看到了什么。
So let's talk a little bit to begin with about the growth of Claude Code. It's been massive, right? I think at a recent event, Dario Amodei, the CEO of Anthropic, talked about how demand for Anthropic's products has been up like 80 times year-over-year. I remember speaking with him last year around this time and he was thrilled that Anthropic was at $4 billion ARR. That seems quaint right now. The numbers right now say maybe it's 45 billion, right? So, a 10x there, 80x demand. And the question is how fast the company can serve the demand here. But talk about the portion of that demand that Claude Code makes up and what you've seen in terms of demand growth and the amount of people using this thing.
对于世界上越来越多的人来说,我认为使用智能体和 AI 的方式,不仅仅是 Anthropic 的产品,尤其是 Claude Code。当然,对于 Anthropic 来说,有很多不同的产品。有 Claude Code、Claude Chat、Claude Design、Co-work,还有 API 产品。有很多方式可以体验 Anthropic。但对很多人来说,Claude Code 是他们的第一次接触。是的,增长简直疯狂。当我们首次内部发布时,它立即飙升。所以在我们向 Anthropic 之外的任何人发布 Claude Code 之前,我们就觉得它很可能会大获成功。大约在去年 5 月我们发布 Opus 4 和 Sonnet 4 的时候,增长变得指数级。我从未见过如此陡峭的增长。然后随着 Opus 4.5(去年 11 月)、4.6(今年 2 月)和 4.7 的发布,它一次又一次地加速。我们团队中有很多人长期在科技行业工作,参与过各种超高速增长的产品。这是科技界一直谈论的事情,比如独角兽和超高速增长,但即使在我们团队,也从未见过这样的增长。所以我们正在努力弄清楚如何让每个人都能继续体验这一点,如何让我们能够以这种速度继续增长,以及我们预期的未来可能比今天更陡峭的速度。我们正在学习如何做到这一点,以及如何继续扩展服务。
For an increasing number of people in the world, I think the way that you use agents and the way that you use AI, it's not just Anthropic products, but it's Claude Code in particular. And you know, of course, for Anthropic, there's a lot of different products. There's Claude Code, there's Claude Chat, there's Claude Design, there's Co-work, there's the API products. There's a lot of ways to experience Anthropic. But for a lot of people, Claude Code is their first introduction. And yeah, the growth has just been insane. It's, you know, when we first released it internally, it just skyrocketed immediately. And so before we even released Claude Code to anyone outside of Anthropic, we felt that it's pretty likely that this is going to be a hit. And around the time that we released Opus 4 and Sonnet 4, this was in May of last year, the growth just went exponential. And I've just never seen growth this steep. And then it just kept going more and more exponential with Opus 4.5 that was November and then 4.6 that was February of this year and then 4.7 it just keeps inflecting over and over. And you know, there's a lot of people on our team that have worked in tech for a long time and we worked on all sorts of hypergrowth products. This is something you talk about in tech all the time, these like unicorns and hypergrowth, but even on the team we've never seen growth like this. And so we're just trying to figure out how do we make it so everyone can continue to experience this, how do we make it so we can continue growing at this pace and the pace that we expect in the future which might be even steeper than it is today. And we're learning a lot about how to do this and how to keep scaling the services.
所以一年前,很明显 Anthropic 的 AI 模型的大部分使用是通过 API 进行的,对吧?比如像咨询公司这样的公司将其应用于银行,银行用它来总结一些计算。我只是举个例子。与 Claude 聊天机器人相比,API 在用量、收入等各方面都占了绝大部分。今天还是这样吗,还是 Claude Code 正在超越它?
So a year ago it was clear that the bulk of usage of Anthropic's AI models was happening through the API, right? That would be like a company like a consulting group for instance putting it into action at a bank and the bank using it to summarize some calculations. I'm just throwing an example out there. Compared to the Claude chatbot, it was far and away the API was the lion's share of usage, revenue, all these things. Does that still the case today or is Claude Code overtaking that?
我们有一个混合体。所以,产品对 Anthropic 来说比一年前重要得多。这确实是事实。产品增长正在加速,增长非常快。API 也在加速,增长非常快。对我们来说,我们正在投资两者。我们必须成为一家产品公司,因为实验室构建产品有很多原因。实际上,这在早期并不明确。在 Anthropic 历史的早期,在我加入之前,这实际上是一个积极的辩论:我们是否应该构建产品?这真的有用吗?结果证明这非常有用。对于思想份额,但也为了安全。从根本上说,我们存在是为了研究 AI 安全,这给了我们更好的工具来做这件事。我们也是少数人,所以世界上大多数东西我们不会构建。
We have a mix. So you know, products play a much bigger role for Anthropic than they did a year ago. That's definitely the case. Product growth is accelerating, it's growing very quickly. API is also accelerating and growing very quickly. And for us, we are investing in both. We have to be a product company because there's a lot of reasons for a lab to build products. And you know, this actually wasn't clear early on. Very early on in Anthropic's history, this is before I joined, this was actually an active debate: should we even build products? Is this actually a useful thing to do? And it turns out it's very useful. For mind share but then also for safety. Fundamentally we exist to study AI safety, this gives us better tools to do that. We're also a small number of people, and so most things in the world we will not build.
对吧?
Right?
所以这也是为什么我们还必须提供一个平台,我们有托管智能体、API 和 SDK 等所有这些产品,以便人们可以在其上构建。你知道,成千上万的企业选择这样做。
And so this is why we also have to provide a platform and we have managed agents and API and SDK all these products so people can build on top. And you know, thousands and thousands of businesses choose to do that.
是的,听到你甚至回答说是混合体,这很有趣。所以我想你现在不会分享哪个更大。
Yeah, it's interesting to hear you even answer the question saying that it's a mix. So I take it you're not going to share which is bigger right now.
也许现在不会。
Maybe not right now.
好的。但事实是,API 更大并不是那么明确。也许它是。但你甚至说它是混合体,这本身就表明 Anthropic 自有和运营的产品正在大规模增长。现在我们设定了背景,这是一个指数级增长的东西。我们显然看到 Anthropic 的收入随着这个产品呈指数级增长。这是你构思、构建并今天运营的产品。我想可能有些观众会想,嗯,Claude Code 是什么?我们的大多数观众显然知道它是什么。我当时想,我如何用一个简单的句子定义它?我写道,它是一种用简单英语构建网站和软件的方式。然后在我来这里的路上,我想,嗯,这有点低估了它。
Okay. But the fact that it's not a clearcut the API is bigger. Maybe it is. But the fact that you even say it's a mix just shows that Anthropic's owned and operated products are just growing massively. And now we've set the stage here that this is something that's growing exponentially. We've obviously seen the Anthropic revenue grow exponentially kind of alongside this product. This is a product that you conceived of and built and run today. I think that there's probably some people watching who are like, well, what is Claude Code? Most of our viewers obviously know what it is. And I was like, how do I write this in a simple one-sentence definition? And I wrote that it's a way to build websites and software in plain English. And then on the way over here, I was like, well, that kind of sells it short a little bit.
我的意思是,你会怎么描述它?
I mean, what would you describe it as?
我认为这实际上是一个相当不错的描述。
I think that's actually a pretty good description.
好吧。我们就用这个。
All right. We'll take it.
我认为当很多人想到 AI 时,他们会想到聊天机器人。
I think when a lot of people think about AI, they think about chatbots.
你知道,对工程师来说,这就是 AI 的样子,大概一年半前,在我们开始做 Claude Code 之前,对大多数人来说 AI 就是那样。后来我们意识到模型在编程和使用工具方面变得非常出色。这些其实一直是我们训练模型去做的事情。这已经是研究方向有一阵子了。大约一年半前,它开始变得有商业价值。所以对于 Claude Code,我们下了这个赌注,偏离了当时所有人写代码的方式。世界上所有人写代码用的都是某种花哨的文本编辑器,我们只是觉得也许可以做得更好,做一些真正不同的事情。这确实是个赌注。所以我们推出了 Claude Code。Claude Code 与当时聊天机器人的不同之处在于,它能使用工具。就这么简单。这就是区别。用聊天机器人,你来回对话,但一个智能体——而 Claude Code 就是一个智能体——它能使用你的工具。
And you know, for engineers, that's what AI was, maybe like a year and a half ago, before we started Claude Code. That's what AI was for most people. And we realized at some point that the model was actually getting really good at coding and using tools. These are things we've kind of always trained the model to do. This has been the research direction for a while. It started to become commercially useful about a year and a half ago. So for Claude Code, we took this bet and deviated from the way everyone wrote code at the time. Everyone in the world wrote code using essentially a fancy text editor, and we just thought maybe we can do much better than this and do something really different. It was very much a bet. So we introduced Claude Code. The thing that made Claude Code different from chatbots at the time was that Claude Code can use tools. That's it. That's the difference. With a chatbot, you're going back and forth talking, but an agent—and Claude Code is an agent—can use your tools.
对。我们能快速定义一下工具吗?工具可以是任何东西——如果我说错了你纠正我——从使用浏览器到登录 Cloudflare,然后以那种方式设置某个智能体,对吧?所以它不再关注这个产品本身能做什么,而是这个产品能登录什么,然后利用你能在线使用的多种产品来做什么。
Right. And can we just quickly define the tools? So tools could be anything—and you tell me if I'm wrong—from using a browser to logging into Cloudflare and then setting up some agent that way, right? So it becomes less about what this product does itself and more about what this product can log into and then do with a multiplicity of products you can use online.
没错。它能连接你所有的不同工具。它能使用你的浏览器。它能使用你的电脑。甚至像编辑电脑上的文件这样简单的事情。一年半前,还没有 AI 产品能真正做到这一点。但这是 Claude Code 能做的第一件事。它可以编辑你桌面上的文件。如果你桌面上有一堆文件,它可以整理它们。所以 Claude Code 和 Claude Work 拥有这种访问权限,只要你选择授予它。
That's right. It can connect all your different tools. It can use your browser. It can use your computer. Even something as simple as editing a file on your computer. A year and a half ago, there was no AI product that could actually do that. But this is the first thing that Claude Code was able to do. It could edit a file on your desktop. If you have a bunch of files on your desktop, it can organize them. So Claude Code and Claude Work have this access, if you choose to give it.
当然。
Granted.
是的。它能做到这一点。这很神奇。这个微小的差异完全改变了人们使用这个产品的方式,也彻底改变了这个产品能为你做什么。
Yeah. And it can do this. This is magical. This tiny difference completely changes the way people can use this product and totally changes what this product can do for you.
是的。我的意思是,根本问题,我想深入探讨的是,AI 似乎已经从擅长自动补全转变了。在底层,AI 只是预测下一个是什么——如果你使用机器学习并将其应用于大数据集,预测你是否可能拖欠抵押贷款以及银行是否应该批准抵押贷款。对于句子,预测下一个词;对于代码,预测序列中的下一段代码。所以我认为那是第一代。但你现在谈论的是,机器实际上能够在你给出自然语言提示后,自己编码,接入工具,然后为你做事。所以如果我说错了请纠正我,但这里的用例已经从开发者接入它并用 Claude Code 编写代码——我们看到这种爆发主要由他们驱动——但随后是第二股力量:非技术人员,像我这样的人,可以通过指挥 AI 智能体(即 Claude Code)来构建软件,为他们构建工作流软件或网站,或者通过类似 Claude Work(可以说是更简单的姊妹产品)来控制你的电脑,然后说:“你有权限访问我的浏览器。你知道我喜欢订哪种航班。我几周后要去印度。订机票吧。”
Yeah. I mean the fundamental thing, I think just to drill down here, is that it seems like AI has shifted from being great at autocomplete. At the fundamental layer, AI is just predicting what comes next—predicting, if you're using machine learning and applying it on a large dataset, predicting whether you might default on your mortgage and whether a bank should grant a mortgage. When it comes to a sentence, predicting the next word; with code, predicting the next bit of code in a sequence. So I think that was Gen 1. But what you're talking about now is the machine is actually able to go, after you give it this natural language prompt, code itself, hook into tools, and then do things for you. So correct me if I'm wrong, but the use cases here have gone from developers hooking into it and writing code with Claude Code—and we've seen this explosion largely driven by them—but then by a secondary force: non-technical folks, people like me, who can build software by directing the AI agent, which is Claude Code, to build a piece of workflow software for them or a website, or to take control of your computer via something like Claude Work, which is sort of the easier sister product, and saying, "Well, you have access to my browser. You know what type of flights I like to book. I need to be in India in a couple of weeks. Book the flight."
是的,完全正确。我实际上刚刚用 Claude Work 订了一堆航班。这个月我要飞很多次,你知道,我们即将在伦敦和东京举办 Code with Claude 活动,沿途还有其他几个站点。我和 Claude Work 来回沟通。我说:“好的,我需要在这些时间到达这些地方。”一共五个站点,很多城市。这是大致的行程。查看我的邮件,查看我的日历,再核对一下,确保我没有遗漏什么。它实际上发现了我遗漏的两个站点,还有我告诉它的几个日期是错的。它在我要求之后通过查看我的邮件发现了这些。然后我让它订机票,我就去写代码做别的工作了。一小时后我回来,它已经订了八张机票和五家酒店。其中一家酒店不太对——位置错了。我让它重新预订并更改,它就搞定了。就是这样。我实际上每次都用 Claude Work 和 Claude Code 做这种测试。我有一些这样的测试用例,一些我常做的事情,随着模型改进,我会用不同模型重新尝试。这是我得到过的最好结果。Claude Work 结合 Opus 4.7 能做到这一点。我认为对我来说最难的事情之一是,随着模型改进,你不得不不断调整对它能力的期望。如果你和人们交谈,特别是那些一年前用过这个模型但之后再没用过的工程师,他们可能会说:“哦,它在编程方面不是很好。我不相信它能一次写超过几行代码。”因为那是一年前模型的样子。它当时还不够好。
Yeah. Exactly. I actually just used Claude Work to book a bunch of flights. I'm going to be flying a bunch this month for, you know, we have Code with Claude coming up in London and Tokyo, and there are some other stops along the way. I went back and forth with Claude Work. I said, "Okay, I need to be in these places at this time." It was five stops, a lot of cities. Here's roughly the schedule. Look through my email, look through my calendar, and just double-check it, make sure I'm not missing anything. It found actually two stops that I was missing and also a couple dates that I told it wrong. It just found this by looking at my email after I asked it to do that. Then I told it to book the flights, and I went and was coding on something else, just doing work. I came back an hour later, and it booked eight flights and five hotels. One of the hotels was kind of incorrect—it was in the wrong area. I asked it to rebook and change it, and it was done. That was it. I actually try this every time with Claude Work and Claude Code. I have these sort of test cases, common things I would do, and I just retry them with different models as the model improves. This is the best result I've ever gotten. There's something about Claude Work combined with Opus 4.7 that enables this. I think one of the hardest things for me has been that as the model improves, you constantly have to readjust your expectations of what it can do. If you talk to people, especially engineers who used the model a year ago and haven't used it since, they might say, "Oh, it's not very good at coding. I don't trust it to write more than a few lines at a time." Because that's what the model was a year ago. It wasn't very good yet.
如果快进到今天,你让这些人坐下来试用新模型,就像越来越多工程师一直在做的那样,这完全是一种不同的体验。能力完全不同。我认为这是我用过的第一种这样的技术,每个月它的能力都会发生阶跃变化。作为这项技术的用户,这相当困难,因为你必须不断重新训练,不断重新尝试。你总是需要保持初学者心态,去重新尝试这项技术,并用它来做以前它不擅长的事情,因为下一个模型可能就能完美地完成它。
And if you fast forward to today and you sit down these people and they try the new model, as a lot of people have been doing, an increasing number of engineers, it's just a completely different experience. The capability is completely different. I think this is the first technology I've used like this where every month there's a step change in what it can do. As a user of this technology, it's just quite hard because you have to keep retraining, you have to keep retrying. You always need this beginner mindset to retry the technology and use it for a thing it was not good at before, because the next model might just do it perfectly right.
所以我认为这就是你描绘的愿景。实际上,以前当你使用技术时,你受制于界面。你有一个为规模化而构建的软件公司,但你会得到很多可能不适用于你的功能。每次你想预订东西时,即使你知道自己想要什么,你也得经历所有这些花哨的步骤,而且没有一个网站会知道你的偏好。现在它改变了范式,它又变成了一个智能体。它是一个为你出去做事的东西,并且可以按照你想要的方式塑造你的在线体验。我认为这就是人们正在抓住的东西,这就是为什么我们看到爆炸式增长。
So I think this is the vision the way that you're outlining it. Effectively, previously when you would use technology, you would be subject to the interface. You would have a software company that built for scale, but you would get a lot of features that maybe weren't applicable to you. You would have to go through all these bells and whistles whenever you were trying to book something, even though you knew what you wanted, and you wouldn't have a website that would know your preferences. Now it sort of shifts the paradigm where you have again it's an agent. It's something that goes out and does things for you and can potentially shape your experience online the way that you want it. And that is I think what people are seizing upon, and that's why we're seeing the explosive growth.
但现在我想稍微压力测试一下这个论点,并提出一些让我好奇的事情。这其中有多少是真实的,又有多少只是对潜力的狂热,但也许我们应该现实地审视一下。首先,需求如此之大。但问题是,这些需求中有多少是纯粹的需求,又有多少是被游戏化的需求。在硅谷内外,有一种做法叫做 token maxing。我相信你听说过。公司有指令,要求人们通过尽可能多地运行他们的 AI 智能体来使用大量 AI token。然后那些使用最多 token 的人会在排行榜上获得奖励,或者达到他们必须采取的 AI 行动目标,而不是物理行动。所以我想听听你对 token maxing 的看法,以及你是否认为这构成了你们正在构建的产品使用量的很大一部分。
But now I want to pressure test the thesis a little bit and bring up some things that make me curious. How much of this is real and how much of this is just unbridled enthusiasm at the potential, but maybe stuff we should have a reality check on. The first thing is that there is such great demand. But the question is how much of that demand is pure demand versus demand that's gamified. There is a practice that's going on within Silicon Valley and outside of it that's called token maxing. I'm sure you've heard of it. It's where companies have a mandate where people are supposed to use lots of AI tokens by running their AI agents as much as they can. And then those who use the most tokens are rewarded on a leaderboard or meet a goal of AI actions that they have to take, as opposed to physical actions. So I want to hear your perspective on token maxing and whether you think that makes up a large portion of the usage of the products that you're building.
是的,我不认为 token maxing 占了很大比例。我的思考方式是,在 Anthropic 之前,我实际上曾在 Facebook 这样的大型科技公司工作。我在 Facebook,它正是那些进行 token maxing 的公司之一。没错。我的职责之一是维护 Meta 所有应用代码的健康状况,比如 Facebook、Instagram、WhatsApp。我们关心代码健康的原因之一是,如果代码质量很高,工程师的生产力就会更高。有一个大团队致力于生产力提升。在模型出现之前,在 Claude 出现之前,你需要工作很长时间,每年每个工程师的生产力可能只能提升 1%、2% 或 3% 左右。那已经是一个相当大的提升了,而且来之不易。你基本上必须尝试很多想法,最终才能找到像这样提升生产力的东西。而 Claude 带来的变化是,现在许多公司,包括 Anthropic 和我们所有最大的客户,都报告了数百个百分点的收益。我记得我们上次报告的数字是,自从我们推出 Claude Code 以来,Anthropic 每位工程师编写的代码量增长了大约 250%,同时代码质量和可靠性等方面保持稳定。所以,在这些方面没有退步的情况下,代码量大幅增长。我认为这种生产力影响非常新颖,人们正在试图弄清楚如何获得它。很多公司在问如何获得这些好处,因为很多公司已经看到了,而有些还在摸索。我的建议几乎总是一样的。第一,给每个人 token,让他们去实验。我不一定推荐 token maxing,但我建议让他们实验,这样他们就不必为每个 token 请求批准。第二,给人们心理安全感,因为很多时候,当人们创新并构建提高生产力的工具时,他们正在改变自己的工作流程以提高生产力。他们尝试一堆想法,有些可能不奏效,有些则有效。所以你想给人们这种心理安全感,让他们觉得可以放心实验并找到这些新流程。然后,很多公司发现,生产力提升和创新并非来自你预期的人。在过去,每个人都能指出哪些是他们最高效的工程师。但我认为现在很多改进来自你完全意想不到的人。可能是你组织角落里某个会计,以工程师从未想过的方式自动化了会计工作。可能是某个营销人员,以你从未想过的方式自动化了营销。也可能是一个刚毕业的软件工程师,构建了一些了不起的东西。这在以前是不会发生的。挑战在于,你无法提前识别这些工程师和这些人。你不知道他们是谁,而且他们几乎总是会让你惊讶。所以你要做的就是让人们实验,给他们安全感,然后一旦有某种用例规模化,那时你才考虑优化它。但你不想提前优化。所以,我不知道以竞争的方式做这件事是否适合某些公司的文化,如果是,那很好。如果其他公司想要的方式只是创造安全感和为工程师创造实验空间,就像我们在 Anthropic 所做的那样,那也很好。这真的取决于公司。
Yeah, I don't think token maxing is a large percent. The way that I would think about it is, before Anthropic actually, I used to work at a big tech company at Facebook. I was at Facebook, which is one of the companies that's token maxing for. That's right. And one of my responsibilities was the health of all of the code across the Meta apps. So this is like Facebook, Instagram, WhatsApp. And one of the reasons that we care about the health of the code, and this is essentially things like code quality, is if the code is really high quality, engineers are more productive. There's a big team of people that worked on productivity. And before models, before Claude, you would work for a really long time and you would see maybe a 1, 2, 3% improvement in productivity per engineer over the course of a year, something like that. And that was a pretty big improvement. It was very hard won. You essentially had to try a lot of ideas and eventually you find something that improves productivity like this. And what happened with Claude is now many companies, including Anthropic and all of our biggest customers, are reporting gains on the order of hundreds of percentage points. I think the last number that we reported is the amount of code written per engineer at Anthropic has grown something like 250% since we introduced Claude Code, and this is while keeping code quality and reliability and all these things kind of stable. So without those things regressing, the volume of code has grown a lot. So this kind of productivity impact I think is just very new, and I think people are trying to figure out how to get this. There are a lot of companies asking how do we get these kinds of benefits, because a lot of companies are seeing it and then some are still figuring it out. My advice is almost always the same. The first thing is just give everyone tokens, let people experiment. I wouldn't necessarily recommend token maxing, but I would recommend let people experiment so they don't have to ask for approval for every token. The second thing is give people psychological safety, because a lot of times when people are innovating and they're building tools that make them more productive, they're changing their own workflows to make them more productive. They try a bunch of ideas, some of them might not work and then some of them work. So you want to give people this kind of psychological safety so they feel okay experimenting with it and finding these new processes. And then the thing that a lot of companies see is the productivity improvements and the innovations do not come from the people you expect. Back in the old days, everyone could point out like these are my most productive engineers. But I think nowadays a lot of the improvements are coming from people you just never would expect. It could be an accountant somewhere in the corner of your org that just automates accounting in a way that no engineer would have thought of. It could be some marketer automating marketing in a way that you never would have thought of. It could have been a new grad software engineer that just built something amazing. And this is something that just didn't happen before. The challenge is you can't identify these engineers and these people ahead of time. You don't know who they are and it's almost always going to surprise you. So the thing you want to do is let people experiment, give them safety, and then once there's some kind of use case that scales up, that's when you think about optimizing it. But you don't want to optimize ahead of time. So I don't know if doing it in a competitive way works for some companies with their culture, then I think that's great. If for other companies the way they want to do it is just kind of create safety and create space for engineers to experiment, which is what we do at Anthropic, then I think that's great too. It really depends on the company.
是的。我要说,我用了很多 token。我一直在使用这些工具。我认为 Claude Code 和 Claude Co-Work 对我的业务都非常有用。
Yeah. And I'll say look, I use a lot of tokens. I'm in the tools all the time. I think Claude Code and Claude Co-Work have both been pretty great for my business.
我是一个独立运营者,虽然这么说有点低估了,因为我背后有一个团队,主要是兼职帮我,但那是另一个节目的事了。但我确实在想,当我读到这些故事时,大公司占了这些预算的很大一部分,还有激励措施,就像我在节目开始时说的,这能持续多久?有些地方的激励措施很糟糕。这是最近《金融时报》的一篇文章。亚马逊员工使用 AI 工具做不必要的任务来夸大使用分数。一些员工说,同事们用软件自动化额外的、不必要的 AI 活动来增加他们的 token 消耗量。他们说,此举反映了在亚马逊为超过 80%的开发者设定了每周使用 AI 的目标后,采用该技术的压力。我和一位亚马逊员工核实过,他们说:“是的,就是这样。”他们告诉我:“我触发了一个自动化程序,每天运行几个小时然后被删除,就是为了达到这些目标。”所以你说你不认为这种 token 最大化是需求的主要部分。你那边有没有什么迹象表明这只是个例,而不是大多数地方的常态?
I'm a solo operator, although that kind of sells it short because I have a team of people behind me that help me mostly on a part-time basis, but that's for a different show. But I do wonder, you know, when I read these stories, the large corporations are largely making up big percentages of these budgets and the incentives, you know, and again like I started the show saying how sustainable is this? The incentives are bad in some of these places. This is from the Financial Times recently. Amazon staff use AI tool for unnecessary tasks to inflate usage scores. Some employees said colleagues were using the software to automate additional unnecessary AI activity to increase their consumption of tokens. They said the move reflected pressure to adopt the technology after Amazon introduced targets for more than 80% of developers to use AI each week. I gut-checked this with an Amazon employee. They're like, "Yep, this is what's happening." They told me, "I triggered an automation that runs for hours and then gets deleted every day in order to meet these targets." So you said you don't think that this token maxing stuff is a big part of demand. Is there anything that you can see on your end to indicate that it's not that this is an outlier and not the rule in most places?
是的,这……我不知道有多少公司在做这种 token 最大化的事情。我听说过一点这个趋势。如果你看看 Claude Codes 的客户,我们有很多很多客户。所以并不是某一家公司在推动使用量。不是这样的。我确实想退一步想想这种变化是如何发生的。因为我认为这些公司试图达到的目标,我不想替他们说话,我建议直接和他们谈。但他们的目标,我认为,很可能是组织变革和业务流程变革。如何让你的公司从 AI 中受益?这通常不明确。这非常依赖于公司,因为每家公司都有不同的业务、文化、组织结构和做事方式。有一篇 90 年代的《哈佛商业评论》文章,我很喜欢,我忘了标题,但大概是“计算机来了,为什么没人看到生产力影响?”这是一个大问题,对吧?对我们来说,计算机显然提高了生产力。今天这非常明显。但在 90 年代,这并不明显。当时个人电脑正在被采用。它们正在取代大型机,而且现在价格实惠。所以普通公司、普通初创公司都能买一台。你不再需要花几百万美元买大型机了。但有一个挑战和悖论。公司采用了电脑,但没有看到生产力提升。怎么回事?所以那篇《哈佛商业评论》文章提出,为了从计算机中获益,你必须围绕计算机重组整个业务流程。它们必须成为你做事方式的核心。如果你还有纸质文件柜,抽屉里塞满了东西,仍然是纸笔的物理流程,而计算机只是外围设备,那你真的不会受益。但如果你扔掉文件柜,扔掉装满文件的抽屉,把计算机放在中心,用它来做所有业务流程,那你就会受益。公司之间出现了分化。有些公司这样做了,经历了相当痛苦的变革,并从中受益,而其他公司则没有。我认为现在也是类似的情况。很多公司都在试图弄清楚如何从 AI 的生产力影响中受益,有很多实验,每个人都在尝试不同的方法来找到受益的方式。我不认为有一种正确的方法。
Yeah, this is... I don't know how many companies are doing this token maxing thing. I've heard of it as a trend a little bit. If you look at Claude Codes' customers, we have many, many customers. So it's not like there's one company driving the usage. It's not like that. I do want to step back a little bit and think about how this kind of change happens. Because I think the goal of what these companies are trying to do, I don't want to speak for them, and I would recommend just talking to them. But the goal of what they're trying to do, I think, is probably organizational change and business process change. How do you make it so your company benefits from AI? And this is often unclear. It's very dependent on the company because every company has a different business, a different culture, a different org, a different way of doing things. There's this old Harvard Business Review article from the '90s, which I just love, and I forget the title, but it was something like "Computers are here, why is no one seeing the productivity impact?" And this was a big question, right? It's like to us it's obvious computers make us more productive. This is just incredibly obvious today. But in that '90s, this was not obvious. And what was happening is personal computers were being adopted. They were replacing mainframes and now they're affordable. So the average company, the average startup can buy one. You don't have to spend millions of dollars on a mainframe anymore. But there was this challenge and there was this paradox. Companies were adopting it, but they were not seeing productivity improvement. What's going on? And so this Harvard Business Review article made the case that in order to get a benefit from computers, you have to restructure your whole business process around computers. They have to be at the center of the way that you do things. And if you still have paper filing cabinets and you have a bunch of drawers full of stuff and it's still a paper and pen kind of physical process and there's a computer somewhere on the periphery, you're really not going to benefit. But if you throw away your filing cabinets, you throw away your desk drawers full of papers and you put a computer at the center of it and that's the way that you do all your business process, then you benefit. And there was this split between companies. Some were doing this and they were doing this fairly painful change and they benefited from it and then others didn't. And I think it's kind of the same thing now. A lot of companies are trying to figure out how to benefit from the productivity impacts of AI and there's just a lot of experimentation and everyone is trying different approaches to figure out how to benefit from it. I don't think there's one right approach.
好的。你看,当我们看到像 Claude Code 和 Anthropic 增长得这么快的时候,聊聊这些事挺好的,听听你的观点。所以,这就是 token 最大化。现在,token 当然是模型的输出,比如模型输出的单词或单词的一部分,以及输入模型的单词或单词的一部分,对吧?这就是这些公司收费的方式,你用得越多,就需要越多的数据中心,等等。你知道,随着模型变得更好,它们还没有……好吧,我这么说吧:有时我怀疑它们是否尽可能高效。这些大模型有时会做很多工作,使用很多 token,即使输出很棒。人们想知道,这是否只是在推高 token 需求,而本来可以是一个简单的过程,模型却消耗了大量 token,没有高效地完成。我给你举个例子。我一直在用 Claude Code 做 PPT。它很擅长这个。我用的是 Opus 4.7 模型,有几次我说:“好了,你在做这个,把它导出为 PDF。”然后它就开始失控了。它循环使用尽可能多的工具,似乎无法导出 PDF。最后我一直告诉它:“不,你在做这个 PPT,它在哪里?导出。”然后它说:“我欠你一个道歉。我钻了牛角尖,担心一个实际上并不阻碍我们的限制。文件在那里。”然后它导出了。谈谈这些模型的效率,以及这是否是一个合理的担忧,因为随着增长,部分原因就是像 Opus 4.7 这样的模型在执行基本任务时可能会陷入这种循环。
Okay. And look, I think that when we see something grow as fast as Claude Code has grown and as fast as Anthropic has grown, it's good to just talk this stuff through and it's good to hear your perspective. So okay, that's token maxing. Now tokens of course are the output of the model, like the words or portions of words that the model outputs and the words and portions of words that go into it, right? And that is how these companies charge and the more you have the more data centers you need, etc. You know, as these models get better, they haven't... Well, let me put it to you this way: sometimes I wonder whether they're as efficient as they can be. These big models can sometimes do a lot of work, use a lot of tokens even if the output is great. People wonder, well, is this just driving up token demand where it could have been a really easy process and the models are expending many many tokens and not getting there as efficiently as they could. Let me give you an example. I've been using Claude Code to make PowerPoint presentations. It's really good at it. And I've been using the Opus 4.7 model and a couple of times I've said, "All right, you're working on this, ship it as a PDF," and it just starts to lose its mind. It cycles and it uses as many tools as it possibly can and it seems unable to ship the PDF. Eventually I kept telling it, "No, you're making this PowerPoint, where is it? Ship it." And it goes, "I owe you an apology. I went down a rabbit hole worrying about a constraint that wasn't actually blocking us. The file's there." And then it shipped it. Talk a little bit about the efficiency of these models and whether that is a legitimate worry that as we've seen the growth, part of it is these loops that a model like Opus 4.7 might find itself in to do basic tasks.
是的,通常当我们考虑模型时,有几个不同的方面。
Yeah, generally when we think about models there's a few different aspects of it.
一个是它的智能程度,另一个是它的速度,还有一个是它的效率。我们通常试图让这些方面一起提升。在这三者中,我认为我们应该优先优化智能——这是最重要的。所以即使它效率稍低,但更智能,能让你做更多事情,那也很有用,因为效率优化是之后的事。我们先提升智能,然后再提升效率。所以基本上是一个接一个地做。
One is just how intelligent it is, another one is how fast it is, and another one is how efficient it is. And we generally try to move all these together. Between these, I think we should probably optimize for intelligence—that's the most important thing. So even if it's a little bit less efficient but it's more intelligent and lets you do more things, that's really useful because the efficiency optimization comes after. After we make it more intelligent, then we can make it more efficient. So it's sort of we do one then we do the other.
我们一直在试验如何让用户对此有控制权,因为我们并不总是知道正确的默认设置。有时候你使用时更清楚。我们提供的一种机制是选择模型,你可以选 Opus、Sonnet 或 Haiku。另一种我们正在试验的机制是努力程度。
We've been experimenting with how exactly we give people control over this, because we don't always know the right default. Sometimes when you're using it, you know better. One mechanism that we had for this is picking a model, so you can pick Opus or Sonnet or Haiku. Another mechanism that we've been experimenting with is effort.
Opus 最大,Sonnet 居中,Haiku 最小。
Opus is the biggest, Sonnet middle, Haiku smallest.
没错。这其实就是模型的大小。
That's right. And this is just the size of the model.
对。然后还有努力程度。努力程度本质上就是你愿意投入多少精力。你可以设置这个;我们有一个推荐的努力程度。例如,为了最大化 Opus 4.7 的智能,你会想用超高或最大努力。但如果你想用更少的 token,你可以选中等或低努力。这是你拥有的一个控制选项。
Right. And then there's effort. Effort is essentially how much effort do you want to put into it. You can set this; we have a recommended effort. For example, to maximize intelligence for Opus 4.7, you want to use extra high or maximum effort. But if you wanted to use fewer tokens, you can pick medium or low effort. This is a control that you have.
是的,我最近在一个节目上谈到这个,有个评论者插话。我原本认为更大的模型会想办法在 PDF 导出这类事情上变得更高效。有个评论者写道:‘Alex,他们无法解决那个 PDF 问题。这是 LM 技术固有的,也是智能体式 AI 有用且广泛传播和使用的最大障碍。’我想他们想说的是,我们之前讨论过预测——这一切都是概率性的,预测下一个词。你从 AI 智能体那里不会得到两次相同的答案。因此这类事情是他们工作方式的一个特性,无法修复。你怎么看?
Yeah, I talked about this on a show recently and we had a commenter that came in. I was of the opinion that bigger models will find a way to become more efficient on things like the PDF export. We had a commenter that wrote, 'Alex, they can't fix things like that PDF problem. It's inherent to LM technology and it's the biggest barrier to useful widespread dissemination and usage of agentic AI.' I think what they were trying to say is we talked about predictions earlier—that this is all probabilistic, predicting the next word. You don't get the same answer from an AI agent twice. And therefore this type of thing is a feature of the way that they work and not fixable. What do you think?
不,我不认为那是正确的。当你思考时,让我们稍微拉远一点。工程师是第一批采用者。工程师大约一年半前就开始使用 Claude Code,那时非工程师还没有有意义地使用智能体,还没有 co-work 等等。如果我回想一年半前的 Claude Code,它并不太好。我可以用它写一点代码,但如果我真的信任它来构建整个功能或整个产品,结果不会很好。它会陷入循环,质量不好,或者它构建了但代码很糟糕或无法运行。但到了某个时候,它开始变好了。随着模型改进,Claude Code 改进,结果越来越好。所以快进到今天,Claude Code 100% 由 Claude Code 编写。Co-work 100% 由 Claude Code 编写。Anthropic 及其产品中越来越多的功能完全由 Claude Code 编写。
No, I don't think that's right. When you think about it, let's zoom out a little bit. Engineers are the first adopters. Engineers started using Claude Code like a year and a half ago, before non-engineers were using agents in a meaningful way, before co-work and so on. If I think back to what Claude Code was a year and a half ago, it wasn't very good. I could use it to write a little bit of code, but if I really trusted it to build an entire feature or entire product, it wouldn't turn out well. It would go in spirals and the quality wasn't good, or it built it and either the code was bad or it didn't work. And at some point, it just started to get better. As the model improved and as Claude Code improved, the result just got better and better. So you fast forward to today, Claude Code is 100% written by Claude Code. Co-work is 100% written by Claude Code. An increasing number of features are fully written by Claude Code across Anthropic and its products.
这也是我们从客户那里听到的。
And this is something that we hear from customers also.
我昨天在 Y Combinator(那个创业孵化器)做了一个演讲。我让大家举手。每个人都在用 Claude Code。我问他们:‘如果今天你 100% 的代码都是用 Claude Code 写的,请举手。’大约一半的人举手了。然后我问:‘如果 0% 的代码是用 AI 写的,请举手。’大概只有一个人举手,房间里大约有一百人。那个人真有勇气。显然这还有空间。其他人都介于中间——他们大部分代码是用 Claude Code 写的,但不是全部。但这大致就是今天模型所处的位置。一年前它还不是这样;一年前它还不够好。所以这正是你现在在 co-work 上看到的情况。它仍然早期。我们发布它才几个月?它会不断改进,随着产品变好、模型变好,它会越来越好。但这只是早期阶段。我认为今天使用 co-work 的每个人都是早期采用者。甚至今天使用 AI 的每个人都是早期采用者。世界上有那么多人,大多数人还没有真正尝试过 AI。所以改进的空间还很大。
I did a talk at Y Combinator, the startup incubator, yesterday. I asked people to raise their hands. Everyone's using Claude Code. I asked them, 'Raise your hand if 100% of your code is written using Claude Code today.' About half the hands went up. Then I asked, 'Raise your hand if 0% of your code is written with AI.' There was like one hand that went up, in a room of about a hundred people. Power to that person. And there's still room for this obviously. Everyone else was somewhere in the middle—most of their code is written with Claude Code but not all of it. But that's kind of the place where the model was at today. It was not there a year ago; a year ago it was not good enough for this. And so this is exactly what you're seeing play out with co-work right now. It's still early. We released it what, like a few months ago? It's going to keep improving, it's going to keep getting better as the product gets better, as the model gets better. But this is early days. I think still everyone using co-work today is an early adopter. Everyone even using AI today is an early adopter. There are so many people in the world and most people have not tried AI in a meaningful sense. So there's just a lot more room to improve this.
是的,我们 6 月 18 日要在旧金山办一个活动,很多营销材料我都是用 co-work 做的。我反复调整。我不让它一次性完成,所以我会看文案,但我会做像上传我们的下载数据来展示播客增长之类的事情。我给它演讲者的名字,它在制作宣传册方面非常出色。这是活动内容,谁会来参加,谁在演讲,你为什么应该来,如何联系。
Yeah, we're hosting an event here in San Francisco on June 18th and a lot of the marketing material I've turned out with co-work. Now, I go back and forth. I don't let it oneshot it, so I'm looking at the copy, but I do things like upload our download statistics to show the growth of the podcast. And I give it the names of the speakers and it is amazing at building a prospectus. Here's what the event's going to be, who's going to be in the audience, who's speaking, why you should be there, how to get in touch.
太棒了。它太好了。
Insane. It's so good.
你第一次使用它、第一次看到智能体使用你的工具时是什么感觉?
What was your feeling the first time that you used it and the first time that you saw the agents use your tools?
嗯,我基本上启用了所有功能。我认为这是很多人都有的体验:有一个 Claude 的浏览器扩展,你意识到只有让 Claude 接管你的浏览器并为你做事,你才能获得这种好处——或者获得最大好处。这种体验几乎和我用 Waymo 时一样。最初几次交互,我紧张地握着拳头,看着,想着‘我应该批准吗?’阅读所有内容。然后你开始有点信任它,你就一直点批准、批准、批准。在 Waymo 中也是一样。你会想:‘好吧,这看起来不会害死我。’然后 5 分钟后,你就在手机上,AI 在干活。这就是我用代码和 co-work 的体验。这大致符合吗?
Well, I've sort of enabled everything. I think this is an experience that many people have had: there's a browser extension for Claude and you realize that you can only get the benefit of this—or you'll get most benefit—by letting Claude take over your browser and do things for you. The experience is kind of almost the same as I had with a Waymo. Those first couple turns I was white-knuckling, watching, thinking 'should I approve?' reading everything. Then you start to trust it a little bit and you just hit approve, approve, approve. In the Waymo, the same thing. You're like, 'Okay, this looks like it's not going to kill me.' And then 5 minutes later, you're on your phone as the AI does the work. And that was my experience with code and co-work. Does that sort of track?
我的体验也是这样。就像任何技术一样。我观察了一个朋友,她一直在学习使用 co-work。她不是工程师。
I mean, this is like my experience, too. It's like any technology. I was watching a friend that's been learning to use co-work over time. She's not an engineer.
前几天有个例子。她的笔记本电脑语言输入出了问题,不知道怎么修。以前她会去谷歌搜索,但这次她直接问了 Co-Work。Co-Work 说:‘让我看看,能用你的电脑吗?’她同意了,然后 Co-Work 接管了电脑,屏幕泛起橙色光。你能看到 Co-Work 打开设置,诊断语言选择器的问题,然后修复它。你仍然在驾驶座上,所以能看到整个过程。这太神奇了。我的本能是打开谷歌,但她直接用了 Co-Work。我经常看到新用户这样——他们用 Co-Work 做我根本想不到的事情。这很惊人,也很有创意。每次看到都能学到很多。
There was this use case the other day. She had a language input issue on her laptop and couldn't figure out how to fix it. Before, she would have Googled it, but this time she asked Co-Work. Co-Work said, 'Let me take a look, can I use your computer?' She said yes, and it took over the computer with an orange glow. You watch as Co-Work opens settings, diagnoses the language picker issue, and fixes it. You're still in the driver's seat, so you can see it happening. It's magical. My instinct was to open Google, but for her, she went straight to Co-Work. I see this often with new users—they use Co-Work for things I wouldn't have thought of. It's amazing and creative. I learn a lot every time.
目前最大的缺点是速率限制。我看到你在 X 上回复过这个问题。当人们说他们试过 Claude Code 但放弃了,通常是因为用完了 token 配额,只能用一小时,然后要等四小时。他们就会找替代品。速率限制对产品增长有什么影响?有没有计划取消这些限制?
The biggest drawback right now is the rate limits. I've seen you reply to people on X about this. When people say they've tried Claude Code but are done with it, it's usually because they hit their token allotment and it only works for an hour, then they have to wait four hours. They look for alternatives. What have rate limits done to your product's growth, and what's the plan to remove them?
我们正在积极解决这个问题。实际上,只有很小比例的用户会达到速率限制——Pro 用户的比例低得惊人,Max 用户稍高一些,但仍然很低。发生了两件事:我们曾短暂降低了峰值速率限制,但已经回滚并翻倍了。另外,Claude Code 支持插件和集成,有些插件使用 token 效率很低。我们正在让用户能看到这些信息,以便决定使用哪些插件。第三,很多人成了重度用户。刚发布 Claude Code 时,一次只能运行一个。现在我自己一次可能运行五个,大多数晚上我会并行运行几百甚至几千个。这消耗了大量 token。这些新工作流已经接近 Max 计划的极限。如果需要更多 token,也可以通过 API 付费,很多企业就是这么做的。
We're actively working on this. The reality is a very small percentage of people actually hit their rate limits—surprisingly low for pro users, a bit higher for max, but still quite low. A couple things happened: we reduced peak rate limits briefly, but that's rolled back and we've doubled them. Also, Claude Code is extensible with plugins and integrations, some of which use tokens inefficiently. We're surfacing this so users can decide which plugins to use. Third, many people have become power users. When we first released Claude Code, you ran one at a time. Now I run maybe five at a time, and most nights I run hundreds or even thousands in parallel. That uses a lot of tokens. These new workflows are at the edge of what a max plan can do. You can also pay via API if you need more tokens, which many enterprises do.
Anthropic 的 CEO Dario 提到在数据中心支出上保持纪律,与 OpenAI 的做法形成对比。但 OpenAI 也在做 Codex,而且他们有大量算力。你怎么看?当用户遇到速率限制时,可能会转向 Codex。Anthropic 内部如何看待这种竞争?
Dario, Anthropic's CEO, mentioned being disciplined in spending on data centers, contrasting with OpenAI's approach. But OpenAI is doing the same with Codex, and they have a lot of capacity. How do you think about this? When people hit rate limits, they might switch to Codex. How does Anthropic view this competition?
我们的增长从未像今天这么快。Claude Code 的增长正在加速。大多数人不常遇到速率限制,所以这不是大问题。我们专注于改善体验——我们将 500 的速率限制翻倍,今天还宣布增加每周速率限制。我们还宣布了新的 Colossus 算力来服务这些新用户。这种增长超出了我们最疯狂的预测。最重要的是服务好用户,让他们满意。我们正在竭尽全力。
Our growth has never been faster than today. Claude Code's growth is accelerating. Most people don't hit rate limits often, so it's not a huge issue. We're laser-focused on improving the experience—we doubled the 500 rate limits, and today we're announcing increased weekly rate limits. We also announced new Colossus capacity to serve all these new users. This growth is beyond our wildest forecasts. What matters most is serving our users and making them happy. We're doing everything we can.
你对 Codex 感到惊讶吗?你如何看待这个竞争对手?
Are you surprised by Codex? How do you view them as a competitor?
总会有模仿者和竞争对手。这很荣幸,也迫使每个人都做得更好。对我来说,最重要的是尽最大努力服务好用户。我们鼓励团队每天与用户交流,让产品每天进步一点点。
There are always copycats and competitors. It's flattering and forces everyone to do better. For me, the most important thing is doing the best job to serve our users. We encourage the team to talk to users every day and make the product a little better each day.
好的,我想休息一下。我们还有很多要聊——这如何扩展到代码之外,聊天机器人的未来,等等。我们真的需要两个小时。先休息一下,然后继续。
Okay, I want to take a break. We have so much more to cover—how this extends beyond code, the future of the chatbot, and more. We really need two hours. Let's take a break and come back.
欢迎回到 Big Technology Podcast,今天我们请到的是 Anthropic 的 Claude Code 负责人 Boris Cherny。Boris,很高兴你能来。我每天都在用你们的产品,所以能和你聊聊它真的很棒。我们之前聊过一些,但我觉得有一点需要强调:这真的会超越聊天机器人。我们聊过订机票、营销演示,而就在这周,你们又推出了一个新用例:Claude Code 可以用于小企业,包括接管 QuickBooks 做记账。这会走向何方?你觉得大致的路线图会带我们去哪里?
And we're back here on Big Technology Podcast with Boris Cherny, the head of Claude Code at Anthropic. Boris, it's great having you here. I'm in your product daily, so it's really fun to speak with you about it. We talked a little bit about this, but I think one thing we should highlight is that this is really going to extend beyond the chatbot. We talked about booking flights, marketing presentations, and the week we're talking, you have a new use case where Claude Code can be used for small businesses, including taking over QuickBooks and doing some bookkeeping. Where does this go? What do you think the broad roadmap takes you?
我们在为 Claude Code 和 co-work 考虑几件事。有几个大主题。一是提升智能。随着模型改进,我们能做更有雄心的任务。对于编程,以前是一次写一行代码,现在是构建整个功能或整个产品。对于 co-work,它最近才起步,但以前是制作文档,现在则是订机票、组合多种工具、做你的 QuickBooks。所以这个前沿正在快速推进。我们还在考虑如何让 Claude Code 执行更长时间的任务。我们最近推出了一个叫 auto mode 的功能。auto mode 本质上替代了权限提示。以前,每当模型使用工具时,Claude 会问你:‘我可以使用这个工具吗?’通常你只会说‘是’,然后一遍又一遍地说‘是’,感到厌倦。
We're thinking about a few things for Claude Code and for co-work. There are a few big themes. One is improving intelligence. As the model improves, we can do more ambitious work. For coding, it used to be writing a line of code at a time; now it's building entire features or entire products. For co-work, it started pretty recently, but it was like making a document, and now it's things like booking flights, combining many tools, doing your QuickBooks. So this frontier is improving and moving very quickly. We're also thinking about how to do longer running tasks for Claude Code. We recently shipped this thing called auto mode. Auto mode is essentially a replacement for permission prompts. Before, whenever the model uses a tool, Claude would ask you, 'Is it okay if I use this tool?' And usually, you just say yes and get tired of saying yes over and over.
总是允许。就是那个按钮。
Always allow. That's the button to hit.
没错。但出于安全考虑,你必须非常谨慎。我们意识到,与其让人们每次提示都深思熟虑,不如直接解决疲劳问题——因为展示太多对话,他们只会说‘是’或‘总是允许’。所以 auto mode 就是答案。这是一种路由工具调用的新方式。每当 Claude 想使用工具时,它会问另一个 Claude:‘使用这个工具安全吗?’Claude 有一些上下文,但不是全部。还有多层安全检查。我们花了几个月迭代,用数千个基准测试和评估来确保安全。我们在实验室和实际环境中都发现,这比以前更安全。所以对用户来说,这是一个很好的好处,因为你不用一遍遍说‘是’。而且结果更好:如果一个大列表里藏着一个不安全命令,你可能会误点‘是’,但第二个 Claude 用 auto mode 就不会。所以这是个大投入。第三个大主题可能是并行运行更多 Claude。我们早期从 Claude Code 用户那里发现一个很酷的现象:很少有人一次只运行一个 Claude Code。大多数人都运行多个,从几个到几千个。在 co-work 上,我们也看到了同样的情况。随着你越来越放心让 co-work 运行,你会开始一个任务,然后开始第二个任务,并行做更多事。有很多机会让这种体验更好,让人们更清楚如何做、何时做。
That's right. But it's actually very important for security that you're very thoughtful about this. What we realized is that instead of being thoughtful about every prompt, because we're showing people so many dialogues, they just got fatigued and would say yes or always. So auto mode is the answer. This is a new way of routing these tool calls. Whenever Claude wants to use a tool, it asks another Claude, 'Is it safe to use this tool?' Claude has some context, but not all. There are also layers of safety checks. We spent months iterating to make it really safe, with thousands of benchmarks and evals. We found both in the lab and in the wild that this is safer than what we had before. So as a user, it's a nice benefit because you don't have to say yes over and over. And the result is better because if there's one unsafe command buried in a big list, you might have accidentally said yes, but a second Claude using auto mode won't. So that's one big investment. Maybe the third big one is running more Claudes in parallel. One cool thing we started seeing early with Claude Code users is that very few people run one Claude Code at a time. Most run many, from a few to thousands. With co-work, we're seeing the same thing. As you get more comfortable letting co-work run, you start a task, then a second task, and you do more in parallel. There's a lot of opportunity to make this experience nice and more obvious for people.
没错。这很可能延伸到你对聊天机器人的使用方式。有趣的是,Anthropic 与聊天机器人的关系很特别。一开始是技术优先,决定构建聊天机器人,推出 Claude,然后转向企业。你看所有图表,Claude 总是在底部,但现在 Claude 的使用量在上升。我有一个想法想跟你确认:聊天机器人的未来不是‘我给你一个问题,你给我一个答案’,而是‘我给你一个问题或跟你聊一个问题,然后聊天机器人会建议一些可以代表我采取的行动’。就像现在,我一直在聊去印度旅行的事,我认为未来我得到的回复会是像你说的那样——不再需要中间步骤,直接去订机票。一个更主动的聊天机器人会说:‘好的,让我帮你处理。’这是正确的方向吗?我这么想对吗?
Right. And it probably extends to the way you use a chatbot. It's interesting because Anthropic had this kind of interesting relationship with the chatbot. Started out as technology first, decided to build the chatbot, ship Claude, and then moved more towards enterprise. You looked at all the charts and Claude was always at the bottom, but now you're seeing Claude's usage rise. I have a thought I'd love to check by you: the future of the chatbot is not 'I give you a question and you give me an answer.' It's 'I give you a question or talk to you about a problem, and the chatbot will then suggest some action you can take on my behalf.' Like right now I'm talking a lot about a trip to India, and what I think I'll get back in the future is this thing being like what you said—not having this secondary step between having to go there and book the flights. A more proactive chatbot that says, 'Okay, let me take care of this for you.' Is that the right direction? Am I thinking about that?
我能看到这一点。是的。
I could see that. Yeah.
你们在朝这个方向努力吗?
Are you working on it?
智能体是未来,我们正在尝试各种不同的实验。有些东西我们正在尝试,就像这样。是的。
Agents are the future, and we're trying all these different experiments. There's some stuff we're trying that's like this. Yeah.
好吧。但这里有一个限制,对吧,关于它能做什么。人们谈论你可以并行运行数千个 Claude 的限制时,一个有趣的方式是看 Anthropic 在招聘谁。我在 Anthropic 网站上最喜欢的职位是你们在招聘 Salesforce 管理员。你们还在招聘顾问来帮助企业部署这项技术。许多人认为这等于默认了这些东西只能走这么远。沃顿商学院教授 Ethan Mollick 对此评论说:‘当 AI 实验室解散他们新成立的咨询——抱歉,前向部署工程团队时,你就会知道他们相信超级智能。只要还需要人来弄清楚 AI 如何有用、进行组织变革和系统集成,工作似乎就相当安全。’你怎么看?
Okay. But there is a limit here, right, to what this can do. A funny way people have talked about the limits of the thousands of Claudes you can run in parallel is looking at who Anthropic is hiring. My favorite job listing on the Anthropic site is that you're hiring Salesforce administrators. You're also hiring consultants to help enterprises deploy this technology. Many view that as a tacit admission that this stuff can only take you so far. Here's Wharton professor Ethan Mollick on it. He says, 'You will know that the AI labs believe in artificial superintelligence when they disband their newly formed consulting, sorry, forward deployed engineering groups. As long as people are required to figure out how AI is useful and do organizational change and systems integrations, jobs seem pretty safe.' What do you think about that?
是的。看看我做的这类工程,我不写代码。我提示 Claude。实际上现在,我主要做的是让一个 Claude 去提示其他 Claude。所以我甚至不和 Claude 对话。我有一个 Claude 在和我的其他 Claude 对话。我认为在工程领域,你看到了一个人拥有的杠杆作用爆炸式增长。问题在于一个人能建立多大的业务?一个人能支持多少产品?现在 Anthropic 一个工程师拥有的杠杆作用简直疯狂。而且我认为我们开始在其他领域也看到这一点。
Yeah. When you look at the kind of engineering that I do, I don't write code. I prompt Claude. And actually nowadays, mostly what I'm doing is I have a Claude that prompts other Claudes. So I don't even talk to Claude. I have a Claude that's talking to my Claudes. And I think in engineering, you've seen just this explosion in the amount of leverage that a single person has. It's about how big of a business can a person build? How many products can one person support? The leverage that one engineer has now at Anthropic is just insane. And I think we're starting to see this across other disciplines too.
所以我们开始看到这一点,营销人员用 Claude 做事,外派工程师用 Claude 构建实现,我们的销售团队也是如此——实际上在 Anthropic,大约一半的市场团队用 Claude,另一半用 Core。每个人都在用这些产品。我们看到的是,个人获得的杠杆在增加,而我们仍然受限于优秀人才的数量。所以即使人均杠杆上升,你仍然招不到足够多的优秀人才,因为需求太疯狂了,还有太多东西要建。所以这仍然是我们的瓶颈。但我想说,如果有人认为这些东西如此强大,你可以说:看看我的销售组织如何运作,然后用一个提示词配置 Salesforce。另一个例子是:如果 Anthropic 让 AI 处理 IPO 文件而不聘请投行,我就相信他们的 AI 很强大。这些测试不公平吗?
So we're starting to see this with the marketers that are using cloud to do things. We're starting to see this also for forward deployed engineers that are using Claude to build implementations. We're seeing this for our sales team because, actually at Anthropic, I think like half the go-to-market team uses Claude and the other half uses Core. Everyone's using all these products. The thing we're seeing is the amount of leverage an individual has goes up, and we are still bottlenecked on the number of good people. So even if the leverage per person goes up, you still just can't hire enough good people because the demand is so insane and there's so much more to build. So that's still the bottleneck for us. But I would say if people would argue that if this stuff was so powerful, you could say take a look at the way my sales organization operates and then configure Salesforce that way with a prompt. Another example people give is: I'll believe that Anthropic has very powerful AI if they let it handle the IPO paperwork and don't hire an investment bank. Are these unfair tests?
嗯,我们开始看到团队里有一个人用 Claude 报税。我不一定推荐这么做,但我承认我让 Claude 处理过我的税务,然后和我的会计师对比,结果相当接近。
Well, we're starting to see there's one person on the team that was using Claude to do their taxes. I would not necessarily recommend this, but I'll admit I've run my taxes through Claude and compared it against my accountant and it was pretty close.
是的,我也做过同样的事。
Yeah, I did the same thing.
各位,不是建议你们也这么做,但这确实是个有趣的用例。
Folks, not suggesting you should do that, but it is an interesting use case.
没错。但我认为人们在这个对话中忽略的根本问题是,最终必须有一个人与 Claude 对话,让 Claude 做这件事。所以即使 Salesforce 被自动配置,是 Claude 在做,也必须有人让 Claude 去做。如果你需要以多种不同方式配置 Salesforce,让 Claude 做这件事可能实际上是一份全职工作。而总有一天 Claude 会变得非常擅长让 Claude 做这件事。那个人会去问那个被问的 Claude。这个链条会越来越深。但最终,你仍然需要人来驾驶这一切。
That's right. But I think fundamentally what people are missing in this conversation is in the end it's a person that has to talk to Claude to ask Claude to do this thing. So even if Salesforce is automatically configured and it's Claude doing it, someone has to ask Claude to do that. And if you have to configure Salesforce in a bunch of different ways, it could actually be a full-time job to ask Claude to do this. And at some point Claude is going to become really good at asking Claude to do this. And that person is going to be asking Claude that asked Claude to do this. This chain will just keep getting deeper. But in the end, you still need people that are piloting this.
但也许他们未来的工作只是问一个问题。
But maybe their job is just asking one question then in the future.
是的。但想象一下那有多大的杠杆——问对问题。
Yeah. But imagine how much leverage that has—asking the right question.
没错。说得好。
That's true. That's a good point.
所以我们谈到了 Salesforce,那就必须谈谈 SaaS 的末日。你对哪些软件公司会安全、哪些会陷入困境有一些有趣的观点,随着编程越来越自动化。你之前谈到过不同的护城河,哪些更重要,哪些不那么重要。你能简单分享一下吗?
So we talked about Salesforce, so we have to talk about the SaaS apocalypse. You have some interesting views on the type of software companies that will be safe as we get more automated programming and those that might be in trouble. You've talked previously about the different moats that exist and which moats are more important and which are less important. Can you just share that briefly?
有一个非常好的框架叫《七种力量》,用来讨论商业中的护城河。这方面的框架很多,但这是我最喜欢的。我上学时学的是经济学,不是计算机科学,所以我仍然用这类框架思考。商业中有很多不同的护城河。有些公司有一个护城河,有些有几个——它们有一个护城河组合。一个是规模经济:随着生产规模扩大,规模报酬递增。另一个是网络效应:比如即时通讯应用,使用的人越多,对每个用户的价值就越大。另一个是转换成本。还有一个是流程能力。我认为大多数护城河仍然重要,但相对而言,有些在未来一年会变得更加重要,有些则会减弱。我认为会变得更加重要的是网络效应,因为谁在写代码并不重要。不管你的产品核心是智能体还是产品中有智能,如果你的产品有网络效应,那仍然重要。有些护城河变得不那么重要,比如转换成本,因为如果你想从供应商 A 切换到供应商 B,你可以直接让 Claude 去做,而 Claude 会越来越擅长这件事。所以我认为作为一家公司,你应该思考你的护城河是什么。很多大公司都有很多护城河——不只是一样,因为你要实现规模并建立可持续的防御,你需要积累这些护城河。你需要多个。但总之,我会思考什么会变得更有价值,什么会变得不那么有价值。
There's this really good framework called the Seven Powers for talking about moats in business. There are so many frameworks for this, but this is my favorite. I actually studied economics in school, not computer science, so this is still the way I think in terms of these frameworks. There are a lot of different moats in business. Some companies have one moat, some have a few—they have a portfolio of moats. One is scale economies: as you scale up production, there are increasing returns to scale. Another is network effects: like a messaging app, the more people on it, the more valuable it is for any person. Another is switching costs. There's another one that's process power. I think most of these moats are still going to matter, and relatively some are going to increase in importance over the next year and some are going to decrease. One that I think will increase in importance is network effects, because it doesn't matter who's writing the code. It doesn't matter if it's an agent at the core of your product or if there's intelligence in your product. If there's a network effect in your product, that's still going to matter. Some moats get less important, for example switching costs, because if you want to switch from vendor A to vendor B, you can just ask Claude to do that, and Claude is going to get better and better over time at it. So I think as a company, you should be thinking about what your moats are. A lot of the largest companies have many moats—it's not just one thing, because the way you get to scale and build a defensible business over time is you accumulate these moats. You need a number of them. But yeah, I would just think what's going to be more valuable and what's less valuable.
我认为当你考虑这些不同的软件公司时,如果你使用 Claude,所有护城河都会消失,因为你可能最终只用一个应用来与所有软件交互,这意味着实际上只有一家软件公司。
I think when you think about these different software companies, if you're using Claude, all moats kind of blend away because you could potentially be in this one app that is interfacing with all software, which means there's really only one software company.
是的。我的意思是,有很多种可能的发展方式。我认为这种情况是可能的,但对我来说有点牵强,因为如果我考虑,比如,假设我用一个即时通讯应用,我如何决定用哪个?我用的是我朋友在用的、我能联系到的那个。所以即使我能为自己构建一个非常棒的应用——我今天就能做到,用 Claude 几个小时就能建一个很棒的即时通讯应用——它仍然没用,因为我无法和朋友们聊天。
Yeah. I mean there are just a lot of ways this could play out. I think something like this is possible, but it seems a little far-fetched to me because if I think about, for example, let's say I'm using a messaging app, how do I decide which app to use? I use the app that my friends are on that I can reach. So it doesn't matter if I can build a really awesome app for myself, which I can do today. I can build a great messaging app with Claude in a few hours. It's still not useful because I can't talk to my friends.
但这就是例子。你的即时通讯应用里会有一个智能体,通知你朋友发消息了。我知道你经常在 iPhone 上用 Claude,对吧?所以你只会看到通知,然后回复别人。只要公司配合,你所有的沟通都可能集中到这些应用里。
But this is the example. You'll have an agent in your messaging apps that will let you know when your friends have messaged you. I know you use Claude on your iPhone a lot, right? So you will just see the notification and you'll speak back to people. All your communication could potentially be centralized in these as long as the companies play ball.
我的意思是,最终可能还是智能体,但通信到底是怎么发生的呢?比如,你看像 Signal 这样的消息应用,它使用一种协议进行通信。我可以构建一个应用,也许能用同样的协议,但我认为它实际上无法给 Signal 上的其他人发消息。不过,是的,我可以让一个智能体使用我的应用,通过一个支持这种功能的现有应用来完成消息发送。
I mean, it could be kind of the agent in the end, but how does the communication actually happen? For example, if you look at a messaging app like Signal, there's a protocol it uses to communicate. I can build an app that maybe uses that same protocol, but I think it actually can't message other people that are on Signal. But yeah, I can have an agent that uses my app to do that messaging using an existing app that supports this.
嗯。
Yeah.
所以,事情会如何发展还不明显。我认为如今人们混合使用应用和智能体。但我从根本上认为,很多这些模式实际上仍会随着时间的推移而增值。
So it's not obvious how it's going to play out. I think today people use a mix of apps and agents. But I do fundamentally think that a lot of these modes are actually still going to increase in value over time.
你可以想另一个例子,比如台积电或某种芯片制造商。想想他们投入了多少工作来打造一个流程,使得成本随规模下降,这是一种基本的经济力量。有很多公司做这类事情,尤其是在制造业,规模扩大成本降低。对于科技公司,基础设施也是如此。所以如果你构建了非常好的基础设施,你可以支持更多用户,每个用户的边际成本随时间下降。所以如果你有这种效应,不管你我能不能构建应用,这仍然是一个非常强大的模式。但我确实认为两件事都在起作用。
You can think of another example, let's say TSMC or some kind of chip manufacturer. If you think about the amount of work they put into making a process where the costs go down with scale, this is a fundamental economic force. There are a lot of companies that do this kind of thing, especially in manufacturing, where with scale the cost goes down. With tech companies, this is the case for infrastructure. So if you build a really great infrastructure, you can support more users and the marginal cost per user goes down over time. So if you have this kind of effect, it doesn't matter if you or I can build apps. That's still a really powerful mode. But I do think for sure both things are in play.
好的,我还有三个问题,十分钟内。看看我们能不能都问到。Anthropic 的创始人之一 Jack Clark 最近说,他认为这些模型有大约 60% 的可能性会在 2028 年开始自我改进。可能百分比或年份有偏差,但大致准确。你就在那个编码自主进行的应用里。你在运行这个应用。你同意 Jack 吗?
Okay, I got three more in 10 minutes. Let's see if we can get to them all. Jack Clark, one of the Anthropic founders, recently said he believes there's like a 60% chance that these models will start improving themselves by 2028. It could be off by a percentage or a year, but ballpark that's accurate. You're in the app where coding happens autonomously. You're running this app. Do you agree with Jack?
看起来是对的。是的。当我看到 Claude Code 的编写方式时,100% 的 Claude Code 都是用 Claude Code 编写的。我认为从去年 11 月,也就是 Opus 4.5 以来,情况就是如此。
Seems right. Yeah. When I look at the way that Claude Code is written, 100% of Claude Code is written using Claude Code. This has been the case since I think November of last year, since Opus 4.5.
那就像是快速起飞的情景。你预料到这一点吗?
It's like a fast takeoff scenario then. Do you anticipate that?
我的意思是,这是可能的,这也是 Anthropic 存在的原因。如果你问任何工程师或研究员为什么加入 Anthropic,他们会告诉你为了 AI 安全。这是因为当我们思考未来,多年以后,最重要的事情,我们想为孩子们做对的事情,就是确保这个东西是安全的,确保它进展顺利,因为那是可能的结果之一。我认为我们还没有看到这种情况。现在 Claude Code 在自我编写,但仍然是人在做提示。Claude 开始为自己生成接下来要构建什么的想法,但并不总是好主意,我仍然产生大部分想法。在某个时候,它会改变。模型会改进,它会变得更像一个自我强化的循环。
I mean, it's possible and this is why Anthropic exists. If you ask any engineer or any researcher why they joined Anthropic, they're going to tell you it's for AI safety. And it's because for us when we think about the future, years from now, the thing that's the most important and the thing that we want to get right for our kids is we want to make sure this thing is safe and we want to make sure it goes well because that is one of the possible outcomes. I think that's not yet what we're seeing. Right now Claude Code is writing itself, but it's still a person that's doing the prompting. Claude is starting to generate its own ideas for what to build next for Claude Code, but it's not always good ideas, and I still generate most of the ideas. At some point, it's going to change. The model's going to improve, and it's going to become more of a self-reinforcing loop.
好的。我确实想听听你对世界模型争论的看法。支持世界模型的人说,大型语言模型不理解后果,你需要构建一个世界模型才能有有效的智能体。这是 Yann LeCun 的观点。他说没有世界模型就无法构建可靠的智能体系统。LLM 没有世界模型。它们无法在行动前预测后果。根据 Yann 的说法,它们只是行动,接下来发生什么就是别人的问题了。我最近和 OpenAI 的 Greg Brockman 聊过,他说他基本上不接受这个论点,他认为 LLM 是直接通往 AGI 的方式,这些文本模型就是通往 AGI 的道路。你站在哪一边?你相信世界模型智能需要被内置,还是认为仅凭 LLM 就足够了?
Okay. I definitely want to get your thoughts on the world model argument here where people who are pro-world model say that a large language model has no understanding of the consequences and you need to build a world model into it to have effective agents. Here's something from Yann LeCun. He says you cannot build a reliable agentic system without a world model. LLMs don't have world models. They can't predict the consequences of their actions before taking them. According to Yann, they just act and whatever happens next is someone else's problem. I was speaking with Greg Brockman from OpenAI recently and he said basically he doesn't accept that argument and he thinks LLMs are the way directly, these text models are the way to AGI. Which side are you on? Are you a believer that world model intelligence needs to be baked in or do you think that LLMs alone are good enough?
我想向 Yann 发出邀请,如果他愿意坐下来和我一起用 Claude Code 一个小时。我很想展示给他看。
I would put out an offer to Yann if he wants to sit down and Claude Code together for an hour. I'd love to show him.
你们应该在这个节目上做这件事。
You guys should do that on this show.
是的。然后我很好奇他会怎么想。也许他会改变主意,也许不会。
Yeah. And then I'm curious to hear what he thinks. Maybe he'll change his mind, maybe he doesn't.
对。但你的观点呢,
Right. But your perspective though,
你知道,我非常坚定地站在产品这边。所以我对此没有真正的观点。
You know, I'm pretty firmly on the product side. So I don't really have a perspective on it.
但好吧,如果你不介意,让我再深入一点。你站在产品这边,但我听到很多人提出这个观点:没有对世界运作方式的概念,比如世界模型,LLM 就不理解世界运作方式和后果等。你用 Claude Code 订了多少航班?八次航班和酒店。你一定认为它有一定理解后果的能力,否则你不会给它你的信用卡,我猜你给了。那么你特别如何看待这个论点?
But okay, let me drill down a tiny bit deeper if you don't mind. You're on the product side, but I've heard multiple people bring out this idea that without a conception of the way the world works, like in a world model, an LLM just doesn't have an understanding of the way that the world works and consequences and stuff. You use Claude Code to book how many flights? Eight flights and hotels. Like, you must think that it has some understanding of consequences, otherwise you wouldn't have given it your credit card, which I presume you did. So what do you think about that argument in particular?
我认为,从我在 Anthropic 从事研究的人那里读到的内容来看,这些模型的智能程度令人惊讶,因为就像你一开始说的,它们基本做的事情就是预测下一个词元。所以你会觉得这有点蠢,这怎么可能导致智能?但我们实际上发表了很多工作,关于模型如何能够规划,它们实际上能够推理。有很多非常令人惊讶的行为,你实际上不会期望一个只预测下一个词元的模型会有这些行为。所以我不知道,我不会低估它。
I think from what I've read from folks working on research at Anthropic, it is surprising the degree to which these models are intelligent because like you said at the beginning, the thing that they fundamentally do is they predict the next token. And so you think this is kind of a stupid thing, like how can this possibly lead to intelligence? But we've actually published a lot of work about how the models are able to plan, they're able to actually reason. There were all these very surprising behaviors that you actually wouldn't expect from a model that just predicts the next token. So I don't know, I wouldn't discount it.
我的意思是,我最喜欢的是当它们写诗时,在写第一行时,你可以在模型中看到——这是 Anthropic 的研究——它们已经在思考下一行了。
I mean I think my favorite is when they write poetry as they're writing the first line you can see in the model this is Anthropic research that they're already thinking about the next line.
没错。
That's right.
这怎么可能呢?
Which is like how is that even possible?
没错。我的意思是,这就是我对此的看法。如果我在写诗,我也会这样做。这很疯狂,你教这个东西预测下一个词,而如果下一个词足够难,它就必须学会真正提前规划,必须学会如何做所有这些。
That's right. I mean that's kind of how I think about it. Like if I were writing poetry, that's how I would do it too. And it's crazy like you teach this thing to predict the next word and somehow if the next word is hard enough, it has to learn to really plan ahead and it has to learn how to do all of this.
好的,最后一个问题。有时候我看到正在发生的重大技术变革,在我报道这些东西的职业生涯中,有些成功了,有些没有。我总是要问自己,我们如何确定这是未来,而不是一场狂热梦。我认为数据表明这是真实的东西。
Okay, last one for you. Sometimes I wonder when I see big tech changes underway and in my career covering this stuff, some have worked out and some haven't. I always have to ask myself how are we sure that this is the future and this is not a fever dream. And I think the data indicates that this is a real thing.
但我也在想,你不得不质疑,从未来发展的角度来看,我们能推断出多少。有一种观点认为这只是一场狂热梦,也许人们只想要简单的界面,他们不介意点击操作,而用类似代码的语言交流感觉太技术化了,不会像在开发者中那样吸引普通用户。你怎么看?
But I also wonder, you have to question how much you can extrapolate towards the future in terms of how this will continue to progress. The argument that this is a fever dream is that maybe people just want simple interfaces and they don't mind tapping through things, and speaking in a code-like language feels a bit too techy and it just won't appeal to the everyday user as much as it's taken off with developers. How would you answer that?
我们最近为 Opus 4.7 举办了一场黑客马拉松,其中一位获奖者是一名医生,他构建了一个应用。还有一位电工、一位木匠,这些人很多都没有编程经验,但他们用代码构建了有用的东西。有一个人因为我们的黑客马拉松而创建并出售了一家初创公司。毫无疑问,我们最初构建代码时是为工程师设计的,工程师们学会了如何使用它。但很快,非工程师的人也学会了用它来构建有经济价值的东西。实际上,看看现在的使用情况,很多用户都不是工程师。它太有用了,以至于人们不遗余力地使用它。甚至在桌面应用出现之前,人们就在终端里安装代码。对很多人来说,这是他们第一次使用终端。现在我们有了桌面应用、iOS 应用、Slack 应用——多种交互方式。但人们之前就愿意克服困难去使用它,因为它太有用了。对我来说,作为产品人员,这是最终的市场检验:这个东西有用吗?有很多人每天使用并持续使用吗?答案是肯定的,用户很多,而且还在增长。我经常对人们使用它的方式感到惊讶。
We had a hackathon for Opus 4.7 recently, and one of the winners was a doctor that built an app. There was an electrician, a carpenter, and a lot of these people didn't have coding experience, but they used code to build something useful. One person built and sold a startup as a result of one of these hackathons. Undoubtedly, when we first built the code, it was for engineers, and engineers figured out how to use it. But very quickly, people that were not engineers figured out how to use this to build economically useful things. Actually, if you look at a lot of the usage today, it's not engineers. It's just so useful for people that they were going out of their way, jumping through hoops. Even before the desktop app, people were installing code in a terminal. For a lot of people, this was their first time using a terminal. Now we have a desktop app, an iOS app, a Slack app—many ways to interact with it. But people were jumping through hoops to use it because it was so useful. For me as a product person, this is the ultimate market test: is this thing useful? Are there a lot of people that use it every day and keep using it? And yes, it's a lot of people, and it just keeps growing. I'm constantly surprised by the way people use this.
是的,我得说我自己使用这些工具的方式也让我感到惊讶。我不知道,我们拭目以待接下来会发生什么。非常兴奋能继续使用它,也很高兴有机会和你交流。希望我们能再次对话。
Yeah, I will say I've been surprised by the way I've found myself using the tools. I don't know, we'll see what comes next. So excited to keep using it and thrilled to have a chance to speak with you. I hope we can do it again.
谢谢邀请我。
Yeah, thanks for having me on.
好的,谢谢你 Boris。很高兴和你交流。各位听众和观众,非常感谢你们的收听和观看,我们下次在 Big Technology Podcast 再见。
All right, thank you Boris. Great speaking with you. All right, everybody. Thank you so much for listening and watching and we'll see you next time on Big Technology Podcast.