AI Coding on Phone: Boris on 8B Tokens & ROI
打开互动全文版(中英对照 + 朗读 + 问答)→Boris 分享他自去年 11 月起 100%用 Claude 在手机上写代码、使用 80 亿 Token 的经历,并讨论企业应如何平衡 AI 模型成本与 ROI,建议给全员发放 Token 进行实验。
Boris shares his experience writing 100% of code with Claude on phone, using 8B tokens since March, and discusses how companies should balance AI model cost vs ROI by giving everyone tokens to experiment.
大家早上好。今天我和 Boris 一起。在开始之前,我们得赶紧拍张自拍。如果大家能加入的话。我快速拍一张,好吗?好了,开始。大家说茄子。开玩笑的。好了。
Good morning, guys. I'm here joined with Boris. Before we get started, we need to take a selfie real quick. If you guys can join us. I'm going to take one real quick, all right? All right, here we go. Everyone say cheese. Just kidding. Okay.
我们有 40 分钟时间,和我了不起的朋友 Boris 一起。他不需要介绍,但如果你没听说过他,他是 Claude Code 的负责人。我是 Meta Dev Infra 团队的产品总监,负责 AI 开发者生产力相关的工作。我们将进行一场友好的炉边谈话,讨论一些想问 Boris 的问题。有一个二维码,我想之前屏幕上分享过。我们稍后会开放提问。我这里有一个 iPad,上面有大家提交的问题。我们会挑一些投票最高的问题来回答。但在大家填写的同时,我想快速问一下,有多少人用 AI 写代码?请举手。好的,不少人。如果你们 100%的代码都是由 AI 写的,请继续举手。看看房间里有几位?好的。这比我之前问这个问题时多多了。所以,现在谈论这个话题真的很有意思,因为我们都在见证行业的变革。那么,我们开始吧。我先问几个问题。然后等大家填写并投票选出最热门的问题后,我再问那些。先来几个热身问题。Boris,你今年写了多少行代码?
We have 40 minutes joined by my amazing friend Boris. Who doesn't need any introduction, but in case you haven't heard of him, he's the head of Claude Code. I'm a director of product on the Dev Infra team here at Meta supporting our AI developer productivity stuff. And we're going to just have a friendly fireside chat to chat about a bunch of different questions that we have for Boris. There's a QR code, I think that was shared up screen earlier or somewhere. We are going to have a later portion of the talk to be open mic. So, I have an iPad here with a bunch of questions that you guys are submitting. We'll go through some of those top voted questions. But, as you guys are going to fill those out, I want to quickly ask folks, how many people here use AI to write their code? Just a quick raise of hands. Okay, good amount of people. Keep your hands up though if you 100% of your code is written by AI. Let's look around the room. How many people is that? Okay. It's a lot more than before when I asked these questions. So, it's really kind of an interesting time to be here talking about this stuff because all of us are witnessing this transformation in our industry. And so, I think we're going to get this question started. I'm going to have a few questions that I'm going to start off first. And then, as you guys fill this out and vote on the top questions, I'll get to those afterwards. So, just a few quick warmer questions. Boris, how many lines of code have you written this year?
多少行代码?
How many lines of code?
对。
Yeah.
我其实提前准备了答案,因为我知道你会问这个问题。好的。1700 个 PR。新增了 40 万行,删除了 25 万行。
So, I actually pulled the answer for this in prep because I knew you were going to ask that question. Okay. So, 1.7 thousand PRs. Added 400,000 lines, deleted 250,000 lines.
好的。
Okay.
我记得去年我删的比加的多。但今年我加的多一些,我还想看看 token 数。不幸的是,数据因为保留策略之类的原因被删了。但从三月份以来,我用了 80 亿个 token。
I think last year it was actually I deleted more than I added. But this year I added a little bit more than I deleted and I also tried to get the token count. Unfortunately, the data kind of gets deleted due to retention and stuff. But since March, I used 8 billion tokens.
80 亿个 token。那你删掉和添加的那些代码,都是你自己写的还是 Claude code 写的?
8 billion tokens. And so, those lines of code that you deleted and added, those are all written by you or Claude code?
从 Opus 4.5 开始,我 100%的代码都是由 Claude code 写的。那大概是去年 11 月。
100% of my code has been written by Claude code since Opus 4.5. So, that's like November of last year.
太疯狂了。那你现在是用手机还是笔记本电脑写代码?比例大概是多少?
That's crazy. And then, do you use your phone or your laptop? Like what's the split nowadays when you code?
这大概是最疯狂的事了。如果六个月前你问我我们会在哪里写代码,我绝对想不到。但现在我大部分代码都是在手机上写的。六个月前如果有人告诉我这个,我会觉得他疯了。但事实就是这样。
Yeah, this is sort of like the craziest thing. I would not have predicted this if you asked me like 6 months ago, where are we going to be doing our coding? But yeah, like most of my coding now is on my phone. I would have said you're crazy if you told me that 6 months ago. But yeah, here we are.
是啊,很疯狂对吧?好的,我想很多人一直在问我的一个问题是,他们开始考虑效率和 ROI。我们听说 Uber 和其他公司开始设定年度或月度预算,比如每个工程师 1500 美元。但与此同时,Anthropic 你们和其他前沿实验室也在推出能力更强但也更贵的模型。你怎么看?公司应该如何平衡性能更强但成本更高的模型与需要展示更高效率、token 效率之间的张力?
Yeah, it's crazy, right? Okay, I think one of the questions that a lot of people have been pinging me about and we really want to get your thoughts on is they're really starting to think about like efficiency and ROI. We kind of hear of Uber and other companies starting to set like annual like monthly budgets, for example, like $1,500 per engineer. But at the same time, Anthropic, you guys and other frontier labs are pushing out these more capable but also more expensive models. What do you think about this? Like what are your thoughts on how companies should balance this tension between like more performing but costly models versus like having to start to demonstrate more efficiency, token efficiency?
完全同意。幸运的是,我每天都能和很多公司交流,包括我们的客户和潜在客户。大体上,人们有两种思考方式。一些公司从成本角度考虑,另一些从 ROI 角度。我认为 ROI 绝对是正确的框架,因为你不能只考虑成本,你投入了东西,也会得到回报。当你考虑 ROI 和部署 Claude Code 时,我认为思考一个好的部署方式很有用。我们看到最成功的公司,他们给每个人分配 token,不仅仅是工程师,还有产品经理、设计师、数据科学家。他们给所有人 token,让公司去实验,看看想法从哪里来,因为往往不是你想的那个人。很多最有趣的想法和最创新的流程改进或新产品创意,可能来自组织角落里某个会计,或者 CEO 从未听说过的市场人员。很多创新就来自那里。不一定总是最资深的工程师。因此,我认为公司应该让团队去实验。方法是给人们 token,并给他们实验的安全感,让他们觉得可以尝试而不会受到惩罚。一旦你找到有效的内部用例,你就想控制成本,而且要在后端控制,而不是前端。一旦某个用例起飞,消耗大量 token,你就需要考虑如何优化。有很多方法,比如在 Claude Code 中,我们有按席位成本控制。可以使用顾问模型。你可以为整个公司更换模型。你可以控制整个公司的努力程度。你可以按部门或按后端设置预算。有很多方法。第二个思考角度是部署方面,以及如何实际推广。一旦公司使用 Claude Code 一段时间,你真的需要考虑 ROI。ROI 有两部分:投资和回报。过去,投资很容易衡量,就是 token。回报过去我们衡量的是 AI 写的代码占比,或者代码行数增长百分比。我们刚才看到很多人举手说 100%的代码是 AI 写的。
Yeah, totally. So, you know, like I luckily I get to talk to so many companies every day, you know, that are our customers and kind of future customers. And broadly, people think about it in two ways. Some companies think about it in terms of the cost. And other companies think about it in terms of ROI. And I'm like ROI is absolutely the right framing because you don't want to just think about cost because you kind of spend something on it and you get something back. When you think about ROI and kind of like deploying Claude Code in general, I think it's useful to think about like what's a really good way to deploy it. And I think like the most successful companies we've seen, they sort of give everyone tokens, not just engineers, but also like the product managers, the designers, the data scientists. They give everyone tokens and let the company experiment to see where the ideas come from because often it's just like not the person that you would expect. Often some of the most interesting ideas and the most innovative like ways to improve processes and kind of like new product ideas, it's going to come from like, you know, like an accountant somewhere in the corner of the org or like a marketing person that you've never heard that the CEO has never heard of. That's where a lot of the innovation comes from. It won't necessarily be like the most senior engineer always. And so I think for this reason you kind of want companies to experiment and you want your team to experiment. The way to do this is give people tokens and give them safety to experiment so they feel like they can try stuff and they're not going to get kind of penalized for it. Once you find these internal use cases that kind of work, then you want to control the costs and you want to do that on the back end, not on the front end. And you know, like once there's some use case that takes off, uses a lot of tokens, then you want to think about how to like optimize that. There's all sorts of ways like you know, like in Claude Code we have like per-seat cost controls. There's ways to use like advisor models. You can change the model for your entire company. You can control the effort level for your entire company. You can set budgets like based on department, based on our back. So there's just like all sorts of ways to do this. I think the maybe like the second way to think about it. So this is kind of like the deployment side and like how to actually get it out and like when you're like at the very beginning. Once companies have been using Claude Code for a while, you really do want to think about ROI. And you know, like for ROI, there's kind of like two parts. There's the investment and there's the return. In the past, the way that, you know, investment is easy to measure. That's just tokens. Return in the past, the way that we used to measure it is what percent of the code is written by AI. Or maybe like percent increase in the lines of code or something. And yeah, like, you know, we just saw a bunch of hands go up that said, you know, like 100% of your code is AI.
所以,一旦达到 100%,你实际上如何衡量回报?这就变得有点难了,因为我们以前做过 Devin。在 Devin 上,如果一年有 2-3%的生产力提升,那就算很不错了。现在我们看到的生产力提升是几百个百分点。在 Anthropic,自今年年初以来,每位工程师的代码量增长了 8 倍。所以在这种模式下,你如何看待回报?我认为一部分是让 Claude 写出 100%的代码。然后思考每位工程师的代码量在加速多少。第三件事是还有什么其他瓶颈在阻碍?因为一旦工程师能写出大量代码,瓶颈就会变成好想法。那么如何解除这个瓶颈,让公司更快地产生想法?这可能意味着更多的产品经理或用户研究员。然后你必须思考如何更快地将这些想法推向市场。所以这是在市场推广和营销方面解除瓶颈。这大致是我思考的顺序。当我观察我们的客户时,每个人都处于这个采用曲线的不同阶段。
So, once it gets to 100%, how do you actually measure that return? And that's where it gets kind of hard because, you know, we worked on Devin back in the day. And on Devin, if you had a 2-3% productivity improvement for a year, that used to be pretty good. Now, the productivity improvements we're looking at are hundreds of percentage points. At Anthropic, we've seen an 8x increase in code per engineer since the beginning of this year. So in this kind of regime, how do you think about the return? I think part of it is get to 100% of your code written by Claude. Then think about how much the code per engineer is accelerating. And then the third thing is what are the other bottlenecks getting in the way? Because once you get to this point where engineers are just writing a lot of code, the bottleneck is going to be good ideas. So how do you unblock that so your company can generate ideas faster? This could mean more PMs or user researchers. Then you have to think about how to get these ideas to market faster. So this is unblocking on the GTM and marketing side. This is roughly the order I think about it. When I look at our customers, everyone is at a different step of this adoption curve.
是的。我觉得思考你说的是如何从编码开始,越来越多的人举手表示他们 100%的代码由 AI 编写,这很有趣。但然后你开始意识到,对于公司最终想要的——交付更多商业价值——存在上游和下游的瓶颈。所以解除这些想法瓶颈,让代码更快、更安全、更准确地部署。我认为这些都是整个行业正在解决的问题。就像,好吧,假设编码基本上被 AI 解决了。那么现在阻碍协调和协作的是什么?我实际上看到排名最高的问题之一就是关于这个的。所以我们一会儿会讨论。但我认为这就是我们现在面临的有趣问题。一年前如果我们做这个炉边谈话,我们会讨论如何更多地使用 AI 进行编码?如何做到这些?但现在这更像是下一个问题:既然编码越来越多地由 AI 编写,接下来是什么?所以我认为我们现在处于一个非常有趣的情况。好的,下一个问题。背景是,Boris 和我实际上曾在 Threads 一起工作。当然,Threads 昨天刚刚宣布月活跃用户达到 5 亿。
Yeah. I think it's really interesting to think about how you said it started with coding and just getting more and more hands up where 100% of their code is AI. But then you start to realize there are upstream and downstream bottlenecks to the ultimate thing companies want, which is delivering more business value. So unblocking those ideas, getting code to be deployed faster, safely, and more accurately. I think those are all things the whole industry is figuring out. It's like, okay, let's say coding is largely getting solved with AI. What are the things now blocking coordination and collaboration? I actually see one of the top voted questions is about that. So we'll talk about that in a moment. But I think that's the interesting question we're in now. A year ago if we were doing this fireside chat, we'd be talking about how can we use AI more for coding? How can we do all that? But now it's kind of the next question: what's next now that coding is more or less going to be more and more written by AI. So I think that's a really interesting situation we're in right now. Alright, next question. For context, Boris and I actually used to work together on Threads. Of course, and Threads just hit 500 million monthly actives yesterday, we announced.
恭喜。
Congratulations.
谢谢。所以我不得不向 Threads 上的一些人提问。我有一个来自我们前同事 Peter 的问题。他问:Loops 是下一个炒作周期还是真实的?也许你可以为不熟悉的人解释一下什么是 Loops,以及你个人如何使用它。
Thank you. So I had to ask some folks on Threads for questions. I have one from our ex-colleague Peter. He asked: is Loops the next hype cycle or is it real? Maybe you can explain what Loops are for people who aren't familiar, and then how you use it personally.
是的。在座有人尝试过 Loops、routines 吗?
Yeah. Have people here tried Loops, routines?
还比较早期,是的。
Still pretty nascent, yeah.
好的。为了让我知道解释的技术深度:如果你是写代码的工程师,请举手。好的。对于在座的工程师,思考方式是:两年前我们手动编写源代码。我们开始过渡到让智能体写代码。现在我们正在过渡到智能体提示其他智能体来写代码。这可能是非技术性的解释。更技术性的解释:我的思维模型是,如果源代码是基础层,就像编程中的语句,那么写代码的智能体就像编程中的函数,而 Loops 就像高阶函数。所以我们正在抽象阶梯上再上一层。Loops 的重要性不亚于从源代码到智能体的那一步;Loops 是从智能体到下一步的步骤。同样重要和巨大。对我来说,现在在 Loops 的世界里,我们就像一年半前的智能体。所以还相当早期,但我们开始看到它起作用的迹象。想法是:假设作为一名工程师,我的很多工作是代码审查。我可以手动做代码审查。另一种方式是我有一个智能体,提示它做代码审查。Loops 版本是我有一个智能体在循环中运行,完成所有代码审查。另一个例子:我阅读 Threads 来查看人们的反馈。我可以手动做,或者让智能体做,或者让一个智能体每 5 到 10 分钟循环读取反馈并提交修复的 PR。一年半前我们处于第一步;现在我们处于第三步。当你思考工程师和非工程师(设计师、数据科学家、市场营销人员)的所有工作时,很多可能都可以分解成 Loops。我认为我们开始达到一个点,越来越多的代码以 Loops 形式表达。就我个人而言,平均每天大约 30%的代码由 Loops 编写。如果我非常努力,有些天可以达到 100%,但还没有完全成熟。
Okay. And just so I know how technical to make the explanation: raise your hand if you're an engineer that codes. Okay. So for the engineers in the room, the way to think about it is: two years ago we wrote source code by hand. We started to transition to make it so agents write the code. And now we're transitioning to the point where agents are prompting agents that then write the code. Maybe that's the less technical explanation. The more technical explanation: my mental model is sort of like if source code is like the basic level, like a statement in programming, then the agent writing the code is like a function in programming, and then loops are like a higher-order function. So we're stepping up the abstraction ladder, one more step up. Loops are sort of as big as the step from source code to agents was; loops are the step from agents to the next thing. It's just as important and as big a step. To me, it feels like right now in the world of loops, we're where agents were maybe a year and a half ago. So it's still fairly early, but we're starting to see signs that it's working. The idea is: let's say as an engineer, a lot of my work is code review. One thing I could do is do the code review by hand. Another thing is I can have an agent and prompt it to do the code review. The loop version is I have an agent running in a loop that does all the code reviews. Another example: I read Threads to see people's feedback. I could do this by hand, or have an agent do it, or have an agent constantly looping over every 5 or 10 minutes reading the feedback and putting up PRs for fixes. A year and a half ago we were at the first step; now we're at the third step. When you think about all the work engineers and non-engineers do—designers, data scientists, marketing people—a lot of it can probably be chunked into loops. I think we're starting to get to the point where a bigger percentage of the code is expressed as loops. For me personally, maybe 30% of my code is written by loops on an average day. If I try really hard, I can get to 100% for some days, but it doesn't totally click yet.
不错。我觉得如果我们问有多少工程师写 Loops,也许一年后我们会看到更多举手。所以在某些方面,你总是比团队其他成员领先 3 到 6 个月。我认为 Loops 在内部一直很吸引人。人们开始探索各种方式来编排所有工具以创建这种自动循环,因为现在人类还在循环中——你必须参与每一个环节。
Nice. I feel like if we ask the question of how many engineers write loops maybe in a year from now we'll see more hands. So in some ways you're always living 3 months, 6 months ahead of the rest of the team. I think loops is something that's been fascinating internally. People have been starting to explore the various ways you can orchestrate all the tools to create this automatic loop, because right now humans are in the loop—you have to be every single part of the way.
我们刚才在第一个问题里聊到如何让这个过程越来越自动化。让这个循环运转起来确实是解决问题的关键,也能提高生产力。另一件我们一直在讨论的事情是,Anthropic 最近在 co-work 上投入了很多。我相信在座很多人主要用的是 Claude Code。我想如果你能告诉我们为什么大家应该用 co-work,以及你最兴奋的一些用例,特别是跟我们刚才第一个问题相关的,比如人们除了编程之外还怎么用它,那就太好了。
We were just talking about in the first question how to kind of get this process more and more automated. And so getting this loop going is really sort of the key to figuring this out and getting more productivity. The other thing we've been talking about for a while is that Anthropic's been investing a lot in co-work as of late. I'm pretty sure a lot of the audience here primarily uses Claude Code. I think it'd be great if you could tell us sort of why folks here should use co-work, what are some of the use cases you're most excited about, especially related to the first question we talked about, just like how people are using it beyond just coding.
没错。对于不了解的人来说,使用 co-work 的方式就是下载 Claude 桌面应用。这个应用里同时有聊天、Claude Code 和 co-work。你下载就能用,支持 Mac 和 Windows。Co-work 就是面向非工程师的 Claude Code。它底层就是 Claude Code,实际上用的是和 Claude Code 相同的 Claude 智能体 SDK。所以基础设施是一样的,你也可以自己基于这个 SDK 进行开发。完全是一回事。我们说它不是给工程师用的,而是给非工程师用的,是因为它内置了更多的安全护栏。Co-work 有一个完整的虚拟机,还有相当复杂的机制。我们还会接入操作系统,确保你不会意外删除东西。对提示注入也有很强的保护。我们做了各种事情,让你更难搬起石头砸自己的脚。
Yeah, totally. So for folks that don't know, the way that you try co-work is you download the Claude desktop app. That's the same app that has chat and Claude Code in it. It also has co-work in it. So you just download it. You can use it. Works on Mac, works on Windows. Co-work is Claude Code for non-engineers. It's just Claude Code under the hood and it actually uses the same Claude agent SDK that Claude Code is built on. So the infra is the same and you can actually build on the SDK yourself, too. It's exactly the same thing. The reason we say it's not for engineers, it's for non-engineers, is there are a bunch more guardrails built in. So co-work has like an entire virtual machine. It has this actually quite sophisticated... And then we hook into the operating system to make sure you don't accidentally delete stuff. There's a lot of protection for prompt injection. There's all sorts of things that we do that make it a little bit harder to shoot yourself in the foot.
说到我自己怎么用 co-work,我实际上用它来做所有非工程的事情。举个例子,我用它做项目管理。我们以前每天早上开站会,讨论团队每个人做了什么。现在呢,我让 co-work 在浏览器里打开一个电子表格,这个表格包含了本周所有的工作流。它会自动在 Slack 上给每个工程师发消息,询问最新状态。很多时候,他们的 co-work 智能体会回复。
When I think about how I use co-work, I actually use it for everything that's non-engineering. So, one example is I use it for project management. We used to have stand-up every morning where we talk about what everyone on the team did. And now what I do is I have co-work and it opens up this spreadsheet in the browser, and the spreadsheet has essentially all of the work streams for the week. And it'll message every engineer in Slack for me automatically to ask like what's the latest status. And often their co-work agents will respond.
这是智能体跟智能体对话吗?
What's agents talking to agents?
就像 co-work 智能体在互相交流。
It's like the co-work agents talking to each other.
对,co-work 智能体。
Yeah, co-work agents.
但有时候工程师会回复,然后 co-work 看到后就会在电子表格里更新状态。整个过程就是 co-work 打开浏览器帮你完成。几乎零设置。你只需要 co-work 和 Claude Chrome 扩展,它们就会协同工作。这就是组合工具的神奇之处。我觉得这就是人们使用 co-work 时体验到的神奇时刻:'天哪,这个东西能用我的工具,还能像我自己一样把所有工具组合起来。' 感觉太棒了,就像第一次使用 AI 聊天应用一样,是一种启示。
But sometimes engineers will respond, and then co-work sees this and it'll fill out the status in the spreadsheet. All this is, is co-work will open a browser and it'll do this for you. It's like zero setup. You just need co-work and the Claude Chrome extension, and then they'll just kind of work together to do this. So this is the kind of magic of combining tools. And I think this is the magic moment that people see when they use co-work. It's like, 'Oh my god, this thing can use my tools and it can combine all my tools for me in the way that I would.' And it just feels incredible. It feels like using an AI chat app for the first time. It's like a revelation.
还有更高级的用法。我以前用 co-work 预订所有旅行。我会说:'这是我的行程,我需要某天到某地,某天到某地。你能帮我订机票吗?' 它会打开浏览器,进入 Anthropic 常用的旅行网站,填写信息并订票。现在呢,我进一步自动化了。Co-work 有一个定时任务,每天查看我的邮件,寻找我在 Google 日历上接受的活动。如果活动地点不在旧金山,它就会自动订票,然后发给我。它知道订酒店,也知道我所有的航班和酒店偏好,全自动完成。我之前去东京参加 Code with Claude,之前还去了伦敦和柏林,它都帮我订好了所有票。那是多段航班加酒店,它全搞定了。我完全没参与。
There's this more sophisticated usage. One thing that I used to do is I would use co-work to book all of my travel. So what I would say is, 'Okay, here's my itinerary. I need to be here on this day, here on this day. Can you go out and book the plane tickets for me?' And it would open a browser and go to the travel site that we used to book stuff at Anthropic, and it would just fill it out and book the tickets. And what I've done now is I've actually automated it a little bit more. So what co-work does now is it has a scheduled task where every day it'll look through my email. It'll look for any events that I've accepted on my Google Calendar. If the event is in a different location, like not in San Francisco, it'll go and automatically book the tickets. And then it'll send it to me. It'll book the tickets, it knows to book hotels, and it knows all my flight and hotel preferences. And it'll just do this automatically. So I was just in Tokyo for Code with Claude, and then I was in London and Berlin for Code with Claude before that. And it just fully booked all the tickets for me. It was a multi-leg flight and hotel thing, and it just booked everything. I wasn't involved at all.
完全没参与?
At all?
完全没有。它就在我邮件里找到信息,我确认后它就订好了。太疯狂了。
Not at all. Yeah, it just found it in my email and then it booked it after I confirmed. Crazy.
好的。在座有多少人试过 co-work?只是好奇。其实不少,对吧?希望更多。太好了。另一个我收到很多私信的问题,想听听你的看法:我们很多人只用了 Fable 大概三天左右。在等待后续消息的同时,我猜你肯定用了更久,分享一下你用 Fable 编程的经验吧。你们现在有一系列不同的模型,你怎么考虑针对不同的软件工程用例使用不同的模型?
Okay. Yeah, I mean, how many people here has tried co-work? Just curious. Good amount, actually, right? Hopefully more. That's great. I think another question that I got a lot of pings about, would love to hear your take on this, is that many of us only got access to Fable for like maybe 3 days or so. So while we wait to hear what happens next, obviously I'm assuming you had access to them for much longer, share maybe some of your experiences with how you use Fable for coding and like, now you guys have a family of different models. How do you think about using different models for different software engineering use cases?
没错。首先关于 Fable,我想说从我们的角度看,这是一个误解,我们正在努力尽快恢复它。我对 Fable 的看法是,在编程方面,去年十一月有一个时刻。不知道大家是否记得,当时所有人都在 Twitter 上发帖,讨论这个模型在编程上变得多好。那是因为 Opus 4.5 发布了。从之前的模型到 Opus 4.5 的飞跃非常大,很多人第一次开始用 Claude 写所有代码。对我来说,那也是一个时刻,我直接卸载了 IDE,因为我不再用了。哇。所以对我来说,那是 Claude 开始写所有代码的时刻。从 Opus 4.5 到 Fable 的飞跃,我感觉至少是同样的大小。而且我认为在模型能力上可能是一个更大的飞跃。对我来说,Fable 有这种细微差别和维度,以及思考问题的方式,就像我最聪明的同事。它不像之前的模型那样是个粗钝的工具,不理解细微之处。它实际上能够深入处理问题。
Yeah, totally. And first of all, with Fable, just want to say from our point of view, it is a misunderstanding and we're working to hopefully get it back really soon. The way that I think about Fable is with coding, there was this moment back in November of last year. I don't know if everyone remembers this, but everyone started posting on Twitter and talking about how good the model had gotten at coding. And what happened is Opus 4.5 came out. And the leap from the previous model to Opus 4.5 was so big that for a lot of people they just started writing all of their code using Claude for the first time. And that was actually the moment for me where I just uninstalled my IDE because I wasn't using it anymore. Wow. And so for me that's the same moment when Claude started to write all the code. The leap from Opus 4.5 to Fable feels to me like a leap of at least that same size. And I think it might be even a larger leap in model capability. For me, Fable is it just has this nuance and dimensionality and kind of way of thinking about things that is similar to my smartest co-workers. It's not just a blunt instrument the way the previous models were where it doesn't understand nuance. It's actually really able to grapple with a problem.
这在数据分析等各种场景中都非常有用,因为其中有很多细微之处,你需要反复追问为什么才能触及问题的本质。Fable 天生就能做到这一点。在调试时也很有用,你需要形成假设并追踪下去,寻找证据。它在这方面表现得非常好。至于编程,我觉得我已经没有难题可以给它了。我实在想不出更难的问题。我给的每个问题,它基本上一次就能解决,或者稍微提示几次就能搞定。我真的想不出更难的问题了,而且很多团队成员也这么说。所以,这确实是一个很大的进步。
And this is useful for all sorts of things from data analysis, because there's a lot of nuance there and you have to ask why like three times to get to the bottom of a question. Fable just naturally does this. It's useful in debugging where you have to form a hypothesis and chase it down, look for evidence. It's able to do this really well. And for coding, I feel like I actually ran out of hard problems to give it. Like I just couldn't think of a harder problem. Every problem I gave it, it essentially one-shotted or maybe few-shotted with a little bit of prompting. I just kind of ran out of hard problems, and this is something I heard from a lot of the team also. So yeah, it's quite a big step.
Claude Code 现在基本上 100% 是由 Claude Code 自己编写的吗?
Is Claude Code just basically 100% written by Claude Code at this point?
是的。从去年 11 月开始就一直是这样的。
Yes. That has been the case again since November of last year.
哦,所以甚至不是用 Fable,而是用 Opus 4.5。
Oh, so it wasn't even with Fable, it was with Opus 4.5.
是的,用 Opus 4.5。Claude Code 是这样,co-work 也是这样,Anthropic 越来越多的产品都是如此。我认为在整个 Anthropic,平均 80% 到 90% 的代码是由 Claude Code 编写的。但实际上,对于越来越多的团队来说,这个比例就是 100%。
Yeah, with Opus 4.5. That's true for Claude Code, that's true for co-work, it's true for an increasing number of Anthropic's products. I think across all of Anthropic, 80 to 90% of the code is written by Claude Code on average. But actually, for a larger and larger share of teams, it's just 100%.
我还听说 Fable 更贵,也稍微慢一点,但就像你说的,它也非常智能。你是怎么用的?是所有任务都用 Fable,还是根据不同场景降级到 Opus 或 Sonnet?
I also heard that Fable is more expensive, but it's also a little slower, but it's also incredibly intelligent, like you said. How do you use it? Do you just use Fable for all your tasks, or do you drop down to Opus or Sonnet for different use cases?
是的,我所有事情都用 Fable。所有事情都用 Fable。
Yeah, I use Fable for everything. Just Fable for everything.
是因为 Anthropic 没有预算限制吗?
Is it because Anthropic has no budget?
其实我们确实会考虑 token 的使用。我们本质上是在生产 token,但它们对我们来说并不是免费的,因为我们每使用一个 token,就意味着少给客户一个 token。所以存在机会成本。当我思考这个问题时,它实际上可能归结为 ROI。所以当你考虑像 Fable 这样的 ROI 时,如果你使用 Fable 配合 advisor 模型,或者默认使用 Opus 并在需要时调用 Fable,你可以将投资减少 50% 左右。有各种优化使用的方法,随着新模型的出现,你需要不断调整优化方式。跟上这种变化并确保其良好运行其实需要不少工作,你必须运行评估来验证。虽然你可以使用 advisor,这是推荐的即用方法。但我实际上认为,如果你考虑 ROI,可能只有 50% 的机会减少投资,但可能有 1000% 甚至 10 万或 1 万倍的机会增加回报。所以我的想法是,直接使用最贵的模型,专注于如何从中获得更多价值,如何提高回报。不要专注于削减成本。这项技术的采用还处于非常早期的阶段,不值得过多考虑成本。你可能需要花一些精力在成本控制上,并且一定要确保成本可控且治理良好。但我建议你把几乎所有精力都放在提高回报上。
So, you know, we do actually think about token use, right? Like we essentially make the tokens. But they're not free for us because every token we use is a token we do not give to a customer. So there's an opportunity cost. When I think about it, it actually maybe comes back to ROI. So when you think about ROI for something like Fable, you can reduce the investment maybe 50% or something if you use Fable with an advisor model or use Opus by default and have a call out to Fable when needed. There are all sorts of ways to optimize usage, and as new models come out, you have to keep tuning the way you optimize this usage. It's actually quite a bit of work to keep up with it and to keep that you have to run evals to make sure it works pretty well. Although you can use advisors, that's the recommended out-of-the-box approach. But I actually think that if you think about ROI, there's probably a 50% chance to reduce the investment, but probably a thousand percent opportunity or even hundred thousand or ten thousand percent opportunity to increase the return. So what I would think about is just use the most expensive model and focus on how do I get more out of it? How do I increase the return? Do not focus on cost cutting. It's just so early in the adoption of this technology to think a lot about this. And you probably want to spend some effort on cost cutting, and you definitely want to make sure costs are under control and well-governed. But I would focus almost all your effort on increasing returns.
是的,本质上就是更关注上行空间,而不是潜在的下行风险。
Yeah, focus on the upside more than the potential downsides, essentially.
完全正确,因为现在的上行空间远远大于优化下行空间的空间。
Exactly, because there's so much more upside right now than there is room to optimize the downside.
是的。希望我们其他人也能尽快用上 Fable,这样我们也能探索这些可能性。好了,我们收到了不少问题。实际上有 107 个问题。我们显然不可能全部回答,但我会挑选并主持一些投票最高的问题。Andrew,这里投票最高的问题是:'Claude Code 在改善协作方面做了什么?目前它感觉像是一个单人工作产品,只有像 GitHub 这样糟糕的东西用于与他人协作。' 这是个有趣的问题。你怎么看?
Yeah. Hopefully the rest of us get access to Fable soon so we can all explore that as well. Okay, I think we got a decent amount of questions. Actually there are 107 questions. We're obviously not going to get through all of them, but I'm going to handpick and MC some of the top voted questions. So Andrew, the top voted question here is: 'What is Claude Code doing to improve collaboration? Currently it feels like a solo work product with only awful things like GitHub for working with others.' That's an interesting question. What do you think?
是的,我们正在筹备很多东西。所以希望很快能有一些成果。与此同时,我建议使用 MCP 将 Claude Code 连接到 Slack、Teams、G Chat 或任何你用于协作的工具。
Yeah, you know, we have a bunch of stuff in the fire that we're cooking up. So hopefully we'll have something there soon. In the meantime, the thing that I would recommend is use an MCP to hook up Claude Code to Slack or to Teams or to G Chat or whatever you use for cooperation.
听起来不错。下一个投票最高的问题是:'既然智能体现在能完成大部分编码工作,工程师应该重点放在哪里?' 之前我们稍微提到过,但还是很好奇。
Sounds good. The next voted question was: 'Since agents can do most of the code now, where should engineers focus most?' Kind of touched on this earlier, but yeah, curious.
是的。当我们思考工程师的工作时,编码是其中一部分,但我们还做很多非编码的工作。我们与客户交流,提出想法,与设计师和产品经理一起讨论,进行数据分析,决定下一步构建什么,与组织其他部分对齐。工程师要做各种各样的事情。我认为随着时间的推移,模型将能够比我们更好地完成所有这些事情。但我们现在还没到那一步。所以目前,模型负责编码,但需要有人来提示模型。而决定给模型什么提示,实际上涉及很多工作。你必须弄清楚下一步是什么。你必须做市场调研。你必须与团队沟通。你必须做所有这些其他工作。所以,是的,就是这些其他事情。
Yeah, yeah. So as we think about what engineers do, one of the things we do is coding, but we do a lot of stuff that's not coding. We talk to customers, come up with ideas, jam on stuff with designers and PMs, do data analysis, figure out what to build next, align with other parts of the org. There are all sorts of things we do as engineers. I think over time the model will be able to do all of these things better than we can. We are not there yet. So in the meantime, the model does the coding, but someone has to prompt the model. And figuring out what prompt to give the model, there's actually a lot that goes into that. You have to figure out what's the next thing. You have to do market research. You have to talk to your team. You have to do all this other work. So yeah, it's all these other things.
是的,而且我认为就像你说的,随着循环等机制,也在逐步提升抽象层,也许在某些情况下,你甚至不需要提示。所以,我觉得很有趣,就像我在这次会议上早些时候说的,我们正处于一个转型阶段。整个行业都在共同探索。我们在 Meta 也看到,就像你说的,普通软件工程师实际上只有少数时间花在编码上。其他都是上下游的事情,无论是管理代码部署,还是协作,在文档和规划中工作。所以,我认为这正是 Claude 和其他 AI 应用与产品思考如何加速整个开发生命周期的地方。
Yeah, and then I think like you said with loops, too, that's also progressively moving up the abstraction layer, where maybe at some point for some certain cases, you don't even actually need a prompt. And so, I think it's interesting, like I said earlier in this conference, is that we're in this transformational stage. We're all figuring this out together as an industry. We also have seen at Meta that like you said, it's actually a minority of an average software engineer's time that is actually spent coding. It's all the other things that happen upstream and downstream with it, whether it's managing your code deployments, or collaboration, working in docs and planning. And so, I think this is where Claude and other AI apps and products are thinking about how we can accelerate the whole development life cycle.
这是我们现在主要通过编码来解决的问题,而且我们会继续这样做,但开发流程的上游和下游部分仍然是我们正在探索的领域。
This is something that we're largely solving now with coding, and we're going to continue doing so, but then the upstream and downstream parts of the development lifecycle are still the things we're discovering.
对,对。而且编码,我觉得从来都只占少数时间,对吧?
Yeah, yeah. And like coding, I think was always the minority of the time, right?
确实,是的。
Been, yeah.
对我来说,有时候手动写一堆代码确实很有趣,但有时候又完全是苦差事。而且你知道,我一开始就不想手动做这些。
And for me, it's like sometimes it's really fun to code and write a bunch of code by hand, and then sometimes it's just a total slog. And you know, I never wanted to do it by hand in the first place.
嗯。
Yeah.
呃,我听到推特上的笑声了。
Um, I heard the Twitter laughing.
对,对。呃,所以你知道,有时候在飞机上没有 Wi-Fi 之类的时候,我还是会手动写代码,纯粹为了好玩。但这其实并不是我个人真正怀念的事情。对我来说,Claude Code 就像一个喷气背包,随着模型变得更好,我的喷气背包里就有了更多喷气引擎,我可以越来越快。而现在,我的瓶颈纯粹在于我提示的速度有多快。而且我现在大部分提示实际上只是音频,就像和 Claude 对话一样。瓶颈在于好点子。但编码已经不再是瓶颈了。
Yeah, yeah. Um, and so like, you know, sometimes when I don't have Wi-Fi on an airplane or something, I'll still write code by hand just for fun. But it's actually not something that I personally really miss. It feels to me like Claude Code is just like this jetpack, where as the model gets better, I get more jets or something in my jetpack, and I can go faster and faster. And at this point, I'm purely bottlenecked on how fast I can prompt. And most of my prompting now is actually just audio, like talking to Claude. And being bottlenecked on good ideas. But coding is just no longer the bottleneck.
没错。这正是很多对话在讨论的:软件工程的角色是什么?它越来越不关乎编码本身,而更多是关于你如何监督这些智能体,如何端到端地管理智能体,并帮助生成和产品化你的一些想法。我认为与此相关的是,随着 AI 生成代码的爆发,每个人都在用 AI 飞快地写代码。你或 Anthropic 如何处理下游影响?比如代码审查,对吧?我认为很多公司都要求人工代码审查。由于上游供应增加,这种范式可能正在被打破或转变。你们是怎么考虑的?
Right. And this is kind of what a lot of conversations have been about: what is the role of software engineering? It's less and less about the coding aspect and more about how you supervise these agents, how you manage the agent end-to-end, and help generate and productize some of the ideas that you have. I think related to that is, with the explosion of AI-generated code, everyone's just writing code so quickly now with AI. How do you or Anthropic handle the downstream impacts of that? So like with code review, right? I think a lot of companies require a human code review. That paradigm may or may not be actually breaking or shifting because of the increased supply from upstream. How are you thinking about that?
是的。当我们考虑代码时,本质上你写代码,然后它进入生产环境,然后希望它能驱动业务指标,比如收入或使用量,或者任何你关心的业务指标。当你思考这个过程时,有各种不同的瓶颈,而最大的瓶颈曾经是编码。我们现在已经解决了那个瓶颈,在 Anthropic 和许多已经采用 Claude Code 一段时间的客户那里,他们也达到了这个点。所以现在我们正在考虑下一个瓶颈,对我们来说,下一个瓶颈是代码审查。你写了大量代码,现在需要有人来审查。我们的答案是为此构建一个产品,所以我们构建了这个产品,叫做 Claude Code Review。它对所有人开放。你可以使用它。它和我们内部在 Anthropic 用于每个拉取请求的代码审查产品完全相同。它不同于市场上的任何其他代码审查产品,因为它贵得多。而它贵得多的原因是它使用了大量的 token 来完全自动化代码审查。所以当我作为工程师看到一个拉取请求时,基本上可以保证所有 bug 都已经被捕获了。虽然不是 100%,我们仍在改进,但大约能捕获 98% 到 99% 的 bug。所以当我查看代码时,基本上我不再寻找 bug 了。我知道没有 bug,因为 Claude 已经捕获并修复了它们。接下来要考虑的是:这个拉取请求应该存在吗?这是个好主意吗?我们的下一个瓶颈是安全审查,因为你运行所有这些代码,需要确保它是安全的。智能体可能像人一样引入漏洞。那么,如何确保代码既好又安全?我们的答案是 Claude 安全产品。同样,我们在内部构建了它来解决我们自己的问题,解决我们自己的瓶颈。它的作用是每周运行一次,扫描我们所有的代码库,发现问题,并自主修复它们。实际上,我们现在已经到了这样一个阶段:对于每一个大的新功能发布,我们都会进行红队测试和渗透测试,以确保安全。而且我们实际上已经达到了 Claude 安全能够捕获甚至渗透测试人员都没有发现的问题的程度。
Yeah. So when we think about code, essentially you write code and then it gets to production and then hopefully it drives a business metric for you like revenue or usage or whatever business metric you care about. When you think about this process, there are all sorts of different bottlenecks and the biggest bottleneck used to be coding. We've now solved that bottleneck and, at Anthropic and for a lot of our customers that have adopted Claude Code for a bit, they're getting to this point also. And so now we're thinking about the next bottleneck and, for us, the next bottleneck was code review. You write a lot of code, now someone has to review it. And our answer was to build a product for this and so we built this product, it's called Claude Code Review. It's available to anyone. You can use it. It is the exact same code review product that we use internally at Anthropic for every single pull request. It is different than any other code review product on the market because it's a lot more expensive. And the reason it's a lot more expensive is it uses a large number of tokens to fully automate the code review. So by the time that I see a pull request as an engineer, there's essentially a guarantee that all the bugs have been caught. And it's not 100%, we're still working on improving it, but it's like 98, 99% of the bugs. So when I look at the code, essentially I'm not looking for bugs anymore. I know there's no bugs because Claude caught it and Claude fixed it. The next thing is like, is this a pull request that should exist? Is this a good idea? The next bottleneck for us was security review because you're running all this code, you need it to be secure. Agents can introduce vulnerabilities the same way that people can. And so, how do you make sure the code is really good and secure? And the answer for us is the Claude security product. And again, we built this internally to solve our own problem, to solve our own bottleneck. And what it does is every week we run it and it scans over all of our codebases, it finds issues, and it fixes them autonomously. And we're actually at the point now where we do red teaming and we do penetration testing for every big new feature launch to make sure that it's secure. And we're actually getting to the point where Claude security is catching issues that even pen testers didn't catch.
哦,哇。
Oh, wow.
以前并非如此。我们拥有这个产品已经有一段时间了,但由于 Opus 4.8 模型,它现在开始达到那个水平。所以,这是下一个瓶颈。我们解决了它。同样,我们将其提供给我们的产品,所以客户也可以使用相同的产品。这就是 Claude 安全产品。现在我们正在考虑下一个瓶颈。下一个瓶颈可能是创意生成。也可能是优化 CI 以使其更好地扩展。例如,我昨晚做的一件事是我注意到我们的 CI 有点慢。所以我让 Claude Code 运行,我说:“使用工作流查看我的数据,从数据集中查看真实的 CI 时间,并优化 CI 使其更快。”那就是我的提示。那就是整个提示。
This didn't used to be the case. We've had this product for a bit, but because of the model with Opus 4.8, it's now starting to get to that point. And so, this was the next bottleneck. And so, we solved it. And again, we make this available to our product, so customers can use the same product, too. And this is just the Claude security product. And now we're thinking about the next bottleneck. And the next bottleneck might be idea generation. It might be optimizing CI to make it scale better. So, for example, a thing that I did last night is I noticed that our CI was a little bit slow. And so, what I did is I had Claude Code run and I said, 'Use a workflow to look at my data, look at real CI timings from the data set, and optimize CI to make it much faster.' That was my prompt. That was the entire prompt.
就这些?
That's it?
就这些。它使用了一个动态工作流,这是我们几周前推出的新功能,本质上模型动态编排数十、数百、数千个子智能体。这本质上是一种新的测试时算力形式。它用了大概几百万个 token,运行了几个小时。它产生了四个拉取请求,将 CI 时间减少了 50%。所以我昨晚合并了它们。而这项工作在过去需要几天、几周甚至几个月才能完成所有的性能分析,而这只是下一个瓶颈,我们也可以用 Claude 来解决它。
That's it. It used a dynamic workflow, which is this new feature that we launched a few weeks ago where essentially the model orchestrates dozens, hundreds, thousands of sub-agents dynamically. It's essentially a new form of test-time compute. And it used I think something like a few million tokens and it ran for a few hours. And it produced four pull requests that reduce the CI time by 50%. And so I landed those last night. And this is work that would have taken days or weeks or months in the past to do all this kind of profiling, and this is just the next bottleneck and we can use Claude for it, too.
太疯狂了。所以你基本上,我的意思是,你从我们开始就描述的方式就是不断追逐下一个瓶颈,不断追逐下一个瓶颈。到目前为止,基础模型只是不断拥有底层能力来逐步解决它。这真的很酷。
That's crazy. So you basically, I mean, the way you just described everything since we started is like just keep going after the next bottleneck, keep going after the next bottleneck. So far the base model just keeps having the underlying capabilities to be able to progressively solve it. So that's really cool.
我觉得这正好引出了我想问的下一个问题:你和你的团队下一步的大动作是什么?你对未来一年的愿景是什么?你刚才提到了遇到瓶颈就解决瓶颈,但 Claude Code 未来一年的愿景是什么?
I think it's actually a good segue into the next question I wanted to ask, which is what is the next big thing for you and your team? What is your vision for the next year? I mean, you kind of mentioned solving bottlenecks as you come across, but what is the vision for Claude Code over the next year?
嗯,先说一点:我们按周或月规划,没有一年计划。这个领域变化太快了,指数级增长太疯狂,你只能抓紧,一点点规划。总的来说,我们的方向和过去一两年非常相似:我们想成为最强大的智能体,想在任何地方工作。你的团队在哪里工作,Claude 就在哪里工作,你不需要切换到我们的全栈来使用它。我们想让人们以其他产品无法提供的方式体验新模型的能力。这是几年前我们意识到的:Sonnet 3.5 在编码上有了很大飞跃,但很少有产品能让你充分体验这一点。所以 Claude Code 就是为了实现这个目标。想法是:不再有源代码,你直接使用智能体,这就是体验方式。展望未来几个月和一年,模型在长周期任务上会做得更好。Claude 已经是长周期任务中最好的,而且我认为这个领先优势会扩大。代码会更安全、质量更高,模型的对齐也会更好。所以无论你作为用户——工程师、产品经理还是设计师——有什么意图,模型都能更好地表达你的意图。我们会关注这些能力,并构建产品让人们以最简单的方式体验它们。
Yeah, so as a caveat, we plan on a weekly or monthly cycle. We don't have a one-year plan. This space is changing too fast. The exponential is crazy. You just have to hang on and plan a little bit at a time. Broadly, the direction for us is very similar to the direction from before, from the last year or two. We want to be the most capable agent. We want to work anywhere. So wherever your team works, Claude works. You don't have to switch to our full stack to use it. And we want to let people experience the capabilities that the new model brings in a way that other products don't really let you experience. This is something we realized a couple years ago: Sonnet 3.5 was a big leap in coding, but there weren't many products that let you fully experience that. So Claude Code was a way to get at that. The idea was, no more source code, you're just using an agent. That's how you experience this. As we think about the model over the next few months and year, it's going to get even better at long-running work. Claude is already by far the best at long-running work, and that lead is going to increase, I think. The code is going to become more secure and higher quality. The model is going to become even better aligned. So whatever your intent is as the user—engineer, product manager, or designer—the model will be even better at expressing that intent. We're going to look at these capabilities and build products to let people experience them in the easiest way possible.
好的。Reza 问了一个我觉得很好的问题:大型项目中的编码不是最大的问题,维护才是。代码应该如何长期维护?你们是怎么处理维护的?
Nice. Reza asks a question here that I thought was really good too: coding in large-scale projects is not the biggest problem; maintenance is. How should the code be maintained in the long term? How do you guys deal with maintenance?
嗯,我一直在尝试的一个方法是用循环来做维护。
Yeah, so something I've been experimenting with is actually using loops for maintenance.
用循环做维护?能详细说说吗?
Loops for maintenance. Can you share more?
嗯,举个例子:你可以让 Claude Code 循环运行,检查代码库并改进架构,或者检查代码库找出测试套件不稳定的地方并改进,消除不稳定因素,或者找出没用的测试并删除,或者检查代码库中重复的抽象并统一成一个抽象。这些其实都是我一直在运行的循环。
Yeah, so one example is you can have Claude Code running in a loop to look at the codebase and improve the architecture, or look at the codebase and find places where the test suite is flaky and improve it to get rid of the flakes, or find instances of tests that are not useful and delete them, or look at the codebase for duplicated abstractions and unify them into a single abstraction. These are actually all loops that I have running.
那么在那些情况下,你是先检查代码再删除或修改,还是它们直接发 PR,你在循环运行后审查?
So in those cases, do you check or review what the code does before they delete or make any changes, or do they just send out PRs and you review them after they run the loop?
是第二种。我只看 PR。
Yeah, it's the second one. I just look at the PRs.
你只看它们修改后的变化。
You just look at the change after they made it.
没错,没错。对于这类问题,Claude 其实很容易理解这些结构性问题,所以通常效果很好。如果你用最新模型,结果不理想的话,你只需要说“寻找改进代码库质量的机会”,然后加上魔法词“使用工作流”。
Exactly, exactly. For problems like this, it's actually quite easy for Claude to wrap its head around these shape problems, so it's usually quite good. If you use the latest model and if the result isn't good, all you have to do is say, 'look for opportunities to improve the quality of the codebase' and then add the magic words 'use a workflow'.
我还以为你会说“不犯错误”。我以为那会是答案。
I thought you were going to say 'make no mistakes'. I thought that was going to be the answer.
但基本上,你说“使用工作流”,它就会投入更多的测试时算力,给你更好的结果。
But yeah, it's like essentially you say 'use a workflow' and it'll throw more test-time compute at it and give you a much better result.
这太酷了。工作流是个挺新的概念。我好像看到过相关公告。不知道在座有多少人熟悉它,但感觉是另一个值得关注的东西。能简单总结一下工作流和循环的区别吗?还是说它们差不多?
That's really cool. Workflows is a pretty new concept. I think I saw an announcement about it. I don't know how many people here are familiar with it, but it seems like another one we should be watching out for. Is there a quick summary of the difference between workflows and loops, or is it kind of the same thing?
嗯,区别很大。我的理解是:AI 中有传统的缩放定律。之前的缩放定律论文提出了一个观点:Transformer 和 LLM 的扩展方式取决于数据、神经网络大小和训练所用的算力。这种指数特性就是智能持续指数级增长的原因。过去几年,我们实际上增加了第四个因素:测试时算力。本质上,测试时算力就是模型生成 token 数量的一个花哨说法。有一种方法可以让模型高效地生成更多 token 以获得更好的结果。有几种方式可以实现。一种是努力程度设置。对于 Claude 模型,有低、中、高、极高、最高。这配置了你希望模型输出的 token 数量,调整了测试时算力行为。token 越多,结果越好。第二种方式,我们刚刚引入的,是动态工作流。它让 Claude 编写一个在虚拟机中运行的小程序,来协调其他 Claude 解决问题。这是一种我们仍在探索的新形式的测试时算力,但本质上,它是 Claude 启动几十、几百、几千个智能体来完成工作的一种方式。
Yeah, it's fairly different. The way I would think about it is: there are these traditional scaling laws in AI. The scaling laws paper a while back introduced the idea that the way transformers and LLMs scale is a function of data, the size of the neural network, and the amount of compute used to train it. This exponential property is why intelligence keeps increasing exponentially. Over the last couple years, we've actually added a fourth factor: test-time compute. Essentially, test-time compute is just a fancy way of saying how many tokens the model generates. There's a way to make the model productively generate more tokens to achieve a better outcome. There are a few ways to do this. One is effort settings. For Claude models, there are low, medium, high, extra high, max. This configures how many tokens you want the model to output, tuning the test-time compute behavior. More tokens, better result. The second way, which we just introduced, is dynamic workflows. This uses Claude to write a little program running in a virtual machine to orchestrate other Claudes to solve a problem. It's a new form of test-time compute we're still exploring, but essentially it's a way for Claude to launch dozens, hundreds, thousands of agents to get work done.
太厉害了。好的,下一个问题。Claude 有哪些难以解决的难题?
Wild. Awesome. Okay, next question. What are some of the hard problems that Claude has had a hard time solving?
嗯,我认为我们的模型并不完美,还有很多需要改进的地方。其中之一是产品直觉。我仍然能想出比 Claude 更好的产品创意。
Yeah, I think our models are not perfect, and there are many places where they still need to improve. One of them is product sense. I still come up with better product ideas than Claude does.
所以是创意生成。
So idea generation.
创意生成。是的,它还没达到那个水平。
Idea generation. Yeah, it's not there yet.
好的。
Okay.
它的代码现在比我写的代码还要好。
Its code is now better than the code I would have written.
它的前端设计比我的设计好。另一个我觉得自己仍然比 Fable 强很多的地方是分布式系统设计。比如思考有哪些服务?如何组织?数据如何流动?如何考虑负载因素等等?我认为 Fable 在这方面还有很多改进空间。
Its front end design is better than my design. Another place where I think I'm still a lot better than Fable is distributed system design. So, kind of thinking through like, what are the services? How do we organize it? How does data flow? How do we like think about load factors and things like that? This is a place where I think Fable still has a lot of opportunity to improve.
再过几个月这个就不成立了?还是几周?或者几天?
How many more months before that's not true anymore? Or weeks? Or days?
不,我不喜欢做预测,但我想大概到今年年底,它会变得相当不错。
No, I don't like giving predictions, but I would say by probably by the end of the year, it'll be quite good.
好的。我们听到了。
Okay. All right. We heard it here.
我们快到时间了,所以我大概用最后一个问题来结束。这是 Andrew 的问题,他问:“你如何防止工程师变得懒惰,接受 Claude 输出的所有内容?”
We're approaching the last few minutes, so I'm going to probably close this out with this last question here. This is a question from Andrew. He asks, "How do you prevent engineers from getting lazy and accepting everything they output in Claude?"
有几种思考方式。我认为这个问题可能有两部分。一部分是如何确保输出真的很好,并且人们在做正确的事。我们考虑的一个方法是让 Claude 做正确的事,这样你就不必操心了。一个例子是,从一开始我们就为 Claude Code 设置了权限提示。每当 Claude 想在电脑上运行命令时,它会问:“可以运行这个吗?是或否。”比如“我可以运行这个 bash 命令吗?我可以使用这个 MCP 吗?我可以在浏览器中获取这个 URL 吗?是或否。”工程师必须批准或拒绝。我们发现,随着时间的推移,你会变得有点懒,至少对我来说,我一直说“是”,并没有真正阅读命令。我相信很多人都是这样,不知道你是否会向老板承认。
There's a couple ways to think about it. I think there's maybe two parts to this question. One part is how do you make sure that the output is really good and that people are kind of doing the right thing. One way we think about this is how do we get Claude to do the right thing so you don't have to. An example of this is since the beginning we've had these permission prompts for Claude Code. Anytime Claude wants to run a command on your computer, it asks you, "Is it okay to run this? Yes or no." Like, "Can I run this bash command? Can I use this MCP? Can I fetch this URL in a browser? Yes or no." And an engineer sitting there has to approve it or decline it. Something we found is over time it felt like you get kind of lazy, at least for me I just kept saying yes. I wasn't really reading the commands. I'm sure this is true for a lot of people, I don't know if you admit this to your boss.
我们的安全人员看到了这一点,他们说:“嘿,这里有人类参与。人类参与的目的是提高安全性。但实际上,它正在损害安全性,因为人们有提示疲劳。他们只是不断说‘是’,而没有真正阅读细节。”所以这促使我们构建了自动模式,这是 Claude Code 中的一种新权限模式。我们在 Anthropic 内部使用它,现在绝大多数用户也在使用。具体是:每个权限提示都路由给一个模型,模型根据你在对话中说过的话来决定是或否。这不仅更安全——我们已经证明结果比危险模式更好,也比默认的是/否权限模式更好,因为提示疲劳——而且作为工程师,我少了一件事要做。因此,它解锁了长时间运行的智能体,因为你不需要坐在那里说是或否。我现在可以让 Claude 运行几个小时甚至几天。各种基准测试表明,Claude 在长时间运行任务方面是最好的。当然,多年的研究才使自动模式得以工作,因为如果你看 Claude 模型,它们基本上不再抵抗提示注入。它们不再容易受到提示注入的影响。这有时会让人惊讶,但如果你看我们的系统卡,100 次尝试的成功率大约为 1%。这绝对是业界最好的。当你把这个与提示注入分类器结合起来——我们现在对大部分流量运行这个分类器——模型基本上不再容易受到这种攻击。这使我们能够推出自动模式。所以作为工程师,我不需要坐在那里说是或否。这是问题的第一部分答案:让 Claude 做更多,并想办法安全地解锁 Claude 做更多,而不是试图控制自己。
Our security people actually saw this and they were like, "Hey, there's a human in the loop. The goal of that human in the loop was to improve security. But actually what's happening is that it is hurting security because people have this prompt fatigue. They just keep saying yes without actually reading the details." So this led us to build auto mode, which is a new permission mode in Claude Code. We use this internally at Anthropic and now the vast majority of our users use it too. What this is: every permission prompt is routed to a model and the model decides yes or no based on what you've said so far in that conversation. Not only is it more secure—we've shown the results are better than dangerous mode and better than the default yes/no permission mode because of prompt fatigue—but it's actually one less thing I need to do as an engineer. So it unblocks very long running agents because you don't have to sit there saying yes or no. I can now run Claude for hours or for days. There are all sorts of benchmarks that show Claude is the best at these kind of long-running tasks. Of course, years of research went into making auto mode work, because if you look at Claude models, they're essentially not resistant to prompt injection anymore. They're not susceptible to prompt injection anymore. This surprises people sometimes, but if you look at our system cards, the success rate at 100 attempts is around 1%. It's by far the best in the industry. When you take this and combine it with a prompt injection classifier, which we run for a large share of traffic now, essentially the models are not susceptible to this kind of attack anymore. This enables us to ship auto mode. So as an engineer, I don't have to sit there and say yes or no. That's part one answer to the question: let Claude do more and figure out how to unblock Claude to do more safely rather than trying to control yourself.
我认为第二件事是,当你不再写代码时,作为工程师的日常感受是什么?你如何继续学习?如何保持参与?我发现一个非常强大的工具是 Claude Code 中的输出风格。每当有新工程师加入我们团队,我们都会告诉他们使用探索性输出风格。就像/config output style equals exploratory。你在 Claude Code 中运行这个,或者让 Claude 为你设置。它的作用是,每当 Claude 做出更改时,它会向你解释:“嘿,这是架构的工作原理。如果你以前没用过这种语言,这是它的工作方式。这是代码库这部分的工作方式。”它会解释给你听,这样你就能学习。还有一种学习输出风格,适用于非编码人员。它会向他们解释:“嘿,这是这种语言的工作方式。在非常基础的层面上,它会教你如何做这件事,而不是直接做。”所以它会说:“好的,在 JavaScript 中,例如,这是某件事的工作原理。我不会为你做更改。我会带你一步步完成。第一步是打开这个文件并这样编辑。第二步是运行这个命令。好的,我看到你完成了。第三步是这样做。”所以我认为输出风格和更多地使用 Claude 是一个超级强大的学习工具。它帮助我作为工程师理解正在发生的事情,即使我们的技术栈变化、基础设施变化,尤其当我在使用新语言时。
I think the second thing is kind of like what is the feeling day-to-day as an engineer when you're not writing the code anymore? How do you still learn? How do you still stay in the loop? One thing I found really powerful for this is output styles in Claude Code. Whenever any new engineer joins our team, we tell them to use the exploratory output style. So, this is just like /config output style equals exploratory. You run this in Claude Code or you can ask Claude to set this for you. What it does is that anytime Claude will make a change, it'll explain to you, "Hey, here's how the architecture works. Here's how this language works if you haven't used it before. Here's how this part of the code base works." It'll explain to you so that you can learn. Then there's also a learning output style which you can use, and this is for people that are not coders. It'll explain to them, "Hey, here's how this language works. At a very basic level, it's going to teach you how to do the thing instead of doing the thing." So it'll say, "Okay, in JavaScript, for example, here's how something works. I'm not going to make the change for you. I'm going to walk you through it. Step one is open this file and edit it in this way. Step two is run this command. Okay, I see you done that. Step three is do this." So I think output styles and using Claude even more are a super powerful learning tool. It helps me as an engineer understand what's going on even as our stack changes, as our infra changes, and especially when I'm working in new languages.
太棒了。我们时间到了。正如我所说,这是我们行业非常有趣的时期,很荣幸邀请到 Boris 在这里分享他关于 Claude Code 的一些想法。让我们给他们鼓掌。非常感谢。
Awesome. Well, we're at time. It's been like I said, a very interesting time in our industry, and it's been a privilege to have Boris here working on Claude Code share some of his thoughts. And yeah, let's give a round of applause to them. Thank you so much.
嗯。
Yeah.