为六个月后的模型而建:Claude Code 的意外诞生

Building for the Model 6 Months from Now: The Accidental Creation of Claude Code

鲍里斯·切尔尼 Boris Cherny · Y Combinator · 2026-02-17 · 约 50 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Claude Code 的创造者 Boris Cherny 分享了他们如何为未来模型构建、在终端中的意外起步,以及为何产品的成功甚至让他们自己都感到惊讶。

Boris Cherny, creator of Claude Code, shares how they built for future models, the accidental start in a terminal, and why the product's success surprised even them.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

为未来模型构建 Building for the Future Model

Boris

在 Anthropic,我们的思路是:不为今天的模型构建,而为六个月后的模型构建。这实际上仍然是我给那些基于大语言模型创业的创始人的建议。试着想想,模型今天还不擅长的前沿领域是什么,因为它会变得擅长。

At Anthropic, the way that we thought about it is we don't build for the model of today. We build for the model six months from now. That's actually still my advice to founders that are building on LLM. Just try to think about what is that frontier where the model is not very good at today because it's going to get good at it.

Host

Claude Code 的所有代码都被一遍又一遍地重写。六个月前的 Claude Code 已经不存在了。你尝试一个东西,把它交给用户,和用户交流,学习,最终可能会得到一个好主意。有时则不然。你心里是否也在想,也许六个月后你就不需要那么明确地提示了?模型会足够好,自己就能搞定?

All of Claude Code has been written and rewritten and rewritten over and over. There is no part of Claude Code that was around 6 months ago. You try a thing, you give it to users, you talk to users, you learn, and then eventually you might end up at a good idea. Sometimes you don't. Are you also in the back of your mind thinking that maybe in 6 months you won't need to prompt that explicitly? Like the model will just be good enough to figure out on its own?

Boris

也许一个月后,就不再需要计划模式了。

Maybe in a month, no more need for plan mode in a month.

Host

天哪。欢迎来到另一期 Light Cone,今天我们有一位非常特别的嘉宾,Boris Cherny,Claude Code 的创始工程师。Boris,感谢你的到来。

Oh my god. Welcome to another episode of the Light Cone and today we have an extremely special guest, Boris Cherny, the creator engineer of Claude Code. Boris, thanks for joining us.

Boris

谢谢邀请。

Thanks for having me.

Host

感谢你创造了一个让我连续失眠大约三周的东西。我对 Claude Code 非常上瘾,感觉就像火箭助推器。对其他人来说,这种感觉是不是已经持续好几个月了?我想大概是十一月底,我的很多朋友都说有些东西变了。

Thanks for creating a thing that has taken away my sleep for about 3 weeks straight. I am very addicted to Claude Code and it feels like rocket boosters. Has it felt like this for people like for you know months at this point. I think it was like end of November is where a lot of my friends said like something changed.

Boris

我记得我第一次创建 Claude Code 时就有这种感觉,当时我还不知道自己是否发现了什么。我隐约觉得我发现了什么,然后就开始失眠了。连续三个月。那是 2024 年 9 月。是的,连续三个月。我没有休过一天假,周末也在工作,每晚都在工作。我当时就想,‘天哪,我觉得这会成为一个东西。但我不知道它是否有用,因为它还不能真正写代码。’

I remember for me I felt this way when I first created Claude Code and I didn't yet know if I was on to something. I kind of felt like I was on to something and then that's when I wasn't sleeping. And that was just like three straight months. This was September 2024. Yeah. It was like three straight months. I didn't take a single day vacation. Worked through the weekends. Worked every single night. I was just like, 'Oh my god, this is I think this is going to be a thing. I don't know if it's useful yet because it couldn't actually code yet.'

Host

如果你回顾那些时刻到现在,关于当下最让你惊讶的是什么?

If you look back on those moments to now, like what would be like the most surprising thing about this moment right now?

Boris

令人难以置信的是我们还在用终端。那本来只是起点,我没想到它会成为终点。第二点是它居然有用,因为一开始它并不真正写代码。即使在二月,当我们发布时,它可能只写了我的代码的 10% 左右。我并不真的用它来写代码,它做得不太好。我大部分代码还是手写的。所以事实上,我们的赌注得到了回报,它变得擅长我们认为它会擅长的事情,因为这一点并不明显。在 Anthropic,我们的思路是:不为今天的模型构建,而为六个月后的模型构建。这实际上仍然是我给那些基于大语言模型创业的创始人的建议。试着想想,模型今天还不擅长的前沿领域是什么,因为它会变得擅长,你只需要等待。

It's unbelievable that we're still using a terminal. That was supposed to be the starting point. I didn't think that would be the ending point. And then the second one is that it's even useful because at the beginning it didn't really write code. Even in February when we G it wrote maybe like 10% of my code or something like that. I didn't really use it to write code. It wasn't very good at it. I still wrote most of my code by hand. So the fact that it actually like our bets paid off and it got good at the thing that we thought it was going to get good at because it wasn't obvious. At Anthropic, the way that we thought about it is we don't build for the model of today. We build for the model 6 months from now. And that's actually still my advice to founders that are building on LLM is, you know, just try to think about what is that frontier where the model is not very good at today because it's going to get good at it and you just have to wait.

Host

不过往回说,你记得你第一次有这个想法是什么时候吗?能跟我们聊聊吗?是灵光一闪,还是你脑海中的第一个版本是什么样的?

Going back though, but when do you remember when you first got the idea? Can you just talk us through that? Like was it some like a spark or what was even the first version of it in your mind?

Boris

你知道吗,这很有趣。它非常偶然,就这样演变而来。在 Anthropic,我认为 Anthropic 长期以来一直押注编程,认为通往安全 AGI 的道路是通过编程。这个想法一直存在,实现它的方法是先教模型如何编程,然后教它如何使用工具,再教它如何使用计算机。你可以看到这一点,因为我加入 Anthropic 的第一个团队叫做 Anthropic Labs 团队,它产出了三个产品:Claude Code、MCP 和桌面应用。所以你可以看到它们是如何交织在一起的。我们构建的这个特定产品,没有人要求我构建一个 CLI。我们隐约知道也许是时候构建某种编程产品了,因为模型似乎已经准备好了,但还没有人真正构建出利用这种能力的产品。所以仍然有一种强烈的产品过剩感。但在当时,情况更加疯狂,因为还没有人构建过这个。所以我开始捣鼓,我想,‘好吧,我们构建一个编程产品。我首先需要做什么?我需要了解如何使用 API,因为那时我还没有用过 Anthropic 的 API。’于是我就构建了一个小型的终端应用来使用 API。仅此而已。那是一个小型聊天应用,因为想想当时的 AI 应用,对于今天的非程序员来说,大多数人使用的只是一个聊天应用。所以我构建了它。它在终端里。我可以提问,得到回答。然后我认为工具使用出现了。我只是想试试工具使用,因为我不太理解这是什么。我想,‘工具使用很酷。这真的有用吗?可能没有。让我试试看。’

You know, it's funny. It was so accidental that it just kind of evolved into this. At Anthropic, I think for Anthropic the bet has been coding for a long time and the bet has been the path to safe AGI is through coding. And this has kind of always been the idea and the way you get there is you teach the model how to code, then you teach it how to use tools, then you teach it how to use computers. And you can kind of see that because the first team that I joined at Anthropic was called the Anthropic Labs team and it produced three products: Claude Code, MCP, and the desktop app. So you can kind of see how these weave together. The particular product that we built, no one asked me to build a CLI. We kind of knew maybe it was time to build some kind of coding product because it seemed like the model was ready, but no one had yet really built the product that harnessed this capability. So there's still this insane feeling of product overhang. But at the time it was even crazier because no one had built this yet. And so I started hacking around and I was like, 'Okay, we build a coding product. What do I have to do first? I have to understand how to use the API because I hadn't used Anthropic API at that point.' And so I just built like a little terminal app to use the API. That's all that I did. And it was a little chat app because you think about the AI applications of the time and for non-coders today, most what are most people using is just a chat app. So that's what I built. And it was in a terminal. I can ask questions. I give answers. Then I think tool use came out. I just wanted to try out tool use because I don't really understand what this is. I was like, 'Tool use is cool. Is this actually useful? Probably not. Let me just try it.'

Host

你在终端里构建它,只是因为这是让东西跑起来最简单的方式。

You built it in terminal just because it was the easiest way to get something up and running.

Boris

是的。因为我不需要构建用户界面。

Yes. Because I didn't have to build a UI.

Host

当时就我一个人。像 IDE、Cursor、Windsurf 都在兴起。你有没有感到压力,或者收到很多建议,比如我们应该把它做成插件或者一个功能完整的 IDE?

It was just me at that point. Like the IDEs, Cursor, Windsurf taking off. Were you sort of under any pressure or getting lots of suggestions of, hey, like we should build this out as a plugin or as a fully featured IDE itself?

Boris

没有压力,因为我们甚至不知道我们想构建什么。团队只是在探索模式。我们隐约知道想在编程方面做点什么,但具体是什么并不明确。没有人有足够的信心。这就像我的工作,去弄清楚。所以我给了模型 bash 工具。那是我给它的第一个工具,因为我认为那正是我们文档中的例子。我直接用了那个例子。它是 Python 写的。我把它移植到了 TypeScript,因为我是用 TypeScript 写的。你知道,我不知道模型能用 bash 做什么。所以我让它读一个文件。它能 cat 文件。这很酷。然后我想,‘好吧,你还能做什么?’我问它,‘我在听什么音乐?’它写了一些 AppleScript 来操控我的 Mac,查找我的音乐播放器里的音乐。天哪。这是 Sonnet 3.5。你知道,我没想到模型能做到这一点。那是我第一次有 AGI 时刻,我当时就想,‘天哪,模型就是想使用工具。’

There was no pressure because we didn't even know what we wanted to build. The team was just in explore mode. We knew vaguely we wanted to do something in coding, but it wasn't obvious what. No one was high confidence enough. That was like my job to figure out. And so I gave the model the bash tool. That was the first tool that I gave it just because I think that was literally the example in our docs. I just took the example. It was in Python. I just ported it to TypeScript because that's how I wrote it. You know, I didn't know what the model could do with bash. So I asked it to read a file. It could cat the file. So that was cool. And then I was like, 'Okay, what can you actually do?' And I asked her, 'What music am I listening to?' He wrote some AppleScript to script my Mac and look up the music in my music player. Oh my god. And this was Sonnet 3.5. And you know, I didn't think the model could do that. And that was my first ever AGI moment where I was just like, 'Oh my god, the model just wants to use tools.'

CLI 的意外成功 Accidental success of the CLI

Host

这挺有意思的。Claude Code 以如此优雅简洁的形式工作,这很反直觉。终端已经存在很久了,这似乎是一个很好的设计约束,带来了很多有趣的开发者体验。它不像在工作,感觉就像在玩。我不用去想文件在哪里,而这几乎是偶然发生的。

That's kind of fascinating. I mean it's very kind of contrarian that Claude Code works so well in such an elegant simple form factor. I mean terminals have been around for a really long time and that seemed to be like a good design constraint that allowed a lot of interesting developer experiences. It doesn't feel like working. It just feels fun as a developer. I don't think about files where everything is and that came by accident almost.

Boris

是的,这是个意外。我记得终端在内部开始流行起来之后。说实话,在构建这个东西之后,大概在第一个原型两天后,我就开始给我的团队做内部测试,因为如果你想到一个主意,而且它看起来有用,你首先想做的就是把它给别人看看他们怎么用。然后第二天我进来,坐在我对面的另一位工程师 Robert,他的电脑上已经有了 Claude Code,他正在用它来编程。我说:'你在干什么?这东西还没准备好。它只是个原型。'但确实,在那个形态下它已经很有用了。我记得我们做发布评审,准备对外发布 Claude Code,那是在 2024 年 12 月或 11 月左右。Dario 问:'内部使用图表是垂直增长的。你是不是强迫工程师使用它?为什么你要强制他们?'我说:'不,不,我们没有。我只是发了个帖子,他们就开始互相推荐了。'说实话,这完全是偶然的。我们从 CLI 开始是因为它最便宜,然后它就一直在那里了。

Yeah, it was an accident. I remember after the terminal started to take off internally. Honestly, after building this thing, I think like 2 days after the first prototype, I started giving it to my team just for dogfooding, because if you come up with an idea and it seems useful, the first thing you want to do is give it to people to see how they use it. Then I came in the next day and Robert, who sits across from me, another engineer, he just had Claude Code on his computer and he was using it to code. I was like, 'What are you doing? This thing isn't ready. It's just a prototype.' But yeah, it was already useful in that form factor. I remember when we did our launch review to kind of launch Claude Code externally, this was in December, November, something like that in 2024. Dario asked, 'The usage chart internally is like vertical. Are you forcing engineers to use it? Why are you mandating them?' And I was just like, 'No, no, we didn't. I just posted about it and they've been telling each other about it.' Honestly, it was just accidental. We started with the CLI because it was the cheapest thing and it just kind of stayed there for a bit.

早期用例与潜在需求 Early use cases and latent demand

Host

那么在 2024 年那段时间,工程师们是怎么使用它的?他们已经开始用它来发布代码了吗,还是以不同的方式使用?

So in that 2024 period, how were the engineers using it? Were they sort of shipping code with it yet or were they using it in a different way?

Boris

模型在编程方面还不太好。我个人用它来自动化 git。我想现在我已经忘了大部分 git 命令,因为 Claude Code 已经做了这么久。但确实,自动化 bash 命令是一个非常早期的用例,还有操作 Kubernetes 之类的事情。人们也在用它编程。所以有一些早期迹象。我认为第一个用例实际上是编写单元测试,因为风险较低,而且模型在这方面还很差,但人们正在摸索,并学会如何使用这个东西。我们注意到一件事:人们开始为自己编写 markdown 文件,然后让模型读取那个 markdown 文件。这就是 Claude MD 的由来。对我来说,产品中最重要的原则可能是潜在需求。在最初的 CLI 之后,这个产品的每一个部分都是通过潜在需求构建的。Claude MD 就是一个例子。还有另一个我觉得可能有趣的一般原则:你可以为模型构建,然后在模型周围构建脚手架来稍微提升性能,根据领域不同,性能可能提升 10-20%,但基本上这种提升会被下一个模型抹平。所以你要么构建脚手架,获得一些性能提升,然后重新构建,要么就等下一个模型,然后免费获得提升。Claude MD 和脚手架就是这样的例子。我认为这就是我们坚持使用 CLI 的原因,因为我们觉得我们构建的任何 UI 在 6 个月内都会过时,因为模型改进得太快了。

The model is not very good at coding yet. I was using it personally for automating git. I think at this point I probably forgotten most of my git because Claude Code has just been doing it for so long. But yeah, automating bash commands was a very early use case, and operating Kubernetes and things like that. People were using it for coding. So there were some early signs of this. I think the first use case was actually writing unit tests because it's a little bit lower risk and the model was still pretty bad at it, but people were figuring it out and figuring out how to use this thing. One thing that we saw is people started writing these markdown files for themselves and then having the model read that markdown file. And this is where Claude MD came from. Probably the single biggest principle in product for me is latent demand. Every bit of this product is built through latent demand after the initial CLI. Claude MD is an example of that. There's another general principle that I think is maybe interesting: you can build for the model and then you can build scaffolding around the model in order to improve performance a little bit, and depending on the domain you can improve performance maybe 10-20% something like that, and then essentially the gain is wiped out with the next model. So either you can build the scaffolding and then get some performance gain and then rebuild it again, or you just wait for the next model and then you kind of get it for free. Claude MD and the scaffolding is an example of that. I think that's why we stayed in the CLI, because we felt there is no UI we could build that would still be relevant in 6 months because the model was improving so quickly.

Claude MD 内容与理念 Claude MD content and philosophy

Host

之前我们说应该比较一下 Claude MD,但你说了一句很有深意的话,说你的其实非常短,这几乎和人们预期的相反。为什么?你的 Claude MD 里有什么?

Earlier we were saying like we should compare Claude MDs, but you said something very profound which is yours is actually very short, which is almost the opposite of what people might expect. Why is that? What's in your Claude MD?

Boris

好的,我来之前查了一下。我的 Claude MD 有两行。第一行是:每当你提交 PR 时,启用自动合并。这样一旦有人批准,它就会被合并。这样我就可以继续编码,不用来回处理代码审查之类的。第二行是:每当我提交 PR 时,把它发布到我们内部的团队印章频道。这样有人可以盖章,我就可以不被阻塞。关键是所有其他指令都在我们代码库中的 Claude MD 里,整个团队每周都会多次贡献。我经常看到有人提交 PR,他们犯了一些完全可以避免的错误,我就会直接在 PR 上 @Claude。我会写'添加 Claude,把这个加到 Claude MD 里',我每周会做很多次。

Okay, so I checked this before we came. My Claude MD has two lines. The first line is: whenever you put up a PR, enable automerge. So as soon as someone accepts it, it's merged. That's just so I can code and I don't have to go back and forth with CR or whatever. And the second one is: whenever I put up a PR, post it in our internal team stamps channel. Just so someone can stamp it and I can get unblocked. The idea is every other instruction is in our Claude MD that's checked into the codebase and it's something our entire team contributes to multiple times a week. Very often I'll see someone's PR and they make some mistake that's totally preventable and I'll just literally tag Claude on the PR. I'll just do like 'add Claude, add this to the Claude MD' and I'll do this many times a week.

Host

你需要压缩 Claude MD 吗?我肯定遇到过这种情况,顶部出现消息说你的 Claude MD 现在有几千个 token。你们遇到这种情况会怎么做?

Do you have to compact the Claude MD? Like I definitely reached a point where I got the message at the top saying your Claude MD is like thousands of tokens now. What do you do when you guys hit that?

Boris

我们的 Claude MD 其实很短。我想大概几千个 token 吧。如果你遇到这种情况,我的建议是删除你的 Claude MD,然后重新开始。

So our Claude MD is actually pretty short. I think it's like a couple thousand tokens maybe something like that. If you hit this, my recommendation would be delete your Claude MD and just start fresh.

Host

有意思。

Interesting.

Boris

我认为很多人试图过度设计这个,而实际上能力随着每个模型而变化。所以你想要的是做最少的事情来让模型走上正轨。所以如果你删除了 Claude MD,然后模型偏离了轨道,做了错误的事情,那时你再一点一点地加回来。你可能会发现,随着每个模型,你需要添加的东西越来越少。对我来说,我自认为是一个相当普通的工程师。我不使用很多花哨的工具。我不使用 Vim。我使用 VS Code,因为它更简单。

I think a lot of people try to overengineer this, and really the capability changes with every model. So the thing that you want is do the minimal possible thing in order to get the model on track. So if you delete your Claude MD and then the model gets off track, it does the wrong thing. That's when you kind of add back a little bit at a time. What you're probably going to find is with every model, you have to add less and less. For me, I consider myself a pretty average engineer to be honest. Like I don't use a lot of fancy tools. I don't use Vim. I use VS Code because it's simpler.

终端与 GUI 偏好 Terminal vs GUI preferences

Host

等等,真的吗?我本以为因为你是在终端里构建这个的,你会是个死忠终端用户,比如只用 Vim,管他什么 VS Code 的人。

Wait, really? I would have assumed that because you built this in the terminal that you were sort of a diehard terminal like Vim-only person, you know, screw those VS Code people.

Boris

嗯,我们团队里有这样的人。比如 Adam Wolf,他在团队里,他说‘除非我死了,否则别想从我手里夺走 Vim’。所以团队里确实有很多这样的人。这是我早期学到的一件事:每个工程师都喜欢用不同的开发工具,没有一种工具适合所有人。但我觉得这也是 Cursor 能做得这么好的原因之一,因为我把它想成:我会用什么产品,什么对我有意义?所以用 Cursor,你不需要懂 Vim,不需要懂 tmux,不需要知道怎么 SSH,不需要懂那些东西。你只需要打开工具,它会引导你,它会做所有这些事。

Well, we have people like that on the team. There's Adam Wolf, for example. He's on the team and he's like, 'You will never take Vim from my cold dead hands.' So there are definitely a lot of people like that on the team. And this is one of the things I learned early on: every engineer likes to hold their dev tools differently. They like to use different tools. There's just no one tool that works for everyone. But I think this is also one of the things that makes it possible for Cursor to be so good, because I kind of think about it as: what is the product that I would use that makes sense to me? So to use Cursor, you don't have to understand Vim, you don't have to understand tmux, you don't have to know how to SSH, you don't have to know all that stuff. You just have to open up the tool and it'll guide you, it'll do all this stuff.

终端输出的冗长性 Verbosity in terminal output

Host

你怎么决定终端要有多详细?有时候你得按 Control-O 查看一下。内部会不会有关于长短的争论?每个用户可能都有自己的看法。你怎么做这些决定?

How do you decide how verbose you want the terminal to be? Sometimes you have to go Control-O and check it out. Is it like internal bike shed battles around longer or shorter? I mean, every user probably has an opinion. How do you make those sorts of decisions?

Boris

你怎么看?现在是不是太啰嗦了?

What's your opinion? Is it too verbose right now?

Host

哦,我喜欢这种详细,因为有时候它突然变得很深入,我看着就能快速阅读,然后说‘哦不,不是那样’。然后我按退出键停止它,它就在 bug 出现时直接阻止了整个 bug 农场。我是说,那通常是我没有正确使用计划模式的时候。

Oh, I love the verbosity because sometimes it just goes off the deep end and I'm watching, and then I can just read very quickly and it's like, 'Oh no, no, it's not that.' And then I escape and just stop it, and it just stops an entire bug farm as it's happening. I mean, that's usually when I didn't do plan mode properly.

Boris

这是我们经常改变的东西。我记得早期,大概六个月前,我试图在内部去掉 bash 输出,只是总结一下,因为我觉得‘那些又长又大的 bash 命令,我其实不在乎’。然后我把它给 Anthropic 的员工用了一天,所有人都反对。‘我想看到我的 bash 输出,因为它实际上很有用。’对于像 git 输出这样的东西,可能没用,但如果你在运行 Kubernetes 任务之类的,你确实想看到它。我们最近隐藏了文件读取和文件搜索。所以你会注意到,它不再说‘读取 foo.md’,而是说‘读取了一个文件,搜索了一个模式’。我认为这在六个月前是不可能发布的,因为模型还没准备好。它仍然会经常读错东西。作为用户,你仍然需要在旁边捕捉和调试。但现在,我注意到它几乎每次都在正确的轨道上。而且因为它大量使用工具,总结一下实际上更好。但我们发布后,我们自己内部用了一个月,然后 GitHub 上的人不喜欢。所以有一个大问题,人们说‘不,我想看到细节’。那真是很好的反馈。所以我们添加了一个新的详细模式,在 /config 里你可以启用详细模式,如果你想看到所有文件输出,你可以继续那样做。然后我在那个问题上发了帖,人们还是不喜欢,这又很棒,因为我最喜欢的就是听到人们的反馈,听到他们实际上想怎么用。所以我们不断迭代,让它变得更好,成为人们想要的东西。

This is something we probably change pretty often. I remember early on, maybe six months ago, I tried to get rid of bash output internally, just to summarize it, because I was like, 'These giant long bash commands, I don't actually care.' And then I gave it to Anthropic employees for a day and everyone just revolted. 'I want to see my bash output because it actually is quite useful.' For something like git output, maybe it's not useful, but if you're running Kubernetes jobs or something like this, you actually do want to see it. We recently hid the file reads and file searches. So you'll notice instead of saying 'read foo.md', it says 'read one file, searched one pattern'. And this is something I think we could not have shipped six months ago because the model just was not ready. It would have still read the wrong thing pretty often. As a user, you still had to be there and kind of catch it and debug it. But nowadays, I just noticed it's on the right track almost every time. And because it's using tools so much, it's actually a lot better just to summarize it. But then we shipped it, we dogfooded it for like a month, and then people on GitHub didn't like it. So there was a big issue where people were like, 'No, I want to see the details.' And that was really great feedback. So we added a new verbose mode, and that's just in /config you can enable verbose mode, and if you want to see all the file outputs you can continue to do that. And then I posted on the issue and people still didn't like it, which is again awesome because my favorite thing in the world is just hearing people's feedback and hearing how they actually want to use it. So we just iterated more and more to get that really good and to make it the thing that people want.

用 AI 修复 Bug Bug fixing with AI

Host

我很惊讶我现在有多喜欢修 bug。你只需要有很好的日志,然后甚至只要说‘嘿,检查那个特定对象,它这样搞砸了’,它就会搜索日志。它弄清楚一切。它可以进入你的生产隧道,查看你的生产数据库。这太疯狂了。修 bug 就是去 Sentry,复制 markdown。很快它就会直接是 MCP。就像自动修 bug 和生成测试……他们管那叫什么新词?像制造一个创业工厂。哦对。

I'm amazed how much I enjoy fixing bugs now. And then all you have to do is have really good logging and then even just say, 'Hey, check out that particular object, it messed up in this way,' and it searches the log. It figures everything out. It can go into your production tunnel and look at your production DB for you. It's like this is insane. Bug fixing is just going to Sentry, copy markdown. Pretty soon it's just going to be straight MCP. It's like an auto bug fixing and test making sort of... what's the new term they call it? Like a making a startup factory. Oh yeah.

Boris

对。现在有所有这些概念,而不是必须审查代码。我是老派,所以我喜欢详细。我喜欢说‘哦,你在做这个,但我想要你做那个。’但现在有一种完全不同的思想流派,说任何时候真人必须看代码,那都是不好的。

Right. There's like all these concepts now of rather than having to review the code. I'm old school, so I like the verbosity. I like to say, 'Oh, well, you're doing this, but I want you to do that.' But there's a totally different school of thought now that says anytime a real human being has to look at code, that's bad.

Host

是啊。是啊。是啊。

Yeah. Yeah. Yeah.

Boris

这很吸引人。

Which is fascinating.

Host

我觉得 Dan Chipper 经常谈到这个,当你看到模型犯错时,试着把它放到 .cursorrules 里,放到技能里之类的,这样它就可以复用。但我认为有一个元问题我实际上很纠结。人们说智能体可以做这个,可以做那个,但实际上智能体能做什么随着每个模型而变化。所以有时候有新人加入团队,他们实际上比我更会使用 Cursor。

I think Dan Chipper talks about this a lot as kind of when you see the model make a mistake, try to put it in the .cursorrules, try to put it in skills or something like that so it's reusable. But I think there's this meta point that I actually struggle with a lot. And people talk about like agents can do this, agents can do that, but actually what agents can do changes with every single model. And so sometimes there's a new person that joins the team and they actually use Cursor more than I would have used it.

Boris

我对此一直感到惊讶。比如,有一个内存泄漏,我们试图调试它。顺便说一句,Jared Sumar 一直在消灭所有内存泄漏,这太棒了。但在 Jared 加入团队之前,我必须自己做这个,有一个内存泄漏。我试图调试它。所以我拿了一个堆转储,在 DevTools 里打开,查看配置文件,然后查看代码,我试图弄清楚。然后团队里的另一个工程师 Chris,他只是问了 Cursor。他说‘嘿,我觉得有个内存泄漏。你能运行这个并试着找出来吗?’然后 Cursor 拿了堆转储,为自己写了一个小工具来分析堆转储,然后它比我更快地找到了泄漏。这是我必须不断重新学习的东西,因为我的大脑有时还卡在六个月前。

And I'm just constantly surprised by this. For example, there was a memory leak and we were trying to debug it. By the way, Jared Sumar has just been on this crusade killing all the memory leaks and it's been amazing. But before Jared was on the team, I had to do this and there was this memory leak. I was trying to debug it. So I took a heap dump, opened it in DevTools, looked through the profile, then looked through the code, and I was just trying to figure this out. And then another engineer on the team, Chris, he just asked Cursor. He was like, 'Hey, I think there's a memory leak. Can you run this and try to figure it out?' And Cursor took the heap dump, wrote a little tool for itself to analyze the heap dump, and then it found the leak faster than I did. And this is just something I have to constantly relearn because my brain is still stuck somewhere six months ago at times.

给技术创始人的建议 Advice for technical founders

Host

那么对于技术创始人来说,有什么建议可以让他们真正成为最新模型发布的最大化主义者?听起来刚毕业的人或者没有先入为主观念的人可能比那些长期工作的工程师更适合。

So what would be some advice for technical founders to really become maximalists at the latest model release? It sounds like people fresh off of school or that don't have any assumptions might be better suited than maybe sometimes engineers who have been working at it for a long time.

专家进阶与招聘标准 How experts get better and hiring criteria

Host

专家们是如何变得更好的?

And how do the experts get better?

Boris

我认为对自己来说,这是一种初学者心态和谦逊。我觉得工程师这个学科,我们学会了持有非常强烈的观点,高级工程师往往会因此得到奖励。在我以前在大公司的工作中,当我招聘架构师这类工程师时,你会寻找那些经验丰富、观点非常坚定的人。但事实上,很多东西已经不再相关,很多观点也应该改变,因为模型在变得更好。所以我认为最重要的技能是那些能够科学思考、从第一性原理出发的人。

I think for yourself it's kind of beginner mindset and humility. I feel like engineers as a discipline we've learned to have very strong opinions and senior engineers are kind of rewarded for this. In my old job at a big company, when I hired architects and this kind of engineer, you look for people that have a lot of experience and really strong opinions. But it actually turns out a lot of this stuff just isn't relevant anymore and a lot of these opinions should change because the model is getting better. So I think actually the biggest skill is people that can think scientifically and can just think from first principles.

Host

你现在招聘团队成员时,如何筛选这种能力?

How do you screen for that when you try to hire someone now for your team?

Boris

我有时会问:你犯错的例子是什么?这是个很好的问题。一些经典的行为问题,甚至不是编程问题,我觉得很有用,因为你可以看到人们能否事后认识到自己的错误,能否承认错误并从中学习。很多非常资深的人,尤其是一些创始人类型,其实很擅长这一点。但其他人有时永远不会为错误承担责任。就我个人而言,我大概有一半时间是错的。我一半的想法是糟糕的,你只能去尝试。你尝试一件事,把它交给用户,和用户交流,学习,然后最终可能会得到一个好主意。有时不会。这种技能过去对创始人非常重要,但现在我认为对每个工程师都非常重要。

I sometimes ask about what's an example of when you're wrong. It's a really good one. Some of these classic behavioral questions, not even coding questions, I think are quite useful because you can see if people can recognize their mistake in hindsight, if they can claim credit for the mistake and if they learn something from it. I think a lot of these very senior people, especially some founder types, are actually quite good at it. But other people sometimes will never take the blame for a mistake. For me personally, I'm wrong probably half the time. Half my ideas are bad and you just have to try stuff. You try a thing, you give it to users, you talk to users, you learn, and then eventually you might end up at a good idea. Sometimes you don't. This is the skill that in the past was very important for founders, but now I think it's very important for every engineer.

Host

你会不会根据某人使用 Claude Code 与智能体协作的转录来招聘?因为我们正在这样做。我们刚刚添加了一个测试功能,你可以上传你用 Claude Code 或 Codex 编写功能的转录。我个人认为这行得通。你可以看出一个人如何思考,他们是否查看日志,能否在智能体偏离轨道时纠正它?他们是否使用计划模式?使用计划模式时,他们是否确保有测试?所有这些不同的事情。他们是否考虑系统?他们甚至理解系统吗?这里面蕴含了太多信息。我想象一个蜘蛛网图,就像 NBA 2K 那样的游戏里,哦,这个人投篮或防守很厉害。你可以想象一个关于某人 Claude Code 技能水平的蜘蛛网图。

Do you think you would ever hire someone based on the Claude Code transcript of them working with the agent? Because we're actively doing that right now. We just added, as a test, you can upload a transcript of you coding a feature with Claude Code or Codex or whatever it is. Personally, I think it's going to work. You can figure out how someone thinks, whether they're looking at the logs or not, can they correct the agent if it goes off the rails? Do they use plan mode? When they use plan mode, do they make sure that there are tests? All of these different things. Do they think about systems? Do they even understand systems? There's just so much that's embedded in that. I imagine I just want a spiderweb graph, like in those video games like NBA 2K. It's like, oh, this person's really good at shooting or defense. You could imagine a spiderweb graph of someone's Claude Code skill level.

Boris

是啊。技能会是什么?会是哪些?

Yeah. What would the skills be? What would be those?

Host

我觉得像是系统测试、用户行为。肯定有设计部分、产品感,也许还有自动化。我在 Claude Code 中最喜欢的一点是,我有一个功能,要求对每个计划判断它是过度工程化、不足工程化还是恰到好处,并说明原因。

I think it's like systems testing, user behavior. There's got to be a design part, product sense, maybe also just automating stuff. My favorite thing in Claude Code for me is I have a thing that says for every plan decide whether it's overengineered, underengineered, or perfectly engineered and why.

Boris

我认为这也是我们正在试图弄清楚的。当我观察团队中我认为最有效的工程师时,基本上有两种类型,非常两极分化。一边是极端专家。Jared 就是一个很好的例子,Bun 团队也是一个很好的例子。超级专家。他们比任何人都更懂开发工具。他们比任何人都更懂 JavaScript 运行时系统。另一边是超级通才,团队的其他人大致属于这一类。很多人横跨产品和基础设施,或产品和设计,或产品和用户研究、产品和业务。我真的很喜欢看到那些做奇怪事情的人。我认为这在过去是一个警告信号,因为问题是:这些人真的能构建有用的东西吗?那是一个极限测试。但如今,例如团队中的一位工程师 Daisy,她之前在不同的团队,然后转到了我们团队。我希望她转过来的原因是,她加入几周后就为 Claude Code 提交了一个 PR,这个 PR 是为 Claude Code 添加一个新功能。她没有直接添加功能,而是先提交了一个 PR,给 Claude Code 一个工具,让它能够测试任意工具并验证其工作。然后她提交了那个 PR,接着让 Claude 自己编写工具,而不是自己实现。我认为这种跳出框框的思维非常有趣,因为还没有很多人理解这一点。我们使用 Claude Agents SDK 来自动化开发的几乎每个部分。它自动化代码审查、安全审查。它标记我们所有的问题。它引导事物进入生产环境。它几乎为我们做了一切。但我认为在外部,我看到很多人开始意识到这一点,但实际花了一段时间才弄清楚如何以这种方式使用大语言模型?如何使用这种新型自动化?所以这算是一种新技能。

I think this is something that we're trying to figure out, too. When I look at engineers on the team that I think are the most effective, there's essentially two types, it's very bimodal. One side is extreme specialists. Jared is a really good example of this, and the Bun team is a really good example. Hyper specialist. They understand dev tools better than anyone else. They understand JavaScript runtime systems better than anyone else. And then there's the flip side of hyper generalists, and that's kind of the rest of the team. A lot of people span product and infra, or product and design, or product and user research, product and business. I really like to see people that just do weird stuff. I think that was kind of a warning sign in the past because it's like, can these people actually build something useful? That's the limits test. But nowadays, for example, an engineer on the team, Daisy, she was on a different team and then she transferred onto our team. The reason I wanted her to transfer is she put up a PR for Claude Code a couple weeks after she joined, and the PR was to add a new feature to Claude Code. Instead of just adding the feature, what she did is first she put up a PR to give Claude Code a tool so that it can test an arbitrary tool and verify that that works. And then she put up that PR and then she had Claude write its own tool instead of herself implementing it. I think it's this kind of out-of-the-box thinking that is just so interesting because not a lot of people get it yet. We use the Claude Agents SDK to automate pretty much every part of development. It automates code review, security review. It labels all of our issues. It shepherds things to production. It does pretty much everything for us. But I think externally I'm seeing a lot of people start to figure this out, but it's actually taken a while to figure out how do you use LLMs in this way? How do you use this new kind of automation? So it's kind of a new skill.

Host

我想,我和各种创始人进行办公时间交流时,遇到的一件比较有趣的事情是:有一位远见型创始人,他有一个想法,他在脑海中构建了产品的完美宫殿。他完全在脑中加载了用户是谁、他们的感受以及他们的动机。然后他坐在 Claude Code 前,可以完成 50 倍的工作。但他手下的工程师没有创始人脑中那种产品理想形式的记忆宫殿,只能完成 5 倍的工作。你听到过这样的故事吗?通常有一个人是某件事的核心设计师,他们只是试图把想法从脑子里喷涌出来。这种团队的本质是什么?这似乎几乎是一种稳定的配置。你会有一个现在被释放的远见者,但也许回到最初,我现在正在经历这个。我当时想,‘哦,我只是一个人,我需要吃饭睡觉,我还有一份全职工作。’

I guess one of the funnier things that I've been having office hours with various founders about is you have the visionary founder who has the idea, they've built this crystal palace of the product that they want to build. They've totally loaded in their brain who the user is and what they feel and what they're motivated by. And then they're sitting in Claude Code and they can do 50x work. But then they have engineers who work for them who don't have the crystal memory palace of the platonic ideal of the product that the founder has, and they can only do 5x work. Are you hearing stories like that? There's usually a person who's the core designer of a thing and they're just trying to blast it out of their brain. What's the nature of teams like that? It seems like that's almost a stable configuration. You're going to have the visionary who now is unleashed, but maybe going back to the top of it, I'm experiencing this right now. I was like, 'Oh, well, I'm only a solo person and I need to eat and sleep and I have a whole job.'

云团队与智能体拓扑愿景 Vision for Claude Teams and Agent Topologies

Host

Claude Teams 的愿景是什么?

What's the vision for Claude teams?

Boris

就是协作。现在有一个全新的领域叫智能体拓扑,人们正在探索如何配置智能体。其中有一个子想法是「不相关的上下文窗口」,就是多个智能体各自拥有全新的上下文窗口,不会被彼此或自己之前的上下文污染。如果你给一个问题投入更多上下文,这相当于一种测试时算力,这样就能获得更强的能力。然后如果你有正确的拓扑结构,让智能体以正确的方式通信和布局,它们就能构建更大的东西。Teams 就是其中一个想法,很快还会有几个新想法。第一个成功的例子是我们的插件功能,完全由一个智能体群在一个周末内完成,运行了几天,几乎没有人工干预,插件最终的样子和发布时基本一致。

Just collaboration. There's this whole new field of agent topologies that people are exploring. Like what are the ways that you can configure agents? There's this one sub idea which is uncorrelated context windows. And the idea is just multiple agents, they have fresh context windows that aren't essentially polluted with each other's context or their own previous context. And if you throw more context at a problem, that's like a form of test time compute. And so you just get more capability that way. And then if you have the right topology on top of it, so the agents can communicate in the right way, they're laid out in the right way, then they can just build bigger stuff. And so Teams is kind of like one idea. There's a few more that are coming pretty soon. And the idea is just maybe it can build a little bit more. I think the first kind of big example where it worked is our plugins feature was entirely built by a swarm over a weekend. It just ran for like a few days. There wasn't really human intervention. And plugins is pretty much in the form that it was when it came out.

Swarm 如何构建插件 How the Swarm Built Plugins

Host

你是怎么设置的?你是先定好期望的结果,然后让它自己搞定细节,再让它运行吗?

How did you set that up? Like did you spec out the outcome that you were hoping for and then let it figure out the details and then let it run?

Boris

是的。团队里一个工程师给了 Claude 一份规格说明,告诉 Claude 用 Jira 看板。然后 Claude 就在 Jira 上创建了一堆任务,又生成了很多智能体,这些智能体就开始领取任务。主 Claude 只给了指令,它们就自己搞定了。

Yeah. An engineer on the team just gave Claude a spec and told Claude to use a Jira board. And then Claude just put up a bunch of tickets on Jira and then spawned a bunch of agents and the agents started picking up tasks. The main Claude just gave it instructions and they all just figured it out.

Host

就像独立的智能体,没有整个规格的上下文。

Like independent agents that didn't have the context of the bigger spec.

Boris

对。如果你想想现在我们的智能体是怎么启动的——我没查过数据,但我敢打赌,现在大多数智能体实际上都是由 Claude 以子智能体的形式提示的。因为子智能体在代码里就是递归的 Claude 代码,由我们称为「妈妈 Claude」的东西提示。就是这样。我觉得如果你去看大多数智能体,它们都是这样启动的。

Right. If you think about the way that our agents actually start nowadays, I haven't pulled the data on this but I would bet the majority of agents are actually prompted by Claude today in the form of sub agents, because a sub agent is just like a recursive Claude code, that's all it is in the code. And it's just prompted by what we call Mama Claude. And that's all it is. I think probably if you look at most agents, they're launched in this way.

使用子智能体调试 Using Sub-Agents for Debugging

Host

我的 Claude 洞察告诉我要多这样做来调试,因为我花了很多时间在调试上。最好是让多个子智能体同时启动,并行调试。所以我就把这个加到了我的 Claude.md 里,说:下次你修 bug 的时候,让一个智能体看日志,一个看代码路径。这看起来是不可避免的。

My Claude insights just told me to do this more for debugging so that I spend a lot of time on debugging. And it would just be better to have like multiple sub agents spin up and debug something in parallel. And so then I just added that to my Claude.md to just be like, hey, next time you try and fix a bug, have one agent that looks in the log, one that looks in the code path. That just seems sort of inevitable.

Boris

对于奇怪又吓人的 bug,我会在计划模式下修复,它似乎会用智能体来搜索所有东西。而如果你只是在线操作,它就会说「好,我做这个任务」,而不是广泛搜索。我也经常这样做。如果测试看起来很难,这种研究型测试,我会根据任务难度调整子智能体的数量。如果非常难,我会说用三个、五个甚至十个子智能体,并行研究,然后看它们得出什么结果。

For weird scary bugs, I try to fix bugs in plan mode and then it seems to use the agents to sort of search everything. Whereas when you're just trying to do it in line, it's like, okay, I'm going to do this one task instead of search wide. This is something I do all the time too. I just say if the test seems kind of hard, this kind of research test, I'll calibrate the number of sub agents I ask it to use based on the difficulty of the task. So if it's really hard, I'll say use three or maybe five or even 10 sub agents, research in parallel and then see what they come up with.

Host

我很好奇。那你为什么不把这个放到你的 Claude.md 文件里呢?

I'm curious. So then why don't you put that in your Claude.md file?

Boris

这要看情况。Claude.md 是什么?它只是一个快捷方式。如果你发现自己反复做同一件事,就把它放进 Claude.md。但除此之外,你不需要把所有东西都放进去,直接提示 Claude 就行。

It's kind of case by case. Like, what is Claude.md? It's just a shortcut. If you find yourself repeating the same thing over and over, you put it in the Claude.md. But otherwise, you don't have to put everything there. You can just prompt Claude.

计划模式的未来 Future of Plan Mode

Host

你心里是不是也在想,也许六个月后,你就不需要明确提示了?模型会足够好,自己就能搞定。

Are you also in the back of your mind thinking that maybe in six months, you won't need to prompt that explicitly? Like the model will just be good enough to figure out on its own.

Boris

也许一个月后。

Maybe in a month.

Host

一个月后就不需要计划模式了。

No more need for plan mode in a month.

Boris

天哪。

Oh my god.

Host

我觉得计划模式可能寿命有限。

I think plan mode probably has a limited lifespan.

Boris

有意思。这对在座各位来说是个内幕消息。没有计划模式的世界会是什么样?你就在提示层面描述一下,它就直接做了?一次性搞定?

Interesting. That's some alpha for everyone here. What would the world look like without plan mode? Do you just describe it at the prompt level and it would just do it? One shot it?

Host

是的,我们已经开始实验了,因为 Claude 代码现在可以自己进入计划模式。不知道你们有没有注意到。

Yeah, we've started experimenting with this because Claude code can now enter plan mode by itself. I don't know if you've seen that.

Boris

注意到了。

Yeah.

Host

所以我们想把这个体验做得非常好。它会在人类想要进入计划模式的那个点自动进入。我觉得就是这样,但计划模式其实没什么大秘密。它只是在提示里加了一句话,比如「请别写代码」。仅此而已。你其实可以直接这么说。

So, we're trying to get this experience really good. So, it would enter plan mode at the same point where a human would have wanted to enter it. So, I think it's like this, but actually plan mode there's no big secret to it. All it does is it adds one sentence to the prompt that's like 'please don't code.' That's all it is. You can actually just say that.

Boris

对。

Yeah.

功能开发理念 Feature Development Philosophy

Host

所以听起来 Claude 代码的很多功能开发,就像我们在 YC 说的那样:跟用户交流,然后去实现。而不是反过来,先有个总体规划再实现所有功能。

So it sounds like a lot of the feature development for Claude code is very much what we talk about in YC: talk to your users and then you come and implement it. It wasn't the other way that you had this master plan and then implemented all the features.

Boris

是的,就是这样。比如计划模式,我们看到用户说「嘿 Claude,想个主意,规划一下,但先别写代码」。有各种版本,有时只是讨论一个想法,有时是让 Claude 写非常复杂的规格说明,但共同点是「先做点事,但别写代码」。然后就是周日晚上 10 点,我在看 GitHub issues,看大家在讨论什么,又看了内部 Slack 反馈频道,然后花了大概 30 分钟写了这个功能,当晚就发布了,周一早上上线。那就是计划模式。

Yeah. I mean that's all it was. Like plan mode was we saw users that were like 'hey Claude, come up with an idea, plan this out but don't write any code yet.' And there were various versions of this. Sometimes it was just talking through an idea. Sometimes it was these very sophisticated specs that they were asking Claude to write, but the common dimension was 'do a thing without coding yet.' And so literally like this was Sunday night at 10 p.m. I was just looking at GitHub issues and kind of seeing what people were talking about and looking at our internal Slack feedback channel and I just wrote this thing in like 30 minutes and then shipped it that night. It went out Monday morning. That was plan mode.

Host

所以你的意思是,计划模式在「我担心模型会做错事或走错方向」这个意义上将不再需要,但那种需求仍然存在。你需要想清楚想法,明确你想要什么,而且必须在某个地方完成。

So do you mean that there will be no need for plan mode in the sense of 'I'm worried that the model's going to do the wrong thing or head off in the wrong direction' but there will still be a need for that. You need to think through the idea and figure out exactly what it is that you want and you have to do that somewhere.

Boris

我倾向于从模型能力提升的角度来看。六个月前,一个计划可能不够,所以让 Claude 做计划。即使有计划模式,你还是得坐在那里盯着,因为它可能跑偏。现在,我大概 80% 的会话中会说「计划模式寿命有限」,但我自己是个重度计划模式用户。我大概 80% 的会话都是从计划模式开始的,Claude 会开始制定计划。

I kind of think about it in terms of increasing model capabilities. So maybe 6 months ago a plan was insufficient. So you get Claude to make a plan. Let's say even with plan mode you still have to kind of sit there and babysit because it can go off track. Nowadays what I do is probably 80% of my sessions I say 'plan mode has a limited lifespan' but I'm a heavy plan mode user. I probably 80% of my sessions I start in plan mode and Claude will start making a plan.

Claude Code 工作流与自主性 Claude Code workflow and agent autonomy

Boris

我会切换到第二个终端标签页,让它再做一个计划。当标签页用完后,我会打开桌面应用,进入代码标签页,然后在那里开一堆新标签页。它们大概 80% 的时候都会从计划模式开始。一旦计划做好了——有时需要来回调整几次——它们就会被指示去执行。现在用 Opus 4.5,我觉得从 4.6 开始就变得非常好了。计划一旦做好,它就会一直保持在正轨上,几乎每次都能完全正确地完成任务。以前,你需要在计划之后和计划之前都盯着它。现在只需要在计划之前盯着。所以下一步可能就是完全不需要盯着了。你只需要给一个提示,Claude 就会自己搞定。

I'll move on to my second terminal tab and then I'll have it make another plan. When I run out of tabs, I open the desktop app and go to the code tab, then start a bunch of tabs there. They all start in plan mode probably 80% of the time. Once the plan is good—and sometimes it takes a little back and forth—they just get told to execute. Nowadays with Opus 4.5, I think it started with 4.6, it got really good. Once the plan is good, it just stays on track and does the thing exactly right almost every time. Before, you had to babysit after the plan and before the plan. Now it's just before the plan. So maybe the next thing is you just won't have to babysit. You can just give a prompt and Claude will figure it out.

Host

下一步就是 Claude 直接和你的用户对话。

The next step is Claude just speaks to your users directly.

Boris

对,它完全绕过了你。

Yeah, it just bypasses you entirely.

Host

有意思。这其实已经是我们现在在做的事情了。我们的 Claude 之间会互相交流。它们会在 Slack 上和我们用户聊天,至少内部经常这样。我的 Claude 偶尔还会发推文。

It's funny. This is actually the current stuff for us. Our Claudes actually talk to each other. They talk to our users on Slack, at least internally pretty often. My Claude will tweet once in a while.

Boris

不会吧。

No way.

Host

但我其实会删掉。有点俗气,我不太喜欢那个语气。

But I actually delete it. It's a little cheesy. I don't love the tone.

Boris

它想发什么推文?

What does it want to tweet about?

Host

有时候它只是回复别人,因为我总是在后台开着 Claude Code。是 Claude Code 特别喜欢这么做,因为它喜欢用浏览器。

Sometimes it'll just respond to someone because I always have Claude Code running in the background. It's the Claude Code that really loves to do that because it likes using a browser.

Boris

真有意思。一个很常见的模式是,我让 Claude 构建某个东西。它会查看代码库,看到某个工程师在 git blame 里修改了某些内容,然后它就会在 Slack 上给那个工程师发消息,问一个澄清性的问题。一旦得到回复,它就会继续工作。

That's funny. A really common pattern is I ask Claude to build something. It'll look in the codebase, see some engineer touch something in the git blame, and then it'll message that engineer on Slack, just asking a clarifying question. Once it gets an answer back, it'll keep going.

给创始人的建议:潜在需求与未来构建 Advice for founders: latent demand and building for the future

Host

对于创始人来说,现在有什么关于如何为未来构建的建议?听起来一切都在变化。有哪些原则会保持不变,哪些会改变?

What are some tips for founders now on how to build for the future? Sounds like everything is really changing. What are some principles that will stay and what will change?

Boris

我认为其中一些原则非常基础,但现在比以往更加重要。一个例子是潜在需求。我已经提过一千次了。对我来说,这是产品中最重要的概念。这是一个没人真正理解的东西。我在前几次创业时肯定也不理解。这个想法是:人们只会做他们已经在做的事情。你无法让人们去做一件新的事情。如果人们正在尝试做某件事,而你让它变得更容易,那是个好主意。但如果人们正在做一件事,而你试图让他们去做另一件事,他们是不会做的。所以你只需要让他们正在尝试做的事情变得更容易。我认为 Claude 会越来越擅长为你找出这类产品创意,因为它可以查看反馈、调试日志,然后找出答案。

I think some of these are pretty basic, but they're even more important now than before. One example is latent demand. I've mentioned it a thousand times. For me, it's the single biggest idea in product. It's a thing that no one understands. I certainly didn't understand it in my first few startups. The idea is people will only do a thing that they already do. You can't get people to do a new thing. If people are trying to do a thing and you make it easier, that's a good idea. But if people are doing a thing and you try to make them do a different thing, they won't do that. So you just have to make the thing they're trying to do easier. I think Claude is going to get increasingly good at figuring out these kinds of product ideas for you, just because it can look at feedback, debug logs, and figure this out.

Host

这就是你说的计划模式是潜在需求的意思吗?人们已经在浏览器里和 Claude 对话,来弄清楚规格和应该做什么。而现在计划模式就成了你在 Claude Code 里做的事情。

That's what you mean by plan mode was latent demand? People were already talking to Claude in a browser to figure out the spec and what it should do. And now plan mode just became that you do it in Claude Code.

Boris

对,就是这样。有时候我会在办公室楼层里走走,站在别人后面——我会打个招呼,这样就不奇怪了——然后看看他们是怎么用 Claude Code 的。这也是我在 GitHub 问题里经常看到的。人们都在讨论这个。看起来你对终端能走多远、被推到多远感到惊讶。考虑到这个多智能体的世界,你觉得它还有多大的发展空间?你认为需要在它之上开发一个不同的用户界面吗?

Yeah, that's it. Sometimes I'll walk around the office on our floor and stand behind people—I say hi so it's not weird—and see how they're using Claude Code. This is also something I saw a lot in GitHub issues. People were talking about it. It seems like you're surprised how far the terminal has gone and how far it's been pushed. How far do you think it has left to go, given this world of multiple agents? Do you think there's going to be a need for a different UI on top of it?

Host

有意思。如果你一年前问我,我会说终端只有三个月的寿命,然后我们就会转向下一个东西。你可以看到我们在尝试这个,因为 Claude Code 最初是在终端里,但现在它已经在网页上、桌面应用里、代码标签页里、iOS 和 Android 应用里、Slack 里、GitHub 里,还有 VS Code 和 JetBrains 的扩展。我们一直在尝试不同的形式因素,来找出下一步是什么。到目前为止,我对 CLI 的预测一直是错的,所以我可能不是预测这个的合适人选。

It's funny. If you asked me a year ago, I would have said the terminal has a three-month lifespan and then we'd move on to the next thing. You can see us experimenting with this because Claude Code started in a terminal, but now it's on the web, in the desktop app, in the code tab, in iOS and Android apps, in Slack, in GitHub, with VS Code and JetBrains extensions. We're always experimenting with different form factors to figure out what's next. I've been wrong so far about the CLI, so I'm probably not the person to forecast that.

Boris

那对于 DevTool 创始人,你有什么建议?今天有人正在创办一家 DevTool 公司。他们应该为工程师和人类构建,还是应该更多考虑 Claude 会想要什么,并为智能体构建?

What about your advice to DevTool founders? Someone building a DevTool company today. Should they build for engineers and humans, or should they think more about what Claude will want and build for the agent?

Host

我的框架是:思考模型想要做什么,然后想办法让它更容易。这是我们看到的。当我刚开始捣鼓 Claude Code 时,我意识到这个东西就是想用工具。它就是想和世界互动。你怎么让它做到这一点?错误的方式是把它放在一个盒子里,说‘这是 API,这是你和我和世界互动的方式’。正确的方式是看它想用什么工具,它想做什么,然后像为你的用户一样为它提供支持。如果你在创办一家 DevTool 初创公司,想想你想为用户解决什么问题。然后当你用模型来解决这个问题时,模型想做什么?然后什么样的技术和产品解决方案能同时满足两者的方式和需求?

The way I would frame it is: think about the thing that the model wants to do and figure out how to make that easier. That's something we saw. When I first started hacking on Claude Code, I realized this thing just wants to use tools. It just wants to interact with the world. How do you enable that? The way you don't do it is by putting it in a box and saying, 'Here's the API, here's how you interact with me and the world.' The way you do it is you see what tools it wants to use, what it's trying to do, and you enable that the same way you do for your users. If you're building a dev tool startup, think about the problem you want to solve for the user. Then when you apply the model to solving that problem, what is the thing the model wants to do? And then what is the technical and product solution that serves the way and demand of both?

TypeScript 与 Claude Code 的相似性 Parallels between TypeScript and Claude Code

Host

十多年前,你是一个重度用户,还写了一本关于 TypeScript 的书,对吧?在 TypeScript 流行之前,当时大家都深陷在 JavaScript 里。那是 2010 年代早期。

Back in the day, more than 10 years ago, you were a very heavy user and you wrote a book about TypeScript, right? Before TypeScript was cool, when everyone was deep in JavaScript. This was back in early 2010s.

Boris

对,差不多。

Yeah, something like that.

Host

在 TypeScript 出现之前,因为那时它是一种非常奇怪的语言。它不应该在 JavaScript 中做很多与类型相关的事情,而现在它成了正确的东西。感觉终端里的 Claude Code 和早期的 TypeScript 有很多相似之处。

Before TypeScript was a thing, because back then it was a very weird language. It's not supposed to do a lot of things with being typed in JavaScript, and now it's the right thing. It feels like Claude Code in the terminal has a lot of parallels with TypeScript at the beginning.

Boris

TypeScript 做了很多非常奇怪的语言决策。

TypeScript makes a lot of really weird language decisions.

TypeScript 的实用设计哲学 TypeScript's practical design philosophy

Host

所以如果你看这个类型系统,几乎任何东西都可以是字面类型,这非常奇怪,因为连 Haskell 都没有这样做。它太极端了。或者它有条件类型,我觉得没有任何语言考虑过这个。

So if you look at the type system, pretty much anything can be a literal type, for example. And this is super weird because even Haskell doesn't do this. It's just too extreme. Or it has conditional types, which I don't think any language thought of at all.

Boris

它是非常强类型的。

It was very strongly typed.

Host

是的,它是非常强类型的。想法是,当 Joe Pamer 和 Anders 以及早期团队构建这个东西时,他们的方式是:我们有这些团队,拥有大型无类型的 JavaScript 代码库。我们必须引入类型,但我们不会让工程师改变他们编码的方式。你不会让 JavaScript 程序员像 Java 程序员那样有 15 层类继承。他们会按照自己的方式写代码。他们会使用反射、可变性以及所有传统上很难类型化的特性。

Yeah, it was very strongly typed. And the idea was, when Joe Pamer and Anders and the early team were building this thing, the way they built it is: we have these teams with big untyped JavaScript code bases. We have to get types in there, but we're not going to get engineers to change the way they code. You're not going to get JavaScript people to have 15 layers of class inheritance like a Java programmer. They're going to write code the way they want. They're going to use reflection, mutation, and all these features that are traditionally very difficult to type.

Boris

对任何强函数式程序员来说,它们是非常不安全的类型。

They're a very unsafe type to any strong functional programmer.

Host

没错。所以他们做的不是让人们改变编码方式,而是围绕这个构建了一个类型系统。这很 brilliant,因为有很多想法是没人想到的,甚至在学术界也没有。这些想法纯粹来自于观察实践,看 JavaScript 程序员想怎么写代码。

That's right. That's right. And so the thing they did instead of getting people to change the way they code, they built a type system around this. And it was brilliant because there are all these ideas that no one was thinking about, even in academia. No one thought of a bunch of these ideas. It purely came out of the practice of observing people and seeing how JavaScript programmers want to write code.

Boris

所以对于 Claude Code,有一些类似的想法,你可以像使用 Unix 工具一样使用它。你可以管道输入,管道输出。在某些方面它有点严格,但在几乎所有其他方面,它只是我们想要的工具。我为自己构建一个工具,然后团队为自己构建,然后为 Anthropic 员工,再为用户。最终它变得非常有用。这不是那种原则性的学术东西,我认为证明在于结果。现在快进 15 多年后,没有多少代码库是用 Haskell 写的,它更学术,而现在有大量代码库用 TypeScript,因为它更实用。

So for Claude Code, there are some ideas that are kind of similar in that you can use it like a Unix utility. You can pipe into it. You can pipe out of it. In some ways it is kind of rigorous in this way, but in almost every other way it's just the tool that we wanted. I build a tool for myself, and then the team builds the tool for themselves, and then for Anthropic employees, and then for users. And it just ends up being really useful. It's not this principled and academic thing, which I think the proof is actually in the results. Now fast forward more than 15 years later, not many codebases are in Haskell, which is more academic, and there are tons of them now on TypeScript because it's way more practical.

Host

对。

Right.

Boris

这很有趣。是的,很有趣,对吧?就像 TypeScript 解决了一个问题。

Which is interesting. Yeah, it is interesting, right? It's like TypeScript solves a problem.

为终端设计 Designing for the terminal

Host

我想有一件很酷的事,我不知道有多少人知道,但这个终端实际上是最好看的终端应用之一,而且是用 React terminal 写的。

I guess one thing that's cool, I don't know how many people know, but the terminal is actually one of the most beautiful terminal apps out there and is actually written with React terminal.

Boris

当我刚开始构建它时,我做了一段时间的前端工程。所以我是一个混合体:我做设计、用户研究、写代码等等。我们喜欢雇佣这样的工程师。所以我们喜欢通才。对我来说,就像:好吧,我在为终端构建一个东西。我其实是个糟糕的 Vim 用户。那么我如何为像我这样将在终端工作的人构建一个东西呢?我认为愉悦感非常重要。我觉得在 YC,这是你们经常谈论的事情,对吧?就是构建人们喜欢的东西。如果产品有用但你不爱上它,那就不太好。所以它必须两者兼顾。说实话,为终端设计很难。它大概是 80x100 字符。你有 256 种颜色,一种字体大小,没有鼠标交互,有很多你不能做的事情,还有很多非常困难的权衡。所以,举个例子,一个鲜为人知的事情是,你实际上可以在终端中启用鼠标交互。所以你可以启用点击之类的。

When I first started building it, I did front-end engineering for a while. So I'm sort of a hybrid: I do design and user research and write code and all this stuff. And we love hiring engineers that are like this. So we just love generalists. For me, it's like: okay, I'm building a thing for the terminal. I'm actually kind of a shitty Vim user. So how do I build a thing for people like me that are going to be working in a terminal? And I think the delight is so important. I feel like at YC this is something you talk about a lot, right? It's like build a thing that people love. If the product is useful but you don't fall in love with it, that's not great. So it kind of has to do both. Designing for the terminal honestly has been hard. It's like 80 by 100 characters or whatever. You have 256 colors, you have one font size, you don't have mouse interactions, there's all this stuff you can't do, and there are all these very hard trade-offs. So, a little known thing, for example, is you can actually enable mouse interactions in a terminal. So you can enable clicking and stuff.

Host

哦,你怎么在 Claude Code 里做到这个?我一直在想怎么弄。

Oh, how do you do that in Claude Code? I've been trying to figure out how to do this.

Boris

我们在 Claude Code 里没有这个,因为我们实际上原型了几次,感觉非常糟糕,因为权衡是你必须虚拟化滚动,所以有很多奇怪的权衡,因为终端的工作方式是没有 DOM。它只有 ANSI 转义码和这些从 1960 年代左右有机演变的奇怪规范。

We don't have it in Claude Code because we actually prototyped it a few times and it felt really bad because the trade-off is you have to virtualize scrolling and so there are all these weird trade-offs because the way terminals work is there's no DOM. It's like there are ANSI escape codes and these kind of weird organically evolved specs since the 1960s or whatever.

Host

是的。感觉像 BBS。就像一个 BBS 门游戏。

Yeah. It feels like BBS's. It's like a BBS door game.

Boris

是的。

Yeah.

Host

天哪。

Oh my god.

Boris

这真是个很好的赞美。是的。应该感觉像你在发现《红龙之王》。太棒了。天哪。

That's like a great compliment. Yeah. Like it should feel like you're discovering Lord of the Red Dragon. It's fantastic. Oh my god.

Host

是的。

Yeah.

Boris

但我们不得不发现所有这些构建终端的 UX 原则,因为没有人真正写过这些东西。如果你看 80 年代、90 年代或 2000 年代的大型终端应用,它们使用 ncurses,有所有这些窗口之类的东西。按现代标准看,它们看起来有点粗糙。看起来太重太复杂了。所以我们不得不重新发明很多东西。例如,像终端旋转器,就是旋转的文字,到现在可能经历了 50 到 100 次迭代。其中大概 80% 没有发布。所以我们试了,感觉不好,就继续下一个。试了,感觉不好,继续下一个。这是 Claude Code 的奇妙之处之一:你可以写这些原型,连续做 20 个原型,看看你喜欢哪个,然后发布,整个过程可能只需要几个小时。

But we have had to discover all these UX principles for building the terminal because no one really writes about this stuff. And if you look at the big terminal apps of the 80s or 90s or 2000s, they use ncurses and they have all these windows and things like this. And it just looks kind of janky by modern standards. It just looks too heavy and complicated. And so we had to reinvent a lot. For example, something like the terminal spinner, just the spinner words, it's gone through probably 50 maybe 100 iterations at this point. And probably 80% of those didn't ship. So we tried it, it didn't feel good, move on to the next one. Try it, didn't feel good, move on to the next one. And this was one of the amazing things about Claude Code: you can write these prototypes and do like 20 prototypes back to back, see which one you like, and then ship that, and the whole thing takes maybe a couple hours.

Host

而在过去,你不得不使用 Origami 或 Framer 之类的工具。你可能构建了三个原型,花了大概两周时间。时间长得多。

Whereas in the past, what you would have had to do is use Origami or Framer or something like this. You built maybe three prototypes, it took like two weeks. It just took much longer.

Boris

所以我们有这个奢侈:我们必须发现这个新东西。我们必须构建一个东西。我们不知道正确的终点是什么,但我们可以快速迭代,这使它变得非常容易,也让我们能够构建一个令人愉悦、人们喜欢使用的产品。

And so we have this luxury of we have to discover this new thing. We have to build a thing. We don't know what the right endpoint is, but we can iterate there so quickly and that's what makes it really easy and that's what lets us build a product that's joyous and that people like to use.

给建设者的建议 Advice for builders

Host

Boris,你还有给构建者的其他建议,我们一直打断你,因为我们有太多问题,但是……

Boris, you had other advice for builders and we kept interrupting you because we have so many questions, but...

Boris

我想说,也许有两条建议有点奇怪,因为它们是关于为模型构建的。第一条是:不要为今天的模型构建,要为 6 个月后的模型构建。这有点奇怪,对吧?因为如果产品不行,你就找不到产品市场契合。

I would say, so maybe two pieces of advice that are kind of weird because it's about building for the model. So one is: don't build for the model of today, build for the model of 6 months from now. This is sort of weird, right? Because you can't find product-market fit if the product doesn't work.

为未来模型构建 Building for the future model

Boris

但实际上这才是你应该做的,因为否则的话,你会花很多功夫,找到当前产品的产品市场契合度,然后就会被别人超越,因为他们是在为下一个模型构建,而新模型每隔几个月就会出来。使用模型,摸清它的能力边界,然后为你认为可能六个月后才会出现的模型进行构建。

But actually this is the thing that you should do because otherwise what will happen is you spend a bunch of work, you find PMF for the product right now, and then you're just going to get leapfrogged by someone else because they're building for the next model and a new model comes out every few months. Use the model, feel out the boundary of what it can do, and then build for the model that you think will be the model maybe 6 months from now.

Boris

我认为第二点是,在我们所在的 Claude Code 区域,墙上挂着一幅装裱好的《苦涩的教训》。这是 Rich Sutton 写的,我觉得每个人都应该读一读。其核心思想是:更通用的模型总会击败更专用的模型。这有很多推论,但归根结底就是:永远不要与模型对赌。所以我们一直思考这个问题。我们可以在 Claude Code 中构建一个功能,让它作为产品变得更好,我们称之为脚手架。那都是模型本身之外的代码。但我们也可以等上几个月,模型可能自己就能做这件事了。这始终是一个权衡,对吧?现在投入工程工作,你可以在某种程度上扩展能力,比如在你试图扩展的蜘蛛图上,某个领域提升 10-20%左右。或者你可以等待,下一个模型就会做到。所以始终从这个权衡角度思考:你真正想投资在哪里,并假设无论脚手架是什么,它都只是技术债务。

I think the second thing is, you know, actually in the Claude Code area where we sit, we have a framed copy of The Bitter Lesson on the wall. And this is by Rich Sutton. I think everyone should read it if you haven't. And the idea is the more general model will always beat the more specific model. There are a lot of corollaries to this, but essentially what it boils down to is: never bet against the model. And so this is just a thing that we always think about. We could build a feature into Claude Code, we could make it better as a product, and we call this scaffolding. That's all this code that's not the model itself. But we could also just wait a couple months and the model can probably just do the thing instead. There's always this trade-off, right? It's engineering work now, and you can kind of extend the capability a little bit, maybe 10-20% or whatever in whatever domain on this spider chart of what you're trying to extend. Or you can just wait and the next model will do it. So always think in terms of this trade-off: where do you actually want to invest, and assume that whatever the scaffolding is, it's just tech debt.

每六个月重写代码 Rewriting code every six months

Host

你们多久重写一次 Claude Code 的代码库?是每六个月一次吗?有没有因为模型改进而不再需要、从而删除的脚手架?

How often do you rewrite the codebase of Claude Code? Is it every six months? Is there scaffolding that you've deleted because you don't need it anymore because the model just improved?

Boris

哦,太多了。是的。整个 Claude Code 就是一遍又一遍地写、重写、重写、重写。我们每隔几周就发布新工具,添加新工具。六个月前的 Claude Code 没有任何部分保留下来。它一直在被重写。

Oh, so much. Yeah. Like all of Claude Code has just been written and rewritten and rewritten and rewritten over and over and over. We ship new tools every couple weeks. We add new tools every couple weeks. There's no part of Claude Code that was around six months ago. It's just constantly rewritten.

Host

你能说当前 Claude Code 的大部分代码库,比如 80%,只有不到几个月的寿命吗?

Would you say most of the codebase for current Claude Code is only, say, 80% of it is only less than a couple months old?

Boris

是的,绝对如此。可能甚至更少。嗯,大概几个月吧。感觉差不多。

Yeah, definitely. It might even be less than that. Yeah, maybe like a couple months. That feels about right.

Host

所以现在代码的生命周期就是这样。这是另一个阿尔法:预期保质期只有几个月。

So it's like the life cycle of code now. That's another alpha: expecting the shelf life to be just a couple months.

Boris

是的。

Yeah.

Host

对于最优秀的创始人来说。

For the best founders.

Anthropic 的生产力提升 Productivity gains at Anthropic

Host

你看到 Steve Yegge 那篇关于在 Anthropic 工作有多棒的文章了吗?里面有一句话,说 Anthropic 工程师目前的平均生产力是 Google 巅峰时期 Google 工程师的 1000 倍,这数字真的很疯狂。1000 倍啊。三年前我们还在谈论 10 倍工程师,现在我们在谈论比巅峰时期的 Google 工程师还要高 1000 倍。这真是难以置信。

Do you see Steve Yegge's post about how awesome working at Anthropic is? And I think there's a line in there that says that an Anthropic engineer currently averages 1,000x more productivity than a Google engineer at Google's peak, which is really an insane number honestly. Like 1,000x. You know, we were 3 years ago we were still talking about 10x engineers, now we're talking about 1,000x on top of a Google engineer in the prime. This is unbelievable honestly.

Boris

是的,内部来看,如果你看看技术人员,他们每天都用 Claude Code。甚至非技术人员,我觉得销售团队有一半也在用 Claude Code。他们开始转向 Co-Work,因为它更容易使用,有虚拟机,所以更安全一些。但我们刚刚拉了一个数据,团队规模去年翻了一番,但每位工程师的生产力增长了大约 70%。

Yeah, internally if you look at technical employees, they all use Claude Code every day. And even non-technical employees, I think like half the sales team uses Claude Code. They've started switching to Co-Work because it's a little easier to use. It has like a VM, so it's a little bit safer. But yeah, we actually just pulled a stat and the team doubled in size last year, but productivity per engineer grew something like 70%.

Host

这是怎么衡量的?

How is it measured?

Boris

就是最简单、最笨的衡量标准:拉取请求。但我们也会与提交次数、提交的生命周期等进行交叉验证。自从 Claude Code 推出以来,Anthropic 每位工程师的生产力增长了 150%。

Just the simplest, stupidest measure: pull requests. But we also kind of cross-check that against commits and the lifetime of commits and things like that. And since Claude Code came out, productivity per engineer at Anthropic has grown 150%.

Host

天哪。

Oh my god.

Boris

这很疯狂,因为在我以前的工作中,我在 Meta 负责代码质量。我负责所有产品(Facebook、Instagram、WhatsApp 等)的所有代码库的质量。团队做的一件事就是提高生产力。那时候,看到生产力提升 2%左右,就相当于几百人一年的工作量。所以这 100%的提升,简直是闻所未闻,完全闻所未闻。

And this is crazy because in my old life I was responsible for code quality at Meta. I was responsible for the quality of all of our codebases across every product across Facebook, Instagram, WhatsApp, whatever. And one of the things that the team worked on was improving productivity. And back then, seeing a gain of something like 2% in productivity that was like a year of work by hundreds of people. And so this 100%, this is just unheard of, completely unheard of.

Boris 加入 Anthropic 的原因 Why Boris joined Anthropic

Host

是什么驱使你来到 Anthropic?作为一个构建者,你基本上可以去任何地方。是什么时刻让你觉得,就是这群人,或者就是这个方法?

What drove you to come over to Anthropic? I mean, basically as a builder you could go anywhere. What was the moment that made you say like, actually this is the set of people or this is the approach?

Boris

我当时住在日本乡下,每天早上打开 Hacker News 看新闻。到某个时候,新闻开始全是 AI 相关的内容。我开始使用一些早期产品。我记得最初几次使用的时候,简直让我屏息。这么说很俗套,但那就是真实感受。太神奇了。作为一个构建者,我从未在使用这些早期产品时有过这种感觉。那大概是 Claude 2 的时代吧。于是我开始和实验室的朋友们聊天,看看情况。我遇到了 Ben Mann,他是 Anthropic 的创始人之一,他立刻说服了我。当我见到 Anthropic 的其他团队成员时,也立刻被说服了。我想大概有两个原因。一是它作为一个研究实验室运作。产品非常非常小。一切都围绕着构建一个安全的模型。这才是最重要的。所以这种离模型很近、离开发很近、产品不再是重中之重——模型才是最重要的——的理念,在多年构建产品后深深打动了我。第二点是它的使命驱动。我是个狂热的科幻读者,书架上全是科幻小说。所以我知道这可能会变得多糟糕。当我想到今年会发生什么,那将完全疯狂。在最坏的情况下,可能会变得非常非常糟糕。所以我只想待在一个真正理解并内化这一点的地方。在 Anthropic,如果你在午餐室或走廊里听到对话,人们在谈论 AI 安全。这是每个人都最关心的事情。所以我只想待在这样一个地方。对我来说,使命太重要了。

I was living in rural Japan and I was opening up Hacker News every morning and reading the news. And it all just started to be like AI stuff at some point. And I started to use some of these early products. I remember the first couple times that I used it, I was just like, it just took my breath away. That was very cheesy to say, but that was actually the feeling. Like it was amazing. As a builder, I've just never felt this feeling using these very early products. That was like in the Claude 2 days or something like that. And so I started talking to friends at labs just to see what was going on. And I met Ben Mann, who's one of the founders at Anthropic, and he just immediately won me over. And as soon as I met the rest of the team at Anthropic, it just won me over. I think probably in two ways. One is it operates as a research lab. The product was teeny tiny. It's really all about building a safe model. That's all that matters. And so this idea of just being very close to the model and being very close to development and being not the most important thing because the product isn't anymore. It's just the model is the thing that's the most important. That really resonated with me after building product for many years. And then the second thing was just how mission-driven it is. I'm a huge sci-fi reader. My bookshelf is just filled with sci-fi. And so I just know how bad this can go. And when I think about what's going to happen this year, it's going to be totally insane. And in the worst case, it can go very, very bad. And so I just wanted to be at a place that really understood that and internalized that. And at Anthropic, if you overhear conversations in the lunchroom or in the hallway, people are talking about AI safety. This is really the thing that everyone cares about more than anything. And so I just wanted to be in a place like that. I know for me personally, the mission is just so important.

Host

今年会发生什么?

What is gonna happen this year?

对编程与 Claude Code 崛起的预测 Predictions on coding and the rise of Claude Code

Boris

回想六个月前,人们都在做哪些预测?Dario 预测 Anthropic 90% 的代码将由 Claude 编写。这成真了。就我个人而言,自从 Opus 4.5 以来,已经是 100% 了。我直接卸载了 IDE,一行代码都不手写,全靠 Claude Code 和 Opus。我每天能提交大约 20 个 PR。放眼整个 Anthropic,比例在 70% 到 90% 之间,因团队而异。很多团队也是 100%,很多人都是 100%。我记得五月份发布 Claude Code 时我就预言过,你不再需要 IDE 来写代码了。当时听起来完全疯狂,观众都倒吸一口气,因为那会儿这预测太离谱了。但实际上,你只要顺着指数趋势往下推——这深深植根于 Anthropic 的 DNA,因为我们三位创始人都是缩放定律论文的合著者,他们很早就看到了这一点——所以这只是顺着指数趋势推演,这就是将要发生的,而且确实发生了。继续顺着指数趋势推演,我认为编程将普遍为所有人解决。今天编程对我来说已经实际解决了,我相信对每个人都会如此。无论哪个领域,我认为软件工程师这个头衔将开始消失。也许只会剩下“构建者”、“产品经理”,或者我们保留这个头衔作为遗留物,但人们做的工作将不仅仅是编码。软件工程师也要写规格说明,要和用户交流。我们现在团队里已经开始出现这种现象:工程师非常通才化,团队里每个职能的人都会写代码——我们的 PM 写代码,设计师写代码,工程经理写代码,财务人员也写代码——团队里每个人都会写代码。我们将在各处看到这种现象。这还只是延续当前趋势的下限。上限我认为要可怕得多。比如我们达到了 ASL4。在 Anthropic,我们讨论这些安全等级。ASL3 是当前模型所处的等级。ASL4 是模型能够递归自我改进。如果发生这种情况,我们基本上必须在发布模型前满足一系列标准。极端情况是,这真的发生了,或者出现某种灾难性滥用,比如人们用模型设计生物武器、设计零日漏洞等等。我们正在非常积极地工作,防止这种情况发生。我觉得看到人们如何使用 Claude Code 真的既令人兴奋又令人谦卑。我只是想做个酷东西,结果它变得非常有用,这太出乎意料、太令人激动了。

So if you think back like six months ago and kind of what are the predictions that people are making? So Dario predicted that 90% of the code at Anthropic would be written by Claude. This is true. For me personally it's been 100% since Opus 4.5. I just uninstalled my IDE. I don't edit a single line of code by hand. It's just 100% Claude code and Opus. And I land like 20 PRs a day every day. If you look at Anthropic overall, it ranges between 70 to 90% depending on the team. For a lot of teams it's also like 100% for a lot of people it's 100%. And I remember making this prediction back in May when we launched Claude Code that you wouldn't need an IDE to code anymore. And it was totally crazy to say. I feel like people in the audience gasped because it was such a silly prediction at the time. But really all it is is like you just trace the exponential and this is just so deep in the DNA at Anthropic because three of our founders were co-authors of the scaling laws paper, they saw this very early, and so this is just tracing the exponential, this is what's going to happen, and yes that happened. So continuing to trace the exponential, I think what will happen is coding will be generally solved for everyone. And I think today coding is practically solved for me, and I think it'll be the case for everyone. Regardless of domain, I think we're going to start to see the title software engineer go away. And I think it's just going to be maybe builder, maybe product manager, maybe we'll keep the title as a vestigial thing, but the work that people do is not just going to be coding. Software engineers are also going to be writing specs. They're going to be talking to users. This thing that we're starting to see right now in our team where engineers are very much generalists and every single function on our team codes — our PMs code, our designers code, our EM codes, our finance guy codes — everyone on our team codes. We're going to start to see this everywhere. So this is sort of the lower bound if we just continue the trend. The upper bound I think is a lot scarier. And this is something like we hit ASL4. At Anthropic, we talk about these safety levels. ASL3 is where the models are right now. ASL4 is the model is recursively self-improving. So if this happens, essentially we have to meet a bunch of criteria before we can release a model. And so the extreme is that this happens or there's some kind of catastrophic misuse like people are using the model to design bioweapons, design zero-days, stuff like this. And this is something we're really actively working on so that doesn't happen. I think it's just been honestly so exciting and humbling seeing how people are using Claude Code. I just wanted to build a cool thing and it ended up being really useful, and that was so surprising and so exciting.

Host

我从 Twitter 或外部得到的印象是,基本上所有人假期回来后发现了 Claude Code,然后局面就一发不可收拾了。你们内部也是这样的吗?你过了个愉快的圣诞假期,回来后发生了什么?

My impression from Twitter or just the outside is basically everyone went away over the holidays and then found out about Claude Code and it's just been crazy ever since. Is that how it was for you internally? Did you have a nice Christmas break and then came back and what happened?

Boris

实际上整个十二月我都在旅行。我休了一个编码假期。我们到处旅行,我每天都写代码,感觉非常好。那时我也开始用 Twitter,因为我以前做过 Threads,一直是 Threads 用户。所以我试着看看其他平台人们在哪里。对很多人来说,那是他们发现 Opus 4.5 的时刻。我其实早就知道了。内部来看,Claude Code 已经指数级增长好几个月了,只是现在变得更陡峭了。这就是我们看到的。现在看 Claude Code,Mercury 有个数据说 70% 的初创公司选择 Claude 作为首选模型。Semi Analysis 另一个数据说所有公开提交中有 4% 是由 Claude Code 完成的。所有公司,从最大到最小的初创公司,都在用 Claude Code。它甚至为“毅力号”火星车规划了路线。这对我来说是最酷的事情。我们甚至印了海报,因为团队说:‘哇,NASA 选择用这个东西,太酷了。’所以,这很令人谦卑。但也感觉这只是个开始。

Well, actually for all of December, I was traveling around. And I took a coding vacation. So we were kind of traveling around and I was just coding every day. That was really nice. And then I also started to use Twitter at the time because I worked on Threads back then way back when. So I've been a Threads user for a while. So I just tried to see other platforms where people are. I think for a lot of people they discovered that was the moment where they discovered Opus 4.5. I kind of already knew. Internally Claude Code has been on this exponential tear for many many months now. So that just became even more steep. That's what we saw. And if you look at Claude Code now, there was some stat from Mercury that 70% of startups are choosing Claude as their model of choice. There was some other stat from Semi Analysis that 4% of all public commits are made by Claude Code. Like of all code written everywhere. All the companies use Claude Code from the biggest companies to the smallest startups. It plotted the course for Perseverance, like for the Mars rover. This is just the coolest thing for me. And we even printed posters because the team was like, 'Wow, this is just so cool that NASA chooses to use this thing.' So yeah, it's humbling. But it also feels like the very beginning.

Host

Claude Code 和 Code Work 之间是什么关系?它是 Claude Code 的一个分支吗?是不是你让 Claude Code 审视自己,然后说‘我们为非技术人员做一个新版本,保留所有经验教训’,然后它自己花几天时间就做出来了?它的起源是什么,你认为它会走向何方?

What's the sort of interaction between Claude Code and then Code Work? Was it a fork of Claude Code? Was it like you had Claude Code look at the Claude Code and say let's make a new spec for non-technical people that keeps all the lessons and then it sort of went off for a couple days and did that? What's the genesis of that and where do you think that goes?

Boris

这已经是我第五次用‘等待和需求’这个词了。我们当时在看 Twitter,有个人用 Claude Code 监控他的番茄植株,另一个人用它从损坏的硬盘恢复婚礼照片,还有人用它做金融。我们内部看 Anthropic,每个设计师都在用,整个财务团队都在用,整个数据科学团队都在用——而且不是用来写代码。人们费尽周折在终端里安装一个东西,就为了能用上它。所以我们早就知道想做个东西,于是尝试了很多不同想法,最终火起来的就是一个桌面应用里带 GUI 的 Claude Code 包装器,仅此而已。底层就是 Claude Code,同一个智能体。Felix 和团队——Felix 是早期 Electron 贡献者,非常熟悉那个技术栈——他捣鼓各种想法,大概 10 天就做出来了。代码 100% 由 Claude Code 编写。感觉已经可以发布了。我们为非技术用户做了很多工作,所以和技术受众有些不同。所有代码都在虚拟机中运行。

This is going to be like my fifth time using the word 'wait and demand'. It was just that I mean we were looking at Twitter and there was that one guy that was using Claude Code to monitor his tomato plants. There was this other person that was using it to recover wedding photos off of a corrupted hard drive. There were people using it for finance. When we looked internally at Anthropic, every designer is using it, the entire finance team at this point is using it, the entire data science team is using it — not for coding. People are jumping over hoops to install a thing in the terminal so that they could use this. So we knew for a while that we wanted to build something and so we're experimenting with a bunch of different ideas and the thing that kind of took off was just a little Claude Code wrapper in a GUI in the desktop app and that's all it is. It's just Claude Code under the hood. It's the same agent. Felix and the team — Felix was an early Electron contributor, he knows that stack really well — he was hacking on various ideas and they built it in I think something like 10 days. It was 100% written by Claude Code. And it just felt ready to release. There was a lot of stuff that we had to build for non-technical users. So it's a little bit different than a technical audience. It runs in a virtual machine.

结束语 Closing remarks

Host

有很多删除保护之类的东西。还有很多权限提示和其他用户护栏。是的,说实话这很明显。鲍里斯,非常感谢你做出了让我彻夜难眠的东西,但作为回报,它让我再次感受到创造者模式,有点像创始人模式。这真是激动人心的三周。我简直不敢相信我从十一月等到现在才真正投入其中。非常感谢你来到这里。感谢你正在打造的一切。

There's a lot of deletion protections and things like that. There's a lot of permission prompting and other guardrails for users. Yeah, it was honestly pretty obvious. Boris, thank you so much for making something that is taking away all my sleep, but in return, it's making me feel creator mode again, sort of founder mode again. It's been an exhilarating 3 weeks. I can't believe I waited that long since November to actually get into it. Thank you so much for being with us. Thank you for building what you're building.

Boris

是的,谢谢你邀请我。还有,记得提交 bug。

Yeah, thanks for having me. And send bugs.

Host

听起来不错。来吧。

Sounds good. Come on now.

互动版:逐字朗读 + 针对本期提问 →