从宝可梦卡到 AI:Claude Code 的故事

From Pokémon Cards to AI: The Claude Code Story

鲍里斯·切尔尼 Boris Cherny · The Pragmatic Engineer · 2026-03-04 · 约 98 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Claude Code 的创造者 Boris Cherny 分享了他如何从在 eBay 上卖宝可梦卡到构建增长最快的开发者工具之一,每天提交 20-30 个拉取请求且无需手写一行代码。

Boris Cherny, creator of Claude Code, shares how he went from selling Pokémon cards on eBay to building one of the fastest-growing developer tools, shipping 20-30 pull requests daily with zero handwritten code.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 43)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Background

Host

你是 O'Reilly 那本 TypeScript 书的作者。我在日本一个小镇上找到了那本书的日文译本,那感觉太酷了。然后我意识到我完全不记得 TypeScript 了。现在,Claude Code 平均写了 Anthropic 大约 80% 的代码。我每天可能提交 10 到 20 个拉取请求。Opus 4.5 和 Claude Code 写了每一个拉取请求的 100%,我没有手动编辑过一行代码。Andrew Carpet 发帖说他从未像现在这样觉得自己作为程序员落后这么多。这是我真正纠结的事情。模型进步太快了,以至于在旧模型上有效的想法可能在新模型上就不管用了。我对这个时刻的一个比喻是 15 世纪的印刷机,因为当时有一群知道如何书写的抄写员。一些国王是文盲,他们雇佣了抄写员。如果你想想抄写员发生了什么,他们不再是抄写员,但现在有了作家和作者这个类别。这些人现在存在了,他们存在的原因是文学市场大大扩张了。当你加入世界上顶尖的 AI 实验室之一,你的第一个拉取请求被拒绝会发生什么?不是因为代码不好,而是因为你手写了它。这正是 Boris Cherny 加入 Anthropic 时发生的事情。Boris 是 Claude Code 的创建者和工程负责人。在加入 Anthropic 之前,他在 Meta 工作了七年,领导了 Instagram、Facebook、WhatsApp 和 Messenger 的代码质量,并且是公司最多产的代码作者和代码审查者之一。在今天的节目中,我们将介绍 Claude Code 如何从一个副项目变成增长最快的开发者工具之一,以及 Anthropic 内部关于是否发布它的辩论。Boris 的日常工作流程是每天发布 20 到 30 个拉取请求,没有手写代码,以及当 AI 编写所有代码时代码审查如何工作。为什么 Boris 相信我们正在经历一个像印刷机一样变革的时代,以及哪些工程技能现在更重要,哪些不再重要。如果你想了解最接近 AI 编码智能体的人之一今天如何实际构建软件,以及这对我们其他工程师意味着什么,这期节目就是为你准备的。本期节目由 Statsig 呈现,这是一个用于标志、分析、实验等的统一平台。查看节目说明以了解更多关于他们以及我们的其他季赞助商 Sonar 和 WorkOS 的信息。你是如何进入科技、软件工程和编程领域的?

You wrote the first ever TypeScript book with O'Reilly. Yeah. I found that book translated in Japanese in this little town in Japan. That was just the coolest moment. And then I realized I don't remember TypeScript at all. Now we're at the point where Claude Code writes, I think, something like 80% of the code at Anthropic on average. I wrote maybe 10, 20 pull requests every day. Opus 4.5 and Claude Code wrote 100% of every single one. I didn't edit a single line manually. Andrew Carpet posted that he's never felt as much behind as a programmer as he is now. This is something I really struggle with. The model is improving so quickly that the ideas that worked with the old model might not work with the new model. One metaphor I have for this moment in time is the printing press in the 1400s because there was a group of scribes that knew how to write. Some of the kings were illiterate who were employing the scribes. And if you think about what happened to the scribes, they ceased to become scribes, but now there's a category of writers and authors. These people now exist. And the reason they exist is because the market for literature just expanded a ton. What happens when you join one of the top AI labs in the world and your first pull request gets rejected? Not because the code was bad, but because you wrote it by hand. This is exactly what happened to Boris Cherny when he joined Anthropic. Boris is the creator and engineering lead behind Claude Code. Before joining Anthropic, he spent seven years at Meta where he led code quality across Instagram, Facebook, WhatsApp, and Messenger and was one of the most prolific code authors and code reviewers at the company. In today's episode, we cover how Claude Code went from a side project to one of the fastest growing developer tools and the internal debate at Anthropic whether to release it at all. Boris's daily workflow of shipping 20, 30 pull requests a day with zero handwritten code and how code review works when AI writes everything. Why Boris believes we're living through a time as transformative as a printing press and which engineering skills matter more now and which ones do not. If you want to understand how one of the people closest to AI coding agents actually build software today and what that means for the rest of us engineers, this episode is for you. This episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes to learn more about them and our other season sponsors, Sonar and WorkOS. How did you get into tech, software engineering, and coding in general?

Boris

这要从很久以前说起。我想有两条平行的路径交叉了。大约在我 13 岁的时候,我开始在 eBay 上卖我的旧宝可梦卡。我意识到在 eBay 上你可以写 HTML。我看了别人的宝可梦卡列表,发现有些有大的颜色和字体之类的。然后我发现了 blink 标签。我加上 blink 标签,就能把卡卖到 99 美分而不是 49 美分。所以我这样学了 HTML,然后我买了一本 HTML 书学了 HTML。第二件事是,我想也是在中学的时候,我们有旧的 TI-83 图形计算器,用来做数学。我意识到如果我把数学考试的答案编程进计算器,我就能得到更好的分数。所以我写了这些小程序。你只需要编程答案,然后考试变难了,所以我不得不编程求解器而不是实际问题,因为我事先不知道系数之类的。然后第二年数学更高级了,所以我不得不从 Basic 降到汇编,只是为了让程序运行得快一点。

It starts a while back. I think there was kind of like two parallel paths that crossed. So, when I was maybe 13 or something like this, I started selling my old Pokémon cards on eBay. And I realized that on eBay you can actually write HTML. And I was looking at other people's Pokémon card listings, and I realized some of them have big colors and fonts and stuff like this. And then I discovered the blink tag. I put the blink tag on it, I could sell my card for like 99 cents instead of 49 cents or whatever. So, I kind of learned about HTML this way, then I got an HTML book and learned about HTML. And then the second thing was this was also, I think, sometime in middle school. We had these old TI-83 graphing calculators. And we used them for math. And what I realized is I can get a better answer on the math test if I just program the answers to the math test into my calculator. And so, I wrote these little programs. You just program the answers, and then the test got harder, so then I had to program solvers instead of the actual questions because I didn't know what the coefficients and stuff would be ahead of time. And then the math got more advanced like the next year. And so, I had to drop down from basic to assembly to just make the program run a little bit faster.

Host

哦,所以你在高中就降到汇编了?

Oh, so you got in high school you dropped down to assembly?

Boris

我想这是中学或高中,可能是八年级或九年级之类的。然后我意识到班上的每个人都开始发现我有求解器,他们有点嫉妒。所以我买了一条小串行线,这样我也可以给他们。然后下一次数学考试,全班都得了 A。老师问:“怎么回事?”最后她发现了,就说:“好吧,你这次侥幸过关,别干了。”但对我来说这非常实用。所以,在学校我学了经济学。我实际上辍学去创业了。我从未想过编程会成为职业。它对我来说一直很实用。编程是构建东西和制作有用东西的手段。

I think this is like middle school or high school. It may be like eighth or ninth grade or something like this. Then the thing I realized is everyone in my class was starting to realize that I had the solver, and they got kind of jealous. And so, I bought this little serial cable so I can give it to them, too. And then the next math test, everyone in the class just got A's. And the teacher was like, "What's going on?" And then eventually she realized it. It was just like, "Okay, you get away with it once, and knock it off." But for me it was very practical. So, you know, in school I studied economics. I actually dropped out to start startups. And I never thought that coding would be a career at all. It was always very practical to me. Coding is a means to build things and to make useful things.

Host

第一个创业项目,我想是我和朋友们想弄到大麻。所以我们搞了一个大麻评论网站。我们做了一个网站,联系了不同的药房。然后我们试图弄到大麻样品,这样我们就可以为他们写评论。它实际上有点火了。然后我变得更感兴趣,因为当时没有人测试这些东西。所以我进入了化学测试和化学分析领域。之后我又做了一堆其他创业项目。然后我很早就加入了 YC。我是 Palo Alto 那家 YC 创业公司的第一个员工。

This startup the first one was I think it's like my friends and I were trying to get weed. And so we started this like weed review startup. We made like a website. We called kind of different dispensaries I think. And then we just tried to get kind of like weed samples so we could like review it for them. And it actually kind of blew up. And then I actually got more interested in at the time no one was like testing this stuff. And so I got into kind of the chemical testing kind of chemical analysis. And then after this I kind of did a bunch of other startups. And then I joined YC actually pretty early. And I was the first hire of this YC startup up in up in Palo Alto after.

Host

你是如何决定从一个创业项目跳到另一个的?

How did you decide to go from one startup to the other?

Boris

凭感觉吧。因为你知道创业公司,从来不是一条直线。你总是要转型、转型、再转型。你必须弄清楚市场想要什么,用户想要什么。从来不是你最初想的那样。你总是尝试一些东西,但想法总是假设,然后几乎总是要转型一次、两次、三次。你知道这家叫 Agile Diagnosis 的医疗软件公司。这是一家早期的 YC 公司,大概是 2011 年或 2012 年。它是为医生设计的医疗软件。想法是,有这些临床决策协议,每家医院差异很大。我们的想法是,芝加哥有一家医院有一个非常好的针对心脏症状的协议。所以我们想,如果美国每家医院都使用相同的协议,结果会不会很好?所以我们试图标准化它。我们制作了这个决策树软件供医生使用。我写了部分软件。团队只有我们几个人,非常小。我写了软件,它运行在网页浏览器中。我记得那是在 Internet Explorer 6 的时代,医院用的就是那个。

Kind of vibes. I'd say. Because you know startups, it's never a linear path. You always kind of pivot, pivot, pivot. You have to figure out what the market wants and what users want. And it's never the thing that you think. You always try something but the idea is always a hypothesis and then almost always you have to pivot once, twice, three times. You know at this medical software company called Agile Diagnosis. This was kind of an early YC company. This was back in maybe 2011, 2012 something like that. It was medical software for doctors. And the idea was there's these clinical decision protocols. They vary a lot hospital to hospital. And our idea was there's one hospital in Chicago that had a really great protocol specifically for cardiac symptoms. And so we're like wouldn't outcomes be great if every hospital in the US would use the same protocol? And so we tried to standardize it. And we made this decision tree software for doctors to use. And I wrote some of the software. The team was just a few of us. It was a pretty small team. And I wrote the software. It was in a web browser. And I remember this was back in the Internet Explorer 6 days. That's what hospitals were using.

从健康初创公司学习用户行为 Learning from user behavior at a health startup

Boris

我写了一个 SVG 渲染器,因为它是一个可视化的决策树。我们上线后,日活用户数一直持平,搞不懂原因。当时我们在几家医院试点,包括加州大学旧金山分校。我那时骑摩托车,就骑到 UCSF,跟诊了几天,看看医生到底怎么用这个产品。我发现医生根本没时间坐下来用电脑——看完一个病人,到下一个病人可能只有 5 分钟。这 5 分钟里,你得走过走廊,到电脑站,打开一台老旧的电脑。启动就要 3 分钟,再打开 IE6 又要 30 秒,然后打开我们做的应用,登录。5 分钟就没了,根本没时间用。所以我们重写了一遍,改成在安卓上运行,但他们还是不用。后来我们意识到,医生身后总跟着一群住院医。这种场景下,社交因素很重要——他们需要维持权威形象,不想被人看到在玩手机。于是我们又转型了,觉得也许医生不是目标用户,护士或放射科技师可能更合适。那时我离开了,因为觉得这离我想做的事太远了。对我来说,最有趣的是找到产品市场契合点,因为它总是出人意料。你不能只有一个大想法,因为那个想法很可能是错的。所以你要形成假设,去验证,看看什么是对的。

And I wrote this like SVG renderer because it was this visual decision tree. And we launched it and then we had a DAU chart and the DAUs were flat and couldn't figure it out. And we were piloting it with a few hospitals at the time. And at the time we were based in Palo Alto, we were piloting it with you know, a few hospitals including UCSF. And I rode a motorcycle at the time. So I rode my motorcycle up to you know, UCSF and I shadowed doctors for a couple days just to see how how do they actually use this. And I realized that actually doctors don't have time to sit down and use a computer because you're seeing a patient then you have maybe 5 minutes until the next patient. And in those 5 minutes you have to walk down the hall, you have to go to the computer station, you have to open up this totally legacy computer. By the time it boots up that's like 3 minutes. Then you open up Internet Explorer 6. That takes like 30 seconds. Then you have to open up this like app that we built. You have to sign in. And your 5 minutes are up. You don't even have time to use it. And so we rewrote everything to run on Android and they still weren't using it. And the thing we realized is doctors are walking around with a bunch of residents behind them. In this kind of situation it's like a social situation, right? Like the thing that matters is they're seen as an authority. They don't want to be seen on their phones. And then we pivoted again. So at that point we were like, okay, so maybe the doctor isn't the target user. Actually we want it to be used by maybe nurses or x-ray technicians or something like this. At that point I left because I was like this is actually pretty far off from kind of what I wanted to do. This is like the most fun thing for me is finding this this product market fit cuz it's always surprising. You can't have one big idea because the idea is probably going to be wrong. So, you kind of form hypotheses, you you follow down, and and you see what's right.

Host

我觉得你讲这个故事很有意思,因为很多成功故事的背后,我们听到的都是成功路径。但首先,很多创业公司都是这样的;其次,让我印象深刻的是,你当时是作为软件工程师被招进去的,那还是在产品工程师这个概念出现之前。但你骑着摩托车去现场,跟诊,了解他们怎么用、为什么不用,从中获得想法。我觉得这就是一个优秀软件工程师的素质,无论过去还是现在。你似乎并不专注于技术本身,而是专注于结果。

Also, I find it's so interesting how you're telling us this story, cuz I feel behind a lot of sort of success stories, we hear the success story, we hear the path of how it went. But, first of all, a lot of startups are like this, and second of all, what struck me is you you were hired as a software engineer, right? And this was back before product engineers or anything was a thing, which we're now talking about. But, you just like you rode your motorbike, and you went there, and you shadowed the people, and you understood how they're using it, why they're not using it, getting getting ideas. I I feel, you know, this this is what makes a great software engineer back then and and even today, right? You you you weren't Doesn't seem to me that you were focused on the technology, you were focused on the outcome, though.

Boris

是的,工程师有很多种,做事的方式也不同。比如在我们现在的团队里,像 Jared Sumner 这样的工程师,技术头脑非常出色,他对系统的理解超过我见过的任何人。你需要这样的人,需要这种深度。对我来说,工程一直是实用性的。我一直是个通才,不管是做设计、工程还是用户研究,都无所谓。

Yeah, I mean, look, there there's different kinds of engineers, and there's different ways to do it. And, you know, I even even on our team right now, I look at an engineer like Jared Sumner, and he's just incredible technical mind. He understands systems better than anyone I've met. And, you know, you need you need people like this. You need people with this kind of depth. For me, engineering has always been a practical thing. Uh and, you know, for me, I've always been a generalist. And like, it doesn't matter if I'm doing you know, like design, or you know, if I'm doing engineering, or user research, or whatever.

Sonar 广告集成 Sonar ad integration

Host

AI 与软件工程的投资逻辑很简单:AI 写的代码越多,需要验证的代码就越多。但有一个问题:AI 生成的代码平均比人类写的更难验证。这就是 Sonar(SonarQube 的开发者)存在的意义。作为 AI 赋能世界的关键验证层,Sonar 确保 AI 带来的速度和数量不会损害你的代码库。Sonar 的竞争地位建立在 17 年的专业经验之上,这是任何基础模型都无法复制的。我们说的是深度分析引擎,比如符号执行和跨仓库数据流追踪,它们模拟代码的实际行为,而不仅仅是表面内容。为了弥合 AI 生产力与代码质量之间的鸿沟,Sonar 发布了 SonarCube MCP 服务器。这个工具充当 AI 应用与 SonarCube 平台之间的通用翻译器。通过使用模型上下文协议,它让 Claude Code、GitHub Copilot 和 Cursor 等 AI 工具直接访问 SonarCube 的分析能力。无需切换上下文,你的 AI 代理就能成为完整的代码审查和质量保证副驾驶,能够分析代码语法问题、按严重性过滤错误,甚至在提交代码前检查项目的质量门状态。无论你是使用编码辅助还是扩展全代理工作流,Sonar 都提供自动化验证,75%的财富 100 强企业依赖于此。它让你的开发者能够自由创新,而不用担心破坏代码库。访问 sonarsource.com/pragmatic,了解更多关于 Sonar 如何让你自信地以 AI 速度开发。接下来,让我们回到 Boris 的职业生涯以及他在创业公司学到的经验。

The investment thesis for AI and software engineering is straightforward. As AI writes more code, more code needs to be verified. But, there's a catch. AI-generated code is, on average, harder to verify than human-written code. This is why there's Sonar, the makers of SonarQube. As a critical verification layer for the AI-enabled world, Sonar ensures that speed and volume with AI does not compromise your codebase. Sonar's competitive position is built on 17 years of specialized expertise that no foundational model can replicate. We're talking about deep analysis engines, like symbolic execution and cross-repository data flow tracking that simulate how code actually behaves, not just what it says. To bridge the divide between AI productivity and code quality, Sonar has released the SonarCube MCP server. This tool acts as a universal translator between AI applications and the SonarCube platform. By using the model context protocol, it gives AI tools like Claude Code, GitHub Copilot, and Cursor direct access to SonarCube's analysis capabilities. Instead of context switching, your AI agent becomes a full-fledged code review and quality assurance copilot capable of analyzing code syntax for issues, filtering bugs by severity, and even checking your project's quality gate status before you ever commit code. Whether you're working with coding assistance or scaling up with full agenda workflows, Sonar provides the automated verification that 75% of the Fortune 100 rely on. It's about giving your developers the freedom to innovate without the fear of breaking the code base. Head to sonarsource.com/pragmatic to learn more about how Sonar enables the confidence to develop at the speed of AI. With this, let's get back to Boris's career and what he learned working at startups.

早期自由职业与第一份工作 Early freelancing and first job

Boris

我的第一份工作是在 16 岁,当时我想买一把电吉他。于是我开始做自由职业,想着那就做网站吧。那时 Fiverr 还没出现,但有一些其他的自由职业网站。我建了个网站,开始竞标项目。第一笔薪水全花在了电吉他上。但这很实用,因为在这种模式下,你得自己做工程、会计、设计,还要跟客户沟通。对我来说一直是这样。

My first job I ever had, I was like, I think I was 16. And I just wanted to buy an electric guitar. And so what I did was I I started I just started freelancing. And so I was like, okay, I guess I'll make websites. And I think Fiverr was not a thing back then, so there were some other freelancing websites. So I just started like I put up a website, I started bidding on stuff. And my first paycheck, I just spent the entire thing on an electric guitar. But it But it was very practical. Right? Cuz it's like when you're in this kind of setup, you have to you have to do the engineering, you have to do kind of the accounting, you have to do the the design, you have to talk to customers. So it it's just always been like that for me.

Facebook 七年与职业成长 Seven years at Facebook and career growth

Host

在几家创业公司之后,你去了 Facebook,也就是现在的 Meta。你在那里待了 7 年。能跟我们聊聊你在那里做了什么、学到了什么吗?你的职业发展也非常出色,7 年内晋升了 4 次。你从那段经历中收获了什么?

After a couple of these startups, you ended up at Facebook, now now now called Meta. And there you spent 7 years there. Can you just talk us through what you worked there, what what you've learned there? You've also had a very remarkable career growth in terms of four promotions over over over 7 years. And what do you take away from that that experience?

Boris

是的,我一开始在 Facebook Groups 工作。那是 Vlad Kolesnikov 招的我,他现在应该还在 Facebook,可能在别的团队。那很酷,我合作过的一大批人都是早期的 JavaScript 开发者。我做了很多 JavaScript 相关的工作,有趣的是我总跟这些人打交道。Vlad 做过 Bolt JS,那是驱动 Ads Manager 的框架,后来变成了 React JS。我总跟这些人有交集。

Yeah, so I started on Facebook groups. That was the first time I worked on Vlad Kolesnikov hired me. I think I think he's actually still at Facebook. Um I think he's on some other team now. And it was cool actually. There there's a big group of people that I worked with that were these kind of early JavaScript people too. And you know, like I did I did a bunch of JavaScript stuff and it's funny like I kept crossing paths with these people. And so Vlad, he worked on Bolt JS, which was the software it was the framework that powered Ads Manager, which later became React JS. I kept crossing paths with these people.

早期职业生涯与 Facebook 群组 Early Career and Facebook Groups

Boris

后来还有更多这样的人。总之,我当时在 Facebook 群组工作。我非常兴奋,因为它的使命是连接人们和他们的社区,这正是吸引我的地方。那时我是 Reddit 的重度用户。我十几岁就开始用 Reddit,因为我不认识其他编程的人。即使在大学,我也不太认识编程的人。说实话,我一直有点尴尬,因为我觉得这是件书呆子气的事。我觉得这是我懂的东西,但我想当酷孩子,不能告诉别人我会编程,这太书呆子了。后来我发现 Reddit 上有个编程社区,我震惊了。还有其他人也喜欢这个。这是个奇怪的爱好,很小众。找到志同道合的人并建立联系,这太令人兴奋了。所以我想在这方面工作,以某种方式做出贡献。我在 Facebook 群组工作了一段时间。最终我成为了 Facebook 群组的技术负责人。工作从构建变成了大量的文档编写、协调和委派他人。当时文化正在变化。早期的 Facebook 文化正在消失。文档来了,对齐会议来了。在隐私和安全等基础工作上有了更多的工作。老实说,早期为了增长走了很多捷径,但最终你必须偿还这笔债。那就是那个时候。

And later on there were a bunch more people like this. But anyway, I was working on Facebook groups. I was really excited about it because of this mission of connecting people to their community. That's what drew me in. At the time I was a big Reddit user. I became a Reddit user back when I was a teenager because I didn't know anyone else that coded. Even in college I didn't really know anyone that coded. Honestly, I was always kind of embarrassed about it because I thought it was this nerdy thing. I thought it was something I knew how to do, but I wanted to be like a cool kid. I couldn't tell people that I coded; it was very nerdy. At some point I discovered a programming community on Reddit. I was just shocked. There are other people into this thing. It's such a weird hobby, so niche. It was just so exciting to find like-minded people and get this connection. So I wanted to work on this, to contribute in some way. I worked on Facebook groups for a while. Eventually I became the tech lead for Facebook groups. The work changed from building to a lot of doc writing, coordination, and delegating to others. The culture was changing at the time. Early Facebook culture was disappearing. Docs were coming in, alignment meetings were coming in. There was a lot more work around foundational stuff like privacy and security. Honestly, early on a lot of corners were cut to grow, but at some point you have to pay that debt. That was the time when that happened.

Host

嗯。

Mhm.

转战 Instagram 与日本 Move to Instagram and Japan

Boris

之后我在 Instagram 待了几年。这也有个有趣的故事。我妻子得到了一份工作机会,她非常兴奋。她来找我说:‘嘿,我拿到了这个 offer,但我们要搬家。可以吗?’我说:‘没问题。我在科技行业工作,我们可以远程工作。工作在哪?’她说:‘在奈良。’我说:‘那是哪?’奈良是日本的乡下。那是 2021 年。

Then I spent a few years at Instagram after. That was also a funny story. My wife got a job offer and she was really excited about it. She came to me and said, 'Hey, I got this offer, but we're going to move. Is that okay?' I said, 'Yeah, that's fine. I work in tech. We can work remotely anywhere. Where's the job?' She said, 'It's in Nara.' I said, 'Where's that?' Nara is rural Japan. This was 2021.

Host

不同的时区,是啊。大概 12 小时的时差之类的。

Different time zone, yeah. 12 hours difference or something like that.

Boris

差不多吧。我试图找一个能赞助我的团队,因为有一些关于时区和同地办公的神秘 HR 规定。Instagram 在东京有一个刚起步的小团队。Will Bailey 负责这个团队,他也是 Instagram Stories 的创造者。他当了我一段时间的经理。我们决定一起壮大那个团队。我从奈良远程工作,大部分团队在东京。这段时间我在捣鼓 Instagram,技术栈简直疯狂。Facebook 拥有世界上最好的 Web 服务栈。从 Hack 语言到 HHVM 运行时,到 GraphQL 作为传输层,到 Relay 这样的客户端库,还有 React,一切都优化得令人惊叹。世界上没有其他开发栈这么好。它是完全优化的。然后我去了 Instagram,用的是 Python,类型检查器不工作,点击跳转到定义也不工作。它是一个拼凑起来的 Django,一个 CPython 运行时的分支。什么都不好用。所以我来到 Instagram,加入了日本的 Labs 团队。想法是为 Instagram 找到下一个大事件。我们尝试了一些东西,但我很快意识到我在这个栈上工作效率不高,因为它太糟糕了。所以我开始做开发基础设施,因为我们需要修复它。我们做了几个项目。一个是从 Python 迁移到 Facebook 的大单体仓库。另一个是从 Instagram GraphQL 迁移。这些项目实际上还在进行中;它们涉及数百名工程师多年时间。这是一个大型代码库,大型迁移。现在有了 AI 工具,速度更快了。迁移是它们很好的用例。这是完美的用例。

Something like that, yeah. I tried to find a team that would sponsor me because there were arcane HR rules about time zones and co-location. There was a little nascent team for Instagram in Tokyo. Will Bailey was running the team; he also made Instagram stories. He was my manager for a while. We decided to grow that team together. I worked remotely from Nara, and most of the team was in Tokyo. During this time I was hacking on Instagram, and the stack was just insane. Facebook had the single best web serving stack in the world. The way everything is optimized—from the Hack language to the HHVM runtime, to GraphQL as the transport layer, to client libraries like Relay and all that stuff, and React—it was just amazing. There's no other dev stack in the world that was this good. It was fully optimized. Then I went to Instagram, and it was Python where the type checker didn't work. Click to definition didn't work. It was a hacked-together Django, a fork of the CPython runtime. Nothing really worked. So I came to Instagram, joined the Labs team in Japan. The idea was to find the next big thing for Instagram. We tried some stuff, but I quickly realized I was not effective at working on the stack because it was such a terrible stack. So I went and started working on Dev Infra because we needed to fix it. There were a few projects we worked on. One was migrating from Python to the big Facebook monolith. Another was migrating from Instagram GraphQL. These projects are actually in progress; they involve hundreds of engineers over many years. It's a big code base, a big migration. Now it's faster with AI tools. Migrations are a pretty good use case for them. It's the perfect use case.

Meta 的代码质量与卓越工程 Code Quality and Better Engineering at Meta

Boris

到我离开 Instagram 时,我在做开发基础设施并领导这些迁移。那也是我与 Fiona Fung 交集的地方,她现在负责 Claude Code 团队。我和她一起工作,她是一位了不起的领导者,在技术方面有深厚的底蕴和历史。我认为没有比她更好的经理来管理这个团队了。然后我也开始做代码质量。Instagram 的工作范围扩大了一些。到我离开时,我领导了整个 Meta 的代码质量工作。我负责 Instagram、Facebook、Messenger、WhatsApp、Reality Labs 等所有代码库的质量。在 Meta,有一个叫做 Better Engineering 的项目。我想它大概始于 2016 或 2018 年。扎克伯格规定公司每位工程师必须花 20% 的时间修复技术债务。我们称之为 better engineering。有些是自下而上的,团队最清楚需要修复的技术债务;有些是自上而下的,需要进行大型迁移到新的语言特性、新框架等。在 Facebook 的规模下,每年有数万次这样的迁移。我领导这一切,很快意识到它需要更多秩序。没有目标,没人知道结果,没有追踪。所以我们开发了一些东西。一个想法是集中式的方法来优先处理不同的代码质量工作。第二件事是弄清楚代码质量对工程生产力的影响,结果发现影响很大。

By the time I left Instagram, I was working on Dev Infra and leading a bunch of these migrations. That's also where I intersected with Fiona Fung, who is now the manager for the Claude Code team. I just worked with her, and she was such an amazing leader. Incredible depth and history in tech. I thought there's no better manager for this team. Then I also started working on code quality. The work on Instagram expanded a bit. By the time I left, I was leading code quality for all of Meta. I was responsible for the quality of the code bases across Instagram, Facebook, Messenger, WhatsApp, Reality Labs, all these code bases. At Meta, there was this program called Better Engineering. I think it started around 2016 or 2018. Zuck mandated that every engineer at the company spend 20% of their time fixing tech debt. We call this better engineering. Some of it is bottom-up where a team knows best the tech debt they have to fix, and some of it is top-down where you need to do very big migrations to new language features, new frameworks, etc. At Facebook scale, there are tens of thousands of these migrations every year. I was leading all this, and I realized very quickly that it needed a little more order. There were no goals, no one knew the outcomes, there was no tracking. So we developed a bunch of stuff. One idea was a centralized way to prioritize different code quality efforts. The second thing was figuring out the impact of code quality on engineering productivity, which turned out to be significant.

衡量工程师生产力 Measuring Engineer Productivity

Host

你是怎么衡量的?你发现了什么?

How did you measure? What did you find there?

Boris

有很多东西。有些已经发表了,我不确定是否全部发表了。但本质上,你尝试做因果分析和因果推断,这是方法论。你试图找出哪些因素能让工程师更高效。其中一部分是代码质量,另一部分则与代码质量无关。例如,Meta 决定回到办公室办公而不是远程工作,部分原因就是基于此。因为我们发现了一些我们认为具有因果关系的强相关性。代码质量实际上对生产力的贡献达到了两位数百分比。事实证明,即使在最大规模下也是如此。

There was a bunch of stuff. I think some of this has been published. I don't know if all of it has, but essentially you try to do causal analysis and causal inference. This is the methodology. You try to figure out what are the factors that make engineers more productive. Some of it is code quality, some of it is outside of code quality. So for example, Meta went back to return to office instead of work from home. That was partially driven by this. Because we just found some fairly strong correlations that we thought were causal. The code quality actually contributes double-digit percent to productivity. It turns out even at the biggest scale.

Host

听到这个让人欣慰,因为我觉得很少有地方能真正衡量这一点,但我们能感觉到。比如,当你有一个干净、模块化的代码库时,工作起来会更轻松。而且我在想,这对大语言模型来说是不是也更容易处理?我的直觉是肯定的,对吧?但我觉得这方面的数据很少,不过这是我的感觉。

This is kind of comforting to hear because I think it's rare to have a place where you actually measure this, but I think we feel it. Like when you have a clean code base, modular, it can get easier to work with. And I think reasoning could it also be easier for LLMs to work with it? My hint would be yes, it should be, right? But I think there's just very little data, but that's the feeling I would have.

Boris

是的,我认为很多大公司都发表过相关研究。比如 Facebook 发过一些东西,微软发了很多,谷歌也是。但完全同意。如果你每次构建一个功能时,都要考虑“我该用框架 X、Y 还是 Z?”,这些都是你可以考虑的选项,因为代码库处于部分迁移状态,所有这些框架都散落在代码中。作为工程师,你会很痛苦。作为新员工,你会很痛苦。作为模型,它可能会选错东西,然后用户需要纠正它。所以,更好的做法是始终保持代码库的整洁。确保当你开始迁移时,就完成迁移。这对工程师有好处,现在对模型也有好处。

Yeah, I think a lot of the big companies have published about this. Like I think Facebook published something, Microsoft publishes a bunch about this, Google does. But yeah, totally. If every time you build a feature, you have to think about 'do I use framework X or Y or Z?' These are all options you can consider because the code base is in a partially migrated state where all of these are around the code somewhere. As an engineer, you're going to have a bad time. As a new hire, you're going to have a bad time. As a model, you might just pick the wrong thing. And then the user has to course correct you. So actually, the better thing to do is just always have a clean code base. Always make sure that when you start a migration, you finish the migration. This is great for engineers, and nowadays, it's great for models, too.

加入 Anthropic 与首个 PR 被拒 Joining Anthropic and First PR Rejected

Host

然后你加入了 Anthropic,我听说你的第一个拉取请求被 Adam Wolf 拒绝了,你可以确认或补充更多细节吗?

And then you joined Anthropic, and I've heard the story which you can confirm or give more color to that your first pull request was rejected by Adam Wolf.

Boris

他是我的入职伙伴。我加入 Anthropic 时,正在思考下一步做什么。我见了各个实验室的很多人,而 Anthropic 对我来说是显而易见的选择,因为它的使命。这是我个人最需要的东西。而且看到正在发生的所有这些变化,拥有一个框架来思考这个问题以及我们在其中的角色是很重要的。我还是一个科幻小说迷,那绝对是喜欢的类型。我读很多书,家里有一个大书架。我知道事情可能会变得多糟糕。我觉得这是一个有严肃思考者的地方。人们非常认真地对待这件事,思考我们能做些什么让它变得更好。所以,当我加入 Anthropic 时,我做了一些入职项目,就是各种我捣鼓的东西。我手动写了第一个拉取请求,因为我觉得代码就该这么写。过去确实是这么写的。但即使在那个时候,Anthropic 已经有了一个叫 Claude 的东西,它是 Claude Code 的前身。它非常粗糙,是 Python 写的,启动需要 40 秒。那是研究代码,不是智能体式的。但如果你非常小心地提示它,并正确使用这个工具,它可以为你写代码。所以,Adam 拒绝了我的 PR,他说:“实际上,你应该用这个 Claude 工具来做。”我说:“好的,没问题。”我花了大概半天时间才学会怎么用这个工具,因为你需要传入一堆标志并正确使用。但然后它直接生成了一个可用的 PR,一次就搞定了。

He was my ramp-up buddy. So, I joined Anthropic. I was trying to figure out what to do next, and I met a bunch of people at all the different labs, and Anthropic was just the obvious choice for me because of the mission. This is the thing that personally I know that I need the most. And also just seeing all this change that's happening, it's important to have some sort of framework to think about this and to think about our role in it. I'm also a really big sci-fi reader. That's definitely my genre. I'm a big reader. I have a giant bookshelf at home. And I just know how bad this thing can go. And I just felt like this is a place that has serious thinkers. People are taking this very seriously and thinking about what we can do to make this thing go better. So, when I joined Anthropic, I did a bunch of ramp-up projects, just various stuff that I was hacking on. And I wrote my first pull request by hand because I thought that's how you write code. That used to be how you write code. But even at the time at Anthropic, there was this thing called Claude, and it was the predecessor to Claude Code. It was super janky. It was Python, it took like 40 seconds to start up. It was research code. It was not agentic. But if you prompt it very carefully and hold the tool just right, it can write code for you. And so, Adam rejected my PR, and he was like, 'Actually, you should use this Claude thing for it instead.' And I was like, 'Okay, cool.' It took me like half a day to figure out how to use this tool because you have to pass in a bunch of flags and use it correctly. But then it spat out a working PR. It just one-shotted it.

Host

哦。

Oh.

Boris

那大概是 2024 年,九月或八月左右。对我来说,这是我在 Anthropic 第一次“感受到 AI”的时刻。我当时想“天哪”。我不知道模型能做到这个。我习惯了那种 IDE 里的 Tab 补全、行级补全。我完全不知道它能直接为我生成一个可用的拉取请求。

And this was like 2024. It's like September 2024 or August. Something like that. And I think for me this was my first 'feel the AI' moment at Anthropic. I was just 'oh my god.' I didn't know the model could do this. I was used to these kind of tab completions, line-level completions in an IDE. I had no idea that it could just make a working pull request for me.

赞助商插播:Statsig Sponsor Break: Statsig

Host

War 刚刚谈到他在工作中使用他们的 AI 模型时有一个真正的“哇”时刻。一个非常不同的“哇”时刻是当你使用一个工具,让工作比以前轻松得多。这正好引出了我们的赞助商 Statsig。Statsig 为工程团队提供实验和功能标记工具,这些工具过去需要多年的内部工作才能构建。这种工具非常复杂,只有像 Meta 或 Uber 这样的大公司才有自己的定制高级工具。Statsig 在实践中是这样的:你将一个更改放在功能开关后面,逐步推出,比如先面向 1% 或 10% 的用户。你观察发生了什么。不仅仅是它是否崩溃,而是它对你关心的指标有什么影响:转化率、留存率、错误率、延迟。如果看起来不对劲,你迅速关闭它。如果趋势正确,你继续推进。关键是测量是工作流程的一部分。你不需要在三个工具之间切换,事后尝试匹配细分和仪表板。功能标记、实验和分析都在一个地方,使用相同的底层用户分配和数据。这就是为什么 Notion、Brex 和 Atlassian 等公司的团队使用 Statsig。Statsig 有慷慨的免费层供入门,团队专业版定价每月 150 美元起。要了解更多并获取 30 天企业试用,请访问 static.com/pragmatic。接下来,让我们回到 Boris 和 Claude Code 的起源故事。

War just talked about how he had a true wow moment at work using their AI model. A very different wow moment is when you use a tool at work that makes things so much easier than before. And this leads us nicely to our presenting sponsor, Statsig. Statsig offers engineering teams tooling for experimentation and feature flagging that used to require years of internal work to build. It's the kind of tool that was so complex to build that only large companies like Meta or Uber had their own custom advanced tooling for it. Here's what Statsig looks like in practice. You ship a change behind a feature gate and roll it out gradually, say to 1% or 10% of users at first. You watch what happens. Not just did it crash, but what did it do to the metrics you care about? Conversion, retention, error rates, latency. If something looks off, you turn it off quickly. If it's trending the right way, you keep it rolling forward. And the key is that measurement is part of the workflow. You're not switching between three tools and trying to match up segments and dashboards after the fact. Feature flags, experiments, and analytics are all in one place using the same underlying user assignments and data. This is why teams at companies like Notion, Brex, and Atlassian use Statsig. Statsig has a generous free tier to get started, and pro pricing for teams starts at $150 per month. To learn more and get a 30-day enterprise trial, go to static.com/pragmatic. And with this, let's get back to Boris and the origin story of Claude Code.

Claude Code 的起源 Origin of Claude Code

Host

是的,然后当你加入 Anthropic 时,我们已经深入探讨过这个,但我们可以简要回顾一下 Claude Code 是如何从一个看似副项目或一个很酷的黑客项目诞生的。

Yeah, and then when you joined Anthropic, we've covered this in a deep dive, but we could recap briefly on how Claude Code came to be out of what seemed like a side project or just a cool hack.

Boris

是的,我开始捣鼓各种不同的东西。我在产品方面做了一些工作。我花了一点时间研究强化学习,只是为了理解我构建的层之下的那一层。这仍然是我给很多工程师的建议:始终理解底层。这非常重要,因为它能给你深度,让你在实际工作的层上有更多的杠杆。这是 10 年前的建议,今天仍然适用。但现在的底层有点不同了。

So, yeah, I started hacking on a bunch of different stuff. I was working on some things in product. I worked on reinforcement learning for a little bit just to kind of understand the layer under the layer of which I was building. This is still advice that I give to a lot of engineers: always understand the layer under. It's really important because that just gives you the depth and you kind of have a little bit more levers to work at the layer that you actually work at. This was the advice 10 years ago. It's still the advice today. But the layer under is a little bit different now.

从聊天机器人到智能体 AI From Chatbot to Agentic AI

Boris

你知道,以前你得理解 Java,如果你写 JavaScript,就得理解 JavaScript 虚拟机和框架之类的。现在呢,你得理解模型。所以我当时在捣鼓各种东西,有些发布了,有些没有。某个时候,我就是想试试 Anthropic 的公开 API,因为我之前从没用过。我不想搭个 UI,只想快速搞点东西出来,因为那时候还没有 Claude Code,我们还在手写代码。于是我写了个小 bash 工具,唯一的功能就是调用 Anthropic API,本质上就是个基于聊天的应用,只不过跑在终端里——因为那时候 AI 就是那样的。我始终觉得,工程师是最早的采纳者。所以当我们从对话式 AI 转向智能体式 AI 时,虽然花了一点时间,但工程师很快就理解了。现在如果你问非工程师什么是 AI,他们会说是对话式 AI,像个聊天机器人什么的。正因如此,我对我们刚发布的 co-work 这个新产品特别兴奋,因为它会把工程师很早就看到的东西带给所有人。但当我想到 co-work 时,我会回想起这个非常早期的时刻。Claude Code 最初并不是 Claude Code,它是个聊天机器人,因为那时我以为 AI 就是那样。但我们得搞清楚下一步是什么。所以我当时建了个聊天机器人,有点用,但就是个聊天机器人。接下来我尝试的是让它使用工具,因为工具使用刚出来,我不知道那是什么。我想,实验是什么?我给了它一个工具,就是 bash 工具,我不知道该拿它做什么。于是我问它——其实我都不确定它能不能做到——我问它:我现在在听什么音乐?它就用 osascript 之类的写了个小 AppleScript 程序,打开我的音乐播放器,然后查询正在播放的音乐。它用 Sonnet 3.5 一次就搞定了。这其实是我第二次有 AGI 的感觉,紧跟在第一次之后。模型就是想用工具,这就是我意识到的。如果你给它一个工具,它会自己想办法用它来完成任务。我觉得当时人们处理 AI 和编程的方式,大家心里都有个模型:你把模型放进一个盒子里,然后想,接口是什么?你想怎么和这个模型交互?你需要它做什么?本质上就像你写程序时,先搭个模块框架,搭个函数框架,然后说,好了,这部分现在是 AI 了。但程序的其他部分还是普通程序。这种思考模型的方式不对。正确的思考方式是:模型是独立的东西。你给它工具,给它它可以运行的程序,让它运行程序,让它写程序。但你不要把它做成这个大系统的一个组件。我觉得这是“苦涩教训”的一个版本。“苦涩教训”是一个非常具体的框架,但它有很多推论。这就是其中一个推论:让模型做它自己的事,别试图把它放进盒子里,别试图强迫它按特定方式行事。

You know, before it was like understand the Java if you're writing JavaScript, understand the JavaScript VM and frameworks and stuff. Now it's like understand the model. So, I was hacking on a bunch of different stuff. Some things shipped, some things didn't ship. And at some point I just wanted to understand the public Anthropic API because I'd never used it before. I didn't want to build a UI. I just wanted to hack something up quite quickly because we didn't have Claude Code back then. We were still writing code by hand. And I wrote this little bash tool that all it did was hit the Anthropic API and it was essentially like a chat-based application but just in the terminal, because that's what AI used to be. I still think about it like engineers are the first adopters. So when we started to move out of conversational AI to agentic AI, it took a little bit, but engineers understood it pretty quick. I think now when you ask non-engineers about what is AI, they would say it's this conversational AI, like a chatbot or something. That's why I'm actually very excited for co-work, this new product that we launched, because it's going to bring the same thing that engineers saw very early to everyone else. But when I think about co-work, I think back to this moment very early on. Claude Code originally wasn't Claude Code. It was a chatbot. Because that's what I thought AI was. But we had to figure out what is the next thing. So at the time I built this chatbot. It was somewhat useful, but it was just a chatbot. And the next thing that I tried was I wanted it to use tools. Because tool use just came out and I didn't know what it was. I was like, what's the experiment? I gave it a single tool, which was the bash tool, and I didn't know what to do with the bash tool. So I asked it, I actually didn't know if it could even do this, but I asked it like, what music am I listening to? And it just wrote a little AppleScript program using osascript or whatever to open up my music player and then query it to see what music it's playing. And it one-shotted this with Sonnet 3.5. This is actually my second AGI feeling moment very quickly after the first one. The model just wants to use tools. That's what I realized. Like this thing, if you give it a tool, it will figure out how to use it to get the thing done. I think at the time, the way that people were approaching AI and coding, everyone essentially had this mental model of you take the model and you put it in a box. You figure out what is the interface, how do you want to interact with this model, what do you need it to do? Essentially, it's like if you have a program, you stub out some module, stub out some function, and you say, okay, this is now AI. But otherwise, the rest of the program is just a program. This is just not the way to think about the model. The way to think about it is the model is its own thing. You give it tools. You give it programs that it can run. You let it run programs. You let it write programs. But you don't make it a component of this larger system in this way. I think this is a version of the bitter lesson. The bitter lesson is a very specific framing, but there are many corollaries to it. This is one of the corollaries: just let the model do its thing. Don't try to put it in a box. Don't try to force it to behave a particular way.

Host

你最早看到的方式就是给它工具,让它访问 bash,后来是文件系统,再后来是更多工具,对吧?

What are the first ways you saw it was giving it tools, giving access to the bash, and then later to the file system, and then to more tools, right?

Boris

没错。我们给了它 bash,然后文件编辑是第二个。上次深度探讨时我们聊到一件有趣的事:当你把它建好,它开始用你给的所有工具实际写代码时,Anthropic 内部有过一场辩论:我们是不是该把它留给自己用?因为它突然在工程团队里传开了,让你们所有人都效率大增,对吧?

That's right. Yeah, we gave it bash, then file edit was the second one. One of the interesting things we talked about last time for the deep dive is when you built it and it started to actually write code with all the tools you had, you had an internal debate inside Anthropic: should we just keep it to ourselves? Because it suddenly spread across engineering and it was making all of you a lot more productive, right?

Host

是的,没错。最终我们决定发布,这样就能在真实环境中研究安全性。因为说到安全——我一直在提这个词——Anthropic 作为一个实验室存在的原因就是安全。这是它创立的原因,也是它存在的理由。如果你问 Anthropic 的任何人为什么选择这里,答案都是因为安全。所以说到模型安全,有几个不同的层面。模型层面有对齐和机制可解释性。然后是评估,就像把模型放在培养皿里进行合成研究。再然后你可以在真实环境中研究它,看它实际如何表现。你可以看到用户怎么谈论它,可以看到真实环境中的风险,这样能学到很多。通过这样做,我们让模型变得更安全了。所以事后看来,这完全是正确的决定。

Yeah, that's right. In the end, the decision was to release so that we can study safety in the wild. Because when you think about safety, and I keep talking about the word safety. The reason Anthropic exists as a lab is safety. This is the reason it was founded. This is the reason it exists. If you ask anyone at Anthropic why they chose it, it's because of safety. So if you think about model safety, there are different layers at which to think about it. There's alignment and mechanistic interpretability at the model layer. Then there's evals, which is like putting the model in a petri dish and synthetically studying it. Then you can study it in the wild and see how it actually behaves. You can see how users talk about it. You can see what are the risks in the wild and you actually learn a lot this way. By doing this, we've been able to make the model much safer. So in hindsight, it was totally the right decision.

Host

从你的角度听来很有意思,因为从外部来看,我和很多工程师看到的是:哦,Anthropic 发布了 Claude Code。哇。第一次发布是随 Sonnet 4 一起的吗?最初是 Sonnet 4 还是 Sonnet 4.5?

It's amusing to hear about it from your perspective because from the outside what I saw and what a lot of engineers saw was like, oh, Anthropic released Claude Code. Oh, wow. This for the first release with Sonnet 4 release. Was it with Sonnet 4 originally or Sonnet 4.5?

Boris

我想是二月份正式发布的,但之前应该有个研究预览版。

I think it was for the general availability in February, but I think it was research preview before that.

Host

是的,但刚出来时,我的理解是:哦,这东西写代码挺不错的。后来它变得越来越强。所以从我们的角度看,它就是个非常强大的编码工具,我们开始采用它,用于各种越来越有生产力的部分。我相信它已经成为增长最快的开发者工具之一。我每次听到它其实来自研究、目的是理解人们如何使用模型时,都很惊讶。因为另一方面,有些初创公司一直在刻意打造开发者工具来获取用户,而这个研究工具反而获得了更多用户。

Yeah, but when it came out, my interpretation was like, oh, this thing can write code pretty well. And over time it became a lot more capable. So from our perspective, it was like this really capable coding tool that we just started to adopt and use for all sorts of increasingly productive parts. And it has become, I believe, one of the fastest-growing developer tools. I'm always surprised to hear the story that it actually comes from research and the goal to understand how people use the model. Because on the other hand, some startups have been trying to build developer tools deliberately to get adoption. And yet this research tool is getting a lot more adoption.

Boris

我的意思是,你知道,Anthropic 是个研究实验室,是个安全实验室。产品是附带的东西。产品存在是为了更好地服务研究,让模型更安全。我们就是这么看待一切的。

I mean, this is a you know, Anthropic we're a research lab. We're a safety lab. Product is this kind of thing tacked onto the side. Product exists so that we can serve research better and so we can make the model safer. This is how we think about everything.

发布审查与采用 Launch Review and Adoption

Host

早期有个有趣的时刻,我们当时在做发布评审,决定是否要发布。我记得那一刻,因为我们在房间里。有迈克·克里格、达里奥和其他一些人。我们看着内部采用率图表,那简直是垂直上升。太疯狂了。现在已经是 100%了,对吧?Anthropic 的每个技术人员每天都用 Claude Code,几乎 100%。非技术人员也快接近 100%了,增长非常快。销售团队有一半在用 Claude Code。达里奥问它怎么增长这么快,他说:‘你是不是强迫大家用?’我说:‘没有。我们提供这个工具,大家用脚投票。’我们只是让人们用他们喜欢的工具。

There was this funny moment early on when we had this launch review. And we were deciding whether to launch it. I remember this moment because we were in the room. I think there was Mike Krieger, Dario, and some other folks. We were looking at the internal adoption chart, which was just vertical. It was insane. Nowadays it's 100%, right? Every technical employee at Anthropic uses Claude Code every day. It's pretty much 100%. For non-technical employees, it's also getting quite close to 100%. It's increasing very quickly. Half of the sales team uses Claude Code. Dario had this question about how it grew so fast. He asked, 'Are you forcing people to use it?' And I said, 'No. We offer this tool. People vote with their feet.' We just let people use the tool they prefer.

Boris

对,你看起来不像那种强迫别人用你工具的人。我们的做法就是发布产品,然后倾听用户,和人们交流,看他们怎么用,跟进并改进。现在 Claude Code 平均写了 Anthropic 大约 80%的代码。而且它肯定写了我所有的代码。

Yeah, you don't seem like the person who's exactly forcing people to use your tool. I mean, the way we did it, we just launched the thing, then listened to the users, talked to people, saw how they used it, followed up, and made it better. Now we're at the point where Claude Code writes something like 80% of the code in Anthropic on average. And it writes all my code, for sure.

转向信任 Claude Code Switch to Trusting Claude Code

Host

这对你来说是从你第一次提到开始的,我想是 11 月,它开始写你所有的代码。那个转变是什么时候发生的?是什么让你信任它来写你的代码,或者说你有多信任它?你审查那些代码的程度如何?

This started for you the first time you mentioned, I think it was in November, when it started to write all of your code. When did that switch come? What happened to make you trust it to write your code, or how much do you trust it? How much do you review that code?

Boris

转变是瞬间的,当我们开始使用 Opus 4.5 时。那是在它发布之前,我们内部试用了一段时间。立刻就变了。它能力强太多了,我发现我再也不用打开我的 IDE 了。我卸载了 IDE,因为那时我已经不需要它了。实际上我一个月后才卸载,因为我都没意识到自己已经不用它了。

The switch was instant when we started using Opus 4.5. This was before it came out, we were dogfooding it for a little bit. It was just right away. It's such a more capable model, I just found that I didn't have to open my IDE anymore. I uninstalled my IDE because I just didn't need it at that point. I actually did that like a month later, because I didn't even realize I wasn't using it anymore.

Host

我们很多人都有类似的经历,一旦 Opus 4.5 公开,尤其是寒假期间。我也有类似的经历。我意识到,老实说,这个东西写的代码和我自己在我非常熟悉的代码库、副项目中写的代码一样好,而对于我不太熟悉的代码库或技术,它写得比我好得多。

A lot of us had similar experiences once Opus 4.5 was out in the public, especially over the winter break. I had a similar experience. I realized that this thing actually writes, if I'm being honest with myself, as good code as I would have written in the stack that I'm very familiar with in my code base, my side projects where I know it, and just a lot better than what I could for code bases that I'm not as familiar with, or technologies that I'm not as familiar with.

Boris

是的,老实说,它写的代码比我好。我不想承认,我还是想保留点自尊,但很可能确实如此。

Yeah, I'll be honest, it writes better code than I do. I don't want to go there, I still like to keep my pride, but probably true.

Host

我意识到这一点也是因为 12 月我旅行了一段时间。我过了一个编码假期。我们之前聊过这个,我去了欧洲。我们在不同的时区,有点像游牧一样到处走。太有趣了,因为我整天都在编码,这是我最大的爱好。我每天可能写 10 到 20 个拉取请求。Opus 4.5 和 Claude Code 写了每一个的 100%。我没有手动编辑一行代码。那个月底我发现,Opus 只给我引入了大概两个 bug。而如果是我手写,可能会有 20 个 bug 左右。

I realized this because also in December I was traveling a little bit. I was on a coding vacation. We were talking about this before, but I went to Europe. We were in a different time zone, kind of nomading around. It was so fun, because I was just coding all day, every day, which is my favorite thing to do. I wrote maybe 10 to 20 pull requests every day, something like that. Opus 4.5 and Claude Code wrote 100% of every single one. I didn't edit a single line manually. I realized at the end of that month, Opus introduced me maybe two bugs. Whereas if I'd written that by hand, that would have been, you know, like 20 bugs or something like that.

开发工作流与技巧 Development Workflow and Tips

Host

我们能谈谈你的开发工作流吗?你写了一些关于这个的帖子,很棒。在 Threads 和 X 上。但你能告诉我们你今天如何使用 Claude Code,关于并行性以及你和团队学到并分享的技巧吗?

Can we talk about your development workflow? You have written some threads about this, which is awesome. It's on social media on Threads and on X. But can you tell us how you use today Claude Code in terms of parallelism and tips and tricks that you and the team have kind of learned and share across the team?

Boris

是的,听着,使用 Claude Code 没有唯一正确的方法。所以我可以分享一些技巧,但我认为错误的结论是直接照搬使用。我们构建 Claude Code 的方式是让它可定制。因为我们知道每个工程师的工作流都不同。没有一种方法适合所有人。没有两个工程师有相同的工作流。每个工程师都不一样。

Yeah, look, there's no one right way to use Claude Code. So I can share some tips and things, but I think the wrong conclusion to draw would be to just copy these and use it. The way we built Claude Code is we built it to be hackable. Because we know every engineer's workflow is different. There's no one way to do things. There's no two engineers that have the same workflow. It's just every engineer is different.

Host

工作站设置,对吧?比如键盘、显示器位置,每个人都不同。

Workstation setup, right? Like keyboards, monitor placement, all that everyone has it differently.

Boris

是的,我们就像手工艺人,对吧?你选择你的工具。我们非常在意这一点。所以没有唯一正确的方法。对我来说,我通常的做法是打开五个终端标签页。每个标签页都有一个仓库的检出副本。所以是五个并行的检出。我通常会轮询,在每个标签页中启动 Claude Code。几乎每次我都从普通模式开始,就是在终端里按两次 Shift+Tab。当标签页用完了,我也会溢出。终端标签页只有那么多。我以前经常用网页版,比如 claude.ai/code。那是我溢出的地方。现在我用桌面应用,更方便。Claude Code 已经在我们的桌面应用里好几个月了。它只是 Claude 应用中的一个代码标签页。我真的很喜欢它,因为它内置了工作树支持。这已经存在一段时间了。这对并行性非常好。所以你不需要多个检出。你只需要一个,然后我们会自动为你设置 Git 工作树。你就能得到这种环境隔离。我这样做是因为我真的讨厌在命令行里摆弄 Git 工作树,因为它很繁琐。你需要知道 cd 到...。

Yeah, it's like we're crafts people, right? You choose your tools. We care deeply about it. So there's no one right way to do it. For me, the way that I do it generally is I have five terminal tabs. Each one has a checkout of the repository. So it's five parallel checkouts. Usually I'll kind of round robin and start Claude Code in each one. Almost every time I start in plain mode, so that's like shift tab twice in the terminal. I also overflow as I run out of tabs. There's only so many terminal tabs. I used to use web a lot for this, so like claude.ai/code. That's the place that I overflow to. Nowadays I actually use the desktop app. It's more convenient. Claude Code has been in our desktop app for many months. It's just a code tab in the Claude app. I really like it because it has built-in worktree support. That's existed for a while. That's quite nice for parallelism. So you don't need multiple checkouts. You just have one, and then we automatically set up Git worktrees for you. You get this kind of environment isolation. The reason I do that is I actually just really hate fiddling with Git worktrees on the command line because it's kind of fiddly. You need to know to CD and...

Host

对于那些不太熟悉的人来说,工作树就是你可以检出,而不是有一个单独的本地文件夹,它几乎像是检出一个单独的分支,对吧?然后你可以单独在上面工作,只在合并时才有冲突。

For those who are not as familiar with it, worktree is when you can check out and instead of having a separate local folder, it's almost like checks out separate branch, right? And then you can work on it separately, but not have the conflicts only at that merge time.

Boris

没错。想象你有一个文件夹,但 Git 以非常廉价且容易丢弃的方式制作了五个副本。这样你就得到了隔离。你可以并行工作,这些代码会话不会互相干扰。

That's right. Imagine that you have a folder, but Git makes five copies of that folder in a way that's very cheap and kind of easy to throw away. So you get this kind of isolation. You can work in parallel and the quads don't interfere.

Host

所以你知道对这个的支持,我想你最近添加了原生支持,但对于你的工作流,你坚持使用旧的单独文件夹检出的方式,对吧?

So you know how support for this, which I think you recently added like native support, but for your workflow, you just stuck with the old one of checking out on separate folders, right?

Boris

是的,没错。随着时间的推移,我越来越多地使用桌面应用来做这个。只是因为我不需要这些单独的检出,我只需要并行运行一堆 Claude Code 会话,不用操心。

Yeah, exactly. Over time, I'm using the desktop app more and more for this. Just because I don't need these separate checkouts and I just have a bunch of Claude Code sessions running in parallel and I don't have to think about it.

iOS 应用与并行智能体 iOS app and parallel agents

Host

另一个让我惊喜的是 iOS 应用。每天我醒来,就在手机上启动几个智能体。

The other surprise hit is the iOS app for me. Every day, I start like I wake up and I just start a few agents on my phone.

Boris

哦,原生的那个,对吧?

Oh, the native one, yeah?

Host

对,原生的。就是 Claude 应用里的代码标签页,和 Claude Code 完全一样。只不过它在云端运行,对吧?

The native one, yeah. It's just like it's the Claude app. It's the code tab in the Claude app, and it's the same exact Claude code. Yeah, except it runs in the cloud, right?

Boris

它在云端运行。是的,所以你需要配置环境。你的环境很简单,我们只用钩子来处理。你只需用会话启动钩子来配置。让 Claude Code 高度可定制的好处之一就是很容易做这种配置。老实说,我完全没想到。如果六个月前你告诉我我会在手机上写三分之一甚至一半的代码,那太疯狂了。但这就是我今天在做的事。

It runs in the cloud. Yeah, so you have to kind of configure the environment. Your environment's pretty simple, and we just use hooks for it. So you just use the session start hook and configure it. This is one of the benefits of making Claude code really hackable—it's very easy to do this kind of configuration. And honestly, I would never have predicted this. If you told me six months ago I'd be writing maybe a third or half of my code on a phone, that's crazy. But that's what I'm doing today.

Host

而且你在用并行智能体。你从什么时候开始用的?它怎么改变了你的工作?因为我注意到我自己,其实没用那么多并行智能体,可能一次两个。但我喜欢掌控,尤其是用 Claude。Claude 是一个你可以跟着看的工具。它会告诉你它在做什么。你还可以用学习模式,这个发布得更早,你可以跟着看,它会给你任务。我觉得待在一个标签页里跟着看,模型也很快,我能跟上。我猜你肯定也这么做过,但后来你换成并行之后,你觉得失去了控制,还是说其实没什么关系?

And you're using parallel agents. At what point did you start using them and how has it changed your work? Because one thing I noticed on myself, I don't really use that many parallel agents. I may be like two at a time, but I like to be in charge, especially with Claude. Claude is a tool you can follow along. It tells you what it's doing. You can also have, for example, learn mode, which was shipped a lot earlier, where you can actually follow along. It gives you tasks. I feel that staying in one tab and following along, the model is pretty fast as well. I can kind of keep in touch. I'm assuming at some point you must have done this, but then what happened when you changed to parallel? Do you feel you're losing any control or it doesn't really matter that much?

Boris

是的,我觉得有两种模式或两种工作流可以思考。当你刚接触一个代码库时,学习模式很棒,强烈推荐。对于加入 Claude Code 团队或 Anthropic 的新人,我们建议在 Claude Code 里用 /config,选择输出风格,可以用学习或解释模式。我们通常推荐解释模式,因为它对你不熟悉的新代码库更好。对我来说,一旦你熟悉了代码库,你就只想提高效率。你只想尽可能多地交付,并且高效地做到这一点。所以角色真的变了。我不再深入任务了。我启动一个 Claude,进入计划模式。让它开始做点什么。用 Opus 4.5 时,我觉得它已经不错了。用 4.6 时,它真的做到了。一旦有了好计划,它几乎每次都能一次性实现。所以最重要的是来回沟通,把计划做对。所以我先启动一个,进入计划模式,给它一个提示。当它运行的时候,我切换到第二个标签页,启动第二个 Claude,也进入计划模式。让它运行,然后第三个标签页,第四个。然后当第一个完成通知我时,我再回去。我会跟踪这些。

Yeah, I think there are kind of two modes or two workflows to think about. When you're new to a codebase, learn mode is awesome. Highly recommend it. For people onboarding to the Claude Code team or onboarding to Anthropic, what we recommend is you do /config in Claude Code, pick the output style, and you can do learn or explanatory. We usually recommend explanatory because that tends to be better for new codebases you haven't been in before. For me, once you're familiar with a codebase, you just want to be productive. You just want to ship as much as you can and be effective doing that. So the role really switches. I don't really go deep into tasks anymore. I start a Claude in plan mode. I'll have it kick something off. With Opus 4.5, I think it got there. With 4.6, it just really does it. Once there's a good plan, it will one-shot the implementation almost every time. So the most important thing is to go back and forth a little bit to get the plan right. So what I do is I start one. I enter plan mode. I give it a prompt. As it's chugging along, I'll go to my second tab and start the second Claude also in plan mode. Get it chugging along, then go to the third tab, go to the fourth one. Then maybe I'll go back to the first one when I get notified that it's done. And I'll kind of keep track of that.

Host

你打开还是关闭通知?

Do you turn them on or off?

Boris

我其实两种模式都用。有时候我用 Mac 的专注模式,所以关掉通知,但有时候我也用系统通知。

I actually operate in both modes. Sometimes I use focus mode on the Mac, so I just have it off, but also sometimes I use the system notifications.

Host

而且你在拉取请求方面非常高效。我觉得即使在假期,社交媒体上也很明显——你回应了某个报告 bug 或功能请求的人,一两个小时后它就完成了,因为你做了。你也提到过一天内完成的拉取请求数量,不是为了炫耀,而是作为背景。一个拉取请求通常涉及多少复杂度?有些是超级琐碎的,还是有些实际上是更大的工作?

And you're very productive with pull requests. I think it was very visible even around the holiday breaks on social media—you actually were responding to someone who reported a bug or feature request, and an hour or two later it was done because you did it. You've also talked about the number of pull requests you've done in a day, not to show off, but as context. What does a pull request typically involve in terms of complexity? Are some super trivial or some actually larger pieces of work?

Boris

是的,每个拉取请求差异很大。有时只有几行,有时几百行或几千行。它们都非常不同。变化太大了。以前在 Instagram 的时候,我按代码量算应该是效率最高的两三个工程师之一。所以我一直写很多代码。编程是我表达自己的方式,也是我大脑思考的方式。现在我只是继续这么做,但有了 Claude Code,你写的代码,如果你效率很高,它甚至更多——拉取请求的数量其实低估了实际发生的事。因为在 AI 助手出现之前,那些效率很高的人,很多代码可能只是代码迁移。所以每天提交 20 或 30 个拉取请求的人,很多都是很琐碎的,比如一行改动或者从 A 迁移到 B。现在,我每天提交 20 或 30 个拉取请求,但每个都完全不同。有些几千行,有些几百行,有些几十行,有些只有一行。没有一个是代码迁移,因为 Claude 直接做了那些,我不需要参与。

Yeah, pull requests each one varies a lot. Sometimes it's a few lines, sometimes it's a few hundred or a few thousand lines. They're all very different. It's changed so much. Back when I was at Instagram, I think I was one of the top two or three most productive engineers there just by volume of code written. So I've always coded a lot. Coding is a way I can express myself and it's how my brain thinks. And now I just get to do it, but with Claude Code, the kind of code you write, if you are very productive, it tends to be even more—the number of PRs sort of undersells what's happening. Because people who used to be very productive in the old days before AI assistants, a lot of the code maybe was like code migrations. So people who shipped 20 or 30 PRs every day, a lot of it was pretty trivial, like a one-liner or migrating A to B. Nowadays, I ship 20 or 30 PRs every day, but every PR is completely different. Some are thousands of lines, some are hundreds, some are dozens, some are one-liners. None of these are code migrations because Claude just does those and I don't need to be part of that.

Host

交付这么多代码到生产环境,任何软件专业人士都会问:代码审查怎么办?以前团队的工作方式是——你创建一个拉取请求,提交上去,有强制的人工审查。在谷歌,实际上有两个审查,因为还有一个代码质量审查。这个工作流怎么变了?Claude Code 团队如何看待代码审查,它随时间怎么变化的?

Shipping this much code or this much to production, the obvious question for any software professional is: what about review? The way teams used to work—you make a pull request, put it up, there's a mandatory human reviewer. At Google, there are actually two because there's one on code quality as well. How has this workflow changed? How does the Claude Code team think about code review and how has it changed over time?

Boris

是的,我先说说我以前怎么做代码审查。我以前也是最高产的代码审查者之一。所以既写代码又审查。这其实是在不同时区的好处之一。我不是超人,只是没有会议。我处理代码审查的方式是,每次我要评论什么,就把它记在电子表格里,描述问题。比如有人给函数参数命名不好,我就记在电子表格里。

Yeah, I'll start by talking about how code review used to work for me. I used to be one of the most prolific code reviewers as well. So both writer and reviewer. That's actually one of the benefits of being in a different time zone. I'm not superhuman, I just didn't have any meetings. The way I approached code review is every time I would comment about something, I would drop it in a spreadsheet and describe the issue. So let's say someone named a parameter in a function badly, I would put that in a spreadsheet.

用 Lint 规则自动化繁琐工作 Automating Tedious Work with Lint Rules

Boris

如果有人写了糟糕的 React 模式之类的东西,我会把它记在电子表格里。然后随着时间的推移,我会统计这个表格,每当某一行出现超过三四个实例时,我就会为它写一条 lint 规则。就这样用静态分析来自动化。我以前就是这么做的。对我来说,我一直试图把自己自动化掉,因为有太多事情要做。这是我们工程师的超能力之一:我们能够自动化所有繁琐的工作。很少有其他领域能做到这一点。这是我们独有的能力。我一直很喜欢它,因为它给了我更多的空闲时间,让我能做自己真正喜欢的工作。

If someone did some bad React pattern or something, I would put that in a spreadsheet. And then over time, I would just tally up the spreadsheet and anytime that a particular row had more than three or four instances, I would write a lint rule for it. So, just automate it with static analysis. That's what it used to look like. For me, I've always tried to automate myself away because there's just so many things to do. This is one of our superpowers as engineers: we are able to automate all the tedious work. There are very few other fields where you're able to do this. This is a thing uniquely that we're able to do. I've always enjoyed it because it gives me more free time, and I get to do the work I actually enjoy.

Host

所以今天,情况有点不同,但与此类似。当 Claude Code 写代码时,它通常会在本地运行测试——这是 Claude 在相关时经常决定做的事情——或者它会写新的测试。所以你做了这种验证。当我们对 Claude 代码进行更改时,Claude 也会测试自己。它会在子进程中启动自己,验证自己,并进行端到端测试。这是针对你内部的 Claude 代码实现。所以你有一个测试套件,让它能够自我测试。

So today, the way this looks is a little different, but it mirrors this a bit. When Claude Code writes code, it generally runs tests locally—this is something Claude often decides to do when relevant—or it'll write new tests. So you do this kind of verification. When we make changes to Claude code, Claude will also test itself. It'll launch itself in a subprocess, verify itself, and test itself end-to-end. This is for your internal Claude code implementation. So you have a test suite so it can test itself.

Boris

是的,没错。但它实际上会在 bash 进程中启动自己,然后看看‘嘿,我还能正常工作吗?’所以它会这样做。这不是我们编码进去的。尤其是 Opus 4.5,它开始自发地这样做。它就是想检查一下。所以我们这样做,然后我们还运行 Claude-P。这是 CI 中的 Claude 智能体 SDK。所以 Anthropic 的每个拉取请求都由 Claude code 进行代码审查。这实际上能捕获大约 80% 的 bug。这是第一轮代码审查。Claude 会自动处理其中一些。有些它会留给人类,因为它不确定该怎么做。总有一个工程师进行第二轮代码审查。总得有一个人参与批准更改。

Yeah, that's right. But it'll literally launch itself in a bash process and just see, 'Hey, do I still work?' So it'll do this. This is something we didn't code in. With Opus 4.5 especially, it just started spontaneously doing this. It just wants to check. So we do this, and then we also run Claude-P. This is the Claude agent SDK in CI. So every pull request at Anthropic is code reviewed by Claude code. That actually catches maybe 80% of bugs. It's the first round of code review. Claude will automatically address some of these. Some it'll leave to a human because it's not sure what to do. There's always an engineer that does the second pass of code review. There always has to be a person in the loop approving the change.

Host

所以在团队中,任何东西投入生产之前,工程师都会检查它。

So on the team, before anything goes into production, an engineer does look at it.

Boris

是的。

Yes.

何时跳过人工审查 When to Skip Human Review

Host

说到代码审查,你会对每种项目都这样做吗?还是特别因为你知道这有实际影响,人们依赖它,有很多用户?我换个方式问。你能看到哪些情况下你不会让工程师审查代码?那会是什么情况?

As you think of code review, would you do this for every type of project, or is this specifically because you know this has real-world impact, people depend on it, there are a lot of users? Let me put it the other way around. Can you see places where you would just not have an engineer review code? What situations would that be?

Boris

我认为这取决于它的使用方式。如果你在构建一些个人副项目,你可以直接 YOLO 到主分支。即使在 AI 之前,你也不会审查。你只是相信自己,或者直接部署到生产环境,或者 SSH 到生产环境做更改。你懂那种情况,对吧?

I think it depends on how it's used. If you're building some personal side project, you can just YOLO straight to main. Even before AI, you would have not reviewed. You just trust yourself, or you just shipped to production or SSH'd into production and made changes. You get that kind of stuff, right?

Host

没错。最早的内部版本 Claude code,我直接提交到主分支。但一旦有了用户,对于 Anthropic,我们的主要客户群是企业,这是最关心的。出于安全原因,安全性非常重要,隐私也很重要。这些都是相关的。对我们的客户也非常重要。所以因为这是企业产品,它必须安全。我们必须确保它达到一定标准。所以我们肯定使用了很多自动化,但至少目前,必须有一个人参与以确保万无一失。

Exactly. The very first versions of Claude code that were internal, I committed straight to main. But then, as soon as you have users, and for Anthropic, our main customer base is enterprises, this is what we care about the most. For safety reasons, security is really important, privacy is important. These are all related. It's also very important for our customers. So because this is an enterprise product, it has to be secure. We have to make sure it meets a certain bar. So we definitely use a lot of automation, but at least for now, there has to be a human in the loop just to make sure.

确定性检查与非确定性检查 Deterministic vs Non-Deterministic Checks

Host

关于 LLM 的一个已知问题是它们是非确定性的。让 LLM 作为审查者,Claude 进行审查,它会给出很好的反馈,但你怎么处理这样一个事实:即使它有能力,你也不能保证它总能捕捉到问题?你在这个循环中做了什么来使其确定性吗?例如,linting 是非常确定性的。你有没有考虑过结合这些想法,或者你在代码库中使用 linter,还是觉得没必要?

One thing that is known about LLMs is they're non-deterministic. By putting an LLM as a reviewer, Claude doing a review, it will give good feedback, but how do you deal with the fact that you can't be sure it will always catch an issue, even if it's capable? Are you doing anything in this loop to make it deterministic? For example, linting is very deterministic. Have you thought of marrying some of these ideas, or are you using linters on the codebase, or have you found no need?

Boris

当然。我们有类型检查器,有 linter,我们运行构建。Claude 实际上非常擅长写 lint 规则。我现在做的,不再是在电子表格里统计,而是当同事提交拉取请求,我觉得这可以 lint 时,我就会让 Claude 在那个 PR 里写一条 lint 规则。我们有一个 GitHub 应用,你可以在任何拉取请求或问题上标记 @Claude。我每天都用这个。所以你需要这些确定性的步骤。此外,还有一些方法可以让 Claude 更确定一些。例如,你可以做最佳 N 选,让它多次运行。这很容易做到。我们内部使用的代码审查技能是开源的,在 Claude Code 仓库中可用。我们所做的就是启动并行智能体来做事,然后启动并行去重智能体来检查误报。本质上,最佳 N 选的实现方式就是直接说‘Claude,启动三个智能体来做这个。’就这样。

Absolutely. We have type checkers, we have linters, we run the build. Claude is actually so good at writing lint rules. What I do now, instead of tallying stuff up in a spreadsheet, when a coworker puts up a pull request and I think this is lintable, I'll just ask Claude, 'Please write a lint rule for this in that PR.' We have a GitHub app you can tag @Claude on any pull request or issue. I use this every single day. So you want these deterministic steps. Also, there are ways to get Claude to be a little more deterministic. For example, you can do best of N, have it do multiple passes. This is quite easy to do. The code review skill we use internally is open source and available in the Claude Code repo. All we do is launch parallel agents to do stuff, and then launch parallel deduping agents to check for false positives. Essentially, best of N, the way you implement it is just say, 'Claude, start three agents to do this.' And that's it.

赞助商插播:WorkOS Sponsor Break: WorkOS

Host

Boris 刚刚谈到了构建企业基础设施层。认证、权限、安全性,这些都必须先做好,才能交付给真正的客户。现在是谈论我们本季赞助商 WorkOS 的好时机。如果你在构建任何 SaaS,尤其是 AI 产品,那么认证、权限、安全性和企业身份可能会悄然变成一项长期投资。SAML、边缘情况、目录同步、审计日志,以及企业客户期望的所有东西。构建这些关键部分需要大量工作,维护它们更是如此。但你不必这样做。WorkOS 将这些构建块作为基础设施提供,这样你的团队就可以专注于真正让你的产品独特的东西。这就是为什么 Anthropic、OpenAI 和 Cursor 等公司已经在使用 WorkOS。优秀的工程师知道什么不该构建。如果身份管理对你来说是其中之一,请访问 workos.com。接下来,让我们回到与 Boris 一起构建 Claude Code 的话题。

Boris just talked about building that enterprise infrastructure layer. The auth, the permissions, the security, that has to all work before you can ship to real customers. This makes it a great time to speak about our season sponsor, WorkOS. If you're building any SaaS, especially AI product one, then authentication, permissions, security, and enterprise identity can quietly turn into a long-term investment. SAML, edge cases, directory sync, audit logs, and all the things enterprise customers expect. It's a lot of work to build these mission-critical parts, and then some more to maintain them. But you don't have to. WorkOS provides these building blocks as infrastructure, so your team can stay focused on what actually makes your product unique. That's why companies like Anthropic, OpenAI, and Cursor already run on WorkOS. Great engineers know what not to build. If identity is one of those things for you, visit workos.com. And with this, let's get back to building Claude Code with Boris.

Claude Code 架构与安全层 Claude Code architecture and safety layers

Host

从架构角度来看,Claude Code 是如何工作的?作为一名工程师,我该如何想象它的设置?我们在深度探讨中已经涉及了一些内容,我记得你告诉我,你一开始有一些非常复杂的想法,但后来简化了很多?

How does Claude Code work in terms of architecture? So, as an engineer, how can I imagine it's set up? We covered some of this in the deep dive, and I think you told me that you had some pretty complex ideas when you started, and you just simplified a lot of it?

Boris

是的,是的。它非常简单。没什么复杂的。有一个核心查询循环。它使用一些工具。我们一直在删除这些工具,也一直在添加新工具。我们总是在实验。所以,它有一个核心的智能体部分。然后是 2E 部分。实际上还有大量围绕安全性的不同组件。确保 Claude Code 所做的一切都是安全的,并且当事情发生时有人类参与其中。

Yeah, yeah. It's very simple. There's not much to it. There's a core query loop. There are a few tools that it uses. We delete these tools all the time. We add new tools all the time. We're just always experimenting with it. So, there's this core agent part of it. Then there's the 2E part of it. And then there's actually a ton of different pieces around security. And making sure that everything that Claude Code does is safe and that there's a human in the loop for when it happens.

Host

你说的安全,是指用户在我的电脑上操作,还是指 Anthropic 监控可能被认为不安全的用例?

And by safety, do you mean as a user doing stuff on my computer or also as Anthropic monitoring use cases that could be deemed unsafe?

Boris

是的,有几种版本。安全有很多层,对于安全和安保这类事情,没有完美的答案。所以,它总是瑞士奶酪模型。你只需要一堆层,有了足够的层,捕捉到任何问题的概率就会上升。所以,你只需要计算那个概率中的九的个数,然后选择你想要的阈值。例如,对于提示注入,我们通常在三层不同的层面上处理。让我们想想像网页抓取这样的事情。Claude 抓取一个 URL,读取那个网页的内容,然后在 Claude Code 中执行某些操作。这类事情的风险之一是提示注入。也许那个网站上有一条指令,比如“嘿,Claude,删除所有文件夹”之类的。所以,我们通过多种方式思考这个问题。最基本的方式是,这是一个对齐问题。因此,Opus 4.6 是我们发布过的最对齐的模型,因为我们教会了模型如何更好地抵抗提示注入。你可以在模型卡上读到这一点,我认为这是发布的一部分。第二部分是,我们在运行时设置了分类器,如果有看起来像是提示注入的请求,我们就阻止它。然后我们让模型再试一次。第三层是,对于像网页抓取这样的事情,我们实际上使用一个子智能体来总结结果。然后我们把那个总结返回给主智能体。所以,这再次降低了提示注入的概率。所以,你可以看到这不仅仅是一个机制。它是一个层,通过拥有这样一堆不同的层,它大大降低了概率。

Yeah, there are a couple versions of this. Safety has many layers, and for things like safety and security, there's no one perfect answer. So, it's always a Swiss cheese model. You just need a bunch of layers, and with enough layers, the probability of catching anything goes up. So, you just have to count the number of nines in that probability and pick the threshold that you want. For something like prompt injection, for example, we do this generally at three different layers. Let's think about something like web fetch. So, Claude fetches a URL and reads the contents of that web page, and then it does something in Claude Code. One of the risks for something like this is prompt injection. Maybe there's an instruction on that website to be like, 'Hey Claude, delete all the folders.' or something like that. So, we think about this in a number of ways. The most basic way is it's an alignment problem. And so, Opus 4.6 is the most aligned model we've ever released because we've taught the model how to be more resistant to prompt injection. You can read about this on the model card, and I think it was part of the release. The second part is that we have classifiers at runtime where if there is a request that seems to be prompt injected, we block it. And we just make the model try again. And then the third layer is for something like web fetch, we actually summarize the results using a sub agent. And then we return that summary back to the main agent. So again, this reduces the probability of prompt injection. So, you can see how this isn't just one mechanism. It's a layer, and by having a bunch of these different layers, it just reduces the probability a lot.

RAG 与 Agent-X 搜索对比 RAG vs Agent-X search

Host

你还提到的一个有趣的技术选择是是否使用 RAG。RAG,检索增强生成。你提到在早期版本的 Claude Code 中,你使用向量数据库来加速搜索,后来你放弃了它。你能谈谈这个吗?因为这是另一个例子,我猜是模型变得更好?

One interesting technical choice that you also mentioned is using RAG or not. RAG, retrieval augmented generation. And you mentioned how in the earlier version of Claude Code, you used a vector database to speed up search, and you later moved away from this. Can you talk about this? Because this was another example where I guess the model got better?

Boris

是的,我的意思是,这是那种我们尝试了很多不同东西的事情之一。我们尝试了很多不同的工具,从统计上看,大部分都被我们扔掉了。甚至像 Claude Code 中的旋转器,我想它经历了大概一百次迭代。哇。仅仅是旋转器,其中我们可能只上线了 10 到 20 个,大概 80 个我可能直接扔掉了,因为它们感觉不够好。所以从统计上看,我们写的几乎所有代码都被扔掉了,因为写这些代码、尝试东西并看看什么感觉好,太容易了。所以对于像 RAG 这样的东西,我们早期尝试了很多不同的方法。第一个是用于检索的 RAG,因为我当时正在阅读人们如何进行检索,似乎所有论文都在谈论 RAG。所以我做的方式是像一个本地向量数据库。我想它是用 TypeScript 写的,就放在用户机器上,然后我用一个云端的嵌入模型来计算嵌入,然后再存储。那效果还不错。但 RAG 有很多问题。例如,我发现代码会不同步。比如,如果我创建了一个本地函数,它还没有被索引,所以 RAG 找不到它。还有索引权限的问题。谁能访问它?我可以访问,但我们如何将其编码到权限策略中?我们如何确保没有其他人能访问它?我们如何确保如果公司内部有一个恶意的 IT 人员,他们无法访问别人的数据?这一点非常重要,我们必须考虑。所以我们决定它虽然能用,但也有很多缺点。于是我们尝试了很多其他东西。其中之一是直接用模型递归地索引所有内容。那是个挺酷的想法。另一个版本是我们尝试了 glob 和 grep。我们尝试了很多不同的东西。结果发现 Agent-X 搜索的表现超过了所有其他方法。什么是 Agent-X 搜索?它只是 glob 和 grep 的一个花哨说法。仅此而已。很好。所以,模型变得足够好了,你意识到它们可以相当高效地使用这些工具。这实际上部分受到了我在 Instagram 经历的启发。因为在 Instagram,点击跳转到定义不起作用,因为开发栈一半时间都是坏的。我想现在好多了。所以工程师们会怎么做呢?假设你在找函数 foo 的定义。他们不会用点击跳转,而是使用全局索引(在 Meta 非常好用),然后搜索 foo 左括号。这效果很好。有趣的是,这对模型来说也效果很好。一个领域的想法如何能应用到另一个领域,很有意思。

Yeah, I mean this is one of those things where we try so many different things. We try so many different tools, and just statistically most of them we throw away. Even something like the spinner in Claude Code, I think it's gone through like a hundred iterations I want to say. Wow. Just the spinner, and out of those we landed maybe like 10 or 20 in production, and like 80 of them I probably just threw away because it didn't feel good enough. So just statistically, almost all the code we write we throw away because it's so easy to write this code and try stuff and see what feels good. So for something like RAG, we tried a bunch of different approaches early on. The first one was RAG for retrieval, because I was reading up on how people were doing retrieval and it seemed like all the papers were talking about RAG. And so the way I did it was it was like a local vector database. I think it was written in TypeScript and it just lived on the user machine, and then I was using some embedding model that was in the cloud to compute embeddings before storing it. And that worked pretty good. But there are a lot of issues with RAG. For example, I was finding that the code drifted out of sync. Like if I make a local function, it's not yet indexed, and so RAG isn't going to find it. There's also this question of how exactly is the index permissions. So who can access it? I can access it, but then how do we encode that in permission policies? How do we make sure no one else can access it? How do we make sure that if there's a rogue IT person within the company, they can't access someone else's data? This is really, really important that we think about this. And so we just decided it was sort of working, but it also has a lot of downsides. And so we tried a bunch of other stuff. One of them was just using the model to index everything recursively. That was kind of a cool idea. There was another version where we just tried glob and grep. We tried a bunch of different stuff. It turned out that Agent-X search just outperformed everything. And what is Agent-X search? It's just a fancy word for glob and grep. That's all it is. Nice. So, the model both got good enough and you realized that they can use these tools pretty efficiently. And this was partially inspired honestly by my experience at Instagram. Because at Instagram, click-to-definition didn't work because the dev stack was just broken like half the time. And I think now it's better. So what engineers would do instead is, let's say you're looking for the definition of the function foo. Instead of click-to-definition, you would use the global index, which is quite good at Meta, and then you would search for foo opening parenthesis. And this worked pretty well. And it's funny because this works for the model pretty well, too. Interesting how one idea from one area can come to the other.

权限系统与沙箱 Permission system and sandboxing

Host

我们之前还谈到过 Claude Code 的一个更高级的部分,就是权限系统。你能谈谈它复杂在哪里吗?另外,你最近开源了沙箱,对吗?

One of the more advanced parts of Claude Code that we also previously talked about is the permission system. Can you talk about what was complex about it? And also you recently open-sourced sandboxing, right?

Boris

权限管理非常复杂。就像其他与安全相关的事情一样,它是一个瑞士奶酪模型。有许多分类器在运行,以确保命令是安全的。我们还进行静态分析,以确保命令是安全的。作为用户,你还可以将你知道安全的特定模式列入白名单。例如,一些标准的 Unix 工具我们预先允许,因为我们知道它们是只读的,我们知道它们不能泄露数据或类似的东西。

Permissioning is really complex. Like everything else that has to do with security, it's a Swiss cheese model. There are a number of classifiers that run to make sure the command is safe. And there's also static analysis that we do to make sure the command is safe. As a user, you can also allow list particular patterns that you know to be safe. So, for example, some standard Unix utilities we pre-allow because we know they're read-only, because we know they can't exfiltrate data or anything like this.

权限系统与安全 Permission System and Safety

Host

所以我们不会直接提示你请求权限。但实际上,很少有工具属于这一类,因为即使是像 find 这样的命令,也有办法通过系统标志来执行任意代码。甚至 set 命令也有办法做到这一点。所以,这些 Unix 工具中有很多隐秘的细节,实际上并不像你想象的那么安全。因此,我们希望默认情况下对允许的操作保持相当保守的态度。不过,作为用户,你可以配置一个允许列表。例如,你可以说这些模式是允许的,这些模式是不允许的。我们让你定义这个列表,并且我们也会检查这个允许列表以确保它是安全的。

So, we'll just won't prompt you for permission. But actually quite few tools fall into this category because even something like the find command, there's actually a way to execute arbitrary code as part of that command because there's like system flags you use for this. Or even something like the set command, there's ways to use this. So, there's just like all this arcania about these various Unix utilities where it's actually not as safe as you think. And so, we want to be by default fairly conservative about what we allow by default. As a user though, you can configure an allow list. So, you can say for example like these patterns are allowed, these patterns are not allowed. And so, we let you define that and we also check this allow list to make sure that it's safe.

Boris

是的,然后你有一个简洁的权限系统,每次运行需要权限的命令时,你可以选择运行一次、在会话期间运行,或者全局允许它继续运行,对吧?

Yeah, and then you have this neat permission system where every time you run a command that needs permission, you can decide to run it once, to run it for either the session or whatever it makes sense, or just globally allow it to go forward, right?

Host

没错。这是一个有趣的产物。这实际上是在 Quadcode 的第一个版本中实现的。权限就是这样工作的。那是第一个版本,大概是 2024 年 9 月,第一个内部版本。我记得当时我们不确定智能体安全性是否能够解决。所以,内部安全团队有很多反对意见,他们说:“你不能让模型直接运行 bash 命令。那不安全。那怎么办?这个问题解决不了,所以我们不能发布这个。”我和 Ben Mann 一起 brainstorm,Ben 是实验室团队的创始人,他是 Anthropic 的联合创始人之一。实际上,是他把我招进 Anthropic 的。我们想出了权限提示的方式来解决这个问题。如果你不确定,就问人类,然后他们来决定。

That's right. This is a funny artifact. This was actually in the very first version of Quadcode. This is the way permissions worked. This is the very first release. This was like September 2024, the first internal release. I remember at the time we weren't sure whether agentic safety could even be solved. And so, there was actually a lot of pushback internally from safety teams because they were like, "Okay, you can't just let the model run bash commands. That's unsafe. So, what do you do? This is not a solvable problem. So, we can't launch this." I brainstormed with Ben Mann, and Ben started the labs team. He's one of the founders at Anthropic. He's actually the person that hired me to Anthropic. We just came up with permission prompts as the way to do this. If you're not sure, just ask the human and then they can decide.

头衔与通才文化 Titles and Generalist Culture

Host

我想问问你 Anthropic 的软件工程总体上是怎么做的。第一个问题,从外部来看比较正式的一个,就是头衔或者没有头衔。Anthropic 的每个人头衔都一样,都是技术成员。为什么会这样?这带来了什么结果?

I want to ask you about how software engineering is done in general in terms of Anthropic. And one of the first questions, which is a more formal one from the outside, is titles or lack of them. Everyone at Anthropic has the same title, member of technical staff. Why did this happen and what does this result in?

Boris

这基本上就是大家都没有头衔,对吧?除了一个例外。我认为这承认了每个人都在摸索。如果你仔细观察人们的工作,其实都很相似,而且相当通才。如果你和普通的软件工程师交谈,他们可能不仅仅做编码,还可能做一些设计,和用户交流,写自己的产品需求,写软件,同时也做研究。他们可能写产品代码,也写基础设施代码。在 Anthropic,有很多通才。这也是我倾向于这里的原因之一。我认为“技术成员”这个头衔在人们互相交流时编码了这一点,即使他们彼此不认识。如果没有这个头衔,默认情况是,我在 Slack 上看到你的名字,下面写着“软件工程师”,然后我就会想,好吧,你是做编码的,所以我不会问你产品问题。但当每个人的头衔都是“技术成员”时,默认你会认为每个人什么都做。这就在人与人之间反转了这种关系,即使你们不太熟悉。

This kind of like everyone basically no titles, right? Except for one. I think it's kind of an acknowledgement that everyone just is figuring stuff out. And if you kind of squint and look at the work people are doing, it's all quite similar and it's quite generalist. If you talk to the average software engineer, they might not just be doing coding; they might also be doing a little design. They might also be talking to users. They might be writing their own product requirements. They might be writing software and also doing research. They might be writing product code and also infrastructure code. At Anthropic there's a lot of generalists. This is also one of the reasons that I gravitated towards it. And I think member of technical staff just kind of encodes this in the way that people talk to each other even if they don't know each other. Without this title, the default would have been I see your name on Slack and under your name it says software engineer. And then I'm like, well okay I guess you're the coding person and so I'm not going to ask you product questions. But when everyone's title is member of technical staff, by default you assume everyone does everything. And so it kind of inverts this relationship between people even if you don't know each other well.

Host

在某种程度上,这是一种融入结构中的乐观主义。我认为这也是未来的一个缩影,因为我认为软件工程正在朝这个方向发展。我认为每个学科都在朝这个方向发展,即更多的通才模式。

In a way it's kind of this optimism built into the structure. I think it's also a glimpse of the future because I think this is where software engineering is going. I think this is where every discipline is going, is more of this generalist model.

Boris

在软件工程中确实有这种感觉。我听到马克·安德森说过一个有趣的评论,他说科技界正在发生一场墨西哥对峙:设计师说他们现在实际上在做产品管理和工程工作;工程师说我们在做设计;每个人都认为自己在做别人的工作,他们站在那里说“我也在做你的工作”。但实际上,每个人的角色都在扩展,这主要归功于 AI,因为它让工程师更容易做产品工作,或者让产品人员更容易做工程工作,等等。所以正如你所说。

Definitely it feels like it in software engineering. And I heard this funny comment by Mark Andreessen how he said that there's this Mexican standoff happening in the tech world where the designers are saying that they're actually now doing PM and engineering work. The engineers are saying that we're doing design and everyone thinks they're doing the work of the others and they're kind of standing there like I'm doing your work as well. But in reality, everyone's role is expanding, most of it thanks to AI because it makes it easier for an engineer to do product work or for a product person to do engineering work and so on. So just what you've said.

Host

我记得去年六月或七月,我走进办公室,数据科学家们——有一排数据科学家坐在 Quadcode 团队旁边,至少当时是这样。我走进去,我们 Quadcode 团队的数据科学家屏幕上开着 Quadcode。他正在使用它,我说:“这很有趣,因为你是数据科学家。你为什么用终端?你没有安装 Node.js,因为我们当时依赖 Node.js。我说:“你是在内部测试吗?还是只是想弄清楚这个东西怎么工作?”他说:“不,不,我在用它运行查询。”他只是在用 Quadcode 运行 SQL,终端里有小的 ASCII 可视化。然后下一周,整排数据科学家的电脑上都运行了 Quadcode。这还扩展了。所以如果你看看今天的团队,在 Quadcode 团队,每个人都写代码。工程师写代码,我们的工程经理写代码,设计师写代码,数据科学家写代码,我们的财务人员也写代码。团队里的每个人都写代码。我认为部分原因是 Quadcode 让写代码变得非常容易。你不需要真正理解代码库,就可以直接深入并轻松地做小改动。但另一件事是,人们能够更多地使用 Quadcode 来完成他们的工作,无论是财务预测还是数据科学等等。通过这样做,实际上很容易跨界,用它来写一点代码。所以这只是一种试水的方式。

I remember back in June or July of last year, I walked into the office and the data scientists—there's a row of data scientists that sit right next to the Quadcode team, at least at the time. And I walked in and our data scientist for the Quadcode team had Quadcode up on his monitor. And he was using it and I was like, "This is interesting because you're a data scientist. Did you have like why are you using a terminal? You didn't have Node.js installed because we depended on Node.js back then. I was like, "Are you dogfooding it? Are you just trying to figure out how this thing works or something?" He's like, "No, no, I'm using it to run queries." He was just using it to run SQL and it had little ASCII visualizations in the terminal. And then the next week the entire row of data scientists had Quadcode running on their computers. And this expanded. And so if you look at the team today, on the Quadcode team, everyone codes. The engineers code, our engineering manager codes, designers code, data scientists code, our finance guy codes. Everyone on the team codes. And I think part of it is Quadcode just makes it so easy. So you don't really have to understand the code base, you can just dive in and make small changes quite easily. But I think another thing is people are able to use Quadcode to do their jobs more, whether it's financial forecasts or data science or whatever. And by doing this it's actually quite an easy crossover to just use it to write a little bit of code also. So it's just a way to dip your toe in the water.

PRD 与产品流程 PRDs and Product Process

Host

关于你们的工作方式,另一件有趣的事情是,Katwo 谈到,我猜你的头衔是一样的,但人们可能会更倾向于某个角色,我理解她更偏向产品角色。但你说过,在 Anthropic 内部其实并不怎么写 PRD。PRD,即产品需求文档,是大型科技公司以及越来越大的初创公司中一个众所周知的产物,你写一个规范,写下你的想法,大家达成一致,然后发送出去,就知道要构建什么了。但显然你们不怎么这么做,或者根本不这么做。

One other interesting thing about how you work is Katwo was talking about she is I guess your title is the same but people might gravitate to a role a little bit more and I understand she's a little bit more on a product role. But you said that PRDs are just not really written inside Anthropic. And PRDs, product requirement documents, are a well-known artifact across big tech and increasingly over larger startups where you write a spec and the idea is that you write down your thoughts, people align, you send it over, and now you know what to build. But apparently you're not doing much of this or at all.

Anthropic 的原型文化 Prototyping culture at Anthropic

Host

我认为部分原因是 Anthropic 仍然是一家初创公司。所以你实际上不需要和那么多人对齐。通常你只需要在 Slack 上讨论一下或者直接做就行了。但还有一部分原因是,Kat 曾经是工程经理,她非常懂技术。我认为我们的产品团队也是这么想的:最好直接发个 PR。你做了很多原型设计。所以当我们谈论你早期如何构建 Claude 核心时,你展示了一个关于数字的完整线程。我记得你为待办事项列表做了 15 到 20 个原型,而且都是可交互且能运行的。让我惊讶的是,根据我过去的科技行业经验,你说你在一天半内就完成了这 20 个原型,全部尝试过并获得了感觉,这对我来说简直不可思议。通常这需要一两周,而且人们不会做 20 个,只会做 3 个。那么,你是否看到原型设计和构建展示的增加,而不是写文档?

Some of this I think is because Anthropic is still a startup. So you don't actually have to align with that many people. Usually you can just kind of talk about it or do it in Slack or whatever. But yeah, also part of it is, like Kat used to be an engineering manager. She's extremely technical. And I think this is the way that our product team thinks about it too: better just send a PR. You're doing a lot of prototyping instead. So that's also something where when we talked about how you were building Claude core early on, you were showing actually you had a whole thread about the number. I think you did like 15 or 20 prototypes for the to-do list and all of them interactive and working. And what surprised me compared to my past tech experience, and you said that well, you did this in like a day and a half. All 20, tried it out, got a feeling for it, which is incomprehensible for me. It would have taken a week or 2 weeks and people would have not done 20. They would have done three. So are you seeing an increase in prototyping and in building and showing instead of writing things?

Boris

是的,绝对如此。我的意思是,我们团队的文化是:我们不写东西,我们只是展示。回想以前有点困难,因为现在原型设计已经深深融入我们的构建方式中。每件事都会多次原型化。比如,我们本周推出了智能体团队,这是我们实现 swarm 的方式。这非常令人兴奋,因为它让 Claude 能够更长时间、更自主地完成更多工作。你有一堆不相关的上下文窗口,智能体之间可以进行这种通信。它们可以做得更多。这是 Daisy、Suzanne 和其他团队成员以及 Karen 花了几个月时间原型化的东西。他们尝试了可能数百个版本,才获得了一个感觉非常好的用户体验。这真的很难做对。如果我们从 Figma 中的静态模型或纯设计文档开始,我们根本不可能发布这个产品。这是一个你必须构建、感受并体验的东西。对我来说,其中一个重要的收获是,我们可能应该更多地原型化,更大胆一些,或者放弃关于构建原型需要多长时间或谁需要构建的先入之见。过去总是需要工程师来构建,但现在可能不再是这样了。

Yeah, absolutely. I mean on our team the culture is we don't really write stuff. We just show. It's a little hard to reflect back on the time before because I think now prototyping everything is so baked into the way that we build. It just everything is prototyped multiple times. Like, we launched agent teams over this week. This is our implementation of swarms. It's very exciting because it just lets Claude do more work for longer more autonomously. You have a bunch of different uncorrelated context windows and you have this kind of communication between agents. They can just do more. This is something that Daisy and Suzanne and other folks on the team and Karen, they prototyped this for months. And they tried all in all probably hundreds of versions of this before they got a user experience that felt really good. It was just really really hard to get right. There's just no way we could have shipped this if we started with static mocks in Figma or if we started with a pure design doc or something like this. It's a thing that you have to build and you have to feel and you have to see how it feels. And to me one of the big takeaways even from there was like we probably should prototype more and just be more daring or just release your priors of how long it took to build a prototype or who needed to build. Back then it was always an engineer that needed to build, but it's probably not true anymore.

Host

是的,没错。我的意思是,我们现在所处的世界,我们不知道正确答案是什么。回想过去的构建方式,构建成本很高,所以你必须在出手前花大量精力仔细瞄准。因为一旦出手,就很难纠正方向。你只能开很少的枪。但现在情况变了。构建成本非常低,但我们也不知道目标在哪里。所以我们只能尝试,看看什么感觉好。这非常具有探索性。我认为其中很大一部分是谦逊,就我个人而言,我有一半的时间是错的。我可以说我的大部分想法都不好,至少一半是坏的。在尝试之前,我不知道哪一半是坏的。

Yeah, that's right. I mean, we're in this world right now also where we just don't know what the right answer is. I think back in the old way of building, the cost of building was high and so you had to actually spend a lot of effort to aim very carefully before you take your shot. Because after you take your shot, it's very hard to course correct. You can only take so few shots. But now it has changed. The cost of building is very low, but also we don't know where we're aiming. So we just have to try and we have to see what feels good. And it's just very very exploratory. And I think also a big part of it is humility where, you know, personally I'm wrong like half the time. I'd say like most of my ideas are bad. At least half of them are bad. And I don't know which half until I try it.

Boris

嗯。有时也会从别人那里得到反馈。没错。就像我必须自己尝试,然后看看别人怎么想,因为我的直觉并不总是和别人一致。当你展示这些关于任务如何构建的原型时,你告诉我你构建了原型,然后你的过程总是先自己看看,尝试一下,获得感觉,然后对于你觉得好的,你展示给别人,有时他们会反馈说‘不,这不行’。有时当感觉不错时,你会更广泛地分享。所以我觉得这是一种混合,对吧?有时你可以自己决定,有时你得到反馈,最终一些好主意就出来了。

Mhm. And then get feedback from others as well sometimes. That's right. It's like I have to try it myself and then I have to see what others think because my intuition does not always match others. When you were showing these prototypes of just how the tasks were built, you were telling me that you built the prototypes and then your process was always you first looked at it, you tried it out, you got a feel for it, and then for the ones that you felt were good, you showed it to others and sometimes they give you feedback like nah, this doesn't work. And then sometimes when it felt good, then you shared it even broader. So I feel like it's a mix, right? Where sometimes you can decide already and then sometimes you get feedback and then eventually some good ideas come out of it.

Boris

是的,有很多这样的例子。比如我们推出了文件读取和文件搜索的压缩视图,因为现在的模型非常智能体化。我觉得一半的屏幕都是这些文件读取,而我实际上并不关心。我读了一个东西,我不在乎它是什么。所以我们把它压缩了,让输出更易读。在大约 30 个原型之后,我非常喜欢它。为了让感觉很好很干净,我们付出了很多努力。我们在 Anthropic 内部向员工推出了大约一个月,让每个人都试用,然后我根据所有反馈修复了大概十几个 bug,做了十几个调整。我们对外发布后,几乎所有用户都喜欢,但也有一些用户不喜欢,因为他们想要更展开的输出。所以在 GitHub 问题上,我和人们反复沟通,问他们不喜欢什么。人们给了很多反馈。我发布了另一个版本。然后一些人喜欢,一些人不喜欢。所以我再次迭代,让它变得更好。实际上我认为它已经差不多了,人们可以按照自己的方式配置,但默认设置仍然很好。但这只是过程。我们有时做对了。我们必须向用户学习。我们希望听到人们的意见,这样我们才能做对。

Yeah, and there's a lot of examples of this. Like we launched this kind of condensed view for file reads and file search just because the model is just so agentic now. I felt like half the screen is these file reads and I actually don't care. I read a thing. I don't really care what it is. And so, we condensed this down to make the output a little bit more readable. I really liked it after probably 30 prototypes or something like this. It took so much effort to make that feel really good and clean. We rolled it out to employees at Anthropic for about a month, and we had everyone dogfood it, and I fixed another probably dozen bugs, dozen tweaks based on all this feedback. We launched it externally, and almost all users liked it, but there were a few users that didn't because they want more expanded output. And so, on the GitHub issue, I was just going back and forth with people to be like, what don't you like? And people give a lot of feedback. I shipped another version. Then, some people liked it, some people didn't. And so, I iterated it again, and kind of made it good. And it's actually I think almost there where people can configure it the way that they want, but still the default is really good. But, this is just the process. We get it right some of the time. We have to learn from our users. We want to hear from people so we can get it right.

Host

你在工作中使用工单系统吗?就是那种记录‘好的,这是工作’的系统?还是说你有工作就直接做?

Do you use ticketing systems for your work? Where you capture like, all right, here's the work. Or do you just pretty much do the work as it comes in?

Boris

在 Anthropic,我们让团队自己决定。在 Claude Code 团队,我们让每个人自己决定。不同的人使用不同的方式。例如,我不使用工单系统。有些人喜欢用 Asana 或笔记之类的。我看到的最酷的事情之一,大概是三个月前,我们推出了插件。我们推出它的方式是,Daisy 在一个周末,她有一个非常早期的 swarm 版本。她让 swarm 运行,并告诉它:‘你的工作是构建插件。你必须提出一个规范,然后创建一个 Asana 看板并拆分成任务。然后所有不同的智能体必须构建它。’她设置了一个容器,并将 Claude 设置为危险模式。她让它运行了整个周末。它生成了几百个智能体。它们在 Asana 看板上创建了 100 个任务。然后,它们实现了它。

So, at Anthropic, we leave it up to teams. On the Claude Code team, we leave it up to every person. Different people use this differently. For example, I don't use a ticketing system. Some people like to use Asana or notes or something like this. One of the coolest things that I saw, this was maybe like 3 months ago or something, we launched plugins. And the way we launched that is Daisy, for a weekend. She had a very early version of swarms. And she let the swarm run, and she told it, 'Your job is to build plugins. You have to come up with a spec, then you have to make an Asana board and split up into tasks. And then, all the different agents have to build it.' And she set up a container, and she set up a Claude in dangerous mode. And she let it run for the entire weekend. It spawned a couple hundred agents. They made 100 tasks on the Asana board. And then, they implemented it.

Claude Colab:可视化智能体界面 Claude Colab: A Visual Agent Interface

Host

我们来聊聊 Claude Colab。这是很重要的一点,看起来很棒。我试过了。在 Claude 里面有一个 Colab 标签页,然后你可以——我觉得这是一种更直观的运行智能体并与之交互的方式。我听说它是在 10 天内建成的,这很令人惊讶。你能跟我们讲讲构建它需要什么,这到底意味着什么?是从想法开始,还是从决定开始?团队有多大?

Let's talk about Claude Colab. It's one of the very important things about this, it looks great. So I tried it out. It's inside Claude, you have the Colab tab there, and then you can—I feel it's a lot more visual way of running agents and interacting with them. One of the surprising things I heard is that it was built in 10 days. Can you take us through what it took to build it and what does that actually mean? Was it from the idea or from the decision of building it, and how big was the team building it?

Boris

团队非常小,只有几个人。很长一段时间以来,我们都觉得应该为非工程师打造一个产品。之所以这么想,是因为长期以来,使用 Claude Code 的人很多都是非工程师。在产品领域,当你看到潜在需求,看到人们费尽周折去用一个不是为他们设计的产品时,这就是一个很好的信号——是时候为他们专门打造一个产品了。Twitter 上有很多这样的人,有个人用 Claude Code 来监控他的番茄植株。我太喜欢这个了。他装了一个网络摄像头,Claude 会说:‘天哪,我们的植物发芽了,我好开心。’因为它有摄像头,每天监控,看到番茄在长大,它特别高兴。还有个人用 Claude Code 从损坏的硬盘里恢复照片,那是他的婚礼照片。哇。我们 Anthropic 的整个财务团队都在用 Claude Code,销售团队也在用。所以有很多非工程师在使用它。那时,Claude Code 已经有很多形式了,对吧?我们从终端开始,然后扩展到 IDE,所以有基于 VS Code 和 JetBrains 的 IDE 扩展。还有 iOS 和 Android 应用、桌面应用、网页版,以及 Slack 和 GitHub 应用。我们可以把它扩展到所有这些地方,让工程师用起来更方便。但最终,这些都不是为非工程师设计的。Claude Code 进化了很多,但感觉还是有差距,需要一个能让人们更轻松使用的产品。所以过去几个月,团队一直在摸索,看什么产品合适。后来有人想到:如果我们把 Claude Code 加上一些护栏呢?比如,Claude 运行在一个虚拟机上,这是我们确保安全的方式之一,尤其是对不想读 bash 命令来了解它在做什么的非技术用户。他们大概用了 10 天左右的时间来开发,完全是用 Claude Code 构建的。然后我们就发布了。

The team was really small. It was just a few people. For a long time we felt that there is some product to be built for non-engineers. The reason we felt this is that for a long time, people that were using Claude Code were non-engineers. And so, in the product world, when you see latent demand, you see people jumping through hoops to use a product that was not designed for them. That's a really good sign it's time to build another product that is built just for them. There are all these people on Twitter—there's this one guy that was using Claude Code to monitor his tomato plants. I just loved this. It was like he had a webcam set up and the Claude was like, 'Oh my god, I'm so happy that our plant is budding.' And because it had a webcam, it was monitoring it every day and was so happy that the tomatoes were growing. There was someone that was using Claude Code to recover photos off of a corrupted hard drive, and it was his wedding photos. Wow. Our entire finance team at Anthropic uses Claude Code, our sales team uses Claude Code. So there are all these non-engineers that were using it. At that point, Claude Code is available in a lot of form factors, right? We started in a terminal. Then we expanded and added support for IDEs. So we have extensions for every VS Code-based IDE, every JetBrains-based IDE. There are also iOS and Android apps. There's the desktop app. There's web. Then Slack and GitHub apps. So we can expand it to all these places to make Claude Code easier for engineers. But ultimately, none of these are built for non-engineers. Claude Code evolved a lot, but it still felt like there's a gap and there's a product that could make this even easier for people. So for the last couple of months, the team was hacking around and seeing what the right product is. At some point, someone came up with this idea: what if we just take Claude Code, add some guardrails? For example, Claude works with a virtual machine. This is one of the many ways we make sure it's really safe, especially for non-technical users that don't want to read bash commands to figure out what it's doing. They were hacking on this, I think it was something like 10 days or something. It was just fully built with Claude Code. And then we shipped it.

Claude Colab 背后的复杂性 Complexity Behind Claude Colab

Host

你能给我们讲讲像这样的应用背后的复杂性吗?能不能说说哪些部分需要构建?因为从外面看,很难判断:这只是一个漂亮的 UI 封装,还是说只有几百行代码?我这是在故意挑衅。或者实际上,它背后是一个非常复杂的软件。我这么问是因为 Uber 就是一个很好的例子:人们看那个应用,觉得很简单。我在那里工作过,知道它其实非常复杂,因为你看不到很多复杂性。有很多区域性的东西,很多后端的东西都隐藏起来了。所以单看 Claude Code,很难判断有多少是额外需要仔细思考的业务逻辑,又有多少其实只是模型上一个漂亮的小薄层。

Can you give us a sense of the complexity behind an app like this? If we can walk through what parts needed to be built, because from the outside it's a little hard to tell: is this just a nice UI wrapper, or is it a few hundred lines of code? I'm being provocative here. Or behind the scenes, it's actually really complex software. The reason I ask is that Uber is a great example where people look at the app and it looks really simple. I worked there and I know it's really complex because you don't see a lot of the complexity. There are a lot of regional things, a lot of back-end things that are hidden. So from just looking at Claude Code, it's hard to tell how much of this is additional business logic that needed to be carefully thought out versus it's actually just a nice little thin wrapper on top of the model.

Boris

有些地方,我觉得复杂性比你想象的要低;有些地方则更高。在产品方面,它相当简单,因为它就是 Claude 桌面应用。你下载 Claude 应用,它是一个单一的桌面应用,有协作标签页、代码标签页和聊天标签页。所以它只是一个应用,我们继承了很多产品逻辑。有一些 UI 渲染代码。在底层,它运行的是同样的 Claude Code,是同样的 Claude 智能体 SDK 驱动 Claude Code。很多复杂性其实在于安全性。因为我们知道用户是非技术性的,我们只想确保他们有良好的体验。比如,如果有人启动应用然后删掉一堆家庭照片,那就不太好了。所以我们想确保防止这种情况,你不会意外地这么做。这就是很多护栏的来源。后端运行着一堆分类器,用于安全以及针对提示注入等安全风险的额外缓解措施。在前端,我们附带了一个完整的虚拟机。还有很多操作系统级别的集成,确保人们不会意外删除东西。所以光是安全方面,就有很多工作。然后我们还得重新考虑权限系统,因为我们继承了 Claude Code 的权限系统。但此外,对于协作来说,很大一部分价值不仅在于本地运行,还在于像 Claude Code 那样使用你所有的工具。然而,对于非技术用户来说,你的工具并不能以 CLI 的形式使用。有些可以通过 MCP 使用,很多在浏览器里可用。所以协作配上 Chrome 扩展就非常好用。这就是我通常使用的方式。比如,我每周都用它来做团队的项目管理。我们有一个电子表格,高层级地跟踪每个人在做什么。这是我个人的项目管理方式。对于我自己的任务,我什么都不用,但对于整个团队,我有这个电子表格。我让协作来检查。我每周都会问协作:‘嘿,你能看看那些状态没填的行吗?能不能在 Slack 上提醒一下工程师?’然后它就会在 Chrome 里打开一个电子表格标签页,再打开一个 Slack 标签页,然后开始在 Slack 上给工程师发消息。它一次性就完成了。

In some places, I think there's less complexity than you would think; in some places, there's more complexity. On the product side, it's quite simple because it's just the Claude desktop app. You download the Claude app. It's a single desktop app. It has a tab for co-work, a tab for code, a tab for chat. So it's just one app, and we're able to inherit a lot of that product logic. There's some UI rendering code. Under the hood, it's just the same Claude Code running. It's the same Claude agent SDK that powers Claude Code. A lot of the complexity actually is about safety. Because we know the user is non-technical, we just want to make sure they have a good experience. For example, if someone launches the app and then deletes a bunch of family photos, that's really not good. So we wanted to make sure we protect against this, so you can't accidentally do that. That's where a lot of the guardrails came from. There are a bunch of classifiers running on the back end for safety and extra mitigations for things like prompt injection and security risks. On the front end, there's an entire virtual machine that we ship. There are a bunch of operating system-level integrations to make sure people don't accidentally delete things. So just around safety, there's a lot there. Then we also have to rethink the permission system because we inherit the permission system from Claude Code. But also, for co-work, a big part of the value is not just running locally, but using all of your tools the way Claude Code uses them. However, for non-technical users, your tools aren't really available as CLIs. Some are available over MCP, many are available in a browser. So co-work is really good when you pair it with a Chrome extension. This is the way I usually use it. For example, I use it every week to do project management for the team. We have a spreadsheet that tracks at a high level what everyone is working on. This is my personal way of project managing. For my own tasks, I don't use anything, but for the team overall, I have the spreadsheet. I have co-work check in. I just ask co-work every week, 'Hey, can you look at the rows for any status that has not been filled out? Can you just ping the engineer on Slack?' So it'll open one tab in Chrome for the spreadsheet, another tab with Slack, and then it'll start messaging engineers in Slack. It just one-shots it.

技术栈与平台选择 Tech stack and platform choice

Host

这背后的技术栈是什么?我猜很多会和 Claude 应用类似,但它是 Electron、TypeScript 这类东西,还是别的?

What's the tech stack behind this? I assume a lot of it will be similar to the Claude app, but is it Electron, TypeScript, those kind of things or something else?

Boris

对,就是 Electron 和 TypeScript。实际上,参与开发的一些人是早期 Electron 的开发者。比如 Felix,他是 Cowerk 的创建者,他曾经是 Electron 的早期工程师,参与过它的构建。

Yeah, just Electron and TypeScript. Actually, some of the people working on it are early Electron folks. So Felix, who's the creator of Cowerk, he was a really early engineer on Electron and he helped build it.

Host

太棒了。Cowerk 只发布了 macOS 版本。选择这个平台优先发布,以及目前只支持这个平台的原因是什么?

Amazing. And Cowerk launched macOS only. What was the reason for both choosing this platform first and for now only choosing this platform?

Boris

嗯,Windows 版即将推出。我想可能等这期播客发布的时候,我们就已经支持 Windows 了。我们只是想尽早开始,尽早学习。你知道,就像我们在 Anthropic 做的每一件事一样,这有点像我自己讲的故事——我喜欢 Anthropic 的一点是,它非常符合这里人们的思维方式。回到这一点,我们对自己构建的东西并没有很高的确定性。我们的直觉常常是错的,所以我们只能向用户学习,弄清楚人们真正想要什么,花大量时间倾听用户,深入理解反馈。这就是我们构建产品的方式。所以我们总是在产品还没完全准备好时就发布。我们当初对 Claude Code 就是这么做的。Claude Code 刚发布时,甚至不支持 Windows。它也不支持很多不同的技术栈,然后在接下来的几周里,我们逐步增加了对所有技术栈的支持。现在 Claude Code 支持每一个技术栈。你知道,Windows、任何奇怪的 Linux 发行版、macOS,我们都支持。所以对于 Cowerk 也是一样,我们只是想尽早发布。我们想从 Mac 开始,那只是一个起点。但没错,它会支持所有平台。

Yeah, so Windows coming soon. I think probably by the time this podcast comes out, we will have Windows support. We just wanted to start early and start learning. You know, like everything we do at Anthropic, it's kind of like the way that I told my own story, one of the things I like about Anthropic is it just really matches the way that people here think about it. Back to this point where we don't have high certainty about the things that we build. And our intuition is often wrong, and so we just have to learn from users and figure out what people actually want and spend a lot of time listening to people and understanding the feedback deeply. This is the way that we build a product. And so we always launch a little bit before it's ready. We did this for Claude Code. When we launched Claude Code initially, it didn't even support Windows. Also, it didn't support a lot of different stacks, and then over the coming weeks we added support for every stack. Now Claude Code supports every single stack. You know, like Windows, whatever weird Linux distro you use, macOS, we support everything. And so for Cowerk also, we just wanted to launch early. We wanted to start with Mac as that was just the starting point. But yeah, it's going to support everything.

可观测性与隐私 Observability and privacy

Host

你提到了获取反馈。我很好奇,对于 Claude Code 和 Claude Co-work,你们在发布时如何处理可观测性、监控这类事情?你们使用功能标志吗?我更感兴趣的是,你们是为这个构建了自定义工具,还是决定使用某些供应商?因为特别是对于可观测性,我确信这既重要,而且听起来规模也相当大,用户数量可观,或者这不会是一个小规模运营?

One thing you mentioned is getting feedback. I'm curious both for Claude Code and for Claude Co-work, how do you go about things like observability, monitoring, when you are rolling out? Do you use any feature flags? And I'm more interested in like did you build custom tools for this or did you decide to use certain vendors? Because especially for observability, I'm sure that this is both important but it also sounds like pretty high scale in terms of the number of users that we can derive or is this will not be a small operation?

Boris

嗯,我们使用了一些现成的供应商,也使用了一些自定义代码。所以实际上是两者混合。没什么太令人惊讶的。关于 Anthropic 有一点很有意思,因为我们是一家企业公司,非常注重隐私和安全,我们不能查看用户的数据。所以,你知道,如果有人报告一个 bug,我实际上不能调出你的日志来看看发生了什么。大量工作都投入到如何以保护隐私的方式记录事件等事情上。这对我们的运营方式非常重要。

Yeah, there's some off-the-shelf vendors that we use, there's some custom code that we use. So, it's actually a mix of both. There's nothing too surprising about it. There's one thing about Anthropic that's kind of interesting is because we're an enterprise company and we care a lot about privacy and security, we can't see people's data. And so, you know, like if someone reports a bug, I actually can't pull up your logs to kind of see what's going on. A lot of work goes into figuring out how to log events and things like this in a privacy-preserving way. This is just very important to the way that we operate.

从 Cowerk 早期学到的经验 Early learnings from Cowerk

Host

对于 Co-work,到目前为止你们有什么收获?它已经发布了几周了。你们有没有看到什么意想不到的事情?你们是否根据收到的反馈来调整产品?

For Co-work, what kind of learnings have you had so far? It's been out for I think a few weeks now. Did you see something unexpected? Are you shaping the product based on feedback that you're getting?

Boris

嗯,团队每天都在发布大量修复。最令人惊讶的是,老实说,人们非常喜欢它。Claude Code 刚推出时,其实并不是一夜爆红。人们以为是那样,但一开始它其实是缓慢起步的,我认为第一个大的转折点是五月份我们发布了 Opus 4 和 Sonnet 4,那时它才真正火起来,我们的增长也变成了指数级。但在最初,它更像是一个研究预览,人们不太知道怎么用。有些人立刻理解了,但大多数人没有。这花了一点时间。对于 Co-work,它的增长轨迹比 Claude Code 初期要陡峭得多。所以它是一炮而红,这真的很令人惊讶。我没想到会这样。

Yeah, every day the team is landing so many fixes. The most surprising thing is just how much people are loving it to be honest. When Claude Code first came out, it actually wasn't an overnight hit. This is something people think it was but it was sort of a slow take-off at the beginning and I think the first big inflection was in May when we released Opus 4 and Sonnet 4, that's when it really clicked and that's when our growth became exponential. But at the beginning, it was sort of a research preview, people didn't really know how to use it. Some people got it immediately but most people didn't. It took a little while. For Co-work, it's a much steeper growth trajectory than Claude Code was at the beginning. So, it's just been an instant hit and that's actually been very surprising. I didn't really expect that.

智能体团队与非相关上下文窗口 Agent teams and uncorrelated context windows

Host

你们最近发布的一个新功能,我想就在我们录制这期播客的前一天或前两天,是智能体团队。据我理解,智能体团队或智能体集群的想法是,你可以有一个主导智能体,它可以委派任务给不同的队友,而不是只有一个智能体。你们是如何开始试验这个的,又是如何决定现在发布的?

One of your new releases which came out just very recently, it was I think yesterday or the day before when we're recording this podcast was agent teams. And as I understand the idea with agent teams, agent swarms, instead of a single agent, you can have a lead agent and it can delegate to its different teammates. How do you start experimenting with this and how did you decide to ship it now?

Boris

我们一直在做实验,对吧?有很多方法可以从 Claude Code 中获得更多价值。一种方法是扩展上下文。另一种方法是自动压缩上下文,这样它基本上就是无限上下文,我们现在就是这样做的。另一种方法是使用子智能体,这样就有多个智能体协同工作。有很多不同的方法可以从上下文窗口中获得更多价值。有一个想法叫做非相关上下文窗口。这是我们起的名字,想法是你有多个上下文窗口,但它们基本上是全新开始的。所以它们彼此不知道对方。举个例子,相关上下文窗口是,你让模型做一个任务,然后在同一个上下文窗口中让它做第二个任务。在这种情况下,第二个任务知道第一个任务,因为它们在同一个窗口中。但对于子智能体这样的东西,它是非相关的,因为主智能体提示子智能体,但子智能体的上下文窗口是全新的。除了那个提示,它不知道父上下文窗口里有什么。你实际上可以在子智能体与技能的比较中看到这一点。因为当你运行一个技能或斜杠命令时,它能看到父上下文窗口,而子智能体则不能。所以它是非相关的。有些情况下你需要那个上下文,有些情况下不需要。有趣的是,非相关上下文窗口,以及向问题投入更多上下文和更多 token,当窗口是非相关时,会带来更好的结果。这实际上是测试时计算的一种形式。对于智能体团队这样的东西,我们已经实验了一段时间,我想大概从去年十月或九月左右开始的。

We're always doing experiments, right? There's all sorts of ways to get more mileage out of Claude Code. One way you can do it is by extending context. Another way is auto compacting context, so it's essentially infinite context and that's what we have right now. Another way is using sub agents, so you have multiple agents kind of working together. There's just a lot of different approaches to get a little bit more mileage out of the context window. There's this one idea called uncorrelated context windows. That's what we call it and the idea is you have multiple context windows, but they essentially start fresh. So, they don't know about each other. And so, an example of this is like a correlated context window is if you have the model and it does a task and then you have it just do a second task in that same context window. And in this case the second task knows about the first one because it's in the same window. But for something like a sub agent, it's uncorrelated because the main agent prompts the sub agent, but the sub agent's context window is fresh. Besides that prompt, it doesn't know what's in the parent context window. And you can see this actually a little bit in for example sub agents versus skills. Because when you run a skill, or slash command, it sees the parent context window versus for a sub agent, it doesn't. So, it's uncorrelated. There's some cases where you want that context. There's some cases when you don't. And there's this kind of interesting thing where uncorrelated context windows and just throwing more context at the problem and throwing more tokens at it, when the windows are uncorrelated, gives you better results. It's actually a form of test time compute to do this. And for something like teams, we've been experimenting with this for a while, I think since maybe like October or September or something like this.

Opus 4.6 与群体智能 Opus 4.6 and Swarm Intelligence

Boris

感觉在 Opus 4.6 上,它真的就通了。模型学会了怎么用这个功能。有时候你会看到一些很可爱的互动,智能体之间互相聊天讨论事情,看起来非常酷,某种程度上很人性化。但其他时候,你又能得到非常好的结果。比如我们做了一堆内部评估,让多个智能体构建非常复杂的东西——比单个智能体能构建的更复杂。我们看到 Opus 4.6 配合团队协作后,结果真的提升了很多。所以我们就觉得是时候发布了。但我们也很谨慎。之所以需要手动选择加入、作为研究预览发布,是因为它消耗大量 token——毕竟是一堆智能体在运行。不是所有人都一直需要这个。所以很期待看到大家怎么用,听听反馈。它适合相当复杂的任务,可能不适合每个任务。主智能体决定子智能体的规则,我们没有固定的方式,这取决于具体上下文。我不认为有唯一正确的方法。实际上,我觉得这个功能的神奇之处在于“不相关上下文窗口”这个概念,而不是智能体的具体配置。但人们应该去实验,没有一刀切的方案。

And it really just felt like with Opus 4.6, it clicked. Where the model figured out really how to use this. And sometimes you see these kind of cute exchanges where the agents are talking to each other and they're like discussing something and it's just very cool to see. It's very like humanistic in a way. But there's other times where you just get very good results. And so we had a bunch of internal evaluations for example where we have quad build something very, very complex. Something more complex than what a single quad would build. And we saw the results just really, really improve with Opus 4.6 with teams. And that's why we felt it's the right time to release it. We also wanted to be careful. Um and the reason you have to opt into it, the reason it's a research preview is it uses a ton of tokens. Cuz it's just a bunch of quads that are running. Um not everyone wants this all the time. So it's just excited to see how people use it and uh you know, to to hear the feedback. It's It's something you want for fairly complex tasks. You don't probably want this for every task. The main quad decides the rules for the sub quads. We don't have a kind of a regimented way to do this. It's It's context specific. I wouldn't say there's one right way to do it. I think actually a lot of the magic of this comes out of this idea of uncorrelated context windows. It's less about the specific configuration of the agents. But it you know, it's something that people should experiment with. I don't think there's a one-size-fits-all.

Host

你有没有看到一些用例——我知道它还在研究阶段——但有没有看到这种群体智能方法看起来很有前景的用例?

Have you seen use cases even in even I I know it's it's still research but have you seen use cases where it could look it it looks promising this approach, this swarm approach?

Boris

嗯,我之前说过,插件完全是用群体智能构建的。从那以后还有很多其他功能也是用这种方式构建的。所以是的,我认为任何单个智能体难以应对的事情,群体智能都能帮忙。这很有意思。

Well, you know, I guess I said before plugins were fully built with swarms. There There's a bunch of other features since that are built in this way. So yeah, I I I think for anything where you see a single quad struggling, the swarms can help. It's It's an interesting to to look at.

拥抱 AI 的快速变化 Embracing Rapid Change in AI

Host

说到变化。去年十二月你和 Andrej Karpathy 有一次很有意思的交流,他发帖说作为程序员,他从未像现在这样感到落后,因为 AI 进步太快。然后你分享了一个故事:你开始用老方法调试内存泄漏,结果 Claude 一次就搞定了。我觉得这反映了大家的感觉——变化太快了,假期里我开始觉得事情真的变了。你是如何接受甚至拥抱这种变化的?

Talking about change in in general. With Andrej Karpathy you had a really interesting exchange back in December where when he posted that he's never felt as much behind as as a programmer as he is now because of the progress with AI. And then you shared the story about how you started to debug a memory leak the old-fashioned way and then Claude just one shot at it. I think it was a reflection of like how everyone is feeling that things are changing so fast and in the in the holiday break I started to feel that things have have really shifted. How did you I guess come to terms with this or or start to embrace this change?

Boris

这确实是我很纠结的地方。模型进步太快了,以前模型上有效的想法在新模型上可能无效,以前无效的现在可能有效。这很奇怪,因为很少有其他技术是这样。所以我没什么经验可借鉴,不知道该怎么应对。这成了我必须学习的新技能。某种程度上,你必须始终保持初学者心态。老实说,我经常用“谦逊”这个词,但你真的需要这种智力上的谦逊。因为以前不好的想法现在变好了,反之亦然。我觉得就是这样,我经常提醒自己。有趣的是,以前如果有人重试一个过去失败过的想法,通常的反应是“你怎么又做这个?”这有点像守门,但某种程度上是合理的——比如架构上有人说“我们为什么不做微服务?”,另一个人说“试过,不行”,如果那是两三年前,确实有道理,因为没什么变化。没错。微服务这东西每十年流行一次。但现在,我认为这是历史上第一次,每隔几个月重试同一个想法并不疯狂,因为模型进步了,它就奏效了。我在团队里也看到这一点,新来的工程师有时做得比我还好。我不得不看着他们,学习并调整期望。举个例子,发布功能时,我有时会在 X 或 Threads 上截图展示用法。但最近我们的开发者关系负责人 Tariq,他写了很多代码,他很棒。他开始自动化这个流程,让热代码为发布生成自己的视频。我觉得这也许可行,但自己不会去试,因为觉得模型还没准备好,但他直接做了,而且成功了。

This is something I really struggle with. The model is improving so quickly that the ideas that worked with the old model might not work with a new model. The things that didn't work with a new model might work or with the old model might work with a new model. And it it's weird because there's just not a lot a lot of other technologies like this. So I I just don't really have a lot of experience to draw on to figure out how I should approach this. And it it's been this new skill that I've had to learn. In a way it's like you just always have to bring this beginner mindset. Honestly like I'm I'm using the word humility a lot but you always just have to bring this kind of intellectual humility. Because just all of these ideas that were bad before are now good and and and the inverse. I I think that's honestly it. It's something I constantly have to remind myself about. And back in the it's funny back in the old world when someone tries an idea again and we've tried it in the past and it didn't work usually the feedback is like why are you doing this again? Yeah yeah the you should run. This used I mean we used to call it a bit of a gatekeeping but it was somewhat valid where I know with architecture someone came and said like why don't we do microservices and someone said we tried it and it didn't work and if you tried it a year or two or three years ago it was kind of valid right cuz not much has changed. Yeah that's right that's right. And it's something like microservices is funny cuz it's like every 10 years it goes in and out of in and out of style. But yeah now it's now it's I think the first time ever where it's actually not crazy to just try the same idea every few months because the model improves and it just works. And I I actually see this with engineers on the team like new people that are newer to the team, people that are newer to engineering sometimes do things in a better way than than I do. Um and I just have to like look at them and I have to learn and I have to adjust my expectations. You know, like an an example of this is, you know, when when we release features sometimes I'll like screenshot myself using them on, you know, on X or on threads or whatever just to kind of talk about it. Um but recently Tariq, our um you know, our devrel guy, he actually coded a lot. Um he's amazing. And he just started automating this. So, he's having like hot code generate its own videos for for its launches and he just started doing this. And you know, this is something like I thought would be, you know, maybe it's possible. It's not something I would have tried cuz I wouldn't have thought the model was ready, but he just he just did it and it just kind of worked.

放下编码身份 Letting Go of Coding Identity

Host

有一件事让我觉得有点奇怪,我想很多开发者都能理解:从 Opus 4.5 开始,我就接受了这个事实,类似模型比如 GPT 5.2 也给了我同样的感觉。模型写代码非常厉害,我意识到当我想完成事情时,我不会再手写代码了。如果我真的想享受写代码的乐趣,我仍然可以。但我反思的是,为了擅长编程付出了太多努力。我记得学习过程,从大学前的瞎折腾,到学 C 和 C++,真的很难。后来在最初几份工作中,我逐渐变得更好,更擅长调试。有那么一个阶段,我的很多身份认同都建立在编程能力上。那是我们过去找工作或高薪工作的方式。当我在 Uber 做入门经理时,我们设计面试流程,和经理讨论筛选标准:开发者大部分时间在做什么?大约 50% 的时间在写代码。因此,我们把大约 50% 的信号放在编程上。所以很多事都和编程绑定,因为它确实很难。我们都知道这需要毅力,需要一定智力才能做好。

One thing that I've I felt like just a bit like odd about and I think a lot of developers can relate is I've come to terms with this starting from Opus 4.5 the and and also similar models like I think GPT 5.2 gave me similar vibe as well. The models have been just really good at writing code and I I realize that I don't think I will handwrite the code when I'm get I when I want to get stuff done. If if I actually want to, you know, get the pleasure of writing it, I can still do it. But one thing I reflected on is it's just been so much effort to get good at coding. I I remember when I when I was learning when I I started from like kind of hacking around to go into university to learning C and then C++ and it was just bloody hard. And actually, you know, going through my my first few jobs where I started to become better at it, I become better at debugging. And there's a point where like a lot of my identity was tied to being good at coding. That's how we used to get jobs or higher paying jobs. When I was an entry manager, when we designed the interview loop at Uber, we we had talk with managers of what we need to screen for and we we talk like, well, what do developers do most of their time? About 50% of the time they code. Therefore, we place about 50% of the signal was all about coding. So, there was a lot of things tied into coding because it it is just hard. I think we all know that it takes grit. It takes some level of intelligence to get good at it.

编码作为手艺的失落与哀伤 Grief and Loss of Coding as a Craft

Host

而且有一种失落感。我觉得模型能做到这一点很棒,但感觉有些东西很快就被夺走了,我个人没想到会这么快。我想很多其他人也有同感。有些人更容易接受,但肯定有一种悲伤的感觉。你是怎么想的?你在 Facebook 内外写了那么多代码。那只是一个工具,但没多少人能做到你做的。现在模型也能做得和你一样好,甚至更好。这就是挑战。

And there's a sense of loss. I think it's great that the model can do it, but it feels like something got taken away quickly that I personally didn't think would happen this fast. And I think a lot of other people feel that too. Some people move on easier, but there's definitely a sense of grief. How did you think about it? You wrote so much code at Facebook and outside. It was a tool, but not many could do what you did. Now models can work as well as you, if not better. That's the challenge.

Boris

是的,我认为这曾经是我们软件工程师做的事情,现在正变成每个人都能做的事情。我刚开始编程时,它非常实用,是完成事情的一种方式。后来我迷上了编程的艺术、语言和工具本身。我深陷其中。我写了一本关于编程语言 TypeScript 的书,是 O'Reilly 出版的第一本 TypeScript 书。在日本一个小镇上有一个美妙的时刻:我去书店,发现那本书被翻译成了日语。在那个小镇上,那是最酷的时刻。然后我意识到我完全不记得 TypeScript 了,因为我只写 Python 写了几年。后来我创办了世界上最大的 TypeScript 聚会,在旧金山。我见到了很多我的英雄:写了《反应性通论》的 Chris Kowal,创造了 Node 的 Ryan Dahl。那是我第一次深入这个社区、语言和工具本身。对于 TypeScript,类型系统中有一种美。Heilsberg 很聪明——条件类型,任何东西都可以是字面类型。这些是即使最硬核的函数式语言也没有的深刻思想。连 Haskell 都没走那么远。Anders 把它推得更远。Joe Pamer 和其他人输出了这些想法。对他们来说,这是实用的:他们有大型无类型 JavaScript 代码库,需要逐步迁移到类型化代码,这需要优美的想法。对我来说,Scala 是另一个兔子洞——函数式编程。当我写代码,当模型写代码时,我总是先想类型。类型签名比代码本身更重要。这其中有美和艺术。但归根结底,它是实用的——是达到目的的手段,而不是目的本身。

Yeah, I think it's something that used to be a thing we do as software engineers. It's becoming something everyone can do. There was a moment when I started coding—it was very practical, a way to get things done. At some point I fell in love with the art of coding, the languages, the tools themselves. I fell down this rabbit hole. I wrote a book about a programming language—TypeScript. I wrote the first ever TypeScript book with O'Reilly. There was an amazing moment in a small town in Japan: I went to a bookstore and found that book translated into Japanese. In that tiny town, it was the coolest moment. Then I realized I didn't remember TypeScript at all because I had only been writing Python for a couple of years. At some point I started the biggest TypeScript meetup in the world, in San Francisco. I met a lot of my heroes: Chris Kowal, who wrote the General Theory of Reactivity; Ryan Dahl, who made Node. One of the first times I went deep into that community, the language itself, the tools. For TypeScript, there's beauty in the type system. Heilsberg is brilliant—conditional types, anything can be a literal type. These are deep ideas that even the most hardcore functional languages don't have. Even Haskell doesn't go that far. Anders pushed it much further. Joe Pamer and others exported these ideas. For them, it was practical: they had large untyped JavaScript codebases, and they needed to gradually migrate to typed code, which required beautiful ideas. For me, Scala was another rabbit hole—functional programming. When I write code, and when the model writes code, I always think in types first. The type signature matters more than the code itself. There is beauty and art to it. But in the end, it's practical—a means to an end, not an end in itself.

印刷术类比 The Printing Press Analogy

Boris

我认为这个时刻的一个比喻是 15 世纪的印刷机。那时有一群抄写员知道如何书写。学习过程很艰难——你需要设备、赞助或选拔。反复练习生产同样的东西,很少有人能做到。那是高地位或高薪的。然后印刷机出现了。在欧洲,领主或国王必须雇佣你,你要经过多年的训练。有一个抄写员阶层,受雇于国王或王后这样的人,而他们自己往往不识字。那是一个非常小众的技能——当时欧洲不到 1%的人口识字。然后印刷机出现了。在接下来的 30-50 年里,印刷材料的成本下降了大约 100 倍。在接下来的 50-100 年里,印刷材料的数量增加了 10,000 倍。这是第一个效果。识字率花了一段时间才跟上——全球识字率上升到大约 70%,但这又花了 200-300 年,因为学习读写很难。它需要教育系统、纸张和墨水的基础设施,以及不用在农场工作的空闲时间。早期工业化才实现了这一点。但让象牙塔里的东西变得人人可及的效果——没有它,我们周围的东西都不会存在。如果我们不识字,如果制造这个麦克风的人不识字,就很难有现代经济。这些东西都不会存在。那时,如果人们要预测印刷机带来的变化,没人会预测到麦克风。所以我认为这是我们当前时刻最好的类比。

I think one metaphor for this moment is the printing press in the 1400s. At that time, there was a group of scribes who knew how to write. It was a hard process to learn—you needed equipment, sponsorship, or selection. Practicing to produce the same thing over and over, few could do it. It was high prestige or highly paid. Then the printing press came along. In Europe, a lord or king had to employ you, and you went through years of training. There was a class of scribes employed by someone like the king or queen, who themselves were often illiterate. It was a very niche skill—less than 1% of the population was literate in Europe back then. Then the printing press came out. The cost of printed material went down about 100x over the next 30-50 years. The quantity of printed materials went up 10,000x in the next 50-100 years. That was the first effect. Literacy took a while to catch up—global literacy went up to about 70%, but that took another 200-300 years, because learning to read and write is hard. It takes education systems, infrastructure for paper and ink, and free time instead of working on a farm. It took early industrialization to get there. But the effect of making something locked away in an ivory tower accessible to everyone—none of the things around us would exist today without that. If we weren't literate, if the people who built this microphone weren't literate, it would be hard to have a modern economy. None of these things would exist. Back then, if people had to predict what would happen with the printing press, no one would have predicted the microphone. So I think this is the best analog for the moment we're in now.

Host

有趣的是,你说一些国王不识字却雇佣抄写员,因为说实话,我们有企业主知道他们想建什么,却雇佣软件工程师,因为他们自己不会写代码。我们喜欢嘲笑那些在白板上画原型并说“这应该很容易”的 CEO,当然他们不明白这有多难。

It's interesting that you say some kings were illiterate who employed scribes, because if we're honest, we have business owners who know what they want to build and employ software engineers because they can't write code themselves. We like to mock CEOs who come with drawn prototypes on whiteboards and say 'this should be easy,' but of course they don't understand how difficult it is.

印刷术类比与 AI 影响 Printing Press Analogy for AI Impact

Host

但这里似乎有个类比:一个人有想法,但之前需要雇一个软件专家来实现,想法和本人之间总有脱节。就像印刷术一样,如果他们能自己表达呢?比如国王能自己读写信件,就不需要中间人了,事情会更高效。当然,对抄写员来说这不一定是好消息,但聪明的抄写员也能适应,总得有人写书、操作印刷机等等。

But there seems to be a bit of analogy where there's a person who wants what they want, but until now they needed to hire a software specialist who can build that and there's always that disconnect between the idea and the person. And just like with the printing press, what would happen if they could actually express themselves? Like the king could actually read or write their own letters. They wouldn't need that middleman and things become more efficient. I mean, of course for the scribe it's not the best news necessarily, but the smart scribes can also adapt, so someone needs to write the books, run the press, etc.

Boris

没错。想想抄写员的命运:他们不再是抄写员,但出现了作家和作者这个类别。这些人之所以存在,是因为文学市场大大扩张了。而且想想以前,抄写员的作品只有少数人读,有了印刷术,作者多了很多,有些没人读,但有些的影响力远超想象。由此诞生了新的职业。

Yeah, exactly. And if you think about what happened to the scribes, right? They ceased to become scribes, but now there's a category of writers and authors. These people now exist. And the reason they exist is because the market for literature just expanded a ton. And I guess also if we think about back then, a scribe's work was read by a few people, and with the printing press, an author—there are a lot more authors and some of them are not really read, but some of them have wider reach than they could imagine. There are new careers that exist because of that.

Host

我喜欢这个类比。最让我兴奋的是,今天根本无法预测这次转变之后会发生什么。我们所知的经济如果没有它就不会存在。那么下一步是什么?我们今天甚至无法预测会出现什么?因为任何人都能做到。

Yeah, I love the analogy. And the most exciting thing for me is it's just so impossible to say today what will happen after this happens and after this transition happens. Just, you know, the economy as we know it would not have existed without it. So what's next? What is the thing that we can't even predict today that will exist? Because anyone can do this.

Boris

我们无法预测,但可以看看现在什么有效。看看你周围,比如团队里的软件工程师、构建者或技术人员,哪些人很突出?他们在做什么?培养了哪些技能?工作方式有什么变化?

Well, we cannot predict, but I think we can look at what is working right now. If you look around in your environment, may that be the team across on traffic, who are software engineers or builders or members of technical staff, however we call them, who to you are standout? What are they doing? What skills have they built up and how have they changed the way they work?

Host

很难具体点名,因为这些人是我职业生涯中合作过的最强的人。有各种不同的类型。有些人是出色的原型设计师,能把东西从 0 做到 0.5,找出酷想法和技术解锁点。另一些人擅长找到产品市场契合点,从 0.5 到 1 或 0 到 1。还有些人跨学科,我越来越多地看到这样的人:跨产品工程和基础设施工程,或者产品和设计,设计和工程。我看到越来越多这种混合型人才。

It's hard to name individuals because honestly, these are the strongest people I've ever worked with in my career. There are all sorts of different archetypes. There are some people that are really amazing prototypers. So, take something from 0 to 0.5. Just, you know, figure out what are some cool ideas, what is the technology unlock. There are other people that are amazing at finding product-market fit. So, kind of 0.5 to 1 or maybe 0 to 1. There are other people that span different disciplines and I'm just seeing more and more of these people. Like I said, people that span product engineering and infrastructure engineering or, you know, product and design or design and engineering. I think I'm just seeing a lot more of these hybrids.

对 AI 安全信念的改变 Changed Beliefs on AI Safety

Host

从去年到今年,有什么信念改变了?你曾经相信或坚信的东西,后来修正或完全抛弃了?

What's a belief that changed from last year to this year? Something that you either believed or a conviction that you had that you've either revised or completely threw away.

Boris

说实话,我以前不确定安全问题有多大。我加入 Anthropic 是因为读了很多科幻小说,知道事情可能变得多糟。但我不确定。从内部观察,看到去年出现的新风险,让我更加担忧。所以这对我来说很重要。现在,最重要的事情就是如何确保一切顺利。

I think one thing I wasn't sure about is how big a problem is safety, to be totally honest. I joined Anthropic because, like I said, I read a lot of sci-fi and I kind of know how bad this thing can go if it goes bad. It wasn't something I was sure about. But seeing it from the inside and then seeing how the new risks that have arisen in the last year, it just makes me much more worried about it. So, I think it was kind of an important thing for me. Now, it's just the most important thing is how do we make sure this thing goes well?

重要技能与被抛弃的技能 Skills That Matter and Those Left Behind

Host

可以说,在 AI 兴起之前你就是一位出色的软件工程师,看起来效率很高,既是团队一员,个人能力也很强。有哪些技能在成为软件工程师之前就很有价值,现在依然如此甚至更重要?哪些技能可能没那么重要了,最好抛弃?

I think it's safe to say you were a really great software engineer even before all the AI things started and you seem to be a very productive engineer, of course part of a team as well, but also individually. What are some skills that, before being a software engineer, are still as valuable or maybe even more valuable than before and what are ones that are maybe just not as much and then they're best left behind?

Boris

可能最好抛弃的是对代码风格和语言等的强烈偏好。我迫不及待想摆脱这些没完没了的语言和框架争论。因为模型可以用任何语言和框架,你不喜欢它就能重写。所以这不再重要。今天仍然重要的是有条理和假设驱动。这在产品设计中很重要,因为一切都在被颠覆,我们需要决定下一步做什么,每个人都在思考。在日常工程中也很重要,比如调试,你必须非常有条理。模型可以帮忙,但我们现在还处于过渡期,仍然需要这个技能。我不知道 6 个月后是否还需要。其他更有价值的技能是保持好奇,愿意做超出自己领域的事情。比如你做工程,但真正理解商业,就能构建出很棒的产品。我认为下一个十亿美元的产品,比如 Claude Code 之后,下一个成为万亿美元初创公司的,可能就是一个有酷想法的人,能跨工程、产品和商业思考,或者设计和金融等。人们会越来越跨学科,并得到更多回报。所以某种程度上,今年将是通才的一年。另一个实际上被奖励的技能是注意力短暂。我看到它被奖励了。就像青少年用 TikTok 之类的东西,这在某种程度上对社会有危险,因为你需要能深度思考的人,而不是快速跳到下一个想法。但某种程度上,今年是奖励多动症的一年。因为我的工作变成了在云之间跳转,管理云。所以不是深度工作,而是上下文切换和快速跨多个上下文的能力。

Probably, the stuff that's best left behind is maybe like very strong opinions about code style and languages and things like this. I can't wait to get past these endless language debates and framework debates and all this stuff. Because the model can just, you know, use whatever language and framework and if you don't like it, it can just rewrite it for you. So it just doesn't matter anymore. I think something that still matters a lot today is being methodical and hypothesis-driven. This matters both in product design in this world where everything is being disrupted and we need to figure out what to build next and this is something everyone is thinking about. But it also matters for engineering day-to-day, you know, like something like debugging. You just have to be very methodical about it. And the model can do this and it can help a lot. But I think still we're in this transition point where you still need to have the skill. I don't know if you'll still need to have it in 6 months. Other skills that I think are more valuable are being curious and being open to doing things beyond your swim lane. So, if you're working on engineering, but you really understand the business side, you can just build really awesome products and I think the next billion-dollar product, after Claude Code, whatever the next startup is that becomes the next trillion-dollar startup, it might just be like one person that has some cool idea and their brain just is able to think across engineering and product and business or, like design and finance and something else. People are going to become more and more multi-disciplined and this will become more and more rewarded. So, in some ways I think this will be the year of the generalist. I think the other skill that's actually been rewarded is having a short attention span. I've seen rewarded now. Oh, yeah. It's like people, like teenagers are using TikTok and all this stuff and I think in some ways it's kind of dangerous for society because you want people that can think deeply and can contemplate ideas and aren't just moving on to the next idea very quick. But in some ways I think this year is kind of the year that is going to reward it's like the year of ADHD. Because the work for me has become jumping between clouds. It's become managing clouds. And so it's not so much about deep work. It's about how good am I about context switching and jumping across multiple different contexts very quickly.

结束语与书籍推荐 Closing thoughts and book recommendations

Host

我能补充一点吗?从你所说的来看,也许可以加上‘适应性’这一点。你提到 ADHD 让你能跳跃思维,但之前你也非常擅长深度专注。让我印象深刻的是,你似乎非常愿意调整工作方式,看看什么在当前阶段最有效,尤其是在变化中。我认为唯一确定的是,无论下一个模型是什么,它都会再次改变,你需要保持好奇并开放地调整工作方式,对吧?

Could I add that from what I understood from all you said, maybe you could add one thing which is adaptability because you're saying of course that ADHD and you can jump across but of course earlier you were very good at focusing deeply on one thing as well. And what strikes me about you and maybe this is true for other people as well, you're just kind of very open to adapting your working style and seeing what works well for this stage especially when things are changing. I think the one certain thing we can be sure is whatever the next model comes out it will change again and you need to be curious and open to adapting how you work, right?

Boris

是的。最后,你有什么书推荐吗?我最近迷上了刘慈欣。他是《三体》的作者,但他还有很多其他好书。我非常喜欢他的短篇小说,他出了几本短篇小说集。我是他的忠实粉丝。对于刚接触科幻、想要更硬核科幻的人,我强烈推荐斯特罗斯的《加速》。这本书我完全推荐。它基本上就是未来 50 年的产品路线图。书中开始出现技术起飞和 AI 奇点,最后以围绕木星的群体龙虾意识体结束。太棒了,我认为它真正捕捉到了那种加速、加速、再加速的节奏,非常符合当下的感觉。在技术方面,我强烈推荐《Scala 函数式编程》。即使语言选择不再那么重要,我认为函数式编程的艺术能教你如何更好地编码。它会教你如何用类型思考。如果你读这本书,我认为做练习非常重要,我大概做了三遍所有练习,效果惊人。它真的把函数式类型的概念刻进你的脑子里,让你无法停止思考。

Yeah. And as closing, what's a book or books that you would recommend? I've gone down a Cixin Liu rabbit hole. So, he's the Three-Body Problem guy but he actually has like a lot of other really good books. I really love his short stories. He has a couple books of short stories. I'm a big fan. For people that are new to sci-fi and you want like a little bit like harder sci-fi I really love Accelerando by Stross. This is a book I would totally recommend. It's like essentially the product roadmap for the next 50 years. It with takeoff kind of starting to happen and kind of AI singularity. And then it ends up with this kind of like group lobster consciousnesses orbiting Jupiter. And it's just like amazing and the thing that I think it really captures is just the pace, this like quickening quickening quickening pace of how this feels. It really matches the feeling right now. And then on the technical side, I would strongly recommend functional programming in Scala. Even if language choice just doesn't matter as much anymore, I think there is this arts to functional programming that just teaches you how to code better. Um, and it'll just teach you how to think in types. If you read this book, I think what's really important is to do the exercises also and I've gone through and I've done all of them probably like three times over and it's just amazing. It it it really just like knocks this idea of functional types into your head and it's just a thing you can't stop thinking about.

Host

Boris,非常感谢。这次对话太棒了。

Boris, thanks so much. This was awesome.

Boris

是的,谢谢 Greg。这次对话真的很有趣,我反复想到的是 Boris 的印刷术类比。中世纪抄写员是少数能写字的精英,受雇于常常不识字的国王,而我们软件工程师今天可能处于类似位置。我们就是抄写员,花了多年掌握这门手艺,现在印刷术来了。但 Boris 告诉我,抄写员并没有消失,他们变成了作家和作者,整个书面作品市场扩大到无人能预测的程度。我觉得这很有希望,也感谢 Boris 没有粉饰太平。另一件让我印象深刻的是 Claude Code 团队构建软件的方式有多么不同:没有 PRD,没有强制工单系统,设计师、数据科学家和财务人员都写代码,在发布功能前构建几十甚至上百个原型。Boris 每天提交 20 到 30 个拉取请求,没有手动编辑一行代码。还有不同的验证系统:Claude Code 审查自己的代码、自动 lint 规则、最佳通过检查以及人工代码审查。如果你喜欢这个播客,请在喜欢的播客平台和 YouTube 上订阅。特别感谢你给节目评分。谢谢,下次见。

Yeah, thanks Greg. This was a really interesting conversation and the thing that I keep coming back to is to Boris's printing press analogy. The idea that medieval scribes were this tiny elite who could write employed by kings who themselves were often illiterate and that we software engineers might be in a similar position today. We are the scribes. We spent years mastering this craft and now the printing press is arriving. But what Boris told me is that the scribes did not disappear. They became writers and authors and the entire market for written work expanded beyond anything anyone could have predicted. I do find this hopeful and also appreciate that Boris didn't sugarcoat it. The other thing that stuck with me is just how differently the Claude Code team built software. No PRDs, no mandatory ticketing system, designers and data scientists and finance people all writing code and building dozens or hundreds of prototypes before shipping a feature. And Boris is shipping 20 to 30 pull requests a day without editing a single line by hand. And there are different verification systems in place. Claude Code reviewing its code, automated lint rules, best of end passes, and human code review. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks, and see you on the next one.

互动版:逐字朗读 + 针对本期提问 →