Claude Code 与 Fable:编程智能体如何重塑软件工程

Claude Code and Fable: How Coding Agents Are Rewriting Software Engineering

凯特·吴 Cat Wu · AI Engineer · 2026-07-15 · 约 52 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 的 Thariq Shihipar 与 Cat Wu 探讨 Claude Code 与 Fable 如何改变日常工程实践,从逐条权限确认到一句话生成功能。

Anthropic's Thariq Shihipar and Cat Wu discuss how Claude Code and Fable have transformed day-to-day engineering, from permission prompts to one-shot features.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 25)

全文 · Full transcript(中英对照)

Welcome and Introductions

Host

欢迎来到这场炉边谈话。今天和我一起的是来自 Anthropic 的 Thariq Shihipar 和 Cat Wu。我们会深入聊聊 Claude Code,可能也会稍微谈谈最近新闻里一直在说的 Fable 这件事。说到 Fable,就在一分半钟前,Fable 恢复了。Fable 现在对我可用了。所以如果你们想跑出房间去用掉你们的 Fable 额度,我不会怪你们。不过我们会有一场很棒的对话,所以请留下来。总之,请和我一起欢迎 Cat。

Welcome to this fireside chat. I have with me Thariq Shihipar and Cat Wu from Anthropic. We are going to be diving deep into Claude Code and we'll probably talk a little about this Fable thing that's been going on, that's been out there in the news. Actually, on the subject of Fable, literally a minute and a half ago, Fable came back. Fable is now available to me. So if you all want to run out of the room and start using up your Fable credits, I wouldn't hold that against you. But we're going to have a great conversation, so please stick around. But yeah, so please welcome Cat for me.

Cat

谢谢邀请我们。

Thanks for having us.

Thariq

是啊,我们肯定是掐着这场对话的时间来的。

Yeah, we timed it for the chat for sure.

Host

对,对。这就是为什么一切都赶上了。

Yep. Yep. This is why it's all happening.

How Coding Agents Changed Daily Work

Host

今年有点离谱。太惊人了。Claude Code 是去年二月发布的。它还不到一岁半,当时只是 Claude Sonnet 3.7 发布时的一个要点。我很想听听你们的看法。既然现在我们有了这些真正能替我们干活的编程智能体,你们过去一年里的日常工作发生了什么变化?

This year has been somewhat absurd. It's amazing. Claude Code came out in February of last year. It's under a year and a half old and it was a bullet point on the Claude Sonnet 3.7 launch. I'd love to hear from you. How has what you do on a day-to-day basis changed in the past year now that we have these coding agents that actually work for us?

Cat

我记得我们刚推出 Claude Code 和 Sonnet 3.7 的时候,你给它一个任务,就得密切盯着它想做的每一件小事。我记得我会极其仔细地读每一个权限提示。我经常说不。我总是说不行不行不行,你检查这个文件了吗?你检查那个文件了吗?而现在,随着每一代模型,情况变得不可思议。我觉得我们都有机会退后一步,把更多琐碎的实现工作委托给 Claude。这释放了我们大量时间,去思考更有创造性的工作,比如既然我们知道 Claude Code 能实现其中很多部分,那我们应该为用户提供什么样的正确体验。而现在有了 Fable,这又是完全不同的一次阶跃式提升。我们很多用例里,你现在真的可以用 Fable 一次性完成大量功能。所以看到这种转变,并和社区里的大家一起经历这一切,真的很棒。

I remember when we first came out with Claude Code and Sonnet 3.7, you would give it this task and you would have to closely monitor every single little thing that it tried to do. I remember I would read every permission prompt extremely carefully. I would frequently say no. I would always say no no no, like did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like we've all gotten a chance to just take a step back, delegate a lot more of the menial implementation to Claude. And it just freed up a lot of our time to think about more creative work, like what is the right experience that we should be providing to our users now that we know Claude Code can implement a lot of it. And now with Fable, it's just a totally different step change improvement. We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now. So it's been amazing to see the transition and to go through this with all of you in the community.

Thariq

是啊,我记得我收到关于 Claude Code 的第一条短信。我最好的朋友之一说:“哦,你得去试试 Claude Code。”那大概正好是 Opus 4 出来的时候,我试了试,然后我心想:“哦,靠。”我需要现在就去 Anthropic 工作,你知道吗?那是 Opus 4,我是说模型很棒,但当时还是有权限提示。我觉得我们的健忘程度有点离谱。我会想,哦,自动模式不是一直都有吗?我甚至都不记得按过“是”和“允许”。对我来说,我一直在逼自己的一件事就是:哦,我们必须做出比以往任何时候都更高质量的工作。你知道,输出质量高得不可思议。我一直在用它大量剪辑视频,我会想,好吧,它必须在几个小时内满足我们品牌团队极其苛刻的要求,否则我们就做不了。所以我觉得这就是我试图随着 Fable 转变的方式:好吧,做出我们做过的最好的作品,而且比以往任何时候都快。

Yeah, I mean I think I remember the first text I got about Claude Code. One of my best friends was like, "Oh, you need to go try Claude Code." And it was about just when Opus 4 came out and I tried it and I was like, "Oh, shit." Like I need to work at Anthropic now, you know? And that was Opus 4, which I mean great model but like yeah you were permission prompts and yeah I think it's kind of crazy how much amnesia we have. I think where I'm like oh like auto mode has always been here, right? I don't even remember pressing yes and allow. And yeah I think for me the big thing that I'm trying to push myself is like oh we have to do higher quality work than we've ever done before. You know like the outputs are like incredibly high quality. I've been using it to edit videos a bunch and I'm like okay it has to meet the very exacting demands of our brand team and in a couple hours or we just can't do it, you know. And so yeah I think that's how I'm sort of trying to shift with Fable where it's like okay the best work we've ever done, faster than we've ever done it before.

Host

我自己确实也有这种感觉。软件工程变得更难了,因为我们现在能承担的事情的野心水平提高了。有了这些工具支持我,我对自己有了高得多的期望,这很有趣,但工作量也很大。

I've certainly been finding that myself. Software engineering is getting harder because the level of ambition of the stuff we can take on has gone up. I have such higher expectations of myself now that I have these tools to back me up, which is fun but it's a lot of work.

Thariq

是啊。全是思考,你知道。很累。

Yeah. It's all the thinking, you know. It's tiring.

What Conventional Wisdom No Longer Holds

Host

那么,一年前还成立、但在这个新世界里你认为不再成立的一条传统软件工程观念是什么?

And so what's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?

Cat

我认为我们在工程技能组合中看到的最大转变之一是,你知道,两年前,产品经理通常会去和一堆客户聊,然后在六个月里和跨职能团队就某个 PRD 对齐,再写出详尽的规格和文档,说明在第一行代码写出来之前我们到底要怎么实现。而现在情况完全反过来了。对很多工程师,我想给在座各位的建议是,更多地培养你的商业意识和产品意识,去判断我们应该构建什么,因为现在从有想法到把它做出来的时间线短太多了。从 6 到 12 个月缩短到可能一周。这意味着我们所有人都需要对什么值得构建、什么会真正改变我们所服务的业务有更好的品味。所以我认为,在大多数产品领域,产品品味和商业意识的价值上升了,执行的价值稍微下降了。当然,对基础设施来说,仍然非常强调确保所有细节都正确。

I think one of the biggest shifts that we're seeing in the eng skill set is, I think, you know, two years ago it was pretty typical for a product manager to go talk to a bunch of customers and over the course of six months align with cross-functional teams on some PRD and then write this thorough spec and doc on how exactly we'll implement this before the first line of code gets written. And now things are completely turned the opposite way. I think for a lot of engineers, the push I would give to a lot of folks in the room is to develop more of your business sense and product sense on what is it that we should build, because now that the timeline between having this idea and building it is so much shorter. It's down from six to 12 months to maybe even a week. That means all of us need to have better taste on what is it that is worth building, what is it that will actually inflect the businesses that we're working on. So I think it's an increase in value on product taste and business sense and a bit lower on execution in most product domains. Of course for infra there's still a very heavy emphasis on making sure all the details are right.

Thariq

是啊。对我来说,就是重写现在变成好事了,你懂我意思吧?我觉得就像……

Yeah. I think for me it's like rewrites are now good, you know what I mean? I think that like...

Host

以前你能做的最糟糕的事,现在其实没问题了。

The worst thing you could do is now actually fine.

Thariq

是啊。没错。没错。就像所有那些,尤其是《人月神话》里说的“永远不要重写”。我现在支持重写,你知道。如果你有好的测试套件,我觉得重写实际上会迫使你确保自己有好的测试套件。但我觉得人们低估的是,代码库就是一份规格,也许它是你拥有的唯一一份规格副本,对吧?因为没人知道代码库的每一个分支部分。你可以把它当作一个产物,去提炼它,或者创建它的其他版本。显然,我们把 Bun 用 Rust 重写了,你知道,它运行得很好,现在就在我这儿跑着。

Yeah. Exactly. Exactly. Like all the like especially, yeah, mythical man stuff like never rewrite. I'm a pro-rewriting now, you know. Like if you have a good test suite, I think actually the rewrite forces you to make sure you have a good test suite. But I think that like what people underestimate is like a codebase is a spec and maybe it's the only copy of the spec that you have, right? Because like no one knows every branching part of the codebase. And yeah you can take this as like an artifact and like distill it or create other versions of it. Obviously like yeah we rewrote Bun in Rust and you know it works great, like you know it's live for me right now.

Host

但你们还没有把 Claude Code 跑在用 Rust 重写的 Bun 上,对吧?还是说……

But you're not shipping Claude Code on Bun in Rust yet, right? Or is that...

Thariq

内部已经跑了。

Internally we have.

Host

哇。哦,那太令人兴奋了。是啊。

Wow. Oh, that's exciting. Yeah.

Thariq

是啊。是啊。

Yeah. Yeah.

Host

我觉得你刚才说的关于,是啊,那个……抱歉,我刚刚思路断了。不过,重写这件事对我来说也非常有意思,因为你几乎可以先想出好的测试套件,然后同时启动三个实现,再挑出其中最准确的那个。我现在做更多原型了。我一直是个爱做原型的人,现在我在会议期间用手机做原型,这样之后就能接着用,而且现在真的能跑,这有点非凡。那么最近另一个大发布是 Claude Tag,到现在大概一周了吧,至少对我们其他人来说是这样。据我了解,Claude Tag 在 Anthropic 内部被非工程师大量使用。

I think what you're saying there about, yeah, the... I'm sorry, I just lost my train of thought. But yeah, the rewrites thing has been really interesting for me as well because you can almost come up with a good test suite and then spin up three implementations and pick which of those implementations was the most accurate. I'm doing a lot more prototyping now. I've always been a prototyper and now I prototype things on my phone during the conference just so I've got something that I can pick up later on and that's working now, which is kind of extraordinary. So the other big launch recently was Claude Tag, which that's what, a week old now I think, at least for the rest of us. Claude Tag, I understand, is being used in Anthropic by non-engineers a great deal.

Introducing Claude Tag

Host

非工程师用 Claude Tag 做些什么事情?

What kind of things are non-engineers doing with Claude Tag?

Thariq

Claude Tag 是一个住在你团队协作工具里的 Claude。我们上周在 Slack 里发布了它。Claude Tag 的不同之处在于它默认就是多人的。你把 Claude Tag 加进一个 Slack 频道后,你可以插话,你的队友也可以插话,大家可以一起协作处理 PR。另一个很大的不同是它是主动的,而不是被动的。你可以告诉 Claude Tag:嘿,监控这个频道里的每一条 bug 报告,提一个 PR 把它修好,并 @ 最后碰过这块代码的工程师,它会在频道的整个生命周期里一直这么做,不需要你手动把它拉进来。第三个大的变化是我们加入了团队记忆。如果你在频道里告诉 Claude Tag 你的偏好,它会在之后的每一条帖子里都记住。比如你总是想让它调试线上故障,但不想让它调试警告,你只要在频道里用自然语言告诉它,它就会为你和团队里的其他人记住。在我们内部,我们把 Claude Tag 看作 Claude Code 的进化。我们认为这是内部工作方式的一次巨大转变。Claude Tag 目前承担了我们 65% 的产品 PR。

So Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about Claude Tag is it's multiplayer by default. So once you add Claude Tag into a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. So you can tell Claude Tag, hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase, and it'll do it for the lifetime of the channel without you having to manually tag it in. And then the third big shift that we've seen is we've added team memory into this. So if you tell Claude Tag your preferences in the channel, it'll remember this for every future post. So if you always wanted to debug outages, but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team. Internally we see Claude Tag as the evolution of Claude Code. So we see this as a large shift in how we work internally. Claude Tag currently lands 65% of our product PRs.

Host

是整个 Anthropic,还是只针对 Claude Code?

For all of Anthropic, and for Claude Code, or just for Claude Code?

Thariq

这只是针对我们的产品工程团队。我们内部版本的 Claude Tag 目前承担了 65% 的产品 PR。这是一次巨大的转变,超过了我们 PR 的一半。我们看到的大家拆分 Claude Code 和 Claude Tag 工作的方式是:当你需要和智能体交互式地反复迭代时,Claude Code 仍然是最适合处理最复杂任务的地方。而 Claude Tag 则非常适合让它主动替你干活,这样你就不必再为正在开发的功能冒出来的所有 bug 报告手动启动 Claude Code 了。

This is just for our product engineering team. So our internal version of Claude Tag lands 65% of our product PRs right now. And this is a huge shift. This is more than 50% of our PRs. And the way that we actually see people split work between Claude Code and Claude Tag is Claude Code is still the best place for your most complex tasks when you're interactively iterating with the agent. But Claude Tag is great for having it work proactively on your behalf so that you no longer need to manually kick off Claude Code for all of the bug reports that might come up for features that you're working on.

Non-Coding Use Cases

Cat

是的。在非编程的场景里,我们见过大家用 Claude Tag,比如在这次分享之前,我们问 Claude Tag:嘿,Fable 什么时候发布?我们想确保能和发布公告对齐。于是 Claude Tag 会搜索我们的 Slack,看看谁说了什么。所以作为公司的搜索引擎,它非常有价值。它掌握着你产品的全部上下文,你可以问它指标相关的问题。很多时候你在做决策时,希望它能参考指标怎么说,然后你把它接到你的事件存储上。我见过我们的市场团队做这样的事:嘿,给我讲讲这个功能。他们不是程序员,但 Claude 是程序员。它可以克隆代码库,然后说:哦对,这就是那个功能,它长这样,这是我使用这个功能的录屏。所以它开启了一大堆可能性。我觉得我们还在很早期地摸索这些。

Yeah. And for non-coding cases, I think we've seen people use Claude Tag just, for example, before this talk we asked Claude Tag, hey, when is Fable releasing? We wanted to make sure that we'd line it up with the announcement. And so Claude Tag would search our Slack and look at who's been saying what. So as a search engine for your company it's really valuable. It has all the context for your product. So you can ask it metrics-related questions. And oftentimes when you're making decisions, you want it to be informed by what the metrics say. And then you hook it up to your event store. I've seen our marketing team do things like, oh, hey, tell me about this feature. And they're not programmers, but Claude is a programmer. It can clone the codebase and be like, oh yeah, this is the feature. This is what it looks like. This is a recording of me using the feature. So yeah, it just enables a whole wide variety of things. And I think we're still early on in figuring that out.

Claude Code Building Claude Code

Host

我觉得这是 Claude Code 故事里最迷人的一点:你们用 Claude Code 来构建 Claude Code,而且大概从一年半前 Claude Code 公开发布之前就在这么做了。我自己用编程智能体遇到的一个问题是,我明白作为个人怎么用它们,但我不太清楚在团队环境里该怎么用。听起来 Claude Tag 就是你们目前对这类团队协作层的答案。

Well, I feel like this is one of the fascinating things about the Claude Code story: you use Claude Code to build Claude Code, and you've been doing this since presumably before the public launch of Claude Code a year and a half ago. And one of the problems I've had with coding agents is I get how to use them as an individual, but I'm not really clear on how I use that in a team environment. It sounds like Claude Tag is your current answer to that sort of team collaborative layer for this stuff.

Thariq

正是如此。我们现在有很大一部分会话实际上是多人的。也就是说,也许我说:嘿,我觉得我们应该在 co-work 里实现这个新功能。然后我会 @ Claude Tag 让它先做第一版。接着我告诉 Claude Tag:嘿,把你最终实现的录屏分享出来。然后我会 @ 设计团队来看一眼,他们会微调一下,再把它交给 edge 团队收尾并发布到生产环境。所以这是一种非常流畅的体验。我们还在摸索如何协调同一个会话里的社交动态,但我们发现大家会观察别人怎么用,然后遵循这些社交规范,把 Claude Tag 融入团队对我们来说其实相当容易、相当直观。

Exactly. And a large percentage of our sessions are actually multiplayer right now. So that means maybe I say, hey, I think we should implement this new feature in co-work. And I'll tag in Claude Tag to do a first pass at it. And then I'll tell Claude Tag, hey, just share a recording of your final implementation. And then I'll tag in design to take a look and they'll nudge it, and then they'll pass it on to edge to take it to the finish line and get it out to prod. And so it's been this very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session, but we've found that people just observe how others use it and then follow those social norms, and it's actually been pretty easy, pretty intuitive for us to integrate Claude Tag into our teams.

Cat

是的,我觉得它很适合教别人,也能降低 SLO,因为如果你看到有人直接说:嘿 @Claude 修一下这个,你会想,呃,我觉得有一种社会性的——不是社会性的,就是大家都在看着你一起用 Claude,这本身也会提升你用 Claude 的水平。

Yeah, I think it's great for teaching people and also kind of reducing SLO because if you see someone just be like, hey @Claude fix this or something, you're like, uh, I think there's some societal—not societal, just the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well.

Host

对,你想做那种你愿意公开拿出来的工作,质量不会断崖式下跌。

Right, you want to do work that you're proud to do in public, where the quality doesn't fall off the cliff.

Prioritization and Dogfooding

Host

既然我们在聊用 Claude 构建 Claude,那构建 Claude 的流程是怎样的?你们怎么处理整个工程领域最难的问题——优先级排序,对吧?当构建一个功能现在便宜得多的时候,你们怎么决定哪些功能值得做、值得发布?这才是难的地方。

So since we're talking about using Claude to build Claude, well, how's the process for building Claude? How do you deal with the hardest problem in all of engineering, it's prioritization, right? How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now? This is the hard thing.

Thariq

我们有几种方式。一是我们每天都吃自己的狗粮。每当我们在产品里想做某件事却做不到时,我们不会去找别的解决方案,而是去修我们的产品,让它支持这个场景。我们内部有非常浓厚的吃狗粮文化。所以在把产品分享给全世界之前,我们先分享给 Anthropic 内部的每个人,也分享给一些早期客户,他们会给我们非常诚实的反馈。越狠越好。我们不断迭代,直到大家喜欢它。我们对一个功能在分享给全世界之前必须达到的活跃用户数和留存率有一个内部标准。因为这个标准非常清晰,每个工程师都知道自己要达到什么。我觉得这也提升了我们的打磨程度,因为如果功能没打磨好,用户就会流失,那我们就不该发布这个功能。

So there's a few ways we approach it. One is we dogfood our products every single day. Whenever there's something that we want to be able to do in our products that we're not able to, instead of finding a different solution, we fix our product so that it can support this case. We have a very heavy dogfooding culture internally. So before we are able to share our products with everyone in the world, we share it with everyone within Anthropic, and we share it with some early customers who give us very honest feedback about it. The more brutal the better. And we iterate until people love it. So we have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world. And because this bar is very clear, every engineer knows what they're trying to hit. And I think this also levels up our polish, because if the feature isn't polished, people will churn and then we shouldn't ship that feature.

A Surprising Feature

Host

你有没有一个让你意外的功能例子?你把它推出来,参与度高得离谱,它变成了——本来不太可能发布的东西,结果真的变成了一个真正的产品。

Do you have an example of a feature which surprised you? You rolled it out and the engagement was off the charts, and it became—was something that was unlikely to be shipped that actually turned into a real product thing.

Thariq

我确实有一个。我们团队里很多人喜欢远程控制。远程控制让你可以用移动设备或浏览器里的 Claude,连接到你在 CLI 里运行的本地 Claude Code 会话。

I do have one. So a lot of folks on our team love remote control. So remote control lets you connect—use your mobile device or Claude in the web browser to connect a local Claude Code session running in your CLI.

Remote control workflow

Cat

我从来没有这个需求,因为我直接在手机上启动任务,它在云端会话里运行,不使用我的本地环境。我想这是因为我做的都是很简单的编码任务。但这件事我一开始没有完全理解。我当时想,嘿,大家应该直接搭好自己的远程开发环境。但实践中,我们推出远程控制之后,我聊过的很多人都说:“好,现在我每天晚上做的就是——把笔记本电脑插上电源,合上屏幕,或者开一堆远程控制会话,锁屏,然后坐在沙发上用手机控制 Claude Code。”所以这就是我们现在正在拥抱的流程,我一开始没搞懂,但现在懂了。我就是这么做的。因为现在能远程控制了,我能在更舒服的环境里在笔记本上完成好多工作。这真的很好玩。

I never have this need because I just kick off the task directly on mobile and it runs in a cloud session and doesn't use my local environment. I think this is because I'm doing very easy coding tasks. But this is something where I didn't totally understand it. I was like, hey, people should just set up their remote dev environments. But in practice, once we rolled out remote control, so many people who I talk to are like, "Okay, now what I do every night is I plug my laptop into a power charger, close the screen or open a bunch of remote control sessions, lock the screen, and then use my mobile phone from my couch to control Claude Code." And so this has been this flow that we're now leaning into that I didn't originally get, but now I do. I do exactly that. I get so much work on my laptop done from more comfortable environments because I can remote control it now. That's really fun.

Host

代码审查是怎么运作的?你们会审查——进入 Claude Code 的每一行生产代码都有真人审查吗?如果没有,你们是怎么做的?怎么保证质量?

How does code review work? Are you reviewing — does a human being review every line of production code that makes it into Claude Code? And if not, what are you doing? How do you keep the quality up?

Cat

当然。是的。这很大程度上取决于任务。对于重要的区域,我们有代码负责人(code owner),对吧?系统提示词就是一个例子,我们有代码负责人——你必须得到他们的批准。

Sure. Yeah. It varies on the task a lot. So for important areas we have code owners, right? And so the system prompt is kind of an example where we have a code owner — you really need to get their approval.

Host

所以我想,代码负责人直接对那块代码的质量负责。

So I guess the code owner is directly responsible for the quality of that area of the code.

Cat

没错。是的。是的。

That's right. Yeah. Yeah.

Host

而且任何触及那块代码的 PR 都需要他们批准。

And they need to approve the PR that touches it.

Cat

对。没错。我们有代码审查——我们基于 Git 的代码审查会审查一切——所以每个 PR 都会走这个流程,很多时候它承担了大部分的审查工作。我在团队里看到的一点是,对于更复杂的 PR,你可能会做一个 artifact 来解释这个 PR,方便其他人来审查。而且我们投入了很多在验证、CI/CD 之类的事情上,确保任何时候有东西失败,我们都有测试,有一个非常健壮的环境,让 Claude 能控制 Claude Code 并测试它,你懂我的意思吧。所以代码审查是一个多管齐下的方法。你有什么要补充的吗?

Right. That's right. We have code review — our code review, Git-based, you know, reviews everything — and so that goes on every PR and oftentimes that's doing the bulk of the review. I think something I've seen on the team is that for more complex PRs you might make an artifact to explain the PR so that other people can then review. And yeah, we just invest a lot into verification, CI/CD things like that to make sure that anytime anything fails, we have a test, we have a really robust environment where Claude can control Claude Code and test it, you know what I mean. So yeah, there's just a multi-pronged approach to code review. Do you have anything to add?

Thariq

总的来说,我们正努力走向一个人类不需要在环内的世界。所以对于 Claude Code 核心以及其他产品核心最关键的改动,始终有代码负责人,他们会手动审查所有改动。但对于越来越外层的变化,我们实际上让 Claude Code 的代码审查完全接管。这听起来挺吓人的,但我们花了六个多月才走到这一步,我认为建立对代码审查的信任需要一步步来。一开始我们对所有东西都做人工审查,然后逐渐地我们会说,好,对于触及这些文件的代码改动,代码审查能 100% 抓住所有问题,所以我们实际上不需要人工去审查那些了。另外,当我们做事故复盘时,我们会看导致事故的 PR,然后说,好,我们怎么更新代码审查来抓住这个问题?然后我们还会把这些 PR 加进一个 eval 集,确保我们未来对代码审查的改动永远不会让那个指标回退。所以把人类从代码审查环里移除是一大步。我觉得这听起来可能吓人,也不是一夜之间能做到的,但你可以通过好几个月对基础设施的投入来获得信心,相信代码审查能抓住你在意的一切。

In general we are trying to move to a world where humans don't need to be in the loop. And so for the most critical core changes to the core of Claude Code and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly for the changes that are at the outer layers, we actually have Claude Code review fully review those. That sounds pretty scary, but we've had this six-plus-month-long process to get here, and I think there are baby steps that you take to build up trust with code review. So in the beginning we would have human review for everything, and then increasingly we would say, okay, for code changes that touch these files, code review is catching 100% of the issues there, so we actually don't need a human to be manually reviewing those. And then also when we have incident review, we look at the PRs that cause the incident and we say, okay, how do we update code review to catch that? And then we also take those PRs and add them to an eval set to make sure that our future changes to code review never regress that metric. So it is a big step forward — removing humans from the code review loop. I think it can sound scary and it's not something that you can do overnight, but it is something that you can do through many months of investment in the infrastructure to give you the confidence that code review is catching everything that you care about.

Host

你提到建立对模型的信任,这很有意思,因为这也是我发现的一点——我知道 Opus 4.8,如果我让它给我做一个运行 SQL 查询并输出 JSON 的 JSON 端点,它肯定能做对。那不是我需要仔细审查的东西。但然后一个新模型出来了,我还是不知道——我怎么快速建立对 Fable 的信任,相信它不会搞砸 Opus 不会搞砸的事?这是你们需要多考虑的事情吗,那些新模型——新模型如何影响你对它能做什么、不能做什么的直觉?

So it's interesting you mentioned building trust in the models, because that's something I found — I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, it's just going to get it right. That's not something I have to review closely. But then a new model comes along and I still don't know — how do I build trust in Fable quickly that it's not going to mess things up that Opus didn't? Is that something that you have to think about much, those new models — how does the new model affect your intuition for what it can do and what it can't do?

Thariq

我们之所以随着时间积累这个 eval 库,主要原因就是让新模型可以即插即用。因为我们有了新模型之后,会跑整个 eval 集,确保比如 Fable 严格优于 Opus 4.8,这给了我们把它换上去的信心。

So the main reason that we're building up this eval base over time is so that new models can be a drop-in replacement. Because what we do when we have a new model is we run the whole eval set and we make sure that, for example, Fable is strictly better than Opus 4.8, and that gives us the confidence to drop it in.

Host

那些模型 eval 是整个 Anthropic 的,还是你们 Claude Code 团队专用的 eval?

And are those model evals for Anthropic as a whole, or are these Claude Code team-specific evals that you're using?

Thariq

两者都有。我们团队有自己的 eval,而且我们在 Anthropic 内部的每个仓库都跑代码审查,所以我们也为此有 eval。对于像自动模式这样的东西,我们不仅在 Anthropic 内部每个用户上都有 eval,还委托了多个外部测试者来做红队测试,构造带有提示注入和恶意输入的环境,确保自动模式不会让其中任何一个通过。

We have both. So we have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic, but we've also commissioned multiple external testers to red team this, to create environments with prompt injections and malicious inputs and make sure that auto mode doesn't let any of those pass.

Host

那么对于 Claude Code 本身——这也是我在自己构建东西时遇到的一个挑战。我想知道我做的系统提示词改进是否真的改进了产品,对吧?这是最基础的产品专属 eval 形式。我到现在也没有很好的感觉该怎么做。你们在做这件事吗,以至于你们完全有信心,你们对系统提示词做的这个微调确实带来了更好的输出?

So for Claude Code itself — and this is a challenge I've had with stuff I'm building. I want to know if the system prompt improvement I made actually improved the product, right? That's the most basic form of product-specific eval. And I still don't have a great feel for how to do that. Is that something that you're doing, such that you have complete confidence that this tweak that you've made to the system prompt does result in better output?

Thariq

我们没有完全的信心,但我们做了很多来确保性能不回退。我们的起点是,我们有一套我们信任的外部 eval,并用一套更大、我们信任的内部 eval 来补充。一开始,我们主要优化的是能力。也就是说,给定一个完整定义的任务和完整的代码库,Claude 是否做出正确的决策,完全修好 bug 并通过所有测试?这是起点,也是我们优化的目标,因为它最直接地对应了用户想要的。但有很多行为会影响用户使用 Claude Code 时的感受。比如,人们非常不喜欢 Claude Code 说“该睡觉了”这种话。或者人们非常不喜欢它说“嘿,我完成了五部分中的两部分。你想让我继续吗?”——就像“是的,请继续”。所以我们正在建立一套行为 eval 来抓住这些。

We don't have complete confidence, but we do a lot to make sure that we don't regress performance. So the starting point that we have is we have a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, we mainly optimize for capability. So given a complete definition of a task and the full codebase, does Claude make the right decisions and fully fix the bugs and pass all the tests? So that's the starting point and that's the thing that we optimize for, because it is most directly what users want. But there's a lot of behaviors that impact how users feel when they work with Claude Code. For example, people really don't like it when Claude Code says it's like time to go to sleep. Or people really don't like it when it says, "Hey, I finished two out of five parts. Do you want me to continue?" Like, "Yes, please continue." And so we're building up a set of behavioral evals to catch these.

User Feedback and Evaluation Priorities

Thariq

当我们收到用户反馈时,请大声告诉我们你的用户反馈。收到反馈后,我们就排序,好,这些是优先问题。然后我们一个一个往下,为每个问题构建评估。所以不是 100% 覆盖,但我们努力——提高覆盖率是我们的优先事项。

And as we get user feedback, please be loud with us about your user feedback. As we get user feedback, we just rank, okay, these are the priority issues. And we go down one by one and build evals for each of them. So, it's not 100% coverage, but we try to—it is a priority for us to increase the coverage.

Host

那么,Claude Code 团队和 Anthropic 最初训练模型的团队之间有多少重叠——有多少互动?是相当紧密的合作吗?

And how much overlap is there between the—how much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place? Is that quite a close collaboration?

Thariq

在 Anthropic 内部,我们都合作得很紧密。所以我们经常开会讨论,比如我们期望下一代模型能做什么。我觉得我们的研究团队也很棒,会公开分享这些。所以我们在博客文章里经常谈到我们如何瞄准越来越长跨度的工作,如何训练 Claude 本身做到诚实、无害、有帮助。我们也投入了很多努力,确保它符合你的意图,即使你的意图表达得比较模糊。当然,尽量具体说明你想要什么。这样 Claude 就有全部上下文,但即使你不具体,我们也教 Claude 做出好的假设,是的,我认为这是一段富有成效的合作关系。

Now across Anthropic, we all work quite closely together. So we meet often to talk about like what do we expect the next generation of models to be able to do. I think our research team has also been amazing about just like showing this publicly. So we often talk in our blog posts about how we're targeting ever-increasing longer horizon work, how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure that it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want. So Claude has all the context, but even when you're not specific, we teach Claude to make good assumptions and yeah, I think it's been a productive partnership.

System Prompt Reduction and Model Judgment

Host

那么,Derek,你今天早上提到 Claude Code 的系统提示词减少了 80%。因为 Claude Fable,你能再详细说说那是什么样子吗?你们能够去掉哪些东西?

And so, Derek, you this morning you mentioned that the system prompt for Claude Code has been reduced by 80%. Because of Claude Fable, can you go into a little bit more detail about what that looks like? What kind of things have you been able to drop?

Thariq

是的。所以不只是 Fable,还有 Opus 4.8。是的,未来模型也是,但我们确实——我们现在为不同模型准备了不同的系统提示词。我认为我们看到的一些模式是,我们过度约束了 Claude,对吧?所以我觉得最初像 Opus 4 左右的模型想要很多例子,而移除例子非常有帮助,因为它比我们给的例子更有创造力。

Yeah. So it wasn't just Fable, it was Opus 4.8 as well. And yeah, going forward the future models, but we do sort of—we have different system prompts for different models now. I think that like some of the patterns we saw is that we were over constraining Claude, right? So I think the initial like maybe Opus 4ish kind of models wanted a lot of examples and removing examples was extremely helpful because it was just more creative than like you know the examples we gave it.

Host

那真是——因为我给人们的首要提示技巧之一就是给它例子,例子是最简单的方式。如果这不再成立,那有点打破了我的提示模型。

That's really—because one of the top prompting tips I give people is give it examples like examples are the easiest ways. If that's no longer true that kind of breaks my prompting model a little bit.

Thariq

同——是的,我也一样。我觉得听到这个很惊讶。我认为现在更多是关于你给 Claude 的工具的形状,以及你的系统提示词之类的。我们做的另一件事是——我们尝试给它更多上下文,更少“不要做这个”,因为我觉得这对 Claude 来说是一个很强的冲动,尤其是如果这后来与用户指令冲突,那对 Claude 来说会非常困惑,对吧?因为你会想,“哦,我有这个技能说这个,系统提示词说那个。”所以我们尝试减少硬约束,更多只是上下文,整体上更少的指令。是的,我认为这绝对是一门科学。构建它需要一堆评估。我不确定你在精简系统方面还有没有别的要补充。

Ex—yeah same same here. I think I was surprised to hear that. I think that now it's more about like sort of the shape of what you give the tools to Claude and like yeah your system prompt and things like that. The other thing we did is we—we try and give it more context and fewer like do not do this you know because like I think that it's just a very strong impulse to Claude and especially if that conflicts with user instructions later on that can be like extremely confusing to Claude, right? because you're like, "Oh, like I've got this skill that says this and the system prompt says this." And so we try and like have fewer hard constraints and more just like sort of context and just like fewer instructions overall. Yeah, I think it's definitely a science. It took a bunch of like evals to build. I'm not sure if you had anything else on the lean system.

Cat

我认为一般来说,当你给这些模型写提示词时,你应该总是思考,我给的指令有没有边缘情况?当我们回去审查 Claude Code 系统提示词中的所有指令时,我们发现有一些情况,是的,这个陈述 90% 是真的,但有 10% 的情况不是真的。我们不想约束模型或让它困惑,以为它应该总是这样做。一个很好的例子是验证。这里每个人都希望 Claude 验证它的工作。我们在提示词中有一些指令只是说,如果你做前端更改,总是验证。但你知道,这是有限度的。例如,如果它只是把文案从一个字符串改到另一个字符串,用户说只是快速修复并更新测试,也许你不想验证。所以我们调整了措辞,从说总是验证验证验证验证,变成嘿,大多数时候当你做前端工作时,你不能总是通过访问后端端点来理解完整体验。所以当你对用户体验做更多更改时,请在本地运行应用,实际上,那个指令可能甚至不好,因为什么是大更改,也许它想对小更改也测试。一般来说,每当你给模型一个提示词,你应该总是思考它可能被一个善意的其他用户或人类误解的方式,以便更好地理解模型可能如何解释它,并确保你可以软化提示词,使其实际上 100% 准确,因为你 100% 的时间都在给模型这个提示词。

I think in general when you're prompting these models, you should always think about like are there edge cases to the instruction that I'm giving it? And when we went back and we reviewed all the instructions in the Claude Code system prompt, we found a few cases where yes, this statement is like 90% true, but there's like a real 10% of cases where this is not true. And we didn't want to constrain the model or like confuse it into thinking, hey, it should always do this. Like one good example is verification. Everyone here wants Claude to verify its work. And we had some instructions in the prompt that just said, if you make a front-end change, always verify. But you know, there is a limit to it. Like for example, if it's changing copy from one string to another string and the user tells says it says like just make a quick fix and update the test, maybe you don't want to verify. And so we've also adjusted our wording from saying always verify verify verify verify to like hey most of the time when you're doing front-end work you can't always understand the full experience by hitting the backend endpoints. So like when you make like more changes to the user experience please run the app locally and actually in fact that instruction probably isn't even good because what is a large change like maybe it wants to change it test it for small changes too. In general, whenever you give a prompt to the model, you should always think about the ways in which it could be misinterpreted by like a well-intentioned other user or human in order to better understand how the model might interpret it and in order to make sure that you can soften the prompt such that it is actually 100% accurate because you are giving this prompt to the model 100% of the time.

Host

但有趣的是,你依赖模型的判断力。这必须是 Opus Fable 级别的东西。一年前的模型没有必要的判断力来决定是否要测试一个更改。这绝对迷人。但如果你为广泛的模型构建,并尝试用更便宜的模型处理更便宜的任务,这就会失效。

But what's fascinating about that is you're relying on the model's judgment. And that's got to be a Opus Fable level thing. Like models a year ago did not have the levels of judgment necessary to decide if they were going to like test a change or not. That's absolutely fascinating. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks and stuff.

Thariq

实际上,正因为这个原因,我们现在为每个模型准备了不同的系统。所以,只有我们最前沿的模型才有这 80% 的 token 减少,而较旧的模型实际上仍然有完整的系统提示词。

We actually have a different system per model now because of this very reason. So, it's only our most frontier models that have this 80% token decrease and the older models actually still have the full system prompt.

Host

你认为 Fable 和 Opus 足够聪明,能够给 Haiku 写更详细的提示词,因为它们理解 Haiku 判断力更少、品味更少吗?

Do you think Fable and Opus are smart enough to be able to prompt Haiku with more details because they understand that Haiku has less judgment, has less taste?

Cat

我们还没能评估,但我们没有任何硬数据来证明。我认为小模型有时有个难题,因为你知道我们看到过,有时大模型在难题上比小模型更 token 高效。所以你知道,需要建立一点那种直觉,比如有时你几乎总是想要前沿智能,但是的,就像平行曲线移动,所以很难找到。是的。

We haven't been able to evaluate, but we don't have any hard data to show it. I think there's a tough thing with smaller models sometimes because like you know we saw this with just like sometimes the larger models can be more token efficient on a hard problem than the smaller models. And so you know there's like a little bit of that intuition to build about like you know sometimes you really just want frontier intelligence almost all the time you know but it's yeah like the parallel curve shifts you know and so it's hard to find. Yeah.

Host

我的意思是,这是我发现很迷人的事情。我觉得一年前我不信任模型来写提示词。就像今天,好的模型非常擅长写提示词。就像我的很多提示词是由模型写的,这感觉荒谬,但实际上效果很好。帮助我接受这一点的是思考子智能体,这完全是一个 Claude 模型为另一个 Claude 模型设置提示词,以便它知道该去做什么。

I mean that's something I found fascinating. I feel like a year ago I did not trust a model to write a prompt. Like today the good models are very good at prompting. Like a lot of my prompts are written by models which feels absurd but it actually works really well. And something that helped me come to terms with that was thinking about sub agents which is entirely about a Claude model setting up a prompt for another Claude model so that it knows what to go and do.

Workflows and Claude Prompting Claude

Thariq

对,我觉得工作流其实就是一个很好的例子,因为它不只是给单个子智能体写提示,而是给多个子智能体的编排写提示,每个子智能体都会收到非常详细的提示。所以这几乎比单纯生成一个子智能体高了一个层级。所以,它在这方面确实很擅长。我在自己的机器上也一直在用,比如给它 Gemini API,然后说“来,生成图片”。它做得非常好,比我给图像模型写提示要勤快多了。所以,就是 Claude 一路提示 Claude。

Yeah, I think workflows are actually a really good example of this, because it's not just prompting a single sub-agent, but it's prompting the orchestration of many sub-agents, and each one of them gets a very detailed prompt. So it's almost like a level above just spawning a sub-agent. So yeah, it's quite good at that. I've also been using it on my personal machine, like giving it the Gemini API and being like, "Oh, here, generate images." And it's so good at it. It's way less lazy than I am at prompting an image model. So yeah, it's just Claude prompting Claude all the way down.

Host

我觉得 Claude 也写了工作流工具的提示。

I think Claude also wrote the prompt for the workflow tool.

Thariq

工作流工具的提示我读过,写得很好。其实这是我对 Anthropic 普遍感到不满的一点:你们会发布 Claude 的提示——有个网页上有——但不包括工具提示和 Claude Code 的提示。我还是得跑个代理去拦截它们。如果 Claude Code 的提示能被有意发布出来就好了,因为它们就是文档。它们能让你知道工具能做什么、怎么工作。

For the workflow tool, I've read that prompt. It's a good prompt. I mean, that's actually a frustration I have with Anthropic generally: you publish the prompts for Claude—there's a web page with them on—but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. I would love it if the Claude Code prompts were deliberately published, because they're the documentation. They're how you know what the tool can do and how it works.

Cat

我会记下这个功能请求。

I'll write down that feature request.

Thariq

请务必。

Please do.

Cat

我会让 Claude 标签来做。

I'll have Claude tag do it.

Thariq

还有差异对比。比如我时不时会对比新旧提示的差异,这样就能了解新模型的能力。我真的很期待看到这个 80% 的缩减到底是什么样子。

And also the diffs. Like every now and then I'll diff the older and the newer prompt, and that's how I learn the capabilities of the new model. I'm really looking forward to seeing what this 80% reduction actually looks like.

Cat

对,这是我的责任。我得详细写一篇关于这个的文章。

Yeah, this is on me. I have to make a post about this in detail.

Tool Design Philosophy

Host

那么,你的标准是什么?我们稍微聊聊工具。Claude Code 基本上就是一大袋工具。你引入新工具的标准是什么?你怎么决定什么时候值得在那个层面做额外的工程?

So, what's your bar? Let's talk about tools a little bit. Claude Code is basically just a big bag of tools. What's your bar for introducing a new tool? How do you decide when it's worth doing that additional engineering at that level?

Cat

你想来回答吗?因为你引入了我们最好的工具之一。

Do you want to take it? Because you introduced one of the best tools we have.

Thariq

对。我是说,引入“询问用户问题”工具时,我的职业生涯达到了顶峰。我觉得是这样。这真的很难。尤其是像“询问用户问题”这样的工具——它是 Claude 用来问你的工具——所以很难评估,有时更多是用户偏好问题。所以尤其是当时我们的评估还比较少,非常依赖“蚂蚁吃狗粮”——抱歉,是“吃狗粮”——对,“蚂蚁吃狗粮”是我们蚂蚁版的叫法。但总的来说,我们一直在趋向于减少工具数量。我记得我们最后引入的一组工具是任务工具,试着给 Claude 更通用的版本来做这些事。

Yeah. I mean, yeah, it's like my career peaked when I introduced the ask user question tool. I think so. It's really hard. I think especially for some tools like ask user question—it's Claude's tool to ask you—so it's hard to eval, and sometimes it's more of a user preference thing. So especially back then we had fewer evals. It was very ant fooding based—or sorry, dog fooding is—yeah, ant fooding is our ant version of that. But yeah, I mean, I think overall we've been trying to trend towards fewer tools. I think the last set of tools we introduced were like the task tool, I think, and try and give Claude more general versions to do this.

Host

对,最有趣的工具之一是文件编辑工具,但你可以把文件编辑作为一个工具,或者教它用 sed 和 grep 之类的方式来做。你们文件编辑工具的最新演进是什么?

Right, one of the most interesting tools is the file editing tool, but you can have file editing as a tool or you can teach it to use sed and grep and do things that way. What's the latest evolution of your file editing tool?

Thariq

我觉得我们还有一个,但比如我们移除了 grep 和其他搜索工具、glob 工具,改用原生 bash。所以,我们还有一个。我觉得这就像我之前演讲里说的:模型更像生物学而不是物理学。所以很难——尤其是工具设计,我觉得相当难。我不确定 Cat 是否不同意,觉得“哦,我们应该——不,评估它是有科学的”,但我倾向于认为,工具设计更像一门艺术,或者像生物学。

I think we still have one, but for example we removed our grep and other search tools, glob tools, for just like native bash. So yeah, we still have one. I think this is kind of like I said in my talk earlier: the models are kind of more of a biology than a physics. So it's hard—especially tool design, I think, is quite hard. And I'm not sure if Cat disagrees and is like, "Oh, we should—no, there's like a science to the eval of it," but I'm sort of like, yeah, tool design is more of an art, maybe, or like a biology.

Cat

对,我大体同意,但我觉得总的来说,随着我们引入更多工具,我们尽量保持基数很低,并确保我们添加的每个工具与其他工具都有不同的功能,这样 Claude 就能很容易地区分何时调用哪个。对于文件编辑,实际上,我们有文件编辑是因为我们可以渲染它。以前我们展示——或者我猜我们现在仍然展示——当 Claude 修改文件时,我们会向人们展示,有一个漂亮的专用 UI 说“你批准对这个文件的编辑吗?”我们有一个专用文件编辑工具的原因,是为了能确定性地知道 Claude 在修改文件,这样我们就能向人们展示这个漂亮的 UI。对于很多新加入的用户,我觉得他们仍然很喜欢这种体验。所以我们保留了它。但对于我们很多现在处于自动模式的人——或者希望你不是在 YOLO 模式——但总之,对于我们现在很多人来说,我觉得它其实不重要,我们可能可以直接移除文件编辑,也完全没问题。

Yeah, I think I largely agree, but I think in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that every tool we add has a distinct function from every other tool so that Claude can very easily distinguish when to call each. For file edit, actually, the reason that we have file edit is because we can render it. Back in the day we used to show—or I guess we still do—so we show people when Claude makes a file change, and there's this nice dedicated UI that just says, "Do you approve this edit to this file?" And the reason that we had a dedicated file edit tool was so that we could deterministically know that Claude was making a file so we could show people this nice UI. And for a lot of the new users who are onboarding, I think they still really like this experience. So we've kept it around. But for a lot of us who are on auto mode right now—or hopefully you're not on YOLO mode—but anyway, for a lot of us right now, I don't think it actually matters, and we probably could just remove file edit and we'll be totally fine.

Safety and Auto Mode

Host

那么,我们聊聊自动模式,或者聊聊总体上的安全与保障。我深知提示注入的风险,如果别人告诉我的 Claude Code 该做什么,可能会发生很多糟糕的事情。我仍然大多在 YOLO 模式下运行 Claude Code,并为此感到非常内疚。Anthropic 内部对于安全运行 Claude Code 有什么建议?比如你们告诉人们该怎么做?

So, let's talk about auto mode or let's talk about safety and security in general. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. What's the advice within Anthropic for safely running Claude Code? Like what do you tell people to do?

Cat

为什么不用自动模式?

Why not auto mode?

Host

我开始用自动模式了,但我对它了解不够,不知道它有多安全。不过,大概从三周前开始,我默认用自动模式了。

I am starting to use auto mode and I don't understand it enough to get how safe it is. But yeah, that's as of maybe three weeks ago, I'm defaulting to auto mode.

Cat

好,所以在 Anthropic 内部,几乎每个人都用自动模式。这是在 Claude Code 中安全地做长时间运行工作的最佳方式。我们做了大量的测试。我们有数千个评估。我们委托了许多红队成员创建对抗性环境,试图诱骗 Claude Code 做坏事,我们已经缓解了他们发现的每一个问题。所以我们将在未来几周发布一些评估,但我们基本上已经缓解了所有攻击。

Okay, so broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found. And so we're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack.

Host

这是一个很大的声明。非常令人兴奋。如果这能站得住脚,

That is a big claim. That's very exciting. If that holds up,

Cat

我们会分享评估结果,让大家可以评估,但我们非常努力地识别 Claude 可能出错的所有方式,然后更新自动模式来应对。它不能 100% 捕捉所有问题。我不认为任何——对,那会是一个太强的声明。但对于我们担心的主要风险类别,比如提示注入、数据外泄,风险远低于普通人类审查者。

We will share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things. I don't think any—yeah, that would be way too strong of a claim. But for the main categories of risks that we're concerned about, like prompt injection, data exfiltration, the risks are far lower than the average human reviewer.

Host

那么,对了,稍微讲讲自动模式如何工作。我觉得建立这个心智模型很有用。

So, and oh yeah, a little bit on how auto mode works. I think it's useful to build this mental model.

Auto Mode and Permissions

Thariq

So, whenever Claude is doing a turn, there's a bash call, there's a Sonnet classifier that is judging the tool and also the context of the conversation, your instruction, right? And so there are some things around like permissions which are dependent on your request, right? So you don't want to give git push permissions all the time, but if you say hey push this to GitHub, you want it to do it, right? And so auto mode will — or if you say don't push, you want it to deny it, right? And so auto mode will do that. That particular thing happens to me a lot where it's like, oh, like auto mode — like Claude tried to do this because it's very helpful and proactive, and auto mode saw like, oh, you know, don't do this, and it surfaced it. So it's good at the dynamic permissions that you yourself give inside of the prompt, which I think is really important. It also works well with our sandboxing infrastructure because sandboxing is one of those things where there are so many different edge cases. And it's hard for us to deterministically follow them. But if we have a sandbox and something needs to escape the sandbox, like a network request, auto mode can then look at that request and be like, oh hey, does this make sense, right? And just allow that in.

Thariq

So, whenever Claude is doing a turn, there's a bash call, there's a Sonnet classifier that is judging the tool and also the context of the conversation, your instruction, right? And so there are some things around like permissions which are dependent on your request, right? So you don't want to give git push permissions all the time, but if you say hey push this to GitHub, you want it to do it, right? And so auto mode will — or if you say don't push, you want it to deny it, right? And so auto mode will do that. That particular thing happens to me a lot where it's like, oh, like auto mode — like Claude tried to do this because it's very helpful and proactive, and auto mode saw like, oh, you know, don't do this, and it surfaced it. So it's good at the dynamic permissions that you yourself give inside of the prompt, which I think is really important. It also works well with our sandboxing infrastructure because sandboxing is one of those things where there are so many different edge cases. And it's hard for us to deterministically follow them. But if we have a sandbox and something needs to escape the sandbox, like a network request, auto mode can then look at that request and be like, oh hey, does this make sense, right? And just allow that in.

Host

So I hadn't realized auto mode is interacting with the networking sandbox as well.

Thariq

Yeah, exactly. So it's also part of sandbox. Yeah.

Host

It interacts with any permission prompt that the user would otherwise see.

Thariq

And how old is auto mode? Like I feel like as a feature that I had access to, it's only a couple of months old, right?

Host

We've been using it within Anthropic since January.

Thariq

Okay.

Host

So we've been hardening it for quite a while. And it's obviously Anthropic is extremely focused on safety and security. And so we've been working broadly across our alignment and safeguards teams in order to enable the rollout internally, build out these evals even more robust before sharing it out with the world. I think my only main problem with auto mode is I don't understand it deeply enough. Like for anything that's looking after my security, I want to know as much as I can about how it works and what it protects me against and what it doesn't so I can decide how much I can trust it.

Thariq

I think yeah Dell was working on a post about this. So a little bit just a little bit more about auto mode. This is also the reason Claude Tag is so good, right? Because Claude Tag uses auto mode and like you can imagine that like one of — I've heard a lot of like build versus buy questions on a Slackbot. I'm like please you probably shouldn't build your own AI Slackbot, you know, like there's so many attack vectors, you know what I mean? And like you have a feedback channel that like users can post feedback into and now your bot is reading it, right? And so I think that like this — the work we've put in with auto mode and you know we have a general Swiss cheese defense of like security, right? We also like yeah you know like RL against this stuff and things like that. I think this is really what makes Claude Tag work. It like just works seamlessly with your permissions and yeah like you know we you don't want to be prompt injected in your Slack. Yeah. Yeah.

Host

Do you have any — are there any more security things in the pipeline beyond that go beyond auto mode?

Thariq

I think — I mean I think we're very secure. Like so we with Claude Tag you can — Peri can provision your own like sort of credentials for Claude so it doesn't need to act on your behalf. You can have like Claude as an identity and that also makes it easier to audit and inspect what Claude is doing.

Host

Well I guess cuz Claude Tag is influenced by anyone who can talk to it. So it's got a much wider pool of people who are telling it what to do.

Thariq

That's right. Yeah. Yeah. And of course we have probes as well like with like Mythos and uh sorry with Fable and that's also like a downstream effect of our safety and research work. And I think this is the moment where you sort of see AI like Anthropic being an AI safety company really paying off when you like you know we really want Claude to be able to run in an aligned way over long periods of time and like yeah auto mode has to be basically flawless for this to work right and it's sort of like all downstream of our like you know our being an AI safety company. We also launch trusted devices for the remote control users out here who want to be safer. And for all of our remote environments, we support credential injection. So if you want like Claude Code to be able to access DataDog but you don't want Claude Code itself to hold the DataDog credential you can set up our identity credential management system so that the DataDog credentials are only usable by the agent but not accessible by the agent. So we insert it on the fly when the agent tries to make a DataDog request.

Host

This is that token proxying trick, right? Where the proxy knows anytime somebody calls this an API.Dog.com address with a token, replace the token with the real thing. I love that pattern. I'm seeing that in a whole bunch of places. It feels so obviously right to me.

Human Element and Craft

Host

Let's talk a little bit about the human element. And you touched on this in the keynote this morning, but a lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models. How do you think about that? Like how firstly how has the past year and a half changed the way you think about your own craft and the value that you add?

Thariq

Yeah, I think for me and I think Kat is always such a good reminder. Ken and Boris are such good reminders of like you have to be more ambitious. They're always like, you know, you have to like — we're growing so fast, we have to be on the edge, we have to do like the best work we can. I think that that's kind of like a constant reminder for me where I'm like anytime I'm like kind of like slow on something. I'm like, okay, can I do it faster? Can I be more ambitious here? I think the like point on and like oftentimes the answer is Claude because Claude is getting better as you go. So, I'm like, 'Oh, the last time I tried this, it was with the previous model or something.' I think with your point on loss, I think this is real. And I do feel that like if you're only trying to do the same work you were doing before LLMs and now it's like a prompt, it is like I think kind of a sad feeling.

Offsetting AI with ambition

Thariq

我觉得你抵消这种影响的方式就是变得更有野心,对吧?我特别喜欢 Jared 这个例子。他在奥克兰的公寓里花了一年时间手写所有 Zig 代码,几乎不出门,而且他做得特别开心。现在我又看到他把整个 Bun 用 Rust 重写,他也做得特别开心,对吧?这野心大得多,这就是他抵消那种影响的方式。我觉得总体上就是,好吧,我怎么去做更大的事、做更多的事?我觉得成功是件有趣的事,你知道,这就是我大概——

And I think the way you offset that is by being more ambitious, right? I love that Jared is such a good example. He hand-wrote all of the Zig code in his Oakland apartment in like a year, barely left his house, and he had so much fun doing that. And now I see him rewrite all of Bun into Rust and he's having so much fun doing that, right? It's so much more ambitious, and that's how he offsets that. I think just generally being like, okay, how do I do the bigger thing and do more? And I think success is fun, you know, and that's how I kind of like—

Host

就是改变你的野心。改变你做的事,因为你以前做的事,打个引号说,要容易得多,但现在我们可以承担这些更大的挑战。

It's changing your ambition. It's changing what you do, because what you did before is a lot easier in quotes, but now we can take on these bigger challenges.

Thariq

对,我觉得平均来说,每个人都有一些希望自己做过的事,你懂我意思吧,还有一些希望自己更擅长的事。我觉得现在就是,那就去做吧,你知道,就像我们——对。

Yeah, I think there's just like, on average, everyone has things they wish they did, you know what I mean, and that they were better at. And I think now it's, let's do it, you know, like let's— Yeah.

Product management in the AI era

Host

Cat,那从产品管理的角度来看,这又是什么样子呢?

And Cat, what does that look like from a sort of product management perspective?

Cat

我觉得产品这个角色每个月都在变,很大程度上就是去识别,好吧,现在是什么——我们团队里所有 PM 都是工程师、设计师、PM 的混合体。我们团队里大多数工程师其实以前都做过全职工程师。所以对我们来说,这真的意味着哪里有缺口就补哪里。所以如果是,好吧,我们有了这个想法,但没有激发任何工程师去把它做出来,那我们就应该自己把它做出来,然后放进一个 notebook 里,激发别人把它推进到生产环境。或者如果设计看起来有点不对,那我们就找一张类似的页面,先做一版初稿设计,然后拉一个非常注重细节的人来填补缺口。或者如果我们注意到,嘿,现在我们团队和产品在公司的采用率大了一点,更多人需要知道 Claude Code、Claude Tag 和 Co-work 接下来要出什么,那我们做的就是,好吧,让我们把整个发布日历自动化,让我们把那些状态更新异步化,这样就不用去打扰别人,然后让我们弄清楚,好吧,这是我们三个内部公告渠道,确保我们在那里的更新足够详细、切中要点。所以对我们来说,很大程度上就是理解现在从一个好想法到把东西送到客户手里之间的缺口是什么,然后我们如何尽可能把它自动化。

I feel like the product role just changes every single month, and it's very much just identifying, okay, what are— like all the PMs on our team are this mix of engineer, designer, PM. Most of the engineers on our team actually used to be full-time engineers in the past. And so for us it really means plugging in whenever there's any kind of gap. So if it's like, okay, we have this idea and we didn't inspire any engineer to go build it, then we should just build it and then put it into a notebook and inspire people to take this to production. Or if the designs look a little off, well, let's take a page that's similar and do a first-pass design and tag in someone who's very detail oriented to fill in the gaps. Or if we notice that, hey, now our team and our product adoption is a bit bigger within the company and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Co-work, what we do then is like, okay, let us automate figuring out our whole launch calendar, let's automate getting those status updates asynchronously so we're not bugging people, and then let's figure out, okay, these are our three internal announce channels and make sure that our updates there are fully detailed and to the point. And so for us, it's very much just understanding what is the gap right now between a great idea and getting something to our customers, and then how do we automate it as much as possible.

When Claude surprised you

Host

听起来产品管理总是有更多事要做,对吧?让我感觉不错的一件事是,我从没在一家没有上千件想做却没资源做的事的公司工作过。那么,Claude 什么时候让你惊讶过?就是模型做了某件真正让你惊讶、你没想到它能做到的事?

It sounds to me like with product management, there's always more to do, right? One of the things that makes me feel good is I've never worked at a company that didn't have a backlog of a thousand things they wanted to do and didn't have the resources to take on. So what's a moment when Claude has surprised you? Like when the model has done something that genuinely surprised you that you didn't think it would be able to do?

Cat

对,我发过很多关于 Claude 视频剪辑的内容,但最近我在 ACM Agentic 大会上做了个演讲,我说,嘿各位,你们有剪辑好的视频吗?我想发出来分享给我的传播团队。他们说,哦,这个要花很久。我说,好吧,能把原始文件发给我吗?于是他们把我台上演讲的视频、幻灯片的视频和音频文件发给我,说,祝你好运。然后我把这个交给 Claude。我也把我的 HTML 幻灯片给它,我说,嘿,你能把这个剪到一起吗?它做出来的东西真的不可思议。我都准备好发布了。它把整个视频转录了。它注意到有时候我幻灯片的视频有点怪——中间有个自动更新的弹窗——它就说,哦,我大概不该用你幻灯片的视频。其实,我要做的是切分并弄清楚你在哪一页,然后用你的 HTML 源文件。于是它显示 HTML 源文件。然后它有我的一段视频,但我在舞台上只占一小部分。所以它动态地裁剪我在舞台上的位置。而且我在来回走动。所以它跟着我走动来追踪我。然后我有了我的裁剪画面、幻灯片,接着它在转录我说的话。

Yeah, I mean, I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference and I was like, hey guys, do you have the edited video? I'd love to post it and share with my comms team. And they're like, oh, it's taking so long. And I'm like, okay, could you send me the raw files? So they send me the video of me talking on stage, the video of the deck, and the audio file, and they're like, good luck. And so I give this to Claude. I give it my HTML deck as well, and I'm like, hey, can you just edit this together, you know? And what it does is honestly incredible. I'm ready to ship it. So it transcribes the entire video. It notices that sometimes the video of my deck is a little bit weird—there's a popup of an auto update in the middle—and it's like, oh, I probably shouldn't use the video of your deck. Actually, what I'm going to do is I'm going to slice up and figure out which slide you're on and instead use your HTML source, right? And so it's displaying the HTML source. Then it's got a video of me, but I'm only taking up a small part of the stage. And so it's cropping dynamically where I am on the stage. And I'm pacing. So it's tracking me as I'm pacing. And I've got this crop of me, the deck, and then it's transcribing what I'm saying.

Host

这是 Fable 做的,对吧?

This was Fable, right?

Cat

这是 Fable。对。对。绝对是。对。

This is Fable. Yeah. Yeah. Absolutely. Yeah.

Cat

而且就是,你知道,那是个好提示词,但只是一次性提示词。然后我让它加一些有趣的动画和图形,我就被震撼到了,你知道,它就把这些全做了。它用 ffmpeg,用 Remotion,它——

And it's just like, you know, it was a good prompt, but it was a one-shot prompt. And then I asked it to add some interesting animations and graphics, and I was just blown away, you know, and it just does all this stuff. It does ffmpeg, it does Remotion, it does—

What Claude still can't do

Host

那我得追问一句。它做不到什么?还有哪些事让你失望?你在等 Claude Fable 6 帮你解决?

So I have to ask the follow-up. What can't it do? What are the things where you're still disappointed? You're waiting for Claude Fable 6 to figure out for you?

Cat

我希望它有更好的设计和 UX 品味。

I wanted to have better design and UX taste.

Host

嗯哼。

Uh-huh.

Cat

我觉得它现在已经到了这个程度:如果我给它——如果我写一个提示词,详细说明我希望某个功能怎么表现,它通常会那样表现。但内边距可能不对,或者界面还不够令人愉悦。我觉得它有点依赖现有的应用最佳实践,依赖应用设计的既有方式,但我觉得对于前沿 AI 产品,还有太多新的交互体验等着我们去设计。

Like I feel like it's now at the point where if I give it—if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off or the interface is just not delightful yet. I think it kind of leans on existing best practices for apps, for how apps are designed, but I feel like for frontier AI products, there are so many new interaction experiences that we still have yet to design.

Host

有一种 Opus 美学。你看一样东西就能说,对,这是 Opus 设计的。如果我们能超越那个就好了。

There's an Opus aesthetic. You can look at something and go, yeah, that was designed by Opus. It'd be good if we could move beyond that.

Cat

对。对。我非常期待未来的模型能成为交互设计的思想伙伴。

Yeah. Yeah. I'm very excited for future models to hopefully be like interaction design thought partners.

Host

嗯。它做不到什么?

Hm. What can't it do?

Cat

我觉得我很想看到它更多地与真实世界互动,比如,好吧,它能做这个吗,它能解决科学问题吗,对吧?它能编排实验吗?这里面有一定量的编程,但它还需要对更广阔世界的另一种品味。所以——

I think I would love to see it interact more with the real world, like, okay, can it do this, can it solve science, right? Can it orchestrate the experiments? And there's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs. So—

Host

Claude Science 是几天前刚出的新产品,对吧?

Claude Science is a new product that just came out a few days ago, right?

Cat

对,但我对它没有了解。

Yeah, but I have no context on it.

Host

我正想问,那是 Claude Code 的一部分,还是单独的一块?

I was going to ask, is that part of Claude Code or is that a separate section?

Cat

那是我们的合作团队做的。

It's our partner team.

Host

明白了。不过还是试试吧。

Gotcha. Try it out, though.

Cultural hacks worth stealing

Host

那么我们有——我有几个收尾问题。你觉得 Anthropic 公司文化中哪些部分独特地帮助 Anthropic 用这些工具保持高效,其他公司应该借鉴?人们应该从你们这里学走哪些文化上的窍门?

So we've got—I've got a couple of closing questions. Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools that other companies should steal? What are the cultural hacks that people should be adopting from you?

Cat

我先分享一个,然后你来。

I'll share one and then you go.

Claude Tag Best Practices

Cat

我来分享一个关于 Claude Tag 的用法。Claude Tag 在公共频道里、并且你的大多数频道都是公开的时候效果最好。Claude Tag 能够搜索所有公开频道,获取尽可能多的上下文,从而给出准确度最高的回答。而它只有能访问所有内容时才能做到这一点。

I'll share one for Claude Tag. So Claude Tag works best when you have it in a public channel and when most of your channels are public. Claude Tag is able to search across all public channels to get as much context as possible to give you the highest accuracy answer. And it's only able to do this if it has access to everything.

Ambition and Trade-offs

Thariq

是的,我在主题演讲里提到过,但我觉得这对我太重要了,我想再强调一下:我认为联合创始人们说过,我们不要跟自己谈判。这真的很重要——你可以在脑子里想象各种权衡,然后说服自己不去做有野心的事。或者你可以直接去尝试做那件有野心的事。我觉得我们常常会想,好吧,如果我们直接做了会怎样?比如,这真的是一个权衡吗?如果是,为什么?有什么证据证明这是一个真实的权衡,而不只是听起来合理?所以我觉得,就让权衡自己显现出来吧。尽可能有野心。

Yeah, I mentioned this in my keynote, but I think it's so important to me. I want to reemphasize: I think the co-founders say we don't negotiate against ourselves, you know, and I think this is really important where you're like you can imagine trade-offs in your head and talk yourself out of doing something ambitious, you know? Or you can just try and do the ambitious thing. And I think that we're just so often being like, okay, what if we just did it? Like what if, you know, is this a real trade-off or not, right? Or if so, why? Where's the proof that it's a real trade-off and not just it sounds reasonable, right? And so I think just yeah, make the trade-offs show themselves to you. Be as ambitious as you can.

Host

这太有意思了,因为这和我的经验完全相反。我有 25 年的软件经验,它告诉我默认答案应该是“不”。一切都是权衡,一切都有成本。而现在我们不得不重新想象所有这些直觉。这挺迷人的。好,那么最后一个问题,问你们两位:你们用 Claude 做过的最喜欢的荒诞事情是什么?仅仅因为你能做就做了。

That's so because that goes against I've got 25 years of software experience that says the default answer should be no. Everything is a trade-off. Everything has a cost. And now we're having to reimagine all of those intuitions. It's kind of fascinating. Okay. And so final question for both of you. What is something, what's one of your favorite absurd things that you've built with Claude just because you could build it?

Absurd Projects with Claude

Thariq

我先来。我在做一个 2D 街头霸王格斗游戏,角色是我自己,还有我的朋友们。它用 Claude Code 来提示 Gemini,说实话那个场景舞蹈模型做视频动画相当不错。效果很好。它很擅长提示。它能验证帧,检查这是不是一段好的动画。

I can go. I'm working on a 2D Street Fighter fighting game with me as a character and my friends as well. And it uses Claude Code to prompt, you know, Gemini and honestly the scene dance model is pretty good to make video animations. And it works great. It's so good at prompting. It can verify the frames to check if this was a good animation.

Host

你生成的是《街头霸王 2》那种级别的 2D 精灵图吗?

Is this Street Fighter 2 level 2D sprites that you're generating?

Thariq

对,没错。就是 2D 精灵图。动画看起来很棒。它还能算出判定框。它会说,哦,你的拳头在这里,我来画 JSON 判定框。是的,太不可思议了。所以,我不知道会不会发布这个,但它……

Yeah, exactly. Yeah. Yeah. Like 2D sprites. The animation looks amazing. And it can also figure out hitboxes. It can be like, oh, your fist is here, I'll draw the JSON hitbox. Yeah. It's like incredible. So, I don't know if I'll put this out, but it's

Host

我觉得至少得来个截图。这听起来……

I feel like we need a screenshot at least. This sounds

Thariq

我可以搞个截图。是的。

I can make a screenshot happen. Yeah. Yeah.

Cat

我的要简单得多。我是个攀岩爱好者,我很多朋友也攀岩,所以我们用 Claude Code 做了个小应用,用来记录我们正在攻克的所有线路,我们也经常一起去户外。所以我们让 Claude 用工作流做所有研究。工作流太棒了。我们把它包装成编码工具,但它做旅行深度研究非常出色。我还用它来规划团队团建,它很擅长找到能容纳我们所有人的场地。它有很多附带好处。总之,我还用工作流来研究我们可能想去的所有攀岩目的地,从我们各自所在地有直飞航班的地方。它会去 Mountain Project 找所有适合我们难度等级的线路。它找 Airbnb,还会实际规划路线——我不喜欢徒步,所以我很在意接近路线要非常短。从停车的地方到岩壁的步行距离要非常短。所以它会按这个筛选。用现有应用我得手动在 Mountain Project 上点来点去,但用这个我只要输入我们所有的偏好,它就是我们专属的定制应用。

Mine is much more simple. I'm a big rock climber and a lot of my friends climb and so we have this little app that we built with Claude Code where we just log all the projects that we're working on and we also go outdoors together a lot. So we have Claude do all this research with workflows. Workflows is amazing. We brand it as a coding tool but it's amazing for doing deep research for travel. I also plan our team offsites and it's good at finding venues that can fit all of us. It has a lot of side benefits. But anyway, I also use workflows to just research all the climbing destinations that we might want to go to, what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs that are in our grade level. It finds the Airbnb and it actually maps out — I don't like hiking and so I care a lot about it having a very short approach. So very short walking distance from where the car parks to where the rock actually is. And so it filters for this and so with existing apps I have to manually click through Mountain Project but with this I just put in all of our preferences and it's just a custom app for us.

Host

所以你基本上是在用 vibe coding 给攀岩做一个 Jira。

So you're basically vibe coding Jira for mountain climbing.

Cat

没错。那太棒了。

Exactly. That's pretty fantastic.

Audience Q&A

Host

这样吧,我们还有时间回答几个观众提问。如果你想上前来对我说,我会重复一遍让大家都能听到。但好的,这样吧,谁先到这儿谁先问。后面的朋友抱歉了。

You know what? We have time for a couple of audience questions. If you want to come forward and say them to me and I will repeat them so everyone can hear them. But yeah, tell you what, anyone who gets here first gets to ask a question. Sorry for people at the back.

Audience

嘿,呃,其实,是的,我的问题是:你们近期有没有计划构建更多评估工具,让我们能构建评估数据集之类的,以及更多可观测性工具来监控智能体和工作流的性能?

Hey, um, actually yeah, uh my question is that do you have any near plan to build more eval tools for us to build eval data set or anything like that and more observability tools to monitor the performance of agents and workflows?

Thariq

我们考虑过构建评估工具,但我认为限制因素实际上往往是客户需要很长时间才能构建出真正高质量的评估。所以我认为工具本身不是主要约束,更多是技能问题:如何构建一个好的评估?这是我们很兴奋地在内部投入的领域,也希望我们能在外部分享一些最佳实践。

We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high quality evals. And so I think the tooling is less of the constraint and more the skill set of how do you build a great eval? And that's an area where we're excited to both invest internally and also hopefully we can share some of the best practices externally.

Audience

嘿 Cat,我叫 Sai。我的问题是我对记忆和多玩家更感兴趣。记忆是如何设计的?所以两个问题:第一,今天记忆是如何设计的?我猜是围绕文件。第二,你们有没有考虑过另一个方向,即实际上需要一个数据存储来存储这些记忆,而不是文件,以便更好地扩展?这就是我的问题。

Hey Cat, my name is Sai. So my question was because I'm more interested in the memory and the multiplayer. How is memory being designed? So two questions right, so how is memory being designed today? I assume it's around files. So and second part of it is have you thought about thinking in an orthogonal direction where you would actually need a data store to store these memory instead of files to scale it better. So I think that's my question.

Cat

是的,目前 Claude Tag 的记忆是频道特定的。所以那个频道里的每个 Claude 都有一个共享记忆,然后实例有一个会话,但会话可以回馈到主记忆。我们做了很多记忆研究,而且什么才是正确的记忆方式可能有点反直觉。但没错,我们一直在研究这个。

Yeah right now for Claude Tag the memory is channel specific. So every Claude in that channel has a shared memory and then the instances have a session but the session can contribute back to main memory. We do a lot of memory research and it can be kind of unintuitive like what is the right way to do memory. But yeah we're always working on this.

Thariq

是的,我的意思是我们一直在做记忆实验。我没有什么——是的,就像现在 Claude Tag 里的工作方式是每个频道一个 markdown 文件。

Yeah, I mean we're always running memory experiments. I don't have anything to — yeah, like how it works right now in Claude Tag is a markdown file per channel.

Audience

谢谢。

Thank you.

Host

恐怕我们时间到了。请和我一起感谢 Cat 和 Thariq,我们会在走廊里继续回答更多问题。

So I'm afraid we're out of time. Please join me in thanking Cat and Thariq and we will be around for more questions in the hallway.

Thariq

谢谢大家。

Thanks guys.

Cat

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →