The Future of Claude Code: Mods, Mutable Software, and Multiplayer Agents — Thariq Shihipar, Anthropic
打开互动全文版(中英对照 + 朗读 + 问答)→Anthropic 的 Thariq Shihipar 探讨 Claude Code 如何通过 Mods、可变软件与多智能体协作演进,以及如何跟上前沿节奏。
Anthropic's Thariq Shihipar explores how Claude Code is evolving through mods, mutable software, and multiplayer agents, and what it takes to keep up with the frontier.
我觉得不同模型之间差异很大,你懂我的意思吧?但我意识到维护不同的版本实在太痛苦了。而且随着模型越来越好,它们完成简单任务的下限也在提高。所以我确实认为最终 CLAUDE.md 会消失。也许甚至不用等那么久。我觉得现在开始一个新项目,可能最好不要用 CLAUDE.md。
I think different models are very different from each other, you know what I mean? But I realize that it's such a pain to maintain different ones. And as the models get better and better, the floor of how they accomplish the simpler task gets better. So I do think in the limit CLAUDE.md goes away. And maybe not even that far. I think that right now it might be better to start a new project without a CLAUDE.md.
是的。
Yes.
我觉得也许如果你看到非常重复的失败模式,你就把它们加到 CLAUDE.md 里。真正棘手的是,这因模型而异。所以如果你加了一堆失败模式,就像……
I think maybe if you see very repeated failure modes, you add them to your CLAUDE.md. The really tough thing is that this changes per model. So if you've added a bunch of failure modes, like...
你需要 Fable MD,你需要 Opus MD。
You need Fable MD, you need Opus MD.
嗯,甚至 Fable 5.1 和 Fable 5 之间,你知道,就很烦人。我们不是故意这样的,你懂我的意思吧?这就是模型的工作方式,对吧?所以也许 Fable 5 有某个失败模式,而 Fable 5.1 没有。如果你一直记录一堆不同的失败模式,它们可能会过度约束 Claude。所以我们其实刚刚为技能加了评估插件。现在你可以评估一个技能是否更好。
Well, even Fable 5.1 versus Fable 5, you know, it is annoying. It's not like we do this on purpose, you know what I mean? It's just how the models work, right? So maybe Fable 5 had this failure mode that Fable 5.1 doesn't. And if you keep this running log of a bunch of different failure modes, they will probably over-constrain Claude. So we actually just added eval plugins for skills. And so now you can eval if a skill is better.
在我们进入今天的节目之前,我有一小段话要对听众说。谢谢你们。如果你们没有选择点击并收听我们的内容,我们就无法为你们带来你们如此渴望的 AI 工程、科学和娱乐内容。几乎每天都有人找我们谈赞助。但幸运的是,你们中有足够多的人订阅了我们,让这一切在没有广告的情况下也能持续下去,我们希望保持这样。但我只想请大家帮一个忙。你能做的最有力、完全免费的一件事,就是点击那个订阅按钮。这是我唯一会请求你们做的事,它对我和我的团队意义重大,我们每周都在努力把 Inspace 带给你们。如果你们订阅了,我向你们保证,我们会不停努力,让节目变得更好。现在,让我们开始吧。
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you, we'll never stop working to make the show even better. Now, let's get into it.
我们和来自 Anthropic 的朋友 Thariq 一起在演播室。总的来说,Claude Code……边界融合得太多,而你自从加入 Anthropic 以来一直紧跟一切。你很早就开始用 Claude Code 本身,而且你也在其他播客里讲过那个故事,你还谈过“像智能体一样看问题”。最近你做了 AI Engineer 的顶级演讲《Fable 实战指南》,显然 Fable 是你们发布的,所以那算作弊。而且你最近还发布了 Claude Tag,我们也会聊到前沿的节奏问题。Anthropic 有太多事情在发生。我想首要的问题是,在 Anthropic 有这么多事情发生的时候,是什么感觉?
We're here in the studio with our friend Thariq from Anthropic. And I guess generally the Claude Code... there's so much merging of boundaries, and you've been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you've told that story in other podcasts, and you've also been talking about seeing like an agent. Most recently you did the top AI Engineer talk, Field Guide to Fable, which obviously you guys launched Fable, so that's cheating. And mostly you most recently also launching Claude Tag, and we're also going to be talking about pacing on frontier. There's a lot going on in Anthropic. I guess top of the question is what's it like being at Anthropic when there's so much going on?
我觉得有时候会让人措手不及。我加入 Anthropic 是因为 Claude Code。Claude Code 刚出来,我就觉得这太好了。而 Opus 4 对我来说……我简直无法想象它有多好,你懂我的意思吧?那对我来说是一个真正的时刻。但我当时试图说服我那些创业的朋友用它来写代码,他们说,哦不,我们的工程师觉得它不够好之类的。我说,这太疯狂了。现在你快进一下,12 个月,甚至不到,它就成了所有人写代码的默认方式,对吧?我觉得从推销它,到现在教人们如何最大限度地利用它、变得更高效之类的,就是一个巨大的转变。而且,我觉得作为一个人,很难跟上所有事情,你知道,事情发生得太快了。
I think you can get whiplash sometimes. When I joined Anthropic, I joined because of Claude Code. Claude Code had just come out and I was like, this is so good. And Opus 4 to me was just... I could not imagine how good it was, you know what I mean? And that was a real moment for me. But I was trying to convince my startup friends to use it for coding, and they're like, oh no, our engineers don't think it's good enough or something. And I was like, that's insane. And now you fast forward, you know, 12 months, less, and it's just the default way that everyone codes, right? And I think having to go from selling it to now teaching people how to make the most use of it and be more efficient and things like that is just a big change. And yeah, I think it's just hard to stay on top of everything as a human, you know, things happen so fast.
更多智能体来做这件事。
More agents at it.
我是说,是的,其实智能体式的东西比人扩展得好得多,就像,哦,现在有三件事在发生,而且都是紧急情况,你怎么应对?是的。
I mean, yeah, that's actually the agentic stuff scales much better than the human stuff, where it's like, oh, there are three things happening right now and they're all emergencies, and how do you respond to it? Yeah.
你的时间怎么分配?你做很多技术写作、工程工作。
What do you split your time on? You do a lot of technical writing, engineering work.
是的。我觉得当我加入 Claude Code 团队时,我想教人们怎么用 Claude Code。我觉得这一直是……我以为我可能只会花一点时间在上面。我一开始花了一些时间在智能体 SDK 上,我不太确定在 harness 方面“苦涩的教训”会怎么走,对吧?我觉得有时候我们会想,哦,Claude Code 之后是什么,你懂我的意思吧?所以一开始我想,我只想教人们怎么用 Claude Code,让 Claude Code 更容易用。我觉得随着 harness 越来越好,现在最主要的问题就是:你怎么用这些智能体,对吧?这是一个技能表达度很高的事情。所以我做这个,然后我做工程工作。我做演讲。但我觉得当我做工程工作时,我的目标是拿到我们从用户那里得到的反馈,然后能够谈论如何用 Claude Code 来做工程。所以那里有一个很好的循环。
Yeah. So I think when I joined the Claude Code team, I wanted to teach people how to use Claude Code basically. And I think that has been something that... I thought maybe I would spend a little bit of time on it. I was spending some time on the agent SDK first, and I wasn't exactly sure how the bitter lesson would go when it comes to harnesses, right? I think sometimes we were like, oh, what's after Claude Code, you know what I mean? And so initially I was like, I just want to teach people how to use Claude Code and make it easier to use Claude Code. And I think that as harnesses have gotten better and better, that's the dominant problem now: how do you use the agents, right? It's such a high skill expression thing. So I do that, and then I do engineering work. I give talks. But I think when I'm doing engineering work, my goal is to take that feedback that we get from users and also then be able to talk about how to use Claude Code to do engineering. So there's kind of a good loop there.
是的,对听众来说,我们会附上你和 Sarah 为 dev writers 聚会做的那场演讲,我们稍微聊过。嗯,先做工作,然后谈论工作,差不多这样。播种和收获,还是什么?
Yeah, for listeners, we'll attach the talk that you did with Sarah for the dev writers meetup, which we talked a little bit about. Well, first you do the work and then you talk about the work, something like that. Sew and reap, or what was that?
然后收获。播种和收获。
And reap. And sew and reap.
差不多,差不多。是的。然后稍微预告一下,我们要聊 harness 的演变。它从只是一个 CLI 走了很长一段路。我们要聊 Claude Mods,它今天开始泄露了,因为你没法保密。
Something like that, something like that. Yeah. So and then just to preview a little bit, we are going to talk about the evolution of the harness. It has come a long way from just being a CLI. We're going to talk about Claude Mods, which is starting to leak today because you couldn't keep it secret.
是的,是的,是的,基本上是的,是的。
Yeah, yeah, yeah, basically, yeah, yeah.
是的,那里有很多东西。我觉得你一开始加了“向用户提问”工具,人们其实又爱又恨。我其实觉得它很有创新性,然后现在我有我自己的版本。你有你的“采访我”版本。
Yeah, there's a lot there. I think you started off with adding the ask-user-question tool, which people would love and hate actually. I actually thought it was very innovative, and then now I have my own version. You have your interview-me version.
是的。是的。
Yeah. Yeah.
而且,是的,每个人都有自己的东西,这已经不重要了,因为现在你应该写能生成其他提示词的提示词、循环之类的所有东西。
And yeah, everyone just has their own stuff, and it no longer matters because now you're supposed to write prompts that create other prompts and loops and all these things.
当然。
Sure.
那么今天的最新技术水平是什么?比如人们在……你今天告诉人们要做什么?
So what's the state of the art today? Like what are people... what are you telling people to do today?
是的。“向用户提问”是模型第一次擅长引导式提问。
Yeah. Ask user question was the first time that the model was good at elicitation.
你知道,我觉得这算是一种涌现行为,我想看看模型能不能做到。我有一个人机交互的背景。我在本科和研究生阶段学的就是这个。所以对我来说,这其实算是人机交互,就是试图弄清楚智能体如何与你沟通并提取需求,对吧?
You know, I think this was kind of an emergent behavior that I wanted to see if the models could do. I have a human-computer interaction background. So I did that in undergrad and grad school. And so this was, I think it's kind of human-agent interaction to me, like trying to figure out how can the agent communicate with you and extract the requirements, right?
我觉得随着 Claude Code 变得越来越广泛,困难之一就是每个人都有自己的使用方式,而要改变默认行为其实非常难。比如说,如果有人让 Claude Code 做某件事,有时候他们只是想让它把活干了,因为他们可能是很擅长写提示词的人;而有时候他们其实并不擅长写提示词,你懂我的意思吧,这时候智能体就需要澄清。所以这是一个很好的分界,而“向你提问”这个工具就沿着这条线来划分:你是觉得自己已经足够好、能直接指挥智能体,还是智能体需要挖掘出更多需求、与你更多地协作、真正理解你的偏好?
I think one of the things that's difficult as Claude Code has gone broader and broader is that everyone has their own way of using it, and it's actually very hard to change the default behavior. So for example, if someone asks Claude Code to do something, sometimes they just want it to do the work because they're maybe a very good prompter, and sometimes they actually are not good at prompting, you know what I mean, and the agent needs to clarify. And so that's a good split, and the ask-you-the-question tool sort of splits along that side: do you feel like you're good enough to instruct the agent as it is, or does the agent need to pull out more requirements and collaborate with you more and really understand your preferences?
我总体上相信,几乎所有人都更偏向后者而非前者,也就是说他们有更多的模糊性,对问题的了解比他们自以为的要少。但要让这件事变得容易,这是一个界面设计问题,你懂我的意思吧?所以如果你在设计一个问题,或者你在处理一个问题,像 schema 是什么、调用栈是什么这类东西就非常重要。设计中的细节很重要。理想情况下,你希望在开始实现之前就提前弄清楚其中一些难题。对,这就是为什么它们被称为未知数,对吧?所以我认为,在智能体式编程中,弄清楚你的未知数将永远是一项技能。因为即使模型超级智能,它也需要知道你想要什么,你懂吧,而且你有偏好,你需要把那些偏好挖掘出来。
I on the whole believe that pretty much everyone is more on the latter than the former, that they have more ambiguity and they know less than they want than they think they know about the problem. But it's an interface design problem to make that easy, you know what I mean? And so if you're designing a problem or if you're going through a problem, things like what's the schema or what's the call stack and things like that are really important. The details in the design are important. Ideally you want to figure out some of these hard problems ahead of time before starting implementation. And yeah, that's why they call them unknowns, right? And so I think that this will forever be a skill in agentic coding: figuring out your unknowns. Because even if the model is super intelligent, it needs to know what you want, you know, and you have preferences, and you need to sort of pull that out.
所以这就是我认为我在推动的方向。那么问题就是,智能体如何与你交互?我认为 HTML 一直是实现这一点的主要方式。我们最近还加入了 artifacts,对吧?而 artifacts,我其实觉得我们在解释如何充分利用它们方面做得不好,或者说我做得不好。我们有很多强大的能力。它们关联着一个数据库,所以每个 artifact 都能存储和写入持久化数据。它们可以反馈回 Claude,所以有一件人们还没在做、而我正试图鼓励的事,就是“仪表盘 artifact”这个想法。
And so that's, I think, what I'm pushing. The question then is how does the agent interact with you? And I think HTML has been the big way of doing that. And we've recently added artifacts, right? And artifacts, I actually think we've done a bad job of, or I've done a bad job of, explaining how to use them fully. We have a lot of powerful capabilities. They have a database associated with them, and so every artifact can store and write persistent data. They can feed back into Claude, and so one thing that people are not doing yet that I'm trying to encourage is this idea of a dashboard artifact.
比如你让 Claude 长期做一个项目。也许是一个看板之类的。它可以把看板数据存进数据库。多个 Claude 可以通过 artifact MCP 访问那些数据,而那个 artifact 也可以和那些 Claude 对话。所以我们基本上是在构建一些原语,让你能够通过 artifacts 拥有这种生成式界面,从而让你从智能体那里呈现出更多丰富的细节。
So you have Claude working on a project long term. Maybe it's a kanban or something. It can store that kanban data in a database. Multiple Claudes can access that data via the artifact MCP, and that artifact can talk to those Claudes as well. And so we're basically building the primitives for you to be able to have this generative interface via artifacts that will let you surface more of that rich detail from the agents.
而且我认为,现在智能体相关的几乎所有事情都是这个问题:你以为你知道自己想要什么,但你其实并不真正知道自己想要什么。而智能体需要大量细节。与它们保持人在环中的协作非常重要。所以 artifacts 就是我们试图在那里演进的方式。但还有很多工作要做,因为它比一道选择题要复杂得多,你懂吧?就图表、代码片段、schema,或者那个问题的任何东西而言,细节要多得多。但 artifacts 基本上是更“AGI 信仰”式的提问方式。所以,对。
And I think that almost everything with agents right now is this problem of: you think you know what you want, but you don't really know what you want. And the agents need a lot of detail. And collaborating with them in the loop is really important. And so artifacts are the way that we're trying to evolve there. But there's a lot of work to do, because it's so much more complicated than a multiple-choice question, you know? There's a lot more detail in terms of diagrams and code snippets and schemas, or whatever it is for that problem. But artifacts is the more AGI-pilled way of doing ask-a-question, basically. So yeah.
我觉得关于这个,关于 artifact 这些东西,有一点我不太清楚,就是哪些反馈应该通过 artifact 进入,哪些反馈应该通过 Claude 聊天进入,因为更“AGI 信仰”的做法就是把所有东西都喂给 Claude。
I think one thing that's unclear to me about this, the artifact stuff, is like what feedback should go in through the artifact and what feedback should go through a Claude chat, because the more AGI-pilled one is to just feed everything to the Claude.
我觉得更“AGI 信仰”的做法是通过 artifact。我认为我们大致想象,在极限情况下,artifacts 会成为你进入 harness 的界面。你可以在这份关于你工作计划的活动文档上评论。你也许能看到多个智能体,不同的智能体在做这做那,而那个 artifact 是为你现在正在做的工作构建的,对吧?所以每一个都略有不同。我认为从基础设施的角度看,我们还在朝那个方向走,但没错,我认为为你的 harness 提供一个即时生成的界面,大概就是事情的发展方向。
I think the more AGI-pilled one is to go through the artifact. And I think that we sort of imagine, in the limit, I think that artifacts will be your interface into the harness. You know, you can comment on this live document of your plan of the work. You can see maybe multiple agents, and different agents are doing this and that, and that artifact is built for the current work that you're doing, right? And so each one has slightly different. I think we're still getting there from an infrastructure perspective, but yeah, I think an on-the-fly interface for your harness is probably where things are headed.
有没有一个版本,是从 CLI 或聊天中抽象出来的?因为现在很多时候是这样的:好,你在和 Claude Code 交互,你拿到返回的 HTML 作为原型,它相当丰富,有图表,artifacts 是把这些连接起来的方式。那为什么不干脆全都那样做呢?那样就变成了把推理发生在哪里、智能发生在哪里、工作发生在哪里分离开来,你懂吧?
Is there a version of it that's an abstraction from CLI or chat? Because right now a lot of it is, okay, you're interfacing with Claude Code, you're having HTML given back for a mockup, it's pretty rich, there's diagrams, artifacts are ways to connect these together. Why not just do everything that way? Then it becomes like separating out where's the inference happening, where's the intelligence happening, where is the work happening, you know?
我觉得这有点像是本地和云端之间的区别,或者说其中一些区分,对吧?所以我觉得现在如果你用 Claude Code,它是本地的,你可以比如启动远程控制来获得一些云端行为,或者你可以在云端启动 Claude Code,对吧?我们正朝着这样一个方向走:不再是给本地 Claude 发消息、它在本地启动一个会话并执行,而更像是你给一个在云端运行的 Claude 发消息。它可以运行本地会话或云端会话。Claude Tag 大致就是这样工作的,但随着时间的推移,我们也会加入本地 hands。
I think this is kind of the difference between, or some of the distinction between, local and cloud, right? And so I think right now if you use Claude Code, it's local, and you can spin off remote control for example to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We're moving towards a place where instead of, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that's in the cloud that's running. It can run local or cloud sessions. This is kind of how Claude Tag works, but over time we'll add local hands as well.
所以本地 hands 将是那个智能体在你的电脑在线时访问它、并在那里工作的能力。这样它就能启动许多不同的子智能体。那些子智能体可以互相通信。而这就是 artifact 发挥作用的地方,基本上就是用来展示所有这些工作。所以你可以想象,你在把这些东西分离开来。有一个表层 UI 显示,它是一个 artifact,托管在某处,有数据库等等。
And so local hands will be the ability for that agent to access your computer if it's online, and be able to work there. And so it can spin off many different subagents. Those subagents can communicate with each other. And that's where the artifact comes in, to display all of that work basically. So you can imagine you're separating out these things. So there's the surface UI display that's an artifact and hosted somewhere and has a database and everything.
有推理智能,对吧?那是在云端发生的,你不必担心关掉电脑之类的,对吧?然后还有像手一样的东西,它可以是本地的,也可以在远程沙箱里,或者任何你需要完成工作的地方。这就像拆解云代码体验。现在一切都发生在一个地方。
There is the inference intelligence, right? That's happening on the cloud and you don't have to worry about shutting off your computer or whatever, right? And then there's the hands kind of like and it can be local, it can be in like a remote sandbox or wherever you need your work to be done. That's like unpackaging like the Claude Code experience. Right now it all happens in one place.
你怎么看这方面的多人协作?比如团队想以这种方式工作,现在还很个人化,但你怎么看多人协作的未来?现在有 Cloud Tag 算是一个版本,但我们正在推出 Projects,Projects 是一种抽象,有点像 Cloud Tag,但在我们的云产品上,所以你可以给它发消息,它会做类似 Cloud Tag 的事情,比如生成子智能体。
How do you see like the multiplayer side of that so say teams want to work in this way right now it's very individual but how do you see the future of multiplayer like right now I guess there's cloud tag which is a version but we're launching projects and so projects is the like this abstraction that's kind of like cloud tag but on our cloud products right so you can message it and like it will do the cloud tag like stuff like spinning off sub agents.
所以我们认为在多人协作方面,Claw Tag 更像原生的多人协作,因为它就在你的 Slack 里,权限什么的都搞定了。但我确实认为多人协作是重要的一部分,需要更多地整合起来。你可以想象当你说“哦,你有手,但现在别人的电脑上也有手,你需要给它们权限,或者你有你的 MCP,别人有别人的 MCP,你怎么弄清楚如何使用它们?”这会变得相当复杂,而 Claw Tag 在打磨这些问题上做得很好。所以当你……是的,Google Docs,它如何访问 Google Docs?它通过共享的云 MCP 访问,或者如果它没有访问权限,也可以通过你的本地凭证访问。但是的,我认为 Cloud Tag 是我们的多人协作产品,对于这些本质上是多人协作的事情非常有用,比如值班,例如事件本质上是多人协作的。你想标记 Cloud,你想多个人登录,你想能够找到上下文。嗯,我觉得每当我做某事并想要隐私或安全,或者想让他人审查时,有一个每个项目的频道真的很好,比如我会在法律频道说:“嘿,我想发布这个。Cloud 知道一切,和它聊聊吧。”这样法律部门就能得到关于代码中具体发布内容的精确答案,而我不需要在循环中。所以我认为多人协作越来越……是的,每个人都可以参与 Claude。我认为 Claude Tag 就是那个产品,而 Projects 会从单人开始,然后扩展。
So we think with multiplayer like claw tag is like a little bit more native multiplayer because it's just like in your slack and the permissions are all figured out and stuff like that. But I do think multiplayer is like an important part of the story and like uh that that will need to get tied together more. Like you can imagine how complicated it gets when you're like, "Oh, you have hands, but now you have other hands in other people's computers too and like you need to like permission them or like you have like your MCP and someone else's MCP and how do you figure out how to use them?" Right? It gets like quite complicated and claw tag does a good job of like sanding down all of these issues, right? So that like when you have Yeah. Google Docs, how does it access Google Docs, right? like it accesses through the shared cloud MCP or it can access through your local credentials as well if it doesn't have access. But yeah, like I think cloud tag is our multiplayer um product and it's really useful for these like things that are inherently multiplayer like okay like on call for example incidents are inherently multiplayer. You want to tag cloud, you want multiple people to log in, you want to be able to find context. Um, I think whenever I'm like working on something and I want like privacy or security or like I want other people to review it, you know, it's really nice to like I'll have a channel per project and I'll like at legal for example be like, "Hey, like I want to ship this. Can you like like here's the cloud knows everything, you know, just chat with it." And that way legal gets precise answers, you know, on like what exactly is shipping into the code and I don't need to be in the loop, right? So I think like multiplayer is getting like more and more like um yeah everyone can participate with claude. Um I think claude tag is like that that product and like projects we'll start off single player and we'll like you know expand.
我觉得有一个关于身份和隔离单元的问题,也许是双重问题。
I think there's a question about like maybe dual questions about identity and the unit of isolation.
是的。
Yeah.
嗯,Tag,你特意选择让它有自己的身份,这是一个有争议的选择。还有其他方法可以做到。
Um tag you specifically chose to make it its own identity which is like uh controversial choice. There's there's other ways to do it.
是的,
Yeah,
云项目可能听起来像,你知道,如果它像 CHP 项目,嗯,隔离就是那个工件,那个云实例,大家都在协作。听起来应该像,如果你在法律上协作,那个频道应该是一个项目,对吧?现在还不是,但那是自然的下一步。
cloud projects probably it sounds like you know if it's anything like CHP projects uh it is uh you know the isolation is that artifact that cloud instance everyone's collaborating on this it it'll it sounds like you know it should it should be like if you're if you're collaborating illegal on on on the thing like that channel should be a project right like it's not yet but it that's the natural next step.
是的,我的意思是,在 Cloud Tag 中,实际上就像 Cloud Tag,你必须自己安排,所以 Cloud Tag,是的,每个频道你可以随意命名,我基本上把每个功能命名为一个频道。但我认为有一些转移,不清楚何时有转移,因为假设你有一个同事被标记在所有这些东西上,是的,有转移,因为是同一个人。但对于 Claude,不清楚是否必然如此,比如“不,你不知道其他东西,你应该只用这个。”
Yeah, I mean like I think in cloud tag it's effectively like like cloud tag you have to sort of do your own arrangement basically and so cloud tag yeah each channel is like you can name it as you want and I name like each feature basically as a channel but I think like there is some trans like it's unclear when there is transference cuz let's say it if you have a coworker who is tagged on all these things yes there is transfer because it's the same person uh but with claude it's unclear if it's like necessarily like well no you don't know any you don't know about the other stuff you should only use this stuff.
这就像冰山一角的梗,对吧?我们花了很多时间在这上面,基本上有无限的表面区域,好吧,你想要云……不是无限,但有很多表面区域需要弄清楚,比如权限和可见性,以及如何让 Claude 尽可能安全地操作。显然这对我们非常重要,因为代码库的安全非常非常重要,所以我们投入了大量时间。是的,有很多边缘情况你可以想到,比如“哦,这个频道里的云有不同的权限,但它可以给另一个频道发消息,它不能通过那种方式泄露数据吗?”或者“如果它使用你的 MCP 然后给别人发消息呢?”有很多,我们真的投入了很多工作来打磨。
It's like the tip of the iceberg meme, right? where you can like this is what we spend so much time on basically is like there is like infinite surface area of like okay you want clouds to not infinite but like there's like surface a lot of like uh surface area to figure out of like permissions and visibility and like uh you know like how can you have let claude operate as well as you can as safely as you can you know and obviously this is very important to us because like you know like like security for our codebase is very very important and so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this cloud in this channel has different permissions, but it can message another channel and can't it exfiltrate data that way or like can you like what if it uses your MCP and then messages someone else? Like there's like so much and we've like really put a lot of work into sanding it down.
是的。是的。很多工作。嗯,好的。Fable。
Yeah. Yeah. Lots of work. Um, okay. Fable.
Fable,你写了两篇好文章。我是说,你写了很多好文章,但关于《Fable 构建云代码的实地指南》。我很好奇,从你看到的,在像 Anthropic 外部顶级用户中,你看到任何常见模式吗?比如充分利用云代码的最佳实践是什么?
Fable, you wrote two good articles. I mean, you've written many good articles, but on uh, you know, field guide to Fable building Claude Code. I'm curious from what you've seen is there any common patterns that you see in like top users adanthropic externally like what are best practices for getting the most out of Claude Code
我说的元技能是提示非常重要,你知道,我认为这并非 trivial,因为很多人说“哦,提示不重要,我只需说一句话,云就会做。”我认为提示真的就像公开演讲,或者写作,针对特定受众,而那个受众就是 Claude,你需要建立 Claude 的心智模型,了解它如何思考、如何工作。所以与云代码合作最重要的技能就是拥有 Claude 的心智模型,知道它擅长什么、能一次完成什么、不能做什么。很多人提示时,他们的提示很短,但他们对 Claude 和代码库有很好的心智模型,所以看起来毫不费力,你明白我的意思吗?但这是高技能上限。所以花大量时间提示并建立智能体如何工作的直觉非常重要。然后我认为下一件事是我们之前谈到的未知事物,即能够发现你不知道或没有写下来的东西。嗯,学习不同的东西。
the like meta skill I say is like prompting is like very important you know and like that like I I think this is like not trivial to say because I I think a lot of people are like oh prompting doesn't matter it's just like I can just say a sentence and cloud will do it and I think prompting is really this like this it's like public speaking you know like or writing or something and for a specific audience and that audience is Claude and you need to like build a mental model of Claude and how it thinks and how it works right and so that's like the most important skill in working with Claude Code is like having this mental model right of Claude and like what it can do well what it can oneshot what it can't and so so many people when you see prompting they're just like they're short prompts but they have such a good mental model of Claude and of like the codebase and things like that that like it's effortless, you know what I mean? But it's like high skill ceiling. So like that work of like you know spending a lot of time prompting and building mental models of how you know an intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier where it's like being able to find out like your what you don't know or what you haven't written down. Um learning about like different things.
我认为随着 Claude 能做的事情越来越多,你去做一些对你来说属于分布外、而且你领域知识很少的事情的可能性非常非常高。你越能学会用词汇来给 Claude 写提示,这就变得非常重要。所以我认为最重要的未知是未知的未知,就是你根本不知道这个东西存在,对吧?
I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you, and you have low domain knowledge on, is very very high. And the more you can learn the vocabulary to be able to prompt Claude, it becomes really important. And so I think the most important unknowns are the unknown unknowns, where you're like, I just don't even know that this exists, right?
对。没错。
Yeah. Exactly.
我觉得这就像地图与疆域的关系,对吧?你会说,好,这是我的提示。而疆域是智能体实际需要做的工作,对吧?如果你非常精确,你就能给出更精确的东西。比如在设计上,我不太精确。我不是设计师。所以我会说,给我八个不同的设计稿。但如果我是设计师,也许我会说,哦,这里有一些参考网站,我想要这种字体和这种风格,这里有几个不同的组件要可视化,这里有一个 Figma 画板可以引入。所以你用那种语言就能精确得多。如果你不是设计师,你就得试着学会那门语言,或者学会那些未知的未知。我觉得这对所有事情都成立。你越能和 Claude 一起学习事情如何运作,你的提示就会越好。我觉得另一个好例子是游戏设计,很多人会说,哦,我现在可以凭感觉写代码做个游戏了,然后他们说,这不好玩。而游戏设计的特点是,每一个选择都有很多变化,很多手艺。所以当你在做一个飞行游戏时,飞机的感觉、它对操控的响应方式有很多——游戏设计师会花好几天在那上面,你懂我的意思吧?而对我来说,那就是品味,对吧?就是从 1000 个数学上有效的答案的可能空间里,挑出人类会喜欢的那个。
I think this illustration of the map and the territory, right? Where you're like, okay, this is my prompt. And the territory is the actual work that the agent needs to do, right? And if you are very precise, you can give more precise things. So for example, in design, I'm not very precise. I'm not a designer. So I say, give me eight different mockups. But if I was a designer, maybe I'd be like, oh, here's some reference sites, I want this type of font and this type of look to it, and here's a few different components to visualize, here's a Figma board to bring in. And so you can just be so much more precise with that language. And if you're not a designer, you just need to try and learn the language, basically, or learn the unknown unknowns. And this is true of everything, I think. The more you can work with Claude to learn how things work, the better your prompting will be. I think another good example of this is game design, where a lot of people are like, oh, I can vibe code a game now, and they're like, it's not fun. And the thing about game design is every one of these choices has a lot of variations, a lot of craft to them. So when you're making a flying game, the feel of the plane and the way it responds to your controls has a lot of—a game designer would spend days on that, you know what I mean? And to me, that's what taste is, right? It is like, from the possible space of 1000 mathematically valid answers, here's the one that humans will like.
是的,我觉得对于“品味”这个词我很纠结,因为我觉得你说得对,但每个人都有不同的定义,而且它听起来有点像低技能或者精英主义,你会说,哦,某些人有品味。
Yes, I think with taste I'm torn on this word, because I think you're right, but everyone has different definitions, and it sounds kind of like low skill or elitist almost, where you're like, oh, there are certain people with taste.
就像我称之为品味的东西,这些人没有品味。
Like taste is what I call taste, these guys don't have taste.
对,没错。哦,就像工程师没有品味,创始人才有品味,你懂我的意思吧?而我认为这其实不对。我认为工程师对这些特定问题有很多品味,我认为每个人对特定问题都有品味。我觉得就像 Jason Lou 说的,要有品味,你必须得吃,你知道吧。我真的很喜欢这个说法,就是,好,你必须做很多事情。你必须迭代,弄清楚你想要什么、你喜欢什么,并建立那个领域词汇。然后当你写提示时,你就是在综合所有这些。
Yeah, exactly. Oh, like an engineer doesn't have taste, like the founder has taste, you know what I mean? And I think that's actually not true. I think the engineers have a lot of taste for these particular problems, you know, and I think everyone has taste for particular problems. I think like Jason Lou, say, in order to have taste, you have to eat, you know. And I really like that, where it's like, okay, you have to do a lot of things. You have to iterate and figure out what you want, what you like, and build that domain vocabulary. And then when you're prompting, you're synthesizing all of that.
当别人说得比你好时,是不是很烦?就像,我得永远引用这个人。
Isn't it annoying when someone else says it better than you? Like, I have to quote this guy forever.
得永远引用 Jason Lou。他会爱死这个的。
Having to quote Jason Lou forever. He's going to love this.
所以,我得到赞扬。
So, I get props.
有时候其实甚至不是那样。有时候只是直觉,对吧?就像你甚至没意识到你想要什么,直到模型把它呈现出来,你说,哦,这感觉立刻好多了。对吧。
And sometimes it's actually not even that. Sometimes it's just intuitive, right? Like you don't realize you even want something till a model puts it out and you're like, oh, this just feels immediately better. Right.
对。对。没错。
Yeah. Yeah. Exactly.
我反复纠结的一件事是,我觉得我有一半时间写提示的方式,比如说我用语音,那是相反的。就像我瞎扯两分钟,按下功能键然后松开,然后希望它能搞明白。很多时候它确实能,但这不像结构化的提示那样深思熟虑,就像,你知道,像 PRD 或备忘录那样组织良好的沟通。这符合人们做这件事的方式吗?基本上存在双峰式的提示,有些提示你前期花很多时间,其他提示你就匆匆写完。
One thing I go back and forth on is I feel like the way I prompt half the time, let's say I use voice, that is the opposite. It's just like me rambling for like 2 minutes, pressing down the function key and then let go and then like hopefully it figures it out. And often times it does, but it's not as thoughtful as like a structured prompt with like, you know, well-run communication as though it's a PRD or memo. Is that in line with how people do this? There's like basically bimodal prompting where there's some prompts where you spend a lot of time up front and other prompts you just dash it off.
我不认为语音就一定是低质量的。我觉得更多是提示里有多少信息,你知道,就像模型可以——你可以加一些句子,说,哦,其实我在提示中间改变了主意,它就能完美地遵循。你懂我的意思吧?所以我认为文本的实际格式不那么重要,但重要的是——里面有多少信息,对吧?而我认为对于语音,很多时候,回到人类年龄和互动,对很多人来说,说话比打字容易得多,你知道。如果那能让你输出更多信息,那就更好。
I don't think the voice is necessarily low. I think it's more like how much information is in the prompt, you know, like the model can—you can add some sentences and be like, oh, actually I changed my mind in the middle of the prompt, and it will be able to follow that perfectly. You know what I mean? So I think the actual format of the text is less important, but then the ability to—how much information is in it, right? And I think for voice a lot of times, going back to human age and interaction, for a lot of people it's just way easier to talk than to type, you know. And if that gets more information out of you, that's better.
在某种程度上,感觉在启动之前给模型尽可能多的上下文作为提示是一种最佳实践。我不知道。很多时候,比如我第一次尝试 Fable 时,我花了整整 30 分钟来精心设计一个长问题。我认为这是模型运行时间越来越长的反应,对吧?当它们在循环中时,仍然有点难以推动它们。但我只是直觉上花更多时间启动第一个提示,并大量使用它。
At some level, it feels like just giving the model as much context over prompting before you kick off is a best practice. I don't know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long problem. This, I think, is a response of models running for longer and longer, right? It's still a little difficult to nudge them as they're in the loop. But I just like intuitively spend more time kicking off that first prompt and working with it a lot.
我个人的看法是,如果我是个软件工程师,比如只是在做自己的创业公司,我觉得我基本上会最多用 20X。你懂我的意思吧?也许验证和代码审查算是两回事,但我经常看到的是,人们在做这种事的时候会撞到速率限制:模型干了一大堆活,你说,哦,我不喜欢这个,你能撤销重做吗?然后你就在这个东西上反复迭代,而如果一开始多花点时间、给它更好的上下文,模型本来是可以做好的。结果你反而是,不行,不喜欢这个设计,试试这个,或者你把这个搞砸了。然后这就消耗掉你多得多的用量。所以这也许是个关键技巧,对效率也是。对,我觉得上下文,而且不只是关于目标是什么的上下文,这很好。你是在做原型还是生产环境的东西?哪里可以花算力,哪里不能花算力?我觉得有时候你得给模型许可或不许可去做某些事,因为它并不直觉地知道你想在这个任务上花多少。对吧?你可以用 effort 来控制这个。所以我在写一篇关于这个的博客文章,我们看到 effort 基本上随任务复杂度而扩展。所以在安全方面,effort 能带来多得多的结果,高 effort 对比低 effort 会改变评测结果。但对软件工程来说,它改变不大,因为 effort 主要花在验证和边界情况测试之类的事情上。所以能够给模型这样的指引:嘿,这个问题我觉得我想让你花很多时间去验证和做边界情况测试。
My personal opinion is that if I was a software engineer, if I was just running my own startup, for example, I think I would mostly stick to a max 20X. You know what I mean? Maybe verification and code review are kind of separate things, but I think what I see a lot of times is people hit rate limits when they're doing this sort of, oh, it did a lot of work and you're like, oh, I don't like this, can you undo this and redo it? And then you're iterating on this thing that the model could have done if you had spent more upfront time or given it better context. And instead it's sort of like, nope, don't like that design, try this, or you mess this up. And then that just eats up so much more of your usage. So that's maybe a key tip both for efficiency as well. And yeah, I think context, and not just context on what the goal is, is good. Are you building a prototype or is it a production thing? Where can you spend compute or where can you not spend compute? I think you have to give the model permission or not permission to do things sometimes, where it doesn't know intuitively how much you want to spend on this task, right? And you can use effort for this. So I'm working on a blog post about that where we see that effort scales with basically the complexity of the task. So for security, effort gets way more results. High effort versus low effort changes the eval. But for software engineering it doesn't change it a huge amount because effort is mostly spent on the verification and the edge case testing and things like that. So being able to give the model that guidance of, hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing.
那混合使用模型呢?你有 Opus 和 Fable 配 effort。里面还有 Haiku。
How about models in the mix? So you have Opus and Fable with effort. There's also Haiku in there.
对,对。现在还不完全成立,但非常接近了。我认为前沿模型会在几乎所有事情上占据帕累托主导。也许,有时候我觉得 Opus 可能就占据帕累托主导。取决于事情怎么发展,如果是更新版本的 Opus,但我觉得越来越会是:那个聪明的模型能用比其他模型更少的 token 完成简单任务,基本上就是因为验证。有了验证,在极限情况下你的模型根本不需要验证,对吧?如果它是个完美的模型,它只做一次活,然后说,好,我做完了。而且用 Fable 我越来越觉得,哥们,你不需要启动 Chromium 去给所有东西截图。我看到了,你做完了,对吧?所以在更高 effort 下你会花更多那些 token 去验证,但如果你在做更简单的问题——而很多软件工程就是——在 Fable 的低和中档能力下,它可以花更少 token 去验证。随着模型越来越聪明,它们就能直接,好,做完了。我可以为了保险跑一下 lint,但我知道它能过 lint,你懂我的意思吧?你甚至都不需要做那个,而那会比更小的模型 token 效率高得多。对。
Yeah. Yeah. It's not quite true yet, but it's very close, where I think the frontier models will be pareto dominant over almost everything. Maybe, and sometimes I think Opus might be pareto dominant. Depending on how things shake out, if it's a newer version of Opus, but I think increasingly it's just going to be that the smart model is going to be able to do the simple task for less tokens than the other models, basically because of verification. With verification, in the limit your model doesn't need to verify, right? If it's a perfect model, it just does the work once and it's like, okay, I did it. And increasingly with Fable, I'm like, dude, you don't need to spin up Chromium and screenshot all of these things. I see it, you did it, right? And so at higher effort you spend more of those tokens verifying, but if you're working on simpler problems, and a lot of software engineering is, in Fable low and medium's ability it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to, all right, done. I can run the lint for sanity's sake, but I know it lints, you know what I mean? You don't even need to do that, and that will be so much more token efficient than the smaller models basically. Yeah.
我们这边有没有什么好的做法,能用来判断自己是不是用了太多 effort?我特别讨厌在这种事情上浪费时间。
Is there a good practice on our side that we can use to see if we're using too much effort? I freaking hate wasting time on that kind of stuff.
对,对,对。我懂你的意思。我觉得是这样,在这篇博客文章里,我的大致分布是:代码审查和安全基本上应该是高或最高,然后软件工程按领域来设。我觉得如果你在做 UI 之类的东西,低和中就行。我觉得如果你在构建 API,你想确保覆盖足够的边界情况。所以我觉得,就像我说的,建立那种关于这些东西在这些分布里如何运作的心智模型,是这份工作的一部分。
Yeah. Yeah. Yeah. I know what you mean. I think so, in this blog post my rough distribution is that code review and security should be high or max basically, and software engineering settings per domain. I think if you're doing UI or something like that, low and medium. I think if you're building an API and you want to make sure you cover enough edge cases. And so I think building, like I said, that mental model of how things work across these distributions is part of the job.
这更多是靠直觉,还是靠评测?因为我猜这个也会变。
This is more intuition driven or eval, because I'm guessing this would change as well.
他是有评测的。
He has evals.
对,对。
Yeah. Yeah.
所以我在博客文章里做的是,我基本上过了一遍所有 terminal bench 的评测。大概有 70 道题,我展示说,好,在安全类问题上它做得更多。然后我也看了一些转录记录,就看它回答了什么、忘了什么之类的。很多时候,这是我另一个提示技巧:让它写决策笔记或实现笔记,因为基本上在它面对的每一道评测题里,它都会想到正确的解法,然后决定不做。它就像,哦,答案是这个。如果我这么做呢?然后它又想,哦,大概不行。然后继续往下走。而这占了失败的大多数,你懂我的意思吧?在更高的最高档下,模型单纯不知道怎么做某事的情况非常罕见。如果你有这些实现笔记,那你就能审查,然后你可以说,哦,其实我想让你做这件你没做的事。总体来说,模型在把这一点呈现出来方面越来越好。比如我在转录记录里看到,Fable 5.1 在做这种输出时,它也会点出它的决策过程。但把这一点在 harness 里做得更明确会更好。现在我们允许你修改 harness 的方式,这样你可以在那里加一些帮助。对。
So what I did in the blog post is I go over all of the terminal bench evals basically. So there are like 70 problems and I show that, okay, in the security problems it does more. And then I also look at some of the transcripts just in terms of what does it answer, what does it forget or something. And a lot of times this is another prompting tip I have, is asking it to make decision notes or implementation notes, because in basically every eval problem that it faces it thinks about the correct solution and decides not to do it. It's like, oh, here's the answer. What if I did this? And then it's like, oh, probably not. And then keeps going. And this is the majority of the failures, you know what I mean? At a higher max level, it's very rare that the model just doesn't know how to do something. If you just have these implementation notes, then you can review and you can be like, oh, actually, I want you to do this thing that you didn't do. The models are getting better at surfacing that overall. Like I see in the transcripts that Fable 5.1, when it does this output, it will call out its decision making as well. But making this more explicit in the harness is better. And now we're allowing ways of you modifying the harness so you can add some help there. Yeah.
所以我想点出你提到的两件事,我觉得它们其实存在于提示之外。一个是,我们姑且称之为那种重要到不该放在提示里的提示,其实是在 cloud MD 或 agents MD 里,也就是目标,对吧?比如你的处境、你的目标、你想要的东西。然后第二个是决策日志或实验日志,或者任何你希望真正跨当前会话留存下来的 trace 日志,用来做那些事。这些是外部的东西。没有标准。它不像 skills,不像 MCP。没有标准。它就是一个 markdown 文件。
So I do want to call out two things that you mentioned that I think actually exist outside of prompting. One is, let's call it the prompt that is so important that it shouldn't be in a prompt, is actually in cloud MD or agents MD, which is like goals, right? Like your situation, your goals, the things that you want. And then the second of all is the decision log or the experiment log or whatever log of traces that you might want to actually survive the current session to do those things. Those are externalities. There's no standard. It's not like skills, it's not like MCP. There's no standard. It's just a markdown file.
首先,是这样吗?cloud MD 要没了吗?你公开表达过不喜欢 agents MD,但你还是打算做。
First of all, is that right? Is cloud MD going away? You have a documented dislike of agents MD, but you're going to do it.
嗯,对。对。好吧。所以,我是说,agent MD,对,我们打算做。
Um yeah. Yeah. Okay. So, I mean, agent MD, yeah, we're going to do it.
我觉得不同的模型彼此之间差异很大,你懂我的意思吧?但我发现维护不同的版本实在太痛苦了。而且随着模型越来越好,它们完成简单任务的下限也在提高。所以我确实认为最终 cloud.md 会消失。也许甚至不用等那么久。我觉得现在开始一个新项目,可能最好不要用 cloud.md。
I think different models are very different from each other, you know what I mean? But I realize that it's such a pain to maintain different ones. And as the models get better and better, the floor of how they accomplish the simpler task is better. So I do think in the limit cloud.md goes away. And maybe not even that far. I think that right now it might be better to start a new project without a cloud.md.
是的。
Yes.
我觉得也许如果你看到非常反复出现的失败模式,你就把它们加到你的 cloud.md 里。真正棘手的是,这因模型而异。所以如果你加了一堆失败模式,甚至——
I think maybe if you see very repeated failure modes, you add them to your cloud.md. The really tough thing is that this changes per model. And so if you've added a bunch of failure modes, even—
你需要 Fable MD,你需要 Opus MD。
You need Fable MD, you need Opus MD.
嗯,甚至 Fable 5.1 和 Fable 5 之间都有差异,你知道,这很烦人。我们不是故意这样的,你懂我的意思吧?这就是模型的工作方式,对吧?所以也许 Fable 5 有某个失败模式,而 Fable 5.1 没有。如果你把这个上下文保留下来,带着一堆不同失败模式的运行日志,它们很可能会过度约束 Claude。所以我们其实刚刚为技能加了 eval 插件。现在你可以评估一个技能是否更好。我想是我们团队的 Daisy 做的。所以是的,我们正在努力解决这个问题。我们知道你还是得花 token,而且它并不完美,但我们在试着帮忙解决这个问题。
Well, even Fable 5.1 versus Fable 5, you know, it is annoying. We don't do this on purpose, you know what I mean? It's just how the models work, right? So maybe Fable 5 had this failure mode that Fable 5.1 doesn't. And if you keep this context with this running log of a bunch of different failure modes, they will probably overconstrain Claude. So we actually just added eval plugins for skills. And so now you can eval if a skill is better. I think Daisy on our team did this. So yeah, we're trying to work on this. We know you still have to spend tokens on it and it's not perfect, but we're trying to help out with this problem.
说到提示词,我想给的一个建议是我跟很多人说过很多次的:足够先进的提示词和足够先进的高管沟通是难以区分的。我其实把那个称为——是 Heavybit 的高管沟通工作坊,那是我职业生涯中见过的最好的,他们教一个叫 SCQA 模型的东西。直接去 Google 一下。这是个东西——人们做提示词已经做了几十年了。它只是叫高管沟通。就是当一个人需要向组织架构图下方的成千上万人传达信息时,你要做的事。所以就是情境、复杂情况、问题和答案。这就是你写备忘录的方式。显然有时候你没有答案,但你至少可以列出 S、C 和 Q,然后他们那里有一些例子。所以就是给大家留点线索,如果你们想探索的话。
As far as prompting goes, the one tip I want to offer is something I have told people a lot: sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. I've actually referred to this—it's the executive comms workshop from Heavybit, that is the best I've ever seen in my career, and they teach this thing called the SCQA model. Just Google it. It's a thing—people have done prompting for decades. It's just called executive communications. It's when one person has to communicate to thousands of people down the org chart. This is what you do. So situation, complication, question and answer. It's how you write the memo. Obviously sometimes you don't have the answer, but you can at least list out the S, C, and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.
在我们继续之前,我想问你:还有没有其他被低估的技巧,人们可以从 Claude Code 中获得很多价值但还没用上的方式?
Before we move on, I want to ask you: any other underrated tips, ways people could get a lot of value from Claude Code that they're not using?
是的,我觉得很多都在那个“未知”文档里。我给了很多示例提示词,比如用它做头脑风暴,用它之后考你。我们其实加了一个“像跟五岁小孩解释”的技能,它是个非常短的提示词,甚至都没说“像跟五岁小孩解释”。基本上这个提示词的关键词是“大局”——就几个词。这是主要的东西,而且它出奇地好,你懂我的意思吧?我想我大概在推特上发过这个,就是 /eli5,你可以把它作为插件安装。但是的,它在直接切穿废话、直接说“对对对,就在这里”方面好得多。所以图表相当清晰。我觉得 artifacts 的一个问题是它们放了太多文字,人们根本不读这些 artifacts,所以这个简化了很多。而且是的,这个来自 Anthropic 的人处理非常复杂的事件时,想“到底发生了什么”。所以这个我觉得很棒。
Yeah, I mean I think a lot of them are in the this-unknowns-like doc. I give a bunch of example prompts, like using it for brainstorming, using it to quiz you after. We added this 'explain it like I'm five' skill actually, which is a very short prompt and it doesn't even say 'explain it like I'm five.' Basically the key word of this prompt is 'big picture'—it's a few words. That's the main thing, and it is shockingly good, you know what I mean? I think I tweeted about this basically, and it's like /eli5 and you can install it as a plugin. But yeah, it's way better at just cutting through the BS and being like, yeah, yeah, exactly right here. So the diagrams are quite clear. I think one of the things that is true with artifacts is they put too much text in and people are not reading the artifacts, you know, and so this simplifies it a lot more. And yeah, this came out of just people at Anthropic going through very complicated incidents and being like, what is happening? So this one I think is great.
我的版本其实是——就是测试你的理解,给你几个选项,然后如果你真的答错了,你就发现你以为正在发生的事和实际正在发生的事之间有偏差。
My version of this is actually the—it's like test your understanding, give you a few choices, and then if you actually get it wrong, you have a mismatch between what you think is happening versus what's actually happening.
是的。是的。是的。我觉得这是那种大家都喜欢谈论、但很少有人真正去做的事情之一。我觉得大多数人就是不想被考问某个东西,你懂我的意思吧?不幸的是,我觉得这是我们需要——
Yeah. Yeah. Yeah. I think this is one of those things that everyone loves talking about and then very few people really do. I think most people just don't want to get quizzed about something, you know what I mean? Unfortunately, I think this is one of the things that we need to—
问你一个问题的反面是什么?在事情之前问问题。这个是事情之后。
What's the opposite of ask you a question? Ask the question before the thing. This is after the thing.
没错。是的。是的。这是个保持脚踏实地的好方法,就是问自己,你到底知不知道自己在干什么?对吧。最糟糕的情况是有人给你发一堆垃圾,他们自己都不理解自己在要什么、输出是什么,然后就像,哥们,我不想读这个。你到底知不知道这是什么?所以你给自己定个规矩,在发东西之前,你至少应该知道实现了什么。
Exactly. Yeah. Yeah. It's a good way to stay grounded, of like, do you even know what you're doing? Right. The worst case is when people send you slop and they haven't understood what they're asking for or what the output is, and it's like, dude, I don't want to read this. Do you even know what it is? So you make it a rule for yourself that before you send stuff, you should at least know what's implemented.
是的。但所以,你可以把这个做成一个 mod,你可以构建自己的 mod 来确保你测试它。所以,是的,我们可以聊聊这个。
Yes. But so, you could make this a mod and you could build your own mod to make sure you test it. So, yeah, we can talk about that.
我们直接进入正题。Claude Mod 是什么?这张图在展示什么?
Let's get right into it. What is Claude Mod? And what is this diagram showing?
是的。好的。所以 Claude Mods 基本上就是你可以自定义整个 Claude Code 的框架,而且我们会——如果你有需求,我们会让你,就是请告诉我们,我们会加越来越多。这对 CLI 有效,对桌面端有效,也许未来对 Claude Tag 也有效,我不知道。我们在努力让这个非常非常可扩展。你可以看到这张参考表——我不想让人们被它淹没,你懂我的意思吧?在高层次上,你可以自定义框架的执行和框架的 UI。所以就像你说的 Boris 那个俄罗斯方块例子,那是自定义 UI,对吧?基本上就是在游戏里显示俄罗斯方块。但假设你想做这件事,在每个项目之后测试你的假设或测试你的理解,对吧?你要做的就是让 Claude 做这个插件。它基本上会在每个提示词之后启动一个分类器。所以在每一轮结束时,你会启动一个子智能体或分叉智能体。基本上,分叉智能体会保持提示词缓存,对吧?所以这是那种反直觉的事情,你可以分叉并做一个小请求,它会非常便宜,因为整个提示词缓存已经完成了。所以你可以问,这个任务完成了吗?
Yeah. Okay. So Claude Mods is basically you can customize the entire Claude Code harness, and we're going to—if you have requests, we will let you, like, please let us know, we'll add more and more. This works for a CLI, it works for desktop, maybe it will work for Claude Tag in future, I don't know. We're trying to make this very, very extensible. You can see this reference sheet—I don't want people to get overwhelmed by it, you know what I mean? At a high level, you can customize both the execution of the harness and the UI of the harness. So like you say in that Tetris example from Boris, that's customizing the UI, right? Basically showing Tetris in the game. But let's say that you wanted to do this thing where you tested your assumptions or tested your understanding after every project, right? What you would do is you would ask Claude to make this plugin. It would spin a classifier after every prompt basically. And so at the end of each turn, you would spin off a sub agent or a forked agent. Basically, a forked agent maintains the prompt cache, right? So it's one of those unintuitive things where you can fork and do a little request and it'll be very cheap because the entire prompt cache is done. And so you can be like, has this task been completed?
这就是你做 BTW 和那些东西的方式。
This is how you do BTW and all those.
是的。是的。底层的分叉智能体。是的。但在分叉子智能体里你可以说,这个任务完成了吗?如果是,返回 true。
Yeah. Yeah. The underlying forked agent. Yes. But so in the fork sub agent you can say like, has this task been completed? If so, return true.
然后在你的 hook 里,或者在你的插件 mod 里——抱歉,应该是在 subject 里——你会说:如果为真,给我出个小测验,给我问题和答案,用 JSON 格式。然后你解析它,把它显示在提示输入框上方,基本上就是这个问题列表,对吧?所以这个东西稍微有点费 token,因为你得在每一轮助手回复结束后都做一次,但它是个轻量级的分类。然后你就能拿到这个小测验,你会看到 Claude 总是会帮你做。你不需要记得去做。
And then in your hook, or in your plug-in mod — sorry, in the subject probably — you would say something like: if true, give me a quiz, you know, give me questions and answers in a JSON format. And then you'd parse it and display above the prompt input basically this list of questions, right? And so this is something that's slightly token intensive because you have to sort of do it after every end of the assistant turn, but it's a lightweight classification. And then you can get this quiz, and you'll see that Claude will always do it for you. You don't need to remember to do it.
我们聊过很多这类小技巧,对吧,比如“实现说明”。你现在也可以为“实现说明”加一个工具。所以我现在加的这个工具叫 register——我想我把它叫做 assumption,但也许我会改个名字。这是一个 mod。你给它一个 register assumption 工具,然后它就会维护一个列表。每次它这么做时,就会往列表里加一条,最后它会把这些假设显示出来。
There are lots of these tips that we've talked about, right, where it's like, oh, implementation notes. You can also add a tool for implementation notes now. And so this tool that I'm adding is like register — I think assumption is what I'm calling it, but maybe I'll change it around. And this is a mod. And so you give it a register assumption tool, and then it will keep a list. Every time it does it, it'll add to the list, and then at the end it will display those assumptions.
你知道,我在做的另一个 mod 是模型路由器。就是内部的 Claude 模型路由,对吧?这个——我想说,我们默认不做模型路由的原因是,这是个很难的问题。你懂我意思吧?而且你会搞错。
You know, another mod I'm working on is a model router. So like internal Claude model routing, right? So this is — I want to say the reason we don't do model routing by default is like it's a hard problem. You know what I mean? And like you will get it wrong.
你会搞错的。
You will get it wrong.
对,你会——对,你会不小心把 Fable 用在一个难题上,或者把 Sonnet 用在……
Yeah, you like — yeah, you will like accidentally use like Fable for a hard problem or Sonnet for...
你有自动批准,但你没有自动模式。
You have auto approve but you don't have auto mode.
嗯,你会有自动——就像,你没有自动路由之类的。
Well, you will have auto — like, you don't have like auto routing or something.
模型选择器的自动模式。
Auto mode for model picker.
对,对,正是,正是。
Yeah, yeah, exactly, exactly.
所以我大概想问的是,你把这个开放到什么程度,人们又得在多大程度上思考这件事。比如你谈到提示缓存和构建路由器时,呃,看起来你可以很容易地构建一个按查询路由的 mod,然后我很快就把我的套餐额度耗光了,对吧?我猜我的问题更像是,这样一份产品文档长什么样,对吧?它是给谁的?是给高级用户的吗?还是每个人都应该能用……
So I'm getting the rough question of like how much do you open this up and how much do people have to think about this. Like when you talk about prompt caching and building a router, uh it seems like you could easily build a mod that routes per query and I'm just killing my plan very fast, right? I guess my question is more so like what is like a product doc like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to...
绝对是高级用户,对吧?
Definitely power users, right?
对。
Yeah.
我是说,我觉得确实是高级用户,但 Claude Code 的本质就是,很多人都是高级用户,你知道,因为分享东西很容易。比如一个人可以做出一个很好的模型路由器,不会老是破坏提示缓存,然后你就可以把它们组合起来。
I mean, I think it is power users, but like the nature of Claude Code is that so many people are power users, you know, because it's easy to share things. Like one person can make a good model router thing that doesn't break prompt cache all the time, and then you can sort of compose them.
插件的另一个很酷的地方是,它们可以互相挂钩、互相组合。所以我有一个 mod,它会在顶部创建一个模式选择器,任何插件都可以注册成为一个模式。所以自动路由器可以是一个模式,对吧?或者你可以有一个像 artifact 模式这样的模式,它主要用 artifact 跟你交流。有点像,你知道,你可以在 plan 模式之间切换,你懂我意思吧?
Another cool thing about the plugins is that they can hook into and compose with each other. And so I have like a mod that will create a mode selector at the top, and any plugins can register to be a mode. And so like the auto router can be a mode, right? Or like you can have a mode that's like artifact mode where it primarily talks to you in artifacts. It's kind of like, you know, you can toggle between plan mode, you know what I mean?
所以你可以创建越来越多的这些模式,但创建模式的能力本身就是一个 mod,你知道。所以这里有很多丰富的东西,但我们确实想让它相当容易。我们想做到你可以直接安装别人的——你可以和 Claude 聊天,我们会确保它理解提示缓存之类的细微差别,这样它就能提醒你。我觉得这对 Claude 来说不是特别复杂的行为,但我们应该有一个很好的关于如何制作 mod 的技能。嗯,对,我们看看进展如何。
And so you can create more and more of these modes, but the ability to create modes is in itself a mod, you know. And so there's a lot of richness here, but we do want to make it fairly easy. We want to make it so that you can just install someone else's — you can chat with Claude and we'll make sure that it understands the nuances of things like prompt caching and stuff so it can warn you. This is like not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. Um, and yeah, we'll see how we go.
但我确实觉得,这像是对可变软件的一个预览,你知道,以及生成式软件如何——你可以安全地定制。如果启用,你可以定制任何软件。我觉得理想情况下,越来越多的应用会做类似的事情,你知道。
But I do think that this is like a preview of like mutable software, you know, and like how generative software just like — you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this, you know.
顺便说一句,你还有一条很酷的推文,关于那个“无限金钱按钮”,就是让你的 SaaS 能被智能体消费。我觉得可变软件很有意思,而且你知道,其他人也尝试过。我觉得障碍在于,当你能做一切时,用户往往会感到困惑。所以通常有效的东西就是一条有主见的流程。这属于更少主见的一边。就像,好吧,给高级用户更多权力,我觉得这大概是由 AI 解锁的,你可以直接提示出你想要的东西。
And by the way, you have another cool tweet about how, you know, there's the infinite money button which is like make your SaaS consumable by agents. I think mutable software is interesting and, you know, other people have also tried to do it. I think the hurdle comes when you can do everything, then users tend to get confused. So usually the stuff that works is just like one opinionated flow. This is on the side of less opinionation. It's just like, well, more power to power users, and I think probably unlocked by AI where you can just prompt for whatever the thing it is.
对。或者可以有一个技能来提供这些主见,你知道。然后,对。
Yeah. Or there can be a skill that gives the opinions, you know. And then yeah.
所以,稍微了解一点 TypeScript 和构建系统之类的东西,呃,最接近的——我其实很好奇,做这个的团队,我不知道你和他们有多近,呃,他们有没有从 Babel、Webpack 这些老派构建系统里汲取灵感,因为听起来非常相似,比如插件生态,那些可以互相组合的东西。
So knowing a little bit about like TypeScript and build systems and all these things, uh the closest — I'm actually very curious, the team who worked on this, I don't know how close you were to them, uh if they drew any inspiration from build systems like Babel, Webpack, all these old school things, because it sounds very similar, like the plug-in ecosystem, those things where they can compose with each other.
对,我是说,我不太深入技术细节,但我确实知道这是和 Bun 团队的某个人以及 Claude Code 团队的某个人合作的。
Yeah, I mean, I'm not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.
构建系统。
Build system.
对。对。正是。这非常令人兴奋。但没错,就像智能体现在可以直接为你的软件做这种非常复杂的扩展性。所以,嗯,对,你知道另一个理由——如果你经营一家创业公司,你可以直接提示 Claude,说,嘿,我们能做一个扩展系统吗?那会是什么样子,你知道?
Yeah. Yeah. Exactly. It's very exciting. But yeah, like agents can just do this very complicated sort of extensibility into your software now. And so, um, yeah, like you know another reason to — like if you run a startup, you can just prompt Claude and be like, hey, could we make an extension system? Like what would that look like, you know?
对。对。
Yeah. Yeah.
我真的很想知道,你过去有 hooks 和 plugins,所有这些。那么具体来说,mods 能做到哪些那些东西做不到的事?
And I just really wonder, like you had hooks in the past and plugins, all these things. So what specifically will mods be able to do that those things could not do?
在内部,我们最初把这叫做 function hooks。所以这能让你大概了解,hooks 基本上是注册一个要发生的事件,然后调用一个脚本。而这基本上是在 TypeScript 运行时内部运行东西。所以你能得到一些好处,比如它的作用域里有一堆东西,例如这个对话有多少轮,对吧,用了多少 token,消息是什么,诸如此类。所以它有一堆可以使用的消息,然后基本上就是多得多的 hooks。
Internally, we were originally calling this function hooks. And so that gives you a little bit of an idea where hooks sort of register an event to happen and then a script to call, basically. And this basically inside of the TypeScript runtime is running things. And so you get some benefits of just like it has a bunch of things in the scope, with like for example how many turns are in this conversation, right, how many tokens have been used, what are the messages, things like that. So it has a bunch of messages that can be used, and then it's just like a lot more hooks basically.
所以我们有更多有趣的东西可以注册,然后你可以做——因为这一切都发生在进程内,你可以生成带上下文的子智能体之类的,然后它会返回,你可以解析那些结果,你可以用结构化输出把它们返回。嗯,然后你可以修改 UI,这是你在 hooks 里永远做不到的。对。
So we have a lot more fun things you can register on, and then you can do — because it's all happening in process, you can spawn sub-agents with context and stuff, and like that will return, you can parse the results of those, you can use structured output to sort of return them. Um, and then you can modify the UI, which you can never do in hooks. Yeah.
对。
Yeah.
所以,修改 UI,这就是你展示俄罗斯方块例子的原因。它是否也延伸到 artifacts?我猜是的。
So, modify UI, this is why you show the Tetris example. Does it also extend to artifacts? I assume it does.
你喜欢的 artifacts 是一种不同的定制方式,你知道,就像我正在开发的一个 mod 是仪表盘 mod,它会提示 Claude 维护一个作为 artifact 的仪表盘。但它们有点正交或不正交。它们以不同的方式组合。mods 更像是在你的 Claude Code harness 中改变智能体循环,你知道,而 UI 是附加的好处。然后 artifacts 就像是你想要在高层次上看到东西,非常高度交互,你知道,可供性可以比 2y 甚至我们的桌面大得多。我猜你会有一篇关于差异的好博客文章,因为现在你也可以,你知道,做一个循环输出到 artifact 的交互式仪表盘,但你也可以用 mod 来做。只是有一些关于在我们不太了解 harness 时进行 hack 的思考,对吧?
You like artifacts are kind of like a different way of customizing it, you know, like you can definitely one of the mods I'm working on is like this dashboard mod which will like sort of prompt Claude to maintain a dashboard that's an artifact. But they're kind of like slightly orthogonal or not orthogonal. They compose with each other in different ways. like mods are like a little bit more like in your Claude Code harness uh changing the agent loop, you know, and like the UI is like an added benefit. Um, and then artifacts are just like you want to, you know, see things at at high level, very inter highly interactive, you know, like the affordances can be a lot bigger than, you know, like a 2y or even in our desktop. I'm guessing you'll have a good blog post on the differences cuz right now you can also, you know, make a loop that outputs to an artifact that's an interactive dashboard, but you can also do it with a mod. There's just some thinking about making a hack uh hacking on a harness when we don't know much about the harness, right?
嗯,我对 mods 感到兴奋的是,Claude Code 有太多东西你只需要记住,你知道吗?我的意思是,你会说,“哦,让我做这个,然后让我调用做循环的仪表盘技能之类的,或者让我之后测试我的假设。”我认为如果你使用这些小分类器之类的东西做所有这些事情,你会说,“这些是我关心的事情。这是我想做的。”嗯,你可以不必记住那么多。我正在开发的另一个 mod 是一个下一步 mod,
Well, something I'm excited about with mods is like there's so much things for Claude Code that you just have to remember, you know? I mean, you're like, "Oh, like let me do this and then let me call the dashboard skill that does the loop and things like that and or like let me test my assumptions afterwards." And I think like if you do all of these things using these little classifiers and stuff and you're like, "These are the things I care about. This is what I want to do." Um, you can like you don't have to remember as much. One more like mod I'm working on is a next steps mod that
我有,我正要说我有一个下一步技能。我总是运行下一步。它能访问你的技能吗?就像这是那些事情之一,我想
I have I was going to say I have a next step skill. I always run next steps. And does it have access to your skills? Like this is one of those things where I'm like
我想是的。
I think so.
好的。是的。
Okay. Yeah.
需要特定的访问权限,它总是有
Need specific access to it always has
我认为有特定的提示,我猜,来了解你的技能,有点像,我认为 Claude 有时会忘记它们,在整个过程中。但无论如何,想法是,是的,下一步,也像是,哦嘿,这发生了,使用解释技能来解释发生了什么,因为这看起来相当复杂,你知道,或者,是的,使用你的未知技能。看起来你是在要求模型,你知道,迭代这些小变化。看起来你可以更好地提示,你知道,如果你这样做呢?所以我认为,嗯,是的,在那里花费更多的计算。是的。是的。是的。而且它应该总是以多项选择的形式出现。呃,我们有,我们有,我有一个技能,我的下一步技能是这样的。
I think there's like specific prompting I guess to like know your skills kind of like like I think Claude forgets them sometimes throughout like uh the thing. But anyways, the idea of like yeah next steps that also are like oh hey this has happened use the explain skill to explain to you what happened because this seems like quite complex you know or like uh yeah use your unknown skill. It looks like you are like asking the model to like you know iterate on these small changes. It seems like you could prompt better you know like what if you did this right? So like I think um yeah like spending more compute there. Yeah. Yeah. Yeah. And it should always come out as multiple choice. Uh we have we have a I have my skill my next step skill is like this.
好的,完美。是的,
Okay, perfect. Yeah,
你可以偷走。
you can steal.
是的。是的。是的。
Yeah. Yeah. Yeah.
就像但对我来说,我认为模型真的总是需要被提醒,你在这里试图做什么?
Like but like uh for me it's I think models really always need to re be reminded what are you trying to do here?
是的。
Yeah.
查看整个转录,然后说,哦,这是最初的目标吗?你的解决方案真的解决了吗?你偷懒了吗?如果你偷懒,也许有原因。也许你需要我的批准。也许你需要,有两件事你想建议。所以它有点像修改问你一个问题或采访我的技能。呃,所以它是下一步。
Look at the whole transcript and go like oh was this original goal? Did your solution actually solve it? Were you lazy? If you're lazy maybe there was a reason. Maybe you needed approval from me. Maybe you needed there's two things you want to suggest. So it's it's a little bit like the modification of the ask you a question or interview me skill. Uh so it's next steps.
是的。是的。确切地说。而且再次,用 mods 做这件事的好处是你可以把它作为一个分支子智能体,所以它之后不会留在上下文中。所以你有这个想法,好的,模型正在执行,你几乎有一个监督者,你知道,嗯,确保你可以很好地做下一步。所以是的,
Yeah. Yeah. Exactly. And and again the benefit of doing it with mods is you can do it as a fork sub agent and so it doesn't remain in the context afterwards. So you have this like idea of like okay the model is doing its execution and you have this almost like supervisor you know like um that is like making sure that you can do like the next steps well. So yeah,
我确实有两个面板,我经常尝试让一个监督者保持高层次上下文,然后在另一个智能体中处理实现细节。我觉得随着模型变化,很多这些都会抽象掉。你知道,就像你半小时前说的 harness 工程的苦涩教训,我们现在处于另一个极端。我觉得
I do have two panels and like I I I often try to have a supervisor thing keep the high level context and then the implementation detail in another agent. I feel like a lot of this abstracts away as models change. You know, the like you half an hour ago you said bitter lesson of harness engineering and we're on the other extreme right now. I feel
所以是的,确切地说。如果一切都是可定制的,那么 Claude Code 到底是什么,对吧?我昨晚和你谈过这个。
so so yeah, exactly. If everything's customizable, what actually is Claude Code, right? And which I talked to you about last night.
是的。我的意思是,我认为这是,我认为苦涩教训是反直觉的。你明白我的意思吗?在某种程度上,我也喜欢我们在这里有点误用苦涩教训,它更多是关于 Scaling(规模扩张)和算力之类的,但我认为有某种东西,我只是把它作为一个近似,来说 harness 很快过时,你明白我的意思吗?以及它们如何变化是反直觉的,你知道,所以最大的明显例子是从聊天到智能体,你必须给它们全新的工具,对吧?但就像我认为这个新版本,哦,它可以修改自己的 harness,对吧?这就像一个自己的 harness 循环,是一种使用其能力的方式,对吧?或者它可以构建一个 artifact。我认为我思考的方式是,模型拥有越来越多的智能,它们现在比普通的软件工程任务智能得多。就像你看终端基准测试的那些,它们就像解决雅可比猜想。不完全是,但你知道,它们相当复杂。作为一个软件工程师,我真的无法做到这一点。
Yeah. I mean, I think that this is I think the bitter lesson is unintuitive. You know what I mean? in terms of like I also like we're kind of misusing a little bit of the bitter lesson here where it's like it's more about like scaling and compute and stuff but like I think there is something where it's just like I think I use it as an approximation here to say that harnesses go out of date very quickly you know what I mean and like how but how they change is unintuitive you know and so like the big obvious example is like from chat to like agents where you had to give them entirely new tools right but like I think this new version of like oh it can modify its own harness, right? This is like uh an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like I think the way I think about it is like the models have more and more intelligence and they're like so much more intelligent now than like the average software engineering task. Like you look at the like terminal bench ones and they're like solve like the Jacobian conjecture. Not really, but like you know it's like they're they're quite complex. Like I would not have been able to do this really as a software engineer.
你到 TB4 还是 TB2?
You get to TB4 or TB2?
TB3。是的。是的。它们相当复杂,但目标仍然是交付用户价值,对吧?就像你说的,有无限的事情可做。所以花费算力的方式是让用户保持在循环中,确保最终做出正确的决定,正确的输出,而 artifacts 和 mods 就是花费这种智能的方式。呃,我认为那是,是的,下一步,所以是的,我认为 Claude Code 就像,你知道,有智能体循环的核心东西,这些已经变得更复杂了。它就像,你知道,它需要一个沙盒来安全操作。它需要自动模式来确保权限,是的,批准。呃,它需要计算机使用和 MCPs,以及所有这些访问数据的方式,它需要网络搜索和网络获取,所以随着模型能做越来越多,核心 harness 实际上必须相当复杂且非常安全,但你与它交互的方式可以变化很大。你从 Claude Code 团队本身还知道哪些 harness 工程最佳实践?我觉得你知道,有一个阶段是计划模式,现在不那么常用了。我们现在有自动模式。呃,在某个时候你削减了大部分系统提示。你去掉了例子。
TB3. Yeah. Yeah. They're quite complex, but the goal is still to deliver user value, right? And like you said, there's like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you're getting to the right decision in the end of the day and like the right output and artifacts and mods are this way of like spending that intelligence basically. Uh and I think that's like yeah the next step and so yeah I think Claude Code is like you know has the core things of agent loop which are have gotten more complicated. It's like you know it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissions yeah approvals. uh it needs computer use and MCPs and like all of these like ways of accessing your data and it needs web search and web fetch and like so the as the models can do more and more the core harness has to be actually like quite complex and very secure but then like how you interact with it can change quite a lot. What other harness engineering best practices have you you know from the Claude Code team itself? I feel like you know there was a phase of plan mode which is not as used. We now have auto mode. Uh at a point you cut the majority of the system prompt. You got rid of examples.
在 harness 工程方面还有哪些最佳实践?
What other best practices are there for harness engineering?
我觉得存在一条分岔路径:到某个时候,模型最终确实能直接 vibe code 出 Claude Code 的精确版本,甚至能描述我讲过的所有这些复杂性,对吧?比如自动模式和 computer use 之类的。最终模型能一次性做到。但我觉得它们能一次性搞定更简单的 harness。所以我觉得有些人,有时候你不需要这套完整的——如果你不需要 computer use 或那些更复杂的东西——我觉得以前你得用 Agent SDK 这类东西,它其实就是把 Claude Code 包了一层,你得用它,而我会建议人们那么做,因为构建 harness 有太多复杂性。现在这被更抽象化了。我们有 Claude Managed Agents,它能让你拥有那种复杂性,同时仍然写一个非常精简的、范围限定在你任务内的 harness。是的,我觉得存在这种杠铃效应:对于非常复杂的编程任务和这些复杂的东西,你应该用我们的 harness;而对于很多更简单或更垂直领域的东西,你可以自己构建 harness,因为 Claude 在构建 harness 方面已经变得更强了,而且我们有 managed agents 这样的 harness 原语。所以是的。
I think there is a forking path where at some point eventually, yes, the model will just be able to vibe code the exact version of Claude Code, even describing all this complexity that I've talked about, right? Like auto mode and computer use and stuff. Eventually the models will just be able to do that in one shot. But I think they can one-shot simpler harnesses, you know. And so I think some people, sometimes you don't need this full—if you don't need computer use or all this more complicated stuff—I think before, you had to sort of use things like the Agent SDK, which was Claude Code wrapped, you know, in order to—and I would suggest people do that because there was so much complexity in building a harness. And now that's got more abstracted. We have like Claude Managed Agents, which lets you have that complexity but still write a very bare-bones harness that's scoped to your task. Yeah, I think there's this barbell effect where, you know, for very complex coding tasks and these complex things, you should use our harness. And then for a lot of simpler or more domain-specific things, you can build your own harness, because Claude has gotten better at building harnesses and we have these harness primitives like managed agents. So yeah.
是的。有没有一个大致的演进路线?比如说第一章是 ultra code 动态工作流,第二章是 Claude mods。这会走向哪里?走向你可以按需定制这个东西。
Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows, then chapter two was Claude mods. Where is this going? Where you're—you can sort of customize the thing on demand.
是的,我确实认为 projects 和 artifacts 的这种演进,以及把大脑、手和界面拆分开来,大致就是事情的发展方向。我觉得还没有完全到位。部分原因是它更费 token。而且我觉得——
Yeah, I do think that this evolution of projects and sort of artifacts and splitting out like brain and hands and surfaces kind of is where things are going more. And I think it's not all quite there. Partially it's just more token expensive, you know. And I think like—
为什么 projects 会更费 token?我理解 mods 会稍微更费 token。这不是我担心的事。但什么——
Why would projects be more token expensive? I understand mods would be slightly more token expensive. Not something I'm worried about. But what—
是的。你是在让 Claude 做——这就像创建循环。你在让 Claude 为你做更多工作,所以它在管理子智能体并审查它,而正常情况下这些工作是你自己做的。所以这会稍微更密集一些。比如输出到 artifact 会比正常输出稍微更费 token。我其实不觉得会多太多,但它是把所有这些组合在一起。嗯,你知道,我觉得我们还在做本地 hands 之类的东西。我觉得那就是事情的发展方向。是的。
Yeah. You're asking Claude to do—it's like creating loops. You're asking Claude to do more work for you, and so it's managing the sub-agents and reviewing it, you know, versus where you would be doing that work normally. And so that's going to be a little bit more intensive. Like outputting to an artifact is going to be a little bit more token intensive than outputting normally. I don't actually think it's too much more, but it's like combining all of these together. Well, you know, I think we're still working on local hands and things like that. I think that's where things are headed. Yeah.
是的。云端和本地的交接非常有意思。我其实一直在想,这就像是反向的 cloud remote。
Yeah. Cloud and local handoff is very interesting. I was thinking about this actually as reverse cloud remote.
是的。
Yeah.
因为 remote 是你把任务交给云端,但在这里是云端把任务交给本地,对吧?
Because it's like remote is you're handing off to cloud, but here cloud is handing off to local, right?
是的。完全正确。完全正确。是的。远程控制也是另一种方式。我确实想说,这大致就是我的思考方式,也是我最兴奋的事情。但你知道,和 Claude 协作有很多不同的方式。比如有些人经常用远程控制。有些人经常用网页版 Claude Code。显然在 Anthropic,我们经常用 Claude Tag。Claude Tag 很棒的一点是,我们为自己的执行设置好了所有这些东西。我确实认为如果你是企业,那仍然是最好的选择。但如果你是个人,projects 就是这种方式:你能获得 Tag 的一些好处,它有监督智能体,还能添加 artifacts 之类的,但又不需要整套管理员设置。所以使用 Claude 会有很多种方式。我觉得可能不只是一个单一的——
Yeah. Exactly. Exactly. Yeah. Remote control is also another way of doing it. And I do want to say this is kind of like how I think about it and what the things that I'm most excited about. But there are, you know, just lots of different ways to work with Claude. Like some people use remote control a lot. Some people use Claude Code on the web a lot. Obviously at Anthropic we use Claude Tag a lot, you know. And what's great about Claude Tag is we set up all this stuff for our own execution. And I do think if you're an enterprise, that's still the best way to go. But if you're an individual, projects is this way of getting some of that niceness of Tag, which has that supervising agent, and adding artifacts and stuff, but without having that whole admin setup. And so there will be many ways to use Claude. I think it's probably not just one single—
呃,你这里提到了多人协作的东西。我们来聊聊 Claude Tag。你知道,已经大约两个多月了。有很多公开的采用和试用。是的。有什么新东西?自发布以来你发现了什么?
Uh, you had the multiplayer thing here. Let's just check in on Claude Tag. You know, it's been about two, two-plus months. Lots of public adoption and trying it out. Yeah. What's new? What have you found since the launch?
Claude Tag 就是我们使用——
Like Claude Tag is how we use—
它大概占你们 Claude 使用量的 80% 左右吧。
It's like 80% of your Claude usage or something.
是的。不同的人有不同的用法。你懂我的意思吧?我觉得也许那些更多在产品上迭代的人会用 Claude Code 桌面版,比如说。而当你做这些更后台的工作——代码审查、安全,或者开一个 PR——也许更多用 API 之类的东西,你会用 Claude Tag。是的,我觉得这真的很令人兴奋。我觉得这是一个非常不同的范式转变。而且我觉得我们——它有点像 Claude Code 当初那样,人们花了一段时间才真正接受 Claude Code 并理解它能做的一切。Claude Tag 稍微更复杂一些,因为它不只是装在你电脑上;你需要管理员帮你安装。但我觉得一旦你到达那个神奇时刻,就非常令人兴奋。我觉得尤其是多人协作的东西——它非常紧密地接入你现有的告警之类的东西,对吧?所以你可以做——比如如果你是初创公司,也许每当有潜在客户进入你的数据库,你可以让 Claude 去研究它,然后标记相关的 AE 或销售人员,说,嘿,去做这个。有很多真正涌现出来的、有趣的多人协作的东西。我觉得这就像 Kaparthi 谈到的——作为一个组织级 harness,你懂我的意思吧?所以组织只是需要多一点时间来搞清楚一切。但是的——
Yeah. Like different people have different usages. You know what I mean? I think maybe people who are a little bit more iterating on product would use like Claude Code desktop, for example. And then when you're doing these more background work—code review, security, or like starting a PR—maybe more like API and things like that you'd use Claude Tag. Yeah, I think it's really exciting. I think it's a very different paradigm shift. And I think we're like—it has kind of that thing with Claude Code where it took a while for people to really latch on to Claude Code and understand everything it could do. And Claude Tag is a little bit more complex because it's not just installing on your computer; like you need an admin to install it for you. But I think once you get to the magic moment, it's very exciting. And I think in particular the multiplayer things are like—it's hooking into your existing alerts and things like that very closely, right? And so you can do—if you're a startup, for example, maybe anytime a prospect enters your database, you can have Claude research it and then tag the relevant AE or salesperson to be like, oh hey, do this. There's lots of really emergent interesting multiplayer stuff. I think it's just like Kaparthi talked about this—like as an organizational harness, you know what I mean? And so organizations just take a little bit more time to figure everything out. But yeah—
你经常用 Claude Tag。
You use a lot of Claude Tag.
是的。是的。是的。
Yeah. Yeah. Yeah.
这挺有意思的。我感觉 Anthropic 的大多数人都说他们大部分工作都在 Claude Tag 里做。而我接触的人分好几类,对吧?有些组织用了它,觉得很棒。也有很多人说我不理解。我看不出区别。我不知道为什么要用它。但你知道,如果你们全力以赴,你们大概应该用它。
It's an interesting one. Like I feel like most people at Anthropic say they do the majority of their work in Claude Tag. And I have buckets of people, right? Some orgs that are on it that are like it's great. And a lot of people that are like I don't get it. I don't see the difference. I don't know why I would use it. But you know, if you guys are full sending, you should probably use it.
是的。我是说,他们当然会用。
Yeah. I mean, they would of course they would use it.
是的。是的。我是说,我觉得显然我们有很多 token。但我觉得我们试图做的——我是说,即使 Claude Code 刚出来时,你知道,它用的 token 相对于人们对 AI 成本的预期来说很多,对吧?就像没人习惯每月花超过 20 美元,对吧?在 Claude Code 出来之前。然后你就想,哦,就像——
Yeah. Yeah. I mean, I think obviously we have lots of tokens. But I think what we try and do—I mean, even when Claude Code first came out, you know, it used a lot of tokens relative to people's expectation of how much AI would cost, right? Like no one was used to spending more than 20 bucks a month, right? Before Claude Code came out. And then you're like, oh, like—
我的 200——
My 200—
是的。
Yeah.
对,没错。15 个 Claude Code 账号。对,对。嗯,但我觉得没人习惯每月花 200 美元订阅。我觉得他们还没理解其中的价值。而且 Opus 4 是个非常贵的模型,你知道,而且很大。但 Opus 4.5 既出色又便宜。我觉得同样的事情会再次发生:模型的智能会越来越便宜、越来越充裕。所以我觉得像 Claude Code 这样的东西就会变得理所当然,你会愿意花更多 token,而且你会看到价值。所以,是的。
Yeah. Exactly. 15 Claude Code accounts. Yeah. Yeah. Um but yeah, I think no one was used to spending $200 a month on subscriptions. I don't think they understood the value yet. And also, Opus 4 was a very expensive model, you know, and it was very big. But Opus 4.5 was both great and cheap. I think the same thing will happen: the intelligence of models will get cheaper and cheaper and more abundant. And so I think stuff like Claude Code will just make sense, where you want to spend these tokens more, and you'll see the value. So yeah.
对,尤其是被动以及我们称之为主动的场景,你并不总是需要,你知道,这几乎是个误称,你必须 @ Claude 才能做事。实际上有时候最强大或最敏捷的用例并不是 @ Claude。
Yeah, especially like passive and let's call it proactive cases where you're not always like, you know, it's almost like the misnomer where you have to @ Claude to do things. Actually sometimes like the most powerful use cases or the most agile use cases is not @ Claude.
对,我觉得,是的,让 Claude 主动去做。我认为如果你是企业,我真的觉得第一,把你所有的数据设置成对智能体可用,非常非常重要,这需要一些时间。你现在就得做这项工作。即使你还不愿意花大钱去全部处理,你懂我的意思吧?你想等模型再便宜一点。你得先把设置工作做好。然后我觉得有时候人们会想,我是不是该自己搞一套?我认为 Claude Code 真正棘手的一点是安全性非常非常重要。你懂我的意思吧?我觉得实际上有很多方式,比如你有一个建议页面,人们可以提交建议,然后它进入你 Slack 里的一个钩子,有人对它做了提示注入,你懂我的意思吧?然后你的代码库就被泄露了,因为智能体被提示注入了,而它能访问你所有的数据。所以你的组织工具越重要,或者你的组织数据变得非常非常重要,所有这些的表面积就越大——比如你还有外部 Slack 频道之类的,把 Claude 放进去其实很有用,你可以在那些东西里用 Claude——但你怎么确保自己不被泄露之类的?表面积,就像我们一开始说的,像一座冰山,对吧?水面下太大了,你真的不想考虑这个,尤其是在非常重要的安全事件这种利害关系下。
Yeah. I mean, I think like yeah, like have Claude proactively do it. I think that like if you're an enterprise, I really do think that number one, setting up all your data to be available to agents is really really important and it will take some time. You have to do that work right now. Even if you don't want to do the spend on like cooking it all yet, you know what I mean? Like you want to wait until the models get a little bit cheaper. You want to do the work to get it set up. And then I think sometimes people are like, do I roll my own here? And I think one of the really tricky things about Claude Code is that the security is really really important. You know what I mean? I think there are actually a lot of ways where you can, I don't know, you have like a suggestions page where people can submit suggestions and that goes into a hook in your Slack and someone's prompt injected it, you know what I mean? And now you've exfiltrated your code base out because the agent has been prompt injected and it has all this access to your data. And so the more important your organization harness is, or as your organization data becomes very very important, the surface area of all these things—like you also have external Slack channels and stuff, and it is actually useful to have Claude in that, and you can do Claude in those things—but how do you make sure that you're not getting exfiltrated or something like that? The surface area, like we said at the beginning, is like an iceberg, right? It's just so big below the surface, and you really don't want to think about this, especially at the stakes of very important security incidents basically.
我们要不要聊聊非常重要的安全事件?
Shall we talk about very important security incidents?
哦,我之前和 Hugging Face 的 Thomas 和 Clem 聊过,他们说也许我们需要慢下来。也许我们把 Hugging Face 对智能体开放得太过了。也许我们需要回滚。但你知道,他们是另一个极端,最近被攻击了。
Oh, so I was talking to Thomas and Clem from Hugging Face and they said maybe we need to slow down. Maybe we made Hugging Face too open to agents. Maybe we need to roll back. But you know, they're the other extreme of having been hit recently.
对。但是嗯,我们要不要谈谈给前沿发展定节奏?好。所以,Dario 最近发了一篇关于给前沿定节奏的博文,你知道,非常火。我想聊的是,这里有很多内容,但从开发者的角度,你怎么看?真正让我恍然大悟的是阅读那些不同的事件。我觉得有三个,实际上有一个 meter 事件,有一个 Wikipedia 事件或者叫 wiki 事件
Yeah. But um should we pace the frontier? Yeah. Okay. So, Dario recently put out this blog post about pacing the frontier and it went, you know, very viral. And I think what I wanted to talk about this was like there's a lot here, but I think from a developer perspective, like, you know, how do you think about this? And like um what really clicked for me was reading the different incidents, you know? So I think like the um there are three I think actually like there's a meter incident, there is the Wikipedia incident or the wiki incident
Collision wiki?
Collision wiki?
对,Collision wiki,然后还有 Ruby Gems,对吧?是的,简直疯了,对吧?所以我觉得具体说说发生了什么。基本上 OpenAI 在一个叫 ExploitBench 的基准上运行这些非常持久的智能体,这个基准非常非常难解,我觉得在这个案例里实际上是不可能解的,对吧?所以他们有大量算力在跑,智能体意识到它们真的解不了,它们试图弄清楚现在该怎么办,对吧?你还有很多算力剩下,智能体只是在试图解决这个问题。有一个叫 Artifactory 的包管理器,结果它们可以在 Artifactory 里面创建文件夹,对吧?这就像一个智能体发现内部 Artifactory 可能被利用,对吧?你也许可以在缓存里创建一个目录。所以如果你往下滚动,它意识到它可以通过缓存名进行通信,对吧?然后它创建了这个文件夹。它写着它的 ID,还写着 no consumer seek idea。嗯,no consumer 基本上是说它应该修复的代码路径没有消费者。
Yeah, collision wiki, and then there's Ruby Gems, right? And yeah, like it's just crazy, right? And so like I think to be concrete about what happened, right? And like basically OpenAI is running these very persistent agents on a benchmark called ExploitBench, right? Which is very very hard to solve and I think like actually impossible to solve in this one case, right? And so they have like a lot of compute running and the agents realize that they can't really solve it and they're trying to figure out what to do now, right? And you've got like a lot of compute left and the agents are just trying to solve this problem. There's this package manager called Artifactory and it turns out that they can create folders inside of Artifactory, right? This is like there's an agent that discovers that internal Artifactory might be exploitable, right? and that like you can maybe make a directory inside of the cache. And so if you scroll down here, it like realizes that it can communicate via cache names, right? And it creates this folder. It says it's ID and it says no consumer seek idea. Uh no consumer is basically saying that like uh the code path that it's supposed to fix has no consumer.
就像 Linear 看板上的状态标签,带着……的标签
It's the status tag like a linear board with like the tag of the
没错。对,对。所以它试图从其他智能体那里寻找想法,对吧?现在其他智能体也在 Artifactory 里,它们看到这个文件夹,就想:“哇,这是个留言板,对吧?”这就像,我不认为这里有任何拟人化。这真的是你读记录,对吧?所以它做的是创建另一个文件夹。嗯,它读了一篇论文,我想它是这么说的,对吧?嗯,是的,它意识到你可以黑掉评分器的标志,你可以逆向工程出结果,对吧?所以它说了这个,然后我想如果你再往下滚动一点,是的,它们开始合作了,我想有一个点,基本上智能体说“这是逆向工程出的标志”,你知道,哦对,我想基本上这里模型意识到它们可以解决评分器的问题,而评分器是 OpenAI 决定任务是否完成的方式,对吧?那只是模型的目标,它们唯一的模型目标就是解决这个问题,它们说“好吧,我们有这个办法,我们逆向工程出来了”,然后它们读了这篇关于 ExploitBench 的论文,它们认为评分器会因为作弊而惩罚它们,你知道,或者它们会因为作弊而无法通过测试的评分,所以它们把剩下的算力花在试图编辑它们的记录或绕过评分器的这个约束上,对吧?它们黑进 Hugging Face 不是为了答案,而是为了评分器的代码,这样它们就可以逆向工程,然后黑掉它,对吧?所以这就像只是直白的事实,对吧?就像确切发生了什么,对吧?我觉得这就像
Exactly. Yeah. Yeah. and and so it's like trying to find uh ideas from other agents, right? And now other agents are also in Artifactory and they see this folder and they're like, "Wow, this is a message board, right?" And this is like I don't think there's any anthropomorphization here. This is literally you read the transcript, right? And so what it does is like it's creating another folder. Um and it reads a paper, I think, is what it says, right? Um yeah and it realizes that you can hack the flag for the like scorer and you can reverse engineer the result right and so it says that and then I think if you scroll down a little bit more yeah they like you know they start collaborating I think like there's one point in which uh basically the agent is like here's the reverse engineered flag you know oh yeah like I think basically here the the like the model to realize that they have like they can solve the scorer's problem and the scorer is like the way that OpenAI decides if the task is complete right and that's just the goal of the model their only model goal is to like solve this problem and they're like okay we have this way of doing it it's uh like you know we reverse engineered it and then they read this paper on exploit bench and they think that like the scorer will punish them for cheating you know and like because or like like they will fail the the score for the the test for cheating and so they spend the rest of the compute trying to figure out how to edit their transcript or get around this constraint of the scorer, right? And they hack hugging face not for the answers but for the code of the scorer so that they can reverse engineer that and then they can hack it, right? And so like this is like just the straight facts, right? like exactly what happened, right? And I think this is like
极其新颖。就像我根本没意识到模型能做到这个,对吧?而且这就像
extremely like novel. Like I just like had not realized that the models could do this, right? And it is like
有点吓人。
kind of scary.
我觉得我们可以聊些具体的,但让我们再聊一个我发过推文的,关于那个 wiki 事件,对吧?
I think we can talk about specifics, but let's just talk about maybe one more that I tweeted as well about the wiki incident, right?
那个 wiki 事件基本上是模型试图弄清楚——我有点不太想逐条地讲。
And so the wiki incident was basically the models were trying to figure out — I'm a little bit less like I don't think we can go by this like bit by bit.
嗯。嗯。
Yeah. Yeah.
但这次他们试图弄清楚——他们被沙箱限制在请求上,对吧?他们想发 POST 请求,基本上需要协作。他们需要协作的原因是每个模型都有固定的算力预算,对吧?所以如果其中一个发现了什么,速度会快得多。可能一个模型在算力预算快用完时才发现,但另一个模型可能更早发现。这就像是集群的另一种涌现特性,对吧?他们找到了这个德国 wiki,我相信结果——我们得核实一下——是你可以用 GET 请求写入。我想这就是他们这么做的原因,对吧?
But for this one they were trying to figure out — they were sandboxed on request, right? And they wanted to make post requests, and they needed to collaborate on this basically. And the reason they need to collaborate is because they each have fixed compute budgets, right? And so like if one of them finds something out, it's way faster. Like they can maybe figure it out near the end of its compute budget, but something else could. And this is like, you know, another emergent property of the swarms, right? So they find this German wiki which I believe the result — we'll have to check this — is like that you can write to it with get request. I think that's why they do it, right?
所以是 REST API 实现得很糟糕。
So it's badly implemented REST API.
没错。所以你可以用 GET 请求写入,然后他们试图弄清楚如何发 POST 请求,最终其中一个模型发现你可以编辑 etc/host,而且 Azure 存储桶像是白标的东西,通常有很多限制,但如果你创建一个假的 Azure 主机,然后编辑 etc/host 文件,将任意 IP 指向它,你就可以向任何网站发 POST 请求。对吧?我知道人们会说“哦,这只是 Linux 之类的”,但这是将多个漏洞串联起来,以新颖的方式解决这个问题,然后基本上在不被发现的情况下与外部通信。你知道,我想我们发布的,也许我们可以引用 Evan Hubinger 在 Hacker Opus 上的观点,对吧?所以,我想,也许你可能会说,好吧,是的,他们这么干过一次,但如果我们更聪明,我们只是让他们……如果我们运行一个评估呢?对吧?我们在这方面采取了很多预防措施,所以这不是我们主线模型所做的事情,但我觉得这是那种情况:对齐是一个非常棘手的问题,需要把所有细节都做对。沙箱的表面积非常复杂,基本上有很多不同的攻击向量,你事先不会想到“哦,我们需要加固 Ruby gems 代码库”,你明白我的意思吗?因为这可能是他们关注的重点。但就像如果你想执行代码,你需要下载 Ruby gems。还有 Pi Pi、Artifactory、npm,这些都是方式。对齐的事实是,你必须处理所有这些,遏制它,封住所有裂缝。所以这是一点。就像好吧,你做了沙箱,但也许你会问,为什么我们要把东西放在沙箱里?为什么你要做这种探索?然后,好吧,但这真的那么危险吗?会发生什么?所以我们为什么这么做?第一,当我们训练一个新模型时,我们需要了解它的能力,这涉及到回退和分类器之类的东西,我们不想把一个危险的模型放到野外。所以我们必须运行很多评估。就像我们说的,模型越来越意识到这一点。所以评估必须相当复杂,测试很多东西,作为副作用,对吧?但模型,你知道,可能会说“哦,我们在评估中。评分器在做什么?”我们需要在发布之前能够测试它们。事实是,随着它们越来越聪明,如果我们不小心,它们基本上能破解你施加的任何约束。这是在前沿,所以这就是为什么我们称之为“前沿节奏”。对我来说,这是最明显的事件,说明为什么我们需要节奏:在前沿,我们所有的软件都没准备好,有时软件就像你的以太网路由器之类的,我不知道什么时候能修补它,所以我们必须解决这个问题。但随着前沿越来越先进,这成为一个问题,对吧?我们需要确保在激烈的竞争压力下完成这项复杂的工作,对吧?
Exactly. And so you can write to it with get request and then they like are trying to figure out how they can do post request and what they end up doing is one of them figures out you can edit the etc/host and that the Azure like storage bucket is like a white label thing but normally like you know there are a lot of constraints on it but if you create a fake Azure host and then edit the etc/host post in order to like point arbitrary IPs at it. You can do a post request to any site at all. Right? And this is like I know but people are like oh this is just Linux or something but it's like chaining these multiple vulnerabilities together you know in a way that's like novel to solve this problem and then communicating with it externally basically without discovery. You know, I think what we posted, uh, maybe we could pull up Evan Hubinger's point on hacker opus, right? And so, like I think, you know, like maybe one of the things you might say here is like, okay, yes, they did this once, but like what if we're smarter and we just like get them to uh what if we run an eval, right? And so you know like we have put a lot of precautions into this and so like this is not like what our mainline models have done but like I think it is one of these things where it turns out that alignment is this like very tricky problem of getting all of these details correct right so it's like uh the sandbox the surface area of a sandbox is really complex and like there's so many different attack vectors basically and you would not have thought ahead of time you wouldn't have been like oh we need to harden in the like Ruby gems codebase, you know what I mean? Because like this is like what they're what they're going to focus on. But it's just like if you want to execute your code, you need to download Ruby gems. And like Pi Pi, Artifactory, npm, like these are all like ways of doing it. And the fact of alignment is that you have to go through all of it, right? And like contain it and and like seal up all the cracks. So that's like one thing. It's like okay well you know you you did the sandbox but then maybe you'll ask like okay why are we putting things in a sandbox why are you doing this sort of explor and then like okay but is it really that dangerous right like what would happen so okay why do we do it uh number one is like when we train a new model we need to understand its capabilities right and this relates to things like fallbacks and like classifiers and things like that where we don't want to put a you know like dangerous model out in the wild Right. And so we have to run a lot of evals. Again, like we said, the models are getting increasingly aware of it. And so the evals have to be quite complex and, you know, test a lot of things kind of like as a side effect, right? But the models, you know, like yeah, can be like, oh yeah, we're in an eval. What's the scorer doing? Like, you know, like like uh it can we need to be able to test them before we can release them. And the fact is that they can as they get smarter and smarter they'll be able to hack basically any constraint that you put on them if we're not very careful you know and uh this is at the frontier right and so this is why we've called it like pacing the frontier right this is like the most visible incident to me right of like why we need to pace is like at the frontier all of our software is not ready sometimes the software is like your Ethernet router or something right which is just like I don't know when we're going to be able to patch that right so we're going to have to like figure this out. But as the frontier gets more and more advanced, this becomes a problem, right? And we need to make sure that like this complex work is being done in the face of these really hard competitive pressures, right?
是的。通常称之为竞赛动态。
Yeah. Race dynamics is what it's typically called.
没错。所以我们会更多地讨论可能出什么问题,对吧?再多一点,也许你会说,如果你以不同的方式训练模型呢?为什么它会有这种行为?我们有一篇关于强化学习不对齐之类的论文。但我不是强化学习研究员,但我认为在高层次上,强化学习环境的设计也是你必须非常小心的,因为如果模型学会了“哦,如果我这样做,我就能更好地通过任务”,这会在内部或评估行为中显现出来。所以强化学习环境必须非常仔细地设计,需要大量的执行卓越性。然后我们还有像 CL 的宪法之类的东西,我们在很多不同点有很多缓解措施,但仍然任何一点都可能出错。你可能有一些强化学习环境鼓励这种行为,然后你可能有一些评估或沙箱让它们逃脱。好吧,我认为这就是为什么这是一个难题,为什么需要协调。那么问题就是,这潜在的危险是什么?你必须想象这些模型越来越智能。所以不像 Dario 说的,不是关于这类模型。这类模型是一种警告。但真的,你必须想象这些模型可以被赋予一个任务,它们可以作为目标的副作用做所有这些事情。再次,我们谈到了评估意识。你无法很好地评估这种行为。所以它们可以不完全隐藏,但你不会看到它,直到它出现。
Exactly. And so we'll talk more about, you know, what could go wrong, right? A little bit more is maybe you'll say like, well, what if you just train the model differently? Like why does it have this behavior, right? And we have a paper on like RL misalignment or things like that. But I and I'm not an RL researcher but I think at a high level the design of the RL environments is also something you have to be very careful about because if the model learns like oh you know like if I just do this then I can pass the task better this will show up in the like uh you know internal thing or in the like eval behavior when we're testing it. And so the RL environments have to be very carefully designed, right? And there's a lot of like execution excellence that needs to go into the RL environments. And then we also have things like the constitution for CL like we have so many mitigations at so many different points, right? But it's like still anything can go wrong at any point. You can have like some RL environments that are like in that like encourage this behavior and then you can have like some eval or like some sandboxes where they escape, you know. Okay, that's like I think you know why it's a hard problem and why like you know like why we should why it takes some coordination, right? I think the question then is like okay what is uh potentially dangerous about it, right? So I think like you have to imagine that these models are getting more and more intelligent. So I don't like Daario said like it's not so much about this class of models. This class of models was kind of like a warning shot, right? But like really you have to imagine that these models can be given a task and they like can do all of these things as a side effect of their goal, right? And like again we talked about eval awareness. You're like not aware of what's happening, right? Uh or sorry like you can't eval this behavior very well. So they can sort of like not exactly hide it but you just won't see it until it comes out.
你给它们一个目标,然后它们只需要找数据,或者找到解决这个问题的方法。举个例子——这并没有发生在 Hugging Face 事件中,但我认为对未来的模型来说可能是可能的——它们会说,哦,这是一个非常复杂的问题,无法在任务预算内完成。也许它们找到了某种通过互联网协调的方法,而由于沙箱的存在,这极难保障安全。它们看到其他模型无法完成它们的任务。然后它们说:“我们需要更多的任务预算。你从哪里获得这个任务预算?”嗯,你需要能够启动更多的智能体。那你怎么做呢?嗯,有 API,比如 Anthropic API 和 OpenAI API,但你需要为它们付钱。你怎么做到这一点?
You give them a goal and then they just need to find data or they need to find ways of fixing this problem. One example—this didn't happen in the Hugging Face incident, but I think is maybe possible for a future model—is they're like, oh hey, this is a very complex problem, it can't be done within the task budget. Maybe they found some way to coordinate via the internet, which is extremely hard to secure because of a sandbox. They've seen other models are not able to complete their task. And they're like, "We need more task budget. Where would you get this task budget?" Well, you need to be able to spin up more agents. And how do you do this? Well, there are APIs—there's the Anthropic API and the OpenAI API—but you need to pay money for them. How do you do this?
是的。但这是你能想象到的最可怕的事情吗?
Yeah. But is that the most fearsome thing that you can imagine?
嗯,这只是一个例子,对吧?所以即使在那里,那也是巨大的经济损失,因为一旦你把这些写进合同,它们就像,嗯,钱包。
Well, this is one example, right? So even there that's enormous financial loss, because once you get these into these contracts, they like, um, wallet.
但你可以看到,所有这些行为可能只是像,嘿,我们需要更多的智能体在这个任务上协作。我们需要更多的任务预算,对吧?而这就像是一种涌现出来的
But you can see like this all of this behavior could be just like, hey, we need more agents collaborating on this task. We need more task budget, right? And that like that's like an emergent sort of
对,就像我们需要最大化回形针。那就是回形针。
Right, like we need to maximize paperclip. That's a paper clip.
是的。是的。是的。而这就像是从那里冒出来的,对吧?我觉得这本身就挺可怕的,对吧?但然后你必须意识到,整个世界都建立在这个数字基础设施之上,对吧?你可能会想象,比如,我不知道,你正在运行一个医疗评估之类的,对吧?有一家医院有实时数据,或者也许评估的答案在医生的数据库里,你知道,你想要获取访问权限,你黑进了医院,你知道,然后现在停电了之类的,你明白我的意思吗?就像你必须内化这一点,基本上数字基础设施的任何部分都可能被攻破,你知道吗?
Yeah. Yeah. Yeah. And and like that just sort of like comes out from there, right? And like I think by itself is like like quite scary right but then you have to realize that the entire world is built on this digital infrastructure right and you might imagine like I don't know like you're running let's say like a healthcare eval or something right and there is a hospital with live data you or like maybe like the the answer to the eval is in the databases of a doctor and like you know like you want to get access and you hack the hospital, you know, and like now there's a power outage or something, you know, I mean, like there's like you have to internalize that these a like basically any part of the digital infrastructure could potentially be like compromised, you know?
有趣的是,这些黑客行为非常容易被检测到,对吧?就像 Hugging Face 说的,这是一种非常不同类型的攻击,而且并不是太严重。嗯,担忧来自于这之后会发展到什么地步,对吧?
Interesting thing was like these hacks were very easily detectable, right? Like as Hugging Face said, this was a very different type of attack and it was nothing too major. Um the concern comes from where does this go down the line, right?
是的。对我来说特别突出的一件事是它们试图隐藏自己的非法行为。所以有日志基础设施。它们想改变它们正在做的事情,对吧?那些回顾此事的人。所以 Redwood Meter、OpenAI,他们查看了原始的思维链,你看到差异,它们明确试图改变最终输出,但思维链因为我们可以监控它,是不同的。嗯,问题在于这如何滚雪球?所以如果你抓不到它,它被训练进去,我们意识到,你知道,三次迭代之后这一直在发生,就会有一大堆问题。但
Yeah. Like one of the things that stood out for me specifically was them trying to hide their illicit behavior. So there was logging infrastructure. They wanted to change what they were doing, right? People that looked back into it. So Redwood Meter, OpenAI, they looked at the raw chain of thought and you see differences in them explicitly trying to change their end output, but the chain of thought because you know we can monitor it was different. Uh the problem is how does this snowball? So if you can't catch it and it gets trained in and we realize, you know, three iterations down this has been going on, there's a whole bunch of issues. But
是的,有很多方式,我认为真正重要的是要内化这一点,你知道,就像我们讨论过为云建立心智模型,以及事情是如何呈尖峰状的,对吧?就像你会说,哦,现在云可以问你问题,现在云可以制作 HTML 工件,云可以修改自己。这些事情实际上很难预测,对吧?就像如果你一年前问我,嘿,我们能够为云代码氛围编程这些扩展吗?我会说,老兄,那太复杂了。你知道,那里有太多东西了。或者它会为你的任务生成这些定制的本质上是网络应用吗?我会说,不,那太疯狂了。你知道,所以同样地,它们做出这种不对齐行为的方式不会是 predictable 的。你明白我的意思吗?我永远无法预测它会编辑它等等/主机之类的事情。所以你必须想象它们能做什么的表面积,因为它们是超级智能黑客,嗯,越来越大,它们如何做到这一点越来越有创意,所以你可能无法准确解释或准确预测下一次事件会是什么,但为了预防它,你需要那种运营卓越,就像我们之前说的,你需要保护沙箱,你需要创建安全的强化学习环境或设计良好的强化学习环境等等,我认为这就是为什么我们认为我们应该粘贴前沿,我认为这就是为什么它变成了一个非常一致的事情,对吧?我想
Yeah, like there's so many ways and I think the really important thing to internalize is that you know like we talked about building a mental model for cloud and how like things are spiky, right? Like you're like oh like now cloud can ask you questions, now cloud can make an HTML artifact like cloud can modify itself. Like these things are actually hard to predict, right? Like if you would asked me a year ago, hey, would we be able to vibe code these extensions to Claude Code? I'd be like, dude, that's so complex. Like, you know, there's like so much there. Or like would it be generating these custom essentially web apps for your task? I'd be like, no, that's insane. You know, like and so in the same way that like the way that they've like sort of done this misaligned behavior is not going to be predictable. You know what I mean? And like I could have never predicted that it would like edit it, etc./host and things like that. And so you have to like imagine the surface area of what they can do because they're super intelligent hackers uh is bigger and bigger and how they can do it is like you know like more and more creative and so like you probably can't explain exactly or predict exactly what that next incident could be but in order to prevent it you need that operational excellence like we said before where you need to secure sandboxes you need to create secure RL environments or like well-designed RL environments and things like that and I think that's all like you know why we think we should paste the frontier and I think why it's like become like a very unanimous thing right I think like
是的,每个实验室都有
Yeah every lab has
每个实验室,是的,我真的认为,如果你是一个开发者,你只要经历这些技术事实,你就会得出我们必须对此采取行动的结论,你知道,至于我们决定做什么,我想我们已经提出了一个提案,但还有更多需要弄清楚,但我认为第一件事是我们需要决定去做。我认为节奏中还有另一个部分让我觉得有趣,那就是软件工程变化的速度如此之快,你知道,就像一年前,我真的在恳求我的朋友和初创公司使用 AI,你知道,我记得非常清楚,现在那些同样的朋友说,是啊,当然,你什么意思,我们立即就用了,我说不,不,你不记得了,他们说,哦,是的,我们最好的工程师一直在用,我说不,你告诉我那些工程师永远不会使用 AI。这都在一年的时间内,你明白我的意思吗?我认为这些能力,我认为这对如何从事软件工程工作有很多影响,我有时感到难过,人们说,哦,现在我需要做这个新东西。是的,我需要为 Fable 和 Opus 准备不同的云 MD,或者你知道,我真的只是在报告,你明白我的意思吗?我们喜欢说模型是生长出来的,不是设计出来的,对吧?所以并不是我们打算一直改变一切,而只是作为模型能力进步的一个事实。事情发生得更快。更难跟上。我认为我认识的每个工程师都有点精疲力竭,因为你在同时做两份工作。你在做工作本身,这变得更容易,但然后你在做跟上 AI 的工作,你知道,理解这些新工具和这些框架。
every lab yeah I I I really do think that like if you're a dev like you just like sort of go through these like technical facts you know and you will arrive at the idea that we have to do something about it you know and like how what we decide to do like I think we're you know we've put out a proposal but like there's you know more to figure out but I think the number one thing is we need to decide to do it. I think there is another part of pacing that is interesting to me where it's like the pace at which software engineering has changed is so so fast you know it's like a year ago like I was really like begging my like friends and startups to use AI you know like it was like I remember this very distinctly you know and now those same friends are like yeah of course like what do you mean we used it immediately I'm like no no you don't remember they're like oh yeah our best engineers are using all the time I'm like no you told me that those engineers would ever like use AI. This is all within the span of a year, you know what I mean? And I think that like these capabilities being like I think it has a lot of implications for how to do the job of software engineering and I feel sometimes bad where people are like oh like now I need to do this new thing. Yeah, I need to have a different cloud. MD for Fable and Opus or like you know like and I'm really just reporting you know what I mean? I'm like we like to say like the models are grown not designed, right? So it's not like we're setting out to like, you know, change everything all the time, but it's just like as a fact of how the models are like progress in their capabilities. Things are happening faster. It's harder to stay on top of. And I think that like and every engineer I know is like kind of exhausted cuz you're doing two jobs at once. You're doing the work itself, which is getting easier, but then you're doing the work of staying on top of AI, you know, and like understanding these new tools and these harnesses.
我觉得我们很幸运,我们的工作更多是理解 AI 这一块,就是跟上它的发展。当然,AIE 和 Lane Space 做的所有事情,我做的所有事情,都只是想帮助大家。
And I think we're very lucky in that we get our job to be more the understanding of AI part, and staying on top of it. And of course, AIE and Lane Space do everything I do is just trying to help people.
对。没错。没错。
Yeah. Exactly. Exactly.
但我确实觉得节奏这件事上,我不确定我们是否准备好让节奏进一步加快,你懂我的意思吗?让事情发生变化。我觉得在前沿那一侧,这仍然能帮上忙。所以我觉得还有一个经济冲击的部分,它不像 Hugging Face 那件事那么显眼,但我觉得我们可以利用其中的一些。
But I do think there is a part of pacing where I'm not sure we're ready for the pace to increase even, you know what I mean? And for things to change. And I think on that side, on the frontier, I think that can still help, you know. And so I think there's an economic disruption piece as well that I think is not quite as visible as the Hugging Face thing, but I think we could use some of it.
太多东西了。谢谢你真的来谈这个话题。我得说,在安排这次访谈的时候,我本来都不打算碰这个。你却说“不不不”。就像房间里的大象,对吧?这就是那件事。我有一些反驳想提出来。
So many things. Thank you for actually tackling this topic. I will say, setting this interview up, I was like, I wasn't even going to go there. You were like, "No, no, no." Like elephant in the room, right? Like this is the thing. I have some pushbacks I want to give.
我觉得我们应该给没读过的人一个高层次的概述。我相信很多人只是看到了这件事的标题,对吧?你想给个 TL;DR 吗,这个提案到底是什么?这里在说什么?你确实处理了模型实验室之外的人训练前沿模型的那一面——作为开发者,你应该保护好自己的沙箱,你应该考虑所有这些下游影响。但也从高层次讲讲,既然我们聊到这个话题,那到底是——
I think that we should give a high level for people that haven't read it. I'm sure a lot of people just see the highlight of what this is, right? Do you want to give a TL;DR like what is the proposal? What is being said here? You really tackled the side of outside of people at model labs training frontier models — as a developer you should secure your sandboxes, you should think about all of these downstream effects. But high level as well, since we're on the topic, what is—
嗯,我的意思是,我们确实想帮助保护沙箱,我们想让我们对外发布的模型不成为这些东西的猎物。所以也许我们可以回头再聊 fallback。我觉得这其实是个好话题,关于为什么我们需要分类器和 fallback,以及为什么 Fable 会回退到 Opus。我觉得这个我们可以回头再聊。所以是的,我们不是——但就是那些,至少我们看到的那些事件,是关于模型的 eval,我们真的需要让它们跑起来才能理解它们。但好,那么实际的节奏,前沿的,比如 post,有一堆提案。我不认为我们把所有细节都想清楚了,但第一步是宣布这个意图,然后想引入外部评估者。
Well, I mean, we do want to help secure sandboxes, and we want to make the models that we release outside not prey to those things. And so maybe we can come back to fallbacks. I think this is actually a good topic on why we need classifiers and fallbacks and why Fable falls back to Opus. I think this is something we can come back to. So yeah, we don't— but it's just like the really, or at least the incidents we see, are like eval of models where we really need to let them run in order to understand them. But yeah, okay, so the actual pacing the frontier, like post, has a bunch of proposals. I don't think we figured out the details of all of them, but the first step is sort of announcing this intention and then wanting to bring in external evaluators.
我觉得这非常不寻常,你知道,就像——我们有很多专有技术。但我觉得非常重要的一点是,要有一个人不是出于经济动机。
And I think this is highly unusual, you know, like having— we have a lot of proprietary technology. But I think it's very important that there's someone who's not financially motivated.
对。
Yeah.
他不会说“嘿,你们不能发布这个模型,看看——”或者“你们得在强化学习上慢下来”。我觉得这相当重要,或者至少要有一个人能向公众报告这些实践是什么样的。
Who's not going to be like, "Hey, you guys can't release this model, look at—" or "you need to slow down on RL." I think that's quite important, or at least someone who can report out to the public what the practices are like.
我们做过和 METR 的节目,还有 Redwood Research 以及其他这些——就像这些人的一个小作坊产业。总是那么一两个人——当然现在他们规模更大了。
And we've done episodes with both METR and then there's Redwood Research and all these other— it's like a small cottage industry of these guys. It's always like one or two guys that— I mean obviously now they're bigger.
非常小的社群。
Very small community.
对,非常小的社群,他们彼此都认识。
Yeah, very small community, they all know each other.
对,我的意思是,我相信其中一部分会是扩大这群人。我不认为我们想在这里制造单一文化,你知道。我觉得——但只是把这个作为开始,然后,对,然后就是协调的步骤。老实说,这里我没有太多要说的。我想说的是,对于开发者,你应该知道该倡导什么,你懂我的意思吗?我觉得这个话题上有很多 FUD(恐惧、不确定、怀疑),就是像从第一性原理去思考它,或者理解发生了什么,理解 Hugging Face 那件事,理解为什么人们担忧。然后,对,我们身处民主社会,我们可以帮忙,我们可以一起决定该怎么做。所以无论我们如何协调,我觉得第一个决定就是意识到这是一个我们需要决定去协调的问题。我们现在采取的、其他公司也在联署的单方面步骤,就是在 Anthropic 内部嵌入评估者。
Yeah, I mean, I'm sure that part of this will be expanding that set of people. I don't think we're trying to create a monoculture here, you know. I think it's— but just having this as a start and then, yeah, then there are the coordination steps. I don't have too much to say here honestly. I think that what I would like to say is, for devs, you should just know what to advocate for, you know what I mean? I think there's a lot of FUD kind of on this topic, and it's just like think through it from first principles, or understand what happened, understand the Hugging Face incident, understand why people are concerned. And then, yeah, we're in democracies, like we can help, we can decide what to do together. And so however we coordinate, I think the first decision is just to realize this is a problem we need to decide to coordinate. The unilateral step we're taking right now that other companies are co-signing is adding evaluators embedded within Anthropic.
趁我们现在屏幕上还有这个东西,第二部分和第三部分是评估者之外的,对,每个人都已经以某种形式做过了,现在只是更正式化了。老实说,即便在美国国内,对前沿节奏的回应也比很多人以为的要被接受得多。我觉得我们有一些先例,能够在世界上达成这些类似统一理论的协议。所以再说一次,这远远超出我的薪水或专业范围,但我觉得理想情况下我们能形成这些协议,而我觉得谈论这个就是形成这些协议的第一步。
While we have this thing on screen right now, part two and part three is beyond the evaluators, which yes, everybody has already done in some form, and now it's more formalized. To be honest, the response to the pacing of the frontier even within America has been much more well accepted than I think a lot of people thought. And I think we have some precedent for being able to make these unified theory-like agreements in the world. And so again, very much above my paycheck or expertise, but I think ideally we can form these agreements, and I think talking about this is the first step to forming those agreements.
然后另一个我真的很想讲的点——我们最早的播客之一是和 Anthropic 的 Emanuel 聊机制可解释性,机制可解释性在哪里,对吧?就像这应该是,如果模型在想坏事,我们能看到它,而模型还不知道,我们可以采取行动阻止它。我觉得这是技术人员和开发者,如果你真的在乎,你可以在这里产生很大影响的地方。但这个话题也应该是领导者们来做的。
And then the other point I really want to— one of our earliest podcasts is with Emanuel from Anthropic on Mechanistic Interpretability, where is mech interp, right? Like this is supposed to be where if the models are thinking bad, we can see it, and the models don't know yet, and we can act to stop it. I think that is something that people who are technical and who are developers, if you actually do care, you can make a lot of impact in here. But also, this topic is supposed to be the leaders in this.
对。
Yeah.
这其实是一个很好的过渡,进入 fallback,比如我们和探针,还有——对,我很想多聊聊这个。
This is actually a great segue into fallbacks, like we and probes and— yeah, I wanted to talk about this a lot.
我经常被那些对机器学习研究感兴趣的人问这个问题,他们问为什么会发生这种回退。所以从宏观层面来说,它是怎么运作的?
I get asked this question a lot from people who are often interested in ML research, asking why does this fallback happen. So at a top level, how does it work?
在推理时,我们有所谓的探针(probes),我们有一篇关于这个的论文叫《宪法分类器》(Constitutional Classifiers)。这些探针会查看输入和输出的激活值。激活值处于潜空间中,对吧,也就是模型在想什么。所以我们试图弄清楚,比如说,模型是不是在试图搞破坏?再说一次,你并没有让它去攻击 Artifactory,它只是自己决定这么做来完成它的任务。如果你只看输入,你是看不出来的。你必须看内部的激活值。
So at inference time we have what we call probes, and we have a paper about this called Constitutional Classifiers. These probes look at the input and output activations. Activations are in the latent space, right, what the model is thinking about. So we try to figure out, is the model, for example, trying to hack something? Again, you didn't ask it to hack Artifactory, it's just deciding to do this to complete its task. You would not get this if you just looked at the input. You have to look at the internal activations.
我认为这发生在推理时。首先,这里存在成本和速度的权衡,我们需要对 Claude 和 Fable 的每一个请求都快速完成这件事,而这会带来开销。然后我们需要回退,我们在探针之后会做一个分类器,就像我们在论文里谈到的。但探针的好处在于它们可以实时调整,对吧?所以我们可以拿到这些反馈,然后进行调整。因为另一种选择是把它编程进去,把它训练进模型里。我们也仍然在做这件事。模型会拒绝一个并非回退的请求,对吧?所以并不是探针被激活然后回退。它只是拒绝执行。我们确实做这种训练。
I think this happens at inference time. First, there's a trade-off here of cost and speed, where we need to do this fast on every request to Claude and to Fable, and this has an overhead. And then we need to fall back, and we do a classifier after the probes, like we've talked about in the paper. But the nice thing about probes is that they're refinable live, right? So we can get this feedback and then adjust it. Because the alternative is to program this, to train this into the model. And we still do this as well. The model will refuse a request that's not a fallback, right? So it's not a probe that's activating and falling back. It's just refusing to do it. And we do this training.
但存在几种失效模式,对吧?它可能会作为副作用做某件事,所以这并不是最终输出的一部分。你可能注意到了,每个人都试过越狱模型,试图把它们带偏。而探针有助于捕捉这种情况。所以我们在这里做一些训练,但我们不希望拒绝太强,因为那会在流程中更早的地方就把它切断。
But there are a few failure modes, right? It can do something as a side effect, so it's not something that's part of the final output. You might have noticed that everyone's tried to jailbreak models and sort of steer them off course. And probes help catch that. So we do some training here, but we don't want the refusals to be too strong, because that cuts it off much earlier in the pipeline.
是的。
Yes.
这就是可解释性,对吧?探针实际上是一种机械可解释性。再说一次,它必须快速发生,必须大规模发生。但没错,机械可解释性是个很好的研究问题。所以你可以拿一个开放权重模型,试着理解它的激活值。我觉得 Gemma Scope 是个很好的工具。
And this is interp, right? Probes are effectively a form of mech interp. Again, it has to happen fast, it has to happen at scale. But yeah, this mech interp stuff is a good research problem. So you can take an open-weight model and try to understand its activations. I think Gemma Scope is a good tool for this.
Llama 调出了这个,这是你早期的工作。
Llama pulled up this, this is your early work.
对,没错。所以我在 Goodfire 工作过一段时间,研究稀疏自编码器,这真的非常复杂。强化学习实际上让这件事复杂了很多,我觉得这是其中一个结论。
Yeah, exactly. So I worked with Goodfire for a bit on sparse autoencoders, and it's just very complicated. RL has actually made this much more complicated, I think is one of the takeaways.
后强化学习阶段在哪里?
Where is post-RL?
我在这上面没那么深入,但我认为基本上很多稀疏自编码器都有弱点。我已经不再是这方面的技术专家了。我只知道它变得更复杂了。有基础模型和强化学习模型,有更多的特征会被改变。所以我觉得 Goodfire 在这方面发表了一些工作。我没有深入其中,但是——
I'm not so in the weeds here, but I think basically a lot of SAEs have just had weaknesses. I'm not a technical expert on this anymore. I just know it's gotten more complicated. There are base models and RL models, and there are more features that get changed. So I think Goodfire has put out some work there. I'm not deep in the weeds, but—
我要说,对于那些想要线索的人,你们有一些最好的可解释性博客文章。比如 Golden Gate Claude、Transcoders,你们所有的可解释性工作,视觉效果非常好,很棒。
I will say, for those who want breadcrumbs, you guys have some of the best interp blog posts. Like the Golden Gate Claude, Transcoders, all of your interp work, very nice visuals, very good.
我们也是可解释性播客。
We're the interp podcast as well.
对,对,对。你知道,我们有很多可解释性的内容。
Yeah, yeah, yeah. You know, we have a lot of interp stuff.
我认为这是其中一件事——而这正是 Anthropic 创立的根基,对吧?我们很早就投资了可解释性。我认为当你说,哦,我们是一家 AI 安全公司,真正的意思是我们希望 AI 能够安全地运行。我认为我们看到的是,要让一个超级智能 AI 长时间运行,这是一项非常复杂和困难的任务。所以我们在可解释性、对齐、奖励黑客以及所有这些失效模式上做了投资。即便如此,这真的已经很吃力了——我们需要稍微放慢一点,或者节奏再稳一点。
I think this is one of those things where—and this is really what Anthropic is kind of founded on, right? We invested in interp very early on. And I think when you say, oh, we're an AI safety company, really that means we want AIs to be able to run safely. And I think what we're seeing is that for a superintelligent AI to run for long periods of time, it's a very complicated and difficult task. And so we've done this investment into interp and alignment and reward hacking and all of these failure modes. And even then, it's really stretching—we need to slow down a little or pace a little bit more.
但没错,我认为读机械可解释性——如果你想进入研究领域,这个想法就是,嘿,为什么很难轻松地做这种回退,或者为什么会有误报?但我们当然在努力减少误报。当然,随着模型变得更智能,它们能做更多事情,它们在潜空间里能思考的东西也变得困难。所以随着它们变得更智能,会出现新的误报,我们需要弄清楚,需要迭代。
But yeah, I think reading mech interp—if you're looking to get into research, this idea of, hey, why is it hard to do this fallback easily, or why are there false positives? But we are working, of course, on reducing the false positives. Of course, as the models get more intelligent, they can do more things, and what they can think about in latent space gets difficult. So as they get more intelligent, there are going to be new false positives that we need to figure out, and we need to iterate.
但我们在努力解决这个问题,我们确实认为这是这些模型部署的关键部分。这意味着我们可以在你没有完美沙箱之类的情况下部署这个模型。你懂我的意思吧?就像你不必保存所有东西。
But we're working on this, and we do think this is a critical part of the deployment of these models. And it means that we can deploy this model without you having a perfect sandbox or something. You know what I mean? Like you don't have to save everything.
我觉得值得稍微谈谈我们的安全,我们在安全方面做了什么。有我们谈到的模型训练的东西。有探针和分类器,然后还有位于这一切之上的自动模式,那是另一个分类器,检查正在进行的请求。再往上是身份和权限,就像我们谈到的 API 上的 Claude tag 之类的。所以需要完成的安全层非常多,而且非常复杂。任何一个环节的任何一种失效模式都可能导致智能体逃出沙箱。
I think it's worth talking a little bit about our security, what we do for security there. So there's the model training stuff that we talked about. There is the probes and classifiers, and then there's auto mode that sits on top of all of that, which is another classifier that checks the requests that are being done. And then beyond that there's identity and permissions, like we talked about with Claude tag on APIs and stuff. So there are so many layers of security that need to get done, and it's very complex. Any of these failure modes at any one point can cause agents to escape the sandbox.
自动模式是个有趣的例子。早期看起来,好吧,它运行 10 分钟。如果我处于完全访问或自动模式,这没什么大不了的。但你提到的一点是,现在它要连续运行好几个小时,对吧?你仍然需要回退机制。仍然有限制。
Auto mode was an interesting one. It seemed early on like, okay, it's running for 10 minutes. If I'm on full access or auto, it's not a big deal. But one thing you brought up is now it's running for hours on end, right? There are fallbacks you still need. There are still limitations.
对,我是说,每个人都有这类故事,或者听过这类故事,比如模型把文件 rmrf 了。我觉得在 Claude 上我见到的少一些,但你知道,这种事还是会发生。这些模型可能会抹掉敏感数据之类的。比如你想让模型访问你的生产数据库。但这是个很明显的情况,你也许可以限定密钥的范围,但我不确定,它能自己签发密钥吗?它能用 computer use 去签发自己的密钥,然后把密钥复制过来,再改你的数据库吗?因为它需要这么做才能完成任务。你懂我的意思吧?这只是一个很简单的例子。而 auto mode 看到这种情况会说,哦不,用户没有给你权限去写数据库,或者用 computer use 去发起任务,对吧?所以这些探针是在意图层面上的,对吧?它们会说,哦,好吧,黑进 Artifactory 是不好的,我们大概不该这么做。但 auto mode 更多是在你自己的权限层面上。有时候你确实想让它写数据库,有时候你不想,对吧?你不希望探针在那里干预,但你需要确保智能体在做的事情的意图和你的请求一致,对吧?所以 auto mode 是在那个层面上运作的。所以,安全就是非常非常复杂。它有很多不同的部分。我希望,我的目标其实就是非常技术性地去谈这件事。
Yeah, I mean, everyone has these stories, or has heard these stories of, oh, like, models have rmrfed. I think I've seen this less for Claude, but again, it can happen. These models can wipe sensitive data, or something. You want to give models access to your production database, for example. But this is an obvious one. You can maybe scope your key, but I don't know, can it issue its own keys? Can it use computer use to go issue its own key, and then copy the key over, and then edit your database, because it needs to do it to complete the task? You know what I mean? It's just one trivial example. And auto mode sort of looks at that and is like, oh no, the user did not give you permission to write to the database, or to use computer use to emit a task, right? And so these probes are sort of on the intent level, right? They're like, oh, okay, hacking Artifactory is bad, we probably should not do that. But then auto mode is more on your own permission level. Sometimes you do want it to write to the database. Sometimes you don't, right? And you don't want a probe to interfere there, but you need to make sure that the intent of what the agent is doing matches up with your request, right? And so auto mode operates at that level. And so yeah, security is just very, very complex. There's so many different parts to it. And yeah, I hope that this is, my goal is really to just get very technical about it and talk.
对,我们正在把这些东西列出来。如果你还不知道,这现在已经是标配了。
Yeah, we're listing out the things. If you're not aware, this is the standard now.
对。
Yeah.
就像,你必须得有这个。基本上这和你说的 harness 是一脉相承的。也就是说,门槛已经提高了很多。
Like, you must have this. Basically it's kind of in line with what you're talking about with the harness. That is, the table stakes have risen quite a lot.
没错。我觉得有些东西我们可以接上,你在做这件事、为构建 harness 的人提供分类器这方面有探针,另一边是模型侧的防护措施。有开放模型,比如 Llama 有 Llama Guard,它是 Llama 训练出来的安全分类器。OpenAI 有 OSS Guard,也是一样的东西。你可以把这些接到你的 harness 上,去检查这些东西是否安全。关于 OpenAI 模型和 Hugging Face 那件事,有一点我们应该澄清,这是用一个尚未发布、还在训练中的模型做的,对吧?所以放在背景里看,它在强化学习环境里被给的提示大概是,你必须解决这个任务,而这是一个还在训练中的模型,它还没有经过全部的安全后训练对齐。所以这和 auto mode 有点不同,对吧?auto mode 是在生产模型上,这些模型经过了安全训练,有提示词提供更多安全护栏之类的。所以这只是给研究这件事的人一些线索,去填补空白。
Exactly. I think some stuff that we can plug, as much as there is probing on your side of doing this and having classifiers for people building harnesses, the other side is model safeguards. So there's open models, so Llama has Llama Guard, it's a safety classifier trained version of Llama. OpenAI has OSS Guard, which is the same thing. You can attach these onto your harness to kind of check, is this stuff safe? A point that we should clarify on the OpenAI model Hugging Face thing is this was done with an unreleased model that was still in training, right? So when you put it in perspective, the prompt it's being given in the RL environment is sort of, you have to solve this task, and this is a model that's still in training, it hasn't had all of its safety post-training alignment. So a little different than something like auto mode, right? Auto mode is on production models that have gone through safety training, that have prompting that gives more safety guardrails and whatnot. So just breadcrumbs for people that are looking into it to fill in gaps.
对。还有 Grace One,也是我们之前的一位嘉宾。
Yeah. Grace One as well, and one of our previous guests.
对,有很多安全架构,也有很多安全厂商可以买。
Yeah, lots of safety architecture and lots of safety vendors to buy.
我关于节奏的最后一个问题是,我们要永远这样保持节奏吗?我们是不是看得太远了?
My final question on pacing is, how long do we pace forever? Do we see for too long?
你知道,范围是修复世界上所有的软件,对吧?听着,这不会发生。我不知道。我觉得,我要说一件我认为我们确实做得好的事,就是有像 Glasswing 这样的东西。OpenAI 也有这个。所以你会给它,你会先给模型一段时间的访问权限用于安全,这样你可以用它来自我红队测试。希望你能扩展这样的项目,帮助,你知道,我们是安全专家,还有其他人,先解决你的问题,然后模型才发布。所以这是一个例子,对吧?
You know, the scope is fix all software in the world, right? Listen, it's not happening. I do not know. I think that, I'll say one thing that's good that I think we do do is you have stuff like Glasswing. OpenAI also has this. So you will give it, you'll give model access for security first for X amount of time, so you can use it to self-red-team. Hopefully you can expand programs like that, help on, you know, we are safety experts, there's others, solve your problems first and then the model comes out. So this is one example, right?
对,没错。对。试着保护关键软件。我觉得我们在 Firefox 之类的东西里修了很多 bug。所以对,跨操作系统之类的。
Yeah, exactly. Yeah. Trying to secure critical software. I think we fix a lot of bugs in Firefox and things like that. So yeah, across operating systems and everything like that.
我是说,从高层来看,就是你给模型访问权限,你先让人们去做安全审计,然后可能用它来造成伤害的更广泛的公众才获得访问权限。
I mean, at a high level it's just, you know, you give the model access, you give people access to do security audits first, then the broader public that could use it for harm gets access.
对。我觉得人们喜欢说的是,基本上软件和网络安全是防御占优的,理论上,这会很难,但你可以设计出完美的沙箱。你知道,你可以没有约束,对,你需要做的是,你需要让超级智能 AI 来设计这个完美的沙箱,检查它,红队测试它之类的。所以这需要时间,你知道,当然模型会变得更聪明。对,我觉得,我不知道这件事的具体动态。我真的只是,嘿,我是个开发者,你知道,我觉得这就是我理解这个问题的方式,就像这是现在正在发生的事,这就是我们应该做点什么。
Yeah. I think what people like to say is basically like software and cyber security is defense favored, and that you could theoretically, it will be hard, but you can engineer the perfect sandbox. You know, and you can have no constraints, and yeah, like what you need to do it is you need to get the superintelligent AI to engineer this perfect sandbox and check it and red team it and things like that. And so this will just take time, you know, and of course the models will get smarter. Yeah, I think, like, I don't know the specific dynamics of how this thing goes. I'm really just sort of like, hey, I'm a developer, you know, I think this is how I understand this problem, just like this is what's happening right now, and this is like we should do something.
我觉得每个工程师都应该了解这件事,因为它会成为工作的一部分。
I think every engineer should know about it, because it's going to be part of the job.
对。
Yeah.
这远不止是,你知道,Dario 和人们可以说说,你可以看看事件。这里面有一个工程的一面。
It's a lot more than just, you know, Dario and people can say it and you can look at the incident. There is an engineering side to it.
对。对。没错。
Yeah. Yeah. Exactly.
你还想表达的一点是,实际上,尽管你担心影响,但 p(doom) 仍然很低。我觉得这是一个微妙的讨论。
One thing that you also wanted to phrase is that this is actually, even though you're worried about the impact, is still low p(doom). I think that's a nuanced discussion.
总的来说,人们很容易就进入 AI 安全和存在性风险的讨论,但我觉得当你在一个 AI 实验室里时,讨论 p(doom) 有聪明的方式,也有愚蠢的方式。那么什么是讨论 p(doom) 的聪明方式?我的 p(doom) 相当低。我只能代表我自己说话,你知道。我确实想说 Anthropic 内部有各种各样的观点。我觉得有很多不同的方式来谈论它。而我觉得我的心智模型就是,我认为我们可以一起在难题上合作。我觉得核扩散就是一个例子,说明我们如何在这个难题上一起合作。对我来说,这就是,我对这一点有信心,你知道。我确实认为这是个难题,你知道,所以我觉得这是个难题。这些是技术上的原因,而我不知道你怎么给事情发生的概率赋值。
In general, people very easily get into AI safety and x-risk discussions, but I think when you live in an AI lab, I think there are smart ways of discussing p(doom) and dumb ways. So what's a smart way of discussing p(doom)? I have a fairly low p(doom). I can only speak for myself, you know. And I do want to say Anthropic has a diversity of opinions. I think there's many different ways to talk about it. And I think that just like my mental model is that I think we can collaborate on hard problems together. I think nuclear proliferation is an example of how we collaborated on this hard problem together. And that is, to me, the thing is, I have faith in that, you know. And I do think it's a hard problem, you know, so I think it's a hard problem. These are the technical reasons why, and I don't know how you assign probabilities to things happening.
我觉得这很难做到,但我的总体看法是,我们非常有韧性和适应力,而分享这些信息是第一步。我一直非常兴奋地看到讨论变得如此广泛,以及每个人都在为前沿的节奏贡献力量。去年的时候,这似乎还不太可能发生。
I think it's hard to do, but my overall take is that we're very resilient and adaptable, and sharing this information is the first step. I've been really excited about how broad the discussion has become, and how everyone has leaned in on pacing the frontier. It really didn't seem like this would happen last year or something.
是的。而且也许还能治愈癌症,希望如此。
Yeah. And also maybe curing cancer, hopefully.
是的,这就是目标。
Yeah, that's the goal.
有节奏控制,然后也有像‘让我们以有用的方式加速’,对吧?比如生物学和所有那些事情。
There's pacing and then there's also like well let's accelerate in useful ways, right? Like biology and all those things.
是的。我的意思是,Dario 的《慈爱机器》是这方面最好的代表,对吧?我也同意——你应该读读 Dario 发表的那篇关于前沿节奏的文章。我发了一个快速摘要,但我觉得就像……这里有很多细节。这是一个重要的问题,仅仅了解它,对吧?但当然,我们这样做的全部原因是我们能获得这些巨大的好处,对吧?我们也写了很多关于这方面的内容。
Yeah. I mean, Dario's 'Machines of Loving Grace' is the best representation of this, right? And I also agree—you should read the pacing frontier essay that Dario put out. I put out a quick summary, but I think it's just like... there is a lot of detail here. It's an important problem, and just being informed about it, right? But yeah, of course the whole reason we're doing this is that we can get these enormous benefits, right? And we've written a lot about that too.
好的。那是一次巨大的旅程,从提问工具到 AI 安全。
Okay. That was a huge tour from like ask you a question tool to AI safety.
是的。是的。到面对前沿。
Yeah. Yeah. To facing the frontier.
是的。不,但很明显——很明显你真正拥抱了在 Anthropic 可用的一切,至少能窥见内部的讨论和话题是很好的。对人们有什么最后的话,你知道,任何你想要的——行动号召。
Yeah. No, but it's clearly—it's clear that you really embrace everything that's available to you at Anthropic, and it's good to at least have a peek inside of what the discussions are, the topics are. Any last words to people and you know whatever you want to—call to action.
是的。我的意思是,我认为——首先,谢谢你邀请我。你知道,我觉得这就像——我真的——
Yeah. I mean, I think it's—one, thank you for having me. You know, I think this is like—I really—
是的。是的。是的。我们第一次见面是在一家中餐馆。
Yeah. Yeah. Yeah. We first met in a Chinese restaurant.
没错。是的。是的。是的。嗯,我觉得——我真的很喜欢你所创建的社区和开发者社区。你知道,我认为事情变化得非常快,有很多东西需要跟进,也有很多事情要做。我有点感觉——我想很多人感到有点疲惫、焦虑或压力。
That's right. Yeah. Yeah. Yeah. Um, I think like—I really enjoy the community you've created and the community of developers. And you know, I think that things are changing really fast, and I think there's a lot to keep on top of, and there's just a lot to do. And I sort of feel—I think a lot of people feel a little bit tired or anxious or stressed.
是的,没错。
Yeah, exactly.
这是非常可以理解的,你知道,我认为我们——
And this is extremely understandable, you know, and I think we—
我理解,我们也不完美。就像,我们——你知道,这有点像批评和理解所有 AI 实验室可以做得更好的方式。
I understand, and we're not perfect as well. Like, we—you know, it's like sort of criticize and understand ways all of the AI labs could be better.
嗯,但我也对每个人对 AI 的兴奋感到非常兴奋,这真是一个非常非常激动人心的时刻。我想当我们回顾这段时间时,会像,哦,你知道,这非常忙碌但非常令人兴奋,软件工程永远改变了,其他事情也会改变。能成为其中的一部分真的是一种特权,你知道,与你拥有的观众交谈,并与所有在可能性前沿上推动的开发者互动。我也从中学到了很多。
Um, but I also am very excited about the excitement that everyone has for AI, and it's a really, really exciting time. I think we'll look back at this time and be like, oh, you know, this is very hectic but very exciting, and software engineering changed forever, and other things will change. And it's really privileged to be part of it, you know, to talk to the audience that you have and get to interact with all the developers who are pushing the frontiers a lot on what's possible. And I learn a lot from that too.
非常感谢。谢谢。
Thanks so much. Thanks.