Claude Tag:你在 Slack 中的主动队友

Claude Tag: Your Proactive Teammate in Slack

拉米斯·穆克塔 Lamis Mukta · AI Native Dev · 2026-07-07 · 约 58 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 的 Lhamu Dolma 介绍 Claude Tag,一个在 Slack 中工作的主动 AI 队友,具有持久性、记忆能力,并能主动发起对话和长时间完成任务。

Anthropic's Lhamu Dolma introduces Claude Tag, a proactive AI teammate that works in Slack, with persistence, memory, and the ability to initiate conversations and complete tasks over long periods.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 22)

全文 · Full transcript(中英对照)

Claude Tag简介 Introduction to Claude Tag

Host

大家好,欢迎收听新一期的 AI Native Dev 播客。非常荣幸地介绍来自 Anthropic 的技术团队成员 Lhamu Dolma。欢迎来到我们的节目。

Hello and welcome to another episode of the AI Native Dev and it's wonderful and my pleasure to introduce Lhamu Dolma, member of technical staff from Anthropic. Welcome to the podcast.

Lamis Mukta

嘿 Simon,很高兴今天能来。

Hey Simon, so happy to be here today.

Host

太棒了。我们将深入探讨智能体式编程,Anthropic 如何在内部使用 Claude Code 和智能体式开发,同时也要热烈祝贺你们最近发布并推出了 Claude Tag,我们马上就会聊到。但首先,Lamas,请简单介绍一下你自己、你在 Anthropic 的职责以及你具体做什么。

Amazing. We're going to have a great conversation about agentic coding in general, how Anthropic use Claude Code and agentic development in your environment, and also a massive congratulations for the very recent announcement and release of Claude Tag, which we'll go into in just a second. But Lamas, tell us first a little bit about yourself, your role in Anthropic, and what you do.

Lamis Mukta

好的。我是 Lamas,Anthropic 的技术团队成员,隶属于应用 AI 团队。这个团队介于产品、研究和市场推广之间。我们的工作包括直接与客户合作,我主要与初创公司和创始人打交道,同时也与产品和研究部门合作开展一些内部项目。

Yeah, absolutely. So, I'm Lamas. I'm a member of technical staff at Anthropic, and I sit in our applied AI team. So, this is a team which sits between product, research, and go to market. And so, we do a mixture of working directly with customers. I tend to work with startups and founders, and we also work on some internal projects with the product and research departments.

什么是Claude Tag? What is Claude Tag?

Host

太棒了。今天我们有很多话题要聊,但让我们从最激动人心的消息开始。最近 Claude Tag 发布了。显然你们已经在内部使用了好几周。请跟我们讲讲 Claude Tag 是什么,以及你们如何使用它?

Amazing. And we've got a lot to talk about today, but let's start at the very exciting news. Very recently, Claude Tag was announced. You've obviously been using this internally for many weeks now. Tell us a little bit about Claude Tag. What is it, and how do you use it?

Lamis Mukta

好的。Claude Tag 是你的主动式队友,目前它就在你工作的 Slack 中与你相遇。Claude Tag 的特别之处在于,它拥有你在使用 Claude Code 时熟悉的所有连接器、工具和上下文,但额外增加了主动性和持久性,能够更长时间地跟进工作任务。正如我们所说,它能同时与你及你的队友协作。

Yeah, absolutely. So, Claude Tag is your proactive teammate, which meets you right where you work in Slack for now. And what's really special about Claude Tag is it has all of the kind of connectors and tools and context that you're used to when you're using Claude Code, but it's got this extra degree of proactivity and the persistence to see pieces of work through for longer periods of time. And like we said, it is able to collaborate with you and your teammates at the same time.

Slack中@Claude与Claude Tag的区别 Difference between @Claude in Slack and Claude Tag

Host

是的,这太棒了。即使在 Tessell 内部,我们也大量使用 Claude Code。我们做的一件事听起来有点类似,所以我想问一下区别:我们在 Slack 中集成了 Claude 和 Claude Code,经常通过 @Claude 来让它帮我们做一些事情。这已经有一段时间了。那么,在 Slack 中 @Claude 和 Claude Tag 之间有什么区别?

Yeah, and that's amazing. And even internally in Tessell, we use Claude Code very heavily. And one of the things that we do, which sounds a little bit similar, so I'd love to kind of ask how it differs, is we have an integration of Claude and Claude Code within Slack. So, we very often tag Claude to do certain things for us. That's been available for a little while. What's the difference between that, tagging Claude there and Claude Tag?

Lamis Mukta

嗯,这是个好问题,因为表面上看起来非常相似,但实际上底层的一些细微差别区分了这两个产品。首先,Claude Tag 非常主动,它可以长时间执行工作任务,然后回来通知你事情已完成。这充分利用了当前编程或智能体长时间执行任务的能力。这是第一点。第二,通常你会通过提问来触发智能体,然后 Claude in Slack 会回复。而 Claude Tag 有时会主动来找你,告诉你有些事情需要你关注。所以这又体现了主动性。它还有跨频道的记忆和上下文,能够理解整个团队在做什么,因为它同时与所有队友聊天,这些信息会随时间积累。因此,你不再只是进行一次性交互,Claude Tag 能够跳出会话和个人的范围。

Yeah, so I think this is a great question because I think on the surface these things look really similar and it's actually some of the more intricacies underneath that differentiate these two products a bit. But I would say first of all, Claude Tag is incredibly proactive so it can go and execute pieces of work for long periods of time and then come back and let you know when that thing is done. So it's really leveraging the ability of coding or agents these days to do tasks for very long periods of time. So that's one thing. The second is like often you would go and trigger your agents by asking them questions and kind of Claude in Slack would reply. Claude Tag will actually sometimes come and find you and tell you that something needs your attention. So again there's this element of proactivity. It also has memory and context which can span your channels so it can understand what one whole team is working on and because it's chatting to all of your teammates at the same time that's something that builds up over time. So rather than you just having these one-off interactions, Claude Tag is able to move out of the scope of kind of sessions and individuals.

Host

这太酷了。所以它几乎可以主动发起对话来询问?

That's really cool. So it can almost initiate the conversation to ask?

Lamis Mukta

完全可以,是的。

It absolutely can, yeah.

内部使用示例与上下文记忆 Internal usage example and context memory

Host

是的,太棒了。我们在 Slack 中使用 @Claude 的方式通常是:先在 Slack 里讨论一个需求、功能或 bug,然后可能 @Linear 说需要创建一个工单,接着很快我们就会找 Claude 说:“请实现这个。”它会很快回复:“好的,我会使用这个线程的上下文。”它显然也了解代码库,然后会提供一个拉取请求并问:“你能看看这个吗?”然后我们会说:“没问题,看起来不错。合并吧。”所以这里有几个变化。第一个显然是历史记录——它似乎会知道跨 Slack 或更广范围的过往交互。这是你提到的第一个点吗?

Yeah, amazing. So the way we tend to use the @Claude invocation or tag within Slack—the way we tend to use it is very often we'll start discussing a need or a feature or a bug maybe in Slack and then we'll maybe @Linear and say we need a ticket about this and then we'll pretty soon after that just hit up Claude and say, 'Can you just implement this, please?' And it will return pretty quickly and say, 'Okay, I'll use the context from this thread.' It obviously knows the code base as well and then it will provide something like a pull request and say, 'Can you have a look at this and see?' And then we'll say, 'Yeah, absolutely. This looks good. Merge this and we're good.' So there's a couple of changes there then. The first one is obviously the history there—it sounds like the history it will know about past interactions across the Slack or beyond. Was that one of the first ones that you mentioned?

Lamis Mukta

是的,完全正确。关于上下文这一点,Claude Tag 有一套完整的权限系统,你可以真正控制并调整它到你想要的方式,但基本上它的范围限定在频道级别。

Yeah, absolutely. So, I think on that context point, we have this whole permissioning system with Claude Tag, and you can really control that and tune it to how you want it, but essentially it's kind of scoped to a channel level.

Claude Tag的上下文与记忆 Claude Tag's Context and Memory

Lamis Mukta

所以,如果这是一个你经常讨论功能请求的频道,它就有大量关于历史功能请求的上下文,以及流程是怎样的,例如,通过技能来端到端地跟进这些流程。所以,有很多历史和上下文。如果配置得当,它还能搜索 Claude Tag 所在的其他公开频道。所以,也许你有一个客户支持频道,那里有更多关于问题出处的上下文,而 Claude Tag 在获得权限后,也能访问这些信息并将其呈现出来,甚至你可能自己都没有那种可见性。所以,这是一点,就是它的上下文窗口或者说记忆范围更大了。

So, if this is a channel where you're always talking about feature requests, it has a ton of context on historical feature requests, and how like what the process is, for example, through skills for seeing those processes through end-to-end like. So, there's a lot of history and context. If configured as well, it can also have the ability to search other public channels that Claude Tag is in. So, maybe you have like a customer support um channel where there's a bit more context on where this issue arose, and Claude Tag, given the permissions, could also access that information and surface it where maybe you might not even have that kind of visibility. So, that's one thing, it's just this like this context window or the the memory scope that it has is larger.

Host

我想第二点是,你提到之后你会去 Claude Code 实现 Linear 里的工单吗?

I think a second is you mentioned that do you go to Claude Code to implement the tickets afterwards after that in Linear?

Lamis Mukta

我想这通常是我们做的方式。

I think that's typically the way we do it.

Host

好的。

Okay.

Slack中的端到端工作流 End-to-End Workflow in Slack

Lamis Mukta

是的,所以我认为这里要强调的一点是,你可以在 Slack 的单个线程中实现整个端到端的工作流。那么,假设在这种情况下,有一条反馈进来了。在 Claude Tag 的世界里,你可以配置它,让 Claude Tag 先抓取这条反馈,然后@所有它认为应该负责或应该参与回复的人。作为其技能和工作流的一部分,它还可以自动创建 Linear 工单,因为这是你配置好的工作流。嗯,这只需要通过提示和它随着时间的推移在执行工作流中学习就能实现。然后最后一点,我认为是真正的区别,就是你不再需要打开 Claude Code 去让它执行了。所以,它可以直接理解它需要做的工作流是:创建工单、开始编码,这样它就能启动自己的沙箱,如果你有权限的话,它可以在你的仓库中运行代码,然后在 PR 准备好时通知你。它还可以做一些事情,比如验证这些方面是否满足最初工单提出的原始需求。所以,我认为总的来说,如果你希望它能够运行从工单到——可能先问你是否要执行代码(如果你想要人工把关的话)——然后推动 PR,提示人们审查等等,它端到端地更加主动。从这个意义上说,它更加主动,并且有上下文和记忆来做好这项工作。

Yeah, so I think one one thing to surface here is you could achieve that whole end-to-end workflow in just one thread in Slack. So, let's imagine in this case a piece of feedback comes through. In the Claude Tag world, you could have it configured so that Claude Tag actually picks up the feedback first and tags everybody they think should be responsible like should be involved in the response to that. As part of its skills and workflows, it could also then automatically raise the Linear ticket because that's the workflow that you have configured. Um and that's available just through kind of prompting and and it learning over time through doing that workflow. And then the final thing is, which I think is the real difference, is you no longer need to open Claude Code to go and ask it to execute that. So, again, it could just understand that the workflow it needs to do is kind of raise it start coding so it can spin up its own sandbox, it can start running the code in your repository if you've got access to it, and then it can ping you when the PR is ready. It can also do things like verify the aspects of that meet the original requirements from whatever ticket was raised in the first place. So, I think what we see here overall is more proactivity end to end if you do want it to be able to run that process from kind of ticket to from probably asking you if it if you want it to execute the code first if you want a human gate there and then kind of seeing that PR through prompting people to review it etc. etc. It's much more proactive in that sense and it has access to the context and memory to do that job well.

Claude Tag作为编排器 Claude Tag as Orchestrator

Host

而且听起来它也变得更像一个编排者了。所以,以前我通常需要打开 Linear 说“为这个创建工单”。现在听起来,如果我能直接把 Claude Tag 拉进来,它实际上可以在后台完成这些,而我更多考虑的是我想要改变的功能或特性,而不是工作流的机制,比如“好的,先创建工单,然后再把 Claude 拉进来”。

And it sounds as well that it it's becoming more of the orchestrator as well. So, whereas I would need to normally hit up linear and say raise a ticket for this. It sounds like if I can just pull Claude tag in, it can actually do that behind the scenes and I'm more thinking about the functionality or the feature that I want to change versus the mechanism of the workflow of okay, let's raise a ticket first. Now, let's pull Claude in at that at that stage.

Lamis Mukta

是的,完全正确。所以,我认为这又回到了各个团队需要思考的问题:他们希望在多大程度上、在工作流的哪个环节来引导智能体。所以,你可以非常明确地设定:这些是绝对的关卡,要么你绝对没有权限做这些事情,要么我希望你总是在这些领域征求我的意见。而且我认为我们已经看到,因为 Claude Tag 能够随着时间的推移建立记忆,它对这些工作流的样子有了很好的感觉,并会进行调整。嗯,这真的很酷,我想再提一个关于这个新工作流的最后一点区别:如果你去实现这个 Tag,你会看到的是,启用这种跨职能工作流的能力非常酷。例如,你有一个客户支持工单被创建,你可能会@一些负责该客户的客户经理,只是为了让他们看到,以便他们了解情况。然后你可以@并循环你的工程师或产品人员,让他们对某些产品实现提供反馈。也许如果 Tag 带着一个计划回来,然后在你需要代码审查和整个部署工作时,再让工程师参与进来。所以,我认为我们看到的是这种多人协作能力,到目前为止,仅通过 Claude Code 或 Cowerk 之类的东西是很难实现的。所以,我认为这就是我们看到这些效果真正倍增的地方。

Yeah, absolutely. So, I think again, it's really for individual teams and to to think about the degree to which like where where it is in that workflow that they want to steer the agent. So, you can be very clear about like these are the absolute gates where either you absolutely do not have the permissions to do these things or I always want you to ask for my input on these these areas. And I think what we've seen is because Claude tag is able to build up this memory over time, it gets a really good feel for what those workflows look like and it adjusts. Um so, that's really cool and I think just to throw in one last difference on that new workflow, what you'll see if you kind of went and implemented this tag um is just the ability to enable this cross-functional workflow is really cool. So, for example, like you have the customer support ticket raised, you might kind of tag some of the account executives that are on on that account just to see if they like for their visibility so that they know what's going on. You can then kind of tag in looping your engineers or product people for feedback on some of the product implementation. Maybe if if tag comes back to you with a plan and then have your engineers looped in later when you need that code review to go through and the whole deployment to work. So, I think what we see is this like multiplayer ability, which has been quite difficult to achieve so far with maybe just working through Claude Code or Cowerk or something like that. So, I think this is where we see some of those effects really multiplying.

工作流变革:从单人游戏到多人游戏 Workflow Change: From Single-Player to Multiplayer

Host

太棒了。而且会有很多听众可能目前没有在 Slack 环境中使用 Claude 或 Claude Tag。所以,这实际上是一个相当大的变化。那么,当我们考虑流程变化时,我们是从什么状态变到什么状态?假设人们现在是智能体式开发者。它如何改变流程,让我们基本上更多地生活在聊天环境中、更多地生活在 Slack 中,而不是我们更智能体的流程或 IDE 中?工作流如何变化?

Amazing. And there are going to be a lot of listeners who maybe don't use Claude or Claude Tag in the in the Slack environment today. So, it's actually a pretty a pretty major change. So, when we think about the the process change, what are we going from and to? So, let's assume people are agentic developers today. How does it change the flow whereby we're essentially living more in our chat environments, more in our Slack versus potentially our more agentic flows or IDE? How does that workflow change?

Lamis Mukta

是的,所以我认为有几件事。为了从我们内部看到的影响来框架化,自从我们在内部开始使用 Claude Tag 以来,我们的产品团队,比如产品工程团队,有 65% 的 PR 是由 Claude Tag 打开的。哇。所以,这就是规模,显然我们是 Claude Code 的早期采用者。而现在,它常常是我们启动某些工作流以及这些工作流起源的停靠港。所以,是的,这就是我如何框架化我们目前所处的位置。在工作流方面,是的,我认为这个聊天界面绝对不同。我认为它允许我们实现的是,当你使用 Claude Code 时,假设你在做智能体式开发,你有一个概念,就是它相当单玩家。所以,你经常进入你的 Claude Code 实例,你可能在本地运行代码,然后启动单个会话。我们确实有这个边界,即会话的概念,在那里你陈述你试图实现的目标,管理上下文,然后你有一些结果、目标或循环在运行。而有了 Claude Tag,一切都变得有点模糊,所以我们开始不那么以会话为中心,也不那么以单玩家为中心。所以,当你用 Claude Tag 启动这些智能体式编码或智能体流程时,你看到的是它从一开始就更加多人协作。所以,Tag 可以获取工作流可能需要的人的意见。

Yeah, so I think there's a a couple of things. And just to just to frame it in terms of the impact that we've seen internally, since we've started using Claude Tag internally, our product teams like like product engineering teams have 65% of their PRs opened by Claude Tag. Wow. So, that's kind of the the scale of which like we obviously were the early adopters of Claude Code. And now that's like often the the port of call for how we would kick off some of these workflows and how those things originate. So, yeah, that that's how I'd frame it in terms of like where where we're at today. In terms of workflow, yeah, I think this chat interface is definitely different. I think what it allows um what it allows us to achieve is when you're using Claude Code, so let's say you're and you know, doing agentic development, you have these this concept of it's it's quite single player. So, you're often kind of moving into your Claude Code instance, you're potentially like running code code locally, and you're kicking off individual sessions. Like we do have this perimeter which is a concept of a session, and that's where you're kind of stating what you're trying to achieve, managing the context for that, and then you have some kind of like outcomes or or goals or loops that you're running. What we have with Claude Tag is it all becomes a little bit more amorphous, so it we're starting to like be less session focused and it's less single-player focused. So, when you're kicking off these agentic coding um or just agentic processes with Claude Tag, what you're seeing is that it's so much more multiplayer from the get-go. So, Tag can like get the opinions from people that might be needed for a for a workflow.

Claude Tag vs Claude Code:异步与同步 Claude Tag vs Claude Code: Asynchronous vs Synchronous

Lamis Mukta

所以,与其让我作为开发者去联系产品人员或销售人员,我可以把它嵌入到 Claude Tag 的流程中。而且,所有这些工作最终都可能公开进行,这意味着那些人从一开始就能看到这些工作流。他们确切知道工程师给出的产品规格,并且可以在需要时随时补充额外上下文等。也许他们自己的 Claude Tag 也会把这些事情推送给他们。所以,这有一种协作性质,非常不同。另一件事是,它比你在 Claude Code 中实现的要异步得多。所以,我认为使用 Claude Code 时,你几乎要参与智能体的每一步。你希望它回来向你报告它做了什么、需要跟进什么、更新什么等。当你想要更紧密地引导它时,这很好。而使用 Claude Tag,我们看到更多这种长时间运行的异步任务模式,这充分利用了最近模型迭代解锁的一些能力。所以,你看到的是,你触发 Claude Tag,它可能在几小时后回来,已经端到端地构建了你提到的功能。所以,你基本上可以让智能体自己工作,它会在完成时回来找你,这与 Claude Code 非常不同。我的意思是,它受到了我们在 Claude Code 中发布的一些功能的启发,但这就是我认为变化的地方。

So, rather than me as the developer having to go and ping some product folks or sales folks on something, I can get that embedded in the Claude Tag flow. Also, ultimately all of this work is potentially happening in public, which means that those people have visibility on those workflows from the get-go. They know exactly what kind of product specs the engineer is giving and can always chip in if they want to with extra context, etc. And maybe their own Claude Tags are going to ping these things over to them. So, there's this whole collaborative nature, which is really different. The other thing is it's a lot more asynchronous than what you'd achieve with Claude Code. So, I'd say that with Claude Code, you're kind of there for every single turn of the agent. You want it to come back to you with exactly what it did, anything that needs follow-up, updates, etc. And when you really want to steer that thing more closely, that's a great thing to do. With Claude Tag, we're seeing more of a pattern of these kind of long-running asynchronous tasks, and that's really leaning into some of the model capabilities that have been unlocked with recent iterations of models. So, what you see is you kind of ping Claude Tag and it comes back to you maybe a couple of hours later having built end-to-end this feature that you talked about. So, you're kind of able to just leave the agent to itself and it will come back to you when something's done, which is quite different, I think, to Claude Code. I mean, it's inspired by some features that we've released in Claude Code, but that would be how I think things have changed.

Host

行业的发展方式很有趣。大约一年前,我和 Slack 的一位产品副总裁聊过。我当时想,Slack 会成为开发者的新 IDE 吗?当我们思考这种演变时,这非常有趣。开发者显然热爱他们的 IDE,它已成为他们最高效的工作场所。然后,当我们开始使用各种副驾驶工具,将智能体式辅助引入 IDE 时,这逐渐演变得越来越多,支持多文件更改,像 Cursor 这样的工具也出现了。但一旦 Claude Code 真正进入终端 IDE,它就改变了人们的工作方式。我认为这几乎是下一个转变,它确实需要信任。我很想谈谈信任。要真正离开 IDE,因为你不再专注于代码,这确实需要信任。我认为这将是未来,我们达到那种甚至不需要看代码的信任水平。但今天,我认为有一个重大转变:一旦我们进入终端 IDE,我们实际上更依赖测试和验证,而不是查看代码或代码审查的结果。一旦这扩展到 Slack 或聊天环境,我们几乎又向远离代码抽象了一层。我认为像 Claude Tag 或 Slack 内的任何工具,一年前还不会被接受,对吧?但现在我们更加信任 AI 能做正确的事,而且 AI 生成的代码和智能体式代码的结果确实可靠,以至于我们能够从更远离代码中心的环境中使用智能体。Slack 和聊天是新的 IDE 吗?

And the way the industry is evolving, it's funny actually, about a year ago or so, I was having a chat with one of the VPs of product, I believe, at Slack. And I mused, is Slack becoming the new IDE of a developer? And it's very interesting when we think about the evolution. Developers obviously loving their IDEs, it's becoming the most productive place that they're in. Then all of a sudden when we started using, with the co-pilots of the world, where we started introducing agentic assistance into the IDE. That gradually became more and more evolved with multi-file changes and things like Cursor coming in. But then as soon as Claude Code really hit the terminal IDE, it changed people's way of working. And I see this as almost the next shift where it really requires trust. I'd love to talk about trust for a little while. It really requires trust to actually step away from the IDE because you're not focused on code, which I think will be the future, the space whereby we get to that level of trust where we don't even need to look at the code. But today, I think there's this big shift: as soon as we're in the terminal IDE, we're actually relying much more on the tests and the validations versus looking at the code or the results from a code review. As soon as that then gets extended beyond and into Slack or a chat environment, we're almost abstracting away from the code one stage further. I think a tool like a Claude Tag or anything from within Slack, this wouldn't have been accepted a year ago, right? But we're so much more trusting of AI doing the right thing, and actually the results of AI generated code and agentic code being reliable that we're able to use agents from a space which is further abstracted away from a very code-centric environment. Is Slack and chat the new IDE?

Lamis Mukta

这是一个很好的问题。我想很难说。我不认为这些是直接的替代品。我认为每个都有其位置和目的。但我非常认同你提到的趋势,即我们的行为、我们与智能体互动的方式正在改变,原因有几个。所以,我认为一个是关于信任这一点,模型正变得越来越好。我们正处于智能体运行时间的指数增长趋势中。所以,Meter 图表总是显示,大约每 4 个月,智能体能够自主运行的时间就会翻倍。因此,从纯粹的能力角度来看,我们能够更信任这些智能体,因为实际能力在提升。

This is a great question. I guess it's hard to say. I don't think these are like-for-like replacements. I think each has their place and their purpose. But I really resonate with this trend that you mentioned in terms of how our behavior, the way that we interact with agents, is changing for a number of reasons. So, I think one is that to this trust point, models are just getting better and better. We are squarely on this exponential trend in terms of how long agents can run for. So, the meter chart always shows us that roughly every 4 months the amount of time that agents are able to run for autonomously is like doubling. And so from a sheer capability perspective, we're able to trust these agents more because of the actual capability that's improving.

Host

我想深入探讨你刚才说的。每 4 个月,智能体能够更自主地做某事的时间就会翻倍。这是模型的能力,还是因为也有人类方面的因素?比如,如果它去执行任务,我能与它交互吗?是模型和智能体的纯粹能力解锁了这一点吗?

I'd love to just double down on that what you just said. Every 4 months the time that an agent can do something more autonomously is doubling. Is that the capability of that model or is that almost because there's a human aspect to that as well? In terms of if it goes off and does something, am I able to interact with it? Is it pure capability of the model and the agent that's unlocking that?

Lamis Mukta

这是一个有趣的问题。Meter 做了这项研究,涵盖了多个不同领域,本质上是时间跨度。也就是说,智能体能够成功完成特定长度任务的时间,是能力的一种度量。它并不完美。我们还有其他评估和基准来衡量模型能力,但这是一个普遍广泛的指标,确实反映了与这些模型交互的感觉。所以,它们能够完成任务的时间越长,往往反映出它们在处理更复杂的事情。它们能够执行这些多步骤任务,能够在多个阶段验证自己的工作结果,然后回来成功完成这些任务。所以,这是我们看到的一个行业广泛趋势。如果你看它,相当惊人,因为每次你觉得自己跟不上那条漂亮的对数直线图时,你每次都跟上了。这种情况已经持续了近十年。所以这非常有趣。我认为当我们谈论模型能力时,显然运行时间是一方面。但随着时间的推移,我们看到很多以前我们放在束缚中的行为正在被嵌入到模型中。具体来说,现在的模型在验证自己工作方面要好得多。所以,从本能的角度来看,它们会在回来告诉你完成之前检查自己的工作。

So this is an interesting one. Meter produced this research and it covers a couple of different domains and essentially time horizons. So that's how long agents are able to successfully complete a task of a certain length, is a measure of capability. It's not perfect. There are other evals and benchmarks that we use to measure model capability, but it's one generally broad one that has really mirrored what it feels like to interact with these models. So the longer that they're able to complete these tasks for, it often is mirroring the fact that they're doing more complicated things. They're able to do these multi-step tasks, they're able to verify the results of their work at multiple stages and come back and successfully complete these tasks. So this is one kind of industry broad trend that we see. And if you look at it, it's quite shocking because every time you think that you're not going to keep up with that nice log chart on the straight line, you just do every time. And that's been happening for the best part of a decade. So this is really interesting. And I think when we talk about model capability, obviously the amount of time it runs for is one thing. But what we see over time is a lot of behaviors which previously we'd kind of put in the harness are getting embedded into the models. So to make that concrete, models are a lot better at verifying their own work these days. So both from an instinctual perspective, they do just check over their own work before coming back to you and saying it's complete.

模型自我验证工作 Models verifying their own work

Lamis Mukta

而且当模型拥有验证自己工作的工具时——无论是通过前端测试、运行测试还是创建自己的评估——它们在这方面也变得越来越好。所以回到你关于信任的观点,我认为这些是相辅相成的。你给模型提供验证自己工作的工具,它们就能更好地自我验证。我认为在当今的开发中,我们还看到另一个模式:开发者或工程师的行为正在更多地转向如何为你的用例定义成功的样子。这不仅适用于编码,实际上适用于一切。比如,你是否对成功的样子有很好的理解?如果有,就把它交给模型,模型可以循环迭代,或者让另一个智能体审查它的工作,直到审查者认为任务完成。所以我认为这是两件事。在智能体和模型方面,它们只是变得越来越有能力,越来越擅长做这些事情。而在人类方面,我们的行为正在更多地转向:我们能否定义什么是好的?另外,我想补充一点,对于我们很多人来说,在过去一年里真正深入使用这些工具后,我们越来越清楚这些工具在哪些地方真正有用,哪些地方需要更多监督。我们可以调整自己的行为和输入,以便更好地让这些工具协同工作。

And where they have the tools to verify their own work, whether that's through like front end test or running tests or creating their own evals, they're getting a lot better at doing that themselves as well. And so to your trust point, I think these things go hand in hand. Like you give the models the tools to verify their own work and they are better at kind of doing that themselves. And I think another pattern that we see in development today is that the behavior of the developer or the engineer is shifting much more towards how well can you define what success looks like in your case. And that's not just for coding, that's really for everything. Like do you have a really good sense of what success looks like? And if so, hand it to the model and it can loop over or have another agent review its work until they both kind of until the reviewer kind of believes that that thing is complete. So I think these are the two things. It's like on the agent and model side, they're just getting more capable and better at doing these things. And then on the human side, like we are getting our behavior shifting more to just like can we define what good looks like? And I think another thing to just like round that off is that having really like for a lot of us having used these tools in anger for like the last year, we're getting a really good sense of like where are these tools really useful? Where is it that they need a bit more supervision? And we can kind of tune our behaviors and inputs more to like what makes those things work really well in tandem.

Host

嗯。

Mhm.

Slack作为智能体界面 Slack as a surface for agents

Lamis Mukta

所以关于你的问题:Slack 是新的家还是新的 IDE?我认为它是一个界面,能让你的智能体更贴近你一直在工作的地方。它允许你在任何地方标记它们——任何时候你需要更多上下文或想委派任务时,你都可以直接把智能体拉进来。有时是编码任务,有时是构建某个功能或仪表盘,有时只是问这个缩写是什么意思,或者告诉我上周发生了什么之类的事情。所以我认为它提供了更多的灵活性。你不需要在“与我的智能体协作”和“与我的团队协作”之间切换上下文。一切都更加集成。

So to your point on like is Slack the new home or is Slack the new IDE? I think it's a surface that allows your agents to be closer to where you're doing your work all the time anyway. It allows you to tag them in places like anywhere really like anytime you need more context or you want to delegate something, you're able to just loop the agent in and sometimes that is a coding task, sometimes it's like build this feature or this dashboard, sometimes it's just like what does this acronym stand for like can you please just tell me what happened last week or something like that. So I think it just allows a bit more flexibility. You're not kind of switching context between working with my agents and working with my team. It's all a bit more integrated.

聊天与IDE中定义“好”的标准 Defining what good looks like in chat vs IDE

Host

你提到定义“好”的样子,这很有意思。作为人类,我们总是想为手头的任务使用最好的工具、最好的场所。在定义代码中“好”的样子时,当然,我们想描述测试、编写测试用例并构建它们。很多时候我们会依赖 IDE 来实际构建。但当我们要协作定义“好”的样子时,我们自然发现自己处于聊天环境中,而这是定义它的正确场所。所以在这个阶段标记某物,给它关于“好”的上下文,非常重要。让我们稍微退一步,更广泛地看看。因为我认为那些高级用户会看到像 Claude Code 这样的工具,觉得这正是他们需要的,想立即引入。今天在特斯拉办公室,我们有一个黑客马拉松正在启动,现场大约有 100 到 150 人。现有的采用情况非常广泛。有些人对智能体式编码还比较陌生,其他人显然已经使用了很长时间。我想问你关于 Claude Code 的外部采用情况,从更广泛的行业来看。从纯粹的智能体式编码和开发的角度来看,你认为行业在使用 Claude Code 方面处于什么成熟阶段?显然,我们稍后会讨论 Anthropic 内部的采用,但外部来说,人们今天主要如何使用 Claude Code?

And it's really interesting when you mention defining what good looks like. When we as humans, we always want to use the best tools, the best place for the best task at hand. And in the case of defining what good looks like when we focus on code, yes absolutely, we want to describe tests, we want to write some test cases and we want to build them out. And a lot of the time we'd lean into the IDE to actually build that out. But when we want to collaboratively define what good looks like, we naturally find ourselves in a chat environment and it is the right place to define it. And so tagging something in at that stage, giving it the context of what good looks like is super important. Let's step a little bit back and look more broadly because I think people who are the power users will look at something like Claude Code and think this is exactly what I need. I want to bring this in immediately. So here actually in the Tesla office today we have a hackathon kicking off and there's like 100, 150 folks out there. Very broad sets of existing adoption. Some people are more new to agentic coding, others have been using it obviously for a long time. I'd love to ask you about the external adoption of Claude Code from, you know, obviously the wider industry. From a purely agentic coding and development point of view, at what stage of maturity would you say is the industry at in terms of using Claude Code? Obviously, we'll talk about Anthropic's adoption in the second, but externally, how are people mostly using Claude Code today?

Claude Code的外部采用与规模 External adoption and scale of Claude Code

Lamis Mukta

当然。我认为我们看到各种形态和形式:对于单个工作,能够更快地以更高质量完成。这是我们在个人层面看到的。当我们将此扩展到团队和组织时,我们看到了一些非常惊人的事情。比如,我们看到像 Stripe 这样的组织,用几天或几小时完成了原本需要数周或数月的整个代码库重写。这就是当我们大规模部署这些工具并且每个人都参与进来时,我们所谈论的规模。这些雄心勃勃的项目,你可能总是推迟数月或数年,因为找不到资源,现在终于可行了,这让人们能够专注于产品和工程工作的其他部分。我们确实一直看到这种情况,比如其他团队在一个月内交付了数百万行代码。所以我认为当每个人都真正投入时,这就是我们能够看到的开发规模。另一件事是,你开始看到更多团队追求这些登月项目,否则他们根本没有时间、资源或能力。在产品方面,我们看到团队基本上对几个不同的想法或方法进行原型设计,用智能体工具或内部测试一堆,然后全力投入他们看到效果最好的那个。所以,实验的空间更多了,而且我认为在产品开发的方式上,也更多了一些勇气和大胆。

Absolutely. And I think we see this in all sorts of shapes and flavors and forms from being able to, for individual pieces of work, complete that to a higher degree of quality faster. That's something we see on the individual level. And when we scale this to teams and across organizations, we've seen some pretty phenomenal things. Like, we've seen organizations, for example, like Stripe, do entire codebase rewrites that would have taken weeks or months in like days or hours. So, this is the kind of scale we're talking about when we really deploy these things at scale and everyone's kind of on board. These kind of ambitious projects that maybe you just would always put off for months or years because it's just where do you find the resource are finally doable and that's allowing people to focus on other parts of product and engineering work. And we yeah, we really see this consistently like other teams shipping millions of lines of code in just a month or something like that. So, I think when everyone kind of really leans in, that's the scale at which we're able to see development happening. And I think the other thing is you start to see more teams pursuing these kinds of moonshot projects that they just wouldn't have time, resources, or capacity to otherwise. Like on the product side, something that we see is teams basically prototyping a couple of different ideas or approaches to something, testing a bunch of them either with agentic tools or internally, and then just going all in on the one that they see work best. So, there's a bit more room for experimentation and I guess a bit more bravery and boldness in the way that you're approaching product development.

自建与购买及人类适应速度 Build vs buy and human adaptation speed

Host

当然,构建与购买的问题就变得非常有趣了,因为像 Claude 这样的工具会让快速构建变得非常便宜。你可以把一个想法或原型非常快速地变成一个实际可用的应用程序。我想问题在于,团队是否想要长期维护它。而且我认为当我们思考这个行业中什么在快速变化时,人类和人们是适应最慢的。人们能跟上 AI 编码领域正在发生的快速变化和交付吗?

And of course the build versus buy question then becomes super interesting because tools like Claude will make it so much cheaper to build rapidly. You can take an idea or a prototype very quickly to an actual live working application. And I guess the question then is if teams want to continue maintaining that over time. And I guess when we think about what's changing quickly in this industry, it is humans and people that are slowest here in terms of adapting. Are people able to keep up with the rapid change and delivery that's happening in the AI coding space?

Lamis Mukta

这是个好问题。我最近在欧洲几个城市巡回,与一些创始人社区交流。我喜欢问在场每个人一个问题:谁有 FOMO(害怕错过),觉得自己在日常工作和生活中没有足够地使用 AI?

It's a great question. I've been on a bit of a tour around a couple of European cities recently talking to some founder communities. And one question I love to ask everyone in the room is like who here has FOMO that they're not using AI like they're not using they're not AI pulled enough in their day-to-day work and their life etc.

AI开发速度与基础设施挑战 Pace of AI development and infrastructure challenges

Host

而且每次都是满屋子的人举手。Anthropic 的员工也一样,我们都举手了。谁能跟上这个节奏?Karpathy 在他那条著名的推文里也说,他从未觉得自己如此跟不上某个东西。

And it's always just like a full room of hands. And the Anthropic employees as well, we all have our hands up. Like who can keep up with the pace of this? Karpathy famously in his tweet as well said he's never felt more out of touch with something.

Lamis Mukta

如果 Karpathy 都这么说,那我觉得归根结底,事情的规模和速度已经超出了一个人脑或个人的承受范围。我认为这已经不止是一份全职工作了。有时候我甚至发现一些我们自己都不知道的功能,因为谁能跟得上呢?但我们聊了很多关于这些东西发展的速度和节奏。还有另一个故事,对吧?那就是:这在多大程度上转化为对人们的实际影响?作为产品开发者或构建者,你如何将这种指数级增长映射到你能为客户提供的价值上?以及内部流程的影响。我认为这才是更难的问题,也是更难解决的难题。我们可以拥有所有这些原始智能,但我们有基础设施来让这些价值落地吗?这涵盖了很多方面。包括你的工具链,你如何管理上下文和记忆,如何管理工具,如何让这些智能体安全地访问所需的一切。你如何进行权限管理?我们在设计 Claude Tag 时深入思考过这个具体问题。然后还有所有其他方面。所有的基础设施。你如何托管和部署这些模型?你如何处理所有的推理?所以我认为,虽然有很多炫酷的事情发生,但这些才是大家应该真正聚焦的问题。我们总是努力开发工具来帮助人们真正获取这些价值,因为理论上它是存在的。我们能看到它。但为了确保人们真正感受到它,我认为还需要更多的努力。而且我认为很多开发者都有同感。

And if Karpathy's saying that, I just like ultimately the scale of things, the speed of things is like more than one human mind or person can keep up with. I mean it's more than a full-time job at this point I think. Sometimes I even discover features I didn't know that we had because who can keep up with that? But I think that we've talked a lot about the speed and the pace at which these things are developing. And there's another story as well, right? Which is: how much is this translating into actual impact for people? As product developers or builders, how are you mirroring that exponential in the value that you're able to deliver to your customers? And in the impact on processes internally. I think this is the much harder question, and this is the much harder problem to solve. We can have all this raw intelligence, but do we have the infrastructure to bring that value to life? And there's so much that covers. It covers your harnesses and how you're managing your context and your memories, how you're managing your tools, how you're giving these agents access to everything they need, and also in a secure way. How are you permissioning that? It's something we thought a lot about when we were designing Claude Tag, that specific problem. And then there's everything else. There's all the infrastructure. How are you going to host and deploy these models? How are you going to deal with all of the inference essentially. So I think that yes, whilst it's amazing that there are lots of shiny amazing things happening, those are the problems that everyone should be really laser-focused on. We always try to develop tools that help people really access that value because on paper it's there. We can see it. But in order to make sure that people really feel that, I think that's a lot more hard work. And I think it's something that a lot of developers experience.

Host

是的,非常有趣。你提到了记忆,我还想聊聊“做梦”(dreaming),这是你在我们伦敦的 AI Native DevCon 大会上谈到的话题。我们马上就来聊这个。但在此之前,我聊了一点外部社区和行业。现在我想聊聊 Anthropic 内部,以及 Anthropic 自身如何开发软件。

Yeah, very very interesting. And you mentioned memories, and I'd also love to chat a little bit about dreaming as well, which is something that you talked about at AI Native DevCon, our conference in London here. And I'd love to talk about that in just a second. But before we do, I talked a little bit about external communities and industry. I'd love to talk about internally at Anthropic now, and how Anthropic develops software itself.

Lamis Mukta

好的。

Yeah.

Host

然后我还想聊聊 Anthropic 在整个组织中如何使用 Claude Code 以及编码。首先,我们回到最初的那一天吧?Boris 在他地下室里捣鼓,你知道的,在构建一个叫 Claude Code 的东西。给我们讲讲那个故事。

And then I'd love to talk a little bit about how Anthropic uses Claude Code generally in the org as well as coding. So first of all, why don't we go back to day zero? Boris is playing around in his basement playing, you know, building this thing called Claude Code. Take us through that story.

Lamis Mukta

是的,这基本上是 Boris 在做的一个副业项目。首先,在 Anthropic 的文化中,有一种真正的实验文化。人们总是在构建自己的工具并进行实验,这就是 Boris 在做的事情。有趣的是,最初他在 Slack 上分享这个时,只得到了大约六个反应。这总是告诉你数据并不完美。你不可能有完美的流程来理解什么样的产品是好的等等。但有几个人看到了,非常兴奋,并继续推进。在很短的时间内,我们在公司内部看到了惊人的采用率,大约一半的员工每周都在使用它。另一件重要的事情是,我认为在这个领域开发产品时的一个关键原则是:一旦模型改进到足以让你长时间高效处理编码任务时,这个产品的采用率就真正起飞了。早期的迭代感觉智能体性弱得多。它们更像是直接返回代码块,而后来它真的可以开始访问不同类型的工具,在代码库上高效工作,并保持对大量上下文的把握和以目标为导向。所以我们总是对开发者说:要为模型未来的能力而构建,不要为它们今天的能力而构建,因为这些东西发展得太快了。

Yeah, so this is very much like a side project that Boris was working on. And I think first of all, culturally at Anthropic, there's this real experimental culture. People are always building their own tooling and experimenting with things, and this was something that Boris was working on. And the funny thing is that originally, I think when he shared this in a Slack post, it got like six reactions. Which always tells you that data isn't perfect. You can't have perfect processes for understanding what good products look like, etc. But a couple of people saw this and were really excited by it and continued to work on it. And over a short period of time, we saw amazing adoption within the company, like half of the company using this weekly. Another thing that's important here, and I think it's a really key principle when developing products in this space, is that we saw a real takeoff in the adoption of this product once the models improved a little bit more to make it really achievable to work on these coding tasks for long periods of time. So early iterations felt a lot less agentic. They were more similar to just getting chunks of code back, whereas later it really could start to access different kinds of tools, work really efficiently over the code base, and stay on top of a lot of context and stay goal-oriented. So one of the big things that we always say to people when they're developing is: build for where the models are going to be in the future. Don't build for where they are today because as we've said, these things move so quickly.

Host

非常有趣。这里有几个点我想展开聊聊。

Super interesting. And there's a couple of things I want to unpack here.

Lamis Mukta

当然。

Of course.

Host

那我们首先来聊聊“吃自家狗粮”(dog fooding)。因为我认为这是 Anthropic 内部一个真正的“吃自家狗粮”的成功案例,对吧?有记录显示 Claude Code 在内部非常受欢迎,以及它是如何突然被意识到:如果我们的工程师如此广泛地使用它并从中获得如此大的价值,那么这绝对是我们需要产品化的东西。谈谈 Anthropic 是如何知道这非常有价值,并希望与更广泛的受众分享的。

So let's jump into the dog fooding first of all. Because I think this was a real dog fooding success story within Anthropic, right? It was documented how popular Claude Code was internally, and how it was realized all of a sudden that actually this is a huge thing that if our engineers are using this so broadly and getting so much value out of it, this is something we absolutely need to productize. Talk about how Anthropic knew this was super valuable and wanting to share this with the broader audience.

Lamis Mukta

是的,我认为这又回到了那个想法:当时大约一半的团队每周都在使用 Claude Code,这对一个新产品来说相当疯狂。我认为这种内部的产品市场契合感(PMF)让我们意识到是时候更广泛地发布这个产品了。Claude Tag 也是同样的故事。就像我说的,在发布之前,我们 65% 的 PR 都是由 Claude Tag 提交的。我认为到了一个点,感觉这个产品真正引领了我们今天看到的模型能力,它现在是我们日常使用的核心工具。其中一个重要模式是 Claude Code 的情况:最初是所有工程师依赖它来快速交付大量代码。然后我们看到一个模式,Anthropic 的所有团队都完全投入到了 Claude Code 中。有一个疯狂的故事:市场团队有个人,一天开始时还在谷歌搜索什么是终端以及如何使用它。到一天结束时,他们已经自动化了一个原本需要 30 分钟的工作流程,现在只需要 30 秒。

Yeah, so I think it comes back to this idea that I think like half of the team at that time were just using Claude Code every week, which is quite crazy for a new product. I think that sense of internal PMF really made us realize that it was time to release this product more broadly. And it's really the same story with Claude Tag. Like I said, prior to the release 65% of our PRs had been raised by Claude Tag. And I think there comes a point where it feels like this product is really leading into where we see model capabilities today and it's the one that's really our daily driver at this point. And I think one of the big patterns there is what happened with Claude Code: at first it was all of the engineers who were relying on this to ship a ton of code really quickly. And then we saw this pattern where all of the teams at Anthropic were totally going all in on Claude Code. There's this crazy story of someone on the marketing team whose day started with Googling what the terminal was and how to use it. And by the end of the day they'd automated one of their workflows which took them 30 minutes and now it took 30 seconds.

编码之外的智能体工具 Agentic tools beyond coding

Lamis Mukta

就像他们能在那么短的时间内制作出这些广告。所以我认为我们看到的是,每个人都看到了这些工具的力量,并且能够创造性地找到方法将其映射到非编码领域的工作流程中。也许验证和上下文的问题更难一些。你没有像整洁的文件系统、GitHub 来管理版本控制,也没有单元测试的能力。你必须更有创造力,但我认为这就是我们前进的方向,对吧?我们能够设定结果和成功标准,或者为好的文档或好的简报制定评分标准。人们正在改变他们创建和生成数据以及存储数据的行为,以便智能体能够更容易地访问这些数据。所以,在 Anthropic,我们有一种巨大的文化,就是在 Slack 上非常公开地工作。我们故意这样做,因为这意味着我们的智能体能够以任何人都无法拥有的全局视野来连接各个点。有时我甚至会在公共频道中与 Claude Tag 交谈。我与 Claude Tag 的所有工作,除非是非常私密的事情,我都会在公共频道中进行。这意味着有时我团队中或公司里从未见过我的人会给我发消息说:‘我看到你在做这件事。我很想用它。请告诉我它是否可以共享,我们可以合作吗?’所以,我认为能够以这种规模连接一个组织,只有通过我们拥有的这些工具才能实现。

Like they were able to produce these ads in that short period of time. And so I think what we've seen is everyone is seeing the power of these tools and is able to creatively find ways to map that to their workflows in different domains that are not coding. Maybe that problem of verification and context is a bit harder. You're not set up so well with neat file systems, GitHub to manage your version control and the ability to unit test things. You have to be a bit more creative, but I think that's where we're all headed, right? We're able to set out outcomes and success criteria or rubrics for what a good document looks like or what a good briefing looks like. People are changing their behavior around how they create and produce data and where they store it so that agents are able to more easily access that. So, we have this huge culture at Anthropic where we work really, really publicly in Slack. And we do that on purpose because it means that our agents can connect the dots in ways that no person could ever have the visibility over. Sometimes I work to the extent where I talk to Claude Tag in a public channel. All of my work with Claude Tag, unless it's something really private, I do in a public channel. And it means sometimes people on my team or people in my company who I've never met will message me being like, 'I saw you working on this thing. I'd love to use it. Please can you tell me if that's shareable and can we collaborate on this kind of thing?' So, I think the ability to connect an organization at that scale is only possible because of the kinds of tools we have.

Host

非常有趣。这实际上引起了很大的共鸣,因为即使是我们的法务团队,例如,他们使用 Claude Code 构建应用、添加技能,并将技能检查到 Tesla 注册表中。令人惊讶的是,这不仅赋能了工程团队,还赋能了整个组织,我们稍后会深入探讨。我想问一个问题,当围绕 Claude Tag 和早期的 Claude Code 有如此多的内部试用文化时,产品方向在多大程度上是由你们的内部反馈和内部试用驱动的?

Super interesting. It actually resonates a lot because even our legal team, for example, builds apps using Claude Code, adds some skills, checks our skills into the Tesla registry. And it's amazing how much that's empowering not just the engineering community but the whole organization, which we'll touch on in a little more depth. I'd love to ask the question about, you know, when there's so much of a dogfooding culture around Claude Tag, around Claude Code in the early days as well, how much is the product direction driven by your internal feedback and the internal dogfooding?

Lamis Mukta

是的,这是个很好的观点。我认为这非常重要,但这里有几件事需要考虑,因为在智能体时代的产品开发中,有些东西表面上看起来是个绝妙的主意,单人体验感觉也很棒。但当你真正考虑如何将其扩展到企业时,问题的形态就完全不同了。所以,就与 Claude Tag 的交互而言,我认为我们很早就知道这是非常有效的。

Yeah, that's a really good point. I think this is really important but there are a couple of things to think about here because quite often with product development in the agentic era, there are things that on the surface look like an amazing idea and feel like an amazing experience single player. But when you really think about what it takes to scale that thing to the enterprise, it's a really different shape of problems. So, in terms of the interaction with Claude Tag, I think we all knew really early that this was something that was working really effectively.

Host

嗯。

Mhm.

Lamis Mukta

我们也意识到,在 Anthropic,我们彼此分享信息的方式相当自由。显然,对于团队严格私密的信息,我们有非常强的防护栏,但我们在 Slack 工作区等地方都设定了这些。我们有非常好的权限防护栏。这意味着,在那些你知道是可信的信息共享空间里,人们可以非常开放,而这正是我们的智能体能够表现良好的原因。所以,我们在 Claude Tag 上的一个设计原则是,我们非常仔细地设计了如何为每个频道授权。每个频道、工作区和频道都有自己的权限范围,包括可以访问哪些工具、拥有哪些不同服务和连接器的 API 密钥,以及可能可以访问哪些其他频道。所以,我认为显然在某种程度上,我们想分享哪些文化实践让我们能够与智能体良好协作。其中之一就是公开工作。但同时,要确保我们的产品内置了防护栏,这样你就可以合理地实现这种行为,并且能够真正扩展到企业。我们认为这非常有效。我们为了让这更有效而做的另一件事是提出了智能体身份的概念。所以,Tag 与你使用 Claude Code 或 Co-work 的一个非常大的区别是,当你使用 Claude Code 或 Co-work 时,它们会假定使用你自己的权限。它们会使用我的 API 密钥或我的权限系统,而我授予对这些东西的访问权限。而对于 Claude Tag,我们实际上给了那个智能体自己的权限和自己的密钥等,这样它就可以自主地去处理这些事情。它不是代表某个人工作,而是代表团队工作。

We're also aware that at Anthropic, we're pretty liberal with the way we share information with each other. Obviously, there are very strong guardrails for what is strictly private information to a team, but we have all of that set out in our Slack workspaces etc. We have really good guardrails for the permissioning. And what that means is that within those spaces where you know it's a trusted place to share information, people can be really open, and that's what allows our agents to perform really well. So, one of the design principles we had with Claude Tag is that we've really carefully designed how you permission each channel. Each channel, the workspaces and the channels have their own permission scopes in terms of what tools they can access, what API keys they have for different services and connectors, and potentially what other channels they can access. So, I think obviously on one level, we want to share what cultural practices are allowing us to work really well with agents. So, one of these is working in public. But at the same time, make sure that our products come with the guardrails baked in so that you can reasonably achieve this behavior in a way that can actually scale to an enterprise. We think that works really well. Another thing we did to make this work more effectively is we came up with this concept of agent identities. So, one really big difference between Tag and you working with Claude Code or Co-work is when you work with Claude Code or Co-work, they kind of assume your own permissions. They'll work using my API keys or my permission systems, and I grant access to all those things. Whereas with Claude Tag, we actually give that agent its own permissions and its own keys etc. so that it can go off and work autonomously on these things. It's not working on behalf of one individual, it's working on behalf of the team.

Host

嗯。

Mhm.

Lamis Mukta

而且审计起来也容易得多。它不是以你的身份行事,而是以它自己的身份行事。所以,这是我们需要做的另一个关键架构变更,以实现这种多人协作行为。所以,是的,总结一下,肯定有很多产品层面的反馈,所有团队都会参与进来,确保它真正适用于不同类型的用例和用户。但同时,我们认为很多工作都是为了确保这能够真正扩展到企业,让人们真正从中获得价值。

And it's much easier to audit that. It's not doing this as you, it's doing this as itself. And so that's another key architectural change that we needed to do to enable this multiplayer behavior. So yeah, to round that point off, there's definitely a lot of the product level feedback that all of the teams will chip in with and make sure that it really works for different kinds of use cases and different types of users. But at the same time, we think a lot of the work goes into making sure that this is something that actually scales to enterprises and people can really get value out of.

Host

太棒了。我们来谈谈,我知道我们的听众以及整个行业,他们改进和变得更好的方式是通过理解和倾听,不仅是我们做对了什么,还有我们如何绊倒、如何跌倒、不得不爬起来并尝试寻找另一条路。我相信 Anthropic 和每个组织一样,都有尝试过但不太顺利的领域。所以,我想从采用的角度或从与智能体编码协作的方式来看,你们或 Anthropic 在使用 Claude Code 甚至 Claude Tag 的过程中,最大的收获是什么?

Amazing. And let's talk, I know our audience as well as the industry, the way they improve, the way they get better is through understanding and hearing what not just what worked for us, but also how we tripped over, how we fell and had to get up and try and find another path. And I'm sure Anthropic just like every organization have areas that they tried and didn't get on with. So I guess from an adoption point of view or from a ways of working with agentic coding, what were some of your or Anthropic's greatest learnings would you say in the way you were using Claude Code, the way you were using maybe even Claude Tag as well?

Lamis Mukta

当然。所以我认为有几个。有些是在开发方面,有些是在行为方面。

Yeah, of course. So I think there are a few. There are some on the development side and some on the behavioral side.

为未来模型构建并简化框架 Building for future models and simplifying harnesses

Lamis Mukta

在开发方面,我之前说过一点:我们应该始终为模型未来的发展方向而构建,而不是为它们当前的能力。在 Claude Code 这边,我们经常思考的是,每次有新模型时,我们都会讨论这些模型在哪些维度上变得更强大。这意味着我们会非常定期地重新审视那些“脚手架”长什么样。我们很乐意随着时间推移从中删除一些东西,让它变得更简单、更轻量,让模型自己承担繁重的工作。所以随着时间的推移,脚手架实际上会变小,因为我们在某些能力上更信任模型了,只需要给它所需的工具使用和基础设施。这就是一点。新模型并不意味着塞进更多的提示词和更多的架构。有时候少即是多。

I think on the development side, one thing I said before is we should always build for where we think the models are headed, not where they are today. And one thing we think about a lot on the Claude Code side is every time we have a new model, we discuss how those models become more capable in certain dimensions. That means we very regularly revisit what those harnesses look like. And we're very happy to delete stuff from that harness over time to make it simpler and more lightweight, and just let the model do the heavy lifting. So over time, we see something like the harness actually gets smaller because we can trust the model more with certain capabilities, and we just need to give it the tool use and infrastructure it needs. So that's one thing. A new model doesn't mean chuck in way more prompts and way more architecture. Sometimes it means less is more.

Lamis Mukta

另一件事是回顾我之前给你们做的关于“做梦”的演讲。我们尤其在与初创公司合作时,很多人问我关于记忆和上下文基础设施的问题。这没有一刀切的解决方案。人们会想出非常有创意的方式来构建他们的记忆数据库或记忆结构。我们在托管智能体 API 中的解决方案是一个非常简单的记忆文件系统,它只是依赖智能体读写记忆的能力。我在那次演讲中提到,我们过去尝试过很多不同的东西,比如索引记忆存储,或者非常具体规定如何读写记忆的工具。我认为随着时间推移,我们发现这有很多问题。有时我们对智能体应该如何与记忆交互过于固执己见,而随着它们能力变强,最好让它们自己管理。它们非常擅长直接使用文件系统和原生的 bash 和 grep 工具。所以我们学到的一点是,我们实际上可以移除一些抽象层。即使我们对那些记忆结构本身的结构有看法,我们后来也意识到索引并非普遍的最佳实践,简单的文件系统更好。显然,这些都需要通过尝试来学习,而且所有这些都是非常开放的研究和开发领域。我相信未来我们会找到更多最佳实践。但这些是几个例子,说明我们尝试了一些东西并简化了工作流程。

I think another thing is to reference back to that talk I did for you guys on dreaming. We, especially working with startups, I get a lot of people asking me about memory and context infrastructure in particular. It's not a one-size-fits-all solution. People come up with really innovative ways to structure their memory databases or memory structures. The solution we have in our managed agents API is a very simple memory file system that just leans on agents' abilities to read and write to memory. Something I touched on in that talk is that we tried a lot of different things in the past, like indexed memory stores or tools that were very specific about how to read and write to memory. I think what we learned over time was that this had a number of problems. Sometimes we were being too opinionated about how the agents should interact with memory, and they were better left, especially as they became more capable, to just manage that themselves. They were great at just using file systems and the native bash and grep tools. So one thing we learned was we could actually remove some of these abstractions. Even our being opinionated on the structure of those memory structures, we realized over time that indexing wasn't something we thought was best practice across the board, and we thought a simple file system was better. Obviously you have to learn these things by trying, and all of these things are very open areas of research and development. I'm sure we'll find more best practices down the line. But these are a couple of examples where we've tried a few things out and simplified our workflows a bit.

Host

这真的很有意思。我很好奇上下文这部分。或者抱歉,不是上下文,而是智能体模型的变化。无论是智能体的变化还是模型的变化,它确实会影响我们实际需要提供什么,无论是上下文还是记忆,才能让它发挥最佳性能。

It's really interesting. I'm really curious about the context piece. Or not sorry, the context piece, but the agentic model changes. Whether it's the agent change or the model change, it really does affect what we actually need to provide it, whether that's context or memory, in order for it to perform the best it can.

Lamis Mukta

是的。

Yeah.

Host

而我觉得最有趣的是,我们不需要为了重新运行评估而改变代码或上下文,来看:当前状态下它是否仍然有价值?或者因为模型变化或智能体升级,我是否实际上需要提供更少的上下文,因为模型或智能体已经变得更擅长在没有上下文的情况下做这些事?结果,我添加这个技能是不是只是在膨胀上下文?或者我需要为这个模型改变技能?我认为正是这种对我们环境的持续评估,我们需要定期进行,来问:我需要改变什么?是上下文吗?是我的提示词吗?是因为智能体变化而需要改变脚手架吗?

And I think what's most interesting is we don't need to change our code or our context in order for us to need to rerun an eval to see: is this actually still valuable in its current state? Or because of a model change or an agent upgrade, do I actually need to provide it with less context because the model or the agent has actually got better at doing these things without the context? And as a result, am I just bloating context by adding this skill? Or do I need to change the skill for this model? I think it's that continuous evaluation of our environment that we need to do on a regular basis to say: what do I need to change? Is it the context? Is it my prompt? Is it my harness because of an agent change?

Lamis Mukta

是的。这听起来在 Anthropic 内部非常普遍。

Yeah. And that sounds like something that's very commonplace then within Anthropic.

Host

是的,所以我认为当我们测试这些新模型时,尤其是在应用 AI 方面,因为我们与客户紧密合作,在测试的早期阶段,我们会特别关注哪些行为变化需要我们围绕提示词进行调整,以及哪些领域我们现在可以更放松一些,因为模型更好了。所以我们总会发布一些关于使用这个新模型的最佳实践指南,并帮助客户进行迁移。是的,我们总会在模型发布时发布很多有用的资源,帮助人们确保他们能非常轻松地迁移过来。

Yeah, so I think when we're testing these new models, especially on the applied AI side because we're working closely with customers, in our early stages of testing we really look out for what are these changes in behavior that we need to prompt around, and where are the areas that we can be a bit more relaxed about now because the model is just better. So we'll always come out with some guidance on what are the best practices for working with this new model and help customers with those migrations as well. So yeah, there's a lot of helpful resources that we'll always publish around model releases to help people make sure they can move over really easily.

Host

大家好。希望你们到目前为止喜欢这期节目。我们的团队在幕后非常努力地为您带来最好的嘉宾,这样我们就能进行关于智能体开发的最有见地的对话。无论是谈论最新工具、最高效的工作流程,还是定义最佳实践。但是,出于某种原因,你们中很多人还没有订阅这个频道。如果您喜欢这个播客,希望我们继续为您带来最好的内容,请帮我们一个忙,点击订阅按钮。这真的会带来不同,让我们能够继续提高嘉宾的质量,并为您打造更好的产品。好了,回到节目。我们来聊聊你之前简短提到的,比如营销团队创建了一个应用,这很棒。我认为 Claude tag 确实让这变得更容易,并赋能了人们,因为非技术人员或工程团队之外的人通过 Slack 与智能体编码环境互动可能会更自在。Anthropic 在传统工程领域之外是如何使用 Claude Code 和 Claude tag 的?

Hey everyone. Hope you're enjoying the episode so far. Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development. Whether that's talking about the latest tools, the most efficient workflows, or defining best practices. But, for whatever reason, many of you have yet to subscribe to the channel. If you're enjoying the podcast and want us to continue to bring you the very best content, please do us a favor and hit that subscribe button. It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you. All right, back to the episode. Let's talk a little bit about you mentioned very briefly a while back about how the marketing team, for example, created an app, which is wonderful. And I think Claude tag really makes this easier and empowers people because folks who are non-technical or rather outside of the engineering team are probably much more comfortable engaging and interacting with an agentic coding environment when they're doing it through Slack. How has Anthropic used Claude Code, Claude tag outside of the traditional engineering spaces?

Lamis Mukta

是的,老实说,方式太多了。我认为每个团队,整个公司都在 Claude 的轨道上运行。另一个有趣的统计是,在过去一年里,我们公司所有团队愿意委托给 Claude 的工作量翻了一番。所以从我们能够依赖 Claude 的程度来看,从大约 30% 增长到了 60%。所以是的,营销那个很有趣。那是一个用于广告生成或文案生成的管道。我们很多事件响应基础设施也在一定程度上依赖 Claude 来进行分类和通知正确的人。在可能的情况下,它还能开始诊断代码问题等等。所以所有这些不同的解决方案都依赖于你智能体的略微不同的配置。

Yeah, in just so many ways, honestly. I think every team, the whole company is running on the rails of Claude. Ultimately, another fun stat is in the last year, the amount of work which we as a company are comfortable delegating to Claude across all teams has doubled. So it's gone from like 30% to like 60% in terms of where we're able to rely on Claude. So yeah, the marketing one is fun. That one was like a pipeline for ad generation or copy generation. We have a lot of our incident response infrastructure also relies on Claude to some extent to triage things and loop in the right people. And where it can, it can kind of start to diagnose code problems, etc. So all of these different solutions rely on slightly different configurations of your agent.

Claude在销售与产品开发中的主动与定时任务 Claude's proactive and scheduled tasks in sales and product development

Lamis Mukta

比如在销售方面,我们的销售团队有每周简报,Claude 会总结每个人这周的工作,给出所有统计数据和仪表盘更新,这样大家开会时都准备好了,没人需要费力做幻灯片。这是按计划运行的一件事,显然为每个人都节省了时间。还有其他更响应式的事情,比如事件响应,Claude 知道何时介入以及应该多主动。所以我们确实看到了各种形式。例如,当我开发产品时,我有一些界面可以在原型中输入反馈,然后 Claude Code 会在后台处理这些反馈。所以每个人都在这个意义上构建自己的工具。我认为其中很多都融入了我们开发 Claude Tag 的思路,它旨在让所有团队都能轻松使用。例如,主动性这个东西你可以真正地调高或调低。所以你可以有各种设置:Claude 只在被标记时响应,或者 Claude 创建计划来运行任务,或者 Claude 会在认为有相关上下文要分享时主动跳入线程。我们从之前的 Claude Code 产品以及托管智能体中学到的是,在调度和知道多主动方面什么效果好,因为世界上最糟糕的事情就是一个机器人对所有事情都做出烦人的回应。

So for example, on sales, our sales teams have weekly briefs that Claude runs over what everyone did that week and gives all the stats and dashboard updates so everyone is ready for the meeting and nobody had to labor over the slides. That's one thing that runs on a schedule. It's clearly a time save for everybody. Then there are other things that are more responsive, like incident response, where Claude knows when to jump in and how proactive to be. So we really see all flavors of these things. For example, when I'm developing products, I'll have interfaces where I can type feedback into prototypes, and then Claude Code will just work on those in the background. So everyone builds their own tooling in this sense. And I think a lot of this has gone into how we thought about the development of Claude Tag, which is meant to be really accessible across the board for all teams. For example, that proactivity thing is something you can really dial up and down. So you can have everything from Claude only responds when it's tagged, or Claude creates a schedule on which it runs its tasks, or Claude will proactively jump into threads when it thinks it has relevant context to share. And what we've learned from previous products like Claude Code and also on our managed agents is what works well in terms of scheduling and knowing how proactive to be, because the worst thing in the world is a bot that responds to everything with annoying context.

Host

是的。

Yeah.

Lamis Mukta

我们把主动性看作一个真正的频谱,并且随着时间的推移进行了调整,这样它就知道什么合适、在哪里介入、何时不介入,以及何时按程序计划做事。这也是团队可以引导的。如果 Claude 做了你认为不符合你偏好事情,你可以直接告诉它,它会更新记忆,并在未来表现得更像你想要的。

We see this proactivity thing as a real spectrum, and we've tuned over time so that it knows what's appropriate, where to jump in, and when not to, and when to do things on a programmatic schedule. And that's something you can also steer as a team. If Claude does something that you think is not aligned with your preferences, you can just tell it, and it will update its memory and behave more similarly to what you want in the future.

Host

是的。Connor 让我想起之前和 Datadog 的 CEO Livia Polmel 做的一期节目,她谈到人类因为关键问题凌晨 3 点起床——这种情况到底能持续多久?我们能在多大程度上信任智能体去处理事件,做出可逆的修复(希望是可逆的),然后早上我们醒来,注意到发生了某事,再判断是否正确,也许撤销它,换一种方式,但我们可以依赖智能体来做这些。想到你刚才说的,我设想一个场景:可观测性数据进入 Slack 提供信息,然后 Claude Tag 看到这些,当它发现有点不对劲时,会说“我应该上报事件吗?”然后像 Claude Tag 这样的智能体会说“实际上,我要在这里上报一个事件,然后做一些更改,记录我在做什么,做出更改,推送出去。”感觉这真的是我现在想实验的东西,看看这是否是一条有趣的路径。你们在 Anthropic 做这个吗?

Yeah. Connor reminds me of thinking back to a previous episode with Livia Polmel, the CEO of Datadog, talking about how humans getting up at 3:00 a.m. because of a critical issue—how long is that actually going to last? How much do we trust agents to go ahead on an incident, make a fix that could be reversible, hopefully will be reversible, and then maybe in the morning we wake up, notice something happened, and choose if it's the right way, perhaps reverse it, do it a different way, but we can rely upon agents to do this. Thinking about what you were saying, I envision a space where we have observability data going into Slack, providing information, and then I can see something like Claude Tag looking at that, and when it realizes something is a little off, saying, "Should I raise the incident?" And you'll have an agent like Claude Tag saying, "Actually, I am going to raise an incident here, and then I'm going to make some changes, document what I'm doing, make my change, push that." It feels like this is something I'd love to experiment with now and see if that's an interesting path. Is that something you do at Anthropic?

事件响应智能体设计原则与信任建立 Incident response agent design principles and trust journey

Lamis Mukta

是的,事件响应智能体确实是我们看到效果很好的模式。作为一个曾经是值班工程师的人,我知道凌晨 3 点接到寻呼机电话的恐惧,那可不是什么好事。但我认为这非常有趣,它涉及几个重要的设计原则。首先是,要让这个智能体做好工作,你需要给它良好的数据源访问权限。比如你的数据仓库、可能有的任何日志和指标,还有你的代码仓库,这样它才能开始诊断问题。我认为这是这类智能体非常好的入门套件。但另一个非常重要、实际上也影响我们设计托管智能体产品的东西是:你希望在人与智能体交互的哪些环节设置关卡?这完全取决于团队,我们完全理解大规模推广这些东西需要一个信任建立的过程。所以,也许刚开始时,你让 Claude 尝试一下,同时保留传统流程,然后检查它在某个成功阈值下是否做了你希望的事,随着时间的推移,你会更有信心,把越来越多的工作委托给它。或者它可能只是开始诊断修复,然后把结果传给工程团队,如果它认为足够关键就唤醒他们,甚至可能开始起草 PR,或者任何你想要的访问关卡。所以我认为这是工程团队能真正看到价值的地方——这是一个巨大的问题,帮助更快地解决事件,这绝对是我们内部看到的,只需要你仔细考虑在哪些地方希望 Claude 征求你的批准。这是团队自己应该思考和配置的;我们可以建议我们看到的效果好的做法,但这显然是一个高信任度的领域。

So, yeah, the incident response agent is definitely a pattern that we see working really well. And as somebody who used to be an engineer who was on call, I know the fear of the pager duty call coming through at 3:00 in the morning, and it's not a nice one. But I think this is a really interesting one which plays on a couple of important design principles. One is how much—first of all, for this agent to do a good job, you need to give it good access to different data sources. So, your data warehouse, potentially whatever logging and metrics you have, and also potentially your repository so that it can start to diagnose things. This is a very good starter kit, I think, for that kind of agent. But the other thing that's really important and actually something which feeds into how we designed our managed agents product, is where do you want to design those gates between human-agent interaction? So, it's really up to teams, and we totally understand that rolling these things out at scale needs this journey of trust that you need to go on. So, maybe when you start doing this, you just let Claude have a go, and you also keep your traditional process, and you just check that on whatever success threshold it did what you wanted it to, and then over time you get more confident delegating more and more of that work over. Or maybe it just starts diagnosing fixes and passing that over to engineering team, waking them up if it thinks it's critical enough, to potentially starting to raise a draft PR, to whatever other gates of access you want. So I think this is just one of those where you can really see the value for engineering teams of this being a massive problem and helping you resolve incidents much faster, which is definitely what we've seen internally, and just needs you to carefully think about where you want Claude to ask for your approval on these things. And that is something that teams themselves should think about and configure; we can suggest what we've seen work well, but it's obviously a high trust area.

Host

是的,有趣的是,我之前的笔记本电脑上贴着一张贴纸,上面写着“AI 在我睡觉时工作”,现在我觉得——我们会突破信任障碍,变成“AI 在我睡觉时修复生产环境”,我认为这是正确的方向。在诊断、找到根本原因方面,我认为智能体式智能体会更准确、更快地找到数据,而停机造成的损失——有时从停机成本来看,让人工诊断、找到根本原因、提出修复方案等反而更昂贵。所以这里肯定有一个有趣的平衡,我们将看到它会如何发展。看到这种转变将会非常迷人。

Yeah, it's funny actually on my previous laptop I used to have a sticker that said "AI works while I sleep" and now I'm kind of like—we'll break through that trust barrier and it will be "AI fixes production while I sleep" and it's the right path I think. In terms of diagnosing, in terms of getting root causes, agentic agents will find that data more accurately quicker I think, and the loss of the outage—sometimes it's more expensive when you look at it from the cost of an outage to have a human diagnose, find the root cause, propose a fix, etc. So there's definitely an interesting balance there that we're going to see how it's going to be. It's going to be fascinating to see that shift.

托管智能体中的“梦想”功能 Dreaming feature in managed agents

Host

没错。Lamis,几周前你在伦敦 AI Native Dev Con 上做了一个精彩的演讲。实际上,我们还宣布了纽约的 AI Native Dev Con,将于 2026 年 11 月举行,大家也可以关注一下。你提到了一个叫“做梦”的概念,这非常有趣。首先,能跟我们讲讲“做梦”是什么吗?

Absolutely. So, Lamis, you gave a wonderful session at AI Native Dev Con London just a number of weeks ago actually. And in fact, we have AI Native Dev Con in New York, which was announced. That's going to be happening in November 2026. So, take a look at that as well. And now you mentioned a concept called dreaming, which was super curious. First of all, why don't you tell us a little bit about dreaming what it is?

Lamis Mukta

当然可以。“做梦”是我们托管智能体产品中的一个研究预览功能。对于不熟悉托管智能体的人来说,这个产品本质上让你能更快地在生产环境中构建和部署智能体。我们承担了很多工作,比如在 Anthropic 这边管理智能体的框架、部分基础设施和可观测性等。我们真正在做的是,把一段时间以来构建这些智能体所积累的所有经验,转化为非常具体的原语,比如智能体、环境和会话,让你能快速组合这些智能体并迅速部署。这就是托管智能体产品。当然,考虑到上下文在智能体式开发中作为概念的重要性,没有记忆功能是不完整的。所以,这个功能允许智能体在学习过程中读写不同的记忆存储。而且访问权限控制得很好。有非常重要的组织级上下文,只能被读取,比如智能体的草稿区,它们可以在那里存放关于当前工作的上下文,这对它们的工作非常有帮助。正如我所说,人们经常问如何构建这些记忆系统的最佳实践。随着运行时间变长,确实存在信息过时的风险。有些东西已经过时,有些信息缺失,或者写得令人困惑等等。因此,我们设计并引入了“做梦”这个概念,顾名思义,你可以按任意频率运行这些“做梦”任务,输入一些记忆存储和托管智能体的会话记录。这些基本上是智能体执行任务时的轨迹。你把所有这些交给另一个智能体,它会审查这些记录和记忆,寻找任何不一致之处。比如,它可能发现缺少某些信息,如果智能体拥有这些上下文,表现会更好;或者相反,存在误导性信息,降低了性能;或者它只是找到了一种重新组织信息的新方式,使其更容易被智能体搜索和呈现。这个过程相当全面。它会给出修改假设,附上它认为有证据的会话记录。然后你可以决定实施哪些修改。我认为这里非常重要的是,这真正开启了持续学习的道路。比如,你可以今天运行智能体,然后根据可以优化的地方,明天再运行它们,实际上看到它们变得更好。通过“做梦”,你可以在很大程度上将这个过程自动化,然后只需批准你认为相关的修改。这是我们非常兴奋的一点,我们已经看到许多客户在运行这样的流程后,其部署的智能体性能有了显著提升。

Yeah, of course. So, dreaming is a research preview feature that we have on our managed agents offering. And for those who aren't familiar with managed agents, it's called managed agents. This is a product which essentially allows you to build and deploy agents much faster in production. So, we take on a lot of the everything from kind of managing the harness of your agent on Anthropic's side, some of the infrastructure and the observability, etc. And what we're really doing with this product is taking all the learnings that we've got from building these agents over some period of time and building them into really concrete primitives like agents and environments and sessions that allow you to quickly compose those agents and deploy them really fast. So, that's the managed agents product. And of course, given how important context has been as a concept in agentic development, it wouldn't be complete without a memory feature. So, this allows agents to read and write to different memory stores as they learn things. And that's very well access gated. So, there's huge organization level context which is really important and can only be read from to like agent scratch pads that allow them to drop context about the work they're doing, which is just amazing in terms of enabling their work. And like I said, people often ask about what is the best practices for structuring these memory systems. This really starts to run the risk as you run it over longer periods of time that there's stale information in there. Some stuff has gone out of date, there's missing information, it's just confusingly written, etc. And so, what we designed and introduced is this concept called dreaming, which is kind of what it sounds like, I suppose, where essentially you are able to run these dreaming jobs at whatever cadence you like, where you input some of your memory stores, and some of your session transcripts from managed agents. So, these are basically the traces of how your agents have gone and carried out a couple of tasks. And you give this all to another agent, and it basically reviews those transcripts, and it reviews the memories, and it looks for any kinds of discrepancies. So, maybe it finds that some information is missing that the agents would have performed better if they had that context, or vice versa that there's something misleading in there, which is degrading performance, or it just finds a new way to reorganize that information to make it easier to search and surface for the agents. And it does this in a pretty extensive manner. Like, it gives you hypotheses for what to change, gives you attachments to the sessions where it thinks that the evidence is there. And then you can basically decide which of those changes to implement. And I think what's really important here is this really opens the path towards continual learning. Like, this idea that you can run your agents on one day, and then based on whatever could have been optimized, you can run them the next day and actually see that they get better. And with dreaming, you can kind of hand over that to a large degree, allow that process to run in an automated fashion, and then just approve whatever you think is relevant. So, it's something we're really excited about, and we've seen a bunch of customers just see much better performance improvements with their deployed agents when they run processes like this.

Host

太棒了。对于想了解更多内容的听众,你的演讲实际上已经上线了,我们会确保提供链接,让大家看到 Lamis 的完整演讲。

Amazing. And for folks in the audience who want to learn more, your talk is actually online, so we'll make sure we link the audience to that, and you can see Lamis's session in full.

Lamis Mukta

是的,我想再补充一点。那天能去那里真的很开心,而且就像我们说的关于功能的 FOMO(错失恐惧症),有机会谈谈那些可能更复杂、人们使用机会较少的功能,感觉很好。所以,希望大家会喜欢。

And yeah, just add another point there, I guess. I mean, it was such a delight to be there that day, and I think, like we said about FOMO with features, it's nice to have the opportunity to speak about one of those more complicated features potentially that people have less opportunity to use. So, yeah, I hope folks will find that enjoyable.

Host

确实如此。大家很喜欢你的演讲,我们收到了非常好的反馈。非常感谢你。我们时间不多了,但应该做个总结。我们一直喜欢给出实用的建议,让听众能实际去做。那么,对于那些可能正在尝试引入托管工作流、在组织中进一步推广智能体式开发的人,根据你的经验,有哪些日常实践可以极大地帮助他们进入下一阶段?

Absolutely. It was a people loved the session. We got amazing feedback about your session. So, thank you very much for that. So, what are we running out of time, but we should wrap. I'd love we always love, you know, giving practical advice and practical things that our listeners can do. So, what would you say is something that you would say for folks who are maybe, you know, trying to introduce maybe it's managed workflows, introducing agentic development further in their organizations? What would you say are some day-to-day practices that people can do from your experience that will massively unlock the next stage for folks?

Lamis Mukta

当然。考虑到听众可能很广泛,我真的很鼓励大家去设置 Claude 并尝试一下。你只需要让你的 Slack 管理员开启它,做一些配置,但这真的花不了多长时间。给你举个例子,我做的第一件事是设置每日简报。由于 Claude 可以访问的连接器和联系人,它能告诉我过去 24 小时内发生的事情,尤其是与跨国团队合作时。比如我在旧金山的团队做了什么,它会立刻全部呈现给我。这样我就不用醒来面对一堆邮件和 Slack 消息,而是收到一份精心整理的简报,告诉我哪些事情需要关注。

Yeah, of course. And I think just thinking about, you know, potentially broad audience here, I would really encourage people to go and set up Claude and give it a try. You just need to get your Slack admin to turn it on. And do some of the configuration, but it really doesn't take long. And to give you a flavor of the things that I'm doing with this, so the first thing I set up was a daily briefing. And because of the connectors and contacts that Claude has access to, it's able to tell me about things that happened over the past 24 hours, especially working with international teams. Like things that my teams who are in San Francisco did, it just immediately surfaces that all to me. So, I don't have to wake up to a wall of emails and Slack messages, but I just have a nice curated brief about what needs my attention.

每日简报与团队报告 Daily Brief and Team Reports

Lamis Mukta

它还能关联我正在进行的工作流,所以它会告诉我,比如我正在做的这件事,可能是这次演讲,也许你需要在做之前先审阅这些文档等等。所以每日简报非常棒。另外我还设置了一个功能,它知道哪些频道对我重要,会实时提醒我任何需要我关注的紧急事项。所以它看到的任何它认为与我相关的内容都会告诉我,这很棒,因为作为一家公开运营的公司虽然很好,但意味着 Slack 上有很多事情发生,我无法全部跟上。然后在团队层面,我们还有 Tag 做各种事情的每周报告。所以在 Applied AI 团队,我们有一份每周报告,记录团队本周学到和看到的不同事情,这非常好。它鼓励人们不断分享这些背景信息,并真正在整个组织中扩展这些知识。所以从最佳实践的角度来看,这很酷。我认为这些是几个很好的起点,你会很快感受到它的能力。最后一个小技巧是,你可以让 Tag 为你构建一些自定义软件,它可以部署到 Claude Code 的 artifact 中,然后你可以将其作为个人仪表盘分享,或者与团队分享,说‘嘿,这是用来跟踪我们某个工作流的东西’。所以是的,这些是很好的起点。

It has contacts to my ongoing workflows as well, so it can tell me like this thing that I'm working on, maybe this talk, like maybe you want to review these documents before doing that, etc. So, daily brief is great. And then something else I have set up is it knows which channels are important to me, and it pings me about anything urgent that needs my attention on a live basis. So, anything it sees that it thinks is relevant to me, it will tell me about, and that's great because whilst working in public as a company is fantastic, it means there's a lot of things going on on Slack and I cannot keep on top of that. And then some other things like on a team level, what we have is Tag will do weekly reports on various things. So on the Applied AI team, we have a weekly report of different things that the team learned and saw this week, which is really nice. It encourages people to keep sharing that context and really scales that knowledge across the organization. So that's cool in terms of best practices. And yeah, I think these are a couple of good places to get started and I think you'll very quickly get a feel for what the capability of that is. You know, one final flourish you can do is ask Tag to build some custom software for you and it can deploy it into Claude Code artifact and then you can share that as maybe just a personal dashboard or you can share it with your team and be like, 'Hey, this is something to track XYZ workflow that we have.' So yeah, these are some good places to get started.

Host

太棒了。这实际上非常及时,因为就在这周,我非常喜欢像 Getting Things Done 这类工作流。我做这类事情最大的困难之一就是回顾和检查。所以我做的是创建了一个应用并部署了,它基本上做类似的事情。它会查看我的 Slack、我的电子邮件、我的 Todoist 待办列表。它还会查看我的 Granola 笔记,然后添加一堆我从 Granola 中提取的待办事项,它会说,‘哦,这里有一些你应该注意或应该添加到日历的事情。’它会添加进去。它有点像我的个人助理。我告诉你,这真的释放了生产力。所以我完全支持那些回顾、每周回顾、每月回顾以及每日检查和反思。从生产力的角度来看,这真的是一个游戏规则改变者。

Amazing. It's actually super timely because just this week, I'm a big fan of things like Getting Things Done and those types of workflows. And one of my biggest areas of trouble of doing those types of things is the reviews and the check-ins. And so what I did was I created an app which I've deployed and it essentially does very similar. It looks through my Slack, it looks through my email, it looks through my Todoist to-do lists. And it also looks through my Granola notes as well and it will add a whole bunch of to-dos that I am extracting out from Granola and it will essentially say, 'Oh, here are some things that you should be aware of or should add to your calendar.' It adds them in. It's kind of like a little bit like my EA as well. And I'll tell you what, that's really unlocking productivity. So I totally am on board with those reviews, weekly reviews, monthly reviews and daily check-ins and reflections. It's a real game-changer from a productivity point of view.

Lamis Mukta

是的,一定要用 Claude Tag 试试看。

Yeah, no, definitely give it a go with Claude Tag.

Host

我需要切换到 Claude Tag 试试看效果如何。

I need to switch to Claude Tag and try and see how it goes.

Lamis Mukta

我还特别喜欢让这些智能体告诉你这周你做得好的地方。有时候你没有时间反思,这真的很不错。比如告诉我这周我的三个胜利。告诉我一些关于可以优化的事情的反思等等。但有时候你没有时间反思这些,而且你知道,在 AI 的世界里很多事情发生得很快。能够花点时间反思正在发生的事情是很好的。

Some of the ones I really like as well are like just getting these agents to tell you what you did well that week. Like that's a really nice thing sometimes that you don't have time to reflect on. Just like tell me three wins I had this week. Tell me a couple of reflections on things I could optimize, etc. But sometimes you don't have time to reflect on those things and you know, in the world of AI a lot is happening very quickly. It's nice to be able to take moments to reflect on what's going on.

Host

我还做的是,我实际上也把我的——我最近刚做了年度评估。所以我把我的年度反馈放进去,以及其他我觉得我的管理团队等会希望我做的事情。它还会根据我正在做的事情给我反馈。我是否在采取下一步改进?我是否在做团队需要我做的事情?这些实际上非常有价值,因为你在做你想做的事,同时其他人的需求也通过你做的事情得到了满足。所以——我不知道。所以我只需要一个智能体也能做我的工作,然后我就可以放手了。

What I also do is I actually also put my — I've just recently had an annual. So what I do is I put my annual feedback in, as well as other things that I feel like my management team and things like that would want of me. And it also gives me feedback based on what I'm doing. Am I taking my next steps in improvement? Am I doing what the team are needing from me and those types of things which are actually really valuable in terms of you're doing what you want and actually, other people's needs are also being satisfied by some of the stuff that you're doing. So it — I don't know. So all I need is an agent to just do my work as well and then I can just let it go on.

Lamis Mukta

是的。我们可以去享受阳光了。

Yeah. We can go enjoy the sun.

Host

是的。

Yeah.

结束语 Closing Remarks

Host

太棒了。Lamis,时间过得真快。这是一次令人难以置信的讨论。我非常感谢你,不仅是因为 AI Native Dev 的环节非常受欢迎,而且这是一次绝对精彩和迷人的对话。非常感谢你加入我们,也感谢你分享 Anthropic 如何使用 Claude Tag、Claude Code 以及你学到的东西的所有见解。这太棒了。

Amazing. Lamis, this has absolutely flown by. It's been an incredible discussion. I very much thank you not just for the AI Native Dev session which was very well received but absolutely wonderful and fascinating conversation. So really appreciate you joining us and thanks for all the insights as well as to how Anthropic are using Claude Tag, Claude Code, things that you've learned. It's been wonderful.

Lamis Mukta

非常感谢你,Simon。我也非常开心。非常感谢你邀请我。

Thank you so much, Simon. It's been an absolute blast as well. And thank you so much for having me.

Host

太棒了,Lamis。非常感谢你。我相信我们的观众很喜欢这次讨论。请收听下一集。再见。AI Native Dev 由 Tessal 为您呈现,Tessal 是技能和上下文的包管理器。您的主持人是 Guy Pidgeon 和我 Simon Maple。我们的制作人是 Tom Dawler。AI Native Dev 不仅仅是一个播客,它是一个社区,我们每月在伦敦市中心的 Tessal 办公室举办聚会。请访问 tessal.io/community 了解更多信息,希望在那里见到你。

Amazing, Lamis. Thank you so so much. I'm sure our audience enjoyed that discussion. Tune into the next episode. Bye for now. The AI Native Dev is brought to you by Tessal, the package manager for skills and context. Your hosts are Guy Pidgeon and me, Simon Maple. Our producer is Tom Dawler. The AI Native Dev is not just a podcast, it's a community and we host monthly meetups at the Tessal offices in Central London. Visit tessal.io/community to learn more and I hope to see you there.

互动版:逐字朗读 + 针对本期提问 →