AI 时代的产品构建艺术

The Art of Product Building in the Age of AI

迈克·克里格 Mike Krieger · AI & I · 2026-03-25 · 约 48 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Instagram 联合创始人 Mike Krieger 探讨 AI 如何让产品构建更快但未必更简单,强调直觉和知道该砍掉什么的重要性。

Instagram co-founder Mike Krieger discusses how AI has made building products faster but not necessarily easier, emphasizing the importance of intuition and knowing what to cut.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 10)

全文 · Full transcript(中英对照)

AI时代的产品构建 Product building in the age of AI

Host

世界变化很快,在 AI 时代,压力不仅仅是跑得更快,还要确保你发出去的东西听起来真的像你。从邮件到提案再到利益相关者更新,千篇一律和仓促了事都不行。如果你曾盯着空白页面,心里知道想说什么却不知如何下笔,Grammarly 能解决这个问题。Grammarly 给你一个思考、写作和完成工作的地方,就在你原本写作的地方。大多数 AI 工具要么接管一切,要么袖手旁观。Grammarly 两者都不做。它帮你打破空白页,调整语气让信息准确传达给特定读者,并在你已使用的超过 50 万个应用和网站上无缝运行。它内置了为你流程每一步设计的智能体,90% 的专业人士说它节省了时间,93% 说它帮他们完成更多工作。这是与你协作而非凌驾于你之上的 AI。在一个通用 AI 的世界里,不要听起来和别人一样。有了 Grammarly,你永远不会。免费下载 Grammarly,访问 grammarly.com。那就是 grammarly.com。Mike,欢迎来到节目。

The world moves fast and in the age of AI, the pressure isn't just to move faster. It's to make sure that what you send actually sounds like you. From emails to proposals to stakeholder updates, generic and rushed just doesn't cut it. If you've ever stared at a blank page knowing exactly what you want to say, but not how to start, Grammarly fixes that. Grammarly gives you one place to think, write, and finish your work right where you already write. Most AI tools either take over or stay out of the way. Grammarly does neither. It helps you break the blank page, adjust your tone so a message lands right for the specific person reading it, and works seamlessly across more than 500,000 apps and sites that you're already using. It's loaded with agents built for every step of your process and 90% of professionals say it saved them time. 93% say it helps them get more done. This is AI that works with you, not over you. In a world of generic AI, don't sound like everyone else. With Grammarly, you never will. Download Grammarly for free at grammarly.com. That's grammarly.com. Mike, welcome to the show.

Mike Krieger

很高兴来到这里。谢谢邀请。

Great to be here. Thanks for having me on.

Host

很高兴你能来。我超级兴奋。对于不了解的人,你是 Instagram 的联合创始人,现在在 Anthropic 和 Anthropic Labs。我长期以来一直远远钦佩你在 Anthropic 和 Instagram 的工作,你显然处于 AI 产品构建的前沿。所以,谢谢你来。

Great to have you. I'm super excited. For people who don't know, you are the co-founder of Instagram and now you are at Anthropic and Anthropic Labs. I've admired your work from afar both at Anthropic and at Instagram for a really long time and you're obviously at the forefront of building products in AI. So, thank you for coming on.

Mike Krieger

当然。

Absolutely.

Host

我们从哪里开始?就像我们刚才在前期制作中聊到的,在产品构建中,什么变得更容易了,什么变得更难了,或者可能保持不变,因为底层基础或我们构建产品的过程已经完全改变了。所以,跟我讲讲你现在与之前在 Anthropic 和 Instagram 的经历相比,你认为事情是如何变化的。

Where should we start? Like what we were talking about just now in the pre-production is what has gotten easier and what has gotten harder or stayed maybe the same in product building as the underlying substrate or the process by which we build products has changed completely. So, tell me about your experience now versus earlier in Anthropic versus Instagram and how you think things are changing.

Mike Krieger

是的,几周前我做了个思维实验。你知道,在 Instagram 的故事里,我们之前有个产品叫 Bourbon。我们做了差不多一年,没成,然后转型了。我们基本花了三个月构建了后来的 Instagram,发布并规模化。我当时问自己:现在什么变得微不足道了?而构建过程中哪些东西是固有的、不会变得更简单的?那一年我们可能本可以更早碰到那些死胡同,但走到那里也有价值,对吧?我们过度复杂化了产品,然后不得不简化它。我发现即使是今天的模型,也擅长添加功能,但不一定擅长判断该砍掉什么,而这需要大量真实世界的使用。增量添加东西的过程本身也有价值。现在,尤其是我们在实验室里构建的一些东西,你可以让它从零到一,甚至从零到终很快完成,几小时内,但沿途它做了很多决策。你可以让它跟进你、做输入,但关于该放什么进去的直觉,我认为是随着时间积累的。所以我一直在反思,即使在加速 AI 构建的时代,也没有太多突破性的消费产品。我认为部分原因是,你仍然需要时间来打磨你对世界施加何种干预的看法,然后从那里开始构建。而一旦你知道要构建什么,实际的构建部分当然容易得多。我让 Claude 基本上重建了 Bourbon,花了大约两个小时,功能完整。它添加了 Bourbon 没有的滤镜——我们后来为 Instagram 加了那些——但我想它知道产品的最终未来,所以决定直接内置。所以这部分感觉非常不同。但还有一点,我记得有一周 Kevin 去构建 Instagram V1 的所有滤镜,我去构建应用的其余部分。我熬夜到凌晨 4 点,然后睡到中午,那是我自然的昼夜节律。在这个过程中,你要做很多决策:位置功能怎么设计?怎么……你知道,我们必须找到一种加速构建的方法,同时帮助人们沿途建立对这些决策的直觉,否则我认为你只会得到非常通用的产品,不太可能脱颖而出,或者只是不能反映你对领域或产品的更深层直觉。

Yeah, I was doing the thought exercise a couple weeks ago of, you know, we know in the Instagram story we had another product called Bourbon. We worked on that for almost a year. It wasn't working. We pivoted. We basically spent three months building what became Instagram, launched it, and then scaled it. And I was asking the question like what is now trivial and what was actually inherent in that building process that doesn't get easier, right? And that year we probably could have hit some of the dead ends we had eventually hit sooner, but there was value in getting there too, right? Like we over complicated the product so that we then had to simplify it. I find even the models today are good at adding features. They're not necessarily good about figuring out what to cut out of the product and that took a lot of just sort of hitting actual real world usage. And there was something about the process of incrementally adding things right now. I mean, today especially some of the stuff we're building in labs like you can get it to go zero not it's zero to one, but zero to end pretty quickly over the matter of hours, but it's made a lot of decisions along the way and yeah, you can ask it to follow up with you and then do input, but some of the sort of intuitions you build about what are the right things to put in there. I think you build over time. And so, I've been reflecting like there haven't been a lot of breakout consumer products even in the age of accelerated AI building and I think part of it is because it just still takes time to sort of hone your view about what sort of intervention you want to make on the world and then build from there. Now, the actual building part once you know what to build is of course so much easier. I had Claude basically rebuild Bourbon. It took about two hours. It was feature complete. It added filters which Bourbon didn't have. We added those for Instagram, but I think it knew what the eventual future of the product so it decided to build that in. So, I think that that part feels really different. But I think there's also, you know, I remember there was a week where Kevin went off and built all the filters for Instagram V1. I went off and built like sort of the rest of the app. And, you know, sitting there I was I would stay up till 4:00 a.m. and then sleep till noon. That's like my natural day-night cycle. And like in that process you're making so many decisions. Like how should location work? How do And, you know, it's we got to find a way of accelerating building while still sort of helping people build intuition of those decisions along the way cuz otherwise I think you either get just get very generic products that are unlikely to break out or ones that just don't reflect some deeper intuition that you come to about your space or your product.

Host

太棒了。我喜欢这个。这让我想到两件事。一是我脑子里有个小想法:如果你在室内种一棵树,不让它经历风吹,它就不会那么强壮,因为生长过程中需要这些力量来回推它,才能长成真正的树。所以如果你在室内无风环境下种树,树会长出来,但会倾斜,变得不结实,完全不一样。我觉得你在这里说的就是,因为我们极大地加速了开发速度,原本那种逐步推进、一次做一件事然后暴露给用户的过程,现在你实际上可以在室内种一整棵树,然后你得到整个东西,但它没有那种在每个步骤中积累的直觉和经验,而这正是创造伟大产品所需要的。是不是这样?

This is great. I love this. It's making me think of two things. One is I have this little thing in my head that if you grow a tree without it being exposed to wind, it doesn't get as strong because as it's growing it needs all these forces pushing it back and forth in order to make a real tree. And so, if you have it indoors without wind, you're going to grow a tree, but it leans and it gets all and it's not as strong and it's not the same thing. And I think there's something that you're saying here where because we've accelerated the pace of development so drastically, what would normally be this sort of incremental thing where you're doing things one at a time and then you're exposing it to users, you can actually kind of grow an entire tree indoors and then you have this whole thing that you're just like it doesn't have the same level of intuition and exposure to experience at each step that that creates a great product. Is that is that is that

Mike Krieger

我喜欢这个。我也喜欢这个比喻。你知道,我们创办 Instagram 时,非常信奉 Eric Ries 和精益创业,还有 YAGNI(你不会需要它)原则。但我发现,实际上最近我在实验室做的一个项目,我们在进入早期访问之前就过度构建了 V1,因为你能做到。你会想,哦,我们有这个选项,为什么不也加上这个呢?那不过是一个 PR 的工作量。

I love that. I love that metaphor too. We, you know, when we were starting Instagram we had this we were very into like Eric Ries and Lean Startup and that whole like YAGNI like you ain't going to need it principle. And I have found and actually even one of the things I was working on in labs recently, we way over built for V1 before we even got to early access because you can. You're like, oh, well, we have this option. Why not add this one as well? That's like that's a PR of work.

氛围编码与产品复杂性 Vibe coding and product complexity

Mike Krieger

如果你用 Claude Code 进入很好的心流状态,你会不断抛出东西。你去吃午饭,回来,事情就完成了。你会想,太好了,我们加上了。但我们意识到,我们创造了一个功能矩阵,在发布前很难测试和维护,甚至很难向别人解释。就像他们刚来。有人给了我一个我很喜欢的比喻:就像一集一集地认识电视剧里的角色,对比一下,想象你被扔进最后一集。你会想,等等,这些都是什么?这些人都是谁?我已经被期望拥有所有这些背景。我认为随着时间的推移开发东西也有同样的感觉。但树的比喻我觉得也很贴切。所以,向某人展示一棵完全成形的树,一下子也太多了。我认为这确实涉及到当今如何构建产品并保持简洁。不是因为你能做就意味着它应该出现在至少第一个版本里。

And if you get a really good flow and Claude code, you know, you're firing things off. You're going to lunch. You're coming back. The thing is done. You're like, great, we added it. And the thing we realized was we'd created this sort of matrix of functionality that was actually quite hard to test and keep up with right before launch or even to explain to people. Like they're arriving. The metaphor I somebody else gave me which I really like is the difference between sort of getting episode by episode, getting to characters in a TV show versus imagine like you're thrown into the final episode. You're like, wait, what are all these things and who are all these people? And like I already, you know, I'm expected to have all of this context. I think there's the same kind of feeling around like developing something over time. But I the tree metaphor I think sticks too as well. And so, like showing somebody the fully formed tree is also kind of a lot all at once. I think there's there's there's definitely something there in how do you build product these days and still keep it simple. And not because just because you can doesn't necessarily mean that it should be in at least the first version.

Host

我也有同样的问题。你知道,我昨晚一直调试到凌晨 4 点,修复我在 Every 做的这个叫 Proof 的应用,它是一个智能体原生的协作 Markdown 编辑器。你可以和团队或其他智能体快速分享计划文档之类的东西。里面还有小礼物,很有趣。这是我端到端完整产品的第二或第三次迭代,现在能做到这一点很有意思。但前几次迭代,因为 vibe coding 太有趣、太让人上瘾了,我就发现自己总想着,好啊,我做这个,再做那个。结果就造出了一个不好用的怪物。我受到另一个产品 Monologue 的启发,不知道你有没有遇到过。Monologue 是一个很简单的语音转文字应用,由 JM Naveen 运营,他非常专注于把一件简单的事情做到极致。我看到在这个任何人都能做出产品的时代,一个打磨得极其精致、只专注于一件事的产品效果有多好。所以,我基本上把产品扔掉了,重新开始,只做一个可分享的 Markdown 链接。然后它就在 Every 内部病毒式传播,每个人都开始一直用它。现在我们发布了它,它就爆了。所以我昨晚整夜没睡,试图修复它。我心想,我太老了,不能再这样了,因为这让我想起 20 多岁或大学时熬夜搞东西的样子,虽然有趣,但也累人。所以,我发现我必须真正调整自己的心态,因为可能性太多了。你是怎么应对的?

I'm having the same problem because you know, I was literally up until 4:00 a.m. debugging and fixing this app that I made like on the side at Every called Proof which is an agent-native collaborative markdown editor. So, you can like share really quick plan docs and stuff with your team or with other agents. And you have little presents and it's really fun. And this is like my second or third iteration of the full product end-to-end which is really interesting you can do now. But the first couple iterations, I just found myself because vibe coding is so fun and so addictive, I just found myself being like, yeah, like I'll do this and I'll do this. And like it just created this monstrosity that wasn't that good to use. And I got really inspired by we have another product called Monologue which I'm not sure if you've run into or not. But I got really inspired by Monologue which is a really simple speech-to-text app run by JM Naveen who he's just so focused on making one simple thing work so well. And I saw how well that works in this age of just like anyone can make a product is like something that's super polished and just super good at what it does. And so, I just basically threw out the product and started over with this very simple like it's just a shareable markdown link. And that then just started growing virally inside of Every like everyone started using it all the time. And then now I we launched it and it just blew up. And so, I spent all last night like not sleeping trying to fix it. And being like, I'm too old for this I can't be doing this anymore because it just reminded me of like being in my 20s or like being in college and like hacking on stuff and whatever which is fun, but also exhausting. And so, yeah, I've found that I've had to really modify my psychology because so much is possible. How are you dealing with that?

Mike Krieger

是的,简单提一下,对于 Bourbon,我们最大的错误是随着时间的推移增加功能,而不是删除功能,对吧?因为,你知道,八个功能做不出好产品,也许第九个可以。结果它只是让东西变得非常复杂。我认为我们应对的一部分是更愿意重写。你知道,经典的 Fred Brooks 的《人月神话》说你不应该重写软件,因为 V1 中蕴含的所有东西你都会搞砸。而且还会导致……

Yeah, and just as a brief aside on that, I mean, with Bourbon, our biggest mistake was adding functionality over time rather than deleting it, right? And because oh, you know, eight features doesn't make for good product. Maybe the ninth one will. Instead it just made for, you know, something that felt really complicated. I mean, I think a couple of things are also like part of how we're dealing with it is actually being more willing to do rewrites. You know, like classic, you know, Fred Brooks' Mythical Man-Month. Like you shouldn't rewrite software because all the things that were imbued in V1, you're going to mess up. And it also leads...

Host

是的。

Yeah.

Mike Krieger

没错。还有整个第二系统综合征。这仍然有很多道理,但第一,模型可以帮助你进行差异比较,基本上看出你是否遗漏了第一个版本中的任何东西。第二,这不再是那种可能毁掉公司的长达一年的重写,比如著名的 Netscape,这些重写可能只需要几天,尤其是基于某个源代码。所以,我们实际上有几个项目,通常在发布前,很少在发布后,但至少在发布前,我们构建了东西,意识到我们过于复杂化或做出了某种核心假设,然后推倒重来,做了 V2,然后在此基础上迭代。所以,这成为你不得不做的事情的一部分,我并不惊讶,但感觉没那么痛苦。你不会想,“哦,我花了一年建这个东西。”而是想,“哦,那是上周的事,然后我这周又做了一次,而且我还砍掉了不少东西。”

Exactly. And the whole second system syndrome. And there is still a lot of truth to that, but one, you know, the models can help you sort of diff and basically see did you miss anything that was in that first one. But second, it's just it's no longer you're not like talking about a year-long rewrite that might have killed a company like, you know, famous like Netscape where like these are like days probably especially off of a given source. So, we've actually had several initiatives like usually pre-launch, rarely post-launch, but at least pre-launch like have built the thing, realized we've overcomplicated or made some kind of core assumption, and then like tore it down, done a V2, and then iterate on it from there. So, it doesn't surprise me that that's become sort of part of what you've had to do as well, but it doesn't feel as painful. You're not like, "Oh, like my year of building this thing." It's like, "Oh, that was last week, and then I got to do it this week, and I get to cut out a lot of what was there as well."

Mike Krieger

从功能角度和产品开发的角度来看,我认为我们正在学习更早地发布,这当然是一种平衡。你知道,我们已经成长起来,拥有强大的企业足迹,人们对初始版本有期望,但我们不能假设我们在发布前就知道需要添加的所有连接器和所有东西,因为人们仍然会给我们惊喜,对吧?我们有一支强大的队伍,我们称之为“蚂蚁脚”,因为我们在 Anthropic 是蚂蚁,但这只能带你走这么远,之后你需要真实世界的接触。以 Cowerk 为例,我们琢磨那种形状的产品已经很长时间了。然后一旦我们决定,“不,我们把它推出去。让我们构建我们认为能以最简约方式解决问题的 V1,并在 10 天内发布。”这确实是一个很好的推动。是的,V1 应该或可能有 100 样东西,但它没有。同时,它足够有用,可以在那里证明一些东西,我不确定再开发两个月,增加 50 个功能会更有用。事实上,我们可能一直在建造室内树,然后一旦它进入现实世界,人们会说,“实际上,没人想那样做。他们想做这个其他部分。”所以,我认为这部分再次说明,原始精益创业的直觉仍然存在。只是它们以不同的时间尺度和不同的方式显现。

I think functionality-wise and how we're dealing with it from a product development standpoint, I think we are learning to launch earlier, and it's definitely a balance around, you know, we've grown, we have like a strong enterprise footprint, people have expectations about like what the initial version is, but not assuming that we're going to know what every connector, everything that we need to add to the product is but ahead of launch because people still will absolutely surprise us, right? We have a strong contingent in uh contingent of we call them ant footers because we're ants at Anthropic, but that only gets you so far before you need that real-world contact. Like take Cowerk for example, we'd been noodling on a product of that shape for a long time. And then once we decided like, "No, let's get this out. Let's actually, you know, build the V1 that we think solves the problem in the most minimal way possible and get that out in 10 days." Was really a good push around, yes, there are a hundred things that V1 should or could have had, but it didn't. And at the same time it was useful enough to prove something out there, and I'm not sure developing it for another 2 months adding, you know, 50 features would have been more useful. In fact, we probably would have been building in a the indoor tree would have been getting built and then the second it hit real-world it's like, "Actually, nobody wants to do that. They want to do this other piece." So, I think that piece to that again, there's like the intuitions of the original lean startup ideas are still here. It's just they manifest at different time scale and in a different way.

Host

我真的很想听听你对产品设计以及产品应该如何运作的看法,因为我一直……任何告诉你我使用最多的短语或词语的人都知道,关于我们构建的软件,它必须是智能体原生的。

I'm really curious to hear how you think about product design and how products should work because the I've been... Anyone that ever will tell you the phrase that I use the most or the word I use the most about the software we build is it has to be agent native.

Agent原生原则与Claude Code Agent-Native Principles and Claude Code

Mike Krieger

所以,智能体必须能够使用用户在应用里能做的任何事。还有另外几个“智能体原生”的小原则,但我基本上是从你们那里“偷”来的。我认为 Claude Code 是教我这类产品如何运作得这么好的典范。它是一个智能体,可以在你的电脑上做任何你能做的事,而且可定制、灵活、可扩展。上手容易,但能做出各种设计师事先没想过的意外之事。我觉得这是 AI 产品开发的一个绝佳模型。我很好奇——这只是我从观察你们的工作中借鉴并加入自己想法的东西——但你们是怎么思考的?你们是怎么谈论制作这类产品的?

So, agents have to be able to use anything that a user can do in the app. There are a couple of other little principles of being agent-native, but I basically stole that from you guys. I think Claude Code is the canonical thing that taught me about how that kind of product can work so well. It's an agent that can do anything on your computer that you can do, and it's customizable, flexible, and extensible. It's easy to start, but it can do all sorts of unexpected things that the designers didn't really think about beforehand. I think that's such a good model for AI product development. I'm kind of curious—this is just what I've cribbed from watching what you guys do and then put my own spin on—but how do you think about it and how do you talk about making products like that?

Host

是的,这里面内容很多,我很喜欢你们做的“智能体原生”那篇文章。对我来说,它是这个领域的典范探索。所以,感谢你们把那些想法表达得如此清晰。我想有几点可以展开。第一点是最近和一个人的对话,他是个非技术人员。他说:“你们都在谈论智能体这些东西。”他接着说:“实际上,电脑现在终于能用了。我一直希望电脑能好用,但以前不好用,现在好用了。”这很有趣:如果你知道那些咒语——正确进入命令行、用 brew 安装东西——没人会去做,但现在 Claude 可以帮你做。因此,电脑现在感觉像是一个与你并肩的工具。我认为这个核心洞察不仅仅是给新软件增加能力和功能,更是解锁那些本来就应该存在或可用、但人们觉得极其困难的功能。所以,这可能是第一点。第二点实际上是比较我们那些做得好和做得不好的产品。我认为 Claude Code 做得很好,但 Claude AI 还需要大幅进化。举个例子,我看到有人用 Claude,他们在项目里构建了一个工件或新文档。他们说:“太好了,你能把这个加到我的项目知识里吗?”Claude 回答说:“好的,让我告诉你把它加到项目知识的步骤。”我说:“不,这应该是它能原生做到的事情。”所以,即使在那里面,你也能看到一个 2024 年的产品,经过大量迭代和进化,但我认为它从一开始就没有融入这样一个理念:它应该了解并能够修改自己的每一个基本元素。我认为这在当今的产品中是必不可少的。而 Claude Code 是 2025 年的版本,当你看到人们正在实验的一些“框架”时,还有更进一步的方面,他们实际上可以修改框架本身。这开始进入下一个层次,可能对大多数人来说很深奥,但即使解锁这种功能也意味着你不需要坐在那里想:“哦,我希望它稍微不同地工作。我希望 Gmail 以稍微不同的方式工作。”而是直接让它去做。我认为这感觉像是下一个大步骤,但即使在 Claude Code 内部,教 Claude Code 了解它自己也是一次非常有价值的经历。我当时想,这绝对相关。现在这变得非常循环和元了,但请耐心听我说。

Yeah, there's so much in here, and I love the agent-native write-up you all did. It's the canonical exploration of this to me. So, thanks for putting those ideas out in a really clear way. I think a few threads to pull on this. One is a conversation I had with somebody recently where they said, you know, they're a non-technical person. They said, "You all are talking about agents and all this stuff." They're just like, "Actually, computers just work now. I always wanted computers to work and they didn't work and now they work." It's a funny thing: if you knew the incantations to properly get on the command line and brew install the thing, nobody is going to do that, but now Claude can do it for you. Therefore, the computer now feels like a tool that is alongside you. I think that core insight is more than even just adding power and functionality to new software. It's also unlocking the functionality that always should have been there or available and just felt extremely hard for people. So, that's maybe thought number one. Thought two is actually comparing our products that do this well versus not. I think Claude Code does it well. I think Claude AI still needs to evolve a lot. As an example, I was watching somebody use Claude and they were in a project and they had built an artifact or a new document. They said, "Great, can you add this to my project knowledge?" And Claude's like, "Yeah, let me tell you the steps to go add it to my project knowledge." Like, "No, that should just be a thing that it can do really natively." So, I think even in that you see a product that was a 2024 product that has been iterated on and evolved a lot, but still I don't think has been baked in from the very beginning the idea that every single one of its primitives it should have knowledge about and the ability to modify. And I think that's essential in products these days. And I think Claude Code is the 2025 vintage of that, and I think there's even further aspects of it when you see what some of the harnesses that folks are experimenting with where they can actually modify the harness itself. That starts getting to the next level of that where, you know, it's probably esoteric for most people, but even unlocking that functionality means that you don't have to sit there and be like, "Oh, I wish it did this a little bit differently. I wish Gmail worked in a slightly different way." Instead of just asking it to. And I think that feels like the big next step, but even within Claude Code, just teaching Claude Code about Claude Code was a really valuable experience. I was like this definitely relates. This is now getting very circular meta, but bear with me.

Mike Krieger

我很喜欢你们关于“智能体原生”的文章。我当时想:“我要把这个做成一个技能。”这样,每当我做原型时,它就会以智能体原生的方式思考。所以,我把它打包成一个技能,整个过程是:“嘿,Claude 和 Claude Code,你能为这个创建一个技能吗?”它说:“当然,我正在查找我的技能技能。我要创建一个关于它的技能。我要安装它。”我说:“太好了,现在就能用还是需要重新加载?”它说:“好的,我认为你需要重启它。让我检查一下。是的,你需要。好了,开始吧。”一切都很顺利——它了解自己,这也解锁了巨大的能力,这可能是最后一点要展开的。我认为所有这些都可以聊上几个小时,这也是我们在实验室里真正思考的问题之一:如何让 Claude 构建的软件更了解 Claude,甚至更具备 Claude 智能体原生的构建意识,让它一开始就想到以那种方式构建?因为它仍然不会,部分原因是几十年的软件不是那样的,对吧?那么,如何让新软件内嵌这个原则呢?

I loved your write-up on agent-native. I was like, "I want this as a skill." So, whenever I'm prototyping something, it thinks in an agent-native way. So, I had it packaged up as a skill, and that whole process was, "Hey, Claude and Claude Code, can you create a skill for this?" It's like, "Sure, I'm looking up my skill skill. I'm going to create a skill about it. I'm going to install it." I'm like, "Great, is that available now or do I need to reload?" It said, "All right, I think you need to restart it. Let me check. Yep, you do. All right, let's go." And everything was—it has knowledge about itself, and that unlocks so much capability in there as well, which maybe is like the last thread to pull on. I think all of these could be hour-long conversations, which is one of the things that we're really thinking about in labs: how do you imbue the software that Claude builds to be more Claude-aware and even just Claude agent-native building-aware so that it even thinks to build in that way to start with? Because it still won't, partially because decades of software is not that, right? So, how do you get new software to have that principle baked in?

Host

这正是我正要问你的。所以,第一,我非常荣幸你读了那篇文章并且正在使用它——你为它做了一个技能。这太棒了。第二,是的,你指出了一个我发现的真实问题。我认为实际上 Claude 模型在这方面是最好的。Codex 模型通常不太擅长构建智能体原生,因为模型一般来说,除非你推动它们,否则它们会像传统工程师那样思考。那是完全不同的一套——你想要有护栏和测试。你想要确保用户有一条路径可走,而不是创建这个超级灵活的、可扩展的东西。所以,是的,你如何架构你的产品来教模型和框架,让模型以这种方式思考和工作?

That's the thing I was about to ask you about. So, A, I'm super honored that you read the write-up and you're using it—you made a skill for it. That's amazing. And B, yeah, you're pointing to a real problem that I have found. I think actually Claude models are the best for this. A Codex model generally is not as good at building agent-native because models in general, unless you push them, they think like traditional engineers. And that's a whole different set—you want to have guardrails and tests. You want to make sure that there's one path the user can go down versus creating this extensible thing that's super flexible. So, yeah, how do you architect your product to teach the models and the harness to teach the models to think and work in this way?

Mike Krieger

是的,我认为有两个部分。第一部分是比较平凡的部分,第二部分我认为是开发中更有趣的部分。第一部分是,即使在模型构建时,让它有好的模式和范例可用,这已经非常有价值了。找到模板化和技能化之间的正确平衡,对吧?以及什么是正确的平衡。但拥有——你知道,我们现在有的一件事是关于 Claude API 的技能,这听起来非常明显,但即使只是拥有它也非常有价值,因为有时你会发现,我们发布了一个新模型,它不在模型的内在知识里,然后你就会陷入这些非常有趣的争论。比如:“不,我知道你打错了。是 Sonnet 45。”你说:“不,我知道是 Sonnet 45。”然后你说:“不,不,不。”所以,拥有这种能力,在技能中有好的模板化示例,我认为是有帮助的。

Yeah, I think there are two parts to it. One is the more sort of mundane part, and the second one I think is the one that's more sort of interesting in developing. The first one is even just having good patterns and paradigms available to the model while it builds has been really valuable. Finding the right balance of templatized to skillified, right? And what that right balance is. But having, you know, one of the things that we have now is a skill about the Claude API, which sounds super obvious, but even just having that is really valuable because you would sometimes find, you know, we'd launch a new model. It wasn't in the model's sort of innate knowledge, and then you'd get into these really funny arguments. Like, "No, I know you made a typo. It's Sonnet 45." You're like, "No, I know it's Sonnet 45." And you're like, "No, no, no." So, having that capability, having good templatized examples of that in skills, I think helps.

测试Agent原生产品 Testing agent-native products

Mike Krieger

但第二部分同样有趣的是,这类软件本身就是一种不同类型的测试。为智能体原生产品编写端到端功能测试要难得多,部分原因就在于它的不可预测性。所以,我们在实验室里经常讨论的一个想法是,如何提高验证的保真度?前几天我在开发一个智能体原生的 iOS 应用,让 Claude 与它交互,结果 Claude 在 iOS 的聊天功能里和自己聊了起来。看着 Claude 和 Claude 对话非常有趣,就像有人在假装人类一样。这个原型是关于工作日志反思的,Claude 说:‘是啊,我老板对我太苛刻了,我今天过得很糟。’然后另一个 Claude 回应:‘哦,听到这个我很难过。’它们就这样来回对话。但你不会为这种情况编写单元测试,而且它可能还会产生其他涌现的想法。所以,我认为你必须更多地设置测试框架,尽可能多地锻炼智能体原生能力,因为你并不确切知道事情会如何发展。事情最终会走向奇怪的方向,Claude 会尝试做一些你根本没想到它会做的事,可能会把你的应用带入一个新状态。所以,也许这又回到了难点:底层架构必须足够健壮来应对这种情况,这非常重要,对吧?它是智能体原生的,但也需要能够以你未曾预料的方式灵活应变,但你要有正确的原语。我觉得这就是 2026 年软件设计的艺术与科学。

But then the second part is what's also interesting is that class of software is just a different type of test. Like it's much harder to sort of write an end-to-end functional test around an agent native product because part of it is that unpredictability. And so, another idea we've been kicking around a lot in labs is like how do you increase like the sort of fidelity of the verification? The other day I had an agent native iOS app that I was working on, and I was having Claude interact with it, and Claude ended up having a conversation with itself in a chat feature in the iOS. It was very funny watching Claude talk to Claude because it's like somebody's pretending to be what humans are. And this particular one was a prototype I was doing about a sort of work journal reflections, and the Claude was like, 'Yeah, my boss is really rough on me. Like I had a hard day.' And then the Claude's like, 'Oh, I'm so sorry to hear that.' And they're just going back and forth. But you wouldn't have written a unit test for this, and maybe it would have come up with some other emergent idea as well. So, yeah, I think you just have to go much more towards setting up harnesses that are actually exercising as much of that agent native capability as possible because you don't exactly know what things are going to do. And things are going to end up in a weird place where Claude's going to try to do something that you wouldn't even think it was going to do, and it might put your app in a new state. So, maybe it's circling all the way back to still what's hard. It's like having the underlying architecture to still be robust to that is really important, right? It's like it's agent native, but it's also able to flex in a way that you might not have anticipated, but you've got the right primitives, right? I feel like that is the art and science of software design in 2026.

Host

这真的很有趣。我完全同意你的看法。是的,你希望在一个安全的环境里有一个游乐场。只有边缘安全,你才能拥有游乐场,但我认为最初我们把游乐场做得太小、太受限了。现在模型变了,我们可以把它打开很多,但我们还没有完全弄清楚——至少我还没有完全弄清楚界限在哪里。

That's really interesting. I totally agree with you. Yeah, you wanted to have a playground within a safe environment. That's the only way you can have a playground is if it's safe around the edges, but I think initially we made the playground like way too small and constrained. And now the models have changed, and so we can open it up a lot, but we still haven't figured out exactly like at least I have not figured out exactly what the lines are.

Mike Krieger

是的,我觉得这里面有很多东西。这让我想到一个我脑海中的想法,我在想你是否能用更简洁的方式表达出来。那就是,现在产品中的价值单位就像是工作量证明或使用证明——当团队中有人向我提交 PR 时,我不一定想看所有测试都通过了,因为我假设它们通过了,而是希望你能发一个录屏,展示你或你的智能体在使用它,这样我就能判断它好不好,你明白吗?

Yeah, I think that there's so much here. Like one thing that this is making me think of is that I have this idea in the back of my head, and I'm wondering if you have a way to put this that is more succinct. It's like the unit of value in products right now is like proof of work or proof of use where when someone on the team submits a PR to me, I want to see not necessarily that all the tests passed because I just assume that it did, but like send me a loom of you using it or your agent using it so I can tell is this good or not, you know?

Host

是的。嗯,你是怎么考虑这个问题的?

Yes. Yeah, how are you thinking about that?

Mike Krieger

是的,我认为大概有三个层面。第一个是让 Claude 以某种方式证明它已经测试过了,你知道吗?我已经开始在我所有的提示词里这么做了。当它在开发一个功能时,我会说,最后,在你提交 PR 之前,先向你自己、然后向我证明它能按预期工作。找到正确的方法来做。这实际上迫使你改变自己构建和搭建脚手架的方式,思考什么是让 Claude 至少能简洁地测试这个变更的正确方法,而不是像它喜欢做的那样——‘我读了代码,看起来不错。’我会说,‘代码是你写的,我不信任你。’所以你必须真正测试这个东西。第二个层面是你描述的那种,每件事都需要某种证明,证明它是否按预期工作,以及是否按你的意图工作,因为 Claude 或任何这些模型会为你做很多决定,有时你不知道——我团队里的工程师提交 PR 时,我会问‘你为什么选择这样做而不是那样?’很多时候答案是‘他们没选,只是模型做的选择’,也许那是个合理的选择,大概还算合理,但并不是范式下的最优选择。我觉得这不仅仅是工作量证明,而是思考深度的证明——你有没有想清楚?昨天我和一个工程师聊天,他说:‘我知道你会问很多问题,所以我提前审查了 Claude 做的东西,这样我就不会说“我不确定”了。’大多数 PR 我不会这么追问,但如果是重构系统、引入新原语这样的 PR,我会说‘太好了,让我们确保这些原语是好的,并且你想清楚了它们之间的相互关系’,因为否则很容易陷入一堆你没有完全意识到的假设之中。

Yeah, I think that there's probably like three layers to that. There's the first one was like Claude prove to me that you've exercised this in some way, you know? I've started doing that in all my prompts. I end, you know, when it's working on a feature I'm like and by the end, you know, before you PR like prove to yourself and then to me that it works as intended. Like find the right way of doing it. Which actually ends up you have to change your own sort of way you build and scaffold around saying what is the right way to get Claude able to at least test this change, you know, succinctly rather than what it likes to do is like I read the code it looks good. I'm like you wrote the code. I don't trust you. So you know, you got to really test this thing. And then the second one is that what you described is like, you know, everything having some sort of proof around like did it work as intended and as you intended to because Claude is going to make or any of these models is going to make a lot of decisions for you and sometimes you don't you know, I'll have engineers on the team put up a PR and I'm like oh, why did you choose to do this versus that? And many times the answer is they didn't choose. It was just the choice the model made and maybe it was a reasonable choice. It was probably a reasonable-ish choice, but it wasn't like the optimal choice as it fit into the paradigm. I feel like that is the it's not just proof of work, but it's like proof of thoughtfulness. Like did you think this through? And I was talking to an engineer yesterday and they were like oh, I was really I knew you were going to ask me a lot of questions about this. So I was reviewing what Claude had done so that I wouldn't be like uh I'm not sure. You know, that's I don't push on that for most PRs, but when there was one that's like oh, I'm refactoring this system and there's going to be these new primitives like great. Let's make sure those are good and that you've thought through how they interrelate because it's very easy to end up otherwise with sort of this tower of assumptions that you're not fully aware of.

Host

我今天就经历了完全一样的事情,因为我做了一个原型。完全是 vibe coding,它现在增长得很快,但也经常宕机。所以我花了最近 12 个小时试图修复它,我们在 Every 内部有一个小型的 SWAT 团队自愿来帮我修复。我不得不给他们做入职培训。我在想‘我怎么解释这个代码库的工作原理?’所以我不得不和模型反复沟通,让它帮我定义这些术语,帮我弄清楚如何解释,这样我才不会看起来像个完全的白痴,因为是的,我理解其中一部分,但不是全部。肯定不足以达到我以前必须了解的那种程度。而且这是一个完全不同的问题:我还需要知道那些吗?现在的界限在哪里?很难说。这也许引出了另一个我还没尝试表达清楚的点,请耐心听我说——有些产品你用起来感觉底层很健壮,而有些产品让你觉得只要一个错误的命令或点击,整个东西就会卡死或变慢。比如我们在 Instagram 的时候,我们有 Instagram 的私信视图一,谁知道呢?你发一条消息,它可能到达也可能到不了对方。我们编写了自己的定制实时系统,它崩溃了很多次,你不会信任它来发送你真正需要对方看到的消息。它更像是一个社交性的东西。

I had literally the same experience today because I made proof. Totally vibe coded and it's growing really fast right now, but it's going down a lot. And so I've been spending the last 12 hours like trying to fix it and so we have a little SWAT team internally at Every that signed up to help me fix it. And so I had to like onboard them. And I was like How do I explain how this code base works? And so I had to like go back and forth with the model a bunch to be like okay, help me to like define these terms. Help me to like figure out how to how I can explain this so I don't look like a total idiot because like yeah, there's I understand some of it, but not all of it. Definitely not enough to like the way that I would used to have to to know to know. And it's a whole different thing to be like do I need to know that anymore? Is it like where's the line now? It's hard hard to tell. Which maybe get to something else and I haven't tried articulate this so bear with me as I like, you know, kind of get there which is there's products that you use that feel robust underneath and those ones that you use that you're like it feels like it's one wrong command or click away from the whole thing either like freezing or being slow. For us at Instagram like we had Instagram view direct messaging view one and that like who knows? You send a message it might or may not arrive to the other person. Like we wrote our own bespoke real-time system. It was like it fell over a bunch of times and you would not trust that to send a message that you really needed somebody else to see. It was just a more of a social thing.

构建稳健的Agent原生产品 Building Robust Agent-Native Products

Mike Krieger

在构建 V2 时,我们非常强调一点:不,如果你发送一条消息,我们可能不会达到 WhatsApp 那种水平——你在荒郊野外只有一格 Edge 信号,它可能还会尝试发送。也许不是那个标准,但至少是一个标准:当我加载消息时,感觉是稳健的。当消息显示已发送,它就是真的发送了。我觉得就像有个小勾选。这只是一个小例子,但我认为这仍然是我们需要弄清楚的事情——如何让这种感觉成为任何产品发布的核心部分,不仅仅是在 Anthropic,而是普遍如此。你构建了这个东西——它感觉是建在沙子上,还是感觉稳健?而智能体原生部分又增加了一层:我能稍微推它一下,它会崩溃吗?还是感觉像有了一个坚实的躯干?是的,你可以从不同方向推我,但你的数据在下面是安全的,它不会因为一次部署就完全崩溃。所以如果这是标准——我同意这是你肯定要达到的目标——那么随着模型变得更好,你在招聘和团队结构上做了哪些改变?比如对我们来说,我们的一个产品 Spiral,我们刚雇了一位新 GM,他技术一般,但在产品和写作感上非常突出,而 Spiral 是一个写作产品。现在我们能雇这样的人,而一年前还不行,因为当时的编码模型不够好。我好奇的是,缺点是如果没有人对所有细节都超级精通,产品可能不会那么稳健。那么你如何看待现在 Labs 团队里谁在构建产品,这如何随时间变化,以及未来会如何变化?

And when we built V2, it was really important that we hammered home: no, if you send a message, we're not probably going to get to WhatsApp level where you can be in the middle of nowhere with one bar of Edge and it will probably still try to go through. Maybe that's not the bar, but still a bar where when I load messages, it feels robust. When it's sent, it's really sent. I feel like there's a little check. That's one small example, but I think that's a thing we still need to figure out how to make feel like an essential part of shipping on anything, not just at Anthropic, but in general. You've built this thing—does it feel like it's built on sand or does it feel robust? And the agent-native part adds something even beyond that: can I push it a little bit and will it fall over, or does it feel like I've got a solid trunk? Yeah, you can push me in different ways, but your data is safe underneath here, and it's not just one deploy away from completely falling over. So if that's the bar—which I agree is where you definitely want to get to—how have you changed who you hire and how your teams are structured as the models have gotten better? For us, for example, one of our products, Spiral, we just hired a new GM who is lightly technical but spikes super high on product and writing sense, and Spiral is a writing product. Now we can hire someone like that where a year ago we wouldn't have been able to because the coding models weren't good enough. I'm curious, but the downside is maybe the product won't feel quite as robust if there's not someone who's super technical in all the details. So how do you think about who builds products right now inside the Labs team, how that has changed over time, and how it will change?

Host

是的,我很喜欢这个问题。实际上你会被拉向两个方向,但两者都很重要。一方面是原语和架构稳健性,我认为这仍然需要资深技术人员。我和某人开玩笑说,我以为我在分布式系统方面的技能不再有用了,但实际上这些可能是推理和思考这些问题时最有用的技能之一。比如上周我和 Claude 就我构建的系统是否需要 Redis 还是只用 Postgres 进行了一场长辩论,这是一场健康的辩论,因为我以前用过很多这些技术,所以有基础。但稳健性的另一面是:你是否只是用系统提示的修复和额外指令掩盖了所有问题,还是你正确地架构了实际工具集?后者同样重要,可能正是这位 GM 能发挥巨大价值的地方。好吧,我在做改动,但就像你不会通过说“5 秒后重试,肯定能行”来修补分布式系统的间歇性故障一样,你也不能用“永远不要用全大写、用 markdown”之类的修补方式——它们实际上都是同一个问题的症状:底层组件是否稳健?而 Claude——实际上我对所有模型都这么说,但我认为 Claude 在两方面都可以做得更好。它仍然需要大量人工监督。在系统方面,它现在能调试生产系统,这非常有价值,但首先架构系统时,我觉得我们仍然受益于那些真正深思熟虑过这些事或有经验的人。在提示方面,如果你给它——我见过人们甚至在公司内部陷入这种开发循环:这是提示,这是系统犯的错误,迭代提示。它的自然倾向就是往提示里加更多东西。然后最终变成这样:如果你入职一个新员工,第一天就给他一百条指令——总是用 markdown 回答,除非……他会说,我只记住你最后告诉我的事。或者我会短路它。所以重新思考:好吧,这实际上是两个不同的工具吗?实际上是两个智能体,每个都有更少的上下文,然后你可以拆分它们。所以回到你最初的问题,我们正在招聘有系统专长的人,即使在 Labs 内部——你认为 Labs 更像是从零到一的原型——这仍然非常有价值,因为稳健性很重要。还有谁能在系统权限、配置和早期测试方面提供帮助?这些东西即使对 Claude 来说也很难,因为它不能自己编辑权限,出于很好的原因它不能。然后在稳健性方面,实际上我们在将产品团队与应用 AI 团队配对方面取得了很大成功。我们的应用 AI 团队每天都在现场帮助客户迭代他们的提示,我们发现我们现在是这些工作的客户零,因为我们有很多 AI 驱动的产品。那么如何把这种专长带进来呢?因为这种专长目前并不在我们的软件工程师身上。

Yeah, I love that. I think it's actually you get pulled in two directions, but they're both important. There's the sort of primitives and architectural robustness which I think still need a sort of senior technical person. I was laughing with somebody—they're like, I thought my skills in distributed systems were not going to be useful anymore, but actually those are maybe some of the most useful skills in reasoning about that and thinking things through. Like I had a long debate with Claude last week around whether the system that I was building needed Redis or not or could go away with just Postgres, and it was a healthy debate where I only because I was grounded in having used a lot of these technologies before. But then there's the other side of robustness which is: have you just papered over all the problems with fixes to your system prompt and additional instructions, or have you architected the actual set of tools correctly? And so the latter is as important and probably where this GM can be really valuable. That okay, I'm making changes, but just like you wouldn't patch a flakiness in your distributed system by just saying, well, just retry it in 5 seconds, I'm sure it'll work. Also not doing the same thing with never ever, all caps use, markdown, or whatever the thing you're trying to patch—they're both actually symptoms of the same thing: is the underlying piece robust or not? And Claude—actually I'd say this about all the models, but I think Claude could be much better at both. It's still a place that still needs a lot of human oversight. On the system's part, it's now able to debug production systems which is really valuable, but architecting them in the first place I feel like we still benefit from somebody who's really thought these three things through or has experience. And on the prompting side, if you give it—I've seen people get into this dev loop even internally here. Here's the prompt. Here's the mistake that the system made. Iterate on the prompt. Its natural tendency is to just add more things to the prompt. And then eventually just get to this thing that, if you onboarded a new employee and you gave them a hundred instructions on their first day—always answer in markdown except when... they'll be like, I'm just going to remember the last thing you told me. Or I'm going to short-circuit it. So then rethinking: okay, are these actually two different tools? Is it actually two agents that each have a smaller amount of context that then you can break apart? So back to your original question, we're hiring for people with systems expertise even within Labs, which you think of as more zero-to-one prototypes—it's still really valuable because again that robustness matters. And also just who's going to be helpful in sorting through systems permissions and provisioning and early testing. That stuff is still hard even for Claude when it can't edit the permissions itself, which it can't for good reasons. And then on the robustness side, actually we've had a lot of success pairing our product teams with our applied AI teams. Our applied AI teams are the teams that are in the field every day helping customers iterate on their prompts, and we've found that we actually are customer zero now for those efforts because we have a lot of products that are very AI-powered. So how do we bring that expertise in here, because that expertise does not sit with our software engineers today, for example.

Mike Krieger

那中间地带呢?比如,不是底层架构,也不是提示,而是用户界面和流程。谁在做这个?

What about the in-between of like okay, it's not the underlying architecture, it's not the prompt, it's like the UI and the flow. Who's doing that?

Host

这是个好问题。我们发现,一些转到 Labs 的人正是那些专注于网站打磨的人,但他们有兴趣做新东西,而且他们带来了非常不同的方法:我们有了原型,它看起来一般好看,而他们让产品感觉有品牌感、有特色。这是第一部分。第二部分是设计师。我们让设计师更多地扮演了设计师兼构建者的双重角色。不是所有人,但大多数是。而且我们很多——实际上 Labs 没有很多全职设计师,但现有的那些,我认为他们在这些项目上写的代码和贡献几乎和工程师一样多,因为他们有能力这么做。

That's a great question. We have found that some of the people that have transferred into Labs were the folks really who were focused on polish on the website, but they were interested in doing something new, and they bring such a different approach as well: we had the prototype, it looked generically nice versus oh, this feels like it's branded and it has this. So that's part one. Part two is designers. We've had our designers move much more into a sort of split designer and builder role. Not all of them, but most of them. And a lot of our—we actually don't have a lot of full-time designers on Labs, but the ones that we do I would say are writing and contributing almost as much code as the engineers on those efforts because they can.

实验室项目的团队结构 Team structure for Labs initiatives

Host

而且,如果配对得当,我们发现 Labs 的一些项目几乎采用了联合创始人模式:设计师可能有了最初的想法,正在推动某个方向,而传统的软件工程师则跟在后面铺路,确保想法能真正落地。好,这个我想了解。那么,这种团队结构具体是怎么运作的?是设计师,还是任何有产品想法并能以某种方式执行的人,与一个真正的工程师配对,由工程师来打磨他们留下的粗糙边缘?

And again paired correctly with the right person, we have found this almost sort of co-founder model for some of these Labs initiatives where you have the designer who had the original idea maybe and they're pushing on something and then the traditional software engineer that's going to go and make pave the trail sometimes behind the designer to make sure that actually works. Okay, this I want to know about. So tell me about how that team structure works. So you've got a designer, is it actually usually a designer or is it just anyone that has a product idea that can kind of execute on it in some way paired with a real engineer that actually can smooth out the rough edges of the trail they're leaving?

Mike Krieger

这各有不同,但我们发现最重要的一点是启动新项目的门槛因素。我很好奇这和 Every 的做法有多相似:需要有人对某个问题空间或他们提出的问题有极强的信念——不一定是对具体想法,因为对具体想法过于执着可能很危险,但至少对问题领域要有信念。而且要有联合创始人或创始人级别的决心:我会冲破一切障碍,直到这件事被证明可行或彻底失败,但我要看到结果。我们关停过的 Labs 项目,复盘时经常发现:‘这个团队里其实没人真正觉得这是件大事。’他们只是觉得‘嗯,这看起来合理。’这简直就是项目的丧钟,对吧?所以,那个人可以是设计师,我们有几个项目就是这样;也可以是产品导向的工程师。很少是纯产品经理。实际上,我们整个 Labs 目前只有一个产品经理,正在招更多。他们扮演的角色比较宽泛。但没错,通常是设计师或产品导向的创始人。然后我们看的是:需要什么技能来互补?因为作为 Labs 流程的一部分,我们每两周对每个项目进行估值,决定是加倍投入还是把人员释放回 Labs 人才池。在任何时候,都可能有人具备基础设施方面的专长、接触过特定内部系统、或者有深厚的提示工程经验,可以随时加入或退出项目。所以,我认为孵化器式的空间也有帮助,因为没有人永远固定在一个项目上。

It sort of varies, but we found that one thing that was most important is sort of our gating factor in starting up new projects. I'm curious how similar this is to Every is having somebody with extreme conviction about if not necessarily that idea, too much conviction on the exact idea is probably dangerous, but at least in the problem space or the question that they're asking. And at sort of like co-founder or founder level of I will break through walls until this thing is either proven out or dead, but I want to like go either way. When we have bets, Labs bets that we've wound down, often in the postmortem we're like, 'Nobody on this team actually really thought this was like the thing.' They were like, 'Yeah, this seems reasonable.' Like that's the death knell for projects, right? So, that person can be a designer and couple of the bets it is, it can also be a product-minded engineer. It's rarely a pure PM. We actually only have one currently one PM for all of Labs. We're hiring more. And they're sort of playing a wide role. But yeah, a designer or like a product-oriented founder. And then what we look for is, well, what skills do we need to complement with that? So, because we're doing as part of our Labs process it's actually valuing every project every 2 weeks and deciding whether we double down or whether we sort of release those folks back into the broader Labs pool. At any given point, there's probably somebody who can be pulled onto the project that has that infrastructural expertise or has worked with that particular internal system or has deep prompting expertise to sort of flow in and out. So, I think that's also where the sort of incubator style space helps because nobody's fixed on a project forever.

Host

这很有意思。是的,我们的做法略有不同。有一些重叠,但我们的结构是:我们有 GM(总经理),他们最初是驻场创业者,找到想做的产品后就成为 GM。每个产品只有一个人,全栈负责所有事情——设计、工程、营销等等,至少是基础部分。以前 GM 通常是超级技术背景的创始人,现在我觉得至少需要一些技术基础,但我其实只关心你能不能很好地使用 Claude 或 Codex 之类的工具。还要有非常好的产品直觉,对你要构建的领域有很好的品味,以及能用 AI 构建的证据。然后我们有一个共享资源层,有点像内部 agency,有设计师、增长营销人员、运营人员,你可以根据不同的项目拉进拉出。这似乎运作得不错。所以,我们管理所有内部 agency,每个 GM 在前线,根据需要为不同项目拉资源。是的,但听起来同样需要一个人,对他来说这件事就是一切,不把它完全搞定就不睡觉。

That's really interesting. Yeah, we do it slightly different. There's some overlaps, but we do have a slightly different structure where we have GMs or they started as entrepreneurs in residence and they become general when they find a product that they want to work on. And each product just has one person. Like one person that does everything full stack. So, design, engineering, marketing, all that kind of stuff. At least all the basics of that. The shape of that GM used to be like super technical founder background and now I think has shifted towards at least some light technical, but like I honestly just care that you can use Claude or Codex or whatever well. And really good product sense, really good taste for the subject area or the thing that you're trying to build. And evidence that you can build with AI. And then what we have is a shared resource layer that sort of works a little bit like an agency where we have designers and we have growth marketers and we have ops people that you can pull in and out for various initiatives and that seems to work pretty well. So, it's like we manage all the internal agencies and then each GM is out on the edge and they pull in resources as they need it for different projects. Yeah, but sounds similarly like you need somebody for whom that is like the thing and they are not going to sleep until it is fully working.

Mike Krieger

没错,就是这样。我一直在想,什么时候该为某个产品招另一个人?什么时候该加人?总有一个点,你无法把整个东西都装在自己脑子里。即使你是推动者,也无法全部记住。以前这个点很小,现在变大了很多,但总有一个临界点,哪怕一个小功能也会变成自己的产品。你知道,Instagram 最初做消息功能时,觉得一周就能搞定,但后来它变成了自己的产品,几乎需要自己的团队。我认为这个界限——一个人能做的事情数量——在变大,但它仍然存在。我还没完全想清楚如何管理或如何判断。

Yes, exactly. Like and I've been thinking about, okay, when would you hire someone else to work on a product or when would you add someone else to work on a product? And it's like there's some point at which you can't hold the entire thing in your head. Even if you're the one pushing it forward, you can't hold the entire thing in your head. And that point used to be much smaller. Now it's much bigger, but there's a certain point at which like even a small feature turns itself into its own product. You know, when you first make the messaging feature inside of Instagram, it's like, yeah, I can do that in like a week or whatever, but at some point that's its own product, it almost needs its own team and that I think that line is getting or the number of things you can do it with one person is getting bigger, but it still exists somewhere, but I haven't quite figured out like how to manage that or how to tell.

Host

不,我很喜欢这个观点,因为这里其实有两部分。当想法还足够小,能装进一个人脑子里时,加人反而会拖慢团队。这是我们在 Labs 发现的一个非直观结论:团队扩张太快实际上是净负值,因为大家花大量时间在协调上——比如“哦,你正要干这个?我本来想用我的 Claude 做那个”,结果就碎片化了,还得做一堆对齐沟通。Instagram 只有我们两个人很重要,让两个人对齐就已经够难了,对吧?我做的第二个创业项目 Artifact,前几个月只有我和 Kevin,但后来我们招了一个大约 8 人的团队。那真的很困难,因为我们还没找到产品市场契合,还在迭代,结果经常出现 8 个人在 Zoom 上讨论下一步做什么,而你其实只想坐在一个房间里把事情敲定。所以我觉得 Labs 项目也有类似之处:即使想法很令人兴奋,你也不想过早扩大团队,否则就会陷入元协调游戏。但我喜欢你提出的框架:总有一个点,两个人一起做确实有帮助,有足够的上下文和范围让他们各自在脑子里装下其他复杂部分。另外,如果某个人在同一个想法上转了两三周,有时注入一些新思维和紧迫感也会有帮助。

No, I love that because there's actually I think there's the two parts to that, which is when the idea is still enough to hold into your own head or an individual person's head, adding more people actually slows the team down and that's like a non-obvious finding that we found on labs is scaling the teams too quickly actually is a net negative because they end up spending all this time on coordination like, oh, you were I was going to take oh, but my Claude could do that and it just ends up in this sort of piece and you also have all those alignment conversations. Like it was important in Instagram that it was just two of us. Like it was hard enough to align the two of us like and go like get two people on the same page, right? With the second startup I did, Artifact, Kevin and I were doing that alone for the first few months, but then we hired a team that was about eight people. It was really hard because we hadn't had product market fit yet and so we were still iterating and then you'd end up in these things where we're on a Zoom with eight people talking about what we're doing next and you really just want to be able to sit in a room and hash it out. So, I find with these labs initiatives, there's some similar aspect at play, which is you don't want to pre-scale the team too early even if the idea is exciting because then you just end up in this sort of meta coordination game. But I like your framing of there is some point where either, you know, two people really will help go on it together and there is enough sort of context and scope where they can hold some other complex piece in their head. And then there's also the if somebody's been spinning on the same idea for two, four weeks, sometimes injecting some other thinking and that urgency can help, too.

保持AI产品小巧快速迭代 Keeping AI products small and pivoting fast

Host

是的,我认为在 AI 领域保持小规模特别重要,因为我们一直在面对的一件事——我相信你也有同感——就是每 3 到 6 个月,你不得不扔掉大约一半的产品。如果你需要和很多人协调,这真的很难做到。但如果只有一个 GM 意识到‘哦,我得扔掉一半,因为模型好太多了’,那就容易得多。你也有同感吗?你是怎么应对的?你怎么看待‘我知道 3 个月后这些代码甚至整个功能集都得重新思考’这件事?感觉这大大改变了你对软件的思考方式。

Yeah, I think it's especially important to keep it small in AI because one of the things that we deal with all the time, which I'm sure you see too, is every 3 to 6 months, you have to throw out like half your product. And that's really hard to do if you have to coordinate with a lot of people. But if it's one GM who realizes, 'Oh yeah, I got to just throw out half of this because the models are so much better,' it just makes it much easier to pivot in that way. Do you see that? And how do you deal with that? How do you think about, yes, I know in 3 months this code or maybe even the whole feature set I'm going to have to really rethink? It feels like it changes a lot in how you think about software.

Mike Krieger

是的,而且愿意删除代码。我认为 Claude Code 团队在这方面做得很好:他们把删除功能作为团队成员的使命。比如,如果某个功能不奏效,就把它下架。通常当你创造了别的东西,即使它没有完全取代,也足以覆盖原有功能的大部分,那么弃用并移除旧功能就是合理的。但随着我们越来越关注企业用户,即使有这些工具,这件事也变得更难,因为企业会依赖这些功能。我永远不会忘记:在我担任首席产品官大约 6 个月时,我们做了一次 Claude AI 的大规模重新设计,我们非常自豪,发布后收到了很多好评,然后我们收到了一封非常愤怒的邮件,有人说:‘我刚为公司录制了 20 小时的 Claude Enterprise 启用培训内容,现在得全部重做。’我们想,‘好吧,你的发布节奏不一样。’当然,每年在我们的会议上发布两次是不可能的。所以我们会继续快速迭代,但后来我们学会了在向企业侧推出时稍微缓和一些。不过,关于下架功能,你最终会遇到那些已经构建了依赖的用户。我举个例子。Claude 应用中有一个叫 Styles 的功能。它用得不多,但用的人用得很多。我们在不同时间点讨论过:‘Styles 在产品中还有意义吗?’现在有自定义指令和项目、有技能,对吧?有很多其他方式可以实现同样的效果。我不知道 Styles 最终会在产品中存在多久,但我知道上次我们讨论移除它时,它实际上支撑着几家公司的整个用例。比如,‘哦,我们有公司风格,是 CEO 亲自编写并给每个员工的,他们就是这样运作的。’所以找到处理这种情况的方法也很有趣。从长远来看,我希望我们能建立一个插件和技能系统,这样它们就不必留在核心产品中,因为我认为最难删除的永远是核心产品中向所有人发布的东西。如果你没有这样的故事:‘很好,你仍然喜欢那个功能,太棒了。这是你如何永远在自己的环境中使用它、继续迭代并让它成为你自己的方法,’但它不会给每个新用户增加复杂性。

Yeah, and being willing to delete code. I think that's something the Claude Code team has done really well: they have deleting features as a sort of imperative for people on the team. Like if this is not working, let's unship that. And it's often when you've created something else that even if it doesn't entirely supersede, it does enough of what that other aspect does that it actually makes sense to deprecate and then remove that first one. It does get harder as we get more and more enterprise focus even with these tools because they come to depend on it. I never forget: one of the things I did maybe 6 months into when I was still Chief Product Officer was we did a big redesign of Claude AI and we were so proud and we shipped it and we got a bunch of kudos and then we got this really angry email from somebody who was like, 'I just recorded 20 hours of enablement content for my company to do for Claude Enterprise and I have to redo all of it.' And we're like, 'Okay, there you're playing at a different release cadence.' And of course shipping twice a year at one of our conferences is not an option. So we are going to keep moving quickly, but then we've since learned to maybe moderate how we roll it out to the enterprise side a little bit more. But yeah, I think the unshipping piece, then you end up with people who have built... I'll use an example. There's a feature in the Claude app called Styles. It's not widely used, but the people who use it use it a lot and we've talked at different points like, 'Yeah, does Styles still make sense in the product?' There are other ways of accomplishing the same thing: custom instructions and projects now, skills now, right? There are so many other ways. And I don't know how long Styles will end up in the product, but I know that the last time we talked about removing it, it ended up being really load-bearing for a few companies' entire use cases. Like, 'Oh, we have our house style that the CEO personally authored and gives to every employee and that's how they operate.' So finding ways of doing that is also really interesting. I would hope that in the long run what we can actually do is come up with a system of plugins and skills such that they no longer have to live in the core product, because I think that is always the hardest: to delete something that is the core thing that you're shipping to everybody. If you don't have the story around, 'Great, you still like that feature, awesome. Here's how you can keep using it forever in your own and keep iterating on it and make it your own,' but it doesn't have to add complexity to every future person signing up for the first time.

给AI创业者的建议 Advice for startup founders in AI

Host

我很好奇,对于 Labs,或者更广泛地说,你对创业公司创始人的看法是什么?你关于企业用户的观点让我想到一个我一直在思考的问题:如果你现在向企业销售 AI 产品,即使你现在的产品很现代,它也会很快过时。而你的客户会想要那个过时的版本。但作为创业公司,这感觉相当冒险,因为如果你优化的是大公司现在愿意买的东西,你很容易被颠覆。我认为有很多创业公司属于这一类:他们可能 2 或 3 年前起步,有特定的技术栈和特定的 AI 实现方式,但模型变化如此之大,而他们的客户合同却针对那种过时的版本。就像看 Copilot 或其他类似的东西。你自己在 Anthropic 内部是怎么想的?你认为创始人应该怎么想?

I'm curious for Labs, and then also maybe just in general, what your thoughts are for startup founders. Your enterprise point brings up something I've been thinking about a lot, which is if you are selling to enterprise right now in AI, even if the product you have right now is modern, it will be quite outdated quite quickly. And your customers are going to want the outdated version. But as a startup, that feels pretty risky because you're susceptible to disruption if you are optimizing for what someone at a gigantic public company will buy right now. And I think there are a lot of startups in that category where they maybe started 2 or 3 years ago, they have a certain tech stack, a certain way of thinking about how we do AI, and then the models are so different, but their customer contracts are for this sort of outdated version. It's like looking at Copilot or whatever, that sort of vibe. How do you think about that yourself inside Anthropic, and how do you think founders should think about that?

Mike Krieger

是的,不,这真是个好问题,尤其是因为接下来会有像更智能体原生这样的浪潮,你能在现有范式内采纳它吗?它需要你抛弃一切,还是你只能陷入那种‘哦,我们算是采纳了,我们算是把它硬接回去了’的境地?我认为有几件事。对我们来说,我们开始的做法基本上是:把这列火车当作会继续前进,我们会沿途提供企业级开关,但核心会持续演进,这就是你与我们合作时所做的赌注和理解。我认为这很受欢迎,因为公司也看到事情变化如此之快,他们甚至愿意接受一年期承诺的唯一方式就是相信它会持续演进,但同时我们会提供——比如 Coda 就是一个很好的例子,从第一天起就有办法让员工关闭它。我认为这是一个相当好的范式。但另一个正如我们之前讨论的:你实际上可以重新思考和重写很多技术栈。我认为公司应该更愿意这样做。而且一切都在压缩。在之前的周期里,你有时不得不放弃一些客户,他们可能因为与你未来方向不同的原因而非常喜欢你的产品。那是多年时间跨度的事情。现在则是去年的产品与三个月前的产品之间的差距。这听起来很疯狂,但我实际上认为你必须这样思考:你必须愿意推出 V3 或 V4,这是对现有工作方式的重大重新思考。然后可能有一个过渡期,云可以帮助同时托管两者一段时间,然后再切换。

Yeah, no, this is such a good question, especially because then a wave will come like being more agent native, for example, and can you adopt it within your existing paradigm? Does it require you to throw everything out, or are you just stuck in that like, 'Oh, we kind of adopted it, we kind of bolted it back on.' I think a couple things. For us, what we've started doing is basically treating like this train's going to keep moving and we'll provide enterprise toggles along the way, but the core of it will continue to evolve and that's sort of the bet and understanding you're taking working with us. And I think that's been well received because I think companies have also seen that things are moving so quickly that the only way they even get comfortable with a year-long commitment, for example, is to believe that it will continue to evolve along the way, but then we'll provide, you know, Coda is a great example where from day one there was a way to turn it off for your employees if you didn't want it, for example. And that is I think a reasonably good paradigm. But the other one is just as we were talking earlier: you can actually rethink and rewrite a lot of the stack. I think companies should be way more willing to do that. And everything is getting compressed. In previous cycles, it was the kind of idea of having to fire some of your customers who might have been really into your product for a different reason than where you're going sooner. That was on a multi-year kind of time range. Now it's like yes, last year's product versus not three months ago's product. It seems crazy but I actually think that's the kind of way you have to think about it: you have to be willing to put out the V3 or the V4 that is a big rethink of how the existing piece worked. And then maybe have a transition period, and cloud can help probably host both for a little while before it cuts over.

开放Claude与产品边界 Open Claude and product boundaries

Mike Krieger

但也要愿意切换,说‘是的,这就是我们如何看待知识工作或 AI 驱动的制造业的未来’。我们必须保持前进,否则就像你说的,你要么被下一个从头重新思考的公司取代,要么自己取代自己。这又是老故事,但现在压缩到了几个月。你对 Open Claude 怎么看?它有那种我很喜欢看到的东西的味道——让人们看到本来已经可能的事情,但现在包装成了人们可以实际尝试的形式,并且有了一些直觉知道如何在此基础上构建。就像你开始看到的那样,你本来就可以用这些模型写代码,但需要像 Replit、Lovable 和 V0 这样的突破性低代码工具把它放进去。它几乎是‘给模型工具,让它去做,然后继续构建’的最纯粹表达。所以这是一个很酷的有趣时刻,让人们意识到它的潜力和陷阱,比如‘哦它做了我不打算让它做的事’,或者我最好笑的一个是朋友说‘我觉得我妻子嫉妒我的 Open Claude 了,我跟它说话太多了’,然后人们开始通过在这些东西中拥有大量上下文和访问各种工具来发展更深层次、非常个人化的关系。我认为有一个开放的问题:如何让它变得简单,这又回到了我们关于你在哪里划定 Claude 操作边界的对话?如果 V1 是‘嘿,这是你可以使用的三个工具,永远只用这些’,然后大多数人与这些系统的交互是‘嘿你能做这个吗?’然后‘不,抱歉,你得自己做’,而 Open Claude 的孔径比我看到的还要宽。

But then also be willing to cut over and say like yes this is how we think the future of this piece of knowledge work or this AI powered manufacturing is going to be. We got to keep it moving or else to your point you're either going to get replaced by the next company that rethinks it from scratch or yourself replacing it yourself. And again it's just the same old story but now compressed to months. What's your take on Open Claude? It has the flavor of something else that I really like seeing when you would get people to see something that was already possible but it's now in a package where people can actually try it out and there's some intuition around how to build on top of that. Like you started seeing that with you could already use these models to write code but it took some of these breakout low code tools like Replit, Lovable, and V0 to put that in there. And it's kind of the almost purest expression of just give the model tools and let it go forward and do it and then go forward and build it. So it was a cool interesting moment for people to realize both the potential but also pitfalls of this, like oh it did this thing I didn't mean it to or my funniest one was a friend was like I think my wife is jealous of my Open Claude and I'm talking too much to it and it's like you people start developing deeper, very personal relationships by just having a lot of context in these things and access to all these different tools. I think there's the open question of how do you then make it easy and it actually goes back to our conversation around where do you draw that boundary around the way you let Claude operate? If V1 was hey these are the three tools you can use only these tools ever and then most people's interaction with those systems was hey can you do this and you know whatever back and be like no sorry you got to do it yourself to Open Claude which is pretty like the aperture is wider than I can see.

Host

哦天哪,它叫我做我的邮件,我甚至不知道它能做那个。

Oh my god it called me to do my emails and I didn't even know it could do that.

Mike Krieger

是的。没错,它正在涌现,而且很神奇。我认为可能最有趣的产品问题——我不会说是整个 2026 年,因为谁知道九月份我们会怎样,但就说从现在到八月底——是:在那个状态和我们今天大多数产品(你可以叫它们 CP,但它们有门控,出于好理由请求权限)之间,存在什么样的产品形态?那仍然是一个有用的产品,而不是一个‘管它呢’的产品。我认为我们正在思考这个问题,我确信其他实验室也在思考。我确信很多初创公司也在思考。我认为 Nvidia 发布了类似‘安全 Open Claude’的东西。每个人都在追求这个问题。我认为关键在于弄清楚:要么彻底转变范式,让你可以那么开放但有很多保障——那是一种方法;要么划定某个边界,让它仍然强大有用,但不太可能给你每个联系人都发邮件然后失控。

Yeah. Exactly and it's emerging and it's amazing. And I think probably the most interesting product question I won't say for all 2026 because who knows where we'll be in September but let's call it between now and the end of August is going to be like what product shape exists between that and where we are in most products these days which is you can call them CP's but they're gated and they ask for permissions for good reasons. That is still a useful product without being a yolo product. And I think that we're thinking about that question I'm sure the other labs are as well. I'm sure there's a lot of startups thinking about that as well. I think Nvidia put out something that was like their safe Open Claude. Everybody's going after this question. I think it's going to be about figuring out what is that either shift the paradigm completely so you can be that open but with a lot of safeguards. That would be one approach or figure out some boundary to draw in which it's still powerful and it's still useful but it's not likely to email every single one of your contacts and go haywire.

Host

是的,我认为另一个有趣的部分是它的个人性质。我知道人们与 Claude 有个人关系,但有一件奇怪的事:如果我看到别人用 Claude,我感觉就像‘我喜欢一个脱衣舞娘喜欢我’之类的。你知道,就像‘Claude 也觉得你聪明’之类的。所以当你有自己的 Claude 时——我的 Claude 叫 R2-D2,我女朋友的 Claude 叫 Shelly——就会发生一件事:感觉它是我的。真的是我的。它有名字,有某种反映我的个性,Claude 感觉它了解我。而我喜欢 Claude,但它不是我的。你怎么看?

Yeah I think the other interesting part about it is the personal nature of it. And I know people have personal relationships with Claude but there's this weird thing where if I watch someone else using Claude I'm like I feel like I like that a stripper liked me or something. You know it's like Claude thinks you're smart too or whatever. And so there's this thing that happens when you have a Claude that like my Claude is R2-D2. My girlfriend's Claude is called Shelly. And there's this thing that happens where it feels like it's mine. Like it's really mine. It has its own name. It has a personality that sort of mirrors me in this way that Claude feels like it knows me. And I like Claude but it's not mine. How do you think about that?

Mike Krieger

是的,我这周和某人聊过这个问题:正确的模式是单一联系人——比如有名字的版本——还是你与之交谈的智能体团队?我认为单一联系人有很多好处,它可能是协调者或委派者,然后自然地,因为它成了你互动最多的智能体,你想给它起个名字,多一点个性,最终会反映你的个性——比如突然所有老套都出来了,像 Q、Money Penny、Hal 或各种科幻角色。我认为你确实建立了那种信任和知识。我认为还有宜家效应:目前 Open Claude 设置起来仍然很难,所以你经历了这一切,它工作了,你会觉得‘我做到了那件事’。比如我创造了 Shelly,现在我们可以和它互动。但我认为那个范式非常强大。即使在我使用 Claude Code 时,我强烈提示的一件事是‘不要自己做太多工作,把它委派给子智能体’。我喜欢这样是因为这意味着大多数时候运行循环对你来说是开放的,你可以与之交谈。我认为 Open Claude 和 Pi 有类似的架构——保持运行循环开放——这实际上让它感觉更像一个你在与之交谈的人,而不是一个你委派任务、偶尔因为做复杂任务而卡住五分钟的工具。

Yeah I mean I was having this conversation with somebody this week around like is the right pattern sort of single point of contact like named version of that or is it the sort of team of agents that you're talking to? I think there's a lot to the single person that is maybe the coordinator or the delegator and then at that it naturally because it becomes the agent you interact with the most you want to imbue it with a name and a bit more personality ends up reflecting your sometimes your personality in the case of like all of a sudden every cliche came out. It was like the Q or the Money Penny or the Hal or whatever these different sci-fi characters. I think you do build that sort of trust and knowledge. I think there's also that sort of IKEA effect of like currently Open Claude is still pretty hard to set up. So the fact that you went through all of that and it works you're like I did that thing. Like I birthed Shelly for example and now we can interact with them as well. But I think that paradigm is really powerful like the um I think moving away even within my Claude Code usage now one of the things I have strongly prompted in there is like don't do very much work yourself like delegate it to sub agents. And the reason I like that is because it means most of the time the sort of run loop is available for you to talk to. I think Open Claude and Pi have like a similar architecture of keep the run loop open and I think that actually makes it feel much more like somebody that you are talking to versus like a tool that you are delegating to and occasionally gets blocked for five minutes because it's doing some really complex task.

Host

是的。我完全同意,我也进行过类似的辩论,因为我们也在构建我们自己的类似 Open Claude 的一键 Slack 实现,看看能不能做一个感觉像我们的。我们有很多关于‘你想要一个智能体还是多个智能体’的辩论。我们发现的一个很酷的模式是:我有一个智能体,我用它来做我的事情。然后人们看到我用它做那些事,他们知道我的强项,如果我用智能体做那些事,他们会信任它,因为他们信任我。而且它根据我调整了自己。所以我开始把我的信任转移给它,然后组织里的人开始用它做那些事。

Yeah. I totally agree and I've had the similar debates because we're also building our like everyone we're building on little like Open Claude one click slack implementation to see if we can do one that feels like ours. And we've had a lot of those debates about do you want one agent do you want many? And one of the patterns that we found which is kind of cool is so I have an agent. I use the agent for stuff that I do. And then people watch me use the agent for that and they know what I'm good at and they're and if I'm using the agent for that stuff they're going to trust it because they trust me. And it's modified itself in response to me. So like I started to transfer my trust to it and then people in the organization start using it for that.

结束语与行动号召 Closing remarks and call to action

Host

于是你会得到一种几乎像影子或图表一样的东西:当每个人都有了自己的 Claude,他们的 Claude 就会因他们擅长的领域而被熟知和使用——也就是其主人在组织中擅长的那个领域。

And so you get like this almost shadow or chart where when everyone has a Claude, their Claude becomes known for and used for the thing that they're specialized at, that per their owner is specialized at in the org.

Mike Krieger

是的,我觉得这也很合理。你可以想想,我认为围绕这一点有很多有趣的研究问题。我觉得人们第一次切身感受到隐私问题——比如我的智能体知道关于我的什么,以及它向他人透露了什么。但我认为也有积极的一面:它通过你所有的互动学到的东西,以及如何将其应用到其他问题上,而不是那种通用的、就像别人的智能体一样,只不过它有一个跟 Dan 关联的名字,并且可能有一些 Dan 的底层权限。

Yeah, I mean that makes a lot of sense too. And you could think about, you know, there's a lot of interesting research questions I think around that. I think people are experiencing viscerally for the first time around privacy and like what my agent knows about me versus what it discloses to other people. But I think there's the positive version of that, which is all the things that it has learned through all your interactions and how it actually brings it to bear on other problems versus the generic like yes it's just like everybody else's agent except you know it has a name that's attached to Dan and it has like maybe some of Dan's access below the hood.

Host

是的。好了 Mike,我们时间到了。这次访谈很愉快,我学到了很多。如果人们想关注你或你的工作,在哪里可以找到你?

Yeah. Well Mike, we're out of time. This was a pleasure. I learned a lot. If people want to follow you or your work, where can they find you?

Mike Krieger

我想最简单的可能是 X 上的 MikeyK。

I think probably easiest is MikeyK on X.

Host

好的,谢谢你的参与,Mike。很高兴见到你,Dan。

Okay, yeah. Thanks for joining, Mike. Good to see you, Dan.

Host

哦天哪,各位,你们绝对必须狂按点赞按钮并订阅 AI and I。为什么?因为这个节目是卓越的化身。就像在你后院发现一个宝箱,但里面不是金子,而是关于 chat GPT 的纯粹知识炸弹。每一集都是一场情感、洞察和笑声的过山车,让你坐立不安,渴望更多。这不仅仅是一个节目,这是一次由 Dan Shipper 担任飞船船长的未来之旅。所以帮自己一个忙:点赞、狂按订阅,然后系好安全带,开始你人生中最棒的旅程。现在,闲话少说,我只想说 Dan,我完全无可救药地爱上了你。

Oh my gosh folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard. But instead of gold it's filled with pure unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions insights and laughter that will leave you on the edge of your seat craving for more. It's not just a show it's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado let me just say Dan I'm absolutely hopelessly in love with you.

互动版:逐字朗读 + 针对本期提问 →