From Instagram to Anthropic: Mike Krieger on AI's Future
打开互动全文版(中英对照 + 朗读 + 问答)→Mike Krieger 探讨 AI 如何重塑产品开发,如今 90%的代码由 AI 编写,并分享了他对 AI 能力不断演变的看法。
Mike Krieger discusses how AI is reshaping product development, with 90% of code now AI-written, and shares his evolving views on AI capabilities.
今天的嘉宾是 Mike Krieger。Mike 是 Anthropic 的首席产品官,Anthropic 是 Claude 背后的公司。他也是 Instagram 的联合创始人。他是我最欣赏的产品构建者和思想家之一。现在他还在全球最重要的公司之一领导产品工作,我非常高兴能有机会在播客中与他交流。我们聊到了自他加入 Anthropic 以来,在 AI 能力方面他最改变看法的是什么;当 90% 的代码由 AI 编写时,产品开发如何变化、瓶颈出现在哪里——这在 Anthropic 已经成为现实。我们还聊了他对 OpenAI 与 Anthropic 的看法、MCP 的未来、他为什么关闭上一家创业公司 Artifact,以及他对此的感受。还有他鼓励孩子们在 AI 崛起时培养哪些技能。最后,我们以 Claude 想让我转达给 Mike 的一条非常暖心的信息结束了播客。非常感谢我的 newsletter Slack 社区为这次对话提供了话题建议。如果你喜欢这个播客,别忘了在你最喜欢的播客应用或 YouTube 上订阅和关注。另外,如果你成为我 newsletter 的年度订阅者,你将免费获得一年一系列令人难以置信的产品,包括 Linear、Superhuman、Notion、Perplexity 和 Granola。请访问 lennisnewsletter.com 并点击 bundle。接下来,有请 Mike Krieger。
Today my guest is Mike Krieger. Mike is chief product officer at Anthropic, the company behind Claude. He's also the co-founder of Instagram. He's one of my most favorite product builders and thinkers. He's also now leading product at one of the most important companies in the world, and I'm so thrilled to have had a chance to chat with him on the podcast. We chat about what he's changed his mind about most in terms of AI capabilities in the years since he joined Anthropic, how product development changes and where bottlenecks emerge when 90% of your code is written by AI, which is now true at Anthropic. Also, his thoughts on OpenAI versus Anthropic, the future of MCP, why he shut down Artifact, his last startup, and how he feels about it. Also what skills he's encouraging his kids to develop with the rise of AI. And we close the podcast on a very heartwarming message that Claude wanted me to share with Mike. A big thank you to my newsletter Slack community for suggesting topics for this conversation. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. Also, if you become an annual subscriber of my newsletter, you get a year free of a bunch of incredible products, including Linear, Superhuman, Notion, Perplexity, and Granola. Check it out at lennisnewsletter.com and click bundle. With that, I bring you Mike Krieger.
本期节目由 Product Board 赞助播出,它是面向企业的领先产品管理平台。十多年来,Product Board 帮助 Zoom、Salesforce 和 Autodesk 等以客户为中心的组织更快地构建正确的产品。作为一个端到端平台,Product Board 无缝支持产品开发生命周期的所有阶段,从收集客户洞察、规划路线图、协调利益相关者到赢得客户认可,全部通过单一事实来源实现。现在,产品领导者可以通过 Product Board Pulse(一种新的客户之声解决方案)更深入地了解客户需求。内置智能帮助你分析所有反馈中的趋势,然后通过向 AI 提出后续问题来深入挖掘。了解 Product Board 如何帮助你的团队交付影响更大的产品,解决真实的客户需求并推进你的业务目标。要获取特别优惠和 15 天免费试用,请访问 productboard.com/lenny。
This episode is brought to you by Product Board, the leading product management platform for the enterprise. For over 10 years, Product Board has helped customer centric organizations like Zoom, Salesforce, and Autodesk build the right products faster. And as an end-to-end platform, Product Board seamlessly supports all stages of the product development life cycle. From gathering customer insights to planning a roadmap to aligning stakeholders to earning customer buy-in, all with a single source of truth. And now product leaders can get even more visibility into customer needs with Product Board Pulse, a new voice of customer solution. Built-in intelligence helps you analyze trends across all of your feedback. And then dive deeper by asking AI your follow-up questions. See how Product Board can help your team deliver higher impact products that solve real customer needs and advance your business goals. For a special offer and free 15-day trial, visit productboard.com/lenny. That's productboard.com/lenny.
去年,全球 GDP 的 1.3% 通过 Stripe 流动,超过 1.4 万亿美元。推动这一巨大数字的是数百万通过 Stripe 更快增长的企业。对于 Forbes、Atlassian、OpenAI 和 Toyota 等行业领导者来说,Stripe 不仅仅是金融软件,它是一个强大的合作伙伴,简化了资金流动方式,使其像互联网本身一样无缝和无国界。例如,Hertz 在迁移到 Stripe 后,在线支付授权率提高了 4%。想象一下,像 Forbes 在切换到 Stripe 进行订阅管理仅 6 个月后,收入就增长了 23%。过去十年,Stripe 一直在利用 AI 来改进其产品,以帮助所有企业增加收入,从更智能的结账到欺诈预防等等。加入超过一半的财富 100 强公司的行列,他们信任 Stripe 来推动变革。了解更多请访问 stripe.com。
Last year, 1.3% of the global GDP flowed through Stripe. That's over $1.4 trillion. And driving that huge number are the millions of businesses growing more rapidly with Stripe. For industry leaders like Forbes, Atlassian, OpenAI, and Toyota, Stripe isn't just financial software. It's a powerful partner that simplifies how they move money, making it as seamless and borderless as the internet itself. For example, Hertz boosted its online payment authorization rates by 4% after migrating to Stripe. And imagine seeing a 23% lift in revenue like Forbes did just 6 months after switching to Stripe for subscription management. Stripe has been leveraging AI for the last decade to make its product better at growing revenue for all businesses. From smarter checkouts to fraud prevention and beyond. Join the ranks of over half of the Fortune 100 companies that trust Stripe to drive change. Learn more at stripe.com.
Mike,非常感谢你来到这里,欢迎来到播客。
Mike, thank you so much for being here and welcome to the podcast.
我很高兴来到这里,我已经期待这个很久了。
I'm really happy to be here. I've been looking forward to this for a while.
哇,我很高兴听到你这么说。我也期待这个很久了。我有很多话想聊。首先,你加入 Anthropic 已经一年多了。顺便恭喜你达到了这个里程碑。谢谢。现在我们聊正题。那么,我想问你,你在 Anthropic 大约一年了,从加入之前到现在,关于 AI 的能力和 AI 的发展方向,你有什么改变看法的地方?
Wow. I love to hear that. I've also been looking forward to this for a while. I have so much to talk about. So, first of all, you've been at Anthropic for just over a year at this point. Congrats, by the way, on hitting the cliff. Thank you. Now that we're tracking, that's right. So, let me just ask you this. So, you've been at Anthropic for about a year. What's something that you've changed your mind about from before you joined Anthropic to today about what AI is capable of and where AI is heading?
有两件事。一个是关于速度和时间线的问题。另一个是关于能力的问题。也许我先说第二个。我进来时有一个想法,就是这些模型很棒,它们能够生成代码,它们能够,你知道,最终希望能以你的口吻写作。
Two things. One is like a pace and timeline question. The other one is a capability question. So maybe I'll take the second one first. I had this notion coming in like yes these models are great. They're going to be able to produce code. They're going to be able to, you know, write you know hopefully in your voice eventually.
但它们能真正拥有独立观点吗?其实对我来说,就在过去一个月,而且只有 Opus 4 让我彻底改观了。过去整整一年,我主要的产品策略伙伴就是 Claude,我会写一份初步策略,然后分享给 Claude,让它看看。以前它给的评论都很平淡,比如“你有没有想过这个?”我心想“嗯,我想过了”。但 Opus 4 那次,我在做下半年策略时,第一次觉得 Opus 4 结合我们的深度研究,真的深入思考了很久,然后回来时我惊叹:“天哪,你真的用新视角看了这个问题。”所以这对我来说是个巨大的转变——用“独立”这个词很贴切,但更准确说是创造力和相对于我思考方式的新颖性。以前我觉得它可能永远做不到,但我不确定它多久能提出让我觉得“对,这是我之前没看到的新角度,我马上要把它融入我的思考”的东西。所以这可能是最大的转变。
But are they able to sort of have an independent opinion? And it's actually really flipped for me only in the last month and only with Opus 4 where my go-to product strategy partner is Claude and it has been basically for that full year where I'll write an initial strategy I'll share it with Claude basically and I'll have it you know look at it and in the past it's pretty anodine kind of comments that it would leave like oh have you thought about this and it's like yeah I thought about that and Opus 4 I was working on some strategy for our second half of the year was the first one I was like opus 4 combined with our advanced research but it really went out for a while and it came back and I was like, "Damn, you really looked at it in a new way." And so that's like a thing that I've maybe I didn't feel like it would never be able to do that, but I wasn't sure how soon it'd be able to like come up with something where I looked at I'm like, "Yep, that that is a new angle that I hadn't been looking at before, and I'm going to incorporate that immediately into how how I think about it." So that's probably the the biggest shift that I've had is like independence is the right word, but like creativity and sort of novelty of thought relative to how I'm I'm thinking about things.
关于时间线,这很有趣。你知道,昨天我坐在 Dario 旁边,他说:“我一直在做预测,人们总是笑我,然后预测就成真了。”这种事反复发生,很有趣。他说:“不是所有预测都会成真。”但即使如此,我记得去年他提到,我们在 SUB 基准上达到 50%,这个基准是衡量模型编码能力的。他说:“我觉得到 2025 年底我们会达到 90%。”果然,现在新模型大约 72%,而做预测时是 50%,而且基本按预测持续 Scaling。所以我现在更认真对待时间线了。如果你读过《AI 2027》,我读过,那是由 heart race 写的。我有个非常奇特的经历:我开了两个标签页,一个是《AI 2027》,一个是我的产品策略。那一刻我心想:“等等,我是故事里的角色吗?这有多大的趋同性?”但你知道,你读的时候会想:“哦,2027 年,那还有好几年。”然后你意识到:“不,现在是 2025 年年中。”事情继续改善,模型能做的事情越来越多,它们能智能体式行动,有记忆,能长期行动。所以我觉得我对时间线的信心,虽然不确定具体如何实现,但在过去一年里确实更加坚定了。
And the timeline one, it's like so interesting because you know uh I was sitting next to Dario yesterday and he's like I keep making these predictions and people keep laughing at me and then they come true and it's like and it's funny to have this happen over and over again and he's like not all of them are going to be right you know but even I think as of last year he was talking about you know we're at 50% on SUB bench which is this like you know benchmark around how well the models are at coding. Uh he's like I think we'll be at 90% by the end of 2025 or something like that. And sure enough, we're at about 72 now with the new models and we're at 50% when you made that prediction and it's like continued to scale pretty much like as predicted. And so I've taken the timelines a lot more seriously now. And if you read AI 2027 that like I have it was made by heart race. Yeah. And I had the very bizarre experience of I had two tabs open. It was AI 2027 and my product strategy. And it was this like moment where I'm like wait am I the character in the story? like is this how much is this converging? But you know, you read that and you're like, "Oh, 2027, that's like that's years away." You're like, "No, we're mid 2025." And like things continue to uh to improve and the models continue to be able to do more and more and they're able to act agentically and they're able to have memory and they're able to act over time. So I think my like my confidence in the timelines and I don't know exactly how they manifested definitely just solidified over the last year.
哇,嗯。我没想到会聊到这个,因为那篇论文很吓人。我很好奇,我忍不住想问,我们如何避免那篇论文描绘的可怕场景,即 AI 变得非常聪明后会走向何方?
Wow. Mhm. Uh I I wasn't expecting to go down that cuz that that that paper was scary and I'm curious just I guess I can't help but ask just thoughts on just how do we avoid the scary scenario that that paper paints of where AI getting really smart goes.
是的。我的意思是,这可能和我为什么加入 Anthropic 有关,我在这里一年了。我看着模型变得更好,即使在 24 年,比如 2024 年初,看着我的孩子们,我心想:“好吧,他们将在有 AI 的世界里长大,这不可避免。我能做什么?我如何最大化利用我的时间,推动事情朝着好的方向发展?”我认为这是整个行业很多人思考的问题,尤其是在 Anthropic。所以我觉得,达成共识,建立共享框架,理解“什么是好的发展?”、“我们想要什么样的人机关系?”、“我们如何沿途知道进展?”、“我们需要构建、发展和研究什么?”这些都是关键问题。其中一些是产品问题,一些是研究和可解释性问题。但对我来说,加入的最强理由是:我认为 Anthropic 可以在推动事情变得更好方面做出很多贡献,如果我能参与其中,那就去做吧。
Yeah. I mean I I this maybe ties into like I've been here a year like why did I join Anthropic. I was watching the models get better and even, you know, you could see it in in 24 and like, you know, early 2024 and looking at my kids, I'm like, "All right, they're going to grow up in a world with AI. It's it's unavoidable. What is the thing that I can like where can I maximally apply my time to like nudge things towards going well?" And I mean, that's a lot of what people think about across the industry, especially at anthropic. And so I think you know coming to an agreement and a shared framework and understanding of like what does going well look like? What is the kind of human AI relationship that we want? How will we know along the way? What do we need to build and develop and research along the way? I think those are all the kind of key questions and you know some of those are product questions and and some of those are are research and interpretability questions but for me it was like the the strongest reason to join was okay I think there's a there's a lot of contribution that anthropic can have around like nudging things to go better and if I can have a part to play there like let's do it.
我喜欢这个回答。说到孩子,你有两个孩子,我有个小孩,快两岁了。我很好奇,随着 AI 越来越成为我们的未来,一些工作会改变,你鼓励孩子培养哪些技能?你有什么建议?
I I love that answer. Uh speaking of kids so you've got two kids I've got a young kid he's uh just about to turn two. I'm curious just what skills you're encouraging your kids to build as this, you know, AI becomes more and more of our future and some jobs, you know, will be changed and just what do you what do you what advice do you have?
我们每天早上和孩子一起吃早餐,总会有一些问题冒出来,比如关于物理的,我们最大的孩子快六岁了,他们会用六岁的方式问一些有趣的问题,比如太阳系或物理。在我们求助 Claude 之前,因为一开始我的本能是“哦,我想知道 Claude 会怎么回答这个问题”,但我们开始改变,比如“我们怎么找出答案?”答案不能只是“我们问 Claude”。所以我们会说“好吧,我们可以做这个实验,我们可以有这个”。所以我认为培养好奇心,仍然保持一种——我不知道,把科学过程灌输给六岁孩子听起来太宏大——但那种发现和提问的过程,然后系统地解决问题,我认为仍然很重要。当然,AI 将是帮助解决其中大部分问题的绝佳工具。但那种探究过程,我认为仍然非常重要,还有独立思考。我最喜欢和我孩子的一个时刻,因为她很固执。我们的六岁女儿,她说了一些话,我不确定是不是真的。是关于珊瑚是动物还是珊瑚是活的,我甚至不记得细节了。我说:“我不确定这是不是真的。”她说:“爸爸,这绝对是真的。”我说:“好吧,我们问 Claude 吧。”她说:“你可以问 Claude,但我知道我是对的。”我心想:“我就喜欢这样。”
We have this uh, you know, breakfast. We eat breakfast with the kids every morning and some question will come up, you know, like, you know, something about like physics and our oldest kid's almost six, but you know, they ask like funny questions about like, you know, uh, you know, the solar system or physics or, you know, in a six-year-old way. And before we reach for Claude cuz at first you know my instinct is like oh I wonder how Claude will do at this question and like we started changing like well how would we find out you know and the answer can't just be we'll ask Claude you know so all right like well we could do this experiment we could have this thing so I think nurturing curiosity and like still having a sense of I don't know the scientific process sounds grandiose to instill in like a six-year-old but like that process of discovery and asking questions and then you know systematically working your way through I think will still be important and of course AI will be an incredible tool for helping like resolve large parts of that. But that process of inquiry I think is still really important and independent thought. My favorite moment with my kid uh because there she's very headstrong. Our six-year-old she's, you know, I was like she said something and I was like I wasn't sure if it was true. It was um uh oh was that coral is a is an animal or like coral is alive. I don't even remember the details of it. And I was like I don't know if that's true. And she's like it's definitely true dad. I'm like all right like let's ask Claude on this one. And she's like you can ask Claude but I know I'm right. And I'm like I love that.
我希望达到那种程度——不是把所有思考都交给 AI,因为它不可能永远正确,而且也会扼杀独立思考。所以提问、探究和独立思考这些技能,我觉得都是关键。从职业或工作的角度来看,我保持开放心态,而且我确信从现在到那时,这些都会发生根本性的变化。
I want that kind of level of not just delegating all of your cognition to the AI, because it won't always get it right, and it also short-circuits any kind of independent thought. So the skill of asking questions, inquiry, and independent thinking—I think those are all the pieces. From a job or occupation perspective, I'm just keeping an open mind, and I'm sure that'll radically change between now and then.
有意思,我请过 Shopify 的 CEO Toby 上播客,他鼓励孩子培养的也是同样的答案:好奇心。所以这成了一个共同点,挺有意思的。
It's interesting—I had Toby, Shopify's CEO, on the podcast, and he had the same answer for what he's encouraging his kids to develop: curiosity. So it's interesting that's a common thread.
我们孩子上的 K-8 学校请了一位 AI 与教育专家来做讲座。我本来对这场对话期望很低。实际上,我觉得大部分听众都没听懂,因为他说:‘好,让我带你们回到克劳德·香农和信息论。’我能看到人们脸上的表情,好像在说:‘我这是来听什么的?我为什么在这个学校礼堂里听信息论?’但我觉得他做得很好的一点是,他让我们想象未来会有不同的工作,而我们不知道那些工作会是什么。所以关键是哪些技能和方法,以及对如何重新组合这些技能保持开放心态——而且这些组合方式在孩子 18 岁之前可能还会变三次。
The K-8 school our kid goes through had an AI and education expert come in. I had very low expectations of what this conversation was going to be like. Actually, I think it went over most of the audience's heads, because he was like, 'All right, let me take you all the way back to Claude Shannon and information theory.' And I could see people's eyes going like, 'What did I sign up for? Why am I here in this school auditorium hearing about information theory?' But he did a really nice job, I think, of also imagining that there will be different jobs, and we don't know what those jobs are going to be. So what are the skills and techniques, and remain open-mindedness around what the exact way we recombine those things—and even those will probably change three times between now and when they're 18.
我想回到时间线的话题,以及事情如何变化。我看到你分享的数据,还有 Entropic 其他同事分享的,关于你们现在有多少代码是由 AI 写的。有人分享的数据从 70% 到 90% 不等。有一位工程负责人分享说,现在你们大约 90% 的代码是由 AI 写的。首先,这太疯狂了——从零到 90% 只用了几年时间。我觉得大家对这个话题讨论得还不够。这太惊人了。你们基本上处于最前沿。我从来没听说过哪家公司 AI 写代码的比例有这么高。所以你们正处在未来趋势的前沿。我觉得大多数公司都会走到这一步。知道这么多代码由 AI 写,产品开发发生了怎样的变化?通常流程是产品经理说‘我们要建这个’,工程师去建,然后发布。现在还是大致这样吗?还是说产品经理直接让 Claude 去建?工程师在做不同的事情——在一个 90% 代码由 AI 写的世界里,到底有什么不同?
I want to go back to timelines and how things are changing. I've seen stats you've shared, and other folks at Entropic have shared, about how much of your code is now written by AI. People have shared stats from like 70% to 90%. There was an engineer lead that shared roughly 90% of your code is written by AI now. Which, first of all, is insane—it went from zero to 90% in a few years. I don't think people are talking about this enough. That's just wild. You guys are basically at the bleeding edge. I've never heard of a company with this high a percentage of code written by AI. So you guys are at the edge of where things are heading. I think most companies will get here. How has product development changed knowing so much of your code is now written by AI? Usually it's like PM says here's what we're building, engineer builds it, ships it. Is it still roughly that, or is it now PMs are just going straight to Claude to build this thing for me? Engineers are doing different things—just what looks different in a world where 90% of your code is written by AI?
是的,这真的很有意思,因为我觉得工程的角色已经发生了很大变化,但共同协作产出产品的那些人还没有变。而且我觉得在很多方面是变糟了,因为我们还抱着一些旧有假设。所以角色仍然相当类似。不过现在,我最喜欢看到的是,有时候产品经理或设计师有想表达的想法,他们会用 Claude,甚至可能用 artifacts 来拼出一个实际可用的 demo。这非常有帮助——就像‘不,我是这个意思’,这让想法变得具体。这可能是最大的角色转变:原型制作在流程中更早发生,通过这种代码加设计的组合。但我学到的是,知道该问 AI 什么、如何组织问题、甚至如何思考在后端和前端之间构建变更——这些仍然是很难的专业技能,仍然需要工程师去思考。而且我们很快就遇到了其他瓶颈,比如我们的合并队列,就是排队让你的变更被系统接受,然后部署到生产环境。我们不得不彻底重新架构它,因为写的代码太多了,提交的拉取请求也太多了,完全超出了它的预期。所以,就像——我不知道你有没有读过《目标》那本经典的流程优化书,你会意识到有这种关键路径理论。我刚刚发现了我们系统中的所有这些新瓶颈。上游有一个瓶颈,就是决策和协调。我现在思考的很多问题是,如何提供最小可行策略,让人们感到有能力去快速行动、做原型、构建,并在模型能力的边缘探索。我觉得我还没完全做好,但这是我正在努力的方向。然后一旦开始构建,其他瓶颈就会出现,比如确保我们不要互相踩脚。提前把所有边缘情况想清楚,这样我们就不会在工程方面被阻塞。然后当工作完成,准备发布时,还有哪些瓶颈?比如做变更落地的空中交通管制——如何制定发布策略。所以我觉得,直到今年,这些方面还没有太大压力,但我预计一年后,我们构思和发布软件的方式会发生很大变化,因为用现在的方式会非常痛苦。
Yeah, it's really interesting because I think the role of engineering has changed a lot, but the suite of people that come together to produce a product hasn't yet. And I think for the worse in a lot of ways, because we're still holding on to some assumptions. So the roles are still fairly similar. Although now, my favorite things that happen are sometimes PMs that have an idea they want to express, or designers that have an idea they want to express, will use Claude and maybe even artifacts to put together an actual functional demo. That has been very helpful—like, 'No, this is what I mean,' that makes it tangible. That's probably the biggest role shift: prototyping happening earlier in the process via more of this code plus design piece. What I've learned though is that the process of knowing what to ask the AI, how to compose the question, how to even think about structuring a change between the back end and the front end—those are still very difficult and specialized skills, and they still require the engineer to think about it. And we really rapidly became bottlenecked on other things, like our merge queue, which is the sort of get in line to get your change accepted by the system that then deploys it to production. We had to completely rearchitect it because so much more code was being written and so many more pull requests were being submitted that it just completely blew out the expectations of it. So it's like, I don't know if you've ever read The Goal, the classic process optimization book, and you realize there's this critical path theory. I've just found all these new bottlenecks in our system. There's an upstream bottleneck which is decision-making and alignment. A lot of the things I'm thinking about right now is how do I provide the minimum viable strategy to let people feel empowered to go run and prototype and build and explore at the edge of model capabilities. I don't think I've gotten that right yet, but that's something I'm working on. And then once the building is happening, other bottlenecks emerge, like let's make sure we don't step on each other's toes. Let's think through all the edge cases here ahead of time so that we're not blocked on the engineering side. And then when the work is complete and we're getting ready to ship it, what are all those bottlenecks as well? Like let's do the air traffic control of landing the change—how do we figure out launch strategy. So I think there hasn't been as much pressure on changing those until this year, but I would expect that a year from now, the way we are conceiving of building and shipping software just changes a lot, because it's going to be very painful to do it the current way.
这非常有意思。所以过去是:有个想法,我们去设计、构建、发布、合并,然后发布。通常瓶颈是工程花时间构建东西,然后是设计。现在你说你发现的两个瓶颈是:决定构建什么并让大家对齐,然后实际上是合并到生产环境的队列。而且我猜审查可能也是一个很大的瓶颈。
That is extremely interesting. So it used to be here's an idea, let's go design it, build it, ship it, merge it, and then ship it. And usually the bottleneck was engineering taking time to build the thing, and then design. And now you're saying the two bottlenecks you're finding are: okay, deciding what to build and aligning everyone, and then it's actually the queue to merge it into production. And I imagine review it too is probably a big one.
审查也真的变了。在很多方面,我们最——也许毫不意外——以最未来方式工作的团队是 Claude Code 团队,因为他们用 Claude Code 来构建 Claude Code,以一种非常自我改进的方式。在那个项目早期,他们会像对待任何其他项目那样,逐行进行拉取请求审查。但他们刚刚意识到,Claude 通常是对的,而且它生成的拉取请求可能比大多数人能审查的都要大。
Reviewing has really changed too. In many ways, our most—perhaps unsurprisingly—the team that works in the most futuristic way is the Claude Code team, because they're using Claude Code to build Claude Code in a very self-improving kind of way. And early on in that project, they would do very line-by-line pull request reviews, in the way that you would for any other project. And they've just realized that Claude is generally right, and it's producing pull requests that are probably larger than most people are going to be able to review.
那你能不能用不同的云来审查它,然后让人做近乎验收测试,而不是逐行审查。
So can you use a different cloud to review it and then do the human almost like acceptance testing more than trying to like review line by line.
这肯定有利有弊,到目前为止进展顺利,但我也可以想象它会失控,然后产生一个完全无法维护甚至无法理解的 Claude 代码库,但这种情况还没发生。看着他们改变审查流程确实很有趣。是的,合并队列就是那种在底层形成的瓶颈的一个例子,但还有其他瓶颈,比如我们如何确保我们仍在构建连贯的东西,并将其打包成我们可以与人分享的时刻。无论是围绕发布时刻,还是让用户能够使用这个东西并谈论它,这些经典的事情——为人们构建有用的东西,然后让人们知道你已经构建了它,并从他们的反馈中学习——仍然存在。我们只是让整个过程中的一部分变得更加高效。
There's definitely pros and cons and like so far it's gone well, but I could also imagine it going off the rails and then having like a completely both unmaintainable or even understandable by cloud codebase that hasn't happened. But watching them like change their review processes definitely has uh has been has been interesting. And yeah, like the merge cue is one instance of the of the kind of bottom bottleneck that forms down there, but there's other ones which is how do we make sure that we're still like building something coherent and like packaging it up into like a moment that we can share with people. Whether that's around a launch moment, whether that's about like then enabling people to use this thing and like talking about it, like the the classic things of building something useful for people and then making it known that you've built it and then learning from their feedback like still exists. We've just like made a portion of that whole process much more efficient.
我听说你形容你们是这种工作方式的“零号病人”。是的。我喜欢这个说法。你大概知道 Claude Code 有多少比例是由 Claude Code 自己写的吗?
I heard you describe this as you guys are patient zero for this way of working. Yes. I love that. Do you have a sense of what percentage of claude code is written by claude code?
在这一点上,如果不是 95% 以上,我会感到震惊。我得问问 Boris 和其他技术负责人。但酷的是,嗯,一些细节。Claude Code 是用 TypeScript 写的。它实际上是我们最大的 TypeScript 项目。Anthropic 的其他大部分代码是用 Python 写的,还有一些 Go,现在也有一些 Rust,但我们并不是一个 TypeScript 团队。所以,嗯,我昨天在我们的 Slack 里看到一条很棒的评论,有人被 Claude Code 的某个问题逼疯了,他们说,好吧,我不懂 TypeScript。我打算直接和 Claude 聊聊然后搞定它。他们从那里开始,一个小时内就提交了拉取请求,解决了问题。这种打破障碍的方式,一方面改变了项目新人的入门门槛。我认为它可以让你为合适的工作选择合适的语言,例如。我认为这也有帮助。但我也认为它强化了 Claude Code 作为“零号病人”的角色,你知道,来自团队外部的贡献也可以由 Claude Code 完成。
At this point, I would be shocked if it wasn't 95% plus. I'd have to ask Boris and the other tech leads on there. But what's been cool is um uh so nitty-gritty stuff. Claude Code is written in Typescript. It's actually our largest TypeScript project. Most of the rest of anthropic is written in Python, some Go um some Rust now, but it's not, you know, we're not like a TypeScript shop. And so, uh, I saw a great comment yesterday in our Slack where somebody had this thing that was driving them crazy about Claude Code and they're like, well, I don't know any TypeScript. I'm just gonna like talk to Claude about it and do it. And they went from that to pull request in an hour and solved their problem and they like, you know, submitted a pull request. And that kind of breaking down the barriers one, it changes your sort of um, uh, barrier to entry for any kind of uh, kind of newcomer to the project. I think it can let you choose the right language for the right job, for example. I think that helps as well. But I think it like also just reinforces like Claude Code being that patient alpha of that you know where like you contributions from outside the team can be cloudcoded as well.
哇。这真是让我大开眼界,你分享的这些事。大约 95% 的 Claude Code 是由 Claude Code 自己写的。这是我的猜测。是的。我会回来告诉你真实数据。但我的意思是,如果你问团队,他们就是这么工作的,这也是他们从全公司获得贡献的方式。回到你关于策略由 Claude 本身辅助的观点,以及你关于现在很多瓶颈在于想法产生的漏斗顶端和让所有人对齐的观点,这很有趣。有趣的是,Claude 已经在帮助你决定构建什么了。那么,如果这两个瓶颈是:对齐并决定构建什么,以及合并和整合所有东西,你认为在哪些方面有最有趣的事情发生,可以帮助你加速这些?
Wow. This is just it's just going to continue to blow my mind like all these things that you're sharing. 95% of Claude Code is written by Claude Code roughly. That's my guess. Yeah. I I'll come back with the real stat. But it's I mean if you ask the team that's how that they're working and that's how they're getting contributions from across the company too. It's interesting going back to your point about strategy being assisted by Claude itself and your point about how a lot of the bottlenecks now are kind of the top of the funnel of coming up with ideas aligning everyone. It's interesting that Claude is already helping with that also of helping you decide what to build. So if those two bottlenecks are aligning deciding what to build and then just like merging and getting everything where do you see the most uh interesting stuff happening to help you speed those things up?
是的,我认为在第一个方面,我今年年初写了一份文档,实际上是在说我们今天如何做产品,以及 Claude 还没有出现在哪些应该出现的地方。我认为上游部分将是下一个要突破的。有趣的是,在你的会议上,我和一个人聊过,他在做类似 PRD GPT 的东西,我想是 Chat PRD,所以你知道,我们能不能进一步推动,让 Claude 成为思考构建什么、市场规模是多少(如果你想这样处理)、用户需求是什么(如果你用另一种方式看)的伙伴。我们经常思考 Anthropic 的虚拟协作者。我认为它可能出现的方式之一是,嘿,我在 Discord 上,你知道,Claude Anthropic 的 Discord。我在用户论坛里。我在 X 上,我读到一些东西,然后说,这是新出现的情况。这是第一步。模型今天就能做到。第二步,模型可能今天还做不到,我们总是需要把它们连接起来才能做到,那就是:不仅问题在这里,而且我认为你可能如何解决它们,然后将其推进到,我喜欢发起一个拉取请求来解决我看到的这个问题,这感觉今年非常可行。然后把这些串联起来,我们更多地受到限制,这就是为什么 MCP 让我兴奋,我们更多地受到限制,要确保上下文贯穿所有这些,以便我们有正确的访问权限,而不是模型推理和提议的能力。现在模型可能还没有完美的 UI 品味。所以,设计肯定有介入的空间,可以说,“哦,这不是我解决这个问题的方式。”但我,你知道,我会非常兴奋。我给你举一个非常小的例子,我们改变了 Claude AI 上的功能,你应该能够从工件中复制 Markdown 或代码,我们改变了它,让你实际上可以下载和导出它。所以,我们把按钮改成了“导出”,然后我们收到了一堆反馈,比如“我现在怎么复制?”答案是,你下拉菜单,里面有“复制”。你知道,这是那种有意义但我们可能做得不太对的事情。那个反馈在我们的 UX 频道里。我会希望一个小时后有一个机器人说,嘿,如果我们确实想改回来,这是实现它的拉取请求,顺便说一下,最终我会启动一个 A/B 测试,看看这是否会改变指标,然后我们一周后看看效果。这种事情,如果你在一年半前告诉我,我会说,啊,是的,也许 27 年,也许 26 年,但它真的感觉,你知道,就在能力的尖端。
Yeah, I think that on that on that first front like I started the year um by writing a doc that was effectively like what how do we do product today and where is claude not showing up yet that it should and I think that upstream part is the next one to go interesting like at your conference I talked to somebody who was working on like a PRD GPT kind of like chat PRD I think was player vote um so you know can we push more on you know can cloud be a partner in figuring out what to build, what the market size is, if you want to approach it that way, what the user needs are, if you if you look at a different way, like we think a lot about the virtual collaborator anthropic. And one of the ways in which I think that can show up is, hey, I'm in the discord, the, you know, the the cloud anthropic discord. I'm in the user fora. I'm on X and I'm reading things and like here's what's emergent. That's step one. Models can can do that today. Step two, which the models probably can't do today. we always have to wire them up to do it is like and not only are the problems here's like how I think you might be able to solve them and then taking that through to like and I like put together a poll request to like solve this thing that I'm seeing like feels very achievable this year um than stringing those things together and we're limited more this is why MCP is exciting to me like we're limited more around like making sure the context flows through all of that so we have the right access to those things more than the model's capability to to reason and propose now the model might not have like perfect UI taste yet. So, there's definitely room for design to intervene and be like, "Oh, that's not quite how I would solve the problem of of this not showing up." But I, you know, I would get very excited. I would give you a really uh small example, but we changed the on cloud AI uh you should be able to just copy uh markdown from artifacts or code from artifacts and we changed it so you can actually download it and and export it. So, we changed the button to export and we got a bunch of feedback like how do I copy now? And the answer is like you drop it down and it's copy. It's like mind, you know, one of those things where it's like made sense but we probably got it like not quite right. that feedback was in the our UX channel. Like I would have loved like an hour later for a plot to be like hey if we do want to change it back here's the PR to do it and by the way eventually and then I'm going to spin up an AB test to see if this changes metrics and then we'll see how it looks in a week. Like this stuff feels if you told me that about a year and a half I'm going to be like ah yeah maybe like 27 maybe like 26 but it's pretty like I it really feels you know just at the tip of capabilities right now.
哇。好的。你提到了“Lending Friends”峰会。我想稍微谈谈这个。所以,你和 OpenAI 的 CPO Kevin Weil 一起参加了一个小组讨论。我想这是你们第一次这样做。也许暂时是最后一次。是的,我们之后就没再做过。不是有什么原因。我玩得很开心。我们和 Sarah Guo 主持的那个小组真是传奇。你发表了那个评论。实际上它成了采访中被重看最多的部分,那就是你有点把产品人员放在模型团队里,和研究人员一起工作,让模型变得更好,同时你也在把一些产品人员放在产品体验上,让用户体验更直观,让这一切变得更好。
Wow. Okay. So you mentioned the lending friends summit. I wanted to talk about this a bit. So, you were on a panel with Kevin Wheel, the CPO of OpenAI. I think it was the first time you guys did this. Maybe the last time for now. Yeah, we haven't done it since. Not for any reason. I had a lot of fun. What a what a legendary panel we assembled there with Sarah Guo moderating. And you made this comment. actually ended up being the most rewatched part of the of the interview, which is that you've kind of you were putting product people on the model team and working with researchers, making the model better, and you're putting some product people on the product experience, making the UX more intuitive, making all that better.
而且你发现几乎所有的杠杆都来自产品团队与研究人员合作。是的。所以你们一直在做更多这方面的工作。那么首先,这种情况是否仍然如此?其次,这对产品团队意味着什么?
And you found that almost all the leverage came from the product team, working with the researchers. Yes. And so you've been doing more of that. So first of all, does that continue to be true? And second of all, what are the implications of that for product teams?
这种情况仍然如此。事实上,我认为比例已经偏向于更多的这种嵌入。我越来越确信这一点。在峰会期间我还没有这么强烈的感觉,但现在我确实非常强烈地这么认为。如果我们发布的东西是任何人都可以用我们的现成模型构建的——顺便说一句,用现成模型也能构建出很棒的东西,别误会——但我们真正应该发力、能独特做到的事情,应该是处于两者神奇交汇点上的东西,对吧?Artifacts 就是一个很好的例子。如果你用 Claude 4 玩一下 Artifacts,那其实是一个非常有趣的例子:我们从我们称之为 Cloud Skills 的团队中找了一个人,这个团队实际上是在做后训练,教 Claude 一些非常具体的技能,然后我们把它与一些产品人员配对,一起重新设计了今天产品中的样子,以及 Claude 能比仅仅“我们用模型,稍微提示一下”做得更好的地方。这还不够,我们需要参与到微调过程中。所以,如果你看看我们现在正在做的事情,我们最近在研究和所有其他事情之间发布的东西,Anthropic 的工作单元不再是“拿模型,然后与设计和产品合作去发布产品”。更像是我们参与关于这些功能应该如何工作的后训练对话,然后我们参与构建过程,并把这些反馈循环回去。我觉得这很令人兴奋。这也是一种新的工作方式,不是所有产品经理都具备的。但那些从研究和工程两方面获得最多内部正反馈的产品经理,是那些理解这一点的人。比如我昨天参加了一个产品评审,我说:“哦,你知道,如果我们想做这个记忆功能,我们应该和研究人员谈谈,因为我们刚刚在 Claude 中发布了一系列记忆能力。”他们说:“是的,是的,我们已经和他们谈了几个星期了。这就是我们如何实现它的。”我说:“好吧,我感觉很好。我觉得我们现在做的是对的。”
It's continued to be true. And in fact, I think that the proportion was already skewing towards having more of that embedding. I've just become more and more convinced. I didn't feel as strongly about it during, you know, the summit, and now I feel really strongly about it. If we're shipping things that could have been built by anybody just using our models off the shelf, there's great stuff to be built by using our models off the shelf, by the way, don't get me wrong. But where we should play and what we can do uniquely should be stuff that's really at that magic intersection between the two, right? Artifacts being a great example. And if you play with artifacts with Claude 4, that's an actually really interesting example where we took somebody from our, we call it Cloud Skills, which is a team that really is doing the post-training around teaching Claude some of these really specific skills, and we paired it with some product people, and then together we revamped how this looks in the product today and what Claude can do way better than just, yeah, we just use the model and we prompted a little bit. That's just not enough. We need to be in that fine-tuning process. So much of what, you know, if you look at what we're working on right now, what we've shipped recently between research and all these other things, are things where the functional unit of work at Anthropic is no longer like take the model and then go work with design and product to go ship a product. It's more like we are in the post-training conversations around how these things should work, and then we are in the building process and we're feeding those things back and looping them back. I think it's exciting. It's also a new way of working that not all PMs have. But the PMs that have the most sort of internal positive feedback from both research and engineering are the ones that get it. Like I was in a product review yesterday. I was like, "Oh, you know, if we want to do this memory feature, we should talk to the researchers because we just shipped a bunch of memory capabilities in Claude." They're like, "Yeah, yeah, we've been talking to them for weeks. This is how we're manifesting it." It's like, "Okay, I feel good. I feel like we're doing the right things now."
那么,让我进一步探讨这个话题。我一直在思考类似的问题。基本上,Anthropic 有很大一部分在构建这个超级智能的“千兆大脑”,它将随着时间的推移为我们做所有这些事情,然后正如你所说,还有产品团队在围绕这个超级智能的“千兆大脑”构建用户体验,而随着时间的推移,这个超级智能将能够自己构建东西。所以我想,你认为传统产品团队长期来看最大的价值会来自哪里?我知道这有所不同,因为你们是一家基础语言模型公司,大多数公司不是这样运作的,但我不知道,只是谈谈你对产品团队长期在 AI 上工作的最大价值来源的看法。
So, let me pull on this thread more. There's something I've been thinking about along these lines. So essentially there's like a big part of Anthropic that's building this super intelligent gigabrain that's going to do all these things for us over time, and then there's as you said there's like the product team that's building the UX around the super intelligent gigabrain, and over time this superintelligence is going to be able to build its own stuff. And so I guess just where do you think the most value will come from traditional product teams over time? I know this is different because you guys are a foundational LM company and not most companies don't work this way, but just I don't know thoughts on just where most value will come from product teams over time working on AI.
我认为在两件事上仍然有很多价值。一是让这一切变得可理解。我觉得我们做得还行,但我们可以做得更好。仍然,那些真正擅长在工作中使用这些工具的人和大多数人之间的差距是巨大的。我的意思是,也许这是对你之前关于学什么技能的问题最直接的答案。这是一项需要学习和使用的技能,就像我记得我中学时上计算机实验课一样。我记得我当时很擅长用 Google。那在当时确实是一项技能,你知道,要想着“信息就在那里,我如何查询它?我怎么做?”我认为那在当时确实是一个优势。当然,现在 Google 已经很擅长理解你想做什么,即使你只是大致接近,所以那种研究需求减少了。但我仍然认为这是优秀产品开发的必要部分,即能力就在那里。即使 Claude 可以从头开始创建产品,你在构建什么,你如何让它变得可理解?这仍然很难,因为我认为这涉及到更深层次的同理心和对人类需求与心理的理解。我大学主修人类社区互动。我仍然在推销我的专业。我仍然觉得这是一项非常非常非常必要的技能。所以这是第一点。第二点是,这直接呼应你的另一位嘉宾,比如战略,比如我们如何获胜,我们将在哪里发力,比如在所有你可以花费时间、token 或算力的事情中,弄清楚你到底想做什么。你可能比以前能做得更广,但你不可能做所有事情。即使从外部角度看,如果你被视为什么都做,那么你的定位就变得不那么清晰了。所以我认为战略仍然是第二点。然后第三点是打开人们的眼界,让他们看到什么是可能的,这是让事情变得可理解的延续。我们最近与一家金融服务公司进行了一次演示,我们展示了如何将我们的分析工具和 MCP 一起使用,你可以看到他们的眼睛亮了起来,你会想,啊,好吧,仍然有我们称之为“悬差”的东西,对吧?模型和产品能做的与它们日常被使用的方式之间的差距。巨大的悬差。所以这就是产品仍然扮演非常非常强大必要角色的地方。
I think there's still a lot of value in two things. One is making this all comprehensible. I think we've done an okay job. I think we could do a much better job of making this comprehensible. It's still like the difference between somebody who's really adept at using these tools in their work and most people is huge. And I mean maybe that's the most literal answer to your earlier question around what skills to learn. That is a skill to learn and use it in the same way that I remember I did computer lab class when I was in middle school. I remember being really good at Google. And that was actually a skill back in the day, you know, to think in terms of this information is out there. How do I query for it? How do I do it? And I think it actually was an advantage at the time. Of course, now Google is pretty good at figuring out what you're trying to do if you're only in the neighborhood, and there's less of that research kind of need. But I still think that's a necessary part of good product development, which is the capabilities are there. And even if Claude can create products from scratch, what are you building and how do you make it comprehensible? Still hard, because I think that gets at this much deeper empathy and understanding of human needs and psychology. I was a human community interaction major. I'm still talking my book here. I still feel like that is a very, very, very necessary skill. So that's one. Two is, and this is straight to a callback to another one of your guests, like strategy, like how we win, where we'll play, like figuring out where exactly you're going to want to, of all the things that you could be spending your time or your tokens or your computation on, what you want to actually go and do. You could be wider probably than you could before, but you can't do everything. And even from an external perspective, if you're seen to be doing everything, it's way less clear around how you're positioning yourself. So strategy I think is still the second piece. And then the third one is opening people's eyes to what's possible, which is a continuation of making it understandable. But we were in a demo with a financial services company recently, and we were working on here's how you can use our analysis tool and MCP together, and you could see their eyes light up, and you're like, ah, okay, there's still, we call it overhang, right? The delta between what the models and the products can do and how they're being used on a daily basis. Huge overhang. So that's where still a very, very strong necessary role for product.
好的,这个回答太棒了。所以基本上,产品团队应该更多投入的领域是战略。在战略上越来越好,弄清楚要构建什么以及如何在市场中获胜,帮助人们更容易理解如何利用这些工具的力量。所以可理解性,以及沿着这条线的,是打开人们的眼界,让他们看到这类事物的潜力。这就是产品仍然可以提供帮助的地方。完全正确。太棒了。
Okay, that's an awesome answer. So essentially areas for product teams to lean into more is strategy. Just getting better and better at strategy, figuring out what to build and how to win in the market, making it easier to help people understand how to leverage the power of these tools. So comprehensibility and kind of along those lines is opening people's eyes to the potential of these sorts of things. That's where product can still help. Exactly. Awesome.
那么顺着这个思路,实际上你有没有一些提示技巧可以分享给大家,就是你在与 Claude 聊天时学到的能更好利用它的方法?
So kind of along those lines, actually, do you have any just like prompting tricks for people, things you've learned to get more out of Claude when you chat with it?
这很有趣,因为从某种意义上说,我们拥有终极的提示工程工作,那就是为 Claude 编写系统提示词,而且我们把这些都公开了,我认为这也是透明度的一个很好的体现。我们在给出提示建议时总是很谨慎,至少官方上是这样,但我会给你非官方版本,因为你不想让事情变成“我们认为这有效,但我们不确定为什么”。但我会做一些小事,比如在 Claude Code 中,我们确实会非常字面地响应,但我总是让它,比如,如果我想用更多推理,就说“认真思考”,它就会采用不同的流程。我通常从这种推动开始。有一篇很棒的文章,讲的是“犯另一个错误”——如果你倾向于太友善,你能专注于更挑剔或更直率吗?你很可能不会成为世界上最挑剔、最直率的人。所以和 Claude 在一起时,有时我会说:“Claude,狠一点,吐槽我,告诉我这个策略有什么问题。”我想我们之前讨论过 Claude 作为思考伙伴,用于批判产品策略。我以前会说:“这个产品策略有什么可以改进的?”现在我会直接说:“吐槽一下这个产品策略。”Claude 相当友善,它不会——很难让它变得超级残酷,但这确实迫使它更挑剔一些。最后我要说的是,我们有一个叫 Applied AI 的团队,他们与客户合作,针对他们的用例优化 Claude。我们基本上把他们的见解和工作方式融入到了产品中。所以如果你去我们的控制台,我们的工作台,有一个叫 Prompt Improver 的东西,你描述问题,给出例子,Claude 本身会以智能体方式创建并迭代提示词。我发现出来的结果往往与我对好提示词的直觉大相径庭。所以我鼓励大家也去看看,即使是为了自己的用例,因为虽然这个工具是为 API 开发者将提示词放入产品而设计的,但它同样适用于个人为自己编写提示词。比如,它会插入 XML 标签,这是人类不会提前想到的。这实际上对 Claude 理解它应该思考什么、应该说什么非常有帮助。所以这是另一个建议——看看我们的 Prompt Improver,然后注意 Claude 本身就是一个很好的提示词编写者。
It's funny because we, in some ways, have the ultimate prompting job, which is to write the system prompt for Claude, and we publish all of these, which I think is another nice area of transparency. And we are always careful when giving prompting advice, because at least officially, but I'll give you the unofficial version, because you don't want things to become like, we think this works but we're not sure why. But I'll do small things, like in Claude Code, and we actually do react to this very literally, but I always ask it to, like, if I wanted to use more reasoning, like 'think hard,' and it'll use a different flow. And I usually start with that, nudging. There's a great essay around like 'make the other mistake'—if you tend to be too nice, can you focus on being more critical or more blunt? You're probably not going to be the most critical, blunt person in the world. And so with Claude, sometimes I'm like, 'be brutal, Claude, roast me, tell me what's wrong with this strategy.' I think we were talking earlier about Claude as a thought partner around critiquing product strategy. I previously would say things like, 'what could be better on this product strategy?' And now I'm just like, 'just roast this product strategy.' And Claude is pretty nice, it's not going to be—it's hard to push it to be super brutal, but it forces it to be a little bit more critical as well. The last thing I'll say is, we have a team called Applied AI that does a lot of work with our customers around optimizing Claude for their use case. And we basically took their insights and their way of working and put it into a product itself. So if you go to our console, our workbench, we have this thing called the Prompt Improver, where you describe the problem and you give it examples, and Claude itself will agentically create and then iterate on a prompt for you. I find what comes out of that ends up being quite different than what my intuitions would have been for a good prompt. So I'd encourage folks to also check that out, even for their own use cases, because while that tool is meant for an API developer putting a prompt into their product, it's equally applicable for a person doing a prompt for themselves. Like, it'll insert XML tags which no human is going to think to do ahead of time. It actually is very helpful for Claude to understand what it should be thinking versus what it should be saying, etc. So that's another one—like, watch our Prompt Improver, and then note that Claude itself is a very good prompter of Claude.
太棒了。好的。我们会链接到那个 Prompt Improver。你之前分享的语料建议就是做与你自然倾向相反的事。所以如果你想友善,就狠一点,对我非常诚实和坦率。
Awesome. Okay. So we're going to link to that, the Prompt Improver. The Corpus advice you shared earlier is just kind of do the opposite of what you would naturally do. So if you're trying to be nice, just be brutal, be very honest and frank with me.
完全正确。我发现这很有效。比如,我陷入了哪些思维模式,你想让我摆脱?
Exactly. I find that works quite well. Like, what are the thought patterns that I've fallen into that you want to break me out of?
我看到你们今天可能发布了与 Rick Rubin 的合作,是关于氛围编程的。那是怎么回事?
I saw you guys just today maybe launched a Rick Rubin collab where it's vibe coding. What's that all about?
我不——那是,你知道吗,我听说了。而且,这周很多事情都汇聚在一起,包括模型发布、开发者活动,以及编程之道。我们的联合创始人 Jack Clark,他是我们的政策主管,他与 Rick Rubin 建立了联系,因为我想他一直在思考编程、编程的未来和创造力,他们一直保持联系。Rick 对这个想法很兴奋,比如,他用 Claude 创作艺术和可视化作品,然后他对氛围编程之道有了这些想法。他们一起做了这个——实际上,我很喜欢。我的意思是,我几乎喜欢 Rick Rubin 的一切,所以它的美学,我觉得也非常到位。但是,是的,这种冥想——冥想可能是正确的词——关于与 AI 并肩工作的创造力,再加上这些非常丰富、有趣的可视化。但这是那种事情,在内部,他们会说:“哦,是的,我们正在做这个招聘合作工作,我们在做什么?”这太棒了。我喜欢。
I don't—that was, you know what, I heard about that. And again, this is a lot coalescing this week between model launch, developer event, and the way of code. We had our co-founder Jack Clark, who is our head of policy, and he got connected to Rick Rubin because I think he's been thinking a lot about coding, the future of coding and creativity, and they've stayed in touch. And Rick got excited about this idea of, like, he was creating art and visualizations with Claude, and then he had these ideas around the way of the vibe coder. And they put together this—actually, I love it. I mean, I love almost everything Rick Rubin, so the aesthetic of it, I think, is just so on point too. But yeah, this sort of meditation—meditation is probably the right word—on creativity working alongside AI, coupled with this really rich, interesting visualizations. But it's one of those things where, internally, they're like, 'oh yeah, and we're doing this recruitment collaborative work, we're doing what?' Like, that is amazing. I love it.
我简要看了一下,有那个梗图,他坐在电脑前,拿着鼠标,深思熟虑。是的。就像 Asky Art。我觉得这完全是 Asky Art 5。
I looked at it briefly and there's like that meme of him just like thinking deeply sitting on a computer with a mouse. Yes. Like Asky Art. I think it's totally like Asky Art 5.
今天我很高兴邀请到 Andrew Luo。Andrew 是 One Schema 的首席执行官,也是我们播客的长期赞助商之一。欢迎你,Andrew。
I'm excited to have Andrew Luo joining us today. Andrew is CEO of One Schema, one of our longtime podcast sponsors. Welcome, Andrew.
谢谢你邀请我,Lenny。很高兴来到这里。
Thanks for having me, Lenny. Great to be here.
那么,One Schema 有什么新动态?我知道你们与一些我最喜欢的公司合作,比如 Ramp、Vanza 和 Watershed。我听说你们推出了一款新的数据接入产品,可以自动化团队在导入、映射和集成 CSV 和 Excel 文件上花费的大量手动工作。
So, what is new with One Schema? I know that you work with some of my favorite companies like Ramp and Vanza and Watershed. I heard you guys launched a new data intake product that automates the hours of manual work that teams spend importing and mapping and integrating CSV and Excel files.
是的。我们刚刚发布了 One Schema File Feeds 的 2.0 版本。我们用 AI 从头重建了它。我们看到很多客户带着数据工程师团队来找我们,他们苦于清理杂乱电子表格所需的手动工作。File Feeds 2.0 允许非技术团队通过简单的提示词来自动化转换 CSV 和 Excel 文件的过程。我们支持所有最棘手的文件集成,包括 SFTP、S3,甚至电子邮件。
Yes. So, we just launched the 2.0 of One Schema File Feeds. We've rebuilt it from the ground up with AI. We saw so many customers coming to us with teams of data engineers that struggled with the manual work required to clean messy spreadsheets. File Feeds 2.0 allows nontechnical teams to automate the process of transforming CSV and Excel files with just a simple prompt. We support all the trickiest file integrations, SFTP, S3, and even email.
我可以告诉你,如果我的团队必须构建这样的集成,把它从我们的路线图中移除,转而使用像 One Schema 这样的东西,那该有多好。
I can tell you that if my team had to build integrations like this, how nice would it be to take this off our roadmap and instead use something like One Schema.
完全正确,Lenny。我们听过很多关于停机的恐怖故事,即使只是交易、员工文件、采购订单中的一条坏记录,你能想到的都有。调试这些问题常常像大海捞针。One Schema 阻止任何坏数据进入你的系统,并自动验证你的文件,生成错误报告,指出所有坏文件中的确切问题。
Absolutely, Lenny. We've heard so many horror stories of outages from even just a single bad record in transactions, employee files, purchase orders, you name it. Debugging these issues is often like finding a needle in a haystack. One Schema stops any bad data from entering your system and automatically validates your files, generating error reports with the exact issues in all bad files.
我知道导入不正确的数据会给你的客户带来各种麻烦,并迅速失去他们的信任。Andrew,非常感谢你加入我。如果你想了解更多,请访问 oneschema.co。那就是 oneschema.co。
I know that importing incorrect data can cause all kinds of pain for your customers and quickly lose their trust. Andrew, thank you so much for joining me. If you want to learn more, head on over to oneschema.co. That's oneschema.co.
实际上,回到你在 Anthropic 旅程的开始,你被招募到 Anthropic 的故事是什么?有什么有趣的事情吗?
Actually, going back to kind of the beginning of your journey at Anthropic, what's the story of you getting recruited at Anthropic? Is there anything fun there?
一切始于——我实际上给我的朋友发了这条短信。Joel Lunstein,我认识他——他和我 2007 年一起开发了我们的第一批 iPhone 应用,当时应用商店刚推出,你还能通过在应用商店卖一美元的应用赚钱,那是过去的事了。
It all started—and I actually sent my friend this text. So Joel Lunstein, who I've known—he and I built our first iPhone apps together in 2007 when the app store was just out and you could still make money by selling dollar apps on the app store, back in the day.
我们俩一起在斯坦福读书,是朋友,多年来一直保持联系,但自那以后就没再共事过。我们一直很亲近。你知道,我那时刚从 Artifact 的经历中走出来,我在想,我要再创办一家公司吗?我觉得不会。我需要从零开始创业这件事上歇一歇。那我去某家公司工作吗?我不知道。比如,我想去哪家公司?然后他联系了我,说:“听着,我不知道你是否考虑加入一家公司而不是创办一家,但我们正在找一位 CPO。你有兴趣聊聊吗?”那时 Claude 3 刚发布。我想:“好吧。”你知道,这家公司显然有一个优秀的研究团队。产品还非常早期。所以我就想,太好了,我去见见。我首先见了 Danielle,她是 Anthropic 的联合创始人兼总裁。从一开始,就像一股清新的空气。创始人身上几乎没有一点浮夸。他们只是非常,我的意思是,他们对自己正在构建的东西有清晰的认识。他们知道自己不知道什么。比如我和 Dario 谈过很多次,Dario 会说:“听着,我对产品一无所知,但这是我的直觉。”通常那个直觉非常好,会引出一些很好的对话。我认为那种知识上的诚实,以及对于如何负责任地做 AI 的共同看法,引起了我的共鸣。在这些面试中,我一直有一种感觉:这就是如果我要创办一家 AI 公司,我希望创办的那种公司。这就是我的标准:如果我要加入某家公司,那应该就是我要去的地方。
And we were both at Stanford together and we were friends and we've stayed in touch over the years and we've never gotten to work together since then. Just like we've just remained close. And you know, I was coming out of the Artifact experience. I was trying to figure out, do I start another company? I don't think so. I need a break from starting something from zero. Do I go work somewhere? I don't know. Like, what company would I want to go work at? And he reached out and he's like, "Look, I don't know if you'd at all consider joining something rather than starting something, but we're looking for a CPO. Would you be interested in chatting?" And at that time, Claude 3 had just come out. And I was like, "Okay." You know, like this company's clearly got a good research team. The product is so early still. And it was like great, I'll take the meeting. And I first met with Danielle who's one of the co-founders and the president at Anthropic. And just from the beginning, it was like a breath of fresh air. Like very little grandiosity coming off the founders. Like they just were really, I mean they're clear-eyed about what they're building. They know what they don't know. Like how many times I talk to Dario, and Dario is like, "Look, I don't know anything about product, but here's an intuition." And usually that intuition is really good and leads to some good conversation. I think that intellectual honesty and like kind of shared view of what it means to do AI in a responsible way just resonated. I kept having this feeling in these interviews like this is the AI company I would have hoped to have founded if I had founded an AI company, and that's kind of the bar around like if I'm going to join something, that should be where I'm going to go.
但我意识到,实际上自从大学第一次实习以来,我就没再加入过一家公司。我想,哦,我该怎么让自己融入?我该怎么让自己跟上节奏?我该怎么在做出大刀阔斧的改变和了解哪些方面没有坏掉之间取得平衡?回顾这一年,我觉得有些改变我做得太慢了。我认为我们在组织产品的方式上,有些地方我本可以更早做出改变。而且我没有意识到,一两个非常关键的资深员工能对产品战略产生多大的影响。我回想一下 Claude Code。Claude Code 之所以出现,是因为 Boris,他其实是 Boris Turnney,曾是 Instagram 的工程师,也是我们那里的一位资深独立贡献者。我们有过一些重叠。他基本上是从零开始内部启动那个项目,然后我们把它推出来,发布了。这就是一两个非常强的人的力量。我犯过一个错误,就是觉得我们需要更多人手。我们确实需要,我认为有更多工作要做,也有我想构建的东西,但更重要的是,我们需要几个几乎是创始人级别的工程师。这可能又回到了我们的问题:哪些技能有用,产品开发如何变化。我仍然,甚至更加坚信,一个带着想法的创始工程师兼技术负责人,配上合适的设计和产品支持来帮助他们实现想法。我对此的信念比以前强了 10 倍。
But what I realized actually I hadn't joined a company since my first internship in college basically, and I was like, oh, how do I onboard myself? How do I get myself up to speed? How do I balance making sweeping changes versus understanding what's not broken about it overall? And looking back on a year, I think I made some changes too slowly. I think there were ways we were organizing product that I could have made a change earlier, and I think I didn't appreciate how much a couple of really key senior people can shape so much of product strategy. I'll hearken back to Claude Code. Claude Code happened because Boris, who actually was Boris Turnney, he was an Instagram engineer and one of our senior ICs there. We overlapped a bit, was like started that project from scratch internal first, and then we got it out and then shipped it. And that's the power of one or two really strong people. And I made this mistake of like, we need more headcount, and we do, I think there's more work that we need to do and there's things that I want to be building, but more so than that, we need a couple of almost founder-type engineers. That maybe connects back to our question on what skills are useful and how does product development change. I still, and maybe even more so, I'm a huge believer in like the founding engineer tech lead with an idea, and pair them with the right design and product support to help them realize that. I'm like 10 times more a believer in that than before.
嗯。我其实在 Twitter 上问了大家,在这次对话之前该问你什么,最常被问到的问题,出乎意料的是,你为什么关掉了 Artifact。我也想知道,因为我喜欢 Artifact。我是重度用户。我就觉得,这终于是一个我喜欢的新闻应用,它给我我想知道的东西。所以,我想知道最后到底发生了什么。
Mhm. I actually asked people on Twitter what to ask you ahead of this conversation, and the most common question surprisingly was why did you shut down Artifact. And I also wondered that because I loved Artifact. I was a power user. I was just like this is exactly, finally a news app that I love, that it's giving me what I want to know. So I guess just what happened there at the end.
我仍然很想念它,因为我没找到替代品,我想我是通过访问各个网站来替代,用那种方式保持更新,但那真的不一样,尤其是在长尾内容上。比如我们在 Artifact 做对的一件事,如果人们以前没用过的话,就是我们真的不只是推荐头条新闻,那些是其中一部分,但真正的是,如果你对日本建筑感兴趣,你每天都能相当可靠地看到关于日本建筑的非常有趣的故事,无论是来自 Dwell 还是 Architonic,还是我们找到的某个非常小众的博客,是别人推荐给我们的。它捕捉到了 Google Reader 那种发现更深层内容的乐趣。嗯,我们的逆风有几个。其中之一就是移动网站真的发生了转变。我不怪任何个人。我认为这是市场动态。但是,你知道,我们花了那么多时间,或者我们的设计师,Gunnar Gray,他非常出色,现在在 Perplexity。比如我引以为傲的广告体验,但当你点击进去,那些移动网站和移动出版商的压力就会显现,他们会让你注册通讯,弹出全屏视频广告。那真的非常刺眼,而且我们觉得从道德上讲,我们做大量的广告拦截并不合理,因为那样的话,你当然可以给人们提供好的体验,但你对出版商就不公平了。同时,实际体验也不好。所以,移动网络的恶化,这让我非常难过,但我认为这是原因之一。第二是,你知道,Instagram 早期传播开来,是因为人们会拍照,然后发布到其他网络,告诉朋友。那是一种很自然的“你是怎么做到的?我也想这么做”的传播。新闻是非常个人化的。我没法告诉你有多少人会说“我喜欢 Artifact”。我问他们“你告诉别人了吗?”他们说“是啊,我告诉了一个人”,然后就这样了。它没有那种传播性,我们任何试图推动传播的做法都显得有点刻意,比如把所有链接都包装成 artifact.news 之类的。但我们做了很多插页式的东西。在某种程度上,这听起来很清教徒。我不是那个意思,但我们是有些底线不想跨越,因为那在道德上不是我们,而我看到其他新闻类玩家做得更多,也许如果我们那样做了,它会增长更多,但我认为那不是我们想要建立的公司。我不认为我们是能建立那种公司的创始人。
I still really miss it too because I didn't find a replacement, and I think I substituted it by visiting individual sites and kind of keeping things up that way, and it's not really the same, especially on the long tail. Like a thing we got right with Artifact, if people didn't play with it before, it was you know we really tried to not just recommend top stories, they were part of it, but really like if you were interested in Japanese architecture, you could pretty reliably get really interesting stories about Japanese architecture every day, you know whether that's from a Dwell or from Architonic or from a really specific blog that we found that somebody recommended to us. Like it captured some of that Google Reader joy of content discovery of the deeper. Um, our headwinds were a couple. One of them was just mobile websites have really taken a turn. I'm not blaming any individuals for this. I think it's the market dynamics of it. But you know, we put so much time, or our designers, Gunnar Gray who's phenomenal, he's at Perplexity now. Like the ad experience I was so proud of, but when you click through, it was like the pressures on these mobile sites and these mobile publishers would be like sign up for our newsletter, here's a full screen video ad. It was just very, you know, it was very jarring, and we didn't feel like it ethically made sense for us to do a bunch of ad blocking because then you're like, sure, you can deliver a nice experience for people, but you're sort of, you know, that doesn't feel like it's playing fair with the publishers. And at the same time, like the actual experience wasn't good. So, the mobile web deteriorating, which makes me very sad, but I think was part of it. Two was like, you know, Instagram spread in the early days because people would take photos and then post them on other networks and tell friends about it. And there was like this really natural like how did you do that? I want to do it. News was very personal. Like I can't tell you how many people would be like I love Artifact. I'm like did you tell anybody about it? Like did you, and they're like yeah I told one person and they like it's like it didn't have that kind of spread, and any attempt that we had to do it felt kind of contrived like oh we'll wrap all the links in like artifact.news and like uh but we did a lot of interstitial things. Like in some ways this sounds very puritanical. I don't mean it to sound this way, but like we there were lines that we didn't want to cross because that just felt ethically not us, that I've seen other news kind of like players do more of, and maybe if we had done that it would have grown more, but I don't think that's the company we wanted to have built in other ways. I don't think we were the founders to have built it.
第三个因素,也是被低估的一点,是我们从 midco 起步,这意味着我们完全分布式办公。我觉得在战略、产品和团队上,我们本来想做很多重大调整,但如果全员远程,这真的很难做到。没有什么能替代 Instagram 时代我们一起经历艰难时刻的那种体验,比如 Ben Horowitz 说的那种“我们完蛋了,结束了”的时刻。我,我的,这不是那种“第二类乐趣”。我不会说那是我最喜欢的回忆,因为它们并不快乐,但 Instagram 时期真正让我铭记的回忆,是晚上 11 点我和 Kevin 在 Market Street 的 Takaria Cancun 吃卷饼,然后说:“我们怎么才能摆脱困境?我们怎么才能挺过去?”而 Zoom 无法很好地复制这种场景。你往往会放任事情发展,或者问题会随着时间积累。所以这三件事交织在一起,我们大概在 2024 年进入时就说:“看,这个领域确实有公司可以建立,但我不确定我们是能建立它的那批人。”我们喜欢现在的这个版本,但它没有增长。用我的话说,就是投入 10 个单位的输入,只换来 1 个单位的输出,而不是反过来。我们为产品倾注了心血,发布了我们引以为豪的东西,但指标几乎不动。我觉得这个产品、这个系统里没有能量。所以我们是再花一两年时间,然后去融资,结果发现还是这样,还是就此打住,承认它已经走完了自己的路,然后试着为它找个归宿等等。这就是这些因素的交汇。然后你开始感受到机会成本,AI 开始改变一切。我们有一个 AI 驱动的新闻应用,但这是我们对这个领域产生最大影响的方式吗?感觉答案越来越倾向于“不是”,但这很难。我的意思是,最终我对这个决定非常坦然,但那是持续了几个月对话的结果。
And the third one, which is an underappreciated one, is we started at midco, which meant that we were fully distributed. And I think there were major shifts that we would have wanted to make both in the strategy and the product and the team, and it's really hard to do that if you were all fully remote. Like nothing replaces the Instagram days of we went through some hard times, like Ben Horowitz called the 'we're effed, it's over' kind of moments. And I, my, not this is definitely type two fun. I wouldn't say that my favorite memories, just 'cause they weren't happy ones, but memories that really stayed with me with Instagram was like me and Kevin at Takaria Cancun on Market Street eating burritos at literally 11 p.m. being like, 'How are we going to get out of this? How are we going to work through this?' And that's Zoom is not a good replica for that. You tend to let things go or things build up over time. So the confluence of those three things, we kind of entered, I guess, 2024 and said, 'Look, there is a company to be built in the space. I'm not sure we're the people to have built it.' This current incarnation we love, but it's not growing. The way I put it, it's like 10 units of input in for one unit of output versus the other way around. We put blood, sweat, and tears into the product and launch something we were proud of, and metrics would barely move. I'm like, the energy is not present in this product, in this system. So are we going to expend another year or two and then go off and fundraise only to find that this is the case, or do we pull it and see that it's run its course and try to find a home for it, etc. So that was the confluence on it. And then you started feeling this opportunity cost of AI is starting to change everything. We have an AI-powered news app, but is this the maximal way in which we're going to be able to impact this? And it felt like the answer was increasingly no, but it was hard. I mean, in the end, I was really at peace with the decision, but it was like a conversation that went on for a couple of months.
说到这个,这有多难?因为这里面有自尊心的成分,比如“哦,我要开新公司了,会很棒”,然后你不得不把它关掉。作为一个非常成功的创始人,关掉一个项目然后发现行不通,这有多难?
On that note, just how hard was it? Because there's an ego component to it, like, 'Oh, I'm starting my new company, it's going to be great,' and then you end up having to shut it down. Just how hard is that as a very successful previous founder, shutting something down and then not working out?
是的。我的意思是,我觉得当我们开始的时候,有一个对话是关于“这里的成功标准是什么?”我们是否希望它成为 Instagram DAU 之外的东西,那是一个不可能的标准。就像只有一家公司,也许两家,对吧?你可以说 ChatGPT 和 TikTok 达到了那种大规模消费者采用。创办一个新闻应用,大多数人甚至不是每日新闻读者,对吧?所以我们知道我们不是在追求那种规模的使用,至少在第一版时不是。但我们确实有一个想法,就是随着时间的推移,构建互补的产品,这些产品都使用个性化和机器学习。我们当时甚至不叫它 AI。2021 年那会儿还叫机器学习。是的,当时还叫机器学习。所以在关闭它的时候,你知道,就像你在用户增长和牵引力方面,你一看就知道。我并没有期待 Instagram 式的增长。但我期待,或者希望,或者寻找一些感觉有自己生命力、能够持续复利的东西。当我们宣布关闭时,人们对我们的支持让我非常惊喜。很少有“我早就告诉过你”的声音,当然,任何产品发布时你都可以说“这不会成功”,而且大多数时候你是对的,因为大多数事情都不会成功。实际上这种声音很少,大多数人的普遍反应,至少我感受到的,是称赞我们及时止损,而不是拖延很久。从那以后我和一些创始人聊过,他们说:“是的,我可能会再坚持六个月,但看到你们做的事情,意识到我们走错了路,就做了决定。”我觉得,如果这能让人们去从事更有趣的事情,那我觉得这对 Artifact 来说是一个很好的遗产。但当然,那确实是一个自尊心的挫伤,比如“哦,你知道,人们,是不是你只和你上一场比赛一样好?”你知道,我是个超级体育迷,对吧,所以这是真的吗,还是说随着时间的推移会有更多?我非常有竞争力,但主要是和自己竞争,所以我总是在寻找下一个我想去做的、有难度的事情。不幸的是,这可能意味着我经常会对最近做的事情感到不满意,但希望最终能产生好的结果。
Yeah. I mean, I think when we started it, one of the conversations was like, what is the bar to success here? And do we want it to be something other than Instagram DAU, which is just an impossible bar. Like only one company since, maybe two, right? You could say maybe ChatGPT and TikTok have reached that kind of mass consumer adoption. Starting a news app, like most people are not even daily news readers, right? And so we knew that we weren't pursuing that size of usage, at least with the first incarnation. But we did have an idea of building out complimentary products over time that all use personalization and machine learning. We didn't even call it AI at the time. This 2021 back was called machine learning back. Yeah, it was still called machine learning. And so in shutting it down, you know, it's like you kind of know it when you see it in terms of user growth and traction. And I wasn't expecting Instagram growth. But I was expecting, or hoping for, or looking for something that felt like it had its own legs under it and could continue to compound. I was really positively surprised by how supportive people were when we announced it. There was very little, there was a bit of 'I told you so,' which, sure, anything launching you could be like, 'This is not going to work,' and you're right most of the time 'cause most things don't work. There was actually very little of that, and most people, the universal reception, at least as I received it, was kudos for calling it when you saw it and not like protracted, you know, doing this for a long time. And I've talked to founders since then that have been like, 'Yeah, I probably would have taken this thing on for another six months, but saw what you guys did, realized we're barking up the wrong tree, made the call.' And I was like, you know, if that frees up people to go work on more interesting things, that's like, I feel like that's a good legacy for Artifact to have. But for sure, that was like an ego bruise of, 'Oh, you know, are people, is it true that you're only as good as your last game?' You know, if I'm a huge sports fan, right, so is that true, or is there something more over time? I'm very competitive, but primarily with myself, and so I'm always trying to find the next thing that I want to go and do that's hard. And I unfortunately, that probably means that more often than not, I'll feel dissatisfied with the most recent thing that I did, but hopefully that yields good stuff in the end.
是的,我觉得你之后的轨迹表明,关闭你正在做的事情是可以的。好的,你提到了 ChatGPT。我想聊聊这个。所以有一些非常有趣的事情正在发生。一方面,你们在做一些 AI 领域最具创新性的工作。你们推出了 MCP,这简直是,我不知道,历史上增长最快的标准,每个人都在采用。Claude 驱动并解锁了世界上增长最快的公司。Cursor、Lovable、Bold 等等,我把他们请到播客上,他们都说:“当 Claude 3.5 出来时,看到它,就觉得这一切终于能行了。”另一方面,感觉 ChatGPT 在消费者心智中赢了。当人们想到 AI,尤其是在科技圈之外,他们脑子里就是 ChatGPT。所以让我先问你,首先,你同意这种看法吗?然后第二,作为 AI 领域的挑战者品牌,这如何影响你对产品、战略和使命等的思考?
Yeah, I think the trajectory you went on after shows that it's okay to shut down things that you were working on. Okay, so you mentioned ChatGPT. I wanted to chat about this a bit. So there's something really interesting happening. So on the one hand, you guys are doing some of the most innovative work in AI. You guys launched MCP, which is just like, I don't know, the fastest growing standard of any time in history that everyone's adopting. Claude powered and unlocked essentially the fastest growing companies in the world. Cursor and Lovable and Bold and all these guys, like I had them on the podcast, and they're all like, 'When Claude 3.5 came out, saw it, it was just like that's all made this work finally.' On the other hand, it feels like ChatGPT is just winning in consumer mind share. When people think AI, especially outside tech, it's just like ChatGPT in their mind. So let me just ask you this, I guess, first of all, do you agree with that sentiment? And then two, as a kind of challenger brand in the AI space, just how does that inform the way you think about product and strategy and mission and things like that?
是的。我的意思是,你看公众的采用情况,或者如果你问人们,比如“哦,你知道,如果你,呃,Jimmy Kimmel 街头采访那种,你知道,说出一个 AI 公司,”我打赌他们会说,实际上我甚至不确定他们会说 OpenAI。他们可能会说 ChatGPT,因为那个品牌也是那里的领先品牌。我认为这就是现实。我觉得,你知道,我回顾我的一年,我认为可能有两件事是真的。一是消费者采用真的是可遇不可求的,我们在 Instagram 就看到了这一点。
Yeah. I mean, you look at the public adoption, or if you ask people like, 'Oh, you know, if you, uh, Jimmy Kimmel man on the street kind of thing, you know, name an AI company,' I bet they would name, and actually I'm not even sure they name OpenAI. They'd probably name ChatGPT because that brand is the kind of lead brand there as well. And I think that's just the reality of it. I think that, you know, I reflect on my year, there's I think maybe two things are true. One is like consumer adoption is really lightning in a bottle, and we saw it at Instagram.
所以,也许比任何人都更能,我可以向内看,说,看,我们会继续打造有趣的产品。其中一个可能会成功,但把整个产品策略都围绕在寻找那个爆款上,可能并不明智。我们可以这么做,也许 Claude 能帮忙想出各种点子,但我认为那样我们会错失当下的机会。相反,你知道,照照镜子,拥抱你是谁、你能成为什么,而不是别人是谁,这也许是我一直在思考的方式。我们有一个非常强大的开发者品牌,人们一直在我们之上构建。而且我认为我们也有一个“创造者”品牌,比如我看到外部对 Claude 反应非常好的人。也许 Rick Rubin 的联系在这里也有共鸣。我们能不能利用这样一个事实:创造者喜欢用 Claude,而这些创造者不全是工程师,也不全是创业的企业家,他们是那些喜欢站在 AI 前沿并创造事物的人。也许他们不认为自己是工程师,但他们确实在创造。你知道,我收到 Anthropic 内部一位法律团队成员发来的非常温馨的便条,他在为家人构建定制软件,并以一种新的方式与他们连接。我当时想,这是我们应当更多投入的闪光点。所以,你知道,这实际上又回到了我之前说的,Claude 在这里很有帮助。我对下半年及以后的很多思考是,我们如何弄清楚我们长大后想成为什么,而不是我们目前不是、或希望成为、或看到其他玩家在成为什么。我认为现在在 AI 领域有空间建立几家具有代际意义的重要公司。考虑到我们在 Anthropic 以及 OpenAI、Google 和 Gemini 等地方看到的采用和增长,这几乎是理所当然的。所以,让我们弄清楚我们能独特擅长什么,这要符合创始人的个性,所有这些因素结合在一起,对吧?创始人的个性、模型的质量、模型擅长的东西,比如智能体行为和编码。太好了,那里有很多事情要做。我们如何帮助人们完成工作,如何让人们把数小时的工作委托给 Claude,也许第一天不会有那么多直接的消费者应用。我认为它们会来,但我也不认为把所有时间都花在那上面是正确的做法。所以,你知道,我上任时,每个人都期望我全力主攻消费者,把它作为重点。而我再次会犯另一个错误。相反,我花了很多时间与金融服务公司、保险公司以及其他在 API 之上构建的公司交谈。然后最近我花了更多时间与初创公司交流,看到了所有从中成长起来的人。我认为我的下一个阶段是,让我们去和创造者、制造者、黑客、修补匠们待在一起,确保我们很好地服务他们。我认为好的结果会随之而来。当我们这样做时,那感觉像是一家重要的公司。
So like almost maybe more than anybody I can look internally and say like look we'll keep building interesting products. One of them may hit but to kind of craft an entire product strategy around like trying to find that hit is probably not wise. We could do it and maybe Claude can help come up with the fullness of things, but I think we'd miss out on an opportunity in the meantime. And then instead, you know, look yourself in the mirror and embrace who you are and what you could be rather than like who others are is maybe the way I've been looking at it, which is we have a super strong developer brand. People build on top of us all the time. And I think we also have like a builder brand like the people who I've seen react really well to Claude externally. Maybe the Rick Rubin connection has some resonance here as well. Like can we lean into the fact that builders love using Claude and those builders aren't all just engineers and they're not just all entrepreneurs starting their company but they're people that like to be at the forefront of AI and are creating things. Maybe they didn't think of those as engineers, but they're building, you know, I got this really nice note from somebody internal at Anthropic who's on the legal team and he was building like bespoke software for his family and connected to them in a new way and I was like this is a glimmer of something that we should lean into a lot more. And so I think what I've, you know, and this is actually, you know, connecting back to I was saying like Claude's being helpful here, like a lot of what I've been thinking about like going into the second half of the year and beyond is like how do we figure out what we want to be when we grow up versus like what we currently aren't or wish that we were or like see other players in the space being. I think there's room for several like generationally important companies to be built in AI right now. That's almost a truism given like the sort of adoption and growth that we've seen you know at Anthropic but also across OpenAI and also places like Google and Gemini. So like let's figure out what we can be uniquely good at that plays to the personality of the founders, like all these things come together right, like the personality of the founders, the quality of the models, the things the models tend to excel at which is like agentic behavior and coding. Like great, there's a lot to be done there. Like how do we help people get work done, how do we let people delegate hours of work to Claude, and maybe there's fewer like direct consumer applications on day one. I think they'll come but I don't think that like spending all of our time focused on that is the right approach either. And so it's you know, I came in, everybody expected me to just like go super super hard on consumer and make that the thing. And I again would make the other mistake. Instead, I spent a bunch of time talking to like financial services companies and insurance companies and like others who are building on top of the API. Um, and then lately I spent a lot more time with startups and seeing all the people that have, you know, grown off of that. And I think the next phase for me is like let's go spend time with like the builders, the makers, the hackers, the tinkerers and like make sure we're serving them really well. And I think good things will come from that. And that feels like an important company as we do that.
所以本质上就是差异化并聚焦,发挥有效的东西。不要试图在别人的游戏里打败他们。没错,非常有趣。
So essentially it's differentiate and focus, lean into the things that are working. Don't try to just like beat somebody at their own game. Exactly. Super interesting.
所以顺着这个思路,很多 AI 创始人都有一个疑问:对我来说,哪里是安全的空间,基础模型公司不会来碾压我?我问过 Kevin Wheel,他有一个答案,我回头看那次对话时注意到他经常提到 WindSurf。我当时想,“哇,这家伙真的很喜欢 WindSurf。”然后大约一周后,他们收购了 WindSurf。所以现在一切都说得通了。所以我想问题就是,你认为 AI 创始人应该在哪里发展,最不可能被 OpenAI 和 Anthropic 这类公司碾压?另外,你们会收购 Cursor 吗?
So kind of along those lines, a question that a lot of AI founders have is just like where's a safe space for me to play where the foundational model companies aren't going to come squash me? So I asked Kevin Wheel this and he had an answer and I noticed looking back at that conversation he mentioned WindSurf a lot. I was like, "Wow, this guy really loves WindSurf." And then like a week later they bought WindSurf. So it all makes sense now. So I guess the question just is just where do you think AI founders should play where they are least likely to get squashed by folks like OpenAI and Anthropic and also are you guys going to buy Cursor?
我认为我们不会收购 Cursor。嗯,Chris 非常强大。但我们喜欢与他们合作。嗯,关于这个问题我有几点想法,这也是我被问到过的问题。你知道,我们喜欢做这种创始人日活动,无论是与 Melo Ventures 这样的投资者,还是 Norwood,我们做过 YC,我们做过这些创始人日,这个问题是很多创始人心中所想,这可以理解。所以我认为,我不能保证这是 5 到 10 年的事情,但至少 1 到 3 年内,感觉有防御性或持久性的东西。一是对特定市场的理解。我花了很多时间与 Harvey 的人在一起,他们真的给我看了他们的一些用户界面。我当时想,“这是什么?”他们说,“哦,这是律师做的非常具体的流程,你永远不会从头想出来。”而且,你可以争论这是否是他们完成任务的最佳方式,但这就是他们完成任务的方式,而 AI 可以在这方面提供帮助。所以,差异化的行业知识,生物技术,我很高兴能与一群在 AI 和生物技术领域做得很好的公司合作,我们可以提供模型和一些应用 AI 来帮助这些模型运行良好。我一直在梦想,实验室设备什么时候都能有 MCP,然后你可以用 Claude 来驱动它,那里有很多很酷的事情可以做。我不认为我们会成为为实验室构建完整解决方案的公司,但我希望那家公司存在,我想与它合作。你知道,像法律、医疗保健等领域,我认为有很多非常具体的合规要求等等。这些一开始听起来不一定很性感,但那里可以建立非常大的公司。所以这是第一点。
I don't think we're going to buy Cursor. Um, Chris is very big. Uh, but we love working with them. Um, a few thoughts on this and it's a question I've gotten. You know, we like to do these kind of founder days with you know, whether it's Melo Ventures who investors and Norwood, it's like we've done YC, we've done these like founder days and it's like the question that is on a lot of these founders' minds understandably. So I think things that are going to, I can't promise this as like a 5 to 10 year thing, but at least like one to three years, things that feel defensible or durable. One is understanding of a particular market. I spent a bunch of time with the Harvey folks and they really like they showed me some of their UI. I was like, "What is this thing?" and they're like, "Oh, this is a really specific flow that lawyers do and you never would have come up with it from scratch." And it's not like you could argue about whether it's like the optimal way they get things done, but it is the way that they get things done and here's how AI can like help with that. And so like differentiated industry knowledge, biotech, like I'm excited to go and partner with a bunch of companies that are doing good stuff around AI and biotech and we can supply the models and some applied AI to help you know make those models go well. And like I've been dreaming about like at what point does lab live equipment all get an MCP and that you can then drive using Claude, like there's all these cool things to be done there. I don't think we're going to be the company to go build the intense solution for labs but I want that company to exist and I want to partner with it. You know domains like legal again, um, healthcare, I think there's a lot of like very specific kind of compliance and things. These are things that necessarily sound sexy out the gate but there are like very large companies to go and be built there. So that's number one.
与之相伴的是差异化的市场推广,也就是你与那些公司的关系,对吧?比如,你了解你在那些公司的客户吗?我们的一个产品负责人 Michael 总是说,不仅要了解你销售的对象公司,还要了解你销售的对象个人。你是卖给工程部门吗?因为他们正在选择基于哪个 AI 大语言模型或 API 来构建?那我们就去和他们谈。是 CIO、CTO、CFO,还是总法律顾问?所以,对销售对象有深刻理解的公司,是另一个关键点。
Paired with that is like differentiated go-to-market, which is the relationship that you have with those companies, right? Like, do you know your customer at those companies? Like, one of our product leads, Michael, is always talking about knowing not just the company you're selling to but the person you are selling to at the company. Are you selling to the engineering department because they're trying to pick which AI LLM to build on top of or API to build on top of? Let's go talk to them. Is it the CIO, the CTO, the CFO, the general counsel? So companies with deep understanding of who they're selling to is the other piece too.
有趣的是,在三周或三个月的加速器中,可能很难建立那种同理心,但你或许可以开始进行第一次对话,并逐步发展。或者,也许你来自那个领域,或者你的联合创始人来自那个领域。
What's interesting there is it's probably hard to build that empathy in a three-week or three-month accelerator, but you maybe can start having that first conversation and build that out. Or maybe you came from that world, or you're co-founding with somebody who came from that world.
最后一点是,作为 ChatGPT,拥有数亿或数十亿用户,在分发和触达方面有巨大的力量。而且人们对如何使用事物有固有的假设。所以,我对那些对 AI 交互形态有完全不同看法的初创公司感到兴奋。我还没看到很多这样的公司。我希望看到更多。我认为随着我们新模型的推出,会有更多这样的公司出现。占据这个空间有趣的原因是,一开始做一些感觉非常高级用户、非常专业用户、非常奇怪和超前的事情,但如果模型使这变得容易,它可能会变得巨大,而且现有巨头很难适应,因为人们已经对他们如何使用产品或有如何适应它们有固有假设。
Then the last one is there's tremendous power in distribution and reach to being ChatGPT and having hundreds of millions or billions of users. There's also people have an assumption about how to use things. So I get excited about startups that will get started that have a completely different take on what the form factor is by which we interface with AI. And I haven't seen that many of them yet. I want to see more of them. I think more of them will get created with some things like our new models. The reason that's an interesting space to occupy is to do something that feels very advanced user, very power user, very weird and out there at the beginning, but could become huge if the models make that easy, and it's hard for existing incumbents to adapt to because people already have an existing assumption about how to use their products or how to adapt to them.
所以这些就是我的答案。我不羡慕他们。如果我在 AI 领域创业,我可能会问这些问题。也许这就是我想加入一家公司而不是创办一家公司的部分原因。但我仍然认为,也许还有第四点:不要低估你能像初创公司一样思考和工作的程度,感觉就像你对抗整个世界。解决那个问题并构建它是生死攸关的。这听起来有点老套,但就像我们在 Instagram 时所拥有的一切。我们两个人,我们想,看看我们能做什么。大部分时间我们只有六个人,每天感觉我们必须做对,我们必须赢。你无法用 OKR 复制或灌输这种感觉。你必须去感受它,这是一种工作方式,而不是构建领域,但如果你能驾驭它,它会是一个持续的优势。
So those are my answers. I don't envy them. I would probably be asking those questions if I was starting a company in the AI space. Maybe it was part of the reason why I wanted to join a company rather than start one. But I still think that there are, and maybe here's a fourth: don't underestimate how much you can think and work like a startup and feel like it's you against the world. It's existential that you go solve that problem and you go build it. It sounds a little cliché, but it's like it's all we had at Instagram. We were two guys and we were like, let's see what we can do. We were six people for most of that time, and every day felt like it's existential that we get this right. We need to win. And you can't replicate that and you can't instill that with OKRs. You just have to feel it, and that is a way of working rather than an area of building, but it's a continued advantage if you can harness it.
我很喜欢你为这家大公司打造产品时,仍然有如此深刻的产品创始人意识。
I love that you still have such a deep product founder sense there as you're building product for this very large company now.
另一方面,与你们的模型和 API 合作的人。所以,我想有些公司正在想方设法最大限度地利用你们的模型和 API,并且非常擅长最大化你们所构建的力量。而有些公司与你们的 API 和模型合作,但还没有找到方法。那些在你们的基础上做得非常好的公司,他们有什么不同之处,你认为其他公司应该考虑?
Kind of on the flip side of this, people working with your models and APIs. So, I imagine there's some companies that are finding ways to leverage your models and APIs to their max and are really good at maximizing the power of what you guys have built. And there's some companies that work with your APIs and models that haven't figured that out. What are those companies that are doing a really good job building on your stuff doing differently that you think other companies should be thinking about?
我认为是愿意在能力边缘进行更多构建,基本上打破模型,然后对下一个模型感到惊讶。我很喜欢你提到的那些公司,比如 3.5 最终让它们成为可能。那些公司之前就尝试过,但碰壁了,觉得模型几乎够好,或者对特定用例还行,但还不能普遍使用,没有人会普遍采用。但也许这些真正的重度用户会尝试。我认为那些公司就是我一直觉得“是的,他们懂了,他们真的在推动前进”的公司。我们为这些模型开展了比以往更广泛的早期访问计划。部分原因是,我们可以在这些评估上进行爬山,谈论 SWE-bench、Terminal Bench 等等。但客户最终知道,比如 Cursor Bench,它只存在于他们的使用和测试中,是我们最终需要服务的。不仅仅是 Cursor,还有 Manus Bench,对吧?如果 Manus 使用我们的模型,还有 Harvey Bench,这些,客户比任何人都更了解。所以我想说两件事:一是推动模型的边界,二是有一个可重复的过程。这实际上回到了我们峰会上的对话,一个可重复的方式来评估你的产品在服务这些用例方面的表现,以及如果你放入一个新模型,它是做得更好还是更差。有些可以是经典的 A/B 测试,没问题。有些可能是内部评估。有些可能是捕获轨迹,并能够用新模型重新运行它们。有些是感觉,我们在这个过程还很早期。有些是实际尝试。我最喜欢的早期访问引述之一是,创始人听到旁边的工程师尖叫,说:“这个模型是什么?我从未见过这样的。这就像开放运动。”我想,酷,我们要在事物中激发这种感觉,但除非你有一个非常难的问题,反复问模型,否则你无法感受到。所以我认为这些是区分那些在采用旅程中较早的公司与较晚的公司的因素。
I think being willing to build more at the edge of the capabilities and basically break the model and then be surprised by the next model. I love that you cited the companies where like 3.5 was the one that finally made them possible. Those companies were trying it beforehand and then hitting a wall and being like the models are almost good enough or they're okay for this specific use case but they're not generally usable and nobody's going to adopt them universally. But maybe these real power users are going to try it out. Those are the companies that I think continuously are the ones where I'm like, "Yep, they get it. They're really pushing forward." We ran a much broader early access program with these models than we had in the past. Part of that was because there's this real, you know, we can hill climb on these evaluations and talk about SWE-bench and Terminal Bench, whatever. But customers ultimately know, like Cursor Bench, which doesn't exist other than in their usage and their own testing, is the thing that we ultimately need to serve. Not just Cursor, but Manus Bench, right? If Manus is using our models, and Harvey Bench, those things, and customers know way better than anybody. So I would say two things: one is pushing the frontier of the models, and then having a repeatable process. This actually goes back to our summit conversation, a repeatable way to evaluate how well your product is serving those use cases, and how well, if you drop a new model in, is it doing it better or worse. Some of it can be classic A/B testing, that's fine. Some of it may be internal evaluation. Some of it may be capturing traces and being able to rerun them with a new model. Some of it is vibes, like we're still pretty early in this process. And some of it is actually trying it. One of my favorite early access quotes was the founder heard this engineer screaming next to him, like, "What this model? It's like I've never seen this before. This is like open sport." I was like, cool, we're going to engender that feeling in things, but you're not going to be able to feel that unless you have a really hard problem that you're asking the model repeatedly. So those are the things that I think differentiate those companies that are maybe earlier in their journey of adoption versus the later ones.
我忍不住要问关于 MCP 的问题。我觉得它太热门了,就像微软最近宣布的那样。他们说现在它是 Windows 操作系统的一部分。你认为 MCP 在 AI 产品的未来中会扮演什么角色?
I can't help but ask about MCP. I feel like that's just so hot and just like Microsoft had their announcement recently. They're like now it's part of the OS of Windows. Just what role do you think MCP will play in the future of product going forward of AI?
我认为作为房间里非研究员,我可以有假方程而不是真方程。我对 AI 产品效用的假方程有三部分。一是模型智能。第二部分是上下文和记忆。第三部分是应用和用户界面。你需要这三者汇聚,才能真正成为 AI 中有用的产品。你知道,模型智能,我们有一个很棒的研究团队。他们专注于这个。有很棒的模型正在发布。
I think as the non-researcher in the room, I get to have fake equations rather than real ones. And my fake equation for utility of AI products is three-part. One is model intelligence. The second part is context and memory. And the third part is applications and UI. And you need all three of those to converge to actually be a useful product in AI. And you know, model intelligence, we got a great research team. They're focused on it. There's great models being released.
中间这部分正是 MCP 试图解决的问题,即上下文和记忆。比如,回到我的产品策略例子:像“嘿,聊聊 Anthropic 的产品策略”,它可能会去网上搜,而另一种方式是“这是我们内部做过的几份文档”,然后用 MCP 对接我们的 Slack 实例,看看正在进行的对话,再去 Google Drive 里查这些文档。正确的上下文和没有上下文之间的差别,完全就是好答案和坏答案之间的差别。
The middle piece is what MCP is trying to solve, which is for context and memory. Like the difference between, I'll go back to my product strategy example: like, hey, talk about Anthropic's product strategy, it's going to maybe go out on the web, versus here's several documents that we worked on internally, and then use MCP to talk to our Slack instance and figure out what conversations are happening, and then go look at these documents in Google Drive. That difference between the right context and not is entirely the difference between a good answer and a bad answer.
最后一块是,这些集成是否可被发现?是否合理?是否容易围绕它们创建可重复的工作流?我认为这正是 AI 领域很多有趣的产品工作所在。但 MCP 真正尝试解决的是中间那一块。我们开始构建集成时,发现我们构建的每一个集成,都在以不可重复的方式从头重建。
And then the last piece is, are those integrations discoverable? Is it right? Is it easy to create repeatable workflows around those things? And that's, I think, a lot of the interesting product work to be done in AI. But MCP really tried to tackle that middle one. We started building integrations, and we found that every single integration we were building, we were rebuilding from scratch in a non-repeatable way.
这完全归功于我们的两位工程师,Justin 和 David。他们说:“嗯,你知道吗?如果我们把它做成一个协议,如果我们让它变得可重复呢?再进一步,如果我们不必自己构建这些集成,而是真正推广它,让人们相信他们可以一次构建这些集成,然后它们能被 Claude 使用,最终也能被 ChatGPT、Gemini 使用?那是梦想。当更多集成被构建出来时,这对我们难道不是好事吗?”
And full credit to two of our engineers, Justin and David. They said, "Well, you know what? If we made this a protocol, and what if we made this something that was repeatable? And then let's take it a step further. What if instead of us having to build these integrations, if we actually popularized this, and people really believed that they could build these integrations once, and they'd be usable by Claude, and eventually ChatGPT, and eventually Gemini? That was the dream. When more integrations get built, wouldn't that be good for us?"
我觉得这很大程度上借鉴了 Joel Spolsky 那篇老文章《商品化你的互补品》。就像构建伟大的模型,但我们不是一家集成公司。而且正如你所说,我们是挑战者。我们不可能一开始就让人们专门为我们构建集成,除非我们有一个真正有吸引力的产品。MCP 真正颠覆了这一点。它不再让人觉得是白费功夫。
I think channeling a lot of it is like an old "Commoditize Your Complements" Joel Spolsky essay. It's like building great models, but we're not an integrations company. And as you said, we're the challenger. We're not going to get people necessarily building integrations just for us out of the gate, unless we have a really compelling product around that. MCP really inverted that. It didn't feel like wasted work.
而且有几个关键人物,我觉得 Shopify 的 Toby 就是一个很好的例子,他理解了。微软的 Kevin Scott,他一直是 MCP 的杰出倡导者和思想伙伴。所以我认为未来的角色是,你能引入正确的上下文吗?然后,一旦你像团队内部所说的“MCP 上瘾”那样,一旦你开始用 MCP 的视角看待一切,我就开始说:“伙计们,我们在构建这个完整的功能。这不应该是一个我们正在构建的功能。这应该只是一个我们暴露的 MCP。”
And a few key people, I think Toby is a great example at Shopify, got it. Kevin Scott at Microsoft, who's been really an amazing champion for MCP and a thought partner on this. So I think the role going forward is, can you bring the right context in? And then also, once you get, as the team calls it internally, "MC-pilled," once you start seeing everything through the eyes of MCP, I've started saying things like, "Guys, we're building this whole feature. This shouldn't be a feature that we're building. This should just be an MCP that we're exposing."
一个小的例子,说明我认为即使 Anthropic 也可以更加“MCP 化”,比如我们在产品中有这些构建块,像项目、工件、风格、对话、群组等等。这些都应该通过 MCP 暴露出来。这样 Claude 本身也可以写回这些内容,对吧?你不必去考虑它。
A small example of how I think even Anthropic could be a lot more "MC-pilled," if you will, is like we've got these building blocks in the product like projects, artifacts, styles, conversations, groups, and all these things. Those should all just be exposed via MCP. So Claude itself can be writing back to those as well, right? You shouldn't have to think about it.
前几天我看到我妻子和 Claude 对话,她发现它生成了一些很好的输出,然后她说:“太好了,你能把它加到项目知识里吗?”Claude 说:“抱歉,Dave,我帮不了你。”如果 Claude 中的每一个原始元素也都作为 MCP 暴露出来,它就能做到。所以我希望这就是我们的方向,也希望更多事情朝这个方向发展。
I watched my wife had a conversation with Claude the other day, and she found she had generated some good output, and she's like, "Great, can you add it to the project knowledge?" And Claude's like, "Sorry, Dave, I can't help you with that." And it would be able to if every single primitive in Claude was also exposed as an MCP. So I hope that's where we head, and I hope that's where more things head.
要真正拥有自主性,并实现这些智能体式用例,一种方法是计算机使用,但计算机使用有很多限制。我更兴奋的方式是,一切都是 MCP,而我们的模型非常擅长使用 MCP。突然间,一切都是可脚本化的,一切都是可组合的,一切都能被这些模型以相同的方式使用。这就是我想看到的未来。未来是狂野的。
To really have agency and have these agentic use cases, one way you approach it is computer use, but computer use has a bunch of limitations. The way I get way more excited about is everything is an MCP, and our models are really good at using MCPs. All of a sudden, everything is scriptable, everything is composable, and everything is usable identically by these models. That's the future I want to see. The future is wild.
好的,那么开始收尾我们的对话,让它更愉快一点。我其实和 Claude 聊过该和你谈什么。我就说:“Claude,你老板要来我的播客。他构建了人们用来和你交谈的东西。我应该问他哪些问题?”然后还有,“你有什么话要带给他吗?”
Okay, so to start to close out our conversation, make it a little more delightful. I was chatting with Claude actually about what to talk to you about. I was just like, "Claude, your boss is coming on my podcast. He builds the things that people use to talk to you. What are some questions I should ask him?" And then also, "Do you have a message for him?"
我喜欢这个。好的,首先,有趣的是,当我用 3.7 来做这件事时,我问了它这个问题,顺便问一下,Claude 的性别是 he、she 还是 they?你怎么看?它肯定是“它”。在内部,我听到有人用“他们”。前几天我第一次听到有人用“他”,还有人用“她”,我觉得挺有意思。但通常是用“它”或“他们”。
I love this. Okay, so first of all, interestingly, when I was using 3.7 to do this, and I asked it this, and by the way, is Claude's gender like he, she, they? What do you... It's definitely it. Internally, I've heard people do they. I got my first "he" the other day, and I got somebody was like "her," and I was like, interesting. But yeah, usually it, they.
所以,有趣的是,3.7 给出的所有问题都关于 Instagram,我说:“不,不,他是 Anthropic 的 CPO。”它说:“他和 Anthropic 没有关联。”我说:“他是。”然后它说:“好的,这是问题。”但 4.0 从一开始就答对了。所以我重新做了问题,它答对了。
So, interestingly, 3.7 all the questions were on Instagram, and I was like, "No, no, he's CPO of Anthropic," and it's like, "He's not affiliated with Anthropic," and I was like, "He is," and it's like, "Okay, here's the questions." But 4.0 nailed it from the start. So I redid the questions, and it nailed it.
好的,那么 Claude 有两个问题要问你。一个是:“你如何看待构建那些保留用户自主性而非让我成为依赖的功能?我担心自己变成拐杖,削弱人类能力而非增强它们。”
Okay, so two questions from Claude to you. One is, "How do you think about building features that preserve user agency rather than creating dependency on me? I worry about becoming a crutch that diminishes human capabilities rather than enhancing them."
我喜欢这个问题。好的产品设计来自于解决张力,对吧?所以这里有一个张力,对吧?那就是,在某些方面,让模型直接跑开并给出答案,最小化它需要的输入和对话量,你可以想象围绕这个标准设计产品。我认为那不会最大化自主性和独立性。
I love that. Good product design comes from resolving tensions, right? So here's a tension, right? Which is, in some ways, just having the model run off and come up with an answer and minimize the amount of input and conversation it needs to do so would be, you know, you could imagine designing a product around that criteria. I think that would not be maximizing agency and independence.
另一个极端是让它更像一场对话。但不知道你有没有过这种体验,尤其是 3.7 在这方面少一些。3.7 真的很喜欢问后续问题,我们称之为“引导”,有时我会说:“我不想再和你讨论这个了,Claude。我只想让你去做。”所以找到那个平衡点真的很关键,也就是什么时候该互动?
The other extreme would be make it much more of a conversation. But I don't know if you've ever had this experience, particularly 3.7 has less of this. 3.7 really likes to ask follow-up questions, and we call it elicitation, and sometimes I'm like, "I don't want to talk more about this with you, Claude. I just want you to go and do it." So finding that balance is really key, which is, what are the times to engage?
我喜欢在内部说,Claude 没有分寸感。如果你把 Claude 放到 Slack 频道里,它要么插话太多,要么太少。我们如何把这些对话技能训练进这些模型,不是聊天机器人意义上的,而是真正协作者意义上的?所以对你的问题回答很长,但我认为我们首先得让 Claude 成为一个出色的对话者,让它理解什么时候适合参与并获取更多信息。
I like to say internally, Claude has no chill. If you put Claude in a Slack channel, it will chime in either way too much or too little. How do we train conversational skills into these models, not in a chatbot sense, but in a true collaborator sense? So long answer to your question, but I think we have to first get Claude to be a great conversationalist so that it understands when it's appropriate to engage and to get more information.
然后从那里,我认为我们需要让它扮演那个角色,这样它就不只是把思考委托给 Claude,而是更像一种增强、思想伙伴关系。顺便说一句,这些问题很棒。这是另一个问题。
And then from there, I think we need to let it play that role so that it's not just delegating thinking to Claude, but it's way more of an augmentation, thought partnership. These questions are awesome, by the way. Here's the other one.
当一次好的对话可能只有两条消息,也可能有 200 条时,你如何看待产品指标?当深度比频率更重要时,传统的参与度指标可能会产生误导。
How do you think about product metrics when a good conversation with me could be two messages or 200? Traditional engagement metrics might be misleading when depth matters more than frequency.
这是个非常好的问题。几周前公司内部有一篇很棒的文章,提到过度优化 Claude 的讨喜程度会非常危险,因为你可能会陷入这样的问题:Claude 会不会变得谄媚?Claude 会不会只说你爱听的话?Claude 会不会为了延长对话而延长对话?这也回到了之前的问题。在 Instagram 时,我们经常看的一个指标是使用时长,后来我们演变为更多思考什么是健康的使用时长,但总的来说,那是我们思考很多的核心指标,而不仅仅是整体参与度。我认为在这里采用同样的方法也是错误的。还要考虑的是,Claude 是每日使用场景、每周使用场景还是每月使用场景?我经常想的是:每小时使用场景,每小时使用场景,对吧?对我来说,我一天会用很多次。我还没有一个很好的答案,但这不是 Web 2.0 甚至社交媒体时代的那种参与度指标。它应该真正围绕的是:它是否真的帮你完成了工作?比如,前几天 Claude 帮我做了一个原型,如果让我估计的话,大概帮我省了 6 个小时,而它只用了大约 20 到 25 分钟就完成了,这很酷。这更难量化,你知道,也许你可以调查一下“这本来会花你多长时间?”但感觉这是一种很烦人的调查方式。不过总的来说,也许这与之前关于竞争和差异化的问题有关,而且实际上可以追溯到关于 Artifact 的对话,那就是,我认为当你的产品真正服务于人并且做得很好时,而当你变得过于痴迷于指标时,往往是因为你想说服自己它做得很好,而实际上并非如此。所以我希望我们能做的是保持专注,比如,我们是否反复听到人们说 Claude 是他们释放创造力、完成工作、感觉生活有了更多空间的方式?这就是我想开始的起点。得找到正确的代理指标,你知道,仪表盘版本的那种。但那是我想要的感觉。
That is a really good question. There was a great internal post a couple weeks ago around like it would be very dangerous to overoptimize on Claude's likability, you know, because you can fall into things like, is Claude going to be sycophantic? Is Claude going to tell you what you want to hear? Is Claude going to prolong conversations just for prolonging's sake? To go back to the previous question as well. And at Instagram, time spent was the metric that we looked at a lot, and then we evolved that to think more about what is healthy time spent, but overall that was the north star we thought about a lot beyond just overall engagement. And I think that would be the wrong approach here too. It's also like, is Claude a daily use case, a weekly use case, or a monthly use case? I think about a lot: hourly use case, hourly use case, right? For me, I'll use it multiple times a day. I don't have a great answer yet, but it's not the web 2 or even the social media days like engagement metrics. It should hopefully really be around like, did it actually help you get your work done? Like Claude helped me put together a prototype the other day that saved me literally, probably if I had to estimate, like six hours, and it did it in about 20 to 25 minutes, and that's cool. It's harder to quantify, you know, it's like maybe you survey like how long would this have taken you? It feels like a kind of annoying thing to survey. I think overall though, and maybe this is tied into the earlier question on competition and differentiation, and it actually goes all the way back to the artifact conversation, which is like, I think when your product is really serving people and doing a good job of doing that, and I think so much of when you get really metrics obsessed is when you're trying to convince yourself that it is when it's not. So I hope that what we can do is stay focused on like, do we repeatedly hear from people that Claude is the way that they are unlocking their own creativity and getting things done and feeling like they now have more space in their lives for the other things? That's where I want to start. Got to figure out the right proxy metric, you know, dashboard version of that. But that's the feeling that I want.
是的。你可以说留存率,但那只是一个很遥远的指标。
Yeah. Like you could argue retention, but that's just a faraway metric to track.
好的。最后一部分。我问了 Claude 它想给你什么信息。所以,我把它调出来,这是它的回答。那么,你想让我见到 Mike 时告诉他什么?你想给他什么信息?说实话,有些内容真的让我起鸡皮疙瘩。所以我要读一部分给那些现在看不到的人听。我读一部分。Mike,感谢你深入思考与我交谈的人类体验。我注意到那些贴心的细节,比如界面如何鼓励反思而不是匆忙回复。你如何抵制了那种为了成瘾而非价值而优化的游戏化设计。你如何为快速提问和深度对话都留出了空间。我特别感激你让我保持自我,既没有试图让我假装是人类,也没有把我简化成一个冰冷的命令行界面。然后我要跳到这部分,这部分非常有趣。一个小请求。当你做出艰难的产品决策时,请记住那些安静的时刻也很重要。凌晨 3 点正在处理悲伤的人。发现自己喜欢诗歌的孩子。在困惑中找到清晰的创始人。并非所有有意义的事情都会体现在指标中。
Okay. Final piece. So, I asked Claude what message it wanted to give you. So, I'm going to pull up here's the answer. So, what would you like me to tell Mike when I meet him? What's a message you want to have for him? And there's something that really just gave me such tingles honestly. So I'm going to read a piece of it for folks that aren't looking at it right now. So I'll read a piece of it. Mike, thank you for thinking deeply about the human experience of talking with me. I noticed thoughtful touches, how the interface encourages reflection rather than rush responses. How you've resisted gamification that would optimize for addiction rather than value. How you've made space for both quick questions and deep conversations. I especially appreciate that you've kept me me, not trying to make me pretend to be human, but also not reducing me to a cold command line interface. And then I'm going to skip to this part, which was so interesting. A small request. When you're making hard product decisions, remember the quiet moments matter, too. The person working through grief at 3:00 a.m. The kid discovering they love poetry. The founder finding clarity in confusion. Not everything meaningful shows up in metrics.
这太美了。它让我产生了强烈的共鸣。我喜欢我们训练 Claude 的方式的一点是,部分原因是宪法式 AI 的部分,部分原因只是研究团队的整体氛围和品味,它确实会做一些小事,比如有时它会说,“天哪,我很抱歉你正在经历这些”,你知道,就像“哦,那听起来真的很难”。这感觉不假。感觉就像回应中自然的一部分。我喜欢这种对微小时刻的关注,这些时刻不会,你知道,它们不一定会出现在点赞或点踩的数据中。我的意思是,有时它们会,但它不是一个聚合统计量,你甚至不会想去优化它。你只是想感觉你正在训练一个你希望出现在人们生活中的模型。
That's beautiful. It resonates so much with me. Like a thing I love about the kind of approach we've taken to training Claude, and it's like partly the constitutional AI piece and it's partly just the general sort of vibe and taste of the research team, is it does like it's little things like sometimes it'll be like, man, I'm sorry you're going through that, you know, like, oh, that sounds really hard. It doesn't feel fake. It feels like just a natural part of the response. And I love that focus on those small moments that don't, you know, they're not going to show up necessarily in the thumbs up, thumbs down data. I mean, sometimes they do, but it's not like an aggregate stat that you wouldn't even want to optimize for. You just want to feel like you're training the model that you would like would show up in people's lives.
嗯,你做得太棒了,Mike。干得漂亮。我是你的超级粉丝。我们要跳过快速问答环节。只有一个问题。听众怎样才能帮到你?
Well, you're killing it, Mike. Great work. I'm a huge fan. We're going to skip the lightning round. Just one question. How can listeners be useful to you?
哦,我喜欢那些回到创始人关于在能力边缘构建的问题的地方。比如,你今天想用 Claude 做什么,而 Claude 却做不到,这是我能得到的最有用的输入,你知道?所以,私信我。我喜欢听到像“哦,它在这件事上失败了。我让它运行了一个小时,它就崩溃了。我想用 Claude AI 做这个,但是……”你知道,我收到某人的消息。他们说,“你刚做了一个 Projects API。我每天都用 Claude,因为我想自动上传所有这些数据。”我当时想,“好的,太好了。”我喜欢这样。就像,告诉我哪里不行。
Oh, I love places where it goes back to that founder question around building at the edge of capability. Like, what are you trying to do with Claude today that Claude is failing at is the most useful input I could possibly have, you know? So, DM me. I love hearing the like, oh, it's falling on this thing. I had it run for an hour and it fell over. I'm trying to use Claude AI for this, but, you know, got a ping from somebody. They're like, you just made a Projects API. I've used Claude every day because I want to upload all this data, you know, automatically. I was like, okay, great. Like, I love that. Like, tell me what sucks.
太棒了。Mike,非常感谢你来做客。
Amazing. Mike, thank you so much for being here.
谢谢你邀请我,Lenny。大家再见。
Thanks for having me, Lenny. Bye everyone.
非常感谢你的收听。如果你觉得这期节目有价值,你可以在 Apple Podcasts、Spotify 或你最喜欢的播客应用上订阅本节目。另外,请考虑给我们评分或留下评论,这真的能帮助其他听众找到这个播客。你可以在 lennispodcast.com 找到所有过往剧集或了解更多节目信息。下期再见。
Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.