Anthropic Labs:AI 产品前沿探秘

Anthropic Labs: Inside the Frontier of AI Products

迈克·克里格 Mike Krieger · Alex Kantrowitz · 2026-06-25 · 约 41 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Mike Krieger 探讨 Anthropic Labs、Fable 模型争议以及 AI 发展的快速节奏。

Mike Krieger discusses Anthropic Labs, the Fable model controversy, and the rapid pace of AI development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 30)

全文 · Full transcript(中英对照)

引言 Introduction

Host

Anthropic 不仅是一家大规模模型构建公司,也是一家大规模产品构建公司,像 Claude Code 和 Co-work 这样的产品在过去几个月里发展得如火如荼。

Anthropic is not just a massive model builder, it's a massive product builder as well with products like Claude Code and Co-work that have taken off like crazy over the past few months.

Host

Claude Code 出自一个我们很多人都不太了解的地方,叫做 Anthropic Labs。Anthropic Labs 是 Anthropic 内部的一个组织,致力于以 AI 为核心构建下一代前沿产品。今天我们有幸请到了负责该实验室的人,Mike Krieger,他是 Instagram 的联合创始人,现在是 Anthropic 旗下 Anthropic Labs 的负责人。我们将欢迎他上台,同时还有《连线》杂志的 Lauren Goode 与我共同主持采访。Mike 和 Lauren,让我们听听你们两位的见解。

And Claude Code came out of a place that few of us know much about called Anthropic Labs. And Anthropic Labs is an organization within Anthropic that is working on building the next level of frontier products with AI at the center. And so today we are lucky to hear from the person running that lab, Mike Krieger, who is the co-founder of Instagram and now the lead of Anthropic Labs at Anthropic. We're going to welcome him on stage along with Lauren Goode of Wired who will join me as a co-interview. Mike and Lauren, let's hear from both of you guys.

开场闲聊 Opening Banter

Host

嘿,让我看看。你在这边。

Hey, let me see. You're over here.

Host

好的。好的,Mike。Anthropic 这边还挺平静的?

All right. All right, Mike. So chill times in Anthropic land?

Mike

没什么大事。

Nothing going on.

Host

这周比较平淡。Lauren,你想先开始吗?

Slow week. You want to start, Lauren?

角色与白宫局势 Role and White House Situation

Host

是的。首先,我们想谈谈你在 Labs 的工作,并向大家解释你的角色。但我想先问你,你现在与白宫那边的情况有多接近?

Yeah. Uh, first of all, I mean, we want to talk about what you're working on at Labs and explain your role to folks. But I want to ask you first, how close are you right now to the situation with the White House?

Mike

比我担任 CPO 时要少一些。大约 5 个月前我转到了这个 Labs 的职位。我觉得如果我还是 CPO,我会深陷其中。现在,显然,我们想要恢复访问,作为产品人员,我希望确保能恢复访问,但不像我之前担任高管职位时那样接近。

Less than in my CPO role. So I transitioned about like 5 months ago into this Labs role. I think in the CPO role I think I would have been deep in it. Now I'm more, you know, obviously, you know, we want to restore it and as a like product person I want to make sure that gets access to, but less close to it than, you know, in that sort of C-level role that I had before.

寓言引发的反弹 Fable Backlash

Host

好的。Alex,你有追问吗?我大概有八个追问。

Okay. Alex, do you have a follow-up? I have like eight follow-ups.

Host

当然。嗯,X 上有个叫 Ben 的人说:“如果 Anthropic 能为我重新启用 Fable,我会把我的长格式出生证明原件寄给他们。”我现在听起来就像那些痴迷于 4o 的疯子。你会接受 Ben 的长格式出生证明吗?

Definitely. Well, we There was a guy named Ben on X who said, "I will mail Anthropic an original copy of my long-form birth certificate if they will enable Fable for me again." I sound like those lunatics who were obsessed with 4o now. Will you take Ben's long-form?

Mike

我不知道我们会不会接受他的长格式出生证明,但这确实很有趣。我的意思是,Fable 只上线了几天,但自从那以后,我每次发推,他们都不看我发了什么,大多是在说“把 Fable 带回来”。就像在 Instagram 上你下架了 Gotham 滤镜,你还记得 Gotham 吗?那之后接下来的 8 年里,我听到的都是“把 Gotham 带回来”。所以,这触动了大家的神经。不过,Fable 会回来的,而且会比 Gotham 回来得更快。但,是的,显然那些用过它并开始将其融入工作的人,这真的很有趣。我学会了不完全相信模型发布当天甚至当周的反应。你得真正测试过才知道。所以,在新模型发布的最初几天,我几乎完全屏蔽噪音,因为每个人可能都有自己喜欢的新模型玩具示例,但除非你真的用它完成实际工作,否则很难真正测试它。我觉得人们刚开始这样做,然后我们就不得不撤回 Fable。但,我记得 12 月我们发布 Opus 4.6 的时候,那是一个有趣的时期,大家都回家过节,很多人在圣诞和新年之间那一周休假。然后他们回来后说:“哦,我花了很多时间,我真正明白了为什么 Opus 好,我要用它。”所以,我认为 Fable 还没有得到那样的机会。

I don't know that we'll take his long-form, but it has been interesting. I mean, Fable was only available for a few days, but I definitely, every time I've tweeted since then, they've not read whatever I was tweeting and they've mostly been like, "Bring back Fable." Which like in Instagram where you got rid of Gotham, do you remember Gotham, the filter? This is like and then for the rest of like the next 8 years all I heard was bring back Gotham. So, it struck a nerve. But, Fable will come back before Gotham did. But, yeah, it's clearly the folks that had gotten to use it and started incorporating it. It's actually really interesting. I've learned to not really trust day of or even week of model reactions. You don't really know until you've put it through its paces. And so, like I almost just completely block out the noise in the first couple of days of any new model release cuz I don't know, everybody has maybe their like toy example thing that they like to do with the new model, but it's hard to actually put it through its paces until you've actually had real work done with it. And I think people were just starting to do that and then, you know, we had to sort of pull back Fable. But, I remember in December when we put out Opus 46. It was like this interesting time where everybody went home for the holidays and a lot of people had that week off between Christmas and New Year's. And then they came back and were like, "Oh, I spent a lot of time and I really get why Opus is good and I'm going to do it." So, I don't think Fable has had that opportunity yet.

特朗普政府的反应 Reaction from Trump Administration

Host

但尽管如此,特朗普政府对此反应相当大,我想在座的各位,尤其是听过 Alex Stamos 那场演讲的人,都明白发生了什么。但这发生在模型发布后的几天内。如果我可以给《连线》一些功劳,《连线》昨晚刚刚报道说,这是因为 Anthropic 与 SK 电信的关系,或者说给了 SK 电信模型访问权限,这可能引起了政府内部的警觉。你对这种反弹来得如此之快感到有多惊讶?

But, despite that though, I mean this was a pretty big reaction from the Trump administration and I think everyone here, especially if you were listening to the Alex Stamos session, understands what's going on. But, this happened within a few days of the model release. If I can give Wired some credit, Wired just reported last night that it was due to Anthropic's relationship with or having given access to the model to SK Telecom that could have raised flags within the administration. How surprised were you by how immediate that backlash was?

Mike

是的,我认为这个反应决定令人惊讶,我们当时也立即与他们接触,试图恢复访问。同时,我们内部经常思考的一件事是,当我们还在 Facebook 时,墙上有一张海报,上面写着“每一天都像一周”。我认为这在 AI 领域正成为现实,我认为整个行业需要提醒自己的是,我们正面临前所未有的时代。我们面对的是新情况,而且它们可能发展得非常快。因此,我认为培养能力和联系,确保这些对话能够快速进行,真的非常重要。

Yeah, I think the sort of reaction decision was surprising and we were sort of immediately engaged with them to, you know, restore access as well. And so, at the same time, you know, with the one thing that we think a lot about internally is, you know, there used to be a poster on the Facebook wall when we were still there that was like, "Every day feels like a week." And I think that's becoming true in AI, and I think a good thing to remind ourselves of in general in the industry is we're dealing with unprecedented times. We're dealing with new situations, and they can develop really quickly as well. And so, I think also developing the capabilities and the connections to make sure those conversations can happen quickly is really, really important.

Host

我为你墙上的海报提个新标语:快速行动,越狱一切。

I have a new motto suggestion for you for the wall. Move fast and jailbreak things.

Mike

快速行动,越狱一切。我觉得他们不会用这个的,Lauren。

Move fast and jailbreak things. I don't think they're going to use that, Lauren.

为何Anthropic被单独针对 Why Anthropic Was Singled Out

Host

好的。你知道,我们刚刚听 Alex 说,其中一些能力在之前的模型上就已经有了,或者你可以在今天可用的其他模型上找到漏洞。那么,你认为为什么 Anthropic 在这方面被单独挑出来?

Okay. You know, we just heard from Alex that some of these capabilities have been available on previous models or you can find bugs with other models that are available today. So, why do you think Anthropic got singled out on this front?

Mike

我不知道为什么 Fable 被单独挑出来,而且 Fable 也是非网络攻击意图的模型。我认为随着时间推移真正变化的是,如果你知道该提示什么、知道你在找什么,那是一种能力;还有另一种能力,我不确定我是否认同 Alex 那个“给高中运动员打兴奋剂”的比喻,但你知道,我们考虑模型安全时经常思考“提升”这个概念。比如,如果你看我们的模型卡,我们在生物领域评估风险的方法之一,是比较普通人使用模型与专家或普通人仅使用互联网所获得的提升,看看对比如何。所以,这是一个随着模型能力增强而不断发展的趋势。因此,可能不是为什么 Fable 被单独挑出来,而更多是整体趋势的问题。

I don't know why Fable specifically was singled out as, you know, again, Fable being the non-cyber intention model as well. I think the thing that does change over time is there's the capabilities if you knew what to prompt and knew what you're looking for, and then there's the capabilities I think, you know, I'm not sure I resonate with the like juicing high school athletes metaphor from Alex, but like, you know, the uplift that you get like uplift is a thing that we think about a lot when we think about model safety. So, if you look at our model cards, for example, one of the ways that we look at risk in the bio domain is comparing the uplift from sort of lay person using the model versus, you know, an expert or lay person just using the internet and seeing what the comparison there is. And so, that is one trend that has been progressing as the models get more capable. And so, it may be less why Fable is singled out and maybe more like what the overall trajectory is.

支出数据一览 Ramp Data on Spending

Host

有趣。我们刚刚在播客中邀请了 Ramp 的首席经济学家 Our Karaziyan。他说的一个有趣的事情是,在五角大楼事件期间,有一些头条新闻说某家公司不再使用 Anthropic 了。但实际上,Ramp 的数据显示,对 Anthropic 模型的花费实际上增加了。显然这对公司来说是一个很好的宣传时刻。

Interesting. So, we just had Our Karaziyan, the lead economist on Ramp on the podcast. And one of the interesting things that he said was, you know, we had during the Pentagon situation, there were these headlines, okay, this company will not use Anthropic anymore. But actually the data from Ramp shows that spending actually increased to Anthropic models. It was apparently a good publicity moment for the company.

营销与真实关切 Marketing vs. Real Concern

Host

所以,我觉得这正好切中了我们节目里一直在争论的一个话题:这里面有多少是真正关心这些问题,又有多少是来自 Anthropic 的营销。而我们今天正好请到了一位 Anthropic 的人,可以给我们讲讲。我们刚刚听了 Themos 从安全角度谈了一点,但今天很荣幸能请到你。那么,这到底是实质性的,还是营销,还是两者兼有?

So, I think it sort of plays into a debate that we have here on the show about how much of this is real concern for the issues and how much of it is, you know, from Anthropic's marketing. And we have somebody from Anthropic here who can actually shed some light on that. We just had Themos talking a little bit about it from his perspective on the security side, but we're lucky to have you here today. So, is it material or is it marketing or some combination?

Mike

我觉得,最难的事情之一就是真正深信某件事是真的而不是营销,而且这不只是针对 Anthropic。我认为人们普遍对任何公司的说法持怀疑态度是对的,你也应该用自己的标准去过滤。但就我个人而言,我总是觉得“不,这是真的”,我们既深切关心安全,而且我们在任何问题上发声,要么是为了描绘很可能即将到来的图景,要么是我们相信会到来的,要么是我们已经看到并发现的。比如在 Mythos 的案例中,我们确实是在做漏洞扫描和缺陷查找,并且与那些参与 Project Glasswing 初始公告的公司合作。所以,这项技术也确实在做一些令人难以置信的事情,因此即使我们把看到的事情说出来,我觉得也可能显得像是在炒作。我希望我能按下一个按钮,让所有人都相信我们没有在炒作。我意识到我们身处的现实并非如此,但至少从我个人的角度来看,我们尽量实话实说。

I mean, I think it was one of the hardest things to really deeply believe that something is true and not marketing, and that's not even just Anthropic. I think people are generally right to be skeptical of any company saying anything, and you should put it through your own filters as well. But for me personally, I'm always like, 'No, but it's real,' and we both deeply care about safety, and to the extent that we are being vocal about anything, it is to either help paint the picture of what very likely is coming, or we believe is coming, or what we've already seen and spotted. For example, in the Mythos case, really just looking at vulnerability scanning and bug finding, and doing it in partnership with companies that were in that kind of Project Glasswing initial announcement. And so, the technology is also doing really incredible things, and therefore even calling out what we see as what is happening, I think can seem hypie. I wish I could press a button and make everybody believe that we are not being hypie. I realize that's not the reality that we operate in, but at least from my perspective, we try to call it like it is.

禁令对开发的影响 Impact of the Ban on Development

Host

据我所知,在实验室内部,尤其是在 Anthropic 的研发部门,你们实际上在用 AI 模型来构建新产品、制作原型,并验证你们的论点。那么,这个禁令现在是否实质上限制了你们在实验室内部这样做的能力?

My understanding is that within labs, in particular in research and development at Anthropic, you're using the AI models to actually build new products, to prototype new products, and sort of test out your thesis. So, is this ban essentially now limiting your ability to do that within labs?

Mike

是的,我的意思是,Fable 绝对是我用过的最好的模型,这并不是说工作已经停止了,但它肯定不如我们之前用的其他模型。或者抱歉,我们之前用的模型也不如它。而且,这也许是对第一周不信任的逆向反应,你知道,当你没有 Fable 时会发生什么。显然,那些已经接触过这个模型并开始使用它的人在 Twitter 上的反应很强烈。但我想说,即使在我个人的使用中,我也会想,“哦,我现在用的是 Opus 4.8。”它很好。我仍然很有生产力。我在工作。但我们可以深入谈谈我的工作是如何随着这些类似 Fable 的模型而改变的,但肯定是很明显的。

Yeah, I mean, definitely Fable is the best model I've ever used, and it's not to say that work has stopped, but it's definitely less good than the other models that we were using. Or sorry, we were using models that are less good than that as well. And it is also, I mean, maybe the reverse of the distrusting the first week, you know, response is what happens when you don't have Fable. And obviously the Twitter reaction for the people that had already kind of gotten into the model and raising it was strong. But I'd say even in my personal use, I'm like, 'Oh, I'm on Opus 4.8.' And it's good. I'm still productive. I'm doing work. But we can go into sort of how my work changed with these Fable-like models, but it is noticeable for sure.

与先进模型的差异 Differences with Advanced Models

Host

是的,我想我们很想知道这一点。我的意思是,公众接触 Fable 的时间大概只有半分钟。像 Project Glasswing 这样的团队已经接触到了 Mythos。我们真的不知道使用你今天能用的模型和使用这些超级 Anthropic 模型之间有什么区别。那么,有了 Fable 或 Mythos,你到底能做出哪些不同的事情?

Yeah, I think we'd like to know that. I mean, you know, we the public had access to Fable for like a half a minute. Groups in this like Project Glasswing have had access to Mythos. We don't really know what the difference is between using a model that you can use today and using one of these, you know, super Anthropic models. So, what actually could you do differently with a Fable or a Mythos?

Mike

对我来说,关键是委派的范围和规模。而且这些东西其实都不完美。比如,人们会说,“哦,这现在是一个五级软件工程师还是六级?”但任何广泛使用过这些模型的人都知道,它们的能力仍然参差不齐,对吧?在某些方面,它们在很多方面是比我更好的工程师。但在其他方面,比如,我今天还在抱怨它漏掉了一个下伸部,就是字母 G 的底部那部分叫下伸部。我说,“你怎么把它放到 UI 里的?”结果被裁剪了。当然,还有视觉能力需要改进,调试能力,有时甚至只是人类常识,模型比我强得多。但总的来说,我认为对我来说最大的转变是,这很有趣,因为它恰好发生在我重新回到构建者角色的时候。所以我真的看到了从作为高管使用这些模型,你试图最大化利用它们,但不会让它写你所有的电子邮件,我认为战略仍然需要来自你,然后你可以用模型来测试它。但回到构建者角色,我从委派小块任务,比如“请修复这个 bug”或“我想实现这个功能,我们来来回回讨论”,变成了更像是“好吧,我收到了用户的 bug 报告,或者我有一个想构建的东西的想法,你能画出两三种实现方式吗?”好吧,这看起来可行。我经常发现,有时它会给我解释或提议,我会说,“好吧,这超出了我的理解。你显然比我聪明得多,请向我解释,不是把我当五岁小孩,但至少不是对你。”它有时会这样解释,但然后去构建,并且以非常非常高的成功率做对。我认为这开始真正改变你的运作方式。我更多地转向,在睡觉前,确保我已经为 Fable 排队了足够多的块状工作,我称之为整夜的工作,然后我稍后检查,它在一小时内完成了,就像客人接下来七个小时都在闲逛。但真的,委派的是目标,而不仅仅是任务。

I think for me, it's the sort of scope and scale of delegation. And all these things are really imperfect. Like, you know, people say like, 'Oh, is this now a level five software engineer or a level six?' But anybody who's used these models extensively knows that they're still spiky in capabilities, right? In some ways they're, you know, in many ways they're better engineers than me. And in other ways, like, you know, I was complaining today that it had missed a descender, like the G's bottom part is called the descender. I'm like, 'How did you put that in the UI?' And it's clipped. And of course, there's vision capabilities that need to improve, and there's debugging capabilities, and there's sometimes even just sort of human common sense that the models are way better at than I am. But overall, I think the big shift for me working, and it was really interesting because it sort of coincided with me going back into a builder role. So I really got to see going from using these models as an executive, where you're trying to get the most out of them but you're not going to have it write all your email, and I think strategy still needs to come from you, and then you can use the models to sort of test it. But going back into a builder role, I went from delegating chunks like 'please fix this bug' or 'I'm thinking of implementing this feature, let's go back and forth' to something that ends up being much more like, 'All right, I got this bug report from one of our users, or I have this notion of something that I want to build, can you sketch out two or three ways in which we could do it?' All right, that seems plausible. Often I find actually sometimes it'll give me the sort of explanation or proposal and I'll be like, 'Okay, that actually is over my head. You are clearly way smarter than me, explain it to me like I'm not five, but at least not you.' And it'll sometimes explain it that way, but then go build it and get it right at a very, very high rate. And I think that starts really changing how you operate. I moved much more to, before going to bed, making sure that I had queued up for Fable enough chunky work to last, I would call it the whole night, and I would check in later and it got it done in an hour, and it was like the guest hanging out for the next seven hours. But really, delegating much more of a goal than just a task.

示例任务 Example Task

Host

那么,举个例子,一个任务。

So one task, for example.

Mike

是的。

Yeah.

Host

给我们举一个你会交给它的任务的例子。

Like give us an example of one task you would hand to it.

Mike

我的意思是,这里有一个有点疯狂的例子,是给观众中的程序员的。我用 Python 写了我们实验室的一个项目。这是我很熟悉的语言;Instagram 全是 Python。由于一些不太令人兴奋的原因,我们实际上需要它用 TypeScript 来部署。我当时想,“好吧,那会像 Instagram 一样,我们多年来一直谈论从 Python 迁移到 PHP 或 Hack,即收购后的 Facebook 语言,但至少我在那里的时候,从未做过。”

I mean, here's a kind of crazy one, which is for the programmers in the audience. I'd written one of our labs projects in Python. It's like the language I know well; Instagram was all Python. And for some not super exciting reasons, we actually needed it to be in TypeScript to deploy it. And I was like, 'All right, that's going to be like an Instagram, we for years talked about moving from Python to PHP or Hack, the Facebook language after the acquisition, and then, at least when I was there, never did.'

动态工作流与任务执行 Dynamic Workflows and Task Execution

Mike

我们有一个叫动态工作流的功能,可以让它把任务分解成很多子任务。我信任它不只是做单个动作,而是处理整个语言转换,涉及几十万行代码,让它去规划、执行、验证工作,双重验证,然后我回来时工作已经完成了。所以那个级别的任务是一个很大的、块状的任务。

We have a feature called dynamic workflows where you can have it break down the task into a lot of subtasks. I trusted it to not just do the individual action, but handle a whole language conversion of hundreds of thousands of lines of code, go off and do it, plan it, execute it, verify the work, double verify the work, and then I came back to the work being complete. So that level of task is a big, chunky task.

Host

嗯。所以你基本上是说它更快了。一个小时就完成了?你是在和以前的情况做比较猜测。

Hm. So you're basically saying it was faster. Did it in an hour? You're sort of guessing compared to what it would have been before.

Mike

我觉得主要区别在于,过去它会说“太好了,我完成了。”然后你会说“真的吗?你在这里走了捷径,或者这不太对,我需要去验证,或者你偷工减料了。”

I think the main difference is in the past it would be like, "Great, I did it." And you'd be like, "Well, did you? You took a shortcut here or this is not quite right or I need to go verify it or you cut this corner."

Host

这就像过去一年大家一直在说的管理实习生那回事。

It's like the managing interns thing that everyone's been saying for the past year.

Mike

是的,顺便说,这对实习生来说很冒犯,但确实如此。

Yeah, it's very offensive to interns, by the way, but yes.

Host

是的,没错。你说它更正确了。所以它更快、更准确、更可靠,然后,根据美国政府的说法,很危险。

Yeah, that's true. And you're saying it was more correct. So it's faster, more accurate, more reliable, and then, according to the US administration, dangerous.

Mike

我觉得另一个方面是它有更强的心理理论——这个词不对——但有点像项目理论。所以它不会只是“哦,我要做这个改动”,然后它说“好的,我来做这个改动”。但实际上,尤其是如果你做过大规模软件工程,最好的工程师会记住这个东西的所有不同部分,而且他们还能预见潜在问题。比如,“我可以做这个改动,但如果我不以这种方式做,那么下一个改动就会越来越难。”我认为这就是我在那类模型中看到的显著差异。

I think the other piece is it has a greater theory of mind—that's the wrong word—but sort of theory of project. So it's less, "Oh, I'm going to make this change," and it'll say, "Great, I'll make this change." But really, especially if you've done software engineering at scale, the best engineers keep in mind all the disparate parts of how this thing is and they also see around the corners. Like, "I can make this change, but if I don't do it in this way, then the next change is going to be incrementally harder." And I think that's been a significant difference I've seen in that kind of class of models.

Host

好的,酷。

Okay, cool.

Anthropic实验室及其目的 Anthropic Labs and Its Purpose

Host

所以当我们谈到 Anthropic Labs 时,人们会想到 Claude Code,因为它确实是你的突破性产品。听起来你的任务基本上就是找出下一个 Claude Code 是什么。你觉得这是对你在 Labs 所做工作的准确描述吗?还有,为什么 Anthropic 需要 Labs?

So when we talked about Anthropic Labs, people think of Claude Code because it is really your breakout product. And it sounds like you've been tasked with basically figuring out what the next Claude Code is. Would you say that's an accurate description of what you're doing at Labs? And also, why does Anthropic need Labs?

Mike

Labs,是的。也值得想想为什么我在 2024 年到达时需要 Labs,以及为什么今天需要 Labs,因为我觉得答案有所转变。我在 Anthropic 的第三周和联合创始人 Ben Mann 一起创建了最初的 Labs 团队,这件事之前一直在酝酿。当时的原因真的很不同。我们所有的产品工程团队只有 25 人,而且我们还没有真正好的模型。我加入时,我们有 Claude Sonnet 3 和 Opus 3;那些在当时是不错的模型,但它们甚至还不如实习生,对吧?它们甚至不是 IC3 工程师。所以如果你只有 25 或 30 个工程师的团队,他们都在做下一个增量的事情。我们觉得模型开始变得更好,但我们没有任何产品能展示这一点。对我来说一个好的试金石是,当我们准备发布一个模型时,我们是否有一个产品、演示或其他能展示非常不同东西的示例?而且随着时间推移这变得更难。比如用 Fable,甚至展示那个周末任务或那种更长时间的工作。所以当时 Labs 的真正目的是确保我们的产品不会落后于正在发生的模型指数级增长。Claude Code 就来自那个最初的团队,因为产品团队的其他人都没有——人们在考虑编码,但没有人有空间去思考,如果我们彻底改变形态因素,并接受模型会以这种方式进化,会怎样?我们在 Labs 做的两个最有用的思维练习:一个是可视化模型今天能做什么和大多数人如何使用它之间的差距,我们能缩小这个差距吗?这是第一个。另一个是想象模型现在不擅长但 6 个月后会很擅长的东西,并确保我们到那时准备好产品。我认为这就是 Labs 的两个指导性问题。从最初的版本中诞生了 computer use。但 computer use 不同,因为我们构建它时它真的很差。我们尝试了很多产品,那是在 Sonnet 3.5 时期。你会说“Claude,你能帮我清理桌面吗?”然后它会点击那个东西并删除文件。你会说“这对发布不安全。我们肯定不会去构建或发布这个。”但我们有那个产品,所以我们每次发布新模型时,都会先在内部检查并说“computer use 变好了吗?”我们会告诉研究团队它变好还是变坏,直到我们说“足够好了。我们真的要围绕它发布产品了。”所以它也给你一个通向未来的灯塔,你可以用它来衡量未来的产品。但然后和现在比较。我们有一个驾驶产品团队,有 co-work,Claude Code 已经发展了很多,我们有我们的平台。现在我觉得这更多是关于——这些产品团队都没有在做这种思考——而更多的是模型进步非常快,甚至我们与它们交互的能力也需要进化。所以我们今天与 Labs 和 Claude Code 合作发布的东西之一是 Claude Code Artifacts。让 Claude Code 不仅能打字回复你,还能画图或给你插图。这部分来自在 Labs 花了很多时间说“仅仅一个文本框和一个大文本响应已经不够了。”当我提到模型和我说话时感觉比我聪明得多时,有时我会说“你能给我画张图吗?因为我真正需要这个来完全理解。”但这真的是我们一直在思考的:是的,我们有很多产品,而且我们实际上有很多产品需要整合。那是我们另一个计划。

Labs, yeah. It's also worth thinking about why we needed Labs in 2024 when I arrived and why we need Labs today, because I think the answer kind of shifts. I started the original Labs team with Ben Mann, one of the co-founders of Anthropic, in my third week at Anthropic, and it had been something that had been bubbling under. At the time, the reason was really different. All of our product engineering team was 25 people, and we didn't have the models really. When I joined, we had Claude Sonnet 3 and Opus 3; those were good models for the time, but they weren't even interns, right? They weren't even IC3 engineers. So if you have a team of only 25 or 30 engineers, they're working on the next incremental thing. And we were feeling like the models are starting to get better, but we don't have any products that show that off. A good litmus test for me is when we get ready to release a model, do we have either a product or a demo or some other illustration of something that is very different? And it gets harder over time. Like with Fable, even illustrating that weekend task or that longer amount of work. So really Labs at the time was to make sure our products don't fall behind the model exponential that's happening. And Claude Code came out of that initial one because nobody in the rest of product work—people were thinking about coding, but nobody had the space to go and think about, well, what if we totally change the form factor and we embrace the fact that the models were going to evolve in this way? A lot of the two most useful thought exercises we do in Labs: one is visualize the gap between what the models can do today and how most people use it, and can we close that gap? That's one. The other one is imagine what the models are bad at now that they're actually going to be really good at in 6 months, and let's make sure we have a product ready for that by then. I think those are the two guiding questions for Labs. And out of that first incarnation came computer use. Computer use was different though because when we built it, it was really bad. We tried a bunch of products with it, and this was around Sonnet 3.5. You'd be like, "Claude, can you help me clean up my desktop?" And it would click the thing and delete the file. You're like, "This is not safe for release. We're definitely not going to build this or ship this." But we had that product so that every new model we released, we'd first check it internally and say, "Did computer use get better?" And we'd tell the research team how it had gotten better or worse until the moment where we said, "It's good enough. We're actually going to put a product out around this." So it also gives you this sort of beacon into the future that you can measure your future products against. But then compare it to now. We have a driving product team, there's co-work, Claude Code has grown a lot, we have our platform. And now I think it's much less about—none of these product teams are doing this sort of thinking—and it's much more that the models are advancing really quickly, and even our capability to interact with them needs to evolve. So one of the things we collaborated with Labs and Claude Code that we shipped today is Claude Code Artifacts. So having Claude Code not just be able to type back to you but also draw a picture or give you an illustration. That partially came from spending a lot of time in Labs saying, "Just a text box and a big text response is not going to cut it anymore." When I mentioned that the models feel like they're way smarter than me when they talk to me, sometimes I'm like, "Can you draw me a picture because this is what I actually need to fully understand this?" But it's really what we've been thinking about: yes, we have a lot more products, and we actually have a lot of consolidation to do in our products. That's another initiative that we have.

可访问性与平台产品张力 Accessibility and Platform-Product Tension

Mike

但在那之中,我们仍有机会让那些不会把所有时间都花在思考提示词、指数增长以及高、中、低努力程度差异上的人更容易使用这些工具。我们在这方面还能做很多。

But within that, we still have an opportunity to make things much more accessible to a person that does not spend all of their time thinking about prompting and the exponential and the difference between high, low, and medium effort. Like there's a lot we can still do there.

Host

但是 Mike,这确实让使用 Anthropic 模型的人处于一个有趣的境地,对吧?你知道,Cursor 我想刚刚以 600 亿美元卖给了 SpaceX。然后有人在推特上发了个梗图,说如果没有这个人,Cursor 本可以卖到 3000 亿美元。图片上是 Boris Tcherny,也就是创建 Claude Code 的人。所以对于那些打算在 Anthropic 技术之上构建产品的公司来说,他们会想:我是想和 Anthropic 合作,还是 Anthropic 会直接去构建我想构建的产品,甚至可能是在合作之后?

But Mike, so there's a it's it puts people using Anthropic models in an interesting place, right? Um there you know, Cursor I think just sold for 60 billion to SpaceX. And someone put this meme on Twitter that like, you know, Cursor would have sold for 300 billion if it wasn't for this guy. And it's a picture of Boris Tcherny, the person who created Claude Code. Um and so for companies that are going to build on top of Anthropic technology, you know, they're going to wonder do I want to partner with Anthropic or is Anthropic going to go ahead and and build the product that I'm going to want to build, potentially even after partnering with them?

Mike

是的。我的意思是,我们以智能体式编程为例,我认为同时作为平台和产品这个更广泛的层面非常有趣。当我们接手项目时,目标往往是推动那个行业领域向前发展。所以,你知道,之前有 AI 编程编辑器,有些确实很好,但没有人像我们用 Claude Code 那样以自由形式去思考它。现在很多产品都有了那种味道,我认为这比原本可能的情况要好。所以,如果我们进入一个行业,只是做和别人一样的事情,但顶着 Anthropic 的品牌,我觉得那是浪费我们的时间,也浪费实验室或产品团队的时间。如果我们进入某个领域,应该是想说:“好吧,我们认为前进的方向是这样的。我们可以构建一个这样的产品。”而且顺便说一句,不存在也不应该存在所有产品都是 Anthropic 产品的世界。那会是一个糟糕的世界,对吧?所以,这希望要么是为公司创造新的空间,要么是展示其他产品也可以融入这些特性的方式。

Yeah. And I mean, we'll take the like agentic coding side and I think the broader sort of aspect of of, you know, being both a platform and a product I think is really interesting. Um when we take on projects it's the goal is often to sort of push that area of the industry forward. So, um you know, there were AI coding editors and some of them were really good and, you know, uh but nobody was quite thinking about it in as sort of free form a way as we got to think about it with Claude Code. And now a lot more products have that flavor than I think would have otherwise. And so, I think if we're ever You can call me out on this, Alex, if like if we're ever entering an industry where like all you're doing is the same thing everybody else is doing, but like you've got the Anthropic brand, I feel like that's a bad use of our time and a bad use of our either labs or product team time. Like if we're going in somewhere it should hopefully be to say, "All right, we think that the direction of travel is this way. We can build a product of that." And then by the way, there's no world, nor should there be a world where like all the products are Anthropic products. That'd be a bad world, right? So, like that is hopefully either creating new space for for companies or sort of showing the way where other products can incorporate that too.

Host

那几乎就像在一家拥有社交、消息、视频的科技公司工作。全都包了,对吧?

It would almost be like working for a tech company that has like social, messaging, video. It's all too Right?

Mike

是的。

Yeah.

Host

Mike?是的。好的。

Mike? Yeah. Okay.

Mike

呃

Uh

Host

我会离开。

I'd leave.

Host

嗯,比如有个问题,你知道,Anthropic 推出了一款被视为与 Figma 竞争的产品,而你之前是 Figma 的董事会成员。然后你辞职了。是这样吗?

Well, there was some question for example when you know, Anthropic launched a product that was seen as competitive to Figma and you had been on the Figma board prior to that. And I I you stepped down. Is that correct?

Mike

是的,所以嗯,这是 Alex 提出的一个好问题,我认为,硅谷以这种非常健康、充满活力、容忍风险的创业生态系统而闻名。当大公司带着大量风险投资和大量资源进来时,人们会说:“等等,他们是不是基本上要偷走我的想法?”

Yeah, and so um it's a good question that it's Alex has brought up, I think, where Silicon Valley is known for this really healthy, vibrant, risk-tolerant startup ecosystem. And when the big start coming in with tons of venture capital and, you know, a lot of resources, people say, "Well, wait, are they just are they essentially just going to steal my idea?"

Host

是的。

Yeah.

Mike

现在,我认为我们的双重存在,这也是其他公司必须应对的。我们在之前的讨论中谈了很多关于亚马逊的事情,比如他们必须应对这个世界,他们既是基础设施提供商,显然也有非常大的电子商务业务。他们做视频,但也提供视频服务。然后,总的来说,客户可以生活在那个双重世界里,比如“好的,我在使用他们的基础设施”,同时也知道他们也在使用自己的基础设施来做这些。我认为你可以和我们的客户谈谈,看看我们在这方面做得如何。我总是试图做的事情是,至少以非常透明的方式处理。所以,Cursor 的例子很有趣,Michael 和我在那段时间里谈了很多关于事情发展方向的问题。嗯,你知道,对于其他我们考虑的产品,类似地,我认为有几件事:透明度和共享构建模块。我认为总的来说,我实际上不认为有任何情况是像我们试图建立在其他地方也可用的相同能力之上。上次我站在 Commonwealth Club 的舞台上是我们年初的医疗保健日。我们没有只发布我们独有的云医疗保健,别人都没有。我们发布了很多插件、技能、MCP 和互补能力。所以,这就是我如何应对这个公认复杂的情况,我不是说这很容易或直接,但这就是我们试图导航的方式。

Now, I think I think our dual existence, and it's something that other companies have to navigate. We talked I'll talked about Amazon a lot in the in the previous panel, like they have have to navigate this world where they are both infrastructure provider. They obviously have a very large e-commerce thing. They do video, but they also serve video. And and then, you know, by and large, customers can live in that dual world of like, "Okay, I'm using their infrastructure." Also knowing that they they are also using their infrastructure to do that. And I think the uh you can talk to our customers and see how well we're doing at this. Like, the thing I always try to do is like at least approach it with a lot of transparency. So, the cursor example is an interesting one where like Michael and I talked a lot over the, you know, time around here's where things are heading. Um and uh you know, with similarly with with the other products that we think about, like, can we I think it's a couple things. It's transparency, and then it's shared building blocks. Like, I think um in general, and I actually don't think there's any cases where this is even true. Like, we're trying to build on top of the same capabilities that are available elsewhere. The last time I was here in the Commonwealth Club on the stage was our healthcare day at the beginning of the year. And we didn't ship like cloud healthcare only we have it, like nobody else has it. We shipped a bunch of like plugins and skills and MCPs and like complimentary abilities. So, that's how I'm not claiming it's easy or that it's a straightforward thing, but it is how we're trying to navigate what is like admittedly a complicated sort of situation.

Anthropic估值与道德定位 Anthropic's Valuation and Ethical Positioning

Host

说到创业公司,Anthropic 严格来说还是一家创业公司。但你值很多钱。我的意思是,最新估值是多少?是

Speaking of startups, Anthropic is still technically a startup. But you're worth a lot of money. I mean, what what's the latest valuation? Is it

Mike

965。

965.

Host

9650 亿美元左右。所以,

965 billion dollars or something like that. So,

Mike

他们以 10 亿美元卖掉了 Instagram,对吧?

They sold Instagram for a billion, right?

Host

创业公司。

Startup.

Mike

按 2010 年的钱算。

In 2010 money.

Host

对,对。自那以后,财务状况发生了很大变化。然而,Anthropic 已经定位了自己。它是一家公益公司,并把自己定位为围绕构建 AI 的更道德的公司。我想知道你能不能谈谈你如何看待这种定位,特别是 Anthropic 在改变硅谷文化方面的作用。比如,我记得谷歌在 2000 年代初如何以多种方式真正改变了硅谷的文化。那你如何看待 Anthropic 的文化现在主导这个新时代?

Right. Right. Financials have changed quite a bit since then. And yet, Anthropic has positioned itself. It is you know, a PBC and it's positioned itself as sort of a more ethical company around building AI. And I'm wondering if you could talk a little bit about how you see that positioning in Anthropic's role in particular changing the culture of the valley. Like I know I think back to how Google in the beginning of the 2000s really changed the culture of Silicon Valley in so many ways. And how do you see Anthropic's culture now dictating this next era?

Mike

是的,这是一个非常有趣的问题。也许我先从内部说起,我认为也有外部因素。我最初加入的原因是我正在结束我的第二个创业公司,知道我想去前沿实验室工作,因为我开始使用这些模型进行编程,它们在编程方面很糟糕,但我能看到它们在编程方面已经糟糕到了极点。它们会改进的。而且我已经开始在这些 API 之上构建。所以我做的创业公司叫 Artifact,我们做的是 AI 驱动的新闻推荐,实际上当时通过 Artifact 阅读了很多大型科技新闻。那是我们添加的功能之一。所以嗯

Yeah, that's a really interesting question. Maybe I'll start like inside and I think there's an external component too. I think the reason I joined in the first place so I was winding down my second startup and knew I wanted to go work at a frontier lab because I'd started to use these models for coding and they were bad at coding but I could see that they were as bad as they were ever going to be at coding. They were going to improve. And I had started building on top of these APIs. So the startup I was doing was called Artifact and we did sort of AI-powered sort of news recommendations and actually read a lot of big tech via Artifact back in the day. That's one of the things we added. So um

Host

但不包括 Wired。

But not wired.

Mike

不包括 Wired。你知道,说实话,你们的付费墙真的很硬。

Not wired. You know, you guys had a really hard paywall to be honest.

Host

说得对。

Fair enough.

Mike

我们在这方面做得不太好。

We didn't do very We didn't do a great job on that.

Host

需要给我的订阅打个折吗?因为我可以帮你弄一个?好的。

need a discount on my subscription cuz I can get one for you? Okay.

Mike

其实真的很有趣,比如做交易

It's actually really funny like the making deals

Host

是的,做交易。就是登录 cookie。实际上很难留住像你这样的人,你知道所以

Yeah, making deals. It's the It's the login cookies. It's actually really hard to keep people like you know so

Mike

我知道。请把这个升级到康泰纳仕。我知道。

I know. I just please escalate this to Conde Nast. I know.

加入Anthropic与公司文化 Joining Anthropic and company culture

Mike

嗯,而且在应用里做邮箱登录非常难。但我在 API 之上构建东西时就想,“哇,好吧。他们能做出非常有趣的东西。”但最终让我加入 Anthropic 的是,他们言行一致,真心深信要努力让 AI 对人类有益。这在公司内部是深入人心的,我认为这也是为什么即使我们成长了,公司依然保持如此凝聚力。我觉得这也证明了联合创始人们,他们多么频繁地谈论这一点。这对我来说是个惊喜,因为我来自一个每周开全员大会、95% 时间谈产品、也许 5% 时间谈公司周围世界其他事情的世界。可能低估了我们的市场推广;也许是 80/20,但绝对是非常非常侧重产品。我记得大约 6 个月后,我和销售组织的一位负责人 Kate Jensen 一起主持了一次全员大会,我们谈到了如何一起做产品和市场推广。人们说,“这太棒了。我终于理解了我们的产品策略和我们一直在做的事情。”我就想,“哦,对。这不是一家产品公司。这是一家使命驱动的 AI 公司,对自己为何存在于世有非常强烈的意识。”

Um, and email login is very hard to do in an app. But I was building on top of the APIs and thinking, "Wow, okay. They're able to do really interesting things." But what ultimately made me go to Anthropic was that they walked the walk and they really deeply believe in trying to make AI go well for humanity. That's in the water internally, and I think it's why the company has remained as cohesive as it has even as we've grown. I think it's also a testament to the co-founders there, how often they talk about this as well. It was a surprise for me coming from a world where at Instagram we did a weekly all-hands and we talked about product 95% of the time and maybe 5% of the time about something else going on in the world around the company. Probably underselling our go-to-market; maybe it was 80/20, but it was definitely very, very heavy product. I remember about 6 months in, myself and Kate Jensen, one of the leaders in the sales organization, did a joint all-hands where we talked about how we're doing product and go-to-market together. People were like, "This is so great. I finally understand our product strategy and what we've been doing." It's like, "Oh, right. This is not a product company. It is a very mission-driven AI company with a very strong sense of why it exists in the world."

对硅谷的影响 Impact on the valley

Mike

我认为对硅谷的整体影响还有待观察。我看到的积极迹象,或者说有趣的迹象,是各界对慈善事业重新燃起的兴趣。这已经有人写过,我认为这将是一个有趣的溢出效应。再说一次,谁知道这一切会如何发展,但取决于结果,可能意味着很多有趣的新慈善部署。然后我认为另一部分是,关于 AI 可能或应该如何发展的对话,是随着技术实时进行的,而不是事后回顾,我认为其他技术浪潮都是事后回顾,而这是件好事。

I think in terms of the overall impact on the valley, it remains to be seen. I think positive signs I've seen, or interesting signs, are a renewed interest in philanthropy across the board. That's something that has been written about, and I think it will be an interesting sort of outflow. Again, who knows how all of this goes, but depending on how it goes, it could mean a lot of interesting new philanthropic deployment. And then I think the other piece is that the conversation around how AI could or should go is happening in real time with the technology versus retrospectively, which I think has been the case for other technology waves, and I think that is a good thing.

赞助广告 Sponsor ad

Host

大家好,我是 Alex Kantrowitz。我想告诉大家我和 Gravity 合作制作的一部纪录片,探讨 AI 智能体安全的未来。为了了解我们是否真正为自主智能体做好准备,我与 MIT 教授 Ramesh Raskar、前白宫 CIO Teresa Payton、米其林集团首席数据和 AI 官 Ambika Rajgopal,以及阿里巴巴前高管 Sharon Guy 进行了对话。他们各自对这个不断发展的领域提供了独特见解。最后,我们与 Gravity 的 CEO Rory Blundell 一起讨论了前进的道路。在 Gravity 的引领下,加入我们的旅程吧。你可以在节目说明中的链接观看完整纪录片。

Hi everyone, Alex Kantrowitz here. I want to tell you about a documentary I've made with Gravity to explore the future of AI agent security. To find out if we're truly ready for autonomous agents, I sat down with MIT Professor Ramesh Raskar, former White House CIO Teresa Payton, Michelin's Group Chief Data and AI Officer Ambika Rajgopal, and Sharon Guy, a former executive at Alibaba. They each offer unique insights into this evolving landscape. We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward. With Gravity leading the way, join us on this journey. You can watch the full documentary at the link in the show notes.

Anthropic实验室与两大主题 Anthropic Labs and two themes

Host

Mike,你谈到 Anthropic 看到了模型能力与大家构建产品之间的差距。通过 Labs,你们试图提前布局,以便向人们展示 AI 现在和 6 个月后可能做到的事情。那么,请告诉我们你在构建什么,你看到哪些潜力,以及人们应该关注什么。

Mike, you talked a little bit about Anthropic having this gap that it sees between the capabilities of the models and where everybody is building products. And with Labs, what you try to do is get ahead of that so you can show people what AI might be able to do now and 6 months from now. So, please tell us what you're building, where you see the potential, and what people should be on the lookout for.

Mike

嗯。

Mhm.

Host

如果可以的话,我还想问,你现在在构建什么,但如果你有像 Elon Musk 在太空建数据中心那种天马行空的野心,我也想知道。把一切都告诉我们。

And if I can throw in, what you're building now, but also if you have a pie-in-the-sky, like Elon Musk data centers in space type ambition, I want to hear about that too. Tell us everything.

Host

你还有 13 分钟。开始吧。

You have 13 minutes left. Go.

Mike

没错。剩下的 13 分钟就是我产品团队的独白。

Exactly. The rest of the 13 minutes is the monologue in my product team.

主题一:自主性与自我认知 Theme one: agency and self-knowledge

Mike

我想也许有两个我非常兴奋的主题,我们一直在大量探索。第一个是给 Claude 一个环境,让它有更多的自主性,也有更多的自我认知。我来解释一下,因为那有很多 AI 术语,但我会给你一个我们目前做得不好的例子。比如,如果你在 Claude 项目中用 Claude 创建了一个文件,你会说,“太好了。你能把它添加到我们的项目里吗?”Claude 会说,“不行。你得去下载文件,然后拖放到这个日期里。”你会说,“什么?”直到昨天,我对 Claude 设计和 Claude Code 也会这么说,你在 Claude Code 里,你会说,“酷。我需要为我们正在构建的东西设计一下。”或者你在 Claude 设计中做了一个模型,你想去构建它,它会说,“酷。这是一个 zip 文件。”你会说,“什么?”所以这有点互操作性的问题,但总的来说,这个主题是给 Claude 很多关于其环境的概念。我实际上和一个客户,一个 API 客户谈过,他们正在试验的一件事是,在 Claude 在其产品的智能体循环中运行时,给它一个他们源代码的安全版本,这样如果它遇到问题,它不会说,“我不知道,我遇到了问题,”它可以说,“嗯,可能是这个东西,”至少当它和软件维护者之一交谈时。所以这个整体主题,当然你必须用保障措施来做,并且要非常小心你解锁了什么。这听起来有点明显,但实际上在这些产品最终能表达多少方面,是天壤之别。你甚至可以看到从 Claude AI 中的核心聊天或经典聊天,到像 co-work 这样的东西,它有更多的自主性,有一个运行时,能够稍微了解它的环境。但我认为我们只走了大约 10% 的路。实际上,我认为人们为什么对像 open Claude 这样的东西感到兴奋的原因之一,是看到一个可修改的 harness,你可以和它谈论事情,你永远不会感觉到,“哦,对不起,我不能那样做,你得去这个设置屏幕打开它,”它只是它能访问的东西,希望有正确的保护和权限。所以这是第一个主题,我非常兴奋,我认为如果我们做对了,它应该会彻底改变我们所有的产品。

I think maybe two themes I'm really excited about that we've been exploring a lot. The first one is giving Claude an environment where it has more agency and it also has more self-knowledge. I'm going to unpack that because that's a lot of AI words, but I'll give you an example of where we are currently doing a bad job of this. Like, if you are in a Claude project and you make a file with Claude, you're like, "That's great. Can you add it to our project?" Claude will be, "No. You have to go download the file and drag and drop it into this day." And you're like, "What?" Until yesterday, I would have said the same thing about Claude design and Claude code, where you're in Claude code, you're like, "Cool. I need a design for this thing that we're building." Or you're in Claude design and you make a mock-up and you want to go build it, be like, "Cool. Here's a zip file." And you're like, "What?" So that's a little bit of interoperability, but in general, this theme of giving Claude a lot of notion of its environment. I was talking to actually a customer, an API customer, and one of the things they were experimenting with is actually giving Claude a secure version of their source code while it's running in the agent loop in their product, so that if it hits an issue, it doesn't go like, "I don't know, I hit an issue," it can be like, "Well, it's probably this thing," at least when it's talking to one of the maintainers of the software. So that overall theme, and of course you have to do it with safeguards and be really careful about what you unlock with it. It sounds kind of obvious, but it's actually night and day in terms of how expressive these products end up being able to be. You can even see it going from core chat or classic chat in Claude AI to something like co-work, where it's got a little bit more agency and it has a runtime and it's able to understand a little bit about its environment. But I think we are at like 10% of the journey about where we could go. Actually, one of the reasons I think people got excited about things like open Claude is seeing how a harness that is modifiable and you can talk to it about things, and you don't ever get the sense of, "Oh sorry, I can't do that, you're going to have to go to this settings screen and turn it on," it's just a thing it has access to, and hopefully with the right guarding and permissions. So that's theme one that I'm extremely excited about, and I think if we do it right, it should actually transform all of our products from head to toe.

弥合能力与现实差距 Closing the gap between capabilities and reality

Mike

另一部分是,我可能会分享——如果不是我们正在做的内部产品——我从反馈中得到的一句话是“缩小差距”。我说过要缩小能力与现实之间的差距。我认为这也是缩小人们对自己工作的理解与实际日常工作之间的差距。我和我们隐私团队的一个人聊过,要把一个工单通过任务追踪器从一个队列移到另一个队列,需要八步复制粘贴,手动移动。这很烦人,而且容易出错。我得不断抽查。我们用实验室的一个项目帮她解决了这个问题。她说:“啊,这是我职业生涯中的第一次。”她工作了 30 年。我们说:“我脑子里的想法和我现在用的东西,现在是一回事了。”差距已经弥合了。我想把这种感觉带给每个人。当然,Claude 让很多非技术人员能够编程。但我们仍然要求人们理解太多概念,比如我的沙盒环境和生产环境有什么区别,或者以我自己或他人的身份连接 MCP,或者我应该如何存储数据?当然,你不能抽象一切,但如果你把这两个主题结合起来——给 Claude 很多自我知识,并创造一个环境,让你能以可重复的方式为人们解决复杂问题——我会非常非常兴奋。

The other piece is, and I'll maybe share, if not the internal product we're working on, the phrase I got out of feedback was closing the gap. I talked about closing the gap between capabilities and reality. I think it's also closing the gap between how people understand their own work and then how the actual day-to-day is to do that work. I was talking to somebody internally who's on our privacy team, and to move a ticket from one queue through another one via the task tracker into another one was like eight different steps of copy and pasting, manually moving it. Pretty annoying to have to do, probably error-prone. I have to keep spot checking it. And we helped her with one of our Labs projects to basically make that not a pain. And she's like, "Ah, this is the first time in my career." And she's been working for 30 years. We're like, "What's in my head and what I am using is now this." Like it is now closed. And I want to bring that feeling to everybody. Of course, Claude unlocked a lot of non-technical people being able to code. But we're still asking people to understand way too many concepts, like what is the difference between my sandbox environment and production, or connected MCP as myself or others, or how should I store data? And of course you can't abstract everything, but if you combine both of those themes—if you give Claude a lot of self-knowledge and you're creating an environment where you can actually solve complex problems for people in repeatable ways—I get very, very excited about that.

登月计划与SpaceX Moonshot and SpaceX

Host

那你的登月计划呢?

And your moonshot?

Mike

嗯。

Yeah.

Host

不让你蒙混过关。你的登月计划是什么?

Not letting you off the hook. What's your moonshot?

Mike

登月计划?

Moonshot?

Host

对。

Yeah.

Mike

呃,太空方面没什么。不过我想我们正在和 SpaceX 谈一些太空相关的事情。

Uh, no nothing in space. Although I guess we're talking to SpaceX about spacey things.

Host

等等,你们在和 SpaceX 谈?

Wait, you're talking to SpaceX?

Mike

我是说,有一个内部……

I mean, there's an internal...

Host

当然。你知道,为了算力。对。

Absolutely. You know, for compute. Yeah.

Mike

嗯,抱歉。

Um, is it sorry.

Host

是在探索轨道外吗?那句话怎么说来着?关于探索……

Was it exploring extra orbital? What was the phrase? Something about exploring...

Mike

探索,对。

Exploring, yeah.

Host

轨道后的世界之类的事情。

Post-orbital world things.

Mike

绝对不是我负责的部门。但是,是的,有东西在……

Definitely not my department. But yeah, there's stuff in...

Host

你是说——所以实验室并没有专门和团队合作处理算力?

Are you—so Labs isn't working specifically with the team on compute?

Mike

对,没错。那些是独立的算力……

Right, exactly. Those are separate compute...

Host

完全分开。好的。好的,那你的登月计划。

Totally separate. Okay. Okay, so your moonshot.

Mike

嗯。

Yeah.

Host

你个人相信太空数据中心吗?

Do you personally believe in data centers in space?

Mike

我聊过一次。我远不是数据中心专家,但我和一个往太空送东西的人聊过,不是埃隆·马斯克。你会这么说。

I had a conversation. I'm far from a data center expert, but I talked to somebody who is a person who sends things to space who's not Elon Musk. And that's what you would say.

Mike

他们非常看好。我试图理解为什么,基本上就是,我想如果你转换得好,就有几乎无限的电力,还有无限的土地。我说:“好吧,你信这个。”然后我觉得他们对必须做的屏蔽也很有信心。再说一次,这显然不是我的专业领域,但那次谈话之后我想:“好吧,我明白了。”即使需要几年时间。起初我承认总体上觉得这是个疯狂的想法,但现在我想:“哦,我真的能理解为什么这可能说得通。”

And they're really bullish. And I was trying to talk about why, and it was basically like, I guess effectively infinite power if you convert it well, and infinite land. And I was like, "Okay, you buy that." And then I think they feel good about the shielding you have to do. Again, clearly not my area of expertise, but after that talk I was like, "Okay, I see it." Even if it's going to take a few years. At first I admittedly thought it was a crazy idea in general, but now I'm like, "Oh, I actually really can understand why this might make sense."

代币经济学与衡量 Token economics and measurement

Host

当你之前谈到 Claude 的工作将被压缩以及所有这些步骤时,我不禁想到 token,以及你知道,如果人们必须采取这么多步骤并使用这么多 token,短期内对你的商业模式可能有利。但 token 已经成为我们现在用来描述这个行业的经济单位,人们在 token 最大化,现在又在 token 化。第一,我想听听你在这个光谱上的位置,你是不是一个 token 最大化者。第二,在不久的将来,行业是否不再以 token 来衡量?你知道,它会像 MIPS 或拨号上网那样,或者有其他衡量单位真正定义这个时代的经济?

When you were talking earlier about the ways that the work in Claude is going to get compressed and all those steps, I couldn't help but think of tokens and how, you know, maybe it's good for your business model in short term if people have to take so many steps and use so many tokens. But tokens have become this unit of economics that we're using to describe the industry now, and people are token maxing and now they're tokenizing. And one, I want to hear how you sit on that spectrum if you're a token maxer. And two, is there a near future in which the industry is not actually measured by tokens? You know, it goes the way of MIPS or dial-up or some other unit of measurement that actually defines the economics of this era?

Mike

是的,我觉得这两个都是非常有趣的问题。今年早些时候,当你开始听说有公司用仪表盘显示谁用得最多时,我们当然也有内部指标。我们发现,使用 token 最多的人和那个我会——这其实是个有趣的想法。也许你们公司可以试试:写下你认为最有生产力的 10 个人,然后看看 token 使用量前 10 的人,看看相关性有多高。至少对我们来说,相关性没那么高。纯粹推崇最大使用量似乎很危险。显然这很容易被操纵,但即使除此之外,我认为——是的,你可以让 Claude 做 10 个不同的变体,但如果你深入思考,也许你会选择两个你认为最有希望的,如果之后有迭代,再加第三个。所以我不会说自己是 token 最大化者。实际上,最 token 化的事情是我做的那个转换——那只是几百万个 token。把那个东西从 Python 转换成 TypeScript 花了很多 token。但我认为人们对这些不同部分越来越深思熟虑。我们每次看模型发布时,看的不仅仅是模型智能,我们还在认真考虑模型智能、努力和 token 效率的组合。我认为这是我们必须改进的一个大杠杆:我们如何继续在给定任务上越来越 token 高效,这样你也不必太费心。我们可以自动做到这一点,但我们能更好地调整解决方案以适应问题。然后回答你的第二个问题,是的,当我还在 CPO 职位上时,我考虑了很多基于结果的定价,如果能做到的话会非常有趣。当然,如果你和世界上那些 Sierra 和 Finn 们谈,他们有非常清晰的——比如我们保住了这个,我们能够解决这个客户请求而不升级——那非常清晰。但在我们如今实际让 Claude 做的这些任务上,就模糊得多了。比如我有一份战略文档,我用 Claude 来批评我的战略文档。结果是什么?就像,嗯,我不知道。告诉我六个月后战略会怎样。

Yeah, I think both of those are really interesting questions. It was interesting earlier this year when you started hearing about companies that have dashboards showing who used it the most, and we of course have internal metrics as well. And we found that there's not a lot of correlation between the person who's using the most tokens and the person that I would—it was an interesting thought actually. Maybe do it at your companies: write down your 10 most productive people that you think are most productive, and then get your top 10 token users and see how closely they correlate. At least for us, it wasn't that correlated. It seemed dangerous to sort of purely glorify the maximum uses. Obviously it's very gameable, but even beyond that, I think it's—yes, you can ask Claude to do 10 different variants on something, but if you thought about it deeply, maybe you would choose two that you thought were most promising, and a third one if you then had some iteration on that as well. So I would not say I'm a token maxer. Actually, the tokeniest thing was that conversion thing I did—it was just a couple million tokens. It's a lot of tokens that it took to convert the thing from Python to TypeScript. But I think people are being more thoughtful about these different pieces. And one of the things we look at whenever we look at a model launch is not just model intelligence, but we're also really thinking about model intelligence and effort and token efficiency as that combination. And I think that's a big lever we have to improve: how do we continue to be more and more token efficient for a given task, so that you can also hopefully you don't have to think very hard about this. We can do this automatically, but we're able to tune the solution to the problem a little bit more. And then to your second question, yeah, when I was still in the CPO seat, I was thinking a lot about outcome-based pricing as something that would be really interesting to do if you could do it. And of course, if you talk to the Sierras and Finns of the world that have a really clear—like we kept this, we were able to solve this customer request and not have it go escalated—that's really clear. It gets so much fuzzier on these tasks that we actually ask Claude these days. Like I had a strategy document, I used Claude to critique my strategy document. What was the outcome? It's like, well, I don't know. Tell me how the strategy goes six months from now.

基于结果的定价与Claude托管代理 Outcome-based pricing and Claude managed agents

Mike

我觉得要捕捉这一点也会非常困难。但我希望看到更多实验,看能否更好地捕捉它对个人和公司的价值,然后找到最佳方式。我们在这方面最具体的进展是一款叫 Claude 托管智能体的产品,我们会为你运行所有基础设施,包括智能体式框架和工具调用等。你可以用普通模式,给它任务,它会消耗 token,完成后告诉你。或者我们有基于结果的模式,你可以说,这是好的标准,这是评分标准,去做吧,它会朝着结果去执行。所以如果大家都转向那个 API,我想也许我们可以采用不同的基于结果的定价,但要看它被采用的情况。

It feels like it's going to be very hard to capture that as well. But I would like to see some more experimentation around can you better capture what it's worth to the individual and then to the company, and can we find the best way to do that as well. The most concrete thing we've moved towards that is a product called Claude managed agents, where we'll run all of the infrastructure for you in terms of doing the agentic harness and calling the tools, etc. You can either do it in the normal mode, which is you give it tasks, it will go through tokens, it'll tell you when it's done. Or we have an outcome-based mode where you can say, here's what good looks like, here's a rubric, go and do it, and it'll go off and make it more outcome. So if everybody had moved on to that API, then I think maybe we could have a different sort of outcome-based pricing, but we'll see how that gets adopted.

随机图片:应用发布与使用 The random image: app releases vs usage

Host

约翰,或者后面的同事,能调出那张随机图片吗?能展示一下吗?如果可以,太好了。

John or guys in the back, can we have the random image? Can we show the random image? If we can, great.

Mike

就是那张随机图片。

It's the random image.

Host

哦,出来了。好的,这只是因为我们没有给它起个好名字,所以就叫它随机图片,因为它可能随时出现。但这是《金融时报》的一张图表。说到效用,它显示了应用发布数量,正在飙升,而使用量显著的应用似乎在下降,应用评价也在下降。所以迈克,我想听听你对这张图的看法。有没有可能大家都在编码和发布,但我们并没有真正看到生产力的巨大提升?

Oh, here it is. Okay, it's just because we didn't have a good label for it, so we just called it the random image because it might come up at any point. But this is a chart from the Financial Times. Speaking of utility, it shows the amount of app releases that have come out, which are skyrocketing, and then apps with significant usage that seem to be going down, and app reviews, which seem to be going down. So Mike, I'd love to hear you respond to what we're seeing in the image here. Is it possible that everybody's coding and releasing, but we're not really seeing a big boom in productivity?

Mike

这真的很有意思。我觉得在整体应用使用上肯定有类似情况。我很想知道这些发布的应用中有没有成为使用量显著的应用。

That's really interesting. I think there's definitely a parallel in app usage in general. I'd be interested in seeing if any of those app releases became one of the apps with significant usage.

Host

好了,可以撤下了。好,你继续。

All right, we can take it down. Yeah, but go ahead.

Mike

我觉得这和我一直在思考的某件事有关。显然我的背景是消费者领域,我一直在想消费者 AI 的突破最终会是什么。我不知道我们是否已经看到了很多。部分原因是,我不知道那张图回溯了多久,但当我们发布 Instagram 时,应用领域还像狂野西部。人们对应用很兴奋,两个随机的人发布一个应用,就能在 3 个月内冲到照片和视频类第一名,对吧?我觉得现在难多了,想想前十名有多集中,人们花多少时间在 TikTok 和 Reels 这类应用上。很多,对吧?所以我认为要获得突破性的消费者体验真的非常非常难。所以这既是一个关于如今消费者产品有多集中的故事,第一;第二,拥有那种数据有多根深蒂固、多强大——人们称之为数据引力。比如你的 Google 文档都在 Google 文档里,数据引力就在那里。所以即使有人有一个好 2 倍的 AI 文档编辑器,你会把所有东西搬过去吗?也许,可能不会。所以我认为这说明了那些有粘性的东西。我经常想的是,困难的事情仍然困难。做出人们想要的东西仍然非常困难。我们内部在这个话题上有很棒的争论。不是我们所有的产品都成功,对吧?所以我认为这对像我这样的产品人来说是个看涨信号,因为这意味着我们希望能继续增加价值。但我认为那张图可能又是一个例子,说明即使你能更快地编码,突破也比以往任何时候都难。我们能不能在一个月内而不是三四个月做出 Instagram?可能。但我们是在经历了漫长曲折的过程后才做到的。

I think it ties into something I've been thinking a lot about, which obviously my background is in consumer, and I've been wondering what the consumer AI breakouts will end up being. I don't know if we've seen a lot of them yet. Part of it is, I don't know how far back that chart goes, but when we were releasing Instagram, it still felt a little bit wild west in terms of the apps. People were excited about apps, and two kind of random people released an app and were able to get to number one in photos and video within 3 months, right? I think that is much harder now when you think about how consolidated the top 10 is, and how much time is spent on the TikToks and Reels of the world. It's a lot. So I think getting that breakthrough consumer experience is really, really hard. So that is as much a story about how consolidated consumer products are these days, number one. Number two, how entrenched or how powerful it is to have that sort of data—people call it data gravity. Like the data gravity of something like your Google Docs are in your Google Docs. So even if somebody has a 2x better AI-powered doc editor, are you going to move all your stuff? Maybe, probably not. So I think it speaks to the things that are sticky. I think a lot about is that the hard stuff is still hard. Making something people want is still really hard. We have amazing debates internally on this topic. Not all of our products work, right? So I think that's a bullish sign for product people like me because it means we hopefully still add value. But I think that chart is maybe another place where it's harder in many ways than ever to break through even if you can code more quickly. Could we have done Instagram in a month instead of three or four? Probably. But we got there after a long, winding process with twists and turns.

Instagram团队规模与AI工具 Instagram team size with AI tools

Host

你卖掉 Instagram 的时候有 18 个人吧?

Had I think 18 people at Instagram when you sold it?

Mike

13 个。

13.

Host

有了这些工具,你觉得你会有多少人?

With these tools, do you think you would have had how many people?

Mike

嗯,这真的很有意思,因为那 13 个人里……

Well, it's really interesting because of those 13...

Host

因为大家都说,哦,10 亿美元,10 亿美元,1 人创业公司。你们能有多接近?

Because everyone's like, oh, 1 billion, 1 billion dollar, 1 person startup. How close could you guys have gotten?

Mike

我觉得我们可能四五个人就能做到。

I think we could have gotten there with like four to six.

Host

好的。

Okay.

Mike

是的。如果我们长大了,我们能做的事情之一就是可以同时做多件事。就像 Instagram——如果你看过我 5 岁儿子踢足球,我的意思是球在那里,每个人都跑向球。那就是我们的产品团队。就像,视频,上。然后每个人都去做那一件事。我们本可以打位置。比如我们为 Instagram 大约一个月就建了 Android 版。有了模型,我们可能一周就能完成。但为了建 Android,我们把所有人都从 iOS 上撤下来,我们都重新学习用 Android 系统编码。然后我们去做那件事。整整一个月,我们几乎没在 iOS 上发布更新。所以我认为你可以更……实际上,一个很好的例子。我有一个内部实验室项目,帮助加速 Anthropic 工程师编码和代码审查。那个项目我维护 iOS 和 Android 两个版本,我基本上让负责 iOS 的云端智能体去联系 Android 那个,说,我实现了这个。抱歉 Android 用户,即使在 AI 世界里它还是第二个,抱歉。然后 Android 版本说,好的,我要做这个。哦,那不算,因为那个功能在这里没意义。我要放弃它。当然,我们在 Instagram 时不可能委托所有这些,但通过平台对齐,我们肯定能做得更多。这种平台接近对齐的梦想现在其实相当可行。

Yeah. The thing that we would have done if we had grown up is we'd be able to do some things in more than a single track. Like Instagram was—if you ever watched my 5-year-old play soccer, by which I mean the ball is there and every single person runs to the ball. That was our product team. It was like, video, go. And everybody goes and works on the one thing. We'd be able to play positions. Like Android we built in about a month for Instagram. We could have done it probably in a week with the models. But to build Android, we took everybody off iOS and we all relearned to code in Android OS. And then we went off and did that. For that whole month, we were barely shipping updates on iOS. So I think you can be a lot more... Actually, a really good example. There's a labs project I have internally that helps accelerate how Anthropic engineers code and do code review. That project I am maintaining an iOS and an Android version of, and I basically have the cloud that works on the iOS one basically ping the Android one and be like, yeah, I implemented this. Sorry Android users, it's still the second one even in the AI world, sorry. And then the Android version is like, okay, I'm going to do this. Oh, that doesn't count because that feature doesn't make sense here. I'm going to drop it. And of course we wouldn't have been able to delegate all of that at Instagram, but we sure could have done a lot by having platform parity. Like this dream of platform close to parity is now actually quite doable.

关于团队裁减与哥谭的玩笑 Jokes about team cuts and Gotham

Host

你现在可能会接到 Instagram 团队剩下那六七个人的电话,问,我是谁,我在新时代里被裁掉了吗?另外,听起来如果你真想的话,现在可能可以把 Gotham 带回来了,用所有这些……

You're probably going to get calls now from the remaining six or seven people on your Instagram team going, who was I, did I make the cut in the new era? Also, it sounds like you probably could bring Gotham back now if you really wanted to with all the...

Mike

我忘了我们最后有没有……我想愚人节那天我们可能把它带回来过一天。是的。

I forget if we eventually... I think for April Fools maybe we brought it back one day. Yeah.

关于道德困境产品的最后问题 Final question on ethically fraught product

Host

好,我们还有时间问最后一个问题吗?

Yeah, do we have time for one more question?

Mike

最后一个问题。好的。

Last question. Yeah.

Host

好的,我最后一个问题是,你曾参与一个产品,随着它的发展,现在在很多方面存在伦理争议,因为人们担心它对儿童的一些伤害。

Okay, my last question for you is you worked on a product that now, as it has evolved, is in many ways ethically fraught because of some of the harms that people are concerned about with children.

消费者AI风险与危害 Consumer AI risks and harms

Host

当你谈到目前还没有真正爆款的 AI 消费级应用时,我认为其实是有的,那就是聊天机器人,对吧?而聊天机器人也给年轻人带来了一些真正的危险和伤害。那么当你们在实验室里开发时,是如何考虑让这项技术变得更好所带来的潜在危害和风险的呢?

And when you talk about the fact that there hasn't really been a big breakout consumer app for AI, I think there has and it's chatbots, right? And chatbots have also led to some real dangers and harms for young people. And so when you are building in labs, how are you thinking about the potential harms and the risks that come with just making this technology that much better?

Mike

是的,我认为确实有一些产品我们做过原型或概念设计,然后觉得“这个产品听起来太夸张了,我讨厌它”。但如果发布的话,它会对世界产生负面影响,或者会把人们引向错误的方向。即使我们做对了,我们认为那个错误的或道德上更有问题的版本也会是主动有害的。所以我认为在内部经常问这个问题会带来不同。而且我们的核心产品和模型都表现非常好,这是一种奢侈,所以如果我们认为某个东西可能获得大量使用,这在某种程度上是个容易的决定。但回到之前的对话,我认为前置思考非常重要,真正想清楚……现在公司里有人专门思考你所构建的东西对世界的影响,这已经更常态化了,尤其是在 Tropic 肯定有经济学家这样做。而我认为在过去的很多年里,在大多数社交媒体公司里,情况并非如此。

Yeah, I mean I think there are certainly products that we have either prototyped or conceptualized and been like, "This product... it kind of sounds so hypey. I hate this." But like, this product if shipped would be bad for the world or would nudge people in the wrong direction. Or even if we did it right, the wrong or more morally fraught version of this would be actively, we think, bad. And so I think asking that question a lot internally makes a difference. And it's a luxury to have core products and models that are doing really, really well, so we don't... that's in some ways an easy decision if we think it could get a lot of use. But yeah, I think going back to an earlier conversation, I think front-loading it is really valuable and really thinking through... it is now more normalized to have people at a company, and definitely at Tropic does, for economists thinking about the impact of the thing that you're building on the world. And that was not the case in along the years on most of social media, I think.

结束语 Closing remarks

Host

Mike,和你聊天总是很愉快。再次感谢你今天带来的见解。我们很快再聊。让我们为 Mike 和 Lauren 鼓掌。谢谢。干得好。谢谢。

Mike, it's always great speaking with you. Thank you again for bringing your insight today. And let's do it again soon. Let's hear it for Mike and Lauren. Thank you. Great job. Thank you.

互动版:逐字朗读 + 针对本期提问 →