Opus 4.5:深入探讨 Anthropic 的研究与产品策略

Opus 4.5: Deep Dive into Anthropic's Research and Product Strategy

黛安·娜·潘恩 Dianne Na Penn · Unsupervised Learning · 2025-12-02 · 约 42 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 研究产品负责人讨论 Opus 4.5 的能力、研究过程以及安全关注如何助力产品开发。

Anthropic's head of product for research discusses Opus 4.5's capabilities, the research process, and how safety focus aids product development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 28)

全文 · Full transcript(中英对照)

引言与Opus 4.5概览 Introduction and Opus 4.5 Overview

Host

Opus 4.5 是一个非常令人印象深刻的模型。它的基准测试表现惊人,很多公司都从中获得了出色的成果。在 Supervised Learning 播客上,我与 Anthropic 研究产品负责人 Diane 坐下来聊了聊 Opus 4.5 的一切。我们谈到了这个模型新解锁的能力,以及 Diane 认为我们在计算机使用方面处于什么阶段。我们讨论了 Anthropic 如何进行研究、如何打造这样的模型,以及他们如何在内部使用这些模型。我们还谈到了模型的下一步发展,以及过去几年脚手架(scaffolding)的演变。最后,我们聊了聊 Anthropic 的安全关注点如何在产品端帮助他们。深入探讨研究过程、模型能力以及其他许多事情,真的非常有趣。我想大家会很喜欢。闲话少说,下面就是我们的对话。非常感谢你来做客播客。

Opus 4.5 is a really impressive model. It's had amazing benchmarks. A bunch of companies are getting great outcomes from it. And on supervised learning, I got to sit down with Diane, the head of product for research at Anthropic to talk about all things Opus 4.5. We hit on what's newly enabled by the models, as well as where Diane thinks we are with computer use. We talked about how Anthropic does research and gets to models like this and how they use these models internally. We talked about what's next for models and the evolution of scaffolding over the past years. And we hit on how Anthropic's safety focus actually helps them on the product side. This is just a lot of fun to get to go deep into the research process, model capabilities, and a bunch of other things. I think folks really enjoy it. Without further ado, here's our conversation. Thanks so much for coming on the podcast.

Dianne Na Penn

谢谢你邀请我。提前祝你感恩节快乐。

Thank you for having me. Happy early Thanksgiving.

Host

是啊,感恩节前夜做一期播客真是一种享受。我觉得你们在节前发布这个大模型,就像是送给我们所有人的礼物。我敢肯定,明天感恩节餐桌上,很多创始人都会说他们对此心怀感激。

Yeah, Thanksgiving Eve pod is like a real treat to get to do. I feel like you guys dropped this big model as a gift to us all right before the holidays. I'm sure a lot of founders will be saying they're grateful for it tomorrow at the Thanksgiving table.

Dianne Na Penn

太棒了。这正是我们想要营造的氛围。

Love it. That's the vibe we're going for.

开始开发Opus 4.5 Starting Work on Opus 4.5

Host

你们之前已经有一些非常令人印象深刻的模型了。那么,开始做 Opus 4.5 这样的工作时,是什么样子的?你是如何思考改进这些模型真正需要什么,以及整个过程最终是怎样的?

You guys already had some really impressive models. What does it look like when you begin the work for something like Opus 4.5? How do you think about what actually is required to improve these models and what does that whole process end up looking like?

Dianne Na Penn

我认为在某种程度上,我们有一个相当雄心勃勃且长期的路线图,围绕我们关心的模型能力以及我们希望改进的方向。这包括更好的指令遵循、编码进步、让模型在记忆方面更出色等等。实际上,每一代 Claude 在某种程度上都是这些能力得以表达的载体。因此,即使我们在设计新版本的 Claude 时,核心也是围绕我们想要确保交付的整体进步,以及如何以与用户当前用例产生共鸣的方式打包、定位和定价,甚至包括那些用户可能尚未意识到 AI 可以帮助他们完成的任务。所以,你会选择一组你想要改进模型的问题或方面,然后看改进这些问题的研究方向是否明确,或者你是否有一大堆不同的东西,需要在早期弄清楚哪些真正能推动进展。

I think in some ways we have a pretty ambitious and long-range roadmap around model capabilities that we care about and that we care about making improvements on. So this includes things like better instruction following for users, coding advancements, making the models better at memory, etc. And really in some ways every generation of Claude is like the vehicle by which these capabilities are expressed. So even when we are designing new versions of Claude, it's really around what are the overall advancements that we want to make sure can be delivered and how do we actually package it, position it, price it in a way that resonates for the types of use cases that users have today, but also they might not even be aware that AI can help them take on. So you kind of pick a set of problems or things that you want to improve the models on, and then to what extent is it clear the research directions to go down to improve that, or are you having a dozen different things and you're kind of figuring out early on which ones actually help move the needle there.

Host

我认为两者兼有。我们通常对这项技术的能力有很强的感觉,认为它可以在各种用例中产生巨大变革,无论是工程还是其他领域。但也有一些事情让我们感到惊讶,比如用户和开发者如何发现它们。例如,让 Claude 非常擅长 Excel 和 PowerPoint——这在年初只是一个相对较小的投入,但我们发现它确实引起了金融服务客户的共鸣,因此我们正在加倍投入,让 Claude 非常擅长这类通用的 Excel 或 PowerPoint 工作,这看起来非常有益。所以两者都有。

I think it's both. I think we have had a very generally strong sense of what we think this technology can do, which is that we think it can be extremely transformative across a wide variety of use cases, whether it's engineering or others. I think other things surprise us with how users and also builders discover them. So things like making Claude really good at Excel and PowerPoint — that was like a relatively small investment earlier on in the year, and what we found is that it really resonated with financial services customers, and so we're doubling down and making Claude really good at that type of general Excel work or PowerPoint work that seems really beneficial. So it's both.

构思加倍投入 Conceptualizing Doubling Down

Host

如何正确理解“加倍投入”?是围绕某个主题获取更多数据,还是在该领域做更多的强化学习?

What's the right way to conceptualize like doubling down? Is that like getting more data around a topic or doing more RL around that area?

Dianne Na Penn

是的,我认为对于实际应用来说,确实是从用户开始的——既包括那些可能带着用例来找我们的用户,也包括当我们思考像计算机使用这样的东西时,人们为什么要在意?我们几乎必须想象一个我们想要走向的世界,所以有一个想象阶段。作为产品经理,我能想到的最接近的东西就是产品愿景文档或 PRD,你要弄清楚“那又怎样”,为什么有人会来使用这个解决方案,然后将其转化为实际的评估——我们称之为 evals,对吧?我们要构建哪些评估来知道模型是否真的擅长或不擅长?也许它已经完成了一半,需要在数据或强化学习上做某些改进来完成最后 50%;也许只完成了 20%,我们实际上需要弄清楚更重大的改变。所以,从这种设想开始,实际上与传统产品管理非常相似,我想这可能会让人感到惊讶。

Yeah, I think it does for practical applications start with users — both users who might be coming to us with a use case, but also when we think of something like computer use, why should people care? We kind of have to imagine almost a world that we want to go towards, and so there's like an imagination phase. As a PM, the closest thing I could think of is like product vision docs or PRDs where you're trying to figure out what is the so what, why should somebody come and use this solution, and then also translating that into practically what are evaluations — we call them evals — right? What are the evals we want to build to know the model's really good or not good at it? Maybe it's already halfway there and it needs to make certain improvements on data or RL for that last 50%; maybe it's 20% and we actually need to figure out much more significant changes we want to make. So starting with that envisioning in mind is actually very similar to traditional product management, which I think might be surprising to hear.

评估与现实价值 Evals and Real-World Value

Host

这真的很有趣,尤其是因为经典的评估感觉越来越脱离现实世界中的实际价值或人们使用这些东西的方式。你既有客户带着当下问题来找你,这很可能非常有帮助;同时,在评估方面,你显然拥有最佳视角,能看到这些模型未来可能做到的事情,并真正想象其中一些用例可能是什么。

It's really interesting, I think especially as the kind of classic evals feel like they're more and more divorced from what actually the value is in the real world or the ways that people are using these things. You have this combination of both customers coming to you with here-and-now problems that are probably really helpful. Evals, you obviously have the best seat to what these models might be able to do in the future and really imagining what some of those use cases might be.

Dianne Na Penn

有没有一些你曾经想象过、Opus 4.5 是第一个能够做到的事情,这些可能出现在一年半前的愿景文档中,你当时想“也许有一天我们能做 X”?或者有哪些模型能力真的让你感到惊讶?

Were there any things that you were imagining that Opus 4.5 is the first model to be able to do that maybe were in the vision doc like a year and a half ago and you were like maybe one day we could do X? Or what were some of the model capabilities that really surprised you?

Dianne Na Penn

是的,我认为其中一些是延续性的,比如能够执行更复杂的智能体编码任务,但也让工作更具迭代性,并且实际上运行时间更长。所以这是一件事——我们开始看到在复杂性和迭代改进交付物方面的一个转折点。我认为计算机使用是另一个例子。我们很早就看到了计算机使用的潜力并一直在投资。去年 11 月,我想是 10 月或 11 月,我们推出了计算机使用 API,此后我们一直在持续投资。我经常使用像 Claude for Chrome 这样的浏览器使用扩展功能,我注意到交互质量有了提升,因为现在 Claude 的视觉能力好多了。所以我认为有时是多种因素共同作用的结果。

Yeah, I think some of them have been like continuations of being able to do more complex agent coding tasks, but also making work that is more iterative and actually longer running. So that's been one thing — we're starting to see more of an inflection point on both complexity and also being able to iterate and continuously improve some of the deliverables. I think computer use has been another one. We saw and have been investing in computer use for a long time. Last year in November, I think November, October, we launched a computer use API and since then we've continued to invest in it. And I use things like Claude for Chrome, which is our browser use extension feature, pretty often, and I saw an improvement just in terms of the quality of that interaction because now Claude's vision is much better. So I think sometimes it's a combination of multiple things working together.

计算机使用演变 Computer Use Evolution

Host

我猜有很多开发者听众对计算机使用感到好奇。你目前有没有一个大致的思维框架,说明计算机用在哪些地方有效、哪些地方无效,以及我们在这条路上走到了哪一步?

I guess there's a lot of builders listening to the podcast that are curious about computer use. Do you have a rough mental framework today of where computer use works and where it doesn't, and where we are on the journey?

Dianne Na Penn

是的,我认为计算机使用已经从我们发布时的早期实验性功能发展起来了。最初我们把它看作第一阶段,像是智能体式编程这类核心功能之上的补充或特性。现在,随着 Opus 4.5 及后续版本的出现,它越来越能独立成为一个端到端的智能体。所以它不仅能在更受限的环境中处理 QA 测试(这是一个流行的用例),还能在网页浏览器中帮助监控、管理并充当智能体。它仍然比开放式的环境(比如你的整个笔记本电脑)更受限,但有一个能帮我在 Google 上重新安排日历的智能体非常有用,因为那往往变得相当复杂。我认为这就是它的发展轨迹:从更受限的环境走向更开放的环境。

Yeah, I think computer use has gone from a very early stage experimental feature when we launched it. Initially we thought of it as probably the first stage, like a complement or a feature on top of something core like agentic coding. Now, more and more between Opus 4.5 and beyond, it's able to be more of an end-to-end agent by itself. So it can not just within a more constrained environment look at QA testing, which was a popular use case, but actually be able to help monitor, manage, and be an agent in a web browser. It's still more constrained than an open-ended environment like your entire laptop, but it's very helpful to have an agent that helps me reschedule my calendar on Google, because that tends to get pretty complicated. I think that's been the arc: moving from more constrained environments towards things that are a bit more open-ended.

Opus 4.5的意外用途 Surprising Uses of Opus 4.5

Host

你第一次内部使用 Opus 4.5 时,有没有什么让你感到惊讶的?

Was there anything you played around with Opus 4.5 when you first had it internally that was surprising?

Dianne Na Penn

从产品团队的角度来看,它实际上非常有助于讨论定价和模型定位等问题,因为它在很多方面都是一个巨大的飞跃。所以从产品经理的角度来看,它不仅是一个优秀的写作者,更是一个出色的思考者。它围绕定价和定位等问题提出了不同的想法,不仅仅是完善我的想法,而是比我见过的其他模型更自发地提出了替代方案。

It was actually really helpful from a product team perspective to debate things like pricing and positioning for the model, because it's such a great leap in many ways. So from a product manager perspective, it was not just a great writer but a great thinker. It came up with different ideas around things like pricing and positioning, where it wasn't just refining my ideas but came up with alternatives more spontaneously than I had seen with other models.

Opus模型的效率与定价 Efficiency and Pricing of Opus Models

Host

传统上,很多 Opus 模型都要贵得多。这个模型既非常好又相当便宜。是什么驱动了这一点,从一开始就有多清楚这个 Opus 模型能够如此高效地提供服务?

Traditionally, a lot of the Opus models were much more expensive. This is both a really good model and quite cheap. What drives that, and how clear was it from the beginning that this Opus model would be able to be served this efficiently?

Dianne Na Penn

从一开始,我们就希望这对 Opus 模型来说是一件好事,我们能够获得更高的效率提升,并将其传递给我们的用户、客户和开发者。除了核心模型训练之外,我们还特意让诸如 effort 参数之类的东西成为可能,我认为这个参数目前被低估了。你实际上可以以一小部分价格获得 Sonnet 4.5 级别的智能。

From the beginning, we were hoping that this would be a thing for Opus models, where we're able to have more efficiency gains and pass them on to our users, customers, and builders. We intentionally also made it possible, in addition to the core model training, for things like the effort parameter, which I actually think is underhyped right now. You could actually get Sonnet 4.5 level of intelligence at a fraction of the price.

Host

我认为这是行业还没有真正做好的事情:仅仅因为一个模型有特定的 token 价格和标价,这并不总是衡量完成一项任务的端到端成本的好方法。我们希望通过 Opus 和诸如 effort 参数之类的东西让它更易获取,这样你实际上可以以更低的成本获得更高的质量。我们从 Opus 开始,我对这个领域的未来感到非常兴奋。

I think this is something the industry hasn't gotten really good about: just because a model has a certain token price and sticker price, that's not always a good measure of the end-to-end cost to achieve a task. What we want to do with Opus and with things like the effort parameter is to make it more accessible, so you can actually achieve higher quality at a lower cost. We're starting with Opus, and I'm really excited about this area going forward.

Host

你觉得大多数人凭直觉就能理解这一点吗?还是你如何看待教育市场,让他们明白不仅仅是每个 token 的成本,还有完成任务所需的 token 数量?

Do you feel like most folks get that intuitively, or how do you think about educating the market around the idea that it's not just the per-token cost but the amount of tokens it takes to complete tasks?

Dianne Na Penn

是的,我认为这是我们作为模型提供商可以做得更多的事情。我从开发者那里听到的一件事是,较小的模型实际上需要更长的时间来完成一项任务,甚至可能无法完成任务,但你却花费了比直接使用 Opus 模型更多的 token。在今年的大部分时间里,我们传统上专注于 Sonnet 并拥有一个核心旗舰产品。随着我们进入不同层次的模型,教育开发者和用户这一点实际上变得更加重要。所以我认为这是我们将从营销角度进行投资的领域。

Yeah, I think it's something we as model providers can be doing more around. One thing I hear from builders is that smaller models actually just take longer to do a task, or might not even get the task done, but you've spent a bunch more tokens than if you just use an Opus model. For a lot of this year, we were traditionally focused around Sonnet and having a core flagship offering. As we move into different tiers of models, it actually becomes more important for us to educate developers and users on it. So I think this is an area we're going to be investing in from a marketing perspective.

早期发布的惊喜 Early Release Surprises

Host

你们对这些模型的能力有很好的直觉,但一旦发布出去,它们就会以你们内部无法想象的无数种方式被使用。在发布的最初几天里,有哪些事情最让你惊讶?

You all have a very good intuitive sense of what these models can do, but then when you put it out there, they start getting used in a million different ways you couldn't have conceived of internally. In these first days of release, what have been some of the things that have surprised you the most?

Dianne Na Penn

是的,现在还非常早。我们也正处于感恩节那一周,所以随着人们测试系统,我的答案会有所变化。我想说可能有两件事。第一,在我们与客户的早期访问中,我们投资了让模型在办公交付物方面变得更好的领域。看到 Opus 4.5 对像 Shortcut(一个销售智能体)这样的客户来说准确度提升如此之大,真的很令人惊讶。客户说,仅仅在不改变工具链或其他改动的情况下,准确度就提高了大约 20%。这引起了巨大共鸣,因为他们可以将这种智能提升传递给他们的用户。第二,我在早期普遍看到的其他事情:人们倾向于在游戏用例上测试模型,特别是 3D 游戏变得更好了。这总是很酷——可视化其中的一些内容一直是人们看到智能边界的简单方式。我也对我们在质量方面收到的反馈感到非常兴奋。客户和用户说:“哦,它真的帮我清除了整个 bug 积压。”

Yeah, it's very early. We're also in the middle of Thanksgiving week, so my answer will change a bit as people test the systems. I'd say maybe two things. One, in our early access with customers, we invested in areas like making the models better at office deliverables. It was really surprising to see how big of an accuracy jump Opus 4.5 was for customers like Shortcut, which is a sales agent. Customers were saying something like 20% accuracy improvement just without changing harnesses or making other changes. That resonated a ton because they can pass that intelligence improvement to their users. Two, other things I've seen generally in the early days: people tend to test the models on gaming use cases, and 3D games have gotten better in particular. That's always really cool—visualizing some of this has been an easy way for people to see intelligence bounds. I'm really excited also for the feedback we're getting around quality. Customers and users are saying, "Oh, it's actually helped me clear out my entire backlog of bugs."

模型的产品市场契合 Product-Market Fit for Models

Host

对于 Opus 4.5,当你思考更广泛的应用生态系统时,我很好奇你的思维模型:在这些模型之上,什么已经有了产品市场契合度,或者今天真正有效?

With Opus 4.5, as you think about the broader ecosystem of applications, I'm curious about your mental model of what has product-market fit or really works today on top of these models.

Dianne Na Penn

我认为主要的几个是:智能体式编程显然会持续存在。我们实际上继续有更多企业就 Claude Code 和编程能力等解决方案进行接洽。我认为同步智能体在编程之外普遍有很强的产品市场契合度理由。

I think the big ones are: agentic coding is very visibly something that's here to stay. We continue to actually have more enterprises reach out about solutions like Claude Code and coding capabilities. I think synchronous agents have pretty strong reason for product-market fit generally beyond coding.

代理用例的转折点 Inflection Point for Agentic Use Cases

Dianne Na Penn

我认为更多是我们整个行业还没找到合适的工具和产品功能来构建上层应用。智能体式编程在网页监控和个人智能体场景下该是什么样子?很多用例仍然非常以聊天为中心。所以我觉得我们正处于一个转折点,需要更多智能来真正推动大量用例。

I think it's more that we as an industry haven't figured out the right harness and product features to build on top. What is agentic coding but for web monitoring and personal agents? A lot of use cases are still very chat-focused. So I feel like we're at an inflection point where we need a bit more intelligence to really boost a large set of use cases.

Host

那你认为我们现在已经具备针对这些用例的智能水平了吗?

And do you think we have that level of intelligence now for those use cases?

Dianne Na Penn

我认为人们会对 Opus 4.5 的优秀程度感到惊讶。我相信 Opus 4.5 会催生新的事物——新的功能和产品。

I think people will be surprised by how good Opus 4.5 is. I think there will be new things born from Opus 4.5—new features and products.

Host

所以更多是主动式的体验,或者智能体主动为你执行任务然后呈现结果,而不是你通过聊天来提示?

So more proactive experiences or agents going out and doing things on your behalf and then surfacing results, versus you prompting via chat?

Dianne Na Penn

是的,这是我们在内部经常听到的一点:模型在没有明确人类指令的情况下就能理解。再加上上下文窗口质量越来越高,以及记忆等功能开始更好地工作,你就得到了一个不仅能完成单一交付物,还能做以前不可能做到的事情的系统——比如监控和维护。这些实际上非常有价值。这不仅仅是发布一个 MVP,还关乎如何维护你构建的东西。

Yeah, that was one thing we heard a lot internally: the model just got it without explicit human instruction. When you couple that with the fact that the context window is getting higher quality, and things like memory start working better, you get a system that can do not just one deliverable but things that were not possible before—like monitoring and maintenance. Those are actually really valuable. It's not just about shipping an MVP; it's also about how you maintain something you build.

优先模型改进与客户反馈 Prioritizing Model Improvements and Customer Feedback

Host

我很好奇你提到的 Shortcut 和改进 Excel 的例子。我想肯定有很多人来找你,希望你在某个领域或任务上改进模型。你们如何确定优先级?基础模型是否会因为优先事项不同而走向不同方向?

I'm curious about the example of Shortcut and improving Excel. I imagine many people come to you asking to improve the model on X domain or Y task. How do you prioritize that? Do foundation models end up going different directions based on what they prioritize?

Dianne Na Penn

我觉得有点道理。确实有客户问我们关于图像生成或视频的事情,而我们一直非常有意——我在这里大约两年半了——专注于扩展智能。所以我认为你在实验室里已经能看到这一点。关于反馈,这是一个双向飞轮。有时我们会听到客户的痛点。今天我们发布速度的好处是,如果客户来找我们,我们可以很快做出改变,因为如果它在 4.5 里,也许就会出现在 Claude 5 里。你有多次机会让你的反馈被听到。然后我们还会自动化更多流程:如何构建系统来有机地找出合适的环境,并自己建立那个循环?以及如何更广泛地交付新进展——比如计算机使用等?所以这是双向的。

I think a little bit, yes. We have customers asking about things like image generation or video, and we've been very intentional—I've been here about two and a half years—about focusing on expanding intelligence. So I think you do see that a bit with the labs already. In terms of feedback, it's a two-way flywheel. Sometimes there are pain points we hear from customers. What's good about our shipping velocity today is that if a customer comes to us, we can make changes pretty quickly because if it's in 4.5, maybe it's in Claude 5. You have multiple chances to get your feedback heard. Then we also automate more of that: how do we build systems to organically figure out the right environments and build that loop ourselves? And how do we deliver new advancements more generally—things like computer use and beyond? So it's bidirectional.

Host

你谈到构建内部工具来更轻松地搭建环境,并降低在特定类别上改进模型所需的工作量,这很有意思。我相信内部工具在这方面确实很有帮助。

It's interesting you talk about building internal tools to more easily spin up environments and lower the effort required to improve models in specific categories. I'm sure internal tooling really helps with that.

Dianne Na Penn

是的,这绝对是一个持续投入的领域。我们有很多事情可以做,所以必须有意地选择投入方向。正如你所说,我们如何让 Claude 来帮助我们?

Yeah, it's definitely a continuous investment area. We have a lot we could be doing, so we have to be intentional about where we spend. To your point, how do we get Claude to help us?

Host

是的。

Yeah.

Dianne Na Penn

当我们完成时。

When we're done.

明确选择:多模态与商业聚焦 Explicit Choices: Multimodal and Business Focus

Host

听起来多模态显然是一个重点领域。还有哪些事情是你们明确选择不花太多时间的?

It sounds like multimodal has obviously been a focus area. Are there other things you've explicitly chosen not to spend too much time on?

Dianne Na Penn

我认为这些是其中一些大的方面。从产品角度来看,我们一直非常有意识地专注于商业用例。这意味着我们在产品上的很多关注和投入都围绕数据安全、隐私、满足企业采用和使用 Claude 的要求,而不是消费级用例。

I think those are some of the big ones. From a product perspective, we've been very intentional about focusing on business use cases generally. That means a lot of our focus and investment in product is around things like data security, privacy, meeting enterprise requirements for them to adopt and use Claude, rather than consumer use cases.

企业代理与模型进展讨论 Discourse on Enterprise Agents and Model Progress

Host

最近关于企业智能体的讨论很有意思,很大程度上是由播客驱动的。Andrej Karpathy 说我们距离真正的企业智能体还有十年,Ilya Sutskever 说当前范式只能带我们走到这里。你如何将这些与 Opus 4.5 这样的改进相协调?你的看法是什么?Anthropic 内部的感觉如何?

There's been interesting discourse around enterprise agents recently, largely driven by podcasts. Andrej Karpathy said we're still a decade away from real enterprise agents, and Ilya Sutskever said this current paradigm only gets us so far. How do you square that with improvements like Opus 4.5? What's your opinion and the feeling within Anthropic?

Dianne Na Penn

我认为智能和模型改进并不像你在评估基准中看到的那样是一条平滑的曲线。它们更加参差不齐。根据不同的评估或个人测试,可能感觉像是一个小跳跃,也可能是一个大跳跃。所以这取决于框架。从客户和用户的角度来看,我们从乐天或 Lovable 等公司听到的是,能力持续提高了他们团队的生产力。在 Anthropic 内部,我们较少谈论 AGI,更多谈论变革性 AI——我们是否在制造具有变革性的技术?我认为我们正走在正确的道路上。Claude 每一代都在改变 Anthropic 员工的工作方式。我们感受到这一点是因为我们以不同的程度采用了它。如果你没有以同样的程度采用,可能就不会觉得有显著的进步。

I think intelligence and model improvements are not always a smooth line like you see in eval benchmarks. They're more jagged. Depending on the eval or personal test, it might feel like a small jump or a large jump. So there's a level of framing. From a customer and user perspective, what we hear from companies like Rakuten or Lovable is that capabilities have continued to improve their team's productivity. Internally at Anthropic, we talk less about AGI and more about transformative AI—are we making technology that is transformative? I think we're very much on that path. Claude transforms every generation of how Anthropic employees work. We feel that because we adopt it to a different degree. If you're not adopting it to the same degree, it might not feel like meaningful advancements.

未来方向:更长期的智能 Future Directions: Longer-Running Intelligence

Host

你们已经取得了很大进展,一些计划中的事项已经完成。下一步的计划是什么?

You've gotten very far, and some things on the board are off now. What's on the next board?

Dianne Na Penn

没有人能预测未来,但我非常期待的一些领域包括向更长时间运行的智能迈进。

Nobody can predict the future, but some areas I'm really excited about include this move towards longer-running intelligence.

长期代理与开放式任务 Long-running agents and open-ended tasks

Dianne Na Penn

所以,不只是人类给 Claude 分配一个具体任务,而是 Claude 承担更开放式的责任。不只是“给我建网站的这部分”,而是维护它、在认为正确时重构代码,不需要那么多手把手的指导。我还对 Office 改进等能力感到兴奋,比如 Excel、PowerPoint——还有很多工作要做。还有计算机使用。我觉得计算机使用在采用率和质量上正在发展,它可能真正带来变革,从企业到普通用户的角度,就像编程那样。计算机能操作计算机是我们大多数人互动的方式。我们现在就在做播客,对吧?所以让 Claude 能在人们工作的环境中交互,会让它更有用。我对计算机使用的未来非常期待。

So again, not just that a human gives and delegates a specific task to Claude, but actually Claude taking responsibilities that are more open-ended. So not just 'build me this portion of my website,' but maintain it, refactor the code when you think it's correct and not needing so much handholding. I'm also really excited about capabilities like Office improvements that I mentioned, things like Excel, PowerPoint — there's a lot more to go. And also computer use. I feel like computer use is actually on the way of adoption and quality where it could really be transformative from an enterprise and also just general user perspective, the way that coding could be. Computers being able to navigate computers is how most of us interact with each other. We're on a podcast right now, right? So being able to have Claude that can interact in these environments where people work makes them much more useful. So I'm really excited about where computer use will go.

Host

你提到了长期运行的智能体,我觉得这确实是世界发展的方向,甚至可能有一群智能体替你工作。在很多方面,显然模型会变得更好,但部分问题在于产品端:长期运行在后台的智能体实际是什么样子?管理十几个智能体又是什么样子?你显然是产品专家——你觉得产品界面可能是什么样的?你有没有看到什么特别有趣的东西?

I mean you mentioned long-running agents and I feel you know it feels like this is certainly where the world's headed and you know maybe even having a fleet of agents doing work on your behalf. It feels like in many ways, obviously the models will get better, but in many ways part of this is solving the product side of it and what does it actually look like to have a long-running agent going on in the background and what does it look like to have a set of a dozen of them that you're managing. You're obviously a product guru — how do you think about what the product surface area might look like here and have you seen anything that you thought was particularly interesting?

Dianne Na Penn

是的。我认为长期运行的智能体和长期运行的智能本身并不是一个用例,对吧?我们需要一个适合它的用户问题。所以我认为真正有价值的是能够维护和迭代改进。比如,假设你是一个投资者,你想了解最新的股票走势或如何调整你的投资组合。这不是一次性的任务。我认为我们今天缺乏的是以易于评估质量改进的方式处理这些长期任务。所以我认为我们开始追踪的最接近的一个——可能不是最合适的,因为没有完美的评估——是 Vending Bench,它……

Yeah. I think long-running agents and long-running intelligence generally is not a use case in itself, right? We need to actually have a user problem that it's a good fit for. So I think things that are really valuable are around being able to maintain and iteratively improve. For example, let's say you're an investor and you want to understand the latest stock movements or how you should adjust your portfolio. That's not a one-time thing. And I think what we kind of lack today is having these really long-horizon tasks in a way that makes it easy to evaluate quality improvements. So the closest one that I think we're starting to track — and it might not be the right one because no eval is perfect — is Vending Bench, which is...

Host

……Claude 经营自动售货机生意,对吧?

...the Claude runs a vending machine business, right?

Dianne Na Penn

我喜欢你们做这个。

I love that you guys do this.

Host

或者 Claude 玩宝可梦,对吧?这些——Claude 在我们最初的时候玩过宝可梦……

Or Claude plays Pokemon, right? These just — Claude played Pokemon when we first...

Dianne Na Penn

我有点怀念宝可梦评估。我不知道是不是我在报告里漏掉了。我喜欢它曾是人们讨论的主要评估之一的时候。

I kind of missed the Pokemon eval. I don't know if I'm just missing it in the reports. I liked when that was one of the main ones people were discussing.

Host

好的,我会记下来,可能会把它带回来。

Okay, I will take a note for potentially bringing it back.

Dianne Na Penn

但也不只是完成任务,而是需要多长时间,对吧?你能不能因为不是暴力解决问题,而是记得小智当时在真新镇,已经和某个用户聊过,从而用更少的步骤完成任务?抱歉,让我宅一下。我们可以深入聊宝可梦。我完全没问题。

But it's also not just about completing the task, but how long does it take, right? And can you actually complete the task in much less steps because you're not brute forcing the problem and you remember the fact that Ash Ketchum was in Pallet Town at this time and already talked to some user, right? Sorry to nerd out for a second. We can go deep on Pokemon. I'm all game for that.

Host

好的。是的。但我想知道的是,智能不仅仅是“我能完成任务并打勾”——越来越多地,它也是判断的质量和模型在情境中拥有直觉的质量。那么我们能否大幅减少达到某个结果所需的时间或精力?有时这确实体现在长期任务中,对吧?但我对这个领域真的很兴奋。我认为整个行业更需要的是更好的评估。

Okay. Yeah. But I guess the part of what I want to know here is just that intelligence isn't just about 'can I complete a task and check the box' — more and more it's also the quality of the judgment and the quality of what intuitions the models can have in a situation. And so can we actually dramatically reduce the time it takes or the amount of effort to get to a certain outcome? And sometimes this really shows up in things like long-running tasks, right? But really I'm just excited about this area. I think what we need more as an industry is better eval around it.

Dianne Na Penn

那么评估是什么?

And what are the evals?

Host

哦,这是个很好的问题。我认为我们需要超越像 SWE-bench 和 TA 这样的东西。我认为特别是用 Opus 4.5,我们达到了大约 80.9%,这些 DC 评估已经非常饱和了。

Oh, it's a very good question. I think we would need to evolve beyond things like SWE-bench and TA. I think especially with Opus 4.5, we got to like 80.9% and these DC evals are so saturated.

Dianne Na Penn

也喜欢那个通过改变票价类别来升级机票的例子——那是一种很有趣的绕过问题的方式。

Also love the example of finding a way to upgrade the ticket through changing a fare category — that was a very fun way to get around things.

Host

是的,这很疯狂,因为去年我推出工具使用 API 时,准确率可能低于 50%。进步太大了。问题是如何衡量和产品化。我认为对我们来说,重要的评估仍然是:人们使用 Claude 的领域以及它的表现如何。所以我认为最终用户的质量衡量仍然重要,或者我们从客户那里得到的反馈。我确实认为我们可能会看到更多向开放式评估的转变,比如 Vending Bench 这类。它不会完全一样,因为我不知道经营自动售货机有多现实,但方向是——它是开放式的。有一些可量化的方式来衡量质量,但不仅仅是是/否,因为世界上有多少任务是是/否的,对吧?如果我们超越编程进入其他类型的影响,很难说生物学中是否有是/否的 SWE 等价物。

Yeah, it's crazy because when I shipped the tool use API last year, I think we were at maybe below 50% on accuracy. It's just again there's so much advancement. It's really how do we measure and how do we productize it. I think the evals that matter continue to be for us: what are the areas where people are using Claude and how well does it work. So I think end-user measurements of quality continue to matter, or feedback that we get from customers. I do think we will probably see more movement towards evals that are more open-ended, like a Vending Bench sort. It's not going to look exactly like that because I don't know how realistic running a vending machine is all the time, but that type of direction — it's open-ended. There is some quantifiable way to measure quality, but it's not just yes/no, because how many things in the world are yes/no tasks, right? If we're moving beyond coding into other types of impact, it's hard to say if there's a yes/no SWE equivalent in biology.

Dianne Na Penn

但你几次提到了人们围绕这些模型构建的“缰绳”,显然你与使用 Opus 4.5 的前沿团队合作密切。你怎么看——你觉得今天人们围绕模型构建的脚手架有典型的一套吗?

But you've kind of alluded a few times now to the harnesses that people are putting around these models, and obviously you work very closely with teams that are at the cutting edge of using Opus 4.5. How would you — do you feel like there's a typical kind of set of scaffolding people are building around models today?

Host

我认为类似于模型智能,脚手架也在进化。我会说在 2022 年、2023 年,甚至 2024 年,很多脚手架更像是训练轮,让模型保持在分布内,它们往往是“不要这样做,总是这样做”的形式,对吧?就像指令和 20 条规则。我认为今年越来越多地,我们看到脚手架更多地围绕智能增强而不是训练轮。最好的脚手架往往是迭代地移除那些不再增强智能的部分。例如,在 Claude Code 上,我们的脚手架相对轻量。

I think similar to model intelligence, scaffolds have evolved. I would say in 2022, 2023, even 2024, a lot of the scaffolds are more like training wheels to keep the model on distribution, and they tend to be of the form 'do not do this, always do this,' right? Just like instructions and 20 rules. And I think more and more this year, what we've seen is scaffolds become much more around augmentations for intelligence rather than training wheels. The best scaffolds tend to be iteratively removing the parts of the scaffold that are no longer intelligence-amplifying. So for example, on Claude Code, our scaffolds are relatively lightweight.

AI代理的工具与框架 Tools and scaffolds for AI agents

Dianne Na Penn

我们给它的工具类型是像批处理工具这样,不是非常特定和独特的。目的是最大化模型在工作时的自主性。我认为会继续有价值的脚手架类型可能是那些增强智能的东西,比如给它通用工具集。多智能体系统今年开始变得更可行了,不只是有一个模型,而是协调一组模型来放大和改进上下文质量等。

The types of tools we give it are things like batch tools that are not specific and very unique. The point is to maximize autonomy of the work as it's being done by the model. The types of scaffolds that I think will continue to be valuable might be things that are intelligence amplifying, like giving it generic sets of tools. Multi-agent systems have started to become more viable this year, where it's not just having one model but orchestrating a set of models to amplify and improve on things like context quality, etc.

Host

我想很多开发者都在问自己:这些东西中哪些会被下一代模型淘汰,哪些会继续为下一代模型增强智能。从内部看是显而易见的,还是说也要等到我们有了那些模型才能知道?

I think a bunch of builders are asking themselves: what set of this stuff becomes obviated by the next generation of models versus continues to amplify intelligence for that next set. Is it obvious from the inside, or is it also like we'll see when we get to those models?

Dianne Na Penn

我觉得有一点:这就是为什么更薄的 harness 和脚手架层很重要。我认为变化程度有点难以预测,但我们确实看到通过模型对脚手架进行一定程度的迭代,质量有所提升。部分原因也是用户需求在变化,对吧?我们看到的是,当我使用 Claude Code 或 Claude 时,如果我发现它非常擅长编辑文档,我可能会给它一大堆任务,比如“嘿,想个替代策略”。所以我们作为用户,也在不断推动别人构建的产品。这不仅仅是为了新模型而更新脚手架,实际上是为了满足用户行为。用户自然会推动你的产品变得更复杂一些,而作为开发者如何满足这种需求,才是正确的思考方式。

I think there is a bit of this: that's why having thinner layers of harnesses and scaffolds is important. I think how much it changes is a little hard to predict, but we do see improvements in quality by some level of iteration on scaffolds by model. Part of it also: user requests change, right? What we see is, when I use Claude Code or Claude, if I see that it's really good at editing a document, I might give it a large set of things like, 'Hey, come up with an alternative strategy.' So we're constantly, as users, pushing the products that people build. It's not just in service of new models to update your scaffold, but actually in service of user behavior. Users will naturally push your product to be slightly more complex, and how you as a developer meet that demand is the right way to think of it.

Anthropic的公司文化 Culture at Anthropic

Host

你从早期就在 Anthropic 了,所以我觉得过去两年半一定是一段迷人的旅程。我想,首先,你如何将 Anthropic 的文化与你工作过的其他地方进行比较?

You've kind of been at Anthropic since the early days, and so I feel like it must have been a fascinating journey these past two and a half years. I wonder, to start, how do you compare the culture of Anthropic to other places you've worked?

Dianne Na Penn

我想说,很大程度上体现在我们的领导者如何表现,以及人们如何谈论使命和目标。实际上,这很大程度上就是我们内部做决策的方式。就是言行一致。我认为 Anthropic 是我加入过的公司中最真实的,从我加入的第一天到现在都是。我想不用说,人才水平和人才密度都非常惊人。这是我工作过的人才密度最高的环境。人们都深度拥有主人翁精神,深思熟虑,友善但直接,并且都是为了把产品、模型、能力、研究做得更好。

I would say a lot of how our leaders show up and how people talk about the mission and the goal. It's very much actually how we internally make decisions. It's a lot of walking the walk. I think Anthropic has been the most authentic company I've been in, from the first day I joined to now. I think it goes without saying, but the talent caliber and talent density has been incredible. It's been the most talent-dense environment I've gotten to work with. People are deeply taking radical ownership, deeply thoughtful, kind, but direct, and also in the service of making products, models, capabilities, research better.

日常角色的演变 Day-to-day role evolution

Host

我相信我们的听众会好奇:你日常的工作是什么样的?

I'm sure our listeners will be curious: what does your day-to-day look like?

Dianne Na Penn

我觉得每 3 个月我的工作就会变,因为几乎我们变成了一家不同的公司。我加入时大约有 150 人,两年前的这个时候,我正在搭建我们的第一个 A/B 测试,并给潜在客户发邮件让他们试用新模型。那是非常亲力亲为的。作为产品经理,你会不惜一切代价把事情做成。现在,随着我们有更多的产品经理、研究员和工程师,我大部分时间都在帮助团队。我的很多产品经理更深入地与他们的研究伙伴合作,所以指导和支持他们是我一天中的重要部分。大约三分之一的时间用于与客户通话,更好地了解所有不同的用例,特别是那些正在出现的、今天还不奏效但可能非常关键的东西。所以花很多时间与公司或初创公司在一起,然后思考未来的方向。

I think every 3 months my job changes, because it's almost like we're a different company. We were like 150 people when I joined, and two years ago around this time I was setting up our first A/B tests and emailing prospective customers to try new models. It was extremely hands-on. As a PM, you do whatever it takes to get the thing done. Now, as we have more product managers, more researchers, more engineers, a lot of my day is spent helping the team. A lot of my PMs are more embedded with their research counterparts, so coaching and supporting them is a big part of my day. Maybe a third of the time with calls with customers, just getting a better sense of all the different use cases, particularly what's emerging, what is not working today that could be really pivotal. So spending a lot of time with either businesses or startups, and then thinking about what's ahead.

Anthropic历程的关键决策点 Key decision points in Anthropic's journey

Host

显然,过去两年半真是疯狂。我想当你回顾这段旅程时,有没有一些关键决策点让你印象深刻,比如“哦,那真的改变了局面”或者相当关键?

Obviously, what a wild last two and a half years. I guess as you reflect back on that journey, are there some key decision points that stick out to you, like 'oh, that really tipped things' or was pretty pivotal?

Dianne Na Penn

我认为我们非常有意为之。早期,我想说在 2023 年,我们收到的最常见的用户请求是“你们需要有一个嵌入模型,因为我们在做 RAG,而 RAG 很流行。”我认为我们对于 LLM 和 AI 能做什么以及 Claude 能成为什么非常明确。我们专注于像投资智能体式编码这样的事情。这有点像用户引导与用户中心之间的区别,对吧?他们甚至没有要求的东西是什么?2023 年没有人来敲我们的门说“给我们智能体式编码”。所以我认为这是一个:早期非常清楚并专注于更大的机会。另一个我个人认为的关键时刻是我们选择在 API 上推出计算机使用功能,并且作为 beta 功能发布。我们知道它还不完全好用,但我们认为这对展示 AI 的能力非常有帮助,以及它只是一种不同的形式因素。我认为我们仍然在这条路上,但把它带给世界。当时有很多关于如何确保一切非常安全等的讨论。所以我们做了很多安全工作,但不可能覆盖所有边缘情况。我们不得不把它稍微放到世界上,然后随着能力扩展,弄清楚还需要做什么来确保安全。我认为那是一个非常真实的时刻,做出了一个我认为是正确的大胆决定。这可能是两件大事。还有很多其他有趣的时刻,但在我心中并不关键。

I think we were very intentional. In the early days, I would say in 2023, the most common user request we got was 'you guys need to have an embedding model because we're doing RAG and RAG is all the rage.' And I think we were very intentional about what LLMs and AI could do and what Claude could be. We focused on things like investing in agentic coding. This is a little piece of being user-led versus user-centric, right? What is the thing that they're not even asking for? Nobody's coming down our door in 2023 being like 'give us agentic coding.' So I think that's one: the early days of being very clear and focused around what is the bigger opportunity. I think another one personally is when we chose to ship computer use on the API, and we shipped it as a beta feature. We knew that it didn't quite work completely yet, but we thought that this was really helpful for showcasing what AI could do, and how it's just a different form factor. I think we're still on that journey, but bringing that to the world. There was a lot of discussions around how to make sure everything is very safe, etc. So we did a lot of safety work, but it's impossible to capture every edge case. We had to put it out into the world a bit to then figure out what else we need to make safe as the capabilities expand. I thought that was a very authentic moment, making a bold decision that I think was the right one. Those are probably two of the large things. There's a lot of other fun moments, but they're not pivotal in my mind.

Host

那么,有什么有趣的时刻让你印象深刻?

Well, what sticks out as a fun one?

Dianne Na Penn

我和我的团队做了 Golden Gate Claude,这是我们可解释性工作的一个分支。那非常酷,因为公司当时还不到 500 人,但已经开始感觉有点大了。我们在不到一天的时间里就把 Golden Gate Claude 从模型推到了 UI。

My team and I worked on Golden Gate Claude, which was an offshoot of our interpretability work. That was very cool because the company was still less than 500 people, but it was starting to feel a bit bigger. We shipped Golden Gate Claude from model to the UI within less than a day.

Host

你知道它会变得这么火吗?

Did you know it was going to be this kind of viral thing?

Dianne Na Penn

它在内部有点火了。所以我认为通常 Anthropic 的员工有很好的产品感和品味。所以我们想,“好吧,如果我们喜欢它,就试试吧。也许会有几百个人和我们一起狂热。”但我们不知道。是的,它非常草根。一些工程师、研究员、一个产品经理和一个设计师。

It went a little viral internally. So I think usually Anthropic employees have good product sense and taste. So we were like, 'Okay, if we like it, let's just try it. Maybe we get like a couple hundred people who geek out with us.' But we didn't know. Yeah, it was very grassroots. Some engineers, researchers, a PM, and a designer.

最自豪时刻:向用户展示模型 Proudest moment: showing model to users

Dianne Na Penn

我们当时想,好吧,我们来想办法如何向用户展示这个,而且我们必须在一周内完成,因为我们发表了一篇论文。那可能是最自豪的时刻之一。

We were like, okay, let's figure out how do we show this to users and we needed to do it in the week because we had published a paper. So that was probably one of the proudest moments.

快问快答:AI时间线观点转变 Quick fire: changed mind on AI timelines

Host

这是一场非常精彩的对话。我们总是喜欢在采访结束时进行快速问答,我会在时间用完前尽可能多地提问。那么首先,过去一年里你在 AI 方面改变了什么看法?

It has been a fascinating conversation. We always like to end our interviews with a quick fire around where I basically just stuff in as many questions as I can before we run out of time. So maybe to start, what's one thing you've changed your mind on in AI in the last year?

Dianne Na Penn

我实际上认为我们比年初预期的更接近变革性的长期运行 AI。感觉基础模块已经就位了。我现在比年初更强烈地感受到这一点。

I actually think we are closer to transformative long-running AI than I expected even starting the year. Like it actually feels like the building blocks are kind of there. I feel that more now than in the beginning of the year.

有效使用模型的建议 Advice for effective model use

Host

显然,你们的云 API 终端用户水平参差不齐。那些最老练的用户有没有什么做法,是你希望我们所有听众都能效仿,从而更有效地使用这些模型的?

You obviously have a range of sophistication of end users of the cloud API. Are there things that the most sophisticated folks do that you wish all of our listeners would do with these models to use them more effectively?

Dianne Na Penn

我认为大概有两件大事。我认为最大胆的构建者和用户不仅考虑当前有效的东西,还会拿出之前没成功的东西,或者一些不太完善的原型。他们仍然会组合起来,用新版本的模型进行测试。我们称之为原型设计,就像建立一个以用户为中心的原型库。我认为这非常有价值,因为很多时候这些系统并不是计划出来的。你几乎必须去发现能力。如果你没有原型,你就无法发现新东西。你总是会想,不知道它什么时候能真正擅长药物发现,但如果你没有某种方式每次去测试,那永远都太抽象了。所以我认为构建者拥有非常雄心勃勃的原型、产品创意或功能——即使过去可能不成功——但举办黑客马拉松这样的活动来实际测试这些想法非常重要。其次,当你有新技术版本时,愿意投入资源去改变产品体验,以顺应智能发展的趋势。

I think there are maybe two big things. I think the boldest builders and users are constantly thinking about not just what's working today, but actually might have things off the shelf that did not work before or some prototype that doesn't quite work. They still kind of put together and test with new versions of models that come out. So we call this kind of prototyping, like having a library of user-centric prototyping. I think that has been something really valuable to see, because a lot of times these systems are not planned. It's almost like you have to discover the capability. If you don't have a prototype, then you're not going to discover the thing. You're always going to wonder, oh, I wonder when it's going to get really good at drug discovery, but if you don't have some way to actually test that each time, it's always going to be too abstract. So I think builders having very ambitious prototypes, product ideas or features that might not have worked in the past, but just having things like hackathons where you could actually test these ideas is really important. And then I think being willing, when you have a new version of a technology, to invest in actually potentially changing your product experience to meet the intelligence tailwinds.

什么是模型品味? What is model taste?

Host

什么是模型品味?

What is model taste?

Dianne Na Penn

我认为模型品味就像产品直觉。产品品味是随着时间培养出来的。它实际上是一种不断磨练出来的对模型能力的感知,以及通过亲自动手、理解用户需求并持续迭代来发现这些能力的意愿。

I think of model taste as just like product sense. Product taste is developed over time. It's really a continuously honed sense of model capabilities and a willingness to discover the capabilities by being hands-on and also by understanding what users are trying to do and continuously iterating on that.

Host

你觉得模型品味是什么?

What do you think model taste is?

Dianne Na Penn

是的,我喜欢亲手实践模型的想法,对吧?就是培养一些直觉,了解它们能做什么、不能做什么,以及如何正确地推动它们或围绕它们构建脚手架,从而最大化利用它们。有些人,你给所有人同样的工具,他们似乎总能从中获得更多。

Yeah, no, I mean I love the idea of getting your fingernails dirty with the models, right? And just developing some intuitions to both what they can and can't do and also the right ways to push them or build scaffolding around them in order to get the most out of them. There's just some people that you give everyone the same tools and they always seem to just get way more out of them.

Dianne Na Penn

是的。我认为这是你愿意去尝试。创造性解决问题可能不是最准确的词,但那些愿意把模型当作新发现、不断尝试新事物的人,这就是培养模型品味的一种方式。

Yeah. I think it's your willingness to experiment. I think creative problem solving is not the perfect word, but people who are willing to approach models as a new discovery and constantly try new things, that's a way of developing model taste.

模型更新的框架与直觉 Scaffolding and intuition for model updates

Host

我还注意到,你提到人们能够在模型发布时尝试新事物。我想这不仅仅是接入 Opus 4.5 然后看它是否有效那么简单,因为显然需要构建很多脚手架。所以我认为这门艺术的一部分在于能够有一些直觉:我不是仅仅更换 API 密钥然后发现它不工作,而是实际上知道要调整哪些东西才能让它工作。

Well, I'm struck by too, you talked about folks having the ability to try new things as the models come out. I imagine it's not as simple as just plugging in Opus 4.5 and seeing if it works, because obviously there's a lot of scaffolding one has to build. So I imagine part of the art is also being able to have some intuition as to, I'm not just switching out the API key and it doesn't work, but I'm actually having some idea of things to tweak that might make it work.

Dianne Na Penn

这就是为什么我们内部会定期举办黑客马拉松。我们在 Anthropic 内部大概每三到四个月举办一次黑客马拉松,因为人们有很多积压的想法,比如“不知道 Claude 现在能不能做 X”,你需要给构建者、工程师、产品人员一个在日常工作之外进行创造性工作和探索的空间。因为确实,没有人会主动要求像智能体式编程这样的功能,它需要被一点点发现。它需要亲自动手才能弄清楚是否可能。给人们提供这样的空间非常重要。

This is why we actually internally have things like hackathons pretty regularly. I think we do a hackathon maybe every three or four months internally at Anthropic, just because people have these pent-up ideas of like, I wonder if Claude can now do X, and you really need to give builders, engineers, product people a place to do that creative work and do that discovery work outside of their day-to-day. Because yeah, nobody asks for a thing like agentic coding to work, and it has to be a bit discovered. It has to be a bit hands-on to figure out if something is possible. Giving people the place to do that is really important.

被低估的影响:安全的好处 Under-discussed implications: benefits of safety

Host

我想你谈到了长期运行的智能体,以及这些进展。你认为作为社会,我们有哪些方面讨论得不够,或者这些影响中哪些可能被低估了?

I guess you talked about obviously long-running agents, being closer and all these things. What's one thing you think of as a society we're not talking about enough, or one of the implications of all this that maybe is under-discussed?

Dianne Na Penn

我认为我们对安全的好处讨论得不够,不仅仅是“确保模型不做坏事”,更重要的是对齐模型的好处。其中一个活跃的研究领域是谄媚问题。对吧?模型只会说你想听的话。而实际上,一个对齐良好的安全模型恰恰相反。它实际上是一个独立思考者,对吧?而独立思考者正是我们取得突破和产生更好想法的方式。我认为我们没有充分讨论为什么安全实际上对更高价值的智能也有好处。它不仅仅是约束 AI,如果运作良好,它实际上是提升智能质量的一种方式。回到我一开始的例子,我问克劳德:“嘿,我们怎么考虑定价?这里有两个选项。”它提出了第三个非常好的选项,推动了我的思考。如果我们有一个非常谄媚的 Opus 4.5 版本,它可能就不会这么做。它只会同意我,对吧?我认为谄媚问题远未解决,但通过投资于对齐的 AI,我们实际上可以获得更好的智能。我觉得我们没有讨论这一点。

I think we don't talk enough about the benefits of safety, not just the fact of hey, it's to make sure the model doesn't do bad things, but more also like what is the benefits of an aligned model. One of the key issues that is an active research area is around things like sycophancy. Right? That models just tell you what you want to hear. And actually a well-aligned safe model, it's not, it's actually the opposite. So it's actually an independent thinker, right? And independent thinkers are actually how we have breakthroughs and how we have better ideas. And I don't think we really talk enough about why safety is actually really good for higher value intelligence too. It's not just to constrain AI. It's actually a way to amplify quality of intelligence if it works well. Coming back to my example in the beginning, I asked Claude, 'Hey, how do we think about pricing? Here's two options.' And it came up with a third that was really good that pushed my thinking. And if we had a very sycophantic version of Opus 4.5, it might not have done that. It would have just agreed with me, right? And I don't think sycophancy is solved by any means, but just by investing in making AI that's aligned, we could actually get to a better intelligence. I feel like we don't talk about that.

个人ASI时间线 Personal ASI timelines

Host

我喜欢这个例子。你知道,Anthropic 的人以非常激进的超级智能时间线而闻名。你属于那一类吗?或者你怎么看?

I love that example. You know, I think folks at Anthropic famously have very aggressive ASI timelines. Do you fit into that bucket or how do you think about that?

Dianne Na Penn

我会说我的时间线今年可能提前了,基于我在 Opus 4.5 这样的模型上看到的情况。

I would say my timelines have probably moved up this year based on things I'm seeing with models like Opus 4.5.

构建模块与产品过剩 Building blocks and product overhang

Dianne Na Penn

我觉得基础模块其实比我们想象的要更近。而且这更像是一种产品积压,或者说表达它的产品机会。

I feel like the building blocks are actually closer than we think. And that it's actually more of like a product overhang or product opportunities to express it.

Host

那么,你如何再次构建合适的支架,或者构建合适的方法来利用模型质量呢?

And how do you like again build the right scaffolds or build the right ways to harness like the model quality?

Dianne Na Penn

我更倾向于杠铃策略。我认为短期内概率很大,长期内概率也很大。

I'm more of a barbell. I think there's a very large probability in short term and then a very large probability in long term.

结语与更多学习资源 Closing remarks and where to learn more

Host

嗯,这真是太棒了。我觉得如果再留你太久,我会推迟未来一代模型的发布。但我想确保把最后一句话留给你。人们可以去哪里了解更多关于 Opus 45、关于你,或者任何你想指引他们去的地方?话筒交给你了。

Well, this has been fascinating. I feel like I will delay the launch of the future generation of models if I keep you any too much longer here, but I just want to make sure to leave the last word to you. Where can folks go to learn more about Opus 45, about you, anywhere you'd like to point them? The mic is yours.

Dianne Na Penn

是的。我觉得是我们的博客和网站。我认为,你知道,我们相当——我们喜欢让作品自己说话。所以,关注我们的网站是正确的选择。

Yeah. I think our blog and our website. I think that, you know, we are pretty like, we like to let the work speak for itself. So, following on our website is the right place.

Host

是的。我猜你们以不发布“明天发布”或模型发布前的神秘推文而闻名,而其他公司似乎都这么做。

Yeah. I guess you guys famously don't do the like coming tomorrow or the cryptic tweets before model launches that everyone else seems to do.

Dianne Na Penn

不。

No.

Host

嗯,太棒了。这真是太棒了。非常感谢你的时间。

Well, awesome. This has been fascinating. Thanks so much for the time.

Dianne Na Penn

非常感谢。聊得很愉快。

Thank you so much. It was great chatting.

互动版:逐字朗读 + 针对本期提问 →