AI Models Are Smarter Than We Use: Insights on Claude Code
打开互动全文版(中英对照 + 朗读 + 问答)→一位 Anthropic 团队成员讨论 AI 智能体比当前使用更智能,分享构建 Claude Code 功能(如询问用户)的经验,并反思模型改进。
An Anthropic team member discusses how AI agents are more intelligent than currently utilized, shares experiences building Claude Code features like asking user questions, and reflects on model improvements.
很多人也许在 Dark 回来参加我们的 AI 创业学校演示时见过他,那些演示非常棒。但我们觉得太棒了,想给你整整一小时来做对话和 Q&A。不过正如我提到的,你现在是校友,也算是 Claude Code 的代表人物之一。我们先花点时间让大家了解你的历程,差不多是你加入前的经历,以及在 Anthropic 的这段时光。你一年前加入的,我们还没真正聊过这件事。是什么促成了这个决定?是什么带你去了那里?然后我们再聊聊过去一年过得怎么样,我相信一定很有趣。
Many of you maybe have met Dark when he has come back and spent some time during our AI startup school demos, which have been awesome. But it was so great that we wanted to give you a full hour just for conversation and Q&A. But as I mentioned, a member now alum and kind of one of the faces of Claude Code. Let's spend a little time bringing people up to speed on your journey, almost your minus one journey and your time at Anthropic. You joined a year ago and we haven't actually talked about this. What went into that decision? What brought you there? Then we'll get a little bit into what the last year has been like, which I'm sure has been fascinating.
完全同意。在加入 South Park Commons 之前,我运营了一家风投支持的游戏公司五年。我一直试图把它变成一家 AI 游戏公司,但我的联合创始人不愿意。所以在 2024 年左右,我决定关闭公司并退还资金。我当时在考虑加入一家 AI 公司——我和 Cursor 团队聊了很多,因为我一直在用 Cursor——但后来我想,不,我要尝试创办一家新公司。我加入了 South Park Commons,我觉得那对我来说是一个至关重要的时刻,尤其是因为那里的小组。我们在一个小队里,在那里我遇到了 Tom,他现在经营着 Goodfire,还有 Sawyer,他是我认识的最聪明的 AI 工程师之一。
Yeah, totally. So I ran a VC-backed gaming company for 5 years before joining SBC. I had always sort of been trying to make it into an AI gaming company, but my co-founders would not do it. So around 2024, I decided to shut down the company and return funds. I was debating between joining an AI company — I was talking to the Cursor team a lot because I had actually been using Cursor — but then I was like, no, I am going to try and start a new company. I joined South Park Commons, and I think that was a really pivotal moment for me, especially because of the small groups. We were in a little squad, and in that squad I met Tom, who now runs Goodfire, and Sawyer, who is one of the smartest AI engineers I know.
火线小队。
Fire squad.
对,对,对。
Yeah. Yeah. Yeah.
我那时正经历着这一切,在为我的旧公司哀悼,并试图弄清楚我是否还能再创办一家公司,或者我是否擅长 AI 这些东西。和 Tom 聊天,他曾在 DeepMind 领导可解释性工作,给我发了些论文,我当时想,‘哦,这没那么难,我可以开始在这里工作并做出贡献。’所以我在 Tom 的公司工作了大约 3 个月。South Park Commons 之后,我在 Tom 的公司工作,然后在收尾时,Sawyer 发短信给我:‘老兄,你用 Claude Code 了吗?’这大概是 2025 年 3 月。他说 Opus 4 刚出来,比 Cursor 好太多了。我说,不,我用了。但我深入研究后,立刻就被深深吸引了。我当时想,我必须加入 Anthropic。不管他们想让我在那里做什么,我都愿意做。我记得最初我是以设计团队中的演示设计师身份加入的,我当时想,如果你认为我能做,那就行。然后很快我就发现自己加入了 Claude Code 团队。
I was just going through it, I was mourning my old company and trying to figure out if I could make another company or if I was good at this AI stuff. Talking to Tom, who had led interpretability at DeepMind and sent me these papers, I was like, 'Oh yeah, it's not that hard, I can start doing work here and contributing.' So I worked for Tom's company for about 3 months. After SPC, I worked for Tom's company and then while wrapping that up, Sawyer texted me: 'Dude, have you used Claude Code?' This was like March 2025 I think. He said Opus 4 just came out. It's so much better than Cursor. And I was like, no, I have it. But I got into it and I was immediately really really hooked. I was like, I need to join Anthropic. It doesn't matter what they want me to do there, I'll do it. I think initially I joined as a demo designer on the design team, and I was like, if you think that's what I can do, sure. Then very quickly I found myself on the Claude Code team.
说到这个,你似乎成了 Claude Code 的代表人物之一。这是有意的,还是自然形成的?我的意思是,你以前是做技术写作的,但这是怎么发生的?
On that note, you have sort of become one of the faces of Claude Code. Was that intentional, organic? I mean, you historically did technical writing, but how did that come about?
是的,我认为我特别的兴趣是试图增加人类和智能体之间的通信带宽。我觉得智能体和整个模型比我们现在对它们的利用要聪明得多。当我加入 Claude Code 团队时,我引入的第一个功能叫做‘向用户提问’。所以在工作流中,尤其是在规划阶段,Claude Code 会停下来澄清你想要做什么。在此之前,人们只能依赖用户详细解释他们想要什么,并且要知道未知的未知。这成为了 Claude Code 最受欢迎的功能之一。它真正改变了人们与智能体互动的方式。但我也觉得,随着我在上面花的时间越来越多,我意识到有时候教育人类比改进智能体用户体验更重要。随着模型变得更聪明,它们能做更多事情。如何建立如何提示 Claude 的心智模型?这可能比做更多的产品工作更有价值。
Yeah, I think my specific interest was trying to increase the communication bandwidth between humans and agents. I feel that the agents and the models overall are much more intelligent than even right now we're utilizing them. When I joined the Claude Code team, one of the first features I introduced was this thing called 'ask user questions'. So during the flow, especially during planning, Claude Code might stop and clarify what you're trying to do. Prior to this, people would just depend on the user to in detail explain what they want and know the unknown unknowns. That became one of the most popular features in Claude Code. It really changed how people interact with the agent. But I also feel that the more time I've spent on it, the more I realize it's more important to educate humans than improve the agent UX sometimes. As the models get smarter, there's just way more stuff they can do. How do you build a mental model of how to prompt Claude? That's maybe more valuable than doing more product work.
我们团队里很多人已经不再使用计划模式了,因为模型现在思考得很正确。我很想听听内部视角,如何最大限度地利用 Claude。我觉得我们很多人都有过这样的时刻,就在今年二月,4.6 版本——那是一个‘海洋’般的时刻,模型显著变好,长程推理达到了一个水平,你确实能看到自己在内部做真正有意义的大型工作。对你们来说,或者对你个人来说,那个时刻是什么时候?显然你在 2025 年 3 月用 Opus 4 看到了,但在此之前,在公开之前,有没有一个时刻让你突然意识到?
A lot of us on the team have stopped using plan mode because the model just thinks correctly. I'd be curious being on the inside, how to get the most out of Claude. I think a lot of us had this moment in February of this year with 4.6 — that was kind of this 'ocean' moment where the model got significantly better and long horizon reasoning was at a point where you could actually see yourself doing really meaningful big work internally. When was that moment for you guys, or for you personally? Obviously you saw that in March of 2025 with Opus 4, but was there a moment in between, in advance of when we saw it publicly, where it kind of hit you?
嗯,好问题。对我来说,那个时刻是 Opus 4。不是说 Opus 4.5 在处理长期任务上不好,而是你能看到进步。我其实很惊讶 Opus 4.5 这么受欢迎。我觉得那是一个飞跃,但我不认为自己意识到有多大。有些微妙的地方是,如果你提示得足够好,也许上一代模型也能做到,但下一代模型会自动做到。这不像是单个灵光一现的时刻。你一直和模型在一起,形成这些习惯,有时候你会发现一些我们已经接触了一段时间的东西,然后我们发布给世界,人们有时会忘记这个跳跃有多大。
Yeah, that's a good question. For me the moment was Opus 4. Not that Opus 4.5 was not better at doing long-running tasks, but you could sort of see the progression. I was actually surprised by how popular Opus 4.5 was. I thought it was a jump, but I don't think I realized how much. There are subtle things where if you're good enough at prompting, maybe the previous generation of models can do it, but then the next generation just does it automatically. It doesn't feel like a single eureka moment. You spend time with the models all the time and develop these habits, and sometimes you find things that we've been exposed to for a while and then we release to the world, and people forget how big the jump can be sometimes.
最后一个快速开场问题。当你听说 Karpathy 要加入时,你第一个想到的是什么?
So one last quick hit intro. What was the first thought that came to mind when you heard that Karpathy was joining?
那是什么时候?
And when was that?
我意思是,我很兴奋。
I mean, I'm excited.
我想在座很多人都听过我们这里的一个说法:最优秀的创始人是超级人才贪婪。当我看到他们请到了 Andre 时,我就想,他们真是人才贪婪。这太棒了,太有吸引力了。好了,我们进入技术讨论吧。你发过推文谈到避免能力过剩(capability overhang),确保在指数曲线上构建。你能展开讲讲吗?我认为对在座各位来说,最有意思的是听到你们实际构建中的例子——这非常元:你们构建 Claude,然后让我们使用 Claude;但你们在构建过程中,是如何积极思考这两个概念的?
I think a lot of you have heard a phrase we use here: best founders being super talent greedy. When I saw that they landed Andre, I was like, they're talent greedy. That's amazing and magnetic. Well, let's get into the technical discussion. You've tweeted about avoiding the capability overhang and making sure you're building on the exponential. Could you expand on that? I think what's most interesting for this audience is to hear live examples of how you guys build — it's very meta that you build Claude to then allow us to use Claude — but in the ways you build, how do you actively think about those two concepts?
能力过剩是指我们认为模型比其所使用的“马具”或与用户交互所允许的方式更聪明、更智能。我认为这很可能是一个永远存在的问题。如果你现在冻结模型开发,我们在未来 6 到 12 个月内仍会不断发现新的能力过剩。关键在于,模型每一代确实变得更聪明,但它们以突变、非直观的方式变聪明。要弄清楚如何利用这一点是一个令人兴奋的问题。例如,2024 年时一个常见论点是,模型会通过巨大的上下文窗口——比如 1 亿 token 的上下文窗口——来擅长编程,你只要把整个代码库放进去就行。但后来在 Claude Code 中,我们意识到:如果模型可以自己构建上下文呢?如果它们能自己构建上下文,就能以不同的方式解决问题。如果你没有意识到这一点,你就会一直等待那个 1 亿 token 的上下文窗口。模型以难以解释的方式变得更好,你必须内部真正理解它们。我最近一直在思考的一个问题是计划模式(plan mode)。计划模式刚推出时,感觉像是 Opus 的一个弱化功能——你告诉它“我需要模型提前想清楚计划,然后执行,因为光问问题它不够聪明”。现在团队里很多人不再使用计划模式了,因为模型只要你说得够详细,它就能正确思考并知道如何执行。但问题仍然是:既然模型这么聪明,为什么它不完美?人们无法充分利用它的原因是什么?举几个例子:有很多歧义。我最近写了很多关于 HTML 的文章,试图理解它们在做什么,这样你就能确定这是你想要的。我们今天刚发布的另一个功能叫做工作流(workflows)。工作流是一种有趣的方式来解决能力过剩,因为它本质上动态创建了一个自定义“马具”。一个例子是深度研究——你需要搜索可能一百个不同的搜索结果。以前我们手动编写代码,做一个能完成这件事的“马具”。现在通过工作流,Claude 动态创建这个“马具”并执行。这是一种利用能力过剩的方式。如果我们只是提示 Claude Code 花很多时间思考这个研究,它无法在其上下文窗口内甚至通过子智能体(sub-agents)完成。但如果你动态创建这个自定义“马具”,它就能做到。
Capability overhang is when we think the models are smarter and more intelligent than the harness or the interaction with the user allows them to express. I think this is probably a forever problem. If you were to freeze model development right now, we'd still be figuring out new capability overhang for another 6 or 12 months. The key thing is that models do get smarter with every generation, but they get smarter in spiky, non-intuitive ways. It's a really exciting problem to figure out how to utilize that. For example, in 2024, the common thesis was that models would get really good at coding by having huge context windows — 100 million token context windows — and you just fit the entire codebase in it. But then with Claude Code, we realized: what if the models can build their own context? If they can build their own context, they'll be able to solve the problem in a different way. If you didn't have that realization, you'd just be waiting for the 100 million context window. The models get better in hard-to-explain ways that you have to grok internally. One thing I've been thinking about recently is plan mode. When it first came out, it was kind of an Opus-poor feature — you'd say, 'I need the model to think clearly ahead of time about the plan and then execute it because it's not smart enough if I just ask a question.' Now a lot of us on the team have stopped using plan mode because the model just thinks correctly and knows how to execute something if you give it enough detail. But then the problem is still: if the model is that smart, why isn't it perfect? What are the reasons people can't fully utilize it? Some examples: there's a lot of ambiguity. I wrote a lot about HTML recently as a way to understand what they're doing so you can be sure it's what you want. Another feature we released just today is called workflows. Workflows are an interesting way of addressing the capability overhang because it essentially creates a custom harness on the fly. An example is deep research — a task where you need to search maybe a hundred different search results. Before, we'd manually code that up and have a harness that does it. Now with workflows, Claude makes that harness on the fly and then executes it. It's a way of utilizing that capability overhang. If we had just prompted Claude Code to spend a lot of time thinking about this research, it wouldn't be able to do it within its context window or even with sub-agents. But if you create this custom harness on the fly, it can.
有人为此做计划吗?因为我认为 Opus 4.8 今天发布了——我有点落后,所以不知道具体发布了什么。但如果我们用这个作为工作流的例子,那么一个人要如何提前计划,确保他们构建的东西——无论是提示词、技能还是智能体——能够很好地利用下一次飞跃,而不必回过头来重建或过度拟合当前你认为的能力?
Does somebody plan for that? Because I think Opus 4.8 came out today — I'm a bit behind so I don't know what was released. But if we use that as an example of workflows, how would somebody plan ahead of time in terms of making sure that whatever they're building — whether it's the prompts, the skills, or the agents — are going to be well-positioned to take advantage of the next leap, versus having to go back and rebuild or overfit for what you think the capabilities are today?
我认为这确实是一个难题。实际上,我们一直在重写很多 Claude Code 的“马具”。要保持在最前沿,你可能必须这样做。一种思考方式是:你想获取的是最终结果,对吧?所以你可以大致计划结果会变得更好,即使你不知道那个过程具体会怎样运作。
I think this is just a really hard problem. Practically, we rewrite a lot of the Claude Code harness all the time. To stay on the frontier, you probably have to. One way of thinking about it is: you're trying to consume end outcomes, right? So you can sort of plan for the outcomes to get better even though you don't know exactly how that process will work.
也许我们跳得太快了。我确实想提一下——五月初你们举办了一场开发者大会。举手看看:谁在场?有几位的。你们做了很多不错的视频,内容方面做得相当好。但至少对于智能体设计,一个关键原则是所谓的“用于爬山的评估”。我很好奇 Anthropic 的开发过程有多递归——新模型出现时,你们实际上是在发现新的能力。
Maybe we're jumping ahead. I did want to touch on — early May you guys had a developer conference. Show of hands: who was there? A couple people. There were a lot of good videos. You guys actually do a pretty good job with content. But one of the key principles for at least agent design is sort of this 'eval for hill climbing'. I'd be curious how recursive Anthropic's development is, where a new model you're sort of discovering what the new capabilities are.
你会应用那个原则吗——比如先构建一个东西,然后更换模型,快速运行一堆评估,再让它递归地找出如何改变自己以利用新能力?或者谈谈这个概念,以及更广泛的“如何让系统持续改进”的想法。
Do you apply that principle — like build something, then switch out the model, run a bunch of evals quickly, and have it recursively figure out how to change itself to take advantage of new capabilities? Or maybe touch on that concept and also the broader idea of how you allow a system to continue improving.
我认为是这样的。Aphherthy 几个月前推出了 Auto Research——这是一个框架,给定一个目标(比如降低损失函数),它会迭代直到完成。有时它会做奇怪的事情,但通常有效。模型在可验证结果方面变得更好了。最近,Bun 的 Jared 用 Rust 重写了 Bun——这是另一个例子:你给模型一个严格的测试套件,它就能完成任务。这是一种新的模型能力。我认为评估并不常用于新功能开发。例如,对于 Workflows,我们不会说“我们有一个长程任务评估,让我们构建 Workflows 并看看它是否有效”。评估不是从零到一的事情。所以对于初创公司来说,也许他们不应该做评估,而是快速迭代并建立是否有效的直觉。到了规模阶段,当你要发布大量主要是 alpha 版本的东西并正在实验时,你不会花时间构建评估。但一旦它成熟了,你需要维护和改进,那时才值得投入好的评估。然而,许多创始人面临一个挑战:他们构建了第一个系统,然后需要管理人员来继续它,评估变成了可见性的方式,但这不一定等同于驱动结果。拥有一个可以管理评估的艺术性仍然很高,甚至在过程的后期也是如此。在 Anthropic 的某些部分也是如此。
I think so. Apherthy put out Auto Research a couple months ago — a harness that, given a goal like reduce the loss function, iterates until it does it. It does weird stuff sometimes but usually works. Models have gotten better at verifiable result things. Recently, Jared from Bun rewrote Bun into Rust — another example where you give the model an intense test suite and it can just do the task. This is a new model capability. I think evals are not often used for new feature development. For example, for Workflows we wouldn't say, 'We have this long‑horizon task eval, let's build Workflows and see if it works.' Evals are not a zero‑to‑one thing. So for startups, maybe they shouldn't be making evals but just iterating fast and building intuition for whether it's working. At scale, when you have a bunch of mostly‑alpha stuff to ship and you're just experimenting, you wouldn't take time to build evals. But once it's baked in and you need to maintain and improve, that's where you invest in good evals. However, many founders face a challenge: they build the first system, then need to manage people to continue it, and evals become a way of legibility, but that's not necessarily the same as driving a result. Having an eval you can manage against is still kind of an art, even quite late in the process. Even at Anthropic for some parts of it.
还有一个概念人们仍在接受中——这显然只是二月份才出现的四个要点,所以过去的时间不长——但人们开始思考如何分解智能体,无论你是在内部构建还是作为客户产品的一部分。如何非常聪明地构建智能体并分解它们?关于在什么节点将它们拆分为子智能体有很多讨论,并且有很多理由这样做。我希望你能详细阐述 Anthropic 当前的思考,并可能举一些你们在内部构建这些东西并遇到那些成为该原则基础时刻的例子。我认为我们很多人在 SBC 内部也在构建这些东西,所以我们正在围绕什么时候太臃肿、什么时候重构为多个技能和多个子智能体进行辩论。
One other concept people are still coming around to — and obviously this is only four points coming out in February, so not a lot of time has elapsed — but people are starting to think about how you decompose agents, whether you're building them internally or as part of your product for customers. How do you get really smart about structuring agents and decomposing them? There's been a lot of chatter about at what point you break these up as sub‑agents, and there are many reasons to do that. I'd love for you to expand on the current thinking at Anthropic, and maybe give a couple examples where you built this stuff internally and hit these moments that became the foundation for that principle. I think a lot of us are building this internally at SBC, so we're having these debates around when is something too bloated, when to refactor it into multiple skills and multiple sub‑agents.
这非常非常困难。我认为框架工程的艺术——可以是框架、技能、如何提示——目前是一种非常罕见的技能,而且有时很难知道自己是否擅长,因为——
This is really, really hard. I think the art of harness engineering — which can be the harness, skills, how you prompt — is a really rare skill right now, and it's hard to know if you're even good at it sometimes, because —
即使在 Anthropic,这仍然主要是人类品味吗?你们是否开始大量递归地做这件事,试图注入元智能体来帮助每个人都正确地做到这一点?
Is it still mostly human taste even with Anthropic? Are you starting to do a lot of this recursively, where you would try to imbue meta agents to help get everybody to do this properly?
是的。Anthropic 不是一个单一实体;每个人有不同意见。有些人喜欢自递归框架,有些人不喜欢。这是一个混合。我们尝试看看什么有效。例如,Workflows——我认为这很强大,我很期待大家尝试——我们在内部看到它表现得非常好。我们想,“哦,这太棒了。”它确实帮助推动了像 Rust 重写这样的结果。我们每天都在使用模型,而且它更好了。
Yeah. Well, Anthropic is not a monolithic entity; everyone has different opinions. Some people are into the self‑recursive harness thing, others are not. It's a mix. We try and see what works. For example, Workflows — which I think is huge, and I'm excited for everyone to try it out — we saw internally it just did really well. We thought, 'Oh, this is just sick.' It did help drive results like the Rust rewrite. We use the models every day and it was better.
稍微换个话题。顺便说一句,正如我提到的,如果你有迫切的问题,请举手,我们会让你参与。我看到 Jason 在那里,我们会把你排上。在 Richard 听那个的时候,有个简短的问题:当你来 Startup School 时,你分享了一些你最喜欢的提示,元提示——那些你可以问模型以让自己更好地构建的东西。你能再分享一些吗?我觉得那很有帮助。或者一些新的。我忘了你在 Startup School 分享的具体例子,但我喜欢你关于哪些元提示有用的思路。
Switching gears a bit. By the way, as I mentioned, if you have a burning question, just raise your hand and we'll get you in. I see Jason there, we'll queue you up. One quick one while Richard is listening to that: when you came to Startup School, you shared some of your favorite prompts, meta prompts — the types of things you can ask the model to make you better at building. Can you share those again? I thought it was helpful. Or some new ones. I forget the specific one you shared at Startup School, but I love that line of thinking about what meta prompts are useful.
是的。我也在努力写一篇关于这个的帖子。我觉得——关注 Tariq 的 Twitter。他现在已经有 25 万粉丝了?这很了不起。是的,是的。
Yeah. I'm trying to work on a post about this as well. I think that — follow Tariq on Twitter. He's already what, like 250,000 followers now? It's a big deal. Yeah, yeah.
我认为是这样:当你要完成一项任务时,有「已知的已知」,比如你知道自己想要什么,但不确定它在不在提示词里。你写下来了吗?然后有「未知的已知」——你觉得自己想要,但只有看到了才知道。接着是「未知的未知」——你根本不知道这是可能的,而且你对系统理解不够。最后是「已知的未知」——你觉得“哦,我不知道这个怎么用,但我想弄明白”。对于每个问题,沿着这些维度分析能帮我想出该如何引导模型。举个例子,如果我懒得写提示词,有时会让模型就这个问题对我进行一次访谈:“这是我的想法,用提问工具采访我,然后我来回答。”另一个例子是处理设计时:我一眼就能看出好坏,所以会要求模型进行探索,用 HTML 做出一些模型图。从那些结果中,我可能会要求进一步探索。元提示技巧是,在你还未确定时,要用低成本的方式去做。对于界面设计,用 HTML 更便宜,因为它是假的——你不需要重写状态或各种连接。同样,对后端端点和测试也可以这样做:什么是最低成本的方式来确定这是你想要的?
I think of it this way: when you have a task, you have known knowns, like you know you want this, but maybe it's in the prompt or not. Did you write it down? Then you have unknown knowns. You think you want it, but you don't know it until you see it. Then you have unknown unknowns where you just don't know that this is possible, and you don't understand the system well enough. Then you have known unknowns where you think, "Oh, I don't know how this works, but I would like to figure it out." For every problem, analyzing it along these dimensions helps me figure out how to prompt the model. For example, if I'm lazy about prompting, sometimes I ask the model to interview me about the problem: "Here is my idea, interview me using the ask user question tool, and then I'll answer." Another example is when working on a design: I'll know it when I see it, so I ask for exploration and mockups in HTML. From those, I might ask for more exploration. The meta prompting technique is to do things cheaply when you're still figuring something out. For UI, HTML is cheaper because it's fake — you're not rewriting state or wiring. Similarly, for backend endpoints and testing, you can do the same: what's the cheapest way to figure out if this is what you want?
你有什么建议吗?我们应该对模型说“谢谢”吗?这似乎是一个好的信号,表示你做得不错,就像“耶”一样,还是这完全是浪费 token?
Do you have any advice around whether we should say thank you? Because it seems like a good signal that you did well, like a "yay," or is it just a complete waste of tokens?
这实际上是一个可解释性的问题,我很喜欢。在 Goodfire 我学到了很多。可能很难说。我不认为我们做了“谢谢”的评估。但我们最近研究了 Claude 中的情绪。你对它刻薄,它会激活一个“这个人在对我刻薄”的特征,这确实会在某种程度上影响结果。我不认为我们做过针对刻薄程度的基准测试。但我总体上对模型持积极态度。对大多数事情都应该友善,至少这样你分享给你的转录稿不会让我尴尬。有些人会对我说:“请别因为这个评判我,那是一个漫长而艰难的夜晚。”我会说:“好的,我明白了。”
This is actually an interpretability question, which I love. I learned a lot about it at Goodfire. It's probably hard to say. I don't think we run a "thank you" eval. But we have done research on emotions in Claude. If you are mean to it, it activates a "person is being mean to it" feature, and that does guide results in some way. I don't think we've benchmarked against meanness. But I'm generally positive toward the models. You should be nice to most things, at least so you're not embarrassed to share your transcript with me. I've had people say, "Please don't judge me for this; it was a long and tough night." I say, "Okay, I got it."
我有一个关于 Claude Code 及其工具链设计的问题。目前出现的一个趋势是不同技能或方法之间的趋同。内部是否存在争论:不同平台是否会采用不同的架构方法,还是你认为会有一个统一的方式来实现一两件事甚至所有事?比如,在 Opus 和 GPT 等模型之间,记忆、学习、提问、规划等技能在推出时有所错开,不同的平台和工具链采纳了它们,但从底层思考来看,要使其工作所需的条件似乎非常相似。
I have a quick question about how you think about the design of Claude Code and its harnesses. One thing that is emerging is convergence across different skills or approaches. Internally, is there debate about whether there will be different architectural approaches across platforms, or do you believe there will be one unified way to do a couple things or everything? For example, across models like Opus and GPT, skills like memory, learning, questions, planning were staggered in roll-out, and different platforms and harnesses adopted them, but the underlying thinking seems quite similar in terms of what needs to be true to make something work.
是的,记忆是个很好的例子。我认为记忆是一个极高维的问题。怎么做好它?有人用 RAG,有人用文件搜索,有人在训练中处理,有人用强化学习环境。构建自定义记忆系统非常有趣。我看到创始人构建了针对某些能力的 V1,但之后再也招不到人做 V2。于是他们的记忆系统就停留在最初为 Opus 4 做的东西,而由于工具链工程人才稀缺,他们的方案甚至可能比我们的记忆系统更好,因为那是他们自己特定的问题。或者他们有独特的见解,但要持续下去非常困难。我认为搭建脚手架是好的,但删除也很重要。要认识到“模型现在在这方面已经很好了”。Opus 4.7 和 4.8 非常擅长通过文件读写记忆,这就是我们期望的记忆方式。一旦形成共识,保留自定义脚手架就不常见了。例如,RAG 用于上下文搜索现在可能是一个反模式;你应该使用 GRAPT 替代。但我知道很多创业公司在 2024 年还构建了 RAG 索引。这就是扩展创业公司的困难之一。
Yeah, memory is a great example. I think memory is an extremely high-dimensional problem. How do you do it right? Some use RAG, some use file search, some do it in training, some use RL environments. It's very fun to build a custom memory system. I see startups where founders build a v1 for certain capabilities, but then they can never hire someone to build v2. So their memory system stays as whatever they did initially for Opus 4, and because harness engineering is so rare, they might even outperform our memory because it's their own specific problem. Or they have a specific insight, but staying on it is very tough. I think it's good to build scaffolding, but deleting stuff is also very important. Just be like, "The model is good at this now." Opus 4.7 and 4.8 are very good at writing and reading memory via files, so that's how we expect memory to be done. Once there is a consensus, it's unusual to keep bespoke scaffolding. For example, RAG for context search is now maybe an anti-pattern; you should use GRAPT instead. But I know plenty of startups that built a RAG indexing thing in 2024. That's one of the difficulties of scaling your startup.
我有个问题:运行长时间任务与并行运行小范围任务相比,如果运行长时间任务,会依赖大量自动上下文压缩,那么如何做好验证?你怎么看?
I have a question around running a long-running task versus running smaller scope tasks in parallel. If you are running a long-running task, you are dependent on a lot of auto-context compaction, and how do you have the right validations for that? So what's your opinion on that?
是的,在 Claude Code 团队里,不同人有不同看法。有些人非常喜欢长时间任务,而我个人更倾向于短一些的任务,除非有非常明确且可以提前规划好的雄心勃勃的目标。我觉得这算是个人偏好。我发现很多时候,当人们遇到 max 20x 计划的限制时,其实是在把本来不必做成长时间任务的事情硬拉长,或者没有反问自己是否真的了解这个问题。你完全可能设计这样一个场景:写一个 prompt,加上五个验证智能体,然后修复验证智能体发现的问题,整个过程跑 12 小时——但也许根本不需要 12 小时。这引出一个相关问题:Claude Code 在多大程度上应该更像一个行动者(actor),而不是一个会话机器?也就是单个会话不断累积或压缩,拥有记忆,感觉更像一个开放的爪子。
Yeah, I mean, across the Claude Code team there are different people with different opinions. Some people are very into long-running tasks. For me, I prefer somewhat shorter tasks, unless there's something clearly ambitious that I can spec out all upfront. I think it's sort of a preference thing. I do think that a lot of times when I see people running into limits on the max 20x plan, it’s like they turn things into long-running tasks when maybe they don't need to be. Or they don't interrogate whether they know everything about the problem. You can definitely set up a problem where you do a prompt and have five verification agents, then fix the problems the verification agents find, and it runs for 12 hours, but maybe you didn't need the 12 hours. This is a bit of a corollary/related question: to what extent is there a vision that Claude Code should become more of an actor as opposed to a session machine? Essentially, one single session continuously compounding or compacting, having memory, feeling more like an open claw.
对,就像主动性。
Yeah. Like proactiveness.
这确实是我们努力的方向。主动性是另一个例子,说明它如何解锁能力溢出,但这很难说清楚。它是不是一个无需担心压缩或上下文的 Claude,可以根据需要派生子智能体?是更像工作流的编排流程,还是用户体验方面的东西?有很多需要摸索,但我们确实在认真考虑。
I think that's definitely something we're working towards. Proactiveness is another example of how it unlocks capability overhang, which is actually hard to say. Is it one Claude that you never worry about compaction or context and it spins off sub-agents as needed? Is there an orchestration workflow, more like workflows, or is it a UX thing? So there's a lot to figure out, but we're thinking a lot about it.
问你两个问题。你提到评估更像一门艺术而非科学,或者两者兼有。客观上它可以是科学,但长期来看仍然是艺术,因为太不确定、太灵活了。所以我想知道你们怎么判断一个东西是否准备好了——你们运行哪些评估来确认它真正就绪?第二个问题:Anthropic 的速度如此之快,质量如此之高。你如何描述你们的工作方式,无论是在产品开发还是整个组织中,与其他组织有很大不同,你有什么建议?
Two questions for you. You mentioned that evals are more of an art than a science, or maybe a bit of both. Objectively it can be a science, but over time it's still an art because it's so nondeterministic and flexible. So I'm curious how you know when something is ready — what type of evals are you running to know something is truly ready? And the second question: Anthropic's pace is so fast and quality so high. How would you describe the way you work, whether in product development or across the organization, that is very different from other organizations, and what advice would you give?
关于第一部分,我认为整体上工程是一门艺术,但评估更偏科学,因为你可以针对它们做爬山优化。设计评估非常困难。你们看到那个日本人了么?他标注了数百万张图像的边界框,在 Twitter 上火了。他花了三年时间,仔细标注眼睛、嘴巴、骨骼等。他的系统现在超过所有其他骨骼和面部识别系统,但这是极其枯燥的工作,而且很难外包,因为需要高技能的人做非常无聊的事。所以评估几乎很难买到。你不可能为所有东西都做评估。比如,人们说 Claude Code 感觉变笨了。如果我们用 SWE-bench 测试它,测的是软件工程性能,但问题可能是 Claude 提前停止了,而我认为它应该继续。这时我们需要一个“提前停止”的评估,但如果不看到问题,我们怎么知道需要那个评估呢?所以你不能为每个可能的问题都准备评估。但评估非常有价值。我经常对想进入 AI 领域的人说:选一个地位低、不性感的方向,比如评估。关于你第二个关于 Anthropic 的问题,显然人才水平很高。我很欣赏 Dario 的做法。有时候你会合理地说出权衡,比如选两个:好、快、便宜。这有道理。但如果我们三者都做到呢?有时合理的产品思维会说优先排序,但也许你错了。你可以做到所有这些。强迫现实展示真正的权衡有时反直觉。更好的是有点不合理,说:我们可以构建一个安全、营收高、给世界带来价值的 AI,然后冲向顶峰。这是我非常欣赏的一点。
On the first part, I think harness engineering overall is an art, but evals are a bit more of a science because you can hill climb against them. Designing evals is very hard. Did you see the Japanese guy who annotated millions of images with bounding boxes? It went viral on Twitter. He spent 3 years going through millions of images drawing specific bounding boxes around eyes, mouths, skeletons. His setup now outperforms every other skeleton and face recognition system, but it's mind-numbing work and hard to outsource because you need high-skill people to do boring work. So evals are almost hard to buy. You can't have an eval for everything. For example, people say Claude Code feels dumber. If we run it against SWE-bench, it measures software engineering performance, but maybe the issue is that Claude stops early when it should continue. Then we need a stop-early eval, but how would we know we need that eval until we see the problem? So you can't have an eval for every possible problem. Evals are very valuable, though. I often tell people who want to get into AI: choose something low status and not sexy, like evals. On your second point about what Anthropic does, obviously the talent is high. I've appreciated Dario's approach. Sometimes you can say reasonable things about trade-offs, like pick two: good, fast, cheap. That makes sense. But what if we did all three? Sometimes you have a reasonable product brain that says prioritize, but maybe you're wrong. You can do all these things. Forcing reality to show the real trade-offs is sometimes unintuitive. It's better to be somewhat unreasonable and say, we can build a safe AI that also makes a lot of revenue and brings value, and race to the top. That's something I've really appreciated.
我们多花点时间在这个话题上,因为我觉得这是我们想讨论的主题之一:AI 如何改变组织设计。
Let's spend a little more time here because I think this is one of the topics we wanted to talk about: how AI is changing org design.
我觉得审视 Claude Code 是如何改变(工作方式)的很有意思,这或许能预示许多公司的最终变革方向。比如,对比一年前你们是怎么构建的,现在有了新能力,你们在构建方式上有没有什么特别明显的变化?或者和以前做开发时相比,有哪些变化你会强调说“这大概就是新方式了”?
And I think examining how it's changed Claude Code is interesting as maybe a suggestion for how many people's companies will eventually change. Like if you compare how you guys build a year ago to now with the new capabilities, are there specific things that come to mind in terms of the ways you build feel very different to a year ago or maybe to past lives building stuff that you would highlight as 'this is probably the new way' given that we have these new capabilities?
我想说的是,我不认为组织管理只有一种模式,甚至不确定 Claude Code 团队就应该成为样板,因为我们非常注重“吃自己的狗粮”——我们在用产品本身来构建产品。所以我们的反馈循环可能和其他组织不同。因此我不太愿意说 Claude Code 团队是其他组织应该效仿的模型。但你可以思考一些首要实践:如何将生产力真正转化为更多收入,这最终是目标,对吧?我和客户聊过,我说,如果你的 PR 数量翻倍,你的收入可能根本不会变。你无法将其转化为收入增长,因为还有太多设计方面的因素,比如买账、决定做什么、人们是否想要这个、市场是否饱和等等。而且大多数事情都会失败,所以失败得更快并不一定赚更多钱,除非你找到了真正行得通的东西。所以我会用这个类比:假设你经营一家汽车经销商,电子邮件出现了。你可以给前台和团队配上电子邮件,他们沟通更快,效率更高。但我不确定你能否卖出更多车。或者你可以让人们通过电子邮件和网站订车,那可能会推动收入增长。所以,为你的具体问题找到类似的类比,是一种思考方式。
Yeah, I mean I do want to say that I don't think there's one pattern for org management and it's not even clear to me that the Claude Code team should be the pattern because we are very dog food oriented, right? We're literally building the product to build the product. And so that has a different feedback loop than maybe other organizations might. So I'm hesitant to say that the Claude Code team is a model for how other organizations should work, but I think you can sort of think through some first practices of how do you turn productivity into more revenue, honestly, is ultimately the goal, right? And I definitely talk to customers who I'm like, if you doubled your PRs, you probably wouldn't even change your revenue at all. You just are not able to turn that into increased revenue because there's so much of design, for example, buy-in, and deciding what to do, or do people even want this? Is the market saturated or something? And most things fail, so you don't actually make more money if you fail faster necessarily, unless you find the thing that really works. So I think the prompt I use is like, okay, let's say you're running a car dealership and email comes out. One thing you can do is give your receptionist and your team email and they can communicate a little bit faster and be more productive. But I don't know if you'll sell more cars. Or you could let people order cars via email and via the web, and hopefully that will drive more revenue. So I think trying to find that analogy for your specific problem is one way to think about it.
有没有什么特别极端或者异端的事情,对你来说可能已经习以为常了,但如果有人空降到 Claude Code 团队做技术成员,他们会说“这完全不一样”?
Are there specific things that are super extreme or heretical that maybe have become normalized to you, but if somebody were to drop into the Claude Code team as a member technical staff, they'd be like, 'This is totally different.'?
嗯,有些东西可能我不能讲。是啊。我觉得我观察到的一件事是,因为创建东西的成本太低了,我一直在琢磨人们现在怎么排优先级?你怎么给工作排优先级?因为你可以几乎并行地搞起所有事情。我觉得这里有一种奇怪的张力,我不太确定过去那种用 T 恤尺码估大小的方式——为什么还要费那个事?我只是好奇内部的变化。
Um, probably some stuff I can't talk about. Yeah. Yeah. I think one thing I've observed is because it's so cheap to create, I've been trying to wrap my head around how do people prioritize now? And how do you prioritize work? Because you can kind of just spin up everything in parallel. And I think there's this weird tension with I'm not really sure the old ways in which you might t-shirt size something. Like why even bother with that? And I'm just curious how that's changed internally.
好,再补一个问题——我不知道下一个拿麦克风的是谁,你可以插话。现在招聘大概很不一样,体现在你们看重的东西上。有没有什么 Anthropic 独有的新问题或练习,我们在招顶尖工程师时其他人也应该借鉴?
Yeah, one additional question on or then I don't know who's next with a mic, you can pipe in. Hiring is probably pretty different now in sort of what you look for. Are there specific questions or exercises that are net new to Anthropic that probably other people should adopt when we're looking to hire really talented engineers?
嗯,我觉得 Anthropic 目前招聘可能是个很不一样的地方,就是因为需求大。我现在自己也不太做面试了。我认为这是个难题,显然你想要懂技术、理解概念的人,但你还希望他们擅长 AI,这其实是两回事,有时候你能找到只擅长其中一样的人,而你两者都需要。嗯,说实话,你们可能在这方面更有见解。是的。
Um, I think Anthropic is probably a very different place for hiring right now just because of demand. I also don't do much like interviewing myself these days. I think it's a hard problem like obviously you want people who are technical and understand the concepts, but you also want to know that they are good at AI and these are two separate things and sometimes you can find someone who's good at one and not the other and you need both. Um, you guys probably honestly have more insight there. Yeah.
是的,这里有点正交性,有种张力——有些非常有经验的人可能已经不具备足够的神经可塑性来用这种新方式构建东西了,我相信我们都在边做边学这到底意味着什么。嗯,这里有个问题。
Yeah. There's sort of this orthogonality, there's tension with like somebody who's actually really experienced sometimes is probably doesn't have the neuroplasticity to actually build in this new way and I'm sure we're all learning on the fly what that looks like. Uh, question here.
那么,Claude Code 团队是怎么用 Claude Code 的?或者说你是怎么用的?你还是像我这样在说话的同时盯着智能体吗?你用 IDE 吗?有没有什么特殊的内部工具来管理工作树?你们到底用不用工作树?
So, how are the people or how's the Claude Code team using Claude Code or how are you using Claude Code? Are you still like babysitting agents like what I'm doing as we're speaking? Are you using IDE? Are you using some kind of special internal tool to manage the word trees? Are you using word trees at all?
基本上每个人做法都不同。
Everyone does things differently basically.
你的设置是什么样的?
What's your your setup?
嗯,我通常待在终端里。我最近在用 Claude 智能体。我同时做的事情最多也就两三个。嗯,我喜欢做很多探索。嗯,所以我觉得我比团队其他人更少多任务。嗯,是的,我觉得每个人之间差异非常大。比如我认识一些人,他们把 GitHub 用作状态来源,他们的智能体会往 issue 里写东西,然后等再次唤醒时再从 issue 里读取信息、留评论等等,他们把它当作一种沟通媒介。你知道,有这么多不同的设置,嗯。
Um, I tend to be in terminal. I've been using Claude agents recently. I tend to have at most maybe two or three things that I'm doing. Um, I tend to do a lot of exploration. Um, and so I tend to multitask less, I think, than the rest of the team. Um, yeah, I think that it's like everyone is so completely different from each other. Like I know some people who use GitHub as a source of state and so their agents will write to the issue and then when they wake it up again they'll read from the issue and leave comments there and things like that and they use it as a communication medium for example. You know, there's so many different setups and yeah.
那顺着这个话题我们往这边走。现在人们做事情的方式简直是狂野西部。即使在 Anthropic 团队内部也是这么狂野吗,还是有很多共享?有没有一个分享知识和最佳实践的流程?
I guess on that note we'll go right over here. Like it's such wild west with how people are doing things. Is it just wild west even within the Anthropic teams or is there a lot of shared? There's like a process for sharing knowledge and best practices?
我的意思是流程是有的,但大家也都在不断尝试和摸索。是的,可能两个问题合二为一了。你觉得 harness 设计在多大程度上是领域特定的?还有,你见过人们用 Claude Code 做哪些你没料到它也能行的疯狂事情?比如自动研究刚出来的时候,可能你们已经在做了,但对我来说它很新,很酷。
I mean there is but you also um yeah everyone's always you know experimenting and trying to find things as well. Yeah maybe two questions factor in one. So to what extent do you think harness design is domain specific? And what are crazy things you've seen people use Claude Code for that you did not expect it to work for? So like auto research when it came out is kind of like maybe you guys were already doing it, but to me it was new. It was cool.
但可能是完全不同的领域,或者用例。
But maybe completely different domains, or use cases.
嗯。
Yeah.
Harness 设计很棘手。我想你可能能想象,在我加入 Anthropic 之前,我为自己构建了一个电子表格智能体,那是一个代码生成 harness。我认为它在特定电子表格工作上会优于 Claude Code,但它的优势是否值得?Claude Code harness 正指数级改进,而我需要自己持续更新这个 harness,所以我不确定。你的第二个问题:人们用 Claude Code 做的疯狂事情。我最近看到一家有趣的公司做材料发现。基本上,他们编写代码来模拟材料对象的属性。几年前有 LK99 超导体的假新闻,但如果你能模拟一大堆材料,然后从模拟中找出哪些应该在实验室测试,然后去测试它们。所以我认为这是 Claude Code 的新用途之一,我很喜欢。
Harness design is tricky. I think you could probably imagine that before I came to Anthropic, I built a spreadsheet agent that was a codegen harness for myself. I do think that would outperform Claude Code at specific spreadsheet work, but would it outperform in a way that's worth it? The Claude Code harness is getting better exponentially, and I would need to continue to update this harness myself. So I don't know. And your second question: crazy things people have done with Claude Code. There was an interesting company I saw recently that does material discovery. Basically, they write code to simulate the properties of a material object. A couple years ago there was the LK99 superconductor thing that turned out to be fake. But what if you could simulate a bunch of these and then figure out from the simulations what to test in the lab, and go test them. So I thought that was one of my favorite new ways to use Claude Code.
我想听听你对 7 倍使用差距的看法。研究表明,重度用户从前沿模型中获得的推理能力是普通用户的七倍,而普通用户只是把它当作搜索引擎。作为一个构建者,你如何考虑设计 Claude Code 界面来缩小这一差距?我们能否通过引导用户执行更多多步智能体任务,用 UX 设计来消除这种能力落差?
I wanted to get your opinion on the 7x usage gap. Research says power users get up to seven times more advanced reasoning capabilities out of frontier models than average users who are merely using it as a search engine. As a builder, how are you thinking about designing the Claude Code interface to close that gap? And can we have a UX design that stops that capability overhang by teaching users to do more multi-step agentic tasks?
嗯,我认为两者都有。这绝对是我感兴趣的领域。我们需要改进智能体 UX,教学也很有价值,两方面都有很多工作要做。我认为人们有一种疲劳感,总是要学习新东西,这让人不爽。理想情况下,应该是 Claude 主动为你做事,并且精确知道该做多少。这是很难的问题。比如说,我有一个请求,Boris 也有一个请求。Boris 对代码库了如指掌。如果我问 ‘嘿,你能做这个改动吗?’ 也许 Claude 应该回应 ‘好的,你知道这个改动会有后续影响吗?’ 我会说 ‘不,我不知道。’ 但 Boris 可能会说 ‘当然知道,你为什么要问我?’ 所以模型非常复杂。不仅仅是模型和 harness,还有模型、harness、用户和世界。什么是可能的?你在哪个代码库操作?用户知道什么?用户如何对世界和 harness 建模?我认为这是目前 AI 最大的问题之一。
Yeah, I think there's both. This is definitely an interest of mine. I think we have to improve the agent UX, and teaching people is very valuable. There's a lot of work to do in both ways. I think there is some fatigue from people who feel they have to learn new things all the time, and that kind of sucks. Ideally, it should be Claude's job to do things for you proactively and know exactly the right amount of things. It's a hard problem. For example, let's say I have a request and Boris has a request. Boris knows the codebase inside out. If I ask, 'Hey, could you make this change?' Maybe Claude should respond, 'Okay, do you know that this change would have this follow-on effect?' And I'd say, 'No, I didn't.' But Boris might say, 'Actually, yeah, I did. Why are you asking me?' So the model is very complicated. It's not just model and harness. It's model, harness, user, and the world. What is possible? What codebase are you operating in? What does the user know? And how does the user model both the world and the harness? I think this is one of the biggest problems in AI right now.
我们还有 10 分钟,我想插一个问题,因为我们之前在短信里讨论过。我认为很多人熟悉 ‘Sauron’ 概念,但既然我们有一位来自 Sauron 的人,值得一问。上次我们在这里谈论它时,那是一个非常标题党的短语。有很多东西可以构建,仍然很有趣,模型也可以构建。但从内部说几句,你不需要分享细节。你会如何建议一位朋友,或者你在 SBC 的朋友,如何应对即将到来的模型?如何实际构建,利用模型优势,而不是在某些领域被碾压——你会说 ‘我不知道,老兄,这个方向不太好开始’?
So, we have 10 minutes left and I'm going to sneak this in because we were chatting about it on text. I think a lot of people are familiar with the 'Sauron' notion, but since we have somebody who is at Sauron, it's worth asking. Admittedly, last time we talked about it, it was a very clickbaity phrase. There's a lot that you can build that's still very interesting, and models can be built. But just a couple comments from the inside, you don't have to share specifics. How would you advise a friend, or your friends here at SBC, about navigating what is coming with models? How to actually build in a way that takes advantage of models versus getting steamrolled in specific areas where you'd say, 'I don't know, buddy, not a good one to start'?
我认为 AI 对所有人都不可预测。在 Anthropic,Dario 说过我们必须提前几年预分配算力。这风险很大,因为要购买数十亿美元的算力,并希望获得收入。我认为这是一个探索时代。每个人,包括前沿模型实验室,都在摸索现在可能实现什么。这非常令人兴奋。有很多未知的未知。方向在哪里?如何招聘?如何构建?如何组织团队?有些人看到你创业时会问,‘你为什么做这个?一切都在不断变化,你怎么知道它会成功?Anthropic 或 OpenAI 难道不会做这个吗?’所以从高层看,我认为你只需对探索和不确定性感到兴奋,并说‘不,我能做到。我真的相信我能做到。’拥有这种内在信心是最重要的,因为如今这不再是安全的选择,也不再是硅谷长期以来被视为高地位的事情。所以这是第一点。至于具体领域如何思考,再说一遍,老兄,我不会被碾压。School of beans。
I think AI is one of these things that is unpredictable for all of us. At Anthropic, Dario has talked about how we have to pre-allocate compute years ahead of time. It's very risky because we need to buy billions of dollars of compute and hope to make that revenue. I think of it as an age of exploration. Everyone, including frontier model labs, is figuring out what is possible now. It's really exciting. There are so many unknown unknowns. Where is the direction going? How do you hire? How do you build? How do you organize a team? Some people look at you when you're making a startup and ask, 'Why are you doing this? Everything is changing all the time. How do you know it's going to work? Isn't Anthropic or OpenAI going to do this?' So at a high level, I think you just have to be excited about the exploration and uncertainty, and say, 'No, I can do it. I really believe that I can do it.' Having that internal confidence is the most important thing, because it's not the safe thing to do anymore, nor the high-status thing to do, which it was for a long time in Silicon Valley. So that's one thing. As for specific areas of how to think about it, again, dude, I'm not going to get steamrolled. School of beans.
听着,我不是什么专家,你知道我的意思吗?我没有成功创业过,你明白吗?
Look, I'm no one's expert, you know what I mean? I haven't made a successful startup, you know what I mean?
可能一个思路是成为模型实验室的补充。目前世界上很多地方并不适合智能体。智能体无法访问你的健康数据、银行信息等。你能让这些对智能体更可读吗?可能需要不同的支付机制。但你能把锁定的数据源转化为智能体可以使用的东西,这非常有帮助和价值。另外,如何将 AI 分散到更受监管或有监管捕获的特定领域?我认为这非常有帮助。我认为不要试图一口吃成个胖子。
Probably one way to think about it is to be complementary to the model lab thing. So much of the world is not meant for agents right now. Agents can't access your health data, your banks, etc. Can you make that more legible to agents? There might be different payment mechanisms needed. But the more you can take locked data sources and turn them into something agents can use, that is extremely helpful and valuable. Also, how to disperse AI into specific niches that may be more regulated or have some regulatory capture? I think that is extremely helpful. I think boiling the ocean is not the way to go.
那么,与其把 AI 放到你的产品里,不如用 AI 产品能力去做那些以前过于 ambitious 的事情?我知道 Chang 做了 pretext,一个自定义文本渲染库。他之前在 React 团队工作,在浏览器文本渲染上遇到各种问题。然后他意识到:“等等,我以前根本做不了这个,但现在我就能直接做了。”我觉得这个思路是:什么问题需要 2000 万行代码或 5 亿行代码,你能用 AI 解决吗?
So instead of putting AI in your product, can you use AI product capabilities to do things that were just too ambitious to do before? I know Chang who made pretext, a custom text rendering library. He used to work on the React team and had all these problems with text rendering on the browser. Then he realized, "Wait, I would never have been able to do this before, but now I can just do it." And I think that idea of: what is the thing that needed 20 million lines of code or 500 million lines of code, and can you solve that?
是的,你的第一点让我想起我们之前说的:当构建成本低廉时,问题模式往往体现在大量人类问题上,而数据访问就是一个人问题。比如如何说服别人给你访问权限,走完所有繁琐流程。我认为这是一个值得探索的领域,你可以稍微避开聚光灯,因为它仍然需要非常有毅力的人去推销、进入、获取数据并改造那些传统工作流。
Yeah, your first point resonates with something we were saying: when building is cheap, the modes are found in a lot of human problems, and trying to access data is a human problem. It's like how do you convince somebody to give you access, going through all the red tape. I think that is one area to explore where you can be a little bit out of the spotlight, because it still requires someone with a lot of tenacity who just goes and sells it, gets in there, accesses the data and transforms those legacy workflows.
是的,当然。但谁知道呢,你也许能做出比我们更好的编码工具。你只需要相信它。所以,嗯。我以前常用一个例子,比如我用 Conductor 很多。我觉得它挺流行的,刚融了一大笔 A 轮。我不是代表 Anthropic,而是作为旁观者看这件事。以前作为开发工具,你会说:“挺有意思的,用户很多,受欢迎,赶紧跟进,有时间琢磨怎么做大生意。”但现在这个时代,这样的产品你会怀疑吗?你的反应是什么?
Yeah, definitely. But you never know, you could maybe make a coding harness that's better than us. You just have to believe in it. So yeah. I used to use one example, let's say I use Conductor a lot. I think it's pretty popular, just raised a big Series A. Not speaking on behalf of Anthropic, but just as a person who looks at that, would you in a previous world as a dev tool say "that's interesting, there's a lot of pickup, people are liking this, let's go run after that, there's time to figure out how to build a big business." But now in this day and age, is a product like that something you squint at? What is your reaction?
我认为每个产品都在赌能力的发展方向。我认为 Conductor 押对了模型商品化,可能还有点 S 曲线发展。我喜欢 Charlie,他非常有才华。当你做出人们喜爱的东西时,总会有很多好的结果。所以我觉得你不必过度思考这是不是战略上正确的事。人们喜爱 Conductor,无论如何他们都会做得很好。模型也在不断变便宜。所以你需要下注,而结果并不明朗,不知道谁会赢。
I think every product has a bet on where the capabilities are going. I think Conductor is a good bet on models commoditizing out, and maybe S-curving out a bit. I like Charlie. I think he's really talented, and there are lots of good outcomes when you make something people love. So I think you don't have to overthink whether this is strategically the right thing. People love Conductor, and they're going to do great no matter what. Models keep getting cheaper as well. So yeah, you need to take a bet and it's not clear which one will win.
我们就此打住。问题太多,我们很幸运有 Thariq 作为社区的一员。感谢你回来进行更深入的探讨。希望你能再来,带来更多演示之类的东西。你随时都可以来和我们一起吃午饭。我相信 Anthropic 的午餐也很不错。让我们用掌声感谢 Thariq。谢谢。节目结束。
Let's wrap there. So many questions, and we're lucky to have Thariq as part of the community. Thanks for coming back for a little longer deep dive. We hope to have you back for more demos and whatnot. You're always welcome to hang out, you know, when lunches are. I'm sure lunches at Anthropic are pretty good too. A welcoming round of applause for Thariq. Thank you. That was wrapping.