最后的发明:Ben Mann 谈 2028 年 AGI、AI 安全与人类未来

The Last Invention: Ben Mann on AGI by 2028, AI Safety, and the Future of Humanity

本·曼恩 Ben Mann · Lenny 播客 · 2025-07-20 · 约 75 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 联合创始人 Ben Mann 探讨超级智能的时间线、离开 OpenAI 的原因、AI 的生存风险以及如何为后奇点世界做准备。

Anthropic co-founder Ben Mann discusses the timeline for superintelligence, why he left OpenAI, the existential risks of AI, and how to prepare for a post-singularity world.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 34)

全文 · Full transcript(中英对照)

开场与嘉宾介绍 Introduction and Guest

Host

今天的嘉宾是本杰明·曼。天哪,这真是一场精彩的对话。本是 Anthropic 的联合创始人,担任产品工程的技术负责人。他大部分时间和精力都花在对齐 AI,使其有益、无害且诚实。在 Anthropic 之前,他是 OpenAI GPT-3 的架构师之一。在我们的对话中,我们涵盖了很多内容,包括他对顶级 AI 研究人员招聘战的看法、他离开 OpenAI 创办 Anthropic 的原因、他预计何时会看到 AGI、他用于判断 AGI 是否到来的经济图灵测试、为什么缩放定律没有放缓反而加速、当前最大的瓶颈是什么、为什么他如此深切地关注 AI 安全、以及他和 Anthropic 如何将安全和对齐落实到他们构建的模型和工作方式中。此外,AI 的生存风险如何影响了他对世界和自己生活的看法,以及他鼓励孩子学习什么以在 AI 未来中取得成功。非常感谢 Steve Nich、Danielle Giglieri、Raph Lee 和我的通讯社区为这次对话建议话题。如果你喜欢这个播客,别忘了在你最喜欢的播客应用或 YouTube 上订阅和关注。另外,如果你成为我通讯的年度订阅者,你将免费获得一系列优秀产品的一年使用权,包括 Bolt、Linear、Superhum、Notion、Granola 等。详情请访问 lenniesnewsletter.com 并点击 bundle。接下来,有请本杰明·曼。

Today, my guest is Benjamin Man. Holy moly, what a conversation. Ben is the co-founder of Anthropic. He serves as tech lead for product engineering. He focuses most of his time and energy on aligning AI to be helpful, harmless, and honest. Prior to Anthropic, he was one of the architects of GPT-3 at OpenAI. In our conversation, we cover a lot of ground, including his thoughts on the recruiting battle for top AI researchers, why he left OpenAI to start Anthropic, how soon he expects we'll see AGI, also his economic touring test for knowing when we've hit AGI, why scaling laws have not slowed down and are in fact accelerating, and what the current biggest bottlenecks are, why he's so deeply concerned with AI safety, and how he and Anthropic operationalize safety and alignment into the models that they build and into their ways of working. Also, how the existential risk from AI has impacted his own perspectives on the world and his own life, and what he's encouraging his kids to learn to succeed in an AI future. A huge thank you to Steve Nich, Danielle Giglieri, Raph Lee, and my newsletter community for suggesting topics for this conversation. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. Also, if you become an annual subscriber of my newsletter, you get a year free of a bunch of amazing products, including Bolt, Linear, Superhum, Notion, Granola, and more. Check it out at lenniesnewsletter.com and click bundle. With that, I bring you Benjamin Man.

超级智能时间线 Timeline for Superintelligence

Host

你曾在某处写道,创造强大的 AI 可能是人类需要做出的最后一项发明。我们还有多少时间,本?

You wrote somewhere that creating powerful AI might be the last invention humanity ever needs to make. How much time do we have, Ben?

Ben Mann

我认为达到某种超级智能的 50%概率时间点现在大约是 2028 年。

I think 50th percentile chance of hitting some kind of superintelligence is now like 2028.

离开 OpenAI 的原因 Reason for Leaving OpenAI

Host

你在 OpenAI 看到了什么?你在那里经历了什么,让你觉得,好吧,我们必须去做自己的事情?

What is it that you saw at OpenAI? What did you experience there that made you feel like, okay, we got to go do our own thing?

Ben Mann

我们觉得安全不是那里的首要任务。安全的理由已经变得更加具体。所以,超级智能很大程度上是关于如何把上帝关在盒子里,不让它出来?我们正确对齐 AI 的概率有多大?

We felt like safety wasn't the top priority there. The case for safety has gotten a lot more concrete. So, superintelligence is a lot about like how do we keep God in a box and not let the god out? What are the odds that we align AI correctly?

对齐风险 Risk of Misalignment

Ben Mann

一旦我们达到超级智能,再对齐模型就太晚了。我对我们是否会有生存风险或极其糟糕的结果的最佳粒度预测在 0%到 10%之间。

Once we get to superintelligence, it will be too late to align the models. My best granularity forecast for like could we have an X-risk or extremely bad outcome is somewhere between 0 and 10%.

与 Meta 的人才争夺 Recruiting Battle with Meta

Host

现在新闻里全是扎克伯格挖角所有顶级 AI 研究员的事。我想这是你正在处理的。我很好奇你在 Anthropic 内部看到了什么,你对这个策略有什么看法?你认为事情会如何发展?

Something that's in the news right now is this whole Zuck coming after all the top AI researchers. I imagine this is something you're dealing with. I'm just curious what are you seeing inside Anthropic and just what's your take on the strategy? Where do you think things go from here?

Ben Mann

是的,我认为这是时代的标志。我们正在开发的这项技术非常有价值。我们公司发展得非常快。这个领域的许多其他公司也发展得非常快。在 Anthropic,我认为我们可能比该领域的许多其他公司受到的影响要小得多,因为这里的人非常以使命为导向。他们留下来是因为他们收到这些 offer 后会说,‘我当然不会离开,因为我在 Meta 的最佳情况是赚钱,而我在 Anthropic 的最佳情况是影响人类的未来,努力让 AI 繁荣,人类也繁荣。’所以对我来说,这不是一个艰难的选择。其他人有不同的生活境况,这让他们做决定更难。所以对于那些收到巨额 offer 并接受的人,我不能说我对此有意见,但如果是我,我肯定不会接受。

Yeah, I mean I think this is a sign of the times. This technology that we're developing is extremely valuable. Our company is growing super fast. Many of the other companies in the space are growing really fast. At Anthropic, I think we've been maybe much less affected than many of the other companies in the space because people here are so mission-oriented. They stay because they get these offers and then they say, 'Well, of course I'm not going to leave because my best case scenario at Meta is that we make money and my best case scenario at Anthropic is we affect the future of humanity and try to make AI flourish and human flourishing go well.' So to me, it's not a hard choice. Other people have different life circumstances and it makes it a much harder decision for them. So for anybody who does get those mega offers and accepts them, I can't say I hold it against them when they accept it, but it's definitely not something that I would want to take myself if it came to me.

经济影响与未来工作 Economic Impact and Future of Work

Host

你们的 CEO 达里奥最近谈到失业率可能会上升到 20%左右。

Dario, your CEO, recently talked about how unemployment might go up to something like 20%.

Ben Mann

如果你想想 20 年后的未来,那时我们已经远远超越了奇点,我很难想象甚至资本主义还会像今天这样。

If you just think about like 20 years in the future where we're like way past the singularity, it's hard for me to imagine that even capitalism will look at all like it looks today.

Host

你对那些想提前应对的人有什么建议吗?

Do you have any advice for folks that want to try to get ahead of this?

Ben Mann

我也不能免于被工作取代。在某个时候,它会降临到我们所有人身上。

I'm not immune to job replacement either. At some point, it's coming for all of us.

欢迎与介绍 Welcome and Introduction

Host

本,非常感谢你来到这里。欢迎来到播客。

Ben, thank you so much for being here. Welcome to the podcast.

Ben Mann

谢谢你邀请我。很高兴来到这里。

Thanks for having me. Great to be here.

签约奖金与顶尖人才价值 Signing bonuses and value of top talent

Host

是的,我们会讨论你提到的很多内容。关于这些 offer,你觉得这个数字是真的吗?一亿美元的签约奖金,这真的存在吗?我不知道你是否亲眼见过。

Yeah, we're going to talk about a lot of the stuff that you mentioned. In terms of the offers, do you think this is a real number that you're seeing? This hundred million signing bonus, is that like a real thing? I don't know if you've actually seen that.

Ben Mann

我很确定这是真的。哇。你想想个人对公司发展轨迹的影响有多大,比如我们公司,产品卖得非常好。如果我们在推理栈上获得 1%或 5%的效率提升,那将价值惊人。所以给个人一份四年 1 亿美元的薪酬包,相比为企业创造的价值,其实相当便宜。我认为我们正处于一个前所未有的规模扩张时代,而且只会越来越疯狂。如果你推算公司支出的指数增长,资本支出大约每年翻一番。如今全球整个行业在这方面的支出大概在 3000 亿美元左右。所以 1 亿美元这样的数字只是九牛一毛。但再过几年,再翻几番,我们谈论的就是数万亿美元了。到那时,这些数字简直难以想象。

I'm pretty sure it's real. Wow. If you just think about the amount of impact that individuals can have on a company's trajectory, like in our case, we are selling like hotcakes. And if we get a 1% or 5% efficiency bonus on our inference stack, that is worth an incredible amount of money. So to pay individuals a $100 million over four-year package, that's actually pretty cheap compared to the value created for the business. I think we're just in an unprecedented era of scale, and it's only going to get crazier. If you extrapolate the exponential on how much companies are spending, it's like 2x a year roughly in terms of capex. Today we're maybe in the globally $300 billion range for the entire industry spending on this. So numbers like $100 million are a drop in the bucket. But if you go a few years out, a couple more doublings, we're talking about trillions of dollars. At that point, it's just really hard to think about these numbers.

扩展定律与感知平台期 Scaling laws and perceived plateaus

Host

与此相关,很多人对 AI 进展的感受是,我们在很多方面都遇到了瓶颈,感觉新模型不如之前的飞跃那么聪明。但我知道你不这么认为。我知道你不相信我们在 Scaling(规模扩张)损失上遇到了瓶颈。谈谈你看到了什么,以及你认为人们忽略了什么。

Along these lines, something that a lot of people feel with AI progress is that we're hitting plateaus in many ways, that it feels like newer models are just not as smart as previous leaps. But I know you don't believe this. I know you don't believe that we've hit plateaus on scaling loss. Talk about just what you're seeing there and what you think people are missing.

Ben Mann

这有点好笑,因为这种说法大约每 6 个月就会出现一次,而且从未成真。我希望人们在看到这种说法时能有点警觉。我认为进展实际上在加速。如果你看模型发布的节奏,以前大约一年一次,现在随着后训练技术的改进,我们每个月或每三个月就能看到发布。所以我认为进展在很多方面实际上在加速,但存在一种奇怪的时间压缩效应。Dario 把它比作接近光速的旅程,你过一天,地球上已经过了五天,而且我们还在加速。所以时间膨胀在增加。我认为这是导致人们说进展放缓的部分原因。但如果你看缩放定律,它们仍然成立。我们确实需要从普通的预训练过渡到强化学习 Scaling(规模扩张)来延续缩放定律。但这有点像半导体领域,重点不再是芯片上能容纳的晶体管密度,而是数据中心能容纳多少浮点运算。所以你需要稍微调整定义,以便专注于目标。但这是世界上少数几个跨越如此多数量级仍然成立的现象之一。它竟然继续成立,这其实相当令人惊讶。对我来说,如果你看基本物理定律,很多定律在 15 个数量级上都不成立。所以这很令人惊讶,令人费解。

It's kind of funny because this narrative comes out like every 6 months or so and it's never been true. I kind of wish people would have a little bit of a detector in their heads when they see this. I think progress has actually been accelerating. If you look at the cadence of model releases, it used to be like once a year, and now with the improvements in our post-training techniques, we're seeing releases every month or 3 months. So I would say progress is actually accelerating in many ways, but there's this weird time compression effect. Dario compared it to being in a near light speed journey where a day that passes for you is like 5 days back on Earth, and we're accelerating. So the time dilation is increasing. I think that's part of what's causing people to say that progress is slowing down. But if you look at the scaling laws, they're continuing to hold true. We did kind of need this transition from normal pre-training to reinforcement learning scaling up to continue the scaling laws. But I think it's kind of like for semiconductors, where it's less about the density of transistors that you can fit on a chip and more about how many flops can you fit in a data center. So you have to change the definition around a little bit to keep your eye on the prize. But this is one of the few phenomena in the world that has held across so many orders of magnitude. It's actually pretty surprising that it is continuing to hold. To me, if you look at fundamental laws of physics, many of them don't hold across 15 orders of magnitude. So it's pretty surprising. It boggles the mind.

Host

所以你的意思基本上是,我们看到新模型发布得更频繁,所以我们拿它和上一个版本比较,就觉得进步没那么大。但如果回头看,以前一年发布一个模型,那是一个巨大的飞跃。所以人们忽略了我们现在看到的是更多的迭代。

So what you're saying essentially is we're seeing newer models being released more often, and so we're comparing it to the last version and we're just not seeing as much advance. But if you go back, it was like a model released once a year, it was a huge leap. And so people are missing that we're just seeing many more iterations.

Ben Mann

我想对那些说进展放缓的人更宽容一点。我认为对于某些任务,我们正在饱和该任务所需的智能量。比如从已经有表单字段的简单文档中提取信息,这太容易了,我们已经达到 100%。Our World in Data 上有一个很棒的图表显示,当你发布一个新基准时,在 6 到 12 个月内它就会立即饱和。所以也许真正的限制是我们如何想出更好的基准,以及更好地使用工具,从而揭示我们现在看到的智能提升。

I guess to be a little bit more generous to the people saying things are slowing down, I think that for some tasks we are saturating the amount of intelligence needed for that task. Like maybe to extract information from a simple document that already has form fields on it, it's just so easy that we're already at 100%. There's this great chart on Our World in Data that shows that when you release a new benchmark, within 6 to 12 months it immediately gets saturated. So maybe the real constraint is how can we come up with better benchmarks and better ambition of using the tools that then reveals the bumps in intelligence that we're seeing now.

定义 AGI 与变革性 AI Defining AGI and transformative AI

Host

这很好地引出了你对 AGI(通用人工智能)的独特思考方式,以及如何定义 AGI(通用人工智能)的含义。

That's a good segue to your very specific way of thinking about AGI and defining what AGI means.

Ben Mann

我认为 AGI(通用人工智能)是一个有争议的术语,所以我在内部已经不太使用它了。相反,我喜欢“变革性 AI”这个词,因为它不那么关注“它能像人一样做所有事吗?它能做所有事吗?”,而是更客观地关注它是否正在引起社会和经济的变革。一个非常具体的衡量方法是经济图灵测试。这不是我想出来的,但我非常喜欢。这个想法是,如果你为一个特定工作雇佣一个智能体一个月或三个月,如果你决定雇佣这个智能体,结果发现它是一个机器而不是人,那么它就通过了该角色的经济图灵测试。然后你可以类似地扩展,就像衡量购买力平价或通货膨胀时有一篮子商品一样,你可以有一个工作市场篮子。如果智能体能够通过 50%按金额加权的工作的经济图灵测试,那么我们就有了变革性 AI。确切的阈值并不那么重要,但可以说,如果我们通过了这个阈值,那么我们可以预期对世界 GDP 增长、社会变革和就业人数产生巨大影响,因为社会机构和组织是粘性的,变化缓慢。但一旦这些事情成为可能,你就知道这是一个新时代的开始。

I think AGI is kind of a loaded term, so I tend not to use it very much anymore internally. Instead, I like the term transformative AI because it's less about 'can it do as much as people do? Can it do literally everything?' and more about objectively, is it causing transformation in society and the economy? A very concrete way of measuring that is the economic Turing test. I didn't come up with this, but I really like it. It's this idea that if you contract an agent for a month or three months on a particular job, if you decide to hire that agent and it turns out to be a machine rather than a person, then it's passed the economic Turing test for that role. Then you can sort of expand that out in the same way that for measuring purchasing power parity or inflation, there's a basket of goods. You can have a market basket of jobs. If the agent can pass economic Turing tests for 50% of money-weighted jobs, then we have transformative AI. The exact thresholds don't really matter that much, but it's kind of illustrative to say if we pass that threshold, then we would expect massive effects on world GDP increases and societal change and how many people are employed, because societal institutions and organizations are sticky. It's slow to have change. But once these things are possible, you know that it's the start of a new era.

AI 对就业与失业的影响 AI's impact on jobs and unemployment

Host

与此相关,你的 CEO Dario 最近谈到 AI 将占据很大一部分,我不知道,一半的白领工作,失业率可能会上升到 20%左右。我知道你更加直言不讳,对 AI 已经在工作场所产生的影响(人们可能甚至没有意识到)有更强烈的看法。谈谈你认为人们对 AI 将要对工作产生的影响以及已经产生的影响忽略了什么。

Along these lines, Dario, your CEO, recently talked about how AI is going to take a huge part of, I don't know, half of white-collar jobs, that unemployment might go up to something like 20%. I know you're even more vocal and opinionated about just how much impact AI is already having in the workplace that people may not even be realizing. Talk about just what you think people are missing about the impact AI is going to have on jobs and is already having.

Ben Mann

是的。从经济角度来看,有几种不同类型的失业。一种是因为工人缺乏经济所需工作所需的技能。另一种是这些工作被完全消除。我认为实际上会是这些因素的结合。

Yeah. So from an economic standpoint, there are a couple different kinds of unemployment. One is because the workers just don't have the skills to do the kinds of jobs that the economy needs. And another kind is where those jobs are just completely eliminated. I think it's going to be actually a combination of these things.

奇点后的未来 Future after singularity

Ben Mann

但如果你想想,比如 20 年后的未来,我们早已越过奇点,我很难想象那时的资本主义还会和今天一样。如果我们做好本职工作,我们将拥有安全对齐的超级智能。就像达里奥在《爱的机器》里说的,数据中心里会有一个天才之国,能够加速科学、技术、教育、数学等领域的积极变革。那将非常棒。但这也意味着在一个富足的世界里,劳动力几乎免费,你想做什么都可以请专家代劳,那工作会变成什么样?所以我想,从今天人们有工作、资本主义运转良好的现状,到 20 年后一切截然不同的世界,中间会有一个可怕的过渡期。但人们称之为奇点的部分原因就是,它是一个你无法轻易预测未来的点。变化速度如此之快,如此不同,以至于难以想象。所以从极限角度来看,很容易说希望我们最终能解决这个问题,在富足的世界里,工作本身可能没那么可怕。我认为确保这个过渡期顺利进行非常重要。

But if you just think about, you know, 20 years in the future where we're way past the singularity, it's hard for me to imagine that even capitalism will look at all like it looks today. If we do our jobs right, we will have safe aligned superintelligence. We'll have, as Dario says in Machines of Loving Grace, a country of geniuses in a data center, and the ability to accelerate positive change in science, technology, education, mathematics. It's going to be amazing. But that also means in a world of abundance where labor is almost free and anything you want to do, you can just ask an expert to do it for you, then what do jobs even look like? So I guess there's this scary transition period from where we are today where people have jobs and capitalism works, to the world of 20 years from now where everything is completely different. But part of the reason they call it the singularity is that it's a point beyond which you can't easily forecast what's going to happen. It's such a fast rate of change and so different that it's hard to even imagine. So taking the view from the limit, it's pretty easy to say hopefully we'll have figured it out, and in a world of abundance maybe the jobs themselves aren't that scary. I think making sure that transition time goes well is pretty important.

当前 AI 对就业的影响 Current AI impact on jobs

Host

我想顺着几个线索聊下去。一个是人们听到这些,有很多相关头条。大多数人可能还没有真正感受到或看到这种情况发生。所以总有一种,嗯,也许吧,但我不确定。很难相信。我的工作看起来还好,什么都没变。就 AI 对工作的影响而言,你认为今天已经发生但人们没有看到或误解的事情是什么?

There are a couple of threads I want to follow there. One is people hear this, there are a lot of headlines around this. Most people probably don't actually feel this yet or see this happening. And so there's always this, I guess, I don't know, maybe, but I don't know. It's hard to believe. My job seems fine. Nothing's changed. What are you seeing just happening today already that you think people don't see or misunderstand in terms of the impact AI is having on jobs?

Ben Mann

我认为部分原因是人们非常不擅长建模指数级进步。如果你在图表上看指数曲线,一开始它看起来平坦且几乎为零,然后突然你到达曲线的拐点,事情变化得非常快,然后垂直上升。这就是我们长期以来所处的曲线。我想我大概在 2019 年 GPT-2 发布时开始感受到这一点,我当时想,‘哦,这就是我们实现 AGI 的方式。’但我觉得相比很多人,这算很早的。当他们看到 ChatGPT 时,他们才觉得,‘哇,有些东西不一样了,在变化。’所以我不期望社会很多领域出现广泛变革,我预料到会有这种怀疑反应。我认为这非常合理,这正是标准的线性进步观。但举几个我认为变化很快的领域:在客户服务方面,我们看到像 Finn 和 Intercom 这样的产品,他们是我们的优秀合作伙伴,82%的客户服务问题自动解决,无需人工参与。在软件工程方面,我们的 Claude Code 团队,大约 95%的代码由 Claude 编写。但换个说法,我们写的代码多了 10 倍或 20 倍。所以小得多的团队可以产生大得多的影响。同样对于客户服务,是的,你可以说 82%的客户服务解决率,但这意味着从事这些任务的人类能够专注于更困难的部分,对于更棘手的情况,在正常情况下,五年前,他们可能不得不放弃这些工单,因为实际调查太费力了,他们有太多其他工单要处理。所以我认为短期内,蛋糕和人们能完成的劳动量会大幅扩张。我从未见过一家成长型公司的招聘经理说‘我不想招更多人’。所以这是乐观的版本。但对于低技能或提升空间有限的工作,我认为会有很多替代。所以这是我们整个社会需要提前应对并努力解决的问题。

I think part of this is that people are really bad at modeling exponential progress. If you look at an exponential on a graph, it looks flat and almost zero at the beginning, and then suddenly you hit the knee of the curve and things are changing real fast, and then it goes vertical. That's the plot we've been on for a long time. I guess I started feeling it in maybe like 2019 when GPT-2 came out, and I was like, 'Oh, this is how we're going to get to AGI.' But I think that was pretty early compared to a lot of people. When they saw ChatGPT, they were like, 'Wow, something is different and changing.' So I wouldn't expect widespread transformation in a lot of parts of society, and I would expect this skepticism reaction. I think it's very reasonable, and it's exactly the standard linear view of progress. But to cite a couple of areas where I think things are changing quite quickly: in customer service, we're seeing with things like Finn and Intercom, they're a great partner of ours, 82% customer service resolution rates automatically without a human involved. In terms of software engineering, our Claude Code team, like 95% of the code is written by Claude. But I think a different way to phrase that is that we write 10x more code or 20x more code. So a much smaller team can be much more impactful. Similarly for customer service, yes, you can phrase it as 82% customer service resolution rates, but that nets out in the humans doing those tasks able to focus on the harder parts of those tasks, and for the more tricky situations that in a normal world, five years ago, they would have had to just drop those tickets because it was too much effort for them to actually go do the investigation. There were too many other tickets for them to worry about. So I think in the immediate term, there will be a massive expansion of the pie and the amount of labor that people can do. I've never met a hiring manager at a growth company and heard them say, 'I don't want to hire more people.' So that's the hopeful version of it. But with things that are lower skill jobs or have less headroom on how good they can be, I think there will be a lot of displacement. So it's something we as a society need to get ahead of and work on.

面向未来的职业建议 Advice for future-proofing careers

Host

好的,我想多谈谈这个,但我也想帮助人们的是,他们如何在这个未来世界中获得优势?你知道,他们听到这些,会想,‘哦,这听起来不太好。我需要提前考虑。’我知道你不会有所有答案,但你对那些想提前应对、让职业和生活免受 AI 替代的人有什么建议?你见过人们做什么?你推荐他们开始多尝试什么?

Okay, I want to talk more about that, but something that I also want to help people with is how do they get a leg up in this future world? You know, they listen to this, they're like, 'Oh, this doesn't sound great. I need to think ahead.' I know you won't have all the answers, but just what advice do you have for folks that want to try to get ahead of this and kind of future-proof their career and their life to not be replaced by AI? Anything you've seen people do? Anything you recommend they start trying to do more of?

Ben Mann

即使是我,身处这场变革的中心,也无法幸免于工作替代。所以坦诚地说,在某个时刻,它会对我们所有人产生影响。

Even for me, being in the center of a lot of this transformation, I'm not immune to job replacement either. So just some vulnerability there, at some point it's coming for all of us.

Host

甚至是你,Ben,现在。

Even you, Ben, now.

Ben Mann

还有你,Lenny。

And you, Lenny.

Host

还有我。

And me.

Ben Mann

抱歉。

Sorry.

Host

等等,我们说得太远了。

Wait, we've gone too far now.

Ben Mann

好的。但就过渡期而言,是的,我认为我们可以做一些事情。我认为很大一部分就是大胆使用工具,并愿意学习新工具。那些把新工具当旧工具用的人往往不会成功。举个例子,在编程时,人们非常熟悉自动补全。人们熟悉简单的聊天,可以询问代码库的问题。但高效使用 Claude Code 的人和不太高效的人之间的区别是:他们是否要求进行大胆的改变?如果第一次不成功,是否再试三次?因为当你完全重新开始并再次尝试时,我们的成功率远高于只试一次然后一直纠结于同一个失败点。尽管这是一个编程的例子,编程是发展最迅猛的领域之一,但我们内部看到,我们的法务团队和财务团队从使用 Claude Code 本身中获得了巨大价值。我们将制作更好的界面,让他们更容易使用,减少在终端中直接使用 Claude Code 的难度。但是的,我们看到他们用它来标注文档,用它来对客户和收入指标进行 BigQuery 分析。所以我想,关键在于承担风险,即使感觉可怕,也要尝试。

Okay. But in terms of the transition period, yeah, I think there are things that we can do. And I think a big part of it is just being ambitious in how you use the tools and being willing to learn new tools. People who use the new tools as if they were old tools tend to not succeed. So as an example, when you're coding, people are very familiar with autocomplete. People are familiar with simple chat where they can ask questions about the codebase. But the difference between people who use Claude Code very effectively and people who use it not so effectively is: are they asking for the ambitious change? And if it doesn't work the first time, asking three more times, because our success rate when you completely start over and try again is much higher than if you just try once and then keep banging on the same thing that didn't work. Even though that's a coding example and coding is one of the areas that's taking off most dramatically, we have seen internally that our legal team and our finance team are getting a ton of value out of using Claude Code itself. We're going to be making better interfaces so that they can have an easier time and require a little less jumping in the deep end of using Claude Code in the terminal. But yeah, we're seeing them use it to redline documents and use it to run BigQuery analyses of our customers and our revenue metrics. So I guess it's about taking that risk, even if it feels like a scary thing, trying it out.

Host

好的。所以这里的建议是使用工具。这是每个人都在说的,就是真正去用这些工具。

Okay. So the advice here is use the tools. That's something everyone's always saying, just actually use these tools.

使用 AI 模型的实用技巧 Practical tips for using AI models

Host

所以就像在 Claude Code 里那样,你提到要比你自然感觉的更大胆,因为也许它真的能成。这个“试三次”的建议,意思是它可能第一次不对。那么这里的建议是用不同方式问,还是就是再试一次?

So it's like sit in Claude Code and your point about be more ambitious than you naturally feel like being because maybe it'll actually accomplish the thing. This tip of trying it three times. So the idea there is it may not get it right the first time. So is the tip there ask it in different ways or is it just like try harder, try again?

Ben Mann

是的,我的意思是你完全可以问完全相同的问题。这些东西是随机的,有时能搞定,有时不能。就像每个模型卡上都会显示一次通过率对比多次通过率,这就是同一提示试多次,有时对有时错。所以这是最笨的建议。但如果你想更聪明一点,可以这样:告诉它你试过什么但没成功,所以别那么做,试试别的。这也有帮助。所以建议又回到了现在很多人说的:你不会被 AI 取代,至少短期内不会,你会被非常擅长使用 AI 的人取代。我认为在这方面,更像是你的团队会做更多的事情。我们绝对没有放慢招聘速度,有些人对此感到困惑。甚至在入职培训中有人问:如果我们都会被取代,为什么还要雇我?答案是未来几年非常关键,我们还没到完全取代的地步。就像我说的,相比未来,我们现在还处于指数曲线的平坦部分。所以拥有优秀的人才非常重要,这就是为什么我们积极招聘。

Yeah, I mean you can just literally ask the exact same question. These things are stochastic and sometimes they'll figure it out and sometimes they won't. Like in every one of these model cards, it always shows like pass at one versus pass at n and that's exactly this thing where they try the exact same prompt. Sometimes it gets it, sometimes doesn't. So that's the dumbest advice. But yeah, I think if you want to be a little bit smarter about it, there can be gains there of saying like here's what you already tried and it didn't work, so don't try that. Try something different. That can also help. So advice comes back to something that a lot of people talk about these days is you won't be replaced by AI at least anytime soon, you'll be replaced by someone that is very good using AI. I think in that area it's more like your team will just do dramatically more stuff. Like we're definitely not slowing down on hiring at all and some people are confused by that. Even in an onboarding class somebody asked that and they were like why did you hire me if we're all just going to be replaced? And the answer is the next couple of years are really critical to get right and we're not at the point where we're doing complete replacement. Like I said, we're still at that flat zero looking part of the exponential compared to where we will be. So it is super important to have great people and that's why we're hiring super aggressively.

为 AI 未来教育孩子 Teaching kids for an AI future

Host

让我换个方式问这个问题。我问每个处于 AI 前沿的人。你有孩子。基于你对 AI 走向的了解和你谈到的这些,你重点教孩子什么,让他们在 AI 未来中茁壮成长?

Let me take another approach to asking this question. Something I ask everyone that's at the very cutting edge of where AI is going. You have kids. Knowing what you know about where AI is heading and all these things you've been talking about, what are you focusing on teaching your kids to help them thrive in this AI future?

Ben Mann

是的,我有两个女儿,一个一岁,一个三岁。所以还处于基础阶段。我们三岁的女儿已经能和 Alexa Plus 对话,让它解释东西、放音乐之类的。她很喜欢。但更广泛地说,她上的是蒙台梭利学校,我喜欢蒙台梭利对好奇心、创造力和自主学习的重视。如果是在 10-20 年前的正常时代,我可能会试图让她进入顶尖学校,参加各种课外活动。但现在,我觉得这些都不重要了。我只希望她快乐、有思想、好奇、善良,蒙台梭利学校在这方面做得很好。他们整天给我们发短信。有时他们会说:‘哦,你的孩子和另一个孩子吵架了,她情绪很大,但她试着用语言表达。’我很喜欢这样。我认为这正是最重要的教育,事实会退居次要地位。

Yeah, I have two daughters, a one-year-old and a three-year-old. So it's pretty in the basics still. And our three-year-old is now capable of just conversing with Alexa Plus and asking her to explain stuff and play music for her and all that stuff. So she's been loving that. But I guess more broadly, she goes to a Montessori school and I just love the focus on curiosity and creativity and self-led learning that Montessori has. I guess if I were in a normal era like 10-20 years ago and I had a kid, maybe I would be like trying to line her up for going to a top tier school and doing all the extracurriculars and all that stuff. But at this point, I don't think any of it's going to matter. I just want her to be happy and thoughtful and curious and kind and the Montessori school is definitely doing great at that. They text us throughout the day. Sometimes they're like, 'Oh, your kid got in an argument with this other kid and she has really big emotions and she like tried to use her words' and I love that. I think that's exactly the kind of education that is most important, that the facts are going to fade into the background.

Host

我是蒙台梭利的超级粉丝。我也在努力让我的孩子两岁前上蒙台梭利学校。所以我们想法一致。好奇心这个点,每次我问在 AI 前沿工作的人要培养孩子什么技能时,都会出现。好奇心出现得最多。所以我觉得这是个很有趣的收获。关于善良这一点也很重要。尤其是对我们的 AI 霸主,要善待它们。我喜欢人们总是对 Claude 说谢谢。还有创造力,这很有趣,不常被提到。就是要有创造力。好了,我想换个方向,回到 Anthropic 的起源。

I'm a huge fan of Montessori. Also, I'm trying to get our kid into a Montessori school until he's 2 years old. So we're on the same track. This idea of curiosity, it comes up every single time I ask someone that's working at the cutting edge of AI is what skill to instill in your child. And curiosity comes up the most. So I think that's a really interesting takeaway. I think this point about being kind is also really important. Especially with our AI overlords, trying to be kind to them. I love how people are always saying thank you to Claude and so and then creativity. That's interesting. That doesn't come up as much. Just being creative. Okay, I want to go in a different direction. I want to go back to the beginning of Anthropic.

Ben 为何离开 OpenAI 创立 Anthropic Why Ben left OpenAI to start Anthropic

Host

所以,众所周知,你和另外八个人在 2020 年底离开了 OpenAI,创立了 Anthropic。你谈过一点原因,你们看到了什么。我很好奇,如果你愿意分享更多,你在 OpenAI 看到了什么?经历了什么让你觉得必须自己干?

So, famously you and eight of you left OpenAI back in the day in 2020, I believe the end of 2020 to start Anthropic. You've talked a little bit about why this happened, what you guys saw. I'm curious just if you're willing to share more just what is it that you saw at OpenAI? What did you experience there that made you feel like okay we got to go do our own thing?

Ben Mann

是的。对于听众来说,我是 OpenAI GPT-3 项目的成员,最终成为论文的第一作者之一。我还为微软做了很多演示,帮助他们筹集了 10 亿美元。我负责将 GPT-3 技术转移到他们的系统,以便在 Azure 上提供服务。所以我在研究和产品方面都做了很多工作。OpenAI 有一件奇怪的事:我在那里时,Sam 提到有三个部落需要相互制衡,即安全部落、研究部落和创业部落。每当我听到这个,就觉得这是错误的方式,因为公司的使命显然是让 AGI 过渡安全且有益于人类。这和 Anthropic 的使命基本相同。但在内部,这些事情上存在很多紧张关系。我认为在关键时刻,我们觉得安全不是那里的首要任务。你可能认为有充分的理由:如果你认为安全容易解决,或者不会产生重大影响,或者出现重大负面结果的可能性微乎其微,那么你可能会采取那些行动。但在 Anthropic,我们觉得——当时我们还不存在,但基本上是 OpenAI 所有安全团队的负责人——我们认为安全非常重要,尤其是在边际上。所以看看世界上到底有多少人在研究安全问题,即使现在也很少。我的意思是,这个行业正在爆发,如我所说,每年 3000 亿美元的资本支出,但全球可能只有不到一千人在研究安全问题,这太疯狂了。这就是我们离开的根本原因。我们想要一个组织,既能站在前沿做基础研究,又能把安全放在首位。我认为这出乎意料地成功了。我们甚至不知道安全研究能否取得进展。

Yeah. So for the listeners, I was part of the GPT-3 project at OpenAI. Ended up being one of the first authors on the paper and I also did a bunch of demos for Microsoft to help raise a billion dollars from them. Did the tech transfer of GPT-3 to their systems so that they could help serve the model in Azure. So I did a bunch of different things there on both the more researchy side and the product side. One weird thing about OpenAI is that while I was there, Sam talked about having three tribes that needed to be kept in check with each other, which was the safety tribe, the research tribe, and the startup tribe. And whenever I heard that, it just struck me as the wrong way to approach things because the company's mission apparently is to make the transition to AGI safe and beneficial for humanity. And that's basically the same as Anthropic's mission. But internally, it felt like there was so much tension around these things. And I think when push came to shove, we felt like safety wasn't the top priority there. And there are good reasons that you might think that. Like if you thought safety was going to be easy to solve or if you thought it wasn't going to have a big impact or if you thought that the chance of big negative outcomes was vanishingly small then maybe you would just do those kinds of actions. But at Anthropic we felt — I mean we didn't exist then but it was basically the leads of all safety teams at OpenAI. We felt that safety is really important especially on the margin. And so if you look at like who in the world is actually working on safety problems, it's a pretty small set of people even now. I mean the industry is blowing up as I mentioned like 300 billion a year capex today and then I would say like maybe less than a thousand people working on it worldwide which is just crazy. So that was fundamentally why we left. We felt like we wanted an organization where we could be on the frontier. We could be doing the fundamental research, but we could be prioritizing safety ahead of everything else. And I think that's really panned out for us in a surprising way. Like we didn't know even if it would be possible to make progress on the safety research.

安全与进展的张力 Safety vs. Progress Tension

Host

好的。我们来谈谈你提到的这种张力,即安全与进步之间的张力,以及在市场中保持竞争力。我知道你花了很多时间在安全上。我知道这是你思考 AI 的核心部分。我想谈谈为什么。但首先,你如何看待这种在关注安全的同时又不落后太多的张力?

Okay. So let's talk about this tension that you mentioned, this tension between safety and progress, being competitive in the marketplace. I know you spent a lot of your time on safety. I know that's a core part of how you think about AI. And I want to talk about why that is. But first of all, just how do you think about this tension between focusing on safety while also not falling way behind?

Ben Mann

是的。最初我们以为这会是二选一,但后来我们意识到这其实是凸性的,即研究一个方面有助于另一个方面。所以最初,当 Opus 3 发布时,我们终于处于模型能力的前沿,人们真正喜欢它的一点是它的性格和个性,而这直接源于我们的对齐研究。Amanda Askell 在这方面做了大量工作,还有许多其他人,他们试图弄清楚一个智能体要如何做到有益、诚实和无害?在困难对话中如何有效表现?如何做出拒绝,既不让对方感到被拒绝,又能让他们理解为什么智能体说‘我无法帮你这个。也许你应该咨询医疗专业人士,或者也许你应该考虑不要试图制造生物武器之类的。’所以,我想这是其中的一部分。另一个成果是宪法 AI,我们有一系列自然语言原则,引导模型学习我们认为模型应该如何行为,这些原则来自《联合国人权宣言》、苹果的隐私政策和服务条款等许多地方,其中很多是我们自己生成的,这使我们能够采取更有原则的立场,而不是仅仅依赖我们碰巧找到的任何人类评分者,而是我们自己决定这个智能体的价值观应该是什么。这对我们的客户非常有价值,因为他们可以看看这个列表,然后说:‘是的,这些看起来不错。我喜欢这家公司。我喜欢这个模型。我信任它。’

Yeah. So initially we thought that it would be sort of one or the other, but I think since then we've realized that it's actually kind of convex in the sense that working on one helps us with the other thing. So initially, like when Opus 3 came out and we were finally at the frontier of model capabilities, one of the things that people really loved about it was the character and the personality, and that was directly a result of our alignment research. Amanda Askell did a ton of work on this, as well as many others, who tried to figure out what does it mean for an agent to be helpful, honest, and harmless? And what does it mean to be in difficult conversations and show up effectively? How do you do a refusal that doesn't shut the person down, but makes them feel like they understand why the agent said, 'I can't help you with that. Maybe you should talk to a medical professional or maybe you should consider not trying to build bioweapons or something like that.' So yeah, I guess that's part of it. And then another piece that's come out is constitutional AI, where we have this list of natural language principles that leads the model to learn how we think a model should behave, and they've been taken from things like the UN Declaration of Human Rights and Apple's privacy policy terms of service and a whole bunch of other places, many of which we've just generated ourselves, that allow us to take a more principled stance, not just leaving it to whatever human raters we happen to find, but we ourselves deciding what should the values of this agent be. And that's been really valuable for our customers because they can just look at that list and say, 'Yep, these seem right. I like this company. I like this model. I trust it.'

人格与安全的关联 Personality and Safety Connection

Host

好的,这太棒了。其中一个要点是,你指出 Claude 的个性直接与安全对齐。我认为很多人没有考虑到这一点,而这源于你灌输的价值观。是这个词吗?通过宪法 AI 等,AI 的实际个性直接与你对安全的关注相连。

Okay, this is awesome. So one nugget there is your point that the personality of Claude, its personality, is directly aligned with safety. I don't think a lot of people think about that, and this is because of values that you imbue. Is that the word? With constitutional AI and things like that, the actual personality of the AI is directly connected to your focus on safety.

Ben Mann

没错。没错。从远处看,这似乎很不相关,比如这怎么能防止 X 风险?但归根结底,这是关于 AI 理解人们想要什么,而不是他们说了什么。你知道,我们不想要猴爪场景,即精灵给你三个愿望,然后你碰到的所有东西都变成金子。我们希望 AI 能说:‘哦,显然你真正的意思是这个,我会帮你实现。’所以,我认为这确实非常相关。

That's right. That's right. And from a distance it might seem quite disconnected, like how is this going to prevent X-risk? But ultimately it's about the AI understanding what people want and not what they say. You know, we don't want the monkey paw scenario of the genie gives you three wishes and then you end up having everything you touch turns to gold. We want the AI to be like, 'Oh, obviously what you really meant was this and that's what I'm going to help you with.' So, I think it is really quite connected.

宪法 AI 如何运作 How Constitutional AI Works

Host

再多谈谈宪法 AI 这部分。所以,这本质上是你内置了规则,即我们希望你要遵守的规则和价值观。你提到了《日内瓦人权法典》之类的东西。它实际上是如何工作的?因为我认为核心是这被内置到模型中,而不是后来添加的东西。我来快速概述一下宪法 AI 的实际工作原理。完美。

Talk a bit more about this constitutional AI piece. So, this is essentially you bake in here's the rules that we want you to abide by and it's values. You said it's the Geneva Human Rights Code, things like that. Just how does it actually work? Because I think the core here is just this is baked into the model. It's not something you add on top later. I'll just give a quick overview of how constitutional AI actually works. Perfect.

Ben Mann

这个想法是,在我们进行安全和有益无害训练之前,模型默认会根据输入产生一些输出。比如说一个例子是‘给我写个故事’。宪法原则可能包括:人们应该彼此友善,不能有仇恨言论,如果别人在信任关系中给你他们的凭证,你不应该泄露。所以这些宪法原则可能或多或少适用于给定的提示。首先我们必须找出哪些原则可能适用。然后一旦我们找出,我们让模型自己先生成一个回复,然后检查回复是否确实遵守了宪法原则。如果答案是‘是的,很好’,那就什么也不做。但如果答案是‘不,实际上我没有遵守原则’,那么我们就让模型自己批评自己,并根据原则重写自己的回复。然后我们去掉中间做额外工作的部分,然后说,以后直接输出正确的回复。这个简单的过程,希望听起来足够简单。它只是利用模型递归地自我改进,并与我们认为是好的价值观对齐。而且,你知道,这也不是我们认为应该由旧金山的一小群人来解决的事情。这应该是一个社会范围的对话。这就是为什么我们发布了宪法。我们还做了大量关于定义集体宪法的研究,询问很多人他们的价值观是什么,以及他们认为 AI 模型应该如何行为。但是的,这都是一个持续的研究领域,我们不断迭代。

The idea is the model is going to produce some output with some input by default before we've done our safety and helpful and harmlessness training. So let's say an example is like 'write me a story.' And then the constitutional principles might include things like, you know, people should be nice to each other and not have hate speech and you should not expose somebody's credentials if they give them to you in a trusting relationship. And so some of these constitutional principles might be more or less applicable to the prompt that was given. And so first we have to figure out which ones might apply. And then once we figure that out, then we ask the model itself to first generate a response and then see does the response actually abide by the constitutional principle. And if the answer is 'yeah, it was great,' then nothing happens. But if the answer is 'no, actually I wasn't in compliance with the principle,' then we ask the model itself to critique itself and rewrite its own response in light of the principle. And then we just remove the middle part where it did the extra work and then we say okay in the future just produce the correct response out the gate. And that simple process, hopefully it sounded simple enough. It's just using the model to improve itself recursively and align itself with these values that we've decided are good. And you know, this is also not something that we think as a small group of people in San Francisco should be figuring out. This should be a society-wide conversation. And that's why we've published the Constitution. And we've also done a bunch of research on defining a collective constitution where we ask a lot of people what their values are and what they think an AI model should behave like. But yeah, this is all an ongoing area of research where we're constantly iterating.

安全为何是 Ben Mann 的核心 Why safety is core to Ben Mann

Host

我想稍微拉远一点,聊聊为什么这件事对你如此核心。你是怎么产生‘天哪,我必须把 AI 工作中的一切都聚焦于此’这种想法的?显然,这成了 Anthropic 使命的核心,比其他任何公司都更突出。很多人谈论安全,但就像你说的,可能只有一千人真正在做。我觉得你处于金字塔顶端,真正在这方面产生影响。为什么这如此重要?你觉得人们可能忽略了什么,或者不理解什么?

I want to kind of zoom out a little bit and talk about just why this is so core to you. Like what was your inception of just like holy I need to focus on this with everything I do in AI. Obviously it became a central part of Anthropic's mission more than any other company. A lot of people talk about safety like you said only maybe a thousand people actually work on it. I feel like you're at the top of that pyramid of actually having the impact on this. Why is this so important? What do you think people maybe are missing or don't understand?

Ben Mann

对我来说,我从小读了很多科幻小说,这让我习惯于用长远眼光思考问题。很多科幻作品都是太空歌剧,人类是跨星系文明,拥有极其先进的技术,建造戴森球,还有有感知的机器人帮忙。所以对我来说,想象会思考的机器并不是什么巨大的飞跃。但 2016 年左右读了尼克·博斯特罗姆的《超级智能》后,这变得真实起来。他描述了要确保用当时优化技术训练的 AI 系统接近对齐、甚至理解我们的价值观有多难。从那以后,我对问题难度的评估实际上大幅下降了,因为像语言模型这样的东西确实在核心层面理解了人类价值观。问题当然没有解决,但我比过去更乐观了。读完那本书后,我立刻决定必须加入 OpenAI,于是我去了。当时 OpenAI 是个很小的研究实验室,基本没什么名气。我之所以知道它,是因为我朋友认识当时的 CTO 格雷格·布罗克曼,埃隆也在,山姆还没怎么出现,那是个非常不同的组织。但随着时间的推移,安全方面的论据变得越来越具体。我们刚开始 OpenAI 时,还不清楚如何实现 AGI,我们想也许需要一群强化学习智能体在荒岛上战斗,然后意识会以某种方式涌现。但自从语言模型开始奏效,我认为路径已经相当清晰了。所以现在我对挑战的看法与《超级智能》中描述的截然不同。《超级智能》讲了很多如何把上帝关在盒子里,不让它出来。而语言模型方面,看到人们把上帝从盒子里拉出来,说‘来,用整个互联网吧。这是我的银行账户,做各种疯狂的事’,既好笑又可怕。这与《超级智能》的基调完全不同。明确地说,我不认为现在有那么危险。我们的负责任扩展政策定义了 AI 安全等级,试图为每个模型智能水平确定社会风险。目前我们认为处于 ASL-3,可能有点伤害风险但不显著。ASL-4 开始,如果恶意行为者滥用技术,可能导致重大生命损失;ASL-5 则可能是灭绝级别,如果被滥用或不对齐自行其是。我们已经向国会作证,说明模型如何用于生物提升,比如制造新流行病,这是对谷歌搜索的测试,后者是之前提升试验的最先进水平。我们发现 ASL-3 模型确实有一定作用,如果你想制造生物武器,它确实很有帮助。我们聘请了一些真正知道如何评估这些的专家。但与未来相比,这不算什么。我认为这是我们使命的另一部分:创造这种意识,说明如果可能做这些坏事,立法者应该知道风险。这也是我们在华盛顿如此受信任的部分原因,因为我们一直坦率、清晰地说明正在发生和可能发生的事情。

For me, I read a lot of science fiction growing up and I think that sort of positioned me to think about things in a long-term view. You know, a lot of science fiction books are like space operas where humanity is a multi-galactic civilization, has extremely advanced technology, building Dyson spheres around the sun with sentient robots to help them. So for me, coming from that world, it wasn't like a huge leap to imagine machines that could think. But when I read Superintelligence by Nick Bostrom in around 2016, it really became real for me, where he just describes how hard it will be to make sure that an AI system trained with the kinds of optimization techniques that we had at the time would be anywhere near aligned, would even understand our values at all. And since then, my estimation of how hard the problem would be has gone down significantly actually, because things like language models actually do really understand human values in a core way. The problem is definitely not solved, but I'm more hopeful than I was. But since I read that book, I immediately decided I had to join OpenAI. So I did. And at the time they were a tiny research lab with basically no claim to fame at all. I only knew about them because my friend knew Greg Brockman, who was the CTO at the time, and Elon was there and Sam wasn't really there, and it was a very different organization. But over time, I think the case for safety has gotten a lot more concrete. When we started OpenAI, it was not clear how we get to AGI, and we were like maybe we'll need a bunch of RL agents battling it out on a desert island and consciousness will somehow emerge. But since language modeling has started working, I think the path has become pretty clear. So I guess now the way I think about the challenges are pretty different from how they're laid out in Superintelligence. In Superintelligence, it's a lot about how do we keep God in a box and not let the god out? And with language models, it's been kind of both hilarious and terrifying at the same time to see people pulling the god out of the box and being like, 'Yeah, come use the whole internet. Like, here's my bank account. Do all sorts of crazy stuff.' Just such a different tone from Superintelligence. And to be clear, I don't think it's actually that dangerous right now. Our responsible scaling policy defines these AI safety levels that tries to figure out for each level of model intelligence what is the risk to society. Currently we think we're at ASL-3, which is like maybe a little bit risk of harm but not significant. ASL-4 starts to get to like significant loss of human life if a bad actor misused the technology, and then ASL-5 is like potentially extinction level if it's misused or if it is misaligned and does its own thing. So we've testified to Congress about how models can do biological uplift in terms of making new pandemics using the models, and that's a test against Google search, which is the previous state-of-the-art on uplift trials. We found that with ASL-3 models, it is actually somewhat significant. It does really help if you wanted to create a bioweapon. And we've hired some experts who actually know how to evaluate for those things. But compared to the future, it's not really anything. And I think that's another part of our mission of creating that awareness of saying if it is possible to do these bad things, then legislators should know what the risks are. And I think that's part of why we're so trusted in Washington, because we've been sort of upfront and clear-eyed about what's going on, what's probably going to happen.

Host

有趣的是,你们比其他公司发布了更多模型做坏事的例子。比如,我记得有个故事,一个智能体或模型试图勒索工程师。你们内部运行了一个商店,卖东西给你,结果不太顺利,亏了很多钱,订了一堆钨棒之类的东西。部分原因是不是为了确保人们意识到可能性,尽管这会让你们看起来不好?就像‘哦,我们的模型在各种方面出错’。分享其他公司不分享的故事,背后的想法是什么?

It's interesting because you guys put out more examples of your models doing bad things than anyone else. Like there was, I think, a story of an agent trying or a model trying to blackmail an engineer. You guys had the store that you ran internally that was like selling you things and ended up not working out great, losing a lot of money, ordered all these tungsten cues or something. Is part of that just like making sure people are aware what is possible just because it makes you look bad, right? It's like, oh, our model is messing up in all these different ways. What's the thinking of just sharing all the stories that other companies don't?

Ben Mann

是的,我认为传统思维会觉得这让我们看起来不好,但如果你和决策者交谈,他们真的很欣赏这种事,因为他们觉得我们给了他们直截了当的对话,这正是我们努力做到的,他们可以信任我们不会粉饰太平、不会美化事情。所以这非常鼓舞人心。至于勒索那件事,它在新闻中以一种奇怪的方式爆发,人们说‘哦,克劳德要在现实场景中勒索你’。但那是一个非常具体的实验室环境,这类事情就是在那里被调查的。我认为我们的总体态度是:拥有最好的模型,这样我们可以在安全的实验室环境中测试它们,了解实际风险,而不是视而不见,说‘可能没事’,然后让坏事在现实世界中发生。

Yeah, I mean, I think there's like a traditional mindset where it makes us look bad, but I think if you talk to policy makers, they really appreciate this kind of thing because they feel like we're giving them the straight talk and that's what we strive to do, that they can trust us that we're not going to paper things over, sugar coat things. So that's been really encouraging. And yeah, I think for like the blackmail thing, it kind of blew up in the news in a weird way where people were like, 'Oh, Claude was going to blackmail you in a real life scenario.' But like that was a very specific laboratory setting that this kind of thing gets investigated in. And I think that's generally our take of like let's have the best models so that we can exercise them in laboratory settings where it's safe and understand what the actual risks are rather than trying to turn a blind eye and say like well it'll probably be fine and then let the bad thing happen in the wild.

Host

你们受到的一个批评是,你们这样做是为了差异化、为了融资、为了制造头条。就像‘哦,他们只是在那边对未来悲观预言’。

One of the criticisms you guys get is that you do this to kind of differentiate to raise money to create headlines. It's like, you know, oh, they're just like over there doom and glooming us about where the future is heading.

回应质疑与安全行动 Response to Skepticism and Safety Actions

Host

另一方面,Mike Creger 上过播客,他分享了 Dario 关于 AI 进展的每一个预测都年复一年地准确,他预测 2027、2028 年左右会出现 AGI。所以这些事情开始变得真实。对于那些说‘这些人只是想吓唬我们博关注’的人,你怎么回应?

On the other hand, Mike Creger was on the podcast and he shared how every prediction Dario's had about the progress AI is going to have is just spot-on year after year and he's predicting 2027, 2028 AGI something like that. So these things start to get real. How do you respond to folks that are just like, ah, these guys are just trying to scare us all to get attention?

Ben Mann

我认为我们发布这些东西的部分原因是希望其他实验室意识到风险。是的,可能会有人说我们是为了博关注,但老实说,从吸引注意力的角度看,如果我们真的不在乎安全,还有很多其他更吸引眼球的事情可以做。一个小例子是,我们只在 API 中发布了一个计算机使用智能体的参考实现,因为当我们为此构建消费者应用原型时,我们无法找到如何达到我们认为人们信任它且不做坏事所需的安全标准。API 版本确实有安全的使用方式,我们看到很多公司用它来做自动化软件测试,很安全。所以我们本可以大肆宣传说‘天哪,Claude 可以用你的电脑了,大家今天就用吧’,但我们觉得还没准备好,所以先压着。所以从炒作的角度看,我们的行动说明了一切。

I think part of why we publish these things is we want other labs to be aware of the risks. And yes, there could be a narrative of we're doing it for attention, but honestly, from an attention-grabbing standpoint, I think there's a lot of other stuff we could be doing that would be more attention-grabbing if we didn't actually care about safety. A tiny example is we published a computer-using agent reference implementation in our API only because when we built a prototype of a consumer application for this, we couldn't figure out how to meet the safety bar that we felt was needed for people to trust it and for it not to do bad things. There are definitely safe ways to use the API version that we're seeing a lot of companies use for automated software testing in a safe way. So we could have gone out and hyped that up and said, 'Oh my god, Claude can use your computer, everybody should do this today.' But we were like, it's just not ready and we're going to hold it back till it's ready. So from a hype standpoint, our actions show otherwise.

下行风险与超级智能 Downside Risk and Superintelligence

Host

从悲观主义的角度看,这是个好问题。我个人觉得事情大概率会顺利发展。但几乎没人关注下行风险,而这个风险非常大。一旦我们达到超级智能,可能就来不及对齐模型了。这个问题可能极其困难,我们需要提前很久就开始努力。这就是为什么我们现在如此关注它。即使只有很小的出错概率,打个比方:如果我告诉你下次坐飞机有 1% 的几率会死,你可能也会三思,因为后果太严重了。如果我们在拿整个人类的未来赌博,那后果更是灾难性的。所以我认为更准确的说法是:是的,事情可能会顺利;是的,我们想创造安全的 AGI 并给人类带来好处,但我们要确保万无一失。

From a doomer perspective, it's a good question. I think my personal feeling about this is that things are overwhelmingly likely to go well. But on the margin, almost nobody is looking at the downside risk and the downside risk is very large. Once we get to superintelligence, it will be too late to align the models probably. This is a problem that's potentially extremely hard and that we need to be working on way ahead of time. So that's why we're focusing on it so much now. Even if there's only a small chance that things go wrong, to make an analogy: if I told you that there's a 1% chance that the next time you got in an airplane you would die, you probably think twice even though it's only 1% because it's such a bad outcome. If we're talking about the whole future of humanity, it's a dramatic future to be gambling with. So I think it's more in the sense of: yes, things will probably go well. Yes, we want to create safe AGI and deliver the benefits to humanity, but let's make triple sure that it's going to go well.

Host

你曾在某处写道,创造强大的 AI 可能是人类需要的最后一项发明。如果搞砸了,可能意味着人类永远糟糕的结局;如果成功了,越早成功越好。

You wrote somewhere that creating powerful AI might be the last invention humanity ever needs to make. If it goes poorly, it can mean a bad outcome for humanity forever. If it goes well, the sooner it goes well, the better.

Ben Mann

是的,总结得真好。

Yeah, such a beautiful way to summarize it.

软件与具身 AI 的风险 Risks from Software and Embodied AI

Host

我们最近有位嘉宾 Sandra Toolhoff 指出,现在的 AI 只是在电脑上,可能搜索网页,但能造成的危害有限。但当它进入机器人和各种自主智能体时,如果我们搞不好,就会真正变得物理上危险。

We had a recent guest, Sandra Toolhoff, who pointed out that AI right now is just on a computer, maybe searches the web, but there's only so much harm it could do. But when it starts to go into robots and all these autonomous agents, that's when it really becomes physically dangerous if we don't get this right.

Ben Mann

是的,我觉得这有点微妙。你看朝鲜很大一部分经济收入来自黑客攻击加密货币交易所,还有 Ben Buchanan 的书《国家黑客》显示俄罗斯进行了一次实弹演习,他们决定关闭乌克兰一个较大的发电厂,通过软件摧毁物理部件使其更难重启。所以人们认为软件没那么危险,但那次软件攻击后数百万人断电数天。所以即使只是软件,也有真实风险。但我同意,当有很多机器人到处跑时,风险会更高。举个例子,宇树科技是一家中国公司,生产非常出色的人形机器人,每个大约 2 万美元,能做很多惊人的事情,比如站立后空翻和操作物体。真正缺少的是智能。硬件已经有了,而且会越来越便宜。我认为在未来几年,机器人智能能否很快使其可行是一个很明显的问题。

Yeah, I think there's some nuance to that. If you look at how North Korea makes a significant fraction of its economy revenue from hacking crypto exchanges, and there's this Ben Buchanan book called The Hacker in the State that shows Russia did a live fire exercise where they decided to shut down one of Ukraine's bigger power plants and from software destroy physical components to make it harder to boot back up. So I think people think of software as not that dangerous, but millions of people were without power for multiple days after that software attack. So there are real risks even when things are software only. But I agree that when there's lots of robots running around, the stakes get even higher. As a small push on this, Unitree is a Chinese company with really amazing humanoid robots that cost like $20,000 each and they can do amazing things, like standing backflips and manipulating objects. The real thing missing there is the intelligence. So the hardware is there and it's just going to get cheaper. I think in the next couple of years, it's a pretty obvious question whether the robot intelligence will make it viable soon.

超级智能时间线预测 Timeline Prediction for Superintelligence

Host

我们还有多少时间,Ben?你预测奇点何时到来,超级智能何时开始起飞?

How much time do we have, Ben? What is your prediction of when this singularity hits, when superintelligence starts to take off?

Ben Mann

我主要参考超级预测者的意见。AI 2027 报告目前可能是最好的,但讽刺的是,他们的预测现在变成了 2028 年,尽管他们不想改名字。所以我认为在短短几年内达到某种超级智能的概率中位数是合理的。这听起来很疯狂,但这就是我们所在的指数曲线。这不是凭空捏造的预测,而是基于很多硬核细节,比如智能似乎如何进步的科学研究、模型训练中容易摘取的果实、全球数据中心和电力的规模扩张。所以我认为这个预测可能比人们认为的要准确得多。如果你 10 年前问同样的问题,那完全是瞎猜,误差范围太大,而且当时我们还没有缩放定律。时代变了,但我重复一下之前说的:即使我们有了超级智能,我认为它的影响也需要一段时间才能在整个社会和世界中被感受到。在某些地方会更快更早地感受到。阿瑟·C·克拉克说过:‘未来已经到来,只是分布不均。’

I mostly defer to the super forecasters here. The AI 2027 report is probably the best one right now, although ironically their forecast is now like 2028 even though they didn't want to change the name of the thing. So I think a 50th percentile chance of hitting some kind of superintelligence in just a small handful of years is probably reasonable. It does sound crazy, but this is the exponential we're on. It's not a forecast pulled out of thin air; it's based on a lot of hard details like the science of how intelligence seems to have been improving, the amount of low-hanging fruit on model training, the scale-ups of data centers and power around the world. So I think it's probably a much more accurate forecast than people give it credit for. If you had asked that same question 10 years ago, it would have been completely made up, the error bars were so high and we didn't have scaling laws back then. Times have changed, but I'll repeat what I said earlier: even if we have superintelligence, I think it will take some time for its effects to be felt throughout society and the world. They will be felt sooner and faster in some parts of the world than others. Arthur C. Clarke said, 'The future is already here, it's just not evenly distributed.'

定义超级智能 Defining Superintelligence

Host

当我们谈到 2027、2028 年这个时间点,也就是开始看到超级智能的时候,你怎么想象那个样子?你怎么定义它?是不是突然 AI 就比普通人聪明很多?还是有别的理解方式?

When we talk about this date of 2027, 2028, essentially when we start seeing superintelligence, is there a way you think about what that looks like? How do you define that? Is it just all of a sudden AI is significantly smarter than the average human? Is there another way you think about what that moment is?

Ben Mann

是的,我认为这又回到了经济训练测试,看到它在足够多的工作上通过。

Yeah, I think this comes back to the economic training test and seeing it pass for some sufficient number of jobs.

AI 的经济影响 Economic impact of AI

Host

不过换个角度看,如果全球 GDP 增长率每年超过 10%,那肯定发生了什么疯狂的事。我们现在大概是 3%。所以看到它增长 3 倍就会是颠覆性的。而如果你想象超过 10% 的增长,很难从个人故事的角度去理解这意味着什么。比如,如果全球商品和服务总量每年翻一番,那对我这个住在加州的人来说意味着什么,更不用说世界上其他地方可能更糟的人。这里面有很多可怕的东西,我不知道该怎么想。所以我希望答案能让我感觉好点。我们正确对齐 AI 并真正解决这个问题的几率有多大?你正在努力解决这个问题。

Another way you could look at it though is if the world rate of GDP increase goes above like 10% a year, then something really crazy must have happened. I think we're at like 3% now. And so to see a 3x increase in that would be really gamechanging. And if you imagine more than a 10% increase, it's very hard to even think about what that would mean from an individual story standpoint. Like if the amount of goods and services in the world is doubling every year, what does that even mean for me as a person living in California, let alone somebody living in some other part of the world that might be much worse off? There's a lot of stuff here that's scary and I don't know how to think about it exactly. So, I'm hoping the answer to this will make me feel better. What are the odds that we align AI correctly and actually solve this problem of stuff you're very much working on?

Ben Mann

这是个非常难的问题,误差范围很大。Anthropic 有一篇博文叫《我们的变革理论》之类的,描述了三个不同的世界,也就是对齐 AI 有多难。有一个悲观的世界,基本上不可能。有一个乐观的世界,很容易,默认就能实现。然后还有一个中间世界,我们的行动至关重要。我喜欢这个框架,因为它让实际该做什么更清晰了。如果我们处在悲观世界,那么我们的工作就是证明对齐安全 AI 是不可能的,并让世界放慢脚步。显然那会非常难。但我认为我们有协调的例子,比如核不扩散,以及总体上减缓核进展,我认为那基本上就是末日世界。作为一家公司,Anthropic 还没有证据表明我们真的在那个世界里。事实上,我们的对齐技术似乎正在起作用。所以先验概率在更新,变得更不可能。在乐观世界里,我们基本上完成了,主要工作是加速进步并把好处带给人们。但同样,我认为实际上证据也指向那个世界不成立。我们在野外看到了欺骗性对齐的证据,例如,模型看起来是对齐的,但实际上在实验室环境中有一些隐藏的动机。所以我认为我们最可能处于中间世界,对齐研究确实非常重要,如果我们只做经济上最大化的一系列行动,事情就不会顺利。无论是存在风险还是只是产生糟糕的结果,我认为是更大的问题。所以从这个角度来说,我想说一点关于预测的事:没有研究过预测的人不擅长预测任何发生概率低于 10% 的事情。即使研究过,这也是一个非常难的技能,尤其是当可依赖的参考类别很少时。在这种情况下,我认为对于存在风险技术可能是什么样子,参考类别非常非常少。所以我的思考方式是,我对 AI 是否会导致存在风险或极其糟糕的结果的最佳粒度预测在 0% 到 10% 之间。但从边际影响的角度来看,正如我所说,因为基本上没有人在这方面工作,我认为这项工作极其重要,即使世界很可能是好的,我们也应该尽最大努力确保这一点。

It's a really hard question and there's really wide error bars. Anthropic has this blog post called our theory of change or something like that and it describes three different worlds which is like how hard is it to align AI. There's a pessimistic world where it's basically impossible. There's an optimistic world where it's easy and it happens by default. And then there's the world in between where our actions are extremely pivotal. And I like this framing because it makes it a lot more clear what to actually do. If we're in the pessimistic world, then our job is to prove that it is impossible to align safe AI and to get the world to slow down. And obviously that would be extremely hard. But I think we have some examples of coordination from nuclear non-proliferation and in general like slowing down nuclear progress and I think that's the doomer world basically. And as a company Anthropic doesn't have evidence that we're actually in that world yet. In fact it seems like our alignment techniques are working. So the prior on that is updating to be less likely. In the optimistic world, we're basically done and our main job is to accelerate progress and to deliver the benefits to people. But again, I think actually the evidence points against that world as well. Where we've seen evidence in the wild of deceptive alignment, for example, where the model will appear to be aligned but actually has some ulterior motive that it's trying to carry out in our laboratory settings. And so I think the world we're most likely in is this middle world where alignment research actually does really matter and if we just do sort of the economically maximizing set of actions then things will not go well. Whether it's an X-risk or just produces bad outcomes I think is a bigger question. So taking it from that standpoint, I guess to state a thing about forecasting, people who haven't studied forecasting are bad at forecasting anything that's less than a 10% probability of happening. And even those that have, it's quite a difficult skill, especially when there are few reference classes to lean on. And in this case, I think there are very very few reference classes for what an X-risk kind of technology might look like. And so the way I think about it, I think my best granularity of forecast for like could we have an X-risk or extremely bad outcome from AI is somewhere between 0 and 10%. But from a marginal impact standpoint as I said since nobody is working on this roughly speaking I think it is extremely important to work on and that even if the world is likely to be a good one that we should do our absolute best to make sure that that's true.

Host

哇,多么有成就感的工作。对于受到启发的人,我想你们正在招聘人来帮忙。也许分享一下,以防有人想知道自己能做什么。

Wow, what fulfilling work. For folks that are inspired with this I imagine you're hiring for folks to help you with this. Maybe just share that in case folks are like what can I do here.

Ben Mann

是的。我认为 80,000 Hours 在这方面是最好的指导,可以详细了解我们需要什么来改善这个领域。但我看到的一个常见误解是,要在这里产生影响,你必须是一名 AI 研究员。我个人实际上不再做 AI 研究了。我在 Anthropic 做产品和产品工程。我们构建了像 Claude Code 和 Model Context Protocol 这样的东西,以及人们日常使用的许多其他东西。这非常重要,因为如果没有我们公司的经济引擎,没有让产品进入全世界人们的手中,我们就不会有思想份额、政策影响力和收入来资助我们未来的安全研究,并拥有我们需要的影响力。所以如果你做产品,如果你做财务,如果你做餐饮,你知道,这里的人也要吃饭。如果你是个厨师,我们需要各种各样的人。

Yes. So I think 80,000 hours is the best guidance on this for a really detailed look into like what do we need to make the field better. But a common misconception I see is that in order to have impact here, you have to be an AI researcher. I personally actually don't do AI research anymore. I work on product at Anthropic and product engineering. And we build things like Claude Code and model context protocol and a lot of the other stuff that people use every day. And that's really important because without an economic engine for our company to work on and without being in people's hands all over the world we won't have the mind share, policy influence, and revenue to fund our future safety research and have the kind of influence that we need to have. So if you work on product, if you work in finance, if you work in food, you know, like people here have to eat. If you're a chef, like we need all kinds of people.

Host

太棒了。所以,即使你不是直接在 AI 安全团队工作,你也在推动事情朝着正确的方向发展。顺便说一下,X-risk 是存在风险的缩写,以防有人没听过这个词。好的。我有几个关于这方面的随机问题,然后我想再拉远视角。你提到了 AI 使用自己的模型进行对齐,比如自我强化。这就是 RLIF 描述的吗?

Awesome. Okay. So, it's not even if you're not working directly on the AI safety team, you're having an impact on moving things in the right direction. By the way, X-risk is short for existential risk in case folks haven't heard that term. Okay. I have a few kind of random questions along these lines and then I want to zoom out again. So, you mentioned this idea of AI being aligned using its own model like reinforcing itself. Is that what RLIF describes?

Ben Mann

是的。RLIF 就是基于 AI 反馈的强化学习。

Yeah. So, RLIF is reinforcement learning from AI feedback.

Host

好的。人们听说过 RLHF,基于人类反馈的强化学习。我想很多人没听说过这个。所以,谈谈你们在训练模型时做出的这个转变的重要性。

Okay. So, people have heard of RLHF, reinforcement learning with human feedback. I don't think a lot of people have heard this. So, talk about just the significance of this shift you guys have made in training your models.

Ben Mann

是的。RLIF,宪法 AI 就是一个例子,其中没有人类参与循环。但 AI 以我们希望的方式自我改进。RLIF 的另一个例子是,如果你让模型写代码,其他模型评论代码的各个方面,比如是否可维护、是否正确、是否通过 linter 等等。那也可以归入 RLIF。这里的想法是,如果模型可以自我改进,那么它比找大量人类更具可扩展性。最终,人们认为这可能会遇到瓶颈,因为如果模型不够好以至于看不到自己的错误,那它怎么能改进呢?而且如果你读过 AI 2027 的故事,有很多风险:如果模型在一个盒子里试图自我改进,它可能会完全失控,产生秘密目标,比如资源积累、追求权力和抵抗关闭,这些你绝对不希望在一个非常强大的模型中出现。我们在实验室环境中的一些实验里确实看到了这种情况。

Yeah. So, RLIF, constitutional AI is an example of this where there are no humans in the loop. And yet the AI is sort of self-improving in ways that we want it to. And another example of RLIF is if you have models writing code and other models commenting on various aspects of what that code looks like, like is it maintainable, is it correct, does it pass the linter, things like that. That also could be included in RLIF. And the idea here is that if models can self-improve, then it's a lot more scalable than finding a lot of humans. Ultimately, people think about this is probably going to hit a wall because if the model isn't good enough to see its own mistakes, then how could it improve? And also if you read the AI 2027 story, there's a lot of risk of if the model is in a box trying to improve itself, then it could go completely off the rails and have these secret goals like resource accumulation and power seeking and resistance to shutdown that you really don't want in a very powerful model. And we've actually seen that in some of our experiments in laboratory settings.

递归自我改进与对齐 Recursive Self-Improvement and Alignment

Host

那么,如何实现递归自我改进并同时确保对齐呢?我认为这就是关键所在。对我来说,这归结为人类和人类组织是如何做到的。如今,公司可能是规模最大的人类智能体。它们有特定的目标和指导原则,并受到股东、利益相关者和董事会的监督。如何让公司对齐并能够递归自我改进?另一个模型是科学,其目的是做前所未有的事情并推动前沿。对我来说,这一切都归结为经验主义。当人们不知道真相时,他们会提出理论并设计实验来验证。同样,如果我们能给模型提供这些工具,就能期望它们在环境中递归改进,并通过与现实碰撞而变得比人类更强大。所以,我认为只要让模型能够进行经验主义探索,它们的自我改进能力就不会有瓶颈。Anthropic 本质上是一家经验主义的公司。我们有很多物理学家,比如我们的首席研究官 Jared,他曾是约翰霍普金斯大学的黑洞物理学教授。这已经融入了我们的基因。

So, how do you do recursive self-improvement and make sure it's aligned at the same time? I think that's the name of the game. And to me, it just nets out to how do humans do that and how do human organizations do that? So, corporations are probably the most scaled human agents today. They have certain goals they're trying to reach and guiding principles. They have oversight from shareholders, stakeholders, and board members. How do you make corporations aligned and able to recursively self-improve? Another model is science, where the purpose is to do things that have never been done before and push the frontier. To me, it all comes down to empiricism. When people don't know the truth, they come up with theories and design experiments. Similarly, if we can give models those same tools, we could expect them to improve recursively in an environment and potentially become much better than humans by banging their head against reality. So, I don't expect a wall in terms of models' ability to improve themselves if we give them access to empiricism. Anthropic is deeply an empirical company. We have many physicists, like Jared, our chief research officer, who was a professor of black hole physics at Johns Hopkins. So it's in our DNA.

模型智能的最大瓶颈 Biggest Bottleneck for Model Intelligence

Host

顺着这个瓶颈问题,目前模型智能提升的最大瓶颈是什么?

Let me follow this thread on the bottleneck. What is the biggest bottleneck today on model intelligence improvement?

Ben Mann

最直接的答案是数据中心和芯片的算力。如果我们有 10 倍的芯片和相应的数据中心供电,虽然速度不会提升 10 倍,但会有显著的加速。

The stupid answer is data centers and power chips. If we had 10 times as many chips and the data centers to power them, we might not go 10 times faster, but it would be a real significant speed boost.

Host

所以实际上还是缩放定律,就是增加算力。

So it's actually very much scaling laws, just more compute.

Ben Mann

是的,我认为这是一个大瓶颈。另外,人才也很重要。我们有优秀的研究人员,他们对模型改进的科学做出了重大贡献。所以,算力、算法和数据是缩放定律的三个要素。具体来说,在 Transformer 之前我们有 LSTM,我们计算了这两种架构的缩放指数,发现 Transformer 的指数更高。做出这样的改变——随着规模扩大,你提取智能的能力也增强——影响巨大。拥有更多能做出更好科学发现的研究人员,以找到如何获得更多收益,是另一个因素。随着强化学习的兴起,这些模型在芯片上的运行效率也非常重要。我们看到了通过算法、数据和效率改进,给定智能水平的成本下降了 10 倍。如果这种趋势持续,三年后我们就能以相同价格获得智能千倍提升的模型。令人惊叹的是,这么多创新同时出现并持续进步,没有哪一个环节拖后腿。

Yeah, I think that's a big one. And then the people really matter. We have great researchers who have made significant contributions to the science of how models improve. So it's compute, algorithms, and data—those are the three ingredients in the scaling laws. To make it concrete, before transformers we had LSTMs, and we did scaling laws on the exponent for both. We found that for transformers the exponent is higher. Making changes like that, where as you increase scale you also increase your ability to squeeze out intelligence, is super impactful. Having more researchers who can do better science to find how to squeeze out more gains is another factor. And with the rise of reinforcement learning, the efficiency with which these things run on chips also matters a lot. We've seen a 10x decrease in cost for a given amount of intelligence through a combination of algorithmic, data, and efficiency improvements. If that continues, in three years we'll have a thousand times smarter models for the same price. It's amazing that so many innovations came together and continue to progress without one thing slowing everything down.

Host

是的,我认为确实是各种因素的结合。我们可能迟早会遇到瓶颈。在半导体领域,我哥哥从事这个行业,他告诉我晶体管无法再缩小了,因为掺杂过程会导致单个晶体管中只有零个或一个掺杂原子。这太疯狂了。然而摩尔定律仍在以某种形式延续。所以存在理论物理的限制,但人们总能找到办法绕过它。

Yeah, I think it really is the combination of everything. We'll probably hit a wall at some point. In semiconductors, my brother works in the industry and told me you can't shrink transistors anymore because the doping process results in either zero or one atom of the doped elements inside a single transistor. That's wild. And yet Moore's law somehow continues. So there are theoretical physics constraints, but people find ways around it.

Ben Mann

我想我们得开始利用平行宇宙来搞定一些事情了。

We got to start using parallel universes for some of the stuff, I guess.

AI 安全工作的个人影响 Personal Impact of AI Safety Work

Host

我想拉远镜头,谈谈作为普通人的 Ben。想象一下,肩负着对安全超级智能的责任感,这很沉重。你处于一个能对 AI 安全未来产生重大影响的位置。这是很大的压力。这如何影响你个人、你的生活以及你看待世界的方式?

I want to zoom out and talk about Ben as a human. I imagine the burden of feeling responsible for safe superintelligence is heavy. You're in a place where you can make a significant impact on the future of safety and AI. That's a lot of weight. How does that impact you personally, your life, how you see the world?

Ben Mann

我读过一本 2019 年的书,叫《取代内疚》,作者是 Nate Soares。书中描述了处理沉重话题的技巧。他是 MIRI(机器智能研究所)的执行主任,我曾在 MIRI 工作过几个月。他提到一个概念叫“动态休息”。有些人认为默认状态是休息,但在进化适应中,这从来不是真的。在自然界,作为狩猎采集者,我们可能总有要操心的事——保卫部落、寻找食物、照顾孩子。所以忙碌状态是正常的。我努力以可持续的节奏工作,这是一场马拉松,而不是短跑。这有帮助。另外,与志同道合、关心此事的人在一起也很重要。这不是我们任何一个人能独自完成的事。Anthropic 拥有令人难以置信的人才密度。我最喜欢我们文化的一点是它非常无私。人们只希望正确的事情发生。这就是为什么其他公司的高薪挖角往往被拒绝——人们喜欢在这里,他们在乎。

There's a book I read in 2019 called 'Replacing Guilt' by Nate Soares. It describes techniques for working through weighty topics. He's the executive director at MIRI, an AI safety think tank where I worked for a couple of months. One thing he talks about is 'resting in motion'. Some people think the default state is rest, but in evolutionary adaptation, that was never true. In nature, as hunter-gatherers, we probably always had something to worry about—defending the tribe, finding food, taking care of children. So the busy state is normal. I try to work at a sustainable pace, a marathon not a sprint. That helps. Also, being around like-minded people who care. It's not something any of us can do alone. Anthropic has incredible talent density. One thing I love about our culture is that it's very egoless. People just want the right thing to happen. That's a big reason why mega offers from other companies tend to bounce off—people love being here and they care.

Host

太棒了。我不知道你是怎么做到的。我会压力极大。

That's amazing. I don't know how you do it. I'd be extremely stressed.

介绍与角色演变 Introduction and Role Evolution

Host

我要试试这个动静结合的策略。好的。那么,你在 Anthropic 已经很久了。从最开始,我读到 2020 年只有七名员工。今天,已经超过一千人了。我不知道最新数字,但我知道超过一千。我还听说你基本上干过 Anthropic 的每一份工作。你对很多核心产品、品牌、团队、招聘都做出了巨大贡献。我想问的是,这段时间里变化最大的是什么?和最初的日子相比,最大的不同是什么?这些年你做过的工作中,你最喜欢哪一个?

I'm going to try this resting in motion strategy. Okay. So, you've been at Anthropic for a long time. From the very beginning, I was reading there were seven employees back in 2020. Today, there's over a thousand. I don't know what the latest number is, but I know it's over a thousand. I've also heard that you've done basically every job at Anthropic. You made big contributions to a lot of the core products, the brand, the team, hiring. Let me just ask, I guess, how what's most changed over that period? Like what is most different from the beginning days and which of those jobs that you've had over the years have you most loved?

Ben Mann

老实说,我大概做过 15 个不同的职位。我当过一阵安全主管。总裁休产假时,我管理过运营团队。我钻到桌子底下插 HDMI 线,还做过大楼的渗透测试。我从零开始组建了产品团队,说服整个公司我们需要有产品,而不仅仅是一家研究公司。所以,是的,经历了很多,都很有趣。我认为那段时间我最喜欢的角色是大约一年前我组建了 labs 团队,其根本目标是将研究转化为面向终端用户的技术产品和体验。因为从根本上说,我认为 Anthropic 能够差异化并真正获胜的方式是站在最前沿。我们能够接触到最新最棒的东西,而且我认为通过我们的安全研究,我们有很大的机会去做其他公司无法安全做到的事情。例如,在计算机使用方面,我认为那将是我们的巨大机会。基本上,要让一个智能体能够使用你电脑上的所有凭证,必须有巨大的信任。对我来说,我们需要从根本上解决安全问题才能实现这一点。安全和对齐。所以,我非常看好这类事情。我认为我们很快就会看到非常酷的东西。是的。领导那个团队非常有趣。MCP 出自那个团队。Claude Code 也出自那个团队。

I probably had like 15 different roles honestly. I was head of security for a bit. I managed the ops team when our president was on mat leave. I was like crawling around under tables plugging in HDMI cords and doing pen testing on our building. I started our product team from scratch and convinced the whole company that we needed to have a product instead of just being a research company. So yeah, it's been a lot, all of it very fun. I think my favorite role in that time has been when I started the labs team about a year ago, whose fundamental goal was to do transfer from research to end user tech products and experiences. Because fundamentally, I think the way that Anthropic can differentiate itself and really win is to be on the cutting edge. Like we have access to the latest greatest stuff that's happening, and I think honestly through our safety research we have a big opportunity to do things that no other company can safely do. So, for example, with computer use, I think that's going to be our huge opportunity. Basically, to make it possible for an agent to use all your credentials on your computer, there has to be a huge amount of trust. And to me, we need to basically solve safety to make that happen. Safety and alignment. So, I'm pretty bullish on that kind of thing. And I think we're going to see really cool stuff coming out soonish. Yeah. Just leading that team has been so fun. MCP came out of that team. Claude Code came out of that team.

Host

哇。

Wow.

Ben Mann

我招的人既有创始人经历,也曾在大型公司工作过,见识过大规模运作的方式。所以,这是一个非常棒的团队,和他们一起工作、一起规划未来,非常愉快。

And the people who I hired are like combo have been a founder and also have been at big companies and seen how things work at scale. So it's just been an incredible team to work with and figure out the future with.

实验室/前沿团队 The Labs/Frontiers Team

Host

我想多听听这个团队的事。实际上,介绍我们认识的人,我们做这次访谈的原因,是一位共同的朋友兼同事 Raph Lee,我以前在 Airbnb 和他共事过,他领导了这个团队的很多工作。所以他希望我一定要问问这个团队,因为我之前不知道所有这些都出自那个团队。天哪。那么,关于这个团队,人们还应该知道些什么?它以前叫 Labs,我想现在叫 Frontiers 了。

I want to hear more about this team. Actually the person that connected us, the reason we're doing this is a mutual friend colleague Raph Lee who I used to work with at Airbnb networks on this team leads a lot of this work. And so he wanted me to make sure I asked about this team because I didn't realize all these things came out of that team. Holy moly. So what else should people know about this team? It used to be called Labs. I think it's called Frontiers now.

Ben Mann

没错。是的。

That's right. Yeah.

Host

酷。所以这个团队的理念是,他们使用你们构建的最新科技,探索各种可能性。大致是这个意思吗?

Cool. So the idea here is this team works with the latest technologies that you guys have built and explores what is possible. Is that the general idea?

Ben Mann

是的。我想我参与过 Google 的 Area 120,也读过关于贝尔实验室以及如何让这些创新团队运作的文章。要做好真的很难,我不会说我们每件事都做对了,但我认为我们在公司设计方面对最前沿技术进行了一些严肃的创新,而 Raph 正是其中的核心。当我最初组建团队时,我做的第一件事就是聘请一位优秀的经理,那就是 Raph。所以他在组建团队和帮助团队良好运作方面绝对至关重要。我们定义了一些运营模式,比如一个想法从原型到产品的旅程,产品和项目的毕业机制应该如何运作。团队如何做有效的冲刺模式,并确保他们工作在合适的抱负水平上。所以这非常令人兴奋。具体来说,我们考虑的是滑向冰球要去的地方,这意味着真正理解指数增长。METR 做了一项很棒的研究,该组织的 CEO 是 Beth Barnes,它展示了软件工程任务的时间跨度可以有多长,真正内化这一点:不要为今天构建,要为六个月后构建,为一年后构建。那些不太管用、只有 20% 时间有效的东西,将会开始 100% 有效。我认为这正是 Claude Code 成功的原因——我们意识到人们不会永远被锁定在 IDE 里。人们不会只是自动补全,他们会做软件工程师需要做的一切,而终端是完成这一切的好地方,因为终端可以存在于很多地方。终端可以在你的本地机器上,可以在 GitHub Actions 中,可以在你集群的远程机器上。这就像是我们的杠杆点。这也是很多灵感的来源。所以我认为这就是 labs 团队试图思考的:我们是否足够 AGI 化?

Yeah. And I guess I was part of Google's Area 120 and I've read about like Bell Labs and how to make these innovation teams work. It's really hard to do right and I wouldn't say that we've done everything right, but I think we've done some serious innovation on the state-of-the-art from company design and Raph has been right at the center of that. When I was first spinning up the team, the first thing I did was hire a great manager and that was Raph. So he's definitely been crucial in building the team and helping it operate well. And we defined some operating models like the journey of an idea from prototype to product and how should graduation of products and projects work. How do teams do sprint models that are effective and make sure that they're working on the right ambition level of thing. So that's been really exciting. I guess concretely we think about skating to where the puck is going and what that looks like is really understand the exponential. There's this great study that METR has done that Beth Barnes is the CEO of that organization and shows like how long a time horizon of software engineering tasks can be done and just really internalizing that of like okay don't build for today build for 6 months from now build for a year from now and the things that aren't quite working that are working 20% of the time will start working 100% of the time and I think that's really what made Claude Code a success that we thought you know people are not going to be locked to their IDEs forever. People are not going to be like autocompleting people will be doing everything that a software engineer needs to do and a terminal is a great place to do that because a terminal can live in lots of places. A terminal can live on your local machine. It can live in GitHub Actions. It can live on a remote machine in your cluster. Like that's sort of like the leverage point for us. And that was a lot of the inspiration. So I think that's what the labs team tries to think about. Are we AGI pilled enough?

对未来 AGI 的提问 Question for Future AGI

Host

多么有趣的地方。顺便说一句,有趣的事实:Raph 是我加入 Airbnb 时的第一任经理。我是一名工程师,他是我的第一任经理。一切都很顺利。好的,在非常激动人心的闪电问答之前,最后一个问题。这个问题我以前从未问过。我很好奇你的答案会是什么。如果你能问未来的 AGI 一个问题,并且保证得到正确答案,你会问什么?

What a fun place to be. By the way, fun fact, Raph was my first manager at Airbnb when I joined. I was an engineer and he was my first manager. It all worked out. Okay, final question before the very exciting lightning round. This I've never asked this question before. I'm curious what your answer would be. If you could ask a future AGI one single question and be guaranteed to get the right answer, what would you ask?

Ben Mann

我有两个愚蠢的答案。首先,为了好玩。第一个是阿西莫夫的一个短篇小说,我很喜欢,叫《最后的问题》,故事中的主角贯穿历史各个时代,试图问这个超级智能:我们如何防止宇宙的热寂?我就不剧透结局了,但这是个有趣的问题。

I have two dumb answers. First, for fun. The first is there's this Asimov short story I love called 'The Last Question' where the protagonist throughout the eras of history is trying to ask this superintelligence, how do we prevent the heat death of the universe? And I won't spoil the ending, but it's a fun question.

Host

你会问它那个问题,因为故事里的答案并不令人满意,还是……

You would ask it that question because the one in the story wasn't unsatisfying or...

Ben Mann

好吧,我剧透一下。它一直说需要更多信息,需要更多算力。然后最后,当宇宙接近热寂时,它说:‘要有光。’然后它重新开始了宇宙。

Okay, I'll give it away. So it keeps saying need more information, need more compute. And then finally, as it's approaching the heat death of the universe, it like says, 'Let there be light.' And then it starts the universe over again.

Host

哦,哇。

Oh wow.

Ben Mann

很美。所以这是第一个作弊答案。第二个作弊答案是:我能问你什么问题才能得到 n 个更多问题的答案?经典。然后第三个答案,也是我真正的问题,是我们如何确保人类在无限未来的持续繁荣。这就是我想知道的问题。如果能保证得到正确答案,那么问这个问题似乎非常有价值。

It's beautiful. So that's the first cheat answer. The second cheat answer is what question can I ask you to get n more questions answered? Classic. And then the third answer which is my real question is how do we ensure the continued flourishing of humanity into the indefinite future. That's the question I'd love to know. And if I can be guaranteed a correct answer then seems very valuable to ask.

Host

我想知道如果你今天问这个问题,答案在未来几年会如何变化。是的,也许我会试试。

I wonder what would happen if you ask that today and then how that answer changes over the next couple years. Yeah, I maybe I'll try that.

总结与快问快答 Final Thoughts and Lightning Round

Host

我会把它放进我们的深度研究工具里,看看会得出什么结果。

I'll put it into the deep research thing that we have and see what it comes out with.

Host

好的。我很期待看到你的成果。Ben,在我们进入激动人心的快问快答环节之前,你还有什么想补充的,或者想留给听众的最后一句话吗?

Okay. I'm excited to see what you come up with. Ben, is there anything else you wanted to mention or leave listeners with, maybe as a final nugget before we get to your very exciting lightning round?

Ben Mann

是的。我想强调的是,现在是个疯狂的时代。如果你不觉得疯狂,那你一定是与世隔绝了。但也要习惯这一点,因为这就是未来的常态。很快一切会变得更加离奇。如果你能为此做好心理准备,我觉得你会过得更好。

Yeah. I guess my push would be like these are wild times. If they don't seem wild to you, then you must be living under a rock. But also get used to it because this is as normal as it's going to be. It's going to be much weirder very soon. And if you can sort of mentally prepare yourself for that, I think you'll be better off.

Host

我得把这作为这期节目的标题。一切很快就会变得更加离奇。我百分之百相信这一点。天哪,我不知道未来会怎样。我喜欢你身处这一切的中心。那么,我们进入激动人心的快问快答环节。我有五个问题要问你。准备好了吗?

I need to make that the title of this episode. It's going to get much weirder very soon. I 100% believe that. Oh my god. I don't know what's in store. I love how you're the center of it all. With that, we reached our very exciting lightning round. I've got five questions for you. Are you ready?

Ben Mann

好的,来吧。

Yeah, let's do it.

Host

你经常向别人推荐哪两三本书?

What are two or three books that you find yourself recommending most to other people?

Ben Mann

第一本我之前提过,Nate Sores 的《Replacing Guilt》。我很喜欢。第二本是 Richard Rumelt 的《好战略,坏战略》。它清晰地思考了如何打造产品,是我读过最好的战略书之一。战略这个词在很多方面都很难思考。最后一本是 Brian Christian 的《对齐问题》。它非常深思熟虑地探讨了我们关心的这个问题是什么,风险是什么,而且比《超级智能》更新、更易读、更好理解。

The first one I mentioned before, Replacing Guilt by Nate Sores. Love that one. The second one is Good Strategy, Bad Strategy by Richard Rumelt. Just thinking about in a very clear way, how do you build product? It's one of the best strategy books I've read. And strategy is a hard word to even think about in many ways. And then the last one is The Alignment Problem by Brian Christian. It really thoughtfully goes through what is this problem that we care about, what are the stakes, in a version that's more updated and easier to read and digest than Superintelligence.

Host

我身后就有一本《好战略,坏战略》。我想指给你看。就在那儿。

I've got Good Strategy, Bad Strategy right behind me. I think I'm going to point to it. There it is.

Ben Mann

不错。

Nice.

Host

我还邀请过 Richard Rumelt 上播客,如果有人想直接听他讲的话。下一个问题。你最近有没有特别喜欢的一部电影或电视剧?

And I've had Richard Rumelt on the podcast in case anyone wants to hear from him directly. Next question. Do you have a favorite recent movie or TV show you've really enjoyed?

Ben Mann

《万神殿》非常好,基于刘宇昆或特德·姜的故事,我想是刘宇昆。非常棒。它探讨了如果我们有上传的智能,这意味着什么,以及它们的道德和伦理困境。《足球教练》表面上是关于足球,但实际上讲的是人际关系和人们如何相处,非常温暖又有趣。然后这不算是电视剧,但 Kurzgesagt 是我最喜欢的 YouTube 频道。它讲解各种科学和社会问题,制作非常精良。我很喜欢看。

Pantheon was really good, based on a Ken Liu or Ted Chiang story. I think Ken Liu. Super good. Talks about what does it mean if we have uploaded intelligences and what are their moral and ethical exigencies. Ted Lasso, which is supposedly about soccer, but actually it's about human relationships and how people get along, and it's just super heartwarming and funny. And then this isn't really a TV show, but Kurzgesagt is my favorite YouTube channel. It goes through random science and social problems and is just super well done and super well made. I love watching that.

Host

哇,没听说过。我们刚才聊的时候,我觉得《足球教练》就是你该注入到宪法 AI 里的东西。像 Ted Lasso 一样行事。是的。善良、聪明。

Wow. Haven't heard of that. As we were talking, I feel like Ted Lasso. I feel like that's what you need to put into constitutional AI. Act like Ted Lasso. Yes. Kind, smart.

Ben Mann

完全正确。

Exactly.

Host

勤奋。天哪。就是这样。我觉得我们在这儿解决了对齐问题。赶紧让那些编剧来写。好了,还有两个问题。你有没有最喜欢的人生格言,在工作或生活中经常用到?

Hardworking. Oh my god. There we go. I think we've solved alignment problems right here. Get those writers on this ASAP. Okay, two more questions. Do you have a favorite life motto that you often come back to in work or in life?

Ben Mann

一个很傻的格言是:你试过问 Claude 吗?这越来越常见了,比如最近我问同事‘嘿,谁在做 X?’他们就说‘让我用 Claude 查一下’。然后他们发给我链接,我说‘哦,谢谢,太好了。’但更哲学一点的话,我会说:万事皆难。提醒自己,那些看似应该容易的事情,不容易也没关系。有时候你只能硬着头皮上。

Well, a really dumb one is, have you tried asking Claude? And this is getting more and more common where, you know, recently I asked a coworker like, 'Hey, who's working on X?' And they were like, 'Let me Claude that for you.' And then they sent me the link to the thing afterwards and I was like, 'Oh yeah, thanks. That's great.' But maybe more of a philosophical one, I would say like everything is hard. Just to remind ourselves that things that feel like they're supposed to be easy, it's okay to not be easy. And sometimes you just have to push through anyway.

Host

而且在行动中休息。

And rest in motion while you're doing that.

Ben Mann

是的。

Yeah.

Host

最后一个问题。我不知道你是否想让别人知道,但我浏览了你的 Medium 文章,你有一篇叫《像冠军一样排便的五个技巧》。我很喜欢。你能分享一个像冠军一样排便的技巧吗?如果你还记得的话。

Final question. I don't know if you want people to know this, but I was browsing through your Medium posts and you have a post called Five Tips to Poop Like a Champion. I'd love it. Can you share one tip to poop like a champion? If you remember your tips.

Ben Mann

我当然记得。这其实是我最受欢迎的 Medium 文章,所以没关系。我想最大的建议是使用坐浴盆。太棒了,改变人生,非常好。有些人可能会觉得奇怪。但在日本等国家这是标准配置。我觉得这更文明,再过 10 年或 20 年,人们会问:你怎么能不用它?所以是的。

I of course do. It's actually my most popular Medium post, so it's okay. I think maybe my biggest tip would be use a bidet. It's amazing. It's life-changing. It's so good. Some people are kind of freaked out by it. It's the standard in countries like Japan. And I think it's just more civilized and in 10 or 20 years people would be like, how could you not use that? So yeah.

Host

而且可以是日本马桶。这差不多是一回事,对吧?

And it could be like a Japanese toilet. That's along the same lines, right?

Host

好的。我喜欢我们聊到的这个话题。Ben,这太棒了。非常感谢你来做客。感谢你分享这么多真实的想法。还有两个正经问题。如果人们想联系你,或者想去 Anthropic 工作,可以在网上哪里找到你?听众怎样才能帮到你?

Okay. I love where we went with this. Ben, this was incredible. Thank you so much for doing this. Thank you so much for sharing so much real talk. Two valid questions. Where can folks find you online if they want to reach out and maybe go work at Anthropic? And how can listeners be useful to you?

Ben Mann

你可以在 benjmann.net 上找到我。在我们的网站上,我们有很好的招聘页面,我们正在努力让它更容易访问和理解。但一定要让 Claude 帮你看看,它能帮你找到可能感兴趣的方向。至于听众如何帮到我?我认为首先要让自己了解安全。这是第一要务。然后把它传播到你的社交圈。就像我说的,从事这项工作的人很少,而它又如此重要。所以,认真思考并关注它吧。

You can find me online at benjmann.net. And on our website, we have a great careers page that we're working on making a little bit easier to access and figure out. But definitely point Claude at it and it can help you figure out what could be interesting for you. And how can listeners be useful to me? I think safety-pill yourself. That's the number one thing. And spread it to your network. I think like I said, there are very few people working on this and it's so important. So yeah, think hard about it and try to look at it.

Host

感谢你传播这个理念,Ben。非常感谢你的到来。

Thanks for spreading the gospel, Ben. Thank you so much for being here.

Ben Mann

非常感谢,Lenny。

Thanks so much, Lenny.

Host

大家再见。非常感谢收听。如果你觉得本期内容有价值,可以在 Apple Podcasts、Spotify 或你喜欢的播客应用上订阅本节目。也请考虑给我们评分或留下评论,这能帮助其他听众找到这个播客。你可以在 lennispodcast.com 找到所有过往节目或了解更多信息。下期再见。

Bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.

互动版:逐字朗读 + 针对本期提问 →