与达里奥·阿莫迪小酌一杯

A cheeky pint with Dario Amodei

达里奥·阿莫迪 Dario Amodei · Cheeky Pint · 2025-08-06 · 约 63 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

就着一杯啤酒:创办 Anthropic、Scaling,以及 AI 的经济学。

Over a beer: building Anthropic, scaling, and the economics of AI.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 27)

全文 · Full transcript(中英对照)

与兄弟姐妹创业 Starting a company with sibling

Host

我很兴奋终于能了解和兄弟姐妹一起创办公司是什么感觉。

I'm excited to finally learn what it is like to start a company with your sibling.

Dario

嗯,我不知道你为什么问我这个问题,因为你知道……模型想要学习。模型想要在市场上取得非凡的成功。

Well, I don't know why you're asking me that question because you know... it's like the models want to learn. The models want to be extraordinarily successful in the market.

Host

是的,没错。除了这种学习冲动,模型还有这种资本主义冲动。有时人们想到 API 业务,会说它粘性不大或者会……API 业务。我喜欢 API 业务。

Yes. Right. In addition to having this learning impulse, the models have this capitalistic impulse. Sometimes people think of the API business and say it's not very sticky or it's going to be... API business. I love API business.

Dario

不,不,正是如此。我认为我们将进入一个模型犯错频率远低于人类,但错误会更奇怪的世界。

No. No. Exactly. Exactly. I think we're going to be in a world where the models will make mistakes much less often than humans, but they'll be stranger mistakes.

Host

所以,我们需要为 Lens 发明含糊不清的表述。

So, we need to invent slurring for Lens.

Dario

那是不正确的,很糟糕。

And that's incorrectly poor.

Host

哦,哇。Dario 是 Anthropic 的 CEO,Anthropic 是当今前沿 AI 实验室之一。他从几年前的一名 AI 研究员,变成了现在经营着世界上增长最快的企业之一。干杯。所以,我很兴奋能谈谈 Anthropic 的业务。你学过物理和计算神经科学。

Oh, wow. Dario is CEO of Anthropic, one of today's frontier AI labs. He's gone from being an AI researcher just a few years ago to now running one of the world's fastest growing businesses. Cheers. So, I'm excited to talk a bit about the Anthropic business. You studied physics and computational neuroscience.

Dario

是的。

Yes.

Host

然后你在 BYU 工作过,接着是 Google Brain,然后是 OpenAI,最后创办了 Anthropic。

You then worked at BYU, then Google Brain, then OpenAI, and then started Anthropic.

Dario

是的。

Yes.

Host

我们会深入探讨 Anthropic 的业务,但我很兴奋终于能了解和兄弟姐妹一起创办公司是什么感觉。

And we'll get into Anthropic business, but I'm excited to finally learn what it is like to start a company with your sibling.

Dario

我也可以问你同样的问题,但你知道,经营公司几乎需要做两件事:你需要执行运营,也需要有好的战略,并看到最重要或别人看不到的东西。所以我的工作是后者,Daniela 的工作是前者,我们都擅长各自的事情。我认为这让我们每个人都能把大部分时间花在自己最擅长的事情上。

I could ask the same question of you, but you know, it's almost like there's two things you need to do when you're running a company. You need to operationally execute, and you need to have a good strategy and kind of see the most important thing or the thing that no one else sees. So my job is the second and Daniela's job is the first, and we're both good at the things we do. I think it's allowed us each to spend most of our time on the thing we're best at.

Host

大概还有信任方面的问题,科技界和 AI 界的联合创始人团队通常是不稳定的组合。有一个你长期深度信任的人很重要。

Presumably there's something about the trust side of things as well, where co-founder teams in general in tech, and AI as well, are unstable pairings. Just having someone where you have a long running and deep trust.

Dario

是的,拥有完全彻底的信任。我认为甚至不止于此,Anthropic 有七位联合创始人。当我们创办时,几乎所有人的建议都是七个联合创始人是一场灾难。公司很快就会分崩离析。每个人都会互相争斗。对于我决定给每个人相同股权,负面评价更多。但我们发现,我认为是因为显然我和 Danielle 是兄弟姐妹,而且我们七个人中,有些人认识很久了,或者有工作历史,不仅仅是认识,而是过去一起工作过。我认为这真的让我们始终步调一致。尤其是随着公司发展,有七个人真正承载公司的价值观并将其传递给更多人,这让你能够将公司扩展到更大规模,同时保持我们的价值观和团结。

Yeah. Where you have total and complete trust. I think even beyond that, Anthropic has seven co-founders. When we founded it, the advice from pretty much everyone was that seven co-founders is a disaster. The company will fall apart before you know it. Everyone will be fighting with each other. There was even more negativity on my decision to give everyone the same amount of equity. But what we found, and I think it was because obviously me and Danielle are siblings, but then all seven of us, some of us knew each other for a long time or had a history of work, not just knowing each other but working together in the past. I think that really allowed us to always be on the same page. Especially as the company grows, the idea that you have seven people who really carry the values of the company and project them to a wide set of people, it allows you to scale the company to a much larger size while holding on to the values and the unity we have.

Anthropic 业务与 AI 市场 Anthropic business and AI market

Host

我想问问 Anthropic 的业务,因为这又是一个不可思议的故事。最近有报道说你们的年化收入突破了 40 亿美元,所以关于你们正在开发的技术有很多正确的讨论,但这也是历史上增长最快的企业之一。我想谈谈 AI 市场。也许从这个问题开始:大家都在用 AI 做什么?比如有编程、客户服务工作,但所有这些收入来自哪里?

I want to ask about the Anthropic business because again it's an incredible story. It was reported recently that you'd blown through $4 billion in ARR, and so there's a lot of discussion correctly about the technology you're developing, but also this is just one of the fastest growing businesses in history. I want to talk a bit about the AI market. Maybe the place to start is: what is everyone doing with AI? Like there's coding, there's customer service work, but where does all this revenue come from?

Dario

是的,有各种各样的用途,而且随着时间推移有所变化。我可以说,增长最快的应用绝对是编程,尽管它远不是唯一的应用。我认为它增长如此之快的原因,除了我们专注于编程且模型擅长编程之外,实际上反映了社会扩散。如果我们看今天的 AI 模型,我认为在每个领域,它们的能力与实际部署之间都存在巨大差距,因为存在一些摩擦。大型企业的人不熟悉这项技术。我看银行或保险公司在做什么,即使模型不再进步,即使我们停止在模型之上构建产品,单个企业仍有巨大的数十亿美元潜力。我交谈过的公司 CEO 们通常非常理解这一点,但如果公司有 1 万或 10 万人,他们的运营方式已经固定,改变需要时间。但在编程领域,写代码的人在社交和技术上都与开发 AI 模型的人非常接近,所以扩散非常快。他们也是那种早期采用者,习惯新技术。所以我认为编程的巨大增长,最大的原因是从事编程的人和专注于编程的初创公司都是快速采用者,非常了解这项技术。但这绝不仅限于编程。如果你看,有很多公司做工具使用之类的事情。还有你提到的客户服务。我们与 Intercom 等公司密切合作。我们开始看到生物学方面的一些事情。我们与制药和医疗保健公司合作,也在基础科学研究方面努力。例如,我们与 Benchling 等公司合作,但也与一些非常大的制药公司合作。不久前,我们与诺和诺德合作撰写临床研究报告。临床研究报告就像你做了临床试验,然后写出结果:这些是不良事件,这些是统计数据。临床研究报告通常需要九周时间。Claude 可以在五分钟内完成,然后人类花几天时间检查。所以你真的可以看到加速的机会,随着模型变得更好,它们也会深入到深度研究中。

Yeah, there's a wide range of things and it's kind of changed over time. I would say definitely the application that has grown the fastest, although it's not the only application by far, is definitely coding. My theory on why it's grown so fast, other than that we focused on coding and the models are good at coding, is that it's really a statement about societal diffusion. If we look at today's AI models, I think in every area there's a huge overhang in terms of what they could do compared to how they're actually being deployed today, because there's some friction. People at large enterprises are not familiar with the technology. I look at what a bank does or what an insurance company does, and there's huge potential even if the model stopped getting better, even if we stop building products on top of the model, there's still huge billion dollar potential in an individual enterprise. Often the CEOs of companies that I talk to understand that perfectly well, but if the company is a 10,000 or 100,000 person company, they are set up to operationally do a certain thing a certain way, and it takes time to change them. But in code, the people who write code are very socially and technically adjacent to the folks who develop AI models, so the diffusion is very fast. They are also the kind of people who are early adopters, used to new technology. So I think the big growth in code, I would say the biggest cause is just that the people doing it and the startups devoted to it are fast adopters who understand the technology super well. But it's by no means limited to code at all. If you look at there are a bunch of companies that do things like tool use. There are as you mentioned customer service. We work closely with companies like Intercom. We're starting to see some things on the biology side. We're working both with pharmaceutical and healthcare companies, and we're working on the side of basic scientific research. We work with companies like Benchling for example, but we also work with some of the very large pharma companies. There was something done a while back where we worked with Novo Nordisk to write clinical study reports. Clinical study reports are like you've done a clinical trial and then you write up the results: these are the adverse events, these are the statistics. The clinical study report normally takes like nine weeks. Claude could do it in like five minutes, and then it took a human a few days to check it. So you can really see the opportunity for acceleration, and as the models get better, they'll reach into the deep research as well.

代码作为 AI 采用早期指标 Code as early indicator of AI adoption

Host

所以我想总结一下就是,代码领域领先,但我们看到大量其他用例的长尾,包括一些非常重要的用例。

So I guess a way to summarize it would be to say that code is out in the lead, but we see a long tail of quite a lot of other stuff including some very significant use cases.

Dario

我认为代码可能是一个早期指标,像是其他地方将要发生的事情的预兆。同样的指数增长,只是发生得更快。

I think code is maybe an early indicator, like a premonition of what's going to happen everywhere else. It's the same exponential, it's just happening faster.

Host

没错。所以很多领域都有显著的 AI 提升,但工程师习惯于采用。想想 Hacker News 上人们争论最佳工具,他们对此充满热情。我们发布 Claude Code 两小时后,就有人用它尝试了一万种不同的事情,并把它接入所有框架,Twitter 上两小时形成一种观点,两小时后又修正观点。想想这个速度,相比制药公司将其用于研究的速度,或者传统零售公司的速度。我们希望把一切带给所有人,世界上一些最大的好处涉及实体经济,我们想达到那里,但本质上它不会以同样的速度发生。

Right. So there are many places where there's significant AI uplift, but engineers are used to adopting. You think about Hacker News and people arguing over the best tools. People are passionate about it. And like two hours after we released Claude Code, there's some person out there who has tried 10,000 different things with it and plugged it into all the frameworks, and Twitter forms one opinion after two hours and then revises an opinion in two hours. You think of the speed of that compared to the speed that a pharmaceutical company can use it in research, or a traditional retail company. We want to bring everything to all, some of the biggest benefits in the world are touching the physical economy, and we want to get there, but it intrinsically does not happen at the same speed.

自研 vs 平台策略 First-party vs platform approach

Host

你如何决定哪些垂直领域自己来做,哪些交给平台?比如你们有 Claude Code,显然也有像 Windsurf 和 Cursor 这样的平台公司。你们推出了 Claude for Financial Services。大概还有其他垂直领域你们会说,我们不在那里构建工具。只是,嗯,你们有像 Claude for Enterprise 这样的东西,这不是一个垂直领域,而是面向企业的一般性策略。

How do you decide which verticals to do yourself versus which to allow platform? Like you have Claude Code, and obviously there are also platform companies like Windsurf and Cursor and everyone like that. You launched Claude for Financial Services. Presumably there are other verticals where you say, well, we're not building a tool there. Just yeah, you know, we have things like Claude for Enterprise, which is not a vertical but a general play to go with enterprise.

Dario

我认为我们喜欢这样思考:我们首先把自己看作一个平台公司。这里的类比可能是云服务之类的。如果你想象一个我们希望在几年内达到的规模非常大的平台业务,有很多理由说明为什么你也需要一些第一方的东西。一个原因是当你希望直接接触用户时。终端用户让你了解他们到底如何使用它?他们最想要什么?如果你是一个纯平台,没有那种直接联系,你在各种硬产品上可能会处于劣势。很难构建最好的产品。甚至可能很难知道模型真正需要往哪个方向发展。人们说像编码这样的东西,但有很多模型似乎擅长编码,但它们并不以实际相关的方式擅长。我们实际上已经成功让 Claude 以与人们实际使用相关的方式变得擅长。所以我认为这是一个原因。另一个原因回到大型企业,对于更传统的公司来说,在 API 上构建有时更具挑战性,你需要给他们一些更容易使用的东西,要么是一个帮助他们构建的工具包,要么是一个应用。所以企业也喜欢 Claude Code,我们正在逐步将 Claude for Enterprise 发展成我们所谓的虚拟同事。

I think the way we like to think about it, I think we think of ourselves as a platform company first. The analogy here would be maybe clouds or something. If you think of a really large platform business of the size we're trying to get to in hopefully a small number of years, there are a number of reasons why you would also want to have things that are first-party. One is when you want to have direct exposure to the users. The end user gives you some sense of how exactly are they using it? What are they most looking for? If you're a pure platform and you don't have that direct connection, you can be disadvantaged in various hard products. It's hard to build the best products. It may even be hard to know where the model really needs to go. People say things like coding, but there are many models that seem to be good at coding but they aren't good in the way that's actually relevant. We've actually managed to make Claude good in a way that's relevant to what people actually use. So I think that's one reason. Another reason goes back to the large enterprises where building on an API sometimes is more challenging for a more traditional company to do, and you need to give them something that's a little bit easier to use, either a kit to help them build things or you need to give them an app. So enterprises have also liked Claude Code, and we're gradually developing Claude for Enterprise into what we call a virtual coworker.

优先用例:国防、科学、发展中世界 Prioritizing use cases: defense, science, developing world

Host

但我很难想象 Anthropic 会开发 Claude for Oil and Gas Exploration。为什么?也许事实上这是下一个发布。

But I find it hard to picture Anthropic developing Claude for oil and gas exploration. Why is that? Maybe in fact it's the next launch.

Dario

我们目前没有在开发 Claude for Oil and Gas Exploration。我会区分我们不允许的事情,比如非法的事情,以及一些用例,比如,好吧,我们是一个平台,人们会做很多事情,但我们对此没有热情,我们不会在其他用例之前去实现它。所以我认为有一部分是这样的:我们可能以超出其直接盈利能力的比例投入科学和生物医学,因为我们认为这是值得的。我们对发展中国家的事情也有同样的感觉。我给你一个有争议的例子:人们以相反的方式思考。我们在国防和情报方面的工作,人们常常说,啊,这些家伙在出卖自己。我恰恰相反地看待它。有一个与国防部和情报界签订的 2 亿美元上限的合同。人们说,哦,Anthropic 在出卖自己。恰恰相反。从某个编码初创公司再拿 2 亿美元比拿到那个合同要少一个数量级的努力。国防非常值得,我们这样做是因为我们想捍卫民主,而且我们在界限内行事。有一些事情我们担心。我深切关注国内政府权力的滥用。我们更多考虑外向型的一面。但这是一个例子,说明我们优先考虑的是我们认为好的事情,不一定是感觉良好或外部舆论会积极的事情。我们实际上对某些事情有信念,并且无论如何都会去做。

We're not currently working on Claude for oil and gas exploration. I would draw a distinction between things we just don't allow, things that are illegal, and there are a number of use cases that it's like, okay, you know, we're a platform, people are going to do a bunch of things, but we're not passionate about it, we're not going to go out and make this happen before the other use cases. So I think there is a component of that where we probably work on things like science and biomed out of proportion to its immediate profitability, because we think it's worthwhile. We feel the same way about things in the developing world. One I'll give you that's controversial: people think about it the opposite way. The work we do on defense and intelligence, people are often like, ah, these guys are selling out. I think about it the opposite way. There was this contract with a ceiling of 200 million with the DoD and intelligence community. People are like, oh man, Anthropic is selling out. Exactly the opposite. Getting another 200 million from some coding startup would take an order of magnitude less effort than getting that contract. Defense is very worthwhile, and we're doing it because we want to defend democracies, and we do it within bounds. There are some things we're concerned about. I'm deeply concerned about abuse of government authority on the domestic side. We think more on the kind of outward directed side. But that's an example of the things we prioritize are things that we think are good, not necessarily things that feel good or that external buzz will be positive. We actually have conviction around some things and we do them regardless.

3-5 年业务愿景 Business aspirations in 3-5 years

Host

你提到了你想要建立的那种业务。你对 Anthropic 业务在未来三到五年的期望是什么?

You reference the kind of business you want to build. What are your aspirations for the Anthropic business in say three to five years time?

Dario

AI 在很多方面都很奇怪。我认为其中一个奇怪之处在于,因为它是指数级的,我们很难准确校准业务会有多大。所以我们有以下经历。2023 年,我从未从机构投资者那里筹集过资金。2023 年初我们的收入为零,因为我们还没有发布产品。所以我在整理一些东西时想,哦,我认为我们第一年可能能获得一亿美元的收入。这导致一些投资者说这太疯狂了,这在资本主义历史上从未发生过,你在我这里失去了所有信誉,再见。然后我们真的做到了。所以第二年我想,好吧,我认为我们可以从一亿到十亿。

AI is strange in a number of ways. I think one of the ways it's strange is that because it's an exponential, we have a hard time calibrating exactly how big the business will be. So we had the following experience. In 2023, I had never raised money from institutional investors before. Our revenue was zero at the beginning of 2023 because we had not released a product. So I was putting together something and I'm like, oh, I think we can probably get a hundred million of revenue in the first year. And this caused some investors to say this is crazy. This has never happened in the history of capitalism. You've lost all credibility with me. Goodbye. And then we actually did it. So then the next year I was like, oh well, I think we can go from 100 million to a billion.

业务指数增长 Exponential Growth in Business

Dario

实际上,第一次做到的时候,人们觉得没那么疯狂了,但通常还是被认为很疯狂。然后我们又做了一次。今年已经过半,我们的收入远超 40 亿美元。所以在对数空间里,要再增加一个数量级。未来有很多种可能。一种是当事情达到一定规模后,曲线会放缓。但也有一种挑衅性的世界,指数增长持续下去,两三年后这些就会成为世界上最大的企业。在像 Anthropic 这样的公司工作或运营的基本体验和不确定性之一就是你不太确定。你做出这个指数级预测,听起来很疯狂,可能确实疯狂。但也可能并不疯狂,因为这条趋势线之前已经验证过。我在训练 AI 模型、AI 模型认知能力的技术方面也说过类似的话。但现在我们在商业方面也看到了同样的持续增长曲线。

And actually, having done it the first time, people were like, it was a little bit less dismissed as crazy, but still often dismissed as crazy. And then we did it again. This year we're halfway through the year. We're well past 4 billion in revenue. So in logarithmic space, to add another order of magnitude. There's a bunch of different futures. There's one where once things get to a certain size, the curve slows down. But there's a provocative world where the exponential continues and in two or three years these are the biggest businesses in the world. One of the fundamental experiences and uncertainties of working at or running something like Anthropic is you kind of don't know. You make this exponential projection. It sounds crazy. It might be crazy. But it also might not be crazy because that trend line has followed before. I've said much the same thing in the context of training AI models, in the context of the cognitive capabilities of AI models on the technological side. But now we're seeing the same kind of continuous lines on the business side.

Host

那么这与缩放定律的类比是什么?你并行扩展模型质量的相关输入,就能得到更好的模型性能。是否存在这样一种情况:你投入更好的模型,有一条曲线,你花费 5 倍或 10 倍的成本来训练模型,或者有 5 倍或 10 倍的数据,无论缩放定律怎么说,收入也有某种转移曲线?我在模型上多花 10 倍的钱,模型从聪明的本科生变成聪明的博士生,然后我去一家制药公司问这值多少钱。他们通常会说值 10 倍左右。这些幂律分布出现在很多情境中。在技术方面,当你训练模型时,你捕捉到的相关性尾巴越来越长。语言结构、世界、模式中的相关性,这种相关性被认为导致了缩放定律,因为存在这种对数分布。随着模型在认知任务上越来越强大,从经验上看,如果你考虑模型在经济中的用途,公司的组织方式,组织架构图存在幂律结构。这几乎就像你在攀登价值的幂律分布。我想我对产品和市场推广的看法是,模型想要处于收入的指数曲线上,而产品和市场推广就像擦干净窗户,让光线透进来。有一种方式可以打开光圈,让指数增长发生。

So what's the analogy to scaling laws here? You scale up the relevant inputs for model quality in parallel and you get much better model performance. Is there something where you put better models in and there's some curve where you spend 5x or 10x more to train a model, or you have 5x or 10x more data, whatever the scaling laws say, and there's some transfer curve for revenue? Where I spend 10 times more on the model and the model goes from being a smart undergrad to a smart PhD student, and then I go to a pharmaceutical company and ask how much more that is worth. Often they end up saying that's worth about 10x. These power law distributions occur in a bunch of contexts. On the technical side, when you train the model, there's a longer and longer tail of correlations that you're capturing as you train the model. Correlations in the structure of language, in the world, in patterns, and that correlation is what's thought to lead to the scaling laws because there's this logarithmic distribution. As the model gets more capable in cognitive tasks, empirically so far, if you think of the uses of the model in the economy, the way companies are organized, there's a power law structure of org charts. It almost feels like you're climbing that power law distribution of value. I guess the way I think about product and go-to-market is that the model wants to be on that exponential of revenue, and product and go-to-market are a way to clean the window and let the light shine through. There's a way to open the aperture and let the exponential happen.

Dario

是的。除了这种学习冲动,模型还有这种资本主义冲动,它们想要体现出来,除非配上糟糕的产品或销售,因为它们真的很有用。那种智能对人们非常有用,所以它就像被从你身上拉出来一样。

Yes. In addition to having this learning impulse, the models have this capitalistic impulse that they want to embody unless they're given a bad product or bad sales to go with it, because they're really useful. That intelligence is really useful to people and so it kind of gets pulled out of you.

Host

是的,是的,是的。这是一种思考方式。

Yes. Yes. Yes. That is a way to think about it.

终端市场结构 Terminal Market Structure

Host

这里的最终市场结构是什么?是少数大型参与者,还是我们会不断看到针对特定领域的新兴企业?

What is the terminal market structure here? Is there a few large-scale players or do we keep seeing new upstarts for specific spaces?

Dario

很难说。很难确定。两三年前还有很多不确定性,但我认为我们可能已经相对接近最终的参与者阵容,尽管不一定是最终的市场结构或参与者的角色。大概有三到六个参与者,取决于你怎么算。这些参与者有能力构建前沿模型,并且有足够的资本来支撑自己。

Very hard. It's hard to tell for sure. There was quite a lot of uncertainty 2 or 3 years ago, but I think we might be relatively close to the final set of players, if not necessarily the final market structure or the roles of the players. There's probably somewhere between three and six players, depending on how you count. Those are the players that are capable of building frontier models and have enough capital to plausibly bootstrap themselves.

Host

我很想了解模型业务是如何运作的:你前期投入大量资金进行训练,然后得到一项快速折旧的资产,尽管可能有用处的长尾,并希望收回成本。到目前为止,我认为外界人们的印象是资本支出越来越大,以及如何被烧钱。

I would love to understand how the model business works where you invest a bunch of money up front in training and then you have this fast depreciating asset, though maybe with a long tail of usefulness, and hopefully you pay that back. Thus far, I think the image people have from the outside is ever larger amounts of capex and how you get burned.

Dario

有两种不同的方式可以描述模型业务目前的情况。假设 2023 年你训练了一个成本 1 亿美元的模型,然后在 2024 年部署,它带来了 2 亿美元的收入。同时,由于缩放定律,2024 年你还训练了一个成本 10 亿美元的模型。然后在 2025 年,你从那个 10 亿美元的模型中获得 20 亿美元的收入,而你花费 100 亿美元训练模型。所以如果你以传统方式看公司的损益表,第一年亏损 1 亿美元,第二年亏损 8 亿美元,第三年亏损 80 亿美元。看起来越来越糟。如果你把每个模型看作一个公司,2023 年训练的模型是盈利的。你支付了 1 亿美元,然后它带来了 2 亿美元的收入。模型推理有一些成本,但在这个简化的例子中,即使你把这两项加起来,情况也不错。所以如果每个模型都是一个公司,模型实际上是盈利的。实际情况是,在你从一个公司获得收益的同时,你又在创办另一个更昂贵、需要更多前期研发投资的公司。最终的结果是,这种情况会持续下去,直到数字变得非常大,模型无法再变大。

There are two different ways you could describe what's happening in the model business right now. Let's say in 2023 you train a model that costs $100 million and then you deploy it in 2024 and it makes $200 million of revenue. Meanwhile, because of the scaling laws, in 2024, you also train a model that costs a billion dollars. Then in 2025, you get $2 billion of revenue from that $1 billion, and you spend $10 billion to train the model. So if you look in a conventional way at the profit and loss of the company, you've lost $100 million the first year, you've lost $800 million the second year, and you've lost $8 billion in the third year. So it looks like it's getting worse and worse. If you consider each model to be a company, the model that was trained in 2023 was profitable. You paid $100 million and then it made $200 million of revenue. There's some cost to inference with the model, but let's assume in this cartoon example that even if you add those two up, you're in a good state. So if every model was a company, the model is actually profitable. What's going on is that at the same time as you're reaping the benefits from one company, you're founding another company that's much more expensive and requires much more upfront R&D investment. The way it's going to shake out is this will keep going up until the numbers go very large, the models can't get larger.

投资与商业模式 Investment and Business Model

Dario

那么它将成为一个规模庞大、利润丰厚的业务。或者,在某个时候,模型会停止进步,迈向 AGI 的进程会因某种原因停滞,然后可能会出现一些过度投资,导致一次性的「天哪,我们花了很多钱却一无所获」,之后业务又回到原来的规模。另一种描述方式是风险投资支持的典型模式:事情先花很多钱,然后开始赚钱。在这个领域,这种情况在同一家公司内反复发生。所以我们现在处于指数增长阶段;在某个时候我们会达到平衡。唯一相关的问题是我们在多大的规模上达到平衡,以及是否会出现过度投资。

Then it'll be a large, very profitable business. Or at some point the models will stop getting better, the march to AGI will be halted for some reason, and then perhaps there'll be some overhang, so there'll be a one-time 'oh man, we spent a lot of money and we didn't get anything for it,' and then the business returns to whatever scale it was at. Another way to describe it is the usual pattern of venture-backed investment: things cost a lot and then you start making money. That is happening over and over again in this field, within the same companies. So we're on the exponential now; at some point we'll reach equilibrium. The only relevant questions are at how large a scale we reach equilibrium and whether there is ever an overshoot.

Host

对。你提到了云公司作为比较,但云公司的数据中心资本支出更持续,他们总是在建新数据中心。而这几代模型是离散的,有点像发动机制造商不断推出新技术,比如 F-16,或者有点像药物研发,是一种研发密集型的事情。

Right. And you referenced the cloud companies as a point of comparison, but there's something about the cloud companies where their data center capex is more continuous. They're always doing new data centers, whereas there's something about how discrete these generations are that maybe it's like the way engine manufacturers keep coming up with new technologies, like the F-16, or it might be a little bit like drug development, a kind of R&D heavy thing.

Dario

是的。那你什么时候真正投入精力去训练呢?

Yes. And when do you actually go to the effort of training?

Host

是的。这几乎就像一家制药公司:你开发一种药,如果有效,就开发 10 种药;如果有效,就开发 100 种药。药物开发市场在数量上并非如此,但感觉上就是这样。

Yeah. It's almost like a drug company where you develop one drug, and if that works, you develop 10 drugs, and if that works, you develop 100 drugs. The drug development market does not work like that numerically, but it is as if it did.

Dario

对。所以我们可以把每个模型看作独立的项目,看它们的独立损益表。你是说这些模型的回报计算,至少在我们行业看到的模型中,其实并不难。当你获取客户时,如果回收期是 9 个月,你会一直这么做;这是一个很容易承保的回收期。你说回收期大概是 9 个月、12 个月,非常……

Right. So we can look at each of these models as individual programs and look at their individual P&Ls. You're saying that the payback math on those, at least in the models we've seen in the industry, is not actually that challenging. When you're acquiring a customer, if you have a 9-month payback, you'll do that all day long; that's a very easy payback to underwrite. And you're saying the paybacks are kind of 9 months, 12 months, they're very...

Dario

我不想做任何具体声明,但从定性角度看,如果你这样逐个模型地看业务,它看起来非常可行。

I don't want to make any specific claims, but qualitatively, if you look at the business this way, model by model, it looks very viable.

Host

因为不断增长的资本支出掩盖了模型业务的潜在质量。

Because the ever-growing capex is masking the underlying quality of the model businesses.

Dario

是的。

Yes.

数据墙与强化学习 Data Wall and RL

Host

2023 年,每个人都在谈论数据墙。我们就是这样解决数据墙问题的吗?

In 2023, everyone was talking about the data wall. Is this how we solved our way out of the data wall?

Dario

是的。嗯,我不知道。人们在公开场合谈论事情,有时是谣言或猜测。我甚至不一定认为存在数据墙。我要说的是,使用强化学习的想法已经存在一段时间了。追溯到谷歌 DeepMind 用 AlphaGo 击败世界围棋冠军时,那首先是强化学习。然后我们构建了这些语言模型,现在我们把两者结合起来,在语言模型之上应用强化学习。这就是思维链或推理的全部:它只是强化学习的一种花哨说法,其中强化学习环境是模型写一堆东西然后给出答案。仅此而已,只是有个花哨的名字。所以我认为这是两种关键的学习方式:基础大语言模型训练是模仿学习,强化学习是试错学习。这是两种学习风格。如果我是一个孩子,有两种学习方式:我观察父母并尝试学习他们做的事,或者我可以自己实验并学习。发展心理学很清楚人们两种方式都用。所以我们现在在语言模型中看到了这一点。我们有一个阶段进行模仿学习,另一个阶段进行试错学习。这看起来很自然。

Yeah. So, I don't know. People talk about things in public and sometimes they're rumors or suppositions. I wouldn't even necessarily assume that there's a data wall. One thing I will say is that the idea of using RL has been around for a while. If we go all the way back to when Google DeepMind beat the world Go champion with AlphaGo, it was RL first. Then we built these language models, and now we're uniting the two by putting RL on top of the language models. That's all chain of thought or reasoning is: it's just a fancy way of saying RL where the RL environment is that the model writes a bunch of stuff and then gives an answer. There's nothing more to it than that. It just has a fancy name. So I think of these as the two key ways of learning: base LLM training as learning by imitating, and RL as learning by trial and error. Those are the two styles of learning. If I'm a child, there are two ways for me to learn: I look at my parents and try to learn what they do, or I can experiment with the world and learn things. It's very clear in developmental psychology that people use both. So we're now seeing that recapitulated in the language models. We have a stage where we do the imitative learning and a stage where we learn by trial and error. It seems very natural.

人才与知识产权保护 Talent and IP Protection

Host

另一个对非 AI 行业人士来说显而易见的事情是人才争夺战,以及你的知识产权每晚都会随员工离开公司。你在最近的一次采访中提到,价值 1 亿美元的机密只是几行代码。显然,我认为你是在国家安全背景下说的,但也可以从人才角度考虑。像制药行业用专利保护秘密,华尔街也有价值 1 亿美元的秘密,只是一个非常简单的想法,对冲基金文艺复兴科技公司非常成功地锁定了员工。在当前 AI 环境下,如何保持商业领先地位?

The other thing that's obviously notable to people not in the AI industry is all the talent wars and the fact that your IP walks out the door each evening. You referenced in a recent interview you gave $100 million secrets that were a few lines of code. Obviously I think you were talking about that in a national security context, but you can also think about it in a talent context. How does one like in the pharma industry protect their secrets with patents, in Wall Street where they also have $100 million secrets that are just a very simple idea, Renaissance Technologies the hedge fund just very successfully locks up its employees. How do you make keeping a commercial lead work in the current AI environment?

Dario

是的。我要说的是,有些事情确实如此,但我认为随着领域成熟,越来越多地是关于专有技术和构建复杂对象的能力。我们处理的一些想法很简单,但简单的想法,比如「哦,调整 Transformer 的这个元素」,往往会被独立发现,或者很快就会被所有人知道。但有些事情,比如「天哪,这个东西从工程角度真的很难实现,我们还没实现」,或者「这个东西做起来很麻烦」,而且有诀窍。这些往往是更集体性的东西,更难泄露,所以我认为这些东西更具防御性。话虽如此,仍然存在泄露,我们也不希望它发生,既出于商业竞争原因,也出于国家安全原因。两者都是问题。所以我们做几件事:一是我们倾向于隔离信息。如果你和任何情报机构交谈,他们就是这样运作的:只告诉你需要知道的内容。我认为 Anthropic 内部的每个人,但这可能和典型的硅谷文化非常不同,在那里一切都在公司里飞来飞去。我们实际上在保持非常开放文化的同时也这样做。

Yeah. So one thing I will say is that there are some things that are like that, but I think more and more as the field matures, it starts to be more about knowhow and ability to build complex objects. Some of the ideas we work with are simple, but the simple ideas, the ones like 'oh yeah, twiddle this element of the transformer,' those tend to be independently discovered or anyone knows them before too long. But there are things like, 'oh man, this thing is actually really hard to implement from an engineering sense and we haven't implemented it,' or 'this thing is just kind of a pain to do,' and there's a knowhow to doing it. Those tend to be more collective things that are more difficult to leak, and so I think those things are substantially more defensible. That said, there's still leakage and we still don't want it to happen, both for commercial competitive reasons and for national security reasons. Both are problems. So a few things we do: one is we tend to compartmentalize information. If you talk to any intelligence agency, that's how they operate: you're only told what you need to know. I think everyone within Anthropic, but that's probably quite different to a normal Silicon Valley culture where everything's just flying around the company. We actually do that at the same time as we have a very open culture.

公司文化与留人 Company Culture and Retention

Dario

我对公司说的话,也许别人会用公关语言包装。但如果有秘密,我认为这反而会让人们相信这是你真正需要知道的事情。

I say things to the company that maybe another person would put in PR speak. But when there is a secret, I think that actually leads to people trusting that it's something you actually need to know.

Host

最后,拥有更高的留存率和更少的人员流失是这里最重要的事情之一。

And then finally, having better retention rates and losing fewer people is one of the most important things here.

Dario

所以我们拥有所有 AI 公司中最高的留存率。我认为差异甚至更明显,因为每家都有一定比例的非遗憾流失率,可能是个常数。如果减去这个,差异就更大了。有时候人们离开后还会回来。我见过这种情况。如果你看看公开的去 Meta 超级智能实验室的人员名单,即使按我们的规模调整,也不多……而且很多人拒绝了他们。

So we have the highest retention rate of all the AI companies. I think the differences are even starker because everyone has a non-regretted attrition rate that's maybe constant. So if you subtract that off, the difference is even larger. Sometimes when people leave, they come back. I saw that. If you look at the publicly available list of people who went to the Meta super intelligence lab, even if you normalize for our size, it's not... and then many turn them down.

Host

所以在大家都在谈论的疯狂亿万美元薪酬战中,你们并没有太难熬。

So in the crazy hundred million dollar comp wars that everyone's been talking about, you guys have not had too hard a time of that.

Dario

我认为相对于其他公司,我们做得不错。我们甚至可能相对有优势。这是对使命的真正信念和对股权上升空间的信心的结合。我认为 Anthropic 已经建立了说到做到的声誉,有时承诺更少但兑现承诺,并且多年来立场清晰、保持一致。这围绕公司创造了一种团结,我认为这是对犬儒主义的一种良好防御。

I think relative to other companies, we've done well. We may even have been relatively advantaged. It's a mixture of true belief in the mission and belief in the upside of the equity. I think Anthropic has developed a reputation for doing what it says it will do, in some cases making fewer promises but keeping those promises, and being very clear on what we stand for and being consistent over the years. That creates a unity around the company and I think it's a good guard against cynicism.

业务推介 Pitching the Business

Host

当你谈到股权的上升空间,当你向投资者或候选人推销时,你如何推销 Anthropic 的业务?比如我们正在建立一个非常大的业务。这是一个好的开始。还有什么?

When you're talking about the upside of the equity, when you're pitching investors or maybe candidates, how do you pitch the Anthropic business? Like we're building a very large business. That's a good start. What else goes into it?

Dario

我经常谈论平台和模型的重要性。出于某种原因,有时人们想到 API 业务,会说它粘性不高或者会……

Often I'll talk about the platform and the importance of the models. For some reason, sometimes people think of the API business and say it's not very sticky or it's going to be...

Host

API 业务。我喜欢 API 业务。

API business. I love API business.

Dario

不,完全正确。还有比我们两家都更大的。我会再次提到云服务。那些是千亿美元的 API 业务。当资本成本高且玩家很少时,相对于云,我们制造的东西差异化更大。这些模型有不同的个性,就像和不同的人交谈。我经常开的一个玩笑是:如果我坐在一个房间里有 10 个人,这是否意味着我被商品化了?还有九个人有类似的大脑,差不多身高。那么谁需要我?但我们都知道人类劳动不是这样运作的。我对这个也有同样的感觉。所以我认为 API 业务是一个伟大的业务。但我们想走得更远。我的想法是,其他玩家如 OpenAI 和现有巨头如谷歌非常关注消费者端。为企业提供 AI 是我们正在努力做得更好的事情。我认为我们在这方面取得了早期领先。我不确定,因为我不确切知道其他玩家的收入,但我认为目前我们可能拥有 API 市场和企业 AI 市场的多数份额。

No, exactly. There are even bigger ones than both of ours. I would point to the clouds again. Those are hundred billion dollar API businesses. When the cost of capital is high and there are only a few players, and relative to cloud, the thing we make is much more differentiated. These models have different personalities, they're like talking to different people. A joke I often make is: if I'm sitting in a room with 10 people, does that mean I've been commoditized? There are nine other people with a similar brain, about the same height. So who needs me? But we all know human labor doesn't work that way. I feel the same way about this. So I think the API business is a great business. But we want to go broader than that. The way I think about it is other players such as OpenAI and existing incumbents such as Google are very focused on the consumer side. The idea of providing AI to businesses is something we are trying to get better and better at. I think we're out to an early lead in that. I'm not sure because I don't know for sure what the revenues of the other players are, but I think we probably at this point have the plurality of the API market and AI for business market perhaps.

商品化争论与差异化 Commodity Argument and Differentiation

Host

当你谈到商品化论点时,很有趣。我们成长过程中一直面对这个怀疑论。我记得当 AWS 终于在 2015 年不得不拆分出他们的数字时,我觉得非常惊人。它们以前被包裹在亚马逊的数字里,然后他们拆分了出来。人们一直说云是商品,无趣。然后他们拆分出来,结果是有史以来最伟大的业务之一。一个业务可以有竞争对手和关心价格的买家,但这与商品化非常不同。正如你所说,产品运作方式不同。

It's funny when you talk about the commodity argument. We grew up facing this as a skeptical argument. I remember finding it so striking when AWS finally had to break out their numbers in 2015. They used to be wrapped up in Amazon's numbers and they broke them out. People had been saying cloud is a commodity, uninteresting. Then they broke it out and it's one of the greatest businesses of all time. There's something where a business can have competitors and buyers who care about price, but that's very different from being a commodity. As you say, products work differently.

Dario

完全正确。我的意思是,我们是云服务最大的客户之一,而且我们使用不止一家。我可以告诉你,云服务的差异化远小于 AI 模型。

Exactly. I mean, we're one of the biggest customers of the clouds, and we use more than one. I can tell you the clouds are much less differentiated than the AI models.

Host

确实。因为感觉上,行为是非确定性的,这不是故意设计得困难,但这自然意味着我们用这个模型得到更喜欢的客服答案,而另一个模型则不同,而且我不知道为什么。

For sure. Because it feels like one, the behavior is non-deterministic, which not by design trying to make it hard, but that just naturally means that we get the customer service answers we prefer with this model versus that model, and I don't know why.

Dario

完全正确。你不知道为什么。有点像烤蛋糕。你放入原料,它就以某种方式出来。一个厨师这样做,另一个那样做。如果你说完全像那个厨师那样做,你做不到。就是做不到。

Exactly. You don't know why. It's a little like baking a cake. You put in the ingredients, it just comes out a certain way. One chef makes it this way and the other chef makes it that way. If you say make it exactly like that chef makes it, you can't. You just can't.

Host

而且,让我害怕的是,目前没有 AI 产品那么个性化,但感觉个性化将是一件大事。

And presumably, it's frightening to me none of the AI products are that personalized right now, but it feels like personalization will be a huge deal.

Dario

将是一件大事。

Will be a huge deal.

Host

并且将成为粘性的重要来源,因为你不会想切换产品。我不确切知道那会是什么样子,但考虑到数量……对于消费者和企业用例都是如此。

And will be a big source of stickiness because you won't want to switch products. I don't know exactly what that looks like, but given the amount of... for both the consumer and the business use case.

Dario

绝对。我认为我们才刚刚开始触及表面,模型以各种方式定制,用于与特定企业或企业内的特定人员合作。所以我认为我们只是看到了 API 业务的开始。但我不认为企业 AI 只是 API。通过像 Claude Code 这样的东西,我们不仅向个人开发者销售,也向企业销售。

Absolutely. I think we've just started to scratch the surface in terms of models that are customized in various ways for working with a particular business or particular person within the business. So I think we're just seeing the beginning of the API business. But I don't think AI for business is just about API. With things like Claude Code, we're selling that to not just individual developers, but enterprises as well.

Host

他们发现它很有用。Claude for enterprise 正在向许多企业销售。我确实看到了这一点,你可以在一些云服务中看到,它们有一堆不同的服务。有些是应用,有些是底层云本身。它们的方式是,AWS、GCP 或 Azure 展示自己的方式,以及我们开始展示自己的方式是:嘿,我们想成为你的一站式 AI 或云服务商。你可以购买所有这些,你可以和我们讨论哪个用于什么。

And they find it something useful. Claude for enterprise is selling to a lot of enterprises. I actually see it, and you see this with some of the clouds where they have a bunch of different services. Some of them are apps, some is the underlying cloud itself. What they are is, the way that AWS or GCP or Azure will present themselves, and the way that we are starting to present ourselves is: hey, we want to be your one-stop shop for AI or for cloud. You can buy all of these things and you can talk to us about which to use for what.

大公司 AI 采用 AI Adoption in Large Companies

Host

如果你想想典型的财富 500 强公司,它们采用 AI 的程度与应有的程度相比如何?

If you think about a typical Fortune 500 company, how AI adopted are they compared to how much they should be?

Dario

当然,远低于应有的水平。

Well, certainly much less than they should be.

Host

大概是 5%还是 30%?

What is it like 5%? 30%?

Dario

我的看法是,高层通常很有信心。你和 CEO 谈,CEO 懂;你和 CTO 谈,CTO 也懂。他们的困难在于,公司有 10 万名员工,他们的工作是做别的事情,比如银行、保险或药物开发。他们听说过 AI,但并非专家。挑战往往在于与领导层合作,让这 10 万人真正熟悉并使用这项技术。代码相关的东西进展最快,因为开发者最贴近并关注趋势。客户服务和流程类的东西其次。但你真的有种直觉,即使使用今天的模型,其应用规模也可能比现在大 100 倍。

So what I would say is there is very often conviction at the top. You talk to the CEO, the CEO gets it. You talk to the CTO, the CTO gets it. The struggle they have is that they have a 100,000 people whose job is to do something else, like banking or insurance or drug development. They've heard about this AI stuff, but they're not experts in it. The challenge is often working with the leadership to get the 100,000 people really familiar with and using the technology. The code stuff goes fastest because developers are most adjacent and watching the trend. Customer service and process stuff is next. But you really have the instinct that even with today's models it could be 100 times bigger than it is.

Host

我的直觉是,我们会从初创公司看到 AI 采用的模式,因为它们不受现有组织的约束,可以做任何有意义的事情;而大型组织则僵化了,因为它们有那么多职责固定且需要咨询的人。所以我们会从小型初创公司看到新行为,然后大公司——如你所说,CEO 和 CTO 们很敏锐、很聪明——他们会移植这些新想法,就像我们看到的云或其他技术趋势的采用一样。你看到的是这样吗?

My intuition is that we will see the patterns of AI adoption from startups because they're unconstrained by existing organizations, so they can do whatever makes sense, versus large organizations which are calcified because they have all these people whose job is to do X and need to be consulted. So we'll see new behaviors from small startups, and then large companies, as you say, the CEOs and CTOs are switched on and smart, they'll port the new ideas, like the adoption we saw with cloud or other tech trends. Is that what you're seeing?

Dario

小公司的新想法,或者小公司本身会对它们构成威胁并颠覆它们,这会给它们带来紧迫感,推动事情发生。我见过一个效果不错的模式,实际上我建议大公司这样做:成立一个独立于公司其他部门的突击队或特别小组,开发这些原型,然后就能获得势头。之后总是有整合到公司其他部门的艰苦工作,但如果你有足够的势头,并且证明了事情是有效的,那就更容易做到。

The new ideas from the small companies or the small companies will become threatening to them and disrupt them, and that will give them the urgency to drive things through and make them happen. A pattern I've seen that works pretty well, that I actually recommend if you're a large company, is to make a strike team or strike force that's separate from the rest of the company and develop these prototypes, and then you can get momentum behind something. Then there's always the hard work of integrating into the rest of the company, but if you have a lot of momentum and you've shown the thing works, it's easier to do that.

持续学习与 AI 能力 Continual Learning and AI Capabilities

Host

你读过他最近关于 AI 时间线的博客文章吗?

Did you read his recent blog post on his AI timelines?

Dario

哦,关于持续学习的。是的。

Oh, on continual learning. Yeah.

Host

他谈到,他对许多用于生产力的 AI 模型的基本看法是,它们就像一个 5 分钟前入职的超聪明虚拟同事,但始终停留在 5 分钟前的状态。它们不会随时间学习。我们将如何解决这个问题?

He talked about how his fundamental issue with many AI models for productivity is that they're like the super smart virtual coworker who started 5 minutes ago, but they remain the coworker that started 5 minutes ago. They don't learn over time. How will we solve that?

Dario

我在 AI 研究和技术方面看到的模式是,我们一次又一次地看到那里似乎有一堵墙。看起来 AI 模型做不到这个,对吧?比如 AI 模型不能推理。最近又有 AI 模型不能做出新发现。几年前,AI 模型不能写出全局连贯的文本,现在显然可以了。再往前几年,它们能正确使用语法但无法理解语义,而所有这些障碍都被突破了。

The pattern I've seen in AI on the research and technical side is that what we've seen over and over again is that there's what looks like a wall there. It looks like AI models can't do this, right? It was like AI models can't reason. And recently there's this AI models can't make new discoveries. A few years ago it was like AI models can't write globally coherent text, which of course now they obviously can. You go back a few more years and it was like they can get syntactics right but they can't get semantics right, and every one of those has been blown through.

Host

那么在新发现方面我们突破了什么?这是人们最近常说的。

So what have we blown through on the new discoveries? This is a thing that people have said recently.

Dario

实际上,我的观点是,就像许多其他事情一样,关于新发现,它并不是二元的。它们不会在论文上署名。但什么是新发现?什么是天才?我记得一本发展心理学书里说,我们神化天才,但假设一张桌子摇晃,我拿一个杯垫垫在桌腿下,桌子就不晃了。那也是一个想法,在某种程度上是一个新发现,即使我以前没见过别人这么做。它和诺贝尔奖级别的发现之间的区别是程度问题,而非本质不同。所以我会说 AI 模型一直在做发现。我有家人遇到医疗问题,AI 模型诊断出来了,而医生却漏诊了。那就是一个新发现。你可以说它们只是在匹配过去发生过的事情,但新发现就是这样。写全新小说的作家也是在混搭并加入新元素。这一切都是连续的。这就是我关于持续学习想说的。我认为那种认为它不存在的观点……我会说它存在一点。

Actually, my view on this, like many of the other things, on new discoveries is that it's not really a binary. They don't get to have their name in the paper. But what is a new discovery? What is genius? I remember a developmental psychology book saying something like, we lionize genius, but let's say a table's wobbly and I take a coaster and put it under the table and it's not wobbly anymore. That's an idea, a new discovery in a way, even if I've never seen someone do that before. The difference between that and a Nobel Prize winning discovery is a matter of degree, not fundamentally different. So I would say AI models make discoveries all the time. I've had family members where they had a medical problem and the AI model diagnosed it when doctors missed it. That's a new discovery. You could say they're just pattern matching to things that happened before, but new discoveries are like that. Writers who write novels that are totally new are remixing together and adding a new element. It's all more continuous. And that was the thing I was going to say about continual learning. I think this idea that it isn't present is... I would say it's present a little bit.

Host

是的,这让人安心。我们会找到方法获得更多。

Yes, it is comfortable. We're going to find a way to get more of it.

Dario

例如,模型在上下文中学习。你和它们交谈,它们吸收上下文。最终上下文将达到 1 亿个 token,也许我们会以某种方式训练模型,使其专门在上下文中学习。你甚至可以在上下文中更新模型的权重。有很多非常接近我们现有想法的思路可能做到这一点。我认为人们非常执着于相信存在某种根本性的障碍,某种无法做到的不同之处。这让我想起 19 世纪的活力论,即人体和生物体是由与无生命物质根本不同的材料构成的。

So for instance, the models learn within the context. You talk to them and they absorb the context. Eventually the context is going to be 100 million tokens, and maybe we'll train the model in such a way that it is specialized for learning over the context. You could even during the context update the model's weights. There are lots of ideas that are very close to the ideas we have now that could perhaps do this. I think people are very attached to the idea that they want to believe there's some fundamental wall, something different that can't be done. It kind of reminds me of the 19th century notion of vitalism, the idea that the human body and living organisms are made of fundamentally different material than inanimate matter.

活力论与心智本质 Vitalism and the nature of mind

Host

当然,我们现在科学上知道这不是真的,但这是人们非常愿意相信的东西,而且你的常识似乎也暗示了这一点。比如我不太像一张桌子,我由与金属或玻璃等完全不同的材料构成。但当我们深入到基本单位时,当然我们都是由同样的东西构成的。

Which of course we know scientifically now is not true, but it's something people very much want to believe and your common sense seems to suggest it. Like I'm not very much like a table, I'm made of very different materials than metal or glass or whatever. But when we actually go down to the fundamental units, of course we're all made of the same thing.

Host

但你认为人们现在有一种现代版的活力论,认为某种根本的人性存在,并说模型做不到。我认为有一种倾向相信这一点。而且我认为,就像活力论一样,解决之道是认识到心智就是心智,无论它由什么构成。认知或感知的尊严或特殊性,并不是说它不特殊,而是它可以由任何东西构成。

But you think people now have this kind of modern concept of vitalism in whatever the fundamental humanity is, and they're saying models can't do it. I think there's some tendency to believe it. And I think as with vitalism, the way around it is to recognize that a mind is a mind no matter what it's made of. The notion of the dignity or the specialness of cognition or sentience is not that it isn't special. It's that it can be made out of anything.

Host

你提到了医疗用例,我认为这是一个非常酷的用例。显然是因为所有因此解决医疗问题的人。但另一个是你在《爱与优雅的机器》一文中谈到的,我真的很喜欢,关于智能的边际收益。智能的限制因素是什么?我对流行医疗用例的理解是,它显然是一个有魅力的用例,但对大多数普通人来说,他们都有某种医疗问题,无论轻微还是严重,实际上社会非常受智能限制。并不是说你无法接触到聪明的医生,希望你能,但他们给你的时间非常有限。他们花 10 秒钟思考你的问题,结果发现测试时算力正是我们在医疗方面需要的。但这是你的看法吗?

You referenced the medical use case, which I think is a very cool use case. Obviously because of all the people who have fixed medical issues as a result. But another one is you talked in your 'Machines of Love and Grace' post, which I really enjoyed, about the marginal returns to intelligence. What are intelligence's limiting factors? My read of the popular medical use case is obviously it's a kind of charismatic use case, but also for most normal people, they have some kind of medical issue, low level or serious, and actually society is very intelligence limited. Not that you don't have access to a smart doctor, hopefully you do, but they give you very limited time. They think for 10 seconds about your problem, and it turns out test-time compute was actually what we needed there on the medical stuff. But is that your take on this?

Dario

我也是这么想的。我和诺贝尔奖得主生物学家聊过,他们说,我只——这听起来有点精英主义,但他们说我只找前 1%的医生,因为剩下的 99%,我从 LLM 那里能得到更好的建议。这确实是真的。医生很忙,工作过度。而且医疗数据和医疗信息的本质就是大量的模式匹配,很多都是相同的东西。一致性和整合许多不同事实的能力,我认为这是 LLM 相当擅长的。

That is how I think about it as well. I have talked to Nobel Prize-winning biologists who say, I will only — it sounds a little elitist, but they'll say I'll only go to the top 1% of doctors because the rest of the 99% I can get better advice from an LLM. It really is true. Doctors are busy. They're overworked. And just the nature of medical data and medical information, it's a lot of pattern matching. It's a lot of the same things. The level of consistency and the ability to put together many different facts, I think it's something that LLMs are quite good at.

医学之外的智能受限领域 Intelligence-limited areas beyond medicine

Host

所以你在《爱与优雅的机器》一文中谈到了人类层面的一些大领域,我们在这些领域受智能限制,但个人医疗用例又是一个很好的例子,说明社会受智能限制,如果你给很多人更多关于他们具体问题的智能,这非常有价值。在消费者用例或商业用例中,你认为还有哪些领域我们明显受智能限制?

So you talked in the 'Machines of Love and Grace' post about some of the big humanity-level areas where we're intelligence limited, but again the personal medical use case is a good example of one where society is intelligence limited, and if you give lots of people much more intelligence on their specific issues, it's very valuable. What are other areas either in the consumer use case or in the business use case where you think we're just very obviously intelligence limited?

Dario

是的,至少今天的 AI 模型最能帮助的地方。其特点是事情是重复的,但每个例子都略有不同,对吧?在 AI 之前的自动化,如果你能精确编程事情如何发生,你就能做到。所以如果你一遍又一遍地做同样的事情,但客户服务就像——就拿客户服务为例——有长尾的东西,但很多就像你接到一堆电话。每个电话都不同,但每个电话基本上都是关于 10 件事之一。而且是一个不同的人用不同的声音,基本上以不同的方式说这 10 件事之一。这种事情重复且相似但不相同,每个都有其特点的情况,我认为正是 AI 最能介入的地方。

Yeah, the places where at least the AI models of today can help the most. The characteristic quality is something is repetitive, but every example is a little different, right? Automation before AI, if you could program exactly how it happened, you could do it. So if you were doing the same thing over and over again, but customer service is like, just to take customer service as an example, there's a long tail of stuff, but a lot of it is like you get a bunch of calls. Each call is different, but each call is basically about one of 10 things. And it's like a different person in a different voice saying basically one of these 10 things in a different way. And that situation where things are repetitive and similar but not the same and each has its own thing, that's where AI can come in the most, I think.

AI 做税务预测 Prediction on AI doing taxes

Host

是的。Doresh 在同一篇博文中预测,你还不能把所有的财务数据给现有的 AI,转发所有邮件,让它帮你报税。他预测了哪一年可以做到,也就是你第一次通过把所有东西通过电子邮件发送给你使用的任何 AI 来完成报税的年份。他的预测是 2028 年。你怎么看这个预测?

Yes. Doresh had in the same blog post the prediction that you can't yet give an existing AI all of your financial data and forward all the emails and have it do your taxes. And his prediction for the year in which you can, what is the year where your first tax return is done by just emailing everything to whatever AI you use. His prediction was 2028. What do you make of that prediction?

Dario

可能更早。我不知道是 2026 年还是 2027 年。部分原因是模型的准确性。我认为模型今天就能做到,但会犯太多错误。所以研究让模型检查自己的工作并减少错误的方法是一部分。还有接口部分。但如果要那么久,我会感到惊讶。

Probably sooner than that. I don't know if it's 2026 or 2027. Some of that is model mostly accuracy. I think the model could do that today, but it would make too many mistakes. And so working on ways to have the model check its own work and do fewer mistakes is one part. There's kind of an interface part of it as well. But I would be surprised if it takes that long.

幻觉与人类对比 Hallucinations and human comparison

Host

好的。2026 年或 2027 年。当你谈到错误时,实际上你正在列举人们认为 AI 永远无法解决的事情。感觉幻觉应该在这个列表上。虽然没有完全解决,但已经好多了。

Okay. 2026 or 2027. And when you say about mistakes, actually you were running through the list of things that people thought we would never solve in AI. It feels like hallucinations should be on that list. Not totally solved, but they've gotten a lot better.

Dario

它们已经好多了,而且我认为人们已经更习惯了。他们大致知道该信任模型什么,不该信任什么。模型也通过引用得到了支撑。我的意思是,我们在 claude.ai 上做到了这一点。我们在 Enterprise Claude 上也做到了。所以部分解决方案是引用。部分解决方案是算法上模型现在幻觉更少了。还有部分解决方案是人们已经适应并理解了模型的弱点。我对幻觉这类事情的看法一直是,有一类批评者指出模型奇怪或比人类差的地方,然后说「看,它们根本不像我们」或「它们永远达不到」。我有点理解这种直觉从何而来,也许他们是在看我们是否完全匹配了人脑。他们说「哦,这太不同了,不可能像人脑。」但我基本上认为这是一个谬误。有一种通用智能的概念,但它由许多不同的东西组成。你可以拥有大部分东西,但在某些方面差很多,在其他方面好很多。比如,如果我们看看人类,你知道,你见过人类吗?对吧。比如,如果你看看自闭症患者与精神分裂症患者,如果你看看人类面对而机器不会被愚弄的视错觉,很明显我们也有这些弱点。就像模型的幻觉一样。只是它们看起来非常不同,而且我们更习惯它们,因为我们整天被人类包围。

They've gotten a lot better, and I think people have gotten more used to it. They kind of know what to trust the model for and what not to trust the model for. The models have also been grounded in citations. I mean, we've done that with claude.ai. We've done that with Enterprise Claude. So part of the solution is citation. Part of the solution is algorithmically the models hallucinate less now. And part of the solution is people have adapted and understand the weaknesses of the model. My view on things like hallucinations has always been there's a certain class of critic who points to something where models are weird or worse than what humans do and say 'see, they're not like us at all' or 'they'll never get there.' And I kind of get where the instinct comes from, where maybe they're looking to see if we've matched the human brain exactly. They're saying 'oh this is so different, it can't be like a human brain.' But I basically just think it's a fallacy. There's a notion of kind of general intelligence, but it's made up of a bunch of different things. And you can simply have most of the things and be much worse on some and much better on others. Like if we look at humans that are, you know, have you met humans? Right. Like if you look at humans who are autistic versus humans that are schizophrenic, if you look at the optical illusions that humans face that machines are not fooled by, it's very clear that we have some of these weaknesses. Just like the models' hallucinations. It's just that they look very different and we're much more used to them because we're surrounded by humans all day.

自动驾驶双重标准 Double standard for autonomous vehicles

Host

自动驾驶汽车的双重标准感觉是最明显的例子。

The autonomous vehicle double standard feels like the clearest example of this.

Dario

最明显的例子。是的。

Clearest example of this. Yes.

Host

人们有更高的标准。

Where people have much higher standards.

Dario

人们有更高的标准。但我认为这将是这项技术的一个特点,并且对商业方面有影响。我认为我们将进入一个模型犯错频率远低于人类,但错误会更奇怪的世界。实际上,这需要一些适应,因为想象你是一个最终用户。如果你和人类一起工作,你会习惯并有一些概念,对吧?所以,如果人类有 5%的时间犯错,你可能很好地理解原因,比如如果我和一个客服代表通话,他们听起来语无伦次、口齿不清,他们可能喝多了,工作做得不好。那是一个糟糕的错误,但如果我和这个人说话,我大概知道发生了什么,并且知道不要相信他们说的话。而大语言模型犯错的频率可能低五倍,但更具欺骗性。模型听起来和说正确内容时一样清晰、连贯。这是一个适应问题,不是根本问题。当我们和客户交谈时,我们会告诉他们这一点。我们告诉他们需要习惯这一点。

People have much higher standards. But I think it's going to be a feature of this technology and it has implications on the business side. I think we're going to be in a world where the models will make mistakes much less often than humans, but they'll be stranger mistakes. And actually, that takes some adaptation because imagine you're an end user. If you work with humans, you get used to it and you have some notion, right? So, if a human makes a mistake 5% of the time, you might have a good understanding of why, like if I'm talking to a customer service agent and they sound incoherent and they're slurring their speech, they probably had too much to drink and they're not doing their job very well. That's a bad mistake, but also if I'm talking to this person, I kind of know what's going on and I know not to trust what they're saying. Whereas an LLM might make a mistake five times less often, but it's more deceptive. The model sounds just as articulate, just as coherent as when it's saying something that's right. That's an adaptation thing. That's not a fundamental thing. And that's something that when we talk to our customers, we tell them about that. We tell them they need to get used to that.

Host

所以你是说我们需要为大语言模型发明口齿不清。

So we need to invent slurring for LLMs is what you're saying.

Dario

对,对,对。正是。

Right. Right. Right. Exactly.

从研究员到 CEO 转型 Transition from researcher to CEO

Host

所以你最初是一名研究员,但现在你是一家公司的 CEO,从事销售 AI 的业务。关于市场推广或与客户打交道,你学到了什么?

So you started out as a researcher, but now you're the CEO of a company and you're in the business of selling AI. What have you had to learn about go-to-market or dealing with customers?

Dario

是的,当然。我认为我的观点是,我创办公司并不是因为我最初对销售或商业之类的东西感到兴奋。我见过其他一些公司的运作方式以及他们试图构建的东西的规模和重要性,我只是有点担心那些人和动机可能不是最好的。我知道这个领域会有很多参与者,但感觉至少有一个在做事方式上有强烈指南针的参与者可能对生态系统产生积极影响。我们会以不同的方式构建东西,以不同的方式部署它们。最重要的是,我们会有一个简短的准则清单,并尽可能坚持它们。所以我认为这是最初的动机。当然,我对构建技术感到兴奋。随着事情的发展,我和其他联合创始人不得不学习如何思考商业和战略。我认为我自然而然地就对商业方面感兴趣。实际上,我惊讶于自己这么快就产生了兴趣。主要原因是,我对作为我们客户的所有行业感到好奇。有点像云,也许像你的业务,我们服务的业务遍及所有可能的行业。所以你了解到经济中你从未想过的部分。即使在我名义上很了解的领域,比如我以前是生物学家,所以我对制药业务了解很多,但我从未想过科学之外的东西。我从未想过投资组合方面,从未想过临床试验如何运作以及如何降低成本。我从未详细考虑过国防和情报业务。所以当你经历这些,我发现理解人们的问题以及 AI 如何帮助解决这些问题非常有趣。实际上,产品方面是我最初更不情愿的。我觉得我对商业方面有自然的兴趣和好奇心,但构建应用程序最初从未吸引我,即使在我创办公司之后。但最近,随着我看到哪些产品成功、哪些失败,我认为如何设计产品使其成为我们所说的 AGI-proof(抗 AGI)的想法,即产品的方向是持久的,并且是通向未来有用事物的桥梁。我们都听说过包装公司或包装产品的概念。想法是,你制作了 Claude N,有人制作了一个产品来解决 Claude N 的缺陷,但然后你推出了 Claude N+1,它就把那个产品吃掉了。我经常给出的建议,我认为所有 AI 公司的人都会给出,就是不要那样做。看清领域的方向,尝试制作互补的东西。思考如何以新的方式、以 AGI-proof 的方式制作产品,这极大地引起了我的兴趣。

Yes, absolutely. I think my view on this was, I started a company not because I was initially excited about selling things or business or any of that. I'd seen the way that some of the other companies had run and the magnitude and gravity of what they were trying to build, and I was just a bit concerned that the people and the motivations were maybe not the best ones. I knew that there would be a number of players in this space, but it felt like having at least one player that had a strong compass in how we do things could have positive effects on the ecosystem. We would build things in a different way. We would deploy them in a different way. And above all, we would have a short list of principles, but we would stick to them as well as we could. So I think that was the initial motive. And of course I was excited about building the technology. As that has happened, I and the other co-founders have had to learn how to think about the business and the strategy. I think I've been very naturally interested in the business side of it. Actually, I was surprised at how quickly I became interested in it. The primary reason was that I was curious about all the industries that are customers of us. Somewhat like the clouds and perhaps like your business, the businesses that we serve run across every possible industry. So you learn these things about parts of the economy that you've never thought about. Even in areas where nominally I know a lot, like I used to be a biologist, so I know a lot about the pharmaceutical business, but I never thought about it beyond the science. I never thought about the portfolio side of it. I never thought about how clinical trials work and how they could be made cheaper. I never thought about the defense and intelligence business in any great detail. So you run through those and I just find it super interesting to understand what people's problems are and how AI can help with those problems. I very naturally actually the product side was one where I was initially more reluctant. I felt like I just had a natural interest and curiosity in the business side of it, but building apps it was somehow initially it was never a thing that drew me in even after I started the company. But I think more recently as I've seen what products have succeeded and what products haven't, I think this idea of how to design products so that they are what we call AGI-proof, so that the direction of the product is durable and is kind of a bridge to things that are useful in the future. We've all heard this idea of wrapper companies or wrapper products. The idea is, you make Claude N and someone makes a product that basically addresses the deficiencies of Claude N, but then you come out with Claude N+1 and it just eats it. The advice I always give that I think all the folks at the AI companies give is, don't make that. See the direction of the field and try to make something that's complementary. And thinking about how to make products in a new way, in a kind of AGI-proof way, that has caught my interest a great deal.

AI 界面缺失与拟物化 Lack of AI UIs and skeuomorphism

Host

好的。很高兴你提到这个。你不觉得我们现在没有 AI 用户界面吗?就像我们仍然在文本框中输入文字,和 20 世纪 70 年代的终端一模一样。我是说,圆角多一点之类的。就像我们仍然对着手动触发的语音伴侣模式说话,这和 Transformer 之前的 Siri 一样。所以用户界面完全……

Okay. So, glad you brought this up. Doesn't it feel like we have no AI UIs right now? Like we still enter text into text boxes, literally same as terminals from the 1970s. I mean, a bit more rounded corners and everything. Like we still talk into voice companion modes that are manually triggered, which is the same as pre-transformer Siri. So UIs are just completely...

Dario

是的,有些不对劲。我基本同意你的看法。这让我想起互联网早期,人们会制作那些结构看起来像物理世界的网站。打开壁橱,像角色一样。有个术语,我忘了。拟物化。是的。感觉这里也有类似的情况。我想说的是,随着我们更多地转向智能体,我们将进入一个 AI 模型可以端到端完成某些事情的世界。我们几乎做到了,Claude 可以端到端完成某些事情,并且大部分时间都能做对。

Yeah, there's something not quite right about it. I basically agree with you. It reminds me a little bit of the early days of the internet, where people would make these websites that had structures that looked like they were in the physical world. Open the closet and like the characters. There's some term for this. I forget. Skeuomorphism. Yes. It feels like there's some of that going on here. A thing I would say is that as we move more towards agents, we're going to be in a world where the AI model can do something end to end. We're almost there with Claude can do something end to end and get it right most of the time.

Host

是的。

Yes.

AI 代理的界面问题 Interface problem with AI agents

Dario

人类的主要工作是偶尔检查。但有趣的是,检查往往意味着深入了解发生了什么。所以这里存在某种阻抗不匹配,需要某个产品或界面来解决。你希望东西尽可能流畅,自动运行,大多数时候无需关注,但出问题时却需要深入参与。

And a human's main job is to check sometimes. But interestingly, checking often means getting really into the details of what happened. So there's some kind of impedance mismatch here. That some product or interface is the solution to. You want something that's as slick as possible and just goes off and does something, and you don't want to have to pay attention most of the time, but when something's wrong, you might actually need to get quite involved.

Host

我觉得现在没有任何产品或界面按照这个原则运作或处理这个问题。

And I don't feel like any products or interfaces operate on this principle now or handle this problem now.

Dario

是的。是的。

Yes. Yes.

Host

我不知道这有没有道理,但……

I don't know if that makes sense, but...

Dario

不,有道理。我同意。我认为你希望你的智能体去为你做好工作,然后带着成果回来让你审查、引导、决策。但你不能被淹没,因为它要做的事情远多于你有时间查看的。如果你总是盯着它,可能比你自己做还慢。所以这在我看来是一个界面问题。

No, it does. I agree. I think what you want is your agent to go away and do really good work for you, then come back with its work product to let you review, steer it, decide. But you can't be overwhelmed because it's going to do so many more things than you have time to look at. If you're always looking at it, it can be slower than if you just did it yourself. So it strikes me as an interface problem.

Host

是的。是的。概括来说,我觉得 AI 最令人兴奋的一点是,当前能力有巨大过剩,可以转化为好产品,即使 AI 进步现在停止,我们也有大约 10 年的好产品可做。

Yes. Yes. The generalization of this is that it feels to me one of the most exciting things about AI is we have such an overhang of current capabilities turning them into good products, where even if AI progress was frozen right now, we'd have like 10 years of good product.

Dario

哦,我完全同意。实际上,整个行业构建产品的方式——我们也是这么想的——非常不同,因为进步在持续。如果模型进步停止,我们构建产品的方式会立刻改变。原因在于,我认为我们从未遇到过在构建产品时技术如此快速变化的情况。所以长期产品路线图或常规产品规划的想法,我又开始明确强调。在 Anthropic 早期,我就像个对产品一无所知的傻瓜。但现在我总会跟新来的人说,这不像在非 AI 领域构建产品,对吧?

Oh, I completely agree. And actually, the way products are being built by everyone in the industry, but we've thought about it this way, is very different because the progress is continuing. If the progress in models stopped, the way we built products would change instantly. The reason is I don't think we've ever had a situation in which the technology is changing under you so fast as you're building the product. So this idea of long-term product roadmaps or the usual way of product planning, I've started explicitly again. Early in Anthropic, I was like, I don't know anything about product. I'm a doofus. But now I always try to talk to people when they come in and say this is not like building products in the non-AI space, right?

Host

它们需要更……

They need to be more...

Dario

你可能擅长构建这些,但技术在你脚下移动。所以快速迭代的理念比通常更加重要。

You may be the expert at building these, but the technology is moving under you. So these ideas about fast iteration are even more true than they are normally.

Host

有什么具体例子吗?

What's a specific example of this?

Dario

我认为如果你想做某个东西,计划 6 个月后完成,这比孤立构建更不合理。所以你需要更紧凑的发布周期。你需要尝试。很难预测什么会流行,因为新模型可能突然擅长某件事,从而使产品成为可能。所以最重要的是,你在尝试从未尝试过的东西。有一个新模型,只在公司内部可用。所以你应该做的是在上面构建东西,让内部人员试用。这有种「永恒九月」的感觉。就像你第一次发现数据库技术,然后想,你能在上面构建什么?永远是第一天。这就是不同之处。

I think that if you're trying to make something and it's going to be ready in 6 months, that makes even less sense here than building in isolation. So you need tighter ship schedules. You need to try things. It's very hard to tell, even harder to tell what's going to catch on, because a new model may have come out and suddenly be good at something that makes a product possible. So much more than anything else, you're trying something that's never been tried. There's a new model, only available within the company. So the thing you ought to do is just build something on it, let people internally try it. There's this eternal September vibe to it. It's as if you discovered database technology for the first time and you're like, well, what could you build on this? And it's always the first day. That's what is different.

Host

你提到了数据库技术,这或许为思考开源提供了一个有趣的类比。第一个成功采用的关系数据库是专有的,但后来开源追赶上来。你如何保持与开源选项的差距?

You mentioned database technology and maybe that provides an interesting analogy as we think about open source. The first relational databases that were successful in terms of adoption were proprietary, but then the open source guys caught up. How do you keep the gap with the open source option?

Dario

是的。我认为开源在 AI 模型中与其他领域含义不同。因此有人称之为开放权重模型以作区分。我认为主要区别在于,如果你看到模型权重,你无法理解实际发生了什么。没有那种可组合性。我无法阅读源代码,无法生成一个略有不同的版本。现在,Anthropic 实际上在研究机械可解释性,这让你能看到模型内部。所以我们正在研究能实现某些特性的东西,但还没做到,差得远。有些事可以做。例如,如果你能访问模型,你可以微调它。我们现在通过接口让人们微调模型。所以问题在于,访问实际模型权重比一个允许你做某些事的厚 API 有多大价值?这涉及经济学问题,但请注意,在云端运行模型成本很高。必须有人托管,运行快速推理,然后你回到边际成本或部分边际成本。

Yeah. Open source, I think, has a different meaning in AI models than in other areas. For this reason, some have called it open weights models to distinguish. I think the main difference is that if you see the weights of the models, you can't understand what's actually going on. There's not that kind of composability. I can't read the source code, I can't produce a trivially different version of it. Now, Anthropic is actually working on mechanistic interpretability which allows you to see inside the models. So we're working on things that would allow some properties, but we're not there yet, not anywhere close. There are some things you can do. For example, if you have access to the model, you can fine-tune it. We're now through interfaces allowing people to fine-tune the model. So there is a question of how valuable access to the actual model weights is over and above some thick API that lets you do something. There's some question of economics, but note that it costs a significant amount to run the models on the cloud. Someone has to host it, run fast inference, and then you're back to the margin or some portion of the margin.

Host

所以你认为开放权重模型没那么有用,而完全开源模型差距很大。我想说的是,与之前技术的类比只是部分成立。这是一种我们仍在发现的不同事物。但从我们的角度看,当新模型出现,当竞争对手的模型出现时,我们并不真正考虑它是否是开放权重模型。我们考虑它是否是一个强大的模型。所以如果有人做出了一个强大的模型,在我们擅长的方面表现出色,那就是竞争,对我们不利,无论它是否是开放权重模型。两者之间没有太大区别。

So you think open-weight models are not that useful, and fully open source models there's just a big gap. I guess what I would say is that the analogy to previous technologies is only partial. It's kind of a different thing that we're still discovering. But I can say from our perspective that when a new model comes out, when a competitor model comes out, we don't really think about whether it's an open weights model or not. We think about whether it's a strong model. So if someone makes a strong model that is good at the things we do, that's competition, that's bad for us, whether it's an open weights model or not. There's not a huge difference between the two.

Host

Anthropic 如何比其他组织更注重 AGI?一是更快,比如更紧凑的产品发布节奏,但可能更广泛地贯穿整个组织,而不仅仅是产品开发。

How is Anthropic more AGI-pill than other organizations? One is faster, like a tighter product release cadence, but maybe more broadly across the organization, not just within product development.

Dario

是的。我提到过,每隔几周我会站在组织面前描述我的愿景。我认为目的之一是让人们专注于使命。

Yeah. I mentioned this thing that every couple of weeks I get up in front of the organization and describe my vision. I think one of the purposes of that is to keep people focused on the mission.

世界怪现状与公司理念 The Strange State of the World and Company Thesis

Dario

这是一个奇怪的世界状态,我对此总是表示不确定,但如果要我下注,我会押注这一点:在一两年或三年内,我不确定具体多久,我们将在数据中心拥有一个我所说的「天才之国」。这很奇怪。它将改变经济,加速科学进步,带来全球对齐和国家安全风险,也可能引发经济问题。上行空间巨大,破坏潜力也同样巨大。我试图反对的是这样一种想法:员工加入后认为,「哦,我曾在某个行业工作,曾在某类公司工作,现在我要去一家 AI 公司,也许几年后我会去别处。」这与以往截然不同。这是一件完全不同的事情。在整个组织中,我们要确保当财务人员考虑财务预测时,他们明白这不仅是指数增长,而且可能出现极端结果。当招聘人员考虑薪酬时,他们明白疯狂的薪酬可能发生。当产品人员思考时,他们制造 AGI 产品。当政策人员互动时,他们理解其中的利害关系。我工作的一大部分是让组织围绕这个核心论点保持一致。不是每个人都必须相信这个论点,这不是洗脑,但基本想法是公司建立在这样一个假设之上:这些巨大变化是可能的,甚至很可能发生。业务和社会效益的每个方面都应围绕这一可能性来构建。

It's a strange state of the world and I always express uncertainty about it, but if I were to bet, I would bet in favor of this: in one, two, or three years, I don't know exactly how long, we'll have what I've described as a country of geniuses in the data center. This is weird. It's going to change the economy, accelerate the pace of science, pose global alignment and national security risks, and may pose economic problems. The upside is huge, and the potential for disruption is also huge. What I'm trying to fight against is the idea that employees join and think, 'Oh, I worked in this industry, I worked at this kind of company, and I'm going to work at an AI company, and maybe a couple years later I'll go to this.' This is categorically different from previous. This is a really different thing. Up and down the organization, we want to make sure that when our finance people think about financial projections, they understand this is not just exponential but wild outcomes are possible. When our recruiting thinks about comp, they understand crazy comp could happen. When the product people think, they make AGI products. When the policy people interact, they understand the stakes. A big part of my job is keeping the coherence of the organization around this central thesis. Not that everyone has to believe the thesis, it's not indoctrination, but the basic idea is that the company is built around the hypothesis that it is possible and perhaps likely that these large changes will happen. Every aspect of the business and social benefit should be constructed around the strong possibility that this may happen.

Host

具体来说,你曾谈到 AI 可能带来 10%的年经济增长率。这是否意味着当我们谈论 AI 风险时,通常指的是 AI 的危害和滥用?难道最大的 AI 风险不是我们稍微监管过度或减缓了进展,从而因为 AI 不足而损失了大量人类福祉吗?

To put numbers on this, you've talked about the potential for 10% annual economic growth powered by AI. Doesn't that mean that when we talk about AI risk, it's often harms and misuses of AI? Isn't the big AI risk that we slightly misregulate it or slow down progress and therefore there's a lot of human welfare missed out on because you don't have enough AI?

Dario

我有过家人死于疾病的经历,而这些疾病在他们去世几年后就被治愈了。所以我真正理解进展不够快的代价。我要说的是,AI 的一些危险有可能严重破坏社会稳定或威胁人类文明。所以我们不能对这种级别的风险掉以轻心。我绝不主张停止或暂停这项技术。出于多种原因,我认为这根本不可能。我们有地缘政治对手,他们不会停止开发这项技术,而且涉及的资本量巨大——哪怕你提出最轻微的监管,也会有数万亿美元的资本反对我,因为那不符合他们的利益。这显示了可能性的极限。但是,与其考虑放慢速度还是全速前进,有没有办法引入安全措施,要么不减缓技术发展,要么只减缓一点点?如果我们可以实现 9%的经济增长而不是 10%,并为所有这些风险购买保险,我认为这才是真正的权衡。正因为 AI 有可能快速发展并解决许多问题,我认为更大的风险是事情过热。所以我不想停止反应,而是想引导它。

I've had the experience where family members died of diseases that were cured a few years after they died. So I truly understand the stakes of not making progress fast enough. I would say that some of the dangers of AI have the potential to significantly destabilize society or threaten humanity or civilization. So we don't want to take idle chances with that level of risk. I'm not at all an advocate of stopping the technology or pausing it. For a number of reasons, I think that's just not possible. We have geopolitical adversaries who won't stop making the technology, and the amount of money involved—if you even propose the slightest regulation, I have many trillions of dollars of capital lined up against me for whom that's not in their interest. That shows the limits of what is possible. But instead of thinking about slowing it down versus going at maximum speed, are there ways to introduce safety and security measures that either don't slow the technology down or only slow it a little? If instead of 10% economic growth, we could have 9% and buy insurance against all these risks, I think that's what the trade-off actually looks like. Precisely because AI has the potential to go so quickly and solve so many problems, I see the greater risk as the thing overheating. So I don't want to stop the reaction; I want to focus it.

Host

你曾说如果到了 2025 年 12 月还没有 AI 法律,你会非常担心。你现在感觉如何?

You said if we hit December 2025 and there's no AI law, you'll be really worried. How are you feeling?

Dario

加州确实有一项法案,SB53。去年我们有 SB 1047。我们对 SB 1047 感受复杂。最初的版本过于激进。我的意思是,技术发展很快,过于具体的规定没有帮助,最终无助于安全。我担心如果这样的法案通过,规定的测试会看起来很愚蠢,然后 AI 行业的人会想,「哦,这就是安全和监管的样子。真蠢。」他们不会认真对待,只会字面遵守而非精神上遵守。作为深思熟虑监管的倡导者,我对此感到担忧。我们提出了一些修改意见,直到我们觉得满意,并试图在行业和安全倡导者之间达成妥协。但我们没有成功。不过今年我们正在制定一个更温和的法案,特别关注实践的透明度,安全和安保实践的透明度,这是 Anthropic 一直积极推动的。其他公司也开始这样做,但并非所有公司,而且无法判断他们披露的内容是否真实。

There is actually something in California. There's a bill out SB53. Last year we had the whole SB 1047 thing. We had mixed feelings on SB 1047. The initial version was too aggressive. When I say that, I mean the technology is moving fast and it's unhelpful if you're too prescriptive. It ends up not contributing to safety. I was worried that if something like that passes, the prescribed tests would look stupid, and then people in the AI industry would think, 'Oh, this is what regulation for safety and security looks like. It's really stupid.' They wouldn't take it seriously and would comply only in letter, not in spirit. As an advocate of thoughtful regulation, I was concerned. We offered some changes to the bill to a point where we felt good about it and tried to make a compromise between industry and safety advocates. We didn't really succeed. But this year we're making a bill that is more moderate, focused particularly on transparency of practices, transparency of safety and security practices, which is something Anthropic has been very forward about. Other companies are starting to do it, but not all, and there's no way to tell if folks are telling the truth about what they're revealing.

Host

加州的监管就足够了,因为所有公司都在这里有业务联系。

And California regulation is enough because all the companies have nexus here.

Dario

是的。大多数法案都是围绕在加州开展业务来组织的。所以很难切断。这里的人非常关注 AI。

Yeah. Most of these bills are organized around doing business in California. So it would be difficult to shut off. People are very AI here.

监管策略 Regulatory Approach

Dario

我不确定会发生什么,但我们一直持这样的态度:我们支持护栏,包括针对该技术的立法护栏。但我们认识到需要谨慎。我们不想杀死下金蛋的鹅。我们只是想阻止它过热或偏离轨道。

I'm not sure what's going to happen, but we've always had this approach that we are in favor of guardrails, including legislative guardrails on the technology. But we recognize the need to be careful. We don't want to kill the golden goose. We just want to stop it from overheating or running off the road.

Host

是的。也许像现代银行监管这样的东西,尽管人们抱怨,却是一个很好的例子,说明一种固有高风险的活动。

Yeah. Maybe something like modern bank regulation, for all people complain, is a good example where there's an inherently very risky activity.

Dario

是的。不,危险是相当明显的。我的意思是,就像银行挤兑之类,对吧?但一旦我们弄清楚了监管环境,它在现代时代运行得相当好。

Yeah. No, the dangers are pretty clear. I mean, it's like bank runs or not, right? But it all works pretty well in the modern era once we figured out the regulatory environment.

个人 AI 使用 Personal AI Usage

Host

最后一个问题。你个人的 AI 工具集是什么?你使用 AI 的方式与科技界的其他人有何不同?

Last question. What is your personal AI stack? How do you use AI differently to maybe other people in tech?

Dario

是的。有趣。我基本上写很多东西。也许我对自己写作的自尊心太强了。我用 Claude 来生成很多想法。我把它当作研究工具,但到目前为止我自己写作。Claude 实际上可能比其他模型更接近,但仍然不够好。对于商务邮件我会放心使用,但如果我在写一篇文章或一些我想真正做对的东西,它还不够好,但也许一年左右就会达到。

Yeah. Interesting. I basically write a lot. Perhaps I have too much pride in my own writing. I use Claude to generate lots of ideas. I kind of use it as research, but so far I've done the writing myself. Claude is actually maybe closer than the other ones, but it's still not there. I'd be comfortable with it for business emails, but if I'm writing an essay or something I want to really get right, it's not quite there yet, but maybe it will be in a year or so.

Host

是的。非常酷。那么,Dario,这太棒了。谢谢你的到来。

Yeah. Very cool. Well, Dario, this is awesome. Thanks for coming by.

Dario

谢谢你邀请我。

Thank you for having me.

互动版:逐字朗读 + 针对本期提问 →