超大规模资本支出与 AI 实验室算力扩展

Hyperscaler Capex and AI Lab Compute Scaling

迪伦·帕特尔 Dylan Patel · Dwarkesh 播客 · 2026-03-13 · 约 151 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Semi Analysis 的 Dylan 解析 6000 亿美元超大规模资本支出、20GW 部署时间线,以及 Anthropic 和 OpenAI 为何需要大规模算力增长。

Dylan from Semi Analysis breaks down the $600B hyperscaler capex, 20GW deployment timeline, and why Anthropic and OpenAI need massive compute growth.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 45)

全文 · Full transcript(中英对照)

引言与资本支出概览 Introduction and Capex Overview

Host

好的。这是《我的室友教我半导体》这一集。也是这一系列的告别篇。是啊,你用完之后,我就觉得,我不能再用了。我得离开这里。那些二手货给书呆子。是的。好的。Dylan 是 SemiAnalysis 的 CEO。Dylan,我有个核心问题。如果你把四大巨头——亚马逊、Meta、谷歌、微软——加起来,你最近发布的他们今年预测的资本支出总额是 6000 亿美元。考虑到每年租用这些算力的价格,那将接近 50 吉瓦。显然我们今年不会投入 50 吉瓦。所以这很可能是在为未来几年陆续上线的算力买单。所以我想问的是,如何思考这些资本支出上线的时间线?类似的问题也适用于实验室:OpenAI 刚刚宣布他们筹集了 1100 亿美元。Anthropic 刚刚宣布他们筹集了 300 亿美元。如果你看看他们今年即将上线的算力——你应该告诉我具体是多少——但今年他们总共不是还有另外 4 吉瓦吗?感觉 OpenAI 和 Anthropic 今年租用算力的成本,按每吉瓦 130 亿美元算,光是这些单独的融资就足以覆盖他们全年的算力支出,这还不包括他们今年将赚取的收入。所以请帮我理解:第一,大型科技公司的资本支出实际上什么时候上线?第二,如果 1 吉瓦数据中心的年租金是 130 亿美元,这些实验室筹集这么多钱是为了什么?

All right. This is the episode of My Roommate Teaches Me Semiconductors. It's also the sendoff for this current set. Yeah, after you use it, I'm like, I can't use this again. I got to get out of here. Those sloppy seconds for dork. Yes. Okay. Dylan is the CEO of SemiAnalysis. Dylan, the overarching question I have for you. If you add up the big four—Amazon, Meta, Google, Microsoft—their combined forecasted capex that you published recently this year is $600 billion. And given yearly prices of renting that compute, that would be close to 50 gigawatts. Now obviously we're not putting on 50 gigawatts this year. So presumably that's paying for compute that is going to be coming online over the coming years. So I have a question about how to think about the timeline around when that capex comes online. Similar question for the labs: OpenAI just announced they raised $110 billion. Anthropic just announced they raised $30 billion. And if you look at the compute they have coming online this year—you should tell me how much it is—but isn't it another 4 gigawatts total that they'll have this year? It feels like the cost to rent the compute that OpenAI and Anthropic will have this year to sustain their compute spend at, say, $13 billion per gigawatt—those individual raises alone are enough to cover their compute spend for the year, and this is not even including the revenue they're going to earn this year. So help me understand: first, when is the time scale at which the big tech capex is actually coming online? And second, what are the labs raising all this money for if the yearly price of a 1-gigawatt data center is like $13 billion?

Dylan Patel

当你谈论这些超大规模企业的资本支出,大约 6000 亿美元,再看看供应链的其他部分,总额会达到大约 1 万亿美元。其中一部分是今年立即用于上线的算力——芯片和其他今年确实要支付的资本支出。但也有大量前期资本支出。当我们谈论美国今年大约 20 吉瓦的增量容量时,其中一部分并不是今年花的;一部分资本支出实际上是在前一年花的。所以当你看到谷歌的 1800 亿美元时,实际上很大一部分花在了 2028 年和 2029 年的涡轮机定金上。一部分花在了 2027 年的数据中心建设上。一部分花在了购电协议和首付以及他们为未来更远时期所做的所有其他事情上,以便能够实现这种超快速扩张。这适用于所有超大规模企业和供应链中的其他人。所以,今年大约部署了 20 吉瓦。其中很大一部分来自超大规模企业,一部分不是。所有这些公司最大的客户是 Anthropic 和 OpenAI。Anthropic 和 OpenAI 目前大约在 2 吉瓦、2.5 吉瓦和 1.5 吉瓦的范围内。他们正试图扩展到更大的规模。如果你看看 Anthropic 在过去几个月做了什么——增加了 40 亿、60 亿美元的收入——如果我们简单地画一条直线,他们每个月还会再增加 60 亿美元的收入。人们会认为这很悲观,他们应该更快。这意味着他们在未来 10 个月内将增加 600 亿美元的收入。按照 Anthropic 目前(至少媒体上次报道的)的毛利率,600 亿美元的收入意味着他们需要大约 400 亿美元的算力支出来支持这 600 亿美元收入的推理。这 400 亿美元的算力,按每吉瓦大约 100 亿美元的租赁成本计算,意味着他们需要增加 4 吉瓦的推理容量才能仅仅实现收入增长,而且这还假设他们的研发训练集群保持不变。所以,从某种意义上说,Anthropic 需要在今年年底前达到远高于 5 吉瓦的水平,这对他们来说非常困难,但也是可能的。

So when you talk about the capex of these hyperscalers, on the order of $600 billion, and you look across the rest of the supply chain, that gets you to on the order of a trillion dollars. A portion of this is immediately for compute going online this year—the chips and the other parts of capex that do get paid this year. But there's a lot of setup capex as well. When we're talking about 20 gigawatts this year in America roughly—incremental added capacity—a portion of this is not spent this year; a portion of that capex is actually spent the prior year. So when you look at Google's $180 billion, actually a big chunk of that is spent on turbine deposits for 2028 and 2029. A chunk is spent on data center construction for 2027. A chunk is spent on power purchasing agreements and down payments and all these other things they're doing further out into the future so that they can set up this super fast scaling. And this applies to all the hyperscalers and other people in the supply chain. So, 20 gigawatts roughly deployed this year. A big chunk of that being hyperscalers, a chunk not. And all of these companies' biggest customers are Anthropic and OpenAI. Anthropic and OpenAI are in the 2 gigawatt, 2.5 gigawatt, and 1.5 gigawatt range roughly right now. They're trying to scale to much larger. If you look at what Anthropic has done over the last few months—$4 billion, $6 billion revenue added—and if we just draw a straight line, they'll add another $6 billion of revenue a month. People would argue that's bearish and that they should go faster. What that implies is that they're going to add $60 billion of revenue across the next 10 months. And $60 billion of revenue at the current gross margins that Anthropic had, at least last reported by media, would imply that they have roughly $40 billion of compute spend for that inference for that $60 billion of revenue. That $40 billion of compute at roughly $10 billion per gigawatt rental cost means that they need to add 4 gigawatts of inference capacity just to grow revenue, and that's saying that their research and development training fleet stays flat. So, in a sense, Anthropic needs to get to well above 5 gigawatts by the end of this year, and it's going to be really tough for them to get there, but it's possible.

Host

我能问个问题吗?所以,如果 Anthropic 今年年底前无法达到 5 吉瓦,但它需要这么多算力来服务比预期更疯狂的收入(可能还会更多),再加上研究和训练以确保其模型明年足够好,那么这些算力从哪里来?你知道,Dario 在你的播客上时非常保守。他说,我不会在算力上发疯,因为如果我的收入以不同的速度膨胀,在某个时间点,我不想破产。我想确保我们在扩展时负责任。但实际上,他肯定错过了像 OpenAI 那样的机会,后者就是签下这些疯狂的交易。而 OpenAI 到今年年底获得的算力比 Anthropic 多得多。那么 Anthropic 必须做什么才能获得算力?他们必须去找以前不会用的低质量供应商。理想情况下,Anthropic 历史上一直拥有最好的供应商,比如谷歌和亚马逊。而历史上,世界上最大的公司——现在微软——他们正在扩展供应链,去接触更新的参与者。OpenAI 在接触多个参与者方面更加激进。是的,他们从微软获得了大量容量。他们也有谷歌和亚马逊,但他们还与 CoreWeave 和 Oracle 有大量合作,并且他们去了随机公司——或者人们认为的随机公司——比如软银能源,这家公司一生从未建过数据中心,但现在正在为 OpenAI 建设数据中心。他们还去了许多其他公司,比如 Nscale 等,从中获取容量。所以这对 Anthropic 来说是一个难题,因为他们在算力上太保守了,不想发疯。从某种意义上说,去年下半年很多金融恐慌都是:OpenAI 签了所有这些交易,但他们没有钱支付。甲骨文股票要暴跌。CoreWeave 股票要暴跌。所有这些公司的股票都暴跌了。信贷市场也疯了,因为人们认为最终买家付不起。现在,哦等等,他们筹集了一大笔钱。好吧,他们能付得起了。

Can I ask a question about that? So, if Anthropic was not on track to have 5 gigawatts by the end of this year, but it needs that to serve both the revenue that's gone crazier than expected and maybe it's going to be even more than that, plus the research and training to make sure its models are good enough for next year. How is that going to come from? You know, Dario when he was on your podcast was very conservative. He's like, I'm not going to go crazy on compute because if my revenue inflates at a different rate, at a different point, I don't want to go bankrupt. I want to make sure that we're being responsible with this scaling. But in reality, he's definitely missed the pooch in terms of going like OpenAI, which was let's just sign these crazy deals. And OpenAI is kind of got way more access to compute than Anthropic by the end of the year. And so what does Anthropic have to do to get the compute? Well, they have to go to lower quality providers that they would not have gone to before. Optimally, Anthropic at least historically has had the best quality providers, like Google and Amazon. Whereas, at least historically minded, the biggest companies in the world, now Microsoft and now they're expanding across the supply chain and going to other players that are newer. OpenAI has been a bit more aggressive on going to many players. Yes, they have tons of capacity from Microsoft. They have Google and Amazon as well, but they also have tons with CoreWeave and Oracle, and they've gone to random companies—or one would think random companies—like SoftBank Energy, who has never built a data center in their life, but they're building data centers now for OpenAI. So they've gone to many others like Nscale and others that they're going and getting capacity from. So there's this conundrum for Anthropic because they were so conservative on compute because they didn't want to go crazy. And in some sense, a lot of the financial freakouts in the second half of last year were like: OpenAI signed all these deals, but they don't have the money to pay for them. Oracle stock's going to tank. CoreWeave stock's going to tank. All these companies' stocks tanked. And credit markets went crazy because people were like, the end buyer can't pay for this. Now it's like, oh wait, they raised a ton of money. Okay, fine. They can pay for it.

Anthropic保守计算策略vs OpenAI Anthropic's conservative compute strategy vs OpenAI

Host

但从某种意义上说,Anthropic 要保守得多。他们会说,我们会签合同,但我们会坚持原则,故意低估我们认为可能做到的事情,保持保守,因为我们不想可能破产。

But in the sense, Anthropic was a lot more conservative. They're like, we'll sign contracts, but we'll be principled and we'll purposely undershoot what we think we can possibly do and be conservative because we don't want to potentially go bankrupt.

Host

但我想理解的是,在紧急情况下获取算力意味着什么?是不是必须去找 NeoCloud 这样的公司?他们的算力更差吗?差在哪里?是不是因为临时抱佛脚,不得不向云提供商支付原本不需要支付的毛利率?谁建造了备用容量,让 Anthropic 和 OpenAI 能在最后一刻拿到?基本上,如果到 2027 年他们最终拥有相似的算力规模,OpenAI 到底获得了什么具体优势?是不是只是他们今年年底的千兆瓦数不同?如果是,Anthropic 和 OpenAI 到今年年底各会有多少千兆瓦?

But the thing I want to understand is, what does it mean to have to acquire compute in a pinch? Is it that you have to go with like NeoClouds? Is it that they have worse compute? Or in what way is it worse? Is it that you had to pay gross margins to a cloud provider that you wouldn't have otherwise had to pay because you're coming in at the last minute? Who built the spare capacity such that it's available for Anthropic and OpenAI to get last minute? And basically, what is the concrete advantage that OpenAI has gotten if they end up at similar compute numbers by 2027? Is it just like they're going to end this year with different gigawatts? If so, how many gigawatts is Anthropic and OpenAI going to have by the end of this year?

Dylan Patel

是的。要获取多余的算力,超大规模云服务商确实有容量。而且并非所有算力合同都是长期的,对吧,5 年。有些算力,比如 2023 或 2024 年的 H100,2025 年的,签的就不是 5 年合同。没错,OpenAI 的绝大部分算力签的是 5 年合同,但他们可以,你知道,还有很多其他客户签了 1 年、2 年、3 年、6 个月的按需合同。随着这些合同到期,谁是最愿意出价的参与者?从这个意义上说,我们看到 H100 的价格大幅上涨,人们愿意签长期合同,甚至超过 2 美元,对吧?比如我看到一些 AI 实验室(我出于原因会说得模糊一点)签了高达 2.40 美元、为期 2 到 3 年的 H100 合同。想想利润率,Hopper 刚推出时成本是 1.40 美元,或者分摊到 5 年,现在两年过去了,你签的是 2 到 3 年、2.40 美元的合同,利润率高得多。对吧,所以现在你可以挤掉所有其他供应商,无论是 Amazon、CoreWeave、Together AI、Nebius 还是其他公司。这些 NeoCloud 公司通常拥有更高比例的 Hopper,因为他们在这方面更激进,而且他们倾向于签短期合同。你知道,不是 CoreWeave,但其他公司倾向于签短期合同。所以,嘿,如果我想要 Hopper,外面还有一些容量。另外,虽然 Oracle 或 CoreWeave 的大部分 Blackwell 容量都签了长期合同,但本季度上线的任何东西都已经卖掉了。在某些情况下,他们甚至没有达到承诺的销售数字,因为有些数据中心延迟了。不只是这两家,还有 Nebius 和其他公司,微软、亚马逊、谷歌。但有很多 NeoCloud 以及一些超大规模云服务商正在建设尚未售出的容量,或者原本打算分配给某些内部用途(不一定专注于超级 AGI)的容量,他们现在可能转而出售。或者,在 Anthropic 的情况下,他们不必直接拥有所有算力。对吧,Amazon 可以拥有算力,他们提供 Bedrock 服务;或者 Google 可以拥有算力,提供 Vertex 服务;或者微软可以拥有算力,提供 Foundry 服务,然后与 Anthropic 进行收入分成,反之亦然。

Yeah. So to acquire excess compute, yes, there is capacity at hyperscalers. And not all contracts for compute are long-term, right, 5 years. There's compute that in 2023 or 2024 H100, 2025 that were signed at not 5-year deals. Right, OpenAI, the vast majority of their compute is signed at 5-year deals, but they can, you know, there were many other customers that had one-year, two-year, three-year deals, six-month deals on demand. And as these contracts roll off, who is the participant most willing to pay price? And in this sense, right, we've seen H100 prices inflect a lot and go up, and people willing to sign long-term deals for, you know, as above $2 even, right? Like I've seen deals where certain AI labs, I'll be a little bit vague here for a reason, have signed at as high as $2.40 for two to three years for H100s. Which, if you think about the margin, $1.40 for Hopper when you release it, or Hopper to build it across 5 years, and now 2 years in, you're signing deals that are two to three years that are at $2.40, those margins are way higher. Right, and so now you can crowd out all of these other suppliers, whether it's Amazon had these, or CoreWeave had these, or Together AI, or Nebius, or whoever it is. Right, you know, these NeoClouds are the firms that had a higher percentage of Hopper in general because they were more aggressive on it, and they tended to sign shorter term deals. You know, not CoreWeave, but the others tended to sign shorter term deals. And so, hey, if I want Hopper, there is some capacity out there. And then also, while most of the capacity at like an Oracle or a CoreWeave is signed for a long-term deal in terms of Blackwell, anything that's going online this quarter is already sold. And in some cases, they're not even hitting all the numbers that they promised they would sell because there's some data center delays. Not just those two but like Nebius and all the other folks, Microsoft, Amazon, Google. But there is a lot of NeoClouds as well as some of the hyperscalers who have capacity they're building that they did not sell yet, or capacity that they were going to allocate to some internal use that is not necessarily super AGI focused, that they may now turn around and sell. Or they may, you know, in the case of Anthropic, they don't have to have all the compute directly. Right, Amazon can have the compute, they can serve Bedrock, or Google can have the compute and serve Vertex, or Microsoft can have the compute and serve Foundry, and then do a revenue share within Anthropic or vice versa.

Host

基本上,你是说 Anthropic 不得不支付大约 50% 的溢价,要么是收入分成,要么是最后一刻的现货算力,如果他们早点购买算力,就不必支付这些费用,对吧?

Basically, you're saying Anthropic is having to pay either this like 50% markup in the sense of the revenue share or in the sense of last minute spot compute that they wouldn't have otherwise had to pay had they bought the compute early, right?

Dylan Patel

你知道,这里有一个权衡。但同时,整整 4 个月,每个人都像在说:“OpenAI,我们不会和你签合同。”这听起来很疯狂,对吧,因为你们没有钱。现在每个人都像在说:“是的 OpenAI,我们一直相信你。我们可以签任何合同,因为你已经筹集了所有这些钱。”但从某种意义上说,Anthropic 在这方面受到了限制。算力的增量买家还不多,因为 Anthropic 首先达到了能力层级,他们的收入……这很有趣,因为否则我们会说,拥有最好的模型是一种极度贬值的资产,3 个月后你就没有最好的模型了。但重要的是,你可以签这些合同,然后提前锁定算力,获得更好的价格。

And you know there's a trade-off there. But also at the same time, you know, for a solid like 4 months everyone was like, 'OpenAI, we're not going to sign deals with you.' That sounds crazy, right, because you guys don't have the money. Now everyone's like, 'Yeah OpenAI, we believed you the whole time. We can sign any deal because you've raised all this money.' But in a sense, Anthropic is constrained in that sense. There are not that many incremental buyers of compute yet because Anthropic hit the capabilities tier first where their revenues. That's interesting, like that's this, you know, because otherwise we're like, well having the best model is an extremely depreciating asset that, you know, 3 months later you don't have the best model. But like the reason it's important is that you can sign these deals and then lock in the compute in advance, get better prices.

Host

顺便问一下,这难道不也意味着——也许这是个显而易见的观点——但至少直到最近,人们还在大谈特谈 GPU 的折旧周期,空头们,比如 Michael Burry 之类的人,说人们说这些 GPU 能用四五年,而实际上,也许是因为技术改进太快或其他原因,对这些 GPU 采用两年折旧周期可能更合理,这会增加给定年份的报告摊销资本支出。因此,建造所有这些云服务在财务上可能不那么有利可图。但事实上,你指出折旧周期可能甚至超过 5 年,因为如果我们使用 Hopper,尤其是如果 AI 真的起飞,到 2030 年我们可能得启动 7 纳米工厂,我们可能得回到 A100,重新开启 A100。那么实际上折旧周期非常长,所以我觉得这是你所说的一个有趣的财务含义。

Doesn't this also imply, by the way, and maybe this is an obvious point, but there's at least until recently people had made this huge point about, oh what is the depreciation cycle of a GPU, and the bears, Michael Burry or whatever, have said look, people are saying that four or five years for these GPUs, and in fact if you maybe it's because the technology is improving so fast or whatever, it might make sense to have two-year depreciation cycles for these GPUs, which increases the sort of like reported amortized capex in a given year. And so makes it maybe financially less lucrative to building all these clouds. But in fact, you're pointing at like maybe the depreciation cycle is even longer than 5 years, because if we're using Hoppers, and then especially if AI really takes off and in 2030 we're like we got to like get the 7 nanometer fabs up and we got to like we got to go back to the A100s, like turn on the A100s again. Then it's like actually the depreciation cycle is incredibly long, and so I feel like that's an interesting financial implication of what you're saying.

Dylan Patel

这里有几个线索可以展开。一个是 GPU 的折旧会怎样,对吧?我想我没有回答你之前的问题,关于 Anthropic,我认为他们到年底能够达到大约 5 千兆瓦,可能再多一点,通过他们自己以及通过 Bedrock、Vertex 或 Foundry 提供的产品。我认为他们能达到 5 到 6 千兆瓦。

There are a few strings to pull on there. One is what happens to depreciation of GPUs, right? And I guess I didn't answer your prior question, which is like Anthropic, I think will be able to get to like 5 gigawatts-ish, maybe a little bit more by the end of the year, through themselves as well as their product being served through Bedrock or through Vertex or through Foundry. I think they'll be able to get to five or six gigawatts.

GPU折旧与TCO模型 GPU depreciation and TCO model

Dylan Patel

这远远超出了他们最初的计划,对吧?总之,这有点像是一个开始。它大致相同,可能略高一点,实际上根据我们的数据会更高一些。但无论如何,GPU 的折旧周期,对吧?Michael Burry 说它是三年或更短,对吧?这就是他的论点。有两种方式或视角来看待这个问题。从机械角度看,有一个 TCO 模型,对吧?GPU 的总拥有成本,我们大致预测 GPU 的价格并构建集群的总成本。但有很多成本,对吧?有你的数据中心成本,对吧?有你的网络成本。有你的智能手和人在数据中心更换东西的成本。有你的备件成本,对吧?有你的实际芯片成本。有你的服务器成本。所有这些各种成本被整合在一起,还有一些折旧周期。还有一些信贷成本。然后你得到,好吧,这就是你如何构建的。嘿,一个 H100 在 5 年折旧期内的批量部署成本是每小时 140 美元。然后如果你签署一份为期五年的每小时 2 美元的协议,你的毛利率大约是 35%。略高于此,但如果你签的是每小时 1.90 美元,毛利率大约是 35%。然后在第五年,GPU 就报废了,对吧?它死了。在某些情况下,人们提出的论点是,如果你没有签署长期协议,因为每两年,Nvidia 的性能就会翻三倍、翻四倍,而价格只翻倍或增加 50%,那么 H100 的价格,当然,也许 2024 年市场价值是每小时 2 美元,毛利率 35%,但在 2026 年,当 Blackwell 达到超高产量并每年部署数百万个时,你实际上现在只值每小时 1 美元。而当 Rubin 在 27 年达到超高产量时,对吧?尽管它今年开始出货,但明年还不是超高产量,每年部署数百万个芯片到云中,你又获得了 3 倍的性能和 50% 或 2 倍的价格提升。实际上,Hopper 只值每小时 70 美分。所以 GPU 的价格会继续下降。这是一种视角。

Uh which is way above their initial plans, right? And anyways, that's sort of like an opening eye. It will be a little roughly the same, maybe a little higher, actually a little bit higher based on our numbers. But anyways, the depreciation cycle of a GPU, right? Michael Burry was saying it's three years or less, right? That's sort of his argument. And there's sort of two ways or lenses to look at this. Mechanically, there's a TCO model, right? Total cost of ownership of a GPU, where we sort of project pricing out for GPUs and build up the total cost of a cluster. But there's a number of costs, right? There's your data center cost, right? There's your networking cost. There's your smart hands and people in the data center swapping stuff out. There's your spare parts, right? There's your actual chip cost. There's your server cost. All these various costs get lumped together, and there's some depreciation cycles on it. There's certain credit costs on it. And you get to, okay, that's how you build up. Hey, an H100 costs $140 an hour to deploy at volume across 5 years if your depreciation is 5 years. And then if you sign a deal at $2 an hour for those five years, your gross margin is roughly 35%. It's a little bit above that, but if you sign it for $1.90, it's 35% roughly. And then at that fifth year, the GPU falls off a bus, right? It's dead. And in some cases, the argument people are making is, well, if you didn't sign a long-term deal because every two years, Nvidia's tripling, quadrupling the performance while only 2xing the price or 50% increasing the price, then the price of an H100, sure, maybe the value in the market was $2 at 35% gross margins in 2024, but in 2026 when Blackwell is in super high volume and deploying millions a year, you're actually now worth a dollar an hour. And when Rubin in 27 is in super high volume, right? Even though it starts shipping this year, isn't super high volume next year, doing millions of chips a year deployed into clouds, you've got another 3x in performance and another 50% or 2x in price. Actually, the Hopper is only worth 70 cents an hour. And so the price of a GPU would continue to fall. That's like one lens.

Host

这太疯狂了。

That's crazy.

GPU效用价值与模型改进 Utility value of GPUs and model improvements

Dylan Patel

另一种视角是你从芯片中获得的效用,对吧?因为如果你能制造无限的 Rubin 或无限的最新芯片,那么是的,这正是会发生的情况。随着新芯片的推出和每价格性能的提升,Hopper 的价格会在现货或短期合同利率下下降。但由于你在半导体和部署时间表等方面受到很大限制,最终给这些芯片定价的不是“嘿,我今天能买到什么比较的东西?”而是“我今天能从这块芯片中获得什么价值?”从这个意义上说,我们以 GPT-5.4 为例。GPT-5.4 的运行成本比 GPT-4 便宜得多。它的活跃参数更少。它小得多,在活跃参数的意义上,再加上它是一个稀疏模型,而 GPT-4 是一个密集模型。在训练、强化学习、模型架构、数据质量等方面还有很多其他进步,所有这些都使 GPT-5.4 比 GPT-4 好得多,而且服务成本更低。所以当你看到 H100 时,它每 GPU 可以服务更多 GPT-5.4 的 token,比你在上面运行 GPT-4 要多。对吧?所以在某种意义上,它正在产生更多更高质量模型的 token。所以在某种意义上,显然 GPT-4,其 token 的最大 TAM 是多少?也许是几十亿,也许是几百亿美元。采用需要时间。对于 GPT-5.4,这个数字可能超过 1000 亿,但存在采用滞后和竞争,所以其他人也在获得它,而且每个人都在不断改进。所以如果改进在这里停止,H100 的价值现在取决于 GPT-5.4 能从中获得的价值,而不是 GPT-4 能获得的价值,以及这些实验室所做的利润和所有这些东西,而且他们处于竞争环境中,所以他们的利润不能无限大。所以你有一种相当有趣的动态,即 H100 今天比三年前更值钱。

The other lens is what is the utility you get out of the chip, right? Because if you could build infinite Rubin or infinite of the newest chip, then yes, that's exactly what would happen. The price of a Hopper would fall at a spot or a short-term contract rate as the new chips come out and the per price performance goes up. But because you are so limited on semiconductors and deployment timelines and all these things, you end up with actually what prices these chips is not, hey, what's the comparative thing I can buy today? It's actually what is the value I can derive out of this chip today. And in that sense, let's take GPT-5.4. GPT-5.4 is both way cheaper to run than GPT-4. It has fewer active parameters. It's much smaller, in that sense of active parameters, plus because it's a sparse model versus GPT-4 being a dense model. There's also been so many other advancements in training, RL, model architecture, etc., data qualities, all these things that have made GPT-5.4 way better than GPT-4, and it's cheaper to serve. And so when you look at an H100, it can serve more tokens per GPU of GPT-5.4 than if you had run GPT-4 on it. Right? So in some sense, it's producing more tokens of a model that is of higher quality. And so in some sense, obviously GPT-4, what is the maximum TAM for its tokens? Maybe it was a few billion, maybe it was tens of billions of dollars. Adoption takes time. For GPT-5.4, that number is probably north of 100 billion, but there's an adoption lag and there's competition, so other people are getting it, and there's the constant improvements that everyone else is having. So if improvement stopped here, the value of an H100 is now predicated on the value that GPT-5.4 can get out of it instead of the value that GPT-4 can get out of it, and the margins and all that stuff that these labs are doing, and they're in a competitive environment, so their margins can't go to infinity. So you sort of have this dynamic that is quite interesting in that an H100 is worth more today than it was 3 years ago.

Host

这太疯狂了。我的意思是,从向前推进的角度来看,这也很有趣。如果我们开发了真正的 AGI 模型,如果我们有真正的人类在服务器上,以每秒浮点运算次数为基础,H100,这些关于大脑能做多少次浮点运算的数字非常粗略,但以每秒浮点运算次数为基础,H100 估计为 1E15,这大约是一些人估计人脑的浮点运算量。显然在内存方面,人脑要多得多。H100 大约是 80 GB,而大脑可能有 PB 级。

That's crazy. And I mean it's also interesting from the perspective of just take that forward. If we had actual AGI models developed, if we had like genuinely human on a server, and a human like on a flop basis, an H100, these are such handwavy numbers about how many flops can the brain do, but on a flop basis, an H100 is estimated to one E15 is like how much some people estimate the human brain does in flops. Obviously in terms of memory, the human brain has way more. H100 is like 80 gigabytes and brain might have petabytes.

Dylan Patel

哦,是的,你有 PB 级。给我说出一个 PB 的 1 和 0,兄弟。给我一个字符串。

Oh yeah, you've got petabytes. Name me a petabyte of ones and zeros, bro. Name me a string.

Host

嗯,这实际上是关键点,实际上在……

Well, this is actually the point where like actually in...

Dylan Patel

不,我们只是拥有了有史以来最好的稀疏注意力技术。

No, we've just got the best sparse attention techniques ever.

Host

真的,对吧?就像压缩的信息量,可能是 PB 级,但实际上,你知道,它极其稀疏。但无论如何,想象一下,如果我们有一个人类知识工作者每年能产生六位数的价值。那么如果一个 H100 能产生接近这个价值的东西,如果我们有真正的人类在服务器上,H100 的价值就像它可以在几个月内收回成本。

Genuinely, right? Like in the sort of amount of information that is compressed, it might be petabytes, but like the actual, you know, it's like extremely sparse. But anyways, imagine if we had a human knowledge worker can produce six figures a year of value. And so if an H100 can produce something close to that, if we had actual humans on a server, the value of an H100 is like it can repay itself in the course of like a couple of months.

Mercury广告 Mercury ad

Host

在我为报税做准备的过程中,我意识到去年我与超过 50 个不同的承包商合作过,从电影摄影师到音频技术人员再到编辑,我欠他们所有人 1099 表格。过去,我只是用电子表格和一大文件夹的发票来弄清楚我需要向谁收集税务表格。但面对这么多承包商,这需要很多时间,而且我差点漏掉一些人。不过今年,Mercury 让我的流程更加直接。每当我在 2025 年付款给某人时,我只需点击一个开关,让 Mercury 向他们索取 W9 表格。正因为如此,我需要开具 1099 表格的所有信息都直接发送给了 Mercury。我 literally 只是点击了一个按钮,Mercury 就生成并发送了所有表格。这只是众多我从未想过银行平台能为我处理的事情之一。Mercury 有很多这样的功能,它们将在这个报税季总共为我节省好几天的时间。你可以在 mercury.com 了解更多。Mercury 是一家金融科技公司,不是 FDIC 保险银行。银行服务通过 Choice Financial Group 和 column NA(FDIC 成员)提供。

As I've been going through everything to prep for taxes, I realized that I worked with over 50 different contractors last year, from cinematographers to audio technicians to editors, and I owed all of them 1099s. In the past, I've just used a spreadsheet and a big folder of invoices to figure out who I need to collect tax forms from. But with so many contractors, this takes a bunch of time, and I've almost missed some people. This year though, Mercury made my process way more straightforward. Whenever I paid somebody in 2025, I just hit a toggle to have Mercury request a W9 from them. Because of that, everything that I needed to issue 1099s got sent directly to Mercury. I literally just clicked a button and Mercury generated and sent them all out. This is just one of the many things that I never would have assumed that a banking platform could just handle for me. Mercury has a bunch of features like this which are going to collectively save me multiple days this tax season. You can learn more at mercury.com. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and column NA members FDIC.

关于计算的不一致陈述 Inconsistent statements on compute

Host

所以当我采访 Dario 时,我想表达的观点并非我认为奇点两年后就会到来,因此 Dario 迫切需要购买更多算力——尽管营收确实摆在那里,他需要购买更多算力。但我想表达的是,鉴于 Dario 似乎一直在说的话,鉴于他声称我们距离一个天才数据中心还有两年时间,肯定不超过五年,而一个天才数据中心应该能赚取数万亿甚至更多的收入。这根本说不通,为什么他一直在发表关于在算力上更保守的言论,或者按你的说法,在算力购买上比 OpenAI 更不激进。我想这个观点被误解了,因为后来人们都在吐槽我,说这个播客试图说服一个价值数千亿美元公司的 CEO 去“梭哈”。但并非如此,我是想说他的内部言论是自相矛盾的。总之,能厘清这一点很好。

So when I interviewed Dario, the point I was trying to make is not that I think the singularity is two years away and therefore Dario desperately needs to buy more compute, although the revenue is certainly there that he needs to buy more compute. But the point I was trying to make is that given what Dario seems to be saying, given his statements that we're two years away from a data center of geniuses, certainly not more than 5 years away, and a data center of geniuses should be earning trillions upon trillions of dollars of revenue. It just does not make sense why he keeps making these statements about being more conservative on compute, or to your point, buying being less aggressive than OpenAI on compute. And I guess that point got lost because then people were like roasting me about like, oh this podcast was like trying to convince this multi-hundred billion company CEO like why don't you YOLO it bro? But no, I was trying to say that internally his statements are inconsistent. Anyway, so it's good to iron it out.

Dylan Patel

是的,我认为回到之前的观点:如果模型如此强大,GPU 的价值会随时间上升。随着我们越来越接近某个点——比如目前只有 Anthropic 持有这种观点——实际上每个人,即使是使用开源模型,也会开始看到每块 GPU 的价值飙升。因此从这个意义上说,你现在就应该承诺投入算力。但有趣的是,以 Anthropic 的风格,对吧?有个梗说他们存在承诺问题,有点多角恋的意思。不是指 Dario,但这确实是个梗。

Yeah, I think going back to the earlier view that if the models are so powerful, the value of a GPU goes up over time. As we approach closer and closer to, let's say, a point where right now only Anthropic has that viewpoint, as we approach further and further out, actually everyone is going to, even with open source models, be able to start to see that value skyrocket per GPU. So in that sense, you should commit now to compute. But interestingly, in Anthropic fashion, right? There's a bit of a meme that they have problems with commitment issues and they're sort of polyamorous. So not Dario, but this is a bit of a meme.

Host

这就解释了一切。

Explains everything.

AI计算的Alchian效应 Alchian effect on AI compute

Dylan Patel

顺便提一下,有一个有趣的经济学效应叫阿尔钦效应,意思是如果你提高不同商品的固定成本,其中一种质量较高,另一种质量较低,这会使人们在边际上选择质量较高的商品。举个具体例子:假设更好吃的苹果售价 2 美元,较差的苹果售价 1 美元。现在,假设你对它们征收进口关税,那么现在好苹果卖 3 美元,中等苹果卖 2 美元,对吧?

By the way, there's this interesting economics effect called Alchian, which is the idea that if you increase the fixed cost of different goods, one of which is higher quality and one is lower quality, that will make people choose the higher quality good on the margin. So to give a specific example: suppose the better tasting apple costs $2 and the shittier apple costs $1. Now, suppose you put an import tariff on them, so now it's $3 versus $2 for great apple, medium apple, right?

Host

这是因为它们都涨了 1 美元,还是应该涨 50%?

Is that because they both increase by a dollar, or should it be like a 50% increase?

Dylan Patel

不是。因为如果它们都涨了 1 美元,整体效果是:如果两者都有固定的供应成本,那么相对价格、它们之间的价差、比率都会发生变化。所以之前的情况是:较贵的那个贵一倍。现在它只贵 1.5 倍了。

No. Because if they both increase by a dollar, the whole effect is that if there's a fixed cost of supply to both, the relative price, the price difference between them, the ratio changes. So previously it was like this: the more expensive one was 2x more expensive. Now it's just 1.5x more expensive.

Host

所以我想知道,如果应用到 AI 领域,这是否意味着:如果 GPU 变得更贵,算力价格会出现固定成本上涨。结果,这将促使人们愿意为稍好一点的模型支付更高的溢价,因为算盘是这样的:反正我要花这么多钱在算力上,不如多花一点确保用的是最好的模型,而不是稍差一点的模型。

So I wonder if applied to AI that would mean that look, if GPUs are going to get more expensive, there will be a fixed cost increase in the price of compute. As a result, that will push people to be willing to pay higher margins for slightly better models because the calculus is: I'm going to be paying all this money for the compute anyways, I might as well just pay slightly more to make sure it's the very best model rather than a model that's slightly worse.

Dylan Patel

没错。所以 Hopper 从 2 美元涨到了 3 美元。如果一个 Hopper 可以生成 100 万个 Opus token,也可以生成 200 万个 Sonnet token,那么 Opus 和 Sonnet 之间的价格差就缩小了,因为 GPU 的价格从 2 美元涨到了 3 美元。有趣。

Right. So the Hopper went from $2 to $3. And if a Hopper can make a million tokens of Opus and it can make 2 million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3. Interesting.

Host

我认为这非常有道理。而且,我们看到今天所有的流量都集中在最好的模型上,所有的收入也来自最好的模型。在一个算力受限的世界里,会发生两件事,对吧?那些已经锁定了算力、没有承诺问题的公司,他们签了五年期的算力合同,锁定了巨大的利润率优势,因为他们以五年前、三年前或两年前的交易价格锁定了算力。而如果你现在正处于那个五年合同的第三年,别人的两年或三年合同到期了,现在你想以当前价格购买,当价格与模型价值挂钩时,价格会高得多。所以从某种意义上说,早期承诺的人通常拥有更好的利润率。而且市场中处于长期合同的比例远大于处于短期合同的比例,短期合同可以作为最后时刻增加的弹性容量。

I think that makes a ton of sense. Also, I think we just see all of the volumes are on the best models today. All the revenues are on the best models today. And in a compute-limited world, there are sort of two things that happen, right? Companies that have locked up, you know, and don't have commitment issues, have these 5-year contracts for compute, they've kind of locked in a humongous margin advantage because they've locked in compute for 5 years at a price of what it transacted at 5 years ago or three years ago or two years ago, whatever it is. Whereas, if you're now three years into that five-year contract and someone else's two-year contract or three-year contract rolled off and now you're trying to buy that at modern pricing, when you're priced to the value of models, the price is going to be up a lot more. So in a sense, the person who committed early has better margins in general. And the percentage of the market that is in long-term contracts is much larger than the percentage of the market in short-term contracts that can be this sort of flex capacity that you add at the last second.

Dylan Patel

与此同时,利润率流向哪里了呢?因为模型变得更值钱了。云厂商能在多大程度上灵活定价?实际上,如果你看看 CoreWeave,他们目前的平均合同期限超过三年。他们 90% 以上的算力合同都超过三年。所以他们面临一个难题:他们实际上无法灵活定价,但每年他们都在增量增加远超以往的容量。仅今年一年,他们增加的算力就相当于 2022 年用于服务 WhatsApp、Instagram 和 Facebook 以及做 AI 的所有计算机的总和。同样,Meta、CoreWeave、Google 和 Amazon 等公司每年都在增加海量算力。这些新增算力以新价格交易。所以从某种意义上说,只要处于起飞阶段,你就锁定了价格,对吧?OpenAI 去年从 600 兆瓦增加到 2 吉瓦,今年从 2 吉瓦增加到 6 吉瓦以上,明年从 6 吉瓦增加到 12 吉瓦,对吧?所有成本都来自新增算力,而不是之前的长期合同。那么,掌握定价权的是基础设施提供商,对吧?所以云厂商、Neocloud 或超大规模云服务商可以收取溢价。哦,他们不能,或者说在一定程度上可以,但当你向上游看,谁掌握了所有的内存和逻辑产能?主要是 Nvidia。他们签了很多长期合同。他们目前有大约 900 亿美元的长期合同,而且他们正在与内存供应商谈判三年期的合同。显然,Amazon 和 Google 通过 Broadcom 以及 Amazon 直接等方式,还有 AMD 这些公司,都掌握了主动权,因为他们锁定了产能。而台积电没有涨价,但内存供应商在某种程度上大幅涨价,对吧?所以他们准备再次翻倍或三倍涨价。

And at the same time, right, so where does the margin go? Because models get more valuable. How much can the cloud players flex their pricing? Well, if in fact, if you look at CoreWeave, their average term duration is over three years right now. For like 90% plus of their compute, it's over three years. And so they end up with this conundrum of like, well, they can't actually flex price, but every year they're adding incrementally way more capacity than they had previously. This year alone, the capacity fleet of computers for all purposes for serving WhatsApp and Instagram and Facebook in 2022 and doing AI, they're adding that alone this year. So in the same sense, you know, you talk about Meta doing that, CoreWeave, and Google and Amazon, all these companies are adding insane amounts of compute year on year on year. That new compute gets transacted at the new price. So in a sense, yes, you've locked in as long as we're in a sort of takeoff, right? OpenAI went from 600 megawatts to 2 gigawatts last year, and from 2 gigawatts to 6 plus this year, and 6 to 12 next year, right? The incremental added compute is where all the cost is, not the prior long-term contracts. So then who holds the card is the infra providers for charging margin, right? So now the cloud players, the Neoclouds or the hyperscalers can charge the margin. Oh, they can't because, or they can to some extent, but then as you go upstream to, well, who has access to all the memory and logic capacity? Well, it's Nvidia for the most part. They've signed a lot of long-term contracts. They've got like $90 billion of long-term contracts today, and they're negotiating three-year deals with the memory vendors today. You've got, obviously, Amazon and Google through Broadcom and their, you know, Amazon directly and all these companies, sort of AMD, these companies hold all the cards because they've secured the capacity. And TSMC is not raising prices, but memory vendors are just sort of to some extent raising a lot of price, right? So they're going to double or triple price again.

利润率与产能限制 Margins and Capacity Constraints

Host

但他们也在签这些长期合同。所以实际上,谁能攫取所有的利润呢?可能是云厂商,可能是芯片供应商,还有内存供应商。直到台积电或 ASML 跳出来说:“不,我们要大幅提价。”但与此同时,模型厂商能收取疯狂的高利润吗?我认为至少今年我们会看到模型厂商的利润率大幅上升,对吧?因为他们产能严重受限,必须抑制需求,对吧?他们不可能继续——Anthropic 不可能以当前速度继续下去而不抑制需求。

Uh but then they're also signing these long-term deals. So who is able to capture all the margin dollars is actually, you know, potentially the cloud, potentially the chip vendors, and the memory vendors. Until TSMC or ASML break out and they're like, 'No, actually we're going to charge a lot more.' But at the same time, do the model vendors get to charge crazy margins? I think at least this year we're going to see margins for the model vendors go up a lot, right? Because they're so capacity constrained, they have to destroy demand, right? There's no way they can continue—Anthropic can continue at the current pace without destroying demand.

Dylan Patel

是的。

Yeah.

英伟达在逻辑与内存上的策略 Nvidia's Strategy in Logic and Memory

Host

好的。我们来谈谈逻辑和内存。英伟达具体是如何锁定这两者的大量产能的?根据你的数据,到 2027 年,英伟达将占据 N3 晶圆产能的 70% 以上,或者差不多这个比例。然后我忘了 SK 海力士和三星等内存厂商的份额数字。但如果你看看 Neocloud 业务如何运作,英伟达如何与之合作,或者强化学习环境业务如何运作,Anthropic 如何与之合作,在这两种情况下,英伟达都刻意试图分化互补行业,以确保自己拥有尽可能多的杠杆。他们向各种 Neocloud 分配产能,确保没有一家公司拥有所有算力。类似地,Anthropic 或 OpenAI 在与数据提供商合作时,也会说“不,我们要培育一个庞大的行业,这样我们就不会被任何一家数据环境供应商锁定。”我想知道,在 3 纳米工艺上,会有 Tranium 3、TPU v7 以及其他加速器。为什么台积电要把所有产能都让给英伟达,而不是试图分化市场呢?

Yeah. Let's get into logic and memory. How specifically Nvidia has been able to lock up so much of both. So if you—I think according to your numbers by '27 Nvidia is going to have like 70 plus percent of N3 wafer capacity or something like that, or around that area. And then I forget what the numbers were for memory at SK Hynix and Samsung and so forth. But if you look at how the Neocloud business works and how Nvidia works with that, or how the RL environment business works and how Anthropic works with that, in both those cases Nvidia is purposely trying to fracture the complementary industry to make sure that they have as much leverage as possible. So they're giving allocation to random neoclouds to make sure that there's not one person that has all the compute. Similarly, Anthropic or OpenAI when they're working with the data providers, they say no, we're going to just seed a huge industry of these things so that we're not locked into any one supplier for data environments. And I wonder why on the 3 nanometer process that's going to be Tranium 3, that's going to be TPU v7, other accelerators potentially. And why is TSMC just giving it all up to Nvidia rather than, you know, trying to fracture the market?

台积电的考量与英伟达的早期承诺 TSMC's Calculus and Nvidia's Early Commitment

Dylan Patel

是的。所以我认为这里有几点。关于 3 纳米,如果我们回顾去年,3 纳米的绝大部分产能都给了苹果。苹果正在转向 2 纳米。内存价格上涨,所以苹果的产量可能会下降。随着内存价格上涨,他们要么削减利润率,要么转向。由于长期合同,存在一些时间滞后,但基本上苹果可能会减少需求或更快转向 2 纳米。目前 2 纳米仅适用于移动芯片,未来 AI 芯片也会转向那里。所以苹果有这种情况,而且苹果也在与第三方供应商洽谈,因为他们有点被台积电挤出了,因为台积电在高性能计算(HPC)AI 芯片上的利润率高于移动芯片,因为他们在 HPC 上的优势比移动更大。但无论如何,当你审视台积电的算计时,实际上他们正在向做 CPU 的公司提供非常好的产能分配。所以当你想到亚马逊有 Tranium 和 Graviton,两者都在 3 纳米上——Graviton 是他们的 CPU,Tranium 是他们的 AI 芯片——台积电更愿意给 Graviton 分配产能,而不是 Tranium,因为他们认为 CPU 业务更稳定,长期增长更可靠。作为一家保守、不想过度追逐增长周期的公司,你实际上应该先分配给更稳定、增长率较低的市场,然后再把增量产能分配给高增长市场。这通常是这样的。所以当你看到 AMD,他们在 CPU 上获得的分配——台积电对 CPU 的兴奋程度远高于 GPU。同样对于亚马逊和英伟达。英伟达有点独特,因为他们有 CPU、交换机、网络、NVLink、InfiniBand、以太网,所有这些不同的产品、网卡。总的来说,随着 Rubin 的发布以及该系列的所有芯片(GPU 是最重要的),到今年年底,这些东西大部分都将采用 3 纳米工艺。然而英伟达获得了大部分供应。部分原因是台积电和其他公司通过多种方式预测市场需求。但这也是市场信号。市场发出了信号:我们明年需要这么多产能。我们需要这么多。我们会签不可取消、不可退货的合同。我们甚至可能支付定金。英伟达比谷歌或亚马逊早得多地做到了这一点。在某些情况下,谷歌和亚马逊遇到了障碍。其中一款芯片延迟了几个季度——Tranium 等等。所以在这种情况下,出现了巨大的“好吧,这些家伙在延迟,但英伟达想要更多、更多、更多”的情况。我们正在检查供应链的其他部分:产能足够吗?所以他们去找所有 PCB 供应商,说:“嘿,PCB 产能够吗?”Vitec Giant 是英伟达最大的 PCB 供应商之一,他们是一家中国公司。所有 PCB 都来自中国,来自他们,或者说很多都来自他们。无论如何,他们问:“你们有足够的 PCB 产能吗?太好了。哦,嘿,内存供应商,谁有所有的内存产能?好的,英伟达有。太好了。”所以当你以同样的方式看待时,谁足够“AGI 上脑”以至于愿意在长期时间线上以对非 AGI 上脑者来说荒谬的水平购买算力,但他们仍然愿意支付相当高的利润率并现在签约,因为他们认为未来这个比例会搞砸。半导体供应链也发生了同样的事情。英伟达——虽然我不认为英伟达完全“AGI 上脑”,黄仁勋并不相信软件会被完全自动化等等。

Yeah. So I think there are a couple of points here. On 3 nanometer, if we go back to last year, the vast majority of 3 nanometer was Apple. Apple is moving to 2 nanometer. Memory prices are going up, so Apple's volumes may go down. As memory prices go up, they have to either cut margin or move on. There's some time lag because they have long-term contracts, but basically Apple likely reduces demand or moves to 2 nanometer faster, where 2 nanometer is only capable for mobile chips today, and in the future AI chips will move there. So Apple has that, and Apple is also talking to third-party vendors because they're getting squeezed out of TSMC a little bit, because TSMC's margins on high-performance computing (HPC) AI chips are higher than for mobile, because they have a bigger advantage in HPC than in mobile. But anyway, when you look at TSMC's calculus, actually they're providing really good allocations to companies that are doing CPUs. So when you think about Amazon has Tranium and Amazon has Graviton, both on 3 nanometer—Graviton being their CPU, Tranium being their AI chip—TSMC is much more excited to give allocation to Graviton than to Tranium, because they view the CPU business as more stable long-term growth. As a company that is conservative and doesn't want to ride cycles of growth too hard, you actually want to allocate to the market that is more stable and lower growth rate first, before you allocate all the incremental capacity to the fast growth rate market. That is generally the case. So when you look at AMD, the allocations they get on their CPUs—TSMC is much more excited about those than for GPUs. Likewise for Amazon and Nvidia. Nvidia is a bit unique because they have CPUs, switches, networking, NVLink, InfiniBand, Ethernet, all these different products, NICs. By and large, most of these things will be on 3 nanometer by the end of this year with the Rubin launch and all the chips in that family, the GPU being the most important one. And yet Nvidia is getting the majority of supply. Part of this is because TSMC and others forecast market demand in many ways. But also it's the market signal. The market signaled: we need this much capacity next year. We need this much. We'll sign non-cancellable, non-returnable. We may even pay deposits. Nvidia just did it way earlier than Google or Amazon. And in some cases, Google and Amazon had stumbling blocks. One of the chips got delayed slightly by a couple of quarters—Tranium and all these sorts of things happened. So in that case, there was a huge sort of 'okay, these guys are delaying, but Nvidia is wanting more, more, more.' And we are checking with the rest of the supply chain: is there enough capacity? So they go to all the PCB vendors and say, 'Hey, is there enough PCB capacity?' Vitec Giant is one of the largest suppliers of PCBs to Nvidia, and they're a Chinese company. All the PCBs come from China, from them, or many of them. And anyway, they're like, 'Do you have enough PCB capacity? Great. Oh hey, memory vendors, who has all the memory capacity? Okay, Nvidia does. Great.' So when you look at it in the same way, who is AGI-pilled enough to buy compute in long timelines at levels that seem ridiculous to people who aren't AGI-pilled, but nonetheless, they're willing to pay a pretty good margin and sign it now because they view in the future that ratio is screwed up. The same thing happens with the supply chain for semiconductors. Nvidia was—while I don't think Nvidia is quite AGI-pilled, Jensen doesn't believe software is going to be automated fully and all these things.

Host

加速计算,不是 AI 芯片,对吧?

Accelerated computing, not AI chips, right?

Dylan Patel

是 AI 芯片,对吧?

It's AI chips, right?

Host

但他是这么叫的,对吧?

But that's what he calls it, right?

Dylan Patel

是的。因为我认为有一个更广泛的术语,对吧?AI 包含在其中,但还有物理建模、模拟等等——

Yeah. Because I think there's a broader term, right? AI is within that, but like physics modeling and simulations and like—

Host

或者说,他只是没有拥抱那种主要用例。

Or but just like he's not embracing the sort of main use case.

Dylan Patel

我认为他是在拥抱它,但我不认为他像 Dario 那样“AGI 上脑”,对吧?但他仍然更“AGI 上脑”,原因很简单。你可以看到所有的数据中心建设。

I think he's embracing it, but I just don't think he's AGI-pilled like Dario, right? But he's still way more AGI-pilled, and the reason is pretty simple. You can see all the data center construction.

谷歌TPU与GPU的困境 Google's TPU vs GPU dilemma

Dylan Patel

他就像说:“好吧,我要拿下这个市场份额。”我们基本上追踪了所有数据中心,你可以看到很多数据中心可以说,嗯,它们可能是这样或那样。所以在某种程度上,谷歌和亚马逊,尤其是谷歌,尽管部署自己的 TPU 对他们更有利,但他们还是不得不部署大量 GPU,因为他们没有足够的 TPU 来填满数据中心。他们无法生产出足够的 TPU。

He's like, 'Okay, I want to have this market share.' We sort of have all the data centers tracked, and you can see there's a lot of data centers that you could say, well, they could be one or the other. So to some extent, Google and Amazon, Google especially, even though their TPU is just better for them to deploy, they have to deploy a crapload of GPUs because they don't have enough TPUs to fill up their data centers. They can't get them fabbed.

Host

等等,我能问个问题吗?谷歌卖给了 Anthropic,我记得是一百万块,是 V7 还是 Ironwood?你意思是说,总的来说,现在或明年存在一个大瓶颈?我是说,我觉得从现在开始永远都会是逻辑内存,这些制造芯片所需的东西。而谷歌有 DeepMind,这是另一个第三大知名 AI 实验室。如果这是个大瓶颈,他们为什么不直接给 DeepMind,反而要卖掉呢?

Wait, can I ask a question about that? Google sold, I think a million, was it the V7s, the Ironwoods to Anthropic. And you're saying in general there's this big bottleneck right now, this year or next year? I mean, I guess going forward forever now is going to be the logic memory, the stuff that it takes to build these chips. And Google has DeepMind, this is the other third prominent AI lab. If this is the big bottleneck, why would they sell it rather than just giving it to DeepMind?

Dylan Patel

对,这又是一个问题,你知道,DeepMind 的人会说:“这太疯狂了,我们为什么要这么做?”但谷歌云的人和谷歌高管有不同的想法。基本上,你和我都知道计算团队。有一个人,其实他们两个都来自谷歌,是 Anthropic 计算团队的主要成员。他们看到了这个错位。他们谈判了一笔交易,在谷歌意识到之前就拿到了这些算力。所以事件经过,至少从我们找到的数据来看,是在 Q3 初,我们看到在大概六周的时间里,Anthropic 的,抱歉,是 TPU 的容量大幅增加,在那六周里增加了好几倍。有多次请求。谷歌甚至不得不去找台积电,解释他们为什么需要增加产能,因为太突然了。但那些产能增加大部分是为了卖给 Anthropic。

Right, so this is again like a problem of, you know, DeepMind people were like, 'This is insane, why do we do this?' But then Google Cloud people and Google executives saw a different thought process. Basically, you and I for the compute team. There's one guy from, you know, both of them actually came from Google, the main people on the compute team at Anthropic. They saw this dislocation. They negotiated a deal and they were able to get access to this compute before Google realized. So the chain of events, at least from our data that we found, was in early Q3, we saw over the course of two, over the course of like six weeks, we saw capacity on Anthropic, or sorry, on TPUs go up by a significant amount over the course of those six weeks, and it went up like multiple times in those six weeks. There were multiple requests. Google even had to go to TSMC and explain to them why they needed this increase in capacity because it was so sudden. But a lot of that capacity increase was for selling to Anthropic.

Host

嗯。

Yeah.

Dylan Patel

因为 Anthropic 比谷歌更早看到了这一点。然后谷歌推出了 Nano Banano 和 Gemini 3,导致用户指标飙升,谷歌领导层说:“哦”,然后他们开始声明必须每 6 个月(或者我不记得他们说的确切数字)翻倍算力。但他们真的醒悟了很多,然后他们说:“哦,嘿,台积电,我们想要更多。我们想要更多。”台积电说:“嗯,抱歉,各位。我们明年的产能已经卖光了。我们可以努力一下明年,也许 2026 年能多拿 5-10%,但真正要等到 2027 年。”在我看来,实验室之间存在信息不对称。我不知道是不是完全这样,但这是我从供应链数据、晶圆订单、Anthropic 和 Fluid Stack 签署的数据中心情况中自己得出的叙事。很明显谷歌搞砸了,你可以从谷歌 Gemini 的年度经常性收入(ARR)看出来。Q1 几乎为零,Q3 开始推理后有一点,但 Q4 退出时 ARR 大约 50 亿美元。所以 Q4 的营收按 ARR 算大约是 50 亿美元。很明显谷歌没有看到收入飙升。从某种意义上说,Anthropic 在 AR 爆发之前有点承诺问题,尽管他们有更多的信息不对称,能看到未来的趋势。谷歌会比 Anthropic 更保守。第一,第二,谷歌的 ARR 更少。所以他们似乎就是不愿意做,然后意识到应该做。所以从那以后,谷歌变得极其激进。他们收购了一家能源公司。他们为涡轮机支付定金。他们购买了极高比例的有电力供应的土地。他们去找公用事业公司谈判长期协议。他们在数据中心和电力方面非常激进。所以,我认为谷歌在去年年底醒悟了,但花了一些时间。

Because Anthropic saw it before Google. And then Google had Nano Banano and Gemini 3 which caused their user metrics to skyrocket, and leadership at Google was like, 'Oh,' and then they started making the statement of we have to double compute every, is it 6 months or I don't remember the exact number that they said. But they really woke up a lot more, and then they're like, 'Oh, hey TSMC, we want more. We want more.' And it's like, 'Well, sorry guys. Like, we're sold out for next year. We can work on next year. We can maybe get like 5-10% more for '26, but really we're going to work on '27.' There's sort of like this information asymmetry of the labs in my mind. And I don't know if this is exactly it, it's the narrative I've spun myself from seeing all the data in the supply chain and like wafer orders and like what's going on with the data centers that Anthropic signed and Fluid Stack signed and all this. It's pretty clear to me that Google screwed up, and you can see this from Google's Gemini ARRs. They had next to nothing in Q1, Q3 a little bit once they started inferencing, but Q4 they were at like $5 billion ARR exiting or something like this. So it's like $5 billion revenue for Q4 on an ARR basis. So it's clearly like Google didn't see revenue skyrocket. And in a sense, Anthropic was not willing, you know, kind of had like a little bit of commitment issues before their AR exploded, even though they have far more information asymmetry and see what's coming down the pipe. Google is going to be more conservative than Anthropic is. A, and B, Google had even less ARR. So they sort of were like, I think just not willing to sort of do it, and then they realized they should do it. And so now since then, Google has gotten absurdly aggressive in terms of what they're doing. They bought an energy company. They're putting deposits down for turbines. They're buying a ridiculous percentage of the powered land. They're going to utilities and negotiating long-term agreements. They're doing this on the data center and power side very aggressively. So, I think Google woke up towards the end of last year, but it took them some time.

Host

那你认为到明年年底谷歌会有多少吉瓦?

And how many gigawatts do you think Google will have by the end of next year?

Dylan Patel

买我的数据。

Buy my data.

Host

这种信息你是收费的。

You charge for that kind of information.

Dylan Patel

是的。

Yes.

Host

我觉得每年阻碍我们扩展 AI 算力的瓶颈都在变。几年前是成本。去年是电力。今年,你会告诉我今年的瓶颈是什么。但我想了解五年后,什么会限制我们部署奇点。

I feel like every year the bottleneck for what is preventing us from scaling AI compute keeps changing. A couple years ago was cost. Last year it was power. This year, you'll tell me what the bottleneck is this year. But I want to understand 5 years out what will be the thing that is constraining us from deploying the singularity.

Dylan Patel

是的,我认为最大的瓶颈是算力,而其中交货周期最长的供应链不是电力或数据中心,实际上是半导体供应链本身。瓶颈从电力和数据中心又转回了芯片。在芯片供应链中,有很多不同的瓶颈。有内存,有台积电的逻辑晶圆,有晶圆厂本身。建造晶圆厂需要几年时间,两到三年,而数据中心不到一年。我们看到亚马逊在 8 个月内就建成了数据中心。所以交货周期差异很大,因为建造实际制造芯片的晶圆厂很复杂。还有工具,它们的交货周期也很长。

Yeah, I think the biggest bottleneck is compute and for that the longest lead time supply chains are not power or data centers. They're actually the semiconductor supply chain themselves. It switches back from being power and data center as a major bottleneck to chips. And in the chip supply chain, there's a number of different bottlenecks. There's memory, there's logic wafers from TSMC, there's fabs themselves. Construction of the fabs takes a couple years, three to three years, versus a data center takes less than a year. We've seen Amazon build data centers in as fast as 8 months. So there's a big difference in lead times because of the complexity of building the fab that actually makes the chips. And then the tools, those also have really long lead times.

AI计算扩展的瓶颈 Bottlenecks in AI compute scaling

Dylan Patel

随着规模扩张,瓶颈已经从“供应链目前无法做到什么”转移到了 CoWoS、电力和数据中心,但这些都是较短前置时间的项目,对吧?CoWoS 是封装芯片的简单得多的过程。电力和数据中心最终比芯片的实际制造简单得多。因此,移动或 PC 到数据中心芯片的产能有一些转移,但这在一定程度上是可替代的。而在 CoWoS、电力和数据中心方面,这些基本上必须作为供应链重新开始。但现在,移动和 PC 行业(曾经是半导体行业的大部分)已经没有更多产能可以转移到 AI 了,对吧?英伟达现在是台积电的最大客户,英伟达也是最大内存制造商 SK 海力士的最大客户,对吧?所以这种规模扩张或资源从普通人那里转移,PC 和智能手机再向 AI 芯片转移几乎是不可能的。那么现在,我们如何扩大 AI 芯片生产?这就是到 2030 年最大的瓶颈。

And so the bottlenecks as we've scaled have shifted from 'Hey, what is the supply chain currently not able to do?' which was CoWoS and power and data centers, but those were all shorter lead-time items, right? CoWoS is a much more simple process of packaging chips together. Power and data centers are ultimately way more simple than the actual manufacturing of the chips. And so there's been some sliding of capacity across mobile or PC to data center chips, but that's been somewhat fungible. Whereas on CoWoS and power and data centers, those have sort of had to start anew as supply chains. But now there's sort of no more capacity for the mobile and PC industries, which used to be the majority of the semiconductor industry, to shift over to AI, right? Nvidia is now the largest customer at TSMC, and Nvidia is the largest customer at SK Hynix, the largest memory manufacturer, right? So it's sort of impossible for this scaling or the sliding of resources away from the common person, right? PCs and smartphones to shift anymore towards the AI chips. And so now, how do we scale the AI chip production? And that's the biggest bottleneck as we go to 2030 is those.

Host

如果存在一个绝对的吉瓦上限,你可以仅根据“我们无法生产超过这么多 EUV 机器”来预测到 2030 年,那会非常有趣,对吧?

It'd be very interesting if there's an absolute gigawatt ceiling that you can project out to 2030 based just on, hey, we can't produce more than this many EUV machines, right?

Dylan Patel

所以要进一步扩展算力,今年、明年会有一些不同的瓶颈,但最终到 2028-2029 年,瓶颈会落到供应链的最底层,也就是 ASML,对吧?ASML 制造世界上最复杂的机器,即 EUV 光刻机。这些机器的售价是 3-4 亿美元。目前他们每年能生产大约 70 台。明年会达到 80 台。即使在非常激进的供应链扩张下,到本十年末也只能达到略超过 100 台。这意味着什么?好吧,到本十年末他们能生产 100 台这样的机器,现在有 70 台。这实际上如何转化为 AI 算力呢?我们看到 Sam Altman 和供应链上许多其他人提出的数字:吉瓦、吉瓦、吉瓦,对吧?我们增加了多少吉瓦?我们还看到 Elon 说“每年在太空中 100 吉瓦”。这些数字的问题或挑战实际上不在于电力,也不在于数据中心。我们可以深入探讨,但问题在于芯片制造,对吧?所以,一吉瓦的英伟达 Rubin 芯片,对吧?Rubin 是在 GTC 上发布的,我相信就在本期播客上线的那一周。要制造一吉瓦数据中心容量的英伟达最新芯片(他们将在今年年底发布),你需要几种不同的晶圆技术,对吧?你需要大约 55,000 片 3 纳米晶圆。你需要大约 6,000 片 5 纳米晶圆,然后你需要大约 170,000 片 DRAM 晶圆,对吧?内存。所以在这三个不同的类别中,每个都需要不同数量的 EUV 光刻,对吧?当你制造晶圆时,有成千上万的工艺步骤,涉及沉积材料、去除材料,但关键步骤(至少在先进逻辑中占芯片成本的 30%)实际上并不在晶圆上添加任何东西,对吧?你取晶圆,涂上光刻胶(一种化学物质,暴露在光线下会发生化学变化),然后把它放进 EUV 光刻机,机器以特定方式照射光线。它进行图案化,对吧?因为有一个叫做掩模的东西,实际上是设计的模板。所以当你观察晶圆时,领先的 3 纳米晶圆有大约 70 个掩模,对吧?大约 70 层光刻,但其中 20 层是最先进的 EUV,对吧?具体来说,如果你考虑一下,好吧,如果一吉瓦需要 55,000 片晶圆,如果每片晶圆进行 20 次 EUV 曝光,那么你可以计算一下,那就是一吉瓦需要 110 万次 EUV 曝光。所以实际上很简单。然后加上其余部分,最终达到 200 万次,对吧?加上 5 纳米和所有内存,一吉瓦大约需要 200 万次 EUV 曝光。你知道,这些机器非常复杂。所以当你考虑它在晶圆上的操作时,它扫描并步进,对吧?它扫描、步进,在整个晶圆上重复数十次或数百次。所以当你问“需要多少次 EUV 曝光?”时,那就是整个晶圆以一定速率被曝光。一台 EUV 光刻机每小时大约可以处理 75 片晶圆。机器大约 90% 的时间在运行,对吧?所以最终,一吉瓦需要大约 3.5 台 EUV 光刻机来完成 200 万次 EUV 晶圆曝光。所以 3.5 台 EUV 光刻机满足一吉瓦。所以想想这些数字很有趣,对吧?因为我们说一吉瓦成本是多少?大约 500 亿美元,对吧?而 3.5 台 EUV 光刻机成本是多少?大约是 12 亿美元,对吧?实际上是一个低得多的数字,这很有趣:50 吉瓦的经济性,数据中心资本支出,以及在此基础上构建的代币价值甚至更大,对吧?可能价值 1000 亿美元的 AI 价值进入供应链,却被这 12 亿美元的工具所阻碍,而这些工具根本无法快速扩张其供应链。

So to scale compute further, right, there's some different bottlenecks this year, next year, but ultimately by 2028-2029, the bottleneck falls to the lowest rung on the supply chain, which is ASML, right? ASML makes the world's most complicated machine, i.e., an EUV tool. And the selling price for those is $300-400 million. And currently they can make about 70. Next year they'll get to 80. Even under very aggressive supply chain expansion, they only get to a little bit over a hundred by the end of the decade. And so what does that mean? Okay, they can make a hundred of these tools by the end of the decade and, you know, 70 right now. How does that actually translate to AI compute, right? We see all these numbers from Sam Altman and many others across the supply chain: gigawatts, gigawatts, gigawatts, right? How many gigawatts are we adding? And we see, you know, Elon saying, 'Hey, the 100 gigawatts in space a year.' The problem with any of these numbers or the challenge to these numbers is, you know, actually not the power, not the data center. We can dive into that, but it's manufacturing the chips, right? So, a gigawatt of, you know, Nvidia's Rubin chips, right? So, Rubin is announced at GTC, I believe the week this podcast goes live. And to make a gigawatt worth of data center capacity of Nvidia's latest chip that they're releasing at the end of this year, towards the end of this year, you need, you know, a few different wafer technologies, right? You need about 55,000 wafers of 3 nanometer. You need about 6,000 wafers of 5 nanometer, and then you need about 170,000 wafers of DRAM, right? Memory. And so across these three different buckets, each of these requires different amounts of EUV, right? So when you manufacture a wafer, there's thousands and thousands of process steps where you're depositing material, removing them, but the sort of key critical step, which at least in advanced logic is like 30% of the cost of the chip, is something that doesn't actually put anything on the wafer, right? You take the wafer, you deposit photoresist, which is like a chemical that basically chemically changes when you expose it to light, and then you stick it into the EUV tool, which shines light at it in a certain way. It patterns it, right? Because there is what's called a mask, which is a stencil effectively for the design. And so when you look at a wafer, you know, leading edge 3 nanometer wafer has 70 or so masks, right? 70 or so layers of lithography, but 20 of them are the most advanced EUV, right? And that specifically, you know, if you think about, okay, well, if I need 55,000 wafers for a gigawatt, if I do 20 EUV passes per wafer, you then you can do the math, that's like, okay, that's 1.1 million passes of EUV for a single gigawatt. So, actually like it's pretty simple. And then once you add the rest of the stuff, it ends up being 2 million, right? Across 5 nanometer and all the memory, you're at roughly 2 million EUV passes for a single gigawatt. You know, these tools are very complicated. So, when you think about what it's doing across a wafer, it's taking the wafer and it's scanning and it's stepping across, right? It's scanning, stepping across, and it does this hundreds of times across the entire or dozens of times across the whole wafer. And so when you're talking about, hey, how many EUV passes? That's the entire wafer being exposed at a certain rate. A wafer EUV tool can do roughly 75 wafers per hour. And the tool is up roughly 90% of the time, right? So in the end you end up with actually I need about three and a half EUV tools to do the 2 million EUV wafer passes for the gigawatt. So three and a half EUV tools satisfies a gigawatt. So it's funny to think about the numbers, right? Because we're talking about oh what's a gigawatt cost? It costs like $50 billion roughly, right? Whereas what does three and a half EUV tools cost? That's like $1.2 billion, right? It's actually like quite a lower number, which is interesting to think about like oh 50 gigawatts of economic, you know, sort of capex in the data center and what gets built on top of that in terms of tokens is even larger, right? It might be $100 billion worth of AI value into the supply chain is held up by this $1.2 billion worth of tooling that simply just cannot expand its supply chain quickly.

Host

我想你最近读到一篇文章,你说过去三年台积电的资本支出达到了 1000 亿美元。所以大概是 300 亿、300 亿、400 亿,如果你想想,其中一小部分被英伟达用于 3 纳米或之前用于其芯片的 4 纳米。但英伟达将其转化为——你上一季度的收益是 400 亿美元,乘以四就是 1600 亿美元。所以仅英伟达一家就将 1000 亿美元资本支出中的一小部分(这些支出将在多年内折旧,而不仅仅是这一年)转化为一年 1600 亿美元。然后当你沿着供应链向下到 ASML 时,情况更加极端,ASML 用价值 10 亿美元的机器生产一吉瓦。当然,这些机器可以使用超过一年,对吧?所以它产生的价值远不止这些。

And I think so you read this article recently where you're saying over the last three years TSMC has done $100 billion of capex. So it's like 30, 30, 40, and if you think of I mean a small fraction of that is sort of like being used by Nvidia for the 3 nanometer that it's going to or you know previously 4 nanometer that it's using for its chips. But Nvidia has turned that into what was what are it's like your earnings last quarter was like $40 billion into $40 billion times four. So $160 billion. So Nvidia alone is turning some small fraction of $100 billion in capex that's going to be depreciated over many years not just this one year into $160 billion in a single year. And then that gets even more intense when you go down the supply chain to ASML which is taking $1 billion worth of machines to produce a gigawatt. And of course those machines last for more than a year, right? So it's doing more than that.

EUV工具数量与Sam Altman的千兆瓦目标 EUV tool count and Sam Altman's gigawatt target

Host

所以现在我想了解:到 2030 年,如果不仅算当年售出的机器,还包括前几年累积的,总共会有多少台这样的机器?这对 Sam Altman 提出的 2030 年每周一个吉瓦的目标意味着什么?把这些数字加起来,是否与那个目标兼容?

So now I want to understand: okay, well, how many such machines will there be by 2030, if you include not just the ones that are sold that year but have been compiling over the previous years? And what does that imply about Sam Altman's goal of a gigawatt a week in 2030? When you add up those numbers, is that compatible with that?

Dylan Patel

对,这完全兼容,对吧?因为如果你想想台积电和整个生态系统,他们已经有大约 250 到 300 台 EUV 光刻机了。然后今年加上 70 台,明年 80 台,到 2030 年增长到 100 台。到本十年末,你会有大约 700 台 EUV 光刻机。700 台 EUV 光刻机,每吉瓦需要 3.5 台机器,假设全部用于 AI——虽然并非如此——但每吉瓦 3.5 台机器,就能得到 200 吉瓦的 AI 芯片供数据中心部署。所以是 200 吉瓦。Sam 想要 50 吉瓦,对吧?每年 52 吉瓦。那他只占了 25% 的份额。显然,有一部分会分配给移动设备和 PC,假设出于某种原因我们还能有消费品,并且不会被挤出市场,但大致上,他说的是芯片制造总量的 25% 市场份额。这相当合理,因为仅今年一年,我认为他就能获得已部署的 Blackwell GPU 的 25%。所以这并不疯狂。

Right. That's completely compatible, right? Because if you think about TSMC and the entire ecosystem, they have something like 250 to 300 EUV tools already. Then you stack on 70 this year, 80 next year, growing to 100 by 2030. You're at like 700 EUV tools by the end of the decade. 700 EUV tools, three and a half tools per gigawatt, assuming it's all allocated to AI — which it's not — but three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy. So 200 gigawatts. Sam wants 50 gigawatts, right? 52 gigawatts a year. He's only taking 25% share then. Obviously there's some share given to mobile and PC, assuming that for some reason we're allowed to still have consumer goods and we don't get priced out of them, but roughly, he's saying 25% market share of the total chips fab. That's kind of very reasonable, given that this year alone I think he's going to have access to 25% of the Blackwell GPUs that are deployed. So it's not that crazy.

使用十年老EUV工具的意外 Surprise about using decade-old EUV tools

Host

我觉得这很令人惊讶——第一台 EUV 光刻机是什么时候出货的?当 7 纳米开始时,所以我不确定具体时间——但你说在 2030 年,他们将使用最初于 2020 年出货的机器。所以 10 年了,你还在使用这个世界上最先进技术行业中最重要的机器。我觉得这很令人惊讶。

I find it surprising that — when was the first EUV tool shipped? When 7 nanometer started, so I don't know when that was exactly — but you're saying in 2030 they're going to be using machines that initially were shipped in 2020. So 10 years, you're using the most important machine in this most technologically advanced industry in the world. I find that surprising.

Dylan Patel

所以,ASML 出货 EUV 光刻机已经有大约十年了,但直到 2020 年左右才进入大规模量产。这些机器并不相同。当时,机器的吞吐量更低。它们有各种规格,比如套刻精度(overlay),对吧?我之前提到过,你要一层层堆叠。你会做一些 EUV 步骤,然后进行一系列不同的工艺步骤:沉积、蚀刻、清洗晶圆——在下一个 EUV 层之前有几十个这样的步骤。有一个规格叫套刻精度,意思是:你完成了所有这些工作,在晶圆上画出了这些线条。现在我要画这些点来连接这些金属线,然后打孔,再下一层是另一组垂直的线条。所以你要连接互相垂直的导线。你必须让它们精确对准。这就是套刻精度,ASML 在这方面改进得非常快。晶圆吞吐量也提升得很快,同时机器的价格也上涨了,但涨幅不及性能提升。最初 EUV 光刻机大约 1.5 亿美元,随着时间的推移,到 2028 年我预计是 4 亿美元左右。但机器的性能也翻了一倍多,尤其是在吞吐量和套刻精度上——即使中间有大量步骤,也能精确对准后续的层。所以 ASML 进步非常快。还有一点值得一提:ASML 可能是世界上最慷慨的公司之一。他们拥有这个关键部件。没有人有竞争力。也许中国到本十年末会有一些 EUV 光刻机,但其他任何人都没有接近 EUV 的技术。然而,他们并没有疯狂地提高价格和利润率。你去问问我们经常聊的其他一些人,比如 Leopold,他们会说:“为什么不涨价?”因为他们可以。利润率就在那里。你可以像英伟达那样赚取利润。存储厂商也在赚取利润,但 ASML 从未将价格提高到超过工具性能提升的程度。所以从某种意义上说,他们总是为客户提供净收益。并不是说工具停滞不前;只是这些工具老了。是的,你可以对它们进行一些升级,而且新工具也在推出。为了简化,我们在本期播客中忽略了每台工具在套刻精度或吞吐量上的进步。

So, ASML has been shipping EUV tools for roughly a decade, but it only entered mass volume production around 2020. The tools are not the same. Back then, the tools had even lower throughput. There are various specifications around them called overlay, right? I was mentioning you're stacking layers on top of each other. You'll do some EUV, you'll do a bunch of different process steps: depositing stuff, etching stuff, cleaning the wafer — dozens of those steps before you do another EUV layer. There's a spec called overlay, which is: okay, you did all this work, you drew these lines on the wafer. Now I want to draw these dots to connect these lines of metal, and then do holes, and then the next layer up is another set of lines going perpendicular. So now you're connecting wires going perpendicular to each other. You have to be able to land them on top of each other. So it's called overlay, and overlay is a spec that's been improved rapidly by ASML. Wafer throughput has been improved rapidly by ASML, and also the price of the tool has gone up, but not as much as the capabilities of the tool. Initially, the EUV tools were like $150 million, and over time they're now like $400 million as I look out to 2028. But the capabilities of the tools have more than doubled as well, especially on throughput and overlay accuracy, which is the ability to accurately align the subsequent passes on top of each other even though you do tons of steps between. And so ASML is improving super rapidly. I think it's also something noteworthy to say: ASML is maybe one of the most generous companies in the world. They have this lynchpin thing. No one has anything competitive. Maybe China will have some EUV by the end of the decade, but no one else has anything even close to EUV. And yet they haven't taken price and margins up like crazy. You go ask some other folks we talk to all the time, like Leopold, and they're like, "Why don't you let the price go up?" Because they can. The margin is there. You can take the margin like Nvidia takes the margin. Memory players are taking the margin, but ASML has never raised the price more than they've increased the capability of the tool. So in a sense, they've always provided net benefit to their customer. It's not that the tool is stagnant; it's just that these tools are old. Yes, you can upgrade them some, and new tools are coming. And for simplicity's sake, we're kind of ignoring the advances in overlay or throughput per tool for this podcast.

ASML产能限制 Constraints on ASML's production capacity

Host

所以你说我们今年生产 60 台这样的机器,然后接下来几年是 70 台、80 台。如果 ASML 决定将其资本支出翻倍或三倍,会发生什么?是什么阻止他们在 2030 年生产超过 100 台?为什么你如此确信,即使五年后,你也能相对确定他们的产量?

So you say we're producing 60 of these machines this year, and then 70, 80 over subsequent years. What would happen if ASML just decided to double its capex or triple its capex? What is preventing them from producing more than 100 in 2030? Why are you so confident that even 5 years out you can be relatively sure what their production will be?

Dylan Patel

所以我认为有几个因素。ASML 并没有决定要“豁出去”尽可能快地扩张产能。总的来说,半导体供应链没有这样做;它经历过繁荣和萧条。我们可以多聊一点,但基本上没有人——直到最近一些参与者才醒悟过来——但总体上没有人真正看到每年 200 吉瓦的 AI 芯片需求或每年数万亿美元的半导体供应链支出。他们就是没有被 AI 洗脑,对吧?他们不是 AI 信徒。

So I think a couple factors here. ASML has not decided to just go YOLO and expand capacity as fast as possible. In general, the semiconductor supply chain has not; it's lived through booms and busts. We can talk a bit more about it, but basically no one — some players as of very recently have woken up, but in general no one really sees demand for 200 gigawatts a year of AI chips or trillions of dollars of spend a year in the semiconductor supply chain. They're just not AI-pilled, right? They're not AI.

Host

我们今年就要达到一万亿美元了。

We're going to get to a trillion dollars this year.

Dylan Patel

是的,我理解,但我只是说供应链中没有人真正理解这一点。我们不断被告知我们的数字太高了,然后当它们正确时,他们又说:“哦,是的,是的,但你明年的数字还是太高了。”但无论如何。ASML 的工具有四个主要部件:光源,由圣地亚哥的 Cymer 制造;掩模版台,在康涅狄格州威尔明顿制造;晶圆台;以及光学系统,即透镜等。后两者在欧洲制造。当你审视这四个部件中的每一个时,它们都是极其复杂的供应链,a) 他们没有尝试大规模扩张,b) 当他们尝试扩张时,时间延迟相当长。所以再说一次,这是人类制造的最复杂的机器,没有之一,而且是批量生产。但让我们专门谈谈光源。光源是做什么的?它滴下这些锡滴,然后用激光完美地连续击打三次。

Yeah, I feel you, but I'm just saying no one really understands this in the supply chain. Constantly we're told our numbers are way too high, and then when they're right, they're like, "Oh, yeah, yeah, but your next year's numbers are still too high." And it's like, but anyways. ASML's tool has four major components: the source, made by Cymer in San Diego; the reticle stage, made in Wilmington, Connecticut; the wafer stage; and the optics, the lenses and such. Those two are made in Europe. When you look at each of these four, they're tremendously complex supply chains that a) they have not tried to expand massively, and b) when they try to expand them, the time lag is quite long. So again, this is the most complicated machine that humans make, period, at any sort of volume. But let's talk about the source specifically. What does the source do? It drops these tin droplets and hits it three subsequent times with a laser perfectly.

EUV光刻工艺 EUV Lithography Process

Dylan Patel

第一个脉冲击中锡滴,使其膨胀。它再次击中,所以它膨胀成完美的形状,然后以超高功率轰击它,锡滴被激发到足以释放 13.5 nm 的 EUV 光。然后它进入这个装置,基本上收集所有光线并将其引导到透镜组中。对。然后你有透镜组,这是卡尔蔡司的,正如你提到的。还有其他一些公司,但蔡司是最重要的部分。

So the first one hits this tin droplet, expands out. It hits it again. So it expands out to this perfect shape and then it blasts it at super high power and the tin droplets get excited enough that they release EUV light at 13.5 nm. Then it's in this thing that is like basically collecting all the light and directing it into the lens stack. Right. Then you have the lens stack which is Carl Zeiss, right, as you mentioned. And some other folks, but Zeiss being the most important part of it.

Dylan Patel

他们也没有尝试扩大产能,因为他们看不到任何需求,你知道,他们就像,“哦,是的,是的,因为 AI 我们增长了很多。我们从 60 增长到 100,对吧?”实际上,“不,不,不。我们需要达到几百台,但没关系,随便吧。”

They also have not tried to expand production capacity because they don't see any, you know, they're like, "Oh, yeah, yeah, like we're growing a lot because of AI. We're growing from 60 to 100, right?" It's like, "No, no, no. We need to go to like a couple hundred, but it's fine, whatever."

Dylan Patel

每台这样的设备,我认为有 18 个这样的透镜,实际上是反射镜。它们是多层反射镜,由钼和钌(如果我没记错的话)完美层叠而成。然后光线完美地反射,但它不像我们通常想的透镜那样有形状并聚焦光线。这是一种既是反射镜又是透镜的东西,所以非常复杂。这些超薄沉积层中的任何缺陷都会破坏它。任何曲率问题,扩大生产面临很多挑战。从这个意义上说,它相当手工化,因为你每年不是生产数万个,而是数百个,数千个。你知道,每年 60 台设备,每台 18 个,最终你仍然在数百或数千的量级,这些透镜和投影光学器件大约在千这个数字。

Each of these tools has, I think, 18 of these lenses, effectively mirrors. They are multi-layer mirrors which are perfect layers of molybdenum and ruthenium, if I recall correctly, stacked on top of each other in many layers. Then the light bounces off of it perfectly, but it's not just like when we think about a lens, it's like in a shape and it focuses the light. This is like a mirror that's also a lens, so it's pretty complicated. Any defect in this perfect layer of stack in these super thinly deposited stacks will mess it up. Any curvature issues, there is a lot of challenges with scaling the production. It's quite artisanal in this sense, right, because you're not making tens of thousands of these a year, you're making hundreds, you're making thousands. You know, talk about 60 tools a year, 18 of these per tool, you end up with you're still in the hundreds of tools or thousand, you're at the thousand number roughly for these lenses and projection optics.

Dylan Patel

然后你来到掩模版台,这也是非常疯狂的东西。这个东西的移动速度,我想说 9 个 g,它会以 9 g 的加速度移动,因为当你在晶圆上步进时,设备会移动,而晶圆台是互补的。它是晶圆部分。所以你要对齐这两样东西。你让所有光线通过聚焦的透镜,这里是掩模版。这里是晶圆,你让掩模版向一个方向移动,晶圆向相反方向移动,同时扫描晶圆上 26x33 毫米的区域。然后它停下来,移到晶圆的另一部分,再次扫描,这只需要几秒钟。每个都以 9 g 的加速度向相反方向移动。所以每样东西都是化学、制造、机械工程、光学工程的奇迹,因为你需要对齐所有东西并确保它们完美。所有这些都有大量的计量设备,因为你需要完美地测试一切,因为如果任何东西出错,良率就会降到零,对吧?因为这是一个如此精细调谐的系统。

Then you step forward to the reticle stage, which is also something really crazy. This thing moves at I want to say 9 g's, like it will shift at 9 g's because as you step across a wafer the tool will go, and the wafer stage is complementary. It's the wafer part. So you line these two things up. You're taking all the light through the lenses that's focused and here's the reticle. Here's the wafer and you're passing the reticle moving one direction and the wafer moving the other direction as it scans a 26x33 mm section of the wafer. Then it stops, it shifts over to another part of the wafer and does it again, and it does that in just seconds. Each of them are moving at 9 g's in opposite directions. So each of these things is like a wonder and marvel of chemistry, fabrication, mechanical engineering, optical engineering because you have to align all these things and make sure they're perfect. All these things have crazy amounts of metrology because you have to perfectly test everything because if anything is messed up, the yield goes to zero, right? Because this is such a finely tuned system.

Dylan Patel

顺便说一句,它太大了,以至于你在荷兰费尔德霍芬的工厂里建造它,然后拆解它,用多架飞机运送到客户现场,然后在现场重新组装并再次测试。这个过程需要好几个月。所以供应链中有很多步骤,无论是蔡司制造他们的透镜和投影光学器件,还是 ASML 旗下的公司 Cymer 制造 EUV 光源。每个都有自己的复杂供应链。ASML 曾评论说,他们的供应链中有超过 10,000 人。

By the way, it's so large that you're building it in the factory in Veldhoven, Netherlands, and then you're deconstructing it and shipping it on many planes to the customer site, and then you're reassembling it there and testing it again. That process takes many many months. So there's just so many steps in the supply chain, whether it's Zeiss making their lenses and projection optics or Cymer, which is an ASML owned company making the EUV source. Each of these has its own complex supply chain. ASML has commented their supply chain has over 10,000 people in it.

Dylan Patel

如果你想想,你在谈论两个物理移动的物体,这么大和这么大,晶圆的大小,对吧?而且它必须精确到个位数纳米甚至更小,因为整个系统,套刻精度,层与层之间的变化必须在 3 nm 的量级。所以如果套刻精度是 3 nm,那就意味着每个单独部件,其物理运动的精度必须甚至更小,在大多数情况下必须低于 1 nm,因为这些误差会累积。所以不可能打个响指就增加产量。

If you just think about okay, you're talking about two physically moving objects that are this large and this large, the size of a wafer, right? And it has to be accurate to the level of single-digit nanometers or even smaller because the entire system, the overlay, layer to layer variation has to be on the order of 3 nm. So if the overlay is 3 nm, that means each individual part, the accuracy of its physical movement has to be even less than that, it has to be sub one nanometer in most cases because the error of these things stacks up. So there's no way to just snap your fingers and increase production.

Dylan Patel

像电力这样简单的事情,对吧?美国从 0% 的电力增长到 2% 的电力增长,尽管中国已经达到 30%,这对美国来说非常困难。而且这是一个非常简单的供应链,供应链中的人很少,他们制造困难的东西。美国大概有十万名电工,或者在电力供应链中工作的人更多。而当你看看 ASML,它雇佣的人很少。卡尔蔡司可能雇佣了不到一千人从事这项工作。而且所有这些人都非常专业。所以你不可能在瞬间培训普通人来做这个。你不可能让整个供应链都动员起来。英伟达已经做了很多工作,才让整个供应链甚至能够提供他们今年要生产的产能。

Things as simple as power, right? The US going from 0% power growth to 2% power growth even though China's already at 30% was so hard for America to do. And that's a really simple supply chain with very few people in the supply chain, who make difficult things. There's probably what, a hundred thousand electricians, people who work in the supply chain of electricity in the US or more. And when you look at ASML, employs like so few people. Carl Zeiss probably employs like less than a thousand people working on this. And all of those people are super specialized. So you can't just train random people up for this in the snap of a finger. You can't just get your entire supply chain to get galvanized. Nvidia has had to do a lot to get the entire supply chain to even deliver the capacity they're going to make this year.

Dylan Patel

尽管当你去和 Anthropic 谈时,他们会说,“嗯,我们缺 TPU,我们缺训练,我们缺 GPU。”当你去和 OpenAI 谈时,他们会说,“我们缺这些东西。”所以 OpenAI、Anthropic,他们知道他们需要 X。英伟达并没有那么“AGI 中毒”,他们在建造 X 减一。然后你沿着供应链往下走,每个人都在做减一,在某些情况下他们在做除以二,对吧,因为他们就是没有“AGI 中毒”。所以你最终会看到这个鞭子反应的时间滞后。这种“AI 中毒”程度和增加产量的愿望需要很长时间。然后一旦他们终于明白,“嘿,我们需要快速增加产量,”并且他们认为他们明白了,“哦,AI 意味着我们从 60 增加到 100,此外工具都在变得更好更快,光源功率从 500 瓦增加到 1000 瓦,以及供应链的所有其他方面在技术上进步加上产量增加。”他们认为他们实际上在大量增加产量。

Even though when you go talk to Anthropic, they're like, "Well, we're short of TPUs, we're short of training, we're short of GPUs." When you go talk to OpenAI, they're like, "We're short of these things." So OpenAI, Anthropic, they know they need X. Nvidia is not quite as AGI-pilled, and they're building X minus one. And you go down the supply chain, everyone's doing minus one, and in some cases they're doing divided by two, right, because they just don't get AGI-pilled. So you end up with the time lag for this whip to react. The sort of AI-pilledness and desire to increase production is so long. And then once they finally understand, "Hey, we need to increase production rapidly," and they think they understand, "Oh, AI means we go from 60 to 100 in addition to the tools all just getting better and faster, the source getting higher power from 500 watts to 1000, and all these other aspects of the supply chain advancing technically plus increase of production." They think they're actually increasing production a lot.

计算供应链瓶颈 Compute supply chain bottlenecks

Host

但如果你看看这些数字,埃隆想要什么?他想要到 2028 年或 2029 年每年在太空部署 100 吉瓦。山姆·奥特曼想要到本十年末每年 50 吉瓦。Anthropic 也需要同样的量,谷歌也需要。你纵观整个供应链,就会发现,供应链根本不可能为每个人建造足够的算力容量。

But if you look at the numbers, what does Elon want? He wants 100 gigawatts a year in space by 2028 or 2029. And Sam Altman wants 50 gigawatts a year by the end of the decade. Anthropic needs the same, and Google needs that too. You go across the supply chain and it's like, wait, the supply chain can't possibly build enough capacity for everyone to get what they want on the compute side.

Host

所以我觉得,在过去几年里,数据中心供应链中一直有人争论某个具体环节是瓶颈,因此 AI 算力无法扩展到超过 X。但正如你写过的,如果电网是瓶颈,我们就在现场做表后燃气轮机等。如果那行不通,还有很多其他替代方案。我想问你,我们能否想象半导体供应链也发生类似的事情?如果 EUV 成为瓶颈,我们能不能回到 7 纳米,像中国现在做的那样,用 DUV 机器通过多重图案化生产 7 纳米芯片?看看像 A100 这样的 7 纳米芯片,从 A100 到 B100 或 B200 有很多进步。但其中有多少进步仅仅来自数值格式?如果保持 FP16 不变,从 A100 到 B100,B100 略高于 1 petaflop,而 A100 大约是 300 teraflops。所以从 A100 到 B100 大约有 3 倍的提升。其中一部分是工艺改进,一部分是加速器设计改进,这些我们未来可以再次复制。所以看起来从 7 纳米到 4 纳米的工艺改进效果非常小。如果我们每月有 15 万片 3 纳米晶圆,最终 2 纳米也有类似数量,而 7 纳米也有类似数量,并且由于每片晶圆的比特面积少了 50%,这些旧晶圆有大约 50% 的折扣,那么仅仅拿出 7 纳米晶圆似乎也没那么糟糕。那会再给你 50 或 100 吉瓦。告诉我为什么这很天真。

So I feel like in the data center supply chain for the last few years, people have been arguing that this specific thing is a bottleneck, therefore AI compute can't scale more than X. But then, as you've written about, if the grid is a bottleneck, we just do behind-the-meter on-site gas turbines, etc. If that doesn't work, there are all these other alternatives. And I want to ask you whether we can imagine a similar thing happening in the semiconductor supply chain. So if EUV becomes a bottleneck, what if we just went back to 7 nanometer and do what China is doing, producing 7 nanometer chips with multi-patterning using DUV machines? If you look at a 7 nanometer chip like the A100, there's been a lot of progress from the A100 to the B100 or B200. But how much of that progress is just numerics? If you hold constant FP16, from A100 to B100, the B100 is a little over 1 petaflop and the A100 is about 300 teraflops. So you have basically a 3x improvement from A100 to B100. Some of that is process improvement, some is accelerator design improvement, which we could replicate again in the future. So it seems like the effect from process improving from 7 nanometer to 4 nanometer is very small. So if we have, say, 150k wafers per month of 3 nanometer and eventually similar amounts for 2 nanometer, but there's a similar amount for 7 nanometer, and if you have all those old wafers with maybe a 50% haircut because the bits per wafer area are 50% less, then it doesn't seem that bad to just bring out 7 nanometer wafers. That gives you another 50 or 100 gigawatts. Tell me why that's naive.

Dylan Patel

是的。所以我认为我们确实可能疯狂到这种程度,因为我们只需要增量算力,而且这些算力值得更高的成本和功耗。但在某种程度上,很大程度上也不太可能,因为我认为仅仅比较其中一些是不公平的。例如,从 A100(312 teraflops)到 Blackwell(大约 1000 FP16,或者可能是 2000),再到 Rubin(大约 5000 FP16),这种比较并不公平,因为这些芯片的设计目标截然不同。在 A100 上,那是英伟达优化的目标:FP16、BF16 数值。到了 Hopper,他们没那么关心那个了;他们关心 FP8。到了 Rubin,他们不太关心 FP16 和 BF16;他们主要关心 FP4 和 FP6。所以数值格式就是他们设计芯片的目标。所以,假设我们重新设计,在 7 纳米上做一个新的芯片设计,针对现代数值格式进行优化。性能差异仍然会比你提到的 flops 差异大得多。通常很容易把问题归结为每瓦 flops 或每美元 flops,但这实际上不是一个公平的比较。所以这就是你可以引入的地方,看看 Kimi K1 或 DeepSeek。当你观察 Kimi K2.5 和 DeepSeek,并查看它们在 Hopper 与 Blackwell 上、在高度优化的软件下的性能时,你会得到截然不同的性能。而这大部分并不归因于 flops。很多是数值格式的原因,因为这些模型实际上是 8 位的。所以并不是说 Blackwell 和 Hopper 都针对 8 位优化,而 Blackwell 没有利用其 4 位优势。性能鸿沟实际上要大得多。比较和思考它们的方式是:当然,缩小工艺技术、让晶体管更小是一回事,每个芯片有 X 个 flops,但你忘记了主要的制约因素:这些模型不是在一个芯片上运行,而是同时在数百个芯片上运行。如果你看看 DeepSeek 的生产部署,那已经是一年多以前了,他们当时在 160 个 GPU 上运行。那就是他们服务生产流量的方式。所以他们把模型拆分到 160 个 GPU 上。每次你跨越一个芯片到另一个芯片的边界,都会有成本。

Yeah. So I think we potentially do go crazy enough that this happens because we just need incremental compute and the compute is worth the higher cost and power of these chips. But it's also unlikely to some extent, to a large extent, because I think just comparing some of these is not a fair comparison. For example, from A100 which is 312 teraflops to Blackwell which is like a thousand ish of FP16, or maybe it's 2000, and then Rubin is like 5000 or so FP16. It's not a fair comparison because these chips have vastly different design targets. At A100, that's what Nvidia optimized for: FP16, BF16 numeric. When you look at Hopper, they didn't care as much about that; they cared about FP8. When you look at Rubin, they don't care about FP16 and BF16 as much; they care mostly about FP4 and FP6. So the numeric is what they designed their chip for. So let's say we redesign and make a new chip design on 7 nanometer, optimized for the numerics of the modern day. The performance difference is still going to be much larger than the flops difference you mentioned. Often it's easy to boil things down to flops per watt or flops per dollar, but that's actually not a fair comparison. So this is where you can bring in, let's look at Kimi K1 or DeepSeek. When you look at Kimi K2.5 and DeepSeek, and you look at their performance on Hopper versus Blackwell on very optimized software, you get vastly different performance. And most of this is not attributed to flops. A lot of this is numeric, because those models are actually 8-bit. So it's not like Blackwell and Hopper are both optimized for 8-bit and Blackwell is not really taking advantage of its 4-bit there. The performance gulf is actually much larger. The way you can compare them and think about them is: sure, it's one thing to shrink process technology and make transistors smaller, and each chip has X number of flops, but you forget the big gating factor: these models don't run on a single chip; they run on hundreds of chips at a time. If you look at DeepSeek's production deployment, which is well over a year old now, they were running on 160 GPUs. That's what they serve production traffic on. So they split the model across 160 GPUs. Every time you cross the barrier of a chip to another chip, there's a cost.

数据移动中的效率损失 Efficiency Loss in Data Movement

Dylan Patel

存在效率损失,因为现在需要通过高速电路传输,而且随着制程节点缩小,会产生延迟成本、功耗成本以及各种动态问题。单个芯片上的算力增加了。现在,芯片内部的数据传输速度至少是每秒几十 TB,甚至几百 TB。芯片之间是每秒 TB 级别。不同机架中的芯片之间是每秒几百 Gb,大约每秒 100 GB。所以有一个巨大的阶梯:芯片内通信速度极快,机架内慢一个数量级,机架外更慢。一旦突破芯片边界,就会产生性能损失。我解释这个的原因是,对比 Hopper 和 Blackwell,即使两者都使用一机架的芯片,Hopper 也明显更慢,因为在每个域内——晶体管或处理单元之间每秒几十 TB,处理单元之间每秒 TB——所利用的性能更高,因此性能更高。对于推理,比如 DeepSeek 和 Kimi 2.5 每秒 100 个 token,Hopper 与 Blackwell 的性能差异大约为 20 倍,有趣的是,并非 FLOPS 差异所显示的 2 倍或 3 倍,尽管它们采用相同的制程节点。网络技术和各自的工作重点有所不同。当你看 3 纳米的 Rubin 时,有些东西根本无法移植回 7 纳米的 A100。某些架构改进可以移植,有些则不能。因此性能差异不仅仅是 FLOPS,而是累积的:每芯片 FLOPS、芯片间网络速度、芯片与系统的 FLOPS 比例、单芯片和整个系统的内存带宽。所有这些因素叠加在一起。

There is an efficiency loss because you now have to transmit over high-speed electrical circuits, and there's a latency cost, a power cost, and all these dynamics that hurt as you shrink the process node. You've increased the amount of compute in a single chip. Now, on-chip data movement is at least tens of terabytes per second, if not hundreds. Between chips, it's on the order of terabytes per second. Between chips in different racks, it's hundreds of gigabits per second, roughly 100 gigabytes per second. So you have this huge ladder: on-chip communication at super fast speeds, within the rack at an order of magnitude slower, outside the rack even slower. As you break the bounds of chips, you end up with performance loss. The reason I explain this is because when you look at Hopper versus Blackwell, even if both use a rack worth of chips, Hopper is significantly slower because the performance leveraged within each domain—tens of terabytes per second between transistors or processing elements, and terabytes per second between processing elements—is much higher, so performance is much higher. For inference at, say, 100 tokens per second for DeepSeek and Kimi 2.5, Hopper versus Blackwell shows a performance difference on the order of 20x, interestingly not 2x or 3x as the flops difference indicates, even though they are on the same process node. There are differences in networking technologies and what they've worked on. When you look at Rubin on 3 nanometer, some things are just not possible to port back to A100 even on 7 nanometer. Certain architectural improvements can be ported, others cannot. So the performance difference is not just flops; it's cumulative: flops per chip, networking speed between chips, how many flops on a chip versus a system, memory bandwidth on a single chip and on an entire system. All these things compound.

单芯片上的芯片扩展 Scaling Dies on a Single Chip

Host

我能问一个非常天真的问题吗?今年,去年 B200 在一个芯片上有两个 die,这样就能在单个芯片上获得那种带宽,而无需通过 NVLink 或 InfiniBand。明年 Rubin Ultra 将在一个芯片上有四个 die。是什么阻止我们用旧工艺做到这一点?一个芯片上最多能放多少个 die,还能获得每秒几十 TB 的带宽?

Can I ask you a very naive question? So this year, last year the B200 has two dies on a single chip, so you can get that bandwidth on a single chip without having to go through NVLink or InfiniBand. And then next year Rubin Ultra will have four dies on one chip. What is preventing us from just doing that with an old process? How many dies could you have on a single chip and still get these tens of terabytes per second?

Dylan Patel

即使在 Blackwell 内部,芯片内通信与跨芯片通信也存在性能差异。这些界限比跨整个芯片小得多,但每个 die 与封装内相比仍有损失。当芯片数量增加时,会有一些性能损失——不是完美的,但比不同完整封装好得多。先进封装能扩展到多大?Nvidia 使用 CoWoS,Google 与 Broadcom、MediaTek,Amazon Trainium 都使用类似方法。但看看 Tesla 的 Dojo:Dojo 是一个晶圆大小的芯片,上面有 25 个 die。有取舍——他们无法放置 HBM。但好处是一个封装上有 25 个 die。至今,它可能仍然是运行卷积神经网络的最佳芯片,只是不擅长 Transformer,因为形状、内存、算术等规格不适合 Transformer,而适合 CNN。随着封装变大,你会有其他约束:网络速度、内存带宽、散热能力。这并不简单,但你会看到封装上芯片更多的趋势。你可以在 7 纳米上做到这一点——事实上,华为在他们的 Ascend 910C 或 D 上就是这么做的。他们最初是一个 die,然后做了两个。他们专注于扩大封装,因为这是他们可以比制程技术更快进步的领域,而制程技术他们无法缩小。但归根结底,这也是可以在领先芯片上做的事情。在封装方面,你在 7 纳米上能做的,很可能在 3 纳米上也能做。

Even within Blackwell, there are differences in performance when communicating on the chip versus across chips. Those bounds are much smaller than when going out of the entire chip, but each die versus within the package still has some loss. When you scale the number of chips up, there is some performance loss—it's not perfect, but it is way better than different entire packages. How large can advanced packaging scale? Nvidia uses CoWoS, and Google with Broadcom and MediaTek, Amazon Trainium all use similar approaches. But look at what Tesla did with Dojo: Dojo was a chip the size of an entire wafer with 25 chips on it. There were trade-offs—they couldn't put HBM on it. But the positive side was 25 chips on one package. To date, it is still probably the best chip for running convolutional neural networks, just not great at transformers because the shape, memory, arithmetic, and other specifications are not well suited for transformers. They are well suited for CNNs. As you make packages bigger, you have other constraints: networking speed, memory bandwidth, cooling capabilities. It's not simple, but yes, you will see a trend of more chips on the package. You can do that on 7 nanometer—in fact, that's what Huawei did with their Ascend 910C or D. They started with one die and then did two. They are focusing on scaling packaging up because that is an area where they can advance faster than process technology, where they can't shrink. But at the end of the day, that's something you can do on leading-edge chips too. Anything you do on 7 nanometer, you can probably also do on 3 nanometer in terms of packaging.

中国2030年半导体未来 China's Semiconductor Future by 2030

Host

如果到 2030 年,西方拥有最先进的制程技术但产量没有大幅提升,而中国——我不知道到 2030 年他们是否会有 EUV 和 2 纳米之类的——但他们是半导体强国,大规模生产。我想知道哪一年会出现交叉点,我们在制程技术上的优势消退到一定程度,而他们在规模上的优势增长到一定程度,再加上他们拥有整个供应链和愿景的一个国家,而不是德国、荷兰等零散供应商,这意味着中国在产生大规模 FLOPS 的能力上会领先。

If we end up in this world in 2030 where the West has the most advanced process technology but has not ramped it up as much, whereas China—I don't know if by 2030 they would have EUV and 2 nanometer or whatever—but they are a semiconductor powerhouse producing in mass quantity. I'm wondering what year there is a crossover where our advantage in process technology has faded enough and their advantage in scale has increased enough, and also their advantage in having one country with the entire supply chain and vision rather than random suppliers in Germany and Netherlands, would mean that China would be ahead in its ability to produce mass flops.

Dylan Patel

迄今为止,中国仍然没有完全本土化的半导体供应链。

To date, China still does not have an entire indigenized semiconductor supply chain.

Host

但到 2030 年他们会吗?

But would they by 2030?

Dylan Patel

到 2030 年,他们有可能做到。但迄今为止,中国所有的 7 纳米和 14 纳米产能都使用 ASML 的 DUV 设备。他们能从 ASML 进口的数量很大,但 ASML 的绝大部分收入,尤其是 EUV,全部来自中国以外。

By 2030, it's possible that they do. But to date, all of China's 7 nanometer and 14 nanometer capacity uses ASML DUV tools. The amount they can ship and import from ASML is large, but the vast majority of ASML's revenue, especially on EUV, all of it is outside of China.

规模优势与中国半导体进展 Scale advantage and China's semiconductor progress

Host

所以规模优势仍然在西方加上台湾、日本等一方。

So the scale advantage is still in the favor of the West plus Taiwan, Japan, etc.

Dylan Patel

他们正在尝试制造自己的 DUV 和 EUV 光刻机,对吧?他们正在做所有这些事情。问题在于他们能以多快的速度推进并提升产量和质量。到目前为止,我们还没有看到这一点。不过,我非常看好他们在未来 5 到 10 年内能够做到这些,真正扩大生产,全力加速。他们有更多的工程师在投入,也有更强的意愿投入资本。

They're trying to make their own DUV and EUV tools, right? They're trying to do all these things. The question is how fast can they advance and scale up production as well as quality. And to date, we haven't seen that. Now, I'm quite bullish that they're going to be able to do these things over the next 5 to 10 years, right? Really scale up production, really kick it into high gear. They have more engineers working on it. They have more desire to throw capital at it.

Host

那么到 2030 年,他们能实现完全自主的 DUV 吗?

So by 2030, do they have fully indigenized DUV?

Dylan Patel

我认为肯定可以。肯定。DUV。是的。

I think for sure. For sure. DUV. Yes.

Host

那到 2030 年完全自主的 EUV 呢?

And fully indigenized EUV by 2030?

Dylan Patel

我认为他们会有可用的工具。但我不认为他们能马上大规模生产,对吧?让设备工作是一回事,量产地狱是另一回事。最终,就像 ASML 在 2010 年代初期就让 EUV 以某种能力工作了一样,但当时工具精度不够,没有为高产量制造进行规模化,可靠性也不足。然后他们必须提升产量,这都需要时间。量产地狱需要时间,这就是为什么又花了 5 到 7 年才让 EUV 在晶圆厂进入量产,而不仅仅是在实验室里工作。

I think they'll have working tools. I don't think that they'll be able to manufacture a bunch yet, right? There's having it work and then there's production hell. Ultimately, like ASML had EUV working in the early 2010s at some capacity, right? Now the tools were not accurate enough. They were not scaled for high volume manufacturing, reliable enough. And then they had to ramp production and that all took time. Production hell takes time, which is why it took another 5 to 7 years to get EUV into mass production at a fab rather than just working in the lab.

Host

那么你认为中国在 2030 年能制造多少台 DUV 光刻机?

So how many DUV tools do you think China will manufacture in 2030?

Dylan Patel

哦,这是个好问题。尤其是要深入了解这个供应链有点挑战。我们非常努力。但在某些情况下,他们从日本供应商那里购买东西,如果他们想要完全自主的供应链,就不能从日本供应商那里购买这些镜头、投影光学器件或工件台。他们需要内部制造。所以真的很难说他们能达到什么程度。老实说,我觉得这就像瞎猜。但他们每年大概能生产 100 台 DUV 光刻机,这并非不可能。而 ASML 目前每年生产数百台 DUV 光刻机。

Oh, that's a great question. It's a bit of a challenge to look into this supply chain especially. We try really hard. But in some instances, they're buying stuff from Japanese vendors, and if they want a fully indigenized supply chain, they need to not buy these lenses or projection optics or stages from Japanese vendors. They need to build it internally. So it's really tough to say where they'll be able to get to. I honestly think it's like a shot in the dark. But it's probably not unlikely that they'll be able to do on the order of 100 DUV tools a year. Whereas ASML is doing hundreds of DUV tools a year currently.

Host

还没有人实现一个月生产 100 万片晶圆的工艺节点,对吧?马斯克说他想做到,中国显然也会去做。我不认为台积电在尝试这个。存储制造商也可能达到每月 100 万片晶圆的水平,但不是在单个晶圆厂里。想到这个规模就让人难以置信,而且很难看到供应链为此动员起来。所以我不确定,我不想怀疑中国的规模化能力。

No one's made a process node where they make a million wafers a month, right? Elon says he wants to do it and China's obviously going to do it. I don't think TSMC is trying to do that. The memory makers may get there as well, to the million wafers a month, but not in a single fab. It's mind-boggling to think of that scale and challenging to see the supply chain galvanized for that. So I'm not sure, I don't want to doubt China's capability to scale.

Dylan Patel

我觉得这是个有趣的问题。我想某个时候 SemiAnalysis 会对此进行深入分析。但这个问题是:中国何时能够……如果把你模型的所有输入加起来,自主的中国半导体生产可能比西方其他国家的总和还要大。当他们拥有可扩展的 DUV 光刻机,当他们拥有可扩展的 EUV 光刻机时。因为我认为有一个问题:如果你对 AI 的时间线看得比较长,长到 2035 年——这在宏观上并不算长——你是否应该预期一个中国主导半导体的世界?我认为在旧金山,这个问题没有被足够多地提出。我们只考虑几周的时间尺度。而如果你不在旧金山,你根本不会考虑 AGI。所以这个问题是:好吧,如果我们有了 AGI,如果我们有了这个变革性的东西,它驱动着数十万亿或数百万亿美元的经济增长和 token 输出等等,但它在 2035 年才发生?这对西方与中国意味着什么?我认为 SemiAnalysis 必须为此写出权威的模型。

I guess this is an interesting question. I think at some point SemiAnalysis will do a deep dive on this. But this question of by when would China be able... Indigenized Chinese production could be bigger than the rest of the West combined if you just add up all the inputs of your model. When they'll have DUV machines that scale, when they'll have EUV machines that scale. Because I think there's this question around if you have long timelines on AI, by long meaning 2035 which is not that long in the grand scheme of things, should you expect a world where China is dominating in semiconductors? I think that doesn't get asked enough in San Francisco. We're just thinking on time scales of weeks. And if you're outside of San Francisco, you're not thinking about AGI at all. So this question of: okay, what if we have AGI, what if you have this transformational thing that is commanding tens of trillions of dollars or hundreds of trillions of dollars of economic growth and token output and so forth, but then it happens in 2035? What does that imply for the West versus China? I think SemiAnalysis has got to write the definitive model on this.

Dylan Patel

是的。所以我认为当你把时间尺度拉得那么远时,真的很有挑战性。我们通常关注的是追踪每个数据中心、每个晶圆厂、所有设备以及它们的去向。但这些事情的时间滞后相对较短。我们只能根据土地购买、许可证、涡轮机购买等来对数据中心容量做出相当准确的估计。我们知道这些东西的去向,这就是我们销售的数据。但当你展望 2035 年时,情况会变得截然不同,误差范围变得非常大,很难做出估计。但归根结底,如果起飞或时间线足够慢,那么我当然看不出他们为什么不能大幅追赶。从某种意义上说,我们有一个低谷,大约 3 到 6 个月前,中国模型达到了它们有史以来最具竞争力的水平。也许现在也是。我认为 Opus 4.6 和 GPT 5.4 已经真正拉开了差距,让差距变大了一点。但我相信一些新的中国模型会出现。随着我们从公司销售 token(提供整个推理链)转向销售自动化白领工作、自动化软件工程师——发送请求,他们返回结果,后台有一堆思考过程不展示给你——从美国模型中蒸馏到中国模型的能力将变得更难。另外,随着实验室拥有的算力规模——OpenAI 去年以大约 2 吉瓦的容量结束。Anthropic 今年将达到 2 吉瓦以上,到明年年底,它们都将达到大约 10 吉瓦的容量。中国在 AI 实验室算力上的扩展速度远没有那么快。所以在某个时候,当你无法从这些实验室中蒸馏出知识到中国模型,再加上 OpenAI、Anthropic、Google、Meta 都在进行的算力竞赛,最终模型性能应该开始出现更大的分化。然后所有这些资本支出都花在数据中心等等上,亚马逊 2000 亿,谷歌 1800 亿,等等。所有这些公司都在花费数千亿美元的资本支出。

Yeah. So I think it's really challenging when you move time scales out that far. What we tend to focus on is tracking every data center, every fab, all the tools, and where they're going. But the time lags for these things are relatively short. We can only make reasonably accurate estimates for data center capacity based on land purchasing, permits, turbine purchasing, and all these things. We know where all these things are going, and that's what the data we sell is. But as you go out to 2035, things are just so radically different and your error bars get so large, it's kind of hard to make an estimate. But at the end of the day, if takeoff or timelines are slow enough, then certainly I don't see why they wouldn't be able to catch up drastically. In some sense, we've got this valley where, call it 3 to 6 months ago, Chinese models were as competitive as they've ever been. Maybe even now. I think Opus 4.6 and GPT 5.4 have really pulled away and made the gap a little bit bigger. But I'm sure some new Chinese models will come out. As we move from companies selling tokens where they provide the entire reasoning chain to selling automated white-collar work, automated software engineer, send them the request, they give you the result back, and there's a bunch of thinking on the back end that they don't show you. The ability to distill out of American models into Chinese models will be harder. Also, as the scale of the compute that the labs have, OpenAI exited the year with roughly 2 gigawatts last year. Anthropic will get to 2 plus gigawatts this year, and by the end of next year they'll both be at like 10 gigawatts of capacity. China is not scaling their AI lab compute nearly as fast. So at some point, when you can't distill the learnings from these labs into the Chinese models, plus this compute race that OpenAI, Anthropic, Google, Meta are all racing on, at some point they end up getting to a point where the model performance should start to diverge more. And then all of this capex being spent on data centers and all that, Amazon on 200 billion, Google 180, so on and so forth. All these companies are spending hundreds of billions of dollars of capex.

中美AI分化 US vs China AI divergence

Host

今年美国数据中心资本支出接近一万亿美元,对吧?那么投资资本回报率是多少?你我会认为数据中心资本支出的投资资本回报率非常高。至少从 Anthropic 的收入来看,一月份他们增加了大约 40 亿美元。二月份是短月,他们增加了大约 60 亿美元。我们看看三、四月能做什么,因为算力限制是增长的瓶颈。Claude Code 的可靠性实际上很低,因为他们算力太紧张了。如果这种情况持续,这些数据中心的投资资本回报率会非常高。在某个时间点,美国经济会因为所有这些资本支出和模型产生的收入以及下游供应链而开始加速增长,而中国还没有做到这一点。他们还没有建立足够规模的基础设施来投资模型,以获得能力并大规模部署这些模型。看看 Anthropic,他们年化收入大约 200 亿美元。利润率低于 50%,至少根据 The Information 上次的报告。那么,他们运行的算力按租赁成本算大约是 1340 亿美元,这实际上相当于有人为 Anthropic 投入了约 500 亿美元的资本支出,以产生当前的收入。中国还没有做到这一点。如果 Anthropic 的收入再次增长 10 倍——我认为答案是“何时”而非“是否”——那么中国就没有算力来大规模部署。所以有一种感觉,我们正处于快速起飞阶段。不是说某天建戴森球,而是收入以这样的速度复合增长,确实影响经济增长,这些实验室积累资源的速度非常快,而中国还没有做到。所以在这种情况下,美国和西方实际上在拉开差距。另一方面,这些基础设施投资回报平平,也许没有预期的那么好。也许谷歌错了,想把自由现金流降到零,明年花 3000 亿美元资本支出。也许他们就是错了。华尔街看空的人和不懂 AI 的人是对的。在这种情况下,美国建设了所有这些产能,但没有获得很好的回报,而中国能够建立完全垂直的本土化供应链,而不是美国、日本、韩国、台湾、东南亚、欧洲等国家一起建设这种不那么垂直的供应链。从某种意义上说,如果 AI 需要更长时间才能达到某些能力水平,中国最终可能超越我们。你播客上的绝大多数嘉宾认为,时间线短则美国赢,时间线长则中国赢,对吧?但我不确定“时间线短”是什么意思。我认为你不必相信 AGI 才能有美国赢的时间线。

There's nearly a trillion dollars of capex being invested in data centers in America this year roughly, right? What's the return on invested capital here? You and I would think that the return on invested capital for data center capex is very high. At least if we look at Anthropic's revenues in January, they added like $4 billion. In February, which was a shorter month, they added like $6 billion. We'll see what they can do in March and April, given compute constraints are what's bottlenecking their growth. The reliability of Claude Code is actually quite low because they're so compute constrained. If this continues, then the ROIC on these data centers is super high. At some point, the US economy starts growing faster and faster over the next year because of all this capex and all this revenue that these models are generating, and downstream supply chain, versus China doesn't have that yet. They have not built the scale of infrastructure to then invest in models to get to the capabilities to then deploy these models at such scale. When you look at Anthropic, they're at call it $20 billion ARR. The margins are sub-50%, at least last reported by The Information. So then you're at like $134 billion of compute that it's running on rental cost-wise, which is actually like $50 billion worth of capex that someone laid out for Anthropic to generate their current revenue. China has just not done this. If and when Anthropic 10x's revenue again, and I think our answer would be when, not if, then China doesn't have the compute to deploy at that scale. So there is some sense of like, oh, we're in fast takeoff-ish. It's not like we're talking about Dyson sphere by X day. It's more like the revenue is compounding at such a rate that it does affect the economic growth, and the resources these labs are gathering are going so fast, and China hasn't done that yet. So in that case, the US and the West is actually diverging. The flip side is actually these infrastructure investments have middling returns. Maybe they're not as good as hoped. Maybe Google is wrong for wanting to take free cash flow to zero and spend $300 billion on capex next year. Maybe they're just wrong. People on Wall Street who are bearish and people who don't understand AI are correct. In which case, the US is building all this capacity, it doesn't get really great returns, and China is able to build the fully vertical indigenized supply chain, not US, Japan, Korea, Taiwan, Southeast Asia, Europe, all these countries together building this less vertical supply chain. In a sense, at some point China is able to scale past us if AI takes longer to get to certain capability levels. The vast majority of your guests on this podcast believe fast timelines, US wins, long timelines, China wins, right? But I don't know what fast timelines means. I don't think you have to believe in AGI to have the timelines where the US wins.

内存短缺与替代内存方案 Memory crunch and alternative memory solutions

Host

我们回到内存问题,因为我认为华尔街的人和行业人士明白这有多重要,但也许普通人并不理解。正如你所说,我们面临内存紧缩。之前我问过,能否通过回到 7 纳米来解决 EUV 工具短缺。现在我问一个类似的内存问题。HBM 由 DRAM 制成,但每晶圆面积的比特数比其基础的 DRAM 少 3 到 4 倍。未来的加速器能否直接使用普通 DRAM 而不是 HBM?这样我们可以从 DRAM 中获得更多容量。我认为这可能的原因是,如果我们有智能体自主工作,而不是同步聊天机器人应用,那么你就不再需要极高的低延迟。所以也许你可以用低带宽,因为将 DRAM 堆叠成 HBM 是为了更高带宽。是否可能转向 HBM 加速器,并基本上实现与 Claude Code 快速相反的效果,比如让 Claude Code 变慢?

Let's go back to memory because I think people on Wall Street and people in the industry understand how big this is, but maybe generally people don't understand how big a deal this is. So we've got this memory crunch as you're talking about. Earlier I was asking about could we solve for the EUV tool shortage by going back to 7 nanometers. So let me ask a similar question about memory. HBM is made of DRAM but has 3 to 4x less bits per wafer area than the DRAM it's made out of. Is it possible that accelerators in the future could just use commodity DRAM and not HBM? And so we can make much more capacity out of the DRAM we get. The reason I think this might be possible is look, if we're going to have agents that are just going off and doing work and it's not a synchronous chatbot application, then you don't necessarily need extremely high fast latency kinds of things anymore. So maybe you can have the low bandwidth because the reason you stack DRAM into stacks and make HBM is for higher bandwidth. Is it possible to go to HBM accelerators and basically have the opposite of Claude Code fast, like have Claude Code slow, and do that?

Dylan Patel

是的,我认为归根结底,愿意为 token 支付最高价格的增量买家也是价格不那么敏感的人。在资本主义社会中,算力应该分配给价值最高的商品,私人市场通过支付意愿来决定。所以某种程度上,Anthropic 确实可以发布慢速模式。他们可以发布 Claude 慢速模式,大幅提高每美元的 token 数量。他们可能将 Opus 4.6 的价格降低 4 到 5 倍,速度降低大约 2 倍。推理吞吐量与速度之间的曲线在 HBM 上已经存在,但他们没有这样做,因为没人想用慢速模型。此外,在这些智能体任务中,模型能在几小时的时间范围内运行很好。但如果模型运行得更慢,那几小时就会变成一天,对吧?反之,如果模型运行更快,几小时就变成一小时。但没人愿意等待一整天,因为最高价值的任务也有时间敏感性。所以我认为很难,是的,你可以用 DDR。但有几个挑战。你可以用普通 DRAM。一是你仍然受限。芯片的核心限制之一是,尽管芯片有一定尺寸,所有 IO 都从芯片边缘引出。通常,芯片的左右两侧是 HBM。芯片到 HBM 的 IO 在侧面,而顶部和底部是到其他芯片的 IO。所以如果你从 HBM 换成 DDR,那么边缘的 IO 带宽会显著降低,但每芯片的容量会显著增加。

Yeah, I think at the end of the day, the incremental purchaser who's willing to pay the highest price for tokens also ends up being the one that's less price sensitive. The compute should be allocated in a capitalistic society towards the goods that have the highest value, and the private market determines this by willingness to pay. So to some extent, sure, Anthropic could actually release a slow mode. They could release Claude slow mode and have an increase in tokens per dollar by a significant amount. They could probably reduce the price of Opus 4.6 by 4x or 5x and reduce the speed by maybe just 2x. The curve on inference throughput versus speed is there already just on HBM, and yet they don't, because no one actually wants to use a slow model. Furthermore, on these agentic tasks, it's great that the model can run at this time horizon of hours. That's kind of like, okay, well if the model was just running slower, that hours would become a day, right? Or vice versa, if the model's running faster, that hours becomes an hour. Yet no one really wants to move to that day-long wait period because the highest value tasks also have some time sensitivity to them. So I struggle to see, yes, you could use DDR. But then there are a couple of things that are challenging with this. You could use regular DRAM. One is you're still limited. One of the core constraints of chips, even though they're a certain size, all of the IO escapes on the edges of the chip. Often times, what you see is the left and the right of the chip are HBM. The IO from the chip to the HBM is on the sides, and then the top and bottom are IO to other chips. So if you were to change from HBM to DDR, then all of a sudden this IO on this edge would have significantly less bandwidth, but it had significantly more capacity per chip.

内存带宽与容量限制 Memory Bandwidth vs Capacity Constraints

Host

因为限制算力的瓶颈就是让下一个矩阵进出,为此你只需要更多带宽。

Because the thing that is constraining the flops is just getting in and out the next matrix. And for that you just need more bandwidth.

Dylan Patel

对。把权重取出来,把 KV 缓存取出来,对吧?所以在很多情况下,这些 GPU 并没有跑满内存容量。这显然是个系统设计问题——模型、硬件、软件、代码设计——我要做多少 KV 缓存?多少留在芯片上?多少卸载到其他芯片,需要时再调回来(比如工具调用之类的)?我在多少芯片上并行处理?搜索空间非常广,所以我们才有像开源模型那样的推理引擎,在八种不同芯片和模型上搜索所有推理最优解。关键在于,你不一定总是受限于内存容量;你可能受限于算力、网络带宽、内存带宽或内存容量。简化来说有四个约束,每个还可以细分。在这个例子里,如果换成 DDR,是的,每片 DRAM 晶圆能产出 4 倍的比特数,但突然间约束条件大变,系统设计也大变。你变慢了。市场变小了吗?也许。但现在所有这些算力都浪费了,因为它们只是闲置着等内存。你不需要那么多容量,因为你没法真正增大批处理大小,因为 KV 缓存的读取时间会更长。

Yeah. Getting out the weights and getting in and out the KV cache, right? And so in many cases these GPUs are not running at full memory capacity. It's obviously a system design thing—model, hardware, software, code design—of how much KV cache do I do? How much do I keep on the chip? How much do I offload to other chips and call when I need it for tool calling or whatever? How many chips do I parallelize this on? The search space is very broad, which is why we have inference engines like open-source models that search all the optimal points for inference on a variety of eight different chips and models. The point is you're not always necessarily constrained by memory capacity; you can be constrained by flops, network bandwidth, memory bandwidth, or memory capacity. There are four constraints if you simplify it down, and each can break out into more. In this case, if you switch to DDR, yes, you produce 4x the bits per DRAM wafer, but all of a sudden the constraints shift a lot and your system design shifts a lot. You go slower. Is the market smaller? Maybe. But now all these flops are wasted because they're just sitting there waiting for memory. You don't need all that capacity because you can't really increase batch size since the KV cache would take even longer to read.

Host

有意思。HBM 和普通 DRAM 的带宽差距有多大?

Interesting. What is the bandwidth difference between HBM and normal DRAM?

Dylan Patel

以 HBM4 的 HBM 堆叠为例,就说 Rubin 里的配置吧,因为我们一直在拿它做基准。它的接口宽度是 2048 位,连接区域大约 13 毫米宽。所以 2048 位,传输速率大约 10 Giga transfers/s。那么 HBM,一个 HBM4 堆叠,在约 13 毫米(或 11 毫米)宽的区域内提供 2048 位带宽,以 10 Giga transfers/s 传输。相乘再除以 8 位/字节,每个 HBM 堆叠大约 2.5 TB/s。再看 DDR,同样面积下,接口宽度可能只有 64 或 128 位。DDR5 的传输速率在 6.4 到 8 Giga transfers/s 之间。所以带宽低得多。64 * 8,000 / 8 = 64 GB/s。即使按宽松的 128 * 8 Giga transfers 算,同一 shoreline 也只有 128 GB/s,而 HBM 是 2.5 TB/s。每边缘面积的带宽差了一个数量级。如果你的芯片是正方形,或者尺寸是 26 mm × 33 mm(单芯片裸片的最大尺寸),你的边缘面积有限,芯片内部放的是所有算力。你可以做一些事情来改变——比如更多 SRAM、更多缓存等等——但归根结底,你严重受限于带宽。

So, an HBM stack of HBM4, let's talk about the stuff in Rubin since that's what we've been indexing on, is 2048 bits across connected in an area that's like 13 millimeters wide. So 2048 bits and it transfers memory at around 10 gig transfers per second. So HBM, a stack of HBM4, is 2048 bits on an area that's 13 mm wide roughly or 11. That's the shoreline you're taking on the chip. In that shoreline, you have 2048 bits transferring at 10 giga transfers per second. Multiply those together and divide by eight bits to bytes, you're at roughly 2.5 terabytes per second per HBM stack. When you look at DDR, in that same area, it's maybe 64 or 128 bits wide. And DDR5 is transferring at anywhere from 6.4 giga transfers per second to maybe 8 giga transfers per second. So your bandwidth is significantly lower. It's 64 * 8,000 divided by 8, you're at 64 gigabytes per second. Even if you take a generous interpretation of 128 * 8 gig transfers, you're at 128 gigabytes per second for the same shoreline versus 2.5 terabytes per second. There's an order of magnitude difference in bandwidth per edge area. If your chip is a square or it's 26 by 33 mm, the maximum size for a chip individual die, you only have so much edge area, and on the inside of that chip you put all your compute. There are things you can do to try and change that—more SRAM, more caching, etc.—but at the end of the day, you're very constrained by bandwidth.

Host

有意思。那么问题来了,你可以在哪里摧毁需求,为 AI 腾出足够空间?我想情况尤其糟糕,因为正如你所说,如果为了获得同样比特数的 HBM 需要多花 4 倍的晶圆面积,那么为了给 AI 腾出一个比特,你就得摧毁 4 倍的消费者需求(来自笔记本电脑、手机等)。那么这对未来一两年意味着什么?我记得你在新闻通讯里说,2026 年大型科技公司的资本支出中有 30% 将流向内存。

Interesting. So then there's a question of where can you destroy demand to free up enough for AI. And I guess the picture is especially bad because as you're saying, if it takes 4x more wafer area to get the same bits for HBM, you had to destroy 4x as much consumer demand for laptops and phones and whatever in order to free up one bit for AI. So what does this imply for the next year or two? I think in your newsletter you said 30% of the capex in 2026 of big tech is going towards memory.

Dylan Patel

这很疯狂,对吧?在 6000 亿或别的数字里,你说 30% 只是流向内存。显然英伟达做了一些利润叠加。如果你把他们的利润分摊到内存和逻辑上,到头来,是的,他们资本支出的三分之一流向了内存。

That's insane, right? Of the 600 billion or whatever, you're saying 30% is going just to memory. And obviously there's some level of margin stacking that Nvidia does. So if you separate out and apply their margin to the memory and the logic, at the end of the day, yeah, like a third of their capex is going to memory.

Host

太疯狂了。好吧,那么随着内存紧缩的到来,未来一两年我们应该期待什么?

That's crazy. Okay, so what should we expect over the next year or two as this memory crunch hits?

Dylan Patel

内存紧缩会越来越严重。价格持续上涨,这对市场不同部分的影响也不同。这就引出了一个问题:人们会不会越来越讨厌 AI?是的,因为现在智能手机和 PC 不会再逐年变好了;事实上,它们会逐年变差。如果你看一部 iPhone 的物料清单,内存占多大比例?如果内存价格翻倍,iPhone 会贵多少?我相信一部 iPhone 有 12 GB 内存。每 GB 过去大约 3 到 4 美元,所以是 50 美元。但现在内存价格涨了三倍。我们按 DDR 每 GB 12 美元算吧。那么现在是 150 美元对 50 美元,苹果成本增加了 100 美元。苹果也有利润率;他们不会就这么吃掉利润。所以成本增加了 100 美元。这还只是 DRAM。NAND 也有类似的市场。所以 iPhone 的成本可能增加了 150 美元。苹果要么转嫁给消费者,要么自己消化。我不认为苹果会大幅降低利润率。也许他们消化一点,但到头来,终端消费者要多付 250 美元买一部 iPhone。这是去年内存价格和今天的对比。苹果感受到压力会有一些滞后,因为他们通常签三、六或一年的内存合同,但苹果最终会受到很大冲击。他们要到下一代 iPhone 发布才会真正调整。但那是高端市场。每年只有几亿部手机。苹果每年卖多少?2 到 3 亿部。市场的主体是中低端。以前每年卖 14 亿部智能手机。

Memory crunch will continue to be harder and harder. Prices continue to go up, and this affects different parts of the market differently. It gets to the question of whether people are going to hate AI more and more. Yes, because now smartphones and PCs are not going to get incrementally better year over year; in fact, they're going to get incrementally worse. If you look at the bill of materials of an iPhone, what fraction is memory? How much more expensive does an iPhone get if memory is 2x more expensive? I believe an iPhone has 12 gigabytes of memory. Each gig used to cost roughly three or four dollars, so it's $50. But now the price of memory has tripled. Let's call it $12 per gig for DDR. So now you're talking about $150 versus $50, a $100 increase in cost on Apple. Apple also has some margin; they're not just going to eat the margin. So that's a $100 cost increase. That's just on the DRAM. The NAND also has the same sort of market. So it's probably a $150 increase on the iPhone. Apple has to either pass it on to the consumer or eat it. I don't see Apple reducing their margin too much. Maybe they eat a little bit, but at the end of the day, the end consumer is paying $250 more for an iPhone. That's last year's memory pricing versus today's. There's some lag for Apple to feel the heat because they have tended to have three, six, or year-long contracts for a lot of memory, but Apple gets hit pretty hard by this. They won't really adjust until the next iPhone release. But that's the high end of the market. That's only a few hundred million phones a year. Apple sells what, 200-300 million phones a year. The bulk of the market is mid-range, low end. Used to be 1.4 billion smartphones sold a year.

内存价格动态与智能手机影响 Memory price dynamics and smartphone impact

Dylan Patel

现在大概是 1.1,但我们的预测是今年可能降到 8 亿,明年降到 6 亿或 5 亿。我们看了来自中国的一些数据点,是我们亚洲、新加坡、香港和台湾的分析师提供的。他们一直在追踪,发现小米和 OPPO 正在将低端和中端智能手机的产量削减一半。因为没错,对于 1000 美元的智能手机来说,只是涨价 150 美元,或者对于 1000 美元的 iPhone 来说,BOM 成本增加 150 美元,而苹果的利润率更高。但如果我们看更小的手机,内存和存储占 BOM 的比例要大得多,而且利润率更低。所以它们消化涨价的能力更弱。而且它们通常不会签订长期内存协议。为什么这很重要?因为如果智能手机销量减半,减半实际上会发生在低端和 mid-range,而不是高端。所以释放的比特数并不会减半,对吧?目前消费类产品占内存需求的一半以上。即使智能手机销量减半,由于减半的形态,低端削减超过一半,高端削减不到一半,因为你我会购买价格超过 1000 美元的高端手机,即使它们稍微贵一点我们也会买。苹果的销量下降幅度不会像低端智能手机供应商那么大。PC 也是如此。这对市场的影响相当剧烈。DRAM 被释放给 AI 芯片,它们愿意签订更长期的合同,愿意支付更高的利润率等等,因为最终它们从终端用户那里提取的利润要大得多。所以这可能会导致人们更加讨厌 AI,因为他们会开始像现在这样——你已经在 PC 子版块和 PC Twitter、游戏 PC Twitter 上看到所有那些梗图,比如猫跳舞视频,配文是“这就是内存价格翻倍、你买不到新游戏 GPU 或新台式机的原因”。当内存价格再次翻倍时,情况会更糟,尤其是 DRAM。

Now we're at like 1.1, but our projections are we maybe get down to like 800 million this year and next year like 600 or 500 million. And we look at some data points out of China from some of our analysts in Asia and Singapore and Hong Kong and Taiwan. They've been tracking this and they see Xiaomi and OPPO are cutting low-end and mid-range smartphone volumes by half. Because yes, it's only a $150 price increase on a $1,000 smartphone, or $150 bomb increase on $1,000 iPhone where Apple has some larger margin. But if we look at the smaller phones, the percentage of the BOM that goes to memory and storage is much larger and the margins are lower. So there's less capacity to even eat the margins. And they have generally tended not to do as long-term agreements on memory. Why this is a big deal is if smartphone volumes halve, the halving will frankly happen in the low and mid-range, not in the high end. So it's not like the bits released are halving, right? Currently consumer is more than half of memory demand. Even if you halve the smartphone volumes because of the shape of the halving, low-end gets cut by more than half. High-end gets cut by less than half because you and I will buy the high-end phones that cost north of $1,000, we'll buy them even if they get a little more expensive. And Apple's volumes will not go down as much as a low-end smartphone provider. The same applies to PCs. What this does to the market is quite drastic. DRAM gets released to AI chips who are willing to do longer term contracts, willing to pay higher margins, etc., because at the end of the day, the margin that they extract is much larger from the end user. So this probably leads to people hating AI even more, because they're going to start being like, today you already see all the memes on PC subreddits and PC Twitter, gaming PC Twitter, like cat dancing videos and it's like this is why memory prices doubled and you can't get a new gaming GPU, or you can't get a new desktop. And it's going to be even worse when memory prices double again, especially DRAM.

Dylan Patel

另一个非常有趣的动态是,不仅仅是 DRAM,还有 NAND。NAND 也在涨价。这两个市场在过去几年中产能扩张都非常缓慢。NAND 几乎为零,但智能手机中,用于手机和 PC 的 NAND 比例大于用于手机和 PC 的 DRAM 比例。所以当你摧毁需求时,你会释放出更多——主要是为了 DRAM 的目的——你会释放出更多 NAND,这些 NAND 可以被分配到其他市场,因此 DRAM 的价格涨幅将大于 NAND,因为你从消费类产品中释放了更多。实际上,你为 AI 生产了更多内存。

Another dynamic that's quite interesting is it's not just DRAM, it's also NAND. NAND is also going up in price. Both of these markets have expanded capacity very slowly over the last few years. NAND almost zero, but smartphones, the percentage of NAND that goes to phones and PCs is larger than the percentage of DRAM that goes to phone and PC. So as you destroy demand, you unlock more, mostly for the DRAM purposes, you unlock more NAND that gets allocated and can sort of go to other markets, and so the price increases of DRAM will be larger than those of NAND because you've released more from the consumer. In fact, you've produced more memory for AI.

Host

抱歉,NAND 的部分我可能刚解释过但错过了。是因为数据中心大量使用 SSD 吗?

Sorry, but the NAND is I maybe just explained it and I missed it. Is it because SSDs are being used in large quantities for data centers?

Dylan Patel

是的,但数量不如 DRAM 大。好吧。但你是说它们也会涨价,因为用了一些,但需求不如 HBM 那么大。有道理。

They are, but not as large quantities as DRAM. Okay. But so you're saying they will also increase because they're using some quantity, but there's not as much a need as there is for HBM. Makes sense.

Dylan Patel

有一件事我直到读了你们的一些通讯才意识到,那就是未来几年阻碍逻辑 Scaling 的约束,与阻碍我们生产更多内存晶圆的约束非常相似。事实上,完全相同的机器——EUV 工具——也是内存所需的。所以我想也许有人现在会问,为什么我们不能生产更多内存?

One thing I didn't appreciate until I was reading some of your newsletters is that basically the same constraints that are preventing logic scaling over the next few years are quite similar to what's preventing us from producing more memory wafers. In fact, literally the same exact machine, this EUV tool, is needed for memory. So I guess maybe there's a question that somebody could be asking right now like, well why can't we just make more memory?

Host

是你吗?

Is that somebody you?

Dylan Patel

是啊,谁知道呢。所以我认为,正如我之前提到的,约束并不一定是现在的 EUV 工具或明年的。它们会在本十年后期成为约束。但目前,约束更多的是他们根本没有建造晶圆厂。过去三四年里,这些供应商根本没有建新厂。那是因为内存价格非常低,利润率低,实际上 2023 年他们在内存上亏损。所以他们想,‘哦,我们不建新厂了。’然后市场慢慢复苏,但直到去年才真正好转。2024 年我们一直在敲锣打鼓地说,推理意味着长上下文,意味着大的 KV 缓存,意味着你需要大量内存需求。我们谈论这个已经一年半、两年了。懂 AI 的人当时就做多内存。所以你看到了那种动态,但现在它终于在定价中体现出来了。显而易见的事情花了这么长时间:长上下文,KV 缓存变大,你需要更多内存,而加速器一半的成本是内存。所以他们当然会开始疯狂投入。这花了一年时间才真正反映在内存价格上。一旦内存价格反映出来,又花了 6 个月、3 个月内存供应商才开始建厂。而这些厂需要两年才能建成。所以直到 2027 年底或 2028 年,我们才有真正有意义的晶圆厂来放置这些工具。相反,你看到的是为了获得产能而采取的一些非常疯狂的措施。美光从一家台湾公司买了一个生产落后芯片的晶圆厂。海力士和三星正在做一些相当疯狂的事情来扩大现有晶圆厂的产能,这些也对经济产生了非常大的连锁反应。那么,为什么我们不能建造更多产能?因为没有地方放工具。而且不仅仅是 EUV,DRAM 和逻辑还涉及其他工具。对于逻辑,N3 大概占成本的 28% 是 EUV(最终晶圆)。当你看到 DRAM 时,它在百分之十几。它在上升,但仍在百分之十几。所以它占成本的比例要小得多。这些其他工具也是瓶颈,尽管它们的供应链不像 ASML 那么复杂。所以你看到应用材料、泛林研究和其他公司也在大幅扩张产能。总之,你没有地方放工具,因为人们制造的最复杂的建筑就是晶圆厂,而晶圆厂需要两年才能建成。

Yeah, who knows. So I think the constraints as I was mentioning earlier are not necessarily EUV tools today or to next year. They become that as we get to the latter part of the decade. But currently, the constraints are more so they physically just haven't built fabs. Over the last 3 to 4 years, these vendors have just not built new fabs. That's because memory prices were really low. Their margins were low and in fact they were losing money in 2023 on memory. So they were like, 'Oh, we're not building new fabs.' And then the market slowly recovered over time, but never really got amazing until last year. In 2024 we were like banging on the drums that reasoning means long context, which means large KV cache, which means you need a lot of memory demand. We've been talking about that for like a year and a half, two years. People who understand AI went long memory then. So you've seen that sort of dynamic, but now it finally played out in pricing. It took so long for what was obvious: long context, KV cache gets bigger, you need more memory, and accelerators half their cost is memory. So of course they're just going to start going crazy on it. It took a year for that to actually reflect in memory prices. Once memory prices reflected, then it took another 6 months, 3 months for the memory vendors to start building fabs. And those fabs take two years to build. So we don't have really meaningful fabs that you can even put these tools in until late 2027 or 2028. Instead, what you've seen is some really crazy stuff to get capacity. Micron bought a fab from a company in Taiwan that makes lagging edge chips. Hynix and Samsung are doing some pretty crazy things to try and expand capacity at their existing fabs, that also have very large knock-on effects in the economy. So, why can't we build more capacity? There's nowhere to put the tools. And it's not just EUV, there's other tools involved in DRAM and logic. For logic, N3% or so of the cost, 28% of the cost is EUV of the wafer of the final wafer. When you look at DRAM, it's in the teens. It's going up, but it's in the teens. So it's a much smaller percentage of the cost as DRAM or as EUV. These other tools are also bottlenecks, although their supply chains are not as complex as ASML's. And so you see Applied Materials and Lam Research and all these other companies also expanding capacity a lot. Anyways, you don't have anywhere to put the tool because the most complex building that people make is fabs, and fabs take two years to build.

Jane Street广告 Jane Street Ad

Host

他们的基础设施团队构建了世界上最大的研究集群之一,拥有数万块高端 GPU、数十万个 CPU 核心和 EB 级存储。这些算力正是 Jane Street 从极其嘈杂的市场数据中挖掘所有隐藏模式的方式。即使抛开噪音,信号本身的性质也在不断变化,以应对疫情、选举、新法规甚至情绪变化。这是一场永不停歇的游戏:试图判断你的旧模型是否仍能反映真实世界,如果不能,又该如何应对。如果你对这类工作感兴趣,Jane Street 正在招聘机器学习研究员和工程师。他们也在接受夏季机器学习实习项目的申请,地点包括伦敦、纽约和香港。如果你恰好在剧集发布后一周参加 GTC,Jane Street 的 GPU 性能团队将发表演讲。请访问 janestreet.com/swash 了解更多。

Their infrastructure team has built some of the biggest research clusters in the world with tens of thousands of high-end GPUs and hundreds of thousands of CPU cores and exabytes of storage. This compute is part of how Jane Street surfaces all the hidden patterns that are embedded in incredibly noisy market data. Even beyond the noise, the nature of the signal changes constantly in reaction to things like pandemics and elections and new regulations and even changes in sentiment. There's this unremitting game of trying to figure out whether your old models still reflect the real world and if not, what to do about it. If you're interested in working on this sort of thing, Jane Street is hiring ML researchers and engineers. They're also accepting applications for their summer ML internship program with spots in London, New York, and Hong Kong. And if you happen to find yourself at GTC, which is happening the week after this episode drops, Jane Street's GPU performance team is giving a talk. Go to janestreet.com/swash to learn more.

Elon的晶圆厂计划 Elon's Fab Plans

Host

我最近采访了 Elon,他的整个计划是,我猜他们要建一个千兆工厂、万亿工厂,10 的某次方,他们要建洁净室。我甚至不会问你关于脏房间的事,但假设他们建了洁净室。我有几个问题。第一:你认为这是 Elon 能比传统建设快得多的事情吗?这不是关于建造最终工具,只是关于建造设施本身。建一个洁净室并极快地完成有多复杂?如果今年或明年我们的瓶颈就在洁净室,Elon 这种快速行动的风格能快很多吗?第二:如果两年后你的观点是我们不再受洁净室空间限制,而是受工具限制,那这还重要吗?

I interviewed Elon recently and his whole plan is that I guess they're going to build this gigafab, terafab, some power of 10, and they're going to build the clean rooms. I won't even ask you about the dirty rooms thing, but like let's say they build the clean rooms. I have a couple questions. One: do you think this is the kind of thing that Elon could build much faster than people are conventionally building it? This is not about building the end tools. This is just about building the facility itself. How complicated is it to just build a clean room and do it extremely fast? Is this something that like Elon with this move fast thing could do much faster if that's what we're bottlenecked on this year or next year? And two: does that even matter if in two years your view is that we're not bottlenecked on clean room space but we're bottlenecked on the tooling?

Dylan Patel

所以我认为,就像任何复杂的供应链一样,这需要时间,而且瓶颈会随时间推移而变化。即使某件事不再是瓶颈,也不意味着市场不再有利润,对吧?例如,能源在未来几年不会成为大瓶颈。但这并不意味着能源增长不快、没有利润。它只是不是关键瓶颈。在晶圆厂领域,洁净室是今年和明年的最大瓶颈。而到了 2028、2029、2030 年,那里仍然会有约束。关于 Elon,我认为他拥有巨大的能力来获取物理资源和真正聪明的人来建造东西。他招募真正优秀的人的方式就是尝试建造最疯狂的东西,对吧?在 AI 领域,这不太奏效,因为每个人都在试图构建 AGI,每个人都很有野心。但在像我们要去火星、我们要制造能自己着陆的火箭、我们要制造全自动电动汽车、或者我们要制造人形机器人这些方面,对吧?这些方法是招募那些认为这是世界上最重要问题的人来研究这个问题,因为他是唯一真正努力的人。在半导体领域,我想建造一个每月生产 100 万片晶圆的工厂。没有人有这么大的工厂。这就是他说的,对吧?他想每月生产 100 万片晶圆。他有可能招募到很多非常优秀的人,让他们投入到这项英雄般的疯狂任务中,试图建造一个每月生产 100 万片晶圆的工厂。第一步是建造洁净室。我认为他可能能做到,对吧?我认为他有一种心态,比如删除东西。它可以脏一点。没关系。可能不对。实际上,我认为 100% 不对。你需要晶圆厂非常干净。我认为晶圆厂内的所有空气每 3 秒更换一次。就是这么快,而且颗粒物非常少。但我认为他可以建造洁净室。这需要一两年时间。也许一开始不会很快,但随着时间的推移,你会越来越快。但真正复杂的部分是开发工艺技术和制造晶圆。我不认为他能很快开发出来。我认为那需要很多积累的知识。这又是最复杂的集成,涉及非常昂贵的工具和供应链,由台积电、英特尔或三星完成。而其中两家公司甚至不那么出色,而且它们极其复杂。

So I think, as with any complex supply chain, it takes time and constraints shift over time. Even if something isn't a constraint anymore, that doesn't mean the market no longer has margin, right? So, for example, energy will not be a big bottleneck as we get to a couple years from now. But that doesn't mean energy is not growing super fast and there's no margin there. It's just not the key bottleneck. In the space of fabs, clean rooms are the biggest bottleneck this year and next year. And as we get to 2028, 2029, 2030, there will still be constraints there. The thing about Elon is I think he's had a tremendous capability to garner physical resources and really smart people to build things. And the way he's able to recruit really amazing people is just try and build the craziest stuff, right? In the case of AI, that's not really worked because everyone's trying to build AGI, everyone's very ambitious. But in the case of like we're going to go to Mars, we're going to make rockets that land themselves, or we're going to make fully autonomous cars that are electric, or we're going to make humanoid robots, right? These are methods of recruiting the people who think that's the most important problem in the world to work on that problem because he's the only one trying really hard. In the case of semiconductors, I want to make a fab that's a million wafers per month. No one has a fab that big. That's what he stated, right? He wants to make a million wafers a month. It's possible that he's able to recruit a lot of really awesome people and get them on this heroically crazy task of trying to build a fab that does a million wafers per month. Step one is to build the clean room. And I think that he probably can do it, right? I think there's some mindset, his mindset around like delete things. It can be dirty. It's fine. Probably not right. Actually, I think 100% it's not right. You need the fab to be very clean. I think the entire air in the fab gets replaced every 3 seconds. It's that fast and there's so few particles. But I think he can build the clean room. It'll take a year or two. Maybe initially it won't be super fast, but then over time you'll get faster and faster at it. But then the really complex part is actually developing the process technology and building wafers. And I don't think he can develop that quickly. I think that has a lot of built-up knowledge. It's again the most complicated integration of very expensive tools and supply chain that's done by TSMC or Intel or Samsung. And some of these two other companies aren't even that great and they're tremendously complex.

Host

如果到 2030 年,突然出现某种彻底颠覆,我们不再使用 EUV,而是使用效果更好、生产更简单、产量更大的东西,你会感到多惊讶?我确定作为业内人士,这听起来像是一个完全天真的问题,但你明白我在问什么吗?我们应该给那种完全出乎意料、让这一切都不再相关的东西分配多少概率?一种非常简单且易于扩展的东西。

How surprised would you be if in 2030 there just happened to be some total disruption where we're not using EUV, we're using something that has much better effects, much simpler to produce, we can produce in much bigger quantities. I'm sure as an industry insider that sounds like a totally naive question, but do you see what I'm asking? Like what probability should we put on something totally out of left field comes out and none of this is relevant? Something that's very simple and easy to scale.

Dylan Patel

我给出的概率非常非常低。有很多公司正在研究类似粒子加速器或同步加速器的东西,产生 13.5 nm 的 EUV 光,甚至 X 射线,甚至更窄的波长如 7 nm,然后用于光刻工具。但这些东西就像巨大的粒子加速器,然后产生这种光。建造起来非常复杂。所以有几家公司,我认为那可能对行业造成比 EUV 更大的颠覆。我不一定认为我们会神奇地建造出某种直接写入、超级简单且可以大规模制造的新东西,尽管有一些尝试在做类似的事情。

I have very, very low probability. There are a number of companies working on effectively like particle accelerators or synchrotrons that generate light that's either 13.5 nm like EUV or even X-ray, even narrower wavelength like 7 nm, whatever wavelengths of light to then use in lithography tools. But those things are like massive particle accelerators that are then generating this light. It's a very complicated thing to build. So there are a couple companies and I think that could be a big disruption to the industry beyond what EUV is. I don't necessarily think that we're going to just magically build something new that is direct write and super simple and can be manufactured at huge volumes, although there are some attempts to do things like this.

Host

是的,因为我问是因为如果你想想 Elon,过去火箭技术是那种被认为……我的意思是它极其复杂。

Yeah, because I asked because if you think about Elon, in the past rocketry was this thing that thought... I mean it is incredibly complicated.

Dylan Patel

听着,与 Elon 相比,我只是一个天真的空谈者,对吧?我建过什么?所以也许有可能,对吧?

Look, I'm just a naive yapper compared to Elon, right? What have I built? So maybe it's possible, right?

Host

是的,是的。为了将来能够制造更多内存,可以像做 3D NAND 那样构建 3D DRAM,然后回到 DUV。

Yeah, yeah. In order to be able to build more memory in the future, could build 3D DRAM the way we do 3D NAND and then go back to DUV.

Dylan Patel

这是希望所在。目前每个人对 3D DRAM 的路线图是仍然会使用 EUV,因为你想要更紧密的重叠。现在当你进行后续处理步骤时,你希望一切垂直堆叠,层数更多,间距更小,等等。

This is the hope. Currently everyone's roadmap for 3D DRAM is that you'll still use EUV because you want to have that tighter overlay. Now when you're doing these subsequent processing steps, you want it to be, you know, everything's vertically stacked, you have more layers on top of each other, and you want the pitches to be tighter and all these things.

3D RAM与EUV需求 3D RAM and EUV demand

Host

所以,总的来说,人们仍在尝试 EUV,但 3D 的作用是:单次 EUV 光刻能制造的比特数会大幅增加。这是希望所在,但目前每个人的路线图大致是:从当前的 6F² 单元到 4F² 单元,最终在本十年末或下个十年初实现 3D RAM。因此,还有大量的研发、制造和集成工作要做。我不认为这不可能。我认为它很有可能发生。但这也会要求晶圆厂进行大规模改造,对吧?晶圆厂中工具的构成非常不同。实际上,光刻工具是唯一没有太大区别的东西,但它们的数量相对于不同类型的化学气相沉积、原子层沉积、干法刻蚀或不同化学工艺的刻蚀腔体——所有这些。针对不同工艺节点,你有各种不同类型的工具。你不可能在短时间内将逻辑晶圆厂改造成 DRAM 晶圆厂,反之亦然。同样,现有的 DRAM 晶圆厂从 1b 或 one alpha 到 one beta 再到 one gamma 工艺节点就需要大量改造,因为现在它们必须添加 EUV,并改变使用 EUV 时的沉积和刻蚀化学材料堆栈,而且 EUV 工具必须到位。此外,当转向 3D RAM 时,变化会更大。因此,这些晶圆厂在工具方面需要进行大量改造。这将是一个巨大的颠覆,会使 EUV 需求总体降低。但随着时间的推移,EUV 需求占晶圆成本的比例一直在上升。最初,在 2014 年左右,光刻成本约占晶圆成本的 16-17%,在过去 15 年里上升到了 30%。对于 DRAM,它处于中低两位数,现在已趋向于高两位数。在 3D 到来之前,它可能会突破 20% 的范围,但如果 3D 实现,EUV 占总晶圆成本的比例会再次下降。

So, generally people are still trying to do EUV, but what 3D would do is it would take a single EUV pass and drastically increase the number of bits it can make. That is the hope, but right now everyone's roadmap is sort of: you go from current 6F² cell to 4F² cell, and then finally 3D RAM by the end of the decade or early next decade. So there's still a lot of R&D, manufacturing, and integration to be done. I wouldn't call that out of the cards. I think it's very much likely going to happen. It also is going to require a huge retooling of fabs, right? The breakdown of tools in a fab is very different. Actually, the lithography tool is the only thing that isn't that different, but the number of them relative to different types of chemical vapor deposition, atomic layer deposition, dry etch, or different kinds of etch chambers with different chemistries—all these things. You have all these different kinds of tools for different process nodes. You can't just convert a logic fab to a DRAM fab or vice versa in a short amount of time. And in the same way, existing DRAM fabs require a lot of retooling just to go from 1b or one alpha to one beta to one gamma process nodes because now they have to add EUV and change the chemistry stacks for deposition and etch when using EUV, and the EUV tool has to be there. Furthermore, when you change to 3D RAM, there's going to be an even larger shift. So there's a lot of retooling of these fabs that needs to happen in terms of the tools. That would be a big disruption that would make EUV demand generally lower. But as we've seen across time, EUV demand as a percentage of wafer cost has trended up. Initially, lithography in the 2014 era was like 16-17% of wafer cost, and it's gone to 30% over the last 15 years. For DRAM, it was in the mid-teens or low teens, and now it's trended towards the high teens. Before we get to 3D, it'll likely cross into the 20s percentage range, but then if we get to 3D, it tanks again in terms of the total wafer cost as a percentage of EUV.

Host

是的,我想你更关心的是瓶颈有多大,而不是成本占比。但成本占比某种程度上是一个代理指标。对。所以,如果你是 Jensen 或 Sam Altman,或者任何从 AI 算力 Scaling 中获益巨大的人,有传闻说他们会去台积电问:“嘿,为什么我们不能做 X、Y、Z?”但我认为你这里的观点是,台积电做什么在某种意义上并不重要。事实上,即使英特尔和三星长期内建造更多晶圆厂,你也会被 ASML 和其他工具及材料制造商卡住。所以首先,这个理解对吗?其次,那么硅谷的人为什么应该去荷兰,试图说服 ASML 制造更多工具,以便在 2030 年获得更多 AI 算力?

Yeah, I guess you care less about the percent of cost and more about how much it bottlenecks. But the percentage of cost is sort of a proxy. Yeah. So if you're Jensen or Sam Altman or whoever stands to gain a lot from scaling up AI compute, there are these stories that they'd go to TSMC and say, 'Hey, why can't we do X, Y, and Z?' But I think the point you're making here is it doesn't really matter in some sense what TSMC does. In fact, even if Intel and Samsung build more foundries in the long run, you're going to be bottlenecked by ASML and other tool makers and other material makers. So first, is that correct interpretation? And second, then why should basically Silicon Valley people be going to the Netherlands to try to pitch ASML to make more tools so that in 2030 they can have more AI compute?

Dylan Patel

你知道,这是一个有趣的动态。我们在 2023、2024 和 2025 年看到,那些比其他人更早看到能源瓶颈的人,不对称地去找西门子、三菱和 GE Vernova,买断了涡轮机产能。现在他们能够因为能源问题而高价部署这些涡轮机。同样的事情也可以发生在 EUV 上,只不过 ASML 不会随便相信任何想买 EUV 工具的傻瓜。这些涡轮机比 EUV 工具便宜得多,而且产量也大得多,尤其是工业燃气轮机——不仅仅是联合循环,还有更便宜、更小、效率较低的型号。人们为这些支付定金。所以从某种意义上说,有人可以这样做,对吧?有人应该去荷兰说:“我给你十亿美元。你给我两年后购买 10 台 EUV 工具的权利。我两年后排在第一位。”然后在两年内,你等着所有人都意识到:“哦,糟了,我没有足够的 EUV 工具。”然后你试图以溢价出售你的期权。但你实际上只是在说:“ASML,你真笨。你在这上面赚的利润不够多。我来赚这个利润。”问题是,ASML 会同意吗?我不这么认为。

You know, it's a funny dynamic. We saw in 2023, 2024, and 2025, people who saw the energy bottleneck before others asymmetrically went to Siemens, Mitsubishi, and GE Vernova and bought up turbine capacity. Now they're able to charge excess amounts for deploying these turbines because of energy. In the same sense, this could be done for EUV, except ASML is not just going to trust any random bozo who wants to buy EUV tools. These turbines are much cheaper than EUV tools, and there are many more of them produced, especially once you get to industrial gas turbines—not just combined cycle, but the cheaper, smaller, less efficient ones. People put down deposits for these. So in a sense, someone could do this, right? Someone should go to the Netherlands and be like, 'I'll pay you a billion dollars. You give me the right to purchase 10 EUV tools two years from now. I'm first in line two years from now.' Then over those two years, you go around and wait for everyone to realize, 'Oh crap, I don't have enough EUV tools.' Then you try to sell your option at some premium. But all you're effectively doing is saying, 'ASML, you're dumb. You weren't making enough margin on these. I'm going to make a margin.' And the question is, will ASML even agree to this? I don't think so.

Host

但存在一种可能,他们至少能从那里得到需求信号,从而增加产量。

But there's a world where they at least get the demand signal from that to increase production, potentially.

Dylan Patel

有可能。我同意。但这听起来像你在说:“哦,他们即使想增加产量也做不到。”这正是那种市场:如果他们无法增加产量——就像台积电无法那么快增加产量,而需求却在飙升——那么显而易见的解决方案就是套利,因为你我都知道需求远高于他们的预测和建造能力。所以你可以通过锁定产能并签订远期合约来套利,然后在其他人意识到我们没有足够产能时再出售。这样你就能获得 ASML 和台积电本应收取的巨额利润。但问题是,我不知道 ASML 和台积电是否会同意这样做。

Potentially. I agree. But it sounds like you're saying, 'Oh, they couldn't even increase production if they wanted to.' That's exactly the market in which if they can't increase production—just like TSMC cannot increase production that fast, and yet demand is mooning—then the obvious solution is to arbitrage this because you and I know demand is way higher than their projections and their capability to build. So you arbitrage this by locking up the capacity and doing a forward contract, then trying to sell it at a later date once other people realize we don't have enough capacity. Then you'll have this insane margin that ASML and TSMC should have been charging. But the thing is, I don't know if ASML and TSMC will ever agree to this.

Host

好的,现在让我问问电力的问题。听起来你认为电力可以任意扩展。不是任意,但可以超过这些数字。如果我没记错的话,你关于电力的博文《How I Love Increasing Power》中暗示,GE Vernova、三菱和西门子每年能生产 60 吉瓦的燃气轮机,还有其他来源,但不如涡轮机重要,而且其中只有一部分用于 AI,我猜。那么,如果到 2030 年我们有足够的逻辑和内存来支持每年 200 吉瓦,你认为这些东西会 ramp up 到每年超过 200 吉瓦吗,还是你怎么看?

Okay, let me ask about power now. So it sounds like you think power can be arbitrarily scaled. Not arbitrarily, but yes, beyond these numbers. I think if I'm remembering correctly, your blog post on power—'How I Love Increasing Power'—you were implying that GE Vernova, Mitsubishi, and Siemens could produce gas turbines at 60 gigawatts a year, and then there are other sources but they're less significant than the turbines, and only a fraction of that goes to AI, I assume. So if in 2030 we have enough logic and memory to do 200 gigawatts a year, do you just think that these things are on a path to ramp up to more than 200 gigawatts a year, or what do you see?

Dylan Patel

是的。我的意思是,现在我们处于 30 吉瓦,对吧?或者 20-20-20。顺便说一句,这是关键 IT 容量,对吧?这一点很重要。当我说这些吉瓦时,我指的是关键 IT 容量——服务器插电。那是它消耗的功率。但整个链条上有损耗,对吧?传输有损耗。

Yeah. So I mean, right now we're at 30, right? Or 20-20-20. This is critical IT capacity by the way, right? This is an important thing to mention. When I'm talking about these gigawatts, I'm talking about critical IT capacity—server plugged in. That's how much power it pulls. But there are losses along the chain, right? There is loss on the transmission.

发电挑战与解决方案 Power Generation Challenges and Solutions

Dylan Patel

而且还有转换损耗、冷却损耗等等。所以你应该把这个系数放大,从今年的 20 吉瓦或到本十年末的 200 吉瓦,再提高 20% 到 30%。然后还有容量因子,对吧?涡轮机不会 100% 运行。事实上,如果你看 PJM,我认为是美国最大的电网,覆盖中西部、东北部地区,但不是整个东北部——总之,PJM 在他们的模型中估算,比如涡轮机,我们想要多少超额容量,大概是 20% 的容量。此外,在这 20% 的超额容量中,我们让所有涡轮机以 90% 的容量运行,因为出于可靠性考虑它们被降额了,比如设备故障、维护等等。所以实际上,能源的铭牌容量总是远高于最终的关键 IT 容量,因为所有这些因素。

And there's losses on the conversion, there's losses on cooling, etc. And so you should gross this factor up, you know, from 20 gigawatts for this year or 200 gigawatts by the end of the decade, to some number 20-30% higher. And then you have capacity factors, right? Turbines don't run at 100%. In fact, if you look at PJM, which is the largest grid I think in America, sort of the Midwest, sort of Northeast kind of area, not the full Northeast, but anyways, PJM, they rate in their models for like, hey, turbines, how much capacity we want to have excess, roughly 20% capacity. In addition, in that 20% excess capacity, we're running all the turbines at 90% because they are derated some for reliability, things go down, maintenance, etc. So in reality, the nameplate capacity for energy is always way higher than the actual end critical IT capacity because of all these factors.

Host

嗯,但不仅仅是涡轮机,对吧?

Um, but it's not just turbines, right?

Dylan Patel

如果你只是用涡轮机发电,那很简单、无聊、容易,对吧?我们是人类,资本主义要高效得多。所以那篇博客的重点是,是的,只有三家公司制造联合循环燃气轮机,但我们可以做的还有很多,对吧?我们可以用航改型燃气轮机,对吧?我们可以把飞机发动机改造成涡轮机。市场上甚至还有新进入者,比如 Boom Supersonic 正在尝试这样做,对吧?他们正在与 Kratos 合作,而且市场上还有其他已有的方案。还有中速往复式发动机,对吧?就是旋转的发动机,有点像任何柴油发动机。大约有 10 家公司这样制造发动机,对吧?比如康明斯,你知道,至少我来自佐治亚州,我们那里的人过去常说,“哦,伙计,你装了一台康明斯发动机。”比如在 Ram 卡车上,但实际汽车制造业正在下滑;这些公司都有产能,可以扩大规模并转化为数据中心电力,对吧?装上所有这些往复式发动机。是的,它不像联合循环那么清洁,如果你愿意,也许可以把它们从柴油改为燃气。但归根结底,这些旋转发动机——哦,那船用发动机呢,对吧?所有这些用于大型货船的发动机都很棒。Nebus 正在为微软在新泽西的一个数据中心做这件事,对吧?他们用这些船用发动机发电。哦,还有 Bloom Energy 做燃料电池。我们已经看好他们大约一年半了,因为他们有很强的增产能力,而且增产的投资回收期非常快。即使成本比联合循环(成本和效率最佳)略高一些。然后还有太阳能加电池,随着成本曲线持续下降,这些也可以上线。还有风能,当然还有降额——你知道,当你安装风力涡轮机时,你可能会说,“哦,我只期望最大功率的 15%,因为东西会波动,”但你可以加电池,有所有这些办法。另一件事是,电网的规模是为了——我们不会在峰值用电时断电,比如夏天最热的那天。但实际上,那是一个比平均负荷高 10%、15%、20% 的负荷尖峰。如果你只安装足够的公用事业规模电池,或者只运行一小部分年份的调峰电厂,那么突然之间——那些可以是燃气、工业燃气轮机、联合循环,或者我提到的任何其他电源。它们可以是电池。然后突然之间,你就为数据中心解锁了美国电网 20% 的容量,因为大多数时候这些容量是闲置的,它们真的只是为了那个峰值,对吧?也就是一两天,对吧?全年只有几天的那几个小时是峰值。所以你只要有足够的容量来吸收那个峰值负荷,突然之间你就把一切都转移了。如今数据中心只占美国电网电力的 3-4%,到 28 年将达到 10%。但如果你能像这样解锁美国电网 20% 的容量,那并不疯狂。美国电网是太瓦级别的,而不是几百吉瓦级别,对吧?所以我们可以增加很多能源。这并不容易。我不是说它容易。这些事情不会很难。有很多艰苦的工程。人们必须承担很多风险。人们必须使用很多新技术,但埃隆是第一个做表后燃气的人。从那以后,我们看到人们为了获取电力而做的各种事情呈爆炸式增长,它们并不容易,但人们将能够做到,而且供应链比芯片简单得多。

If you were just making power from turbines, that's simple, boring, easy, right? We're humans and capitalism is far more effective. So the whole point of that blog was yes, there's only three people making combined cycle gas turbines, but there's so much more we can do, right? We can do aeroderivatives, right? We can take airplane engines and turn them into turbines as well. And there's even new entrants in the market like Boom Supersonic trying to do that, right? And they're working with Kratos, and also there's all the other ones that already exist in the market. There's medium-speed reciprocating engines, right? Engines that spin in circles, right? So sort of like any diesel engine, right? There's like 10 people who make engines that way, right? So Cummins, you know, at least I'm from Georgia and we, you know, people used to be like, 'Oh man, you got a Cummins engine in there.' You know, like regarding Ram trucks, but it's like, well actually, auto manufacturing is going down; these companies all have capacity and could scale and convert that to data center power, right? Stick all these reciprocating engines. Yes, it's not as clean as combined cycle, maybe you can convert them from diesel to gas if you want. But at the end of the day, these spinning engines—oh, what about ship engines, right? All these engines for these massive cargo ships, those are great. Nebus is doing that for a data center in Microsoft in New Jersey, right? They're running these ship engines to generate power. Oh, there's Bloom Energy doing fuel cells. We've been very positive on them for like a year and a half now because they have such a capability to increase their production, and their payback period for production increase is very fast. Even if the cost is a little bit higher than combined cycle, which is the best cost and efficiency. And then there's solar plus battery, which as these cost curves continue to come down, those can come online. There's wind, and of course the derating of those—you know, when you put on a wind turbine, you might say, 'Oh, I'm only going to expect 15% of the maximum power because things just oscillate,' but you add batteries, there's all these things. And the other thing is that the grid is scaled for, you know, we are not going to cut off power at peak usage, which is like the hottest day in the summer. But in reality, that's a load spike that is 10, 15, 20% higher than the average. Well, if you just put enough utility-scale batteries or you put peaker plants that only run a small portion of the year, then all of a sudden—and those could be gas, they could be industrial gas turbines, they could be combined cycle, they could be any of the other sources of power I mentioned. They could be batteries. Then all of a sudden, you've unlocked 20% of the US grid for data centers because most of the time that capacity is sitting idle and it's really only there for that peak, right? Which is a day or two, right? And it's a few hours of maybe a few days of the full year is that peak. And so you just have enough capacity to absorb that peak load and all of a sudden you've transferred it all. And today data centers only use 3-4% of the power of the US grid, and by '28 they'll be 10%. But if you can just unlock 20% of the US grid like this, it's like not that crazy. The US grid is terawatt level, not hundreds of gigawatts level, right? So we can add a lot more energy. It's not easy. I'm not saying it's easy. These things aren't going to be hard. There's a lot of hard engineering. There's a lot of risks that people have to take. There's a lot of new technologies people have to use, but Elon was the first to do this behind-the-meter gas. And since then we've seen an explosion of different things that people are doing to get power, and they're not easy, but people are going to be able to do them, and the supply chains are just way more simple than chips.

Host

有意思。所以,我猜他在采访中提出,他看的那台特定涡轮机的特定叶片,交货时间已经排到 2030 年以后了。而你的观点是……

Interesting. So, I guess he made the point during the interview that the specific blade for the specific turbine he was looking at, the lead times for that go out beyond 2030. And your point is that...

Dylan Patel

那很好。还有很多其他方式可以产生能源。只要效率低一点就行。没关系。

That's great. There's so many other ways to make energy. Just be inefficient. It's fine.

Host

对吧?所以你现在——我猜联合循环燃气轮机的资本支出是每千瓦 1500 美元,而你说你可以——要么采用比这贵得多的技术,要么其他东西变得足够便宜,使其具有竞争力。

Right? So you're like right now I guess combined cycle gas turbines have capex of $1,500 per kilowatt and you're saying you could just—it would make sense to have either technologies that are much more expensive than that or other things are getting cheap enough to make it competitive.

Dylan Patel

完全正确。完全正确。你知道,它甚至可能高达每千瓦 3500 美元。对吧。所以可能是联合循环成本的两倍,而 GPU 的总成本,在总拥有成本基础上,每小时只上涨了几美分,对吧?再说一次,如果我们——因为我们一直在讨论 Hopper 定价 140 美元,现在变成了,哦,电价翻倍,好吧,原来 140 美元的 Hopper 现在成本是 150 美元。这就像,哦,我不在乎,因为模型改进得太快了,它们的边际效用远高于那 10 美分的能源成本增加。

Exactly. Exactly. You know, it can be as high as $3,500 per kilowatt even. Right. So it could be twice as much as the cost of combined cycle, and the total cost of the GPU, you know, on a TCO basis has gone up a few cents per hour, right? Again, if we're—because we've been talking about Hopper pricing $140, now becomes, you know, oh, the power price doubles, okay, the Hopper that was $140 is now $150 in cost. It's like, oh, I don't care because the models are improving so fast that the marginal utility of them is worth way more than that 10-cent increase in energy.

Host

好的。然后你说电网的 20%——那么冬天呢——其中 20% 可以通过增加公用事业规模电池直接上线?你愿意投入多少……

Okay. And then you're saying 20% of the grid—so winter, what about—20% of that can just come online from utility scale batteries increasing? What you would be comfortable putting...

Dylan Patel

顺便说一句,那里的监管机制并不容易。

Regulatory mechanism there is not easy, by the way.

Host

但那是 200 吉瓦,假设这种情况发生。但你说仅从你提到的不同燃气发电来源——不同种类的发动机和涡轮机——加起来,到本十年末它们能解锁多少吉瓦?

But like that's 200 gigawatts, like if that hypothetically happens. But you're saying on just from the different sources of gas generation you mentioned—the different kinds of engines and turbines—combined, how many gigawatts could they unlock by the end of the decade?

Dylan Patel

是的。

Yeah.

发电与表后容量 Power generation and behind-the-meter capacity

Dylan Patel

根据我们追踪的数据,仅天然气发电设备就有超过 16 家制造商。是的,联合循环只有三家涡轮机制造商,但我们追踪了 16 家不同的供应商,并掌握了他们所有的订单。结果发现,各类数据中心的订单总量高达数百吉瓦。到本十年末,我们认为新增容量中约有一半将采用表后模式。我们观察发现,表后模式几乎总是比并网更贵,但并网面临诸多问题,比如许可、互联排队等等。因此,尽管成本更高,人们还是选择表后模式。表后模式采用的技术五花八门:往复式发动机、船用发动机、衍生型号、联合循环(虽然联合循环不太适合表后)、Bloom Energy 燃料电池、太阳能加电池,等等。

So we're tracking in some of our data where there are over 16 different manufacturers of power generating things just from gas alone. Yes, there are only three turbine manufacturers for combined cycle, but we're tracking 16 different vendors and we have all of their orders. It turns out there are just hundreds of gigawatts of orders to various data centers. As we get to the end of the decade, we think something like half of the capacity being added will be behind the meter. When we look at a lot of this, behind the meter is almost always more expensive than grid connected, but there are just a lot of problems with getting grid connected, permits, interconnection queues, and all that sort of stuff. So even though it's more expensive, people are doing behind the meter. What they're doing behind the meter ranges widely: reciprocating engines, ship engines, derivatives, combined cycle (though combined cycle is not that great for behind the meter), Bloom Energy fuel cells, solar plus battery, any of these things.

Host

你是说这些技术中每一种都能达到几十吉瓦的规模?

You're saying any of these individually could do tens of gigawatts.

Dylan Patel

每一种都能达到几十吉瓦,而整体上将达到数百吉瓦。

Any of these individually will do tens of gigawatts, and in the whole they will do hundreds of gigawatts.

Host

仅此一项就应该足以……我是说,电工工资可能会再翻一倍或两倍。会有很多新人进入这个领域,很多人会赚钱。但我不认为这是主要瓶颈。

So that alone should more than... I mean, it's going to take... electrician wages probably double or triple again. There's going to be a lot of new people entering that field and a ton of people who make money. But I don't see that as the main bottleneck.

劳动力限制与模块化 Labor constraints and modularization

Host

目前,Crusoe 为 OpenAI 在 Abene 建造的 1.2 吉瓦数据中心,高峰期有 5000 名工人。如果扩大到 100 吉瓦,虽然效率会提高,但大概需要 40 万人来建造。想想美国的劳动力,有多少电工?有多少建筑工人?我猜大约有 80 万电工。我不确定他们是否都能以这种方式替代。建筑工人有数百万。但如果我们每年新增 200 吉瓦,最终会不会面临劳动力短缺?还是说这其实不是真正的制约因素?

Right now in Abene, the 1.2 gigawatt data center that Crusoe is building for OpenAI, I think they had 5,000 people working there at peak. If you scale that to 100 gigawatts, and I'm sure things will get more efficient over time, that would be like 400,000 people to build 100 gigawatts. If you think about the US labor force, how many electricians are there? How many construction workers? I guess there are about 800,000 electricians. I don't know if they're all substitutable in this way. There are millions of construction workers. But if we're in a world where we're adding 200 gigawatts a year, are we going to be crunched on labor eventually, or do you think that is not a real constraint?

Dylan Patel

劳动力是巨大的制约因素。人们需要培训。同样,我们可能会开始以这种方式引进高技能劳动力,因为现在有理由让欧洲那些曾从事电厂退役工作的高技能电工来到美国,建设数据中心的高压电力系统。人形机器人或机器人技术可能会开始发挥作用,但减少用工人数的主要途径是模块化,并在亚洲的工厂生产——虽然对美国来说不太理想,但韩国、东南亚,以及许多方面也包括中国。这些地区将越来越多地运输预制好的数据中心模块。也许现在你运输的是服务器或机架,然后连接到来自不同地方的部件,但未来你会把它们运到工厂,整合成完整的模块。可能是一个 2 兆瓦的模块,从高压电直接转换为机架所需的电压和直流电,而不是交流高压电。或者冷却系统:你运输一个集成了大量冷却子系统的完整单元,因为水管工也是很大的制约因素。此外,不再是一个个机架由工人布线,而是将一整排服务器放在一个滑橇上,从工厂直接运出。目前单个机架可能为 120-140 千瓦,但到了下一代 Nvidia Kyber 等产品,单个机架接近 1 兆瓦。如果是一整排,机架、网络、冷却和电源机架都集成在一起。这样,现场需要布线的部分大大减少,无论是光纤网络、电力还是管道。这能大幅减少数据中心所需的工人数量,从而大幅提升建设能力。在此过程中,有些人适应新事物更快,有些人更慢。谷歌、Meta 等公司一直在大力提倡这种模块化。其他公司则会慢一些。归根结底,适应快的人可能会遇到更多延误,而适应慢的人则会面临劳动力问题。市场总会出现错位,因为这是一个非常复杂的供应链,但最终它足够简单,我们能够通过资本主义和人类的聪明才智在所需的时间尺度上解决。

Labor is a humongous constraint in this. People have to be trained. Likewise, we probably start importing the highest skilled labor in this way, because now it makes sense that a really high-skilled electrician in Europe who was working on decommissioning power plants now comes to America and is building data center high voltage electricity, power moving across the data center. Humanoid robots or robotics might start to help, but the main factor for reducing the number of people is modularizing things and making them in factories in Asia, unfortunately, but at least for America, Korea, Southeast Asia, in many ways China as well. These areas are going to ship more and more built-out sections of the data center. Maybe today you ship servers or a rack in and then plug that into different pieces from different places, but now you'll ship it to a factory and integrate the entire thing. Maybe this is a 2 megawatt block that goes from high voltage power to the voltage and DC that you deliver to the rack, instead of AC and high voltage. Or cooling: you ship a fully integrated thing that has a lot of the cooling subsystems already put together, because plumbers are also a big constraint. Furthermore, instead of a single rack with people wiring up all these racks, you take a skid and put an entire row of servers that is shipped from the factories. Today a single rack may be 120-140 kilowatts, but as we get to next generation Nvidia Kyber and things like that, it's almost a megawatt. In addition, if you do an entire row, it'll have the rack, networking, cooling, and power racks all integrated together. So now when you come in, you have much less stuff to cable, whether it be networking with fiber, power, or plumbing. This drastically reduces the number of people working in data centers, and therefore the capability to build these will be much larger. Along the way, some people move faster to new things, some move slower. Google, Meta, and many others have been talking a lot about this modularization. Others are going to be slower to do it. At the end of the day, people who move faster may have more delays, or people who are slower have labor problems. There will always be dislocations in the market because this is a very complex supply chain, but it's still simple enough that we will be able to solve it through capitalism and human ingenuity on the time scales that are required.

太空GPU与许可挑战 Space GPUs and permitting challenges

Host

说到需要解决的大问题,埃隆·马斯克非常看好太空 GPU。如果你说得对,地球上的电力不是制约因素,那么我想另一个合理的理由是,尽管你可以在物理上获得足够的燃气轮机等设备在地球上建造,但埃隆的下一个论点是,你无法获得许可在地球上建造数百吉瓦的设施。你认同这个论点吗?

Speaking of big problems to solve, Elon Musk is very bullish on space GPUs. If you're right that power is not a constraint on Earth, I guess the other reason they would make sense is that even though you can physically get enough gas turbines or whatever to build it on Earth, I think Elon's next argument is that you can't get the permitting to build hundreds of gigawatts on Earth. Do you buy that argument?

Dylan Patel

从土地角度看,美国很大,数据中心占不了多少空间。这可以解决。从许可角度看,空气污染许可是个挑战,但特朗普政府让事情变得容易多了。你去德克萨斯州,可以跳过很多繁琐手续。埃隆在孟菲斯和跨境建电厂时不得不处理很多复杂问题。

Landwise, America is big, data centers don't take that much space. You can solve that. Permitting wise, air pollution permits are a challenge, but the Trump administration made it much easier. You go to Texas and you can skip a lot of this red tape. Elon had to deal with a lot of complex stuff in Memphis and then building a power plant across the border.

数据中心选址与许可 Data Center Location and Permitting

Host

但既然埃隆住在得克萨斯,他为什么不直接去得克萨斯呢?

But why, given that Elon lives in Texas, why didn't he just go to Texas?

Dylan Patel

我认为部分原因是他们暂时过度依赖了电网电力,对吧?因为他们当时认为需要更多电力。

I think it was partially like they overindexed on grid power for a temporary period of time, right? Because that's just what they thought they needed more of.

Host

他们说过那里有一个连接到电网的铝冶炼厂。

They said an aluminum refinery connected to the grid there.

Dylan Patel

不,那是一个闲置的家电工厂。

No, it was an appliance factory that was idled.

Dylan Patel

但我认为他们可能更看重电网电力,也可能更看重水资源和天然气接入,因为实际上我认为他们买下那里时就知道天然气管道就在旁边,打算接入。水资源也一样。那是一系列不同的约束条件。那里可能更容易找到电工之类的人。但说到底,我不太确定他们为什么选择那个地点。我打赌如果埃隆能重来,他会选择得克萨斯州的某个地方。但没错,因为他面临的监管挑战。最终,许可审批是个难题,但美国很大,有 50 个州,事情总能办成。有很多小辖区,你可以把需要的工人全部运过去,临时工作 6 个月到 1 年,根据承包商类型,甚至可能只要 3 个月,给他们安排临时住房,支付高额费用,因为劳动力相对于 GPU、网络以及最终产生的 token 价值来说非常便宜。所以所有这些都有足够的空间来支付。而且现在人们也在多元化:澳大利亚、马来西亚、印度尼西亚、印度,这些地方的数据中心建设速度要快得多。但目前仍有 70% 以上的 AI 数据中心在美国,这一趋势仍在继续。所以我认为人们正在摸索如何建设这些东西,而许可审批和繁琐手续在得克萨斯、怀俄明或新墨西哥的偏远地区,可能比把东西送入太空要容易得多。

But I think they may have indexed more to what was grid power. They may have indexed more to like water access and gas access because actually I think they bought that knowing that the gas line was right there and they were going to tap it. Same with water. It was a whole host of different constraints. It was probably an area where electricians and things like that were easier to find. But at the end of the day, I'm not exactly sure why they chose that site. I bet Elon would have chosen somewhere in Texas if he could have gone back. But yeah, because of the regulatory challenges he's faced. Ultimately, permitting is a challenge, but America is a big place and there are 50 states and things will get done. There are a lot of small jurisdictions where you can just transport in all the workers that you need for a temporary period of 6 months to a year, depending on the type of contractor, it can be even 3 months, and put them in temporary housing, pay out the butt because labor is very cheap relative to the GPUs and the networking and so on and so forth and the end value of the tokens it's going to produce. So all of these things have plenty of room to be paid for. And also people are diversifying now: Australia, Malaysia, Indonesia, India, these are all places where data centers are going up at a much faster pace. But currently still 70% plus of the AI data centers are in America and that continues to be the trend. And so I think people are figuring out how to build these things and permitting like ultimately permitting and red tape in middle of nowhere Texas or middle of nowhere Wyoming or middle of nowhere New Mexico is probably a hell of a lot easier than sending stuff into space.

对太空数据中心的怀疑 Skepticism About Space Data Centers

Host

除了考虑到能源只占数据中心总拥有成本的一小部分,经济论证就不那么合理之外,你还有其他怀疑的理由吗?

Well, other than the fact that the economic argument makes less sense once you consider the fact that energy is a small fraction of the cost of ownership of a data center, what are the other reasons you're skeptical?

Dylan Patel

是的。显然,太空中的电力基本上是免费的。

Yeah. So, obviously power is free in space basically.

Host

不,那正是这么做的理由。

No, that's the reason to do it.

Dylan Patel

是的,那是理由。但还有其他反对论点,对吧?因为即使电力成本翻倍,它仍然只占 GPU 总成本的一小部分。主要的挑战是我们看到的分散问题。我们有 Cluster Max 来评估所有的新兴云服务商,我们测试了超过 40 家云公司,包括超大规模云和新兴云。除了软件之外,这些云服务商最大的区别在于它们部署和管理故障的能力。GPU 的可靠性极差。即使在今天,大约 15% 的 Blackwell 在部署后需要退货授权(RMA)。你必须把它们取出来,也许只是重新插拔,但有时你必须把它们运回英伟达或它们的合作伙伴那里进行 RMA。

Yeah, that's the reason to do it. But then there are all the other counter arguments, right? Which is because even if power costs double, you're still at a fraction of the total cost of the GPU. The main challenges is what we've seen that disperses. We have Cluster Max which rates all the neoclouds and we test them. We test over 40 cloud companies including the hyperscalers and neoclouds. What differentiates some of these clouds the most outside of software is their ability to deploy and manage failure. GPUs are horrendously unreliable. Even today, 15% of Blackwells or so that get deployed have to be RMA. You have to take them out, you know, maybe just plug them and plug them back in, but sometimes you have to take them, ship them to Nvidia or rather their partners who do these RMAs.

Host

你怎么看埃隆的那种说法,即一旦过了初始阶段,它们实际上不会出那么多故障?

What do you make of Elon's kind of argument that once you have the initial phase, they actually don't fail that much?

Dylan Patel

当然。但现在你做了这些:你测试了所有 GPU,拆解它们,装上飞船,送入太空,然后再重新上线。这需要几个月,对吧?如果你的论点是 GPU 有 X 年的使用寿命,比如 5 年,而这个过程多花了 3 个月,可能是 6 个月,就算 6 个月,那就是你集群使用寿命的 10%。而且由于我们严重受限于算力,理论上这些算力在最初 6 个月最有价值,因为现在比未来更受限制,现在的算力可以贡献给未来的更好模型,或者立即产生收入,用来筹集更多资金。现在永远是最重要的时刻。所以你可能会将算力部署延迟 6 个月。而这些云服务商的区别在于,我们看到有些云在地球上部署 GPU 需要 6 个月,而有些远少于 6 个月。那么问题来了,太空方案如何融入其中?我看不出你在地球上测试所有 GPU、拆解、运输、发射到太空,然后还能比直接放在测试地点更快。

Sure. But now you've done this, you've tested them all, you deconstructed them, put them on a spaceship, put them into space, and then put them online again. That's months, right? And if your argument is that GPUs have a useful life of X years, right? If a GPU has a useful life of 5 years and it takes three additional months, probably six, let's say six additional months, then that is 10% of your cluster's useful life. And because we're so capacity constrained, that compute is most valuable theoretically in the first 6 months you have it because we're more constrained now than in the future because that compute now can contribute to a better model in the future or contribute to revenue now which you can use to raise more money. Now is always the most important moment. And so you've delayed your compute deployment by 6 months potentially. And the thing that separates these clouds is we see clouds that take six months to deploy GPUs today on Earth, right? We see clouds that take a lot less than six months. And so the question is where does space get in there? I don't see how you would test them all on Earth, deconstruct them and ship them and shoot them into space and it not take longer than just putting them in the spot that you were testing them.

太空通信拓扑 Space Communication Topology

Host

所以我想问的问题是太空通信的拓扑结构。目前星链卫星之间的通信速度是 100 Gbps,你可以想象通过优化的光学星间激光链路,速度会高得多。这实际上已经非常接近 InfiniBand 的带宽了,后者大约是 400 GB/s。

So the question I wanted to ask is the topology of space communication. So right now Starlink satellites talk to each other at 100 gigabits per second and you can imagine that being much higher with optical inter-satellite laser links that are optimized for this. And that actually ends up being quite close to the InfiniBand bandwidth which is like 400 gigabytes a second.

Dylan Patel

但那是每 GPU 的带宽,不是每机架的。

But that's per GPU not per rack.

Host

我明白了。好的。

I see. Okay.

Dylan Patel

所以还要乘以 72,就像 Hopper 那样。到了 Blackwell 和 Rubin,带宽会翻倍再翻倍。

So multiply that by 72 also like that was Hopper. When you go to Blackwell and Rubin, that doubles and doubles again.

Host

好吧。但在推理过程中,有多少计算是在单个 scale-up 域内完成的?不同的 scale-up 域还在协同工作吗,还是只是在一个 scale-up 域内进行批处理?

All right. But how much compute is happening per like during inference? Are the different scaleups still working together or is it just happening it's a batch within a single scaleup?

Dylan Patel

很多模型可以放在一个 scale-up 域内,但很多时候你会将它们拆分到多个 scale-up 域。我认为,随着模型变得越来越稀疏(至少这是总体趋势),你希望每个 GPU 只访问少数几个专家。而如果当今的领先模型有数百甚至数千个专家,那么你就需要在数百或数千个芯片上运行。即使我们继续向前发展,最终你会遇到这个问题:现在你还需要在通信方面将所有卫星连接起来。

A lot of models fit within one scaleup domain but many times you split them across multiple scaleup domains. I think that you really have to, as models become more and more sparse, at least this is like the general trend, then you want to ping just a couple experts per GPU. And if leading models today have hundreds if not thousand experts, then you want to run this across hundreds of chips or thousands of chips. Even as we continue to advance into the future, and so then you end up with this problem of well now you need to connect all these satellites together communication-wise as well.

太空数据中心挑战 Space data center challenges

Host

所以这会很困难,因为我曾想象,如果存在一个世界,你可以在单个 scale-up 上对一批数据进行批量推理,那也许更可行,但如果不是这样,那么……

So that would be tough because I was imagining if there's a world where you could do batch inference for a batch on a single scale-up, then maybe it's more plausible, but if not, then it's...

Dylan Patel

将这些卫星联网是个问题,而且你不能把卫星造得无限大,对吧?物理上有很多挑战让卫星变得非常大。所以你需要卫星之间的互连。这些互连比集群中的更贵——大约 20% 或 15% 的成本是网络。突然间,你就要用太空激光器,而不是那种以百万计产量制造、带有可插拔收发器的简单激光器。而且这些东西也很不可靠——顺便说一句,在集群的整个生命周期中,它们比 GPU 更不可靠。你得一直拔插、清洁,无缘无故地拔插。这些东西就是没那么可靠。所以你还有这个问题:你得用更昂贵、更复杂的太空激光器来通信,而不是那种超高产量的可插拔光收发器。

Networking these satellites together is a problem, and you can't just make the satellite infinitely large, right? There are a lot of challenges with physics to making a satellite really big. So then you need these interconnects between the satellites. Those interconnects are more expensive than in a cluster—like 20% of the cost or 15% of the cost is networking. All of a sudden, now you're making it like space lasers instead of pretty simple lasers that are manufactured in millions of volumes with pluggable transceivers. And those things are very unreliable as well—more unreliable than the GPUs, by the way, across the life of a cluster. You have to unplug, clean it all the time, unplug, replug it just for random reasons. These things are just not as reliable. So you've got that problem as well: you've got a more expensive, complicated space laser to communicate instead of this pluggable optical transceiver that's been in super high volume.

Host

好的。那么,总的来说,这对太空数据中心意味着什么?

Okay. So, all in all, what does that imply for space data centers?

Dylan Patel

所以,太空数据中心实际上并不受限于“我们拥有能源优势”这一点。它实际上只受限于同样的竞争资源。到本十年末,我们每年只能制造 200 吉瓦的芯片。那么,我们要如何获得那部分产能呢?无论是在陆地还是太空,这都不重要。你可以建造那种电力。我认为人类的能力和容量可以达到全球每年增加 1 太瓦各种类型电力的阶段。在某个时刻,我们会跨越鸿沟,太空数据中心变得有意义,但不是在当前这十年。那要远得多,一旦能源约束真正成为大瓶颈,一旦空间、土地、许可成为更大的瓶颈,因为它占据了越来越多的经济领域。而芯片不再是瓶颈,因为芯片是最大的瓶颈。所以你希望它们一制造完成就立即部署用于 AI 工作。所以人们正在做很多事情来加快这一速度,无论是模块化数据中心,还是模块化机架,你实际上在数据中心只放入芯片,而其他一切都已布线并准备就绪。所以人们正在做诸如此类的事情来缩短时间,而这些在太空中是无法做到的。归根结底,在芯片受限的世界里,唯一重要的是让这些芯片尽快开始产生 token。在一个世界,也许是 2035 年,一旦半导体行业和 ASML、Zeiss 以及所有其他供应商、L Research、Applied Materials、晶圆厂制造商,像钟摆一样摆动,他们能够制造足够的芯片,我们真正在优化每一个旋钮,那么优化 10% 或 15% 的能源成本才有意义,或者当我们转向 A6 时,英伟达的利润率不再是 70% 以上,也许能源成本占集群的 30%,晶圆厂建设、所有这些事情、数据中心建设——这些才是要优化的东西。但这不是,你知道,埃隆·马斯克不会通过 20% 的收益获胜。他从不那样赢。他赢的时候是全力出击,实现 10 倍的收益,对吧?这就是 SpaceX 的意义,这就是特斯拉的意义,这就是他所有成功的意义。不是追逐那 20%。所以我认为太空数据中心最终会带来 10 倍的收益,随着地球资源变得越来越有争议,但这不是当前这十年。

So, space data centers effectively are not limited by, you know, 'hey, we have this energy advantage.' It's actually just limited by the same contended resource. We can only make 200 gigawatts of chips a year by the end of the decade. So, what are we going to do to get that capacity? It doesn't matter if it's on land or in space. You can build that power. And I think human capabilities and capacity could get to the period where we're adding a terawatt a year globally of various types of power. At some point, we do cross the chasm where space data centers make sense, but it's not this decade. It is much further out once you have energy constraints actually being a big bottleneck. Once you have space, land, permitting being a much bigger bottleneck as it subsumes more and more of the economy. And chips are no longer the bottleneck because chips are the biggest bottleneck. And so you want them deployed working on AI the moment they're done being manufactured. And so there's a lot of things people are doing to increase that speed faster and faster, whether it be modularizing data centers or even modularizing racks where you actually put the chip in at the data center, but only the chip and everything else is already wired up and ready to go at the data center. So there's things like this that people are doing to decrease that time that you cannot do in space. And at the end of the day, all that matters in a chip-constrained world is get these chips working on producing tokens ASAP. In a world, maybe 2035, once the semiconductor industry and ASML and Zeiss and all these other suppliers, L Research, Applied Materials, fab manufacturers, like pendulum swings and they're able to make enough chips and really we're optimizing every dial and like it makes sense to optimize the 10% of energy costs or 15% of energy cost or as we move to A6 potentially and Nvidia's margins aren't 70% plus, maybe that energy cost is 30% of the cluster, and fab construction, all these things, data center construction—these are the things to optimize. But that's not, you know, Elon doesn't win by doing 20% gains. Elon never wins that way. Elon wins when he swings for the fences and does 10x gains, right? That's what SpaceX is about, that's what Tesla was about, that's what all of his success has been about. It's not been about chasing the 20%. So I think space data centers will eventually be a 10x gain, potentially, as Earth resources get more and more contentious, but that's not this decade.

Host

是的,我的意思是,我想只是提供一些关于地球上有多少土地的直觉。显然芯片本身,特别是如果我们进入一个机架以兆瓦充电的世界,实际上它甚至不是一个限制因素。

Yeah, I mean, I think just to drive some intuition about how much land there is on Earth. Obviously the chips themselves, especially if we move to a world where you have racks that are megawatt charge, like literally it's not even a ring factor.

Dylan Patel

这是另一件事,对吧?功率密度。你知道,如果芯片和制造是当前的约束,那么对于 AI 芯片等,大约是每平方毫米 1 瓦。

That's the other thing, right? The power density. You know, if chips and manufacturing is the constraint right now, roughly it's one watt per millimeter squared for AI chips and such.

Host

一个简单的方法是将它提高到每平方毫米 2 瓦。现在你可能不会获得 2 倍的性能,可能只获得 20% 的性能提升,而这需要更奇特的冷却,对吧?它需要更复杂的冷板和非常复杂的液冷,或者可能需要浸没式冷却之类的东西。但在太空中,更高的每毫米瓦数非常困难。而在地球上,这些都是已解决的问题。其中一件事能让你获得更多的 token。也许每制造一个晶圆能多出 20% 的 token。那是一个巨大的提升。

One easy way is to pump that to two watts per millimeter squared. Now you may not get 2x the performance, you may only get 20% more performance, and that requires much more exotic cooling, right? It requires more complicated cold plates and very complicated liquid cooling, or maybe it requires things like immersion cooling. But in space, higher watts per millimeter is very difficult. Whereas on Earth, these are solved problems. And one of these things enables you to get a lot more tokens. Maybe it's 20% more tokens per wafer that's manufactured. And that's a huge way.

Host

所以你说的毫米,是指芯片面积?

So by millimeter, you mean of die area?

Dylan Patel

是的,芯片面积。芯片面积的平方毫米。是的,我的意思是,这对太空更有利,因为如果你能运行更高的每毫米瓦数,芯片会变得更热。芯片越热……我想这是计算机芯片工程的问题,但根据斯特藩-玻尔兹曼定律,散热与温度的四次方成正比。所以如果你能运行一个非常热的芯片,因为你不能让它更热,你只能让它更密集。

Yeah, of die area. Square millimeters of a die area. Yeah, I mean it would be better for space because if you can run more watts per millimeter, the chip runs hotter. And the hotter the chip... I guess this is a question of computer chip engineering, but it cools to the power of fourth by Stefan-Boltzmann's law. So if you can run a very hot chip because you can't run it hotter, you can only run it denser.

Host

问题在于,将热量从那个密集区域排出意味着你必须从标准的风冷和液冷转向更奇特的液冷形式,甚至浸没式冷却,以达到更高的功率密度。这在太空中比在地球上更困难。

And the problem is getting the heat out of that dense area means you have to move away from standard air cooling and liquid cooling to more exotic forms of liquid cooling or even immersion to get to higher power densities. And that's more difficult in space than it is on Earth.

Dylan Patel

是的。也许现在值得解释一下 scale-up 到底是什么,以及它在英伟达、Tranium 和 TPU 中分别是什么样子。

Yeah. And maybe it's at this point worth explaining what exactly a scale-up is and what it looks like for Nvidia versus Tranium versus TPUs.

Host

是的。所以之前我提到过,芯片内部的通信非常快。同一机架内芯片之间的通信很快,但没那么快。然后你知道,那是在 TB 级别。而非常远距离的通信是在 GB 级别,几百 GB,对吧?所以随着距离增加,计算量级会变化,也许跨国家是每秒 GB 级别,对吧?Scale-up 域就是这个紧密的域,芯片之间以每秒 TB 级别通信。所以对于英伟达来说,以前这意味着一个 H100 服务器有 8 个 GPU,这 8 个 GPU 可以以每秒 TB 的速度相互通信。到了 Blackwell NVL72,他们实现了机架级 scale-up,这意味着机架中的所有 72 个 GPU 可以以每秒 TB 的速度相互连接,并且速度逐代翻倍。

Yeah. So earlier I was mentioning how communication within a chip is super fast. Communication within chips that are in the same rack is fast but is not as fast. And then you know it's on the order of terabytes. And then communication very far away is on the order of gigabytes, hundreds of gigabytes, right? So this order of magnitude as you get further distance, compute and maybe across the country it's on the order of gigabytes a second, right? Scale-up domain is this tight domain where the chips are communicating on the order of terabytes a second. And so for Nvidia previously this meant an H100 server had eight GPUs and those eight GPUs could talk to each other at terabytes a second. With Blackwell NVL72, they implemented rack-scale scale-up and that meant all 72 GPUs in the rack would connect to each other at terabytes a second speed and the speed doubled gen on gen.

扩展域拓扑差异 Scale-up domain topology differences

Dylan Patel

但他们做的最重要的创新是从 8 个扩展到 72 个。看看 Google,他们的 Scale-up 域完全不同,一直都是数千的量级。TPU v4 的 Pod 有 4000 个芯片,TPU v5 的 Pod 在 7000 到 9000 个左右。关键区别在于,这和 Nvidia 不一样,不是对等的。Google 的拓扑是环形(Torus),每个芯片只连接六个邻居;而 Nvidia 的 72 个 GPU 是全互联(All-to-All),它们之间可以每秒传输 TB 级数据,任意两个芯片都能直接通信。Google 这边,TPU 1 要和 TPU 76 通信,就必须经过多个芯片跳转,这总会造成资源阻塞,因为每个 TPU 只连接六个其他 TPU。所以拓扑和带宽都有差异,两者各有优劣。Google 能拥有巨大的 Scale-up 域,但代价是芯片间通信需要跳转,只能直接和六个邻居通信。Amazon 则改造了他们的 Scale-up 域,介于 Nvidia 和 Google 之间,试图做出更大的 Scale-up 域。他们部分采用全互联(用交换机,和 Nvidia 一样),部分采用环形拓扑(和 Google 一样)。随着下一代产品的发展,这三家都在越来越多地转向蜻蜓拓扑(Dragonfly),也就是部分全互联、部分非全互联,这样 Scale-up 域可以扩展到成百上千个芯片,同时跳转时不会争抢资源。

But also the most important innovation they did was going from 8 to 72 in the domain. When we look at Google, their scale-up domain is completely different. It has always been on the order of thousands. With TPU v4, they had pods the size of 4,000 chips. With TPU v5, they have pods in the 7,000 or 8,000 to 9,000 range. And what's relevant here is that it's not the same as Nvidia. It's not like for like. Google has a topology that is a torus. So every chip connects to six neighbors, rather than Nvidia where the 72 GPUs connect all-to-all, so they can send terabytes a second to each other to any arbitrary other chip in that pod of scale-up. Whereas with Google, you have to bounce through chips. So this means if TPU 1 needs to talk to TPU 76, it has to bounce through various chips, and there is always some blocking of resources when you do that, because that one TPU is only connected to six other TPUs. So there's a difference in topology and bandwidth. And there are trade-offs and advantages of both. Google gets to have a massive scale-up domain, but then they have the trade-off of you have to bounce across chips to get from one chip to another. You can only talk to six direct neighbors. And Amazon has mutated their scale-up domain. They're somewhere in between Nvidia and Google, effectively trying to make larger scale-up domains. They try to do all-to-all to some extent, which is with switches, which is what Nvidia does, but also to some extent they use torus topologies like Google does. As we advance forward to next generations, all three of them are moving more and more towards a dragonfly topology, which means there are some fully connected elements and some elements that are not fully connected, so you can get the scale-up to be hundreds or thousands of chips but also have it not contend for resources when you're bouncing through chips.

参数扩展与硬件容量 Parameter scaling and hardware capacity

Host

相关问题。我听到有人声称,参数 Scaling 之所以缓慢,直到现在 OpenAI 和 Anthropic 才推出越来越大的模型,是因为原始 GPT-4 已经超过一万亿参数,而现在模型才重新接近这个规模。有个理论说原因是 Nvidia 的 Scale-up 域内存容量一直不够。具体来说,如果你有一个 5 万亿参数的模型,以 FP8 运行,那就是 5 万亿 GB。再加上 KV 缓存,假设一个 batch 大小相同,那么一次前向传播就需要 10 TB。直到 GB200 和 VL72,Nvidia 的 Scale-up 域才有 20 TB,之前都小得多。而 Google 一直有巨大的 TPU Pod,虽然不是全互联,但单个 Scale-up 域就有数百 TB 的容量。这能解释为什么参数 Scaling 慢吗?

Related question. I heard somebody make the claim that the reason that parameter scaling has been slow and only now are we getting bigger and bigger models from OpenAI and Anthropic is that original GPT-4 is over a trillion parameters and only now are models starting to approach that again. And I heard a theory that the reason is that Nvidia's scale-ups have just not had that much memory capacity. So what was the claim exactly? If you have say a 5 trillion parameter model running at FP8, that's 5 trillion gigabytes. And then you have the KV cache. Let's say it's the same size for one batch. So you need 10 terabytes to be able to run a single forward pass. And then only with the GB200 and VL72 do you have an Nvidia scale-up that has 20 terabytes, and before that they were much smaller. Whereas Google on the other hand has had these huge TPU pods that are not all-to-all but still have hundreds of terabytes of capacity in a single scale-up. So does that explain why parameter scaling has been slow?

Dylan Patel

我认为部分原因是容量和带宽,但另一方面,模型越大,部署速度就越慢。对终端用户来说推理速度并不关键,真正关键的是强化学习(RL)。我们看到,在实验室里算力的分配主要有几种方式:分配给推理(即收入)、分配给开发(即制造下一代模型)、以及分配给研究。在开发中,又分为预训练和强化学习。所以实际情况是,研究带来的算力效率提升非常大,你其实希望把大部分算力投入研究,而不是开发。因为研究人员不断产生新想法、尝试、测试,推动缩放定律的帕累托最优曲线不断向前。至少从经验上看,模型成本每年下降 10 倍甚至更多。同等规模下成本降 10 倍,而要达到新前沿,成本可能持平或更高。所以你不应该把太多资源分配给预训练、后训练和强化学习,而应该把大部分资源分配给研究。中间是开发阶段。如果你现在预训练一个 5 万亿参数的模型,你就得花大量时间在强化学习的 rollout 上。5 万亿参数模型的 rollout 比 1 万亿参数模型大 5 倍。如果你想做同样数量的 rollout,也许大模型样本效率更高——假设是 2 倍。好,那你就需要 2.5 倍的强化学习时间才能让模型更聪明。或者你可以对小模型做 2 倍时间的强化学习,那么大模型虽然样本效率高 2 倍、做 x 次 rollout,但小模型只有 1 万亿参数,样本效率低一些,却可以做两倍数量的 rollout,而且完成得更快。这样你就能更快得到模型,做了更多强化学习,然后用这个模型帮助构建下一个模型,帮助工程师训练,实现各种研究想法。所以这个反馈循环在任何情况下都偏向小模型,无论硬件如何。至于 Google,他们确实部署了所有主要实验室中最大的生产模型,Gemini Pro 比 GPT-4 更大。

I think it's partially the capacity and bandwidth, but also as you build a larger model, the ability to deploy it is slower, right? In terms of inference speed for the end user, that's kind of irrelevant. What's really relevant is RL. And what we've seen with these models and allocation of compute at a lab is there are a few main ways you can allocate compute: you can allocate it to inference (i.e., revenue), you can allocate it to development (i.e., making the next model), and you can allocate it to research. In development specifically, you split it between pre-training and RL. So when you think about what exactly is happening, the compute efficiency gains you get from research are so large that you actually want most of your compute to go to research, not to development. Because all these researchers are generating new ideas, trying them out, testing them, and continuing to push the Pareto-optimal curve of scaling laws further and further. At least what we've seen empirically is that model cost gets 10x cheaper every year, or even more than that. At the same scale, it gets 10x cheaper, or to reach new frontiers, it costs the same amount or more. So you don't want to allocate too many resources to pre-training and post-training and RL; you actually want to allocate most of your resources to research. And then in the middle is this development period. If you pre-train a 5 trillion parameter model now, you have to spend all this time on rollouts in RL. The rollouts for a 5 trillion parameter model versus a 1 trillion parameter model are 5 times larger. If you wanted to do as many rollouts, maybe the larger model is more sample efficient—let's say it's 2x more sample efficient. Okay, great, now you need 2.5 times as much time of RL to get the model smarter. Or you could RL the smaller model for 2x the time, and you'd still have a 25% difference in the big model which is 2x more sample efficient and doing x number of rollouts, versus the small model which is 1 trillion parameters, although it's less sample efficient, is doing twice as many rollouts. It's still done faster, so you get the model faster sooner, and you've done more RL. Then you can take that model to help you build the next models, help your engineers train, and do all these research ideas. So this feedback loop is actually weighed towards smaller models in every case, no matter what your hardware is. And as you look to Google, Google does deploy the largest production model of any of the major labs, with Gemini Pro being a larger model than GPT-4.

模型大小与RL速度权衡 Model size vs. RL speed trade-off

Dylan Patel

或者它比 Opus 更大的模型,结果就是……是的,谷歌这么做是因为他们拥有单极化的算力,对吧?几乎全是 TPU。而 Anthropic 用的是 H100、H200、Blackwell、Trillium 以及各代 TPU。OpenAI 目前主要用 Nvidia,但也在转向 AMD 和 Trillium。像谷歌这样的算力集群可以针对更大的模型进行优化,利用上千个芯片在扩展域内大幅提升强化学习的速度,从而让反馈循环变快。但归根结底,孤立来看,你几乎总是希望用更小的模型,这样强化学习更快,能更快部署到研发中,从而构建下一代产品,获得更多算力效率优势。然后产生复合效应:我做了个更小的模型,强化学习更多,更早投入研发,训练本身消耗的算力更少,因为我能把更多算力分配给研究。这种越来越快做研究的复合效应可能带来更快的起飞,而这些公司追求的就是尽可能最快的起飞。

Or it's a larger model than Opus, and so you end up with... Yes, Google does this because they have a unipolar set of compute, right? Almost all TPUs. Whereas Anthropic is dealing with H100s, H200s, Blackwell, Trillium, TPUs of various generations, right? And OpenAI is dealing with mostly Nvidia right now, but going towards having AMD and Trillium as well. The fleets of compute like Google can just optimize around a larger model, and they can leverage a thousand chips in a scale-up domain to get the RL time speed much faster, so that you can actually have this feedback loop be fast. But at the end of the day, in isolation, you almost always want to go with a smaller model that gets RL faster and gets deployed into research and development, so you can build the next thing and get more compute efficiency wins. And then this compounding effect of: oh, I made a smaller model that I RL more, that I then deployed into research and development earlier, and I spent less compute on the training itself because I was able to allocate more compute to the research. This compounding effect of being able to do the research faster and faster and faster is potentially a faster takeoff, and that's all these companies want: the fastest takeoff possible.

Host

嗯。

Yeah.

Leopold为何赚大钱 Why Leopold makes outrageous money

Host

好,来个辛辣问题。你知道,你解释说 Semi Analysis 卖这些电子表格,你总是说,“啊,六个月前或一年前,我们告诉人们内存危机,现在你告诉人们洁净室危机,未来还有工具危机。”为什么只有 Leopold 用你的表格赚了大钱?其他人在干嘛?

Okay, spicy question. You know, you're explaining you make the Semi Analysis sells these spreadsheets, and you're always like, "Ah, six months ago or a year ago, we told people the memory crunch, or now you're telling people the clean room crunch, and then in the future the tool crunch." Why is Leopold the only person that is using your spreadsheets to make outrageous money? What is everybody else doing?

Dylan Patel

我认为有很多人以各种方式赚钱。显然,Leopold 开玩笑说他是唯一一个告诉我我们数字太低的客户。其他所有人都说我们的数字太高,几乎令人作呕。比如,某个超大规模云厂商说,“嘿,那个超大规模云厂商,他们的数字太高了,”我们说,“不,就是那样,”他们就说,“不不不,不可能,”等等。然后我们最终必须用所有事实和数据说服他们,当我们与超大规模云厂商或 AI 实验室合作时,实际上那个数字并不高,是正确的。但最终,有时需要六个月或一年他们才意识到。我认为其他客户,比如交易方面的,也使用我们的数据,对吧?我们把数据卖给很多……大概 60% 的业务是行业客户。所以 AI 实验室、数据中心公司、超大规模云厂商、半导体公司,整个 AI 基础设施供应链。但大约 40% 的收入来自对冲基金,对吧?我不会评论我们的客户是谁,但我认为很多人使用这些数据。问题在于你如何解读,以及你如何看待数据之外的东西。我要说的是,Leopold 几乎是唯一一个总是告诉我数字太低的人。有时他太高,有时我太低,对吧?但总的来说,我认为其他人也在做这件事,你可以查看某些……你可以看看整个领域的对冲基金,查看他们的 13F 文件,会发现他们持有的可能不完全是 Leopold 持有的。因为问题总是:什么是最受约束的东西?什么是最超出预期的东西?那才是你真正想要利用的:市场低效。从某种意义上说,我们的数据通过让基础数据更准确来使市场更高效。但从另一种意义上说,我认为许多基金确实根据公开信息进行交易,我不认为 Leopold 是唯一的人。不过,我认为他对整个 AGI 起飞最有信心。

I think there are a lot of people making money in many ways. I think obviously Leopold jokes that he's the only client of mine that tells me our numbers are too low. Everyone else tells me our numbers are too high, almost ad nauseam. You know, whether it's a hyperscaler saying, "Hey, that other hyperscaler, their numbers are too high," and we're like, "No, that's it," and they're like, "No, no, no, no, it's impossible," blah blah blah. And then you finally have to convince them through all these facts and data when we're working with hyperscalers or AI labs that in fact no, that number isn't too high. That's correct. But eventually, sometimes it's like six months later it takes them to realize, or a year later. I think other clients, like on the trading side, also use our data, right? We sell data to a lot of... I think roughly 60% of my business is industry. So AI labs, data center companies, hyperscalers, semiconductor companies, the whole supply chain across AI infrastructure. But then like 40% of our revenue is like hedge funds, right? And I'm not going to comment on who our customers are, but I think a lot of people use the data. It's just how do you interpret it and then what do you view as beyond it? And I will say Leopold is pretty much the only person who tells me my numbers are too low always. And sometimes he's too high, sometimes I'm too low, right? But in general, I think other people are doing that, and you can check certain... You can look across the space at hedge funds and look at their 13Fs and see actually they own maybe not exactly what Leopold does. Because it's always a question of: what is the most constrained thing? What's the thing that's going to be most outside of expectations? And that's what you're really trying to exploit: inefficiencies in the market. And in a sense, what our data shows is making the market more efficient by making the base data of what's happening more accurate. But in a sense, I think many funds do trade on information that is out there, and I don't think Leopold's the only person. I think he has the most conviction on the entire AGI takeoff though.

Host

对。但我的意思是,这些赌注不是关于 2035 年会发生什么。你下的赌注——至少从包括 Leopold 在内的不同基金的公开回报可以看出——是关于过去一年发生的事情,而过去一年的事情可以用你的电子表格预测,对吧?所以这更像是……

Right. I mean, but the bets are not about like what happens in 2035. The bets that you're making that are at least exemplified by public returns we can see for different funds including Leopold's are about what has happened in the last year, and the last year stuff could be predicted using your spreadsheets, right? So it's like it's less about...

Dylan Patel

这是关于购买下一年的电子表格。

It's about buying like the next year spreadsheets.

Host

它们不只是电子表格,你知道的?还有报告、API 数据访问,很多数据。但无论如何,我觉得……

They're not just spreadsheets, you know? There's reports, there's API access to the data, there's a lot of data. But anyways, you know, I think...

Dylan Patel

你明白我的意思吗?这不是关于某种疯狂的奇点事件,而是关于,“哦,你要买内存危机吗?”

Do you see what I mean? Like it's not about some crazy singularity thing, it's about like, "Oh, do you buy the memory crunch?"

Host

但一个简单的例子是:你只有在相信 AI 会大规模起飞时才会买入内存危机。而内存危机很大程度上是基于——至少对于湾区考虑基础设施的人来说——很明显:随着上下文变长,KV 缓存爆炸,所以你需要更多内存。然后你算一下,还需要对供应链有深入了解:哪些晶圆厂在建,哪些数据中心在建,多少芯片,等等。所以我们非常紧密地追踪所有这些不同的数据集。但归根结底,需要有人完全相信这会发生。我认为一年前,如果你告诉某人内存价格将翻四倍,智能手机销量将在未来一两年下降 40%,人们会说,“你疯了,这从未发生过。”除了少数人相信,那些人确实交易了内存,对吧?而且人们确实……我不认为 Leopold 是唯一购买内存公司的人。我认为有很多人购买内存公司。他当然在规模、定位和操作上比某些人——也许是大多数人——做得更好,对吧?我不想评论谁的回报如何。但他确实做得很好。但其他人也做得很好,对吧?你试图这样……哇,你让我第一次变得外交了。

A simple one though is like you only buy the memory crunch if you believe AI is going to take off in a huge way. And the memory crunch a lot of it was predicated on, like, at least for people in the Bay Area who think about infrastructure, it's obvious: KV cache explodes as context lengths go longer, so you need more memory. And then you do the math, and you also have to have a lot of supply chain understanding of what fabs are being built, what data centers are being built, how many chips, and all these things. So we track all these different data sets very tightly. But at the end of the day, it takes someone to fully believe that this is going to happen. I think a year ago, if you told someone memory prices will quadruple and smartphone volumes are going to go down 40% over the year or two after that, people were like, "You're crazy, that never happened." Except a few people did believe that, and those people did trade memory, right? And people did... I don't think Leopold is the only person buying memory companies. I think there were a lot of people buying memory companies. He of course sized and positioned and did things in better ways than some, maybe most, right? I don't want to comment on whose returns are what. But certainly did well. But other people also did really well, right? You're trying to be like this... Wow, you've made me diplomatic for the first time ever.

Host

不,不,你很好。你很好。我觉得这很有趣,对吧?我在当外交官,而通常我很辛辣。

No, no, you're fine. You're fine. I think this is hilarious, right? I'm being a diplomat, you know, whereas usually I'm like spicy.

Dylan Patel

嗯。

Yeah.

台积电会为AI将苹果踢出N2吗? Can TSMC kick Apple off N2 for AI?

Host

好,也许快速问答来收尾。如果像你说的,内存逻辑等等,N3 主要是 AI 加速器,但还有 N2,目前主要是苹果。未来,我猜 AI 也会想用 N2。如果 Nvidia、亚马逊和谷歌说,“嘿,我们愿意为 N2 产能付很多钱,”他们能把苹果踢出去吗?

Okay, maybe some rapid fire to close out. Can TSMC, if you're saying look, the memory logic, etc., the N3 is mostly going to be AI accelerators, but then there's N2, which is mostly Apple now. And then in the future, I guess AI would also want to go on N2. Can they kick out Apple if Nvidia and Amazon and Google say, "Hey, we really are willing to pay a lot of money for N2 capacity."

Dylan Patel

所以我认为这里的挑战是芯片设计周期很长。

So I think the challenge with this is chip design timelines take a long while.

台积电2nm节点与苹果角色变化 TSMC's 2nm node and Apple's changing role

Dylan Patel

所以那已经是一年多以后了,而基于 2nm 的设计还要一年多才能出来。

And so that's more than a year, and the designs that are on 2nm are more than a year out.

Host

是的。

Yeah.

Dylan Patel

所以真正会发生的是,苹果——抱歉,英伟达和其他公司会这样说:“嘿,我们打算预付产能费用,你为我们扩产。”然后苹果可能会……台积电会赚一点利润,但不会太多。他们不会完全把苹果踢出去,对吧?他们会做的是,当苹果下订单 X 时,他们可能会说:“嘿,我们预计你只需要 Y 或 X 减一。”所以我们就给你 X 减一。然后那部分弹性产能苹果就有点被坑了。而传统上,苹果总是多订大约 10%,然后在一年内再削减 10%,有些年份他们确实用满了那 10%……你知道,销量会随季节和宏观因素波动,等等。

And so what would really happen is Apple, or sorry, Nvidia and all these others will be like, "Hey, we're going to prepay for the capacity and you're going to expand it for us." And then Apple would be... and maybe TSMC takes a little bit of margin but not a ton. They're not going to kick Apple out entirely, right? What they're going to do is when Apple orders X, they may say, "Hey, we project you only need Y or X minus one." And so that's what we're going to give you is X minus one. And then that flex capacity Apple's kind of screwed on. Whereas traditionally Apple's always overordered by like 10% and cut back by 10% over the course of the year, and some years they hit the entire 10% just... you know, volumes vary right based on the season and macro, blah blah blah.

Host

是的。

Yeah.

Dylan Patel

所以我不认为台积电会踢掉苹果。我认为苹果在台积电营收中的占比会越来越小,因此台积电迎合其需求的重要性也会降低。台积电最终可能会开始说:“嘿,你得提前两年预订明年的产能,而且必须预付资本支出,因为英伟达、亚马逊和谷歌都是这么做的。”

And so I don't think TSMC would kick out Apple. I think Apple will become a smaller and smaller and smaller percentage of TSMC's revenue and therefore be less relevant for TSMC to cater to their demands. And TSMC could eventually start saying, "Hey, you got to pre-book your capacity for next year for two years out and you have to prepay for the capex because that's what Nvidia and Amazon and Google are doing."

Host

是的。我想知道是否值得谈一下具体数字,比如……我手头没有数据,像苹果在 N2 晶圆上占了多少比例,或者未来几年与 AI 相比如何?

Yeah. I wonder if it's worth going to specific numbers on like... I don't have any of them on hand of like how many N2 wafers or what percentage of N2 does Apple have its hands on versus over the coming years versus AI?

Dylan Patel

是的,我的意思是,今年苹果将占据 N2 产能的大部分。AMD 有一点份额,他们正在尝试早期制造一些 AI 芯片和 CPU 芯片。份额很小,但大部分是苹果的。嗯,到了明年,随着其他人开始上量,苹果的份额仍然接近一半,但随后会急剧下降,对吧?就像 N3 一样,他们占了一半。嗯,我们会看到的。我说 N2 时,也包括 A16,它是 N2 的一个变种。随着时间的推移,这些节点将成为主流。还有一点有趣的是,传统上苹果总是第一个采用新工艺节点。2nm 实际上是他们第一次不是第一个。嗯,除了华为,对吧?华为在 2020 年及之前曾与苹果并列第一,但两者都在制造智能手机。现在到了 2nm,AMD 正试图在同一时间框架内制造 CPU 和 GPU 小芯片,并使用先进封装将它们封装在一起,这与苹果的时间线相同。嗯,这对 AMD 来说是一个很大的风险,可能导致潜在的延迟,因为这是一种全新的工艺技术,难度很大。但归根结底,这是一个赌注,他们想以此比英伟达更快地扩展规模,并试图击败他们。

Yeah, I mean this year Apple has the majority of N2 that's going to get fabricated. There's a little bit from AMD. They are trying to make some AI chips and CPU chips early. There's a little bit, but for the most part, it's Apple. Um, and as we go forward to the year after that, Apple still, you know, gets closer to like half of it as other people start ramping, but then it falls drastically, right? Just like for N3, they were half. Um, we'll see. And when I say N2, that includes A16, which is a variant of N2. Over time those nodes will be the majority. And what's also interesting is traditionally Apple's been the first to a process node. 2nm is actually the first time they're not. Well, besides Huawei, right? Huawei back in 2020 and before was the first with Apple, but they were both making smartphones. Now with 2nm, you've got AMD trying to make a CPU and a GPU chiplet that they used advanced packaging to package together in the same time frame as Apple. Um, and this is a big risk for AMD that causes potential delays potentially because it's a brand new process technology. It's hard. But at the end of the day, this is a bet that they want to do to, you know, scale faster than Nvidia and try and beat them.

Dylan Patel

实际上,当我们进入 A16 节点时,第一个客户甚至不是苹果,而是 AI。随着我们向前发展,这种情况会越来越普遍。苹果不仅不会是第一个采用新节点的,也不会是新节点的主要产量贡献者。然后他们就会像任何老客户一样。而且由于台积电的资本支出规模不断膨胀,而苹果的业务增长却没有跟上,他们变得越来越不重要。他们还会削减订单,因为供应链中的各个环节都在排挤他们,无论是封装、材料、DRAM 还是 NAND。这些成本都在上升,他们很可能无法将所有成本转嫁给客户,因为消费者并不那么强劲,最终你会陷入这样的困境:他们不再是台积电历史上那样的最佳伙伴了。

As we move forward, actually, when we move to the A16 node, the first customer there is not even Apple. It's AI. And as we move forward, that will become more and more prevalent. Not only will Apple not be the first to a node, they will also not be the majority of the volume to the new node. And then they'll just be like any old customer. And because the scale of TSMC's capex keeps ballooning but Apple's business is kind of not growing at the same pace, they become a less and less relevant customer. And they also will just cut their orders because things in the supply chain are kicking them out, whether it be packaging or materials or DRAM or NAND. These things are increasing in cost, they can't pass on all the cost to customers likely because the consumer is not that strong, and you end up with like this conundrum where they are just not Apple TSMC's best bud like they have been historically.

Host

你认为如果华为能获得 3nm 工艺,他们会拥有比 Reuben 更好的加速器吗?

Do you think if Huawei had access to 3nm they would have a better accelerator than Reuben?

Dylan Patel

有可能。是的。我认为华为……他们也是第一个推出 7nm AI 芯片的公司。他们是第一个推出 5nm 手机芯片的公司,但他们也是第一个推出 7nm AI 芯片的。华为昇腾比 TPU 早了大约两个月,比英伟达的……我想说是 V100 还是 A100?应该是 A100。所以,你知道,我的意思是,那只是工艺上的进步。不,这并不意味着软件或硬件设计等其他方面。但华为可以说是世界上唯一一家拥有所有腿的公司,对吧?华为拥有顶尖的软件工程师。华为拥有顶尖的网络技术。这实际上是他们历史上最大的业务,对吧?而且他们拥有顶尖的 AI 人才。但此外,超越英伟达,他们实际上拥有更好的 AI 研究人员。此外,超越英伟达,他们有自己的晶圆厂。此外,超越英伟达,他们有自己的终端市场,比如销售代币等。华为往往能够吸引最顶尖的人才。英伟达也能,但没有那么集中。而且华为在中国有更大的人才库。很有争议的是,如果华为拥有台积电,他们会比英伟达更好。而且中国在某些领域有优势,这些领域英伟达不容易进入,对吧?不仅仅是规模,还有一些技术方面,比如某些光学技术,中国实际上非常擅长。所以我认为非常合理的是,如果在 2019 年那个问题……不是华为被禁止使用台积电,华为本已超越苹果成为台积电最大的客户,而且华为在网络、计算、CPU 等领域都占有巨大份额。他们会继续获得份额,并且很可能成为台积电最大的客户。

Potentially. Yeah. I think Huawei... they were the first with a 7nm AI chip as well. They were the first with a 5nm mobile chip, but they were the first with a 7nm AI chip. The Huawei Ascend was like two months before the TPU and like four months before Nvidia's... I want to say was it V100 or A100? A100 I think. And so, you know, I mean, that's just moving to a process. No, that doesn't imply software. It doesn't imply hardware design, all these other things. But Huawei is arguably the only company in the world that has all the legs, right? Huawei has cracked software engineers. Huawei has cracked networking technologies. That's in fact their biggest business historically, right? And they have cracked AI talent. But furthermore, beyond Nvidia, they actually have better AI researchers. And furthermore, beyond Nvidia, they have their own fabs. And furthermore, beyond Nvidia, they have their own end market of like selling tokens and things like that. And Huawei tends to be like... they're able to get the top top top talent. Nvidia is as well, but not as concentrated. And Huawei has a bigger pool in China. It's very arguable that Huawei, if they had TSMC, would be better than Nvidia. And there are areas where China has advantages outside of... in areas that Nvidia can't access as easily, right? Around not just scale, but also like some things around tech, you know, certain optical technologies China's actually really good at. So there's certain... I think it's very reasonable that if in 2019 that issue that... that was not that Huawei was not banned from using TSMC, Huawei would have already eclipsed Apple as the biggest TSMC customer, and Huawei has huge share in networking and compute and CPUs and all these things. They would have kept gaining share and they'd likely be TSMC's biggest customer.

Host

哇,这太疯狂了。我有个有点随意的最后一个问题。埃隆采访的另一部分是关于机器人的。

Wow, that's crazy. I've got a kind of a random final question for you. So the other part of the Elon interview was robots.

人形机器人的云端与终端计算 Cloud vs. On-Device Compute for Humanoids

Host

那么,如果人形机器人的普及速度超出预期,到 2030 年有数百万台人形机器人到处跑,每台都需要本地算力,你觉得这意味着什么?需要什么条件?人们在机器人上部署 VLM 等东西有很多困难,但在某种程度上,你不需要把所有的智能都放在机器人里。不这样做会更高效,对吧?因为在云端你可以做批处理等等。所以你可能想做的是:大量的规划和长周期任务由云端一个能力更强、运行在很高批处理规模的模型来决定,然后它把这些指令推送给机器人,机器人再在后续每个动作之间进行插值。或者给它一个指令,比如“拿起那个杯子”,然后机器人上的模型可以拿起杯子,在拿起的过程中,它可能意识到“实际上,重量和力之类的可能需要由机器人上的模型来决定”。但并不是所有事情都需要像“嘿,拿起那个……”或者“嘿,那是个耳机”这样。实际上,云端的超级模型知道这些耳机是索尼 XM6s,这不是一个愚蠢的广告植入,但你知道……

And so if humanoids take off faster than people expect, if by 2030 there's millions of humanoids running around, each needing local compute, any thoughts on what that implies, what would be required for that? There's a lot of difficulties with VLMs and all these things that people are deploying on robots, but to some extent you don't need to have all the intelligence in the robot. It would be much more efficient to not do that, right? Because in the cloud you can batch process and all these things. So what you may want to do is: a lot of the planning and longer horizon tasks are determined by a much more capable model in the cloud that runs at very high batch sizes, and then it pushes those directions to the robots, who then interpolate between each subsequent action. Or it's given like, 'Hey, pick up that cup,' and then the model on the robot can pick up the cup, and as it's picking up, it's like, 'Oh, in fact, things like weight and force may have to be determined by the model on the robot.' But not everything needs to be like, 'Hey, pick up the...' or 'Hey, that's a headphone.' Actually, the supermodel in the cloud knows that these headphones are Sony XM6s, which is not a dorkish ad spot, but you know...

Dylan Patel

就像,这家伙为什么这么卖力地推销这东西?它就在桌子上。我们俩一起进 Satia 的时候它就在他脖子上。他是不是收了索尼的钱?

Like, why is this guy plugging this thing so hard? It's like on the table. It's like on his neck when we're entering Satia together. Like, is he getting paid by Sony?

Host

嗯,很不幸,没有。但不管怎样,它可能会说:“嘿,头带很软,这是它的重量等等。”然后机器人上的模型可以没那么智能,接收这些输入并执行动作。它可能每秒被云端模型告知一次,或者每秒十次,取决于动作的频率。但很多处理可以卸载到云端,否则,如果你在设备上做所有处理,我认为会更贵,因为你无法批处理。第二,你无法拥有像云端那么多的智能,因为云端的模型只会更大。第三,我们处于半导体短缺的世界,你部署的任何机器人都需要最先进的芯片,因为机器人的功耗很成问题,对吧?你需要低功耗和高效率,然后突然间你把原本用于 AI 数据中心的功耗和芯片放到了机器人里。

Um, unfortunately not. Unfortunately not. But anyways, like you know, it might say, 'Hey, the headband is soft and this is the weight of it and all these things.' And then the model on the robot can be less intelligent and take these inputs and do the actions. And it may get told by the model in the cloud every second, every 10 times a second maybe, depending on the hertz of the action. But a lot of that can be offloaded to the cloud because otherwise, if you do all of the processing on the device, I believe it would be more expensive because you can't batch. Two, you couldn't have as much intelligence as you do in the cloud because the models will just be bigger in the cloud. And three, we're in a semiconductor shortage world, and any robot you deploy needs leading edge chips because the power is really bad for robots, right? You need it to be low power and efficient, and all of a sudden you're taking power and chips that would have been for AI data centers and you're putting them in robots.

Host

是的。

Yeah.

Dylan Patel

所以现在,如果你部署数百万台人形机器人,那 200 吉瓦就会减少。我认为这非常有趣,因为人们可能没有意识到未来智能在物理意义上会有多集中。现在对于人类来说,你的计算机——就像有 80 亿人类,他们的算力就在他们的脑袋里、身上。在未来,即使是有实体机器人在世界上——我的意思是,显然知识工作会以集中的方式在数据中心完成,有数十万甚至数百万个实例——但即使对于机器人技术,你所暗示的未来也是一个更集中的思考和计算驱动着世界上数百万台机器人的未来。我认为这只是关于未来的一个有趣事实,人们可能没有意识到。

So now that 200 gigawatt gets lower if you're deploying millions of humanoids. I think this is very interesting because something people might not appreciate about the future is how centralized in a physical sense intelligence will be. Right now with humans, your computer—like there's 8 billion humans and their compute is on their heads, on their person. In the future, even with robots that are out physically in the world—I mean obviously knowledge work will be done in a centralized way from data centers with huge like hundreds of thousands of instances or maybe millions of instances—but even for robotics, the future you're suggesting is one where there's more centralized thinking and centralized computation that's driving millions of robots out in the world. I think that's just an interesting fact about the future that people might not appreciate.

Host

我认为埃隆意识到了这一点,这就是为什么他去不同的地方采购芯片,对吧?他和三星签了巨额协议,在得克萨斯制造他的机器人芯片,因为他认为——我个人认为他认为台湾风险很大。正因如此,加上资源集中在台湾,他把机器人芯片放在得克萨斯,并且拥有一个独立的供应链,不那么受限——除了英伟达下周要推出的新 LPU 之外,真的没有人在三星上制造 AI 芯片,但我们是在那周之前录制的。

I think Elon recognizes this, which is why he's going to different places for his chips, right? He signed this massive deal with Samsung to make his robot chips in Texas because he thinks—I personally think he thinks that Taiwan risk is huge. And because of that and the centralization of resources in Taiwan, him having his robot chips in Texas and also being a separate supply chain that is not as constrained—no one's making AI chips really on Samsung besides Nvidia's new LPU that they're launching next week, but we're recording it the week before.

Dylan Patel

哦,这期节目在那之前播出。太棒了。所以,他们下周要推出这款新的 AI 芯片,基于三星制造,但这只是英伟达最近的一个进展。那是那里唯一的其他 AI 需求。而在台积电,一切都在竞争。所以,他既获得了地缘政治多元化,也为他的机器人获得了供应链多样性。而且他不用和那些数据中心天才们无限的支付意愿竞争太多。

Oh, this episode's coming out before. Sick. So, they're launching this new AI chip next week, which is built on Samsung, but that's like sort of a recent development from Nvidia. And then that's the only other AI demand there. Whereas on TSMC, everything is competing. So, he gets this like both geopolitical diversification, but also supply chain diversity for his robots. And he's not competing as much with the willingness to pay of infinity of the data center geniuses.

台湾风险与去风险策略 Taiwan Risk and De-risking Strategies

Host

好的,最后一个问题。关于台湾。如果我们相信工具是最终的瓶颈,那么仅仅通过制定一个计划,在情况发生时——比如台湾被封锁之类的——把台积电的每一位工艺工程师空运出来,我们能在多大程度上降低台湾在半导体供应链中的地位?还是说你实际上仍然需要运出 EUV 工具,每个工具就需要多架飞机运输,而且不现实?如果你运出所有工艺工程师,并假设情况足够紧急以至于你摧毁了晶圆厂,现在没有人拥有台湾所有的晶圆厂,这是一个很大的风险,对吧?

Okay, final question. On Taiwan. If we believe that tools are the ultimate bottleneck, how much of Taiwan's place in the semiconductor supply chain could we de-risk simply by having a plan to airlift every single process engineer at TSMC out when things come to—if they get blockaded or something? Or do you actually still need to ship out the EUV tools, which would be multiple plane loads per single tool and would not be practical? If you ship out all the process engineers and assuming it's hot enough that you destroy the fabs, no one has all the fabs in Taiwan now, which is a big risk, right?

Dylan Patel

你知道,这些工具实际上使用了大量在台湾制造的半导体。所以这就像一个蛇咬自己尾巴的梗,因为没有台湾的芯片你就无法制造工具,而没有工具你也无法在台湾使用芯片。显然那里有一些多元化,但归根结底,还是有点尾巴咬龙的意思。仅仅运出所有工程师并炸毁晶圆厂意味着中国拥有比世界其他地区更强的半导体供应链,对吧?就垂直整合而言,既然你移除了台湾,你拥有了所有技术诀窍,但你必须在比如亚利桑那州或其他地方为台积电复制它,而要建立台积电多年来积累的所有产能需要很长时间。所以你大幅减缓了美国和全球的 GDP——不仅仅是增长,你大幅缩减了 GDP——而且你面临更大的问题。你增加算力的增量能力几乎降为零,对吧?到本十年末,原本每年几百吉瓦,假设本十年末台湾出了事。现在你可能只有英特尔和三星的 10 吉瓦或 20 吉瓦。这几乎等于零,对吧?

You know, these tools actually use a lot of semiconductors which are manufactured in Taiwan. So it's like a snake eating its own tail sort of meme because you can't make the tools without the chips from Taiwan, which you can't use without the tools in Taiwan. There's obviously some diversification there, but at the end of the day, there is some tail eating the dragon. Just shipping out all the engineers and blowing up the fabs means China has a stronger semiconductor supply chain than the rest of the world, right? In terms of verticalization, now that you've removed Taiwan and now you've got all the knowhow, but you've got to replicate it in, let's say, Arizona or wherever for TSMC, and it's going to take a long time to build all the capacity that TSMC has had built over the years. So you've drastically slowed US and global GDP—not just growth, you've shrunk the GDP massively—and you've got a lot bigger problems. And your incremental ability to add compute goes to almost zero, right? Instead of hundreds of gigawatts a year by the end of the decade, let's say by the end of the decade something happens to Taiwan. Now you're at maybe like 10 gigawatts across Intel and Samsung or 20 gigawatts. It's like nothing, right?

Host

嗯,但现在你突然在 AI 中引发了一些疯狂的动态。

Um, but now all of a sudden you've really caused some crazy dynamics in AI.

产能扩张 Capacity expansion

Dylan Patel

当然,现有的算力是有的,但跟正在扩张的算力相比,简直是小巫见大巫。

Of course you have all the existing capacity, but that existing capacity pales in comparison to the capacity that's being expanded.

Host

好的,Dylan,讲得太棒了。非常感谢你来做客。

Yeah. Okay, Dylan, that was excellent. Thank you so much for coming on the podcast.

Dylan Patel

谢谢邀请。今晚见。

Thank you for having me. And see you tonight.

互动版:逐字朗读 + 针对本期提问 →