From Forum Poster to AI Infrastructure Analyst: The Story of Semi Analysis
打开互动全文版(中英对照 + 朗读 + 问答)→Dylan Patel 分享了 Semi Analysis 从青少年时期论坛发帖到如今 90 人研究公司的历程,专注于 AI 和半导体供应链。
Dylan Patel shares the journey of Semi Analysis from his teenage forum posts to a 90-person research firm covering AI and semiconductor supply chains.
好的,大家好,欢迎回到我们下一期的“下一件大事”播客。今天和我一起的是我的同事 Klay Hyman,以及我们的新合伙人、Semi Analysis 研究集团的创始人 Dylan Patel。今天很激动,能有机会和 Dylan 一起梳理 AI 基础设施领域的现状。你可能在很多不同的播客上都听过 Dylan 的分享。我也一直关注他的内容,以及 Semi Analysis 网站上的新闻通讯。最近我看到的其中一篇是关于太空数据中心的。如果有人对这个话题感兴趣,他们有一篇非常棒的文章,细节很丰富。但 Dylan,我很想听听 Semi Analysis 这个想法是怎么来的。我知道在 Substack 社区里,现在大家都在谈论你们公司的收入和取得的成功,但我感觉很多时候人们只看到现在的成功,却忘记了背后的历程、起源以及为之付出的努力。
All right. Hello everyone and welcome back to our next episode of the next big thing podcast here. I'm joined by my colleague Klay Hyman and Dylan Patel, our new partner, founder of Semi Analysis, the research group. It's an exciting day that we get to sort of go through the current state of the AI infrastructure landscape with Dylan. You've probably heard Dylan on many different podcasts. I follow his content all the time as well as the newsletter on the Semi Analysis site. Most recently, one of the ones I saw was about data centers in space. So if anyone is interested in that topic, they have a great piece with a lot of detail out on that. But Dylan, I would love to hear a bit about where the idea for Semi Analysis sort of came from. I know in the Substack community these days, people are talking a lot about the revenue of the firm and the success the firm has achieved, but I feel like a lot of times you sometimes see the current level of success and forget about the journey and the origins and how much work probably went into all of it.
是的。我认为 Semi Analysis 的起源真的来自于“发帖”,以一种不太严肃的方式在网上发帖,对吧?回想我最早关于半导体的帖子,那是我十几岁的时候,在我甚至还没有智能手机之前,我就在网上发帖讨论芯片、智能手机、手机屏幕和手机 SoC 等等。我就是对这些东西着迷,同样还有游戏硬件、PC 硬件、主机硬件。我一直在各种论坛上发帖谈论这些东西。到 12 岁的时候,我已经在管理和创建很多与 Android、Apple、Google、Intel、Nvidia、AMD 等硬件话题相关的论坛了,还有 Reddit 上的各种相关论坛。这就是一切的起源。我一直是个发帖者。我总是发表我的观点,总是回复,总是接受评论。而且我认为,现在我们的团队有 90 人,所以实际上我手下有市场部的人,他们会说:“Dylan,别在网上跟那些无名小卒瞎扯了。你让我们看起来很糟糕。”但我就是有那种强烈的欲望,当有人试图批评我时,我一定要在网上回应他们。所以这可能是个坏习惯,但基本上,我的整个青少年时期都在管理这些论坛。我在十几岁后期开始赚钱后就开始投资。我做了两年量化交易员,然后创办了自己的公司,但整个过程我一直在发帖、发帖、发帖。我有过匿名博客,发过匿名帖子。但到了 2020 年,我厌倦了我的工作。对量化交易员工作的幻想破灭了,它并没有看起来那么美好。是的,你能赚钱,但并没有看起来那么美好。我有点辞职了,创办了自己的公司。我当时并不确定它会走向何方,但我在自己做的 WordPress 网站上发帖,用的是真名,内容混合了技术、商业、金融、供应链,这些是我最感兴趣的东西,因为我是在一个小生意环境中长大的。我是在汽车旅馆长大的。我父母在佐治亚州乡下经营一家汽车旅馆,我们就住在里面,所以我懂生意。后来我们还有了加油站,所以我算是生活在其中,在其中成长。所以我一直热爱商业。从投资的角度,从东西如何制造的角度来看,供应链总是很有趣。我觉得我一直有那种 knack,就是想知道东西是怎么造出来的。技术方面当然非常令人兴奋,金融方面也非常令人兴奋。所以结合这些,我最早的一些帖子是关于当时美国禁止华为使用台积电的,所以我第一篇帖子实际上是关于联发科是最大赢家。因为华为在中国智能手机芯片和智能手机整体市场占有率第一。显然,由于无法再使用台积电,华为的份额会暴跌,所以美国市场认为,“哦,高通会赢。”但我认为,联发科这家台湾公司实际上会赢得更多份额,因为从地缘政治角度看,既然我们刚刚制裁了华为,中国更愿意从台湾公司而不是美国公司购买。所以两家公司都受益了,但联发科受益更多。这就是技术、供应链、金融、地缘政治这些东西融合在一起的例子。在接下来的几年里,我把我的 WordPress 转换成了 Substack,开始收费,我写的是关于整个半导体和 AI 供应链的所有这些话题。因为我非常关注 AI,在做量化交易员时也做过一点 AI 相关的工作,同时出于热情关注半导体。然后这就不断发展壮大。四年来,我周游世界,参加了世界上所有的会议。我一年参加 40 场会议。我没有固定的住处;我只是去参加我能去的每一个会议,无论是像 NeurIPS、ICML、ICLR 这样的 AI 会议——这些是 AI 研究者的会议,主要是研究人员和一些公司去那里展示他们的研究——一直到下游的、关于半导体供应链中化学品投入的随机会议。所以我会上上下下地跑遍整个产业链,无论是服务器、网络、制造、AI,整个产业链。我一年参加 40 场会议。有些会议非常小众,只有 300 人参加,而且除了大约五个人之外,其他人只说日语,我会想,“好吧,无所谓,就是这样。”还有一些会议有 1 万、2 万人,规模很大。这就是整个光谱和连续体。所以我能够跨越整个生态系统。一个会议你去三次,现在你就懂那种语言了。
Yeah. So I think the origin of Semi Analysis really comes from ship posting, posting online in a less serious manner, right? When I go back to what were my first posts about semiconductors, they were when I was in my tween, you know, and I was just posting on the internet about chips, about smartphones, about smartphone displays and smartphone SoCs and all these other things before I ever even had a smartphone. I was just obsessed with them, and same with gaming hardware, PC hardware, console hardware. I was just posting about this stuff all the time on the internet on various forums. By the time I was 12, I was moderating and creating a lot of forums related to Android and Apple and Google and Intel, Nvidia, AMD, like the hardware topics of the world, various forums on Reddit related to these things. So that's sort of where the origin of it all comes from. I was just always a poster. I've always posted my opinion, always replied, always taken comments. And I think one of the things the people in my team who, now we're a 90-person organization, so I actually have people in marketing, they're like, "Dylan, stop replying to random bozos on the internet. You're making us look bad." It's like I just have that burning desire to respond to anyone on the internet when they try and criticize me. So that's maybe a bad thing, but basically throughout my teenage years I was moderating these forums. I started investing as soon as I started making money in my later teenage years. I was a quant for two years and then I started my firm, but the whole time I was posting, posting, posting. I had anonymous blogs, anonymous posts. But then in 2020, I got fed up with my job. Disillusionment of being a quant is not as amazing as it seems. Yes, you make money, but it's not as amazing as it seems. And I kind of quit my job and started my company. I wasn't exactly sure where it was going to go, but I was posting on a WordPress website that I made, and I was posting under my real name, and I was posting this mix of technology, business, finance, supply chain, sort of the things I was most interested in because I grew up in a small business. I grew up in a motel. My parents owned a motel in rural Georgia and we lived in it, and I kind of just know business. We later had gas stations too, so I sort of lived in it, grew up in it. So I always loved business. Supply chain was always interesting from the perspective of investing, from the perspective of how things are made. I think that's always just been a knack, like how are things made. The technology aspect of course is super exciting, and the finance aspect is super exciting. So taking the combo of that, my first posts were things like, at the time, China was banning or the US was banning Huawei from access to TSMC, and so my first post was actually about how MediaTek was the biggest winner. Because Huawei was number one market share in China for smartphone chips and smartphones in general. Obviously, that was going to tank because they no longer had access to TSMC, and so the US market thought, "Oh, Qualcomm's going to win." But MediaTek, which is the Taiwanese firm, I thought they're actually going to win a lot more share because it turns out geopolitically China would rather buy from a Taiwanese firm than from a US firm given we just banned Huawei. So both firms benefited, but MediaTek benefited a lot more. So it was that mix of technology, supply chain, finance, geopolitics, all these things sort of melded together. And over the coming years, I converted my WordPress into a Substack at one point, started charging for it, and I was writing about all these topics across the entire semiconductor and AI supply chains. Given I followed AI a lot, I did a little bit of AI when I was a quant as well as following semiconductors out of a passion. And that just grew and grew and grew. For four years I traveled all around the world and I went to every conference in the world. I was going to 40 conferences a year. I didn't live anywhere; I was just going to every conference I could, whether it be an AI conference like NeurIPS, ICML, ICLR — these are sort of AI researcher conferences, mostly researchers and some companies go there and present their research — all the way downstream to random conferences for chemicals that are inputs into the semiconductor supply chain. So I'd go up and down the stack, whether it be servers, networking, fabrication, AI, just all up and down the stack. I'd go to 40 conferences a year. There'd be some that are super niche and there'd be 300 people there and they only spoke Japanese except for like five people, and I'd be like, "Well, whatever, right? That is what it is." And then there'd be others where it'd be 10,000, 20,000 people and it'd be huge. And it'd be the whole spectrum and continuum. And so I was able to go across the whole ecosystem. You go to a conference three times, you actually know the language now.
你在那里认识人,可以建立所有这些背景,然后向他们提问。我建立了整个生态系统和学习体系,覆盖了每一次转折点。我对技术非常好奇,但一旦某件事脱颖而出,无论是技术还是供应链,我从会议中学到它会在供应链或财务层面走向何方。所以我写的是整个混合体。有时报告以技术为中心,金融界没人关心;其他时候他们会说,‘哇,这是瓶颈,或者这是正在发生的转折,或者这家公司会因为下一代技术而获得大量市场份额’,而我会在华尔街任何人、任何对冲基金、任何人之前预测到。所以那就是开始。然后随着那个 Substack 在 2022 年不断增长,我开始招人。我的前两个员工是我在 Discord 上认识多年的朋友。但之后,我的第三个员工是 Myin,他之前在对冲基金工作,正要搬到日本和妻子一起生活,所以他算是个自由人。我发了一篇文章,说——那在当时是一篇很有趣的文章——在 23 年初,内存是 AI 的最大输家,原因是 AI 芯片和 AI 服务器使用的内存量比普通服务器少得多。普通服务器大约一半的物料成本是内存;在 AI 服务器中,内存占比小得多。部分原因是英伟达的利润率更高,还有其他几个因素。当然,英伟达的下一代芯片大幅增加了内存含量,现在内存占比高多了。但当时,文章说内存是最大输家。在付费部分,我说,‘嘿,我在招人。’Myin 联系了我,他是第一个有对冲基金背景的人。另外两个人是技术背景。他一加入公司,我们就开始制作所有这些模型,真正把业务从——我们仍然做 newsletter,仍然发布大量精彩内容,比以往更多——但我们将它转变为围绕销售信息服务、销售这些报告、销售这些数据集。随着这一切开始发生,球开始滚下山坡。从 2023 年到 2024 年,我从 2 人增加到 7 人。然后从 24 年底到 25 年初,我从 7 人增加到 20 人。从 25 年到 26 年,我从 20 人增加到 60 人。现在今年我们有 90 人。所以我们今年增加了 30 人。球一直在滚下山坡,我们不断添加新领域。我一直对所有事情感兴趣,但现在我能招到专家了。所以我认为半导体分析最令人兴奋的是,我不知道还有哪家公司有我们这样的专业水平和密度。我有在 ASML、应用材料、泛林半导体等设备公司工作过的人,这些公司制造晶圆,一直上游到在英特尔、台积电、英伟达、微软和亚马逊工作过的人,当然还有在 OpenAI 做模型的人,在特斯拉做 FSD 的人,在 Cohere 工作过的人。我们有在模型层工作过的人,在另一个垂直领域,我们有在数据中心工作过的人。我公司里有人曾在哈萨克斯坦建过发电厂。我们就是拥有这种最疯狂的人才密度,太棒了。所以公司一半是来自整个行业的工程师,另一半是前对冲基金或来自互联网的超级热情的普通人。我在 Twitter 或 Discord 上找到他们,然后说,‘你很聪明。来为我工作吧。’而且这很管用。这就是我们建立起来的,现在 SemiAnalysis 有很多不同的业务线。显然有数据服务、咨询、信息服务,我们有 newsletter。我们在做所有这些媒体。我们很快要举办一个大型会议。还有各种各样我们做的事情。这真是一段疯狂的旅程。
You know people there, you can build all these contexts so you can ask them questions. I developed this whole ecosystem and learnings, covering the inflections at each of these. I was very curious technologically, but then once something stood out, whether it be technology or supply chain, I learned from a conference I knew where it would lead to on a supply chain or finance basis. So I was writing about the whole mix. Sometimes reports would be centered around technology and no one in finance would care, and other times they'd be like, 'Wow, this is the bottleneck or this is the inflection that's happening, or this company is going to gain a ton of market share because they're next-generation technology,' and I would call it before anyone else on the street, before any hedge fund, before anyone. So that was like the start. And then as that Substack grew and grew in 2022, I started hiring people. My first two hires were people I knew from Discord for years. But then after that, my third hire was Myin, and he worked at a hedge fund before and he was moving to Japan to live with his wife there, so he was sort of a free agent. I'd put in a post about how, you know, it's interesting—this is a very interesting post at the time—in early '23, memory was the biggest loser from AI, and the reason was because the amount of memory that AI chips used and AI servers used versus regular servers was a lot less. Right, regular servers about half the BOM was memory; in AI servers it was much less. Part of that was Nvidia's margins are much higher, part of that was there's a couple different factors. Of course, Nvidia's next-generation chips have increased the memory content dramatically, and now it's like way more. But at the time, it was saying memory is the biggest loser. And in the paid section, I said, 'Hey, I'm hiring.' And Myin reached out, and he was the first person from a hedge fund background. The other two people were technical backgrounds. And so as soon as he joined the firm, we started making all these models, and we really converted the business from—we still do the newsletter, we still post a lot of amazing content there more than ever before—but we converted it from that to being one around selling information services, selling these reports, selling these data sets. And as that started to happen, the ball started tumbling down the hill. And from 2023 to 2024, I went from 2 to 7. And then from '24 to the end of '25, to the beginning of '25, I went from 7 to 20. And from '25 to '26, I went from 20 to 60. And now this year we're at 90. So we've added 30 people this year. So it's just been a ball tumbling down the hill, and we've just added sectors. I've always been interested in everything, but now I've been able to add people that are experts. So I think the most exciting thing about semi-analysis is I don't know another firm that has the level of expertise and concentration we do. I have people who have worked at ASML, Applied Materials, Lam Research, the equipment companies that build wafers, all the way upstream to people who have worked at Intel and TSMC and Nvidia and Microsoft and Amazon, of course, but also people who have worked at OpenAI on models, someone who's worked at Tesla and FSD, someone who's worked at Cohere. We've got people who worked on the model layer and then on another vertical we've had people who have worked on data centers. There's someone at my company who built a power plant in Kazakhstan. It's like we just have this most insane talent density that's awesome. So half the firm is sort of people who have done engineering across the industry, and the other half is either ex-hedge fund or just random people from the internet who are super passionate. I found them on Twitter or Discord and I'm like, 'You're smart. Come work for me.' And it works. And so this is what's built up, and now SemiAnalysis has many different lines of businesses. Obviously data services, consulting, information services, we have the newsletter. We're doing all these media. We're having a conference soon that's going to be big. And all sorts of different stuff that we do. And it's just a hell of a ride.
说到这段旅程,我有一个时刻,Dylan。WisdomTree 和 SemiAnalysis 已经合作了好几个月。英伟达 GTC 每年三月举行。我坐在北卡罗来纳州夏洛特市看直播。我想还有 55,000 人在看直播。突然——体育场里有 20,000 人。体育场里有 20,000 人,老兄。他直接提到了你,说你像在‘放水’,你说他在某个数字上放水,你的图表就在舞台上。所以坦白说,我替你感到激动,看着全球最大公司的 CEO 引用你的研究,以及你如何批评他对某些数字的看法。我想请你——听起来你当时就在体育场里。
Speaking of the ride, I had a moment, Dylan. So WisdomTree and SemiAnalysis have been on this journey for a number of months working together. And so Nvidia GTC happens in March as it does every year. And I'm sitting watching the live stream here in Charlotte, North Carolina. And I think there's 55,000 other people watching the live stream. And suddenly—and there's 20,000 people in a stadium. There's 20,000 people in the stadium, dude. And he's referring to you directly, that you were like sandbag, you said he was sandbagging a certain number, and your charts are right up on the stage. And so admittedly, I had a moment I guess probably on your behalf, basically watching the CEO of the biggest company in the world basically refer to your research and how you were criticizing his take on some of the numbers. I'd love for you to sort of—I guess you were in the stadium from what it sounds like.
是的。那个时刻相当超现实。基本上,SemiAnalysis 做的事情之一是我们有一批工程师,我们对所有开源 AI 模型以及所有硬件进行开源基准测试。这是一个非常了不起的工作。我们这边有很多工程师,但我们与行业紧密合作。所以我们获得硬件——我们有超过 5000 万美元的硬件来自 OpenAI、微软、亚马逊、谷歌、CoreWeave、Nebius、Crusoe 等公司,你能想到的所有主要云服务商都捐赠过。甲骨文也捐赠了硬件,我们用它来运行这些基准测试。所以我们有八种不同的 GPU——所有 H100、H200、Blackwell、AMD 的各种 GPU。此外,我们还有谷歌的 TPU 和亚马逊的 Trainium。我们每天对最新版本的软件运行基准测试。原因是每晚都可能发布新的 CUDA 版本、PyTorch 版本、驱动更新、推理引擎 vLLM 版本等等。所以我们每晚都在整个曲线上运行这些基准测试,权衡你想要多快的 token 生成速度与多高的成本效益,在最优场景下。我们每天都运行所有这些。
Yeah. So that moment was quite surreal. So basically one of the things that SemiAnalysis does is we actually have a number of engineers and we do open-source benchmarking of all the AI models that are open source as well as all the hardware. It's a pretty awesome effort. So there are a number of engineers on my side, but we collaborate with the industry heavily. So we get hardware—we have over $50 million of hardware donated to us from companies like OpenAI, Microsoft, Amazon, Google, CoreWeave, Nebius, Crusoe, all of the major clouds you can think of that have donated to us. Oracle has donated us hardware that we run these benchmarks on. And so we have eight different kinds of GPUs—all H100s, H200s, Blackwell, AMD's various GPUs. In addition, we have TPUs from Google and Trainium from Amazon. And what we do is we run benchmarks on the latest version of software every single day. And why that is is because every night a CUDA version could release, a PyTorch version could release, a driver update could release, inference engine vLLM version could release, etc. And so we run these benchmarks every single night on the entire curve of how fast do you want the tokens versus how cost-effective do you want them on the optimal scenarios. And we run all of this every single day.
这就像一个自动化的基准测试套件。当初 Jensen 发布 Blackwell 时,声称会有 25 倍的提升。当时没人相信他。连我们都说‘哦,好吧’。我们比以往更乐观,根据模拟认为可能是 15 到 20 倍。我们有一个性能模拟器。但当我们构建这个名为 Inference X 的推理基准测试时,发现在 DeepSeek V3 上,Blackwell 在某些连续任务上比 Hopper 快 30 倍。所以我拿到结果就立刻给 Jensen 发了邮件。结果自动发布到了开源 GitHub 上;这是一个与 Nvidia 人员合作的开源项目。我特意告诉他:‘嘿 Jensen,2024 年你发布 Blackwell 时说 25 倍,所有人都吐槽你。连我都吐槽了。我说不可能 25 倍,最多 15 到 20 倍。但我错了。你故意保守了。实际上是 30 倍。’他接受了这个结果,但我不知道他会拿它做什么。我从几个客户那里听说,Meta 有人告诉我,在一次会议上 Jensen 用这个作为证据,证明他没有虚报数字。他当时在谈论下一代芯片。总之,我没想到他会直接在台上展示。在 Inference X 中,我们制作了一条 WWE 冠军腰带,上面写着‘推理之王’,并送给了所有合作者:Nvidia、AMD、SG Lang、VLM,以及其他帮助基准测试或捐赠硬件的人。这是一个开源项目,我每年花几百万美元在工程师薪资上,其他人则捐赠数百万的硬件或工程师薪资。我把腰带寄给了他们。他把腰带放在幻灯片上,举起来,旁边是我们的图表。他讲了五分钟,说‘Dylan 说我虚报,但我没有,我们的性能是最好的’。那感觉太不真实了。他在整个演讲中谈论我们的时间比任何人都长。唯一被同样提及的是 OpenClaw,它正在席卷世界。所以那是个不可思议的时刻。
It's like an automated benchmarking suite that runs. When Jensen originally launched Blackwell, he claimed it would be a 25x improvement. At the time, no one believed him. Even we were like, 'Oh, okay.' We were more bullish than ever, thinking it could be 15 to 20x based on our simulations. We have a simulator for performance. But as we built out this inference benchmarking called Inference X, we found that in DeepSeek V3, Blackwell is 30x faster than Hopper on some continuum. So I emailed Jensen as soon as we had the results. They automatically published onto the open source GitHub; it's an open source collaboration with Nvidia people. I highlighted to him, 'Hey Jensen, back in 2024 when you launched Blackwell, you said 25x and everyone gave you crap. Even I gave you crap. I said there's no way it's 25x, maybe 15 to 20x. But I was wrong. You were sandbagging it. It was 30x.' He took that and I didn't know he was doing anything with it. I heard from a couple customers; someone at Meta told me there was a meeting where Jensen used that as proof that he doesn't sandbag numbers. He was telling about the next generation chip. Anyway, I didn't expect it to happen on stage. In Inference X, we created this belt that looks like a WWE belt, saying 'Inference King', and sent it to all our collaborators: Nvidia, AMD, SG Lang, VLM, and others who helped with the benchmark or donated hardware. It's an open source effort where I spend a couple million dollars a year on engineer salaries, and others spend millions on hardware or engineer salaries donating to this open source effort. I sent them this belt. He had it on the slide, held it up, and there were our charts. He talked for five minutes about how Dylan said I was sandbagging but I wasn't, and our performance is the best. It was such a surreal moment. He talked about us longer than anyone else in the entire presentation. The only other person he talked about as much was OpenClaw, which is taking the world by storm. So it was an incredible moment.
是的 Dylan,你提到了几件事。我觉得很有意思,你提到了开源,现在可能正在转向一些最新发展和市场。关于开源模型与闭源模型的推理效率一直有讨论。而且直到今天,还有很多投资者在质疑这一切的投资回报率。就在过去一两周,一位彭博经济学家谈到,很多 AI 项目可能在一些公司中失败。我知道你强调过你的公司正在广泛使用 AI,给员工大量 token 访问权限,而且你们还在招聘。所以我很好奇:你对推动这一大规模建设、并伴随着各种限制(这些限制在过去一个月里推高了相关股票)的终端需求有什么看法?
Yeah Dylan, you mentioned a couple things. I think it's interesting that you mentioned open source and now there's maybe a transition towards recent developments and the markets. There have been discussions about the inference efficiency of open source models versus closed source models. And there are still many investors questioning the ROI on all this. Just in the last week or two, a Bloomberg economist talked about how a lot of AI initiatives may be failing at some firms. I know you've highlighted how your firm is using AI extensively, giving employees lots of access to tokens, and you're hiring. So I'm curious: what's your take on the end demand that drives this big buildout, with all the constraints that have been driving up stocks tied to these themes?
是的。我想说几点。当你审视投资回报率这个根本问题——公司是否从 AI 中赚到了足够的钱?这能持续吗?人们使用 AI 是否真的获得了价值?——有几种分析方式。首先,Anthropic 在第二季度实现了自由现金流为正并盈利。甚至在 4 月份,他们结账时就是盈利且自由现金流为正。5 月份同样如此。6 月份看起来也会一样。虽然还没完全结账,但三个月中有两个月都是自由现金流为正且盈利。他们的经常性收入已飙升至超过 500 亿美元的年度经常性收入(ARR)。他们做得非常出色。但这是一方面:Anthropic 在印钱。显然,还有很多公司没有印钱,但它们正在接近。OpenAI 的收入随着 Codex 的采用而开始增长。所以这些公司都在变得更加盈利。Anthropic 的毛利率非常高,超过 70%。
Yes. I would say a few things. When you look at the overarching question of ROI—are companies making enough money from AI? Will this continue? Are people using AI actually getting value?—there are a few ways to dissect it. First and foremost, Anthropic is free cash flow positive and profitable in Q2. Even in April, they closed April's books profitable and free cash flow positive. In May, they were free cash flow positive and profitable. June looks like it will be the same way. It's not fully closed yet, but for two of the three months, they've been free cash flow positive and profitable. Their recurring revenue has soared past $50 billion ARR. They're doing fantastic. But that's one side of the coin: Anthropic is printing. Obviously, there are a lot of companies that aren't printing, but they're getting there. OpenAI's revenue has started to inflect as Codex's adoption has grown. So these companies are all getting much more profitable. Anthropic's gross margins are really high, above 70%.
说到底,这只是硬币的一面。人们开始提及的另一面是:公司在 AI 上的支出到底如何?在 Semi-Analysis,我们的年度经常性支出——我喜欢叫它 ARS——去年 11 月,在 Claude Code 真正起飞之前,我们的年度经常性支出还不到 10 万美元。我们给每个用户都订阅了所有模型,或者说是 ChatGPT 的 200 美元套餐。差不多就这样。支出不到 10 万美元。如果有人想要 xAI 或 Claude,我们也会提供,但标准配置是给每人一份 200 美元的 OpenAI 订阅。那是 11 月的状况,我觉得当时我们已经处于最前沿了。但后来 Claude Code 随着 Claude Opus 4.5 和 4.6 等版本真正迎来了转折点。到 1 月底,我们的年度经常性支出已经达到了 400 万美元,因为大家都在用 Claude Code。今天大约是 1100 万美元。最高的一天——如果我们按一周乘以 52 周来算——曾经达到过 1400 万美元。这个数字会根据大家的工作内容大幅波动。目前,对于一个 90 人的公司,平均每年的支出大约是 100 万美元。这太疯狂了。我们在 AI 上的支出已经超过了员工支出的三分之一。到今年年底,我们可能会达到一半,具体取决于 Methos 和其他模型的进步。这是一笔巨大的支出。问题是:投资回报率如何?我认为回报巨大,因为我们能够构建产品、增加销售、提高整个公司的效率。我看到了回报,但很多公司都在质疑:如果我有一个优秀的开发者,年薪 30 万美元或更多,他们的 AI 支出正在接近一比一。对于非开发者,支出可能更低,但在 Semi-Analysis,我们最大的支出者中有很多是不会编程的人——他们只是告诉模型想要什么,然后不断迭代直到得到结果。所以你会看到人均支出飙升。很多公司都在问:我们全年的 AI 预算在第一季度或第二季度就花光了,现在怎么办?是削减 AI 支出还是削减其他方面?很多公司开始削减其他 SaaS 产品。他们说,‘我们可以增长得更快,所以就这么做吧。’他们说,‘在 AI 上花钱没问题;我们暂时承受这个冲击。AI 越来越便宜。’六个月前我做的事情,现在用 AI 做要便宜得多。当然,我现在用 AI 做的事情范围也广得多。有些人甚至在裁员而不是削减 AI 支出。有些人在收紧 AI 支出,但这些公司将在生产力提升方面被甩在后面。
Ultimately, that's one side of the coin. The side people are starting to allude to is what about the company's spending on AI? At Semi-Analysis, we went from our annual recurring spend—I like to call it ARS, annual recurring spend—in November, before Claude Code really took off for us in December last year, our annual recurring spend was less than $100k. We had a subscription to every model, or a subscription to the $200 tier for ChatGPT for every user. That was about it. We were spending less than $100,000. If people wanted xAI or Claude, we'd give it to them as well, but standard was giving everyone the $200 OpenAI subscription. That was the state in November, which I think was on the bleeding edge even then. But then Claude Code really started to hit its inflection point with Claude Opus 4.5 and 4.6, and so on. By the end of January, our annual recurring spend had hit $4 million, because people were using Claude Code. Today, it's about $11 million. The highest day—if we take a week and multiply by 52—we've ever had was $14 million. It oscillates a lot based on what work people are doing. Right now, the average looks to be about $1 million of spend per year for a 90-person firm. That's insane. We're spending more than a third of employee spend on top of that is AI. We'll probably get to half by the end of the year, depending on how Methos and other models improve. That's a huge amount of spend. The question is: what's the ROI? I think there's been huge ROI because we've been able to build product, sell more, and increase efficiency across the company. I see the ROI, but many companies are questioning: if I take a good developer making $300,000 a year or more, their AI spend is starting to approach one-to-one. For non-developers, the spend can be lower, but at Semi-Analysis, many of our biggest spenders are people who don't know how to code—they just tell the model what they want and iterate until they get it. So you see soaring spend per employee. Many companies are asking: our entire AI budget was blown through in Q1 or Q2; now what do we do? Do we cut spend or cut elsewhere? Many companies are starting to cut other SaaS products. They say, 'We can grow faster, so we'll just do it.' They say, 'It's okay to spend on AI; we'll take the hit temporarily. AI keeps getting cheaper.' For what I used to do six months ago, AI is much cheaper today. Of course, what I'm doing with AI today is much more extensive. Some people are even cutting employees instead of cutting AI. Some are clamping down on AI, but those companies will be left in the dust in terms of productivity gains.
明白了。缓解部分增量成本的一种方法是选择更便宜、有时智能程度较低的模型——不必总是处于最前沿。我很好奇:像你们这样的公司,会不会在某个点上决定,某些用例更适合用 DeepSeek V4 这类模型来完成特定工作,而对于需要更高智能的任务,则依赖成本更高的方案?这是不是计算的一部分?
Gotcha. One way to mitigate some of the incremental cost is choosing cheaper, sometimes less intelligent models—not always being at the leading edge. I'm curious: is there a point where firms like yours decide that some use cases are more optimal for using a DeepSeek V4 type model for certain work, while you might need to rely on things that cost more for tasks requiring more intelligence? Is that part of the calculus?
当然,这对一些人来说是计算的一部分。你需要把 AI 工作负载分成两类。一类是集成到现有流程中的 AI——比如客户发来一份文件,我检查其中的 XYZ,把文件输入模型,由模型检查。这种情况下,我只需要达到一定的质量水平,然后就可以停止改进模型,通过等待更新或更便宜的模型来降低成本,或者追求成本效率。我们看到 AI 模型的成本每年大约降低 60 倍。你达到某个质量水平,一年后成本就便宜 60 倍。DeepSeek 让人震惊,因为它比 GPT-4 便宜 600 倍。那是在 GPT-4 发布大约两年后,所以 60 倍乘以 60 倍是 360 倍,而实际达到了 600 倍。所以曲线上的某个点——不管是每年便宜 60 倍还是 90 倍——DeepSeek V3 对比 GPT-4 在两年内便宜了 600 倍。如果你有一个工作流并集成了 AI,你达到某个质量水平,然后就可以转向更便宜的方案。另一类工作是 AI 助手。在这种情况下,成本优化实际上不是转向更便宜的模型。
Absolutely, that's part of the calculus for some folks. You have to break out AI workloads into two types. One is AI integrated into a process I have—like when a customer sends me a document, I check it for XYZ, put it into the model, and the model checks it. There, I just need to hit some level of quality, and from there I can stop improving the model and start decreasing cost by waiting for newer or cheaper models, or cost efficiency. We've seen AI models improve at a rate of about 60x per year in cost. You take a quality level, and a year later it's 60x cheaper. People freaked out about DeepSeek because it was 600 times cheaper than GPT-4. That was about two years after GPT-4, so 60x times 60x is 360x, and it ended up being 600x cheaper. So somewhere on the curve—whether it's 60x cheaper per year or 90x—DeepSeek V3 versus GPT-4 was 600x in two years. If you have a workflow and integrate AI, you get to a quality level, then go cheaper. The other range of work is AI assistant. There, cost optimization is not actually going to a cheaper model.
成本优化往往意味着使用最新模型。比如,Claude 4.6 Opus 完成一项任务需要 10 万个 token,而且可能需要来回对话几次,耗时 10 分钟;而 Claude 4.8 Opus 只需要四分之一的 token,即 2.5 万个,可能只需要一次往返。所以实际成本更低,因为生成的 token 数量更少,我花费的时间也更少。因此,当我观察开发者或从事智力工作的人时,如何降低成本?实际上,不是通过使用更便宜的模型,而是通过将现有任务——那些有时需要与模型反复磨合才能完成的任务——交给更新、更先进的模型,现在它只需一次迭代或一次调用就能完成整个工作流,使用的 token 更少。我们从 4.6 Opus 到 4.7 Opus 看到的情况是,我的成本实际上先下降了一周,然后才飙升,因为人们用得越来越多。为什么飙升?因为人们需要适应新的工作流:之前的工作做完了,那就做更多。所以成本又上去了。同样,从 4.7 到 4.8,成本先下降大约一周到一周半,然后飙升,因为人们觉得“哦,现在我能做更多工作了”。所以必须将生产力与成本一起衡量。对于 AI 助手来说,token 效率非常重要。这大概就是 Anthropic 一直胜过 OpenAI 的原因:他们的模型 token 效率更高。实际上,OpenAI 的模型在边缘案例——顶尖科学、数学、代码——上往往能完成 Anthropic 模型无法完成的任务,但需要 3 倍的时间和 4 倍的 token,因此成本高得多,而且人与 AI 的反馈循环不够快,所以在客户感知上反而更差。你说“嘿,模型,做这个任务”,然后回来检查是否完成,这是一回事;你说“嘿,我有四个小时做这个任务”,无论是调用模型一次让它工作四小时,还是调用四次来回交互,哪种效果更好?事实证明,当有人的参与时,Anthropic 的反馈循环更快、更好,因为 token 效率更高。这就是我们仍然主要使用 Anthropic 产品的原因。有些任务人们确实用 OpenAI,通常是那些可以运行一整夜的任务,但大多数任务他们还是用 Claude Code。所以这是模型和 token 效率方面一个有趣的因素:成本很难拆解,但有些任务你会固定模型质量,等待模型降价;而另一些任务,你实际上就是想要最聪明的模型,因为它更便宜。
Cost optimization is often taking the newest model because the newest model—whereas Claude 4.6 Opus would take 100,000 tokens to do a task and it might take a couple turns, me talking to it back and forth, so it might take 100,000 tokens and 10 minutes of my time—Claude 4.8 Opus can do it in a quarter of the tokens, 25,000 tokens, and it might only take one back and forth. So the cost is actually less because the number of tokens being generated is less, and the amount of time I'm using is less. So when I look at a developer or someone doing intelligence work, how do I reduce the cost? Actually, it's not by using a cheaper model. It's by taking an existing task that could be done sometimes with the model after fighting with it, and going to newer and newer models, and now it's able to either just one iteration or one-shot the entire workflow. It's able to do it in fewer tokens. So what we've seen from 4.6 Opus to when 4.7 Opus came out, my cost actually fell for a week before it soared back up because people were using it more and more. And why did it soar back up? Because people were like, okay, they have to adjust to the new workflow, which is: the work I was doing is done, let me do more. So it goes back up. Likewise, when 4.7 came out to 4.8, the cost fell for about a week, a week and a half, and then it soared back up because people were like, "Oh, yeah, now I can do more work." So you have to measure the productivity alongside the cost. When it's an AI assistant, token efficiency is really important. This is sort of why Anthropic has been beating OpenAI: their models are more token efficient than OpenAI's. Actually, OpenAI's models on the edge cases—leading science, leading math, leading code—can often do a task that Anthropic models cannot, but they take 3x as long and 4x as many tokens, and therefore cost a lot more, and the feedback loop of human and AI is not as rapid, so it actually ends up being worse on a customer perception basis. It's one thing to say, "Hey model, do this task," and then you come back and check if the task is done. It's another thing to say, "Hey, I have four hours to do this task," and whether it's one call to the model and it does work for four hours, or it's four calls to the model and it goes back and forth—which one does it better? It turns out Anthropic, when you have this human-in-the-loop feedback loop, is actually way faster and better because it's more token efficient. So that's the main reason why we still remain a majority Anthropic shop. Some tasks people do use OpenAI, and often the tasks that they let run overnight are the ones they give to OpenAI codex, but most tasks they keep with Claude Code. So this is one of the interesting factors of what's going on with the models and token efficiency: cost is a bit hard to parse out, but some tasks you freeze the model quality and wait for the models to get cheaper, and others it's actually "I just want the smartest model because it is cheaper."
Dylan,我想听听你对硬件方面的看法。我知道今年早些时候,有一份新闻通讯——如果听众还没意识到的话,我多年来一直是它的忠实粉丝——里面有一篇文章谈到内存。内存通常是有周期的:可能 18 到 24 个月上涨,18 到 24 个月下跌。而且我们知道,现在几乎所有东西都短缺。如果你参与数据中心组件的供应,感觉问题不是你能不能拿到组件,而是你要等多久?因为如今世界上几乎任何组件都很难拿到。以你的经验,纵观硬件领域,你认为像内存这样过去一直是商品化产品的东西——你经历上涨,经历下跌,过去 40 年如此循环往复——将会发生什么变化?
Dylan, I was curious on your thoughts shifting it a bit to the hardware side. I know earlier this year, one of the newsletters—I've been a big newsletter fan for multiple years, if the audience hasn't realized it—but there was an article talking about memory. Memory has usually been a cycle: maybe it's 18 to 24 months you go up, 18 to 24 months you go down. And we know it feels like almost everything is in shortage. If you're involved in a component that goes into a data center, it feels like it's not a question of can you even get the component; it's more, how long are you going to have to wait? Because it feels like in the world today you can barely get any component. So with your experience having looked across the hardware side, what do you expect is going to be changing with something like memory, which used to always be this commoditized product—you ride the upswing, you go through the downswing, and it just repeats going back the last 40 years?
是的。我不是说以后不会有周期了。我认为周期还会发生。显然,我们正处于一个超级周期,上涨势头疯狂,也会有一些下跌,而且会很残酷。但下跌周期,从谷底到谷底,仍然有很大的增长。所以我认为现在内存和其他组件的关键点是正在发生的转变。历史上,上涨周期终端市场增长 50%,因此对于内存这类定价弹性较大的商品市场,股票会涨 2 到 3 倍。但如今,我们看到的不是增长 50%,而是过去几年支出已经翻倍,而且还会再翻倍。总支出翻倍了,当你观察不同终端市场的弹性时,内存的价格已经上涨了约 4 倍,而且还会再涨 2 到 3 倍,此外还有产能增长。所以股票在下跌之前疯狂上涨。我认为内存真正令人兴奋的地方在于,它不仅仅是终端市场的问题——不仅仅是“市场火爆,商品弹性大,内存是商品,因此价格随终端市场需求弹性变化”。真正有趣的是,我们在 2024 年 o1 发布时就写过——OpenAI 发布了 o1,这是第一个推理模型——它催生了推理模型的新热潮,OpenAI、Anthropic、DeepSeek 等许多公司都在利用这些模型实现长周期智能体任务。所以当我们观察这一点时,有趣的是,o1 一发布,我们立刻注意到工作负载发生了巨大变化。以前做聊天时,与 ChatGPT 对话,你可能会发送一个 50 字或 500 字的提示,然后它会给你回复,这个比例——上下文长度——是几千。所以上下文长度大概在 2000 左右。
Yeah. So I'm not saying there's not going to be cycles anymore. I think cycles will happen. Obviously, we're in a super cycle where the upswing is crazy, and there will be some downswing and it'll be brutal as well. But the downswing, obviously, trough to trough, there's still a lot of growth. So I think what's relevant now about memory and other components is the shifting faces of what's happening. Historically, we had upcycles would be up 50% for the end market, and therefore for commodity markets like memory where pricing is more elastic, you'd end up with those stocks 2-3x. But what we've gotten today is, instead of up 50%, we've gotten over just the last few years, spend has already doubled and it's going to double again. So total spend has doubled, and when you look at the elasticity of different end markets, memory we've had the pricing go up like 4x and it's going to go up another 2x-3x again, in addition to capacity growth. So you've got the stocks just ripping like crazy before going back down. So I think what's really exciting about memory is it's not just an end market thing—it's not just "oh the market's ripping and it's a very elastic good and memory is a commodity and therefore the pricing is very elastic with the end market demand." What's actually interesting, and this is something we wrote in 2024 when o1 came out—OpenAI released o1, which was the first reasoning model—and it created a new boom of reasoning models that OpenAI, Anthropic, DeepSeek, and many others have been exploiting to get the models to go long-horizon agent tasks. So when we look at that, what's interesting is when o1 came out, the immediate thing we noticed is the workload changed dramatically. When we were doing chat, talking to ChatGPT, you may send a prompt that is maybe 50 words or 500 words, but you're going to send a prompt and it's going to give you a response back, and that ratio—the context length—is a few thousand. So you might have a context length of let's call it 2,000.
这意味着,当你运行推理时,每次生成一个 token,你都要把所有权重读入芯片,把所有上下文读入芯片,处理一个 token,然后再次迭代。你读取所有 token、上下文和权重。上下文被称为 KV 缓存,对吧?它创建了所有这些 token 之间的关系。有趣的是,当你运行模型推理时,在权重方面,如果上下文长度是 1000 或 100,000,你仍然需要读取所有权重。所以,推理时的内存强度在权重方面是相同的,但在 KV 缓存方面,当你读取 1000 个 token 与 100,000 个 token 时,内存差异巨大,尽管计算量大致相同。由于 KV 缓存的内存缓存等因素,计算成本不会飙升,但内存成本会飙升。因此,我们在 01 报告中强调,2024 年 12 月我们讨论了缩放定律,以及预训练缩放定律如何让位于推理缩放定律,而 01 是一个巨大的阶跃变化。我们讨论了 KV 缓存如何因推理而爆炸,因此内存将成为最大赢家。我们在 2024 年 12 月做了这个预测。在 2025 年,我们多次对内存感到兴奋。但在 2026 年 1 月,我们写了一份报告,当时人们说,好吧,内存已经涨了 50%。这是周期的顶部吗?我们需要继续吗?我们写的报告基本上是:不,不,不。我认为你们不明白,对吧?内存容量在未来三年每年仅增长 20-30%。然而,需求却在翻倍。所以最终会发生的是内存价格将持续飙升。对价格弹性较低或适应能力较差的用户将退出市场。智能手机、笔记本电脑——因为成本飙升太多,它们将退出市场,这一切都将让位于 AI。这意味着价格将不得不持续飙升,直到这种情况发生,因为容量增长不够快。所以我们的观点是,内存不是短缺,这不是短期短缺。这是一个将持续数年的短缺。因此,我们在第一季度和第二季度看到的是内存表现强劲。它一直在飙升。有些日子它因随机原因下跌 7-8%,但最终图表是向上向右的。我们认为它将继续飙升,因为价格持续上涨。我们还没有看到——一些中国中低端智能手机制造商,如小米,表示其出货量下降了 40%。但我们还没有看到高端市场受到影响。明年 iPhone 价格必须上涨。明年 MacBook 价格必须上涨。现在,如果 MacBook 或 iPhone 价格上涨 100 美元,市场不会调整太多。但内存将变得越来越昂贵,直到 AI 得到满足。这意味着智能手机价格不仅会上涨 100 美元,它们将不得不涨几百美元。所以,在某个时候会达到一个平衡,AI 获得所需的需求,而移动和消费硬件被压缩到足够低。但显然,在某个时候人们仍然需要新手机和新笔记本电脑,所以他们仍然会购买。因此,我们将不得不达到一个新的平衡,因为内存的终端市场容量增长不够快。随着我们扩展到整个生态系统,真正重要的是许多不同的组件都短缺。谁有弹性?谁有价格弹性,谁没有?例如,台积电在定价上没有弹性。他们是一家相当不错的公司,对客户相当公平,长期合作。他们表示,我们会提价 5-10%。内存公司处于商品市场。他们让现货市场和合约市场的供需平衡真正调整价格。所以你看到价格有两到三个轴,总有一天价格会减半,对吧?因为内存不一定值得 85% 的利润率,尽管它正朝着这个方向发展。我们还没有达到内存 85-90% 的毛利率,但我们会达到的。然后在某个时候,它会从那里减半回到 70% 左右甚至更低。因此,我们将看到内存的这种波动。在台积电,你看不到这么大的波动。在其他领域,如 ASML,我们在定价上看不到太多波动。他们制造设备,但生态系统的不同部分会根据终端 AI 需求流向他们的程度而不同地波动。供应链的不同部分,每花在 AI 上的 1 美元,可能只有 1 美分用于这个产品,但可能有 5 美分用于那个产品。所以显然,基础设施供应链中的不同终端市场将受益不同。但就需求而言,此外,那里的市场动态是什么?是垄断还是寡头垄断?是竞争激烈的大市场?是定价相当稳定且有很多长期协议的市场?还是相当商品化且定价基于供需的市场?所有这些因素决定了特定终端市场——无论是内存还是现在人们谈论的 MLCC 短缺、PCB 钻头短缺、PCB 箔、铜箔短缺,或所有这些随机组件。你会在网上看到这是下一个短缺,这是下一个短缺。你会看到所有这些不同的东西,但重要的是实际的需求流量有多大。这个终端市场是翻倍?增长 50%?还是翻四倍?以及基于市场结构,价格会上涨多少?
And what that means is, when you're running inference, every time you generate a token, you read all the weights into the chip, you read all the context into the chip, you process a token, and then you iterate again. You read all the tokens, the context and weights. The context is called the KV cache, right? It creates this relationship between all these tokens. What's interesting is, when you're running a model inference on the weights side, if the context length is a thousand or if the context length is 100,000, you still have to read all the weights. So, memory intensity on inference is the same on the side of the weights, but on the side of the KV cache, the memory when you have a thousand tokens that you're reading in versus 100,000 tokens is a humongous difference even though the compute amount is roughly the same. The compute amount, because of KV cache caching on memory and things like that, you can sort of get away with it—your compute costs don't soar, but your memory costs soar. So what we highlighted in our 01 note was, in December of 2024 we talked about the scaling laws and how pre-training scaling laws were giving way to reasoning scaling laws, and 01 was a big step function change. We talked about how KV cache was going to explode because of reasoning, and therefore, memory was going to be the biggest winner. And so, we did that in December 24. And multiple times in 25, we were really excited about memory. But in January of 26, we wrote the note which was talking about, at the time people were like, okay, memory has gone up 50%. Is it the top of the cycle? Do we need to keep going? And we wrote a note that was basically like, no, no, no. I don't think you guys get it, right? Memory capacity is only growing 20-30% a year for the next three years. And yet, demand is doubling. And so what's going to end up happening is memory prices are going to keep soaring. Users of memory who are less elastic or less capable of adapting to the elasticity of pricing will drop out of the market. Smartphones, laptops—because the costs are going to soar so much, they're going to drop out of the market, and that's going to all give way to AI. And what that means is the price is just going to have to soar and soar until that happens, because capacity is not going up enough. So ultimately our point there was memory isn't a shortage, and this is not a short-term shortage. It's a shortage that's going to last years. And so what we've seen so far over the last Q1 and now Q2 is memory has just been gangbusters. It's been soaring. There have been days where it's gone down 7-8% for some random reason, but ultimately the chart has been up and to the right. And where we see it continuing to go is it's going to continue to soar because pricing continues to go up. We still haven't had—we've had some Chinese smartphone makers in the mid-range and low end like Xiaomi say their shipments are down 40%. But we haven't seen the high-end market get impacted yet. And next year iPhone prices have to go up. Next year MacBook prices have to go up. And right now, if MacBook prices or iPhone prices go up a hundred bucks, that market's not going to adjust too much. But that memory is going to keep getting more and more expensive until AI gets its fill. Which means smartphone prices aren't just going to go up 100 bucks. They're going to have to go up a few hundred bucks. And so, at some point there's going to be an equilibrium where AI gets the demand it needs and mobile and consumer hardware gets pushed down enough. But obviously at some point people still need new phones and new laptops, so they'll still buy. So we're going to have to reach a new equilibrium because this end market capacity of memory doesn't grow fast enough. And as we extend across the ecosystem, what really matters is a lot of different components are in shortage. Who is elastic? Who has an elastic price and who doesn't? An example is TSMC is not elastic on pricing. They're a pretty good company, pretty fair with their customers, partner long term. They're like, we'll take up price 5-10%. Memory companies, they're in a commodity market. They let the spot market and contract market supply-demand balance really adjust pricing. And so you see these two to three axes in pricing, and someday you'll see pricing half, right? Because memory necessarily doesn't deserve a margin of 85%, which is where it's headed to though. We're still not at 85-90% gross margins for memory, but we'll get there. And then at some point from there, it'll also half down back to like 70s or maybe even lower. And so we'll see this oscillation in memory. In TSMC, you don't see so much oscillation. In other areas like ASML, we don't see much oscillation on the pricing. They make equipment, but different parts of the ecosystem will oscillate differently based on how much of the end AI demand flows through to them. Different parts of the supply chain are going to have, for every dollar spent on AI, it might be one cent on this product, but it might be 5 cents on this product. So obviously there's different end markets in the infrastructure supply chain will benefit differently. But then in terms of demand, in addition, what are the market dynamics there? Is it one where there's a monopoly or an oligopoly? Is there one where there's a very competitive large market? Is there one where the pricing is pretty stable and it's a lot of long-term agreements? Or is there one where it's quite a commodity market and pricing is based on supply-demand? And so all of these factors determine whether a specific end market, whether it be memory or now people are talking about shortages of MLCCs or PCB drill bits or PCB foil, copper foil, or all these random components. You'll go online and see this is the next shortage, this is the next shortage. You'll see all these different things, but what matters is how much flowthrough of demand is there actually. Is this end market doubling? Is it going up 50% or is it quadrupling? And how much is the pricing going to go up based on the market structure?
这些才是真正决定基础设施供应链走向的因素。如果你套用这个框架,感觉每年市场都会突然意识到你所说的那种新的“潜在短缺”。今年早些时候,OpenAI 的爆火让各个平台的人们开始关注 AI 智能体及其各种可能性。用你刚才描述的框架,我很好奇你对 CPU 市场的看法——在 AI 的前三年里,我几乎没听过“CPU”这个词,但今年到处都在提 CPU。
These are what really determines what happens in the infrastructure supply chain. And if you take that framework, because it feels like year by year the market wakes up to exactly what you said, a new quote unquote potential shortage. So earlier this year we have the open claw virality on various sites awakens people to the world of AI agents and all the possibilities. Take the framework that you just described and I'd be curious to hear your take on the CPU market, which for the first three years of AI I don't think I heard the word CPU, and this year I'm hearing CPU everywhere.
是的。关于 CPU,有趣的是,去年 11 月我们在为客户做机构研究时就开始大量讨论它,因为 OpenAI 和 Anthropic 已经开始与亚马逊、谷歌、微软等公司达成协议,购买他们所有机群中的 CPU 并租用出去。然后从去年年底到今年,CPU 需求一直在攀升。原因是什么?我们先说原因。AI 最初在训练和推理阶段——推理大多是短上下文——主要依赖算力和网络。但随着预训练转向强化学习,以及聊天式推理转向智能体式,我们迎来了一个重大转折:现在对 CPU 的需求更大了。为什么?从预训练到强化学习,预训练是将整个网络数据集训练到模型中。而强化学习是模型生成一些合成数据或推理轨迹,然后对照环境进行检查。这个环境可能是对代码运行单元测试,可能是一个类似网站的沙箱,也可能是一个类似工程系统或其他平台的沙箱——无论是网站、购物网站还是其他什么。它可能是你在互联网上使用的东西,比如编译代码。这些环境需要大量 CPU,而之前的预训练中,实际处理 token 并不需要太多 CPU,主要是环境检查。对吧?“我生成了这些 token,它们有效吗?在 Python 或 C 编译器里看起来怎么样?或者在一个网站里,如果我想买东西——电商之类的。”作为一个智能体工作流,我不断测试这些东西,这需要大量 CPU。另一方面,现在进行实时推理时,聊天模式是:“我告诉它一些东西,它给我一个答案,结束。我可能会再问几个问题,但仅此而已。”但现在我谈论的是智能体工作流,模型会进行工具调用——比如:“我要去搜索这个。哦,我要查一下数据库。哦,我要去问 Python 解释器,写一点代码来检查我的工作。哦,我要写一些代码,编译并部署它。”这些智能体流程最终需要越来越多的 CPU,因为它们必须与真实世界交互。人类与模型交互是一回事:我告诉模型一些东西,模型给我回应,我读一下,然后“复制粘贴到任何地方”。但当模型与互联网世界交互时,情况就不同了。这会导致循环中有更多的算力、更多的 AI——或者更确切地说,更多的 CPU——在来回传递答案。所以无论是强化学习还是智能体工作流,都需要大量 CPU。
Yeah. So on the side of CPUs, what's interesting is in some of our institutional research for our clients in November last year, we started talking a lot about it, and that's because OpenAI and Anthropic had started striking deals with Amazon, Google, Microsoft, etc. on buying all the CPUs they had in their fleets, renting them out. And then over the course of late last year and now this year, CPU demand has just been inflecting. And the reason why—let's talk about the reason first. So AI initially, when it was training and inference, and inference was mostly a short context thing, it was mostly just predicated on compute and networking. But as pre-training shifted to reinforcement learning, and as chat-style inference turned into agentic, we had this big inflection: now CPUs are demanded more. Now why is that? In the case of pre-training to reinforcement learning, pre-training is all about training the entire web dataset into your model. Whereas reinforcement learning is the model generates some synthetic data or generates a reasoning trace and then checks it against an environment. That environment may be running unit tests on code. That environment may be a sandbox that looks like a website. That environment may be a sandbox that looks like an engineering system or some other platform that you would use—whether it be a website or shopping site or what have you. It would be something that you utilize on the internet, might be compiling the code. And those environments require a lot of CPUs, whereas before, pre-training the actual processing of tokens didn't require much CPU; it was all the environment checking. Right? "I've generated these tokens, now are they valid? What do they look like inside a Python or a C compiler, or within a website if I'm trying to buy something—e-commerce, whatever it is." As an agentic workflow, I'm testing these things constantly, that requires a lot of CPU. And then the flip side is when you're doing live inference now. When you're doing chat, it's like: "Okay, I tell it something, it gives me an answer back, done. I might ask it a few more questions, but that's it." But now when I talk about agentic workflows where the model is making tool calls—so it's like: "Okay, I'm going to go search for this. Oh, I'm going to look this up in a database. Oh, I'm going to go ask the Python interpreter and I'm going to write a little bit of code to check my work. Oh, I'm going to write some code and compile it and deploy it." These agentic flows end up requiring more and more CPU because they have to actually interact with the regular world. It was one thing if the human is interacting with the model: I'm telling the model something, the model gives me a response, I read it, I'm like, "Okay, copy and paste into whatever it is." It's a different thing when it's the model reacting with the internet world. And there ends up being a lot more compute in the loop, a lot more AI in the loop—or a lot more CPUs in the loop, sorry—that are bouncing back and forth the answers. So both of these, whether it be reinforcement learning or agentic workflows, need a lot of CPU.
那么现在发生了什么?因为我们需要大量 CPU,但让我们用框架评估一下之前的情况。市场结构如何?市场上有几家参与者:英特尔和 AMD,现在 ARM 也发布了 CPU,因此 ARM 的股票因此大涨——他们是市场的新进入者,看起来相当有竞争力。还有亚马逊,它是这个领域的领导者,但微软和谷歌也在发布他们内部开发的 CPU。还有英伟达,他们也在发布自己的 CPU。所以市场上有许多不同的竞争者,但直到两年前,整个市场还只有英特尔和 AMD。现在亚马逊已经获得了相当大的份额。英伟达和 ARM 也开始获得更多份额。但最终在终端市场上,英特尔实际上能够提高价格,AMD 也能提高价格。所以他们都提价了,需求也大幅上升。亚马逊能够从 CPU 中榨取惊人的利润,因为他们不是制造并销售 CPU,而是制造并出租 CPU。所以他们的 Graviton CPU 租用火爆,订单量大幅增加。英伟达以前只销售与 GPU 捆绑的 CPU,现在通过 Vera 独立销售 CPU。他们给出了 200 亿美元的 CPU 收入指引。对英伟达来说,这其实只是皮毛——大概几个百分点的增长。不,我开玩笑的。但我认为,当你看到英特尔、AMD、ARM、亚马逊这些公司——谁获得了收入,而不仅仅是销售收入——那里正在发生巨大的变化。
Now what has ended up happening? Because okay, we need a lot of CPU, but let's evaluate the prior things in our framework. What is the market structure? Well, there are a few people in the market: there's Intel and AMD, and now ARM is releasing a CPU, so ARM stock has gone gangbusters because of that—they're a new entrant into the market that looks pretty competitive. And then you've got Amazon, who's the leader of this, but also Microsoft and Google releasing their own CPUs that they've developed internally. And then you've got Nvidia who's releasing their own CPU. So you've got a lot of different competitors in the market, but really, up until two years ago, all of the market was Intel and AMD. And now Amazon's gotten a good amount of share. Now Nvidia and ARM are starting to get more share. But ultimately, what happens in the end market is: Intel's actually able to increase their price. AMD is actually able to increase their price. And so they've both increased their pricing. They've obviously gotten demand to go up a lot. Amazon is able to extract incredible margins out of CPUs because they don't make them and sell them; they make them and rent them. So their Graviton CPUs are renting like crazy, and they've increased the orders massively. Nvidia, who was previously only selling CPUs attached to their GPUs, is now selling CPUs standalone with Vera. And so they've given the guidance of $20 billion of CPU revenue. Now for Nvidia, that doesn't really scratch the surface—it's like a few percentage growth. No, I'm just kidding. But I think when you look at other companies like Intel, AMD, ARM, Amazon—who gets the revenue instead of just sales revenue—there are huge things happening there.
Dylan,也许基于刚才关于 CPU 的讨论,我听到的一些讨论是,用于智能体的 CPU 在某些方面与传统的 CPU 不同。我记得 Jensen 在谈到 Vera CPU 时说过或暗示过,其核心针对智能体活动进行了优化。此外,还有很多关于 GPU 与 CPU 比例的讨论,这显然突出了 CPU 需求的方向。你能就这两个话题多谈一些吗?因为我认为这个概念在高层面上对人们来说很有道理,但有一些技术细节可能被掩盖了。我不确定这仅仅是营销还是确有其事。
Dylan, maybe on the back of that with CPUs now, some of the discussion at least that I've heard has been that the CPUs for agents are different than say the historical CPUs in some regards. So the cores are more optimized for agentic activity, is what I remember hearing Jensen kind of saying or implying around the Vera CPU. And then there's also a lot of discussion around this GPU to CPU ratio, which obviously highlights maybe the direction of the demand and need for CPUs. Can you maybe give us a little more color on each of those topics? Because I think the concept makes a lot of sense at the high level to people, but then there are some technical things that are probably pushed under the rug if there are any. I'm not sure if it's just marketing or if there's a reality to this.
所以说到智能体式工作流,CPU 的使用情况差异很大。有些智能体工作流是这样的:模型运行后,我把所有 token 发送给某个 CPU 工作流,然后等待 CPU 处理,再送回模型继续工作。问题是,模型运行的算力在等待 CPU 时是否会停顿?有些情况会,有些不会。在停顿的情况下,运行模型的算力会闲置等待 CPU 响应,这时 CPU 的架构就需要非常不同。基本概念是:我想要更多核心还是更快核心?CPU 架构中有个规律:如果你把 CPU 核心做大一倍,意味着芯片上核心数量减半,但每个核心的性能并不会翻倍,可能只提升 50%。显然工程上有很多细节,权衡没那么简单,但简化来说,这是个简单的思考方式。
So when it comes to agentic workflows, use of CPUs varies a lot. You have some workflows in agentic where it's like, okay, if my model's running and then I send a response to send all the tokens to some CPU workflow and then I'm waiting on the CPU to do something and then I send it back to the model and the model works some more. The question is, did the compute that the model's running on stall while you're waiting for the CPU? In some cases it does, in some cases it doesn't. In the cases where it does stall, the compute that's running the model just stalls waiting for the CPU's response, then the CPU needs to be architected very differently. The basic concept is, do I want more cores or do I want faster cores? There's sort of a law within CPU architecture which is basically if you make the CPU core twice as big, which means I have half as many CPU cores on the chip, my performance doesn't go up 2x per CPU core. My per CPU core performance may only go up 50%. Obviously there's a lot of engineering and the trade-off is not that simple, but to simplify, that's a simple way to think about it.
如果看 Nvidia 为 Vera 设计的 CPU,它不到 100 个核心,但这些核心比 AMD 的核心更快。AMD 领先的 CPU 有 256 个核心。所以核心数量差距很大,但 Nvidia 的核心更快,不过并没有 AMD 核心的两倍快。因此人们在设计空间上做权衡。对于某些工作负载,AI 算力在等待 CPU 时必须停顿,那么——谁在乎我核心减半且只快 50%?总体 CPU 性能更低,但单核性能更高,因此我不需要频繁等待 CPU 核心。我不需要超并行工作负载,我真正需要的是这个工作负载立刻完成。所以在 AI 算力停顿的情况下,我想要尽可能快的核心,愿意牺牲多核性能。这就是某些类型的智能体工作流。
If I look at an Nvidia CPU for Vera, it has less than 100 cores, but those cores are faster than the AMD cores. AMD's leading CPU has 256 cores. So you've got this big delta in number of CPU cores, but the Nvidia core is faster, but it's not twice as fast as an AMD CPU core. So there's this trade-off that people are making on the design space. For some workloads where the AI compute has to stall while waiting for the CPU, then you need—who cares if I have half as many cores and they're only 50% faster? In total the performance of the CPUs is less but the per core performance is higher, and therefore I'm not waiting on the CPU cores as often. I don't need a super parallel workload; what I really need is this one workload done now. So in that case where the AI compute is stalling, I want to have the fastest core possible and I'm willing to sacrifice multi-core performance. That's some types of agentic workflows.
其他类型的智能体工作流:如果说到我日常如何使用 Claude,或者团队如何使用 Claude,我们每年在 Claude 上花费 1100 万美元(按 ARS 计算)。那是什么情况?我在调用 Claude,Claude 处理大量 token,但他们不只服务我一个人。他们把所有算力上的数十万用户批量处理。所以如果我得到响应后,它等待我去执行,无论是等待我还是某个 CPU 核心去执行,这没问题,因为计算机仍在运行,只是不为我运行,而是为其他人运行。所以如果 CPU 较慢但核心数量更多,那就是不同类型的任务。
Other types of agentic workflows: if I talk about how I use Claude day-to-day or how the team uses Claude, how we spend $11 million a year on Claude on an ARS basis. What is that? Well, I'm calling Claude. Claude is processing a bunch of tokens, but they're not just using only me. They're batching hundreds of thousands of users together across all of their compute. So if I get the response back and now it's waiting on me to implement it, whether it's waiting on me or a CPU core to implement it somewhere, you end up with—that's okay because the computer is still running, just not for me. It's running for other people. So if the CPU is slower but I get way more of them, it's a different sort of task.
另一个问题是:这是 AI 的主动使用,还是使用 AI 生成的内容并部署它?美妙之处在于,如果我们看全球 GitHub 提交量,比去年增长了好几倍。不是增长 10% 或 50%,而是好几倍。这意味着大量代码被生成并部署到世界各地。很多代码是粗制滥造的,但也有很多被部署。当代码被部署时,它运行在 CPU 上。这些是标准代码——可能只是一个网页爬虫,可能是一个分析引擎,可能是业务流程自动化。这些不一定需要超快的 CPU 核心,可以在成本效益高的 CPU 核心上运行。
Another one is: is it the active use of AI, or is it the use of what AI generated and then you're taking what AI generated and deploying it? The beauty is, if we look at GitHub commits globally, they're up multiplefold versus last year. Not just 10% up or 50% up, they're up multiple X. What that means is all this code is being generated to the world and people are deploying a lot of the code. A lot of the code is slopped, but a lot of code is being deployed. And when it gets deployed, it's being put on CPUs. And it's standard code—it might just be like a web scraper, it might be like some analytical engine, it may be some business process automation. That doesn't necessarily need to be on a super fast CPU core. It can be on a cost-effective CPU core.
当你审视这个连续谱时,Nvidia 构建了性能最高的 CPU 核心,但这不一定意味着——如果我有芯片,最大核心数乘以单核性能是多少?他们在这方面其实并不出色。而如果看 AMD 和 Amazon,他们有更多核心,数百个,但单核性能较低。ARM 也在那一端。那么在这个连续谱中你想选哪一端?有些工作负载你确实需要 Vera,有些则需要 Graviton 或 AMD CPU。我不认为事情那么简单。
When you look at the continuum, Nvidia has built the highest performance CPU core, but it's not necessarily giving you—if I have a chip, what is the maximum number of CPU cores times the performance per core? They're actually not so great at that. Whereas if I look at AMD and Amazon, they have a lot more cores, hundreds, but they have less per core performance. And ARM is on that end too. So where in that continuum do you want to go? For some workloads you do want Vera, and for some workloads you want the Graviton or the AMD CPU. I wouldn't say it's as simple.
至于你提到的另一个问题,即比例,CPU 需求上升是无可争议的。我们是去年年底在机构研究报告中第一个指出这一点的,今年 1 月又在我们的新闻简报中提及。自我们发布以来,一些 CPU 股票大涨。ARM 涨了好几倍。Intel、Nvidia——Intel 涨了好几倍。AMD 也大涨。这些股票都涨了。但现在,根本不了解技术的卖方分析师只是在胡编乱造,他们声称 CPU 与 GPU 的比例,或者 CPU 与 AI 算力的比例已经失衡到 CPU 比 AI 算力更受青睐的程度。这是错误的。
As far as the other question you mentioned, which is ratio, it is indisputable that CPU demand is going up. We were the first to call it out late last year in our institutional research and in January this year in our newsletter. Since we published that, some of these CPU stocks have ripped. ARM has gone up multiple X. Intel, Nvidia—Intel's gone up multiple X. AMD's ripped. These stocks have ripped. But now the sellside, who doesn't really understand technology at all, is just making up stuff and is getting to the point where the ratio of CPUs to GPUs or ratio of CPUs to AI compute is getting lopsided to the point where it's more in the favor of CPUs than it is AI compute. That's false.
重申一下,如果你看 Blackwell 全配置,每颗芯片大约 5 万多美元。如果比例是 1:1,CPU 大约 5000 美元。那么,如果 Blackwell 销售额达到 3000 亿美元或 5000 亿美元,CPU 销售额只有 300 亿或 500 亿美元。这是人们忽略的另一点。是的,这个终端市场在暴涨。但最终,大部分资金仍然流向 AI 算力和内存。这个市场之前被低估了,现在定价更合理。我认为人们需要认识到:CPU 的需求并不会持续增长到超过 AI ASIC 的程度。这更像是一次规模调整。
Just to reiterate, if you look at a Blackwell full all-out Blackwell, it's like $50,000 something per chip. If you had a 1:1 ratio, CPUs cost like $5,000. So you end up with, okay, well then for $300 billion of Blackwell to sell or $500 billion of Blackwell to sell, you would only get $30 billion or $50 billion of CPU sales. So that's another thing that people are missing. Yes, this end market is ripping. Ultimately, the majority of the dollars are still going to AI compute and memory. This market was underpriced and it's more fairly priced now. I think that's something that people need to recognize: it's not like CPUs are going to keep growing in demand beyond that of AI ASICs. It's a bit of a right sizing.
基本上在 2023 和 2024 年,我们卖出了数百万颗 AI 芯片,但 CPU 很少。现在突然 CPU 需求转向——比例不应该在这里,而应该在那里。人们处于追赶模式。所以现在我需要购买大量 CPU,以追赶历史上购买的所有算力,同时还要加上当前购买的算力。一旦我追上了之前购买的所有 AI 芯片的积压,并为它们配齐 CPU。
Basically in '23 and '24, there were years of selling millions of AI chips and very few CPUs. Now all of a sudden the CPU demand has inflected to—the ratio shouldn't be here, it should be here. And people are in catch-up mode. So now I need to buy a bunch of CPUs to catch up all the compute I've historically bought, as well as in addition to the compute I'm currently buying. Well, once I catch up that backlog of all these AI chips that I had bought previously and I catch them up for CPUs.
嗯,那种需求已经不存在了,对吧?我已经追上了,现在只剩下增量部分。所以,你想一下,如果比例是 1 个 CPU 对 2 个 GPU,每个 GPU 大概 5 万美元,每个 CPU 大概 5000 美元。那么,每花 10 万美元在 GPU 上,我只花 5000 美元在 CPU 上。这对 CPU 增长来说,其实并不是一个很好的市场动态。对 CPU 来说仍然很好,比以前好多了。但反过来看,如果过去三年我出货了 1000 万个没有配 CPU 的 AI6 GPU,那现在这 5000 美元就有巨大的追赶空间,对吧?所以我们正在经历的就是一个巨大的追赶过程,比例上移了,积压订单正在被消化,需求看起来非常疯狂,但最终会趋于稳定。我们正处于一个 CPU 的小周期中。
Well, that demand isn't there anymore, right? I've already caught it up and now it's only the incremental. And so, if you think about it, if there's a ratio of, let's say, one CPU to two GPUs, and each of those GPUs cost, let's say, $50,000 and each of those CPUs cost $5,000. Well, then for every $100,000 I'm spending on GPUs, I'm spending $5,000 on CPUs. And that's actually not that great of a market dynamic in terms of CPU growth. It's still great for CPUs. It's way better than it used to be. But if you flip back to, well, what happens if I have 10 million GPUs in AI6 that I shipped over the last three years that don't have any CPU attached really. Well, now that $5,000 has a huge catch-up, right? And so that's what we're experiencing right now is a huge catch-up as well as the ratio has shifted up and then that huge backlog is getting caught up and so you're seeing demand be ridiculous, but it will sort of subside and then it'll be at a steady state, right? We're in sort of a mini cycle of CPU.
不,这个背景非常好,非常有帮助。那么也许我们转向网络,进入堆栈的另一个领域。我认为这引起了很多投资者的关注,尤其是当他们深入研究光学供应链及其中的一些限制时。我们看到一些估计认为,共封装光学虽然被广泛讨论,但实际部署可能要等到 2028 年左右,2027 或 2028 年。当你思考“能用铜就用铜,不得不用才用光”这个概念,以及从光学到铜的转变时,还有黄仁勋在 Computex 上也谈了很多。还有很多其他讨论让 Marvell 这样的公司备受关注。你对光学有什么额外的想法吗?你认为未来两年数据中心在网络领域的架构会如何演变?
No, that's excellent context. Really really helpful. And then maybe just moving to networking to kind of move into another area of the stack. I think this is one that's come to a lot of investors' attention, particularly as they dive down into the optics supply chain and some of the constraints there. And we're seeing some estimates that co-packaged optics is maybe something that's talked about a lot, but is really probably going to be deployed 2028 or so, 2027, 2028. As you think about the concept of using copper when you can, optics when you must, and this kind of transition from optics to copper. And then we also had Jensen talking a lot about it at Computex as well. And a lot of other discussions bringing a lot of attention to firms like Marvell. Is there any additional thoughts you have around optics and how you see the architecture of the data center within the networking domain evolving over the next two years?
是的,很明显,随着模型变大,我们如何跨模型运行它们?如何训练模型?光学堆栈中有很多不同的领域,对吧?有电信光学,像 Ciena 这样的公司表现强劲,它们周围的许多供应链也是如此。再看数据通信,芯片到芯片的通信,目前既有铜领域也有光学领域,这些都在快速增长,因为网络内容的增速在百分比上超过了其他任何内容。网络支出占 AI 芯片相关支出的比例从不到 10% 上升到超过 10%,而当我们进入 CPO 时代时,网络占比会进一步增长到 20-30%。所以网络内容有巨大的提升。但另一方面,CPO 是行业的一个巨大阶跃变化,现在每个人都意识到了 CPO。但我认为人们现在有点过于兴奋了。目前,人们对 CPO 有点过于乐观。在我看来,它不会在 2027 年到来。真正在 2028 年底,但 2029 年才是共封装光学大规模量产的真正起点。这有很多问题。这是一个制造问题。如果我们今天能以合理的成本部署它,那太棒了,每个人都会这么做。但这非常困难。制造量不够,良率不够,芯片也还没有真正设计好。这是一个非常复杂、难以量产的东西。所以人们会尽可能长时间地使用铜。这意味着 Rubin 全是铜。Fineman 在 GPU 上仍然是铜,对吧?那是下一代英伟达 GPU。在 Rubin 之后,是 Rubin Ultra,然后是 Fineman。我们甚至还没有到 Rubin 出货的阶段。Rubin 刚刚开始出货。所以我们在 GPU 上实现共封装光学之前还有几代芯片。交换机上的共封装光学会比 GPU 或 ASIC 上的更早到来。但最终,即使没有它,随着集群规模变大,每个 GPU 需要更多的光学器件或更多有源电缆之类的东西。所以我们看到了这种巨大的动态变化和转变。实际上,周一我们在 SemiAnalysis 为我们的机构研究订阅者发布了一份报告,那是针对本地时间的,不是终端市场。显然,技术方面,我们一直认为 CPO 会发生,我们推动这个已经很长时间了。我们的观点是铜会随着时间的推移被取代,但从中期来看,我们实际上非常看好铜,也非常看好非 CPO 的光学器件,而我们实际上有点看空 CPO,因为我们看到下游芯片的某些延迟。Fineman 没有完全采用 CPO,以及其他类似的情况,像 Amphenol 这样的铜厂商,他们生产所有的背板连接器和电缆,未来几年的表现将比之前预期的好得多,因为我们之前认为 CPO 会更快量产,但现在它被推迟了。供应链中就会发生这样的事情。但最终,光学器件是一个如果你今天闭上眼睛,五年后睁开,它会变得大得多的领域。其中很多已经反映在股价中,很多还没有,我想说存在一些局部的脱节。但这正是我们研究的一部分,也是我们与你们合作的工作的一部分:我们如何权衡这些?如何权衡 CPO 偏好的光学器件与非 CPO 光学器件、传统光收发器与铜之间的比例?因为铜实际上还有很长的路要走。铜行业有很多创新正在发生,并推迟了 CPO。就像,我为什么要做 CPO?因为归根结底,集成光学比电传输贵得多。除非我必须用电传输,但距离不够远,除非我添加中继器或光学器件。所以这是一种权衡和连续体,CPO 会发生,但看起来它被推迟了一点。
Yeah, I mean it's pretty clear as models get bigger, how do we run them across models? How do we train models? There's a lot of different domains within the optical stack, right? There's telecom optics, companies like Ciena have been ripping and many of the constituent supply chains around them. You look at datacom, chip-to-chip communications that's got a copper domain and an optics domain today, and those are all ripping because the growth of networking content is faster than the growth of any other content in terms of percentage. Networking is going from sub 10% to above 10% of spend associated with AI chips, and when we get to CPO, networking grows even further, it's like 20-30%. So we've got this huge uplift in networking content. But on the flip side, CPO is such a huge step function change in the industry and everyone sort of recognizes CPO now. But I think people are getting a little exuberant today. Currently, people are a little bit too excited on CPO. It's not coming in 2027 in my view. Really in the tail end of 2028, but 2029 is the real ramp for scale-up co-packaged optics. There's been a lot of problems. It's a manufacturing thing. If we could deploy it today at good cost, amazing, everyone would do it. But it's really hard. The manufacturing volumes are not there. The yields aren't there. The chips aren't really designed there yet. It's a very complex, difficult thing to ramp. And so people are going to stay in copper as long as they can. And that means Rubin is all copper. Fineman on the GPU is still copper, right? Which is next generation Nvidia GPU. After Rubin, Rubin Ultra, then Fineman. And we're not even at Rubin shipping yet. So Rubin just starting to ship. So we've got a few generations of chips before we get to co-packaged optics on the GPU. There's co-packaged optics on switches which is coming earlier than on the GPU or the ASICs. But ultimately, even without that, as the cluster size gets bigger, you need more optics per GPU or more active electrical cables and things like that. So we've seen this big dynamic and shift. Actually on Monday, we released a note at SemiAnalysis for our institutional research subscribers which was on a localized time, not saying end market. Obviously the technology, our agreement of CPO is going to happen, we've been pushing that for a long time. Our agreement is copper will subsume over time, but on a medium-term basis, we're actually very bullish copper and we're very bullish optics that aren't CPO, and we're actually kind of bearish on CPO because of certain delays on chips that we see downstream. Fineman not being full of CPO and other things like that where copper names like Amphenol, who make all the backplane connectors and cables, are actually going to do way better over the next few years than previously expected because we previously thought CPO would ramp sooner but now it's delayed out. So things like this happen in the supply chain. But ultimately, optics is one which if you close your eyes today and open it 5 years from now, it's going to be way bigger. A lot of that is priced into stocks. A lot of that isn't, and I'd say there's some local disjunctions. But this is sort of part of the research that we do and work that we've been doing with you folks as well is how do we weight that? How do we weight what amount is CPO favored optics versus non-CPO optics, typical optical transceivers versus copper? Because copper has actually got a long ways to go. There's a lot of things happening in the copper industry that are innovating and pushing back CPO. It's like why would I do CPO? Because at the end of the day, integrating optics is so much more expensive than sending something electrically. Except if I have to send something electrically, I can't go that far unless I add repeaters or optics. And so there's this trade-off and continuum, and CPO will happen, but it looks like it's getting pushed out a little bit.
我们可能不应该回避数据中心里那个显而易见的问题:你们怎么获取电力,又怎么把电力转换成合适的形态?我知道你在新闻简报里写过直流电和交流电的区别,以及某些相关要素。当你看到那些超大规模企业花这么多钱建数据中心,甚至可能把发电厂建在站点内、放在电表后面,我们该怎么看待电力需求——电网供电和非电网供电?我知道这是个很大的话题,但我觉得至少应该提一下。
We'd probably be remiss to not mention the elephant in the room at any data center: how are you getting the electricity, and how are you getting that electricity into the right form? I know you've written through the newsletter about direct current versus alternating current and certain elements. When you see the hyperscalers spending all this money building these data centers, potentially even putting power plants on site behind the meter, how should we think about the electrical demand, the grid versus non-grid? I know it's a big topic, but I feel like we'd be remiss to not at least mention it.
是的。我想说,数据中心的增长非常巨大。今年我们部署了 20 吉瓦的数据中心。明年这个数字会增长 50%,抱歉,是 30 吉瓦,然后后年将达到 50 吉瓦。数据中心容量的增长是巨大的。人们不得不应对很多局部的不协调。能源是最大的问题之一。其次是政治问题,第三是建设问题。建设数据中心、获得许可和备案在政治上很困难;人们试图阻止它。但最终限制它的主要因素还是能源。能源可以分解为几个方面:发电——我从哪里产生电子;输电——我如何将电子从发电地传输到数据中心;以及转换——因为传输过来的电力形态芯片无法直接使用,芯片需要另一种形态。这个转换管道是什么样的?在这三个方面,都有非常看好的点。在输电方面,这是最难被看好的,因为建设更多输电容量面临监管和政治困难,而且公用事业的地方垄断方式:如果他们建一条公用线路,必须将成本分摊给所有用户,而不仅仅是单个用户。输电存在各种奇怪的错位,所以在输电基础上建设更多电网容量相当困难。但在发电和转换方面,有两个有趣的事情。发电方面,显然电网上的发电量在增加。还有一个重大转变是为数据中心发电。我们预测几年后,数据中心新增电力的一半将在现场发电,而不是场外。所以电表后的发电正在飙升。我们从数据中心和能源模型中的电表后追踪器看到了这一点。我们团队中的一个人,Ellie,在哈萨克斯坦建了一个发电厂;她领导我们的能源模型。她一直在追踪,我们一直在构建整个电网的模型——每个发电资产、每个输电资产、所有负载资产,以及所有电表后的工作。有趣的是,我们看到电表后发电的巨大繁荣。在许可和监管方面有很多斗争——无论是人们不允许空气许可,还是不允许将天然气管道建到现场等。我们在 Oracle 的一个数据中心看到了这一点。有很多不同的方面正在发生,但最终的结果是电表后发电正在飙升。其中很多是天然气——来自 GE Vernova、三菱或西门子的联合循环燃气反应堆。但除此之外,还有许多不同类型的能源:往复式发动机、工业燃气轮机、各种类型的柴油发动机、火车发动机。人们把火车发动机、船用发动机、卡车发动机改装成数据中心发电。所以我们看到了大量的创新。并不是我们没有工业能力。美国每年生产数百万台往复式发动机。这些只是燃烧燃料并旋转的发动机。将这些发动机改装成使用天然气而不是柴油是相当简单的。但即使是柴油也没问题。然后你装上电动机,基本上反向驱动它,就能发电。所以你可以大量这样做来发电。我们看到超过 10 吉瓦的数据中心将使用这样的技术建造——将柴油卡车发动机改装成使用天然气(在生产时非常简单),装上反向驱动电动机,把它们放在数据中心现场,成百上千台这样的发动机支持一个数据中心。然后你从汽车修理店雇一群人,这些发动机需要维护,所以他们整天跑来维护这些柴油发动机。你有一些缓冲,这样当它们停机时,你可以维护它们并保持运行,获得最大功率。显然你需要一些电池作为中间缓冲,因为你不希望数据中心的波动炸毁发动机。所以你有整个电表后的供应链,这很令人兴奋。此外,大约两年后,太阳能加电池将比天然气更便宜。太阳能加电池的供应链很困难,这取决于你想要什么样的可靠性水平。如果你只有足够过夜的电池,那更便宜。但如果你需要足够三天的电池,因为可能连续下雨两天呢?你想要多少个九的可靠性?但由于中国的制造优势和某些补贴,太阳能加电池正以惊人的速度变得越来越便宜。在某个时候,太阳能加电池会变得更便宜。
Yeah. So, I would say data center growth is massive. This year we're deploying 20 gigawatts of data centers. Next year that number goes up 50% sorry, 30 gigawatts, and then it'll be 50 gigawatts the year after that. The growth in data center capacity is massive. There are a lot of local disjunctions that people are having to deal with. Energy is one of the biggest ones. The other is political, and third is construction. Building data centers and getting permits and filings is tough politically; people are trying to stop it. But the primary factor gating it is really the energy at the end of the day. What's happening there is energy can be broken down into a few things: generation—where do I generate the electrons from; transmission—how do I transmit the electrons from where it was generated to the data center; and conversion—because the power that gets transmitted is in a form factor that the chips cannot consume, so the chips need it in a different form factor. What does that conversion pipeline look like? In all three of these, there are very bullish aspects. On the transmission side, that's the hardest to be bullish on because of the regulatory and political difficulties with building more transmission capacity, and the way local monopolies of utilities work: if they build a utility line, they have to amortize it across all users, not just individual users. There are all these various weird dislocations with transmitting power, so building more grid capacity is kind of difficult on a transmission basis. But generation-wise and conversion-wise, there are two interesting things. Generation-wise, obviously there's more generation happening on the grid. There's also this big shift to generating for the data center. We predict in a couple years, half of the power for data centers for incremental new power will be generated on-site, not offsite. So behind the meter is soaring. We see this with the behind-the-meter tracker we have in our data center and energy models. Someone on our team, Ellie, built a power plant in Kazakhstan; she's leading our energy model. She's been tracking, and we've been building this model of the entire grid—every generation asset, every transmission asset, all the load assets, as well as all the behind-the-meter work. What's interesting is we've seen this huge boom in behind the meter. There's been a lot of fight on the permitting and regulatory side—whether it be people not wanting to allow air permits, or not allowing the gas pipeline to be built to the site, etc. We've seen that with an Oracle data center. There are a lot of different aspects happening, but ultimately the end state is that behind the meter is soaring. A lot of that is gas—dual combine cycle gas reactors from GE Vernova, Mitsubishi, or Siemens. But beyond that, there have also been many different types of energy sources: reciprocating engines, industrial gas turbines, various types of diesel engines, train engines. People have taken train engines, boat engines, truck engines and converted them into power generation for data centers. So we see a sea of innovation happening there. It's not like we don't have the industrial capacity. The US makes millions of reciprocating engines a year. These are just engines that burn fuel and spin. It's pretty trivial to retool those to be for gas rather than diesel. But even if it's diesel, that's fine. Then you stick an electrical engine on it and basically back-drive it, and that generates electricity. So you can do this in large volumes to generate power. We see 10 gigawatts plus of data centers that are going to be built with technologies like this—taking diesel truck engines, converting them to gas (which can be done at the time of production very simply), putting a back-driving electrical motor, sticking them on site on a data center, and you have hundreds of these backing a data center. Then you hire a bunch of people from car mechanic shops, and these things need to be serviced, so they run around servicing these diesel engines all day. You have some buffer so that when they go down, you can service them and keep them going and have max power. Obviously you need some batteries in between because you don't want the up and down of the data center to blow up the engines. So you have this entire supply chain of behind the meter, which is exciting. In addition, in about 2 years, solar plus battery will be cheaper than gas. Supply chains for solar plus battery are difficult, and it depends on what level of reliability you want. If you have just enough battery to get through the night, it's cheaper. But what if you need enough batteries to get through three days because it might rain for two days? How many nines of reliability do you want? But solar plus battery are getting cheaper and cheaper at an incredible pace because of China's manufacturing excellence and some subsidies. At some point, it's going to get cheaper to do solar plus battery.
然后还有太空数据中心,对吧,你甚至不需要电池,把它放在太空里,装上太阳能板就行了。所以有一整套发电方式,从把柴油发动机改成燃气发动机,到使用双联合循环发动机,再到干脆把芯片送到太空里去。所以这是一个完整的连续谱系,里面有很多钱可以赚,有很多有趣的动态事情可以做。这就是为什么 SemiAnalysis 最大的数据集和研究垂直领域,你以为是半导体,其实是数据中心和能源。是数据中心、工业和能源。我们称之为 DEI 团队,数据中心能源工业团队。内部是个双关语,标签是 @DEI 团队。Jeremy 领导那个团队,他起的名字。总之,数据中心、能源和工业实际上是我们最大的研究垂直领域,因为我们追踪每一个数据中心和每一座发电厂。所以当我们识别出延迟、某个建设正在发生、或者某家公司本季度将有这么多数据中心上线时,这是业内其他任何人都做不到的。这就是为什么它成为我们最大的垂直领域之一,每个人都对此感兴趣。谷歌对 Meta 能部署什么感兴趣,Meta 对 OpenAI 能部署什么感兴趣,而且所有这些公司都在关注供应链能做什么、谁有产能,还有所有投资者也在关注。所以那是我们最大的数据集。但我要说这是一个非常分散的市场。内存方面只有三家公司,很简单。加速器方面也只有几家。半导体晶圆制造设备方面也只有几家。但在这个领域,供应链里有成百上千家公司,制造各种随机的小零件。有几十家公司在建数据中心。还有几十家公司试图做独立电力生产商,或者做表后业务,提供某种电池服务等等。这是一个非常复杂的供应链,但充满活力,最终我认为有很多创新在发生。所以虽然数据中心会继续成为某种约束,但它们也不会成为约束,因为这取决于你愿意多疯狂。但就像我说的,你可以直接把卡车发动机改装一下,雇一群机械师,就这样运营一个站点。那不会是最好的。很多人会说:“真恶心,那能有多可靠?那会很烦人。”但人们正在这么做,而且它会起作用。虽然很麻烦,但会管用。一直到把东西发射到太空。那会很麻烦,很难让它工作,但它会管用。所以对于数据中心问题,你有解决方案,无论是完全肮脏的方式还是完全太空的方式。而供应链的其他部分则没有。所以我认为这就是这个市场如此动态的原因:你会看到人们起起伏伏很多。
And then you've got space data centers, right, where you don't even need a battery, you just stick it in space and you've got a solar panel and that's it. So you've got this whole continuum of ways to generate power, whether it be taking diesel engines and making them gas engines, or using dual combined cycle engines, all the way to let's just ship the chips into space instead. So there's a whole continuum, and there's a lot of money to be made there. There's a lot of interesting dynamic things to do there. That's why actually the largest data set and research vertical for SemiAnalysis, which you think is semiconductors, is actually data centers and energy. It's data centers, industrials, and energy. We call it the DEI team, the Data Center Energy Industrial team. It's a pun internally. The tag is at DEI team. Jeremy leads that team; he came up with the name. Anyway, data centers, energy, and industrials is actually our biggest research vertical because we're tracking every data center and every power plant. So when we identify a delay or that this build is happening, or that this company is going to have this many data centers go online in this quarter, it's something that no one else in the industry can do. That's why it's one of our biggest verticals and everyone's interested in that. Google's interested in what Meta is able to deploy. Meta's interested in what OpenAI is able to deploy, but also all of these guys are looking at what the supply chain is able to do and who has capacity, and then also all the investors are looking. So that's our biggest data set. But I would say it's a market that is very decentralized. In the case of memory, there are three names; it's pretty simple. In the case of accelerators, there are just a few names. In the case of semiconductor wafer fabrication equipment, there are just a few names. In the case of this, there are hundreds of names in the supply chain, making all these random little widgets. There are dozens of companies building data centers. And there are dozens of companies trying to do whether you're an independent power producer or you're doing it behind the meter, offering some sort of battery service or all these different things. It's a very complex supply chain, but one that has a lot of dynamism and ultimately I think there's a lot of innovation happening. So while data centers will continue to be a constraint of sorts, they will also not be a constraint because it depends on how crazy you're willing to go. But like I said, you can just take truck engines and convert them and hire a bunch of mechanics and run a site like that. It's not going to be the best. A lot of people say, 'That's disgusting. How reliable is that going to be? That's going to be really annoying to do.' But people are doing it and it will work. It's a pain in the ass, but it will work. All the way to shipping it into space. It's going to be a pain in the ass. It's going to be really hard to make it work, but it will work. So you've got solutions to the data center problem, whether it be going full dirty or going fully into space. Whereas other parts of the supply chain, you don't. So I think that's what makes this market so dynamic: you're going to see people go up and down a lot.
然后在发电和输电方面,转换是另一件事:如何把电力从发电或输电的地方变成芯片需要的?那里有一整套供应链,无论是 IGBT、碳化硅、各种 MOSFET、氮化镓 MOSFET,所有那些名字。当我们从 12 伏到 54 伏再到 800 伏直流电时,转换供应链里会发生什么?随着固态变压器的创新,会发生什么?所有这些都在这个领域发生。UPS 会怎样?电池备份和超级电容器以及其他各种平滑电力的方式,把左边产生的脏的、变化的电力变成右边超级干净但使用也变化的电力。如何匹配?整个转换管道非常令人兴奋。我们上周刚写了一篇关于 800 伏的博客。最近我们向机构订阅者谈到了一些延迟,Nvidia 那边在 Kyber 上延迟了。Reuben Ultra Kyber 不再有 800 伏了。那对供应链意味着什么?嗯,它被推迟了一点。
And then on the generation and transmission side, on the conversion side is the other thing: how do you get the power from where it is generated or transmitted to what the chips want? There's an entire supply chain of stuff going on there, whether it be IGBT, silicon carbide, various types of MOSFETs, GaN gallium nitride MOSFETs, all the names there. What happens when we go from 12 volt to 54 volt to 800 volt DC and the conversion supply chain in there? What happens with solid state transformers as those get innovated? All these things are happening in the space. What happens with UPSs? Battery backups and super capacitors and all these other different ways to smooth out the power, make it from the dirty variable power that gets created on the left side to the super clean power but also variable usage on the right side. How do you match that? That entire conversion pipeline is one that's super exciting. We had a blog on that and 800 volt just last week. And we've talked more recently to our institutional subscribers about some delays happening there on the side of Nvidia as they delay it out of Kyber. Reuben Ultra Kyber doesn't have 800 volt anymore. So what does that mean for the supply chain? Well, it gets pushed out a little bit.
Dylan,我想代表我们这边非常感谢你。如果按章节来算,这是第一期节目。这是我们第一次请 Dylan 上播客,但肯定不是最后一次,因为还有更多信息,而且正如他多次直接说到的,整个技术栈都在变化。一切都在不断变化,很难跟上。
So Dylan, I want to thank you profusely on our side. This is the first episode if we think in terms of chapters. This will be the first time we've had Dylan on the podcast, but certainly not the last, because there's a lot more information and like he said directly numerous times across the entire stack. Everything's changing all the time. It's a bear to keep track.
另一件事是,这个供应链太疯狂了。很多时候我们谈论大东西:内存、CPU、数据中心。但当你深入供应链时,局部波动非常小且很多。有几个月我们一直在讨论 PCB 钻头,就是在 PCB 上钻孔的钻头。就像 PCB 上的随机铜箔。供应链中有所有这些随机的小东西,它们也有这些错位,而其中的公司遍布各地。它们可能在台湾、日本、韩国或世界各地交易,投资者不容易接触到。所以我认为我们的合作以及我们合作的方式非常令人兴奋:我们能够影响正在发生的事情,我们能够大量讨论这些供应链中断,而且在我之前提出的框架以及我们试图覆盖的整个图景中,也有很多有趣的东西。
The other thing I would say is this supply chain is so freaking crazy. A lot of times we talk about the big ones: memory, CPUs, data centers. But actually when you drill down to the supply chain, the local bumps are very small and many. For a couple months we were talking about PCB drill bits, the drill bits that drill into PCBs for the holes. It's like random copper foil that goes on PCBs. There are all these random small things in the supply chain that also have these dislocations, and the companies that exist in them are all over the place. They could be trading in Taiwan, Japan, Korea, or all parts of the world. It's not just easily accessible to investors. So I think that's what's really exciting about our partnership and the way we're working together: we get to influence what's going on, we get to talk a lot about these supply chain disruptions, but also what's really interesting in the framework that I laid out earlier and the entire landscape that we're trying to cover.
那么,期待你再次来到我们的节目,以及我们其他的合作。
And so you know looking forward to coming back on the show more and our other collaborations.
当然。在此,我必须为我们的合规团队说明一点:本播客中表达的观点和意见仅代表 Wisdom Tree,且可能发生变化。本播客中的任何内容均不应被视为预测、研究、投资或税务建议。本播客中表达的信息和意见不构成对任何证券的购买或出售推荐、要约或招揽,听众应自行决定是否依赖这些信息。请记住,过往业绩并不预示未来结果。感谢大家今天抽出时间与我们交流,我们期待未来再次相聚。保重。
Absolutely. And with that I do have to state one thing here for our compliance team to clarify the views and opinions expressed in this podcast are those of Wisdom Tree and are subject to change. Anything we present in this podcast is not intended to be relied upon as a forecast, research, nor as investment or tax advice. The information and opinions expressed in this podcast are not a recommendation, offer or solicitation to buy or sell any securities and reliance upon them is at the sole discretion of the listener. Please remember past performance is no indication of future results. Thank you everyone for taking some time with us today and we look forward to coming back in the future. Take care.