从汽车旅馆到半导体分析:Dylan Patel 的创业之路

From Motel to Semi Analysis: Dylan Patel's Journey

迪伦·帕特尔 Dylan Patel · Training Data · 2026-06-30 · 约 70 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Dylan Patel 分享了他从在家族汽车旅馆和加油站长大,到创立半导体分析领域顶级研究公司 Semi Analysis 的历程。

Dylan Patel shares his journey from growing up in a family motel and gas station to founding Semi Analysis, a premier research firm in the semiconductor space.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 21)

全文 · Full transcript(中英对照)

引言与半分析文化 Introduction and Semi Analysis Culture

Host

我觉得在 Semi Analysis 内部特别有意思,因为我们有 90 个人,其中很大一部分是覆盖整个供应链的技术工程师,还有很大一部分是前对冲基金的人。你会看到这样的争论:有人说“哦,那不重要”,然后另一个人说“但是成本呢”,然后工程师说“不不不,这个技术最酷了”。你会看到这种自发的交锋。我们相当随意。考虑到我以前是版主,你可以想象我有多享受。不要和猪摔跤,因为猪乐在其中,对吧?

I think it's really fun inside of Semi Analysis because we have 90 people and a big chunk of them are technologists engineers across the whole supply chain. And then a big chunk is people who are formerly at hedge funds. And you see these arguments like people are like, "Oh, well that doesn't matter." And it's like, then someone's like, "Well, but cost." And then the engineers like, "No, no, no, but this technology is the coolest." And you see this organically fight it out. And we're pretty informal. Given the fact that I was a moderator, you can imagine what the enjoying it. You don't wrestle with a pig because a pig enjoys it, right?

Host

我们现在在 Semi Analysis 的办公室,和 Dylan Patel 在一起。我是 Sequoia 的 Sean,还有我的合伙人 Sonia Huang。你取得的成就太惊人了。五年前,半导体在西方并不性感,在东方很性感,但西方这里的人几乎忘了它们。但你没有忘记。你大力投入,创建了可能是这个领域首屈一指的研究公司,从非常技术性的细节到供应链再到宏观图景,一直在教育世界。有传言说 Semi Analysis 最近营收突破了 1 亿美元,我不知道这些传言有多准确。但不管数字是多少,你们都在大杀四方。

We're here in the Semi Analysis office with Dylan Patel. I'm Sean from Sequoia. My partner Sonia Huang. It's pretty insane what you've done. Semis 5 years ago were not very sexy in the west. They were sexy in the east but people here in the west had kind of forgotten about them. You did not forget about them though. You went very long. You created probably the premier research company in the space that's been educating the world on the state of the art from very technical details to supply chain to the bigger picture. There's rumors that Semi Analysis recently passed 100 million of revenue. I don't know how accurate those are. Whatever the numbers are you guys are crushing.

Dylan

和信息的准确度一样。

It's as accurate as the information is.

Host

酷。你永远不知道。还有传言说你可能要创办一个风投基金。我在生态系统中一直听到人们想要与 Semi Analysis 关联。你建立了这个值得信赖的品牌,无论你做什么都很成功。这显然只是你旅程的开始。恭喜你。但这一切是怎么发生的?你最初是怎么走到现在这一步的?

Cool. You never know. There's also rumors that you might start a venture fund. I hear all the time in the ecosystem people wanting affiliation with Semi Analysis. You've built this trusted brand and whatever you do it's working. This is clearly just the beginning of the journey for you. Congratulations on all of that. But how did this happen? How did you first get to where you are now?

早年生活与首个神经网络 Early Life and First Neural Network

Dylan

好吧,当我还是个小男孩,刚从娘胎里出来的时候。不。所以,我在一个小生意家庭长大。我父母有一家汽车旅馆,我们就住在汽车旅馆里。我们还有一个加油站。所以我一直在卖东西。我经常开玩笑说,我训练的第一个神经网络是根据人们走进加油站时的种族和视觉特征来识别他们该拿哪种香烟。基本上,香烟都陈列在货架顶部,我太矮了够不着,而且从技术上讲,那个年龄卖香烟是不合法的,但不管了。我得把台阶凳挪到正确的位置。

Well, when I was a young boy coming out of the womb. No. So, I grew up in a small business. My parents had a motel. We lived in the motel. We had a gas station. So, I was selling. I joke a lot of times the first neural network I trained was racially and visually profiling people based on when they enter the gas station, which cigarette to pick. Basically, the cigarettes were all extruded across the top and I was too short to actually reach them and technically it wasn't legal to sell cigarettes at that age, but whatever. I had to move the step stool over to the right area.

Host

我第一份工作也是在合法年龄之前就开始干了。所以,但这是很好的经历。

I started working my first job before it was legal, too. So, but it's good experience.

Dylan

好吧,我没拿到工资,对吧?这是家族生意。

Well, I didn't get paid, right? It's a family business.

Host

一样。

Same.

Dylan

但没错,我们有汽车旅馆,街对面就是加油站。所以,有时候有人走进来,如果是一个卷发的老白人女士,我就把梯子或台阶凳挪到骆驼牌香烟的位置。不同的年龄、人群、职业、种族等等,我都会挪动台阶凳。我开玩笑说这是我训练的第一个神经网络,因为如果我等他们告诉我,我就得先挪凳子再站上去,而不是提前准备好。所以薄荷醇和 100s Slim 之类的,我开玩笑说那是我训练的第一个神经网络。但我在家族生意中长大,住在汽车旅馆里。

But yeah, we had our motel and then across the street was our gas station. So, sometimes someone would walk in and if an old white lady with curly hair walked in, I'd move the ladder or the step stool over to where the Camels are. And different age, demographic, profession, race, etc., I would move the step stool over. I joke this is the first neural network I trained because if I waited for them to tell me I'd have to move it over and then step up versus just being ready. So menthols versus 100 slims and all these things, I joke that's the first neural network I trained. But I grew up in family businesses. Lived in a motel.

Xbox 360与硬件兴趣 Xbox 360 and Interest in Hardware

Dylan

这一切都要追溯到我的 8 岁生日。我的生日在五月。四月的时候 Xbox 360 发布了。生日那天我没有要 Xbox,也没有要生日礼物。我父母问我要什么,我说要等到圣诞节。我们庆祝圣诞节,但至少在当时,我觉得他们不可能在圣诞节给我 Xbox 360,所以我就为圣诞节要了它。不管怎样,圣诞节到了,我得到了它。几个月后,我住在阿拉巴马州的表弟要来春假,他们也是住在汽车旅馆里,我们打算在我家玩。他的年龄在我和我哥哥之间,我哥哥更爱运动,所以他对 Xbox 不太在意,偶尔玩玩,但不上心。但我表弟,我想让他觉得我很酷,对吧?所以我在电话里吹嘘了好几次,说“是啊,我有 Xbox 了”。然后 Xbox 坏了。有一个硬件缺陷叫“三红”。但长话短说,我不得不打开它,短路温度传感器,然后修好了。但我之前试了很多其他方法,都没用。这就是我如何开始接触硬件的。我打开了潘多拉魔盒。

It all really goes back to when I was like, it was my 8th birthday. My birthday's in May. And it was April when the Xbox 360 was announced. For my birthday, I didn't ask for the Xbox or I didn't ask for a birthday gift. My parents asked what I wanted. I asked for it for Christmas. We celebrated Christmas, but there was no way, at least at the time, I thought there was no way they would give me the Xbox 360 for Christmas and so I asked for it for Christmas. Anyways, Christmas comes around, I get it. Fast forward a couple months, my cousin who lives in Alabama, they also lived in a motel, was going to come over for spring break, for his spring break and we were going to just hang out at my house. He's in between me and my older brother in age, brother's a bit more jockey. So he didn't really care too much about the Xbox. He played sometimes, but he didn't really care. But my cousin, I wanted to think I'm him to think I was cool, right? So I bragged many times on the phone. I was like, "Yeah, I got an Xbox." And then the Xbox broke. There was a hardware defect called the red ring of death. But long story short, I had to open it up and short the temperature sensor and it fixed it. But there were many other tricks I tried first and none of them worked. And so that's sort of how I got into hardware. I opened Pandora's box.

Dylan

到 12 岁时,我经常在这些论坛上阅读和发帖。大约在那个时候,Reddit 吞噬了所有其他形式,所以我成了 Android、Apple、Google 以及硬件板块的版主,关注着 Intel、Nvidia 和 AMD 等所有其他论坛。我组装了一台 PC。所有这些论坛我都在看、读、发帖,有些我还担任版主。所以,智能手机,看着它们从非常简单发展到速度竞赛,再到在架构上比 PC 更先进,GPU 也一样,追踪、观察、阅读每一条评论。总是带着经济学的色彩,因为我在小生意中长大。所以我总是关注经济学。

By the time I was 12, I was on these forums a lot reading and posting a lot. This is around the time when Reddit ate all other forms and so I became a moderator of Android, Apple, and Google as well as hardware and was watching Intel, Nvidia, and AMD and all these other forums. I built a PC. All these forums I was watching, reading, posting a lot, but some of them I was moderating a lot. And so, smartphones, watching smartphones develop from very simple to speed racing to being technologically more advanced than PCs in many ways architecturally, and same with GPUs, just tracking and watching that, reading every comment. Always having the economic tinge because I grew up in a small business. So I was always looking at the economics.

Dylan

曾经有一段时间,网上所有的技术宅都喜欢 AMD GPU,我个人也买过 AMD GPU,因为性价比。但说到技术上哪个更好,我总是说“不不不,Nvidia 更好,因为他们用更小的芯片实现了更好的性能和能效,而且利润率更高”。所以我总是谈论 Nvidia 的利润率如何优于 AMD 的 GPU,这很有趣。

There was a time where all the neck beards on the internet loved AMD GPUs and I personally had bought an AMD GPU too because price performance but then when it came down to what's technically better I'd always be like no no no Nvidia is better because they use a smaller chip to get better performance at better power efficiencies and their margins better. And so I would always talk about how Nvidia's margins were better than AMD's in the GPU landscape and so it's very fun.

Host

而你当时才 12 岁。

And you were 12 at the time.

Dylan

我 12 岁开始当版主,但这贯穿了我的青少年时期和高中时代,对吧?

I started moderating when I was 12, but this is all through my teenage tween age and high school years, right?

Host

你还有其他奇怪的爱好吗,还是只有半导体?

Do you have any other weird hobbies or was it just semis?

Dylan

我玩了很多《星际争霸》。有一段时间,我在北美天梯上是宗师级别。《星际争霸 2》。

I played a ton of Starcraft. At one point, I was grandmaster on the North American ladder. Starcraft 2.

Host

非常认真。

Very serious.

Dylan

所以,你在多个事情上都痴迷地做到了极致。

So, you've gotten obsessively good at multiple things.

痴迷与成绩 Obsession and Grades

Host

是的,我的意思是,痴迷是好事。

Yeah. I mean, obsession is good.

Host

你的成绩怎么样?

How were your grades?

Dylan

嗯,还算不错。我大部分科目都是 A,但那些我觉得很无聊或者不喜欢的课,比如西班牙语,成绩就不太好。不过顺便说一句,我西班牙语说得很流利,所以这很蠢。但就是……

Um, they were decent. I would say I had mostly A's, but they were classes I thought were really boring or I just didn't enjoy. Like Spanish, I got not the greatest grades. But I speak fluent Spanish by the way, so it's really dumb. But it's just sort of...

Host

也许这就是你成绩不好的原因。

Maybe that's why you didn't get a good grade.

Dylan

公平地说,我后来才学的西班牙语。但没错,我的成绩还行,对亚洲父母来说足够好了。我比学校里大多数人都好,但也不是那种拼命拿全 A 的。

I didn't learn Spanish till later, to be fair. But yeah, my grades were fine. They were fine enough for Asian parents. I was better than most of the school, but it wasn't like tryhard maxing for all A's.

创办半分析 Starting Semi Analysis

Host

所以你很大程度上是互联网的学生,这就是你培养专业知识的方式。你是在什么时候决定创办 Semi Analysis 的?自公司成立以来,最大的惊喜是什么?

So you're very much a student of the internet then, this is how you develop this expertise. At what point do you decide to start Semi Analysis and what's been the biggest surprise since starting the company?

Dylan

是的,我上过学,拿了几个与半导体无关的学位。我在一家小型量化风险公司做了两年量化分析师。然后,一系列事件达到了顶点。一是我的奖金被坑了。我利用市场中的风险点,为公司创造了数百万美元的无风险收入,我想超过 1000 万,然后别人却抢了我的功劳。但最终我还是被合理化了。不过我失去了与公司的社会契约。另外,我的祖父母和我们一起住在家里,住在汽车旅馆里。他们和我们住在一起,所以我和他们很亲近。我祖母得了痴呆症,她忘记了我,还从楼梯上摔下来,发生了一场悲剧,去世了。这一切都发生在 2020 年初。此外,还有一些感情问题。所以有几件事让我非常难过。所有这些事情都汇聚到了一起。然后新冠疫情爆发了,我哥哥说:“兄弟,来和我住吧。”他住在纳什维尔,所以我就去和他住了。我们当时想:“哦,封锁只会持续几周。你可以和我住在一起,等封锁结束再回家。”经典的最后遗言。封锁持续了更长时间。但和我哥哥住了几个月,感觉就像,好吧,我不知道自己在做什么。我现在在我哥哥家。一切都是他的规矩。他和他当时的未婚妻,现在的妻子,都在那里。所以我基本上得小心翼翼,但我不在乎我的工作。所以我发帖比平时更多。我一直在网上发很多帖子。我一直大量交易股票,但我做空新冠疫情和做多新冠疫情赚了很多钱。半导体短缺也差不多在那时发生。总之,我非常痴迷于发帖。最终在那段时间,我在网上和某人发生了争执,他们人肉了我。他们公开了我匿名账户的真实身份。当时我想:“哦,不,我好害怕。”我停了大约三周没发帖,然后想:“我在干什么?我为什么在乎?”于是我就开始用真名发帖。我也有博客之类的东西。我创建了一个真正的博客,Semi Analysis,在我 24 岁生日那天,我发了两篇博客。从那以后,它就起飞了。那不是一份通讯,但我获得了巨大的关注,因为现在不是用匿名名字发帖,而是真名,而且我在这两篇文章上投入了比平时多得多的精力。不是像在网上发帖那样,而是真正用心写博客。你可以回去读读那些文章,如果你想的话。它们没那么好,但在当时算不错了。它们是当时网上能找到的关于半导体最好的内容。我就一直发帖。我开始接到很多咨询业务。2020 年,我也有些崩溃了。不知道自己想做什么。所以我收拾好所有东西。我开上我的卡车,买了一个适合卡车后斗的帐篷,买了一个气垫床,然后开车去了美国各地的国家公园。每周有两三天,我会住在随便一家汽车旅馆,我把价格谈到每晚 30 美元左右。然后我会工作。周末我会读书,经常在某个国家公园或徒步时读教科书,听关于半导体、人工智能以及所有我关心的事情的有声书,在这六个月里我学到了更多东西,我去了每一个国家公园。整个过程我都是一个人,发着博客。每个人都问:“Dylan,你到底在干什么?”

Yeah, so I went to school, got a few degrees in stuff that wasn't related to semiconductors. I was a quant for two years at a small quant risk firm. And then basically, there was a culmination of events that happened. One was that I got screwed out of a bonus. I'd made my company many millions of revenue of risk-free revenue because I exploited a risk thing in the market. I think well over 10 million, and then someone else took credit for my work and all this sort of stuff. But eventually I did get rightsized. But I lost a social contract with the company I was working with. Also, my grandparents grew up in my house with us, in the motel with us. They lived with us, so I was very close with them. My grandmother got dementia and she forgot who I was, and she fell down some stairs and had a tragic accident and passed away. All of that happened in early 2020. Additionally, there were some girl things. So there were a few things that happened that made me very sad. And so all of those things sort of culminated. Then COVID happened and my brother said, "Dude, just come stay with me." He lived in Nashville, so I came and stayed with him. We were like, "Oh, lockdowns will be a few weeks. You can stay with me while they happen and then you can go back home." Famous last words. Lockdowns lasted much longer. But living with my brother for a few months was like, okay, I didn't know what I was doing. I was now at my brother's home. Everything was his rules. Him and his fiance at the time, now wife, were there. So I basically had to tiptoe around, but I didn't care about my job. And so I was posting even more than normal. I'd always been posting a lot on the internet. I'd always been trading stocks a lot, but I made a lot of money shorting COVID and longing COVID and all this stuff. Semiconductor shortages happened around then too. And anyways, I was very much obsessed with posting. And eventually around that time, someone I got into an argument with on the internet doxed me. They publicly revealed my identity for my anonymous account. At the time I was like, "Oh no, I'm scared." I stopped posting for like three weeks and I was like, "What am I doing? Why do I care?" So then I just started posting under my real name. I had blogs and stuff as well. I made a real blog, Semi Analysis, and on my 24th birthday, I posted two blogs. And from there, it just took off. It was not a newsletter, but I got so much traction because now instead of posting on an anonymous name, it was a real name and I put a lot more effort into those two posts than I usually did. Instead of like posting on the internet, it was real effort into the blog. You can actually go back and read those if you want. They're not that great, but they were good for the time. They were the best stuff you could find on the internet about semis. And I just kept posting. I started getting a lot of consulting business. In 2020, I also sort of crashed out. Didn't know what I wanted to do. So I packed everything up. I took my truck, bought a tent that fits on the back of the truck, bought an air mattress, and drove around all these national parks all around America. So two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like $30 a night for a room. And I'd work on stuff. And then the weekends I'd read books and oftentimes read textbooks while in some random national park or hiking and listen to audiobooks about semiconductors, about AI, about all the things I cared a lot about, and got way more educated over these six months where I'm just going to every national park. And the whole time I was alone, posting blogs. Everyone was like, "Dylan, what the hell are you doing?"

Host

是在星链之前还是星链早期?

Pre-Starlink or the very early days of Starlink?

Dylan

在星链之前。是的,所以感觉就像“你在干什么?”我先和一个朋友在拉丁美洲旅行了一年,然后和我的前女友又旅行了大约一年。然后 2022、2023、2024,2021 年底、2022、2023 和 2024,我从 2020 年中期开始就完全无家可归了。但我旅行去参加世界各地的每一个会议。我每年参加 40 多个会议,不管它在供应链的哪个环节。我会想:“哦,那个看起来很有趣。我想我会去。”我去了一次会议,觉得:“哇,太棒了。你可以和专家交谈,他们会和你聊,因为你很兴奋。”在半导体领域,每个人都是婴儿潮一代,所以这很棒。他们看不到对此兴奋的年轻人,所以他们很乐意分享。你只需要问。

Pre-Starlink. Yeah, so it was very much like "What are you doing?" I traveled around Latin America again for a year initially with my friend and then with my ex for about a year. And then 2022, 2023, 2024, end of 2021, 2022, 2023, and 2024, I'm completely still homeless since mid-2020. But I'm traveling around to every conference in the world. I go to 40 plus conferences a year no matter where in the supply chain it is. I'm like, "Oh that looks interesting. I guess I'll go to that." I went to one conference and was like, "Wow, this is amazing. You get to talk to the experts and they'll talk to you because you're so excited." In the case of semiconductors, everyone's a boomer, so it's great. They don't see young people who are excited about it, so they're really happy to tell stuff. So you just have to ask.

Host

供应链的某个部分或这些会议中,有没有哪个特别改变了你对半导体世界的看法,或者你当时或现在觉得特别被低估的?

Was there a part of the supply chain or one of these conferences that particularly changed your view of the semi-world or that you felt then or feel now is particularly underrated?

Dylan

我认为展会和会议的范围非常广泛。显然,我最喜欢的一些包括 NeurIPS。为什么?因为那里有 2 万名 AI 研究人员,而且他们通常和我的年龄相仿。所以很有趣,但他们也是顶尖的 AI 研究人员。这很有趣,你能学到很多东西。

I think the trade shows and conferences range really widely. Obviously some of the ones I have the most fun at include NeurIPS. Why is that? Because it's 20,000 AI researchers and they're generally in my age range. So it's a lot of fun, but they're also leading AI researchers. It's a lot of fun and you learn a lot.

会议经历与供应链洞察 Conference experiences and supply chain insights

Dylan

嗯,还有很多派对,然后各种类型都有,比如日本有个随机化学会议,有 300 个日本大叔。大概 20 个来自 ASML,20 个来自台积电,20 个来自英特尔,只有这些人说英语。其他人只说日语,你会想,“嗯,我觉得还是挺有意思、挺好玩的。”我觉得我的一项技能是,无论对方背景如何、是谁,我都能和他们打成一片。我能和他们聊天,找到有趣的话题。通常都是技术相关,但你知道……所以我觉得最有趣的会议往往是那些大型会议,因为那里发生着最重要的事情。但我觉得真正令人兴奋的细分领域是像 SPIE 这样的。有 IEEE(国际电气电子工程师学会),还有 SPIE,这是另一个生态系统。SPIE 会议非常非常深入细节。我参加的每一次,尤其是 SPIE 高级光刻或 SPIE 光掩模会议,第一次去的时候,我连 90% 的内容都听不懂。然后我读了很多资料,当然有了一些背景知识,第二次去的时候,我大概听懂了一半。第三次去的时候,我大概听懂了 75%。即使现在,我去参加,还是不能完全理解所有内容。而像 NeurIPS 这样的会议,你去几次就能理解,比如“什么是神经符号推理?”、“这是什么?”、“那是什么?”——你很快就能对整个领域有个大致了解。但供应链的某些部分非常神秘、非常深入、非常技术性。你需要花很多时间才能理解到底发生了什么。而且对于每篇研究论文,你去参加会议有几个原因:你了解研究,但所有研究都是已发表的。你真正关心的是,这些研究如何与技术结合,以及这些研究与现有技术有何不同。而这些研究论文没有一篇告诉你当前的技术现状。但你可以直接问人,建立人脉,学习,然后了解供应链,比如“这家公司供应那家公司”,即使这些信息没有公开披露。或者你会知道某种化学品成本大约是多少,某个工具消耗多少。你还会听到一些恐怖故事,比如某种化学品短缺,完全打乱了供应链的某个环节,结果发现全世界只有三家公司生产那种化学品。我最喜欢的一个故事是,在那个几乎没人说英语的日本会议上,一个日本大叔用非常蹩脚的英语告诉我,他父亲在 1980 年代在这个行业工作,当时世界上唯一生产那种化学品的工厂烧毁了,导致内存价格翻倍甚至三倍。我当时想,“哇,和今天没什么两样。”

Um, there's also a lot of parties, and then it ranges all the way to like, you know, there's a random chemical conference in Japan where it's 300 Japanese dudes. It's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Everyone else speaks only Japanese, and you're like, 'Huh, I guess they're still pretty interesting and fun.' I think one thing that I have as a skill set is that I'm able to bond with anyone regardless of their background and who they are. I'm able to talk to them, find something interesting to talk about. Oftentimes, it's the tech stuff, but, you know, it's—and so I think the most interesting conferences are oftentimes, like, you know, the really big ones because that's where the biggest stuff is happening. But I think the niches that are really, really exciting is like, you know, SPIE. So there's IEEE, which is International Electrical Engineering something, and there's SPIE, which is another ecosystem. SPIE conferences are super, super deep in details. Every single one that I went to, especially like SPIE Advanced Lithography or SPIE Photomask, I went to them the first time and I didn't even understand 90% of what I heard. And then I read, read, read, I had made some context, of course, and then next time I went, I understood like half of what I went to. Third time I went, I understood like 75% of what I went to. Even now, I went and I was like, I still don't understand everything that's going on. Whereas like you go to like NeurIPS, you know, a couple times you can understand, okay, what's neurosymbolic reasoning? Okay, what's this? What's that? Like you can kind of get a mapping of what everything is pretty quickly. But some parts of the supply chain are so arcane and so deep and so technical. It takes a lot of time for you to even understand what's happening. And on everything, right, for every research paper, it doesn't necessarily mean you—you know, you go to a conference for a few reasons, right? You understand the research, you understand like, but it's all the research that's being published. But what you really care about is understanding how does that research intersect with technology, also how does that research differ from what's there today. And none of these research papers tell you what's happening today. But then you just ask people, and you build contacts, and you learn, and then you learn about the supply chain, and 'Oh, this company supplies this company' even though it's not publicly stated anywhere, or like, you know, you learn that this chemical costs about this much and a tool uses about this much. And you hear the horror stories of like this chemical had a shortage and it totally threw off this part of the supply chain, and then it turns out there's only three companies in the world that make that chemical. And it's like—my favorite one is I learned from a Japanese guy at that specific Japanese conference that I went to where almost no one spoke English, in very broken English, he told me about how his father worked in this industry in the 1980s, that the only factory in the world that built this chemical burned down, and that caused memory prices to like double or triple. And I was like, 'Wow, not too different from today.'

Host

和今天没什么两样。太疯狂了。

Not too different from today. Crazy.

Dylan

嗯,推理将成为地球上最大的市场,甚至超越地球的最大市场。同意还是不同意?

Um, inference is going to be the biggest market on earth, biggest market beyond earth. Agree or disagree?

Host

嗯,我的意思是,显然 token 的使用将成为最大的市场,token 创造的价值也将成为最大的市场。但我认为 token 经济学,也就是 token 的使用、AI 的采用,是正在发生的最重要的事情。而推理,无论是开源模型还是闭源模型,都将成为世界上最大的市场之一,我认为比石油大得多,比很多其他领域都大。AI 推理将占据 GDP 的多个百分点,没错。

Um, I mean, obviously use of tokens is going to be the biggest market, and the value that's created from tokens is going to be the biggest market. But I think tokconomics, sort of the use of tokens, adoption of AI, sort of is the most important thing that's happening. And inference, whether it's open models or closed models, will be like one of the biggest markets in the world, much bigger than oil, I think, much bigger than like, you know, many other parts. Like inference of AI will be many percentage points of the GDP, yeah, right.

Dylan

我认为你在 Inference X 上所做的工作是行业标准。也许说说你为什么创办它,它做什么,以及人们对推理性能基准测试有什么误解?

What you've done with Inference X, I think, is industry standard. Maybe say a word on why you started it, what it does, and what do people misunderstand about performance benchmarking on inference?

Host

是的。所以,往回看,Semi Analysis 做了很多事情,很多是为机构客户做的研究,以及我们的订阅产品,但也有很多是像“嘿,这个搞明白会很酷。我们想想怎么弄明白,然后公开发布。”这样的事。然后规模越来越大。我们在 GPU 基准测试、训练性能和推理性能的测试方面做了很多工作。但最终我们发现推理基准测试是时间点式的。你测试一下,花点时间,发布结果,但结果又慢又神秘又过时,因为模型一直在变。我觉得每周都有新模型,不管是中国的模型,还是今天 Mythos 5、Fable 发布了,新模型层出不穷。在软件层面,PyTorch、VLM、SG lang、新驱动、新东西不断发布。事实上,大多数这些库的更新周期是每周两次。所以软件一直在更新,性能也随之变化。新的推理优化不断出现,它们也会更新。所以我觉得这是一连串无休止的突破,不断推动效率和成本下降,这就是为什么我们看到同等质量的模型成本每年下降 60 倍。太不可思议了。但要跟上这个节奏,你不能用时间点式的基准测试。你需要让基准测试保持鲜活,即持续在最新硬件和最新模型上运行。所以我们启动了一个项目,得到了生态系统的广泛支持。这之所以可能,是因为我们在生态系统中积累了一定的声誉,能够获得 CoreWeave、Crusoe、Nebius、Oracle、Microsoft、Amazon、Google 和 OpenAI 为我们贡献算力。然后我们能够与 SG Lang、VLM,以及现在的 Radix Arc 和 InRact 合作,这些是领导这些开源工作的私营公司。我们还能够获得 Nvidia、AMD、Google 和 Amazon 的合作,因为我们正在加入 TPU 和 Trainium。现在所有这些公司都在合作。我们获得了超过 5000 万美元的硬件捐赠。一旦我们推出 TPU 和 Trainium,硬件捐赠应该会超过 1 亿美元。

Yeah. So, to zoom back, right, like Semi Analysis, we do a lot of stuff that's like, you know, a lot of it is like research for institutional clients and our subscription versus products, but a lot of it is also like, 'Hey, you know, this would just be cool to figure out. Let's figure out how to figure it out and just post it publicly.' And that gets, you know, more and more scale. And so we've done this with a lot of GPU benchmarking and testing and training performance and inference performance. But, you know, ultimately we saw like inference benchmarking was like point in time. You know, you test it and you take some time, you release it and it's like slow and arcane and outdated because models change all the time. Every I feel like every week there's a new model, whether it's a Chinese model or, you know, today Mythos 5, Fable dropped, and new models are coming out all the time. On the software layer, PyTorch, VLM, SG lang, new drivers, new something drops. You know, in fact, the update cycle for most of these libraries is twice a week. So you basically have the software updating all the time and therefore performance changing. You know, new inference optimizations are coming out and those get updated. And so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down, which is why we've seen model cost drop for equivalent quality by like 60x a year. It's incredible. But to stay on top of that, you can't have point-in-time benchmarking. You need to have benchmarks be living and breathing, i.e., constantly running on the latest hardware on the latest models. And so we embarked on a project and we got a lot of buy-in from the ecosystem. This was only possible because we had enough aura with some of the ecosystem where we're able to get CoreWeave and Crusoe and Nebius and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us compute. And then we were able to work with SG Lang and VLM and now Radix Arc and InRact, which are the private companies who are sort of leading those efforts, the open source efforts, to collaborate with us. We're able to get Nvidia and AMD and Google and Amazon now because we're adding TPUs and Trainium to collaborate. Now we've got all these people collaborating. We've got over $50 million of hardware donated to us. Once we launch TPUs and Trainium, it actually should be over $100 million of hardware.

推理基准测试与吞吐-交互曲线 Inference benchmarking and the throughput-interactivity curve

Dylan

嗯,大概有 15 种不同的芯片类型,每天都在所有最新模型上运行这些基准测试,对吧?Moonshot 最好的模型、阿里巴巴最好的模型、还有大约五个不同的中国模型、最好的开源模型、最好的中国实验室。我们每天对他们的模型运行基准测试,还有美国最好的开源模型,GPT-OSS、Neotron 等等。所以我们每天以自动化方式运行这些基准测试,它们运行在专门用于推理基准测试的服务器上,我们遍历了非常多的配置和优化类型。然后它产生的结果——所有结果和配置都是公开的。现在我们有了帕累托最优曲线,因为很多时候人们比较推理性能时,会拿别人的次优曲线或点来对比自己的最优曲线。这就像,嗯,是的,我可以——如果我开保时捷而对方是赛车手,显然我开得更慢。推理基准测试也是一样。所以我们做的是创建了开源容器,覆盖交互性(即响应速度)与批大小(即同时服务的用户数)曲线上的每个最优点。现在任何想要最优点的人都可以去 InferenceX 下载,然后作为最优点运行。他们可以每天检查,甚至可以自动下载该模型的最优解,推理性能就会接近峰值。

Um, you know, maybe about 15 different chip types all running these benchmarks every single day on all the latest models, right? The best model from Moonshot, the best model from Alibaba, the best model from... there's about five different Chinese models, the best open source models, the best Chinese labs there. We run benchmarks on their models every day and then also the best US open source models, GPT-OSS, Neotron, etc. So we're running these benchmarks every day in an automated fashion and they run on these servers that are dedicated to us for inference benchmarking and we sweep across so many different configurations and optimization types. And then what it creates is—and all the results are public and all the configurations are public. So now we have the Pareto optimal curve because a lot of times when people are comparing inference performance, they're taking a suboptimal curve or point for someone else and comparing it to their optimal one. And it's like, well, yeah, I can make—I can stick, you know, if I drove a Porsche versus some race car driver, obviously I'd drive it slower. The same thing with inference benchmarking. And so what we did is we created open-source basically containers for the optimal points across every point on the interactivity, i.e., how fast is it responding to me versus batch size, i.e., how many users am I simultaneously serving curve? And so now anyone who wants the optimal point can just go to InferenceX, download it, and run that as the optimal point. And they can check every day if they want, or they can even auto-download the most optimal point for that model, and their inference performance will be near peak.

Host

在你看来,那条曲线是最重要的吗?吞吐量-交互性曲线是最重要的。

Is that curve like the most important curve in your opinion? The throughput-interactivity curve is the most important one.

Dylan

是的,我认为硬件基础设施、模型应用层等几乎所有东西都取决于那条曲线,对吧?是要求极快、极低延迟?我不太在乎成本,所以把批大小设得很低,大量使用推测解码或多词预测等技术,有很多可能的技术。还是说实际上我在批量处理大量文档,根本不在乎这些?我不使用那些在成本效率上更差但能提升单个用户速度的技术,因为我只想打包一堆用户。我不在乎文档处理一整晚,对吧?目前我们对待 AI 基础设施的方式是一刀切。但随着时间的推移,我们会达到这样的状态:有批量工作负载,也有需要即时响应的场景,整条曲线对用户都很重要。我们在 Anthropic 也看到了这一点,对吧?Claude Code 快速模式比普通模式贵得多。OpenAI 的优先级队列也是如此。

Yeah, I think most things in hardware infrastructure, model application layer, everything is downstream of that curve, right? Is it something that needs to be super fast, super low latency? And I don't really care about the cost, so I make batch size very low and I use techniques like speculative decoding or multi-token prediction heavily, and there's so many possible techniques there. Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things? I don't use these techniques that actually are worse on cost efficiency but help you with speed for an individual user because I just want to pack a bunch of users. I don't care if the document takes all night to process, right? And right now the way we treat AI infrastructure, it's like one-size-fits-all. But over time, we're going to get to the point where there's stuff where you have batch workloads or you need instant response, and there's the whole curve that's going to matter for users. And so we see this with Anthropic, right? Claude Code fast mode costs way more than regular mode. And same with OpenAI's priority queue thing.

Host

抱歉,问个笨问题。成本如何体现在图表中?假设一个虚构的例子,批大小为 100,每个用户每秒 10 个 token。那么总计每秒 1000 个 token 来自那一个计算单元。这是曲线的一端,非常慢,每秒 10 个 token。另一端是每秒 500 个 token,但只有一个用户。所以可能是每秒 250 个 token,一个用户。中间有一些更帕累托最优的点,对吧?普通人实际上想要每秒 50 或 100 个 token,以及我能批量服务的用户数。所以曲线是:每秒总计 1000 个 token 或 250 个 token,取决于我批处理多少用户,中间有一条曲线。最终,某些工作负载会想要 4 倍的成本降低,因为同样的硬件单元可以处理 1000 个 vs 250 个 token,而有些用户会多付 4 倍,因为我不在乎价格,我在乎时间,因为使用 token 的人很贵,或者这里的反馈循环很贵。

Sorry, dumb question. How does cost factor into the chart? So if I have, let's say imaginary example, I have a batch size of 100, okay? And I can do 10 tokens per second per user. So in total I'm doing a thousand tokens per second off of that one piece of compute. That's one side of the curve. Super slow, 10 tokens per second. The other side is I have 500 tokens per second, but I only have one user. And so maybe 250 tokens per second, one user. And then there's points in the middle that are more Pareto optimal, right? The average person actually wants like 50 or 100 tokens a second and maybe the number of users I can batch together. So the curve is: okay, a thousand tokens total per second or 250 tokens total per second depending on how many users I batch, and there's a curve in the middle. And so ultimately some workloads will actually want the 4x cost decrease because the same unit of hardware can do a thousand versus 250, and some users will pay 4x more because I don't care about the price, I care about time because the person using the tokens is expensive or the feedback loop that I have here is expensive.

Host

如果你必须猜,选择 10 年或 15 年的时间框架。你认为百分之多少的推理算力会在太空中进行?可以是 0%、50%、99%?这很难。你选一个时间框架,比如 10 年,随便什么,然后你……

If you had to guess, you choose the time frame 10 years or 15 years. What percent of inference compute do you think will happen in space? Can be 0%, 50%, 99%? This is a tough one. You choose the time frame like 10, whatever time frame, and you're...

Dylan

所以我认为非共识,或者至少与 SpaceX 相反的观点——顺便说一句,我喜欢 SpaceX,如果能买股票我绝对会买 IPO。不是投资建议。谢谢。不是投资建议。从任何一方来说——我不认为太空数据中心在未来 3 到 5 年内会真正重要。话虽如此,我认为在 20 年内,绝大多数算力将在太空中进行。所以真正的因素在于成本、时间框架、在地面建设电力的成本,以及你在地面能获得多少电力。显然,我对推理算力——你知道,多少吉瓦或太瓦用于推理——的看法,对我来说是一条疯狂的曲线。

So I think the non-consensus, or at least against SpaceX thing—you know, I love SpaceX by the way and I totally would buy the IPO if I could buy stocks. Not investment advice. Thank you. Not investment advice. From either—I don't think that space data centers will really matter in the next 3 to 5 years. With that said, I think in 20 years, I think the vast majority of compute will be going in space. And so the real factor there is sort of the cost, the time frame, the cost of building power on terrestrial land, and how much power you're going to be able to do on terrestrial land. And I think obviously my views of where inference—you know, how many gigawatts or terawatts are devoted to inference—it's a crazy curve for me personally.

Host

你的预测是什么?多少吉瓦?

What's your forecast? How many gigawatts?

Dylan

是的,我认为到 2030 年,仅 OpenAI 和 Anthropic 合计就将超过 100 吉瓦。然后再加上 Meta、谷歌等等。这将是一个巨大的推理算力投入。到 2040 年,将达到太瓦级别,对吧?我们将获得的生产力曲线和推理部署规模将非常巨大。所以如果看 2040 年,我认为超过一半的新增算力将在太空中。但如果看 2030 年,我认为不到 1%。

Yeah, I think by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined. And then you'll add Meta and Google and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. And by 2040, it'll be terawatts, right? The curve of productivity that we're going to get and inference deployments is going to be huge. So if you look at 2040, I think probably more than half of the incremental compute will be going in space. But if you look at 2030, I think it's sub 1%.

Host

你认为每瓦特智能在增加吗?而且我们目前的每瓦特智能与人类生物学之间似乎仍有巨大差距。那么,如果我们正在缩小差距,你认为会缩小吗?如果是,这种提升将从何而来?

Do you think intelligence per watt has been increasing? And it seems like there's still a giant gap between where we are intelligence per watt versus human biology. So if we are, do you think we are to close that gap? And if so, where is that gain going to come from?

Dylan

是的,我认为这通常也取决于你在做什么,对吧?比如 TI-84 在数学计算方面的每瓦特智能远高于我们,而且它已经 30 年了,对吧?显然这是个愚蠢的……

Yeah, I think it often depends on what you're doing too, right? Like a TI-84 is way more intelligence per watt in terms of doing math than us, and it's like 30 years old, right? Obviously this is a dumb...

Host

通用智能。

General intelligence.

Dylan

是的。但就通用智能而言,InferenceX 做的事情之一是我们也测量所有这些硬件的功耗和成本。所以我们不仅提供吞吐量与交互性的对比,还提供成本与交互性的对比。

Yeah. But general intelligence wise, so one of the things InferenceX does is we also measure the power and cost of all of this hardware. And so we offer not just throughput versus interactivity, we offer cost versus interactivity.

每瓦智能与协同优化 Intelligence per watt and co-optimization

Host

我们提供的是算力与交互性。那么,你知道,智能每瓦特一直在提升吗?我提到过,在相同基准水平下,成本下降了 60 倍。我们在智能每瓦特上也看到了同样的趋势。虽然没有正好 60 倍,但接近 40 倍。有些效率提升并非来自功耗方面,但智能每瓦特每年都有巨大改进,至少今年、去年、前年、大前年都是如此。我预计这还会持续。至于与人类大脑的差距,我们还差好几个数量级。幸运的是,这并不重要。我们可以给计算机投入大量电力。给计算机供电比给人类大脑供电容易得多。比如,我们会有疾病、食物偏好和睡眠。

We offer power versus interactivity. And so as far as, you know, intelligence per watt been increasing? I mentioned it's been a 60x cost decrease for same benchmark level. We've also seen the same on intelligence per watt. It's not been exactly 60x, it's been closer to like 40x. Some of the efficiencies are non-power ways, but there's been a humongous improvement in intelligence per watt on an annual basis at least so far this year, last year, year before, year before. And I expect that to continue. As far as where we are from the human brain, we're many orders of magnitude away. Thankfully, it doesn't really matter. We can devote a lot of power to computers. Much easier to power computers than human brains. Like, you know, we have sickness, disease, and food preferences, sleep.

Host

嗯,确实如此。

Yeah, exactly.

Host

我再问一个关于这个总体主题的问题。在我看来,就智能每瓦特或智能每美元这类指标而言,我认为有三个层面的投入。你可以获得硬件层面的改进,让硬件更高效;你可以获得底层系统优化,比如内核级改进、矩阵乘法库之类的;或者你可以获得高层模型级的算法改进。在我看来,过去三年里,大部分收益来自硬件层面,部分来自模型层面。你同意吗?你觉得未来也是这样吗?你认为内核层面还有很大的挖掘空间吗?

Let me just ask one more question on the general theme. In my opinion, in terms of intelligence per watt or intelligence per dollar, any of these metrics, I think there's kind of three levels of input. You can get hardware improvements where the hardware is more efficient. You can get low-level systems optimizations like kernel-level improvements, matrix multiplication libraries, things like that. Or you can get high-level model-level algorithmic improvements. It seems to me that in the last three years, most of the gains have come from hardware level and some from the model level. Do you agree with that? Do you think that's what it looks like in the future? Do you think there's a bunch of juice to squeeze in, say, kernel level?

Dylan

嗯,肖恩,我完全不同意你的看法。很好,这就是我问这个问题的原因。

Yeah, Sean, I completely disagree with you, by the way. Great, that's why I'm asking this question.

Dylan

好的,我认为一种方法是看这三个不同的层面。从这个意义上说,从 Hopper 到 Blackwell,也就是过去三年我们所有的进展,在 DeepSeek 上最优化部署大约提升了 30 倍。在推理方面,你可以看到大约 30 倍的改进。但在过去三年里,智能每瓦特的提升要大得多,其中很大一部分来自模型层面。对吧?如果你回顾三年前,那是 GPT-4。现在呢,比如 Qwen,一个较小的 Qwen 模型,总共 27B 参数,活跃参数 2B,表现要好得多。所以模型层面有巨大改进,硬件层面也有相当可观的改进,但真正重要的是协同设计层面。如果你看看这些模型的架构,但 DeepSeek 是最著名的公开模型,人们都见过。

Okay, so I think one way is to look at these three different layers. In that sense, from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeek on the most optimized deployment. You can see on inference there's about a 30x improvement. But over the last three years, we've had way more improvement in intelligence per watt, a lot of that coming from the model layer. Right? If you look back three years, it's GPT-4. Now it's like, you know, maybe Qwen, one of the smaller Qwen models that's like 27B parameters total and 2 billion active, is way better. So you've got this huge improvement on model layer, you've got this pretty sizable improvement on hardware, but it's that co-design layer that's important. If you look at the architecture of any of these models, but DeepSeek is the most famous one that's public and people have seen.

Host

是的,DeepSeek 通过协同优化或内核级内存优化获得了巨大的效率提升。

Yeah, DeepSeek got huge efficiency gains from co-optimization or kernel-level optimizing memory.

Dylan

是的,我认为这当然是内核层面的,但实际上你是为芯片构建硬件架构。如果你看看 DeepSeek V3 中所有专家的形状,它们都是针对 Hopper 优化的。而 V4 则针对 Blackwell 和华为芯片进行了优化。有趣的是,尽管 TPU 客观上是一款出色的芯片,它运行了 DeepMind 的所有模型,也为 Anthropic 做预训练,但 TPU 在运行 DeepSeek 时表现很差,而它们在运行其他在 NVIDIA 上表现不佳的模型时却非常出色。这种深度优化已经达到了一定水平,无论是形状、网络 IO 模式、如何做集合通信,还是围绕注意力机制的算术强度。所有这些不同方面都在模型、硬件和中间的基础设施软件之间进行了协同优化。很难说你能把收益分离开来。

Yes, I think it's kernels of course, but it's actually you build the hardware architecture for the chip. So if you look at the shapes of all the experts in DeepSeek V3, they were all optimized for Hopper. And if you look at V4, they're optimized for Blackwell and Huawei's chip. What's interesting is despite the fact that TPUs are objectively an amazing chip, and they run all of DeepMind and they do all the training for Anthropic as well on the pre-training side at least, TPUs suck at running DeepSeek, but they are really great at running other kinds of models that don't run well on NVIDIA. There is some level of such deep optimization that has been done, whether it be shapes, network IO patterns, how you do the collectives, how you do things around the arithmetic intensity of the attention mechanism. All these different things are co-optimized between the model and the hardware and the infra-software in between. It's hard to say you can disentangle the gains.

Host

你认为,我的理解是,过去几年中国在这方面做得比西方好得多吗?比如 DeepSeek 是最早真正这样做的模型之一。

Do you think that, like my understanding is that China has done this a lot better than the West the last few years? Like DeepSeek was one of the first models to really do this.

Dylan

我不一定这么认为。我认为更多的是西方不告诉别人他们做了什么。对吧?OpenAI 没有告诉人们 GPT-4 有多稀疏,形状大小是多少,所有这些事情。但 GPT-4 的大小与 DeepSeek V3 大致相同,略小一些,而且我记得 GPT-4 发布得更早一点。

I don't necessarily think so. I think it's more so that the West doesn't tell people what they do. Right? OpenAI didn't tell people that GPT-4 was how sparse it was, what the shape size was, all these things. But GPT-4 is roughly the same size, slightly smaller than DeepSeek V3, and GPT-4 came out a little bit earlier, if I recall correctly.

Host

那么你的观点是,这三件事一直在以大致相同的速度同时发生,而最大的收益来自于协同优化?

So is your view that all three of these things have been happening simultaneously at roughly the same rate, and the biggest gains are when you just co-optimize?

Dylan

我想说,模型层面的收益比协同优化层面、软件基础设施层面和硬件层面都要多。但每个层面都有创新,而真正最大的收益和最优秀实验室的妙处在于他们协同优化了所有三个层面。比如,Anthropic 虽然使用了多种不同的硬件,但他们并不怎么在 TPU 上做推理。他们主要在 TPU 上训练,而在 Cranium 和 GPU 上做大量推理。GPU 更像是一个万金油,但他们优化了硬件、优化了模型、优化了一切,从而能够做到这一点。而 OpenAI,之前的模型更多针对 Hopper 优化,现在则更多针对 Blackwell 优化。随着时间的推移,这些实验室,Google 也一样。他们优化了 Gemini 2,它真正针对 TPU v6e 优化,然后是 Gemini 3,而即将推出的下一代 Gemini 则真正针对 TPU v7 优化。所以很多这些东西都在被协同优化,实际上当你把那个模型放到旧硬件上运行时,效果并不好。所以我认为这种协同优化是最重要的。它被称为软硬件协同设计,这才是真正令人兴奋的地方。比如,我认为我的日常工作很棒,你可以看到一个层面,这里有很多创新,每个层面都有很多创新。真正的突破性创新是当你跨越几个层面,对它们进行协同优化和协同设计时,突然间,你本来可能在这里有 2 倍、那里有 2 倍、那里有 2 倍,原本相乘是 8 倍,但实际上却变成了 100 倍,因为你优化了所有三个层面。

I would say there's been more gains on the model layer than on the co-optimization layer, than on the sort of software infrastructure layer and the hardware layer. But there's been innovations on every layer, and really the biggest gain and the beauty of the best labs is when they co-optimize all three. That's what, like, when Anthropic is, even though they used many different kinds of hardware, they don't really inference too much on TPUs. They mostly train on TPUs, and they inference a lot on Cranium and GPUs. GPU is more a jack of all trades, but they've optimized their hardware, they've optimized their model, they've optimized everything so they can do that. Whereas OpenAI, prior models were optimized for Hopper more, now they're more optimized for Blackwell. And you step forward through time, these labs, and the same with Google. They've optimized Gemini 2 was really optimized for the TPU v6e, and then Gemini 3 was, and then the next Gemini that's coming out is really optimized for TPU v7. So a lot of these things are being co-optimized, and actually when you pull that model and run it on the old hardware, it's really not that great. So I think a lot of this co-optimization is the most important thing. It's called software-hardware co-design, and that's what's really exciting. Like, you know, I think my day-to-day is great, you get to look at one layer, there's all these innovations happening here, there's all these innovations happening on every layer. The real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x because you've optimized across all three layers.

全栈协同优化 Co-optimization across the stack

Host

这就是为什么实验室里看到的景象如此令人兴奋——比如像英伟达这样的公司,它并非严格在模型层进行协同优化,而是从模型层一路向下延伸到硅片层面。或者看看台积电,他们不仅优化制造工艺,还从组件、耗材、工具一直向上游延伸到芯片设计,客户告诉他们的是这种跨多个抽象层的协同优化。

And so that's what's really exciting about sort of like what you see at the labs, which you see at like a company like Nvidia who's not co-optimizing on the model layer per se, but a little bit from the model layer all the way downstream to, you know, silicon. Or you look at a company like TSMC, they're co-optimizing not just, you know, fabrication, but all the way from the components and the consumables and the tools all the way upstream to what the designs, their chips are, the customers are telling them is this co-optimization across many layers of the abstraction stack.

Dylan

不过,这种优化中总会有一些瓶颈滞后,需要被拉上来,然后打上补丁才能运作。

There will always be bottlenecks somewhere in that optimization though that are like lagging behind and then need to get pulled forward, you know, and band-aids to act.

Host

如果你要预测,在堆栈的任何层级——它可能出现在任何地方——未来一年你最密切关注哪些瓶颈?不一定是供应链或规模方面,而是实际的技术层面,当然也可以是供应链。比如是内存改进?还是说就是 Scaling(规模扩张)本身?

If you had to predict like what are at any level of the stack, it can be literally anywhere. What are some of the bottlenecks you're most like you're kind of tracking most acutely the next year? And not necessarily in the supply chain, not in like scale, but in terms of the actual um and it can be in the supply chain too, but just like you know, is it memory improvements? Is it is it that like just like scaling?

Dylan

内存是个老生常谈的话题,但我不打算从供应链角度讲,而是从技术角度。内存容量和带宽提升得非常缓慢。NAND 单元是大约 25 年前发明的,DRAM 单元是大约 40 年前发明的,之后在单元结构上就没有重大突破了——你知道 NAND 单元其实就是个很简单的门,DRAM 单元也是。未来可能会有一些极具创新性的东西出现。但即便在过去五年里,我们真正做的也只是让 HBM 堆叠更多层、速度更快。不过未来几年会有新的创新:不再把 HBM 单独堆叠在芯片旁边,而是直接把内存堆叠在芯片上,这样带宽会爆炸式增长。这个领域有一些有趣的公司,也在做一些有趣的 PC 尝试。我认为内存带宽是最大的瓶颈之一。

So memory is an easy one that everyone's talked about, but I'm not going to talk about from a supply chain angle. I'm talking about from a technology angle, right? Memory capacity and bandwidth have been improving very slowly. The NAND cell was invented like 25 years ago. The DRAM cell was invented like 40 years ago and there's been no major breakthrough in cell like you know how what a NAND cell is. Obviously NAND is like a very simple gate or DRAM cell. There is stuff that could come down the pipeline that could be hugely innovative. But even over the last, you know, five years, all we've really done is make the HBM, you know, more stacks, faster, but actually there's like new innovations coming in the next few years where instead of, you know, stacking the HBM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode. Um, and so there's interesting companies in that space and interesting PCs that companies are trying to do there. I think like memory bandwidth is one of the biggest.

Dylan

另一个瓶颈是,至少在过去二十年里,硅芯片的功耗基本可以一眼看出:数据中心或桌面芯片的功耗峰值大约是每平方毫米 1 瓦。所以如果芯片是 100 平方毫米,功耗通常在 100 瓦左右或略低。看看最新的英伟达和谷歌 TPU 芯片,仍然在这个范围内——每平方毫米 1 瓦。现在芯片功耗已经达到 1400 瓦,下一代英伟达芯片(比如 Reuben)将达到 2000 瓦,再往后 Reuben Ultra 可能达到 4000 瓦左右。但实际上,这主要是通过增加硅片面积实现的。令人兴奋的是,我们现在终于开始尝试——而且已经在研发中——让芯片实际承受的功率远超每平方毫米 1 瓦。这意味着你需要的硅片更少。显然,它运行在更高功率下,某些情况下效率更低,但你减少了硅片用量,同时还要处理热问题、电气干扰等各种问题。这就是为什么这是个棘手的工程难题,也是我们一直卡在 1 瓦左右的原因。但令人兴奋的是,全世界都在努力改变这些。

Another one is um for the history of like silicon basically for the last two decades at least you know how many watts a chip is can be easily predicted just by looking at it for for a data center or desktop chip it it peaks up at one watt per millimeter squared and so if a chip is 100 millimeter squared generally the power consumption is around 100 or a little bit less um and if you look at the newest Nvidia silicon the newest TPU silicon it's still on that range of one watt per millimeter squared so you know chips are now getting to you know, 1400 watts. Next generation is 2,000 watts for Nvidia. Um, with Reuben and such. Uh, and and you move forward to Reuben Ultra, it's going to be like 4,000 watts or something like that. But really, there's increasing the amount of silicon. What's exciting is we're now finally doing things and it's in development right now where you actually can pump the amount of power into the silicon uh to be way more than one watt per millimeter squared. And now that all of a sudden means you need less silicon. Obviously, it's running at higher power. It's less efficient in some cases, but you reduce the amount of silicon and you're able to like thermal issues, thermal issues. Um there's uh interference of like electrical interference issues. There's all sorts of different issues uh that crop up and that's why it's a hard engineering problem. That's why we've stuck at about one. But what's exciting is the world is trying to change these things.

Dylan

有趣的是,在供应链的另一端,人们常说能源是个难题,存在能源瓶颈。但确实有一些非常简单的解决方案。比如,美国有制造数百万台卡车柴油发动机的能力,你可以很轻松地在生产线上把它们改装成燃气发动机,然后连接到一个电机上反向驱动——让电机发电,而不是让电机驱动车轮。这样,通过向美国能制造数百万台的设备中注入燃气,你就产生了电力。然后,你可能会说,维护起来很麻烦吧?因为一个数据中心站点需要数百台这样的设备。但实际上,你只需要从汽车修理厂招些人,让他们到处维修卡车发动机就行了。其实这相当简单——我不想说它简单,我自己做不了。

I think interesting like in a different part of the supply chain it's sort of like you know people people will talk about like energy is hard and you know we have energy bottlenecks and it's like yeah but there's actually like very simple solutions you know one could think of right um take the millions of diesel engines for trucks that the US has the capacity to make um you can very trivially convert them to be using for gas uh in the assembly line and then stick them up to a electrical motor like back driving it so the electrical motor generates electricity rather than the electrical motor causing the the rotation of the wheel, for example, but doing it the opposite direction. And now you've generated electricity by pumping gas into something that us can make millions of. Um, and then, okay, well, that sounds like a pain in the ass to uh service, right? Because now you have to have hundreds of these on a data center site. Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines. Actually, it's actually pretty trivial to not I don't want to say it's trivial, I couldn't do it. Um

Host

我觉得你说得很对。因为过去二三十年西方并没有真正关注半导体乃至更广泛的硬件领域,所以创新不多。最聪明的人都在想怎么改进这些——但为什么要去搞硬件呢?当你可以做广告的时候……没错,就是这样。

I think you're making a really good point which is that like because the west wasn't really thinking about semic even hardware more broadly the last 20 30 years we didn't have like much innovation we'd have the best minds like thinking about how do you improve these why would you why would you want to go work in hardware when you can uh make ads to ads yeah exactly

Host

嗯,好吧,我特别想问:英伟达对比 TPU,你怎么看?

Um okay I'm dying to ask Nvidia versus TPU what are your thoughts

Dylan

嗯,我觉得每个人都想在这两者中选一个,但这其实取决于具体情况。你看,两年后,谷歌将通过其供应链制造超过 1000 万块 TPU,而英伟达将制造数千万块 GPU。两者都将达到千亿美元级别——谷歌每年创造的 TPU 价值超过 1000 亿美元,英伟达可能是 5000 亿或更多。我不是在做具体预测。

Um I think I think like everyone wants to pick one or the other for this, but it's really like a function of like look, you know, you look two years from now, Google's going to make 10 plus million TPUs and through their supply chain and Nvidia is going to make, you know, many more million tens of millions of GPUs and both are going to be 100 plus billion dollar, you know, well, Google's going to be 100 plus billion dollars, you know, of TPU created a year and and Nvidia will be, you know, 500 plus or, you know, whatever. I'm not making a specific estimate.

Host

这不是收入预测,只是个思想实验。

This is not revenue forecast. This is just a thought experiment.

Dylan

对,或者说研究。

Yeah. Or research.

Host

你一直在做媒体转换。

You've been media trans.

Dylan

当然。你知道,正在为 SpaceX 的想法做准备。

Absolutely. You know, getting ready for the SpaceX idea.

Host

嗯,你们在 SpaceX 投了很多吗?好吧,那说得通了。嗯,

Um, are you guys big in SpaceX? Okay, so that makes sense. Um,

Dylan

我们很幸运,是很大的投资者。

We're very lucky to be very large investors.

Host

太棒了。太棒了。嗯,所以我想说,谷歌 TPU 和英伟达 GPU 各有优势,对吧?英伟达会说:“我们有交换机,而且是通用型的。”而 TPU 会说:“我们更优化,实际上更节能,我们的网络针对某些网络架构做了更多优化。”

Awesome. Awesome. Um, so I would say um the the case of sort of like Google TPUs versus uh Nvidia GPUs, they both have like points that are really like in their favor, right? You know, Nvidia will be like, "Oh, well, we have switches and we're general purpose." And and TPUs will be like, "Well, we're more optimized. actually more energy efficient and our network is actually more um optimized for certain types of network architectures."

软硬件协同设计与模型架构分化 Hardware-Software Co-Design and Model Architecture Divergence

Host

所以你看,这些对立观点都能深入探讨,我可以一本正经地跟你争论 GPU 比 TPU 好得多,或者 TPU 比 GPU 好得多,但这归根结底是硬件和软件的协同设计。实际上,按照 OpenAI 模型的发展方向,他们用 TPU 可能是个糟糕的决定。而按照 Anthropic 和 Google 模型的发展方向,他们用 GPU 训练可能也是个糟糕的决定。那根本区别是什么?

And so you have like these counterpoints that both would really get into and you know I could with a straight face argue with you like that GPUs are way better than TPUs or TPUs are way better than GPUs but it comes down to hardware software codesign. So actually the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs potentially. And the way that Anthropic and Google's models are headed, it's actually a terrible decision potentially for them to train with GPUs. I mean, it'd be fun to what's the fundamental difference there?

Dylan

有很多不同点,比如矩阵乘法单元的大小就是一个很简单的差异。因此,你做的矩阵乘法的形状、使用的注意力机制、注意力机制的结构、专家的结构都会不同。所以你认为 OpenAI 和 Anthropic 正在走向非常不同的模型架构。我认为它们的模型架构确实很不同。实际上,OpenAI 的模型更稀疏,这有好处。而 Anthropic 的模型虽然也稀疏,但总体上更稠密,这有另外的好处。还有很多其他因素,比如网络拓扑。Nvidia 的所有芯片都通过 NVLink 交换机连接。而 Google 没有交换机。但他们的做法是,Nvidia 的 NVLink 只能连接 72 个 GPU,而 Google 的 ICI 可以以超高带宽连接 8000 个芯片,但数据必须经过其他芯片才能到达目的地,因为没有交换机。所以这里有权衡取舍,有优点也有缺点,这会影响模型架构。不一定非要说哪个更好,因为归根结底,你无法孤立地衡量它们,因为这种影响会延伸到模型层。

There's various things, right? Like the size of the matrix multiply unit is different as a very simple thing. And therefore, the shape of the matrix multiply you do, the attention mechanism you use, the way that attention mechanism is structured, the way the experts are structured. So you think OpenAI and Anthropic are converging on very different model architectures. I think they have quite different model architectures. In fact, OpenAI's are much more sparse, and that has benefits. And then Anthropic's, they're still sparse, but more dense in general, and that has different benefits. And there's many other things, right? The network topology, right? Nvidia, all of their chips are connected to switches, NVLink switches. For Google, they have no switch. But what they've done is they've been able to, you know, Nvidia, the NVLink can only connect 72 GPUs. For Google, their ICI can connect 8,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch. And so there's trade-offs there. There's positives and negatives and that influences the model architecture. It's not necessarily that you should claim one is better than the other because at the end of the day, how do you say that this is better than that when you can't measure them in isolation because it also extends up to the model layer, right?

CUDA护城河与生态变迁 CUDA Moat and Changing Ecosystem

Host

但我记得很长一段时间以来,我一直认为 Nvidia 的可编程性和 CUDA 是一个巨大的护城河。在我看来,至少在过去三到六个月里,这种说法已经改变了。模型公司不再关心是否必须为其他芯片编写自定义内核,如果需要,他们可以同时使用四五种芯片。Claude 和 Codex 实际上非常擅长做很多优化工作。而且,现在并不是有一万家模型公司每家都需要可编程性,大概只有几十家模型公司。所以在我看来,如果基本前提是成千上万的大客户需要 CUDA 兼容性,那么这种论点似乎正在改变。

But I remember for a long time thinking you know one the programmability of Nvidia and just CUDA as such a big moat. It seems to me that narrative has kind of changed at least in my mind for the last three six months like model companies no longer care about if we have to write custom kernels for you know this other chip so be it. We'll work with four or five chips if we have to. Claude and Codex are actually quite good at doing a lot of that optimization work. And so it seems like some of the and then it's you know it's not like there's 10,000 model companies that are each you know each need programmability. There's on the order of tens maybe model companies and so it seems to me that like if you the fundamental premise of like tens of thousands of big customers that need CUDA compatibility like it seems that kind of thesis is changing in the last.

Dylan

是的。我的意思是,CUDA 护城河和软件护城河至少部分被解构了,因为模型非常擅长编码,所有软件在这种情况下都变得商品化了。我确实认为存在一定程度的开源,人们所说的 CUDA 护城河实际上与 CUDA 无关,而是因为 DeepSeek、Kimi、智谱、阿里巴巴、腾讯、小米(最近有个很棒的模型)等公司的模型都是为 GPU 协同设计的,所以如果我想在 TPU 上运行它们,实际上在某些情况下它们运行得并不好。现在 Google 必须创建自己的开源模型生态系统或开源模型本身,所以他们有 Gemma 模型。所以最终,这其实不是 CUDA 作为护城河,而是下游产品更针对 Nvidia 优化。在这些情况下,这些公司只是开源了模型,或者像 Neotron 那样开源了模型,然后用户(例如推理 API 提供商、试图将开源模型定制化用于企业业务用例的强化学习公司)都受制于这样一个事实:好吧,我想我需要用 Nvidia,因为生态系统用的是 Nvidia,即使我并不特别关心编写 CUDA 内核,因为模型很擅长这个。但问题在于,专家的维度是这样,隐藏维度是那样,所以它在 Nvidia GPU 上运行得比 TPU 好,反之亦然。如果 Google 真的开源了非常好的模型,情况也会一样:人们会拿他们的模型,然后说“哇,这些在 Nvidia GPU 上运行得不太好,我应该租 TPU 或买 TPU 来跑”。

Yeah. I mean certainly the CUDA moat and software moat is at least partially disentangled because models are just great at coding and all software gets commoditized in that case. I do think there is some level of like open source and you know what people call the CUDA moat is not actually anything to do with CUDA but it's like the fact that DeepSeek, Kimi, Zhipu, Alibaba, Tencent, all these companies, Xiaomi had an awesome model recently, their models are co-designed for GPUs and therefore if I want to run them on TPUs actually in some cases they don't run really well on TPUs. Now Google just has to create their own open source model ecosystem or open source models themselves so they have the Gemma models and so you end up with like well that's not really CUDA as a moat it's that the downstream product is more optimized for Nvidia and in these cases these companies are just open sourcing them or like Neotron is just open sourcing it and then the users of it for example the inference API providers, the RL companies that are trying to take open models and customize them for company's business use cases, all these different companies are downstream of the fact that like okay well I guess I need to use Nvidia because the ecosystem uses Nvidia even though I don't particularly care about writing CUDA kernels because the models are great at that, but it's like the shape of like well this expert the dim is this and you know the hidden dimension blah blah blah is this right and so therefore it's better to run on Nvidia GPUs than it is on TPUs and vice versa right if Google were to actually open source really good models you know this would be the same thing right people would take their models and they'd be like oh wow these don't run that well on Nvidia GPUs I should actually just rent TPUs or buy TPUs and do it on there.

Host

对于小团队来说,你会想用所有的开源软件,比如 vLLM、SGLang、PyTorch 等等。但大实验室不一定需要所有这些。OpenAI 很久以前就 fork 了 PyTorch,Anthropic 和其他公司也不一定严重依赖这些开源实现,他们已经 fork 或自己构建了。所以他们不需要依赖开源。因此现在更像是,我会选择最好的硬件,然后为那个最好、最具成本效益的硬件从头到尾协同设计我的模型和基础设施软件。

For small teams you're going to want to use all the open source software like vLLM, SGLang, PyTorch, all that stuff but the big labs they don't necessarily need to use all that right. OpenAI forked PyTorch long ago and you know Anthropic and all these other people don't necessarily rely heavily on the open-source implementation of these things they forked things or built it on their own already and so they don't need to rely on the open source and therefore now it's more like you know I'll choose the best hardware and I'll co-design my model and infrastructure software through and through for that hardware that is the best and most cost efficient.

Dylan

而且我会让 AI 帮我编写所有这些软件。

And you know I'll have AI help me write all that software.

对Cerebras与快速推理的思考 Thoughts on Cerebras and Fast Inference

Host

你怎么看 Cerebras?

What do you think of Cerebras?

Dylan

我认为 Cerebras 是一家非常有创新力的公司。在市场的某些领域,他们非常出色。推理速度非常快。我认为这是一个很大的市场。我们在 SemiAnalysis 几乎只使用快速模式。

I think Cerebras is a really innovative company. I think in some spots of the market they're really really good. Very fast inference. I think that's a big market. We use fast mode almost exclusively at SemiAnalysis.

Host

顺便说一句,我很喜欢你们在核算方面有多严谨——我不知道那是你们做的一个展示还是你们一贯的做法——但你们会核算每项任务的美元花费和投资回报率。很棒的分析。

By the way I love how disciplined you've been about accounting for I don't know if that was one exhibit you did or if you do it consistently but accounting for the dollar spent and the ROI on each task. Awesome analysis.

Dylan

是的。我们做得很认真,谢谢。那是我们写的“暗黑 GDP”文章。我们还按天追踪每个人的 token 消耗,如果有人突然飙升,我会问“你做了什么?”然后说“好的,谢谢告诉我,这看起来值得。”然后继续我的一天。我认为快速模式显然对高端任务很有价值。我可以看到很多不同的用例,超快 token 是值得的。我也能看到另一面,有很多用例不需要超快 token,因此市场不会为此付费,他们会改用 GPU 和 TPU。

Yeah. We do it pretty diligently and so thank you. That was the dark GDP article that we wrote. And also like track everyone's token spend by day and if someone's like spiked up I'm like what did you do? It's like okay thank you for telling me that that seems worth it. Cool. On with my day. I think fast mode is obviously worth a lot for high-end tasks, right? I could just see so many different use cases where super fast tokens are worth it. I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and therefore the market won't pay for them and they'll use GPUs and TPUs instead.

Cerebras与大模型推理挑战 Cerebras and large model inference challenges

Dylan

我认为 Cerebras 面临的最大风险是:最好的模型才是你想用快速模式的,而小模型可能不需要。在金融市场,比如 Jane Street 的高频交易或中频交易,情况可能不同。但归根结底,在基于 SRAM 的芯片(如 Cerebras 和 Groq)上运行超大规模模型并处理超长上下文非常困难。如果模型变得太大怎么办?如果 OpenAI 的模型不是数千亿或数万亿参数,而是 10 万亿以上,我认为 Cerebras 将无法容纳。再加上长上下文,比如百万级上下文,这就更难证明了。到目前为止,我们看到实验室的大部分收入和用量都集中在它们最好的模型上,即使模型价格上涨也是如此。有数据显示,尽管 Fable 今天才发布,但很多人已经转向了 Fable 和 Mythos 这类更高级的模型,尽管它们贵得多。

I think the big risk for Cerebras is that the best models are the ones you want to use fast mode on, and small models you might not. I could see that being wrong for financial markets or something like Jane Street high-frequency trading or medium-frequency trading. But ultimately, running really large models at very long context is very difficult on SRAM-based chips like Cerebras and Groq. So what happens if models get too big? If OpenAI's model is not hundreds of billions or low trillions of parameters but actually 10+ trillion parameters, I don't think that will fit on Cerebras. And with a long context length, like a million context, that makes it really difficult to justify. So far, we've seen the bulk of revenue and usage at the labs be on their best model, even when the model price has gone up. There's data showing that even though Fable just released today, many people switched to Fable and Mythos, the next tier model, even though it's way more expensive.

Host

那是按美元计量的量还是按 Token 计量的量?

Is that volume by dollars or by tokens?

Dylan

嗯,谁在乎按 Token 计量的量?关键是美元。如果我不在乎卖出了 20 万辆 Mini Cooper 或丰田凯美瑞,而福特 F-150 的均价是它们的 5 倍,销量只有一半,那么美国最赚钱的市场是皮卡。虽然有点开玩笑,但……

Well, who cares about volume by tokens? It's about the dollars. If I don't care that there are 200,000 Mini Coopers or Toyota Camrys sold if Ford F-150s have 5x ASP and sell only half as much, then the most lucrative market is pickup trucks in America. Mostly being facetious, but...

Host

我认为这正是你做得非常好、并且让你与众不同的地方:除了技术,你还非常关心经济学。很少有人能很好地把这两者结合起来。

I do think this is one of the things you've done so well and differentiates you from almost everyone else: you care so much about the economics in addition to the technology. Very few people bridge those two things well.

Dylan

我觉得在 Semi Analysis 内部非常有趣,因为我们有 90 个人,其中很大一部分是覆盖整个供应链的技术工程师,还有很大一部分是前对冲基金人士。你会看到这样的争论:有人说“哦,那不重要”,然后有人说“但成本呢”,接着工程师说“不不,这项技术最酷”。你会看到这种有机的争论。我们相当不拘礼节,考虑到我曾经是版主,你可以想象这有多有趣。

I think it's really fun inside Semi Analysis because we have 90 people, a big chunk are technologist engineers across the whole supply chain, and a big chunk are people formerly at hedge funds. You see these arguments: people say 'oh that doesn't matter,' then someone says 'but cost,' then the engineers say 'no no, this technology is the coolest.' You see this organically fight it out. We're pretty informal, and given that I was a moderator, you can imagine how enjoyable it is.

Host

不要和猪摔跤,因为猪乐在其中。

You don't wrestle with a pig because a pig enjoys it.

Dylan

没错。就这个话题,半导体领域有没有什么让你特别恼火的话题?比如有人说“内存是瓶颈”——虽然这是事实,但最让我恼火的是有人说“AI 没有 ROI”。这让我很愤怒。或者否认模型进步:有人说模型没有变得更好,它们不会推理,会走向死胡同和平台期。但能力曲线一直在向右上方延伸。他们说“看,这个基准测试没有提升”——那是因为它已经达到 90% 了,看看新的基准测试,你把它饱和了,现在它们正在飙升。我认为这才是问题所在。半导体非常复杂,我不怪人们缺乏理解。我每天都从别人那里学到关于半导体供应链的新东西,我从 12 岁开始管理论坛,已经研究了 18 年。但即便如此,抽象层还有很多层。昨天我了解到一种新的化学品,销售额达 1 亿美元,我想“哇,不知道还有这种东西”。它是必不可少的,每个芯片都需要它。有上千个工艺步骤。我觉得最有趣的是,当人们掌握了所有事实,却得出了完全错误的结论。

Exactly. Just on this topic, are there trigger topics in semis for you? Like if someone says 'memory is the bottleneck' — I mean it's true, but the one that really gets me is people saying 'AI has no ROI.' That infuriates me. Or denying model progress: people say models aren't getting better, they can't reason, they're going to dead-end and plateau. But the line has been up and to the right in terms of capabilities this entire time. They say 'look, this benchmark didn't improve' — that's because it was at 90%, look at the new benchmark, you saturated it, now they're skyrocketing. I think that's more the issue. Semis are really complex, and I don't fault people for lacking understanding. I learn stuff every day about the semiconductor supply chain from people, and I've been studying it for 18 years since I started moderating forums when I was 12. But even then, there are so many layers of the abstraction stack. I learned about a new chemical that does a hundred million dollars of sales yesterday, and I thought 'whoa, didn't know this one existed.' It's essential, and every chip requires it. There are a thousand process steps. What I think is most funny is when people have all the facts in front of them and then get the conclusion completely wrong.

Host

这在我们的工作中也经常发生。

That happens in our job all the time too.

Dylan

是的。我认为我的态度不是因为你那样做而生气,而是尽快去做。

Yeah. I think my attitude is not to be mad that you do that. It's to do it as fast as possible.

Host

这个行业现在非常重要,而且有很多短期瓶颈。我们谈了很多短期的事情。有没有什么长期的事情让你很兴奋,比如 10 年时间框架?我们谈到了轨道数据中心,但硅基呢?在 10 年时间框架内,它们是被低估还是被高估了?

The industry is so important right now, and there are so many near-term bottlenecks. We talk a lot about the near-term. Are there longer-term things you're really excited about, say on a 10-year time frame? We talked about orbital data centers, but what about siliconics? Are they underrated or overrated on a 10-year time frame?

Dylan

是的,我认为太空在 10 年时间框架内非常酷——太空数据中心、小行星采矿等等。我对 SpaceX 的愿景非常兴奋。再次强调,不是投资建议。在半导体方面,当事情提前或推迟一年发生时,市场会出现巨大的变动。共封装光学:每个人都知道它会在本十年末实现,争论在于 2027、2028、2029、2030 年,但终将实现。更有趣的是像 Navian Ral 的公司这样的——你们投资了吗?

Yeah, I think space is super crazy awesome in the 10-year time frame — space data centers, mining asteroids, all these things. I'm super excited about the vision of SpaceX. Again, not investment advice. On the semiconductor side, tremendous market movements can happen when things happen one year later or sooner. Co-packaged optics: everyone knows it's going to happen by the end of the decade, the debate is 2027, 2028, 2029, 2030, but at some point it will happen. The more interesting thing is companies like Navian Ral's company — I did you guys invest in that?

Host

我们投了。

We did.

Dylan

好的。所以我认为他正在同时尝试在硅层、软件抽象层和模型层进行创新,而且他完全明白这不是一个两年时间框架。

Okay. So I think he's trying to innovate on the silicon layer, the software abstraction layer, and the model layer simultaneously, and he fully understands that it's not a two-year time frame.

Host

不是几年时间框架。

It's not a few year time frame.

长期押注与遇见Deaveen Long-term bet and meeting Deaveen

Dylan

这是一个长期赌注。像这样的东西,比如我们要一次性引入模拟计算、基于能量的模型等等所有这些疯狂的东西,这很令人兴奋。可能不会成功,但你知道,这很令人兴奋,我真的很期待。

It's a long-term bet. And stuff like that is like, okay, we're going to bring potentially analog compute with energy based models and all this crazy stuff all at once. That's exciting. Probably won't work, but you know, that's exciting and I really look forward to.

Host

肯定不能很快成功。

Definitely won't work quickly.

Dylan

对,我应该说肯定不能很快成功。我相信 Deaveen。我很早就认识他了,他是我在行业里最早认识的人之一,有趣的是,在 2020 年或 2021 年。实际上是 2020 年。这很能说明他的为人。他总是试图帮助年轻一代,发掘人才。

Yeah, definitely won't work quickly is what I should say. I believe in Deaveen. I met him very early, I think he's one of the first people I met in the industry, funnily enough, in 2020 or 2021. Actually 2020. It says something about him. He's always trying to help the younger generation, trying to identify talent.

Host

我在网上钓他,就这样联系上了。

I baited him on the internet. That's how we connected.

Dylan

他在 Mosaic 项目上也非常超前。我记得有人向我推销过。

He was also so ahead of his time with Mosaic. I remember getting pitched.

Host

不,那是 2019 年。我当时其实还是匿名。我在网上钓他,他开始回复,然后我转到私信,再转到电话。那是我在整个半导体行业里第一个真正重要的对话。挺有意思的。

No, it was 2019. I was still anonymous then actually. I baited him on the internet and he started replying, then I took it to DMs and then to a call. That was the first really important person I talked to in the entire semiconductor industry. Funny.

Dylan

抱歉打断一下。

Sorry to interrupt.

Host

真有意思。你认为生态系统的终局是什么?你觉得每个实验室、每个超大规模云服务商都会有自己的芯片吗?训练现在似乎已经可行了,对吧?所以你认为最终每个实验室、每个超大规模云服务商至少推理会用自研芯片,而训练可能还是找英伟达或其他厂商?你认为终局是什么?

That's funny. What do you think is the end state of the ecosystem? Do you think every lab, every hyperscaler just has its own chips? Train seems like it's now working, right? So do you think we end up with every lab, every hyperscaler having its own chips at least for inference and then maybe for training you go to Nvidia or whoever? What do you think is the end state?

Dylan

我认为每个人都会尝试,然后放弃尝试。最终,供应链很重要。你能引入什么技术很重要。随着行业规模扩大,供应链会多元化。现在每个人的芯片或多或少看起来都一样:中间一个大逻辑计算 die,左右是 HBM,顶部是网络,底部是 PCIe 和其他 IO。Trainium、TPU、英伟达芯片都是完全相同的结构。大多数初创公司(Groq 和 Cerebras 在做奇怪的东西,这很酷)也类似。随着我们向前发展,硬件架构和模型架构会进一步分化,因此人们会进行协同优化。有些会陷入局部最小值。如果这像梯度下降,人们试图达到最优解,有些人会冲向局部最小值,然后问题是如何跳回全局最小值。在某种程度上,英伟达的芯片总是比任何其他芯片更通用,至少在并行 AI 计算方面是这样,因为他们有太多关心不同事物的客户,会在设计中提供反馈。最小值总是比他们好,但那个最小值是局部最小值吗?比如,TPU、Trainium、Groq 或 Cerebras 的设计是针对这里优化的,但最终状态实际上需要去那里,所以它们错了?也许它们在一段时间内很好,但最终会出错。这才是真正的问题。我认为通用 AI 计算会有一个大市场。实验室里的人甚至不知道一年后他们会用什么架构。他们有赌注,很多研究赌注,但他们不知道方向。他们知道自己有什么硬件,并试图协同优化,但如果模型架构出现新突破——比如用别的东西替换注意力机制——最好的硬件就会改变。那么人们会仅仅基于更专门的 ASIC 进行五年期的硬件投资吗?还是他们会保留一些通用计算?你可以看到谷歌以每小时 11 美元的价格向 xAI 出租 GPU。这很疯狂。尽管他们有 TPU,但为什么这么做?谷歌实际上有三个不同的 TPU 设计项目。他们与博通合作制造 TPU,与联发科合作的 TPU 架构不同,第三个与前面两个非常不同。所以人们认识到局部最小值可能发生。我认为每个人都会有自己的 ASIC 项目,投入数十亿甚至数百亿美元。谷歌每年在自己的 ASIC 上投入数千亿美元。但他们也会有不用 TPU 的工作负载。谷歌的一些非 Gemini 或 DeepMind 的赌注主要使用 GPU,而不是 TPU。有些也主要使用 TPU。对于药物发现或 Waymo,你可能不想用 TPU。有不同的架构赌注和不同的 AI 路径。科学 AI 可能具有与通用智能 AGI 模型不同的算法模式。所以我认为我们会看到多样性继续扩散。因为市场变得如此之大,利基市场会被开辟出来,使得公司能够拥有自己的利基并真正赚钱,即使大部分份额被英伟达、TPU 和 Trainium 占据。

I think everyone will try and then stop trying. Ultimately, supply chains matter. What technology you can bring in matters. As the industry gets bigger, supply chain diversification happens. Right now everyone's chip more or less looks the same: a big logic compute die in the center, HBM on the left and right, networking on top, PCIe and other IO on the bottom. That's the exact same structure for Trainium, TPU, Nvidia chips. Most startups (not Groq and Cerebras, which are doing weird stuff, that's cool) are similar. As we step forward, we'll get more bifurcation of hardware architecture and model architecture, so people will co-optimize them. Some will end up in local minima. If this is like gradient descent, people try to go to the most optimized solution; some will race to a local minima, and then the question is how to leap back to the absolute minima. To some extent, Nvidia will always be more general purpose than anyone else's chip, at least on a parallel AI compute basis, because they have so many customers who care about different things and give feedback in the design. The minima will always be better than them, but is that minima a local minima? Like, is the TPU or Trainium or Groq or Cerebras design optimized for here, but in the end state you actually need to go over there, so they're wrong? Maybe they're great for a little bit of time, but then they end up being wrong. That's the real question. I think there will be a big market for general purpose AI compute. People at labs don't even know what architecture they'll be doing in a year. They have bets, many research bets, but they don't know where it's going. They know what hardware they have and try to co-optimize, but if a new breakthrough happens on model architecture—like replace the attention mechanism with something else—the best hardware will change. So will people make five-year investments on hardware solely on a more specialized ASIC? Or will they have some bucket of more general purpose compute? You see this with Google paying $11 an hour per GPU to xAI for GPUs. That's insane. Despite having TPUs, there are questions why they do that. Google actually has three different design programs for TPUs. They're making a TPU with Broadcom, a different architecture than the TPU with MediaTek, and a third one that is very different from the first two. So people recognize that local minima can happen. I think everyone will have their own ASIC program, deploy billions or tens of billions of dollars. In Google's case, hundreds of billions a year of their own ASICs. But they'll also have workloads that don't use TPUs. Some of Google's bets that are not Gemini or DeepMind primarily use GPUs, not TPUs. Some also primarily use TPUs. For drug discovery or Waymo, you might not want to use TPUs. There are different architecture bets and different paths for AI. AI for science may have different algorithmic patterns than general intelligence AGI models. So I think we'll see diversity continue to proliferate. Because the market has gotten so big, niches will be carved out, making it possible for companies to have their niche and actually make money even if the majority of the pie goes to Nvidia, TPU, and Trainium.

数据中心建设与算力紧缺 Data center buildout and compute crunch

Host

好。说得太好了。我们能谈谈数据中心建设吗?从各方面来看,如果你看图表,每算力小时的美元成本,我们正处在一个疯狂的算力紧缩之中。

Yeah. Okay. Love that. Can we talk about the data center buildout? It seems like by all accounts, if you look at the charts, dollars per compute hour, we are in the middle of a crazy compute crunch.

算力紧缺与模型改进 Compute crunch and model improvement

Host

嗯,这看起来既是需求侧也是供给侧的问题,对吧?对长智能体的需求飙升,而供应方面,所有这些数据中心的建设都延迟了。嗯,你认为我们在可预见的未来会处于算力短缺中,还是会在某个时候缓解?

Um and it seems like it's both a demand and supply side crunch, right? Demand for long agents skyrocketing, supply, all these data center buildouts are delayed. Um, do you think we're in a compute crunch for the foreseeable future or do you think it alleviates at some point?

Dylan

是的,每个季度我们部署的算力都比上一季度大幅增加,建设的数据中心也比上一季度多。嗯,今年将有 20 吉瓦,即使考虑到延迟,明年将有超过 30 吉瓦,也考虑了延迟。嗯,当然,任何事情都可能出现延迟,对吧?任何硬件都可能延迟。这就是生活的现实。我们会在余生都面临算力短缺吗?这取决于模型的发展。但就像 Mythos 的 TAM,你知道,Mythos 5、Fable 5 不仅仅是 Opus 的两倍,对吧?模型好得多,能完成的任务多得多,所以 TAM 大得多。然而,世界上的算力在过去六个月并没有翻倍,对吧?从 Opus 45 发布到现在大概七八个月。巨大的改进,46、47、48 都是改进,但 Fable 和 Mythos 是阶跃式的改进。世界算力在同一时期并没有翻倍或翻四倍,但 AI 能完成的有用任务的需求,有用任务的数量和价值却增加了。所以现在的问题是会发生什么。显然,Anthropic 在第二季度实现了盈利,扣除股权激励后的净利润盈利。嗯,我认为到第三季度,他们甚至可能包括股权激励在内也盈利。这就是他们盈利的程度。他们在 Opus token 上的利润率,至少 Opus 48 token 的 API 价格利润率超过 80%。他们有很多交易,由于 Bedrock 和 Vertex 等合作方式,公司整体毛利率被略微压低。但最终,他们的每 token 利润率非常高。那么,如果没有上限,他们有能力最终以高于市场价的价格购买每一块 GPU。你知道,他们还从 SpaceX 以高于市场价的价格购买了 GPU,这低于谷歌的价格,但那是因为他们签得早。嗯,你知道,其他公司,比如风险投资支持的公司或利润率不高的公司,不一定能做到,对吧?成本效益比是怎样的?比如,我因为算力不足而租用的每一块 GPU,我可以立即将其转化为 token 销售,或者每一块 TPU 或每一块 Tranium,我都可以立即以正利润率销售 token。如果我的毛利率是 75%,而算力成本翻倍,那也没问题,我仍然有 50% 的毛利率。对他们来说,增加更多计算节点不一定是需要人力的任务,如果他们是租用的话。所以最终,我的净营业收入仍然上升,对吧?所以我会以任何价格租用 GPU,在某种程度上,我想付多少钱就能付多少钱。

Yes, every quarter we're deploying vastly more compute than the prior quarter and there's more data centers built than the prior quarter. Um, this year there's going to be 20 gigawatts, even accounting for the delays, and next year there's going to be more than 30 gigawatts accounting for the delays. Um, of course delays happen on everything, right? Anything hardware can have a delay. That's just the reality of life. Are we gonna have a compute crunch for the rest of our lives? It depends on what happens with models. But like the TAM for Mythos, you know, Mythos 5, Fable 5 is not just like 2x that of Opus, right? The model is so much better and it can do so many more tasks that the TAM is way larger than that. And yet compute in the world did not double in the last, you know, six months, right? From, you know, Opus or maybe like seven or eight months since Opus 45 launched to now. Huge, you know, 46, 47, 48 were improvements but Fable and Mythos were like a huge step function improvement. The world's compute did not double in that or quadruple or whatever in that same time frame, but the demand for useful tasks that can be done by AI, the number of useful tasks and the value of them that can be done by AI has. And so now the question is what happens. Well obviously Anthropic in Q2 is profitable, their net income profitable, excluding stock-based compensation. Um, and I think by Q3 they may even be profitable including stock-based compensation. That's like how profitable they're getting. And their margins on an Opus token, at least Opus 48 token, is like north of 80% for the API price. They've got a lot of deals where their total corporate gross margins get clawed down a little bit because of like how they do Bedrock deals and Vertex deals and things like that. But ultimately their per token margin is so high. Well, then if you don't have the cap, they have the capability to pay ultimately every GPU they buy at above market rate. You know, they also bought GPUs at above market rate from SpaceX, which is below the rate of Google, but that's because they signed earlier. Um, you know, it's something that other companies, maybe a venture-backed company or a company that's not really got positive margins, can't necessarily do, right? What is the cost benefit ratio? Like every GPU I rent because I'm out of compute capacity, I can immediately turn around and sell tokens on it, or every TPU or every Tranium I can immediately sell tokens on it at a positive margin. And if I'm running 75% gross margin and I double the cost of the compute, it's fine, I'm still running 50% gross margin. And spinning up more compute nodes is not really necessarily a human requiring task for them if they're renting them. So ultimately it's like, well my NOI still goes up, right? And so I'm going to rent GPUs at whatever price, at some level whatever price I want to pay, I can pay.

Host

我几乎有一个相反的问题,比如在某个时候,这种算力建设会不会在夜间突然出问题?今天早些时候,我想有一条推文说 Cuso 公开表示,他们的一位客户要求暂停一个数据中心的建设。似乎生态系统中的每个人都如此杠杆化,觉得我们必须建设,必须去建设,必须建设。高杠杆高增长对我来说,作为投资者,让我非常非常紧张。

I have almost the reverse question of like at some point does this compute build out go bump at night? Earlier today I think there was a tweet like Cuso publicly said one of their customers had asked to halt construction on one of their data center buildouts. Like it seems like everybody in the ecosystem is so levered right now to like we got to build, we got to go build, we got to build. High leverage high growth to me is like makes me very very nervous as an investor. Like

Dylan

等等,等等。高杠杆高增长意味着少量股权有巨大上行空间。你不是债务投资者。

Wait hold on. High leverage high growth means small amount of equity has huge upside. You're not a debt investor.

Host

你是信贷,你是股权投资者,对吧?

You're a credit, you're an equity investor, right?

Dylan

来吧。

Let's go.

Host

嗯,

Um,

Dylan

听着,你得去私募股权学校。只做杠杆收购。

Look, you got to go to the school of private equity. Levered buyouts only.

Host

我其实来自私募股权学校。

I actually come from the school of private equity.

Dylan

哦,太棒了。

Oh, awesome.

Host

她忘了学校。做 VC 太久了。

She forgot the school. It's been a VC for too long.

Dylan

是啊。

Yeah.

Host

不,我只做收入倍数。不,但你看到任何迹象吗?你担心吗?

No, I just do revenue multiples. No, but are you, do you see any signs of that? Are you worried about that?

Dylan

我明白你的意思。对。这又回到了模型的问题,对吧?显然,如果模型扩展的总经济价值工作,也就是我们做的暗 GDP 报告和你之前提到的,如果这些模型能做的工作没有比算力增长更快,那么趋势就会逆转,对吧?在过去六个月里,趋势非常倾向于模型能做更多工作,或者说它们能完成的工作的 TAM 增长快于算力增长。所以价格在上涨。模型进步突然停止是非常可能的。你问 Anthropic 或 OpenAI 的任何人,也许他们被洗脑了,但你基本上问他们所有人,他们都会说,‘不,不,不,不。模型进步仍在上升。’嗯,所以,你知道,最终,当前的方法可能会在某处停滞。我不确定会在哪里。似乎我们能看到模型改进的前景,快速的模型改进。事实上,模型改进的速度比六个月前或一年前更快,因为我不称之为递归自我改进,但基本上工程师们,模型正在帮助编写所有信息,并越来越快地推出下一个模型。所以你有了一个伪递归自我改进循环,模型变得越来越好,越来越快。嗯,但最终,资本是一个大问题,这就是为什么谷歌筹集了资本。你知道,他们拥有数量惊人的 SpaceX 股份,对吧?他们拥有公司大约 5% 的股份。

I see what you mean. Right. And that sort of goes back to the model point, right? Obviously if the models expanding the total economic valuable work, sort of the dark GDP report that we did and you mentioned earlier, if the work that these models can do does not expand faster than the compute capacity, then that tide turns, right? And over the last six months that tide has been very much levered in this direction of the models can do more work or is expanding their TAM of work they can do faster than the compute is increasing. And so prices go up. It's very possible that all of a sudden model progress stops. You talk to anyone at Anthropic or OpenAI, maybe they're drinking the Kool-Aid, but you talk to basically all of them, they're like, 'No, no, no, no. Model progress still go up.' Um, and so, you know, ultimately, current methods could stall somewhere. I'm not sure where that would be. It seems like we have line of sight to model improvement, rapid model improvement. And in fact, models are improving faster than they were six months ago or a year ago because there's I wouldn't call it recursive self-improvement, but basically the engineers, the models are helping write all the info and launch the next model sooner and sooner. So you've got this like pseudo recursive self-improvement loop going and so the models are getting better and better and better faster. Um, and so but ultimately, capital is a big problem which is why Google raised capital. You know, they've got an ungodly amount of SpaceX, right? They own like 5% of the company.

Host

我认为多一点,但是的。

I think a little more, but yeah.

Dylan

是的,也许。我认为他们一度拥有大约 10%。拉里·佩奇以 100 亿美元的估值投资了 10 亿美元,获得了公司 10% 的股份,后来被稀释了,等等。但这是有史以来最伟大的投资之一。干得好,拉里。这家伙。所以,他们知道自己在银行里有大约 1000 亿美元,可以在锁定期后九个月左右出售。他们还有所有的毛利润,但他们仍然建模,觉得需要筹集资本,于是进行了发行,这太疯狂了。这告诉你他们认为需要花多少钱。但资本真的很,你知道,Meta 确实宣布他们将进行融资。股价暴跌。人们不喜欢,但你知道,所有这些公司都将筹集资本,无论是债务还是股权。在某个时候,资金水龙头将不得不放缓。

Yeah, maybe. I think at one point they had like 10%. Larry Page invested a billion dollars at a $10 billion valuation, got 10% of the company, it got diluted, like all this. But that was one of the greatest investments of all time. Good job, Larry. The guy. So, they know they have like a hundred billion dollars in the bank that they can sell in, you know, nine months or whatever from the lockup. And they have all the gross profit they do, and yet they still modeled that and they were like we need to raise capital and so they did an offering and it's like that's insane. So that tells you how much they think they need to spend. But capital is like really, you know, Meta did announce that they're going to do a raise. Stock tanked. People don't like it, but you know, that's all these companies are going to raise capital, whether it be debt or equity. At some point, money spigots will have to slow down.

数据中心质量与定价差异 Data center quality and pricing disparity

Host

但现在,亚马逊每增加一块 GPU,他们就能获得更高的收入,或者每个 TPU 或 Trainium,无论谁增加,都能获得毛利润。我稍微铺垫一下,然后问你一个问题。当我们讨论这个时,我脑子里想的是关于 Crusoe 案例的一个替代假设。我用石油来类比:在石油领域,沙特阿拉伯每桶石油的生产成本远低于许多其他国家。还有石油的纯度;沙特的石油通常杂质很少,这使得精炼更容易。我的问题是:当你看到每千兆瓦的投入,比如今天有 20 千兆瓦上线,这些千兆瓦之间有多大的同质性?谷歌的千兆瓦是否比大多数新云服务商的价值高两倍,因为他们有光交换机,他们做这个很久了,而且他们知道如何进行功率平滑?这可能是一个替代假设:擅长建设数据中心的人应该全力以赴,因为需求巨大,而且他们做得更好,但也许我们开始看到那些不太擅长的人受到一些打击的早期迹象。我不知道实际情况如何;我只是好奇你怎么看。

But right now, every GPU that Amazon adds, they're making higher revenue or every TPU or Trainium, whoever adds is making gross profit. I do a little bit of a setup to turn into a question for you. As we talk about this, the thing going through my head is an alternative hypothesis for the Crusoe example. I'll use an oil analogy: in oil, Saudi Arabia has way lower cost per barrel to produce oil than many other countries. There's also the purity of the oil; Saudi generally has very low contaminants, which makes refining easier. The question for me is: when you look at every gigawatt being put in the ground, say 20 gigawatts coming online today, how much homogeneity do you see in those gigawatts? Are Google's gigawatts two times more valuable than most neoclouds because they have optical switches, they've been doing it for a long time, and they know how to do power smoothing? This could be the alternative hypothesis: people who are good at building data centers should just do it to the max because there's so much demand and they're so much better, but maybe we're starting to see early signs of those who are not as good getting hit a little. I don't know the reality; I'm just curious how you think about this.

Dylan

到目前为止,有这方面的指标。Trainium 以每千兆瓦低于 100 亿美元的租金价格卖给 Anthropic 和 OpenAI。GPU,至少在最近六个月的疯狂之前,通常每千兆瓦在 120 到 130 亿美元左右。所以租金价格,这是从新云服务商与亚马逊的比较来看,现在亚马逊卖 GPU 时,也大概是 130 亿美元左右。

So far, there are metrics for this. Trainium sells at sub-10 billion per gigawatt rental rate to Anthropic and to OpenAI. GPUs, at least before the craziness of the last six months, usually went around 12 to 13 billion per gigawatt. So the rental rate, and this is from a neocloud versus Amazon even, and now when Amazon sells GPUs, they'd also be 13 or so.

Host

据我所知,亚马逊对此有一点补贴,所以差距甚至更大。

And my understanding is that Amazon subsidized that a little bit, so the disparity was even more.

Dylan

低于 100 亿。低于 100 亿,但有一些奇怪的情况。

It's less than 10. It's less than 10, but there's some weird basically.

Host

而且,据我所知,Anthropic 在让 Trainium 变得有用方面发挥了很大作用,比如编写所有库等等。

And look, my understanding is that Anthropic played a big role in making Trainium useful in terms of writing all the libraries, etc.

Dylan

我听到的一切都说 Trainium 是非常好的硬件,而且越来越好,显然 Anthropic 现在大量使用它,所以希望我们会看到价格上涨。他们做的交易实际上有一个保底机制:如果表现不好,价格会更低,甚至可能取消;如果表现很好,价格会更高。但实际结果是,Trainium 的价格低于 100 亿美元每千兆瓦,而 GPU——我是说 SpaceX 的交易又是每千兆瓦 250 亿美元之类的疯狂数字,或者与谷歌的每年每兆瓦 2500 万美元的租金。这是一个巨大的差异。显然,如果亚马逊今天卖 Trainium,它可能会比 100 亿美元更贵,因为算力短缺,但你已经从数据中心看到了这一点。通常,如果你做托管,数据中心的租金——不是里面的算力,只是电力——通常按每千瓦每月多少美元来定价。以前是每千瓦每月 60 美元,现在你看到交易价格在 120 到 160 之间。但不同质量的数据中心:我见过数据中心高达 200 美元,当客户信用评级不太好而数据中心质量很好时。我也见过低至 100 美元的,或者在印度低至 80 美元,因为电网不可靠,互联网连接不好,而且是一个相当一般的数据中心,但至少是个数据中心。所以你已经看到了巨大的差异。在数据中心建设方面,通常的陷阱是:他们就是失败。很多人失败——四个人说,‘是的,我买了一些涡轮机。我付了定金。我要建一个数据中心。’然后他们一拖再拖,最后失败。所以你必须概率加权、时间加权、考虑时间滞后,区分好的团队和差的团队。我们的数据中心模型就是这样做的。我们跟踪每个数据中心,并根据他们使用的设备等所有因素对每个数据中心进行这样的分析。你提到的关于谷歌的一点:在一个千兆瓦的数据中心,他们实际上会放 1.5 千兆瓦的硬件,因为他们从工作负载到电力都有深刻的理解,他们能够调配电力。而不是一个千兆瓦的算力通常以 60% 或 70% 的功耗利用率运行——不是硬件利用率,而是有人一直在租用——他们现在以这样的方式运行:那个 60% 到 70% 意味着它在一个千兆瓦上,你使用了全部千兆瓦。你会看到人们与谷歌和公用事业公司做交易,他们说,‘哦,我知道这个电网可以持续承受一个千兆瓦,但除了每年三天,你实际上可以做两个千兆瓦,所以给我两个千兆瓦,然后告诉我关掉。’所以他们就这样做。这类技巧需要卓越的工作负载管理、备用电源、所有这些,现场发电机来弄清楚如何可持续地保持 2 千兆瓦。当人们这样做时,他们能够收取更多费用。无论是实际上我卖两个千兆瓦尽管只有一个千兆瓦,因为那三天我可以通过电池、天然气等处理,还是我找到了在现场发电的方法。现在我有一个别人没有的千兆瓦,所以我能够快速完成。这不一定是更高的交易价格;而是我卖了更多的千兆瓦。有时有一些杠杆,你卖更多的千兆瓦,每个千兆瓦以不同的价格出售。我认为更多是在数据中心和能源层面;更多的是关于拥有它与否以及是否延迟。它更二元化。但在算力方面,我确实认为有更多有趣的工作。一个千兆瓦给 Anthropic 客观上比给 OpenAI 产生更多收入。

Everything I hear is that Trainium is really freaking good hardware and it's getting way better, and obviously Anthropic is now using it a lot, so hopefully we would see that price go up. The deal they did actually had a floor mechanism: if it didn't do well, it would be cheaper, to the point where it's cancellable; if it did really well, the price is kind of higher. But effectively, less than 10 is where Trainium shakes out, whereas GPUs — I mean the SpaceX deal again was like 25 or something crazy billion dollars per gigawatt, or $25 million per megawatt per year rental rate with Google. That's a crazy divergence. Obviously, if Amazon was selling Trainium today, it would probably be more expensive than 10 because of the compute shortage, but you do see this already in the sense of data centers. Often, a rental price of a data center if you're doing colocation — not compute in there, but just power — you price it generally on dollars per kilowatt per month. They used to be $60 per kilowatt per month, and now you see things transacting at anywhere from 120 to 160. But different quality data centers: I've seen data centers go as high as 200 when the customer is not such a great credit rating and the data center is a pretty good one. And I've seen stuff go as low as 100 still, or in India go as low as 80 because the grid's not reliable, the internet connection's not great, and it's a pretty mid data center but at least it's a data center. So you see this huge discrepancy already. In the case of data center construction, usually the pitfalls: they just fail. There are a lot of people who fail — four guys who say, 'Yeah, I bought some turbines. I put the money down for them. I'm gonna build a data center.' And then they get delayed, delayed, delayed, and fail. So you have to probability-weight, time-weight, time lag, the teams that suck versus don't. Our data center model does that. We kind of track every data center and try to do this for every single one based on equipment they're using and all these things. One of the things you mentioned about Google: in a gigawatt data center, they'll actually put like 1.5 gigawatts of hardware because they have such understanding all the way from workload to power, they're able to slosh the power around. Instead of a gigawatt of compute which typically runs at 60 or 70% utilization in terms of power consumption — not utilization of the hardware, someone's always renting it — they're now running it at you know, that 60 to 70% means it's at a gigawatt and you're using the full gigawatt. You see people doing deals with Google with utilities where they say, 'Oh, I know this grid can sustainably take a gigawatt, but except for three days of the year, you can actually do two gigawatts, so give me two gigawatts and then just tell me to turn off.' And so they'll do that. These sorts of tricks require supreme management of workload, backup power, all these things, generators on site to figure out how to actually keep it at 2 gigawatts sustainably. When people do this, they're able to charge more. Whether it be I'm actually selling two gigawatts despite only having one gigawatt because those three days I'm able to deal with via battery, gas, etc., or I figured out how to build power on site. Now I have a gigawatt where no one else does, so I'm able to do it quickly. It's not necessarily transacting for a higher price; it's that I'm selling more gigawatts. And sometimes there are levers where you're selling more gigawatts where each gigawatt is selling at a different price. I think it's more on the data center and energy layer; it's more about just having it versus not and that being delayed or not. It's more binary. But on the compute side, I do think there's a lot more interesting work there. A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI.

AI实验室的供需约束 Demand and supply constraints at AI labs

Host

而且看起来,他们两家现在每千兆瓦的算力都能卖光,考虑到 OpenAI 和 Anthropic 的速率限制、token 上限等问题,尤其是 Codex 5.5 发布后,效果更好了。同样,如果你给 SpaceX 一千兆瓦,你知道,他们会……

And it seems that both of them could sell every gigawatt that they have right now, given rate limit problems and token max limit and all these sorts of things at OpenAI and Anthropic, especially since Codex 5.5 came out, it's much better. And then likewise, if you gave a gigawatt to SpaceX, you know, they turn...

Dylan

我猜,我怀疑他们可能比大多数人更善于利用硬件。就像我认为人们低估了他们从 Starlink 获得的网络经验,以及从 Tesla 获得的电源管理经验。

My guess, my suspicion is that they probably make better use of the hardware than most people. Just like I think people underestimate how much networking experience they have from Starlink in particular, and also how much power management experience they have from Tesla.

Host

是的。像 Brett Mayo 这样的人非常厉害。

Yeah. People like Brett Mayo are incredible.

Dylan

相当不错。

Pretty good.

Host

是的。所以对我来说,这可能是很多人分析中遗漏的一点。我其实不知道答案,但我觉得可能被忽略了。

Yeah. And so I think that for me, that's actually probably the thing that might be missing from the analysis a lot of people are doing. I don't actually know the answer, but I think that might be missing.

Neocloud与超大规模云优势 Neocloud vs hyperscaler advantages

Dylan

我认为还有一个事实:当 CoreWeave 建设一千兆瓦时,尽管他们的 GPU 算力在性能上客观优于 Amazon、Google 或 Microsoft。我们测试过性能和可靠性。问题是 Google 在建成前六个月就销售,他们需要拿着签好的合同去获取债务融资,然后才能支付已经发出的采购订单。而 SpaceX 则说:‘不,不,不,这个已经在运行了,现在就买。’有资产负债表和没有资产负债表之间差距很大,这也使得每兆瓦收入高得多。

I think it's also the fact that when CoreWeave builds a gigawatt, even though their GPU compute is objectively better than Amazon or Google or Microsoft's in terms of performance. We've tested the performance and reliability. The problem is Google sells it six months before they have it up, and they need to turn around and take that paper that they signed to get debt with that credit backing, and then turn around so they can actually pay for the PO that they've already issued for the order. Whereas SpaceX was like, 'No, no, no, this is running now, buy it right.' And it's a big discrepancy when you have a balance sheet to do that versus not, and that also helps your revenue per megawatt be much higher.

Host

为什么 Neocloud 的机会还存在?因为如果你五年前问我,我会说超大规模云厂商会主导这一切。你刚才提到 CoreWeave 性能比超大规模云厂商更好。那么,为什么这个机会还存在?也许从宏观层面和执行层面来看?

Why does the Neocloud opportunity even exist? Because if you had asked me five years ago, I would have said the hyperscalers are going to own this. And you mentioned just now CoreWeave has better performance than the hyperscalers. Like, why does this opportunity exist, maybe at the macro level and then in the execution level?

Dylan

是的。2023 年我写了一篇报告,让 Amazon 很恨我。标题是《Amazon 云危机》。我谈到 Amazon 曾经是最好的云,因为他们有 Nitro 网卡,提供租户隔离,所有虚拟机管理程序都在网卡上运行,这样你可以出售所有核心,他们还有自制的 SSD,购买原始 NAND 以降低成本,以及自研的 Graviton CPU,降低了每核心成本。他们拥有所有这些优势,可以销售更多核心,提供更好的安全性和网络,但这都是针对传统 CPU 和传统云世界的存储。但在 AI 云中,这些东西很多都损害了性能。对吧?这些 Nitro 网卡对性能不利,仍然更差。虽然他们通过几次迭代改进了很多,但性能仍然较差。很多安全特性并不重要,因为不是时间片分割用户或把插槽分给多个用户,对吧?没有人会租用 8 GPU 服务器中的单个 GPU,也没有人租用 72 GPU 机架中的单个 GPU。他们租用整个机架,实际上租用多个机架。而且没有‘我租六小时就归还’的情况。每个人都签长期合同。所以 GPU 租赁市场的机制意味着超大规模云厂商的很多专长都失效了。而且他们拥有的很多专长实际上是有害的,对吧?Google 和 Amazon 的网络性能:他们有定制网络,对传统 CPU 和他们正在做的事情更好,但实际上对 AI 有效。在其他情况下,比如 Microsoft 通过自建数据中心节省成本,但他们的数据中心团队实际上并不那么出色。所以当需要运行时,当建设可预测时,没问题。但当需要实际将年度预测翻倍时,他们就摔了个跟头,不得不去获取大量 Neocloud 容量。我认为性能,我认为上市时间是另一个因素。对吧?这些大型组织,没有人因为更快地建设数据中心而致富。但看看 Crusoe,比如 Chase 和团队中的其他人。我本来想提一些名字,还是算了。如果这些人更快地交付算力,他们就会致富。他们是高杠杆的股权所有者。

Yeah. So in 2023 I wrote a report that had Amazon really hate me. It was called 'Amazon Cloud Crisis'. So I talked about how Amazon was the best cloud because they had their Nitro NICs which offered tenant isolation, all the hypervisor ran on the NIC, and then you could sell all the cores, and they had custom SSDs that they made, and they'd buy the raw NAND and have lower cost because they'd buy the raw NAND and build their own SSDs, and they had their custom Graviton CPUs, and that drove down cost per core. So they had all these things that enabled them to sell more cores, have better security, good networking, but this was all for the traditional CPU, better storage for the traditional cloud world. But in the AI cloud, a lot of this stuff hurt performance. Right? These Nitro NICs are bad for performance, still are worse performance. Although they've caught up a lot because they've had a couple iterations to improve them, but they're still worse for performance. A lot of the security stuff doesn't matter because it's not like I'm time-slicing users or slicing a socket into many users, right? It's like no one rents a single GPU in an 8-GPU server. No one rents a single GPU in a 72-GPU rack. They rent the whole rack, and in fact, they rent many of the racks. And then there's no like, 'Oh, I rent for six hours and I give it back.' Everyone has these long-term contracts. So the mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away. And a lot of the expertise that they did have were actually detrimental, right? Network performance for Google and Amazon: they had custom networks that were better for traditional CPU and for the stuff that they were doing, but actually worked for AI. And then in other cases, it's like, well, Microsoft would save money by building their own data centers, but their data center teams were actually not that great. So when it came time to run, when it was predictable building, it was fine. When it came time to actually double your forecast for the year, it's like they fell on their face, and they had to go get a bunch of Neocloud capacity. I think performance, I think time to market is another one. Right? These massive organizations, no one's getting rich from building this data center faster. But you look at Crusoe, for example, Chase and all the other people on the team. I was going to name some people, I'd rather not. All these people are getting rich if they deliver this compute faster. They're hyperlevered equity owners.

Host

嘿,你看,他们也都来自比特币。你知道,你不该这么说。

Hey, look, they're also all coming from Bitcoin. And you know, you're not supposed to say that.

Dylan

我是说,很多数据中心,比如他们的主要数据中心负责人来自 Microsoft。

I mean, a lot of the data center, like their main data center guy came from Microsoft.

Host

我不知道。我只是开玩笑。但你在一个波动很大的市场里能学到很多。

I don't know. I'm just teasing. But it's like you learn a lot when you're in a very high fluctuation market.

黄仁勋的多极世界策略 Jensen's strategy for a multipolar world

Host

你认为 Jensen 在每步棋中打的是什么算盘?

What do you think was Jensen playing for each chess?

Dylan

Jensen 绝对讨厌所有超大规模云厂商掌握所有权力的世界。他之所以在随机 AI 实验室上砸钱,我甚至不知道是否合理,但他就是砸钱并吹捧它们,然后对全世界每个人说你应该投资这家公司,因为他想创造一个多极世界。这就是为什么他喜欢中国实验室,因为他想创造一个多极世界。一个只有 OpenAI、Anthropic 和 Google 模型的世界,是他完蛋的世界。是的。

Jensen absolutely hates a world where all the hyperscalers have all the power. There's a reason he's like blowing money on random AI labs that I don't even know if it makes sense to, but he's blowing money and pumping them up and going to everyone around the world and saying you should invest in this company because he wants to create a multipolar world. That's why he loves Chinese labs, because he wants to create a multipolar world. A world where OpenAI, Anthropic, and Google models are the only models is one in which he's screwed. Yep.

Host

对。一个只有超大规模云厂商建设算力的世界,是他完蛋的世界。

Right. A world in which the hyperscalers are the only ones building compute is one he's screwed in.

Dylan

所以,当然他需要把分配枪口指向 Neocloud,帮助支持他们的集群,做任何事,因为今天卖给 Crusoe、CoreWeave、Google 和 Amazon 的 GPU 对他来说价格都一样,但五年后,Crusoe 和 CoreWeave 的存在意味着 Google TPU 会更弱,Amazon Trainium 会更弱,更多推理由非封闭模型实验室完成,对他更有利。

And so, of course he needs to point the allocation gun at Neoclouds, help backstop their clusters, do anything and everything because while today a GPU sold to Crusoe and a GPU sold to CoreWeave and a GPU sold to Google and Amazon are all the same price for him, five years from now Crusoe and CoreWeave existing means Google TPU will be weaker and means Amazon Trainium will be weaker, and more inference being done with non-closed model labs is better for him.

Neocloud生态与新实验室 Neocloud ecosystem and neolabs

Host

所以我认为,新云生态系统就像狂野西部。这些新实验室,很多都有英伟达的投资。有些会失败,很多会失败,但有些会脱颖而出,成为真正优秀的团队。无论是 Crusoe,一群加密货币玩家开始建数据中心、做废气发电,还是 CoreWeave,最初是一群纽约对冲基金和加密货币玩家。但有很多人差不多同时起步,最后却失败了。

So I think the neocloud ecosystem is the wild west. These neolabs, a lot of them have investments from Nvidia. Some will fail, many will fail, but some will emerge as really great teams. Whether it be Crusoe, a bunch of crypto guys who started building data centers and doing flared gas stuff, or CoreWeave, initially a bunch of New York hedge fund and crypto guys. But there were a lot of people who started around the same time and just failed.

Dylan

我得说,这两个团队都非常出色。他们值得很多赞誉。

I gotta say both those teams are phenomenal. They deserve a lot of credit.

Host

是的,我的意思是,就像你往水里扔一堆饵料,最好的鱼会想办法存活下来。新云也一样,希望新实验室也是如此。我们会看看有没有新实验室真的冒出来。Thinking Machines 有几亿美元的 ARR,对吧?这相当令人印象深刻,尽管媒体说他们流失了所有人才。但 Tinker 的产品不到六个月,就做到了几亿美元 ARR。我们希望其他新实验室也能如此。他想要一个多极世界。

Yeah, my point is like you throw a bunch of bait into the water and the best fish will figure out and survive. Same with the neoclouds and hopefully the neolabs as well. We'll see if any of the neolabs really bubble up. Thinking Machines has a few hundred million dollars of ARR, right? That's pretty impressive even though in the media it's like they've lost all this talent. But Tinker is doing a few hundred million of ARR for a product less than six months old. We hope the same happens to other neolabs. He wants a multipolar world.

Dylan

真心祝贺你的成功。谢谢。

Truly, congratulations on the success. Thank you.

Host

最后我想说,我亲眼见证了一些。公众大概能从你的言谈中听出你有多努力,但很明显,你过去十多年一直在拼命工作,这才有了最近几年天时地利人和的成果。你取得的成就令人难以置信,我知道这还只是开始。

Just the last thing I'll say is I've seen a little bit of this. I think the public can probably tell from listening to you how hard you work, but it's clear you've just been working your ass off for more than a decade and it led to the last few years of being in the right place, right time. It's unbelievable what you've accomplished and I know it's just the beginning.

Dylan

非常感谢。

Thank you so much.

Host

谢谢你接受采访。

Thank you for doing this.

Dylan

太棒了。

Awesome.

互动版:逐字朗读 + 针对本期提问 →