Which GPU Clouds Are Actually Good? ClusterMAX 3.0
打开互动全文版(中英对照 + 朗读 + 问答)→SemiAnalysis 的 Dylan 解读 ClusterMAX 3.0,这套 NeoCloud 评级从性能、安全、可靠性和定价等维度评测 GPU 云服务商。
SemiAnalysis's Dylan explains ClusterMAX 3.0, the NeoCloud ratings that benchmark GPU cloud providers on performance, security, reliability, and pricing.
嗨,Jordan,恭喜你们发布 ClusterMAX。先简单介绍一下你自己的背景,以及你打造 ClusterMAX 的历程吧?
Hi Jordan, congrats on launching ClusterMAX. What's your brief history as a personal introduction, and also your journey into building ClusterMAX?
好的。我去年夏天、去年六月加入 SemiAnalysis,一直在技术团队工作。过去一年我们成长了很多。我的个人背景是,我在惠普企业(HPE)工作了 10 年。在 ChatGPT 发布之前,我就在设计这些 Neocloud 采购的很多硬件系统,后来 22、23 年那波大热潮,很多厂商涌入市场。我当时是 SemiAnalysis 的重度读者,每篇都读,私下在群聊里认识了不少人,看到 Dan 和团队发布 ClusterMAX 1 时,里面有很多观点和我一致。而在 SemiAnalysis 找工作最好的办法,就是找到他们在做的东西然后批评它,接着 Dylan 就会雇你来修好它——我的情况就是这样。
Sure, yeah. I joined SemiAnalysis last summer, last June, and have been working on the technical staff. We've grown a lot over the past year. My personal background is that I spent 10 years working at Hewlett Packard Enterprise. I was designing a lot of the hardware systems that these Neoclouds were buying back before ChatGPT launched, and then with the big craze in '22, '23 when a lot of them came into the market. I was a big SemiAnalysis reader, read every word, knew a bunch of the guys in group chats on the side, and saw a lot of my opinions represented in ClusterMAX 1 when Dan and the team put it out. And the best way to get a job at SemiAnalysis is to find something that SemiAnalysis is doing and criticize it, and then Dylan hires you to fix it, which was what happened in my case.
那你得跟我们说说,对 ClusterMAX 1 的批评是什么?
You got to tell us what were the criticisms of ClusterMAX 1?
嗯,有两篇是我自己一个人写的,所以我可以给你讲讲这些批评。
Well, I did two all by myself, so I can give you the critiques.
哦,那里可没什么批评。
Oh, no criticisms there.
呃,
Uh,
从一代到二代有什么变化?然后我们来看看有哪些新东西。
What changed from one to two? And then, you know, let's click into what's new.
各位,屏幕上你们看不到,它本身就很完美。我不知道你们会批评什么。嗯,好吧。对,ClusterMAX 1。当时就是漏掉了很多厂商,总共只有 24 家,而且很多大厂本该被测试却没测。小团队没时间。另外当时很侧重 Slurm,而我觉得 Kubernetes 非常重要,也需要测试。在标准等方面,我们没改太多,但回到 ClusterMAX 1,排行榜页面看起来差别很大,名字少了很多。发布市场全景图时,我记得当时总共约 120 家厂商,中间有 111 家新兴 Neocloud。这次我们没有发布完整的市场全景图,因为把它放到了付费的 Neocloud 模型里,我们在那里追踪所有这些厂商。很多是上市公司,很多即将 IPO,所以我们现在有了整个行业的财务模型。但 3.0 版本,我们现在追踪的厂商总数达到了 323 家。今天早上我收件箱里还有人联系我,说他们有 3000 块 GPU 想让我们测试,我们得……总之,这是一次大演进。我觉得测试方面最大的变化是,ClusterMAX 1 时所有测试都是手动的。那时还没有 Claude Code,也没有 Codex 认真起来。所以我们很多测试都是手动跑基准脚本。而现在,过去一年半我们开发了一整套仓库,包括行业标准基准、自定义基准、可靠性和老化测试的东西。现在非常自动化了,我们可以用编码智能体启动这些测试,让它们自主运行并修复一堆问题,不需要我们参与。而且,覆盖的广度和深度,我觉得都随着时间提升了。
You can't see it on screen, guys. It's perfect as is. I don't know what you would criticize. Um, okay. Yeah, ClusterMAX 1. I mean there was just a lot of providers missed, like 24 total providers, and a lot of them big ones, you know, they should have been testing and weren't. Small team didn't have the time. And then there's a big focus on Slurm; to me Kubernetes was really important, need to test that as well. In terms of criteria and stuff, we haven't changed a lot, but yeah, all the way back to ClusterMAX 1, you know, the rankings page looks a lot different, there's just a lot less names. And releasing the market view also, I believe at the time it was around 120 total providers, 111 emerging Neoclouds there in the middle. And we didn't release the full market view this time because we paywalled that behind our Neocloud model, where we track all of these providers. A lot of them are public, a lot of them are going IPO soon, so we've got this whole financial model for the industry now. But yeah, 3.0, I know we're up at 323 total providers that we track. And I have some people this morning in my inbox reaching out saying that they've got some 3,000 GPUs and they want us to test, and we got to, you know, anyway, it's been a big evolution. I think the big thing in terms of testing is that with ClusterMAX 1, all the testing was hands-on. This is pre-Claude Code, pre-Codex getting serious. And so we did a lot of the testing hands-on by just running benchmark scripts manually. And now we've been able to develop this whole repo over the past year and a half of industry standard benchmarks, custom benchmarks, stuff for reliability and burn-in. It's like very automated now where we can launch this stuff with the coding agents and have them run and fix a bunch of issues autonomously without us involved. And yeah, I mean just the breadth of coverage as well as the depth of coverage I would say has improved over time.
我想,先介绍一下吧。ClusterMAX 是什么?NeoCloud 评级和排名是什么?你们主要测试的是哪些方面?
I mean I guess introduce it. What is ClusterMAX? What is NeoCloud ratings and rankings? What are like the main things you guys are actually testing for here?
好的。网站上描述了 10 项标准。基本上,我们把 Neocloud 定义为任何出租算力访问权、出售算力访问权的厂商。所以 Neocloud 意味着在 ChatGPT 之前就存在超大规模厂商和传统云。比如 AWS、Google、Oracle、Azure,还有一些老牌云,像 Rackspace 或 Digital Ocean,甚至 Cloudflare AI。任何拥有大型数据中心并出租访问权的,那就算是云,而 Neocloud 是后来出现的新玩家。我们说所有人都在变成 Neocloud,因为 Neocloud 的不同之处在于它们专门聚焦 AI。它们部署 GPU 和其他加速器用于 AI 工作负载,包括训练、推理,越来越多还有强化学习。我们试图测试的就是:谁能给你最好的性能、安全性、可靠性,以及易用性、定价、可用性。我们尝试做全面评估,而不是单个基准。ChatGPT 刚出来时,很多人联系 SemiAnalysis 说:“嘿,我在五家厂商之间做选择,该怎么考虑?该怎么做?”他们反复用同样的问题烦 Dylan,因为有人说:“我刚融了 2000 万美元,马上要把其中 1000 万、500 万给一家厂商。这是我整个种子轮。这对我来说是极其关键的决定,我需要所有人的反馈。”所以与其让 Dylan 和团队每次单独定制回答,我想,不如我们写篇文章给出粗略的第一想法,这是我们的快速分级排名、领奖台,然后你可以就具体需求联系我们做更详细的东西,比如咨询、尽调之类的,我们可以进一步深入。
Yeah. So there's 10 criteria described on the website. Basically a Neocloud we define as anybody who rents access to compute, sells access to compute. So Neocloud implying that there were hyperscalers and legacy clouds that existed before ChatGPT. This is AWS, Google, Oracle, Azure, also some of the old school clouds like Rackspace or Digital Ocean or even like a Cloudflare AI. Anybody who's running a big data center footprint and renting access out, that would be a cloud, and then the Neoclouds are the new ones that have come along. We say everybody's becoming a Neocloud because what makes Neoclouds different is that they're specifically focused on AI. So they deploy GPUs and other accelerators for AI workloads. This is training, inference, increasingly RL. And that's what we try and test on: who's going to give you the best performance, security, reliability, and just like ease of use, pricing, availability. We try to go through comprehensive assessments instead of individual benchmarks. Back when ChatGPT came out, a bunch of people were reaching out to SemiAnalysis saying like, "Hey, I'm trying to decide between five providers, like how should I think about this? How should we go about it?" They were bothering Dylan with the same questions over and over because they're like, "Hey, I just raised $20 million. I'm about to go give 10, 5 million of it to a provider. You know, it's my whole seed round. Like, this is a really critical decision for me. I need to get everybody's feedback." And so instead of Dylan and team answering that question individually custom every single time, I was like, why don't we write an article to give our coarse-grained first thoughts, here's our quick tier list rankings, podium, and then you can engage us for more detailed stuff, you know, consulting, DD, whatever it is, for your specific requirements and we can dig in further from there.
这里面信息很多。我想到的是那 10 项标准里的一个。我想到的是你们引用的那句话,Ilya Sutskever 说它们全都有糟糕的网络安全。Neocloud 的网络安全现状如何?
There's a lot there. The one that comes to mind is the criteria of the 10 criteria. The one that comes to mind is the one that you guys quoted, which is Ilya Sutskever saying all of them have terrible cyber security. What is the state of cyber security in Neoclouds?
我是说,你说得对,很糟糕。对。这取决于我们说的是谁。你看,呃,我们……对。
I mean you said it, terrible. Yeah. It depends who we're talking about. Look, uh, we... Yeah.
另外我也跟 Dylan 打个招呼,他刚加入我们。Dylan,
And also I'll just say hi to Dylan, who has just joined us. Dylan,
很抱歉我迟到了。
So sorry for being late.
Dylan,Neocloud 的网络安全现状现在如何?
Dylan, what's the state of Neocloud cyber security right now?
嗯,我是说,我们,你知道,我觉得安全性真的很差。甚至从第一代 ClusterMAX 开始,就有相当严重的安全问题。人们没有正确实施网络安全控制。所以我们能看到别人的任务、别人的数据、别人的存储等等。在第一代 ClusterMAX 中,我们向公司以及 Nvidia 上报了,当然把那家公司评为表现不佳,并强烈建议人们不要用他们。在第二代 ClusterMAX 中,我们看到了类似的情况。我觉得很多人变好了,因为我们指出了问题。但在第三代 ClusterMAX 中,我们也做了测试,我想 Jordan 和他的团队测试了,看到了相当重大的安全问题。所以我们发了一篇通讯,标题是“早告诉过你:大多数 Neocloud 的安全性很烂”。
So, I mean, we, you know, I think security is really bad. Even back from the first ClusterMAX, there was pretty significant security issues. People not properly implementing network security controls. And so we were able to see other people's jobs, see other people's data, see other people's storage, etc. In the first generation of ClusterMAX, we escalated to the firm as well as Nvidia, you know, and ranked that company underperform of course and strongly recommended people don't use them. And in the second generation of ClusterMAX, we saw similar stuff. I think a lot of people got better because we highlighted stuff. But in the third generation of ClusterMAX, we also tested, I think Jordan and his team tested this and saw pretty major security issues. So we posted a newsletter titled "Told You: Most Neoclouds Suck at Security."
几个小时后,Ilia 发推说大多数 Neocloud 的安全性都很糟糕。他们显然是一家非常注重安全的公司,有很多人来自网络安全领域。所以我想他们已经在多家云上做过自己的测试和验证,结果很清楚,对吧?
A handful of hours later, Ilia tweeted about most Neocloud security sucking. They're obviously a very security-conscious firm. They have a lot of folks who are from the former cyber security world. So I imagine they've done their own work on a number of clouds there to see and test, and it's pretty clear, right?
什么?这么糟吗?
What? It's so bad.
有一次,我们居然能直接看到某个国家的国家情报内容,对吧?那个 Neocloud 上跑着好几个国家安全机构和军方风格的情报系统,我当时就想,天哪,我们只是来测集群的,居然能看到这些东西。想象一下如果我们有恶意会怎样。我敢肯定那个集群上也有恶意的人,或者本来可能有恶意的人。所以谢天谢地,我想他们后来修好了。但我们不该是那个必须去发现谁的安全做得好、谁做得差的人。不过我觉得整个行业确实需要这项服务。所以这是我们评判标准里很大的一块。
In one case, we were literally able to see national intelligence of a certain country's stuff, right? They had multiple national security agencies and military-style intelligence stuff running in that Neocloud, and it's like, good god, we're just here to test the cluster and we can see this stuff. Imagine if we were malicious. I'm sure there were people who were malicious or could have been malicious on that cluster as well. So thankfully they fixed it, I think. But we shouldn't be the one that has to find whether or not someone is good or bad at security. But I think the industry generally needs that service. So that's a big area of our criteria.
对,我觉得在技术层面,我们的检查非常简单。你的软件版本是最新的吗?驱动是最新的吗?配置正确吗?这些都是相对简单的检查,但当人们跑着两三年前的软件版本,而网上又有记录详尽的 CVE 时,你根本不需要前沿模型就能做出 PC 漏洞利用。你只需要知道去哪里找。所以作为文章的一部分,我们还发布了一个免费的 CLI。大家可以安装 ClusterMAX,也就是 CMAX,然后用它检查自己集群上的软件版本,并识别出问题——终端输出里会给你一个链接,告诉你该升级什么。就是想帮大家保持软件更新。这就是 Neocloud 安全的问题所在。
Yeah, and I think on the technical side, our checks are very simple. Is your software version up to date? Are your drivers up to date? Is it configured correctly? These are relatively simple checks, but when people are running versions of software that are 2, 3 years old and they have well-documented CVEs online, it's like you don't need a frontier model to create a PC exploit of this. You just need to know the right place to look. And so as part of the article, we also put out a free CLI. People can install ClusterMAX, CMAX, and then use it to check the software versions on their clusters and identify — you get a link in the terminal output to show you where to upgrade stuff. Trying to help people keep the software up to date. This is the thing with Neocloud security.
另一件事是网络密钥,对吧?InfiniBand 和以太网有各种——至少在 InfiniBand 这边有 M keys、P keys 这些东西。它们有没有被正确配置和部署?不过大多数 Neocloud 糟糕的程度真是令人震惊。真的非常糟。
The other thing is the network keys, right? InfiniBand and Ethernet have various — at least on the InfiniBand side it's like M keys, P keys, all these things. Are those properly configured and deployed? It's shocking how bad most Neoclouds are though. Like it is really bad.
那你们——我想问问有没有什么有意思的,有没有大的跃升,从第一代、第二代到第三代有没有什么改进?有什么你想特别提一下的吗?
Well, you guys — I guess anything interesting, any major jumps, any improvements that they've made from one, two, three? Anything you want to highlight?
我们在 ClusterMAX 上的门槛是——第一代非常低。每一代我们都把门槛抬得越来越高。所以趋势就是大家被降级。当然也有少数人被升级了。Google 在跨越 ClusterMAX 各代的过程中做得非常好。我记得他们一开始是铜级,然后升到银级,现在是金级。或者只是银级,抱歉。哦不,是金级。对。所以 Google 一直在升级,而有不少云在跨代过程中被降级了,因为他们停滞了。他们改进了一点点,但没有大幅改进。而我们一直在努力抬高门槛,并告诉大家,嘿,这些就是你们应该做的所有事情,让自己变得更好。
Our bar on ClusterMAX is — the first generation was super low. And each generation we've made it harder and harder. Hence, the trend has been for people to get downgraded. There have been a few folks who have been upgraded, of course. Google's done a really good job at getting upgraded across the ClusterMAX generations. I think they started out as bronze and moved to silver and now gold. Or just silver, sorry. Oh no, gold. Yeah. So Google's consistently upgraded, while there have been a number of clouds who have been downgraded across the generations because they sort of stalled. They improved a little bit, but they didn't improve a lot. And we keep trying to raise the bar and teach everyone, hey, these are all the things you should do to be better.
对,而且这一次的门槛就是最新最强。那么,你在 2026 年部署 GB300 NVL72 了吗?你知道,这是一款极其热门的芯片。所有前沿实验室都想买,所有 Neocloud 也都想买。而典型的客户来找我们说:“我融到了 1 亿美元,我想花在算力上。”他们想要一个 GB300 NVL72 集群。所以一些原本在金级的供应商并不提供 GB300 NVL72。为什么不提供?因为你得应对一大堆不同的东西。有新的网络。NVL72 意味着机架内的 scale-up 域。你得转向 ARM CPU。它是 Blackwell GPU。你有 800G 网络。强制直接液冷。所以从物理数据中心到如何部署操作系统,再到如何管理网络中的驱动,几乎一切都不同。当人们说,哦,我们只是做了一个战略性的商业决定,不做 GB 系列,那么明年 Vera Rubin 出来的时候怎么办?你打算信任谁?所以趋势就是:如果人们出去部署了 1 万台 GB300,我们测试集群,它能正常工作,我们和他们的支持团队沟通,就一些问题给出反馈,他们一小时内就修好了,而不是两周,那他们的排名就会上升。其他人要么没有部署这些最新最强的 GPU,要么部署了但完全坏掉,修起来要很久,要么我们发现了安全问题,要么我们从客户那里听到巨大的可靠性问题——而这可能是这些顶级客户最看重的第一标准——那他们就会往下掉。所以大致就是这样的动态。
Yeah, and the bar for this time was like latest and greatest. So, are you deploying GB300 NVL72 in 2026? You know, this is an incredibly popular chip. It's what all the frontier labs want to buy. It's what all the Neoclouds want to buy. And the typical customer that comes to us and says, "I've raised $100 million. I want to spend it on compute." They want a GB300 NVL72 cluster. And so some of the providers that were in gold originally do not offer GB300 NVL72. Why not? Well, you've got to contend with a whole bunch of different things. There's a new network. NVL72 implies it's a scale-up domain in the rack. You've got to go to ARM CPUs. It's a Blackwell GPU. You have an 800 gig network. It's mandatory direct liquid cooling. So from the physical data center to how you deploy the operating system to how you manage the drivers in the network, just about everything's different. And when people say, oh, we just made a strategic business decision not to do GBs, well, what happens when Vera Rubin comes next year? Who are you going to trust? And so the trend is like, if people are going out and they're deploying 10,000 GB300s and we test the cluster and it works, and we engage with the support team and we give them feedback on some things we have issues with and they fix it in an hour, not two weeks, they're moving up the rankings. Other people that are not deploying these latest and greatest GPUs, or they're doing it and it's completely broken, it's taking them a long time to fix it, or we find security issues, or we hear from customers huge issues with reliability, which is probably the number one criteria that these top customers care about — they're moving down the list. So that's a bit of the dynamic there.
那么假设这些铜级、银级、荣誉提名的供应商大多都值得租用,那可用性、成本、算力的现状如何?比如说你是一个新实验室,刚融到几千万美元,想开始租算力。这个市场现在和一两三年前相比怎么样?是变好了还是变差了?
So assuming most of these, you know, bronze, silver, honorable mentions are good to rent from, what's the state of availability, cost, compute? So say you are a neolab, you just raised tens of millions of dollars, you want to start to rent. How does that market look now compared to one, two, three years ago? Things getting better, worse?
祝你好运。
Good luck.
对,差多了,老兄。
Yeah, way worse, man.
真的很难。现在获取算力是史上最难的时候。显然 Anthropic 和 OpenAI 一直在买越来越小的集群规模。但与此同时,所有推理提供商都越来越赚钱。以前 Base10 在推理上跑 10% 的毛利率,我记得大概是六个月前,或者九个月前,那还只是一回事。但当 Base10、Fireworks 以及许多其他公司,甚至像 Morph 这样的小公司,在开源模型上用 vLLM 或 SGLang,再加上一点点优化,就能跑到 60% 的毛利率,还能做得非常出色,那就是另一回事了。还有像 Incept 和 Radix Arc 这样的公司。所以任何能拿到 GPU 的人都能立刻赚钱。所以归根结底就是,好吧,我是一个想训练一个有趣、酷炫新模型的新实验室,但我同时在和那些现在就能赚钱的人竞争。而以前我主要是在和亏钱的、你知道,亏钱的开源、亏钱的 Anthropic、亏钱的实验室竞争。以前很少有人靠 GPU 赚钱。现在人人都能靠 GPU 赚钱。所以想拿到任何算力都非常非常困难。
It's really tough. It's the toughest it's ever been to get compute. Obviously Anthropic and OpenAI have been buying smaller and smaller cluster sizes. But also all of the inference providers are more and more profitable. It was one thing when, say, Base10 was running at 10% gross margins, which I think might have been like six months ago, right, on inference, or nine months ago. But it's another thing when Base10, Fireworks together, and many other companies, even small companies like Morph, are running at 60% gross margins on open models using vLLM or SGLang, using just a little bit of optimization beyond that, and able to do an awesome job. And then you know companies like Incept and Radix Arc as well. So anyone who can get access to a GPU can immediately make money. And so ultimately it's like, well, great, I'm a neolab who wants to train an interesting, cool new model. Well, I'm also competing against people who make money now. Whereas before I was mostly competing against money-losing, you know, open, money-losing Anthropic, money-losing labs. Very few people were making money off of GPUs. Now everyone can make money off of GPUs. And so it's just very, very difficult to get any sort of compute.
和大型集群相比,这些实验室现在降到什么规模了?
What level are the labs going down to compared to like major?
他们以前只从供应商那里租 8000 块 GPU 左右。这个数字一直在稳步下降。我们现在甚至看到实验室租用少至 1000 块 GPU 的集群。
They used to only go like 8k, you know, 8k GPUs from a provider. And it's steadily gone down. We've even seen as small as 1,000 GPUs get rented by the labs now.
所以 1000 块 GPU 更接近一家新实验室想要的规模,于是市场被挤压得很厉害。然后你看 Vast、Fireworks、Together 这些推理服务商,它们对四个节点就很满意,就像 Modal 那样,对吧?这些公司哪怕在云上只有四个节点也很满足。所以归根结底,它们的编排确实很难。Modal 谈了很多这方面的事,对吧?它们要跨 20 个不同的云做编排。
So 1,000 GPUs is closer to what a neolab wants, and so you've squeezed out the market a lot. And then when you look at the Vast 10s, the Fireworks, the Togethers and all these other inference providers, they're cool with four nodes, like they're modal, right? These guys are cool with even four nodes at a cloud. So ultimately, obviously their orchestration is tough. Modal's talked a lot about that, right? Their orchestration across like 20 different clouds.
是啊。
Yeah.
所以归根结底,如果 Modal 靠一家公司的四个节点就能赚得盆满钵满,那任何人想拿到算力都非常困难。
So ultimately it's very challenging for anyone to get any sort of compute if Modal can make money hand over fist off of even four nodes from a company.
是啊,情况很糟。我是说,连个人——你现在都没法按需租到单个节点了。整条链路都完蛋了。那实验室这边的情况呢?你怎么划分这些大玩家——OpenAI、Anthropic、Google?谁拥有最多算力,这个分布是怎样的?关于路线图,你觉得谁投入最多,有什么有意思的地方吗?
Yeah, it's bad. I mean, even individual—you can't spot rent single nodes anymore. Like, down the whole stack is cooked. What about the state of the labs? How do you segment the large players—OpenAI, Anthropic, Google? Where's the split of who has the most compute? Anything interesting on the roadmaps of where you think who's investing the most?
我是说,显然是 OpenAI 和 Anthropic 在领跑,无论是从大规模裸金属超大规模云厂商那里租用。换句话说,OpenAI 是 CoreWeave 业务的驱动力,也是微软的。他们还在和其他超大规模云厂商探索。他们也在自建,比如和 Oracle 搞的大规模项目。然后 Anthropic 显然和 AWS 有 Rainier 这个大合作。他们在 Google Cloud 上也有很多 TPU。他们和 CoreWeave 签了协议,也和其他方签了。所以他们是主要的驱动力。你就看 ARR,总得有什么东西在驱动它。那就落到算力这一侧。
I mean, it's clearly OpenAI and Anthropic that are leading the way in terms of both renting from large-scale bare metal hyperscalers. In other words, OpenAI is a driver of CoreWeave's business, also Microsoft. They're also exploring with other hyperscalers. They're doing self-build as well, like massive stuff with Oracle. And then Anthropic obviously got their big AWS partnership with Rainier. They've also got a lot of TPUs with Google Cloud. They signed stuff with CoreWeave. They've signed stuff with others. And so they're the big driver. I mean, you just look at ARR and something has to drive that. And so it goes to the compute side.
在此基础上,另一个方面是,像 Crusoe 这样的很多服务商有很棒的裸金属,对吧?但那不是我们在 ClusterMAX 里测试的东西。我们不测裸金属。我是说,它确实是标准的一部分,但很多标准是关于托管集群的,对吧?托管 Slurm、托管 Kubernetes,诸如此类。所以这是人们在 GPU 之上想要的另一个层级的服务。
With that, another aspect of this is a lot of these providers like Crusoe have fantastic bare metal, right? But that's not what we're testing in ClusterMAX. We're not testing bare metal. I mean, it is part of the criteria, but a lot of the criteria is a managed cluster, right? Managed Slurm, managed Kubernetes, all of that sort of stuff. So it's a different level of service that people want on top of the GPU.
最初的 ClusterMAX 标题——很大一部分叫《如何租一块 GPU》。我们也是在教人们租 GPU 时该怎么做。在我们看来,大多数小玩家——如果你规模很小,你不在乎节点是不是裸金属。如果你规模超级大,你想要裸金属,因为你要部署自己 fork 的 Slurm/Kubernetes,在托管服务上做自己的事。但大多数客户想要托管 Slurm、托管 Kubernetes,以及云在某个节点故障时为他们热切换节点。最大的客户会自己管理这些东西,但中型和较小的客户希望这些由别人托管。所以 ClusterMAX 就是围绕这个展开的。
The original ClusterMAX was titled—also like a big chunk of it was called how to rent a GPU. We're also just teaching people what they should do to rent a GPU. And in our view, most of these smaller folks—if you're super small, you don't care about your node being bare metal. If you're super, super big, you want bare metal because you're going to put your own fork of Slurm/Kubernetes and do your own stuff on managed services and such. But most customers want a managed Slurm, managed Kubernetes, and for the cloud to hot swap nodes for them when one fails. The biggest customers will manage stuff like this, but the midsize and smaller customers want this sort of stuff managed for them. And so ClusterMAX is focused around that.
所以像 Crusoe 这样的公司会租裸金属,对吧?他们和公司签了巨额合同。或者他们在数据中心做得很好,或者他们开始在推理 API 上做得好,但在托管服务上,他们做得不好。而这正是 ClusterMAX 瞄准的——托管服务层,对吧?因为否则就太复杂了。人们想租用的领域有很多不同。有些人想要裸金属,有些人想要托管服务,有些人想要直接输出 token。所以归根结底这是一系列选项。
So companies like Crusoe will rent bare metal, right? They have massive contracts with companies. Or they'll do great in data centers, or they're starting to do well in inference API, but on managed services, they're not doing well. And that specifically is what ClusterMAX is targeting—the managed service layer, right? Because otherwise it's too complicated. It's many different areas people want to rent from. Some people want bare metal, some people want managed services, some people want tokens out. And so ultimately it's a range of options.
就连 Google 也在从别人那里租一些算力,对吧?他们和 CoreWeave 签了协议,和 SpaceX 签了协议。所以他们签下的算力相当大——我想他们从各方签下了超过一千亿美元规模的算力。Nebius 也和实验室有协议之类的。但这些人只做裸金属。OpenAI 和 Anthropic 肯定占了最大份额,买得最多。Google 在这一点上肯定是落后了。OpenAI 和 Anthropic 拥有的算力不相上下,甚至略多一点。到今年年底,两家投入到研发的算力都会超过 DeepMind。所以是的,Google 目前总算力仍然比 OpenAI 多,但其中很大一部分卖给了别处、用于其他产品。但 DeepMind 的研发算力对比 OpenAI 和 Anthropic 的研发算力——这两家实验室现在算力更多,因为他们达到了那个规模。所以所谓‘哦,Google 算力最多,他们会赢’这种说法,其实已经不成立了。他们仍然有成本优势,因为他们自建 TPU,成本低得多,但归根结底 DeepMind 的算力比 OpenAI 和 Anthropic 少。
As far as even Google is renting some compute from folks, right? They signed a deal with CoreWeave. They signed a deal with SpaceX. So pretty large amounts of compute that they've signed over—I think over a hundred billion dollars worth of compute that they've signed from various customers. Nebius also has deals with the labs and such. But these guys are only going for bare metal. OpenAI and Anthropic are definitely lion's share buying the most. Google is definitely falling behind at this point. OpenAI and Anthropic have as much compute, if not a little bit more. By the end of the year they'll both have more compute dedicated to R&D than DeepMind. So yes, Google still ultimately has more compute than OpenAI for now, but a lot of that is sold out to other places and dedicated to other offerings. But DeepMind's R&D compute versus OpenAI and Anthropic's R&D compute—the two labs have more compute now because they've reached that scale. So this whole 'oh, Google has the most compute, they'll win' is actually not really the case anymore. They still have a cost advantage because they build the TPU, and it's much lower cost, but ultimately DeepMind has less compute than OpenAI and Anthropic.
是啊,有个很大的转变。我在 AIE 最大的收获之一就是碰到一些 DeepMind 的研究员,他们非常沮丧,因为他们在说,你看,我们的领导层一直把我们想要的算力卖给别人。
Yeah, there's a big shift. One of the biggest insights for me at AIE was running into some DeepMind researchers and they were very depressed because they were like, look, our leadership keeps selling the compute that we want to other people.
我是说,想象一下 xAI 的处境,对吧?
I mean, imagine being xAI, right?
嗯,是啊。你知道,我还以为 SpaceX 是一家新云呢。搞什么?为什么我们——为什么我们拿不到?
Well, yeah. You know, I thought SpaceX was a neocloud. Like, what the hell? Why are we—why are we unavailable?
他们在这上面。
They're on here.
是啊。等我们能测了,他们就会往上排,你知道。他们会成为一个竞争者。
Yeah. Well, as soon as we can test, they'll move up the rankings, you know. They'll be a competitor.
我觉得 SpaceX 和 Google 的协议就非常有意思,Anthropic 的也是,因为这印证了 Dylan 说的——实验室正在考虑以这种机会主义的方式变成新云。Meta 也在考虑同样的事。你还能看到 Mistral 和 Poolside 也在里面。这些公司具备建设和运营数据中心的能力,然后当有机会把算力变现时——无论是租给别人还是提供推理服务——他们就想去做,但只做到一定程度,而且他们始终想要保留把算力收回、用于研究或英雄级训练的可选项,只要负担得起。他们想现在就尽可能把能负担的算力用于研究。所有额外的部分都是在机会主义地赚钱,来资助那些研究。所以 SpaceX 和 Google 的合同,比如,有双方 90 天的取消权。所以如果 SpaceX 的研究开始见效,他们可以把算力收回来。如果 Google 在某个推理端点上不再靠那部分算力赚钱,他们就把算力收回来。
I think SpaceX's and the Google deal, for example, is super interesting, as well as the Anthropic one, because it goes to what Dylan was saying about the fact that the labs are considering becoming neoclouds on this opportunistic basis. Meta's considering the same thing. You see Mistral and Poolside in there as well. These companies have this skill set of building data centers and managing them, and then when there's an opportunity to monetize the compute—either through renting it to other people or through serving inference—they want to pursue that, but only to a certain level, and they always want the optionality to bring the compute back and use it for those research or hero runs as soon as they can afford to. They want to use the maximum amount of compute they can afford to research right now. And all the extras are opportunistically making the money to fund all that research. And so SpaceX's contract with Google, for example, has mutual 90-day cancellation rights. So if SpaceX's research starts to work, they can pull the compute back. If Google stops making money off of that compute on some inference endpoint, they pull the compute back.
这个结构真的很有意思,双方有点像在对峙。谁会先眨眼?谁更相信 AGI?至少对 SpaceX 来说,对吧?如果你是 xAI 的员工,然后被 SpaceX 收购了,接着 SpaceX 把所有算力都卖了,至少你不会沮丧,因为你会想:“好吧,看,Elon 在建一大堆。我能用上这些,他这么做是为了让我们财务上更稳健。”所以对 Google 的 IPO 来说,就像:我他妈为什么要卖算力?我已经是一家非常赚钱的公司了。为什么我们不能直接用算力来构建模型,然后以高得多的利润率卖这些模型,而不是把算力给 Anthropic 或 OpenAI 或其他公司?SSI 也是 Google 的客户。你怎么看 Google 资助那些离开去搞 NeoLab 的人来买 GCP 这种动态?换句话说,Dave Silver、Jeff Dean,他们分拆出去,从 Google Ventures 拿到资金,然后花在 GCP 上。
It's a really interesting structure where both are kind of in this little standoff. Who's going to blink first? Who believes in AGI more? At least for SpaceX, right? If you're an xAI employee and you got bought by SpaceX and then SpaceX sold all the compute, at least you're not depressed because you're like, "Okay, look, Elon's building a shitload more. I'm gonna have access to this and he's doing this because it helps us be more financially solvent." So for the IPO, um, with Google it's like, why the hell am I selling the compute? I'm already a very profitable company. Why can't we just use the compute to build models and sell those models at way higher margin instead of giving it to Anthropic or OpenAI or other firms? SSI is also a customer of Google. What do you think about the dynamic of Google funding the people that leave to go do a NeoLab to buy GCP? In other words, like Dave Silver, Jeff Dean, they spin out, they get funding from Google Ventures to then go spend on GCP.
我觉得如果只看 GCP 作为一家独立公司,这是个绝妙的举措,对吧?Thomas Kurian 继续推高他的营收、营业利润率和营业利润。但对 Google 整体来说,这真的很蠢,对吧?你不应该激励你最好的研究员离开。你不应该激励以每兆瓦 2000 万到 3000 万美元的价格出售算力,而你本可以用它跑模型赚到每兆瓦 5000 万美元。
I mean, I think it's a fantastic move if you look at just GCP as its own company, right? Thomas Kurian continues to pump his revenue, his operating margin, his operating income. But for Google as a whole, it's really dumb, right? You should not incentivize your best researchers leaving. You should not incentivize selling at $20–$30 million a megawatt when you could run models on it for $50 million a megawatt.
那么如果 Discovery Loop 有重大突破,你觉得 Google 能让他们留在自己的轨道上吗?还是你觉得 Jeff Dean 会带着那个突破去做点什么,而 Google 完全无法从中获益?比如他们会是投资者。他们显然是云服务提供商。
So if Discovery Loop has a big breakthrough, do you think Google is able to keep them in their orbit or do you think Jeff Dean goes and takes that breakthrough and does something without Google being exposed to any of the upside? Like they would be an investor. They're obviously the cloud service provider.
看看 Anthropic 就知道了,对吧?Google 确实投资了 Anthropic。但如果 Google 没投资 Anthropic,Anthropic 今天还会存在吗?因为 Google 是很早期的信徒。像 Opus 3 和 Sonnet 3、3.5,我觉得那些肯定是在 Google 的基础设施上训练的。然后他们后来引入了 Amazon,Amazon 也很棒,但至少 Amazon 没有像 Google 那样训练和销售模型的业务。他们有一些人做,但没什么大不了的。但 Google 有。所以就像,Google 是不是最终扶植了他们最大的竞争对手?现在 Anthropic 不仅从 Amazon 和 Google 租算力,还从许多其他云租。我们现在有 10 多个云的交易。然后 Anthropic 也在建自己的数据中心,租自己的数据中心,开始填满它们,与 Fluid Stack 合作在自己的数据中心部署 TPU。所以当然,Google 还在获得收入,但就像,我以前从你这里租,现在我从你这里买 TPU,这对你的利润率更低。而未来,我要建自己的芯片。我要做一切,不需要你。所以 Anthropic 的情况很清楚。Jeff Dean,如果他在 Discovery Loop 做出惊人的东西,他们为什么要继续?为什么要困在 Google?他们可以用任何云。切换云最终并非不可能。数据引力之类的东西,人们在 AI 前时代常谈的,在 AI 时代已经不那么相关了,因为 AI 可以帮你迁移一切。而 GPU——GPU 不只是 GPU,但如果你够强,GPU 在很大程度上可以作为跨云的可替代资产。
I mean, go look at Anthropic, right? Google did invest in Anthropic. But if Google didn't invest in Anthropic, would Anthropic be here today? Because Google was a really early believer. Like Opus 3 and Sonnet 3 and 3.5, those I think were just trained on Google's infra for sure. And then they later brought on Amazon, which was also great, but at least Amazon is just like—they don't have a business of training models and selling them in the same way. They have some folks, but nothing crazy. But Google does. So it's like, did Google stand up ultimately their biggest competitor? And now Anthropic is not only renting from Amazon and Google but also from numerous other clouds. We have deals with like 10 plus clouds now. And then Anthropic is also building their own data centers, leasing their own data centers, starting to fill them up, working with Fluid Stack to deploy TPUs in their own data centers. So sure, Google's still getting revenue, but it's like, well, I used to rent from you. Now I'm buying TPUs from you, which is less margin for you. And now in the future, I'm building my own chip. I'm gonna do everything without you. So it's pretty clear what happened with Anthropic. Jeff Dean, if he does something amazing in Discovery Loop, then why would they continue? Why would they be stuck to Google? They can use any cloud. Switching clouds is ultimately not impossible. It's not—data gravity and all this stuff that people used to talk about with clouds in the pre-AI era is a lot less relevant in the AI era because AI can just help you migrate everything. And a GPU—a GPU is not just a GPU, but a GPU is largely workable as a fungible asset across clouds if you're good.
那么,你觉得呢?你觉得 Google 完蛋了吗,老兄?
So, what do you think? You think Google's cooked, man?
我觉得你们比我专业多了。而且——是的,Google 就像一堆不同业务的 conglomerate,对吧?所以我确实认为数据优势——我是说作为 NeoCloud 有数据引力——但 Google 的数据优势是大多数人会反复提到的,对吧?而且还有组织挑战,现在有了新领导层。上帝保佑 Cororey,他有着世界上最难的工作来修复这个,但他是个非常非常能干的人。而且他们——我觉得他们知道需要重置,这不是秘密。对我来说,Demis 基本上从领导位置上退下来就像一只完全的黑天鹅。我从未想过这会发生。也许我只是不亲自认识他,他可能会说:“是的,其实我只是很关心 Isomorphic。”但他创立了它。他有点像 DeepMind 的精神领袖。所以这可能是第一个主要前沿实验室的完整领导层过渡。
I mean, I think you guys are much more of an expert than I am. And there's—yeah, Google's just like a conglomerate of a bunch of different things, right? So I do think the data advantage—I mean there's data gravity as a NeoCloud—but the data advantage of Google is the thing that most people keep coming back to, right? And there's just organizational challenges and now there's new leadership. God bless to Cororey who has the hardest job in the world to fix this, but he's a very, very capable guy. And they—I think they know it's not a secret that they do need a reset. And to me, Demis basically stepping out of the leadership seat was just like a complete black swan. I just never thought it was possible. Maybe I just don't know him personally where he's like, "Yeah, actually, I just care a lot about Isomorphic." But he founded it. He's sort of like the spiritual leader of DeepMind. And so this is probably the first major Frontier Lab leadership transition fully.
嗯,我们能说 Sam 离开过两天吗,对吧?但是——
Um, can we say that there were like two days where Sam left, right? But—
我是说他在人们心中从未离开。但是的,这就像 Tim Cook 的——
I mean he never left in anyone's like hearts. But yeah, this is the Tim Cook of like—
是的。
Yeah.
Ilya 离开 OpenAI。
Ilya leaving OpenAI.
不,不是领导者。
No, not the leader.
是的。是的。有点像——是的。是的。我是说,是的。好吧。所以,显然人们分开了,对吧?总之,这就是我的粗略看法。我不觉得我会有比你们更劲爆的观点。我确实想提一件事。这就像是个机会来补上过去六个月的疯狂新闻。
Yeah. Yeah. It was like kind of like the—yeah. Yeah. I mean, yes. Okay. So, like obviously people split, right? Anyway, so that's my rough take. I don't think I'm going to have spicier takes than you guys. I did want to pick up on one thing. This is like just a chance to catch up on the last six months of crazy news.
你们提到了 Poolside。你们从未真正写过 Poolside,因为他们没那么大。但是——他们确实有——他们被收购了——他们的员工被 Nvidia 以 120 亿美元收购了。这是怎么回事?
You guys mentioned Poolside. You've never actually written about Poolside because they're not that big. But there's—they do have—they got bought by—their employees got bought by Nvidia for 12 billion. What's going on?
好吧。是的。这真的很酷,对吧?因为 Nvidia 付钱给所有这些研究员,让他们能把 Neotron 做得更好。
Okay. Yeah. It's really cool, right? Cuz Nvidia is paying for all these researchers so they can make Neotron better.
不,这是防御。Poolside 是 Poolside——你知道从一开始,他们就像——所有云 NeoCloud 都有点烂,所以他们开始大量自建基础设施,他们越来越擅长自建基础设施,然后他们试图部署自己的集群,并与 CoreWeave 合作等等。但他们一直在慢慢向下堆栈移动——也许是第一个 AI 实验室开始运行自己的基础设施,甚至早于 OpenAI 和 Anthropic 尝试自建基础设施。
No, this is defense. Poolside is Poolside—like you know from the beginning, they just like—all the cloud NeoClouds kind of sucked and so they started building a lot of their own infra and they got better and better at building their own infra and then they were trying to deploy their own clusters and had a partnership with CoreWeave on this and stuff like that. But they've just been slowly moving down the stack—maybe the first AI lab to become running their own infra even before OpenAI and Anthropic were trying to build their own infra.
这作为副项目太疯狂了。就像人们必须做这个。这是全职工作吗?这几乎是他们的副项目。
Which is wild as a side project. It's like people have to do this. Is it a full-time job? This is a side project for them almost.
嗯,现在这是他们的全职工作了,对吧?因为他们把所有研究人员都卖给了 Nvidia 来搞 Neotron。Nvidia 显然希望 Neotron 能好得多。我觉得他们对很多开源人士有点失去信心了。当然,Reflection 还在尝试,Thinking Machines 也开源了一些东西,但最终我觉得 Nvidia 对美国开源有点失去希望了。所以,他们认为必须自己来做。于是,他们买了一大堆员工。是的,我以为那是 70 亿美元,不是 120 亿。但不管怎样,我以为大概是 70 亿美元,然后 120 亿里还有 10 亿的投资。
Well, now it's their full-time job, right? Because they sold all their researchers to Nvidia for Neotron. Nvidia clearly wants Neotron to be a lot better. I think they've sort of lost hope in a lot of the open source folks. I mean, sure, Reflection is still trying, Thinking Machines open source some stuff, but ultimately I think Nvidia is losing sort of hope in American open source. So, they think they just have to do it themselves. So, they've bought a bunch of employees. And yes, I thought that was $7 billion, not 12. But anyways, regardless, I thought it was like a $7 billion and then a $1 billion investment out of 12.
我们有细节,是 120 亿。
We have the details in 12 billion.
我选更大的数字。
I take the bigger number.
许可和员工都很厉害。数字是 6。
License and employees are sick. The number is six.
总是选更大的数字。但你知道,我认为最终,Nvidia 可以说,实际上,我们只付了四分之一的价格,因为所有这些钱都只是回去买 GPU 了。那么,他们付了 70 亿美元吗?60 亿给员工,10 亿投资,还是实际上只付了 15 亿,因为那 70 亿都用来买 GPU 了?事实上,他们这样做会激励其他人向他们投资一大笔钱。所以,实际上,雇佣所有这些员工可能是一件积极的事情。就像,他们实际上没花任何钱,因为他们现在创造了一个新客户,一个新的云,将建立一个巨大的集群,对吧?想象一下,他们现在有 70 亿美元的现金。他们再筹集几十亿美元。假设他们资产负债表上有 100 亿美元的现金。他们能够获得 75% 的贷款价值比,你知道,25% 首付,75% 贷款。所以现在他们有 400 亿美元。他们可以建一个吉瓦。太棒了。这是另一个吉瓦规模的 Neocloud。所以他们刚刚创造了一个真正的新竞争对手。
Always take the bigger number. But, you know, I think at the end of the day, it's like Nvidia can say, well, actually, we only paid 1/4 the price because all of this money is just going back to buying GPUs. So, did they pay, you know, $7 billion, six for the employees and one in an investment, or did they actually pay, you know, one and a half because all seven of that is going into buying GPUs? And in fact, them doing this is going to incentivize other people to invest a bunch of money into them. And so, actually, it's probably a positive thing to hire all these people. Like, they actually didn't spend any money because now they've created a new customer, a new cloud that's going to build a massive cluster, right? Like imagine, you know, they've got $7 billion now of cash. They raise another few billion dollars. Let's just say they have $10 billion of cash on the balance sheet. And they're able to get a loan to value of like 75%, you know, 25% down and 75% loan. So now they have $40 billion. They can build a gigawatt. Awesome. This is another gigawatt scale Neocloud now. So they've just made a real new competitor.
我们在文章后面有提到。是的,我的意思是,我不知道那是不是,我的意思是,如果这是剧本,老实说,后面,后面,有领导这件事的人 Robert Bonar,如果这是剧本,应该有更多人这样做。
We have that lower down in the article. It's, yeah, I mean, I don't know if that, I mean, this is the playbook, if this is the playbook, honestly, lower down, lower down, with the guy who's leading it, Robert Bonar, and like if this is the playbook, a lot more people should be doing this.
应该做什么?应该建设,应该和 Nvidia 做这样的交易,因为 Nvidia 会把你捧成王者,只是因为他们想要更多云。是的,我的意思是,不是每个人都发布了美国第二好的开源模型。
Should be doing what? Should be building, should be doing deals with Nvidia like this, because Nvidia will kingmake you just 'cause they want more clouds. Yeah, I mean, not everybody releases the second best open American model.
没错。
That's true.
并且花了几年时间构建它。我的意思是,我们有点忽略了 Poolside 确实构建了一些有趣的东西。对我来说,这次收购和 Hugging Face 都是 Nvidia 的防御性收购,他们不希望竞争对手抢走一个真正有优势的实验室,并且有势头去构建真正的训练和推理集群所需的经验。换句话说,对于构建基础设施公司来说,你真正理解工作负载如何运作,并且有基于训练模型和运行模型形成的观点,这非常重要,而不是像加密矿工那样,他们决定,嘿,我找到了一个站点,我有电力,我有钱,但他们没有能力执行,不像那些已经证明他们能训练出世界上第二好的美国开源模型的人,Poolside 做到了。对吧。
And spends years building it. I mean, we're kind of brushing over the fact that Poolside did build something interesting. And to me, both this and Hugging Face are defensive acquisitions by Nvidia, who does not want their competition snapping up a lab that actually has something going for them and having some momentum where they can build the real experience that's required to build real training and inference clusters. In other words, it's so important for building an infrastructure company that you actually understand how the workload works and you have opinions that are informed by training models and running to get represented in the not crypto miners who decide, hey, I found some site, I got some power, I have some money, they don't have the ability to execute like somebody who has proven they can train the second best American open model in the world, which Poolside did. Right.
是的。
Yeah.
多说点关于
Say more on the
我想说你在那里提到了 Hugging Face。你们怎么看 Hugging Face 的收购?有什么想评论的吗?
I was going to say you mentioned Hugging Face in there. How do you guys see the Hugging Face acquisition? Anything you want to comment on there?
我觉得这很相似。我认为这是防御性的。我不认为 Nvidia 希望其他公司来影响生态系统,拥有像这样关键的公司。
I think it's quite similar. I think it's defensive. I don't think that Nvidia wants the influence in the ecosystem of other companies coming in and owning somebody that's like critical in
这里有一个银河大脑级的观点,对吧,现在 Hugging Face 可以起诉 OpenAI 黑他们。现在 Nvidia 可以 hammered?不,我不这么认为。我不这么认为。但我认为 Hugging Face 显然,你知道,远远领先于其他一切。我团队的一些人已经开始为 Model Scope 提交一堆 PR,以便在整个生态系统中工作。
There's a Galaxy brain take here, right, which is now Hugging Face can sue the out of OpenAI for hacking them. And now Nvidia can has that hammered? No, I don't think so. I don't think so. But I think Hugging Face is obviously like, you know, leagues ahead of everything else. Some folks on my team have started submitting a bunch of PRs for Model Scope to work across the ecosystem.
在什么意义上遥遥领先?
By leagues ahead in what sense?
就像作为开源仓库,为
Just like for being the open source repository for
我明白了。
I see.
是的,就像 Hugging Face 做的那样。但最终,很多都是,你知道,问题是,Hugging Face,你知道,有一种方法,如果你愿意,你可以在 Amazon Trainium 上部署 Hugging Face 模型,对吧?你只需点击一个按钮,它就会部署它们。这会不会消失,从现在开始只有 GPU?或者有没有任何封闭发生?我们拭目以待。但最终,我不知道。我也喜欢价格是表情符号代码。
Yeah, just like what Hugging Face does. But then ultimately a lot of them are, you know, like it's the question is like does Hugging Face, you know, there's a way you can just like deploy Hugging Face models on Amazon Trainium if you want, right? You can just click a button and it'll deploy them. Does that go away and is it only GPUs from here on out? Or is there any closing up that happens? We'll see. But ultimately it's, I don't know. I also love that the price was the emoji code.
太酷了。
So sick.
嗯,团队非常出色。看,我认为我们有点忽略了 Hugging Face 和 Poolside 在这两个组织中都有不可思议的人才。很多人会认为这是 Nvidia 收购他们的动机。我说这是防御性的,因为如果你看 Nvidia 的业务,他们每季度产生近 500 亿美元的自由现金流。所以一年 2000 亿美元。Nvidia 有一个公司战略团队,获得其中一部分钱,说我们需要部署它,对吧?他们可以进行投资,可以进行收购,用自由现金流的 10% 来收购 Hugging Face 和 Poolside。我的意思是,对我来说,这些是你团队中的好资产,至少不为你的竞争对手工作,这是显而易见的。假设 Hugging Face 保持独立,他们遵守法律和收购,他们继续支持所有其他芯片等等。如果他们被其他人收购,他们会积极反对 Nvidia 的利益。我认为这不是他们想要的。所以,我认为你会看到更多这种情况。如果公司成长一点,他们建立这个团队,100 或 200 个不可思议的人,他们想把他们带进来,确保他们留在 Nvidia 生态系统中,而不是去和别人竞争。是的,有道理。不是要贬低 Poolside 所做的任何事情。我们是他们的超级粉丝。我们在他们被收购前一个月请他们上过播客。对于感兴趣的人,模型工厂的工作,训练基础模型的方法,他们构建这一切的方式非常非常好。
Well, and the team's incredible. Look, I think we're kind of brushing over the fact that Hugging Face and Poolside have incredible talent in both these organizations. And like a lot of people would see that's the motivation for Nvidia to acquire them. I say it's defensive because if you look at Nvidia's business, they're generating almost $50 billion of free cash flow a quarter. So $200 billion a year. And there is a corporate strategy team at Nvidia that gets some amount of that money and says what we need to deploy it, right? They can make investments, they can make acquisitions, and 10% of free cash flow to acquire Hugging Face and Poolside. I mean, it's a no-brainer to me that these would be good assets on your team and working at least not for your competition. Let's say Hugging Face stays independent and they follow the letter of the law and the acquisition and they keep supporting all these other chips and stuff. If they were acquired by somebody else, they would be actively working against Nvidia's interests. And I think that's not what they want. So, I think you're going to see more of this. If companies grow up a little bit, they build this team, 100 or 200 incredible people, they want to bring them into make sure that they stay in the Nvidia ecosystem as opposed to going and competing against them with somebody else. Yeah, makes sense. Not to discredit anything Poolside's done. We're pretty big fans. We had them on the pod a month before they got acquired. For people interested, the model factory work, the approach to training foundation models, very very good way that they built all this out.
是的。对我来说,你知道,也许这让我们回到 Clustermax,所有 NeoClouds 应该证明像 custom 4 的方式就是训练一个前沿模型,拜托。
Yeah. To me, you know, coming, you know, maybe maybe if this brings us back to Clustermax, the way that you that all NeoClouds should be proving like custom 4 is like just train a frontier model, please.
我们来聊聊要达到那个目标需要什么。如果要训练前沿模型,我认为 Poolside 和 Hugging Face 过去可能都考虑过被收购,金额也差不多。
Let's talk about what it takes to get there. If you're going to train a frontier model, I think Poolside and probably Hugging Face both entertained acquisition options in the past to the tune of similar money.
是的。具体来说,英伟达曾试图收购 Poolside——抱歉,是英伟达在今年年初试图收购 Hugging Face,可能有些人不知道。
Yeah. Specifically, Nvidia tried to buy Poolside—sorry, Nvidia tried to buy Hugging Face at the start of the year, in case people didn't know.
嗯,我的意思是,我觉得英伟达有点收购狂热。有传言说 Thinking Machines 以及一大堆英伟达的收购。
Well, I mean, I think also Nvidia is on a bit of an acquisition craze. There were rumors of Thinking Machines and a whole bunch of Nvidia acquisitions.
是的。对于很多打造了这些公司的人来说,创始人可以选择在想要的时候套现,对吧?那么,你什么时候决定套现?我认为 Poolside 可能正在观望,心想:“好吧,我已经建好了推出 Lagona 所需的 1 万 GPU 集群。”但如果我想竞争,我需要一个 10 万 GPU、40 万 GPU 的研究集群。就像,我如何跟上明年的开放前沿?我觉得这很难持续下去。钱来了,他们就说:“好吧,当然。我们就套现,和其他人一起建。这似乎没问题。”但另一件值得指出的事情是,在技术方面,Poolside 的论文在基础设施方面真的很好。我和团队谈过——如果你要训练前沿模型,你需要对集群进行健康检查。你需要高可靠性。你不能让模型在训练时没有进展并且不断失败。他们有一套令人难以置信的——他们只是拥有令人难以置信的基础设施。所以,拥有这种人才的人能够将其作为服务提供给其他人,我认为这显然是正确的。而且只有有限的人有这种经验。正如我们所说,即使我们进入银级和铜级,大多数银级和所有铜级的人都没有在训练前沿模型。很多推理,一点强化学习,很多研究,但坦率地说,除了 Azure,我不会信任那些人中的任何一个来构建,对吧?但要构建一个 10 万 GPU 集群,谁已经证明他们能做到?
Yeah. Founders kind of have the option to cash in when they want for a lot of these people that have built these things, right? So, when do you decide to cash in? I think Poolside is probably looking out there and going, "Okay, I've built the 10,000 GPU cluster I needed to get Lagona out." But if I want to compete, I need a research cluster that's 100,000 GPUs, 400,000 GPUs. Like, how do I keep up with the open frontier into next year? I think that's pretty hard to keep going. And the money comes and they're like, "Okay, sure. Let's just cash in and build it with the other people. That seems fine." But the other thing worth pointing out is on the technical side, Poolside's papers are really good on the infrastructure stuff. I've talked to the team—if you're going to train a frontier model, you need to have health checks on your cluster. You need to have high reliability. You can't just have this model not make progress as you're training and constantly be failing. They have an incredible suite of—they just have incredible infrastructure. So the idea that people with that talent would be able to provide that as a service to other people I think is just obviously true. And there's only limited people that have that experience. As we're saying, even when we get into the silver and bronze tier here, people are not training frontier models on most of the silver tier and all the bronze tier, let's say. A lot of inference, a bit of RL, a lot of research, but I wouldn't trust any of those guys frankly to build except Azure, right? But to build a 100,000 GPU cluster, like who has proven they can do that already?
是的。
Yeah.
五家公司。
Five companies.
是的。
Yeah.
是的。
Yeah.
好的。我确实想深入探讨一两个你可能没有实际覆盖的事情。也许鉴于我们剩下的时间,最好的方式就是直接回应批评。你知道,你已经——显然这是一份非常有价值的列表,人们确实会根据它做出购买决定。Periodic Labs 确实站出来说:“是的,实际上,这帮助我们决定我们的支出。”是的,我的意思是,你知道,大多数人有什么批评?
Okay. I do want to maybe dive in on one or two things that maybe you don't actually cover. Maybe the best way given the time that we have left is to just address the criticism. You know, you've—obviously this is a very valuable list, like people do make purchasing decisions based on it. Periodic Labs did come out and say like, "Yeah, actually, this helps us decide what we spend on." Yeah, I mean, you know, what is the criticism that most people have?
我认为有很多针对很多事情的批评。首先,这是在测试托管集群。所以这并不意味着一家公司在其他方面不好,对吧?Crusoe 因为托管集群被降级了。但他们是世界上最好的数据中心建设者之一。如果不是最好的。他们拥有几乎最多的已签约——或者最多的已签约吉瓦在管道中,以及一堆他们正在快速建设的站点。但这不是对他们数据中心建设能力的批评。这只是对他们托管集群产品的批评。他们仍然是较好的之一,对吧?但他们被降级了。同样,其他公司——Together 有很好的推理端点,这与托管服务无关。所以我认为围绕这个测试的内容有很多困惑,对吧?这是针对托管集群的,不是针对数据中心建设或裸金属集群或推理端点,这些界限变得模糊,因为每个人都在做这些事情,对吧?另一个主要的批评是:“哦,他们是不是收钱来做这个?”就像,不,我们不收钱来做任何排名。那会是非法的。我自己在 Crusoe 有个人投资。我在 Fluid Stack 有个人投资,这两家公司——我在这两家都做了 SPV。这两家公司都不在前五云中。因为那不是——我不是来偏袒的,对吧?Jordan 会说他有自由做他想做的事。技术上,他甚至不知道我们签合同的所有公司。如果我看世界前 20 家公司,有五家不是我们的客户。像礼来和沙特阿美之类的。但最终,那些是比 CoreWeave 或 Nebius 大得多的客户,对吧?因为他们是更大的公司,有更大的预算。所以比如 AWS 是比 CoreWeave 或 Nebius 更大的客户。然而这与排名无关。基础设施供应链中的每个人都买我们的东西。就是这样。如果你在跟踪数据中心,你想知道其他人在他们的数据中心做什么。如果你在跟踪加速器供应链,如果你想了解各种冷却器的成本,如果你想了解网络和光学供应链,你会买我们的研究和数据服务。这并不意味着——你经常会委托我们做更多的工作。但这并不意味着我们会偏袒我们的排名。所以我认为围绕这个有很多批评,但最终,首先,如果我们收钱并排名而不披露,那会是非法的。但其次,这没有意义。我们对行业的价值在于我们说出真相和我们的信念。如果我们不这样做,那么没人会在意我们说什么,我们就会对着虚空尖叫。所以我们的声誉是最重要的,胜过快速赚钱。
I think there's a lot of criticism for a lot of things. First and foremost, this is testing managed clusters. So that does not mean that a company is not good at other things, right? Crusoe was downgraded for managed clusters. But they're one of the best data center builders in the world. If not the best. They have nearly the most contracted—or the most contracted gigawatts in their pipeline and a bunch of sites that they're building on really rapidly. But this is not a criticism against their data center building capabilities. It's only a criticism against their managed cluster offerings. They were still one of the better ones, right? But they were downgraded. Likewise, other companies—Together has great inference endpoints that has nothing to do with managed services. And so I think there's a lot of confusion around what this is testing, right? This is for managed clusters, not for data center construction or bare metal clusters or for inference endpoints, which become a blurrier line because everyone is doing these things, right? The other big major criticism we have is like, "Oh, are they taking money for this?" It's like, no, we do not take money for any ranking. That would be illegal. I myself have personal investments in Crusoe. I have personal investments in Fluid Stack, and neither of those companies—I did SPVs in both of them. Neither of those companies are in the top five clouds. Because that's not what—I'm not here to be biased, right? Jordan will say he has free reign to do whatever he wants to do. Technically, he doesn't even know all the companies we do contracts with. If I look at the top 20 companies in the world, there are five companies that are not our customers. Like Eli Lilly and Saudi Aramco and stuff like that. But ultimately, those are much bigger customers than CoreWeave or Nebius could ever be, right? Because they're way bigger companies with way bigger budgets. And so like AWS is a bigger customer than CoreWeave or Nebius, for example. And yet that has nothing to do with the rankings. Everyone in the infra supply chain buys our stuff. It just is what it is. If you're tracking data centers, you want to know what other people are doing in their data centers. If you're tracking accelerator supply chains, if you're trying to understand the cost of various kinds of chillers, if you're trying to understand the supply chain for networking and optics, you're going to buy our research and our data services. That does not mean—and you'll often commission us to do work beyond that. But that doesn't mean that we're going to ever bias our rankings. So I think there's a lot of criticisms around that, but ultimately, first of all, if we were taking money and ranking people and not disclosing it, that'd be illegal. But second of all, it makes no sense. The value we have to the industry is that we tell the truth and what we believe. If we didn't do that, then no one would care what we would say and we'd be screaming into the void. So our reputation is what matters the most over a quick buck.
是的。也许一种说法是,你是一个买方服务,人们指责你收卖方钱,他们不明白买方在收入方面可能比卖方大 1000 倍。
Yeah. Maybe one way to put it is like you're a buy-side service and people are accusing you of taking money for the sell side and they don't understand the buy side is probably 1,000 times larger in terms of revenue for you than the sell side.
我的意思是,我觉得人们试图挑刺说这是一份付费榜单,你知道,那只是一个角度。更广泛的观点是,好吧,人们误解了你在这里试图排名的是什么。很多说法只是,嗯,这太疯狂了,这个云比 Y 更好,对吧?但那并不是你实际测试的东西。
I mean, I think people trying to poke at it's a paid list, you know, and that's just one IQ. The broader point is like, okay, people are misunderstanding what you're trying to rank here. And a lot of the claims are just, well, it's crazy that this is a better cloud than Y, right? But that's not what you're actually testing.
我听到两种关于写作的明确批评,我倾听这些批评,因为我尊重他们的工作,而且他们不是那种匿名的、对着虚空喊叫的人。这些批评是:你们写得不够多,你们应该发布更多数据,你们应该写更多;或者这是 30,000 字,我读不完。所以我们在两者之间平衡真的很难。
I hear two clear criticisms about the writing from people where I listen to the criticism because I respect their work and they're not like anonymous just shouting into the void. And those criticisms are: you guys didn't write enough, you should have released more data, you should have written more; or this is 30,000 words, I can't read this all. So it's really hard for us to balance between those two.
但这是一个新的……是的。你当然是在评估新云,就像……
But it's a new... Yeah. You're evaluating neo clouds of course like...
所以对于前者,我对那种批评感兴趣,我们想做得更好。所以我们有这个研究产品叫做云 TCO 模型,任何人都可以来授权。这是年费。我们每周写关于行业趋势的笔记。我们有一个仪表板,我们在下面放了一张经过编辑的截图,你可以获取我们所有的数据。我们所有的性能数据都在上面。是的,就像你让我们全年每周为你做研究,而不是仅仅为这个每年一次或每六个月一次的发布。所以这是为那些想要更多细节的人准备的。
So for the former, I'm interested in that criticism and we want to do better. And so we have this research product called the cloud TCO model that anybody can come and license. It's a yearly fee. We write notes every week about the trends in the industry. We have a dashboard which we took a redacted screenshot of below where you can get all of our data. All of our performance data is on there. Yeah, it's like you get us working on the research for you year round weekly rather than just for this individual once a year release or every six months or something is what we do. So that's for the people who want more detail.
还有另一群批评,比如,你知道,通常是那些持有某只股票的人,而我们把它们排名很差或很好,对吧?就像我记得上次人们说,顺便说一句,我们对一家公司在股市表现更好或更差的看法与 Jordan 和他的团队无关。你知道,我们确实对所有公司进行前瞻性收入估算,然后你可以授权我们的数据,看看你认为这应该意味着你应该交易什么。我们从不推荐交易。但有时候,比如我们认为 Nebius 比 CoreWeave 更好。我们认为 Nebius 比 CoreWeave 更好的原因,尽管我们认为 CoreWeave 仍然是比 Nebius 更好的云,是因为 Nebius 的合同期限更短,所以当 GPU 价格飙升时,Nebius 能更好地获得上涨空间,对吧?就像我们在研究中发布的东西,是在另一个团队,不是 Jordan 和那些在 cluster max 上工作的技术人员。或者嘿,就像 Iris Energy 有这个巨大的数据中心即将上线,我们认为它会被租赁,我们会发布一份报告,同时说你不应该使用 Iris Energy 的管理服务。所以就像完全不同的东西,对吧?你知道,他们把那个数据中心作为裸金属卖给了微软。所以,你知道,我们在机构研究方面的研究,对于可能购买股票的人来说,与嘿,cluster max 真的是为像 Periodic Labs 或 Applied Compute 这样的公司,或者那些公开表示喜欢 cluster max 和我们所做事情的公司而做的,完全不同。你知道,甚至 OpenAI 也说他们喜欢 cluster max,我们有一句他们的引述。我们有各种不同的买家说他们真的很欣赏这种方法论,并且我们正在为所有云提高标准。所以我认为,你知道,如果你试图根据 cluster max 来买股票,那完全错了。它与股票表现无关。
There's another crowd of criticisms which is like, you know, it's oftentimes someone who like owns one of the stocks and we rank them poorly or positively, right? Like I remember last time people were like, by the way, like our view on a company doing better or worse in the stock market has nothing to do with Jordan and his team. There's, you know, we do estimate revenue for all these companies going on a forward basis and then you can like license our data and see what you think that should mean you should trade. We never recommend us to trade. But there have been times where like for example we thought Nebius was better than CoreWeave. And the reason we thought Nebius was better than CoreWeave even though we think CoreWeave is still a better cloud than Nebius is because Nebius's contracts are shorter term and so with prices of GPU spiking, Nebius gets to take the upside better, right? Like that's something that we've released in our research in a separate group of people, not Jordan and the technical staff that are working on cluster max. Or hey, like Iris Energy has this massive data center that's going to come online and we think it's going to get leased and we would put a report out about that while at the same time saying you should never use Iris Energy's managed services. And so like it's completely different things, right? You know, they sold that data center as a bare metal to Microsoft. And so we were, you know, our research on the institutional research side for someone who might be buying stocks is completely different from like, hey, cluster max is really made for people like Periodic Labs or Applied Compute or companies like that who have publicly came out and said they like cluster max and what we do there. You know, even OpenAI said they like cluster max and we have like a quote from them on that. We have a variety of different buyers who said they really appreciate the methodology and that we are pushing the bar up for all the clouds. And so I think it's, you know, if you're trying to buy stocks based off cluster max, that's like completely wrong. It has nothing to do with stock performance.
是的,100% 伙计。这严格是为买家准备的。这是为了帮助人们购买算力。这不是为了让人们挑选什么股票来把所有储蓄都梭哈进去。我的意思是,它会在某种限度上相关。我想,是不是有一种倾向,真的……真的从你们那里,你们太……好吧。好吧。好吧。天哪。哦,我的上帝。
Yeah, it's 100% man. This is strictly for buyers. This is for helping people buy compute. It is not for people to pick what stock to yolo all their savings into. I mean, it will correlate at some kind of limit. I think like, isn't there a tendency to literally... literally from you guys are so okay. All right. All right. All right. Geez. Oh my god.
不。真的,对吧?就像有人不擅长运营云,但有一个他们正在建设的吉瓦级数据中心,而且它实际上是可信的,就像一个吉瓦级数据中心。嗯,他们可能是 cluster max 排名最差的,但他们仍然可以赚很多钱。就像在当前股价下,这有关系吗?许多名字便宜,许多名字昂贵,就像信任。
No. Literally, right? Like someone who's terrible at running a cloud but has a gigawatt data center that they're building and it's actually credible like a gigawatt data center. Well, they could be the worst ranked cluster max and they could still make shitloads of money. Like does it matter at the current price at the current price of the stock? Many names are cheap, many names are expensive like trust.
是的,我明白。基本面与,你知道,一个贡献于整体基本面的产品。它贡献于整体估值和方向。我以前是对冲基金的人。我明白。你知道,我只是说,这就像你们是领域专家,人们确实重视那个元素的评级。但显然这更多是买方,买方是指新实验室,只是不断强调这一点,但我们真的有一个核心研究项目,人们报告股票并给出对这些名字的意见,许多从事那项研究的人,他们问我们问题,但然后他们出来说,表现不佳或不可用层级的人对这个名字持积极态度,这与我们公开发布的、说你应该从谁那里租集群不同,也许我需要在未来加另一个免责声明,说这不是对他们股票当前估值的认可。
Yeah, I do get it. Fundamentals versus, you know, one offering that contributes to the overall fundamentals. It contributes to the overall valuation and direction. I used to be a hedge fund guy. I get it. You know, I'm just saying like this is like you guys are the domain experts and people do value the rating on that element. But obviously this is way more buy side and buy as in the neolab to just keep stressing this but we genuinely have a core research project where people report on stocks and they give opinions about the names and many of the guys that work on that research look they ask us questions but then they come out and they say people in the underperforming or unavailable tier were positive on the name this is different than what we put out publicly and say you should rent clusters from and maybe I need another disclaimer on the future one that's like this is not a cosign of the current valuation of their stock.
它可能,你知道,它可能真的很有趣,比如一个 2x2 或二维图表,这里显示 cluster max 排名,然后这里是同样的分析,比如股票排名,差异将是有趣的研究。也许值得说的是,还有更多内容即将到来,那就是 Dylan 提到了这里的个别名称和标志,它们的业务远大于我们在这里排名的管理集群,所以我们将在新闻通讯中讨论推理端点、裸金属、R 基础设施、沙箱、RL 基础设施的一部分,但有点不同,我们将写所有这些,敬请期待,因为我认为人们将开始对这些在不同市场细分中竞争的业务有更全面的看法,包括推理、管理集群和裸金属。我们在这里有很多关于管理集群的关注,但坦率地说,很多人联系我们,帮助决定选择哪个端点提供商,或哪个托管 RL 训练提供商,或哪个沙箱提供商。这与为你的管理集群选择谁是一个不同的问题。
It could be, you know, it could be really interesting for like a 2 by 2 or like a two-dimensional chart where like here's the cluster max ranking and then here's the same analysis like stock ranking and like the differences are going to be interesting research. Maybe worth saying though there is a lot more coming which is that Dylan mentioned it individual names and logos on here have businesses that are much larger than just managed clusters which is what we're ranking here and so we are going to be addressing that on the newsletter for inference endpoints for bare metal for R infrastructure for sandboxes for you know part of RL infrastructure but kind of different we're going to be writing about all of that and stay tuned for that because I think people will start to have a more holistic view of these businesses that compete in different market segments across inference, managed clusters, and bare metal. And we've got a lot of focus on managed clusters here, but frankly, so many people are reaching out to us for help deciding on which endpoint provider to go with or which hosted RL training provider to go with or which sandbox provider to go with. And that's a different question than who to go with for your managed cluster.
是的。
Yeah.
所以解决方案是更多的 maxis、更多的基准测试吗?
So is the solution more maxis, more benchmarks?
对,没错。是的。
Yeah, that's right. Yeah.
好。你们现在有几个了?你们有 InferenceX。
Okay. How many do you have? You have InferenceX.
还有一个就像,你去购物、去杂货店的时候,你饿了。道理是一样的,对吧?我们什么都想做。再看吧。再看吧。希望我们都能做出来。
You have one on like, when you go shopping, when you go to the grocery store when you're hungry. It's the same way, right? We want to do everything. We'll see. We'll see. Hopefully, we can execute on all of it.
是的。不过也许应该说清楚。今天的 InferenceX,名字是因为它测的是推理性能,但它其实是关于芯片的。它是关于比较芯片和系统的。我们接下来会发布一个叫 EndpointX 的东西,用来测试无服务器推理端点提供商,深入那些细节。因为我觉得方法论非常相似。我们知道应该套用类似的方法论,不只停留在性能上。我们看成本,看可靠性,看安全性,看监控基础设施。一个推理端点涉及的东西太多了。仅仅因为某家在 batch one 这种人为设计的测试里跑出了更高的数字,并不代表当你需要那些额外能力时,它就是那个能帮你扩展推理需求的合适伙伴。
Yeah. Maybe should be clear though. InferenceX today, the name is because it's for inference performance, but it's about chips. It's about comparing chips and systems. And we're going to be releasing something called EndpointX where we test the serverless inference endpoint providers and kind of dig into those details. Because I think there's a very similar methodology. We know there's similar methodologies that we should apply where we go beyond just performance. We look at cost, we look at reliability, we look at security, we look at monitoring infrastructure. There's all of this stuff that goes into an inference endpoint. Just because somebody hits a higher target on batch one, contrived tests, doesn't necessarily mean they're the right partner to scale up your inference requirements with when you need all that extra.
对,有些提供商性能更好,就是原始性能更好,优化工程师更强,但他们的 API 稳定性更差,更容易出故障。有各种各样的原因导致它对端点的终端客户来说未必更好,尽管它每用户每秒的 token 数更多、成本更低。
Yeah, there's providers with better performance, like raw better performance, better optimization engineers, but then their API is less stable, it's flakier. There's all sorts of reasons why it may not be as good for the endpoint end customer even though it's like more tokens per second per user and lower cost.
是的。太棒了。你们会怎么覆盖,我们姑且称之为非英伟达阵营?也就是 Ryzen。
Yeah. Amazing. Do you ever, how will you cover the let's call it the non-Nvidia universe? So which is Ryzen.
你们这里已经有一些了,对吧?比如 TensorWave 就是 AMD 云。
You have some in here, right? Like TensorWave is AMD cloud.
对,这也是另一个我明确有投资的,顺便说一句各位。但我们还是把它们排得很低。你知道,Time Intellect,我在第一个 cluster max 里有投资。我们把它列为表现不佳,对吧?我觉得很多这种偏见,是的,我可能会投资,我会去跟他们聊、试着帮他们,但归根结底排名就是排名。所以我们确实想往外扩展,但说到底,你知道,Fluid Stack 不会把 TPU 租给我们,对吧?因为他们把所有东西都租给 Anthropic 了,对吧?
Yeah, that's another one where I explicitly have an investment, by the way, guys. And we still rank them poorly. You know, Time Intellect, I had investment in the first cluster max. We put them as underperform, right? I think like a lot of this bias against, it's yes, I might make an investment, I'm going to talk to them and try and help them, but at the end of the day the ranking is the ranking. So we do want to extend beyond, but ultimately like, you know, Fluid Stack is not going to rent TPUs to us, right? Because they're renting everything to Anthropic, right?
也许吧。
Maybe.
再看吧。我希望。我希望。
We'll see. I hope. I hope.
我在付钱给他们。
I'm paying them.
是啊,他们不可用有点说不过去。拜托,各位。
Yeah, it's kind of inexcusable that they're unavailable. Like, come on, guys.
不,我是说,这同时也说得通,对吧?就像,你看,它对我们不可用,所以我们租不到。你也租不到,除非你真的真的特别牛。因为他们把一切都卖光了。就像 SpaceX 不会租给我们。他们不会租给你。他们不会租给,你知道,随便什么公司。他们会租给巨头。
No, I mean, it's valid at the same time, right? It's like, look, it's unavailable to us, so we can't rent it. Won't be able to rent it either unless you're really really freaking cool. Because they're selling everything. Like SpaceX is not going to rent to us. They're not going to rent to you. They're not going to rent to, you know, random companies. They're going to rent massive.
Dylan。我觉得你需要练练你的,我觉得你需要,这是我第一次给你这个反馈,但我觉得你需要练练你的自信,老兄。我们能挤进去的。我觉得我们能跟这些家伙挤进去。是的。
Dylan. I think you need to work on your, I think you need to, this is the first time I'm ever giving you this feedback, but I think you need to work on your self-confidence, man. We can get it in there. I think we can get it in there with these guys. Yeah.
我觉得随着时间推移,就像今天,很多这些提供商,甚至 Fluid Stack 和我们认识的现在在 SpaceX AI 的人,都觉得读这篇文章有价值,觉得获得外部意见有价值,即使他们不一定会把提供给我们的服务公开提供。他们喜欢第三方的,你知道,来自公司外部的、诚实的技术意见,能给他们真实反馈的意见。我们就是来帮忙做这些的。
I think over time, like today, a lot of these providers, even Fluid Stack and people that we know who are at SpaceX AI now, see value in reading the article and they see value in getting an outside opinion even if they're not necessarily going to provide the service publicly that they offer to us. They enjoy a third, you know, a technical opinion that's outside the company, that's honest, that's going to give them real feedback. And we're here to help for all of that stuff.
回答关于多芯片的问题,明年会非常庞大。就像英伟达的 Vera Rubin 这一代,我们会看到它,因为他们会把 Groq 的 LPU 引进来,就是 LPX 系统,对吧?所以即使我们只测英伟达,我们也会有多种芯片可测。AMD 显然有自己的东西。Cerebras 正在成为 OpenAI 的新云。他们在建自己的数据中心,对吧?从单纯的推理端点业务,转向真正为其他所有人管理他们芯片的集群。我们看到 SambaNova 进入市场。很多芯片初创公司,像 Positron、Etched、MatX,很多人在推出这类东西。有传闻,你知道,显然 Fluid Stack 在 GCP 之外拿到了 TPU。那边还有更多东西要来。Trainium,AWS 能生产很多芯片。也许未来他们不会全部自建。他们已经有关于异构 Trainium 与 Cerebras 等的公告。所以我觉得明年这方面会有更多,我们将不得不把多芯片作为领先新云综合方案的一部分来应对。未来他们需要能够支持异构算力。
To answer the question about multi-silicon, it's going to be huge next year. Like with the Vera Rubin generation with Nvidia, we're going to see it because they're going to bring the LPUs in from Groq, the LPX systems, right? So even if we're just testing Nvidia, we're going to have multi-silicon to test. AMD obviously has their own thing. Cerebras is becoming a neocloud for OpenAI. They're building their own data centers, right? Moving from just inference endpoint stuff to like really managing clusters of their chips for everybody else. We see SambaNova coming to market. Lots of the chip startups like Positron, Etched, MatX, like lots of people are launching stuff like this. There's rumors, you know, obviously Fluid Stack's got TPU outside of GCP. There's more stuff coming there. Trainium, lots AWS can produce a lot of chips. Maybe they won't self-build all of it in the future. They've got announcements on heterogeneous Trainium with Cerebras and stuff like that. So I think there's going to be a lot more for that next year, and we're going to have to contend with multi-silicon as part of a comprehensive offering from the leading neoclouds. They will need to be able to support heterogeneous compute in the future.
而这一切主要还是聚焦在推理上,对吧?因为它是,你知道,就像
And that's all focused on inference mostly, right? Like because it's, you know, it's like
我是说,比如 Anthropic 大部分主要用 TPU 做训练。所以那种情况下有点相反,对吧?他们其实不是,他们大部分推理是在 Trainium 和 GPU 上。他们大部分训练其实是,尤其像,你知道,Fable 5 就是在 TPU 上训练的,用于预训练,然后后训练和推理主要是在 Trainium 和 GPU 上。
I mean, most Anthropic, for example, uses TPUs for training primarily. So in that case it's sort of the opposite, right? They don't actually, most of their inference is on Trainium and GPU. Most of their training is actually, especially like, you know, Fable 5 was trained on TPUs for example, for the pre-training, and then the post-training and inference is mostly Trainium and GPUs.
所以这是多芯片。我会说大多数情况下其他芯片方案是用于推理的,但有些情况下人们也用它做训练。
So it's a multi-silicon. I would say in most cases other silicon offerings are for inference, but there are some cases where people are using it for training.
是的。但我觉得关键在于,推理和训练现在不是这种硬性划分了,因为当你在做研究、在真正对模型做后训练时,强化学习过程中有大量的前向传播。所以
Yeah. But I think the key thing is inference versus training is not this hard split now, where there's lots of forward passes during RL when you're doing research and when you're actually post-training the model. And so
事实上大部分工作负载都是前向传播,对吧?所以你可能会把它经典地称为推理,即使它是用于,你知道,强化学习的前向传播。
In fact most of the workload is forward passes, right? And so you might call that classically inference even though it's for, you know, forward passes for reinforcement learning.
是的,推理是循环的一部分。
Yeah, inference is part of the loop.
使用,比如说,更便宜、能高吞吐的芯片来做强化学习的前向传播,有一大堆好处。但如果你用两种不同的芯片,就会有一大堆数值正确性方面的挑战。哪怕只是两代不同的英伟达 GPU,B200 对 B300,你可能内存量不同、批大小不同。有所有这些效应。所以我绝对预期训练和推理中都会用到异构芯片。预训练大概是同构集群。人们大概不会在 40 万 GPU 这种规模上去冒那个险,但再看吧。
There's a whole bunch of benefits of using, let's say, cheaper chips capable of high throughput for the forward pass in RL. But then there's a bunch of challenges with numerical correctness if you're using two different chips. Even just two different generations of Nvidia GPUs, B200 versus B300, you might have different amounts of memory, different batch sizes. There's all these effects. So I absolutely expect there to be heterogeneous silicon used across training and inference. Pre-training probably a homogeneous cluster. People probably not going to risk that at like 400,000 GPU scale, but we'll see.
是啊,但你知道现在强化学习的工作负载跟预训练差不多大,有时甚至更大,对吧?所以这实际上意味着你可能需要一种新的解耦形式,把不同种类的算力放在同一个数据中心里来回调度。
Yeah, but like you know RL is like as big or sometimes bigger than the pre-trained workload right now, right? And so that actually means that you need like maybe a new form of disaggregation where like you have the different kinds of compute in the same data center that you shuttle back.
甚至都不一定非得是同一个数据中心,对吧?人们正在很多很多站点上做强化学习。这个数据中心在做 token 生成,那个数据中心也在做 token 生成。而且在那里用不同类型的芯片其实也没问题。这只是一个非常复杂的信息问题,因为你要用不同的方式对模型做分片,权重更新的时间点也不一样,还有这一大堆事情。
It doesn't even have to be the same data center, right? People are doing RL across many many sites, right? This data center is doing token generation. That data center is doing token generation. And it's actually fine to do different types of chips there. It's just a very complex info problem because you're going to shard your model differently and weight updates happen at different times and there's all this.
异步——现在听这期节目的人估计都在想,天哪,千万别。
Async for people listening to this right now are like god no please no.
你得在数据中心里做多芯片的强化学习实时更新,天哪,求求你别。
You have to data center multisilicon RL updates on the fly god please no.
你是在跟光速较劲,你怎么可能把这所有数据都传过去。
You're fighting speed of light how how are you going to transfer all this this data.
我是说,在解耦式、去中心化强化学习方面,确实有一些有意思的架构工作在推进。刚出来的 MIMO 2.6 就是一篇很好的公开论文,讲他们如何把强化学习这一侧超大规模地扩展。不过说到这一点,你也提到了很多定制芯片。华为、中国芯片方面有什么可说的吗?你们会关注它们吗?
I mean there's interesting architecture stuff being done on disaggregated like decentralized RL. Well, the MIMO 2.6 that just came out, it's a good open paper on how they super scale up the RL side of things. Um, but on that point as well, you mentioned a lot of the custom chips. Anything on Huawei, Chinese chips? Um, do you guys look at them?
是的,我们在跟踪它们。华为的芯片确实进步得快多了。有一段时间,你知道,基本上它们被美国政府禁运之后,花了几年才重新调整过来。你知道,还有昇腾 910 B 和 C,但现在它们已经相当快地接连推出了 940、950 和 960。你知道,我们正试着获取它们,并在各个方面进行测试,对吧?不管是租用它们来测试性能,还是去看规格和制造产能之类的。也试着拿到它们,这样我们就能在俄勒冈的逆向工程实验室里做逆向工程分析。我们正试着在所有这些维度上开展工作。你知道。
Yeah, we're tracking them. The Huawei chips are really getting better much faster. For a while, they were like, you know, basically when they were banned by the US government, it took them like a few years to reset. um you know and the Ascend 910 B and C uh but now they've launched the 940 950 and 960 in pretty quick succession. Um they you know we're trying to sort of get access to them and test them along all bounds, right? Whether it be renting them so we can test the performance, whether it be um you know looking at the specs and the manufacturing capacity and things like that. um also trying to access them so we can do a reverse engineering analysis out of our Oregon reverse engineering lab. Um we're trying to do work across all of these um dimensions. Um you know
可用性怎么样?有没有集群——那些实验室没有英伟达也还行吗?总体上有没有什么早期迹象?
How's the how's the availability? Are there clusters like are the labs okay without Nvidia? Any any early signs in general?
是的,我是说人们非常——我是说那些实验室在用,你知道,Anthropic 用了四种不同的
Yeah, I mean people are super I mean the labs are using you know anthropic uses four different
不,我是说那些中国实验室用它们自己的芯片。
No, I mean like the the the Chinese labs using their own chips.
是的。所以我觉得,比如 GLM 就明确说过,对吧,他们有 5 万块不是英伟达的芯片。那是中国制造的 AI 加速器,用于他们最新模型的推理。当我们把范围扩展到其他公司,你知道,Cameracon、Ilvatar、Inflame 以及很多其他公司,有那么几家不同的公司。它们都在往上追赶。你知道,显然它们还落后英伟达不少。华为绝对是它们当中最好的。然后,BU 正在把它的芯片业务分拆成一家完全独立的公司。所以中国芯片生态里有很多进展,尤其是在推理方面,我觉得你知道,随着前沿模型——甚至现在开源模型——在编写内核方面变得越来越好,移植东西也越来越容易,最终你会有更多的多样性和异构性,这也是为什么英伟达转向了这种超快的节奏,试图每年发布新芯片、每年发布新系统,因为他们知道自己必须以光速尽可能快地跑,否则就会被追上,因为 CUDA 的护城河正在迅速瓦解。
Yeah. So I think it was like um GLM said explicitly right that they had uh 50,000 chips uh that were not Nvidia. They were Chinese-made uh AI accelerators for inference of their newest model. Um as we extend out to other firms um you know Cameracon, Ilvatar, Inflame, and many others, there's there's a handful of different companies. Um they're all pushing up the curve. You know, obviously they're further behind Nvidia. Huawei is definitely the best of them all. Um and then uh BU is spinning out their chip business uh as a completely separate company. Um so there's a lot of developments in the Chinese chip ecosystem and and especially for inference I think you know the Chinese you know as as frontier models um become and even even now the open source models become better and better at writing kernels uh it's easier and easier to port stuff um and and so ultimately you're you're going to have more diversity and heterogeneity uh which is why Nvidia's moved to this super fast pace of trying to release new chips every year and new systems every year because they know they have to run as fast as possible at the speed of light otherwise they will get caught up to um because the CUDA mode is rapidly being deteriorated
但这只是产量问题,对吧 Dylan,这些芯片今天确实能用,但他们生产不出足够多的量来真正满足所有国内需求。
But it's just volume right Dylan like the um these chips do work today but they can't produce enough of them to really serve all the domestic demand.
所以
So
是的。
Yeah.
我是说,你知道,如果你看——如果你做 H100 的等效产品,或者假设做 GV200 的等效产品,英伟达在造数以百万计的芯片,谷歌在造数以百万计的芯片,亚马逊在造数以百万计的芯片,AMD 在造数以百万计的芯片,对吧?但如果你看中国公司,实际上,华为是唯一一家即便按总芯片数算也能造出 100 万块以上的。然后华为的芯片还落后。所以你还得为此打个折扣。其他所有公司都还停留在几十万块的量级,而扩大生产非常困难,尤其是考虑到出口管制之类的因素。我认为中国会在明年和后年迅速提升产量。到 2028 年,他们甚至会达到生产数千万块芯片。但这确实需要时间。
I mean it's it's you know you know if you look if you do an H100 equivalent or let's say if you do an GV200 equivalence um Nvidia's making you know millions and millions of chips, Google's making millions and millions of chips. Amazon's making millions and millions of chips. AMD is making millions of chips, right? But then if you look at the Chinese companies, really, Huawei is the only one that even in gross chips, they're making a million plus. Um, and then the Huawei chips are are behind. So then you have to discount them for that. Everyone else is still in the hundreds of thousands of units and scaling production is very difficult, especially given export controls and stuff. Um, I think China will ramp production rapidly across next year and the year after. Um and they will get to even tens of millions of chips produced by 2028. Um but it does take time.
是的。我是说,别忘了他们是第一个围绕这一切制定国家计划的国家。你知道,当他们动起来的时候,他们动起来的规模是人类历史上从未见过的。嗯,好,这里有很多可聊的。恭喜 cluster max,也恭喜你们正在做的所有事情。其实我觉得我们算是聊到了今年到目前为止的很多主要话题,这很棒。也了解一下你们在俄勒冈建的那个硬件实验室。我其实很想去看看。我们应该安排一次行程之类的。
Yeah. I mean uh let's not forget that they they're the first to have a national plan around all this stuff. You know when when they get going they get going in real scale like unlike we've ever seen in human human history. Um okay lots to discuss here. Uh congratulations on cluster max uh and all the stuff that you guys are doing. Actually I think we managed to hit on like many of the major topics of like the year so far which is great. get a catch up on like the the hardware lab that you guys are building in Oregon. I'd actually love to see it. We should we should do like some kind of trip.
好啊。飞过来。飞过来。我们会带你参观一下。
Yes. Fly out. Fly out. We will uh give you a tour.
好的。
Yeah.
好的。你要去的时候提前跟我说一声,我们就安排参观。我想
Yeah. Give me a heads up whenever you are going like we'll do a tour. Uh I want to
Dylan 会开他那辆橙色吉普载你一程。
Dylan will give you a ride in his orange jeep.
嗯,我们——我想去 Abalene。咱们就把这些事都安排上吧,
Um we I I want to go to Abalene. Like let's let's just do all these things,
你知道吗?嗯,我是说显然你们也去过那里。总之,你们得走了。很高兴能叙叙旧。恭喜你们的一切,而且我觉得还有很多工作要做。Jordan 给我们透露了一点你们在做的事情。endpoint X 听起来很有意思。而且市场真的在细分,对吧?对我来说,这就是你们早早识别出来、并且在为所有人提高标准的事情。所以我觉得,代表所有人,谢谢你们做这些。
you know? Uh I mean obviously you guys have been there as well. Anyway, uh you guys got to go. Uh it's great to catch up. Uh, congrats on everything and also like I think there's just a lot of more work to do. Jordan gave us a sneak peek at the the kind of things you're doing. Uh, endpoint X sounds interesting. Uh, and like the the market is really segmenting, right? Like to me that's that's what you guys are identified early and are raising the standard for everyone. So I think for you know on behalf of everybody like thank you for doing this.
谢谢你,兄弟。
Thank you man.
谢谢。
Thank you.
好的。很高兴跟你们聊。再见。
Yeah. Good to talk to you guys. Bye.