Will Nvidia’s moat persist?
打开互动全文版(中英对照 + 朗读 + 问答)→算力之王谈 GPU、AI 基建,以及物理 AI。
The compute kingpin on GPUs, the AI buildout, and physical AI.
我们看到许多软件公司的估值暴跌,因为人们预期人工智能会商品化软件。有一种可能天真的想法是:你看,英伟达把 GDS2 文件发给台积电。台积电制造逻辑芯片,制造开关。然后把它和 SK 海力士、美光、三星制造的 HBM 封装在一起。再发给台湾的 ODM 组装成机架。所以英伟达本质上是在制造由他人生产的软件。如果软件被商品化,英伟达也会被商品化吗?
We've seen the valuations of a bunch of software companies crash because people are expecting AI to commoditize software. And there's a potentially naive way of thinking about things which is like: look, Nvidia sends a GDS2 file to TSMC. TSMC builds the logic dies. It builds the switches. Then it packages them with the HBM that SK Hynix, Micron, and Samsung make. Then it sends it to an ODM in Taiwan where they assemble the racks. And so Nvidia is fundamentally making software that other people are manufacturing. And if software gets commoditized, does Nvidia get commoditized?
最终,必须有人将电子转化为 token。这种转化——将电子转化为 token,并让这些 token 随时间变得更有价值——我认为很难完全商品化。从电子到 token 的转化是一段不可思议的旅程。制造 token 就像让一个分子比另一个分子更有价值,让一个 token 比另一个更有价值。让 token 变得有价值所需的艺术、工程、科学和发明——我们显然正在实时目睹这一切。所以这种转化、制造、以及其中的所有科学远未被深刻理解,旅程也远未结束。因此我怀疑它会被商品化。当然,我们会让它更高效。英伟达的全部——事实上,你提问的方式正是我对我们公司的思维模型:输入是电子,输出是 token。中间是英伟达,我们的职责是尽可能少做,但必须做足够多,以惊人的能力实现这种转化。所谓「尽可能少做」,意思是我不需要做的部分,我就与别人合作,让它成为我生态系统的一部分。看看今天的英伟达,我们可能拥有最大的合作伙伴生态系统,包括供应链上游和下游。所有计算机公司、所有应用开发者、所有模型制造者——人工智能就像五层蛋糕,我们在全部五层都有生态系统。所以我们尽量少做,但事实证明,我们必须做的那部分极其困难。我不认为那会被商品化。
Well, in the end, something has to transform electrons to tokens. That transformation — the transformation of electrons to tokens and making those tokens more valuable over time — I think that it's hard to completely commoditize. The transformation from electrons to tokens is such an incredible journey. Making that token — it's like making one molecule more valuable than another molecule, making one token more valuable than another. The amount of artistry, engineering, science, invention that goes into making that token valuable — obviously we're watching it happen in real time. And so the transformation, the manufacturing, all the science that goes in there is far from deeply understood and the journey is far from over. So I doubt that it will happen. We're going to make it more efficient, of course. The whole thing about Nvidia — in fact, the way you frame the question is my mental model of our company: the input is electron, the output is tokens. In the middle is Nvidia, and our job is to do as much as necessary, as little as possible, to enable that transformation to be done at incredible capabilities. And by "as little as possible," I mean whatever I don't need to do, I partner with somebody and make it part of my ecosystem. If you look at Nvidia today, we probably have the largest ecosystem of partners, both in supply chain upstream and downstream. All the computer companies, all the application developers, all the model makers — AI is a five-layer cake, if you will, and we have ecosystems across the entire five layers. So we try to do as little as possible, but the part that we have to do, as it turns out, is insanely hard. I don't think that gets commoditized.
我也不认为企业软件公司、工具制造商——如今大多数软件公司都是工具制造商,有些不是,但有些是工作流编码系统。但对很多公司来说,它们是工具制造商。例如,Excel 是工具,PowerPoint 是工具,Cadence 制造工具,Synopsys 制造工具。我实际上看到了与人们所见相反的情况。我认为智能体的数量将呈指数级增长。工具用户的数量将呈指数级增长,所有这些工具的实例数量很可能会飙升。Synopsys 设计编译器的实例数量很可能会飙升,使用平面规划器、所有布局工具和设计规则检查器的智能体数量也会飙升。如今智能体的数量受限于工程师的数量。明天,这些工程师将得到一群智能体的支持。我们将以前所未有的方式探索设计空间,并希望使用我们今天使用的工具。所以我认为工具的使用将导致这些软件公司飙升。这还没有发生的原因是智能体在使用工具方面还不够好。所以要么这些公司自己构建智能体,要么智能体变得足够好以使用这些工具。我认为两者都会发生。
I also don't think that the enterprise software companies, the tools makers — most of the software companies today are tools makers, some of them are not, but some are workflow codification systems. But for a lot of companies, they are tool makers. For example, Excel is a tool, PowerPoint is a tool, Cadence makes tools, Synopsys makes tools. I actually see the opposite of what people see. I think the number of agents is going to grow exponentially. The number of tool users is going to grow exponentially, and it's very likely that the number of instances of all these tools are going to skyrocket. It is very likely the number of instances of Synopsys design compiler is going to skyrocket, and the number of agents that are going to be using the floor planners and all of our layout tools and design rule checkers. The number of agents today is limited by the number of engineers. Tomorrow, those engineers are going to be supported by a bunch of agents. We're going to be exploring the design space like you've never seen before and want to use the tools that we use today. So I think tool use is going to cause these software companies to skyrocket. The reason why it hasn't happened yet is because the agents aren't good enough at using their tools yet. So either these companies are going to build the agents themselves, or agents are going to get good enough to be able to use those tools. I think it's going to be a combination of both.
在你们最新的文件中,与代工厂、内存、封装相关的采购承诺接近一千亿美元。而 SemiAnalysis 报道称,你们将有 2500 亿美元的这类采购承诺。所以一种解读是,英伟达的护城河在于你们锁定了这些稀缺组件多年的供应。别人可能有加速器,但他们真能拿到内存来制造吗?他们真能拿到逻辑芯片来制造吗?这确实是英伟达未来几年的大护城河。
In your latest filings, you had almost a hundred billion dollars in purchase commitments with foundries, memory, packaging. And then SemiAnalysis has reported that you will have $250 billion of these kinds of purchase commitments. So one interpretation is Nvidia's moat is really that you've locked up many years of these scarce components. Somebody else might have an accelerator, but can they actually get the memory to build it? Can they actually get the logic to build it? And this is really Nvidia's big moat for the next few years.
嗯,这是我们能做而别人难以做到的事情之一。我们之所以能做到——是因为我们在上游做出了巨大的承诺。有些是显性的,比如你提到的这些承诺;有些是隐性的。例如,上游的很多投资是由我们的供应链做出的,因为我对 CEO 们说:「让我告诉你这个行业会有多大,让我解释为什么,让我和你一起推理,让我展示我所看到的。」通过这个告知、激励、与各行业上游 CEO 对齐的过程,他们愿意进行投资。那么,为什么他们愿意为我投资而不是为别人?原因在于他们知道我有能力购买他们的供应并通过我的下游销售。英伟达的下游供应链和下游需求如此之大,他们愿意在上游投资。所以如果你看 GTC,人们惊叹于 GTC 的规模和参会人数。这是一个 360 度的视角——整个人工智能宇宙汇聚一处。他们聚在一起是因为需要互相看见。我把他们聚在一起,让下游能看到上游,上游能看到下游,所有人都能看到人工智能的所有进展。非常重要的是,他们都能见到人工智能原住民、正在构建的人工智能初创公司以及所有正在发生的奇妙事情,这样他们就能亲眼看到我告诉他们的一切。所以我花了很多时间直接或间接地告知我们的供应链、合作伙伴和生态系统面前的机会。我的大多数主题演讲——有些人总是说:「Jensen,在大多数主题演讲中,就像是一个接一个的公告。」
Well, it's one of the things that we can do that is hard for someone else to do. The reason why we could — we've made enormous commitments upstream. Some of it is explicit, these commitments that you mentioned; some of it is implicit. For example, a lot of the investments that are upstream are made by our supply chain because I said to the CEOs, "Let me tell you how big this industry is going to be and let me explain to you why and let me reason through it with you and let me show you what I see." And as a result of that process of informing, inspiring, aligning with CEOs of all different industries upstream, they're willing to make the investments. Now, why are they willing to make the investments for me and not someone else? The reason is because they know that I have the capacity to buy their supply and sell it through my downstream. The fact that Nvidia's downstream supply chain and our downstream demand is so large, they're willing to make the investment upstream. So if you look at GTC, people are marveled by the scale of GTC and the people that go. It's a 360° view — the entire universe of AI all in one place. And they're all in one place because they need to see each other. I bring them together so that the downstream could see the upstream, the upstream could see the downstream, and all of them could see all the advances in AI. And very importantly, they can all meet the AI natives and all the AI startups that are being built and all the amazing things that are happening, so that they could see firsthand all the things that I tell them. So I spend a lot of my time informing directly or indirectly our supply chain and our partners and our ecosystem about the opportunity that's in front of us. Most of my keynotes — some people always say, "Jensen, in most keynotes, it's like one announcement after another after another."
我们的主题演讲总是有点折磨人,因为它看起来像教育。事实上,这正是我的想法。我需要确保整个供应链,上游和下游,生态系统都理解什么在向我们袭来,为什么它会发生,何时发生,规模有多大,并且能够像我一样系统地推理。所以我认为,正如你描述的,我们能够为未来构建。如果未来几年是万亿美元的规模,我们有供应链来实现。没有我们的影响力,没有我们业务的流速,就像有现金流一样,也有供应链流。如果业务周转率低,没有人会为一个架构构建供应链。我们维持规模的能力仅仅是因为下游需求如此巨大,他们看到了,都听说了。他们看到一切正在到来。这让我们能够以我们能够做到的规模去做我们能做的事情。
Our keynotes are always a bit torturous in the sense that they come across like education. And in fact, that's exactly on my mind. I need to make sure that the entire supply chain, upstream and downstream, the ecosystem understands what is coming at us, why it's coming, when it's coming, how big it's going to be, and be able to reason about it systematically just like I reason about it. So I think the mode, as you describe it, we're able to build for a future. If our next several years is a trillion dollars in scale, we have the supply chain to do it. Without our reach, the velocity of our business, just as there's cash flow, there's supply chain flow. Nobody's going to build a supply chain for an architecture if the business turns are low. So our ability to sustain the scale is only because our downstream demand is so great, and they see it and they all hear about it. They see it all coming. And that allows us to do the things that we're able to do at the scale we're able to do.
我确实想更具体地了解上游能否跟上。多年来,你们一直实现收入同比翻倍。你们每年向世界提供的算力翻了三倍多。
I do want to understand more concretely whether the upstream can keep up. For many years now you guys have been doubling revenue year-over-year. You guys have been more than tripling the amount of flops you're providing to the world year over year.
在现在的规模下翻倍确实令人难以置信。
And doubling at the scale now is really incredible.
没错。那么看看逻辑芯片,你是台积电 N3 节点最大的客户,也是整个 AI 领域最大的客户之一。根据一些分析,今年将占 N3 的 60%,明年将达到 86%。如果你已经是多数,你怎么翻倍?而且每年都这样?那么我们现在是否处于一个由于上游原因 AI 算力增长率必须放缓的体制?你看到绕过这些的方法了吗?最终我们如何每年建造两倍的晶圆厂?
Exactly. So then you look at logic, say you're the biggest customer on TSMC's N3 node, and you're one of the biggest on AI as a whole. This year is going to be 60% of N3. It's going to be 86% next year according to some analysis. How do you double if you're the majority? And how do you do that year-over-year? So are we in a regime now where the growth rate in AI compute has to slow because of upstream? Do you see a way to get around these? How do we build twice as many fabs year-over-year ultimately?
是的,在某种程度上,瞬时需求大于全球上下游的供应。在任何时刻,我们可能受到水管工数量的限制。
Yeah, at some level, the instantaneous demand is greater than the supply upstream and downstream in the world. And at any instant, we could be limited by the number of plumbers.
嗯。
Mhm.
这确实会发生。
Which actually happens.
水管工被邀请参加明年的 GTC。
The plumbers are invited to next year's GTC.
是的。顺便说一句,好主意。但这是一个好的状况。你想要一个市场,一个行业,其中瞬时需求大于行业的总供应。相反的情况显然不太好。如果我们差距太大,如果某个特定项目、某个特定组件太远,显然行业会蜂拥而至。例如,注意人们不再怎么谈论 co-ass 了。
Yeah. By the way, great idea. But that's a good condition. You want a market, you want an industry where the instantaneous demand is greater than the total supply of the industry. The opposite is obviously less good. If we're too far apart, if one particular item, one particular component is too far away, obviously the industry swarms it. So for example, notice people aren't talking very much about co-ass anymore.
是的。
Yeah.
原因是在过去的两年里,我们全力以赴地攻克它,我们翻倍、翻倍、再翻倍,现在我认为我们处于相当好的状态。台积电现在知道 co-ass 供应必须跟上其他逻辑需求和内存需求,所以他们正在扩大 co-ass 规模,并以与逻辑芯片相同的水平扩展未来的封装技术,这太棒了,因为长期以来 co-ass 相当特殊,HBM 也相当特殊,但它们不再是特殊产品了。人们现在意识到它们是主流计算技术。当然,我们现在更有能力影响更大范围的供应链。过去,在 AI 革命初期,我现在说的所有事情五年前我就说过,有些人相信并投资了。例如,Sanjay 和美光团队。我仍然清楚地记得那次会议,我清楚地说明了将要发生什么以及为什么发生,以及今天的预测,他们真的全力以赴,我们与他们合作,涉及 LPDDR、HBM 内存。他们真的投资了,这对公司来说显然是巨大的。有些人来得晚一些,但他们现在都在这里了。所以我认为每一个瓶颈都得到了大量关注,现在我们提前几年预取瓶颈。例如,过去几年我们与 Lum 和 Coherent 以及整个硅光子学生态系统的投资,我们真正重塑了硅光子学的生态系统和供应链。我们围绕台积电建立了一整套供应链。我们与他们合作开发了 coupe,发明了大量技术。我们将这些专利授权给供应链,保持开放。所以我们通过发明新技术、新工作流程、新测试设备、双面探测、投资公司、帮助他们扩大产能来准备供应链。所以你可以看到我们正在努力塑造生态系统,使其准备好,供应链准备好支持规模。似乎有些瓶颈比其他更容易。
And the reason for that is because for two years we swarmed the living daylights out of it and we doubled, doubled, doubled on several doubles, and now I think we're in fairly good shape. And TSMC now knows that co-ass supply has to keep up with the rest of the logic demand and the memory demand, so they're scaling co-ass and their scaling future packaging technologies at the same level as scale logic, which is terrific because for a long time co-ass was rather specialty and HBM was rather specialty, but they're not specialties anymore. People now realize they're mainstream computing technology. And of course, we're now much more able to influence a larger scope of our supply chain. In the past, in the beginning of the AI revolution, all the things that I say now I was saying five years ago, and some people believed in it and invested in it. For example, Sanjay and the Micron team. I still remember the meeting really well where I was clear about exactly what's going to happen and why it's going to happen, and the predictions of today, and they really doubled down on it, and we partnered with them across LPDDR, across HBM memories. They really invested in it, and it obviously has been tremendous for the company. Some people came a little bit later, but they're all here now. So I think each one of these bottlenecks gets a great deal of attention, and now we're prefetching the bottlenecks years in advance. For example, the investments that we've done with Lum and Coherent and all of the silicon photonics ecosystem in the last several years, we really reshaped the ecosystem and the supply chain of silicon photonics. We built up an entire supply chain around TSMC. We partnered with them on coupe, invented a whole bunch of technology. We licensed those patents to the supply chain, kept it nice and open. And so we're preparing the supply chain through invention of new technologies, new workflows, new test equipment, double-sided probing, investing in companies, helping them scale up their capacity. So you could see that we're trying to shape the ecosystem so that it's ready, the supply chain so that it's ready to support the scale. It seems like some bottlenecks are easier than others.
顺便说一句,我去了最难的那个。
I went to the hardest one by the way.
哪个?
Which is?
水管工。
Plumbers.
是的,没错。我确实去了最难的那个。是的。
Yeah, it's true. Yeah. I actually went to the hardest one. Yeah.
是的。水管工和电工。原因是这是我对所有描述工作终结和就业消失的末日论者的担忧之一。如果我们劝阻人们成为软件工程师,我们就会缺少软件工程师。十年前同样的预测,一些末日论者说:「无论你做什么,都不要成为放射科医生。」你可能还会在网上看到一些这样的视频。放射学将是第一个消失的职业。不再需要放射科医生了。猜怎么着?我们缺少放射科医生。
Yeah. Plumbers and electricians. And the reason for that is because this is one of the concerns that I have about all the doomers describing the end of work and killing of jobs. One of the things that if we discourage people from being software engineers, we're going to run out of software engineers. And the same prediction ten years ago, some of the doomers were saying, 'Whatever you do, don't be a radiologist.' And you might hear some of those videos are still on the web. Radiology is going to be the first career to go. Nobody's going to need any more radiologists. Guess what? We're short of radiologists.
哦,但好吧。回到这一点,有些东西你可以扩大规模,其他东西比如你如何每年制造两倍的逻辑芯片?最终受限于内存和逻辑,受限于 UV 光刻机。你如何每年获得两倍的 UV 光刻机?年复一年。
Oh, but okay. So going back to this point about well some things you scale, other things like how do you actually manufacture twice the amount of logic a year? Ultimately that's bottlenecked by memory and logic, bottlenecked by UV. How do you get to twice as many UV machines a year? Year over year.
这些都不是不可能快速扩大规模的。你只需要……你可以在两三年内完成所有这些。
None of that is impossible to scale quickly. You just need to... you could do all of that within two or three years.
你只需要一个需求信号。一旦你能造出一个,你就能造出十个;一旦你能造出十个,你就能造出一百万个。所以这些东西并不难复制。你要沿着供应链走多远?你会去找 ASML 说,「嘿,如果我展望三年后,英伟达要年收入达到两万亿,我们需要更多的 EUV 光刻机」?有些是我直接去谈,有些是间接的。如果我能说服台积电,ASML 也会被说服。所以我们必须考虑关键的瓶颈点。但如果台积电被说服了,几年内你就会有足够的 EUV 光刻机。没有一个瓶颈会持续超过两三年。一个都没有。与此同时,我们正在将计算效率提升 10 倍、20 倍——从 Hopper 到 Blackwell,甚至提升了 30 到 50 倍。因为 CUDA 非常灵活,我们不断提出新算法。除了增加产能,我们还在开发各种新技术来提高效率。这些都不让我担心。
You just need a demand signal. Once you can build one, you can build 10, and once you can build 10, you can build a million. So these things are not hard to replicate. How far down the supply chain do you go? Do you go to ASML and say, 'Hey, if I look out three years from now, for Nvidia to be generating two trillion in a year in revenue, we need way more EUV machines'? Some of them I have to directly, some of them indirectly. And if I can convince TSMC, ASML will be convinced. So we have to think about the critical pinch points. But if TSMC is convinced, you'll have plenty of EUV machines in a few years. None of the bottlenecks last longer than two or three years. None of them. Meanwhile, we're improving computing efficiency by 10x, 20x—in the case of Hopper to Blackwell, some 30x to 50x. We're coming up with new algorithms because CUDA is so flexible. We're developing all kinds of new techniques to drive efficiency in addition to increasing capacity. None of that worries me.
问题在于我们下游的东西。能源政策阻碍了能源……没有能源你无法创造产业。没有能源你无法创造全新的制造业。我们希望美国再工业化。我们希望恢复芯片制造、计算机制造、封装,我们希望建造电动汽车和机器人等新东西,我们希望建造 AI 工厂。没有能源你什么都建不了,而这些事情需要很长时间。但更多的芯片产能?那是两三年内能解决的问题。更多的算力?两三年内能解决的问题。
It's the stuff downstream from us. Energy policies that prevent energy from... you can't create an industry without energy. You can't create a whole new manufacturing industry without energy. We want to re-industrialize the United States. We want to bring back chip manufacturing, computer manufacturing, packaging, and we want to build new things like EVs and robots, and we want to build AI factories. You can't build any of these things without energy, and those things take a long time. But more chip capacity? That's a two- or three-year problem. More compute capacity? Two- or three-year problem.
有意思。有时候我的嘉宾会告诉我完全相反的事情,而这次我没有足够的技术知识来判断。
Interesting. I feel like I have guests tell me the exact opposite thing sometimes, and in this case I just don't have the technical knowledge to adjudicate.
嗯,美妙之处在于你正在和专家对话。
Well, the beautiful thing is you're talking to the expert.
好的。我想问问关于你的竞争对手。如果你看看 TPU,可以说世界上排名前三的模型中有两个——Claude 和 Gemini——是在 TPU 上训练的。这对英伟达的未来意味着什么?
Okay. I want to ask about your competitors. So, if you look at TPU, arguably two out of the top three models in the world, Claude and Gemini, were trained on TPU. What does that mean for Nvidia going forward?
嗯,我们建造了非常不同的东西。英伟达建造的是加速计算,而不是张量处理单元。加速计算用于各种事情:分子动力学、量子色动力学、数据处理、数据框、结构化数据、非结构化数据、流体动力学、粒子物理,此外我们还用它做 AI。加速计算要多样化得多。虽然 AI 是当今的话题,显然非常重要且影响深远,但计算远不止于此。英伟达所做的是重新发明计算的方式,从通用计算转向加速计算。我们的市场覆盖范围远超任何 TPU 或 ASIC 可能达到的。我们是唯一一家加速各种应用的公司。我们有一个巨大的生态系统,所以各种框架和算法都在英伟达上运行。因为我们的计算机设计成可由他人操作,任何操作员都可以购买我们的系统。大多数自建系统要求你自己当操作员,因为它们从未设计得足够灵活让别人操作。结果,我们存在于每一个云中,包括谷歌、亚马逊、Azure 和 OCI。无论你是想运营来出租还是自用,我们都有能力帮助你。例如,对于埃隆的 xAI,我们可以让任何公司或行业的操作员为礼来公司的科学研究和药物发现建造超级计算机。我们可以帮助他们运营自己的超级计算机,用于药物发现和生物科学的所有多样性。我们可以解决一大堆 TPU 无法解决的应用,因为英伟达也将 CUDA 构建成了一个出色的张量处理单元,但它处理数据、计算和 AI 的每一个生命周期。我们的市场机会要大得多。我们的覆盖范围要大得多。因为我们有如此庞大的生态系统,你可以在任何地方建造英伟达系统,并且知道会有客户。这是非常不同的。
Well, we built a very different thing. What Nvidia built is accelerated computing, not a tensor processing unit. Accelerated computing is used for all kinds of things: molecular dynamics, quantum chromodynamics, data processing, data frames, structured data, unstructured data, fluid dynamics, particle physics, and in addition, we use it for AI. Accelerated computing is much more diverse. Although AI is the conversation today and is obviously very important and impactful, computing is much broader than that. What Nvidia has done is reinvent the way computing is done, from general-purpose computing to accelerated computing. Our market reach is far greater than any TPU or ASIC can possibly have. We're the only company that accelerates applications of all kinds. We have a gigantic ecosystem, so all kinds of frameworks and algorithms run on Nvidia. Because our computers are designed to be operated by other people, anyone who's an operator can buy our systems. Most of these homebuilt systems require you to be your own operator because they were never designed to be flexible enough for others. As a result, we're in every cloud, including Google, Amazon, Azure, and OCI. Whether you want to operate it to rent or for yourself, we have the ability to help you. For example, for Elon with xAI, we can enable operators in any company or industry to build a supercomputer for scientific research and drug discovery at Lilly. We can help them operate their own supercomputer for the entire diversity of drug discovery and biological sciences. There are a whole bunch of applications we can address that you can't do with TPUs because Nvidia built CUDA as a fantastic tensor processing unit as well, but it does every lifecycle of data processing, computing, and AI. Our market opportunity is just a lot larger. Our reach is a lot greater. Because we have such a large ecosystem, you can build Nvidia systems anywhere and know there will be customers for it. It's a very different thing.
这大概会是个长问题,但你有惊人的收入,而且这些收入大多不是来自制药或量子计算;你之所以能赚到钱,是因为 AI 是一项前所未有的技术,增长速度快得前所未有。所以问题是:什么对 AI 最有利?我不了解细节,但我跟我的 AI 研究员朋友聊过,他们说当我使用 TPU 时,它是一个巨大的脉动阵列,非常适合矩阵乘法,而 GPU 非常灵活,适合分支和不规则内存访问。但 AI 就是这些非常可预测的矩阵乘法,一遍又一遍。你不需要为 warp 调度器或线程与内存库之间的切换牺牲任何芯片面积。所以 TPU 真正针对的是目前即将上线的计算收入和使用案例的主体部分。你怎么看?
This is going to be sort of a long question, but you have spectacular revenue, and this revenue is mostly not from pharma or quantum; you're making it because AI is an unprecedented technology growing unprecedentedly fast. So the question is: what is best for AI specifically? I'm not in the details, but I talked to my AI researcher friends, and they say when I use a TPU, it's this big systolic array perfect for matrix multiplies, whereas a GPU is very flexible, great for branching and irregular memory access. But AI is just these very predictable matrix multiplies again and again. You don't have to give up any die area for warp schedulers or switches between threads and memory banks. So the TPU is really optimized for the majority of the bulk of this growth in revenue and use case for compute coming online right now. How do you react to that?
矩阵乘法是 AI 的重要组成部分,但不是唯一的部分。如果你想提出新的注意力机制,或者想以不同的方式解耦,或者想提出一种全新的架构——例如混合 SSM,或者某种融合扩散和自回归的模型——你需要一个普遍可编程的架构。我们能运行你能想象的一切。这就是优势。它使得新算法的发明容易得多。
Matrix multiplies are an important part of AI, but it's not the only part. If you want to come up with a new attention mechanism, or if you want to disaggregate in a different way, or if you want to come up with a whole new type of architecture altogether—for example, a hybrid SSM, or a model that fuses diffusion and autoregressive somehow—you want an architecture that's just generally programmable. We run everything you can imagine. That's the advantage. It allows for invention of new algorithms a lot more easily.
正因为这是一个可编程的系统,发明新算法的能力才是推动 AI 快速进步的关键。TPU 和其他技术一样受摩尔定律影响。我们知道摩尔定律每年大约提升 25%。所以真正实现 10 倍、100 倍飞跃的唯一方法,就是每年从根本上改变算法及其计算方式。
And so because it's a programmable system, the ability to invent new algorithms is really what makes AI advance so quickly. TPUs like anything else is impacted by Moore's law. And we know that Moore's law is increasing about 25% per year. And so the only way to really get 10x leaps, 100x leaps, is to fundamentally change the algorithm and how it's computed every single year.
这正是 Nvidia 的根本优势。我们之所以能让 Blackwell 比 Hopper 快 50 倍——我之前说是 35 倍,当我第一次宣布 Blackwell 能效比 Hopper 高 35 倍时,没人相信。后来 Dylan 写了篇文章,说我其实保守了,实际上是 50 倍。仅靠摩尔定律不可能合理实现这一点。我们解决这个问题的方法是新的模型,并行化、解耦并分布在整个计算系统中。如果没有能力深入并用 CUDA 开发新内核,这几乎不可能做到。我们架构的可编程性、Nvidia 作为极端协同设计公司甚至能将部分计算卸载到网络本身(例如 NVLink、Spectrum X),以及我们能够同时影响处理器、系统、网络、库和算法——所有这些都同步进行。没有 CUDA,我甚至不知从何入手。
And that's Nvidia's fundamental advantage. The only reason why we were able to make Blackwell to Hopper 50 times — I said it was 35 times, and when I first announced that Blackwell was going to be 35 times more energy efficient than Hopper, nobody believed it. And then Dylan wrote an article. He said in fact I sandbagged it; it's actually 50 times. And you can't reasonably do that with just Moore's law. The way we solve that problem is new models, parallelized, disaggregated, and distributed across a computing system. Without the ability to really get down and come up with new kernels with CUDA, it's really hard to do. The combination of the programmability of our architecture, the fact that Nvidia is an extreme codesign company where we could even offload some of the computation into the fabric itself — NVLink for example, into the network Spectrum X — and that we could affect change across the processors, the system, the fabric, the libraries, the algorithm. All of that was done simultaneously. Without CUDA to do that, I wouldn't even know where to start.
我的赞助商 Crusoe 是最早提供 Nvidia Blackwell 和 Blackwell Ultra 平台的云服务商之一,他们刚刚宣布了今年晚些时候部署 Nvidia Vera Rubin 的计划。但获得最先进的硬件只是故事的一部分。例如,大多数推理引擎已经为单个用户的前向传播做了 KV 缓存,但 Crusoe 跨用户和 GPU 进行缓存。所以如果一千个智能体运行在相同的系统提示上,Crusoe 只需计算一次 KV 缓存,集群中的每个 GPU 都能使用。随着系统变得更加智能体化,需要更长的前缀来使用工具和访问文件,这一点尤为重要。在最近的基准测试中,Crusoe 的首词生成时间比 VLM 快 10 倍,吞吐量高 5 倍。这只是你应该在 Crusoe 上运行推理工作负载的众多原因之一。如果你需要 GPU 进行训练,也无需切换云服务商。Crusoe 同样能满足你的需求。访问 crusoe.ai/torcashe 了解更多。
My sponsor Crusoe was among the first clouds to offer Nvidia's Blackwell and Blackwell Ultra platforms, and they just announced their Nvidia Vera Rubin deployment scheduled for later this year. But access to state-of-the-art hardware is only part of the story. For example, most inference engines already do KV caching for a single user's forward passes, but Crusoe does it across users and GPUs. So if a thousand agents are running on the same system prompt, Crusoe only has to compute the KV cache once for it to become available to every single GPU in the cluster. This is especially important as systems get more agentic and require much longer prefixes in order to use tools and access files. In a recent benchmark, Crusoe was able to deliver up to 10 times faster time to first token and up to five times better throughput than VLM. This is just one among many reasons that you should run your inference workload with Crusoe. And if you need GPUs for training, you don't need to switch clouds. Crusoe's got you covered there, too. Go to cruso.ai/torcashe to learn more.
这就引出了一个关于 Nvidia 客户群的有趣问题。如果 60%的收入来自五大超大规模云服务商,而在另一个时代,客户是教授做实验,他们需要 CUDA,不能用其他加速器,他们只需要运行 PyTorch 和 CUDA,一切优化好。但面对这些超大规模云服务商,他们有资源编写自己的内核。事实上,他们必须这样做,才能获得针对其特定架构所需的额外 5%性能。Anthropic、Google 主要运行自己的加速器或 TPU 和 Trainium,但即使是使用 GPU 的 OpenAI 也有 Triton,他们表示需要自己的内核。所以他们深入 CUDA C++,而不是使用 cuBLAS 和 cuDNN 等,他们有自己的栈,也能编译到其他加速器。那么,如果大多数客户能够并且确实在替换 CUDA,那么 CUDA 在多大程度上真正是让前沿 AI 在 Nvidia 上实现的关键?
So, this gets at an interesting question about Nvidia's clientele. If 60% of your revenue is coming from these big five hyperscalers, in a different era where customers were professors running experiments, they needed CUDA. They couldn't use another accelerator. They needed to just run PyTorch with CUDA and have everything optimized. But if you've got these hyperscalers, they have the resources to write their own kernels. In fact, they have to, to get that extra last 5% that they need for their specific architecture. Anthropic, Google are mostly running their own accelerators or running TPUs and Trainium, but even OpenAI using GPUs has Triton, which they're like we need our own kernels. So they've gone down to CUDA C++ instead of using cuBLAS and cuDNN and everything; they've got their own stack which compiles to other accelerators as well. And so if most of your customers can and do make replacements for CUDA, to what extent is CUDA really the thing that is going to make Frontier AI happen on Nvidia?
CUDA 是一个丰富的生态系统。如果你想在任何计算机上优先构建,首先基于 CUDA 构建是非常明智的,因为生态系统如此丰富。我们支持每一个框架。如果你想创建自定义内核,例如,我们为 Triton 贡献巨大,Triton 的后端包含了大量 NVIDIA 技术。我们乐于帮助每个框架发挥最大潜力。框架非常多:Triton、VLM、SG lang 等等。现在还有一大批新的强化学习框架涌现:Verl、NeMo RL 等等。随着后训练和强化学习的发展,整个领域正在爆发。所以如果你想基于某个架构构建,基于 CUDA 是最合理的,因为你知道生态系统很棒。你知道如果出问题,更可能出在你的代码里,而不是底层的大量代码中。别忘了构建这些系统时你处理的代码量。当某些东西不工作时,是你还是计算机的问题?你希望总是你的问题,并且能够信任计算机。显然我们自己也有很多 bug,但我们的系统经过充分验证,你至少可以在其基础上构建。所以第一点是生态系统的丰富性、可编程性和能力。第二点是,如果你是一名开发者,无论构建什么,你最想要的就是安装基数。你希望你的软件能在大量其他计算机上运行。你不是只为自己构建软件。你为你的集群或所有人的集群构建软件,因为你是框架构建者。Nvidia 的 CUDA 生态系统最终是巨大的财富。我们现在有数亿个 GPU。每个云都有,从 A10、A100、H100、H200 到 L 系列、P 系列,各种尺寸和形状。如果你是机器人公司,你希望 CUDA 栈能在机器人本身中运行。我们无处不在。安装基数意味着一旦你开发了软件和模型,它在任何地方都有用。安装基数极其宝贵。最后,我们存在于每一个云中,这使我们真正独一无二,因为作为 AI 公司和 AI 开发者,你不确定与哪个 CSP 合作以及在哪里运行。我们在任何地方运行,包括本地部署,如果你愿意的话。
CUDA is a rich ecosystem. If you want to build on any computer first, building on CUDA first is incredibly smart because the ecosystem is so rich. We support every framework. If you want to create custom kernels, for example, we contribute enormously to Triton, and the back end of Triton has huge amounts of NVIDIA technology. We're delighted to help every framework become as great as it can be. There are lots and lots of frameworks: Triton, VLM, SG lang, and more. Now there's a whole bunch of new reinforcement learning frameworks coming out: you've got Verl, NeMo RL, a whole bunch of new ones. And now with post-training and reinforcement learning, that entire area is just exploding. So if you want to build on an architecture, building on CUDA makes the most sense because you know that the ecosystem is great. You know that if something happens, it's more likely in your code and not in the mountain of code underneath. Don't forget the amount of code you're dealing with when building these systems. When something doesn't work, was it you or was it the computer? You would like it always to be you and to be able to trust the computer. Obviously we still have lots of bugs ourselves, but our system is so well wrung out that you can at least build on top of the foundation. So that's number one: the richness of the ecosystem, the programmability of it, the capability of it. The second thing is, if you were a developer building anything at all, the single most important thing you want more than anything is install base. You want the software that you run to run on a whole bunch of other computers. You don't want to build software just for yourself. You're building software for your fleet or for everybody else's fleet because you're a framework builder. Nvidia's CUDA ecosystem is ultimately its great treasure. We are now, I don't know how many, several hundred million GPUs. Every cloud has it, going back to A10, A100, H100, H200, the L series, the P series. There's a whole bunch of them in all kinds of sizes and shapes. If you're a robotics company, you want that CUDA stack to actually run in the robot itself. We're literally everywhere. The install base says that once you develop the software, once you develop the model, it's going to be useful everywhere. The install base is just too incredibly valuable. And lastly, the fact that we're in every single cloud makes us genuinely unique because you're an AI company and an AI developer. You're not exactly sure which CSP you're going to partner with and where you would like to run it. We run it everywhere, including on prem for you if you like.
所以我认为生态系统的丰富性、安装基础的广泛性以及我们当前所处位置的多样性,这种组合使得 CUDA 变得不可或缺。
And so I think that the richness of the ecosystem, the expansiveness of the install base, and the versatility of where we are, that combination makes CUDA invaluable.
这很有道理。我好奇的是,这些优势对你的主要客户是否真的很重要。有很多人可能在乎这些,比如那些能够自己构建软件栈的人,他们构成了你大部分的收入。尤其是如果我们进入一个 AI 在那些有严格验证循环的事情上变得特别擅长的世界,你可以对它们进行强化学习,那么如何编写一个在规模扩展中最有效地处理注意力或 MLP 的内核,这是一个非常可验证的反馈循环。那么,所有超大规模云厂商都能自己编写这些定制内核吗?他们可能仍然会,英伟达仍然有很好的性价比。所以他们可能仍然倾向于使用英伟达。但问题是,这是否就变成了一个谁在给定价格下提供最佳规格、最佳算力和内存带宽的问题?历史上,英伟达因为 CUDA 模式,在 AI 硬件和软件上拥有最好的利润率,超过 70%。问题是,如果你的大多数客户实际上能够负担得起构建而不是使用 CUDA 模式,你能否维持这些利润率?
That makes a lot of sense. I guess the thing I'm curious about is whether those advantages matter a lot to your main customers. There are many people who they might matter for, for the kind of person who can actually build their own software stack, who make up most of your revenue. Especially if you go to a world where AI is getting especially good at things which have tight verification loops where you can RL on them, and then this question of how do you write a kernel that does attention or MLP the most efficiently across a scale up, it's a very verifiable sort of feedback loop. So, can everybody, can all the hyperscalers write these custom kernels for themselves? And they might still, Nvidia still has great price performance. So they might still prefer to use Nvidia. But then the question is, does it just become a question of who is offering the best specs, the best flops and memory and memory bandwidth for a given dollar? Where historically Nvidia has just had and still has the best margins in all of AI across hardware and software, 70% plus, because of this CUDA mode. And the question is, can you sustain those margins if for most of your customers they can actually afford to build instead of the CUDA mode?
我们分配给这些 AI 实验室的工程师数量是惊人的。与他们合作,优化他们的软件栈。原因在于没有人比我们更了解我们的架构。而这些架构不像 CPU 那样通用。CPU 就像一辆凯迪拉克,是一辆舒适的巡航车。它永远不会开得太快。每个人都能很好地驾驶它。它有巡航控制,一切都很简单。但在很多方面,英伟达的 GPU 是加速器,有点像 F1 赛车。我可以想象每个人都能以 100 英里的时速驾驶它,但要把它推到极限需要相当多的专业知识。我们使用大量 AI 来创建我们拥有的内核。我相当确定在相当长一段时间内我们仍然会被需要。所以我们的专业知识帮助我们的 AI 实验室合作伙伴轻松地从他们的软件栈中获得额外 2 倍的性能。很多时候,当我们完成优化他们的软件栈或特定内核时,他们的模型速度提升了 3 倍、2 倍、50%。这是一个巨大的数字,尤其是当你考虑到他们拥有的所有 Hopper 和 Blackwell 的安装基础时。当你将其提升两倍时,收入就翻倍了。这直接转化为收入。英伟达的计算堆栈是世界上每总拥有成本性能最好的,没有例外。没有人能向我证明当今世界上有任何平台具有更好的性能 TCO 比。没有一家公司。事实上,基准测试就在那里。Dylan 的 Inference Max 就在那里供大家使用,没有一个 TPU 会来,Trillium 也不会来。我鼓励他们使用 Inference Max 来展示他们令人难以置信的推理成本。这真的很难。没有人愿意出现。MLPerf,我欢迎 Trillium 来展示他们一直声称的 40%。我很想听他们展示 TPU 的成本优势。在我看来这毫无意义。从基本原理上讲这绝对没有意义。这说不通。所以我认为我们如此成功的原因仅仅是因为我们的 TCO 如此出色。
The number of engineers we have assigned to these AI labs is insane. Working with them, optimizing their stack. And the reason for that is because nobody knows our architecture better than we do. And these architectures are not as general purpose as a CPU. The reason why a CPU is so, you know, a CPU is kind of like a Cadillac, it's a nice cruiser. It never goes too fast. Everybody drives it pretty well. It's got cruise control, and everything is easy. But in a lot of ways, Nvidia's GPUs are accelerators, kind of like F1 racers. And I could imagine everybody's able to drive it at 100 miles an hour, but it takes quite a bit of expertise to be able to push it to the limit. And we use a ton of AI to create the kernels that we have. And I'm pretty sure we're going to still be needed for quite some time. And so our expertise helps our AI labs partners get another 2x out of their stack easily. Often times it's not unusual that by the time we're done optimizing their stack or optimizing a particular kernel, their model sped up by 3x, 2x, 50%. That's a huge number, especially when you're talking about the installed base of the fleet that they have of all the Hoppers and Blackwells that they have. When you increase it by a factor of two, that doubles the revenues. That directly translates to revenues. Nvidia's computing stack is the best performance per TCO in the world, bar none. Nobody can demonstrate to me that any single platform in the world today has better performance TCO ratio. Not one company. And in fact, the benchmarks are out there. Dylan's Inference Max is sitting out there for everybody to use, and not one TPU won't come, Trillium won't come. I encourage them to use Inference Max and demonstrate their incredible inference cost. It's really really hard. Nobody wants to show up. MLPerf, I would welcome Trillium to demonstrate their 40% that they claim all the time. I would love to hear them demonstrate the cost advantage of TPUs. It makes no sense in my mind. It makes absolutely zero sense on first principles. It makes no sense. And so I think the reason why we're so successful is simply because our TCO is so great.
你说 60%的客户是前五大,但大部分业务是外部的。例如,AWS 上的大部分英伟达实例是给外部客户使用的,而不是内部使用。我们在 Azure 上的大部分客户,显然都是外部的。我们在 OCI 上的所有客户都是外部的,不是内部使用。他们青睐我们的原因是因为我们的覆盖范围如此之大。我们可以为他们带来世界上所有伟大的客户。他们都建立在英伟达之上。所有这些 AI 公司都建立在英伟达之上的原因是因为我们的覆盖范围和多样性如此之大。所以我认为飞轮效应在于安装基础、我们架构的可编程性、生态系统的丰富性,以及世界上有如此多的 AI 公司,现在有数万家。
You said 60% of our customers are the top five but most of that business is external. For example, most of AWS is most of Nvidia in AWS is for external customers not internal use. Most of our customers at Azure, obviously all of our customers are external. All of our customers at OCI are external, not internal use. The reason why they favor us is because our reach is so great. We can bring them all of the great customers in the world. They're all built on Nvidia. And the reason why all these AI companies are built on Nvidia is because our reach and our versatility is so great. And so I think the flywheel is really install base, the programmability of our architecture, the richness of our ecosystem, and the fact that there's so many AI companies in the world, there's tens of thousands of them now.
如果你是那些 AI 初创公司之一,你会选择什么架构?你会选择最丰富的架构,世界上最多的地方,拥有最大安装基础和丰富生态系统的架构。这就是飞轮效应。这就是为什么结合以下几点:第一,我们的每美元性能如此出色,以至于他们拥有最低成本的 token。第二,我们的每瓦性能是世界上最高的。所以如果这些公司中的一家,如果我们的合作伙伴建造了一个 1 吉瓦的数据中心,那个 1 吉瓦的数据中心最好能提供最大量的收入和 token 数量,这直接转化为收入。你想生成尽可能多的 token,最大化该数据中心的收入。我们拥有世界上每瓦 token 最高的架构。最后,如果你的目标是租用基础设施,我们在世界上拥有最多的客户。这就是飞轮效应起作用的原因。
And if you were one of those AI startups, what architecture would you choose? You would choose an architecture that's most abundant, where the most abundant in the world, the one has the largest installed base, and one that has a rich ecosystem. And so that's the flywheel. That's the reason why between the combination of one, our perf per dollar is so great that they have the lowest cost tokens. Second, our perf per watt is the highest in the world. And so if one of these companies, if our partners built a 1 gigawatt data center, that 1 gigawatt data center better deliver the maximum amount of revenues and number of tokens, which directly translates to revenues. You wanted to generate as many tokens as possible, maximize the revenues for that data center. We have the highest tokens per watt architecture in the world. And then lastly, if your goal is to rent the infrastructure, we have the most customers in the world. And so that's the reason why the flywheel works.
有趣。我想问题归结为这里的实际市场结构是什么。因为即使有其他公司,也可能存在一个世界,有数万家 AI 公司大致拥有相等的算力份额,但即使通过这些五大超大规模云厂商,真正在亚马逊上使用算力的人,Anthropic、OpenAI,以及这些大型基础模型实验室,他们自己能够负担并有能力让不同的加速器工作……
Interesting. I guess the question comes down to what is the actual market structure here. Because even if there's other companies, there could have been a world where there's tens of thousands of AI companies that have roughly equal share of compute, but if even through these five hyperscalers really the people on Amazon using the compute, Anthropic, OpenAI, and these big foundation labs who can themselves afford and have the ability to make different accelerators work...
不,我认为你的假设是错误的。
No, I think your assumption is wrong.
也许,让我问你一个稍微不同的问题,那就是……
Maybe, let me ask you a slightly different question, which is...
回来让我纠正你的前提。
Come back and make me correct your premise.
好吧,让我问一个不同的问题,那就是如果一切……
Okay, let me just ask a different question, which is okay if everything...
但还是要确保让我回来纠正,因为这对 AI 太重要了,对科学的未来太重要了,对行业的未来太重要了,那个前提……
But still make sure that make me come back and fix because it's just too important to AI, it's too important to the future of science, it's too important to the future of the industry that that premise...
前提,听着,让我先问完问题,然后我们可以一起讨论。
The premise, look, let me just finish the question and then we can address it together.
那么,如果所有这些关于性价比和每瓦性能等说法都是真的,你认为为什么像 Anthropic 这样的公司,几天前刚宣布与 Broadcom 和 Google 达成一项多吉瓦的 TPU 交易,而且他们的大部分算力显然来自 Google 的 TPU?所以,当我观察这些大型 AI 公司时,似乎曾经全是 Nvidia,但现在不是了。我很好奇,如果这些说法在纸面上成立,他们为什么还要选择其他加速器?
So what do you think if all these things are true about price performance and performance per watt, etc., why do you think it is the case that say Anthropic, for example, just announced a couple days ago they have a multi-gigawatt deal with Broadcom and Google for TPUs, and majority of their compute obviously for Google it's TPU majority compute? So if I look at these big AI companies, it seems like there was some point where it was all Nvidia and now it's not. And so I'm curious how to square if these things are true on paper, why are they going with other accelerators?
是的,Anthropic 是一个特例,不是趋势。没有 Anthropic,TPU 的增长从何而来?100% 是因为 Anthropic。没有 Anthropic,Tranium 的增长又从何而来?100% 是因为 Anthropic。我认为这一点众所周知,也很好理解。并不是说 ASIC 机会很多,只有一个 Anthropic。
Yeah, Anthropic is a unique instance and not a trend. Without Anthropic, why would there be any TPU growth at all? It's 100% Anthropic. Without Anthropic, why would there be any Tranium growth at all? It's 100% Anthropic. And I think that's fairly well known and well understood. It's not that there's an abundance of ASIC opportunities. There's only one Anthropic.
但 OpenAI 与 AMD 有合作,他们还在打造自己的 Titan 加速器。
But OpenAI deals with AMD. They're building their own Titan accelerator.
是的。但我们可以承认,他们主要还是用 Nvidia,我们仍然会一起做很多工作。
Yeah. But they're mostly, we could all acknowledge they're vastly Nvidia and we're going to still do a lot of work together.
是的。我们并不介意别人使用其他东西并尝试新事物。如果他们不尝试这些,怎么会知道我们的有多好?有时候你需要被提醒。我们必须不断赢得我们现在的地位。总有人声称要替代我们。看看有多少 ASIC 被取消了。就算你要造 ASIC,你也必须造出比 Nvidia 更好的东西。而造出比 Nvidia 更好的东西并不容易,实际上并不明智。Nvidia 肯定有什么过人之处。说真的,因为我们的规模、我们的速度,我们是世界上唯一一家每年都推出重大突破的公司。每年都有巨大飞跃。
Yeah. And we're not offended by other people using something else and trying things. If they don't try these other things, how would they know how good ours is? And sometimes you got to be reminded of it. And we have to continuously earn the position that we're in. There are always claims. Look at the number of ASICs that have been cancelled. Just because you're going to build an ASIC, you still have to build something better than Nvidia. And it's not that easy building something better than Nvidia. It's not sensible actually. Nvidia's got to be missing something. Seriously, because our scale, our velocity, we're the only company in the world that's cranking it out every single year. Big leaps every single year.
我猜他们的逻辑是,不需要更好,只要不比 Nvidia 差 70% 以上就行,因为他们付给你 70% 的利润率。
I guess their logic is that it doesn't need to be better. It just needs to be not more than 70% worse because they're paying you 70% margins.
不,不,不。别忘了,即使是 ASIC 的利润率也相当高。假设 Nvidia 的利润率是 70%,但 ASIC 的利润率是 65%。你真正省下了什么?
No, no, no. Don't forget even an ASIC margin is really quite high. Nvidia's margin is 70% let's say, but an ASIC margin is 65%. What are you really saving?
哦,你是说从 Broadcom 那里?你总得付钱给别人。
Oh, you mean from Broadcom or something? You got to pay somebody.
是的。所以我认为 ASIC 的利润率从我所知来看非常高,他们自己也这么认为。他们对自己惊人的 ASIC 利润率感到自豪。所以你问为什么。很久以前,我们根本没有能力做到这一点。当时我没有深刻认识到建立一个像 OpenAI 和 Anthropic 这样的基础 AI 实验室有多困难,以及他们需要来自供应商自身的巨额投资。我们当时无法向 Anthropic 投资数十亿美元,让他们使用我们的算力。但 Google 和 AWS 可以,他们在初期投入了巨额投资,作为回报,Anthropic 使用他们的算力。我们当时没有这个能力。我的错误在于,我没有深刻认识到他们真的别无选择,风投永远不会向一个 AI 实验室投资 50 到 100 亿美元,期望它成为 Anthropic。这是我的失误。但即使我理解了,我也不认为我们当时有能力做到。但我不会再犯同样的错误。我很高兴投资了 OpenAI,并帮助他们扩展规模,我相信这至关重要。当 Anthropic 来找我们时,我很高兴成为投资者,帮助他们扩展。但我们当时确实无法做到。
Yeah. And so I think the ASIC margins are incredibly good from what I can tell and they believe it too. And so they're quite proud of their incredible ASIC margins. So you ask the question why. A long time ago we just didn't have the ability to do it. And at the time I didn't deeply internalize how difficult it would be to build a foundation AI lab like OpenAI and Anthropic. And the fact that they needed huge investments from the supplier themselves. We just weren't in a position to make the multi-billion dollar investment into Anthropic so that they could use our compute. But Google and AWS were and they put in huge investments in the beginning so that Anthropic in return use their compute. We just weren't in a position to do so at the time. Nor did I, I would say my mistake is I didn't deeply internalize that they really had no other options, that a VC would never put in 5-10 billion of investment into an AI lab with the hopes of it turning out to be Anthropic. And so that was my miss. But even if I understood it, I don't think we would have been in a position to do that at the time. But I'm not going to make that same mistake again. And I'm delighted to invest in OpenAI and I'm delighted to help them scale and I believe it's essential to do so. And then when Anthropic came to us, I'm delighted to be an investor, delighted to help them scale. But we just weren't at the time able to do so.
如果我能倒带一切,Nvidia 当时如果能像现在这么大,我会非常乐意这么做。这其实很有趣,多年来 Nvidia 一直是 AI 领域赚钱的公司,赚了很多钱,现在你开始投资。据报道,你向 OpenAI 投资了高达 300 亿美元,向 Anthropic 投资了 100 亿美元。但现在它们的估值已经上涨,我相信还会继续上涨。所以,如果这些年来,你给他们提供算力,你看到了发展方向,而它们在几年前甚至一年前还只值现在十分之一,而你又有这么多现金。那么,要么 Nvidia 自己成为一个基础实验室,进行巨额投资来实现这一点,要么在更早的时候以当前估值达成你现在所做的交易。你当时有现金,所以我很好奇为什么没有更早这么做。
If I could rewind everything, Nvidia could have been as big back then as we are now, I would have been more than happy to do it. This is actually quite interesting, which is for many years Nvidia has been this company in AI making money, making lots of money, and now you're investing. It's been reported that you've done up to 30 billion in OpenAI and 10 billion in Anthropic. But now their valuations have increased and I'm sure they'll continue to increase. And so if overall these many years, you were giving them the compute, you saw where it was headed, and then they were worth like one-tenth what they are now a couple years ago or even a year ago in some cases, and you had all this cash. There's a world where either Nvidia themselves becomes a foundation lab, does the huge investment to make that possible, or has made the deals you've made now at current valuations much earlier on. And you had the cash to do it. So I am curious actually why not have done it earlier.
我们一有能力就做了。我们一有可能就做了。如果我能更早做,我会的。在 Anthropic 需要我们的时候,我们当时没有能力去做。这不符合我们的理念。
We did it as soon as we could. We did it as soon as we could have. And if I could have, I would have done it even earlier. At the time that Anthropic needed us to do it, we just weren't in a position to do it. It wasn't in our sensibility to do so.
那是资金问题还是……
How's that like a cash thing or just...
是的,投资规模太大了。我们当时从未对外投资,也没有那么多资金。我们没有意识到需要这么做。我总以为他们可以去找风投,就像所有公司一样。但他们想做的事情无法通过风投完成。OpenAI 想做的事情无法通过风投完成。我现在明白了,但当时不知道。但这就是他们的天才之处,这就是他们聪明的地方。他们当时就意识到必须这么做。我很高兴他们做到了。尽管我们导致 Anthropic 不得不去找别人,我仍然很高兴事情发生了。Anthropic 的存在对世界是件好事。我为此感到高兴。
Yeah, the level of investment, you know, we never invested outside the company at the time and not that much. And we didn't realize we needed to. I always thought that they could just go raise VCs for God's sakes, like all companies do. But what they were trying to do couldn't have been done through VCs. What OpenAI wanted to do couldn't have been done through VCs. And I recognize that now. I didn't know it then. But that's their genius. That's why they're smart. And so they realized it then that they had to do something like that. And I'm delighted that they did. And even though we caused Anthropic to have to go to somebody else, I'm still happy that it happened. Anthropic's existence is great for the world. I'm delighted for it.
我想你仍然在赚很多钱,而且每个季度赚得更多。
I guess you still are making a ton of money and you're making way more money quarter after quarter.
有遗憾也是可以的。
It's still okay to have regrets.
那么问题仍然存在:现在我们到了这里,你拥有这么多不断赚来的钱,英伟达应该用它做什么?有一种答案认为,出现了一个完整的中介生态系统,将这些实验室的资本支出转化为运营支出,以便他们可以租用算力,因为芯片非常昂贵。它们在整个生命周期中能赚很多钱,因为模型越来越好,它们从代币中产生的价值在增加,但建立成本很高。英伟达有钱做资本支出。事实上,据报道你正在支持 CoreWeave。你拥有高达 63 亿美元,并已投资了 20 亿。但为什么英伟达不自己成为云服务商?为什么不自己成为超大规模云服务商并运行这些计算?你有这么多现金可以做到。
So then the question still arises: now that we're here and you have all this money that you keep making, what should Nvidia be doing with it? There's one answer which says there's this whole middleman ecosystem that has popped up for converting capex into opex for these labs so that they can rent compute, because the chips are really expensive. They make a lot of money over their lifetime because the models are getting better, the value that they generate from their tokens is increasing, but they're expensive to set up. Nvidia has the money to do the capex. In fact, it's been reported you're backstopping CoreWeave. You have up to 6.3 billion and have invested 2 billion. But why doesn't Nvidia become a cloud themselves? Why doesn't it become a hyperscaler themselves and run this compute out? You have all this cash to do it.
这是公司的理念,我认为是明智的。我们应该做必要的事,尽可能少做。这意味着我们在构建计算平台方面所做的工作。如果我们不做,我真心相信它就不会完成。如果我们不承担我们所承担的风险,如果我们不按照我们的方式构建 NVLink,如果我们不构建整个堆栈,如果我们不按照我们的方式创建生态系统,如果我们不致力于 20 年的 CUDA 而大部分时间都在亏损,如果我们不做,没有人会做。如果我们不创建所有 CUDA X 库,使它们都是领域特定的——这是几十年前的事了,我们推动领域特定库,因为我们意识到如果我们不创建这些领域特定库,无论是用于光线追踪、图像生成还是 AI 的早期工作,这些模型,如果我们不为数据处理、结构化数据处理或向量数据处理创建它们,没有人会做。我完全确定这一点。我们创建了一个用于计算光刻的库叫 cuLitho。如果我们不创建它,没有人会做。如果我们不做我们所做的,加速计算就不会像现在这样进步。我们应该全心全意地将公司所有的力量投入到这件事上。然而,世界上有很多云服务商。如果我不做,别人会出现。遵循做必要的事但尽可能少做的理念,这个理念今天存在于我们的公司,我所做的一切都带着这个视角。在云服务方面,如果我们不支持 CoreWeave 的存在,这些新云、这些 AI 云就不会存在。如果我们不帮助 CoreWeave 存在,它们就不会存在。如果我们不支持 Nscale,它们就不会有今天的成就。如果我们不支持 NBS,它们就不会有今天的成就。现在,它们做得非常好。这是商业模式吗?不,我们应该做必要的事,尽可能少做。我们投资于我们的生态系统,因为我希望我们的生态系统蓬勃发展。我希望架构和 AI 能够连接尽可能多的行业、尽可能多的国家,并使地球能够建立在 AI 之上,建立在美国技术栈之上。这个愿景正是我们追求的。
This is a philosophy of the company and I think is wise. We should do as much as needed, as little as possible. What that means is the work that we do with building our computing platform. If we don't do it, I genuinely believe it doesn't get done. If we didn't take the risk that we take, if we didn't build NVLink the way we built, if we didn't build the whole stack, if we didn't create the ecosystem the way we did, if we didn't dedicate ourselves to 20 years of CUDA while losing money most of that time, if we didn't do it, nobody else would have done it. If we didn't create all the CUDA X libraries so that they're all domain specific—this is several decades ago, we pushed into domain specific libraries because we realized that if we didn't create these domain specific libraries, whether it's for ray tracing or image generation or even the early works of AI, these models, if we didn't create them for data processing, structured data processing or vector data processing, if we didn't create them, nobody would. I am completely certain of that. We created a library for computational lithography called cuLitho. If we didn't create it, nobody would have. Accelerated computing wouldn't advance the way it has if we didn't do what we did. We should dedicate our company all of our might wholeheartedly to go do that. However, the world has lots of clouds. If I didn't do it, somebody would show up. Following the philosophy of doing as much as needed but as little as possible, that philosophy exists in our company today and everything I do, I do it with that lens. In the case of clouds, if we didn't support CoreWeave to exist, these neo clouds, these AI clouds wouldn't exist. If we didn't help CoreWeave exist, they would not exist. If we didn't support Nscale, they wouldn't be where they are today. If we didn't support NBS, they wouldn't be where they are today. Now, they are doing fantastically. Is that a business model? No, we should do as much as needed, as little as possible. We invest in our ecosystem because I want our ecosystem to thrive. I want the architecture and I want AI to be able to connect with as many industries as possible, as many countries as possible, and make it possible for the planet to be built on AI and to be built on the American tech stack. That vision is exactly what we're pursuing.
你提到的一件事是,有很多很棒的基础模型公司,我们试图投资所有它们。这是我们做的另一件事。我们不挑选赢家,我们需要支持每个人,这是我们这样做的乐趣的一部分。这对我们的业务是必要的,但我们也会刻意不挑选赢家。所以当我投资其中一家时,我投资所有。为什么你刻意不挑选赢家?
One of the things that you mentioned, there are so many great foundation model companies and we try to invest in all of them. This is another thing that we do. We don't pick winners and we need to support everyone and it's part of our joy of doing so. It's an imperative to our business, but we also go out of our way not to pick winners. So when I invest in one of them, I invest in all of them. Why do you go out of your way not to pick winners?
因为这不是我们的工作。第一。第二,当英伟达刚起步时,有 60 家图形公司,60 家 3D 图形公司。我们是唯一幸存下来的。如果你拿那 60 家公司问自己哪一家会成功,英伟达会是那个最不可能成功的名单上的第一名。这远在你之前,但英伟达的图形架构是完全错误的。不是有点错误。我们创建了一个完全错误的架构。开发者不可能支持它。它永远不会成功。我们从良好的第一性原理推理,但最终得到了错误的解决方案。每个人都会认为我们不行,但我们在这里。我有足够的谦逊认识到你不挑选赢家。要么让它们都自己照顾自己,要么照顾所有。
Because it's not our job to. Number one. Number two, when Nvidia first started, there were 60 graphics companies, 60 3D graphics companies. We are the only one that survived. If you would have taken those 60 companies and asked yourself which one was going to make it, Nvidia would be the top of that list not to make it. This is long before you, but Nvidia's graphics architecture was precisely wrong. It's not a little bit wrong. We created an architecture that was precisely wrong. It was an impossible thing for developers to support. It was never going to make it. We reasoned about it from good first principles, but we ended up in the wrong solution. Everybody would have counted us out, and here we are. I have enough humility to recognize that you don't pick winners. Either let them all take care of themselves or take care of all of them.
有一件事我不明白,你说:「看,我们优先考虑这些新云并不是因为它们是新的云,我们想支持它们。」但你也列出了一些新云,并说如果没有英伟达它们就不会存在。这两件事如何兼容?
One thing I didn't understand is you said, 'Look, we're not prioritizing these neo clouds just because there are new clouds and we want to prop them up.' But you also said you listed a bunch of new clouds and you said they wouldn't exist if it wasn't for Nvidia. How are those two things compatible?
首先,它们需要想要存在,并且它们来向我们寻求帮助。当它们想要存在,并且有商业计划、专业知识和热情时,它们显然必须自己有一些能力。但如果最终它们需要一些投资来启动,我们会支持它们。但一旦它们的飞轮开始转动,你的问题是:我们想从事融资业务吗?答案是否定的。我们不想,因为有人从事融资业务,我们宁愿与所有从事融资业务的人合作,而不是自己成为融资者。我们的目标是专注于我们做的事情,保持我们的商业模式尽可能简单,支持我们的生态系统。当像 OpenAI 这样的公司需要 300 亿美元规模的投资,因为它还在 IPO 之前,我们深深相信它们,我深深相信它们今天已经是一家非凡的公司,它们将成为一家不可思议的公司。世界需要它们存在。世界希望它们存在。我希望它们存在,它们拥有一切,顺风顺水。让我们支持它们,让它们扩张。对于那些投资我们会做,因为它们需要我们这样做。但我们并不试图做尽可能多的事。
First of all, they need to want to exist and they come to ask us for help. When they want to exist and they have a business plan and they have expertise and they have the passion for it, they obviously have to have some capabilities themselves. But if at the end of the day they need some investment in order to get it off the ground, we would be there for them. But the sooner they get their flywheel going, your question was do we want to be in the financing business? The answer is no. We don't want to be, because there are people in the financing business and we rather work with all of the people who are in the financing business than to be a financier ourselves. Our goal is to focus on what we do, keep our business model as simple as possible, support our ecosystem. When someone like OpenAI needs an investment of $30 billion scale because it's still before their IPO and we deeply believe in them, I deeply believe that they are going to be an extraordinary company already today, they are going to be an incredible company. The world needs them to exist. The world wants them to exist. I want them to exist and they have everything, they have the wind at their back. Let's support them and let them scale. To those investments we will do because they need us to do it. But we're not trying to do as much as possible.
我们尽量少做事。我花了太多时间在 Google Docs 和聊天机器人之间复制粘贴文本。所以我建了一个基本上是用于写作的 cursor,它按照我认为 AI 共同研究员应该的方式运作。我可以标记它,它可以通过内联评论线程与我交谈,帮助我深入挖掘和头脑风暴。我整个周末用 cursor 和他们的新 composer 2 模型写了这个。对于很多智能体式编码工具,我觉得我对底层发生了什么一无所知。我只能放弃控制,希望最好。但 cursor 让我尝试了很多不同的想法,同时保持对实现的掌控。我大部分头脑风暴是在 agents 窗口中完成的。在放好一些基本文件后,我用 diff 窗口跟踪变化。少数几次我需要手动快速调整时,我就用编辑器。如果你想亲自尝试我的 AI 代码研究员,我在描述中链接了 GitHub 仓库。如果你有一个一直想构建的工具,你应该让它发生。访问 cursor.com/cash 开始。
We're trying to do as little as possible. I spend way too much time copy pasting text back and forth from Google Docs to chatbots. And so I built what's basically a cursor for writing which operates the way I think an AI co-researcher should operate. I can tag it and it can talk with me through inline comment threads and help me dig deeper and brainstorm. I wrote this entire thing over the weekend with cursor and their new composer 2 model. With a lot of agentic coding tools, I feel like I have no idea what's going on under the surface. I just have to relinquish control and hope for the best. But cursor let me try a bunch of different ideas while staying on top of the implementation. I did most of my brainstorming in the agents window. And after I got some basic files in place, I used a diff window to track changes. The few times that I needed to make a quick tweak by hand, I just used the editor. If you want to try my AI code researcher yourself, I've linked the GitHub repo in the description. And if you have a tool that you've been wanting to build, you should make it happen. Go to cursor.com/cash to get started.
这可能是个显而易见的问题,但我们多年来一直处于 GPU 短缺的状态,现在因为模型越来越好,短缺更严重了。
This may be sort of an obvious question, but we've lived many years in this situation where there's a shortage of GPUs and it's grown now because models are getting better.
我们确实有 GPU 短缺。
We have a shortage of GPUs.
是的。
Yes.
嗯。
Yeah.
而英伟达以分割稀缺分配而闻名,不仅仅基于最高出价者,而是基于「嘿,我们想确保这些新云存在。给 Core 一些,给 Crusoe 一些,给 Lambda 一些。」这对英伟达有什么好处?首先,你同意这种市场分割的描述吗?
And Nvidia is known for diving up the scarce allocation not just based on highest bidder but rather on hey we want to make sure that these neo neo clouds exist. Let's give some to Core. Let's give some to Crusoe. Let's give some to Lambda. Why is it good for Nvidia? First of all, would you agree with this characterization of fracturing the market?
不。你的前提就是错的。
No. Your premise is just wrong.
嗯。
Yeah.
我们对这些事情足够谨慎。首先,如果你不下采购订单,说再多也没用。所以在收到订单之前,我们能做什么?第一件事是我们和每个人都努力做预测,因为这些东西需要很长时间来建造,数据中心也需要很长时间。所以我们通过预测来匹配供需。这是第一要务。第二,我们尽量和尽可能多的人做预测,但归根结底,你还是得下订单。也许出于某种原因你没下单,我能怎么办?所以到某个时候,就是先到先得。但除此之外,如果你没准备好,因为你的数据中心没准备好,或者某些组件没准备好让你建立数据中心,我们可能会决定先服务另一个客户。这只是为了最大化我们工厂的吞吐量。所以我们可能会做一些调整。除此之外,优先级就是先到先得。
We are sufficiently mindful about these things. First of all, if you don't place a PO, all the talking in the world won't make a difference. And so until we get a PO, what are we going to do? The first thing is we work really hard with everybody to get a forecast done because these things take a long time to build and the data centers take a long time to build. So we align ourselves with demand and supply through forecasting. That's job number one. Number two, we've tried to forecast with as many people as possible, but in the final analysis, you still had to place an order. Maybe for whatever reason, you didn't place your order, what can I do? So at some point, first in first out. But beyond that, if you're not ready because your data center is not ready or certain components aren't ready to enable you to stand up a data center, we might decide to serve another customer first. That's just maximizing the throughput of our factory. So we might do some adjustments there. Aside from that, the prioritization is first in first out.
是的。你必须下订单。如果你不下订单,当然有故事,比如这一切都始于一篇关于 Larry 和 Elon 和我共进晚餐的文章,他们乞求 GPU。
Yeah. You got to place a PO. If you don't place a PO, now of course there are stories about that, like for example, all of this kind of started from an article about Larry and Elon having dinner with me where they begged for GPUs.
那从未发生过。我们确实吃了晚餐。那是一顿很棒的晚餐。他们从未乞求 GPU。他们只需要下订单,一旦下了订单,我们会尽力把产能给他们。我们不复杂。
That never happened. We absolutely had dinner. It was a wonderful dinner. At no time did they beg for GPUs. They just had to place an order and once they place an order we do our best to get the capacity to them. We're not complicated.
好的。所以听起来有一个队列,然后根据你的数据中心是否准备好以及你下采购订单的时间,你在某个时间得到它们。但这听起来仍然不是最高出价者就能得到。有什么理由这样做吗?
Okay. So it sounds like there's a queue and then based on whether your data center is ready and when you place a purchase order, you get them a certain time. But it still doesn't sound like highest bidder just gets it. Is there a reason to do it?
我们从不那样做。
We never do that.
好的。
Okay.
我们从不。
We never do.
为什么不直接给最高出价者?
Why not just do highest bidder?
因为这是糟糕的商业实践。你定好价格,然后人们决定买不买。我知道芯片行业的其他公司在需求高时会调整价格。但我们不。这从来不是我们的做法。你可以信赖我们。我更愿意成为可靠的人,成为行业的基础。你不需要猜疑。如果我给你报了价,那就是最终价。即使需求飙升,也如此。
Because it's a bad business practice. You set your price. You set your price and then people decide to buy it or not. I understand that others in the chip industry change their prices when demand is higher. But we just don't. That's never been a practice of ours. You can count on us. I prefer to be dependable, to be the foundation of the industry. You don't need to second guess. If I quoted you a price, that's it. And if demand goes through the roof, so be it.
另一方面,这就是为什么你和台积电有富有成效的关系,对吧?
And on the other end, that's why you have a productive relationship with TSMC, right?
是的。英伟达已经经营了,我们和他们做生意大概快 30 年了,英伟达和台积电没有法律合同。总有一些粗略的公平,有时我对,有时我错。有时我得到更好的交易,有时我得到更差的。但总的来说,关系非常好,我完全信任他们。我完全依赖他们。你可以信赖英伟达的一点是,明年 Vera Rubin 会很棒。后年 Vera Rubin Ultra 会来。再后年 Feynman 会来,再后年我还没公布名字。所以每年你都可以信赖我们。你必须在世界上找另一个 ASIC 团队。选一个你能说「我可以把整个业务押注在你每年都会为我服务」的 ASIC 团队。你的成本,你的 token 成本每年会降低一个数量级。我可以像相信时钟一样相信这一点。我刚才说了关于台积电的事。历史上没有其他代工厂能让你这么说。今天你可以这么说英伟达。每年你都可以信赖我们。如果你想买价值十亿美元的 AI 工厂算力,没问题。如果你想买一亿美元,没问题。你想买一千万美元或仅仅一个机架,没问题。或者仅仅一张显卡,没问题。如果你想下一个一千亿美元的 AI 工厂订单,没问题。今天我们是世界上唯一一家你可以这么说的公司。我也可以这么说台积电。我想买十亿,没问题。
Yeah. Nvidia has been in business, we've been doing business with them for I guess coming up on 30 years and Nvidia and TSMC don't have a legal contract. There is always some rough justice and sometimes I'm right, sometimes I'm wrong. Sometimes I got a better deal, sometimes I got a worse deal. But overall the relationship is incredible and I can completely trust them. I completely depend on them. One of the things that you can count on with Nvidia is that next year Vera Rubin is going to be incredible. The year after Vera Rubin Ultra will come. The year after that Feynman will come and the year after that I haven't introduced the name yet. So every single year you can count on us. You're going to have to go find another ASIC team in the world. Pick your ASIC team where you can say I can bet my entire business that you will be here for me every single year. Your cost, your token cost will decrease by an order of magnitude every single year. I can count on it like I can count on the clock. I just said something about TSMC. No other foundry in history can you possibly say that. You can say that about Nvidia today. You can count on us every single year. If you would like to buy a billion dollars worth of AI factory compute, no problem. If you like to buy $100 million, no problem. You'd like to buy $10 million or just one rack, not a problem. Or just one graphics card, no problem. If you would like to place an order for a hundred billion dollar AI factory, no problem. We're the only company in the world where you can say that today. I can say that about TSMC as well. I want to buy one billion, no problem.
我们只需要经历规划的过程,以及成熟人士所做的所有事情。所以我认为,英伟达能够成为全球 AI 行业的基础,这个地位我们花了二十年来达成,付出了巨大的承诺和奉献。我们公司的稳定性和一致性非常重要。
We just got to go through the process of planning for it and all the things that mature people do. So I think this ability for Nvidia to be the foundation of the world's AI industry is a position that has taken us a couple of decades to arrive at, enormous commitment, enormous dedication. And the stability of our company, the consistency of our company is really, really important.
好的。我想问关于中国的问题。我总是喜欢扮演魔鬼代言人,与我的嘉宾唱反调。当 Dario 来的时候,他支持出口管制,我问他为什么美国和中国不能各自拥有一个天才数据中心。但既然你持相反立场,我就反过来问你。一种思考方式是,Anthropic 几天前发布了一个模型 Mythos,他们甚至没有公开发布,因为他们说它具有如此强大的网络攻击能力,以至于他们认为世界还没有准备好,直到这些零日漏洞被修补。但他们说,它发现了数千个高严重性漏洞,覆盖所有主要操作系统和浏览器。它在 OpenBSD 中发现了一个漏洞,这个操作系统是专门设计来避免零日漏洞的,而且它存在了 27 年。所以,如果中国公司、中国实验室和中国政府能够获得 AI 芯片来训练像 Claude Mythos 这样具有网络攻击能力的模型,并用更多的算力运行数百万个实例,问题是,这对美国公司和美国国家安全构成威胁吗?
Okay. I want to ask about China. And I always like to play devil's advocate against my guest. When Dario was on, who supports export controls, I asked him why can't America and China both have a country of geniuses in a data center. But since you're on the opposite side, I'll ask you in the opposite way. One way to think about it is Anthropic actually announced a couple days ago a model, Mythos, they are not even releasing publicly because they say it has such cyber offensive capabilities that we don't think the world is ready until we make sure these zero days are patched up. But they say it found thousands of high severity vulnerabilities across every major operating system, every browser. It found one in OpenBSD, which is an operating system specifically designed to not have zero days, and it found one for 27 years it's existed. So if Chinese companies, Chinese labs, and the Chinese government had access to the AI chips to train a model like Claude Mythos with these cyber offensive capabilities and run millions of instances of it with more compute, the question is, is that a threat to American companies and American national security?
首先,Mythos 是由一家非凡的公司用相当普通的算力和数量训练出来的。训练它所需的算力类型和数量在中国非常充足。所以你必须首先意识到,芯片在中国是存在的。他们制造了全球 60%的主流芯片,可能更多。这对他们来说是一个非常大的产业。他们拥有一些世界上最伟大的计算机科学家。如你所知,所有 AI 实验室中的大多数 AI 研究人员都是华人。他们拥有全球 50%的 AI 研究人员。所以问题是,如果你担心他们,考虑到他们已经拥有的所有资产——丰富的能源、充足的芯片、大多数 AI 研究人员——创造安全世界的最佳方式是什么?把他们变成受害者、变成敌人,可能不是最好的答案。他们是竞争对手。我们希望美国获胜。但我认为,进行对话和研究交流可能是最安全的做法。由于我们目前将中国视为对手的态度,这一领域明显缺失。我们的 AI 研究人员和他们的 AI 研究人员必须进行交流。我们必须就 AI 在软件漏洞发现方面的用途达成一致。当然,这正是 AI 应该做的。它会在大量软件中发现漏洞吗?当然。有很多漏洞,包括 AI 软件中的。所以这就是 AI 应该做的。我很高兴 AI 已经达到了能够帮助我们提高生产力的水平。被低估的一点是围绕网络安全、AI 安全、AI 隐私和 AI 安全的生态系统的丰富性。整个 AI 初创公司生态系统正在努力创造一个未来,其中有一个强大的 AI 智能体,周围有成千上万个 AI 智能体保护它的安全。这个未来一定会到来。而让一个 AI 智能体在无人监管的情况下到处乱跑的想法是疯狂的。所以我们很清楚,这个生态系统需要蓬勃发展。事实证明,这个生态系统需要开源、开放模型、开放堆栈,以便所有这些 AI 研究人员和伟大的计算机科学家能够构建同样强大的 AI 系统,并确保 AI 安全。所以我们需要确保的一件事是保持开源生态系统的活力,这一点不容忽视。其中很多来自中国。我们不能扼杀它。关于中国,我们希望美国拥有尽可能多的算力。我们受到能源的限制。但我们有很多人在研究这个问题,我们不能让能源成为我们国家的瓶颈。但我们也希望确保全球所有 AI 开发者都在美国技术栈上进行开发,并将 AI 的贡献和进步,尤其是开源方面的,提供给美国生态系统。创建两个生态系统将是极其愚蠢的:一个只运行在中国技术栈上的开源生态系统,和一个只运行在美国技术栈上的封闭生态系统。我认为这对美国来说将是一个可怕的结果。
First of all, Mythos was trained on fairly mundane capacity and a fairly mundane amount of it by an extraordinary company. The amount of capacity and the type of compute it was trained on is abundantly available in China. So you just have to first realize that chips exist in China. They manufacture 60% of the world's mainstream chips, maybe more. It's a very large industry for them. They have some of the world's greatest computer scientists. As you know, most of the AI researchers in all of these AI labs, most of them are Chinese. They have 50% of the world's AI researchers. So the question is, if you're concerned about them, considering all the assets they already have—they have an abundance of energy, plenty of chips, most of the AI researchers—what is the best way to create a safe world? Victimizing them, turning them into an enemy, likely isn't the best answer. They are an adversary. We want the United States to win. But I think having a dialogue and research dialogue is probably the safest thing to do. This is an area that is glaringly missing because of our current attitude about China as an adversary. It is essential that our AI researchers and their AI researchers are actually talking. It is essential that we try to both agree on what not to use AI for with respect to finding bugs in software. Of course, that's what AI is supposed to do. Is it going to find bugs in a lot of software? Of course. There are lots of bugs, including in AI software. So that's what AI is supposed to do. And I'm delighted that AI has reached a level where it could help us be so much more productive. One of the things that is underemphasized is the richness of ecosystem around cybersecurity, AI security, AI privacy, and AI safety. That whole ecosystem of AI startups is trying to create a future where you have one AI agent that's incredible surrounded by thousands of AI agents keeping it safe and secure. That future surely is going to happen. And the idea that you're going to have an AI agent running around with nobody watching after it is kind of insane. So we know very well that this ecosystem needs to thrive. It turns out this ecosystem needs open source, open models, open stacks so that all these AI researchers and great computer scientists can go build AI systems that are as formidable and can keep AI safe. So one of the things we need to make sure we do is keep the open-source ecosystem vibrant, and that can't be ignored. A lot of that is coming out of China. We must not suffocate that. With respect to China, we want the United States to have as much computing as possible. We're limited by energy. But we have a lot of people working on that, and we must not make energy a bottleneck for our country. But what we also want is to make sure that all the AI developers in the world are developing on the American tech stack and making contributions and advancements of AI, especially when it's open source, available to the American ecosystem. It would be extremely foolish to create two ecosystems: an open-source ecosystem that only runs on the Chinese tech stack, and a closed ecosystem that runs on the American tech stack. I think that would be a horrible outcome for the United States.
由于有很多事情,让我梳理一下回应。我认为回到算力差异和黑客攻击的担忧是:是的,他们有算力,但有一些估计认为,由于他们使用 7 纳米工艺,并且由于芯片制造出口管制而没有 EUV,他们实际能产生的算力大约只有美国的十分之一。那么,他们最终能训练出像 Mythos 这样的模型吗?是的。但问题是,因为我们有更多的算力,美国实验室能够首先达到这些能力水平。而且因为 Anthropic 首先做到了,他们说,「好吧,我们将保留它一个月,同时我们让所有美国公司访问它,他们将修补所有漏洞,然后我们再进一步发布。」即使他们训练了这样的模型,大规模部署的能力——如果你有一个网络黑客,拥有一百万个比拥有一千个要危险得多。所以推理算力确实非常重要。事实上,他们拥有这么多优秀的研究人员,这才是真正可怕的地方,因为让工程师研究人员更高效的是算力。如果你和任何美国实验室交谈,他们会说制约他们的是算力。
Since there are a lot of things, let me just triage the response. I think the concern going back to the flop difference and the hacking is: yes, they have compute, but there are some estimates that because they are at 7 nanometer, they don't have EUV due to chip-making export controls, the amount of flops they are able to actually produce is like one-tenth the amount of flops that the US has. So with that, could they eventually train a model like Mythos? Yes. But the question is, because we have more flops, American labs are able to get to these level capabilities first. And because Anthropic got to it first, they said, 'Okay, we're going to hold on to it for a month while all these American companies we give access to it, they're going to patch up all their vulnerabilities, and now we release it further.' Even if they train a model like this, the ability to deploy it at scale—if you had a cyber hacker, it's much more dangerous if they have a million of them versus a thousand of them. So that inference compute really matters a lot. And in fact, the fact that they have so many researchers who are so good is the thing that makes it so scary, because what makes engineer researchers more productive is compute. If you talk to any lab in America, they say the thing that's bottlenecking them is compute.
所以有 DeepSeek 创始人或领导层的引述说,我们受瓶颈限制的是算力。那么问题来了,美国公司因为拥有更多算力,先达到能力水平,让我们的社会做好准备,然后中国再达到,这样不是更好吗?我们应该永远第一,永远拥有更多。
So there are quotes from DeepSeek founder or leadership saying the thing we're bottlenecked on is compute. So then the question is, isn't it better that American companies, because they have more compute, get to the level of capabilities first, prepare our society for it before China can get there because they have less compute? We should always be first and we should always have more.
但要让这个结果成立,你必须把它推到极端。他们必须没有算力。如果他们有一些算力,问题就在于需要多少。中国拥有的算力是巨大的。我的意思是,你说的是全球第二大计算市场。如果他们想部署和聚合算力,他们有大量算力可以聚合。
But for that outcome to be true, you have to take it to the extremes. They have to have no compute. If they have some compute, the question is how much is needed. The amount of compute they have in China is enormous. I mean, you're talking about the country being the second largest computing market in the world. If they want to deploy and aggregate their compute, they have plenty of compute to aggregate.
但这是真的吗?我的意思是,有人做这些估算,他们说中芯国际在工艺节点上实际上落后了。
But is that true? I mean, people do these estimates and they say SMIC is actually behind on process nodes.
我正要告诉你。他们拥有的能源量是惊人的。AI 是一个并行计算问题。他们为什么不能把四倍或十倍数量的芯片堆在一起?因为能源是免费的。他们有那么多能源。他们有完全空置、通电的数据中心。他们有鬼城、鬼数据中心。他们有如此多的基础设施容量。如果他们想,他们可以堆更多芯片,即使是七纳米的。而且他们的芯片制造能力是世界上最大的之一。半导体行业知道他们垄断了主流芯片。他们产能过剩。所以认为中国无法获得 AI 芯片的想法完全是胡说八道。当然,如果你问我,如果整个世界都没有算力,美国会不会更领先?但那不是真实的情景。他们已经拥有大量算力。你担心的那个门槛,他们已经达到了甚至超过了。所以我认为你误解了 AI 是一个五层蛋糕。最底层是能源。当你拥有丰富的能源时,它可以弥补芯片的不足。如果你有丰富的芯片,它可以弥补能源的不足。例如,美国能源稀缺,这就是为什么英伟达必须不断推进我们的架构并进行极致的代码设计,这样我们出货的少量芯片,因为能源非常有限,我们的每瓦吞吐量是惊人的。但如果你的瓦特数完全充足且免费,你还在乎每瓦性能干什么?所以七纳米芯片基本上就是 Hopper。今天的模型主要是在 Hopper 代上训练的。Hopper 七纳米芯片已经足够好了。能源丰富是他们的优势。
I'm about to tell you. The amount of energy they have is incredible. AI is a parallel computing problem. Why can't they just put four or ten times as many chips together? Because energy is free. They have so much energy. They have data centers sitting completely empty, fully powered. They have ghost cities, ghost data centers. They have so much infrastructure capacity. If they wanted to, they just gang up more chips, even if they are seven nanometer. And their capacity for building chips is one of the largest in the world. The semiconductor industry knows they monopolize mainstream chips. They have overcapacity. So the idea that China won't be able to have AI chips is completely nonsense. Now, of course, if you ask me, would the United States be further ahead if the entire world had no compute at all? But that's not a scenario that's true. They have plenty of compute already. The threshold they need for the concern you're worried about, they've already reached that threshold and beyond. So I think you misunderstand that AI is a five-layer cake. At the lowest layer is energy. When you have abundant energy, it makes up for chips. If you have abundant chips, it makes up for energy. For example, the United States is scarce on energy, which is why Nvidia has to keep advancing our architecture and do extreme code design so that with the few chips we ship, because the amount of energy is so limited, our throughput per watt is off the charts. But if your amount of watts is completely abundant and free, what do you care about performance per watt? So seven nanometer chips are essentially Hopper. Today's models are largely trained on Hopper generation. Hopper seven nanometer chips are plenty good. The abundance of energy is their advantage.
但接下来有个问题,考虑到他们的限制,他们是否真的能制造足够的芯片。
But then there's a question of whether they can actually manufacture enough chips given their constraints.
但他们确实做到了。证据是什么?华为刚刚度过了公司历史上最大的一年。
But they do. What's the evidence? Huawei just had the largest single year in the history of their company.
他们出货了多少芯片?
How many chips did they ship?
大量。数百万。数百万比 Anthropic 拥有的多得多。
A ton. Millions. Millions is way more than Anthropic has.
所以问题在于中芯国际和长鑫存储能提供多少逻辑芯片,以及多少内存。
So there's a question of how much logic from SMIC and CXMT, and how much memory.
我告诉你实际情况。他们有大量的逻辑芯片和大量的 HBM2 内存。
I'm telling you what it is. They have plenty of logic and plenty of HBM2 memory.
对。但如你所知,训练和推理的瓶颈通常是内存带宽。HBM2 与最新的 HBM 相比,可能有近一个数量级的差异。
Right. But as you know, the bottleneck in training and inference is often memory bandwidth. HBM2 versus the newest HBM can be almost an order of magnitude difference.
华为是一家网络公司。
Huawei is a networking company.
但这改变不了你需要 EUV 来制造最先进的 HBM 的事实。
But that doesn't change the fact that you need EUV for the most advanced HBM.
不对。完全不对。你可以把它们堆在一起,就像我们用 NVLink72 做的那样。他们已经展示了用硅光子学将所有计算连接成一个巨型超级计算机。你的前提完全错了。事实是,他们的 AI 发展得很好。而且因为世界上最好的 AI 研究人员在算力上受限,他们也想出了极其聪明的算法。记住我说的:摩尔定律每年进步约 25%。然而,通过优秀的计算机科学,我们仍然可以将算法性能提升 10 倍。我的意思是,优秀的计算机科学才是杠杆。毫无疑问,惊人的注意力机制减少了算力需求。我们必须承认,AI 的大部分进步来自算法进步,而不仅仅是原始硬件。现在,如果大部分进步来自算法和计算机科学,那么告诉我,他们庞大的 AI 研究人员队伍难道不是他们的根本优势吗?我们看到了。DeepSeek 不是微不足道的进步。而 DeepSeek 首先在华为上发布的那一天,对我们国家来说将是一个可怕的结果。
Not true. Not at all true. You can gang them together just like we do with NVLink72. They've already demonstrated silicon photonics connecting all of this compute into one giant supercomputer. Your premise is just wrong. The fact of the matter is their AI development is going just fine. And because the best AI researchers in the world are limited in compute, they also come up with extremely smart algorithms. Remember what I said: Moore's law is advancing about 25% per year. However, through great computer science, we can still improve algorithm performance by 10x. What I'm saying is great computer science is where the lever is. There is no question that incredible attention mechanisms reduce the amount of compute. We have to acknowledge that most advances in AI came from algorithm advances, not just raw hardware. Now if most advances came from algorithms and computer science, tell me that their army of AI researchers is not their fundamental advantage. And we see it. DeepSeek is not an inconsequential advance. And the day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation.
为什么?因为目前像 DeepSeek 这样的模型如果是开源的,可以在任何加速器上运行。为什么未来会不一样呢?
Why is that? Because currently you can have a model like DeepSeek that runs on any accelerator if it's open source. Why would that stop being the case in the future?
好吧,假设不是这样。假设它针对华为进行了优化。假设它针对他们的架构进行了优化。那会让我们处于劣势。你描述了一个我认为是好消息的情况:一家公司开发了一个 AI 模型,它在美国技术栈上运行得最好。我认为那是好消息。你把它当作坏消息的前提。我来告诉你坏消息:世界各地的 AI 模型被开发出来,它们在非美国硬件上运行得最好。这对我们来说是坏消息。
Well, suppose it doesn't. Suppose it's optimized for Huawei. Suppose it's optimized for their architecture. That would put us at a disadvantage. You described a situation that I perceived to be good news: that a company developed an AI model and it runs best on the American tech stack. I saw that as good news. You set it up as a premise that it was bad news. I'm going to give you the bad news: AI models around the world are developed and they run best on non-American hardware. That is bad news for us.
我想我只是没有看到证据表明存在这些巨大的差异,会阻止你切换加速器。美国实验室在所有云上运行他们的模型,在所有……
I guess I just don't see the evidence that there are these huge disparities that would prevent you from switching accelerators. American labs run their models across all clouds, across all...
证据就是。你拿一个针对英伟达优化的模型,试着在其他东西上运行。
The evidence. You take a model optimized for Nvidia and try to run it on something else.
但美国实验室确实这样做。
But American labs do that.
而且它们运行得并不更好。英伟达的成功就是完美的证据。事实是,AI 模型在我们的技术栈上创建,并在我们的技术栈上运行得最好。这有什么难以理解的?
And they don't run better. Nvidia's success is perfect evidence. The fact that AI models are created on our stack and run best on our stack. How is that illogical to understand?
我只是在观察。你看,Anthropic 的模型在 GPU 上运行,也在 Trainium 上运行,在 TPU 上运行。
I'm just looking. Look, Anthropic's models are run on GPUs, they're run on Trainium, they're run on TPUs.
要改变它需要大量的工作。
A lot of work has to go into it to change.
但去全球南方、去中东,一开箱就这样。如果所有 AI 模型都在别人的技术栈上运行得最好,那你现在就得提出一个荒谬的主张,说这对美国是好事。
But go to the global south, go to the Middle East, coming out of the box. If all of the AI models run best on somebody else's tech stack, you've got to be arguing some ridiculous claim right now that that's a good thing for United States.
但我想我不理解这个论点。如果中国公司先达到下一个神话,他们发现所有安全运行者先发布美国软件,但他们可以在英伟达硬件上做,然后运到全球南方。他们用英伟达硬件做。这怎么是好事?我是说,它在硬件上运行。
But I guess I don't understand the argument. Like if Chinese companies get to the next mythos first, they find that all the security runner releasing American software first, but they can do it on Nvidia hardware and they ship it to the global south. They do it on Nvidia hardware. How is that good? I mean, it runs on hardware.
这不是好事。
It's not good.
对吧?
Right?
这不是好事。所以我们别让它发生。
It's not good. So let's not let it happen.
你为什么认为如果你不给他们运电脑,他们就会完全被华为取代?他们落后,对吧?他们的芯片比你的差。
Why do you think it's perfectly fungible that if you didn't ship them computers, they would exactly be replaced by Huawei? They are behind, right? They have worse chips than you.
这完全……现在有证据。他们的芯片产业是巨大的。
It's completely... there's evidence right now. Their chip industry is gigantic.
你只要看看 H200 和华为 910C 之间的算力、带宽或内存比较。差不多是一半,一半。
You can just look at the flop or bandwidth or memory comparisons between the H200 and the Huawei 910C. It's like half, half.
他们用更多。他们用两倍的数量。
They use more of it. They use twice as many.
我想你的论点似乎是他们有这么多现成的能源,对吧?他们需要用芯片来填满它。
I guess it seems like your argument is they have all this energy that's ready to go, right? And they need to fill it with chips.
而且他们擅长制造。
And they're good at manufacturing.
而且我确信最终他们能制造超过所有人,但这有关键的几年。
And I'm sure eventually they would be able to just out manufacture everybody, but there's these few critical years.
你说的关键年份是什么?
What is the critical year you're talking about?
未来几年我们会有这些模型进行所有网络攻击。如果关键年份,未来关键年份是关键,那么我们必须确保世界上所有 AI 模型都建立在美国技术栈上。这些关键年份。
These next few years we've got these models that are going to do all the cyber attacks. If the critical years, the next critical years is critical, then we have to make sure that all of the world's AI models are built on American tech stack. These critical years.
好吧,如果它们建在美国技术栈上,那怎么能防止它们如果拥有更先进的能力,发起类似神话的网络攻击……
Okay, how would that prevent if they're built on American tech stack, how would that prevent them from if they have more advanced capabilities from launching the mythos equivalent cyber attacks on...
两种方式都没有保证。
There's no guarantee either way.
但如果你更早拥有它,我们可以做准备。
But if you have it earlier, we can prepare for it.
听着,你为什么让 AI 产业的一个层面失去整个市场,以便让另一个层面受益?有五个层面,每个层面都必须成功。最需要成功的层面实际上是 AI 应用。你为什么如此执着于那个 AI 模型,那一家公司?什么原因?因为那些模型使得这些极其攻击性的能力成为可能,而你需要算力、芯片、AI 研究者的生态系统使之成为可能。
Listen, why are you causing one layer of the AI industry to lose an entire market so that you could benefit another layer of the AI industry? There's five layers and every single layer has to succeed. The layer that has to succeed most is actually the AI applications. Why are you so fixated on that AI model, that one company? For what reason? Because those models make possible these incredibly offensive capabilities and you need computer energy, the chips, the ecosystem of AI researchers make it possible.
几个月前,Jane Street 花了大约 2 万 GPU 小时,将后门植入三个不同的语言模型。然后他们挑战我的观众找出触发短语。我刚刚联系了设计这个谜题的 Rickson,了解 Jane Street 收到的一些解决方案。如果你认为基础模型在这里,后门模型在这里,你可以线性插值权重来调整后门的强度,但也可以外推使后门更强。在某些情况下,如果你让它足够强,模型就会直接吐出响应短语应该是什么。所以,如果你不断放大基础版本和后门版本之间的差异,最终它应该会吐出触发短语。但这种技术只对三个模型中的两个有效。连 Ricken 也不确定为什么对另一个无效。能够验证一个模型只做你认为它做的事,是 AI 安全中最重要的开放问题之一。如果这类问题让你兴奋,Jane Street 正在招聘研究者和工程师。访问 janestreet.com/thorcash 了解更多。好了,退一步说,中国必须能够建造足够的 7 纳米产能。记住,他们仍然卡在 7 纳米,而你会推进到 3 纳米,然后 2 纳米或 1.6 纳米。所以当你在 1.6 纳米时,他们仍然在 7 纳米,他们必须生产足够的量来弥补短缺,而且他们有这么多能源,你给他们的芯片越多,他们拥有的算力就越多。所以最终的问题是,他们在训练和推理中获得了更多的算力输入。
A few months ago, Jane Street spent about 20,000 GPU hours trading back doors into three different language models. Then they challenged my audience to find the trigger phrases. I just caught up with Rickson who designed the puzzle about some of the solutions that Jane Street received. If you think the base model was here and the back door model was here, you can kind of linearly interpolate the weights to adjust the strength of the back door, but you can also extrapolate it to make the back door even stronger. And in some cases, if you make it strong enough, the model will just regurgitate what the response phrase was supposed to be. So, if you keep amplifying the difference between the base version and the back door version, eventually it should spit out the trigger phrase. But this technique only worked on two out of the three models. Even Ricken isn't sure why it didn't work on the other. Being able to verify that a model only does what you think it does is one of the most important open questions in AI security. If this is the kind of problem that excites you, Jane Street is hiring researchers and engineers. Go to janestreet.com/thorcash to learn more. Okay, stepping back, it has to be the case that China is able to build enough 7 nanometer capacity. And remember, they're still stuck on 7 nanometer while you will move on to 3 nanometer and then 2 nanometer or 1.6 nanometer with fineman. So while you're on 1.6 nanometer they're still going to be on 7 nanometer and they have to produce enough of it to make up for the shortfall and they have so much energy that the more chips you give them the more compute they'd have right like so I just there's it comes to the question of ultimately they are getting more compute as input to training and inference.
我只是觉得你说话太绝对。我认为美国应该领先。美国的算力是其他任何地方的 100 倍。美国应该领先。好吧,美国是领先的。英伟达制造最先进的技术。我们确保美国实验室最先听说并最先有机会购买。如果他们钱不够,我们甚至投资他们。美国应该领先。我们想尽一切办法确保美国领先。第一点。你同意吗?我们正在尽一切努力做到这一点。
I just think you speak in absolutes. I think that United States ought to be ahead. The amount of compute in United States is 100 times more than anywhere else in the world. The United States ought to be ahead. Okay, the United States is ahead. Nvidia builds the most advanced technologies. We make sure that the US labs are the first to hear about it and the first chance to buy it. And if they don't have enough money, we even invest in them. The United States ought to be ahead. We want to do everything we can to make sure the United States is ahead. Number one point. Do you agree? And we're doing everything we can to do that.
但把芯片运到中国怎么让美国保持领先?他们被禁了。我们有 Vera Rubin 给美国。现在,美国。我在美国吗?你认为我是美国的一部分吗?
But how is shipping chips to China keeping the US ahead? They're banned. We have Vera Rubin for United States. Now, United States. Am I in United States? Do you consider me part of the United States?
是的。
Yes.
英伟达,你认为英伟达是美国公司吗?好吧。第一,为什么我们不制定一个更平衡的法规,让英伟达能在全球取胜,而不是放弃世界?你为什么想让美国放弃世界?芯片产业是美国生态系统的一部分。它是美国技术领导力的一部分。它是 AI 生态系统的一部分。它是 AI 领导力的一部分。为什么?为什么你的政策、你的哲学导致美国放弃世界市场的很大一部分?
Nvidia, you consider Nvidia a United States company? Okay. Number one, why is it that we don't come up with a regulation that's more balanced so that Nvidia can win around the world instead of giving up the world? Why would you want United States to give up the world? The chip industry is part of the American ecosystem. It's part of American technology leadership. It's part of the AI ecosystem. It's part of AI leadership. Why? Why is it that your policy, your philosophy leads to United States giving up a vast part of the world's market?
这里的说法是 Alfred Dario 有句话,他说这就像波音吹嘘我们在向朝鲜出售核武器,但导弹外壳是波音制造的,这 somehow 增强了美国技术栈。就像你从根本上给了他们这种能力。
The claim here is Alfred Dario had this quote where he said it's like Boeing bragging that we're selling North Korea nukes but the missile casings are made by Boeing and that's somehow enabling the US technology stack. Like fundamentally you're giving them this capability.
把 AI 和你刚才提到的任何东西相比是疯狂的。
Comparing AI to anything that you just mentioned is lunacy.
但 AI 类似于浓缩铀,对吧?它可能有正面用途,也可能有负面用途。我们仍然不想把浓缩铀送到其他国家。
But AI is similar to enriched uranium, right? And then it can have positive uses, it can have negative uses. We still don't want to send enriched uranium to other countries.
谁在送浓缩铀?
Who's sending enriched uranium?
这个类比是浓缩铀。
The analogy is enriched uranium.
因为这是个糟糕的类比,不合逻辑的类比。但如果那台电脑能运行一个模型,可以对所有美国软件进行零日漏洞利用,那怎么不是武器?
Because it's a lousy analogy, it's an illogical analogy. But if that computer can run a model that can do zero day exploits against all American software, how is that not a weapon?
首先,解决这个问题的方法是跟研究者对话,跟中国对话,跟其他国家对话,确保人们不以那种方式使用技术。这是必须进行的对话。第一点。
First of all, the way to solve that problem is to have dialogues with the researchers and dialogues with China and dialogues with other countries to make sure that people don't use technology in that way. That's a dialogue that has to happen. Number one.
第二,我们还需要确保美国领先。Ruben Vera Rubin Blackwell 的一切在美国都很充裕。堆积如山。显然,我们的结果会证明这一点。大量的算力。我们有很棒的 AI 资源。这很好。我们必须保持领先。然而,我们也必须认识到 AI 不仅仅是一个模型。AI 是一个五层蛋糕。AI 产业在每一层都很重要。我们希望美国在每一层都获胜,包括芯片层。放弃整个市场不会让美国在计算堆栈的芯片层长期赢得技术竞赛。这是事实。
Number two, we also need to make sure that the United States is ahead. Everything that Ruben Vera Rubin Blackwell is available in the United States in abundance. Mounds of it. Obviously, our results would show it. Abundance of tons of it. Tons of it. The amount of computing we have is great. We have amazing AI resources here. It's great. We have to stay ahead. However, we also have to recognize that AI is not just a model. That AI is a five-layer cake. That AI industry matters across every single layer. And we want the United States to win at every single layer, including the chip layer. And conceding the entire market is not going to allow the United States to win the technology race long-term in the chip layer in the computing stack. That is just a fact.
那么关键就在于,现在向他们出售芯片如何帮助我们长期获胜?就像特斯拉长期向中国销售非常优秀的电动汽车。iPhone 在中国销售,非常优秀。它们并没有导致某种锁定。中国仍然会制造他们自己的电动汽车,并且他们正在主导,或者智能手机主导。
I guess then the crux comes down to how does selling them chips now help us win in the long term. Like Tesla sold extremely good electric vehicles to China for a long time. iPhones are sold in China, extremely good. They didn't cause some lock-in. China will still make their version of EVs and they're dominating or smartphones dominating.
当我们今天开始对话时,你会承认,并且你确实承认了 Nvidia 的地位非常不同。你用了像「护城河」这样的词。对我们公司来说,最重要的事情是我们生态系统的丰富性,这关乎开发者。50% 的 AI 开发者在中国。我们不想,我们不应该,美国不应该放弃这一点。但我们在美国有很多 Nvidia 开发者,这并不妨碍美国实验室将来也能使用其他加速器。事实上,现在他们也在使用其他加速器,这很好。我不明白为什么如果你向他们出售 Nvidia 芯片,中国就不会出现同样的情况,就像 Google 可以使用 TPU 和 Nvidia 一样。
When we started the conversation today, you would acknowledge and you acknowledged that Nvidia's position is very different. You use words like moat. The single most important thing to our company is our richness of our ecosystem which is about developers. 50% of the AI developers are in China. We don't want to, we shouldn't, the United States should not give that up. But we have a lot of Nvidia developers in the US and that doesn't prevent American labs from also being able to use other accelerators in the future. In fact, right now they're using other accelerators as well, which is fine and great. I don't see why that wouldn't be the case in China as well if you sell them Nvidia chips, just the same way that Google can use TPUs and Nvidia.
我们必须持续创新,而且你可能知道,我们的份额在增长,而不是减少。那种即使我们在中国竞争,我们也会失去那个市场的前提。我没有,你不是在跟一个醒来就是失败者的人说话。那种失败者的态度,失败者的前提对我来说毫无意义。我们不是一辆车。我们不是一辆车。我可以今天买这个牌子的车,明天用另一个牌子的车。很容易。计算不是这样的。x86 仍然存在是有原因的。ARM 如此有粘性是有原因的。这些生态系统很难被取代。这需要大量的时间和精力,大多数人不想这么做。所以我们的工作是继续培育这个生态系统,不断推进技术,这样我们才能在市场上竞争。基于你描述的前提放弃一个市场,我根本无法认同。这毫无意义,因为我不认为美国是失败者。我们的行业不是失败者。那个失败的主张,失败的心态对我来说毫无意义。
We have to keep innovating and, as you probably know, our share is growing, not decreasing. The premise that even if we competed in China, we're going to lose that market anyways. I don't, you're not talking to somebody who woke up a loser. And that loser attitude, that loser premise makes no sense to me. We are not a car. We are not a car. The fact that I can buy a car, this car brand one day and use another car brand another day. Easy. Computing is not like that. There's a reason why the x86 still exists. There's a reason why ARM is so sticky. These ecosystems are hard to replace. It costs an enormous amount of time and energy and most people don't want to do it. And so it's our job to continue to nurture that ecosystem, to keep advancing the technology so that we could compete in the marketplace. Conceding a marketplace based on the premise you described, I simply can't acknowledge that. It makes no sense because I don't think the United States is a loser. Our industry is not a loser. And that losing proposition, that losing mindset makes no sense to me.
好的,我换个话题。我只是想确保……
Okay, I'll move on. I just want to make sure...
你不必换话题。我很享受。
You don't have to move on. I'm enjoying it.
好的,很好。那我感谢你。但我想也许关键,谢谢你跟我绕圈子,因为我觉得这有助于揭示关键所在。
Okay, great. Then I appreciate that. But I think the maybe the crux, and thanks for walking around the circles with me, because I think it helps bring out what the crux here is.
关键是你走向了极端。你的论点从极端开始,认为如果我们在这个狭窄的时刻给他们任何算力,我们就会失去一切。
The crux is you're going to extremes. Your argument starts from extremes that if we give them any compute at all in this narrow moment, we will lose everything.
不,我认为我的论点是……
No, I think what my argument is...
那些极端,很幼稚。
Those extremes, they're childish.
想法不是存在某个关键的算力阈值,而是任何边际算力都有帮助,对吧?所以如果你有更多算力,你可以训练更好的模型。
The idea is not that there is some key threshold of compute, it's that any marginal compute is helpful, right? So if you have more compute, you can train a better model.
我只是想让你承认,对美国科技行业来说,任何边际销售都是有益的。
And I just want you to acknowledge that any marginal sales for the American technology industry is beneficial.
我其实不这么认为。我的意思是,如果运行在这些芯片上的 AI 模型……
I actually don't. I mean, if the AI models that run on those chips...
嗯。
Yeah.
……具备网络攻击能力,或者训练模型具备网络防御能力,在这些实例上运行更多模型。它不是核武器,但它促成了一种武器。
...are capable of cyber offensive capabilities, or training models are capable of cyber defense, running more models at those instances. It is not a nuclear weapon, but it enables a weapon of a kind.
你用的逻辑,你大可以把它用在微处理器和 DRAM 上。你大可以把它用在电上。
The logic that you use, you might as well say it to microprocessors and DRAMs. You might as well say it to electricity.
但事实上,我们确实对制造最先进 DRAM 的相关技术有出口管制,对吧?我们对中国的各种运输有各种出口管制。
But in fact, we do have export controls on the technology that is relevant to making the most advanced DRAM, right? We have all kinds of export controls on China for all kinds of shipping.
我们向中国出售大量 DRAM 和 CPU。我认为这是对的。
We sell a lot of DRAM and CPUs into China. And I think it's right.
我想这回到了根本问题:AI 是否不同,对吧?如果你拥有那种能在软件中发现零日漏洞的技术,我们是否要最小化中国先达到、领先的能力?
I guess this goes back to the fundamental question of is AI different, right? If you have the kind of technology that can find these zero days in software, is that something where we want to minimize China's ability to get there first, to be ahead?
我们可以控制那个。
We can control that.
如果芯片已经在那里,他们用它们来训练那个模型,我们如何控制?
How do we control that if the chips are already there and they're using that to train that model?
我们有大量的算力。我们有大量的 AI 研究人员。我们正在尽可能快地竞赛。
We have tons of compute. We have tons of AI researchers. We're racing as fast as we can.
再说一次,我们拥有的核武器比任何人都多,但我们不想把浓缩铀送到任何地方。
Again, we have more nuclear weapons than anybody else, but we don't want to send enriched uranium anywhere.
我们不是浓缩铀。这是一个芯片,而且是一个他们自己能制造的芯片。
We're not enriched uranium. It's a chip and it's a chip that they can make themselves.
但他们从你这里买是有原因的,对吧?我们有中国公司创始人的引述,说我们在封锁那项技术……
But there's a reason they're buying it from you, right? And we have quotes from the founders of Chinese companies that say that we're bottling that technology...
因为我们的芯片更好。总的来说,我们的芯片更好。这毫无疑问。没有我们的芯片,没有我们的芯片,你能承认华为创下了纪录的一年吗?你能承认一大批芯片公司已经上市了吗?你能承认吗?
Because our chips are better. On balance, our chips are better. There's just no question about it. In the absence of our chip, in the absence of our chip, can you acknowledge that Huawei had a record year? Can you acknowledge that a whole bunch of chip companies have gone public? Can you acknowledge that?
你能承认吗?你也能承认我们曾经在那个市场拥有非常大的份额,而现在不再拥有大份额这个事实吗?我们也可以承认中国约占世界科技产业的 40%。那个市场,放弃那个市场,为美国科技产业放弃那个市场,是对我们国家的损害。是对我们国家安全的损害。是对我们技术领导地位的损害。全部为了一个公司的利益。这对我来说毫无意义。我想我搞糊涂了。感觉你在说两个不同的陈述。一个是,如果允许我们竞争,我们将赢得与华为的竞争,因为我们的芯片会好得多。另一个是,没有我们,他们反正也会做同样的事情。对吧?这两件事怎么能同时成立?
Can you acknowledge that? Can you also acknowledge that the fact that we used to have a very large share in that market and we no longer have the large share in that market. We can also acknowledge that China is about 40% of the world's technology industry. That market, to leave that market, concede that market for the United States technology industry is a disservice to our country. It is a disservice to our national security. It is a disservice to our technology leadership. All for the benefit of one company. It makes no sense to me. I guess I'm confused. It feels like you're making two different statements. One is that we're going to win this competition with Huawei because our chips are going to be way better if we're allowed to compete. And another is that they would be doing the same exact thing without us anyways. Right? How can those two things be true at the same time?
这显然是真的。在没有更好选择的情况下,你会选择你唯一的选择。这怎么会不合逻辑?这非常合逻辑。
It's obviously true. In the absence of a better choice, you'll take the only choice you have. How is that illogical? It's so logical.
他们想要 Nvidia 芯片的原因是它们更好。
The reason they want Nvidia chips is they're better.
更好就是更多算力。更多算力意味着你可以训练出更好的模型。更好是因为它更容易编程。我们有更好的生态系统。无论更好是什么。当然我们会把算力给他们。那又怎样?事实是我们获得了利益。别忘了我们获得了美国技术领导力的利益。我们获得了开发者在美国技术栈上工作的利益。当这些 AI 模型扩散到世界其他地方时,我们也获得了利益。因此美国技术栈是最好的。我们可以继续推进和扩散美国技术。我相信这是积极的。这是美国技术领导力非常重要的一部分。现在你主张的政策导致美国电信行业基本上被排挤出世界市场,以至于我们不再控制自己的电信。我不认为这是明智的。这有点狭隘,并且导致了意想不到的后果,我现在正在向你描述,而你似乎很难理解。
Better is more compute. More compute means you can train a better model. It's better because it's easier to program. We have a better ecosystem. Whatever the better is. And of course we're going to send them compute. So what? The fact of the matter is we get the benefit. Don't forget we get the benefit of American technology leadership. We get the benefit of developers working on the American tech stack. We get the benefit as those AI models diffuse out into the rest of the world. The American tech stack is therefore the best for it. We can continue to advance and diffuse American technology. I believe that is a positive. It's a very important part of American technology leadership. Now the policy that you're advocating resulted in the American telecommunication industry being policied out of basically the world to the point where we don't control our own telecommunications anymore. I don't see that as smart. It's a little narrow-minded and it led to unintended consequences that I'm describing to you right now that you seem to have a very hard time understanding.
好的,我们先退一步。看起来核心在于有潜在收益和潜在成本,我们试图弄清楚收益是否值得成本。我想让你承认潜在成本:算力是训练强大模型的输入。强大模型确实有强大的攻击能力,比如网络攻击。美国公司首先获得神话级能力是好事,然后现在他们将推迟这些能力,以便美国公司和美国政府能在这一能力水平公布之前更好地保护他们的软件。如果中国有更多算力,如果我们能更早制造出神话级模型并广泛部署,那将非常糟糕。这没有发生的原因之一是我们有更多算力,这要感谢像英伟达这样的美国公司。这就是向中国输送算力的成本。所以暂时把收益放在一边。你承认这是一个潜在成本吗?
Okay, let's just step back. It seems like the crux here is there's a potential benefit and there's a potential cost and we're trying to figure out is the benefit worth the cost. I guess I'm trying to get you to acknowledge the potential cost that compute is an input to training powerful models. Powerful models do have powerful offensive capabilities like cyber attacks. It is a good thing that American companies got to claim mythos level capabilities first and then now they're going to hold off on those capabilities so that the American companies and American government can make their software more protected before this level cap announced if China had had more compute, if we could have had made a mythos level model earlier and deployed it widely that would have been very bad. One of the reasons that hasn't happened is that we have more compute thanks to companies like Nvidia in America. That is a cost of sending to China. So let's leave the benefit aside for a second. Do you acknowledge that this is a potential cost?
我还要告诉你,潜在成本是我们允许 AI 堆栈中最重要的层之一——芯片层——放弃整个市场,世界第二大市场,这样他们就能发展规模,建立自己的生态系统,使得未来的 AI 模型以与美国技术栈截然不同的方式优化。随着 AI 扩散到世界其他地方,他们的标准、他们的技术栈将变得比我们优越,因为他们的模型是开放的。
I will also tell you the potential cost is we allow one of the most important layers of the AI stack, the chip layer, to concede an entire market, the second largest market in the world, so that they could develop scale so that they could develop their own ecosystem so that future AI models are optimized in a very different way than the American tech stack. As AI diffuses out into the rest of the world, their standards, their tech stack will become superior to ours because their models are open.
我想我只是足够相信英伟达的内核工程师和 CUDA 工程师,认为他们能够优化。
I guess I just believe enough in Nvidia's kernel engineers and CUDA engineers to think that they could optimize.
如你所知,AI 不仅仅是内核优化。
AI is more than kernel optimization as you know.
当然,但你可以做很多事情,从蒸馏到适配你芯片的模型。
Of course, but there's so many things you can do from distilling to a model that's well fit for your chips.
我们会尽最大努力。
We're going to do our best.
你们有所有这些软件。我只是很难想象会长期锁定在中国生态系统。他们可能暂时有稍微好一点的开源模型。
You have all this software. I just find it hard to imagine that there's a long-term lock-in to Chinese ecosystem. They have this like slightly better open source model for a while.
中国是世界上开源软件的最大贡献者。事实,对吧?中国是世界上开放模型的最大贡献者。事实。今天它建立在美国技术栈上,事实。AI 技术栈的所有五个层都很重要。美国应该去赢得所有五个层。它们都很重要。当然最重要的是 AI 应用层。那个扩散到社会的层,使用它最多的层将从这场工业革命中获益最多。但我的观点是每一层都必须成功。如果我们吓唬这个国家,让 AI 看起来像核弹,以至于每个人都讨厌 AI、害怕 AI,我不知道你如何帮助美国,你是在帮倒忙。如果我们吓唬所有人,让他们不做软件工程工作,因为 AI 会消灭所有软件工程工作,结果我们没有软件工程师,我们是在帮倒忙。如果我们吓唬所有人,让他们不做放射科医生,因为计算机视觉完全免费,没有 AI 会比放射科医生做得更差。我们误解了工作和任务的区别。放射科医生的工作是病人护理,任务是阅读扫描。如果我们如此深刻地误解这一点,并吓唬所有人不去上放射科学校,我们将没有足够的放射科医生和足够好的医疗保健。所以我要说的是,当你做出如此极端的假设时,一切都从零到无穷。我们最终以不真实的方式吓唬人。生活不是那样的。我们是否希望美国第一?当然。我们是否需要在该堆栈的每一层都成为领导者?当然。今天你在谈论神话级,因为神话级很重要。当然。那很棒。但几年后,我预测,当我们希望美国技术栈、美国技术扩散到世界各地,到印度、中东、非洲、东南亚,当我们的国家想要出口,因为我们想要出口我们的技术、我们的标准。那一天,我希望你和我再次进行同样的对话。我会准确地告诉你今天的对话,你的政策和你的想象如何导致美国毫无理由地放弃了世界第二大市场。我们不应该放弃它。如果我们失去它,那是失去。但我们为什么要放弃它?现在,没有人主张全有或全无。没有人主张全有或全无,意思是随时把所有东西运到中国。没有人主张我们应该总是把最好的技术留在这里。我们应该总是拥有最多的技术和最先的技术。但我们也应该尝试在世界范围内竞争并获胜。这两件事可以同时发生。这需要一些细微差别,一些成熟度,而不是绝对化。世界不是绝对的。
China is the largest contributor to open source software in the world. Fact, right? China is the largest contributor to open models in the world. Fact. Today it's built on the American tech stack and fact. All five layers of the tech stack for AI is important. United States ought to go win all five of them. They're all important. The one that is the most important of course is the AI application layer. The layer that diffuses into society, the one that uses it most will benefit from this industrial revolution most. But my point is that every layer has to succeed. If we scare this country into thinking that AI is somehow a nuclear bomb so that everybody hates AI and everybody's afraid of AI, I don't know how you're helping the United States, you're doing a disservice. If we scare everybody out of doing software engineering jobs because it's going to kill every software engineering job and we don't have any software engineers as a result of that, we're doing a disservice to United States. If we scare everybody out of radiology, so nobody wants to be a radiologist because computer vision is completely free and no AI is going to do a worse job than a radiologist. And we misunderstand the difference between a job and the task. The job of a radiologist is patient care, task is to read a scan. If we misunderstand that so profoundly and we scare everybody out of going to radiology school, we're not going to have enough radiologists and good enough healthcare. So I'm making the case that when you make a premise that is so extreme, everything goes from zero or infinity. We end up scaring people in a way that's just not true. Life is not like that. Do we want United States to be first? Of course we do. Do we need to be a leader in every layer of that stack? Of course we do. Is today you're talking about mythos because mythos is important. Sure. That's fantastic. But in a few years time, I'm making you the prediction that when we want the American tech stack, when we want American technology to be diffused around the world, out to India, out to the Middle East, out to Africa, out to Southeast Asia, when our country would like to export because we would like to export our technology, we would like to export our standards. On that day, I want you and I to have that same conversation again. And I will tell you exactly about today's conversation about how your policy and what you imagined literally cause the United States to concede the second largest market in the world for no good reason at all. We shouldn't concede it. If we lose it, we lose it. But why do we concede it? Now, nobody is advocating all or nothing. Nobody is advocating all or nothing, meaning we ship everything to China at all times. Nobody is advocating that we should always have the best technology here. We should always have the most technology here and the first. But we should also try to compete and win around the world. Both of those things can simultaneously happen. It requires some amount of nuance, some amount of maturity instead of absolutes. The world is just not absolutes.
好的。
Okay.
这个论点基于他们构建了针对其架构优化的模型,也就是他们几年后制造的最好的芯片,这些芯片出口到世界各地,这设定了一个标准,因为 EUV 出口管制。正如我们所说,你将推进到 1.6 纳米,即使几年后,他们仍将停留在 7 纳米。在国内,他们可能会倾向于说,嘿,我们有这么多能源,可以大规模制造,我们会继续使用 7 纳米。但出口方面,他们的 7 纳米芯片必须与你的 1.6 纳米芯片竞争,他们的模型必须针对 7 纳米进行深度优化,以至于在 7 纳米上运行他们的模型比在你的 1.6 纳米上运行更好。
The argument hinges on they've built models that are specified for their architecture, the best chips that they make in a few years, and those chips get exported around the world. That sets a standard because of EUV export controls. As we said, you're going to move on to 1.6 nanometer, there's still going to be on 7 nanometer even after a few years from now, and it might make sense that domestically they would prefer, hey, we got so much energy, we can manufacture at scale, we'll still keep using 7 nanometer. But the exporting thing, their 7 nanometer chips have to be competitive against your 1.6 nanometer chips, and their models have to be so far optimized for the 7 nanometer, it's better to run their models on 7 nanometer than to run their models on your 1.6 nanometer.
那我们能不能看看事实?Blackwell 的光刻技术比 Hopper 先进 50 倍吗?有 50 倍吗?差远了。我一直在说,摩尔定律已死。从晶体管本身来看,Hopper 到 Blackwell 的进步,大概 75%。相隔 3 年,75%。Blackwell 是 Hopper 的 50 倍。我的重点是架构很重要,计算机科学很重要,半导体物理也很重要,但计算机科学更重要。AI 的影响很大程度上来自计算栈,这就是为什么 CUDA 如此有效,为什么 CUDA 如此受欢迎。它是一个生态系统,一个计算架构,提供了极大的灵活性,如果你想完全改变架构,创建像扩散模型这样的东西,或者创建解耦的东西,你都可以做到,很容易。所以事实是,AI 既关乎上层的栈,也关乎下层的架构。当我们拥有针对我们的栈和生态系统优化的架构和软件栈时,这显然很好,因为我们今天一开始就谈到 Nvidia 的生态系统多么丰富,为什么人们总是喜欢先在 CUDA 上编程。他们确实如此,中国的研究人员也一样。但如果我们被迫离开中国,那将是一个政策错误。显然,这会有反噬。显然,这对美国来说结果很糟糕。它推动并加速了他们的芯片产业,迫使他们的整个 AI 生态系统专注于内部架构。现在还不算太晚,但无论如何,这已经发生了。未来你会看到他们不会停留在 7 纳米。显然他们擅长制造,他们会从 7 纳米继续前进。5 纳米和 7 纳米之间有 10 倍的差距吗?答案是没有。架构很重要,网络很重要,这就是为什么 Nvidia 收购了 Mellanox。网络很重要,能源很重要。所有这些都很重要,不像你试图简化的那样简单。
Can we just look at the facts then? Is Blackwell 50 times more advanced lithography than Hopper? Is it 50 times? Not even close. I just kept saying it over and over again. Moore's law is dead. Between Hopper and Blackwell from the transistors themselves, call it 75%. It was 3 years apart. 75%. Blackwell is 50 times Hopper. My point is architecture matters. Computer science matters. Semiconductor physics matter as well. But computer science matters. The impact of AI largely comes from the computing stack, which is the reason why CUDA is so effective, which is the reason why CUDA is so beloved. It's an ecosystem, a computing architecture that allows for so much flexibility that if you wanted to change an architecture completely, create something like diffusion, create something that's disaggregated, you could do so. It's easy to do. And so the fact of the matter is AI is about the stack above as much as it is about the architecture below. To the extent that we have architectures and software stacks optimized for our stack, for our ecosystem, it is obviously good because we started the conversation today about how Nvidia's ecosystem is so rich, why people always love programming on CUDA first. They do. And so do the researchers in China. But if we are forced to leave China, it would be a policy mistake. Obviously, it has backlash. Obviously, it has turned out badly for the United States. It enabled and accelerated their chip industry. It forced all of their AI ecosystem to focus on their internal architectures. It's not too late, but nonetheless, it has already happened. You're going to see in the future they're not stuck at 7 nanometer. Obviously they're good at manufacturing. They will continue to advance from seven and beyond. Now is there a 10x difference between 5 nanometer and 7 nanometer? The answer is no. Architecture matters. Networking matters. That's why Nvidia bought Mellanox. Networking matters. Energy matters. And so all that stuff matters. It's not simplistic like the way you're trying to distill it.
我们可以不谈中国了,但这引出了一个有趣的问题,关于台积电和内存等的瓶颈。如果我们处在一个你已经占据 N3 大部分产能,未来也会占据 N2 大部分产能的世界里,你是否认为可以回到 N7,利用旧工艺节点的闲置产能,说,嘿,AI 需求如此巨大,而我们扩展前沿工艺的能力跟不上,所以我们要基于今天对数值计算的所有了解以及你描述的其他改进,来制造 Hopper 或 Ampere?你认为这种情况会在 2030 年之前发生吗?
We can move on from China, but that actually raises an interesting question about the bottlenecks at TSMC and memory and so forth. If we're in this world where you're already the majority of N3, at some point you'll be N2, you'll be a majority of that. Do you see that you could go back to N7, this spare capacity at an older process node, and say, hey, the demand for AI is so great and our capacity to expand the leading edge is not meeting it, so we're going to make a Hopper or Ampere about everything we know about a numeric today and all the other improvements you described. Do you see that world happening within before 2030?
没有必要,原因是每一代架构不仅仅是晶体管尺寸的问题。你还做了大量的工程、封装、堆叠、数值计算和系统架构。当你无法轻易回到另一个节点时,那是一种没人能负担得起的研发水平。我们可以承受向前推进,但我不认为我们能承受倒退。现在,如果世界说,那一天,让我们做个思想实验。那一天,我们说,听着,我们再也不会有更多产能了,我会毫不犹豫地回去用 7 纳米吗?是的,我当然会。
It's not necessary, and the reason for that is because with every generation, the architecture is more than just the transistor scale. You're also doing so much engineering and packaging and stacking and the numeric and the system architecture. When you run out of capacity to easily go back to another node, that's a level of R&D that no one could afford. We could afford to lean forward. I don't think we could afford to go back. Now, if the world simply says, on that day, let's do the thought experiment. On that day, we go, listen, we're just never going to have more capacity ever again, would I go back and use seven in a heartbeat? Yeah, of course I would.
我和某人聊天时有一个问题,为什么 Nvidia 不同时运行多个完全不同架构的芯片项目?你可以做 Cerebras 风格的晶圆级芯片,可以做 Dojo 风格的大封装,也可以做一个没有 CUDA 的。你有资源和工程人才并行做所有这些。那么,既然没人知道 AI 和架构会走向何方,为什么要把所有鸡蛋放在一个篮子里呢?
One question somebody I was talking to had is why Nvidia doesn't run multiple different chip projects at the same time with totally different architectures. So you could do a Cerebras style wafer scale, you could do a Dojo style huge package, you could do one without CUDA. You have the resources and the engineering talent to do all these in parallel. So why put all the eggs in one basket given who knows where AI might go and architectures might go?
哦,我们可以。只是我们没有更好的想法。是的,我们可以做所有这些事情,只是它们并不更好。我们模拟了所有方案,它们在模拟器中被证明更差,所以我们不会去做。我们正在做的正是我们想做的项目。如果工作负载发生巨大变化——我不是指算法,而是指工作负载本身,这取决于市场形态——我们可能会决定添加其他加速器。例如,最近我们加入了 Grok,并打算将其融入我们的 CUDA 生态系统。我们现在正在这样做,因为 token 的价值已经变得非常高,以至于你可以对 token 进行不同定价。在过去,就在几年前,token 要么免费,要么几乎不贵。但现在你可以有不同客户,这些客户想要不同的答案。而且,因为客户赚了很多钱,比如我们的软件工程师,如果我能给他们响应更快的 token,让他们比现在更高效,我愿意为此付费。但这个市场是最近才出现的。所以我认为我们现在有能力基于响应时间将同一模型划分为不同细分市场,这就是为什么我们决定扩展前沿,创建一个推理细分市场,提供更快的响应时间,尽管吞吐量较低。直到现在,更高的吞吐量总是更好。我们认为可能存在一个世界,token 的平均售价非常高,即使工厂吞吐量较低,平均售价也能弥补。这就是我们这样做的原因。
Oh, we could. It's just that we don't have a better idea. Yeah, we could do all of those things. It's just not better. And we simulate it all. They're in our simulator provably worse, and so we wouldn't do it. Yeah, we're working on exactly the projects that we want to work on. And if the workload were to change dramatically, and I don't mean the algorithms, I actually mean the workload, and that depends on the shape of the market, we may decide to add other accelerators. For example, recently we added Grok, and we're going to fold Grok into our CUDA ecosystem. We're doing that now because the value of tokens has gone up so high that you could have different pricing of tokens. Back in the old days, just a couple years ago, tokens were either free or barely expensive. But now you can have different customers and those customers want different answers. And so, because the customers make so much money, for example, our software engineers, if I can give them much more responsive tokens so that they're even more productive than they are today, I would pay for it. But that market has only recently emerged. And so I think that we now have the ability to have the same model based on the response time have different segments, and that's the reason why we decided to expand the frontier and create a segment of inference that is faster response time even though it's lower throughput. Until now, higher throughput is always better. We think that there could be a world where there could be very high ASP tokens, and even though the throughput is lower in the factory, the ASPs make up for it. That's the reason why we did it.
但从架构角度来看,我认为英伟达的架构是……我更愿意把更多资金投入到架构上。我认为这种极其优质的词元以及推理市场解耦的想法非常有趣。
But otherwise from an architecture perspective, I think Nvidia's architecture is... I would rather put more money behind the architecture. I think this idea of extremely premium tokens and just the disaggregation of the inference market is very interesting.
最后一个问题:假设深度学习革命没有发生。英伟达会做什么?显然是游戏,但考虑到……
Final question: supposed deep learning revolution didn't happen. What would Nvidia be doing? Obviously games but given...
加速计算。我们一直在做同样的事情。我们公司的前提是,摩尔定律和通用计算对很多事情有好处,但对很多计算来说并不理想。所以我们把一种叫做 GPU 的架构与 CUDA 结合到 CPU 上,从而加速 CPU 的工作负载。不同的代码内核或算法可以卸载到我们的 GPU 上,结果应用程序的速度提升了 100 倍、200 倍。这可以用在哪里?工程、科学、物理、数据处理、计算机图形学、图像生成,各种领域。即使今天没有 AI,英伟达也会非常非常大。原因很根本:通用计算持续扩展的能力已经基本走到尽头。而实现扩展的方式是通过特定领域的加速。我们最初涉足的领域之一是计算机图形学,但还有很多其他领域:粒子物理、流体、结构化数据处理,各种受益于 CUDA 的算法。我们的使命是将加速计算带给世界,推动通用计算无法实现的应用,并扩展到足以突破某些科学领域的水平。早期的一些应用包括分子动力学、能源勘探的地震数据处理、图像处理。所有这些领域,通用计算都太低效了。所以如果没有 AI,我会很伤心。但由于我们在计算方面取得的进步,我们使深度学习民主化了。我们让任何研究人员、科学家或学生都能使用 PC 或 GeForce 显卡进行出色的科学研究。这个基本承诺一点都没有改变。如果你看 GTC,整个开头部分都不是 AI。计算光刻、量子化学、数据处理,所有这些都与 AI 无关,但仍然非常重要。AI 非常有趣且令人兴奋,但有很多人在做与 AI 无关的重要工作。张量并不是唯一的计算方式。我们想帮助所有人。
Accelerated computing. The same thing we've been doing all along. The premise of our company is that Moore's law and general purpose computing is good for a lot of things, but for a lot of computation it's not ideal. So we combined an architecture called a GPU with CUDA to a CPU so that we can accelerate the workload of the CPU. Different kernels of code or algorithms could be offloaded onto our GPU, and as a result you speed up an application by 100x, 200x. Where can you use that? Well, engineering, science, physics, data processing, computer graphics, image generation. All kinds of things. Even if AI doesn't exist today, Nvidia would be very, very large. The reason for that is fairly fundamental: the ability for general purpose computing to continue to scale has largely run its course. The way to do that is through domain specific acceleration. One of the domains we started with was computer graphics, but there are many other domains: particle physics, fluids, structured data processing, all kinds of algorithms that benefit from CUDA. Our mission was to bring accelerated computing to the world and advance applications that general purpose computing can't do, and scale to a level that helps break through certain fields of science. Some of the early applications were molecular dynamics, seismic processing for energy discovery, image processing. All those fields where general purpose computing is simply too inefficient. So if there was no AI, I would be very sad. But because of the advances we made in computing, we democratized deep learning. We made it possible for any researcher, scientist, or student to access a PC or a GeForce card and do amazing science. That fundamental promise hasn't changed, not even a little bit. If you watch GTC, the whole beginning part of it, none of it is AI. That part with computational lithography, quantum chemistry, data processing, all of that is unrelated to AI and still very important. AI is very interesting and exciting, but there are a lot of people doing important work that is not AI related. Tensors is not the only way you compute. We want to help everybody.
它没有。非常感谢。
It doesn't. Thank you so much.
不客气。我很享受。我也是。太好了。
You're welcome. I enjoyed it. Me too. Sweet.