Google's AI Infrastructure Chief on the Physics and Economics of Frontier AI
打开互动全文版(中英对照 + 朗读 + 问答)→谷歌 AI 基础设施负责人 Amin Vahdat 阐释 AI 数据中心如何与硬件协同定制设计,以及为何 FLOPS 只是虚荣指标。
Amin Vahdat, Google's head of AI infrastructure, explains how AI data centers are purpose-built and co-designed with hardware, and why FLOPS is a vanity metric.
在硬件方面,如你所知,有一个机会:你越针对特定工作负载进行专门化,灵活性就越低,但硬件会更快、更节能。所以这是一门艺术,也是对你设计目标以及该工作负载持续性的预测。换句话说,如果它在一两个月或三个月后就会消失,即使在那三个月里规模很大,你拦截它的窗口也非常窄。因此它必须有一定持久性。而且你必须能够准确预测专门化能带来多少收益。
In hardware again as you know there is this opportunity where the more you specialize to a particular workload the less flexible it is the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload. In other words if it's going to go away after a month or two months or 3 months even if it's big for those three months you got a really narrow window to intercept it. So it has to be somewhat durable. And you have to be able to project exactly what win you can get for specializing to it.
非常高兴欢迎 Amin Vad 来到节目。Amin,感谢你今天加入我们。
Thrilled to welcome Amin Vad to the show. Amin thank you for joining us for today.
能来这里真的很兴奋。我非常期待。
It's really exciting to be here. I'm really looking forward to it.
我对今天的话题非常兴奋,因为我们正处于人类历史上最大的资本支出建设之中。你身处其中。仅谷歌今年预计就将在资本支出上花费超过 2000 亿美元,其中大部分用于建设数据中心。你处于这一切的中心。你在去年年底被任命为谷歌 AI 基础设施负责人。所以你是人类历史上最资本密集型建设之一的领军人物,因此我非常期待今天与你深入探讨。
I am very excited for today's topic because we are in the middle of the biggest capex buildout in human history. You are in the middle of it. Google alone is expected to spend more than $200 billion on capex this year. Most of which is going into building data centers. You're at the center of it all. You were named the head of Google's AI Infra at the end of last year. So you are the man spearheading the efforts of one of the most capital intensive buildouts in human history and so I am really excited to get into it with you today.
这无疑是巨大的一年,当然整个行业包括谷歌都是如此。老实说,我不认为我们见过类似的事情,当然在谷歌肯定没有,但正如你所说,我认为在人类历史上,就建设规模和转型速度而言,真的难以置信。在我们深入之前,也许先给观众做个速成课。什么是 AI 数据中心,它与非 AI 数据中心有何不同?
It's a huge year for sure of course across the industry including at Google really have I don't think we've ever seen anything like this honestly certainly not at Google but as you said I think in the history of humanity in terms of the buildout and the pace of transformation really unbelievable. Before we get into it, maybe just just crash course for the audience. What is an AI data center and how how is it different from a non-AI data center?
这是一个很好的问题,因为 AI 数据中心和非 AI 数据中心实际上有很多相似之处,它们并没有根本性的不同。我的意思是,它由混凝土组成,一个外壳,包括电气场、机械场、冷却系统,你有一排排的电力分配。那里有大量的网络基础设施。换句话说,我们将大量计算资源相互连接。相当多的存储基础设施也放在那里。我认为我们在 AI 基础设施中看到的最大区别是专门化。过去,当我们建造数据中心时,它实际上是一个 20 到 25 到 30 年的建筑投资。我们考虑的是它在这 20 到 25 到 30 年期间将如何演变。你可能有服务器进去,可能有网络存储,可能有一些加速器、GPU、TPU,无论它们是什么,但这是一个 25 到 30 年的规划视野。所以我们必须规划很多代,而硬件的寿命可能只有六年。AI 数据中心通常更倾向于专用化。换句话说,我们经常与可能放入其中的硬件共同设计建筑。我们可能会说,你知道吗,实际上我们不会在这个建筑里放很多存储。为什么?因为一个存储机架可能有 10、20、30、40 千瓦的功率。你把它放在一个 TP 机架或 GP 机架旁边,今天这些机架很容易达到数百千瓦。未来几年人们还在谈论可能达到兆瓦级。设计一个可能容纳 30 个存储机架的建筑,而同一排中只有一两个 AI 机架。非常非常不同的设计。只需从尺寸、功率以及它如何分布在整个建筑等方面考虑。网络也是如此。存储机架所需的网络量很小。我的意思是,特别是如果它是硬盘驱动器,与 AI 机架相比。所以,如果你试图让它完全通用,你可能会把它建得太大、过度建设。在 30 年期间肯定通用。AI 数据中心可能更加专用化,与硬件共同设计,甚至到冷却和电力分配等方面。所以,更多的协同优化。
It's a very good question in that AI data center and a non-AI data center have do have a lot of similarities actually and they're they're not radically different. I mean it consists of concrete right an enclosure it consists of electrical yards mechanical yards cooling you have row after row of power being distributed. There's a huge amount of network infrastructure there. In other words, we connect large amounts of computes to one another. A fair amount of storage infrastructure goes in there. I think that the big difference that we're seeing with AI infrastructure is specialization. In the past, when we're building a data center, it really is a 20 25 30-year building investment. And we're thinking about how it's going to be evolving over that 20 25 30-year period. You might have servers go into it. You might have networking storage. You might have some accelerators, GPUs, TPUs, whatever they might be going into it, but it's a 25 30 year planning horizon. So we have and the lifetime of the hardware might be six years. So many generations that we have to plan for. An AI data center often times is going to be more purpose-built. So in other words, we are oftenimes co-designing the building with the hardware that might go into it. We might be saying, you know what, actually we're not going to put a lot of storage into this building. Why? Because a storage rack might have 10, 20, 30, 40 kilowatts of power. You put that next to a TP rack or a GP rack that easily is hitting hundreds of kilowatts today. Might hit me. People are talking about that in the next few years. Designing a building that might take 30 storage racks in a row versus one or two AI racks in that same row. Very very different design. Just think of it in terms of the size. Think of it in terms of the power and how that would be distributed across the building etc. Networking would be the same. The amount of networking that you would need for a storage rack tiny. I mean, if especially if it's hard drives compared to an AI rack. So, if you're trying to make it fully fungeible, you'll probably make it too big and too overbuilt. Fungeible over certainly a 30-year period. An AI data center likely is going to be much more purpose-built, co-designed with the hardware, even to the point of cooling and power distribution, etc. So, a lot more co-optimization.
你们向我的一个投资组合公司 Ineffable Intelligence 交付了一个大型 Ver Rubin 集群,我看到了它交付时的照片。那是一件美丽的东西。
You all delivered a big Ver Rubin cluster to one of my portfolio companies, Ineffable Intelligence, and I saw the photographs of it as it went out. That is a thing of beauty.
确实。确实。
It is. It is.
我的意思是,这就像人类历史上最大的建筑项目一样,你看着那个数据中心,就像,哇,那是人类能力的纪念碑。
It's it's I mean, that's one of those things where, you know, similar to looking at the biggest construction projects in human history, you look at that data center, it's like, wow, that is a monu monument to what what mankind can do.
不,我们拍了一张漂亮的照片,这只是几个机架。你知道,它的布线和光纤分布真的很美。我的意思是,美不在于观者,但对于像我这样的人,可能还有你自己和观众中的其他人来说。它是一件美丽的东西。我们把这个照片放在社交媒体上,人们喜欢 Ineffable,他们是他们所做事情的超级粉丝。优秀的团队,引起了很多关注。好的,为 Ineffable 准备的伟大机架。但实际上,人们只是喜欢看到光纤的照片和光纤的分形性质。实际上,我们曾经最受欢迎的帖子之一就是关于这个的。所以,是的,绝对。我们对此非常非常兴奋。
No, we got a nice picture of it and it um and this is just you know a couple of racks. You know the the cabling the fiber distribution for it is is really be I mean you got beauty is not the beholder but for folks like me and probably yourself and others in the audience as well. Um it is a thing of beauty. We put this picture up on um social media and um people love the fact that ineffable by their huge fan of what what they're doing. Fantastic team that got a lot of attention. Okay, great rack for an effable. But actually, people just loved seeing the the pictures of the fiber and the sort of the fractal nature of the fiber. One of the most popular um posts that uh we've ever had actually uh was around that. So, so yeah, absolutely. We're really really excited about that.
是的,看到照片我起了鸡皮疙瘩。
Yeah, I had goosebumps seeing the picture.
在这个大型建设过程中,你如何衡量并对自己负责?我们在节目开始前谈到,FLOPS 是一个虚荣指标,你更喜欢另一个指标。你能详细说说吗?
How do you measure and hold yourselves accountable in this in the middle of this big buildout? Um we were talking right before the show about how flops is a vanity metric and you prefer an alternate metric. Can you can you say more?
是的。所以我认为需要注意的一点是,无论是 FLOPS 还是你喜欢的其他以芯片为中心的指标,这些都是理论上的。换句话说,在某些条件下,对于你拥有的任何芯片,这是你能提供的最大 FLOPS。我们最终真正关心的是工作负载交付的性能。这就是我们所关注的。而且很少,实际上相对很少,性能是由单个芯片决定的。它涉及 FLOPS、HBM、SRAM 容量等。这些都非常重要,但可能还涉及如何将 2、4、8、16、1000、10000 个这样的芯片组合在一起,不仅仅是加速器(无论是 TPU 还是 GPU),还有可能为它们提供数据的 CPU,以及连接它们的网络。所以,你运行的工作负载是什么,该工作负载的性能如何?一个有趣的衡量标准是你的 FLOPS 利用率。例如,对于特定工作负载,如果你理论上具有 teraflop 或 petaflop 能力,你实际交付了其中的多少比例?
Yeah. So I think uh one thing to um note is that you know whether it's flops or pick your other favorite uh chip centric uh uh metric these are in theory. In other words under some conditions for whatever chip you have this is the maximum amount of flops that you can deliver. What we really care about in the end is what's the performance delivered by workload. That's that's what we're looking at. And rarely it it happens but actually relatively rarely is that performance determined by a single chip. It's flops, it's HBM, it's SRAMM uh capacity etc. Those are all super important but it might be how 2 4 8 16 a thousand 10,000 more of these chips composed together with not just the accelerators again whether TPUs or GPUs but then the CPUs that might feed them the data the network that connects them all together. So what is the workload that you're running and what is the performance of that workload? One measure that is interesting is what's your flop's utilization. So for example for a particular workload if you have in theory uh teraflop or a pedlop capable for your workload what fraction of that are you delivering?
现在这是一种有效吞吐量的衡量。
Now that's a measure of goodput.
什么是有效吞吐量?
What is goodput?
所以你可以想到吞吐量,这是一个众所周知的术语。
So you can think of throughput which is a well-known term.
是的,这是可能的吞吐量。但现在让我们考虑一些其他因素。一是工作负载的减速?同样,这只是工作负载固有的,但其他方面确实影响了我们的可靠性。所以就我们如何对自己负责而言,如果我们有一个芯片在同步工作负载中失败——而许多这些工作负载,无论是训练、服务还是智能体式工作负载,它们都是同步的。有许多许多组件同时协同工作。现在如果你有,比如说,一千、一万、十万个这样的组件同时工作,并且它们真的需要在微秒或毫秒粒度上协调。其中一个失败了,它实际上可能会让整个事情停止,对吧?因为每个人都依赖其他人完成他们的部分工作,以便得出一个非常棘手问题的答案。其中一个停止了。好吧,现在我们必须弄清楚发生了什么,哪一个停止了。我们在之前某个时间点有什么检查点?我们如何恢复那个检查点?我们如何重启?最坏的情况,我们必须从头开始。那会非常糟糕。但这在某些情况下可能发生,尤其是在推理方面。这里的重点是,如果你现在必须回去重做一堆计算,如果你必须暂停并等待弄清楚失败发生了什么,所有这些工作都是工作。它实际上并没有帮助你得到答案。对吧?就比如说你在纸上解决一个问题。第一步,第二步,第三步,第四步。如果你必须回到第一步,因为你犯了一个错误,是的,你仍然在做工作。那是吞吐量。你在做工作,但在交付你的答案方面的有效产出是什么?基本上就是你解决问题所花费的总时间。这才是重要的。如果你有失败,如果你有故障恢复,如果有任何中断你工作的东西,那就是问题的一部分。所以现在,我们如何对自己负责?是交付的有效产出,不是理论基准吞吐量或理论上的 goodness,而是实际上对于工作负载在真实故障条件下,发生了什么?不幸的现实是,在 10 万加速器规模下,
Yep, that's the throughput that is possible. But now let's consider some other considerations. One is what is the slowdown of the workload? Again, just inherent to the workload, but other aspects of it that really hit our reliability. So in terms of how we hold ourselves accountable, if we have a chip failing for a synchronous workload — and many of these workloads, whether training or serving or agentic workloads, they're synchronous. There are many, many components that are working together simultaneously. Now if you have, let's say, a thousand, 10,000, 100,000 of these components working together simultaneously and they're really needing to coordinate at microsecond or millisecond granularity. One of them fails, it might actually bring the whole thing to a stop, right? Because everyone is counting on everyone else to do their part of the job in order to come up with the answer to a really tough question. One of them stops. Okay, now we have to figure out what happened, which one stopped. What's a checkpoint that we have of the computation at some previous state in time? How do we restore that checkpoint? How do we restart? Worst case, we have to restart from the beginning. That would be really bad. But that can happen in certain cases, especially on more on the inference side. The point here is that if you now then have to go back and redo a bunch of computation, if you then have to pause and wait to figure out what happened for the failure, all that work is work. It's not actually helping you get the answer. Right? In terms of, let's say that you're working out a problem on paper. Step one, step two, step three, step four. If you have to go back to step one because you made a mistake, yes, you're still doing work. That's throughput. You're doing work, but what is the goodput in delivering your answer? It's basically the total amount of time it took you to solve the problem. That's what matters. And if you have failures, if you have failure recovery, if you have whatever it is that's interrupting your work, that's part of the issue. So now, how do we hold ourselves accountable? It's delivered goodput, not theoretical benchmark throughput or goodness in theory, but actually for a workload for real failure conditions, what's happening? And the unfortunate reality is that at a 100,000 accelerator scale,
我想说在那个规模下,总有东西在出故障,对吧?
I was going to say at that scale, something's filling all the time, right?
一直如此。而且每一个,如你所知,每一个芯片都是自然的奇迹。我的意思是,它们处于制造可能性的最前沿,对吧?所以再次,对我来说这令人惊叹。而现在实际上这些芯片不仅仅是一个芯片,它们是封装,由许多情况下两个、四个、八个,也许更多小芯片组成,对吧,组合在一起。然后当然还有旁边的 HBM、网络连接,也许还有共封装光学之类的东西。没有批评。很多东西都可能失败。现在如果你有 10 万个,某个东西会失败。你必须为此做好准备。近乎实时地检测到它。近乎实时地从它恢复。做所有这些,真的就像大海捞针——遥测问题巨大。就像连续地大海捞针,肯定跨越秒和分钟。但对于某些任务,小时、天甚至周,就像连续在线。我们如何对自己负责,就是我们在数据中心真正重要的工作负载上交付的有效产出是什么。
All the time. And each one of these, as you're aware, each one of these chips is a wonder of nature. I mean, they're at the very bleeding edge of what's possible to manufacture, right? And so again, to me it's stunning. And now actually these chips aren't just one chip, they're packages made up of in many cases two, four, eight, maybe more chiplets, right, that are composing together. And then of course HBM off to the side, network connectivity, maybe things like co-packaged optics. No criticism. Lots of things can fail. And now if you got 100,000 of them, something is going to fail. You have to be prepared for that. Detect that in near real time. Recover from that in near real time. Doing all that, it's really like finding the — the telemetry problem is massive. It's like finding a needle in the haystack continuously across certainly seconds and minutes. But for some of these jobs, hours and days and even weeks, like just continuous and online. How we hold ourselves accountable is what's the goodput that we deliver for the workloads that actually matter in the data center.
有效产出是谷歌的术语还是行业术语?
Is goodput a Google term or is it an industry term?
这是谷歌的术语,但我认为行业中越来越多的人开始采用它。
It's a Google term but I think more and more people across the industry are starting to pick it up.
然后只是为了校准我,10 万加速器规模,我们是在说它每分钟失败一次,每天一次,像怎样?
And then just to calibrate me, 100,000 accelerator scale, are we talking it's going to fail once a minute, once a day, like how?
10 万个加速器。所以让我看看,我认为在那个规模下肯定每天会多次,也许取决于具体配置每小时多次,某个东西会失败。
100,000 accelerators. So let me see, I think that it is definitely going to be at the scale multiple times a day and perhaps depending on the exact configuration multiple times an hour, something is going to fail.
最常见的失败原因是什么?
And what are the most common failure reasons?
这就是问题所在。实际上,这是一个非常好的问题。如果有一个常见的失败原因,我们就会弄清楚并修复它。这是一个不断发现的长尾。你知道,每当我们推出新产品时,总会有一些东西打击我们。在许多情况下,坦率地说,因为再次,它处于最前沿,可能是网络相关的。可能是我们如何以超高速连接这些东西。那可能是与硬件有关的事情。这些事情我们会解决。但然后很多问题可能是软件。所以这是我们如何对自己负责的另一个方面。再次,芯片可能能够达到一定水平的 flops。但如果你有一个编译器错误、运行时错误、模型问题、其他东西、操作系统问题,那没关系。那会影响整个系统性能。所以你可能有完美的硬件完全可靠,但然后你可能有软件问题伤害你。
This is the problem. Actually, it's a very good question. If there were a common failure reason, we'd have figured it out and fixed it. It is a long tail of constant discovery. You know, whenever we have a new product that's being introduced, there's going to be something that hits us. In many cases, frankly, because again, it's at the very bleeding edge, it might be network related. It might be how we connect these things together at super high speed. That might be something to do with the hardware. Those things we work through. But then a lot of the issues could be software. So this is the other aspect of how we hold ourselves accountable. Again, the chip might be capable of a certain level of flops. But if you have a compiler bug, a runtime bug, a model issue, something else, operating system issue, it doesn't matter. That's going to impact the entire system performance. So you might have perfect hardware fully reliable but then you might have software issues that hurt you.
加速器公司,比如 Nvidia 或 TPU 团队,有没有一个标准的参考栈,你知道,这里是围绕我们的加速器构建的最佳系统,只要你按照那个系统构建你就没问题?或者你必须做多少自己的数据中心设计,超出半导体公司给你的?
Is there a standard reference stack from the accelerator companies, from Nvidia or from the TPU team, of you know here is the optimal system to build around our accelerators and as long as you build to that system you're good? Or how much of your own data center design do you have to do above and beyond what the semiconductor companies give you?
对。所以有一个参考栈,我的意思是,我会说 Nvidia 是一家令人难以置信的整体系统公司——我的意思是他们显然是一家半导体公司,但他们不仅仅是一家半导体公司。他们给你一个非常非常强大的参考栈。但我会说许多——我们发现,我会这样表述,我们的大多数客户利用那个参考栈,但许多也进行专门化。换句话说,他们可能会发现,对于他们的特定用例来说,这是自然的,他们有一个优化机会或需要做一些不同的事情,他们就会去做。同样对于 TPU 方面,我们有一个参考栈。但然后大多数人会利用那个。不过许多也会对其进行专门化。
Right. So there's a reference stack and I mean I would say that Nvidia is an incredible whole systems company that — I mean they're obviously a semiconductor company but they're not just a semiconductor company. They give you a really, really strong reference stack. But I would say that many — we find that, the way I'd put it is most of our customers leverage that reference stack but many also specialize. So in other words, they might find that as natural for their particular use case, they have an optimization opportunity or something different that they need to do and they're going to do that. Similarly for the TPU side, we have a reference stack. But then most people will leverage that. Many will specialize it as well though.
明白了。然后我猜就整体容量而言,你告诉你的团队谷歌必须大约每六个月左右将服务容量翻倍。对吗?
Got it. And then I guess in terms of overall capacity, you've told your teams that Google has to roughly double the serving capacity every six months or so. Is that right?
嗯,所以澄清一下,这是从有效可用容量的角度,从我们最终看待它的方式来看,是从服务的角度。这是 token 生成能力。所以容量,这就是我会说它是软件和硬件的结合。硬件可能有一些所谓的固有 flops 水平。我在这里说的不一定是你必须每六个月将 flops 数量翻倍。那是一条路径。
Well, so to be clear, this is in terms of effective available capacity from the way in the end that we look at it is from a serving perspective. It's token generation capability. So capacity and this is where I would say it's a combination of software and hardware. The hardware might have some quote unquote inherent level of flops. What I'm saying here is not necessarily you have to double the number of flops every six months. That's one path.
你必须每六个月将硬件生成 token 的能力翻倍。而其中来自软件的部分,至少和来自硬件的部分一样多,甚至更多。换句话说,可能是模型优化带来了这种提升,也可能是某种运行时优化。实际上,很可能要靠几十甚至几百项单独的优化不断落地、反复叠加,才让这一切成为可能。但没错,能力提升的速度是惊人的。
You have to double the capability of that hardware to generate tokens every six months. And as much or more of that is going to come from software as it is from hardware. So in other words, it could be a model optimization that delivers that. It could be some runtime optimization that delivers that. It's really probably going to be dozens or hundreds of individual optimizations that are just landing again and again and again to make it all possible. But yes, the rate of capacity improvement is incredible.
我猜我们现在已经经历了几年的数据中心建设,也有了几年的模型进步和软件进步。从经验上看,如果以每瓦智能来衡量能力,能力提升中有多少来自芯片本身,多少来自模型,多少来自其他软件,以及任何其他大的组成部分?
I guess we've had a few years of data center buildout now and a few years of model progress and software progress. What has been the empirical breakdown of how much capacity — me as measured by intelligence per watt — how much capacity increase has come from the silicon itself versus the models versus other software and any other kind of big components?
是的,这是个非常好的问题。我没有确切的分解数据,但我会说,根据我们的经验,大部分收益来自模型侧的改进,以每瓦智能来衡量。顺便说一句,每瓦智能是个极好的指标。真正聚焦于——对我们来说也是每瓦有效产出(goodput per watt),我们可以回头再谈为什么瓦特需要作为分母。有效产出很可能就是每瓦所交付智能的一种度量。换句话说,有效产出是一个与工作负载相关的指标。但我会说,大部分收益往往来自模型侧,软件系统也贡献相当大。为什么?因为它们能真正确保你拥有的硬件被有效利用。然后硬件本身,它相当惊人。换句话说,我们现在所处的世界里,每年 2 倍甚至更高的性能提升是绝对可能的。所以硬件确实支持这一点。而且它是一个——如果你愿意,它是一个免费的乘数。不是免费的,但位于其上的每个人都可以指望它,作为一个年复一年提升所有人的乘数。
Yeah, it's a really good question. I don't have the exact breakdown, but I would say that in our experience, most of the benefits come from model side improvements in terms of intelligence per watt. And by the way, intelligence per watt is a fantastic metric. Really focusing — we also for us it's goodput per watt, and we can come back to why the watts need to be the denominator. Goodput could well be a measure of the intelligence delivered per watt. In other words, goodput is a workload specific metric. But I would say most of the gains often come from the model side, software system quite a bit. Why? Because they're able to actually make sure that the hardware that you have is being used effectively. And then the hardware, it is pretty stunning. In other words, we're living in a world right now where 2x or more year-over-year performance improvements is absolutely possible. And so the hardware does support it. And it is a — if you want, it's a free multiplier. Not free, but everyone else above it can count on it in terms of that being a multiplier that lifts everyone else up year-over-year.
我很想谈谈 TPU 项目和协同设计。谷歌十多年前就开始构建定制芯片。当时这是一个反共识的决策。TPU 项目是如何演变的?
I'd love to talk a bit about the TPU program and co-design. So Google started building custom silicon more than a decade ago. It was a contrarian call at the time. How has the TPU program evolved?
变化很大。我的意思是,当这个项目在 2013 年启动时,那确实是一个反共识的决策,我觉得当时很难把自己放回那个时刻。2013 年的传统智慧,你知道所有最聪明、最有智慧的人都会说,你不会为单一工作负载构建定制加速器。为什么?因为摩尔定律还在,性能每 18 或 24 个月翻一番还在,你可以利用标准编程模型,你所有的 C++ 代码等等,Java、Python,不管是什么。就像芯片领域的苦涩教训。
Quite a lot. I mean when the program started in 2013 it was really a contrarian call and I mean I think that at the time it's hard to put yourself back into the moment. 2013 conventional wisdom, you know all the smartest wisest people would say you don't build a custom-built accelerator for a single workload. Why? Because Moore's law is there, doubling of performance every whatever it is 18 or 24 months is there, you get to leverage standard programming models, all your C++ code etc., Java, Python, whatever it might be. Like the bitter lesson of chips.
是的,没错。就像芯片领域的苦涩教训是,专业化永远不会赢。
Yes, exactly. It's like the bitter lesson of chips is that specialization never wins.
但在这个特定情况下,我们有这一个应用,或者少数几个应用,会从中极大受益,而支持它们需要难以想象的大量通用 CPU。所以真的在 2013 年,这是一个赌注,而且我可以说,即使在公司内部也有相当多的人不确定它能否成功。所以这是一个赌注。结果证明这是一个极其成功的赌注。第一块芯片完全是关于推理的。第二块芯片是,嘿,实际上我们可以用同样的思路构建一块训练芯片。从那里开始,它被越来越多的用例采用。大约在第二块芯片问世时,Transformer 被发明了。我的意思是,这是一个重大时刻,实际上彻底改变了 TPU 项目。我们当然了解到推荐系统可以在 TPU 上运行得非常好,所以广告和相关用例也随之而来。所以我会说,这种变化和演变一直是范围和影响力的扩展。换句话说,它始于一个非常有影响力的用例,几个用于推理服务的应用。主要是语言翻译和语音识别,然后是训练,然后是 Transformer、推荐系统,并随着我认为 GenAI 时刻的到来而继续泛化,扩展到越来越大、越来越可扩展的系统。
But in this particular case, it was we have this one application or a few, a small number of them, that would benefit tremendously and that would require an unimaginable amount of general purpose CPU to support. So really in 2013 it was a bet and there were quite a few people even within the company I would say who were not sure that it would work out. So it was a bet. It turned out to be a massively successful bet. So the first chip was all about inference. The second chip was hey actually we can take the same idea and build a training chip. From there it got picked up for more and more use cases. Around the time the second chip came out transformers were invented. I mean this was a major major moment and that actually completely shifted the TPU program. We of course learned that recommender systems could run really really well on the TPUs and so ads and related use cases came along. So I would say that change and evolution has been one of expansion in scope and impact. In other words, it started with a really impactful use case of a couple of applications for inference serving. This was language translation primarily and voice recognition, to then training, to then transformers, recommender systems, and continuing to generalize as I guess the GenAI moment took a hold, to ever larger, ever more scalable systems as well.
你提到 Transformer 的出现是一个重大时刻。我很好奇想探讨一下,TPU 像是一种特定架构,但它并不是专门针对 Transformer 的。你如何把握这条微妙的界线,即你希望芯片对工作负载有多具体?
You mentioned the moment the transformer came through being a major moment. I'm curious to explore the TPU was like a specific architecture but it wasn't you know specific to the transformer specific. How do you kind of straddle the fine line of you know how specific you want your chip to be for the workload?
这是一个很好的问题,我认为这真的归结为适用性。我的意思是,芯片对特定工作负载的适用性。换句话说,最终我们每一代都在思考这个问题。我们是否要进一步专业化?我们大约两年前或更早面临的问题是,在 2026 年,我们应该有两块芯片还是一块?我们可以有一块芯片同时做推理和训练,并且两者都做得很好,或者我们可以有两块芯片。一块进一步专门用于推理,另一块进一步专门用于训练。这项分析和这项工作导致了今年发布两块芯片。8 I 用于推理,8 T 用于训练,最终我们意识到,虽然几年前不一定如此,但到 2026 年我们看到推理和服务真正起飞,所以拥有一块在服务方面显著更快的芯片,我们认为它在其生命周期内可能占据 30%、40%、50%、60% 的市场,这对我们来说开始变得非常合理,相对于如果我们预测推理只占市场的 2% 或 5%,即使那块专用芯片快 2 倍,也可能不合理,对吧,因为你实际上会选择那块通用芯片,它没有完全优化,但没关系,因为你现在有了统一性等等。所以真的,这归结为一个计算问题:好吧,你要进一步专业化吗?那个工作负载有多大?那个工作负载预计在两三四年后有多大?这是硬件中特定工作负载的持续增长吗?再次如你所知,有一个机会,即你越专门针对特定工作负载,它就越不灵活,硬件就越快、越节能。所以这是一门艺术,是对你正在设计什么以及那个工作负载有多持久的预测。换句话说,如果它在一两个月或三个月后消失,即使在那三个月里很大,你也有一个非常窄的窗口来拦截它。所以它必须有一定的持久性。你必须能够准确预测通过专门化能获得什么收益。所以大致的权衡是,支持一个新项目有一笔固定成本——一大笔固定成本。
This is a great question and I think that it really comes down to the applicability. I mean I think that of the chip to a particular workload. So in other words in the end we've been thinking about this for every generation. Would we further specialize? The question that we were facing let's say two years ago or a little bit more was do in 2026 should we have two chips or one? We could have one chip that could do inference and training simultaneously and do both quite well or we could have two chips. One that was further specialized for inference and a second one that was further specialized for training. This analysis and this work led to the release of two chips this year. 8 I for inference, 8 T for training, and in the end what we realized is that it wasn't the case necessarily a couple years ago but by 26 we saw inference and serving really taking off and so having a chip that would be significantly faster for serving that we thought might be 30, 40, 50, 60% of the market in its lifetime started making a lot of sense for us relative to if we projected inference to be 2% or 5% of the market even if that specialized chip is let's say 2x faster it might not make sense right because you'd actually go for the one general purpose chip that's not fully optimized but that's okay because you now have the uniformity etc. So really it does come down to a calculus question of okay would you further specialize? How big is that workload? How big is that workload projected to be in a couple of three four years? Is this a sustained growth for a particular workload in hardware? Again as you know there is this opportunity where the more you specialize to a particular workload the less flexible it is, the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload. In other words, if it's going to go away after a month or two months or three months, even if it's big for those three months, you got a really narrow window to intercept it. So it has to be somewhat durable. And you have to be able to project exactly what win you can get for specializing to it. So the rough trade-off is there's a fix — a large fixed cost to support a new program.
是的。
Yes.
基本上你必须考虑,对那个特定程序的需求是否足够大,以证明成本的合理性。
And you basically have to think that there's going to be enough demand for that specific program to justify the cost.
没错。在某种程度上,比如我们的 8i 和 8t 芯片,两款芯片都能处理另一种工作负载。这一点也很关键。如果 8i 只能做推理,完全不能做训练——换句话说,它在训练上的性能为零——反之亦然,如果它在训练上非常出色,但推理性能为零,那也会是一个关键限制。为什么?因为我们必须在硬件六年的生命周期内提前预测每种工作负载的需求量。在这种情况下,这很好,因为两款芯片在各自专长的领域都更出色。但如果你需要的话,如果某一处有剩余容量,两者实际上都能做对方的工作。根据你做的专业化程度,你可能会变得过于专业化,以至于限制了自己的灵活性。
Exactly. And to some extent, for example, for our 8i and 8t chips, both chips can do the other workload. This is also key. If 8i could only do inference and it couldn't do training at all—in other words, it had zero performance for training—and then vice versa, if it was amazing at training and had zero performance for inference, that would have also been a key limitation. Why? Because we would have to predict ahead of time over a six-year period lifetime of the hardware exactly how much we need of each. In this case, it was nice because both chips are better at what they're specialized for. But if you needed to, if you had leftover capacity in one place or the other, both can actually do the other's job. Depending on the specialization you do, you might get so specialized that you actually limit yourself from being flexible.
那么,如果粗略的权衡是另一端程序规模的大小,在我看来,现代 AI 市场大部分都是基于 Transformer 的。所以也许问一个挑衅性的问题:为什么不直接把 Transformer 架构烧录到芯片里?
And if the rough trade-off then is kind of size of program on the other end, like it seems to me that so much of the modern AI market is transformer-based. And so maybe just asked a provocative question, why not just burn the transformer architecture into the chip?
是的。所以我认为它是基于 Transformer 的,但下一个层次的问题是:Transformer 归根结底是向量和矩阵乘法运算,以及 softmax。所以本质上有一系列线性代数原语。我们已经把这些大致烘焙到硬件中,其他人也是如此。这一切都是基于 Transformer 的,但接下来是你的模型架构——好吧,你有多少层?你如何穿过这些层?你如何跨层?对于每个维度,你应用的确切矩阵和向量形状是什么?你可以更进一步,不仅针对 Transformer,还针对你的模型进行专业化。那是下一个层次的专业化。我认为有很多公司正在考虑这一点。我认为这也是一个非常非常有趣的方向。
Yeah. So I think that it's transformer based but then the next level of question would be: transformers in the end are about vector and matrix multiply operations and you know a softmax. So there's a range of essentially linear algebra primitives. We have those and others do as well roughly baked into the hardware. It's all transformer based but then it's your model architecture—okay, how many layers do you have? How do you go through the layers? How do you go across the layers? How for each dimension, what exact shape of matrices and vectors are you applying? You could go further and specialize to not just the transformer but your model. That's the next level of specialization. I think there are a number of companies out there that are thinking about that. I think it's a very, very interesting direction as well.
好的,那么现在,你的客户是否将 TPU 和 GPU 视为大致可互换的等价物,还是有一组问题更适合其中一种?
Okay so at this point in time do your customers view TPU and GPU as roughly fungible equivalents or is there a certain set of problems that is better suited for one or the other?
肯定有一组问题更适合其中一种。我的意思是,GPU 比 TPU 更通用,这一点很清楚。Google 和 Google Cloud 有令人难以置信的产品。我们卖很多 GPU。我们在内部也使用 GPU。但我认为这真的取决于你问题的具体细节。所以我认为我们的客户——两者之间有重叠——但我们的客户会评估他们的工作负载并评估他们的选择。我们在 Google 喜欢做的是给客户选择。换句话说,我们希望为他们的需求提供正确的解决方案,当然也提供最能满足他们需求的解决方案。
There's for sure a better set of problems that are better suited for one or the other. I mean GPUs are more general purpose than TPUs for one. That is clear. There are incredible products at Google and Google Cloud. We sell a lot of GPUs. We use GPUs internally. But I think it really then comes down to the specifics of your problem. So I think that our customers—there is overlap between the two—but our customers then evaluate their workload and evaluate their options. What we like to do at Google is give our customers choice. In other words, we want to have the right solution for their needs and of course provide the solution that best meets their needs for them.
从芯片到网络再到软件的协同设计,支持的理由是什么,反对的理由又是什么?
What is the case for co-design from the chip to the network to the software and then what is the case against co-design?
是的。所以协同设计有巨大的优化机会。你可以想象,如果你有一个端到端的层栈,并且你希望能够挑选任何你想要的组件,假设你跨多个云运行,或者跨多个硬件、多个软件运行。你可以设计抽象层,基本上说:我可以在任何硬件上运行。我可以在任何软件上运行。我可以在任何网络拓扑上运行,你给我任何数量的网络都没问题,我的系统很可能会完全自适应。所以现在你有了一个很棒的能力。你可以移动到任何地方,就像新容量一夜之间可用。你就能启动并运行,因为你以这种方式设计了系统,能够利用任何东西。你没有硬编码或专门针对任何人的特定基础设施。缺点是,如果你变得完全可互换,你可能会在性能上留下很多。如果你对任何人的硬件、软件、网络、存储、计算等栈都完全灵活。所以协同设计的理由是,如果你试图获得完全通用性,每一层之间都有很大的阻抗不匹配。所以每一层可能有 10%、20%、2 倍的优化机会。你开始把这些优化机会乘起来,突然之间你就有了一个巨大的端到端机会,无论你想说每瓦智能还是每瓦有效吞吐,你都可以利用,甚至一直向下到电力输送和电力可用性软件优化等等。所以优点是你可以随时随地运行,没有锁定。缺点是你留下了显著的、很可能大量的性能。
Yeah. So there's huge optimization opportunities with co-design. And so what you can imagine is that if you have an end layer stack and you want to be able to pick and choose whatever component you want, let's say you're running across many clouds or you're running across many pieces of hardware, many pieces of software. You could design abstraction layers that would basically say I can run on any hardware. I can run on any software. I can run on any network topology, any amount of network that you give me no problem and my system is going to be fully adaptive very likely. So now you have a great capability. You can move anywhere like new capacity becomes available overnight. You're up and running because you've designed your system that way to actually be able to take advantage of anything. You have not hardcoded or specialized at all to anyone's particular infrastructure. Downside is you probably leave a lot of performance on the table if you become fully fungible. If you become fully flexible to anybody's hardware, software, network, storage, compute, etc. stack. So the case for co-designing is between each one of those layers there's a big impedance mismatch if you try to get full generality. And so there might be 10%, 20%, 2x across each of these layers. You start multiplying those optimization opportunities through and all of a sudden you're left with a big end-to-end opportunity in terms of whatever you want to say intelligence per watt or goodput per watt that you can leverage even again all the way down to power delivery and power availability software optimizations etc etc. So pros are you can run anywhere anytime you have no lock in. Cons are you're leaving significant and in all likelihood significant amount of performance on the table.
所以我的理解是,如果你看 OpenAI 和 Anthropic,OpenAI 主要是在同构计算栈上构建,而 Anthropic 则是在大致更异构的计算栈上构建。你认为这是协同设计的一部分原因,导致他们收敛于传闻中非常不同的模型架构吗?
So my understanding is that if you take an OpenAI and Anthropic, OpenAI was kind of primarily building on a homogeneous compute stack and Anthropic was building on a roughly more heterogeneous compute stack. Do you think that's part of the reason that co-design is part of the reason why they converged on what is rumored to be very different architectures for their models?
我不能推测。我不能——我不想推测 OpenAI 和 Anthropic 在做什么。这是一种可能性。但我认为,在不知道他们正在做什么的细节的情况下,我想象他们最终采用不同架构可能有很多原因。
I can't speculate. I can't—I don't want to speculate in terms of what OpenAI and Anthropic are doing. It is one possibility. But I think that without knowing the details of what they're doing, I would imagine that there could be many reasons for why they wind up with different architectures.
那么,对 Google 来说呢?我很好奇你的团队和 DeepMind 之间的工作关系是什么样的。在模型开发的哪个阶段谁在场,你们如何共同决定协同设计?
Well, what about for Google then? I'm curious what the working relationship looks like between your team and DeepMind. Kind of who's in the room at what stage of model development and how are you making decisions together on co-design?
在 Google 最有趣、坦率地说最令人满足的部分之一,就是有机会与 DeepMind 团队真正并肩工作,共同设计我们的硬件和模型。还有第三个元素,我们还可以将其扩展到消费者服务和云。我暂时把这个放在一边。我可以稍后再谈。但关于 DeepMind,这确实是一个深度合作伙伴关系。让我——我的意思是,从过去我可以给你例子,他们提出了模型优化,比如对 Transformer 或他们可能想做的特定数学运算,而我们可能有一个芯片正在进行中,还没完全完成,但然后他们说:哦天哪,如果我们有硬件支持这个,我们的端到端训练或服务可能会显著更快、更高效。要真正改变我们可能执行的硬件定义以容纳这一点,需要什么?
It's one of the most fun and frankly gratifying parts of being at Google is the opportunity to work really shoulder-to-shoulder with the DeepMind team in terms of co-design of our hardware and models. There's a third element to it in terms of that we also get to extend that with the consumer services and cloud. I'll put that aside for a moment. I can come back to that. But with respect to DeepMind, it really is a deep partnership. So let me—I mean from the past I can give you examples where they have come up with model optimizations, let's say to transformers or to particular math that they might want to do, and we might have a chip in progress not quite done, but then they say oh my gosh if we had hardware support for this our end-to-end training or serving might get significantly faster more efficient. What would it take to actually now change our hardware definition that we might be an execution on to accommodate?
这就会让我们的工程师和研究员密集地聚在同一个房间里,待上几天、一周、两周,说:“好,是的,我们能做这个。更可能的是,我们没法完全做到你想要的,但我们可以做另一件事。”然后也许你可以把模型架构朝另一个方向改,给你 98%——给我们 90% 我们想要的东西。是的,然后我们可以回去拦截硬件。或者我们可以说,你知道吗,我们要把流片推迟一两周,因为哇,为了这种程度的收益,完全值得。
And this then leads to our engineers and researchers getting together intensely in the same room for a few days, a week, two weeks, saying, "Okay, yes, we can do this. More likely, we couldn't quite do what you wanted, but we can do this other thing." And then maybe you can change your model architecture in this other direction that gives you 98%—that gives us 90% of what we were looking for. And yes, we can then go back and intercept the hardware. And/or we can say, you know what, we're going to delay the tape out by a week or two weeks because wow, for this level of benefit, it's totally worth it.
同样,当我们规划路线图时,是时候——我们在任何时间点都有很多代芯片在推进。我们有正在生产的芯片,这是第一类。我们有正在努力投入生产的芯片,它们已经从制造商那里回来了,我们正在调试它们,让它们能工作。我们有正在实现、即将流片并送往制造商的芯片。我们有处于设计阶段的芯片,然后我们有处于概念阶段的芯片。所以这真的是一个五、六个阶段、很多很多年的流水线,从生产到在你脑海中,等等。
Similarly, when we're projecting our road map out and it's time for—we have many generations of chips that are essentially in progress at any point in time. We have the chips that are in production. That's one. We have the chips that we're working on getting into production. They're already back from the manufacturer and we're debugging them, making them work. We have the chips that are in implementation that are about to tape out and go to the manufacturer. We have the chips that are in design phase and then we have the chips that are in concept phase. So it's really this five, six stage, many, many year pipeline from in production to in your mind, etc.
与 DeepMind 在生产阶段的合作很重要,因为我们可以一起最大化交付的智能或每瓦交付的有效产出,我们确切知道模型中发生了什么,确切知道硬件中发生了什么,以及两者之间的一切。我们还可以在即将流片的芯片上深度合作。为什么?因为我们可以拦截,我们可以在飞行中对芯片——字面意义上的芯片架构——进行修改,如果跨公司边界工作,这介于困难和不可能之间。不是不可能。那会难得多,对吧,我们要说:“天哪,我们离完成这颗芯片只有几周或几个月了,现在让我们肩并肩地待在同一个房间里,弄清楚我们是否应该打乱这个项目。”这是可能的,但我会说更难。
The collaboration with DeepMind for in-production is significant because we get to work together in maximizing delivered intelligence or delivered goodput per watt, and we know exactly what's happening in the model and we know exactly what's happening in the hardware and everything in between. We also get to collaborate deeply on the chips that are actually just about to tape out. Why? Because we can intercept and we can make changes to the chip—literally the chip architecture—in flight, which would be somewhere between hard and impossible to do if we were working across company boundaries. Not impossible. It would be much harder, right, for us to say, "Oh my gosh, we're a few weeks or a few months away from getting this chip done and now let's get in the same room shoulder-to-shoulder and figure out if we should disrupt the program." It's possible, but harder is what I would say.
然后当然,对于处于设计阶段的芯片,我们可以一起评估许多架构。我们可以问 DeepMind 的同事,你认为模型架构在未来 2、3 年会走向何方?这是我们可以做的事情的帕累托前沿。这是模型架构走向的帕累托前沿。我们实际上有深入而重要的仿真基础设施,可以预测工作负载将如何映射到不同的硬件架构。深度、快速的迭代。这真的不是团队各自独立工作。而是在同一栋楼里,很多时候在同一个房间里。深度的日常互动。我每周多次与 Koray 或 Demis 交谈,等等。所以这真的是工作中超级有趣的一面。
And then of course for the chips that are in design, we have many architectures we can evaluate together. We can then ask our colleagues in DeepMind, where do you see model architectures going in 2, 3 years time? Here's the pareto of things that we can do. Here's the pareto of where model architecture is going. We have actually deep and significant simulation infrastructures that can predict how the workloads are going to map to different hardware architectures. Deep, deep and fast iteration. It really isn't the teams are working separately. It's in the same building, same rooms many times. Deep daily interaction. I talk to whether that's Koray or Demis multiple times a week, etc. So it really is a super fun aspect of the work.
太棒了。最终他们的模型会帮助芯片设计。
That's awesome. And eventually their models will help with chip design.
我们也在用 Gemini 为未来的 Gemini 设计硬件。
We're using Gemini to design hardware for future Geminis as well.
真的很酷。我想稍后再回到合作本身。我猜是否仍然存在一种阻抗不匹配——我想当你处理硬件时,周期时间可能比你的 DeepMind 同事在软件方面处理的要慢,而且我想你们的规划周期要提前得多。
That's really cool. I want to come back to that a little bit later on the collaboration itself. I guess is there an impedance mismatch still of just—I figure the cycle times are probably just slower when you're dealing with hardware than than your DeepMind folks get to deal with on the software side, and I figure your planning cycles are much further ahead.
当然。
For sure.
那么你们实际上有多少调整空间?
Um and so how much room do you actually have to adjust?
是的,这是个很好的问题。我们提前两、三、四、五年规划硬件。毫无疑问。换句话说,我们刚才谈到了 TPU v8 和我们已经宣布的,但你可以想象 9、10 以及可能其他一些正在概念执行到其他阶段。所以它们可能还有几年。如果你在做模型架构,默认情况下你不会想几年后的事。
Yeah, it's a really good question. We are planning hardware two, three, four, five years in advance. No question. In other words, we talked about TPU v8 and at—just now that we announced, but you can imagine that 9, 10 and maybe some others are in concept execution to something else. And so they might be years out. If you're working with model architecture, by default you're not going to be thinking years out.
是的。
Yeah.
但我认为这也是公司共同成长的好处,你知道有 Google Research 和 DeepMind,Transformer 被发明——所有这些也都发生在这些重叠的房间里。换句话说,有一整代研究员已经习惯了能够影响硬件,并且知道硬件是在多年周期上运作的,他们也知道,看,如果他们有一个小的调整,能为即将流片的芯片带来 1% 或 0.5% 之类的提升,他们可能不会来找我们,因为他们足够了解,实际上这不像软件,你只要做一个变更列表,两周内就会投入生产。实际上停止流片是件大事,但他们也知道,如果他们有一个大的——一个真正好的机会——是的,我们绝对会一起努力,看看能不能把它放进去。
But I think this is also the great thing about how the company has grown up together and you know there was Google Research and DeepMind and Transformers being invented—all of this was happening in these overlapping rooms as well. So in other words, there's a whole generation of researchers who've grown accustomed to being able to influence the hardware and knowing that the hardware is operating on multi-year cycles, and they also know that look, if they have a small tweak that's going to deliver like 1% or 0.5% or something like that for a chip that's about to tape out, they're probably not going to come to us because they know enough to know that actually it's not like software where you can just do a change list and it's going to ship to production in two weeks time. Like it actually stopping a tape out is a big deal, but they also know if they've got a big—like a really good opportunity—yes, we're absolutely going to work together to figure out if we can get it in there.
所以,正如我所说,收益来源是乘数性的。硬件确实能水涨船高,对吧?因此,DeepMind 团队的很大一部分——我们非常感激,这是一个了不起的团队——但很大一部分人在思考我如何影响路线图,因为它实际上是一个流水线。就像我两年前的 idea,现在已经在生产中,帮助 Google 的所有工作负载跑得更快。那种感觉很好。
So, as I said, it's multiplicative in terms of where the benefits come from. The hardware does lift all the tides, right? And so, significant portions of the DeepMind team—and we're so grateful for it, it's an amazing team—but significant portions of it are thinking about how do I influence the road map because it's actually a pipeline. Like the idea I had two years ago, like it's in production now and it's helping all the workloads at Google go faster. Like that's a good feeling.
完全同意。不过,预测五年后哪些工作负载最常见、哪些算法突破会发生,似乎是一项不可能的任务。但似乎也不是不可能。
Totally. It seems like an impossible task though to predict in five years what workloads will be most common and what algorithmic breakthroughs will have occurred. Doesn't seem like an impossible task.
我听到你说这似乎是一项不可能的任务,但这里有个很棒的事情——实际上我们正在详细地写这个,做起来很有趣。令人惊讶的是,TPU 架构在中等细节层面上,不是超级高的细节层面,在中等细节层面上,自 TPU v1 以来并没有真正改变。一种看待它的方式是 CPU 的指令集架构。你有加载、存储、加法、减法、分支,然而在那段时间里,它上面的软件做了什么?TPU 也一样。我们有一些基本指令和基本原语。当然,我们扩展了它。并不是说指令集完全没有改变,但那里的原语——专门用于数值超大矩阵乘法单元、管理向量操作和散射-聚集操作的稀疏核心等等。有五六件事真正定义了一个加载、远程加载存储。那实际上是另一个。我们可以通过 ICI 网络读写远程内存。
I hear you that it would seem like an impossible task, but here's the awesome thing—and actually we're working on writing this in great detail and it's a lot of fun to do it. The stunning thing is that the TPU architecture at a medium level of detail, not at a super high level of detail, at a medium level of detail hasn't really changed since TPU v1. Like one way to look at it is the instruction set architecture for a CPU. Like you've got loads and you've got stores and you've got adds and you've got subtracts and branches, and yet what has the software on top of it done right over that period of time? Same thing with TPUs. We have some fundamental instructions and fundamental primitives. Of course, we've extended it. It's not like the instruction set hasn't gotten changed at all, but the primitives there in terms of specializing the numeric very large matrix multiply units, a sparse core that manages vector operations and scatter-gather operations, etc. There's five or six things that really define a load remote load store. That's another one actually. We can read and write remote memory through our ICI network.
好,说到工作负载的转变,过去一年左右最大的变化之一——我觉得真正是从今年年初、这个自然年开始的——大概就是长时程智能体的兴起。
Okay, speaking of shifting workloads, it seems like one of the biggest changes in workloads over the last year or so — I think it really started at the beginning of this year, this calendar year — was kind of the rise of the long-horizon agent.
是的。
Yes.
我猜这跟过去那种快速往返的 LLM 对话是完全不同形态的工作负载。就数据中心的需求而言,这意味着什么?
And I would guess that's a very different shape of workload than the quick-turn LLM conversations of years past. What does that mean in terms of data center needs?
对。我觉得这里面有两个巨大的方面。第一,它不再是人与加速器之间的交互。换句话说,当你在网页浏览器或手机上敲一个提示词时,当然,针对你的提示词会有一堆计算发生,但等响应回来之后,你得读它、得思考它,然后可能还有一个后续追问,这中间就是好几秒的交互时间。而在长时程智能体里,循环中没有人,也就没有天然的限流来决定请求以多快的速度打到模型上。所以,原本以秒、甚至几十秒计的交互时间,现在可能变成了毫秒级。对吧?我一拿到响应,就能解析它,也许还能稍微推理一下,然后就能确定我的下一个请求是什么。这是第一个大变化。第二个大变化是,所有这些推理和解析大概率会发生在 CPU 上。而这个 CPU 接下来还得考虑:好,在我把下一个提示词发回模型之前,我还需要去收集哪些其他状态?换句话说,我从这个响应里学到了一些东西,我要再走一步,但我实际上需要去抓一些上下文——也许来自我本地的 DRAM,也许来自另一个 CPU 上别人的 DRAM,也许来自 SSD,或者来自别处的 HDD。所以现在还得进行大量的编排。于是设计其实正在发生相当大的变化:对加速算力的需求在上升,但对 CPU、网络和存储——也就是传统数据中心 CPU 等等——的需求也在暴涨。
Yeah. So I think there are two huge aspects to this. One is it's no longer human-to-accelerator interaction. In other words, when you are typing a prompt on a web browser or on your phone or whatever, of course, in response to your prompt, there's going to be a bunch of work that happens, but then when the response comes back, you've got to read it. You've got to think about it, and then maybe you have a follow-up that's going to be multiple seconds of interaction time. Now, in deep horizon agents, there's no human in the loop that is going to naturally rate-limit how quickly requests are going to go to the model. So this — what went from seconds, maybe tens of seconds, in terms of interaction time — is now going into perhaps milliseconds. Right? As soon as I get a response back, I can parse it, I can reason about it perhaps a bit, and I can figure out what my next request is going to be. So that's big change one. Big change two is all that reasoning and all that parsing is going to probably happen on a CPU. And that CPU is then going to have to probably think about, okay, what other state do I need to go gather before making my next prompt back into the model? In other words, I've learned something from this response. I'm going to take another step, but I actually need to go grab some context — maybe from DRAM local to me, maybe from someone else's DRAM on another CPU, maybe on SSD or maybe on HDD somewhere else. So now a huge amount of orchestration has to take place as well. So the design actually is changing pretty significantly, where the demand for accelerated compute is going up, but the demand for CPU and networking and storage — the traditional data center CPU, etc. — is also going through the roof.
那这是否意味着你们要在 GPU 机架旁边放更多 CPU?也就是——对,GPU、TPU 机架。
Does that mean you're putting more CPUs alongside your GPU racks, then? So this is — yeah, GPU, TPU racks.
这就是关键问题——回到优化和专门化这个问题上。如果我们开始在 GPU 和 TPU 旁边放大量 CPU 机架,那就意味着我们实际上无法针对 TPU 机架相对于 CPU 机架的密度和网络需求做完全的专门化。TPU 机架的密度会比 CPU 机架更高,需要的网络大概也比 CPU 机架更多。所以换句话说,我们的建筑设计就要变了。另一个选择是保持统一性:把 TPU 或 GPU 全放在一栋楼里,CPU 放在隔壁那栋楼——顺便说一句,硬盘可能还得放在另一边另一栋楼里,因为它们又有一套不同的需求。这样一来,你就需要在这些楼之间做相当大规模的网络互联。所以一旦跨出单栋楼,网络复杂度就会大幅上升——从可靠性角度、从成本角度都是如此。延迟会上升,大概还在可接受范围内,但现在可能会到几百微秒,加上组件之间的排队甚至可能更高。所以这些考量确实在以相当有意思的方式发生变化。
This is the key question — going back to this question of optimization and specialization. If we start putting lots of CPU racks next to the GPUs and TPUs, that means that actually we can't fully specialize to the density and network requirements, let's say, of a TPU rack relative to a CPU rack. A TPU rack is going to be more dense than a CPU rack. It's probably going to need more networking than a CPU rack. So in other words, now our building design is going to change. Another option is maintain your uniformity. Put your TPUs or your GPUs all in one building and the building next door — but maybe put your CPUs, and maybe, by the way, the hard drives have to be in another building on the other side, because they have yet another set of requirements. Now you need networking between these buildings at a pretty significant level. So once you leave a building, the networking complexity goes up significantly — from a reliability perspective, from a cost perspective. Latency goes up, probably acceptably, but still now you might go to hundreds of microseconds, potentially more with queuing between the components. So the considerations do change in pretty interesting ways.
你这活儿真难。
Your job is hard.
这很好玩。
That's fun.
很好玩。对。网络这边有什么进展?我听说 Google 一直处在最新网络技术的最前沿,包括光网络。你能谈谈光网络现在处于什么状态吗?
It's fun. Yeah. What's happening on the networking side? I've heard that Google's always been at the forefront of the newest in networking, including optical. Can you say a word on what is — like, this — the state of optical networking?
你知道,我们 Google 大概在十五六年前,是最早引入所谓波分复用的公司之一,也就是可以在数据中心内的一根光纤上承载多路信号,我们实际上把它用在了所有机架之间的通信上。与此同时,我们还随波分复用一起引入了一项叫光电路交换的技术。本质上,光电路交换与传统的电分组交换不同,它完全在光域中传输和搬运数据。它的好处在于——让我先描述一下支撑它的技术。在分组交换机里,你会拿到一个带报头的分组,在电域里查看它,从报头里判断这个分组要去哪里。它可能有一个 IP 地址,说:好,我该往哪儿送?你去查表,好,对于那个目的地,我该从哪些端口转发出去?所以基本上,每秒有几十亿、几十亿、几十亿,甚至上万亿个分组以超高速度涌进来,你把它们转发出去。而光电路交换说的是:我不在电域里碰这些比特。我要做的是,对于一个输入端口,判断该把光送到哪个输出端口。好。现在我有——做这件事有多种方式。我们最早用的是所谓的 MEMS 开关,也就是微型电机,基本上是在三维空间里控制镜面。于是我们现在可以以可编程的方式,拿一个可能有——你选——128 个端口、256 个端口,差不多这个数量级的盒子。我们可以为每一根接入光纤的输入端口做配置,把它映射到一个输出端口,然后改变镜面的旋转角度,让光 literally 照在这些镜面上,被反射到正确的输出端口。最初我们做这件事有两个原因:一是在机架组之间建立局部性。比如说我有一个计算集群和一个存储集群,两者都在支撑——还是那句话——做搜索。我们知道这两个集群之间会大量通信。于是我们会配置镜面,纯粹以光的方式在这两个集群、这两个机架组之间建立捷径。这是第一个原因。第二个原因是,我们希望能够在不实际移动任何光纤的情况下扩展和收缩网络。不展开细节的话,我可以在白板上画出来。光电路交换机让你能够真正重新配置网络的骨干,来扩展或收缩它,而不需要人做任何事,只需要一个控制器来管理它。现在快进到 TPU。TPU 有一个环面拓扑,把所有 TPU 直接互相连接起来。
You know, we at Google, this was probably 15, 16 years ago, were among the first to bring essentially what's called wave division multiplexing, where you could put multiple signals on a single fiber within the data center, and we actually leveraged that for all of our communication between racks. At the same time, we introduced a technology along with the wave division multiplexing called optical circuit switching. And essentially what optical circuit switching does is, in contrast with traditional electrical packet switching, it transmits and moves data entirely in the optical domain. The great thing about that is — and so let me describe the technology that underpins it. In a packet switch, you would take a packet that has a header. You would look at it in the electrical domain. Figure out where the packet is headed in its header. It might have an IP address that says, okay, where do I head it? You look up, okay, for that destination, you look up in a table, what ports do I forward it along? So basically, billions and billions and billions, perhaps trillions of packets coming through at super high speed per second. You're forwarding them along. Optical circuit switching says I'm not touching these bits in the electrical domain. What I'm going to do is I'm going to figure out, for an input port, which output port to send the light to. Okay. And now I have — there's multiple ways to do this. The one that we started with is called MEMS switches, micro electrical motors that basically control mirrors in 3D. So we can now programmatically take a box that might have, you pick, 128 ports, 256 ports, some number like that. And we can configure every input port for a fiber that comes into it, map it to an output port, and then change the rotation of mirrors where the light literally shines down on these mirrors and gets reflected to the right output port. Initially, the reason we did this was twofold: to create locality between groups of racks. Let's say I had a compute cluster and a storage cluster, and both of them were in support of — again, making up search. We knew that these two clusters would talk to each other a lot. So we would configure the mirrors to create shortcuts purely optically between those two clusters, clusters of racks. That was reason number one. Reason number two was we wanted to be able to expand the network and contract the network without actually moving any fiber. And without working into the details, I could draw this on a board. The optical circuit switch would allow you to actually reconfigure the spine of the network to expand it or to shrink it without a human being having to do anything other than — it'd be literally a controller that would manage that. Now fast forward to TPUs. A TPU has a torus topology that connects all the TPUs to one another directly.
我之前谈到了吞吐量和有效吞吐量。我们能做的一件事是,如果有一个 TPU 机架出现故障,我们可以在不移动任何光纤的情况下用另一个 TPU 机架替换它。同样,我得比划一下或者画个白板。但本质上,我们可以说我们随时都有一个备用机架,当某个机架故障时,我们把光重定向到那个新机架。
I talked about throughput and goodput earlier. One of the things that we can do is if we have a TPU rack that fails, we can replace it with another TPU rack without moving any fiber. Again, I'd have to wave my hands or draw a whiteboard. But essentially, we can say we have a spare rack available at all times and when a rack fails, we redirect the light to that new rack.
而且这可以在毫秒内完成。
And that can be done in milliseconds.
是的。
Yes.
那你们为什么还要用光纤,而不是完全用自由空间?
Why do you have the fiber at all then, as opposed to fully free space?
是的,这是个很好的问题。衰减损耗和带宽都会大幅下降。此外,在一栋非常大的建筑范围内,在没有光纤把它们在三维空间中连接起来的情况下,要对准所有东西会很有挑战性,甚至可能不可能。我们讨论过这个。我们其实讨论过,而且有过一些非常有趣的讨论。但没错,主要还是用光纤,然后当它们到达光交换机时,光实际上就直接照射在这些微小的芯片上。所以这是网络的一个大方向。有很多——我的意思是,坦白说,网络在数据中心里的能力和需求都在爆发式增长。
Yeah, it's a very good question. Attenuation loss and the bandwidth would drop dramatically. Furthermore, across the range of a very large building, aiming everything without the benefit of fiber to connect it all together in 3D would be challenging to possibly impossible. We've talked about it. We've talked about it actually and there have been some really interesting discussions. But yes, it's mostly in fiber but then when they hit the optical circuit switch, essentially the light then literally shines down on these tiny chips. So that's one big direction of networking. There's a lot — I mean I would say networking is exploding in terms of its capabilities and need frankly in the data center.
太酷了。我们可以就这个话题聊一整场。
So cool. We could have an entire conversation on that.
是的,非常非常酷。
Yes, it is very very cool.
甚至就像回到那张来自 Ineffable 的照片,全是那些刺眼的线缆。
Even like the going back to the picture from ineffable, it was all the cables that cut everybody's eyes.
没错。是的。那些线缆最终有一些会接到我们数据中心里的光交换机上。
Exactly. Yeah. And those cables eventually some of them wind up at optical circuits which in our data centers.
有道理。
Makes sense.
好。我想转到电力这个话题。你一直在谈每瓦有效吞吐量和其他每瓦单位。这让我觉得每瓦意味着电力在某种程度上是约束性的、稀缺的或昂贵的约束。
Okay. I want to flip to talking about power. You've been talking about goodput per watt and other units per watt. Makes me think per watt means power is kind of the binding constraint or the scarce constraint or the expensive constraint in some way.
你知道,我经常被问到这个问题:我们面临的最大约束是什么?而现实是,我们并没有单一的最大约束。它们全都是约束。它们都超级难。而且它们还在不断变化。但如果我必须从根本上回答,我会说电力是我们面临的最根本的单一约束,其他一切似乎——就像我们知道怎么解决它们,只是需要在一段时间内解决。电力,是的,我觉得你的表述非常好。这是一个约束性的长期问题,我们还没有——我的意思是,你知道核能,也许核能会带来丰富的清洁能源,这肯定会解决很多问题,当它实现的时候,以及当它规模化实现的时候,仍然未知。
You know, so I get asked this question, what is the biggest constraint that we face? And the reality is there is no single biggest constraint that we face. They're all constraints. They're all super hard. And then they shift continuously. But if I had to answer fundamentally, I would say that power is the single most fundamental constraint that we face, like everything else seems — like we know how to solve them and it's a question of solving them over some period of time. Power, yeah I think the way that you put it is really nice. It's a binding long-term issue that we do not have a — I mean there you know nuclear, you know perhaps nuclear is going to abundant clean energy which would solve a lot of problems for sure when that happens and at when it happens at scale, still still unknown.
那这在实践中是怎么运作的?你知道,你要建一个新的数据中心,它需要 1 吉瓦的电力。我想你不能直接打电话给 PG&E 说,嘿,请送 1 吉瓦的电力过来。那么电力供应实际上是什么样的,你们是否必须一直垂直整合到比如自己做涡轮机,还是你们怎么解决电力瓶颈?
And so how does it work in practice? You know you're standing up a new data center, it needs a gigawatt of power. I imagine you can't just call PG&E and say hey please send a gig a lot of power. So what does provisioning power actually look like and are you having to actually vertically integrate all the way down to like, you know, doing your own turbines or how do you solve the power bottleneck?
是的,这又是一个大问题,重要的问题。我们——在 Google,我们偏好的模式始终是接入公用事业电网,接入电网。所以,这相当于打电话给你的——
Yeah, it's again it's a big question, important question. We do — our preferred model at Google is always to be utility connected, to be grid connected. So, it's the equivalent of calling your —
哦,所以你不能直接打给 PG&E。
Oh, so you can't just call PG&E.
没错,你非常礼貌地打电话给你数据中心所在地最喜欢的公用事业公司,当然你要提前很多很多年通知他们。换句话说,如果我们谈的是 1 吉瓦规模,这不是你能说“嘿,我明天需要 1 吉瓦。你什么时候开始计费?”的事情。这是我们共同规划的事情。你知道,对我们来说,从确保我们与这些公用事业公司合作时,建设这些基础设施的成本——我们可以就这个话题聊很久——我们承担这些成本的角度来看,这也是我们非常认真对待的事情。因为按照计费方式,实际上可能是,公用事业公司为我们建设产能的行为,比如说,其他人的费率可能会上涨。理论上,我们确保的是,比如说必须建设、升级的输电线路、额外的公用事业变电站等等,我们也为这些付费。但所以这是一个漫长的规划过程。完全可能是这样的情况:比如说我们需要 1 吉瓦——我随便编个日期——2028 年,而公用事业公司能在 2029 年给我们 1 吉瓦,能在 2028 年给我们比如说 700 兆瓦。所以现在我们可能面临一个问题:我们怎么补上那 300 兆瓦。一个答案就是——另一个答案是,好吧,我们能不能自己想办法发一部分电,也许用太阳能电池,或者用电池作为备用等等。会不会用其他来源?然后我们再次与公用事业公司合作。所以也可能是一种组合:我们可能在本地保留一些发电能力,即使在公用事业公司完全上线、比如说达到 1 吉瓦规模时,我们实际上还能向电网反向供电。这样在他们需要的时候也能提供,对吧?换句话说,一种看待方式是,一年中可能有那两周,也许是一年中最热的两周,住宅需求巨大。如果我们有本地发电,我们也可以把它回馈给电网。所以对我们来说,这确实是多年与公用事业公司合作的过程。
Exactly, you call very politely your favorite utility wherever it is that your data center is and of course you give them many many years of notice. So in other words, if we're talking about a gigawatt scale, it's not something that you can say, "Hey, I need a gigawatt tomorrow. When can you start billing me?" It's something that we co-plan together. You know for us it's something that we also take very seriously from the perspective of ensuring that when we work with these utilities, the costs of putting that infrastructure in place — we could get into a long conversation just on this topic as well — that we cover those costs because with the way that billing works it actually could be that by the act of the utility building out capacity for us, let's say other people's rates could go up. In theory what we ensure is that actually, let's say the transmission lines that have to be built, upgraded, additional utility-based stations etc., that we pay for those as well. But so it is a long planning process. It can absolutely be the case that in let's say that we need a gigawatt in — I'll make up a date — 2028, and the utility can get us a gigawatt in 2029, that can get us let's say 700 megawatts in 2028. So now we might be left with a question of how do we cover those 300 megawatts. Well one answer is just — Another answer is to say okay well would we figure out how we generate some subset of that power ourselves and maybe that would be with solar cells or with batteries as a backup etc. Would it be for other sources? So then we again work with the utility. So it might also be a combination where we might maintain some power generation local and have the capability even when the utility comes online fully at let's say the gigawatt scale where we could actually provide power back to the grid. So having that when they need it as well, right? So in other words, there might be the one way to look at it is there might be the two weeks of the year, maybe it's the hottest two weeks of the year where there's huge amount of residential demand. If we have some local generation, we can then provide that back to the grid as well. So it really is working over multiple years with the utilities for us.
为什么你们偏好这样做,而不是与公用事业公司合作,而不是自己垂直整合?
Why is your preference to do that versus to kind of go like to work with the utilities as opposed to to kind of vertically integrate yourselves?
是的。所以,主要原因是灵活性和双方的提升。一种看待方式是统计复用,或者如果你愿意,大数定律。如果我们需要 1 吉瓦的电力,假设我们希望有 99.99% 以上的可靠性,可能意味着我们必须建设 2 吉瓦的电力,对吧?在那个水平,99.99 或 99.999,你必须有一加一冗余。那会变得昂贵。然后还要确保理想情况下是清洁能源,而且可能就在我们数据中心旁边,这也可能很有挑战性。现在,如果我们与数据中心合作——同样,也许我们自己能带来一部分,在他们需要较少时我们可以给电网,我们可以取用他们的电力。所以换句话说,通过在更大的基础上利用这种统计复用,实际上每个人都赢。我们赢,电网赢,住宅用户也赢,等等。
Yeah. So, the main main reason is flexibility and uplift on both sides. So, one way to look at it is statistical multiplexing or if you want the law of large numbers. If we need to have a gigawatt of power, let's say we want to have that with 99.99% plus reliability, probably means we have to build two gigawatts of power, right? At that level at 99.99 or 99.999, you have to have one-plus-one redundancy. That gets expensive. And then making sure that that's going to be ideally a clean energy source probably right next to our data center, that can also be challenging. Now if we partner with the data center — again maybe we have some amount of that that we can bring ourselves that we can give to the grid when they needed less, that we can take their power. So in other words, by leveraging that statistical multiplexing over a much larger base, actually everybody wins. We win, the grid wins, residences win, etc.
我们更倾向于并网,只有在极少数情况下才会做表后供电,但即便如此,我们也是在假设会与公用事业公司合作的前提下做的,可能一年后我们就会想接入电网。所以这既是我们的提升,也是电网的提升。
We prefer that, and in rare, rare cases we will do it behind the meter, but even then we're doing it under the assumption that we're going to work with the utility where it might be a year later we want to be connected to the grid. So again, it's an uplift for us, it's an uplift for the grid.
有道理。你们如何决定数据中心建多大?
Makes sense. How do you decide how big to make a data center?
是的,这是一门艺术,也是激烈争论的焦点。曾经我们有过争论。我记得甚至 10 多年前、15 年前,在 Google 有过大辩论:我们是不是应该把所有东西都放进一个数据中心?当时,那会是一个吉瓦,那可是……
Yeah, that's an art and a source of significant debate. At some point, we had a debate. I remember even 10 plus years ago, 15 years ago, big debate at Google: should we just put everything into one data center? Back then, it was going to be a gigawatt, which was...
更简单的时代。
Simpler times.
是的。不同的时代,就像 10 到 15 年前,我们要建一个吉瓦的数据中心。所以一个明显的担忧是单点故障,对吧?如果你从 30 年的角度来看,那是个巨大的担忧。另一方面,如果你谈论今天的训练工作负载,越大越好。换句话说,拥有更多,因为从网络的角度来看,实际上你希望把东西集中在尽可能小的距离内。但话又说回来,两个问题:单点故障,还有现在的电力可用性。虽然 10 或 15 年前一个吉瓦很大但还可想象,但现在在美国或世界任何地方,都不可能把 Google 的总需求建在一个地方。这根本不可能。那么现在的最佳规模是多少?我们为此有模型、模拟器等。但这也取决于情况。看,在一些地方,我们会在网络边缘,或者在一个可能只有几十兆瓦的国家。我们实际上与 ISP 合作,可能是在伊拉克,真的可能是伊拉克。这是一个训练集群。好吧,也许那会接近一个吉瓦。其他一些站点可能是几百兆瓦,等等。
Yeah. Different times, like 10-15 years ago, we're going to have a gigawatt data center. So, one obvious concern there is single point of failure, right? If you're looking at things from a 30-year perspective, that's a huge concern. On the other hand, if you're talking about training workloads today, bigger is better. In other words, having more because from a networking perspective, actually you want to have things concentrated in as small a distance as possible. But then again, two issues: single point of failure, but then also now power availability. While a gigawatt was huge 10 or 15 years ago but imaginable, there's no way anywhere in the country or the world we're going to get the total requirements of Google built in one place. It's just not going to be possible. So now what's the optimum size? Again, we have models for this, simulators, etc. But it also depends. Look, in some places we're going to be at the edge of the network, or we're going to be in a country that might be tens of megawatts. We actually partner with ISPs that might be Iraq, like literally it could be Iraq. This is a training cluster. Okay, maybe that's going to be closer to a gigawatt. Some other sites might be hundreds of megawatts, etc.
有意思。我很想了解你如何考虑组合生命周期管理。我猜最大的最新集群用于训练最新的前沿模型,然后你回收旧的东西并运行推理。这是一个合理的框架吗?你们是否也建立专门的推理集群?这一切是如何运作的?
Interesting. I'd love to understand how you think about the portfolio life cycle management. So I guess my guess would be that the biggest newest clusters are used for training the latest frontier model and then you kind of recycle the older stuff and run inference on it. Is that a fair framework? Are you also standing up inference specific clusters? How does that all work?
你考虑的框架非常合理,我认为很有道理。但我想说,推理的需求如此之大,我们不能仅仅依赖那些不再完全用于训练的旧训练集群作为推理的基础。而且如果你想一想,我们可能在某一年再次将训练集群集中到少数几个大型站点,并且保持它们之间的网络距离小是有好处的。所以现在,无论它们位于世界何处,它们可能在同一大陆,甚至在某一年可能在同一大陆的同一部分。所以现在你可能在其他大陆没有足够的服务容量,等等。那么好吧,实际上我们必须在世界其他地方建立专门的推理。所以我认为你的直觉很准确,但不完全。换句话说,确实必须这样:这里是训练集群。是的,可能几年后它们会用于服务。但我们也必须建立服务集群。
Very reasonable framework that you're thinking about, and I think makes a lot of sense. But I would say that the demand for inference is such that we can't just rely on whatever older training clusters that are no longer being fully used just for training as the basis for inference. And also if you think about it, we might centralize again in a particular year into a small number of large sites our training clusters, and there is benefit to keeping the network distance between them small. So now let's say wherever they're located in the world, they're probably going to be on the same continent or they might be even in the same portion of a continent in a particular year. So now you might be left with not enough serving capacity on the other continents, etc. So then okay, actually we then have to go build the specialized inference elsewhere across the world. So I think your intuition is spot on but not wholly. In other words, it really does have to be okay here are the training clusters. Yes, probably they're going to be used for serving in some number of years. But then we're also having to build the serving clusters as well.
你们的服务集群与训练集群不同吗?比如它们更小吗?每兆瓦更便宜吗?
And your serving clusters are they different than the training clusters? Like are they smaller? Are they cheaper per megawatt?
实际上不一定更便宜,因为对于服务来说,更需要将计算、网络和存储放在一起,因此我们无法专门化。所以对于训练,你实际上有这种大密度均匀部署等。而对于服务,你必须混合存储、计算和加速器。此外,对于服务,你实际上不希望在一个地方有太多。你希望从全球各地提供工作负载,但现在我们最终有单独的模型端点,这实际上可能是模型的变体。所以现在我们必须将这些模型分散到世界各地,还要考虑本地性。所以它们会更小。在推理方面,它们实际上可能垂直整合程度更低。
Not necessarily cheaper actually because for serving there is more need to colocate compute, networking, and storage, and so that then we get to the inability to specialize. So for training you actually have this big density uniform deployment, etc. Whereas for serving you're going to have to have the mix of storage, compute, and accelerators. Furthermore, for serving, you actually don't want to have too much in one place. You want to be serving your workloads from all over the planet, but now we wind up having individual model endpoints and that actually can be variations of models. So now we have to spread these models out across the world also accounting for locality. So they will be smaller. They'll actually be less vertically integrated probably on the inference side.
超级有趣。所以如果你 5 年前建了一个数据中心,你知道,5 年前最先进的加速器非常不同,效率远低于今天制造的。
Super interesting. So if you made a data center 5 years ago, you know, state-of-the-art accelerator 5 years ago was very different, vastly less efficient than the ones that are being made today.
是的。
Yes.
你们真的会回去更换那些旧数据中心的芯片吗?我知道这有点关联到一直存在的争论,我认为,比如芯片的实际使用寿命是多久?
Do you actually go back and like swap out the chips in those old data centers? And I know this kind of relates to there's been an ongoing debate, I think, of like what is the actual useful life of a chip?
是的。是的。所以,你知道,我曾公开说过这一点,而且我比较惊讶它引起了这么多关注,但你知道,我说过我们 7 年和 8 年的 TPU 仍然 100% 利用率。
Yeah. Yeah. So, you know, I've been on record of saying this and one of my more, I was surprised by how much pickup this had, but you know, I said that our seven and 8 year old TPUs are still at 100% utilization.
哇。
Wow.
我想我看到了这个。我想那是……是的,我没想到它会成为一个有意义的声明,但显然它是有意义的。是的。所以,我们的旧 TPU、GPU,但我们的旧 TPU 正在得到大量利用。我们最终会更换它们。有一个问题,好吧,它们是否,折旧寿命大约是 6 年,所以一旦它们完全折旧,并且考虑到新几代的能效等,更换和升级它们确实有意义。不是我们更换芯片,而是整个系统。换句话说,我们以 pod 为单位思考。所以一个 TPU,比如说一个 8 pod 可能是 9600 个芯片,可能是 140 多个,152 个机架等。所以我们会说,好吧,我们要拔出那个 pod,我们要用 TPU 12 或 13 或 14 pod 替换它。它可能不是完美匹配,然后我们必须,换句话说,新 pod 的占地面积可能不完全适合旧 pod 的空缺。我们只需要考虑到这一点。我们必须弄清楚如何改造。我们无法计划,因为我们不知道那么多代 TPU 会是什么样子。所以这又是一门艺术,弄清楚我们将如何,实际上很难弄清楚我们将如何退役,然后尽快用新的 TPU 替换。
And I think I saw this. I think that was... Yeah, I didn't mean for it to be a meaningful statement, but it was a meaningful statement apparently. Yeah. So, our older TPUs, GPUs, but our older TPUs are seeing significant utilization. We do replace them in the end. There's a question of okay have they, depreciation lifetime is approximately 6 years, so once they're fully depreciated and given the power efficiency of newer generations etc., it does make sense to replace them and upgrade them. It's not that we replace the chips but it's really the systems. So in other words, we think in terms of pods. So a TPU, let's say an 8 pod might be 9600 chips and it might be 140 something, 152 to racks, etc. So then we would say, okay, we're going to pull out that pod and we're going to replace it with a whatever it is, TPU 12 or 13 or 14 pod. It might not be a perfect fit and then we have to, in other words, the new pod's footprint might not be a perfect fit for the old pod's vacancy. We then just have to account for that. We have to figure out how we would retrofit. We can't plan that because we don't know what that many generations of TPUs are going to be. So then it's again a bit of an art to figure out how we would and it's actually hard work to figure out how we would decommission and then replace as quickly as possible with new TPUs.
你公开写过关于开放标准的文章。你能谈谈这个吗?
You've written publicly about open standards. Can you say a word on that?
是的。
Yeah.
所以我认为这就是我们谈到的互操作性问题。对我们来说,虽然我们支持垂直整合,也让你可以按需榨取尽可能多的性能,但对我们而言,非常重要的是不强制锁定,不强制在系统端到端运作方式上搞一种围墙花园。我举一个例子:在 Google,我们开发了一个模型开发框架叫 JAX。我们很喜欢它,觉得它非常非常好,内部也广泛使用。我们的很多客户喜欢 PyTorch。我们可以说,嘿,如果你想在 TPU 上运行,你就得用 JAX——它是最好的。我有点在开玩笑,但也许它是,也许它不是;我们认为它是最好的,而且这是你唯一的选择,因为它太棒了。或者我们可以说,看,如果你喜欢 JAX,我们也喜欢 JAX;如果你喜欢 JAX,你可以用它,但如果你喜欢 PyTorch,我们有 Torch TPU,让你未经修改的模型也能运行。我们在历史上见过很多这样的例子。我用过 IP 的例子:为什么互联网协议赢了?实际上有很多与 IP 竞争的协议——这是古老的历史,在 70 年代和 80 年代初。为什么 IP 赢了?因为它是一个开放标准,是沙漏的细腰,软件层面任何东西都能跑在上面,硬件层面任何东西都能跑在下面。所以开放标准、可互操作:任何带来 IP 的人都能插进路由器端口,就这样成为互联网的一部分。这很美,也正是这一点让互联网爆炸式增长并触及全世界。所以我们真的相信那些开放标准作为接入点。现在如果你想接入高度专业化的东西,你可以——如果你认为你有更好的东西可以接入我们的框架,你绝对可以。但我们想真正支持开放标准,最好围绕它也支持开源。
So I think that's the interoperability question we talked about. For us, while we support vertical integration and we make it possible to extract just as much performance as you want, for us it's really important to not force lock-in and not force a sort of walled garden in terms of how the system works end to end. I'll give one example: at Google we've developed a model development framework called JAX. We like it a lot, we think it's really good, and we use it extensively internally. Many of our customers like PyTorch. One thing we could say is, hey, if you want to run on TPUs, you've got to use JAX—it's the best. I'm being facetious a little bit, but maybe it is, maybe it isn't; we think it's the best, and that's your only choice because it's so great. Or we could say, look, if you like JAX, we like JAX; if you like JAX you can use it, but if you like PyTorch we have Torch TPU where your unmodified models can run. We've seen this throughout history, actually, with many examples. I've used the example of IP: why did the Internet Protocol win? There were actually many competing protocols to IP—this is ancient history, in the '70s and early '80s. Why did IP win? Because it was an open standard and it was this narrow waist to the hourglass where anything could run on top in terms of software and anything could run underneath it in terms of hardware. So open standard, interoperable: anyone who brought IP could plug into a router port and become part of the internet just like that. It was beautiful, and that's what allowed the internet to explode in its growth and reach across the world. So we really do believe in those open standards in terms of plug-in points. Now if you want to plug in something highly specialized, you can—if you think you've got a better thing to plug into our framework, you absolutely can. But we want to really support open standards, ideally open source around it as well.
这个建设规模太大、太重要了,不能成为任何一家供应商的封闭专有技术栈。
It's too big and too important of a buildout to be any one vendor's closed proprietary stack.
是的,我们真的相信这一点,而且它必须是一个有选择的东西。这也是为什么,例如,我们完全支持并拥有 TPU、GPU、其他加速器等等。
Yeah, we really believe that, and it's got to be one of choice. That's also why, for example, we fully support and have TPUs, GPUs, other accelerators, etc.
随着 AI 的发展,你团队的日常工作生活发生了怎样的变化?它在哪个方面对你的职能改变最大?
How has day-to-day life for your team changed with AI? Where is it changing your function the most?
我认为最简单的答案是在软件工程方面,这方面外部已经有很多记录了。我认为在 Google,我和我的团队正在非常有效地利用它来提升软件开发能力,而且坦白说也包括测试发布,甚至帮助设计等等。但在硬件方面——这方面外部报道可能少一些——也发生了重大变化。换句话说,我的硬件工程师们——事实上我今天早些时候刚看了数据——正在使用 AI,用 token 数来说,虽然不是最好的指标,但仍然是一个指标。我们的硬件工程师使用 AI 的程度和软件工程师一样多。生产力提高了。从设计启动到流片的时间在缩短。流片后调试的时间也在缩短。所以这是另一个生产力显著提升的领域。但还有一些也许不那么意料之中的变化:我们做数据中心设计的方式已经发生了显著变化。换句话说,我们如何评估——你提到,嘿,我们是做一个吉瓦级建筑,还是 100 兆瓦园区,还是 200 兆瓦园区。过去,这会是非常详细的、非常依赖电子表格的、人工驱动的流程,现在在一定程度上仍然如此,但现在有很多 AI 参与其中,真正简化了规划和开发方面的流程。
So I think the easiest answer is on the software engineering side, where it's been documented externally quite a bit. I think that we at Google, and my team, are using it to great effect in terms of our software development capabilities, but also frankly test rollouts, even helping with design, etc. But there have been on the hardware side, which has maybe gotten less external coverage, also significant change. In other words, my hardware engineers—I was just looking at the numbers earlier today, in fact—are using AI, using whatever the token counts if you want, not the best metric, but it's still a metric. Our hardware engineers are using AI as much as the software engineers are. And productivity has gone up. Time from design kickoff to tape-out is shrinking. Time for bring-up is shrinking. So another area where productivity is going up significantly. But then other maybe less expected changes: how we do data center design has changed significantly. In other words, how we evaluate—you mentioned, hey, do we do a gigawatt building or a 100-megawatt campus or a 200-megawatt campus. In the past, this would be very detailed, very spreadsheet-driven, human-driven processes, and it still is to some extent, but there's a lot of AI now involved that really streamlines the process in terms of planning and development as well.
有意思。所以,像是推理模型还是——
Interesting. So, like reasoning models or—
我还不至于说是推理模型。它没有取代人类判断,但它让把必要信息汇集到一处变得容易得多,基本上就是把正确的信息摆在做决策的人面前。
I wouldn't say quite yet reasoning models. It's not replacing human judgment, but it's making it much easier to bring the necessary information together in one place and basically put the right information in front of the humans making the decisions.
好,我要用两个有点天马行空的趣味问题来收尾。
Okay, I'm going to bring us home with two kind of out-there, fun questions.
好。
Yeah.
第一个问题:轨道数据中心。我看到 Google 在相当认真地考虑这件事,你也想让我对此发表评论,但我的意思是,人们认真地在算轨道算力的账,这是否意味着地球上存在某种严重的硬性约束?你怎么看轨道数据中心?
Question number one: orbital data centers. I've seen that Google is thinking about this quite seriously, and you want me to comment on that, but I mean, does the fact that people are seriously running the numbers on orbital compute mean that there are kind of serious binding constraints on Earth? And what do you think of orbital data centers?
是的,这是一个令人兴奋的方向。我们实际上正在推进它,而且我们毫不开玩笑地把它称为一个 moonshot,是我们兴奋地投入的宏大努力之一。这要回到对话早些时候你提出的——我认为你说得对——从根本性约束来看,能源和能源生产是一个关键挑战。所以底线是,从基本原理看,在太空中你能获得大约多 40% 的电力,因为大气中没有衰减等等。换句话说,就是更多的太阳能容量——大概 1.4 倍?所以这是其中一部分。但在太阳同步轨道上,你的太阳能电池有 98% 到 100% 的阳光覆盖,而在地面上是 28%、30%,也许 35%,对吧?所以换句话说,就是巨大量的电力。显然太阳——你算上 1.4 倍,再算上每天小时数的 3 到 4 倍,基本上就把电池从等式中去掉了,现在你就有潜力做出真正——当然还有无碳、很多好处,也有很多挑战。
Yeah, it's an exciting direction. We are actually pursuing it, and we've referred to it, with no tongue in cheek, as a moonshot, as one of our big efforts that we're excited about investing in. It goes back to the earlier part of the conversation where you raised, I think correctly, that in terms of fundamental constraints, energy and energy production is a key challenge. So bottom line is, from a fundamentals perspective, in space you have something like 40% more power available because of lack of attenuation in the atmosphere, etc. In other words, it's just more solar capacity—what, 1.4x? So that's one part of it. But in a sun-synchronous orbit, you have 98 to 100% coverage of sunlight on your solar cells, relative to 28, 30, maybe 35% on land, right? So in other words, just this enormous amount of power. Obviously the sun—you take the 1.4x and you take the 3 to 4x in terms of number of hours per day, you remove batteries more or less from the equation, and now you have the potential for something that can really—and of course carbon-free, lots of benefits, lots of challenges.
是的。
Yep.
对。很多很多挑战。所以换句话说,好,那冷却呢?天真地想,你可能会觉得太空里冷却更容易。实际上更难。可靠性——我们谈过这些东西有时会出故障。所以维修变得更难,不是不可能,但在太空中变得更难。用机器人上去处理集群——
Right. Lots and lots of challenges. So in other words, okay, now what about cooling? Naively you might think cooling in space is easier. It's actually harder. Reliability—we talked about how these things fail sometimes. So repairs become harder, not impossible, but it becomes harder in space. Robot up with the cluster—
正是如此,所以也许会有——这个模式是一个有前景的方向。
And exactly so perhaps there'll be—this model is a promising direction.
冗余可能会是你的朋友。你之前关于自由空间光学的问题现在要成为现实了。我们可能不会在这些组件之间拉光纤。所以实际上会是自由空间,我们会让激光指向接收器并实时校准。这里没有拦路虎。没有根本性的拦路虎。
Redundancy will probably be your friend. Your pre-question on free-space optics is now going to become reality. We're probably not going to be stringing fiber between these components. So it will actually be free space, and we're going to have lasers pointing at the receivers and calibrating in real time. There are no showstoppers here. No fundamental showstoppers.
最后一个问题。整个过程中,我脑海里一直有你们建造的那台难以言喻的大型超级计算机的形象。所以这里有一个问题。
Last question. I've had this image of the ineffable big supercomputer that you built in my head this whole time. And so here's a question.
10 年后,最前沿的超级计算机长什么样?
10 years from now, what will the most frontier supercomputer look like?
哦,天哪。是啊。10 年正好处在一个边缘,我真诚地说,可能有些人在思考 2036 年的计算机会是什么样。我们可能只有少数人会想那么远,但哇,那里的不确定性锥体实在太宽了。换句话说,如果你看趋势,集成度将会非常巨大。所以,你知道,我们看到了漂亮的光纤和我们为 ineffable 组装的那批 VR200 机架。我的猜测是它会更加集成,光纤更少——我不会说 2036 年没有光纤,但我认为从机架的角度来看,机架会看起来更加紧密集成,光纤的量——可能只有一小束光纤从机架里出来。然后,换句话说,从模块化制造的角度来看,一个观点是,2036 年这些机架很可能会集中制造。所以无论它们有 72 个还是 144 个还是 288 个,或者可能更多,你知道,576、52,选你最喜欢的 GPU 倍数——既然我们在谈 ineffable 的情况,但 TPU 也会一样——深度集成到一个机架中。一个机架会是数兆瓦吗?你知道,我们不知道具体怎么实现,但 2036 年,我们可以想象单个机架可能是数兆瓦。现在你引入水、引入电力、引入光纤,对吧?你把那个机架推入,插上这三样东西,你就可以开始比赛了。
Oh my gosh. Yeah. 10 years is right at that edge where it's, I would say in all sincerity, probably some people are thinking about what the 2036 computer looks like. We might have a few thinking that far ahead, but wow, the cone of uncertainty there is way, way too wide. In other words, if you look at the trends, the level of integration is going to be tremendous. So, you know, we saw the beautiful fiber and the VR200 racks that we put together for ineffable. My guess is it's going to be more integrated, less fiber—I won't say no fiber in 2036, but I think that probably from a rack perspective, the rack will look much more tightly integrated, and the amount of fiber—there'll be maybe a small fiber bundle coming out of the rack. And then, so in other words, one view you might have from a modular manufacturing perspective is these racks likely in 2036 are going to be manufactured centrally. And so whether they have 72 or 144 or 288 or probably more, you know, 576, 52, pick your favorite multiple of GPUs—since we're talking about the ineffable case, but then be the same for the TPUs—deeply integrated into a rack. Would a rack be multiple megawatts? You know, we don't know exactly how, but 2036, we get to imagine things could be multiple megawatts in a single rack. And now you bring water and you bring power and you bring fiber, right? You wheel that rack in, you plug those three things in, and you're off to the races as well.
好吧。所以,到 2036 年,它不会像太空中一个巨大的外星球体那样。
Okay. So, it's not going to be like a big alien orb then up in space by 2036.
我不会说 2036 年不可能直接发射到太空,然后也许被某个空间站的机械臂抓住,再插入正确的模块。是的,2036 年也许可以那样。
I wouldn't say that it's not possible for 2036 for it to be then directly launched into space and then maybe taken by a robotic arm at some space station and then plugged into the right module. Yeah, maybe that for 2036.
想想就很有趣。
Fun fun fun to think about.
嗯,我真的很享受这次对话,这是历史上最激烈、最大的资本支出,但我也认为这是一场技术革命,一场美丽的革命,我认为你对技术的美、约束条件以及如何平衡这一切有着深刻的理解。谷歌在可靠的人手中。感谢你抽出时间分享你正在做的事情。
Um I really enjoyed this conversation where this is the most intense biggest capback spilled out in history, but I also think it's a technology revolution, a revolution that is a thing of beauty and I think you have such a grasp over, you know, the beauty of the technology and the binding constraints and how to balance this all. Uh Google is in good hands. Thank you for taking the time to share what you're doing.
非常有趣。非常有趣。我们活着。我们正在经历这一切,并且我们有机会定义它。所以非常感谢。这是一次很棒的对话。
Ton of fun. Ton of fun. We get to live. We get to live through this and we get to define it. So thank you very much. This is a great conversation.
谢谢。
Thank you.
是的。谢谢。
Yeah. Thanks.