The Largest Infrastructure Build in History: Inside OpenAI's Compute Strategy
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 工业计算负责人 Sachin Katti 探讨 AI 数据中心的惊人规模、电网限制以及 AI 如何开始设计自己的芯片。
Sachin Katti, Head of Industrial Compute at OpenAI, discusses the staggering scale of AI data centers, power grid constraints, and how AI is beginning to design its own chips.
每当我们觉得算力够了、可以放慢脚步时,现实总会给我们一个负面的惊喜——我们本不该放慢。如今需求远远超过算力供给,所以我们能上线的一切都会被立即消耗掉。这仍然是我们最大的担忧。在我们试图获取和构建算力的规模上,物理世界并没有那么快的速度。我们确实相信,递归的世界并不遥远——AI 将设计它所需用来训练和运行下一代 AI 的系统,包括芯片。
Anytime we have thought we have had enough compute, we can slow down. Always negatively surprises like we should not have slowed down. Demand far outstrips compute supply today. So, anything we can bring online we consume immediately. Our biggest worry is that still. At the scale at which we are trying to get compute and build compute, the physical world does not move that fast. We do believe that the world of recursion is not that far where AI will design the systems it needs to train and run the next generation of AI including chips.
大家好,我是 Matt Turk。欢迎收听 Matt Podcast。今天的嘉宾是 Sachin Katti,他拥有当下科技界最迷人且最相关的头衔——OpenAI 工业计算负责人。Sachin 的背景令人惊叹:他曾是斯坦福大学教授、多次创业者,最近还担任过英特尔 CTO。现在,他正在领导许多人称之为人类历史上最大的基础设施建设项目。在这一集中,我们将跳出模型层,深入探讨算力以及 AI 热潮的物理现实。我们会谈到正在建设的数据中心的惊人规模,深入探讨液冷超级计算机、电网限制、核能的潜力,以及 OpenAI 通过“jalapeno”项目进军定制芯片。我们还会讨论更广泛的启动策略,以及一个略显清醒的现实——AI 已经开始帮助设计自己的芯片。这是一次绝佳的幕后揭秘,看看为智能的未来提供动力究竟需要什么。请享受我与 Sachin Katti 的对话。
Hi, I'm Matt Turk. Welcome to the Matt Podcast. My guest today is Sachin Katti who holds what might be the most fascinating and relevant title in tech right now, Head of Industrial Compute at OpenAI. Sachin has an incredible background. He was a professor at Stanford, a multi-time founder, and most recently the CTO at Intel. Now, he's leading what many are calling the largest infrastructure build out in the human history. In this episode, we step away from the model layer and dive deep into compute and the physical reality of the AI boom. We talk about the staggering scale of the data centers being built. We get into the weeds on liquid-cooled supercomputers, power grid constraints, the potential of nuclear energy, and OpenAI's move into custom silicon with jalapeno. We also discuss the broader start strategy and the slightly sober reality that AI is now beginning to help design its own chips. It is a fantastic look behind the curtain at what it actually takes to power the future of intelligence. Please enjoy my conversation with Sachin Katti.
好的,Sachin,欢迎你。很高兴能进行这次对话。我们是在巴黎 Race 会议的间隙录制的,所以感谢你顶着酷暑前来。
All right, Sachin, welcome. Uh excited to do this. Uh we are recording this on the sidelines of the Race conference in Paris. Uh so, thank you for braving the the heat.
是的。谢谢。很高兴来到这里。
Yeah. Thank you. Great to be here. Great to be here.
这里又是一波热浪。首先,有些人将当前计算和数据中心领域发生的事情描述为历史上最大的基础设施建设,比高速公路、比铁路还要大。我很好奇,第一,你是否同意;第二,从内部看是什么感觉?你如何看待 OpenAI 目前正在建设的东西?
Uh it's another heatwave here. To start, some people describe what's currently happening in the world of compute data centers as the largest infrastructure build out in history, bigger than uh the highway, uh bigger than the railroads. And I'm I'm curious one, if you agree, and and two, what it feels like on from the inside. Like what's How do you view what you're currently building at OpenAI?
是的,这确实感觉像是人类有史以来建造的最庞大的工程之一。绝对比我听说过的许多工程都要大。我年纪不够大,没有经历过高速公路建设。但没错,这感觉就像听起来那样——我可以说是在“巨兽的腹中”。每一天我们都在做决策,这些决策的规模在我之前的角色(比如在英特尔)可能需要几个月才能做出。但需求如此贪得无厌,而且增长如此迅速,我们必须非常快速地行动。所以这是一个紧张的时期,但可能是工程师最想参与的最激动人心的事情。
Yeah, it definitely feels like one of the largest things humanity has ever built, effectively. Definitely bigger than many of the things that I've heard of. I'm not old enough to have experienced the highway buildout. But no, it's feels exactly like what it sounds I'm in the in the belly of the beast, so to speak. Every day is we are making we are making decisions. We are compute that historically from my previous role, for example, at Intel, we'd probably take months to make given the magnitude of the those decisions. But the demand is so insatiable that and it is growing so rapidly that we have to move very quickly. So it's an it's an intense time, but it's probably the most exciting thing an engineer would want to be part of.
是的,我在某处读到 OpenAI 计划今年在算力上花费约 500 亿美元。这个数字方向性上还大致正确吗?
Yeah, and I read somewhere that OpenAI was planning on spending about 50 billion compute this year. Is that Is that still the rough number directionally?
实际上这听起来差不多。是的,整个行业今年的算力支出也将达到 7000 亿美元。所以这数字简直疯狂。
It actually that sounds about right. Yeah, and the whole industry itself was going to be 700 billion in compute spend this year as well. So like insane insane numbers. It seems.
是的,而且可能还在继续增长,对吧?很多建设正在进行中。所以其中很大一部分也会在一两年内转化为像我们这样的人的算力使用。
Yeah, and it's probably continuing to grow, right? And a lot of build happening. So a lot of that is also going to translate to compute usage from people like us in a year or two.
这样理解是否正确:对 OpenAI 来说,这有点像是一个新世界?显然,所有 AI 研究都在全速推进,但在公司内部建设一个全新的业务。这样讲公平吗?人们是这样想的吗?因为显然构建模型是一回事,构建数据中心则完全是另一个世界。
Is the right way to think about this that for OpenAI it's it's a bit of a new world, right? There's obviously not quite a a because obviously the all AI research is going you know full speed ahead but like building a whole new business within the the the company. Is that Is that fair? Is that how people think about it? Because obviously building models is one thing, building data centers is a whole different world.
是的,我认为 OpenAI 一直有一个基本信念:计算机是一切的基础,对吧?计算机是智能的基础。我们持续扩展和分发智能的方式就是拥有算力。这一点从未改变,这一直是我们的信念。我认为越来越清楚的是,要构建我们需要的算力,并且达到我们必须的规模,不能仅仅依赖从合作伙伴那里获取算力。我们越来越需要扮演更积极的角色,去构建和获取我们需要的算力。所以这绝对感觉像是我们在公司内部锻炼的一块新肌肉。
Yeah, I mean I think OpenAI has always has had a fundamental belief that computers are the foundation of everything, right? Computers are the foundation for intelligence. And the way we keep continuing to scale intelligence and distribute intelligence is by having compute. And so that has never been different. That has never always been the belief. I think what's becoming clear is to build the kind of compute we need and and at the scale we have to not just rely on getting compute from our partners. We increasingly have to take a much more active role in building and getting that compute that we need. So it does absolutely feel like a new muscle that we're building in the company.
也许从一开始就锚定对话,谈谈数据中心在现实中到底是什么,会非常有帮助。我想大家都知道数据中心正在建设,但说实话,我不确定每个人都能说清楚到底在建什么,因为几十年来我们一直在为云计算建设数据中心。那么,我们今天为 AI 建设的数据中心,有什么根本性的不同和新颖之处呢?
And maybe to anchor the the conversation from from from the beginning, it would actually be very helpful to talk about what a data center is in reality. So I think like everybody knows that data centers are being built but you know going to one's head like I'm not sure that everybody could say well what is actually being built because we've been building data centers over the industry for cloud for for decades at this point. So what is fundamentally different and new about the data centers that we're building for AI today?
我认为最大的不同可能是规模,对吧?所以我们本质上是在建造大型超级计算机,正如我们对 AI 的思考。随着我们构建智能、交付智能,模型变得越来越强大,我们将其用于越来越复杂的任务。我们需要越来越大的计算机。所以我认为我们将数据中心视为巨大的工厂,对吧?它们将电子转化为 token。这是现在流行的一句话,但它确实很有道理。那么,我们如何获取电力?如何利用这些电子来驱动芯片,从而有效地交付智能?我将其想象成大型足球场,采用液冷,因为这些芯片运行温度非常高。芯片上的温度非常高,所以你必须用液体冷却它们,不能用空气。所以,很多液冷设备——基本上就是冰箱——放置在建筑旁边。
I think the biggest probably is is a scale, right? So we are essentially building large supercomputers as we think about AI. And as we build intelligence and deliver intelligence and models become more capable, we use it for more more and more complex tasks. We need more and more bigger computers effectively. And so I think the way we visualize data centers is giant factories, right? That are turning uh electrons into tokens. Uh that's a popular phrase nowadays. But that it's it actually has a lot of a ring of truth to it. So, how do we take power? How do we take those electrons and actually use it to power chips that effectively are delivering intelligence? Uh but the way I visualize it is large football fields, uh liquid cooled because these chips run really hot. Uh the temperatures on these chips are very very high. And so you have to cool them with liquids. You can't cool them with air. So, a lot of liquid cooled uh basically refrigerators effectively uh that are sitting in the alongside the building.
那么,冷却是在数据中心层面进行,还是在芯片层面?还是两者都有?
And you know, on that a big wall where it's uh the the the cooling happens at the data center level or does it happen at the chip level? Or both?
两者都有。是的。所以你需要冷却数据大厅,但也需要单独冷却芯片,因为只做其中一项是不够的。而且你还需要冷却连接芯片的东西,对吧?这就是为什么现在几乎到处都需要冷却。甚至分配电力的电缆和变压器也会变得过热,所以它们也需要冷却。因此,任何处理能量的东西都会产生热量。
Both. Right. So, you need to cool the data halls, but you also need to need to cool the chips individually because it's not going to be enough to do one or the other. And uh you also have to cool the things that connect chips, right? And that's why you need cooling pretty much everywhere uh nowadays. Even the cables that are the transformers that distribute the power become too hot. So, they also need to be cooled. So, everything that processes energy produces heat.
那么,目前使用的冷却技术是已经成熟、正在部署的,还是说冷却领域正在发生根本性的新变化?
And is cooling technology that's being used something that's well understood and is just getting deployed, or is there fundamental new things happening in cooling right now?
我认为液冷技术已经存在一段时间了,但从未在如此大规模下部署。因此,创新更多在于如何使其可靠、更便宜、更具可扩展性。这方面有很多创新。还有很多新的创新,比如新型液体、新型材料,它们能更好地吸收热量,因为任何能提高热传递效率的东西对数据中心都非常重要。这样我们就可以让芯片运行得更热。芯片运行温度与计算机性能之间存在直接关联。芯片越热,你获得的内存带宽就越多,每秒浮点运算次数也越多。所以回报很丰厚。如果你能很好地冷却,那也意味着你能产出更多智能。
I think liquid cooling has been around for some time, but has never been deployed at this scale. So the innovation is more around how to make it reliable, how to make it cheaper, more scalable. There's a lot of innovation around that. There's also a lot of new innovation, new kinds of liquids, new kinds of materials that can absorb heat better, because anything that can improve the efficiency of heat transfer is very important for data centers. So we can then run the chips hotter. There is a direct correlation between running a chip hotter and how powerful the computer is. The hotter the chip, the more memory bandwidth you get, the more flops you get. So there's a strong payoff. If you can cool well, that also means you can produce more intelligence.
好的。所以是巨大的工厂,大量的冷却。另一个在我看来对任何讨论都至关重要的部分是电力和能源。那么,从高层次来看,这是如何运作的?你们是接入电网,还是自己发电?
All right. So gigantic factories, lots of cooling. The other part that seems to me very critical to any discussion is power and energy. So how does that work starting at a high level? Do you connect to the grid? Do you have your own power generation?
我认为早期我们都接入电网,而且我们仍然希望接入电网。目前我们开始遇到限制,我们正在投资电网的发电基础设施和输电基础设施。因此,无论我们在哪里建设数据中心,我们都坚定承诺不会从电网中夺走电力。事实上,我们正在投资电网以产生新的电力,这样我们就可以将其用于数据中心。
I think in the early days we all connected to the grid, and we still all would want to connect to the grid. At this point we are beginning to hit limits, and we are investing in generation infrastructure for the grid, transmission infrastructure for the grid. So whenever we build a data center anywhere, we make it a hard commitment that we are not taking power away from the grid. In fact, we are investing in the grid to generate new power so that we can consume it for data centers.
这在实际中意味着什么?
What does that mean practically?
某个地方有一个电网。它有一定的发电和配电能力,一定的兆瓦数。显然,一个数据中心出现了。如果有闲置容量,那么数据中心当然可以使用。但如果没有闲置容量,那么我们就必须为电网增加新的天然气、太阳能或水力发电基础设施。所以我们投资并资助这种建设。然后你必须建设输电线路,投资变压器、变电站来分配电力。因此,无论我们在哪里建设数据中心,我们都在资助所有这些基础设施的开发。如果没有这些数据中心,这些基础设施本来不会得到资助。这种大规模数据中心建设的一个附带好处是,美国乃至全世界的电网基础设施正在迅速升级。这就是电力部分。只要我们能做到,我们就这么做,我们从电网消耗电力,但我们也成为电网的好公民,因为我们为每个人改善了基础设施,不仅为数据中心,也为家庭。在一些地方,我们开始达到电网电力建设和消耗的极限。所以每个人都在关注表后发电。我们也在进行一些表后发电,即我们拥有不来自电网的现场发电和配电能力,实际上数据中心在电力方面变得自给自足。
You have a grid somewhere. It has a certain generation and distribution capability, a certain number of megawatts. Obviously, a data center shows up. If there is spare capacity, then of course the data center can use it. But if there isn't spare capacity, then we have to add new gas or solar or hydro generation infrastructure to the grid. So we are investing and funding that build-out. And then you have to build transmission lines, invest in transformers, substations to distribute that power. So wherever we build data centers, we are funding the development of all of that infrastructure. This is infrastructure that would otherwise not have been funded if not for these data centers. One of the side benefits of this big data center build-out is the grid infrastructure of America and the whole world, for that matter, is getting upgraded very quickly. So that's the power piece. Whenever we can do that, we do that, and we consume power from the grid, but we are also being good citizens of the grid because we are improving the infrastructure for everyone, not just for data centers, but also for households. In some places, we are beginning to hit the limits of how much grid power we can build and consume. So everyone is looking at behind the meter. We are also doing some behind the meter generation, where we would have on-site power generation and distribution capability that did not come from the grid, but in fact, the data center becomes effectively self-sufficient in terms of power.
它使用燃气轮机吗?
Does it use gas turbines?
目前,是燃气轮机,尤其是在美国,因为这是最密集、最可运输的能源形式,而且在美国也相当广泛可用。但那里,我们说的是次要的。
Today, it's gas turbines, especially in the US, because that's the most dense, transportable form of energy, and also one that is quite widely available in the US. But there, we are talking about the sub-mention.
你认为核能讨论有趣吗?也许我们在法国录制,这里有很多核能发电。核能系统在美国也重新回到了讨论中。这是你考虑的事情吗?你觉得它有趣吗?
Do you think that the nuclear conversation is interesting? Maybe we're recording this in France, which has a bunch of nuclear power generation. Nuclear systems have come back to the discussion in the US as well. Is that something that you think about? Do you think it's interesting?
当然。它来得越快越好。我认为这是我们可以生产和消费的最密集的能源形式,而且也很清洁。所以我认为它绝对是数据中心大规模可扩展能源的良好来源。显然,在法国之外,世界其他地区在建设这种基础设施方面还有很多追赶工作要做,但我认为它将在数据中心领域发挥非常重要的作用。
Absolutely. It can't come soon enough. I think it is the densest form of energy we can all produce and consume, and it's also clean. So I think definitely it would be a good source of massive scalable energy for data centers. Obviously, outside of France, the rest of the world has a lot of catching up to do in building this infrastructure, but I think it's going to play a very important role in the data center world.
好的,这是对数据中心的精彩介绍。你们最近另一个有趣的新闻是 Jalapeno。那么现在对于 OpenAI,除了应用业务(消费者和企业)、模型 AI 研究业务以及计算数据中心业务之外,似乎 OpenAI 也进入了芯片业务,如果这么说公平的话。所以完全是全栈的,但我很好奇整体战略。这适合哪里?
Okay, so that's a great introduction on data centers. The other interesting bits of news that you guys recently had is Jalapeno. So now for OpenAI, in addition to being the application business (consumer and enterprise), and then being in the model AI research business, and the computing data center business, it seems that OpenAI is in the chip business, if that's fair. So completely full stack, but I'm curious about the overall strategy. Where does that fit?
随着我们开始,世界人口的很大一部分,AI 的使用正在爆炸式增长。推理显然成为我们工作负载的很大一部分,并且消耗大量算力。另一个认识是,因为我们确切知道工作负载是什么,我们要运行的模型是什么,我们可以共同设计硬件,使其对这些模型非常高效。因此,Jalapeno 背后的战略论点是,我们如何利用对最终工作负载和模型本身的了解,设计出在服务这些模型方面非常高效的芯片。它确实让我们能够推动效率优势,提高每瓦特产生的 token 数。Jalapeno 优化的关键指标是最大化每瓦特能产生的 token 数。因为当今世界受限于电力,所以在相同电力下能产生更多 token 对每个人都更好。因此,我们将其视为扩展我们向世界交付智能的关键要素。
As we begin, a pretty big fraction of the world's population, AI usage is exploding. Inference is obviously becoming a big fraction of our workload, and it's consuming a lot of compute. One of the other realizations is because we know exactly what the workload is, what the model we want to run, we can co-design the hardware to be super efficient for those models. So the strategic thesis behind Jalapeno is how do we take advantage of knowing what the end workload is, what the model itself is, and design chips that are very efficient in serving those models. It really allows us to drive the efficiency advantage, drive more tokens per watt. The key metric that Jalapeno is optimizing is maximizing the number of tokens you can produce per watt. Because the world is constrained by power today, so the more tokens you can produce for the same amount of power, it's better for everyone. So we look at it as a very critical ingredient in scaling how we deliver intelligence to the world.
不具体评论 OpenAI 内部的情况,就目前而言,推理在算力使用上是否同样远大于训练?我们是否已经从那些非常繁重的预训练运行作为算力的主要用途,转向了推理占多数?
Without commenting on necessarily what's going on at OpenAI specifically, is inference equally being much bigger than training these days in terms of usage of compute? Are we shifted from those very heavy pre-training runs as the major use case for compute to now just inference being the majority?
不,推理规模很大,甚至可能占算力的大多数。而且我认为我们不喜欢区分训练和推理,因为现在很多训练本身就是推理。例如,当我们训练新模型时,我们在生成合成数据,那是推理;当我们训练新模型时,我们在做后训练,那也是推理;当你训练模型时,你在做测试时算力,那全都是推理。所以当我们说训练时,很多算力实际上就是推理,即使是在那个工作阶段。因此推理是一个基础构建模块。
No, inference is big, perhaps even the majority of compute. And I think one of the things is we don't like to make a distinction between training and inference because a lot of training is now inference. So when we train a new model, we are generating synthetic data for example. That's inference. When we train a new model, we are doing post-training and that's inference. When you train a model, you're doing test time compute. That's all inference. So when we say training, a lot of the compute actually is inference even in that phase of the work. So inference is a fundamental building block.
是的。显然我忍不住要问一个不可避免的问题,关于过度建设的潜在风险,考虑到需求和使用之间的滞后以及建设数据中心需要多长时间。你曾在某处提到,你故意对未来三年的问题非常偏执。我们正在为未来的意外做准备,这似乎是一个非常健康的方法。那么,你是怎么想的?有没有办法缓解这种风险,还是说就像我要玩一样?你是否相信这就是未来,我们只会一直向前冲?
Yeah. Obviously I cannot resist asking the inevitable question around the potential risk of overbuilding given the lag between demand and usage and how long it takes to build a data center. And you mentioned somewhere that you were deliberately very paranoid about the problems ahead in the next three years. We're planning on that the surprises ahead, which seems like a very healthy approach. So, how do you think about that? Is there any way to mitigate that or is it just like I'm going to be playing? Do you have any belief that this is the future and we're just all going to go go go?
我们对 Scaling(规模扩张)有坚定的信念,对吧?历史已经证明了这一点。例如,我们的收入实际上与之同步:我们增加三倍算力,收入也增加三倍。我们相信这将继续成立。今天需求远远超过算力供应。所以我们能上线的一切都会立即被消耗。至少对我们来说,没有算力会被浪费。所以我认为这种信念完全没有改变。而且,我们看到研究和训练上的缩放定律仍然成立,并且我们进行研究的步伐可能正在加速,对吧?因为 AI 本身。现在 AI 正在做很多 AI 研究。其中一个微妙的影响是,以前我们的研究人员运行实验需要算力,但他们能运行的实验数量受限于人类研究人员的数量,而这是世界上稀缺的资源,对吧?没有多少人能做 AI 研究。现在 AI 本身可以做 AI 研究,我们能利用的工具数量激增,因此研究所需的算力也爆炸式增长。所以我们看不到在可预见的未来我们会拥有无限算力的世界,对吧?当我提到意外时,我的担忧更多是负面的:我们实际上无法建设所有你想要的算力。而这正是我们的另一面,对吧?因为一直以来,每当我们认为算力不足时,我们可以放慢速度。但总是出现负面意外,比如“哦,我们不应该放慢”。对吧?所以我们最大的担忧仍然是这个。而且在我们计划获取和建设算力的规模上,物理世界不会移动得那么快,对吧?物理供应链、工厂不会那么快移动,也无法那么快增加产能。所以对我们来说,意外更多是来自那个方向,而不是另一个方向。
We have deep conviction in scaling, right? And history has borne us out. So, effectively our revenue, for example, has tracked it. We triple compute and we triple revenue. And we believe that continues to be true. Demand far outstrips compute supply today. So, anything we can bring online we consume immediately. So, there's no compute that is going to waste for us at least. So, I think that conviction has not changed whatsoever. And if anything, we are seeing that scaling laws on research and training continue to hold and potentially the pace at which we are doing research is accelerating, right? Because of AI itself. So, AI is doing a lot of AI research now. And so, one of the subtle implications of that is previously our researchers used to run experiments and they needed compute to run experiments. But the number of experiments they could run was limited by the number of human researchers they have, which is a scarce resource in the world. Right? There's not a lot of people who can do AI research. Now AI itself can do AI research, the number of experiments we can draw next tools. And therefore, the amount of compute you need for research also explodes. So, we don't see a world where we will have unlimited compute for the foreseeable future, right? When I was referring to surprises, my worry is more on the downside of we are not able to actually build all the compute that you want. And this is where the other way for us. Right? And because that is consistently the case, anytime we have thought we have having less compute, we can slow down. Always negatively surprises like, oh we should not have slowed down. Right? And so our biggest worry is that still. And at the scale at which we are planning to get compute and build compute, the physical world does not move that fast. Right? Physical supply chains, factories don't move that fast or cannot add capacity that fast. So for us the surprise is more on that direction than the other direction.
这很有趣。你刚才提到了社区,这显然是一个关键辩论。所以很好奇你的观点,在一个谱系上,一方面一个极端你会说,嗯,AI 行业和计算机行业有公关问题,没有问题,只是我们无法解释得足够好;另一个极端实际上那些社区有道理。你认为现实是什么?
That's fascinating. You alluded to communities a minute ago and obviously that's a key debate. So curious about your perspective on a spectrum where on the one hand one extreme you'd say, well, the AI industry and computer industry has a PR problem and there's no problem it's just that we cannot explain it well enough to the other extreme actually those communities have a point. What do you think the reality is?
我认为每当有像这样革命性的新技术出现时,总会发生颠覆。但你知道,我们从历史中学到,这总是为社会带来更好的结果,对吧?那么我们如何从今天的位置画一条线到那个结果呢?并向世界解释为什么这是我们都需要走的轨迹。我的意思是,这是我们的责任。在社区方面,有一点局部与全局的问题,对吧?我认为数据中心即使在今天对每个社区都是净正面的。例如,我们在美国农村地区建设这些数据中心,那里没有其他东西在建。我只是开玩笑。所以我们出现在德克萨斯州农村。我们建设一个数据中心,为社区带来新的财产税,资助学校,资助医院。我们出现并投资于新的电网基础设施,否则这永远不会发生,因为没有需求。所以那个地区可以享受现代化的电网。基本上我们创造了就业。很好。所以我认为我们大力投资的一件事是,每次在某个地方建设数据中心时,解释当地的好处,并确保人们充分理解这带来的好处。而且数据中心一旦建成,本质上是非常清洁的居民,对吧?它们不产生任何气体或有毒化学物质,对吧?它们是自包含的,只产生智能。
I think anytime there's new technology which is as revolutionary as this technology is, there is always disruption that's going to happen. But you know, we have learned this over history that this always leads to better outcomes for society, right? And so how do we draw a line from where we are today to that outcome, right? And explain to the world why this is the trajectory we all need to be on. I mean, it's our responsibility to do that. On the communities front there's a little bit of a local versus global issue, right? The communities, I think data centers are even today a net positive to every community. It is we are building these data centers in rural areas of America for example, right? Where there's nothing else that is being built. I'm just kidding. So we show up in rural Texas. We build a data center that produces new property taxes into the community. That funds schools, that funds hospitals. We show up and we invest in new grid infrastructure. Which otherwise would never have happened. Because there's no demand. So there's a modernized grid that that area can enjoy. Basically we produce jobs. Nice. And so I think one of the things we are investing a lot in is explaining the local benefits every time we build a data center somewhere. And making sure that it is well understood the kind of upside that this has. And data centers, once they are built, are essentially very clean citizens. Right? They don't produce any gases or toxic chemicals or anything. Right? They're self-contained. They just produce intelligence.
是的。一个典型的问题是水。我认为这已经被研究驳斥了不少,但也许给我们讲讲你对水问题的看法。
Yeah. Typical question that comes up is water. And I think that's been debunked quite a bit by research, but maybe give us just color on your outlook on the water question.
我们使用液冷,液体是循环利用的。所以实际上数据中心的耗水量相对于家庭用水来说小得惊人。所以我认为正如你所说,这已经被驳斥了。认为数据中心消耗大量水是一种误解。相反,它们做的事情消耗的水非常少。而且所有的水都是循环利用的。所以一旦达到某个点,我们不会净消耗新水。水在液冷过程中被循环利用。
We use liquid cooled and the liquid is recycled. So we do actually the water consumption of a data center is shockingly small relative to household water consumption. So I think as you put it, it's been debunked. It's a misperception that data centers consume a lot of water. If anything, they consume so little water for what they do. And all of that water is recycled. So we don't net consume new water once we get to a particular point. The water just gets recycled as it is liquid cooled.
是的。所以所有那些关于“棕水”的故事都没有意义,因为数据中心的水是在一个封闭回路里。
Yeah. So all those stories about like brown water is that just don't make sense because the water at the data center is in like a contained circuit.
这是一个闭环。是的,一个闭环。
It's a closed loop. Yes, it's a closed loop.
顺便说一下,你提到了德克萨斯州的农村地区。既然我们在对话开始时讨论了数据中心,为什么 OpenAI 和其他公司选择农村地区?比如你们如何选择数据中心的地点?
And by the way, you mentioned Texas in rural areas. Since we talked about data centers at the beginning of this conversation, why do OpenAI and other companies pick rural areas? Like how do you select a site for a data center?
很多因素。其中一个当然是土地,比如充足的土地。
So many factors. So one is of course land, like plentiful land.
第二点是许可审批。我们能否建造这些设施?我们希望建造时不影响周边社区。因此,选址理想的是相对偏远的土地。当然,电力接入也很重要,比如强大的电网和充足的天然气供应。这些都是重要因素。第四点是劳动力。你能多快建成?劳动力的可用性,包括建筑工人、合格电工、水管工等,都很关键。所有这些因素都会影响每个站点的选址决策。我知道德克萨斯州很受欢迎,因为它符合很多条件,但并非唯一选择。我们在全国都有数据中心,包括洛杉矶等地。
Number two would be permitting. Can we build these things? And we want to build them such that they are not affecting any neighborhoods. So land that is somewhat removed is the ideal candidate. Of course, access to power is important. So a strong grid, strong gas availability. All of those are important factors. And then four is labor. So how quickly can you build these things? Availability of labor, construction labor, qualified electricians, plumbers, all this stuff comes into play. So all of those factors go into every single site selection decision. I know Texas is popular because it fits a lot of these criteria, but it's not the only state. We have data centers all around the country in LA plus.
好的,太好了。我们稍后会详细讨论这些,但先聊聊你和你的经历。你是 OpenAI 的工业计算负责人。顺便说一句,“工业计算”这个头衔对于我们当下的时代来说非常贴切。但这具体意味着什么?这个角色是做什么的?在 OpenAI 内部,这项工作是如何组织的?在你能透露的范围内谈谈。
Okay, great. We're going to go into all of this in more detail, but let's talk about you and your journey. You're the head of industrial compute at OpenAI. By the way, the title 'industrial compute' is such a perfect title for the moment we're in. But what does that mean? What is the role and how is this effort organized within OpenAI, to the extent you can talk about it?
可以把我的角色和团队的角色理解为:我们如何以工业规模将算力投入使用?这就是我们实际在做的事情。这涉及整个生命周期。所以,我们如何找到算力所需的要素:土地、电力、机架、芯片、散热?我们如何为它们融资?因为这些都是巨额资金。所以,我们如何确保为电网基础设施、基础设施建设和芯片提供资金?然后是如何运营这一切?如何确保这些事情按时完成并保持运行?如何运营所有这些基础设施?这就是整个生命周期。当然,还有如何实际使用算力。所以,我的角色很大一部分是 OpenAI 内部的容量分配。算力始终是一种稀缺资源。
Think of it as my role and our team's role: how do we bring compute online at industrial scale? That's effectively what we're doing. It's the entire life cycle. So, how do we find the ingredients that are going to compute: land, power, shelves, chips, heat? How do we finance them? Because these are massive dollars. So, how do we make sure that we finance the grid infrastructure, the construction of the infrastructure, and the chips? Then it's about how do you operationalize all of this? So, how do you actually make sure these things happen on time and stay up? How do you operationalize all of this infrastructure? So, it's that entire life cycle. And then, of course, how do you actually use the compute? So, a big part of my role is capacity allocation inside of OpenAI. It is always a scarce resource.
听起来你人缘很好。
You sound very popular.
并不受欢迎。无论你做什么决定,总会有人不满意。但我们的团队为容量分配决策提供输入。我们揭示不同的选择点,以及不同分配方案下的假设问题。所以,容量规划,然后当然,利用这些来预测我们在哪里需要多少容量。因为不仅仅是更多的算力,还包括在哪里、什么类型、什么形态、什么芯片、你想在那里运行什么工作负载。所以,这些都是这个团队要弄清楚的:预测和规划应该是什么。这为决策提供信息并形成闭环。这指导我去哪里寻找下一块土地、电力和芯片来部署。
Not very popular. There is always someone who is unhappy with whatever decisions you make. But our team provides the input to make the capacity allocation decisions. We surface what are the different choice points and what are the what-if questions on different allocation choices that we have. So, capacity planning and then, of course, using that to forecast how much capacity we need where. Because it's not just more compute, it's also where, what kind, what shape, what chip, what workload you want to run there. So, all of those are what this team figures out: what should be the forecasting and planning. And that informs and closes the loop. So, that informs where I go find the next chunk of land and power and chips to put in there.
在不涉及机密细节的前提下——虽然我猜当你们上市后,这一切很快就会公开——但现阶段是数千人吗?是多个不同的团队吗?还是你们有很多事情在推进,或者有很多承包商?
Without going into confidential details, although I guess when you guys go public, all of this will be public soon. But is that thousands of people at this stage? Is it multiple different teams? Or do you guys have to ask a bunch of things in works, or a bunch of contractors?
这是一种组合策略。我们永远不会把所有事情都外包或完全依赖外部。这始终是一种混合模式,因为这样做才合理。你不能把所有鸡蛋放在一个篮子里。所以,我们会让超大规模云服务商提供很大一部分算力,甚至大部分算力。我们也会将新的云服务纳入组合。我们会与能够建造所需算力的设计建造公司合作。当然,我们可能也会自己建造一部分。我们始终会采用组合策略,因为在我们所需的规模下,我们需要利用所有算力来源。我们不能只依赖某一种机制。
It's a portfolio approach. We are never going to be in a world where we outsource everything or we let everything be outsourced. It's always going to be a mix because that's the reasonable thing to do. You don't want to put all your eggs in one basket. So, we will have hyperscalers providing a big chunk of compute, a majority of compute. We will have new clouds as parts of our portfolio. We will be partnering with design-build firms that can build the compute that we need. And of course, we may build some of it ourselves. We're always going to have a portfolio approach because at the scale which we need, we will need to tap into all sources of compute. We can't just rely on one particular mechanism.
在此之前你的背景:你既是斯坦福大学的教授,又是企业家或创始人,或者主要是学者。请带我们回顾一下你的经历。
Your background before all of this: you're both a professor at Stanford and an entrepreneur or founder, or mostly an academic. Just walk us through your journey.
以上都有一点。本质上,我是一名学者。我从 2010 年起就是斯坦福大学的教授。
A bit of all of the above. At my heart, I'm an academic. I've been a professor at Stanford since 2010.
你在那里专注于什么?
What do you focus on there?
我是计算机科学和电气工程系的教授。
I was a faculty in computer science and electrical engineering.
有特定的研究方向吗?
With a particular interest?
我的研究领域是网络和光纤网络,包括移动无线网络等。所以,实际上就是网络。
My area of research was networking and optical fiber networks, both mobile wireless networks and the like. So, only networks actually.
好的。
Okay.
三四年前,我在斯坦福期间创办了几家初创公司。最后一家被 VMware 收购,我因此认识了 Pat。Pat 后来成为英特尔 CEO,我也因此加入了英特尔。在加入 OpenAI 之前,我担任英特尔 CTO。所以我经历了各种不同的事情:学术界、初创公司、英特尔的企业环境。当然,在 OpenAI 则是以上所有混合体,因为他们同时拥有研究实验室、初创公司和快速成长的公司。
Three or four years ago, while I was at Stanford, I did a couple of startups. The last startup got acquired by VMware, and that's how I got to know Pat. Pat became Intel's CEO, and that's how I ended up at Intel. Most recently before coming to OpenAI, I was Intel's CTO. So I've seen all the different things: academia, startups, corporate at Intel. And then, of course, a mix of all of the above at OpenAI because they have a research lab, a startup, and a fast-growing company all mixed into one.
这就是你接受这份工作的原因吗?
Is that why you said yes to the job when it came up?
实际上就是我刚才说的。这使这份工作非常独特。其他地方很难找到。你总是需要做出选择,但这里既有世界级的研究环境,又有最困难的技术问题。我的意思是,我们正在建造世界上最大的计算机。有很多新问题需要解决。同时还有快速发展的业务。所以,我喜欢处于商业、技术和战略交叉点的问题。这是一个非常独特的历史时刻和独特的角色,非常有吸引力。
It's actually what I just said. That makes this so unique. It's hard to find anywhere else. You always have to choose, but having a world-class research environment coupled with the hardest technical problems. I mean, we're building the largest computer in the world. There are a lot of new problems that we need to solve. But also a fast-growing business. So, I like problems that sit at the intersection of business, technology, and strategy. This is a very unique time in history and a unique role, which is very attractive.
非常酷。好的,那么我们来更具体地谈谈 OpenAI 的算力策略。也许先总结一下你们目前拥有的。我想有一些来自微软,你们与 Cerebras 达成了 200 亿美元的大交易。还有很多事情在进行。还有 Stargate。也许先给我们介绍一下现状,然后我们再谈谈下一步计划。
Very cool. All right, so going into a bit more specifics about OpenAI's compute strategy. Maybe let's summarize what you guys currently have. I think there's some Microsoft, you did this big $20 deal with Cerebras. There's a bunch of things going on. There's Stargate. Maybe just give us the lay of the land of what you currently have and then we'll talk about what you're building next.
我们实际上从多个来源获得算力。微软显然是一个重要的合作伙伴。我们还从 AWS、ION 和 Google 获得算力。所以,我们实际上拥有所有超大规模云服务商的算力。
We have compute from effectively many sources. Microsoft is obviously a big partner, an important partner. We also have compute from AWS, ION, and Google. So, we have compute from all of the hyperscalers effectively.
我们也有来自 CoreWeave 的算力,比如一家新的云服务商。当然,还有芯片合作伙伴现在直接提供的算力,比如 Cerebras,为我们提供任何 26 算力。所以,我认为这大致就是今天的组合。展望未来,我们显然会继续发展所有这些合作关系。但同时,也在考虑更多选择,比如我们自己设计算力、数据中心,甚至可能自己建造数据中心。
We also have compute from CoreWeave, for example, a new cloud. And then, of course, compute that chip partners are supplying now, like Cerebras, directly for any 26 compute for us. So, I think that's the mix roughly today. As we go forward, obviously, we'll be building on all of these relationships. But also, looking at more options where we design the compute, the data center itself ourselves, or potentially even build a data center ourselves.
是的。
Yeah.
所以,这些都是扩展我们算力规模的方式。
So, all of those are ways of scaling the amount of compute that we have.
嗯。
Mhm.
所以,回到我之前的回答,答案可能总是尝试,但以上所有方式都会采用。我同意你的看法。
So, I think coming back to my earlier answer, it's the answer always will probably be try, but it's all of the above. I'm with you.
是的。不,这很有帮助。显然,考虑到稀缺性,多元化完全合理。然后,Stargate 似乎已经演变。从原本打算与 Oracle 和 SoftBank 成立的合资企业,到现在似乎更像是算力战略的总称。这样描述公平吗?
Yeah. No, that's helpful. And obviously, diversification makes all the sense in the world, given the scarcity. And then, it seems that Stargate has evolved. From what was going to be a joint venture with Oracle and SoftBank, to what now seems like it's more like an umbrella term for the compute strategy these days. Is that a fair way to describe it?
是的。我认为我们将 Stargate 视为我们的算力战略。它不同程度地涉及我们自己设计或建造算力。例如,与 Oracle 的紧密合作中,我们帮助他们设计,帮助他们了解如何运营 AI 算力,这对我们来说是一种新型算力。我们与公开上市的 SoftBank Energy 合作。我们基本上与他们共同设计了外壳,他们正在执行这些外壳,我们将自行研究如何在这些数据中心中使用我们的新芯片来运营。所以,对我们来说,Stargate 是涵盖所有这些不同方面的总括战略。把它看作一个持续演进的进程,因为我们永远不会明天醒来只做一种我们正在建造的东西。
Yeah. I think we look at Stargate as our compute strategy. And it is varying degrees of us designing or building the compute ourselves. For example, with Oracle close partnership, we help them design, we help them with how we do operate AI compute, which is effectively new kind of compute for us. We work with SoftBank Energy, which is public. We basically have co-designed the warm shell with them and they're executing on that warm shells and we will be kind of figuring out how to operate our chips in these data centers, ourselves, using our new chips. So, Stargate to us is that umbrella strategy for across all of these different things. And think of it as an evolution that we continuously be on because it's never going to be tomorrow we wake up and do only one kind of we are building it.
是的。
Yeah.
对我来说,Stargate 是一种持续学习如何扩展算力的方式,我们不断增加更多能力。
I've been Stargate to us is a continuous way of learning how to scale compute and we adding on more and more capability.
作为其中的一部分,现在正在德克萨斯州阿比林等地建造数据中心。那么,请带我们了解一下。目前正在建造什么,让人们有一些情境意识?
And as part of that, there are data centers being built right now in like Abilene, Texas. So maybe walk us through that. What is currently being built for people to have some situational awareness?
我们显然与 Oracle 有重大合作。那就是阿比林数据中心。例如,我们正在那里训练我们最新的模型。所以,对此非常兴奋。那是一个非常大的 GB 黑带集群,满足我们的需求。
We obviously have a big partnership with Oracle. That's the Abilene data center. That's where, for example, we are training our newest models. So, very excited about that. That's a very big GB black belt cluster for our needs.
是的。所以,它已经投入运行了。
Yeah. So, that's up and running.
它已经投入运行。它被用于训练最近的两个模型,所以非常棒。
It's up and running. It's being used for training the last two models more at your Um, so it's super trip.
是的。
Yeah.
很多,而且你看到了结果。比如你看到模型能力提升得有多快。
A lot and you're seeing the results. Like you're seeing how quickly the models are becoming more capable.
是的。
Yeah.
正是因为这类算力。
It's because of these kinds of compute.
是的。目前是否有正在建造的数据中心,是的,那些合适的,是的。
Yeah. Are there data centers that are currently being built that are Yes. That are the right Yeah.
是的。所以,Oracle 正在建造多个数据中心,所有这些都是公开的,分布在密歇根、德克萨斯等地。这些将在未来几年内上线,随着它们建成并配备当时的芯片,用于训练,但这些实际上旨在成为非常大的集群,既用于训练,也用于产品推理类竞争。
Yes. So, Oracle is building a number of data centers all of which are public so across Michigan and Texas and other places. So, these are coming online in the next couple of years as they get built and being put whatever chips on that time training rather than but these are really meant to be very big clusters that Amel is new both training, but also product inference kind of compete.
对。按照这些交易的结构,Oracle 是主要建造方,
Right. With the way those deals are structured, Oracle is the prime building those and
Oracle 是云服务商。
Oracle of the cloud.
是吗?
Is it
而我们是消费方公司
And we're consuming company
所以你是双方的核心租户,有趣。对于你们正在建造的东西,融资策略如何运作?你知道,你们最近筹集了 1220 亿,是这个数字吗?所以现金并不短缺,尽管考虑到所有这些支出,我不确定,但全球核心资源以战略性地使用债务而闻名。这也是融资策略的一部分吗?你们怎么考虑?
so you're the core tenant from both the Interesting. How does the for the stuff that you're building, how does the financing strategy work? You know, obviously you raised what was it? The 122 billion was it the number recently, so there's no shortage of cash, although, given all those expenses, I don't know, but the core resource of the world are famous for being very strategic users of debts. Is that part of the financing strategy as well? How do you think about it?
我的意思是,对于我们现在拥有的所有这些算力,我们有出色的合作伙伴实际上在为我们提供渠道。所以,微软、谷歌、亚马逊、Oracle 都是我们的承购方。所以,正如你所说,我们是租户。我们承诺消耗这些算力,购买这些算力,无论是否在线。
I mean, with all of these compute that we have today, we have amazing partners who actually are channeling that for us. So, Microsoft, Google, Amazon, Oracle are we are off the off-take. So, we are the tenants, as you put it. So, we commit to consuming that compute, to buying that compute, whether it's online
而且,全面来看,你们也在建造东西。
Also, across the board, like everything you're building stuff as well.
所以,你们在建造,合作伙伴也在建造
So, you're building you the partners are building
合作伙伴是某些人。
Partners somebody.
而且你们始终是全面的租户,而不是所有者
And you're always the across the board the the tenants, not the owner at the
正确。
Correct.
好的。所以,因此融资外包给了你们的合作伙伴。
Okay. So, therefore, financing is outsourced to your partners.
好的。我们谈到了 jalapeno,让我们更详细地讨论一下,因为特别地,你们似乎进展得非常快。我读到从设计到流片只用了 9 个月?所以,请带我们了解一下,为什么这么快?
Okay. We talked about the jalapeno, let's go into a bit more detail there because in particular, it seems that you guys went incredibly quickly. And is that anything that I read somewhere in 9 months from design to tape out? So, maybe walk us through that and what was the reason it went so quick?
是的,非常快。9 个月非常非常快,可能是我职业生涯中见过最快的。我认为有几个原因。一是团队强大。团队中许多人过去在谷歌设计过 TPU 芯片。所以,经验非常丰富的团队。我们有一个出色的合作伙伴博通。他们在交付定制 ASIC 方面有非常强的记录。所以,我认为与博通的强大合作促成了这一点。第三,我认为也许是 OpenAI 独特的资金。在大多数芯片公司或设计芯片时,你不知道你在为什么设计。因为你是一个供应商,最终运行工作负载的客户是不同的人。这里有一个独特的优势,我们知道未来模型可能是什么样子,因此能够缩短许多你需要做出的设计决策。所以,这非常有帮助。最后,AI 本身越来越多地帮助设计和优化芯片。这通常是耗时最长的部分,因为你基本上受限于人类处理所有数据和运行实验的时间,而我们可以更快地完成许多迭代。
Yes, it was incredibly quick. 9 months is very very fast, probably the fastest I've seen in my career. I think several reasons. One is it's a strong team. They have many of the team have designed TPU chips at Google in the past. So, very well experienced team. We have a great partner in Broadcom. That have a very strong track record of delivering get skills ASICs. So, I think a strong partnership with Broadcom made this happen. Three, I think perhaps OpenAI unique fund. In most chip companies or when you design chips, you don't know what you're designing it for. Because you are a vendor, the customer who eventually runs the workload is someone different. There's a unique advantage here of us knowing what the future models might look like and therefore being able to short circuit a lot of the decisions you need to make design decisions you need to make on the chip side. So, that's super helpful. And finally, increasingly AI itself helping design and optimize the chip. That is usually the one that takes the longest time because you're basically limited by how much human time there is to process all this data and run the experiments and we can do a lot of those iterations much faster.
用 AI?
With AI?
使用 AI。
Using AI.
是的。所以,AI 现在正在建造自己的芯片。
Yeah. So, AI is building its own chips now.
是的。我认为那个世界并不遥远。我的意思是,AI 现在正在辅助芯片设计。但我们确实相信递归的世界并不遥远,AI 将设计它需要用来训练和运行下一代 AI 的系统。
Yes. I think that world is not very far. I mean, AI right now is assisting in chip design. But we do believe that the world of recursion is not that far where AI will design the systems it needs to train and run the next generation of AI.
包括芯片。
And including chips.
包括芯片。你几周前还发布了 MRC,这是一个网络协议。请给我们讲解一下。它是什么?为什么它很重要?
Including chips. You also released a few weeks ago MRC, which is a networking protocol. Walk us through that. What is it and why is it a big deal?
这是一种新的网络协议路由技术,用于扩展这些超大型集群网络。想象一下,你有 10 万个 GPU,它们需要连接在一起。在进行大规模训练时,它们会不断相互通信,因为这些模型非常大,处理过程遍布整个 10 万 GPU 集群。你可以想象连接所有这些芯片所需的链路、交换机和网卡数量。在这种规模下,故障很常见,随时可能发生。你甚至无法列举所有可能的故障方式。因此,MRC 背后的策略是如何设计算法和协议,优雅地掩盖所有这些故障,确保训练工作负载不受影响。网络是一个抽象系统,训练任务从不需要担心它。它始终存在,即使链路故障也能找到路径。所以,这完全关乎可靠性和可用性。我们如何设计协议来驯服如此庞大集群的复杂性,确保不会因常见故障而停滞?MRC 是一种多路径喷洒协议,你可以将数据包喷洒到多条路径上。任意两个芯片之间有很多条路径,就像城市中任意两点之间有很多条路一样。我们不只选择一条路径,而是将流量发送到所有路径上,哪条成功就用哪条。这样,即使其中一条失败,也不会造成阻碍。这就是基本思路。显然,它要复杂得多,但在这种规模和速度下实现它很难,这也是它极具创新性的原因。
It's a new networking protocol routing technology, if you will, to scale these really large cluster fabrics. So, imagine you have 100,000 GPUs. They need to be connected together. And when you're doing large training runs, they're constantly communicating with each other because these models are so large that the processing is happening over the entire 100,000 GPU cluster. You can imagine the number of links, switches, and NIC cards needed to connect all these chips. At this scale, failures are common. It happens all the time. You can't even enumerate all the ways things could fail. So, the strategy behind MRC is how to design algorithms and protocols that can gracefully mask all these failures and ensure the training workload is not impacted. The network is an abstracted system that the training job never worries about. It's always going to be there, always going to find a path even if a link fails. So, it's all about reliability and availability. How do we design protocols that can tame the complexity of such a big cluster and ensure we don't get stopped because of failures, which are very common? MRC is a multi-path spraying protocol where you can spray packets over multiple paths. Between any two chips, there are many routes, like between any two points in a city. Instead of picking just one route, we send traffic across all of them, and whichever one succeeds, you take that. That way, even if any one fails, it's not a showstopper. That's the basic intuition. It's obviously a lot more sophisticated, but doing this at scale and speed is hard, which is why it's quite innovative.
你最近遇到了哪些瓶颈?计算机行业中瓶颈的性质一直在变化。现在人们经常谈论内存。这是其中之一吗?还是其他什么?
What are the bottlenecks you experience these days? The nature of the bottleneck keeps changing in the computer industry. People talk a lot about memory these days. Is that one of them? Or what else?
老实说,我认为瓶颈无处不在,从供应链开始。所以,我不认为只有一个瓶颈。我们在数据中心建设本身就有瓶颈,涉及许可、燃气轮机和变压器的可用性。这些行业在过去十年左右的时间里并没有增加太多产能,却突然经历了需求冲击。
I think there are bottlenecks everywhere, to be honest, from the supply chain. So, I don't think there's any one. We have bottlenecks in the data center building itself, around permitting, availability of gas turbines, transformers. Those industries historically have not added much capacity over the last decade or so, and they've suddenly experienced a demand shock.
是的。
Yeah.
而且需要数年时间才能具备生产更多涡轮机和变压器的能力。所以,我们正在努力追赶。
And it takes years before you have capacity to produce more turbines and transformers. So, we're trying to play catch-up.
还有你之前提到的工作问题,电工和懂这些技术的技工是否也存在短缺?是否存在人力瓶颈?
And to the job thing you mentioned earlier, is there a shortage of electricians and technical trade people who know how to build those things as well? Is there a human bottleneck?
绝对是的。电工、水管工等各种技工都短缺,你能想到的都有。我们所能做的任何培训更多人来从事这些工作的努力——这些都是高薪工作,所有超大规模企业和实验室都会积极招聘,只要你有资质。所以,我认为这绝对是一个瓶颈。它正变得越来越严重,因为我们都在试图用无限的需求建造更多东西。
Absolutely. There is definitely a shortage of electricians, plumbers, all kinds of trades, you name it. Anything we can do to train more folks to be able to do those jobs—they are very well-paying jobs that all the hyperscalers and labs would actively hire for if you had the qualifications. So, I think that is definitely a bottleneck. It's becoming an increasingly bigger bottleneck because we're all trying to build more and more with unlimited demand.
我们来谈谈商业方面。你最近推出的另一项业务是向客户保证容量以锁定算力。这很有趣。感觉 OpenAI 也正在成为一家为他人提供算力的公用事业公司。这背后的故事和策略是什么?
Let's talk for a minute about the business side. Another thing you launched recently is guaranteeing capacity for customers to lock in compute. That's interesting. It almost feels like OpenAI is also becoming a utility company providing computing to others. What's the story behind that and the strategy?
是的,保证容量就是保证 token。我们实际上是在说,我们会保证你一定金额的智能 token。这很合理。在算力短缺的世界里,token 总是很珍贵。即使我们拥有的算力有限,能生产的 token 也是短缺的。随着这成为企业的基础投入,企业将需要越来越多的智能来运营。这是企业确保所需智能 token 可用的一种方式,这样它们就不必承担业务风险。我认为这是良好的商业惯例。如果你有一种关键的供应资源——而智能是每个企业最重要的供应品——确保这种供应是合理的。这就是我们试图满足的需求。我知道这是一个新概念,为智能提供保证容量意味着什么,但我认为智能正在变成这样:每个数字企业的供应单位。
Yeah, guaranteed capacity is guaranteed tokens. We are effectively saying we will guarantee you a certain dollar amount of tokens of intelligence. It makes sense. In a world where compute is in shortage, tokens are always going to be at a premium. There's a shortage of tokens we can produce even with the limited compute we have. As this becomes a fundamental input to the enterprise, enterprises are going to need intelligence, more and more of it to run. This is a way for enterprises to gain assurance that the tokens of intelligence they need will be there, so they don't have to take business risk. I think it's good business hygiene. If you have a critical supply resource—and intelligence is the most important supply item for every enterprise—it makes sense to secure that supply. That's the demand we're trying to fulfill. I know it's a new concept, what it means to have guaranteed capacity for intelligence, but I think that's what intelligence is becoming: a supply unit for every digital enterprise.
也许以一个有趣的问题结束,你可以以 OpenAI 的身份回答,也可以不回答。太空中的数据中心。这令人兴奋吗?是科幻小说吗?有必要吗?还是人们只是因为酷才喜欢谈论它?你的立场是什么?
Maybe to end on a fun one, you can answer with your OpenAI hat on or not. Data centers in space. Is that exciting? Is that science fiction? Is it needed? Is it something people like to talk about just because it's cool? Where do you land?
对于我内心的极客和工程师来说,这绝对令人兴奋。看到一组卫星产生算力,那将是非常酷的事情。我实际上认为随着我们解决工程问题,它会变得可行。这些问题可以通过足够的时间和投资来解决。至于是否必要,我认为轨道算力是有空间的。我不认为它能解决所有算力需求,但它肯定会成为武器库中的补充。我们正在等待的是发射卫星的经济性何时改变,以及硬件的经济性何时改变。我们需要达到一个点,即发射硬件很便宜,如果出现故障,丢弃它也很便宜。你不能像在地面上那样上去修理它。我认为这个转折点有望很快到来,到那时它就变得可行了。
For the geek in me and the engineer in me, it's definitely exciting. It's one of those things that would be super cool to see a constellation of satellites producing compute. I actually think it will become feasible as we solve engineering problems. They can be solved with enough time and investment. Whether it is needed, I think there is room for orbital compute. I don't think it's going to solve all compute needs, but it will definitely be a complement in the arsenal. What we are waiting for is when the economics of launching satellites change and when the economics of the hardware change. We need to get to a point where it's cheap to launch the hardware, and if something fails, it's cheap to throw it away. You can't go up and fix it, unlike on the ground. I think that inflection point will hopefully happen soon, and at that point it becomes viable.
没错。Chirag,这真是太棒了。非常感谢你抽出时间与我们交流。我们很感激。
Correct. Chirag, it has been wonderful. Thank you so much for spending time with us. We appreciate it.
谢谢。
Thank you.
这次聊天非常愉快。
It's been great to have this chat.
酷。嗨,我是 Matt Turk。感谢收听本期 Mad Podcast。如果你喜欢这期节目,我们非常感激你考虑订阅(如果还没订阅的话),或者在观看或收听本期节目的平台上留下好评或评论。这真的能帮助我们发展播客并邀请到优秀的嘉宾。谢谢,下期再见。
Cool. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.