Cerebras CS4:让晶圆级计算成为主流

Cerebras CS4: Making Wafer-Scale Mainstream

肖恩·利 Sean Lie · Latent Space · 2026-09-02 · 约 44 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Cerebras CTO 讨论新一代 CS4 架构,其功率和互联带宽翻倍,在 GPT-OS 上实现每秒超过 4400 个 token,并谈及 Hot Chips 上硬件创新的黄金时代。

CTO of Cerebras discusses the new CS4 architecture, which doubles power and interconnect bandwidth, achieving over 4,400 tokens per second on GPT-OS, and the golden age of hardware innovation at Hot Chips.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

开场与赞助商信息 Opening and Sponsor Message

Host

在进入今天的节目之前,我有个小请求要对听众们说。谢谢你们。如果不是你们选择点击并收听我们的内容,我们就不可能为你们带来如此渴望的 AI 工程、科学和娱乐内容。几乎每天都有赞助商找上门,但幸运的是,有足够多的你们真正订阅了我们,让我们能在没有广告的情况下维持运营,我们也希望保持这样。但我只想请你们帮一个忙。你们能做的最有力、完全免费的一件事,就是点击订阅按钮。这是我唯一会请求你们做的事。这对我以及每周辛勤工作为大家带来 Inspace 的团队来说意义重大。如果你们订阅了,我保证我们会永不停止地努力让这个节目变得更好。现在,让我们开始吧。

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now, let's get into it.

Host

好的,我们现在在 Crisis 总部,和 CTO Sha Lee 在一起。欢迎。

Okay, we're here at Crisis HQ with CTO Sha Lee. Welcome.

Sean

谢谢。感谢你们的邀请。

Thank you. Thank you for having us.

Host

嗯,不,谢谢你邀请我。

Yeah. No, thank you for having me.

Host

今天是 Hot Chips 大会的第二天。有很多东西发布,甚至苹果也发布了 M6 系列。但你们显然在 Supernova 上发布了 CS4,我们稍后会聊聊 CS5。OpenI 也提到了 Jalapeno。Hot Chips 上全是热门话题。你在这一行干了这么久,今年有什么看法?

And it is the day after Hot Chips. Lots of things launching. Even like Apple launched the M6 stuff. But you guys obviously had we were at Supernova. You talked about CS4. We're going to talk a little bit about CS5. OpenI talked about Jalapeno. All the hot stuff in Hot Chips. What's your take this year having been in this industry?

Sean

我觉得,首先,谢谢你邀请我。你说得完全对。现在是一个非常激动人心的时刻。Hot Chips 是整个社区聚在一起的好时机。你知道,过去——我还记得 10 到 20 年前参加 Hot Chips 时,那还是一群计算机架构极客在讨论速度和参数。而现在,这里变成了发布改变行业硬件的场所。所以它已经走了很长的路。这是一个非常令人兴奋的社区。我昨天离开 Hot Chips 后最大的收获之一,就是芯片设计和硬件行业正处在一个多么激动人心的时代。现在真的是硬件的黄金时代。就像你说的,我在这个行业已经有一段时间了。这是一个非常独特的时期,不仅仅是因为 AI 正在起飞,显然有很多人在做自己的硬件,而且现在在多个领域发生的创新数量——我们在半导体行业历史上从未见过这种在各个层面上的创新,对吧?在芯片、互连、系统设计、软件、光学、方法论、工具链——人们都在全面推动。所以作为一个技术专家和骨子里的计算机架构极客,看到这个社区和行业处于现在这个阶段,真的非常有趣。我们在 Hot Chips 上也感受到了——我交谈的每个人,那种兴奋感都在。太棒了。

I think, well, first of all, thank you for having me. And you're absolutely right. It's a super exciting time right now. Hot Chips is a really good time when the whole community comes together. And you know, what used to be—I still remember going to Hot Chips 10, 20 years ago when it was a bunch of computer architect geeks geeking out on speeds and feeds. And now it's like this is where industry-changing hardware is being revealed. So it's come a long way. It's a really exciting community. And I think one of the main takeaways I had after leaving Hot Chips yesterday is what an exciting time it is for the chip design and the hardware industry. This is really the golden age of hardware right now. And like you said, I've been in this industry for some time. This is a very unique time, not just because AI is taking off and there are obviously a lot of people doing their own hardware, but the amount of innovation happening now across multiple fronts—we've never seen in the history of the semiconductor industry this level of innovation across all the levels, right? In the chip, the interconnect, the system design, the software, the optics, the methodology, the tooling—people are pushing across the board. So as a technologist and a computer architecture geek at heart, it's really fun to see this community and this industry at this stage right now. And we felt that at Hot Chips—everybody I talked to, that buzz is there. It's amazing.

CS4架构与性能 CS4 Architecture and Performance

Host

我们想逐一聊聊这些芯片。先说说你们的芯片吧。最令人兴奋的是什么?CS4 发布时带来了一些疯狂的数字——比 GPU 快 30 倍。但请给我们讲讲新的 CS4 有什么新东西。

We want to go through the chips. Let's talk about your chip first. What's most exciting? CS4 came out with some crazy numbers here—30x faster than GPUs. But walk us through what's new with the new CS4.

Sean

当然。所以我们设计了下一代 CS4 架构,搭配全新的系统平台,目标真正是让晶圆级芯片成为主流,将其带入超大规模数据中心,并解决我们每天面临的高密度数据中心问题。所以我们设计了这种模块化平台,为晶圆提供的功率是上一代的两倍,互连带宽也是两倍,延迟减半。所有这些都是在系统架构层面完成的。你在这里看到的就是结果:通过能够为晶圆提供显著更多的功率,并在机架中封装更高的密度,我们能够进一步提升性能。我们喜欢说,Cerebras 凭借我们现有的产品,已经是超快推理领域无可争议的领导者。而令人惊叹的是,CS4 产品将把这一前沿推得更远,速度再提升两倍。这是一个完美的例子——在 Hot Chips 上我们做的这个演示中,我们展示了 GPT-OSS 以超过每秒 4400 个 token 的速度运行,这简直令人难以置信。感觉就像假的一样,对吧?我们认为这将真正彻底改变整个行业,因为我们不仅能够继续推动这些模型运行速度的前沿,它还将开启各种新的应用。用户体验开始变得截然不同——过去是批处理和离线的应用开始变得实时。而所有新的智能体流程和智能体框架的增长也意味着,你坐在那里等待这些智能体在它们的智能体循环中一遍又一遍地运行。而突然间,如果你的模型以每秒超过 4000 个 token 的速度运行,现在你可以做更多的智能体循环,你可以做更多的推理,最终你会得到显著更强大、更智能的智能体。所以这才是我们真正兴奋的地方——不仅仅是我们展示了这些非常大的数字,这当然也很酷,而是你可以真正开始做那些否则无法做到的事情,并开启全新的能力。这就是让我对昨天的发布如此兴奋的原因。

Sure. So we designed our next-generation CS4 architecture with a brand new system platform, with the goal really to make wafer scale mainstream, to bring it to hyperscale, and to solve a lot of the high-density data center problems that we're facing every single day. So we've designed this modular platform that provides twice the amount of power to the wafer than we have in our previous generation, twice the amount of interconnect bandwidth, half the latency. And all of it is done at the system architecture level. And what you see here is the result: by being able to provide significantly more power to the wafers and packing more density into the rack, we're able to drive up the performance even further. We like to say that Cerebras, with our current product, is already the undisputed leader in ultra-fast inference. And what's amazing is this CS4 product will actually push that frontier even further by taking it another two times faster. And this is a perfect example—in this demo we gave at Hot Chips, we're showing GPT-OSS running at over 4,400 tokens per second, which is just mind-blowing. It almost feels like it's fake, right? And we think that this is going to really revolutionize the entire industry, because not only are we able to continue to push the frontier on how fast these models can run, it will enable all sorts of new applications. User experiences start to become extremely different—what used to be batch and offline applications start to become real-time. And all of this growth in new agentic flows and agentic frameworks also means that you're sitting there waiting for these agents to go over and over in their agentic loop. And all of a sudden, if you're running your model at over 4,000 tokens per second, now you can do more agentic loops, you can do more reasoning, and ultimately you get significantly more capable, more intelligent agents. And so this is what really excites us about this—not just the fact that we show these really big numbers, which is also really cool, but the fact that you can really start to do things that you can't really do otherwise, and start to enable brand new capabilities. That's what makes me so excited about this reveal we had yesterday.

晶圆级与AI热潮 Wafer-scale and the AI boom

Host

你知道,晶圆级的一切——你一直说的那些——人们之前就是不当回事,直到他们不得不认真对待,现在他们真的、真的超级重视了,对吧?

You know, wafer-scale everything—everything that you've always said—it's just like people didn't take it that seriously until they had to, and now they're really, really taking it super seriously, right?

Host

嗯,我确实觉得,作为联合创始人、CTO、整个项目的架构师——你们是怎么应对 AI 热潮之后的新一代产品的?我想象 CS1、2、3 的时候可能更平静一些。现在,OpenAI 真的在和你们一起推出超快模式,这几乎是——怎么说呢——你们几乎是在和芯片共同设计你们的模型。

Um, I do feel like, as you know, co-founder, CTO, the architect of this whole thing—how are you guys approaching the new generations post-AI boom? Like, I imagine CS1, 2, 3 was a bit more, uh, sort of calm. Now, like, literally OpenAI is launching the ultra-fast mode with you guys, and it is a matter of—I don't know—like you're co-designing your model with the chip almost.

Sean

不,绝对。我觉得过去几年对我们来说发生的主要转变是,第一代芯片主要是技术演示,对吧?我们必须向世界、向自己证明,你确实能造出这样的东西。我们告诉很多人这是未来,他们大多数人都觉得‘你们疯了,这不可能’。所以一开始,我们主要是搞清楚基础技术,让它成为可能。现在我们看到的是,既然我们已经证明了这东西能造出来,而且能达到这种架构所能达到的性能,那它能实现什么,对吧?你提到了我们最近和 OpenAI 的超快发布——这就是个完美的例子。我们在运行前沿级别的、最智能的模型之一——OpenAI 目前最大、最强、最智能的模型——速度比他们正常的 GPU 快 14 倍。当你把最智能的模型和硬件带来的速度结合起来,这就变成了一件彻底变革性的事情。从技术演示到解决实际问题、实现新的真正能力,这个转变正是我们公司正在经历的。我们现在应对的方式是,真正把重心转移到确保我们能规模化。我们能提升所需的产能。我们能运行用户最终关心的模型。然后我们继续这个飞轮,确保持续推出更快的硬件,不断推动前沿。这就是我们这边的转变。而且它的展开方式几乎比我们想象的还要好。这是硬件能力已经得到验证,现在又和模型及应用的能力完美契合。所以我们现在的目标就是尽可能快地规模化这个东西。

No, absolutely. I feel like the main shift that has happened over the last few years for us is that the first-generation chip primarily was a technology demonstration, right? We had to demonstrate to the world, to ourselves, that you can actually build such a thing. And we told a lot of people that this is the future, and most of them were just like, 'You guys are insane, this is not possible.' So at the beginning, it was really about figuring out the foundational technologies to make it possible. And what we're now seeing is that now that we have in fact demonstrated that this can be built and it can have the kind of performance that this type of architecture can have, what are all the things that it can enable, right? And you alluded to our recent ultra-fast launch with OpenAI—that's a perfect example. We're running frontier-level, one of the most intelligent models—OpenAI's largest, most capable, most intelligent model right now—at 14 times faster than their normal GPU speeds. And it's becoming this completely transformational thing when you can combine the most intelligent models with the speed that the hardware brings. And this transition from technology demonstration to solving real problems, enabling new real capabilities, is this transition that we as a company are going through. And how we're coping with that right now is we're really shifting our focus to ensuring that we can make this scale. We can bring up the capacity that we need. We can run the models that the users ultimately care about. And then we continue this flywheel by making sure that we continue to put out faster hardware and keep pushing the frontier. And that's kind of how things have shifted for us. And it's unfolding in a way that is almost better than we could have imagined. It's this perfect confluence of having the hardware with the capabilities demonstrated now matching with the capabilities of the models and the applications. And so our goal right now is just to scale this thing as fast as we can.

预览CS5 Previewing CS5

Host

那也许我们聊聊昨天的事。对。呃,预览 CS5。你能给那些可能还没跟上、也许是第一次听说这个故事的人回顾一下吗?

And so maybe we'll talk about yesterday. Yeah. Uh, previewing CS5. Can you recap for people who maybe are not yet caught up and, you know, maybe are new to the story for the first time as well?

Sean

好,没问题。所以我提到,我们刚推出的 CS4 系统基于一个全新的 RAS 平台,我们称之为 Nexus 平台。这个平台的特别之处在于模块化。前面有模块化电源,后面有模块化服务器,我们称之为背包。这个平台实现了我们刚刚看到的 2 倍性能。但它也是多代产品的基础。我们从零开始设计,以支持多代产品。特别是,我们是和明年即将推出的下一代晶圆芯片一起设计的。有了那个芯片,我们会把性能推得更高。所以我们今年在 CS4 上看到了 2 倍的提升。我们还会再推一个 2 倍的性能提升。这最终意味着,你将能运行中等规模的模型,比如 GPOSS 或 Gemma,速度高达每秒 10,000 个 token,甚至像 Kimmy、DeepSeek 和 GPT-5.6 Soul 这样的前沿模型,速度也能达到每秒 5,000 个 token。这完全是颠覆性的。这一切再次得益于我们为多代产品设计的这个全新平台。所以这就是我们现在的轨迹。既然世界看到了超快推理的价值,我们打赌这还只是开始。

Yeah, no problem. So I mentioned that our CS4 system that we just launched is based off a brand new RAS platform that we call the Nexus platform. And what's special about this platform is the modularity. It has a modular power supply in the front. It has a modular server in the back that we call the backpack. And this is a platform that enables that 2x performance that we just all saw the numbers for. But it's also the foundation for multiple generations of products. We've designed it from the ground up to support multiple generations of our products. And in particular, we designed it together with our next-generation wafer chip that will be coming next year. And with that chip, we'll be pushing the performance even further. So we just saw a 2x improvement this year with CS4. We're going to push it even further with another 2x improvement in performance. And what this ultimately means is you'll be able to run medium-sized models like GPOSS or Gemma at speeds up to 10,000 TPS, and even frontier-level models like Kimmy or DeepSeek and GPT-5.6 Soul up to 5,000 TPS. Completely game-changing. All again enabled by this brand new platform that we've designed for multiple generations. And so this is the trajectory that we're on right now. And now that the world sees the value of ultra-fast inference, our bet is that even this is just the beginning.

推出过程与OpenAI合作 Rollout process and OpenAI partnership

Host

推出过程是怎样的?你们和 OpenAI 有未来几年的协议。嗯,我们什么时候能用上,你知道吗?现在每个人都想要超快。当然。有 Gemma OSS。我们什么时候能用上大的 Kimi、DeepSeek?我们什么时候能真正开始用?

What's the rollout process like? So you have a deal with OpenAI over the next few years. Um, when do we get access, you know? So everyone wants ultra-fast right now. Sure. There's Gemma OSS. When do we get the big Kimies, the Deep Seeks? When can we start to actually use them?

Sean

这是个好问题。所以现在我们基本上把我们正在建造的一切都卖光了,而且我们非常有策略地确保每一兆瓦的部署都尽可能有战略意义。特别是,相当一部分给了 OpenAI。我们对此非常公开——他们是我们最大的合作伙伴,不仅因为商业安排,还因为我们有这种共同设计的精神,你们也经常听到 OpenAI 谈论这个,我们相信这基本上能让我们继续这个飞轮,不仅让模型更快,而且利用这种速度让它们更智能,等等。所以今天,很多产能都流向了 OpenAI,而在 OpenAI 内部,他们非常有策略地决定把相当一部分留给自己用。所以他们现在就在用。没错。所以现在他们内部把它用在很多速度至关重要的关键用例上。比如他们用在事件响应团队里,对吧?当他们的服务出现中断时,每一秒、每一分钟都很重要。所以他们从中获得了巨大的价值。

That's a great question. So right now we are basically sold out of everything that we're building, and we are very strategically making sure that we're deploying every single megawatt in the most strategic way possible. And in particular, quite a bit of it is going to OpenAI. And we've been very public about this—they're our biggest partner, not just because of the commercial arrangement, but also because of the fact that we have this co-design spirit, which you guys have heard OpenAI talk about a lot as well, which we believe will basically allow us to continue this flywheel of not just making models faster, but using that speed to make them more intelligent, and so on. So today, a lot of that capacity is going into OpenAI, and within OpenAI they have very strategically decided to use quite a bit of it for themselves. So they're using it right now. That's correct. So right now internally they're using it for a lot of really critical use cases where the speed really matters. Like they're using it in their incident response teams, right? When there's an outage in their service, for example, every single second, every single minute matters. And so they're getting a tremendous value out of that.

超快推理能力与企业接入 Ultra-fast inference capacity and enterprise access

Sean

他们正把它用在一些最关键的研究应用中,额外的推理能力带来了明显更智能的响应。所以目前相当一部分算力都用在了那里。然后,正如我们在与 OpenAI 的发布中所提到的,我们现在也向企业客户开放,他们能够将其用于一些要求最高、价值最高的应用。随着时间的推移,OpenAI 和 Cerebras 已承诺带来足够的算力,让超快推理惠及更广泛的受众。我们的 CS4 和 CS5 发布正是这一承诺的一部分,继续推动更多、更快的 token 和更高的吞吐量,以满足各种不同的用例。

They're using it in some of the most critical research applications where the extra reasoning is enabling significantly more intelligent responses. So right now, quite a bit of the capacity is going there. And then, as we mentioned in our launch with OpenAI, we are now also making it open to enterprise customers who are able to use it for some of the most demanding and high-value applications. Over time, OpenAI and Cerebras have committed to bringing enough capacity to make ultra-fast inference available to a much wider audience. Our CS4 and CS5 announcements are very much part of that commitment to continue to drive more, faster tokens and more throughput to satisfy all the various use cases out there.

Host

顺便说一句,听你这么一说,我意识到他们其实不必向企业客户开放,完全可以留着自己用,没人会知道。这几乎是个有趣的商业决策。我不知道你对此有何看法,这更像是商业分析师的观点。

By the way, the way you framed it, I realized that they didn't have to expose it to enterprise customers. They could have just kept it for themselves. Nobody knew that. That's an interesting business decision almost. I don't know if you have any weighing on that. That's more like a business analyst point of view.

Sean

我想说的是,显然,我无法明确评论他们的想法。最终,如何使用这取决于 OpenAI 的决定。但我认为从外部来看,这有一定道理。我的意思是,他们的使命很大程度上是继续推动可能性,而内部使用有助于他们做到这一点。但他们现在也是一家企业,对吧?有关于他们 IPO 的讨论等等。即使从外部看,你也能看到他们在平衡这两方面。我认为我们在超快推理领域也清楚地看到了这一点。我的意思是,他们在平衡很多事情,最近他们的热门……

I would say that, obviously, I can't explicitly comment about their thinking. Ultimately, it's OpenAI's decision how they want to use this. But I think if you look at it from the outside, it kind of makes sense. I mean, their mission is very much to continue to push what's possible, and by using it internally, that helps them do that. But they're also a business now, right? There are talks about them going IPO and all this. Even externally, you can see they're balancing both of these. And I think we very much see that playing out in the ultra-fast space as well. I mean, they're balancing a lot, most recently their hot...

Host

我们得谈谈这个,得深入聊聊。那么你怎么看?预填充、Jalapeño、解码、Cerebras,你有什么看法?

We got to talk about this as well. We got to go into them. So what are we thinking? Prefill, Jalapeño, decode, Cerebras. What are your takes?

Sean

我认为这是非常理性的结论。在所有热门芯片发布中,Jalapeño 可能是最让我兴奋的,但也许原因和其他人不一样。

I think that's a very rational conclusion. I think of all the hot chip announcements, probably Jalapeño was the most exciting to me, but maybe not for the same reason as everyone else.

Host

我想深入探讨一下,你是芯片专家,对吧?有人关注每秒 token 数,你觉得什么有趣?

I guess to double-click on that, you're the expert in chips, right? There are people who see tokens per second. What do you see as interesting?

Sean

我认为他们在性能上下了很大功夫,而且他们明显优于 Nvidia 的 GPU。但我看到的是,他们打造了一款明显更好的 GPU,这本身就是一项巨大成就。Nvidia 知道自己在做什么,对吧?他们主导市场是有原因的,他们不是傻瓜。所以能够一鸣惊人,打造出明显更好的 GPU 是巨大成就。但对我来说,Jalapeño 如此令人兴奋的原因甚至不是这些参数,而是其背后的设计方法论。他们显然采用了截然不同的方法来构建这款芯片。AI 优先的方法论使他们能够更快地构建芯片,并取得这些令人瞩目的成果。这 100% 是我们行业的未来。看到 OpenAI 在这方面引领潮流并不奇怪,因为这正是他们的行事风格。但回到你刚才说的我们如何使用它,我认为 Jalapeño 正在推动吞吐量和延迟两方面的可能性边界。对我来说,即使我摘下 Cerebras 的帽子,我也认为这对行业来说是件好事。这将抬高所有船只,每个人都会受益。他们还成功将延迟推入传统 GPU 无法企及的范围。这也很好,因为我们相信速度,那里有很多价值。真正有趣且让我超级兴奋的是——OpenAI 是我们最大的客户,所以我认为这对我们具有重要战略意义——当明年 Jalapeño 上市、我们的下一代 CS5 上市时,我们将共同提供一个完整的快速推理产品组合,它比今天已有的产品有本质上的不同和更好,而今天的产品已经与基准 GPU 有本质不同。在此基础上,还有机会进一步集成,比如预填充和解码分离。

I think that they pushed a lot on performance, and the fact that they're significantly better than Nvidia's GPU. But what I see is that they've built a significantly better GPU, and that in its own right is a big achievement. Nvidia knows what they're doing, right? They own the market for a reason. They're not dopes. So to be able to come out of the gate and build a significantly better GPU is a big achievement. But to me, the reason why Jalapeño is so exciting isn't even all these parameters. It's really the design methodology behind it. They very clearly took a drastically different approach to building this chip. Having an AI-first methodology enabled them to build the chip faster and achieve some of these very impressive results. And that is 100% the future of our industry. It's not surprising to see that OpenAI is kind of leading the way here because this is very much their M.O. But coming back to what you said about how we're going to use this, I think that Jalapeño is pushing the boundaries of what's possible for both throughput and latency. And for me, even if I take the Cerebras hat off, I think that's awesome for the industry. This is going to lift all the tides. Everyone is going to benefit. They've also now been able to push the latency into regimes that traditional GPUs can't hit. And that's also great because we believe in speed. There's a lot of value there. And what's going to be really interesting, and what I'm super excited about—and OpenAI is our biggest customer, so this is one of the things that I think is very strategically important to us—is that when Jalapeño is available next year, when our next-generation CS5 is available, together we will enable a full fast inference portfolio that is substantially different and better than what's already available today, which is already substantially different than your baseline GPU. And on top of that, there's opportunity to integrate even further, like pre-fill and decode disaggregation, for example.

Host

他们实际上并没有专门为这个优化 Jalapeño,对吧?

They didn't actually specialize Jalapeño for that, right?

Sean

他们没有专门为那个优化 Jalapeño,但他们为吞吐量做了专门优化。所以他们获得了巨大的吞吐量。因此,作为一名计算机架构师,你可以开始用它做很多不同的事情。然后我们还有惊人的延迟,就像我提到的,CS5 中高达 10,000 TPS。你可以开始想象我们可以一起构建的一些非常酷的产品。这就是为什么有他们作为合作伙伴让我如此兴奋。我们还对 AI 工具方面的合作感到兴奋,因为他们从 AI 工具基础设施中获得的许多好处——现在可能每个人都会这么说,但我们显然在 AI 上做了很多工作,同时我们也在与 OpenAI 紧密合作,使用他们的工具帮助我们继续推动芯片设计和软件等方面的可能性。所以这两者结合在一起,我觉得这是一个真正无敌的组合。

They did not specialize Jalapeño for that, but they specialized it for throughput. So they get a tremendous amount of throughput. And so just as a computer architect, there are so many different things you can start to do with that. And then we have the insane latency, like I mentioned, up to 10,000 TPS in CS5. You can start to imagine some really cool products that we can build together. And that's what really excites me about having them as a partner. And we're also excited about collaborations on the AI tooling front, because much of the benefits that they're seeing from the AI tooling infrastructure—everybody probably says this now, but we're obviously doing a lot with AI, but we're also collaborating very closely with OpenAI to use their tools to help us continue to push what's possible in our chip design and our software and all that. So both of those together, I feel like this is a really unbeatable combination.

争议话题:能效与Grok芯片 Contentious topics: power efficiency and Grok chip

Host

我觉得人们在谈论的一件事——我现在想触及热门观点中的分歧。所以关注每秒 token 数的人并没有真正提到功耗。我确实认为一个似乎一致的主题是每瓦性能,而不是每秒 token 数。这个主题有什么变体,或者人们在台下谈论什么更具争议性的话题?

I think one thing that people are talking about—I'm trying to get to the disagreements of the hot takes now. So people focusing on tokens per second haven't really mentioned power. And I do think that something that seems to be a consistent theme is performance per watt rather than tokens per second. Any variation of this theme, or what are people sort of talking about offstage that is more contentious?

Sean

嗯,我觉得有几件事。在一些更具争议性的话题中,我在会议上讨论中经常出现的一个主题是关于 Grok 的发布。你知道,Grok,显然,他们有新芯片并不令人意外。

Well, I think there are a few things. In terms of some of the more contentious things, one of the themes that came out quite a bit in my discussions at the conference was around the Grok announcement. You know, Grok, obviously, not surprising they have a new chip.

SRAMM设计成为主流 SRAMM designs becoming mainstream

Sean

总的来说,我认为 SRAMM 设计现在变得越来越主流,这很棒。我们已经谈论这个很久了,看到行业开始接受它,真是令人惊叹。

In general, I think it's awesome that SRAMM designs are becoming more mainstream now. We've been talking about it for a long time, and it's amazing to see the industry starting to embrace it.

Host

收购后他们感觉有什么不同吗?我的意思是,你已经和他们竞争了一段时间了。

Do they feel different post acquisition? I mean, you've been competing with them for a while.

Sean

首先,没有。我认为看到 SRAMM 架构变得更容易获得,真的很棒。即使是这里最大的人物,英伟达,也在拥抱 SRAMM 设计,承认传统的 GPU 设计确实无法达到超快速度的领域。

To first order, no. I think it's really awesome to see that SRAMM architectures are becoming more accessible. Even the biggest of the big guys here, Nvidia, is embracing SRAMM design, acknowledging that traditional GPU designs really can't hit the ultra-fast regimes.

30B模型发布疑云 Suspicious launch on 30B model

Host

台下讨论的很多都是自然明显的问题:他们为什么在 300 亿参数的模型上发布?还有,当 Jensen 在 GTC 上花了很多时间谈论注意力 FFN 分解时,为什么没有提到这一点?

A lot of what was being discussed offstage was the natural obvious questions: why did they launch on a 30 billion parameter model? And how come when Jensen spent so much time at GTC talking about attention FFN disaggregation, there was no mention of that?

Sean

我认为这很能说明问题。我的意思是,有一个单独的 Reuben 策略。

I think that's pretty telling. I mean, there's a separate Reuben strategy.

Host

嗯,所以有 Reuben,但 LPX 本身应该是它的所在。它应该是 Reuben LPX 一起的,对吧?

Well, so there's Reuben, but the LPX itself was supposed to be where it was. It was supposed to be Reuben LPX together, right?

Sean

如果你还记得,Jensen 在 GTC 上花了大约半小时解释注意力在这里运行,电影在这里运行,等等。我不知道这是不是个热门观点,但非常可疑的是,他们的产品已全面投产,却只展示了非分解的 310 亿参数模型的性能数据。

If you recall, Jensen spent like half an hour at GTC explaining attention runs here and the movies run here, and so on. I don't know if this is a hot take, but it's very suspicious that their product that's in full production, they've only shown performance numbers on a non-disaggregated 31 billion parameter model.

Sean

对我来说,这表明在非晶圆级 SRAMM 设计上运行肯定存在一些挑战,因为每个芯片中的内存不够。我认为这就是正在发生的事情,我们看到了证据。在很多方面,这非常验证了我们所做的设计选择。如果你想想,要运行一个前沿级别的模型,比如几万亿参数,你需要成千上万个 Grock LPU 才能容纳权重。

To me, this shows that there are definitely some challenges in running on a non-wafer-scale SRAMM design because there's just not enough memory in each of the chips. And I think that's what's happening, and we're seeing the evidence of that. In many ways, it's very much validating the design choice that we had. If you think about it, to run a frontier-level model of, let's say, a few trillion parameters, you need thousands and thousands of Grock LPUs just to hold the weights.

Host

是的。

Yeah.

Sean

所以当你开始这样想的时候,他们展示的唯一性能数据是在 30B 上,这令人惊讶吗?

So when you start to think about it that way, it's like, is it surprising that the only performance numbers they're showing are on 30B?

Host

所以他们要把这个渐变到这个。

So they're going to gradient this into this.

Sean

我认为是相反的。最终会发生的是,他们将专注于明显更小的模型。

I think it's the other way. What's going to end up happening is they're going to end up focusing on significantly smaller models.

Host

好的。

All right.

Sean

如果你的架构有那个限制,那么最终就会发生这种情况。而在我们的案例中,我们运行的是世界上最大的模型,在某种程度上这很简单,因为我们的一颗芯片的内存比他们的一颗芯片多大约 100 倍。所以你免费获得了两个数量级的规模差异。所以我认为这绝对是讨论的话题之一,再次独立于 Cerebrus。在一个如此小的模型的性能数据上推出一个全新产品,这有点奇怪。但当你稍微深入一点,这就有道理了。我长期在这个领域工作。我们需要晶圆级集成来聚合足够的 SRAMM 以使其对大型模型有用,这是有原因的。

If you have that limitation in your architecture, then that's what ends up happening. Whereas in our case, we're running the world's largest models, and in some ways it's kind of simple because one of our chips has order 100 times more memory than one of their chips. So you have two orders of magnitude difference in scale kind of for free. So I think that's definitely one of the things that was a topic of discussion, again independent of Cerebrus. It's just kind of odd that you'd launch a brand new product on performance numbers of such a small model. But when you peel it back a little bit, it kind of makes sense. I've been living in this space now for a long time. There's a reason why we needed the wafer-scale integration to be able to aggregate enough SRAMM to make it useful for large models.

市场规模与异构生态 Market size and heterogeneous ecosystem

Host

是的。看,市场很大,对吧?你和他们有不同的市场。你知道你显然是这个领域里经营最久的现有者。

Yeah. And look, the market's large, right? You have a different market than them. And you know that you're clearly the longest running incumbent now in this space.

Sean

不,绝对。市场很大;有很多不同的机会让不同的硬件扮演不同的角色。事实上,在 Cerebrus,我们非常坚信异构、分解的生态系统。不仅仅是预填充-解码分解,而且我认为我们在这个领域才刚刚开始探索可能性。随着这些部署和模型的规模,推理不再只是一个工作负载。它里面有很多子工作负载。你真的想为问题使用正确的工具,为问题使用正确的硬件。所以绝对,所有不同类型的架构都有位置。我们选择了瞄准前沿。

No, absolutely. The market's large; there are a lot of different opportunities for different hardware to play different roles. In fact, at Cerebrus, we believe very strongly in a heterogeneous, disaggregated ecosystem. Not just prefill-decode disaggregation, but I think we're just at the beginning of what's possible in this space. With the scale of these deployments and these models, inference is no longer just one workload. There are many sub-workloads within it. You really want to use the right tool for the problem, the right hardware for the problem. So absolutely, there's a spot for all the different types of architectures out there. We've chosen to target the frontier.

Host

是的。前沿超快。

Yeah. Frontier ultra fast.

Sean

正是。

Exactly.

工作负载类别与客户视角 Classes of workloads and customer perspective

Host

我本来要谈谈今年其他一些顶尖公司,但这给了我机会。我想跟进一件事:你如何看待工作负载的类别?对我来说,超快前沿工作负载显然只是整个推理市场中增长最快的部分之一。我想要它,但我无法为它支付足够多的钱——OpenAI 提供它甚至可能是个错误,但对我们来说无论如何都是好事。所以,根据你的经验,你说你是异构推理解决方案的坚定信徒。有哪些类别,你从与客户的交谈中如何看待它?

I was going to go into some of the other companies that were top of town this year, but this gives me the opportunity. I want to follow up on one thing: how do you think about the classes of workloads? To me, the ultra-fast frontier workload is just very clearly one of the fastest growing segments of the entire inference market. The fact that I want it and I cannot pay for enough of it—it might have been a mistake for OpenAI to offer it even, but it's good for us anyway. So just in your experience, you said you're strong believers in a heterogeneous inference solution. What are the buckets, and how do you see it from talking to your customers?

Sean

嗯,我认为这里可能有两种不同的观点。一种是产品观点,另一种是技术计算机架构师观点。从产品观点来看,实际上非常简单:带来更快的速度会开启重要的机会、应用、不同的用例、不同的能力、更智能的模型、更智能的智能体。你越能推动这一点,你就能启用更多。事实上,当你比我们今天所称的超快还要快时,可能有很多我们甚至无法想象的东西你可以构建,这就是我们不断推动的原因。然后另一个角度实际上归结为容量。就像,是的,我想要最快的,但你需要足够的容量来实际服务你的用例。所以真的就这么简单。然后找到正确的架构和正确的架构组合来提供这一点,这才是关键。

Well, I think there are probably two different views here. One is the product view, and the other is the technical computer architect view. From the product view, it's actually very simple: bringing more speed opens up significant opportunities, applications, different use cases, different capabilities, more intelligent models, more intelligent agents. The further you can push that, the more you can enable. In fact, there are probably all sorts of things we can't even imagine that you can build when you're even faster than what we call ultra-fast today, which is why we keep pushing that. And then the other angle really comes down to capacity. It's like, yes, I want the fastest, but then you need enough to actually be able to serve your use case. So it really is that simple. And then finding the right architecture and the right mix of architectures to provide that is really the name of the game.

异构分解架构 Heterogeneous Disaggregated Architecture

Sean

从计算机架构师的角度来看,在很多方面甚至更简单了,对吧?比如我在 Hot Chips 上跟人聊过这个,他们问我:如果你做解聚,那是不是意味着你得把数据中心分区,部署一定数量的这类硬件和另一类硬件,这难道不是限制吗?我说,你知道,当我们设计模块化时,它就是模块化的。而当我们设计芯片时,每天我们都在决定:我要把硅片面积用在内存上,还是算力上,还是 I/O 上?所以从计算机架构师的角度看,我看到的是工作负载中有非常不同的部分。最简单的就是预填充和解码,但再深入一层,这些阶段到底在发生什么?哦,有加载 KV 缓存,有做注意力机制,有把专家分布到各种不同的硬件上并做负载均衡。所有这些现在都可以用更定制化的硬件解决方案和架构来应对。在小规模下,把一堆不同的硬件凑在一起解决一个问题并不划算。但在我们讨论的规模下——几百兆瓦到吉瓦,甚至多吉瓦的规模——这很容易就回本了。到那时,就好比你把整个数据中心看作一台计算机,就像一块芯片,你要弄清楚:我需要这么多这种能力来快速跑注意力机制,我需要那么多能力来高效跑预填充。能够把这些拼在一起,就像计算机架构师的梦想,工具箱里有所有这些工具,对吧?所以这就是我认为你能从这个异构解聚生态系统中获得价值的方式。

Now, from the kind of computer architect's point of view, in a lot of ways, it's even a little bit simpler, right? Like I was having a conversation with somebody actually at Hot Chips about this and they were asking me, well, you know, if you're doing disaggregation, doesn't that mean you have to partition the data center and deploy a certain amount of this kind of hardware and a certain amount of another type of hardware? Isn't that restrictive? And I'm like, you know, when we design modular, it's modular. And when we design a chip, every single day we're deciding, am I going to use the silicon real estate for memory, or for compute, or for I/O? And so from a computer architect standpoint, what I see is there are significantly different parts of the workload. The simplest ones are prefill and decode, but then you go one level deeper. It's like, well, what's actually happening during these phases? Oh, well, there's loading the actual KV cache. There's doing the actual attention. There's spreading the experts across the various different hardware and balancing them. All of these things can now be addressed with more tailor-made hardware solutions and architectures. And at small scale, it doesn't really make sense to bring together a bunch of different hardware just to solve one problem. But at the scale we're all talking about—the hundreds of megawatts to gigawatts, the multi-gigawatt scale—it easily pays off. And then at that point, it's as if you're thinking about the entire data center as if it's like one computer. It's like one chip that you're trying to figure out: okay, I want this amount of this capability so I can run attention fast. I want this amount of capability so I can run prefill really, really efficiently. And being able to piece all these things together is like the computer architect's dream, to have all of these tools in our toolbox, right? So that's how I think you get the value from this heterogeneous disaggregated ecosystem.

对模型设计的影响 Impact on Model Design

Host

这在多大程度上影响了模型设计和模型架构?所以与 OpenAI 合作,你希望能看到模型的进展,比如之前那一整套推理模型。

How much of that plays into model design, model architecture? So working with OpenAI, you get to hopefully see what's going on with models, like there was the whole reasoning models thing.

Sean

这还不够。还有语音扩散。如果有更多交互式模型呢,比如思维机器那类东西?

It's not enough. There's also voice diffusion. What if there are more interactive models, like the thinking machine stuff?

Host

完全正确。而且我认为随着我们开始进入更多交互式用例,通过 Ultra 实现的,我们已经看到很多这样的情况。比如,如果你尝试交互式地做平面设计——你提到了扩散——图像生成或视频生成现在可能内联到你的创意流程中,成为交互式用例。所以有这么多不同的用例正在涌现。而我认为目前最未被开发的机会,坦率地说,对 Cerebras 而言,也对整个非 Nvidia 环境——非 Nvidia 社区而言——就是我们都在运行为 Nvidia GPU 设计的模型。而且不仅仅是 Nvidia GPU,通常它们是为某一款特定的 Nvidia GPU 设计的,对吧?比如,这东西是为在 B200、GB200、NVL72 上运行而设计的。所以反过来,比如我们展示过你已经能以 14 倍的速度运行模型,但那是为完全不同的架构设计的模型。所以如果你开始开放哪怕稍微调整模型架构的可能性,你就能获得巨大的收益。如果你开始做更多的协同设计,我认为机会几乎是无限的。而这正是让我对与 OpenAI 这样的客户和合作伙伴紧密合作感到非常兴奋的地方。

Absolutely. And I think as we start to get into more interactive use cases that we enable through Ultra, we're even seeing a lot of this right. So if you're trying to interactively, for example, do graphic design—now you mentioned diffusion—image generation or video generation might now be inline in your creative inflow, interactive use case with the thing. And so there are so many of these different use cases that are coming. And what I think is the most untapped opportunity right now, frankly for Cerebras but frankly for the entire non-Nvidia environment—and the non-Nvidia community—is that we're all running models that were designed for Nvidia GPUs. And it's not just Nvidia GPUs, but usually they're designed for one particular Nvidia GPU, right? Like, okay, this thing was designed to run on B200, GB200, NVL72. And so as a reverse, for example, here we were showing that you're running the model like 14 times faster already, but it was a model that was designed for a completely different architecture. And so if you start to open up the possibility of adjusting that model architecture even slightly, you can get massive gains. And if you start to do even more co-design, I think the opportunities are almost limitless. And that's really what excites me a lot about working closely with customers and partners like OpenAI.

AI代码生成内核 AI Codegen for Kernels

Host

是的。我们没时间深入这个,因为我们要继续,但外面的人真的低估了 AI 代码生成在内核上的应用。这对你们来说其实非常有利。

Yeah. And we don't have time for this because we want to move on, but the people out there are really sleeping on AI codegen for kernels. Which actually is very good for you guys.

Sean

是的。当然,绝对。我对此深信不疑。

Yeah. No, absolutely. I'm a huge believer in that for sure.

对AMD和Nvidia的热评 Hot Takes on AMD and Nvidia

Host

我想在进入 Etched 和超级劲爆芯片之前,你对 AMD、Nvidia 有什么看法?他们在做什么?他们有没有可能做点不一样的?有什么劲爆观点吗?

I think before getting into Etched and super spicy chips, any takes on AMD, Nvidia, what they're doing? Could they be doing anything different? Any hot takes there?

Sean

我认为 Nvidia 正在做的正是我们所有人都预料到 Nvidia 会做的事情,对吧?就像我说的,他们不是傻瓜。他们完全知道自己在做什么,并且继续推进。在很多方面,这绝对是正确的,对吧?因为我们需要更多的 token,我们需要更便宜的 token。这是现实。然而,曾经被认为每秒 100 或 200 个 token 就算快的,现在很快变成了新的批处理模式,对吧?很快变成了新的隔夜任务。所以我认为这有它的用武之地。它非常适合提示处理,非常适合非常并行的负载。

I think that Nvidia is doing exactly what we all expected Nvidia to be doing, right? Like I said, they're no dopes. They know exactly what they're doing and they're continuing to push through. And in many ways, it's absolutely the right thing, right? Because we need more tokens. We need them cheaper. That's a reality. However, what used to be considered fast at 100 or 200 tokens per second is quickly becoming the new batch mode, right? Quickly becoming the new overnight job. And so I think there's a place for that. It's great for prompt processing. It's great for very, very parallel workloads.

Host

但我的评估要花 20 小时。我按个按钮就能在两小时内跑完。

But my evals take 20 hours. I can hit a button and run it in two hours.

Sean

完全正确。完全正确。但很快,能够做提示处理还不够,对吧?那是市场的一大部分,但还不够。所以我认为所有传统架构都非常像是在那台跑步机上,对吧?以一种好的方式,不一定是坏的方式。是好的方式,对吧?所以 AMD、Trainium,在很多方面 TPU——所有这些在我看来都在试图造一个更好的 Ruben,对吧?这其中有巨大的价值,但我认为在其他方向上推动边界也有很多价值,这显然就是我们在 Cerebras 这里试图做的。

Exactly. Exactly right. But quickly, being able to do prompt processing isn't quite enough, right? That's a big part of the market, but that's not quite enough. And so I think all the traditional architectures are very much all on that treadmill, right? In a good way, not necessarily a bad way. It's in a good way, right? And so AMD, Trainium, in many ways TPU—all of these in my mind are all trying to build a better Ruben, right? And there's a huge amount of value in that, but I think that there's also a lot of value in trying to push the boundaries in other vectors as well, which is obviously what we're trying to do here at Cerebras.

对Etched的批评 Etched Criticism

Host

好的。劲爆,劲爆的芯片。你们网站上有个页面,Etched 这家公司就是没有。没有数字,没有基准测试。有什么看法吗?这是他们唯一的批评点。但那是 Etched。

Okay. Spicy, spicy chips. There's this page that you have on your website that this company Etched just doesn't have. No numbers, no benchmarks. Any takes? That's their one criticism. But that's Etched.

Sean

嗯,他们对自己在做什么说得很少。但他们说过的内容,与他们最初说的相比,也改变了很多。

Well, so they have said very little about what they're doing. But what they have said has changed also a lot compared to what they originally said.

对Groq和Etched的反应 Reaction to Groq and Etched

Sean

他们承认环境变化很大,所以转型并寻找自己的定位是合理的。我看到这类图片时的反应是,图形设计很惊艳,但我也没看到他们在构建超越传统 GPU 的东西,或者试图做得更好。他们声称拥有分布式存储,因为分布式所以更快,这我不太理解。如果他们能解释这类事情会很有帮助。很难看出他们想推动的差异化在哪里。我觉得他们在 Hot Chips 上展示了一个机架之类的,所以显然他们在构建东西,但感觉还有点遥远。

They acknowledge the environment is changing a lot, so it makes sense they're going to pivot and try to find their space. My reaction when I see pictures like this is that it's very impressive graphic design, but I also don't see them building anything beyond or trying to build something better than just a traditional GPU. They've made claims about having a distributed storage that is somehow faster because it's distributed, which I don't quite understand. It would be very helpful if they could explain those kinds of things. It's hard to see where the differentiation is that they're trying to push. I think they had a rack or something at Hot Chips, so clearly they're building something, but it feels like it's a little bit far off for now.

Host

是的。但与此同时,现在有 200 亿美元,可不是好惹的。我觉得也许最有力的辩护是:你希望他们朝哪个有趣的方向发展?最理想的情况……

Yeah. But at the same time, $20 billion now, it's not a pushover. I think maybe the steel man for this is: what interesting direction would you want to see them pursue? In the best case...

Sean

最理想的情况,我希望能看到他们或其他人,坦率地说,在芯片之外推动更具创新性的解决方案。我坚信,将 AI 算力提升到新水平的下一前沿在于芯片外更好的集成方式:机架、节点、封装、互连、工厂。他们可能在做类似的事情,很难说。但我认为这不仅是针对 Etched 的问题,而是整个行业的问题。即使我暂时摘下 Cerebras 的帽子,最让我兴奋的是单一芯片之外的想法和创造性思维。我们都知道,应用、问题、模型现在如此庞大,你无法在单个芯片上完成所有事情。所以一切都归结于集成:如何汇聚更多算力、更多内存、更多互连?

In the best case, I would love to see them or others, frankly, push the boundaries of more innovative solutions outside of just the chip. I really believe that the next frontier of taking AI compute to the next level is all about better ways to integrate outside the chip: the rack, the node, the package, the interconnect, the factory. They may be doing something like that; it's hard to tell. But I think this is not just a point about Etched; I think it's for the entire industry. Even if I take my Cerebras hat off for a second, what excites me the most is ideas and creative thinking outside of the single chip. We all know that the applications, the problems, the models are now so large that you can't just do anything on a single chip. So everything comes down to integration: how do you bring together more compute, more memory, more interconnect?

Host

调度、流水线,所有这些算法……

Scheduling, pipelining, all those algorithms...

Sean

所有这些。所以昨天发布的内容中,我觉得可能最有趣的一个——虽然不是新消息——是 DMatrix,他们真的在押注 3D DRAM 封装。我认为这是我们整个行业需要做得更多的。

All of those things. So of the reveals yesterday, I think probably one of the more interesting ones to me—it wasn't new news—was DMatrix, the fact that they're really leaning into 3D DRAM packaging. This is the kind of stuff that I think we as an industry need to be doing much more of.

Host

我正想说,你们的 backpack 技术让我想起内存行业在做的事情。这个类比……

I was going to say, your backpack stuff reminds me of what the memory people are doing. Is that analogy...

Sean

有一点类似,对吧?因为我们用 backpack 和供电做的事情,本质上是一个非常紧凑的 3D 封装,可以把电力直接送到芯片或晶圆表面。在很多方面,DRAM 3D 堆叠集成也在做类似的事情,但不是供电,而是内存。昨天我们还宣布,两年前就启动了 DRAM 堆叠项目,原因完全相同。找到做这类集成的方法,是解锁下一个重大进步的关键。所以我不太清楚 Groq 在做什么,或者抱歉,Etched 在幕后做什么,但当你问我行业中什么最让我兴奋时,其实是更有创意的集成、更有创意的封装、更有创意的芯片外思维。

There's a little bit of that, right? Because what we're doing with our backpack and power delivery is very much a very compact 3D package where we can bring power directly to the face of our chip, of our wafer. In many ways, the DRAM 3D stacking integration is doing something similar, but not with power—now with memory. Yesterday we also announced that we started our DRAM stacking program two years ago for the very same reason. Figuring out ways to do this kind of integration is the key to unlock the next major step forward. So I don't know exactly what Groq is doing, or sorry, what Etched is doing underneath the covers, but when you ask what excites me the most about what we're seeing in the industry, it's really more creative integration, more creative packaging, more creative outside-the-chip thinking.

其他知名公司与内存焦点 Other notable names and memory focus

Host

是的。除了 DMatrix,再给人们留点线索吧。沿着这个思路,还有哪些名字让你印象深刻?

Yeah. Let's leave some hints for people other than DMatrix. A couple other names that stood out to you just along these lines?

Sean

我觉得三星有一些关于他们的 ZHBM(高带宽内存)的讨论。

I think Samsung has some discussions around their ZHBM (HBM).

Host

都是内存厂商……

All memory guys...

Sean

你看,归根结底,构建 AI 芯片只有三个主要部分:算力、互连,然后是内存。而且越来越重要的是,正如我们所知,至少在低延迟领域,内存是关键——内存带宽是关键。

Look, in the end, there are only three main things to building an AI chip: there's the compute, there's the interconnect, and then there's memory. And more and more importantly, as we all know, at least in the low-latency space, memory is the key—memory bandwidth is the key.

Host

那供电和散热呢,还不是主要问题?

And power and cooling, not as major still?

Sean

不,绝对重要。这些是支撑这一切的。

No, absolutely. Those are what's powering all of this.

Host

是的。我只是说,物理极限……

Yeah. I'm just saying, physical limits...

Sean

但这其实是个很好的观点。当我们最初开始做晶圆级时,大多数人认为最大的挑战是:如何把这些裸片连接在一起?如何保证良率?这些确实是我们必须解决的根本性挑战。但最终成为最大推动力的其实是:如何供电?如何散热?如何大规模可靠地制造?我认为很多 3D DRAM 技术也会经历同样的过程。在 PowerPoint 幻灯片上画 DRAM 和逻辑是一回事,真正让它工作起来——弄清楚如何供电、如何散热——又是另一回事。这就是为什么我们一直在做的很多事情,是尝试利用我们在那方面积累的所有专业知识,并将其投入到我们的 DRAM 堆叠方案中。例如,我们已经解决了大规模良率问题。我们已经解决了如何真正进行三维封装的问题。而这些问题,三星、DMatrix 以及其他所有人,也都需要随着时间的推移去解决。

But that's actually a really good point. When we originally started with wafer scale, most people thought the biggest challenges were: how do you connect all of these die together? How are you going to yield this thing? And those were absolutely fundamental challenges we had to solve. But what turned out to be the biggest enablers was ultimately: how do you power it? How do you cool it? How do you build it reliably at scale? I think a lot of the 3D DRAM technologies are going to go through the same thing. It's one thing to draw a PowerPoint slide with DRAM and logic, and another to actually make it work—to figure out how you're going to power it, how you're going to cool it. That's why a lot of what we've been doing is trying to leverage all the expertise we've built up there and put it into our DRAM stacking approach. We already have solved yield at scale, for example. We've already solved how you can actually package in a three-dimensional way. And those are problems that Samsung, DMatrix, and everybody else are also going to have to solve over time.

美国供应链与中国独立 US supply chain and China independence

Host

感谢你的评论。最后一个话题,因为我们得走了:美国供应链,半导体供应链的问题。你知道,中国正在慢慢变得完全独立于我们。

Thank you for those comments. Last topic, because we got to go: US supply chain, semiconductor supply chain stuff. You know, China is slowly becoming completely independent of us.

中国开源模型 Chinese open-source models

Host

Vision GLM 简直就是在炫耀。华为什么的,在不知道它跑在什么上面就被大肆炒作。模型本身表现很好,你知道,在 OpenRouter 上发布,甚至不是跑在美国芯片上。所以问题是,大家都在谈论它吗?我们在做什么?

Vision GLM is like really literally just like they're bragging. Huawei and what have you, hyped outside of knowing what it was on. Just model is good doing good in, you know, OpenRouter came out not even running on US chips. So the TDR is, are people talking about it? What are we doing?

Sean

我觉得绝对是这样。当然大家都在谈论它,对吧?因为我们现在处境非常困难。开源模型市场现在 100% 是中国人的。

I mean, I think that absolutely. I mean, of course people are talking about it, right? Because we're in a very difficult situation right now. The open-source model market is 100% Chinese right now.

Host

不是 100%,但是……

Not 100%, but...

Sean

几乎是。好吧,95%,对吧。大多数大模型,大多数高质量的开源大模型,都来自中国实验室。从某些方面来说,共享正在发生,全球社区都在受益,这很好。但显然,对它们的模型如此依赖,这是一个非常具有战略挑战性的处境。

Almost. Okay, 95%, right. Most of the big models, most of the big open models that are high quality, are coming from the Chinese labs. In some ways it's great that sharing is happening, and the global community is benefiting. But obviously that's a very strategically challenging place to be, to have such dependence on them for the models.

Sean

另外,正如这次事件所证明的,我们独立地拥有所有基础设施,硬件基础设施,正在后台慢慢建设,以支持所有这些中国模型。我认为这绝对是一个问题,无论是全球范围,还是作为美国的国家利益,我们都必须继续突破边界,以便我们能够继续竞争并保持领先。这并不容易。这些人知道自己在做什么。但我相信这不是一家公司或一家晶圆厂能解决的,而是需要政府层面、国家利益层面的倡议才能实现。

Separately, as is evidenced by this, we have independently all the infrastructure, the hardware infrastructure, slowly being built up in the background to support all these Chinese models. I think it's absolutely a problem that globally, as well as for the US and as a national interest, we have to continue to push the boundaries so that we can continue to compete and continue to be ahead. It's not easy. These guys know what they're doing. But it's something that I believe is not solved by one company or one fab, but it needs government-level, national-interest-level kind of initiative to be able to make this happen.

Host

所以我想这应该在 Hot Chips 上发生。这就像你们所有人的密室。你们是我们行业的领导者,对吧?你们应该是那些人……

So I imagine that should take place at Hot Chips. It's like the secret room of all you guys. You're leaders of our industry, right? That you would be the guys to...

Sean

我的意思是,我们绝对一直在推动这件事,我们非常支持任何能继续美国在这一领域主导地位的倡议,100%。

I mean, we have absolutely been pushing for this, and we've been very supportive of any initiatives that will continue the US dominance in this space, 100%.

Host

好的。你得走了,但你非常慷慨地付出了时间。恭喜你取得的所有成功。我的意思是,巨大的成功,你 IPO 了。

Okay. You got to go, but you've been very generous with your time. Congrats on all your success. I mean, enormous, you IPOed.

Sean

是啊。是啊。是啊。有时我还得掐自己一下,提醒自己我们 IPO 了,而且是历史上最大的半导体 IPO,就像……

Yeah. Yeah. Yeah. Sometimes I still have to pinch myself to remind me that we IPOed, and it was like the largest semiconductor IPO in history, and it's like...

Host

还早。

Still early.

Sean

那还不错。那还不错。

That was not bad. That was not bad.

Host

不,恭喜。期待未来一年与你见面。

No, congrats. And look forward to meeting up with you in the future year.

Sean

我们得要求你,给我们超快。给我们更大,给我们更多。

We got to ask you, give us ultra fast. Give us bigger, give us more.

Host

不,不,我的意思是,也给人们超快。比如……

No, no, I mean, give the people ultra fast too. Like...

Sean

我们大概能让你在 OpenAI 排队。

We can probably get you in line at OpenAI.

Host

我希望你没有让 OpenAI 太偏向于,你知道,只把它留在内部。不,人们需要它。我们需要更好的模型。

I hope you didn't bias OpenAI too much to, you know, just keep it internally. No, people need it. We need better models.

Sean

这是 OpenAI 的决定。我们只是提供基础设施。

It's OpenAI's decisions. We're just providing the infrastructure.

Host

嘿,但给我们 GLM。给我们 Dec。

Hey, but give us GLM. Give us Dec.

Sean

只要给我们自己的机架,然后我们自己跑。

Just give us our own rack and then we'll run it.

Host

好的。再次感谢你的到来。真的非常感激。

Okay. Thank you again for coming. Really, really appreciate it.

Sean

嗯。

Yeah.

互动版:逐字朗读 + 针对本期提问 →