The Largest Chip Ever and the Speed of AI
打开互动全文版(中英对照 + 朗读 + 问答)→Cerebras 联合创始人兼 CEO Andrew Feldman 讨论为什么大芯片是加速 AI 处理的最佳方式,以及他的公司如何制造出计算史上最大的芯片。
Andrew Feldman, co-founder and CEO of Cerebras, discusses why big chips are the best way to process AI faster, and how the company built the largest chip in computer history.
这是计算机行业历史上造出的最大芯片,面积是普通 GPU 的 58 倍。对 AI 来说,更大的芯片能更快地处理信息,因此你得到答案的时间更短。对于 AI 世界来说,大芯片无疑是正确的方向。推理领域没有摩尔定律。从 GPU 切换到我们的云服务,只需要敲八个键。我们解决了计算史上没人解决的问题。我们在 2020 年将它交付,却无人问津。没人买账,也没人在意。所有人都说我们疯了,这根本行不通。于是我们造出了下一代。
This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly and therefore you get answers in less time. For the AI world, big chips are undoubtedly the best way to go. There's no Moore's law in inference. It takes you eight keystrokes to move from a GPU to us in the cloud. We solved a problem that nobody in the history of computing solved. And we delivered it in 2020, and nobody cared. Nobody bought it, and nobody cared. Everybody said we were crazy. It would never work. So then we built the next one.
嗨,我是 Matt。欢迎收听 Matt 播客。我今天的嘉宾是 Cerebras 的联合创始人兼 CEO Andrew Feldman。Cerebras 造出了计算史上最大的芯片,并且刚刚完成了有史以来规模最大的半导体 IPO。Andrew 最近到处都在聊头条新闻:与 OpenAI 的 200 多亿美元交易、这次 IPO。但这场对话有点不一样。我们从什么是晶圆讲起,一步步往上聊。为什么 GPU 难以做到快速推理、三个没人谈论的短缺问题、十年无人理会他的芯片的艰难岁月,以及为什么 Andrew 认为 CUDA 对 Nvidia 来说不再是护城河。如果你真的想了解当下芯片格局,以及 AI 推理在硅片层面是如何工作的,这期节目就是为你准备的。请欣赏我和 Andrew Feldman 的这场精彩对话。
Hi, I'm Matt. Welcome to the Matt podcast. My guest today is Andrew Feldman, co-founder and CEO of Cerebras, the company that built the largest chip in the history of computing and just pulled off the biggest semiconductor IPO of all time. Andrew has been everywhere talking about the headlines: the $20 billion-plus OpenAI deal, the IPO. But this conversation is a bit different. We started from what is a wafer and built up step by step. Why GPUs struggle with fast inference, the three shortages nobody talks about, the decade in the desert when nobody wanted his chip, and why Andrew believes that CUDA is no longer a moat for Nvidia. If you want to actually understand the current chip landscape and how AI inference works at the silicon level, this episode is for you. Please enjoy this fantastic conversation with Andrew Feldman.
我想从速度这个有趣的话题开始。那么,速度是否已经成为如今 AI 讨论中的主旋律?
I thought a fun place to start would be to talk about speed. So, has speed become the dominant conversation for AI today?
没错。我认为,很长一段时间里,AI 只是一种新鲜玩意儿,就像变戏法一样,酷但没什么用。后来到了某个节点,AI 变得足够聪明,人们开始使用它。要知道,我们用训练构建 AI,但用推理来使用 AI。突然之间,人们想用了;一旦你想用,它就有生产力了。速度很重要。快速的 token 更有生产力。于是整个讨论的焦点都变成了:我们如何让推理更快?如何更快地生成 token?因为这些 token 更有生产力。我们在更短的时间里做更多的事,因此它更有价值。
Yeah. What happened, I think, was for a long time, AI was sort of a novelty. It was like a parlor trick. It was cool but not useful. And at some point, AI got smart enough that people began to use it. And remember, we make AI with training, but we use it with inference. Suddenly people wanted to use it, and the minute you want to use it, the minute it's productive. Speed matters. Fast tokens are more productive. And so the conversation moved from everything else to: How do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens. We get more done in less time, and therefore it's more valuable.
太好了。那速度意味着什么?是 token 速度的公式吗?是任务完成时间吗?正确的衡量标准是什么?
Great. And what does speed mean? Is that an equation of token speed? Is that completion of the task? What's the right metric?
正确的衡量标准是每个用户每秒生成的 token 数。也就是从响应的第一个 token 到最后一个 token 的速度。聊天查询是如此,智能体式的负载也是如此,对吧?如果是多轮式的任务,等待的代价会被放大。所以你希望的是极快的响应,让 AI 感觉像实时互动一样,你才能跟它持续对话。
The right metric is tokens per second per user. That's how fast you get from the first token all the way through the last token in your response. It's true for queries from chat, but it's also true for agentic loads, right? If they're sort of multicycle turns, waiting is amplified. So what you want is blisteringly fast responses so that the AI feels like it's in real time. You can engage with this.
所以,这就是 AI 的“大脑时刻”。
So it's the brain moment for AI.
我觉得没错。而且我认为这是个很好的类比。想想 Netflix:互联网慢的时候,Netflix 是靠信封邮寄 DVD 的。你会收到一个装着 DVD 的信封。当互联网变快之后,他们并没有在邮寄 DVD 上变得更高效,而是变成了一家电影制片厂。速度让他们变成完全不同的东西。这就是速度的一般作用,尤其对于 AI:它打开了一个全新的领域,让你以不同的方式使用 AI。你会停留更久,会更频繁地回来,也会处理更难的问题。
I think that's right. And I think that's a very good analogy. If you think of something like Netflix: when the internet was slow, Netflix delivered DVDs in envelopes. You would get a DVD in an envelope. And when the internet became fast, they didn't get more efficient at delivering DVDs in envelopes. They became a movie studio. The speed enabled them to become something completely different. And that's what speed does in general, and in particular for AI. It opens up a whole new domain. It allows you to use the AI differently. You will stay longer, you will come more often, and you work on harder problems.
是的。所以这实际上就是用户体验的问题,对吧?没人愿意等几秒钟。
Yep. So it's literally a question of the UX, right? That's just nobody wants to wait a few seconds.
确实。我想说,慢搜索的市场能有多大?拨号上网的市场有多大?是零。你会等一个网站加载多久?会等八秒吗?没人会等。AI 也是一样的道理。
That's great. I mean, how big is the market for slow search? How big is the market for dial-up? Zero. How long will you wait for a website to resolve? Will you wait eight seconds? Nobody waits. It's the exact same with AI.
是的。所以不会再有人们开着笔记本电脑,傻等智能体——
Yep. So no more people waiting with their laptops open while the agent —
没错。它就是跑啊跑啊跑。我觉得这不是人们想要的。
That's right. Well, it's running and running and running. I think that is not what people want.
好的,太棒了。我想聊聊当下芯片行业的格局,帮助大家直观了解你们所在的位置。过去基本只有一个概念:一块芯片包打一切。显然现在情况已经发生了翻天覆地的变化。有 GPU,有人听说过 Trainium,也有人听说过 TPU。那么,请你帮我们打个比方,比较一下:谁做什么,用来做什么?
Okay. Great. Wonderful. So I'd love to talk about the landscape of the chip industry right now to help people visualize where you guys are. So there used to be basically this concept of one chip to do it all. And obviously this is evolving dramatically. There are GPUs, there are people who have heard of Trainium, people who have heard of TPU. So help us compare and contrast: who does what for what?
嗯,我认为过去确实有一块包打天下的芯片,那就是 CPU。没错,后来出现了用于独立显卡的协处理器。随着 AI 负载变得有趣,我们越来越把芯片聚焦于这类特定负载。如今,有几家公司做传统的 GPU,Nvidia 和 AMD 做的就是非常标准的 GPU。还有一组公司——超大规模云厂商——会自己做一些芯片。比如 TPU——
Well, I think there used to be one chip to do it all called the CPU. And right, there emerged a co-processor to do discrete graphics. And as the AI workload became interesting, we focused chips more and more on that particular workload. Today, several companies make traditional GPUs. So Nvidia and AMD make very standard GPUs. There are also a group of companies — the hyperscalers — that make some of their own parts. So the TPU —
所以 TPU 是谷歌的。
So TPU is Google.
TPU 是谷歌的,Trainium 是 AWS 的。然后还有一些公司,比如 Cerebras。我们是率先从头打造芯片的先行者之一,专门针对 AI 优化,不做别的。我们不在超大规模云厂商内部;我们不是针对某个云厂商的问题或某个实验室的问题来优化。我们是在为一组 AI 问题设计芯片,我们所有的思考都围绕如何加速 AI。
TPU is Google, and Trainium is AWS. And then there were a group like Cerebras, and we were among the pioneers to build a part from scratch, optimized for AI and nothing else. And we weren't inside of a hyperscaler; we weren't optimized for a hyperscaler's problem or for one lab's problem. We were building a chip for a collection of AI problems, and all our thinking was around how to accelerate AI.
到目前为止,芯片行业大致就是这样一幅格局。
And that's sort of the landscape to this day.
是的。
Yeah.
还有更专门化的芯片,对吧?比如专门为 Transformer 开发的。这些也叫专用集成电路(ASIC)。你能定义一下这个词的意思吗?
There's even more specialized ones, right? Like one developed specifically for transformers. Then those are called ASICs as well. Could you maybe define what that term means?
ASIC 就是专用集成电路(application-specific integrated circuit)。这个词现在的含义很广。它的意思是,你做出了一系列选择,从通用走向更窄的一类问题求解,并且在架构中做出了一些决定,让你在某些方面表现很好,而在另一些方面表现很差。这些选择遍布整个光谱。比如 TPU 就做了类似的选择:它不能做图形,但非常擅长矩阵乘法;对于我们在数学中做的许多其他事情,它并不擅长。GPU 也一样。我们各自做出了不同的选择。目前在生产中的芯片有一批:谷歌有 TPU,AWS 有 Trainium,微软的 Maia 即将问世,还有 Cerebras,以及另外一家被 Nvidia 收购的芯片公司。
ASIC is an application-specific integrated circuit. It's a word that now has a wide range of meaning. It means that you have made a series of choices away from general towards a narrower class of problem solving, and that you've made some decisions in your architecture that make it much better at some things and much, much worse at others. These choices are made across the spectrum. So the TPU has made choices like that. It can't do graphics. It's very good at matrix multiply. It's not very good at a collection of other things that we do in mathematics. Same for the GPU. We've all made different choices. Right now in production, there's a collection: there's the TPU for Google, Trainium for AWS, just coming up is Maia, a part from Microsoft, Cerebras, and there was one other that got acquired by Nvidia.
关于通用芯片与专用芯片之争,有意思的是,当 Groq 被收购时,你和你们 CTO 都庆祝了一番。那是不是 Nvidia 承认了你们一直以来的想法是对的?
And on the general chip versus specialized chip, it was actually interesting because you and your CTO celebrated when Groq was acquired. Was that a recognition by Nvidia that your vision was right all along?
是的,我认为 Nvidia 最根深蒂固的迷思之一,就是认为 GPU 无所不能,是 AI 唯一需要的东西。而他们以 200 亿美元收购 Groq,以及他们做出这一决定的结构和速度,向所有人证明了事实并非如此。GPU 架构无法快速推理,而这个市场庞大且增长迅速。我们是其中速度最快、规模最大的。我们的销售额是 Groq 的 10 倍以上,而他们为第二名支付了 200 亿美元。
Yeah, I think one of Nvidia's most durable myths was the perception that the GPU could do everything and it was all you needed for AI. And the acquisition of Groq for 20 billion and the structure and the speed with which they chose to do it made clear to everyone that that wasn't true. The GPU architecture couldn't do fast inference, and that this market was large and growing quickly. We were the fastest at it and the largest. You know, our sales were more than 10 times Groq's, and they paid 20 billion for the number two player.
所以那是美好的一天。
So that was a good day.
确实是美好的一天,对吧。
That was a good day. Right.
最近 OpenAI 与博通之间关于 Jalapeño 的公告——OpenAI 是你们的大客户——这在这个图景中处于什么位置?这是那种高度专用的 ASIC 吗?对吧,用于推理?
And the recent announcement of Jalapeño between OpenAI — which is a very large customer of yours — and Broadcom. Where does it fit in that picture? Is that one of those highly specialized ASICs? Right? For inference?
没错。Jalapeño 是一个酝酿已久的项目。它最初是与博通一起公布的,后来 OpenAI 才和我们做了那笔交易。记得我们做了一笔大交易——这可能是硅谷历史上最大的交易,超过 200 亿美元。我认为他们对硅片有着巨大的需求。OpenAI 在行业中做得最好的事情之一,就是正视指数级的采用曲线,并且不怕它意味着什么。其他人为了获得算力不得不去签一些很糟糕的合同,而 OpenAI 预见到了这一切。他们为内存、算力、和我们以及和其他人都做成了大笔交易。他们在理解指数曲线外推的含义方面确实很有远见。而那条曲线就是 AI 的使用量,对吧?它增长得难以置信地快。整个行业都在追逐需求。对吧?我们通常——嗯,很多时候恰恰相反。通常人们是先构建出来,希望需求会到来。而我们的情况是,所有人都在追逐人们已经想做的事情,更不用说他们未来可能想做的事情了。
That's right. Jalapeño is a part that has been a long time coming. It was announced with Broadcom before OpenAI did the deal with us. Remember, we did a huge deal — this is probably the largest deal in Silicon Valley history, north of 20 billion dollars. I think they have a yawning need for silicon. And one of the things OpenAI has been the best in the industry at is looking at an exponential adoption curve and not being afraid of what it says. Others have had to go out and strike really bad deals to get capacity, where OpenAI saw this coming. They struck big deals for memory, for compute, with us, with others. They've really been visionary in understanding what it means to extrapolate from an exponential curve. And that curve is AI usage, right? It is growing so unbelievably fast. The whole industry is chasing demand. Right? We usually — well, often it's the other way. Often people are building it hoping it will come. In our case, all of us are chasing what people already want to do, let alone what they might do in the future.
不过,嗯,既然这是个很有趣的话题,最后再问一下 GPU、推理与 ASIC 的对比——从你的角度看,我们会永远处于这种多芯片环境吗?
But, um, because that's such an interesting topic, to finish on GPU versus inference versus ASIC — is that the permanent situation from your perspective, that we're going to be in this forever multi-silicon environment?
是的,多芯片环境是一个健康的生态,对吧?我不认为有人会说 x86 生态是一个健康的地方。20 年来只有两个玩家。而如果你注意的话,当新的工作负载出现——比如手机工作负载——它非常相似,但需要低功耗和电池,而两家都失败了。对吧?英特尔拿零份额,AMD 拿零份额,而一家之前没人听说过的公司成了全球最大的算力销售商——那就是 Arm。健康的生态有很多不同的方式来解决问题。
Yeah, multi-silicon environment is a healthy ecosystem, right? I don't think anybody would say that the x86 environment was a healthy place. There were for 20 years there were two players. And if you notice, when a new workload came around — cell phone workload — it was very similar but required low power and battery, and both lost. Right? Intel got zero share, AMD got zero share, and a company nobody previously heard of became the largest seller of compute in the world — and that's Arm. And healthy ecosystems have lots of different ways to solve problems.
中国在这一切中处于什么位置?我们看到华为昇腾芯片用于 DeepSeek,以及一个全栈式中国 AI 工厂的出现——找不到更好的词。这种分化正在发生,还是被夸大了?
Where does China fall in all of this? So there's the Huawei Ascend chips for DeepSeek, and this emergence of a full-stack Chinese AI factory, for lack of a better term. Is that divide happening, or is that overstated?
不,不,这种分化是真实存在的。我认为他们是一个产业对手。我觉得他们做了非常有趣的投资——在电力、电网方面的投资,而电网恰恰是美国的软肋。如果你在法国,你有核电,结果核电相当便宜又干净。但中国在电网上投入巨大。他们在芯片上落后,但他们的方式更上一层——开源模型。他们正在生产一些非凡的模型,虽然不如 GPT、Anthropic 或 Google 的 Gemini,但已经非常好。而且我认为他们没有美国那样的芯片制约,但他们在其他领域有领导力,可能是在电力方面,而电力正是我们数据中心所需要的。
No, no, that divide is real. I think they are an industrial adversary. I think they have made really interesting investments — investments in power, in their grid, which is a real weakness in the US. If you're in France, you have nuclear, which turns out to be pretty cheap power and pretty clean. But in China they made huge investments in their grid. They are behind in chips, but their approach was at the next level — open-source models. They're producing some extraordinary models, not as good as GPT or Anthropic or Google's Gemini, but very good. And I think they don't have the same chip constraints that the US does, but they've got leadership in other domains, potentially in power, which is what we need for data centers.
那中国在你们的战略中是什么位置?
Where does China fall in your strategy?
我们不向中国销售——出于地缘政治原因、商业原因、还有……
We don't sell in China — for geopolitical reasons, for business reasons, for...
监管原因?
Regulatory reasons?
出于监管原因,还有一些地缘政治原因。
For regulatory reasons and for some geopolitical reasons.
好的。本地 AI 和本地芯片有发挥空间吗?Nvidia 有一个关于为 Windows 电脑制造芯片的公告。那是用于推理吗?你们认为那应该是多芯片生态的一部分吗?
Okay. Is there a role for local AI and local chips? Nvidia had some announcement around just building chips for Windows computers. Is that for inference? Is that something that you guys think should be part of the multi-ecosystem?
是的,我认为如果你看看应用和云生态是如何兴起的,你会发现凡是你能做的事情,都应该在手机或笔记本上做。但要把真正的处理能力放到手机或笔记本上,会受到限制,因为它们通常靠电池工作。所以你会想尽可能多地做,尽可能靠近数据去做。但事实是,在大多数情况下,对于真正的算力需求,AI 也会是这个路子。我们会在手机、笔记本上做很多工作,但对于大型任务,你还是要回到数据中心。这就是我们关注的重点。我们的重点是数据中心算力。
Yeah, I think if you look at the way the ecosystem for apps and cloud emerged, everything you can do you should do on your phone or your laptop. But the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you want to do as much as you can, as close to the data as you can. But the truth is, in most situations, for real compute, that's exactly the way it's going to be with AI. We're going to do a lot of work on the cell phones, on the laptop, but for the big work you're going to go to the data center. And that's where our focus is. Our focus is data center compute.
很好。你提到了需求有多么强劲,以及生产端的需求。那么不可避免要问一下市场均衡——泡沫与否。最近大家都看到,6 月初出现了一次闪电式暴跌,市值蒸发了 1.5 万亿美元,Nvidia 当天就损失了 3000 亿美元。不过 48 小时内就恢复了。所以市场在这个概念上看起来非常紧张。仅基于你所说的,听起来你是站在这种担忧的另一边。
Great. So you mentioned how powerful the demand was, and demand out of production. So just to ask the inevitable question around market equilibrium — bubble or not. So the most recent thing we all saw is that there was a bit of a flash crash earlier in June where $1.5 trillion in market value was lost, and Nvidia had lost $300 billion that day. Now that was corrected within 48 hours. So the market seems very jittery around that general concept. Just based on what you said, sounds like you reserve a place on the other side of that concern.
我认为,如果你花大量时间盯着公开市场,每天看着它的起起落落,那你就是在看错东西。我从公开市场投资的伟大思想者那里得到的智慧是:短期来看,公开市场是投票机——看谁最受欢迎;但长期来看,它是称重机——看谁创造了最大的价值、最重的分量。所以,作为一家新上市公司,我们能做的就是不去看日内的波动,专注于打造非凡的技术、赢得新客户、让现有客户满意、确保他们购买更多、每周都变得更好。所以我认为日内的波动不是我们能控制的。那就是投票的部分。
I think if you spend a lot of time staring at the public markets and watching their ups and downs every day, you're watching the wrong thing. I think the wisdom received from the great thinkers of public market investing is that in the short term, the public markets are voting mechanisms — who's most popular. But in the long term, they're weighing mechanisms — who's created the most value, the most weight. And so what we can do as a newly public company is we can not look at the day-to-day fluctuations, and we can focus on just building extraordinary technology, winning new customers, making our existing customers happy, ensuring they buy more, being better every week. And so I think the day-to-day fluctuations are out of our hands. And that's about the voting.
嗯。
Yeah.
我想唱个反调:看起来很大一部分需求来自那些大型实验室,而实验室本身是由风险投资、私募股权、对冲基金,或者说主权投资者资助的。那么,芯片需求来自这些可能被资本人为堆出来的实验室,你会不会觉得,如果交易这里那里多绕几圈,就能感到某种脆弱性?
And I mean, just to play devil's advocate on the demand—it seems like a lot of the demand comes from the big labs, which themselves are financed by venture capital, private equity, hedge funds, or whatever you call it, and sovereign investors. Is there any concern that the demand for chips comes from labs that are maybe artificially financed, and if you add a few circles of deals here and there, you sense any fragility?
嗯,这跟我们的实际体验不太一样。我们确实有来自 OpenAI 的巨大需求,但还有几十个其他客户在尝试下非常大的订单。
Um, that's not exactly our experience. I mean, obviously we have enormous demand from OpenAI, but we have dozens of other customers who are trying to place very big orders.
从历史上看,泡沫是供给跑到了需求前面。90 年代我们建设电信基础设施,提前好几年铺设光纤,花了六到八年才全部用上,那是“建好了自然有人来”的心态。而现在的 AI 不同,我们都在努力追赶。我们想更快地建设数据中心,想为今天已经存在的需求扩充供应链。而全世界只有极小一部分人在以接近 AI 潜力的方式使用它,我们就已经快被算力淹没了。内存需求的紧缺是真实的短板,GPU 并不是我们要面对的问题,整个生态左右都受限。这不像泡沫。
Historically, bubbles were when supply got out ahead of demand. In the '90s, we built out telco infrastructure—we built out fiber years before it was going to be used. It took six or eight years, and it all got used, but it was a sort of 'if you build it, they will come' mentality. What's different about AI right now is we're all trying to catch up. We're trying to build data centers faster, we're trying to increase our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. We're already sort of overwhelmed with compute. We're overwhelmed with the demand for memory, which is a real weakness. The GPUs are not a problem we face—the ecosystem has constraints left and right. That doesn't feel like a bubble.
你能再展开讲讲吗?这太有意思了。瓶颈似乎在不断转移。那目前的内存短缺问题是什么?大家可能都听说过一些。
Do you want to actually double-take on that? Because that's so interesting. It seems like the bottleneck keeps moving. So what's the current problem of a memory shortage that people may have heard of?
现在有三大瓶颈。第一,除了我们之外,所有 GPU 和所有 ASIC 都使用一种叫 DRAM 的内存,以及一种叫 HBM 的特殊类型。这种内存全世界只有三家公司生产:三星、SK 海力士和美光。三家公司,而且已经卖光了。这是第一大问题。而我们不用它。
So there are three major bottlenecks right now. The first is that all GPUs and all ASICs except us use a type of memory called DRAM, and a particular flavor of that memory called HBM. That memory is made by three companies in the world: Samsung, SK Hynix, and Micron. Three companies, and they're sold out. That's the number one problem. We don't use it.
所谓卖光是指——像硬件领域通常那样,只有建新工厂才能增加供给?
And it's sold out as in—like the only way to increase supply would be to build new factories?
建更多工厂。于是 GPU 就面临这个巨大的问题,现在连 CPU 也这样。你拿不到这种内存。而由于我们的架构选择,我们不用它。第二个制约是台积电的一项工艺,叫 CoWoS。它把硅片当作主板,英伟达、AMD 等厂商都用它做基板。他们把芯片和内存芯片放在上面,这比传统主板更高效。这项工艺也已经卖光了,受限非常严重。我们不用它。第三个制约是台积电的 3 纳米工厂产能。这是最先进的工厂。我们的芯片是 5 纳米,所以我们不用 3 纳米技术。
Build more factories. And so you have this huge problem for GPUs, and even now for CPUs. You can't get this memory. Because of our architectural choices, we don't use it. The second constraint was a process at TSMC called CoWoS. This used silicon as a motherboard, and this is what NVIDIA and AMD and others use as a motherboard. So they put their chips and the memory chips on it, and it's more efficient than a traditional motherboard. That process is sold out, highly constrained. We don't use it. Third limitation in the space is 3-nanometer factory space at TSMC. This is the most aggressive factory, with the most advanced technology. Our chips are at 5 nanometer, so we don't use the 3-nanometer technology.
这太有意思了。那么问个外行问题:所谓“工厂”是什么意思?是有一家专门做这个的工厂吗?
And that's so fascinating. So to ask the layman's question—what does that mean, 'factory'? There's a factory that's solely dedicated?
对。他们建厂,工厂蚀刻出间距一定大小的晶体管。间距越小,每平方毫米的晶体管就越多,能做的事情也就越多。计算机行业制造芯片的历史,就是我们不断把晶体管放得越来越近的历史。
That's right. They build factories, and the factory etches transistors that are a certain distance apart. The smaller the distance apart, the more transistors per square millimeter, and the more you can do per square millimeter. And so the history of the computer industry making chips has been we've gotten better and better at putting transistors closer and closer.
嗯。
Yep.
过去晶体管间距是 16 纳米,后来到了 10、7、5、3 纳米,现在他们正在做 2 纳米和 1.8 纳米。目前 GPU 领域大部分都挤在 3 纳米,工厂里拥挤不堪。
And so we used to be at 16 nanometers apart, then we went to 10, to 7, to 5, to 3, and they're working on 2 and 1.8. So right now the bulk of the GPU world is all 3, and there's tremendous congestion at the factory.
也就是说工厂里用的设备完全是另一套?
And that's literally different tooling at the factory?
对,工厂的设备。而我们用的是 5 纳米工厂。这就是我们的芯片。
Tooling at the factory. And we use the 5-nanometer factory. And this is our chip here.
嗯。
Yes.
这是计算机行业历史上制造的最大芯片。它比 GPU 大 58 倍,内存带宽多出 2500 到 3000 倍。对 AI 来说,更大的芯片处理信息更快,因此你得到答案所需的时间更短。显然,笔记本电脑、手机或很多其他设备不需要更大的芯片,但对 AI 领域来说,大芯片无疑是最佳路径。
This is the largest chip built in the history of the computer industry. It is 58 times larger than a GPU. It has two and a half or three thousand times more memory bandwidth. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. Obviously you don't want a bigger chip for a laptop or a cell phone or lots of other things, but for the AI world, big chips are undoubtedly the best way to go.
好,那我们先——我等会儿想深入聊一下这个芯片本身。但先把这块收个尾:三大瓶颈。你刚才还提到了 CPU,似乎也出现了 CPU 短缺这样的主题。是这样吗?第一,是真的吗?第二,原因是什么?
Okay, all right. So let's—I want to do a bit of a deep dive on that, on the chip itself, in a second. But to close on this, so three bottlenecks. You also mentioned CPUs a couple minutes in this conversation, and there seems to be a theme around the emergence of like a CPU shortage as well. Is that so? One, is that true? Two, what causes it?
智能体式 AI 是一个这样的世界:AI 不只是提供答案,它还会发起行动。这个行动可能是去访问一个网站,可能是去学习、从网站收集数据、取回来,再采取另一个行动。这些行动都是由 CPU 完成的。随着 AI 越来越擅长做事、越来越擅长发出指令调用,我们需要越来越多的 CPU。这就推高了 CPU 的消耗,进而推高了 CPU 的需求。这种对更多 CPU 的巨大推动,正来自像我们这样的机器和 GPU 上运行的智能体式 AI 工作——它们让 CPU 去执行动作:访问网站、订一份卷饼、查找一条信息、从存储中取出来。所有这些工作都由 CPU 完成。
So agentic AI is a world in which AI doesn't just provide answers—it initiates action. So that action might be go to a website, it might be learn, gather some data from a website, bring it back, take another action. Those actions are done by CPUs. And so as AI gets better and better at doing things, at making instruction calls for things to get done, we're using more and more CPUs, right? And that is driving up the consumption of CPUs, and therefore the demand for CPUs. So this huge push for more CPUs is being driven by AI on machines like ours and GPUs doing agentic work, asking the CPUs to take an action—go to a website, order a burrito, find a piece of information, pull it from storage. All that work is being done by the CPU. Yes.
包括在你们这样的系统里——也就是我们这样的系统里吗?
Including in a system like yours—in a system like ours?
哦,这很有意思。所以,所以——
Oh, it's interesting. So, so—
这是什么意思?你的芯片和 CPU 放在一起,组成一个系统?
What does that mean? So you have your chip and CPUs on the side, and those will be a system.
意思是,当我们做智能体式 AI 时,AI 处理器就像大脑,而 CPU 就像身体。它们在数字世界里采取行动、执行事情,受 AI 的指挥——AI 运行在服务器系统里的加速器上。随着我们做越来越多 AI 工作、越来越多智能体式工作,我们就会越来越多地调用 CPU,因此 CPU 的需求简直爆棚。
It means that when we do agentic AI, the AI processor is like the brains, and the CPUs are like the body. They're taking action. They're doing things in the digital world under the direction of the AI, which is running on the accelerator in the server system. So as we do more and more AI work and more and more agentic work, we're making more and more calls to CPUs, and therefore the demand for CPUs is just through the roof.
CPU 也面临同样的内存短缺吧?它们用的是一样的内存?
And CPUs experience the same memory shortage? They use the exact same memory?
没错。
That's right.
好。我们先回到产品上,不过稍微聊一下你和你的经历,给大家一些背景。你知道,很明显你做了一场令人难以置信的 IPO,时机非常完美。而且我知道那有点——
Okay. Yeah. Let's go back to the product in a second, but let's talk about you and your journey a little bit, for people to have context. So, you know, obviously you pulled off an incredible IPO with perfect timing. And I know it was a little bit—
你拥有完美时机的方式,就是先经历 10 年糟糕的时机。完美时机就是这么来的。
You make the way you have perfect timing is to have horrible timing for 10 years. That's the way you have perfect timing.
说得特别到位。
Very much to this point.
那么,你们在十多年前就开始构建这款专注于推理的更大芯片,对吗?
So, you guys started building this bigger chip focused on inference in, you know, over a decade ago?
2016 年。
2016.
2016 年,嗯,回想起来这完全说得通,因为构建这样一项技术需要很长时间。但当时的愿景是什么?对,2016 年,我想那是 ImageNet 四年之后,深度学习已经兴起了。
2016, uh, which makes perfect sense in retrospect, given how long it takes to build a technology like this. But what was the vision? Right, 2016, I guess it was like four years after ImageNet, deep learning was a thing.
当时一切都是,一切都是视觉,都是革命。因此我们认为,正确做法不是把最新最酷的模型嵌入到芯片电路中,那是错误的。你知道,我们在 Transformer 上是全球最快的,而我们的架构在 Transformer 出现之前就已经定型了,对吧?我们在扩散模型上也是全球最快的,而我们的架构在扩散模型出现之前就定型了。你要做的是抓住这些模型的底层基础,这样当市场转向时,你也能同样擅长。否则,因为模型的生命周期很短。
It was all, it was all vision, it was all revolutions. And that's why we don't believe the right thing to do is to embed the latest and coolest model into your circuitry. That's a mistake. You know, we're the fastest in the world of transformers and our architecture was set before transformers existed, right? We're the fastest in the world of diffusion before and our architecture was set before diffusion. What you want to do is get the underpinnings of those so that when the market moves you can be good at that as well. Otherwise, because it's short-lived.
回到之前关于 ASIC 的问题,这就是其他公司的做法。他们把模型架构做进……
To the ASIC question earlier, so that's what others do. They build the architecture of the model into the...
有些公司确实这么做了,而且从历史上看,这是一个结构性错误。
Some have. Some have, and historically that's been a structural mistake.
这确实是结构性错误。所以你们是做一种全新的、适用于任何模型的横向平台。
It's been a structural mistake. So you're doing the brand-new horizontal platform on any kind of model.
你要思考底层的计算是什么。所有这些工作的底层计算就是稀疏线性代数。
You want to think about what is the underlying calculation. The underlying calculation of all this work is sparse linear algebra.
对。
Yeah.
如果能加速这个,那么无论模型开发者发明什么,你都能让它更快。好的,这就是我们的思路。
And if you can accelerate that, whatever the model makers invent, you can make faster. Okay. And that was our approach.
对。你们从 2016 年就开始做了。嗯,你之前有一家公司卖给了 AMD。你在那里学到了哪些经验,并带到了 Cerebras?
Yeah. So you started in 2016. Um, you had a prior company that you sold to AMD. What were some of the lessons you learned there that you took into Cerebras?
我认为这些经验既广泛又复杂。我认为经验就是犯过错误并从中学习的另一种说法,对吧?作为团队,过去 25 年里我们制造过几十款芯片,芯片设计方面的经验回报是巨大的。我们在 C Micro 构建了一种不同类型的计算机,一种针对低功耗优化的计算机,并且针对一个与 AI 截然不同的工作负载进行优化,比如网页浏览。但根本的基础,作为一个计算机架构师要问的问题始终是一样的:我怎样才能让这项工作更快?以及是否有足够的需求让这件事值得去做?这就是我们问的两个问题。嗯,我们是否应该为此制造一个部件?我们是否,我们怎样才能构建一个为 AI 优化的芯片,并且是否会有足够多的 AI 来让我们围绕它建立业务?这些就是我们在 2016 年问的问题。而且你知道,另一面是,我是说,如果 GPU——这个为图形学优化了 20 年、一直把像素推向显示器的产品——突然在一个新领域表现出色,那岂不是太巧了?我们会认为这不是合适的架构。它只是比 CPU 更好,而我们能够构建一个架构,速度大大提升,功耗更低,还能降低成本。
I think the lessons are large and mixed. I think experience is another name for having made mistakes and learning from them, right? As a team, we've built dozens of chips over the past 25 years, and the returns to experience in chip design are enormous. We built a different type of computer at C Micro, a type of computer optimized for low power and optimized for a workload that was very different from AI, something like web browsing. But the fundamental underpinnings, the questions you ask as a computer architect, are always the same: what can I do to make this work faster, and is there enough of it to make it worthwhile? These are the two questions we ask. Um, should we build a part for it? Should we, what could we do to build a chip optimized for AI and will there be enough AI so that you can build a business around it? Those were the questions we asked in 2016. And you know, the flip side of that was, I mean, wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years and been pushing pixels to a monitor, was suddenly good at a new world? Would that be serendipitous? And we came to believe that it wasn't the right architecture for it. It was just better than the CPU, and that we could build an architecture that would be vastly faster, use less power, and drive down the cost.
这就是我们的旅程。所以 Cerebras 起步很早,但随后我们经历了一段荒漠期。在荒漠里,我们漫无目的地游荡。
And that was the journey. So Cerebras was very early, but then we spent time in the desert. In the desert. Then we wandered in the desert.
对,对。也许你可以带我们回顾一下那几年,给所有创始人,尤其是听这档节目的深科技创始人。那么首先,市场时机的问题是什么?是技术不奏效吗?然后你们作为团队是如何应对的?如何让团队和投资者继续支持,并筹集更多轮融资?因为你们大概还没有那种可以证明自己的里程碑。
Yeah. Yeah. Maybe walk us a little bit through those years, for any founders, especially deep tech founders listening to this. So what was, first of all, what was the issue with market timing? Was it the technology was not working? And then how did you go about it as a team, getting your team on board and your investors, and raising more rounds, since you presumably didn't have the proof points that one would need?
对我们来说,一开始我们就对风投很坦诚,告诉他们我们要攻克一个非常难的问题。我们不会去构建一个只比 GPU 好一点点的东西。我们的策略是,你永远无法通过做得比对方好一点点来击败像英伟达这样伟大的公司。他们会以更低的价格买到所有东西,他们有定价压力,他们能够捆绑销售。正确的策略是去做一个在工程上极其困难、但性能要好得多的事情——10 倍、15 倍、20 倍、30 倍、50 倍地好。但要做到这一点,那些普通而显而易见的路径都被堵死了,其他人已经把路走了。于是我们观察到,推理速度将取决于内存,而内存有两种。HBM 这种 VRAM 能存储大量数据,但速度慢。另一种内存叫 SRAM,速度快得惊人,但单位面积能存储的数据很少。而在图形领域,大家一直用的是 DRAM,用的是 HBM。AI 工作负载则根本不同。在图形处理中,你把数据搬到 GPU,然后长时间运算,最后把结果送出去。所以移动加运算的总时间由运算主导。而在推理和 AI 中,恰恰相反。你要移动海量数据——所有的权重,从内存到计算单元,然后做一次计算生成下一个词,然后又得再来一遍。所以所有时间都花在了数据移动上。这就是为什么 GPU 很难变得很快。好吧。所以我们观察到,如果我们采用 SRAM 策略,我们可以更快。但我们必须克服 SRAM 的缺点——它存不了太多。这让我们想到了解决方案:如果我们能制造一个比历史上任何芯片都大得多的部件,对吧?达到餐盘那么大。我们就可以塞满 SRAM,从而获得 SRAM 快速的优点,同时克服它容量小的缺点。这让我们得出了晶圆级芯片策略。这款芯片由一片晶圆直接制成,那块晶圆来自台积电。
So for us, at the beginning we were honest with our VCs and told them we're going to attack a really hard problem. We weren't going to build something that was a little bit better than a GPU. And our strategy was that you will never beat a great company like Nvidia by doing something a little bit better than they do. They're going to buy everything for less. They're going to have pricing pressure. They're going to be able to bundle. And the right strategy would be to do something incredibly hard in engineering that was way better. 10, 15, 20, 30, 50 times better. But to do that, the ordinary and obvious paths are all closed. Everybody else has taken the path. And so what we observed was that speed in inference was going to be a function of memory and that there were two types of memory. There's this V-RAM of HBM and they can store a lot but they're slow. There's another type of memory called SRAM that is unbelievably fast but per unit area can't store very much. And for graphics, everybody had always used DRAM, they used HBM. And the AI workload was fundamentally different. In graphics, you move data to the GPU and then you work on it for a long time and then you send the result. So the total time spent on movement plus work is dominated by work. In inference and AI, it's the exact opposite. You move a huge amount of data, all the weights, from memory to compute and you do one calculation to generate the next word and then you have to do it again. So all the time is dominated by the movement of data. So that's why GPUs have so much trouble being fast. Okay. So we observed that if we chose a strategy using SRAM, we could be faster. But then we have to overcome the tradeoff with SRAM. It can't store very much. That led us to the solution. If we could build a part vastly larger than any part in history, right? At the size of a dinner plate. We could stuff it to the gills with SRAM and thereby get the benefit of SRAM, that it was fast, and overcome the weakness that it can't store very much. And that led us to a strategy called wafer scale. This chip is made from a single wafer that comes out of TSMC.
你想定义一下什么是晶圆吗?
Do you want to define what a wafer is?
晶圆?所有芯片都是从晶圆上切下来的。晶圆是一片直径 300 毫米的圆形硅片,芯片制造的过程就像你妈妈用饼干模具压出饼干一样,压出一块块芯片。在我们之前,最大的芯片是 800 平方毫米,确切地说是 840。而这款是 46,000。所以我们必须发明各种新技术来制造这么大的芯片。造出来之后,还得发明给它供电、散热的方法。没有供应商等着给你配套,对吧?因为你做的这个东西前所未有。这花了好几年时间。这是一个深科技难题,以前从来没人解决过。我们有一个 18 个月的时期,每月花费 800 万美元,却造不出来。所以对你们这些深科技创始人来说,你们有……
A wafer? All chips are cut from a wafer. A wafer is a circular piece of silicon that's 300 mm across, and the process of chip making stamps out, like your mother does with a cookie cutter, stamps out chips. The biggest chip that had ever been built before us was 800 square millimeters. 840 to be exact. And this is 46,000. So we had to invent all sorts of new technology to build a chip this big. And once you build it, you have to invent ways to power it, to cool it. There are no vendors waiting for you. Yeah. Right. Because you look like nothing else ever made. And that took years. And it was a deep tech problem. It had never been solved before. And we had an 18-month period where we were spending 8 million a month and we couldn't build them. So for your deep tech founders, you have...
为什么会花这么多?光是搞清楚这些业务怎么运作,成本是多少?
And why, why so much? What was the cost just to understand how these businesses work?
因为大家都以为很难的部分,我们很快解决了。而真正难的,是别人毫不知情的那部分——因为他们从没真正做过。想象一下,第一支要去攀登珠峰的队伍,在大本营和一支刚失败的队伍喝茶。失败那队说:‘爬到一半有个地方,难得难以置信,我们过不去。’你们队爬上去了,一路登顶,回来后再和那队喝茶,成功的那队对没成功的那队说:‘中间那一段,其实不是最难的部分。’
Because what everybody thought was hard, we solved quickly. And what nobody else knew about, because they'd never actually done it, turned out to be really hard. Imagine the first group that was going to climb Everest, sitting at base camp having tea with a group that just failed. The group that failed says, 'Halfway up, there's this part. It's unbelievably hard. We couldn't do it.' Your team climbs up, makes it all the way to the top, comes back, and the team that made it leans over to the team that didn't and says, 'That part in the middle wasn't the hard part.'
因为没人越过那些坎。没人越过某些环节,所以他们甚至不知道该怕什么。我们现在知道了。那个环节叫封装——怎么把晶圆固定到主板上,怎么供电,怎么散热。以前没人做过。在那 18 个月里,我们从几乎为零做到了全世界最好。
Because nobody had gotten past it. Nobody had gotten past certain things, so they didn't even know what to be afraid of. We now know. And it was a step called packaging: how you fix a wafer to a motherboard, how you deliver power to it, and how you cool it. Nobody had done it before. And over that 18-month period, we became the best in the world at it from approximately zero.
我们靠的是反复迭代、严谨的工程方法论,以及每次失败都做失效分析。所以我们一次次以不同的方式失败。我们每六周跟董事会开一次会,对他们说:‘这就是战略。以前没人做过,这就是我们要做的事。’我们不得不发明新材料、新技术。最终,我们做出了别人需要靠合作伙伴才能做出来的东西。2019 年 8 月,我们宣布我们解决了。
We did that by iterating again and again, using good engineering methodology, and doing a failure analysis on every single failure. So we failed differently again and again. We told our board every six weeks, 'This is the strategy. Nobody's ever done this before. This is what we're going to do.' We had to invent new materials and new techniques. We ended up building things that everybody else had partners who could do. But in August of 2019, we announced we had solved it.
第一次成功时,我和联合创始人坐在一个很小的实验室里,简直不敢相信。我们是最早尝试把它做出来的几个人。我们就那样盯着它——看服务器运行就像看油漆变干一样无聊。但那一刻,我们五个人就盯着这台机器,不敢相信它竟然真的能跑。我们让它工作了。
My co-founders and I, the first time it worked, were sitting in a tiny little lab, and we couldn't believe it. We were the first few people who had tried to make one work. And we just stared at it. Watching a server run is about as exciting as watching paint dry. And there we were, the five of us, staring at this machine, not believing it might have worked. We made it work.
是啊,嗯。和真正敲钟相比,那是个更大的时刻吗?还是说情绪上完全不同?
Yeah. Um. Was it a bigger moment than actually ringing the bell, or completely different, like emotionally?
情绪上,那是完全不同的一件事。它意味着我们把想法真正做成了。我觉得尤其是深科技创业者,心里总有一个小小的声音在说:‘也许这样做是对的?也许我们其实疯了。’
Emotionally, it was a completely different thing. It was that we had made our idea work. I think for deep tech founders in particular, there's always this little thing in the back of your mind that says, 'Maybe it's right? Maybe we're actually crazy.'
没错。
That's right.
也许它不会失败,也许它不会成功。也许我没时间了。也许我们要没钱了。也许,也许,也许。而另一面,是一种喜悦——因为那是我联合创始人的想法。这些想法在世界上落了地,这是非常了不起的事:当一个人的想法获得物理形态。再下一步,是看着别人的工作建筑在你的想法之上。我们喜欢做基础设施。所以当我们今年 5 月 14 日敲钟上市——成为史上最大的半导体 IPO——我们做了一件不寻常的事:邀请了所有从一开始就跟着我们的工程师——每个一起超过九年的人——和他们的家人,我们一起敲钟。我们一起上台。那是一种不一样的骄傲:我们一起走到了这里,而且没有中途倒下。人们不会告诉你这些。我们犯过很多错误,但避开了致命的错误。我们走到了一个成功的平台,让我们有机会去追求更高一层的成功。这就是 IPO 的意义。它不是终点。你到了一个高原,可以在公开市场追逐新的成功。我承认,那种感觉挺好的。
Maybe it's a no-fail thing, maybe it's not going to work. Maybe I don't have time. Maybe we're going to run out of money. Maybe, maybe, maybe. And the flip side of that is the joy that these were my co-founders' ideas. These are their ideas manifest in the world, and that is extraordinary: when someone's ideas take physical shape. And then the next step is when you watch other people's work sit on top of your idea. We love making infrastructure. So when we rang the bell and went public on May 14th this year, in the largest semiconductor IPO in history... We did something unusual. We invited all our engineers who'd been with us since the start — everyone who'd been with us more than nine years — and their families, and we all rang the bell together. We all got up on stage. That was a moment of a different pride: that we had done this together, and managed not to die along the way. People don't tell you that. We made plenty of mistakes, but we avoided the fatal ones. We made it to a level of success that gave us the opportunity to pursue a new level of success. That's what an IPO is. It's not the end. You've gotten to a plateau from which you can chase a new level of success in the public market. And that felt pretty good, I'll admit.
太棒了,太棒了,感谢分享。我们再深入聊聊产品和技术方面。做更大的芯片有什么取舍?你可能会想到失效模式。
Amazing. Amazing. Thanks for sharing. Let's go a little deeper on some of the product and technical stuff. What are the tradeoffs of building a bigger chip? One that may come to mind is failure mode.
当然。如果你在一大块晶圆上放很多小芯片,就能隔离问题。但如果是一颗大芯片,那所有东西会同时失效。所以你必须非常仔细地考虑失效模式。我们发明了一种技术,大约有一百万个相同的 tile,如果其中一个失效,我们可以关掉它,用冗余的 tile 顶上,继续运行。要往大了做,你必须在计算机的最底层架构里想清楚怎么管理失效。GPU 的失效率很高——我相信你们聊过这个。早期失效率惊人,它们一直在坏。有好数据:Facebook 发过一篇论文,讲他们在大型集群里遇到的故障数量。因为我们在一个点上有这么多算力,我们可以投入更多去散热。所以我们开创了 AI 系统的水冷,运行温度比 GPU 低得多。电子设备的失效模式和温度相关,所以跑得更冷,我们就更可靠。这是一个优势。但同样,我们不得不改进技术,来给这么大的芯片散热。我觉得我们有一种系统性思维:要做大事,就要做取舍,就得想清楚一个允许冗余和修复的架构。你必须在架构的每个方面权衡利弊,这就是做出不同、做出新东西的方法。
Sure. If you have a lot of little chips on a big wafer, you can isolate problems. If you have a big one, then everything fails at the same time. So you have to think very carefully about failure mode. We invented a technique with about a million identical tiles; if one fails, we shut it down and use a redundant one and keep going. If you're going to go big, you have to think, in the very architecture of the computer, how you're going to manage failure. The GPUs have a huge failure rate — I'm sure you guys have spoken about this. Infant mortality is enormous, and they fail all the time. There's good data: Facebook put out a paper on the number of failures they get in a big cluster. Now, because we have all this compute in one spot, we can invest more to cool it. So we pioneered water cooling in AI systems, and we run these much colder than GPUs. The failure mode in electronics is temperature, so running them colder makes us more reliable. That was an advantage. But again, we had to advance the technique to cool off a chip this big. I think we have a systemic mentality: if you're going to do something big, you've made tradeoffs, and you have to think about an architecture that allows for redundancy and repair. You have to think about the pros and cons of every aspect of the architecture, and that's how you do something different and new.
好。为了把这点讲透,确保大家从这段对话里获得明确的结论——请把我当成 15 岁的人来解释吧。为什么这比 GPU 更快?从根本上说,是什么让它更快?
Okay. And again, just to drive the point home, to make sure there's a clear takeaway for people from this conversation: explain it to me like I'm maybe not five but fifteen. Why is this faster than a GPU? What fundamentally makes it faster?
正是因为推理时词元如何生成,它才更快。要生成一个词——我们的答案是一整串词,可能是代码,可能是像素,但都叫它们词——模型的权重要从内存搬到算力单元,进行一次计算,然后生成这个词。
How tokens are generated in inference is why it's faster. To generate a word — and our answers are a whole stream of words; it can be code, it can be pixels, but call them words — the weights of the model are moved from memory to compute, a calculation occurs, and that generates the word.
这是指 prefill 和 decode 的区别吗?
Is a prefill versus decode?
两个步骤都是这样。
That is both steps.
好。你能解释一下什么是 prefill 和 decode 吗?
Okay. And you want to maybe explain what prefill and decode are?
好。做推理所需计算中有两步。当你在 ChatGPT 里输入‘给我讲讲这个村庄在二战前的历史’,然后它给出回答,这时发生了两件事。第一件事是你的提示词被处理了,这是第一步。
Okay. There are two steps in the computation necessary to do inference. When you type into ChatGPT, 'Explain to me the history of this village prior to World War II,' and it comes back to you, two things have happened. The first thing is your prompt has been processed. That's step one.
第二步,你的答案已经生成。生成答案的方式叫做解码(decode)。解码是顺序执行的。处理你的提示词,我们称之为预填充(pre-fill),可以并行化,所以能同时处理很多提示词。但你得到答案的速度取决于解码,而解码是一步一步按顺序来的,这是无法改变的。你执行这一步的方法,是把权重——也就是模型的智能——从内存移到计算单元去生成一个词。你做一次计算,得到一个词,然后这个词又用来生成下一个词,权重再次从内存移到计算。所以这个过程就是把权重从内存搬到计算。那么,像 70B 参数这样的一个小模型,权重有多大呢?权重的大小大约相当于 100 GB。所以生成一个词,你要把 100 GB 从内存搬到计算单元,然后下一个词你还得再搬一遍。你想要这样重复一千次才能得到一个好答案——一个千词答案。这正是 HBM 慢的地方。这个步骤就是 HBM 慢的根源。而恰恰在这个步骤,因为这里全是 SRAM,我们快得惊人。在这里,把权重搬到计算的速度大约比 GPU 快两千五百倍。这就是我们在这里所做的事情的本质,也是为什么我们快这么多。
And step two, your answer has been generated. The way your answer is generated is called decode. Decode is sequential. Processing your prompt, which we call pre-fill, can be parallelized, so you can process many prompts simultaneously. But the speed with which you get an answer is a function of the decode, and it is step by step in sequence, and that can't be changed. How you do that step is you move weights—which are the intelligence of the model—from memory to compute to generate a word. You do a calculation and you get a word, and then that word is used to generate the next word where weights are moved from memory to compute. So the process is one of moving weights from memory to compute. Well, how big are the weights in a little model like a 70-billion-parameter model? The weights are about the size of 100 GB. So to generate a single word, you move 100 GB from memory to compute, and then you have to move them again for the next word. And you want to do this a thousand times to get a good answer—a thousand-word answer. This is where HBM is slow. This exact step is where HBM is slow. And at that exact step, by having all the SRAM here, we're blisteringly fast. The speed of moving weights to compute is about two and a half thousand times faster here than on a GPU. So that's the essence of what we've been able to do here and why it's so much faster.
嗯。
Yeah.
有意思。你觉得模型是否也需要在那个快速 AI 世界中演进,还是那只是一个问题?
Fascinating. Do you think that models need to evolve in that fast-AI world as well, or is that just a problem?
不,恰恰相反。因为记住,有两件事在同时发生。第一,我们希望模型更聪明。第二,模型变聪明的方式之一是强化学习,而强化学习在训练内部使用了推理。你训练内部的推理做得越快,你的训练就越快。
No, exactly. Because remember, two things are happening. One, we want the models to be smarter. And two, one of the ways models are getting smarter is with RL, and RL uses inference inside of training. The faster you can do the inference inside of training, the faster your training.
我很好奇。这对你们来说是一个市场吗?
I'm interested. Is that a market for you guys?
这对我们来说也是一个市场。
That's a market for us as well.
所以你们不只是做推理吧?
So you're not just in inference?
我们也做强化学习和传统训练。不是为最大的实验室,而是为下一梯队的实验室。
We do RL and we do traditional training too. Not for the largest labs, but for the next tier.
嗯。做纯粹的预训练?
Uhhuh. For pure pre-training?
预训练、微调、全流程。
Pre-training, fine-tuning, full set.
那么如果你们做预训练——我指的是带强化学习的后训练,以及为其他实验室做的一些预训练——我想说的是,GPU 在什么工作上仍然比你们强?
So if you do pre-training—I mean post-training with RL, some pre-training for the other labs—what I mean is, GPUs are still better than you at what job?
GPU 在训练中存在一些只有极少数社区成员才能解决的挑战。GPU 是一个非常小的芯片,而训练中需要做的计算却非常庞大。训练中最复杂的部分之一是把计算拆分并分散到多个 GPU 上,这叫做分布式计算。这在历史上是超级计算领域的范畴,而且非常困难,不仅仅因为把一个问题拆开来让很多人同时做很难,而且他们必须不断共享信息。而这种信息共享正是他们需要收购 Mellanox 的原因——他们需要控制一个网络结构,所有的共享为了得到答案都要在这个结构上进行。把一个大的矩阵乘法、大的计算拆分,叫做张量模型并行。你把张量拆开,分散开。世界上最好的实验室擅长这个,但其他人都不会。而我们跑训练时,不需要那样做。我们用的是所谓的模型并行和数据并行。数据并行非常简单。所以它能让那些做得很好的团队快速在训练中测试想法。我们更易用、也更快,因为我们让他们使用一种简单得多的技术。
GPUs in training have some challenges that have been solved by a very narrow selection of the community. GPU is a very small chip, and the calculations that we need to do in training are very large. One of the most complicated parts of training is breaking up the calculations and spreading them apart on multiple GPUs, and that's called distributed compute. That has historically been the domain of the supercomputing world and is very difficult, not just because cracking a problem and having lots of others work on it is hard, but they have to constantly share information. And that sharing information is why they needed to buy Mellanox—they needed to control a fabric over which all this sharing, in order to get an answer, would happen. That breaking up a big matrix multiply, a big calculation, is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And the best labs in the world are good at that, but nobody else is. When we run training, we don't have to run that way. We run what's called model and data parallel. And data parallel is very simple. And so it allows teams who are good, and very good, to quickly test ideas in training. So we are easier to use and faster because we allow them to use a technique which is much simpler.
我们来谈谈智能体世界吧,也许从推理开始。我记得看过一篇博客,你们说推理并不总是适合所有问题的解决方案。你是怎么看的?这又适用于哪里?
Let's talk about the agentic world, maybe starting with reasoning. I think I saw a blog post where you guys stated that reasoning was not always the right solution for all problems. How do you think about this, and where does that fit in?
我们把它当作一种技术使用,就有点像你八年级时写一篇论文的不同草稿。单次推理:你写一个查询,得到一个答案。理解推理最简单的方法是,你会做几份草稿。它会分解问题,分部分解决,把各部分合在一起,检查结果,改进结果,然后给你答案。所以这会消耗更多算力。如果你的计算机慢,那就不只是烦人了——可能会致命。所以最好的模型——无论是国内的还是中国的,无论你是 OpenAI、Anthropic 还是 Gemini,无论是哪个中国模型——它们都转向了推理方法。但那意味着推理期间使用了更多的算力。这对我们来说是个巨大的优势,因为我们更快,而且这会让 GPU 的慢速不断累积,让我们的速度优势变得更加明显。所以这对我们来说是一大福音。我们认为这将继续成为这些模型当前运行方式的基本要素。
We use it as a technique, a little bit like when you were in eighth grade and you wrote different drafts of a paper. Single-shot inference: you write a query, you get an answer. The easiest way to think about reasoning is you're going to do several drafts. It's going to break the problem up, solve it in parts, bring the parts together, review the results, improve the results, and then give you an answer. So that is going to take more compute. And if your computer is slow, that's going to be more than an irritant—it could be crippling. And so, as the best models—all of them, whether they're domestic or Chinese, whether you're OpenAI or Anthropic or Gemini, whether you're any of the Chinese models—they all move to a reasoning approach. But that meant more compute was being used during inference. That was a huge advantage for us because we were faster, and it made the GPU slowness stack up, and it made our speed have an even bigger advantage. So this was a huge boon for us. We think this is going to stay as a fundamental essence of the way these models are run right now.
你还写了关于验证的文章——智能体中的瓶颈到底是推理够不够好,它们够不够聪明,还是存在验证问题。
And you also wrote about verification—whether the bottleneck in agents was whether the reasoning was good enough, whether they were smart enough, or whether there was a verification problem.
当然。我认为验证问题有点像护栏问题。你想做的是,在写完两三份答案草稿之后,你要确保它不是错的。这就是你的验证步骤。你可以用另一个模型来做。你可以用不同的方式向现有模型问一个类似的问题。这些都可以给你的结果做压力测试。护栏也可以这样工作。你要么用模型,要么用另一种技术通过评分机制来审查,确保这个问题没有越界,确保这个答案不是关于如何制造生物武器,或者调用信息指示你向 FBI 举报。所有这些都需要耗费算力时间。所以无论你是想通过推理来改进,还是想运行护栏,只要更快,你就能在更短的时间内得到结果。
Sure. I think the verification problem is a little bit like the guardrail problem. What you'd like to do is after you have written two or three drafts of your answer, you would like to be sure that it wasn't wrong. And that's your verification step. You can do that maybe with a different model. You can do that by asking your existing model a similar question in a different way. All of these are ways you can pressure-test your result. Guardrails can work the same way. You want to review, either with a model or with another technique through a scoring mechanism, that this question isn't out of line, that this answer isn't about how to make biological weapons or calling upon information that directs you to the FBI. All of that takes compute time. So whether you're trying to improve through reasoning or whether you're running guardrails, by being faster, you can get results in less time.
嘿,你觉得这个世界将如何为智能体进化?是一堆更小的模型跑得更快、做更多验证,还是一个大型模型?
Hey, so what do you think the world is evolving for agents? Is it a bunch of smaller models running faster, doing more verification, versus a large model?
嗯,我觉得它们是协同工作的。大模型给出一个答案,然后你只需再核对一下数据。
Well, I think those work together. I think your big model produces an answer, and then you just want to double-check your data.
我的意思是,在新闻行业,人们会写一些有数据核查员的报道,对吧?会有人去核实——过去,记者会核查数据,担任数据核查员,对吧?那是一个职业。每一条说法都会被独立核查。
I mean, in the journalism industry, people would write papers that had data checkers, right? Somebody would go and make sure—back then, a journalist would check data, be a data checker, right? That was a job. Each claim was checked independently.
那是另一种模式,对吧?主模型写了文章,然后一个小模型检查了部分答案。我觉得这是一种非常好的做法。
That's a different model, right? The main model wrote the piece, and then a little model checked some of the answers. And I think that's a very, very good way to go about it.
对于更大的模型来说,多模态在你这里处于什么位置?
Where does multimodality for the larger models fall in your world?
嗯,我们刚刚宣布,我们在谷歌的一个多模态模型上做到了全世界最快。我认为事实是,几乎没有文字是不带图表和图形的,对吧?你必须理解文字,还得能理解插图、图形、图表——所以这算是第一步,也是最容易的部分。然后你应该能同时生成这两者。接下来你应该能理解图像。我觉得新模型在这方面显然非常非常擅长。紧随其后的是视频,因为视频不过是一组图像的集合。但现在视频需要海量算力,这也是它被领先实验室暂时搁置的原因之一——它实在太吃算力了。
Well, we just announced that we were the fastest in the world on one of Google's multimodal models. I think the truth is, there's very little text that doesn't have charts and graphs, right? You must understand text, be able to understand illustrations, graphs, charts—so that's sort of the first and easiest part. And then you ought to be able to create both. And then you ought to understand images. And I think the new models are very, very good at that, obviously. What follows that is video, because a video is just a collection of images. But that takes an enormous amount of compute right now, and that's one of the reasons it's been sort of set aside by the leading labs—being so unbelievably compute-intensive.
太好了。我们稍微聊聊服务器业务吧,因为你们显然是一家芯片制造商/供应商,正如我们之前谈到的。你们也是云服务商、数据中心提供商。那么这几个部分是怎么分工的?
Great. Let's talk about the server business a little bit, because you guys are obviously a chip maker/provider, as we discussed. You're also a cloud provider, data center provider. So what are the different parts?
我们制造算力,而且这些算力是为 AI 优化的——是世界上在 AI 方面最快的。如果你有数据中心,我们会把硬件卖给你,部署在你的数据中心里。如果你没有数据中心,想按月或按年租用,我们有数据中心。所以你可以通过我们的数据中心和云来租用我们的设备。
We make compute, and that compute is optimized for AI—it's the fastest at AI in the world. If you have a data center, we will sell you hardware for deployment in your data center. If you don't have a data center and you'd like to rent it by the month or by the year, we have data centers. So you can rent our equipment through our data centers and through our cloud.
这样一来,我们就能服务 AI 原生公司,以及大型企业和政府。
And so that allows us to get AI natives as well as large enterprises and governments.
那比例是什么——
And what's the—
去年是 50/50。我觉得上个季度大概是——也许是 75/25,偏向硬件销售。我觉得今年可能会回到 50/50。随着我们与 OpenAI 的协议继续推进,到 2026 年大概会变成 30/70——30% 是本地部署的硬件,70% 是云。
It was 50/50 last year. And I think this last quarter it was—or maybe 75/25 in favor of hardware sales. I think this year it might be 50/50. And as our OpenAI deal continues, in 2026 it will probably be 30/70—with 30% on-premise deployments of hardware and 70% cloud.
太好了。我们来谈谈那个 OpenAI 协议——这是一个重大的历史性里程碑,破纪录的。所以它提供高达 750 兆瓦……
Great. Let's talk about that OpenAI deal—it's such a major historical milestone, record-making. So it's providing up to 750 megawatts...
兆瓦——顺便说一句,这是一个有趣的度量单位,因为我们是芯片供应商,但这是电力。所以这是……简写吗?是一种简写。我的意思是,现在事实证明——我们之前没谈过这个——在相邻的供应链里,我们经历了内存短缺,也经历了名为 CoWoS 3 纳米产能的工艺短缺。目前我们行业面临的另一个限制就是数据中心的可用性。我是说,这对所有人来说都是一个限制因素。所以 Anthropic 跟 Elon 做了一笔巨额的、非常昂贵的交易,来获取数据中心容量。我们与 OpenAI 的协议也是因为数据中心容量是一个限制性约束,而衡量数据中心的方式就是兆瓦。这笔交易是 750 兆瓦——2026 年 250 兆瓦,多年租约;2027 年再增加 250 兆瓦,多年租约;2028 年再增加 250 兆瓦,多年租约。
Megawatts—which is interesting by the way as a metric because we're a chip provider, but this is power. So is that shorthand for...? It's a shorthand. I mean, it turns out right now—and we didn't talk about this—there's sort of an adjacent supply chain. We went through a shortage of memory, we went through a shortage of a process called CoWoS 3-nanometer capacity. The other limitation in our industry right now is data center availability. And I mean, that is a limiting factor for everybody. That's why Anthropic did a huge—which is a sort of very expensive deal with Elon—for data center capacity. Our deal with OpenAI was because data center capacity is a limiting constraint, measured the way data centers are measured—in megawatts. The deal is 750 megawatts—250 megawatts in '26 on a multi-year lease, an additional 250 megawatts in '27 on a multi-year lease, and an additional 250 megawatts in '28 on a multi-year lease.
那你是为他们建数据中心,还是提供数据中心里的芯片?
And are you doing the data center for them, or are you providing the chips that go in the data center?
为他们提供数据中心。我们交付的是一个完整的云解决方案,所以他们通过 API 连接到我们。
Data center for them. We're delivering a full cloud solution, so they connect to us via an API.
基本上是这样。
Basically.
设想 2026 年——我马上就——那是……
Imagine 2026—do I immediately—that's...
我正在找数据中心——我的下一个会议就是和数据中心供应商谈。我们现在在欧洲有很多业务。
I am looking for data centers—my next meeting is with a data center provider. We're doing a lot in Europe right now.
哦,有意思。
Oh, interesting.
很多在北欧。
A lot in the Nordics.
是因为那里更靠近能源吗?
Is that because it's closer to power sources?
是的,因为那里有低成本的电力——而且是清洁的低成本电力。
Yes, it's because there's low-cost power—and clean, low-cost power.
是啊,低成本。
Yeah, low cost.
就像在时间业务里那样,数据中心的位置在……方面重要吗?
Does it matter—just like in the time business—where the data center is located in terms of...
在成本方面。
In terms of spend.
是的,还有一种额外的延迟叫传输延迟。那就是光通过光纤从赫尔辛基到纽约的速度,如果你的客户在纽约,而你的数据中心在赫尔辛基,你就必须考虑这一点。好了,我的意思是,通常大约是光速的三分之二——想想看要花多长时间。你会希望数据中心在同类平台上。
Yeah, there is an additional latency called transport latency. That's the speed of light through fiber to get from Helsinki to New York, and you have to account for that if your customers are in New York and your data center is in Helsinki. Okay, I mean, it's usually about two-thirds the speed of light—curious how long it takes. You would like data centers on the same kind of platforms.
那就是 OpenAI 交易。还有一笔和 AWS 的令人兴奋的交易——你们有一个联合设计的芯片解决方案。
That's the OpenAI deal. There was an exciting deal with AWS as well—where you have a co-designed chip solution.
没错。这就是你之前提到的分离式解决方案,他们的 Trainium 部分负责 prefill——也就是可并行化的步骤。所以 Trainium 做 prefill 步骤,我们的芯片做 decode。这样你就能得到多得多的 token,而且它们流式输出很快。
That's right. It's the disaggregated solution you mentioned before, where their Trainium part is doing the prefill—the parallelizable step. So Trainium is doing the prefill step, and our chip is doing the decode. So you get a lot more tokens, and they stream fast.
对我们来说是一笔好交易。它用的是他们的数据中心。所以这些是部署在 AWS 数据中心里的。
It's a good deal for us. It uses their data centers. So these are deployments in the AWS data center.
是啊,这很吸引人,对吧?谈到这个行业,灵活性太重要了。就像交付解决方案——你需要芯片,就能拿到芯片;你需要数据中心,每个人都在向不同的供应商购买,以减少依赖。
Yeah, it's fascinating, right? Talking about this industry is about how flexibility is so important. Like delivering the solution—you need chips, you get chips; you need data centers, everybody's buying from different suppliers to reduce dependency.
我认为这就是我们从传统的芯片/系统供应商转变为也提供数据中心的原因之一——因为客户想要的是快速的 token,而我们能做任何让快速 token 的交付变得更简单的事。对一些人来说,那是在他们的数据中心里;对另一些人来说,那是通过 API。只要把你的流量指向我们,我们就会把快速 token 的水龙头对准你。
I think that's one of the reasons why we went from being a traditional chip/system provider to also offering data centers—because what our customers want is fast tokens, and anything we can do to make the delivery of fast tokens easier. For some of them, that's in their data center. For some of them, it's with an API. Just point your traffic to us, and we'll point the fire hose with fast tokens at the back.
是啊。只是想了解一下你们从哪里开始,到哪里为止——或者不做,或者至少目前提供的是云版本。所以如果我想运行 Kimmy——我知道你们在速度方面对 Kiwi 和 Gemma 有惊人的数据,值得一提——但你们不把这些作为服务提供,还是说你们也提供?
Yeah. And just to get a sense for where you start and where you stop—or do not, or at least currently provide the cloud version. So if I want to run Kimmy—I know you have incredible stats for Kiwi and Gemma in the speed field, to mention them—but you don't run those as a service, or do you?
我们也提供。
You do.
好的。所以你们有一项服务,和基线以及 Fireworks 竞争。
Okay. So you have a service competing with the baselines and Fireworks.
是的,我想我们有一个按需服务,你可以来我们的网站按月订购。我想你甚至可以为 Kimmy 或 GLM 或其中一些模型购买 token 包。我们的很多客户来到这里,感到兴奋,然后转向专用方案——他们会租用数百台机器,使用一年、两年、三年或四年,一旦他们验证了这对他们工作的好处。他们经常做 A/B 测试。这并不奇怪——人们喜欢更快的。
Yeah, I think we have an on-demand service where you can come to our site and book a month. I think you can even buy buckets of tokens for Kimmy or GLM or some of these models. Many of our customers come there, get excited about it, and then move to a dedicated offering where they take hundreds of machines for a year, two, three, or four, once they've proven out the benefit for their work. Often they do A/B tests. Not surprising—people like faster.
好的,所以它更像是一个试验场。
Okay, so it's more like a testing [ground].
这是一个完整的环境。你可以去用用看。网址是 three.ai。随便玩。
It's a full environment. You can go and use it. It's at three.ai. Play around.
但这很迷人。但这可能会变成又一个大云业务。好,所以你有芯片,你有数据中心,而且你还有一个云业务跑在前面。真是令人着迷。
But fascinating. But that could become like yet another big cloud business. Okay. So you have chips, you have data centers, and you have a cloud business running in front of them. It's fascinating.
嗯,想想护城河。你知道,著名的是,英伟达有 CUDA 这个经常被讨论的护城河。你的等价物是什么?
Um, thinking about moats. You know, famously, Nvidia has CUDA as a well-discussed moat. What's your equivalent?
嗯,我不认为我们必须……是的,我们应该谈谈这个。
Well, I don't think we must... Yeah, we should talk about that.
嗯。
Yep.
我觉得两年前,每个最先进的模型都是在 CUDA 环境下训练的。现在,Gemini 在没有 CUDA 的情况下训练。Anthropic Claude 在没有 CUDA 的情况下训练。OpenAI 是在有 CUDA 的情况下训练的。所以在一两年内,他们失去了 70% 的训练模型份额。护城河仍然存在,但数据显示护城河正在明显缩小。在推理方面没有护城河。
I think two years ago, every state-of-the-art model was trained in a CUDA flow. Right now, Gemini is trained without CUDA. Anthropic Claude is trained without CUDA. OpenAI is trained with CUDA. So in a one or two year period, they lost 70% share of training models. The moat is still present, but the data shows the moat is clearly shrinking. There's no moat in inference.
呃,从 GPU 转移到我们在云端的服务,只需要八次按键。
Uh, it takes you eight keystrokes to move from a GPU to us in the cloud.
就是八次。
It is eight.
哦,就这些?把你的流量从 GPU API 转移到我们这里。所以很明显,CUDA 在我们行业的创立中起到了巨大作用,它让这里的图形处理比单纯的图形处理更加通用。但从 2023 到 2024 年,我认为它作为持久护城河的能力已经……而且你在思考你的护城河时,有一个完整的生态系统战略。在这个行业里,如果存在任何护城河,你建立了一个完整的生态系统。是吗……?
Oh, that's it? To move your traffic from GPU API to us. And so obviously, CUDA was enormously important in the creation of our industry and in allowing the graphics processing here to be more general than graphics processing. But since 2023-2024, I think its ability to serve as a durable moat is... And you have a whole ecosystem strategy as you think about your moat. To the extent that any moat can be present in this industry, you built a whole ecosystem. Is that...?
是的。
Yeah.
是的,我们已经建立了一个生态系统。我觉得我们的护城河来自于我们的架构,我们正在做别人做不到的事情。并不是说他们能花更多钱,或者他们不能用 GPU 为这个付更多钱。如果你想要快,你是得不到的。嗯,我的意思是,它就是不行。所以这就是我们建立优势的地方,也是我们为客户创造价值的方式。
Yeah, we've built an ecosystem. I think our moat comes from the fact that by virtue of our architecture, we are doing things no one else can do. And it's not that they can spend more money, or they can't pay more for this with GPUs. If you want fast, you can't have it. Well, I mean, it just doesn't work. And so that's where we're building our strength, and that's how we're delivering value to customers.
你对供应链怎么看?我们提到了别人的供应链限制,但你的供应链限制是什么?你全用台积电吗?
What do you think about supply chain? We mentioned supply chain constraints for others, but what are your supply chain constraints? Are you all TSMC?
我们全用台积电。我们有过非常紧密的合作。他们是我们的投资者。他们一直是出色的合作伙伴。我告诉你一个不寻常的故事。2017 年,我们出现的时候总共大约 30 人。我们八月份去的。台湾那种可怕的天气。不要在八月份去台湾。你知道,那里很残酷。
We are all TSMC. We had a very close collaboration. They were investors in us. They've been exceptional partners. I'll tell you an unusual story. In 2017, we showed up as a company of about 30 guys total. We showed up in August. Horrible trend in Taiwan. Don't go to Taiwan in August. You know, it is brutal.
虽然这里今天也不怎么好。
Not that it's so nice here today.
这里只有 90 华氏度。
It's only 90 here.
湿度。
Humidity.
是的。但湿度很大……我们见了台积电的领导层。我们说:‘我们是一家不起眼的小公司。我们相信我们能解决历史上没人解决的问题,下面是我们如何改变你们制造芯片的方式来实现这一点。’他们想了想,然后说:‘我们同意,就这么办。’
Yeah. But all the humidity... We met with the leadership of TSMC. We said, 'We are a little pipsqueak company. We believe we can solve a problem that nobody solved in history, and here's how we would modify the way you make chips to make this possible.' They thought about it and they said, 'We agree. Let's do it.'
在会议上?
In the meeting?
就在会议上。
In the meeting.
在会议上,而不是‘一个月后再来’,就在会议上。
In the meeting, it wasn't 'go away for a month' in the meeting.
那是因为他们早有准备,还是他们反应特别快?
Was that because they had a prepared mind, or were they just exceptionally fast on their feet?
嗯,首先,销售人员已经把决策者都召集来了。其次,我们的提案非常擅长让他们利用自己擅长的东西。这不需要他们做巨大的改变,但确实需要他们做出真正的改变。我觉得他们认为这足够大胆,可以在做中学。他们也知道 AI 在大芯片上表现更好。所以这是开放心态、承担风险的意愿,以及一家非常大的公司的大胆思考的结合。
Um, first, the salesperson had gathered the decision makers. Second, our proposal was really good at allowing them to use what they were good at. It didn't require them to change a huge amount, but it did require them to make real changes. I think they saw this as sufficiently bold that they would learn as they did it. They also knew that AI was better on big chips. So a combination of fair mind, a willingness to take risk, and bold thinking from a very large company.
太迷人了。
Fascinating.
确实迷人。我是说,这就是大公司取胜的方式,对吧?那有多罕见?那是非同寻常的。
It is fascinating. I mean, that's how big companies win, right? And how rare is that? It was extraordinary.
那接下来发生了什么?比如,在那样一个似乎特别快的会议决策后,到……需要多长时间?
And what happened next? Like, how long does it take between a decision in a meeting like that, which seems exceptionally fast, to...
两年。芯片制造是一个漫长而艰难的过程。大多数时候,你的第一颗芯片不会成功。现在有很多初创公司,其中一些有非常聪明的人,但他们的第一颗芯片不会……TPU 呢?谷歌有业内一些最优秀的人才。第一颗芯片没成,第二颗、第三颗也没成。第四颗非常好。现在他们做到第八颗了,那是一颗非常好的芯片。好吗?AWS 的 Annapurna 团队。第一颗芯片不怎么样。第二颗、第三颗就非常好了。
Two years. Chipmaking is a long, hard process. And most of the time, your first chip isn't a winner. There are lots of startups now, some of them with really smart guys, but their first chip will not be... The TPU? Google had some of the best guys in the industry. First chip wasn't a winner, nor the second, nor the third. Fourth was really good. Now they're on their eighth, and it's a really good chip. Okay? The Annapurna team at AWS. First chip wasn't great. Second, third chip really good.
这需要时间。所以我们造了一颗芯片,交付了。而这涉及到你之前问的一个问题。你知道,我们解决了一个计算史上没人解决的问题。我们 2020 年交付了,但没人关心。没人关心。没人买,也没人关心。就像:‘天哪,每个人都说我们疯了。它永远不会成功。现在它成功了,却没人想要。’所以我们就造了下一颗。
It takes time. And so we built a chip, we delivered it. And this gets to an earlier question you asked. You know, we solved a problem that nobody in the history of computing solved. And we delivered it in 2020, and nobody cared. Nobody cared. Nobody bought any, and nobody cared. It was like, 'Oh god, everybody said we were crazy. It would never work. And now it works, and nobody wants it.' So then we built the next one.
我是说,第一颗,我们可能卖得挺……
I mean, the first one, we probably sold quite a...
没人想要它,是因为市场还没准备好,还是因为产品不够好?
And nobody wanted it because the market was not ready, or because the product was not good enough?
嗯,没人想要它,是因为当时 AI 还是一种爱好。如果你的爱好真的很快,谁在乎呢?当它进入生产环境时,你才会在乎速度。当你每天使用时,你才会在乎速度。
Um, nobody wanted it because AI was a hobby at that time. And who cares if your hobby is really fast? You care about fast when it's in production. You care about fast when you use it every day.
所以我们又造了一颗,你知道,那一颗我们卖了三五百个。
And so we built another one, and you know, that one we sold three or five hundred.
嗯。
Yeah.
然后我们造了第三颗,卖了几万个。
And we built the third one, and we sold tens of thousands.
太惊人了。
Amazing.
是啊。是不是很有意思?
Yeah. Isn't that interesting?
嗯,所以回到供应链。你需要考虑本土化多元化吗?
Um, so going back to supply chain. Do you need to think about onshore diversification?
所以很难从台积电多元化出去。芯片非常难。实际上,当你设计一颗芯片时,设计的一部分要符合那家工厂的规则,对吧?所以你无法把你的设计从台积电带到别家去,因为有大量的工作是为了确保你的设计符合他们的规则。因此,即使在历史上,我想只有一两个例外,每一代芯片都只去一家晶圆厂。所以我们下一代也会继续和台积电合作。
So it's very hard to diversify away from TSMC. Chips are so hard. And actually, when you design a chip, part of the design is for the rules of that factory, right? So you can't take your design from TSMC and go to somebody else because a huge amount of the work was to be sure your design is within their rules. And so even in, I think, only with one or two exceptions in history, each chip generation goes to one fab. So we're going to be with TSMC for our next generation as well.
我觉得有些……呃,我们的供应链由很多部分组成,但我们把芯片从台积电运回美国。我们在美国封装,在美国组装。我们在美国进行制造,然后从美国发货。我觉得当你像我们现在这样快速增长时,会有一系列普通的供应链挑战。你知道,某个供应商搞砸了一批货,在海关卡住了。供应链中可能出错的方式多到难以置信。但我们每天都在管理这些,而且我们的制造吞吐量正在呈指数级增长,所以这部分业务我们确实做得很好。
I think some... Uh, we have a supply chain that is built in many parts, but we bring the chips back from TSMC to the US. We package in the US, and we assemble in the US. We do our manufacturing in the US, and then we ship from the US. I think when you're growing this fast, as we are right now, there are a range of garden-variety supply chain challenges. You know, a vendor screws up a batch, it gets stuck in customs. The number of ways that things can go wrong in the supply chain is unbelievable. But we manage these every day, and we're increasing our manufacturing throughput exponentially, so we're really doing that part of the business well.
太不可思议了。
Incredible.
那作为最后一个问题,我们不妨把视野拉远一些——你觉得这一切会走向何方?显然,没人能料到未来几年 AI 会如何发展,但就未来一两年来说呢?
So maybe to zoom out as the last question—what's your best guess about where all of this is going? Obviously, who knows in AI in the next few years, but in the next year or two?
嗯,有些事情我们是确定的。我们现在用的模型——你今天正在用的,也许是 GPT-5 或 GPT-6——将会是你用过的最差的模型。无论你现在觉得它哪里酷,六个月后它都会变得无聊又落后。而这正是让人兴奋的地方。我观察我们年轻工程师使用它的方式,和我自己很不一样。这是一个有意思的时刻,你能从年轻的团队成员身上学到东西——他们使用 AI 的方式完全不同。
Well, we know some things. We know that the model we use—you're using it today, maybe GPT-5 or GPT-6—will be the worst model you ever use. And whatever you think is cool about it right now is going to be boring and backwards in six months. And that is so exciting. I watch the way our young engineers use it, and it's very different from the way I'm using it. It's a fun time where you can learn from your young team members—they're using AI very differently.
我认为仪表盘这类业务——以及它给 SaaS 造成的冲击——已经不可修复了。你知道,你可以让你的 AI 说‘给我做一个像 Salesforce 一样的工具’。三十秒后你就得到一个能用的工具,简直不可思议。而所有那些过去因为跨越组织内部孤岛而难以做到的事情,对吧?
I think the business of dashboarding—and the damage it's doing to SaaS—is, I think, unrepairable. You know, you could ask your AI, 'Build me a tool like Salesforce.' Thirty seconds later you have a working tool that is just unbelievable. And all the things that were difficult because they cut across your internal organizational silos, right?
大概有五套系统。你可能在用 Workday,还有股票计划系统,还有……
There are like five systems. You're in Workday, you're in your stock plan, you're in...
对。
Yeah.
还会用到 Carta。这些系统之间根本无法互通。而 CEO 想知道的正是这个。
You're in Carta. None of them can talk to each other. And that's what a CEO wants.
我的核心员工还有多少未归属的股票?
How much holding power for my top guys?
对啊。
Yeah.
我以前为此写过一些小工具。然后‘砰’的一下,我就做出了一个随手搞定的小应用。
And I used to have little tools I wrote for this. And boom, I've got a little app that I had put together in a snap.
你看吧。
You go.
哇,真是个好故事。听你讲这些真是太不可思议了。何等的一段旅程,何等激动人心的未来。非常感谢你,我学到很多,这次访谈太棒了。谢谢你,Andrew。
Well, what a story. It's just incredible to hear all of this from you. What a journey, and what an exciting future. So, thank you very much. I learned a lot, and this was terrific. Thank you, Andrew.
谢谢你邀请我来节目,我真的很感激。
Thank you for having me on your show. I really appreciate it.
大家好,我是 Matt Turk。感谢收听本期 Mad Podcast。如果你喜欢这期节目,我们将非常感激你能考虑订阅——如果还没订阅的话——或者在你观看或收听本期节目的平台上留下好评或评论。这真的能帮助我们把这个播客做起来,并请到更多出色的嘉宾。谢谢,我们下期再见。
Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks, and see you at the next episode.