The Future of AI Hardware and Architecture with Tri Dao
打开互动全文版(中英对照 + 朗读 + 问答)→Flash Attention 主要作者、Together 首席科学家 Tri Dao 探讨 AI 硬件的未来、NVIDIA 的竞争对手以及模型架构的演变。
Tri Dao, lead author of Flash Attention and chief scientist at Together, discusses the future of AI hardware, competitors to NVIDIA, and the evolution of model architectures.
我觉得当我开始这个播客时,你就是我的梦想嘉宾名单上的人。是什么推动了成本降低和延迟改善方面的这么多进步?
I feel like when I started this podcast, you were on my dream guest list. What's driven so much of the improvements on the kind of lowering cost and improving latency side?
部分原因我认为是更好的架构。模型模仿人类的能力好得离谱。我们刚刚看到那里技术的爆发。
Part of it is I think better architecture. It's kind of crazy how good the models are at acting like human. We've just been seeing an explosion of techniques there.
我们会看到 NVIDIA 生态系统不同部分的竞争对手吗?
Are we going to see competitors to different parts of the NVIDIA ecosystem?
我认为肯定有很多竞争对手试图进入这个领域。
I think certainly there's lots of competitors are trying to enter the space.
那个你最想知道的突出问题是哪个,或者最会影响你工作方向的是哪个?未来几年我想回答的问题是,Freda 在过去几年推动了 AI 中许多最重要的突破。他是 Flash Attention 的主要作者,这是模型推理成本大幅下降的关键原因。他通过 Mamba(Transformer 的替代架构)的工作深度参与了空间模型研究。他是 Together 的首席科学家,他的工作影响巨大,以至于 Semi Analysis 指出 NVIDIA 模式很大一部分是因为 Tree 在 NVIDIA 生态系统中做了所有这些工作。无监督学习能有机会与像 Tree 这样的人坐下来,问他们最关心的问题,包括两三年内有多少工作负载会运行在 NVIDIA 芯片上,这真是莫大的荣幸。我们讨论了 AI 硬件的未来以及机会在哪里。我们讨论了 AI 编码产品如何使 Tree 的工作效率提高了 50%,这远高于我的预期。我们还讨论了他对 Transformer 替代架构未来的看法,以及什么可能真正有效。与这个领域如此杰出的人交谈是一次迷人的机会。我真的很享受这次对话,我想大家也会喜欢。闲话少说,有请 Tree。
What's that question that's outstanding that you're most curious to know or would most impact the direction of your work? The question I'm trying to answer for the next couple years is Freda has driven many of the most important breakthroughs in AI over the past years. He was a lead author of flash attention which has been a key reason why model inference costs have gone down so much. He was super involved in space model work via his work at Mamba the alternative architectures to transformers. He's the chief scientist at together and his work's been so impactful that semi analysis basically noted that a huge part of the NVIDIA mode is the fact that Tree is doing all this work in the NVIDIA ecosystem. It's a real privilege of unsupervised learning to get to sit down with someone like tree and just ask them anything that's top of mind including the percent of workloads that are going to run on Nvidia chips in two three years. We talked about the future of AI hardware and where the opportunities are. We talked about how AI coding products have made tree 50% more efficient in his work uh which was way higher than I'd expected. Uh and we also talked about you know his thoughts on the future of alternative architectures to transformers and what might really work. Just a fascinating opportunity to talk with someone who's so brilliant in this field. really really enjoyed this conversation and I think folks will too. Without further ado, here's Tree.
Tree,非常感谢你来做客播客。真的很感激。
Well, Tree, thanks so much for coming on the podcast. Really appreciate it.
是的,当然。我很兴奋。
Yeah. Yeah, sure. For sure. I'm really excited about this.
我觉得当我开始这个播客时,你知道,两三年半前,你就在我的梦想嘉宾名单上,所以我想那是在 Flash Attention 最初流传的时候,人们对你在该领域的影响赞叹不已,我就想,如果我们能随着时间的推移把节目做起来,也许我们能说服 Tree 来上节目。所以,
I feel like when I started this podcast, you know, two and a half, three years ago, you were on my like dream guest list and so I think that was like when kind of Flash Attention, you know, was initially making the rounds and people were just ooing and aing over the impact you were having on the space and I was like, if we really can build this up over time, maybe we'll convince Tree to come on. And so,
是的。我很荣幸来到这里。这应该会很有趣。
Yeah. Yeah. I mean, I'm very honored to be here. I mean, this should be really fun. Yeah.
所以,我想我们会涉及很多不同的话题。嗯,你知道,我认为这些会对我们的听众感兴趣。我想一个起点就是一个经典的 VC 问题。所以,你得原谅我。NVIDIA 在过去几年的故事,你知道,令人难以置信地令人印象深刻,达到了他们现在的规模。显然,你为那个生态系统做出了巨大贡献,让硬件越来越好用。我们会在芯片方面、系统方面、封装 GPU 方面看到 NVIDIA 生态系统的不同部分的竞争对手吗?显然,谷歌和亚马逊有自己的芯片,但对于行业中的绝大多数其他公司,我确定你经常思考这个问题。
So, I think we'll hit on a bunch of different things. Um, you know, that I think will be of interest to our listeners. I figured one place to start would just be like a classic VC question. So, you'll have to forgive me. The NVIDIA story over the past years has just been, you know, unbelievably impressive to get to the scale they're at. Um, and obviously you've, you know, contributed a tremendous amount to that ecosystem and making, uh, you know, that hardware, you know, better and better for folks to use. Are we going to see competitors to different parts of the Nvidia ecosystem on the chip side on the kind of systems the package GPUs obviously you've got you know Google and Amazon with their own chips but uh you know for the for the vast majority of the rest of the industry uh I'm sure this is something you think about a lot
是的,我花了很多时间思考芯片。我认为肯定有很多竞争对手试图进入这个领域。我的意思是 AMD 已经存在一段时间了。显然 NVIDIA 占主导地位有几个原因:他们设计非常好的芯片,并且构建非常好的软件,这创建了一个人们在其上构建的生态系统。但我认为随着工作负载开始在架构方面(比如 Transformer 等)趋于一致,我们肯定会看到很多竞争对手进入这个领域。为那种工作负载设计芯片变得稍微容易一些。所以如果你只关注推理,我认为 AMD 有某些优势。你知道,他们有更大的内存等等。所以我们开始看到人们采用它。在训练方面,这有点困难。我的意思是,网络是主要瓶颈之一,NVIDIA 当然领先,但我认为人们理解构建好的训练芯片的挑战是什么,构建好的推理芯片的挑战是什么。所以这归结于执行。所以我会说这个领域实际上非常令人兴奋。我和很多设计新芯片的人聊过,无论是推理还是训练。所以我预计在未来几年,一些工作负载可能会变成多硅片。它们可能会运行在不同的芯片上,而不是现在,我不知道,90% 在 NVIDIA 上。
Yeah, I spend a fair amount of time thinking about the chips. I think certainly there's lots of competitors trying to enter this space. I mean AMD has been here for a while. Obviously Nvidia is dominant for a couple reasons: they make very good chips and they build very good software, and that creates this ecosystem where people build on that. But I think certainly we'll see a lot of competitors entering the space as the workload starts to coalesce around the architecture side, you know, transformer and so on. It becomes a little bit easier to design chips for that kind of workload. So if you just focus on inference, I think AMD has certain advantages. You know, they have larger memory and so on. So we're starting to see folks picking that up. On the training side, it's a little bit more difficult. I mean, networking is one of the main bottlenecks and NVIDIA is certainly ahead there, but I think people understand what are the challenges to build good training chips, what are the challenges to build good inference chips. So it comes down to execution. So I would say this space is actually really exciting. I talked to lots of folks who are designing new chips either for inference or training. So I would expect in the next couple years maybe some of the workload will become multi-silicon. They'll probably run on different chips rather than right now I'd say I don't know 90% on Nvidia.
你认为架构足够稳定吗?因为显然我认为芯片设计的一个迷人之处在于你基本上是在对未来两三年推理工作负载和训练工作负载的样子下注。所以显然你对此越确定,就越容易进行优化。感觉有足够的稳定性来下这些赌注吗,还是说基本上有几十家初创公司和公司在押注,其中一两个可能会成功?
Do you think the architectures are like stable enough? Because obviously I think one of the fascinating parts of chip design is you're basically making a bet two three years in the future of what you know inference workloads and training workloads will look like. Um and so obviously the more certainty you have around that the easier it is to make optimizations. Does it feel like there's enough stability to make those bets or is the idea that basically look there's dozens of startups that are and companies that are making bets and one or two of them may work out
对。对。是的。所以我认为在架构方面,从非常高的层面看,它似乎已经围绕 Transformer 稳定下来了。但一旦你更仔细地看,实际上有很多变化。最近几年最新的一个是混合专家模型。你让模型变得更大,参数更多,但更稀疏。这就有某些权衡:你需要更多内存,但也许你做的计算量稍微小一些。所以即使这样也给一些芯片制造商带来了困难,因为他们可能是在为密集模型设计,计算非常均匀,而现在你有了稀疏的东西,所以设计起来更难一些。或者一些围绕注意力机制的变化,注意力机制从 20 年前就存在了,你知道,超过 10 年了,但这里那里的变化实际上让一些东西变得困难。
Right. Right. Yeah. So I think on the architecture side from a very high level it looks like it is sort of stabilized around transformer. But once you look a little bit closer there's actually quite a bit of change around that. So the most recent one in the last couple years has been mixture of experts. So where you make the model much larger, much more number of parameters, but sparser. And so that has certain trade-offs: you need more memory but maybe the amount of compute you do is a little bit smaller. And so even that presents difficulty for some of the chipmakers because maybe they were designing for a dense model where it's very uniform compute and now you have something sparse, so it's a little bit harder to design for. Or some of the stuff changing around attention has been around since 20, you know, for 10 plus years, but there are changes here and there that actually make some of this stuff difficult.
所以 DeepSeek 有这种多头潜在注意力机制,看起来和普通注意力有点不同。比如,它们的头维度非常大,如果你的脉动阵列(矩阵乘法引擎)有特定大小,可能就不太匹配。一旦你深入看,就会发现诸如此类的奇怪问题。所以我想说,这是架构方面的一个点。然后在工作负载方面,我认为人们正在用这些模型处理非常不同类型的任务。我们还有传统的聊天机器人吗?过去两年里传统到什么程度?
So DeepSeek has this thing, multi-head latent attention, that looks a little bit different than attention. For example, they have very large head dimension, and if your systolic array, the matrix multiplication engine, has a certain size, maybe it doesn't fit right. Weird things like that once you look underneath the cover. So I would say that's one on the architecture side. And then on the workload side, I think people are using these models for very different kinds of workloads. Do we have the traditional chatbot? How traditional has been in the last two years?
大概三年了。这是传统 AI 世界,对吧?但出现了新的工作负载,比如编码任务,像 Cursor 和 Windsurf 等等。还有一种智能体工作负载,不仅仅是运行模型,还包括调用工具、运行 Python 解释器、进行网络搜索等等。这给芯片设计带来了挑战。如果芯片过于专注于尽可能快地运行模型,你可能会忽略诸如如何连接到主机以运行网络搜索之类的问题。所以我想说,尽管从高层看架构已经趋于稳定,但底层还有很多变化,部分原因是工作负载也在变化。所以这是一场关于你能多快执行并适应这些新工作负载的持续竞赛。
It's been around 3 years. It's traditional AI world, right? But there are new workloads like coding workloads, such as Cursor and Windsurf, and so on. There's a kind of agent workload where it's not so much just running the model but also making tool calls, running Python interpreter, doing web search, and so on. That presents challenges of how you design the chip. If the chip is very much focused on just running the model as fast as possible, you might neglect things like how you connect to the host to run a web search. So I would say even though from a high level the architecture has kind of stabilized, underneath the cover there's still a lot going on, and part of it is the workload is also changing. So it's a constant race of how fast can you execute and adapt to these new workloads.
那么,如果今天 90% 的工作负载都在 NVIDIA 芯片上,你认为两三年后我们会处于什么位置?
So if 90% of these workloads are on NVIDIA chips today, where do you think we are in two or three years?
是的,我认为在推理方面,部分会多样化。我们开始看到像 Cerebras、Groq 和 SambaNova 这样的公司带来了真正的挑战。他们宣称‘我们可以实现极低延迟的推理’,这对某些用例来说非常棒。我们和一些客户聊过,他们非常关心尽可能低的延迟,并愿意为此支付更多。所以我认为我们开始看到更多这样的细分领域,有些人非常关心低延迟,有些人则非常关心大批量高吞吐的推理,比如大规模数据处理或合成数据生成,或者强化学习训练,你需要生成尽可能多的轨迹。所以我认为它会多样化,仅仅因为工作负载会多样化,无论是低延迟、高吞吐还是其他。也许视频生成会需要不同的算力和内存配置。
Yeah, I think on the inference side, some of it will diversify. We're starting to see companies like Cerebras, Groq, and SambaNova really presenting a serious challenge. They were pitching, 'Hey, we can get very very low latency inference,' and that's fantastic for some use cases. We talked to customers who really care about as low latency as possible and are willing to pay more for that. So I think we're starting to see more of these niche areas where people really care about low latency, or some people really care about very large batch high throughput inference, like massive data processing or synthetic data generation, or RL training where you have to roll out and generate as many trajectories as possible. So I would say it's going to diversify simply because the workloads are going to diversify, whether low latency or high throughput or something else. Maybe video generation will require a different profile in terms of how much compute and memory you need.
如果你今天是一家初创公司,当下似乎很清楚:嘿,针对所有这些用例,人们可能想要各种不同的优化,然后你决定,嘿,我们要制造一款针对这些优化的芯片,然后你等待流片和实际生产能力,等这一切完成时,已经是几年后了。你认为这些不同的优化方向是否足够清晰,以至于你可以做出前瞻性的押注,还是存在这样的风险:实际上两三年后,如果我们再做这期节目,你会说,实际上重要的优化是这五件我们没谈到的事?
If you were a startup today, it kind of feels clear in the moment, hey, here's all the different optimizations one might want to make given all these use cases, and then you kind of decide, hey, we're going to make a chip that optimizes for these things, and you wait for tape out and actual ability to produce, and by the time you've done that, it's a few years down the line. Do you think that it's clear enough what the different optimizations are going to be such that you can kind of make that forward-looking bet, or is there such a risk that actually two or three years from now, if we're doing this episode, you're like, well actually the optimizations that matter are these five things that we didn't talk about?
对,所以我想说,如果你是一家初创公司,你必须下注。你投资,你知道这一点,而且你必须下重注。你可以押注聊天机器人可能会消失,人们真正关心的是,我不知道,视频生成模型、世界模型或机器人技术之类的,然后你下注并说,好吧,这可能会占工作负载的 50%,我们该如何为这种工作负载设计?然后你希望你的赌注是正确的。我的意思是,这就是初创公司的角色,对吧?我认为这可能是你产生影响的方式。我认为如果你不下注,只是说我要优化通用工作负载,那么现有巨头就会在执行上超越你。
Right, so I would say if you're a startup, you have to make a bet. You invest, you know this, and you have to make an outsized bet. So you could bet that maybe chatbots will go away and the thing people really care about is, I don't know, video generation model or world model or robotics or something, and you make that bet and say, okay, that's probably going to be, I don't know, 50% of the workload, how would we design for that workload? And you hope that your bet will be correct. I mean, that's the role of a startup, right? And I think that's probably the way you could make an impact. I think if you just don't place a bet and say, I'm just going to optimize for the general workload, then the incumbents are going to just out-execute you.
我很兴奋我们今天做这期节目,因为昨天我刷 Twitter 时看到 SemiAnalysis 发了一条很有趣的推文,回应你最近的一些工作。显然,你在 NVIDIA 芯片上优化模型方面做了很多工作,我觉得人们可以尝试量化你为整个 NVIDIA 生态系统增加的价值,但这显然是巨大的。而且我认为 SemiAnalysis 的人说,‘天哪,你为这个生态系统增加了这么多价值。为什么 AMD 或谷歌或其他有芯片的公司不付钱让你也在那些领域做类似的优化?’这让我想起了大语言模型世界,那里有巨额合同给少数真正能构建和训练这些模型的研究人员。我想知道你对那条推文的反应,以及你对硬件领域类似动态的看法。
I was excited we're doing this episode today because yesterday I was scrolling Twitter and I saw this really interesting tweet from SemiAnalysis reacting to some of your recent work. Obviously you've done so much on optimizing models on NVIDIA chips, and I feel like one could try and quantify the value you've added to the broader NVIDIA ecosystem, but it's clearly massive. And I think the SemiAnalysis folks were like, 'God, you've added so much value to this ecosystem. Why wouldn't AMD or Google or some of the other folks that have chips try to pay you to get you to do similar optimizations in those spaces?' And it reminded me of the LLM world where you've seen massive packages for the few researchers that are really able to build and train these models. I wonder what your reaction was to that tweet and what you think about similar dynamics in the hardware space.
是的。我个人与来自 NVIDIA、AMD、谷歌、亚马逊等公司的很多人合作。我花了很多时间在 NVIDIA 芯片上,仅仅因为那是我们现在拥有的,也是大多数人使用的。我认为它们设计得非常好,而且他们为此设计了非常好的软件。这让我能够做一些有趣的事情。这就是我所追求的:你能做有趣的事情吗?例如,我们与 AMD 的人合作过,他们有一个 FlashAttention 的版本,我们与他们合作将其集成到公共仓库中等等。所以我们确实与其中一些公司合作过。我不知道最好的安排是什么,但现在我更多地在思考我们需要什么样的抽象层,不仅针对 NVIDIA 芯片,而是针对 GPU 和加速器整体。我花了很多时间在底层,试图从这些芯片中获得最大性能。但随着我们在 Together 的规模扩大,我们思考如何让我们雇佣的其他人变得高效,其中一部分就是构建能在 NVIDIA 芯片以及可能其他芯片上工作的抽象层。
Yeah. I personally work with lots of folks from NVIDIA, AMD, Google, Amazon, and so on. I spend a lot of my time on NVIDIA chips simply because that's what we have right now, that's what most people use. I think they are very well-designed chips and they design very good software for that. So that allows me to do interesting things. That's kind of what I'm after: can you do interesting things? We've worked with the folks at AMD, for example, they had a version of FlashAttention, and we worked with them to integrate that into the public repo and so on. So certainly we work with some of them. I don't know what the best arrangement would be, but nowadays I'm thinking a lot more about what are the abstractions we need, not just for NVIDIA chips but for GPUs and accelerators in general. I spend a bunch of time at the lowest level trying to get the maximum performance out of these chips. But as we scale at Together, we think about how to get other people we hire to be productive, and part of it is building abstractions that will work on NVIDIA chips and potentially other chips as well.
另一件让我兴奋的事情是,我们如何构建抽象层,让 AI 能为我们分担一些工作。答案还不明确,但作为人类和技术领导者,关键在于构建正确的抽象,让其他人能快速上手,让你所做的工作能跨不同芯片或不同工作负载运行。你认为我们现在有能够跨不同芯片工作的抽象层吗?
The other thing I'm excited about is how do we build abstractions so that AI can do some of this work for us. The answer isn't quite clear yet, but as humans and technical leaders, it's about building the right abstraction so that other people can get onboarded really quickly, and so that the things you do could work across different kinds of chips or different workloads. Do you think we're at a place where there are abstractions that allow you to work across the different set of chips out there?
我认为有一些,但这是常见的权衡。Triton 非常好,支持 Nvidia、AMD、Intel GPU 等。这需要他们设计前端和后端,不同公司为后端贡献代码。我认为 Triton 相当不错,很多公司都在押注它。例如,Meta 使用 PyTorch 编译生成 Triton 代码,然后依赖 Triton 为 Nvidia 或 AMD 生成代码。但这是常见的权衡:如果你不控制最底层,可能会牺牲一些性能。关键在于权衡点在哪里。如果牺牲 5% 的性能但效率提升 3 倍,那是可以接受的。但如果牺牲太多性能,人们可能会选择更专门化的方案。
I think we have some, but it's a usual trade-off. Triton is really good and supports Nvidia, AMD, Intel GPUs, and so on. That requires them to design a front end and a back end, and different companies contribute code to the back end. I think Triton is quite good, and a lot of companies are betting on it. For example, Meta uses PyTorch to compile and generate Triton code, then rely on Triton to generate code for Nvidia or AMD. But it's the usual trade-off: if you don't control at the lowest level, maybe you give up some performance. It's about where that trade-off point is. If you give up 5% performance and become 3x more productive, that's an acceptable trade-off. But if you give up too much performance, then maybe people would go with something a bit more specialized.
尤其是在竞争激烈的推理市场。
Especially in a pretty competitive inference market.
没错。所以我要说,为人类设计真的很难。我认为硬件可移植性有点神话色彩。即使是 Nvidia 芯片,代际之间变化也很大。这是我们让这些芯片获得更高性能的唯一方式。不像 CPU,年复一年可能快 5-10%,旧代码还能用。即使是 Nvidia,每一代他们基本上都得重写所有代码。因为获得更多算力的方法是构建更专门化的组件,可能更低精度,可能不同的芯片间同步方式。所以 Nvidia 也大约每两年重写一次软件。甚至代际之间的可移植性都不存在。所以这些抽象层即使只是为了帮助同一制造商后续每一代芯片,也很有价值。
Right. So I'd say designing for humans is really hard. I would say hardware portability is kind of a myth. Even for Nvidia chips, generation to generation they change a lot. That's the only way we get more performance out of these chips. It's not like a CPU where year to year you get 5-10% faster and old code just works. Even for Nvidia, every generation they essentially have to rewrite all the code. Because the way to get more flops is to build more specialized components, maybe lower precision, maybe a different way to synchronize between different parts of the chip. So Nvidia also essentially rewrites their software every two years or so. Even portability between generations isn't there. So those abstractions would be valuable even just to help with each subsequent generation of chips from the same manufacturer.
对。
Right.
是的,所以我认为 Triton 有非常吸引人的抽象层。他们甚至有一些更低层次的东西,比如 Gluon,它暴露更多硬件特性,但代价是通用性降低。Modular 的人一直在构建 Mojo 语言。你怎么看他们在做的事?
Yeah, so I think Triton has very compelling abstractions. And they even have something a bit lower level, like Gluon, which exposes more of the hardware at the cost of being less general. The folks at Modular have been building this language, Mojo. What do you think of what they're doing?
我觉得非常酷。我认为他们有一些正确的抽象。关键在于执行。人们会看它并问:“你在 Nvidia 芯片上有多快?”这在某种程度上是不公平的问题,但这就是人们关心的。所以他们设计抽象层,然后必须做一些定制工作,让代码在 Nvidia 芯片上运行得非常好,再为 AMD 芯片做一点定制。这取决于你愿意做多少定制来权衡性能。
I think it's super cool. I think they have some of the right abstractions. It's just about execution. People are going to look at it and ask, 'How fast are you on Nvidia chips?' which is an unfair question in some sense, but that's what people care about. So they would design the abstractions and then have to do some custom work to make their code run really well on Nvidia chips, and then a bit of custom work for AMD chips. It's just about how much customization you want to do for trading off performance.
你看,市场要求 Nvidia 芯片上最好的推理。我们看到越来越多的这类库或领域特定语言。斯坦福的人有 ThunderKittens,试图抽象一些东西。还有 TinyGrad。Google 有 Mosaic GPU。我肯定还漏了一些。但人们意识到一个问题:我们还没有正确的抽象层。训练新工程师编写高性能 GPU 内核实际上很痛苦。所以答案是构建抽象层。我认为我们正处于快速迭代的阶段。这就是为什么我们看到这么多领域特定语言涌现。随着 AI 模型变得更好,我在思考的一件事是如何为语言模型设计领域特定语言或抽象层,因为它们与人类操作方式有些不同,我们不知道正确答案。所以我认为在未来一两年内,情况会变得更加清晰。现在,每个人都在尝试许多不同的方向。
See, the market demands the best inference on Nvidia chips. And we're seeing more and more of these libraries or domain-specific languages. The folks at Stanford have ThunderKittens, trying to abstract some of this. There's TinyGrad. Google has Mosaic GPU. I'm sure I'm forgetting some. But people realize there's a problem: we don't have the right abstractions yet. It's actually kind of painful to train new engineers to write very performant GPU kernels. So the answer is to build abstractions. And I think we're very much in the phase where we're iterating pretty quickly. That's why we see so many domain-specific languages coming out. And as AI models get better, one thing I'm thinking about is how to design domain-specific languages or abstractions for language models, because they operate somewhat differently from humans, and we don't know the right answer. So I would say in the next one or two years, it's going to become a lot clearer. Right now, everybody is trying lots of different directions.
你认为这些抽象层最可能来自哪里?
Where do you think those abstractions are most likely to come from?
我认为人们从两个角度入手。一个是从机器学习角度:他们考虑我们有什么样的工作负载,以及表达这些工作负载需要哪些原语。例如,推理很大程度上关乎如何尽可能快地移动内存,因为推理通常是内存受限的,或者如何尽可能快地做矩阵乘法,以及在那里表达什么原语。另一个角度是从硬件角度:他们在芯片上有非常酷的专门化组件,并思考如何通过抽象层来暴露它们。Nvidia 非常擅长的一件事是设计更异步的芯片,因为矩阵乘法变得如此之快,以至于其他一切都慢得多。所以重叠矩阵乘法和其他计算变得更重要。那么如何设计允许这种异步执行的抽象层,比如流水线和同步?所以我认为抽象层将来自工作负载侧或硬件侧。我认为可能在一两年内会变得更加清晰。
I think people approach it from two angles. One is from the machine learning side: they think about what kind of workload we have and what primitives are necessary to express those workloads. For example, inference is very much about how to move memory as fast as possible, since inference is usually memory-bound, or how to do matrix multiply as fast as possible and what primitives to express there. The other angle is from the hardware side: they have really cool specialized components on the chips and are thinking about abstractions to expose them. One thing Nvidia has been really good at is designing chips that are more asynchronous, because matrix multiplication is getting so fast that everything else becomes much slower. So it's much more important to overlap the computation of matrix multiplication and everything else. So how do you design abstractions that allow this kind of asynchronous execution, like pipelines and synchronizations? So I'd say that's where the abstractions will come from, either from the workload side or from the hardware side. I think it will become a lot clearer in maybe one or two years.
你之前几次提到过,所以我必须问一下:你谈到让这些抽象适合大语言模型,以及在这些过程中使用 AI 本身。你在多大程度上使用 AI 来确定这些东西?你觉得未来几年这会如何变化?
Well you've alluded to a few times so I have to ask you you talked about kind of making these abstractions suitable for LLMs maybe as well as kind of just using AI in these processes themselves. How much are you using AI itself in determining this stuff and how do you kind of see that changing over the next few years?
是的。我认为模型开始对这类事情变得有用了。实际上是非常最近的事。这让我很惊讶。一方面,人们追求的是全自动的 GPU 内核编写,对吧?你描述问题,语言模型就会为你生成内核。这有点像我们在其他领域已经达到的水平,比如简单的 Python 脚本、数据分析或 Web 前端。你可以做到。那么我们能对 GPU 做到这一点吗?
Yeah. I think models are starting to become useful for this stuff. Actually very recently. This really surprised me. On one level, what people are after is fully automatic GPU kernel writing, right? You describe the problem and the LM will just generate the kernel for you. This is kind of maybe we're there with some of the other areas like simple Python script or data analysis or web front end. You can do this right. So could we do that for GPU?
就是写一些内核。
Just write some kernels.
对,对。我想说,如果你想要的是那个,那我认为我们还非常非常早期。这些模型可以生成一些简单的内核,比如逐元素操作:你输入一个数组,对每个元素做某种操作,对吧?或者做某种归约,比如求和和一些归一化之类的事情。所以模型可以做得相当不错。但一旦变得稍微复杂一点,这些模型就无法生成正确的代码了。我认为这只是因为我们没有足够的训练数据。训练数据对这类东西来说非常困难,因为如果你从互联网上抓取内核代码,你会得到课程项目,你会得到可能是三代前 GPU 的文档,那些是你现在不应该做的事情。所以训练数据非常困难。所以我认为答案是,你可能需要从一些专家级数据开始,然后从中生成合成数据,或者连接到一个像编译器和性能分析器这样的工具,这样你就可以获得大量的训练数据,或者获得正确的环境。我认为这会在一年或两年内解决,但这肯定是一个难题。
Right, right. I would say if that's what you want, then I think we're very, very early. So these models can generate some simple kernels, like element-wise: you take in an array and you do some operation on each element, right? Or you do some kind of reduction like summing and some normalization things like that. So the models can do a reasonably good job. But once it gets a little bit more complicated, these models just don't generate correct code. I think it's just a function of we don't have enough training data. Training data is really tough for this stuff because if you scrape the internet for kernel code, you'll get class projects, you'll get documentations that were meant for GPUs maybe three generations ago that are things you shouldn't be doing now. So training data is really difficult. So I think the answer to that is you probably have to start with some expert-level data and then generate synthetic data out of that, or hook up to a tool like a compiler and profiler so that you can get lots of training data, or getting the right environment. I think it will be solved in a year or two, but it's certainly a difficult problem.
谁有那些数据?
Who has that data?
我不认为这类数据是私有的。有几个地方有专家级代码,但我认为更多的是工作流程:如何从少量专家级数据开始,生成大量合成数据?对吧?所以 Discord GPU 模式的一些人一直在努力做这件事。他们采用了 PyTorch 编译器,它从 PyTorch 代码生成 Triton 代码,这是一种更低级的内核代码,他们能够生成,我想,15000 对这样的 PyTorch 和 Triton 程序。你必须有点创意。我认为互联网上没有那么多数据。所以你必须有点创意,关于如何生成这类数据。所以我认为,如果你想要全自动的内核生成,这是一个角度。我说我们在这方面非常早期。另一方面是:它们能否与人类一起工作?模型能否与人类一起工作?我惊喜地发现这些模型实际上非常有用。
I don't think this kind of data is private. There are a couple of places where you have expert-level code, but I think it's more about the workflow: how do you start from a small amount of expert-level data and generate lots of synthetic data? Right? So some of the folks at the Discord GPU mode have been really trying to do this. They took the PyTorch compiler, which generates from PyTorch code to Triton code, which is this lower-level kernel code, and they could generate, I think, 15,000 pairs of these programs of PyTorch and Triton. You have to be a little bit creative. I think there's just not that much data on the internet. So you have to be a little bit creative about how you're going to generate this kind of data. So I'd say that's one angle if you want fully automatic kernel generation. I say we're super early there. The other side is: can they work alongside humans? Can models work alongside humans? And I've been pleasantly surprised that these models are actually quite useful.
有没有一个特定的时刻,你心想:‘哇,这些模型已经足够好,对我很有帮助了。’
Was there a specific moment where you were like, 'Wow, these models have gotten good enough to be quite helpful to me.'
是的。我想说可能有两个最近的里程碑。一个是 o3。o3 在推理方面变得非常好。我和 o3 以及 GPT-5 头脑风暴的一些问题是:‘嘿,我有这个函数,我该如何优化它?我需要注意哪些事情?’它在高层次上出奇地好。另一个是 Claude Code。不知何故,它在编写 Triton 内核方面实际上相当不错,这太棒了。尽管我喜欢编写内核,但我的很多时间都在思考设计,思考我们应该设计什么样的架构才能充分利用硬件。实现部分——设计真的很有趣,但实现通常很繁重。而 Claude Code 在这里被证明非常有帮助。我想说它让我生产力提高了大约 1.5 倍。
Yeah. I would say maybe there were two recent milestones. One was o3. o3's gotten really good at reasoning. Some of the questions I was brainstorming with o3 and GPT-5 were like, 'Hey, I have this function, how would I optimize it? What are the things I would pay attention to?' And it's surprisingly good at the high level. The other one is Claude Code. Somehow it's actually pretty decent at writing Triton kernels, which is fantastic. As much as I love writing kernels, a lot of my time is thinking about the design, thinking about what kind of architecture we should design so that we can take advantage of the hardware. The implementation part of it is—design is really fun but the implementation is usually quite heavy. And Claude Code turned out to be quite helpful here. I would say it makes me maybe 1.5x more productive.
哇。那真是相当不错。
Wow. I mean that's pretty good.
是的。所以我一直是 Claude Code 的重度用户。如果你让这些模型与人类一起工作,也许它们会更有帮助,而不是仅仅依赖它们全自动生成内核。
Yeah. So I've been a heavy user of Claude Code. If you have these models working alongside humans, maybe they are a lot more helpful rather than just relying on them fully automatically generating kernels.
你在等待的下一个里程碑是什么?我猜当新模型出来时,有没有一些你测试的东西,你会想:‘天哪,如果模型能达到这个阶段,我就能从 1.5 倍提升到 2 倍。’
What are the next milestones you're waiting for? I guess when a new model comes out, are there things you test, you're like, 'God, if the models could just get to this stage, I'd go from 1.5x to 2x.'
对。所以我认为 Claude Code 是一个很好的例子,它是一个阶跃变化,它更加智能体化,对吧?而且他们通过后训练让 Claude 在这方面做得非常好,我相信其他人,OpenAI 和 Google,很快也会达到非常相似的水平。这里的智能体化只是意味着它能很好地使用工具,并且知道何时使用工具。所以它知道:‘嘿,我可能没有正确的 API,那么我该如何查找 API?或者程序无法编译,或者程序不够快,我该如何从性能分析器获取信息?’诸如此类。所以我认为对于新模型,我会看它们如何知道自己不知道,对吧?比如它们何时需要寻找新信息?这有点模糊。我认为人们开始为这种智能体能力推出基准测试,但我们在这方面还非常早期。
Right. So I would say Claude Code is a good example of a kind of step change where it's a lot more agentic, right? And somehow they post-trained Claude to do really well there, and I'm sure other folks, OpenAI and Google, will soon get to a very similar point. Agentic here just means it can use tools really well and it knows when to use the tools. So it knows that, 'Hey, I'm probably not having the right API, so how do I look up the API? Or the program is not compiling, or the program is not fast, how do I get information from the profiler?' That sort of thing. So I think for new models, I would see how well they know whether they don't know, right? Like when do they need to seek out new information? That's kind of a vague thing. I think people are starting to come out with benchmarks for this kind of agentic capabilities, but we're super early there.
自从 ChatGPT 发布以及这些模型被广泛使用以来,你一直是推理市场众多改进的核心部分。所以我想,也许对我们的听众来说,最好能梳理一下过去三年,是什么推动了成本降低和延迟改善方面的巨大进步。而且我想确保给你一个机会谈谈 Flash Attention 的工作。
You've been such a central part of so many of the improvements in the inference market since the launch of ChatGPT and the broader usage of a lot of these models. So I figured, maybe just for our listeners, it would be great to contextualize these last three years and what's driven so much of the improvements on the kind of lowering cost and improving latency side. And I want to make sure to give you an opportunity to talk about a lot of the Flash Attention work as well.
当然。当然。是的。我认为在过去几年里,推理成本可能下降了大约 100 倍。
For sure. For sure. Yeah. I think in the last couple years, inference cost has probably come down maybe 100x.
至少,对吧?自从 ChatGPT 问世以来,我认为这很可能也反映在 API 成本上。一方面,在建模方面,人们用相同数量的参数训练出了好得多的模型。部分原因是数据更多,部分原因是架构更好。当然,新的高效注意力机制也有帮助。另一方面,推理优化方面也涌现了大量技术。早期我们并不清楚推理的瓶颈在哪里。后来人们意识到,很大程度上是数据移动的问题:将权重移入和移出内存,移动 KV 缓存——注意力操作中存储历史信息以进行下一个预测的部分。所以很多优化都围绕减少数据移动展开,比如量化模型。两三年前,每个参数 16 位是常态;现在 8 位很常见,一些新模型用 4 位。甚至还有用 1 位或 2 位的工作。这种权衡很有吸引力:量化通常不会损失质量,这要归功于复杂的技术。例如,OpenAI 的 GPT-4o 版本大部分层用 4 位——总共 1200 亿参数,大约只需 60GB。这带来了非常好的推理性能。
At least, right? Since ChatGPT debuted, I think it's probably reflected in the API cost as well. So on one side, on the modeling side, people have trained much better models for the same number of parameters. Part of it is much more data, part of it is better architecture. Certainly, new efficient attention mechanisms help. On the other side, there's inference optimization, and we've seen an explosion of techniques there. In the early days, we didn't understand the bottlenecks of inference. After a while, people realized it's a lot about data movement: moving the weights in and out of memory, moving the KV cache—the part of the attention operation that stores the history to make the next prediction. So many optimizations focus on reducing data movement, like quantizing the model. Two or three years ago, 16 bits per parameter was typical; now 8 bits is common, and some new models use 4 bits. There's even work with 1 or 2 bits. This trade-off is compelling: quantizing often doesn't lose quality, thanks to sophisticated techniques. For example, OpenAI's GPT-4o release uses 4 bits for most layers—120 billion parameters total, fitting in about 60 GB. That translates to really good inference performance.
量化是一种技术。另一种是模型架构与硬件的协同设计。双方沟通更多了。Flash Attention 就是一个例子:我们意识到内存访问是主要瓶颈,于是重写了注意力算法以减少内存访问。这在推理中很常见。比如 DeepSeek 的多头潜在注意力,通过将 KV 缓存投影到更小的空间来压缩它,使其小得多。这对 DeepSeek 非常有效——他们可以非常高效地服务模型。另一个例子是混合专家模型:每个 token 并不使用所有参数,存在稀疏性。过去几年,模型变得越来越稀疏。Mistral 的第一个开源模型每层激活 8 个专家中的 2 个——比例 25%。现在 DeepSeek 和 OpenAI 的 GPT-4o 激活 128 个中的 4 个——比例是 32 倍。这对服务大量用户非常有利。人们正在理解工作负载,并与推理协同设计架构。很多改进都来自这里。
Quantization is one technique. Another is the co-design of model architecture and hardware. People on both sides are talking more. Flash Attention is an example: we realized memory access was the main bottleneck and rewrote the attention algorithm to reduce memory access. We're seeing that a lot for inference. For instance, DeepSeek's Multi-Head Latent Attention compresses the KV cache by projecting into a smaller space, making it much smaller. That worked really well for DeepSeek—they could serve the model very efficiently. Another example is mixture of experts: for each token, you don't use all parameters; there's sparsity. In the last couple years, models have become sparser. Mistral's first open-source model activated 2 out of 8 experts per layer—a 25% ratio. Now DeepSeek and OpenAI's GPT-4o activate 4 out of 128—a factor of 32. That's great for serving many users. People are understanding the workload and co-designing the architecture with inference. That's where a lot of improvement has come from.
展望未来,你认为持续改进会来自哪里?按照趋势外推,应该还有 10 倍左右的提升。我只是不知道我们是否已经摘完了低垂的果实,还是仍有更多空间。
Looking forward, where do you think continued improvement will come from? There's got to be another 10x or so, extrapolating the trend. I just didn't know if we've picked off the low-hanging fruit or if there's still more.
我认为我们已经摘了很多低垂的果实,但还有很多工作要做。一个领域是硬件。有一段时间,你无法预测两年后的工作负载,所以无法很好地专业化。但随着架构趋于稳定,芯片设计者正在针对推理进行优化:低精度得到良好的原生硬件支持,网络对于更大的混合专家模型变得重要。硬件可能一年内就能带来 2-3 倍的提升。在软件方面,进一步推动模型架构。
I think we've picked off a lot of low-hanging fruit, but there's still lots to do. One area is hardware. For a while, you couldn't predict the workload two years out, so you couldn't specialize well. But as architectures stabilize, chip designers are optimizing for inference: low precision with good native hardware support, and networking becomes important for larger models with mixture of experts. Hardware might give 2-3x in just one year. On the software side, pushing on model architecture further.
所以我研究过像 Mamba 这样的东西。这是一种方法,不是将整个历史存储为 KV 缓存,而是让模型将历史压缩成一个更小的状态向量。当然这会有一些权衡,但我们看到了相当不错的成功。我认为这对于大规模批量推理尤其重要,这是目前一些工作负载所要求的,比如测试时计算和推理,你真正想要的是模型同时探索大量轨迹或思维链。我认为谷歌提出了 Gemini Deep Think,就是这个想法。他们赢得了 IMO 金牌,同时探索多条路径。这意味着你同时在数百个序列上进行模型推理,而 KV 缓存就成了一个更大的问题。所以你可以尝试让模型压缩 KV 缓存;这就是我们使用 Mamba 的方向,并且已经有很多其他变体沿着这个思路出现。所以我认为在模型方面,会有 2-3 倍的提升。
So I've worked on things like Mamba. It's a way to instead of storing the entire history as a KV cache, you could have the model compress that history into a smaller state vector. Of course that's going to have certain trade-offs, but we've seen pretty good success with this. I think this is going to be especially important for large batch inference, which is what some of the workloads are demanding these days, where things like test-time compute and reasoning, what you really want is the model to explore lots of trajectories or chains of thought at the same time. I think Google came up with this thing, Gemini Deep Think, which was that idea. They were winning IMO gold medals, exploring multiple paths at the same time. That means you're doing model inference on hundreds of sequences at the same time, and there the KV cache becomes a much bigger problem. So you could try to get the model to compress that KV cache; that was the direction we were going with Mamba, and there have been a bunch of other variants coming out along that idea. So I think on the model side, it's going to be 2-3x.
是的。
Yeah.
对。然后在内核层面,很多人对内核方面非常感兴趣,其中一些非常有才华。所以我们得到了非常好的内核,这大概又是 2 倍的提升。所以综合来看,我认为即使仅仅一年内,我们可能又会得到 10 倍的提升。
Right. And then on the kernel level, lots of people are getting really interested in the kernel side, and some of them are very talented. So we're getting very good kernels, and that's probably another 2x or so. So taken together, I think even in just one year, we would probably get another 10x.
让我印象深刻的是,随着用例的广泛多样化,以及人们想要如何运行推理和这些模型的架构,似乎已经扩展到许多不同的最终用例,每个用例的瓶颈或优化都不同。那么这意味着什么呢?作为推理提供商,你希望在所有方面都做到最好。你认为生态系统会保持在一个供应商擅长所有事情,还是你会看到随着时间的推移出现专业化,比如如果你试图运行深度研究,有一个推理提供商在这方面很出色,而如果你试图运行更多的聊天机器人用例,你认为这会如何发展?
One thing I'm struck by is with this broader diversifying of use cases and how people want to run inference and the architecture of these models, it seems like it's expanded into many different end use cases, and the bottleneck or optimizations for each are different. So what are the implications of that? As an inference provider, you want to be the best at all those things. Do you think the ecosystem remains with one vendor who is really good at all those things, or do you see specialization over time, like if you're trying to run deep research, there's one inference provider that's amazing at that, and if you're trying to run more of a chatbot use case, how do you see that playing out?
对,是的。我认为可能会有三种工作负载模式,我认为所有推理提供商都会理解这一点并为此优化,但大规模运行有某些优势。所以在工作负载方面,有传统的聊天机器人,你需要一定的交互性。你不能太慢,但也不需要超级快。
Right, yeah. I think there will probably be maybe three kinds of workload patterns, and I think all inference providers will understand this and optimize for it, but there are certain advantages to running at scale. So on the workload side, there's the traditional chatbot, where you need a mix of interactivity. You can't be too slow, but you also don't need to be super fast.
是的。
Yeah.
有时如果回复得太快会有点 creepy;你希望体验上感觉至少有人在另一端思考一下。
Sometimes it's a little creepy if it comes right away; you want the experience to feel like there's someone at least doing a little bit of thinking on the other end.
对,对。
Right, right.
所以有那种工作负载。还有那种非常低延迟的需求,比如像 Claude Code 这样的东西。我总是希望如果这能快 3 倍或 5 倍,我当然愿意付更多钱。
So there's that kind of workload. There's the kind of really low latency requirements, like things like Claude Code. I always wish if this is like 3x faster or 5x faster, I'm certainly willing to pay more.
是的。
Yeah.
对。所以我们会看到更多这样的情况,如果推理速度快 2 倍,使用模型的人的生产力就会提高 2 倍。这就是保持心流状态和不保持的区别。
Right. And so we'll see more of that where if the inference is 2x faster, the person working with the model is going to be 2x more productive. It's the difference between staying in flow state and not.
对,对。我的意思是人们通过同时运行多个这样的模型来解决这个问题,比如同时运行四个 Claude Code。但对于我个人更喜欢深度工作的人来说,我通常只用一个,我伴侣因此对我大喊大叫。她说你应该同时使用四个 Claude Code。但对于这种工作负载,人们可能愿意为这种低延迟付更多钱。
Right, right. I mean people are getting around this problem by running multiple of these models, so having like four Claude Codes running at the same time. But for someone who personally prefers deep work, I usually just use one, which my partner yells at me for. She's like you should be using four Claude Codes at the same time. But for this kind of workload, people probably will be willing to pay more for this low latency.
所以有低延迟的智能体式工作负载。然后还有非常大的批次,我不太关心延迟,只想要尽可能高的吞吐量。这对于生成合成数据之类的事情很重要。正如我提到的,现在人们训练模型的方式是,他们有一小部分专家级数据或人工标注数据。假设你是一家航空公司,试图构建一个 AI 智能体来解决客户投诉。他们有一小部分非常好的数据,然后你可以从中生成大量合成数据。模型在模仿人类方面有多好,这有点疯狂。你可以说,‘嘿,你能扮演一个来自纽约的客户,因为航班从拉瓜迪亚延误而恼火吗?’模型在扮演人类方面出奇地好。
So there's the low latency, agentic kind of workload. And then there is the very large batch, I don't care so much about latency but I just want as high throughput as possible. This is important for things like generating synthetic data. As I mentioned, a lot of the way people are training models now is that they have a small amount of expert-level data or human annotation data. Let's say you're an airline and you're trying to build an AI agent to resolve customer complaints. They have a small amount of really good data, and then you can generate tons of synthetic data out of that. It's kind of crazy how good the models are at acting like humans. You can say, 'Hey, can you act as if you're this kind of customer from New York who is annoyed that their flight is delayed out of LaGuardia.' And models are surprisingly good at acting as humans.
是的。我担心我们教会了他们成为机场愤怒的纽约人的能力。这可不是人类最好的一面。
Yeah. I'm scared that we've taught them the capabilities to be angry New Yorkers at an airport. That's not like humanity at its finest.
是的。但不知何故,互联网上有大量这样的数据,模型已经学会了,然后模型可以利用这些数据。它们内部有这种世界模型,然后可以生成大量数据,可能不如人类数据好,但你可以生成海量数据。对于那种推理用例,你只关心吞吐量。另一个是用于强化学习训练用例,一方面你在训练一个智能体做某事,它改变策略,但训练循环的一部分是,当模型有某个策略时,你需要评估该策略有多好。所以假设我在训练一个 AI 工程师,我如何知道当前的 AI 工程师有多好?这需要我从模型中采样,称为 rollout。
Yeah. But somehow there's lots of that data on the internet where the model has learned that, and the model can then use that. They have this kind of world model internally that they can then generate lots of data, probably not as good as human data, but you can generate a massive amount. For that kind of inference use case, you really only care about throughput. The other one is for RL training use case, where on one hand you're training an agent to do something, it changes policy, but part of the training loop is that as the model has some policy, you need to evaluate how good that policy is. So let's say I'm training an AI engineer, how would I know how good the current AI engineer is? That requires me to sample from the model, called rollout.
你从模型中采样大量补全,然后评估其质量。作为训练的一部分,你确实需要非常大的批次和高吞吐量的推理。所以我会说这是第三个用例:非常大的批次。对于这三个用例,人们开始认识到这些模式,作为提供商,我们当然会进行不同的优化。
You sample tons of completions from the model and then evaluate how good that is. As part of this training, you really need very large batch, high throughput inference. So I would say that's the third use case: very large batch. For these three use cases, people are starting to recognize these patterns, and as providers, we certainly do different optimizations.
你如何考虑在这三者之间分配资源?
How do you think about allocating your resources across the three?
这就是规模化运行真正有帮助的地方。我们称之为集群级优化。如果你在数千个 GPU 上运行推理,你可以动态改变集群分配。例如,运行批量推理:OpenAI 有那个选项,我们也有。如果我们看到集群不忙于交互式查询,我们可以发送批量查询来吸收算力。结果,我们可以提供批量 API 五折优惠,我认为 OpenAI 也这样做。DeepSeek 可能也这么做。
This is where running at scale really helps. We call it fleet-level optimization. If you're running inference on thousands of GPUs, you can dynamically change the cluster allocation. For example, running batch inference: OpenAI has an option for that, and we do too. If we see the cluster isn't busy with interactive queries, we can send in batch queries that soak up the compute. As a result, we can provide a 50% discount on batch API, and I think OpenAI does the same. DeepSeek probably does that too.
规模化确实有帮助。
Having it at scale really helps.
是的。
Right.
当你思考推理市场的演变时,是否感觉随着时间的推移有无限的优化可以做?因此总是有可能保持领先于其他玩家,或者在某些时候,每个方面的低垂果实都被摘完了,实际上是要构建一个更广泛的平台,在这些工作负载之上做其他事情?
As you think about the evolution of the inference market, does it feel like there are endless optimizations to do over time? So it's always possible to stay a bit ahead of other players, or at some point have the low-hanging fruit been plucked on each of these, and it's actually about building a broader platform to do other things on top of these workloads?
早期有很多低垂的果实。如果你编写合理的核函数,构建合理的推理引擎,你会得到比现有方案好得多的结果。但现在开源工具变得非常好。像 vLLM 和 SGLang 这样的东西非常流行,并且已经达到了生产级质量。我们与这些人合作并做出贡献。所以基线已经好多了。但也有越来越新的用例。客户来找我们说,‘我们非常关心前缀缓存’或‘我们非常关心低延迟,因为我们的应用需要这个要求。’或者‘我们不是做文本,而是做视频,这是一个不同的权衡,更面向吞吐量。’我们与这些客户合作。即使开源库变得更好,工作负载也在快速演变,以至于总有新的事情要做。模型变得如此之好,以至于有无数种方式提取价值,这就是为什么我们看到这么多初创公司基于这些模型构建。结果,工作负载将迅速演变。
Early on, there were lots of low-hanging fruits. If you write reasonable kernels, if you build a reasonable inference engine, you get much better results than what was out there. But now the open-source tooling is getting very good. Things like vLLM and SGLang are very popular and have reached production-level quality. We work with these folks and contribute. So the baseline has gotten much better. But there are also new and newer use cases. Customers come to us and say, 'We really care about prefix caching' or 'We really care about low latency because in our application we need this requirement.' Or 'We're not doing text, we're doing video, which is a different tradeoff, much more throughput-oriented.' We work with those customers. Even as open-source libraries get better, the workloads are evolving so fast that there's always new stuff to do. There's an explosion of models getting so good that there are many ways to extract value, which is why we see so many startups building on these models. As a result, the workload will evolve rapidly.
你认为快进一两年后,会有 27 种或大量不同的方式,每种都需要自己的优化吗?
Do you think fast-forwarding a year or two from now, it's like 27 or a ton of different ways, each requiring its own optimizations?
我认为它仍然会汇聚。我认为智能体式工作负载可能是一个杀手级应用。ChatGPT 在应用方面是一个阶跃变化——用户第一次接触到可以与之对话的语言模型,用来调试代码、查找信息、分析和综合信息。但下一组应用将是智能体式的:AI 模型能否自己采取行动并收集信息?这将需要不同的优化。现在不仅仅是在 GPU 上快速运行模型,而是如何连接人类通常使用的工具,比如网络搜索。随着这些工作负载进入不同的垂直领域——比如你是一名设计飞机的工程师——你希望模型能够访问你的设计软件。或者如果你是一名金融分析师,你有一个数据库,你希望模型能够接入。所以我认为这种工作负载将在未来一年左右成为主要工作负载之一。这是我的预测。
I think it will still coalesce. I think the agentic workload might be a killer use case. ChatGPT was a step change in terms of application—the first time users were exposed to language models they can talk to, to debug code, look up information, analyze and synthesize information. But the next set of applications is going to be agentic: can AI models take actions and collect information by themselves? That will require different optimizations. Now it's not just running models really fast on GPUs; it's how to connect with tools humans usually use, like web search. As these workloads move to different verticals—say you're an engineer designing airplanes—you want the model to access your design software. Or if you're a financial analyst, you have a database you want the model to tap into. So I think that kind of workload will be one of the main workloads in the next year or so. That's my prediction.
围绕智能体的系统级工作变成了一整套新问题需要解决。感觉即使原始工作负载有很好的开源工具,现在也有一整套全新的工作负载需要优化和系统工作,比如访问外部数据库。这与仅仅优化推理不同,但似乎是你们提供的产品的自然演变。
That systems-level work around agents becomes a whole new set of problems to solve. It feels like even if the original set of workloads has great open-source tooling, there's now an entirely new set of workloads that also require optimizations and systems work, like accessing external databases. That's different from just optimizing inference, but it seems like a natural evolution of the product you're providing.
是的,我认为这将成为专业或企业级的主要用例。在消费者层面,我的一个赌注是我们将获得实时视频生成。我们开始看到一些这样的东西,我认为它将真正改变消费者格局,类似于 TikTok 改变那个格局。它非常吸引人。在消费者方面,我们合作的一些公司,比如 Pika Labs 和 Hedra,专注于实时视频生成——这就是赌注。有一整套新问题需要解决,而且需要更多的算力。视频生成在算力方面要求非常高,所以这可能会推动更多的芯片和推理优化。
Yeah, I think that's going to be the main use case for professional or enterprise level. On the consumer level, one of my bets is that we'll get real-time video generation. We're starting to see some of that, and I think it will really change the consumer landscape, similar to how TikTok changed that landscape. It's very engaging. On the consumer side, some companies we work with, like Pika Labs and Hedra, are focusing on real-time video generation—that's the bet. There's a whole new set of problems to solve, and it's a lot more compute. Video generation is very demanding on the compute side, so that might fuel even more chips and inference optimization.
非常有趣。显然你在进行大量研究,并在这些重大未决问题之前下注,关于这个领域将走向何方。
Super interesting. Obviously you're conducting a lot of research and making these bets ahead of the big outstanding questions of where the space is headed.
我很好奇,如果你能快进到三年后,得到 AI 基础设施领域一个问题的答案,这个问题会真正帮你确定今天的方向。你最想知道或最影响你工作方向的那个问题是什么?
I'm curious if you could fast forward to the future 3 years from now and get the answer to one question in the AI infrastructure world that would really help you shape your direction today. What's that question that's outstanding that you're most curious to know or would most impact the direction of your work?
未来几年我试图回答的问题是:如何让 AI 达到专家水平?目前,我认为模型在一些任务上达到了人类中位数水平,比如前端编程——它们已经相当不错了。模型在前端编程或数据分析方面比我强得多。对于互联网上有大量数据的任务,这些模型会表现得很出色。所以它们在某些任务上达到了中位数或略高于平均水平。但经济价值高的任务仍然存在。我们花很多钱请人类专家设计飞机、硬件、医生、律师等等。这些人之所以成为专家,是因为他们花时间使用专业工具。他们没有互联网那么庞大的数据——这正是他们成为专家的原因。那么,我们如何让模型达到那个水平,与人类在那个水平上协作?这才是大量经济价值的来源。
The question I'm trying to answer for the next couple years is how do we get AI to expert level? Right now, I think the models are at the median human level on some tasks like front-end programming—they've gotten quite good. Models are much better than me at front-end programming, or analyzing data. For tasks where there's lots of data on the internet, these models are going to crush it. So they've gotten to sort of the median or maybe slightly above average level on some tasks. But the economically valuable tasks are still there. We pay lots of money for human experts in designing airplanes, hardware, doctors, lawyers, and so on. Those people become experts because they spend time working with specialized tools. They don't have an internet's worth of data—that's why they're experts. So how can we get models to that level, to work alongside humans at that level? That's where a lot of the economic value is going to come from.
显然,你需要围绕实现这一目标的方式做硬件优化。你研究过状态空间模型,花了很多时间思考替代架构、Transformer。你的合作者 Albert Gu 说过,Transformer 本身不会是最终解决方案。你认为我们需要架构创新才能达到那个水平吗?
Obviously, the hardware optimizations you have to do around whatever the way to do that is. You've worked on state space models, spent a lot of time thinking about alternative architectures, transformers. Your collaborator Albert Gu said that transformers are not going to be the final solution by themselves. Do you think we need architectural innovation to get us to that level?
对。我喜欢 Albert,但也许这是我们有点分歧的地方。我认为要达到 AGI 或 ASI,我们现有的架构可能就足够了。但代价是什么?如果你有更好的架构,也许你能提前一两年达到——这很值得——或者以十分之一的成本达到,这最终会影响到 ASI 的可及性。这当然至关重要,因为我们每年在 AI 基础设施上花费大约 5000 亿美元。别引用我的话,但我觉得这个数字差不多。那么,我们需要花费 10 倍于此(这似乎不太现实),还是通过更好的架构以当前甚至更少的支出达到目标?我认为这就是架构将发挥作用的地方:我们能否通过更好的架构达到 AGI?我认为当前架构拥有所有正确的要素,如果你继续扩展(人们一直在做),你可以达到目标,但成本可能高得惊人。
Right. I love Albert, but maybe this is where we disagree a little bit. I think to get to AGI or ASI, it's possible that the current architecture we have is sufficient. But at what cost? If you have a better architecture, maybe you get there one or two years sooner—that's probably worth it—or you get there with 10x less cost, which ultimately flows to the accessibility of that ASI when you get there. That's certainly crucial because we're spending maybe $500 billion on AI infrastructure every year. Don't quote me on that, but I think that's the right ballpark. So do we need to spend 10x that, which seems somewhat unrealistic, or with better architecture can we get there with the current amount of spending or even less? I think that's where architecture is going to play a role: can we get to AGI with a better architecture? I think the current architecture has all the right ingredients, and if you keep scaling, which people have been doing, you could get there, but maybe the cost is just astronomical.
你还在关注哪些其他架构?
What other architectures are you paying attention to?
我对混合专家模型(MoE)非常兴奋,尤其是让它变得越来越稀疏。我们正在尝试推动这个极限:能稀疏到什么程度?我们正在研究这个。我认为这是一个很有前景的方向。DeepSeek 做了一些非常重要的工作,表明你可以让这些东西非常稀疏。甚至更早的 DeepMind 工作也在朝这个方向推进。所以我认为这是一种用相同算力获得更多智能的引人注目的方式。最终,我们想要优化每美元的推理性能。这意味着它可以分解为每浮点运算的推理性能和每美元的浮点运算次数。每浮点运算的推理性能更多关乎架构设计、数据和算法。另一方面,每美元的浮点运算次数关乎硬件和内核优化。所以在架构方面,我们试图从相同的计算中提取尽可能多的智能。混合专家模型是其中之一。与 Albert 合作的一些状态空间模型工作非常有趣,人们正在采用它。我们与 Nvidia 的一些人合作训练模型。他们发布了几款模型,展示了使用这种架构——Transformer 和 Mamba 的混合体——可以获得非常高质量的模型,但成本更低或推理性能更高。所以我认为架构对推理非常重要。如今我经常思考所谓“推理优先的架构设计”:反正大部分浮点运算都花在推理上,所以你真正想要的是设计让推理非常高效的架构。
I'm pretty excited about mixture of experts, especially getting sparser and sparser. We're trying to push that limit: how sparse can you go? We're working on that. I think it's been a compelling direction. Some really important work done by DeepSeek showing that you can make these things really sparse. There was even earlier work from DeepMind pushing in that direction as well. So I think that's one compelling way to get more intelligence out for the same compute. Ultimately, I think we want to optimize inference per dollar. That means it can be factored into inference per flop and flops per dollar. Inference per flop is more about architecture design, data, algorithms. The other side, flops per dollar, is hardware and kernel optimization. So on the architecture side, we're trying to extract as much intelligence out of the same computation. Mixture of experts is one. Some of the state space stuff done with Albert has been really fun, and people are picking that up. We've collaborated with a bunch of folks from Nvidia actually training models. They released several of these where they show they can get really high quality models using this kind of architecture—a hybrid of transformer and Mamba—but with much smaller cost or much higher inference performance. So I think architecture is very important for inference. Nowadays I think very much about what I call inference-first architecture design: most of the flops are spent on inference anyway, so you really want to design architecture that makes inference really good.
你通过 FlashAttention 和 Mamba 推动了推理方面的发展,是状态空间模型的先驱。我相信大家很好奇:你现在对哪些研究领域感兴趣?你在哪些方面投入时间?我们期待你的下一篇重要论文是什么?
You've pushed forward the inference side with FlashAttention, with Mamba, you've been a pioneer on state space models. I'm sure folks are curious: what research areas are interesting to you now? Where are you spending your time? What's the next big paper we should expect from you?
我仍然在这些领域工作,对它们非常感兴趣。但我正在探索一些新方向,其中一些涉及下一批真正有影响力的应用。我认为机器人技术是其中之一。机器人技术:我们距离拥有能在家庭中工作的优秀人形机器人还有多远?也许五年,也许十年。
I'm still working in these areas, very much interested in them. But I'm exploring some new directions, some of which involve working on the next set of really impactful applications. I think robotics is one of them. Robotics: how far are we away from having really good humanoid robots that work in homes? Maybe five years, maybe ten years.
在机器人研究领域,你最感兴趣的是什么?
What's most interesting to you within the robotics research area?
对于机器人,我们可以用已有的一些基础模型来初始化控制机器人的模型。你可以用语言模型来做规划——比如如果你让机器人拿起一个咖啡杯,语言模型可以说你走到那张桌子前,拿起咖啡,等等。但我认为缺少的是这种与世界的互动和行动,因为我们没有这方面的数据。我们有语言数据,但没有物理互动的数据。
So with robots, we can initialize a model controlling a robot with some of the foundation models we already have. You can use a language model to do the planning—like if you ask a robot to pick up a coffee cup, the language model can say you go walk over to that table and pick up the coffee, and so on. But what's missing, I think, is this kind of interacting and acting in the world, because we just don't have data for that. We have language data, but not for physical interaction.
对。你显然看到有些人试图扩展模拟数据,其他人则在用远程操作,但实际驱动侧的数据问题依然存在。
Right. You've obviously seen some folks trying to scale simulation data, other folks are doing teleop, but it's a data problem on the actual actuation side.
对。这肯定是其中之一。另一个是,我认为机器人必须以多分辨率、多时间尺度的方式处理信息。比如控制关节时,你必须非常快速地行动。但如果你在为机器人规划路线,你可以慢得多。所以这就像你可能想要使用不同的架构。
Right. So that's certainly one. The other one is I think robots have to process information in a multi-resolution, multi-time-scale way. Some of which is if you're controlling the joints, you have to act very rapidly. But if you're planning a route for the robot, you can do it much slower. So it's like different architectures you might want to use.
我觉得这似乎与你之前做的基于状态的模型工作非常相似。
I mean, I'm struck that seems very analogous to some of the state-based model work you've done.
对。对。所以明确考虑时间尺度——我是想做非常轻量的计算来仅仅控制关节,还是想做更重的推理来为机器人规划最优路线?所以我认为这将是一个由语言模型、视觉模型、音频模型、世界模型初始化的复合系统,但如何将它们拼接在一起?这是个大问题。
Right. Right. So explicitly taking time scale into account—do I want to do very lightweight computation to just control the joints, or do I want to do much heavier-weight reasoning to plan the optimal route for the robot? So I think it's going to be this composite system initialized from language models, vision models, audio models, world models, but how do you stitch them together? That's the big question.
还有一些迷人的硬件问题,要能在本地足够快地运行这些东西。
Some fascinating hardware questions to be able to run the stuff locally and fast enough.
是的。是的。
Yeah. Yeah.
我觉得你大概可以在任何地方工作。一个非常有趣的事情是,你是 Together 的首席科学家,但你也走了学术路线,在普林斯顿担任助理教授。你是如何看待留在学术界的?以及在学术界与工业界分别适合研究哪些问题?
I feel like you could probably work anywhere you want. One really interesting thing is you're chief scientist at Together, but you've also gone the academic route as an assistant professor at Princeton. How did you think about staying in academia? And what set of problems make sense to work on in academia versus industry?
对。这是个好问题。我的答案是因人而异——对每个人来说都不同。就我个人而言,我其实喜欢同时待在初创公司和当教授。它们呈现了不同的思维和执行模式。在初创公司这边,非常有趣,因为我们行动很快。如果我们想做什么事情,可以在几天、几周或最多几个月内完成。这就是我们规划的时间跨度。而且非常有趣。你有一支非常优秀的团队,他们执行得非常快,你能完成伟大的事情。我为 Together 团队所做的工作感到非常自豪。所以那非常有趣。在学术界这边,时间跨度更长一些,我们思考的问题更具推测性。我们不解决需要一个月内解决的问题。我们思考的是:如果世界在两三年后朝那个方向发展,那么有哪些有趣的问题要问,以及在一两年内要解决哪些挑战?我们按那个时间尺度思考。与学生合作也非常有趣且有回报,因为我们能思考这些本身就有趣的问题。显然,有一些权衡——学术界的算力要小得多,评估方式是你的想法是否有趣,而不是它是否有效且运行得很快。那是不同的评估标准。所以我认为在学术界,你有更多自由去深入思考这些更长周期的问题。我恰好喜欢这两种运作模式。这就是为什么我仍然在普林斯顿当教授,并且仍然参与初创公司。
Right. This is a great question. My answer is it's going to be personal—for every person it's different. For me personally, I actually like being both in a startup and being a professor. They present different modes of thinking and execution. On the startup side, it's very fun because we move fast. If things we want to do, we could get it done in days or weeks or at most months. That's the horizon we plan for. And it's very fun. You get a team of really good people, they can execute very fast, and you accomplish great things. I'm very proud of the work the team has done at Together. So that's been very fun. On the academia side, it's a little bit longer time horizon, and the kind of problems we think about are more speculative. We don't solve problems that need a solution in a month. We think about: if the world is moving in that direction in two or three years, what are the interesting questions to ask and what are the challenges to solve in one or two years? We think in that time scale. Working with students has been very fun and rewarding as well, because we get to think about these questions that are intrinsically interesting. Obviously, there are trade-offs—the amount of compute in academia is much smaller, and the way you evaluate is whether your ideas are interesting rather than whether it works and runs really fast. That's a different set of evaluation. So I think in academia, you have some more freedom to think deeply about these longer-horizon problems. I happen to like both modes of operation. That's why I'm still at Princeton as a professor and still involved in the startup.
对。
Right.
而且我认为最终这可能是探索与利用的结合。如果你从宏观层面看,学术界的角色更偏向探索。资金通常来自政府,政府关心的是探索大量想法,看看其中也许 5%或 10%会成功。作为投资者或风险投资人,你运作的模式也类似——探索大量想法,也许 5%或 10%会变得极其重要。一个例子是注意力机制。尽管它因谷歌的论文而闻名,但它实际上出自麻省理工学院的工作,由 Dzmitry Bahdanau、Yoshua Bengio 和其他人完成。它来自学术界。如果你想想当前架构中的所有其他组件,优化器、Adam 优化器,是多伦多的 Jimmy Ba 和其他人做的。LayerNorm 也是 Jimmy Ba。基本上,我们现在拥有的很多东西都来自学术界,因为学术界的角色是探索大量想法。其中一些会成功。然后大公司和初创公司的角色是,他们采纳其中一些想法,进行创新,执行得非常快,但他们了解市场需求以及应该构建什么产品,并且他们有更大的财务实力来实现这些更大的想法。
And I think ultimately it's maybe a mix of exploration and exploitation. If you think at a macro level, academia's role is more on exploration. The funding comes from the government, and the government cares about exploring tons and tons of ideas to see maybe five or ten percent of those ideas will work out. As an investor or VC, that's kind of the mode you operate in as well—explore lots of ideas, maybe five or ten percent will become incredibly important. One example is attention. Even though it got famous with the Google paper, it actually came out of work at MIT with Dzmitry Bahdanau, Yoshua Bengio, and other folks. It came out of academia. If you think of all the other components in the current architecture, the optimizer, Adam optimizer, was Jimmy Ba at Toronto and other folks. LayerNorm also Jimmy Ba. Essentially, lots of what we have now came out of academia because the role of academia is to explore lots of ideas. Some of them will work. Then the role of big companies and startups is they take some of these ideas, they innovate on them, they execute really fast, but they understand what the market needs and what product they should build, and they have much bigger financial power to make some of these bigger ideas happen.
这个融资环境的有趣之处在于,我觉得它几乎扭曲了某些方面——有大量风险投资可供人们进行更像研究实验室的工作。比如,如果你提出‘嘿,机器人技术未来几年需要不同的架构和硬件’,有人可能会说:‘好的,酷。你十年内不会盈利,但这想法很有意思,给你一笔风投资金去干吧。’
The interesting thing about this funding environment is I feel like it's almost distorting some of that where there's plenty of venture funding available for folks to do more research lab style work. Like if you pitched, 'Hey, robotics is going to need different architectures and different hardware over the years,' someone might say, 'Okay, cool. You're not going to monetize for 10 years, but that's a really interesting idea, and here's a bunch of venture money to go out.'
是啊。人们拿到的融资简直疯狂,对吧?比如 Ilya SSI,他明确说‘我们不打算做任何产品’。但人们还是想给他钱,因为他就是 Ilya,对吧?还有你提到的机器人领域的一些人也在这么做。所以我觉得,在 AI 领域,这些风险似乎开始有回报了。投资方更愿意砸钱了。
Yeah. It's kind of crazy the kind of funding that folks are getting, right? Like Ilya SSI, he was explicitly like, 'We're not going to build any product.' But yeah, people want to give him money because he's Ilya, for sure, right? And some of the folks like you said in robotics are doing that. So yeah, I think people in AI, it seems like some of these risks are starting to pay off. So on the investing side, people are much more willing to put in money.
完全同意。好了,这场对话非常精彩。我们习惯在采访结束时来个快问快答环节,听听你对一些问题的看法。那么首先,过去一年里,你对 AI 的哪个看法发生了转变?
Totally. Well, this has been a fascinating conversation. We always like to end our interviews with a quick fire round where we get your take on some questions. So maybe to start, what's one thing you've changed your mind on in AI in the last year?
这些模型出奇地有用,甚至在我日常的高级专业工作中也是如此。它们在数学和编程方面出奇地好。
These models are surprisingly useful even for my daily work at the advanced and expert level. They're surprisingly good at math and coding.
是啊,1.5 倍比我预想的高多了。这相当令人印象深刻。你认为一年后,开源模型在质量上会离闭源模型更近还是更远?
Yeah, 1.5x is way higher than I would have expected to say. That's pretty impressive. Do you think open source models are going to be closer or further in quality from closed source models in a year?
我会说更近。我认为现在的 Scaling(规模扩张)更多在强化学习方面,而这实际上更依赖工具而非原始算力。所以我认为开源在这方面会做得很好。
I would say closer. I think the scaling now is more on the RL side, and that relies actually more on tooling than just raw compute. So I think open source is going to do really well there.
你认为在更广泛的 AI 领域,有哪些发展是人们目前关注不够的?
What development in the broader AI space do you think people aren't paying enough attention to right now?
绝对是数据。我认为数据总是被低估了。数据方面发生了很多事。合成数据,用模型来改写数据,影响巨大,但人们可能关注不够。
Definitely data. I think data is always a little bit underhyped. I think lots has happened on the data side. Synthetic data, using models to rephrase that, has huge impact that maybe people have paid less attention to.
你有没有最喜欢的基于 Together 构建的应用?
Do you have a favorite application that you've seen built on Together?
有,我们和 Pika、Hedra 等视频生成公司合作过,他们用基于 Together 训练的模型生成病毒式传播的 TikTok 视频,并在 Together 上进行推理。效果非常棒。
Yeah, we worked with some video generation companies like Pika and Hedra, and they're generating viral TikTok videos with the models they train on Together and doing inference on Together. So it's been amazing.
是啊。这感觉蓄势待发,会有很多发展。好了,太棒了。这场对话非常精彩。我想把最后一句话留给你。人们可以去哪里了解更多关于你的信息?无论是你在 Together 的工作,还是你的学术工作,任何你想指出的地方,话筒交给你。
Yeah. And that feels like poised for a lot of development. Right. Well, amazing. This has been a fascinating conversation. I want to make sure to leave the last word to you. Where can folks go to learn more about you? Either your work at Together, your academic work, anywhere you want to point folks, the mic is yours.
当然。我在 Together 的工作,我们在 Together 上发布博客文章。我还在 Twitter 上,账号是 tri_dao。偶尔我也会在我的网站 tri-dao.me 上写博客。
For sure. So my work at Together, we put out blog posts on Together. I'm also on Twitter, tri_dao. And once in a while I write blog posts on my website, tri-dao.me.
太棒了。非常感谢。这太迷人了。
Amazing. Well, thanks so much. This is fascinating.
是啊,这真的很有趣。谢谢。
Yeah, this has been really fun. Thanks.