How JLM52 Writes GPU Kernels and Long Query Routing
打开互动全文版(中英对照 + 朗读 + 问答)→探讨使用 JLM52 自动编写 GPU 内核,以及推理引擎中处理长查询的过程。
A discussion on using JLM52 to auto-write GPU kernels and the process of handling long queries in inference engines.
JLM52 is very, very good at writing GPU kernels. It was very funny internally—we had a JLM52 endpoint that we plugged into our Claude Code harness. So every engineer on the team used our JLM52. It would do a forward pass on the JLM52 instance of the node, then get the profile trace, analyze it, find the kernels that are the bottlenecks in that shelling, then write new kernels, then do another profiling trace, and once that's done, it uploads to our thing, and then we can pull that image down and repeat the cycle. Some of the GPU kernels that we run on JLM52 within our inference engine are written by JLM52.
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now let's get into it. Okay, we're here in the studio with Philip, old friend from Inference Engineering the book, as well as Baseten, and everything that you've done—you and I have done before—as well as Ali. Welcome.
Pleasure to meet you.
Waterloo intern.
Waterloo intern Ali.
When did you get Waterloo intern as a handle?
As a handle?
Handle?
I think the rebranding happened like mid-March. When I saw it was open, I was like, I have to take it or for grabs.
The problem is that Ali is really good at his job and is not going to be an intern much longer. So we have to figure out who's going to get the handle.
Pass the torch over to the—
Oh, okay. It can be like you just pass it to another Waterloo grad.
Yeah.
Intern.
You got to get an intern from Waterloo.
But they have to— Oh, it could— But it could come from Base 10. So it's like wherever Baseten gets from Waterloo—
Right. Has the title of Waterloo—
You either get it or you're out. You should also do like a big graduation ceremony where you change the handle.
I mean, you guys are good at ceremonies, clearly. You know, we had a nice launch of the book, very successful. But before we get into all that, I want to start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say 200,000 tokens, into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?
With a long query specifically, the first thing that I'm going to ask is, have you sent me this query before, or at least part of it? And I really hope you have, because it's going to be a lot easier for me and a lot cheaper for you. So the first thing that we're going to look at is some kind of cache-aware routing, where we're going to see—we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of this 200,000 tokens. If you're doing 200,000 tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're going to have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode. So you're going to have one set of GPUs that's solely going to process the input query, that KV cache, and get you your first token. And then that's going to be passed over to a separate set of GPUs which is going to run decode. We'll go into iteratively make those tokens. We'll probably have some kind of speculative model in front of that. I'm going to assume that you do encoding, and because of that, a speculative model which assumes you do encoding is going to have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's going to be slower. And then we stream that output to you and account for it, charge you some number of couple of pennies, and say, 'Hey, would you like to send another one?'
Except these tenders don't charge by pennies.
Well, yeah, we charge—I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.
Yeah, I mean, one of the key differentiators when I was talking with Baseten initially was that actually people who want very, very high volume just need to rent by the box, because then it's up to you to figure out how to saturate the box.
And more often than not, it's like way cheaper if you're pushing like millions of tokens per hour if you just pay per hour instead of pay per token.
Yeah, they do. I think that we've increasingly seen a lot of demand for the sort of per-token APIs just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.
A best practice on when it's time to swap over?
Couple reasons. Yeah, reliability, that's a big one, right?
Like they have a very specific use case. They want you to train something specifically for them. Like they want their own spec doc, for instance, for their own traffic.
Spec doc is speculative decoding.
Speculative decoding.
Yeah, yeah, yeah. You like this—sorry. Like spec—the way spec works is basically if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach like this little kind of parasite—this layer that goes on top of the model. And this model just has to predict those three very fast autoregressive passes, and it will predict like three certain tokens, and then you do one forward pass over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them.
现在这个草稿模型是特定于草稿的,所以如果你喜欢,就像 Philip 说的,如果你在总结《哈利·波特》系列,我可以专门在那个草稿模型上训练《哈利·波特》系列,我可以保证我每次都会接受那三个 token。这样一来,我就提高了你的解码速度。如果你想要一个共享端点,我就没法提供这个,因为我不知道你是在做《哈利·波特》还是编程还是英语——我们不知道。还有,书里提到过,如果他们真的关心某个特定阈值,我记得是第四章。你还记得吗?
Now this draft model is draft-specific, so if you like, as Philip said, if you're summarizing Harry Potter books, I can train that draft model exclusively on Harry Potter books, and I can guarantee you that I'm going to accept the three tokens every single time. And so in that case, I increase your decode speed. I wouldn't be able to provide this to you if you want a shared endpoint, because I have no idea if you're doing Harry Potter or coding or English—we don't know. Also, there's a thing in the book that mentioned that if they really cared about a specific threshold, chapter four I think. Do you remember that?
是的,你能做的事情是,你可以设置特定的批处理大小、特定的并行策略。如果你试图优化吞吐量而不是延迟,你可能会使用一个没通过基准测试的 NVFP4 量化,而你想以更高精度运行模型——你可以这么做。有很多原因让你想要自己的端点,当然最大的一个就是,当你碰巧在服务用户时,你不需要处理别人在端点上跑一亿 token 的基准测试流量。
Yeah, the things that you can do is you can set a specific batch sizing, a specific parallelism strategy. If you're trying to optimize for throughput versus latency, you can maybe use an NVFP4 quant that doesn't pass your benchmarks, and you want to run a model at higher precision—you can do that. There's just a bunch of reasons why you might want to have your own endpoint, and the biggest one of course being that you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.
是的。我觉得有一件事是经典旅程,你知道,就像人们问在浏览器里输入 Google 会发生什么。工具调用——那只是生成 JSON 吗,还是说背后还有更多复杂性?
Yeah. I think one thing that is a classic journey, you know, like it's basically people asking what happens when you type Google into the browser. Tool calling—is that just you know you're generating JSON, or is there more complication beyond that?
我们的一些客户,他们有自己的后训练模型,所以他们要求的工具调用不仅仅是解析文件或查天气——那是非常具体的东西,你必须对此进行后训练。如果模型的后训练不好,或者后训练后的量化(为了让推理变快)不好,模型就会在读取 JSON 文件和工具调用时遇到困难。但这不需要它自己的沙箱。它不会用那个工具调用来逃出沙箱,也不需要被遏制。它可以只是一个普通的专用部署。工具调用面临的挑战越来越多,似乎是公司想要特定的工具调用,而这是训练中非常敏感的东西。因为你处理的是所有的 JSON 输出,如果它没有以非常特定的方式关闭请求的结尾,那个做了工具调用和思考的模型——结果就是,它没有看到结果,只是在解码时幻觉出了一个结果。这似乎是工具调用中最具挑战性的部分,而不是沙箱问题。
Certain customers that we have, they have their own post-trained models, and so they demand tool calling that's not just like parse a file or go find the weather—it's something that's very specific, and you have to do post-training on this. And if the post-training on the model is not good, or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own sandbox. It's not like it's going to use that tool calling to escape a sandbox or doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that companies want certain tool calling, which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't close the end of the request in a very certain manner, the model that did the tool calling and the thinking—as a result of that, it didn't see the result and just hallucinated a result as it decoded. That seems to be the most challenging thing with tool calling. Not really the sandbox problem.
是的,这是训练方面的挑战,而在推理方面,你可以做一些工作来限定可能的输出。所以,我们实际上在差不多两年前发表了解决这个问题的方案——你基本上做一个状态机,用它来约束输出为特定格式。所以,这就是结构化输出问题。如果你还记得以前的语法……
Yeah, that's a challenge on the training side, and then on the inference side, there's work that you can do to scope the possible output. So, we published this actually at this point close to 2 years ago—the solution to this problem, which is you basically make a state machine and you use that to constrain the output to a specific format. So, this is the structured output problem. If you remember back in the grammar...
是的,这回到语法——Gmail 就有这个。
Yeah, this back in the grammar—Gmail had this thing.
是的。所以,这就像老派的‘确保这只是 JSON,否则我的语法会死’这类问题。
Yeah. So, it's like the old school 'make sure this is only JSON or my grammar's going to die' type of problem.
某个时候 OpenAI 发布了一个东西,说,是的,如果你想约束你的输出,写 BNF 语法——巴科斯-诺尔范式。
At some point OpenAI released a thing that was like, yeah, if you want to constrain your output, write BNF grammar—Backus-Naur.
在我们的推理系统中,它只是一个指定的输出格式,你得到保证你的输出会按照那个格式结构化。所以,把它应用到工具调用上可以帮助减少——显然,你仍然可能调用错误的工具或根本不调用工具。它不能解决确定性问题,但至少解决了工具调用的输出结构化问题。
In our inference system, it's just a specified output format, and you get the guarantee that your output's going to be structured along that format. And so, applying that to tool calls can help cut down on—obviously, you can still call the wrong tool or call no tool. It doesn't solve the certainty problem, but it at least solves the output structuring problem with the tool calls.
而 MCP 只是工具的另一种形式,对吧?我的意思是……
And MCP is just another form of tool, right? I mean...
是的。
Yeah.
完全正确。那里没有什么特别的东西。
Exactly. There's no special thing there.
我总是向人们解释的是,LLM 实际上没有能力做任何事情。它只能提出做什么的建议。然后,如果这些建议以某种方式格式化,并应用于一个知道如何处理它们的系统,那么一个行动就发生了。
The thing I'm always explaining to people is the LLM is actually not capable of doing anything. It's only capable of making suggestions of what to do. And then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.
是的。
Yeah.
有趣的部分是,你知道,这也在工具调用之外解决了。比如在智能体循环中,如果输出不正确,你是对的。就像推理工具调用是在推理销售中完成的。就像,‘哦,我不知道该怎么办。让我再试一次。’你知道,它可能试几次就成功了。关于你提到的训练,有时这在较小的模型中更难。所以,当你从大模型切换过来时,你不会得到完全相同的质量输出,对吧?
Part of the fun stuff is, you know, this is solved outside of tool calling, too. Like in an agent loop, if the output is not correct, you're right. Like reasoning tool calling was done in the reasoning sale. Just be like, 'Oh, I don't know what to do. Let me just try again.' And you know, it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models. So, you don't have the same exact quality output when you just swap from a big model, right?
是的。我要说的是,在我们——我觉得我们需要回到推理工程本身——但我曾预期会有东西取代 JSON,因为流式传输 JSON 很难,因为 JSON 必须完全——你必须有大括号的开和闭等等。所以,在流式传输时很难解析或验证某些东西。所以,人们发明了各种各样的东西,比如——我忘了其中一些替代品的名字,但基本上像 TOML,像 YAML。但 JSON 似乎仍然占主导地位。
Yeah. I will say that before we—I think we need to go back to inference engineering proper—but I had expected that something would replace JSON, because it's hard to stream JSON, since JSON must be completely—you must have open and close brackets and everything. So, it's hard to parse something or validate something while it's being streamed. So, people invented all sorts of things that are like—I forgot the name of some of these alternatives, but it's basically something like TOML, something like YAML. But JSON seems to be dominant still.
JSON 输出有那么长吗,对吧?我想你可能会有很长的——因为工具调用也包含参数,也许对于某些工具,你可能会传递很长的参数。但我的印象是,中位工具调用是相对较少的 token 数量,对吧?所以,我预计投机者通常相当擅长像 JSON 这样格式化的东西。所以,你会有相当快的解码步骤,而流式传输就不会那么有价值。但也许我错了。
The JSON output's all that long, right? Like I guess you could have a long—because tool calls also contain the arguments in them, and perhaps for certain tools you might pass a very long argument. But my impression of the sort of median tool call is that it's a relatively small number of tokens, right? So, I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there, and that the streaming wouldn't be as valuable. But maybe I'm wrong about that.
它受限于模型将要集成的软件。如果软件是用 JSON 构建的工具调用,或者如果你的客户说这就是我们软件的工作方式,我们的工具用 JSON 接口,你可以要求他们改变他们的软件,然后说,是的,这对模型会更好。但对于湿训练,应该不会有太大区别。而且如果输出更多 token,可能也更有利可图。
It's bounded by the software that the model is going to integrate with. If the software is built with JSON for the tool calls, or if the company that you're—if your customer says that this is how our software works and our tools interface with JSON, you can ask them to change their software and then say, yeah, this is going to be better for the model. But with the wet training, shouldn't be that much of a difference. Also more profitable if it outputs more tokens, probably.
这取决于你的商业模式。你知道,这真的取决于情况。但我要说的是,作为一个作家,我经常用 AI 生成的输出做实验,我确实尝试从文本转向 JSON 文本,那是非常长的 JSON,对吧?就像每个字段里都有段落,因为我试图结构化它。
Depends on your business model. You know, it really depends. But I will say that as a writer, I experiment a lot with AI-generated output, and I do try to move from text to JSON text, which is very long JSON, right? Like there are paragraphs in every field because I'm trying to structure it.
我希望你先做事实陈述,然后做观点陈述,再做要点总结,要有日期、有实体引用、有参考来源,所有这些。总之,这些我觉得真正尝试结构化输出的人必须非常在意。但让我们稍微往上一层递归。在我们开始录音之前,你提到了一件很酷的事,就是当新的模型提供商发布新模型时,会有大量的工程工作,推理工程。对吧?比如 GLM 5.2、Kimi K3。我之前以为,尤其是像 GLM 5 到 5.1 到 5.2 这样,你们之前已经支持过它们了。工作量真的那么大吗?
I want you to first make factual statements, then make opinion, then make bullet point summaries, have dates, have entity references, have your sources for references, all these things. Anyway, so these are things that I think people who really experiment with structured output have to really care about. But let's sort of recurse up the stack a little bit. Before we started recording you actually mentioned something which is really cool, which is that there's a lot of engineering, inference engineering, that goes on when a new model provider releases a new model, right? So let's call it GLM 5.2, Kimi K3. I had previously assumed, especially if it's like well GLM 5 to 5.1 to 5.2, that you know you supported them before. Is it that much work?
工作量很大。
It's a lot of work.
是啊。好吧,所以你知道很多人,你们所有人,对吧,每当新模型发布,人们就急着说,哦,Hugging Face 支持这个,Fireworks 支持这个,Baseten 支持这个。我就想,是啊,他们当然支持。但这背后是什么?是什么……
Yeah. Okay, so you know a lot of people, all you guys, right, whenever a new model launches, people rush to say like oh Hugging Face supports this, Fireworks supports this, Baseten supports this. And I'm like yeah of course they support it. But what goes into that? What goes into...
我觉得这不仅仅是支持的问题,对吧?这对消费者好处很大。比如我记得是 Kimi K2.5 或者最新的 GLM 5.2,当时有一场推理速度之战,对吧?某家提供商每秒 90 个 token。第二天我们就到了 150。
I think it's more than just supported too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM 5.2 the latest, there was sort of an inference war, right? X provider is at 90 tokens a second. The next day we're at 150.
我算是用 GLM 5.2 挑起了那场竞争。我写了一篇 Twitter 文章,大概有 50 万浏览量。
I kind of kicked that off with GLM 5.2. I wrote a Twitter article about that, got like half a million views.
Baseten 第一。
Baseten number one.
在社交网络上。
For social networks.
是啊,这让大家都很兴奋,想着,嘿,我们怎么才能再往前推一点基准。而且,支持一个模型,指的是我能从这个模型生成 token,和支持一个模型,指的是我有一个生产就绪的 API,这两者是有区别的。要做到能从这个模型生成 token 并不难,因为通常开源推理引擎,比如 vLLM、SGLang 这些,很多时候甚至能提前拿到权重,维护者会,或者模型开发者会合并 PR 来确保支持。所以大多数情况下,你通常能比较轻松地在标准开源栈上跑起来。挑战在于,每家推理公司都有自己的专有栈,有些开源组件,有些自研的东西。而且对于任何任意的模型,总会有一些新东西。有时候你运气好,比如 K2.5 到 K2.6 就挺相似的。
Yeah, which then got everyone really excited about, hey, how can we benchmark a little bit further. And there was a difference between supporting the model as in I can make a token out of this model, and supporting a model as in I have a production ready API from this model. Getting to the point of I can make a token out of this model is not that hard, because generally the open source inference engines, your vLLMs, SGLangs of the world, often times even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can just kind of get it working on the standard open source stack without too much pain in most cases. The challenge is, you know, every inference company's going to have a proprietary stack. You know, some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K2.5 to K2.6 was pretty similar.
是啊,如果我没记错的话,那纯粹是持续的后训练。
Yeah, it was pure continued post-training if I remember correctly.
即使在这种情况下,还是有事情要做。你得重新做量化工作。你把模型从……通常这些模型发布时不是 NVFP4 格式,而我们希望它们是 NVFP4 格式,以获得最佳的 Blackwell 兼容性。所以我们必须执行量化,并校准量化,确保不会导致模型智能的任何退化。然后我们还得训练投机解码器,就像我们之前讨论的。通常,我们有……显然,我们的模型 API 有 ZDR,零数据保留,所以我们不完全知道人们发送的流量,但我们知道什么受欢迎。我们知道编码用例很受欢迎。我们知道智能体式用例很受欢迎。所以我们可以获取代表这类流量的公开数据集,并训练通用的投机解码器。现在,对于投机解码器,你需要用基础模型本身来训练,因为你从模型在这些特定提示上运行推理得到隐藏状态,那就是你用来创建投机解码器的训练数据。所以这个过程需要真正的模型权重。然后当然还有搭建所有基础设施、加载所有东西、测试的过程。而且当新模型有新架构时,我觉得显然 DeepSeek 模型往往最具挑战性,因为它们有最多新颖的架构,一个模型接一个模型。但每个新模型都有点新东西。我的意思是,Kimi K2 有……哦,抱歉。GLM 5.2 有……
Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from... Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization and calibrate the quantization to make sure that we're not causing any kind of regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have... Obviously, we have ZDR, zero data retention, on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agentic use cases are popular. So we can get public datasets that are representative of that kind of traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself, because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process, which you need the real model weights for. And then there's of course just the process of standing up all the infrastructure behind it, loading all the stuff in, testing it. And then when there's a new model with a new architecture, I think that obviously the DeepSeek models tend to be the most challenging, as they have the most novel architectural stuff going on, model over model. But every new model has something. I mean, Kimi K2 had... Oh, sorry. GLM 5.2 had...
Spores。
Spores.
对,DSA。
Yeah, the DSA.
对,这是从 DeepSeek 带来的。
Right, which is brought from DeepSeek.
对,对,而且……
Yeah, yeah, and...
所以你可以复制粘贴然后做那个。
So you can copy-paste and do that.
我不知道这是怎么运作的。
I don't know how this works.
你知道,所以我们必须在我们的运行时中构建对它的支持。你说得对,所有开源实验室互相借鉴的方式真的很有趣。比如,GLM 5.2 没有视觉能力。所以我们团队的一个家伙 Haley,如果我们能看看这个,他有点像把 Kimi 的视觉编码器嫁接到了 GLM 5.2 上。
You know, so like we had to build support for that into our runtime. And you're right, it actually is really interesting the way that all of these open source labs borrow from each other. For example, GLM 5.2 doesn't have vision. So something that Haley, a guy on our team, if we could take a look at this, he kind of grafted the Kimi vision encoder onto GLM 5.2.
只训练投影器。
Only training the projector.
正是。所以如果你想想编码器,编码器是看图像并将其转化为潜在信息的部分。然后还有投影器,它有点像……
Exactly. So if you think about the encoder, there's the encoder which is the part that looks at the image and turns it into latent information. And then there's the projector which kind of like...
空间。没关系。
Space. It's okay.
然后还有投影器,它把它映射到模型本身,然后还有模型权重。你不想动模型权重,因为为了给它视觉,你有可能会让模型在其他方面变笨。所以相反,他一开始只用一个投影器,只有几百万个参数。
And then there's the projector that maps it onto the model itself, and then there's the model weights. You don't want to mess with the model weights because you run a chance of making the model dumb at something else for the purpose of giving it vision. So instead, he started with just a projector, which is only a handful of millions of parameters.
那一年。
The year.
对,而且……
Yeah, and...
你能展示一下训练过程吗?它顿悟的方式非常非常有趣。
Can you show the training one? Like the way it groks is very very interesting.
而且也许,你知道,也许 Ali,你应该从这里接手。你对此的理解比我要好。
And maybe, you know, maybe Ali, you should take it from here. You've got a better understanding of this than I do.
是啊,你可以看到他训练这个的方式真的非常酷。一开始他训练它时,只是像这样:这里有一张山的图片,你能描述一下这座山里有什么吗?那导致了它第一次学习。但在这里你可以看到,我们试图教它的只是翻译编码后的信息,它已经从 Kimi Kaze 那里拿来了编码。它拿了图像。它是冻结的,冻结的,冻结的,带有适配器。所以他们理解大脑是冻结的,眼睛是冻结的。我们只是试图在眼睛和大脑之间建立连接,对吧?所以是投影器。所以你拿这些 token,然后他说,哦,你能描述一下这张图片里有什么吗?然后它说,哦,这是一座山,或者是一个人,或者是一个人类,诸如此类。但那并没有导致完全的理解。所以他改变了方式,让每张图片都关联一个数据集的问题。
Yeah, you can see like the way he trained this is really really cool. At the beginning he was training it using just like here's a picture of a mountain, can you describe what's in this mountain? And that caused it just like the first learning was. But here you can see this all we're trying to teach it is to translate the encoded, like it's already taken the encoded from Kimi Kaze. It's taken the image. It's frozen, frozen, frozen with adapter. So they're understanding the brain is frozen and the eyes are frozen. It's just we're trying to interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, oh, can you describe what's in this image? And he's like, oh, it's a mountain or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a dataset of questions.
比如,这张图里有没有一个白人男性?这张图左上角有没有鸟?这张图里有没有科学家?诸如此类。它必须正确回答这些问题。而且不只是训练它描述图像,还要能回答一个问题、再回答一个问题、再回答另一个问题,持续作答。你能看到那种顿悟(grokking)现象,把视觉能力嫁接到大型语言模型上居然能学到这种程度,真的不可思议。即使对于它表现不佳的图像,比如你给它一张斯蒂芬·霍金的照片,问这是谁,它可能答不上来,但会说这是阿尔伯特·爱因斯坦。它仍然明白这是一位科学家、一位男性、有重大成就等等。所以这真的非常酷。
Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer a question, answer question, answer question, other question, answer over time. Like you can see the grokking, which is like genuinely insane that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, who is this? Maybe it doesn't get it, but it will say something like this is Albert Einstein. Like it still understands this is a scientist who is a man who has you know, some significant achievements, all that stuff. So that's like really really cool.
是的,我们之前讲过 Tian,他是 Lava 论文的作者,那篇论文之前就做过这个。我觉得对于任何没做过视觉工作的人来说,这都是非常基础的工作。
Yeah, and so we've covered how Tian before, who was the author of the Lava paper that did this a while ago. And I think that that's very foundational work for anyone who hasn't done vision work before.
CLIP 和 Meta CLIP 也一样,从单纯的图像描述转向基于图像构建问题,性能提升非常明显。
Same with CLIP and Meta CLIP where you go from just captioning to building out questions off the image and how much better you can get performance.
对对对。
Right, right, right.
是的,但最令人兴奋的是,如果你看这样一个模型。显然,这更像一个研究项目。它还没达到,我记得 MMLU 上大概 56% 吧。所以还不太够格。但是,如果你运行这个模型,在没有图像输入时,你的 GLM 52 质量不会有任何损失,它的行为会和以前完全一样。
Yeah, but what's so exciting about this is if you look at a model like this. Now, obviously, this is a little bit more of a research project. It's not you know, it got to a 56% on MMLU for I think. So, not quite fun too. But, if you're running this model, you haven't suffered any loss on your GLM 52 quality if you don't have an image, it'll just behave exactly the way it used to.
在推理代码中,你实际上不会包含另一部分,对吧?
Which in the inference code, you literally do not include the other part, right?
是的,是的,我的意思是,如果没有图像输入,你只需跳过编码器。
Yeah, yeah, I mean you would just skip the encoder if you don't have an image input.
嗯,只是确认一下。这对整体推理端影响大吗?你并没有增加太多东西。你只是加了一个很小的视觉编码器。这些通常不到十亿参数。
Um, just confirming. Does it affect a lot on the overall inference side? Like you're not adding much. You're adding a very small vision encoder. These are typically like less than a billion parameters.
是的,我是说,视觉编码器之间的标准化程度稍低一些。所以支持矩阵可能有点稀疏,但总的来说,嗯,是的,它是整个系统中相当小的一个组件。最终,你从这个系统中得到的是,突然之间,你把 Kimmy 视觉、GLM 权重和 DeepSeek 注意力机制都整合到一个模型里。我认为开源的力量和美妙之处就在于,你可以把所有这些不同的组件组合在一起,形成一个比任何单个组件都更好的系统。
Yeah, it's I mean there's a little bit less standardization among vision encoders. Um, so the sort of support matrix can be a little bit sparse, so but overall, um yeah, it's it's a pretty it's a pretty minor component of the overall system. And ultimately, what you get out of the system is all of a sudden you have Kimmy vision, GLM weights, and DeepSeek attention all in one model. And that's I think a lot of the power and beauty of open source is that you can take all of these different components and combine them together into a system that's better than anyone can be individually.
人们以前说你会做弗兰肯斯坦式合并,从每个模型里取一些层。现在还有人这么做吗?
People used to say that you would also do franken merges where you would take like layers from each model. Does anyone do that anymore?
嗯,回到你之前提到的,当一个模型刚发布时需要做的支持工作,比如 GLM 52 或 MiniMax M3 之类的。有时你确实需要更换一些东西。比如,MiniMax M3 使用了全注意力机制。使用全注意力机制时,你会遇到一个巨大的瓶颈,因为你在做自回归 token 生成,对序列中的所有 token 进行 O(N²) 的操作。你的 KV 缓存非常大,因为它不是稀疏的,也不是 top K。所以我们发现更好的做法是,好吧,我们要用另一个模型的层来替换这一层,比如使用 GQA 的层。然后通过正确的训练,你可以让它达到相同的接受率。所以从其他模型移植层是非常可能的,实际上也非常必要。如果一个层效率低下,训练就成了挑战。比如,如何确保你正确地训练它?这又回到了之前的观点,即训练和推理之间的衔接。也就是说,你需要非常好的训练才能实现快速推理。我觉得这越来越成为事实。
Well, to your point previously when you were mentioning like the work that goes into supporting a model when it first comes out, like GLM 52 or MiniMax M3 or whatever the case is. Sometimes you do have to like you do have to switch out some things. Like for instance, the MiniMax M3 had uses full attention. And with full attention, you end up with this like insane bottom that can expect it cuz you're doing auto regressive token generation for three tokens, and you're doing this like O of N squared over all of the tokens that are in your sequence. Um your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're going to replace this you know, we're going to replace this layer with a layer from another model that's using like GQA for instance. And then just with the right training you can get it to have the same acceptance rate. So it is it is very possible to to retrofit layers from other models and very much needed actually. If a layer is like inefficient, the training just becomes the challenge. Like how do you ensure that you train it properly? Which again to earlier points is like the the mesh between training and inference. As in like like you need very good training in order to do fast inference. That's like I feel like more and more becoming true.
是的。你说让它完全准备好投入生产,在支持方面还有什么其他问题吗?
Yeah. Anything else on the support side when you when you say like get it to fully production ready?
是的,我觉得还有一个问题,你知道,我们可以对模型进行相当广泛的测试,但我们想尽快把它推出去,然后你会看到很多人测试它,得到有趣的结果。GLM 曾经短暂出现过一些问题,比如在某些提示和特定温度下,模型会陷入模式崩溃,反复输出同一个 token。一旦你把端点暴露给真实世界,就会有各种各样的输入,你就能发现并修复问题。所以这不是一个一天就能完成的过程。在模型发布后的第一周、第一个月,如果它仍然受欢迎,你既要修复 bug,又要继续推动性能提升。
Yeah, I think that there's also a question of just you know, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was a an issue with um GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to to the real world, there's going to be, you know, so many more varieties of of things given to it that that you're able to, you know, discover and and patch things. So it's not just a you know, day zero process. It's then like for the first week, for the first month if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?
你什么意思,你不想让你的模型输出 SSSSSS 吗?
What do you mean you don't want your model outputting SSSSSS?
顺便问一下,有循环检测机制吗?这种情况仍然经常发生,真是令人惊讶。
Is there loop detection on that stuff by the way? It still happens like quite a lot which is surprising.
我们在端点上有这样的机制,如果模型连续输出同一个 token 四次以上,我们就直接截断生成。我们会说,抱歉,请重试。或者我们会重新处理请求。因为我们知道,如果连续四次输出同一个 token,很可能是崩溃了。
We have like we in our in our endpoint like if a model was to output the same exact token like four plus times, we just cut the generation. We say like sorry, this like try again. Or like we will reprocess the request. Cuz we know then like if it like if yeah, it's four times the same token, it's probably collapse.
是的。有没有办法选择退出,以防我真的想要那样?
Yeah. Is there a way to opt out in case I really actually want that?
实际上,我觉得我们有办法处理。我不太确定,但我觉得在某些模型中,比如它们输出类似表格的东西,它们想画 12 个破折号、12 个破折号。是的,我认为有办法实现。我想我们只对某些 token 这样做。比如我们会排除某些特殊字符。
You're actually I think I think there's a way that we have to handle it. I'm not exactly sure, but I feel like in certain models like when they output something like you can imagine a like a table for instance. And so they want they want to draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.
是的。
Yeah.
我们只对某些 token 这样做,比如 S 几乎是最常见的。J 和 52。
We only do it on like certain like like S is the most common almost. J and 52.
哦。
Oh.
嗯,我想 DeepSeek V4 也有这个问题。就像你遇到循环问题,就像你只是……
Um and I think it was D S V 4 as well. Like you just have like looping issues where like you just have like
是的,S 有什么特别之处吗?不,只是随机。
Yeah, is there a special Is there anything special about S? No, just random
似乎就是那个 token。
seems to be the one token.
是的。呃,而且只在温度为零时出现?不,不,即使在其他数值下也会。
Yeah. Uh and it's and it's only temperature zero or No, no, even at other numbers
0.9 或任何其他值,它仍然会崩溃。这很奇怪,对吧?说实话,这是一个推理问题。就像一个软件问题。比如很多时候,你所在的镜像,比如在视频中,我们会发布一个镜像。如果我们把最新的 TRT LM 镜像的更改上游到我们的技术栈中,我们会发现它修复了这个问题。
0.9 or whatever, it will still it will still collapse. That's weird, right? It's an inference It's an inference problem to be honest. Like a software problem. Like oftentimes um the image you're on like in video we'll release an image for instance. And if we will upstream the changes from the latest TRT LM image into our stack, we'll find that it it fixes it.
或者这种情况通常只会在你使用的推理引擎中出现,比如 Azure 系列。但如果你切换到 V LM,就不会这样。所以这看起来像是一个非常非确定性的软件问题,而不是模型问题。不是权重的问题。比如我们会说:“哦,这是量化的问题,我们 PTQ 做错了。”对吧?但这说不通,因为同样的权重用在不同的推理引擎上就不会重复出现这个问题。有时候是后端使用的内核有非常微妙的竞态条件,如果你把这个模型托管在一个集群上,你永远不会遇到这个问题。
Or oftentimes this will only happen in an inference engine that you're using like Azure line. But if you were to switch to V LM, that isn't the case. So it seems to be like an extremely non-deterministic kind of software issue, not really a model issue. It's not like a weights problem. Like we'll say, "Oh, it's a problem with the quant. We did PTQ wrong." Right? But that doesn't make sense because the same exact weights used with a different inference engine does not repeat the problem. And sometimes it's the kernels that are being used in the back end have these very subtle race conditions where if you were to use this model hosted on one cluster, you will never get this problem.
天哪。
Oh my god.
你把它托管在另一个集群上,就会遇到。原因是那个集群中节点间的 KV 缓存传输使用了比另一个集群更慢的互连。所以这就暴露了竞态。而在另一个集群中则不会。所以最后你只能这样:“好吧,这个模型不能托管在这个集群上,我们要把它托管在另一个集群上,因为那个集群会暴露这个问题。”但最后就变成:是软件的问题,是模型权重的问题,还是硬件的问题?
You host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race. Whereas in another cluster, it doesn't. So then you end up just like, "Okay, this model is not going to be hosted on this cluster, we're going to host it on another cluster. Because that cluster exposes that problem." But then it ends up with like, okay, is it the software, is it the model weights, or is it the hardware?
关于这个,有个说法是温度为零时仍然不是确定性的,对吧?可能是因为硬件。即使在温度为零时,同一个模型,你也不总是得到相同的输出。
There was a thing about this with temperature zero still not being deterministic, right? Possibly because of hardware. Even at temperature zero, same model, you won't always get the same output.
嗯。
Mhm.
但即便如此,我还是对竞态条件感到惊讶,因为我以为 PyTorch 是一个图,至少能保证你按正确的顺序执行。
But even but I'm surprised by the race condition one because I thought PyTorch was a graph that like guarantees that you at least execute things in the right order.
嗯,他们不是那样做的。我不是说存在风险。嗯,你有像 PJ 优化这样的东西,你可以在前一个内核结束之前启动一个内核。你想这样做是因为没有开销。
Well, they don't do it like that. I guess I'm not saying that there's a risk. Well, you have things like PJ optimizations where you can start a kernel before the end of the previous kernel. And that's like you want to do that because there's no expense.
没错,没错。
Exactly. Exactly.
但做得并不干净。你会重叠一点执行。不,我觉得内核本身很可能存在竞态条件,比如那个本应在此时运行的块,那个内核本身就有竞态条件。比如缺少屏障。通常如果你在设计内核时想让它非常快,如果没有广泛测试,某些线程会在其他线程写入之前访问寄存器中的数据点。
But it's not done cleanly. You overlap a little bit of the execution. No, I guess it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Often if you're designing a kernel and you want it to be very fast, if you don't test it extensively, you'll have certain threads access data points from registers before they've been written to by other threads.
嗯。
Yeah.
因为你的屏障错了,或者同步错了。但确实,在这种情况下测试本身非常非常困难。
Because your barrier is wrong or your synchronization is wrong. But yeah, the testing itself is very, very difficult in those cases.
而且没有像借用检查器那样的东西来
And there's no like borrow checker for
那是什么意思?
What's that mean?
嗯,就像 Rust。如果你想实现内存安全,这听起来是一个类似的问题。
Uh like Rust. Like that if you're trying to have memory safety, it sounds like a comparable problem.
嗯,是的,但你是在 CUDA 中工作,对吧?在视频 GPU 中,比如
Well, yes, but you're working in CUDA, right? In video GPUs, like
你只需要一个更高级的语言,比如 Modula。也许这就是 Modula 应该做的。我不知道。
You just need a higher-level language like Modula. Maybe that's what Modula is supposed to do. I don't know.
你怎么看待保持模型质量?所以,你谈到了所有这些步骤:好的,你得做量化,训练你自己的投机解码器,在不同的硬件上运行。看看其他模型提供商,好吧,你在消费端掀起了一场推理速度竞赛,那么如何保持质量一致呢?当然,你可以运行基准测试,但你怎么确定量化的程度或标准,实际上是什么?
How do you see keeping quality of the model? So, you talked about all these steps of okay, you got to do quantization, train your own speculative decoder, run on different hardware. Looking at other model providers, okay, you kicked off an inference speed race on the consumer end, what goes into keeping quality the same across them, right? Sure, you can run benchmarks, but like how do you determine how much quantization or the standards, what actually goes into
关于质量有几点。大多数推理优化是无损的。例如 KV 缓存,你只是重新计算或防止重新计算相同的值。投机当然,如果草稿词元错误,它会被拒绝。主要的有损优化是量化。这归结为第一,数据格式。第二,你选择量化模型的哪些部分,哪些层。第三,对量化后的权重进行大量校准,以确保你保留了所有异常值。不过你还可以做其他技巧。一个重要的就是长上下文。因为你一开始问的是“如果我发送一个 20 万词元的请求会怎样?”所以,显然对于长输入序列,你需要存储更多信息。你需要处理更多词元。因此,即使模型有特定长度的上下文,作为推理提供商,你可能会选择构建一个上下文长度更短的 API。当然,也会有一个完整长度的。因为如果某人不需要完整的百万词元上下文,例如,你可以为他们提供更好的性能。我不知道这是否完全是模型质量。我思考质量的方式是,我们在多大程度上忠实地服务于原始模型?如果你想到一个模型的黄金实现,它完全按照模型设计的方式执行,我认为质量就是我们离那个 100% 保真度有多近。当然,你也可以从训练的角度考虑质量,以及如何超越 100%?但当我想到纯粹的推理优化时,就是在尽可能接近 100% 保真度的同时变得更快。当然,我们内部的标准是,你应该无法区分我们的 API 和官方 API 之间的区别。我认为 Kimmy 在供应商基准测试方面做得特别好。
There's a few things on quality. Most inference optimizations are lossless. KV caching, for example, you are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to number one, data format. Number two, which parts of the model you choose to quantize, which layers. And number three, doing a lot of calibration on the quantized weights to ensure that you're preserving all the outliers. There's other sort of tricks that you can do though. A big one is long context. Because one thing you asked at right at the beginning is "Oh, what's going to happen if I send a 200,000 token request in?" So, obviously with a long input sequence, you need to store a lot more information. You need to process a lot more tokens. And so, even if a model has a context of a certain length, you might as an inference provider choose to build an API with a shorter context length. And of course, a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a sort of golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, you know, 100% fidelity of the model. You can also, of course, think about quality from the training side and how do you push yourself past 100%? But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that you should not be able to tell the difference between our API and a, you know, sort of official API. I think Kimmy in particular does a good job of vendor benchmarking here.
是的,他们发布了一个实际的
Yes, they released an actual
没错。
Exactly.
因为他们指责了一些人,亚马逊。有一个提供商在 Kimmy 的测试上做得不太好。
Because they accused some people, Amazon. There was some provider that was not doing it very well on Kimmy's
是的,所以
Yeah, so
它反映了,可能大概
It reflected it reflected maybe probably
这是很久以前的事了吧?
This was a long time ago, right?
不,就像三四五个月前。
No, like three four five months ago.
这也发生过,我不记得是哪个模型了,但他们撤下了不少,然后他们开始了一个完整的图表。可能是
This also happened with I don't remember which model, but they pulled out quite a few and then they started a whole chart about this. It might have been
Kimmy 供应商验证器。
Kimmy vendor verifier.
嗯。
Yeah.
嗯,因为首先,如果我是消费者,我使用亚马逊的端点,我用了 Kimmy,我会觉得,天哪,这太糟糕了。我不会说亚马逊量化模型的方式不好。我会说 Kimmy 很烂,对吧?所以看起来
Well, because first of all, if I'm a consumer and I'm using like Amazon's endpoint for instance, and I used Kimmy and I'm like, oh my god, this is bad. I'm not going to say oh Amazon quantized the model in a bad way. I'm going to say oh Kimmy sucks, right? So it seems like that
是的,他们在乎。他们在乎。
Yeah, they care. They care.
嗯
Um
有道理。
Justifiably.
是的,只是有趣。
Yeah, just fun.
嗯,这可能是个愚蠢的问题,但只是确认一下。量化有什么改进吗?比如量化总是严格更差吗?
Uh this is probably a stupid question, but just checking. Has anything improved from being quantization? Like is quantization always strictly worse?
不。
No.
嗯,从技术上讲,量化是有损的,是一种有损实现。
Well, technically it's a lossy quantization is a lossy implementation.
速度有提升吗?
Is speed improved?
显然从来不会……
It obviously never like...
我一直在寻找逆向缩放定律。这是我从 Noam Brown 那里学到的,通常朝一个方向作用的事物,有时它们会……
I always look for inverse scaling laws. This is something I learned from Noam Brown, where things that normally act in one direction, sometimes they...
嗯,严格来说,当你运行基准测试时,因为这些模型是非确定性的,有时 FP4 量化版本会比精确版本高两个基点。是的,这就在……这就是为什么我总是说在误差范围内。我其实已经不再这么说了,因为大家都以为我的意思是,在某个误差范围内,我们勉强处于最差的那一端,所以我们在说……但没错,有时它只是给你一个更高的输出分数,但就像 Ali 说的,那是噪声。据我所知,你并不一定是在让结果变得更好。你只是在努力,再次强调,让你的保真度尽可能接近原始模型的 100%。
Well, technically when you run a benchmark, because these models are non-deterministic, sometimes the FP4 quant is two basis points higher than your exact. Yeah, it's within... That's why I always say within margin of error. And I actually stopped saying that because everyone assumes that what I mean is well, within some margin of error we're barely inside of that to the worst, so we're saying but yeah, sometimes it's just like gives you a higher output score, but like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to again, like, keep your fidelity as close to 100% to the original model.
关于你的观点,我们在 MP 上做过研究。我不知道你是否能调出我们发的一条推文。我们的研究实习生之一 Joshua,我想是那条关于我们比 Nvidia 的量化 JLN-52 好 20% 的推文。基本上,我们通过这两个月的研究发现:量化是有损的。你是在把数据从占用 16 位压缩到占用 4 位,比如说。所以,你显然会丢失一些信息,而你在努力最小化这种损失。所以,当我说我要量化模型时,我的工作就变成了如何找到可以量化的层,以及如何找到不能量化的层。例如,对于图像模型,我不会量化调制层,也不会量化输出投影,因为这两者中,输出投影是用户看到的,调制是模型看到或理解的。对吧,所以,我想对他的论文,你有没有……我想它没有……是的,这是一篇很长的论文。我不知道我能不能找到它。
There is, to your point, research that we did on MP. I don't know if you are able to pull a tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized JLN-52 than Nvidia. Essentially, what we found through this two-month research is: quantization is lossy. You're compressing the data from occupying 16 bits to occupying four bits, for instance. And so, you're obviously losing some information and you're trying to minimize that. And so, when I say that I'm going to quantize the model, my job becomes how do I find the layers that I can quantize and how to find the layers to not. For instance, with image models, I don't quantize modulation layers and I don't quantize out projections because those two are like out projection is what you see as the user, modulation is what the model sees or understands. Right, and so, I guess to his paper, do you have the... I guess it doesn't have the... Yeah, it's a long paper. I don't know if I can find it.
可以搜索一下,或者它可能在帖子里。
A part to search or it's probably in the thread.
在帖子里。是的。
In the thread. Yeah.
但基本上,长话短说,量化模型的更多部分很可能让结果更好。如果我有一个模型,我量化了第 1、5、10 层,而另一个模型我只量化了第 1、2 层,那么量化了更多信息的模型有可能表现更好,因为量化误差已经相互抵消了。所以,Joshua 在他的数学证明中展示的,他有一个验证器,就是你可以预测哪些层的量化误差会相互抵消,然后你选择量化那些层。所以,做这种数学量化的结果是,你最终得到一个比其他提供商量化程度高 20% 的模型。所以,你获得 20% 更多的吞吐量,因为有更多层在 N 之前运行。而且你的质量比那个其他量化模型更好,因为你选择量化的那些层的误差相互抵消了。比如一层偏右,一层偏左,一层偏右。你最终的 logit 分布更接近模型的原始分布。所以,你有更好的保真度。我们证明这一点的方式是使用 KL 散度。所以,我们不仅仅是在基准测试上打分,我们还计算了量化模型的 logit 分布与原始全精度模型的 logit 分布之间的 KL 散度。我们展示了用这种技术我们得到……如果你的 logit 切换 token 的概率分布与原始模型更一致,你很可能最终会忠于原始模型。所以,是的。所以,看起来以前,在这之前,业界似乎认为:你量化得越多,结果就越差,因为你引入了更多损失。这并不完全正确。所以,是的。它不会改善,但可以抵消。
But basically, the long and short is it is very possible that quantizing more of the model makes the results better. If I have a model that I quantize layers 1, 5, and 10 and another model where I only quantize layers one and two, it is possible that the model in which I quantize more information is going to perform better because the quantization errors have canceled out. And so, what Joshua showed in his mathematical proof where he had like a verifier is that you can predict which layers are going to have quantization errors that will cancel out with each other and you choose to quantize those layers. And so, the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider. So, you get 20% more throughput because there's more layers that are running in N before. And your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out. Like one layer skewed to the right, one layer skewed to the left, one layer skewed to the right. Your final logit distribution is more similar to the original distribution of the model. So, you have better fidelity. And so, the way we proved this was with KL divergence. So, instead of just scoring on the benchmarks, we scored the KL divergences between the logit distribution of the quantized model and the logit distribution of the original full precision model. And we showed that with this technique we get... If your probability distribution on the logit switch token it wants to select is more of the same as the original model, you're probably going to end up staying true to the original model. So, yeah. So, it seems like previously, before this, it seemed like the industry was: well, the more you quantize, the worse it's going to be cuz the more loss you introduce. That's not exactly not necessarily true. So, yeah. Doesn't improve it, but can cancel out.
我想可能是这个,但这让我想起了剪枝,实际上你可以剪掉某些层。但是,非常有趣。不知道你们还发了这么一篇完整的论文。
I think it might be this, but reminds me a good bit about pruning actually, where you can prune off certain layers. But, very interesting. Didn't know this was a whole paper you guys put out.
有个有趣的事实,这篇论文最初有 72 页。然后我们决定我们不能……我们没法缩减它。所以,现在有 425 页。
It was a fun fact, it was originally 72 pages, this paper. And then we decided we can't tell. We couldn't reduce it. So, it's now 425.
还是有 39 页。非常充实。我们谈到了评估和所有这些事情。那么在加速方面有什么可能性?我想这可能是人们最关心的事情,也是你在文章中写到的。比如官方 API 每秒 70 个 token,你把它提升到了 90。这是常见的事情吗?
Still 39 pages. Very substantive. We talked about evals and all these things. And like what's possible in terms of speedup? I guess like probably the number one thing that people do want to care about and is something that you wrote about in your post. Like official API 70 tokens per second and you push it up to 90. Is that like a normal thing?
在 Inflection 工作的酷之处,我认为 Inflection 会是一个长期做工程的好地方,是因为如果你看看高度优化的领域,比如金融,如果你在金融行业,你会用基点来衡量你提高了多少。就像“哦,我提高了五个基点,也就是提高了 1% 的 1/20。”那是大新闻,因为一切都高度优化了。当我们发布优化成果时,是 20%、100%、200%。所以,老实说,可能还有很长的路要走。当我们开始发布关于如何让某个东西快 1% 的成果时,你就会知道 Inflection 已经基本解决了问题。
So, what's cool about working at Inflection, the reason that I think Inflection is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like say finance, if you're in finance, you measure how much better you got in basis points. It's like, 'Oh, I got five basis points better, like 1/20 of 1% better.' That's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100%, it's 200%. So there's still probably like a lot further to go honestly. You'll know that Inflection is pretty much solved when we search for start publishing about how they got 1% faster at something.
顺便说一下,因为我来自金融背景,在 70 年代,那是当时的利润率,当你做量化金融研究时,你会发现……
Which by the way because I am from the finance background in the 70s that was the margin at the time when you did quantitative finance research you would find...
就像 20%……
And like the 20%...
10%……
10%...
是的。对。现在它很小了,对于……
Yes. Yeah. And now it's tiny for...
对于感兴趣的人,可以看看 Andrew Lowe 的论文。他有一个非常有趣的图示,展示了量化统计套利分布从 70 年代那种 20% 的差异缩小到今天几乎为零。这非常非常酷。
For those people interested, look up Andrew Lowe's paper. He had a really interesting illustration of quant stat arb distribution narrowing down from like those kinds of 20% differences in the 70s down to nothing today. Which is very very cool.
完全正确,我们正处于同类事情的起点。现在基准测试很难。我想任何人都会告诉你,基准测试提供商的速度很难,因为有很多变量会影响它。你使用什么硬件?系统上有多少负载?提示词以及输入和输出序列长度的确切性质是什么,诸如此类。但总的来说,当你开始叠加这些改进时,你看到的是倍数。你可以看到最常见的形式当然是 TPS,即每秒 token 数,这是我们行业里一个糟糕的命名,因为实际上有两种每秒 token 数。有作为吞吐量数字的每秒 token 数,和作为延迟数字的每秒 token 数。比如 GPU 输出的总每秒 token 数作为吞吐量数字。
Exactly and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths all that kind of stuff. But overall when you start stacking these improvements you're looking at multiples. You can look at it the most common form of course is TPS tokens per second which is bad naming by us in the industry cuz there's actually two tokens per second. There's tokens per second the throughput number and the latency number. Like total tokens per second out of the GPU as a throughput number.
大多数人只关心每秒 token 数作为延迟指标,我们应该称之为 ITL,即 token 延迟,但我们不这么叫。总之,你可以想象一个没有太多优化的标准 API,对于一个 1 万亿参数的模型,在合理的流量配置下,运行速度大约在每秒 30 到 50 个 token 的范围内。我们通常看到的目标是将其提升 10 倍。但不一定非得是 DaisyLow,而是通过叠加足够的优化,比如你有四个优化,每个都能让性能翻倍,或者抱歉,三个优化每个翻倍,那么叠加起来就是 8 倍的提升。这就是我们在这个领域努力的数量级。我们试图让事情大幅加快,而不仅仅是从 70 提升到 90。
Most people only care about tokens per second as the latency number, which we should call ITL, it's token latency, but we don't. Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for a reasonable traffic profile. And we generally see the goal of pushing to 10x that. But not necessarily DaisyLow, but by stacking enough optimizations, if you have say like four optimizations each of which doubles performance, or sorry, three optimizations each of which doubles performance, then you stack that up, that's an 8x gain. That's kind of the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.
你是说你们已经做到了?
Are you saying you have done that?
假设你有一个合理的基线,每秒 30 或 40 个 token。你可以实现 10 倍的提升。比如在 GLM 5.2 上,如果你想获得未量化的版本,也许甚至在 Happos 上,而且你只是使用现成的推理引擎,没有特别的优化,没有投机解码,没有 KV 路由之类的额外功能,也没有分离。你可能会看到每秒 30 到 40 个 token。你觉得这是一个合理的基线吗?
So let's say you have as a reasonable baseline 30 or 40 tokens per second. You can achieve 10x that. So like on GLM 5.2, if you want to get unquantized, perhaps on Happos even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation. You probably are looking at that like 30 to 40. Do you think that's a reasonable baseline?
对,对。
Right. Right.
要达到 10 倍的提升,你需要做出很多权衡。如果我们运行在每秒 300 到 400 个 token 的范围内,显然你使用了最好的硬件。你有一个优化的投机解码器,你完成了所有的量化工作。你的缓存命中率很高。你以相当小的批处理大小和针对延迟与吞吐量调整的并行配置运行。但这是可能的。所以你在 Artificial Analysis 或 OpenRouter 上看到的差距,从最差的提供商到最好的提供商,往往可以达到那种范围。10 倍当然非常激进。通常可能更像是 4 到 6 倍的提升。但正是当我们能获得这些巨大收益时,而不是仅仅从 70 到 90 个 token,才让我们真正兴奋。
To get to something like 10x, there's a lot of tradeoffs that you're making. If we're running at sort of more like a 300-400 tokens per second range, obviously you are using the best hardware possible. You have an optimized speculator, you have done all of your quantization work. You are seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput. But it is possible. So the spreads that you see if you go on Artificial Analysis or you go on OpenRouter and you look at the worst provider to the best provider, often times can hit that kind of range. 10x is of course very aggressive. It's often times maybe more of a 4 to 6 times improvement. But that's the kind of performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.
这也取决于硬件。比如你显然有这种情况,你在一个 H100 节点上提供服务,然后你把你的短模型分布在四个 B200 节点上,你肯定可以通过投入更多硬件来提高速度。但如果是同样的硬件和相同数量的 GPU,那就另当别论了。
It's also like hardware dependent. Like if you obviously have a thing where you're serving it on just like a node of H100s and then you throw like your short model across like four nodes of B200s, you can definitely increase the speed with just throwing more hardware at it. But like normalizing for the same exact hardware and the same number of GPUs.
是的,那么你可能会看到 2 到 4 倍的提升,具体取决于推理优化。所以,部分取决于需求,部分取决于驱动者。
Yeah, then you're looking at like a two to four x improvement depending on the inference optimizations. So yeah, some of it's what's the call and some of it's who's the driver.
如果你分解这 2 到 4 倍的提升,比如在 B200 单节点上运行 GLM 5.2,对吧?为了榨取最后一点性能,成本权衡或努力是什么,而人们应该怎么想?
If you break down the two to four x, say the example is run GLM 5.2 on B200s single node, right? What's like the cost tradeoff or effort to get like the last bit of juice out versus what should people just think of right?
投机量化和投机解码。
Speculative quantization.
那大概占 95%。
That's like 95%.
那能带来多大的提升?对普通人来说有多容易做到?比如现在我想把 GLM 5.2 的权重放到一个 B200 节点上。找到投机解码器模型或已经量化的模型有多容易?需要做多少工作?
And how far does that get you? And how easy is that for the average person to do? So say right now I want to throw the weights of GLM 5.2 on a node of B200s. How easy is it to find a speculative decoder model or already quantized model? How much work goes into it?
如果你从头开始做,那是相当多的工作。如果你今天做,会有人已经发布了东西,你可以直接获取一些 NVFP4 权重。你可以获取一个投机解码器。是的,如果我们想想我们叠加的 2 倍是什么,从 BF16 到 NVFP4 并不是完全的 2 倍,对吧?我认为从 16 位到 8 位大约有 30% 到 40% 的提升,然后从 8 位到 4 位再乘以 30% 到 40%。所以这并不能完全达到 2 倍,但大约接近 2 倍。投机解码器大约 2 倍。在此基础上,如果你有足够的硬件和足够的流量,分离式推理又能带来大约 2 倍。然后你再加上一些两位数的百分比提升,来自更好的运行时,比如最新的内核等等。这就是叠加的方式。所以构建这些,比如构建量化权重,对于真正懂行的人来说,需要几个小时到几天的工作。构建投机解码器,同样需要几个小时到几天。分离式设置也需要几个小时到几天。好吧,一旦你有了... 是的,是的。第一次让分离式工作,当然是非常困难的,但边际实现是...
If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about like what are the 2x's we're stacking, going from BF16 to NVFP4 is not quite a 2x, right? It's like I think it's about 30 to 40% from 16 to 8 and then another 30 to 40% multiplied from 8 to 4. So that doesn't quite get you a 2x, but like roughly a 2x. Speculator, roughly a 2x. Disaggregation on top of that, if you are able to get enough hardware and put enough traffic through it, another roughly a 2x. And then you add in some double-digit percent increase from having just a better runtime with the latest kernels and stuff behind it. And that's kind of how it stacks up. So building each of those, like building the quantized weights, is for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like hours to days of work. And the disaggregation setup, hours to days. Well, okay. Once you have a... Yeah, yeah. Getting disaggregation working for the first time, I'm saying, of course, is very difficult, but the marginal implementation is...
如果你只是获取,比如你是一个普通人,一个普通消费者,可以使用一个 B200 节点,你想知道,我怎么能自己托管它?你不需要自己量化模型。总会有开源的量化检查点,如果没有其他人发布,供应商也会发布一个。通常,提供商也会有他们自己训练的投机解码技术。你不需要训练自己的投机解码技术,你也可以直接使用那个。
If you're just grabbing, like if you are a person, just a normal consumer who has access to like a node of B200s, and you're wondering, how can I just host it myself? You don't need to quantize the model yourself. There's always going to be like an open-source quantized checkpoint, and a vendor is going to push one out if no one else does. Usually, the providers will have their own spec tech that they've trained as well. You don't need to train your own spec tech, you can just use that as well.
是的。比如 Kimmy GLM 5.2 有自己的 MTP。
Yeah. Like Kimmy GLM 5.2 has its own MTP.
对,对。多 token 预测。
Right. Right. Many maps.
多 token 预测。
Multi-token prediction.
是的。
Yes.
我只是替你说,以防我说错,你知道的。我可以替你说,以防我说错。
I'm just saying it for you in case I get it wrong, you know. I can do it for you in case I get it wrong, you know.
实际上,如果我说错了你应该纠正,但他们的多 token 预测可以用于自投机解码。
Actually, you should correct if I'm wrong, but their multi-token prediction can be used for self-speculative decoding.
我其实不确定。
I'm actually not sure.
好的。
Okay.
我对此半信半疑,但有人可以查证。但描绘这样一个故事是有用的:不只是普通人,比如一家公司想从无服务器推理切换到,我想把它部署到,你知道,我想租一些 GPU,把它部署上去。这些是你采取的步骤,比仅仅把它放在 vLLM 后面要快得多。
I'm semi-confident in it, but someone can check. But it's useful to paint the story of okay, not just the average person, but say a company wants to switch from serverless inference to I want to throw this up on, you know, I want to rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.
对。我一直在等你们提到 Dynamo。我觉得那应该是你衡量的基线。
Right. I was waiting for a mention of Dynamo. I feel like that's supposed to be the baseline that you measure against.
我认为 Dynamo 与其说是一个开箱即用的系统,不如说是一个用于构建的工具包。所以当我们谈论 KV 路由、KV 卸载、PD 分离时,Dynamo 基本上... 顺便说一下,Dynamo 是 Nvidia 的一个开源库。
I would think of Dynamo as less of a sort of out-of-box system and more of a toolkit for building with. So when we talk about doing KV routing, when we talk about doing KV offloading, when we talk about doing PD disaggregation, Dynamo fundamentally is... By the way, Dynamo is an open-source library from Nvidia.
我们已经和 Kyle Klein 讨论过那部分了。
We've done that part with Kyle Klein.
酷。那么你的听众就会知道它支持所有不同的推理框架,而且它实际上是多硬件的,这很有趣。
Cool. So then your listeners know then that it supports all the different inference frameworks, and it actually is kind of multi-hardware, which is interesting.
但它只是一个路由器。它不像一个优化器。
But it's just a router. It's not like an optimizer, though.
是的。
Yeah.
它所做的,就像 Dynamo 擅长的那样,是一个在你的集群、你的硬件之间移动信息的库。所以,如果你有,你知道,KV 缓存在一个地方,你需要它到别处,Dynamo 会协调 Nixel 帮你移动它。这并不意味着开箱即用,你只要说,你知道,pip install Dynamo,然后就能获得巨大的性能提升。它更像是一个开发者工具包。
All it does, like what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So, if you have, you know, KV cache on one place and you need it to be somewhere else, Dynamo coordinates Nixel for you to move that around. That doesn't mean that out of the box you just say, you know, pip install Dynamo, and then you get a massive performance speedup. It's more of a developer toolkit.
是的。我本来会说它带有一组默认设置,你可以随后替换掉。
Yeah. I would have said it comes with a set of defaults that you can then swap out.
确实如此。如果整个行业,我想,都在标准地推出所有这些部署,那么我认为它会是一个可信的基线,但我们必须针对我们在实际中看到的情况进行基准测试。
It does. If the industry at large, I think, was like rolling out all of these deployments standard, then I think it would be like a credible baseline, but we've got to benchmark against like what we're seeing in the wild.
我确实想多谈一点关于 PD 分离,因为那可能是继量化和投机解码之后的第三点。不过在你的书里,我正要把书拿出来。是的。比如 5.2.2 节关于 Medusa,5.2.3 节关于 Eagle,5.2.4 节关于 Indian。5.5 会是,嗯,会是分离。
I did want to talk a little bit more about PD disagg, because that's probably like number three after Quantize and Speculative Decoding. In your book, though, I was just going to pull out the book. Yeah. Like section 522 on Medusa, 523 on Eagle, 524 on Indian. 55 would be Mhm. would be disaggregation.
嗯,不,我只是想稍微多谈一下其他的。那么,你选择包含什么?你选择不包含什么?因为我想,还有所有这些其他技术。
Well, no, I just wanted to dwell a little bit on the other. So, what do you choose to include? What do you choose to not include? Because there were all these other techniques, I guess.
是的。
Yeah.
这些仍然相关吗?因为我认为它们可能是在大约一年半前出现的。
Are these still relevant? Because I think they came out like a year and a half ago, maybe.
Medusa 相当老了。
Medusa is quite old.
是的,Medusa 很老了。
Yeah, Medusa is old.
但它在书中是作为一个好的基线,但你知道,你应该知道这个。就像我读了论文,然后我想,“是的,这太有道理了。”
But is it in the book as a good baseline, but you know like you should know this. Like I read the paper and I'm like, "Yeah, it makes so much sense."
是的,所以写这本书我有几个目标。一个是给人们提供整个领域的实用词汇,另一个是让他们对这些技术各自如何运作有一些直觉。正如我在 AI Engineer 演讲中提到的,这算是这本书的第一个公开附录,投机领域的发展比其他一切都快得多。所以,是的,即使在我写这本书的时候,Medusa,我非常多地把它作为让人们理解这个领域如何演变的方式,而不是什么是最现代的技术。现在,当然,有 Deflash、Despoke,还有比 Eagle 更新的技术,尽管 Eagle 仍然非常常用。
Yeah, so with the book I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is kind of the first public addendum to this, the speculation space has moved much faster than everything else. So, yeah, even at the time that I wrote the book, Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now, of course, there's Deflash, Despoke, there's newer techniques even than Eagle, although Eagle is still very commonly used.
Sparkspecter。
Sparkspecter.
是的,Sparkspecter 或投机解码。
Yes, Sparkspecter or speculative decoding.
什么?你能……
What? Can you...
这是 Trudau 的论文,基本上就是在做投机解码。
It's the paper by Trudau, and it's basically doing speculative decoding.
嗯哼。
Uh-huh.
为投机……
for the speculative...
天哪。
Oh my god.
这真的只是另一个。就像,是的,这是最重要的解释方式。而且他似乎在那里获得了不小的加速,但训练的复杂性似乎。至少在我们看来,几乎和训练 GAN 一样复杂。就像一个非常微妙的平衡,而且很多时候你……就是不值得。但是的,这真的是在投机之上做投机解码……
It's literally just another. It's like, yeah, that's the most important to explain it. And it seems like he got non-trivial speedups there, but it seems that the complexity with training. It's almost like, in our mind at least, it's almost as complex as training GANs. Like a very delicate balance, and often times you... It's just not worth it. But yeah, it's literally speculative decoding on speculative...
投机投机。
speculative speculative.
是的,我们看过这篇论文。
Yeah, we saw this paper.
这很有趣,对吧?我甚至不会期望它训练起来很特别。
It's interesting, right? I wouldn't even expect it to be very particular to train.
对,对。
Right, right.
我天真的部分想,“好吧,训练投机解码器。”
The naive part of me is like, "Okay, train speculative decoder."
但是的,这有道理。就像投机解码的整个想法是,你……就像 iPhone 的自动预测版本,但针对普通模型,对吧?就像你只是生成三个词元,然后你想,“好吧,对它们做预填充。”所以你为你的原始模型节省了那三个回合。现在你的投机解码器在做三个回合的自回归。那么为什么不干脆用一个更小的模型来预测那些词元呢?
But yeah, like it makes sense. Like the whole idea of speculative decoding is you... It's like the iPhone auto-predict version, but for normal model, right? Like you're just generating three tokens and you're like, "Okay, do pre-fill on them." And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto-regression. So why not just have an even smaller model predicting those tokens?
另一个问题是投机者的大小是多少?比如说对于 GLM……
The other question there is what are the size of speculators? So say for GLM...
对。就像十亿参数。比如对于 MiniMax 来说……是的,是的,就像一层。大约是原始模型的 1/60。
Right. It's like a billion parameters. Like for MiniMax it's... Yeah, yeah, it's like one layer. It's like 1/60 of the original model.
是的。实际上,我想我们回到办公室后应该写一篇论文。投机投机投机解码。
Yeah. Actually, I think we should do a paper when we get back to the office. Speculative speculative speculative decoding.
不,这确实看起来像,你什么时候停止?但然后它也看起来像,如果你能够训练投机解码,比如说,对吧?就像如果你能有一个小模型,准确预测中间投机者会预测什么,而它能预测原始目标模型会预测什么,那么为什么不直接使用那个最小的模型呢,对吧?
No, it does seem like, how when do you stop? But then it also seems like, if you're able to train spec decode for instance, right? Like if you're able to have a small model that accurately predicts what the intermediate speculator is going to predict, that is able to predict what the original target model is going to predict, then why not just use that smallest model directly, right?
是的,这是一个路由……
Yeah, this is a routing...
这是一个路由问题。
It's a routing problem.
对。是的,对。
Right. Yeah, right.
关于投机者,使用它们的实际约束之一是,你必须在运行大模型的同一硬件上运行一个小模型。这本身就存在编排和资源竞争的问题。这是对投机的一般性约束之一,即草稿词元需要资源来创建,需要软件复杂性来管理。所以如果你有某种无限递归的投机者,你会在推理引擎的实际实现中增加相当多的复杂性,而不仅仅是在训练过程中。
The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is an orchestration and resource competition problem inherent in that. And that is one of the sort of constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have a sort of infinitely recursive speculators, you're adding quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.
我正要说,我想知道你是否可以做类似的蒸馏和剪枝,你知道,这是同一回事。它只是一个模型。我们不能蒸馏掉很多权重,量化投机者,但这超出了我的领域。我想出现的问题还有,这都是针对大型服务器工作负载的,对吧?其中有多少适用于,比如我有这台 MacBook,我想非常高效地运行 Gemma。类似的问题,但不相同?
I was going to say I would wonder if you could do similar like distillation and pruning of, you know, it's the same thing. It's just a model. Can we not just distill a lot of the weights, quantize the speculator, but out of my domain. I guess the question that also comes up is this is all for big server workloads, right? How much of this applies to say I have this MacBook, I want to run Gemma really efficiently. Similar problems, not the same?
相当不同。几周前我在 Sal 的播客上和他谈过这个。数据中心和生产工作负载的推理工程与本地 AI 的推理工程之间的区别在于,我们从根本不同的约束和目标开始。对于本地 AI,问题是如何把这个模型装到我的硬件上,然后让它不那么笨?而对于数据中心推理,问题是如何加载这个模型,然后让它不那么慢。显然,你知道,我们关心不那么笨,他们关心不那么慢,但本地 AI 推理工程生态系统,我认为实际上有很多值得我们数据中心领域学习的地方。他们是各种量化形式的专家,包括动态量化,而我们在剪枝、蒸馏、层移除方面几乎不碰这些。
Pretty different. I talked to Sal about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI is that we start with fundamentally different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center inference, it's how do I load this model and then make it less slow. And obviously, you know, we care about less dumb and they care about less slow, but the local AI inference engineering ecosystem, I think actually has a lot for us to learn from in the data center space. They are experts in various forms of quantization including dynamic quantization that we just kind of don't touch in the pruning, in the distillation, in the layer removal.
他们较少移除层。
They remove layers less.
是的,他们确实不太做剪枝。
Yeah, they don't do as much pruning really.
是的,嗯,但他们确实做那个。
Yeah, well but they do do that.
令人惊讶,对吧?但那是另一回事。
Surprising, right? But that's a different thing.
只是为了在笔记本电脑上装点东西。
Just to fit something on the laptop.
对,对,对。
Right, right, right.
所以,是的,我的意思是这是一个有趣的空间。不一定是因为他们的技术对我们来说在数据中心里有意义,因为显然我们有不同的资源和不同的目标,而更多的是那个领域的过程以及开放性值得钦佩。
So yeah, I mean it's an interesting space. Not necessarily that their techniques make sense for us to do in the data center, because obviously we have different resources and different goals, but more that the process as well as the openness of that field is something to admire.
是的,就像你说的,某些优化,比如 Turbo Quant,如果你熟悉的话。它在 next 上引起了巨大的轰动,我们在 Twitter 上做了深入的探讨,我当时想,这是什么?它如何工作?为什么好或不好?然后它火了,并在本地设备上实现,因为例如在 MacBook 上,你的内存带宽非常慢。但试着把同样的东西放在 Nvidia GPU 的 B100 上,TurboQuant 就不会被使用。Nvidia 明确表示这不是一个好的优化,我们亲眼看到,在 TurboQuant 8 位内核中,进行反量化和量化的开销实际上比从带宽中节省的时间要慢得多,因为在 B100 上你有大约每秒 3.5 TB 的带宽。你不需要减少那么多存储。你不需要做 FP4 KV 缓存。你不需要使用 TurboQuant。有更好的优化可以做。但在边缘设备上,它极其重要,极其有用。所以看起来那里有不同的优化。但它们都独特地结合在一起,比如,哦,你想量化模型,你想做稀疏编码,以及某些常见的
Yeah, like to your point, certain optimizations like Turbo Quant, if you're familiar. It made such huge hype on next, and we did a whole deep dive on Twitter, and I was like, what is it? How does it work? Why is it good or not? And it took off and was implemented on local devices because your memory bandwidth is so slow on a MacBook, for instance. But try putting the same thing on an Nvidia GPU on a B100, TurboQuant would not be used. Nvidia made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization quantization in the kernel itself for the TurboQuant kernel 8-bit is actually much slower than the time you save from the bandwidth, because on a B100 you have like 3.5 terabytes per second. You don't need to decrease the storage that much. You don't need to do FP4 KV cache. You don't need to use TurboQuant. There are better optimizations to be made. But on edge devices it's extremely important. It's extremely useful. So it seems like different optimizations there. But then they're all uniquely combined with like, oh, you want to quantize the model, you want to do sparse coding, and like certain common
前缀原则
prefixes principles
是的,完全正确,完全正确。
Yeah exactly exactly exactly.
他们还在模型并行方面做了很多工作,尤其是在异构拓扑上,你有一些节点,它们通过以太网 DGX 节点连接在一起。
They also do a lot of work on model parallelism, especially over heterogeneous topology where you have some spokes and they're wired together with Ethernet DGX spokes.
是的,是的,这是 Axle Labs 的伙计们。
Yeah yeah this is the Axle Labs guys.
是的,你有一堆 Mac mini 堆叠起来。
Yeah you have a number of Mac minis stacked up.
机器之间的互连,这就是为什么我们经常做的一件事是使用张量并行,即使用所有 8 个 GPU 并将模型分片到其中。张量并行不适合本地 AI,因为它假设像 NVLink 这样的高带宽互连。他们可能被迫做类似流水线并行的事情,除非我们做某种多节点推理,否则我们永远不会这样做。
There's the interconnect between machines, which is why one thing that we do a lot is work with tensor parallelism, and that's where you are using all eight GPUs and sharding the model across it. Tensor parallelism is not a good fit for local AI because it assumes a very high bandwidth interconnect like NVLink. They might be forced to do something like pipeline parallelism, which we're never going to do unless we're doing some kind of multi-node inference.
既然你提到了,我其实不确定我们是否会涵盖它,但让我们简要解释一下张量并行和专家并行,因为你有很好的图片。
Since you mentioned it, I actually wasn't sure if we're going to cover it, but let's briefly explain tensor parallelism and expert parallelism since you have very nice images.
书?
the book?
是的,是的,我们开始吧。我只是想炫耀一下你的图片。
Yeah, yeah, let's go. I just want to show off your images.
是的,感谢 Baseten 设计团队的 Luke 制作了这些漂亮的图片。哦,在我们进入正题之前,还有另一个区别。我们经常谈论混合专家模型的激活参数,对本地推理的人来说这很重要,因为如果你的批大小为 1,你只激活那么多参数。当我们做
Yeah, shout out to Luke from Baseten's design team for making these beautiful images. Oh, that's actually before we get into this, just one other difference. We talk a lot about the active parameters of a mixture of experts model, and for local inference folks that matters a lot because if you have a batch size of one, you're only activating that many parameters. When we do
是的,我正想把那点带入讨论中。是的。
Yes, I was going to bring that into the discussion conversation. Yeah.
是的,当我们处理一个 MOE 模型并为 API 托管它时,我们假设所有参数都会被激活,因为你会覆盖所有内容。酷。所以,广义上,张量并行可以用于任何模型。专家并行只能用于 MOE 模型。实际上,今天所有足够大以至于你需要在多个 GPU 上并行化的模型都是 MOE 模型。所以那个细微差别现在不那么重要了。对于专家并行,其思想是将整个专家放在一个 GPU 上。通常专家数量多于 GPU,所以你可能每个 GPU 放 N 个专家,比如每个 GPU 放 8 个专家之类的。然后你在每个 GPU 上复制路由器,路由器非常小。然后通过将生成从专家移动到专家,每个专家都在 GPU 内部,它们不会竞争资源,你大幅提高了吞吐量,GPU 到 GPU 的连接不那么重要,因为通信量不大。张量并行要求你能够进行类似全收集全归约的操作。所以你基本上将模型完全分片到 GPU 上,然后每一步你都要合并每个 GPU 的结果,这就是为什么互连很重要。当然,这是一个非常高级的概括。在很多地方这并不正确,但通常 TP 有助于降低延迟,在许多情况下,你会在模型中使用这两种并行的组合,而不是只选择一种。你想补充一些色彩吗?
Yeah, when we go through a MOE model and we host it for an API, we assume that all parameters are going to be active because you're going to hit everything. Cool. So, broadly tensor parallelism you can do with any model. Expert parallelism you can only do with MOE models. Effectively all models today are MOE models that are large enough that you would care to parallelize them across multiple GPUs. So that nuance is less important now. With expert parallelism, the idea is you put the entire expert on a GPU. Generally you have more experts than GPUs, so you might put like N experts per GPU, like eight experts per GPU or whatever. And then you replicate the router, which is very small, across each of the GPUs. And then by moving the generation from expert to expert, with each expert being inside a GPU, they're not competing for resources, you massively increase the throughput that you're capable of doing, and the GPU to GPU connection is not as important because there's not as much communication. Tensor parallelism requires that you are able to do this like all gather all reduce. So you basically shard the model across the GPUs entirely, and then for each step you're combining the results of each of the GPUs, which is why the interconnect matters a lot. And it is generally, of course, this is a very high-level generalization. There's a lot of places where this is not correct, but generally TP is helpful for latency, and in many cases you will use some combination of these two parallelisms across the model rather than just picking one or the other. You want to add some color there.
比如在一个模型中,它们不是互斥的。你有张量并行,也会做专家并行。流水线并行较少单独使用,在我看来我们从不使用 PPN
Like in a model, they're not mutually exclusive. You have tensor parallelism and you'll do expert parallelism. Pipeline parallelism less solely, it seems to me like we never use PPN
是的,你不得不做流水线并行的唯一原因,也就是你将不同层分开,把一半层放在一个硬件上,另一半放在另一个硬件上,是因为你被迫进行多节点推理,因为模型比你必须处理的要大。假设你出于某种原因在 H100 上部署,并且你在上面放一个万亿参数模型。你必须使用多个 H100 节点,由于节点之间的互连非常慢,那里唯一可行的并行方式是流水线,但然后你会在每个节点内做专家并行和张量并行。
Yeah, the only reason you would have to do pipeline parallelism, which is where you separate different layers and put half the layers on one hardware and half on another, is if you are forced to do multi-node inference because a model is bigger than you have to. Let's say you're doing a deployment on H100s for whatever reason and you're putting a trillion parameter model on there. You have to use multiple nodes of H100, and so because the interconnect is so slow between the nodes, the only viable way to parallelize there is pipeline, but then you would do expert and tensor within each node.
H100 的限制因素是 HBM 吗?
And the limiting factor for H100s is HBM?
是的,它们就是不够
Yeah, they just don't have enough
多少?我们需要的神奇数字是什么?
How much? What's the magic numbers that we need to
比如在 B100 上,每个 GPU 是 180 GB,然后一个 8 卡节点,你谈论的是 180 * 8。而 V100 每个参数占用半字节,所以那是 800 GB。在 H100 上大约是 140?
Like on a B100 it's 180 gigabytes per GPU, and then a node of eight you're talking like 180 * 8. And V100s each parameter takes half a byte, so that's 800 gigabytes. On a H100 it's like 140?
是 80。
It's 80.
是的。是的,你看我多老了?我老了,我干这行很久了。我确实记得 H100 的规格。是的,不,所以关于 T4 的一件事。让我告诉你当年在 T4 上运行模型是什么感觉。
Yeah. Yeah, you see how old I am? I'm old, like I've been doing this a long time. I actually remember H100 specs. Yeah, no, so one thing about the T4s. Let me tell you what it was like to run a model on a T4 back in the day.
嗯,有一件事让我惊讶的是,更多人没有做 Jamba。我不知道你们是否记得 AI 21 的 Jamba。
Well, one thing I was surprised to see that more people didn't do Jamba. I don't know if you guys remember Jamba from AI 21.
他们实际上会专门挑选硬件,然后针对硬件设计架构维度,显然会充分利用硬件。这很合理。但不知为何,所有这些模型都没有这样做。
They would actually specifically pick a hardware and then they designed the arc dimensions for the hardware and then they would obviously saturate the hardware. Like it makes sense. And like somehow all these models don't do that.
不过他们在训练端不是这样做的吗?
Don't they do this for the training side though?
我不知道。
I don't know.
什么?抱歉。
The what? Sorry.
训练,就是决定用哪个 GPU。
Training for training, like deciding which GPU to use.
是的,嗯,怎么……是的,是的,他们确实这样做。训练更像是一个数学问题,你可以通过计算来最大化浮点运算。推理则更像是一种自动调优。不知道你熟不熟悉 GPU 内核自动调优。基本上就是,你定义说我有两块 GPU,我可以做 TP1、TP2、EP1、EP2 之类的组合,对吧?这样就有大概两个平方的组合,然后你就像影子一样跑同样的流量,看看哪种配置能带来最好的 TPM、TPS,然后就用那种。我不喜欢的是,你无法推理哪种配置会带来最佳性能,也不存在一个总是最优的固定配置。但似乎自动调优就是找到最佳配置的方式。对于内核和 GPU 内核来说也是如此,设计好内核和配置之后——启动多少线程,用多少共享内存——你就自动调优,扫描参数空间,凭经验决定哪个最好。但确实,它们是结合在一起的,不是分开的。
Yeah, well how to... Yeah, yeah, they do. And with training it's more of like a math. Like you can run the math and say the flops and maximize it. With inference, it's more of like an auto tuning. I don't know if you're familiar with GPU kernel auto tuning. But like it's basically like you define that oh I have two GPUs. I can do TP1, TP2, EP1, EP2 for instance, right? And so that gives you like sort of like two square combinations and then you just like you shadow the same traffic like real full traffic and you just see which configuration gives you the best TPM, TPS and then you just use that. I don't like the fact that it's you cannot reason about which one's going to give you the best performance or that there isn't one specific configuration that's always best. But it seems like auto tuning is just the way that you find the best one. And with kernels and GPU kernels it's much of the same after you you design your kernel and you you design your configuration. How many threads do you launch? How many you know, how much shared memory do you use? You just you just auto tune. You just sweep the parameter space and decide on this is the best one empirically. But yeah, but they are they are combined. They're not just like separate.
训练中有一些部分有点像针对硬件优化的。比如你看 Nvidia 的 Nemo 模型,它们在 Blackwell 上运行得非常好,这并不意外。所以确实有一定程度的针对性,但我认为大多数开放实验室都在努力让模型能在尽可能广泛的硬件上运行,而不是只针对单一芯片。
There's a few bits of training that are kind of like hardware targeted. If you look at for example Nvidia Nemo models, they run very, very well on Blackwell. That's that's unsurprising. So there's some degree of that, but I think that most open labs are trying to make models that can be run on as wide of hardware as possible rather than targeting just like a single chip.
我明白了,为了实用性。
I see. For usefulness.
是的。
Yeah.
呃,好的,趁这张图还在,还有一件事。All gather 和 all reduce 很昂贵。硅谷的一个趋势是 mega-kernels(巨型内核)。就是一直用内核。我不知道。
Uh, okay, one more thing while this chart is still up. All gather all reduce is expensive. One of the things that is a movement in Silicon Valley is mega-kernels. Just keep using kernels. I don't know.
就这么简单吗?
Is it that simple?
嗯,我的意思是,融合内核救不了你。比如在张量并行中,一半矩阵在一个 GPU 上,另一半在另一个 GPU 上。如果下一步需要整个矩阵来做全线性操作,比如做注意力时,我需要 softmax 或者指数运算,我需要整行数据。所以,我需要知道 GPU 2 的部分结果和 GPU 1 的部分结果,才能在下一阶段做 softmax。所以,即使有融合内核,我也必须让它们相互通信,因为每一步都有非线性。另外,对于 mega-kernels,说实话,我非常看空,我直说了。
Well, I mean like a fused kernel can't save you. Like here with with with tensor parallelism, you're then half the matrix is on one GPU and the other half is on another. And if I need the entire matrix in order to do like an all-linear operation on the next step, which is for instance like if I'm doing the tension, I need the soft max or I need to like like exponentiation. I need to have the entire row. So, I need to know what the partial result was from GPU 2 and what the partial result was from GPU 1 in order to be able to do the soft max in the next stage. So, I like I have to make them communicate with each other even if I had a fused kernel because of the non-linearities within each one. Also with with like mega-kernels, like honestly I'm I'm I'm very bearish on on I'll be honest.
请讲,请讲。
Please please please.
不,只是 mega-kernels 这个方向。这是一个好的研究方向,直觉上理论上都很不错。比如,启动一个内核有很多启动开销,我们不断使用数据,就把所有东西融合在一起。但是,内核本身的复杂性使得编写一个高度优化的 mega-kernel 非常困难,非常非常难。而且,不点名公司,即使是我工作过的公司,或者我交谈过的人所在的公司,那些做融合 mega-kernels 的公司,他们往往最终不会在生产中运行这些内核,因为 TRTL 和模块化内核的启动速度更快,因为你可以优化每个单独的组件,并让它们相互并行。
No, it's just like mega-kernels. It was a good research direction and it seems like a very like like intuitively theoretically it's nice. Like oh my like you have a lot of launch overhead from launching one kernel. We keep using data. Just fuse everything together. But yeah, but like like the the kernel complexity itself is is is very difficult to write a very optimized mega-kernel. It's it's it's very very difficult to do so. And even the like not to name any companies, but like even the companies that have have worked at or people that I've spoken to who work at companies that do fused mega-kernels, they very very often don't end up running those in production because the TRTL and modular kernels that launch are faster because you can optimize each individual components and you can just have them parallelize with each other.
关于 Rubin,不知道你们昨天有没有看到 Rubin 的推特帖子,他们也……
With the Rubens, I don't know if you guys saw the Rubens Twitter post yesterday, but they're also um...
他 Rubin 像……
He Rubens like...
不,不,看 GPU。
Look no no look at the GPU.
他们有一个专门为 Rubin 开的推特账号?
They have a Twitter account for Rubens only?
不不不不不。
No no no no no.
好吧。
Okay.
我当时想,你在说什么?
I was like what what are you talking about?
是的,抱歉。Nvidia 的一位技术负责人发了一条推特,说我们揭开了 Rubin 的面纱,这是数据。第三条推文显示,不过多深入技术细节,我得多读读,但这款 GPU 的设计方式基本上让我的内核无用武之地。你不再需要那么多我的内核了。所以,似乎整个研究领域都不会继续了。但是,是的。
Yeah, sorry. One of one of the tech leads at Nvidia has like launched a Twitter post that like we're pulling the curtain on Ruben and here's the here's the stats. And the and the third tweet showed like not to get too technical into it. I had I need to read it much more, but the GPU is is is designed in such a way that it basically kills my kernels. You don't need to use my kernels that much anymore. So, it seems like that entire research field goes into like won't be continued. But, yeah.
我能对 Rubin 推测一下吗?所以,你……
Can I speculate about Rubin for a minute? So, you...
请说。
Go.
嗯,你知道,我已经经历了……
Um you know, I've been through now...
顺便说一句,书里提到了它们。
And by the way, they are covered in the book.
是的,是的。
Yeah, yeah.
哦,好吧。我的意思是,书里提到它们,是因为我有一大堆博客文章。是的,Rubin 将来会出现的。
Oh, well. I mean, they're covered in the book in the sense of like I have a whole bunch of blog posts. Yeah. Rubin is going to happen in the future.
你甚至还有那个名字……
And you even had the the name of the one...
费曼。
Feynman.
是的,就像,“嘿,这会像……”
Yeah, it's like, "Hey, this is this is going to be like..."
这非常新。
This is very up-to-date.
我在努力让这本书面向未来,好吗?我不想明年之前再出新版。呃,总之,我们刚才在讨论我有多老。嗯,你知道,我现在已经经历了三个硬件发布周期。我经历了 Ampere 发布周期、Hopper 发布周期,还有 Blackwell 发布周期。我说发布周期,不一定指硬件实际发货。比如,Ampere 在我进入这个行业之前就已经上架了。但是,从硬件上架到硬件适合推理,中间有很长一段时间。所以,你看最初的 VLLM 和 SG 系列,尤其是 VLLM,那是针对 Ampere 编写的,后来不得不为 Hopper 更新,再为 Blackwell 更新。每个周期都变得更快、更紧迫,但也复杂得多。当我展望 Rubin 会带来什么新东西时,我觉得 Dynamo 给了我很多技术提示,关于什么样的工作会非常有价值。显然,我们在延续 Blackwell 的一些趋势,对吧?NVFP4 很重要。他们在 NVFP4 张量核心背后的算力非常庞大。嗯,我想我们某个时候会讨论视频,那是那里的巨大价值。你知道,内存带宽快得多,这也是 Blackwell 如此出色的原因,但更重要的是更多的系统思维。
I'm trying to future-proof this thing, okay? I don't want to publish a new one until like next year or something. Uh anyway, so we were discussing the degree to which I am old. Um and, you know, I've now been through three hardware launch cycles. I've been through the Ampere launch cycle, the Hopper launch cycle, and the um Blackwell launch cycle. Now, when I say launch cycle, I don't necessarily mean like the uh the actual shipping of the hardware. Like, Ampere's were racked up well before I got in this industry. But, there was a lot of time between hardware being racked up and hardware being sort of feasible for inference. So, if you look at like the original VLLM and SG line VLLM especially like that was written targeting Ampere and then had to be updated for Hopper, updated for Blackwell. With each of these cycles, it becomes faster and more urgent, but also substantially more complicated. When I look ahead to, you know, what's going to be new with with Rubin, I think that like Dynamo gives me a lot of technical hints around like what kinds of work is going to be very valuable. Obviously, we're continuing some trends from Blackwell, right? NVFP4 is big. The amount of compute that they have behind NVFP4 tensor cores is is massive. Um we're we're we're to talk about video, I think, at some point, and and that's the the big value there. You've got, you know, much much faster memory bandwidth, um but which was the same thing that that made Blackwell so good, um but the the big thing is more systems thinking.
你会更强调 CPU 到 GPU 的互连,更强调 GPU 之间的互连,而当你看到 Dynamo 时,它是一个完全围绕“我如何把 KV 缓存移动到它需要去的地方、在它需要到达的时候到达”而设计的系统。所以,我认为像 KV 缓存卸载、KV 感知路由和分离式架构这些主题,在 Rubin 时代会变得重要得多,这意味着推理工程不再只是一个 CUDA 内核问题,也是一个非常传统的硬件基础设施问题,这是我们长期以来一直在构建的方向,也是让我非常兴奋的事情,因为我们会看到多个领域发生碰撞,而从内核层面到硬件层面再回到内核层面的推理能力将非常有价值。
You have more emphasis on the CPU to GPU interconnect, more emphasis on the interconnect between GPUs, and when you look at Dynamo, it's a system entirely designed around how do I move the KV cache to where it needs to be when it needs to get there. So, I think that themes around KV cache offloading, KV aware routing, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a CUDA kernel problem, but also a very traditional hardware infrastructure problem, which is something we've been building toward for a long time, and something that's really exciting to me because we're going to see multiple domains colliding, and the ability to reason from the kernel level up to the hardware level and back down is going to be very valuable.
我会把 Phil 说的话再往前推一步。我认为它正趋向于完全变成一个基础设施问题,比如 PD 分离训练谱系的问题,但编写内核不会是什么大问题,因为 GPU 正越来越像 ASIC,你只是试图编排 GPU 上发生的事情,而不是在逐个线程的层面控制它。你在 QTalk QCSL 上也能看到这一点,你只是在数据块(tiles)的层面工作,而不再控制 GPU 上每个线程做什么。那已经被替你处理好了。所以,我猜,根据你和其他人的交流,你是否同意 GPU 和未来的 GPU 正越来越趋向于变成只需要被启动、然后进行数据操作的 ASIC?
I will take what Phil said one step further actually into that. I think it's trending towards becoming exclusively an infrastructure problem, where the problems of PD disaggregate training spectrum, but writing kernels is not going to be much of a problem because the GPU is moving more towards being an ASIC, where you're just trying to orchestrate what happens on the GPU, but you're not actually controlling it thread by thread level. And you see this with QTalk QCSL, like you're just working at levels of tiles of data, but you're no longer working at controlling what each thread does on the GPU. That's being taken care of for you. So, I guess do you agree that a GPU and future GPUs are trending more and more towards becoming ASICs that just need to be launched and then need to the data operation based on your conversations with other people?
哦,我的意思是,是的,不,那是市场的一部分。
Oh, I mean, yeah, no, that is a section of the market.
对。
Right.
而且显然,ASIC 在只针对它们的工作负载时能提供更高的性能。
And obviously ASICs can do a lot more performance for only their workload.
对。
Right.
而 GPU 中的 G 让它们继续保持非常通用。
And the G in GPU makes them continue to be very general.
是的。我觉得这里有一个谱系,
Yeah. The I think that there's like a spectrum
它是图形,但
it's graphics, but
对。
Yeah.
我一直这么说。我得纠正自己,免得你因为我搞错 G 而找我麻烦。
I keep saying this. I have to correct myself in case you come at me for getting the G wrong.
对,这就像一个谱系,对吧?从非常通用的计算到像 Talos 这样的东西,硬件是为特定的一组模型权重而构建的。
Yeah, it's like a spectrum, right? Of very general-purpose compute to something like a Talos where you've got the hardware built for a specific set of model weights.
烧录进芯片里。
Burned into the chip.
对。
Yeah.
没有加载。
No loading.
我不,我不会说我们会完全走到那一步。更像是沿着这个谱系,朝着硬件内更专业化的方向迈出一步。
I don't I wouldn't say that we're going all the way there. It's more like along the spectrum, it's a step in the direction of more specialization within the hardware.
是的。我很好奇。我觉得他是在引导某个方向。
Yeah. I'm curious. I feel like he was driving towards something.
我想我的观点是,你对除了把权重烧录进芯片之外的一切都持悲观态度。把权重烧录进芯片是不切实际的,因为你想微调、想优化、想量化、想发布模型的新检查点。如果烧录进芯片,芯片在一两个月内就没用了,对吧?我想我的观点是,你怎么能不喜欢看到 Nvidia 越来越专业化,比如把它的 GPU 从通用编程范式(它只是一台通用计算机,你可以用它来编程线程)中带出来。随着每一代新产品的推出,你加入越来越多专门的指令、专门的张量核心、专门的 MMA 指令,这些东西让你几乎可以像控制 ASIC 一样控制它,几乎像控制一组 ASIC。你怎么能看着这个趋势,却仍然看好那些为 AI 开发 ASIC 的公司呢?
I guess my point is being bearish on like you say everything else apart from burning the weights into the chip. Burning weights into the chip is impractical because you want to fine-tune, you want to optimize, you want to quantize, you want to release new checkpoints of the model. If it's burned into the chip, the chip's useless in like a month or two, right? I guess my point is how can you not like seeing Nvidia more and more specialize it like take its GPUs from a general programming paradigm where you're just it's a general computer that you can use to program threads. And with every new generation, you're putting more and more specialized instructions, specialized tensor cores, specialized MMA instructions, things that will allow you to control it almost as an ASIC, almost as a collection of ASICs. How can you look at this trend and then still be bullish on companies that are coming up with ASICs for AI?
从某种意义上说
In the sense that
是的,因为它们有点在朝那个方向进化。
Yeah, because they're sort of they're evolving towards that direction.
朝那个方向进化。就像在 Reuben 中,我猜与 Ampere 或 T4 相比,Reuben 基本上就是一个 ASIC。它基本上就是一个几乎像……一样使用的东西,它相当高。显然你可以编程,说它基本是很有争议的。它是一个 GPU。它是通用的。它确实有线程。我可以写 CUDA 来控制它并改变它的操作。但它有固态核心、张量核心、TMA 和张量内存。它有这些几乎专门用于加载模型权重的东西。它有张量核心指令,这些指令几乎完全围绕当今市场上存在的模型的头部维度而设计。说你要推出一个 ASIC,然后你要在上面蚀刻一些东西。好吧,下一代架构基本上就会让它没用。是的,我不知道。我认为
Evolving towards it. And like as in Reuben, I guess compared to Ampere or T4, Reuben is basically an ASIC. It is basically just a thing that is used almost like an It's like pretty high. Like you can program obviously, it's very controversial to call it basic. It is a GPU. It is general. It does have threads. I can write CUDA to control it and change its operations. But it has the solid cores and tensor cores and TMAs and tensor memory. And it has these things that are almost exclusively useful for loading model weights. It has tensor core instructions that are almost exclusively shaped around the head dimensions of models that exist in the market today. To say that you're going to come up with an ASIC and you're going to etch something into it. Well, with the next architecture it's basically going to be useless. Yeah, I don't know. I think that
要记住的是这些硬件周期有多长。所以,如果一款芯片今天问世,那意味着它的设计过程是在多年前启动的。而他们在 Nvidia 做得非常好,预测了市场的发展方向,你知道,
The thing to remember is just how long these hardware cycles are. So, if a chip is coming out today, that means the design process for it was kicked off years ago. And they've at Nvidia they've done a very good job of predicting where the market is going to go and you know,
他们肯定拥有最多的信息。
They have the most information for sure.
当然。但如果你看看,你知道,公开的开源模型架构,它们或多或少看起来像今天模型的早期版本,Reuben 老实说是第一颗完全在那个世界里构建的芯片。所以,你可以看到很多关于这颗芯片将被要求执行的工作负载形态的理解,体现在它的设计方式中。
Of course. But if you look at, you know, there being public open source model architectures that look more or less like early versions of the one today, Reuben's honestly the first chip that was fully built in that world. And so, you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.
对。
Yeah.
好吧,我不是直接回答这些问题的最佳人选。我认为这些问题非常合理,老实说,这是第一个基于 Reuben 的、我听到被阐述得这么好的问题。我确实认为,我会为垂直整合的模型实验室 ASIC 辩护。所以,就像 OpenAI 和 Broadcom 合作的什么 jalapeno 芯片,这完全合理。我们第一次在播客里和 Martin Casado 讨论这个,他说:“看,如果你有一个万亿美元或 5000 亿美元的训练运行,那么从中拿出 500 亿来做 ASIC。没问题。你不会从 ASIC 获得超过 10% 的效率,这说得通。对吧?”所以,特定模型的芯片,是的。但 ASIC 公司,有趣的是,我觉得你过于关注像你说的 Talos 那些东西。他们做了更多表面工程,或者像内存和硬件的实际分配,以及芯片之间的通信,这些可能仍然不会被 Rubin 触及,但我不知道细节。
Okay, so I'm not going to be the best person to directly answer those questions. I think these are very fair questions that are honestly the first one that's based on Reuben that I've heard articulated so well. I do think that I will make a case for vertically integrated model lab ASICs. So, like the OpenAI Broadcom whatever jalapeno chip which totally makes sense. Like so, we first had this on the pod with Martin Casado where he was like, "Look, if you have a trillion-dollar or 500-billion-dollar training run, then take 50 billion of that and make an ASIC. Like it's fine. Like you won't get more than 10% efficiency from the ASIC and that makes sense. Right? So a model specific chip, yes. But ASIC companies, the interesting thing is I feel like you are hyper-focusing on like you say the Talos stuff. They are doing a lot more sort of surface area engineering or like the actual allocations of memory and hardware and like the communication between chips that probably still won't be touched by Rubin, but I don't know the details.
我明白了。我明白了。
I see. I see.
他们通常谈论的事情,我预期会比 Rubin 所做的任何可编程实现都要大几个数量级,但谁知道呢。
They typically talk about things that I would expect to have bigger orders of magnitude than would be programmably accomplished by whatever Rubin does, but who knows.
不,我明白。我明白。是的,看起来是这样。
No, I see. I see. Yeah. It seems so.
是的,想想真正阻碍推理速度提升 10 倍到 1000 倍的因素是什么?不是那些可以在现有 GPU 设计内部重新安排的东西。
Yeah, like think about what are the real blockers to 10x to 1,000x faster inference? It is not the stuff that can be rearranged just within the existing GPU design.
互连通信。
Intercommunication.
是的。比如这些家伙的目标是每秒 30 万 token。他们还没达到。
Yeah. Like these guys are aiming for 300,000 tokens per second. They're not around.
比 ASIC 的性能提升更重要。
More than performance from ASICs.
也许吧。我觉得有意思的是,你最近做了这么多内核工程,却对这么多内核工程如此看空。
Maybe. I think it is interesting to me that you're so bearish on so much of this kernel engineering given how much of it you've been doing recently.
对,对。所以我做得越多,就越觉得这不是什么大不了的事。我还要补充一点,模型已经有好几代了,对吧?我想在你们那边,你们看到很多,好吧,今天是 Gemini,Kimi,DeepSeek,MiniMax,再加上其他的。有些在做完全不同的事情,对吧?Gemma,没有编码器,最新的 thinking machines 完全是从零开始。但当你看看另一边,我们在 GPT-5 这一代上已经多久了,对吧?他们服务那个模型已经相当久了。当然,也许有更多的预训练,有不同的检查点,但你实际上可以从中榨出不少东西,而且你做了数十亿美元的训练。如果你能让它效率提高 X%,他们会服务一段时间。Claude 5 系列也一样,对吧?
Right. Right. So the more I do it, the more it just seems to me that it's not mega. I would also add that there are generations of models being out, right? I think on your guys' end you see a lot of okay, one day it's Gemini, Kimi, DeepSeek, MiniMax, throw in the others. Some are doing completely different stuff, right? Gemma, no encoder, the latest thinking machines is all from scratch. But when you look at the other side, like how long have we been on the GPT-5 generation, right? They've been serving that thing for quite a while. Sure, there's maybe more pre-training, there's different checkpoints, but you actually can squeeze quite a bit out and you do a multi-billion dollar train run. If you can make it X% more efficient, they serve it for a while. Same with say the Claude 5 family, right?
比如他们现在发布新模型,像 GPT-6 之类的,他们每年发布一个新模型,我们不知道,我们假设他们改变了一些架构细节,而不只是做后训练。就像你每年要花 500 亿美元为模型设计新的 ASIC,然后把前一年的 ASIC 扔掉。
Like they release a new model like GPT-6 now or whatever and they release a new model every year and we don't know, we assume that they're changing some bits of the architecture and not just doing post-training. Like you're going to be spending 50 billion dollars a year every single year coming up with new ASICs for the model and throwing out the ASICs of the previous year away.
是的。是的。容易。
Yeah. Yeah. Easy.
所以我想,好吧,我稍微不同意,基于我的——再说一次,这都是二手的——关于模型的寿命。仍然有人在用 40。
So I think okay, I would slightly disagree based on my again it's all second hand on the longevity of a model. There's still people out there using 40.
是的。
Yeah.
是的,Llama,不是 Llama 2,而是 Llama 3。我仍然看到 Llama 3 的工作负载。
Yeah, Llama not Llama 2 but Llama 3. I still see Llama 3 workloads.
是的。
Yeah.
因为如果它完成了,如果它被信任,就不要改变它。
Because if it's done, if it's trusted, don't change it.
如果它有效。
If it works.
这是开源的一个承诺,对吧?就像整个“拯救 40”运动,你不需要“拯救 Llama 3”运动。你只需要在某处有一个 8100。
Which is one of the promises of open source, right? Like the whole 40 save 40 movement, you don't got to have a save Llama 3 movement. You just got to have an 8100 somewhere.
我觉得这有道理。还有一个问题是,如果一个模型能做足够多的事情,能使用足够的工具调用,并且足够庞大,它能不能直接进行网络搜索、工具搜索、写代码?你真的需要不断压榨更多吗?我们会,因为你们会把它做得便宜、快速、更小,我可以换进去。但在某种程度上,你今天给我 52,或者说随便什么 120B 模型,我可以运行相当长一段时间,对吧?
I think there's a point there. There's also the question of if a model can do enough and use enough tool calls and be a gigantic enough, can it just web search, tool search, write code? Do you really need to keep squeezing more? We will because you guys will make it cheap and fast and smaller and I can swap it in. But at some level, you give me 52 today or say whatever 120B model, I can run with it for quite a while, right?
这是假设你不需要智能。
This is assuming you don't need intelligence.
我认为有很多智能和可预测性。比如我是一家企业,这是经过试验和测试的。它得到了我 5000 个利益相关者的批准。没有搞砸。它每天运行一个批处理作业,我喜欢结果。结果是可预测的。
I think there's a lot of intelligence and predictability. Like I'm an enterprise, this is tried and tested. It is signed off by my 5000 stakeholders. Not butchering it. It runs a batch job every day and I like the results. The results are predictable.
对。
Right.
继续使用它们没有意义,因为东西变得更稀疏、更便宜、更好,但这并不意味着旧模型 GLM 50 不可用,对吧?如果我们遇到停滞,不管什么原因,仍然有很多可以榨取的。
It doesn't make sense to keep using them like stuff gets sparser, cheaper, better, but that doesn't mean that old models GLM 50 isn't usable, right? If we hit a stall say for whatever reason, there's still a lot that can be squeezed out.
我们快没时间了。我确实想确保,是的,实际上我们碰巧有这个图表。把它和任何 Cerebras 的图表比较,对吧?我认为 Edge 和 Matics 还没有发布公开图表,但完整的面积非常不同。大小非常不同,对吧?这不是晶圆级,对吧?这可能是,我不知道,一个晶圆上有几百个这样的。我不知道比较有多大,但这是一个非常明显的面积分配差异。
We're going to run out of time. I did want to also make sure, yes, actually we happen to have this diagram. Pull compare this versus any Cerebras diagram, right? I don't think Edge and Matics have put out public charts yet, but the complete real estate is very different. The size is very different, right? This is not wafer scale, right? This is probably like, I don't know, a few hundred of these on a wafer. I don't know how big the comparison is, but it is a very real estate allocation difference.
好的,是的。差异。我会说几十个。几十个,是的。
Okay, yeah. Difference. Few dozen, I would say. Few dozen, yeah.
在我们离开硬件话题之前,我有两个快速问题。第一,最新的 Kimi,非常大,3 万亿参数,在大多数单节点硬件上放不下。
Before we move from hardware, I have two quick questions. One, the latest Kimi, which is really big, 3 trillion, doesn't fit on most hardware on single node.
是的。你需要 GB300。你需要 GB300。
Yes. You need GB300. You need GB300.
或者 AMD。
Or AMD.
这是简单的数学。NVFP4,2.8 万亿参数,1.4 TB。GB300 每个有 288 GB,所以八个这样的,你就有足够的空间放模型。而且说实话,GPU VRAM 数学的另一件事是你必须为 KV 缓存留出空间。这在一定程度上取决于上下文长度。所以当一个模型既有非常大的参数数量又有非常长的上下文长度时,你就是在争夺空间,这就是为什么 KV 缓存卸载会成为更突出的主题,我认为,对于这些巨大的模型,因为你真的空间紧张。
It's simple math. NVFP4, 2.8 trillion parameters, 1.4 terabytes. The GB300s have 288 gigabytes each, so across eight of those, you have enough room for the model. And honestly, the other thing with GPU VRAM math is you have to leave space for the KV cache. And that's going to depend to some degree on the context length. So when a model has both a very large number of parameters and a very long context length, you're kind of fighting over space, which is why the KV cache offloading would become a more salient topic, I think, with these huge models, because you're really crunched for space.
有了 Rubens,你现在有什么,NVL72 机架,多少,20 TB 的……
With the Rubens, you now have what, NVL72 rack of what, 20 terabytes of...
现在你在 Blackwell 上也有 NVL72,但你不能假设你会在那上面做推理。世界上的 8X 机架比 NVL72 多得多。
Now you still have NVL72 on Blackwell as well, but you can't necessarily assume you're going to do inference on that. There's a whole lot more 8X racks in the world than there are NVL72s.
是的,我想我最后一个关于硬件的好问题是,你注意到硬件世代对新的预训练基础模型有什么影响吗?所以,你提到的效率提升之一是你可以更换硬件。那是 2 倍收益之一。当我们看到 Rubins 上训练方面的新东西出来时,锁有什么变化吗?这是否会影响当这些硬件更可用时我们会看到的模型类型?你能……
Yeah, I guess my last good question on hardware was do you notice anything with hardware generations for new pre-trained base models? So, one of the things you said for efficiency is you can swap hardware. That's one of the 2x gains. When we see new stuff coming out training wise on Rubins, any changes on locks? Does this affect what type of models we will be seeing when these are more available? And can you...
它们变得更大。就像人们理解,鉴于最新的推理硬件,你能运行的模型参数数量有一个上限。这形成了一个天花板。所以例如当 DeepSeek 1 出来时,它有 6710 亿参数,当时真的很大,我认为它极大地推动了我们迅速采用 Blackwell 并擅长在 Blackwell 上服务。所以,是的,在我看来主要是关于模型大小,然后是关于将架构和原生量化与目标硬件匹配,就像我们谈到的,比如所有 Nemotron 模型都是 NVFP4。
They get bigger. Like people understand the ceiling that you have in terms of how many parameters of a model you can run given the sort of latest inference hardware. And that kind of forms a ceiling. So for example when DeepSeek 1 came out, it was 671 billion parameters, which at the time was really huge, and I think did a lot to push us to really quickly adopt Blackwell and get good at serving on Blackwell. So, yeah, it's mostly in my mind about model size and then about matching the architecture and the native quantization to the target hardware, like we talked about with, you know, all Nemotron models are NVFP4 for example.
所以,我们谈了很多关于 LLM 的话题。
So, we talked a lot about LLMs.
嗯。
Mhm.
你在书里还有更多内容。
You have a lot more in the book.
那音频视频呢?推理工程的另一面是什么?阿里,你在视频扩散方面很在行。
What about audio video? What's the other side of inference engineering? Ali, you're pretty big in video diffusion.
视频扩散模型与 LLM 的推理方式有很大不同,LLM 是自回归的,而视频扩散不是。比如,你不做批处理,每个请求只进一个 GPU,服务一个 GPU,不需要分片。模型小得多,比如 1.2 是 200 亿参数的模型,你不用担心,它比最好的 LLM 小几个数量级。而且在这个领域,开源模型的情况是:在 LLM 中,我们看到 Gemini 3 几乎可以媲美 Miethos 或 GPT-5.5,最好的开源 LLM 和最好的闭源 LLM 之间的差距非常小,以前是 6 个月,现在我觉得不再是 6 个月了,基本上已经持平了。但视频模型绝对不是这样,差距巨大。如果你看看今天用开源模型如 1.2.2 生成的最佳视频,与 Kling 或 Veo 相比,差别是天壤之别。这就造成了这种差距,媒体公司大多数时候会选择闭源模型。如果我告诉你:“嘿,我可以用这个模型为你生成一部完整的 3 小时电影,而且我会优化它,让你只需付我 10 美元。”但如果他们用闭源模型做,就得付 1000 美元,也就是 100 倍。我便宜 100 倍,但还是要 1000 美元。他们仍然会选择用 Veo 和 Kling 做所有的剪辑。所以这就像一个鸡生蛋蛋生鸡的循环:需求减少导致创新减少,进而导致发布的开源检查点减少。一些发布开源模型如 1.2 的实验室已经将最新模型闭源了,比如 1.2.7 不是开源的,我们还在 1.2.2。
Video diffusion models are shaped a lot differently from the stuff you can think about reasoning about with LLMs being autoregressive. With video diffusion, it's not the case. For instance, you don't do batching. Every request just comes in on one GPU and it serves one GPU. You don't have to shard. The models are a lot smaller, like 1.2 is a 20 billion parameter model. You don't need to worry about it; it's orders of magnitude smaller than the best LLMs. And it's one of those spaces where the open-source models are like with LLMs, we see Gemini 3 is almost comparable to, you know, Miethos or GPT-5.5. The difference between the best open-source LLM and the best closed-source LLM is very small. It used to be 6 months. I don't think it's 6 months anymore. I think it's basically almost on par. Video models are definitely not. There's a huge gap. If you look at the best video you can generate today with an open-source model like 1.2.2 versus something like Kling or Veo, the difference is night and day. So it creates this disparity where media companies will choose to go most of the time to closed-source models, for instance. If I were to tell you, "Hey, I can generate an entire 3-hour movie for you with this model, and I'll optimize it so that you only have to pay me $10." But if they were to do it on a closed source, they'd have to pay $1,000, which is 100x. Like I'm 100x cheaper. But it's still $1,000. They're still going to choose to do all of their cuts with Veo and Kling. So it's like a chicken and egg cycle where less demand causes less innovation in the field, which causes less open-source checkpoints to be released. And some of the labs that were releasing open-source models like 1.2 will have closed source their latest models. Like 1.2.7 is not open-source. We're still in 1.2.2.
视频模型尤其面临的挑战是 token 的数量。视频模型,你想生成高质量的视频。假设你每秒 16 帧,这是绝对的最低限度。再假设你做 480p 的视频。所以你可以考虑你的尺寸,我有一个很好的图展示了 token 的绝对数量。假设我们只看一个视频,比如《斯巴达》或《斯巴达 300》之类的。假设我们看四帧,对吧?那四帧视频,如果你做全注意力,如果你稍微往上一点,你看到的是 480p 乘以 720 乘以 81 帧,仅仅 5 秒,因为 16 FPS 乘以 5,对吧?然后你把它压缩到潜空间,但你仍然在做 30 乘以 50 乘以 21 个 token。
The challenge with video models, especially, is the number of tokens. Video models, you want to generate a high-quality video. So let's say you're doing 16 frames per second, that's the absolute minimum you'll do. And let's say you'll do 480p video. So you can think about your dimensions, and I have a good diagram that shows the sheer number of tokens, right? Let's say you're looking at just one video like Sparta or Sparta 300 or whatever. So let's say we're looking at four frames, right? Those four frames of that video, if you're doing full attention, if you go a little bit up, you're looking at 480p by 720 by 81 frames in just 5 seconds because 16 FPS by 5, right? And then you compress it down to latent space, but you're still doing 30 by 50 by 21 tokens.
嗯。
Yeah.
这意味着对于注意力来说,仅仅 5 秒,你就要在 35,000 个 token 上运行注意力,对吧?所以注意力成了巨大的瓶颈。而且因为它是平方级的,如果你把它扩展到 10 秒,那就是平方。20 秒,30 秒。所以要生成 1 分钟的好过场动画,在同样的算力时间内几乎是不可能的。它变得不可行,你做不到。所以你最终会走向两个方向。要么你决定对整个视频一次性做注意力,在这种情况下你被迫做稀疏注意力。所以如果你回滚到原始视频图像,你可以看到左边,比如,我会做全注意力,其中稀疏注意力中的每一个 token 都关注其他每一个 token。你可以看到大量的红色块。在右边,我只让每个 token 只关注对它重要的 top K 或 top 12.5%,这可以是空间的。比如代表皇冠的 token 关注头部、脸部,然后前一帧的头部,时间局部性、空间局部性之类的。这导致视频质量很差。这篇文章的重点是展示你可以训练并做所有这些事情,但你仍然会损失一点视频质量。所以你最终会面临两种情况之一。要么你咬紧牙关,拥有巨大的算力,对大约一百万个 token 做全注意力,因为你试图生成 2 分钟的视频。要么你转向自回归视频。在我看来,自回归视频是未来押注的方向,但今天还没有好的开源自回归视频模型。如果你想要一个小时的电影,如果你想看到视频模型生成好莱坞级别的电影,它们必须是自回归的,才能超越那 5 秒的帧。或者必须在算力上发生某种疯狂的飞跃,允许我们同时对数百万个 token 做全注意力,并且以高效的方式。
Which means that for attention, for just 5 seconds, you're running attention on 35,000 tokens, right? So attention becomes such a huge bottleneck. And because it's over in square, if you extend that to 10 seconds, well, it's just square. 20 seconds, 30 seconds. So to generate a good cutscene of like 1 minute, it's almost impossible to do within the same compute time. And it just becomes unfeasible. You can't do it. And so you end up moving towards two directions. Either you decide to do attention on the entire video at once, in which case you're forced to do sparse attention. So if you scroll back down to the original video image, you can see where on the left, for instance, I would be doing full attention where every single token in that sparse attention of 200 scene attends to every single other token. As you can see the sheer number of red patches. On the right, I'm only attending to each token only attends to the top K or top 12.5% that's important to it, which can be spatial. So like the token that represents the crown attends to the head, the face, and then the head on the other frame in the previous frame, temporal locality, spatial locality, that kind of thing. This results in terrible video quality. And the whole point of the post or the article here is to show how you can train and do all these things, but you will still suffer video quality a little bit. So you end up with one of two things. Either you bite the bullet, you have huge compute, and you do full attention over like a million tokens because you're trying to generate like 2 minutes of video. Or you move towards autoregressive video. Autoregressive video seems to me like that is the bet that the future is going to be making, but there are no good open-source autoregressive video models out there today. And that seems to be the if you want to get like an hour movie, if you want to see video models generating like Hollywood level movies, they have to be autoregressive in order to exceed that 5-second frame. Or there has to be some insane leap that happens in compute that allows us to do full attention over like millions of tokens at the same time, and in an efficient manner.
即使是一百万个 token,那也是平方级的,所以你会很快达到极限。我想你能解释一下自回归的优缺点权衡吗?我想到的一个是跨帧的一致性。你在生成自回归扩散 10 分钟后,你会忘记。但这样做的优缺点是什么?
Even millions of tokens, that's like you're quadratic, so you're going to get there really quick. I think can you explain the pros and cons trade-offs of autoregressive? So one that comes to mind is the consistency across frames. You will 10 minutes into generating autoregressive diffusion, you're going to forget. But what are pros and cons of this?
嗯,就像自回归 LLM 一样,你可以把我们讨论过的很多 LLM 优化应用在那里,比如 Specter 之类的。如果你有一个高质量、规模化的模型,我没有理由不能流式输出,比如我可以先给你看第一帧,然后就像 2023 年的 GPT 一样,它几乎一次性生成文本,但你可以边读边生成。对于视频模型,你可以边看边生成。它生成帧。所以逐 token 生成将允许我们大幅扩展并应用注意力机制。缺点是每一个自回归视频模型的质量都很差。如果你把任何 Open-Sora 1.2 的质量与任何其他自回归模型相比,你会看到 Open-Sora 1.2 生成的视频就像猫狗打架,而自回归模型会给你劣质的《猫和老鼠》质量。我不知道。那么生成长输出的解决方案就变成了:“好吧,我们不会使用自回归模型。”
Well, like autoregressive LLMs, you can take a lot of your optimizations that we discussed with LLMs, like Specter and stuff like that, and you can apply it there. And if you have a very high-quality scaled-up model, there is no reason why I can't stream the outputs as in I can show you the first frame, and then I'm like kind of like GPT back in 2023 when you're like now it just almost one shots the text, but then you could read and it's generating as you read. With video models, you can watch and it's generating as you watch. It generates the frames. And so token-by-token generation will allow us to scale a lot up and apply the attention mechanisms there. The downside is every single autoregressive video model is just terrible quality. If you put the quality of any Open-Sora 1.2 versus any other autoregressive model, you can see like a video generated by Open-Sora 1.2 is like a cat and dog fighting. Autoregressive model will give you like degraded Tom and Jerry quality. I don't know. The solution to generating long output then becomes, "Okay, we're not going to use autoregressive model."
如果你看看 Grok Imagineer 视频做的一些事情,他们做得非常好,他们会尝试把这些 7 秒的片段拼接起来。你生成 7 秒,然后你会说:“好,你能扩展这个视频吗?”他们会把两个视频拼接在一起。开源似乎没有他们那些技巧。而且从定义上讲,它是闭源的,我们不知道他们在做什么。但你能做到的最接近的方法是取视频的最后一帧,输入到一个文本和图像到视频的模型中,它会接受文本提示和最后一帧的图像,然后你让它生成接下来的 5 秒。这大概就是你能扩展这种模型来生成一部电影的方式,你只是不断地逐帧流式生成。但会出现漂移。你从图像开始,然后生成一个视频,那下一个 5 秒的视频质量就更低。第三段更低,第四段更低。有时你会看到新视频只是比第一个稍微暗一点,下一个比第二个更暗,直到 25 秒时变成黑屏。我们试着做一个演示来展示这个,但展示起来非常尴尬,所以我们决定不做了。但我认为模型会达到那个水平。它们只需要大幅扩展规模,并朝着自回归的方向发展。但训练技术似乎还不明确。
If you look at some of the things that Grok Imagineer video does, and they do it really well, they'll try to stitch these 7-second chunks together. So you generate 7 seconds, and then you're like, 'Okay, can you extend this video?' And they'll chunk two videos together. Open source doesn't seem to have the tricks that they have there. And by definition, it's closed source. We don't know what they're doing. But the closest you can get is taking the last frame of a video and feeding it into a text and image-to-video model, where it will take the text prompt and the image of the last frame, and you'll ask it to generate the next 5 seconds. That's kind of how you can extend this level of model to generate a movie where you're just constantly streaming frame by frame. But you get drift. So you start with the image, then you generate a video, and that next 5-second video is lower quality. The third chunk is even lower, and the fourth chunk is even lower. Sometimes you'll see things where the new video is just ever so slightly darker than the first one, and the next one is darker than the second, until at 25 seconds you have a black screen. We tried to have a demo that would show this, but it was extremely embarrassing to show, so we just decided not to. But I think models will get there. They just need to scale up significantly and move towards being auto-regressive. But the training techniques don't seem to be clear there.
对于那些对岩石成像感兴趣的人,我们和那个团队的 Ethan 做了一期节目。
For those who are interested in rock imaging, we did a part with Ethan from that team.
对。
Right.
他透露了一些线索,但不足以让我们完全重建一切。
Who dropped a few hints, but not enough that we can fully reconstruct everything.
具体在这部分,他解释了一点……是的,我说过我们谈到了记忆和更长的上下文等等。但据我所知,它不是自回归的,尽管业内没有人是自回归的。
Specifically on this part, he explains a bit of... Yeah, I said we talked about memory and longer context and all these things. But as far as I know, it's not auto-regressive, even though no one in the industry is auto-regressive.
是的,看起来是这样。
Yeah, it seems to be yeah.
理解自回归模型和扩散模型之间的关键区别在于,扩散注意力是双向的,而自回归只向前推进序列。所以这就是为什么你会看到这种偏离轨道的行为。如果你天真地把视频生成模型构建为简单地生成一帧帧的线性序列,你就不能回到序列中去修复某些东西来使整体保持一致。当然,我们需要为视频模型保留所有这些潜在空间的原因,就像你说的,我们把所有 token 都保存在内存中,我们迭代整个序列,你可以调整过去以使未来合理。所以如果我们考虑能让我们达到这些更长、更丰富序列的架构,它可能会是自回归和扩散的混合体,共同发挥各自的长处。
The key thing to understand between an auto-regressive model and a diffusion model is that diffusion attention goes in both directions, while auto-regression only goes forward in the sequence. So that's why you see this sort of going off the rails behavior. If you naively construct a video generation model as simply generating a linear sequence of frames, you can't then go back in that sequence and fix something to make the whole thing consistent. Of course, the reason we need all this latent space for the video model is, like you said, we keep all the tokens in memory, we iterate over that full sequence, and you can adjust the past in order to make the future make sense. So if we think about the architecture that's going to get us to these longer, richer sequences, it's probably going to be a mix of the auto-regressive and the diffusion working together to do what each piece is good at.
嗯,你凭直觉理解了我的意思。比如英语或任何书写语言,就是从左到右。你可以流式输出 token,可以流式输出思维链。即使作为人类,你写,然后思考接下来要生成什么,然后写下来,然后思考你的想法,然后向前生成。当然,你可以争辩说写作时需要回去编辑一些东西。但你需要的次数比你想象的要少。而视频则没有顺序性——左上角的像素和右下角的像素都需要相互关注,以几乎同等程度地理解视频质量。而文本则不需要那么多。
Well, you intuitively got what I'm saying. For instance, English or just writing languages, it's just left to right. You can stream your tokens. You can stream your chain of thought. Even as a human, you write, then you think about what the next thing you're going to generate, then you write that, then you think about your ideas, then you generate forward. And sure, you can argue that as you write, you need to go back and edit some things. But you need to do that less often than you think. Whereas with video, there is no sequential—the pixel in the top left corner and the pixel in the bottom right corner both need to attend to each other to understand how the video quality is going to be almost equally. Whereas with text, you don't need that as much.
音频有类似的情况吗?我不是 100% 确定,但大约一年前有音频 LM,音频有扩散和自回归。就你提到的点,主要是在推理方面,尽管片段较短——大多数音乐是 3 到 5 分钟——我们基本上已经转向自回归了。
Is there a parallel to audio? I'm not 100% confident on this, but there was a point about a year ago where there was audio LM, there's diffusion for audio and auto-regressive. And for the points you mentioned, mostly on the inference side, even though they're shorter clips—most music is 3 to 5 minutes—we've basically swapped over to auto-regressive.
是的,我不能说音乐,但语音是自回归的。这甚至回到一年半前明显的架构选择——你只需在词汇表中添加一堆波形,这样 LM 就能输出代表这些波形的 token,然后你构建语音,这就是你流式输出的方式。
Yeah, I can't speak to music, but speech is auto-regressive. This was even back with the obvious architectural choice a year and a half ago—you just add a bunch of waveforms to the vocabulary so that the LM can output tokens that represent those waveforms, and then you construct speech, and that's how you stream it.
就是这样。
That's it.
嗯,这就是我 2025 年的 AI 演讲。
Well, that's my AI talk from 2025.
不错,不错,不错。
Nice, nice, nice.
不错。
Nice.
但音频不是同样的挑战,对吧?因为音频是用一个生成一切的 LM 解决的。对于音频,它仍然是一个可以用 LM 生成的转录文本。所以你的音频模型只需要转录它——文本到语音。
But it's not the same challenge with audio, is it? Because audio is solved with an LM that generates everything. With audio, it's still a transcript that you can generate with an LM. So your audio model just needs to transcribe it—text to speech.
对于音乐,有一个阶段在扩散和自回归之间权衡,两者相当。各有更多利弊。我只是想试探一下,因为你有见解。
For music, there was a phase of a trade-off between diffusion for music and auto-regressive, and they were both pretty on par. There are probably more pros and cons to either. I just wanted to poke, since you had takes.
是的,我特别不了解音乐。关于你说的编辑——你写作——显然我认为我的编辑会告诉我,我实际上需要更频繁地这样做,回去修复东西。我可以想象音乐或诗歌,例如,你有押韵方案,你可能想回去做一个改变,以便更容易设置你以后想要的押韵。能够双向关注会有一些优势。但据我所知,我非常把这个影响问题分为自回归模型(有一套约束和技术)和扩散模型(有一套约束和技术)。我认为文本、嵌入、语音输入和输出属于自回归侧,而图像和视频属于扩散侧。两者之间有一些重叠。这不是一个完美的划分,但这是我使用的广泛分类。
Yeah, I don't know about music specifically. With what you said about editing—you writing—obviously I think my editor would tell me I actually need to do that more often and go back and fix things. I can imagine music or poetry, for example, where you have a rhyming scheme and you might want to go back and make a change to make it easier to set up a rhyme that you want to make later on. There'll be some advantage to being able to attend in both directions. But yeah, to my knowledge, I very much bifurcate this influence problem into the auto-regressive models, which have a set of constraints and techniques, and the diffusion models, which have a set of constraints and techniques. I think of text, embedding, voice in and voice out as being on the auto-regressive side, and then image and video being on the diffusion side. There's some overlap between the two. It's not a perfect split, but that's the broad categorization I use.
我应该指出,我认为已经确认了,对吧,Nano Banana 和 GPT 图像是自回归图像。这有点像我们讨论的混合方法,但在图像领域,它还没有进入视频领域,至少在开源世界是这样。
I should point out, I think it's confirmed, right, Nano Banana and GPT image are auto-regressive image. It's kind of this blended approach that we're talking about, but in the image space, it hasn't made its way over to the video space, at least in the open source world.
是的,但我认为那不会太远,如果可能的话。至少 Queen 图像的人正在尝试。
Yeah, but I assume that's not too far away, if that is possible. At least the Queen image guys are trying it.
是的,是的,和……
Yeah, yeah, with...
是的。
Yeah.
我对 Queen Queen 图像三真的很兴奋。我希望他们能开源它。
I'm really excited for Queen Queen image three. I hope they open source it.
然后我会展示,我也会提到文本扩散方面有一些进展。不是很多。
And then I'll show, I'll also mention on the diffusion for text side there's been some movements. Not a lot.
是的,我们有 Mercury。
Yeah, we've got Mercury.
你托管 Mercury?
You host Mercury?
是的。
Yeah.
哦,不错。不错。
Oh, nice. Nice.
还有 Gemma,对吧?扩散 Gemma?
There's Gemma as well, right? Diffusion Gemma?
扩散 Gemma 是开源的。
Diffusion Gemma's open source.
是的。
Yeah.
而且我们在科学方面刚刚发布了一些使用扩散的虚拟细胞模型。
And we on the science part we just have been releasing some virtual cell models that use diffusion as well.
是的,他们已经构建了。这肯定还处于廉价快速 token 的世界。我们正在尝试……
Yeah, they have built. It's definitely still in the sort of cheap fast tokens world. We're trying to...
我觉得这是错误的营销。我之前告诉过他们。我说,你看,你无法击败其他大语言模型会做的优化,但你可以有不同的 API。你应该能够以不同于聊天响应的方式使用它。因为它是扩散,你可以做类似,文本的扩散无分类器引导是什么样的?比如给我一首诗,给我一个逐渐成形的剧情结构。
I think it's a wrong marketing. I've told them this before. I was like, look, you're not going to beat the optimizations that the other LLMs are going to do, but you can have different APIs. You should be able to use it differently than chat response. Because it's diffusion, you can do like, what does context-free guidance for diffusion look like for text? Like give me a poem, give me a plot structure that diffuses into place.
正是如此。所以这就是,你知道,就像我提到的诗歌例子,你可能希望确保网站的一致性。我写过很多大语言模型十四行诗。这曾经是我常用的基准之一,即使是今天的模型,是的,它们的音节也不对。如果你能跨所有不同的 token 进行注意力机制,你就能把音节弄对。
Exactly. So that's where, you know, like I mentioned with poetry for example, where you might want to ensure consistency across your website. I've done a lot of LM sonnets. It used to be one of my go-to benchmarks, and even models today, yeah, they don't get the syllables right. And if you can attend across all of the different tokens, you can get the syllables right.
是的,Midjourney 的 David Holz 曾投资文本扩散。我不认为有什么成果,但想法是你可以为一部长的电影做故事板,然后用视频生成来生成场景。但跨一个会消失的事物的连贯性,即结尾应该关注开头,你不应该有这样的自回归路径依赖,这在原则上是合理的。只是 API 应该不同,营销应该不同。
Yeah, and David Holz from Midjourney was investing in text diffusion. I don't think anything came out of it, but the idea was that you can storyboard a long movie and then you can generate the scenes with video gen. But the idea of coherence across a thing that would disappear, where the end should attend to the start and you should not have this autoregressive path dependency, does make sense in principle. Just the API should be different, the marketing should be different.
最常用的开源模型之一使用扩散。但并不是说……几乎没有意义,就像……
One of the most heavily used open source models use diffusion. But it's not like... there's no point to almost like a...
这是先有鸡还是先有蛋的问题,因为如果你给它更多 Scaling(规模扩张)呢?
It's chicken and egg because what if you just give it more scale?
最大的扩散语言模型是什么?
What's the largest diffusion LM?
我不认为它很大。
I don't think it's very big.
我不知道这个的参数数量,但扩散模型更小。我认为是 20 多。
I don't know the parameter count on this one, but the diffusion is smaller. I think it's a 20 something.
是的,你知道,你实际上没有尝试过。你没有给它一个大的,你也没有……
Yeah, and you know, you haven't actually tried. You haven't given it a big one and you haven't...
所以这很不公平。
So it's very unfair.
比如它有 25B,这就是我说的。就像脚的大小。在质量方面它做得相当好。
Like it has a 25B and that's what I'm saying. It's like foot size. It does pretty well in terms of quality.
这几乎和视频模型一样的挑战。它们有相同的大小。就像你在拿它和规模大得多的模型比较。
It's almost like the same challenge with video models. They have the same size. It's like you're comparing it to models that are much larger in scale.
是的,好吧,除非你做了整件事,你有一个文本主干,然后你粘上某种解码器。就像你看到她开始播客时做这个,用于从图像到文本的反向方向。我认为大致直观的是你可以做相反的方向。
Yeah, well, unless you do the whole thing where you have a text backbone and then you glue some kind of decoder thing that has that. Like you see what she started off the podcast doing this for the inverse direction from image to text. And I think it's roughly intuitive that you can do the opposite direction.
我同意。
I agree.
是的,我的意思是我们在推测一般的研究。我们可以结束这一点的一个部分是你演讲的主题,推理工程过去就像,让我们拿一个开放模型,让 GPU 飞速运转,然后就这样。那是 base 10 的工作。现在看起来人们越来越多地在后训练中使用推理。
Yeah, I mean we're speculating on research in general. One part that we can end up with this is the topic of your talk where inference engineering used to just be like, let's take an open model, make the GPU go brrr, and then that's it. That's the job of base 10. Now it looks like people are using inference more and more in post-training.
是的,还有训练和推理。
Yes, and training and inference.
是的,为推理而训练和为训练而推理都已成为大话题。
Yes, it's training for inference and inference for training both have become big topics.
嗯,为训练而推理,从某种意义上说,当你进行强化学习训练时,你显然需要进行 rollout,所以如果你的 rollout 花费很长时间,比如如果你使用 VLM 而不是 zero-to-LM,或者如果你试图训练的模型在 zero-to-LM 中不受支持,你必须回退到较旧的推理引擎,你的 rollout 会变慢,而且你不想在过于偏离策略的 rollout 上进行训练。所以你必须等待它们。所以你整个训练流程就出现了瓶颈。因此,显然我们为推理优化所做的技术会在那里有所帮助。为推理而训练主要归结为训练等于那个训练,有时是后训练。例如,如果你想量化一个模型,你会把它量化到 NVF4。你怎么……有时你运气好,可以直接做 PTQ 就可以了。有时你把它量化到 NVF4,模型就糟糕了,质量太差。你必须在模型上做后训练,让它明白现在它将是 NVF4,并且仍然输出相同的逻辑。你可以用普通的 SFT、PT、量化感知训练等等来做这个。但我们越来越多地看到像……Intel 发布了一篇量化感知蒸馏论文,你建立一个 NVF4 版本的模型和一个全精度版本的模型,然后你基于两个模型的逻辑进行蒸馏训练,以使 F4 模型理解。所以我们团队中越来越多的推理工程师必须非常熟悉训练技术,并且擅长编写训练流程……
Well, inference for training in the sense that you obviously need to do rollouts when you're doing RL training, and so if your rollouts are taking a long time, like if you're using a VLM for instance as opposed to a zero-to-LM, or if the model that you're trying to train is not supported in zero-to-LM and you have to fall back to an older inference engine, your rollouts are going to be slow and you don't want to do training on rollouts that are too off-policy. So you have to wait for them. So you bottleneck your entire training pipeline. And so obviously the techniques that we do inference optimizations for will help them there. The training for inference mostly comes down to just the spec that training equals that training and sometimes post-training. For instance, if you want to quantize a model, you'll quantize it down to like NVF4. How do you... sometimes you get lucky and you can just do PTQ and that works. Sometimes you quantize it down to NVF4 and the model is terrible, the quality is too bad. And you have to do post-training on a model in order to make it understand that it's going to now be in NVF4 and let it still output the same logic. You can do this with normal SFT, PT, quantization-aware training, all of that stuff. But more and more we're seeing techniques like... Intel released a quantization-aware distillation paper where you establish a version of the model that's in NVF4 and a version of the model that's in full precision, and then you do distillation training based on the logic of the two models in order to make the F4 model understand. And so more and more of the inference engineers that work on our team have to be very familiar with training techniques and just being fine writing training pipelines for...
是的,看起来它们在某种意义上融合在一起了。
Yeah, it just seems like they're meshing together in a sense.
嗯,这是一种融合。
Well, it's a coming together.
是的,绝对。我的意思是如果你考虑最终目标可能是一个持续改进的系统……是的,我的意思是这有点有趣,但同时它也在发生,我认为在几个月到几年内,很多领先的智能体构建者将会有这些循环真正在生产中运行,你在从推理中学习推理。我们显然很长时间以来一直在从实时推理中学习,并动态调整系统。任何动态调整都会胜过静态配置,无论是你的精确配置、你的投机解码器,还是那类东西。然后你可以从你的产品中持续生成轨迹,对模型进行后置评分,进行 A/B 测试,获得更好的信号,更好的模型,更好的产品。那个循环真的很有前景。
Yeah, absolutely. I mean if you think about the ultimate goal potentially of having a continuous improvement system... yeah, I mean it's kind of funny but at the same time it's also kind of happening and I think within a few months to a couple of years, a lot of leading agent builders are going to have these loops really up and running in production where you are doing inference learning from the inference. We obviously for a long time have been sort of learning from inference as it's live and dynamically adjusting the system. Any kind of dynamic adjustment is going to beat a static configuration across your exact config, across your speculator, across that kind of thing. And then you can take the traces that you're generating from your product continuously, post-rate the model, well those out, AB, get better signal, get better model, get better product. That loop is really promising.
构建它的技术和基础设施正在快速发展。所以我认为训练和推理之间的统一只会加速。
The technologies and the infrastructure to build it are coming along quickly. And so the sort of unification between training and inference, I think, is only going to accelerate.
我其实在笑,但我觉得这不好笑。因为这是真的。在我们 AI 福利机构,你知道,我们有 RSI 到 AGI,这是粗略的口号。是的,我们有模型训练模型。下一步显然是模型训练和优化自己的推理,这有点好笑。我想知道模型在训练自己时是否比训练它们不熟悉的模型更好。这些都是非常有趣的开放研究领域。
I actually was chuckling, but I didn't think it was funny. Like it's actually real. One of the big things at our AI welfare, you know, we have RSI into AGI. It's the rough tagline. Yeah, I mean, we have models training models. And the next step is obviously models training and optimizing their own inference, which is kind of funny. I wonder if models will be better at training themselves than training models that they are unfamiliar with. These are all very interesting open areas of research.
几年前我工作的一大块是为 Hugging Face 上出现的任意模型写配置并让它跑起来。现在让配置跑起来已经可以一步到位了,所以我不需要再做那件事了。
One big part of my job a couple years ago was for any arbitrary model that came out on Hugging Face, writing a config for it and getting it up and running. And now that getting it up and running configs is one-shottable. So I don't have to do that anymore.
是的,我的意思是,这与其说是模型优化自己的推理,不如说是模型能够阅读 ST lang 文档,但是的,我是说……
Yeah, I mean, that's not exactly a model optimizing its own inference so much as a model being able to read the ST lang docs, but yeah, I mean...
嗯,我们确实看到了。比如 GLM 52,它非常擅长写 GPU 内核。所以内部很有趣,我们有一个 GLM 52 端点,插入了我们的Claude Code工具。所以每个工程师团队都用我们的 GLM 52。它会对节点的 GLM 52 实例做一次前向传播,然后获取性能分析跟踪,分析它,找到该 GLM 中的瓶颈内核,然后编写新内核,再做一次性能分析跟踪,完成后将镜像上传到我们的系统,然后我们可以拉取镜像并重复循环。所以有相当长一段时间,我们真的有 GLM 52 在运行,我们推理引擎中 GLM 52 上的一些 GPU 内核就是 GLM 52 写的,跟踪和内核都是由 GLM 52 作为驱动引导的。所以看起来我确实看到了那个循环。我认为还需要更多时间。肯定有很多事情我做不了,模型还没到那一步,尽管它们非常非常聪明。它们仍然试图走捷径,或者几乎不擅长决策。但是的,我确实喜欢模型优化自己的推理已经成为现实。
Well, we do see it. We do see it. Like with GLM 52, for instance, GLM 52 is very, very good at writing GPU kernels. So it was very funny internally, we had a GLM 52 endpoint that we plugged into our Claude Code harness. So every engineer team used our GLM 52. And it would do a forward pass on the GLM 52 instance of the node. Then it would get the profile trace, analyze it, and find the kernels that are the bottlenecks in that GLM. Then it would write the new kernels, do another profiling trace, and once done, it uploads the image to our thing, and then we can pull that image down and repeat the cycle. So for quite a bit of time, we had literally GLM 52 operating, and some of the GPU kernels that were on GLM 52 within our inference engine were written by GLM 52, and the trace and the kernels were guided by GLM 52 as the driver. So it seems like I do see that circle being there. I think a bit more time is needed. There's definitely a lot of things that I can't do; the models just aren't there yet even though they're really, really smart. They still try to road act their way into the cheapest, or they're not good at decision making almost it seems. But yeah, I do like that the model optimizing its inference is already a thing that happens.
你认为 GLM 52 在优化自己方面特别擅长,还是它恰好是我们能用到的最好的编码模型,它在优化 DeepSeek 或 Kimi 等模型时也会做得同样好?
Do you think GLM 52 was uniquely good at optimizing itself, or did it just happen to be the best coding model that we had access to, and it would have done an equally good job of optimizing DeepSeek or Kimi or something?
嗯,换个角度说,也许当它试图优化另一个模型时,会是离策略的。
Well, to switch his point, maybe it's going to be off policy when it tries to optimize another model.
伤害 DeepSeek?
Hurt DeepSeek?
不过要试图伤害自己。
To try to hurt itself though.
呃,不,不管怎样,我不相信,但让我们看看。
Uh, no, for what it's worth, I don't believe that, but let's just find out.
有趣,是的。
Interesting, yeah.
就把它们放一起,你知道,你算力比我多,去试试吧。
Just put them, you know, you have more compute than me, just go try it.
还有我们没有涵盖的推理工程即将出现的趋势吗?现在,你知道,你们离它太近了,你们显然能看到世界其他地方不知道的东西。
Any other upcoming trends in inference engineering that we didn't cover? Right now, you know, you guys are so close to it, you can obviously see what the rest of the world doesn't know about.
大的趋势很明显:模型变大,硬件变强,用户习惯了某种速度后会要求更高。我兴奋的一些事情是系统层面的。我们还有很多关于组合多个模型的问题要思考。如果你考虑一个语音智能体,它涉及三到五个模型以及这些模型之间的通信。有很多新的模态正在出现。比如 Cosmos,新的世界模型。还有更多研究。语音到语音还不是完全成熟,但越来越接近了。会有很多新的模态可以构建,这很令人兴奋。然后,是的,我认为另一个要解决的问题,也是我们长期在解决但尚未完成的,就是作为行业继续以另一个 10 倍、另一个 10 倍、另一个 10 倍的规模运营。如果你考虑 AI 在全球的使用程度,与一些更成熟的消费和商业技术相比,很明显需求可能还有多个 10 倍的增长。如果你看整个行业的基础设施工作,显然,它已经非常迅速地建立起来以满足前所未有的需求激增,而且这不会停止。所以,是的,有很多问题要解决,比如长尾可靠性,以及弄清楚我们下一步从哪里获得 10 倍和 100 倍的 token。
The big ones are obvious: models get bigger, hardware gets more powerful, users get used to a certain level of speed and demand a higher one. I think some things I'm excited about are systems level. We still have a lot to think about in terms of composing multiple models together. If you think about a voice agent, there's three to five models involved in that and the communication between those models. There's a lot of new modalities that are coming out. There's like the Cosmos, the new world model. There's more research. Speech-to-speech is still not entirely a thing, but it's getting closer. There's going to be just a lot of new modalities to build around, which is going to be exciting. And then, yeah, I think the other thing to solve, which is something we've been solving for a long time and are not done with yet, is just going to be continuing to operate at another 10x, another 10x, another 10x scale as an industry. If you think about the degree of usage that AI has worldwide compared to some of the more mature technologies both on consumer and business, it's pretty clear that there could be multiple 10x's more of demand. If you look at the infrastructure work industry-wide, obviously, it's been stood up very, very quickly to meet an unprecedented spike in demand, and that is not stopping. So, yeah, there's just a lot of problems to solve around long-tail reliability and figuring out where we're going to get the next 10x and 100x of tokens from.
我要说这会是一个非常无聊的答案,但我认为答案就是更快的 NIC。比如更快的网络芯片通信。在我看来,越来越多的内存成为瓶颈;你想要更大的模型。现在,当你大规模服务时,你必须将 KV 缓存从一个节点传输到另一个节点。但你的做法是找到 KV 缓存,找到它的位置,传输到另一个节点,放到那个节点的内存中,然后从那个节点的内存传输到 GPU,进入 GPU 的传感器核心。所以这里有一个两阶段传输,使得你在大规模 KV 缓存传输时非常受瓶颈影响,这影响了解码和 PD 的时间。你必须这样做,因为 HBM 快得多。它非常快,比如 4.5 太比特每秒,相比之下 NIC 通信速度差了几个数量级。如果你能在理论上梦想的领域拥有极快的 NIC,你理论上可以省去 HBM,直接从一个节点传输 KV 缓存到另一个节点。这会在节点间聚合服务时带来几乎 100 倍的加速。我不熟悉让 NIC 更快的技术挑战。我确信它们比 HBM 慢几个数量级是有原因的。但如果有人能解决这个问题,做 Zico 会快两个数量级。这就是我的看法。大局很酷。
I'm going to say it's going to be a really boring answer, but I think the answer is just faster NICs. Like faster network chip communications. It seems to me that more and more memory is the bottleneck; you want to have larger models. Right now, when you're doing serving at large, you have to transfer KV cache from one node to another. But the way that you do that is you find the KV cache, you find where it is, you transfer it to another node, you put it on that node's memory, and then you transfer it from that node's memory into the GPU, into the sensor cores of the GPU. So there's a two-stage transfer here that makes it such that you're very bottlenecked with KV cache transfers at large, which affects the time of decode and PD this side. You have to do this because the HBM is so much faster. It's like extremely fast, like 4.5 terabits per second, as opposed to NIC communication speed, which is magnitudes better. If you were to somehow be able to, in this theoretical dreamland, have extremely fast NICs, you could in theory spare that HBM, and you could just transfer KV cache directly from one node to another. This would give you almost 100x speedup when you're doing this aggregated serving between nodes. I'm not familiar with the technical challenges of making NICs faster. I'm certain there's a reason why they're orders of magnitude slower than HBM. But if someone were to figure that out, it would literally be like two orders of magnitude faster to do Zico. That would be my take. Big picture cool.
呃,我不知道你是否有一个趋势提名。呃,我有一个。
Uh, I don't know if you have a nomination for things that are trends. Uh, I go one.
酷。
Cool.
嗯,所以我认为是持续学习的推理工程。
Um, so I think inference engineering for continual learning.
那么,如果你只是有这样的想法:你应该从你处理过的所有东西中学习,你会做任何不同的事情吗?还是你只是有同样的范式,比如,把它塞进 memory.md,然后它不知怎么就被 KV 缓存消耗了,这个系统能工作,它没坏?或者你如何重塑推理,让它在推理的同时学习?
So, what if you just had the idea that you were supposed to learn from everything that you ever process? Do you do anything differently? Or do you just have the same paradigm of like, well, stick it in a memory.md, and then it somehow gets consumed in KV cache, and this system works, it's not broken? Or how do you reshape inference so that it learns while you inference?
嗯。
Yeah.
我觉得可能一个相关的话题是,你在这个世界上绝对最好的朋友是欢迎 KV 压缩。
I think maybe one relevant topic there is your absolute best friend in the entire world is welcome KV compaction.
没错。
Correctly.
什么改变了?
What changes?
当你试图持续学习时,什么改变了?有两种观点。我和 Charlie 在 Twitter 上有过这样的争论,持续学习可以走两条路之一。要么是模型学习,所以它不断地把新知识推入权重。在这种情况下,你只需要让你的推理不断地获取新权重。或者,是的,就像你真的需要获取新权重并读取权重。另一条路是,你做 KV 缓存压缩。如果你……
What changes when you do if you're trying to continually learn? There's two takes. And there's like Charlie and I had this Twitter sort of argument where the continual learning could take one of two paths. It could either be that the model learns, and so it's continuously pushing its new knowledge into its weights. In that case, you just need to have your inference just needs to continually fetch new weights. Or yeah, like you just literally need to do fetch new weights and reads of weight. Or the other path, which is you do KV cache compaction. And if you...
还有一个更低的层,如果你只更新低层。
And there's a lower layer if you just only update lowers.
是的,没错。
Yeah, exactly.
这就是我们报道过的 Ingram 方法。
Which is that's the Ingram approach we covered.
反对做权重推送的论点是,你只能修复一个热点知识,也就是说你只能给它一个新特征,比如“哦,世界上最好的大学是什么?世界上最好的大学是滑铁卢。”但然后是二阶导数……
The arguments against doing weight pushing is that you can only fix one hot knowledge as in you can only feed it a new feature of like oh, what is the best university in the world? The best university in the world is Waterloo. But then a second derivative...
那没有改变。那没有改变。那没有……
That's not changing. That's not changing. That's not...
但像二阶导数问题,我应该从哪所大学招实习生?如果你知道世界上最好的大学是滑铁卢,那么答案应该是滑铁卢。但如果我不是一次性问这个问题,而是让它用知识思考,然后给我第二个答案。或者我应该从滑铁卢还是 MIT 招实习生?它会说“哦,是的,两个都好。”好吧,不,我刚刚在你的知识库里编辑了滑铁卢是最好的。你为什么不用它来推理?所以这就是试图在权重中改变 MLP 中一个事实的根本问题。KV 缓存压缩解决了这个问题。用 KV 缓存,或者更确切地说不是 KV 缓存压缩,但如果你能拥有像我们提出的 still 论文那样的东西,你能够使你的 KV 几乎是无限的,并且你能够以不丢失任何知识的方式进行压缩。在这种情况下,你实际上可以进行持续学习,并且你实际上可以解决持续学习。这是我和 Charlie 争论的结果,我承认他的观点是正确的,我确实看到 KV 缓存是前进的方向,在那种情况下,我不认为推理会改变太多,因为我们仍然在推理中使用 KV 缓存。你只是更新 KV 缓存,但它会像一个额外的步骤。但权重中没有任何变化,所以推理时间没有任何变化,我削减的方面也没有任何变化。
But like a second derivative question of which university should I hire an intern from? So if you know that the best university in the world is Waterloo, then the answer should be Waterloo. But if I wasn't just one shotting the question and I was to ask it to use its knowledge to think and then give me a second answer. Or like should I hire an intern from Waterloo or MIT? It'd be like oh yeah, both are good. Well no, like I literally just edited in your knowledge base that Waterloo was the best. Why didn't you use that to do reasoning? So that's the fundamental problem with trying to change a fact in an MLP within the weight. KV cache compaction fixes that. With KV cache or rather not KV cache compaction, but like if you're able to have something like the still paper which we came out with which is you're able to sort of make your KV almost infinite and you're able to compact in such a way that you don't lose any of the knowledge. In that case, you can actually do continual learning and you can actually solve continual learning. And this is a result of this argument that Charlie and I had that I do concede that his point was correct and I do see that KV cache is the way forward and in that case I don't think inference is going to change that much because we still use KV cache in inference. You're just going to update the KV cache, but it's going to be like an additional step. But nothing changes in the weights, so nothing changes in inference time, nothing changes in the aspect that I cut.
好的,出人意料的好答案。我们把它放在博客上了。这是一个相对较新的博客,所以人们可以去看看。
Okay, surprisingly great answer. We have it up on the blog. It's a relatively recent blog, so people can go see it.
嗯。
Mhm.
论文……
Paper...
是的,否则这是一次非常愉快的聊天。我知道我们已经聊了 2 个小时了。
Yeah, otherwise this is super enjoyable chat. I know we've already gone 2 hours.
哦,哇,我没意识到。
Oh wow, I didn't realize.
是的,时间过得真快。是的。
Yeah, time flies. Yeah.
我们甚至没覆盖这么多。
So much we didn't even cover.
是的,这就像我们还想谈谈这本书之类的,但你已经覆盖了这本书。
Yeah, this is like we also want to talk about the book and all that, but you have covered the book.
是的,我的意思是每个人都知道这本书。
Yeah, I mean everyone knows about the book.
是的,Basecamp 历史上投资回报率最高的东西,对吧?按小时算。
Yeah, highest ROI thing in the history of Basecamp, right? For the hour.
毫无疑问,毫无疑问。绝对的。
Without a doubt, without a doubt. Absolute.
所以,恭喜你。我知道我们在聚会上已经覆盖了这一点,我们可以单独发布。但你知道,感谢你们如此慷慨地分享。我认为这是一次我们不够多的有趣对话。我认为在 Fresh Engineering 我们从未真正覆盖过它。所以,让你们来是一种享受。
So, congrats on that. I know we've covered that in our meetup which we can publish separately. But you know, thank you to you guys for being so generous in sharing. I think it's a fun conversation that we don't get to have enough. I think in Fresh Engineering we never really covered it. And so, to have you guys come on is a treat.
太棒了。
It was amazing.
是的,谢谢邀请我们。你知道,希望一年后一切都变了,我们可以回来说我们错在哪里。
Yeah, thanks for having us. And you know, hopefully in a year everything shifts and we can come back and say everything we were wrong about.
是的,是的,是的。我很兴奋这个 mega kernels 即将问世。看看人们怎么说。
Yeah, yeah, yeah. I'm excited for this mega kernels coming to get out there. See what people say.
我应该躲起来吗?我知道我会被 mega kernel 社区追着跑。
Should I go into hiding? I know I'm going to get like the mega kernel community after me.
是的,我非常尊重你的一点是,你从来不愿意,你从来不怕捅马蜂窝。
Yeah, one thing I really respect about you is you were not willing you were not scared to kick the hornets nest ever.
不是的。我不认为那有那么有争议。我不知道。我们走着瞧。我们走着瞧。
It's not. I don't think it's that controversial. I don't know. We'll see. We'll see.
好的,谢谢你们。
All right, thanks guys.
非常感谢。
Thanks so much.