AI 中最快的时间线:推理工程

The Fastest Timeline in AI: Inference Engineering

菲利普·基利 Philip Kiely · TWIML AI 播客 · 2026-04-30 · 约 54 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Philip Kiely 讨论了为什么推理是 AI 中最重要且最具粘性的工作负载,以及 Baseten 如何转向专注于推理。

Philip Kiely discusses why inference is the most important and stickiest workload in AI, and how Baseten pivoted to focus on it.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 16)

全文 · Full transcript(中英对照)

开场与签书会 Introduction and Book Signing

Host

好的,各位。欢迎收听 Twilio AI 播客的另一期节目。我是主持人 Sam Charrington。今天请到的是 Philip Kiely。Philip 是 Baseten 的 AI 教育负责人。在开始之前,请务必花点时间在您收听本期节目的任何平台上点击订阅按钮。Philip,很高兴再次见到你。欢迎来到播客。

All right, everyone. Welcome to another episode of the Twilio AI Podcast. I am your host, Sam Charrington. Today, I'm joined by Philip Kiely. Philip is head of AI education at Base 10. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Philip, great to see you again. Welcome to the podcast.

Philip

嘿,Sam,谢谢你邀请我。

Hey Sam, thanks for having me on.

Host

当然,当然。我们最近在 GTC 大会上见过,当时你在签售那本书,就在你左肩上方的那本。

Absolutely, absolutely. We met at the recent GTC conference where you were signing copies of the book that you have there, over your, I guess, your left shoulder.

Philip

是的,我桌上也有一本,就在我面前。

Yeah, and I've got one here on my desk in front of me.

Host

没错。那真是一次有趣的会议。我感觉自己像个摇滚明星,你知道吗,到处发书。

There you go. That was a really fun conference. I felt like a rock star, you know, handing out books.

Philip

说实话,我觉得部分原因是我们展台还有免费冰淇淋,我送出了很多书,但我在展会结束时算了一下,冰淇淋更受欢迎。

Honestly, I think part of it was we had free ice cream at the booth as well, and I gave away a lot of books, but I did run the numbers at the end of the show—the ice cream was more popular.

Host

你知道,我觉得请你来聊聊推理工程、你在书中写的内容,以及你在书之外看到的东西,还有如今从推理角度来看真正产生影响的东西,会很好。不过在我们进入正题之前,我想请你谈谈你的经历,是什么让你来到 Baseten 并从事推理工程。

You know, I thought it'd be good to get you on and talk through, like, a little bit about inference engineering, what you're writing about in the book, but also kind of what you're seeing beyond the book, and what is really making a difference from an inference perspective nowadays. So, before we get to that though, I'd love to have you talk a little bit about your journey, like what brought you to Baseten and inference engineering.

Philip

所以,回到 2022 年,那是在 ChatGPT 公开发布前 10 个月。我当时对 AI 行业有一个非常强烈且新颖的论点。那个强烈且新颖的论点是:我需要一份工作。于是我申请了一堆工作,最终在 Baseten 和一家洗衣配送初创公司之间做选择。那真是个艰难的决定,因为另一家也是好公司,他们现在还在。但最终,我决定去一家当时叫 Baseten 的小初创公司,解决机器学习领域一些非常有趣的问题。到现在我已经在那里干了 4 年多了。我非常幸运,能在工作中学习,亲身经历 AI 行业的崛起——从预测性建模到生成式建模,从建模主要用于内部工作流,到建模成为销货成本的一部分,建模进入 AI 原生公司的收入路径。也正因如此,从非常非常小的用例,发展到非常非常大且复杂的用例。

So, back in 2022, this was 10 months before ChatGPT was launched to the public. I had a really strong and novel thesis about the AI industry. That strong and novel thesis was: I needed a job. So I applied for a bunch of jobs, and I ended up choosing between Baseten and a laundry delivery startup. It was a really tough decision, because, hey, the other one's a good company, they're still around. But ultimately, I decided to go work on some really interesting problems in the ML space with a tiny startup at the time called Base 10. And I've been at it for more than 4 years now. I've had a really fortunate position to be able to learn on the job and experience the rise of the AI industry firsthand—going from predictive to generative modeling, going from modeling being something mostly done for internal workflows to modeling being part of the cost of goods sold, modeling being in the revenue path for AI-native companies. And going from very, very small use cases because of that, to very, very large and complex ones.

推理公司为何存活 Why Inference Companies Survived

Host

是的,关于你的经历,有一点很有趣,我特别想到 Baseten——就像你的时间线所表明的,Baseten 早于 ChatGPT,当时它是众多在 ML Ops 理念上竞争的公司的之一。随着生成式 AI 的兴起,很多这样的公司被收购,很多消失了。但那些仍然存在且表现不错的公司,很多都在做推理。这是为什么呢?

Yeah, one thing that's interesting about your experience, and I'm thinking about Baseten in particular—like Baseten, as your timing indicates, predated ChatGPT, and they were one of a great number of companies at the time kind of competing on ML Ops set of ideas. And with the rise of generative AI, a lot of those companies got acquired, a lot of those companies went away. But those companies that are still around and doing well, like a lot of them are in inference. Like, why is that?

Philip

所以,我相信推理是 AI 中最重要的负载,而且它绝对是最具粘性的。这也是我们把它作为进入 AI 基础设施市场的楔子的原因之一。推理非常非常复杂,做好它对于任何真正依赖生成模型来驱动产品的公司都至关重要。所以,我们一开始就是一家推理公司,但我们当时推理的是非常不同的模型。我当时拿着 CPU,有时拿一些 T4 GPU 或 A100(当我感觉特别高级的时候),在上面跑 XGBoost 模型,跑非常早期的 Wave2Vec 和 GPT-J 类模型,说实话玩得很开心。我们在构建简单的分类器、仪表盘之类的东西。后来随着时间的推移,模型变得越来越大、越来越复杂——我记得 Whisper 发布时,那太震撼了,因为,天哪,现在模型有 10 亿参数?你能想象一个模型就有 10 亿参数吗?我简直不敢相信我会把所有的参数都花在一个地方。但这个模型现在真的很有能力,但它突然需要 GPU,需要很多 GPU,而且很难让它跑起来,很难让它足够快。运行模型突然出现了很多有趣的问题。所以,我认为很自然地,随着这些模型变得更大,更广泛的 ML 服务栈中最复杂的部分就变成了推理。然后当真正的生成模型出现时,我们已经在以推理为先的视角看世界了,自然地将其扩展到越来越大、越来越有能力的模型上。

So, inference is, I believe, the most important workload in AI, and it's definitely the stickiest. That's been one of the reasons we've treated it as a wedge into the AI infrastructure market. Inference is very, very complicated, and doing it well is really critical to any company that actually relies on generative models to power their product. So, we started as an inference company, but we were just inferencing very different models. I was, you know, taking CPUs and I was taking some T4 GPUs or A100s when I was feeling real fancy, and running XGBoost models on them, running really early Wave2Vec and GPT-J type models, and honestly just having a lot of fun with it. We were building simple classifiers, dashboards, that kind of stuff. And found over time that as the models were getting a little bit bigger and more complicated—like I remember when Whisper came out, that was huge because, oh man, now the model is a billion parameters? Like, can you imagine 1 billion parameters in just one model? Like, I can't believe I'm going to spend all my parameters in one place. But this model now, it's really capable, but now it needs GPUs, you need a lot of them all of a sudden, and it's hard to get it running, it's hard to make it fast enough. There's a lot of interesting problems all of a sudden in running models. And so, I think it just kind of naturally occurred that the most complicated piece of the broader ML serving stack became inference as these models got larger. And then as true generative models came around, we were already thinking about an inference-first view of the world, and it felt natural to just extend that to larger and larger and more capable models.

推理与模型服务对比 Inference vs. Model Serving

Host

你提到了模型服务和推理作为模型服务的一个组成部分。那么,你在哪里划定推理的起点和终点,以及更广泛的模型服务需求之间的界限?

You reference both model serving and inference as a component of model serving. Like, where do you draw the line between the beginning and end of inference and the broader model serving set of requirements?

Philip

我可能用这些术语用得有点太随意、太混用了,因为技术定义上,模型服务是从用户请求到用户获得响应的端到端的所有事情。而推理是发生在 GPU 上的那部分。但当我将推理视为一门学科时,我考虑的是端到端。我认为作为一个推理系统,你必须拥有整个体验。所以,要做到这一点,你不仅要考虑 GPU 上发生的事情,还要考虑用户旅程中每一步发生的事情。

I probably use those terms a little too loosely and a little too interchangeably, because the technical definition would be that model serving is everything that happens end to end from a user request to a user getting a response. And inference is the piece that happens on the GPU. But when I think about inference as a discipline, I am thinking about end to end. I think as an inference system, you have to own the entire experience. And so, to do that, you have to think not only about what's happening on the GPU, but what's happening every step of the way on the user journey.

推理中的挑战 Challenges in Inference

Host

是的,那让我们更深入地谈谈是什么让推理变得困难。你知道,显然,我们知道模型越来越大。我们知道我们没有足够的算力,无论是 CPU、GPU 还是其他形式。请谈谈你在做好推理方面看到的其他一些挑战。

Yeah, so let's talk a little bit more deeply about what makes inference difficult. You know, clearly, we know the models are getting bigger. We know we don't have enough compute in the form of CPU, GPU, etc. Talk a little bit about some of the other challenges that you see with regards to doing inference well.

Philip

是的,推理是一个非常有趣且困难的话题,我在这里要稍微进入元领域一下。那就是,我是一名武术家,我一生都是。你熟悉 UFC 和 MMA 这类东西吗?

Yeah, so inference is a really fun and difficult topic, and I'm going to go into meta territory for a second here. Which is that I'm a martial artist, I have been my entire life. Are you familiar with UFC and MMA and all those kind of things?

Host

嗯。

Mhm.

Philip

所以,你知道,在那里你不能只擅长一件事就期望成为冠军。

So, you know, there you can't just be an expert in one thing and expect to become a champion.

推理如综合格斗 Inference as Mixed Martial Arts

Philip

你不能只做一个伟大的摔跤手,也不能只做一个伟大的拳击手。关键在于,你要掌握很多不同的技能,每一项都可能需要一生去精通,而你必须全部做到出色,才能成为一个全面的综合格斗选手。你知道,就我的背景而言,我肯定不是 UFC 冠军,但我练过各种流派的打击技,也练过各种流派的摔跤,我开始理解这些东西是如何融合在一起的。我认为推理(inference)非常相似。你需要拥有涵盖许多截然不同的复杂主题的广泛专业知识,才能构建一个真正有效的推理系统。有 GPU 上发生的事情,对吧?需要理解 CUDA 级别的编程,以及其上的 PyTorch,再往上是推理引擎。还要能够应用研究,理解不同的量化技术、不同的投机采样算法、不同的 KV 缓存复用机制。如何在多个 GPU 之间并行化模型?如何在不同类型的硬件上进行分离?所以,有所有这些应用研究。除此之外,还有运行超大规模分布式系统的所有传统问题。你过去 20 年在云架构中熟悉的任何东西,以及如何妥善处理来自全球的超大流量工作负载。所以,每一块后端 Web 基础设施,每一块 GPU 编程,都在推理中汇聚在一起,而这些窗口和包的要求非常苛刻。比如你经常要处理几百毫秒的延迟 SLA。因此,正是对大量非常复杂的系统的编排,才让这个问题如此困难。

You can't just be a great wrestler, you can't just be a great boxer. The idea is that you have a lot of different skills, each one of which can take a lifetime to master and somehow you have to be excellent at all of them in order to be a well-rounded mixed martial artist. And you know, in my background, I'm certainly no UFC champion, but I've done striking of various disciplines, I've done wrestling in various disciplines and I've started to understand how those things blend together. I think inference is very similar. You need to have a wide range of expertise on a lot of very different complicated topics to build a truly effective inference system. There's what happens on the GPU, right? There's understanding CUDA level programming and PyTorch on top of that and then the inference engines on top of that. There's being able to take and apply the research, the understanding of different quantization techniques, different speculation algorithms, different KV cache reuse mechanisms. How do you parallelize models across GPUs? How do you disaggregate across different types of hardware? So, there's all that applied research. And then on top of that, there all the traditional problems of running very large scale distributed systems. Anything that you're familiar with from the last 20 years of cloud architecture as well as understanding how to, you know, take very very large volume workloads that are coming from around the world and handle them appropriately. So, every piece of back end web infrastructure, every piece of GPU programming, it all comes together in inference in windows and packages that are very demanding. Like you often have a couple hundred milliseconds latency SLAs that you're dealing with. So, it's the orchestration of a large number of very complex systems together that makes it such a difficult problem.

推理从研究到生产 Rapid Research to Production in Inference

Host

在 AI 领域,研究与实现之间有着非常紧密的关系。但特别是在推理方面,我多次注意到,比如你某天听说投机解码的工作,两周后我就到处听到它,而且每个人都已经在用了。你觉得推理领域从研究到生产的周期,相对于 AI 的其他方面,是不是特别快?

In AI broadly, like there's this very tight relationship between research and implementation. But in inference in particular, like I've noted on several occasions, like you know, hear about the speculative decoding work one day and like I hear about it all over in 2 weeks and everybody's already using it. Like do you find that that research to production timeline in inference is, you know, particularly rapid relative to other aspects of AI?

Philip

我的意思是,这可能是世界上最快的时间线。比如想想医药行业,研究可能需要几十年才能进入药房。再想想物理学或工程学,应用一个新概念可能需要数年,材料科学也是如此。即使在 AI 内部,什么更快?比如训练,如果你想用新技术训练模型,仍然需要数周或数月来微调超参数,找到表达该技术的正确方式。但在推理方面,时间线通常是几小时。新的模型架构一出来,你必须在第一天就支持它。Polo Quant 那篇研究论文发布后,我们模型性能团队的一名工程师在 31 小时后就用 CUDA 内核实现了它。所以,应用研究的步伐非常非常快,因为这是一个竞争激烈的行业,每个人都在寻找下一个优势。也许唯一比推理更快将研究投入生产的行业可能还是交易。

I mean, it might be the fastest timeline in the world. If you think about medicine, for example, it can take decades for research to reach a pharmacy. If you think about, you know, physics or engineering, it can take years to apply a new concept, material science. Even within AI, what moves faster? Training ones, for example, if you want to train a model off of a new technique, it can still take weeks or months to fine-tune the hyper parameters and find the exact right way to sort of express that technique. But with inference, the timeline is often hours. A new model architecture comes out, you have to figure out how to support it day zero. We had Polo Quant come out, that research paper, and an engineer on our model performance team had it implemented 31 hours later as a CUDA kernel. So, the pace of applied research is very very fast because it is a highly competitive industry and everyone is sort of searching for that next edge. Maybe the only industry where research goes into production faster than inference might still be like trading.

Host

澄清一下,是交易(trading),不是训练(training)。

To be clear, trading with a D, not training.

Philip

是的。

Yeah.

Polo Quant 背景故事 The Polo Quant Backstory

Host

这可能有点太深入细节了,但关于 Polo Quant,论文已经发表一年了,然后突然在上周或几周前变得非常流行。这背后的故事是什么?

This might be a little bit too in the weeds, but with regards to Polo Quant, like the paper was a year old and then all of a sudden it got really popular last week or a couple weeks ago. Like what was what's the backstory there?

Philip

是的,它是以博客文章的形式重新发布的,引起了大家的注意。我想稍微绕一下,你知道,我在爱荷华州长大。所以,我深受中西部最好的经济学思想流派的影响,那就是芝加哥经济学派,即著名的有效市场假说的创造者。我认为你能让一个易受影响的年轻人接触的最危险的谬误就是有效市场假说。那种认为一切都已经搞清楚、每个人都已经知道一切的想法,因为事实并非如此。你知道,随着大量研究技术的涌现,我认为很容易忽略一些东西。但同样,推理世界中很多容易摘的果实已经被摘掉了。所以现在,当我们试图找出下一步来突破极限时,有时那些即使在一年前看起来还不可行或不值得的技术,突然就变成了,哦等等,我们现在准备好了。现在是时候把它加入行业标准,推动推理栈前进了。只是那封邮件中的提醒让它重新引起了大家的注意,时机正好。它奏效了,你知道,一个没有病毒式传播的发布的好处是,如果你不再看它,就没人知道它。

Yeah, so it was sort of republished as a blog post and sort of caught everyone's attention. I think like as a little detour, you know, I grew up in Iowa. So, I was very influenced by sort of the best Midwestern school of thought in economics, which is the Chicago School of Economics. Famed creator of the efficient market hypothesis. I think the most dangerous falsehood that you can expose an impressionable young person to is the efficient market hypothesis. And the idea that like everything is already figured out and everyone already knows everything because that's just not true. You know, with the volume of research techniques that are coming out, I think that it's easy to lose track of things. But it's also a case where there's a lot of low-hanging fruit in the inference world that has already been picked. And so now, as we try to, you know, figure out what's next to push the envelope, sometimes techniques that might not have seemed, you know, either feasible or worthwhile even as recently as a year ago, all of a sudden become like, oh wait, now we're ready for this. Now it's time to add this to the sort of industry standard and move the inference stack forward. There was just that, you know, the bump in the email that brought it back to everyone's attention and the timing was just right. It worked, you know, the great thing about a launch that doesn't go viral is that no one knows about it if you don't watch it again.

谁该关注推理工程 Who Should Care About Inference Engineering

Host

没错。你知道,我们来谈谈谁应该关心或谁确实关心推理工程,比如对很多人来说,那是别人的问题,是少数大公司工程师的问题。但我怀疑有很多理由让更广泛的人群应该关心它。你怎么看?

Exactly. You know, let's talk a little bit about who should care or who does care about inference engineering, like for many you know, that's somebody else's problem, you know, and it's, you know, the problem of engineers at like a handful of really large companies. But I suspect that there are a number of reasons why a broader set of folks should care about it. Like what's your take on that?

Philip

是的。嗯,关心它的一个原因是,这是一个非常有趣的领域,你可能会考虑从事其中。你知道,3 年前,世界上可能只有几百个推理工程师。那时他们可能甚至不称自己为推理工程师。而现在有数千、数万,可能取决于你怎么界定。软件工程的工作方式当然正在经历巨大的变革,最优秀的工程师的生产力大幅提升。关于未来我们需要多少工程师、我们的工作会是什么样子,有很多问题。我相信,即使有了 AI 辅助代码生成的进步,即使个体工程师的生产力提高了,几年后对推理工程师的需求仍将是今天的 10 到 100 倍。我的意思是,我知道在 Base10,我们招不到足够多的懂推理的人。

Yeah. Well, one reason to care about it is that it's a really interesting field that you might consider working in. There was, you know, 3 years ago, there were maybe a couple hundred inference engineers in the world. They probably didn't even call themselves inference engineers back then. And now there are thousands, tens of thousands, probably, depending on, you know, where you make the cut-off. And software engineering is of course undergoing a massive change in the way it works and the best engineers are becoming substantially more productive. There's a lot of questions around how many engineers we're going to need in the future and what our jobs are going to look like. I believe that even with the advances in AI-assisted code generation, even with the increase in productivity on an individual engineering basis, there's still going to be demand for 10 to 100 times more inference engineers in a couple years than we have today. I mean, I know at Base10, like we can't hire people who are knowledgeable about inference fast enough.

推理知识为何重要 Why inference knowledge matters

Host

即使你不想在推理平台工作,关心它的另一个原因是,每个垂直 AI 应用公司最终都必须弄清楚他们的推理策略是什么,并且需要懂技术的员工,他们理解推理中涉及的实际技术和权衡,以便制定并执行一个称职的策略,因为出色的推理可以成就一个让用户感觉神奇的超快产品,也可能导致一个缓慢、漏洞百出、不断流失用户的体验。那么,我们来谈谈什么时候、为什么以及如何,了解推理会影响我做出的决策和我采取的策略或方法。很多人会把推理委托给提供商,无论是第一方模型提供商,还是像 Base10 或 Firefly 这样的第三方推理提供商。在哪些情况下,我需要深入了解推理?

And even if you're not wanting to work at an inference platform, the other reason to care about it is that every vertical AI application company is eventually going to have to figure out what their inference strategy is and is going to need knowledgeable members of technical staff who understand the actual technologies and trade-offs involved in inference to develop and execute a competent strategy, because great inference can be the difference between a really fast product that feels magical to users and a slow buggy experience that causes constant churn. So let's talk about when and why and how knowing about inference impacts the decisions I'm making and the strategy or approach that I'm taking. Like a lot of people will be delegating inference to a provider, whether it's a first-party model provider or a third-party inference provider like Base10 or Firefly or someone like you. What are the situations where I'm going to want to have deep knowledge about inference?

Philip

是的。所以我认为第一件事是理解什么是可能的。因为在大多数 AI 工程师都接受过 GPT 和 Claude 模型训练的世界里,你会以某种方式理解世界。我可以访问 N 个模型,也许有三五个或十个。每个模型都有固定的智能水平、固定的速度(以每秒 token 数、首个 token 时间衡量),也许是一个范围,固定的输入输出 token 成本,可能还有一些变量,比如缓存 token 价格或长期批处理任务的折扣。然后还有可靠性,你有一个端点,可能获得几个九的可用性,也许需要围绕这个现实来设计产品的其余部分。推理工程的第一课就是把所有这些视为不可变的离散点,而把它们更多地理解为一个频谱。例如,你可以在给定的推理引擎中通过调整简单的批量大小或添加或移除推测算法来权衡延迟和吞吐量。当你这样做时,当你创建了一个结果的频谱,一个高性能推理的高效前沿,你开始明白,等等,我可以选择改变我消费这些系统的方式。例如,你可能开始在产品中设置优先级队列,付费用户流量优先于免费用户流量。也许这是你以前无法构建的。或者你理解到我有能力量化模型,因为我有自己非常复杂且特定于产品的评估,我可以完全自信地这样做,而不会降低用户体验到的服务质量。然后当你量化自己的模型并开始以更快、更便宜的方式提供服务,并且非常有信心地校准好它时,你就不会受到第三方提供商随机的智能降级的影响,他们可能不了解你工作负载的确切性质。相反,你实现了他们追求的价格和性能目标,但实际上保持了适合你任务的东西。所以如果你对推理了解不多,你可能会想,哦,所有量化模型都不好,而事实上,你可以根据你的工作负载适当地校准它们。有各种各样的例子。甚至简单的事情,比如理解 KV 缓存重用的方式意味着你只需要考虑你的前缀,序列早期的一个 token 差异就可能破坏整个其余部分,这可能会影响你构建聊天模板的方式或你向模型输入的结构。有很多方式,理解推理的可能性可以帮助你构建更好的产品,并提供更好的用户体验。

Yeah. So I think the first thing is understanding what's possible. Because in a world where most AI engineers are sort of trained on GPT and Claude models, you understand the world in a certain way. I have access to N models and there's maybe three or five or 10. And each one of them has a sort of fixed level of intelligence, a fixed level of speed in terms of tokens per second, time to first token, maybe a range of that, a fixed cost in terms of input-output tokens with maybe a couple variables around something like a cash token price or a discount for long-running batch jobs. And then also a reliability, you know, you have an endpoint where you might get a couple nines of uptime and maybe need to engineer the rest of your product around that reality. And the first sort of lesson of inference engineering is going from taking all of these as sort of immutable discrete points and understanding them more as a spectrum. You can trade off between, for example, latency and throughput in a given inference engine by adjusting things as simple as batch size or things like adding or removing a speculation algorithm. When you do that, when you create sort of a spectrum of outcomes, an efficient frontier of high-performance inference, then you start to understand like, wait, I can choose to change the way that I consume these systems. So, for example, you might start having priority queues in your product where paid user traffic gets prioritized over free user traffic. And maybe that's something that you couldn't have built previously. Or maybe you understand like I have the ability to quantize models and because I have my own very sophisticated and product-specific evals, I can do that with complete confidence that I'm not degrading the quality of the service my users are experiencing. And then when you quantize your own model and start to serve it faster and less expensive and with a great deal of confidence that you've calibrated it properly, then you are not subject to sort of random intelligence degradations from a third-party provider who might not understand the exact nature of your workload. Instead, you've achieved the sort of price and performance goals that they're going for with that, but actually kept things appropriate for your task. So if you don't know a lot about inference, you might think like, oh, all quantized models are bad, when in fact, you can calibrate them appropriately to your workload. There's all sorts of examples like this. Even things as simple as understanding that the way KV cache reuse works means that you have to think about only your prefix and even one token difference early in a sequence can throw off the entire rest and maybe that affects how you structure your chat template or how you structure your inputs to your models. There's all sorts of ways where understanding what's possible with inference helps you build a better product and helps you deliver a better user experience.

控制层级与托管选项 Levels of control and hosting options

Host

你最后举的关于 KV 缓存重用的例子清楚地说明,无论你是否构建自己的推理系统,你都能从拥有这些知识中受益。很多其他例子,比如利用改变批量大小或不同服务层级,你是否必须托管自己的推理服务,在你自己的模型上做推理,才能利用这些,还是有提供商允许你这样做?我想象中,就像一切事物一样,当你做推理时,你能得到的旋钮数量是一个频谱。

The last example you gave about the KV cache caching is a clear example of it doesn't matter whether you're building out your own inference system, you benefit from having that knowledge. A lot of the other examples, do they, to take advantage of things like changing batch sizes or different tiers of service, do you have to host your own inference service, do inference on your own models, in order to take advantage of those, or are there providers that allow you? I imagine like everything there's like a spectrum of the number of knobs you get when you're doing inference.

Philip

我们在 Base10 的理念,不是要过度炒作 FOMO,是作为最终用户,你应该能接触到每一个能想到的旋钮,当然还有关于哪些该碰、哪些不该碰、哪些是合理默认值的指导。但如果你正在从使用无法访问权重的封闭模型,过渡到使用开放模型或你自己训练的模型,并进行专用部署——基本上每个达到合理规模的公司都会开始需要这样做——不再是按 token 付费,而是为底层 GPU 硬件付费,那么你就能接触到这些旋钮。作为工程师,你是否拥有调整它们的知识和信心,这正是推理工程这门学科存在的原因。但随着各种公司,尤其是垂直 AI 原生公司,达到如此巨大的规模,每个人都在考虑控制他们的 AI,拥有他们的智能,其中一部分意味着拥有他们的推理结果,掌握这些旋钮,为他们的产品和用例找到正确的定位。

Our philosophy at Base10, not to cut up FOMO too much, is that you as the end user should have access to every conceivable knob, and of course along with guidance on which ones are hot to touch and which ones aren't and what some sensible defaults are. But if you are making the transition from working with a closed model where you don't have access to the weights to working with either an open model or a model that you trained yourself, and doing dedicated deployments, which basically everyone at a reasonable scale starts to need to do, where instead of paying on a per token basis, you're paying for the underlying GPU hardware, then you have access to these knobs. Whether or not you as the engineer have the sort of knowledge and confidence to tone them is why this discipline of inference engineering exists. But as companies of all kinds, especially like vertical AI native companies, are achieving the massive scale that they are, everyone is thinking about taking control over their AI, owning their intelligence, and part of that means owning their inference outcomes, taking these knobs and figuring out the right positioning for their product and for their use case.

Host

是的,让我们更深入地谈谈一个人可能追求的掌控层级,这可能是一个严格的序列,也可能不是,但企业中有哪些事情推动了从一个状态到另一个状态的转变?你谈到了从按 token 付费到为 GPU 时间付费。这个频谱延伸到,我想我想说的是,你可以使用像你们这样的提供商,你可以在基础设施提供商上搭建 Kubernetes 集群,你也可以在自己的大楼里托管一切。

Yeah, let's talk a little bit more deeply about the levels of control that one might pursue and it may or may not be a strict sequence, but what are the things that are happening in a business that drive the transition from one state to another? So you talked about paying per token to paying for GPU time. That spectrum goes to, I guess I'm trying to get at like, you could use a provider like you guys, you could stand up a Kubernetes cluster on an infrastructure provider, you could host it all in your own building.

推理部署成熟周期 Inference Deployment Maturity Cycle

Host

比如你看到的主流模型或主流部署方式是什么?是什么驱动一家公司从一种方式过渡到另一种?

Like what are the dominant models or the dominant setups that you're seeing and what drives a company to go from one to the next?

Philip

我看到的路径在很多情况下与产品成熟度周期一致。我想明确一下,这确实是产品成熟度周期,而不是公司成熟度周期,因为经常你会看到相对较小的公司处于周期的更后期,而相对较大的公司反而更早期,这仅仅是因为他们作为公司更“AGI 上头”,或者因为 AI 组件对他们产品的重要性,比 AI 组件对更大、更成熟产品的重要性更高。所以当我们考虑成熟度周期时,有几个选项。选项一是完全依赖按 token 计费的封闭模型提供商。几乎每个人都是从这里开始的。每个人都应该从这里开始。这非常容易。你用一个 API 密钥就能获得前沿智能。当你起步时,这是很难被击败的。我认为下一个阶段通常会出现两个问题之一:要么是成本问题,要么是容量问题。也就是说,要么我的 token 成本失控了,要么我拿不到足够的 token 来做我想做的事。为此,人们常常开始转向超大规模云服务商,比如 AWS、GCP 等,做诸如购买预置吞吐量、在 Bedrock、Vertex 或 Azure AI Foundry 等平台上启动模型之类的事情。这有点把问题往后推,因为现在你与一个拥有大规模和交付能力的提供商签订了大规模承诺。然后我经常看到的下一步是转向更专门的推理提供商,比如 Baseten,公司拥有自己拥有的权重模型,并开始建立专门的部署。我认为那里的另一个岔路口是走向更内部的方案,要么构建内部平台,配备团队,试图在内部提供一流的推理平台;要么在某些情况下,深入边缘推理,特别是如果你想做大量现场工作,你会把所有分布式系统问题替换成物联网问题,但你在扩展大规模分布式边缘网络的推理方面面临非常相似的挑战。我看到越来越少公司出去大量购买 GPU,把它们放在办公室地下室里,然后内部做所有事情。我确实看到这种情况,尤其是在医疗领域、医学研究和医疗保健公司。我也在金融界和其他一些传统企业里见过。我想另一种选择是出去做非常大的支出承诺。比如你看到 Jane Street 现在与 CoreWeave 合作。这有点像在出去买一堆 GPU 放在地下室里和这种金融化方式之间的区别。我认为所有这些方法在适当的情况下都是有效的。我当然不会说世界上所有的推理都需要通过专门的提供商来运行,但我也会说,这些技术,就像我们之前讨论的,要做得很好并跟上行业持续 Scaling 的节奏是非常具有挑战性的。所以,在输入端,越来越多的人倾向于购买而不是构建,这让我非常感激。

The journey I sort of see is in many cases aligned with a product maturity cycle. And I want to specify this is really a product maturity cycle. This is not a company maturity cycle because often times you'll see relatively small companies further along in the cycle versus relatively large companies just because they're either more AGI pilled as a company or just because the AI component of their product is more critical to what they're doing than the AI component of maybe the larger and more established product is. So as we think about the maturity cycle, there's a few options. Option number one is to rely entirely on per token closed model providers. And everyone starts here just about. Everyone should start here. It's really easy. You get frontier intelligence with an API key. And that's a hard thing to beat when you're starting out. I think the next level often looks like running into one of two problems, either a cost problem or a capacity problem. And saying like either my token costs are just getting out of control and or I just can't get enough tokens to do the thing I want to do. And for that often times people start turning to hyperscalers, AWS, GCP, et cetera, and doing things like provisioned throughput purchases, doing things like spinning up models on like a Bedrock or a Vertex or an Azure AI Foundry or something like that. And this kicks the can down the road a little bit because now you have a large scale commit with a provider who has of course massive scale and the ability to deliver you some capacity. And then the sort of next step that I see a lot of times is going on to more of a dedicated inference provider like a Baseten, where companies have the models with the weights that they own and they start to set up a specialized deployment. I think the sort of other fork in the road there is to go for something that's a little bit more in house, either building an in house platform, staffing that team up and trying to deliver a best in class inference platform internally. Or in some cases like going really deep into edge inference, especially if you want to do a lot of field work, you replace all of the distributed systems problems with internet of things problems, but you have a very similar challenge in scaling inference for a large distributed edge network. I see a lot less of companies going out and making enormous capital purchases of large numbers of GPUs, sticking them in the basement of their office, and doing all that stuff internally. I do see that, especially in the medical field and medical research and health care companies do that. I've definitely seen that in the finance world as well and some of those other traditional enterprises. I guess the alternative to that is also like going out and making a very large spend commit. You saw like Jane Street for example now stay with CoreWeave. That's sort of a financialization difference between going out and buying a bunch of GPUs and sticking them in your basement. I think that all of these approaches are valid in the right circumstance. I certainly wouldn't say that all inference in the world needs to be run through a dedicated provider, but I will also say that these technologies, like we talked about earlier, are very challenging to do well and to keep up with as the industry continues to scale. So it's definitely becoming more and more of a trend for folks to buy instead of build on the input side, which is something that I'm really grateful for.

GPU 寿命与推理需求 GPU Lifespan and Inference Demand

Host

是的。随着 AI 的发展,你听到的关于为其他公司做推理的企业的批评之一是,这些 GPU 是快速贬值的资产,而公司正在公开市场上融资购买,并通过债务和其他机制融资。我只是想知道,你对 GPU 寿命与推理、推理成熟度以及该领域发展方式的关系是否有强烈的看法。

Yeah. One of the critiques that you hear about the businesses that are being set up to do inference for other companies as the AI grows is that these GPUs are rapidly depreciating assets and the companies are financing the purchases on the public markets and financing them through debt and all these other mechanisms. I'm just wondering if you have a strong feeling about the GPU lifespan as it relates to inference and inference maturity and the way the field is going.

Philip

所以我已经经历了三个完整的 GPU 周期:Ampere、Hopper、Blackwell。实际上我还没有完全经历 Blackwell,因为 Blackwell 还没有完全推出。问题是,一个 GPU 代际从首次制造到实际分发到每个人手中需要几个月到一年的时间。然后还有另一个几个月到一年的周期,用于将行业中的所有代码移植到其上运行。所以实际上,Hopper GPU 在推理方面仍然非常非常受欢迎。其中一个重要原因是,很多开源工作来自中国实验室,由于出口管制,他们通常使用 Hopper GPU 而不是 Blackwell GPU。所以你通常会得到 FP8 或 int 4 内核,你会得到为 Hopper 的异步编程范式构建的东西,而不是 Blackwell 内核略有不同的范式。你会得到为 8x H200 节点的大小和限制而构建的模型,而不是 GB300 NVL72 系统。所以还有很多东西是为 Blackwell 构建的。抱歉,是为 Hopper。我认为 Ampere 在推理方面基本上已经失宠了,主要是因为它不支持 FP8 量化,所以你不得不以全精度运行模型,否则会遭受灾难性的质量损失。这使得 Hopper GPU 在推理方面比 Ampere 更具成本效益,但由于 Blackwell 短缺以及行业中大量 Hopper 优化的工作负载,Hopper 拥有巨大的持久力。H100 现在按租赁价格比一年前更贵。我认为,随着我们继续从根本上低估推理需求,即使我们能够把模型做得更大、更便宜、更快,对 Hopper 推理的需求也会持续强劲。我认为即使 Rubin 推出,对 Blackwell 推理的需求也会非常非常强劲。而且,Hopper 的另一个优点是它们也非常适合较小的模型。随着我们获得越来越大的 GPU,我们越来越依赖 MIG 多实例 GPU 来拆分它们,并为小模型调整合适的大小。

So I've been through now three full GPU cycles, Ampere, Hopper, Blackwell. And I actually haven't really been all the way through Blackwell because Blackwell isn't fully rolled out yet. The thing is the time it takes for a GPU generation to first be manufactured and actually distributed to the point where everyone can get their hands on it is months to a year. And then there's another cycle of months to a year of porting all of the code in the industry to run on it. So actually Hopper GPUs in particular still are very very popular for inference. One big reason for that actually is because so much open source work comes out of Chinese labs who due to export controls generally work on Hopper GPUs and not Blackwell GPUs. So you generally get like either FP8 or int 4 kernels, you get things that are built for Hopper's asynchronous programming paradigm instead of the slightly different paradigm of Blackwell kernels. You get things that are models that are built for the size and restriction of say an 8x H200 node instead of say a GB300 NVL72 system. So there's a lot that still is built for Blackwell. Sorry, for Hopper. I think that Ampere has finally fallen out of favor in inference for the most part, mostly because it doesn't have support for FP8 quantization, so you have to run models at full precision or suffer sort of catastrophic quality loss. And that makes Hopper GPUs much more cost effective for inference versus Ampere, but due to both Blackwell shortages as well as the large number of Hopper optimized workloads in the industry, Hopper has had a ton of staying power. An H100 is more expensive now than it was a year ago in terms of a rental basis. And I think that as we continue to radically underestimate the demand for inference, even though we're able to scale models bigger, make them cheaper, make them faster, there's going to be a very strong sustained demand for Hopper inference. I think there's going to be a very very strong sustained demand for Blackwell inference even as Rubin rolls out. And yeah, the other thing about Hopper is that they are really good for smaller models, too. As we're getting bigger and bigger GPUs, we're relying more on like MIG multi-instance GPUs to split them up and right-size them for small models.

GPU 寿命与推理 GPU Lifespan and Inference

Philip

比如嵌入模型、语音进语音出模型,这些可以是 1 到 30 亿参数的模型,很多情况下可能大到 70 到 80 亿参数。所以,你不需要那种全新的、机架里装 72 块 Blackwell GPU 的大家伙来跑这些。你只需要一块 Hopper GPU 的一小部分,就能非常高效、经济地完成。所以,回答有点长,但我非常看好 GPU 世代的有效寿命,尤其是从 Hopper 开始。我确实认为,再说一次,老一点的东西其实,我是说 Lovelace,我们还在跑很多 L4 和 L40 的工作负载。很多时候是用于这些较小的模型,或者对成本更敏感的批量工作负载。但是,在那之前,Ampere,你知道,T4 之类的,以前还有 Kepler GPU,我记得当年。我们现在真的看不到对那类东西的需求了。但是,从 Lovelace、Hopper 这些 2018 年的 GPU 开始,也就是 2018 到 2020 年,然后我们 5 到 6 年后还在用它们,所以。它们确实有持久力。

Stuff like embedding models, voice-in voice-out models, like these can be 1 to 3 billion parameter models, maybe as big as like 7 to 8 billion in many cases. And so, you don't need like the massive brand new 72 Blackwell GPUs in a rack to run these things. You need like a small slice of a Hopper GPU to do them in a very efficient and cost-effective way. So, long-winded answer, but I am very bullish on the sort of valuable lifespan of GPU generations, especially starting with Hopper. I do think that again, like the older stuff actually, I mean, Lovelace, like we still run a lot of L4 and L40 workloads. Again, a lot of times for these smaller models or like more cost-sensitive batch workloads. But, yeah, before that, Ampere, you know, the T4s, like all that sort of stuff, you know, used to have like Kepler GPUs, and I remember back in the day. We don't really see demand for that anymore. But, starting with the Lovelace Hopper like 2018 GPUs, so that's like 2018 2020, and then we're still using them like 5 6 years later, so. It has got staying power.

Host

好的。我们稍微聊了一下推理工程作为职业方向,你也提到了 AI 辅助编程的兴起。我很好奇,你觉得 AI,抱歉,推理工程是不是特别不容易被生成式 AI 中的 LLM 完全自动化?比如,它的系统本质是否让它更稳健,或者,LLM 能不能像其他软件工程领域一样,甚至更容易地学会写 CUDA、立方体控制命令之类的各种东西?

Okay. We talked a little bit about inference engineering as kind of a career direction, and you kind of alluded to the rise of AI-assisted coding. I'm curious, do you think that AI, or sorry, inference engineering is like, you know, particularly more or less resistant to being fully automated via LLMs in GenAI? Like, does the system's nature of it make it, you know, more robust or, you know, can LLM figure out how to, you know, write the CUDA and the cube control commands and like all the random stuff just as easily as or maybe even easier than other aspects of software engineering?

Philip

嗯,现在确实有某些技术栈的部分正在被 AI 加速。我觉得一个很好的例子是 CUDA 内核编写,这是一个非常复杂的领域。

Well, there's definitely certain pieces of the stack that are getting AI accelerated right now. I think one great example is CUDA kernel writing, which is a very sophisticated area.

Host

比如有一大堆开源项目和初创公司,基本上就是按需定制 CUDA 内核。

Like there's a whole bunch of, you know, open source and startups that are like doing custom CUDA kernels on demand, you know, essentially.

Philip

完全正确。完全正确。所以,这很重要。问题是,首先,对更专门的推理有大量需求。推理系统相对于模型来说仍然非常通用。你拿一个特定的架构,用一套相当通用的内核、一套相当通用的框架和引擎,在非常通用的硬件上运行。所以,在让推理系统更针对它所运行的具体工作负载,以及动态调整推理系统方面,还有很多解锁空间。所以,我可以想象一个世界,我们使用的推理系统比今天的复杂 10 倍。

Exactly. Exactly. So, that's big. The thing is, there's first off, like there's a lot of demand for more specialized inference. Inference systems are still very generic relative to the models. You're taking a sort of specific architecture and running it with a pretty general set of kernels, a pretty general set of frameworks and engines on top of that on very very general hardware. So, there's a lot of unlocks to be made in terms of making the inference system more specific to the exact workload it's running, as well as adjusting that inference system on the fly. So, I could imagine a world where the inference systems we use are 10 times more complex than the ones we have today.

Host

你在书里提出了一个非常有趣的观点,关于推理解锁,本质上,我在这里转述一下,推理解锁本质上就是为推理的思考方式添加约束。

You made a really interesting point in the book about how inference unlocks, essentially, I'm paraphrasing here, but inference unlocks are essentially adding constraints to the way you think about inference.

Philip

是的。所以,理解工作负载的本质,然后让东西更针对你真正想做的事。这就像不是要能合理地做好所有事情,那是我们行业今天的系统状态。而是要尽可能好地做你确切想做的事。所以,我希望 LLM 辅助推理解锁的第一件事就是更多这样的能力。也就是创建更针对工作负载的系统。其他被 AI 编程辅助的部分,当然包括部署模型和推理的开发体验。比如,如果你想为一个模型创建 spec deck 算法,你需要大量的

Yeah. So, understanding the nature of the workload and, you know, making something more specific to exactly what you want to do. It's like it's not about being able to do everything reasonably well, which is where our systems are today as an industry. It's about being able to do the exact thing you want as well as possible. And so, the first thing that LLM-assisted inference I hope unlocks is more of that. The ability to, you know, create much more workload-specific systems. The other pieces that are getting assisted by AI programming are certainly, you know, the developer experience of deploying models and inferencing them. The creation of, for example, if you want to create a spec deck algorithm for a model, like you need a large amount

Host

Spec deck 算法?

Spec deck algorithm?

Philip

推测解码。

Speculative decoding.

Host

明白了。

Got it.

Philip

是的。如果你想做一个推测算法,比如像 Eagle 3 模型,你需要创建大量的样本,你知道吗?那种合成数据生成可以由 AI 辅助。在这个推理行业里,有很多苦力活可以让 LLM 替我们接管。而且,最终我确实看到它们端到端地拥有某些系统。但是,我们在 Baseten 喜欢说的一句话是,不能靠 vibe 编码来保证正常运行时间。最终,对于这些依赖它们、承载着数亿或数十亿美元经济价值的关键任务系统,仍然需要有人类所有者来为系统的结果负责。而且,当然,我们会继续开发更好的工具来放大这些人的影响力,但他们不会完全消失。

Yeah. If you want a speculation algorithm, like a say like an Eagle 3 model, you need to create a large amount of samples, you know? That sort of synthetic data generation can be AI-assisted. There's a lot of grunt work in this inference industry for LLMs to take over for us. And you know, eventually I do see them owning certain systems end to end. But, one of the things we like to say at Baseten is like can't vibe code uptime. And ultimately, there is still going to need to be for these mission-critical systems that have, you know, hundreds of millions or billions of dollars of economic value like relying on them. There's going to need to be human owners who can be accountable for the results of this system. And you know, of course, like we're going to continue to develop better and better tooling to amplify the impact of these people, but they're not going to go away entirely.

约束与多模态代理 Constraints vs. Multimodal Agents

Host

我们稍微聊了一下添加约束作为获得更高性能或更高效系统的方法。与此同时,我认为我们行业也在向这些能做所有事情的多模态模型、能做所有事情的智能体发展。这些想法是相互矛盾,还是只是指向了推理的某个特定方向?

We talked a little bit about this idea of like adding constraints as a way to get to higher-performing systems or more efficient systems. At the same time, like we're also moving as an industry, I think, to you know, these multimodal models that can do everything, agents that can do everything. Like, are these ideas at tension or, you know, does it just kind of point to a particular direction for inference?

Philip

我认为对智能体和多步推理的需求,实际上正是推动这种非常专门的推理优化的原因。在聊天世界里,每次用户采取行动,你都是向一个模型发出一个请求。在智能体世界里,当用户采取行动时,你是在发出几十个、几百个,甚至几千个请求,而且经常是跨不同模型的。如果你想让你的智能体快速可靠,那么每一个模型以及它们之间的每一个连接都需要快速可靠。所以,这让推理成为一个更关键的挑战。这也意味着,拥有一个能够非常快速地优化单个模型和单个工作负载的系统,也许通过某种 AI 加速的编程范式,意味着你实际上可以更可靠地构建和扩展这些智能体。

I think that the demand for agents and multi-step inference is in fact what is driving the need for this very specialized inference optimization. In a chat world, you are making every time a user takes an action, you are making one request to one model. In an agent world, when a user takes an action, you are making dozens, hundreds, maybe even thousands of requests, often times across different models. And if you want your agent to be fast and reliable, then every single one of these models, as well as every connection between them, needs to be fast and reliable. And so, that makes inference a more critical challenge. And it also means that having a system where you are able to very quickly optimize individual models and very quickly optimize individual workloads, perhaps with a sort of AI-accelerated programming paradigm, means that you can actually, you know, build and scale these agents much more reliably.

Host

谈谈其中的多模态方面和那种异构工作负载方面。我觉得我在这里问了两个不同的问题。一个是,我们有这些多模态模型。从推理的角度来看,它们有什么不同和有趣之处,如果有的话?另一个更想探讨的是在智能体世界里,你提到了这种向不同类型模型发出的请求扇出。我不知道,这个问题突然出现在我脑海里。

Talk about the multimodality aspect of that and the kind of disparate workload aspect of that. Like, I think I'm asking two distinct questions here. One is like, you know, we've got multimodal models. Like, what is different and interesting about them from an inference perspective, if anything? The other is more trying to get at in the agentic world, you kind of spoke to calling, you know, this fan out of like requests to like different types of models. I don't know, this question is kind of occurring to me.

工具工程与推理范围 Tool Engineering and Inference Scope

Host

比如,你也有工具之类的,你有没有想过,或者人们在多大程度上真正关注工具工程或工具推理工程?我不知道这有没有道理。但我很好奇,这可能是关于大规模工具系统优化和工程的一个独立问题,也许超出了推理的范围,但如果没有,我想听听你的想法。

Like, you also have tools like, you know, have you thought about, or to what degree are people really focusing on tool engineering or tool inference engineering? I don't know if this makes sense. But I'm curious, it's probably a separate question of tool-at-scale system optimization and engineering that maybe is out of scope of inference, but if not, I want to hear your thoughts on that.

Philip

我愿意认为大多数事情都在范围内。我们希望尽可能多地为系统的结果承担责任。而工具的提供,以及如何为模型提供工具,这更多是一个上下文工程问题。但这些工具的实际执行,在智能体式的世界中,是推理系统的重要组成部分。

I like to think that most things are in scope. We want to take as much responsibility for the outcomes of the system as possible. And the provisioning of tools and figuring out how to supply tools for the model, that is somewhat more of a context engineering problem. But the actual execution of those tools is an important part of an influencing system in an agentic world.

Philip

我要说的是,模型模态,就像你提到的,是一个值得关注的焦点,因为这里还有一个例子说明理解推理如何帮助你构建更好的产品。以分类任务为例。如果我想对某些数据进行分类,我可以把数据扔给 Opus 4.6,开启最高强度的思考,我确信能得到非常好的分类结果。但我也要为此付很多钱。而如果我想用传统的基于机器学习(ML)的分类器,就像我在 2022 年初玩的那种,如果我有一个轻量级的,也许是我用智能体编码循环写的,我同样能得到非常好的分类结果,而且几乎免费。

I will say that the model modalities, like you were talking about, is an interesting place to focus just because there is another example here of how understanding influence helps you build better products. Let's take classification as an example of a task. If I want to classify some data, I could throw that data into Opus 4.6 with Max thinking high, and I'm pretty sure I would get a really, really good classification out of that. And then I would pay a whole lot of money for it as well. And if I want to use a traditional ML-based classifier like I was playing with in early 2022, if I have the light one, maybe one that I wrote with an agent coding loop, I can also get a really, really good classification. And I can get it for effectively free.

Philip

所以理解不同模型能胜任的不同任务,意味着你当然可以通过转移到更小的模型来节省成本。但你也可以节省时间,这也许更重要。比如,以命名实体识别作为夸张的例子。没有人真正在做命名实体识别,也就是从句子中提取关键词。没有人真正用前沿大语言模型(LLM)来做这件事,至少我希望没有。但即使你用闪存模型之类的来做,而不是专用模型,也许一个高度优化的小型 LLM 可以在 500 毫秒内完成这个任务。我们刚刚发布了一个命名实体识别模型,它能在 1 毫秒内完成。1 毫秒,不是 500。如果你的智能体每个用户请求要做 100 次这样的操作,突然之间,这就从需要盯着加载圈变成了瞬间完成。

So understanding the different tasks that different models are capable of means that of course you can save cost by moving it over to a smaller model. But you can also save time, and that's maybe even more important. So let's say named entity recognition as an over-the-top example. No one is truly doing named entity recognition, which is sort of extracting keywords from sentences. No one is truly using frontier LLMs to do that, or at least I sure hope they're not. But even if you're using a flash model type of thing for that versus a specialized model, maybe a highly optimized small LLM can do that task in 500 milliseconds. We just released a named entity recognition one that does it in 1 millisecond. One. Not 500. And if you have an agent that is doing this 100 times per user request, all of a sudden this has gone from something where you have to look at the spinner to something that happens instantly.

Philip

所以为你想做的每件不同的事情都拥有优化的运行时,起初可能看起来有些过度,但当你看到许多前沿人工智能(AI)公司运营的规模时,实际上并非如此。在推理领域,有大量的工作负载背后有七位或八位数的支出,它们只是运行在这些非常通用的系统上,需要非常专业化的东西。而这种专业化往往看起来像是特定模态的东西。

So having optimized runtimes for every different thing you want to do, it might seem like overkill at first, but when you look at the scale that many of these frontier AI companies are operating at, it actually very much isn't. There are tons of workloads sitting out there with seven, eight figure spend behind them in the influence world that are just running on these very generic systems and need something that's very specialized. And that specialization can often look like a modality-specific thing.

Philip

很多模型都趋向于两种存在方式。一种是自回归式的词元视角,即使是语音进语音出的模型也是如此。嵌入模型也是如此。另一种是扩散视角,更常见于视频生成和图像生成,但也有扩散 LLM,他们总是把扩散偷偷塞进各种不同的东西里。那里有很多有前景的研究。但在每一种方式中,你仍然需要考虑不同的组件。比如视觉语言模型(VLM),视觉语言模型也可以接受图像作为输入。它们在万亿参数的 LLM 前面有一个很小的 10 亿参数的视觉编码器。老实说,那个小小的视觉编码器往往比 LLM 本身造成更多的运行时问题,只是因为现在的视觉编码器在整个行业中非常异构,而且对各种怪癖的软件支持要少得多。

A lot of models are kind of congregating around two ways of being. One being a sort of auto-regressive token look of the world, and even voice-in voice-out models look like that. Embedding models look like that. The other being a diffusion look at the world, which is more often seen in video generation and image generation, but also diffusion LLMs, they're always sneaking diffusion into all kinds of different things. And there's a lot of promising research there. But within each of those, you still have different components to consider. Like with a vision language model, vision language models can take images as input as well. And they have this tiny little 1 billion parameter vision encoder in front of your trillion parameter LLM. And honestly, that little vision encoder often causes a lot more runtime problems than the LLM itself does, just because the vision encoders right now, they're so heterogeneous across the industry, and there's so much less software support for all the different quirks.

Philip

所以这只是说,推理系统对工作负载的专门化仍然处于非常非常早期的阶段,除了拿一个巨大的 LLM 并尽可能快地运行之外,还有很多东西需要构建。尽管那里也还有很多东西要建。

So that's just to say the specialization of the influence systems to the workload is still in very, very early days, and there is a lot more to build across every single thing other than just take a giant LLM and run it as fast as possible. Although there's still a lot to build there, too.

Host

是的,我一边听你说一边想,当你能把所有东西分解成顺序自回归模型或扩散模型时,那么所有问题都被抽象成这两种类型,但如果 VLM 是一个例子,那么这些模态中的每一种都有自己需要解决的怪癖。

Yeah, I thought as you were talking that you were, you know, when you can break everything down to sequential auto-regressive models or diffusion models, then all of the problems get abstracted away into these two types, but if the VLM is any example, then each of these modalities has its own quirks that you need to address.

Philip

确实如此。我想有时候你会得到一些意想不到的附带好处。其中一个例子就是文本转语音(TTS)模型。许多现代 TTS 模型基本上都是经过极度微调的 LLM。你要做的是扩展 LLM 的词汇表,即它能在词元中表示的不同事物,加入几万个音频波形的编码表示,然后微调模型,让它一次生成一个词元,流式输出并解码。对于这类模型,我们能够直接抓取一个定制版的 TensorRT LLM,把 TTS 模型放进去,然后让它跑起来,突然就获得了非常非常出色的性能。

It does. I think one place sometimes you do get sort of unexpected carryover benefits. One place for that actually is text-to-speech models. Many modern text-to-speech models are basically extremely fine-tuned LLMs. What you do is you extend the vocabulary of the LLM, the different things it's able to represent in a token, with an encoded representation of a couple tens of thousands of audio waveforms, and then you fine-tune the model to generate those tokens one at a time and stream them out and decode them. And for those sort of models we were able to just grab a custom version of TensorRT LLMs, stick the TTS model in it, and let it rip, and all of a sudden get really, really excellent performance.

Philip

但如果你尝试对同样的事情,比如我提到的 VLM,或者如果你尝试对音频输入的 Whisper 做完全相同的事情,它确实有编码器和解码器组件,那么就行不通。所以你会时不时得到免费的午餐,但很多时候你必须做这种非常具体的工作。即使你不需要,即使在像 TTS 这样的情况下,你能够重用现有系统,仍然有特定模态的考虑。对于 LLM,你只想生成尽可能多的词元。你希望每秒词元数尽可能高。对于 TTS,你只需要一定数量的每秒词元数来实现实时语音。通常大约是 80 到 100,取决于模型。

But if you try to do the same thing with again, like VLMs as an example I talked about, or if you try and do the exact same thing with Whisper for audio in, which does have the encoder and the decoder component, then it doesn't work. So you do get a free lunch here and there, but a lot of times you do have to do this very specific work. And even when you don't, even in a case like TTS where you are able to sort of reuse an existing system, there's still modality-specific considerations. With LLMs, you just want to make as many tokens as possible. You want your tokens per second to be as high as possible. With text-to-speech, you only need a certain number of tokens per second for real-time speech. Often times that's about 80 to 100, depending on the model.

性能目标与架构 Performance Goals and Architecture

Philip

所以如果你的系统能达到每秒 100 个 token,突然之间你想做的是增大批处理大小,以便运行更多并发流,或者你可能想为重试和检查之类的事情留出空间。或者想办法在更便宜的硬件上实现每秒 100 个 token。你关心的 TPS 速度是有上限的,这是新情况。所以即使在跨模态的类似架构中,你可能会有不同的性能目标,而不同的性能目标会完全改变你设计系统的方式。

So if you get your system to like 100 tokens per second, all of a sudden what you want to do is increase the batch size so that you can run more concurrent streams, or maybe you want to leave room for retries and checks and that kind of stuff. Or figure out how to get that 100 tokens per second now on cheaper hardware. There's a cap on your TPS speed that you care about, and that's novel. So even in a similar architecture across modality, you might have a different performance goal, and having a different performance goal completely changes the way you architect your system.

Host

你提到了 TensorRT,这让我想到软件方面,比如 VLM、TensorRT,你觉得我们处于这个周期的哪个阶段?我们是在爆炸阶段,选项越来越多,还是处于收窄阶段?有时候我觉得两者都有一点。

So you mentioned TensorRT, and that made me think about the software side of this, like VLM, TensorRT, where do you think we are in the cycle? Are we in the explosion part where we're getting a lot more options, or are we at the narrowing part? It kind of feels like a little bit of both to me sometimes.

Philip

这个市场在开源领域非常集中。比如 JavaScript 框架,在鼎盛时期,有几十个非常可行的框架可以用来构建应用。再看编程语言,可能还有大约八种语言,大多数人认为它们是构建网站的合理选择。当然,整个 AI 社区有大量非常有前途的开源项目。我认为在智能体编排和上下文处理等方面有很多非常有趣的工作。在推理领域,确实集中在三大开源运行时上:VLM、SG Lang 和 TensorRT LLM。我认为部分原因是,从零搭建一个新的运行时非常非常复杂。所以大多数人觉得为现有项目做贡献更有用。除此之外,还有很多非常好的开源推理优化工作,比如优秀的开源内核库、开源量化工具、KV 缓存复用工具、新的投机采样之类的东西。但我认为,正是由于从零构建一个完整的推理引擎的复杂性,即使是推理服务提供商也倾向于采用开源引擎,并在其上进行大量定制,而不是完全从头重新发明,除非他们必须这样做,比如因为他们使用了自己设计的专用硬件,或者创建了自己的编程语言等等。行业内的做法是,取开源之精华,结合自己构建之精华,把它们融合在一起。

The market here is remarkably concentrated in open source. If you look at, for example, JavaScript frameworks, at the height, there were dozens of very viable frameworks you could build an application on. If you look at programming languages, there are still maybe eight languages that most people would say are reasonable picks to build a website with. There are of course a large number of really promising open source projects across the entire AI community. I think there's a ton of really interesting work in agent orchestration and context work and that kind of stuff. Within inference, there's really been a concentration around three major open source runtimes: the VLM, SG Lang, and TensorRT LLM. And I think part of that is just the complexity of standing up a new runtime from zero is very, very high. And so most people find it more useful to contribute to an existing one. There's a lot of really good open source work around inference optimization outside of that, like good open source kernel libraries, open source quantization tools, KV cache reuse tools, new speculation stuff, that kind of thing. But I think it's just due to the complexity of building a complete inference engine from scratch that even the inference providers tend to take an open source engine and build a bunch of customization on top of it versus completely reinventing everything from scratch, unless they have to due to either being on specialized hardware that they designed themselves, or creating their own programming language or whatever. The practice in the industry is to take the best of open source and the best of what you can build yourself and put them together.

Host

嗯。你提到了硬件,你觉得硬件方面会发生什么?

Mhm. And you mentioned hardware, what do you see happening on the hardware side of the equation?

Philip

我是说,我完全支持绿色阵营。书的封面是绿色的,这是有原因的。但我认为硬件层的专业化也有其道理。我最近看到的最酷的演示之一是 Talus,那是一个 ASIC,他们把真正的 Llama 3.1 8B 烧录到芯片上,达到了每秒 16,000 个 token 的惊人数字,这非常酷,也是一个有趣的研究项目,指向这个行业可能的发展方向。我认为我们正在看到的是,也许 2026 年会是很多事情的一年。2026 年可能是“解耦之年”,例如 Nvidia 收购 Groq,拥有更专业的预填充算力与解码算力。我们在 AWS 上也看到了类似的情况,他们将训练芯片与 Cerebras 结合用于解码。我认为这些系统确实有可取之处,尤其是对于超大规模工作负载,比如你有一个模型,要发送数百万个并发请求。所以,硬件日益专业化才刚刚开始,我认为它将为行业带来很多解锁,但它肯定不会完全取代对非常复杂软件的需求。

I mean, I'm team green all the way. The book cover is green for a reason. But I think that there's an argument to be made for the specialization at the hardware layer as well. I think one of the coolest demos I saw recently was Talus, which was that ASIC where they burned the actual Llama 3.1 8B onto a chip and got to some ridiculous like 16,000 tokens per second number, which was really cool and sort of an interesting research project in the direction of where this industry could go. I think that what we're seeing, maybe 2026 is the year of a lot of things. One thing 2026 could be is the year of disaggregation with, for example, Nvidia buying Groq and having more specialized pre-fill compute versus decode compute. We saw the same thing on AWS with combining their training chips with Cerebras for decode. I think that there's definitely something to be said for these systems, especially for very large-scale workloads where you have one model that you are sending millions of concurrent requests to. And so, yeah, the increasing specialization in hardware is really just beginning and I think it's going to offer a lot of unlocks for the industry, but it's certainly not going to entirely replace the need for very sophisticated software as well.

Host

你开始说 2026 年是……之年,你还看到了什么?你对什么感到兴奋?你的水晶球怎么说?

You started going down this path of 2026 is the year of like, what else are you seeing? What are you excited about? What's your crystal ball saying?

Philip

我是说,2026 年是推理之年。就像 2025 年也是,2024 年也是,2027 年也会是。但是,嗯,这是推理之年。还有 Linux 登上桌面。

I mean, 2026 is the year of inference. Like, 2025 was too, 2024 was, 2027 will be. But, well, it's the year of inference. And Linux on the desktop.

Host

>> >>

>> >>

Philip

我认为今年我们会看到智能所有权真正增加。你看到像 Shopify 转向量化模型,在工作负载上节省了数百万、数千万。你看到像 Kosa 这样的公司推出了像 Caposa 这样非常复杂的模型,使他们能够真正为依赖其平台的用户创造新颖的体验。比如,我就是日常用户。所以,我最兴奋的趋势是,公司开始明白,如果他们想构建真正差异化的产品,他们需要在每个层面都实现差异化,而这开始包括模型层面。

I think that this year we're going to see a real increase in ownership of intelligence. You've seen like with Shopify them moving to a quant model and saving millions, tens of millions on their workloads. You see companies like Kosa coming out with very sophisticated models like Caposa that are allowing them to really create a novel experience for the users who are depending on their platform. I am a daily user, for example. So, I think that's the trend that I'm most excited about is companies understanding that if they want to build a really differentiated product, they need to be differentiated at every level and that's starting to include the model level.

Host

对于想深入了解的人,他们能在哪里买到这本书?容易获取吗?

And for folks that want to dig in more, where can they get the book? Is it easily accessible?

Philip

是的,Inference Engineering 这本书可以在 baseten.com/inferenceengineering 获取。有免费的 PDF 和免费的电子书 ePub。我正在和我们的一个客户合作制作有声书,这非常令人兴奋。所以希望几周内能推出。实体书稍微难获取一些。我严重低估了这本书的需求。我已经发放了 20,000 份数字版,并且几乎送出了我第一版印刷的全部。下周我应该会开一个 Shopify 商店,让人们可以自助购买,但在任何会议或贸易展的 Baseten 展位上总有书。所以,来找我吧。

Yeah, so Inference Engineering is available at baseten.com/inferenceengineering. There's a free PDF and free ebook ePub. I'm working on an audiobook with actually with a customer of ours, which is very exciting. So that'll be out hopefully in a few weeks. Physical copies are a little harder to come by. I massively underestimated the demand for this book. I've already done 20,000 digital copies and given away nearly my entire first print run. I should have a Shopify store up next week, which will allow people to self-serve stuff, but there's always copies at any Baseten booth at a conference or trade show. So, come find me.

Host

你要去?

You're going to?

Philip

下周我要去 AIE Miami 和 GTC。今年夏天我们会参加 AI Engineer World's Fair。我们在旧金山和纽约周边举办很多活动,在那里分发这些书。所以,是的,要拿到一本有点像寻宝游戏。但推理已经到来,是时候了解它了。所以,过来吧。

I'm going to AIE Miami and GTC next week. We're going to be at AI Engineer World's Fair this summer. We do a bunch of events around San Francisco and New York where we hand these out. So, yeah, it's a little bit of a scavenger hunt to get your hands on one. But inference is here and it's time to learn about it. So, come through.

Host

太棒了。太棒了。好的,Phil,非常感谢你参加节目,和我们聊推理工程。

Awesome. Awesome. Well, Phil, thanks so much for jumping on and talking with us about inference engineering.

Philip

嘿,Sam,非常感谢你邀请我,很感激这次对话。同样,谢谢你。

Hey, Sam, thanks so much for having me and appreciate the conversation. Same. Thank you.

互动版:逐字朗读 + 针对本期提问 →