推理领域从研究到生产的速度极快,常以小时计,源于激烈竞争。 Research-to-production in inference is rapid, often hours, due to intense competition.
核心观点 · Key points
推理是AI中最重要且最具粘性的工作负载,是进入AI基础设施市场的切入点。 Inference is the most important and stickiest workload in AI, serving as a wedge into the AI infrastructure market.
推理工程需要从GPU编程到分布式系统等多方面的专业知识,是一门复杂的学科。 Inference engineering requires a wide range of expertise, from GPU programming to distributed systems, making it a complex discipline.
推理领域从研究到生产的时间极短,通常只需数小时,因为竞争激烈且需要立即支持新技术。 The research-to-production timeline in inference is extremely fast, often hours, due to high competition and the need for immediate support of new techniques.
公司应将推理视为一系列权衡的谱系,从而优化延迟、吞吐量和成本。 Companies should understand inference as a spectrum of trade-offs, enabling them to optimize for latency, throughput, and cost.
即使有AI辅助编程,未来几年对推理工程师的需求将增长10到100倍。 The demand for inference engineers will grow 10 to 100 times in the coming years, even with AI-assisted coding.
针对特定工作负载定制的专用推理系统,对于构建快速可靠的智能体应用至关重要。 Specialized inference systems, tailored to specific workloads, are crucial for building fast and reliable agentic applications.
反共识 · Contrarian takes
Hopper GPU在推理中仍然非常受欢迎,租金价格上涨,与快速贬值的预期相反。 Hopper GPUs remain highly popular for inference, and their rental prices have increased, contrary to expectations of rapid depreciation.
量化模型并非天生不好;它们可以针对特定工作负载进行校准,带来显著的成本和速度优势。 Quantized models are not inherently bad; they can be calibrated to specific workloads, offering significant cost and speed benefits.
由于复杂性,开源推理运行时市场高度集中,只有vLLM和TensorRT-LLM等少数主要选择。 The open-source inference runtime market is highly concentrated, with only a few major options like vLLM and TensorRT-LLM, due to complexity.
推理系统相对于模型仍然通用;使其更加针对特定工作负载有巨大潜力,可能复杂10倍。 Inference systems are still generic relative to models; there is huge potential for making them more workload-specific, potentially 10 times more complex.
即使有AI辅助编程,关键任务推理系统仍需人类负责;“无法用vibe编码保证正常运行时间”。 Even with AI-assisted coding, human accountability is essential for mission-critical inference systems; 'can't vibe code uptime'.
GPU生命周期比预期更长,像Lovelace和Hopper这样的旧世代仍用于较小模型和成本敏感的工作负载。 The GPU lifecycle is longer than expected, with older generations like Lovelace and Hopper still in use for smaller models and cost-sensitive workloads.
本期章节 · Chapters(共 16)
开场与签书会Introduction and Book Signing
推理公司为何存活Why Inference Companies Survived
推理与模型服务对比Inference vs. Model Serving
推理中的挑战Challenges in Inference
推理如综合格斗Inference as Mixed Martial Arts
推理从研究到生产Rapid Research to Production in Inference
Polo Quant 背景故事The Polo Quant Backstory
谁该关注推理工程Who Should Care About Inference Engineering