Jeff Dean 强调数学和编程 AI 模型的突破以及自主智能体的兴起,Bill Dally 则解释英伟达实现超低延迟推理的策略。
Jeff Dean highlights breakthroughs in math and coding AI models and the rise of autonomous agents, while Bill Dally explains NVIDIA's strategies for achieving ultra-low latency inference.
要点 · TL;DR
超低延迟推理(每秒 1 万以上 token)是智能体 AI 系统的关键。 Ultra-low latency inference at 10k+ tokens/s is key for agentic AI systems.
AI 正在变革芯片设计,从标准单元生成到缺陷归因。 AI is transforming chip design from cell generation to bug attribution.
个性化 AI 导师和健康教练是最具影响力的应用之一。 Personalized AI tutors and health coaches are among the most impactful applications.
核心观点 · Key points
模型在数学和编程等可验证奖励问题上显著提升,在 IMO 和 ICPC 中获得了金牌。 Models have dramatically improved on verifiable reward problems like math and coding, achieving gold medals in IMO and ICPC.
基于智能体的工作流现在可以自主处理持续数小时或数天的长期任务,并在此过程中自我纠正。 Agent-based workflows now handle long-running tasks autonomously for hours or days, self-correcting along the way.
超低延迟推理对智能体系统至关重要;目标是每个用户每秒处理超过 10,000 个词元。 Ultra-low latency inference is critical for agentic systems; targeting 10,000+ tokens per second per user.
AI 正在改变芯片设计:从标准单元库生成到缺陷归因和设计探索。 AI is transforming chip design: from standard cell library generation to bug attribution and design exploration.
个性化 AI 导师和健康教练是对社会最有影响力的应用之一。 Personalized AI tutors and health coaches are among the most impactful applications for society.
反共识 · Contrarian takes
我们并没有耗尽数据;未开发的来源包括视频、音频、机器人技术和合成数据。 We are not running out of data; untapped sources include video, audio, robotics, and synthetic data.
预训练应穿插在环境中采取行动,而不仅仅是被动的下一个词预测。 Pre-training should interleave action-taking in environments, not just passive next-token prediction.
推理现在占数据中心功耗的 90%;硬件必须针对预填充和解码阶段进行专门化。 Inference is now 90% of data center power; hardware must be specialized for prefill vs. decode stages.
能效的关键是不移动数据;存内计算和堆叠 DRAM 很有前景。 The key to energy efficiency is to not move data; compute-in-memory and stacked DRAM are promising.
除了结构化稀疏性外,稀疏性很难利用,因为它破坏了规律性和效率。 Sparsity is hard to exploit beyond structured sparsity because it destroys regularity and efficiency.
网络拓扑的选择取决于工作负载;没有一种网络对所有流量模式都是最优的。 Network topology choice depends on workload; no single network is optimal for all traffic patterns.
本期章节 · Chapters(共 21)
开场与最激动人心的进展Opening and Most Exciting Developments
智能体超低延迟推理Ultra-Low Latency Inference for Agents
大模型低延迟的重要性Importance of Low Latency for Large Models
自我改进的智能体系统Agentic systems for self-improvement
快速演进领域中的硬件需求预测Predicting hardware needs in a fast-moving field
超越 Chinchilla 与数据可用性Scaling beyond Chinchilla and data availability
合成数据与数据增强Synthetic Data and Data Augmentation
训练与推理硬件对比Training vs Inference Hardware
模型架构趋势与注意力改进Model architecture trends and attention improvements
AI 芯片设计:从奇特设计到 LLMAI in chip design
智能体编排与硬件加速挑战AI in chip design: from bizarre designs to LLMs
计算能效Challenges in agent orchestration and hardware acceleration
软硬件协同设计与痛点Energy Efficiency in Computing
软硬件协同设计Hardware-Software Co-Design and Friction Points
持续学习硬件Hardware-Software Co-Design
低延迟推理的硬件变革Hardware for Continual Learning
网络设计权衡Hardware Changes for Low-Latency Inference
AI 对人类的积极影响Network design trade-offs
AI 在教育及其他领域的应用Positive impact of AI on humanity
结束语与社区AI in Education and Other Applications
Closing Remarks and CommunityClosing Remarks and Community