LLM 之后:空间智能与世界模型
After LLMs: spatial intelligence and world models
李飞飞 Fei-Fei Li · Latent Space · 2025-11-25 · 约 61 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
为何仅有语言还不够,以及 World Labs 在造什么。
Why language alone isn’t enough, and what World Labs is building.
要点 · TL;DR
- 空间智能和世界模型是超越大语言模型的下一个前沿,能够实现交互式 3D 理解。
Spatial intelligence and world models are the next frontier beyond LLMs, enabling interactive 3D understanding. - Marble 能从文本或图像生成交互式 3D 世界,并支持精确的相机控制。
Marble generates interactive 3D worlds from text or images, with precise camera control. - Transformer 原生建模集合而非序列;位置编码注入顺序信息。
Transformers natively model sets, not sequences; positional embeddings inject order.
核心观点 · Key points
- 空间智能与语言智能互补,而非替代。
Spatial intelligence is complementary to language intelligence, not a replacement. - 世界模型需要理解 3D 结构和物理规律,而不仅是模式拟合。
World models require understanding 3D structure and physics, not just pattern fitting. - Marble 从文本或图像生成交互式 3D 世界,支持精确相机控制。
Marble generates interactive 3D worlds from text or images, enabling precise camera control. - 学术界应专注于疯狂想法和长期研究,而非与工业界比拼规模。
Academia should focus on wacky ideas and long-range research, not competing with industry on scale. - Transformer 原生建模集合而非序列;位置编码注入顺序。
Transformers natively model sets, not sequences; positional embeddings inject order. - 未来世界模型可能结合学习到的物理规律与经典模拟来处理动态。
Future world models may combine learned physics with classical simulation for dynamics.
反共识 · Contrarian takes
- Transformer 本质上是集合模型而非序列模型;位置嵌入才引入了顺序。
Transformers are natively models of sets, not sequences; positional embeddings inject order. - 语言是空间体验的有损通道;空间智能捕捉具身交互。
Language is a lossy channel for spatial experience; spatial intelligence captures embodied interaction. - 当前深度学习拟合模式,但缺乏对重力等物理的因果理解。
Current deep learning fits patterns but lacks causal understanding of physics like gravity. - 高斯泼溅作为原子单元支持实时渲染和精确相机控制,不同于基于帧的模型。
Gaussian splats as atomic units enable real-time rendering and precise camera control, unlike frame-based models. - 视觉因对人类而言毫不费力而被低估,但进化优化了 5.4 亿年。
Vision is underappreciated because it's effortless for humans, yet evolution optimized it for 540 million years. - LLM 可能准确预测行星轨道,但不会自发推导出 F=ma 这样的牛顿定律。
LLMs may predict planetary orbits accurately but won't spontaneously derive Newton's laws like F=ma.
本期章节 · Chapters(共 19)
- 0. 引言与背景 Introduction and Background
- 1. 世界模型的 AlexNet 等价物 AlexNet Equivalent for World Models
- 2. Marble 作为世界模型 Marble as a World Model
- 3. 生态系统多样性与开放 vs 封闭模型 Ecosystem diversity and open vs closed models
- 4. 硬件扩展限制与替代基元 Hardware scaling limits and alternative primitives
- 5. 图像字幕起源故事 Image captioning origin story
- 6. 实时图像字幕演示 Real-time image captioning demo
- 7. 视觉 vs 语言建模 Vision vs language modeling
- 8. 深度学习 vs 人类智能 Deep learning vs. human intelligence
- 9. 单模型 vs 多模型处理不同任务 One model vs. multiple models for different tasks
- 10. 使用物理引擎 vs 从数据学习 Using physics engines vs. learning from data
- 11. Marble 与空间智能简介 Introduction to Marble and Spatial Intelligence
- 12. 为高斯泼溅添加物理 Adding Physics to Gaussian Splats
- 13. 新兴用例 Emergent Use Cases
- 14. 定义空间智能 Defining Spatial Intelligence
- 15. 空间与语言智能的相互作用 Interplay of Spatial and Linguistic Intelligence
- 16. 空间智能 vs 语言模型 Spatial Intelligence vs Language Models
- 17. 高效学习与心智理论 Efficient Learning and Theory of Mind
- 18. 行动号召与人才招聘 Call to Action and Talent Recruitment
阅读全文双语转录 →