李飞飞和 Justin Johnson 探讨深度学习作为计算扩展的历史,从斯坦福到创立 World Labs 的历程,以及空间智能和世界模型的愿景。
Fei-Fei Li and Justin Johnson discuss the history of deep learning as scaling compute, their journey from Stanford to founding World Labs, and the vision for spatial intelligence and world models.
要点 · TL;DR
空间智能是下一个前沿,Marble 作为生成式 3D 世界模型代表了这一方向。 Spatial intelligence is the next frontier, with Marble as a generative 3D world model.
算力扩展推动了深度学习进步,但未来硬件可能需要全新架构。 Scaling compute has driven deep learning progress, but future hardware may need new architectures.
视觉被低估;世界模型必须理解 3D 空间、物理和交互性。 Vision is underappreciated; world models must understand 3D space, physics, and interactivity.
核心观点 · Key points
空间智能是下一个前沿,与语言智能互补。 Spatial intelligence is the next frontier, complementary to language intelligence.
Marble 是一个 3D 世界生成模型,接受文本或图像作为输入。 Marble is a generative model of 3D worlds, accepting text or images as input.
深度学习的历史就是 Scaling(规模扩张)算力的历史。 The history of deep learning is the history of scaling up compute.
学术界应专注于疯狂的想法和新算法,而不仅仅是 Scaling(规模扩张)模型。 Academia should focus on wacky ideas and new algorithms, not just scaling models.
世界模型需要理解 3D 空间、物理和交互性,才能实现真正的空间智能。 World models need to understand 3D space, physics, and interactivity for true spatial intelligence.
反共识 · Contrarian takes
视觉被低估了;大自然花了 5.4 亿年优化视觉,而语言仅 50 万年。 Vision is underappreciated; it took nature 540 million years to optimize, vs. 500,000 years for language.
Transformer 本质上是集合模型而非序列模型;位置编码注入顺序。 Transformers are natively models of sets, not sequences; positional embeddings inject order.
仅靠数据,LLM 可能永远推导不出牛顿定律;它们拟合模式而非因果理论。 LLMs may never derive Newton's laws from data alone; they fit patterns, not causal theories.
对于世界建模,像素比分词文本更无损。 Pixels are a more lossless representation than tokenized text for modeling the world.
GPU Scaling(规模扩张)正触及极限;未来硬件可能需要截然不同的神经网络架构。 GPU scaling is hitting limits; future hardware may require radically different neural architectures.
本期章节 · Chapters(共 20)
计算规模扩展史History of Scaling Compute
大理石作为世界模型Marble as a World Model
引言与背景Introduction and Background
世界模型的 AlexNet 等效AlexNet Equivalent for World Models
开放科学 vs 集中实验室Open Science vs Centralized Labs
开放 vs 封闭模型与商业压力Open vs. Closed Models and Commercial Pressure
硬件扩展与未来架构Hardware scaling and future architectures
图像描述突破Image captioning breakthrough
实时图像描述演示Real-time image captioning demo
视觉 vs 语言建模Vision vs language modeling
世界模型与潜在物理World models and latent physics
深度学习 vs 人类理解Deep learning vs human understanding
大理石与空间智能导论Introduction to Marble and Spatial Intelligence
3D 场景中的物理与动力学Physics and Dynamics in 3D Scenes
空间智能的新兴用例Emergent Use Cases for Spatial Intelligence
定义空间智能 vs 传统智能Defining Spatial Intelligence vs. Traditional Intelligence
空间与语言智能的相互作用Interplay of Spatial and Linguistic Intelligence