Oriol Vinyals 谈世界模型、多模态 AI 及通往 AGI 之路
Oriol Vinyals on World Models, Multimodal AI, and the Path to AGI
奥里奥尔·维尼亚尔斯 Oriol Vinyals · Unsupervised Learning · 2026-05-22 · 约 60 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
Google DeepMind 的 Oriol Vinyals 探讨世界模型、多模态 AI 进展以及 AI 推理和记忆的未来。
Google DeepMind's Oriol Vinyals discusses world models, multimodal AI advances, and the future of reasoning and memory in AI.
要点 · TL;DR
- 结合语言、视觉和视频的世界模型是通往 AGI 的关键。
World models combining language, vision, and video are key to AGI. - 后训练和强化学习在编码和数学之外还有巨大未开发潜力。
Post-training and RL have huge untapped potential beyond coding and math. - 智能体记忆将采用非参数化文件系统存储,而非权重更新。
Agent memory will use non-parametric file-system storage, not weight updates.
核心观点 · Key points
- 联合建模语言、视觉和视频的世界模型是通往 AGI 的核心路径。
World models that jointly model language, vision, and video are a core path to AGI. - 后训练和强化学习仍是绿野,在编码和数学之外有巨大潜力。
Post-training and RL are still a green field with huge potential beyond coding and math. - 智能体的记忆将通过非参数化的文件系统式存储解决,而非权重更新。
Memory for agents will be solved via non-parametric file-system style storage, not weight updates. - 缩放定律仍然成立,但持续学习和元学习是关键缺失的能力。
Scaling laws still hold, but continual learning and meta-learning are key missing capabilities. - 苦涩的教训表明通用系统最终将胜过专门的脚手架。
The bitter lesson suggests general systems will eventually outperform specialized scaffolding.
反共识 · Contrarian takes
- 视频和图像的 GPT 时刻尚未到来;从视觉的纯迁移仍未解决。
The GPT moment for video and images has not yet arrived; pure transfer from vision is still unsolved. - 在数学和编码上的窄域强化学习出人意料地泛化到其他领域,减少了对广泛训练的需求。
Narrow RL on math and coding generalizes surprisingly well to other domains, reducing need for broad training. - 根据某些定义,AGI 可能已经到来;目标不断后移。
AGI may already be here by some definitions; the goalposts keep moving. - 向 Anthropic 出售算力是战略性选择以再投资收入,并非怀疑的信号。
Selling compute to Anthropic is a strategic choice to reinvest revenue, not a sign of doubt. - 科学创新的能力难以评估,可能是最难实现的能力。
The ability to innovate in science is hard to evaluate and may be the hardest capability to achieve.
本期章节 · Chapters(共 22)
- 引言与世界模型 Introduction and World Models
- 世界模型与表征学习 World Models and Representation Learning
- 世界模型与间接评估 World Models and Indirect Evaluation
- 智能体能力与系统设计 Agent Capabilities and System Design
- 苦涩教训与扩展 Bitter Lesson and Scaling
- 推理模型与智能体可靠性 Reasoning models and agentic reliability
- 智能体中的记忆 Memory in agents
- 个性化与共享知识 Personalized vs shared knowledge
- 持续学习与扩展之争 Continual Learning and Scaling Debate
- 组织聚焦与广度 Organizational Focus vs. Breadth
- 后训练与 RL 领域 Post-Training and RL Domains
- 后训练:蓝海与数据限制 Post-training as green field and data limitations
- 元能力与游戏测试 Meta capabilities and testing with games
- RL 泛化与苦涩教训 Generalization from RL and the bitter lesson
- 推理模型与泛化 Reasoning Models and Generalization
- 模型层与应用层 Model Layer vs Application Layer
- 不确定的研究路径 Uncertain Research Paths
- 研究路径与创新 Research paths and innovation
- AGI 已至? AGI is here?
- 战略计算分配 Strategic compute allocation
- 自有芯片的独特优势 Unique position with own chip
- 结语 Closing remarks
阅读全文双语转录 →