一位教授兼 OpenAI 董事会成员探讨我们是否真的耗尽了用于训练 AI 的互联网数据,认为仅使用了极小一部分。
A professor and OpenAI board member discusses whether we've truly exhausted internet data for training AI, arguing that only a tiny fraction has been used.
要点 · TL;DR
下一个词预测是产生智能系统的重大科学发现。 Next-token prediction is a major scientific discovery yielding intelligent systems.
数据短缺并非迫在眉睫;存在大量未开发的多模态数据。 Data shortage is not imminent; vast untapped multimodal data exists.
AI 安全的最大问题是模型无法可靠地遵循规范。 AI safety's biggest issue is models failing to reliably follow specifications.
核心观点 · Key points
下一个词预测能产生智能系统,这是一项重大科学发现。 Next-token prediction yields intelligent systems, a major scientific discovery.
数据短缺并非迫在眉睫;我们还有大量未利用的多模态数据。 Data shortage is not imminent; we have vast untapped multimodal data.
即使在固定数据集上,更大的模型仍能提升性能。 Larger models still improve performance even on fixed datasets.
AI 安全当前最大的问题是模型无法可靠地遵循规范。 AI safety's biggest current issue is models failing to reliably follow specifications.
开放权重模型对研究至关重要,但在高能力水平时可能需要限制。 Open-weight models are vital for research but may need restrictions at high capability levels.
反共识 · Contrarian takes
我们处于后架构阶段;模型架构并不重要。 We are in a post-architecture phase; model architecture matters little.
数据整理被高估了;原始互联网数据足以用于训练。 Data curation is overrated; raw internet data suffices for training.
虚假信息的真正危害是什么都不信,而非相信谎言。 Misinformation's real harm is disbelief in anything, not belief in falsehoods.
AGI 可能在 4 到 50 年内到来,比学术界传统预期快得多。 AGI will likely arrive within 4-50 years, much sooner than academia traditionally expects.
将 AI 类比核武器并不恰当;AI 有许多有益用途。 Nuclear weapon analogy for AI is poor; AI has many beneficial uses.
越狱是根本性安全缺陷,类似于所有模型中的缓冲区溢出。 Jailbreaking is a fundamental security flaw akin to a buffer overflow in all models.
本期章节 · Chapters(共 29)
引言与背景Introduction and Context
当前 AI 系统基础技术Basic Techniques Behind Current AI Systems
数据稀缺与合成数据Data Scarcity and Synthetic Data
多模态数据挑战Challenges of Multimodal Data
模型性能无平台期No Plateau in Model Performance
优化数据以提取价值Optimizing Data for Value Extraction
多小模型 vs 大模型Many Smaller Models vs. Large Models
模型收益更难获取Harder to See Gains in Models
模型平台期感知Perception of Model Plateauing
模型商品化与格局Model Commoditization and Landscape
算力扩展与收益递减Compute Scaling and Diminishing Returns
AGI 定义与企业目标AGI Definition and Corporate Goals
AGI 时间线与影响Timeline and impact of AGI
虚假信息与深度伪造Misinformation and Deepfakes
AI 监管Regulation of AI
安全关注层级Hierarchy of safety concerns
AI 危险能力与下游影响AI's dangerous capabilities and downstream effects
开源 vs 闭源模型发布Open vs. Closed Source Model Release
实用 AI 安全 vs 遥远场景Practical AI Safety vs. Far-Fetched Scenarios
关联故障与基础设施风险Correlated Failures and Infrastructure Risks
紧迫 AI 安全问题Pressing AI safety concerns
对 AI 未来乐观Optimism about AI future
快问:模型观点转变Quick fire: Changed mind on models
快问:数据观点转变Quick fire: Changed mind on data
快问:加入 OpenAI 董事会Quick fire: Joining OpenAI board
快问:董事会角色与职责Quick fire: Board roles and responsibilities
快问:中美 AI 进展Quick fire: China vs US AI progression
快问:最常被问的不想答的问题Quick fire: Most common unwanted question