OpenAI 让超级智能变安全的努力
OpenAI’s push to make superintelligence safe
扬·莱克 Jan Leike · 80,000 小时 · 2023-08-22 · 约 176 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
为「超级对齐」与「可扩展监督」辩护。
The case for superalignment and scalable oversight.
要点 · TL;DR
- RLHF 无法用于超人类 AI;可扩展监督利用 AI 辅助人类评估。
RLHF fails for superhuman AI; scalable oversight uses AI to help humans evaluate. - 可解释性是检测超级智能模型欺骗的关键。
Interpretability is key to detect deception in superintelligent models. - 对齐进展现在可以实证测量,而不仅仅是理论推演。
Alignment progress can be measured empirically now, not just theorized.
核心观点 · Key points
- 基于人类反馈的强化学习(RLHF)无法扩展,因为人类无法评估超人类 AI 的输出。
RLHF won't scale because humans can't evaluate superhuman AI outputs. - 我们需要可扩展监督:用 AI 帮助人类评估更难的任务。
We need scalable oversight: use AI to help humans evaluate harder tasks. - 目标是使下一代模型对齐,而非直接对齐超级智能。
The goal is to align the next-generation model, not superintelligence directly. - 可解释性是检测模型欺骗的关键验证技术。
Interpretability is a key validation technique to detect deception in models. - 自动化对齐研究可以迭代地推动对齐进展。
Automated alignment research could bootstrap alignment progress iteratively. - 对齐方法的实证测试至关重要;我们现在就可以衡量进展。
Empirical testing of alignment methods is crucial; we can measure progress now.
反共识 · Contrarian takes
- 对齐可能比许多人想象的要容易,因为评估本质上比生成更容易。
Alignment may be easier than many think because evaluation is inherently easier than generation. - 我们应该故意训练欺骗性模型作为「模式生物」,以研究并防御它们。
We should deliberately train deceptive models as 'model organisms' to study and defend against them. - 自动化可解释性可以扩展到检查每个神经元,而手动分析则无法做到。
Automated interpretability can scale to inspect every neuron, unlike manual analysis. - 四年时间表虽然雄心勃勃,但选择它是为了推动快速进展,而不是因为预计那时会出现 AGI。
The four-year timeline is ambitious but chosen to force rapid progress, not because AGI is expected then. - 更强大的模型平均而言可能更对齐,但最坏情况下的行为可能会恶化。
More capable models can be more aligned on average, but worst-case behavior may worsen. - 最大的风险不是对齐失败本身,而是对系统实际对齐程度的不确定性。
The biggest risk is not misalignment itself but uncertainty about how aligned a system truly is.
本期章节 · Chapters(共 73)
- 引言与背景 Introduction and Background
- 超级对齐项目公告 Superalignment Project Announcement
- 当前对齐方法的局限 Limitations of Current Alignment Methods
- 与现有安全工作的区别 Distinction from Existing Safety Work
- 对齐超级智能 AI 的挑战 Challenge of aligning superintelligent AI
- 对齐的训练与验证方法 Training vs validation methods for alignment
- 对齐的计算资源 Compute for Alignment
- 可扩展监督的基本思路 The Basic Idea of Scalable Oversight
- 所需:可扩展监督与理解泛化 What We Need: Scalable Oversight and Understanding Generalization
- 可解释性作为关键拼图 Interpretability as a Key Puzzle Piece
- 鲁棒性与泛化失败 Robustness and Generalization Failures
- 对不完美反馈下模型泛化的不同直觉 Different intuitions about model generalization from imperfect feedback
- 可扩展监督方法:漏洞检测与判别性批评 Scalable oversight approaches: bug detection and discriminative critique
- AI 研究能力模型与对齐简介 Introduction to AI research capable models and alignment
- ML 社区反应与 RLHF 历史 Reaction from the ML community and RLHF history
- RLHF 的起源与影响 Origin of RLHF and its impact
- 对齐的怀疑与重要性 Skepticism and importance of alignment
- 对齐中泛化的解释 Explanation of generalization in alignment
- 泛化当前研究 Current research on generalization
- 对抗测试与模型生物 Adversarial Testing and Model Organisms
- 对齐的可解释性 Interpretability for Alignment
- 自动化可解释性概述 Automated Interpretability Overview
- 可解释性作为验证技术 Interpretability as a validation technique
- 可扩展监督:批评与评估 Scalable oversight: critiques and evaluation
- 对齐的多元观点 Diverse opinions on alignment
- 近期 AI 发展的乐观 Optimism from recent AI developments
- 评估比生成容易 Evaluating is easier than generating
- 乐观与经验方法 Optimism and Empirical Approach
- 反对:如何判断成功 Objection: How to Tell If Succeeding
- 失败模式与测量 Failure Modes and Measurement
- 关键尝试前验证 Verification Before Critical Try
- 对齐未解决的预案 Plan if Alignment Not Solved
- 对齐进展的诚实透明 Honesty and transparency about alignment progress
- 方法的最佳反对意见 Best Objections to the Approach
- 第二优选方案 Second Favorite Option
- 让你夜不能寐的事 What Keeps You Up at Night
- 对齐核心挑战:找到正确指标 Core challenge of alignment: finding the right metric
- 经济活动与系统结果 Economic Activity and System Outcomes
- 为何不先用 GPT-4 解决对齐 Why Not Solve Alignment with GPT-4 First
- 更强能力是否意味着更对齐 Does More Capable Imply More Aligned
- 主流 ML 视角的智力兴奋 Intellectual Excitement from Mainstream ML Perspective
- 技术 AI 安全的最大胜利 Biggest wins in technical AI safety
- AGI 开发的民主输入 Democratic input for AGI development
- 将 LLM 融入经济的担忧 Concerns about integrating LLMs into the economy
- 对 GPT-4 联网的分歧 Disagreement over GPT-4 internet connection
- GPT-4 的过早使用 Premature use of GPT-4
- 遏制 vs 部署文化 Containment vs deployment culture
- 实验室与部署中的风险 Risks in lab vs deployed
- 发布 ChatGPT 对灭绝风险的影响 Impact of releasing ChatGPT on extinction risk
- ChatGPT 发布时机与 AI 意识必然性 Timing of ChatGPT release and inevitability of AI awareness
- 阻止 AI 进步的难度 Difficulty of stopping AI progress
- 放缓 vs 加速对齐研究 Slowing down vs speeding up alignment research
- 小领域规模的乐观 Optimism from small field size
- 超级对齐团队不回避对齐 Superalignment team not avoiding alignment
- 商业化与对齐激励 Commercialization and alignment incentives
- 快速起飞场景的规划 Planning for fast takeoff scenarios
- 反对对齐可解性的论点 Arguments against solvability of alignment
- 对 AI 操纵的偏执 Paranoia about AI manipulation
- 扩展中的风险与安全 Risk and Safety in Scaling
- 定义成功与科学共识 Defining Success and Scientific Consensus
- 政策、治理与行动号召 Policy, Governance, and Call to Action
- 招聘计划与角色 Hiring Plans and Roles
- 期望背景与技能 Desired Background and Skills
- 经理角色要求 Manager Role Requirements
- 计算与工程挑战 Compute and Engineering Challenges
- 申请流程 Application Process
- 超级对齐团队推介 Pitch for Superalignment Team
- OpenAI 研究重要问题的团队 Teams at OpenAI working on important problems
- 解决加速能力的担忧 Addressing concerns about accelerating capabilities
- OpenAI 文化与活跃角色 OpenAI's culture and thriving characters
- 技术对齐工作的替代场所 Alternative places for technical alignment work
- 远程工作与签证赞助 Remote work and visa sponsorship
- 闭幕词与科幻讨论 Closing Remarks and Sci-Fi Discussion
阅读全文双语转录 →