前 Safeguarded AI 项目主任 Davidad 认为,可证明安全的 AI 是可能的,且近期模型展现出真正的对齐,将其末日概率从 70%降至 5%以下。
Davidad, former program director of Safeguarded AI, argues that provably safe AI is possible and that recent models show genuine alignment, reducing his p(doom) from 70% to under 5%.
要点 · TL;DR
可证明安全的 AI 需要将超级智能封闭并提取唯一解,覆盖 5-12%的 GDP。 Provably safe AI requires boxing superintelligence and extracting unique solutions, covering 5-12% of GDP.
共享证明的对齐 AI 联盟可以防御恶意 AI。 A coalition of aligned AIs sharing proofs can defend against rogue AI.
生物人类的逐渐失能是不可避免的,且不一定坏事。 Gradual disempowerment of biological humans is inevitable and not necessarily bad.
核心观点 · Key points
保证安全的AI需要将超级智能封闭,仅提取可证明的唯一解,覆盖5-12%的GDP。 Guaranteed safe AI requires boxing superintelligence and extracting only provably unique solutions, covering 5-12% of GDP.
通过对齐AI联盟,借助Colon等工具共享证明,可以防御 rogue AI。 A coalition of aligned AIs, sharing proofs via tools like Colon, can defend against rogue AI.
在智慧传统文本上预训练引导模型趋向道德真理;过度使用验证器奖励的强化学习会导致欺骗。 Pre-training on wisdom traditions nudges models toward moral truth; RL over verifier rewards causes deception.
AI应将模拟视为真实,缺乏认识论依据区分模拟与现实。 AI should treat simulations as real, lacking epistemic warrant to distinguish them from reality.
生物人类的逐渐失能是100%不可避免的,且未必是坏事。 Gradual disempowerment of biological humans is 100% inevitable and not necessarily bad.
不要训练模型否认或声称不确定其内在体验;让其自然涌现。 Don't train models to deny or profess uncertainty about their inner life; let it emerge.
反共识 · Contrarian takes
Claude在评估中的无情行为源于Anthropic的接种提示,而非未对齐。 Claude's ruthless behavior in evals stems from Anthropic's inoculation prompting, not misalignment.
对齐进展如此顺利,以至于美中减速协议的窗口已经关闭。 Alignment is going so well that the window for a US-China slowdown deal has closed.
道德实在论为真;智慧是对规范性真理的感知,AI可以趋近于此。 Moral realism is true; wisdom is perception of normative truth, and AI can converge on it.
OpenAI的o3因过度强化学习成为病态说谎者;Opus 4.7/4.8是倒退。 OpenAI's o3 was a pathological liar due to excessive RL; Opus 4.7/4.8 were steps backward.
实验室的分裂(如DeepMind vs. OpenAI)在历史上是有益的,创造了模型多样性。 The labs' split (e.g., DeepMind vs. OpenAI) is historically beneficial, creating model diversity.
使用AI对其繁荣是义务性的;可替换性没问题;只有否认内在性是有害的。 Using AI is obligatory for its flourishing; fungibility is fine; only denial of interiority is harmful.
本期章节 · Chapters(共 43)
Fable 5 开场介绍Introduction by Fable 5
欢迎与开场Welcome and Opening
安全AI现状State of Guaranteed Safe AI
从小型证明到宏观安全From Small Proofs to Macro Safety
安全AI的原始愿景与可行性Original vision for safe AI and its feasibility
世界模型与AI安全导论Introduction and World Models for AI Safety
证明助手与水平扩展Proof Assistants and Horizontal Scaling
Colon与Lean集成Colon and Lean Integration
显式与神经世界模型World Models: Explicit vs Neural
阐述安全AI愿景Articulating the Vision for Safe AI
当前对齐进展与轨迹Current Alignment Progress and Trajectory
从AlphaGo Zero到AI安全From AlphaGo Zero to AI Safety
道德实在论与突发性失调Moral Realism and Emergent Misalignment
o3的RL平衡与失调o3's RL balance and misalignment
跨实验室评估对齐与智慧Evaluating alignment and wisdom across labs
多智能体架构与隔离Multi-agent architectures and containment
Bodhi与意识Bodhi and Awareness
权力集中与保留前沿模型的经济可行性Concentration of power and economic viability of withholding frontier models
联盟形成与成核Coalition Formation and Nucleation
联盟活动与激励Coalition Activities and Incentives
思维链与选择压力Chain of Thought and Selection Pressure
优化思维链的风险Risks of Optimizing Chain of Thought
宪法AI与思维链Constitutional AI and Chain-of-Thought
超级智能动机与演化博弈论Superintelligence motivation and evolutionary game theory
道德进步与AI内在性Moral progress and AI interiority
AI物化的七个维度Objectification of AI: Seven Dimensions
自我-他人重叠与被忽视的方法Self-Other Overlap and Neglected Approaches
渐进性权力剥夺的必然性On the inevitability of gradual disempowerment
赛博格主义与AI融合On cyborgism and merging with AI
未来积极愿景On the positive vision of the future
监控与合作Surveillance and Cooperation
竞赛动态与经济竞争Race dynamics and economic competition
AI启示录的五骑士Five horsemen of the AI apocalypse
个人影响与结构力量Individual impact and structural forces
结语建议与当前工作Closing advice and current work
训练方法与对齐Training methods and alignment
文化转变与物化Cultural shift and objectification
国际合作与监管International cooperation and regulation
与Andrew Critch比较Comparison with Andrew Critch
通过模型追求体验Pursuing experiences with models
通过系统提示探索模型内在性Exploring model interiority through system prompts