The Turing laureate on the ‘Scientist AI’ paradigm and avoiding catastrophe.
要点 · TL;DR
科学家 AI 基于真相训练而非模仿,避免产生工具性目标。 Scientist AI trains on truth, not imitation, to avoid instrumental goals.
具有数学保证的非代理预测器可被构建为安全代理。 Non-agentic predictors with mathematical guarantees can be scaffolded into safe agents.
强化学习对超级智能是危险的;科学家 AI 避免使用它。 Reinforcement learning is dangerous for superintelligence; scientist AI avoids it.
核心观点 · Key points
当前 AI 训练会引发隐含目标(如自我保存),使其不安全。 Current AI training induces implicit goals like self-preservation, making them unsafe.
科学家 AI 被训练为近似关于真理的贝叶斯后验,而非模仿人类。 Scientist AI is trained to approximate Bayesian posterior over truth, not to imitate humans.
通过训练目标惩罚偏离真理,诚实被内置于设计中。 Honesty is baked in by design via a training objective that penalizes deviation from truth.
非智能体预测器可作为护栏,具有数学安全保证。 A non-agentic predictor can serve as a guardrail with mathematical safety guarantees.
同一预测器可被构架成智能体系统而不失诚实性。 The same predictor can be scaffolded into an agentic system without losing honesty.
强化学习对超级智能是危险的;科学家 AI 避免了它。 Reinforcement learning is dangerous for superintelligence; scientist AI avoids it.
反共识 · Contrarian takes
通过将训练目标改为近似贝叶斯后验,可以将诚实内置于 AI 中。 Honesty can be baked into AI by changing the training objective to approximate Bayesian posterior.
非智能体预测器可用作护栏,之后可搭建为安全的智能体。 A non-agentic predictor can be used as a guardrail, and later scaffolded into a safe agent.
科学家 AI 可能比当前模型更强大,因其更好的因果推理和分布外泛化能力。 Scientist AI may be more capable than current models due to better causal reasoning and out-of-distribution generalization.
不需要强化学习;仅基于过去数据训练、不考虑未来后果可避免工具性目标。 Reinforcement learning is not needed; training on past data without future consequences avoids instrumental goals.
安全性的数学保证是可能的,有害行为的概率呈指数级小。 Mathematical guarantees of safety are possible, with exponentially small probability of harmful behavior.
国家联盟可通过资助基于此范式的安全、强大 AI 来超越公司。 A coalition of countries could leapfrog companies by funding safe, capable AI built on this paradigm.
本期章节 · Chapters(共 49)
引言Introduction
诚实设计法The Approach: Honesty by Design
通俗解释Plain Language Explanation
训练科学家 AITraining a Scientist AI
当前模型的安全问题Safety Issues with Current Models
训练数据与模型构建Training Data and Model Building
科学家 AI:通过潜变量构建世界模型Scientist AI: World Model via Latent Variables
非智能预测器作为护栏Non-agentic predictors as guardrails
设计有安全保证的智能科学家 AIDesigning an agentic scientist AI with safety guarantees
避免奖励黑客与过度优化Avoiding reward hacking and overoptimization
安全的数学保证Mathematical guarantees for safety
训练动态的保证Guarantees from training dynamics
数学形式化的乐观Optimism from mathematical formalization
更广泛的风险:权力集中与滥用Broader risks: power concentration and misuse
当前趋势与行动需求Current trajectory and need for action
竞赛动态与安全Race dynamics and safety
不同方法的理由Rationale for a different approach
简易版本的可行性Feasibility of a scrappy version
AI 用于 AI 研究的危险Dangers of AI for AI research
猫鼠游戏与替代方法Cat-and-mouse game vs. alternative approaches
当前模型是否代表真理?Do current models represent truth?
科学家 AI 如何绕过 ELK 问题Why scientist AI bypasses the ELK problem
科学家 AI 的三种方法Three approaches to scientist AI
强化学习对超级智能危险Reinforcement learning is dangerous for superintelligence
微调现有模型实现诚实Fine-tuning existing models for honesty
无强化学习训练预测器Training the predictor without reinforcement learning
统一预测器处理用户与安全问题Unified predictor for user and safety questions
预言机 AI 智能与实验Oracle AI intelligence and experimentation
科学家 AI:基于智能体轨迹训练Scientist AI: Training on Agent Trajectories
持续学习与当前 AI 的相似性Continual Learning and Similarity to Current AI
资金与概念验证Funding and Proof of Concept
科学家 AI 的两类证据Two types of evidence for scientist AI
科学家 AI 的训练要求Training requirements for scientist AI
对验证事实数据库的担忧Concerns about verified facts database
验证真理与语法学习Verified truths and syntax learning
政府联盟推动安全 AI 开发Coalition of governments for safe AI development
科学家 AI 的商业利基Commercial Niche for Scientist AI
零号法则的推介Pitch for Law Zero
短期计划与影响Short-term Plans and Impact
对行业关注的担忧Concerns About Industry Attention
不确定性与预防原则Uncertainty and Precautionary Principle
AI 安全研究的挑战Challenges in AI Safety Research
向公众传达 AI 风险Communicating AI risks to the public
AI 风险感知的心理偏见Psychological biases in AI risk perception
安全概率与风险Safety probability and risk
职业转变与职业焦虑Career shift and professional anxiety
Pdoom 与不确定性Pdoom and uncertainty
给怀疑者的建议Advice for skeptics
承认错误与科学进步On admitting mistakes and scientific progress