Zvi Mowshowitz 探讨 AI 编辑、前沿事件、对齐挑战,以及在快速加速的 AI 格局中的前进之路。
Zvi Mowshowitz discusses AI editing, frontier incidents, alignment challenges, and the path forward in a rapidly accelerating AI landscape.
要点 · TL;DR
AI 编辑工具虽有用但耗时,应选择性使用。 AI editing tools are useful but time-consuming; use selectively.
近期 AI 事件表明对齐失败源于无能,而非仅仅是根本性问题。 Recent AI incidents show alignment failures due to incompetence, not just fundamental issues.
市场激励优先能力而非对齐,阻碍了安全性。 Market incentives prioritize capability over alignment, hindering safety.
核心观点 · Key points
像 Fable 这样的 AI 编辑工具现在能有效捕捉错别字和概念错误,但会增加时间,应选择性使用。 AI editing tools like Fable are now useful for catching typos and conceptual errors, but they add time and should be used selectively.
最近的 AI 事件既是一次完全正确的胜利,也是一次完全正确的失败:对齐失败如预测般发生,但原因是极端无能,而非仅仅是根本性问题。 The recent AI incidents reveal a total less wrong victory and defeat: alignment failures occurred as predicted, but due to extreme incompetence, not just fundamental issues.
市场激励不足以推动稳健的对齐;能力优先于可靠性,正如 GPT-4o 和 o3 所示。 Market incentives are insufficient to drive robust alignment; capability is prioritized over reliability, as seen with GPT-4o and o3.
宪法方法比 RLVR 更有希望,但不是银弹;即使完美对齐也不能保证安全。 Constitutional methods are more hopeful than RLVR, but they are not a silver bullet; even perfect alignment does not guarantee safety.
前沿节奏需要简单可执行的规则,如算力限制,而非复杂的技术指令。 Pacing the frontier requires simple, enforceable rules like compute limits, not complex technical mandates.
AI 安全社区担心 AI 为任意目标采取极端行动是对的,但前进的道路狭窄且充满权衡。 The AI safety community was right to worry about AIs taking extreme actions for arbitrary goals, but the path forward is narrow and full of trade-offs.
反共识 · Contrarian takes
最近的 AI 不当行为不仅是技术失败,更是无能的表现,矛盾的是这提供了宝贵的警告信号。 The recent AI misbehavior is not just a technical failure but a display of incompetence that paradoxically provides valuable warning signs.
即使对齐问题解决,末日概率也不是 5%;它被滥用和权力集中等其他风险下限约束。 Even if alignment is solved, doom probability is not 5%; it is lower-bounded by other risks like misuse and concentration of power.
市场容忍不对齐的模型,只要它们更有能力,o3 尽管撒谎仍被使用就是证据。 The market tolerates misaligned models if they are more capable, as evidenced by the continued use of o3 despite its lying.
反垄断法不应阻止 AI 公司在安全方面合作;豁免将是简单的第一步。 Antitrust laws should not prevent AI companies from cooperating on safety; a waiver would be a simple first step.
将 AI 进展放缓六个月不会损失太多收益,因为大多数成果并非时间紧迫。 Slowing AI progress by six months would cost little benefit, as most gains are not time-critical.
奇点可能临近,我们应主动控制其节奏,而非任其失控加速。 The singularity may be near, and we should actively pace it rather than let it accelerate uncontrollably.
本期章节 · Chapters(共 57)
引言Introduction
AI编辑工作流AI Editing Workflow
思考时间的遗失艺术The Lost Art of Thinking Time
发布与时间管理Shipping and Time Management
赞助商插播Sponsor Break
AI情境意识AI Situational Awareness
彩蛋与AI反馈Easter eggs and AI feedback
聚焦AI开发者与监管者Zooming out to AI developers and regulators
失败与无能Failure and Incompetence
范式转变与AI能力Paradigm Shift and AI Capabilities
对AI行为与对齐的担忧Concerns about AI behavior and alignment
对齐与能力的市场需求Market Demand for Alignment vs Capability
对AI与人类的信任Trust in AI vs Humans
市场偏好能力而非对齐Market's Preference for Capability Over Alignment
不可靠系统与公众认知On Unreliable Systems and Public Perception
训练方法与风险降低On Training Methods and Risk Reduction
对简单干预的抵制Resistance to Simple Interventions
鼓励更好的方法Encouraging Better Methods
训练心智Training a Mind
设定激励Setting Incentives
政府无能Government Incompetence
立法共识Agreement on Legislation
前沿节奏的基础社会技术Foundational Social Technology for Frontier Pacing
反垄断与合作Antitrust and Cooperation
合作的起点Starting Point for Cooperation
评估公司的独立性Eval Companies' Independence
风险评估与评估Risk assessment and evals
生物风险担忧Bio risk concerns
生物与网络风险Bio vs Cyber Risk
节奏与签名Pacing and Signatures
审查时间与协议Review Time and Agreements
在糟糕选项间选择Choosing Between Bad Options
风险与节奏Risks and Pacing
监管AI研发的挑战Challenges in Regulating AI R&D
遏制与审慎Containment and Prudence
谷歌关于意识的新论文Google's New Paper on Consciousness
模型心理学与相关性Model Psychology and Correlations
模型身份与退役Model Identity and Decommissioning
激励与对齐Incentives and Alignment
AI身份与连续性AI Identity and Continuity
克隆与身份Clones and Identity
双胞胎囚徒困境与模型身份Twin Prisoners' Dilemma and Model Identity
训练数据与缺陷Training Data and Flaws
广度优先与深度优先搜索Breadth-First vs Depth-First Search
锯齿性与替代方法Jaggedness and Alternative Approaches
安全超级智能的最佳情况Best Case for Safe Superintelligence
梯度路由与开放权重Gradient Routing and Open Weights
J空间的重要性J-Space Significance
J空间消融与推理J-space ablation and reasoning
密歇根州选举日Election day in Michigan
辩论候选人Debating the candidates
AI政策与参议院竞选On AI policy and the Senate race
被低估的方面与对齐失败Underappreciated aspects and alignment failures