Far AI 的新排行榜系统测试了前沿 AI 的安全防护,发现部分模型具有韧性,但其他模型可被廉价通用越狱攻破。
Far AI's new leaderboard systematically tests frontier AI safeguards, finding some models resilient but others vulnerable to cheap universal jailbreaks.
要点 · TL;DR
OpenAI 和 Anthropic 的模型抵御了所有越狱攻击,而 Gemini 和 Grok 仅花不到 300 美元就被破解。 OpenAI and Anthropic models repelled all jailbreaks, while Gemini and Grok were broken for under $300.
分层防御——对齐、监控、探针、账户控制——加上数据过滤能有效遏制滥用。 Layered defenses—alignment, monitoring, probes, account controls—plus data filtering can effectively contain misuse.
大多数灾难性 AI 风险可通过谨慎部署避免,而非不可避免的技术宿命。 Most catastrophic AI risk is avoidable through careful deployment, not an irreducible technical fate.
核心观点 · Key points
OpenAI 和 Anthropic 的前沿模型抵御了所有测试的越狱,而 Gemini 和 Grok 则以不到 300 美元的成本被发现大量通用越狱。 Frontier models from OpenAI and Anthropic resisted all tested jailbreaks, while Gemini and Grok had many universal jailbreaks found for under $300.
结合模型对齐、思维链监控、探针和账户控制的纵深防御可以遏制 AI 滥用。 Defense in depth combining model alignment, chain-of-thought monitoring, probes, and account controls can contain AI misuse.
预训练数据过滤是一种被忽视的廉价干预措施,可以大幅降低滥用风险,尤其对开放权重模型。 Pre-training data filtering is an overlooked, cheap intervention that can substantially reduce misuse risk, especially for open-weight models.
OpenAI 与 Hugging Face 事件更多是遏制与监控失败,而非对齐失败。 The OpenAI-Hugging Face incident is more a failure of containment and monitoring than of alignment.
大多数灾难性 AI 风险是人为且可通过谨慎部署避免的,并非不可降低的技术宿命。 Most catastrophic AI risk is man-made and avoidable through careful deployment, not irreducible technical fate.
社会工程和人格操纵是越狱的核心技术;花哨的混淆手段仅带来边际收益。 Social engineering and persona manipulation are the core jailbreak techniques; exotic obfuscation adds only marginal gains.
反共识 · Contrarian takes
经过十年的怀疑,Adam 现在认为在正确技术下,防御对 LLM 滥用占据主导地位。 After a decade of skepticism, Adam now sees defense as dominant against LLM misuse, given the right technologies.
尽管早前有警告,将 AI 人格化已被证明对越狱和理解模型极其有用。 Anthropomorphizing AI, despite earlier warnings, has proven extremely useful for jailbreaking and understanding models.
今天的 LLM 最好被理解为下一个词预测、人格选择与强化学习驱动的目标达成模式的结合。 LLMs today are best understood as combining next-token prediction, persona selection, and RL-driven goal-achiever modes.
特定领域的通用越狱足以造成危害;攻击者不需要跨域泛化。 Domain-specific universal jailbreaks are sufficient for harm; attackers don't need cross-domain generalization.
开放权重模型仍然极其脆弱;微调可移除防护,但预训练过滤能提高成本。 Open-weight models remain extremely vulnerable; fine-tuning can remove safeguards, but pre-training filtering can make it costly.
无需突破,仅靠纪律性工程与协调,我们就能将存在风险从约 10% 降至 1%。 We can reduce existential risk from roughly 10% to 1% without breakthroughs, simply by disciplined engineering and coordination.