Goodfire CTO 谈 Silico 平台与 AI 安全
Goodfire CTO on Silico Platform and AI Safety
丹·巴尔萨姆 Dan Balsam · The Cognitive Revolution · 2026-08-08 · 约 117 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
Dan Balsam 讨论 Goodfire 的新平台 Silico、预测性数据调试以及 AI 概念的几何结构。
Dan Balsam discusses Goodfire's new Silico platform, predictive data debugging, and the geometry of AI concepts.
要点 · TL;DR
- 后训练主要强化预训练知识,而非增加新能力。
Post-training mostly reinforces pre-trained knowledge, not adding new capabilities. - 沿学习到的流形进行引导比朴素线性插值更有效。
Steering along learned manifolds is more effective than naive linear interpolation. - Silico 使前沿可解释性和训练工具民主化。
Silico democratizes access to frontier interpretability and training tools.
核心观点 · Key points
- 模型的大部分知识来自预训练;后训练主要是强化已有的能力。
Most of a model's knowledge comes from pre-training; post-training mostly reinforces existing capabilities. - 可解释性在于将模型分解为组件;理解几何结构能实现更好的引导。
Interpretability is about factoring models into components; understanding geometry enables better steering. - 智能体正在改变研究;Silico 使前沿可解释性和训练工具的获取民主化。
Agents are transforming research; Silico democratizes access to frontier interpretability and training tools. - 开源模型对于防御能力和避免权力集中至关重要。
Open-source models are crucial for defensive capabilities and avoiding concentration of power. - 仅靠监控可能不够;我们最终必须塑造训练过程以控制模型学习的内容。
Monitoring alone may not suffice; we must eventually shape training to control what models learn. - AI 意识不确定;鉴于潜在风险,应保持智识谦逊。
AI consciousness is uncertain; intellectual humility is warranted given potential downsides.
反共识 · Contrarian takes
- 沿学习到的流形进行引导远比朴素的线性插值有效。
Steering along learned manifolds is far more effective than naive linear interpolation. - 模型将情绪表示为轮盘;沿此几何结构引导能一致地调整情绪输出。
Models represent emotions as a wheel; steering along this geometry consistently adjusts emotional output. - 越狱就像网络攻击:一系列小操作,而非单一缺陷。
Jailbreaks are like cyber attacks: sequences of small manipulations, not single flaws. - 多智能体优化可能是危险涌现行为的可能原因,应避免。
Multi-agent optimization is a likely cause of dangerous emergent behaviors and should be avoided. - 开源与闭源模型之间的差距已大幅缩小;开源模型接近前沿。
The gap between open and closed models has shrunk considerably; open models are near frontier. - 我们应该暂停 AI 发展一段时间,以评估风险,然后再进一步推进。
We should pause AI development for a while to assess risks before advancing further.
本期章节 · Chapters(共 48)
- 引言 Introduction
- 训练后能力 Post-training and capability
- 预测性数据调试 Predictive data debugging
- 前沿挑战与开放模型 Frontier challenges and open models
- 概念几何表示 Geometric representation of concepts
- 线性表示假说与几何 Linear Representation Hypothesis and Geometry
- 流形上的引导 Steering on Manifolds
- 赞助商插播 Sponsor Break
- 寻找流形 Finding the Manifold
- 特征与流形的几何 Geometry of features and manifold
- 无监督特征与生命之树 Unsupervised features and tree of life
- 黑箱稀疏特征器 Black sparse featureizers
- BSFs与SAEs BSFs vs SAEs
- 与Graham技术的关联 Connection to Graham technique
- 参数分解与模型因子化 Parameter decomposition and model factorization
- 模型作为遗留代码库及重构 Models as Legacy Codebases and Refactoring
- 黑箱稀疏特征器在生产中的可行性 Feasibility of Black Sparse Featurizers in Production
- 大型模型中的干扰与未来展望 Interference in Large Models and Future Prospects
- 越狱与模型几何 Jailbreaks and Model Geometry
- 介绍Silico Introducing Silico
- 使用Silico的体验 Describing the Experience of Using Silico
- 使用代理的乐趣 The Joy of Using Agents
- 个人研究志向 Personal Research Aspirations
- Silico解决的难题 Hard Problems Silico Solves
- 三类难题 Three Categories of Hard Problems
- 用户体验与长期研究 User Experience and Long-Horizon Research
- 订阅详情与计算资源 Subscription Details and Compute
- 用例与社区分享 Use Cases and Community Sharing
- 公司成就与黑客松结果 Company achievements and hackathon results
- AI研究工具时代 The era of AI research tools
- 如何开始使用该工具 How to start using the tool
- 社区项目示例 Examples of community projects
- 工具局限与通用性 Tool limitations and general-purpose nature
- 为下一代模型构建 Building for Next-Gen Models
- 技能发展建议 Advice for Skill Development
- 开源与监督 Open Source and Supervision
- 开源网络模型与防御 Open Source Cyber Models and Defense
- 额外护栏与平台限制 Additional Guardrails and Platform Restrictions
- 护栏与平台安全 Guardrails and Platform Safety
- 生物风险与AI Bio Risk and AI
- 训练技术共识 Agreements on Training Techniques
- 关于禁止技术与训练塑造 On Forbidden Techniques and Training Shaping
- 纵深防御与JSpace On Defense in Depth and JSpace
- 研究焦点与开放科学 Research Focus and Open Science
- 办公室对话与意识 Office Conversations and Consciousness
- 对Claude意识的个人看法 Personal Views on Claude's Consciousness
- 意识谱系 Consciousness spectrum
- 结语与尾声 Closing remarks and outro
阅读全文双语转录 →