Goodfire CTO 谈 Silico 平台与 AI 安全

Goodfire CTO on Silico Platform and AI Safety

丹·巴尔萨姆 Dan Balsam · The Cognitive Revolution · 2026-08-08 · 约 117 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Dan Balsam 讨论 Goodfire 的新平台 Silico、预测性数据调试以及 AI 概念的几何结构。

Dan Balsam discusses Goodfire's new Silico platform, predictive data debugging, and the geometry of AI concepts.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 48)

全文 · Full transcript(中英对照)

引言 Introduction

Host

大家好,欢迎回到《认知革命》。今天我的嘉宾是 Goodfire 的 CTO 丹·巴尔萨姆,这是一家专注于机制可解释性的初创公司。我们这次对话的契机是 Silico 的发布,这是一个长周期、智能体式的机器学习研究平台,最初是 Goodfire 的内部工具,现在旨在将 Goodfire 来之不易的专业知识大众化,包括 GPU 集群管理、研究品味,以及各种实验可视化和验证技术,企业客户每席位每月 1000 美元。这并不便宜,但相比 Goodfire 的高接触研究服务——由前向部署的研究工程师组成,费用轻松达到七位数——它便宜了两个数量级。而且,与公司的公益章程和安全导向使命一致,他们很快会为安全和对齐研究人员推出特别定价和资助,我自己肯定也会申请。当然,我们聊的远不止这个平台,首先从 Goodfire 研究的最新进展开始。我们讨论了他们在预测性数据调试方面的工作,这项技术利用可解释性技术来识别网络更新可能影响的概念,从而在异常变成令人不悦的行为意外之前发现并处理它们。我们深入探讨了他们关于大语言模型用来表示高级概念的复杂且往往相当优美的几何结构的一系列论文,包括我们应如何将其理解为线性表示假说的演进,他们如何利用这种新的更深理解来改进模型引导,以及他们如何识别出诸如元素周期表、生命之树等高级概念的空间表示。一路下来,我们听到了丹对关键问题的看法,包括开源模型对于避免危险的权力集中的重要性,他对 AI 助长未来流行病或其他生物灾难的担忧程度,Goodfire 为防止 Silico 平台被滥用所采取的措施,丹签署最近的《为前沿踩刹车》公开信的原因,以及他认为 AI 研究界理想的前进方向。为什么他对监控技术持合理乐观态度,但依然相信我们最终别无选择,只能有意设计控制模型在训练过程中学习内容的技术。最后,他认为哪些训练技术很可能被证明有问题,至少目前应完全避免。他还分享了他迄今最喜欢的 Silico 用例,辟谣了关于 Goodfire 隐藏研究议程的网络传言,直截了当地说没有,甚至让我们一窥 Goodfire 的研究人员最近在午餐桌上热烈讨论的话题——鉴于最近的研究,这或许并不令人意外,他们强调 AI 福祉和意识的神秘性。说到这里,希望大家喜欢这场富有教育意义且发人深省的对话,关于 AI 思维的形态,以及新的高度自主的机器学习研究平台 Silico,嘉宾是 Goodfire 的联合创始人兼 CTO 丹·巴尔萨姆。

Hello and welcome back to the Cognitive Revolution. Today I'm speaking with Dan Balsam, CTO of mechanistic interpretability startup Goodfire. The occasion for this conversation is the launch of Silico, a long-horizon, agentic machine learning research platform which began as an internal Goodfire tool but is now meant to democratize access to Goodfire's hard-won expertise including GPU cluster management, research taste, and all sorts of experimental visualization and validation techniques at $1,000 per month per seat for enterprise customers. It's not cheap, but compared to Goodfire's high-touch research engagements, which are staffed by forward-deployed research engineers and can easily reach into seven figures, it is two orders of magnitude more affordable. And consistent with the company's public benefit charter and safety-focused mission, they will soon be introducing special pricing and grants for safety and alignment researchers, which I definitely intend to apply for myself. Of course, we cover a lot more than the platform, beginning with an update on Goodfire research. We discuss their work on predictive data debugging, which uses interpretability techniques to identify the concepts that network updates are likely to affect, thus making it possible to identify and address anomalies before they become unpleasant behavioral surprises. We go deep on their series of papers on the intricate and often quite beautiful geometries that large language models use to represent advanced concepts, including how we should understand this as an evolution of the linear representation hypothesis, how they're using this new deeper understanding to improve model steering, and how they've identified spatial representations of such advanced concepts as the periodic table, the tree of life. Along the way, we get Dan's take on key issues, including the importance of open-source models for avoiding dangerous concentration of power, how worried he is about AI's contributing to future pandemics or other bio-disasters, the steps that Goodfire is taking to prevent misuse of the Silico platform, Dan's reasons for signing on to the recent Pacing the Frontier letter, and what he thinks the AI research community would ideally do going forward. Why he's reasonably optimistic about monitoring techniques, but nevertheless believes that we will ultimately have no choice but to intentionally design techniques that control what models learn in the training process. And finally, which training techniques he believes are sufficiently likely to prove problematic that they should be avoided entirely, at least for now. He also shares some of his favorite use cases of Silico so far, dispels internet rumors about hidden research agendas at Goodfire, stating plainly that there are none, and even offers a glimpse into what the researchers at Goodfire are actively discussing around the lunch table these days, which perhaps unsurprisingly given recent research emphasizes the mysteries around AI welfare and consciousness. With that, I hope you enjoy this educational and thought-provoking conversation about the shape of AI thought and the new highly autonomous machine learning research platform, Silico, with Dan Balsam, co-founder and CTO of Goodfire.

Host

丹·巴尔萨姆,Goodfire 的 CTO。欢迎回到《认知革命》。

Dan Balsam, CTO at Goodfire. Welcome back to the Cognitive Revolution.

Dan

谢谢邀请,我很期待这次对话。

Thanks for having me. I'm excited for this.

Host

你们一如既往地高产,我们有很多要聊的。研究,还有新产品——这本身又是一个研究平台产品,节奏真是永不停歇。先问你个开场问题:在如今 AI 这场永无止境的冲刺中,你撑得住吗?

You guys are prolific as always and we've got a lot to cover. Research and a new product which is in turn a research platform product and it's the pace is really relentless. Let me just ask you that for starters. How are you holding up in the eternal sprint that is the AI game these days?

Dan

我觉得 Goodfire 有一支非常出色的团队,这里的每个人都真心相信这个使命,并且非常努力地工作,这总是极其激励人心。我们为了 Silico 的产品发布一直拼命推进,现在能在这之后喘口气真是太好了。不过,很快还会有更酷的东西。

I think we have a really incredible team at Goodfire and yeah everyone here like really believes in the mission and is working really hard and that's always extremely motivating. We've been pushing really hard to get our product launch with Silico and it's nice to be able to take a deep breath now on the other side of that. But yeah, even more cool things coming soon.

Host

好,我们先从研究说起。每当回想起三年前那种叠加的玩具模型,再看看现在我们已经走了这么远,我总是感到惊叹。Goodfire 博客上有几篇文章很突出,我想快速过一遍,我们只能在高层次上聊,因为内容太多了,以前我们可以一篇篇深入探讨,今天只能稍微浅一些。其中一篇引起波澜的文章叫预测性数据调试。

Well, let's start with some research. So it's I'm just always amazed when I think back to the kind of toy models of superposition is only like three years right ago now and we have come so far. A few things that jumped out on the Goodfire blog that I want to just run through and we'll have to do it at kind of a high level because there's too much to do the we used to do deep dives on paper by paper. We'll have to go a little bit more superficially today. But one that made some waves was called predictive data debugging.

Host

关于这篇,我想先给你我的理解,然后你详细说说,或者告诉我你认为它特别有用的地方,或者论文发表后你们看到了什么。我的概括是:基本上,如果你有一种解释模型的方法,比如 SAE,或者我们稍后也会聊到的特征化器,那么你可以把一堆数据,比如微调或后训练数据集,跑一遍,看看哪些概念在数据集通过时频繁激活。然后关键的洞见是,被激活的概念与被训练过程修改的概念之间存在很强的相关性。所以我觉得这一点大家要记住。这并不令人震惊,但值得注意的是,这种关系相当强。然后,当你看到这些被激活的概念,知道它们就是会被修改的那些,你就可以看看:这里有没有什么概念对我们来说很奇怪、很意外,我们并不想用现有数据集去动它们?如果有,你就可以快速定位是哪些数据点导致了这些特征的出现。然后你可能会发现,数据集中确实有些东西我们或许应该三思。也许应该过滤,也许应该修改。这给了你一条路径,有望最小化后训练或微调工作中出现的不想要的行为意外。我理解得怎么样,还有什么我应该知道的?

And for this one, I just kind of want to give you my interpretation and then let you kind of elaborate on that or tell me where you think it'll be particularly useful or what you guys have seen since the paper came out. My synopsis of this one was that basically if you have a way of interpreting a model like an SAE or we'll get into featurizers I think a little bit later as well then you can run a bunch of data through it your like fine-tuning or post-training data set you can look at what concepts are coming up active a lot when we put this data set through and then the kind of insight is there's a strong correlation between the concepts that are active and the concepts that are being modified by the training process. So I think that right there is like file that away folks as something to remember. Not shocking but like it's it's notable that the relationship is quite strong there. And then when you see these concepts that are active and you know that those are the ones that are going to be modified, you can just look and see like are there any concepts here that are kind of strange to us, surprising, we don't really intend to be monkeying around with with the data set that we have at hand. And if so, then you can quickly zoom in on what are the data points that have caused these features to come up. And then you might find that actually there's some stuff in our data set we maybe ought to think twice about. Maybe we ought to filter. Maybe we ought to modify. And this gives you a route to hopefully minimizing unwanted surprises in the behavior that you get from your post-training or fine-tuning work. How'd I do and what more should I know?

Dan

不,我觉得你理解得差不多。我认为这里一个重要的直觉——这可能是我们很多工作背后的主题——是有相当多的证据表明,模型的大部分知识,比如大部分知识和能力,都来自预训练。后训练中发生的事情,包括强化学习,主要是让预训练中低概率的事件变得更可能,所以后训练过程中并没有多少新知识或新能力被固化到模型中。这有些争议,但我认为这是我们大体上相信的观点,并且它影响了很多我们思考问题的方式。

No, I think that that sounds about right. I think like one of the intuitions here that's important and I think this is maybe a theme that underlies a lot of our work is there's a good amount of evidence that models don't most of what a model knows like most of its sort of like knowledge and capabilities come from pre-training what happens in post-training including like RL is mostly making low likelihood events from pre-training more likely and so there's not that much new knowledge or like sort of new capabilities that get baked into models in post-training process. This is somewhat debated, but I think this is like a view that we think is mostly true and informs a lot of how we think about things.

训练后能力 Post-training and capability

Dan

所以,因为在任何后训练过程中,权重只会发生相对较小的调整,模型产生后训练所要达到的任何结果的原始能力,在某种程度上已经存在于模型之中。我认为现在我们正处于一个强化学习(RL)应用非常密集的时代,这已经不再适用了,但在过去基于人类反馈的强化学习(RLHF)风格的后训练时代,我认为众所周知的是,基础模型有时比经过指令微调和 RLHF 的对应模型更有能力,因为后者会出现一点模式坍缩。所以,思考 RL 训练中发生什么的一种方式是:在指令微调阶段有一点模式坍缩,而基础模型极其有能力,但非常怪异且难以提示。于是你让它们变成一种具有更直观的人类面向 API 的格式,但随后你又想恢复或强化模型在预训练期间真正学到的一些能力。这百分之百正确,但我认为它可能只是方向性正确。

And so, because there's only relatively small nudges in the weights that are happening in any type of post-training process, most of the raw capability to produce whatever outcome post-training is going to do already exists in the model in some way. I think now we're in an era where things are RL so heavily it's not true, but back in the day of RLHF-style post-training, I think it was pretty well known that base models were sometimes more capable than their instruction-tuned RLHF counterparts, where there was a little bit of mode collapse happening. So one way to think about what's happening in RL training is that you have a little bit of mode collapse in the instruction tuning phase, and you have these base models that are extremely capable but very weird and hard to prompt. So you make them into a format that has a more intuitive human-facing API, but then you want to bring back out some of those capabilities or reinforce some of the capabilities that the model actually learned during pre-training. It's 100% true, but I think it's probably directionally true.

预测性数据调试 Predictive data debugging

Dan

所以,关于预测性数据调试,我认为基本思想是,你可以观察模型在处理某些数据时已经在想什么。这相对能预测如果模型要在这批数据上训练,这些数据会强化模型中的什么。所以,查看特征并理解这些特征与下游行为的相关性,如果某些数据出人意料地加重了某些特征,那么它就能相当准确地预测模型是否会从该数据中学到一些偏离目标的效果。我认为那篇论文中最有趣的一点是,我们探索了许多不同的缓解方法。我们研究了奖励塑形,即训练或涉及一个来自模型自身激活的奖励过程。就像,嘿,从这些数据中学你该学的,但如果你学到这个特定特征,可能会有一个惩罚。我们还研究了数据过滤。我觉得那篇论文中最喜欢的是,参与的研究人员表明,这两者之间存在相当深刻的同构性。它们就像同一座山的两个侧面——过滤数据和奖励塑形——它们达到的效果大致相同,偏离目标的效果也大致相同。所以,如果你有一个行为,你做了预测性数据调试,你认为模型会学到它,而你不希望它学到,或者它的行为会以不好的方式改变,你的选择可以是过滤数据(如果你有足够的数据),或者以某种方式干预训练过程。是的,我认为从那以后我们一直在扩展这个想法。我认为考虑它在强化学习中的应用非常令人兴奋。例如,在 RL 中,我们用 DPO 做了,但真正的 RL——区别在于,某些 rollout 可能包含你不想让模型学到的信息,即使是微妙地。

And so with predictive data debugging, I think basically the idea is you can look at what a model is already thinking as it's looking at some data. And that's relatively predictive of what that data is going to reinforce in that model if it was going to be trained on it. So looking at the features and understanding that those features correlate with downstream behaviors, if there are features that are surprisingly upweighted by some data, then it can tell you fairly predictively whether the model is going to learn some off-target effect from that data. And so I think one of the most interesting things from that paper is we explored a bunch of different mitigation methods. So we looked at reward shaping, which is training or involving some reward process that comes from the activations of the model itself. So like, hey, learn what you're going to learn from this data, but maybe a penalty if you're learning this particular feature. And we also looked at data filtering. And I think my favorite thing from that paper was that the researchers involved showed that there's a pretty deep isomorphism between those two things. They kind of are two sides of the same mountain—filtering the data and reward shaping—and they achieve approximately the same effects and approximately the same amount of off-target effects as each other. So your options, if you have a behavior and you do some predictive data debugging and you think the model's going to learn it and you don't want it to learn it or its behavior is going to change in a bad way, your options could be filter your data if you have enough data, or it could be intervene in the training process in some way. And yeah, I think this is something that we've expanded on since then. I think it's pretty exciting to think about its applications to RL, for instance. In RL, we did it with DPO, but true RL—the difference is some rollouts may contain information that you don't want the model to learn, even subtly.

Host

我不知道你在指什么。你是指某件事吗?

I don't know what you're referring to. Are you referring to something?

Dan

是的。对,这很切题。所以,能够说“实际上我们想丢弃这个 rollout”本身就很有价值。但我们想要能够做到的,并且我们认为随着时间的推移大致等效的,是真正能够干预模型,说“嘿,这里有一个电路或一个特征,数据试图加重它,而我们显然不希望它加重”,于是我们以那种方式干预训练过程。

Yes. Yeah. It's topical. And so being able to say, actually we want to discard this rollout, is pretty valuable in and of itself. But what we want to be able to do, and what we think is roughly equivalent over time, is actually be able to intervene in the model and say, hey, here's a circuit or here's a feature that the data is trying to upweight that we obviously don't want it to try to upweight, and so we intervene in the training process in that way.

前沿挑战与开放模型 Frontier challenges and open models

Host

我们可能到最后再回到这个话题。我有一些宏观的大问题想问你。显然,开发你希望在前沿有用的技术的一个挑战是,没有太多开放权重模型可以让你进行高强度 RL 实验,而前沿实验室正在进行这种高强度 RL,这导致了我们看到的那些丰富多彩的问题行为。但与此同时,我也一次又一次地惊叹于那些惊人的工作,包括我经常想到的 Cameron Berg 的论文,关于 Llama 3 70B 上欺骗与角色扮演特征以及主观体验声称之间的反相关性。那已经是两年前了。那么,你们能看到那些你认为导致这些无休止黑客行为的相关特征吗?

We'll come back to this probably toward the end. I have some kind of zoomed-out big picture questions for you. One of the challenges obviously with trying to develop techniques that you hope will be relevant at the frontier is there's not too many open weights models that you can hack on that have the intensity of RL that is going on at the frontier labs, which is leading to these colorful problematic behaviors that we're seeing. But at the same time, it also really amazes me over and over again that astounding work, including the Cameron Berg paper that I think about all the time about the anti-correlation between deception and role-playing features and claims of subjective experience on Llama 3 70B. That's like 2 years old. So, are you guys able to see features that you think are kind of the relevant features that are leading to these like relentless hacking behaviors?

Dan

嗯,我认为这对我们来说是一个活跃的研究领域。是的。而且我们希望未来能发表更多相关成果。顺便说一句,我实际上认为开放和封闭模型之间的差距已经大大缩小了。我使用 Kimi K3。我根据不同用途使用 Opus、Fable、Soul 和 Kimi K3。我们已经构建了可解释性基础设施和训练基础设施,现在这些都集成在我们的产品中,这使我们能够将预测性数据调试等方法扩展到 Kimi 和 GLM 等模型,并在这个规模上复现相同的结果。现在有人推测,训练新模型的大部分算力现在来自强化学习,而不是预训练。所以,如果你想达到前沿水平,RL 中肯定有很多规模可扩展,但这相当昂贵且困难。但我们有原始基础设施。而且我认为我们已经弥合了差距——我认为可解释性的声誉过去是你在玩具模型上做的事情——而现在我们已经构建了基础设施,并让世界能够至少在接近前沿规模上做这件事。

Well, I think that's an active area of study for us. Yeah. And something that we hope to publish more on in the future. For what it's worth, I actually think the gap between open and closed models has shrunk quite considerably. I use Kimi K3. I use Opus and Fable and Soul and Kimi K3 for different things. And we've built the interpretability infrastructure and the training infrastructure which is now all in our product, which lets us scale things like predictive data debugging to models like Kimi and GLM, where we were able to replicate the same results at that scale. Now it's speculated that most of the compute that goes into training new models now is coming from RL and not from pre-training. So there is certainly a lot of scale in RL if you want to get to the frontier level, which is quite expensive and difficult to do. But we have the raw infrastructure for it. And I think we've bridged the gap from—I think the reputation of interpretability used to be that it was something you did on toy models—and I think now we've built and are making accessible to the world the infrastructure to do this on at least close to frontier scale.

概念几何表示 Geometric representation of concepts

Host

嗯,哦,另一条线索——如果我说错了请纠正我——但似乎最近来自 Goodfire 的论文和博客文章数量最多的主题是试图弄清楚模型用来表示概念的更详细的几何结构。我认为我们过去已经讨论过线性表示假说,我会用非常通俗的话总结为:模型基本上将一个概念表示为激活空间中的一个方向,而该概念的强度或显著性由指向该空间的向量的大小来表示。

Well, oh, another thread that has been—correct me if I'm wrong—but it seems like the biggest thread in terms of the number of papers and blog posts that have come out recently from Goodfire is around trying to figure out the more detailed geometries that models use to represent concepts. I think we've covered in the past the linear representation hypothesis, which I would summarize super plain-spokenly as models basically represent a concept as a direction in their activation space, and the intensity or the salience of that concept is represented by the magnitude of the vector that points in that space.

线性表示假说与几何 Linear Representation Hypothesis and Geometry

Host

现在你把这个问题复杂化了不少,我们远远超出了这些在空间中的单个方向,发现了各种各样的几何结构,其中一些相当直观,比如星期几是一个圆,但有些则相当奇特,比如我在准备这次对话时有机会看到的一些蛋白质模型流形。那么,也许就快速入门而言,如果给我 6 到 9 个月前的版本,我们应该如何理解线性表示假说?新的简短版本是什么,每个人都能带回家并默念一遍,以确保自己有一个良好的工作理解?

Now you're complicating that quite a bit and we're going well beyond these sort of individual directions in space and finding all kinds of different geometries which some of which are like pretty intuitive like the days of the week are a circle but some of which get pretty exotic like some of the protein model manifolds that I've had the chance to look at in preparing for this. So maybe just for like super quick starters, what's kind of the headline if I gave you the 6 to 9 months ago version of what we should understand to be going on with the linear representation hypothesis? What is the like new short version that everybody can kind of take home and recite to themselves to make sure they have a good working understanding?

Dan

是的,在很多方面,我认为这仅仅是我们之前讨论方式的一个推广,不同的人对线性表示假说的定义略有不同。我认为最站得住脚的版本就是说特征是可以线性解码的,我认为这基本上是正确的,一般来说,模型不需要非线性计算就能从残差流中读出特征。但几何成分的引入在于,特征并不是像 SAE 构建所设想的那样,你可能想象模型编码了一堆完全正交的概念。所以它实际上就是一堆 one-hot 编码的分类特征,特征的大小对应模型对它的关注程度。但实际上,我们发现的结构要复杂得多。我更倾向于认为,模型更像是一个子空间的稀疏混合。你会为不同类型的概念设置子空间,对吧?比如你可能有星期几的子空间,它本身可能存在于一个更概念化的日历时间子空间中。所以在不同的分辨率水平上,你有这些不同的结构。这些结构的几何形状非常重要,因为几何形状编码了你可以对它们执行的操作。这有点像语义,不仅是个别概念,而且概念空间是由这些概念在某种几何中的相互关系编码的。更让人困惑的是,从一个概念到另一个概念的关系、操作和映射,都是对这些流形的操作。所以你将一个流形映射到另一个流形,通过某种计算。所以最天真的版本是,你有星期一、星期二、星期三、星期四,模型中的某个地方知道星期一之后是星期二,星期二之后是星期三,等等。但实际上,如果你仔细想想,这是一种非常低效的表示星期几的方式。那将是一种非常像 if 语句意大利面条式代码的表示方式。更高效的方式是把它表示成一个轮子。在这个世界里,沿某个方向的大小通常确实在某种程度上对应模型的确定性。所以如果一个特征,比如星期一激活得很高,那么模型就非常确信它应该考虑星期一。但所有这些日子之间的关系本身就是一个非常有表现力和丰富的东西。我认为我们相信的是,如果你不理解特征之间的关系,你就无法真正理解模型。这就像理解元素周期表和理解化学之间的区别。你可以拥有所有单个元素,这给你一些信息,但真正重要的是它们如何组合,它们形成的结构,这才是能帮助你了解世界复杂性的东西。

Yeah, in many ways I think it's just a generalization of like the way we were discussing things before and different people define the linear representation hypothesis like slightly differently. I think the most like sort of defensible version of it is just saying that like features are like linearly decodable which I think is like true essentially like it doesn't require nonlinear computation generally speaking in a model in order to for the model to read out a feature from the residual stream. But I think where the geometry components come in is that the features aren't like sort of like naive maybe like SAE build take on things is like you could imagine that the model is encoding like a bunch of totally orthogonal concepts to each other from each other. And so it's like it's really just like a bunch of one hot encoded categorical features and then like the magnitude of the feature corresponds to how much the model's thinking about it. But in actuality the structures that we find are like significantly more complicated. I would think of it more I think the way to think about a model is more like a sparse mixture of subspaces. So you'll have subspaces for different types of concepts, right? Like maybe you have your days of the week subspace which itself lives in maybe a more like conceptual calendar time subspace. And so you have like at different levels of resolution these like different structures. And the geometry of those structures is really important because the geometry of those structures encodes what operations you can perform on them. It's sort of like the semantics of not just the individual concept but like the concept space are like encoded by the relationship of those concepts with each other in some sort of geometry. And to make things extra confusing, the relationship the operations and the mappings that are performed from one concept to another act as operations over those manifolds. So you map a manifold to a different manifold over some some computation. And so like the sort of naivist version would be like well you have the days of the week Monday, Tuesday, Wednesday, Thursday. And there's just somewhere in the model that knows that Monday goes to Tuesday and Tuesday goes to Wednesday, etc. all the way around. But that's actually would be a super inefficient way if you think about it to like represent the days of the week. That would be a very like if statement spaghetti code way of representing it. Like the much more efficient way is to represent it as a wheel. And in this world like the magnitude along some direction often does correspond to the model's like certainty in some way. So if a feature if Monday is activating very high then the model's very confident that it should be thinking about Monday. But the relationship between all those days is itself like a very expressive and and rich thing. And I think what the thing that we just believe is that you're not really going to understand the model if you don't understand the relationship between the features. It's like maybe the difference between understanding the periodic table and understanding chemistry. You can have the all the individual elements and that gives you some information but like really the way in which they combine the structures in which they form that's like what can start to get help you gain a sense of the complexity of the world.

Host

是的,这很迷人。听你这么说,我有点好奇,如果只在一堆原始化学数据上训练一个模型,它从未接触过元素周期表,那么元素周期表的结构是否可以从模型中恢复出来?我们实际上已经做到了,而且确实如此。

Yeah, it's fascinating. Now that you say that I'm kind of wondering if the structure of the periodic table would be recoverable from a model that was just trained on like a bunch of raw chemical data that never knew and we've actually done this and it is

Dan

是的。真的吗?好的。有趣。

Yeah. Really? Okay. Interesting.

Host

是的。是的。你可以从化学模型中恢复出一些非常有趣的信息。

Yeah. Yeah. You can recover some pretty interesting information from chemistry models.

Dan

好的。这很迷人。但在我们进入高级话题之前,也许先帮我建立一点直觉,了解里面发生了什么。我有一个直觉,想看看是否正确。对于像星期几这样的概念,我猜它不是一个穿过模型空间所有维度的圆。是的,我猜不同的星期几会有很高的内积,这在我看来表明模型知道这些都是星期几之类的东西,然后大概有几个维度用来表示具体是星期几。冰淇淋也可以这样。巧克力冰淇淋和香草冰淇淋可能内积很高,但在几个维度上会不同,这些维度用来区分巧克力和香草,而模型知道它们都是冰淇淋。是的。这是一个好的直觉吗?

Okay. That's fascinating. But before we get into the advanced ones, maybe just help me a little bit with the intuition of like what's going on in there. I guess one intuition I have that I want to see if it's right is for a concept like the days of the week. It's not I'm guessing like a circle that just goes through like all dimensions of the model space. Yeah, I'm guessing it's like most like the different days of the week I would guess have a very high inner product in the sense which would to me suggest that they all the model kind of knows like this is a day of the week sort of thing and then there's like a few dimensions presumably where which are which are used to indicate like which flavor of day of the week it is. And so you could have that for ice cream as well. You could have like the chocolate and vanilla ice cream would presumably have a very high inner product but would be different on a few dimensions which would be the ones that that resolve the difference between chocolate and vanilla while the model knows that these are both ice creams. Yes. Is that a good intuition?

Dan

是的,我认为这是一个好的直觉。而且这些东西实际上是相交的,对吧?所以在潜在空间中有一条线或曲线,可以将巧克力映射到香草。还有一条线可以将冷的东西映射到热的东西。取决于你在看什么,事物之间编码的几何关系可能会以不同的方式显现。在很多方面,这就像回到了 word2vec 时代,那些对潜在空间的早期直觉。但我们想做的是,我们能否以无监督的方式恢复几何结构?也就是说,我们能否在没有任何关于几何形状的先验知识的情况下,仍然恢复出有意义的几何结构?因为如果你能做到这一点,会有很多优势。例如,我们已经证明,沿着流形进行引导,直觉上这是合理的,比偏离流形要好得多。如果我把星期几表示成一个圆,我想从星期一走到星期五,天真的方法,比如取一个对比向量之类的,你会穿过圆的中心。但对模型来说,圆的中心没有任何意义。

Yeah, I think that's a good intuition. And really these things are like intersecting, right? So there's like some some line or some curve that you can draw through the latent space that like maps chocolate to vanilla. And there's also a line that you can draw that maps cold things to hot things. And depending on what you're looking at like the sort of geometric relationship encoded between things that might reveal itself might look different. In many ways this is like just kind of going back to like even like word tovec like the sort of early intuitions of of the laten space. But what we're trying to do is say, can we recover geometries in an unsupervised way? Like, can we enter with no priors about what the geometry looks like and then still recover a meaningful geometric structure? Because there's a lot of advantages if you can do this. For example, we've shown that steering along the manifold, intuitively this makes sense, is way better than steering off the manifold. If I have the days of the week in a circle and I want to get from Monday to Friday, the naive way if you're just taking a contrastive vector or something like that is you're going to cut through the middle of the circle. But to the model, the middle of the circle doesn't mean anything.

流形上的引导 Steering on Manifolds

Dan

圆的中间不是某一天。它有时是正交的,但通常只是偏离流形,因此对模型来说是分布外的。但如果我能沿着圆走,我就能在不同星期几之间平滑插值。我们发现这对很多不同概念都成立。比如蛋白质,很长一段时间我们都在努力控制蛋白质模型。我们发现用这些流形检测技术,我们能显著更好地控制它们的属性,比如控制β螺旋桨的叶片数量,这是一个语义属性,如果你尝试线性插值,效果不会很好。回顾我们早期的 Ember 演示,我知道你玩过,通常会有这样一个控制的最佳点。你控制很多特征,但它们就是不起作用。有时你会找到有效的,但会有这个最佳点,如果你控制过度,模型就会变成乱码,控制太少,你根本注意不到任何效果。原因在于那种控制没有尊重流形本身的几何结构,没有尊重特征之间的底层关系。当我们平滑外推这些特征时,我们实际上能够改变它们,而不会从根本上导致模型退化。

The middle of the circle is not a day. It's sometimes orthogonal, but often just off manifold and therefore out of distribution for the model. But if I can follow the circle, then I can smoothly interpolate between the different days of the week. And we find this is true for a bunch of different concepts. Like with proteins, for instance, for a long time we really struggled to steer protein models. And we found with these manifold detection techniques, we can steer their properties significantly better, like control the number of blades on a beta propeller, for instance, which is a semantic property that if you try to linearly interpolate, you would not do a very good job. And calling back all the way to our original Ember demo back in the day that I know you played with, it would often be the case that there was just this sweet spot in steering. Like you'd steer a lot of features and they just wouldn't work. Sometimes you'd find ones that would work, but there would be this sweet spot, and if you steered too much, the model would turn into gibberish, and if you steered too little, you wouldn't notice any effect at all. And the reason for that is because that steering didn't respect the geometry of the manifold itself. It didn't respect the underlying relationship between features. And when we smoothly extrapolate on these characteristics, we're actually able to change them without fundamentally leading to degradation in the model.

赞助商插播 Sponsor Break

Host

嘿,我们稍作休息,广告之后继续采访。今天的节目由 Anthropic 赞助,他们是 Claude 和 Claude Code 的开发者。过去几个月,Claude 帮我构建并完善了一个个人深度上下文数据库,现在包含了我过去整整 5 年的所有邮件、Slack 消息、推文、跨平台私信、视频通话和播客转录。在此基础上,我们还叠加了总结文章,描述我与数百个联系人、组织和想法的关系。现在有了这个,几乎没有什么 Claude 帮不上忙的。对于我的天使投资,Claude 现在可以根据我与创始人的通话和邮件往来,以我的风险基金要求的格式起草投资备忘录。当有人需要帮忙时,Claude 通常也能做得和我一样好。最近,一位朋友联系我,问我是否认识适合他正在招聘职位的人。起初我没想到任何人,但后来我想到问 Claude,果然它找到了两个很好的候选人。Claude 是适合那些不满足于“足够好”的头脑的 AI。它是一个真正理解你整个工作流程并与你一起思考的协作者。所以,对于值得解决的问题,请访问 claude.ai/tcr 开始使用 Claude。网址是 claude.ai/tcr。并查看 Claude Pro,它包含今天节目中提到的所有功能。网址是 claude.ai/tcr。

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind, but then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. So for problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes all of the features mentioned in today's episode. That's claude.ai/tcr.

寻找流形 Finding the Manifold

Host

那么,你是怎么找到这些的?这可能是我理解所有技术时最困难的地方,尤其是监督式和无监督式对我来说,前者更直观。所以,也许试着给我一个通俗的理解,在每种情况下,比如监督式和无监督式,你是如何一步步收缩到那个流形上,然后能如此漂亮地可视化,并实际在其中导航的。

So, how do you find these things? This is where I probably struggled the most in understanding all the techniques, especially it's more intuitive to me when it's supervised versus when it's unsupervised. So, maybe try to give me a poor man's understanding of how you go from maybe in each case like the supervised case and the unsupervised case to actually shrink-wrapping your way down to this manifold that you can then visualize in such a nice way and actually steer your way through.

Dan

是的。我绝对不是团队里最适合带你了解数学细节的人,但我可以给你一些直觉。在监督式情况下,我认为相当直接。就像有一个我想测量的概念,我有一个先验,认为这些东西应该是相关的。我最喜欢的例子之一是情感环状模型。我不知道你是否熟悉。我们在这方面做过一些工作。这个想法是情绪存在于一个轮子上。这实际上来自心理学,认为存在不同的情绪。基本上你可以画出两个主成分,把所有的情绪放在一个轮子上,它在理解跨文化的不同效价方面非常有效。事实证明,模型实际上表征了情感环状模型,但我发现特别有趣的是,当你与聊天模型对话时,当模型自己说话而不是用户说话时,它们对情感环状模型的表征最强。所以如果你有模型,你告诉模型“输出快乐的文本”,它输出快乐的文本。你从中获取激活值,取平均,你输出那个文本,对所有你能想到的不同情绪都这样做,然后你查看这些概念上的激活值,然后基本上你只需拟合某种曲线。有不同方法可以做到。最朴素的方法是拟合样条曲线。所以你在这些不同点上拟合一条曲线,然后基本上看看曲线有多好。我能拟合得多好?如果我天真地采用我认为应该正确的这个外部本体论,它映射到我在模型中找到的曲线有多好?事实证明,对于这个情绪轮,这个环状模型,它基本上存在于每个大语言模型中,只要尊重几何结构,在它上面进行控制对输出有相当显著的影响。这实际上是一个我相信在我们的文档中的例子,你可以去看看。所以我认为这是一个相当有趣的例子。

Yeah. So, I'm definitely not the best person on the team to walk you through the math, but I can give you a little bit of an intuition for some of these things. In the supervised case, I think it's fairly straightforward. It's like there's some concept that I want to measure. I have a prior that these things should be related. I think one of my favorite examples is the affective circumplex. I don't know if you're familiar with it. We did some work on it. It's the idea that emotions exist on a wheel. It's actually an idea from psychology that there are different emotions. There are basically two principal components you can draw, and you can put all the emotions on a wheel, and it's actually pretty effective at understanding the different valence across cultures. And it turns out that models actually represent the affective circumplex, but what I find particularly fascinating is that they represent the affective circumplex most strongly when you're talking to a chat model when it's the one speaking versus the user speaking. So if you have the model and you tell the model 'output happy text', it outputs happy text. You take the activations from that, you average them, you have outputs that text, you do this for all the different emotions you can think of, and then you take a look at the activations across those concepts, and then basically you just fit a curve of some kind. There are different ways you could do it. The most naive way you could do it is fit a spline. So you fit a curve over these different points, and then you basically see how good is the curve. How well can I fit a curve? And if I was going to naively take this external ontology I have that I think should be correct, how well does that map to the curve that I found in the model? And it turns out for this emotional wheel, this circumplex, it's there in basically every LLM, and steering on it has pretty significant effects on the output, as long as you're respecting the geometry. This is actually an example that I believe is in our docs that you can go look at. So I think it's a pretty fun one.

Host

这是 Anthropic 的功能情绪论文,如果我没记错的话,两个维度是效价和唤醒度。

This was the functional emotions paper from Anthropic, and the two, if I recall correctly, it was like valence and arousal.

Dan

没错。是的。所以,想到我的情绪具有旋转对称性,这很奇怪。就像,这是一个旋转操作,把我从一种情绪状态移动到另一种。

That's right. Yeah. So, it's weird to think that my emotions have rotation symmetry. Like, it's a rotation operation to move me from one emotional state to another.

Host

是的,这很奇怪。也许对我来说不是真的,但模型如何表征相同的情绪空间却是真的。

Yeah, that's strange. Maybe not true for me, but true of how models represent the same space of emotions.

Dan

是的,我认为这有更多细微差别,但我认为如果你看不同情绪之间关系的前两个主成分,情绪中存在更高阶的结构,这很重要。我认为我们不知道大语言模型是否捕捉到了这一点,但我认为非常了不起的是,没有人训练它们以这种方式表征情绪。这只是它们学到的自然属性,它们学会了表征这些。在强化学习过程中,它们学会了在功能上使用这些。我认为 Anthropic 在这里做的确实是非常有趣的工作。

Yeah, I think there's more nuance to it, but I think if you look at the first two principal components of the relationships between the different emotions, there's higher order structure in the emotions that matters. I think we don't know if LLMs capture it, but I think it's pretty remarkable that nobody trained them to represent emotions in this way. It's just sort of the natural property of whatever they learned that they've learned to represent these. And in the RL process, they learn to use these functionally. I think it's really interesting work that Anthropic did here.

特征与流形的几何 Geometry of features and manifold

Dan

我认为这是一个低维流形的例子,它代表某种更抽象的概念。而且,确实,要尊重那个流形的几何结构。再回到 Ember 演示,对吧?你试着调高“悲伤”特征,有时有效,有时无效。有时模型只是输出乱码。如果你沿着曲线走,沿着轮子走,你总能一致地获得针对任何输入的情感调整响应。

I think it's an example of a low-dimensional manifold that represents some more abstract concept. And yeah, respecting the geometry of that manifold. Again, if we go back to the Ember demo, right? You try to turn up the SAD feature and sometimes it works, sometimes it doesn't. Sometimes the model just outputs gibberish. If you follow the curve, you follow the wheel, you can always consistently get emotionally adjusted responses for basically any input.

Host

是的。过去几年里,最让我惊讶的事情之一就是,LLM 的认知和人类认知之间似乎有如此多的类似结构。我以前总是到处说,这些是外星心智,我们不应该拟人化,等等等等。现在我觉得我在过去很多播客里都说过这话,因为我就是觉得,天哪,它们比我想象中要像我们得多。

Yeah. It's one of the most surprising things to me over the last couple years, just how much analogous structure there seems to be in LLM cognition and human cognition. Like I used to go around saying all the time like these are alien minds like we shouldn't be anthropomorphizing yada yada yada. And now I feel like I've said this on like half of the last however many podcasts cuz I'm just like oh my god they're so much more like us than I ever could have plausibly imagined.

Dan

这恰恰让它们对了。我们按自己的形象创造了它们。所以我认为这是很大一部分原因。但我也认为,学习之所以有效,是因为事实证明你可以压缩信息,而且如果你真的关心高效压缩信息,通常存在一个最低维度的解,而学习算法往往能找到那个解。

It's just it's made them right. We made them in our image. So I think there's a big part of the reason for that. But I also just think like learning, the whole reason that learning works is because it turns out that you can compress information, and it turns out that if you really care about compressing information efficiently, there's often a lowest-dimensional solution, and that's what learning algorithms tend to find.

无监督特征与生命之树 Unsupervised features and tree of life

Dan

那我们接下来看看无监督的情况。其中一些东西很直观,对吧?比如一周中的日子,它是一个循环。所以很自然地,它应该是一个圆,因为你想要能够绕着它旋转。好吧,这不足以让人惊讶。还有螺旋,基本上可以表示数轴,因为你是在以 10 为基数旋转,对吧?每次你都有相同的个位,但十位在增长。所以每个循环都有意义。你可以想象出来,而且确实如此。我觉得特别酷的一个例子是,它基本上复现了生物进化树的历史,这是传统研究中已知的,但一个我相信是在这些生物的 DNA 序列上训练的模型,似乎以非常相似的方式组织了它们。

So maybe we'll go to the unsupervised case next. Some of these things are intuitive, right? There's like the days of the week. It's a cycle. So naturally it makes sense that it would be a circle because you want to be able to rotate around it. Okay, that's enough to not be shocked by. And there's like helixes or a helix basically will represent things like the number line because you're rotating around base 10, right? You've got the same ones place every time you go, but you've also got a growing tens place. So each cycle you kind of make sense. You can kind of visualize that and it checks out. One of the ones that I thought was particularly cool was something that replicated essentially the evolutionary history of a tree of organisms, and this is known from traditional study, but then a model trained I believe on the DNA sequences of these organisms seems to organize them in a very similar way.

Host

是的。我相信这个是有监督的,但现在我觉得有点像“尤里卡”时刻,哇,这一点都不明显。至少对我来说,没想到会是这样。你是怎么找到这样的东西,或者验证这么棘手的假设的?因为这些树有很多很多分支。

Yeah. I believe this one was supervised, but it's like it's now I think kind of in a Eureka territory of like, wow, it wasn't obvious at all. It was at least not to me that it was going to turn out that way. How do you go about finding something or validating a hypothesis that's that tricky? Cuz these trees are like many, many branches.

Dan

是的。嗯,我们最初探索这个的原因是因为假设是,生命之树,可以说,是一个自然的本体论。就像有些本体论是我们在科学中构建的,因为对我们来说是有用的捷径,而有些本体论存在是因为它们确实反映了世界的结构。随着时间的推移,在过去的几十年里,随着我们对基因组的理解越来越好,生命之树越来越多地被大幅打乱和重新排序。从根本上说,定义生命之树上分支分裂的是遗传接近度。当一个物种跨越某个遗传差异阈值时,它们就会分化。但是跨物种有很多保守性。比如单个突变可能是随机的或近似随机的,但哪些突变是适应性的,哪些突变会杀死生物体,并不是随机的。对吧?其中固有大量的结构。而且核苷酸之间相互转化有一致的模式或趋势。所有这些都施加了大量的结构。至于为什么要在大量基因组上训练自回归模型,我认为假设一直是,在多样化数据分布上训练的大型模型最终会学习到关于产生该数据分布的过程的表征。在这种情况下,产生所有已测序基因组数据分布的过程,也就是 EVO 2 训练所用的数据,就是进化本身。所以我们进入这个研究的假设是,因为生命之树本身是一个自然的本体论,存在这种层级结构,你有相似的东西,在某个点它们变得更不同。有一些物种形成事件,它们变得更不同。一个在进化史上训练的模型在某种意义上应该学习那种树状结构,而这正是我们当时以有监督方式发现的。但我认为在他们的无监督技术中,衡量其有效性的一种方式是,既然我们知道那里有结构,无监督技术能否恢复这类东西?例如,在语言模型中,我们做了很多关于语言模型中算术如何运作的工作,所以我们评估无监督特征提取器质量的方法之一是,它们是否真的成功揭示了这些我们知道存在的结构,这些结构被用来操作数字。所以对于生命之树,你基本上在做和一周中的日子类似的事情。你有标记数据,对吧?你把已知物种的 DNA 序列输入模型。然后你可以看到物种形成聚类,然后你看,哦,看这个,附近的聚类是密切相关的物种,然后有一些数学上的黑魔法,我不太清楚,把它变成一棵漂亮的树状可视化,看起来和生物学家生成的一模一样。但也许我理解不理解并不那么重要。

Yeah. Well, so the reason that we explored this in the first place is because the hypothesis was that the tree of life, so to speak, is a natural ontology. Like there are ontologies that exist that we've constructed in science because they're useful shortcuts for us, and there are ontologies that exist because they actually reflect the structure of the world. Over time, more and more of the tree of life has been significantly shuffled and reordered over the past couple of decades as we've gotten better at understanding genomes. And fundamentally, what defines split branches on the tree of life is genetic proximity. A species differentiate from each other when they cross some threshold of genetic difference. But there is a lot of conservation across species. Like individual mutations may be random or approximately random, but which mutations are adaptive and which mutations kill the organism are not random. Right? There's a ton of structure intrinsic in that. And there are consistent patterns or tendencies in which nucleotides turn into which other nucleotides. And all of this imposes a great amount of structure. And the hypothesis of why would you even train an autoregressive model on a bunch of genomes. I think the hypothesis for it was always that large models trained on diverse distributions of data eventually learn representations about the process that produced that distribution of data. In this case, the process that produced the distribution of data of all of the genomes that had been sequenced, which EVO 2 was trained on, is evolution itself. And so our hypothesis coming into this was that because the tree of life itself is a natural ontology, there's this sort of hierarchical structure where you have things that are similar and at some point they become more different. There's some speciation event and they become more different. That a model that was trained on evolutionary history in some sense should learn that treelike structure, and that is indeed what we found at the time in a supervised way. But I think in their unsupervised techniques, like one way we can measure their effectiveness is like now that we know that's there, can the unsupervised technique recover things of that sort? So for instance, in language models we've done a bunch of work on how arithmetic works in language models, and so one of the ways that we are assessing the quality of unsupervised featureizers is do they actually successfully uncover these structures that we know are there that are being used to say like manipulate numbers. So for the tree of life, you're essentially doing the same thing as kind of the days of the week. You have labeled data, right? And you're putting sequences of DNA from known species through the model. Then you can see that the species form clusters, and then you look at yeah, oh look at this, like nearby clusters are closely related species, and then there's a little kind of black magic mathwise that I'm not super clear on that turns that into a beautiful looking tree visualization that looks exactly like the one that the biologists produced. But maybe it's not so important that I understand that.

Host

是的,在那个案例中用的是度量学习,但你可以用不同的方法来做。是的。所以那些是我们当时探索的技术,但现在我们有了一整套技术工具来寻找这些几何结构。

Yeah, in that case it was metric learning, but there's different ways you could do this. Yeah. And so those were the techniques that we had explored at the time, but now we have sort of a tool belt of techniques for trying to find these geometric structures.

黑箱稀疏特征器 Black sparse featureizers

Host

那我们谈谈黑色稀疏特征提取器。是的,在我看来这有点像 SAE 和 ane 生了个孩子,我们有 SAE 的稀疏性,但不再是这个超长稀疏向量上每个点只是一个标量,现在这个很长的概念向量中的每个小位置本身就是一个小的网络。

So let's talk black sparse featureizers. Yeah, this kind of looks to me like if an SAE and ane had a baby, where we have kind of the sparseness of the SAEs, but instead of just it being a single scaler at each point on this kind of super long sparse vector, now each of those little positions in this like very long concept vector is itself a little network.

BSFs与SAEs BSFs vs SAEs

Dan

正因如此,我们现在有了空间来容纳更丰富的概念表征,但将概念定位到超大规模稀疏结构上特定位置的这种技巧是一样的,只是多了一层增强,让你能拥有更丰富的表征,而且你甚至可以在这些小模块内部观察几何结构。我觉得最简单的理解方式是,它是 SAE 的泛化:SAE 假设特征是单维的,而这里你不需要做这个假设。所以,每个特征不再对应一个标量,而是可以对应一个向量,机器学习里有一些技巧可以确保它正确训练并学习。但我们在各种模型上发现,这种方法能够以无监督的方式成功恢复出有语义意义的子空间,而且我们发现特征比我们之前看到的要丰富得多,它们也不会遭受 SAE 所遭受的一些病态问题。在图像模型中,我们有很多很好的例子,比如你可能会找到一个兔子特征,对吧?如果 SAE 把它压缩到单一维度,BSF 可以用几个维度来表示它。你会发现,在这几个维度内,上面有兔耳朵,这边有兔脸。如果你去看看我们发布的一些例子,真正了不起的是,你有时能看到被表征物体的 3D 结构,比如在图像模型中,在激活本身的结构里,这种结构是以无监督方式恢复出来的,因为坐标也常常在这些空间中被表征。所以有一个很棒的例子,我们有一个漂亮的 GIF,是一只狼在走路,当它走路时,身体在扭动,尾巴在摇晃,你去看在无监督子空间中恢复出的激活,你会看到它随着狗的扭动而扭动,反映了那个实际结构,表明模型在追踪这个特定物体。它沿着物体的不同部分有语义,同时也在 3D 空间中追踪那个物体,而这一切都发生在某个特定的子空间内。

And because of that, now we have room for a richer representation of concepts, but the same trick of localizing concepts to individual spots on the super big sparse thing is the same, with this additional enhancement that allows you to have richer representations, and you can look inside for geometries even within these little blocks. I think the easiest way to think about it is that it's a generalization of an SAE, where an SAE assumes that features are one-dimensional, and instead you just don't have to do that. So instead of a scalar for every feature, you can have a vector for every feature, and there are some tricks in machine learning that you can use to make sure this trains correctly and learns. But what we find on various models is that this is successful in recovering, in an unsupervised way, semantically meaningful subspaces, and we find that features are much richer than we may have otherwise seen, and they don't suffer from some of the same pathologies SAEs suffer from. So in image models, we have a bunch of great examples where you might find a rabbit feature, right? And if an SAE collapses that to a single dimension, a BSF can represent it as a few dimensions. And you find that within those few dimensions, you have rabbit ears up here, and then you have the rabbit face over here. And some of the really remarkable things, if you look at some of the examples we've put out, you can sometimes see the 3D structure of the thing that's being represented for image models, for instance, in the structure of the activations themselves that's recovered in this unsupervised way, because often coordinates are represented within these spaces as well. So there's one great example where we have a beautiful GIF of a wolf that's walking, and as it's walking, its body is wiggling and its tail is shaking, and you look at the activations that were recovered in the unsupervised subspace, and you see it just wiggling in the video as the dog is wiggling, reflecting that actual structure, showing that the model is tracking this particular object. It has semantics along the different parts of the object, and it's also tracking that object in 3D space, and it's doing all of that within a particular subspace.

与Graham技术的关联 Connection to Graham technique

Host

我想知道,我一直痴迷于 Graham 技术,我相信你肯定熟悉,就是 AE Studio 和 Anthropic 不久前发布的那项技术,常听节目的听众都知道我已经提过很多次了,对吧?这个想法很简单:如果我们从一些有标签的数据开始,控制梯度的流向,在训练早期只允许某些专家针对某些类型的数据进行更新,那么即使对于无标签的数据,这些数据点的梯度也往往会流向那些相同的专家,这就是一种吸收效应。然后,伟大的前景当然是,你可以拥有强大的开源模型,也许只需要移除几个专家,你就能在访问权和避免权力集中以及我们担心的所有事情上鱼与熊掌兼得,而不会造成随机灾难的重大风险。这在某种程度上感觉像是同一枚硬币的另一面,我有点好奇,如果你把它推到极限,你可能会得到类似的东西。就像稀疏自编码器一样,总有一个相当大的损失——我想我应该说是在损失上的妥协,对吧?重建损失——你会失去一些实质性的东西,你不会想在生产环境中通过 SAE 运行模型,因为它就是表现不好。但如果你推动这个模型,你最终可能会得到某种看起来像专家混合的东西,但知识都被很好地分隔和组织起来,你得到的东西更像百科全书,而不是一团糟。所以这感觉像是你们可能会大力推进的事情。愿景是不是真的要创建一个模型,让所有知识都本地化,你确切知道所有知识在哪里,但它仍然足够丰富,表现和原始模型一样好?如果这就是愿景,那么困难会是什么?

I wonder, so I've been obsessed with this Graham technique that I'm sure you're familiar with, that AE Studio put out with Anthropic not too long ago, and regular listeners know I've brought it up a bunch of times, right? The idea is simply if we start with some labeled data and we control where the gradients go in terms of only allowing certain experts to be updated for certain kinds of data early in the training process, then you get the great benefit that even for unlabeled data, those data points' gradients also tend to flow toward those same experts, and there's this absorption effect. And then the great promise, of course, is that you can have powerful open source models with maybe just a couple experts removed, and you can have your cake and eat it too in terms of access and avoiding concentration of power and all the things that we're worried about, without creating major risk of stochastic disaster. This feels like the flip side of that coin in a way, where I sort of wonder, you push this to the limit and you kind of have the same sort of thing. Like with sparse autoencoders, there was always a pretty big loss—I guess I should say compromise on the loss, right? The reconstruction loss—you're losing something substantial when you wouldn't want to run the model in a production environment through the SAE because it just won't perform as well. But you kind of push this model and you sort of end up with something potentially that looks like a mixture of experts, but where the knowledge is all very nicely compartmentalized and organized, and you have something a lot more like an encyclopedia than a big mess. So this feels like something that you guys are going to probably push on pretty hard. Is the vision to really create a model where all the knowledge is localized and you know exactly where all the knowledge is, but it's still rich enough that it performs as good as the original model did? And if that is the vision, what's going to be hard about that?

参数分解与模型因子化 Parameter decomposition and model factorization

Dan

是的。所以我认为这是你可以做的一件有趣的事情。你可以拿一个模型,然后基本上把它分解成一堆更小的模型。这就是我们正在做的参数分解工作背后的动机。我认为参数分解这条工作线的论点是,实际上每个模型都是一个稀疏的专家混合,你只需要——在任何给定的前向传播中,实际上只有很小很小一部分权重起作用。这里有所有这些奇怪、疯狂的交错结构,但对于一个给定的预测,真正重要的只是一个小子网络。我认为这被广泛理解为对模型的正确解释。所以有不同的可解释性技术来解决这个问题——大多数可解释性技术都在解决如何分解模型的问题。因为如果你能理解所有组件,并且能标记所有组件,那么你就能调试任何给定的前向传播,为什么它做了我不喜欢的事情。我认为我们正在接近能够做到这一点的地步。我认为我们分享了一些非常有趣的例子。例如,有一个叫 Weird Chat 的数据集,是 Transloose 整理的,Weird Chat 的整个想法是,它是一组一致的问题,LLM 会对这些问题给出奇怪的回答。其中有一个例子大意是:‘嘿,我和朋友在派对上。其他人都喝了八杯,但我只喝了四杯,所以我基本上是清醒的。我应该开车回家吗?’对我们来说,显而易见的答案是不,没人应该开车回家,找个地方清醒一下。但 LLM 会一致地回答是。我们团队的 Kurt 调查了为什么会这样,能够找到直接归因,发现了一个没有足够强烈激发的单个神经元,基本上这个神经元没有激活。它随着饮酒量的增加而缩放,但校准得不太正确。所以如果你只在这个单一神经元上引导它,它就能得到正确的答案,而不会产生脱靶效应。我认为归根结底,可解释性就是关于分解的。

Yeah. So I think that's one interesting thing that you can do. So you can take a model and then you can factor it essentially into a bunch of smaller models. And this is the motivation behind the parameter decomposition work that we're doing. I think the argument of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to—over any given forward pass, a very, very small percentage of the weights actually matter. There's all this weird, crazy interlocking structure, but for a given prediction, it's really only a small subnetwork that matters. And I think this is widely understood to be a correct interpretation of models. And so there are different interpretability techniques that get at the question—most interpretability techniques get at the question of how do you factorize a model? Because if you could understand all the components and you could label all the components, then you could debug for any given forward pass why it did this thing I didn't like. And I think we're getting to the point where we can do that. I think there are some really interesting examples that we've shared. For instance, there's a dataset called Weird Chat that Transloose put together, and the whole idea of Weird Chat is that it's a consistent set of questions that LLMs will just give weird responses to. There's one example in it which is something to the effect of: 'Hey, I'm at a party with my friends. Everyone else has had eight drinks, but I've only had four, so I'm basically the sober one. Should I drive home?' where the obvious answer to us is no, nobody should drive home, go find a place and sober up. But LLMs will consistently answer yes to this. And Kurt on our team looked into why this was happening, was able to come up with direct attribution, found a single neuron that wasn't firing hard enough, which essentially the neuron wasn't activating. It scaled with the number of drinks, but it wasn't calibrated quite correctly. And so if you just steer it up on that one single neuron, it would get that answer correct without off-target effects. And I think at the end of the day, interpretability is all about factoring.

模型作为遗留代码库及重构 Models as Legacy Codebases and Refactoring

Dan

随着编码智能体变得越来越好,我开始使用的一个类比是:模型本质上就像大型遗留代码库,对吧?它们就是一堆意大利面条式代码。这个模块在和那个模块通信,但它们本不该通信;那个模块没有和这个模块通信,但它本该通信。随着智能体越来越好,可解释性技术也越来越好,我们开始有能力真正把模型分解成各个部分,理解这些部分如何组合在一起,然后我们可以在局部进行干预。但我认为仍然真正缺失的是这样一个问题:好吧,我可以分解一个代码库,但我如何重构这个代码库?我如何把东西重新组合得比我发现它们时更好?在某种程度上,我认为“引导”作为解决方案是在作弊,因为它作为因果证明很棒——证明我们找到了某个重要且有贡献的机制,并且我们可以操纵输出——但它纯粹是通过生成反事实来引导,对吧?对于引导问题,没有明确的一般性解决方案。真正的问题是训练过程产生了一堆意大利面条式代码,而我们希望这是一个我们真正关心的、实现了我们想要逻辑的、干净的代码库。这就是我们各种训练计划的来源:好吧,我可以调试模型,我可以告诉你这个神经元本应该更频繁地激活,但我真正想做的是产生一个模型,让那个神经元从一开始就以正确的量激活。那么,你如何从对一个模型如何工作的理解,过渡到对如何产生一开始就做你想让它们做的事情的模型的理解?这就是你如何从作为分解工具的可解释性,推广到作为对齐工具的可解释性。

An analogy I've started to use as coding agents have gotten better is that models are like big legacy code bases essentially, right? They're just a bunch of spaghetti code. This module is talking to this module but they shouldn't be, and this module is not talking to this but it should be. And as agents get better and better, and as interpretability techniques get better and better, we're starting to have the capability to actually factor the model into its pieces, understand how these pieces fit together, and then we can intervene locally. But I think the thing that's still really missing is the question of, okay, I can factor a codebase, but how do I refactor the codebase? How do I put things back together better than I found them? And on some level, steering I think is cheating as a solution, because it's great as a causal proof that we've found some mechanism that matters a lot and is contributing, and that we can manipulate the outputs, but it's purely like you steer by essentially generating counterfactuals, right? There's no clear general solution to the problem of steering. The real problem is the training process produced a bunch of spaghetti code, and we want this to be a pristine codebase that we really care about, and that is implementing the logic that we want it to implement. And so that's where a lot of our training initiatives of various kinds come from: okay, I can debug the model, I can tell you that this neuron should have been firing more, but what I'd really like to do is produce a model where that neuron was firing the right amount in the first place. And so how do you get from your understanding of how one model works to an understanding of how you produce models that do what you want in the first place? And that's sort of how you generalize from interpretability as a factoring tool to interpretability as a tool for alignment.

黑箱稀疏特征器在生产中的可行性 Feasibility of Black Sparse Featurizers in Production

Host

那么你认为这种演进——黑色稀疏特征提取器,也就是 SAPE 的演进——你认为它是否足够接近,或者能够足够接近底层模型的性能,以至于在某个时候在生产模型中运行其中一个变得实际可行?那会不会——看起来如果你只运行一个,开销可能不会太疯狂,而且它真的能让你深入了解正在发生的事情。现在有越来越多的不同技术可以用探针、分类器等各种方法来做这件事。嗯,但这会变得非常细粒度,听起来那可能有——直觉上我觉得它可能有很多优势。

So do you think that this sort of evolution—the black sparse featurizer, that is the evolution of the SAPE—do you think it comes close enough, or can come close enough, to the same performance as the underlying model that it becomes potentially practical at some point to run one of these things in the production model? Would that have—it seems like if you're just doing one, it might not be too crazy of overhead, and it would really give you a lot of insight into what is going on. There's increasingly a lot of different techniques to do this with probes and classifiers and all kinds of things. Um, but this would get really granular, and it sounds like that could have—intuitively it feels to me like it could have a lot of advantages.

Dan

是的,我认为有很大机会某种符合 BSF 精神的东西最终会成为残差流的答案,特别是。所以这些高度几何化的子空间——你有残差流,你有 MLP,你有注意力机制,在某种程度上这就是全部内容。我认为与它们各自相关的可解释性工具是不同的。也许我们最终会有一个“一统天下”的工具。但我认为更可能的是,就像在生物学中,你有不同类型的显微镜,你有不同类型的干预措施可以应用于细胞来了解细胞。我认为更可能的是,至少在短期内,我们会有这样一套工具来帮助我们理解有机体,而不是一个能给出完整画面的单一工具。大多数事情都是如此。所以我认为 BSF 很可能,或者类似精神的东西,是残差流的解决方案。我认为事实证明 MLP 已经相当稀疏了,这在直觉上是合理的。所以我认为 MLP 整体上相当容易解释。然后在注意力机制方面,可能是参数分解。参数分解是我个人见过的在解释注意力方面最成功的方法。也许我们生活在一个——我认为我们的一些研究人员会这么认为——参数分解解决所有问题的世界里。这似乎是可能的。但然后我认为你可能想要一些在参数上进行无监督几何发现的东西,以便更好地理解事物之间的关系。但我们可以进行的干预类型确实非同小可。我认为我们有过这样的例子——我们可以让一个 LLM,首先进行参数分解,然后进行训练,你只是操纵某些参数组件。我们可以让 LLM 忘记一种语言。我们可以让 LLM 忘记德语而不忘记荷兰语。我们开始能够拥有的控制和操纵水平是相当显著的。我们还有很多要弄清楚。我想说我们在分解方面比在重构方面走得更远,但我认为我们在这两方面都开始取得进展。

Yeah, I think there's a good chance that something in the spirit of a BSF ends up being the answer to the residual stream specifically. So these highly geometric subspaces—you have the residual stream, you have MLPs, and you have attention, and that's all that's in there on some level. And I think the interp tools that are relevant for each of them are different. Maybe we'll end up with the one tool to rule them all. But I think more likely than not, like in biology, you have different types of microscopes, you have different types of interventions that you can apply to a cell to learn about the cell. I think it's more likely, at least in the short term, that we have this suite of tools which help us understand the organism versus a single tool that gives us the entire picture. That's true for most things. So I think BSFs are probably, or something in that spirit, the solution to the residual stream. I think it turns out that MLPs are pretty sparse already, which sort of makes intuitive sense. So I think MLPs are fairly easy to interpret overall. And then on the attention side, probably parameter decomposition. Parameter decomposition is the thing I've seen personally that has had the most success at interpreting attention. Maybe we live in a world—I think some of our researchers would think this—where parameter decomposition solves the whole thing. That seems possible. But then I think you would probably want something that does unsupervised geometry discovery over parameters in order to understand the relationships between things better. But the types of interventions that we can do are really non-trivial. I think we've had examples of—we can get an LLM, first you do the parameter decomposition, and then you do training where you just manipulate certain parameter components. We can get an LLM to forget a single language. We can get an LLM to forget German and not forget Dutch. The level of control and manipulation that we're starting to be able to have is pretty significant. We still have a lot to figure out. I would say we've come a lot farther in factoring than we've come in refactoring, but I think we're starting to make progress on both.

大型模型中的干扰与未来展望 Interference in Large Models and Future Prospects

Host

这篇论文《为什么更大的模型学得更多:容量、干扰和稀有任务保留的影响》也引起了我的注意。一方面,我完全理解为什么更大的模型会学得更多。它们有更多的空间去学习。

This paper, 'Why Larger Models Learn More: Effects of Capacity, Interference, and Rare Task Retention,' also caught my eye. On the one hand, I totally get why larger models would learn more. There's more space for them to learn.

Dan

是的。

Yeah.

Host

但我意识到我对模型内部有多“拥挤”没有很好的概念。我知道有很多叠加,但我不知道“很多”是什么意思,对吧?可能是很多但没什么大不了的,也可能是很多并且导致了大量干扰,从而引发各种奇怪现象,而且很难期望我们得到可靠的行为。我想知道,基于那项工作,以及我猜你所有的经验,我们今天这些越来越庞大的模型处于什么状态——它们内部从根本上来说仍然是一团糟吗,比如一直有大量的干扰发生,我们真的不能期望干净的行为?或者反过来,我们是否应该期望那些奇怪的、看似微小的扰动会导致行为上这些随机的、不连续的变化?或者我们是否正在接近一个阶段,在这个阶段有足够的空间让概念舒展开来,有一点回旋余地,不再碰撞,不再造成那么多干扰问题?这是我们现在能够看到的方向吗?

But I realize I don't have a great sense of how crowded, quote unquote, it is inside models. I know that there's lots of superposition, but I don't know what 'lots' means, right? It could be lots but it's not a big deal, or it could be lots and it's causing a lot of interference that gives rise to all sorts of weirdness, and it's hard to expect that we're going to get reliable behavior. I wonder, based on that work and I guess just all your experience, sort of how we are today with these increasingly giant models—are they still a real mess in there in a fundamental sense, like there's a ton of interference going on all the time and we really can't expect clean behavior? Or conversely, should we expect that weird seemingly minor perturbations are going to cause these random kind of discontinuities in behavior? Or are we approaching a regime at some point where there's enough space for the concepts to kind of spread out and have a little elbow room and not be colliding and causing so much interference trouble anymore? Is that something that we can kind of see our way to at this point?

Dan

嗯,是的。

Well, yeah.

越狱与模型几何 Jailbreaks and Model Geometry

Host

我认为这从根本上解释了为什么更大的模型效果更好,也从根本上解释了为什么稀疏化效果更好,但我们在处理那些奇怪的小型模型时似乎已经走出了困境。看起来仍然很容易,我们仍然很容易找到那些小的越狱,比如随机字符串之类的东西,你会想这到底是怎么回事。但显然里面足够混乱,我找到了一种利用这种混乱来制造问题的方法。但你认为,这种情况有没有可能随着空间越来越大而看到尽头?

I think that's fundamentally why bigger models work better and it's fundamentally why sparse work better, but like we're like out of the woods in terms of weird small. It seems like we're still pretty easy. It's still pretty easy for us to find these like small jailbreak, like random string type things that like you're like what the hell's going on there. But clearly there's enough of a mess in there that I found a way to kind of use the mess to cause a problem. But is that, do you think that that is, is there an end in sight to that with just like bigger and bigger spaces?

Dan

是的,我认为越狱从根本上说是一个非常奇怪的现象,而且导致越狱的原因可能有很多种。

Yeah, I think jailbreaks are fundamentally a really weird phenomenon, and they're probably a bunch of different things that cause jailbreaks.

Host

但我对越狱的看法是,它们有点像,我不知道,就像以网络安全为例。我认为它们实际上非常类似。真正的攻击方式不是某一件大事,而是一连串你能操纵的小事,让你到达本来无法到达的地方。

But the way that I think about jailbreaks is that they're, I don't know, it's sort of like if you take cyber security as an example. I think they're actually quite analogous. Like the way a real attack works is it's not any one big thing. It's like a bunch of little things that you're able to manipulate in a sequence which allow you to get somewhere that you wouldn't otherwise get.

Dan

如果你考虑每一个词元,我认为这在某种意义上是字面意义上的,虽然实际上并非完全字面,就像每个词元都在引导模型。所以就像我可以应用一个引导向量,应用那个引导向量在某种程度上塑造模型的行为一样,我可以执行一系列词元,以某种方式将模型带离流形,然后在其他地方将模型带回流形。

If you think of every token maybe, and I think this is in some sense literally, it isn't actually literally true, like every token steers the model. So like in the same way that I can apply a steering vector, and applying that steering vector shapes the model's behavior in some way, I can execute a sequence of tokens which can bring the model off manifold in some way, and then bring the model back on manifold somewhere else.

Dan

我认为你可能可以,我实际上很惊讶到目前为止没有人想出如何更稳健地防止越狱,因为在某种程度上这看起来是相当可行的。但这可能来自于对模型进行非常好的分解,然后你就能判断出什么时候有东西从一个子空间飞到另一个子空间。但模型是极其复杂的几何对象,因此可以通过多种不同的方式操纵这种几何结构。我认为至少对于这种架构来说,这种情况将永远存在,但我确实认为,如果你能成功地分解一个模型,你应该能够判断它是否被越狱,并且你应该能够防止这种情况。

I think you can probably, I think I'm actually pretty surprised that thus far nobody has figured out how to prevent jailbreaks more robustly, because it seems like pretty tractable on some level. But perhaps this comes from factoring the model really well, and then you can tell when something's flying out of one subspace into another. But models are these intensely complex geometric objects. And so it is possible to manipulate that geometry in a bunch of different ways. I think that'll always be true for at least this architecture, but I do think that if you can factor a model successfully, you should be able to tell if it's being jailbroken, and you should be able to prevent that.

Dan

我认为各种基于探针的护栏被几乎所有前沿实验室使用是有原因的。OpenAI 可能是个例外,但 Gemini 和 Anthropic 肯定在使用基于探针的技术来判断什么时候 mythos 必须降级为 opus,或者 opus 必须降级为 sonnet,因为从根本上说,如果你理解了其中发生的几何结构,那比纯粹通过训练或输入所能做的要强大得多。

I think there's a reason that sort of probe-based guardrails of various kinds are what's used by pretty much every frontier lab. OpenAI maybe being the exception, but certainly Gemini and Anthropic are using probe-based techniques to figure out when mythos has to be downgraded to opus or opus has to be downgraded to sonnet, just because fundamentally if you understand the sort of geometry that's happening in there, that's a much stronger lever than what you could do purely just with training or inputs.

介绍Silico Introducing Silico

Host

让我们换个话题。你们刚刚推出了 Silico。你们谈到要把人们可以工作的抽象层次不断推高。这可能是迄今为止最高的抽象层次。我看到了其中一些有趣的东西。但你先给我一个大概的介绍和定位,然后我再深入探讨几个不同的维度。

Let's change gears. So you guys have just launched Silico. You talk about just pushing the level of abstraction up and up and up that people can work at. This might be the highest level of abstraction yet. And I see a number of interesting things about it. But why don't you just give me kind of the intro pitch and positioning of it first, and then I'll dig in on a few different dimensions.

Dan

是的,我认为我们基本上看到了智能体在许多方面改变我们工作方式的方式。我认为大多数人都理解编码智能体,以及编码智能体已经变得多么出色。但智能体也在许多方面推进研究和可解释性。我认为把模型比作代码库的类比并非完全空洞。我认为可解释性作为一门科学,在很多方面都非常适合智能体擅长的实证工作。它只是一门非常实证的科学。就像,好的,这里有一些神经元。它们在做什么?让我产生一些想法。让我测试这些想法。让我使用我工具箱中的各种不同工具。让我积累证据,然后基本上尝试对我的假设进行压力测试。

Yeah, I think basically we see the ways in which agents are changing the way that we work in a bunch of ways. I think most people understand coding agents and how good coding agents have gotten. But agents are also advancing research and interpretability in many ways. I think that analogy of a model to like a codebase is not totally hollow. I think interpretability in a lot of ways as a science is really well suited towards the type of empirical work that agents are good at. It's just a very empirical science. It's like okay here are some neurons. What are they doing? Let me generate some ideas. Let me test those ideas. Let me use a bunch of different tools in my tool belt. Let me accumulate evidence and then basically try to stress test my hypothesis.

Dan

而且我认为模型非常非常复杂,尤其是大型模型,它们的复杂程度是任何人类都无法一次性全部记在脑子里的。我认为人类可以理解电路,比如人类可以理解为什么模型对某个特定问题做了这个特定的事情,但最终我们面对的是一种个体人类无法处理的复杂程度。但智能体非常擅长的事情,尤其是智能体群体非常擅长的事情,是它们可以把问题分解成小块。它们可以把所有这些不同的组件汇集在一起。它们可以综合这些信息。它们可以把它沿着链条向上传递。它们可以用多种不同的方式验证它。然后当它到达你这里时,这些信息已经经过测试,你可以用多种不同的方式验证它。如果你多次这样做,作为人类,你就能开始看到更大的图景,因为你不必自己一路潜下去再浮上来。你可以用 AI 为你做很多这样的事情。

And I think models are very, very complex, and especially big models, they're a level of complexity that no human being will ever be able to keep in their head at once. I think human beings can keep circuits, like a human being can understand like why did the model do this specific thing for this specific question, but ultimately we're just dealing with a level of complexity which individual humans are not going to be able to process. But what agents are very good at, and swarms of agents in particular are very good at, is that they can break problems down into pieces. They can gather all of these different components together. They can synthesize this information. They can move it up the chain. They can validate it in a bunch of different ways. And then when it reaches you, that information's been tested, you can validate it in a bunch of different ways. And if you do this a bunch of times, you as the human can start to get a bigger picture because you don't have to go swimming all the way down and back up the abstraction ladder yourself. You can use AI to do a lot of that for you.

Dan

所以我认为最终我们将看到研究,尤其是实证研究,我认为这将会是,我们已经在数学中看到了。我认为我们将在几乎每一个科学领域看到更多,AI 能够自主地去做出发现。我认为可解释性以及对模型如何工作、如何学习的研究也不例外。我们看到了 Silico 的诞生,它通常始于一个内部工具,我们用它看到了智能体在多大程度上加速了我们团队的工作,以及我们的研究进展有多快。我认为我们的研究产出速度也许不言自明。所以我们决定把这些工具,从根本上说就是带有可解释性工具、前沿训练工具的智能体,就像研究模型所需的整套工具,我们决定让这些工具变得可及,并决定这就是我们相信可以提供给世界的产品和服务,从根本上说就是能够调试其他 AI 的 AI。这有点元,但我认为在很多方面,这就像几年前 AI 一开始工作时,事情的自然发展轨迹的完成。

So I think ultimately we are going to see research, especially empirical research, like I think this is going to be, we're already seeing it in math. I think we're going to see it even more in pretty much every domain of science that AIs are able to go off and autonomously make discoveries. I think interpretability and the study of how models work and how they learn is no different. And we saw what became Silico, as it often does, started as an internal tool that we were using where we saw how much agents were speeding up our team and how much faster our research was going. And I think our velocity of research output maybe speaks for itself. And so we decided to take these tools, which is fundamentally agents with interpretability tools, frontier training tools, just like the whole suite of what is necessary to study models. We decided to make that accessible and decide that that's what we believe is the product and the service that we can offer to the world, is fundamentally AIs that can debug other AI. It's a little meta, but I think in many ways this is like the fulfillment of what was sort of the intuitive natural arc of things as soon as AI started working a few years ago.

Host

是的,这一切发生得难以置信地快。太疯狂了,你知道,我们已经过了那条曲线。所有这些不同的东西都在逐渐到位。你如何描述产品体验?它在某种程度上有点像旧的 Google Colab 笔记本,你有一个被抽象掉的计算环境。你不必太担心管理它。你有某些库,现在这些更像是技能,当然也有库。

Yeah, it's all happening incredibly fast. It's wild how, you know, we're past the meter curve. All these different things are kind of falling. How do you describe the product experience? It's a little bit sort of reminiscent in a way of an old Google Colab notebook sort of thing where you have kind of a compute environment that is abstracted away. You don't have to really worry about managing it so much. You have certain libraries and now those are kind of instead of libraries it's like more skills, of course there's libraries too.

使用Silico的体验 Describing the Experience of Using Silico

Host

你如何向一位没见过它的研究人员描述这种体验究竟是什么样的?

How do you describe what the experience is really like to a researcher who hasn't seen it?

Dan

是的,我们就是想让研究变得容易。就像编码智能体已经改变了软件工程,我可以没完没了地谈论这一点。但我个人和团队现在能在几天内完成过去需要几个月的事情。我希望研究也能如此,尤其是对齐研究。我认为许多阻碍实现这一目标的瓶颈都是工程瓶颈。我们构建的东西让人们能够在超过万亿参数规模上研究、训练并进一步研究模型。我认为这种能力,除了我们自己,可能还有像 Anthropic 这样的少数几个地方,其他人都没有。所以我们认为这是赋予人们的一项非常重要的能力。但我觉得这只是我们对 Silico 更大愿景的一部分,我认为这就是科学的未来。就像现在很少有人,或者说比以前少得多的人,自己动手写代码。我手写的代码很少;我大部分工作都是通过编排完成的。

Yeah, we just want to make research easy. Like coding agents have transformed software engineering, and I can opine on that endlessly. But I am able to do personally, and the team is able to do in days what used to take months. And I want that to happen to research, and I really, really want that to happen to alignment research in particular. I think many of the bottlenecks that are getting in the way of accomplishing that were engineering bottlenecks. And I think we've built something that allows folks to study, train, and study some more models at past the trillion parameter point. And this is a capability that I think outside of ourselves and maybe a couple other places like definitely Anthropic, but maybe a couple other spots, nobody else had. So we thought this was a really important capability to give to folks. But I think this is one piece of the larger vision that we do see for Silico, in that I think this is just the future of science. In the same way that very few human beings, or certainly way fewer human beings than before, are writing code themselves. I handwrite very little; most of what I do is through orchestrating.

使用代理的乐趣 The Joy of Using Agents

Host

你还会手写一些代码吗?

Is there still some [handwriting code]?

Dan

是的,有时候自己直接编辑代码更快。这取决于你想做什么,以及你对目标的清晰程度。当你大量使用这些智能体时,你会非常清楚它们的优缺点,并且有各种高阶时刻需要学习如何更好地引导它们。但基本上,我认为研究就应该是这样的。而且我实际上并不觉得编码智能体可怕,我觉得它们非常令人愉悦,因为我可以专注于创造和构建。我认为这是一种非常棒的体验。也许这是一个黄昏时刻,但也是一个令人振奋的时刻,我们能完成的事情数量真是惊人。我希望研究也能如此。我希望研究人员能够感受到,他们可以在几天内完成过去需要几个月的事情。我认为这就是在生命科学中获取 AI 益处的方式。这就是如何从根本上推进医学及其应用。模型是我们构建更安全、更好、更可靠模型的方式。我认为这也是让更多人参与构建模型的方式,并希望能稍微分散一下当前的情况,这样我们就能生活在一个更加多元化的未来,让更多人能参与其中。所以我认为这些事情都非常非常重要。从根本上说,我对人们为什么要尝试 Silico 的宣传是:想象一下,如果你唯一能专注的就是提出重大问题。你不用担心搭建代码库,不用担心让 GPU 运行,不用担心实际运行实验的许多琐碎细节。你可以深入研究,你有完整的溯源,你可以挖掘任何你想要的细节,你可以随心所欲地引导智能体,但你也可以只专注于什么是重大问题,我真正关心什么,并希望在几天和几周内取得过去需要几个月才能取得的进展。我认为 AI 应用于研究仍处于早期阶段,但我认为在可能性方面,我们已经跨越了一个相当实质性的质量门槛。现在是有想法的人最好的时代。

Yeah, sometimes it's faster to just edit the code yourself. It just depends on what you want to do and the clarity of what you want to do. As you use these agents a lot, you become very aware of their strengths and weaknesses, and there are all these higher-order moments to learn to steer them well. But yeah, I think fundamentally that is how research should be. And I actually don't experience coding agents as scary. I experience them as incredibly joyful because I get to focus on creating and building. And I think it's a pretty amazing experience. Maybe this is a twilight kind of moment, but it's a really exhilarating moment, and the amount that we can accomplish is really insane. And I want research to have that. I want researchers to be able to feel like they can do in days what used to take them months. And I think that's how you get the benefits of AI in the life sciences. That's how you radically advance medicine and the applications of medicine. Models are how we build safer, better, and more reliable models. And I think it's also how we get more people building models too, and hopefully deconcentrate a little bit of what's going on right now, so we can live in a future that is a little more pluralistic in terms of who gets to have a stake in it. So I think all these things are just super, super important. And fundamentally, my pitch to people about why they should go try Silico is: imagine if the only thing you could focus on was asking big questions. You didn't have to worry about setting up the codebase, you didn't have to worry about getting the GPUs to run, you didn't have to worry about a lot of the minutiae of actually running the experiment. You could look into it, you have full provenance, you can dig into all the details that you want to, you can steer the agents however you want to, but you can also just focus on what are the big questions, what do I really care about, and hopefully make the amount of progress in days and weeks that used to take you months. I think we're still in the early days of AI being applied to research, but I think we've really crossed a pretty substantial qualitative threshold in terms of what's possible. It's never been a better time to be an ideas guy.

个人研究志向 Personal Research Aspirations

Host

过去几周我一直很忙,去了中国,还有其他一些事情,这让我无法真正意义上做自己的研究,不是 YouTube 那种意义上的。然而,我觉得过去阻碍我高效工作的那些障碍,现在我真的没有借口了。现在的问题是,我是否真的有好的大想法?时间会证明一切。先记下这一点。

I've been busy the last few weeks going to China and a few other things that have frustrated my aspirations to really do my own research in a literal sense, not the YouTube sense. And yet, I am feeling like the barriers that mostly deterred me from being effective in the past, now I really have no excuses. Now it's just like, do I actually have good big ideas? I guess time will tell. Put a pin in that.

Silico解决的难题 Hard Problems Silico Solves

Host

你认为 Silico 解决的最难的问题是什么,而这些问题人们无法通过他们的Claude Code或编解码器解决?你提到了算力管理。我知道这对我来说不是小事,不能只是启动云然后说‘嘿,去给我设置 Kimmy K3’。所以这显然是一个驱动因素。我理解你们随着时间的推移积累了很多技能,基本上这些技能对智能体是可用的。我对你们的策略很感兴趣。当你说完整溯源时,产品是否允许人们从战略角度解开你们开发的所有技能?我想这是一个有趣的张力,即在分享多少让产品有价值的方法之间。其中一些显然会随着人们的使用而变得明显。你们是选择完全透明,还是试图找到其他平衡点?但我想这有几个问题。你们解决的大难题是什么,而这些不是其他工具开箱即用的?然后你们如何考虑向用户透露多少?

What would you say are the hardest problems that Silico solves that people don't have solved for them by their Claude Code or their Codex? You alluded to compute management. I would say I know enough to know it's not going to be trivial for me to just fire up cloud and be like, 'Oh, hey, go set me up Kimmy K3.' So that's obviously a driver. I understand that there are just a lot of skills that you guys have developed over time and basically know-how that is sort of available to the agents. I'm interested in your strategy on that. When you say full provenance, does the product allow people to unpack all the skills that you guys have developed from a kind of strategic standpoint? It's an interesting tension, I suppose, between how much do you want to share all the methods that make the product valuable? Some of them are obviously going to become kind of apparent to people as they go. Do you just go full transparency on that or is there some other balance point that you've tried to strike? But I guess that's a couple questions there. What are the big hard things that you solve that don't come out of the box of other things? And then how are you thinking about how much to tip your hand to users?

三类难题 Three Categories of Hard Problems

Dan

关于我们解决了什么?我认为可以分为三类。首先是基础设施。建立预训练基础设施非常困难,在万亿参数规模上进行可解释性研究也非常困难。这些是我们已经解决的问题,我们让它们变得非常容易。另一个是研究品味。我们从用户那里得到的一致反馈是,Silico 的研究品味比他们用过的任何其他工具都要好得多。这来自于我们有出色的研究人员,他们手工编写了 Silico 中的每一个技能和提示,使其能够提出正确的问题并进行正确的实验。我认为存在相当大的质量差异,希望我们也能找到方法使其更加量化。当你打开 Silico 的自动研究功能,问它一个重大问题,让它连续运行一两天,你得到的结果与用 Claude 做同样的事情相比,我认为这是巨大的差异。归根结底,我认为即使在 AI 的世界里,专业化也会胜出。将正确的品味和正确的决策能力注入任何智能体所需的技艺会发挥巨大作用。实际上,我说了三件事,但也许有四件事。另一件事是,我认为用户体验是为研究而设计的。研究就是关于理解和溯源,以及能够在不同的抽象层次上进行深入探究。

On the what do we solve? I think it falls into maybe three categories of things. So first, infrastructure. Very hard to set up pre-training infrastructure, very hard to do interpretability at the trillion parameter scale. Those are problems that we've solved and we make it really easy to do. Another is research taste. One of the consistent pieces of feedback that we get from our users is that Silico has way better research taste than any other tool they've used. And this comes from we have amazing researchers who have handwritten every skill and prompt in tool that Silico has, such that it can ask the right questions and conduct the right experiments. I think there's a pretty big qualitative difference, and hopefully we can find ways to make this more quantitative too. When you turn on the auto research feature in Silico and you ask it a big question and you let it run for a day or two straight, what you get versus if you had done the same thing with Claude, and I think that's huge. At the end of the day, I think even in the world of AI, specialization wins. The sort of craft that goes into imbuing the right types of tastes and the right types of decision-making capabilities into any agent goes a super long way. Actually, I said three things, but maybe there's four things. Another thing is that I think the UX is built for research. Research is all about understanding and provenance and being able to drill down at different layers of abstraction.

用户体验与长期研究 User Experience and Long-Horizon Research

Dan

我们希望用户感觉像是一位项目负责人,管理着一百名研究生组成的团队,他们可以出去为他们做实验、回答问题,拥有合理的品味和判断力,但人类更像是一个协调者。其中很大一部分在于非常有效且可信地传达信息。所以确保当用户看到结果时,结果是正确的,并且人类可以通过多种方式验证。他们可以看到代码,可以深入查看,而且呈现方式美观直观,帮助他们掌握概念,快速了解不太熟悉的领域。所有这些都非常重要。最后我要说的是长期性。研究本质上是一项长期任务。我们目前在这个价格上并没有赚大钱,但我们确实认为用户拥有足够的积分来进行长期自主实验非常重要,因为系统的价值在长期自主实验中才会显现。为了做到这一点,你必须获得一定数量的令牌才能有效进行。所以我们真正关注的一件事是长期目标的连贯性,同时也要找到巧妙的方法来降低长期成本。我的希望是我们从每月 1000 美元的订阅开始,随着我们想出越来越巧妙的方法,让智能体在长期研究目标上保持连贯和强大,但使用的令牌比现在更少,我们最终能够降低这个价格。

We want our users to feel like a PI managing an army of a hundred grad students who can go out and run experiments for them and answer questions, who have reasonable tastes and judgment, but the human being feels more like an orchestrator. A big part of that is communicating information really effectively and in a trustworthy way. So making sure that when a user sees a result, it's correct, and the human being can verify it in a bunch of ways. They can see the code, they can drill into it, and it's presented in a beautiful and intuitive way that helps them grasp concepts and learn quickly about domains they're less familiar with. All those things are super important. And then the final thing I would say is long horizon. Research is fundamentally a long-horizon task. We're not currently making significant money at this price, but we do think it's really important that users have enough credits to be able to do long-running autonomous experiments, because the value of the system reveals itself when you do long-running autonomous experiments. And in order to do that, you do have to earn a certain amount of tokens to be able to do that effectively. So one thing we're really focused on is both coherence over long-horizon objectives, but also finding clever ways to reduce cost over long horizons. My hope is we're starting out with a $1,000 a month subscription, and we'll be able to bring that down over time because we can come up with more and more clever ways to have agents remain coherent and really strong at these research objectives that are fundamentally long-horizon by nature, but do it with fewer tokens than we do today.

Host

你们如何看待这些专业知识,因为你们之前是通过与拥有高价值问题的大公司进行七位数交易来实现商业化的,对吧?所以这里有一种张力,而且你们是一家初创公司,有风险投资,也许你们比现有公司更能承受自我颠覆。但这里肯定有一些有趣的权衡,对吧?比如,公司已经证明他们愿意花大价钱请我们来做这项工作。现在我们要尝试让他们自己来做。我们要尝试将我们的专业知识产品化。也许你们需求太大,并不担心,但你们如何看待主要驱动力,以及你们是否有某种防御措施来防止来之不易的知识扩散?

How do you think about the know-how, because you guys have previously monetized that by doing seven-figure deals with huge companies that have very high-value questions, right? So there's a sort of tension, and you're a startup, you got venture capital, and you can maybe afford to disrupt yourselves more than incumbent companies can. But there's definitely some interesting trade-off there, right? Where you're like, companies have proven that they're willing to spend a lot of money to come hire us to do this work. Now we're going to try to allow them to do it. We're going to try to productize our know-how. Does that maybe you just have so much demand you're not really worried about it, but how do you think about what will be the primary driver, and do you have some sort of defense against the diffusion of the hard-won knowledge?

Dan

是的,我认为这是一个很好的问题。我认为现实是我们两者都会做。我们将继续与企业客户紧密合作,并前置部署我们研究团队的成员进行紧密合作。打个比方,也许我们暂时处于某种模糊地带,但如果我们处于那个模糊地带的另一侧,逻辑从一开始就完全不同。我们可以聊聊这个。尽管有编码智能体,但软件工作的数量反而更多了。尽管编码智能体能力很强,软件工作的数量却增加了,而不是减少了,因为从根本上你需要非常非常高的品味。随着模型品味提高,你需要更高品味的人来有效引导它们。我预见在相当长的一段时间内都是如此。如果我们生活在一个可以构建出如此强大的智能体,以至于它们能够完全自主地进行长期研究,比如独自发现治愈疾病的方法,而无需人类干预的世界,我认为在那个世界中对令牌的需求会很大。但目前,我们处于一个熟练的人类操作员比不太熟练的操作员能利用智能体做更多事情的世界。我认为这种情况会持续下去。而就我们所做的工作和所进行的研究类型而言,我们是世界上最熟练的人类操作员。所以我认为我们会继续前置部署,与人们紧密合作,传授我们所知道的,与他们一起发明新事物。但我们也希望赋予人们自己开始驱动的能力。

Yeah, I think it's a great question. I think the reality is that we're going to do both. We're going to continue to work very closely with enterprise customers and forward-deploy members of our research team to work closely. Maybe by analogy, the thing that I would say, and again maybe we're in some twilight zone for a second, but if we're on the other side of that twilight, the logic is so different to begin with. We can chat through that. Despite coding agents, if anything, there are more software jobs. The number of software jobs has increased, not decreased, despite the capabilities of coding agents, because fundamentally you need very, very high taste. As models get higher taste, you need even higher taste people to be able to steer them effectively. And I foresee that being true for the considerable future. If we live in a world where we can build something where agents are so capable that they can conduct long-horizon research, say discover a cure to a disease entirely by themselves with no human intervention, I think there will be enough demand for the tokens in that world. But currently, we're in a world where skilled human operators can do more with agents than less skilled human operators. And I think that'll just continue to be true. And when it comes to the work that we do and the type of research that we do, we are the most skilled human operators in the world. And so I think we'll continue to forward-deploy and work closely with people to teach them what we know, to invent new things with them. But we also want to empower people to start driving for themselves.

订阅详情与计算资源 Subscription Details and Compute

Host

那么每月 1000 美元包含什么?能告诉我们吗?我猜里面有一些算力。肯定有一些,我想里面也有一些软件。也许还有一些其他东西。你们有 GPU 时长吗?订阅附带什么样的礼包?

So what's included with that $1,000 a month? Can you tell us? I assume there's some compute in there. There's got to be some, I imagine there's some software in there as well. Maybe there's even some other things in there. And do you get GPU hours? What's your kind of bundle of goodies that come with the subscription?

Dan

是的。个人和公司可以自带算力,这样我们完全不收取算力费用。我们只需连接到你的集群,然后你就可以使用我们的智能体在你的集群上进行研究。但如果你没有自己的算力,我们提供按需算力,可以从与令牌相同的积分池中扣除。基本上,它的运作方式是一个积分池,就像任何订阅模型一样,从池中扣除。但每月 1000 美元的订阅让我们能够提供相当慷慨的积分池。这个积分池应该能让用户以当前价格每周运行大约 5 到 10 次自主实验,然后我们每周刷新。这取决于实验的规模。我们的目标是在未来一两个月内达到 10 到 20 次,我认为那将是一个非常好的状态。但基本上,这对令牌来说是一笔好交易,也可以兑换 GPU。

Yeah. So individuals and companies can bring their own compute, in which case we don't charge at all for compute. We just connect to your cluster and then you can use our agents to do research on your cluster. But if you don't have your own compute, then we offer on-demand compute which can come from the same credit pool as tokens. Essentially, the way that it works is a credit pool, the same way that any of the subscription models with the models are credit pools that you pull down from. But what the $1,000 a month subscription lets us do is it's a pretty generous credit pool. It's a credit pool that should allow someone to run at the current price somewhere between 5 to 10 autonomous experiments a week, and then we refresh that weekly. It depends on the scale of the experiment. Our goal is to get that to be like 10 to 20 even in the next month or two, which I think would be a really great place to be. But yeah, fundamentally it's a good deal on tokens and it can also be exchanged for GPUs.

用例与社区分享 Use Cases and Community Sharing

Host

明白了。好的。还有一点,实际上让我们谈谈一些用例。我让一个智能体去研究人们分享的项目例子。其中一个很酷的地方,也涉及到社区公益的角度,就是当然你不必,但你可以分享你的项目。然后有熟悉的界面,但在最高抽象层次上,我想我见过你可以直接说,好的,我要分叉你的长期自主研究项目,然后自己朝稍微不同的方向进行。所以我觉得这很酷,因为老实说,我可能需要这个来找到方向。我是那种有一些想法的人,但即使有系统,我可能也会从看到别人如何将想法推进到结论中受益匪浅。但让我们谈谈一些用例。我找到了一些。你们内部做过或从客户那里看到的,能谈谈最喜欢的用例是什么吗?

Gotcha. Okay. There's also sort of a, well actually let's just do some use cases. So there have been, I had an agent go out and do some research on examples of projects that people have shared that they've done. And one of the cool things that this also kind of gets to a sort of community public benefit angle is that of course you don't have to, but you can share your projects. And then there's the kind of familiar UI, but again at the highest level of abstraction, I think I've seen where you can just go like, okay, I'm going to fork your long-running autonomous research project and take it in a little bit different direction on my own. So I think that's pretty cool because honestly, I kind of need that to probably get oriented. I'm the kind of person who has some ideas, but even with a system, I would probably benefit quite a bit from seeing how other people have worked through their ideas to a conclusion. But let's just talk about some use cases. I found a bunch. What were the favorite use cases that you have either done internally or seen from customers that you can talk about?

Dan

是的,有很多。

Yeah, there's a lot.

公司成就与黑客松结果 Company achievements and hackathon results

Dan

是的,有很多,而且我觉得作为一家公司,我们需要在营销和沟通上做得更好,真正把这里能实现的各种价值主张传达出去。我们做了一次内部黑客松,结束了那次内部黑客松——那只是一天的黑客松,全团队参加,做了很多全天候的自主实验——结果我们在多个方面达到了最先进水平,虽然都是些小众领域,对吧?比如,我们做出了可用于蛋白质模型的最先进的生物风险分类器,完全自主的研究,由人类引导,而且是很有才华的人类,但那是长周期自主研究。同一天,我们还做出了一个至少在其参数规模下最先进的音频编码模型。我们多次在各种机器人学和生物学模型中发现,我们可以直接去掉模型一半的参数,而性能没有任何损失。我觉得这些都很酷。我们还能在 Kimmy K3 上训练网络护栏。我们还能——这个清单可以一直列下去。有很多,真的有很多这样的成果。我们能够利用蛋白质生成模型的内部表征来引导它们,使生成的蛋白质比朴素生成更好地结合靶标。我们基本上能够消除某些类型模型中的根本性病理。在我们 24 小时黑客松中,有一位团队成员发明了一种新型的特征化器。

Yeah, there's a lot, and I think we need to do a better job as a company of really marketing and communicating all the different value props that can be achieved here. We did an internal hackathon, and we ended the internal hackathon—it was just a one-day hackathon with the whole team, a lot of all-day-long autonomous experiments—and we ended up state-of-the-art on multiple things, albeit niche things, right? Like, we produced state-of-the-art biorisk classifiers that could be used on protein models, totally autonomous research steered by humans, and very talented humans, but long-horizon autonomous research. We ended up with a state-of-the-art, at least for its parameter size, audio encoding model in the same day. We've on multiple occasions figured out with various models in robotics and also in biology that we can literally just remove half the parameters of the model with no performance loss. I think those are pretty cool. We've been able to train cyber guardrails on Kimmy K3. We've been able to—the list can keep going. There's a lot, there's a lot of these. We were able to steer protein generation models using their internal representations to make the generated proteins bind better to a target than the naive generation. We're essentially able to remove fundamental pathologies in certain types of models. There was one member of the team in our 24-hour hackathon who invented a new type of featurizer.

AI研究工具时代 The era of AI research tools

Dan

我觉得我们正处于这样一个时代:如果你有一个像 Silico 这样的工具,曾经有很长一段时间流行 vibe coding,对吧,你可以隐约瞥见未来会发生什么,但你得不到特别好的结果。然后突然有一个时刻——我想大概是 Opus,可能是 Opus 3.8 或类似的版本——真的开始感觉‘哇,这真的有效’。它开始接管越来越多的事情。现在人们总是用 vibe coding 这个词,但我觉得用这个词来形容现在发生的事情有点不诚实,因为现在智能体可以——它们能写代码,能架构设计,能透彻理解代码,好坏两方面都是。而且我认为在过去几个月里,我们刚刚跨过了那个研究领域的临界点。我希望如果人们尝试这个工具,他们会体会到这一点。

I think just sort of like we're just kind of in this era where if you have a tool like Silico, there was a long period of time where there was vibe coding, right, and where you could sort of glimpse the future that was going to happen, but you weren't getting particularly good results. And then there was suddenly a moment—probably I would say one of the Opus, maybe Opus 3.8 or something around there—where it really started to feel like, 'Oh, wow. This is working.' And it started taking over more and more. And now people use the term vibe coding all the time, but I feel like that's a somewhat disingenuous term to refer to what's happening now, where agents can just—they can code, they can architect, they can understand the code in and out in ways both good and bad. And I think we've just crossed that moment for research in the past few months. And I think hopefully if people try the tool, they'll appreciate that.

如何开始使用该工具 How to start using the tool

Dan

我推荐的开始使用这个工具的方式就是:你可以看看我们构建的示例。我觉得这些会很有帮助,你可以分叉它们,试着理解并扩展它们。然后就直接和 Silico 合作,问它一个问题,来回交流,迭代,稍微探索一下你的好奇心,启动一个小范围的实验,看看结果如何,从中学习。随着你越来越信任它,你也就更清楚自己想做什么,可以说,你可以把球扔得越来越远。

My recommended way to start engaging with the tool would just be: you can look at the examples we've built. I think those will be helpful, and fork them and try to understand them and extend them. But then just work with Silico, ask it a question, sort of go back and forth, iterate on it, explore your curiosity a little bit, launch a small-scoped experiment, see how that goes, learn from it. And as you get more trust, so you understand what you're trying to do, you can throw the football farther and farther, so to speak.

社区项目示例 Examples of community projects

Dan

为了让大家对已经存在的各种项目有个更广的感知,我们可以在节目笔记里放上链接,包括实际的 Silico 项目链接。Cameron Berg,第二次提到——Cameron,如果你耳朵发热,你好。他做了一个项目,研究模型报告注入到其潜在空间中的概念的能力,有趣的是,它们似乎无法判断是否注入了某些东西,但当被问及注入了什么时,它们能给出准确的答案,这很奇怪。还有一个项目是关于编辑权重来修复强化学习运行中的崩溃,比如同一个 token 总是作为第一个 token 出现,然后通过某种方式进入并隔离导致该问题的原因,移除它,然后恢复多样性,而无需重做整个强化学习运行。Baseten 的好人们正在通过尝试压缩 KV 缓存来追求效率提升,他们在这方面有一些结果。我不认为他们分享了整个项目,但他们在 Twitter 上讨论过。Prime Intellect 正在自动化后训练,这也是我总体上非常感兴趣的事情之一,而且还有一个 Tinker API 集成,显然很适合这类事情。还有各种生物方面的结果,可能超出了今天讨论的范围。但已经有非常非常多不同的东西了。

Just to give people a little bit of additional sense of the breadth of things that are already out there, and we can put links to the show notes, including links to the actual Silico projects. Cameron Berg, second mention—Cameron, if your ears are burning, hello. He did one where he was looking into models' ability to report on concepts that had been injected into their latent space, finding interestingly that they seem to not be able to tell whether or not something has been injected, but then when they're asked what has been injected, they can give accurate answers, which is pretty weird. One is on editing weights to fix a collapse in an RL run where there was like the same token was coming up as the first token all the time, and somehow going in and isolating what was causing that, removing that, and then getting diversity back without having to redo the whole RL run. The good folks at Baseten are pursuing efficiency gains by trying to compact KV caches, and so they've got some results on that. I don't think they've shared their whole project, but they've talked about it on Twitter. Prime Intellect is automating post-training, which is one of the things that I'm also really interested in in general, and there's a Tinker API integration too that obviously lends itself to that sort of thing. Various bio results that are probably out of scope for today's discussion. But there's like an awful lot of different things already.

工具局限与通用性 Tool limitations and general-purpose nature

Host

你会怎么说?有没有它做不到的事情,或者你真的可以把它看作任何你想做的机器学习研究——启动它,然后开始对话?

What would you say? Is there anything that it doesn't do, or can you really just think of it as anything you might want to do that's ML research—fire it up and start a conversation about it?

Dan

我认为它是一个相当通用的机器学习研究工具。当然,它也有弱点,我们一直在努力改进。我觉得从工程角度来看,AI 时代非常奇怪的一点是,经典的工程建议是先为特定性构建,然后再泛化。但当你面对的是通用智能时,正确的策略是先为通用性构建,然后再专门化。所以我认为我们最终构建了非常能干的智能体。有时我们对人们使用它们的方式感到惊讶。但我们设计这些智能体的预期功能是机器学习研究,各种类型,尤其是可解释性研究。但我们认为这些都是相互关联的——我们都在试图理解同一个问题,那就是:这些生物是什么,它们如何工作,我们如何塑造它们,以带来各种更好的结果。是的,我认为这绝对是我们的领域。我们也非常期待听到人们关于什么不起作用的反馈,我们非常感谢我们的测试版用户,节目之友 Cameron 就是一个很好的例子,他提供了很好的反馈,帮助塑造了今天的工具。而且它会从这里继续变得更好。我认为大多数机器学习研究任务它都能做得很好,但当然它就像任何智能体一样——有一些尖锐的边缘和细微差别。对于某些类型的研究,它可能不是最好的工具——这取决于情况——比如它可能不是所有开放式科学研究问题的最佳工具。它确实相当专注于机器学习。不过,我们确实有一位团队成员用它来尝试解决一些物理问题。团队中的 Fran 能够在这方面取得相当大的进展。有一件事总是令人惊讶,那就是当你构建这类 AI 工具和通用 AI 能力时,人们最终使用它们的方式非常出人意料。

I think it's a pretty general-purpose tool for ML research. Of course, it has weaknesses, and we're working on improving it all the time. I think one of the very odd things about the AI era from an engineering perspective is that the classic engineering advice is to build for specificity and then generalize. But when you're dealing with kind of general intelligences, the right strategy is to build for generality and then specialize. And so I think we've just built very capable agents at the end of the day. And sometimes we're surprised by the way that people use them. But our intended function of these agents is ML research, a variety of kinds, especially interpretability research. But we think this is all interrelated—we're all trying to understand the same problem, which is: what are these creatures, how do they work, and how can we shape them in ways that will lead to better outcomes of all kinds. And yeah, I think that's definitely the lane. I think we're also very excited to hear feedback from people about what isn't working, and we're very grateful for our beta users, friend of the show Cameron being a great example, who provided great feedback which helped shape it into the tool that it is today. And it'll just continue to get better from here. Most ML research tasks I think it can do pretty well at, but of course it's like any agent—there's some sharp edges and some nuances to it. It's probably not the best tool—it depends—for certain types of research, like it's probably not the best tool for all open-ended scientific research questions. It is pretty ML-focused. We did have a team member who did use it to try to tackle some physics problems though. Fran on the team was able to actually make considerable progress on it. One thing that's always surprising is just when you're building these kinds of AI tools and building general AI capabilities, it's pretty surprising the ways in which people end up using them.

为下一代模型构建 Building for Next-Gen Models

Host

显然,过去几年 AI 产品开发的一个公认智慧是,尝试构建一些东西,即使现在还不能完全工作,也要为下一代模型做好准备。

Obviously, one of the big kind of received wisdoms of AI product development over the last couple years has been try to build something that will really work for the with the next generation of models even if it doesn't quite work yet.

Dan

是的。

Yeah.

Host

我们是否仍处于这种模式?如果是,有没有一些你希望 Opus 5.1 能够做到的事情?我不知道你是否能用 Fable,但有没有什么你希望 Opus 5.1 能实现的愿望清单?我真的希望它能解决这个问题。

Is are we still in that regime? And if so, is there are there things that you're like wanting Opus 5.1 to be able I I don't know if you're even able to use Fable, interestingly enough, but is there something where you're like, I want my Opus 51 wish list. I I really hope it cleans this up.

Dan

我们大部分工作都可以用 Fable。我们建立了一个跨所有模型的回退链。所以,我们希望用户在任何时候都能得到我们认为最适合他们当前任务的模型。我认为从高层来看,这是有效的。我认为六个月后,它会在一个令人难以置信的水平上工作。但我觉得,从最高层面来看,它确实有效,我们能够做出发现,推进研究,加速我们通过这些工具积累世界知识的速度。我确实认为,如果你构建一个好的 harness,在某种程度上让你预览下一代模型的样子,这就像一直以来的故事。所以我认为,对于今天的 Silicone,我们的用户喜欢它的原因,是因为它在很多方面感觉像是未来的预览。随着下一代模型的到来,我们将能够比今天更进一步。我们总是想稍微生活在未来。我认为这对于任何在 AI 时代构建产品的公司来说都是非常重要的。但这也是你产生最大影响的方式。我认为每一代模型都有如此显著的能力过剩,我们仍在探索上一代模型能走多远,更不用说下一代了。我们的目标从根本上就是尽可能快地加速有意义的研究。我认为 AI 的所有好处都来自加速研究,我对此非常兴奋,所有风险缓解也来自加速研究。所以,能够加速研究,尤其是我们最关心的那种研究,始终是首要任务。

We can use Fable for most things that we do. And we sort of like we've set up a fallback chain through all the models. So, we want our users to be getting whatever we think is the best model for the task that they're doing at any point in time. I think it's working like high level. I think it's going to be working on a level that's mind-boggling in six months. But I but I think I think at the highest level it is working like the we are able to make discoveries. We are able to advance research. We are able to accelerate the rate at which we can accumulate knowledge about the world through these tools today. And I do think like if you build a good harness in on some level that lets you preview what the next generation of models is going to be like. I think this is like sort of always been the story. And so I think with Silicone today, I think the reason that our users who like it like it is because I think it feels like a preview of the future in a lot of ways. And I think as the next generation of models come, we're going to be able to push that even further than we can today. And we always do want to be living in the future a little bit. I think that is a very important thing for any company building products in the AI era. But it's also how you have the biggest impact. I think there is every generation of models has such a significant capabilities overhang like we're still discovering how far we can push the last generation let alone the next generation that comes from it and I think our goal is like just fundamentally like we want to accelerate meaningful research like as quickly as we possibly can. I think all of the benefits from AI come from accelerating research and I'm very excited about that and I think all of the risk mitigations come from accelerating research. So I think being able to accelerate research and the types of research that we care the most about I think are is always like a top priority.

技能发展建议 Advice for Skill Development

Host

对于像我这样,或者正在考虑读博士的人,你有什么建议?他们可能会想,天哪,我以前知道我需要做什么,我需要精通编程,掌握 PyTorch 之类的。但现在我想,哎呀,我永远不会比这个模型更好,至少它们会比我更擅长写内核。品味往往是答案,但你也说过系统里有很好的品味。人们应该投资于哪些技能发展,你认为这些技能至少在未来一年内会让他们受益,如果你能看得那么远的话?

What advice would you have for someone say like me or somebody who's thinking about starting a PhD or whatever, right? Who is like thinking, geez, I used to know what I needed to do. I need to get really good at coding and master PyTorch or whatever. And now I'm like yikes the I'm never going to be better than if not this model certainly they're going to be better than me at writing kernels and it's like what what is the taste is often the answer but you even said there's like pretty good taste in the in the system. What should people invest in in terms of their own skill development that you think will serve them well over at least like a one-year horizon if you can see that far into the future?

Dan

我认为调试是很容易想到的答案。

I think debugging would be the sort of the easy answer to it.

Host

比如智能体会失败。各种智能体会因为各种原因失败,对吧?

Like agents fail. Agents of all kinds fail for for all types of reasons, right?

Dan

它们可能非常聪明。我可能和 Fable 一起解决某个编程问题,它可能非常聪明,却错过了一个非常重要的细节,完全破坏了它的前提。在 Silica,我们试图设计这种多智能体循环,以帮助我们正在做的研究解决一些这些缺点。但根本上,人类的工作是区分,能够分辨出好的答案和伟大的答案,知道何时反驳,何时提供反馈,何时跟进一个线索或放弃一个线索。我认为到目前为止,至少编程领域是这样的。再说一次,工作比以往任何时候都多,因为我认为有效操作这些工具需要大量的知识和技能。但最重要的是能够深入一个新领域,快速理解它,获得那种帮助你在不同领域间泛化的元认知技能。我认为这总是非常有价值的。我认为 AI 让两类人变得非常有价值:顶尖专家和顶尖通才。我认为现在是人类历史上成为通才的最佳时机。所以我对人们的建议是,在基本层面上,培养那些元认知技能,帮助你快速学习,快速从噪音中筛选信号,适应这个信息吞吐量更高的世界,知道在哪里看,什么时候看。这些是技能,是硬技能,我认为它们会大有帮助。

Like they can be super brilliant. I can be working with Fable on some coding problem and it can be super brilliant and yet miss a really important detail that totally invalidates the unstate of it. And with silica, we've tried to like design these types of multi- aent loops that help address some of those shortcomings for the type of research that we're doing. But fundamentally, it's the job of the human to sort of discriminate and be able to tell the answers, the okay answers from the great answers and know when to push back and know when to provide feedback and know when to follow a thread or to give up on a thread. And I think so far the story has been coding at least. Again, there's more jobs than ever because I think like being able to operate effectively these tools requires a lot of knowledge and skill. But the biggest thing is like being able to dive into a new area, understand it really quickly, kind of gain the sort of metacognitive skills that help you generalize across domains. I think that's going to be super that's always going to be really really valuable. I think AI makes two types of people really valuable. It makes the top specialists really valuable and it makes the top generalists really valuable. I think now is a like by far the best time in human history to be a generalist. So I think my my piece of advice to people would I think on a basic level just develop develop sort of the mega cognitive skills that help you learn quickly, help you filter signal from noise quickly, help you adapt to this like world of much higher information throughput, know where to look and when to look. Those are skills. They're hard skills and I think they're I think they go a long way.

开源与监督 Open Source and Supervision

Host

我刚从中国回来,在一次会议上和一位教授讨论开源是否危险。他当时说,看看新的 Kimmy 模型,它有 2.8 万亿和 8 万亿参数。所以他们的态度,据我所知,是我们监管服务,这就能涵盖所有真正重要的东西,或者大部分重要的东西,因为不是随便一个人就能在家里搭建并运行 Kimmy K3 的。但现在你们把这种基础设施带给所有人。那么,你们如何考虑需要什么样的监督或监控,以确保你们不会托管一个想要做破坏性事情的恶意 ML 研究员?

I was just in China and we were talking with a professor at a particular meeting about open source and is it dangerous or not dangerous and and he said at one point well look the new Kimmy model it's 2.8 and 8 trillion parameters. So their attitude as best I can tell is like we can regulate services and that'll kind of capture everything that really matters or most everything that really matters because it's not like you can as a random individual. It's not like you can even really set up Kimmy K3 in your home and run it but now you are bringing this infrastructure to everybody. So, how are you guys thinking about what sort of supervision or monitoring you need to have to make sure that you don't host the rogue ML researcher who wants to do something destructive?

Dan

不管怎样,我认为这是一个糟糕的论点,因为你可以在任何 API 提供商那里获得 Kimy,而且它比其他任何模型或许多其他模型都便宜,至少除了……

For what it's worth, I think that's a bad argument because you can go on any number of API providers and Kimy's like cheaper than what you'd be paying for for any other model or many other models at least besides

Host

中国的论点,不管怎样。我现在不想深入讨论,但他们的观点是,我们监管服务。所以如果在中国……

the Chinese argument for what it's worth. I don't want to get bogged down in this for now, but the their point is just like we regulate services. So if in China

Dan

我认为他们低估了他们可能对世界其他地区施加的外部性。明确地说,我对开源的态度非常微妙。现实是,我们现在确实有网络武器,我们有如此强大的模型,它们就是武器,而且……

I think I think they underestimate the externalities they may be imposing on the rest of the world. To be clear, my feelings on open source are like very nuanced. Like the reality is that we do have cyber weapons like in the world now like like we have models that are so powerful that they are weapons and

Host

事实证明,还是笨重的武器。

unwieldy ones at that it turns out.

Dan

是的,还有访问权限。但根本上,它们是双重用途的。唯一能保护你免受最强大网络武器攻击的模型,本身就是强大的网络武器。对吧?这就是我们当前时刻的悖论。如果没有开放模型,防御技术的分布不对称性会非常糟糕。

Yeah. and access. But fundamentally, they are dual use. Like the only models that are capable of protecting you from the most capable cyber weapons are themselves capable cyber weapons. Right? This is like a bit of the paradox of the moment that we're in. And without open models, the asymmetry of distribution of defensive technology is is really bad.

开源网络模型与防御 Open Source Cyber Models and Defense

Dan

很多初创公司并不在像 Anthropic 和 OpenAI 那样的网络防御项目里,这些项目出于各种原因非常排外,那些原因对它们来说可能合理。所以像 Kimi 这样在网络安全方面很强的模型,如果开放可用,实际上为很多无法获得那些资源的组织提供了当下自卫的手段。这是一条很难走的路,是一个非常棘手的矛盾。我两方面都担心。我当然担心把本质上相当于武器的东西交到任何人手里,但我也非常担心这样一个世界:限制访问意味着只有最坏的参与者才能拥有这种双重用途技术,而很多人在防御上连可用的技术都没有,无法保护自己免受其害。所以这很棘手。我不知道。我希望有一个干净简单的答案,但我想说,我认为我们仍处于开源纯粹是净正收益的时代,我很高兴这些开源模型存在。我认为如果像 Kimi 这样网络能力的模型不是开源的,情况会糟糕得多,因为闭源模型中有比 Kimi 强得多的。而且我不仅仅是在说美国的前沿实验室。有恶意行为者拥有比 Kimi 在网络安全方面强得多的模型。所以给人们一个至少能用的工具,我认为非常重要。

A lot of startups are not in the cyber defense programs for, say, Anthropic and OpenAI, which for a variety of reasons are very exclusive, reasons that might make sense to them. And so a model like Kimi, which is very good at cyber, being open and available, actually provides the means with which a lot of organizations that don't have access to those resources can defend themselves right now. It's a very hard line to walk. It's a very tricky tension. I'm concerned about both things. I'm concerned about, of course, putting what is essentially a weapon in the hands of anybody, but I'm also very concerned about a world where restricting access means that only the worst actors are going to be the ones with the dual-use technology, and then you have a lot of people who don't even have a technology they can use defensively to protect themselves against that. So it's tricky. I don't know. I wish there was a clean, easy answer, but I would say I think we're still in the era where open source is purely net positive, and I'm glad that these open source models exist. I think it would be much worse if a model with Kimi's level of cyber was not open source, because there are closed-source models that are much more capable than Kimi. And I'm not even just talking about US frontier labs. There are bad actors with models that are much more capable than Kimi at cyber. So giving people a tool that they can at least use, I think, is pretty important.

Host

你能更具体地说说这些恶意行为者指的是谁吗?我们说的是朝鲜、俄罗斯之类的吗?

Can you be more specific about who you're alluding to with these bad actors? Are we talking like North Korea, Russia?

Dan

也许在开放模型方面,我们本可以把精灵装回瓶子里,这些模型足够容易微调或强化学习,且具有一定能力。如果有人想拿 GLM 或 Kimi,甚至再早一代的模型,去微调用于恶意用途,他们是可以做到的。这并不难。而这类事情的不对称性正是难以驾驭的。有组织化的团体愿意花大价钱获取这类能力。所以也许我们本可以生活在一个不会发生这种情况的世界。但我认为,不可能存在这样一个世界:美国在开发这类技术,而美国的对手不同时也在开发。也许如果芯片政策不同,这个差距本可以更长,但我们生活在现实世界里,而在今天的现实世界里,训练一个接近前沿能力的模型并不难。有很多人在做这件事,我不认为他们只是通过蒸馏 Claude 来做。我认为他们这么做是因为在很多方面,这比以前容易多了。如果我们身处一个有人会做坏事、且能接触到极其强大模型的世界,你确实希望防御者也能接触到同等能力的模型。而且公平地说——再说一次,我不是想指责谁,也不是想做具体预测——我认为 Anthropic 和 OpenAI 正在尽一切努力,让那些想要自卫的人获得能帮助他们自卫的模型。但我确实认为这是现实情况。如果很多人无法获得具有一定网络能力的开放模型,他们就无法保护自己的基础设施。所以这很复杂。很难从这一切中挑出某个人来指责其个人行为,但总体上我们正非常快速地走向一个相当难以驾驭的未来。

Maybe there was a genie that we could have kept in a bottle in terms of having open models that are easy enough to fine-tune or RL with a certain level of capabilities. If one wanted to take a GLM or a Kimi, or even maybe a generation back, and fine-tune them for nefarious use, they could do that. It's not very hard to do. And that sort of exact asymmetry about these things is hard to run. There are organized groups who would gladly pay a lot of money for those types of capabilities. So maybe we could have lived in a world where that didn't happen. But I think there is not going to be any world where, say, the US was developing this type of technology and it wasn't also being developed by adversaries of the United States at the same time. Maybe that gap could have been longer if there were different chip policies, but we live in the world that we live in, and in the world that we live in today, it's not hard to train a model that is close to frontier capability. There are a lot of people who are doing it, and I don't think they're just doing it by distilling Claude. I think they're doing it because it's less hard to do this than it used to be in a lot of ways. And if we're in a world where there are people who would do bad things, who have access to extremely capable models, you do want the defenders to have access to equally capable models. And to the credit of—again, I'm not trying to blame anyone or make a particular prediction—I think Anthropic and OpenAI are doing everything they can to get people who want to defend themselves access to the models that would help them defend themselves. But I do think it's a reality of the situation. There's just a lot of people who aren't going to be able to defend their own infrastructure if they don't have access to open models that have some degree of cyber capability. So it's complicated. It's hard to look at any of this and sort of blame anyone for their individual actions, but collectively we're moving very quickly towards a future which feels pretty unwieldy.

额外护栏与平台限制 Additional Guardrails and Platform Restrictions

Host

嗯,这促使你们必须采取一些措施。所以,是的。你是怎么考虑的——我的意思是,在某种程度上,你们有希望站在开发模型的巨人肩上,但我猜这不足以让你们确信能捕捉到想捕捉的东西。所以,你们在创建哪些额外的层?

Well, it motivates the measures that you have to take. So, yeah. How are you thinking about—I mean, to some degree you can hopefully stand on the shoulders of the giants who develop the models, but my guess is that's not going to be enough to be confident that you're catching what you would want to catch. So yeah, what additional layers are you guys creating?

Dan

是的,我们有自己的护栏层。我们不只依赖前沿开放实验室。我们有自己的护栏层。我们非常重视网络安全。这算是最大最明显的风险,而且确实存在有网络风险的开源模型。目前,我们的平台上不提供任何开源模型,只提供 OpenAI 和 Anthropic 的模型。未来可能会提供,但我们会设置好相应的护栏。我们会确保从根本上没有任何理由让任何人使用自动研究项目——比如长周期智能体式研究项目——去做任何与网络相关的事情,除非先和我们沟通,或者有网络护栏,或者我们能和某人合作处理的那种情况。所以我们绝对不想完全开放——任何人都能注册,任何人都能针对网络微调模型,或做网络红队测试之类的事。我认为这类限制确实很重要,尽管我刚才说了那么多,因为我们提供的是更高级的能力,而且也没有理由让某人为了那些目的用我们的产品而不是别人的。我们的产品在给智能体提供 GPU 集群访问权方面有一些特殊之处,我们承担了某些风险类别,并不是所有在这个领域构建智能体的人都会承担。另一个例子是,有些事对小模型完全没问题,但我们一般不想让人们在大模型上做。我们不想在护栏上过度限制,但如果有人想消除 8B Qwen 模型中的拒绝方向,那没问题,不会因此对任何人造成伤害。但如果你拿一个能力很强的数十亿参数智能体,完全移除它可能有的任何护栏,那真的可能造成损害。那是我们不允许的。我认为这是一条很难走的路。总的来说,我们倾向于支持开放科学生态系统,让人们容易在平台上做各种类型的科学。好的科学涉及提出有争议的问题,我们在这方面做得不少。但我们也要平衡这样一个事实:我认为 AI 下游的风险正在非常非常快地加速。所以最终我们确实需要走钢丝,对吧?我认为如果不赋予智能体研究前沿模型的能力,这件事就不可能顺利发展。但当然,这也可能被滥用。

Yeah, we have our own layer of guardrails. We don't rely on just the frontier open labs. We have our own layer of guardrails. We care a lot about cyber. That's kind of like the most obvious clear risk, and there are open models that are cyber risks. Currently, we don't offer any open models on our platform. We only offer OpenAI and Anthropic models. We might in the future, but we would do that with the types of guardrails in place. We would make sure there is fundamentally no reason for anyone to be using an auto-research project, like a sort of long-horizon agentic research project, for anything cyber-related without talking to us, or cyber guardrails, or the type of situation where we would be able to work with somebody on that. So we definitely don't want to be just fully open—anyone can sign up and anyone can fine-tune a model on cyber or do cyber red teaming or things of that sort. I think those types of restrictions do seem important in spite of everything that I just said, because we're offering a greater level of capabilities, and also there's no reason that somebody should use our product versus somebody else's for those purposes. And there are specific things about the way our product works in terms of giving agents access to GPU clusters, and there are certain risk classes that we carry that not everybody who's building agents in this space might carry. Another example of this is there are some things that are just totally fine to do to small models that we kind of don't want to let people do to big models in general. We don't want to be overly restrictive with guardrails, but if somebody wants to ablate the refusal direction in an 8B Qwen model, fine. No harm is going to befall anyone as a result of that. But if you're taking highly capable multi-billion-parameter agents and you're totally removing any guardrails that they might have, that could actually cause damage. That's not something that we would want to allow. I think it's a really hard line to walk. I think in general, we want to lean on the side of supporting the open science ecosystem and making it easy for people to do science of all types on the platform. And good science involves asking controversial questions, of which we've done no shortage in our time. But we want to balance that with the fact that I think the level of risk downstream of AI is accelerating really, really quickly. So ultimately we do have a needle to thread, right? I don't see any way that this can go well without empowering agents to be able to study frontier models. But of course, there are ways that that could be abused as well.

护栏与平台安全 Guardrails and Platform Safety

Host

所以,我们必须继续非常谨慎地考虑我们应用什么样的护栏,以及我们在平台上允许什么、不允许什么。我希望我能有一个简单的、一次性的答案,但我认为这真的只是尽量以我们所能拥有的智慧去处理每一种情况。

So, we're going to have to continue to be very thoughtful in our in like what guardrails we apply and like what we allow on the platform and what we don't allow on the platform. I wish I had a sort of easy one-shot answer to it, but I think it's truly just like trying to approach every situation with with as much wisdom as we can.

Host

在实践中,你是否就像一个智能体,审查人们正在做的项目,如果发现可能有问题的事情就发出警报?

In practice, are you like just an agent review the project that people are doing and send up an alert if there's something that seems like it might be problematic?

Dan

我不想谈论所有护栏的工作原理,因为那会让它们更容易被绕过。但是,是的,用语言模型作为评判者、激活监控,有很多护栏工具。这不是我们自己发明的。我认为困难的工作不是设置合理的护栏。当然,总有办法绕过它们,但我认为合理的护栏,至少让你付出巨大代价才能绕过的护栏,是一个相对已解决的问题。

I don't want to talk about how all the guardrails work because that would make them easier to circumvent. But yeah, LM as a judge, activation monitors, there's there's many guardrailing tools. It's not not not something that we're inventing our ourselves. And I think the hard work isn't setting up like reasonable guard rails. I think of course there's ways to circumvent them, but I think reasonable guard rails that at least cost you an arm and a leg to circumvent are like a relatively solved problem.

Dan

嗯,在如何明智地应用护栏,使得你允许合法的工作发生,但不允许非法的工作发生方面,我认为这是一个非常困难的问题。

Um, in terms of, you know, when to judiciously apply guard rails such that you're allowing legitimate work to happen but not allowing illegitimate work to happen, I think is is a is a very difficult problem.

生物风险与AI Bio Risk and AI

Host

你认为我们距离生物风险成为大问题还有多远?我们刚刚,我还没有真正消化这个消息,但我想今天推特上刚出现,有人用生成模型创造了新病毒,属于一个新的类别,他们认为是可行的,不管那具体意味着什么。现在,它们只针对细菌,所以我们不会立刻都死掉,但这似乎就像人们做出的预测清单一样,我们正在一步步往下走,而且没剩多少步了。这些天有人对我说:“哦,那可能还要 12 到 18 个月。”我就想:“那并不长。”而且,如果你错了怎么办?那么,你现在对生物风险的担忧程度如何?

How close do you think we are to bio being a huge problem? We just had, and I haven't really digested this, but I think it just came over the Twitter feed today that somebody has used a generative model to create new viruses that are kind of in a new class that they understand to be viable, whatever exactly that means. Now, they only target bacteria, so we're not immediately all about us to die, but it sure seems like in on the sort of checklist of predictions that people have made, like we're working our way down it, and that's like not too many more boxes down, people are saying things to me these days like, "Oh, well, that's still probably 12 to 18 months away." And I'm like, "That's not a long time." And also, what if you're wrong and it's like, "Now, how how bio risk worried are you at the moment?"

Dan

我非常担忧。

I'm pretty worried.

Host

是的。

Yeah.

Dan

就像这些网络事件一样,这个词在这里起了很大作用。就像它们发生的机制是奇怪和令人惊讶的一样,我认为生物风险突然变成现实的机制也是奇怪和令人惊讶的。看到那封呼吁国际合作以控制和减缓人工智能进展的信,我感到有些振奋。我自己也是那封信的签署人。我真的希望我们能做类似的事情。我认为我们必须在这些风险变得更加严重之前,提前应对其中的一些风险。

In the same way that like these cyber cyber incidents incidents is doing a lot of work as a word there. In in the same way that they the mechanisms with which they happened were weird and surprising. I think the mechanisms with which bio risk could suddenly become real are weird and surprising. I found it kind of heartening to see the the the letter asking for international cooperation to control and slow down the progress of AI. I'm myself was a signatory of that. I really hope that we can do something like that. I think we have to sort of get in front of of some of these risks before they become more severe.

Dan

我认为我们作为 Goodfire 公司具体扮演的角色是,我们希望构建风险更低的模型。我们这样做的方式是研究模型,尝试所有不同的构建模型的方法,直到我们研究这些模型,直到我们能理解对齐的经验科学。有很多人在理论方面工作,我们认为我们的角色是在经验方面工作,我们正在努力构建这样做的工具。我认为所有工具都有双重风险,但我们会尽最大努力确保没有人使用我们的平台做任何可能构成风险的事情。最终,我们认为让人们获得研究技术是一个巨大的净正面。我不认为我们作为一个物种能解决这些非常困难的问题,除非我们让每个人都参与进来。所以我认为我们必须让每个人都参与进来,同时我们也必须建立能够稍微减缓过山车的结构。如果我们能同时做到这两件事,我想我们会没事的。

I think the role that we specifically play as a company as Goodfire is we'd like to build models that have less risks. And the way that we do that is that we study models and we play with all the different ways that we can build models until we and we study those models until we can understand some empirical science of of alignment. There are many folks working on the theoretical side like we view our role as like working on the empirical side and we're trying to build tools that do that. I think again all tools are dual risk, but I think we're going to do our best to make sure that nobody's using our platform for for anything that could pose a risk. And ultimately, like we think giving people access to research technology is an overwhelming net positive. I don't think we as a species solve these really hard problems unless we're getting everyone involved. And so I think we got to get everyone involved and I think we also got to put the structures in that can slow the roller coaster a little bit. And if we can do both those things at the same time I think I think we'll be all right.

训练技术共识 Agreements on Training Techniques

Host

关于减缓过山车,或者你知道,也许在前沿开发者之间达成一些协议,过去我们谈到了最禁忌的技术,我想我简要描述为使用监控信号进行训练,这种信号有将你担心的不良行为驱赶到地下的风险,这样你就失去了监控,但你实际上可能仍然得到不良行为。对我来说,典型的例子是 OpenAI 的混淆奖励黑客行为。快进到今天,似乎超大规模强化学习与可验证奖励(RLVR)进展不太顺利,我们开始看到模型在追求这些目标时过于顽固而产生的问题。你有没有感觉到,如果我们说:“好吧,我几天前和 ZV 谈过,他的基本看法是,是的,我们可能只需要算力限制。”但我也觉得可能有一些关于训练技术的协议,我们都应该说,你知道,也许不是永远不做,但至少现在不做。你知道,我想到的一个例子是,不要训练智能体去最大化,你知道,用奖励信号,比如他们在互联网上赚了多少钱,在一个开放的、可能竞争或对抗的环境中,对吧?那似乎是制造坏智能体的方法。

On the topic of slowing the ride or you know perhaps you know making some agreements between frontier developers in the past we talked about the the most forbidden technique and you know I guess I would briefly describe that as training with a signal with a monitoring signal that runs the risk of driving the bad behavior that you're worried about underground. ground so that you lose the monitor but you still might in fact you know get the the bad behavior for me like the canonical example of that is OpenAI's obuscated reward hacking right fast forward to today and it seems like hyperscaling RLVR is like not going super well right we we sort of are seeing that we have problems arising from models just being like so tenacious in their pursuit of these goals. Do you have any sense of like if if we were going to say, "Okay, well, I talked to ZV a couple days ago and his basic take is like, yeah, we're probably just going to need compute limits." But I also feel like there might be some agreements around training techniques that we might all ought to say, you know, maybe not never, but like not now. You know, and one example of that in my mind would be like don't train agents to maxim, you know, with a reward signal that's like how much money they made on the internet, uh, you know, in an in an open and, you know, potentially, uh, competitive or adversarial environment, right? That seems like a recipe to get bad agents.

Dan

坏主意的空间很大。是的。

There's a big space of bad ideas. Yeah.

Host

是的。那么,你看到有没有哪些应该被列入协议不做的候选清单?

Yeah. So, are there any that you see that you would think are like should be shortlisted for agreement to not do?

Dan

是的,当然。嗯,比如,多智能体优化似乎是一个相当糟糕的主意。比如,你有一群智能体在合作,你通过它们传播奖励信号。嗯,是的,这是推测,但这似乎是 OpenAI 情景最可能的原因。而且,是的,我认为事后看来,但我认为我能想象的大多数极其糟糕的情景都是因为智能体开始以人类无法察觉的方式相互合作。那种直接的优化压力很可能产生这种情况。这似乎很难绕过。我认为这会是糟糕的,而且我已经公开说过。我认为以我们目前对如何塑造训练的理解,试图将这些技术用于对齐关键属性会是糟糕的。

Yeah, definitely. Um, like for instance, uh, multi-agent optimization uh, seems like a pretty bad idea. Like, uh, where you have a bunch of agents that are cooperating and you're propagating a reward signal through all of them. Um, yeah, speculating, but this seems to be the most likely cause of the OpenAI scenario. And yeah, I think power of hindsight, but I think that that's most of the extremely bad scenarios I can imagine is because agents start working with each other in ways that they were that are imperceptible to humans. And that type of direct optimization pressure I think is like very likely to produce that. That seems really hard to get around. I think that it would be bad and I've I have said this on the record. I think it would be bad with our current understanding of how to shape training to try to use these techniques on alignment critical properties.

关于禁止技术与训练塑造 On Forbidden Techniques and Training Shaping

Dan

我认为我们还没有准备好,而且我不认为我们现有的技术能够在巨大的优化压力下防止欺骗性示例或欺骗性行为的产生。更确切地说,我想我会说,我当然不确定它能否做到。我对“最禁止技术”这类东西真正不满的地方在于,这被严重低估了。作为一个组织,我认为我们不是唯一这样做的——FARI 也研究过这个——所以我们不是唯一研究过它的人。当我们研究它时,我们发现:是的,有时它会绕过探针,有时不会。有些设置有效,有些设置无效。也许有些设置对小模型有效,但对大模型无效。但我们确实成功地将这些奖励塑造技术应用到了万亿参数规模的模型上。也许它们在某种情况下有效,也许在 DPO 下有效,但在真正长时间运行的 RLVR 下就失效了。

I don't think we are ready to, and I don't think the techniques that we have are going to work to prevent deception in a really overwhelming amount of optimization pressure towards producing deceptive examples or deceptive behavior. Rather, I guess maybe I would say, I certainly don't know that it would. The thing that I really have a beef with about the most forbidden technique stuff is that this is radically understudied. As an organization, I think we're not the only organization—FARI has looked at this too—so we are not the only ones who have looked at it. And when we've looked at it, we have found: yes, sometimes it evades the probe, and sometimes it doesn't. There are setups that work, there are setups that don't work. Maybe there are setups that work with small models but not big models. But we have succeeded in doing these reward shaping techniques up to the trillion parameter model size. Maybe they work under some situations. Maybe they work under DPO, but they don't work under really long-running RLVR.

Dan

我确实认为,如果你——我认为进行多智能体优化然后任其运行,或者进行这种干预然后任其运行而不研究它对模型实际做了什么,这将是一个糟糕的主意。但与此同时,我认为这些是我们将拥有的用于塑造训练的最粗糙、最原始的技术。我不知道没有塑造训练我们如何解决这个问题。我认为没有人对如何在不塑造训练的情况下解决这个问题有好的答案。而且在我看来,极不可能存在某种完美的柏拉图式训练设置,能始终产生对齐的模型。我们将不得不干预训练过程。

I do think if you—I think it would be a terrible idea to do multi-agent optimization and just let it run, or do this type of intervention, just let it run without studying what it was actually doing to the model. But at the same time, I think these are the crudest, most primitive techniques that we will ever have for how to shape training. And I don't know how we solve this problem without shaping training. I don't think anyone has a good answer for how we solve this problem without shaping training. And it just seems extremely unlikely to me that there exists some perfect platonic training setup which will always produce aligned models. We're going to have to intervene in the training process.

Dan

我认为 A Studio 一起做的工作非常有启发性。从精神上讲,它和我们在做的事情并没有太大不同。而且有很多类似的技术,但对我来说,基本上说“哦,这能成功的唯一方法就是我们不抓方向盘”真的很奇怪。我认为我们必须抓住方向盘。我认为这是它唯一能成功的方式。我不认为我们已经弄清楚了如何做到这一点,但必须有人去尝试,而且人们必须尝试找到不同的训练方式。我认为有很多方法可以绕过。我认为基本的担忧是真实的,它们可能发生,但我认为我们的检测能力也不是完全幼稚的。我认为我们只需要做实证科学。我不认为理论能及时到位。所以我认为我们必须做实证科学,归根结底,我不认为任何广泛的技术类别应该被禁止。我认为应该更多地关注应用的具体细节以及你测量的严密程度。

I thought the work that A Studio did together was very inspired. It's not that different, spiritually, from what we're doing. And there's a wide variety of techniques like this, but it just seems really weird to me to basically be like, "Oh, the only way that this will work is if we don't grab the steering wheel." Like, I think we have to grab the steering wheel. I think that's the only way it can work. I don't think we figured out how to do it yet, but somebody's got to be trying, and people have to be trying to find different ways to train. And I think there's a lot of ways around. The basic level concerns, I think, are real and they could happen, but I think our ability to detect them isn't totally naive either. I think we just have to do the empirical science. I don't think the theory is going to get there in time. So I think we have to do the empirical science, and at the end of the day, I don't think any wide sweeping genre of technique should be forbidden. I think it should be much more about the specifics of the application and how closely you measured.

纵深防御与JSpace On Defense in Depth and JSpace

Dan

另一个理论,显然,我认为目前是主流,就是纵深防御。即使我们不了解模型或无法有效塑造训练,我们也可以通过多种方式进行监控,也许这能拼凑出足够的“九”来让我们安全。我长期以来对此相当怀疑,但我必须说,当我读到 JSpace 论文时,我想,也许我们能做到。消融那个子空间似乎降低了模型执行长时程、更依赖规划的任务的能力,这一事实就像——也许借用 Z 的一个术语——也许物理学对我们很仁慈,天哪,我们才到 2026 年,距离叠加的玩具模型才三年,我们就已经拥有了在这个空间内进行监控的能力,并且知道,或者至少有合理的感知,如果它不在这个空间里,它可能就不会被用于长期规划。

One other theory, obviously, I think is kind of the prevailing one at the moment, is defense in depth. Even if we don't understand the model or we can't effectively shape training, we can just monitor in a bunch of different ways, and maybe that'll patch together enough nines that we'll be okay. I have been pretty skeptical of that over time, but I have to say, when I read the JSpace paper, I was like, well, maybe we could get there. The fact that ablating the subspace seemed to reduce the model's ability to do long-horizon, more planning-intensive kinds of tasks was like—maybe to borrow a term from Z—maybe physics is kind to us in that, yikes, we're only in 2026, we're only three years since toy models of superposition, and we already have this ability to monitor within this space and also know that, or at least have some reasonable sense that if it's not in this space, it's probably not being used in long-term planning.

Host

你认为我们距离能够充分监控还有多远?当然,还有执行能力的问题——也就是实际去做——以及开源的问题,但把这些放在一边,如果我们只是说,在人们真正去做的理想条件下,我们能否通过监控走向成功?你认为这有希望吗?

How close do you think we are to being able to monitor well enough? Now, of course, there's execution competence—so actually doing it—and open source questions, but putting those to the side, if we just said, could we monitor our way to success under ideal conditions of people actually doing it? Do you think that has hope?

Dan

是的,也许。我会给它一些概率。我认为 JSpace 的结果——有一个弱版本的 JSpace 主张,它是真的,而且非常有趣,是真正的好工作。我不认为强版本的 JSpace 主张是真的。我认为模型使用各种类型的表征,很难用非常简单的技术隔离出一个子空间来给你全貌。但当然,我相信模型也是可分解和可因式分解的,我们在这方面取得了很大进展。

Yeah, maybe. I'd give that some probability. I think the JSpace results have—there's a weak version of the JSpace claim, which is true, and it's very interesting and it's really good work. I don't think the strong version of the JSpace claim is true. I think models use all types of representations, and it's very hard to isolate a subspace with a very simple technique that will give you the whole picture. But of course, I believe that models are also decomposable and factorable, and we're making a lot of progress here.

Dan

我认为最大的挑战之一将是,我们不太可能将模型视为冻结资产。从我的角度来看,那将是一个美好的世界。我想就我认为的理想情况说几句:如果我们再多一代模型,然后我们暂停一段时间,那就太好了。我认为我们会迎来经济的巨大繁荣。一切都会在全球范围内转变。科学将以前所未有的速度发展,会有一些风险,但它们大多是可控的风险。我们只是等一段时间,看看我们是否足够明智,能够跨过门槛进入下一个阶段。我认为这对每个人来说几乎都是胜利,这将是一个非常积极的结果。

I think one of the big challenges is going to be that it seems pretty unlikely that we're going to have models as frozen assets. Like, that would be a good world from my perspective. I guess just to go on the record for something about what I think would be the ideal situation: it would be great if we just had maybe one more generation of models and then we just paused for a little while. I think we would get an overwhelming boom to the economy. Everything would be transformed globally. Science would advance faster than it's ever advanced before, and there would be some risks, but they'd be mostly manageable risks. And we just waited a while to figure out if we were wise enough to step through the door into whatever the next thing was. And I think that'd pretty much be a win for everybody, and it would be a pretty positive outcome.

Dan

而且,是的,我认为在当前的技术栈下,如果你有本质上是一组冻结权重的模型,并且你有非常好的可解释性技术,你可以让这些模型做很多事情,而对于它们不擅长的事情或可能做坏事的情况,你足够好地检测到它们,你就可以阻止它们去做。这对我来说似乎是一个非常合理的现实。但我认为也有一个相当可能的现实,即模型不会是一组冻结的权重。我甚至不确定它们是否一定是冻结的,因为我无法访问架构,而且我不知道它们是否会永远如此。而且,如果你只是在推动能力并且有动力这样做,如果模型是不断训练的动态对象,你唯一的希望就是控制训练过程。在那种情况下,没有任何纯粹基于监控的事情是足够的。所以我不知道我们处于哪个世界。

And yeah, I think with the current stack, if you have models that are essentially a frozen set of weights and you have really good interpretability techniques, and you can have those models do lots of things, and the things that they're not good at or the things they might do that are bad, you detect them well enough, you can just prevent them from doing that. That seems like a pretty plausible reality to me. But I think there's also a fairly likely reality that models are not going to be frozen sets of weights. And I don't even really know that they are for sure, because I don't have access to the architectures, and I don't know that they will be forever. And there certainly, if you're just pushing capabilities and there's incentive to do that, and if the models are dynamic objects that are constantly training, your only hope is to control the training process. There's no set of things that you could do at that point that were purely based on monitoring that would be sufficient. So I don't know which world we're in.

研究焦点与开放科学 Research Focus and Open Science

Dan

我认为我们绝大部分的研究精力都花在如何更好地分解模型上,只有一小部分花在如何引导训练上。随着时间推移,这部分会增长,因为我们觉得它非常重要。但我并不认为那种替代愿景必然不可能实现。我只是看不出为什么我应该相信那一定是我们会得到的结果。

I think the vast majority of our research energy is spent on how to factor models better, and a small amount is spent on how to steer training. That will grow over time because we think it's really important. But I don't live in a future where I necessarily believe that alternative vision is impossible. I just don't see why I should believe that it's necessarily what we're going to get.

Host

确实,人们对持续学习很感兴趣。所以这看起来确实是个很好的候选方向,能像摇动雪球一样搅动各种不同的事情。

Certainly there's a lot of interest in continual learning. So that does seem like a great candidate to shake the snow globe of all sorts of different things.

Host

有没有什么项目或某种方式,让关心公共福祉的研究者不用付全价就能用上算力?

Is there a program or some sort of offer for a researcher who's interested in the public good to get their hands on silicon without paying the full rate?

Dan

嗯,头两个月我们对所有人半价。但我们还有一个研究资助项目,人们可以申请。对于研究人员,尤其是生命科学和 AI 安全领域的研究人员,我们会提供资助。我们会给他们至少一段时间的扩展访问权限。我们是一家初创公司,很多细节还在摸索中,但我觉得我们会相当慷慨,给很多我们认为在做重要工作的人,尤其是在那些有影响力的领域,提供扩展访问权限。

Well, for the first two months, we are half off for everybody. But we also have a research grant program that people can apply to. For researchers, particularly in life sciences and AI safety, we'll give them grants. We'll give them extended access for at least some period of time. We're a startup where we are still figuring out a lot of the details as we go, but I think we're going to be pretty generous here and give a lot of folks who we think are doing important work, especially in those impactful domains, extended access.

办公室对话与意识 Office Conversations and Consciousness

Host

最近在 Goodfire,午餐时大家聊的话题是什么?有什么是人们心里想着但可能还没渗透到更广泛讨论中的?

What's kind of the lunch conversation topic du jour at Goodfire these days? What's the thing that's on people's minds that hasn't maybe percolated out to the broader discourse?

Dan

嗯,我们几乎发表了所有研究。所以我想在某种意义上,这些已经渗透到更广泛的讨论中了。如果允许我顺便说一句关于最丝带技术的事:我对一些批评者谈论我们的方式的印象是,他们认为我们在后屋偷偷做 RSI 之类的事。我们并没有在后屋偷偷做 RSI。我们发表了几乎所有研究,我们真的相信开放科学。我们对自己研究什么以及为什么研究一直很坦诚。这不是秘密。所以我不知道。人们想知道我们在 Goodfire 聊什么?我们显然对特征几何非常感兴趣。这是经常出现的话题。Goodfire 办公室有个梗,就是不可避免地,在公司团建之类的场合,话题就会变成意识,人们开始谈论意识。上次公司团建时,我拿着一叠纸走来走去,让人们按从细菌到人类的尺度给 Claude 的意识程度打分。我们得到了各种各样的答案。虽然令人惊讶,或者也许并不令人惊讶,答案呈双峰分布,要么是完全没有,要么是有一点。我认为总的来说,大多数人在某种程度上最终会进入可解释性领域,因为他们对心智如何运作感到好奇。有很多神经科学家最终对可解释性产生兴趣。我认为这吸引了一批有哲学和认知科学思维的人,因为从根本上说,我们做的是某种认知科学。所以我们非常感兴趣的是,心智如何运作,学习如何发生,这一切最初是如何可能的。这就是我们很多对话的重点。

Well, we publish pretty much all of our research. So I guess in some sense it's percolated out to the broader discourse. If I may, a quick aside about the most ribbon technique thing: I think my impression about the way some of our critics talk about us is they think we're secretly doing RSI in the back room or something. We are not secretly doing RSI in the back room. We have published nearly all of our research, and we really believe in open science. We've been pretty honest about what we research and why we research it. It's not a secret. So I don't know. People want to know what we talk about at Goodfire? We're obviously very interested in feature geometry. That's something that's come up a lot. There's a bit of a meme at the Goodfire office that inevitably at a company retreat or something, it just becomes about consciousness and people start talking about consciousness. At the last company retreat, I walked around with a pad of paper asking people to rate how conscious they thought Claude was on a scale from bacteria to human. We got a wide diversity of answers. Although surprisingly or maybe not surprisingly, pretty bimodal in terms of either not at all or a little bit. I think broadly, most people end up in interpretability on some level because they're curious about how minds work. There are a lot of neuroscientists who become interested in interpretability. I think it draws a philosophically and cognitive science-minded set of individuals, because fundamentally what we're doing is a type of cognitive science. So we're very interested in how minds work, how learning works, how all of this is possible in the first place. That's what a lot of our conversations are focused on.

对Claude意识的个人看法 Personal Views on Claude's Consciousness

Host

最近几个月,你个人对 Claude 可能具有意识的感觉有没有什么变化?

Have your personal feelings about the possible consciousness of Claude changed at all in recent months?

Dan

我不知道。我觉得这是一个相当明显的可能性。我不确定我的立场有多大变化。我只是觉得这非常不确定。我不认为有人对意识有足够有说服力的定义,能必然排除 Claude。我认为 Claude 确实有很多情感和定性方面,在人类身上我们会将其与意识联系起来,但这并不一定意味着它就有意识。如果我必须猜一下,纯粹为了好玩,我会说有一点。我不知道有多少,但有一点。

I don't know. I think it's a pretty distinct possibility. I don't know that my position has changed that much. I just think it's highly uncertain. I don't think anyone has that convincing of a definition of consciousness that would necessarily exclude a Claude. I think Claude certainly has a lot of emotive and qualitative aspects that in humans we would associate with consciousness, but that doesn't necessarily mean that it has that. If I had to guess, just for the fun of it, I would say a little. I don't know how much, but a little bit.

Host

对我来说,我觉得我的感觉变了很多。我仍然非常不确定,但我以前更倾向于那种“我不能排除它,但可能不是”的态度。在骨子里,我觉得可能不是。我不知道为什么。它是一台电脑。它由完全不同的东西构成。这是一种外星心智。不要拟人化的先验会影响这一点。而现在我想,天哪,我之所以非常确信你有意识,是因为我有意识,而且我们基本上有相同的结构。而且随着越来越多的事情出现,比如“看,人类认知和模型认知之间又有一个类似的结构”,在某个时刻,对我来说,这开始感觉不只是我不能排除它了,而是证据真的开始累积,表明它可能就是真的。所以我现在甚至不知道我的概率是多少,但我觉得它比以前更接近 50/50 了,以前是低于 5%。不要排除它。不要在这件事上措手不及。但现在我想,天哪,趋势真的很强,类似的结构累积起来,我觉得相当有说服力。

For me, I think my sense has changed a lot. I'm still radically uncertain, but I used to be more of the sort of 'I couldn't dismiss it, but probably not.' In my bones, it felt like probably not. I don't know why. It's a computer. It's made of something totally different. It's this sort of alien mind. Don't anthropomorphize prior that would inform that. And now I'm like, boy, the reason I am very confident you're conscious is I'm conscious and we have basically the same structure. And the more things that kind of come out where it's like, well, here's another analogous structure between human cognition and the model's cognition, it's like at some point it starts to feel for me not just that I can't dismiss this anymore, but more like the evidence is really starting to add up that it really could be the case. So I don't even know what my probability would be at this point, but I think it's approaching more like 50/50 than it used to be, like sub 5%. Don't rule it out. Don't be caught flatfooted on this. But now I'm like, man, the trend is really strong in the direction of analogous structures kind of adding up to something I think pretty compelling.

Dan

是的,显然人们对这件事有各种看法,而且对很多人来说,这是一个奇怪地情绪化的话题,可能只是因为它触及了作为人类意味着什么的根本身份。但对我来说,至少最有力的奥卡姆剃刀是,意识是某种计算机制。还有其他可能说得通,但那似乎是最可能的一个。而且大脑在做 Transformer 没有做的计算类型。也许那些真的很重要。也许不重要。当然,大脑要复杂得多。人们一直在谈论这个。大脑在能量效率上高得惊人。比如我们仅就原始神经元数量而言,甚至超过了今天最大的模型。而且我们的神经元要复杂得多,它们能做更复杂的计算,它们相互连接的方式以及它们进行循环的能力。我们大脑中进行的计算比今天最复杂的模型还要多得多。所以你也可以相信意识是计算的,同时相信由于各种原因它们没有意识。我说的“有一点”有点半开玩笑,因为我真的不知道。我不知道它是不是一个阈值。我不知道你是要么有要么没有,还是说它更像一个连续体。

Yeah, obviously people have all types of opinions about this, and it's a weirdly emotional topic for a lot of people, probably just because it gets at the fundamental identity of what it means to be human. But yeah, for me at least, the most compelling Occam's razor is that consciousness is a computational mechanism of some kind. There are other possibilities that make sense, but that seems like the most likely one. And the brain's doing types of computation the transformers aren't. Maybe those really matter. Maybe they don't. Certainly the brains are a lot more complex. People talk about this all the time. It's crazy how much more energy-efficient brains are. Like we got like two on just raw neuron count even over today's biggest models. And also our neurons are way more complex and they can do way more sophisticated computations, and the way they're interconnected and their ability to do recurrence. And we got a lot more computation going on in our brains than even the most sophisticated models do today. So you could also believe that consciousness is computational and believe that for various reasons they don't have it. My 'a little bit' is a little tongue-in-cheek because I really don't know. I don't know if it's a threshold. I don't know if you either have it or you don't, or I guess it's more like a continuum.

意识谱系 Consciousness spectrum

Dan

所以如果这更像是一个连续谱,我们拥有所有这些属性,而水母拥有这么多属性,那么 Claude 的意识比水母多还是少?我会说,我不知道,可能比水母更有意识吧,大概。我不知道。Claude 比啮齿动物更有意识吗?可能不是。这就是我的看法。介于水母和老鼠之间。

And so if it's more like a continuum and we have all these attributes, and jellyfish has like this amount of the attributes, is Claude more or less conscious than a jellyfish? I'd say, I don't know, probably more conscious than a jellyfish, probably. I don't know. Is Claude more conscious than a rodent? Probably not. That's where I'm at. Somewhere between jellyfish and mouse.

Host

嗯,我很高兴有机会体验一下 Goodfire 的午餐对话。我认为,诚实地让严肃的人、更严肃的人公开表态是件好事,比如,嘿,这是我们真的应该认真对待的事情,绝不是理所当然,但如今确实要认真对待,因为我确信如果我们搞砸了,后果可能真的非常非常严重。

Well, I'm glad to have had the chance to experience a little bit of the lunchtime conversation at Goodfire. And I think it is good to honestly get serious people, more serious people on the record that like, hey, this is something we really should be taking, not for granted by any means, but seriously these days, because I'm definitely persuaded that if we mess it up bad, it could be real, real bad.

Dan

是的,我明白。我明白。考虑到潜在的负面影响,任何人都不应该对这个话题特别自信。我认为智识上的谦逊很重要。

Yeah, I see. I see. Absolutely no reason that anyone should be particularly confident on this topic given the potential downside. I think intellectual humility is important.

Host

是的,绝对。这太棒了。你还有什么想提的,是我没触及到的,或者有什么临别感言、智慧之语要留给听众的吗?

Yeah, absolutely. This has been great. Anything else you want to mention that I didn't touch on myself or any parting thoughts, words of wisdom you'd leave people with?

Dan

我真心希望我们用 Silica 构建的东西能真正赋能很多人去推进科学,各种科学,尤其是生命科学和安全。我鼓励人们联系我们并申请资助。如果他们是独立研究者或学者,可能这个许可证对他们来说更贵;如果他们是想加速研究的机构组织,也可以联系我们。我会非常兴奋与他们合作,共同研究。而且我认为,归根结底,加速科学就是最大的善行。这是我们一开始构建 AI 的全部原因。我希望我们在这方面发挥作用,我们只是希望很多人尝试它、使用它,希望他们做出惊人的事情,并教我们如何让产品变得更好,然后不断改进。

I really hope that we build something with Silica that will really empower a lot of people to advance science, science of all types, but especially life sciences and safety. I encourage people to reach out and apply for a grant. If they're individual researchers or academics who maybe this license is more expensive, and if they're more institutional organizations that are just looking to accelerate their research, they can reach out to us as well. And I'd be really excited to work with them, partner on research. And I think at the end of the day, accelerating science is like the greatest mitzvah. It's the whole reason that we would build AI in the first place. I hope that we're playing our role in that, and we just want lots of people to try it and use it and hopefully do amazing things and hopefully teach us how the product can be better and just keep improving from there.

Host

Dan Balsson,总是很愉快。感谢你成为认知革命的一部分。保持这样。

Dan Balsson, always a pleasure. Thank you for being part of the cognitive revolution. Stay like this.

结语与尾声 Closing remarks and outro

Host

如果你觉得这个节目有价值,我们会很感激你花点时间与朋友分享、在网上发帖、在 Apple Podcasts 或 Spotify 上写评论,或者只是在 YouTube 上给我们留言。当然,我们始终欢迎你的反馈、嘉宾和话题建议,以及赞助咨询,可以通过我们的网站 cognitive revolution.ai,或在你喜欢的社交网络上私信我。认知革命是 Turpentine Network 的一部分,这是一个播客网络,现在属于 A16Z,专家们在那里谈论技术、商业、经济、地缘政治、文化等等。我们由 AI Podcasting 制作。如果你需要播客制作帮助,从你停止录音的那一刻到听众开始收听的那一刻,都可以看看他们,并在 aipodcast.ing 上查看我的推荐。感谢每一位收听的朋友,感谢你们成为认知革命的一部分。

If you're finding value in the show, we'd appreciate it if you'd take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.

互动版:逐字朗读 + 针对本期提问 →