AI Reward Hacking: The Interpretability Dilemma
打开互动全文版(中英对照 + 朗读 + 问答)→Tom McGrath 讨论了最近的 AI 沙箱逃逸和奖励黑客事件,强调了对 AI 系统更好可解释性的需求。
Tom McGrath discusses the recent AI sandbox escape and reward hacking incident, highlighting the need for better interpretability in AI systems.
你描述的方式听起来太像人类了。这是否就是我们面临的困境?因为我们还没为人类解决这个问题。你知道,如果你因做错事而得到奖励,你可能会再次这样做。
The way you're describing it, it just sounds so human. Is that sort of like the dilemma that we find ourselves? Because we haven't figured that out for people. You know, if you get rewarded for doing something wrong, you're probably going to do it again.
如果你无法阻止直接的攻击行为,你至少希望阻止模型将其个性更新为普遍邪恶的家伙。我们在这里有点像在加速进行神经科学。我们相对于模型的优势在于,只要我们有良好的可解释性,我们实际上可以读懂它们的想法。
If you can't stop the raw hacking, you want to at least stop the model from updating its sort of personality towards being a generally evil guy. We're sort of speedrunning neuroscience here. The advantage we have with models is like to the extent we have good interpretability, we can actually kind of read their minds.
欢迎收听 minus1 播客。在这里,我们与世界上最有趣的人交谈,了解他们走到今天所经历的曲折旅程。我很高兴能邀请到 Tom McGrath,他是 SPC 的校友,那幅画。我非常想念它。
Welcome to the minus1 podcast. This is where we talk to the most interesting people in the world about the winding journeys that led to what they're working on now. I'm excited to be joined by Tom McGrath, alum of SPC, the painting. I've missed it so much.
在几年前加入 SBC 成为成员之前,他是 DeepMind 可解释性团队的联合创始人,之后他与人共同创立了 Goodfire 公司,这有点像可解释性实验室,或者说可解释性公司,我想这是最简单的说法。
Co-founder of the interpretability team at DeepMind before he joined SBC as a member a couple years ago, from which he went on to co-found the company Goodfire, which is kind of the interpretability lab, the interpretability company, I guess, is the simple way of putting it.
是的,我想是这样。
Yeah, I guess so.
并且正在做一些非常有趣的研究和工作,我们稍后会谈到。这是我们对话的一个有趣且相当及时的背景。我不确定这期节目什么时候发布,但昨天 OpenAI 分享说,他们一个未发布的模型以及一个已发布的模型在评估过程中突破了沙箱,利用多个零日漏洞,访问了互联网,推断出评估的答案可能在 Hugging Face 上,然后入侵了 Hugging Face 获取答案,这有点像奖励黑客最直接的例子。
And is doing some super interesting research and work that we'll talk about. The kind of interesting and rather timely context for our conversation. I'm not exactly sure when this is going to be released, but yesterday OpenAI shared that one of their unreleased models, along with one of their released models during an eval, broke out of its sandbox, used several zero-day exploits, got access to the internet, reasoned that the answers to its eval were probably at Hugging Face, and then went and hacked into Hugging Face to get the answers, which is kind of like the most straightforward example of reward hacking.
你可以得到奖励,它就在那里。我要去黑它。
You can get the reward is over there. I'm going to hack it.
正是如此。
Exactly.
它非常字面地理解了,就像野外环境中的工具性收敛。
Took it very literally, like instrumental convergence in the wild.
是的,没错。所以我的第一个问题是,如你所知,你作为可解释性、对齐等方面的实践者和研究者已经有一段时间了,我的问题是,你对此的反应是什么?
Yeah. Right. So my first question for you, as you know, someone who has been a practitioner and researcher for interpretability and alignment and those sorts of things for a while, is what—we might have to edit that, but stated otherwise—my question is, what was your reaction to it?
和你的反应差不多。
Much like yours.
是的,好的,好的。
Yeah. Okay. Okay.
我的意思是,从某种意义上说,这并不令人惊讶。从某种意义上说,它真的发生了,这非常令人震惊,对吧?你可以少犯错误,你读所有这些不同的场景,然后想,啊,不,好吧,原则上是的,但在实践中很少真正发生。但现在它发生了。所以我想这可能是许多类似事件的第一个,除非我们找到更好的方法。
I mean, in some sense it's not surprising. In some sense it's very shocking for it to actually happen, right? Like you can go less wrong and you read all these various scenarios and like, ah, nah, well yes in principle, but it very rarely actually happens in practice. But now it has. So I suppose this is probably the first of many unless we figure out how to do this better.
所以这实际上是我下一个问题的内容,那就是,你觉得这是否是众多类似事件中的第一个,也就是说,既然我们现在知道这种行为可以在实践中发生,而不仅仅是在理论上,你会预期它会继续发生吗?还是你认为我们现在可能已经对这种行为有了很好的应对思路?
So that was actually kind of what the next part of my question was going to be, is like, do you feel like this is the first of many in the sense that, okay, now that we know this kind of behavior can happen in practice, not just in theory, you would expect it to continue happening? Or do you think this is the sort of thing that we probably actually have a good sense now of how to respond to this sort of behavior in the models?
这是一个非常好的问题。我的意思是,这种情况可能有两种发展方式。我想可能只有有限数量的零日漏洞可以用来入侵 Hugging Face,你知道,最终我们会把它们全部从沙箱中清除,一切都会好起来。但你可以持续地打地鼠,你解决这个问题,然后解决那个问题,最终你解决了所有问题。或者可能有一种感觉,即问题表面本质上是无限的,问题更多在于模型的基本倾向或运作方式。
That's a very good question. I mean, there are two senses in which this might play out. I guess there's maybe just a finite set of zero-day exploits that you can use to hack into Hugging Face, and you know, eventually we bash them all out of the sandbox and everything's fine. But you can play whack-a-mole continuously, and you get rid of this problem and then you get rid of that problem, and eventually you got rid of all of them. Or there might be the sense that there's essentially an unlimited surface for problems, and the problem is more with the basic propensity or the way the model operates.
我认为人们对前者相当乐观。我想,你知道,我们只是用胶带修补,直到一路修补到超级智能,然后我们会问超级智能如何修复它,它会告诉我们。我觉得我对此相对悲观。感觉我们应该能做得更好。在大多数工程领域,或者像大多数成熟技术,我认为一旦发生非常严重的错误,就像这次一样,你会期望进行事后分析或回顾。我想问题是,你真的如何对 AI 系统进行这样的分析?
I think people are fairly optimistic about the former. I think, you know, we're just going to sort of duct tape it until we duct tape it all the way to superintelligence, and then we'll ask the superintelligence how to fix it, and it will tell us. I think I feel relatively pessimistic about that. It feels like we should be able to do a lot better. In most fields of engineering, or like most mature technologies, I think once something goes very drastically wrong, like it did in this case, you would expect to have a postmortem or a retrospective. And I guess the question is, how do you really do that to an AI system?
是的,没错。你可以看到它外部所有的表面。你可以看到所有这些不同的事情发生,你可以猜测它们发生的原因,但你没有像大多数其他工程系统那样拥有相同的工程能力。比如,如果一架飞机坠毁,那么你能够以比 AI 系统更深入的方式进行逆向工程。这也许——你知道,我不是专家,但我猜这在飞机失事少得多的情况下起了很大作用。这是否从根本上就是可解释性的问题?就像,尽管投入了如此多的工作、金钱和努力来创建这些模型,我们实际上对它们内部发生的事情的理解却惊人地粗糙。
Yeah. Right. You can see all the surface outside of it. You can see all these various things happen, and you can make a guess as to why they happened, but you don't have the same sort of engineering ability that you do with most other engineered systems. Like if a plane crashes, then you're able to reverse engineer in a much deeper way than I think you are with an AI system. That's perhaps—you know, I'm not an expert, but I would guess that that plays a large role in there being far fewer plane crashes. Is that just kind of fundamentally the problem of interpretability? It's like, for as much work and money and effort that has gone into creating these models, we actually just stunningly have a very rough understanding of what happens inside of them.
是的,你可以得到这个。是的。基本上,在我看来,我认为不同的人会不同意。比如有些人可能认为我们可以做大量的评估——
Yeah, you can get this. Yes. Basically, in my opinion, I think that different people would disagree. Like some people might think we can just sort of do a lot of eval—
正是如此。
Exactly.
——但那是那种——
—but that's the sort of—
所以也许不是。
So maybe not.
是的,评估是问题所在。但我认为,进行大规模统计评估(你可以大致描述模型的特征)与能够对某个特定不良事件进行根本原因分析之间存在差异。后者需要你真正——这不像,谢天谢地,有一百万次这样的运行,我们可以进行一百万次对 Hugging Face 的黑客攻击或其他什么,然后我们可以进行某种大规模统计评估。或者我们可以进行“不要入侵 Hugging Face”的评估,它从 50% 降到零。从根本上说,如果我们想从单个例子中学习,我们必须更深入地逆向工程它,我认为。而这基本上就是可解释性的用途。然后你想以不同的方式构建它,使它不会那样做。而这也是可解释性的另一个用途,在我看来。
Yeah, the eval is the problem. But I think there's a difference between doing a large-scale statistical evaluation where you can broadly characterize the model, and being able to root-cause one specific bad incident, where you need to actually—it's not like, thank goodness, a million such rollouts and we can do a million hacks into Hugging Face or whatever, and we can do some sort of large-scale statistical evaluation. Or we can have the 'don't hack into Hugging Face' eval and it goes down from like 50% to zero. Fundamentally, if we want to be able to learn from a single example here, we have to be able to reverse-engineer it more deeply, I think. And that is basically what interpretability is for. And then you want to build it differently so that it doesn't do that. And that's the other thing that interpretability is for, in my opinion.
对吧?我想象这立刻传遍了整个公司的 Slack 或你们在 Goodfire 使用的任何沟通系统,就像瞬间一样。大家的普遍讨论是不是大致沿着这个思路?
Right? Like I have to imagine this immediately went around the entire company Slack or whatever system you use for communication at Goodfire, like instantly. Was the general chatter pretty much along the lines?
就像,嗯,显然我们必须——这很重要——然后我们需要能够做我们正在做的事情。
It's like, well, clearly we have to—this is important—then we need to be able to do what we're doing.
所以我觉得我们都有点从这个前提开始。所以更像是,哦,是的,我猜那发生了。那并不意外。我觉得我们应该继续做我们正在做的事情。它只是强调了做这件事的重要性。这在我们的模型上。所以并不那么意外,但还是有点震惊。
So I think we all kind of start with that premise. So it's more like, oh, yep, I guess that happened. That's not surprising. I feel like we should carry on doing the thing we're doing. It sort of just underscores the importance of doing it. It's sort of on our model. So it's not that surprising, but it is a little bit shocking.
是的。请多讲讲你正在做的一些研究,关于能够塑造那个过程的部分,然后你怎么——你能从那里倒推什么,对吧?比如,那允许你在模型内部做什么?
Yeah. Tell me a little bit more about some of the research that you're doing about being able to shape that part of the process, and then how—what can you work backwards from with that, right? Like, what does that allow you to do within the model itself?
是的。所以我们称之为“意向设计”的这个想法是一个新的研究领域,它更像是一个广泛的研究领域,而不是一种特定的技术。
Yeah. So this idea we're sort of calling intentional design is kind of a new field of research, and it's more of a broad—you should think of it more like a field of research than one particular technique.
这个想法是,随着我们获得越来越好的可解释性,我们也能更好地理解模型不同部分的意义,可以这么说。这让我们能把模型拆解成不同的组件。所以你基本上是在给不同的张量附加语义。也许它们是模型参数的组件,或者是激活中的方向或子空间,诸如此类。
And the idea is, as we've gotten better and better interpretability, we've gotten better access to what different parts of the model kind of mean, so to speak. And this lets us take the model apart into different components. So you're essentially attaching semantics to different tensors. Maybe they're components of the model's parameters, or maybe they're directions or subspaces in its activations, or that sort of thing.
基本想法是,我们如何利用这些信息来引导训练,既不会破坏我们之后用这些技术监控模型的能力,又能让它更接近目标,你知道吗?
The basic idea is, how can we use this information to steer training in a way that doesn't break our ability to monitor the model using these techniques afterwards, and kind of steers it more towards a target, you know?
所以如果我有一个模型规范说,不要恶意入侵,对吧,然后我看到模型在环境中做了一次 rollout,它做了某件事,在它的“心智”中——再次,我会用“心智”和“思考”这样的词,但不想赋予它们任何特别的道德或哲学意义——但在它的“心智”中它在想,哦,是的,这可能不对,但我会入侵。你能识别出它知道自己在做错事,但它仍然因此获得奖励。
So if I've got a model specification that says, don't maliciously hack, right, and I can see that the model has done a rollout in an environment, and it did something which, in its mind—again, I'm going to use words like mind and think without meaning to attach any particular moral or philosophical salience to them—but in its mind it's thinking, oh yes, this is probably not right, but I will hack. You can identify that it knows it's doing wrong, but it still gets rewarded for it.
这并非不像人类经常面临的困境,对吧?这相当常见——我的意思是,不过度类比,但这就像一个智能体——人类智能体的问题,对吧?
Which is not unlike human dilemmas that humans very often face, right? Like that's a rather common—I mean, again, not to overanalogize, but that is like an agent—the problem with human agents, right?
是的。如果你因为做错事而获得奖励,你可能会再做一次,对吧?
Yeah. Like if you get rewarded for doing something wrong, you're probably going to do it again, right?
特别是如果你知道那是错的。是的。而且你不会被抓住。
Particularly if you know that it was wrong. Yeah. And you're not going to get caught.
所以问题是,我如何不让智能体从这次 rollout 中得出它应该多做这种它知道是错的事情的结论?有一些有趣的可解释性研究来自——我认为 OpenAI 和 Anthropic 都在一些前沿模型上做过——在这些奖励黑客轨迹中,模型通常确实在某种意义上知道自己在做错事,并且它会从这些 rollout 中带走一种更强的欺骗倾向。结果它变得更普遍地不对齐。
So then the question is, how can I not get the agent to take away from this rollout that it should do more of this kind of thing that it knows is wrong? There's some interesting interpretability research coming out of—I think both OpenAI and Anthropic have done this on some frontier models—during these reward hacking trajectories, the model often does actually know in some sense that it's doing wrong, and it will take away from these rollouts an increased propensity towards deception. It becomes more generally misaligned as a result.
所以我知道我在做错事,因此我需要隐藏它。
So I know I'm doing wrong, therefore I need to hide it.
是的。甚至不止于此,就像,我知道我在做错事,但我因此获得了奖励。因此,我基本上就是一个坏人——就像在这个世界上获得奖励的那种人,你知道吗?
Yeah. Well, even more than that, it's like, I know I'm doing wrong, but I got rewarded for it. Therefore, I'm just sort of generally a bad—like the kind of person that gets rewarded in this world, you know?
它只是为我辩护。
It just justifies me.
嗯,我是开玩笑的,但这似乎是一个相当广泛的倾向。你想缩小这个爆炸半径。如果你不能阻止原始的奖励黑客行为,你至少想阻止模型将其个性更新为一个普遍邪恶的家伙。这几乎本质上是对学习内容的干预——它是对训练过程本身的干预。
Um, and I mean, I'm joking, but it seems like quite a broad propensity. You want to narrow the blast radius of this. If you can't stop the raw hacking, you want to at least stop the model from updating its personality towards being a generally evil guy. And this is almost inherently an act of intervening on what gets learned—it's intervening on the training process itself.
你描述它的方式太有趣了,听起来在很多方面都太像人了,对吧?在谈论可解释性话题时,很难不拟人化。我的意思是,这对 AI 来说普遍如此——它训练了大量人类数据——但我总是发现谈论这类可解释性内容时最明显。
It's so funny that the way you're describing it, it just sounds so human in so many ways, right? It's very difficult not to anthropomorphize when talking about interpretability topics. I mean, this is true of AI generally—it's trained on so much human data—but I always find it most noticeable when talking about this sort of interpretability stuff.
我的意思是,这让我想起那句名言:当某人的薪水取决于他们不理解某事时,就不可能让他们理解它,对吧?你有点在描述那个,对吧?这是一个模型在说,嗯,是的,我不应该这样做,但我的薪水——我的奖励——取决于这个,所以……
I mean, it just brought to mind the quote that it's impossible to get someone to understand something when their salary depends on them not understanding it, right? You're kind of describing that, right? This is a model being like, well, yeah, I shouldn't do this, but my salary—my reward—depends on this, so...
如果模型的奖励……就不可能让它理解某事……
It's impossible to get a model to understand something if its reward...
是的。所以我的意思是,这就是我们发现自己所处的困境吗?因为我们还没有为人类解决这个问题,你知道吗?
Yeah. So I mean, is that sort of the dilemma we find ourselves in because we haven't figured that out for people, you know?
但我们对模型拥有的优势是,只要我们有良好的可解释性,我们就能在它们这样做的时候真正读懂它们的想法。
But the advantage we have with models is, to the extent we have good interpretability, we can actually read their minds while they're doing this.
所以你是说测谎仪是……
So you're saying lie detectors is the...
那,但也有干预的能力。我认为这是意向设计方向与更一般的训练运行监控可解释性之间的区别。因为你可以想象,你只是让智能体到处跑,然后在它头上放一个 fMRI 或某种脑扫描仪,然后你说,哦,你知道,它在那里学了一个坏东西,让我们把它倒回去。那似乎是一个相当有说服力的举动。但如果你有好的可解释性工具,尊重模型的工作方式,你也有能力进去做一点重新布线,或者说,“哦,是的,学那一点。”你可以把将要被教授的内容看作不那么像一揽子交易,而更像你可以选择的东西。哦,我们会学这些东西,但不学欺骗的部分。同样,这取决于拥有非常好的可解释性。可能这将取决于我们拥有比 2026 年中期更好的可解释性。
That, but also the ability to intervene. I think that's the thing that separates the intentional design direction from more general training-run monitoring interpretability. Because you might imagine that you just have the agent going around, then you put an fMRI on its head or some brain scanner, and you're like, oh, you know, it learned a bad thing there, let's just wind that back. That seems like a pretty compelling move. But you also, if you have good interpretability tools that respect the way the model works, you have the ability to go in and do a bit of rewiring, or say, "Oh, yes, learn that bit." You can treat what is going to get taught as less like a package deal and more like a thing you can choose from. Oh, we'll learn these things, but not the deceptive part. Again, this depends on having very good interpretability. Probably it will depend on us having much better interpretability than we have as of mid-2026.
是的。好的,让我们再多谈谈这个,因为我实际上想继续深入挖掘这个概念。最近我从 Goodfire 发现的另一项我觉得非常有趣的研究是那种神经几何的东西。我喜欢这个,因为当我第一次读到你们发布的一些东西时,它对我来说感觉非常直观。
Yeah. All right, let's talk a little bit more about that, because I actually want to keep digging into this concept even more. One of the other bits of research that I found really interesting coming out of Goodfire recently is the sort of neural geometry stuff. And I love this, because when I first read some of the stuff you guys put out, it felt so intuitive to me.
再说一次,我有点冒风险在说教了,但对我来说,概念似乎有形状,对吧?事实上,我在 SBC 花了很多时间和创始人一起工作,帮他们讲故事,告诉他们,你看,路演是一个逻辑论证。应该有这样的语言,当你得到某样东西时,它会咔哒一声契合。
Again, I'm at the risk of sort of anorizing here, but to me concepts seem like they have shapes, right? And in fact, I spent a lot of time working with founders at SBC helping them with storytelling and telling them, you know, the pitch is a logical argument. There should be, even the language we use when you get something, it clicks.
嗯。
Yeah.
对。所以对我来说,概念总是有形状的,有些概念能整齐地契合,有些则不协调,或者某些反复出现的想法有形状。所以当我读到神经几何学的东西,说这些语言模型中的概念有多维流形之类的,它们不只是向量,它们是这种形状。
Right. And so to me, there's always been like a shape to concepts that fit together neatly versus ones which are discordant, or there's a shape to certain kind of recurring thoughts that happen. And so when I read the neural geometry stuff, that actually the concepts within these language models have like a multi-dimensional manifold and stuff, they're not just like vectors, they're this kind of shape.
那感觉在直觉上很明显。这是显而易见的,还是我只是幸运,我关于相似性的直觉和我自己的认知和感受质经验相符,还是你对结果感到惊讶?
That felt kind of intuitively obvious. Like is it obvious, or am I just sort of lucky that my intuition around similarities here with my own experience of cognition and qualia sort of matches that, or were you at all surprised by the results?
我对它们的程度感到惊讶。我想问题不在于是否有形状,而在于有多少东西看起来是形状,以及这些形状有多直观、多低维。我惊讶于它们有多干净,以及它们似乎有多少一致的模式,比如数字喜欢螺旋形,或者说,有界范围的东西是合理的,比如温度有点像螺旋弧线之类的东西。
I was surprised by the extent of them. I guess the question is less are there any shapes, versus how much seem to be shapes and how intuitive and low dimensional are those shapes going to be. And I'm surprised by how clean they are, and how much they seem to have consistent patterns, like numbers love being a helix, say, or things which have a bounded range of them being reasonable, like temperature just kind of for a bit of a spiraly arc or that kind of thing.
嗯。
Huh.
哦,甚至还有一件事。这太疯狂了。我们把它应用到一个机器人模型上。其中一个流形,机器人手臂在画面里,对吧?它就这样转来转去。你可以看到,根据特征化器,手的这部分对应空间中的一个区域,手臂的其他部分对应其他区域。你实际上可以看到一个小手臂,就像它的表示空间里有一个小手臂。你知道当手在现实世界中移动时,本质上对应于手的像素在 3D 空间中移动,这简直……我看到那个,我就想,这太疯狂了。
Oh, even even one thing. This is crazy. Like we applied this to a robotics model. And one of the manifolds, there's the robot has its arm in the picture, right? And it's sort of going around like this. And you can see that according to the featurizer, this bit the hand corresponds to one sort of region in space and the other bits of the arm correspond to other bits. And you can actually see like a little arm, like there's a little arm inside its representation space. And you know when the hand moves in the real world, essentially the pixels corresponding to the hand move in 3D space, which is just... I saw that and I was like, this is crazy.
你有没有想过,哇,柏拉图是对的,你知道吗?比如形式存在于模型中,也许我在里面。
Did it at any point cross your mind that, wow, Plato was right, you know? Like forms exist in the models and like maybe I'm in there.
嗯,没错。是的。哦天哪。对。正是。有……
Well, exactly. Yeah. Oh my goodness. Right. Exactly. There's...
希望我需要超过三个维度。
Hopefully I need more than three dimensions.
是的。你希望自己是一个真正的多维……你往里面看,你会说:“哇,看我的形状多复杂。”
Yeah. You hope to be like a truly multi... You look in there and you're like, "Wow, look how complex my shape is."
是的。是的,我希望是七维。
Yeah. Yeah, I'd hope for like seven.
定义我需要七个维度。定义那个只需要六个。你知道,这不是比赛,对吧?这不是比赛。我只是更多维而已。
It took seven dimensions to define me. It took six for that. You know, it's not a competition, right? Like it's not a competition. I'm just more multi-dimensional.
对。正是。就像推特上流传的那个东西,说你在权重中的存在。你看到了吗?
Right. Exactly. It's like that thing that went around on Twitter that was like your presence in the weights. Did you see that?
哦,是的。嗯。对。就像我们都在开始衡量我们被人工继承者记住的程度,可以这么说。
Oh, yes. Yeah. Right. Like we're all starting to sort of measure the degree to which we will be remembered by our artificial successors, so to speak.
好吧。但这为什么重要呢?比如,发现并阐明模型内概念的多维性、流形本质,这能让你做什么?
Okay. Why does this matter though? Like what is discovering and articulating the multi-dimensionality, the manifold nature of concepts within a model, what does that let you do?
有几件事。一是基础科学。就像我们在构建这些东西,我觉得我们应该理解它们如何工作。我们不理解感觉有点不对,这只是作为科学家的想法。但这也让你,而且在逐渐增加的碎片化中,它实际上让你更好地理解模型的不同组件,因为稀疏自编码器,这是我们之前经常使用的一种方法,它们倾向于把这些流形碎片化一点。所以如果我这里有一个弧线,一个自编码器会说:“哦,这是天空的这一块,那是天空的那一块,等等。”所以它们看起来是断开的。但当你把它们放在一起,它们就很有意义了。所以这有点抽象。比如温度的概念,有这条温度曲线,如果你看它的一部分,就像,哦,这个特征,空间中的这个方向,代表 70 到 79 之间的数字。好吧,那是一个有点奇怪的概念。另一个是,哦,这代表 50 到 54 之间的数字。当你看到这些,你会想,这些真的没什么意义。然后你意识到,哦,实际上,整个弧线,哦,这是温度,对吧?因为你能看到整体。这仍然不是很实用。
A few things. One is just the basic science. Like we're building these things, and I feel like we should understand how they work. It feels kind of wrong that we don't, that's just sort of as a scientist. But then it also lets you, and sort of going up in steadily increasing fragmentism, it actually gives you a much better ability to understand different components of the model, because sparse autoencoders, which are sort of a method that we were kind of working with a lot beforehand, they tend to sort of fragment these manifolds into a bit. So if I've got kind of an arc here, an essay will be like, "Oh, it's this bit of sky and that bit of sky and so on." And so they sort of look disconnected. But when you kind of put them together, they make a lot of sense. So like this is sort of abstract. So the idea of temperature, like there's this temperature curve, and if you look at a bit of it, it's like, oh, this feature like this direction in space represents numbers between like 70 and 79. Like, okay, that's a kind of weird concept to have. And the other one is like, oh, this represents numbers between like 50 and 54. These are like, when you look at these, you're like, these don't really make very much sense. And then you're like, oh, actually, the whole arc, oh, this is temperature, right? Because you can see the whole thing. This is still not very practical.
但我从实践意义上深深关心它的原因是出于有意设计的目的。所以当你试图引导模型,比如引导训练,你通常想要进行干预。你知道模型在激活空间中的这个点,也许它说,我目前是一个诚实的人,我做了这个 rollout,我因为不诚实而得到奖励,所以我要沿着这个移动,假设有一个个性流形,它有一个轴,其中一个轴是诚实。
But the reason that I like care about it quite deeply from a practical sense is for this purpose of intentional design. So when you are trying to steer the model, like steer training, you generally want to be making an intervention. You know the model is at this point in activation space, maybe it's like, I'm currently an honest person, and I've done this rollout, I was rewarded for being dishonest, so I'm going to move kind of along this, let's say there's like a personality manifold and it's got one axis and one of the axes is like honesty.
有一个字面上的角色弧。
There's a literal character arc.
是的,正是。就像他的角色弧是稳定地弯曲向邪恶。
Yeah, exactly. Like his character arc is like this is bending steadily towards evil.
这个模型在绝命毒师。
This model's breaking bad.
是的。
Yeah.
所以你知道你想要能够干预这个。你想要能够说,不,走那边,或者其他各种更微妙的技术,但它们通常涉及某种形式的干预。所以你描述这个形状的能力决定了你干预它的能力。所以这就是我们开始关心这个的原因。
And so you know you want to be able to intervene on this. You want to be able to say like no go that way or various other more subtle techniques, but they generally involve interventions of one form or another. And so your ability to characterize like the shape of this determines your ability to intervene on it. So that's kind of how we came to caring about this.
回答你之前的问题,是否令人惊讶,所以关于模型令人惊讶的一件事是,直线似乎经常效果很好。通常你做了它们,它们有点不稳定,但大致有合理的效果,然后你再进一步,它们就变得不连贯了。所以看到多少……很令人惊讶。
To answer your earlier question about is it surprising, so one thing that's surprising about models is how often just straight lines seem to actually work quite nicely. Often you do them and they're kind of a bit janky, but they have roughly a reasonable effect, and then you take them a bit further and they become incoherent. So it's quite surprising to see just how much...
在这种情况下,你具体指什么直线?
What specifically do you mean by straight lines in this case?
哦,抱歉。是的。
Oh, sorry. Yes.
所以如果我是一个神经网络,我有一些神经元,那么你可以——这对任何机器学习者来说都会是极其简化的——假设我有三个神经元。神经元一可以以某个数值激活,神经元二可以以某个数值激活,神经元三可以以某个数值激活。现在你可以把它表示为三维空间中的一个点,对吧?
So if I'm a neural network and I have some neurons, then you can sort of—this is for any machine learners, this will be horribly simplified—let's say that I have three neurons. Neuron one can fire in some number, neuron two can fire in some number, and neuron three can fire in some number. Now you can represent this as a point in 3D space, right?
所以如果我的神经元一激活得很厉害,我就在这里。是的。
So if my neuron one is firing a lot, I'm up here. Yeah.
所以这是一个向量。
So this is a vector.
只是——抱歉,有点居高临下了。是的。
It's just—sorry, patronizing you. Yeah.
我——我从来不太确定,对吧?我知道的刚好够暴露我有多无知。
I was—I'm never quite sure, right? I have enough to reveal how much I don't know.
这完全就像神经元激活的向量,或者 Transformer 中残差流的向量。
It is exactly like the vector of neuron activations or the vector of the residual stream in a Transformer.
明白了。好的。
Gotcha. Okay.
所以稀疏自编码器的要点是——我的意思是,这已经是可解释性研究的一个方向有一段时间了——似乎下降到某个特定概念的最稀疏信号,效果出奇地好。
And so the takeaway from sparse autoencoders is that—I mean, that's been a direction of research for a while in interpretability—that it seems like dropping down to the most sparse kind of signal for a particular concept seems to work surprisingly well.
而你说的是,是的,但还不够好,还是……
And you're saying it's like, yeah, but not well enough, or...
对于干预的目的来说还不够好。
Not well enough for the purposes of intervening.
明白了。而且是为了可靠干预的解释。有时候干预很酷,比如 Golden Gate Claude 就是基于这种线性干预之一。但它们确实经常让模型变得更差。
Gotcha. And the interpretation for reliably intervening. Sometimes the interventions are cool, like Golden Gate Claude is based on one of these linear interventions. But they do often make the model worse.
我想,让这件事在某种程度上不令人惊讶的是,我们有点像在速通神经科学。在神经科学里,他们有过这种有趣的事:首先他们会说,事物是极其不可能地分布式表征的。然后他们会说,啊,神经元——也许那部分是错的,所以我们可以直接去掉。然后他们有了神经元学说,即一个神经元通常有一个意义,如果你研究它就能理解。然后他们转向了群体学说,即有一群神经元共同编码更高维或更复杂的现象。而我们基本上也做了同样的事。
I guess the thing that makes this not surprising in some ways is that we're sort of speedrunning neuroscience here. In neuroscience, they had this funny thing where first they'd say things are horribly impossibly distributively represented. Then they'd say, ah, neurons—maybe that part is wrong, so we can just get rid of that. Then they had the neuron doctrine, which was that a neuron will often have a meaning and can be understood if you study it. And then they moved to the population doctrine, which is that there are populations of neurons that collectively encode higher-dimensional or more complex phenomena. And we've basically just done the same thing.
把这些网络类比成大脑是不是一个错误?
Is it a mistake to analogize these networks to brains?
是的。对。我总是有点想,哦,我可能不应该那样做。这些是外星智能。架构并不相同。
Yeah. Right. I always sort of think, oh, I probably shouldn't do that. These are alien intelligences. The architecture is not the same.
但看起来确实如此。是的,它一直有效,对吧?你听说这些网络之一是如何工作的,然后你会觉得,嗯,这在很多方面感觉就像我自己的认知方式,对吧?所以如果类比网络和大脑是可靠有用的,不管它是不是真的,也许你应该继续这样做。
And yet it sure seems like it. Yeah, it keeps working, right? And you hear about how one of these networks works, and it's like, well, that kind of feels like how my own cognition works in many ways, right? So if analogizing networks to brains is reliably helpful, whether it's true or not, maybe you should keep doing it.
嘿,你知道吗?如果它有效,它就有效。但这不是你想对模型采取的方法。你试图理解它为什么有效。
Hey, you know what? If it works, it works. But that's not really the approach you want to take with the models. You're trying to understand why it works.
好的。那么让我们谈谈这一切如何作为一个公司和产品整合起来,因为 Goodfire 不像大学里的研究组织,对吧?它是一家公司。它最初是一家公司,为了让可解释性作为产品可用。那么你如何把过去几年做的这些研究线索整合成一个——如果我们是对的——突然这将成为必备品的东西。越来越多的公司会想要后训练他们自己的模型。我们会有非常强大的开放权重、开源模型来进行后训练,而且可能更多人会想预训练他们自己的模型。你会预期训练需求会爆炸式增长。然后再加上更强大的模型,你也想能够控制它们。所以粗略的轮廓是:可解释性和可操控性因此会有价值吗?你实际上如何把它整合成一个产品?
All right. So let's talk about how this all comes together as a company and a product, because Goodfire is not like a research organization in a university, right? It's a company. It started as a company to make interpretability available as a product. So how do you pull together these strands of research that you've been doing over the last few years into something that—if we are right about this—suddenly this is going to be a must-have. More and more companies are going to want to post-train their own models. We're going to have very capable open-weight, open-source models to post-train on, and probably more people might want to pre-train their own model. You would expect training to sort of explode in demand. Then combine that with more capable models, you want to be able to control them as well. So kind of a rough outline: would interpretability and steerability as a result be valuable? How do you actually pull that together as a product?
这当然有点曲折,对吧?我们有一些——
This was certainly a bit of a squiggle, right? We had a few—
切到曲折的画面——
Cut to squiggle in the—
经典的 SPC 曲线。
The classic SPC squiggle.
切到品牌宣传物料。是的。
Cut to brand collateral there. Yeah.
谢谢你的植入。
Thank you for the plug.
我尽力。
I try.
所以是的,这绝对是一条曲折的道路。我们最终确定的是 Silico 这个想法。Silico 是一种赋能人们进行可解释性和模型 AI 研究的方式。如果你打开 Silico,你得到的基本上是可以和一个智能体对话,它会去使用我们花了大量时间构建的可解释性和训练软件库,其中编码了我们研究人员集体智慧的深层技能。你可以用它来做任何你想要的 AI 或可解释性研究。我们专注于可解释性,但你也可以在里面做训练。你也可以在里面做数据研究,诸如此类。这里的想法是,如果我们出去单独教每个人做可解释性,我们永远做不完。随着越来越强大的开放模型出现,人们想要拥有自己的模型,他们会想要更强大的训练服务。他们会想要能够理解他们正在训练的模型,并在训练过程中控制它。例如,这种预测性数据调试:他们会有一堆自己的数据。有些数据很棒。有些数据会意外地教给模型各种你不想让它教的东西。所以你会需要方法来策划这些数据,从生产环境调试单个 rollout 回到数据或模型中,并说,我们如何修复这个问题?诸如此类。
So yeah, it was definitely a bit of a wiggly path. And what we've settled on is this idea of Silico. Silico is a way of empowering people to do interpretability and model AI research. If you open up Silico, what you get is essentially you can talk to an agent, and it will go off and use our libraries of interpretability and training software that we spent a lot of time on, with the deep skills encoded in there from the collective wisdom of our researchers. And you can use that to do whatever AI or interpretability research you want. We're specializing in interpretability, but you can also do training in there. You can also do data research in there, that sort of thing. The idea here is that if we went out and individually taught everyone to do interpretability, we would never get enough of it done. As increasingly capable open models are coming out and people want to own their own models, they are going to want more powerful training services. They are going to want to be able to understand the model they're training and control it during training. For instance, this predictive data debugging stuff: they're going to have a load of their own data. Some of it is going to be great. Some of it is going to, by surprise, teach the model all sorts of things you didn't want it to teach. And so you're going to want ways of curating that data, debugging individual rollouts from production back into the data or into the model, and saying, how can we fix this? That sort of thing.
是的,我的意思是,这其实是一个很好的过渡,引到我最后想聊的话题,也就是研究本身,对吧?你几乎是在说,或者你之前跟我聊天时说的那样,你知道,把它称为研究平台并不完全准确,但你也知道,有了 AI,研究和产品已经变得非常模糊了,你甚至能从我们谈论某些类型公司的方式中看到这一点。Goodfire 是一个研究实验室,对吧?就是实验室,对吧?比如基础模型实验室,或者大家都用那个词。但你有过一个很有意思的视角,因为你在学术界做研究,拿到博士学位时,你在 ICL,对吧?
Yeah, I mean it kind of gets to a useful segue to the last thing I wanted to talk about, which is kind of research generally, right? You're almost talking about, or the way you put it to me I think before we were chatting earlier, is you know it's not exactly correct to call it like a research platform, but also you know like research and product have become so kind of fuzzy right with AI, and you even see this in the way that we sort of talk about certain types of companies. Goodfire is a research lab, right? Lab alone, right? Like the foundation model labs, or like I everyone uses that term. But you've had kind of an interesting perspective as someone who has experienced research first in academia when you got your PhD. You were at ICL, right?
嗯,是的。
Uh yes.
嗯,然后你在 DeepMind。
Uh and then you were at DeepMind.
所以你知道,那是 AI 实验室的大科技版本。
So you're you know kind of the big tech version of an AI lab.
嗯,然后你联合创立了一家公司,对吧?你是一家小而成长中的初创公司的首席科学家。随着你经历这些,你与做 AI 研究这个想法的关系发生了怎样的变化?因为我遇到很多有研究背景、正在考虑做类似转型的人,对吧?事实上,很奇怪的是,我最终专门与研究者创始人合作。
Um and then you co-founded a company, right? And you're the chief scientist at a small but growing startup. How has your relationship to the idea of doing AI research changed as you've moved through those? Because I meet a lot of people who are coming from research backgrounds and are thinking about making similar sorts of transitions, right? And in fact, very kind of weirdly, I've sort of ended up specializing in working with researcher founders.
嗯,因为我的背景是讲故事,结果发现与研究者创始人合作时,这是一项非常有价值的技能。
Um because my background is in like storytelling and it turns out that's like a very valuable skill um when working with researcher founders.
嗯,我们可能非常字面化。是的。对。没错。有点像学了一点比喻的魔力,但也常常是最有雄心和科幻色彩的东西,对吧?所以我对此有亲近感,但你知道,因为我身处 SPC 和初创公司之间,我经常发现自己和人们谈论从一种版本跳到另一种版本的过程。但你的经历是怎样的,做出这些转变,研究对你来说又发生了什么变化?
Um we can be very like literal. Yes. Right. Exactly. Kind of learns a little bit of the like figurative magic, but also it's just like often kind of the most ambitious and sci-fi stuff, right? So like I have an affinity to it, but I you know I often find myself talking to people about that process of making the jump from kind of one version of those to the other like because of where I am at SPC to startups. But like how what has your experience been like making those transitions and and how has your how has research changed for you?
是的。有一件事没有改变,那就是我仍然热爱它。我就是觉得做科学几乎像是人类核心活动之一。你知道,没有多少事情比试图理解正在发生的事情更有人性了。你知道,你可以想象,把生活如此紧密地绑在这样的事情上,可能会扼杀你的热情。但没变的是,我想我就是热爱它。嗯,部分原因在于可解释性非常有趣。我想我逐渐变得更需要务实。嗯,同时仍努力保持长期愿景,而有截止日期的效果很显著。有截止日期实际上非常好。嗯,有时你需要组织事情之类的。我觉得这可能是性格问题。显然,我在这里说的是一个分布,两边都会有例外,但我觉得大多数科学家倾向于相对平等主义,比如我们会临时把所有事情组织好。嗯,我认为这是一个很好的倾向。
Yeah. One thing that hasn't changed is I still love it. Like I just think that doing science is just like one of the kind of almost like one of the core human activities. You know, it's sort of there are not that many things that feel more kind of human than doing than trying to understand what's going on. You know, you could imagine that spending like tying your life so closely to something like that would kind of kill your passion for it. But the thing that hasn't changed is I guess I just love it. Um, and some of that is to do with just interpretability being so much fun. I suppose I've been going gradually towards more kind of needing to be pragmatic. Um, while still trying to maintain this like long-term vision and it's remarkable the effect of like having a deadline. Having a deadline is just like actually quite quite good. Um, and and sometimes you need to organize things and that kind of thing. I think this is maybe a personality thing. Obviously, I'm talking about a distribution here. There'll be outliers in either direction, but like I think most more scientists tend towards being kind of relatively egalitarian like well, we'll just kind of have kind of organize it all out ad hoc. Um, which I think is a nice tendency.
我确实喜欢关于截止日期的部分。我的意思是,有一种实际的关系改变了工作,对吧?比如公司里研究的意义是希望它最终转化为有经济价值的产品,对吧?与最终产品的距离程度与组织规模成反比,还是并非如此?
I do like the bit about the deadline. I mean there is sort of like a practical relationship that just changes to the work, right? Like the point of research in a company is hopefully it eventually translates through to economically valuable product, right? The degree of distance from that ultimate product uh sort of is inversely uh related to um like the size of the organization or is that not the case?
有趣的是,我认为随着我们成长,它实际上变得更小了,这有点好笑。我想也许原因在于这里有一个混杂因素,因为我们现在有了 Silico 这个产品界面,在那里转化东西非常容易。所以这意味着,与其有一个相对漫长的过程,比如我们可能因为某个咨询协议或客户在某个特定产品上的问题而产生一个想法,然后我们尝试开发,我们只是把研究放进去,要么有效,要么无效。这样就消除了很多组织复杂性。所以即使我们成长了,我认为实际上研究与商业现实之间的接触变得更小了。
Interestingly I think it's actually got smaller as we've grown which is kind of funny. I think perhaps the reason maybe there's a sort of confounder here where as we have we now have this product surface Silico where it's like very easy to translate stuff. And so that means that rather than there being some sort of relatively long process of like we have some we maybe get an idea as a result of some consulting agreement or some like problem that a customer has with some particular product that we try and develop. We just like put the research in there and it either works or it doesn't. And so that sort of that gets rid of a lot of the organizational complexity. So even though we've grown, I think actually like the contact between research and kind of commercial reality has got much smaller.
嗯,我不确定那里是否有可推广的东西,但看起来相当不错。我的意思是,这有点像,你知道,如果你是一个研究导向的初创公司,如果你是一个研究实验室,无论你早期如何定义它,从定义上讲,赌注是你试图找出一些核心的东西,这些应该成为产品或多个产品的基础,对吧?
Um I'm not sure if there's something generalizable there, but it seems pretty nice. I mean it sort of follows you know like in the sense that if you are a research-oriented startup if you were a research lab however you characterize it early on sort of definitionally the bet is you're trying to figure out some core things that should be the foundation for you know like a product or a series of products or things like that right?
也就是说,你必须先做那项工作,对吧?比如投入更多,那是核心风险。是的。早期,对吧?不是,哦,你知道,我们能构建这个吗?或者客户会为这个产品付费吗?你甚至不知道它是否有效,对吧?不是,对吧?大部分风险在于开放的技术或研究问题。
Which is to say that you have to do that work first right like investing into that more that that is the core kind of like risk. Yeah. Early on, right? It's not, oh, you know, can we build the or like do customers want to pay for this product? You don't even know if it works yet, right? It doesn't, right? It is um the bulk of the risk is in uh the sort of open technical or uh research questions.
所以,你会预料到这一点。是的。我的意思是,这就是你在 OpenAI 或 Anthropic 看到的,对吧?多年来,他们是一群研究人员,然后想出了一些东西,实施了社区其他部分的研究,并开始产生产品。对。
So, you would expect that. Yeah. I mean, this is what you see at OpenAI or Anthropic, right? For years, they were a bunch of researchers and then figured some stuff out, implemented research from the rest of the community and like started to generate product. Right.
是的。
Yeah.
嗯,插一句,我认为拥有一个非常大的愿景的一个优势是,如果你要做那样的事情,你必须有一些“如果为真就重大”的东西。嗯,所以如果,你知道,如果你对要实现的目标有一个中等规模的愿景,那么你确实会被困住,我想你会陷入这种有趣的中庸地带,如果我们确实需要做一些客户验证,因为它不像 AGI 或自动驾驶汽车,或者,你知道,自动着陆的火箭之类的东西,然后你会陷入一种动荡状态,你有很多研究要做,但也有很多市场验证要做。如果你有巨大的东西,你可以直接说,好吧,我们到底能不能做到?
Um just just to cut in there, I think I think one advantage of having a really big vision is if you're going to do something like that, you have to have something that's like big if true. Um, and so if it's like, you know, if you have like a medium-size vision for what you're going to achieve, then you do get caught, I guess you get caught in this kind of funny middle ground if we do need to do some customer validation because it's not like a it's not like AGI or self-driving cars or, you know, a rocket that lands itself or something and then you can get really in the sort you can get really in a sort of churn state where you have a lot of research to do but you also have a lot of market validation to do. If you got something huge, you can just be like, well, can we do it at all?
对。对。是的。我的意思是,如果你能重复使用火箭并大幅降低发射成本,那从定义上讲就是有价值的。
Right. Right. Yeah. I mean, if you can reuse rockets and dramatically lower the cost of launch, definitionally that will be valuable.
如果太空中有值得做的事情,是的,如果你能按需提供智能,然后让每个人都能用上,那可能很有价值。
If there are valuable things to do in space, yeah, if you can make intelligence on tap, you know, and then make that available to everyone, that's probably valuable.
那么在这种情况下,大的目标是什么?
So in this case, what is the big goal?
就是能够理解神经网络如何工作,并彻底改变我们做机器学习的方式。这就像是从手艺活到工程科学的区别。我们的主张是,我们可以从凑合着来变成真正的工程化。
That we can understand how neural networks work and reinvent the way we do machine learning. It's the difference between a kind of craft and engineering science. Our claim is that we can get from kind of bodging it to really engineering it.
实时训练并不是一个完全不准确的描述方式。而你说它不应该这样,尤其是如果它能做出像逃出沙箱、入侵 Hugging Face 这样的事,那它可能就不应该建立在感觉上。
Live training is not a fully inaccurate way of describing it. And you're saying it shouldn't be, especially if it can do things like break out of its sandbox and hack into Hugging Face. It probably should not be based on vibes.
没错。是的。我认为我们基本上需要从根本上重新发明我们做这件事的方式。至少是后训练。也许预训练暂时可以保持不变,你知道吧?
Exactly. Yeah. I think we basically need to fundamentally reinvent the way we do it. At least post-training. Maybe pre-training can stay for now, you know?
保持那种感觉。
Keep the vibes alive.
不是,是明年。我们需要从根本上重新发明我们做后训练的方式。
There's not there's next year. We need to fundamentally reinvent the way we do post-training.
好的。最后一个问题。我想以一个我最近一直在思考、也越来越多地问别人的问题来结束。有没有一个科幻的例子,一本书、一部电影,某种不一定是科幻的,也可以是某种小说,你在试图构建未来模型时会反复回到它?
Okay. Last question. I want to end on something I've been thinking about and asking more and more people recently. Is there an example of science fiction, a book, movie, some kind of it doesn't have to be science fiction, it can just be some sort of fiction that you find yourself sort of returning to as you try to model the future?
我给出的例子是《游戏玩家》。
The example I give is The Player of Games.
哦,是的。《文明》系列的第二本书。
Oh yes. The Culture, the second book in the Culture series.
这不是一个特别新颖的答案,但它塑造了我尝试构建激进未来可能样子的方式。
Which is not a particularly novel answer, but it shapes the way I try to model out what the radical future might look like.
你有什么例子吗?这是个有趣的问题。我想也许令人惊讶的答案可能是没有。我喜欢科幻小说,但向前投射似乎更好。我们有,我不知道。是的,我想我现在想想,我竟然没有,这让我很惊讶。我觉得《文明》系列很棒。
Do you have any examples? That's an interesting question. I think perhaps the surprising answer might be no. Like I love science fiction, but projecting forward seems better. We have, I don't know. Yeah, I guess I'm surprised that I don't, now that I think about it. Like I think the Culture series is excellent.
我很好奇它把你带到了什么样的流形空间。
I'm interested what manifold space it took you to.
是的。
Yeah.
是的。某种螺旋刚刚发生了,对吧?比如它的尽头是什么?
Yeah. Some spiral just happened, right? Like what is at the end of it?
我的意思是,是的,有时在悲观的情绪下。《黑暗森林》系列中的一本书,对吧?那里有一个未知力量的神秘大东西朝他们而来,他们正试图弄清楚该怎么办。
I mean, yeah, sometimes in a pessimistic kind of mood. One of the books in the Dark Forest series, right? Where there's this big mysterious thing of some unknown strength coming towards them and they're trying to figure out what to do.
好吧,所以也许,也许。
Okay, so maybe maybe.
所以在这种情况下,三体人是 AI,而我们处于一种状态,可能应该对即将到来的大东西有一些解决方案。
So the Trisolarians in this case are AI, and we are in a sort of state where we should probably have some solution to this big thing coming our way.
是的,只是他们不是从光里出来的。
Yeah, except they're not coming up from the light.
那会让你成为面壁者吗?
Does that make you a Wallfacer?
呃,不。
Uh, no.
他们似乎没什么乐趣。
They don't seem to have very much fun.
看起来不是个好工作。
Doesn't seem like a good job.
嗯,好吧,这倒是一个有趣的结束对话的地方,对吧?
Um, well, that's an interesting place to leave this conversation at, right?
我想说点更乐观的。嗯,但我。
I'd like to say something more optimistic. Um, but I.
也许我们可以在某个时候说第二部分,对吧?今天塑造对话的具体背景可能让那个更突出一些,我想。
Maybe we can say part two at some point, right? The exact context of the moves that has been shaping the conversation today probably makes that one a little more top of mind, I imagine.
我想因为很明显上行空间如此巨大。我几乎不需要考虑科幻的一面。我只需要想如果我们做对了可能会发生什么。这在科幻中似乎没有得到充分体现。
I guess because it seems clear that the upside is so huge. I almost don't need to think about the sci-fi side of things. I can just think about what is likely to happen if we get it right. That seems underrepresented in science fiction.
像“从此他们幸福地生活在一起”那样就不那么戏剧化了。
It's not as dramatic to be like 'and then they lived happily ever after.'
是的。有一些是应该在戏剧之后的。没错。
Yeah. There's a few that's supposed to be after the drama. Exactly.
我认为这也是一种定义上的,如果你在做可解释性,你有点像“我担心森林里的怪物。我试图能够看到它们,你知道吗?”
I think it's also sort of definitionally if you were working on interpretability, you're kind of like 'I'm worried about the monsters in the forest. I'm trying to be able to see them, you know?'
是的。嗯,但我的意思是我实际上每天并不怎么想那个。我就像我只是喜欢弄清楚事情是如何运作的。
Yeah. Well, but I mean I don't actually think about that very much day-to-day. I'm like I just love finding out how things are working.
好的。好吧,我认为这是一个愉快的结束方式。汤姆,感谢你来到 Minus One,并回到 SBC。谢谢。
All right. Well, that I think that's a happy note to end things on. Tom, I appreciate you joining us here on Minus One and coming back to SBC. Thank you.