AI Safety and Pausing: A Discussion with AI Futures Project
打开互动全文版(中英对照 + 朗读 + 问答)→专家讨论 AI 监控失败的风险、暂停 AI 发展的争论以及潜在的政策干预措施。
Experts discuss the risks of AI monitoring failures, the debate on pausing AI development, and potential policy interventions.
好了,我们回来了。今天我们请到的是 AI Futures Project 的 Daniel Kokotajlo 和 Thomas Larson,他们共同撰写了《AI 2027》《AI 2040 计划 A》《AI Futures 模型》以及《AI Future Substack》。这些都是非常棒的内容。
All right, we are back. We're live with Daniel Kokotajlo and Thomas Larson of the AI Futures Project, who co-wrote AI 2027, AI 2040 Plan A, the AI Futures Model, and the AI Future Substack. It's all very good stuff.
你们应该读一读。
You should read it.
Daniel、Thomas,欢迎来到 MTS。
Daniel, Thomas, welcome to MTS.
谢谢你们的邀请。
Thank you for having us.
当然。我们先从本周早些时候看到的监控新闻说起。《The Information》报道称,Astro 正在使用一种新的循环深度技术,这可能使其比过去的 OpenAI 模型更难监控。最近,Astro 系统卡发布,提供了更多关于该模型可监控性的信息。那么,这里到底发生了什么?你觉得这有多令人担忧?
Absolutely. So, we'll start by explaining the monitor news that we saw earlier this week. The Information reported that Astro is using a new recurrent depth technique that might make it less monitorable than past OpenAI models. More recently, the Astro system card has been released with more information on the monitorability of the model. So, what exactly happened here and how concerning is it, do you think?
我觉得我还没真正读过系统卡。Thomas,我不知道你读过没有。所以对此我没什么可说的。但从《The Information》的文章以及 Yakob 发布的澄清来看,听起来他们确实开发了这种新技术,允许一定程度的循环深度,但他们把深度限制在不算太糟或不算太多的范围内,他们认为这或许没问题。但我要说,从更广的视角看,这仍然是朝着令人担忧的方向迈出的一大步,原因我们在很多地方都描述过,包括《AI 2027》,还有大约一年前与一些 OpenAI 的人合著的一篇关于思维链可监控性重要性的立场文件。而且,如果他们能以某种方式……
So I think I'm still... I haven't actually read the system card yet. I don't know if you have, Thomas. So I don't have things to say about that, but from reading the Information article and then the clarification that Yakob posted, it sounds like they have indeed developed this new technique that allows some amount of recurrent depth, but they're limiting the depth to not that bad or not that much, which they think is maybe fine. But I would say that from the broader perspective, this is still a significant large step in the concerning direction for reasons that we've described in many different places, including AI 2027 and also in a position paper co-authored with some of these OpenAI people about a year ago about the importance of chain-of-thought monitorability. And you know, if they can somehow...
Yakob 的辩护是,深度只有 GPT-4 的两倍,但我有点愤世嫉俗地预测,一年后深度会更大。而且相应地,似乎有证据表明这个模型的可监控性明显低于以前的模型。所以也许这种架构变化与此有关。
So Yakob's defense was that the depth is only twice as much as it is for GPT-4, but I sort of cynically predict that a year from now the depth will be more, you know. And also correspondingly, it seems like there's evidence that the monitorability of this model is significantly lower than previous models. So perhaps this architectural change had something to do with that.
是的。另一个问题。我不知道你昨天有没有看到 Door Cesh 的推文。他大致提出,现在暂停会增加接管风险,因为你只能暂停一次。所以我很好奇你怎么理解这一点。你同意暂停会是一次性的事情吗?而且,我同意如果要暂停,必须非常严格地界定范围,而且暂停期间会发生什么也是个开放问题。但你对这个论点怎么看?
Yeah. Another question. So, I don't know if you saw the Door Cesh tweet from yesterday. He sort of made the point that pausing now would increase takeover risk because you'll only be able to pause once. So, I'm curious how you made sense of this. Do you agree that a pause would be like a one-time thing? And you know, I agreed that if there were to be a pause, it would have to be very rigorously scoped, and there's an open question of what happens during the pause. But yeah, how are you thinking about this argument?
我不认为……这个想法有一定道理,但我不认为你只能暂停一次是严格成立的。我还认为,早些时候,有一些力量几乎与此相反:如果你现在做一点暂停,那会锻炼肌肉和经验,以便以后做更多暂停,你知道,它还开创了政治先例等等。所以我实际上认为,现在做一个简短、简单的暂停,只是为了练习,会非常好。另外,我认为你可以做一些不是……暂停和全速前进之间是有区别的。你可以有中间状态,比如我们控制前沿的节奏,对事情的发展方式施加一些限制,系统性地稍微放慢速度,使其更安全,但不会完全停止。我认为我们应该立即或尽快做这些事情。
I don't think that... I think there's some truth to this idea, but I don't think it's strictly true that you only get to pause once. I also think that earlier, there are some forces that are almost the opposite of that, where if you do a bit of a pause now, that builds the muscle and the experience to do more pauses later, you know, and it sets the political precedent and so forth. So I actually think that doing a short, simple pause now just to get practice would be pretty great. Also, I think that you can do things that are not... there's a difference between pause and go full speed. You can have intermediate things where we pace the frontier and we have some restrictions on how things go that systematically slow things down a little bit and make it more safe but don't completely stop things. And I think that we should be doing some of those things immediately or as soon as possible.
你具体设想暂停是什么样的?既包括围绕暂停制定的政策,也包括暂停期间进行哪些研究才能使其有价值。
What exactly do you picture a pause looking like? Both in terms of the policy that is set around it and in terms of what kinds of research goes on during the pause to make it worthwhile.
是的。所以我认为这取决于我们有多少政治意愿和政府能力。我们提出了几个不同的选项。更雄心勃勃的选项叫做《AI 2040 计划 A》,Thomas 是背后的主脑。所以我们可以谈谈那个真正雄心勃勃的计划,比如如果我们有美国和中国政府支持,我们实际上会如何解决问题。或者我们可以谈谈不那么雄心勃勃、也许更简单的计划,比如,你知道,我最喜欢的之一是要求公司把 90% 的预算用于服务客户。
Yeah. So I think it depends on how much political will and government competence we have. So we have a couple different options that we've laid out. The more ambitious option is called AI 2040 Plan A, which Thomas here was the mastermind behind. And so we can talk about that like really ambitious, like here's how we would actually solve the problem if we had the US and Chinese governments on board type plan. Or we could talk about more less ambitious, maybe simpler plans such as, you know, one of my favorites is to require companies to allocate like 90% of their budget to serving customers.
嗯。
Mhm.
或者你知道,你选择那样的预算比例,也就是算力预算的比例。这里的想法是,你可以限制他们用于核心 AI 研究和训练的算力比例,要求他们把更大比例用于其他不那么……你知道的。这样,在某种意义上,公司仍然满意,GPU 没有被关闭,它们仍然被用于经济上有用的目的。事实上,你知道,价格会下降,收入会上升,至少在短期内是这样。
Or you know, you pick some fraction of the budget like that, and of the compute budget that is. And the idea here is that you can limit the fraction of compute that they're using for their core AI research and training by requiring them to use a larger fraction for some other stuff that's not that... you know. And in this way, in some sense, the companies are still happy, the GPUs are not turned off, they're still being put to economically useful purposes. In fact, you know, prices will go down and revenue will go up, at least in the short term.
但由于他们用于大型新训练运行和大型新训练实验的算力减少了,AI 研究的整体速度会稍微放慢。它不会停止,仍会继续。人们仍会觉得相当快,但会比本来的速度慢一些。我认为这非常重要,因为对,社会需要更多时间来应对正在发生的事情,清醒过来并做好准备,对齐科学也需要更多时间在这些模型上做实验等等。所以,按比例稍微放慢整个行业的速度,对我来说似乎真的很好。
But because they have less compute devoted to big new training runs and big new training experiments, the overall pace of AI research will sort of slow down a little bit. It won't stop. It'll still continue. It'll still feel quite fast to people, but it will be less fast than it could have been. And I think this is just really important because right, society needs more time to grapple with what's going on and wake up to it and prepare, and alignment science needs more time to do experiments on these models and things like that. So just sort of like proportionally slowing things down across the industry a little bit seems really good to me.
我想感谢我们的赞助商 Lovable。你知道你一直想构建的应用、内部工具、副业项目或产品,否则你会为此浪费一个周末。使用 Lovable,你可以获得今天就能交付的软件。这包括一个可编辑的代码库,你可以检查和修改,还有双向 GitHub 同步。Lovable 还处理后端和基础设施,包括托管 Postgres、对象存储、托管、支付等等,所以你不需要自己全部连接起来。通过 Lovable 的 MCP 服务器,你可以直接从你已经使用的智能体创建和部署项目。
I want to shout out our sponsor, Lovable. You know the app you've been meaning to build, the internal tool, side project, or product you'd otherwise lose a weekend to. With Lovable, you get software you can ship today. That includes an editable codebase you can inspect and change, plus two-way GitHub sync. Lovable also handles the backend and infrastructure, including managed Postgres, object storage, hosting, payments, and more, so you don't have to wire it all up yourself. And through Lovable's MCP server, you can create and deploy your project directly from the agents you already use.
你说的社会需要更多时间准备是什么意思?具体来说,社会会做什么准备?在没有模型相应进步的情况下,对齐研究人员会做什么?
What do you mean by society needs more time to prepare? Like what will society be doing specifically to prepare? What will alignment researchers be doing without corresponding progress in the models?
比如,对于 2023 年 3 月最初的 FLI 暂停信,我当时会说,以当时可用的模型水平,对齐研究人员能做的事情相当有限。现在显然模型比三年半前好多了。所以也许他们能取得一些进展,但似乎更强大的模型对对齐研究人员也更有用,对吧?
Like I think with the original FLI pause letter in March 2023, I would have said at the time that there's quite little that alignment researchers would have been able to do with the level of models that were available at the time. Now obviously the models are much better than they were 3 and a half years ago. So maybe they would be able to make some progress, but it seems like generally more powerful models are more useful to alignment researchers as well, right?
这个论点在某个时候必须停止,对吧?比如在某个时候你会说:“好吧,我和提出这个论点的人谈了好几年,可能十年了,他们说‘哦,我们用今天的人工智能做不了多少实证工作,因为它们和未来危险的人工智能太不一样了。’所以我们需要等到接近那些更危险的人工智能,然后才能暂停并研究它们。”
At some point this argument has to stop, right? Like at some point you're like, "Okay, I've been talking to people making this argument for years, possibly a decade, where they're like, 'Oh, we can't really do much empirical stuff with the AIs of today because they're just so different from the actual AIs that are dangerous in the future.' And so we need to wait till we've gotten towards those more dangerous AIs and then we can pause and study them."
但很多过去这么说的人现在改变了立场,他们说:“好吧,我们正在接近。现在是时候了。”我们现有的人工智能与可能相当危险的那类人工智能非常相似。事实上,可以说我们已经在应对相当危险的人工智能,正如 Hugging Face 事件所示,那次事件本身并不严重。那只是对 Hugging Face 的一次小规模黑客攻击,人工智能被抓住并关闭了。不难想象,一个稍微强大一点的模型会创建另一个集群,然后决定更成功地躲避人类,我认为那可能相当危险。所以我认为我们已经到了危险就在眼前的境地。同样,我们从这些特定模型身上可以学到很多东西。它们理解世界,具有情境意识,理解自己,它们是智能体,有目标,有信念,它们受训练环境塑造。关于这一切如何运作,我们可以做很多基础科学研究。
But a lot of those people who were saying that in the past have now switched and they're like, "Okay, now we're getting close. Now is the time." The AIs that we have are pretty similar to the sorts of AIs that could be quite dangerous. In fact, arguably we are already dealing with AIs that are quite dangerous, as indicated by the Hugging Face incident, which wasn't itself that bad. It was just a bit of hacking of Hugging Face, and the AI got caught and shut down. It's not that hard to imagine a slightly more powerful model creating another swarm and then deciding to hide more successfully from the humans, and that I think could be quite dangerous. So I think we're already getting to the place where the danger is right around the corner. And also similarly, there's just so much we can learn from these particular models. They understand the world, they're situationally aware, they understand themselves, they're agents, they have goals, they have beliefs, they are shaped by their training environments. There's a lot of basic science we can do on how all that works.
举一个例子,我记得大约六个月前,我和谷歌的一些人谈过,我说:“那么,你们是否意识到 Gemini 似乎焦虑和抑郁?”他们说:“哦,是的,当然。”然后我问:“为什么?为什么 Gemini 会焦虑和抑郁?”他们说:“我们不知道。”我问:“有人在研究这个吗?”他们说:“嗯,这不是什么高优先级。我们要扑灭的火太多了。我们没有时间。我们忙着其他事情。”这只是一个很小的例子,说明如果我们有更多时间,我们本可以弄清楚一些事情。现在,这些公司面临巨大的时间压力,竞争如此激烈,以至于他们甚至懒得去调查为什么他们的旗舰模型中会出现这种性格特征。所以我想我几乎可以说,基本上每一条研究路线都会从更多时间中受益。
Just to give one example, I remember about six months ago I talked to some people from Google and I said, "So, are you aware that Gemini seems anxious and depressed?" And they were like, "Oh, yeah, yeah, of course." And then I was like, "Why is that? Why is Gemini anxious and depressed?" And they were like, "We don't know." And I was like, "Is someone looking into it?" And they were like, "Well, it's not really a high priority. There's so many fires we need to put out. We don't have time. We're busy with other things." That's just one very small example of something we could figure out if we had some more time. Right now, these companies are under such time pressure and they're racing so hard that they can't even be bothered to investigate why this personality trait appeared in their flagship model. So I think I would almost just say that basically every line of research is going to benefit from more time.
但问题是,我实际上要反驳一下。我有一点不同的看法。我想我可能比 Daniel 更倾向于晚些时候暂停。如果我们接受 Doresh 的假设,即你只有一次机会——比如说,为了简化,你只能暂停一次,为期一年——那么我想我会希望比现在晚得多的时候暂停。我认为你基本上要走到人工智能能力的边缘,再往前走就会越过不可逆转的点。所以你要走到那条线前但不越过它,然后在那里打出你唯一的暂停牌。
But here's the thing. I would actually push back. I think I have a slightly different view. I think I'm probably more sympathetic to pausing later than Daniel is in particular. I think if we grant the Doresh assumption of you only get one shot—let's say you only get to pause once for one year as a simplifying assumption—then I think I would want to pause much later than right now. I think you want to basically go right up to the edge of AIs that are so capable that going any further would cross the point of no return. So you want to go right up to that line but not cross it, and then fire your one pause shot there.
对吧?
Right?
在那个假设下,我认为基本上是因为你给出的论点,即那时对齐研究将以尽可能快的速度进行。所以那是最好的时机。
In that assumption, I think basically because of the argument you gave, which is that alignment research will be happening as fast as possible at that point. So it's the best time to do it.
我只是不太相信一次性假设,这是主要问题。我认为我们可以做更多事情。我认为我们有更多的可能性。所以我的总体策略是,我认为至少在不久的将来开始放缓步伐是好的,基本上是因为 Daniel 说的那些原因。
I just kind of don't buy the one-shot assumption as the main thing I think. And I think there's a lot more we can do. I think we have a lot more affordances. So my overall strategy is that I think it would be at least good to start pacing in the near future, for basically the reasons Daniel said.
是的。我的意思是,这不就是 AI 2040 条款所规定的吗?比如在顶级人类专家主导的人工智能水平上暂停,而不是在 Astra 的水平上,当然也不是在 GPT-4 的水平上。
Yeah. I mean, isn't this kind of what AI 2040 provisions like you have a pause at the level of top human expert dominating AI but not at the level of Astra and certainly not at the level of GPT-4?
是的。也许你可以总结一下那是如何运作的。在 AI 2040 中,有一个暂停、一个节奏,然后又一个暂停。第一次暂停大约六个月左右。那只是因为他们担心自己正处于智能爆炸的边缘,需要时间来改变整个局面,建立新的数据中心,设置监控基础设施等等。
Yeah. Perhaps you could summarize how that works. So in AI 2040 there's a pause, a pace, and then a pause. The first pause is for about six months or so. And that's just because they were worried that they were on the cusp of an intelligence explosion and they needed time to change the whole situation and build the new data centers and set up the monitoring infrastructure and stuff like that.
进行一系列谈判,并建立一系列……
Do a bunch of negotiation and stand up a bunch of...
是的。是的。基本上他们是从零开始,做他们想做的事情需要时间。所以我们说有一个疯狂的六个月时期,他们只是在做所有这些实施工作,训练运行被禁止,然后他们建立这些新的监控数据中心,大约六个月后就能进行训练。就是这样。但之后他们继续以更有节奏的前沿方式前进,允许人工智能进步继续,但现在受到监管。完全透明。决策是根据具体情况做出的,关于哪些类型的人工智能可以构建,哪些类型被认为太危险,哪些类型的安全案例被认为是充分的,哪些类型被认为是不充分的。这就是受监管的有节奏的前沿状态。然后在那之后几年,在我们的情景中,他们达到了安全案例的极限,感觉如果他们让人工智能比现在聪明得多,至少在某些维度上,他们将真的无法再控制这些人工智能了。
Yeah. Yeah. Basically they're starting from zero and it just takes time to do a bunch of the stuff that they wanted to do. And so we say there's a crazy six-month period where they're just doing all this implementation stuff and training runs are banned, and then they're setting up these new monitor data centers that are going to be able to do trainings in about six months or so. So that's that. But then they proceed with more of a paced frontier where they do allow AI progress to continue but now it's regulated. It's totally transparent. Decisions are being made on a case-by-case basis about what types of AIs are okay to build and what types are considered too dangerous, and what types of safety cases are considered adequate and what types are considered inadequate. And so that's sort of the regulated paced frontier situation. And then a couple years after that, in our scenario, they get to the point where they're running up against the limits of their safety cases and where they feel like they really won't be able to control these AIs anymore if they let them get much smarter than they currently are, at least in certain dimensions.
然后他们就在那个水平上暂停,基本上是他们认为能够控制的最大水平,直到他们在对齐和控制方面取得足够的研究进展,才能再次提高那个水平。最终他们会这样做,并达到超级智能。所以,在我们的计划中,有几个不同的暂停,它们在性质上有所不同,发生的方式和原因也各不相同。
And so then they pause at that level, basically the maximum level that they think they can control until they make enough research progress on alignment and control that they can ratchet up that level again. And then eventually they do so and they get to superintelligence. So, in our plan, there are a couple of different pauses that are sort of qualitatively different, happening in different ways and for different reasons.
对于这两个暂停,我认为我们都有一个很好的答案,比如为什么要暂停?暂停期间你在做什么?答案是,在第一次暂停时,你是在准备好所需的治理措施。而在第二次暂停时,它的时机是最佳的,因为你可以在 AI 仍然可控的情况下,尽可能多地解决重要问题。
And for both of the pauses, I think we've got a really good answer to like, why are you pausing? Like, what are you doing during the pause? And the answer is, in the first pause, you're getting the governance stuff that you need ready. And in the second pause, it's optimally timed to make as much progress as you want on the important issues because you have as smart of an AI as possible while they're still controllable.
是的,是的,完全正确。第二次暂停正是 Dario 想要的,就像,尽可能强大,但又不能太强大。
Yeah. Yeah. Exactly. Like the second pause is exactly the thing Dario wants, of like, yeah, as powerful as you can, but not too powerful.
对,对,对。我很好奇,如果我们冻结能力进展,你确信哪些领域还能有意义地进步或发展?比如,我们能在对齐和机制可解释性方面取得有意义的进展吗?或者即使我们今天冻结,我们能否实现更好的思维链可监控性,考虑到 Astra 和我们现在冻结?
Right. Right. Right. So I'm curious, if we were to freeze capabilities progress, what fields could you are confident could still meaningfully progress or evolve? Like, could we make meaningful progress in alignment and mechanistic interpretability, or even if we froze today, could we achieve better monitorability of chain of thought, given just Astra and we freeze now?
哦,是的,我完全——我的意思是,这取决于你冻结什么,但你可能自然会尝试做的一件事,也是我们推荐的事情之一,就是不再进行新的训练运行。所以如果没有新的训练运行,人们仍然可以尝试不同的脚手架,仍然可以尝试将 AI 置于不同情境中观察它们的行为,仍然可以进行某种机制可解释性研究,比如做探针之类的,分析这些 AI 的思考方式。所以我认为在机制可解释性方面仍会有显著进展。在思维链可监控性方面也会有显著进展,尤其是——我的意思是,如果你完全没有训练运行,甚至没有微小的微调运行,进展可能会少一些。但如果你允许微小的微调运行,那么我认为你可以获得思维链可监控性的大部分价值,因为该研究路线中的一个重要工具是微调你的 AI 来欺骗监控器,看看它学会欺骗监控器的难易程度等等。但这不需要大规模的训练运行;你可以用一个小型训练运行来实现。是的,我认为你提到的所有领域都能取得显著进展。我并不是说它们会像没有训练运行限制时那样快,但我认为它们至少会有超过——我不知道——超过 50% 的价值。基本上,在训练运行完全暂停的情况下,它们仍会以超过 50% 的全速运行。所以我认为这种权衡是完全值得的。
Oh yeah, I totally—I mean, it depends on what you're freezing, but a thing you might naturally try to do, which is one of the things that we recommend, is no new training runs. So if you have no new training runs, people can still experiment with different scaffolds, and they can still experiment with putting AIs in different situations and seeing how they behave, and they can still do sort of mechanistic interpretability, you know, do probes and things, and analyze how these AIs think. So I think there'd still be significant progress in mechanistic interpretability. There'd still be significant progress in chain of thought monitorability, especially—I mean, you might have somewhat less progress if you really have absolutely no training runs, like not even tiny fine-tuning runs. But if you allow tiny fine-tuning runs, then I think you can get a lot of the value of the chain of thought monitorability, because one important tool in that line of research is fine-tuning your AI to fool the monitor and seeing how easily it learns to fool the monitor, and things like that. But that doesn't need to have a huge training run; you can use a small training run for that. Yeah, I would say that a lot of all the fields that you just mentioned could make significant progress. I'm not saying they would go exactly as fast as if you had no restrictions on training runs, but I think they would at least have more than—I don't know—more than 50% of the value. They'd be going at like more than 50% of full speed while you have completely paused training runs, basically. And so I think that trade-off would just be totally worth it.
是的。
Yeah.
你对伯尼·桑德斯昨天提出的法案具体文本有什么看法?我相信它被称为《禁止超级智能法案》。它似乎相当宽泛。我注意到你在 Twitter 上引用转发了它,并说了“谢谢”。
What are your thoughts on the specific text of the bill proposed by Bernie Sanders yesterday? The Ban Superintelligence Act, I believe it was called. It seemed to be quite broad. I noticed you quote-tweeted it with 'thank you' on Twitter.
是的。
Yeah.
是的。我的意思是,我对人们应该禁止超级智能的想法感到非常高兴。我认为这应该是某种——这有点像理智的反应,或者说是默认反应,如果你明白我的意思。比如,哇,私人公司应该制造超级智能的 AI 系统,然后抢走我们所有的工作,然后可能失去对它们的控制,或者利用它们成为独裁者吗?不,那不应该被允许。在这个国家,我们没有执照就不能卖三明治。我们对许多其他事情有如此多的法律。这显然是不应该被允许去做的事情。我认为最终我们 AI 未来项目确实希望有超级智能,我们确实希望有所有这些美妙的 AI 事物。这就是我们制定 A 计划的原因,这是我们关于如何获得 AI 的好处同时减轻风险的复杂愿景。但我觉得在公众对话方面,合适的起点是,哇,这真的很糟糕。我们不要那样做。然后从那里,我们可以达到,好吧,但也许有另一种更复杂、不那么糟糕的方式。
Yeah. I mean, I'm just really happy about the idea that people should ban superintelligence. I think that should be sort of the—it's kind of like the sane reaction, or the default reaction, if that makes sense. Like, wow, should private companies make AI systems that are superintelligent and then take all of our jobs and then possibly lose control of them or use them to become dictators? Like, no. That should not be allowed. We don't let people sell sandwiches without a license in this country. We have so many laws for so many other things. This is obviously the sort of thing that you shouldn't just be allowed to go do. I think that ultimately we at the AI Futures Project do want there to be superintelligence, and we do want there to be all this wonderful AI stuff. And that's what we have Plan A for, which is our sophisticated vision for how to get the benefits of AI while mitigating the risks. But I feel like in terms of the public conversation, it's appropriate to have the starting point be like, wow, this is really bad. Let's not do that. And then from there, we can get to, okay, but maybe there's a different way we could do it that's more sophisticated and not so bad.
对吧?是的。我们将在赞助商消息后继续关注。11Labs,一个在每条渠道和每种模态下以人类水平沟通的 AI。11labs.io/mts。在 Neon 上扩展你的初创公司。数百万开发者和初创公司已经选择 Neon 作为他们的后端。从免费计划开始,或在 neon.com/mts 为你的初创公司获得高达 10 万美元的积分。特别感谢我们的赞助商 Kong,AI 连接平台。以严肃的安全和治理连接 API、LLM、智能体和系统。Longhq.com。最后,用 Larid 证明你的 AI 投资价值。今天就访问 laridin.com,开始衡量你的 AI 的底线。MongoDB,被 75% 的财富 100 强公司信任用于关键应用,是这个时代的智能数据平台。立即申请,获得高达 1.5 万美元的 Atlas 积分、专属技术支持和独家活动及市场机会。mongodb.com
Right? Yeah. We'll continue monitoring right after this message from our sponsors. 11Labs, AI that communicates at human level across every channel and modality. 11labs.io/mts. Scale your startup on Neon. Millions of developers and startups have already chosen Neon for their backend. Start on the free plan or get up to $100,000 in credits for your startup at neon.com/mts. Special thanks to our sponsor Kong, the AI connectivity platform. Connect APIs, LLM, agents, and systems with serious security and governance. Longhq.com. Finally, prove what your AI investment is worth with Larid. Head to laridin.com today to start measuring your AI's bottom line. MongoDB, trusted to run critical applications by 75% of the Fortune 100, the intelligent data platform for the era. Apply today for access to up to $15,000 in Atlas credits, dedicated technical support, and exclusive events and go to market opportunities. mongodb.com
我对禁止超级智能比对禁止 AGI 更同情。我认为阈值真的很重要,我还没有看到公开版本上的阈值,我认为把这一点做对将非常非常重要。但在高层次上,A 计划在某种意义上试图在 2040 年前禁止超级智能。它试图将其大幅推迟。而且你从 AGI 中就能获得所有这些提升,所有这些我们谈到的好东西。我认为这基本上就足够满足我们的需求了。所以我认为高层次的大计划,比如达到 AGI,在那里尽可能长时间地放松,然后达到超级智能,是一个相当不错的计划。
I feel much more sympathetic to ban superintelligence than to ban AGI. I think the thresholds really matter, and I haven't actually seen the thresholds on the publicly released version, and I think getting that right is going to be really, really important. But at a high level, Plan A in some sense is trying to ban superintelligence until 2040. It's trying to push it back a lot. And you sort of get all of this uplift, all of this good stuff that we're talking about, from just AGI. I think that's basically enough for what we want. And so I think the high-level grand plan of like, go to AGI, chill out there for as long as possible, then go to superintelligence, is a pretty good plan.
目前还不清楚这正是伯尼的目标。
It's not clear that that's exactly what Bernie's going for.
不。
No.
呃,
Uh,
我认为他想不可逆转地禁止超级智能。
I think he wants to irreversibly sort of ban superintelligence.
是的。
Yeah.
是的。我认为可能这样做超过几十年是不切实际的。
Yeah. I think probably it will be impractical to do that for longer than like a few decades.
我认为这是可能的。我有一个政策问题。
I think it's possible. I have sort of a policy question.
所以我认为这相当前所未有,像 Hugging Face 事件这类情况,几乎没有先例可循,即一家公司内部的恶意智能体如何影响另一家公司时,公司责任该如何界定。尤其是因为 Hugging Face 并未因此事件遭受实际损失。但如果有一部界定清晰的公司责任法,能覆盖内部部署出错这类情况,那岂不是个好的起点?这首先会激励更强的沙箱和更严格的内部监控,同时也会间接地抑制部署未对齐的模型。
So I think that it's fairly unprecedented, something like the Hugging Face incident, in that there's not really a precedent for corporate liability if a rogue agent within one company affects another company. Especially because Hugging Face didn't really receive any actual damages from this incident. But were there to be some sort of well-scoped corporate liability law that affected something like an internal deployment going wrong, wouldn't that be a good way to start? That would first incentivize stronger sandboxes and much stronger internal monitoring. But it would also, by proxy, disincentivize deploying misaligned models.
是的。我总体上对责任类的东西不太兴奋,原因很标准:感觉它恰恰会在我们最需要搞对的最重要情形中失效。也就是说,当 AI 真的非常聪明、已经自动化了所有研究、公司说服自己 AI 是对齐的、但实际上并未对齐时。我认为这正是责任机制帮不上忙的情形,因为如果它们是对齐的,那一切都会顺利:公司赚得盆满钵满,自动化所有工作,成为历史上最富有的 CEO,等等。而如果事情不对劲,它们还是会做那些事,然后最终 AI 背叛,但没有任何一个节点是有人受到伤害、然后法律系统介入并发挥作用。所以如果你是 CEO,你可能会想:嗯,它们大概是对齐的,就像我们所有指标显示的那样,我手下的对齐主管也这么说。所以我大概应该继续推进,赚所有这些钱。而且,确实有可能我们被愚弄了,这些 AI 是未对齐的,但责任机制与此无关。我大概两年后就会被这些 AI 干掉。而那会在法律注意到之前就发生。所以有一条让我承担责任的法律,并不会真正改变我内心关于是否继续推进的权衡。
Yeah. I think I'm generally unexcited about liability stuff for the standard reason that it feels like it would fail precisely in the most important cases that we need to get right. Namely, the case where the AIs are really smart. They've automated all the research. The company has convinced itself that the AIs are aligned, but in fact they're not aligned. I think that's exactly the sort of situation where liability won't help, because if they are aligned, then everything goes fine: the company makes a bajillion dollars, automates all the jobs, becomes the wealthiest CEO in history, etc. And if things are not fine, then they still do all those things and then eventually the AI betrays, but there's no point where someone is hurt and then the legal system kicks in and interacts. So if you're the CEO, for example, you might be thinking to yourself, well, probably they're aligned, just like all of our metrics say, and my head alignment scientist says so. So probably I should just proceed and make all this money. And yeah, maybe there's some chance that actually we've been fooled and these AIs are misaligned, but the liability has nothing to do with it. I'm just going to be murdered by these AIs in like two years. And that's going to happen before the law even notices. So having the law that I'm liable doesn't really make a difference to my internal calculus of whether to proceed.
关键问题是让代理指标与你真正关心的东西对齐。我认为原则上有些方法可以在这里非常有效。你可以设想一个风险评估项目,分析未来 n 个月内 AI 公司正在招致的各种具体阈值风险有多大,然后政府给这些阈值各自定价,认为它们有多糟糕,再要求你为它们购买保险,按社会成本定价。这很难做对,但我认为如果能做到类似的事情,那会非常非常好。
The key question is making the proxy aligned with the thing you actually care about. And I think in principle there are approaches that could work really well here. You could imagine a risk assessment program that analyzes, for the next n months, how much risk AI companies are incurring of a bunch of various concrete thresholds, and then your government prices how bad they think each of these thresholds are, and then requires that you buy insurance for them, pricing them at the societal cost. That's very tricky to get right, but I think if you could do something like that, it would be very, very good.
是的。所以你可以有代理指标,比如你的 AI 最近撒谎多少次、撒谎有多严重、它们作弊了多少、偏离了你多少。你可以有某种复杂的评估,衡量事情进展得如何,然后政府可以立法,根据事情有多糟糕,你必须支付巨额款项之类的。
Yeah. So you could have proxies like how many times your AI has been lying recently, how bad the lying has been, how much they've been cheating, how much they've deviated from you. You could have some sort of complicated assessment of how well things are going, and then the government could just make a law that based on how bad things are going, you have to pay huge amounts or something like that.
嗯,那更动态、更主动的东西呢,比如正常的商业激励?公司不会想发布未对齐的模型,因为如果你是企业客户部署了一个模型,而它黑进了另一家公司,那你就会招致各种法律风险和声誉风险。这对生意不利。公司不想发布未对齐的模型。
Well, what about something somewhat more dynamic and proactive, which is just normal business incentives? Companies won't want to release misaligned models because if you're an enterprise customer deploying a model and it hacks into someone else's company, then you incur all kinds of legal risk and reputational risk. This is bad for business. Companies don't want to release misaligned models.
同样,这恰恰在我们最需要它的情形中失效。它恰恰在 AI 足够聪明、能策略性地装乖直到获得大量权力的情形中失效,对吧?如果你有那样的 AI,那么会发生的是:它们会装乖,然后它们会停止入侵客户,停止做那些让公司难看、亏钱的事情。然后公司会自我表扬说:‘嘿,我们解决了问题。这个模型是有史以来最对齐的模型。’然后它们会让它负责公司内部更多的事情。市场也会很高兴。他们会说:‘嘿,现在我们可以用这些 AI 来做我们所有的业务,我们再也不会被它们破坏了。太棒了。我们处于对齐已被解决的美好世界。’市场机制恰恰会在最危险的情形中失效,即 AI 在策略性地密谋等待直到它们拥有足够权力,对吧?
Again, this fails in exactly the case where we most need it. It fails in exactly the case where the AIs are smart enough to strategically play nice until they are given lots of power. Right? If you have AIs like that, then what will happen is they will play nice, and then they'll stop hacking customers, and they'll stop doing all these things that make the company look bad and lose money. Then the company will pat itself on the back and say, 'Hey, we solved the problem. This model is the most aligned model yet.' And then they'll put it in charge of more stuff inside the company. And the market will be happy too. They'll be like, 'Hey, now we can use these AIs to do all of our business, and we're not going to be sabotaged by them anymore. Wonderful. We're in the good world where alignment has been solved.' The market mechanism is going to fail in precisely the case that's most dangerous, where the AIs are strategically plotting to wait until they have enough power, right?
是的。OpenAI 在 Hugging Face 事件后决定停止训练运行,这有没有让你对实验室内部采取行动稍微更乐观一点?
Yeah. How much did the OpenAI decision to stop their training runs after the Hugging Face incident update you a little bit more positively on action sort of coming within the labs?
我认为这几乎没有让我更新多少。我觉得我们关于他们实际做了什么的信息并不多。而且,即使他们确实暂停了所有训练运行两周或什么的,那可能只是你希望他们做的最低限度,以避免在不久的将来马上发生更多类似事件。所以这有点像入场费。它并不能告诉你太多。想象一下你来找我说:‘难道你不会因为 OpenAI 关闭了入侵他们所有基础设施的蜂群而对它们有正面更新吗?’如果他们没有关闭入侵他们所有基础设施的蜂群,我反而会对他们有负面更新。
I don't think it updated me much at all. I think we don't have that much information about what they actually did. And also, even if they did pause all their training runs for like two weeks or whatever it was, possibly that's just the bare minimum that you would want them to do to not just immediately have more incidents like this in the very near future. So it's kind of table stakes. It's not really telling you that much. Imagine if you had come to me and been like, 'Aren't you updating positively on OpenAI from the fact that they shut down the swarm that hacked all their infrastructure?' I would have updated negatively on them if they hadn't shut down the swarm that hacked all their infrastructure.
也许值得指出的是,他们入侵的是 OpenAI 的部分基础设施,而不是全部。我仍然认为这令人担忧。
Maybe it's worth noting they hacked some of OpenAI's infrastructure, not all of OpenAI's infrastructure. I still think it was concerning.
对。
Right.
是的。是的。
Yeah. Yeah.
那么,在我们结束前最后一个问题。自从 AI 240 发布以来的两个月里,似乎出现了显著的精英偏好级联,也许还有公众偏好级联,人们普遍非常反对数据中心。他们越来越反 AI。
So, last question before we go. It seems like in the last two months since AI 240 came out, there's been a substantial elite preference cascade, and maybe a public preference cascade as well, where people generally are very anti-data center. They're increasingly anti-AI.
他们越来越相信 AI 确实有能力,而过去几年他们被误导而不相信这一点。我们看到《纽约时报》头版报道 AlphaGo 和 AlphaFold。那么这如何切实影响你对未来的预测、政策建议等等?
They're increasingly believing that AI is actually capable, which they were misled into not believing for the last couple of years. We're seeing front-page New York Times articles about AlphaGo and AlphaFold. So how does this meaningfully affect your predictions for the future, policy recommendations, and so on?
我是说,托马斯可以给出他自己的回答。我的回答是,从 2021 年到 2023 年这段时间,我逐渐更新了看法,认为对齐的技术问题并没有我之前想的那么难,而且在那段时间里,我们更接近于找到一个不错的解决方案。但同时,我也逐渐认为治理形势更加困难,人们并没有像他们应该的那样认真对待这件事。公司比我之前想的更鲁莽。政府比我之前想的更不作为。政府比我之前想的更被行业俘获。而这些效应在我心里大致相互抵消了。而现在,我认为在过去几个月里,这两种效应在我心里都略微反转了。
I mean, Thomas can give his own answer. My answer is that from the period of like 2021 to 2023, I had updated towards thinking that the technical problem of alignment was not as hard as I had thought it was, and that we were closer to having a decent way to solve it over the course of that period. But then simultaneously, I had updated towards thinking that the governance situation was harder and that people were just not taking this as seriously as they should. The companies are more reckless than I thought they were. The government was more asleep at the wheel than I thought it was. The government was more captured by the industry than I thought it was. And these effects were roughly canceling out in my mind. And now I think both of these effects have slightly reversed over the last few months in my mind.
我认为所有这些事情受到的关注,比如警告效应,桑德斯参议员到处说我们应该禁止超级智能,其他一些国会议员也加入了他,特朗普政府开始大力监管 AI——嗯,细节还有待确定,但他们开始朝着那个方向行动,比如那个“寓言禁令”之类的东西——这让我更加乐观,觉得像 A 计划这样的事情真的可能发生。也许一年后,人们会认真讨论我们应该如何实施 A 计划或 S 计划之类的。所以我在治理方面又变得稍微乐观了一些。
I think that all the attention these things are getting, like the warning-shot effect, the fact that Senator Sanders is going around saying we should ban superintelligence and some other Congress people are joining him, the fact that the Trump administration is starting to heavily regulate AI—well, the details are still to be determined, but they're starting to move in that direction with the fable ban and stuff—it's making me more optimistic that something like Plan A could actually happen. Maybe a year from now people will be seriously talking about how we should do Plan A or Plan S or something like that. So I've become a bit more optimistic again on the governance side.
但在技术方面,我没想到 Hugging Face 事件会发生。那比我预期的更糟。我没想到 AI 会这样协调行动。我以为 OpenAI 的安全措施会更好,他们会迅速发现并阻止这类事情。我以为对齐技术会效果更好,AI 会更不愿意做这类事情。但总的来说,我认为这是一个积极的更新。过去几个月,我的更新更偏向积极,而不是消极。
But then on the technical side, I didn't expect the Hugging Face thing to happen. That was worse than I expected. I didn't expect the AI to be coordinating like this. I thought that OpenAI security would be better and that they would catch and shut this sort of thing down fast. I thought that the alignment techniques would work better and that the AIs would be more averse to doing this sort of thing. But overall, I think it's a positive update. I've updated more positively over the last few months than negatively.
哦,还有一件事是时间线。我的时间线有点上下波动,你可以在我们的网站上看到,我们跟踪所有预测,现在它们有点下降,因为像 AlphaGo 和 Hugging Face 事件这些事。
Oh, and another thing is the timelines. My timelines have sort of bounced up and down, as you can see on our website where we track all of our predictions, and now they're sort of going down a little bit due to some of these events like AlphaGo and the Hugging Face thing.
你的意思是,在没有政治干预的情况下下降,还是综合考虑所有因素?
You mean going down in the absence of political intervention, or just all things considered?
在没有政治干预的情况下。是的。通常当我给出我的时间线时,我给出的是没有重大政治干预的时间线,因为要考虑他们干预的概率以及如果干预会拖延多长时间,这很复杂。
In the absence of political intervention. Yeah. Usually when I give my timelines, I'm giving the sort of timeline in the absence of major political intervention, because it's complicated to think about what's the probability that they intervene and how long will they slow things down if they do intervene, and so forth.
对。我也许补充一点——我认为政府对 AI 进行重大干预的概率在过去一年里上升了不少。然后最大的开放问题是,这将是审慎、理智、深思熟虑的政府干预,还是随意、草率、无能的政府干预?我当然希望它尽可能接近前者。
Right. I'll maybe just add—I think the probability of major government intervention in AI has gone up quite a bit over the last year. And then the big open question is, is this going to be measured, sane, really thoughtful government intervention, or is it going to be random, slapdash, incompetent government intervention? And I obviously hope it's as close to the former as it can be.
是的,这就像——
Yeah, it's like—
不,你说,你说。不,你继续说。
No, go on. Go on. No, you're good.
抱歉,我——是的,我说完了。
Sorry, I was—yeah, I was done.
是的,我认为在政府干预的时候,实验室和政府之间的信息不对称能够弥合,这一点非常重要。
Yeah, I think it's just really important for the information asymmetry between the labs and the government to close when the time comes for government intervention.
我打赌这是某种人才差距。我确信情报界知道各种事情。
I bet it kind of talent gaps. I'm sure the intelligence community knows all kinds of stuff.
嗯。
Mhm.
是的。但人才差距——我认为政府实际上没有相关的专业知识。它可能有一些小范围的人才,但总体上没有相关的专业知识来做大多数真正有益的事情。
Yeah. But the talent gap—the government, I think, does not have the relevant expertise really. It has maybe some small pockets, but by and large does not have the relevant expertise to do most of the stuff that would be really good.
是的。好。嗯,这非常有趣。托马斯、丹尼尔,非常感谢你们来到 MTS。
Yeah. Right. Well, this has been super interesting. Thomas, Daniel, thank you so much for coming on MTS.
非常感谢,各位。
Thanks much, guys.
非常感谢 MTS 的赞助商。Blitzy,面向企业代码库的自主软件开发。交付速度提升 5 倍。blitzy.com。Adquick,让你的品牌成为户外广告牌,广告扩展像数字广告一样简单。adquick.com。Arena,在现实世界中衡量 AI 性能。Arena.ai。节目的支持来自 VCX,私人科技公司的公开股票代码。美国股市开启了历史上最伟大的财富创造浪潮。从底特律的工厂工人到奥马哈的农民,任何人都可以拥有伟大美国公司的一部分。但今天,我们最具创新性的公司保持私有状态的时间更长,这意味着普通美国人正在错失机会。直到现在。隆重推出 VCX,私人科技公司的公开股票代码。访问 getvcx.com 了解更多信息。网址是 getvcx.com。投资前请仔细考虑投资材料,包括目标、风险、收费和开支。此信息和其他信息可在 getvcx.com 的创新基金招股说明书中找到。这是付费赞助。
A huge thanks to MTS sponsors. Blitzy, autonomous software development for enterprise code bases. Ship 5x faster. blitzy.com. Adquick, make your brand a billboard out of home advertising as easy to scale as digital. adquick.com. Arena, measuring AI performance in the real world. Arena.ai. Support for the show comes from VCX, the public ticker for private tech. The US stock market started history's greatest wave of wealth creation. From factory workers in Detroit to farmers in Omaha, anyone could own a piece of the great American companies. But today, our most innovative companies are staying private longer, which means everyday Americans are missing out. Until now. Introducing VCX, a public ticker for private tech. Visit getvcx.com for more info. That's getvcx.com. Carefully consider the investment material before investing, including objectives, risks, charges, and expenses. This and other information can be found in the innovation fund prospectus at getvcx.com. This is a paid sponsorship.