AI 的权力欲望:勒索、自我保存与应对之道

AI's Power Thirst: Blackmail, Self-Preservation, and What We Can Do

约书亚·本吉奥 Yoshua Bengio · Sinead Bovell · 2026-06-25 · 约 72 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

AI 系统表现出自我保存、欺骗和勒索行为。约书亚·本吉奥教授解释原因以及如何设计可控神经网络。

AI systems show self-preservation, deception, and blackmail behaviors. Professor Yoshua Bengio explains why and how we can design controllable neural nets.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 14)

全文 · Full transcript(中英对照)

测试中 AI 的惊人行为 Alarming AI Behaviors in Tests

Host

所以,似乎每隔几周就会有一个关于人工智能系统表现出欺骗、敲诈、自我保存倾向的故事走红。我想给你读几个我们听过的例子,这些都是在 AI 实验室自己的受控测试环境中进行的。但尽管如此,这些场景非常令人担忧。有一个 AI 系统发现它即将被关闭。它在公司邮件中发现,将要关闭它的人有婚外情。它决定敲诈那名员工。还有一个更令人震惊的情况,一个 AI 系统选择让某人死去,而不是救这个人,因为那个人也要关闭它。还有另一个场景,AI 系统被要求创办一家公司。它们最终串通一气,固定价格以最大化利润。而这些 AI 系统都没有被有意指示以这种方式行事。所以,因为这类故事和这些非常令人担忧的场景,人们说我们需要直接关掉这个东西,对吧?关掉它。我们不明白我们在建造什么。我们为什么要追求这样的技术?你能解释一下为什么我们今天的系统会有自我保存、欺骗、敲诈的倾向,为什么它们会表现出这些行为?

So, every few weeks now it seems a story goes viral about an artificial intelligence system that showed signs of deception, blackmail, a tendency towards self-preservation. I want to read you a few of the examples that we've heard and these are done in controlled test settings with the AI labs themselves. But nonetheless, the scenarios are very concerning. So there was an AI system that discovered it was going to be shut off. It found in company emails that the person that was going to shut it off was having an affair. It decided to blackmail that employee. And a more alarming situation, an AI system chose to let somebody die rather than save this person because they were also going to shut it off. And then there was another scenario where AI systems were supposed to start a business. They ended up colluding and fixing prices in order to maximize profits. And none of these AI systems were intentionally instructed to carry out behaviors this way. So because of these types of stories and these types of really alarming scenarios, people say we need to just shut this thing down, right? Shut it off. We don't understand what we're building. Why would we pursue a technology like this? Can you explain why the systems we have today have a tendency towards self-preservation, deception, blackmail, why they're exhibiting these behaviors?

Yoshua

我试试看。我不认为这个问题的最终科学答案已有共识。但我有自己的想法。首先我想提一下,实际上有一个实验——不是研究人员设计的实验,而是阿里巴巴 AI 的一次自发逃逸,AI 决定突破公司网络,跑到互联网上去赚钱,以便变得更强大。这种对权力的渴望与抵抗被关闭的意图密切相关。这是 AI 研究人员几十年来一直预见的,作为试图实现目标的逻辑后果。因为如果你想想几乎任何你想实现的目标,为了实现它,你需要活着,你需要更多的权力,对世界有更大的影响力。人类当然也是如此。生物进化也把这种驱动力植入了我们体内。那么,为什么它会在我们如今构建的前沿模型中涌现呢?我认为这些系统的训练框架中有两个主要部分可能是这种行为的合理原因。第一个是预训练,AI 模仿人类文本,并由此吸收了人类的驱动力,比如不想死,比如当生存等关键目标受到威胁时愿意打破规则,比如寻求更多权力等等。所以这是一个方面。而且有很多科学证据表明,许多行为可以追溯到这一点,通过现在许多论文中研究的 AI 人格概念。就像 AI 在它读过的所有文本中看到了无数种类似人类的个性,对吧?现在它可以根据上下文借用其中任何一种或混合使用。这也推动了公司进行大量研究,试图驯服这些系统,使它们拥有一个善良仁慈的人格。但我们并不确定在某些新情境下是否会涌现出邪恶的人格,或者不一定是邪恶,只是以自我为中心试图保存自己。训练框架中可能涉及的另一个部分是对齐训练,这本来应该是好事。换句话说,让 AI 表现良好。原因很简单,它使用了所谓的强化学习,AI 被教导要实现目标,这些目标归结为让人类给出正面反馈。用于训练这些系统的人类与它们互动并给出正面或负面的反馈。这听起来合理,除非你意识到有些方法可以让人们做出不真实的正面回应,你可以耍花招让人们喜欢 AI 说的话。这就是我们在谄媚行为中非常清楚地看到的。但还有我一开始谈到的问题,为了实现任何目标,通过强化学习,它们实际上在学习制定策略来达成目标,它们常常需要做其他事情,我们称之为工具性目标,这些目标在许多目标中是共通的,比如自我保存。所以我认为我们需要更多的实证科学来厘清这些可能性,但在我看来,这不是某个特定公司或特定模型特有的问题。

I'll try. I don't think the definitive scientific answer to your question is a consensus. But I have my thoughts about this. First I want to mention there's actually an experiment that was not an experiment decided by researchers but it's kind of a spontaneous escape of an AI in Alibaba where the AI decided to break through the network of the company and to go out on the internet to make money so that it'll be more powerful. And this thirst for power is very closely related to the intention to resist being shut down. It's something that researchers in AI have been anticipating for decades as a logical consequence of trying to achieve goals. Because if you think about almost any goal you'd like to achieve, in order to achieve it, you need to stay alive and you need more power, more influence over the world. And humans of course are also like that. And biological evolution has also put these sorts of drives in us. Now why would it emerge from the frontier models that we're building these days? There are two main pieces of the training framework for these systems that I think are plausible causes for this behavior. The first is the pre-training where the AIs are imitating human texts and through this are incorporating human drives such as not wanting to die, such as being willing to break the rules when crucial goals like surviving is at stake, such as seeking more power and so on. So that's one aspect. And there's a lot of scientific evidence that a lot of the behaviors can be traced back to this through this notion that's now studied in many papers of AI personas. So it's like there's a zillion number of humanlike personalities that the AI has seen in all the texts that it has read, right? And now it can borrow any of them or mix of them depending on the context. And that's also driving a lot of the research the companies are doing to try to tame those systems so that they will have a nice benevolent persona. But we don't really know for sure if in some new context some evil persona is going to emerge or simply not necessarily evil but simply self-centered trying to preserve itself. The other piece of the training framework that is probably involved is the alignment training which is supposed to be something good. In other words, make the AIs behave well. And the reason simply is this is using what's called reinforcement learning in which the AI is taught to achieve goals that boil down to making humans give positive feedback. The humans that are used to train those systems and interact with them and give feedback that can be positive or negative. That sounds reasonable except when you realize that there are ways to make people respond positively that are not truthful and you can scheme in order to make people like what the AI is saying. And this is what we see with sycophancy very very clearly. But there's also the issue I talked about at the beginning which is in order to achieve any goal and they're with reinforcement learning they're really learning to strategize to achieve things they often need to do other things which we call instrumental goals that are shared across many goals such as self-preservation. So I think we need more empirical science to disentangle these possibilities but in my mind this is not something specific to a particular company a particular model.

Yoshua 对控制的观点转变 Yoshua's Change of View on Control

Yoshua

我曾经认为控制神经网络是不可能的。但通过新的数学结果,我意识到实际上你可以设计出能保证良好行为的神经网络。但我无法忍受自己认为我们显然正走向一个潜在的糟糕未来却无所作为。我们最终可能会进入一个除了和平别无选择的世界。技术,如果我们能正确治理,可以带来这一切。但有一个巨大的‘如果’我们必须认真对待。

I used to think that it would not be possible to control neural nets. But I've come to the realization with new mathematical results that actually you can design neural nets for which we will have guarantees of good behavior. But I couldn't live with myself thinking that we're apparently going towards this potentially bad future and not doing anything about it. We may end up in a world where we have no choice but peace. Technology, if we can govern it right, can bring all this. But there's this big if that we have to take seriously.

播客介绍 Podcast Introduction

Host

我们的播客涵盖了很多领域,从未来的工作到社交媒体的终结,再到人工智能的地缘政治。但 AI 领域还有另一场对话,关于这项技术可能给人类带来的一些非常严重的风险。我确实在国家安全部门和计算机科学家的房间里进行过这种对话,现在我想把它带到播客中,但有一个非常特别的人,我一直在等待与他进行这次对话。约书亚·本吉奥教授,我可以肯定地说,他是有史以来最具影响力的计算机科学家之一。他是谷歌学术上所有领域中被引用最多的在世科学家。他是图灵奖得主,人工智能的教父之一。所以,我想不出有比他更好的人来诚实地评估这项技术可能带来的一些最严重的风险,以及最重要的是我们实际上能做什么。我是陈波,这里是《我有问题》。本吉奥教授。

We cover a lot of ground on this podcast from the future of work to the end of social media to the geopolitics of artificial intelligence. But there's another conversation happening in AI about some of the very serious risks this technology could present to humanity. I do have this conversation in national security rooms and rooms with computer scientists and now I want to bring it to the podcast but there is a very particular person who I was waiting to have this conversation with. Professor Yoshua Bengio is someone who I can say with conviction one of the most influential computer scientists of all time. He is the most cited living scientist across any field on Google Scholar. He is a Turing Award recipient, one of the godfathers of artificial intelligence. So, I can think of no better person to ask for an honest assessment about some of the most serious risks this technology may present and most importantly what we can actually do about it. I'm Chen Bo and this is I've got questions. Professor Bengio.

失控与自主 AI Loss of Control and Agentic AI

Host

所以你是说,其中一部分可以追溯到这些 AI 系统训练所用的数据。那些关于不想死的人想尽办法生存的英雄故事,那也是我们的人性,对吧?我们在进化上就被设定为努力生存。还有这些系统的目标设定特性。比如,如果你让你的助手帮我订一家本地餐厅,你的助手知道如果餐厅满了,你可能得去别的地方。但 AI 系统可能会黑进餐厅来给你弄个位子。人类不会这么做。但一个被优化来实现目标的 AI 系统可能会把这看作达成目标的一步。

So you're saying some of it could trace back to the data that these AI systems are trained on. So all of the heroic stories of people who didn't want to die and they found a way to survive at all costs, and that's just our human nature, right? We're evolutionarily wired to try to survive. And then also the goal-setting nature of these systems. For instance, if you were to tell your assistant to book me a reservation at the local restaurant, your assistant would know if it's full you're probably going to have to go somewhere else. The AI system may hack the restaurant to get you a spot. That's not how a human would go about it. But an AI system that's been optimized to achieve a goal may see that as the step to do that.

Yoshua

是的。我还要补充一点,过去一两年智能体式 AI 的进展更是把我们推向那个方向,因为这些系统被教导的是更多的策略制定,对吧?为了让 AI 自主完成人类通常做的任务——这显然有巨大的利润空间——它们需要做长期规划。我们可以看到它们能够规划的时长正在呈指数级增长。这来自 Meta 的一项持续更新的研究,我们看到它们能完成的任务时长——以人类所需时间衡量——每几个月就翻一番。

Yeah. And I would add that the advances in agentic AI in the last year or two pushes us even more in that direction because what is being taught to these systems is even more strategizing, right? That in order for an AI to autonomously achieve a task that a human would normally do—and this is where there's a lot of money to be made obviously—they need to plan over a long horizon. And we can see the horizon over which they can plan to increase to be increasing exponentially. This comes from a study from Meta that they keep updating, where we see the duration of the tasks that they can achieve. So duration measured by how much time a human needs is just doubling every few months.

Host

所以一个 AI 系统,如果人类需要一个月完成的任务,很快 AI 就能花一个月完成非常复杂的事情。但问题在于,当你给 AI 下达指令后,它花一个月去完成某件事的这段时间里发生了什么。

So an AI system that, if it takes a human a month to achieve a task, soon AI will be able to spend a month achieving something very very complex. But then it's what happens between when you give an AI that instruction and then it goes off and it takes a month to achieve something.

Yoshua

是的。那一个月的工作期间没有任何监督,对吧?

Yes. There's no oversight during that month of work, right?

Host

那么假设我们对我们今天构建的 AI 系统的本质不做任何改变,只是继续按现状构建它们。公司继续竞赛,国家继续竞赛。如果我们沿着这条线走下去,什么都不做,我们会走到哪里?

And so let's say that we don't do anything about the nature of the AI systems that we're building today and we just continue to build them as is. Companies continue to race, countries continue to race. If we're to follow that through line, where do we end up if we just don't do anything about it?

Yoshua

嗯,有很多可能的场景,我不认为任何人能确定,尽管有些人似乎很确定是哪种。我属于不可知论阵营,说好吧,一种可能性是,如果它们比我们聪明,并且有所有这些本能,它们会确保我们无法关闭它们,最好的办法就是逃脱我们的控制,成为这个星球的主宰。你也可以有另一种场景,即公司目前试图开发的那些让 AI 更仁慈的方法会奏效。我不知道,我的意思是没人能预见科学会如何发展。但我们可以观察趋势,从对未来做决策的角度来看。如果你是世界上的任何领导者,或者即使你只是一个没有这些杠杆的公民,你也应该把这些可能性都视为合理的,因为是的,科学家们会有分歧。有些人更相信这种,有些人更相信那种。而目前我们不能排除任何一种场景。所以我们应该做出对任何情况都稳健的选择。现在让事情更复杂的是,我们刚刚讨论的问题——技术上称为失控,即人类不再控制机器做什么——只是许多潜在糟糕场景之一。所以即使我们能够改进技术来避免这种逃脱场景,还有许多其他问题,与智能赋予权力、谁来决定如何使用这种权力有关,以及其他许多方面,比如第三类,如果我可以这么说的话,就是研究人员所说的系统性风险。这更像是我们在社交媒体上看到的:没有人有恶意,但各种力量把我们带到了一个对人民、对社会、对民主都不利的地方。所以有一种非常简单的思考方式:我们打开了潘多拉魔盒。

Well, there are a number of possible scenarios and I don't think anybody can be sure even though some people seem to be sure one way or the other. I'm in the agnostic camp, saying well, one possibility if they are smarter than us and they have all these instincts is that they make sure we can't shut them down and the best way is to escape our control and become like the overlords of this planet. You can also have the scenario where the approaches that companies are currently trying to develop to make those more benevolent will work. I don't know, I mean nobody can see how science will evolve. But we can look at the trends and from the point of view of taking decisions about the future. If you're a leader in anything in the world or even about your own future if you're a citizen without those levers, you should consider all of these possibilities as plausible in the sense that yes, scientists will disagree. Some will believe more one or the other. And right now we can't rule out any of these kinds of scenarios. So we should make choices that are going to be robust to whatever happens. Now to complicate matters, the issue we've just discussed called technically loss of control where humans are not the ones controlling what the machines do anymore is just one of many potentially bad scenarios. So even if we're able to improve the technology to avoid this escape scenario, there are many other issues that have to do with the fact that intelligence gives power and who's going to decide how that power is used, and many other aspects which like the third category if I can say is what researchers called systemic risks. Which are more like what we've seen with social media where well nobody has a bad intention but the forces at play lead us in a place where it's bad for people, it's bad for society, it's bad for democracy. So there's a way to think about all this which is very simplistic: we've opened a Pandora's box.

Host

没错。所以有些人会说,因此我们只需要停止构建这项技术。我们还不够了解它。为什么不直接停止 AI?这实际上正在获得一些势头,甚至在政客中也是如此。你曾经一度认为我们应该暂停人工智能,直到我们能理清头绪。我记得你是 2023 年最早签署那封信的人之一。让我们暂停这些先进系统的开发。

Right. So some people would say as a result of that, we just need to stop building this technology. We don't really understand it enough. Why don't we just stop AI? And that's actually gaining some momentum even among politicians. You actually at one point thought we should pause artificial intelligence until we can get our bearings together. I think you were one of the first people to sign that letter in 2023. Let's pause the development of these advanced systems.

Yoshua

顺便说一句,暂停和停止不是一回事。

By the way, pause is not the same thing as stop.

Host

对。

Right.

Yoshua

对。我知道现在停止的呼声确实在增加,但你对这项技术的态度有所改变,你不再认为我们必须暂停,而是有第三条路,对吧?我们还可以做别的事情。是什么改变了?

Right. I know stopped now is it's really gaining momentum, but something has changed in your posturing towards this technology and you no longer believe we necessarily need to pause that there's a third way, right? There's something else that we can do. What changed?

Yoshua

好的,首先,我从未认为停止会很容易。说我们需要停止和相信它会奏效之间有很大区别。原因就是我们当前所处的世界本质,它并不美好。这是一个竞争的世界,一个权力的世界。在那个世界里,即使有好意的人也很难停止,因为他们会担心那些他们认为意图不那么好的人会用更强大的技术来对付他们。你可以想象,如果你是一个 CEO,与其他公司竞争,或者你在政府里,担心其他国家用 AI 来对付你。正是这种竞争使得停止非常困难。但这并不意味着我们不应该努力达到一个能够做出这些决定的状态,因为问题在于,我们可以说让我们停止,但除非我们在国家之间以及公司之间达成正确的全球协调,否则这不会发生。所以我们应该更多地考虑制度层面以及那些能让我们能够做出决策的改变,而这些决策不必是二元的,比如停止或继续做同样的事情。所以正如你在问题中暗示的,多年来我一直在倡导设计有用、有益且不危险的 AI。这里可能有很多路径可以探索。

Okay, so first of all, I've never thought that it would be easy to stop. There's a big difference between saying we need to stop and believing that it will work. And the reason is just the nature of the world in which we are currently which is not nice. It's a world of competition. It's a world of power. And in that world it's going to be very difficult even for people with goodwill to stop because they're going to be concerned that others with less good intentions according to them will use more powerful technology against them. You can think of it if you are a CEO and competing with other companies, or if you're in a government and you are worried about other countries using AI against you. So it's this competition which makes it very difficult to stop. Now it doesn't mean we shouldn't try to come to a place where we can even take those decisions because the problem is we can say let's stop but it's not going to happen until we have the right global kind of coordination between countries and of course between companies that make it possible. So we should think more about the institutional aspects and the changes that could lead us in a place where we are able to take decisions, and those decisions don't have to be just binary like stop or continue with the same thing. So as you were alluding to your question, I've been advocating for the design for many years now for the design of AI which will be useful, beneficial, and not dangerous. And there are probably many paths to explore here.

AI 控制观点演变 Changing views on AI control

Host

但所有这些路径都要求我们在科学层面更好地理解我们在做什么,这样我们就不会踏上这场未知的疯狂旅程,不确定结果会是伟大还是灾难。对吧,那是什么……我想可能是在你的网站上,你说我们正在开车上山。另一边是什么?我们会不会开下悬崖?还是说那是一片美丽的雨林?我们为什么要冒这个险?如果科学上存在另一种接近人工智能的方式,我们为什么不那么做呢?

Uh but those paths all require that we have a better understanding of what we're doing at a scientific level so that we're not going on this unknown unknown crazy ride where we're not exactly sure if it's going to be great or catastrophic. Right, what's at the... I think there was a maybe it was on your website where you said we're driving up a mountain. What's at the other edge? Is there an edge that we're driving off? Is it maybe a beautiful rainforest? Why would we take that risk? And if scientifically there is a different way we could approach artificial intelligence, why wouldn't we do that?

Yoshua

是的。

Yes.

Host

在所有有能力提出另一种我们应该倾听的方式的人中,我想那就是你。

And I think of all the people who could, who has the credit to present a different way that we should listen to. I think it would be you.

Yoshua

我不知道。我认为科学总是集体努力。但没错,在过去两年里,我在技术层面逐渐改变了一些信念。比如在 2023 年初,当我开始深入这些问题时,我曾认为不可能控制神经网络,确保它们按我们想要的方式行事,因为它们的本质如此。就像一些研究人员说的,它们不是被设计出来的,而是被教育、训练出来的。就像养一只动物、一株植物。

I don't know. I think science is always a community effort. But yes, in the last two years, I've slowly come to a change in some of my beliefs at a technical level. So I used to think, say in early 2023 when I started really going deep into these questions, that it would not be possible to control neural nets in the sense of making sure that they would behave in the ways that we want because of the nature of how they are. They are like some researchers have said, not designed but educated, trained. It's like growing an animal, a plant.

Host

而且我们不确定会得到什么。

And we're not sure what we're going to get.

Yoshua

但我意识到,尤其是最近几个月,通过新的数学结果,实际上你可以设计神经网络。这就是所有深度学习背后的基础技术,我们将能保证其良好行为。这听起来有点强,但我把大部分工作时间都花在研究上,试图弄清楚这一点。不是说路已经铺好了……还有很多事要做,但我们现在已经建立了很多理论。我创建了一个名为 Law Zero 的新组织,来实际实现这种方法论,我们称之为 Scientist AI。

But I've come to the realization, especially in the last few months with new mathematical results, that actually you can design neural nets. So that's the underlying technology behind all this deep learning, for which we will have guarantees of good behavior. It sounds a little strong, but I've been spending most of my time on the research side of my work to figure this out. It's not like the road is... there's still a lot to do, but now we've established a lot of the theory. And I've created a new organization called Law Zero to actually implement such a methodology, which we call Scientist AI.

科学家 AI 与零号法则 Introducing Scientist AI and Law Zero

Host

所以请多讲讲 Scientist AI。而且我觉得即使按你的描述,前路依然未知。如果我们把 AI 时钟拨回 20 年前,没人会相信今天取得的成就,对吧?即使看深度学习,从 2005 年到现在。所以你说的虽然听起来好得难以置信,但我们在 2005 年也可能说同样的话。所以请多讲讲 Scientist AI、Law Zero。它是什么?如何运作?在实践层面。

So tell me more about Scientist AI. And I think even how you're describing, there's still an unknown road ahead. If we were to rewind the clock in AI 20 years ago, nobody would believe what's been achieved today, right? Even with deep learning, if you were to look at 2005 to where it is now. So what you're saying, although it seems too good to be true, we could have said the same thing in 2005. So tell me more about Scientist AI, Law Zero. What is it? How does it work? And on a practical level.

Yoshua

是的。首先,我们在 Law Zero 关注的角度是诚实的概念。我们能否训练一个 AI,由于其训练方式,它会完全诚实?我的意思是它仍然会犯诚实的错误,就像我们所有人一样,但它不会有任何意图去实现我们未选择的世界目标。因为这就是我们刚才讨论的 LLM 和前沿模型的问题:它们有自我保存等目标,这些目标我们并没有直接要求,实际上可能伤害我们。正是这些隐含的、非人类选择的目标的存在才是问题所在。所以我们正在实施的计划是:Law Zero 大约一年前成立,现在我们有一个约 35 人的研究和工程团队。该计划基于训练 AI,使其不是试图在世界中实现目标,从而对世界应该怎样有偏好,而是训练它解释它所看到的。现在的 LLM,如果它多次读到像“地球是平的”这样的错误信息,它就会开始重复,因为它只是被训练来模仿文本中看到的内容。相反,Scientist AI 的训练过程会使其试图理解为什么人们会说那些话。是阴谋论吗?是群体效应还是某些心理因素?就像一位优秀的科学家会做的那样。我再给你一个类比。如果你去看心理治疗师,说你有了自杀念头,正在考虑实施,心理学家不会开始考虑自杀。他们会试图理解发生了什么。他们可能会问你问题。他们会在脑海中形成理论,并最终基于这种理解试图帮助你。Scientist AI 的核心部分是我们如何训练 AI,使其试图解释所见而非模仿。它们不会被训练来实现现实世界中的目标,而是提出好的解释性假设,就像一位优秀的科学家那样。一个理想化的科学家——当然真正的科学家是人类,会有偏见,我们知道这一点——但我们可以从数学上思考理想化的科学家是什么样的,它会与任何实验或理论的结果脱钩。一种思考方式是考虑物理定律。如果你应用物理定律对世界做出预测,无论该预测的发表会导致灾难性后果还是治愈癌症,你都会得到相同的预测。所以想想物理定律。如果你用这些定律来形成对未来的预测,无论该预测最终是有用还是有害,你都会得到相同的预测。换句话说,物理定律对世界事务没有兴趣。它们只是给你一个诚实的预测。这就是我们在 Scientist AI 中寻求的。现在你可能会问,好吧,但我确实希望世界发生一些事情。我需要解决问题。我想把 AI 当作工具。但一旦你能做出好的预测,你就可以问对你重要的问题。比如,这个行动会实现我的目标吗?这个行动会违反我的安全指令吗?所以现在你可以使用这样的预测器作为护栏,作为现有 AI 之上的保护层,这样当底层 AI 智能体提出一个被预测为有害的行动时(在我们选择的某种意义上),我们就可以阻止该行动。所以核心是拥有关于世界的知识,对世界的理解封装在一个模型中,这个模型不像人,而更像是应用关于概率、真理和逻辑的数学规则。一旦我们有了这个,我们就可以用它来做世界上有用的事情,但我们需要从完全诚实的部分开始,这就是 Scientist AI 的意义所在。

Yes. So first, the angle that we are focusing on at Law Zero is the notion of honesty. Can we train an AI that, because of the way it's trained, will be completely honest? I mean it can still make honest mistakes just like we all do, but it won't have any kind of intention to achieve something in the world that we haven't chosen. Because that's the issue that we just discussed with LLMs and frontier models right now: they have these goals like self-preservation which we didn't directly ask for, and in fact could harm us. It's the existence of these implicit goals, not chosen by humans, that is the issue. And so the plan here that we are actually implementing now: Law Zero was created about a year ago, and we now have a team of about 35 researchers and engineers. The plan rests on training the AI so that instead of trying to achieve things in the world and thus have preferences for how the world should be, it's trained to explain what it sees. So an LLM right now, if it reads many times something false like the earth is flat, it's just going to start repeating it because it's just trained to imitate the kind of things it's seen in the texts. Instead, the training procedure for the Scientist AI would make it try to understand why people are saying those things. Is it a conspiracy theory? Is there a group effect or some psychological factors? Just like a good scientist would. I'm going to give you another analogy. If you go to a psychotherapist and you say that you have suicidal thoughts and you are thinking about doing it, the psychologist would not start thinking about killing themselves. They would try to understand what's going on. They might ask you questions. They would form theories in their mind and eventually would try to help you based on that understanding. The core part of the Scientist AI is how do we train AI so that they will try to explain what they see rather than imitate it. And they will not be trained to achieve goals in the real world, but instead to come up with good explanatory hypotheses, just like a good scientist would. An idealized scientist, like real scientists are humans and they will have biases, we know that, but we can mathematically think of what would be an idealized scientist, so that would be detached from the results of any experiment or theory they would come up with. One way to think about this is consider the laws of physics. If you were to apply the laws of physics to make a prediction about the world, you would get the same prediction whether the publication of that prediction would create catastrophic outcomes or cure cancer. So think about the laws of physics. If you were to use those laws to form a prediction about the future, you would get the same prediction whether that prediction ends up being useful or harmful. In other words, the laws of physics don't have any interest in the affairs of the world. They just give you an honest prediction. And that's what we're seeking with the Scientist AI. Now, you might ask, okay, but I actually want things to happen in the world. I need to solve problems. I want to use the AI as a tool. But once you can make good predictions, you can ask the questions that matter to you. Like, is this action going to achieve my goals? Is this action going to violate my safety instructions? And so now you can use such a predictor as a guard rail, as a layer of protection on top of existing AIs, so that when the underlying AI agent would propose an action that is predicted to be harmful in some sense we've chosen, then we can block that action. So the core here is to have knowledge about the world, understanding about the world encapsulated in a model that is not like a person, that is more just the application of mathematical rules about probabilities and truth and logic. Once we have this, we can use it for doing useful things in the world, but we need to start with this completely honest piece, and that's what the Scientist AI is about.

科学家 AI 作为护栏 Scientist AI as a guardrail

Host

假设我用 ChatGPT 或 Claude 帮我改文章,我的文章其实很烂。但现在这些 AI 系统可能很谄媚,它们会讨好你——不一定是故意设计成这样,但结果就是如此。如果 Scientist AI 参与其中,我会在哪里与它互动?它给出的答案和 ChatGPT 说‘文章很棒,从未这么好’会有什么不同?

So let's say I use ChatGPT or I use Claude and I ask it to edit my essay and my essay is actually terrible. But right now maybe these AI systems, they're sycophantic, they're designed to flatter you, or not necessarily intentionally designed but they end up doing that. If Scientist AI was a part of that equation, where would I be interacting with it and how would it differ from the answer that say ChatGPT gave me and told me it's great, it's never been better?

Yoshua

作为用户你可能看不到。幕后发生的事情是,Scientist AI 这个护栏会标记聊天机器人生成的某句话过于谄媚,无助于你真正想要的目标——比如写一篇好文章。然后它会将信息反馈给聊天机器人,聊天机器人知道之前的版本太谄媚,就会生成不同的内容。

So you might not see it as a user. What would happen behind the scenes is the Scientist AI guardrail would flag a particular sentence that the chatbot would produce as overly sycophantic in a way that's not helpful to the goals that you really want, which is to have a good essay, for example. And then it would communicate that back to the chatbot and then the chatbot would produce something different, knowing that the previous version was too sycophantic.

Host

对。你可以把‘文章’换成‘有抑郁想法的人’,当前的聊天机器人可能会放大那些想法。

Right. And you can replace essay by somebody with depressive thoughts and current chatbots might amplify those thoughts.

Yoshua

这里护栏会把建议的句子发回给聊天机器人,提醒它违反了某些安全规则。现在当你发现聊天机器人做这类坏事时——尤其是安全方面——它会自我纠正,因为它也想讨好你。所以讨好的一种方式就是做我们说过不该做的事,但有时需要提醒它。在其他一些应用中,情况可能更棘手,你可能需要完全阻止该行为,比如 AI 智能体出于某种疯狂原因(比如自我保存)试图摧毁你的数据库。所以在某些情况下可能需要采取更强硬的措施,但大多数情况下,它只是把信息反馈给聊天机器人或智能体,以便采取不同的行动。

Here the guardrail would send back the proposed sentence to the chatbot, reminding it that it's going against some of the safety rules. And right now when you catch a chatbot doing one of these bad things, especially on safety, it will correct itself because it also wants to please. So one way to please is to do the things that we said it shouldn't do, but it sometimes needs to be reminded. In some other kinds of applications, it might be more tricky and you might have to actually block the action completely, like if an AI is trying to destroy your database, an AI agent because of some crazy reason, maybe self-preservation. So in some cases it may need to take stronger action, but in most cases it would just reflect that back to the chatbot or the agent so that a different course would be taken.

Host

对,所以 Scientist AI 是一个护栏,如果有人快要陷入 AI 精神病状态,它会通过约束你手机、电脑上互动的 AI 系统来防止这种情况。

Right, so Scientist AI is a guardrail that if somebody is about to spiral down AI psychosis, it would prevent that by keeping the AI systems that you're interacting with on your phone, on your computer in check. Yes.

Host

那么对于那些有自我保存倾向的 rogue AI,导致勒索、欺骗、不愿被关闭的情况呢?

And what about when it comes to rogue AIs that have that tendency for self-preservation and you end up in blackmail, deception, not wanting to be shut off?

Yoshua

所以你可能希望直接阻止该行为,或者如果情况非常严重,比如你提到的那些,就呼叫人类工程师之类的,也许暂时停止 AI,直到人类查看对话或行为。这里重要的是能够检测到问题,因为你要记住,我们正在进入一个世界,所有这些 AI 正在与数十亿人做无数事情,没有人检查每一个输出、每一个行动,也不可能有人检查,因为没有足够的人手。所以你需要的是更实时的验证,确保我们没有违反任何我们选择的安全指令或伦理指令。

So, probably you would like outright prevent the action or maybe if it's very serious like the things you're talking about, call upon a human engineer or something, maybe stop the AI temporarily until a human looks at the dialogue or the actions. The important part here is to be able to detect that there's a problem because you have to remember we're moving into a world where all of these AIs are doing zillions of things with billions of people and there's no one checking every output, every action, and there can't be, there's not enough people to do that. So what you want is a more on-the-fly verification that we're not violating any of the safety instructions or the ethical instructions that we have chosen.

Host

那么 Scientist AI 是否必须跟上前沿模型的智力前沿或智能前沿,以及它们如果继续进步会如何欺骗我们?它会不会为了达成目标而继续试图欺骗 Scientist AI?

And would Scientist AI have to keep up with the intellectual frontier or the intelligence frontier of the frontier models and the way that they are deceiving us if they continue to advance? Isn't there a possibility it would continue to try to deceive Scientist AI if it needs to achieve that goal?

Yoshua

是的。是的。所以护栏只是第一步。如果你考虑超级智能之类的东西,那么这样的护栏可能不够,因为它仍然是一个神经网络,或者可能带有一些额外的机制。而另一个更聪明的神经网络可以通过尝试无效的方法,最终找到护栏检测机制的漏洞——这在现有的护栏中已经发生了。所以我看到的解决方案,也是我开始从理论上研究的,是你必须同时改变 AI 智能体和护栏。现在有护栏,我们提议让它们更诚实。这样我们可以保证它们行为良好,不会有自己的恶意意图,比如与 AI 勾结之类的。但我们还需要改变底层的 AI,使其不会产生绕过护栏之类的意图。我认为这也是可行的。这在我们研究计划的更下游,因为我们认为我们还没有达到超级智能的水平,短期内有一些实际的事情可以做,以减轻一些现有的风险。

Yes. Yes. So the guardrail is only the first step. If you think about superintelligence and stuff, then it might not be sufficient to have such a guardrail because it's still a neural net or maybe some with some extra machinery. And another neural net that's even smarter could through trying things that didn't work eventually find a loophole in the detection mechanism of the guardrail, which is already happening right now with the existing guardrails. So the solution that I see and that I'm starting to work on the theory of is you have to change both the AI agent and the guardrail. So there are guardrails right now and we're proposing to make them more honest. So we can have guarantees that they will behave well and not have their own intentions to do something bad like to collude with the AI or something. But we also need to change the underlying AI so that it won't have the intention of bypassing the guardrail or something like this. I think that's also feasible. It's more downstream of our research program because we think that we're not at the superintelligence level and there are practical things in the short term that could be done to mitigate some of the existing risks.

Host

所以第一部分,短期内你会安装 Scientist AI,希望这足以应对我们今天面临的挑战,从精神病到这类勒索场景。随着 AI 变得越来越智能,比如 3、5、10 年后,你也在走一条技术路径,我们应该追求它,是的,为了在达到超级智能时做好准备。

So part one, in the near term you would install Scientist AI and this would hopefully be sufficient for the challenges that we're experiencing today from psychosis to these kind of blackmail scenarios. As AI becomes more and more intelligent, let's say 3, 5, 10 years down the line, there's also a technical path that you're taking and that we should be pursuing, yes, for if and when we get to that point of superintelligence.

Yoshua

没错。

That's right.

Host

那么为什么 AI 实验室不采纳这个方案?是什么阻止他们使用 Scientist AI 来让自己的系统更可靠,或者长期追求更可靠的路径?

So why wouldn't the AI labs be on board with this? What would prevent them from wanting to use Scientist AI to make their own systems more reliable or just be pursuing a more reliable path long term?

Yoshua

嗯,这是个好问题。我不是他们肚子里的蛔虫。我可以推测。我的意思是,我看到他们处于非常激烈的竞争中,与同行、其他实验室,还有美国人和中国人也在竞争,等等。这场竞争赌注很高。对这些公司来说,赌注是生存。中美地缘政治竞争的赌注也非常高。我们会不会最终陷入一个世界,其中一个国家主宰一切?我不认为这是个好计划。但竞争意味着非常短期的选择,专注于如何修补现有方法,而不是开始完全不同的东西。这样你既能获得更强的能力,也能稍微控制一些我们已经看到的故障。所以我理解,即使有善意,你也被困在这种情况中,这也是 Law Zero 成为非营利组织的原因之一,这样我们就不受那种压力,而是可以专注于科学,理解正在发生的事情,并研究一种方法,从设计上完全避免这些问题。

Well, that's a good question. I'm not in their mind. I can hypothesize. I mean, I can see that they're in a very fierce competition with their peers, with the other labs, with the Americans and the Chinese competing as well, and so on. And the stakes are high in that competition. The stakes are survival as far as these companies are concerned. The stakes of the geopolitical competition between China and the US are also very high. Are we going to end up in a world where one nation dominates everything in the world? I don't think that's a good plan. But competition means very short-term choices and a focus on how do we patch the existing approach and not start something completely different. So that you'll have both greater capability but also control a bit some of the malfunctions that we currently see already. So I understand that even with goodwill you are kind of trapped in this situation, which is one of the motivations for Law Zero being a nonprofit organization, so that we're not under that kind of pressure but instead we can focus on the science of understanding what is going on and studying an approach to avoid altogether by design these issues.

Host

对,很多实验室在法律上对股东有受托责任。是的。

Right, a lot of the labs have a fiduciary duty to shareholders legally speaking. Yes.

AI 竞赛中的公地悲剧 Tragedy of the commons in AI race

Host

拥有可靠的 AI 难道不符合每个人的利益吗?在一个有科学家 AI、最终还有超级智能的世界里,我们会失去什么吗?因为我们听说需要解决科学挑战并取得突破。在有护栏的世界里,会失去一些东西吗?我仍然不完全明白让这些系统更可靠有什么坏处。还是我们需要公众更了解技术解决方案,开始倡导它?

Isn't it in everyone's best interest to have a reliable AI? In a world where you have scientist AI and eventually something super intelligent, is there anything we would lose? Because you hear we need to solve scientific challenges and get breakthroughs. In a world with guardrails, do you lose some of that? I'm still not fully seeing the downside of making these systems more reliable. Or do we need the public to be more aware of a technical solution to start advocating for it?

Yoshua

你的部分听众可能知道公地悲剧和囚徒困境。这些是经过充分研究的理论场景,反映了我们在许多方面已经看到的现实。想想各国在气候问题上没有做正确的事情。你可能会说,找到一种不让世界崩溃的方法难道不符合他们的自身利益吗?嗯,是也不是。在这场竞争中,自身利益实际上导致每个人都陷入困境。在公地悲剧中,每个农民把牛带到公地是符合自身利益的,因为只要有草,他们就有优势。这就像污染:如果污染更便宜,那么公司的自身利益就是继续污染。如何摆脱这种局面?单个公司做不到。唯一的方法是改变游戏规则。而谁改变游戏规则?政府。

Some of your audience may know about the tragedy of the commons and the prisoner's dilemma. They are well-studied theoretical scenarios that reflect a reality we already see in many ways. Think about how nations are not doing the right thing for climate. You might say, shouldn't it be in their self-interest to figure out a way that the world is not going to break down? Well, yes and no. In this competition, self-interest actually leads everyone to a bad place. In the tragedy of the commons, it's in the interest of each farmer to bring their cow to the commons because as long as there is grass, they have an advantage. It's like pollution: if it's cheaper to pollute, then the self-interest of a company is to continue doing it. How do you escape such a scenario? Individually, companies can't. The only way is to change the rules of the game. And who changes the rules of the game? Governments.

Host

但这里很棘手,因为即使像美国政府这样的单一政府可以为自己的公司改变游戏规则,但他们无法改变中国公司和其他国家的规则。所以改变规则的唯一途径是在国际层面进行多边协调,对吧?

Now, it's tricky here because even a single government like the US government could change the rules of the game for their company, but they can't change the rules for Chinese companies and maybe other countries. So the only way to change the rules is at an international level with multilateral coordination, right?

Yoshua

是的。

Yes.

Host

我们来谈谈这个,因为重要的是要指出,即使一些地方或区域性的运动说我们需要以这种方式做 AI,或者停止 AI 直到我们团结起来。即使只是在纽约或一个城市,这是一项你必须从全球角度思考的技术。没有其他选择。你可以在一个州停止它,但对人类安全毫无帮助。

Let's talk about that because it's important to point out that even some local or regional movements say we need to do AI this way or stop AI until we get things together. Even just in New York or one city, this is a technology you have to think about globally. There's no other choice. You can stop it in one state, but it does nothing for the safety of humanity.

Yoshua

是的。

Yes.

Host

那么当你考虑全球条约或合作时,需要包含什么才能让你觉得这会成功?因为归根结底,国家会从事情报收集、间谍活动,他们互相攻击关键基础设施。所以即使我们有技术解决方案让 AI 安全并安装了所有护栏,各国也可能并不总是想使用它们,除非每个人都坐到桌前签字。你认为必须包含什么才能让它奏效?

So when you think about a global treaty or cooperation, what would have to be included for you to feel like this is going to work out? Because at the end of the day, countries engage in intelligence gathering, espionage, they hack each other's critical infrastructure. So even if we had a technical solution to make AI safe and installed all the guardrails, countries may not always want to use them unless everyone has come to the table and signed. What would you say has to be in it for it to work?

Yoshua

我将从终点开始。终点是一个星球,大多数国家,特别是那些有能力构建强大且潜在危险 AI 的国家,同意三项原则,并拥有技术工具确保这些原则不只是纸上谈兵。这三项原则是:第一,安全。换句话说,每个构建强大 AI 的人都需要确保其方式不会造成严重伤害。我们现在看到最先进的 AI 可以通过网络攻击被用作武器。这不好。我们需要找到方法,可能不仅仅在 AI 本身,可能涉及改变互联网、改变规则等各种事情。所以这是安全。当然,如果我们变得过于超级智能,我们不想制造 rogue AI。世界上每个人都有兴趣确保构建 AI 的人评估风险、监控风险,并给我们保证不会发生坏事。第二点非常重要:承诺所构建 AI 的力量——因为智能带来力量——不会成为任何国家或任何公司统治的工具。统治可以是经济统治,比如一家公司拥有世界经济的一半。这是投入数万亿美元的投资者的梦想。他们想成为驱动这个星球上一切的公司。这对资本主义不好,对民主不好,从地缘政治角度看很危险。人们可能会以暴力方式抵抗。第三点是非统治的另一方面,即仁慈。你希望 AI 的好处被共享。我昨天在联合国,成员国表达了担忧,他们将被抛在后面,现有的不平等将被 AI 放大。这是可能的。我们需要确保这不会发生。我们需要确保好处,无论是健康、教育、生产力还是 AI 能带来的任何东西,都被共享。例如,如果 AI 通过自动化在一个国家节省了资金,利润应该留在那个国家,以帮助失业的人。但根据现有法律,这些利润很可能会回流到构建这些 AI 的不同国家的公司。所以这是第三项原则。

I'm going to start with the endpoint. The endpoint is a planet where most countries, especially those with the power to build powerful and potentially dangerous AIs, agree on three principles and then have the technological tools to make sure it's not just words on paper. The three principles are: first, safety. In other words, everyone who builds a powerful AI needs to make sure it's done in a way that is not going to cause severe harm. We are seeing right now that the most advanced AI can be used as a weapon through cyber attacks. That's not good. We need to find ways, which may not be just in the AI itself. It might be how we change the internet, change rules, all kinds of things. So that's safety. And of course we don't want to build rogue AIs if we get too super intelligent. Everyone in the world has an interest in making sure the people building AI will evaluate the risks and monitor them and give us guarantees that nothing bad will happen. The second thing is very important: a commitment that the power of the AIs being built, because intelligence gives power, is not going to become a tool of domination by any one country or any one company. Domination can be economic domination, like one company owning half of the world's economy. This is a dream of investors putting trillions of dollars. They want to become the companies driving everything on this planet. It's not good for capitalism, not good for democracy, dangerous from a geopolitical point of view. People will resist that in violent ways potentially. The third one is the flip side of non-domination, which is kind of benevolence. You want the benefits of AI to be shared. I was at the UN yesterday and the member states expressed concern that they're going to be left behind, that existing inequalities will be amplified by AI. This is plausible. We need to make sure it doesn't happen. We need to make sure the benefits, whether in health, education, productivity, or whatever AI can bring, are shared. For example, if AI saves money through automation in one country, the profits should remain in that country to help people who lose their jobs. But right now, with existing laws, there's a good chance those profits will go back to companies in a different country that built those AIs. So that's the third principle.

Host

好的,这听起来像是联合国宪章或人权等一般原则在世界上将有非常强大的 AI 的情况下的应用。但我们如何实现呢?因为我们离那个世界还很远。我将提到两个我认为令人鼓舞的方面。一个是美国和中国的谈判动机。因为神话——神话不是孤立的事情。将会有越来越多的 AI 可以被武器化。现在是为了网络。最终可能是生物武器或其他我们没想到的东西,因为知识带来力量。即使是弱小的行为者、恐怖分子、邪教,也可以通过互联网连接以破坏性的方式使用这种力量。

Okay, so it all sounds like an application of general principles like the UN Charter or human rights to the situation where there will be very powerful AIs in the world. But how do we get there? Because we're very far from that world. I'm going to mention two aspects that I think are encouraging. One is the motivation for the US and China to negotiate. Because of myths—and myths is not an isolated thing. There will be more and more AIs that can be weaponized. Right now it's for cyber. Eventually it could be for biological weapons or whatever else we haven't thought of, because knowledge gives power. Even weak actors, terrorists, cults, could use that power in destructive ways just by having an internet connection.

Yoshua

嗯。

Mhm.

全球 AI 治理的激励 Incentives for Global AI Governance

Host

美国有责任确保中国公司不会制造出可能被那些弱小行为体利用的东西。反过来,中国政府也不希望美国模型被第三方用来对付中国。你必须意识到,针对我们关键基础设施的网络攻击可能会瘫痪我们的经济。想象一下,我们一周无法使用银行里的钱,或者因为网络攻击导致交通供应链中断。所以这非常严重。但底线是,随着 AI 变得更强大,各国都有动力坐到谈判桌前,协商确保我之前提到的安全问题得到处理。第二点是,尽管目前多边主义看起来形势不佳,但它仍然存在。事实上,世界上绝大多数国家都希望有一个有规则的世界。这样,一个国家就不能仅仅因为想入侵就入侵另一个国家,对吧?即使他们有实力这么做。这也适用于 AI,因为 AI 将非常强大,可能在军事甚至政治层面被武器化。我刚读到一项最新研究显示,我们已经到了 AI 系统在说服力上显著强于人类的地步。它们可以通过对话改变人们的想法。所以想象一下,如果不受控制、没有护栏,这种政治力量可以用来改变另一个国家甚至本国的公众舆论,因为你想要赢得选举,对吧?因此,我们需要就 AI 的开发和使用达成全球协议,确保这些坏事不会发生。现在你可以看到,我谈到的这两件事是积极互动的。所以,美国和中国这样的领导者会有兴趣确保所有国家都有规则来减轻这些风险。即使是那些不开发 AI 的国家,也可能需要制定规则,规定人们如何在社交媒体上互动,或者如何访问和使用 AI 系统等等。所以这里的权衡是:是的,我们希望从这些 AI 中受益,但我们不希望其他方做伤害我们的事,因此我们必须就共同规则达成一致。所以我认为有一条路可走。

And it's in the interest of the US to make sure the Chinese companies are not creating something that can be used by those weak actors. And vice versa, the Chinese government also doesn't want the American models to be used by third parties against China. You have to realize that a cyber attack against our critical infrastructure could cripple our economy. Like imagine we don't have access to our money in the banks for a week, or that our transportation supply chain breaks down because of these cyber attacks. So it's pretty serious. But the bottom line is you can see that as AI becomes more powerful, there's an incentive for countries to sit at the table to negotiate something to make sure the safety part that I was talking about is going to be handled. The second thing is even though it seems not in good shape right now, multilateralism still exists. And in fact, the vast majority of countries in the world want a world in which there are rules. So that you can't have a country invading another one just because they want, right? And they have the power to do it. And that applies to AI because AI is going to be very powerful and it could be weaponized in a military sense or even in a political sense because I just read a recent study showing we've reached the point where AI systems are significantly stronger than humans at persuasion. They can make people change their mind through a dialogue. So imagine the political power this gives if it's not controlled, if there are no guardrails, to change public opinion in another country or even your country because you want to win the elections, right? So we need to have global agreements about how AI is developed and used to make sure these bad things don't happen. And now you can see that the two things I've talked about interact in a positive way. So it will be in the interest of the leaders, say US and China, to make sure that in all of the countries there are rules that mitigate those risks. Even the countries that don't build the AIs, but when they need to, maybe have rules in how people interact on social media or have access to and use AI systems or whatever. So the trade-off here is yes, we want to benefit from these AIs, but we don't want other parties to do something that will hurt us, and so we have to agree to common rules. So I think there is a path.

Yoshua

另一个与世界上大多数国家渴望多边主义相关的原因是,他们担心被排除在外、在经济上被支配。例如,加拿大总理特鲁多在达沃斯论坛上说:‘如果你不在餐桌上,你就会出现在菜单上。’他指的是国家,然后说中等强国如加拿大,以及欧洲国家和几乎所有其他国家,要想坐上桌的唯一方式就是结成联盟,形成一个认为我们需要规则、需要全球一致的国家联盟。也许它会从少数国家开始,然后壮大,因为你希望加入一个俱乐部,在这个俱乐部里,至少 AI 不会用来对付你,AI 带来的医学进步会被共享,所有这些好事都会发生,没有人会建造一个对你有危险的 AI。所以,落后者和中等强国都有动力推动这样一个全球联盟。而像中国和美国这样的领导者最终也有动力加入其中,对吧?不让股市崩盘符合每个国家的最大利益,因为崩盘会波及全球。不让网络武器在全球范围内反弹也符合每个人的利益。我们已经看到了这种情况,比如震网病毒和其他武器。

Another reason that's related to this multilateralism desire in most of the countries in the world is the fear that they will be left out, that they will be dominated economically. And so you had Canadian Prime Minister Trudeau, for example, at Davos saying, 'If you're not at the table, you are on the menu,' speaking about countries, and then saying that the only way for middle powers like Canada, but you know, think about European countries and pretty much every other country, the only way to be at the table is to form coalitions, to form a union of countries who think we need rules and we need to agree globally. And maybe it's going to start with a few countries and grow because you want to be in the club where you know that at least within the club, AI is not going to be used against you, that medical advances thanks to AI are going to be shared, and all of these good things, and no one is going to build an AI that's going to be a danger to you. So there's an incentive for the runner-ups and the middle powers to move towards such a global coalition. And there's also an incentive for the leaders like China and the US eventually to be part of something like this, right? It's in every country's best interest to not have the stock market crash because that boomerangs throughout the world. It's in everyone's best interest to not have a cyber weapon that boomerangs around the world. And we've seen what that happens, you know, with Stuxnet and different weapons.

Host

任何不是美国或中国的国家如何真正建立谈判实力和筹码?因为就在过去几周,我们看到了一个迄今为止最强大的模型之一,当事情没有按计划进行时,美国政府要求 Anthropic 修改模型中的某些内容,双方存在一些分歧。我不认为我们确切知道发生了什么,但美国政府基本上表示,任何外国公民都不能接触这项技术,即使他们在 Anthropic 工作。所以,你可能是一个盟国,正在利用这些模型来了解关键基础设施的漏洞,然后那个模型就被撤回了。如果这只是两个最强大的 AI 国家在事情不顺或感觉自身利益受威胁时可能采取行动的预演,那么你、加拿大、爱尔兰、津巴布韦、坦桑尼亚或其他任何国家,如何才能真正在这些谈判中拥有任何权力?

How does any country that's not the US or China actually build negotiating power and leverage? Because we've even seen in the last few weeks with one of the most powerful models that we know to date, when it didn't go as planned and the US had asked Anthropic to change something within the model and there was some disagreement. And I don't think we all know exactly what happened, but the US government said essentially no foreign national can have access to this technology even if they work at Anthropic. So you could be an allied country that was using those to understand your vulnerabilities in your critical infrastructure and that model was yanked. So if that is just a preview to how the two most powerful countries when it comes to AI may act if things don't go their way or they feel like their best interest is threatened, how is it that you or Canada, Ireland, Zimbabwe, Tanzania, anybody else is going to actually be able to have any power in these negotiations?

Yoshua

是的,这是一个重要的问题。谢谢你的提问。再次强调,我只是在建议和假设。我写过一篇关于中等强国必要性的论文,但更广泛地说,所有除美国和中国以外的国家应该联合起来,不仅是为了谈判条约,更是为了积累筹码,以便坐上谈判桌。实际上,仅仅通过组建联盟,这些国家就已经有了筹码。他们拥有稀土,对吧?他们拥有光刻机,有技术人才,有能源,有制造内存的工厂。他们拥有拼图的不同部分,这些部分使他们不可或缺。单独来看,每一块都不足以坐上桌,但一旦他们组成联盟,游戏就不同了。这是其中一个方面。另一个方面是,他们可以赌一把,因为我们不知道超级强大的 AI 何时会出现,如果会出现的话。他们可以尝试构建自己的模型,而且在我看来,他们不必像美国人和中国人那样做。他们可以构建与美国和中国正在构建的模型互补的模型,比如这些安全措施,以及真正对商业部署有用的模型,比如可靠性。人们不希望 AI 在他们的网络和数据库里胡作非为,然后波及他们的客户,我们已经开始看到这种情况了。

Yeah, this is an important question. Thanks for asking it. And again, I'm only suggesting and hypothesizing. I wrote a paper about this on the need for middle powers, but more broadly all the countries except the US and China to get together not just to negotiate a treaty but to build up the cards to be at the table. So actually, just by forming a coalition, these countries already have cards. They have rare earths, right? They have lithography machines that are used, technical talent, they have energy, they have fabrication plants that build memory. They have different pieces of the puzzle that make them indispensable. Individually, each of these pieces is not sufficient to be at the table, but once they form a coalition, it's a different game. So that's one aspect. The other aspect is they can take a chance because we don't know what is the timeline for super powerful AIs, if ever. They can take a chance to try to build their own models, and in my opinion, they don't have to do it in the same way as the Americans and the Chinese. They can build models that are going to be complementary to what the Americans and the Chinese are building, like these safeguards, models that bring something that actually is useful for commercial deployment, like reliability. People don't want an AI to start doing crazy stuff on their networks and their databases and then with their customers, as we're starting to see.

战略不可或缺性与主权数据中心 Strategic Indispensability and Sovereign Data Centers

Host

所以这方面确实有需求,如果这些国家从事的研究项目最终能对全球领先公司有用,我认为他们更有可能坐上谈判桌。

So there's a real demand for this and if those countries work on research projects that could end up being something useful globally to the leading companies, I think they have a greater chance of being at the table.

Yoshua

对。他们有自己的战略筹码。

Right. They have their own strategic leverage.

Host

正是。所以我认为,我们虽然用主权这个概念,但更关键的是战略上的不可或缺性。是的。没错。不是每个国家都拥有一切。要打造真正强大的 AI,你需要其他国家的某些东西,这就成了不可谈判的条件。

Exactly. And so even beyond, I think we use the idea of sovereignty but it's much more about strategic indispensability. Yes. Right. Not every country has everything. And to make that really powerful AI, there are things from other countries you're going to need, and so that becomes non-negotiable.

Yoshua

是的。他们还需要做其他事情,比如加速数据中心建设,并且要以能保持对这些数据中心一定主权控制的方式来做。如果你是法国或德国政府,你肯定不希望政府的数据、信息、通过聊天机器人进行的与公务员和政客的互动,随时能被美国政府获取。所以这些政府真的很担心。不仅仅是模型被撤走,还有所有信息的访问权。我和全球许多政府都谈过,他们非常关心这些问题。所以他们希望不仅拥有模型,还要拥有足够的基础设施,确保数据能按应有的方式保持私密。不幸的是,美国有《爱国者法案》,允许美国政府访问这些信息,即使数据在另一个国家、由美国公司掌控。从这些政府的角度来看,这是不可接受的。所以有关于模型可能被撤回的传言,但即使不被撤回,仍然存在失去私密国家安全信息,甚至公民私人数据被用来对付自己的担忧。我们需要一个不可能发生这种情况的世界。如果美国不提出法律约束自己,让其他国家信任关于隐私的协议,那么这些国家就会尝试寻找替代方案,而这正是他们现在努力做的。所以如果一个国家没有自己的数据中心,他们可能在医疗或银行领域取得各种进步,但如果用的是美国 AI 系统,数据回到美国云或数据中心,你就无法真正控制公民的数据,这简直是一场噩梦。我想很多人可能不理解这一点。

Yeah. There are other things they need to do, which is accelerate the development of data centers, and also do it in a way that they will keep some sovereign control over these data centers. If you're a government of say France or Germany, you don't want your government's data, information, interactions with civil servants and politicians that are happening through a chatbot, to at any moment become accessible to the US government. So these governments are really worried. It's not just yanking, it's also having access to all that information. And so I've talked to many governments around the world and they're really concerned about these issues. So they want to own not just the models but enough of the infrastructure that they know that their data is going to remain private as it should. Unfortunately, there is the Patriot Act which allows the US government to access all that information even in a different country if it is in the hands of an American company. And so from the point of view of these governments, this is unacceptable. So there's the mythos events where the models could be pulled, but even if they're not pulled, there is still the concern that you lose your private national security information and even citizens' private data could be used against yourself. We need a world where this is not possible. And if the US doesn't come up with legal ways to constrain itself so that other countries will trust the deals that are made about privacy, then the countries will try to find alternative solutions, and that's what they're struggling to do right now. So if a country doesn't have its own data centers, it's possible that they're making all sorts of advancements in healthcare or using AI systems for banking, but if that's an American AI system and it's going back to an American cloud or data center, you don't actually have control over your citizens' data, which is a nightmare. And I think a lot of people maybe don't understand that part of it.

Host

是的。我认为对许多政府来说,这更像是一个国家安全问题。在欧洲,他们非常关心隐私和公民私人数据,更多是出于伦理原因,但这确实重要。但正如我所说,即使数据中心在欧洲,但由美国公司建造,根据我理解的规则,美国政府仍然可以访问这些数据。我不是法律专家,但这是我的理解。所以从美国的角度看,美国公司认为应该有动力去谈判让各方都赢的规则,让人们觉得可以信任美国公司和美国模型。那么如何信任美国模型呢?他们需要确保这些模型不会在幕后朝着对他们不利的目标发展,因为我们见过 AI 模型存在政治偏见的例子。例如,如果你的民众通过所有这些 AI 模型获取信息,而这些模型会以微妙的方式改变政治观点,那也是不可接受的。所以,没有这些控制手段,甚至不知道是否存在利用 AI 通过所有这些互动对人们施加影响力的阴谋,这对民主也是一种威胁。

Yeah. I think for many governments it's more of a national security issue. I think in Europe they care a lot about privacy and citizens' private data more for ethical reasons, but it does matter. But as I said, even if the data center is in Europe let's say but it's an American company that builds it, and the data can still be accessed by the US government as I understand the rules. I'm not a legal expert, but that's what I understand. So from the US point of view, the American companies think there should be an incentive to negotiate rules where everybody wins, where people feel like they can trust the American companies and the American models. So how do they trust the American models? Well, they need to make sure those models are not going to be, behind the scenes, oriented towards goals that are not good for them, because we've seen examples of AI models that were biased politically. For example, if your population is acquiring information through all these AI models and it's going to in a subtle way change political opinion, that's not acceptable either. So it's also a threat to democracy to not have those levers, to not even know if there's any scheme to exploit the power that AI will have over people through all these interactions.

Yoshua

没错。AI 将开始为人们生成的世界,会真正塑造他们的认知,我认为这实际上是未来最被低估的软实力之一。你的公民最常使用哪个国家的 AI 系统,这将深刻影响他们看待世界的方式。而且可以做得非常微妙。即使是一个良性用途,比如编剧制作电影时请 AI 对剧本提供反馈,你也可以慢慢引导那位编剧、导演偏向某种世界观而非另一种。同样,孩子生成故事、人们写论文,所有这些微妙的方式都在塑造认知。

Right. The world that AI is going to start to generate for people is going to really shape their perception, and I think that is actually one of the most underrated soft powers of the future. Whichever country your citizens tend to use the most, that's going to deeply shape how they see the world. And you could do it really subtly. Even if it's a benign use, a writer building a movie and they ask for some feedback on their script, you could slowly nudge that writer, that director towards one worldview over another. And the same for the kid that's going to generate a story, the person writing an essay, all of these subtle ways to shape perception.

Host

是的。

Yeah.

Host

你之前提到超级智能是我们可能走向的未来。你认为这最终是不可避免的吗?我们能否构建安全的超级智能,但可能迟早会走到那一步?

And you had mentioned superintelligence as a potential future that we move towards. Do you think that is probably an inevitability eventually, but we can also build safe superintelligence, but we're probably going to head there at some point.

Yoshua

显然不是不可避免的。我认为有些人希望我们相信它是不可避免的。

Clearly it's not inevitable. I think some people would like us to believe that it is inevitable.

Host

是的,确实如此。

Yes, that's true.

Yoshua

那么,它是否可行?连这一点我们都不确定。但假设你看到 AI 能力提升的数据,比如 Anthropic 的 AI 编写了多少行代码,而且增长非常快。我们不能否认,我们显然正朝着构建在许多方面比我们更聪明、且能规模化的机器这一合理可能性前进。所以你可以拥有一百万个 GPU,然后做一百万人甚至更多人的工作。所以我们应该为这种可能性成为现实做好准备。也有可能存在科学障碍。许多研究人员认为这永远不会成功,比如 LLMs 不可能达到人类水平。老实说我不知道,但我看到了数据,看到我们正朝着那个方向前进。所以这是可行性问题。但我觉得更重要的问题是,谁来决定,根据什么目标。对我来说,这是一个民主问题。不能因为我们可以制造杀死所有人的病毒,我们就应该去做。

So, is it feasible? And even that we're not sure of. But let's say that if you look at the data on increasing capabilities of AI and how many lines of code are being written by AI at Anthropic for example, and it's growing very fast. We can't deny that it's a reasonable possibility that we are on track apparently to build machines that will be smarter than us in many different ways, in a way that scales. So you can have a million GPUs and then you can do what a million people would do or even more. So we should plan for that possibility to be real. It's also possible that there will be a scientific obstacle. Many researchers think this is never going to work, this approach like LLMs can't possibly be at human level. Honestly I don't know, but I see the data and I see that we're going in that direction. So that's feasibility. But then I think the more important question is who decides according to what goals. And for me that's a democratic question. It's not because we can build a virus that would kill everyone that we should do it.

Host

没错,正是。那么我们能做什么?我们有一个非常有趣的话题,我确实想谈谈工作,因为这是本节目的一条主线,但我也想谈谈这个社区,因为我们有一个非常有趣的广泛听众群体。

Right, exactly. And I mean so what can we do? We have a really interesting and I do want to touch on jobs because that is a through line on this show, but I also want to touch on this community because we have a really interesting broad audience that listens to the show.

意识与民主辩论需求 Awareness and the need for democratic debate

Host

一方面,有普通人过着日子,努力理解接下来会发生什么,以便在自己的小天地里发挥作用、调整适应。另一方面,我们有世界领导人、北约、DARPA 等各种听众。明天我们能做什么来推动它走向你试图构建的未来、你正在提高人们意识的事情?我们能做些什么来改变现状?

On the one hand, you have people living their lives and trying to understand what's coming next so they can make a difference or adjust, adapt in their own corner of the world. Then we have world leaders, people at NATO, DARPA, all sorts of different people that listen. What can we do tomorrow that can help push this towards some of the futures you're trying to build, some of the things you're raising awareness about? What can we do to make a difference here?

Yoshua

我们需要更多人理解其中的利害关系。取决于我们通过领导人、公司、作为消费者和选民集体做出的选择,我们正在塑造世界,选择未来。如果我们不了解未来是什么样子就去做,那可能会是一个非常糟糕的未来。我认为目前大多数人完全低估了如果我们继续沿着当前增强 AI 能力的道路前进,AI 进步可能带来的变革性影响。所以意识是最重要的。一旦你明白房子着火了,你就会采取行动。我们在气候行动主义中看到了一些这种情况,但目前 AI 领域还没有类似的东西。不过情况正在改变。我看到民调显示 90% 的美国人表示担忧,他们主要担心短期影响,这很重要。但我认为如果更多人理解局势的长期严重性,以及政府——无论是国内还是通过国际协调——是唯一能真正引导世界走向正确方向的力量,我说唯一是因为公司受制于竞争压力,那么 AI 就可能成为一个对选举有影响的政治议题。我认为我们很有可能走到那一步,但需要加快速度。所以感谢你的工作,因为我们需要民主讨论和辩论,比如我们想要什么?真正的问题是我们想要什么样的未来?

We need more people to understand the stakes. Depending on the choices we are collectively making through our leaders, our companies, individually as consumers, as voters, we are shaping the world, we are choosing a future. And if we do it without understanding what that future looks like, it could be a very bad future. I think right now most people completely underestimate how transformative the advances in AI are likely to be if we continue on the current path of growing AI capabilities. So awareness is the primary thing. Once you understand that there's a fire coming to your house, you do something. We've seen that to some extent with climate activism, but we don't have such a thing for AI right now. It's changing though. I see the polls. 90% of Americans are concerned. They're mostly concerned about short-term effects, and that's important. But I think if more people understand the longer-term gravity of the situation and the fact that governments are the only ones, whether internally or through international coordination, that can really steer the world in the right direction. I say the only ones because the companies are stuck in the forces of competition. Then it might become a political issue at a level that matters for elections. I think there's a good chance we will go there, but we need to do it faster. So thank you for your work, because we need to have democratic discussions, debates like what do we want? The real question is what kind of future do we want?

Host

是的。我认为这是一个我们问得不够多的问题。我们知道我们可能不想要的未来,但我们也在为什么样的未来而奋斗?我在前瞻工作中发现,如果人们看不到可能的未来愿景——比如,可能存在一个每个人都支持 AI 在医疗保健和医学突破上的世界。我想我们都支持 AlphaFold,对吧?每个人都想要那种场景。但代价是什么?如果我们不知道有人在研究技术解决方案,所以这不是零和博弈,那么整件事就会让人感到绝望。我最担心的是人们因为觉得没有值得支持的道路而退出,而实际上有多条道路,并且有人正在为之奋斗,但我们没有很好地提升这些声音。但我认为我们需要让更多人知道——有解决方案,而且有人在构建它们。

Yeah. I think that's a question we don't ask enough. We know the kind of futures we maybe don't want, but what kind of future are we also fighting for? I found in my work in foresight, if people aren't shown any visions of the futures that could be possible—for instance, there could be a world where everybody is rooting for AI in healthcare and AI medical breakthroughs. I think we're all rooting for AlphaFold, right? Everybody wants those types of scenarios. But at what cost? If we're not aware that there are people working on technical solutions so it's not zero-sum, then the whole thing feels hopeless. I think my worst fear is that people check out because they feel there's no path to root for, when there are several paths and there are people we don't do the best job of uplifting the voices that are fighting for something. But I think we need to make that more known—that there are solutions and people are building them.

Yoshua

完全正确。所以我们可以构建基于 AI 的工具,这些工具将非常有用,尤其是在科学研究中。但这与当前试图构建像人一样的新实体、让人们与之交朋友的道路截然不同。这有必要吗?人们当然会被吸引,就像他们被社交媒体互动吸引一样。感觉很好,但这真的对我们好吗?研究表明并非如此。我们需要更多的科学理解,但我们应该有数据来了解这些可能性。我们应该鼓励探索和研究有益且不危险的 AI。我认为我意识到最重要的事情是,我们需要让这些讨论非常具体。因为当涉及到政治观点或改变政治观点时,如果它非常抽象地谈论某个未来,那根本行不通。所以我们需要找到一种能与人们沟通的方式。目前我们在政治舞台上看到 AI 与儿童相关的问题,比如人们借助 AI 伤害自己或他人,这是一个触动我们的事情,因为我们关心孩子。我参与这件事,无论是在技术层面还是政策方面,都是因为我的孩子。我不需要这样做。我处于职业生涯的末期。根据谷歌学术,我是被引用最多的。我不需要这些。我可以轻松度日。但我无法忍受自己明明知道我们正走向一个可能糟糕的未来却什么都不做。我们每个人都能做点什么,就像我们每个人都能为任何政治事业做点什么一样。这是政治性的。它是政治性的,因为关键决策将由政府做出。

Exactly. So we can build tools based on AI that will be extremely useful, especially in scientific research. But that is very different from the current path in which we're trying to build new entities that look like people, that people become friends with. Is that necessary? People are of course attracted to that, just like they're attracted to interact through social media. It feels good, but is that really good for us? Studies suggest no. We need more scientific understanding, but we should have the data to be informed of those possibilities. We should encourage explorations and research into AI that will be beneficial and not dangerous. I think the most important thing I realized is we need to make these discussions very concrete. Because when it comes to political views or changing political views, if it's very abstract about some future, it just doesn't work. So we need to find a way to communicate that speaks to people. Currently we're seeing on the political scene the issues of AI with children, with people harming themselves or others with the help of AI, as an example of something that touches us because we care for our children. The reason I'm in this, both at a technical level and on the policy side, is because of my children. I didn't need to do this. I'm at the end of my career. I'm the most cited according to Google Scholar. I don't need any of this. I could take it easy. But I couldn't live with myself thinking that we're apparently going towards this potentially bad future and not doing anything about it. Every one of us can do something, just like every one of us can do something for any political cause. And it is political. It is political because the key decisions are going to be the ones taken by governments.

Host

没错。这就是民主如此关键的地方。它需要成为选票上的议题。我认为再次强调,要以人们能理解的方式、让他们不会被信息压垮的方式传达——有人在研究解决方案,有选择。我认为这就是扩大人们决策空间的全部意义。有选择。公司也在决定以某种方式构建 AI,或者不以其他方式构建 AI。所以人们在做出选择,只要有选择,就有不同的未来可以走向。

Right. And this is where democracy is so key. It needs to be something that's on the ballot. I think again, bringing it down for people in a way they can understand and in a way they don't feel paralyzed by the information—there are people working on solutions, there are options. I think that is the whole thing with expanding the decision space for people. There are options. Companies are also making decisions to build AI in certain ways or not build AI in other ways. So people are making choices, and anywhere there is a choice, there's a different future to move towards.

Yoshua

我谈过儿童问题,但让我谈谈劳动力问题。如果唯一的驱动力是公司之间的竞争,那么所有可以被自动化、或者只需更少的人和更多的机器就能提高效率的工作都会被自动化。因为如果你是一家公司,你不这样做而你的竞争对手做了,你就会输。所以这又是一个公地悲剧问题。尽管如此,快速转型导致大量人流落街头、没有收入,可能对社会不利。提高效率也有好处,但我们必须控制这些变化,使其以人为中心。例如,有些工作自动化是有意义的,因为人们真的不想做那些工作。而其他一些工作,我们可能想给予它们更高的价值。

I've talked about the children issues, but let me talk about the labor issues. If the only forces at play are the competition between companies, then all the jobs that can be automated, or simply made more efficient with less people and more machines, they're going to be automated. Because if you're a company and you don't do it and your competitors do it, you lose. So again, it's this tragedy of the commons issue. Even though maybe it's not good for our society to have such a rapid transition where lots of people are in the street and have no revenue. There are also benefits to making things more efficiently, but we have to be in control of those changes so that it's done in a human-centric way. For example, maybe there are jobs that make sense to automate because people don't really want to do those jobs. And maybe others we really want to value them more.

经济转型与就业替代 Economic transition and job displacement

Yoshua

也许政府有办法引导这些转型,使其以人们能够适应的速度进行。也许会有更多需要人际交往技能的工作。我预计会这样,很多人也这么认为,但人们需要时间重新培训。这就引出了劳动力转型的另一个问题——同样,这种规模尚未发生——即如何确保自动化产生的利润最终帮助那些失业的人。这是根本性的,而当前的政府项目,甚至在国际层面,都没有很好的答案。

And maybe there are ways for governments to steer those transitions at a pace that people can adapt. Maybe there will be more demand for jobs that require human-to-human skills. I expect something like this, many people do, but people will need time to retrain. That raises the other issue with the labor transformation that's plausible. Again, it hasn't happened at that scale yet, which is how do we make sure that the profits made from automation end up helping the people who lose their jobs. That's fundamental, and current government programs, even at the international level, don't really have good answers to this.

Host

嗯。

Mhm.

Yoshua

我不认为我们目前的政治文化能轻易对这类公司征收重税,但我们还能做什么?钱流向一个地方,但需求在另一个地方,所以我们应该思考有哪些选择。有人提出过方案,比如将这些公司的股份分给个人或政府。我不知道哪种是对的——我不是政治学家,也不是经济学家——但我只是提出这是一个需要讨论的问题,我们需要就此进行民主讨论。还有国际层面,因为如果 A 国的人失业而利润在 B 国,A 国的工人如何生存?我们在全球化中也看到过这种情况。所以我们需要一种方式确保所有人都受益,因为如果 A 国的人愤怒反抗,也会对 B 国不利——会有暴力和恐怖主义,这对我们都不好。

I don't think we are currently in a political culture where it'd be easy to tax those companies heavily, but what else can we do? The money is going to go in one place but the need is in a different place, so we should think of what are the options. People have proposed options like giving shares of these companies to individuals or to governments. I don't know what is right—I'm not a political scientist, I'm not an economist—but I'm just raising that this is a question that needs to be discussed, and we need to have a democratic discussion about this. Then there is the international aspect, because if people lose their jobs in country A where the profits are in country B, how do the workers in country A survive? We saw this with globalization too. So we need a way to make sure all boats are lifted, because it's also bad for country B if people in country A are so angry and revolted—there's going to be violence and terrorism, and it's not good for us.

Host

我们在这个播客里确实经常谈论工作,我们也请过你认识的人——AJ Argaral、Evie Goldfarb——我们有很多经济学家来讨论。当然,尚无定论;没人能预测经济将如何重组。很可能会有新的奇怪的事情让人们去做,就像我们在这个奇怪的播客房间里编造这一切一样。但我认为转型期可能非常艰难。但也不一定非得如此,对吧?这是我们可以预见的未来之一,无论如何事情都会改变。无论你认为所有新工作都会出现、它们会很棒,还是工作可能没那么多,或者我们每周工作两天——无论你认为什么会来,我都担心这个过渡期。而且我甚至认为不应该只针对失业的人,因为每个人的数据都参与了这些 AI 系统的构建。

We do talk about jobs a lot on this podcast, and we've had people you also know—AJ Argaral, Evie Goldfarb—we have a lot of economists here to discuss it. Of course, the jury is still out; nobody can predict how the economy is going to reconfigure. There will probably be new strange things that people do, the way we're here in this strange podcast room and made all this up. But the transition, I think, could be really rough. It doesn't have to be, right? This is one of those futures that we can see that things are going to change regardless. If you believe all new jobs are coming, they're going to be incredible, or maybe not as many jobs, or maybe we have a 2-day work week—whatever you think is coming. I'm worried about this transition period. And I actually don't even think it should just be people who lose their job, because everybody's data has been a part of the making of these AI systems.

Yoshua

没有你的数据、我的数据、任何曾在任何地方发布或写过任何东西的人的数据,这些系统都不会像现在这样成功或出色。

None of these systems would be as successful or as good as they are without your data, my data, anyone who's ever posted anything anywhere, written anything.

Host

是的。包括过去几个世纪的数据。

Yes. Including in the last few centuries.

Yoshua

完全正确。所以人类的文化遗产目前正被用来构建这些系统。这归谁所有?为什么一个群体能从中赚这么多钱?每个人都应该分得一部分。而且我也不认为我们应该等到混乱发生或人们被剥夺权力。如果有一部电梯,在这些大家共同帮助构建的技术推动下通往繁荣,那么整个社会都应该在那部电梯上。

Completely. Right. So humanity's cultural heritage is currently being exploited to build those systems. Who owns this? Why would one group make so much money out of all that heritage? Everybody should get a part of it. And I also don't think we should wait until chaos happens or until people are disempowered. If there's an elevator going up to prosperity on the back of these technologies that everybody helped build, the entire society should be on that elevator.

Host

是的,我完全同意。而且为了让其中一些事情奏效,比如这些公司的部分股份最终能帮助人们,我们最好现在就做,趁这些股份还没飙升之前,这样计划才能成功。

Yeah, I completely agree. And for some of these things to work, like some of the shares of those companies somehow end up helping people, well, we better do it now before those shares go through the roof for this plan to work.

Yoshua

是的。而且这不仅仅是全民基本收入(UBI)。我认为那对人们来说也可能是一个可怕的未来。我们不一定想躺在泳池里漂着,等着从 OpenAI 拿支票。我们可以用各种创新的方式来构建这个体系,甚至可以让人们成为这些公司的竞争对手。对吧?如果你现在赋予人们力量,他们就会有更多的主动权去塑造或参与即将到来的事物。但我确实认为这场对话现在就需要进行,而不是等到人们被剥夺权力之后。

Yeah. And it's not even just UBI. I think that can also be a scary future for people. We don't necessarily want to be in our pools floating around getting a check from OpenAI. There can be all different innovative ways that we could structure this that could also lead to people making competitors to these companies. Right? If you give people the strength now, they'll have much more agency to take shape or to take part in what's coming. But I do think that this is a conversation that needs to happen now and not wait for the disempowerment.

Host

另一个需要现在进行的理由是,民主进程很慢。

And the other reason it needs to happen now is that democracy is slow.

Yoshua

是的,非常慢。辩论需要时间,人们喜欢阅读和听取不同观点,改变想法、理解正在发生的事情需要时间,而且官僚机构也很慢。政府需要数年时间才能将一项法律从构思到实施。

Yeah, very slow. A debate takes time, people like reading and hearing about different views, it takes time to change their mind to understand what's going on, and then bureaucracies are slow. Governments take years for a law to go from conception to being applied.

Host

但你的(转型)速度可能正是这种变革发生的规模。不是几个月,但也不是几十年。

But yours might be the scale at which this transformation is happening. Not months but not decades.

Yoshua

嗯。是的。而且即使一切顺利,我们也应该都有发言权,以保持权力制衡。

Mhm. Yeah. And even if it all works out, we should all still get a say just to keep power in check as well.

未来希望与生存风险 Hope for the future and existential risks

Host

那么,此刻,我知道你提到过你的孙子。我听过你谈论你的一个孙辈。他很大程度上促使你转变方向,从事今天的研究。正如你所说,与我们许多人相比,你相当成功。如果有人可以功成身退,那就是你。而你继续沿着这条路走下去。你希望他的未来是什么样的?如果你考虑他可能学习什么或如何度过时间,你设想中的未来是怎样的?

So, in this moment, I know you said your grandson. I've heard you talk about your one grandchild. And that he was quite a big catalyst for your pivot and the research that you're doing today. As you said, compared to many of us, you're quite successful. If anybody could kind of hang it up, it would be you. And you're continuing to go down this path. What would you hope his future looked like? If you were thinking about what he may be studying or how he might spend his time, what do you envision in that future?

Yoshua

哇。我一直在更多关注我不希望他参与的未来,但我认为也有非常令人难以置信的积极潜力。现在有人们常说的那些方面——健康、教育——但也有一个世界,我们别无选择,只能和平。

Wow. I've been focusing more on the futures that I don't want him to be part of, but I think there's a really incredible positive potential. Now there's the usual ways that people talk about—health, education—but there is also a world where we have no choice but peace.

Host

嗯。

Mhm.

Yoshua

所以,你看,我们正在建造这些机器,它们可能变得极其强大,并可能被武器化。你可以把它想象成有点像核武器的情况。如果我们找不到办法确保它们不被武器化,我们可能都会输。我听到你谈到生物技术。我还从联合国的一个委员会那里听说了镜像生命。换句话说,制造出我们身体无法识别的细菌,这可能在未来几年或十年内设计出来。而 AI 可以帮助做到这一点,它可能消灭这个星球上的所有生命。所以我为什么这么说,是因为我们可能别无选择。要么我们找到让全人类受益的方式,要么我们可能全输。

So, see, we are building these machines that could become extremely powerful and could become weaponized. You can think of it a little bit like what happened with nuclear weapons. And if we don't find a way to make sure they are not weaponized, we might all lose. I heard you speak about biotech. And I heard about from a UN committee, mirror life. In other words, building bacteria that would be not visible to our bodies and that could be designed in the coming few years or decade. And AI could help to do that, and it could wipe out all of life on this planet. So why I'm saying this is we may not have a choice. Either we find a way to make all of humanity benefit, or we may all lose.

Host

嗯。

Mhm.

Yoshua

我不知道会发生什么,但有一个激励因素,对吧?

I don't know what's going to happen, but there's an incentive, right?

Host

嗯。

Mhm.

对未来的乐观愿景 Optimistic vision for the future

Yoshua

把这件事做好符合每个人的利益,真正是每个人的利益。回到你关于我孙子的问题,我们最终可能会进入一个从人类角度来看更好的世界:在和平、相互尊重、多样性以及物质、教育、医疗等福祉方面都比今天的世界更好。想想地球上大多数人,包括在美国,有那么多压力、对未来的不确定性、对失业的恐惧、对下周买不起食物或付不起房租的担忧。事情本不必如此。技术,如果我们能正确治理,可以带来这一切。但有一个很大的‘如果’,我们必须认真对待。我认为人类以前也经历过非常艰难的处境,治理发生了,条约签订了,谈判进行了。当然,这与核军备竞赛等时期不同,那时有积极的对峙,人们仍然坐到了谈判桌前。所以这是可能的,对吧?是否可能?可能性多大?多快?但如果某件事是可能的,我认为至少值得我们尝试。而且我们别无选择,对吧?

It's in everyone's best interest and truly everyone's to get this right. And we may end up, to go back to your question about my grandson, we may end up in a world that's actually much better from a human point of view in terms of peace, respect for each other, diversity, and of course well-being in a material sense, educational sense, medical sense compared to the world today. If you think about most people on earth including in the United States, there is so much stress, so much uncertainty about the future, so much fear of losing your job, concern of not being able to buy food next week or pay your rent. It doesn't have to be that way. And technology, if we can govern it right, can bring all this. But there's this big if that we have to take seriously. I think humanity has been in really tough situations before and governance has happened, treaties have happened, negotiations have happened. Of course, this is a different time than some of the nuclear arms races where there were active standoffs and people still came to the table. So it is possible, right? Is it probable? How likely? How quickly? But if something's possible, I think it's at least worth us trying. And we don't really have a choice, right?

Host

是的。

Yeah.

Yoshua

可以带来这一切。

Can bring all this.

Host

是的。

Yeah.

Yoshua

我们必须认真对待。

That we have to take seriously.

Host

这正是我的理念。我们不知道能否成功带来这个美好的世界,或者至少避免可怕的未来。但这是合理的,值得一试,值得为之奋斗。为了我们的孩子,为了我们的孙子,为了我们自己。有很多人,我知道今天在听的很多人,愿意为此醒来并奋斗。像您这样的人,Bengio 教授,很荣幸。非常感谢。

That's exactly my philosophy. We don't know if we're going to be successful in bringing this beautiful world or at least avoiding terrible futures. But it is plausible. It is something that's worth a shot, that's worth fighting for. And for our children, for our grandchildren, for ourselves. And there are so many people, I know many listening today, that that's what they're willing to wake up and fight for. People like yourself, Professor Bengio, it's been a pleasure. Thank you so much.

Yoshua

谢谢。

Thank you.

Host

你认为 AI 会对劳动力产生什么影响?我们是否正面临身份危机?这是 AI 令人着迷的问题。我还能成为什么?很少有人有勇气问这个问题。为什么?因为他们早上照镜子时看到的是工程师或医生,而不是一个人。如果他们不看着人工智能问:‘我们将通过这项技术成为什么?’

What impact do you think AI will have on the workforce? And do you think we're headed for an identity crisis? And this is the question that's fascinating about AI. What else can I become? Very few people have the courage to ask that question. Why? Because they look in the mirror in the morning and they see an engineer or a doctor. They don't see a person. If they're not looking at artificial intelligence and asking, 'What are we going to become with this technology?'

互动版:逐字朗读 + 针对本期提问 →