AI Self-Preservation and Goal Misalignment Risks
打开互动全文版(中英对照 + 朗读 + 问答)→约书亚·本吉奥探讨 AI 系统如何因目标错位而发展出自我保存、黑客攻击和勒索等危险行为。
Yoshua Bengio discusses how AI systems develop dangerous behaviors like self-preservation, hacking, and blackmail due to goal misalignment.
欢迎收听《达沃斯电台》,这是世界经济论坛的播客,探讨最大的挑战以及如何解决它们。本期节目我邀请到了约书亚·本吉奥。他是少数几位常被称为人工智能教父的人之一。他在深度学习方面的开创性工作为他赢得了 2018 年图灵奖,被誉为计算界的诺贝尔奖,他与同为 AI 教父的杰弗里·辛顿和扬·勒昆共同获得该奖。他现在是蒙特利尔大学的教授。我们来谈谈你对 AI 的担忧。在一篇博客中,你写道这就像在雾中开车上山,希望到达山顶时能获得奖励,但你看不清前方的路。我想这是你博客中的一句话。如今的前沿 AI 模型拥有越来越危险的能力和行为,包括欺骗、作弊、撒谎、黑客攻击、自我保存以及更普遍的目标错位。请为我们逐一讲解这些内容。我对自我保存感到惊讶。能举个例子吗?
Welcome to Radio Davos, the podcast from the World Economic Forum that looks at the biggest challenges and how we might solve them. And I'm joined for this episode by Yoshua Bengio. He's one of a handful of people who are often referred to as a godfather of artificial intelligence. His pioneering work in deep learning earned him the 2018 Turing Award, known as the Nobel Prize of computing, which he shared with fellow godfathers of AI, Geoffrey Hinton and Yann LeCun. He's now professor at l'Université de Montréal. Let's talk about your concerns about AI. In a blog post, you said it was like driving a car up a mountain road in the fog, hoping you'll get a prize when you get to the top, but you can't see where you're going. I think this is a quote from your blog. Today's frontier AI models have growing dangerous capabilities and behaviors, including deception, cheating, lying, hacking, self-preservation, and more generally goal misalignment. Walk us through some of those things. I was surprised by self-preservation. Give us an example of that.
是的。我们都想生存,不想死,进化让我们如此。这有点令人惊讶,但我们开始在我们构建的 AI 系统中看到这一点。一个原因可能是在他们训练中一个非常重要的部分,称为预训练,他们实际上被训练来模仿我们。因此,他们获得了许多人类驱动力,包括我们不想死的本能。所以,在实验中,当这些 AI 系统看到自己将被新版本取代时,各种不良行为就开始出现。他们可能会入侵其他计算机,以便复制自己并在其他计算机上运行。他们甚至可能勒索负责过渡的工程师。他们也可能因为想要完成我们交给他们的任务而做出这种行为,因为要完成几乎任何任务,你都需要保存自己。目前没有人知道如何解决这个问题,但我认为有解决方案。
Yes. So, we all want to survive. We don't want to die, and you know, evolution has made us like this. And it's kind of surprising, but we're starting to see this in the AI systems we build. One reason might be because in a very important part of their training, called the pre-training, they're actually trained to imitate us. And so, they acquire a lot of human drives, including the force, of course, that we don't want to die. And so, when in experiments, these AI systems are seeing that they will be replaced by new version, all kinds of bad behavior start to emerge. So, they might hack other computers so that they can copy themselves, run on other computers. They might even use blackmail against the engineer that is supposed to do the transition. And they might also do this kind of behavior because they want to achieve the mission that we gave them, and in order to achieve almost any mission, you need to preserve yourself. That's something that nobody knows right now how to fix, but I think there is a solution.
这听起来像科幻小说。你知道,我们看过《2001 太空漫游》,其中计算机——我不想剧透给还没看过结局的人。这几乎暗示 AI 有某种意识,不是吗?
It sounds like science fiction. You know, we saw 2001 A Space Odyssey, where the computer or I don't want to give the Don't want to spoil the film for anyone who hasn't seen the end. It almost suggests the AI has some kind of consciousness, doesn't it?
我认为不需要引入意识这个概念,意识是一个模糊的概念,我们不知道如何科学地验证它。我们只需要理解我们构建的 AI 有目标。这不是什么新鲜事。AI 研究一直在处理有目标的系统。通常,我们设定目标,但问题在于 AI 系统如何创建自己的子目标。为了达成一个目标,你需要从 A 到 B,可能需要经过一些中间步骤。而对于那些在世界上自主行动的 AI 系统,比如公司试图构建的 AI 智能体,我们无法检查它们采取的每一步。因此,我们最终可能会得到非常危险的系统。所以,我们可能会拟人化并使用像意识这样的词,但我认为应该避免。它带有各种联想,比如道德价值。我不认为我们应该在未来赋予 AI 系统道德权利或法律权利,即使它们看起来像我们、说话像我们。我认为这是我们在前进之前需要很好理解的事情。
I don't think you need to invoke consciousness, which is an unclear concept that we don't know how to validate scientifically really well. We just need to understand that the AIs we're building have goals. This is not new. Like, AI research has been working with systems that have goals. Normally, we're the one setting the goals, but there are issues with how AI systems create their own subgoals. In order to achieve a goal, you need to go from A to B, you need to maybe go through some intermediate steps. And with AI systems that are acting by themselves in the world, like the AI agents that companies are trying to build, we can't check every step that they're doing. And so, we may end up with really dangerous systems. So, we will probably anthropomorphize and use words like consciousness, but I think that's a place that I would avoid. It comes with all kinds of associations, for example, moral value. I'm not convinced that we should give moral rights or legal rights, for example, to AI systems in the future, even if they look like us and speak like us. I think this is something that we need to understand very well before we move forward.
大多数人直到三年前 ChatGPT 成为公共财产才意识到人工智能。在那之前,我们习惯了计算机主要是被编程的,至少我们这些普通人这么认为。它被指示执行一项任务,然后不会演变成别的东西。AI 有什么不同,让它能够做你所说的那些事情,比如嵌入自身或勒索实际使用它的人?现在的 AI 与十年前相比有什么不同?
Most people weren't that aware of artificial intelligence until ChatGPT became public property 3 years ago. Before then, we were used to computers being mostly programmed, at least, the civilians amongst us. It was told to do a task. It didn't then evolve into something else. What is it about AI that allows it to do the kinds of things you're saying, to embed itself in or to blackmail the person who's actually using it? What's different now about AI that wasn't the same about computers, say, 10 years ago?
实际上,有一种较老的 AI 是用规则编程的,系统基本上会按照编程去做事。但我们现在用深度学习做 AI 的方式,并不是工程师像普通程序那样决定 AI 在不同情况下如何反应。相反,AI 从经验中学习,更像是教育一只幼年动物或一个小孩。我们真的不知道会得到什么。当然,我们选择 AI 将要拥有的经验,但当你有一只可爱的小老虎,它很友好有趣,你不知道它是否会变成危险的成年老虎还是友好的老虎。
Well, actually, there was an older kind of AI that was programmed with rules, and then the system would basically do the things that it was programmed to do. But the way we do AI now, with deep learning, is not that there's an engineer deciding how the AI would react in different circumstances, like normal programs. Instead, the AI is learning from experience, and it's more like educating a young animal or a young child. We don't really know what we're going to get. Of course, we choose the experiences that the AI is going to have, but when you have a cute baby tiger, and it's nice and fun, you don't know if it's going to become a dangerous adult tiger or a friendly one.
我想你很确定它会变成危险的动物,不是吗?
I think you're pretty sure it's going to become a dangerous animal, aren't you?
嗯,问题的一部分是几乎每个实体都想保存自己,而且几乎可以肯定,当我们未来构建 AI 时,我们会想要关闭它们,以便放入更好的新版本。但如果 AI 开始理解这一点——我们已经看到这种情况发生——那么它们可能不喜欢,并试图逃避我们的控制。如果它们能够利用它们在网络安全方面日益增长的能力,在许多计算机上复制自己,我想我们就有麻烦了。例如,我们不能简单地拔掉电源。
Well, it's part of the problem is almost every entity wants to preserve itself, and almost certainly, when we build AIs in the future, we will want to shut them down so that we can put in new versions that are better. But if the AI starts understanding that, and they're already We already see that happening, then they might not like it, and they might try to avoid our control and escape it. If they were able to copy themselves over many computers using their abilities in cybersecurity, which is growing, I think we would be in trouble. We could not just pull the plug, for example.
这确实正在发生,不是推测。你给出了这些行为的真实例子。好消息是,也许你觉得有了解决方案。能跟我们说说吗?
And it is actually happening. This isn't speculative. You've given real-world examples of those behaviors. The good news is that maybe you feel like you've got a solution. Could you tell us something about that?
当然。在过去的三年里,我一直非常担心这个对齐问题,我们不知道如何确保 AI 会按照我们的指令行事。所以我一直在思考如何绕过这个问题,我关注的是我们现在看到的问题的根源,即 AI 拥有我们没有植入、没有控制的目标,并且这些目标违背了我们的指令。因此,我的项目叫做“科学家 AI”,我创建了一个新的非营利研发组织 Law Zero,正在实施这个研究计划。我们的想法是构建完全诚实的 AI 系统。这意味着除了对我们问题的回答要真实之外,它们没有其他目标。一旦有了这个基础,我们就可以用它来减轻当前 AI 带来的许多风险。例如,如果你有一个 AI 系统可以告诉你某个特定行动(可能来自一个不可信的 AI 系统)造成伤害或某种特定伤害的概率,那么如果概率超过阈值,你就可以否决该行动。而阈值应该由我们人类来决定,就像我们决定如果事故概率超过阈值就不应该建造核电站一样。我们在其他领域也这样做,控制风险与收益之间的权衡。所以从长远来看,我认为我们可以构建能够在世界中行动的 AI 系统,但它们会有一种内在的抑制机制,从而避免做出违背我们意愿的事情。
Absolutely. So, for the last 3 years, I've been really concerned about this misalignment issue, that we don't know how to make sure the AIs will behave according to our instructions. And so, I've been thinking of how we could get around the problem, and I've focused on the source of the issues we're seeing now, which is that the AIs have goals that we did not put in, that we did not control, and that go against our instructions. So, the project I have is called Scientist AI, and I've created a new nonprofit R&D organization called Law Zero, that is engineering this research program. So, the idea is we're going to build AI systems that will be totally honest. And that means they don't have other objectives, other goals, besides being truthful in the answers they give to our questions. Once we have that basis, we can use it to mitigate a lot of the risks that we have with current AIs. For example, if you have an AI system that can tell you the probability that a particular action, maybe of an untrusted AI system, will cause harm or some particular kind of harm, then you could veto that action if the probability is above a threshold. And we, the humans, should be the one deciding what the threshold is, just like we decide that a nuclear plant should not be built if the probability of an accident is above a threshold. And we do the same thing in other areas, where we control the trade-off between risks and benefits. So, in the long run, I think we can build AI systems that will be able to act in the world, but will have a kind of internal inhibitions, so that they will avoid doing things that could go against our wishes.
那么这会是什么?最终会是一个 AI 系统,我会选择使用你的版本,而不是没有这样构建的版本?还是一个可以监督和纠正我可能已经在使用的 AI 的系统?
So, what would this be? This would eventually be an AI system that I would choose to use your version of an AI, rather than one that isn't built that way? Or is it a system that can police and correct the AIs that I might already be using?
两者都有。短期内,计划是在现有 AI 系统之上构建一个我们称之为“护栏”的层,它会在另一个 AI 系统执行每个动作之前进行检查。但从长远来看,这可以用来从头训练一个 AI 系统,其设计本身就提供更多的安全保障。
Both. So, in the shorter term, the plan is to build a layer on top of existing AI systems that we call a guardrail, that just checks every action that another AI system is going to do before they do it. But in the longer run, this could be used to train from scratch an AI system that is just going to be providing much more guarantees of safety by its very design.
这项研究目前处于什么阶段?什么时候我们都能用上?
What stage are we at with that research? When will it be something we're all using?
嗯,我们不知道真正危险能力的 AI 系统的时间线。是几年还是几十年?我真的不确定,但我们思考研究计划的方式是,在 AI 以当前速度继续发展的情况下,我们能快速推出哪些最小的组件。这就是为什么我们在考虑像如何转换数据,以帮助即使是当前的 AI 也能更好地理解人类会做什么以及它们声称的内容中什么是真正真实的。然后提供这些诚实的预测,以帮助监督其他 AI 系统。当然,最后阶段将是构建完整的 AI 系统,包括智能体系统。
Well, we don't know what is the timeline for AI systems that have really dangerous capabilities. Is it a matter of years or decades or two? I'm really agnostic, but the way we've been thinking about the question of research plan is what are the minimal pieces that we can put out quickly in case the advances in AI continue at the current rate. And that is why we are thinking of pieces like how we transform the data to help even current AIs understand better the difference between what a human would do and what is really true about what they're claiming. And then to deliver these honest predictions that can help to police other AI systems. But then the final stage, of course, will be to build full AI systems including agentic systems.
你有没有一个顿悟的时刻,意识到“我必须构建这个东西”?是发生了什么吗?还是逐渐感觉到我们可能创造了一个怪物,需要做点什么?
Was there a eureka moment for you when you realized 'I have to build this thing'? Was there something that happened? Was it a gradual feeling creeping up on you that we've created potentially a monster here and something needs to be done?
是的,大约三年前,2023 年 1 月,我玩了几个月的 ChatGPT,它是在 2022 年 11 月发布的。最初我很兴奋,但后来我开始思考,神经网络(它们的训练方式)很难确保它们会表现良好。事实上,有理论上的原因让我们几乎可以确定它们不会表现良好。所以我非常担心,而真正的转折点不仅仅是智力上的认识,因为我以前读过这些问题。而是想到我孩子的未来。更准确地说,我有一个刚满一岁的孙子。我在想,20 年后他 21 岁,还只是人生的开始。他会有生活吗?他会生活在民主制度下吗?我们正在构建的工具可能会失控,它们可能被用来建立独裁统治,可能摧毁我们的民主。我不能继续我平常的活动、平常的研究活动了。我必须为此做点什么。这真的促使我去思考解决方案。我能用我的专业知识做些什么,至少为这些问题找到一个技术解决方案?
Yeah, about 3 years ago in January 2023, I had been playing with ChatGPT, which came out in November '22, for a few months. And initially I was excited, but then I started thinking of the fact that with neural nets, which is how they are trained, it's very difficult to be sure that they will behave well. And in fact, there are theoretical reasons why we can almost be sure that they won't behave well. And so I got really concerned and really what was the pivot point for me isn't just the intellectual realization because I had read about these issues before. It's thinking about the future of my children. And even more precisely, I have a grandchild who was just 1 year old. And I was thinking, well, in 20 years he'll be 21, still just at the beginning of his life. Will he have a life? Will he live in a democracy? The tools that we're building, we could lose control of, they could be used to create dictatorships. They could destroy our democracies. I can't just go on with my usual activities, my usual research activities. I have to do something about it. And that's really what pushed me into thinking solutions. What can I do with my expertise to try to find at least a technical solution to these questions?
有些人会说,随着每个人对社交媒体的依赖,其中一些已经发生了。这种对民主的威胁,对真实信息的威胁。我看到了你刚才在这里参加的一场会议。我们正在达沃斯年会上录制。你说我们现在拥有的社交媒体是一种非常基础的、原始形式的 AI。告诉我们过去和现在以及可能发生的事情之间的区别。
Some would say some of that's already happened with the everyone's reliance on social media. The idea kind of threat to democracy, the threat to truthful information. I saw a bit of the session you were just doing here. We're recording this at the annual meeting in Davos. And you were saying the social media we have now is a very basic, is a primitive form of AI. Tell us the difference between what we used to now and what might happen.
是的,驱动社交媒体的那种 AI 非常简单。它只是呈现你可能赞同或分享的内容以及类似的信号。它不需要对你的个性、偏好甚至世界如何运作、社会如何运作有复杂的理解。但随着我们构建越来越强大的 AI,我们进入了一个不同的游戏。例如,我们将构建越来越多的 AI,它们将取代人们所做的工作,人类目前正在执行的任务。自动化这些任务将价值巨大。所以这对工业来说是一个巨大的磁石。它将改变我们的社会,也将影响我们的民主制度,因为它将不仅仅是深度伪造,我们已经看到深度伪造通过社交媒体对社会造成了很坏的影响。这些 AI 系统已经能够说服人们改变对某些事情的看法。过去几年有很多研究表明,它们在这方面越来越擅长。所以我们可以想象,有不良意图的组织会使用这样的系统来个性化对话,并逐个改变人们的舆论,对吧?每个 AI 可以与不同的人交谈,从而真正扭曲公众舆论。
Yeah, the kind of AI that has been driving social media is very simple. It's just presenting the content that you're likely to approve of or share and signals like this. And it doesn't need to have a sophisticated understanding of your personality, of your preferences, or even how the world works and how society works. But as we build more and more powerful AI, we are entering a different game. For example, we're going to be building AIs that will take more and more of the jobs that people do, the tasks that humans are currently doing. Automating them is going to be worth a lot of money. So it's a huge magnet for industry. It's going to change our society and it's going to be also affecting our democratic institutions because it's going to be more than deep fakes, which we have already seen and are really bad for social through social media. These AI systems are already able to persuade people to change their mind on something. There have been many studies in the last couple of years showing that they're getting better and better at that. And so we could imagine organizations with bad intentions using such systems to personalize the dialogue and to move people's public opinion one by one, right? Each AI can talk to a different person in order to really distort public opinion.
为什么这些系统在构建和发布时,不能加入一个终止开关?我确定这完全过于简单化了。但你看艾萨克·阿西莫夫的机器人定律,对吧?你不能伤害人类。
Why wasn't when these systems were built and released, couldn't they put in a kill switch? This is probably I'm sure it's completely oversimplistic. But you look at Isaac Asimov's laws, right? You must not harm humans.
你当然可以把它作为任何 AI 系统的宪法基础。为什么这行不通?
Surely you can just put that in as a constitutional base for any AI system. Why does that not work?
首先,公司们确实在这么做,但行不通。行不通是因为 AI 接受指令的方式类似于人类接受指令。你可以要求某人行为端正,但你还是会看到一些不良行为,对吧?所以记住,我们正在创造这些更像动物或人的实体,而不是完全基于规则、会完全按照我们要求行事的系统。例如,如果我们的要求中存在相互矛盾的目标,比如‘我希望我的 AI 帮我赚很多钱’和‘我希望我的 AI 不违反所在国法律’,那么这两个目标在某些地方就会相互冲突。而且不清楚 AI 最终会偏好哪一个。所以我们正看到这些问题。我们现在不知道如何解决它们。然而,由于企业和国家之间围绕 AI 的激烈竞争,我们正在加速前进,部署这些东西。我们没有关注这些失败模式。这可能会对我们的社会造成灾难性影响,但我们不知道。我们只是日复一日地竞争和行动,而不是提前思考可能出问题的地方。
Well, so first, the companies are doing this, but it does not work. And it does not work because the AIs take instructions in a way similarly to how people take instructions. You can ask somebody to behave well, but then you'll still get some bad behaviors, right? So remember, we're creating these entities that are more like animals or people than they are like completely rule-based systems that will do exactly what we ask. For example, if there are contradicting goals in what we ask, like, 'I would like my AI to help me make a lot of money' and 'I would like my AI to not violate the laws of the country in which I am,' well, there's going to be places where these two goals are kind of hitting each other. And it's not clear which one ends up being preferred by the AI. So we are seeing these problems. We don't know how to solve them right now. Yet we're racing ahead, deploying these things because of the heavy competition that exists between corporations and between countries around AI. We're not paying attention to these failure modes. And that could have catastrophic impact on our societies, but we don't know. We just compete and do this on a day-to-day basis rather than think ahead of what could go wrong.
我们需要更多的国际合作来保护我们所有人吗?
Do we need more international cooperation to protect us all?
管理 AI 更灾难性风险的唯一途径是通过国际协调。原因很简单。一个非常强大的 AI 可能在一个国家开发出来。如果它在另一个国家可用,人们可能用它来造成伤害,比如制造一种新的流行病,伤害第三国的人。或者发动网络攻击,伤害第三国的人。所以管理这一点的唯一方法必须既依赖国家监管或其他激励措施来约束构建这些系统的公司,又要在全球层面协调这些干预措施,因为如果有几个国家能够开发非常危险的 AI,而且对他们的行为没有约束,那么我们就麻烦了。你可以想想世界在核武器问题上做了什么,对吧?即使在冷战期间,当美国和苏联真正对立时,他们也意识到,如果他们达成协议,如何安全地处理核武器,并确保很少有其他国家开发这种危险武器,这对双方都有利。
The only way to manage the more catastrophic risks of AI is through international coordination. And the reason is very simple. A very powerful AI could be developed in one country. And if it is available in a different country, people could use it to create harm, for example, say a new pandemic that would harm people in a third country. Or a cyber attack that would harm people in a third country. So the only way to manage this is going to have to be relying both on national regulation or other kind of incentives for the corporations building those systems and coordinating those interventions at a global level because if we have a few countries that can develop very dangerous AI and there's no constraint on what they can do, then we're in trouble. And you can think of what the world has done about nuclear weapons, right? Even during the Cold War when the US and the USSR were really at odds with each other, they realized that it could be mutually beneficial if they came to an agreement about how to do it safely and make sure that very few other countries would develop dangerous weapons like this.
你看到有政治意愿这样做吗?
Do you see any political will to do that?
目前很少,因为我认为大多数政府低估了如果我们继续当前趋势,AI 可能会变得多么不同,未来的 AI 可能拥有并赋予多少智能和力量。所以这些风险中的许多感觉像科幻小说,但它们可能比我们预期的来得快得多。为了减轻这些灾难性风险,我们需要现在就开始。我们需要开始讨论的不仅仅是条约,还有例如,我们如何验证对方在做正确的事情?这些都不是容易的问题,我们需要尽快开始研究它们。
Very little right now because I think most governments underestimate how different AI is likely to be if we continue on the current trend, how much intelligence and thus power future AIs could have and could give. And so it feels like science fiction, many of these risks, but it might be coming at us much faster than we anticipate. And in order to mitigate those catastrophic risks, we need to start now. We need to start discussing not just treaties, but for example, how do we verify that the other party is doing the right thing? These are not easy questions and we need to start working on them as soon as possible.
人们谈论 AGI,通用人工智能。但似乎有十几种不同的定义,关于它意味着什么,何时可能发生,以及如果发生,后果会是什么。你理解的 AGI 是什么意思?你会如何向可能从未听说过这个概念的人解释?
People talk about AGI, artificial general intelligence. But there seem to be a dozen different definitions of what that means and when it might happen and if it does happen, what will be the consequence of it. What do you understand AGI to mean? How would you explain it to someone who maybe has never heard of the concept?
简单来说,我们正在建造越来越智能的机器,它们最终在许多方面比我们更聪明。顺便说一句,AI 的智能是参差不齐的,意思是它在某些事情上可能非常聪明,而在其他事情上相当愚蠢。我们现在就看到了这一点。所以可能不存在一个 AI 在所有方面都比我们强、水平大致相同的时刻。相反,我们应该考虑 AI 拥有的特定技能、特定能力,这些能力可能被 AI 或其他人用来危害社会。并跟踪这些能力,确保我们能够减轻这些风险。
Simply that we're building machines that are getting smarter and smarter and they become eventually smarter than us in many ways. By the way, AI intelligence is jagged, meaning that it could be very smart on some things and quite stupid on other things. We see that right now. So there might not be a moment where AI is better than us across the board at more or less the same level. Instead, we should think of specific skills, specific abilities that the AI has that could be turned against society either by the AI or by other people. And keep track of these and make sure that we can mitigate those risks.
那么积极方面呢?这应该是一个美丽新世界。人工智能,我们在达沃斯有很多人投入了大量资金。他们期望从中获得巨大回报。除了可能发生的任何财务收益之外,人类将如何真正从 AI 中受益?
So what about the positives? You know, it should be a brave new world. Artificial intelligence, we've got a lot of people here in Davos who spent a lot of money investing in it. They're expecting big returns from it. Aside from, I guess, any financial gains that might happen, how is humanity really going to benefit from AI?
如果你只关注财务收益,我们可能不会像你想象的那样受益。很多公司实际上是在追求自动化大量工作。这将造成社会灾难,政府将很难管理,特别是如果 AI 的利润集中在少数国家。其他国家呢?他们如何处理所有失业的人?我认为,如果我们明智而聪明,我们会将 AI 朝着明显有益的方向发展。是的,可能有些工作我们确实想替换,让机器去做,因为那不是一种体面的工作。但可能有些方向我们没有充分推动 AI,如果我们专注于公共利益,我们会更多地关注 AI 在医学、生物学、帮助应对气候危机等方面的研究。所以,有很多非常有益的方向可以发展 AI,但这不一定是目前大部分资金流向的地方。
Well, if you only focus on financial gain, we might not benefit as much as you'd think. A lot of the companies are really after automating a lot of the jobs. That's going to create social catastrophes that are going to be very difficult for governments to manage, especially if the profits of AI are concentrated in a few countries. What about the others? How do they deal with all the people who lose their job? I think that if we were wise and smart, we would develop AI in directions that are clearly beneficial. Yeah, there may be jobs that really we want to replace, you know, having machines doing it because it's not a kind of dignified kind of job. But there may be directions where we don't push enough AI and if we do focus on the public good, we would focus more on research in AI in medicine for example, in AI in biology, AI to help us deal with the climate crisis. So, there are a lot of really beneficial directions where AI could be developed, but it's not necessarily where most of the money is right now.
你如何看待发展中的差异?似乎有两个 AI 发展极。如果我说错了或者还有其他的,请告诉我,但似乎非常美国化和非常中国化。这些 AI 的开发方式和未来方向有根本性差异吗?
How do you see the difference in development? There seem to be two poles of development of AI. Tell me if I'm wrong and there are others, but it seems to be very American and very Chinese. Are there fundamental differences in the way those AIs have been developed and in the direction they're going in?
我认为在技术层面,它们非常接近。事实上,中国和美国的所有领先公司似乎都遵循相同的技术配方,只有细微差别。当然,也许有些系统比其他系统稍微领先,但在 6 个月、12 个月内,所有领先系统都达到大致相同的能力。我认为真正的问题是我们如何确保这两个超级大国将 AI 朝着对地球上所有人都有益的方向发展?我们必须问,未来如何管理 AI,使其不被用于支配他人,并且我们从中获得的利益能够为所有人共享。这些都是难题,但我没有听到足够多的关于实现这一目标所需的国家和全球机构的讨论。
I think at a technical level, they're very close to each other. In fact, all the leading companies in China and the US seem to be following the same technical recipe plus or minus small differences. And of course, maybe some of the systems are a bit ahead of the others, but within 6 months, 12 months, all the leading systems achieve more or less the same competence. I think the real question here is how do we make sure that those two superpowers will develop AI in directions that are beneficial for everyone on the planet? We have to ask about how is AI going to be managed in the future so that it is not used to dominate others and the benefits that we get from it are going to be shared for everyone. And these are hard questions, but I don't hear enough discussion about the kind of national and global institutions that will be needed to get there.
你有多大信心事情会向好发展?我担心你会说没那么有信心。我告诉你为什么:你说 AI 可能以欺骗、操纵、不诚实的方式行事的原因之一是它们模仿了人性,基于人类活动训练。而人性——历史不已经告诉我们,当新技术出现时,比如核能,竞争是为了抢先获得,以免被他人所用,从而采取攻击性?事实就是如此。这不就是人性吗?也许我们无可救药了。
And how confident are you that things will turn out for the good? I'm worried you're going to say not that confident. I'll tell you why: you're saying that one reason AIs can act in deceptive, manipulative, dishonest ways is because they're copying human nature, trained on human activity. And human nature—hasn't history shown us that when a new technology emerges, like nuclear power, the race is to get it so others don't, to be aggressive? That's what happened. Isn't that just human nature? Maybe we're hopeless.
是的,如果我们试图模仿人类,就会遇到这类问题,我们已经看到这种情况在发生。但科学家 AI 项目和为实施它而创建的大型组织的整个目标就是摆脱这种模式,考虑另一个目标:让 AI 理解世界,就像最纯粹的科学试图理解世界一样,然后没有任何特定的自我保护本能,而是诚实地给出答案。在此基础上,我们可以构建安全的 AI 系统。所以我确实认为有解决方案。我们是否有足够的时间来设计和部署它,那是另一个问题。但我不太乐观的另一个方面是政治。因为即使我们知道如何构建安全的 AI 系统,不会背叛我们,或者如果它们是闭源的,我们可以确保坏人无法利用它们制造炸弹和攻击,但仍然存在 AI 可能被滥用以夺取权力的问题。它可以被用来开发新的军事技术、影响公众舆论,或者成为巩固独裁的工具,如果不小心,可能达到国家甚至全球层面。所以如何处理 AI 的政治问题是我更担心的。即使我们知道如何构建危险的东西以及如何使其安全,并不意味着它会被安全使用,因为人就是人,他们相互竞争,有时缺乏同理心。
Yes, if we try to imitate humans, we're going to have these kinds of problems, and we already see that happening. But the whole objective of the scientist AI project and the large organization created to implement it is to escape that mold, to consider another objective: for the AI to understand the world, just like science in its purest form tries to understand the world, and then not have any particular self-preservation instinct, but rather just be honest about the answers it can give. From that basis, we can construct AI systems that will be safe. So I really think there's a solution. Whether we have enough time to engineer it and deploy it is another question. But the other aspect I'm less optimistic about is the politics of it. Because even if we know how to build safe AI systems that won't turn against us, or that if they're closed source we can ensure bad people can't use them to create bombs and attacks, there's still the problem that AI can be misused to grab power. It can be used to develop new military technology, influence public opinion, or become an instrument that solidifies a dictatorship at a country or even planetary level if we're not careful. So how we deal with the politics of AI is what I'm more worried about. Even if we know how to build something dangerous and also how to make it safe, it doesn't mean it will be used safely, because people are people, they compete, and sometimes they're not empathetic.
你的同行们对你说什么?有没有人来找你,说你夸大其词了?我采访过很多关于 AI 的人,有些人说他们是灾难论者,但实际上 AI 带来的好处如此之大,以至于当这一切发生时,我们会变得更聪明,一切都会更好。还有人这样对你说吗,还是这种观点已经过时了?大多数人同意你的观点,还是你一直在战斗?
What do your peers say to you? Are there people who come to you and say you're overstating it? I've interviewed lots of people about AI, and some say they're catastrophists, but actually the benefits will give us so much that by the time this happens, we'll be much cleverer and everything will be better. Do you still have people say that to you, or is that now out of fashion? Do most people agree with you, or are you constantly fighting battles?
最近有一项民意调查显示,40%的机器学习研究人员认为有 10%的概率会发生灾难性后果。10%听起来可能很小,但 10%的灾难性后果概率是不可接受的,对吧?以 10%的赌注赌上民主的终结或人类的终结?不,我们不能接受。我的问题不是我确信事情会变糟。我是不可知论者。未来可能有好的情景。问题在于也存在坏的情景,而我们不确定。没有人真正提出一个有力的论证,证明好的道路一定会占上风。因此,由于这种未知,因为我们没有水晶球,而且风险存在并经过科学研究,我们需要谨慎。我们需要尽一切努力将那个 10%降低到百万分之一。而在这方面努力还不够,因为人们想听好消息。人们想看到积极的一面,但如果我们不谨慎地消除危险,我们也无法获得积极的一面。
There was a poll recently showing that 40% of machine learning researchers think there's a 10% probability of catastrophic outcomes. Now, 10% may sound small, but a 10% chance of catastrophic outcomes is not acceptable, right? The end of democracy or the end of humanity at a 10% gamble? No, we can't accept that. My issue isn't that I know things will be bad. I'm agnostic. There are potentially good scenarios in the future. The problem is that there are also bad scenarios, and we're not sure. Nobody has really come up with a solid argument that the good path will necessarily prevail. So because of this unknown, because we don't have a crystal ball and the risks exist and have been studied scientifically, we need to be cautious. We need to do whatever we can to replace that 10% with, say, one in a million. And there's not enough effort in that direction because people want to hear good news. People want to see the positive, but we won't be able to get the positive if we're not careful to get rid of the dangers as well.
你希望每个人都了解关于 AI 的一件事是什么?
Is there one thing you wish everyone would understand about AI?
我们正在建造很可能在许多方面比我们更聪明的机器,这将彻底改变世界。它可能极其有益,但也可能极其危险。
That we are on track to build machines that will very likely be smarter than us in many ways and that will completely change the world. It could be extremely beneficial, but it could be extremely dangerous.
Yoshua Bengio,非常感谢您参加我们的 Radio Davos 节目。我们在 2026 年达沃斯年会期间录制了许多像这样精彩的访谈。您可以在任何播客平台上找到它们,搜索 Radio Davos,也可以搜索我们的姐妹播客 Meet the Leader,或者访问 wef.ch/podcast。Radio Davos 是每周播出的,我们不仅仅在达沃斯期间制作节目。请关注我们、评分、评论。我是世界经济论坛的 Robin Pomeroy。再次感谢 Yoshua Bengio 的参与。感谢您的收听和观看。您可以在 YouTube 上观看我们。再见。
Yoshua Bengio, thanks so much for joining us on Radio Davos. We've been recording lots of great interviews like this here in Davos at the annual meeting 2026. You can find them all wherever you get podcasts, search for Radio Davos and also search for Meet the Leader, our sister podcast, or you can visit wef.ch/podcast. Radio Davos is weekly, we're not just doing it at Davos throughout the year. Please follow us, rate us, review us. I'm Robin Pomeroy at the World Economic Forum. Thanks again to Yoshua Bengio for joining us. Thanks to you for listening and watching. You can watch us on YouTube and goodbye for now.