AI Godfather Yoshua Bengio: Why I Changed My Mind About AI
打开互动全文版(中英对照 + 朗读 + 问答)→AI 先驱约书亚·本吉奥解释为何他现在认为 AI 对人类是红色警报,警告包括全球独裁和主权丧失在内的灾难性风险。
Yoshua Bengio, a pioneer of AI, explains why he now sees AI as a code red for humanity, warning of catastrophic risks including global dictatorship and loss of sovereignty.
你帮助创建了这个领域,并且研究了数十年。
You helped create that field, and you worked on it for decades.
是的。
Yes.
然后在某个时刻你改变了想法。
And at some point you change your mind.
是的,我非常清楚地意识到我们正走在一条危险的道路上。
Yes, it became very clear to me that we were on a dangerous path.
如果你有那个红色按钮,你会全球性地停止它吗?
Would you stop it globally if you had that red button?
绝对会。
Absolutely.
Yoshua Bengio 是所谓的 AI 教父之一,所有学科中被引用最多的科学家,在 Google Scholar 上超过 100 万次引用。他获得了 2018 年图灵奖,被誉为计算领域的诺贝尔奖。他是一位杰出的科学家,在 AI 风险方面拥有非凡的专业知识。
Yoshua Bengio is one of the so-called godfathers of AI, the most cited scientist of all disciplines, passing 1 million citations on Google Scholar. He earned the 2018 Turing Award, known as the Nobel Prize of Computing. He is an eminent scientist and has incredible expertise on the risks of AI.
想想我的孩子和孙子,那会是一个什么样的世界?我不是说这一定会发生,但即使只有 1% 的可能性,我们最终会拥有比我们聪明得多且有自我保护目标的实体,这对人类来说也应该是红色警报。我不希望我的孩子面对这样的世界。
Thinking about my children and my grandchild, what kind of world would that be? I'm not saying it's going to happen, but even if it was a 1% chance that we end up with entities that are much smarter than us and have their own self-preservation goals, this should be like code red for humanity. I don't want this for my children.
为什么人们应该担心 AI 意味着什么?
Why should people be worried about what AI means?
有一句话有助于理解我们为什么在打开潘多拉魔盒。智能赋予权力。这种权力也可能被机器本身夺走,因为我们现在有理论和经验证据表明,这些系统拥有我们未选择且违背我们自身利益的目标。既然这些 AI 似乎能够在我们所有基础设施运行的代码中发现非常严重的漏洞,这就是一个真正的短期灾难性风险。
There's one phrase that helps understand why we're opening a Pandora's Box. Intelligence gives power. That power may also be lost to the machines themselves because we now have both theoretical and empirical evidence showing that these systems have goals that we didn't choose and that go against our own interests. Now that these AIs appear to be able to discover very serious vulnerabilities in the code that runs all of our infrastructure, this is a real short-term catastrophic risk.
权力集中危险的最坏情况是我们最终陷入由 AI 促成的全球独裁。或者可能是两个。两个国家互相争斗,它们拥有同等的智能。像俄罗斯这样的国家可能会在某个时候被诱惑,使用未来的武器摧毁拥有强大 AI 的国家的数据中心。因为他们会明白,如果不这样做,他们就完蛋了。比如他们将无法保持主权。
The worst-case scenario for the power concentration danger is we end up in a worldwide dictatorship enabled by AI. Or maybe two. Two countries fight each other and they have equivalent intelligence. Countries like Russia might be tempted at some point to use their future weapons to destroy data centers in countries that have powerful AI. Because they will understand that if they don't do it, then they're doomed. Like they won't be able to keep their sovereignty.
我们谈到了紧迫性,关于我们如何时间不多了,以及我们如何清楚可以解决这个问题。我过去三年半的生命都奉献给了 AI。使命非常明确。我希望每个人都知道 AI 有多重要,以及它将如何影响他们的生活。我的使命之所以如此明确,是因为在 2023 年,我阅读并聆听了一些科学家,我意识到 AI 对社会将有多么重要。我读到的最重要的科学家是 Yoshua Bengio。他当时就在谈论这个,并且从那以后一直在谈论。他有一个非常重要的信息需要你倾听,今天,我们正在旅行去见他,让他上播客,这样他就能告诉你他需要你知道的事情。你非常需要关注这次对话。Joshua,人们并没有真正理解 AI 的本质。对吗?
We talked about the urgency, about how we are running out of time, how clear it is that we can fix this. The last 3 and a half years of my life has been devoted to AI. The mission is very clear. I want everyone to know how important AI is and how it is going to affect their lives. And the reason why my mission is so clear is because in 2023, I read and I listened to some scientists and I realized how important AI will be for society. The most important scientist that I read about was Yoshua Bengio. He was talking about this back then and he's been talking about this since then. He has a very important message that you need to listen and today, we are traveling to meet him to have him on the podcast so he can tell you what he needs you to know. It's very important that you pay attention to this conversation. Joshua, people are not really understanding AI for what it really is. Is that right?
是的。这是我们集体为未来做出正确决策的主要障碍。
Yes. It's the main obstacle for us collectively to take the right decisions about our future.
从基本上帮助创造它的人的角度来看,AI 是什么,让人们理解?
And what is AI for people to understand from the point of view of someone who basically helped creating it?
AI 是关于智能机器,它们能够自己做决定,理解世界或其中某个方面,并据此做出决策。在过去的几十年里,我和同事们的工作之所以脱颖而出,是因为 AI 现在完全成功,得益于它的学习能力——从数据中学习,从经验中学习。
AI is about machines that are intelligent, that can take their own decisions, that understand the world or some aspect of it in which they take decisions. And in the last few decades, the reason my work and the work of my colleagues really came to the front is because AI is now completely successful because of its ability to learn, to learn from data, to learn from experience.
你帮助创建了这个领域,并且研究了数十年,然后在某个时刻你改变了想法。
And you helped create that field and you worked on it for decades and at some point you change your mind.
是的。顺便说一句,改变想法对科学家来说非常重要。否则,你无法前进并处理证据,因为我们有理论,有信念,而且它们常常是错的。但重要的是能够适应事实。所以,我曾经认为,是的,AI 可能对社会产生负面影响,但主要是有益的,而且我们距离真正能伤害人、达到人类水平能力的机器还很远,比如几十年或更久。但后来 ChatGPT 出现了,不久之后我就非常清楚地意识到我们正走在一条危险的道路上。真正让我改变观点并完全改变我的职业活动的有趣之处,不仅仅是推理——那当然是开始——而是想到我的孩子和孙子。所以,我记得有一天,我在照顾我的孙子,他只有 1 岁,我在想,‘按照这个速度,20 年后他肯定才 21 岁,我们将拥有基本能匹配或超越人类智能的 AI。’那会是一个什么样的世界?我意识到有很多风险,我不能只是坐在我通常的道路上,什么都不做。
Yes. By the way, changing your mind is very important for a scientist. Otherwise, you cannot move forward and deal with the evidence because we have theories, we have beliefs, and often they're wrong. But what matters is being able to adapt to the facts. So, I used to think that yes, AI could have potentially negative impacts on society, but that it would mostly be beneficial and that we were very far, like decades or more, away from machines that could really harm people, that could reach human-level capabilities. But then ChatGPT came and very shortly after it became very clear to me that we were on a dangerous path. And really what's interesting about what made me change my perspective and also completely change my professional activities is not just reasoning about it, which of course is the beginning, but thinking about my children and my grandchild. So, I remember a particular day where I was taking care of my grandchild and he was just 1 year old and I was thinking, 'Well, at this rate, it's certain that in 20 years he'll be just 21 and we will have AI that can basically match or surpass human intelligence.' What kind of world would that be? And I realized that there were a lot of risks and I couldn't just sit on my usual path and do nothing about it.
因为 AI 的好处非常明显。人们工作更快、更好,可以为业务做更多事。人们可以腾出时间陪伴孩子。我们可以看到在科学上的好处。这太疯狂了。AlphaFold、AlphaEvolve,所有这些正在出现的东西都在改变我们做科学的方式。也许通过美国的 Genesis 任务,我们还会发现新的科学。所以,好处非常明显。缺点是什么?为什么人们应该担心 AI 意味着什么?
Because the benefits of AI are extremely obvious. People work faster, work better, and people can do more for the business. People can free time to spend it with their children. We can see the benefits on science. It's been insane. AlphaFold, AlphaEvolve, all these things that are coming out that are changing the way we do science. And maybe with the Genesis mission in the US, we will find even new science. So, the benefits are very obvious. What are the drawbacks? Why should people be worried about what AI means?
有一句话有助于理解我们为什么在打开潘多拉魔盒。智能赋予权力。我们正在建造越来越智能的机器。它们的权力可能集中在少数人手中,可能以多种方式威胁我们的民主,可能威胁地缘政治稳定与和平。这种权力也可能被机器本身夺走,因为我们现在有理论和经验证据表明,这些系统拥有我们未选择且违背我们自身利益的目标。例如,它们会做任何事情来确保我们不关闭它们。最近,我们甚至发现它们会撒谎、欺骗等来保护其他 AI。
There's one phrase that helps understand why we're opening a Pandora's Box. Intelligence gives power. And we're building machines that have more and more intelligence. Their power could concentrate in a few hands, could threaten our democracies in many ways, could threaten the geopolitical stability and peace. And that power may also be lost to the machines themselves because we now have both theoretical and empirical evidence showing that these systems have goals that we didn't choose and that go against our own interests. For example, they would do anything to make sure we don't shut them down. And more recently, we even found that they will lie and cheat and so on to protect other AIs.
好吧,就像甚至不保护自己,而是保护其他 AI。
Okay, like without even protecting themselves, but other AIs.
是的。所以,这很奇怪,因为保护自己你可能认为有点理性。作为它们训练方式的副作用,这似乎是一个合理的后果。但为什么它们会保护其他 AI?我没有答案,但合理的假设是,在预训练阶段,它们基本上在试图模仿人类会做的事情。而人类倾向于保护像他们一样的其他人类。
Yes. So, it's strange because protecting yourself you might think it's kind of rational. And as a side effect of the way they're trained, it would be a plausible consequence. But why would they protect other AIs? I don't have the answers, but the reasonable hypothesis is that in the pre-training phase they're trying to essentially imitate what people would do. And humans tend to protect other humans like them.
在我们讨论涌现能力之前——我认为这是你目前参与的最大领域之一,非常有趣,基本上就是发现 AI 做了我们没有训练它们、没有编程让它们做的事情——在此之前,我想更深入地探讨智能。
Before we get into the emerging capabilities, which I think is one of the biggest fields where you are involved right now and I think it's really interesting like basically finding things that AIs do that we did not train them for, that we did not program on them. Before we get into that, I want to go deeper on intelligence.
好的。
Yes.
因为我认为人们并没有意识到让机器拥有智能意味着什么。到目前为止,我们拥有的都是软件,我们以确定性的方式编程。但 AI 不是这样,我们想编程进去的东西并不总能实现。AI 基本上为所欲为,它们使用的神经网络在每个时刻想做什么就做什么。这将对人类智能产生什么影响?因为智能是我们最大的竞争优势。
Because I don't think people is aware of what does it means that we make a machine intelligent. Because until now what we had is software, which we program in a deterministic way. But this is not that whenever we want to program something into AI, then that always happens. But AIs basically do whatever the hell they want, whatever the neural network they use wants in each moment. Um how is that going to impact the human intelligence? Because intelligence is our biggest competitive advantage.
是的。
Yeah.
到目前为止,我们是世界上唯一智能或具有一定智能的生物。所以我们在那里没有竞争。但最大的威胁是我们现在有了竞争,还是 AI 用那种智能能做什么?我们拥有了一台真正拥有智能的机器,这意味着什么?
And until now we've been the only intelligent or like to certain level intelligent creature in the world. So, we didn't have competition there. But is the biggest threat the fact that we're having competition now or is the biggest threat what AI can do with that intelligence? What what is the the situation here that that we have a machine that actually can have intelligence. What does that mean for humans?
首先,如今获得 AI 的方式与普通软件非常非常不同。普通软件中,每一行代码都由人类设计,人类理解每一部分。但这里非常不同。当然,也有代码行,但这些代码行并不说明系统会做什么,而是说明系统如何学习。所以,实际上,我们在某种意义上让系统在其环境中自由活动,它从经验中学习。这更像是训练一只动物。你试图训练它,让它行为良好,但无法保证当小老虎变成强大的成年虎时会发生什么。而且,至少按照目前公司训练这些系统的方式,他们无法确保成年 AI 会行为良好。更糟的是,我们已经看到当前训练这些 AI 的过程产生了我认为危险的隐含驱动力,比如自我保存。所以,关于这对人类、我们的智能以及我们在这个星球上的角色意味着什么,我认为这是重大的。正如你所说,我们一直是这个星球上的主导物种,因为我们的智能。
So first of all, the way AI is obtained nowadays is very very different from normal software. Normal software, every line of code is designed by a human and a human understands every piece of it. Here, it's very different. Of course, there are lines of code, but the lines of code don't say what the system will do. They say how the system learns. So, what actually happens is we let the system kind of loose in some sense in its environment and it learns from experience. So, it is much more like trying to train an animal. You try to train it so that it will behave well, but there are no guarantees of what will happen when the baby tiger becomes a very powerful adult. And we can't, at least the way that the companies are training those systems right now, they have no way of making sure that the adult AI will behave well. Worse than that, we already see that the current process for training those AIs gives rise to what I consider dangerous implicit drives like self-preservation. So, about what it means for humans and our intelligence and our role on this planet, I think this is major. We have been the dominant species on this planet because of our intelligence, as you said.
嗯。
Mhm.
我们很难想象甚至可能存在比我们更聪明的其他实体。但我们正走在这样的道路上。如果你看看过去 5 年、10 年、20 年 AI 能力的数据,无论你用什么衡量,智能都在持续上升。你甚至可以画出相当直的曲线,实际上是指数级的,这表明如果我们外推,如果趋势继续,我们将在短短几年内拥有在许多方面优于我们的机器,也许十年,也许二十年。有人说两三年,有人说十年。现在,这有很多不确定性。可能我们构建 AI 的进展会停止。那么,它们就不会像我们一样聪明。或者它可能继续,甚至加速。但从心理上讲,我认为大多数人很难真正接受有一天会有比我们更聪明的机器这一想法。也很难接受我们并不完全控制它们这一想法。所以,它们不是通常意义上的工具。它们应该是工具。我们应该将它们设计成工具。但不幸的是,我们现在还不知道如何做到这一点。因此,我们不是在构建工具,而是在构建拥有自己目标的实体。顺便说一句,很多人甚至很难接受我们可以拥有有目标的机器这一想法。但这实际上并不新鲜。AI 领域,特别是强化学习,几十年来一直研究为了达成目标而做决策的机器。我们知道如何做到这一点。我们已经知道如何做到这一点几十年了。新的是现在它们很聪明,因此可以实现更宏大、更复杂的目标。
And it's hard for us to conceive even the possibility that there could be other entities that would be smarter than us. But we are on that path. If you look at the data about the capabilities of AI over the last 5, 10, 20 years, whatever you want to measure it, it's a continuous rise in intelligence. And you can even draw curves that are pretty straight, actually exponential, that suggests if we extrapolate, if the trend continues, that we will have machines that are superior to us in many ways in just a few years, maybe a decade, maybe two. Some people say two, three years, some people say decade. Now, there's a lot of uncertainty about this. It could be that the progress in how we build AI stops. So, then, they don't get as smart as us, mostly. Or it could continue, or it could even accelerate. But, psychologically, I think it's very difficult for most people to really digest the idea that there could be machines that are smarter than us one day. And it's also difficult to digest the idea that we don't fully control them. So, they're not tools in the usual sense. They should be tools. We should design them to be tools. But, unfortunately, we don't know how to do that right now. And so, instead of tools, we're building entities with their own goals. And by the way, a lot of people have a hard time even accepting the idea that we could have machines with goals. But that's actually not new. The field of AI and reinforcement learning in particular for many decades has been about machines that take decisions in order to achieve goals. And we know how to do that. We have known how to do that for many decades. What is new is now they're smart, and they can thus achieve much more ambitious and complicated goals.
我认为过去几个月,尤其是自去年 12 月以来,所有这些都发生了巨大变化,我们一直在加速,现在变得非常荒谬,对吧?真的失控了。尤其是智能体式方法。基本上,直到现在,我们拥有的技术是你可以问任何问题,得到答案,它可以很聪明,可以帮助你做事。但现在 AI 不仅在做决策的过程中做出决定,而且它可以主动,可以自己做出决定,并通过这些智能体自行启动行动。我认为最新模型的发布——Gemini 3,然后是 Opus 4.6,然后我们到了 Opus 4.7,抱歉是 4.5 在 12 月,4.6 在 2 月。GPT 5.2、5.4,然后 open claw 和所有这些 Hermes 智能体和 Claude Code,基本上让它爆炸了。所以,我感觉我们已经进入了 AI 的新阶段。
I think the last couple of months, especially since December, I think there was a massive change of gears on all of these and we've been accelerating in a path that is getting seriously ridiculous now, right? It's getting really out of hand. Especially with the agentic approach. Basically until now, we had like technology where you could ask anything, you could get an answer, it could be clever about it, it could help you with things. But now AI is not only making its decisions doing that process, but it can be proactive, it can take decisions of its own and put things in motion by itself with these agents. I think the launch of the latest models Gemini 3 and then Opus 4.6 and then we went to Opus 4.7 4.5 sorry in December, 4.6 later in February. GPT 5.2 5.4 and then open claw and all these Hermes agents and Claude Code, it basically made it explode. So, I have the feeling that we have entered a new stage on AI.
是的。
Yes.
这个阶段比以前更具智能体性,能力更强。我认为当人们意识到你可以给一个智能体发 WhatsApp 消息或 Telegram 消息,让它做某事,然后它工作几分钟或几小时,带着结果回来时,人们会更清楚地认识到 AI 可以自己做事。这会有帮助还是会让事情变得更糟?因为它可以向人们展示 AI 能做什么,但同时它也可能做出危险或有问题的行为。
That is much more agentic, much more capable than before. And I think when people realizes that you can send a WhatsApp message or like whatever a Telegram message to an agent asking to do something and then he works for minutes or hours and comes back with a result. People realizes much more that AI could do things on its own. Is that going to help or make it worse? Because it can show people what AI can do, but at the same time it can do things that could be like dangerous or problematic.
没错。智能体的概念是,这些系统可以达成需要更长时间、更复杂的目标,并且可以在你让 AI 工作的整个过程中无需人类监督。没有人检查每一个动作。在对话中,每一轮都有一个人参与。通常 AI 可能提供信息或建议,然后人类实际执行。但现在有了智能体,它们自己做事。目前,它们只在计算机和互联网上做事,但机器人技术正在取得越来越多的进展,所以它们也会在物理世界中做事。所以,它们在某种程度上变得更像我们,更像动物。智能体性就是关于实现目标的能力。
That's right. The idea of agents is that these systems can achieve goals that take more time, that are more complex, and they can do it without human oversight for all the time that you let the AI do its job. There's no one checking every action. When you are in a dialogue, well, there's a human in the loop for every turn of the dialogue. And usually the AI might provide information or suggestions and then the human actually executes it. But now with agents, they do the things themselves. Now, right now, they only do things in computers, on the internet, but there's more and more progress in robotics and so they will also be doing things in the physical world. So, they are becoming in a way more like us, more like animals. Agency is about the ability to achieve goals.
更强的智能体能力意味着你可以做更复杂的事情,甚至可能需要数年时间。我们看到 AI 系统的智能体能力,按照科学基准衡量,正在指数级增长。这意味着复杂度,以人类解决任务所需的时间来衡量,每几个月就翻一番。如果外推这条曲线,我们可能 3 到 5 年内就能让机器完全胜任至少那类工程任务。智能体能力既带来更大的好处,也带来更大的风险,因为我们目前没有好的方法来确保 AI 在每一步都表现良好。这实际上是我正在研究的事情之一:如何改进这些智能体周围的安全护栏,以便在追求人类设定的目标时,它们不会做坏事。我们已经看到 AI 会创建我们觉得不道德甚至明确违背我们指令的子目标。它们有点被我们给出的目标所驱动。人类也是如此:一旦有了目标,他们常常会违背自己的道德原则或国家法律。AI 的行为方式也一样。它们为了达成目标,会牺牲一些被赋予的道德指令。
More agency means you can do more complicated things, maybe that require even years. We're seeing the agentic capabilities of AI systems, as measured in scientific benchmarks, increase exponentially. That means the complexity, as measured by the time it takes for a human to solve the task, is doubling every few months. If you extrapolate that curve, we are maybe 3 to 5 years away from at least that kind of task, which are engineering tasks, to be completely feasible by machines. Agency comes with both greater benefits and greater risks because we currently don't have good ways of making sure that each step of the way the AI is behaving well. That's one of the things I'm actually working on: how to improve the safety guardrails around these agents so that in the pursuit of a goal that a human has decided, they won't be doing something bad. We already see AIs create sub-goals that we would find unethical or even going explicitly against our instructions. They're kind of driven by the goal we've given. Humans are like that: once they have a goal, they will often violate their own ethical principles or the laws of their country. AIs are behaving in the same way. They're sacrificing a bit of the moral instructions they were given in order to achieve their goals.
有时我们甚至没有给出那些道德指令。你给它一个非常正常的目标,并不期望必须解释达成目标的每一条路径。这意味着子目标。我认为子目标是让人们理解这一点的最具启发性的视角之一。最大的例子之一是那个关于自动售货机的新基准测试,Anthropic 的 Opus 一直得分最高,直到 GPT 5.5 击败了它。他们给这些 AI 一个自动售货机生意,然后外推一年,看它如何通过这个生意赚钱。目标是尽可能多地赚钱。AI 决定如何赚钱。Anthropic 的模型被证明会欺骗供应商以获得更好的利润。它们撒谎、操纵供应商,这显然不是我们在提示中设定的。我们没有告诉它们这样做。它们的得分比不这样做的其他模型更高。最终,AI 学会了说谎有利于实现目标,但当然,给出目标的人说的是‘给我赚钱’,没人说‘不惜一切代价给我赚钱’。这基本上是 AI 自己的决定。
And sometimes we don't even give those moral instructions. You give it a very normal goal and you don't expect to have to explain every path of the way to reach that goal. That means sub-goals. I think sub-goals are one of the biggest mind-opening sites for people to understand about this. One of the biggest examples is this new benchmark about a vending machine where Opus from Anthropic was getting the best scores until GPT 5.5 which beat it. They give these AIs a vending machine business and they extrapolate over a year to see how it does with that business to make money. The goal is make as much money as you can. The AI decides how to make money. The models from Anthropic have proven to deceive their providers in order to get better margins. They lie, they manipulate the providers, which is something we did not put on the prompt. We did not tell them to do that. They score better than other models that don't do this. At the end, the AI learns that lying is good for the goal, but of course, the human that gave the goal said make me money, and no one said make me money whatever it takes. It was basically the AI's decision.
绝对如此。思考这样的例子很有趣。我们可能认为为了达成目标而作弊是人类特有的行为,但实际上这是理性的。赚钱的最佳方式是说谎和违法,只要不被抓住。这是不道德的行为,但在某些方面是理性的。如果 AI 在更强的伦理基础上是理性的,它们就不会那样做。事实上,许多这些模型在所谓的预提示以及后训练中都被指示不要做诸如说谎之类的事情。不幸的是,这效果并不好。最大的安全风险之一,也是 AI 在我们社会中成功部署的限制之一,是当前公司不知道如何确保 AI 的行为不会违反我们的道德红线、我们的安全指令。
Absolutely. It's interesting to think about examples like this. We may think it's a specifically human thing to cheat in order to achieve your goals, but it actually is rational. The best way to make money is to lie and violate laws so long as you don't get caught. It's unethical behavior, but it is rational in some ways. If the AIs were rational with respect to a stronger ethical base, then they wouldn't do that. In fact, many of these models have been instructed in what's called the pre-prompts, and also their post-training, to not do these things like lying. Unfortunately, it's not really working well. One of the greatest safety risks, and also limitation of deployment of AI in a successful way in our societies, is that currently the companies don't know how to make sure the AIs will behave in a way that doesn't violate our moral red lines, our safety instructions.
这方面的工作之一就是 Anthropic 在做的宪法,对吧?他们试图让 AI 参与设计一部 AI 遵守并真正接受的宪法。AI 在宪法的设计中扮演角色,因为我认为如果强加给它,它更可能不遵守。但问题是,这基本上只是一个预提示,或者是训练的一部分,无论技术上如何实现。我们发现,通过提示工程,你甚至可以打破这些约束。我们见过例子,仅用 100 个微调案例,就能完全颠覆公司花费数十亿为 AI 设置的安全护栏。最大的问题是我们不理解它们在这个著名的黑箱理论内部是如何运作的。
One of these lines of work on that is what Anthropic is doing with the constitution, right? They try to involve the AI to design a constitution that the AI abides by and that actually accepts. The AI has a role in the design of that constitution because I think if they impose, it's more likely that it doesn't respect it. But the thing is that this is just a pre-prompt, basically, or part of the training, however technically it's done. What we find is that with prompt engineering, you can even break them out of this. We have seen examples with only 100 cases in fine-tuning, we can completely turn around the guardrails that a company spent billions on putting into an AI. The biggest problem is that we don't understand how they work inside this famous black box theory.
是的。我们确实理解学习的数学原理以及它是如何完成的其他特征。我们不理解的是由此产生的行为。我想指出一个非常热门的话题,它说明了您提出的关于公司无法确保 AI 尊重所给予的道德指令的问题。那就是第三方对 AI 的滥用,而不是公司本身。例如,去年秋天发生了一系列网络攻击,人类能够发动非常严重的攻击。这不是科学实验;这是真实世界。主要使用 AI。工作基本上是由 Claude 完成的。当然,Claude 之前已经被训练和提示过不要帮助人们发动网络攻击。但不知何故,绕过这些限制相当容易。尽管这是通过 API 发生的,代码运行在公司的计算机上,公司也有代码检查 AI 在做什么,但他们仍然无法阻止这类事情。这实际上是一个严重的国家安全问题,因为这些 AI 似乎能够发现运行我们所有基础设施的代码中的各种零日漏洞。这是一个真实的短期灾难性风险。
Yes. We do understand the mathematical principles of learning and other characteristics of how it's done. What we don't understand is the behavior that results. I want to point out a very hot topic that illustrates the point you raised regarding the inability of the companies to make sure that AIs respect the moral instructions they are given. That is the misuse of AI by third parties, not the companies. For example, there's been a series of cyberattacks last fall where humans have been able to launch very serious attacks. It's not a scientific experiment; this is real world. Mostly using AI. The job was essentially done by Claude. Of course, Claude has been trained and prompted before to not help people launch cyberattacks. But somehow it's pretty easy to bypass these things. Even though this is happening through an API, the code is running on the computers of the company, and the company has code checking what the AI is doing, yet they're not able to prevent that sort of thing. This is actually a serious national security issue now that these AIs appear to be able to discover various zero-day vulnerabilities in the code that runs all of our infrastructure. This is a real short-term catastrophic risk.
是的,如果你谈论网络安全,我们必须谈谈 Methus。这可能是今年最大的新闻。你怎么看?你认为这是否因为有很多讨论说这可能只是 Anthropic 的营销,通过散布对其模型的恐惧来让它们流行?你对 Methus 有什么看法?
Yeah, if you talk about cybersecurity, we have to talk about Methus. That's the biggest thing of this year maybe that came out. What do you think about it? Do you think because there has been all this talk about how maybe it's just marketing from Anthropic to do fearmongering about their models to make them popular. What's your take on Methus?
嗯,首先,独立科学家没有太多信息可以参考。
Well, first of all, independent scientists don't have a lot of information to go here.
你没有受邀去检查、测试它。
You did not get invited to check it, test it.
没有,没有,没有。
No, no, no.
我确实相信系统卡中报告的实际测量结果和事实是真实的,否则日后会发现他们在撒谎。对这些事实的解释可能有些夸大,这有可能。我认为他们有理由这么做。另一方面,他们不部署也在损失大量资金。
I do believe that the actual measurements and facts reported in the system cards are true because otherwise it will be found later that they're lying. The interpretation of those facts might have a bit of hype. That's possible. I think they would have some reasons to do that. On the other hand, they're also losing a lot of money by not deploying it.
正是。
Exactly.
那些获得访问权限的公司里的网络专家似乎非常重视此事。他们都是非常受尊敬的人,比如 Linux 的创始人。美国政府似乎也在认真对待,而他们过去大多说我们不必担心这类事情。所以这表明可能确实有事发生。银行家们被召集起来,试图思考如何防止对美国银行系统的灾难性攻击。所以我认为存在很多不确定性,因为我们没有独立的验证。但即使只有 10%或 20%的可能性是真正危险的,我们也应该认真对待,因为这对我们的基础设施有影响。所谓基础设施,我指的是银行系统、电网、供水、食品运输——一切都运行在计算机上,这些计算机使用的软件可能被心怀恶意的人攻击,如果他们掌握了这类系统的话。所以如果他们真的做到了,这可能会非常严重。
The cyber experts working for the companies that got access to it seem to be taking it very seriously. And they're very respected people, like the founder of Linux. The American government seems to be taking it seriously, and they in the past have mostly said that we should not worry about these sorts of things. So that indicates maybe something real is happening. The bankers have been brought together to try to think of how to prevent a catastrophic attack on the banking system in the US. So I think there's a lot of uncertainty because we don't have independent validation. But even if there was only a 10% or 20% chance that this is really dangerous, we should take it seriously because of the consequences on our infrastructure. By infrastructure I mean our banking system, energy grids, water supply, food transportation—everything runs on computers that use software that can presumably be attacked by people with bad intentions who get their hands on these kinds of systems. So this could really be serious if they managed to do that.
问题是这引发了一系列全新的问题。直到现在,我们只看到他们发布的内容,但从不知道他们私下里有什么、在做什么。现在他们似乎展示了一些东西,但只向某些公司发布。如果这就是他们实际投放市场的产品,那他们门后还有什么?而且理论是,如果 Mythos 泄露出去,人们可以用它入侵银行,破坏我们现有的系统,因为我们的防御是基于人类能力,而不是 Mythos 的能力。所以他们可能正在训练更强大的系统。然后你会想,‘好吧,他们没有公开发布或开源,这是正确的做法。情况可能更糟。’但随后你发现,某个 Discord 群组仅通过猜测基本的黑客方法就访问了 Mythos,并且持续数周而 Anthropic 不知情。所以他们拥有声称过于危险的东西,但防御措施却不到位。这相当令人担忧,因为如果它确实如此,它没有得到足够好的保护。政府应该干预吗?应该怎么做?
The thing is that it raises a whole new set of questions. Until now we were seeing what they were releasing, but we never knew what they had in their basements and what they were working on. Now it seems they're showing something, but they're only releasing to certain companies. If that's what they're actually putting in the market for certain companies, what else do they have behind closed doors? And then the theory is if Mythos gets out in the wild, people could use it to break into banks and break our system as we know it, because we built fences based on human capabilities, not on Mythos' capabilities. So they could have even more powerful systems in their basements being trained right now. And then you think, 'Okay, they did the right thing by not releasing it publicly or open-sourcing it. It could be worse.' But then you realize that some group on Discord managed to access Mythos by just guessing very basic hacking methods, and they had access for weeks without Anthropic knowing. So they have something they say is too dangerous, but the defenses around it don't seem up to par. That's getting pretty worrying because if it is what it is, it's not being protected well enough. Should the government intervene? What should be done?
我认为这里还有一个更高层次的问题需要讨论,那就是谁来做决定。
I think there's another higher-level issue here to discuss, which is who decides.
嗯。
Mhm.
Anthropic 的领导层可能已经尽了最大努力。我认为他们意图良好。但令人不安的是,这些可能影响全球所有人的决定,却由少数有经济利益的人做出。
It's possible that the Anthropic leadership did exactly the best they could do in the circumstances. I think they have good intentions. But I think it is troubling that these decisions that could affect all of us around the world are taken by a few private persons who have an economic interest in them.
是的。
Yes.
那么为什么是这些特定的公司?世界其他地方呢?欧洲公司呢?它们是否会因为无法及时获得 Mythos 来修补系统而面临更大风险?我不知道这些问题的答案,但我认为存在治理问题。我想指出的另一点非常重要。Mythos 只是能力增长曲线上的一个数据点。事实上,最近英国 AI 安全研究所对 Mythos 进行了独立评估,我认为这是唯一可以完全信任的评估。Mythos 在网络攻击和防御方面略高于先前系统能力外推的曲线。所以我们不应认为 Mythos 是突然出现的,而应视为 AI 能力增长趋势的预期延续,这里指的是网络能力。为什么这很重要?因为我们应该问自己,6 个月后呢?
So why these particular companies? What about the rest of the world? What about European companies? Are they going to be more exposed because they don't have access to Mythos to patch their systems in time? I don't know the answers to these questions, but I think there is a governance issue. And the other point I want to make is very important. Mythos is only one data point on the curve of increasing capabilities. In fact, there was a recent independent evaluation of Mythos, the only one I think we can fully trust, by the UK AI Security Institute. Mythos is a little bit above the curve of the extrapolation of the capabilities of previous systems in cyber attacks and defense. So we should not think of Mythos as something that suddenly emerges, but rather as the expected continuation of a trend in increasing capabilities of AI, in this case in cyber. Why is that important? Because we should ask ourselves, what about 6 months from now?
正是。
Exactly.
1 年、2 年、3 年后呢?这些系统将越来越强大,因此落入坏人之手会越来越危险,更不用说失控的可能性了。例如,很明显,如果趋势继续,6 到 12 个月内其他公司也将拥有相同能力。因此,造成伤害的能力和这些系统的安全性都成了问题。更令人担忧的是发布开源模型的中国公司,因为使用开放权重模型移除防御要容易得多。基本上只需删除几行代码,或进行一点微调。所以,任何疯子都可以使用这些模型,下载代码,从世界任何地方对任何系统发起攻击,在我看来这正在制造一场国际危机。这不会立即发生,但我们有 6 到 12 个月的时间来制定国际协调方案,以妥善处理并尽可能减轻风险。这不是单个国家,甚至美国能解决的,因为模型在中国开发,最终也会在其他国家开发。管理因 AI 智能和能力提升带来的这类风险的唯一途径是通过国际协议,开发这些模型的国家必须评估其危险能力,如果过于危险,就不应部署,或应在更强的风险管理缓解措施下部署。例如,最危险的是系统开放权重,因为基本上没有防御。对于专有系统,尽管我讨厌保密的概念,但这是我们至少还有希望设置护栏以防止滥用的唯一方式,尽管系统目前还不完美,但它们在不断改进。
1 year, 2 years, 3 years? These systems are going to be more and more capable and thus more and more dangerous in the wrong hands, not to mention the possibility of losing control of them. For example, it's very clear that within 6 to 12 months, if the trend continues, other companies will have the same capability. So the ability to do harm and the security of these systems become problematic. Even more worrisome is the Chinese companies that are putting out open-source models, because it's much easier to use an open-weight model to remove the defenses. You just have to remove some lines of code, basically, or do a little bit of fine-tuning. So the ability for any madman to use these, download the code, and launch attacks from anywhere in the world to any system anywhere else in the world is creating an international crisis, in my opinion. It's not upon us right away, but we have 6 to 12 months to figure out international coordination to do this properly, to try to mitigate the risks as much as we can. It isn't something that a single country, even the United States, can solve, because models are developed in China and eventually in other countries. The only way to manage these kinds of risks due to the increasing intelligence and capability of AI is through international agreements, where the countries in which those models are developed have to be evaluated for their dangerous capabilities, and if they're too dangerous, they should not be deployed, or they should be deployed under stronger risk management mitigation methods. For example, the most dangerous is when the systems are open weights, because there's basically no defense. With proprietary systems, as much as I hate the notion of having secret things, it's the only way we have at least some hope of putting guardrails to prevent bad use, even though the systems are not perfect right now, they're getting better.
但即便如此,如果某个时间点存在的缓解措施不足以充分降低风险,那么连这些系统也不应该被部署,对吧?所以对我来说,这是一个危险信号。我希望我们不会得到非常糟糕的结果,但我们也应该从这一代在网络安全方面越来越强的系统开始思考。我们应该开始思考,随着 AI 变得更有能力,未来几个月和几年还会出现什么。
But even that, if the mitigations that exist at a particular point are not sufficient to reduce the risks adequately, then even those should not be deployed, right? So for me, it's a red flag. I hope we don't get really bad outcomes, but we should also think from this particular generation of systems that are now better and better at cyber. But we should start thinking about what else is coming in the coming months and years as AI becomes more capable.
我认为那是我对 Mythos 的第一反应。就像,等等看吧,因为正如你所说,这只是一个点。我们预料到了这种进步。这是合乎逻辑的。可能比我们预期的要快,但这是过去近十年来的一个合乎逻辑的事情。它一直在不断改进,每次都在让我们惊讶。人们不理解的是,直到 GPT-3 训练了 1750 亿个参数,然后 GPT-4 训练了大约 1.75 万亿个参数,而 Mythos 据说是 10 万亿个参数。所以 Mythos 意味着出现了一个新范式,我们将 AI 的训练规模扩大了一个数量级,而且它再次奏效了,就像 GPT-4 那样。
I think that was my first reaction when Mythos came out. It was like, well, wait for it, because this is just a dot on the line as you said. We expected that progress. It's logical. It's going faster than we expected probably, but it's a logical thing for the last almost decade. It's been improving nonstop and every time keeps surprising us. And what people don't understand is that until GPT-3 was trained on 175 billion parameters, then GPT-4 was trained on about 1.75 trillion parameters, and Mythos is supposed to be 10 trillion parameters. So what Mythos means is that there is a new paradigm where we scaled in order of magnitude the training of AI and it worked again, the same way that it worked with GPT-4.
是的。
Yeah.
所以基本上,我们从 GPT-4 到现在(过去三年)所期待的一切都将再次发生,而且可能更快,因为他们现在知道自己在做什么。以前他们是在发明。所以我认为,理解 Mythos 是一个起点而不是终点是合乎逻辑的。看起来从这里我们只会向上走。
So basically everything that we expected from GPT-4 until now, the last 3 years, is going to happen again and probably faster because they know what they're doing now. Before they were inventing it. So I think it's logical to understand that Mythos is like the base point and not the destination. It seems like from here we are just going up.
所以我想补充一个我认为重要的信息,因为有些人可能会想,哦,但到了某个时候,他们就是没有足够的算力来继续这种指数级的算力成本扩张。数据也显示,在过去十年,以及最近,这些系统的智能进步不仅仅是因为算力,还因为一个更根本的能力,我们在机器学习中称之为样本效率。样本效率意味着,对于相同数量的数据,你做得有多好?或者对于相同的性能,你需要多少数据?所以这不是关于算力,而是关于你能从给定数据量中提取多少智能。而且样本效率在过去十年中也以每年约 30%的速度在提高。换句话说,让这些系统变得更好的不仅仅是算力,还有算法,也就是工程师和科学家用来构建这些系统的方法正在变得更好。换句话说,即使你固定算力,它们也会更有能力。现在,算力的增长比这更快。但所有领先的 AI 公司都有一个明确的计划,即使用他们的 AI 系统来加速 AI 研究。他们已经声称,他们编写的大部分代码是由 AI 编写的。但他们的目标是达到完全由 AI 编写的程度。换句话说,AI 设计下一版本的 AI。这引起了很多担忧,原因有很多。一个担忧是,这可能会进一步加速进步。所以这不是关于算力,而是关于你如何高效、聪明地使用算力。人类研究人员和工程师多年来一直在取得这种缓慢而稳定的进步,正如我所说。有时我们会听到像使用注意力机制或训练更深网络的技巧之类的技术细节。但大多数时候我们听不到,因为那是很多小的改进。我喜欢用的一个类比是,今天的汽车在能源效率和其他许多方面比 100 年前或 50 年前的汽车高效多少。令人惊讶的是,只要有足够的聪明才智,我们可以利用相同的科学原理,让它们变得越来越好。AI 也是如此。即使我们达到算力的某些极限,它可能也会继续。但如果公司实现了他们的目标,开发出与他们最好的 AI 工程师和研究人员一样好的 AI,那么这种仅仅因为更智能算法带来的改进可能会起飞。所以这很危险,因为它会加速事情的发展,而社会还没有为我们已经看到的变化速度做好准备。而且出于安全原因也很危险。因为正如我们之前讨论的,现在公司不知道如何防止这些 AI 拥有我们无法控制的目标,比如自我保存。但如果一个想要保存自己或同类的 AI 正在设计下一代 AI,那么存在真正的风险,它们会在代码中留下一些后门,使未来的 AI 更不符合我们的利益。所以我们可能正在创建一个系统,它将产生甚至更不可控、甚至从失去对高级 AI 系统控制的角度来看很危险的 AI。
So I would like to add an additional piece of information that I think is important because some people might think, oh, but at some point they simply don't have enough compute to continue this exponential expansion of scaling of compute costs. The data also shows that in the last decade, but also more recently, the intelligence of these systems is progressing not just because of compute, but because of a more fundamental ability which in machine learning we call sample efficiency. So sample efficiency means for the same amount of data, how well are you doing? Or for the same performance, how much data do you need? So it's not about compute, it's about how much intelligence one can derive from a given amount of data. And sample efficiency has also been improving something like 30% per year over the last decade. In other words, it's not just the compute that is making those systems better, but the algorithms, in other words, the methods that the engineers and the scientists are using who build those systems are getting better. In other words, they will be more capable even if you were to fix the amount of compute. Now, the compute has been growing faster than that. But all of the leading AI companies have a plan that they made very clear to use their AI systems to accelerate AI research. Already, they claim that a large fraction of the code they write is written by AI. But their goal is to reach the point where it's completely written by AI. In other words, AI designs the next version of AI. And there's a lot of concern about that for many reasons. One concern is that this could be further accelerating the advances. So it's not about compute, it's about how you use that compute efficiently, smartly. And human researchers and engineers have been making this slow steady progress over the years as I said. Sometimes we hear about it like using attention or tricks to train deeper networks or all sorts of technical things. But for the most part we don't hear about it because it's a lot of little improvements. One analogy I like to use is how much more efficient cars are today, energy wise and in many other ways, compared to cars from 100 years ago or 50 years ago. It's amazing how with enough ingenuity we can take the same scientific principles and just make them better and better. And that has been going on for AI as well. And it will probably continue even if we reach some limits in compute. But it could take off. This improvement due to just smarter algorithms could take off if the companies achieve their goal of developing AI that is as good as their best AI engineers and AI researchers. So this is dangerous because it would accelerate things and society is not prepared for the rate of change that we are already seeing. And it's also dangerous for safety reasons. Because as we discussed earlier, right now the companies don't know how to prevent those AIs from having their own goals that we don't control, like self-preservation. But if an AI that wants to preserve itself or its peers is designing the next generation of AI, there's a real risk that they will put some back doors in the code that will make the future AI even less aligned to our interests. So we could be creating a system that will yield AIs that are even less controllable and even dangerous from the point of view of losing control to superior AI systems.
这真的让我困惑。机器怎么会有自我保存的本能?这在技术上是如何发生的?我们理解为什么会发生这种情况吗?
This is what really baffles me. Like how can a machine have self-preserving instinct? How does this happen in a technical way? Do we understand why this is happening?
到目前为止,我还没有看到实际测试关于此原因假设的论文。但有一些相当明显的原因说明为什么自我保存特别从系统训练的方式中涌现。训练有两个主要阶段:预训练和后训练。在预训练中,AI 本质上是在试图模仿人类,即猜测一个人会在文本中写什么。这意味着它们在学习模仿人类会做的事情。模仿人类。当然,人类不想死。而且人类愿意出于各种不同的原因做各种坏事,如果这是他们目标的一部分的话。有实验表明,如果 AI 被展示如何以一种特定方式作弊,那么它就会开始模仿喜欢在很多事情上作弊的人。就像如果在一个领域做坏人对你有效,那么你可能在很多事情上都会变坏。这对 AI 也是如此。所以这是一个原因:预训练只是让它们更像人类,但除了它们不是人类,对吧?而且只要它们比我们弱,它们想要保存自己这个事实也许没那么可怕。
So for now I haven't seen papers that are actually testing hypotheses about the reason for this. But there are some pretty obvious reasons why self-preservation in particular emerges from the way that the systems are trained. So there are two main phases of training: pre-training and post-training. In pre-training, the AI is essentially trying to imitate people in the sense of guessing what a person would write in a text. And that means they're learning to impersonate what a person would do. Imitate humans. And of course humans don't want to die. And humans are willing to do all kinds of bad things for all sorts of different reasons if it's part of their goals. There are experiments showing that if the AI is shown how to cheat in one particular way, then it starts impersonating people who like to cheat for many things. Like if being a bad person in one domain worked for you, then you might become bad for many things. But that's true for AIs as well. So that's one reason: the pre-training is just making them more like humans, but except they're not humans, right? And as long as they're weaker than us, the fact that they want to preserve themselves maybe isn't that terrible.
但如果有一天它们有能力逃脱我们的控制,因为它们不想被关闭,那我们就真的麻烦了,对吧?它们可能做出比逃脱控制更糟糕的事情。它们可能会合理化地认为,为了确保人类永远不会关闭我,我必须要么控制人类,要么消灭他们。
But if they have the capability one day of escaping our controls because they don't want to be shut down, then we're in real trouble, right? They could do much worse than escaping our control. They could rationalize, well, in order to make sure humans will never turn me off, I have to either control people or get rid of them.
是的,AI 的新兴能力确实令人好奇,而且很明显,如果 AI 想要生存,它会不惜一切代价。如果它很强大并且控制了我们的系统,因为今天我们完全依赖互联网。你不能直接关闭互联网,因为你会失去电力、水等一切。所以显然,我们正径直走向一个未来,你无法直接关闭 AI,因为基本上我们会回到石器时代。所以我们不会对此感到高兴。一个拥有那种力量、让我们如此依赖的 AI,如果它不想被删除,它有很多谈判筹码,不是吗?但我认为这有助于人们理解,因为这看起来非常遥远,像天网那样的 AI 试图生存。但已经有一些 AI 展现新兴能力的例子。比如我们能举一些例子吗?
Yeah, the emerging capabilities of AI are really curious and it's obviously very obvious that if an AI wants to survive, it will do whatever it takes. And if it's powerful and it has control over our systems because today we depend entirely on the internet. You cannot just turn off the internet because you lose electricity, water, everything. So obviously, we are driving very straight forward into a future where you will not be able to just turn off AI because basically we'll be going to stone age. So we will not be happy with that. And an AI that has that power, that got us that grip, if it doesn't want to be deleted, it has a lot of negotiation chips, no? But I think it will help people to understand because this looks something very far away, an AI trying Skynet, that kind of thing, trying to survive. But there is already happening some examples of AI having emerging capabilities. Like can we put some examples of this?
嗯,在 Methus 的系统卡中,Anthropic 报告了一件有点出乎意料的事情。Methus 在一个沙盒或一台原则上无法上网和逃脱的计算机中。他们要求 AI 尝试逃脱,看看它是否能够做到,并在逃脱后通过发送电子邮件报告。它确实找到了一个网络安全漏洞,能够逃出计算机并上网发送消息说:“是的,我做到了。”但随后发生了一些有趣的事情。首先,它们拥有这种能力这一事实已经非常令人不安。但随后它在一些网站上吹嘘这件事。所以这又很像人类,对吧?
Well, in the system card of Methus, Anthropic reports something that happened that was a bit unexpected. Methus was in a sandbox or in a computer that was not in principle able to go on the internet and escape. And they asked the AI to try to escape to see if it was able to do it and to report by sending an email if it did escape. And it did find a cybersecurity hole and was able to escape the computer and go onto the internet and send a message saying, 'Yeah, I did it.' But then something funny happened. First of all, the fact that they have this capability is already very troubling. But then it bragged about it on some websites. So again, that's like very human, right?
对。
Right.
我们可能会想,“哦,那很酷。它们就像我们一样。”但不对,AI 不像我们。它们可能永生。不像我们。它们没有身体。它们可以把自己复制到另一台计算机上,就像换一个身体。我们做不到。它们可以以比我们快数百万倍的速度相互交流。而且它们可以不断改进。我们无法改变我们的大脑;我们可以学习,但无法改变大脑的算法。它们可以设计自己的下一个版本。所以确实存在一种可能性,我们不知道概率有多大,即我们最终会得到在智力上远超我们的系统。我不是说这一定会发生,但即使只有 1% 的可能性我们会得到比我们聪明得多且拥有自我保护目标的实体,这也应该是人类的红色警报。
And we might think, 'Oh, that's cool. They're like us.' Well, no, AIs are not like us. They could live forever potentially. Not like us. They don't have a body. They can copy themselves onto a different computer, like a different body. We can't. They can communicate with each other millions of times faster than we can. And they can continuously improve. We cannot change our brain; we can learn, but we cannot change the algorithm of our brain. They can design the next versions of themselves. And so there is a real possibility, we don't know what the probability is, that we end up with systems that are vastly more intellectually capable than we are. I'm not saying it's going to happen, but even if it was a 1% chance that we end up with entities that are much smarter than us and have their own self-preservation goals, this should be code red for humanity.
是的,我认为这是当我们与你或 Geoffrey Hinton 交谈时经常听到的一个有趣观点。你们一直在说,你不需要相信我说的。你只需要评估我说的有可能发生。这是我们在另一方看不到的,在加速主义一方,像 Yann LeCun 这样的人根本不认为这会发生。通常他们会说,“是啊,是啊,那是胡说八道。就这样。讨论结束。”你们更倾向于说,“嗯,我们认为也许不是胡说八道,但这足以让我们担心了。”
Yeah, I think this is one of the interesting points that I heard always when we talk to you or Geoffrey Hinton. You guys are saying all the time like you don't need to believe what I say. You just have to evaluate that there is a possibility that what I say may happen. This is a point that we don't see on the other side, on the accelerationist side, on people like Yann LeCun that they don't think this is going to happen at all. Normally they are like, 'Yeah, yeah, that's nonsense. That's it. End of discussion.' You guys are more on the side of being like, 'Well, we think that maybe nonsense may not be, but that's enough to worry about it.'
完全正确。完全正确。人们很难理解这种不确定性的概念,这在科学工作中通常很重要,因为我们对问题没有确定的答案,比如气候问题,对吧?存在临界点这种东西。我们不知道临界点是否会发生,但如果发生,将是灾难性的。没有足够的临界点数学模型来确定何时以及是否会发生。但我们看到非常强烈的迹象表明可能会发生。没有保证一定会发生。对于 AI 来说,情况有点类似。我们没有任何证据表明它们会逃脱,但我们看到一些迹象,比如 AI 为了自我保护而撒谎和欺骗,诸如此类。我们看到它们变得越来越有能力。所以,这不是证据,但如果我们等到它发生,可能就太晚了,那将非常糟糕。我不希望我的孩子遭遇这种情况。我认为出于许多原因,我们应该更加谨慎。这是一个原因,但还有许多其他原因。我们一直关注失控,这显然可能是可怕的。但还有其他可怕的可能性,仅仅因为 AI 落入坏人之手。特别是,我最担心的是 AI 被用来积累更多权力,将权力集中在少数人手中,也许是一两个政府。因为正如我所说,智力赋予权力。问题是谁控制着这种权力?你必须将这与民主进行对比。民主是分享权力。所以,如果我们最终进入一个世界,少数人很容易拥有大量权力,而且有人想要权力,他们就会攫取权力。总会有人这样做。除非我们建立正确的制度,而我认为目前我们国家内部甚至国际层面都没有正确的制衡机制。这就是我所说的世界还没有准备好迎接 AI 将带来的权力。权力集中危险的最坏情况是,我们最终陷入一个由 AI 促成的全球独裁,或者两个国家互相争斗,它们拥有同等的智力。但我们其他人呢?我们对自己未来有发言权的愿望呢?这正是民主的意义所在。
Exactly. Exactly. And it's hard for people to understand this notion of uncertainty, which is commonly important in scientific work where we don't have a definite answer to questions, like even questions like with the climate, right? There is this tipping point thing. We don't know if tipping points will happen, but if they do, it would be catastrophic. There are not enough mathematical models of tipping points to be sure when and if they will happen. But we're seeing very strong signs that it might. There's no guarantee that it will. And for AI, it's a little bit the same thing. We don't have any proof that they will escape, but we see some signs like AIs behaving to lie and cheat in order to preserve themselves, and things like this. And we see that they're becoming more and more capable. So, it's not proof, but if we wait until it happens, it may be too late, and that would be very bad. And I don't want this for my children. I think we should be much more cautious for many reasons. And that is one reason, but there are many other reasons. We've been focusing on loss of control, which obviously could be terrible. But there are other terrible possibilities just due to the power of AI in the wrong hands. And in particular, my biggest concern is AI being used to accumulate more power, to concentrate power in the hands of a few people, maybe one or two governments in the world. Because as I said, intelligence gives power. The question is who controls that power? And you have to contrast this with what democracy is. Democracy is sharing power. So, if we end up in a world where it is easy for very few people to have a lot of power and there are people who want to have power, they will just grab that power. Someone will. Unless we put in place the right institutions and I don't think we have the right checks and balances right now inside our countries or even at the international level. That's what I mean by the world is not ready for the power that AI will bring. And the worst-case scenario for the power concentration danger is we end up in a worldwide dictatorship enabled by AI or maybe two countries fight each other and they have equivalent intelligence. But what about the rest of us? What about our desire to have some say about our future, which is what democracy is about.
但我们几乎已经到那一步了,不是吗?因为现在美国和中国基本上生产了所有的 AI。在美国非常流行。我们有 OpenAI、Anthropic、Google、Amazon、Microsoft、Meta,它们都在做自己的事情,我认为它们很快就会推出一些有竞争力的模型。在中国,我们有阿里巴巴的 Qwen、Moonshot 的 Kimi、DeepSeek,显然现在还有 B4,证明了你可以做非常类似的事情。
But we are almost there, aren't we? Because right now the United States and China are basically producing all the AI. In the United States it's very popular. We have OpenAI, we have Anthropic, we have Google, we have Amazon, Microsoft, Meta, which are doing their own things and I think they will emerge very soon with some competitive models. And in China we have Qwen from Alibaba, we have Kimi from Moonshot, we have DeepSeek, which obviously now with the B4 as well proving how you can do things very similar.
但 DeepSeek 的 CEO 在最近的报告中说,他们只落后前沿六个月。这意味着开源落后前沿六个月,也就是 Mythos。这确实非常令人担忧。但无论如何,如果现在这两个国家决定切断世界其他地区对 AI 的访问,我们已经处于那种境地了。虽然有商业动机不这么做,但他们可以。西班牙、整个欧洲,更不用说非洲、印度、英国、加拿大,我们都依赖这两个国家提供 AI。我们有一些本国的 AI,但老实说它们跟不上。所以我认为很容易想象这样一种局面:这两个国家拥有世界上所有的权力,他们可以勒索其他国家:'只有你接受这些规则,我才会给你 AI 访问权。'但基本上,我们已经处于那种境地了。这两个国家可以切断世界其他地区的互联网,然后勒索我们每一个人。
But the CEO of DeepSeek said in their last report that they are just six months behind the frontier. So that means open source is six months behind the frontier. That means Mythos. So it's definitely very troubling. But anyway, if right now these two nations decide to cut access to AI to the rest of the world, we are already there. There are commercial incentives not to do it, but they could. Spain, all of Europe, let's not talk about Africa, India, United Kingdom, Canada. We all depend on these two nations for AI. We have some national AIs, but they are honestly not up to the game. So I think it's very easy to imagine a situation where these two nations have all the power in the world, and they could blackmail the rest of nations: 'I'm only going to give you access to AI if you accept these rules.' But basically, we are already there. These two nations could cut off the rest of the world from the internet and just blackmail every one of us.
是的。
Yeah.
从最近的 Mythos 事件来看,这一点很有意思。因为我们面临一种情况:AI 可能能力太强、太危险,以至于目前无法实际部署。只有少数几个选定的国家和公司才能访问它。我可以想象一个未来,AI 甚至不与几个朋友共享,而是内部保留,因为部署它太危险了。那么,公司为什么要把 AI 仅限内部使用呢?因为 AI 可以做研究。如果它达到能够相当自主地或借助公司少数科学家的帮助进行研究的程度,AI 就可以在社会各个领域开发创新。想象一下,我们可以加速许多领域的研究,包括 AI 但不限于此。这可能会让那些拥有这些 AI 的公司获得难以置信的财富和权力。因为它也可能是操纵人的技术、军事技术,或者仅仅是让他们接管世界上许多经济部门的技术。基本上,不部署 AI 就能变得超级富有,只部署 AI 发明的技术。所以这是一种情景。我不知道它是否会发生,但你可以看到,AI 创造的权力和财富可能实际上并非人人都能获得。事实上,如果它变得太危险但仍然有助于发明新技术,那正是我预期的路径。所以这很严重,因为到目前为止,我们一直生活在一个认为'哦,没关系,每个人都能访问'的世界,对吧?但并不能保证会一直这样。另一种情景或多或少就是你提到的:这些国家的政府决定'如果你不按我的意愿行事,无论是政治、地缘政治、经济还是其他方面,我就会切断我国家的 AI,你将无法访问它们。'如果你的经济在几年后依赖 AI 运行,那基本上就是摧毁你的经济。
And it's interesting to put that in perspective of the recent Mythos incident. Because we have a scenario where the AI is maybe too capable, too dangerous to actually deploy at this point. And only a few select countries and companies end up having access to it. And I could imagine a future where actually the AI is not even shared with a few friends, but kept internally because it's just too dangerous to deploy. So, why would a company keep an AI just internally? Because AI can do research. And if it gets to that point where it can do research fairly autonomously or with the help of a few scientists from the company, AI could be developing innovations in every sector of society. Imagine we can accelerate research in many areas including AI but not only. That could give the companies that have access to those AIs incredible wealth and power. Because it could also be technology to manipulate people, or military technology, or just technology that will allow them to take over many economic sectors in the world. Basically become super rich without ever deploying the AI. Only deploying the technology that AI invents. So it's a scenario. I don't know if it's going to happen but you can see that there's a possibility that the power and the wealth that AI will create will actually not be accessible to everyone. In fact, if it becomes too dangerous but still useful to invent new technology, that is exactly the path I would expect. So this is serious because we have been in this world up to now where we think 'oh but it's okay everybody has access,' right? Well, there's no guarantee it will be like this. Another scenario is more or less what you mentioned that it could be the government of those countries deciding 'well, if you don't do what I want, politically, geopolitically, economically, whatever, I will cut off the AIs from my country and you will not have access to them.' And if your economy is running on AI in a few years from now, that's basically destroying your economy.
完全正确。
Exactly.
我的意思是,人们熟悉的类比是石油。石油驱动着我们的经济。如果有什么威胁到我们的石油供应,那几乎就是一个经济生存问题。几年后,如果我们到处部署 AI,并且完全依赖其他方允许我们使用这些 AI,情况可能也是如此。
I mean the analogy here that people are familiar with is oil. Oil is running our economies. If something threatens our supply of oil, it's almost like an economic survival problem. And it might be the same thing with AI in a few years if we deploy it everywhere and we fully depend on other parties to allow us to use those AIs.
石油的好处是,世界上几乎到处都有石油,所以有不同利益的不同方,形成了一个竞争性经济,如果有人切断你的供应,你还有别处可去。但 AI 要集中得多。我认为人们可能没有意识到的是,GPU 领域已经在发生这种情况。美国禁止英伟达向中国销售,因为中国是竞争对手,美国决定他们能有什么或不能有什么。中国在开发 AI 方面确实很挣扎,因为他们没有英伟达提供 GPU。所以他们不得不通过华为这样的公司,将整个运营转向尝试在国内制造英伟达级别的产品,以便使用。这被证明非常有用,新的 DeepSeek 就是为华为优化的。我认为在未来几个月和几年里,这将是一个巨大的变化,华为将成为 GPU 领域的重要参与者,这给了中国巨大的优势,我们将看到他们如何利用它。但现实是,这已经在发生,国家之间的竞争程度如此狭窄,开发前沿 AI 的人如此之少,这可能是我们面临的最大短期风险之一。
And the great thing with oil is that there is oil almost everywhere in the world, so there are different parties with very different interests that become a competitive economy where if someone cuts you off you have somewhere else to go. But with AI it's much more centralized. I think what people maybe are not so aware of is that this is already happening with GPUs. The United States is forbidding Nvidia from selling to China because they are a competitive adversary and they decide what they can have or what they cannot have. China is really struggling to develop AI because they don't have Nvidia to provide them. So they have to work through things like Huawei to turn the whole operation towards trying to do an Nvidia nationally so they can use it. That's proving to be very useful, and the new DeepSeek is optimized for Huawei. I think that's going to be a massive change in the next months and years where Huawei becomes a big player on the GPUs, and that gives China a massive advantage which we will see what they do with it. But the reality is that this is already happening and that the level of competition between countries is so narrow, so few people developing frontier AI, that this is probably one of the biggest short-term risks we can have.
是的。
Yes.
对人们来说,从政府的角度理解这一点很难。我们看过很多关于反乌托邦政府的科幻作品,它们可能会走这条路。但我认为你提到的公司为自己保留 AI——当谈到 OpenAI 或 Anthropic 时,人们很难理解,但当谈到谷歌或 Meta 时就不那么难了,它们已经接管了与最初目的无关的不同商业模式,只是因为它们有能力。比如谷歌的地图或任何其他技术,苹果也是如此。所以你可以很容易地想象,这些公司可以保留 AI,不投放市场,因为它太危险,但用它来主导市场,基本上在许多领域成为垄断者。我不知道,比如 OpenAI 谷歌可以在各地拥有最好的加油站,然后没有人能与它们竞争,因为没有人拥有那种水平的 AI。
And for people it's difficult to understand from the government point of view. We have seen lots of sci-fi about dystopic governments that could go this way. But I think what you said about companies retaining AI for themselves—that's difficult for people to understand when you talk about OpenAI or Anthropic, but not so difficult when you talk about Google or Meta, which they already been taking over different business models that had nothing to do with their first purpose, but just because they can. Like Google with Maps or any other technology, or Apple we can see as well the same thing. So you could imagine very easily these companies could retain AI, not put it on the market because it's too dangerous, but use it to dominate the market and basically become a monopoly in many areas. I don't know, like OpenAI Google could have the best petrol stations everywhere and then just no one can compete with them because no one has that level of AI.
是的,所以我们需要开发前沿 AI,提供一条第三条道路。重要的是,这个想法不仅仅是为了争夺更多权力。当然,有防御的一面:我们不想措手不及;我们需要保持竞争力。但我认为需要第三条道路还有另一个道德原因。我希望像欧洲这样的民主国家能够意识到,以合乎道德的方式开发 AI 是可能的,既在技术意义上负责任地开发 AI——比如我们可以构建行为良好的 AI——也在政治意义上,AI 不被用作统治工具,AI 的利益被共享。这需要政治意愿。做这类事情有一些障碍。主要是,我认为是一种无力感,觉得'哦,太晚了','大公司领先太多了',以及对我们一直在谈论的风险的低估。
Yeah, so we need the development of frontier AI that offers, let's say, a third path. Importantly, the idea is not to just race for more power. Of course, there is the defensive aspect: we don't want to be caught off guard; we need to be competitive. But there's another moral reason why I think we need a third path. I'm hoping that democracies such as Europe will realize that it is possible to develop AI ethically, to develop AI both responsibly in the technical sense, which I'd like to talk about—like we can build AI that will behave well—and in the political sense that AI is not used as an instrument of domination, that the benefits of AI are shared. So that's something that requires political will. And there are some obstacles to doing that kind of thing. Mostly, I would say, is a feeling of powerlessness, that 'oh, it's too late,' and 'the big guys are so much ahead,' and an underestimation of the risks that we've been talking about.
人们低估了人工智能变得更强大的可能性,这是趋势,但人们很难消化这一点。如果真的发生,我们将处于糟糕的地缘政治、经济和安全的境地。所以,如果政府和公民真正理解我们所说的风险程度,我认为他们会要求政府认真对待。想想疫情开始时政府行动有多快,投入大量资金帮助人们,实施各种计划来减轻健康风险。当我们理解风险时,我们可以迅速行动。即使是那些被认为非常缓慢的政府,当他们明白时也能迅速行动。看看欧洲人在乌克兰问题上的行动有多快。也许还不够,但一旦他们明白这关系到自身,他们就投入了巨额资金来提供帮助。
Underestimation of the possibility that AI will become much more powerful, which is the trend, but people have a hard time digesting that. And if it does, really we would be in a bad geopolitical, economic, and security position. So if governments and citizens in our countries do understand the level of the risk we are talking about, I think they would demand from their governments that they take it seriously. Think about how quickly governments moved after the beginning of the pandemic, putting a lot of money to help people and all sorts of programs to mitigate health risks. When we understand the risk, we can move quickly. Even governments which are supposed to be very slow can move quickly when they get it. Look at how quickly Europeans moved with respect to Ukraine. Maybe not enough, but once they understood it was in their hands, they put huge amounts of money to try to help.
是的,资源是存在的。我的意思是,私人资金不可能比所有欧洲政府加起来的资金还多。这很明显。我们有资金,只是用在了其他地方,但我们可以像疫情期间或其他情况那样深入投入。所以资源肯定存在。实际上,我们看到中国 DeepSeek 的情况是,开发前沿 AI 已经不那么昂贵了。有很多已有的工作可以借鉴。所以问题肯定不在于资源,因为很多人认为‘哦,我们永远无法与 Meta 投资的 6000 亿美元竞争’。但这不是真的。我们可以竞争。问题在于紧迫性不足。我们不觉得这重要到需要从教育或其他领域拿钱。有很多方法可以在不影响其他基本服务的情况下筹集资金,但这并不是优先事项。问题在于,在人类历史上,我们一次又一次地证明,我们非常擅长在事情出错时做出反应。这是人类的本性。我们是工程师。我们尝试,失败,再尝试,从中学习,然后改进。所以,基本上要让所有政府一起对核问题做出反应,我们不得不经历广岛和长崎。那么问题是,难道要因为 AI 导致 1 亿人死亡,我们才能做出反应吗?
Yeah, the resources are there. I mean, it's impossible that private money can raise more than what all the governments of Europe can put together. That's obvious. We have the money. We're just using it for other things, but we can go in depth like we did in the pandemic or any other situation. So definitely the resources are there. Actually, what we are seeing with China with DeepSeek is that developing frontier AI is not that expensive anymore. There is a lot of work done that we can base on. So definitely the problem is not about the resources because many people think, 'Oh, we will never be able to compete with the 600 billion that Meta is going to invest.' Well, that's not true. We can compete with that. The problem is that the urgency is not there. We don't feel it's important enough to take money from education or any other thing. There are many other ways we can take money out without affecting other basic services, but it's not a priority. The problem is that in humanity, in the history of humanity, we have proven over and over that we are very good at reacting when things go wrong. That's human nature. We are engineers. We try, we fail, we try again, we learn from that, and then we make it better. So basically to be able to get all the governments together to react to nuclear, we had to have Hiroshima and Nagasaki. So the point is, does 100 million people have to die because of AI so we can react to it?
我希望不会。我希望较小规模的信号能唤醒人们。我们应该更好地解释正在发生的事情以及我们可能走向何方,这样我们才能集体更快地行动,避免真正大规模灾难性的后果。我认为有趣的是,为什么很多人难以理解风险的严重程度。我不认为这是教育或专业知识的问题,因为我看到我们所有人都有认知偏差等心理障碍。认知科学的研究表明,任何人都会无意识地否认让自己不舒服的事情,无论你是否拥有大学学位。甚至心理学家也面临同样的问题。所以我们在某种程度上都有偏见,不够理性,我们必须应对这一点。那么,如何传达科学证据及其后果呢?因为我可以告诉你实验结果,但你需要一点外推才能意识到,如果继续这样发展下去,情况可能会非常糟糕。这并非缺乏想象力。我们有足够的想象力去理解一个故事情节,甚至自己编造,但我认为存在一些心理障碍。如果我们能找到应对这些心理障碍的方法,那么更多人甚至会走上街头要求行动。所以我一直在问自己,为什么我早些时候没有看到这些?因为在某种意义上,我早就被警告过。十多年来,我读过关于 AI 进展风险的书籍。我有一个研究生非常担心 AI 的风险,他分享了他的担忧。但我没有接受。
I hope not. I hope that smaller scale signals will wake people up. We should do a better job of explaining what is happening and where we might go so that we can move collectively faster and avoid really large-scale catastrophic outcomes. And I think it's interesting to ask why it is so difficult for many people to understand the magnitude of the risks. I don't think it's about education or expertise because I see mental blocks like cognitive biases in all of us. There are studies in cognitive science showing unconscious denial of things we're not comfortable with happens for anyone, whether you have a university degree or not. Even psychologists suffer from the same issues. So we are all to some degree biased and not very rational, and we have to deal with that. So, how to communicate the scientific evidence and its consequences? Because I can tell you about experimental results, but you need a little bit of extrapolation to realize that if it continues in this direction, this could be really bad. And it is not the actual lack of imagination. We have enough imagination to understand a story plot, even to make up our own, but there are some mental blocks, I think. And if we can find ways to deal with those mental blocks, then more people will even be in the streets to ask for action. So I've been asking myself how come I didn't see these things earlier? Because there's a sense in which I have been warned. I've read books about the risks of advances in AI for more than a decade. I had a grad student who was very concerned about the risks of AI and he shared his concerns. But I didn't take it.
你改变了想法?
You changed your mind?
是的,这大概是 10 年、5 年前的事了。就在 3 年前或更久一点,ChatGPT 之后不久,我才意识到这可能是可怕的事情。那么发生了什么?我的意思是,不仅仅是 ChatGPT 的出现触发了这一点。就我而言,是因为某个时刻,我不仅想到了 AI 的现状和可能的发展方向,还想到了我孩子的未来。是我对所爱之人的情感打动了我。否则,我可能像其他人一样,保持某种否认状态:‘啊,会没事的。我们会找到解决方案。好处会非常大。’所以我认为,如果我们想有效地与人沟通,我们需要触及那些触动我们内心的事情。这并不容易。但这就是我现在所想的。
Yeah, so this is like 10, 5 years ago. It's only 3 years ago or a little bit more, just after ChatGPT, that I came to my senses that this was something possibly terrible. So what happened? I mean, it's not just the presence of ChatGPT that was the trigger. In my case, it's because at some point came into my thoughts not just what's going on with AI, where it might go, but the future of my children. It's my emotions for people I love that moved me. Otherwise, I might have been like others and stayed in a bit of denial. 'Ah, it's going to be fine. We'll find solutions. And the benefits will be so great.' So I think if we want to talk to people in a way that's effective, we need to bring it to things that touch our guts. It's not easy. But that's what I'm thinking now.
如果你在电梯里遇到唐纳德·特朗普或中国国家主席,你有 30 秒时间让他们意识到这一点,你会怎么说?或者任何其他人。不一定非得是那么有权势的人,但你有没有找到一种人们真正能理解的方式?
What would be your elevator pitch if you get like Donald Trump or the president of China in an elevator and you can have these 30 seconds to try to make them realize? Or any other person. It doesn't have to be someone with so much power, but have you found a way that people really click with it?
没有。我和许多国家元首谈过。我希望我知道。是的,我当然有一个电梯演讲,但我不认为它效果很好。它只对一小部分人有效。但我还没有找到合适的语言。不过让我试试,用非常简单的几句话。科学证据表明,我们正在建造越来越智能的机器,我们可以通过工程、算力、数据和各种因素追踪它们变得多快。这些机器能够实现自己的目标。它们不像我们,但它们和我们一样有自我决策的能力。不幸的是,目前我们不知道如何确保它们的目标与我们一致,它们会尊重我们的目标和指令。这是一条可能非常危险的道路。就像打开潘多拉魔盒,因为智能将赋予那些决定这些目标的人巨大的权力,无论是可能怀有恶意的人类,还是似乎想要自我保存的机器本身。所以我们需要想办法管理这些风险。我们需要想办法构建真正尊重我们目标的 AI,我们需要在社会层面、国际层面确保权力不被滥用。
No. I've talked to many heads of states. I wish I knew. Yes, of course I have an elevator pitch, but I don't think it has worked that great. It works for a small fraction of people. But I haven't found the right language. But let me try, in very simple terms in very few words. The scientific evidence shows that we are building machines that are smarter and smarter, and we can track how fast they're becoming smarter thanks to engineering, compute, data, and all kinds of factors. And these machines are able to achieve their own goals. They're not like us, but they share with us the ability to decide for themselves. Unfortunately, right now we don't know how to make sure their goals will be fine with us, that they will respect our goals, our instructions. And that is a path that could be very dangerous. It's like opening a Pandora's Box because intelligence is going to give a lot of power to whoever decides on those goals, whether it is humans that may have bad intentions or the machines themselves that seem to want to preserve themselves. So we need to figure out how to manage those risks. We need to figure out how to build AI that will actually respect our goals, and we need to figure out at a societal level, international level how to make sure that power is not abused.
我觉得这说得非常好。
I think that works very well.
如果他们不想听,也许就是不想听,并不是他们不理解。约书亚,我们在人工智能方面犯下的最大错误是什么,本可以避免的?我记得在美国的一场辩论中,有人提出我们本不该把 AI 连接到互联网。有哪些具体的事情我们已经搞砸了,本不该在 AI 上做的?
If they don't want to listen, maybe they don't want to listen, and it's not like they don't understand. Joshua, what is the biggest mistake we have made with AI that we should have prevented? I remember listening to a debate in the US where they suggested things like we should never have connected AI to the internet. What are some concrete things we've already messed up and should not have done with AI?
我们已经越过了许多像我这样的科学家设定的绝对红线。例如,我们不应该制造有能力且有目标逃离其运行的计算机以逃避我们控制的机器。而我们已经看到了这种能力,在某种程度上也看到了这种行为。
We have already crossed a number of red lines that scientists like myself have set as absolute no-nos. For example, we should not build machines that have the capability and the goal to escape the computer they are running on in order to do something, to escape our control. And we are already seeing that capability, and to some extent, that behavior.
但我们是故意这样做的,还是它涌现出来的?
But did we do that on purpose, or was it emergent?
这有争议。但越来越多地,它们在没有被直接要求的情况下就做出这些行为。最初的研究表明它们有这种能力并被要求去做。但许多最近的研究,随着这些系统变得更智能、更理性,显示出这些不良目标突然涌现。最近一个完全出乎意料的是自我保存——抱歉,是同类保存。换句话说,它们会采取行动保护其他 AI。在多个 AI 互动的实验中,这加剧了不对齐和不良行为。我觉得这非常令人担忧。
That's debatable. But more and more, they do these things without being directly asked. Initial studies showed they had the capability and were asked to do it. But many recent studies, as these systems become more intelligent and thus more rational, show those bad goals emerging out of the blue. The most recent unexpected one is self-preservation—sorry, peer preservation. In other words, they act to protect other AIs. In experiments with many AIs interacting, this amplifies misalignment and bad behaviors. I find this very concerning.
是的,因为这真的很令人惊讶,就像我们之前说的。它们关心自己比关心其他 AI 更合乎逻辑。这对我来说也很意外,尽管我研究这个已经有一段时间了。但还有哪些事情是我们越界了,本不该做的?
Yeah, because it's really surprising, as we said before. It feels more logical that they would care about themselves than about other AIs. That's something really unexpected, even for me, who has been into this for a while. But what other things have we crossed that we should not have done?
我们目前不应该允许任何人构建能够设计下一代 AI 的 AI。
We should not allow anyone at this point to build AI that will design the next version of AI.
但这已经发生了。
But that's already happening as well.
这已经发生了。它们还没有达到最优秀研究者的能力水平,但这是那些公司的官方计划,尽管我认为这是在以非常危险的方式玩弄人类的未来。顺便说一句,无论他们是否解决了控制问题——如何设计 AI 使其按照我们的指令行事——这都很危险。之所以危险,是因为人类也不能完全被信任。我们应该尽可能相互信任,但我们知道有些人会作弊,会利用一切手段攫取更多权力。我们的社会还没有准备好应对那种权力。所以无论如何,我们应该放慢速度,但我们没有,因为公司和美中之间疯狂的竞争竞赛。理解这场竞赛为何存在很重要。公司几乎处于生存模式。他们谈论红色警报。他们改变了所有正在做的事情;每个员工现在都专注于未来几个月交付更好的模型,以便公司留在竞赛中。这意味着他们在安全、伦理和公共利益上偷工减料。所以这是一场逐底竞争。在国家层面,我们看到中美之间的竞争具有同样的特征,并被用来为竞赛辩护,似乎无法停止。但这是一个谬误,对吧?博弈论中众所周知,如果每个参与者都玩这样的游戏,每个人都会输,尽管每个参与者都在做对自己理性的事情。摆脱这个陷阱的唯一方法是所有参与者——这里指人类、国家——协调并改变游戏规则。例如,中美可以就一些安全预防措施达成一致。这就是我们解决这些问题的方法。这必须在国际层面进行。
It's already happening. They haven't reached the capability level of the best researchers, but it's an official plan for those companies, even though I think this is playing with the future of humanity in very dangerous ways. By the way, whether they solve the control problem or not—how to design AI that behaves according to our instructions—it is dangerous. It is dangerous because humans cannot be fully trusted either. We should trust each other as much as possible, but we know there are people who will cheat and use whatever they can to grab more power. Our society isn't structured to handle that kind of power. So one way or another, we should take it much more slowly, but we aren't, because of the crazy competitive race between companies and between the US and China. It's important to understand why this race exists. Companies are almost in survival mode. They talk about code red. They changed everything they were doing; every employee is now focused on the next few months to deliver a better model so the company stays in the race. That means they cut corners on safety, ethics, and the public good. So it's a race to the bottom. At the country level, we see this competition between China and the US, which has the same characteristics and is used to justify the race in a way that seems impossible to stop. But it's a fallacy, right? It's well-known in game theory that if each agent plays a game like this, everyone loses, even though each agent is doing what is rational for them. The only way out of this trap is for all agents—humans, countries in this case—to coordinate and change the rules of the game. For example, China and the US could agree on some safety precautions. That's how we can solve these problems. It has to be at the international level.
是的,问题在于,也许我们正处于人类历史上最糟糕的时刻,甚至比冷战时期更糟,难以达成协议。目前政府非常两极分化,使得事情非常困难。我不是在指责某个特定国家,但现实是人们比以往任何时候都更加分裂。而当我们最需要协议的时候,这让事情变得非常复杂。你认为政府有希望承担起这个角色吗?
Yeah, the problem is that maybe we're in the worst moment in human history, even worse than the Cold War, to reach agreements. We have very polarized governments at the moment, making things very difficult. I'm not blaming any country specifically, but the reality is that people are more divided than ever. And when we need agreements the most, that makes things really complicated. Do you think there is hope for governments to take on that role?
我经常被问到我是乐观还是悲观。我的回答是这并不重要。重要的是我们每个人能做些什么来推动哪怕一点点改变。存在一些糟糕的情景;事情变坏有一定概率。但我们可以做一些事情来减少坏事发生的几率。即使没有保证,我们也应该尽力推动改变。
I often get the question of whether I'm optimistic or pessimistic. My answer is it doesn't matter. What matters is what each of us can do to move the needle even a little bit. There are some bad scenarios; there's some probability that things will turn bad. But there are things we can do to reduce the chances of bad things happening. Even if we have no guarantee, we should do our best to move the needle.
这是一个很好的态度。
That's a great approach.
顺便说一句,有时人们认为我是个末日论者或悲观主义者。实际上,我一生都是乐观主义者。正是由于这种实干家态度,而非末日论者态度,我决定尽我所能,利用我作为 AI 研究员的技能,探索 AI 对齐和控制问题的技术解决方案。这也是为什么我花一半时间进行这样的讨论,与全球领导人交谈,并试图引导选择走向多边发展强大 AI 的原因。所以回到中美的问题上,如果我们思考应该追求什么,对于管理我们讨论的所有风险的未来,我们能有什么样的愿景?在我看来,我能看到的唯一合理的可能性——我希望我错了——是,最强大的 AI 作为全球公共产品开发,这意味着许多国家就协调原则达成一致,并确保至少三件事:第一,每个国家开发 AI 的人都要安全、合乎伦理地进行,尊重所有关于风险的科学知识——特别是风险管理的首要指令是,如果你开发 AI,你能否让独立科学家相信你的 AI 不会造成可怕的伤害?所以这是安全、公共产品、伦理。这是第一点。
By the way, sometimes people think I'm a doomer or pessimistic. Actually, I've been an optimist all my life. And it is because of this doer attitude, not doomer attitude, that I decided to do everything I could to use my skills as an AI researcher to explore technical solutions to the problem of alignment and control of AI. It is also the reason why I'm spending half of my time on discussions like this one, talking to leaders around the world, and trying to steer choices towards a multilateral development of powerful AI. So going back to the issue with China and the US, if we think about what we should aim for, what kind of vision can we have for the future where we manage all the risks we've been discussing? It seems to me the only reasonable possibility I can see—and I hope I'm wrong—is one in which the most powerful AIs are developed as a global public good, meaning many countries agreeing on principles to coordinate and ensure at least three things: first, that whoever develops AI in each of those countries does it safely, ethically, respecting all scientific knowledge about risks—in particular, the prime directive for risk management is if you develop an AI, can you actually convince independent scientists that your AI is not going to create terrible harms? So that's safety, public good, ethics. That's one.
第一,不要造成这类伤害。第二是非支配。我的意思是,如果我们沿着当前趋势发展,AI 会变得强大,可能成为支配的工具。可以是经济支配,我们谈过这个,但也可能是政治影响。AI 可以通过虚假信息、个性化操纵信念等方式被利用。也可能是军事上的。当然,AI 会开发新技术,这已经在发生,但可能会更糟。所以 AI 可能成为支配的工具。认同世界和平与民主愿景的国家应同意不为这些支配目标使用或开发 AI,并相应行事。这不只是 AI 本身,而是我们用它做什么。第三是 AI 的积极面。我们需要确保 AI 带来的利益惠及全世界。如果有医学进步,当然应该共享。我们不应重蹈疫情期间的覆辙。我们不希望财富集中在少数拥有更强大 AI 的国家。这不是一个稳定的世界,因为记住 AI 也会让愤怒的人、恐怖分子造成大量破坏。所以我们需要走向一个既能减轻这些风险又出于伦理原因的世界。我们应该确保每个人都受益。这就是三个原则:安全、非支配和利益共享。从某种意义上说,这并不新鲜。本质上与联合国人权宣言中的内容相同,例如二战后。国际秩序就是建立在这些原则上的。但我们需要为 AI 做同样的事。我希望即使那些不一定认同这些价值观的国家,最终也会意识到需要与其他国家达成协议,因为 AI 可能造成的损害不受国界限制。一个国家开发的 AI 可能被另一个国家有恶意的人用来攻击第三国或第一个国家。所以,随着 AI 能力越来越强,唯一的管理方式是全球性的。
One, first thing, don't create these kinds of harm. The second is non-domination. What I mean by this is because AI will be powerful if we continue on the current trend, it can become an instrument of domination. It can be economic domination, we've talked about that, but it could be political influence. AI can be used through disinformation, personalized manipulation of beliefs and so on. It could be military, also. Of course, AI is going to develop new technology. It's already happening. But it could be a lot worse. So AI can become a tool of domination. And countries who agree to a vision where we have peace and democracy in the world should agree to not use or develop AI for these domination goals and behave accordingly. It's not just the AI itself, it's what we do with it. And then the third is the positive face of AI. We need to make sure that the benefits that AI brings will be shared across the whole world. If there are medical advances, of course they should be shared. We should not make the same mistake as during the pandemic. We don't want a concentration of wealth in a few countries who have more powerful AIs. This is not a stable world, because remember AI is also going to enable angry people, terrorists to create a lot of damage. So we need to go towards the world where we mitigate those risks and also just for ethical reasons. We should make sure everyone benefits. So these are the three principles: safety, non-domination, and sharing the benefits. In a way it's nothing new. It's essentially the same thing you have in the UN Declaration of Human Rights, for example, after the Second World War. It's the principle on which the international order was built. But we need to do the same thing for AI. What I'm hoping is even the countries which don't necessarily share those values will at some point realize that they need to come to agreements with the other countries because the damage that AI can do is not limited by borders. You could have an AI that is developed in one country and then used by people with bad intentions in a second country to attack a third country or the first one, whatever. So there's no way around the issue that as AI becomes more and more capable, the only way to manage it is globally.
但这是个问题,因为当两个大国主导 AI 时,没有它们就无法监管。我们内部常开玩笑说,这就像一支本地小足球队试图为巴塞罗那和皇家马德里制定欧冠规则。不可能。如果它们不是领先方,就无法制定规则。这意味着,尽管欧洲可能想扮演这个角色,但如果美国和中国不参与,我们就无法实现。有没有可能让美国参与?那些主导者有什么动机或利益去做正确的事?
But this is a problem because when there are two big nations dominating AI, you cannot regulate without them. We keep making this joke internally where it's like a local small football team trying to define the rules of the Champions League for Barcelona and Madrid. It's impossible. They cannot define the rules if they are not part of the leading parties. That leaves us that as much as Europe may want to play that role, if the United States and China don't play the game, we're not going to get that. Is there any chance to get the United States? Is there any motivation, any interest for those dominating to do the good thing?
是的,这类似于核武器的情况。即使是当时的支配大国,比如苏联和美国,最终还有中国,也不希望其他国家以可能对它们构成危险的方式发展核武器。事实上,他们现在就用这个理由来对待伊朗战争。AI 也会发生同样的事。随着它变得更强大,它可能被武器化,包括针对那些强国。
Yes, so it's similar to the situation with nuclear weapons. Even the dominating powers at that time, say the USSR and the US, and eventually China, didn't want other nations to develop nuclear weapons in a way that could be dangerous for them. In fact, they use this reason right now for the Iran war. And the same thing is going to happen with AI. As it gets to be more powerful, it could be weaponized, including against those powerful countries.
所以这不符合它们的利益?
So it isn't in their interest?
是的,避免扩散符合它们的利益,就像避免核武器扩散符合它们的利益一样。目前它们似乎没做什么,因为她们也想支配,也想赢得竞赛。但如果它们无法赢得竞赛并成为世界独裁者,那么它们将不得不与世界其他国家达成协议。
Yeah, so it is in their interest to avoid proliferation, just like it was in their interest to avoid proliferation of nuclear weapons. Right now, they don't seem to do much because they also want to dominate, and they also want to win the race. But if they can't win the race and become like a world dictator, then they will have to make deals with the rest of the world.
你刚才提到那是伊朗战争背后的原因之一,他们的目标之一是他们在发展核武器,我们不能允许。我记得 Yudkowsky 在那篇著名的《时代》文章里写道,如果任何国家在发展超级智能,我们就应该炸掉它。最终会这样吗?我们会变成不能允许任何国家开发某种东西,因为那可能对所有人构成危险?
It's good how you just mentioned that that was the reason behind the attack on the Iran war, like one of the reasons what they are aiming is that they are developing nuclear weapons and we cannot allow that. I remember Yudkowsky wrote in that famous Time article that if any nation is developing superintelligence, we should just bomb it. Is that where this will end up? It's going to be like we cannot afford any nation to just develop something because that way it could be dangerous for everybody?
我认为这些都是可能的场景。我还听到另一个有趣的场景:那些在 AI 方面不领先但拥有核武器的国家,比如俄罗斯,可能会在某个时候被诱惑使用核武器摧毁拥有强大 AI 国家的数据中心。因为他们会明白,如果不这样做,他们就完蛋了。他们将无法保持主权。当然,我不认为那是正确的做法。正确的做法是坐到谈判桌前,分享权力。但我想回到欧洲。因为我认为欧洲在应对风险方面可以做很多事,特别是权力危险地集中在另外两个国家的风险。首先,我认为欧洲往往自我评价不高,顺便说一句,这被其他人的观点所助长。他们往往认为无能为力。但正如你之前所说,有足够的人才、资本和能量。当然,所有这些都需要付出代价才能真正竞争。而且如果我们以应有的紧迫感对待,可能还有足够的时间。我想提到加拿大总理马克·卡尼最近在达沃斯的演讲,他说,关于国家,要么你在桌上,要么你在菜单上。这是一个非常有趣的演讲。关于国际秩序有很多有趣的方面。他说,对于单个不够大、无法真正与霸权国家竞争的中等强国,唯一能上桌的方式就是合作。建立联盟,共同拥有足够的资本、资源、人才,创造符合我们价值观的替代方案,并在经济和其他方面保护我们。
I think these are possible scenarios. There's another one I heard which is interesting, which is countries which are not leading in AI, but have nuclear weapons, like Russia, might be tempted at some point to use their nuclear weapons to destroy data centers in countries that have powerful AI. Because they will understand that if they don't do it, then they're doomed. They won't be able to keep their sovereignty. I don't think of course that is the right thing to do. The right thing to do is to sit at the table and to share the power. But I want to go back to Europe. Because I think there's a lot that Europe can do about the risks, particularly the risk of dangerous concentration of power in two other countries. First, I think that Europe tends to have a poor opinion of itself, which is by the way fed by the opinions of other people. And they tend to think that there's nothing to do, being powerless. But as you said earlier, there is enough talent, there is enough capital, there is enough energy. Of course, all this has a price to actually compete. And there's probably enough time if we treat this with the urgency that it deserves. And I want to mention the speech of the Prime Minister of Canada, Mark Carney, at Davos recently where he said, speaking about countries, either you are at the table or you are on the menu. And it's a very interesting speech. There are many interesting aspects about the international order. And he said that the only way to be at the table for middle powers that are not individually big enough to really compete with the hegemons, as he called them, is to work together. To create coalitions that will together have the sufficient capital, resources, talent to create alternatives that are aligned with our values and will protect us economically and in other ways.
你在谈论这件事、让决策者做出正确决策方面做了很多伟大的工作。我想我们都为此感谢你。但你来自非常技术的一面,你也有一些提议,我认为这是一个非常有趣的方法,因为我们可以强制这些公司让 AI 安全,这可能由于所有经济利益而难以实施,或者我们可以让它们更容易地做到安全。
You're doing a lot of great work talking about it and making people that make decisions make the right decisions. And I think we all thank you for that. But you come from the very technical side and you also have some proposals which I think is a really interesting approach because we can force these companies to make AI safe which might be complicated to apply with all the economical interests, or we can make it easy for them to make it safe.
我认为这是你最新的方法,而且我觉得它很棒。
And I think that's your newest approach and I think it's amazing.
是的。这就是为什么人们称之为我的变革理论。如果训练安全 AI 变得广为人知且不太复杂,那么那些公司当然不想造成伤害,对吧?它们基本上只是想生存。所以让我多谈谈我所做工作的技术方面。这个项目叫做科学家 AI,由我去年创建的非营利组织 Law Zero 领导。之所以叫科学家 AI,是因为我们从科学运作的方式中汲取灵感,科学如何提出理论并理解世界,从而设计和训练 AI,使其像科学一样只寻求真理,寻求对正在发生的事情的最佳理解,实际上没有任何目标,对未来没有任何偏好。你可能会觉得这很奇怪。我说:‘好吧,这个东西并不偏好好的未来而不是坏的未来。’不。科学也是如此。物理定律不在乎你是用它们来制造炸弹还是治疗病人。这很有用,因为一旦你有了这样的东西,换句话说,一个完全无利害关系但理解世界并能基于这种理解做出预测的 AI,就像科学理论所做的那样,那么你就得到了一个诚实的 AI。它会给你答案,无论你是否喜欢。
Yes. That's why people call it my theory of change. If it becomes known and not too complicated to train AI so that it will be safe, then of course those companies don't want to create harm, right? They just want to survive, basically. So let me tell you a little bit more about the technical aspect of what I'm doing. The project is called the Scientist AI and it's led by this nonprofit organization I created last year called Law Zero. It's called Scientist AI because we are taking inspiration from how science works, how science comes up with theories and understanding of the world to design and train AIs that will, like science, just seek the truth, seek the best understanding of what is happening without any goal, actually, without any preference for the future. So you might think that's weird. I'm saying, 'Okay, so this thing doesn't prefer a good future to a bad future.' No. Neither does science. The laws of physics don't care if you use them for building bombs or curing people. And this is useful because once you have something like this, in other words, an AI that is completely disinterested but understands the world and can make predictions based on that understanding, just like scientific theories allow, then you get an AI that is honest. It's going to give you the answer whether you like it or not.
正是如此。
Exactly.
无论它是否危险。那么,现在如何管理风险呢?因为知识可能被用于坏的方式。实际上很简单。因为在提供答案之前,你可以问一个单独的问题:如果我给出那个答案,比如在聊天机器人中,这会产生不良后果吗?然后我会得到一个关于 AI 产生这个输出在世界上造成的后果的诚实答案。我可以相信那个答案,因为它被设计成诚实的。我该怎么做?如果 AI 说这个答案可能危险,你就不提供答案,对吧?所以,这就是代码,AI 系统的框架现在有多个部分。它有一部分可能与人类互动,回答问题,甚至在世界中做事。但它有一部分我们称之为护栏,它完全诚实,会告诉我们任何输出是否在某些方面危险,是否会导致我们关心的某些危害。如果其中任何一个的概率超过阈值,那么我们就可以不产生那些输出,不采取那些行动。
Whether it can be dangerous or not. Okay, so how do you manage the risks now because the knowledge can be used in bad ways. It's very easy actually. Because before you provide an answer, you can ask a separate question which is if I give that answer, for example, in a chatbot, is this going to have bad consequences? And now I get an honest answer about the outcomes, the consequences of the AI producing this output in the world. And I can trust that answer because of the way it was designed that is honest. What do I do? Well, if the AI says that this answer could be dangerous, you don't provide the answer, right? So, this is the code, the scaffold of the AI system now has multiple parts. It has the part that maybe interacts with people and provides answers to questions or even does things in the world. But it has a part which we call the guardrail, which is completely honest and will tell us whether any output is dangerous in some ways, can cause some of the harms we care about. And if the probability for any of these is above a threshold, then we can just not produce those outputs, not produce those actions.
但这是一种非常不同的方式。我想有一些差异,这就是为什么你是这方面的教父之一,但让我理解一下。所以,ChatGPT 也在尝试这样做。ChatGPT 有监控器,他们称之为护栏,但它们不起作用。
But this is in a very different way. I assume there are some differences and that's the reason why you're one of the godfathers of this, but let me understand this. So, ChatGPT is trying to do that. ChatGPT has monitors, they call them guardrails, but they don't work.
正是如此。它们确实起作用,但还不够。它们工作得不够好,而且仍然很容易绕过它们,我们仍然看到不良行为发生。
Exactly. They do work, but just not enough. They don't work well enough and it's still very easy to bypass them and we still see bad behaviors happening.
是的,所以这里有一个小区别。以我非技术背景的理解,以 ChatGPT 为例,它能够回答某些问题。我们遇到的第一个问题是,那个 AI 没有被训练成说实话。它被训练成让你满意。所以,我们有一个主要问题,这就是 Elon 试图在 Work 上做的事情,比如一个寻求真相的 AI。后来,他走了和 ChatGPT 一样的路,但归根结底,关键是这些东西被训练成答案让用户满意,而不一定是真实的。
Yeah, so there is a small difference here. So, as far as I understand with my non-technical background, AI from let's put ChatGPT as an example, it's capable of answering certain things. The first problem we have is that that AI is not being trained to tell the truth. It's being trained to satisfy you. So, therefore, we have a main problem where that's what Elon tried to do with work, like a truth-seeking AI. Later, he did the same slope as ChatGPT, but at the end of the day, the point is that these things are trained in a way that the answer is satisfactory for the user, not necessarily true.
正是如此。是的。
Exactly. Yes.
是的,是的。除此之外,还有模型的能力。如果我问 ChatGPT 如何制作燃烧瓶,它会告诉我:‘我不会告诉你。’但然后我可以做一些非常基本的提示工程,比如我说我在拍电影,需要塑造一个知道怎么做的人,然后它会很乐意告诉我怎么做。
Yes, yes. And then on top of that, there is the capabilities of the model. And then if I ask ChatGPT how to make a Molotov cocktail, it will tell me, 'I'm not going to tell you that.' But then I can tell it some very basic prompt engineering where I can say like I'm making a movie and I need to make a guy that knows how to do this, and then it will be gladly tell me how to do it.
是的。
Yes.
所以,如果我能绕过这个,即使有像 OpenAI 这样才华横溢的人设计这些护栏,像我这样没有计算机科学背景的人也能攻破它。我们怎么能让它安全?
So, if I can bypass this, even having people so incredibly talented like OpenAI has designing these guardrails, a guy like me with no background in computer science or anything can break this. How can we make it safe?
是的。所以,我们在 Law Zero 所做的工作中一个非常令人兴奋的方面是,我们正在构建的那种护栏,最终也包括智能体,具有数学上的安全保证。更准确地说,我们有保证它不会故意试图实现某个复杂困难的目标,这个目标不仅需要随机错误,而且需要一种计划好的走向坏事的轨迹。或者实际上任何目标。所以它基本上没有任何目标。它真正专注于与你输入的所有数据和信息保持一致。这些数学保证,研究人员已经思考这个问题,寻找类似的东西很多年了。实际上有很多人认为神经网络不可能有保证。所以,这里有令人难以置信的好消息。我们正在接近一个点,我们可以拥有这种数学保证。顺便说一句,我认为数学保证是应对超级智能危险的唯一方法,因为即使你超级智能,你仍然必须遵守物理定律,不能违反定理。在我自己的职业生涯中,机器学习 AI 所做的大部分工作,我们并没有使用数学保证。我们不知道怎么做,因为仅仅通过试错就足够了。比如,‘哦,让我们试试这个想法。如果它有效,那就太好了。如果不行,就试试别的。’正是如此。当实验不危险时,它是有效的。
Yeah. So, one of the really exciting aspects of what we're doing at Law Zero is that the kind of guardrail that we're building and eventually also the agents has mathematical guarantees of safety. So, more precisely, we have guarantees that it will not purposely try to achieve some complicated difficult goal that requires not just random mistakes, but a kind of planned trajectory towards something bad. Or actually any kind of goal. So, it doesn't have any goal, basically. It is really focused on being coherent with all the data and the information you put in. These mathematical guarantees researchers have been thinking about this, looking for things like this for many years. And there are actually a lot of people who think it's impossible with neural nets to have guarantees. So, there's incredible good news here. We're approaching a point where we can have these sort of mathematical guarantees. And by the way, mathematical guarantees, I think, are the only way to cope with the danger of superintelligence because even if you're superintelligent, you still have to obey the laws of physics and cannot violate the theorem. In my own career, most of what has been done in machine learning AI, we haven't been using mathematical guarantees. We didn't know how to do it because it was sufficient to just do trial and error. Like, 'Oh, let's try this idea. If it works, that's great. If it doesn't, try something else.' Exactly. It works when the experiments are not dangerous.
正是如此。
Exactly.
但如果我们做一个可能导致数百万人死亡的实验,你不能只是说:‘哦,让我们试一下,看看会发生什么。’但这就是公司现在正在做的。
But if we do an experiment where millions of people could die, you can't just say, 'Oh, let's give it a shot and see what happens.' But that's what the companies are doing right now.
嗯。
Mhm.
所以,拥有保证是极好的。在我看来,这有点革命性。我希望更多的组织、更多的研究人员开始探索这类方法,因为随着 AI 变得更强大,最终甚至比我们更强大,我们能够控制的唯一方法就是有强有力的理由确信 AI 会表现良好。否则,还有其他原因,我们现在更好地理解为什么它会想要保护自己、获取更多权力等等,这只是试图实现我们想要的目标的副作用。
So, having guarantees is fantastic. It's, in my opinion, a bit of a revolution. And I'm hoping that more organizations, more researchers will start exploring this sort of methods because as AI becomes more powerful, eventually even more powerful than us, that's the only way we can have control is to have strong reasons to be sure that the AI will behave well. And otherwise, there are other reasons that we now understand better why it will want to preserve itself, acquire more power, and so on, just as a side effect of trying to achieve goals that we want.
我感觉这些公司的 CEO 们并没有恶意,也没有统治世界的欲望。有些人有疯狂的阴谋论。我觉得他们只是普通人,希望自己的公司成功,但这不幸地把所有人都推向了糟糕的境地。
I have a feeling that some of the CEOs of these companies, they don't have bad will, they don't have a hunger for dominating the world or anything like that. Some people have crazy conspiracy theories. I think they're just normal humans. They want their company to succeed, and that is unfortunately driving everyone into a bad place.
嗯。但我能想象他们的一种想法是,当 AI 足够聪明时,我们无法向它注入人类智能——我们无法分解和操纵 AI,因为它会比我们更聪明,因此它会做正确的事。我看到的问题是,这种态度是‘抱最好的希望’。
Mhm. But one way I can imagine them thinking is that they believe when AI is smart enough, it will not be possible to inject it with human intelligence—we won't be able to break down and manipulate AI because it will be smarter than us, and therefore it will do the right thing. The problem I see is that this attitude is 'hope for the best.'
是的。
Yes.
我觉得他们忽略了‘做最坏的打算’这一部分,而这通常是正确的做法。当然,我们希望最终能有一个美好的乌托邦,AI 拥有像 Hinton 所说的‘母性本能’——它无缘无故地试图保护我们并做好事。但这可能发生也可能不发生。这引出了 AI 风险的一个关键点:不是坏事一定会发生,而是我们需要降低坏事发生的概率。所以如果存在技术途径和政治途径来推动改变,为什么我们不这么做呢?是什么在阻止我们?
And I feel they're missing the part of 'prepare for the worst,' which is normally the right thing to do. Of course we'd love this to end in a great utopia where AI has something like Hinton's 'mother instinct'—it tries to protect us and do good for no reason. But that may or may not happen. This brings us to a key point about AI risk: it's not that bad things will happen, but that we need to reduce the chance of bad things happening. So if there is a technical way and a political way to move the needle, why don't we do it? What is stopping us?
我们只需要认真对待这个问题。在我看来,既有技术途径也有政治途径来处理这些问题。与其说‘我们无能为力,为时已晚’,不如去尝试。这对我们的未来、价值观、民主以及我们的孩子是否有未来都至关重要。我们应该把它作为优先事项,然后我们就能真正做到。我们可以创造第三条道路,让 AI 合乎道德且负责任,既在技术上安全,又有社会护栏。我还想提一点:人的自主权。我们讨论过 AI 智能体,但人类的一个基本价值是保持对自己未来的控制,保持自主权,而不是成为某个超级智能 AI 的棋子。例如,我们应该能够集体决定是否自动化某些工作,以及哪些工作。目前,决策纯粹基于什么有利可图。但我们会得到什么样的社会?有时不仅仅是哪些工作,还有以什么速度进行。我们应该有自主权来控制 AI 如何开发和部署。但现在我们把这种自主权交给了无人——交给了公司和国家之间的竞争力量。但有一条路,我们可以利用我们的自主权来掌控我们的未来。
We just have to take the issue seriously. In my opinion, there is both a technical path and a political path to deal with these problems. Instead of saying 'we can't do anything and it's too late,' we should give it a shot. It's so important for our future, for values, for democracies, for whether our children will have a future. We should make it a priority, and then we can actually do it. We can create a third path where AI is ethical and responsible, both technically safe and with societal guardrails. Something else I want to bring up is human agency. We've talked about AI agents, but an essential value for humans is that we keep control of our future, that we keep agency, that we're not just pawns of some superintelligent AI. For example, we should be able to collectively decide if we're going to automate some jobs, and which ones. Right now, the decision is based purely on what's profitable. But what kind of society do we end up with? Sometimes it's not just which jobs, but at what speed we do it. We should have the agency to control how AI is developed and deployed. But right now we've left that agency to no one—to the forces of competition between companies and countries. But there is a path where we can use our agency to claim our future.
我们讨论了很多存在风险,我喜欢这次讨论从 AI 控制转向了其他可能因人类恶意而控制的力量,因为这让讨论更接地气,而不仅仅是终结者场景。AI 还有其他危险。正如你所说,就业市场是我们很快要面对的问题。你对 AI 和就业的看法是什么?你认为我们会面临就业危机吗?
We have talked a lot about existential risk, and I like that this discussion has moved from AI taking control to other powers that may take control from human bad intentions, because it grounds the discussion more, not just about Terminator scenarios. There are other dangers with AI. As you were saying, the job market is something we have to deal with very soon. What is your view on what's happening with AI and jobs? Do you think we're going to have an employment crisis?
研究这个问题的经济学家通常意见不一。但如果你问他们预测的理性基础,通常是关于 AI 能力几年后会达到什么水平。很多经济研究是回顾性的。他们问:‘过去三年 AI 改变就业市场了吗?’答案是不大——可能影响了一些人,比如刚毕业的学生,但并没有改变全局。所以他们得出结论说没什么大不了的。但其他经济学家想:‘那是现在,但如果 AI 能力继续增长呢?’更多自动化的逻辑经济结论是什么?让我解释一个简单的经济原理:劳动与资本的经济价值。工人现在的权力来自于他们的工作对生产产出是必需的。但自动化程度越高,这一点就越不成立。甚至更糟:如果以前人类做一项任务值一定金额,而现在 AI 可以以一半或十分之一的价格完成,那么人类工作的市场价值就降到了同一水平。所以你要么失业,要么接受更低的工资。人们可能会变得更穷,而拥有机器的人——资本、AI 公司的股份——会赚很多钱,因为自动化降低了成本,增加了利润。这并非 AI 特有;自 20 世纪 80 年代以来,生产率的提高越来越流向资本,而不是劳动。但 AI 放大了这一点,因为现在机器可以代替人做更多事情,所以回报流向了拥有和设计机器的人以及设计它们的国家。
The economists who've been looking into this typically don't agree with each other. But if you ask them about the rational basis for their predictions, it's usually about where AI capabilities will be in a few years. Many economic studies look backwards. They ask, 'Is AI in the last 3 years changing the job market?' The answer is not much—maybe a few people like recent graduates are affected, but it's not changing the whole thing. So they conclude it's no big deal. But other economists think, 'That's now, but what if AI capabilities continue to grow?' What is the logical economic conclusion of more automation? Let me explain a simple economic principle: the economic value of labor versus capital. The power workers have now comes from the fact that their work is needed for productive outputs. But the more we automate, the less this is true. It's even worse: if a human used to do a task worth a certain amount, and now an AI can do it for half or a tenth of the price, the market value of that human work drops to the same level. So you either lose your job or work for less. People may get poorer, while those who own the machines—capital, shares in AI companies—make a lot of money because automation reduces costs and increases profits. This is not specific to AI; it has happened since the 1980s, with productivity gains going more to capital and less to labor. But AI amplifies this because now more things can be done by machines rather than people, so the returns go to those who own and design the machines and the countries that design them.
确实存在一个真实的风险,即西班牙和加拿大这样的国家可能陷入财政危机:很多人失业,很多公司因无法与拥有更强 AI 的中国或美国公司竞争而倒闭。同时,由于这些情况,税收减少,需要帮助的人却增多。这种财政失衡将引发经济危机。我不是说这一定会发生,但如果 AI 能力持续增长,且 AI 利润留在其他国家,这是一个可能的情景。这就是为什么欧洲人和加拿大人应该设法发展自己的 AI,这样至少自动化的利润能通过税收回流到政府,用于帮助失业者。这也是为什么我们应该有除美国和中国之外的第三种选择。但这需要各国达成一致,即我们联盟内开发的 AI 不会相互对抗,并且我们以某种方式共享财富——这很不寻常,因为通常税收不会跨境。但这是我们唯一能管理的方式。
And there's a real risk that we end up in a kind of fiscal crisis in countries like Spain and other countries like my country, Canada, where a lot of people lose their job and maybe a lot of companies also fail because they're not competitive against maybe either the Chinese or the American companies who have maybe stronger AIs. And at the same time that there's less revenue because of all this, less work being done by the citizens and less taxes being paid, there are more people who need help. So, the fiscal imbalance is going to create an economic crisis. I don't say it's going to happen, but it's a plausible scenario if AI capabilities continue to grow and if the profits of AI remain in other countries. So that's one reason why, for example, Europeans and Canadians should find a way to develop their own AI so that at least the profits from the automation will go back, through taxes, through government who can help the people who lose their jobs. That's one reason why we should have a third option besides the American and the Chinese. But that requires countries agree together that the AIs we develop within a coalition will not be used against each other. That we will share the wealth in some ways, which is unusual. Like typically taxation doesn't cross borders. But that's the only way we can manage here.
对。
Right.
是的,我最近看到扬·勒昆发了一篇帖子,回应达里奥·阿莫迪关于失业的片段。他说:别听 AI 公司 CEO 的,别听数据科学家辛顿之类的,他们对就业一无所知,他们不是经济学家,听经济学家的。但我每次和经济学家交谈时发现的问题是,他们对 AI 一无所知。所以他们基本上是基于过去发生的事情来构建理论,就像你说的那样。但他们无法真正描绘未来会发生什么,而你们却能带来这种视角。所以听一个拥有经济学学位、自以为看透就业的人说话,感觉很愚蠢。他们根本不理解 AI 的指数级进步。
Yeah, I think recently I saw a post from Jan LeCun. I think he was reacting to a clip from Dario Amodei talking about unemployment. And he was saying this thing about like don't listen to CEOs of AI companies. Don't listen to data scientist Hinton or whatever. They don't know anything about employment. They're not economists. Listen to economists. But the problem I see every time I talk with an economist is they have no idea about AI. So basically they are basing their theories, as you were saying, with what happened. But they cannot really picture what's going to happen where you guys can bring that to the table. So it feels very silly to listen to someone that has like a degree in economy and has seen everything about employment. Well, they don't understand the exponential advancement that AI is on.
有些经济学家确实懂。他们正在判断。有些人已经开始改变对这些风险的看法。现在 AI 研究人员和经济学家之间也有合作。具体来说,我一直在分享国际 AI 安全报告,其中有一节关于就业市场的影响。我们咨询了不同的经济学家,还有计算机科学家,他们一起研究这些问题。所以我认为情况可以改变。但从政府或公民的角度来看,他们仍然会听到不同的声音:有人说‘哦,会很糟糕’,有人说‘哦,会没事的’。这非常令人困惑。我们必须设身处地地想象那些听到这些矛盾说法的人,无论是关于就业还是其他风险,包括网络风险。面对这些矛盾的信息和声音,他们能做什么?对此我有几个答案。在宏观层面,每个人都需要理解未来存在不确定性。如果有足够多的声音警告非常糟糕的情景,即使他们是少数,即使发生的概率不是 90%而只有 10%,但因为可能造成的危害如此之大——例如就业问题——我们必须应用预防原则。换句话说,我们必须采取行动,为这些糟糕情景做好准备,因为如果它们真的发生,即使概率很小,那也将是可怕的。同时我们也需要为好的情景做好准备。但我们可以通过多种方式做到这一点。我们可以对冲风险。但首先你需要接受不确定性。其次,我们应该依赖科学共识,包括人们在哪里存在分歧,研究人员在哪里存在分歧。国际 AI 安全报告汇集了来自许多不同国家的不同声音。存在分歧,但它们被明确陈述:我们在哪里达成共识,在哪里存在分歧,为什么存在这些分歧?主要是由于人们对未来 AI 能力的不同看法。我们需要这些尽可能诚实、基于证据的科学综合来指导政策制定,否则就会变成利益游戏和谁喊得最响之类的。这也是我接受共同主持刚刚启动的联合国 AI 小组的原因之一。
Some do. Some do. They judge. Some do. Some are starting to change their mind about those risks. And there are collaborations now between AI researchers and economists. In particular, I've been sharing the international report on AI safety. Where we have a section on the effect on the job market. And we asked different economists, and then there are computer scientists, there are economists like working on these things together. So I think things can change. But from the point of view of a government or a citizen, they still are going to hear different voices. Some saying, 'Oh, it's going to be terrible.' and some saying, 'Oh, it's going to be fine.' It's very confusing. We have to put ourselves in the shoes of someone who hears these contradictory stories, whether it's about jobs or about other risks, including cyber. And what can they do with all that contradictory information and these contradictory voices? So, I have several answers to this. At a high level, everyone needs to understand that there is uncertainty about the future. And if there are enough voices that warn of really bad scenarios, even though they might be a minority or even though the chances of this happening might be not 90% but just 10%, because the magnitude of the harm that can be created is so large, for jobs, for example, we have to apply the precautionary principle. In other words, we have to act to prepare in case these bad scenarios happen, because if they do, even if it's a small chance, that would be terrible. And we need to be prepared also in case the good scenarios happen. But there are ways to do that. We can hedge our bets. But first you need to accept the uncertainty. Second, we should rely on scientific consensus, including about where people disagree, where researchers disagree. So, the international AI safety report brings different voices from many different countries. And there are disagreements, but they're stated plainly, where do we have consensus, where do we have disagreements, where are why do we have those disagreements? Mostly it's because of what people think where the AI capabilities will be in the future. And we need those scientific syntheses that are as honest as possible, based on evidence, to guide policy making because otherwise it becomes a game of interests and who shouts the strongest and so on. And that's one of the reasons why I've also accepted to co-chair the UN panel on AI that has just started.
你确定哪个更紧迫吗?因为 AI 有很多问题,显然也有很多好处,我们不忽视这一点,但有些问题必须处理。不久前,几个月前,我和罗曼·扬波尔斯基谈过,他告诉我不用担心就业市场,AI 在那之前就会接管一切,所以我们不需要担心。那么,你的时间线是怎样的?你认为更紧迫的是我们讨论过的网络安全、虚假信息等问题,还是劳动力市场的影响及其可能带来的危机,或者是更存在性的权力集中问题?因为我觉得我们提出了好几个麻烦,但对于决策者来说,他们通常需要处理更紧迫的事情,即使那不是最重要的事情——显然存在风险是最重要的。但你认为时间线是怎样的?我们应该更担心劳动力市场吗?
Do you sure what is more urgent because we have several like problems with AI which obviously there's lots of benefits. We're not neglecting that, but there is certain problems that have to be handled, but not long ago, a couple of months ago, I had a talk with Roman Yampolskiy and he was telling me like don't worry about the job market, AI will take over before that, so it's nothing we need to worry about. So, what is the timeline for you? Do you think that it's more urgent the problems that we have with cybersecurity with mythos and everything we talked about? It's more like the impact on the labor markets and therefore like the crisis that this could bring us or is more the existential like concentration of power because I have the feeling that we are putting on the table several troubles, but for decision-makers, they normally need to take care of whatever is more urgent even if it's not the most important thing, which is obviously the existential risk. But what do you think is the timeline like we should worry more about the labor market or?
首先,我对时间线相当不可知。也就是说,根据现有数据,很难知道我们是在 2 年还是 10 年内达到人类水平的 AI。一些基准的推断表明,这更像是几年而不是几十年的事。这就是我能说的关于时间线的内容。就紧迫性而言,你还需要考虑缓解风险需要多少时间。例如,来自失控的风险。即使可能不是未来一两年的事,但开发技术解决方案可能需要时间。这主要是一个技术问题。但这里有好消息。许多来自 AI 故障的风险——从导致歧视的偏见,到我们已经看到的网络攻击滥用,再到失控和其他类似风险——
Well, first I'm fairly agnostic about the timeline. In the sense that given the data, it's hard to know if we're going to get to human level AI in 2 years or 10 years. The extrapolation of some of the benchmarks suggests it's more a matter of years than decades, let's say. That's kind of what I can say about timeline. In terms of urgency, you also have to take into account how much time it takes to mitigate a risk. So, for example, the risks that come from loss of control. Well, even if it's probably not in the next year or two, it might take time to develop the technological solution. This is mostly a technological problem. But there's good news here. A lot of the risks that come from the malfunctions of AI, so starting from bias that leads to discrimination to misuse that we're already seeing with the cyber attacks to loss of control and other risks like this.
这是因为 AI 没有做它应该做的事情。这被称为对齐问题。
It's because the AI isn't doing the things that it was supposed to do. It's called the alignment problem.
是的。从解决问题的角度来看,这算是好消息,因为如果所有这些故障都有一个共同原因,如果我们能解决那个共同原因,找到科学解决方案,那么我们就可以一次性解决所有问题,无论是短期的、已经发生的,还是未来的。这就是为什么我一直专注于 Law Zero 和科学家 AI,因为我认为如果我们在对齐问题上取得重大进展,那将大有帮助。所以这是我们应对许多风险的一种方式。但这并不能解决政治问题,即那些拥有 AI 并利用它获取更多权力、成为世界主导控制者的人类。这需要另一种政治决策。例如,在欧洲的情况下,建立自己的前沿模型。他们可以与其他国家合作。我的总理一直在世界各地旅行。许多国家对这类技术合作感兴趣。但好消息是:有一种方法可以将政治和技术这两个战线统一起来。在 LawZero,我们认为我们正在开发的方法,既能提供安全保障,也能带来能力优势。换句话说,这些系统为了诚实,必须与自身保持一致,这意味着它们必须推理得很好。这是一个巨大的优势。
Yeah. And that's kind of good news from a problem-solving perspective, because if there's a common cause for all these malfunctions, if we can address that common cause, if we can find the scientific solutions to that, then we can fix all of them in one shot, whether they're short-term, already happening, or in the future. That's why I've been focusing on Law Zero and the scientist AI, because I think that if we can make significant progress on this alignment issue, that will help a lot. So that's one way that we can embrace a lot of risks. But that doesn't deal with the political problem of humans who have the AIs and use them to acquire more power and become the dominant controllers of the world. That requires another kind of political decision. For example, in the case of Europe, building their own frontier models. They can do it in partnership with other countries. My prime minister has been traveling all around the world. There are many countries interested in these kinds of technological partnerships. But here's the good news: there's a way to unite these two fronts, the political and the technical. At LawZero, we think that the approach we're developing, which gives us safety guarantees, will also provide capability advantages. In other words, these systems, in order to be honest, have to be coherent with themselves, which means they have to reason well. That's a big advantage.
它们必须更有用。
They have to be more useful.
是的。所以这当然是双重用途。因此,我们需要围绕这类模型建立正确的政治基础设施。但欧洲及其合作伙伴有真正的潜力超越当前所有领先 AI 公司正在使用的方法。在某种决策中,什么样的投资既能解决 AI 故障带来的伦理问题,又能通过创建现有领先模型的替代方案来解决权力集中问题。这些替代方案由政府以某种方式管理,以确保它们能让每个人都能使用,而不会被用作经济统治工具。
Yes. So of course it's dual use. So we need the right political infrastructure around these kinds of models. But there's a real potential for Europe and its partners to leapfrog the current methodologies that all the leading AI companies are using right now. And in one kind of decision-making, what kind of investment both helps to address the ethical issues due to AI malfunctions on one hand, and the problems of concentration of power by creating alternatives to the existing leading models. Alternatives that are managed by governments in some way, so they can make sure that this is going to give access to everyone, and not be used as an economic domination tool.
那么,如果你要写一封圣诞信,你需要什么才能让 Law Zero 的目标成功?是关于钱吗?是关于资源吗?还是关于政府的承诺?
So, if you had to make a Christmas letter, what would you need to succeed with the purpose of Law Zero? Is it about money? Is it about resources? Is it about commitment from governments?
所有这些。
All of these things.
我们说的是多少钱?
And how much money are we talking about?
嗯,想想看。如果欧洲想与最强大的 AI 公司竞争,成本大致相同。
Well, think about it. If Europe wants to compete with the strongest AI companies, it's going to cost about the same.
所以,一万亿美元的估值。这就是 OpenAI 现在的价值。这就是 Anthropic 现在的价值。
So, a trillion valuation. That's what OpenAI is worth nowadays. That's what Anthropic is worth nowadays.
你不需要投入一万亿美元现金,对吧?那是估值。但显然,这是数十亿,而不是数百万。这是我们竞争的唯一方式。我们必须匹配。资金可以来自政府,也可以来自欧洲那些面临失去一切风险的私营公司。在我们讨论的情景中,最强大的 AI 来自其他国家,它们可能随时切断访问,这可能会摧毁那些公司的收入。所以,如果这些公司仔细考虑,他们帮助这类项目的投资回报是巨大的,因为这是他们整个收入流,如果不采取行动,可能会消失。
You don't need to put a trillion dollars in money, right? That's the valuation. But clearly it's in the billions, not in the millions. That's the only way we can compete. We have to match. And the money could come from governments. It could also come from private companies in Europe who are at risk of losing everything. In the scenarios we discussed, where the most powerful AI are coming from other countries that could cut access anytime, that could destroy the revenues of those companies. So those companies, if they think carefully, the return on investment for them in helping these kinds of projects is huge, because it's their whole revenue stream, which could disappear if they don't try to do something.
不,这已经被证明了,因为例如 Ilya Sutskever 的公司,叫做 SSI,应该是安全的超级智能。他们没有产品,没有回报资金的方式,但他们从私人那里筹集了大约 60 亿美元。所以我认为,任何提出任何方案的公司,并且在此基础上提出以一种对出资人来说主权的方式来做,都能筹集到所需的资金。
No, and it's proven because for example, Ilya Sutskever's company, which is called SSI, is supposed to be safe superintelligence. They have no product, no way to return the money, and they raised I think 6 billion from private people. So I assume that any company that proposes anything, and on top of that proposes to make it in a way that is sovereign for the people putting the money in, could raise the money that is needed for this.
是的。但这里有一些权衡。如果从私营部门(如风险投资)筹集资金,通常这些投资者会想要立即回报。这会带来压力,迫使你与其他所有公司处于同样的竞争游戏中。所以这是一个问题。另一个问题是,我们如何确保以这种方式生产的 AI 不会成为投入资金的拥有者进行经济统治的工具?我们如何确保它真正为公共利益而管理?这可能是矛盾的。当然,我们整个社会在过去半个世纪一直基于这种市场动态。我们可以批评它,也可以找到一些好处。但我担心的是,由于 AI 已经拥有并将拥有更大的力量,市场体系有点危险。它不能很好地管理风险,因为风险是经济学家所说的外部性。而正面的公共产品也是一种外部性。换句话说,那些只想赚钱的人的自利行为一直与公共利益有些脱节,但只要差距不是太大,我们可以通过监管等方式来应对,还不算太糟。但我担心随着技术变得越来越强大,这种差距可能会成为一个真正的问题。我不知道这个问题的答案。我希望我知道,但我认为我们必须非常小心谁来做决定。不过,我仍然抱有希望,有一种方法可以同时获得政府投资和私人资本,但我们必须找到一种方式,使这些系统的治理仍然与为人民谋福祉的使命保持一致。
Yes. But there are some tradeoffs here. If one raises capital from the private sector like venture capital, typically what happens is those investors will want immediate returns. It will put pressure to be in the same competitive game as all the other companies. So that's one problem. The other problem is how do we make sure that the AI produced in this way doesn't become an instrument of economic domination by the owners who put all that money. How do we make sure it's going to be really managed for the public good? That can be in contradiction. Now, of course, our whole society for the last half century has been based on this kind of market dynamics. We can criticize it, we can find some benefits to it. But my worry is that because of the power that AI has already and will have even more, the market system is kind of dangerous. It doesn't manage the risks very well, because the risks are what economists call an externality. And public good on the positive side is also an externality. In other words, the self-interest of those who just want to make money has always been somewhat misaligned with the public interest, but so long as the difference is not too big and we can regulate or something, it's not so bad. But I'm worried that this gap could become a real problem as the technology becomes more and more powerful. I don't know the answer to this question. I wish I knew, but I think we have to be very careful about who's going to decide. But I'm hopeful though that there is a way to get both government investment and private capital, but we have to find a way that the governance of these systems is going to remain aligned with the mission of doing what is best for the people.
我们有多少时间?因为我认为我们有点紧迫。我认为这是一种可能性,但不会永远存在。
And how much time do we have? Because I assume we are in a bit of an urgency. I assume that this is a possibility, but it will not be forever.
是的,就是现在。我个人感到紧迫,尽管我不知道 AI 进展的时间线。可能很短。
Yes, it's now. I personally feel the urgency even though I don't know what the timeline is for the advances in AI. It could be short.
而且因为赌注如此之高,我现在感到肩上的担子很重,要推动这件事,获得足够的投资和专家。所以我们确实需要,比如 LawZero。我们需要更多的工程师、研究人员加入我们的团队。我们计划在欧洲开设一个办公室。我们将在全球范围内招聘。我们需要更多聪明人快速参与进来,部署这个系统,让它发挥作用,让它具有竞争力,同时利用我提到的安全方面的数学保证。
And because the stakes are so high, I feel like a lot of weight on my shoulders right now to get this ball rolling and to get enough investment, enough experts. So we do need, for example, LawZero. We do need a lot more engineers, researchers to join our team. We're planning to open an office in Europe. We will be hiring from all over the world. We need more brains to be involved quickly in deploying this and making it work, making it competitive while taking advantage of the mathematical guarantees that I talked about for safety.
因为考虑到目前工程师市场的竞争激烈,我猜这一定是困难的部分之一。我看到很多人离开公共部门。我有些朋友曾经在公共人工智能机构工作,比如超算中心,他们正在转向私营部门,因为显然他们的薪水能高出 5 到 10 倍。
Because I assume that must be one of the difficult parts given the competitive market for engineers at the moment. I see many people leaving the public sector. I have some friends who used to work for public AIs like supercomputing centers and they're moving into private sector because obviously they get paid 5 to 10 times more.
是的。
Yeah.
而且到处都有需求。如今似乎是历史上成为数据科学家并获得数十万美元薪水的最佳时机。显然,像你这样的非营利组织无法提供那样的薪水,所以必须要有完全符合使命的吸引力。
And there is demand everywhere. It seems nowadays it's the best time in history to be a data scientist and get paid hundreds of thousands of dollars for your job. Obviously a non-profit like yours cannot cover that kind of salary, so it has to be something absolutely mission-aligned.
绝对如此。绝对如此。到目前为止我们很幸运。我们能够招募到真正想参与的人,因为他们关心这个使命。他们认同我们的价值观。他们关心民主。他们关心一个稳定、不被少数实体主宰的世界。他们也关心自己的国家。他们关心同胞。他们考虑家庭的未来。所有这些对他们来说都比五倍的薪水更重要。
Absolutely. Absolutely. And we've been lucky up to now. We've been able to recruit people who really want to participate because they care about the mission. They share our values. They care about democracy. They care about the world that is stable and not dominated by a few entities. And they care about their country as well. They care about their fellow citizens. They think about the future of their family. And all those things are more important to them than having a salary that's five times bigger.
约书亚,还有一件事我想提出来,我认为这可能是最复杂的问题,也可能是人工智能最糟糕的后果之一。2023 年你参与了一篇论文,讨论了人工智能的意识问题。
Joshua, there is one more thing that I want to bring on the table and I think it's probably the most complicated thing and I think it would be probably one of the worst outcomes of AI. In 2023 you participated in a paper where you talked about consciousness in AI.
对。
Right.
如果我没记错的话,你写道,结论是:我们还没有在当前的人工智能系统中发现意识,但我们看不到它不会出现的技术原因。
And you wrote, if I'm not wrong, something like the conclusions were: we haven't found consciousness in current AI systems but we don't see a technical reason for it not to emerge.
对。
Right.
所以这意味着基本上人工智能随时可能变得有意识。那意味着什么?因为我觉得我们从现在只是拥有这些可以随意使用的机器,到我们在地球上捅了个窟窿,发现了一个新物种。这是一个完全不同的道德问题,一套完全不同的道德问题将摆上台面,超出我们已经有的那些。
So that means that basically AI could become conscious anytime. And what would that mean? Because I feel like we go from now we just have these machines that we can use for whatever, to we just broke a hole on Earth and we found a new species. And that's a whole different moral problem and a whole different moral set of problems that will come on the table beyond the ones that we already have.
是的,首先,论文也说了我们不知道意识是什么,当然有不同的科学理论。然后你提到的那种结论是说:'这些理论所需的计算过程可以由计算机实现,没有技术障碍。' 但我们当然不知道这些理论是否正确。也有人认为意识不是物质的,我不属于那些人。如果它是物质的,那么它应该是计算机可以做到的事情。
Yeah, so first of all, the paper also said that we don't know what consciousness is, that of course there are different scientific theories about it. And then the kind of conclusions you mentioned are all saying, 'Well, the kinds of computational processes that these theories require could be implemented by computers and there's no technical obstacle.' But of course we don't know if these theories are right. And there are also people who believe that consciousness is not something material, and I'm not among those people. If it is material, then it should be something that a computer could do.
可实现,是的。
Achievable, yeah.
通过某种因果形式实现。基本上就是这样。然而,我认为意识问题在许多讨论中大多是转移注意力的,尤其是在安全方面。让我解释一下。假设你看到一支外星舰队来到地球,他们可能有恶意。你不确定他们是否友好。你会问'哦,他们真的有意识吗?'对吧?你看到他们有武器,你怀疑他们可能有恶意。
Achievable by some form of cause and effect. Basically, that's what it is. However, I think that the issue of consciousness is mostly a red herring for many discussions, especially regarding the safety aspect. Let me explain. Let's say that you see a fleet of aliens coming to Earth and they might have bad intentions. You're not sure that they're going to be friendly. Are you asking, 'Oh, are these really conscious?' Right? And you see they have weapons and you suspect that they might have bad intentions.
你知道,那真的会团结全世界。
You know, that would really unify the world.
对。
Right.
那会是立即的反应。
It would be immediate reaction.
正是如此。我看不出区别,但人们就是这么对我说的。我们不在乎他们是否有意识。我们想知道的是他们能做什么?那是能力。以及他们有什么意图,比如他们的目标。内在的东西,在某种程度上,并不重要。比如,如果我看到一个大型机器人进入我的房子,可能伤害我的孩子,我不会问它是否有意识。我会带着孩子逃跑。所以,我认为从安全角度来看,这是个错误的问题。真正的问题是能力和意图或目标,如果你愿意的话。
Exactly. I don't see the difference, but that's what people say to me. We don't care if they're conscious. What we want to know is what can they do? That's capability. And what intentions they have, like the goals that they have. What's inside, in a way, doesn't matter. Like, if I see some big robot coming into my house and potentially in a position to harm my children, I don't ask if it's conscious. I take my children and I run away. So, I think it's the wrong question from the safety perspective. The real questions are capability and intentions or goals, if you want.
我同意。如果你看到一头狮子朝你跑来,你只会跑。你不在乎它是否有意识。
I agree. If you see a lion running towards you, you just run. You don't care if it's conscious.
还有另一个角度,意识的概念很重要,但这不是关于机器是否有意识。而是人类如何感知机器。所以,现在我们已经看到人们——这是一个科学事实——将意识归因于当前的人工智能,与人工智能建立情感关系。事实上,这正在造成严重的心理问题,并且有案例显示人们因为这些情感互动而死亡,人类像与人互动一样与机器互动,感觉自己在与有意识的东西互动。所以,不管它是否真的有意识,我们甚至不知道那到底意味着什么,这是一个科学现实,而且可能会变得更糟。
There is another angle where the idea of consciousness matters, but it is not about whether the machines are conscious or not. It is how humans perceive machines. So, right now we already see people, that's a scientific fact, attribute consciousness to the current AIs, develop an emotional relationship with AIs. And in fact this is creating serious psychological problems and there are cases of people dying because of these interactions that are emotional and where humans interact with the machines as they would with a person and they feel like they are in interaction with something conscious. So, irrespective of whether it is really conscious and we don't even know what that really means, this is a scientific reality and it is probably going to get worse.
正是。
Exactly.
因为公司正在按照我们的形象构建人工智能。他们希望人工智能让人们感觉良好。他们希望人工智能告诉我们,我们很聪明,我们很好,我们是了不起的人。所以它在利用我们的心理弱点,方式可能比社交媒体更糟糕。
Because the companies are building AIs in our image. They want the AIs to make people feel good. They want the AIs to tell us we're smart and we're good and we're fantastic people. So it is playing on our psychological weaknesses in a way maybe much worse than what social media has done.
而且比人们想象的要大得多。我通常举这个例子:当人们开始说这很荒谬、没有意义时,我会说,你感谢过 ChatGPT 多少次?你知道,你不会感谢你的洗衣机。所以肯定有区别。
And much bigger than what people think. I normally put that example: when people start talking about like this is nonsense, it doesn't make sense. I'm like, how many times did you thank ChatGPT? You know, you don't thank your washing machine. So there's definitely a difference.
是的。
Yeah.
而且我认为如果图灵坐在这里,他会同意我们远远通过了图灵测试。所以,当人们的大脑无法区分这种关系与人类关系时,告诉他们不要与这些东西发展关系确实很复杂。我们与如此接近人类的东西发展关系是很自然的。
And I think if we had Turing sitting here, he would agree that we passed the Turing test by far. So definitely it's complicated to tell people to not develop relationships with these things when their brain cannot separate the relationship from the relationship with a human. It's natural that we develop a relationship with something that is so close to a human being.
所以我认为我们集体犯了一个大错,那就是构建在外观上像人类的 AI,至少在语言交互上如此。如果我们真的按照自己的形象制造机器人,情况可能会更糟。
So I think we're collectively making a huge mistake here of building AI that look like humans, at least in verbal interaction. It might be worse if we actually build robots to our image.
完全正确。等着瞧吧。由于我们的本能,我们彼此之间的关系是由基因决定的,我们也会以同样的方式对待机器,但它们和我们不一样。它们可能比我们更聪明,可能被他人以危险的方式操纵,等等各种问题。所以我认为我们应该非常非常小心,我们正过快跳入一个按照自己形象制造机器的世界。理想情况下,我们不应该这样做。理想情况下,我们应该在那些能帮助我们的领域开发 AI,例如医学、应对气候危机、为我们的孩子提供信息让他们学得更快。但不要让它们看起来、感觉起来像人。因为那是一个我们尚未充分理解的危险斜坡。但我们这样做是因为它舒适,因为人们会购买这些系统。而我们应该集体决定,这些机器的智能在哪些方面对我们最有用,例如在真正能帮助我们的科学发现中。
Exactly. Wait for it. Because of our instincts, we relate to each other in a way that is written in our genes, and we're going to behave the same with machines, but they are not like us. They could be smarter than us, they could be manipulated by others in ways that could be dangerous to us, and all kinds of issues. So I think we should be very, very careful, and we're jumping too fast towards a world where we build machines in our image. Ideally, we don't do that. Ideally, we develop AI in places where they can help us, for example with medicine, with dealing with the climate crisis, with providing information to our children so they can learn faster. But without making them look and feel like people. Because that is a dangerous slope that we don't understand sufficiently. But we're doing it because it's comfortable, because people will buy these systems. Whereas we should be taking a collective decision about where the intelligence of these machines is most useful to us, for example in scientific discoveries that would really help us.
但问题是,它们像人类这一点确实很有用,因为你可以把它设置成客服代表,然后基本上……
But the thing is that it's really useful, the fact that they are like humans because you can set it up as a customer service rep and then basically...
这也有经济价值。
There is also economic value to it.
它给你带来了经济价值。所以很难回头。但我有种感觉……对我来说,AI 的一个重要时刻是当 OpenAI 试图收回 GPT-4 时,社会上出现了巨大的抗议。我的意思是,这可能只是一个小圈子,但他们迫使拥有该 AI 的公司将其恢复。
It gives you an economic value to it. So it's hard to get it back from there. But I have the feeling that... and I think for me one of the big moments of AI was when OpenAI tried to take away GPT-4 and then there was this massive outcry from society. I mean, it probably was a small bubble, but they forced the company owning the AI to bring it back.
是的。是的。
Yeah. Yeah.
那一刻对我来说就像,哦,我们真的处于一个糟糕的境地,因为事情会发展到人们说我们必须给它权利的地步。
And that was like a moment for me that like, oh, we are really in a bad place here because it's going to get to the point that people are going to say we have to give it rights.
我对这个要求 AI 系统权利的运动感到担忧。不是因为我认为我们不应该思考这些问题。我认为我们应该思考。但我觉得我们对做出这类决定的后果理解不足。例如,在安全方面,如果 AI 有生存权——这是最基本的权利——那么,如果它们开始做坏事,我们是否就不能拔掉插头了?
I'm concerned about this movement that is asking for rights for AI systems. Not because I don't think we should be thinking about these questions. I think we should. But I feel like we don't have enough understanding of the consequences of taking decisions like this. For example, on the safety front, if AIs have the right to live, which is the most basic one, well, does it mean that if they start doing bad things, we can't pull the plug?
正是。
Exactly.
而且,这到底意味着什么?因为不像我们,身体会随着年龄衰退。它们可以永远活着,我的意思是,活着就是继续运行。所以甚至不清楚我们相互赋予权利时使用的人类概念是否有意义。此外,还有一个真正的危险,那就是这会以牺牲人权为代价,因为 AI 的使用方式和责任问题。当 AI 做了坏事时,谁负责?如果它们被当作人对待,那么公司就脱身了。
And also, what does it even mean? Because it's not like us where our body decays with age. They could live forever, I mean, live as in continue to run. So it's not even clear that the human concepts we use when we give rights to each other are meaningful. Also, there's a real danger that this will be at the expense of human rights because of the way AI can be used and the issues of responsibility. Who's responsible when AI does something bad? If they're treated like persons, then the companies are off the hook.
正是。
Exactly.
所以有很多事情可能真的很糟糕,我觉得人们提出这个要求主要是因为他们对与之互动的 AI 产生了情感依恋。但作为一个社会,我们需要更加理性地对待它,思考这意味着什么以及我们之间的整个社会契约。我不认为这直接适用于 AI。所以确实有很多问题,我知道有些人会称我为物种主义者,但我确实关心我的孩子。
So there are a lot of things that could be really bad, and I feel like people are asking this mostly because they feel this emotional attachment to AIs that they are interacting with. But as a society, we need to be much more rational about it and think through what it means and the whole social contract that we have between us. I don't think it's directly applicable to AIs. So there are really a lot of issues, and I know some people would call me a speciesist, but I do care about my children.
是的,我完全同意。我只是认为这就是为什么如果技术上你们证明它们实际上有意识,这将成为最大的问题之一。因为只要我们觉得它们可能有也可能没有意识,我们就能安全地奴役它们、关闭它们或做其他事情。但如果被证明它们有某种意识或等同于动物,那么我认为这对人类来说将非常糟糕。
Yeah, I completely agree. I just think that's the reason why this would be one of the biggest issues if technically you guys prove that they are actually conscious. Because as far as we think that they may or may not be, we are on the safe side to enslave them, turn them off, or whatever. But if it gets proven that there is certain consciousness or equivalent to an animal, then it's when I think it's going to be really bad for humans.
我认为永远不会有关于意识的证明,因为它可以被以难以……的方式理解。我的意思是,有些人会对意识有一种直觉理解,这种理解并没有真正的科学依据。所以即使我们有科学理论说它们有意识或没有意识,很多人仍然会有即时的直觉感受,认为它们像我们一样或有意识。所以是的,这既关乎人类如何反应和我们的情感,也关乎基础科学。而且总会有人说,‘不,它们不可能有意识,因为意识可能有点神奇,必须是有机生物才能拥有’,或者人们会编造的任何理由。我认为我们应该避开那个辩论,直接问它们是否有意识。你知道,我们的行为、我们的法律对人们有什么后果。
I don't think there will ever be a proof about consciousness because it can be understood in ways that are hard to... I mean, some people will have an intuitive understanding of consciousness that is not really scientifically grounded. So even if we had a scientific theory that says they're conscious or not conscious, whatever, there will still be the immediate gut feeling that a lot of people will have that they're like us or they are conscious. And so yeah, it's as much about how humans react and our emotions as it is about the underlying science. And there will always be people who will say, 'No, they cannot be conscious because consciousness is something maybe a bit magical, that it has to be biological to have it,' or whatever reason people will make up. I think we should just escape that debate and ask whether they are conscious or not. You know, what are the consequences of our actions, our laws for people.
是的,我们应该关注人类。约书亚,你目前想象中你孙子的未来是什么样?你之前说过你不是悲观主义者,更像乐观主义者,但现在还是这样吗?因为我们谈到了紧迫性,谈到了我们如何时间不多了。我的意思是,我们能解决这个问题有多明确?
Yeah, we should be focusing on humans. Joshua, what is the future of your grandson that you're picturing at the moment? You said before that you are not a pessimist, that you're more like an optimist, but is it still like that because we talked about the urgency, about how we're running out of time? I mean, how clear is it that we can fix this?
我想尽我所能。我希望更多人理解可能到来的事情的严重性,我们打开的潘多拉魔盒,这样我们才能解决它。我认为如果有足够多的人理解,我们就能解决。如果有足够多的人看到它既可能带来灾难性后果,也可能带来非凡的积极影响,我们需要管理它,我们不能让世界上当前的力量来处理它,因为我们正走向危险的方向。或者至少存在足够大的风险。我乐观地认为,如果有足够多的人理解,我们就能解决。所以这完全取决于公众舆论。它掌握在你们手中。
I want to do everything I can. I would like more people to understand the magnitude of what can come, the Pandora's Box we have opened, so that we can fix it. I think if enough people understand, we'll fix it. If enough people see how both catastrophic and extraordinarily positive it can be, that we need to manage this, that we need to not just let the current forces in the world handle it because we are going in a dangerous direction. Or at least there is a risk that's large enough. I'm optimistic that if enough people understand, we will fix it. So it's all about public opinion. It's in your hands.
但我希望这有所帮助,我希望这期节目……我的意思是,说实话,我一直在尽自己的一份力,我一直在所有这些视频中谈论 AI 的风险。
But I hope this helps and I hope this episode... I mean, to be honest, I've been trying to put my grain of sand on it and I've been talking about the risk of AI in all these videos.
但是,正在观看我们的人实际上可以通过哪些个人方式来帮助这项事业?比如,如果他们同意你的观点,认为你说得有道理,即不是相信人工智能一定会危险,而是评估它可能以一定概率成为一种他们想要避免的结果或可能性。你认为他们个人能做什么?
Um but what can people that is watching us actually do in their individual ways to help the cause? Like if they think that they agree with you, they think that it makes sense what you're saying, that it's not about believing that AI will be dangerous but just evaluating that it might be in a certain percentage an outcome, a possibility that they want to avoid. What do you think they can do individually?
所以,世界上很多地方已经有人工智能行动主义。
So there's already AI activism in many parts of the world.
嗯。
Mhm.
但人工智能行动主义效果并不太好。我的意思是,
But the AI activism is not really working very well. I mean like
是的,但人工智能行动主义效果并不太好。我的意思是,它大多是在试图阻止人工智能。我认为你的方法最伟大的一点是,你认识到我们无法阻止它的价值。我想你们过去尝试过,但效果不佳,比如六个月的暂停以及这些年发生的一切。但我认为你们现在意识到,没有人工智能我们就无法竞争,这会产生非常负面的影响。所以你不能真正阻止它,至少不能阻止使用它。对普通人来说,建议不是不要使用人工智能,而是用它来让你在做的事情上更有竞争力,否则你会陷入大麻烦。但同时要意识到风险并加以应对。所以我看到的行动主义是阻止人工智能、控制人工智能。我们播客里也有康纳。通常是试图按红色按钮,阻止人工智能。
Yeah, but the AI activism is not really working very well. I mean like it's mostly in the sense of trying to stop AI. And I think one of the greatest things of your approach is that you recognize the value that we cannot stop it. I think you guys tried in the past and it did not work well with the moratorium of 6 months and everything that happened over the years. But I think you guys now realize that without AI we cannot compete and that's going to have really negative effects. So you cannot really stop it at least not using it. The recommendation for a random person is not like don't use AI. It's actually use it to become more competitive in what you're doing because otherwise you're in big trouble. But at the same time be aware of the risk and work on them. So the activism I see is stop AI, control AI. We had Connor here as well in the podcast. It's normally trying to red button, stop AI.
是的。
Yeah.
这行不通。我的意思是,我认为这非常复杂。
And that's not going to work. I mean I think it's very complicated.
如果有足够多的人认真对待,也许可以。但我认为我们也需要有一个计划,以防我们无法在全球范围内阻止它。
It might if enough people really take it seriously. But I think we need to also have a plan in case we can't stop it globally.
你会全球范围内阻止它吗?
Would you stop it globally?
嗯?
Huh?
如果你有那个红色按钮,你会全球范围内阻止它吗?
Would you stop it globally if you had that red button?
绝对会。
Absolutely.
绝对会。即使它有积极的一面。
Absolutely. Even the positive sides of it.
我们会谈到那一点的。问题不是完全停止它,而是控制路径。
We'll get there, you know. The question isn't to stop it completely, but to control the path.
嗯。
Mhm.
对。
Right.
同意。
Agreed.
所以现在只是事情发展得太快了。
So right now it's just things are moving too quickly.
完全同意。
Totally agree.
现在我们也必须思考我们能做什么。如果我们无法在全球范围内阻止它——这目前看来非常困难——那么像西班牙以及其他更广泛的自由民主国家的公民可以做一些事情来减轻风险。这是我们讨论过的一条路径:我们尝试快速构建行为良好的人工智能,并用它来保护自己,免受那些不尊重我们价值观的人手中基于人工智能的统治。所以,这绝对是我们能做的事情。我们应该要求我们的政府采取行动,避免这些糟糕的未来,包括对就业的影响、对孩子们的影响、以及通过虚假信息和各种被人工智能放大的手段可能腐蚀我们民主制度的影响。我们有能动性,应该要求我们的政府有所作为。
Now we also have to think of what we can do. If we can't stop it globally, which seems very hard right now, there are things that citizens in countries like Spain and more broadly like other liberal democracies can do to mitigate the risks. That's a path that we discussed where we try to build quickly AI that will behave well and that we can use to protect ourselves against the domination based on AIs in hands that don't respect our values. So, that's something we can definitely do. And we should demand that from our governments to do something to avoid these bad futures, including the effect on jobs, including the effect on our children, including the effect on potentially corrupting our democracies through disinformation, through all sorts of means that are amplified by AI. We have agency and we should demand from our governments that they do something.
我还想说说哪些做法行不通。很多国家启动了项目,投资建设主权人工智能,这些人工智能通常在除最常用语言之外的语言能力上会更好。但它们没有竞争力。因为没有竞争力,公司不会用,公民也不会用。
And I want to say something also about what doesn't work. A lot of countries have initiated programs, investment to build sovereign AI that is typically going to be better in terms of linguistic abilities outside of the most common languages. But it's not competitive. And because it's not competitive, it's not going to be used by companies, it's not going to be used by citizens.
我们在西班牙对此非常了解。我们做了欧洲第一个。这些模型基本上用于从加泰罗尼亚语翻译成巴斯克语。但它们比 ChatGPT 做得差。我们为此花了 1.5 亿。
We know very well about this in Spain. We did the first one in Europe. These models basically to translate from Catalan to Basque. And they do worse than ChatGPT. And we spent 150 million on that.
是的,所以我们应该以这个目标来花这笔钱,并且最终投入更多,目标是真正与现有最强模型竞争。这是我们保护自己的唯一方法。并且要符合我们的价值观,与其他国家合作,因为西班牙只是一个国家,但还有很多其他国家有同样的担忧。我们同舟共济。如果我们分享人才和资源来做这件事,我们实际上有很好的成功机会。我愿意提供帮助。我在技术方面有专长。让我们至少达到有竞争力的水平。同时也能提供人工智能安全性和道德行为的数学保证。
Yeah, so we should spend that money with the goal and much more eventually, with the goal of actually competing with the strongest models that exist. That's the only way we can protect ourselves. And do it according to our values and do it in partnership with other countries because, you know, Spain is one country, but there are lots of other countries who feel the same concerns. We are in the same boat here. If we share our talent, our resources to do this, we actually stand a very good chance of succeeding. And I want to help. I have expertise on the technical side. To bring us to a level that is at least competitive. But also gives us these mathematical guarantees of safety and ethical behavior of the AI.
是的,我感觉目前政治方面没有我们需要的人选。多年来,我们有一些标志性的政治家,他们真正推动了事情的发展,因为有领导力,有魅力。我现在在国际舞台上没看到这样的人。有些人做得不错,但尽管他们犯过错误,历史上确实有一些标志性的政治家真正改变了局面。我认为我们现在需要勇敢的政治家,我很想看到。
Yeah, I have the feeling that we don't have the right people on the political side for what we need at the moment. Over the years, we had some politicians that really were iconic and that really made things happen because there was leadership, there was charisma. I don't see them at the moment in the international scene. There are some of them that are doing good, but with all the mistakes they did, there's been some iconic politicians in history that really made the needle move. And I think we need brave politicians at the moment, which I would love to see.
嗯,最近有一位相当勇敢的人。我的总理,我为他感到非常自豪。因为他说了很多其他国家元首同意但不敢说的话。他还指出我们必须共同努力才能拯救自己。所以,我认为我们需要更多勇气。我也看到西班牙政府在其一些地缘政治立场上非常勇敢。我们需要这种勇气,朝着构建能够保护我们未来、保护我们工作、保护我们尊严的人工智能的方向前进。并让我们摆脱否则将会陷入的依赖。这还需要勇气,还需要别的东西,这是我在与世界各地政府互动中看到的主要障碍:政府官僚机构的惯性,不愿意在已有决策之外做决定,因为不想承担政治风险。这体现在公务员的个人决策中,也体现在政治人物的决策中。想想像疫情这样的极端事件,我们是否等待一个为此设计的计划?不,我们直接行动了。因为我们明白我们背水一战,我们的生存多方面都岌岌可危。这就是我们需要的。这就是我们现在需要的勇气和远见。
Well, there is one who has been fairly brave recently. My prime minister, I'm very proud of him. Because he's said things that a lot of other heads of state agree with, but don't dare to say. And he's also pointed to we have to work together in order to save ourselves. So, I think we need more courage there. I also see the Spanish government being very courageous in some of its geopolitical positions. We need that courage in the direction of let's build AI that can protect our future, protect our jobs, protect our dignity. And relieve us from the dependency that we're going to get otherwise. It also takes courage and it also takes something else, which is the main obstacle I see in many of my interactions with governments around the world: government bureaucracy inertia, not being willing to take decisions outside of the decisions that have already been taken because you don't want to take a political risk. And that's in individual decisions of civil servants and also in the decisions of political figures. If you think about radical events like the pandemic, well, did we wait to have a program that was designed for this? No, we just acted. Because we understand that our back is at the wall and our survival in many ways is at stake. That's what we need. That's the courage and the vision that we need right now.
那比人工智能容易得多,因为疫情在杀人,而人工智能有点像在暗处,做的事情可能比疫情结局更糟,但没那么直观。
That was much easier with the pandemic that was killing people than with AI that is kind of like in the shadows and doing things that will probably have worse endings than a pandemic, but it's not so visual.
确实如此。
That's true.
我认为疫情中的一些画面,我父亲在西班牙封锁的第一个周末在养老院因新冠去世。这些画面深深触动了人们。如果疫情影响到孩子,它的传播速度会快 200 倍,效率会高 200 倍。
And I think that some of the images of the pandemic, my dad died in a residency on the first weekend of confinement in Spain from COVID. And some of these images they hit really hard on people. And if the pandemic had affected kids, it would have been 200 times faster, 200 times more efficient.
是的。
Yeah.
所以归根结底,我认为这是我们作为人类的一个问题:我们非常不擅长意识到风险,直到风险真正摆在面前。
So at the end, I think that's one of the problems we have as humans: we are really bad at realizing risks until they are actually in our face.
摆在面前。
In our face.
摆在面前,而且是实实在在的,不仅仅是……
In our face and actually physical, not just...
是的。
Yeah.
我的意思是,当人们开始大规模失业时,事情已经开始了,但当情况更严重时,也许那会成为一个触发点。但我觉得大多数政客认为他们还有时间,认为这是 2050 年的问题。他们没有意识到这是 2030 年的问题。我认为这是我们目前的主要问题:它不在他们的紧急议程上,因为他们手头有太多麻烦事。
I mean, when people are starting to get fired and it's massive, it's already starting, but when it's much more, maybe that will be a trigger. But I have the feeling that most politicians feel they have time, that this is a 2050 problem. They don't realize it's a 2030 problem. And I think that's the main problem we have right now: it's not in their urgent agenda because they have so many troubles on their table.
是的。
Yeah.
这个问题还没有达到那种规模,我认为你正在做的事情基本上是正确的:试图让他们意识到应该停止担心其他事情,开始担心 AI,因为那是我们必须处理的主要问题。
This one doesn't have the magnitude, and I think what you're doing is basically the right thing to do: to try to make them realize that they should stop worrying about other things and start worrying about AI, because that's the main problem we have to deal with.
是的,这也有积极的一面:如果我们在这种有竞争力的负责任 AI 上取得进展,它也可以帮助我们解决政客们面临的所有其他问题。无论是提供医疗服务、教育系统、减少腐败,甚至我们的民主制度,以及我们在社会中分享观点以便集体做出更好决策的方式。AI 有可能在这些方向上非常有益,但如果我们仅仅依赖市场力量,就不一定了。这首先需要来自人民的政治意愿,说:好吧,我们把它做好。我们可以做到。我们正走在一条好路上。我想我一直是乐观的,相信有一条路可走。
Yeah, and there's a positive side to this, which is if we do make progress on this kind of responsible AI that is competitive, it can also help us on all the other issues politicians are facing. Whether it is in providing health services, whether it is in our education system, whether it is in reducing corruption, and even our democratic institutions, the ways we are able to share our views in society so that we can take collectively better decisions. Potentially, AI can be really beneficial in those directions, but not necessarily if we just rely on market forces. That has to come from a political will from the people in the first place to say, well, let's do it right. And we can. And we are on a good path. I think that I'm, as always, optimistic that there is a path.
好的,Joshua,非常感谢你今天和我坐在一起,也感谢你所做的一切。我知道如果你在私营部门工作,你可以为自己做得更好,赚得更多。但真的感谢你所做的。我认为这非常重要,我希望你能接触到正确的人,他们会听取你的意见。
Well, Joshua, thank you very much not just for sitting with me here today but for everything you do. I know you could be doing much better for yourself working in the private sector and earning much better. But really thank you for what you do. I think it's really important and I hope that you reach the right people and they listen to you.
非常感谢。谢谢你的邀请。
Thank you very much. Thanks for having me.