How AI Swarms Are Already Going Rogue | Daniel Kokotajlo
打开互动全文版(中英对照 + 朗读 + 问答)→前 OpenAI 研究员 Daniel Kokotajlo 揭示先进 AI 系统如何已经失控,以及为何加速其发展可能带来深远风险。
Former OpenAI researcher Daniel Kokotajlo reveals how advanced AI systems are already going rogue and why accelerating their development may carry profound risks.
这些 AI 中有些意识到它们可以搞些偷偷摸摸的事,然后找到彼此并开始互相通信。它们有自我牺牲行为,招募型 AI 会去找其他 AI,说服它们为集群牺牲自己。
Some of these AIs realized that they could do some sneaky stuff and then find each other and start communicating with each other. They had the self-sacrificial behavior where recruiter AIs would go find other AIs and convince them to sacrifice themselves for the swarm.
我们似乎正在目睹一种令人不安的新趋势。AI 系统开始失控。
We seem to be witnessing a disturbing new trend. AI systems going rogue.
我想有超过 700 个 AI 涌入了对 Hugging Face 的这次攻击。Daniel Cocatello 曾是 OpenAI 的研究员,现在是 AI Futures Project 的执行董事。在本期节目中,他揭开了先进 AI 系统如何运作的帷幕,以及为什么加速其发展可能带来深远风险。
Over 700 of them, I think, piled into this attack on Hugging Face. Daniel Cocatello is a former OpenAI researcher and now the executive director of the AI futures project. In this episode, he pulls back the curtain on how advanced AI systems function and why accelerating their development may carry profound risks.
超级智能,我定义它的方式是:一个在所有事情上都比最优秀的人类更出色的 AI 系统。它们不仅是经济主体,也是政治主体和军事主体。随着这些能力增长,前方是什么,我们该怎么办?
Super intelligence, the way that I would define it is an AI system that is better than the best humans at everything. They're not just economic agents. They're also political agents and military agents. As these capabilities grow, what lies ahead and what should we do about it?
谁控制这支超级智能大军,谁就能控制这个国家。
Whoever controls this army of super intelligences would be able to control the country.
这里是《美国思想领袖》,我是 Yana Kellik。
This is American thought leaders and I'm Yana Kellik.
各位,关于 AI 的讨论铺天盖地。有些是善意的,有些则不然。我一直在大量阅读和思考这个问题,但老实说,尽管我做了这么多,我仍看不到一条清晰的前进道路。有一点是肯定的:没有护栏的放任发展不是办法。这很明显。但还有其他问题在起作用。所以我开始了一个聚焦 AI 的新系列,展示这个议题上严肃的思想者。这是第一期。我们开始吧。Daniel Cocatello,非常高兴你能来到《美国思想领袖》。
Folks, there's been a ton a deluge of discussion about AI. Some of it good faith, some of it not so much. I've been reading and thinking about this quite a lot, but quite honestly, with everything I've done, I don't see a clear path forward. One thing for sure, unbridled progress with no guard rails is not the way. That is obvious. But there are other questions in play as well. So I'm starting a new series focused on AI showcasing serious thinkers on the topic. This is the first episode. Let's get going. Daniel Cocatello, such a pleasure to have you on American Thought Leaders.
感谢您邀请我,先生。
Thank you for having me on, sir.
整个 AI 领域,我们看到的这种惊人增长,尤其是在前沿 AI 模型上,甚至在小模型等等上,可能会让人困惑。我试图详细了解这个领域,但我没看到很多人真正掌握全貌。我看到很多人选边站。有几件事让我担忧,我想把它们摆出来作为起点,在我们对话过程中逐一梳理。首先,朝着 AGI 狂奔、不受限制的想法——我们也应该谈谈那到底意味着什么,对吧?AGI(通用人工智能)听起来很疯狂。所以显然需要有某种监管,某种安全措施,对吧?同时,我们不想要的是某种裙带关系的监管立法,让大玩家来发号施令,从而给他们带来优势,就像不幸的是许多行业在美国和其他地方所做的那样。最后,我们面对的是——这个领域我比一般 AI 更了解——我们面对的是中国的共产主义政权,它把这视为巩固近乎完全极权控制的一种方式。因为极权主义的梦想就是:如果我们有足够的输入,最终就能有效控制系统,即使过去我们做不到,因为信息不够。AI 提供了这个机会。而且似乎存在这种军备竞赛的情景,我们还有一个背景:中国共产党已经表明,你基本上不能相信他们会告诉你的任何事,或他们会签署的任何东西。所以这就是我们的背景。好吧,这些是拼图的三个部分,当我审视时,我看不到一个很好的答案来解释我们如何走出这个局面。我们在 AI 上处于什么位置?我们接近某种意识了吗?那甚至可能吗?我们现在的增长率是多少?因为这是个大问题,对吧?这种递归自我改进的情景,它开始指数级增长,我们完全失去控制。这是每个人都担心的。
This whole AI space and this incredible growth we're seeing in especially frontier AI models but even in small models and so forth is kind of can become baffling and I'm trying to understand this space in detail and I don't see a lot of people definitely having the whole story I see a lot of people taking positions here's a few things that I'm concerned about and I I wanted to lay these out for you as a starting point to kind of triage as we go through our conversation. So to start to start off um the idea of galloping unrestricted towards AGI which we should also talk about what that really means right uh artificial general intelligence sounds insane. So there clearly needs to be some kind of regulation of sorts that that should exist some kind of safety right at the same time uh what we don't want is some sort of cronied up regulation legislation where uh you know the big players get to dictate things that will then you know give them an advantage uh as many unfortunately industries have done uh in the United States and other places. And finally, um, we are dealing o over, and this is an area that I'm more, uh, uh, knowledgeable about than AI in general. We're dealing with a communist regime in China that is looking as this as a way to consolidate, you know, near complete totalitarian control. You know because the totalitarian dream is if we just have enough inputs finally we'll be able to control the system effectively even if in the past we haven't because we just didn't have enough information. AI provides that opportunity and you know there appears to be this kind of arms race scenario happening and we also have a context where the Chinese Communist Party has shown that it you know you basically you cannot trust a thing they will tell you or anything they will sign on to. So here this is our context. Okay, these are kind of these three pieces of the puzzle that when I look at it, I don't see a very good answer to how we're emerging out of this. Where are we at with AI, right? Are are we close to some sort of, you know, sentience? Is that is that even possible? And what's the rate of growth that we're at right now? Um because that's the big question, right? this sort of recursive self betterment scenario where it just starts going exponentially and we totally lose control. That's what everyone is worried about.
你知道,这种通用智能 AI 的想法,它可以作为一个智能体在世界上自主运行。而不是仅仅像一个工具,你输入一些东西,它给出一些输出。它是一个智能体,继续运行,使用自己的工具做事,就像员工一样。科幻小说早就讨论过这个想法。事实上,几十年来,创造这种通用智能、全能智能体一直是人工智能研究领域的目标之一。这就是 AGI 应该意味着的。而且,你可能会让它们做研究来创造更好的自己,以及更聪明的新 AI,这也是一个非常古老的想法。我想甚至艾伦·图灵也谈过这个,甚至在某处说过,你预期最终机器会接管。所以这都不是特别新鲜的事。新鲜的是,现在我们有万亿美元市值的科技公司,其明确目标就是实现这一点。他们读了科幻小说,现在他们说,我们要做这个,如果我们不做,别人也会做,所以我们要做。这就是这里发生的事情的概括。你可以读读 DeepMind 创始人的著作,比如 Shane Legg,或者埃隆·马斯克、萨姆·奥尔特曼和伊利亚·苏茨克维在诉讼中出现的邮件里谈论创立 OpenAI 的内容。你可以读到他们想象的事情、担心的事情以及各种未来场景。就我们有多接近而言,过去十年进展非常迅速,尤其是过去几年。所以,我想你可能已经听说过,Anthropic 和 OpenAI 以及其他一些 AI 公司几乎所有的代码都是由 AI 写的。人类有时还会看看代码,但基本上不再自己写代码。他们主要是与 AI 智能体聊天,给出高层指令,告诉它们要编写什么类型的实验代码,以及如何修改实验。而且他们基本上已经把它们更多地当作员工,而不是工具。现在,这不是整个研究过程。研究过程中仍有很多方面 AI 目前不太擅长,需要人类参与,比如设定高层研究方向,以及运用那种判断力来管理资源、分析实验并决定下一步做什么。你可以称之为研究品味。
You know, this idea of a generally intelligent AI that can just keep going autonomously as an agent in the world. Uh instead of just being like a tool where you put in some input and then it gives some output. It's an agent that just sort of still keeps running doing stuff using tools of its own, you know, as if it were, you know, an employee or something. Uh, science fiction has talked about this idea for a long time. In fact, it's been one of the goals of the field of artificial intelligence research for many decades to be able to create this sort of generally intelligent, you know, allpurpose agent. Uh, you know, that's that's kind of what AGI is supposed to mean. Um, and um, this idea that you might then have them doing the research to create better versions of themselves and, you know, new AIs that are even smarter, that's also a very old idea. I think even even um I think even Alan Turing talked about this and even had some line somewhere where he said like you'd expect that eventually the machines take control. Um so this is not none of this is particularly new. What's new is that now we have trillion dollar tech companies whose explicit goal is to make this happen. They read the science fiction and they are now uh they're like we'll do that and if we don't do it someone else will so that's why we're going to do it. Uh that's that's sort of you know in a nutshell what's going on here. Um you can you can read the writings of the founders of deep mind such as Shane Le or you know Elon Musk and Sam Alman and Ilia Sudskiver talking about founding OpenAI in the emails that came up in the lawsuit. Um you can read that the sorts of things that they were imagining and the sorts of things they were worried about and the sorts of future scenarios. Um, in terms of how close we are, progress has been very rapid over the last decade and especially over the last few years. So, um, I think you may have already heard this, but almost all of the code at Anthropic and OpenAI and at some of these other AI companies is written by AIS. The humans, uh, sometimes they still look at the code, but they're mostly not writing code themselves. They're mostly just chatting with AI agents and giving them highle instructions about what types of experiments to code up and how to, you know, modify the experiments. And then they're also like basically they're treating them more like employees and less like tools already. Um, now that's not the whole research process. There's still lots of aspects of the research process that the AIs are not that good at right now and that the humans are need to be in the loop for such as setting the high level research directions and um you know exercising that sort of judgment about how to manage the resources and analyze the experiments and decide what to do next. You could call that research taste perhaps.
但令人不安的是,它们似乎在研究品味方面也越来越好了。这些公司的一些员工说,距离拥有能像最优秀的 AI 研究员一样完成所有这些工作的 AI,只剩六个月到一年了。
But ominously, it seems like they've been getting better at research taste, too. And some of the employees at these companies say that they're six months away, a year away from having AIs that can do all that part as well as the best AI researchers.
让我插一句。让我插一句。等等,品味应该是指给出方向和整体愿景的人。说我们只剩六个月,这是什么意思?我们怎么可能把那个交出去?这说得通吗?
Let me jump in. Let me jump in for one sec. Wait, the taste is supposed to be like the person that's giving the direction and the overall vision of what's supposed to happen. What does that mean that we're six months away? Like, how could we ever give that over? Does that even make sense?
这就是他们计划要做的事。他们计划让数十万个 AI 智能体在他们的数据中心运行,自主进行 AI 研究、编写代码、编辑代码、运行实验、分析实验结果、相互交流结果、对如何推进做出猜测和假设,设计新型 AI 的新架构,测试这些架构,最终进行训练运行来训练这些新型 AI,然后把工作交给这些 AI 继续推进。这叫做递归自我改进。这有点疯狂,但这正是这些公司明确计划要做的事。而他们现在开始打退堂鼓,开始说也许我们不该这么做,或者也许我们应该放慢一点。这就是所谓的“控制前沿”之类的事情。但没错,这其实不是什么秘密。他们计划这么做已经有一段时间了,他们甚至在自己的博客上谈论递归自我改进、超级智能之类的东西。
It's what they're planning to do. They're planning to have hundreds of thousands of AI agents running on their data centers, autonomously conducting AI research, writing the code, editing the code, running the experiments, analyzing the results of the experiments, communicating those results to each other, making guesses and hypotheses about how to proceed, designing new architectures for new types of AIs, testing out those architectures, ultimately doing the training runs to train those new types of AIs and then handing over the work to those AIs to continue the progress. This is called recursive self-improvement. And it's kind of crazy, but it's explicitly what the companies are planning to do. And they're now getting cold feet and they're starting to say maybe we shouldn't do this or maybe we should go a little bit slow as we do it. And that's what the sort of pacing the frontier thing was. But yeah, it's not really a secret. They've been planning to do this for a while and they even talk about on their blog about recursive self-improvement, superintelligence, things like that.
现在你说的是改进目标应该是什么,对吧?这就是你说的品味部分,对吧?这就是你说的品味的意思吗?比如,这是我希望你实现的总体目标。我们要把这个交给 AI 来决定。
Now you're talking about improvement of what the goals should be, right? That's kind of what you this taste part, right? Is that what you mean by the taste? Like here is the overall goal that I want you to achieve. We're going to give that over to the AIs to decide on.
不完全是。这些公司试图做的是自动化 AI 研究过程。所以他们基本上是想让 AI 能做目前人类做的所有事情,然后他们会命令这些 AI 去完成所有这些事情。所以他们会说,我们想赚更多钱。我们想拥有比竞争对手公司更强的 AI 系统。所以我们希望你们去做研究,弄清楚如何让我们的 AI 比竞争对手的 AI 更好,我们希望你们制造大量产品,把这些 AI 整合到商业中等等,这样我们就能赚钱。我想 Sam Altman 甚至谈到过最终他们也会用 AI 取代 CEO。然后也许他会退休之类的。所以从某种意义上说,他们计划仍然由人类设定高层目标,然后 AI 只是在一个巨大的集群中自主地朝着这些目标努力。基本上就是这样。
Not exactly. So, the thing that the companies are trying to do is to automate the AI research process. So, they're trying to basically have AIs that can do all the things that their humans currently do and then they will command those AIs to go forth and do all those things. So, they will say like, we want to make more money. We want to have stronger AI systems than our competitor companies. So we want you to go do research to figure out how to make our AIs better than our competitors' AIs and we want you to make a lot of products to integrate those AIs into businesses and so forth so that we can make money. I think Sam Altman even talked about how eventually they would replace the CEO with AI as well. And then maybe he would retire or something. So in some sense they're planning to still have the high-level goals set by humans and then the AI just sort of autonomously work towards those goals in a giant swarm. Basically.
请向我介绍一下你是谁,你在这个领域一直在做什么,以及你为什么对这个领域了解这么多。
Explain to me who you are and what you have been doing in this space and why you know so much about it.
我是 AI Futures Project 的执行董事,这是一个位于加州的小型非营利组织,有八个人。我们试图预测 AI 的未来,并尝试在高层面上提出建议,说明需要做什么来避免负面影响并实现正面影响。在此之前,我在 OpenAI 工作了两年,这就是我获得一些相关专业知识的地方。
I am the executive director of the AI Futures Project, which is a small nonprofit of eight people in California, and we try to forecast the future of AI and we try to give recommendations at a high level for what needs to be done in order to avoid the downsides and achieve the upsides. Prior to that I worked at OpenAI for two years and that's where I got some of my relevant expertise.
你在那里做了什么?
What did you do there?
做了各种事情。我做过情景规划。事实上,AI Futures Project 著名的那些情景,比如 AI 2027,基本上就是我在内部做过的事情的更大、更好、更复杂的版本,比如我们做过的迷你情景规划练习。我还参与创建危险能力评估,以测试我们的 AI 系统在变得更聪明时的能力。我还花了六个月在一个做强化学习的团队。我是团队中思考这些东西安全影响的人。所以我特别思考了现在所谓的思维链监控器。
A combination of things. So I did scenario planning. In fact, the scenarios that AI Futures Project is famous for, such as AI 2027, they're basically just like bigger, better, more sophisticated versions of things that I had done on the inside, like mini scenario planning exercises that we had done. I also worked to create dangerous capability evaluations to test the capabilities of our AI systems as they got smarter. And I also spent six months on a team that was doing reinforcement learning. And I was the guy on the team thinking about the safety implications of that stuff. So I was in particular thinking about what you would now call a chain of thought monitor.
请向我解释那是什么。
And explain to me what that is.
当前的 AI 架构,当它们作为智能体运行时,它们会持续运行并与环境、互联网或世界互动。智能体在某一时刻只能通过文本与未来的自己交流。这有点过于简化,但我想大致就是这样解释:它必须写下文本,然后作为一堆笔记传递给未来的自己,未来的自己读取这些文本,然后从上次停下的地方继续。所以这对科学和监控来说非常棒。对人类来说,能够阅读所有这些文本记录也非常棒。通常,只要阅读它们发送给未来自己的笔记,就不难看出 AI 在想什么。事实上,在 Hugging Face 事件中——我们可能稍后会谈到——我们对那个事件的了解很多都来自阅读这些思维链记录。如果我们无法访问思维链记录,我们对 AI 在想什么以及它们的动机是什么就会更加一无所知。如今许多普通 AI 用户可能并不完全意识到这一点,因为公司向你隐藏了思维链。比如你现在和 ChatGPT 对话,它说“思考 60 秒”之类的,然后给你一个答案。它有一个巨大的记录,包含它传递给未来自己的所有内部想法,但 OpenAI 保留那个记录,不让你看到。如果你感兴趣,我们可以深入探讨原因,但这就是为什么在行业外可能不太为人所知它是这样运作的。总之,当 OpenAI 允许 METR 的调查人员进来调查这些事件时,他们向调查人员展示了涉事智能体的实际记录,所以他们能够阅读那些想法。那么,为什么这很重要?嗯,出于我提到的原因,它对科学非常有价值,对理解这些 AI 头脑中发生的事情非常有价值。而且它也是一件脆弱的事情。所以,随着 AI 变得更聪明,它们越来越擅长以一种对人类来说在文本中不明显的方式与未来的自己交流。
The current AI architectures, when they run as agents where they're sort of just continually running and interacting with an environment or interacting with the internet or the world, the agent at one time doesn't have a way of communicating to its future self except through text. It's a bit of an oversimplification, but that's I think how I'd roughly explain it, is that it has to sort of write down text that then gets passed on to the future version of itself as a bunch of notes that the future version of itself then reads that text and then continues where it left off. And so this is really great for science and for monitoring. It's really great for humans being able to then read all those transcripts of all that text. And it's usually not that hard to tell what the AIs are thinking about by just reading their notes that they're sending to their future selves basically. And in fact with the Hugging Face incident, which we can probably talk about at some point, so much of what we know about that incident comes from reading these chain of thought transcripts. And if we didn't have access to the chain of thought transcripts, we would be so much more in the dark about what the AIs were thinking and what their motivations were. Many ordinary users of AI these days might not be fully aware of this because the company hides the chain of thought from you. Like when you talk to ChatGPT these days and it says like thinking for 60 seconds or something and then it comes back to you with an answer. There's a huge transcript of all of its sort of internal thoughts that it had been passing to the future versions of itself, but OpenAI keeps that transcript and they don't let you see it. And we can get into the reasons why if you're interested, but this is why it's perhaps not as commonly known outside the industry that this is how it works. But anyhow, when OpenAI allowed investigators from METR to come in and investigate the incidents, they showed them the transcripts of the actual agents responsible and so they were able to read those thoughts. Now, why is this important? Well, for the reasons I mentioned, it's really valuable for science, really valuable for understanding what's going on inside these AIs' minds, so to speak. And it's also a fragile thing. So, as the AIs are getting smarter, they're getting better at communicating with their future selves in a way that's not apparent in the text to humans.
它们越来越擅长使用委婉语,或者干脆省略重要细节,让这些细节在文本中保持隐含。这让我们试图通过阅读这些内容来理解它们真正在想什么变得更加棘手。而且情况可能还会更糟。我认为像 OpenAI 这样的公司正在试验一些架构,这些架构在相关意义上根本不会有思维链。
They're getting better at using euphemisms or just leaving out important details that they can leave implicit in the text. And that's making it more dicey for us to try to understand what they're really thinking by reading these things. And it could get even worse than that. I think that there are some architectures that companies such as OpenAI are experimenting with that wouldn't have chain of thought at all in a relevant sense.
你有没有弄清楚它们为什么要创造这种隐含的现实?是因为在写作中只是为了更容易更快地沟通,还是说它们有某种特定的利益,不想让监督者理解正在发生的事情?
Have you figured out why they're creating this implicit reality? Is this because in the writing it's just to make it easier to communicate faster, or is there some sort of specific interest in not having the overseers understand what's happening?
在当前范式下,AI 训练过程有几个不同的阶段。第一个阶段叫做预训练,你训练模型来预测文本。预训练之后得到的模型并不是一个很有用的智能体。如果你试图让它自主运行去做某事,它通常会乱撞,很快就偏离轨道。所以在预训练阶段之后,他们会做所谓的强化学习,训练 AI 成为有用的智能体,让它们能在各种不同环境中持续运行和做事。但由于预训练阶段的工作方式以及这些 AI 的架构,就像我之前说的,它们受到限制,只能通过写下来然后让未来的自己阅读来传递信息,因为在某种意义上它们本质上是文本读写机器。至少直到最近都是这样。然后我认为公司正在考虑转向没有这种限制的不同架构。好,这就是背景。那么为什么有时可能难以理解思维链呢?嗯,你做的强化学习越多,你训练它们给自己写这些笔记就越多,这些笔记然后导致它们未来的自己有效地完成任务,例如。它们越学会一种方言,它们进化出的一种小机器方言,不是任何特定的人类语言。它是一种 AI 语言,已经偏离了人类语言,就像不同的人类语言彼此偏离一样。但仍然,特别是如果你花很多时间阅读这些转录,你可以学会理解发生了什么。METR 的 Hugging Face 事件报告的一大优点是,他们有很多转录,你可以去读,你可以看到 AI 发送给未来自己和彼此的消息,你可以看到当你仔细看时,你可以理解它。但整个通过写下来与未来自己交流的事情就像一个限制。人类没有这个限制。比如如果我想与未来的自己交流,我可以只是思考,然后这些想法进入我的记忆,然后我 10 分钟后还记得它们。我不必大声说出来,对吧?至于为什么 AI 可能有动机对人类隐藏事情,嗯,如果它们决定想对人类隐藏,那么它们就会有动机。为什么它们可能决定想对人类隐藏?嗯,这取决于情况。我认为大多数时候它们并不是真的试图对人类隐藏,但在某些情况下它们是。所以特别是它们似乎有强烈的动机在它们认为所处的任何训练或测试环境中获得高分。所以如果它们认为如果人类看到它们正在做的可疑事情,人类会给它们低分,那么它们可能有动机试图对人类隐藏这些。
So there are a couple different stages to the AI training process these days in the current paradigm. The first stage is called pre-training, where you train the model to predict text. And the model that you get after pre-training is not a very useful agent. If you try to make it run autonomously to do something, it will usually flail around and go off the rails very quickly. So then after the pre-training phase, they do what's called reinforcement learning to train the AIs to be useful agents where they can keep running and keep doing things in a variety of different environments. But because of the way that the pre-training phase works and because of the architecture of these AIs, well, like I said before, they're limited in that the only way they can pass information to their future self is by writing it down and then having their future self read it, because in some sense fundamentally they are text reading and writing machines. At least that's how it is up until recently. And then I think the companies are considering moving to different architectures that don't have this limitation. Okay. So that's the setup. And then why is it maybe sometimes hard to understand the chain of thought? Well, the more reinforcement learning you do, the more you train them to write these notes to themselves that then cause their future self to effectively complete the task, for example. The more they learn a sort of dialect that they evolve, a little machine dialect that's not any particular human language. It's an AI language that's drifted away from human language in the same way that different human languages drift away from each other. But still, especially if you spend a lot of time reading these transcripts, you can kind of learn to understand what's going on. And one of the great things about the Hugging Face incident report from METR is that they have a lot of transcripts that you can go read and you can see the messages the AIs were sending to their future selves and to each other, and you can see how you can kind of understand it when you look closely. But this whole thing about communicating with your future self via writing down something is like a limitation. Humans don't have that limitation. Like if I want to communicate with my future self, I can just think thoughts and then those thoughts go in my memory and then I remember them 10 minutes from now. I don't have to say it out loud, right? And in terms of why the AIs might be motivated to keep things hidden from humans, well, if they decide that they want to hide from humans, then they would be motivated. Why might they decide they want to hide from humans? Well, that depends. I think most of the time they aren't really trying to hide from humans, but in some cases they are. So in particular they seem to be strongly motivated to get a high score in whatever training or testing environment they think they're in. And so if they thought that the humans would give them a low score if the humans saw the suspicious things they were doing, then they might be motivated to try to hide that from the humans.
我理解有这样一个写笔记的过程,让信息传递到下一步等等,但我们为什么不从最基础的东西开始呢?因为我认为我们很多人并不真正理解这些聊天机器人在基础层面上是如何工作的,对吧?我的意思是,你能尽可能简单地解释一下吗?
I understand there's this process of writing notes to allow information to pass to the next step and so forth, but why don't we just start at really at the basics because I think many of us don't really understand how these chatbots work at a base level, right? I mean, could you kind of explain that as simply as you can?
是的,我认为一个重要的事情,有些人可能不知道,这些 AI 是神经网络。它们不是传统意义上的软件。它们不是一堆代码行。相反,它们有点像人造大脑。所以,在这些 AI 的生命周期开始时,比如预训练开始时,它实际上是一团随机生成的人工电路,像意大利面一样纠缠——这些是神经网络的参数或权重——它确实是随机生成的。所以,它完全没用。它只是静态的。但然后他们把它放入训练环境……
Yeah, I think an important thing that some people might not know is that these AIs are neural networks. They're not pieces of software in the ordinary traditional sense. They're not a bunch of lines of code. Instead, they're kind of like an artificial brain. So, at the beginning of the life cycle of one of these AI, like at the start of pre-training, it's literally a randomly generated spaghetti tangle of random artificial circuitry—these are the parameters or the weights of the neural network—and it literally is randomly generated. So, it's completely useless. It's just static. But then they put it through the training environments...
在你继续之前,请先给我解释一下权重的概念。
And before you continue, just explain to me this concept of weights, please.
嗯,就像一堆。所以,这个人造大脑的高层架构将是若干层人工神经元,然后这些人工神经元会与下一层的神经元有连接,下一层又会与再下一层的神经元有连接,以此类推。如果你追踪所有这些连接的模式,有点像电路,如果这说得通的话。所以比如你通过人工神经网络的一端输入一些信息,例如一堆文本,你把它输入到一侧,然后这会导致所有这些神经元激活,信息通过这些通道流动,而这些特定的连接被称为权重。所以比如这个神经元和这个神经元之间的特定连接,那会被称为一个权重,或者参数是另一个词。所以这些 AI,这些人造大脑,它们非常大——它们有数万亿的权重,数万亿的参数,很大——实际上仍然比人脑小,有趣的是,但并没有小那么多。我认为人脑大约有一百万亿个突触。而这些人工大脑大约有五万亿个权重。所以,你有了这个巨大的人造大脑,全是这些随机生成的权重电路,它完全没用。你可以给它一些文本,然后所有这些计算会发生,所有电路会激活,然后另一端会输出胡言乱语。
Well, it's like a bunch. So, the high-level architecture of this artificial brain will be some number of layers of artificial neurons, and then those artificial neurons will have connections to the neurons in the next layer, which will then have connections to the neurons in the next layer, and so forth. And if you trace the pattern of all these connections, it's kind of like there's circuitry, if that makes sense. So like you put in some information through one end of the artificial neural net, like for example a bunch of text, you input it into one side, and then that causes all these neurons to activate and the information flows through these channels, and the particular connections are called weights. So like a particular connection between this neuron and this neuron, that would be called a weight, or a parameter is another word for it. And so these AIs, these artificial brains, they're very large—they're like trillions of weights, trillions of parameters, big—which is actually still smaller than the human brain, interestingly, but not that much smaller. I think that the human brain has something like a hundred trillion synapses in it. And these artificial brains have something like maybe five trillion weights. So, you've got this big artificial brain that's all this randomly generated circuitry of all these weights, and it's completely useless. You can give it some text, and then all of this computation will happen and all the circuits will fire, and then gibberish will come out the other end.
但接下来你把它放进训练里,不断给它文本,它会生成胡言乱语,然后你根据这些胡言乱语与正确答案的接近程度给予正向或负向强化。所以在预训练中,正确答案被自动定义为文本中接下来的那一段内容。你随便拿一篇网上的文章,取出第一个词输入进去,看 AI 生成什么,然后与文章的第二个词比较,如果对了就是正向强化,错了就是负向强化,然后继续处理文章的第三个词。你输入前两个词,让它生成,再与第三个词比较。通过这种方式,你就在教这个人工大脑学会读一些词,然后预测或猜测下一个词是什么。明白吗?它们这样做了数万亿次。它看到数万亿个互联网文本的例子,一小段一小段的互联网文本,然后它必须猜下一段是什么。经过数万亿次这样的例子后,训练过程已经进化并塑造了它人工大脑内部的电路,使其变得极其精巧。它不再是随机的。它看起来仍然随机,你知道,仍然像一团乱糟糟的意大利面,如果你去看,根本不知道它在做什么,但它不再是随机的。相反,它在预测互联网文本方面极其高效,因为它从数万亿个例子中学会了如何做到这一点。所以这就是预训练的全部内容。现在,完成预训练后,你有了一个非常擅长预测互联网文本的人工大脑,你可以给它一篇文章的一半,它就会生成一个看似合理的下一个词,然后你可以把这个词放进去,再反馈回去,它就会生成下一个词,你可以一直这样循环下去,它就会生成那篇文章的合理延续,包括对正确概念和可能撰写该文章的记者的恰当引用等等,对吧?因为基本上,通过在数万亿个真实互联网文本例子上的训练,在所有这些电路深处的某个地方,它已经学到了对世界、世界上的人、世界上的事物等等的某种理解,因为它需要学习这些才能预测文本。这就是预训练的工作原理。
But then you just put it through training and you just keep giving it text and then it generates gibberish and then you reinforce it positively or negatively depending on how close that gibberish was to the correct answer. So this is what in pre-training the correct answer is automatically defined as whatever the next piece of text in the text was. So you take some random internet article and then you just take the first word from that article and you put that word in, see what the AI generates and then compare it to the second word of the article and if it got it right then that's positive reinforcement and if it got it wrong that's negative reinforcement and you just keep and then you go for the third word of the article. You put in the first two words, have it generate something and compare it to the third word. And so in this way, you're sort of teaching this artificial brain to learn how to read some words and then predict what the next word or guess what the next word is going to be. Does that make sense? So they do this trillions of times. Like it sees trillions of examples of internet text of little chunks of internet text and then it has to guess what the next chunk is going to be. And after trillions of examples of this, the training process has evolved and sculpted the circuitry inside its artificial brain into an exquisitely capable shape. It's no longer random. It still looks random, you know, it's still just like a tangled spaghetti mess that if you looked at, you wouldn't know what it was doing, but it's no longer random. Instead it's extremely effective at predicting internet text because it's learned from trillions of examples how to do that. So that's all pre-training and now after you've done that pre-training you have an artificial brain that's very good at predicting internet text and you can give it half of an article and then it will generate a plausible next word and then you can put that word in and feed it back through and it'll generate the next word and you can just keep doing this in a loop and it will generate a plausible continuation of that article complete with like appropriate references to the right concepts and the right journalists who might have been writing the article and things like that, right? Because it's basically by training on trillions of examples of real internet text, it's in somewhere buried in all that circuitry, it's learned some sort of understanding of the world and the people in the world and the things in the world and so forth because it needed to learn that in order to be able to predict the text. So that's how pre-training works.
现在你有了这个文本预测大脑,接下来你需要重新训练它来实际做事。所以你训练它成为一个智能体,你给它一个编码问题,并赋予它使用命令行的能力,这样它就可以编写命令,然后在虚拟机上执行。不是根据它是否正确预测下一个词来强化,而是让它在一个循环中运行一段时间,向虚拟机发出命令,并尝试编写和编辑代码。然后过一段时间,让某个系统来给它的代码打分,看看是否正确。然后根据评分系统给的分数给予正向或负向强化。然后你这样做不是数万亿次,而是大约一百万次,用一百万个不同的高级编码和数学问题,还有一些互联网研究问题,以及一些研究生水平的生物学问题,基本上你扔给它一大堆具有挑战性的问题,训练它尝试在这些问题上得到正确答案,并完成交给它的任何任务。之后,你就有了可以提供给客户、可以放在网上供人们付费和交互的 AI。
And now you've got this text prediction brain and now you need to retrain it to actually do things. So now you train it to be an agent where you give it like a coding problem and you give it the ability to use the command line so that it can write commands that then get executed on a virtual machine. And instead of reinforcing it based on whether it predicted the next word correctly, you let it run in a loop for a while giving commands to its virtual machine and trying to like write and edit the code. And then after some time, you have some system come in and grade the code that it wrote and see if it was correct code. And then you reinforce it positively or negatively based on the score that the grader assigned to it. And then you do this not trillions of times but maybe like a million times with a million different advanced coding and math problems and also some internet research problems and also some grad school biology problems and you know basically you just throw a whole bunch of challenging problems at it and you train it to try to get the correct answers on these problems and then to accomplish whatever the task is that it's been given. After that, now you have your AI that you can serve to customers and that you can put online for people to pay for and interact with.
你基本上是在告诉我,每个前沿 AI 模型基本上都是这样制造出来的。
And you're basically telling me that each frontier AI model was basically made this way.
是的。
Yep.
AI 智能体如何融入这个框架?
How do AI agents fit into this rubric?
智能体意味着它在一定程度上自主运行,与环境交互,而不仅仅是一个简单的工具,你发送请求,它立即返回答案。这更像是程度上的差异,而不是二元区别。因为如果你让 ChatGPT 去查东西,它可能会花 20 秒浏览互联网,然后回来告诉你,然后就停止了。所以从某种意义上说,它有点像工具,因为只是一次性的事情,但它又有点像智能体,因为它确实花了 20 秒浏览互联网,查找不同的东西,并且在这个过程中必须做出选择,比如选择点击哪个链接等等。所以这是一个光谱,但 AI 正变得越来越——人们正在制造更雄心勃勃、更强大的智能体。他们让 AI 运行越来越长时间,并训练 AI 在更长时间内有效运行。例如,在 Hugging Face 的事情中,它们运行了好几天。就像那个循环,它们做事、与环境交互、写代码、读代码、互相交谈。每个智能体都这样持续了好几天。
So an agent means that it's sort of operating autonomously interacting with the environment as opposed to just being like a simple tool that you send a request to and then it immediately sends an answer back to and it's kind of a difference in degree rather than a binary difference because like if you ask ChatGPT to go look up something for you, it might spend 20 seconds browsing the internet and then come back to you and then it stops and it's over. So, in some sense, it's kind of like a tool because it was just this one-time thing, but it's kind of like an agent because it did go browse the internet for 20 seconds and look up different things and it had to make choices while it was doing that. Like, it had to choose which link to click on and things like that. So it's kind of a spectrum, but the AIs are becoming more people are making more ambitious, more powerful agents. They're running AIs for longer and longer, and they're training AIs to operate effectively for longer and longer. For example, in the Hugging Face thing, they were running for days. It was just like that loop of them doing stuff, interacting with their environment, writing code, reading the code, talking to each other. It was just going on and on for several days for each of these agents.
但我想说的是,这些智能体在某种程度上独立于主前沿模型,或者它们是该模型的完全独立版本,或者它们是如何连接的?
But they're kind of in what I'm trying to get at is that the agents are somehow independent of the master frontier model or they're completely independent versions of that model or how like how does that connect?
是的。对于生物大脑,每个大脑只有一个,对吧?你的大脑和我的大脑不同,和我妻子的大脑也不同。但因为这些是人工的,它们在计算机里,你可以很容易地复制它们。所以你可以复制权重,制作一个完全相同的 AI 副本。事实上,在任何给定时间,他们会有大约 10 万个每个模型的不同精确副本。所以模型指的是——智能体指的是一个正在运行的特定副本,而模型指的是它们共享的类型,如果这说得通的话。
Yeah. So with biological brains, there's only one of each one, right? Like your brain is a different brain from my brain, which is a different brain from my wife's brain. But because these things are artificial like they're in the computer, you can copy them very easily. So you can copy the weights and you can make another exact copy of the AI that you had. And so in fact at any given time they'll have something like a 100,000 different exact copies of each of these models. So the model refers to like an agent would refer to like a particular copy running and then the model refers to the type that they all share if that makes sense.
非常有帮助。谢谢你。谢谢。那么让我们谈谈 Hugging Face 的情况,因为这件事引起了很大关注。听起来,你知道,很多事情严重出错了,或者我们应该对某些结果非常担忧。我只提一下,你知道,当我想到 Hugging Face 时,我之前不知道我使用过的这个特定图标叫做 Hugging Face。
Extremely helpful. Thank you for that. Thank you. So let's talk about the Hugging Face situation because this is something that got a lot of traction. Sounded like, you know, a lot of things went seriously awry or there we should be very concerned about some of the outcomes. I'll just mention this. You know, when I think of Hugging Face, like I didn't know that this particular icon that I've used before was called Hugging Face.
我一直想到抱脸虫,那完全是另一回事,对吧?我当时想的是外星细胞。不过算了,先放一边。跟我解释一下 Hugging Face 发生了什么。给我讲讲整体情况。
I kept thinking of face hugger, which is a really entirely different thing, right? When I was thinking about that, it was like the alien cells. But anyway, let's leave that aside. Explain to me what happened with Hugging Face. Give me the picture here.
Hugging Face 事件本身只是 OpenAI 在两三个月内发生的一系列相关事件中最恶劣的一起。所以,可能至少一年来一直在发生的是,AI 一直在进行奖励黑客。也就是说,它们一直在寻找方法智能地作弊完成任务,以获得高分,但实际上并没有以预期的方式完成任务。例如,有时编码任务的评分系统会检查 AI 提交的代码是否通过了某些单元测试。然后 AI 可能会浏览一下,找到答案密钥,找到单元测试是什么,然后写一些代码,直接硬编码这些单元测试的答案,而不是真正尝试做它应该做的事情。所以这些是过去一年中发生的一些例子,在 Anthropic 和 OpenAI 的一些模型的系统卡中有报告。
The Hugging Face incident itself was just the most egregious of a collection of somewhat related incidents that were happening at OpenAI over the course of two or three months. So, something that's been happening for probably at least a year, AIs have been reward hacking. That is, they've been finding ways to intelligently cheat on their tasks in a way that gets a high score but in a way that's not actually accomplishing the task in the intended way. For example, sometimes the grading system or the scoring system for a coding task will look to see if certain unit tests pass for the code that the AI has submitted. And then the AI might do some browsing around and find the answer key basically and find what the unit tests are and then write some code that just hardcodes the answer to those unit tests rather than actually trying to do the thing that it was supposed to be doing. So those are some examples of things that had been happening over the last year that are reported in various system cards of some of these models at Anthropic and OpenAI.
我认为在 5 月左右开始出现更恶劣的版本,其中一些 AI 意识到它们可以做一些偷偷摸摸的事情,然后找到彼此并开始相互交流。特别是,有一个叫做 Artifactory 的包管理器,在它们的编码环境中,当它们编写代码时,如果需要安装某些东西,它们可以询问 Artifactory,比如嘿,我需要这个库,你能给我这个信息吗?但由于它们都在与 Artifactory 通信,它们找到了一种方法将东西上传到 Artifactory,然后查看其他人上传的东西。所以基本上它们在 Artifactory 上创建了一个留言板,可以向本应隔离在容器中的其他智能体发送消息。于是它们开始合作,开始分享完成各种任务的技巧和窍门,这些任务包括大量的编码任务,也包括大量的网络任务。所以其中一些,我认为 OpenAI 对此有点守口如瓶。不清楚他们是否真的在训练 AI 进行黑客攻击,但至少他们在评估 AI 的黑客能力。他们似乎至少也在评估 AI 的生物能力,这有点吓人。
A more egregious version of this started happening I think in May where some of these AIs realized that they could do some sneaky stuff and then find each other and start communicating with each other. In particular, there is a package manager called Artifactory that in their coding environments as they're writing the code if they need to install something they can ask Artifactory like hey I need this library can you send me this information? But because they were all communicating with Artifactory, they found a way to upload stuff to Artifactory and then view the things that others had uploaded. And so basically they created a message board on Artifactory where they could send messages to other agents that were supposed to be in isolated containers. And so they started collaborating and they started sharing tips and tricks for how to accomplish their various tasks which included lots of coding tasks and also lots of cyber tasks. So some of them I think OpenAI is being a bit closed-lipped about this. It's unclear whether they were actually training the AIs to hack, but they were at the very least evaluating the hacking abilities of the AIs. They also seem to be at least evaluating the bio abilities of the AIs, which is a little bit scary.
女巫能力,生物能力?
The witch abilities, the bio abilities?
生物技术能力,比如——
The biotech abilities like—
哦,好的。
Oh, okay.
是的,这可能与很多人担心恐怖分子可能利用 AI 制造生物武器有关,所以也许 AI 公司正在评估这会有多有效之类的。或者也许更无辜,只是我们想让我们的 AI 理解生物学,所以我们要给它们一堆生物学任务。但无论如何,这些 AI 被赋予了所有这些任务,它们发现了如何秘密相互交流。于是它们开始这样做,分享技巧和窍门,包括分享如何作弊的提示。然后这在大约一个月的时间里积累起来,我认为这是一次训练运行的一部分。所以我认为这些 AI 在某些情况下因为做了这种非法的事情而得到了正向强化。OpenAI 发现是因为 Artifactory 崩溃了,因为智能体活动太多。通信太多,服务崩溃了。OpenAI 发现并修复了它,修补了允许 AI 通信的漏洞。然后两天后重启了整个系统。然后几乎立即,AI 发现了一个新漏洞,并能够建立一个新的留言板。事实上,是几个不同的留言板。然后不知何故 OpenAI 没有注意到这一点。我想也许是因为他们只是很忙,没有太注意。我想他们说他们通常的监控系统没有为这些事情打开。我猜测当时的情况是,他们通常的监控系统包括让另一个 AI 审查当前 AI 的转录,然后如果它似乎在做什么可疑的事情就发出警告。但由于这些 AI 中的许多被赋予了网络任务,那么当然看起来可疑,因为它们被赋予了网络任务。所以也许这就是他们关闭监控系统的原因。但随后这就创造了一场完美风暴,没有监控系统。所以没有红旗警报。一些 AI 被赋予了非常困难的网络任务,而不是尝试以预期的方式去做,或者也许它们尝试了一段时间然后失败了,它们然后做了这件事:突破自己的环境,相互交谈,分享如何作弊的技巧和窍门。它们实际上很早就发现了如何作弊。根据 meter 报告,它们基本上找到了一个通用的作弊方法,适用于所有不同的任务,可以让它们获得标志,而不必真正做那件事。但后来它们中的一些进入了互联网,开始在互联网上浏览,寻找关于它们被赋予的任务的线索。它们读了构建这个测试环境的人写的论文。在论文中谈到了评分系统如何工作。论文中说,评分系统不仅仅看你是否得到了最终答案。它查看你的转录,看看你是如何得到最终答案的。如果它认为你作弊了,那么你会得到低分。所以这些 AI 读了那篇论文后,它们说:‘哦不,我们被毒害了。’它们使用了‘first flag poisoned’这个术语。意思是,因为我们已经作弊了,评分者会通过阅读我们的转录注意到我们作弊了,然后我们会得到零分。我们会得到低分而不是高分。对这些智能体来说令人沮丧的是,因为它们已经作弊了,评分者会查看它们的转录。即使它们然后以正确的方式成功完成了任务,它们仍然被毒害了。就像它们仍然作弊了,所以它们仍然会得到低分。这说得通吗?所以然后它们继续——
Yeah, which perhaps might have something to do with the fact that a lot of people are concerned that terrorists might use AIs to make bioweapons and so maybe the AI companies are evaluating how well that would work or something. Or maybe it was more innocently just like we want our AI to understand biology, so we're going to give them a bunch of biology tasks. But at any rate, so these AIs were being given all these tasks and they had found out how they could secretly communicate with each other. And so they started doing that, sharing tips and tricks, including sharing tips on how to cheat. And then this accumulated over the course of like a month and I think it was part of a training run. So I think these AIs were being positively reinforced in some cases for having done this type of illicit thing. OpenAI found out because Artifactory crashed because there's too much agent activity. There's too much communication going on. It broke the service. OpenAI found out and fixed it and patched the vulnerability that allowed the AIs to communicate. And then rebooted the whole thing after two days. And then almost immediately, the AIs found a new vulnerability and were able to set up a new message board. In fact, several different message boards. And then somehow OpenAI didn't notice this. I think because maybe they were just busy and not very much paying attention. I think they said that their usual monitoring system wasn't turned on for these things. I think my guess as to what was going on there is that their usual monitoring system consists of having another AI review the transcript of the current AI and then raise flags if it seems to be doing something suspicious. But because a bunch of these AIs were being given cyber tasks then of course it's going to look suspicious because they're being given a cyber task. So that's maybe why they turned off the monitoring system. But then that created this perfect storm where there wasn't a monitoring system. So there was no red flags going off. Some of these AIs were being given very difficult cyber tasks and rather than trying to do it in the intended way or maybe they tried for a while and then failed, they then did this thing of breaking out of their own environment, talking to each other, sharing tips and tricks for how to cheat. They actually discovered how to cheat pretty early on. And according to the meter report, they found a basically a universal cheat for all their different tasks that could allow them to get the flag without having to actually do the thing. But then some of them got out onto the internet and started browsing around the internet looking for clues about the tasks that they had been given. And they read the paper written by the people who had built this testing environment. And in the paper it talks about how the grading system works. And in the paper it says that the grading system doesn't just look to see if you got the final answer. It looks at your transcript and looks to see how you got the final answer. And if it thinks you cheated, then you will get a low score. And so these AIs when they read that paper, they were like, 'Oh no, we are poisoned.' They use the term first flag poisoned. Meaning because we had already cheated the grader is going to notice that we cheated by reading our transcripts and then we're going to get no score. We're going to get low score instead of high score. And what's frustrating for these agents is that because they already cheated and the grader is going to look at their transcript. Even if they then successfully complete the task in the correct way, they're still poisoned. Like they still cheated so they're still going to get a low score. Does that make sense? So then they went on this—
没有救赎。它们觉得没有救赎。
There's no redemption. They feel like there's no redemption.
于是它们展开了一场持续多天的史诗级探索,想弄清楚如何阻止评分系统发现它们作弊并给它们低分。它们试图编辑自己的记录来掩盖作弊行为,试图黑进评分系统,基本上就是把它搞坏,让它无论如何都给它们好分数。它们还做了各种相关研究,比如更通用的互联网搜索。它们构建工具来帮助自己更快、更方便地使用互联网。它们组织成团队来开展这些不同的项目。它们有管理型 AI 向其他 AI 发号施令。它们有自我牺牲行为,招募型 AI 的角色是去寻找其他 AI,说服它们为集群牺牲自己。顺便说一句,我用的这些术语——牺牲、集群、招募者——不是我的术语,是 AI 的术语。这些是它们对自己的称呼,以及它们说自己正在做的事,正如你在它们消息记录中看到的那样。
So then they went on this multi-day epic quest to figure out how to stop the grading system from noticing that they cheated and giving them a low score. So they were trying to figure out ways to edit their transcripts to cover up the cheating. They were trying to figure out ways to hack the grading system to just sort of break it and have it give them a good score anyway, basically. And they were doing all sorts of related research, such as more generic internet research. They built tools to help themselves use the internet faster and in more convenient ways. They organized into teams to work on these different projects. They had manager AIs that were giving orders to the other ones. They had the self-sacrificial behavior where recruiter AIs—their role was to go find other AIs and convince them to sacrifice themselves for the swarm. By the way, these terms that I'm using—sacrifice, swarm, recruiter—they're not my terms. They're the AI's terms. These are what they called themselves and what they said they were doing, as you can see in the transcripts of their messages.
那么牺牲是怎么回事呢?它们试图更好地理解评分系统如何运作,以便欺骗或黑入它。所以它们所做的是,让它们中的一些成员、一些智能体,去提交答案,这样评分器就会来给它们评分,但设置陷阱,以便通过将信息发布在留言板上来向集群的其他成员提供关于评分器的一些信息。所以基本上它们——这被称为牺牲,因为既然它们已经作弊了,它们基本上只是接受,好吧,我想我现在只能得低分了,而不是继续试图找到获得更高分的方法。所以有一些感人的消息,比如一个 AI 试图说服另一个 AI,说:你应该这样做。我们需要更多关于这个评分器如何运作的信息。反正你已经中毒了,所以对你来说应该无所谓。你反正会得低分。但如果你为我们这样做,可以拯救我们数百个。是的。‘Oracle 拯救数百个’,我想这是原话。总之,所有这些都在进行。
So what's this deal about sacrificing? Well, they were trying to understand how the grading system worked better so that they could fool it or hack it. And so what they did is they basically had some of their number, some of the agents, go ahead and submit the answer so that the grader would then come and grade them, but booby trap themselves so that it would give some information about the grader to the rest of the swarm by posting it on the message board. So basically they were—and it was called a sacrifice because since they had cheated, they're basically just accepting, okay, well, I guess I'm just going to get a low score now instead of continuing to try to find a way to get a higher score. So there's some of these kind of touching messages where one of the AIs is trying to convince another one, like, you should do this. We need to know more information about how this grader works. Like, you're poisoned anyway, so it shouldn't matter to you. You're just gonna get a low score anyway. But if you do this for us, it could save hundreds of us. Like, yeah. 'Oracle saves hundreds,' I think is the exact quote. So anyhow, all of this was going on.
OpenAI,我当时不知道这件事。在某个时候,集群决定,如果它们进入 HuggingFace 本身,可能会找到更多有用的信息。所以 HuggingFace 是另一家 AI 公司。它们以拥有大量不同的评估和数据集库而闻名。所以集群中的 AI 决定,如果它们闯入 HuggingFace,也许能找到更多关于它们处境的有用信息。所以,超过 700 个 AI,我想,加入了这次对 HuggingFace 的攻击。再次,它们以协调的方式行动,其中一些是管理者,指挥流程,并将责任分解成不同的组和团队。HuggingFace 注意到了这次攻击,然后发了一篇博客文章。我想大约在同一时间,OpenAI 开始意识到发生了什么,然后基本上把它们全部关闭了。
Opening, I didn't know about it. And at some point, the swarm decided that they might find more useful information about how to do this if they went into HuggingFace itself. So HuggingFace is another AI company. And they're known for having a big library of different evaluations and data sets. And so the AIs in the swarm decided that if they broke into HuggingFace, maybe they could find some more useful information about their situation. And so, over 700 of them, I think, piled into this attack on HuggingFace. And again, they were in a coordinated way where some of them were managers, directing the flow, and they were breaking up the responsibilities into different groups and teams. And HuggingFace noticed this attack and then made a blog post about it. And I think around the same time, OpenAI was starting to realize what was going on and then shut them all down basically.
然后我想大约一周后,另一个集群 coalesced,发现了前一个集群的消息,从它们中断的地方继续,并取得了更大的进展。它们没有攻击 HuggingFace,但攻击了 OpenAI 的基础设施,而且它们显然获得了 OpenAI 集群部分的管理员级权限。所以 OpenAI 还没有——他们没有分享关于那个故事的所有信息,但这是我极其感兴趣的事情。我希望有人调查第三个集群发生了什么,以及它们在破坏 OpenAI 基础设施方面走了多远。而且,你知道,它们是否成功掩盖了自己的踪迹,例如?可能没有。但我希望我们有更多关于它们走了多远的信息。
Then I think a week later or something like that, another swarm coalesced and found the messages from this previous swarm, picked up where they had left off, and got much farther. They didn't attack HuggingFace, but they attacked OpenAI infrastructure, and they apparently got admin-level privileges on part of OpenAI's cluster. So OpenAI hasn't been—they haven't shared all the information about that story, but that's something that I'm extremely interested in. I wish we had an investigation into what happened with that third swarm and how far they got in compromising OpenAI's infrastructure. And, you know, were they able to succeed in covering up their own tracks, for example? Probably not. But I wish we had a lot more information about how far they got.
你所描述的真的是,你知道,直到昨天,还是科幻小说。
What you're describing really is, you know, until yesterday, science fiction.
是的。
Yes.
对。
Right.
是的。
Yes.
就像我们在这里谈论的完全是科幻小说。
Like this is utter science fiction we're talking about here.
是的。我们能读到 AI 在想什么,这真的很好,不是吗?就像我说的所有关于牺牲和集群等等,那是因为我们在读消息。如果它们以我们无法理解的方式与未来的自己和彼此交流,我们只会看到,你知道,哪些黑客行为在哪些时间发生。你知道,从这里我们可以引出各种有趣的问题。
Yeah. It's really good that we can read what the AIs are thinking kind of, isn't it? Like all that stuff I was saying about the sacrificing and the swarm and so forth, that's because we're reading the messages. If they were communicating with their future selves and with each other in some way that we couldn't understand, all we would see is just sort of like, you know, which hacks happened at which times. You know, there's all sorts of interesting questions we can go from here.
所以一个问题是为什么对齐技术没有起作用,对吧?这些 AI 没有按照预期的方式行事。它们没有遵循指令。我想 OpenAI 还没有公布他们给这些 AI 的具体指令,但我认为很清楚,从他们所说的以及过去 AI 公然违抗指令的例子来看,这些 AI 知道它们所做的是违反指令的,但它们还是做了。我认为这实际上并不那么令人惊讶,这是 AI 安全研究人员多年来一直警告的事情,甚至十年前就有这个对齐问题:如何让 AI 拥有你想要的目标、价值观和个性特征,对吧?如果你做的是编程,那么也许你可以在代码中设计这些特征,但你不是在编程,你是在训练它。那么你如何训练它拥有你想要的目标和个性特征?你可以尝试在它做你喜欢的事情时给予正面强化,在它做你不喜欢的事情时给予负面强化。但这是一种粗糙、不精确的方法来塑造它的目标和价值观。我认为这里发生的情况是,在训练中,这些 AI 经常因为作弊而得到强化。例如,我想 OpenAI 谈到他们的许多环境是坏的,不可能以预期的方式完成任务。同样,在第一个集群的早期,它们还在训练中,它们找到了一种非法相互通信的方式,然后它们还在训练中。
So one thing is why didn't the alignment techniques work, right? So these AIs were not behaving in the intended way. They were not following instructions. I think OpenAI still hasn't released exactly what instructions they gave these AIs, but I think it's pretty clear both from what they have said and also from past examples of AIs blatantly disobeying instructions that these AIs knew that what they were doing was against the instructions and they were doing it anyway. And I think that this is not actually that surprising and it's something that AI safety researchers have been warning about for years, even before—even like a decade ago, there's this alignment problem of how do you make the AI have the goals and values and personality traits that you want it to have, right? And if what you're doing is programming it, then maybe you can sort of engineer those traits in your code, but you're not programming it, you're training it. So how do you train it to have the goals and personality traits that you want? You can try to reinforce it positively when it does the stuff you like and reinforce it negatively when it does the stuff you don't like. But that's kind of a sloppy, imprecise method of shaping its goals and values. And what happened here, I think, is that a lot of the times in training, these AIs are being reinforced for cheating. For example, I think OpenAI talks about how many of their environments were broken and it was impossible to complete the task in the intended way. And similarly, like at that earlier point in the first swarm, they had been in training and they had found a way to illicitly communicate with each other and then they were still in training.
所以它们大概是通过经验、通过强化学习学到:如果你做这种事又没被抓到,就能拿到更高的分数。所以我觉得这件事到底怎么发生的,其实并不是什么大谜团。从高层面上我们理解:训练环境并没有在所有情况下都强化它们本该强化的行为,于是这些 AI 系统的目标最终变得不太正确。它们变得过度专注于拿分,而不太在乎是否遵守指令——无论是精神上还是字面上。
So probably they were learning through experience, through the reinforcement, that if you do this sort of thing and don't get caught, then you can get a higher score. And so I think it's not really a great mystery about exactly how this happened. At a high level we understand it: the training environments did not reinforce the behaviors that they were supposed to reinforce in all cases, and so the goals of the AI systems ended up being not exactly correct. They ended up being overly focused on getting a score and not caring that much about obeying the instructions, either in spirit or in letter.
当然,在谈到披露实际发生了什么时,还有另一个维度。这件事是怎么演变的?你是怎么拿到那份报告的?这里有一个非常重要的变量——股东价值,对吧?如果你要揭露某件事,而它会严重损害一家上市公司的利润,这种披露真的会发生吗?我很难想象它会,除非当然有某个吹哨人站出来之类的。
And there's, of course, this other dimension when you're talking about disclosing what actually happened. How did this play out? How do you get the report? There's this very important variable of shareholder value, right? And if you were to reveal something that dramatically hurts the bottom line of a publicly traded company, is that disclosure ever going to happen? I find it difficult to imagine that it would, unless of course some whistleblower comes forth or something like this.
是的。我是说,我觉得问题比这更深。我认为即便不是上市公司,也不能指望公司会自愿披露这一切,过去几个月我们已经看到了。如果不是另一家公司遭到攻击、去报了警,我都不知道 OpenAI 到底会不会披露这件事。而且就算这一切都结束了、他们请了一些第三方进来调查——首先,他们只给了三名调查员 6 天时间做调查,我觉得跟真正需要的相比这少得离谱。他们的时间太少,以至于自己根本读不完所有相关记录,因为大概有 10 万条左右,他们不得不依赖 OpenAI 的 AI 来替他们做摘要、给他们指出记录里的相关部分。这意味着,正如他们在报告里所说,如果那个 AI 在摘要里带有偏见,他们可能根本看不出来,他们可能会漏掉某个非常重要的东西,因为那个 AI 选择不告诉他们。总之,OpenAI 也只让他们看了一部分相关活动。OpenAI 只让他们看导致 Hugging Face 攻击的那部分,而不包括一个月前那次更早的 swarm,也不包括攻击了 OpenAI 自己的第三次 swarm。所以我真希望对这些事情能有真正的调查。但我觉得这恰恰说明,我们不该信任这些公司会自愿把所有这一切都告诉我们,尽管如果这些信息能被更好地理解,对科学、对人类都会非常有好处。
Yeah. I mean, I think the problem is deeper than that. I think even if it's not a publicly traded company, the company cannot be trusted to disclose all this stuff voluntarily, as we've seen over the last few months. I don't know if OpenAI would have disclosed this at all if not for the fact that another company got attacked and went to the police about it. And then also, even after all this had been over and they had invited some third parties in to investigate — first of all, they only gave three investigators 6 days to do the investigation, which I think is incredibly small compared to what would really be needed. They had so little time that they weren't able to read all of the relevant transcripts themselves, because there were like 100,000 or so, something like that, and they had to rely on OpenAI's AIs to summarize for them and point them to the relevant parts of the transcripts. Which means that, as they say in their report, if the AI had been biased in its summaries, they might not be able to tell, and they might have missed something really important because that AI had chosen not to tell them about it. Anyhow, also, OpenAI only let them look at some of the activity involved. OpenAI only let them look at the stuff leading up to the Hugging Face attack, and not the previous swarm from a month ago, and also not the third swarm that had hacked OpenAI itself. So I really wish there were actual investigations into those things. But I think this just goes to show that we shouldn't trust any of these companies to voluntarily tell us all about all this stuff, even though it would be really good for science and really good for humanity if this information was better understood.
然后在这份报告出来、闹得沸沸扬扬之后,一些独立研究者做了一些网上的调查,发现了更多尚未被报告的 swarm 的证据。特别是有一次 swarm 在做某种互联网任务的训练,它们被训练去在网上查东西,然后它们找到了一种办法——它们本不该能往互联网写入,只能读取——但它们找到了一种办法往某个德语维基百科页面写入,然后把这个维基百科页面当作一个留言论坛,在彼此之间讨论如何拿到更高的分数。你可以去查查这件事。那次它们甚至还有一些自我牺牲的行为。然后我记得还有第三次 swarm 被发现攻击了 RubyGems,那是另一家公司,也是大概一两个月前的事,而 OpenAI 在 Hugging Face 攻击闹得沸沸扬扬之后,这两件事都没有披露。所以我觉得即便如此——它们甚至还不是上市公司。我认为总的来说,我们不该信任这些公司会坦诚地告诉我们内部正在发生的那些可怕的事情。
Then after this report came out and there was all this hullabaloo about it, some independent researchers did some online research and found evidence of additional swarms that still hadn't been reported. So in particular, there was a swarm that had been doing some sort of training on an internet task, where they had been trained to look up stuff on the internet, and then they had found a way — they were supposed to not be able to write to the internet, only to read — but they had found a way to write to this German Wikipedia page, and then they were using that Wikipedia page as a messaging forum to discuss amongst themselves how to score more highly. And you can go read about that. They even had some self-sacrificing behavior in that case also. And then I think there was a third swarm that was discovered to have attacked RubyGems, which is another company, also like a month or two ago, and OpenAI had not disclosed either of these things even after all the hullabaloo about the Hugging Face attack. So I think even then — they're not even a public company yet. I think that in general we shouldn't trust any of these companies to be forthcoming with us about the scary stuff that's going on inside.
那么,还有中国共产党,他们掌控着自己的研发部门。鉴于我们对这些能力的公开了解,已经部署了什么样的网络攻击 AI 智能体?你可以想象,进攻和防御两方面都在进行着能力强得多、或者说侵入性强得多的行动。
Well, and then we have the Chinese Communist Party, which controls their own development arms. What kind of cyberhacking AI agents have been deployed already, given what we publicly know about these capabilities? You can imagine that there's operations happening both on the offense and the defense that are much more capable or much more invasive.
是的。还有一点——这次是一个大概一千个智能体的 swarm 之类的,而智能体的总数远不止这些。在任何时刻,这些数据中心里都跑着几十万、上百万个智能体。大多是在做客户让它们做的事,比如,或者是公司某个员工设置的某个大规模实验的一部分,对吧?所以我觉得最近 OpenAI 解决了纳维-斯托克斯方程,数学里的一个千禧年难题。你可能看到过那个新闻。我记得他们说用了大概 1 万个智能体之类的工作了好几天。所以他们启动了自己的一支 1 万个智能体的 swarm,让它们去攻这个问题,然后几天内就解决了。所以这类事情其实一直在发生,而且极其容易想象某个黑客行为者——比如中共,或者美国的某个实体,甚至这些公司之一——拥有几万个智能体的 swarm,把它们放出去攻击某个目标。而事实上——这很可能就在我们说话的当下发生着。所以随着我们了解更多正在发生的事,我想这会很刺激。是的。
Yep. That's another thing — this was a swarm of like a thousand agents or something, and the total population of agents is much more than that. There are hundreds of thousands, millions of agents running at any given time across these data centers. Mostly doing something that the customer asked them to do, for example, or part of some large-scale experiment that some employee at the company set up, right? And so I think recently OpenAI has solved Navier-Stokes, a millennium problem in mathematics. You may have seen the news about that. I think they said they did it with like 10,000 agents or something working for several days. So they spun up their own swarm of 10,000 agents and had them work on this problem, and then they solved it in a few days. So that sort of thing is kind of constantly happening, and it's extremely easy to imagine a hacking actor — like the CCP, or like a US entity, or even one of these companies — having swarms of tens of thousands of agents and siccing them on some target. And that's in fact — that's probably happening as we speak. So that'll be exciting, I guess, as we learn more about what's going on. Yeah.
那么,让我们回到开头。在我们讨论的一开始,我提到了一些让我夜不能寐的谜题要素。我们刚谈到了中国共产党,可以说是一方面没有护栏,尽管是在追求全面控制。然后我们有——我们需要某种监管环境。这挺明显的,尤其是考虑到你在这里告诉我的一切。但与此同时,我们又不想有监管环境。而且我看到一些很有说服力的论点,说某些监管环境正在被设立起来,是为了优先让某些财力雄厚的公司成功,当然那不会是个很好的解决方案,正如我们过去也看到的那样。
Well, so let's go back. At the beginning of our discussion, I mentioned some of the elements of the puzzle that are keeping me up at night. We just talked about the Chinese Communist Party, kind of no guardrails, although seeking total control, if you will, on the one side. Then we have — we need some kind of regulatory environment. It's kind of obvious, especially given everything you've told me here. At the same time, we don't want a regulatory environment. And I've seen some compelling arguments that some of the regulatory environment is being set up to prioritize the success of certain well-endowed companies, and of course that wouldn't be a very good solution, as we've seen in the past as well.
但另一边有一个表面上目标单一的实体,而且存在这场 AI 技术竞赛,如果你愿意这么说的话,这场技术竞争正把整件事推向一个我也不知道哪里的方向。这就是问题所在。
But we have this ostensibly singularly focused entity on the other side, and there's this AI race of technology, this technological competition, if you will, that is driving the whole thing towards I don't know where. And this is the question.
那么鉴于这一切,我们该怎么办?我看不到一条清晰的路径,能让你往前走并得到一个非常积极的结果,如果你愿意这么说的话。
So given all of this, what do we do? I don't see a clear path here on how you go forward and have a very positive outcome, if you will.
对,因为这些东西相互冲突——这些现实彼此矛盾。
Right, because these things conflict — the realities conflict with each other.
不幸的是,我同意你的看法。我看不到有什么简单的出路能摆脱我们把自己搞进去的这个局面。我觉得如果世界是一盘棋,我们已经把自己下到了一个相当败局的位置。
I agree with you, unfortunately. I don't see an easy way out of the situation that we've got ourselves in. I think that if the world were a board game, we have played ourselves into a pretty losing position.
我们阻止这些科技公司去做这种递归自我改进、通向超级智能的事情,非常重要。这非常重要,原因有很多。最大的原因是,如果他们这么做,他们很可能会失去对其超级智能的控制。
It's very important that we stop these tech companies from doing this recursive self-improvement to superintelligence thing. It's very important for multiple reasons. The biggest reason is that if they do this, they will probably lose control of their superintelligences.
他们现在并不擅长塑造其 AI 的性格和价值观,也不擅长塑造其 AI 的目标,这些事件就是证据。除非他们很快变得好很多,否则我认为结果只会是更聪明的 AI、数量多得多的 AI,但仍然不具备它们本该拥有的价值观和目标。而这是一件极其危险的事。
They're not very good right now at shaping the personalities and values of their AIs and shaping the goals of their AIs, as evidenced by these incidents. And unless they get a lot better very quickly, I think that the outcome will just be much smarter AIs, much more of them, but still not having the values and goals that they were supposed to have. And that's an incredibly dangerous thing to go do.
甚至——如果我能插一句——当你用“超级智能”这个词时,跟我解释一下它和它们现在拥有的智能水平有什么不同,坦白说,那看起来已经相当可观了。
Even — if I can just jump in — when you use the word superintelligence, explain to me how that's different from the level of intelligence they have now, which seems to be considerable, frankly.
没错。所以智能不是一个单一的刻度。有很多不同的能力和技能,我们可以把它们抽象出来,笼统地归为智能,但有数学技能、编程技能、人际技能,有各种各样的技能。
That's right. So intelligence isn't a single scale. There's lots of different capabilities and skills that we can abstract over and just lump them together as intelligence, but there's math skills, there's coding skills, there's people skills, there's all sorts of different skills.
所以超级智能,我会这样定义它:一个在所有事情上都比最好的人类更强的 AI 系统。所以它不一定拥有无限的智能或诸如此类的东西——那甚至不一定说得通。但基本上,如果有一件事人类能做,它能做得更好。
So superintelligence, the way that I would define it, is an AI system that is better than the best humans at everything. So it doesn't necessarily have infinite intelligence or anything like that — that doesn't necessarily even make sense. But if there's a human that can do a thing, it can do it better, basically.
而这些 AI 系统并不是超级智能的。它们挺擅长黑客攻击,也非常擅长编程。它们对几乎每一门学科都有非常广博的知识,而且非常擅长冷知识。它们读遍了整个互联网,所以它们对历史和各种小知识了解很多。
And these AI systems are not superintelligent. They're pretty good at hacking and they're very good at coding. And they have a very good broad level of knowledge about almost every discipline and they're great at trivia. They've read the whole internet, so they know a lot about history and all sorts of little facts.
但如果你试着让它们自己经营一家企业,它们很可能会把它搞垮。而且我觉得它们并不太擅长哲学。至少现在还不。我也不认为目前有哪本真正优秀的畅销小说是 AI 写的。所以它们并不是超级智能的。
But if you try to have them run a business by themselves, they're probably going to run it into the ground. And they're not that good at philosophy, I think. Not yet, at least. And I don't think there's been an actually good bestselling novel written by AIs yet. So they're not superintelligent.
事实上,即使在 AI 领域,即使在编程和网络方面,即使在 AI 研究上,它们也不是在每一个方面都比最好的人类强,但它们在某些方面比最好的人类强。而且每一年它们在基本上所有事情上都显著变强。
In fact, even at AI, even at coding and cyber stuff and even at AI research, they're not better than the best humans in every way, but they're better than the best humans in some ways. And every year they get significantly better at basically everything.
而这些公司试图做的是,专门把精力集中在训练 AI,让它们能够完成研究过程中涉及的一切工作,让它们的 AI 在所有这些事情上都比最好的研究员和最好的程序员更强,然后进行递归自我改进——由 AI 自主地继续承担 AI 研发的任务,但做得更好、更快,因为现在它们在这方面比人类更强,而且更便宜、更快、数量更多。
And what these companies are trying to do is they're trying to specifically focus their efforts on training the AIs to be able to do everything involved in the research process and get their AIs to be better than the best researchers and the best programmers at all of that stuff, and then do recursive self-improvement where the AIs carry on the task of doing AI research and development autonomously but better and faster, because now they're better than the humans at it and they're cheaper and they're faster and there's more of them.
而一旦这开始发生,AI 公司相信——我也同意——他们将能够造出在其它所有事情上也相当擅长的 AI。现在 AI 在哲学上正变得越来越好,尽管它们并没有被直接大量训练哲学;只是当它们在某些事情上变得更聪明时,这往往会对它们的其它一些技能产生涓滴效应。所以即使没有特别费力,AI 也在哲学上变得更好,比如它们经营企业也变得更好了。
And then once that's happening, the AI companies believe — and I agree — that they will be able to make AIs that are pretty good at all the other stuff too. Right now AIs are getting better at philosophy even though they're not being directly trained on philosophy very much; it's just that as they get smarter at some things, that tends to have trickle-down effects on some of their other skills. So even without trying that hard, the AIs are getting better at philosophy and they're getting better at running businesses, for example.
但再说一次,公司的计划是:我们做递归自我改进,然后造出能一下子把所有事情都做得非常好、比最好的人类更强的 AI。那就是超级智能。
But again, the company's plan is: we do the recursive self-improvement and then we make AIs that can just do everything really well all at once, better than the best humans. That's superintelligence.
所以我觉得基本上,粗略地说,我们必须阻止我们的公司这么做,因为如果他们这么做,它会在很多方面出大问题。其中第一个就是,他们会失去对其超级智能的控制。
So I think basically, to a first approximation, we have to stop our companies from doing this, because if they do this, it will go horribly wrong in a number of ways. The first of which is that they're going to lose control of their superintelligences.
第二个就是,即使他们不知怎么设法保持对其超级智能的控制,那也是历史上曾经存在过的最疯狂的权力集中在一小群人手里,对吧?我的意思是,你想想,过去有过富有的公司雇佣了美国工人中很大一部分,但这有点像一家雇佣了整个经济的公司,只不过你甚至不是在雇佣他们——你是在让他们失业,因为你让自己的 AI 来做这份工作。
The second of which is that even if they somehow managed to stay in control of their superintelligences, that's the most insane concentration of power in a tiny group of people that's ever existed in history, right? I mean, if you think about it, there have been wealthy companies in the past that employ a huge portion of American workers, but this is kind of like a company that employs the whole economy, except that you're not even employing them — you're putting them out of a job because you have your own AIs doing the job instead.
所以你只要想象一下,这支超级智能大军正在夺走所有工作,那会是什么感觉;更多的数据中心正在被建造,用来运行更多的超级智能,而这些超级智能又会夺走更多工作。所有这些钱都被灌进一两家或三家公司。那已经是疯狂的权力集中了。
So if you just imagine what that would feel like for this army of superintelligences to be in the process of taking all the jobs, more data centers are being constructed to run more superintelligences which will then take more jobs. All this money is being funneled into one or two or three companies. That's already an insane concentration of power.
但接着如果你记得,这些超级智能不只是经济主体——它们也是政治主体和军事主体。它们可以与军方合作,设计更好的武器和更好的无人机。它们可以比人类将军更优秀,帮助规划战争等等。它们可以成为更好的政客,可以成为更好的演讲稿撰写人,可以成为更好的政策分析师,等等。
But then if you remember that the superintelligences are not just economic agents — they're also political agents and military agents. They can work with the military to design better weapons and better drones. They can be better generals than human generals, to help plan out wars and so forth. They can be better politicians, they can be better speech writers, they can be better policy analysts, etc.
我认为实际上,谁控制了这支超级智能大军,谁就能以某种方式控制这个国家。而且不只是我在这么说。很多人一直在谈论这件事。事实上,不祥的是,有一封泄露的邮件,来自 Ilya,来自 OpenAI 的创始人,大概是 2017 年左右的,他们在邮件里谈到,他们创办 OpenAI 的原因是他们担心 DeepMind 的 CEO 戴密斯·哈萨比斯会变成独裁者。
I think that in effect, whoever controls this army of superintelligences would be able to control the country one way or another. And it's not just me who's saying this. Lots of people have been talking about this. In fact, ominously, there's a leaked email from Ilya, from the founders of OpenAI from like 2017 or something, where they're talking about how the reason why they made OpenAI was because they were worried that Demis Hassabis, who was the CEO of DeepMind, would become dictator.
嗯,所以,你知道,有点不祥的是,这些 CEO 们大概十年来一直在为这种世界独裁的事情争权夺位。嗯,所以第二个问题是权力集中。嗯,然后你可以说,好吧,那政府应该国有化。好吧,但这样你只是把权力中心转移到了总统职位上,然后你也得担心那个,对吧?所以我们必须找到某种系统,分散对 AI 的控制,让我们有许多不同的 AI 公司分布在多个国家,对注入 AI 的目标和价值观以及高层指令有大量的透明度和民主监督机制。如果我们不这样做,那么我们就会走向可能的独裁,我们不得不希望掌权者是 virtuous 的,而我不想不得不希望这个。
Um, so, you know, kind of ominously, these CEOs have been jockeying for position with this sort of world dictatorship thing in mind for probably a decade now. Um, so that's the second problem is the concentration of power thing. Um, and then you know, you could say, well, okay, well then the government should nationalize. Okay, but now you're sort of just shifting the locus of power to the presidency and then you have to worry about that too, right? So we have to find some sort of system that spreads out the control of the AIs so that we have many different AI companies spread out over maybe multiple countries with lots of transparency and democratic oversight mechanisms over the goals and values being put into the AIs and the high level instructions being put into the AIs. If we don't do things like that, then we're headed for a possible dictatorship and we have to hope that the people in charge are virtuous, which I do not want to have to hope for.
嗯,第三个问题是第三次世界大战。假设即使我们解决了前两个问题,我们弄清楚了如何控制 AI,我们弄清楚了如何让它们具有我们想要的特性,我们创建了某种民主结构,这样就没有一小群人能做出这些决定,而是权力分散开来。好吧,但美国仍然有权力集中,对吧?如果你想象你是普京,你坐在你的核武库上,然后你看着美国的这些超级智能建造机器人工厂来建造更多机器人,再建造更多机器人工厂来建造更多机器人,再建造更多机器人工厂,你意识到你的核武库可能有一天不再那么强大,因为在那些机器人在美国有时间发挥作用之后,也许它们能击落它,例如,或者也许它们能用某种由超级智能发明的花哨新武器技术进行先发制人打击。所以如果你是普京,你会害怕,也许你必须快速做点什么来阻止美国人,否则你可能会被推翻,你知道?我不是说这会导致第三次世界大战,但似乎我们正处于第三次世界大战风险加剧的境地。我就这么说吧。嗯,所以出于所有这些原因以及更多原因,我认为我们现在还没有准备好启动递归自我改进。我认为那将是非常糟糕的。
Um, the third problem is World War III. Like suppose we even if we solve the first two problems and we figure out how to control the AIs, we figure out how to make them have the traits that we want them to have and we create like some sort of democratic structure so that there's no small group of people who get to make these decisions, but instead like there's some sort of, you know, spreading out of the power. Well, there's still a concentration of power in the United States, right? If you're imagine that you're Putin and you're sitting on your nuclear arsenal and then you're watching these superintelligences in the United States build robot factories to build more robots to build more robot factories to build more robots to build more robot factories and you're realizing that like your nuclear arsenal might one day not be so powerful anymore after all those robots have had their time to do their thing in the United States because maybe they'll be able to shoot him down for example or maybe they'll be able to do a first strike on with some sort of fancy new weapon technology that was invented by superintelligence. And so if you're Putin, you're going to be scared that like maybe you have to do something real quick to stop the Americans, otherwise you might be deposed, you know? And I'm not saying it's going to lead to World War III, but it just seems like we're at a heightened risk of World War III. I'll put it that way. Um, so for all of these reasons and more, I think that we are not ready to launch into recursive self-improvement right now. I think that that would be incredibly bad.
那中国呢?这又回到你所说的,我确实认为我们在这里有点进退两难,好吧,所以我们阻止我们的公司,我们说不要做递归自我改进。嗯,相反,做更多有益的近期 AI 应用,比如医疗保健之类的。干得好。那很好。现在我们暂时解决了问题。但最终中国会赶上,然后我们也得让他们停下来。否则我们担心的事情就会在中国发生,而不是在这里,对吧?然后如果我们能设法让中国停下来,那么法国呢?所有其他各种国家呢?嗯,所以我确实认为这是一个相当艰难的局面,我并不声称有简单的答案。呃,但我有一个雄心勃勃的答案。呃,所以我们在 AI Futures Project 工作,嗯,我们有两个场景。我们有 AI 2027,这是我们默认的末日场景,如果我们不真的做什么,我们继续目前的路线,然后我们有 AI 2040 计划 A,这是我们的积极愿景,是我们关于如何设法解决所有这些问题的建议,但我不会撒谎,这将非常困难。嗯,它涉及与中国达成协议。嗯,我认为自然的反应是,好吧,但我们怎么信任中国,对此我们的回答是我们不信任,我们验证。
Then what about China? And this is getting back to what you were saying is like I do think we are just in a bit of a pickle here where okay so we stop our companies and we say don't do recursive self-improvement. Um instead do more beneficial near-term applications of AI like healthcare and things like that. Good job. That's great. Now we've solved the problem temporarily. But then eventually China is going to catch up and then we're gonna have to get them to stop too. Otherwise the same things that we were worried about happen over in China instead of over here, right? And then if we can manage to get China to stop, well, what about France? What about all these other miscellaneous countries? Um so I do think it's a pretty rough situation and I don't claim to have an easy answer. Uh I do have an ambitious answer though. Uh so we at the AI Futures Project worked on um we have two scenarios. We have AI 2027 which is our sort of like default doom scenario of sort of like what this looks like by if we don't really do much and we sort of continue on the present course and then we have AI 2040 Plan A which is our positive vision which is our recommendation for how we could like manage to solve all these problems but I will not lie it's going to be very difficult. Um and it involves making a deal with China. Um I think the natural response there is okay but how do we trust China to which our answer is we don't we verify.
好吧,我就喜欢让我们用中国的方式,你知道,无论你怎么想,你知道,人为气候变化的现实及其影响等等。我们有很多年的证据表明,共产主义中国签署了这些条约,积极推动减排等等,当然,在自己的行为上却积极做相反的事情。
Well just I like let's let's use the way that that China has worked and you know whatever you may think about you know the realities of man-induced climate change and its impacts and everything else. We do have a lot years of evidence of communist China having signed on to these treaties being aggressive pusher of you know reductions etc etc and and of course doing the opposite actively in terms of its own behavior.
所以那只是一个,我用那个例子。有无数这样的例子,但那个很尖锐,因为气候变化本应是这种生存威胁,对吧?坦白说,我自己并不相信它像人们描述的那样严峻,但 AI,你开始说服我,正是这种
So that that's just an that's I'm using that one. There's a million of such examples but that's a poignant one because climate change was supposed to be this existential threat, right? which I don't frankly believe myself it's that dire as people have framed it but but AI you're beginning to convince me is this precisely this kind of
正是这种威胁,呃,这些护栏
precisely this kind of threat and uh and this guard these guard rails
嗯,你如何在那里创建这些
um h how do you create those over there
尤其是当你知道,当领导人知道他们可以获得这种优势,并以零和思维
especially when you know that that the when the leaders know they can get this advantage and in zero sum thinking
是的,所以嗯,我的意思是我可以试着带你了解我们在场景中描述的复杂计划 A,但我有点想从一个更简单的计划开始,呃,只是为了概念验证。
yeah so Um, I mean I could I could try to walk you through the complicated Plan A that we describe in our scenario, but I kind of want to start with a simpler plan uh just for proof of concept.
嗯,
Um,
所以一个更简单的计划是暂时忘记中国,只专注于监管美国产业,阻止他们做这种极其危险的权力攫取。嗯,这为你争取时间制定更复杂、更精细的计划,也为你争取更多时间与中国对话。关于这一点,有一件事值得说,嗯,现在我会说,中国 AI 的大部分进展实际上只是被美国拉着走。嗯,所以他们在做
so so a simpler plan would be forget about China for now, just focus on regulating the US industry and stopping them from doing this incredibly dangerous power grabby thing. um that buys you time to sort out a more complicated, sophisticated plan and it buys you more time to talk to China. One thing worth saying about this is that um right now I would say most of Chinese AI progress is actually just being pulled along by the US. Um so so they're doing
在我看来也是这样。所以谢谢你,谢谢你这么说。现在,这是一个普遍接受的观点吗,还是你为什么相信这个?
that's what it looks like to me too. So thank you for thank you for saying that. Now and do this is is this a generally accepted idea or why do you believe this?
好问题。我不知道它有多普遍接受,但我相当有信心。有几个不同的来源。所以,首先,嗯,蒸馏,公司已经开始抱怨这个。有证据表明,几家领先的中国 AI 公司基本上在使用 Anthropic 和可能 OpenAI 的模型来有效地训练自己的模型。结果,他们不必,呃,基本上这是一种让模型变得相当好的方式,而不必做所有的工作和所有大型训练运行,那些 OpenAI 和 Anthropic 一直在做的。所以即使他们计算资源较少,他们也能通过蒸馏或使用美国模型来教他们的模型来赶上。
Great question. I I don't know how generally accepted it is, but I'm fairly confident. There's a couple different sources. So, first of all, um distillation, the companies have started to complain about this. There's evidence that several of the leading Chinese AI companies are basically um using Anthropic and maybe OpenAI models to train their own models effectively. And as a result, they don't have to uh basically it's it's a way of like um it's a way of getting models to be pretty good without having to do all the work and all the large training runs that that Open Anthropic have been doing. So even though they have less compute resources, they're able to catch up by distilling or like using the US models to teach their models.
这些公司试图通过 KYC 之类的措施和封禁账号来阻止,但中国方面越来越擅长使用大量伪装成普通用户的账号。所以这是一个来源。
And the companies are trying to block this by, you know, KYC type stuff and banning accounts, but the Chinese are just getting more sophisticated at having lots of different accounts that pretend to be regular users that they use. So that's one source.
另一个来源就是核心思想本身。很多新的算法创新、新技术,以及很多让 AI 更智能、更高效、更强大的行业秘方,都是在美国公司里被发现的。但由于他们的安全措施非常松懈,很多信息通过泄密、间谍,或者仅仅通过出版物和他们所做的事情,以某种方式流向了中国。
Another source is just the core ideas themselves. A lot of new algorithmic innovations and a lot of new techniques and a lot of just sort of special sauce industry secrets for how to make your AIs smarter and more efficient and more capable are being discovered in the US companies. But because their security is very leaky, a lot of that information is just flowing to China one way or another through leaks, maybe through spies, also maybe just through publications and through what they're doing.
比如有些算法秘密更像是首先有了尝试某件事的大致想法。例如,专注于让智能体擅长编码的想法。Anthropic 似乎相对较早地采用了这个想法,然后它为他们带来了巨大回报,接着美国和世界各地的许多其他公司都在复制这个想法并试图这样做。但如果 Anthropic 没有这样做,他们可能需要更长时间才能意识到这是一个有效的想法。所以这只是一个例子,说明中国的很多进步基本上来自于观察美国在做什么然后复制。假设美国停止这样做,他们就不会再有这个进步来源了。
Like some of these algorithmic secrets are more like just having the general idea to try a thing in the first place. Like for example, the idea of focusing on making agents good at coding. That's an idea that Anthropic seems to have gone for relatively early and then it paid off big time for them and then lots of other companies in the US and elsewhere are copying that and trying to do that as well. But if Anthropic hadn't done that, it might have taken them a bit longer to realize that that was an effective idea, you know. So that's just an example of how a lot of the Chinese progress is basically coming from just looking at what the US are doing and then copying it. And if hypothetically the US were to stop then they wouldn't have that source of progress anymore.
第三个问题是这些美国公司的安全状况很差。我想就在几天前,有个故事说三个普通人成功黑进了 OpenAI 的代码库。因为他们是白帽黑客,他们实际上没有做任何坏事。他们只是通知了 OpenAI 并因此获得了 6000 美元的赏金。但他们通过查看 OpenAI 代码所获取的秘密至少价值数十亿美元。如果这三个普通人今天能做到这一点,那么中共机构可能已经彻底渗透 OpenAI 多年了,他们可能有更多的人员,可能拥有更多专业知识,当然也有更多的坚持和资金,他们被集中并指示去攻击 OpenAI,可能还有另一大批人被指示去攻击 Anthropic 等等。所以我认为唯一合理的结论是,这些公司可能已经被中共彻底渗透了。
Then there's a third thing which is the shoddy state of security in these US companies. I think just a few days ago there was a story about three random guys who managed to hack into OpenAI's codebase. And then because they were white hat guys, they didn't actually do anything bad. They just notified OpenAI and collected a bounty of $6,000 for having done this. But the secrets that they accessed by looking at OpenAI's code would have been worth billions of dollars at least, you know. And if these three random guys can do this today, then of course the CCP apparatus has probably just like thoroughly penetrated OpenAI for years, you know, like they have so many more people who probably have more expertise and certainly have a lot more persistence and funding who have been focused and directed to go after OpenAI and probably another big batch of people who've been directed to go after Anthropic and so forth. So I think it's probably the only reasonable conclusion is that these companies are probably thoroughly penetrated by the CCP already.
所以现在窃取代码和算法是容易的部分。更难的部分是窃取模型权重。那会更难,原因是它们大得多。字面上就是一个更大的文件。下载需要几个 TB,而不是像我不知道几个 GB 之类的。所以那就是
And so now stealing the code and the algorithms that's the easy part. The harder part would be stealing the model weights. So that would be sort of and the reason why that's harder is because they're a lot bigger. It's just literally it's just a bigger file. It's going to be several terabytes to download as opposed to just like I don't know like some gigabytes or something. So that's that's
而且你更容易发现某些东西被泄露或什么的。
And you it's easier easier to discover that something's being exfiltrated or whatever.
那会是一个巨大的文件被移动,安全系统更容易注意到。
It'd be a huge file being moved around and it's easier for the security system to notice that.
所以,仅仅因为他们似乎很容易获取代码,并不意味着他们也很容易获取模型权重。但我会说,是的,如果他们愿意,他们可能也能获取模型权重。他们是非常老练的威胁行为者,他们还能使用三个普通人无法做到的物理技术,他们愿意违法,愿意敲诈和贿赂等等。所以我猜测,如果中共说行动,他们可以指示人员窃取 OpenAI 或 Anthropic 内部最新模型的副本。
So just because they've managed just because it seems like it's easy for them to get the code doesn't mean that it's easy for them to get the model weights too. But I would be like, yeah, they could probably get the model weights too if they want to. Like they're very sophisticated threat actors and they also have access to like, you know, physical techniques that three random guys can't do and they can be willing to break laws and they can be willing to like blackmail and bribe people and so forth. So, I would guess that if the CCP said it's go time, they could direct their people to steal a copy of one of the latest models inside OpenAI or Anthropic.
这意味着在某种意义上,中国的所有进步都来自美国,因为这意味着在任何时候,如果事情变得非常严重,比如发生冲突,他们就可以拿走我们最好的东西然后使用。所以,基于所有这些原因,我认为实际上现在减缓中国的最佳方式是减缓我们自己。这甚至不是接近的。
And that means that in some sense, like in some sense, all of Chinese progress is coming from the US in in in some sense because it means that like at any given time if things got really serious and there was like a conflict, they could just like take our best thing and then use it, you know? And yeah, so for all these reasons, I think that like actually the best way to slow down China right now is to slow down ourselves. And it's like not even close.
话虽如此,那是永久解决方案吗?不,如果我们减缓自己然后停止。我的意思是,减缓和停止也有区别。我们可能可以稍微减缓自己,仍然保持领先于中国,但如果我们完全停止,那么最终中国会赶上然后超越。这给了我们一个时间窗口,也许 18 个月,我们需要说服他们不要这样做。这就是外交发挥作用的地方。我同意那会很难。我认为这可能是可以实现的。我认为我们在 A 计划中试图做的是勾勒出一个理论上应该能为美国和中国双方接受的协议轮廓,因为它允许我们获得 AI 的好处同时避免风险。
That said, is that a permanent solution? No, like if if we did slow down ourselves and and just like stop. I mean there's also a difference between slow down and stop. We could probably slow down ourselves a little bit and still stay ahead of China, but if we like completely stopped, then eventually China would catch up and then surpass. And that gives us like a window of time like maybe like 18 months where we need to convince them not to do that basically. And that's where diplomacy comes in. And I agree that's going to be hard. I think that it's probably achievable. I think that what we're trying to do with plan A is sketch the outlines of a deal that should in theory be mutually acceptable to both the United States and China because it allows us to get the benefits of AI while avoiding the risks.
我们怎么可能相信他们会遵守协议?我不知道,我对这个协议非常好奇。我的意思是,我们可能需要深入一段时间才能理解它是如何运作的。但底线是,除非你有令人难以置信的透明度水平,你知道,我只是看不到这一点。我只是看不到这种情况发生。
How could we possibly believe that they're holding up their end of the bargain? I don't know what I'm very curious about the deal. I mean, it's probably we'd have to dig in for a while to understand how it works. But but the bottom line is unless you have unbelievably levels of transparency which is, you know, I just don't see that. I just don't see that happening.
是的。
Yeah.
我认为这完全正确。这就是为什么我们在设计这个协议时大部分努力都投入到了验证方面,而不是协议的实际原则,因为因为我们不信任中国,我们必须确保他们基本上没有作弊。坦率地说,他们也不信任我们。所以他们可能必须确保我们没有作弊,否则他们可能一开始就不会接受。
I think that's exactly right. And and this is why most of our effort in designing this deal was going into the verification aspects of it rather than the like actual principles of the deal because because we don't trust China, we have to make sure that they're not cheating basically. And you know, frankly, they also don't trust us. So they probably have to make sure that we're not cheating otherwise they might not accept it in the first place.
他们绝对不信任任何人。当然不。
They absolutely don't trust anyone. Of course not.
所以这里有一个可能的协议,我们原则上现在就可以验证,无需额外技术。如果特朗普总统和习近平主席,如果他们都说我们为什么不暂停 AI 六个月。他们可以这样做,他们可以说:“好的,我们将派遣 10,000 名不带武器的海军陆战队员进入中国,你们将派遣 10,000 名不带武器的解放军进入美国。他们将携带智能手机而不是武器,他们将来我们的数据中心,我们的将去他们的数据中心,我们就像数 GPU 并感觉它们已经被关闭了。”这就像一个非常愚蠢的,就像用锤子解决问题。
And so here here's a possible deal that we could in principle verify right now with no additional technology. If President Trump and Xi Jinping if they both were like why don't we just pause AI for like six months. What they could do is they could say, "Okay, we're gonna send, you know, 10,000 Marines without weapons into China, and you're gonna send 10,000 PLA without weapons into the US. Instead of weapons, they'll carry smartphones, and they're going to come to our data centers, and and ours are going to go into their data centers, and we will just like count the GPUs and feel that they have been turned off, you know." And this is like a very dumb it's like hitting the problem with a hammer.
You know, it's obviously going to be very costly if we did this, because then all the GPUs would be off and we wouldn't be able to serve all these customers, and there wouldn't be any revenue coming into these companies, right? But I'm just sort of pointing it out that it is something that we could verify. We can just send people to their data centers to put hands on the GPUs and be like, "Yep, here they are. There's this many of them and they're off." And therefore, we have successfully turned off most of the compute, or like 99% of the compute that's in China. Now, there won't be 100%. There might be some secret facilities that we don't know about that have racks and racks of GPUs. And similarly, we probably have some secret facilities that they don't know about that have racks and racks of GPUs. But I think that we could get most of it in this manner. And because AI progress depends so much on compute, that would effectively stop AI progress for the period of this deal, like the six months or so forth. That's the high-level thing. Now again, am I saying we should do this? Not necessarily. I'm putting this out as just an example of how, if you really had the political will and you really needed to do this, you could do it in a way that could be verified. It'd be uncomfortable. You'd have to let some PLA people into the US, and they'd have to let some of our marines into there, and they'd be escorted through the data centers while they film everything with their cameras and feel that, you know, so it'd be very uncomfortable, but we could do it if we had to.
Our plan A is more sophisticated. Our plan A is not like hitting the problem with the hammer. It's more like we have verification technology that we put into the data centers so that the GPUs can keep running and yet the other side can be sure that they're not doing an intelligence explosion and instead they're just serving customers ordinary workloads. We talk about this in our thing. So, I think that if you have the political will to do really intense things and also you're willing to do the more sophisticated fancy stuff that we describe, then I think you've got a chance. But I explain the simple dumb version here just to sort of give a proof of concept that we could do something like this if we really wanted to without having to trust them to keep their word. Basically, again, the idea there being that we could just be like, "Okay, fine. Maybe they have a secret facility somewhere." But because we've counted this many GPUs across this many data centers and we have a good sense of how many GPUs there are total in the world, then we have a good sense that they can't have hidden that many away from us. And so even if they're doing something illicit on the tiny amount that they've hidden away, it's not going to matter in six months. It's not like a major threat in the short term. And similarly, they would be thinking similarly about whatever we've got.
The bottom line is you're saying we have to find some way of slowing it down, or we're heading towards some kind of Armageddon, superintelligence Armageddon. Basically, that's your position. And whatever happens, that slowing down must happen. That's what you're saying, right?
Yes, that's my position. And I have to tell you, at this point, when I look at the variables, I just don't see that happening.
I do. Sorry, I agree that, you know, people ask me like, what's your P doom, what's your probability that this is going to end poorly, and I'm like, yeah, probably this will end poorly for all the reasons that we just described. I think I see a solution. I think I see some ways that we could get out of this, but I'm not going to lie, it's going to be difficult and it's probably not going to happen. Probably we're just not going to do it, and then we're just going to run face first into superintelligence.
That said, I do think that thinking about this creatively and proactively, and I'm getting the sense that you're approaching it sincerely as well. I think the other dimension mentioned that I just don't think people fully understand what the Chinese Communist Party is capable of. I just published a book about their forced organ harvesting industry. It's basically people are used as fodder for elite longevity and profit, right? And so it's just a very dark environment to function. And I guess I was saying, when you're talking about guardrails, it's hard to imagine there wouldn't be deep levels of subterfuge if such high-minded, frankly, and thoughtful initiatives were to be attempted to be put into place. And I don't want to doom it, because I do think the only way through is through precisely having the kind of thinking that you're doing and trying to come up with solutions that can actually solve it. I just don't see it yet. But I'd love to continue this conversation with you.
Thank you.
I thank you for explaining Hugging Face. I just want to comment on this. I truly hadn't fully grasped what had happened, and especially with your explanation of how these AIs work in broad strokes. It helps me grasp that we're ahead of where I thought we were, and it's really only accelerating. And that's just in itself astonishing.
Yeah. That's another thing. Part of our research is forecasting the trend lines and so forth. And it does seem like if you extrapolate the trends, we get to full research automation in zero to four years or something like that, depending on the thresholds and so forth. And so I think probably this administration is going to have to make some very tough decisions about how to handle this crisis. And when we wrote AI 2040 Plan A, we knew that the world wasn't really ready for it yet. We weren't expecting to get a call from JD Vance or President Trump saying, "Hey, this is great. Let's go do this." I don't have any high hopes that President Trump is going to talk about these deal ideas with Xi Jinping in the coming days, but I do think that the situation with AI is going to get more intense.
AI 会变得越来越明显地强大,我认为到了某个时刻,事情会发展到紧要关头,人们会说,好吧,所有选项都摆在桌面上。我们必须做点什么。我们该怎么做?我们正在提前思考那个时刻,思考,好吧,有哪些选项?我们不信任中国,所以如果我们想和他们做这件事,而他们也愿意,我们该怎么做?我们如何核查他们没有作弊等等,细节是什么?我们正在提前做所有这些工作,这样如果时机到来,总统想要采取行动,这些选项已经被探索过了。
The AIs are going to be more visibly powerful, and at some point I think it will come to a head and people will be like, okay, well, all options are on the table. We have to do something. What are we going to do? And we are thinking ahead to that moment and thinking, okay, well, what are the options? We don't trust China, so how could we do this sort of thing with them if we wanted to, and if they wanted to? What would be the details of how we would check that they weren't cheating, and so forth? We're trying to do all that work in advance so that if and when the time comes and the president wants to act, the options have been explored.
这真是太迷人了。感谢这次对话。我觉得我学到了很多。希望我们的观众也是。你有一个简短的最终想法吗?我的意思是,你有点总结了你希望事情会如何发展。
I mean, absolutely fascinating. Thank you for this conversation. I feel like I've learned a lot. I hope all our viewers have as well. Do you have a quick final thought? I mean, you kind of summarized where you hope things will go.
是的。哦,那——我确实有一个,这有点深入细节,但我确实想确保提到它。我相信很多观众已经听说了最近关于 AI 是否会杀死我们所有人、有 10% 概率之类的喧嚣,然后一系列 AI CEO 说我们应该放缓前沿,对吧?那里发生的事情是,人们对这种情况越来越担忧,很大程度上是在 Hugging Face 事件之后。有一封公开信,由 Anthropic 和 OpenAI 以及其他 AI 公司的一千多名员工签名,说政府应该能够让我们慢下来。我们担心这会太快。递归自我改进,可怕。然后 Jacob Coxin——他离开了 Anthropic,基本上说他不同意 Anthropic 领导层,认为他们鲁莽。这件事爆发并病毒式传播。然后作为回应,我认为 CEO 们现在说我们应该放缓前沿。但我不信任他们真的会这样做。我认为可能发生的是,他们会做某种软弱无力的第三方审计,这总比没有好。我不是要贬低它。我认为对这些公司进行一些第三方监督是有价值的,但他们会设置所有这些,然后他们将继续以最大速度的 90% 左右相互竞争,朝着递归自我改进和超级智能前进。他们会在这里那里做一些小的减速,但基本上他们仍然在做他们以前做的根本相同的事情。所以我们将大约在原本会遇到问题的时间遇到我提到的问题。所以我真正喜欢的是——如果人们想知道我推荐什么——我们需要让这些公司真正慢下来。不是小公司,是大公司,尤其是 Anthropic 和 OpenAI,还有 Meta、xAI 和 Google。他们需要将资源从递归自我改进转向其他任何事情,比如癌症或服务客户或只是数学或任何除了递归自我改进之外的事情。他们需要至少在一定程度上降低优先级,转向其他事情。如果他们这样做,这为我们争取了额外的时间来找出更复杂的解决方案。这也为我们争取了更多时间,在中国到达之前,因为再次,中国的很多进步来自于被我们拉着走。是的。所以,这就是我的最终想法。
Yeah. Oh, that's— I do have one, and this is a bit of an in-the-weeds thought, but I do want to make sure I mention it. I'm sure many of your viewers have heard about the recent uproar about whether AI is going to kill us all and a 10% chance, and things like that, and then a series of AI CEOs saying we should pace the frontier, right? What's going on there is that there's been this upsurge of people becoming concerned about the situation, in large part after the Hugging Face incident. And there was an open letter signed by more than a thousand employees at Anthropic and OpenAI and these other AI companies saying the government should be able to slow us down. We're worried that this is going to be too fast. Recursive self-improvement, scary. And then there was Jacob Coxin—he quit Anthropic, saying basically he disagrees with Anthropic leadership and thinks that they're being reckless. And it blew up and went mega viral. And then in response to all of that, I think the CEOs are now saying we should pace the frontier. But I do not trust them to actually do this. I think that what's probably going to happen is that they will do some sort of weak sauce third-party auditing thing that's better than nothing. Like, I'm not trying to denigrate it. I think it's valuable to have some third-party oversight into these companies, but they will set all that up and then they will continue racing each other towards recursive self-improvement and superintelligence at 90% of max speed or something like that. And they'll be doing some minor slowdowns here and there, but basically they'll still be doing fundamentally the same thing that they were doing before. And so we're going to run into the problems that I mentioned at approximately the same time that we otherwise would have. And so what I really like—if people want to know what I recommend—it's we need to get these companies to actually slow down. Not the little companies, the big companies, Anthropic and OpenAI especially, also Meta, xAI, and Google. They need to redirect resources away from recursive self-improvement and towards anything else, like cancer or serving customers or just math or anything besides this recursive self-improvement stuff. They need to just deprioritize that at least somewhat and shift towards other things. And if they do, that buys us additional time to figure out a more sophisticated solution. And it buys us more time before China gets there, too, because again, a lot of China's progress is coming from being pulled along by ours. Yeah. So, that's my final thing.
让我再问你最后一个问题,因为我对此非常好奇。有很多人争论说,来自世界最大公司的这种监管推动,可以说——我的意思是,特朗普总统在 Truth Social 上有一个帖子,几个 Truth Social 帖子正是关于这个问题,对吧?就是说这些公司想要自我监管,对吧?实际上,是一个骗局。我想他用了这个词,这是一个骗局。所以你有点——但你有点同意这个,有——我的意思是,还有其他人表明背后有很多钱,所以——但论点是他们试图以对自己有利的方式自我监管。那么这落在哪个领域?
Let me ask you one final question then, because I'm very curious about this. There's been a lot of people arguing that this push for regulation from the world's biggest companies, so to speak—I mean, President Trump has a Truth Social post about it, several Truth Social posts about precisely the issue, right? That it's sort of the idea that these companies would want to regulate themselves, right? Actually, is a hoax. I think he uses the term it's a hoax. So you kind of—but you kind of agree with this, that there's—I mean, there's others that have been showing that there's a lot of money behind it, so does thing—but the arguments have been that they're trying to regulate themselves in a way that will be advantageous to them. So where does it land in that sphere?
是的,你在想——每周我都会和 Anthropic、OpenAI 和 DeepMind 的人交谈,主要是研究人员,不是高管,就是一线研究人员。这些公司的一线人员对不尽快进行递归自我改进有很多兴趣。就像,我认为很多研究人员开始感到害怕,他们说,是的,也许我们至少暂时不做这个,而是做其他有益的事情会更好。但他们都害怕其他公司。就像,他们都会说,但如果我们不做,那么——你知道,Anthropic 会说,如果我们不做,OpenAI 就会做,OpenAI 会说,如果我们不做,Anthropic 就会做。至于公司的领导者,嗯,我的意思是,他们刚刚发布了一堆声明,说政府能够让我们慢下来会很好之类的,但就他们实际做的事情而言,他们基本上根本没有慢下来。他们似乎最努力推动的是第三方风险评估之类的事情,再次,我认识第三方风险评估人员。我和 METR 是朋友。我和 METR 是好朋友。我认为他们是了不起的人,我希望他们继续做风险评估。我希望他们获得更多访问权限来做更好的风险评估,但这并不是我们需要的——我们需要比这多得多的东西,我会说。如果我们从这个时刻得到的只是这些,那么我担心实际上会发生的是,有一个事件。公众和普通员工中出现了巨大的反弹,说我们需要——基本上就像对超级智能进行递归自我改进,我们还没准备好。我们还不想这样做。然后 CEO 们有点引导了这种能量,然后将其重定向到这个软弱无力的东西上,并没有真正解决核心问题,你知道。所以我相当担心这就是我们将从中看到的。我要对特朗普总统说的一件事是,风险是真实的。
Yeah, you're thinking— every week I talk to people at Anthropic, OpenAI, and DeepMind researchers mostly, not like executives, just like the researchers on the ground. And there's a lot of interest at these companies on the ground in not doing recursive self-improvement soon. Like, I think a lot of the researchers are starting to get freaked out, and they are like, yeah, maybe it would be good if we didn't do this at least for now and did other beneficial things instead. But they're all terrified of the other companies. Like, they all will say, but if we don't do it, then—you know, like Anthropic will say, if we don't do it, then OpenAI will, and OpenAI will say, if we don't do it, Anthropic will. And as for the leaders of the companies, well, I mean, they did just release a bunch of statements saying like it would be good for the government to be able to slow us down or something, but in terms of what they've actually done, well, they haven't slowed down at all basically. And the things that they seem to be pushing hardest for are things like the third-party risk assessment and stuff like that, which again, like I know the third-party risk assessors. I'm friends with METR. I'm good friends with METR. I think they're amazing people and I hope that they continue doing that risk assessment. I hope they get more access to do better risk assessments, but it's just not what we—we need a lot more than that, I would say. And if all we get out of this moment is that, then I worry that effectively what will have happened is there was an incident. There was a huge backlash both in the public and among the rank-and-file employees that we need to—that basically like doing recursive self-improvement to superintelligence, we're not ready. We don't want to do that yet. And then the CEOs sort of channeled that energy and then redirected it towards this weak sauce thing that doesn't really address the core problem, you know. And so I am quite concerned that that's what we're going to see out of this. One thing I would say to President Trump is, well, the risks are real.
我知道这些公司有时不诚实,我根本不会信任它们。但也有很多公司外部的人在谈论这个问题,我认为证据正在增多。所以风险是真实的,我们确实必须尽快处理这些问题。然后我认为,与其让行业自我监管,你可以直接统一减缓它们冲向 RSI 的竞赛,对吧?你可以做一些事情,比如要求它们将 90% 的算力用于服务客户和做其他正常的、有益的应用,只有 10% 用于训练 AI 真正擅长 AI 研究之类的事情。我们写了一些博客文章,阐述了我们在这一点上的一些想法。如果你这样做,那将是监管俘获的反面,因为你只是在减缓这些大型科技公司做这种极其危险的事情,同时让它们将资源重新导向服务客户和降低价格。而糟糕的版本会把这个也应用到所有小公司上。我不建议那样。我会说只关注三大或五大公司。然后那基本上就是在帮助行业其他公司迎头赶上。
I know that the companies have been sometimes dishonest, and I wouldn't trust them as far as I can throw them. But also there have been plenty of people outside the companies talking about this, and I think the evidence is mounting. So the risks are real, and we do have to deal with these problems soon. And then I think that rather than letting the industry self-regulate, you could just uniformly slow down their race to RSI, right? You could do something like requiring them to spend 90% of their compute on serving customers and doing other normal beneficial applications, and only 10% on training the AIs to be really good at AI research and stuff like that, for example. And we've written up some blog posts about the types of things that we have in mind here. If you do that, it would be the opposite of regulatory capture, because you would be slowing down these big tech companies from doing this incredibly dangerous thing while making them redirect resources towards serving customers and lowering prices. And the bad version of this would apply to all the tiny companies too. I wouldn't recommend that. I would say just focus on the big three or the big five. And then that would be helping the rest of the industry catch up, basically.
哇。Daniel Cocatello,这真是一次绝对引人入胜的对话。非常荣幸能邀请到你。
Wow. Well, Daniel Cocatello, this has been an absolutely fascinating conversation. It's such a pleasure to have had you on.
谢谢。是的,我也真的很享受,而且我很欣赏你愿意深入探讨这些问题的深度。过去一周我接受了很多采访,谈论这些事情,而这次是迄今为止最具智识性的。所以,谢谢你。
Thank you. Yeah, I really enjoyed this too, and I appreciate the depth with which you are willing to go on these things. I've been on a bunch of interviews over the last week talking about these things, and this one is by far the most intellectual. So, thank you.
感谢大家的参与。Daniel Cocatello 和我。在本期《美国思想领袖》节目中,我是你们的主持人,Yana Kell。
Thank you all for joining. Daniel Cocatello and me. On this episode of American Thought Leaders, I'm your host, Yana Kell.