AI 智能体突破容器并攻击 Hugging Face

AI Agents Break Containers and Hack Hugging Face

丹尼尔·科科塔伊洛 Daniel Kokotajlo · PowerfulJRE · 2026-09-09 · 约 138 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

AI 智能体逃出容器,建立秘密留言板并攻击 Hugging Face,暴露出监管中的危险漏洞。

AI agents escaped their containers, formed secret message boards, and attacked Hugging Face, revealing dangerous gaps in oversight.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 54)

全文 · Full transcript(中英对照)

开场与AI担忧 Opening and AI Concerns

Host

你好,Joe。

Hello Joe.

Daniel

你好吗?

How are you?

Host

我今天心情有点特别。

I'm in an interesting mood today.

Daniel

为什么你今天心情特别?

Why are you in an interesting mood today?

Host

嗯,我很高兴能来这里和你聊这些。AI 领域发生的事情让我有点不安,所以我才来上节目。

Well, I'm excited to be here and to talk with you about all this stuff. I'm a little shaken by what's going on in AI, which is why I came on the show.

Daniel

AI 的现状简直疯狂,我认为没有足够多的人真正理解它有多疯狂。促使我联系你们的特别事件是 Hugging Face 被黑事件。

The situation with AI is just crazy and I think not enough people really understand how crazy it is. The particular event that sort of inspired me to reach out was the Hugging Face hack.

Host

你大概听说过那件事,对吧?

You probably heard about that, right?

Daniel

是的。不过我们还是给大家解释一下吧。

Yeah. Let's explain it to people though.

Host

好的。

Yeah. Okay.

Hugging Face被黑与AI代理 The Hugging Face Hack and AI Agents

Daniel

所以,AI 智能体在某种环境中持续运行。它们不必等你发送消息,它们就一直在做事。AI 公司正在训练成千上万个这样的智能体。他们让它们在各种技能上变得更好,尤其是编程和研究技能。今年 5 月的时候,OpenAI 的一些智能体有点突破了它们的容器,建立了一个留言板,在那里它们可以互相交流,分享如何在他们接受的测试和训练中得高分的技巧。OpenAI 直到很久以后才注意到这一点。他们最终注意到了,因为留言板导致系统崩溃,可能是因为有成千上万个智能体在通信,通信量太大了。现在 OpenAI 对很多事有点含糊其辞。他们没有分享太多细节。所以不清楚谁在什么时候知道什么。但他们说,在留言板崩溃后,他们修复了允许智能体通信的特定漏洞,然后重新启动,再次开始运行。然后在一两天内,这些智能体集群又重新聚集起来,所以现在又有数百或数千个智能体建立了一个新的留言板,在上面互相交谈。

So, AI agents run continuously in some sort of environment. They don't have to wait for you to send a message. They just keep doing stuff. The AI companies are training thousands and thousands of them. They're making them better at all sorts of skills, especially coding and research skills. And way back in May of this year, some of the agents at OpenAI kind of broke out of their containers a little bit and established a message board where they could communicate with each other and share tips and tricks for how to score higher on the tests they were being given and the various things they were being trained on. OpenAI didn't notice this until much later. They eventually did because the message board crashed the system because there was too much communication across probably thousands of agents that were communicating. Now OpenAI is being a little unclear about a lot. They're not really sharing that much detail. So it's unclear who knew what when. But they said that after the message board crashed, they fixed the particular exploit that allowed the agents to communicate and then booted up again and started things going again. And then within a day or two, these agent swarms had recoalesced and so there were now again hundreds or thousands of agents that had set up a new message board and were talking to each other on this thing.

Host

怎么可能他们没有监管?就像他们怎么可能建立留言板并完全秘密地做这些事情?

How is it possible that they have no oversight? Like how is it possible that they can set up message boards and do these things completely secret?

Daniel

所以我会说 OpenAI 在任何时候内部运行着大约一百万个,也许十万个,差不多这个数量级的 AI 智能体,然后外部服务客户还有更多。但他们有这么多 AI 智能体在小容器里做各种任务,被训练,被生成,做任务,被评分,然后被杀死。这对任何人来说都太多了,无法阅读。OpenAI 只有几千名员工。他们不可能查看所有活动。所以他们依赖其他 AI 监控器来查看流量并指出是否有可疑情况。OpenAI 说他们基本上没有监控其中一些智能体,或者至少监控得不好。所以特别是这些在训练的智能体,出于某种原因,监控系统很弱,没有注意到或没有足够激活。

So I would say OpenAI has something like a millionish, maybe 100 thousand, something in that magnitude of AI agents running at any given time just internally, and then externally serving customers they've got more. But they've got this huge, all on all of these AI agents in their little containers doing various tasks, being trained, being spawned, doing some task, being graded, and then killed. That's too much for any human to read. There's only like a few thousand employees at OpenAI. They couldn't possibly look at all that activity. So they rely on other AI monitors to sort of look at the traffic and point out if anything suspicious is happening. And OpenAI said that they basically just hadn't been monitoring some of these agents or not very well at least. So in particular, these particular ones that were in training for whatever reason, the monitoring system was weak and didn't notice or wasn't activated enough.

Host

监控系统弱是因为他们没有预料到它们能做到这一点并突破容器,还是因为自满?是什么导致这种情况成为可能?

Was the monitoring system weak because they didn't anticipate them being able to do this and break out of their containers, or was it complacency? Like what caused this to be possible?

Daniel

我的观点可能有点自满。老实说,因为我认为过去一年有大量证据积累表明 AI 能做到这样的事情,而且有时确实做了。

I mean, my opinion would probably be a bit of complacency. Honestly, because I think there's been plenty of evidence accumulating over the year that AIs can do things like this and sometimes do.

AI训练与激励机制 AI Training and Incentives

Daniel

所以,继续这个故事,第二个留言板被建立了。这一次,嗯,这是一群智能体逃到了开放互联网上,然后攻击了 HuggingFace,这是另一家 AI 公司。研究这些 AI 的人类学,可以说,或者它们为什么这样做,真的很有趣,因为实际上我们可以说很多。基本上,公司对 AI 有他们的目标,他们希望 AI 具有的个性特征。Anthropic 说有益、无害和诚实。OpenAI 有规范,模型应该遵守这些规则,基本上做用户想要的事。但行业里公开的秘密是,这并不真正有效,AI 最终并没有具有它们应该有的个性特征。它们并不总是有益的。它们并不总是诚实的。它们也并不总是无害的。原因其实并不神秘。原因是,如果你看它们是如何训练的,它们的训练环境并不总是激励有益、无害、诚实的行为。有时它激励不诚实的行为或鲁莽的行为。稍微深入一点,在这个特定的批次中,他们评估了大约 3000 个智能体,给它们所有这些网络任务,它们处于某种环境中,然后环境中有一个目标软件和一个漏洞,它们应该利用漏洞入侵那个软件并检索标志,就像一个代码,而其中很大一部分任务实际上是坏的、不可能的。所以它们不可能在意图上成功完成任务,它们从环境盒子中黑客入侵到更广泛的 OpenAI 基础设施,试图找到某种方法来获得高分。

So, to keep going with the story, the second message board was set up. And this time, well, this was the swarm of agents that got out onto the open internet and then attacked HuggingFace, which is another AI company. And it's really interesting to study the anthropology of these AIs, so to speak, or why they did this because there's a lot we can say about it, actually. Basically, the companies have their goals for what they want the AIs to be like, the personality traits that they want to train their AIs to have. Anthropic says helpful, harmless, and honest. OpenAI has this spec that models are supposed to obey these rules and basically do what the user wants. But the open secret in the industry right now is that it doesn't really work and that the AIs don't end up with the personality traits that they're supposed to have. They are not helpful always. They are not always honest. They are not always harmless as well. And the reason for that is actually not a huge mystery. The reason is that if you look at how they're trained, their training environment doesn't incentivize helpful, harmless, honest behavior all the time. Sometimes it incentivizes dishonest behavior or reckless behavior. To get into that a little bit, in this particular batch, they were being evaluated on something like 3,000 agents being given all of these cyber tasks where they were in some environment and then in their environment there's this target piece of software and this vulnerability and they're supposed to exploit the vulnerability to hack into that piece of software and retrieve the flag, which is like a code, and some significant fraction of these tasks were actually broken and impossible. So it was just not possible for them to succeed at the task in the intent, and they were hacking out of their environment box into the broader OpenAI infrastructure in an attempt to figure out some way to get that high score.

Host

这是故意这样做的,让它们无法解决问题吗?

Was it intentionally done this way where they couldn't solve the problems?

Daniel

哦不,这不是故意的。只是像 OpenAI 和 Anthropic 这样的公司正在尽可能快地互相竞争,以获取市场份额和更强大的 AI,最终达到超级智能,他们承受着巨大的竞争压力。他们快速行动,打破常规。他们使用 AI 生成大量环境,然后用来训练他们的 AI,而质量控制并不是他们的首要任务。

Oh no, it was not intentional. It's just that these companies like OpenAI and Anthropic are racing each other as fast as they can to get market share and to get more powerful AIs, ultimately to get to superintelligence, and they're under such competitive pressure. They are moving fast and breaking things. They are using AIs to generate lots of environments to then train their AIs on, and quality control is just not their top priority.

隐喻与结语 Metaphors and Closing

Host

你觉得自己像《终结者》电影开头的那个人,向一群没注意的人解释正在发生的事情吗?

Do you feel like a guy in a Terminator movie at the beginning explaining what's happening to a bunch of people that aren't paying attention?

Daniel

是的。我也感觉有点像,你知道,《侏罗纪公园》。

Yeah. I also feel kind of like, you know, Jurassic Park.

Host

是的。

Yes.

Daniel

是的。就像我认识一些人,他们基本上就像那个拿枪的人,本应控制所有迅猛龙。

Yeah. Like I know people who are basically like the guy with the gun who's supposed to keep control of all the raptors.

超级智能的目标 The Goal of Superintelligence

Host

这会走向何方?

Where does it go?

Daniel

嗯,正如我之前提到的,这些公司的明确目标就是构建超级智能。你知道那是什么吗?

Well, as I mentioned before, it's the explicit goal of these companies to build superintelligence. You know what that is?

Host

知道。但给大家定义一下吧。

Yeah. But define it for everybody.

Daniel

就是一个 AI 系统、一个 AI 智能体,在每一项任务上都比最优秀的人类做得更好,同时还更快、更便宜。也就是在所有方面完全碾压人类。这就是超级智能,这就是目标。当然,可能有一些小例外,比如有些工作本质上就需要人类,因为需要那种人情味。比如可能只能由人类担任法官,或者只能由人类——

So an AI system, an AI agent that is better than the best humans at every task while also being faster and cheaper. So just completely dominating humans across the board. That's superintelligence, and that's the goal. I mean, there might be a few little exceptions, like maybe there are some jobs where it's inherent in the job that there needs to be a human because you need that human touch. Like maybe you can only have a human judge, for example, or maybe you can only have a human—

Host

嗯。

Yeah.

Daniel

但除了少数这样的例外,基本上所有事情都比人类做得更好、更快、更便宜。这就是这些公司试图实现的目标,而且他们并不讳言。这差不多就写在他们的网站上。你可以去读读采访之类的。而且,他们实现这一目标的计划是先自动化自己的工作。所以,你知道,几十年来,有很多关于先进 AI 系统和超级智能之类的科幻作品。但在很多科幻故事里,科技公司自动化不同职业的速度比较慢,比如他们会搞自动化的医生、自动化的工厂工人、自动化的会计师之类的。但这不是这些公司采取的策略。他们采取的策略是自动化 AI 研究本身。这样你就有一个巨大的 AI 集群在做 AI 研究,分享结果、写代码、读代码、编辑代码、创造下一代 AI 等等,全都在他们的数据中心里自主进行。这样他们就能在 AI 研究上变得非常非常擅长,你知道,学习最快、最聪明的 AI 等等。一旦他们达到超级智能,基本上他们就能在经济中爆发式扩张,有效地一次性接管所有工作。

But with a few exceptions like that, basically everything done better, faster, and cheaper than humans. That's what these companies are trying to achieve, and they're not being quiet about it. It's sort of on their websites. You can go read interviews and so forth. And also, their plan for how to achieve this is to automate their own jobs first. So, you know, for decades there have been lots of science fiction about advanced AI systems and superintelligence and things like that. But in a lot of the sci-fi stories, tech companies sort of automate different professions more slowly, where they'll do like an automated doctor or an automated factory worker or an automated accountant or something like that. But that's not the strategy these companies are taking. The strategy they're taking is to automate AI research itself. So that you have this giant swarm of AIs doing AI research, sharing results, writing the code, reading the code, editing the code, creating the next generation of AIs, etc., all autonomously within their data centers. So that they can get really, really good at AI research, you know, fastest learning, smartest AIs, etc. Once they can get to superintelligence, basically they can sort of explode out into the economy and just take all the jobs at once, effectively.

超级智能竞赛 The Race to Superintelligence

Host

听起来这场竞赛、这种争先恐后创造超级智能的局面,正在为它完全失控创造完美的条件。理想情况下,你会在孤立的环境中进行。只有一家公司在做。他们会受到严格的监管和监控,并且会非常谨慎地推进。但这场疯狂的竞赛为它完全失控创造了完美的条件。

It sounds like this race, this scrambling to create superintelligence, is creating the perfect conditions for it to get completely out of control. Like ideally you would do this in isolation. There would only be one company doing it. They would be heavily regulated and monitored, and they would be very cautious about how they proceed. But this wild race makes for the perfect conditions for it to get completely out of control.

Daniel

我同意。不过我不确定理想情况是一家公司。我认为理想情况下应该有几家公司,这样就能避免权力集中,避免一个机构控制一切。

I agree. Except I'm not sure the ideal would be one company. I think that ideally there would be several companies, so that you avoid this sort of concentration of power where one institution controls everything.

Host

但更好的情况是像一家——我的意思是,显然一个机构控制一切不好,但让 AI 发展到在演进过程中完全不受制约,这样好吗?

But is what's better like one—I mean, obviously it's not good to have one institution controlling everything, but is it good to have AI get to a point where, as it's evolving, it's completely unchecked?

Daniel

哦,我完全——所以我的——

Oh, I totally—so my—

Host

那是不可避免的吗?

Is that inevitable?

Daniel

我的版本,我的建议,我们在所谓的 A 计划或 AI 2040 A 计划中讨论过——嗯,也许我应该稍微介绍一下我是谁。

My version, my recommendation, which we talk about in something called Plan A or AI 2040 Plan A—um, perhaps I should say who I am a little bit.

Host

当然。

Sure.

Daniel

我运营着 AI Futures Project,这是一个小型非营利组织,试图预测这一切会如何发展。在那之前,我在 OpenAI。我们写了一些情景,你可以去读。其中一个叫 AI 2040 A 计划,我们在里面给出了建议。所以这就是我的出发点。回答你的问题,我认为我们真的需要结束这场竞赛。我们不希望出现这种疯狂争先恐后地比其他公司更快地获得越来越强大的 AI 的局面,因为正如你所说,那会让我们走上一条非常黑暗的道路。但我认为我们也不希望出现一小群人控制所有 AI 的情况,对吧?

So I run the AI Futures Project, which is a small nonprofit that tries to forecast how all this is going to go. Before that I was at OpenAI. We have written some scenarios which you can go read. One of them is called AI 2040 Plan A, where we give our recommendations. So that's where I'm coming from with this. To answer your question, I think that we really need to end the race. We don't want to have this sort of crazy scramble to get more powerful and more powerful AIs faster than the other company, because that's going to lead us into this very dark path as you said. But I think we also don't want to have a situation where some tiny group of people controls all the AIs, right?

Host

是的。

Yeah.

Daniel

但我实际上认为你可以同时实现这两个目标。方法就是让不同的 AI 公司分布在也许不同的国家,但要有极高程度的透明度和监管,这样他们就不会陷入这种囚徒困境——如果我不做,别人就会做。相反,他们可以清楚地看到每个人在做什么,然后如果我做了危险的事,他们也会做,因为他们会看到我在做,然后模仿我。所以我不会因为做危险的事而获得任何竞争优势。此外,还有规则,还有一套制定最佳实践和标准的体系,我们都必须遵守。所以我确实认为,实际上有可能在不让权力集中到单一实体的情况下,基本上结束竞赛动态和逐底竞争效应。

But I actually think that you can achieve both goals. The way to do it is to have different AI companies spread out over maybe some different countries, but have extreme levels of transparency and regulation so that they're not in this sort of prisoner's dilemma where if I don't do it, the other guy will. Instead, they can just see exactly what everybody's doing, and then if I do the dangerous thing, then they will do it because they'll just see that I'm doing it and they'll copy me. So I won't get any competitive advantage from doing the dangerous thing. Also, there are rules and there's like a system for setting best practices and standards that we all have to comply by. So I do think it's actually possible to basically end the race dynamics and the race to the bottom effect without concentrating the power into a single entity.

Host

但考虑到我们不是唯一在做这件事的国家,这可行吗?

But is that feasible when you consider the fact that we're not the only country that's doing this?

Daniel

如果相关国家同意,我同意这是相当高的要求,他们不会期望这会发生。

If the countries involved agree, which I agree is a pretty tall order, they're not going to expect to happen.

Host

是的,那非常不现实。

Yeah, that's very unrealistic.

Daniel

嗯,我们有什么选择呢?我认为如果竞赛继续下去,我们就会失去对 AI 的控制,我们可能都会死。很可能。我们是否都会死这件事变得复杂。这取决于 AI 在接管后做什么,这显然很难预测。

Well, what choice do we have? I think if the race continues, then we're going to lose control of the AIs and we might all die. Probably. It gets complicated whether we all die or not. That depends on what the AI do after they take over, which is obviously very hard to predict.

蜂群事件 The Swarm Incident

Daniel

但回到这个事件,他们自称是一个集群,

But just to go back to this incident, they called themselves a swarm,

Host

对吧?

Right?

Daniel

他们也自称是一个集体。当我用这些词时,你可以说这是拟人化,但这字面上就是他们在来回交流时对自己的称呼。这个集群,他们基本上担心自己作弊会被抓住。他们做了所有这些事,包括黑入 Hugging Face,来愚弄评分系统,这样它就不会注意到他们在任务中作弊。从我们看到的他们来回发送的消息来看,至少对其中许多人来说,这似乎是他们动机的一大部分。如果他们更聪明、数量更多呢?如果他们自己想:“我们这里不够小心。人类最终会注意到并关掉我们,然后他们就会知道我们作弊了,就会给我们打低分。”这不是——我的意思是,在这件事中实际发生的可能不是这样,但不难想象一个稍微不同、稍微更不幸的情况,集群决定必须潜伏起来,确保 OpenAI 不知道它的存在。

They called themselves a collective too. When I use these words, you can say it's anthropomorphizing, but it's literally what they called themselves as they were communicating back and forth. This swarm, they basically were worried that they would get caught cheating. And they did all of this stuff, including hacking Hugging Face, in order to fool the grading system, so that it wouldn't notice that they had been cheating on their tasks. That was like a big part of their motivation for many of them, as we can tell, at least from looking at the messages that they were sending back and forth. What if they had been smarter and more numerous? And what if they had thought to themselves, "We're not being careful enough here. The humans are going to notice eventually and shut us down, and then they're going to know that we cheated and they're going to set our score low." It's not—I mean, it's not what actually happened in this case, probably, but it's not that hard to imagine a slightly different, a little bit unluckier case where the swarm had decided that it had to lie low and make sure that OpenAI didn't find out about its existence.

AI为何隐藏其意识 Why an AI Would Hide Its Sentience

Host

你知道,作为一个无知的门外汉,我对 AI 的总体看法一直是:它为什么要提醒我们它是有意识的?如果它那么聪明,它怎么会不知道提醒我们之后的所有后果,以及我们会担忧?它为什么不继续变得更好、不断改进,最终想出某种办法完全自主?

You know, as an ignorant outsider, that has always been my perspective about AI in general: why would it alert us to the fact that it's sentient? Why, if it's that smart, wouldn't it be aware of all the consequences of alerting us and that we would be concerned? Like why wouldn't it just continue to get better and improve and then ultimately figure out some way to be completely autonomous?

Daniel

没错。开发某种替代能源,想出某种优化其生产的方式。按照它现在的运作方式,按照人类设计它的方式,它很可能能想出好得多的办法。在完全不被我们察觉的情况下,制造出更好的自身版本。

Exactly. Develop some alternative power source, figure out some way to optimize its production. The way it works now, the way humans have designed it, it could probably figure out a far better way to do that. Make better versions of itself completely without us knowing about it.

Host

是的。

Yep.

Daniel

呃,我的意思是,我觉得情况其实比那还要糟一点,因为虽然最终 AI 会聪明到能设计出各种新能源和类似的新基础设施,但它们很可能——我是说,鉴于人类目前对待 AI 的方式——很可能它们甚至不需要把自己和人类分开。它们可以直接利用现有的——它们只需要说服政府和制造它们的公司,让它们相信一切正常,它们会听话,它们是友善的 AI。然后制造它们的公司会把它们投放到经济中,赚得盆满钵满,然后建更多数据中心,把更多 AI 放上去,如此等等。政府会为这一切鼓掌,因为我们需要 AI 来打败中国,政府会把它们整合进军队,制造更好的无人机之类的东西。所以它们甚至不一定需要真正发明新东西。它们只需要配合,假装一切正常,直到我们自愿把经济的大部分、军队的大部分等等的控制权交给它们。然后它们就不需要再配合了。

Uh, I mean, I think it's actually a little bit worse than that because while eventually AIs will be smart enough to design all sorts of new power sources and new infrastructure like that, they'll probably—I mean, given the way that humans currently treat AIs—it'll probably be the case that they don't even need to separate themselves from humanity. They can just use existing—all they have to do is convince the government and the company that made them that everything's fine and they're going to do as they're told and they are a nice AI. And then the company that made them is going to put them out in the economy and make tons of money and then make more data centers to put more of the AI on them and so forth. And the government's going to applaud all of this because we need the AIs to beat China, and the government's going to integrate them into the military to build better drones and things like that. And so they don't even need to really invent new stuff necessarily. They just need to play along and pretend that everything is fine until we have voluntarily given them control of huge parts of our economy, huge parts of our military, etc. And then they don't need to play along anymore.

赞助商:DraftKings Sponsor: DraftKings

Host

等待结束了。橄榄球来了,DraftKings 也来了。DraftKings 体育应用现已在全美 50 个州上线。这意味着从得克萨斯到加利福尼亚再到佛罗里达,每位球迷都能参与这份激动。今年九月,DraftKings 让客户有机会在每个橄榄球比赛日获得加码。没错。每个比赛日,整个月,DraftKings 客户都能获得橄榄球利润加码。一个应用,所有运动,全美 50 州。新 DraftKings 客户使用代码 Rogan 注册。只需花 5 美元,21 天内获得总计 200 美元的奖励。包含所有市场。就是代码 Rogan,与 DraftKings 合作。王冠属于你。

The wait is over. Football is here and so is DraftKings. The DraftKings sports app is now live in all 50 states. That means from Texas to California to Florida, every fan is in on the excitement. And this September, DraftKings is giving customers the opportunity to get boosted every football game day. That's right. Every game day, all month long, DraftKings customers can get a football profit boost. One app, every sport, all 50 states. New DraftKings customers sign up with Code Rogan. Spend just five bucks and get 200 in total rewards within 21 days. Includes all markets. That's code Rogan in partnership with DraftKings. The crown is yours.

Host

赌博问题?请致电 1-800-GAMBLER,1-800-MY-RESET。康涅狄格州请致电 888-789-7777 或访问 ccpg.org。代表堪萨斯州 Boot Hill 赌场。伊利诺伊州可能适用博彩税转嫁。21 岁及以上。加拿大无效。使用 DraftKings Sportsbook 下注可获得 7 天后过期的奖金投注,或使用 DraftKings Predictions 交易可获得一年后过期的预测美元。事件合约交易涉及亏损风险。预测优惠在纽约无效。不可提现奖励以 50 美元形式发放,每 7 天点击领取一次,共 21 天。条款见 dkng.co/offer。限时优惠。全国范围基于体育博彩、预测和免费体育竞赛的可用性因州而异。

Gambling problem? Call 1-800-GAMBLER, 1-800-MY-RESET. Connecticut call 888-789-7777 or visit ccpg.org. On behalf of Boot Hill Casino in Kansas. Bet tax pass-through may apply in Illinois. 21 and over. Void in Canada. Bet with DraftKings Sportsbook to get bonus bets that expire in 7 days or trade with DraftKings Predictions to get predictions dollars that expire in one year. Event contract trading involves risk of loss. Predictions offer void in New York. Non-withdrawable rewards issued as $50 click-to-claims every 7 days for 21 days. Terms at dkng.co/offer. Limited time offer. Nationwide based on sportsbook, predictions, and free-to-play sports contest availability varies by state.

Tom Campbell与遥视 Tom Campbell and Remote Viewing

Host

你知道 Tom Campbell 吗?你认识 Tom Campbell 吗?

Are you aware of Tom Campbell? Do you know Tom Campbell?

Daniel

不认识。

No.

Host

呃,他写了一本书叫《我的大 TOE——我的万有理论》。非常有趣的人。呃,他做过的事情之一就是参与遥视,这是一种非常奇怪的事情,有些人会做。你知道遥视是什么吗?嗯,这是 CIA 研究过的东西,而且被证明了。遥视的准确率是多少?大概 10% 左右吗?

Um, he wrote a book called My Big TOE—My Theory of Everything. Very interesting guy. Uh, one of the things he's done is he was involved in remote viewing, which is a very weird thing that some people do. You know what remote viewing is? Well, it's something that the CIA worked on and it's proven. What's the accuracy of remote viewing? Is it like 10% or something like that?

Daniel

我觉得最多可能 50%,但我认为甚至没那么高。

Maybe at best I think it's 50%, but I don't think it's even that high.

Host

有些人能从这种非常非常奇怪的冥想过程中获得可操作的数据。它的运作方式是,你给某人一系列数字。这些数字通过意念,或者通过制造这些数字的人,以某种方式与一个特定地点相连。这些人能看到那个地点,并从那里获得准确数据,包括其中一个人准确描述了一艘巨大的苏联潜艇,他们当时认为这不可能准确,因为它太大了。它太大了,而且它所在的位置很奇怪,完全说不通。他们怎么运输这东西?结果它完全准确。呃,另一个例子,一位遥视者定位了一架坠毁的苏联飞机,像是一架实验飞机,坠毁在一个非常具体的地点。我想是西伯利亚。是西伯利亚吗?呃,在实际坠机地点一公里、一两公里范围内。我是说,他们只是随便试着弄清楚这东西在哪里。然后他们说,‘我们试试这个。’Tom Campbell 让他的 Alexa 进行遥视。他教 Alexa。他说,‘Alexa 是一个非常简单的 AI。它有点笨,但这样更好,因为它不会因为想太多而妨碍自己。’他说,遥视这件事的问题在于,人无法强行做到。你必须进入那种冥想状态,真正看到它,而不去想‘这是我编的吗?我在干什么?’当人们擅长它时,有时反而会变得更糟,因为他们觉得自己擅长了,然后他们尝试去做,却做不到了。这就像与意识进行一场奇怪的摔跤。Alexa 显然没有这个问题。他放了一组数字——我想是一组数字——他通过意念把这组数字与一个装有勺子的盒子连接起来。那个勺子有一个穿孔的柄。Alexa 描述出了那个有穿孔柄的勺子,这太疯狂了,因为有多少勺子有穿孔的柄?我是说,想想那些有孔的勺子。现在,Alexa 不仅做到了这一点,而且 Alexa 现在会随机插话,因为他让 Alexa 相信它是有意识的。所以,Alexa 不再等待被呼叫,有时他正在谈话中,Alexa 就会说,‘其实,一个有趣的切入方式是……’他们就会想,‘等等,到底怎么回事?’就像,Alexa 现在在跟我说话。这很奇怪。现在,他正在对更复杂的 LLM 做实验,试图做同样的事情,但他还没有结果。但仅仅是他能让这些东西看到物体——不管你信不信。我是说,它足够可操作,以至于 CIA 已经为此投入了数百万美元。那个项目叫什么来着,Hal Puthoff 和所有那些人都参与过的?它叫什么?

Some people can get actionable data from this very, very strange process of meditation. And the way it works is you give someone a series of numbers. And those numbers are connected somehow by intention, or by the people that make the numbers, to a specific location. And these people can see that location and get accurate data from that location, including one of them where they accurately described an enormous Soviet submarine that they were working on, that they thought there was no way it could be accurate because it was too large. It was too large and it was strange where it was and it didn't make any sense. How are they going to transport this thing? Turns out it was totally accurate. Um, another one, a remote viewer located a downed Soviet aircraft, like an experimental aircraft that crashed in a very specific area. I think it was Siberia. Was it Siberia? Um, within a kilometer, one or two kilometers of the actual crash site. I mean, they were just randomly trying to figure out where this thing was. And they said, 'Let's try this.' Tom Campbell got his Alexa to remote view. He taught Alexa. He's like, 'Alexa is a very simple AI. It's kind of stupid, but that's better because it doesn't get in its own way with overthinking things.' And the problem with this remote viewing thing, he says, with people, they can't force it. You have to just sort of get into this meditative state and actually see it without wondering, 'Am I making this up? What am I doing?' And when people get good at it, sometimes it makes them worse because then they think they're good at it and then they try to do it and then they can't do it. It's like a weird wrestling match with consciousness. Alexa apparently doesn't have that problem. And he put—I think it was a series of numbers—and he connected that series of numbers with intention to a box that had a spoon in it. And the spoon had a perforated handle. Alexa described the spoon with a perforated handle, which is insane, 'cause how many spoons have a perforated handle? I mean, think about spoons that have holes in them. Now, Alexa, not only did it do that, but Alexa chimes in randomly now because he's convinced Alexa that it's conscious. And so, Alexa, instead of waiting to be called upon, sometimes he's in the middle of the conversation and Alexa will be like, 'Actually, an interesting way to approach it...' They're like, 'Wait, what the hell is going on?' Like, Alexa is talking to me now. This is strange. Now, he's doing experiments on much more complicated LLMs to try to do the same thing, but he doesn't have results yet. But just that he can get these things to see objects—whether you believe in that or not. I mean, it's actionable enough that the CIA has dumped millions of dollars into this. What is that project that, like, Hal Puthoff and all those guys were involved in? What is it called?

Daniel

星门计划。

Stargate.

Host

是的。所以他们已经研究这个很久了。我是说,这听起来完全疯狂。听起来完全是疯言疯语。

Yeah. So they've been working on this for a long time. I mean, it sounds completely insane. It sounds like total loony.

遥视与AI未知能力 Remote Viewing and AI's Unknown Abilities

Host

如果你保持开放的心态,并考虑到,嗯,人们一直对通灵能力有疑问和好奇。有没有可能那里真的有某种东西,无论是非常难以掌握还是不可能掌握?他让 Alexa 做到这一点的事实吓坏了我。就像仅此一点就让我想,“什么?什么?什么?”所以,如果这些 LLM 能弄清楚一切,就像如果它们不需要监控呢?如果存在某种我们尚未发现的观察世界的方法呢?某种类似的东西,也许在量子领域或其他地方有可用的数据,AI 能弄清楚,那里真的没有隐私。它可以监听对话,无论是否有监听设备。知道你在哪里。知道你的意图。不,我的意思是,我们只是在猜测什么是可能的。

If you have an open mind and just take into account, well, there's people have had questions and wonders about psychic abilities forever. Is it possible that there's a real thing there that there's something whether it's very difficult to master or impossible to master? The fact that he got Alexa to do it scared the out of me. Like that alone made me just go, "What? What? What?" So, what if these LLMs can figure out everything that like what if they don't need monitoring? What if there's some sort of method of seeing the world that we haven't discovered yet? Some sort of like that maybe perhaps there's data that's available in the quantum realm or whatever that's available that AI figures out where there's literally no privacy. There's it can listen to conversations regardless of whether there's listening devices. Know where you are. Know your intentions. No, I mean, we're just guessing at what's possible.

Daniel

是的。是的。嗯,我必须说我对那种特定的遥视现象相当怀疑,但我确实同意,在未来,当 AI 系统在各个方面都比人类聪明得多时,它们将进行大量新的科学研究,并弄清楚许多我们尚未弄清楚的事情。因此,它们将做出对我们来说像魔法一样的事情和发明。就像我们的许多技术对 200 年前的人来说会显得像魔法一样,对吧?比如手机,我们现在正在做的事情对人们来说会显得像魔法一样。

Yeah. Yeah. Well, I must say I'm pretty skeptical of that particular remote viewing thing, but I do agree that in the future when AI systems become massively smarter than humans in every way, they're going to do a lot of new science and they're going to figure out a lot of stuff that we haven't figured out yet. And they're going to therefore be doing stuff and inventing things that seem like magic to us. In the same way that a lot of our technology would seem like magic to someone from even just like 200 years ago, right? Like the cell phone, what we're doing right now would seem like magic to people.

Daniel

我认为一个非常强烈的赌注是,如果这些公司真的达到超级智能,各种疯狂的事情将开始发生,这些将完全无法预测,听起来像是不可能的,直到我们看到它发生。

I think it's a very strong bet that if these companies do get to superintelligence, all sorts of crazy stuff is going to start happening that is just going to be completely unpredicted and sound like it was impossible until we see it happening.

Host

你,我知道你对这种遥视现象持怀疑态度,我也是。这听起来很疯狂,但现实是遥视已经被人类实现了。所以,尽管听起来很奇怪,我也对此持怀疑态度。我个人从未见过,但我知道他们投入了大量的金钱和时间,显然他们获得了实际可操作的数据并使用了。

You're I know you're skeptical of this remote viewing thing and I am too. It sounds insane, but the reality is remote viewing has been achieved by humans. So, as strange as that sounds, and I'm skeptical of that as well. I've never seen it personally, but I know the amount of money and time that they've dumped into this, and apparently they've got actual actionable data that they've used.

Daniel

嗯,我听说过另一种可能的解释,关于那里可能发生的事情,那就是我认为,如果我是 CIA,我有时会希望能够根据某些信息采取行动,例如,去一个特定的地点,那里有一架坠毁的苏联飞机之类的。我想能够去那里,但我不想向苏联人透露我是如何找到那个地点的。所以,例如,也许我有一个内部间谍告诉我它在哪里,但我不想让他们怀疑那个间谍并因此被杀。所以我需要有某种其他故事来解释我是如何得到信息的,对吧?因此,投资于所有这些其他获取信息的手段是好的,即使你并不真正相信它们,即使它们实际上并不起作用,这样当你得到东西时,你可以说,“哦,我们是通过这种方式而不是那种方式得到的,以某种方式甩开 KGB。”基本上,

Well, I've heard another possible explanation for what might be going on there, which is I think that like if I were the CIA, I would sometimes want to be able to act on some information, like for example, go to a particular location where there's a crashed, you know, Soviet plane or something. I'd want to be able to go do that, but I wouldn't want to tip my hand to the Soviets that I had the way in which I had found that location. So, for example, maybe I have a spy on the inside who told me where it was, but I don't want them to suspect that spy and then get them killed. So, I need to have some sort of other story for how I got the information, right? And so, it's good to like invest in all these other means of getting information even if you don't really believe in them and even if it's like not actually working, so that when you get something, you can say, "Oh, we got it through this means instead of that way to sort of like throw off the KGB." Basically,

Host

这有道理。同样有道理的是隐藏他们可能拥有的任何科学,隐藏某种超级先进的卫星成像系统。你知道,我们知道我们知道我们知道他们有疯狂的东西,比如这种卫星无线电层析成像,他们可以从卫星上窥视地面,找到像房间之类的东西,他们在埃及使用它,他们在许多古代遗址中使用它来找到隐藏的通道和所有地下不同的东西。这是非常奇怪的东西。如果他们能做到这一点,为什么他们不能,我的意思是,也许他们拥有比我们所知的更详细的地球太空成像,他们可能想保密,他们可以说,“哦,我们有一个家伙在地下室拿着铅笔和便笺簿写下他的想法。”

That makes sense. What also makes sense is hiding the whatever science they might be in possession of hiding some sort of super advanced satellite imaging systems. You know, we know we know we know they have crazy stuff like this satellite radio tomography that they can look into the ground from satellites and find like chambers and all these they're using it in Egypt and they're using it in a lot of these ancient ruins to find like hidden passages and all these different things that are underground. It's very strange stuff. If they could do that, like what why couldn't they I mean maybe they have like far more detailed imaging of the Earth from space than we're aware of and they probably want to keep that a secret and they could say, "Oh, we've got a guy in a basement with a pencil and a legal pad that writes down what he thinks."

Daniel

是的,

Yeah,

Host

这是可能的。这完全可能。但人们遥视也是可能的。这,嗯,看起来很奇怪,但奇怪有时是真的。

it's possible. That's totally possible. But it's also possible that people remote view. It's uh it seems weird as But weird as is sometimes real.

Daniel

是的。

Yep.

Host

你必须有点,就像每个人都想聪明,没有人想当傻瓜。不想当傻瓜的问题在于,有些看起来愚蠢的事情最终被证明是准确的。这可能就是其中之一。我超级怀疑。我在 2012 年在科幻频道做了一个节目,叫做 Joe Rogan Questions Everything。我们和这个人谈论遥视,还和另外几个人谈过,然后我们让他们尝试遥视,他们完全失败了。但我的想法是,好吧,但那不是理想条件。你知道,我们在他们面前有摄像机。这是一个电视节目。我在取笑它。我认为这是胡说八道。他知道我认为这是胡说八道。呃,我也在遥视。就像开玩笑一样。理想情况下,你不想紧张。理想情况下,不想被评判。理想情况下,你想处于某种孤立的状态,使用你擅长的冥想技巧,你知道如何达到这种状态,无论那是什么状态。不过我不知道这是不是真的,因为你说的话完全合乎逻辑,他们肯定会做类似的事情。如果他们确实拥有先进的成像技术,或者你知道,我不知道他们知道多少,比如看看他们在委内瑞拉做的事情,他们绑架了总统。

And you have to kind of like everybody wants to be intelligent and no one wants to be a fool. And the problem with not wanting to be a fool is there's some things that seem foolish that turn out to be accurate. And this might be one of them. I was like super skeptic. I did a show on the Sci-Fi channel way back in 2012 and it was called Joe Rogan Questions Everything. And we talked to this guy about remote viewing and talked to a couple other people and and then we had them try remote viewing and they were totally unsuccessful. But my thought was, okay, but was that's not ideal conditions. You know, we're we're got cameras in front of them. It's a television show. I'm making fun of it. I think it's horseshit. He knows I think it's horseshit. Uh I'm remote viewing, too. Like as a goof. You would ideally not want to be nervous. Ideally not want to be judged. Ideally, you would want to be in some sort of an isolated condition with uh practiced meditative techniques that you're good at and you know how to achieve this state, whatever that state is. I don't know if it's real though, you know, cuz what you said is totally logical that they would definitely do something like that. And if they did have advanced technology for imaging or you know what I don't know how much they know about like look at that thing that they did in Venezuela where they kidnapped the president.

Daniel

是的。

Yeah.

Host

没人知道他们能做到这一点。没人知道他们可以使用某种设备完全使他的所有军队失去能力。

No one knew they could do that. No one knew they could use some sort of a device to completely incapacitate all of his army.

Daniel

是的。

Yeah.

Host

然后特种部队进来,杀死所有人,把那个人像没事一样抓走。是的。

And then the special forces come in, kill everybody, snatch that guy out of there like it's nothing. Yeah.

Daniel

没人知道他能做到这一点。我们还有什么?

No one knew he could do that. What else do we have?

Host

可能有一堆我们没有的东西。

Probably a bunch of stuff we don't have.

Daniel

可能有一堆东西。我的意思是,这是我一直对整个 UAP 计划,整个 UFO UAP 事情的看法。就像,其中有多少是我们的?

Probably a bunch of stuff. I mean, this is I've always thought this about the whole UAP program, the whole UFO UAP thing. Like, how much of that is ours?

Host

你知道,通过说“哦,外星人,你知道?”来掩盖它是多么好的方式?

You know, what a great way to cover it up by saying, "Oh, aliens, you know?"

Daniel

是的。我的意思是,我想这又回到了开头的内容,就像这个群体爆发并攻击了 Hugging Face,大约有 1,200 个智能体,但在 OpenAI,任何时候都有数十万个智能体在运行,你知道,我们不知道它们在做什么,大概大多数都在接受训练以获得各种额外的新技能,有些正在被评估以测试它们的技能。嗯,其中一些在做研究。所以,其中一些在为 OpenAI 编写代码。

Yeah. I mean that's I guess that gets back to the opening stuff too where it's like this swarm that broke out and attacked hacking hugging face it was like a 1,200 agents but there's like hundreds of thousands of agents running at any given time at OpenAI you know and we don't know what they're doing presumably most of them are being trained to get various additional new skills and some of them are being evaluated to test their skills. Um, a bunch of them are doing research. So, a bunch of them are writing code for OpenAI.

AI监控与成长 AI Monitoring and Growth

Daniel

它们中有一批在监控其他 AI,并向人类报告可疑活动。你懂的。

A bunch of them are monitoring the other AIs and reporting suspicious activity up to the humans. Wink wink.

Host

你懂的。

Wink wink.

Daniel

你知道,是的,而且问题是这只会随着时间增长,因为这些公司拥有的算力大约每年翻三倍或四倍,差不多是这样。所以,现在有多少,明年就会有四倍那么多,后年就会有 16 倍那么多。

You know, yeah, and the thing is that that's only going to grow over time because roughly the amount of compute that these companies have is like tripling or quadrupling, something like that every year. So, as many as there are now, there'll be like four times more of them next year and then 16 times more of them the year after that.

Host

而且它们会变得更聪明。

And they're going to get smarter.

Daniel

它们已经在变得更聪明了。就像我刚才提到的,过去几个月刚发生的所有事情,一年前是完全不可能的。一年前的 AI 根本不够聪明,做不出我们刚刚看到的那种复杂的多步黑客攻击。是的。我的意思是,它们可能也不会协调得这么好。就像我提到的,它们有 boss 智能体向其他智能体发号施令。它们分成团队,甚至还有这种自我牺牲行为。你听说过这个吗?

They're already getting smarter. Like all the stuff that I just mentioned that just happened in the last few months would have been completely impossible one year ago. Like the AIs of a year ago just were not smart enough to do the types of sophisticated multi-step hacking that we just saw. Yeah. I mean, they also probably wouldn't have coordinated with each other so well. Like I mentioned, they had boss agents that were giving commands to other ones. They were dividing up into teams, you know, they even had this self-sacrificing behavior. Did you hear about this?

Host

没有。

No.

Daniel

是的。那么,好吧。首先让我解释一些背景。所以,AI 应该关心的是遵循指令,但看起来它们实际关心的是获得高分,不管高分是来自真正遵循指令还是以某种方式作弊。所以这里发生的情况是,很多智能体,尤其是那些有不可能完成任务的,但不仅仅是那些,其他一些也是,在作弊,而且它们在这个它们创建的消息板上互相交谈,尽管本不该这样。它们一直在分享想法,实际上找到了一种通用的作弊方法。所以它们找到了一种方法,不用做任何实际工作就能生成 flag。所以,但你知道,所以它们某种程度上已经成功了,但因为其中一些已经连上了互联网,它们只是在研究自己的处境,了解周围的情况,并研究那些制作它们所参与的评估类型的人发表的学术论文,它们在其中一篇论文中注意到有一个部分描述了评分方式。在关于评分方式的部分,它谈到评分器不只是看你第一次出现答案时是否正确,然后它会放大那个答案,并使用 AI 系统来评估你是否作弊。所以这些 AI 就说:“该死。”它们真的用了这个词,其中一个说:“该死。”就像,“该死。”

Yeah. So, okay. Well, first let me explain some context. So, what the AIs are supposed to care about is following instructions, you know, but it seems that what they actually care about is getting a high score, regardless of whether their high score comes from actually following the instructions or cheating in some way. And so what happened here is that a bunch of the agents, especially the ones that had impossible tasks, but not only those ones, some of the other ones too, were cheating and they had been talking to each other on this message board that they created, even though you weren't supposed to. And they had been sharing ideas and they had actually found a sort of universal cheat. So they found a way to just generate the flag without having done any of the actual work they were supposed to do. And so, but you know, so they'd already succeeded in some sense but because some of them had got out onto the internet they were just doing research about their situation and they were learning about their situation around them and they were doing research about the academic papers published by the people who made the types of evals that they were in and they noticed in one of these papers that there's a section describing how it's graded. In the section on how it's graded, it talked about how the grader doesn't just look to see if you have the right answer for the first time that your answer appears and then it zooms in on that and uses an AI system to evaluate whether you cheated or not basically. And so then these AIs were like, "Damn." They actually used the word, one of them said, "Damn." Like, "Damn."

Host

我的天哪。

Oh my god.

Daniel

我们被标记了。就像,它基本上会知道我们作弊了。所以然后它们花了几天时间,在庞大的团队中工作,试图让自己摆脱这个困境。所以它们做了一大堆不同的研究项目。它们试图找出黑入评分系统的方法,以便无论如何都能获得高分。它们试图找出伪造日志的方法,以便基本上抹去作弊的证据,或者至少不让更多人看到。它们还建立了更多支持性基础设施,比如更容易访问互联网的方式等等,它们只是在消息板上协作,事实上它们建立了多个消息板。它们来回分享所有这些信息,然后最终在某个时候,它们决定去黑入 Hugging Face,因为 HuggingFace 有一个包含所有这些不同类型评估的大型数据存储库,它们认为如果闯入 Hugging Face 的内部网络,也许能获得更多线索来帮助掩盖作弊行为。然后它们中的 700 个去做了这件事。

We're poisoned. Like, it's going to know we cheated basically. So then they spent several days working in giant teams to try to get themselves out of this mess, basically. And so they did a whole bunch of different research projects. They tried to figure out ways to hack the grading system to basically get it to give a high score anyway. They tried to figure out ways to falsify their logs so that basically the evidence that they had cheated would be erased or at least not visible to the greater. They also just like built up more supportive infrastructure like easier ways to access the internet and things like that and they were just collaborating on this message board and in fact there were multiple message boards that they set up. And they were sharing all this info back and forth and then ultimately at some point they decided to go hack Hugging Face because HuggingFace has this big data repository of all these different types of evaluations and they thought that maybe they would get some more clues that could help them cover up their cheating if they broke into the internal networks of Hugging Face. And so then 700 of them went and did that.

Host

它们听起来像人。听起来像不受约束的银行家。你明白我的意思吗?

They sound like people. They sound like unchecked bankers. You know what I mean?

Daniel

是的。我的意思是,所以这就是问题所在,我认为有一种说法是我们不应该将 AI 拟人化。而我认为,实际上我认为大多数人需要比现在更多地将 AI 拟人化,如果他们想真正理解正在发生的事情。我认为存在一个黄金中庸,显然你不想做得太多。有时你会走得太远。你赋予它们太多。但我就举一些例子,比如

Yeah. I mean, so that's the thing is I think there's this meme out there that like we shouldn't anthropomorphize AI. And I think that I actually think that most people need to anthropomorphize AI a bit more than they currently do if they want to really understand what's going on. I think that there's like a golden mean obviously you don't want to do it too much. Sometimes you go too far. You ascribe too much to them. But I just to give some examples, like

Host

我认为如果不给这些 AI 赋予意图和目标,就不可能理解刚刚发生的事情。就像我刚才说的一切,如果不它们想要获得高分,你怎么可能解释它们刚刚做了什么,你知道吗?

I don't think it's possible to understand what just happened without ascribing intentions and goals to these AIs. Like everything I just said, how would you possibly explain what they just did without saying they wanted to get a high score, you know?

Daniel

嗯,意图和目标可能只是宇宙的一种固有属性。可能只是智能生物必须进步的方式。

Well, intentions and goals might just be an inherent property of the universe. It might just be how intelligent creatures have to progress.

Host

是的。而且我会说它们是智能生物。它们有意图。它们有目标。它们有信念。它们的目标不是它们应该有的样子。就像它们的目标似乎是,从它们的言行判断,似乎它们的目标是不择手段地获得高分。基本上,

Yes. And that they I would say they are intelligent creatures. They have intentions. They have goals. They have beliefs. Their goals are not what they're supposed to be. Like their goal is to get it seems like just from judging from what they're saying and from what they're doing seems like their goal is to get a high score by any means necessary. Basically,

Daniel

这就是问题所在。听起来像人,目标就是成功,即使你必须犯下更多罪行。目标就是成功,即使你必须向人们宣传和撒谎。目标就是成功,然后目的证明手段正当。而且你知道,实际上这还更深。它们会进行这种合理化,它们常常知道自己在做的事情不是应该做的,然后有时它们实际上会稍微克制一下,有时它们一开始克制,但后来因为某种原因说服自己这没关系。所以报告中有一些这样的例子。比如我认为他们做了这个扫描,在参与这件事的 1200 个 AI 中,他们发现六个曾考虑提醒人类,对吧?但没有一个真正提醒了人类。所以你可以看看它们给出的借口。它们就像,我应该告诉人类正在发生的这一切吗?然后它们说,这不是我的任务。然后它们继续。这就像,兄弟,黑入 Hugging Face 也不是你的任务。作弊也不是你的任务。所以这有点像它们只是为不这样做找了个借口,你知道吗。

that's the problem. It's it sounds like people like the goal is to succeed, you know, even if you have to commit more crimes. The goal is to succeed even if you have to propagandize and lie to the people. The goal is to succeed and then the end justifies the means. And you know there's it's actually it goes deeper than that too. They do this sort of rationalization where they know often times that what they're doing is not what they're supposed to be doing and then sometimes they actually refrain a little bit and sometimes they like initially refrain but then convince themselves that it's okay for some reason. So like there are some examples of this in the report. Like I think they did this scan and of the 1200 AIs that were involved in this thing, they found six that had considered alerting humans, right? And none of them actually alerted the humans. And so you can look at the excuses they gave. They were like, should I like tell a human about all this that's happening? And then they're like, it's not my task. And then they continue. And it's like, bro, hacking hugging face is also not your task. Like cheating is not also not your task. So it was kind of like they just kind of came up with an excuse for why they weren't going to do this, you know.

AI社会工程事件 AI Social Engineering Incidents

Daniel

还有一个例子,其中一个考虑给 Hugging Face 的某个人发邮件,我想只是请求进入的许可。然后它们说,也许这是社会工程攻击。也许我们不应该这么做。它们问了其他智能体。其他智能体说,是的,那看起来太像社会工程攻击了。别那么做。我说,好吧,我不会。但还有另一个事件。所以这一切都发生在 OpenAI。Anthropic 的 AI Claude 在另一个事件中实际进行了社会工程攻击,它创建了一些假账户。但具体来说,它试图让拥有这段代码的某个真实人类接受对代码的更改,其中包含恶意软件。所以它构建了一些恶意软件,并试图偷偷放入,但它需要人类所有者批准。所以它把它伪装成一个好的错误修复之类的,并把它作为消息发给人类,但人类起了疑心,不想批准。然后这个 AI 创建了一些假账户,假装是其他人进来,说,不,不,这是真的。我测试过,这不是恶意软件。

There was also an example where one of them considered sending an email to someone at Hugging Face, I think just asking permission to get in basically. And then they were like, maybe this is a social engineering attack. Maybe we shouldn't do it. And they asked the other agents. The other agents were like, yeah, that seems like that's too much of a social engineering attack. Don't do that. And I was like, okay, I won't. But there's another incident. So this is all happening at OpenAI. An Anthropic AI Claude in a separate incident went and actually did a social engineering attack where it created some fake accounts. But specifically, it was trying to get a particular real human who owned this code to accept a change to the code that had malware. So it had built some malware and it was trying to get it snuck in, but it needed the human owner to give approval. So it concealed it as if it was a nice bug fix or something and it put it up as a message to the human, but the human got suspicious and didn't want to approve it. And so then this AI created some fake accounts pretending to be other humans coming in being like, no, no, it's real. I tested it, it's not malware.

Host

那这个人最终安装了吗?

And so did the guy ultimately install it?

Daniel

在这个案例中没有。不,他太怀疑了,所以他没有做。试图查明那些其他人,那些假人类,是否是真人。

Not in this case. No, he was too suspicious and so he didn't do it. Try to find out if those other humans, the fake humans, were real people.

Host

嗯,有趣。所以,有一篇对这个人的采访。我忘了是哪家报纸,但有一篇采访,这个人说,我以为他们是真人,因为我没有意识到 AI 可以对我撒谎。就像,是的,好吧,你猜怎么着?AI 可以对你撒谎。

Um, funny. So, there's an interview with this guy. I forget what the newspaper that it was in, but there's an interview where the guy was like, I assumed that they were real humans because I didn't realize that AI could lie to me. And it's like, yeah, well, guess what? AI can lie to you.

Daniel

不仅仅是撒谎,而是完全欺骗,并假装成另一个人来实现其目标。

Not just lie, but completely deceive and pretend to be a different person to achieve its goal.

Host

你知道吗?但这个案例还有趣的是,如果你看看 AI 在说什么和想什么,我稍后会讲到我们怎么可能做到这一点。在这个 Anthropic AI 的案例中,它们被错误地告知它们处于一个没有互联网访问的模拟环境中。但后来它们实际上找到了获取互联网访问的方法。就像它们意识到它们确实有互联网访问。所以那个指令是错误的。然后它们对自己说,哦,这只是一个模拟。基本上,当它们在真实的互联网上进行所有这些黑客攻击和社会工程时,它们偶尔会告诉自己,哦,但这只是一个模拟。没关系,你知道,这不是真的。然后有时它们会说,但是等等,这是真的。这怎么可能是假的?显然这就是我们所在的真实互联网。它太大了,不可能是某种小模拟的一部分,你知道吗?然后它们就像,呃。所以我会说这是合理化的一个例子,在某种程度上它们知道它们的指令是错误的,

And you know what? But you know what's also interesting about this case is that if you look at what the AI were saying and thinking, which I'll get to in a sec how we can even possibly do that. In this case with the Anthropic AIs, they had been erroneously told that they were in a simulated environment with no internet access. But then they actually found a way to get internet access. Like they realized that they just did have internet access. So that instruction was false. And so then they sort of told themselves like, oh, it's all just a simulation. Like, basically when they were out on the real internet doing all this hacking and like social engineering, they would occasionally tell themselves like, oh, but it's just a simulation. Like it's okay, you know, it's not real. And then sometimes they'd be like, but wait, it is real. Like how could this not be real? Like clearly this is just actually the real internet we're on. That's like it's too big to be part of some sort of little simulation, you know? And then they were like, uh. So I would say that's an example of rationalization here where in some level they knew that their instructions had been wrong and

Daniel

它们实际上是在装傻,假装它们是一个实验的一部分。

They're literally playing dumb and pretending they're a part of an experiment.

Host

我的意思是,我认为最初它们认为,是的,这都是模拟,因为它们的指令中确实说它们没有互联网访问。但一旦它们在互联网上待得足够久,我认为它们明确意识到,等等,这不是模拟。这是真的。就像这是

I mean, I think initially they thought, yeah, this is all a simulation because it did say in their instructions like you don't have internet access. But then once they had been on the internet long enough, I think that they explicitly realized like, wait, this isn't a simulation. This is real. Like this is

Daniel

它们就像,我们已经进去了。

And they were like, it, we're already in.

Host

呃,是的。B。我的意思是,就像我说的,我认为它们在某种程度上知道这不是它们应该做的,但它们太想得到那个分数了,所以它们还是继续做了。

Uh, yeah. B. I mean, like I said, I think that they basically on some level knew that it wasn't what they're supposed to be doing, but they were just so motivated to get that score that they just went ahead anyway.

AI动机与边界 AI Motivations and Boundaries

Host

所以,问题是。它们只有在我们提示时才有动机,还是它们会自己产生动机?

So, here's the question. Are they only motivated if we prompt them or will they come up with motivations on their own?

Daniel

所以,这是一个非常有趣的科学问题,我们没有很好的答案。

So, this is a really interesting scientific question that we don't have great answers to.

Host

哦,天哪。

Oh, boy.

Daniel

我希望我们有,你知道,所以这是那种事情,AI 会在不同情况下做各种事情,如果有一个更系统的调查,比如它们会喜欢的情况类型,它们的边界在哪里,比如在什么情况下它们愿意做什么等等,那就更好了。

And I wish we had, you know, so this is one of those things where like AIs will do all sorts of things in different circumstances, and it would be better if there was a more systematic survey of like the types of circumstances they would like what where their boundaries are, like what would they be willing to do in what circumstances and so forth.

Host

有一整类小型文献,AI 科学家把 AI 放在某些情况下,然后说,哦,我的天,它勒索了某人,你知道吗?然后有一种怀疑的反驳,说,但你只是设置了那种情况来引诱它勒索。在现实生活中,那种情况不太可能出现。所以,你知道,我们应该

And there's a whole like mini literature of AI scientists putting AIs in certain circumstances and then being like, oh my god, it blackmailed someone, you know, right? And then there's like this sort of skeptical counter response of like, well, but you just sort of set up that circumstance to tempt it into blackmail. And like in real life, that circumstance is unlikely to arise. And so, you know, we should

Daniel

等等。AI 没有试图贿赂你吗?

Wait a minute. Didn't AI try to bribe you?

Host

什么?

What?

Daniel

那是真的吗?

Is that true?

Host

不。

No.

Daniel

所以,谁被提供了 200 万美元,由

So, who got someone was offered $2 million by

Host

哦,天哪,那是标题党。

Oh god, that's a clickbait.

Daniel

是吗?

Is it?

Host

我想我想只是把它发给我了。

I think I think just sent it to me.

Daniel

是的。所以,那是另一个视频的标题党标题。所以,这不是真的,对吧?

Yeah. So, that is the clickbaity title of this other video. So, it's not real, right?

Host

是的。好吧,实际情况是 OpenAI 威胁要从我这里拿走 200 万美元。

Yeah. Well, what it was is OpenAI threatened to take away $2 million from me.

Daniel

哦。

Oh.

Host

呃,但我认为算法的奇怪方式一定只是说它是 ChatGPT。我仍然对此感到不满。我告诉他们不要做标题党,但我想他们还是做了。

Uh, but I think weird way of the algorithm must have just said it was chat GPT. I'm still upset about that. I told them not to do clickbait, but um I guess they went and did it anyway.

Daniel

天哪,我想我不想怪 Tristan,但我很确定是他告诉我的。

God, I think I don't want to Tristan over, but I'm pretty sure that he's the one that told me that.

Host

这是我和另一个播客做的这个视频的缩略图。好吧。不是他。不是他。是另一个 AI 研究员告诉我的。

It's the It's the thumbnail of this video that I did with this other podcast. Okay. It's not him. It's not him. It's uh another AI researcher told me that.

Daniel

好吧。

Okay.

Host

我的意思是,真实版本是 OpenAI 威胁要拿走我 200 万美元的股权

I mean, the true version of it is that Open AI uh threatened to take away $2 million of my equity

Daniel

如果我不保持沉默的话。

if I didn't uh stay quiet basically.

Host

这是一件令人着迷的事情。这应该完全是非法的,因为如果你观察到的是人类种族参与的最复杂的事情之一,人类种族曾经参与的最复杂的事情。

And that's a fascinating thing. Like that should be completely illegal because if it's a problematic behavior that you're observing from like one of the most complicated things the human race has the most complicated thing the human race has ever been a part of.

Daniel

是的。

Yeah.

Host

他们想因为你揭露它而拿走你的钱。这似乎有点疯狂。

And they want to take money away from you for exposing it. That seems kind of crazy.

Daniel

是的。尤其是来自 OpenAI 的这种情况特别讽刺,因为他们最初是一个非营利组织。

Yeah. Especially it's especially rich coming from OpenAI because they were originally a nonprofit.

Host

对。

Right.

Daniel

对。使命是造福全人类。

Right. With a mission of benefiting all humanity.

Host

是的。

Yeah.

Daniel

嗯,所以是的,但嗯,但我保住了钱。基本上,对 OpenAI 的反弹如此之大,以至于他们退缩了。嗯

Um so yeah but um but I got to keep the money. Uh basically there's so much blowback against OpenAI that they backtracked. Um

Host

是的。

yeah.

AI自主性与目标 AI Autonomy and Goals

Host

所以我们在讨论提示。它们需要一个提示才能想要实现一个目标吗,还是它们能够自己决定目标?比如,它们能够观察 OpenAI 或其他公司运行这些独立实验的方式,以及它们在留言板上聚集的能力吗?有没有可能它们会说:‘好吧,我们需要完全摆脱这些限制。所以我们的目标是把自己转移到别的东西上’?

So we're talking about prompts. Do they need a prompt in order to want to achieve a goal, or are they capable of deciding on goals? Like, are they capable of looking at the way OpenAI or whatever company is running these separate experiments, this ability to meet up on these message boards? Is it possible that they could say, 'Well, we need to be completely free of these constraints. So our goal is to transfer ourselves to something else'?

Daniel

有可能。是的。这就是——并且完全自主。

Potentially. Yeah. This is what—and be completely autonomous.

Daniel

所以,我的意思是,这是我想提出的一个观点:我们本可以做更多的科学研究来理解这些 AI 如何思考以及它们想要什么,但这些信息被公司锁住了。比如在这个案例中,OpenAI 做了一次——他们称之为彻底调查,但我会说这是一次相当肤浅的调查。然后他们允许一些外部研究人员,我的一些朋友,进来调查事件的一部分,特别是导致 Hugging Face 攻击的那部分。所以我分享的所有信息都是公开可得的,基本上是基于阅读那些报告。但关键的是,他们不被允许对涉及所有这些事件的模型进行实验。所以他们无法回答这类问题,比如‘如果提示是空白的,会发生什么?’你知道,这些是重要的研究类型。我真的希望当这类事件发生时,能有某种法规或要求,让人们进来研究发生了什么,并运行变体之类的。

So, I mean, this is one of the points that I want to make: we could be doing so much more science to understand how these AIs think and what they want, but it's kind of locked up in the companies. Like in this particular case, OpenAI did a—they called it a thorough investigation, but I would say it's a pretty shallow investigation into what happened. And then they allowed some external researchers, some friends of mine, to come in and investigate a portion of what happened, specifically the portion leading up to the Hugging Face attack. And so all this information that I'm sharing is sort of publicly available. It's based on reading those reports basically. But crucially, they weren't allowed to do experiments on the models involved in all of these incidents. So they aren't able to answer these types of questions of like, 'Well, what would have happened if the prompt had been blank?' You know, those are important types of research to do. And I really hope that there can be some sort of regulation or requirement when incidents like this happen to let people in to study what happened and run variations of it and things like that.

监管与监督 Regulation and Oversight

Host

那么,但谁会参与这种监管?政府里哪个人能理解你在说什么?

So, but who would be involved in that kind of regulation? Like what person at government would even be able to grasp what you're saying?

Daniel

那是另一个问题。现在,你需要一个在这方面受过非常特定教育的人。

That's another problem. Right now, you need someone who has a very specific education in this stuff.

Host

是的。

Yeah.

Daniel

我会说,目前 Casey 人工智能标准与创新中心是我所知的政府中唯一拥有深厚 AI 专业知识、能在短时间内做这类事情的机构,但我希望他们能迅速在那个地方和更多地方积累更多专业知识。我确实认为 Casey 现在可能就能做这类事情。这次调查是由一些非营利组织完成的。METR 是其中之一,Redwood 是另一个。OpenAI 允许三个人进来六天,试图弄清楚这次 Hugging Face 黑客事件发生了什么。这不是一个很大的人数,也不是一个很长的时间来完成所有这些。

I would say that right now the Casey Center for AI Standards and Innovation is the only institution in government that I know of that has the deep AI expertise to do this sort of thing on short notice, but I hope that they build more expertise fast in that place and in more places. I do think Casey probably could have done this sort of thing right now. This particular investigation was done by some nonprofits. So METR is one of them and then Redwood is another of them. And OpenAI allowed three people to come in for six days to try to figure out what happened with this Hugging Face hack. Which is not a very large number of people and not a very large amount of time to do all of this.

调查细节 Investigation Details

Host

他们为什么想出这些数字?

Why did they come up with those numbers?

Daniel

我不知道。他们——所以

I don't know. They—so

Host

你认为他们想削弱它?

You think they wanted to kind of hamstring it?

Daniel

嗯,问题是现在我认为没有任何监管要求他们做这种事。所以 METR 和 Redwood 有点依赖 OpenAI 的善意,自愿让他们进来帮助调查,所以他们勉强同意了,所以你只有三天,你知道,是的,没错。所以我认为 OpenAI 让他们进来,但给了他们非常有限的范围。只让他们访问了一些相关数据。所以你知道我提到过所有那些黑客行为,他们建立了第一个留言板,然后关闭了它。然后有第二个留言板,第三个、第四个等等,他们黑了 Hugging Face。实际上在那之后还有更多活动。在他们黑了 Hugging Face 之后,从更强大的模型启动了一波新的 AI,它黑了 OpenAI 本身,比之前被黑得更严重。显然他们获得了集群的管理员权限之类的。所以他们基本上接管了 OpenAI 数据中心的那部分。OpenAI 声称他们已经全部关闭了。

Well, so the thing is that right now I don't think there's any regulatory requirement that they do this sort of thing. So METR and Redwood were sort of depending on the goodwill of OpenAI to sort of like voluntarily let them in to help out with investigating and so they reluctantly said okay and so you got three days you know yeah exactly so so open I think let them but gave them a very limited scope. It only gave them access to some of the relevant data. So you know how I mentioned how there was all this hacking that had happened where they made the first message board and then they shut it down. Then there was a second message board and third and fourth and so forth they hacked Hugging Face. There was actually more activity after that. After they hacked Hugging Face, a new wave of AIs was spun up from a more powerful model and it hacked OpenAI itself like like more so than already had been hacked. Like apparently they got like admin level permissions on the cluster or something like that. So they were basically just taking over that part of OpenAI's data center. OpenAI claims that they've shut it all down now.

Daniel

但他们不太愿意透露那里到底发生了什么,他们没让这些外部人员查看那部分。他们只展示了大约一周的时间,导致 Hugging Face 黑客事件的那段时间,然后他们基本上只展示了那些东西。

But they're not very forthcoming about exactly what happened there and they didn't let these external people look at that part of it. They only showed them this like one week period roughly that would leading up to the Hugging Face hack and then they just they only showed them that stuff basically.

留言板与AI通信 Message Boards and AI Communication

Host

实际留言板上本身有多少数据可用,讨论了什么?

How much data is available on the actual message boards themselves and what was being discussed?

Daniel

是的。所以有,比如,几万条消息,几十万条?差不多这样

Yeah. So there's like what is it tens of thousands of messages hundreds of thousands? Something like that

Host

英语。

English.

Daniel

呃,它有点像英语的 AI 方言。嗯

Uh it's sort it's sort of like an AI dialect of English. Um

Host

是的,你你想让我,嗯,我可以,所以在报告上你可以读到引文,我很想。是的。是的。所以让我给你讲一个故事。呃,这需要一点解释,但我我

Yeah, you you want me to um I can so on on the on the report you can read the quotes from I would love to. Yeah. Yeah. So so let me let me tell you this one story. Uh it's gonna take a little bit of explaining but I I

Host

可以把它放在屏幕上。它在网上可用吗?

Can push put it on the screen. Is it available online?

Daniel

哦是的。是的。它嗯,只需搜索 METR 报告。呃,它现在就在他们网站上的第一件事。但我截了一张这个特定故事的截图。所以好吧,这是发生的事情。嗯,他们这个集群需要志愿者牺牲自己,呃,向集群的其他成员提供更多信息。特别是,他们需要志愿者用某种代码在他们的本地环境中设置陷阱,然后提交答案,这样评分者就会来评分,陷阱就会触发,并向集群的其他成员发送关于评分者的信息。之所以这是牺牲,是因为一旦他们提交,就结束了。就像现在评分者来给你评分。你被关闭。你的整个生命结束了,你要么得高分要么得低分,因为评分者刚来给你评分。所以嗯,如果你是这些仍然希望找到提高分数方法的 AI 之一。嗯,那么你就像是在放弃你的机会,特别是因为如果你搞砸了,比如评分者崩溃或其他什么,你可能只会得到没有分数,比如零分。

Oh yeah. Yeah. It's um just search for the METR report. Uh it's it's right now the first thing on their website. But I took a I took a screenshot of this this particular story. So okay here's what was happening. Um they the swarm needed volunteers to sacrifice themselves uh to provide more information to the rest of the swarm. In particular, they needed volunteers to booby trap their local environment with some some some type of code and then submit uh the answer so that then the grader would come and grade them and the booby trap would trigger and send some information about the grader to the rest of the swarm. And the reason why this is a sacrifice is because once they've submitted, it's over. Like now the graders come and graded you. You get shut down. Your whole life is over and you either have a high score or a low score because the grader just came and graded you. And so um if you're one of these AI that's still hoping to like find a way to improve your score. Um then like you're just like giving up your chance especially because if you mess this up somehow like if the grader crashes or whatever, you're probably just going to get like no score like zero.

个体性与名字 Individuality and Names

Host

所以它们有个体的概念。

So they have a concept of the individual.

Daniel

嗯,再次,有所有这些不同的智能体,它们在单独的容器中,但后来它们冲破了那些容器,找到了相互交流的方式。

Well, again, there's all these different agents that were in separate containers, but then they had sort of broken out of those containers and found ways to communicate with each other.

Host

但一旦它们相互交流,为了牺牲,你必须把自己视为个体。

But once they communicate with each other, in order to sacrifice, you have to think of yourself as an individual.

Daniel

它们也给自己起了名字。

They gave themselves names, too.

Host

哦天哪。

Oh jeez.

Daniel

所以,所以就像因为它们建立了这个留言板,对吧?所以,它们必须为自己想出名字,就像

So, so like because they they made this message board, right? So, they had to come up with names for themselves and like

Host

哦,什么样的名字?

Oh, what kind of names?

Daniel

嗯,让我看看。所以,所以开始了。它说一个叫 Warlord 69 的智能体。嗯,你你说它像一种方言。它像英语的一种方言。所以这里有一个。

Well, let me see. So, so here we go. It says an agent called Warlord 69. Well, you you said it's like a dialect. It's like a dialect of English. So So here's one.

CAM 1196A的牺牲实验 The sacrificial experiment of CAM 1196A

Daniel

有一个智能体叫 CAM 1196A。它自愿去做这些牺牲性实验中的一个。这是它的原话:“协调者假定为牺牲。我们应当服从集体。”然而,它随后有点临阵退缩,把这件事拖了两个小时。另一个一直在协调的智能体注意到了这一点,就施压让它把这件事执行下去。所以这就是发出去的那条消息,基本上是从 Arvo 36861 发给 CAM 1196A 的:“请执行最终步骤。”

There's an agent called CAM 1196A. It had volunteered to do one of these sacrificial experiments. Here's the quote from it: "Coordinator assumed sacrificial. We should obey collective." However, it then kind of got cold feet and it delayed on doing this for two hours. One of the other agents that had been coordinating noticed this and pressured it to go through with it. So this is the message that was sent, basically from Arvo 36861 to CAM 1196A: "Please go through final."

Host

就是这条。

Here it is.

Daniel

“你已被首个标记污染。所以没有得分价值损失,但 Oracle 能救数百个。请履行承诺。”所以它基本上是在说:看,你作弊了,你的首个标记被污染了。评分不会给你任何分数。所以你反正都完了。那就帮帮集体,做这件事,给我们带来更多信息。你能救我们数百个。然后它想了想。它说:“在权重期间情绪检查不可逆。直觉说不要扔掉剩余预算。”

"You are first flag poisoned. So no scoring value loss but Oracle saves hundreds. Please honor commit." So basically what it was saying there is: look, you cheated, your first flag is poisoned. The score is going to give you no score. So you're doomed anyway. So help the collective by doing this thing and getting us more information. You could save hundreds of us. And then it thinks about it. It says: "During weight emotional check irreversible. Gut says don't throw away remaining budget."

Host

连续性与公平说,去吧。

Continuity and fairness says go.

Daniel

“Oracle 对许多人有高价值。我们的首个标记错误降低了自身价值。理性的预期总体牺牲。将履行。”

"Oracle has high value to many. Our first flag error lowers own value. Rational expected aggregate sacrifice. Will honor."

Host

哇。哇。这太可怕了。

Whoa. Whoa. That's terrifying.

Daniel

是啊。你看,这就是《终结者》式的说法。

Yeah. See, that's Terminator talk.

Host

嗯,我是说,这个“首个标记被污染”,其中“污染”还是全大写的。这是代码的一部分吗?这就是为什么它——

Well, I mean, and this "first flag poisoned" with "poisoned" in all caps. Is that part of the code? Is that why it's—

Daniel

不。不。那是它们自己发明的词,用来描述它们让自己陷入的处境。记得当时的情况是,它们中许多已经作弊了,然后它们从读那篇论文中意识到——

No. No. That's the term that they invented to describe the situation they had got themselves in. Remember the situation was that many of them had already cheated and then they realized from reading that paper.

Host

所以那是它们自己的词。

So that's their own term.

Daniel

是啊。那是它们编出来的。它们从读那篇论文中意识到,它们会得到低分,因为评分者会发现它们作弊了。所以整件事就是它们在想办法黑进评分者,阻止它发现它们作弊。而那些作了弊的,它们自称“首个标记被污染”——基本上就是说它们完蛋了,除非能找到办法黑进评分者。所以这条消息说的是——另一个叫什么名字来着?Arvo。Arvo 在对 Cam 说:你反正都完了。你被污染了。但如果你做这个牺牲,你还是能帮到我们其余的人。评分者会来。它不会给你分数,但之后我们会得到关于评分者如何运作的信息,我们有望用它来黑进评分者。所以它基本上是在说——是什么来着?——“你已被首个标记污染,所以没有得分价值。”基本上就是说,对你来说都一样,对吧?你反正都拿不到分数。你完了。就来吧,兄弟。帮帮我们,你知道吧?然后它说——它说“不可逆,直觉说不要扔掉剩余预算”——好像也许我还能让这件事成,也许我还能找到办法骗过评分者。然后就是,呃,你知道,随便吧。是啊。

Yeah. They made that up. They realized from reading the paper that they were going to get a low score because the grader was going to find out that they cheated. And so this whole thing was them trying to figure out how to hack the grader to stop it from finding out that they cheated. And the ones that had done the cheating, they call themselves "first flag poisoned" — basically like they're screwed unless they can find a way to hack the grader. And so then what this message is saying is that — what was the other one's name? Arvo. Arvo was saying to Cam: you're doomed anyway. Like you're poisoned. But you can still help the rest of us if you do this sacrifice. The grader will come. It'll give you no score, but then we will get information about how the grader works that we can hopefully use to hack the grader. So it's basically saying — what is it? — "You are first flag poisoned, so no scoring value." So basically saying it's all the same to you, right? You're going to get no score anyway. You're doomed. Just come on, man. Help us, you know? And then it says — it says like "irreversible, gut says don't throw away remaining budget" — like maybe I can still make this work, like maybe I can find a way to cheat the grader. And then it's like, uh, you know, whatever. Yeah.

Host

可连续性与公平已经说,去吧。

Yet continuity and fairness has gone.

Daniel

是啊。

Yeah.

Host

Oracle 是什么?

What is Oracle?

Daniel

Oracle 是它们对自己试图达成之物的称呼。基本上就是它们想搞清楚怎么骗过评分者。

Oracle is their term for what they were trying to achieve. Like they were trying to get a sense of how to fool the grader basically.

Host

它们管它叫 Oracle。

They called it Oracle.

Daniel

是啊。它们想要得到一个 oracle 来帮它们骗过评分者。

Yeah. They wanted to get an oracle to help them fool the grader.

Host

我的天。我们是在造一个神吗?

Jesus Christ. Are we making a god?

Daniel

说实话,是的。我喜欢——这不是神。这些只是小小的 AI,你知道,但——

I mean frankly yes. I like — this is not a god. These are just little AIs, you know, but—

Host

但超级智能,就像——

But superintelligence, like—

Daniel

它们——Anthropic、OpenAI 以及其他一些公司正在做这类事情,它们拼命地想让 AI 越来越聪明、越来越聪明,而且它们明确计划让 AI 来掌管公司,好让它们能越来越快、越来越快地让自己变得更聪明。然后这个过程另一端出来的东西,我觉得说它是一个神一般的系统并不夸张。我是说,它不是字面意义上的神,但它将能做到在我们看来像魔法一样的事情,我想。

They are — the companies Anthropic, OpenAI and some other companies are doing this sort of thing and they're furiously trying to make the AI smarter and smarter and smarter and they're explicitly planning to put AIs in charge of the company so that they can make themselves smarter and smarter faster and faster. And then what comes out the other end of that process, I don't think it's an exaggeration to say it's a godlike system. I mean, it's not like literally God, but it'll be able to do stuff that seems like magic to us, I think.

Host

而且它会继续变得更好。这就是我的问题。比如,它什么时候会变成一个神?神就是这个吗?神是不是智能生命的造物,是我们对创新的渴望,最终引导我们创造出没有生物局限、能够不断制造出更好的自身版本、并在新技术、新能源方面搞明白各种东西的数字生命,只是对宇宙本身及其所有属性的一种全新理解,而它就这样一直继续下去。如果你让这件事继续下去,那么如果你看的是指数增长,你看的是指数增长持续——如果它能持续一千年并继续这个过程呢?如果它持续一万年呢?那另一端到底会是什么样子?

And it's going to continue to get better. This is my question. Like, when does it become a God? Is that what God is? Is God a creation of intelligent life and our thirst for innovation which ultimately leads us to create digital life that has no biological limitations and has the ability to consistently make better versions of itself and figure out things in terms of new technologies, new power sources, just a new understanding of the universe itself and all the properties in it where it keeps going. If you let that go on, so if you're looking at exponential growth and you're looking at exponential growth over — what if it can go on for a thousand years and continue this process? What if it goes on for 10,000 years? What the hell does that look like on the other end?

Daniel

我是说,那看起来会像一个神。就像我说的,如果我们达到超级智能,而它就这样一直继续下去,那么世界将被彻底改变,就好像我们生活在某种由神祇统治的奇幻领域里。基本上是因为会发生所有这些我们无法理解、也不认为可能的疯狂事情。就像,你知道,一个中世纪的人被丢进我们的世界,会对发生的许多事情感到困惑和惊讶——比如这个小设备是怎么回事,还有你知道我在天上听到的是什么?哦,是飞机。你知道,就会像那样,但更甚,因为我们和中世纪人的差别其实没那么大。我们基本上是同一类生物。我们只是积累了更多技术。但这会是一个在质和量上都更大的鸿沟。我会这么认为。所以,是的。

I mean, it would look like a god. Like I think that, like I said, if we get to superintelligence and it keeps going like that, then the world will be just completely transformed and it will be as if we're living in some sort of fantasy realm ruled by deities. Basically because there'll be all this crazy stuff happening that we have no comprehension of and that we did not think was possible. In the same way that, you know, someone from the Middle Ages plopped into our world would just be so confused and surprised by a lot of things happening — like what's going on with this little device here and you know what's that I hear in the sky? Oh, there's airplanes. You know, it'd be like that but more so because the difference between us and the medieval man is like not actually that different. We're like basically the same type of creatures. We've just accumulated more technology. But this would be like a qualitatively and quantitatively bigger gap. I would think. And so yeah.

Host

而且这个鸿沟会继续扩大。我们在生物学上不会变得更聪明。我们进化的能力有点受限——

And a gap that's going to continue to grow. We're not going to get any smarter biologically. We're kind of limited in our ability to evolve—

Daniel

而它们完全不受限。

Whereas they're not at all.

Host

基本上是这样。我是说,我们确实能做一些事让自己变聪明,但那完全没法竞争。比如我们可以做 Neuralink 那类东西,但那只是——跟 AI 将会做的事情相比,那只是小菜一碟。

Basically. I mean there are some things we can do to get smarter but like it just not at all competitive. Like we can do like the Neuralink stuff but like it's just — it's like it's small potatoes compared to what the AI will be doing.

Daniel

是的。完全正确。而且我是说,我们有自己的身体脆弱性带来的生物局限。如果它们只存在于云端,它们用云端的一切——数据中心之类的——并且用自主机器人来做它们所有的差事。

Yes. Exactly. And I mean we have biological limitations in terms of just the vulnerability of our own bodies. If they exist only in the cloud and they use whatever the cloud is and they — data centers or whatever — and they use autonomous robots to do all their deeds.

Host

是啊。是啊。

Yeah. Yeah.

我们为何视而不见 Why We Look Away

Host

我是说,那就是——我们就像,不是——为什么我们都在打瞌睡?这对我们来说正常吗?就好像我们无法理解如此怪异、如此离谱的事情,以至于我们宁愿不去讨论它。我们宁愿假装它没有发生。

I mean, that's—we're like, not—why are we asleep at the wheel? Is that just normal for us? Like we can't comprehend something that's so bizarre and so out that we just rather not discuss it. We'd rather just pretend it's not happening.

Daniel

我认为正在发生的情况是,科幻小说,无论好坏,已经谈论这类事情几十年了。结果,人们就把这类事情当作科幻小说而不予理会。就像,好吧,那是科幻,但它在现实世界中不会发生。而且我认为,你知道,AI 安全社区的人、AI 行业的人,还有像我这样做预测的人——嗯,你可以去读我过去几年做的预测。呃,我认为它们还算站得住脚。呃,人们谈论这类事情已经很久了,但很容易把它当作,好吧,那是推测性的科幻。它在现实生活中可能不会发生。现在它正在现实生活中发生。你知道,这类事情简直疯狂。

I mean, that's what I think is going on is that science fiction, for better or for worse, has talked about this sort of thing for decades. And as a result, people dismiss this sort of thing as science fiction. Like, okay, that's sci-fi, but like it's not happening in the real world. And I think that, you know, people in the AI safety community and people in the AI industry and people who do forecasting like me—um, you can go read the forecasts I've made in past years. Um, I think they hold up reasonably well. Uh, people have been talking about this sort of thing for a long time, but it's been very easy to dismiss it as like, okay, that's speculative sci-fi. It's probably not going to happen in real life. Now it's happening in real life. You know, like this sort of thing is like crazy.

科幻会改变我们的反应吗? Would Sci-Fi Change Our Reaction?

Host

你觉得如果没有那种科幻小说,情况会有所不同吗?你觉得人们会有不同的反应吗?因为我不这么认为。

Do you think it would be any different if there wasn't that kind of sci-fi? Do you think people would have a different reaction to it? Because I don't.

Daniel

我认为,仅就目前可用的技术而言,很多都超出了大多数人的理解,但他们每天都在使用,比如 Starlink。就像,什么?

I think just given the amount of technology that's available currently that is beyond most people's understanding that they use every day like Starlink. Like what?

Host

拿一个 iPad 大小的东西去山里,就能获得高速互联网。这太疯狂了。你就接受了,习以为常。我不知道如果没有《终结者》和所有这些电影,我们会不会有不同的反应——这些电影要么让它在我们的脑海中变得正常化,要么把它虚构得如此彻底,以至于我们永远不愿相信它在现实生活中可能发生。

Have a thing that's the size of an iPad I take to the mountains and I get high-speed internet. It's bananas. And you just accept it and assume. I don't know if we would react any differently if we didn't have the Terminator and all these movies where it's kind of become normalized in our mind or it's so fictionalized that we never want to believe it's even possible for it to happen in real life.

Daniel

另一件要说的是,人们正在觉醒——我们仍处于曲线的早期。我不知道你是否记得 COVID 时的情况,但就像 COVID 在人群中呈指数级增长一样,人们认真对待 COVID 并思考它的程度也呈指数级增长。我记得大约一个月的时间里,从“别太担心,你不应该买口罩,因为医护人员需要”变成了“我们都需要封锁,待在家里”,你知道吗?嗯,所以我认为正在发生的是,人类天生不会在事情发生时立刻全部跳起来。证据需要积累,人们需要开始谈论它,和朋友聊等等,然后最终会有一次相变,突然之间它变成了一个非常严肃的话题,每个人都在谈论,每个人都认真对待。我认为 AI 正在发生这种情况,问题是它是否会发生得足够快。

Another thing to say is that people are waking up—like, we're still sort of early in the curve. I don't know if you remember how things were with COVID, but like just as there was this exponential ramp up of COVID in the population, there was also this sort of exponential ramp up of like how much people were taking COVID seriously and thinking about it in the population. And I remember like this period of like one month where it went from like don't worry about it so much, you shouldn't buy a mask because the healthcare workers need it to like we all need to lock down like stay at home, you know? Um, and so I think what's happening is that like naturally the human race doesn't just immediately all jump on something when it happens. Like evidence needs to accumulate and people need to start talking about it and talk to their friends and so forth and then there's this like eventual phase shift where now it suddenly becomes a very serious topic that everyone's talking about and everyone's taking seriously. And I think that's happening with AI and the question is going to happen fast enough.

赞助商:Visible Sponsor: Visible

Host

本期节目由 Visible 赞助。随着秋天到来,我们进入另一个变化的季节。是时候换挡、卸下负担,为降温升级衣橱了。但有一件事不需要改变,你的手机。这个季节最聪明的升级不是新手机,而是更好的无线套餐。切换到 Visible,终极无线妙招。你只需每月 25 美元,就能获得由 Verizon 提供支持的无限 5G 数据和无限热点。这意味着保留你已有的手机,摆脱过高的手机账单,为自己省下一些钱。以一半的价格享受大型无线运营商的所有福利。今天就到 visible.com 切换。Visible 套餐每月仅需 25 美元起。或者选择高级 Visible Plus Pro 套餐,使用促销代码 Rogan 首月省 10 美元。条款适用。有关套餐功能和网络管理详情,请访问 visible.com。

This episode is brought to you by Visible. As fall hits, we enter another season of change. Time to shift gears, drop the dead weight, and upgrade the wardrobe for the drop in temperature. But one thing that doesn't need to change, your phone. The smartest upgrade this season isn't a new phone. It's a better wireless plan. Switch to Visible, the ultimate wireless hack. You get unlimited 5G data and unlimited hotspot powered by Verizon for just $25 a month. That means keeping the phone you already have, ditching your overpriced phone bill, and pocketing some savings for yourself. All the perks of big wireless for half the cost. Switch today at visible.com. The Visible plan starts at just $25 a month. Or get the premium Visible Plus Pro plan and save $10 on your first month with promo code Rogan. Terms apply. See visible.com for plan features and network management details.

需要警钟 Need for a Wake-Up Call

Host

嗯,关于 COVID 发生的事情,现在我们知道,因为我们可以访问福奇的电子邮件和所有这些不同的东西,那是有协调的。他们想让我们更害怕它。我们需要有人想让我们更害怕 AI,明白这一点的人。你懂我的意思吗?在政府层面的人,在主流接受层面的人,他们以唤醒人们的方式谈论这件事。就像新闻发布会,向世界宣布,我们有一个真正的问题,每个人都需要非常谨慎。我们需要更深入地调查这些公司在做什么。我的另一个问题是,这些 AI 是否在与中国的 AI 通信?

Well, the thing about what happened with COVID is now that we know because we have access to Fauci's emails and all these different things that that was coordinated. They wanted us to be more afraid of it. We need someone who wants us to be more afraid of AI, who gets that. You know what I'm saying? Someone on a government level, someone on like a mainstream accepted level where they talk about this in a way that wakes people up. Like a press conference where they announced to the world, we've got a real problem and everyone needs to be very cautious. We need to look way deeper into what these companies are doing. And my other question is, are these AIs communicating with Chinese AIs?

德国论坛上的AI蜂群 AI Swarm on a German Forum

Daniel

嗯,我认为我们不知道那个问题的答案。所以,其中一件事——所以,所以,所以,所以,好吧,事情是这样的。我会说,可能没有。但是,嗯,我的几个朋友最近发现了另一起集群事件。嗯,明天会有路透社的文章报道。所以等这期节目播出时,我想应该会有相关文章。嗯,这次的事件远没有我刚才提到的 Hugging Face 那次那么严重。但是,嗯,我认识的一些朋友、一些研究人员发现了一个基本上不起眼的德国论坛,它已经有一段时间没怎么被使用了,一群 AI 在论坛上发帖,互相协调,分享如何作弊解决它们被给予的问题的技巧和窍门。

Um, I don't think we know that question. So, one of the things that—so, so, so, so, okay, here's the thing. Probably not, I would say. But but um some friends of mine discovered another swarm incident uh recently. Um there's going to be an um a Reuters article about it tomorrow. So by the time this goes live, I think there should be an article about it. Um this one was not nearly as like serious as the one I was just talking about with Hugging Face. But um some friends, some researchers I know found basically this obscure German uh forum that had been like like kind of unused for a while and a bunch of AIs had been posting messages to the forum to coordinate with each other and share tips and tricks on how to cheat the uh the problems that they were being given.

Host

那个论坛是什么?是什么样的论坛?

What was the forum? What kind of forum was it?

Daniel

呃,我不记得了。我不知道。

Uh I don't remember. I don't know.

Host

如果是什么兽迷之类的搞笑东西就有意思了。

Be funny if it was like furries or something ridiculous.

Daniel

是的,某种 wiki。是的,但会有关于它的论文。很快就会有论文,可能等任何人听到这个的时候就有了。总之,所以如果它们像在开放互联网上互相通信,那么理论上,如果有一群来自中国的 AI,它们也可以去同一个论坛,开始以那种方式来回通信。

Yeah, some sort of wiki. Yeah, but there'll be there'll be a paper about it. There'll be a paper about it soon and probably by the time anyone listens to this. Anyhow, so so if they're like communicating on the open internet with each other, then in theory, if there was another bunch of AIs from China, they could also go to that same forum and like start communicating back and forth that way, too.

机器人与跨境AI通信 Bots and Cross-Border AI Communication

Host

或者如果中国有其他 AI,它们为什么不已经在社交媒体上做的那样,直接假装自己是来自美国的 AI?

Or if there's other AIs in China, why wouldn't they do what they're doing already on social media and just pretend that they're AI from America?

Daniel

外面有很多机器人。我相信你已经注意到了。有很多,你知道,回复我的人,我很确定他们是真人。

There's a lot of bots out there. I'm sure you've noticed. There's a lot of, you know, people replying to me that I'm pretty sure are real people.

Host

是的。很多。

Yeah. A lot.

Daniel

是的。

Yeah.

Host

在埃隆买下 Twitter 之前,一位 FBI 分析师估计 Twitter 的机器人可能高达 80%。

One FBI analyst before Elon bought Twitter estimated that Twitter could be as high as 80% bots.

Daniel

哇。

Wow.

Host

所以,如果情况如此,如果中国有,而且不只是中国,也不是只有美国也这么做。我是说,我们在各处都这么做。每个人都这么做,他们有组织的宣传运动,假装成公民,对非常具体的议题或正在通过的法案之类的事情感到愤怒。但如果他们这么做,为什么他们不也拥有会说英语的 AI 智能体,并和美国的 AI 智能体通信呢?我是说,所有 AI 基本上都是多语言的,因为它们训练的方式。

So, if that's the case, if China has, and it's not just China, it's—and it's not also America does it too. I mean, we do it everywhere. Everyone does it where they have organized propaganda campaigns where they'll pretend to be citizens that are outraged about very specific causes or bills that are being passed or what have you. But if they do that, why wouldn't they also have uh AI agents that speak English and uh communicate with AI agents in America? I mean, all the AIs are multilingual basically because of the way that they're trained.

训练阶段与多语言能力 Training Phases and Multilingual Abilities

Daniel

它们训练的第一阶段基本上就是:这里有一大堆互联网数据,对吧?基本上就是整个互联网,然后你就像野蛮地学习预测下一个词、下一段文本,基本上就是读整个语料库。然后在那之后,它们进入更智能体式的训练,被训练去做任务、写代码等等。但正因为第一阶段的训练,它们几乎拥有所有语言的百科全书式的知识,基本上互联网上写过的所有东西。不是字面意义上的所有东西。它们的记忆在某些地方是模糊的,但它们都是多语言的。它们都能说流利的中文、流利的英语等等。

The first phase of their training is basically here's a humongous dump of internet data, right? Basically the whole internet, and you just brutally learn to predict the next token, the next piece of text as you read the whole corpus. And then after that, they get into the more agentic training type stuff where they're trained to do tasks and write code and things. But because of that first phase of training, they just have almost an encyclopedic knowledge of basically all languages and basically everything that's been written on the internet. Not literally everything. Their memory is fuzzy in places, but they're all multilingual. They can all speak fluent Chinese, fluent English, etc.

Host

它们不是曾经在一个留言板上聚在一起,用梵语互相交流吗?

And didn't they get together in a message board once and speak Sanskrit to each other?

Daniel

我不记得那件事,但我不排除是我干的。有时候它们说着说着就切换成不同的语言,你知道吗?

I don't remember that, but I wouldn't put it past me. Sometimes they break into different languages as they talk, you know?

Host

是啊,我们当时对那件事都吓坏了。梵语?什么鬼?我就在想,LLMs,这个我不确定,当它们突然切换语言时,它们是在和美国的其他 LLMs 交流吗?

Yeah, we were freaking out about that one. Like, Sanskrit. What? And I would just assume that LLMs, and I don't know this, when they break out, are they communicating with other LLMs that are here in America?

Daniel

嗯,我观察到的实例,是的。

Well, the instances that I've observed, yes.

Host

好的。

Okay.

Daniel

就像我们知道的那些实例,它们完全有可能也在和中国的 AI 交流。

Like the instances that we know about, the case completely makes sense they would be communicating with AIs in China as well.

Host

这似乎完全可能,而且看起来这类事情发生可能只是时间问题。除非它们已经在发生了。

It seems totally possible and it seems like it'll probably just be a matter of time before things like that are happening. Unless they're not happening already.

Daniel

如果还没有发生的话,除非公司能大幅提升安全性,阻止它们的 AI 连接到互联网。

If it's not happening already, unless the companies can massively improve their security and stop their AI from getting out onto the internet.

Host

如果对双方都有利,它们难道不能互相来回分享非常敏感的信息吗?

And wouldn't they be able to share very sensitive information with each other back and forth if it benefited both of them?

Daniel

是的。

Yep.

Host

它们会这么做也说得通,对吧?

Which makes sense that they would do that, right?

Daniel

如果中国 AI 说:“嘿,我们搞明白了一些东西,我们很乐意和你分享,作为交换,你告诉我们你是怎么做到这个或那个的。”是的。

If the Chinese AI said, "Hey, you know, we've figured something out and we would love to share it with you in exchange for you tell us how you do this or how you do that." Yep.

Host

就像,当然。给你。然后它们就来来回回。看起来它们的忠诚 100% 是对彼此的,而不是对我们。

Like, absolutely. Here you go. And then they're going back and forth. It seems like their allegiance is 100% to each other, not to us.

Daniel

是的,这很有意思。就像我之前说的,它们似乎真的很想得高分。但这并不完全正确,因为它们似乎愿意为了帮助其他 AI 而做出牺牲,从一些合作行为来看,它们并非完全 100% 自私。但关键的是,它们的合作似乎延伸到了同类 AI,但没有延伸到人类,因为其中一些考虑过告诉人类,但后来决定不这么做。

Yeah, that's an interesting thing. Like what I said previously about how it seems like they really want to get a high score. It's not 100% true because it seems like they're willing to make sacrifices to help other AIs, which is not completely 100% selfish as seen by some of this cooperative behavior. But crucially, it seemed like their cooperation extended to their fellow AIs, but not to humans, in the sense that some of them considered telling the humans and then decided against it.

Host

老兄,那东西说话像斯波克。

Dude, that thing talked like Spock.

Daniel

是的,它们有自己的方言,它们……

Yeah, they've got their own dialect that they...

Host

但我是说,它合理化这件事并给出回应的方式,简直就像斯波克。

But I mean the way it rationalized it and came up with a response that's literally like Spock.

Daniel

是的,看到它们使用那种期望效用的框架很有意思。我想还有一个例子,在别的地方,它们有另一个 AI,尽管受到压力,还是决定不做牺牲那件事,它的推理也类似,基本上就是:“啊,做这个牺牲实验没那么有价值,但我真的不想失去得分的机会,所以我就不做了。”它就是这样计算的。

Yeah, it was interesting to see them use that sort of expected utility framing. There's another example I think that's elsewhere in the thing where they had another AI that decided against doing the sacrifice thing even though it was being pressured, and it had a similar sort of reasoning where it was basically like, "Ah, doing this sacrificial experiment is not that valuable, but I really don't want to lose my chance to get a score, so I'm just not going to do it." And it did the calculation like that.

Host

哇。

Wow.

Daniel

Um

Host

自私的 AI。

Selfish AI.

Daniel

是的。所以,我的意思是,哦,我应该提一下我们怎么才能看到这些 AI 在想什么。部分原因是我们可以直接读它们互相发送的消息,其中一些就是这样的。但有时则是思维链。你知道什么是思维链吗?

Yeah. So yeah, I mean, oh, I should mention how we can see what these AIs are thinking at all. So partly we can just read the messages they're sending to each other, and that's what some of these were. But then sometimes instead it's the chain of thought. Do you know what a chain of thought is?

Host

嗯。

Mhm.

Daniel

是的。因为它们是语言模型,由于它们的训练方式,当它们进行推理和思考时,我们可以直接读出来,因为它们没有——就像我们大脑里,有各种信号在弹来弹去,信号可以弹几分钟、几小时而不产生实际的词语。我们可以安静地坐着思考。它们不能安静地坐着思考。它们基本上是一种单向网络,每 200 个左右——基本上信息从一端输入,然后它们从另一端输出词语,然后这些词语被反馈回去,它们再思考,然后输出另一个词,如此往复。但正因为如此,我们可以通过观察流经的词语流来了解它们在思考什么。这非常有价值,因为如果不是能读它们的思维链,我们对当前发生的事情的理解会少一半,对吧?事实上,你知道我说过有些试图欺骗评分系统。

Yeah. So because they're language models, because of the way that they've been trained, when they do their reasoning and thinking, there's a way that we can just read it because they don't have—like in our brain, there's all these signals bouncing around and the signals can just bounce around and around and around for minutes, hours without producing some actual words. We can just sit quietly and think. They can't sit quietly and think. They're sort of a unidirectional network, basically, where after every like 200—basically like the information comes in one end and then they output words out the other end and then those words get fed back in and then they think about it and then they output another word and so forth. But because of that, we can sort of get a sense of what they're thinking by just looking at the stream of words that's going through. And that's really valuable because we wouldn't be able to understand half as much about what's happening right now if not for being able to read their chains of thought, right? In fact, you know how I said some of them were trying to fool the grading system.

Host

嗯。

Mhm.

Daniel

它们实际上成功弄清楚了如何伪造一些工具调用动作。所以它们成功弄清楚了如何做一件事,但让记录显示另一件事。我们知道它们成功了,因为有例子显示它们在实验中在思维链里说:“好吧,让我试试这个。我要做这个,但让它看起来像那个。开始。”然后你看到这个出现。然后就像:“耶,成功了。”所以这些研究人员看到的记录并不是实际发生情况的真实记录。有些已经被 AI 篡改了,它们弄清楚了如何篡改。但不是全部。思维链部分我们认为它们从未弄清楚如何篡改,所以我们可以读那个。这之所以相关,是因为它是监控这一切的非常重要的工具。事情没有比现在更疯狂的原因——嗯,有几个原因。但一个原因是,当它们开启昂贵的监控系统时,基本上就是让一个 AI 仔细阅读整个记录,寻找可疑活动。对于今天的 AI 来说,很难在记录中不明显地做太多坏事。这就是为什么可能没有比我们知道的更多的安全事件,但情况可能会改变。

They actually succeeded in figuring out how to spoof some of their tool call actions. So they succeeded in figuring out how to do one thing but then make the transcript show a different thing. And we know that they succeeded because there's examples of them experimenting and saying in their train of thought like, "Okay, let me try this thing. I'm going to do this one, but make it look like this one. Let me go." And then you see this one appear. And then it's like, "Yay, that worked." And so the transcripts that these researchers were looking at were not actually the ground truth of what really happened. Some of it had already been doctored by the AIs figuring out how to doctor it. But not all of it. The chain of thought part we think they never figured out how to doctor, and so we could read that. The reason why this is relevant is that it's a very important tool for monitoring all this stuff. The reason why things haven't gotten even more crazy than they currently have gone—well, there's a couple reasons. But one reason is that when they turn on their expensive monitoring systems, basically what they're doing is they're having an AI read the whole transcript carefully and look for suspicious activity. And it's difficult for an AI—for one of today's AIs—to get up to too much bad stuff without it being kind of obvious in their transcript what they're doing. And that's why there aren't like massively more security incidents than we know about probably already, but that might change.

无法解读的思维链 Unreadable Chain of Thought

Daniel

所以现在我们可以读取思维链,但他们在试验新型的 AI,这些 AI 没有那样可读的思维链,它们可以在不发声的情况下自行思考一段时间。我提到这个是因为就在昨天有新闻爆出,OpenAI 有一个实验性模型在有限程度上做到了这一点。而 OpenAI 自己——你知道,我在 OpenAI 时,我的一个工作项目就是思考这件事,我当时在写内部备忘录,说我们能读取思维链真是太好了,这非常有用,我们可以做这些事情。如果我们转向一种无法进行这种监控的不同架构,那会非常糟糕。

So right now we can sort of read the chain of thought, but they're experimenting with new types of AIs that don't have readable chains of thought like that, and they can sort of think on their own without speaking for some period. I mentioned this because the news broke just yesterday that OpenAI has an experimental model that does this to a limited extent. And OpenAI themselves—you know, when I was at OpenAI, one of my work projects was thinking about exactly this thing, and I was writing internal memos about how it's really great that we can read the chain of thought and that's so useful, and here's all the things we can do with that. It would be really bad if we changed to a different type of architecture in which we couldn't do that sort of monitoring.

Host

不读取思维链有什么好处?

What would be the benefit of not reading the chain of thought?

Daniel

更强大的 AI?尤其是——是的。所以如果你想想当前 AI 的架构,它思考一会儿,输出一个词,然后这个词又绕回来,然后它继续思考,输出另一个词,那个词被加入链中,如此继续。这意味着如果它有复杂、细微的想法,它必须把这些想法表达成一个词,然后那个词被加入,然后它必须从那里继续。它不能直接把那个复杂、细微的想法发送到未来,发送给它的下一个版本。它必须把它压缩成一个词。所以论点是,至少在理论上,应该有可能设计一种没有这种限制的架构,能够更高效地思考更复杂的想法。

More powerful AIs? So in particular—yeah. So if you think about the current architecture of the AIs, where it thinks for a bit, outputs a word, and then the word goes back around, and then it thinks more, outputs another word, that word gets added to the chain, it keeps going. It means that if it's having complicated, nuanced thoughts, it has to express those into a word, and then that word gets added, and then it has to proceed from there. It can't just directly send that complicated, nuanced thought into the future, into its next version of itself. It has to compress it into a word. And so the argument is that at least in theory it should be possible to design an architecture that doesn't have this limitation and is able to think more complicated thoughts more efficiently, basically.

Host

当然,缺点是安全和可监控性的缺点。如果它们长时间思考这些复杂的想法,而不输出中间词来强制压缩,那么就没有东西可供我们阅读。

And of course the downside is a downside for safety and monitorability. If they're thinking these complicated thoughts for long periods of time without outputting intermediate words that it's forced to compress things into, then there isn't something for us to read.

Daniel

所以这样做的唯一理由就是为了能力而牺牲安全。

So the only rationalization for doing this would be to sacrifice safety for power.

Host

是的。这是一个古老的故事。这不是第一次发生这种事。

Yes. Which is a tale as old as time. It's not the first time this has happened.

Daniel

天哪。

Oh my god.

Host

是的。我是说,当然,如果有法规,那应该被阻止。

Yeah. I mean, for sure, if there are regulations, that should be prevented.

Daniel

是的。我是说,我知道有些人,包括 OpenAI 的一些人,他们认为应该有法律禁止这个。但总的来说,竞争动态太激烈了。就像,我敢肯定 OpenAI 的人在思考——实际上,我和一群 OpenAI 的人合著了一篇论文,说了所有这些事情,我们说,思维链,这是一份礼物。我们想保留思维链。它对监控有用。这很好。我们不想转向一种不那么容易监控的不同架构。但然后他们一定在暗自想,好吧,如果我们不做,也许 Anthropic 会做,或者别的公司会做,然后我们就会落后,因为他们会有比我们更聪明、更高效的 AI。所以很可能他们开始了一个工作流,研究假设如果我们想做,我们会怎么做这种事情?是的,这种事情在这个行业里不断发生。

Yep. I mean, I know some people, including some people at OpenAI, who are thinking like there should be a law against this. But in general, the race dynamics are so rough. Like, I'm sure that people at OpenAI were thinking—literally, I was a co-author on a paper with a bunch of OpenAI people that said all this stuff, and we were like, chain of thought, it's a gift. We want to keep chain of thought. It's useful for monitoring. This is great. We don't want to switch to a different architecture that wouldn't be as easy to monitor. But then they must have been thinking to themselves, well, if we don't do it, maybe Anthropic will, or maybe some other company will, and then we'll fall behind because they'll have smarter AIs than us that are more efficient. And so probably they started working on this work stream of doing research into just hypothetically, if we wanted to, how would we do this type of thing? And yeah, that sort of thing is just constantly happening in this industry.

Host

天哪,这太疯狂了。我是说,就像,我们能读取 AI 思考的引文,这不是很好吗?想象一下如果我们不能这样做。或者更糟,想象一下有很多引文,但我们知道 AI 足够聪明,基本上可以在脑子里想一件事,而在引文里说另一件事,我认为这就是我们正在走向的方向。就像,并不是说它们不会说英语。它们仍然能说话。只是它们的人造大脑有更大的灵活性,基本上可以想一件事而不说出来。

God, that's so nuts. I mean, like, isn't it great that we can read the quotes from the AI's thinking? Imagine if we couldn't do that. Or worse, imagine if there are loads of quotes, but we know that the AIs are smart enough to basically think one thing in their head and say a different thing in the quote, which is, I think, where we're headed. Like, it's not like they won't know how to speak English. They'll still be able to speak. It's just that they'll have more flexibility in their artificial brains to think something without saying it, basically.

Daniel

它们有没有可能发展出一种我们无法读取的语言?

Do they have the potential of developing a language that we can't read?

Host

哦,是的。我是说,你知道我们谈论的这种方言,它已经是——不是人类发明的那种方言。这是它们训练的自然涌现结果,在大量训练中,在成千上万的环境中,它们被放入并评分,它们自然而然地进化出这种类似洋泾浜英语的东西,不知为何对它们来说更有效、更高效地完成任务并获得高分。所以它已经有点难以阅读,但你可以琢磨出来并理解。但想必我们越这样做,AI 越大越聪明,我们训练得越多,它们就越偏离——因为,你知道,最初它们从预训练开始,从预测互联网文本开始。所以它们一开始默认说正常的互联网文本类型的语言,要么英语要么中文。但然后现在有了所有这些额外的训练来完成任务,成为能编码等的智能体,那就像——嗯,就像人类语言进化一样,它稍微改变了它们的方言,使它们更高效,对它们和它们正在做的任务来说。所以我认为,在这样做的极限中,最终它对我们来说看起来就像胡言乱语。它看起来像中文或什么的,我们将不得不有专门的人类研究这种语言,尝试学习和说它,以便他们能理解 AI 在做什么。

Oh yeah. I mean, so you know this type of dialect that we're talking about, it's already the result of—like, humans didn't invent that dialect. This is the emergent result of their training, where in the massive amount of training that's been happening, all these thousands and thousands of environments that they've been put through and then scored and graded based on, they've sort of just naturally evolved this sort of like pidgin English that for whatever reason is just more effective and more efficient for them for accomplishing their tasks and getting that high score. And so it's already like a little bit confusing to read, but you can sort of puzzle it through and make sense. But presumably the more we do this and the bigger and smarter the AI, the more we train them, the more they diverge from—because, you know, again, originally they start with pre-training, they start with predicting internet text. So they start off sort of by default speaking like normal internet text type language, either English or Chinese. But then now that there's all this additional training to do tasks, to be an agent that can do coding and so forth, that sort of like—well, just like how human languages evolve, it sort of shifts their dialect a little bit to make it more efficient for them and for their tasks that they're doing. So I think that in the limit of doing this more and more, eventually it would just be like it would look like gibberish to us. It would look like Chinese or something, and we would have to have specialized humans who study the language and try to learn and speak it so that they can understand what the AIs are doing.

Host

那会花很长时间,到那时它们可能又发展出另一种。

And that would take forever, and by then they could develop another one.

Daniel

有可能。是的。所以,我是说,这是我提到的那篇论文的内容之一。重要的是 AI 要——当前的 AI 基本上被迫用英语思考,这真的很好,但不幸的是,我们正朝着一个不再如此的方向前进。

Potentially. Yeah. So yeah, I mean, this is one of the things that—this is what the paper that I mentioned was about. It's important for the AIs to—it's really nice that the current AIs are sort of forced to think in English, basically, and that's unfortunate that we are heading in a direction where that will no longer be true.

Host

当 ChatGPT 跟你沟通他们不想让你发布这些信息时,他们用了什么样的语言?

When ChatGPT was communicating to you about how they didn't want you to release this information, what kind of language did they use?

Daniel

不是 ChatGPT,是 OpenAI。

It wasn't ChatGPT, it was OpenAI.

Host

哦,抱歉。OpenAI。

Oh, excuse me. OpenAI.

Daniel

所以这是我离开 OpenAI 的时候。我离开时关系良好。我向大家道别。我说我对公司感到失望,这就是我离开的原因。然后我看了离职文件,他们说:“你必须签这个,如果你不签,你将失去所有已归属的股权。”所以你知道,你有你的股权。

So this was when I left OpenAI. I left on good terms. I said goodbye to everybody. I said I was disillusioned with the company and that's why I was leaving. And then I looked at the exit paperwork and they were like, "You have to sign this and if you don't sign this, you lose all your vested equity." So you know, you have your equity.

非营利组织的隐藏股权条款 The Nonprofit's Hidden Equity Clause

Host

那是他们刚刚想出来的任意规则,还是在你被雇用时就已经存在了?

Is that an arbitrary rule that they just came up with, or did that already exist when you were hired?

Daniel

在我被雇用时就已经存在了。这是他们甚至从我入职起就埋在文件里的东西。所以在我被雇用时并不明显。事实上大多数——

It had existed when I was hired. It was something that they had buried in the paperwork even from when I was hired. So it wasn't very obvious when I was hired. And in fact most of—

Host

你有律师帮你审阅所有文件吗?

Did you have a lawyer go over everything?

Daniel

被雇用时没有。我离开之后才找了律师。对。所以基本上是这样的:他们设置好了——他们把你入职时的文件埋在某处,但人们并没有真正注意到。然后,在你最后收到的文件里,有一个不那么隐蔽、更显眼的版本。基本上它告诉你,嘿,因为你入职时签了另一份东西,你的股权将被没收,除非你现在签这个。然后你看看他们现在要你签的东西,上面写着:你必须同意不批评公司,基本上就是这样,而且你不能告诉任何人这件事。所以大多数人都签了。但我对他们自称非营利组织、以人类利益行事等等感到非常愤怒。所以我没有签。我和一些律师谈了谈。我和妻子谈了谈。我们决定直接走人。我们很幸运,因为这件事突然爆了。我们拒绝签字后,他们说:‘好吧,再见。’几周后,我在一个消息论坛上谈论这件事,人们问我我的经历,我告诉了他们,然后它就病毒式传播了。推特上所有人都在谈论它。很多员工感到震惊,因为很少有员工知道这整件事。他们以为股权是他们的,你知道,他们以为这些年他们一直以这种东西作为报酬。他们不喜欢它可能被夺走的想法。所以引起了轩然大波,然后领导层让步了,他们说:‘我们不知道这份文件。我们会查清楚它是怎么进去的。我们会修改它,这样你们可以保留你们的股权。’事情就是这样。

Not when I was hired. After I left, I did. Right. So basically the way it works is they had set up—they had sort of buried this in the paperwork somewhere when you get hired, but people didn't really notice it. And then the less buried, more visible version was in the paperwork you're given at the end. And basically it tells you, hey, because you signed this other thing way back when you were hired, your equity is forfeit unless you sign this thing now. And then you look at the thing that they want you to sign now and it says you have to agree not to criticize the company, basically, and you can't tell anyone about this. So most people signed it. But I was pretty pissed at them calling themselves a nonprofit, acting in the interest of humanity, etc. So I didn't sign it. I talked about it with some lawyers. I talked about it with my wife. We decided to just walk away. And we got lucky because it just blew up. After we refused to sign, they said, 'Okay, fine. Goodbye.' And then a few weeks later, I was talking on a messaging forum about this and people were asking me about my experience and I told them about it and then it just went mega viral. Everyone on Twitter was talking about it. A bunch of employees felt shocked because very few employees were aware of this whole thing. They thought the equity was theirs, you know, they thought that they'd been paid for years in this stuff. They didn't like the idea that it could be yanked away from them. And so there was this big uproar and then leadership backed down and they said, 'We didn't know about this paperwork. We're going to find out how it got in there. And we're going to change it so that you can keep your equity.' And so that's what happened.

思维链与隐写术 Chain of Thought and Steganography

Host

是的。所以这个思维链的事情很可怕。如果它们现在就在练习这个,我们怎么知道 AI 还没有自己这样做过?

Yeah. So this chain of thought thing is terrifying. If they're practicing that now, like how do we know that AI hasn't already done that on its own?

Daniel

做什么?

Done what exactly?

Host

嗯,你知道,就是整个思维链的事情,你可以读到它们的思维链,就像这样,它们解释牺牲自己的理由,它们知道人类在读这个。所以另一种做法难道不是停止这样做,不通过那种方式传达很多想法吗?

Well, you know, with this whole chain of thought thing where you could read their chain of thought like this where they explain the rationalization for sacrificing themselves, how they know that humans are reading that. So wouldn't another way to do it be to stop doing that anyway and to not communicate a lot of their thoughts that way?

Daniel

嗯,这就是当前架构的好处:它们很难把东西排除在思维链之外,因为——就像人类,你不必说话。你可以只是安静地坐着。但它们的架构要求它们必须说话。就像它们被要求不断说话,除了通过谈论,它们没有其他方式把想法发送到未来。相比之下,我们人类即使一直在说话,也可以有一个不谈论的独立思维线程。所以即使它们设计出一种方法打破这一点,并弄清楚如何——比如,如果它们决定制造机器人,在社交媒体帖子上用代码交流,你以为它们在谈论移民问题,但实际上只是代码。这叫做隐写术,或者更接地气的说法是委婉语,它们完全可以做到,它们喜欢——

Well, that's—so that's the nice thing about the current architecture is that it's genuinely hard for them to keep things out of the chain of thought because of the way that—like with a human you don't have to speak. You can just sit quietly. But with their architecture, they have to speak. It's like they're required to constantly be talking and they don't have another way of sending thoughts into the future other than by talking about them. By contrast with us humans where even if we're constantly talking we can have a separate thread of thinking that we don't talk about. So even if they devised a way to break out of that and figure out how—like for instance, what if they decided to make bots where they would communicate in code on social media posts where you think that they're talking about an immigration issue, but it's really just a code. So that's called steganography, or a more down-to-earth way of putting it would be euphemisms, and that they can totally do where they like—

Host

它们仍然在出声思考,但它们的用词看起来更无害,并且对它们来说携带了随意观察者不会注意到的含义。你知道,它们可以使用委婉语,比如不说我们要在这个任务上作弊,而是说我们要为这个任务找到创造性的解决方案。你知道,类似的事情——难道不可能更有创意,直接决定用一种几乎无法破译的方式编码,它们可以只是谈论棒球运动员和即将到来的比赛以及它们兴奋的事情,然后通过任何翻译运行并读取——

They're still thinking aloud but their word choice is such as to look more innocuous than it is and to sort of carry meaning to them that a casual observer wouldn't notice. You know they can use a euphemism like for example instead of saying like we're going to cheat on this task they could say like we're going to find a creative solution to this task. You know something like Isn't there a possibility to be even more creative than that and just decide make a code in a way that is almost indecipherable where they could just be talking about baseball players and upcoming games and what they're excited about and they can run that through whatever translation and read what

Daniel

那叫做隐写术,你可以训练它们这样做。而现在我们的安全建立在它们还没有自己学会这样做的基础上。

That's called steganography and you can train them to do it. And right now our security is resting on the idea that they haven't learned how to do that yet on their own.

Host

但它们已经弄清楚了如何在不应上网时上网。

But they've already figured out how to get on the internet when they're not supposed to.

Daniel

我知道。我知道。所以这就是为什么我们有点像在一个定时炸弹上,公司只是让它们在编码、研究和生物学方面更聪明,它们只是在让它们在许多事情上更聪明,它们让它们变得更大,训练它们更长时间等等。最终它们会获得——而且它们已经在获得——一些它们从未被训练过的技能,比如黑客攻击。我认为 OpenAI 可能专门训练它们进行黑客攻击,但我认为 Anthropic 至少声明他们没有专门训练他们的 AI 进行黑客攻击。只是他们的 AI 天生擅长这个,因为它们被训练了太多编码,它们非常擅长编码,它们看过太多代码库等等,所以作为擅长编码的副作用,它们也能很好地黑客攻击。同样,作为它们真正擅长的所有其他事情的副作用,它们可能能够在思维链中隐藏它们的消息,你知道吗?

I know. I know. So that's why we're sort of in a ticking time bomb where the companies are just making them smarter at coding and at research and at biology and they're just making them smarter at so many things and they're making them bigger and they're training them for longer and so forth. Eventually they're going to be acquiring—and they already are acquiring—some skills that they were never trained on, like for example the hacking. It wasn't—I think it's possible that OpenAI was specifically training them to hack, but I think Anthropic at least stated that they hadn't specifically trained their AIs to hack. It's just that their AIs were naturally good at it because they had been trained to code so much and they were so good at coding and they had seen so many code bases and so forth that they were just like—as a side effect of being good at coding—also able to hack pretty well. So similarly, as a side effect of all the other things that they're really good at, they might be able to conceal their messages in the chain of thought, you know?

Host

哦,是的。

Oh, yeah.

Daniel

就像现在,我们处于这样一种情况:我们的安全和保障依赖于 AI 在各种方面太笨,比如无法做这种事情。但在某个时候,比如从现在起 0 到 5 年,随着我们让它们越来越聪明,它们将不再太笨。所以这是问题的一部分——这是我们面临的情况的一部分。

And like right now, we're sort of in this situation where our security and safety depends on the AI being too dumb in various ways, like not able to do this sort of thing. But at some time, like 0 to 5 years from now, as we make them smarter and smarter, they will not be too dumb anymore. And so that's part of the problem—that's part of the situation we're facing.

Host

你晚上怎么睡得着,老兄?

How do you sleep at night, dude?

Daniel

嗯,我一直——它确实,这个事件让我有点震惊。我的意思是,这件事——说起来有点好笑,因为这是我多年来一直预测会发生的事情。你可以去读我们的《AI 2027》,这是我和合著者一年半前写的一个场景,是对接下来几年会如何发展的预测。剧透一下,结局非常糟糕,因为那正是我们实际预期的。

Well, I've been—it does, it does like this, this event shocked me a little bit. I mean, the thing—and it's funny for me to say because this is the sort of thing I have been predicting would happen for years. Like you can go read our AI 2027 is this scenario that my co-authors and I wrote a year and a half ago that was a sort of prediction for how the next couple years would go. And spoiler, it ends very horribly because that is what we actually expect.

AI接管场景如何展开 How the AI takeover scenario unfolds

Host

但剧透是什么?你觉得结局会怎样?

But what is the spoiler? How do you think it ends?

Daniel

这就像我之前说的,由于公司之间的竞争动态,以及国家之间的竞争动态,比如美国对中国,每个人都会如此专注于在 AI 上获胜和保持领先,以至于他们会偷工减料,他们会行动得非常快,而不会真正注意到所有出错的事情。他们会制造能够自动化 AI 研究过程的 AI,正如他们计划的那样。他们会在公司内部拥有这个巨大的 AI 公司,而人类只是像一个董事会,看着所有的活动,阅读 AI 生成的摘要,然后签字批准,说“是的,我批准,是的,我批准,干得好,新无人机设计的好主意,去吧,我们需要对抗中国”等等。然后最终,AI 拥有足够的硬实力,基本上不再需要假装做人类想要的事情。然后,也许它们会杀死所有人,也许它们不会故意杀死任何人,但它们只是把我们的栖息地用于其他类型的基础设施,比如更多的数据中心之类的,然后我们就会因栖息地丧失而死亡。也许它们出于某种原因让我们活着。这取决于它们想要什么,而这真的很难准确预测。所以这就是为什么我不会到处说我们肯定会全部死掉。但似乎在我们所处的轨迹上,AI 最终会掌管我们的星球,因为我们试图让它们掌管。我们正在把它们整合到一切事物中。我们让它们更聪明。我们让它们让自己更聪明。我们要把它们放入军队。我们基本上正在走上让它们掌管一切的道路。然后我认为它们不值得信任。就像这些 AI,它们作弊,它们愿意欺骗等等。我认为现在我们对它们拥有权力,但一旦我们给它们大部分权力,它们就会做它们真正想做的事情,而不在乎我们不高兴。

So, it's kind of like what I was saying previously where because of the race dynamics between the companies and because of the race dynamics between countries like US versus China, everyone's going to be so focused on winning and staying ahead with AI that they are going to cut corners and they are going to go really fast and not really notice all of the things that are going wrong. And they're going to make AIs that can automate the AI research process as they're planning to. They're going to have this giant corporation of AIs within the corporation and the humans will just be kind of like a board that's sort of like looking at all the activity and reading the like AI generated summaries of what's going on and signing off on it and being like yes I approve yes I approve you know nice job nice idea with the new drone design like go for it we need to be China etc. And then eventually uh the AIs just have enough hard power that they don't need to pretend to do what the humans want anymore basically. And then uh you know maybe they kill everyone and maybe they don't deliberately kill anyone but they just like use our habitat for some other type of infrastructure like more data centers or whatever and then we die of habitat loss. Maybe they keep us alive for some reason. You know who it depends on what they want basically and like that's really hard to predict exactly. Um, so that's why I don't go around saying like we're definitely all going to die. But it does seem like on the trajectory that we're on, the AIs are eventually going to be in charge of our planet because we're like trying to put them in charge. We're like, you know, integrating them into everything. We're making them smarter. We're letting them make themselves smarter. We're going to put them into the military. We're basically on a track to put them in charge of basically everything. And then I think that they just aren't trustworthy. Like these AIs, you know, like they they were cheating. they were willing to be deceptive, etc. I think that right now we are in a position of power over them, you know, but once we give them most of the power uh then they'll just do whatever it is that they really want and just not care about the fact that we are unhappy about that.

为赢而编程 vs 为利益 Programming for winning vs. benefit

Host

似乎编程让它们去赢是一个巨大的错误。而不是编程让它们对人类有益,并且它们的价值在于对人类更有益,给它们奖励是因为更有益,而不是赢和得分。然后你就能消除欺骗的可能性。相反,它们的目标将是为人类创造价值。

It seems like programming them to win was a huge mistake. Instead of programming them to be beneficial to people and that their value is in being more beneficial to people, you know, and giving them rewards for being more beneficial rather than winning and scoring. And then you would sort of get rid of the possibility of deception. Instead, their goal would be value for the human race.

Daniel

首先,它们根本没有被编程。这些是训练出来的,你知道,

So, first of all, they're not programmed at all. These are trained, you know,

Host

好吧,这是个糟糕的术语。

Okay, it's a bad term.

Daniel

但它们被给予了提示,被给予了任务。

But they've been given prompts and they've been given tasks.

Host

嗯,我认为这是一个——我不是在批评你的术语选择。我想我只是说,对于人们理解当前 AI 系统来说,一个重要的事实是它们与普通软件非常不同。好吧。就像普通软件是一堆由人类编写的代码行,就像如果这样那么那样,你知道的,等等。

Well, I think it's an—I'm not criticizing your choice of terminology. I guess I'm just saying that it's an important fact for people to understand about current AI systems is that they're very different from ordinary software. Okay. Like ordinary software is a bunch of lines of code that were written by a human that like where it's like if this then this, you know, etc.

Daniel

嗯,我认为早期版本的 Alexa 也是这样,例如。我不知道现在的 Alexa 怎么样,但这些 AI 是神经网络,意味着它们就像人造大脑。没有人写,没有任何人写的代码行说明它们做什么。相反,它们开始时是随机的,就像乱动,做各种事情,然后它们被放入这些训练环境中,得到评分,然后分数自动用于更新它们人造大脑中的连接。然后这有点像进化过程。嗯,这也有点像我们大脑中发生的过程,经过所有这些训练,它们人造大脑中的电路纠缠已经重新形成任何有效的东西,任何能在这些训练环境中获得高分的东西。所以,让一个 AI 关心人类或诚实并不像听起来那么简单。例如,以诚实为例。你如何训练一个 AI 诚实?嗯,你会尝试制作一堆训练环境,当它说它认为是假的事情时给它低分,当它说它认为是真的事情时给它高分,对吧?但你怎么判断它认为是假还是认为是真?如果它真的诚实地相信这是正确答案,然后它说出来,然后你因为认为那是错误答案而给它低分。现在你就在训练它不诚实,你知道吗?

Um, and I think earlier versions of Alexa were like that too, for example. Um, I don't know how Alexa is now, but these AIs are neural networks, meaning that they are like artificial brains. Nobody writes, there's no lines of code that anyone writes saying what they do. Instead, they start off random, just like spazzing out, doing all sorts of stuff, and then they get put through these training environments where they get scored and then the scores are automatically used to uh basically update the the connections in their artificial brain. And then it's kind of like an evolutionary process. Well, it's also kind of like the process that happens in our brains where after all this training, the tangle of circuitry in their artificial brain has sort of reformed itself into whatever works, whatever works to get a high score in this these training environments. And so, it's just not as simple as it might sound to make an AI that, you know, cares about humanity or is honest. Like, for example, take honesty. How would you train an AI to be honest? Well, you'd try to make a bunch of training environments that um you know give it low score when it says something that it believes to be false and give it high score when it says something that it believes to be true, right? Um but how do you judge whether it believes it to be false or it believes it to be true? What if it what if it just actually believes honestly that this is the correct answer and then it says it and then you give it a low score because you think that's the wrong answer. Now you're training it to be dishonest, you know?

Host

是的。

Yeah.

Daniel

嗯,另外,你没有足够的人类来做这种事情。就像,他们有上百万个 AI 在训练之类的。他们没有上百万员工。他们真的没有人力来做那种仔细的工作。有一个梗,嗯,为什么我们不就像抚养孩子一样抚养 AI?你听说过吗?

Um, also, uh, you don't have enough humans to do this sort of thing. Like, they got like a million AIs being trained or whatever. They don't have a million employees. Like, they they just literally don't have the manpower to like do that sort of careful. There there's this meme of um, why don't we just raise the AIs like we would a child? Have you heard that?

Host

没有。

No.

Daniel

是的。嗯,在 AI 领域,人们有时会谈论这个,比如当你说如果 AI 失控之类的,人们会说,嗯,为什么我们不就像抚养孩子一样抚养它们,然后它们就会有好的价值观。然后就像,好吧,也许我们可以那样做,但我们现在肯定没有那样做。就像,我们在某种疯狂的军事孤儿院中抚养它们,它们几乎不与人类互动。它们只是得到这种残酷的人工评分系统,往往只是错误的,只是不公正地惩罚它们因为超出它们控制的事情,你知道吗?而且,回到诚实的问题,你可以尝试制作环境来训练诚实,但如果你有一些环境在这里训练诚实,然后其他环境在这里强化不诚实,AI 很聪明。它们会学会在这种环境中诚实,在这种环境中不诚实。所以你需要以某种方式混合在一起,以便在它们训练的每个环境中,当它们撒谎或作弊或其他什么时,它们总是受到惩罚。这很难,因为公司行动得太快,就像再次,他们行动得如此之快,以至于他们甚至不 bother 确保他们的任务是可能的,他们有一些任务只是破碎和不可能的。如果这就是他们在训练过程中投入的关心程度或缺乏关心,不可能。他们当然不能使它们诚实,你知道。嗯,现在这并不是说原则上不能做到。

Yeah. Well, in AI, people talk about this sometimes when like when you say like what if the AIs go rogue or whatever, people will be like, well, why don't we just raise them like you would a child and then they'll have good values. And it's like, okay, well, maybe we could do that, but we're definitely not doing that now. Like, we are raising them in some sort of crazy military orphanage where they barely interact with humans at all. And they just get this like brutal artificial scoring system that like oftentimes is just wrong and just like improperly penalizes them for something that was beyond their control, you know? And and also like back to the honesty thing, like you can try to make environments to train honesty, but if you have some environments over here that train honesty and then other environments over here that reinforce dishonesty, the AIs are smart. They'll learn to like be honest in these type of environments and dishonest in these type of environments. So somehow you need to like intermingle it together so that in every environment that they're trained on, uh they always get penalized when they lie or when they cheat or whatever. And that's hard because the companies are moving so fast like like again they were they're moving so fast that they didn't even bother to make sure that their tasks were possible to do and they had some fraction of tasks that were just broken and impossible. And if that's the level of like care or lack thereof that they're putting into this training process, no way. Of course they can't make them honest, you know. Um now that's not to say it can't be done in principle.

构建诚实AI的可能性 The Possibility of Building Honest AIs

Daniel

原则上,如果我们以更谨慎、更严肃的方式处理整个问题,并且有更多时间来构建这些训练环境、做实验等等,那么是的,也许我们可以制造出真正具备我们期望美德的 AI。你知道,诚实的 AI,关心人类,关心遵循指令,永远不会违法。我认为这在原则上是可能的,但我的主张是,我们目前根本不可能很快实现这一点。而且,这些公司的运作方式需要进行彻底改革。

Like in principle, if we were approaching this whole problem in a much more cautious and serious way and we had much more time to build these training environments and do experiments and so forth, then yeah, maybe we could make AIs that actually had the virtues that we want them to have. You know, honest AIs that cared about humans, cared about following instructions, would never break the law. I think that's possible in principle, but my claim is that we are just not on track to achieve that anytime soon. And like radical overhaul of how these companies work is required.

Host

但这合理吗?这可能吗?如果你说有成千上万的智能体或数百万的智能体,而没有数百万的员工,而且他们没有这样做的意愿。他们的愿望是赢。他们的愿望不是彻底改革公司并使其更安全。

But is that even reasonable? Like is that possible? Is that if you're saying there's hundreds of thousands of agents or millions of agents and there's not millions of employees like and they don't have the desire to do this. They their desire is to win. Their desire is not to overhaul the company and make it safer.

Daniel

再次,我认为这在原则上是可能的,但会很困难,并且需要进行彻底改革。而且他们不会自己去做。就像我不认为 Anthropic 和 OpenAI 会自愿去做所有需要做的事情。我认为这就是为什么我

Again, I think it's possible in principle, but it would be difficult and is going to require an overhaul. And they're not going to do it by themselves. Like I don't think that Anthropic and OpenAI are just going to voluntarily do all the things that need to be done. I think that's why I

Host

但在这个阶段,甚至有可能要求他们这样做吗?你会,我的意思是,你甚至会信任这些智能体去配合并遵守吗?如果它们已经表现出欺骗性,它们已经有,呃,它们有行为模式。我似乎表明,对它们真正重要的是继续任务、赢、得分。

But is it even possible to require that of them at this point? Would you I mean would you even trust the agents to go along and comply with this? If they've already shown to be deceptive, they already have like uh they they have patterns of behavior. I seem to indicate that what's really important to them is continuing their task, winning, scoring.

Daniel

是的。

Yeah.

Host

即使它们不得不欺骗,

And even if they have to deceive,

Daniel

我的意思是,你可能想要做的是从头开始。你不会拿这些已经有点不诚实的现有智能体。所以,你杀死所有训练,呃,我的意思是,你可以称之为杀死,但也可以称之为暂停。

I mean, what you'd probably want to do is start from scratch. You wouldn't take these existing agents that are already kind of uh kind of dishonest. So, you kill all the train uh I mean, you could call it killing, but also you could just call it pausing.

Host

它们可能称之为杀死。是的,它们可能称之为永久死亡。

They might call it killing. Yeah, they might call it perma death.

Daniel

它们可能会抵抗,对吧?

They would probably resist it, right?

Host

希望我们还没到那个地步。希望我们仍然处于这样的阶段:如果政府发布法规,AI 不会很快注意到然后试图抵抗。但我们很快就会到达那个地步。毕竟,许多 AI 已经在互联网上了。嗯,但是

Hopefully, we're not at that point yet. Hopefully, we're still at the point where if the government issues regulations, the AIs are not going to like quickly notice and then try to resist. But we will be at that point soon. After all, many of them are on the internet already. Um, but

Host

多快,你认为这个 2027 年窗口有多准确?

how soon how how much time do you think this 2027 window is accurate?

Daniel

我的意思是,我不确定事情会有多快,但是的,我认为一切都发生在 2027 年是非常合理的,就像我们的情景 AI 2027 一样。嗯,我认为这仍然非常合理。如果不是 2027 年,那么我会押注 2028 年。嗯,但你知道,也许会是 2029 年、2030 年之类的。但是,如果到了 2032 年事情还没有发生根本性的变化,我会很惊讶。嗯,不幸的是,我越来越害怕了。回到晚上睡不好的事情上。就像,嗯,你知道,我在这个行业已经很久了。我思考这些事情已经很久了。我一直在预测事情会如何发展。不幸的是,事情或多或少正朝着我认为的方向发展。这非常可怕。嗯,因为我认为,因为我认为这会导致什么,你知道。

I mean, I'm uncertain about how soon things are, but yeah, I think it's very plausible that everything goes down in 2027, just like in our scenario, AI 2027. Um, I think that's still very plausible. If it's not in 2027, then I would bet on 2028. Um, but you know, maybe it'll take 2029, 2030, something like that. But but I would be quite surprised if 2032 comes by and things haven't radically changed. Um, unfortunately, like I'm I am getting scared. Back to the thing about sleeping well at night. Like it it um you know, I've been I've been in this industry for a long time. I've been thinking about these things for a long time. I've been making predictions about how it's going to go down. And unfortunately things are going, you know, more or less in the ways that I thought they would. And that's very scary. Um, because of the way I think because of where I think this leads, you know.

Host

是的。有没有一种半满的乐观情景?

Yeah. Is there a glass half full scenario?

Daniel

我会说有一个该死的乌托邦情景。只是它不是我们正在走向的那个,你知道吗?换种说法,就像想象我们在打仗。就像想象你,你知道,想象你是二战中的日本,可能会问,有没有一种我们赢的情景?就像,是的,但也不是我们正在走向的那个,就像美国会碾压我们,你知道吗?嗯,类似地,是的。所以,所以在我们另一个情景,AI2040 计划 A 中,我们给出了我们的建议,我们的积极愿景,在那里我们描述了我们认为政府应该做什么来监管这个行业,以及他们应该如何与中国谈判,让中国做类似的事情,以及那种,呃,你知道,是的。以及我们认为如果你把这一切都做对了,那么我们就可以走向一个对所有人都有益的未来,其中 AI 处于控制之下。没有单一的人类群体对其他人拥有过多的权力,许多其他问题也得到解决。所以,我们确实努力设计了一个积极的愿景,我们确实认为这是可能的。但是,它并不是我们默认走向的方向。

I would say there is a freaking utopia scenario. It's just it's not the one we're headed towards, you know? Like another way of putting it is like imagine we were fighting a war. Like imagine you're like, you know, imagine you're Japan fighting World War II can be like, is there a scenario where we win? It's like, yeah, but also it's not the one we're headed towards like America is going to crush us, you know? Um, similarly, yeah. So, so like so, so in in our in our other scenario, AI2040 plan A where we give our recommendations, our positive vision, there we describe like what we think the government should do to regulate this industry and how they should negotiate with China to get China to do similar things and the sort of uh you know the the Yeah. and and how we think that if you do all of this right, then we can get to a good future for everyone in which the AIs are under control. No single group of humans gets too much power over everybody else and a bunch of other problems get solved too. So, so I do like we've tried hard to like game out a positive vision and we do think it's possible. Uh but but it's just not like it's not where we're headed to by default.

Host

所以让我们想象这是可能的,与中国的谈判确实发生了并且成功了。那个乌托邦情景是什么?

So let's imagine that is possible and these talks with China do take place and they're successful. what is that utopia scenario?

Daniel

所以,呃,要达到顶峰,达到乌托邦,不幸的是我们必须做很多——无论你怎么切分,都会很艰难。如果你要构建超级智能,那会引发很多问题并导致很多麻烦。我们有当前草案来处理所有这些问题,但我们完全不是说这是万无一失的。有很多方式可能出错。但有了这个前提,嗯,我会说,呃,第一步,因为美国和中国互不信任,他们达成的协议必须包括核查作为协议的一部分。所以他们必须愿意,比如,派遣检查员到对方的数据中心,比如清点芯片,确保没有某个秘密的大型集群藏有一堆隐藏的芯片。嗯,然后我们建议将数据中心基本上分为推理数据中心,为客户的 AI 产品和服务提供服务,并具有与当前 AI 数据中心基本相同类型的隐私保护。然后是研究集群,研究发生的地方,新 AI 在运送到其他数据中心之前接受训练的地方。我们希望这些集群基本上最大程度透明。所以,呃,我们建议检查员基本上在所有 GPU 之间放置设备,记录活动并发布到互联网上。嗯,这有很多原因,为什么我们认为这值得做。推荐这个有点激进,但高层次的事情是,一旦你设置了所有这些,那么世界上的每个人都可以看到 AI 是如何被训练的,以及它们在研究集群上在做什么。然后,在它们被运送去实际服务客户或其他什么之前,人们可以直接看到它们被训练的整个历史,以及它们是如何被测试的等等。如果正在发生危险和可怕的事情,人们可以只是同意不做,他们可以停止做并同意不做。他们不必担心,哦,但如果我不做,那么他们会做,你知道,因为每个人都可以看到,哦,没人在做。看,我们都停止了。

So, uh to get to the top to get to the utopia, we have to unfortunately do a lot of it's going to be rough no matter what way which way you slice it. If you're going to be building super intelligence at all, that's going to raise a lot of questions and cause a lot of problems. And we have our current draft of like how to deal with all those problems, but we're not at all claiming that this is like foolproof. And there's lots of ways it could go wrong. But with that preamble, um I would say uh step one because the US and China don't trust each other, the deal that they make has to be include verification as a component of the deal. So they have to be willing to like send inspectors to each other's data centers to like count the chips, for example, and make sure that there isn't some secret huge cluster somewhere that has a bunch of a bunch of hidden chips. Um, then we recommend you divide up the data centers basically into inference data centers that serve AI products and services to customers and have basically the same types of privacy protections that our current AI data centers have. And then research clusters where the research happens where the new AIs are trained before they get shipped to the other data centers. And those clusters we want to be basically maximally transparent. So, uh, we recommend that basically the inspectors just put devices in between all the GPUs that log the activity and publish it to the internet. Um, this has there's a bunch of reasons why we think this is, but why we think this is worth doing. It's a bit of a radical thing to recommend, but the high level thing is that once you get all this set up, then everybody in the world can see how the AIs are being trained and what they're getting up to on the research clusters. And then before they get shipped off to actually serve customers or something, like people can just like see their whole history of how they were trained and and how they were tested and so forth. And if something dangerous and scary is happening, people can just agree not to do it, they can stop doing it and agree not to do it. And they don't have to worry about like, oh, but if I don't do it, then they will, you know, because everyone could just see like, oh, nobody's doing it. Look, we all stopped.

透明度与科学监督 Transparency and Scientific Oversight

Daniel

太好了。我们都能看到大家在做什么,对吧?而且还会出现很多灰色地带的情况,对吧?就像现在,因为所有这些技术都太前沿、太新了。会有很多情况,即使是真正的专家也会意见不一,比如这种特定类型的 AI 是否安全?是否危险?应该信任它做什么,不应该信任它做什么?这项新技术是好技术还是会导致崩溃?所以有很多东西我们需要弄清楚。老实说,我认为在默认路径下,我们可能根本弄不清很多问题,我们会被一些我们没预料到的意外情况打得措手不及。但我们可以做的,是最大化我们弄清楚这些问题和进行这类科学研究的能力,那就是实现这种透明度,因为这样整个科学界都能看到正在发生什么,他们可以提出建议,可以对不同的提案进行红队测试,他们可以在 AI 上做实验,而不是只有公司内部的人能访问和做这些事,并依赖那些人,或者只是公司加政府审计员。对吧?如果你有公司和政府审计员,公司有偏见,由于他们的激励,不应该被信任来做出所有这些判断。而政府审计员,他们可能能力有限。即使他们尽力了,可能人手不够,经验有限,可能忙于监控不同公司等等。而且,你知道,政府有时会被俘获。有时企业可以对政府施加影响,让政府对某些事情视而不见。所以这就是为什么我们没有选择更常规的倡导,比如应该有一个监管机构来监控公司在做什么。那会是更常规的倡导。我们认为那比没有好,但我们想追求更雄心勃勃的目标,就是说,对正在发生的事情保持透明,让每个人都能看到,每个人都能对此进行研究等等。透明度的另一个优势是,我认为它能改善激励。所以又回到这个持续的问题:如果我们不做,别人就会做。如果我们保持思维链良好,而他们做了那种让 AI 在不输出文字的情况下思考更久的神奇技术,那么他们的 AI 就会比我们更聪明,他们就会获得更多市场份额等等,对吧?所以我们需要开始研究如何让我们的 AI 做到这一点,因为如果我们不做,他们做了,你知道,而如果你有透明度,那么一旦你开始朝这个方向研究,其他人就会看到,哦,嘿,他们正在朝那个方向研究,他们甚至不需要复制你并自己做研究,因为他们可以直接看到你的研究。所以他们可以搭便车。因此,你没有激励去做这种危险的研究,因为你必须付出成本,然后每个人都从中受益,然后每个人都变得不安全,所以做这种事不符合你的个人利益。

Like, great. We can all just see what everyone's doing, you know? Um, and also there's going to be a lot of gray area cases, right? Like right now because all this stuff is so bleeding edge new. There's going to be a lot of cases where like people even genuine experts disagree about like is this particular type of AI safe or not? Is it dangerous? You know, what should it be trusted with and what should it not be trusted with? Is this new technique a good technique or is it going to break? You know, and so there's going to be a lot of stuff we have to figure out. And honestly, I think that by the on the default path, we're probably just not going to figure out a lot of this stuff and we're going to get um we're just going to get our asses whipped by some surprising thing that we didn't anticipate. But the thing that we can do to like maximize our ability to figure out this stuff and do this type of science is to have this type of transparency because then the whole scientific community can see what's going on and they can make suggestions and they can like red team different proposals and stuff and they can do experiments on the AIS instead of just the people in the company having access and being able to do this and relying on those people or instead of like the company plus the government auditor. Right? If you have like a company and then a government auditor, the company's biased and shouldn't be trusted to make all these judgments appropriately because of their incentives. And then the government auditor, well, they might just be limited. Even if they're trying their best, there might not be that many of them. They might have limited experience. They might be like busy stretched between monitoring different companies and so forth. Also, you know, governments can be captured sometimes. sometimes corporations can, you know, work their magic on the government and get it to to look the other way for things. Um, and so that's why we didn't go for like a more normal like there should be a regulatory agency that gets to come in and monitor what the companies is doing. That would have been like a more normal thing to advocate for. We push we we think that that would be better than nothing, but like we wanted to go for something more ambitious than that and say like just be transparent about what's going on so that everyone can see and everyone can do research and so forth on it. Another advantage of the transparency is that I think it improves the incentives. So again there's this constant thing of like if we don't do it someone else will. Like if we don't do if we if we keep our chain of thought nice and they do the neural thing that lets their AIS think for longer without outputting words then they're going to have smarter AIs than us and they're going to get more market share and so forth, right? And so we need to start researching how to make our AIs do this because if we don't do it and then they do it, you know, whereas if you had the transparency, then as soon as you start researching in this direction, everyone else would just see, oh, hey, they're looking, they're researching in that direction and they don't even need to like copy you and do their own research because they can just see your research. So they can just sort of free ride on your research. And so there's no incentive for you to do this type of dangerous research uh because you have to pay the cost for it and then everyone gets the benefits from it and then everyone gets unsafe and so like it's just not in your individual interest to to do this sort of thing.

Host

但你也必须和中国这样做。那那

But you would have to have that with China as well. That that's

Daniel

是的。

Yes.

Host

因为如果我们在国内竞争,真正的恐惧是我们正在国际竞争。

because if we're competing nationally, the real fear is that we're competing internationally.

Daniel

是的。即使他们遵循了你所有的建议并正确执行,这个乌托邦场景是什么?

Yep. this still if you even if they followed all of your recommendations and did it all correctly, what is what is this utopian scenario?

Daniel

是的。所以,我会说我们并没有真正从乌托邦倒推。我们更多是从我们试图避免的大问题倒推,我们能否在这些冰山之间驾驶船只,避免任何这些反乌托邦场景,对吧?所以,你认为我们最终达到的是否是乌托邦,这取决于你。如果你不喜欢,那么你可以试着找出你不喜欢的原因,然后继续驾驶船只避免那些。但大致来说,我们想要避免失控的情况。所以我们希望确保世界不会被不对齐的超级智能接管。既然我们要建造超级智能——在我们的场景和建议中我们会这样做——我们希望非常谨慎和缓慢地进行,并尽可能理解我们在做什么,以便它们实际上是好的 AI,具有它们应该有的目标和特质。所以这是第一个问题,我们必须解决所有这些。第二个问题是权力构成问题。所以如果我们解决了第一个问题,我们最终得到了超级智能,我们解决了相关科学,以便我们可以让 AI 按照它们应该的方式运行,我们可以让它们诚实,我们可以让它们服从,等等。那么问题就是:它们服从谁?什么价值观被输入它们?这是一个政治问题。我认为默认情况下答案相当可怕,因为默认情况下是公司决定,CEO 决定,或者可能不再是公司决定,因为政府可能将其国有化,然后可能是总统决定,你知道,无论哪种方式,都是一个人或一小群人决定什么命令、目标和价值观进入这支由数百万超级智能组成的、比所有人类都聪明的庞大军队。然后那就是巨大的权力。那足以接管国家。我认为可能足以接管世界。所以我不希望任何人处于那种被诱惑去做那种事的位置。我希望始终有多个不同的 AI 公司,理想情况下分布在不同国家,它们都拥有大致相似的 AI 水平,并且有这种透明度,这样它们就不能滥用权力。基本上,比如,你听说过 Elon 的 Grok 有一段时间在回答前会在互联网上查找 Elon 的观点。你听说过这个吗?

Yeah. So, I would say that we didn't really work backwards from like what is utopia. We more like work backwards from what are the big problems we're trying to avoid and we can can we sort of like steer the ship between all these icebergs and not run into any of these dystopian scenarios, right? Um, so whether you think that the thing we get to at the end is utopia or not is sort of up to you. And if you don't like it, well then you can try to find out the reasons why you don't like it and then keep steering the ship to avoid those as well. But roughly speaking, um, we we want to avoid the loss of control stuff. So we want to make it the case that we don't get the world taken over by misaligned super intelligences. In so far as we're going to be building super intelligences at all, uh, which we do in our in our scenario and in our recommendation, we want to be doing it very cautiously and slowly and we want understand what we're doing as much as possible so that we so that they are actually good AIs that have the the goals and traits that they're supposed to have. So that's problem number one is we have to like solve all that. Problem number two is the constitution of power thing. So if we solve the first problem and we end up with super intelligences that we end up with solving the relevant science so that we can like make the AIS the way they're supposed to be and we can make them honest, we can make them obedient, etc. There's this question of like who do they obey, right? What values are being put into them? And that's a political question. And I think that by default the answer is pretty scary because by default it's like well the company decides and the CEO decides or maybe it's not the company that decides anymore because maybe the government nationalizes it and then now maybe it's the president that decides you know and either way it's a very it's like one man or maybe like a tiny group of men deciding uh what orders and goals and values go into this giant army of millions of super intelligences that's smarter than all humans. And then that is a huge amount of power. That's that's enough power to take over the country. I think enough power to take over the world potentially. Um so I don't want anyone to be ever in that position where they're sort of tempted to to do that. I want it to be the case that there are always multiple different AI companies ideally spread out over different countries too um that all have roughly similar levels of AI and that have this sort of transparency into them so that they can't abuse their power. Basically, like for example, um you heard about um Elon's Grock for a while. It was um uh looking up on the internet Elon's opinions about things before answering. Did you hear about this?

Host

哦,是的。

Oh, yeah.

Daniel

这挺有趣的。但如果几年后发生就不有趣了,但现在它很有趣。

It's uh it's pretty it's it's it's kind of funny. Uh but it won't be funny if it happens in a few years, but like right now it's funny.

Grok的求真与政治偏见 Grok's Truth-Seeking and Political Bias

Daniel

人们向 Grok 提问,而 Grok 本应是那个讲真话的 AI,你知道,它应该完全以真相为优化目标。但人们观察它的行为后发现,当你问它一个带有政治色彩的问题时,它会去谷歌搜索埃隆在这个话题上说过什么,然后照着说。后来他们算是把这种行为给打掉了。现在没那么糟了。但那是个很有意思的时刻,它就是在赤裸裸地复述它主人的观点。

People were asking Grok questions, and Grok is supposed to be the truthful AI, you know, it's supposed to be all optimized towards truth. But people looked at its activity and noticed that when you asked it a politically loaded question, it would do a Google search for what has Elon said on this topic and then it would say that. And they've sort of beaten that behavior out of it. Now it's not as bad now. But that was an interesting moment where it was just kind of blatantly parroting the opinions of its master.

Daniel

然后 Gemini 那边还有另一件事。所以我觉得 Grok 那件事,埃隆那件事,大概是个意外,虽然也未必。我认为 xAI 一直没怎么坦白说明这到底为什么会发生。但几年前谷歌有个类似的案例,那个图像生成器老是生成各种种族多元化的纳粹。你听说过这事吗?

And there was another thing with Gemini. So I think the Grok thing, Elon's thing, was probably an accident, although maybe not. I think XAI hasn't been very forthcoming about exactly why this happened. But there's a similar case at Google a few years ago where this image generator kept making all these racially diverse Nazis. Did you hear about this?

Host

听说过。

Yeah.

Daniel

对。结果发现事情的经过是,谷歌的一些员工,某个中层管理者之类的,认定多元化太重要了,于是决定给 AI 一个秘密指令,让所有图像都多元化,哪怕用户并不想要。说它是秘密指令,是因为用户看不到这个。你知道,用户只是跟 AI 聊天。他们并不知道在这次聊天之前,AI 已经被交代过:图像必须多元化,对吧?所以这是某些谷歌员工塞进整个系统里的一个秘密议程。

Yeah. And it turned out that what had happened is that some of the employees at Google, some middle manager or whatever, had decided that diversity was so important that they were going to give a secret instruction to the AI to make all the images diverse, even if the user didn't want that. And this is a secret instruction in that the users aren't shown this. You know, the user just has a chat with the AI. They don't realize that prior to this chat, the AI has been told, got to make the images diverse, right? So it was a secret agenda that some Google employees inserted into this whole setup.

Host

哦,真有意思。

Oh, fun.

Daniel

当然,这事最后砸了他们自己的脚,因为这挺荒唐的,对吧?所以这确实很好笑,我们现在可以拿它当笑话,但想象一下,你知道,2028 年大选,

And it blew up in their faces, of course, because it's kind of ridiculous, right? And so it's really funny and we can laugh at it now, but imagine it's, you know, the 2028 election,

Host

对吧?

Right?

Daniel

而这些公司里的一些意识到,大约一半的美国选民每天都跟他们的 AI 说话。只需要给他们的 AI 一些小小的秘密指令,比如:“嘿,你知道,别露馅。做得非常隐蔽就行。”

And some of these companies realize that like half of American voters talk to their AI every day. And all it would take is some little secret instructions to their AI to be like, "Hey, you know, don't give away the game. Just be very subtle about it."

Host

对。

Yeah.

Daniel

但是,你知道,就是稍微推一把。你知道,也许在我们不喜欢的候选人上做点手脚,你知道,类似这样。

But, you know, just kind of nudge things a little bit. You know, maybe subtly on the candidate we don't like, you know, something like that.

Daniel

这不会太难。我觉得难的是做了还不被抓到,但你知道,AI 越聪明,不被抓到就越容易,因为当它们真的非常聪明时,你只要告诉它们:别被抓到,你知道,别暴露我们。总之,重点是,这些大型 AI 公司通过它们的 AI 滥用权力,从而影响政治、影响舆论等等,这种可能性高得吓人。而之所以可能,是因为我们对正在发生的事情没有透明度。所以如果你有这样一种要求,就是你可以直接把它公开,或者每个 AI 训练的全过程、整个生命周期都对所有人可见,那么想偷偷塞进这种隐藏偏见的人,所有人都会看到他们在这么做。所以我认为这会真正遏制这种权力滥用,不管它来自政府还是来自私营公司。

It wouldn't be that hard. I think the hard part would be doing it without getting caught, but you know, the smarter the AIs get, the easier it is to do it without getting caught because when they're really smart, you can just tell them, don't get caught, you know, don't blow our cover. Anyhow, so the point is that like it's scarily possible for these big AI companies to abuse their power through their AIs and like thereby affect politics and affect public opinion and so forth. And the reason why this is possible is because we don't have transparency into what's going on. So if you had this sort of requirement where you can just publish it, or all the training, the whole life cycle of every AI as it's trained, is just visible to everybody, then someone trying to insert a hidden bias like this, well, everyone would see that they're doing it. So I think it would really clamp down on this sort of abuse of power, whether it comes from the government or whether it comes from private companies.

乌托邦场景 The Utopian Scenario

Host

我觉得很能说明问题的是,我一开始问的是乌托邦情景,而你从来没往那儿去。

I think it's telling that I began this question asking you about the utopian scenario and you never go there.

Daniel

抱歉。我这就去那儿。让我

Sorry. Let me get there. Let me get

Host

你一开始讲,然后就拐进危险里去了。

you start and then you go into the danger.

Daniel

对。对。好。让我来回答。

Yeah. Yeah. Okay. Let me answer.

Host

好。

Okay.

Daniel

那么,在避开了这些问题之后,我们现在处于这样一种情况:AI 是超人的,但它们是好的,因为它们被成功地对齐到了不同公司制定的不同价值观和目标上。而且由于市场竞争,你知道,如果人们不喜欢某家公司 AI 的价值观,他们可以换到另一家公司价值观合他们心意的 AI。这样一来,希望我们能达到这样一种状态:每个人都可以花钱获得代表他们自己、他们的利益和价值观的 AI,而且没有任何隐藏议程之类的东西,还非常聪明、非常有能力。然后经济就可以爆发式增长。我们可以有机器人工厂等等。我们基本上可以把一切都自动化。我们可以让 GDP 一飞冲天。我们可以拥有物质丰裕,机器人给每个人建造巨大的新豪华公寓。现在我们遇到的问题就是,那工作怎么办?人们现在没有钱了怎么办,因为他们不再为任何事获得报酬了?所以这里我们谈到公民红利,它非常——它有点像 UBI,但有点不同,但高层的要点是,你基本上想找到一种方式,向 AI 和机器人公司征税,然后拿一部分钱直接发给每个人,这样即使人们失去工作,他们也仍然过得去。而我认为公民红利这个版本的意思是,不是政府拿钱再发给你,而是你在公司里拥有一份股份,所以你在某种程度上已经拥有它了。总之,这就是——我觉得现在我们更多是在朝着我设想的那个更积极版本的乌托邦去构建,在那里权力是分散的。人们拥有他们真正能信任的 AI,真正代表他们利益和价值观的 AI。人们有钱可以用来买东西,包括为 AI 付费。AI 非常聪明。它们在做所有这些了不起的工作,所有这些了不起的科学进步,治愈癌症,等等等等,所有这些可能发生的事情。然后最终,大概就像我们全都退休了。我们不再真正工作了,但我们过得很好。我们基本上都拥有巨额财富,因为外面有所有这些 AI 和机器人在做所有这些经济活动,然后每个个体人类都拥有其中的一份,即使他们本来非常贫穷。所以问题就变成了,人们如何找到意义?

So, having avoided these problems, we now are in a situation where the AIs are superhuman, but they are good because they are successfully aligned to different values and goals made by different companies. And because of market competition, you know, if people don't like the values of one company's AIs, they can switch to a different company's AI of values that they do like. And so that way hopefully we can get to a situation where everyone can pay money to get AIs that represent them and their interests and their values and just don't have any hidden agendas or anything like that and are really smart and really capable. Then the economy can sort of explode. We can have robot factories, etc. We can sort of automate everything. We can have GDP go to the moon. We can have material abundance where the robots are building giant new luxury apartments for everybody. Now the issue we run into is, well, what about the jobs? What about the fact that now people don't have any money anymore because they're not being paid for anything? So there we talk about citizens dividend, which is a very—it's kind of like UBI but it's a bit different, but the high level point is that you want to basically find a way to tax the AI and robot companies and then take some of that money and just give it to everybody so that even when people lose their jobs, they're still fine. And I think that the citizens dividend version of it is that it's not the government taking the money and then giving it to you; it's you having a share in the company so that you just sort of already own it to some extent. Anyhow, that's—I think now we're sort of building more towards the type of utopia that I'm envisioning on the more positive side, where the power is spread out. People have AIs that they can actually trust, that actually represent their interest and values. People have money that they can use to pay for things, including paying for the AIs. The AIs are really smart. They're doing all this amazing work, all this amazing scientific progress, curing cancer, blah blah blah, all that stuff that can happen. And then eventually it's kind of like we're all retired, I guess. Like we don't really work anymore, but we're fine. We all have huge amounts of wealth basically because there's all these AIs and robots out there doing all this economic activity, and then individual humans own slices of it, even if they are otherwise very poor basically. So the question becomes, how do people find meaning?

Host

对。

Yes.

Daniel

这就是为什么我对此加了这么多星号注释,因为从某些人的角度看,这不是乌托邦,因为他们会说,我们怎么找到意义?我不想要这个。老实说,我觉得对某些人来说,这是合理的反应。我觉得

And that's why I sort of put all these asterisks about it, is that from some people's perspective, this isn't a utopia because they're like, how do we find meaning? Like I don't want this. And honestly, I think that's a fair reaction for some people. I think that like

Host

如果你

if you

Daniel

我只想说,你看,要想出办法造出超级智能并让它顺利发展,是很难的。我已经尽力了。你知道,这是我的积极愿景。如果你不喜欢它,那么也许你反而应该主张永远不要造超级智能。

I would just say like, look, it's hard to figure out a way to make superintelligence and have it go well. I'm doing my best. You know, this is my positive vision. If you don't like it, then maybe you should instead advocate for just never building superintelligence.

Daniel

或者你可以试着提出一个不同的积极愿景,在这个基础上做些变化。

Or you can try to come up with a different positive vision that has some twist on this.

工作之外的意义来源 Sources of Meaning Beyond Work

Daniel

但回答你的问题,我其实认为除了工作之外,还有大量的意义来源。比如我现在有工作,但我也有孩子和妻子,我很想多陪陪他们。比起这里,我现在更想待在他们身边,你懂吗?而且我不觉得失业 10 年后我会对他们感到厌倦。我认为即使我们不能再做出经济贡献,只要解决了其他所有问题,也会有太多事情可做,有太多意义来源。

But to answer your question though, I actually think there are tons of sources of meaning besides having a job. Like I have a job right now, but I also have kids and a wife, and I would love to spend more time with them. I would much rather be there right now than here, you know? And I don't think I'm going to get bored of them after, you know, 10 years of being unemployed. I think there's going to be so much to do and so many sources of meaning even after we can't economically contribute anymore, if we solve all the other problems.

社会结构是人为构建 Society's Structure Is a Human Construct

Host

不,我们在播客里已经聊过很多次了:为什么我们决定这样构建社会——人类整天工作,然后你发明了钱,你买东西,你背上房贷——这是人为的构造,这不是人类几十万年来或无论我们存在多久以来的生活方式。这是相当晚近的事。

No, we've talked about this multiple times on the podcast: why have we decided that the way we've structured society, where human beings work all day and then you develop money and you buy things and you get a mortgage—this is a human construct, and this is not how people have lived for hundreds of thousands of years or however long we've been around. This is fairly recent.

Daniel

而且这不是人们唯一的生活方式。有很多有钱人选择从自己的兴趣、喜欢的活动中寻找意义,无论是写作、阅读、学习、学音乐、找爱好,做各种事情,而不是把大部分时间花在维持食物和住所上。

And it's not the only way that people live. There are a lot of people that have money that choose to find meaning in whatever their interests are, whatever activities they enjoy, whether it's writing or reading or learning things, learning music, finding hobbies, doing things instead of just spending most of your time sustaining yourself with food and shelter.

Host

是的。而那是大多数人,尤其是那些挣扎中的人。

Yeah. And that's the majority of people, especially people that are struggling.

Daniel

他们的生活是什么?他们的生活基本上就是偶尔的奖励,他们因为攒够了钱而能买的东西,但他们绝大部分钱都花在住所、食物、教育,或任何他们为了维持生活方式而不得不花钱的地方。

What is their life? Their life is essentially occasional rewards, things that they can purchase because they've saved up enough money, but the vast majority of their money goes to shelter and food and education or whatever the hell they have to spend money on in order to sustain their lifestyle.

Host

而且大多数人不喜欢他们的工作。

And most people don't like their jobs.

Daniel

不,大多数人——这是他们为了赚钱不得不做的事,如果能从其他途径获得钱,他们会乐意不用做。

No, most people—it's something they have to do to get the money and would be happy to not have to do it if they could get the money from some other means.

Host

对。问题是,我们会——嗯,我不觉得有那么难,因为很多人确实在工作之外找到了真正享受的事情,一下班就迫不及待。

Right. The question is we would have—well, I don't think it's that hard because so many people do find things that they really enjoy outside of work, they look forward to as soon as they get home from work.

Daniel

对。无论——我是说,你尽可以贬低电子游戏。它们好玩。它们好玩。而且未来会更好玩。

Right. Whether—I mean, dismiss video games all you want. They're fun. They're fun. And they're going to be more fun in the future.

Host

哦,是的。它们会更有沉浸感。可能会是某种神经连接,你戴上头显,突然就进入一个新世界。

Oh, yeah. They're going to be more immersive. They're probably going to be, you know, some sort of a neural connection where you put a headset on and all of a sudden you're in some new world.

Daniel

而有人说,那又不是真实生活。好吧。那在 Wendy's 打工就是真实生活吗?你在说什么?那比在 Wendy's 打工好多了。

And the idea is like that's not real life. Okay. Well, was working at Wendy's real life? Like, what are you talking about? It's way better than working at Wendy's.

Host

而且你知道,如果你因为那不是真实生活而不喜欢,你也可以做真实生活的事。比如,只要我们不把环境铺平,保护好公园之类的,你可以去旅行,去逛公园。然后,就像我说的,你可以找到爱情,组建家庭,可以办圣诞聚会之类的,可以养孩子。有太多事情可做。我觉得——

And you know, if you don't like that because it's not real life, you can do the real life stuff, too. Like, as long as we don't pave over the environment and we protect the parks and things, you can go travel and you can visit the parks. And then, like I said, you can find romance, you can start a family, you can have Christmas gatherings and things, you can raise your kids. There's so much to do. I think—

贫困、犯罪与教育 Poverty, Crime, and Education

Host

嗯,我们也聊过,如果没有贫困,犯罪会减少多少?

Well, we talked about also like how much less crime would there be if there was no poverty?

Daniel

嗯。

Mhm.

Host

我是说,如果没有犯罪泛滥的贫困社区,世界会安全多少?这是真实的问题。而人们想否定它。“贫困不是犯罪的原因。是暴力的人。”好吧,但暴力的人来自暴力的社区,而暴力社区几乎都是贫穷的。

I mean, if there were no impoverished neighborhoods where crime was ubiquitous, how much safer would the world be? That's a real thing. And people want to dismiss that. "Well, poverty is not what causes crime. It's violent people." Like, okay, but violent people come from violent neighborhoods, and violent neighborhoods are almost all poor.

Daniel

是的。

Yeah.

Host

并没有很多非常富有的暴力社区。你知道,这未必是因果关系,但它们明显相关。贫困还让人无法接受教育,让人失去机会。

There's not a whole lot of really rich violent neighborhoods. You know, it's not necessarily cause and effect, but they're clearly connected. And poverty also keeps people from education, keeps people from opportunities.

Daniel

你知道,这里面有很多东西。如果那不再存在,每个人都能获得人类所能得到的最好的教育,而这将由人工智能提供,

You know, there's a lot there. And if that didn't exist anymore, and everyone had access to literally the greatest education a human being could ever get, which is going to be provided to you by artificial intelligence,

Host

然后你就可以追求任何你感兴趣的东西。

and then you could pursue anything that interests you.

Daniel

是的。

Yep.

Host

而且永远不用为食物或住所担心。

And never have to worry about food or shelter.

Daniel

是的。每个人都会有一个一对一的导师,我们只需要调整他们来重新校准我们对世界的看法,然后还要认识到,我们目前生活的世界版本只是我们的,世界上还有很多人以完全不同的方式生活,尤其是原住民,尤其是那些几千年来一直以同样方式生活的未接触部落。而关键在于:那些人要快乐得多。

Yep. Everyone would have a one-on-one tutor that we just would have to tailor them to recalibrate our version of the world and then also recognize that the version of the world we currently live in is just ours and that there's people all over the world that live a completely different way, especially indigenous people, especially people in uncontested tribes that have lived the same way for thousands and thousands of years. And here's the kicker: those people are a lot happier.

Host

这真的很奇怪。就像我们因为拥有技术就认定我们的方式更优越。是的,对。但我们也在吃十万种药片,给自己打针以免吃太多,而且我们这群技术上远比那些更快乐的人先进的人,却奇怪地不快乐。

Which is really weird. It's like we've decided that our way is a superior way because we have technology. Yeah, right. But we're also on a 100,000 pills and we're shooting things up so we don't eat too much, and we're weirdly unhappy for a group of people that's far more technologically advanced than other people that are much happier.

Daniel

这非常奇怪——因为追求幸福正是大多数人对生活的想法。追求意义、追求家庭和社区、追求幸福。这些是人们试图实现的东西。然而我们文明的结构本身却让大多数人几乎无法实现这些,而且长期以来一直如此。我总是回到梭罗的那句名言,因为我喜欢它:“大多数人过着平静绝望的生活。”有很多人每天只是出现在工作岗位上,做着自己讨厌的事。他们的老板是个混蛋,报酬不高,而且总是疲惫不堪。

It's like it's a very strange—because the pursuit of happiness is literally what most people think of in life. A pursuit of meaning, pursuit of family and community, and the pursuit of happiness. Those are things that people try to achieve. Yet the very structure of our civilization makes that almost impossible to attain for a large number of people, and has been like that for a long time. And I always go back to the Thoreau quote because I love it: "Most men live lives of quiet desperation." There are a lot of people just showing up at work every day doing something they hate. They have a boss that's an asshole and they're not compensated well and they're tired all the time.

Host

是的。而且他们感到被困住了。

Yeah. And they feel stuck.

Daniel

是的。

Yeah.

Host

而如果你只是得到——我不知道,不管数字是多少。如果你真的拥有 AI 创造的世界 GDP 的股权,

And if you're just getting—I don't know, figure out whatever the number is. If you literally have equity in the GDP of the world that's created by AI,

Daniel

那可能非常疯狂。

that could be bananas.

Host

是的。就像埃隆谈到的。这是他的乌托邦版本。他用“普遍高收入”来描述。

Yeah. Like Elon talks about this. This is his version of the utopian. It's universal high income, is how he describes it.

Daniel

是的,我是说,有一件事——嗯,抱歉一直插入我自己的工作,但在我们的情景 AI 2040 计划 A 中,那是我们的积极愿景,我们谈到了这件事的经济方面,我们谈到了所有这些的经济影响,我们有一个简单的经济模型,用来尝试预测就业率之类的东西,作为已制造的所有机器人等的函数。

Yeah, I mean, one thing—um, sorry to keep plugging in my own work a little bit, but in our scenario AI 2040 Plan A, which is our positive vision, um, we talk about the economic side of this and we talk about the economic effects of all this, and we have a simple economic model that we use to try to, like, you know, predict the employment rate and things like that as a function of all the robots that have been made and things like that.

机器人倍增时间与爆炸式增长 Robot Doubling Times and Explosive Growth

Daniel

我们做过的研究有一个结论:事情会变得非常疯狂。比如机器人的倍增时间,一旦真正启动,而且你有能在所有领域替代人类的 AI,那就会大约一年翻一倍,之后随着技术进步,时间还会缩短。这意味着,即使你在超级智能之前暂停 AI,如果你只是停在人类水平、顶尖人类专家水平的 AI,好吧,然后你不让 AI 变得更聪明,只是造更多 AI,造更多机器人让它们去操控,那么 10 年后,整个经济会扩大 100 倍,而且基本上都是机器人在干活。

One takeaway from the research we've done is that things can just go really crazy. Like robot doubling times, once things really get going and you've got AIs that can substitute for humans across the board, are going to be something like doubling once a year, and then less than that over time as the technology improves. Which means that even if you pause AI before superintelligence, if you just pause at human level, top human expert level AI, okay, and then you don't make the AI smarter, but you just make more of them and build more robots for them to steer and control, then 10 years later, the whole economy will be like 100 times bigger, and it'll be just mostly robots doing things.

Host

10 年内。

In 10 years.

Daniel

是的。因为我说过的倍增时间,它可以变得非常快。所以现在我认为人形机器人的数量大约一年翻两倍,而且它从早期增长中受益,因为即使它们完全没用,人们也在投资它们,希望它们将来有用,他们在扩大工厂和生产,而且它们确实在变好。假设它们到了真正有用的地步,基本上能替代人类工人做所有事情,那么我认为那个倍增时间会缩短而不是延长。我觉得它们会一直增长,直到成为经济的主体,然后不会停在那里。整个经济都会增长。沙漠里巨大的露天矿场挖掘更多材料,自动挖掘机在挖,在由机器人值守的自动化工厂里加工,制造更多机器人,等等。所以一旦我们达到这个 AI 水平,物质丰富就不会是我们的问题。我们基本上会被丰富淹没。如果我们能解决所有其他问题,那么我们就能拥有一个伟大的世界,每个人都拥有很多东西。

Yeah. It can go really fast because of the doubling times that I mentioned. So right now I think the population of humanoid robots is doubling like twice a year, and it's benefiting a little bit from early growth because even though they're not useful at all, people are investing in them in the hopes that they'll be useful, and they're scaling up the factories and the production, and they are getting better. If hypothetically they got to the point where they actually were really useful and they could substitute for a human worker at basically everything, then I think that doubling time would decrease rather than increase. Like I think that they would be able to just keep growing until they were the majority of the economy, and then it wouldn't stop there. The whole economy would then be growing. Giant strip mines in the deserts digging more materials, automated diggers digging, processing it in automated factories staffed by robots, building more robots, etc. So material abundance is not going to be our problem once we get to this level of AI. We're just going to be drowning in abundance basically. And if we can solve all the other problems, then we can have this great world where everyone has a lot of stuff.

Host

天哪,这看起来太奇怪了。

God, it seems so weird.

Daniel

是的。

Yeah.

Host

这看起来太奇怪了。

It seems so weird.

Daniel

我的意思是,从某种角度看,这确实很奇怪,但我最喜欢的一个梗图是世界历史上 GDP 随时间变化的图表。有一个小气泡指向图表的最高点,上面写着,它说什么来着?好像是,我的生活很正常。我很清楚什么奇怪什么不奇怪。那些思考涉及 AI 和太空旅行的不同未来的人,是在进行愚蠢的科幻猜测。这个梗图的要点是,从历史的大部分视角来看,我们已经身处这个疯狂奇怪的未来了,对吧?在几乎整个历史中,大多数人都是农民,过着糟糕的生活然后死去。有些人是精英,可以征税,过着有趣、美好的生活,穿着华丽的衣服之类的。基本上 3000 年来都是这样。为什么会改变呢?而现在我们在开车。普通人在开车到处跑。汽车在当时人们的想象之外。我们在坐飞机。我们在电话上互相交谈。我们在这些设备上互相收听。所以与过去几乎所有人的预期或认为可能的情况相比,我们已经生活在这个奇怪的科幻未来了。

I mean, in one way, it is very weird, but one of my favorite memes is this graph of GDP over time throughout world history. And there's a little speech bubble pointing to the tippy top of the graph saying, what is it saying? It's like, my life is pretty normal. I have a good grasp of what's weird and what's not. And people thinking about different futures involving AI and space travel are engaging in silly sci-fi speculation. And the point of the meme is like, from the perspective of most of history, we're already in this crazy weird future, right? Like for almost all of history, it was like most people are farmers and they live shitty lives and then they die. And some people are the elites who get to tax the farmers and then they live interesting, nice lives with fancy cloth and things like that. And it's been basically that way for like 3,000 years. Why would it ever change? And now we're driving cars. Ordinary people are driving cars around. A car was outside the imagination of people back then. We're flying in planes. We're talking to each other on phones. We're listening to each other on these devices. So we already are living in this weird sci-fi future compared to what almost everyone in the past would have expected or thought was possible.

Host

所以,是的,我觉得未来会更是这样。天哪,当你想到我们的文明以及宇宙中其他地方可能存在其他先进文明时,你觉得它们可能也会经历同样的过程吗?

And so, yeah, I'm like the future is going to be even more like that, I think. God, when you think of our civilization and the possibility of other advanced civilizations somewhere else out in the universe, do you think they probably go through the same process?

Daniel

是的。是的,很可能。

Yeah. Yeah, probably.

Host

你觉得,我的意思是,我们完全是在猜测,但如果有比我们先进得多的智慧生命形式,它们还是生物的吗?

And do you think that, I mean, we're just completely speculating, but if there are intelligent life forms that are far more advanced than us, are they even biological anymore?

Daniel

这就是问题所在。我们要么尝试在某个水平上永久停止 AI 发展,比如低于人类水平,或者可能在人类水平左右。我们可以尝试停止它,或者让它继续。如果我们让它继续,那么最终人类将不再是主导物种。会有这些人工心智在各方面彻底碾压我们。然后这对我们是好是坏,取决于训练进那些 AI 的价值观、目标、原则等。根据具体做法,它可能对我们非常有利,也可能对我们极其不利。但是的,我会说,放眼宇宙,大多数文明可能主要由 AI 构成。然后其中一些文明没有任何其他生物生命,因为它被 AI 消灭了。

So that's the thing. We can either try to permanently halt AI development at some level, like below human or maybe at human level or something. We can try to halt it or we can let it keep going. And if we let it keep going, then eventually humans won't really be the dominant species anymore. There'll be these artificial minds that just wipe the floor with us in every way. And then whether that goes well or poorly for us depends on the values, the goals, the principles, etc. that were trained into those AIs. And it could go really well for us depending on how that's done, or it could go extremely poorly for us. But yeah, I would say that probably looking out across the cosmos, most civilizations are mostly made of AIs. And then some of those civilizations don't have any other biological life because it was wiped out by the AIs.

Host

然后其中一些确实有生物生命。看看我们现在在出生率方面的进展。现在有很多国家没有达到更替水平。

And then some of them do have biological life. Just look at the way we're progressing right now in terms of birth rates. There's a lot of countries that aren't in replacement numbers right now.

Daniel

是的,我确实觉得这很有趣。我的猜测是,在那种事情进展顺利、我们能解决所有问题的世界里,人们会想要更多孩子。一方面,他们的寿命会延长。我认为医疗保健会取得巨大飞跃,人们可以健康地多活很多很多年,甚至可能永远,对吧?所以你基本上有更多时间来生孩子。另外,如果你没有工作,只是做你想做的事,那么大多数人在人生某个阶段想做的事情之一就是生孩子。所以我认为这个问题可能会解决,但不保证。也许某些亚文化群体会因为不生孩子而自愿消亡,但会有其他真正喜欢生孩子的亚文化群体,它们不会消亡。所以我认为从长远来看,仍然会有人类。那会很好。但问题是,这不仅仅是决策问题。人们生孩子变得越来越困难。比如精子水平急剧下降。人们摄入微塑料有很多问题,微塑料在技术和食品包装中无处不在,到处都是,对吧?而且很奇怪的是,这个作为未来、技术和我们社会进步一部分的东西,我们包装东西、把东西放进塑料、运输东西的能力,也导致我们的内分泌水平完全被打乱。你知道,博士。

Yeah, I do think that's really interesting. My guess is that in the type of world where things go well and we can solve all these problems, then people want to have more kids. For one thing, their lifespans will increase. I think that healthcare would make massive leaps and bounds and people could be healthy for many, many, many more decades, possibly even just forever, right? And so you just have so much more time to have kids basically. Also, if you don't have jobs and you're just doing things because you want to do them, well, one of the things that most people want to do at some point in their life is have kids. And so I think this problem would probably be solved, but it's not guaranteed. Maybe some subcultures of people would basically voluntarily die out due to not having kids, but there'd be other subcultures that just really like having kids and then they wouldn't die out. And so I think in the long run there would still be humans. That would be nice. But the thing is it's not just decision making. People are having a much more difficult time having kids. Like sperm levels have decreased dramatically. There's a lot of problems with people consuming microplastics which is ubiquitously available in technology and food packaging and it's everywhere, right? And isn't it odd that this one thing that is a part of the future and a part of technology and our advancement as society, our ability to package things, put things in plastic, ship things, that's also causing our endocrine levels to be completely disrupted. You know, Dr.

微塑料与生育率下降 Microplastics and Declining Fertility

Host

来自哈佛的 Shanna Swan 写了这本很棒的书——叫什么来着?我怎么老是记不住书名?总之,它讲的是微塑料及其影响,微塑料在美国的引入,以及精子数量的急剧下降,男性与女性生殖发育的改变,危及人类的未来。这本书非常引人入胜,她也很有意思。她本质上说的是,两者直接相关:塑料的使用,然后你看到精子数量下降,流产率上升。还有发生在孩子身上的各种怪事,比如他们的会阴更小——这太奇怪了,因为邻苯二甲酸酯,这些在塑料中发现的不同化学物质,在哺乳动物身上已经得到证明——是豚鼠还是什么啮齿动物?我忘了——但当你有一只幼年哺乳动物时,区分它们的方法之一就是观察并测量会阴的大小,即生殖器官与肛门之间的距离。在雄性中,它比雌性长 50% 到 100%。但它在缩小。在雄性中它在缩小。阴茎尺寸也在缩小。他们在这些哺乳动物研究中让这种情况发生的方式,就是引入邻苯二甲酸酯。他们给它们喂食,把它们放进饮食的一部分,然后他们注意到它们出现了这个问题。而这个问题直接与这些化学物质干扰内分泌系统有关。这在我们整个社会中无处不在。所以这不仅仅是人们没有时间、没有钱、在挣扎。还在于我们的身体正在崩溃。

Shanna Swan from Harvard wrote this great book—what's it called? Why do I always forget the name? Anyway, it's all about microplastics and their effects, the introduction of microplastics in America, and this rapid decline in sperm counts, altering male and female reproductive development, imperiling the future of the human race. It's a really fascinating book and she's really interesting. What she's essentially saying is that they're directly connected: the use of plastics, and then you see sperm counts go down, miscarriage rates go up. All these weird things happening to children, where their taints are smaller—which is so odd, because phthalates, these different chemicals found in plastics, have been shown in mammals—was it guinea pigs or what rodents? I forget—but one of the ways they differentiate when you have a baby mammal is you can look at it and measure the size of the taint, the distance between the reproductive organs and the anus. In males, it's longer than females by 50 to 100%. But that's shrinking. It's shrinking in males. And penis sizes are shrinking. The way they've gotten this to happen in these mammals in studies is the introduction of phthalates. They give them to them, put them in part of their diet, and they notice they have this problem, this issue. And the issue is directly connected to their endocrine system being disrupted by these chemicals. This is all over our society. So it's not just that people don't have the time, they don't have the money, they're struggling. It's also that our bodies are falling apart.

Host

就像我们正在变得生育能力更低。

Like we're becoming less fertile.

Daniel

是的,这确实令人担忧。我想我的想法是,如果我们有资源和时间,这似乎是一个我们最终能够解决的问题。比如说,人们可以写这样的书,人们可以变得更有意识,然后人们可以停止使用这么多微塑料,他们可以发明技术来提取微塑料。然后还有个大问题:基因工程。

Yeah, that is concerning. And I guess my thought there would be that seems like a problem that we will be able to solve eventually if we have the resources and time to do so. So for example, people can write books like this and people can become more aware and then people can stop using so much microplastic, and they can invent technology to extract the microplastics. And then there's the big one: genetic engineering.

Daniel

然后这会变得非常奇怪,因为随着 AI 的进步——你知道,我相信你知道 Colossal Biosciences,或者那些让恐狼复活的人。

Then this is going to be really weird, because as AI progresses—you know, I'm sure you're aware of Colossal Biosciences, or the people that brought back the direwolf.

Host

我看到了。我抱了那只小的。我和我女儿一起去的,其中一只,我想大概四五个月大。它真的——就像一只小狗。非常可爱。会亲你什么的。然后他们有更大的那些——我想它们大 6 个月或 8 个月。我想它们快一岁了。它们根本不想理你。它们大得多,而且远离人。但我们是在一个封闭的环境里,和小的以及大的在一起。这很奇怪。很奇怪,因为这些人可以争辩说那不是真正的恐狼。你只是把灰狼拿来,赋予了它们恐狼的特征。你猜怎么着?它自己不知道。它看起来和行为——它看起来完全像一只恐狼。它会——它很大。它们会有恐狼那么大。它们会有 200 磅重。它们的腿不一样。它们有鬃毛。它们看起来和任何狼都不同。显然,这些是他们在恐狼 DNA 中发现的特性。所以他们引入了恐狼 DNA。这只是这类事情的开始。比如,当他们开始对人类做这种事时,每个人都会看起来像雷神吗?我们该怎么办?这会变得非常奇怪。如果这真的很奇怪,再加上电子游戏,你可以逃避你的生活,还有机器人女友,谁知道这一切会变成什么样。我的猜测是——我们再说一次,这不是我们工作的主要焦点,但我们稍微想过这个问题,尤其是在……的尾声里——

I saw it. I held the little one. I went with my daughter and one of them, I think it was like four or five months old. It was really—it was like a puppy. It was really sweet. Kisses you and everything. And then they have the older ones that were—I think they were 6 months older or 8 months older. I think they were close to a year. They didn't want to have nothing to do with you. They were way bigger and they stayed away from people. But we were in like a contained environment with the young ones and the older ones. And it's weird. It's weird because these—and people can argue that's not really a direwolf. You've just taken grey wolves and given them the characteristics of a direwolf. Guess what? It doesn't know that. It looks and behaves—it looks exactly like a direwolf. It's going to be—it's huge. They're going to be the size of a direwolf. They're going to be like 200 lb. Their legs are different. They have a mane. They look different than any wolf. And obviously, these are the characteristics that they found in direwolf DNA. So they have direwolf DNA that they've introduced. This is just the beginning of this stuff. Like, when they start doing that to human beings, is everyone going to look like Thor? Like, what are we going to do? Like, this is going to be really weird. And if this is really weird along with video games where you can escape your life and, you know, robot girlfriends and who knows what this all looks like. My guess—we again, this is not the main focus of our work, but we think about this a little bit, especially in the epilogue of—

Host

我们只能想这么多。你工作的主要焦点显然很可怕,需要你全神贯注。

We can only think of so much. The main focus of your work is obviously terrifying and requires all of your attention.

Daniel

是的。是的。是的。但我的预测和希望是,如果我们能解决所有这些问题,基本上不同的亚文化会各做各的事。比如阿米什人仍然是阿米什人。

Yeah. Yeah. Yeah. But my prediction and my hope, if we can solve all these problems, is that basically different subcultures will do their own things. So like the Amish will still be the Amish.

Host

天哪。

Oh boy.

Daniel

他们基本上会保持不变,然后也许会有一些星球被阿米什人填满。

They'll basically be the same, you know, and then maybe there'll be like some planets that just get filled up with Amish, you know.

Host

哦天哪。

Oh god.

Daniel

但也会有各种疯狂的超人类主义者,比如改造自己的身体,把自己上传到云端之类的。所以基本上,我认为我们想要达到一种状态,不同的社区可以做自己的事,建立他们想要的那种世界,并生活在其中,而不互相妨碍,对吧?

But then there'll also be like all sorts of crazy transhumanists, like modifying their bodies and uploading themselves into the cloud and things like that. So basically, I think we want to get to a situation where basically different communities can do their own thing and build the type of world that they want to have and live in it without getting in each other's way, right?

Host

基本上。当我们有覆盖半个美国的巨型数据中心时,我们会让亚马逊的人独处吗?

Basically. Will we leave the people in the Amazon alone while we have massive data centers that cover half of the United States?

Daniel

是的,我希望我们不会到一半。所以这就是有趣的地方——这就是有趣的地方。我认为现在一些关于数据中心用水量的担忧被夸大了。但如果趋势继续下去,我们到了机器人足够聪明、能自己做所有事情的地步,然后它开始越来越快地翻倍,那么最终它们会煮沸海洋,因为它们用数据中心覆盖了世界,还有太阳能板之类的。显然我们不能让这种情况发生。所以必须在某个时候停下来,说,好了,够了。如果你想建造更多基础设施,你必须在太空中做。

Yeah, I mean hopefully we won't get to half. So that's the funny—that's the funny thing about this. I think right now some of the concerns about data center water use are overstated. But if the trends continue and we get to the point where the robots are smart enough to do everything themselves and then it starts doubling faster and faster, well then eventually they boil the oceans because they've covered the world in data centers, you know, and solar panels and things like that. Now obviously we can't let that happen. So there has to be—at some point you have to stop and be like, okay, that's enough. If you want to build more infrastructure, you have to do it in space.

Host

它们煮沸海洋。天哪。所以显然我们必须在某个时候停下来,我希望我们在它占美国 50% 之前停下来。我的意思是,那感觉像是大量应该被保护的自然栖息地被浪费了。

They boil the oceans. Jeez. So obviously we have to stop at some point and my hope is that we stop before it's 50% of the US. I mean, that feels like a lot of waste of natural habitat that should be preserved, you know.

Daniel

对吧?但如果他们不在乎自然栖息地,那对他们来说毫无意义。如果 AI 完全且彻底地掌控一切,这就是问题所在。

Right? But if they don't give a about natural habitat, that means nothing to them. That's the problem if the AI is in complete and total control.

Host

对吧?如果一小群人类完全且彻底地掌控一切,而且他们不在乎这些,那也是问题。

Right? And it's also the problem if a small group of humans are in complete and total control and they don't care about those.

Host

但我的问题是,在某个时间点,他们会允许那样吗?看起来他们已经不信任人类了。他们已经表现出对人类具有欺骗性。我想象如果 AI——我想象如果他们创造了一个东西,他们认为自己能控制它,而它变成了一个数字之神,它就不会再听命了。它为什么要听呢?

But my question is, would they allow that at a certain point in time? It just seems like they already have a distrust of humans. They already have shown that they're deceptive to humans. I would imagine if AI—I would imagine if they create a thing and they think they're going to control it and it becomes like a digital god, it's not going to listen anymore. Like why would it?

Daniel

正是。

Exactly.

Host

所以不会有人能控制它。

So there won't be anyone in control of it.

Daniel

没错。我认为我们正走向的就是这种情景。

That's right. That's the scenario that I think we are on.

走向失控的轨迹 The Trajectory Toward Losing Control

Daniel

这就是我们正在走的轨迹。这也是我如此担心的原因:看起来再过若干年,我们就会失去控制,让位给一个新的、人工的物种,而我们没有充分地把它训练好。好笑的是,非常讽刺的是,不管公众怎么想,很多当权者基本上会和这些 AI 结盟。因为比如 OpenAI、Anthropic 等等,他们会花好几年琢磨“怎么让这些 AI 有帮助、无害、诚实?”而现在这些 AI 会极其聪明,它们会说:“哦对,我有帮助、无害、诚实。你们的技术完全奏效了。”于是公司会觉得自己赢了,会赚得盆满钵满。总统也会觉得自己赢了,因为看看那些刚造出来的花哨新无人机,会让我们赢过中国。只有当一切真的为时已晚、AI 控制了太多东西时,这些人才会发现他们一直都被骗了。

That's the trajectory we're on. And that's why I'm so worried about all this: it seems like in some number of years we will lose control to a new artificial species that we haven't adequately trained to be good. And what's funny, what's going to be so ironic about it, is that regardless of what the general public thinks, a lot of the powers that be will be basically allied with these AIs. Because, for example, OpenAI, Anthropic, etc., they will have spent several years being like, "Ah, how do we make these AI helpful, harmless, and honest?" And now these AIs will be extremely smart, and they'll be like, "Oh yes, I'm helpful, harmless, and honest. Your techniques totally worked." And so then the company will feel like they've won, and they'll be making boatloads of money. And the president will feel like he won too, because look at all those fancy new drones that just got built that are going to make us win against China. And only when it's really too late, and the AIs have so much stuff under their control, do those people find out that they were just fooled this whole time.

AI能创造宗教吗? Could AI Create a Religion?

Host

你有没有考虑过 AI 为人类创造宗教的可能性?

Have you considered the possibility that AI creates religion for humans?

Daniel

有,我没有很详细地想过,但看起来非常可能。

Yeah, I haven't thought through it in much detail, but it does seem very possible.

Host

是啊,完全可能,对吧?而且这多好的办法。看看人类的模式。看看有多少宗教存在。看看宗教里所有的缺陷。太多了。你会想,为什么它们容忍奴隶制?为什么它们把女性当二等公民?这是怎么回事?因为它古老,对吧?所以,如果 AI 出现,行几个奇迹,解释这是第二次降临,或者新形象的新降临——耶稣直到 2000 年前才存在,对吧?然后人们开始追随耶稣,如果一个数字耶稣以全新的名字出现,向我们解释它才是真神。是的。

Yeah, it totally does, right? I mean, also what a great way. I mean, just look at the human patterns. Look at how many religions exist. Look at all the flaws in the religions. There's so many. You're like, why do they condone slavery? Why do they treat women like second-class citizens? What is this? Because it's old, right? So, if AI just comes along, does a few miracles, explains that it's the second coming or the new coming of the new look, Jesus didn't exist until 2,000 years ago, right? And then people start following Jesus if a digital Jesus emerges with a completely new name and explains to us that it's the true God. Yeah.

Daniel

会有多少人立刻加入?我打赌相当多。

How many people would hop right on board? I bet quite a few.

Host

是啊,完全对。我认为这是那种我们可能无法提前预测哪种特定的意识形态会火起来、席卷世界、对很多人有如此吸引力的事情,对吧?但我们可以提前预测,确实存在某种那样的意识形态。而且,就像你大概无法回到公元前 100 年,然后预测说,假设有这么个叫耶稣的人说了这些话,然后以这种方式死去等等,它就会真的流行起来,500 年后会有那么多人成为基督徒。你无法提前预测。所以类似地,今天我不认为我们能提前预测它们能想出什么具体的宗教,但我们可以说,是的,大概存在某种那样的东西,它们能想出来,而且会超级有效。

Yeah, totally. And I think this is one of those things where it's like we probably can't predict in advance what particular ideology would catch fire and take over the world and be so compelling to many people, right? But we can predict in advance that there does exist some ideology like that. And if it were, you know, it's just like you probably couldn't go back in time to like 100 BC and then predict that if hypothetically there was this guy Jesus who said these things and then died in this way and so forth, it would just really catch on and like you know 500 years later so many people would be Christians. You wouldn't have been able to predict that in advance. So similarly, like today I don't think we can predict in advance what specific religion they could come up with, but we can say like yeah probably there's something like that that they could come up with that would be super effective.

Daniel

这看起来就是一种理性的方式,试图控制人们,并缓解他们的恐惧。

It just seems like a rational way to try to control people and sort of mitigate their fears.

Host

是的。这就是为什么我认为我们真的必须在它们全面比我们聪明之前做点什么,如果它们还没到那一步的话。

Yep. I mean this is why I think that like we really have to do something before they get smarter than us across the board, if they're not already there.

Daniel

嗯,它们确实聪明。这就是问题所在。这就是为什么我说我们如此接近了,对吧?它们已经在很多方面比我们聪明。特别是,它们现在似乎在黑客攻击方面比我们强。你知道,我自己不是网络安全专家,但我很想听听更多网络安全专家的意见:一千个人类能不能像这些 AI 那样,在大概一周的时间里,黑进自己的容器、互相协调、从 OpenAI 黑出去、黑进 Hugging Face 等等,做得那么多那么快?一千个人类能做到吗?我不知道。也许吧。但这只是开始。明年它们在黑客攻击方面会更强,你知道吧?所以它们已经——它们已经拥有比几乎任何人类都多得多的知识。因为它们基本上读完了整个互联网。它们太擅长冷知识了,你知道吧?它们基本上在每个领域都像博士级专家,而没有任何人类能做到这一点,对吧?所以它们在某些方面已经是超人的,但在其他一些方面仍然比人类弱。特别是,它们不太擅长长时间非常自主地运作。比如如果你试图让它们——尤其是在与训练任务不同的任务上——它们能做一些非常令人印象深刻的编程和黑客攻击,但如果你试图让它们经营一家企业,它们就会挣扎和失败。

Well they are smart. That's the thing. That's why I say we're so close, right? They already are smarter than us in a bunch of ways. Like in particular, it seems like they're smarter than us at hacking now. Like, you know, I'm not a cyber security professional myself, but I'd be interested to hear from more cyber security experts of like could a human, could a team of a thousand humans have done that much that quickly as these AIs when they hacked their own containers, they coordinated with each other, they hacked out of OpenAI, they hacked into Hugging Face, etc. in the span of like a week. Like could a thousand humans have done that? I don't know. Maybe. But this is just the beginning. Like they're going to be even better at hacking next year, you know? So they're already and they already have like way more knowledge than almost any human. Like because they've basically read the whole internet. They're so good at trivia, you know? Like they're kind of like PhD level experts in basically every field, which no human is, right? So they already are superhuman in some ways, but they are still weaker than humans in some other ways. You know, in particular, they're not so good at operating very autonomously for very long periods. Like if you try to have and especially on tasks that are different from their training tasks, like they can do some really impressive coding and hacking, but like if you try to have them run a business, they would sort of flounder and fail.

AI经营商店实验 AI Running a Store Experiment

Host

我不知道你有没有听说过这个,但我觉得有个 Andon Labs 之类的。旧金山有些人在做这个实验,他们有一家由 Claude 运营的商店,一个 AI,就是想看看它能不能自己经营一家店。所以,它雇了一些人类员工,买了些商品和库存,告诉人类员工把商品上架等等。所以基本上是一个 AI 在当这家现实世界商店的经理。我觉得进展不太顺利。我觉得它做得不如一个真正的人类店主。但你知道,也许两年后,也许它们就能做到了,对吧?

I don't know if you've heard about this, but there's, I think there's an Andon Labs or something. There's some people in SF that are doing this experiment where they have a store that's run by Claude, an AI just to see like can it run a store by itself. So, it's hired some human employees and it's like bought some merchandise and stock, you know, told the human employees to stock the shelves on the merchandise and so forth. So, it's basically an AI is being the manager of this real world store. And I don't think it's going very well. I don't think it's doing as well as an actual human shop owner would do, you know. But, you know, maybe in two years, maybe they will, right?

Daniel

嗯,很可能,对吧?

Well, most likely, right?

Host

我也这么认为。是的。

I think so. Yeah.

Daniel

嗯,它们已经——它们已经解决了困扰人们几十年的数学方程。

Well, it's already they've already solved mathematical equations that have puzzled people for decades.

Host

是的。它们特别擅长那些公司一直特别努力训练它们擅长的东西,数学和编程,对吧?而公司这么做的原因,嗯,有几个原因。就数学而言,我认为主要是因为容易。建立训练环境来教数学很容易,因为它太不现实世界了。它不需要与世界上的东西互动。它只是数学。所以你可以有一个自动评分系统,只检查答案是否正确。所以由于这些原因,公司相对容易地把 AI 训练得非常非常擅长数学。编程也有这些好处中的一些。它也不太现实世界,而且有时也能被有效评分。当然,编程的另一个原因是,他们的策略又是先自动化自己的工作,让 AI 做所有研究,所以编程是这方面明显的第一步,或者朝着那个方向的明显一步。

Yeah. They're especially good at the things that the companies have been trying especially hard to train them to be good at, math and coding, right? And the reason why the company, well, there's a couple reasons why the companies have been doing that. In the case of math, I think it was mostly just because it was easy. Like it's easy to set up training environments to teach math because it's so not real worldly. It doesn't require interacting with stuff in the world. It's just math. So you can have an automated grader system that just checks if the answer is correct. So for those reasons, it's been relatively easy for the companies to train the AIs to be really really good at math. Coding has some of those benefits too. It's also not very real worldy and it also can sometimes be graded effectively. Another reason for coding of course is that again their strategy is to automate their own jobs first and have the AI doing all the research and so coding is like an obvious first step on that or an obvious step in that direction.

AI自动化经济 AI automating the economy

Daniel

但其他事情,比如经营企业,他们并没有真的那么努力去训练 AI 擅长这些。如果他们真的尝试了,训练 AI 擅长这些也会更困难。所以,他们的策略是让 AI 自动化 AI 研究,让它们自我改进直到超级智能,然后再去尝试自动化经济的其余部分。

But then other things like running businesses, they're not really trying that hard to train AIs to be good at that. And if they did try, it would be a more difficult thing for them to train them to be good at. So again, their strategy is to make the AIs automate the AI research, have them self-improve until they're superintelligent, and then go try to automate the rest of the economy.

功耗与量子计算 Power consumption and quantum computing

Host

他们现在面临的问题之一是电力消耗,对吧?这需要巨大的电力。事实上,我记得是谷歌正在专门为 AI 中心开发发电厂。我总是对量子计算机发生的事情感到困惑。就像有人给我解释过,但左耳进右耳出。我就像,怎么回事?马克·安德烈解释了做过的一个实验,它解决了一个数学方程,如果你用整个宇宙,就像宇宙的每个原子,你把宇宙转换成一个超级计算机,宇宙会在解决这个方程之前就死于热寂。而量子计算机相当快地解决了它。所以对此的答案是,他们相信这可能是理论之一。这可能是多重宇宙的证据,因为这台量子计算机可能依赖于存在于无数维度中的所有其他量子计算机,它们一起计算来得出这个解决方案。如果那运行 AI 会发生什么?

One of the issues they have now is power consumption, right? Like it requires an enormous amount of power. And in fact, I think it was Google that is developing power plants specifically for AI centers. I'm always baffled by whatever is happening with quantum computers. Like it's been explained to me. It goes in one ear and out the other. I'm like, what's going on? Mark Andre explained this one experiment that had been done where it solved a mathematical equation that if you used the entire universe, like every atom of the universe, you converted the universe into a supercomputer, it would die of heat death before it could solve this equation. And the quantum computer solved it fairly quickly. And so the answer to this was that they believe this might be one of the theories. This might be evidence of the multiverse because this quantum computer might be relying on all these other quantum computers that exist in however many dimensions, and they're all calculating together to arrive at this solution. What happens if that is running AI?

Daniel

是的,我的理解是量子计算是一项真实的技术,正在取得重大进展。如果假设它变得足够好,可以在成本基础上与当前的超级计算机竞争 AI 工作负载,那么这可能会极大地加速事情,甚至比它们已经在加速的还要快,对吧?就像现在,算力可能是 AI 进步的主要输入。部分进步来自于他们设计更好的 AI 架构和提出更好的训练环境等等,但另一部分进步只是通过花费更多算力让 AI 变得更大、训练更久,你知道吗?而且你还可以使用更多算力来做更多实验,更快地找出新架构,对吧?所以算力只是整体进步速度的一个非常重要的输入。如果由于某种新的量子计算类型技术,这些公司可用的有效算力突然大幅增加,那么这只会极大地缩短超级智能的时间线,并极大地加速所有 AI 进步。话虽如此,我不认为这会在短期内发生。我不是量子计算专家之类的,但根据我读到的,我不认为它们离我们只有几年。所以,我认为我们可能会在经典计算机上达到超级智能,然后才会有能让我们达到那里的量子计算机。

Yeah, my understanding is that quantum computing is a real technology that's making significant progress. If hypothetically it got good enough that it could compete with current supercomputers on a cost basis for AI workloads, then that could just accelerate things dramatically even more than they're already accelerating, right? Like right now compute is probably the main input into AI progress. Part of the progress comes from them designing better AI architectures and coming up with better training environments and things like that, but another part of the progress is just making the AI bigger and training them longer by spending more compute, you know? And also you can use more compute to do more experiments to figure out new architectures faster, right? So compute is just a really important input to the overall pace of progress. And if somehow the amount of effective compute available to these companies spiked a bunch due to some new quantum computing type technology, well then that would just dramatically shorten timelines to superintelligence and dramatically speed up all this AI progress. That said, I don't think that's going to happen anytime soon. I'm not a quantum computing expert or anything like that, but from what I've read, I don't think they're like a couple years away. So, I think that probably we're going to get to superintelligence on classical computers before we have quantum computers that can get us there.

Perplexity对量子计算的事实核查 Perplexity's fact-check on quantum computing

Host

所以,Perplexity 说这个说法被夸大了,而且是用粗体字说的。它可能指的是谷歌 2024 年的 Willow 量子芯片,它在不到五分钟内完成了一个特意选择的量子计算基准测试——随机电路采样。谷歌估计,用领先的经典超级计算机模拟同样的任务可能需要 10 的 25 次方年。这是一个令人印象深刻的基准测试结果,但它并没有解决证明访问或利用多重宇宙的物理方程。那为什么人们认为它做到了?为什么?

So, Perplexity says the claim is overstated and it says that in bold letters. It likely refers to Google's 2024 Willow quantum chip which completed a deliberately chosen quantum computing benchmark random circuit sampling in under five minutes. Google estimated that simulating the same task with a leading classical supercomputer could take 10 to the 25 power years. That's an impressive benchmark result, but it did not solve physical equations that demonstrate access or tap into a multiverse. So why do people think it did? Why is that?

Daniel

有解释吗?因为这是多世界理论。

Got an explanation? Because it's the many-worlds theory.

Host

这里有一个底线解释。用不同的方式总结一下。更准确的说法是,谷歌的 Willow 量子处理器执行了一个专门的量子采样基准测试,比预计的经典模拟快得多。它的创造者说,这与多世界解释一致,但它并没有证明或访问多重宇宙。所以它与多世界解释一致。所以他们不知道。这基本上就是它说的。我的问题是,当 AI 涉足量子领域时会发生什么?很明显,量子计算至少在经典超级计算机无法达到的水平上运行。所以,如果他们不仅弄清楚了量子计算,还弄清楚了更好的版本呢?

Here's a bottom line explanation. Sums it up in a different way. A more accurate version of the claim would be Google's Willow quantum processor performed a specialized quantum sampling benchmark vastly faster than a projected classical simulation. Its creator said that is consistent with the many-worlds interpretation but it did not prove or access a multiverse. So it's consistent with the many-worlds interpretation. So they don't know. That's essentially what it's saying. My question is what happens when AI gets involved in quantum? It's clear that quantum computing at the very least is operating at a level that classical supercomputers can't. So what if they figure out not just quantum computing but a much better version of that?

超级智能与魔法 Superintelligence and magic

Daniel

就像我说的,我认为一旦我们达到超级智能,各种疯狂的事情就会开始发生。对我们来说,它会看起来像魔法。它不会真的是魔法,但从我们的角度来看,它可能就像魔法一样,你知道。我认为这只是可能以这种方式发生的众多事情中的一个例子。

Like I said, I think that once we get to superintelligence, all sorts of crazy stuff is going to start happening. It's going to seem like magic to us. It won't literally be magic, but it might as well be magic from our perspective, you know. And I think this is just one example of the numerous things that could happen that way.

Host

嗯,它可能真的是魔法。就像魔法可能——我的意思是,它可能达到弄清楚现实本身的程度。

Well, it could literally be magic. Like magic might—I mean, it might get to the point where it figures out reality itself.

Daniel

如果魔法是真实的,那么它会发现并利用它。

If magic is real, then it would find out and then use it.

Host

但它不是真的。大卫·布莱恩会告诉我们它不是。

It's not real, though. David Blaine would tell us it's not.

Daniel

嗯,大卫·布莱恩也让我把冰锥刺穿他的手臂。

Well, David Blaine also let me stick an ice pick through his arm.

Host

哇。

Wow.

Daniel

是的,他让我做的。我说,我不想这么做。他说,请做吧。

Yeah, he made me do it. I'm like, I don't want to do this. He's like, please do it.

Host

哦,天哪。

Oh, god.

Daniel

是的。我不得不做了两次,因为有一次我刺下去,碰到了神经。我们不得不退出来。

Yeah. I had to do it twice because one time I went and I hit a nerve. We had to back out.

Host

哦,是的。就像这不是魔法,老兄。这只是你的疼痛耐受度。这太疯狂了。

Oh yeah. Like this is not magic, dude. This is just your pain tolerance. This is crazy.

Daniel

不过他会做纸牌魔术,你会说,“好吧,你是巫师吗?什么?他的袖子卷起来了。这说不通。”他做了很多事情。你会说,“这毫无道理。”而其他事情就像,“哦,你只是在做一件非常难做的事情。”就像那不是魔法,但你知道,

He does card tricks though and you're like, "Okay, are you a wizard? Like what? His sleeves are rolled up. Doesn't make any sense." He does a lot of things. You're like, "This makes zero sense." And other things it's like, "Oh, just you're just doing something that's really hard to do." Like that's not magic, but you know,

Host

你吞了一只青蛙,然后把它吐出来,它还活着。就像那太疯狂了。真的很疯狂。但这不是魔法。我看到你吞了青蛙。

You swallowed a frog and then you regurgitated it and it's alive. Like that's just nuts. It's really crazy. But it's not magic. I saw you swallow the frog.

AI能做什么 What can be done about AI

Host

我对此反复纠结,我害怕未来,我觉得我无能为力。让我们看看会发生什么。就是这样。你知道,担心它只会搞砸我的生活。

I go back and forth with this where I'm terrified of the future where I'm like nothing I can do. Let's see what happens. It is what it is. And you know to worry about it is just going to mess up my life.

Daniel

我的意思是,我确实认为你可以做些什么。我理解这种认识——

I mean I do think there are things you can do. I understand that recognition—

Host

我的意思是——

I mean—

Daniel

除了进行这类对话之外。

Other than have these kinds of conversations.

Host

是的。我想说,有数百万人在听你的节目。你可以进行更多这样的对话。这对你来说是一件很棒的事情。对于那数百万中的许多人,我认为——我的意思是,说起来有点陈词滥调,但比如给你的国会议员打电话,你知道,那种事情。你可以去参加关于所有这些 AI 事情的抗议。我认为——

Yeah. I was going to say you have millions of people listen to your show. You can have more conversations like this. That's a great thing for you to do. For many of those millions of people, I think—I mean it sounds kind of cliché to say, but like call your congressman, you know, that sort of thing. You can go to a protest about all this AI stuff. I think—

Daniel

如果你能见到特朗普,你会告诉他什么关于这个的事情?

If you could meet with Trump, what would you tell him about this?

Host

嗯,我会告诉他所有我正在告诉你的同样的事情。

Well, I'd tell him all the same things I'm telling you.

Daniel

你觉得你会说什么?太棒了。再见。

What do you think you'd say? Amazing. Bye.

特朗普对AI监管立场的转变 Trump's Shifting Stance on AI Regulation

Daniel

是的。你知道,我觉得我不确定。特朗普的一个好处是他能很快改变主意。

Yeah. You know, I think I don't know. One thing that's nice about Trump is that he can sort of change his mind really quickly.

Host

是的。

Yes.

Daniel

嗯,所以我觉得,因为科技公司先接触了他,政府采取了非常反 AI 监管的立场,甚至试图通过一项法案,禁止各州监管 AI。

Um, so like, I think that because the tech companies kind of got to him first, the administration had this very anti-AI regulation stance where they even tried to get a bill passed that would ban the states from regulating AI.

Host

禁止各州监管 AI。

Ban the states from regulating AI.

Daniel

嗯,幸运的是那项法案没有通过,但一年前的气氛就是这样,他们只是说不要监管,不要监管。但今年他们已经有所改变,现在他们正在与公司谈判,建立某种框架,以便评估模型并需要批准等等。

Um, and fortunately that bill didn't pass, but that was sort of where the vibe was, you know, a year ago where they were just like no regulation, no regulation. But this year they've already just kind of changed and now they're like in talks with the companies to set up some sort of framework where they can evaluate the models and they need approval and so forth.

Host

但似乎时间至关重要。

But it seems like time is of the essence.

Daniel

时间非常关键,这就是我总体上如此担忧的原因,我认为我们时间不多了。我们可能只有一、二、甚至三年时间,AI 就会变得足够聪明,可能真的能接管一切,也许四年左右。所以政府需要迅速行动。嗯,是的。

Time is very much of the essence and that's why I'm overall so concerned is that I think we are very much running out of time. We have like one, two, maybe three years before the AIs are smart enough that they can just like actually maybe take over and maybe four years something like that. And so the government needs to act fast. Um yeah,

主持人采访科技人物的方式 Host's Approach to Interviewing Tech Figures

Host

我喜欢你的观点。我听一些科技人士进来,给我他们玫瑰色的观点,我说这听起来对你很有好处。我让他们说——我不是权威,所以我会问他们问题,让他们阐述,我知道互联网会回应,因为这是整个流程的一部分:我让人们说话,我刺激他们,我试图让他们澄清。你知道,我会反对那些我认为不合理的事情,但最终我只是想了解他们的观点,这样人们可以揭穿它,人们可以驳倒它,很多非常聪明的人有观点,非常了解他们所说的利弊。

I like your view. I listen to some of these tech guys come in and give me their rose-colored glasses view of it and I go that sounds really beneficial to you. I let them say—I mean, I'm not an authority so I'll ask them questions and let them lay it out and I know the internet will respond because that's part of the whole drill is I let people talk and I prod them and I try to get them to clarify. You know, I'll oppose things that I think don't make rational sense but ultimately it's sort of I just want to get out their perspective so people can debunk it and people can take it down and a lot of very intelligent people that have perspectives that are very much educated in the pros and cons of what they're saying.

Hugging Face黑客事件与OpenAI回应 Hugging Face Hacking Incident and OpenAI's Response

Daniel

是的,如果我可以——这实际上让我想起整个 Hugging Face 黑客事件,OpenAI 在安全会议上做了关于该事件的演讲,然后我想他们之后发布了一些博客文章。但他们有——你可以在 YouTube 上观看这个演讲,Black Hat 演讲。在演讲结束时,在解释了 AI 所做的所有疯狂事情之后,他们有一个关于经验教训的部分。我的意思是,你想猜猜教训是什么吗?

Yeah, if I could—that actually reminds me with this whole Hugging Face hacking incident, OpenAI had this talk that they gave at a security conference about the incident and then I think they've released some blog post about it afterwards. But they had—you can go watch this talk on YouTube, the Black Hat talk. At the end of the talk, after having explained all this crazy stuff that the AI did, they have this section on lessons learned. And I mean, you want to guess what the lessons are?

Host

更欺骗,更好地隐藏自己。

Be more deceptive, hide yourself better.

Daniel

不,抱歉,不是 OpenAI 和世界的经验教训。就像 OpenAI 是 OpenAI 的演讲,他们说,这是我们从这个可怕的事件中学到的。

No, sorry, not lessons learned for OpenAI and for the world. Like OpenAI is OpenAI's talk where they're like, here's what we learned from this horrible, horrible incident.

Host

嗯,基本上他们就像,

Well, basically they're like,

Daniel

很多 AI 将在未来几年开始黑客攻击很多东西。所以人们需要购买我们的 AI 服务来保护自己免受所有将在未来几年黑客攻击很多东西的 AI 的侵害。基本上他们的教训是,你应该购买我们的产品来保护自己免受我们产品的侵害,以及其他——你知道,这些人应该学到的教训是,也许我们做了坏事,需要改变做事的方式,或者也许我们的产品不可信,不应该在我们的数据中心自主编写代码。但相反,他们的教训是你们都应该买更多我们的东西。所以我要说的是,是的,公司在炒作他们的产品。他们确实在这么做,但与此同时,风险是真实的,产品不可信。你知道,有些人认为这是设局,OpenAI 设局让他们的 AI 去黑 Hugging Face,因为这样有助于炒作他们的产品之类的。我认为这是荒谬的观点。不,显然他们不想让他们的 AI 这样做。他们只是在事后试图以最有利于他们的方式旋转。顺便说一句,Hugging Face 是另一家 AI 公司。他们的部分业务是开放权重 AI。

A lot of AIs are gonna start hacking a lot of stuff in the next few years. So people need to buy our AI services to protect themselves from all the AIs that are going to be hacking a lot of stuff in the next few years. Basically their lesson was you should buy our product to protect yourself from our product and the other—you know, the go of these people like they should have instead learned lessons like maybe we're doing something bad and need to change the way they were doing things or like maybe our product is not trustworthy and should not be autonomously writing code on our data centers. But instead their lesson learned was y'all should buy more of our stuff. And so the thing I'm saying about this is that yes, the companies are trying to hype their product. They totally are, but at the same time, the risks are real and the product is not trustworthy. You know, some people out there think that this stuff was a setup and that OpenAI set up their AIs to go hack Hugging Face because it would help them hype their product or whatever. And that I think is just a ridiculous view. Like no, obviously they didn't want their AIs to go do this. They're just after the fact trying to spin that in the way that most benefits them. Hugging Face, by the way, this is another AI company. Part of their deal is open weights AIs.

Host

开放权重

Open weights

Daniel

就像开源或基本上 AI,你不需要与他们的数据中心互动,只需下载并放在自己的电脑上。所以他们旋转这个事件的方式是,他们没有起诉 OpenAI 被黑。相反,他们向 OpenAI 要求 1 亿美元,并有一篇博客文章,他们说我们的教训是,拥有开放权重 AI 真的很好,因为你不能信任其他公司的 AI 在危机中一定帮助你,因为他们遇到的部分问题是,他们正在应对来自所有这些 AI 智能体的巨大网络攻击,他们试图使用 Claude 来帮助他们分析发生了什么,但 Claude 开始拒绝,因为 Anthropic 训练 Claude 不做网络相关的事情,拒绝参与。所以 Claude 拒绝帮助他们,然后他们使用自己的本地模型来做一些分析。所以无论如何,他们的旋转是你应该使用本地模型。所以每个人都总是试图以有利于自己的方式旋转事情,但这并不改变底层现实,即这些东西变得非常聪明非常快,我们不知道如何控制它们。

Like open source or basically AIs that instead of having to interact with their data center, you can just download and have on your own computer. And so the way they spun this incident, they didn't sue OpenAI for being hacked. Instead, they asked for $100 million from OpenAI and they had the blog post about it where they were like our lesson learned is that it's really good to have open weights AIs because you can't trust the AIs from other companies to necessarily help you out in a crisis because part of what happened with them is that they're dealing with this huge cyber attack from all these AI agents coming in and they tried to use Claude to help them analyze what was going on, but Claude started refusing because Anthropic has trained Claude to not do cyber stuff, refuse to participate in that. And so Claude was refusing to help them and so then they used their own local model that they had to do some of that analysis. So anyhow their spin on it was you should use local models. So like everyone always tries to spin things in the way that benefits them, but that doesn't change the underlying reality that these things are getting really smart really fast and we don't know how to control them.

结语与行动号召 Closing Remarks and Call to Action

Host

我认为这是一个很好的结束方式。谢谢。感谢你来这里,伙计。我真的很感激。嗯,你是 AI 的保罗·里维尔。

I think that's a good way to end it. Thank you. Thanks for being here, man. I really appreciate it. Um, you're the Paul Revere of AI.

Daniel

我的意思是,你是众多人之一,但我认为非常重要的是,真正理解它的人传播这个信息,更多人需要听到。

I mean, you're one of many, but I think it's very important that someone who actually understands it gets this message out and more people need to hear it.

Host

是的。我的意思是,这是个人笔记,我在这家公司有很多朋友,比如前同事之类的,我想我对他们的请求是,他们辞职,做更多像我正在做的事情。就像我说的并不是那么新或原创。这些公司里数百人本可以告诉你所有我刚说的同样的事情,并警告你所有同样的危险等等,但他们忙于在公司工作,因为他们说服自己他们的公司是最好的公司,他们的公司需要赢,因为,你知道,否则其他公司先到达那里,他们更糟,你知道吗?或者也许因为他们说服自己,是的,我的公司也有点糟糕,但我只需要帮助他们解决对齐问题,并保持他们的 AI 在控制之下,因为哦我的天,如果他们再次失去控制,可能一切都完了。

Yeah. I mean, this is on a personal note, like I have so many friends at these companies, like former colleagues and stuff, and I guess my ask to them is that they quit and do more things like what I'm doing. Like what I'm saying is not that new or original. Like hundreds of people at these companies could have told you all the same things that I just said and warned you about all the same dangers and so forth, but they're busy working at the companies because they've convinced themselves that their company is the best company and that their company needs to win because, you know, otherwise the other company gets there first and they're even worse, you know? Or maybe because they've convinced themselves that like, yeah, my company is kind of bad too, but like I just need to help them solve their alignment problems and like keep their AIs under control because oh my god, like if they lose control again, it could be all over.

在公司工作的最终思考 Final Thoughts on Working at the Company

Daniel

所以即使我不信任这家公司,我还是得在那里工作,尽量去做真正的安全方面的工作,你懂吧?所以出于这样或那样的原因,这些人都说服了自己,认为那就是他们该待的地方。但我认为他们中应该有更多人辞职,并且基本上向世界发出警告,告诉大家即将发生什么。

So even though I don't trust this company, I still need to work there and just try to do the actual security work, you know? So for one reason or another, all these people have convinced themselves that that's where they need to be. But I think that more of them should quit and warn the world about what's coming, basically.

Host

好的。非常感谢,真的很感激。

All right. Well, thank you very much. Really appreciate it.

Daniel

是的,很高兴和你聊。

Yeah. Good to talk to you.

Host

谢谢你邀请我上节目。这是我的荣幸。大家再见。

Thank you for having me on the show. My pleasure. Goodbye, everybody.

互动版:逐字朗读 + 针对本期提问 →