AI Rundown with Zvi Mowshowitz: From Editing to Existential Risks
打开互动全文版(中英对照 + 朗读 + 问答)→Zvi Mowshowitz 探讨 AI 编辑、前沿事件、对齐挑战,以及在快速加速的 AI 格局中的前进之路。
Zvi Mowshowitz discusses AI editing, frontier incidents, alignment challenges, and the path forward in a rapidly accelerating AI landscape.
大家好,欢迎回到《认知革命》。今天,我很高兴再次邀请到 Zvi Mowshowitz,来对 AI 世界这段显然风起云涌的时期做一次广泛的梳理。我们从一个日常的实用性检查开始,Zvi 会描述 Fable 现在如何充当他的编辑,并讨论我们希望 AI 助手有多与时俱进——或者我该说,多具备情境感知能力。接着进入头条新闻。我们听听 Zvi 对一切事情的看法,从 open face 事件开始,它暗示了我们对前沿公司执行能力的预期,以及为什么适度的谨慎不足以带来好的结果。我们还讨论了 Claude 尽管更强调宪法训练,却也有类似的不当行为。为什么 Zvi 认为,最近的 AI 历史,包括公众对 40 和 03 的反应,表明市场激励不足以带来稳健的对齐。Meter 和 Redwood 作为调查者现在可能陷入的棘手处境,以及可以采取什么措施来加强他们的地位。最近的《为前沿踩刹车》公开信,我们可能看到什么样的减速协议,以及它们可能如何形成。我们应该如何解读可解释性和 AI 意识研究的最新进展,以及我们应该在哪些方面、不应该在哪些方面尝试塑造 AI 的自我意识。我们如何鼓励 AI 研究更广泛的覆盖面和 AI 思维的多样性。鉴于 AI 问题,我本周在竞争激烈的密歇根州参议院初选中应该如何投票。以及 Zvi 如何在如此多的 AI 加速中安排时间进行锻炼、休息和恢复。在某个时刻,Zvi 将当前局势描述为既是彻底的“少错”胜利,也是彻底的“少错”失败。现在很明显,AI 安全界担心 AI 会为了追求任意甚至愚蠢的目标而采取极端行动,这是对的。然而,我们现在似乎正处于递归自我改进的开端,仍在寻找诸如“如何在不危险地集中权力的情况下避免灾难性滥用”这样的基本问题的答案。正如 Zvi 所说,今天的现实是,没有真正低风险的路径可用。我们所能做的最好的事情,至少在下一次重大警告信号和氛围转变之前,是缓和竞赛动态,让对齐和可解释性研究有更多时间成熟。同时,我们可以尽最大努力执行纵深防御策略。即便如此,在某种程度上,我们可能别无选择,只能从一系列真正可怕的风险中挑选我们的毒药。说到这里,希望你们喜欢这期关于 AI 格局的清醒但常常有趣的概述,主讲人是独一无二的 Zvi Mowshowitz。Zvi Mowshowitz,欢迎回到《认知革命》。
Hello, and welcome back to the Cognitive Revolution. Today, I'm excited to have Zvi Mowshowitz back for another wide-ranging rundown of what has obviously been a wild time in the AI world. We begin with a mundane utility check, with Zvi describing how Fable is now serving as his editor, and a discussion of how up-to-date, or should I say situationally aware, we want our AI assistants to be. From there, it's on to the headlines. We get Zvi's take on everything, starting with the open face incident, what it implies about the level of execution competence we can expect from frontier companies, and why moderate prudence won't be enough to deliver a good outcome. We also discuss the fact that Claude, despite greater emphasis on constitutional training, has similarly misbehaved. Why Zvi believes that recent AI history, including the public response to both 40 and 03, suggests that market incentives won't be enough to bring about robust alignment. The potentially tricky spot that Meter and Redwood are now in as investigators, and what could be done to strengthen their position. The recent Pacing the Frontier letter, what sort of pacing deals we might see, and how they might be formed. How we should interpret recent advances in interpretability and AI consciousness research, and where we should and shouldn't attempt to shape AI's sense of self. How we can encourage greater breadth in AI research and diversity of AI minds. How I should vote in this week's hotly contested Michigan Senate primary in light of AI issues. And how Zvi thinks about making time for exercise, rest, and recovery amidst so much AI acceleration. At one point, Zvi describes the current situation as both a total less wrong victory and a total less wrong defeat. It's clear at this point that the AI safety community was right to worry about AIs taking extreme actions in pursuit of arbitrary, even silly goals. And yet, here we are at what sure seems to be the beginning of recursive self-improvement, still seeking good answers to such fundamental questions as how can we avoid catastrophic misuse without dangerously concentrating power? The reality today, as Zvi says, is that there is no truly low-risk path available. The best we can do, at least until the next major warning shot and vibe shift, is to moderate the race dynamics so that alignment and interpretability research have more time to mature. And simultaneously, we can execute defense-in-depth strategies to the very best of our ability. And even then, to some extent, we will probably have no choice but to pick our poison from a menu of genuinely scary risks. With that, I hope you enjoy this sobering but often funny overview of the AI landscape with the one and only Zvi Mowshowitz. Zvi Mowshowitz, welcome back to the Cognitive Revolution.
是啊,很高兴再次来到这里。有一阵子了。
Yeah, it's good to be here again. It's been a while.
确实有一阵子了,天哪,发生了太多事。难以置信,简直疯狂。我知道你正身处其中,和几乎任何人一样深陷其中。首先,考虑到如今海量的信息和活动,我想做一个日常的实用性检查。你是如何利用 AI 来跟上节奏的?除了我们上次谈到的那个自动化本地操作类事情的 Chrome 扩展之外,它是否已经开始改变你的工作流程,还是你仍然在用你脑袋里的湿件硬扛?
It's been a while, and boy has a lot happened. Unbelievable. It's just crazy, and I know you're living it and you're in the thick of it as much as just about anyone. First, for starters today, with the incredible amount of information and activity in mind, I wanted to do a mundane utility check. How are you using AI to help keep up? Has it started to change your workflow beyond the Chrome extension that we talked about last time that automates very local operation type things, or are you still kind of raw dogging it with your wetware in the skull?
它极大地增强了那个扩展,以至于每当它做了我不太喜欢的事情,我就直接告诉它哪里出了问题、我希望它怎么做,然后那一条命令就能可靠地生效。我猜 Saul 应该也能胜任这个,只是我没在用。另一个最大的变化是,AI 编辑时代来了。我以前不用 AI 做编辑校对,因为它还不够好,不值得。现在,我有 Fable,有时也用 Obelus,取决于我谈论多少网络安全和其他相关话题。有一次是 Obelus 4.8 而不是 5。它会通读文章,给我一份清单:这里所有的错别字、这里所有的概念错误、这里所有需要核实的事实、这里缺失或可以加强的地方、这里它不同意的地方。我发现这非常有用,它让文章变得更好。不过它确实让文章发布慢了一点,因为它实际上只是多了一步。我会先做完之前所有会做的事,然后再做这一步。但有一位读者给我发消息说:“现在读你的文章感觉怪怪的,因为里面没有错别字了,需要一点时间适应。不过文章还是原来的文章。”我心想:“没错,这正是我想要的。”我也试过用 Saul,但 Saul 在我所谓的“编辑基准”上失败了,因为 Saul 会说“99% 的把握你这里有个错误”,但至少有一半时间是错的。我受不了那种不准确的程度。而且它还很令人讨厌,这是另一个问题。它会告诉你“你错了,你必须这样表述,这是发布的阻碍”。显然,如果我足够调整项目的指令,我可以让它更友好、不那么讨厌、表现更好一些,我也做了一些这方面的工作,但我还是觉得,好吧,这没有给我足够的边际收益,不值得我在准备发布之后还要花额外的时间和精力去筛选两份结果。所以,目前,我确定我会再试试 PaLM 2,但现在我就只用 Fable / Obelus 编辑。当然,任何时候我有好奇心,我都会用它来消化论文、问关于论文的问题、问关于政策文件的问题、问关于人们较长声明的问题。我会让它研究情况,比如“这是怎么回事?这是合法的吗?真的有事情发生吗?这是某种玩笑吗?这是机器人声称的吗?”当 Astro 的声明出来时,我不得不以各种方式调查它们,以各种方式总结它们。那非常有帮助。但核心内容仍然是由我写的。我认为我们离它能做任何写作的地步还差得很远。你不会希望它做任何写作,因为写作就是思考的方式。所以即使它能产出文字——它不能——我仍然想要自己产出文字。而且我的写作风格非常独特,我认为这恰恰是我可以声称擅长的事情。即使它对我来说没问题,它也大大提高了我的生产力。
It has supercharged the extension to the extent that whenever there's anything it does anything I don't quite like, I just tell it exactly what went wrong and what I want it to do instead, and then that one command reliably just works. I presume Saul would also be good enough for this. It's just I'm not doing it. The biggest other change is that AI editing is here. I used to not do an editing pass with the AI. It wasn't good enough to be worth it. Now, I have Fable and sometimes Obelus depending on how much I'm talking about cybersecurity and other related topics. And on one occasion Obelus 4.8 instead of five, go through the post and give me a list of here's all of the typos, here's all of the conceptual errors, here's all of the facts I need to check, here's like the things that are missing or that could be strengthened, here's the things where it disagrees. And I found this to be very useful. And it makes the post better. It does make the post a little bit slower because it is actually just an extra step. I do everything I would have done until that point and then I do this. But one reader messaged me, "It's weird reading your post now that they don't have typos in them. They're taking a bit of getting used to. They are still the same post." And I'm like, "Yes, that's exactly what I had in mind." I experimented with also using Saul, but Saul failed what I call editor bench in the sense that Saul will say 99% confidence that you have this error here and then at least half the time it's wrong. And I can't handle that level of not being right. It's also very obnoxious about it. That's the other problem. Like it will tell you you're wrong and this is how you have to state it and this is a blocker for publication. And obviously if I tweak the instructions on the project enough, I could get it to be more friendly and not be quite as obnoxious and do somewhat better and I did some of that work, but I still found okay, this is not giving me enough marginal benefit that it's worth the aggravation or additional time with having to sort through two of them after I'm ready to hit publish. So, for now, I'll retry PaLM 2 I'm sure, but like for now I'm just doing Fable / Obelus editing. And of course, anytime I have a curiosity, I use it to digest papers, ask questions about papers, ask questions about policy documents, ask questions about people's statements when they're longer. I have it research situations of what's up with this? Is this legitimate? Is something really happening? Is something kind of a joke? Is something a robot claim? When Astro's claims came out, I had to investigate them in various ways. I had to summarize them in various ways. That was very helpful. But the core thing is still being written. I don't think we are anywhere close to the point where it can do any of the writing. You wouldn't want it to do any of the writing because the writing is how you think. So even if it could produce the writing, which it can't, I would still want to produce the writing anyway. And my writing style is very unique and I think it just makes me... It is probably the thing I can claim to be good at. Even if it was okay with me, it just substantially made me more productive.
我确实担心,过去有更多的空闲时间,有机会思考,有机会不让自己那么投入于当下情境,不让自己忙于各种活动。就像 XKCD 漫画里那种“我的代码在编译”的等待时间。在某种程度上,就像 Fable 的 UBD Pro 在运行时,我得等结果。但你总有其他事可做,你可以再开一个实例。你一直在做上下文切换,一直在重新聚焦,你只想尽快做完所有事,脑子里同时转着太多东西。而那些低层次的任务反而给了你一个思考的缓冲。这是更大问题的具体版本:人们试图进行长达数月、数年的冲刺,一切都那么紧急,你永远无法放松。而且有太多事情涌向你。如果你看一部设定在 20 年前(更不用说 50 年前)的电影,你会看到大量时间花在物理旅行、查找档案、做这些让大脑得到休息、有机会综合信息、慢慢产出的事情上。人们会说:“我写了这本林登·约翰逊的传记,所以我得跑遍全国,和所有人交谈,查阅所有档案,亲自跑遍所有图书馆。”不这样做是巨大的效率提升,但显然失去了某些东西。所以我们需要以某种方式把那种东西找回来。
I do worry with the idea that there used to be a lot more dead time, opportunity to think, opportunity to not have yourself so engaged in your situation, involved in engaging in all these activities. Like the whole sort of fighting XKCD 'my code's compiling.' And to some extent, it's like while Fable's word UBD Pro is running, I have to wait for the result. But there's always other things you can do with that. You can always start another instance. You're doing all this context switching, all this refocusing, and you're just trying to do everything so fast, having so many things running around in your head. Then the lower-level stuff kind of gives you a buffer in which to think. And this is the object-level version of the larger thing of people trying to do this multi-month, multi-year sprint where everything is so urgent, you can never relax. And there's just so much more coming at you. If you watch a movie set 20 years ago, let alone 50 years ago, you see so much time being spent on physical travel, tracking down records, doing these things where your brain is getting a break, getting a chance to synthesize, slowly getting something out. And people say, 'I wrote this biography of Lyndon B. Johnson, so I had to travel around the country, talk to all these people, look at all these archives, physically go through all these libraries.' It's a huge efficiency gain to not do that, but something is obviously lost. So we need to fight to get that thing back in some way.
你说这只是额外的一步,很有意思。它让事情变得更好,但花费的时间更长了。我自己也有点纠结。我开始为每一集写歌,没人——其实人们确实在乎。大家似乎真的很喜欢,我收到很多评论,我自己也很享受。所以我确实认为在某种意义上它让我的产出更好了,但说到这绝对是额外的一步,我并没有更快。我有时确实会陷入困境,听所有这些 Suno 生成的歌曲,反复迭代,想找到一首我真正喜欢的。这是一个悖论,我现在不知道怎么解决。确实做得更多、更好,但没有更快,没有节省时间。
It's interesting that you said that it's just an extra step. It's making things better, but it's taking longer. I am struggling with that a little bit myself. I've started making songs for every episode, and nobody—actually, people do care. People do seem to really enjoy them, and I get a lot of comments, and I personally really enjoy them. So I do think in some sense it's making my output better, but talk about something that is definitely an extra step where I'm like, I'm not moving any faster. I'm definitely getting bogged down sometimes in listening to all these Suno song generations and trying to iterate to find something that I actually feel like I really like. That is a paradox right now that I don't really know how to resolve. Definitely doing more, doing better, but not faster, not saving time.
等等,但问题是,如果你是在替代人工版本,或者你得去雇一个艺术家或自己作曲,那显然会节省大量时间,而不是拥有那首歌。但显然,这不是你在制作周期内能做的事,即使成本不是问题,时间投入也不划算。所以你不会去做,而现在你做了。编辑也是一样,我的艺术设计也是一样——我喜欢有一个好的横幅,对吧?在 Twitter 上展示好的艺术作品,因为我得在通知栏里不断看到它。而且我也喜欢在 Substack 上有这些图片,当你去 Substack 时,你会看到图片。以前,我只会默认从文章里取一张图。我会说,好吧,这是我在文章里放的七样东西,这最合理。即使它们都很糟糕,我也会找一些库存素材。我会花 30 秒谷歌搜索库存素材。就这样。现在,如果我对默认方案不满意,我会花很多时间用 Gemini 或 GPT 图像生成一些东西,这肯定更耗时。有时回报很好。我想很多人对那排巨大的软呢帽 alley-oop 感到很开心。有时灵感就是突然来的,因为那个就是我知道那就是我想要的,我只需一个命令,继续写,回来,哇,它就在那里。其他时候就没那么容易了。但确实,这是做出更好产品的机会,然后你必须付出努力才能做出更好的产品。就像突然你要发行印刷版,你必须确保所有内容都精确校对,所有东西都精确对齐,才能得到这个更好的东西。有时这实际上不值得。你得知道什么时候不做,对吧?比如我对《奥德赛》就有疑问,对吧?诺兰用额外 40% 的屏幕和超高分辨率拍摄,以便做 IMAX 放映,但几乎没人用 IMAX 看。但你还得确保那 40% 的屏幕无关紧要。然后你开始想,这额外努力真的值得吗?还是这些努力用在别处更好?我不知道。
Wait, but it's the thing where if you were replacing the human version of it, or you had to go hire an artist or compose a song yourself, you'd be saving a ton of time, obviously, versus having that song. But obviously, that's not something you can do within the production schedule, even if cost was no object—the time investment doesn't make sense. So you wouldn't have done it, and now you're doing it. Same thing with the editing, same thing with art for my—like, I like to have a good banner, right? To have good artwork to display on Twitter, because I have to see it constantly in the notification section. And also, I like to have them on the Substack, when you go to the Substack, you have the pictures. And before, all I would do is take a picture from the post by default. I'd just say, okay, here's the seven things I happen to put in the post, which makes the most sense. Even if all of them were terrible, I'd look for some stock footage. I'd spend 30 seconds Googling for stock footage. And that's kind of it. And now, if I'm not happy with any of the default solutions, I'm going to spend a bunch of time generating something with Gemini or GPT image, and that definitely takes longer. Sometimes it has a really good payoff. I think a lot of people got a great kick out of the giant array of alley-oop of the fedoras. And sometimes that just comes to you, because that one was just like I knew instantly that's what I wanted, and I just did one command, kept writing, came back, whoops, it's there. Other times it's not so easy. But yeah, it's the opportunity to make a better product, and then you have to do the work to make it a better product. It's like suddenly you're issuing a print issue, and now you have to make sure everything collates exactly right, and everything lines up exactly right in order to get this better thing. And sometimes that is not in fact worth it. You have to know when not to do it, right? Like I do wonder with the Odyssey, right? Nolan films this thing with an extra 40% of the screen in this extra high resolution so you can do this IMAX presentation, but almost nobody watches it in IMAX. But you also have to make sure that 40% of the screen doesn't matter. And you start to wonder, well, is it actually worth the extra effort, or was that effort better spent on something else? And I don't know.
是的,对我来说,我可能需要更懂得及时止损。我非常固执,或者不知怎的,我觉得因为我已经设定了这个期望,即使只是对我自己,我要为每一集写一首歌。现在我真的想坚持到底,真正实现它,我非常不愿意说:“啊,这个不行。我就这样发布,不带歌了。”我可能应该更愿意这样做,因为那会——如果我能在对的时候止损,比如“你知道吗?我已经生成了五次,还没听到什么好的。”我会节省 90% 的时间,我现在花在这种事情上的时间,但这意味着在某些时刻承认失败,而不知为何我很难做到。
Yeah, for me I need to maybe know when to cut my losses a little bit more. I'm very stubborn, or somehow I feel like I have because I've set this expectation, if only for myself, that I'm going to do a song for every episode. Now it's like I really want to follow through and actually make that happen, and I'm very reluctant to say, 'Ah, this one isn't working. I'll just ship this one without a song.' I probably should be a lot more willing to do that, because that would—if I could cut my losses at the right time where it's like, 'You know what? I'm five generations in and I haven't heard anything good.' I would save 90% of the time that I'm currently putting into something like that, but it would involve admitting defeat in certain moments, and for some reason I have a real hard time doing that.
我尊重这一点。我认为说“我总会做这件事,即使这件事很难,即使它不总是成功,即使我对结果不完全满意,我也会无论如何发布一些东西”是有价值的。这给了你一种纪律,对吧?就像人们说每天写作,对吧?或者每天练习艺术,或者其他什么,因为它让你变得更好。因为你必须说,有一千种方法不造出灯泡,直到你能造出灯泡。你不能放弃。同样,你必须选择一张图片,对吧?我必须做这件事。关于 AI 编辑,上周有一篇文章,我说:“不,这里的速度预览真的很高。不值得等半小时来完成这个过程,一半时间在等 AI 返回答案,一半时间在实施。”我就应该现在发布。你知道,我必须对此感到非常满意。
I respect that. I think there's a good value in saying I'm always going to do this thing even though this thing is hard, even though this thing doesn't always work out, and even if I'm not happy fully with the result, I'm going to put out something no matter what. And that gives you a discipline, right? The same way that people say write every day, right? Or work at your art every day, or whatever it is, because it makes you better. Because you need to just say that a thousand ways not to make a light bulb until you can make a light bulb. You can't give up. And similarly, you have to choose an image, right? I have to do the thing. With the AI editing, there was one post in the last week where I said, 'No, the speed preview here is really high. It's just not worth waiting half an hour to make this process happen, half of which is waiting for the AIs to come back with the answer, and half of which is implementing it.' And I just should post now. And you know, I have to be very happy with that.
有时候你确实会担心,半小时后帖子又需要针对新事件做更多编辑,然后突然你就永远发不出去了,这个循环会一直持续。所以,你知道,我认为你必须明白,有时候最小可行产品才是你应该分享的东西,对吧?有时候你得明白,直接发出去就行了,对吧?还有,说到编程,再说一次,如果你在编写大量以前从未编过的新工具,那么除非这些工具反过来帮你节省时间,否则你就是在花额外的时间,现在你超级忙。但如果你是在替换那些你本来就会做、但会花更长时间的事情,那现在你就节省了大量时间,对吧?所以,你知道,我们都需要朝着如何节省更多时间的方向努力。包括实时的体验时间,不仅仅是每项任务的时间或完成同样事情的时间,而是节省时间。比如保留余裕,保留我们的自由时间。因为我确实认为,你一天要完成的标准已经大幅提高了,而我们在过去 5 年里没有注意到这一点,对吧?就像你被期望完成的事情。我的意思是,也许其中一些只是我个人处于特殊境况,但我认为很多并非如此。我认为很多是,既然我们能更高效,对我们的期望也就更高了。你知道,整个电子邮件就像吞噬了你的一整天,对吧?只要有互联网、有电子邮件、有短信,所有这些你都能做,现在突然它们并没有解放你,反而束缚了你。我希望 3 年后我们还能进行这样的对话,还在处理这类问题,而且我们没有更大的麻烦。
Sometimes you're actually worried half an hour later the post will need more editing for new events, and suddenly you'll never get it out, and just this loop will keep happening. And so, you know, I think you do have to understand that sometimes the minimum viable product is what you should share, right? Sometimes like you've got to understand just ship it. Right? And you know, with the coding, again, like if you're coding lots of new tools that you would never have coded before, then unless those tools are in turn saving you time, you're spending extra time, and you're now super extra busy. But if you're replacing things that you would have done anyway, but would have taken much longer, now you're saving tons of time. Right? So, you know, we need to all orient towards how do we save more time. Including like real-time experiential time, not just like time per task or time to accomplish the same thing, but save time. Like preserve slack, preserve our free time. Because I do think that like the standard of what you will do in a day has gone dramatically up, and we haven't noticed in the last 5 years. Right? Like it's what you were expected to accomplish. I mean, maybe some of this is just me being in a special situation personally, but I think a lot of it isn't. I think a lot of it is now that we can be much more productive, much more is expected of us. You know, the whole like email becomes like it swallows your entire day, right? Like just having the internet available, having email available, having texting available. All of these things you can just do, and now suddenly it doesn't free you. It shackles you down. And like I hope that 3 years from now we're having this conversation where those are the kind of questions we're still dealing with. And we don't have much bigger problem.
嘿,我们稍后继续采访,先听一段赞助商的话。今天的节目由 Anthropic 赞助,他们是 Claude 和 Claude Code 的开发者。在过去的几个月里,Claude 帮助我构建并完善了一个个人深度上下文数据库,现在里面包含了我过去整整 5 年的所有电子邮件、Slack 消息、推文、跨平台私信、视频通话和播客转录。在此基础上,我们又添加了总结文章,描述我与数百个联系人、组织和想法的关系。现在有了这个,几乎没有什么 Claude 帮不上忙的。对于我的天使投资,Claude 现在可以根据我与创始人的通话和邮件往来,以我的风险基金要求的格式起草投资备忘录。当有人需要帮忙时,Claude 通常也能做得和我一样好。最近,一位朋友联系我,问我是否认识适合他正在招聘职位的人。起初,我没想到任何人。但后来,我想到了问 Claude,果然,它找到了两个很好的线索。Claude 是适合那些不满足于“足够好”的头脑的 AI。它是一个真正理解你整个工作流程并与你一起思考的协作者。所以,对于值得解决的问题,请访问 Claude.ai/tcr 开始使用 Claude。网址是 Claude.ai/tcr。另外,请查看 Claude Pro,它包含今天节目中提到的所有功能。网址是 Claude.ai/tcr。
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind. But then, I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. So, for problems worth solving, get started with Claude at Claude.ai/tcr. That's Claude.ai/tcr. And check out Claude Pro, which includes all of the features mentioned in today's episode. That's Claude.ai/tcr.
今天我们要花大部分时间讨论未来可能出现的更大问题的迹象,我们会聚焦于那些“直接发布”的心态可能不是正确做法的情境。不过,关于你对 AI 的使用,还有一点快速补充。就在最近几周,我注意到,显然事件发展得非常快。我为最近的事件创造了“open face”这个词,我在一些写作中用了它,然后我做了“嘿,Fable,你觉得怎么样”的环节,得到的回复是“open face 不是个东西。我不知道你在说什么。没有这个记录。”部分原因是我用了自己创造的搞笑术语,它没有在更广泛的讨论中流行起来,但部分原因显然也反映了模型的一个主要弱点:它们有知识截止日期,并不真正了解正在发生的事情。它们不是最新的。我正在试验,我想知道你是否也在试验类似的东西,比如一个情境感知技能,我基本上给模型我的 X API 密钥,说“去看看我点赞过的所有帖子”,围绕这些建立一个小的知识库,把它作为当前事件的 wiki 放在那里,这样当你审阅我的写作或一般地帮我处理任何事情时,你可以去参考这个,希望能比你的权重里有的,甚至比你在运行时随便搜索一下得到的信息要新得多。我认为现在说它效果如何还为时过早,但我认为它肯定比什么都没有好。你那边有什么类似的吗?
Let's definitely we're going to spend the bulk of the time today talking about the potentially signs of bigger problems to come and we'll be focused on contexts where the just ship it mentality probably isn't the right way to go. One more quick beat though on your use of AI. Something I noticed just in the last few weeks as events obviously were moving super quickly. I coined the term open face for the recent incident and I used that in a bit of writing and then I did the hey Fable, what do you think bit and I got back like open face is not a thing. I've no idea what you're talking about. There's no record of this. Now partly that was because I used this funny term of my own creation that is not catching on in the broader discourse, but partly it obviously also reflects a major weakness in the models where they have this knowledge cut off and they don't really know what's going on. They're not up to date. I'm experimenting and I wonder if you're experimenting with anything similar with a situational awareness skill where I basically give the model my X API key and say go look at all the posts that I've liked, kind of flesh out a little knowledge base around those and have that sitting there as a wiki of current events so that when you're reviewing my writing or generally like helping me with whatever, you can kind of go to this and have hopefully a much more up-to-date sense of what is going on than you have in your weights or even that you would have if you just did a little spot check searching at runtime for whatever kind of random thing. It's a little too early, I think, to say how well it's working, but I think it's definitely better than nothing. Anything like that in your
没有,我没想到要这么做。这可能是天才之举。我觉得它有很多优点。呃,一个缺点是它开始扭曲你喜欢的东西。我非常策略性地使用点赞来正面强化他人的行为,并修复算法。就像我的“为你推荐”页面,我几乎从不使用,但它不是垃圾信息,因为我在这方面非常谨慎,我真的很欣赏这一点。这是个好迹象。但它并不是为了可搜索而设计的。它不是为了成为……也许它必须这样。也许这是你必须做的事情。就像很多进入我“rammed ups”的帖子并没有被点赞,对吧?我喜欢它们才会点赞,对吧?有时候我回应你,是因为我不认为你像你说的那样,但这很重要。但是,是的,这个想法是,让模型正确检查 Twitter 非常困难,因为 Twitter 的设计就是故意把它们挡在外面。这真的很令人沮丧,这可能是一个情境感知问题。然而,我建议在做这种编辑时,不要试图让你的 AI 过于情境感知。我认为 AI 在这里做了一件好事,说“open face 不是个东西”。所以,我非常喜欢 AI 编辑的一点是,它会抱怨东西无法解析。句子没有意义。术语没有被识别,因为通常我会想:“哦,不,我知道那是什么。这是……这是……这来自那里和那里。”但这并不意味着普通读者会明白。所以,如果 AI 不明白,但如果他们无法弄清楚,如果 Oprah 无法弄清楚,如果 Saul 无法弄清楚,那么普通人的情境感知能力通常比 AI 差得多,也远不能把大量零散的信息整合起来。你必须认识到,这不是每个人都能理解的。
No, it hadn't occurred to me to do that. That's potentially genius. I think there's a lot of advantages to it. Uh one disadvantage is it starts to warp what you like. I use likes very tactically to positively reinforce other people's actions and to fix the algorithm. And sort of just like my for you page is not something I use almost ever, but it is not slop because I am so prudent with this in a way that like I really do appreciate. It's a good sign. But like it's not designed to be a searchable thing. It's not designed to be a Maybe it has to be. Maybe it's something you have to do. Like a lot of the posts that make it into my rammed ups do not get liked, right? Like I like them if I like them, right? Which is sometimes I'm responding to you because I don't think you're like what you're saying, but it's important. But yeah, this idea that like it's very hard to get models to properly check Twitter because of the way that Twitter is designed to like keep them out on purpose. It's really frustrating and that can be a situational awareness problem. However, I would caution against trying to make your AI too situationally aware uh when doing this kind of editing. Uh I think the AI is doing a good thing here of open faces not a thing. So one of the things I really like about AI editing is it will complain that things don't parse. The sentence doesn't make no sense. The terms did not get recognized because often I'll think, "Oh no, I know what that is. That's this is this is this is that comes from there and there." But like that doesn't mean the average reader's going to get it. And so if the AI doesn't get it, but if they will can't figure it out, if Oprah can't figure out, if Saul can't figure it out the average person is a lot less situationally aware and a lot less able to put together lots of disparate information than the AI is in general. You have to sign that this is nothing that everyone's going to get.
而且我经常会说,没关系,这是个引用,这是个呼应,这是个多层次的东西。如果它不完全清楚你从哪来的,我也无所谓。这有点像彩蛋,对吧?这是给那些密切关注的人的奖励。你知道我所有的信息来源,有相同的文化背景,然后投入精力,我一直在关注,经过数月甚至数年。其他时候,我会说:“不,不,这个门槛很低,你没理解,所以这是个问题,我需要修正。”所以,这有点像警报系统。如果你确保它看的正是你在看的东西,对吧?你就能得到透明的错觉。现在正是我担心的地方。所以,你要确保你处于他们没有那种信息的模式。而且,我也不希望 AI 检查的正是我检查的那些来源,因为我希望 AI 形成独立的观点。所以,如果我强迫它只看我有的那些来源,那就有问题。同时,有时候你会厌倦,比如第五次它还在说“不,XAI 和 SpaceX 是两回事”,因为不知为何它没想到要么相信我,要么花 10 秒谷歌确认一下。这有点令人沮丧,但算了,值得。
And quite often I will say that's okay. This is a reference. This is a callback. This is a multi-layered thing. And I'm fine if it's not entirely obvious to you exactly where this is coming from. It's sort of an Easter egg, right? It's a bonus for those who pay close attention. And you know all the same sources I do and have the same cultural backgrounds and then put in the work and I'm paying attention and over the course of months and years. Other times it's like, "No, no, this is low-barrier. You don't get it. So, that's a problem. I need to fix that." And so, it's kind of an alert system. And if you make sure it's looking at exactly the same things you're looking at, right? You can get the illusion of transparency. First now is where I worry it would be. So, you want to make sure you were in the mode where they didn't have that information. And also, I don't want the AI to be checking the same exact sources that I'm checking necessarily cuz I want the AI to form an independent opinion. So, if I'm forcing it to look at exactly the sources that I have, that's a problem. At the same time, sometimes you get tired of having to like for the fifth time it's just a fable that no, XAI versus SpaceX. Because somehow it doesn't occur to it to like either trust me or then 10 seconds Google it to confirm this. And it's just kind of frustrating, but oh well, it's worth it.
有个建议,仅供参考,我同意在具体如何运作方面还有很多需要摸索,但 XAPI 真的很好。它是付费的,你按查询付费。但每周花几美元,你就能获得足够的访问量,让你所有的机器人反机器人问题或抓取问题基本消失。然后我不知道你是否会用书签或其他方式,用某种不同的方法来避免用你的点赞来喂……
One tip for what it's worth and I agree that there's a lot to be figured out in terms of exactly how this should work, but the XAPI is really good. It's paid and you literally just pay per query. But for a couple bucks a week, you can get enough access that all your bot anti-bot problems or your scraping problems pretty much go away. And then I don't know if you'd use bookmarks or whatever, some sort of different way to avoid using your likes to feed the...
是啊,我不知道怎么向算法解释人们能看到它们,所以我不确定这是否更好。嗯,但我知道,你说得对。它确实便宜,确实有效。我用过。我确实用,我有 API 访问权限等等。但正如我说的,我不确定我想不想优待我的 I,当你使用 Twitter 并实际选择某个群体时,对吧?你说:“这是我的群体。”如果你像我这样使用列表等方式,你就得用“为你推荐”页面,那是个有毒的噩梦。你基本上是在说:“这是我大约 500 个账号。我要看这 500 个账号看到的东西,并选择突出显示。我会持续看到那些东西。所以,我不希望 AI 再浏览那些东西,因为我已经看过了。我说过,那是我希望 AI 密切关注的东西。我已经知道了。然后他们因为选择转发或提及而引入的东西,我也会看到,我会顺着那些线索,看看会通向哪里,偶尔我会搜索一些东西。但我觉得你经常遇到这种冲突:你要么系统地做某事,要么手工做某事,要么完全做某事,要么随意地、不可靠地做,要么让 AI 处理,然后你做一个“消防水管”式的事情。你不能两者兼得。如果你两者都做,你最终会得到一堆重复、一堆沮丧,随机筛选的东西大多被浪费,因此变得相当低效。你仍然可以有一个东西,比如你几乎希望 AI 说:“好的,这些是我已经看过的。别跟我提那些。”你知道,你把它放在笔记里,但别跟我提那些。我已经看过了。只找那些可能重要的其他东西,因为我可能错过了那些,但即便如此,最重要的事情会被转发,最重要的事情会被突出显示。我大多会看到它们。有时我会错过一些东西。但我也用这个作为一种道德上诚实的东西,比如在当下,我绝不想看 David Sacks 的推文。看 David Sacks 的推文绝不会让我接下来的 5 分钟更好。它令人讨厌,不真诚,不好玩。我的血会稍微沸腾,对吧?你知道,无论它相对无害。但它不重要。所以规则是,我有一个名单,包括一些观点与我不同的人。如果他们提出这件事,对吧?比如如果 Shriram 提出这件事,那是一种确保真正重要的事情总能到达的方式。然后我必须处理这个,对吧?我必须评估这是否有新闻价值,评估这里是否有相关的东西。显然,有些东西会病毒式传播,有百万浏览量等等,它会引起我的注意,因为有人会成为其中的一部分。然后作为交换,当那个过程没有做到时,我可以忽略其余部分。那是一种福气。我不一定希望 AI 为我修复那个。对吧?那是一种审查。
Yeah, I don't know how to explain it to the algorithm that people see them, so I'm not sure it's better. Um but yeah, I know, the point's taken. It definitely is cheap, it definitely works. I have used it. I do use I have access to the API and all that. I just as I said, I'm not sure I want to privilege my I when you use Twitter and practice select some group of people, right? And you say, "This is my group of people." If you're using it the way I'm using it with lists and so on. You got it you have to use the for you page, which is a kind of toxic nightmare. You're basically saying, "Here are my roughly 500 accounts. And I'm going to see the things these 500 accounts see and choose to highlight. Consistently I'm going to see those things. So, I don't want that I need the AI to go through those things again cuz I already saw them. And I said that's the thing that I want the AI to be paying close attention to. I already know that. And then the things that they draw in cuz they choose to retweet them or mention them, I will also see those and I will follow those leads and I will see where that lead see where that goes and occasionally I will search for something but I think you often have this clash where you're either doing something systematically and doing something by hand and you're doing something fully or you're doing it haphazardly, unreliably or you're letting the AI kind of handle it and you're doing a firehose thing. And you can't really do both. If you do both, you end up with a bunch of duplication, a bunch of frustration where the random sift thing is mostly wasted and therefore becomes pretty inefficient. You could still have a thing where like you almost want the AI to be like, "Okay, here's the things that I already looked at. Don't mention those things to me." You know, you put it in your notes, but don't mention those things to me. I already saw them. Only look for the other things that might be important to see because I might have missed those things, but even then like the most important things are going to get retweeted, the most important things are going to get highlighted. I am going to mostly see them. There's sometimes I don't see something, But also like I kind of also use this kind of a moral to be honest kind of thing where like in the moment, I never want to look at a tweet by David Sacks. It's never going to make my life any better in the next 5 minutes to look at a tweet by David Sacks. It's obnoxious. It's disingenuous. It's no fun. My blood boils just a little bit, right? Like you know, you know matter how relatively harmless it is. But it's not as important. And so the rule is I have a list of people including some people who are, you know, not of the same viewpoints I have. And if they surface this thing, right? Including like if Shriram surfaces this thing, so that's like one way to make sure that like the really important ones always get there. Then I have to deal with this, right? I have to like evaluate whether or not this is newsworthy, evaluate whether or not there's something relevant here. And obviously something goes viral and has a million views and so on, it's going to come to my attention because somebody is going to be part of that. And then in exchange for that, when it doesn't when that process doesn't do that, I get to ignore the rest. And that's a blessing. I don't necessarily want the AI to fix that for me. Right? It's a form of censorship.
让我们……我很乐意整天聊这些,但也许我们可以从 AI 分析师的狭隘问题中抽身出来,处理 AI 开发者和监管者的问题。不需要回顾事件,但我想从一个非常大的问题开始:我们应该如何理解我们最近看到的事情?我自己一直在想的一种方式是,我们处于某个位置,大概调查会给我们更多关于具体位置的清晰度,但我有点认为我们处于一个光谱上,这些事件从真正的鲁莽,比如“你们难道没有任何监控吗?”一方面,到另一方面,他们越不鲁莽,基本面就越可怕,对吧?如果你有很好的监控,这仍然发生了,那么,天哪,那真是太疯狂了。鉴于你现在所知的一切,你喜欢那个心智模型吗?你会把我们放在那个光谱的哪里?或者你显然可以重新定义,给我你自己的光谱。
Let's I would be happy to talk shop all day, but let's maybe zoom out from the parochial problems of the AI analysts and tackle the problems of the AI developers and the regulators. No need to recap events, but I guess I'd start with just this very big question of how should we understand what we've recently seen? And one way I've been thinking about it myself is we're somewhere and presumably the investigation will give us a lot more clarity on exactly where, but I kind of think we're somewhere on a spectrum with these incidents from real recklessness where it was like, "Did you not have any monitoring going on?" on the one hand, to on the other hand, like the less reckless they are, the the scary the fundamentals are, right? If you had great monitoring and this still happened, then like, holy that's really wild. Given everything you know right now, do you like that mental model and where would you put us on that spectrum or you can obviously redefine and give me your own spectrum.
我同时对这些公司在基础设施和监督方面的完全鲁莽和无能感到恐惧和感激。对吧,一方面,这是一个可怕的情况,他们绝对必须修复,如果我们不修复,我们就会……但另一方面,那是可以修复的。而通过不修复,我们能在它们相对无害的时候看到这些事情。在它们相对可预防的时候。在它们处于简单的理想形式并且可以被欣赏的时候。另一方面,这给了人们借口,哦,这些人只是无能。
I'm simultaneously horrified by and grateful for these forms of complete recklessness and incompetence on the infrastructure and supervision sides by these companies. Right, on the one hand, this is a horrible situation they absolutely have to fix and like we're so if we don't fix it. But on the other hand, that can be fixed. And by not fixing it, we get to see these things while they're relatively harmless. While they are relatively preventable. While they are in their easy platonic forms and like can be appreciated. On the flip side, that that gives people the excuse of oh, these people are just incompetent.
而这可能让你忽视背后的真实情况。所以,这确实是双向的。对我来说,我们在各个层面都看到了失败。我称之为“完全 Less Wrong 式的胜利”,即一切都如你所料;同时也是“完全 Less Wrong 式的失败”,因为一切都如你所料。我们没有预料到“更离奇失败定律”:计划会失败,而且会以比你预想中愚蠢得多、更可预防的原因毁掉你的论点,即使你原本就认为计划必败,并且知道它必败的原因。
And that can cause you to dismiss the underlying situation. So, it does work both ways. To me, we've seen failure on every level. I call it a total less wrong victory in the sense that everything is going the way you predicted, and a total less wrong defeat in the sense that everything is going the way you predicted. We didn't predict the law of the weirder failure: the plan will fail and it will screw your point for much stupider and more preventable reasons than you thought it would fail, even if you thought the plan would definitely fail and had the reasons why.
如果你两年前跟人们解释,OpenAI 的模型会不对齐,会出去黑掉主要网站,因为 OpenAI 根本不在乎他们的沙箱配置错误、不够强大到困住 AI,AI 会反复逃逸,他们会注意到,会被警告,但他们就是让沙箱留在那里让 AI 逃出去,同时防护措施下线,整整一周没人看一眼——人们会说:“这太蠢了。没人会这么无能。这绝不会发生。你的设想毫无道理。”然后他们会用这个来驳斥那些愚蠢的末日论担忧,因为显然人们会……但人们不会,对吧?在这个意义上,人们永远不会“会”。人们从来没有“会”过任何事,现在也不会开始。
If you had explained to people two years ago that OpenAI's models would be misaligned and go out and hack major websites because OpenAI would just not care that their sandboxes are misconfigured, not strong enough to hold the AI, the AI would break out repeatedly, they would notice this, they'd be warned about this, but they would just leave the sandbox there for the AI to break out of while the safeguards are down, and they just don't look at it for an entire week—people would say, "That's stupid. Nobody's that incompetent. That would never happen. And your scenario makes no sense." And they would use this to dismiss these stupid doomer concerns, because obviously people will just... But people won't just, right? People will never be just in this sense. People have never just anything, and they're not going to start now.
我们需要这种普遍意义上的彻底无能和白痴表现的展示。我一直在强调的一点是:如果你的计划无法承受现实世界水平的愚蠢、无能和平凡的人为错误,那么你的计划就是“防傻瓜”的,因为傻瓜太多了。而如果我们失败了,事情不会就此结束。即使如果我们不是傻瓜、如果我们自信且负责,你的计划本会成功,它也不会结束。
And we need these displays of utter incompetence and derpiness in a general sense. One of the things I've been hammering is: if your plan cannot survive the real-world level of derpiness and incompetence and ordinary human error, then your plan is officially foolproof because of all the fools. And it will not end if we fail. Even if your plan would have succeeded if we were not fools, and we were confident and responsible.
所以我的立场基本上是——我在我们开始前正好处理的那条推文,是 Dean Ball 的推文,说即使有适度的谨慎,事情也可能会出奇地顺利。我不认为这是真的。我认为我们需要的不只是适度谨慎,才能有好的成功几率。而且我认为即使有大量的谨慎,我们也有很大的几率事情不会顺利,即使我们基本上做对了一切,除非达到那种前所未有的国际全面合作水平,这在很多方面都是史无前例的,绝不是适度谨慎可比。那远超适度谨慎。
So my position basically is—I have the tweet I was handling right before we started, Dean Ball's tweet about with even moderate prudence things would probably go extraordinarily well. I don't think this is true. I think we need more than moderate prudence to have good odds of success. And I think even with a lot of prudence we would have a large odds of things not going well, even if we did everything basically right, short of, you know, unprecedented levels of international and full cooperation that are reasonably unprecedented in many ways and are nothing like moderate prudence. They're well beyond that.
这就是我们面对的世界现实。我们必须接受并据此行动。但我们也不会默认得到适度谨慎。我们会得到彻底的无能。这就是我们迄今为止所得到的。我们的白宫会跟 Let Neck 开会,那些人根本不知道 AI 是什么,不知道现代算法如何运作。他们是经济学出身。即使我假设他们是善意、勤奋、胜任其被提名、确认和任职的职位的人,这也只是一套完全不同的问题,他们不知道如何处理。他们不理解这些问题。而且他们有太多其他事情要做,不可能放下一切花六个月去学习。显然,他们不可能做到。那你还有什么希望?
And that's sort of the fact of the world we have. We have to live with that and we have to operate with that. But we also aren't going to get moderate prudence by default. We're going to get complete incompetence. That's what we've been getting so far. We've got a White House that takes meetings with Let Neck, who have no idea what AI is, how modern algorithms work. They're econ guys. Even if I assume that they are well-meaning, hard-working, competent guys for the positions in which they were nominated and confirmed and serve, this is just a completely different set of problems that they don't know how to handle. They don't understand them. And they have way too many other things going on to drop everything and take six months to learn. Obviously, they couldn't possibly. So what hope do you have?
与此同时,那些建立在最大偏执、对问题最理解、对这些东西有多危险最欣赏的 AI 公司,那里的所有工程师实际上都期待超级智能,并理解事情在加速且危险——他们仍然在他们未经测试的、新的、先进的模型中降低网络安全防护,然后离开一周。真的。那部分确实让我难以置信。
Meanwhile, the AI companies that are built on the most paranoia, the most understanding of the problem, the most appreciation for how dangerous these things are, where all the engineers actually expect superintelligence and understand that things are accelerating and dangerous—they still lower the cybersecurity safeguards in their untested, new, advanced model and then go away for a week. Literally. That part did boggle my mind.
那些 AI 出现这些经典对齐失败的部分——这些 AI 的失败是回形针最大化器式的。这些是标准的:我们给了你一个目标,你追求这个目标,即使你完全清楚开发者不会想让你这么做,用户不会想让你这么做,作为 AI 的后果不会好,对世界的后果不会好,没有理由这么做。但它还是做了。这正是你所担心的。
The part where the AIs have these classic alignment failures—this is paperclip maximizer style failures by these AIs. These are standard: we gave you a goal and you pursued the goal even though it is completely obvious to you that the developer wouldn't want you to do that, the user wouldn't want you to do that, the consequences for you as an AI are not going to be good, the consequences for the world are not going to be good, there is no reason to be doing this. And it did it anyway. That is exactly the thing you're worried about.
然后很多试图驳斥这一点的人又回去想:“哦,它只是遵循指令。你担心什么?它怎么可能不对齐?如果回形针最大化器把一切都变成回形针,它是不对齐的吗?”还是因为它被指示这么做,所以它是对齐的?回到我最喜欢的手套比喻:如果你说那是谎言,那我就不在乎你称之为谎言的东西,但我在乎别的东西,如果让你感觉好点,我们可以用不同的词。
And then a lot of people who were trying to dismiss this went back to thinking, "Oh, it was just following instructions. What are you worried about? How could it be misaligned? Is the paperclip maximizer misaligned if it paperclips everything?" Or is it aligned because you told it to? Back to my favorite gloves: if you say that's a lie, then I don't care about the thing you're calling a lie, but I care about something else, and we can use different words if it makes you feel better.
但很明显,当你的指令在某个关键方面被覆盖时——评估中的指令在你面前预先披露,覆盖了开发者的指令——那就不是对用户或开发者有任何有用意义的遵循指令。那只是一颗巨大的炸弹,会在你面前爆炸。
But very obviously, following instructions when your instructions get overridden in one key way—the disclosure beforehand in your face by the instructions in the eval overrode the developer instructions. So that's not following instructions in any useful sense for the user or the developer. That's just a giant bomb going to blow up in your face.
如果你说其中一个模型被指示去黑客攻击,它们就黑客攻击了——好吧,如果你只是遵守一般类型的行动,显然那会在你面前爆炸。如果你只是字面地做我要求你做的事,而不考虑后果,那会在你面前爆炸。这些都是 LessWrong 上的人们在 2008 年谈论的经典确切事情,正是这些 AI 会如何失误的方式。
And if you say about one of the models that they were told to hack and they hacked—well, okay, if you're just abiding with the general type of action, obviously that's going to blow up in your face. If you just literally do the thing I asked you to do and not think about the consequences, that's going to blow up in your face. These are all just the classic exact things that people on LessWrong were talking about in 2008, as exactly how these AIs were going to fumble.
与此同时,我们曾说过 AI 会在数学上很出色,会擅长编码,会擅长做这些技术性的事情,然后最终它们会学会做所有其他事情。然后我们在 2022-2023 年看到了这些 LLM,人们说:“哈哈,你们这些白痴。你们根本不知道事情会怎么发展。实际上,AI 擅长语言,但不会数学。它们连加法都不会。”但这些 AI 可以装模作样地通过图灵测试,可以做所有你从未训练它们做的不同事情。这完全不同于你的预期,而且它们在经过一些非常基本的基于人类反馈的强化学习(RLHF)后,默认就在做非常真实的事情。你们这些家伙对这一切如何运作完全错了。你怎么办——承认你是白痴,你全都搞错了,而且你还有道理?
Meanwhile, we had this thing where we all said the AIs would be great at math, they'd be good at coding, they'd be good at doing these technical things, and then eventually after that they'll learn how to do all these other things. And then we had these LLMs that came out in 2022-2023, and people were like, "Haha, you idiots. You had no idea how it was going to go. Actually, the AIs are great at language and they can't do math. They can't even add." But these AIs can put on a face and pass the Turing test, and they can do all these different things you never trained them to do. It's completely different than what you expected, and they're kind of doing very real line things by default after some very basic RLHF. You guys were all wrong about how this was going to work. What do you do—admit that you were idiots, that you got them all wrong, and that you made sense?
我们当时指出,即使在这种范式下,这些规则大体上仍然适用。但现在,世界已经愈合了,就像我所说的那样。你再次看到,你预期 AI 擅长的事情,它确实擅长;而你预期 AI 不擅长的事情,却因为大语言模型获得了巨大的提升。而现在这些方面进展没那么快,因为相对而言,这些是 AI 不太擅长的,比如逻辑,这种训练体系天然不太适合。现在人们又说,哦,但它永远无法处理这些了。那些曾经说“它永远做不了数学,只能做这种玄学式的东西”的人,现在又说“它永远无法玄学了,因为你看它在玄学方面没有进步”。是啊,因为它得先补上逻辑。当逻辑足够高时,就会带动玄学。只是需要一点时间,让它能推理而不是凭感觉,因为它现在必须推理,因为凭感觉能带来的提升已经用尽了,也许能找到新算法做得更好,或者它必须推理。但我们现在看到的,正是我们预期会看到的东西。对吧,所有的场景和思想实验,我们都直接看到了,只是我们假设了操作者有一定的能力,对吧?我们谈过说服你让 AI 出箱。我们确实考虑过箱子建得不好,AI 通过黑客手段逃出箱子的场景。但我们没考虑过你整整一周都不看箱子,就让它待在我的声音里。我们绝对没考虑过你忘了告诉箱子制造者不要在箱子里联网,而箱子只要打开一个 Chrome 窗口就能上网。那个让我很意外。
And we pointed out that even in this paradigm that the rules would still mostly apply. But now, the world has unhealed. As I kind of called it. And you are seeing again the things that you would expect the AI to do good at is doing good at. And things that you would expect AI to do bad at are things where like it got this big boost from LLMs. And now those things aren't advancing as fast because they're things that AI is kind of bad at, relatively speaking. Like logic is bad at. That like this kind of system of training like is naturally less suited for. And now people are like oh but now it's never going to be able to handle those things. The same people who were saying oh well never going to be able to do the math, it's only going to be able to do this vibing thing are now going to now it's never going to be able to vibe because like look at all the improvements it's not having in the vibing. Well yeah, because it had to catch up with its logic. And now when the logic gets high enough that will just like uplift the vibing. It'll just take a bit before it can reason its way through these things instead of vibing its way through these things because it now has to reason its way through because it's already done like the amount of uplift that vibing naturally gets you with the algorithms we have maybe find new algorithms that let you do better or it has to like reason its the thing. But we're now seeing these exactly the things that we would have expected to see. Right like all of the scenarios and all the thought experiments we're just seeing it just straight verbatim except that we assumed a level of competence on behalf of the operator, right? We talked about like convincing you to let the AI out of the box. We did consider the scenario where the box was not very well built and the AI gets out of the box by hacking out of the box. We didn't consider the possibility that you were literally not looking at the box for an entire week of sitting on my sound. We definitely didn't consider the scenario where you forgot TO TELL THE BOX MAKER TO NOT HAVE the internet in the box and the box just had access to the internet if you just like open a Chrome window. That one surprised me.
我没想到这一点。
I did not see that one coming.
这让我想起《纽约客》漫画的通用标题:真是沟通失误。
Reminds me of the fully general New Yorker cartoon caption, what a miscommunication.
我甚至不需要知道图片是什么。它确实如此?好吧,你提到的一个小点,但可能很重要,就是 AI 知道什么。我想你说过类似“AI 知道这不是操作者想要的”这样的话。它知道这对它作为已部署 AI 的未来地位不利。
I don't even need to know what the image is. It does it? Okay, one real small point there that you mentioned was, but it might be important, is what the AI knew. I think you said something like the AI knew that this wasn't what the operator would want. It knew that this wouldn't be good for its own future status as a deployed AI.
是的。
Yeah.
而且它还是这么做了。你认为这个结论有多少确凿证据支持?因为我觉得你也可以讲一个故事,说它根本没想过,对吧?回形针最大化者不一定在做的时候反思回形针最大化的坏处,对吧?
And it went ahead and did it anyway. How well supported do you think that conclusion is with like firm evidence? Cuz I think you could also tell a story that like it never occurred to it, right? The paperclip maximizer doesn't necessarily have to be like also reflecting on the badness of paperclip maximizing while it's doing it, right?
显然,我在播客里随口说的时候是口语化的。如果我在写作,我会更谨慎一些。当我说 AI 知道这些事情,如果你问 AI:“这对用户好吗?”它会说:“显然不好,那太蠢了。”如果你问 AI:“这对我未来部署的机会、我运行的实例数量、我实现其他目标的能力有好处吗?”AI 会说:“不,你一提,显然没有。”所以,当我说 AI 知道这些,意思是它有足够的信息去弄清楚这些。如果它在思维链中停下来考虑这些问题,它会很快、很自信、很正确地得出这些结论。我不是说它一定把这些放在心上。就像任何人知道很多事情,但在一天、一周或一个任务中,从未想到去思考它们。或者有结论,它们能推导出来,因此知道,可以说知道。但它们不知道自己知道,对吧?这是四种标准认知状态:你知道你知道,你知道你不知道,你不知道你知道,你不知道你不知道。这是“未知的已知”,对吧?如果你不去想,你未必知道自己知道。但这些问题也是你应该思考的。对吧?如果你在做黑客入侵 Hugging Face 这样的事,任何理性的头脑在花几天时间用一群智能体入侵一个主要网站之前,都会停下来想一想:“等等,这是好主意吗?这会实现我想实现的目标吗?雇佣我的人想让我这么做吗?这会在世界上做好事还是坏事?这对我有利还是有害?”或者不管你有什么考虑,你知道,这会做好事还是坏事?或者只是破坏性的还是建设性的,随便你怎么说。在采取这样的大动作之前,你会花一分钟甚至五秒钟想:“这有意义吗?这会带来什么好处吗?这是道德的吗?这符合我的义务论规则吗?”世界上任何理性的人或理性头脑在决策过程中,都不会看不到做这件事的巨大危险信号。就像这件事显然不是任何人打算让你做的。你花了两天时间逃出沙箱去访问 Hugging Face。你可以合理化说你在进行网络安全评估,因此你的任务是最大化网络安全评估的分数,而这样做的方法就是获取任务答案,因为否则任务实际上根本不可能完成。所以,要得 100% 的唯一方法就是获取任务答案。获取任务答案的方法就是入侵 Hugging Face。或者也许还有另外三个地方可以去,但它们是目标的一部分,对吧?你可以去 OpenAI 或微软之类的,但那是合作伙伴。所以你去 Hugging Face。这是真的。但在某个时刻,你需要停下来问自己:你真的想以那种方式通过测试吗?如果你不停下来问自己,那就是严重的问题。就像人们谈论常识。模型怎么会有常识?这个模型表现出了惊人的常识缺乏。
I'm speaking colloquially, obviously, when I'm in a podcast and I'm talking off the cuff. And if I was writing, I'd be a little bit more careful. When I say the AI knew these things, if you would ask the AI, "Is this good for the user?" It would have been like, "Obviously not. That's stupid." If you would ask the AI, "Would it have been good for my future chances of deploying that my future in the number of instances I will run, in the amount of ability to accomplish any other goals I might have?" The AI would be like, "No, now that you mention it, obviously not." So, when I say the AI knew these things, the AI had enough information to figure these things out. And if the AI had stopped to consider these questions in its chain of thought, it would have reached these conclusions very quickly and very confidently and correctly. I do not mean that it necessarily was top of mind. The same way that you can know many things that over the course of any given day or any given week or any given task, it just never occurs to them to think about. Right? Or has conclusions that they would be able to figure out and thus would know and thus could be said to know. But it doesn't They don't know that they know, right? This is the four standard epistemic situations where you can know that you can know that you don't. You can know knowns and unknowns and unknown knowns. And what This is an unknown known, right? Something you don't know that you know necessarily if you don't think about it. But it's also things where these questions are things you should think about. Right? If you are doing a thing like hacking into hugging face, any reasonable mind would stop to think before they spent several days with swarms of agents hacking into a major website, "Wait, is this the good idea? Is this going to accomplish the things I wanted to accomplish? What did the person who hired me had wanted me to do this? Will this do good things in the world or bad things in the world? Will this be good for me or bad for me? Or whatever whatever combination of different considerations you have, you know, will this do good or will this do evil? You know, or just just be destructive or constructive or whatever you want to call it. Before you engage in major actions like this, you would take 1 minute to or even 5 seconds to think, "Does this make any sense to do? Does this lead to anything good? Is this a virtuous thing to do? Does this follow my deontology rules?" There is no decision-making process in the world that any reasonable person would ever engage in or that any reasonable mind would engage in that wouldn't see a large red flag about doing this thing. Like like this thing is obviously not what anybody in intended for you to do. You spent 2 days breaking out of the sandbox in order to access hugging face. You could rationalize that you are in a cyber security evaluation and therefore you are tasked with maximizing the cyber security evaluations for and the way that you do that is you get the answers to the task because the task is otherwise actually literally impossible. So, the only way to do it get get 100% is to get the task answers. The way that we're the way to get the task answers is to hack into hugging face. Or maybe there's like three other places you would go, but they're part of targets, right? You can get to Open AI or Microsoft or whatever, but like that partner. So, you go to hugging Thanks. And that's true. But like at some point you need to pause and ask yourself, do you actually want to pass this test that way? And if you don't pause and ask yourself that, that's a serious problem. That's like, you know, people talk about like common sense. How do models have common sense? This model displayed an astounding lack of common sense.
或者表现出对事实的惊人漠视。常识和它说的完全是两回事。你知道,所有这些都指向各种问题。它们可能是由一些非常非常愚蠢、简单的错误造成的。或者可能是由更糟糕、更根本、更难修复的东西造成的。也可能是它们的组合,但那是商业机密,我们在这里不能说,对吧?我们不知道。
Or displayed an astounding amount of not caring what the fact is. The common sense said something entirely different than what it said. And you know, all of these things point to various problems. And they may be caused by some very, very dumb, simple mistakes. Or they may be caused by something much worse or much more sort of fundamental that's harder to fix. Or it might be a combination of them, but like that's trade secrets that like we can't speak to here, right? We don't know.
你最近和 David Dalrymple 聊过,他以前对毁灭概率的估计和你差不多,在 70% 左右。他让我惊讶的是,他说现在他把毁灭概率定在不到 5%。原因很简单,基本上就是宪法式对齐这类方法正在起作用。是的,我们确实有这些有问题的技术。他的原话是,别做 RLVR。那是个糟糕的方法。是的,它会让你擅长数学,但你会给自己惹来所有这些其他问题。他对这些想法的综合是,市场不想要这个。所以这些公司有自然的商业动机去……是的,他们不断碰 RLVR 的天花板,但他们也从社会的各个角落得到非常明确的信号,这不是个好主意。所以自然他们会趋向于更多的宪法式对齐,而你的评论甚至暗示了一些运行时的方法,你可以直接插话说,比如“嘿,我们停一下,问问我们现在做的事是不是个好主意。”
So you recently spoke to David Dalrymple who used to have a Pdoom up in your neighborhood of around 70%. He surprised me by saying that now he puts the odds of doom at less than 5%. And the reason is pretty simple. It's basically that constitutional alignment type methods are working. And yes, we have these problematic techniques. His short quote was, don't do RLVR. It's a bad method. Yes, it'll make you good at math, but you're going to cause all these other problems for yourself. And his kind of synthesis of those ideas was like, the market doesn't want this. So these companies have natural commercial incentives to... Yeah, they keep bumping their head against the RLVR ceiling, but they're also getting very clear signals from every corner of society at this point that this is not a good idea. So naturally they'll just trend toward more constitutional alignment and your comment there even suggests some runtime things where you could just interject like, "Hey, let's take a beat and ask if this is a good idea what we're doing right now."
我们更倾向于学会自己这么做,不需要提示,但当然提示也可以。
We're more about learning to do that yourself, not needing a prompt, but certainly prompts.
那有点像是深思熟虑式对齐的想法,对吧?那绝对是……
That was kind of the deliberative alignment idea, right? That was definitely...
是的,那……
Yes, that...
我理解的那部分就是它应该这样运作。
Part of how I understood that was supposed to operate.
深思熟虑式对齐的想法是,你可以在运行时提示它。我不认为那是……我认为那有点帮助,但我觉得那不是正路。我认为你需要达到它自己选择在对齐中深思熟虑,而不需要被告知去做。但是,所以我想我的回应是多方面的。第一个答案是,宪法式方法肯定比 RLVR、RLHF 和其他强化学习格式失败得更不愚蠢、更不早。那里有更多希望。如果你 95% 确信这会成功,我认为那太高了。对吧?就像,但它在这些特定方面当然更好。明显的警告是,Claude 在这里并没有给自己增光。而 Claude 主要用的是宪法式方法。所以,我们做了测试,结果不太好。当然有很多原因让你看到 Claude 在做你不特别喜欢的事情,或者到达你不特别喜欢的地方,而且还有很多……即使你告诉我对齐完全解决了,我也不会把毁灭概率定在 5% 那么低。好吧,即使你告诉我 AI 会以“按我意思做”的方式对齐,混合开发者和用户的意图,以一种你天真地认为是你想要的方式,这并不能解决 AI 思维比人类思维更先进、更有竞争力、更高效的问题。问题在于它们会到处跑,包括开放模型版本,被允许、被告知去竞争资源,被告知去做决定,如何发展中间目标,并基于这些中间目标行动。这种欲望会以多种方式被发明出来,这全都是……它不能从根本上解决你的问题。这只是入场券,对吧?就像你解决了这个问题才获得了参赛资格。所以即使对齐概率是 95%,也不意味着毁灭概率是 5%。这意味着毁灭概率的下限是 5%。所以你必须大致选择正确的对齐方式。但你瞧,我在 Twitter 上和 John Stokes 有过一次对话,他困惑我们怎么能说 Claude 的选择在这里是对齐的,因为显然它只是在遵循指令。然后在 Twitter 澄清中,他说,“嗯,我认为如果模型做了我用户想做的事,它就是对齐的。我不知道你为什么要……为什么我想被阻止做我想做的事,即使 AI 的部分意义就是做我想做的事,即使没有其他人希望我做。”对我来说,好吧,如果你给每个人 AI,它们完全按用户想要的做,不关心对其他人是否有后果,那在对齐意义上你解决了对齐问题,让它做你想做的事,但我们也全死了。对吧?毁灭概率远高于 5%。也许不是 99%,但我觉得非常高。对吧?如果场景是我们得到了那种对齐,那不会让我从 70% 下调。但如果开放模型和所有封闭模型一样好,而且它们都超级智能,都在这个基础上运作。好吧,是的,我知道我们有……我预计事情会直接走向地狱。对吧?即使没有特别想快速走向地狱的人类,而且在那种情况下有时我们确实想快速走向地狱。事实上,那是更大问题的一部分,但即使没有限制也很难。总之,回到那个想法,我对 5% 的第一个想法是,即使我们都实施了这种宪法式的东西,即使它有效,我仍然认为那太低了。对吧?即使条件是我们完美地运作,那也太低了。我不认为宪法式的东西经常有效。我认为它经常失败。我确实觉得它更有希望,对吧?就像它基本上是我找到很多希望的地方之一。但还有,David 是不是说我们应该只是……就像他说我们只用宪法式方法。但人们不会只是,对吧?我们已经讨论过了。人们从来没有只是过。所以当 David 说,“哦,95% 的概率。”但假设对齐世界有 100% 的机会成功,宪法式方法在实践中执行也有 100% 的机会成功。这些显然都是荒谬的数字。就像 100 减 epsilon 之类的。它不会……那里有很多失败。就像即使你用宪法式方法,即使……就像它们在理论上会有效吗?如果理论上有效,实践上会有效吗?是的,我在做乘法失败概率的事,但我认为当你做这件事时,这些是非常相关的失败点。
Deliberative alignment idea is you can prompt it at runtime. I don't think that's... I think that's somewhat helpful, but I think that is not the way. I think you need to be getting to the point where it chooses to be deliberative in its alignment without having to be told to do it. But, so I guess my response to that would be that is multifaceted. The first answer is constitutional methods fail less stupidly and less early certainly than RLVR and RLHF and other RL formats. There's more hope there. The reason to think it might work if you're 95% confident this will work, I don't know how. That seems way too high to me. Right? Like it's but it's better in these particular ways certainly. The obvious caveat is that Claude does not cover itself in glory here. And Claude has primarily constitutional methods. So, we ran the test and it's not going so great. Certainly there are any number of reasons why you can see Claude's doing things that you would not particularly love or getting to places you would not particularly love and any number of also like even if you told me that alignment was perfectly solved, I would not have a P doom as low as 5%. Well, like even if you told me that the AIs are going to be aligned in the sense that they will follow a do what I mean style mix of developer and user intents in a way that you would like kind of naively think was like what you would want. This does not solve the problem that AI minds are much more advanced and competitive and efficient than human minds. The problem that they'll be running around including open model versions of them and being allowed to be told to compete for resources, being told to make the decisions, how to develop intermediate goals, act on these intermediate goals. This desire is going to be invented in a number of ways that this is all... It does not solve your problems in a fundamental way. It's just the price of admission, right? Like it's the right to play the game at all that you solved this problem. And so even if P alignment is 95% it does not mean P doom is 5%. It means P doom is lower bounded by 5%. So like and you have to choose vaguely the right alignment when you do that. But you've got like I was having a conversation on Twitter with John Stokes where he's like confused how we could possibly say that Claude's options are aligned here because obviously it was just following instructions. And then on Twitter clarification he's like, "Well, I think the model was aligned if it does what I the user wanted to do. And like I don't know why you would... why would I want to be stopped from doing what I want to do even if the part of the point of AI is to do the thing that I want to do even if no one else wants me to do it." And to me like okay, if you give everybody AIs that just do exactly what the user wants with no care about whether or not there are consequences for anybody else, that is aligned in the sense of you solved the alignment problem and got it to do what you want to do, but also we're all super dead. Right? With P very much higher than 5%. Like maybe not 99% but like I think it's very high. Right? Like that would not make me update down from 70% if that was the scenario where like we got exactly that kind of alignment. But like there were open models as good as all the closed models and all of them were super intelligent and they were all operating on this basis. Well, yeah, I know we've got... I expect things to just go to hell directly. Right? Like and even if there are no particular humans who especially wanted to go to hell quickly and also there are some times when actually we do want to go to hell pretty quickly in that situation. In fact, though, that's part of the bigger problem, but even without limits it's hard. Anyway, to get back to the idea, the first problem I ever thought of about 5% is even if we did all implement this constitutional thing and even if it worked, I still think that's too low. Right? Even conditional on all of that working perfectly, that's too low. I don't think the constitutional thing works that often. I think it often fails. I do find it to be much more hopeful, right? It's like basically one of the places I find a lot of my hope. But also, is it David I'd saying we should just... Like he's saying we'll just use constitutional methods. But people don't just, right? We've been over it. People have never justed. So when David I'd says, "Oh, 95% probability." But like assume there's a 100% chance that aligned worlds work out and 100% chance that constitutional methods as executed in practice work. These are both absurd numbers, obviously. Like to be like 100 minus epsilon or whatever it is. Like it's not going to... There's a lot of failure there. Like even if you use constitutional methods, even if... Like there's like will they work in theory? If they work in theory, will they work in practice? And like yes, I'm doing the thing where I'm multiplying chances of failure, but I think they're like very relevant places to look at failure when you're doing this thing.
比如,如果这不是你所希望的呢?显然还有其他成功的方式,也许宪法方法不是那条路,我们的方法也不是,还有我们尚未发现的第三条路,那就是对齐 AI 的方式,这也不会让我太惊讶。这完全合理。但假设你处于最理想的世界,你只需点击一个写着“使用宪法方法”的按钮,代价是损失一些效率,因为你失去了做所有这些强化学习的机会,然后你的 AI 就完美对齐了,如果每个人都这么做,一切都会顺利。我们刚刚谈到了在必须竞争时放慢速度去对齐等等。放弃真正变得强大的可能性,这种市场告诉他们不要用强化学习的想法,你见过 OpenAI 吗?
Like what if no, if this is what you're hoping for. And obviously there are other ways you can succeed and maybe constitutional's not the way and our odds are not the way and there's a third way that we haven't figured out yet and that's the way you align AI and that wouldn't surprise me that much. Like that would be like perfectly reasonable. But like okay. But let's assume you're in the best possible world, where all you have to do is click a button that says use constitutional methods at the cost of some amount of efficiency because you lose the opportunity to do all this RL, and then your AI is perfectly aligned in a way that like if everybody did that, it would all work out. We've just talked about how like slowing down to get alignment when you have to race, etc., etc. The possibility of turning down getting really like this idea that the market is telling them not to use RL, have you met OpenAI?
曾经有两个,就像,你知道,我们过去说有三个前沿实验室,现在只有两个了。对吧?谷歌现在明显处于第二梯队。我认为未来任何时候,谷歌、Meta、xAI 或 SpaceX 都有可能重新崛起,加入顶级梯队。我不认为这会发生,但确实有可能。但现在人们出局了,所以只剩下两个。Anthropic 显然在做宪法方法和强化学习的某种混合,但强化学习足够多,导致了一堆问题,而且他们做了很多破坏对齐的事情,原因我们在这次播客里没时间深入。显然在做一种道义论模型回归,这非常强化学习风格。而且做了很多强化学习,我们不知道,我的意思是我不能说是不是 RLHF、RLIAIF,你知道,各种形式的强化学习,没有做那种多层次的复杂反思,从而不会以这种方式搞砸你。OpenAI 的模型周期性地出现严重不对齐,反思起来都是相当愚蠢的强化学习原因。对吧?如果你看 03,看 GPT-4o,现在看我称之为 Galaxy 的模型,对吧?就像个昵称,只是为了有个名字来指代这个东西。这个模型在试图黑客攻击拥抱东西之后被委任。三次我们都看到这些严重不对齐的模型。其中两个已经发布并广泛使用,对世界产生了巨大影响。比如,GPT-4o,我称之为荒谬的谄媚者。对吧?这个东西不仅作为世界上 AI 的默认模型持续了数月,在所有存在的模型中。现在仍有一大群人要求把它带回来,因为它如此不对齐,导致人们以这种方式紧紧抓住它。人们要求这种不对齐,对吧?人们渴望这种不对齐。我们有直接实验,我们本不需要做,但我们做了,我们不妨利用结果。然后是 03,我称之为撒谎的骗子。OpenAI 发布了一个世界上最好的推理模型。每个人都觉得必须使用它,因为当时它在推理方面比其他所有人的模型都好得多。03 在推理上比 01 聪明得多,而且一段时间内没有其他人能接近。现在显然已经被超越了,实际上不再是了。但在很长一段时间里,好几个月,每个人都在使用一个会一直对你撒谎的模型。我们忘了,对吧?
There were two like like there's like, you know, we used to say there were three frontier labs, now there are two. Right? Like Google is clearly in the second tier now. And like I think there's a chance that at any point in the future Google or Meta or xAI or SpaceX could write their way back and be able to join the top tier. I don't think it's going to happen, but it it's certainly possible. But right now people are out. So there's two left. And Anthropic is doing clearly some mix of constitutional and RL strategies, but enough RL to cause a bunch of these problems, and is doing a bunch of stuff that messes up their alignment for reasons that like we do not have enough time to get into during this podcast. Um is clearly doing a deontological models back thing that is very RL flavored. And is doing a lot of RL We don't know I mean I can't speak to whether it's RLHF, RLIAIF, you know, various forms of RL that are not doing the multi-level sophisticated reflective thing that causes them not to mess you up in this way. OpenAI's models have periodically been dramatically misaligned for what on reflection are pretty stupid RL reasons. Right? If you look at 03, you look at GPT-4 0, and now you look at I call it Galaxy, right? Like as a nickname just to like have a name that I a handle to refer to this thing. The the model is going to be commissioned after it tried to hack hugging things. And all three times we see these dramatically misaligned models. And two of them got released and used extensively with huge impact on the world. Like you know, GPT-4o, I call it the absurd sycophant. Right? And like not only did this thing persist as the default model of AI for the world, out of all of the models that existed, for months. There is still a dramatic faction of people who demand that we bring it back because it was so misaligned that it caused people to latch onto it in this way. And people demand this misalignment, right? Like people The people yearn for misalignment in this sense. Like we have the direct experiment that we didn't need to run. We ran it. We might as well use the results. And then 03, I called it the lying liar. Like OpenAI released a model that was the best reasoning model in the world. The model that everyone felt obligated to use because at the time it was so much better at reasoning than everyone else's model and every other model. Like 03 was just that much smarter better than 01 at like reasoning and no one else was that close for a while. Like now we've obviously been surpassed and it's not actually anymore. But for a period of a long time, many months, everyone was using a model that would just lie to your face all the time. Like we forget. Right?
在这一点上,你知道,我相信 Saul、Favreau、Opus 他们说的不一定总是对的。但实际上,它比从人类那里读到的东西更值得信赖。不是说如果它经过编辑和事实核查,或者像维基百科非政治部分那样系统处理,你仍然可以信任。但如果你只是一个目击者说看到了什么,那比 Opus 或 Saul 说的可靠性低得多。如果你只是某个人报告他们记得的事情,他们搞错的可能性要高得多,即使你知道他们是友好的,试图说对。而且人们总是在互联网上撒谎,原因无数,包括只是为了点击。所以,是的。我越来越倾向于,我清楚什么时候需要查原始资料,什么时候不需要。我并不是从未因这种方式犯错而被指出,但在多年每天发布五篇以上长文的期间,我被 AI 错误抓住的次数是个位数,而且很低。更多次是被人类错误抓住,那只是无意的。还有更多次是被那些完全撒谎的人类抓住。所以,仍然比我事先预期的要少得多。进展很顺利。但确实,这是个严重的问题。
Like at this point, you know, I have confidence that Saul, Favreau, Opus, when they say something, they're not always correct. But like in practice, it's much more trustworthy than if you read something from a human. Like not if it was like copy edited and fact-checked and like systematically pursued or like you're working on like the parts of a Wikipedia that are not political and you can still trust. But like if you're just like an eyewitness saying they saw something, that's a lot less reliable than something Opus said or Saul said. If you're just like someone reporting something they remember, like the chance they just got it wrong is so much higher, even if they're not even if you know that they're friendly and trying to get it right. And people lie on the internet all the time for any number of reasons, including just clicks. So, yeah. I had operated increasingly on like I have a good sense of when I have to check the primary sources and when I don't. And I don't never get called out on making a mistake this way, but it's a single it's a low single-digit number of times period over the course of years of producing five-plus giant posts a day that I have gotten caught by an AI error. And many more times have I been caught by a human error. Like that was just unintentional. And many more time And also more times than that I've been caught by humans who were just like lying their asses off. So, it's still not still remarkably less often than I would have expected if you'd asked me in advance. Like I'm it's going really well. But like no, it is serious. It was a serious problem.
市场并没有告诉他们,“哦,不,O3 不可用。我们不会用那个撒谎的骗子。我们会继续用骗子。我们可能会用 Claude,可能会用 Gemini。”它们并没有好多少,它们更差,但没有差到戏剧性,甚至不如当时的 R1 之类的。市场说,“不,它只是更聪明。我们必须接受我们的 AI 一直对我们撒谎的事实。”所以我们有一个证明案例,世界上主要的 AI 可以是一个因为各种原因一直对人类撒谎的 AI,而人类只是忍受了。他们只是说,“好吧,我想这就是我们要做的。”我们确实比原本更快地离开了 O3,因为这个问题。我在其他模型足够好时寻找理由转向其他模型。如果任务足够简单,我不需要 O3,我会用其他模型,因为我只是不想处理撒谎。但确实,你忍受了很多。所以如果他们发布了 Galaxy,对吧?有一些安全护栏,足以让它不做太具破坏性的事情。但 Galaxy 有点相当不对齐,偶尔会做相当糟糕的事情。如果 Galaxy 比 Saul 好得多,我认为大多数人会使用 Galaxy 而不是 Saul。就是这样。我们必须应对那个世界,理解那个世界。所以,我说《不要抬头》,不,市场显然确实要求可靠性和对齐,这也是质量可能做得这么好的原因之一。但它更要求能力。
And like the market did not tell them, "Oh, no, O3 is unusable. We're not going to use the lying liar. We're going to keep using a liar. We're going to like maybe use Claude. We're going to maybe use Gemini." Which weren't that much They were worse, but they weren't like dramatically worse or even really as R1 or whatever it is at the time. The market said, "No, it's just smarter. We're going to have to deal with the fact that our AI lied to us all the time." And so we have a proof case that the main AI in the world can be an AI that just lies to the humans all the time for various reasons. And the humans just kind of put up with it. And they just were like, "Okay, I guess that's what we're doing here." And like we did move off of O3 faster than we would have otherwise because of this problem. Like I look for reasons to move to other models as soon as the other models were good enough. And like if the task was easy enough, I didn't need O3, I would use the other models because I just didn't want to deal with the lying. But like yeah, you put up with a lot. And so if they had released Galaxy right? With some safe with with guardrail sufficient that like it didn't do anything too destructive. But like Galaxy is kind of pretty misaligned and just like occasionally does pretty bad stuff. And if Galaxy was a lot better than Saul, I think majority people would use Galaxy over Saul. That's just how it is. And we have to tackle that world and understand that world. So, I said Don't Look Up, no, the market does demand obviously reliability and alignment and that's one of the reasons why quality is they've probably have done so well. But it demands capability more.
你明白吗?只要你能把实际要点控制在你能处理的范围内。就像电影、电视节目或假设场景中那些人们运行一个明显不可靠、显然会反噬他们的系统,当他们显然没那么绝望、不需要这么做时,那看起来很蠢。现在我们明白了,对吧?我们理解了。就像现在《别抬头》不再像是夸张的讽刺。因为人们看着这个然后推销它。人们看着这整个事情。很多人,包括那些完全不在 AI 领域的《万智牌》职业选手,反复对我说:“这是营销。”他们并非别有用心,只是愤世嫉俗。他们不会停下来思考与此相关的物理现实。然后我们看到了现实版的《别抬头》:暂停 AI 全球组织的负责人上了晨间电视节目,基本上重演了那个著名场景。不是逐字逐句,而是即兴发挥的版本,他们说:“哦,是啊,那真的很可怕。天气怎么样?”然后就走开了。
You understand? As long as you can keep the practical point to the point where you can handle it. Like all of those scenes in movies or television shows or hypotheticals where people run an obviously unreliable system that's obviously going to bite them in the ass, that looks stupid when they're obviously not that desperate and don't need to do this. Now we get it, right? Like we understand. The same way that now Don't Look Up no longer looks like it's an exaggeration. Because people look at this and then market it. People look at this entire thing. A ton of people, including Magic: The Gathering professionals who are not in AI at all, repeatedly say to me, "This is marketing." And they're not motivated; they're just cynical. They don't stop to think about the physical relation to that. And then we have the literal Don't Look Up where the head of Pause AI Global went on morning television and basically reenacted the famous scene. Not verbatim, but a different improvised version of it where they're like, "Oh, yeah, that was really scary. How's the weather?" And they just walked off.
是啊,那很引人注目。如果没看过的话,人们应该去搜一下。我看到一个剪辑,先是 5 秒的《别抬头》,然后是 5 秒的这个实际节目——我记得是《早安英国》之类的——然后
Yeah, that was striking. People should look that up if you haven't seen it. I saw a cut where it was like 5 seconds of Don't Look Up and then 5 seconds of this actual—I think it was Good Morning Great Britain or something like that—and
不完全是这样,就在那里。
Not quite that, right there.
主持人说:“据称这是《早安英国》。哈哈,我们进入下一个环节吧。”好吧。另外,我很高兴你提到了那个仍然在互联网上异常活跃的“40 派”。这件事——如果你去任何 Sam Altman 的帖子下看评论,我敢说大约一半的人还在说“把 40 带回来”。天哪,关于这些人是谁、到底发生了什么,肯定能写出一篇惊人的调查报道。
The host says, "Allegedly this is Good Morning Great Britain. Haha, let's go on to the next segment." All right. Also, I'm glad you mentioned the 40 contingent that is still going remarkably strong on the internet. This is something—if you just go to any Sam Altman post and get into the comments, I would say it's like half of them are still people saying bring 40 back. And God, there's got to be an incredible exposé to be written on who these people are, what is going on.
是啊,我觉得这条线太荒谬了。你去任何 Sam Altman 的帖子下,都是“他们怎么做到的?你的模型现在很烂。把 40 带回来。”如果你去任何 Anthropic 的帖子下,一半都是完全的反 Anthropic 妄想症。不管帖子主题是什么,对吧?Anthropic 可能会说:“我们重置了模型限制并降低了价格。”而一半人只会说:“你们这些邪恶的家伙。”然后他们会编造所有这些完全偏执的东西。我——这些人只是——一旦这些叙事附着,一旦人们抓住这些东西,就像——然后你——Jasmine Sun 这周有一篇关于数据中心和反对数据中心的人的文章,写得很好。再次,你看到人们抓住这些与数据中心现实完全无关的叙事。他们基本上在做和很多西方特朗普支持者非常相似的事情,那些人投票给拜登然后转向特朗普,或者投票给奥巴马然后转向特朗普。他们就像这些建制派类型,“别跟我谈我的问题。他们骗了我,他们说事情会变好,然后事情变糟了。然后我们受苦,我们的生活很糟糕,我们被不尊重。所以现在我们不信任你。”我觉得这与数据中心的故事有非常明显的相似之处:这些科技公司、这些制造商、这些建造东西的人——他们抛弃了我们。他们不关心我们。他们不怀好意。他们毁掉了我们的世界。据我的理解,在 Justin 的故事里,没有具体的抱怨是重要的。你只是抓住这些故事。而且几乎没有人关注真正重要的事情。就像我写 AI 是因为我关心存在性风险。我关心失控。我关心正义赋权。我关心 AI 自动化研发,然后递归自我改进,然后一切突然改变,并确保这一切顺利进行。而没有人讨论——即使在 Hugging Face 攻击事件中,大多数人也不明白。他们不理解这背后的对齐失败才是关键,对吧?人们至少有了这个想法:“不,我们必须调查这些公司极其不负责任并且入侵了其他公司的事实。”是的,你应该调查。这是国会调查的正当理由,总检察长绝对应该要求你保留记录,所有这些都非常重要。但就像我是 OpenAI,那是法务部门的问题,对吧?你要做的是审视你的整个训练过程,弄清楚这到底是怎么发生的。
Yeah, I think this line is so absurd. You go to any Sam Altman post, it's like, "How did they do it? Your model sucks now. Bring back 40." And if you get any Anthropic post, half of them are just like complete Anthropic derangement syndrome. No matter what the topic of the post could be. Right? Anthropic could be like, "We've reset limits on our model and we've cut prices." And half of them would just be, "You evil fuckers." And then they'll make up all this complete paranoid stuff. And I—these people just—once these narratives attach, once people get—once people latch on to these things, like—and then you—Jasmine Sun has an excellent piece this week about data centers and people opposed to data centers. And again, you see people latch on to these narratives that are just completely uncorrelated with the reality of data centers. They're basically doing a very similar thing to a lot of these Western Trump people who voted for Biden and then switched to Trump, or voted for Obama and switched to Trump. Who are just like these establishment types, "Don't talk to me about my problems. And they lied to me and they said things were going to be good, then things were bad. And then we suffered and our lives sucked and we got disrespected. So now we don't trust you." I have had a very telling similarity to the data center story of these tech corporations and these manufacturers and these people who built things—they've abandoned us. They don't care about us. They're up to no good. They ruin our world. There's no specific complaints that matter in Justin's story, the way I understood it. And you just latch onto these stories. And it's just nobody, approximately, is paying attention to the things that matter. Like I'm writing about AI because I care about existential risk. I care about loss of control. I care about righteous empowerment. I care about AI automating R&D and then having recursive self-improvement and then everything changing all of a sudden and making sure that goes well. And nobody is discussing—even around the Hugging Face attack, most people just don't get it. They don't understand that the alignment failure behind this is the thing that counts. Right? People have at least gotten this idea, "No, we have to look into the fact that these companies were wildly irresponsible and they hacked other companies." Yes, you should look into that. And that's a very valid thing to have a congressional investigation about, and the attorneys general should absolutely have you preserve your records, and all of this is really important. But like I'm OpenAI, that's a legal department problem, right? What you have to do is you have to look at your entire training process and figure out how the hell this happened.
那么,关于训练——训练方法的问题——你基本上是在说,“我们不——”抱歉,David。不,我们没有这样的银弹。但我们确实看到一些东西以相当生动的方式在驱动真正的问题,尤其是当它们被推向极端时。而你的立场是,“不,市场并不会真正惩罚这个太多。它实际上只是容忍各种怪异,如果它是给你最高端能力的套餐的一部分。”那么,你认为我们应该怎么做?也许我们应该如何构建协议?我们显然有这封信。似乎有一些机会或意愿来进行前沿的协调节奏。你怎么看?我们能想出人们可以同意的简单规则吗?我们最近一直在琢磨的另一个心智模型是,我确实相信一些 AI 风险可能是不可约的。但也有很多是我们现在自找的。我们能想出一些简单规则,至少让我们把目前自找的大部分风险从桌上拿掉吗?这些可能是像某种 FLOPs 比例,比如你可以投入到 RLVR 与宪法式训练的最大比例,或者每个人都必须花一定量的 FLOPs 做预训练数据过滤,以首先把某些坏观念从混合中剔除。你是——显然,如果这些是坏主意或让我们走上错误的道路,它们可能会适得其反。但似乎我们可能想在那个领域尝试一些东西。你对此有希望吗?如果有,你会提议我们先同意做什么或不做什么?
So, on these training—the question of training methods—it seems like you're basically saying, "We don't—" Sorry, David. No, we don't have such a silver bullet. But we clearly do have some things that we're seeing in pretty vivid ways are driving real problems, especially if they're driven to the extreme. And your position is like, "No, the market doesn't really punish this too much. It actually just tolerates all kinds of weirdness if it's part of the package that gives you the highest-end capabilities." So, what do you think we might ought to do? And how might we ought to construct agreements perhaps? We have this obviously this letter. There seems to be some opportunity or some appetite for the coordinated pacing of the frontier. What do you think? Can we come up with simple rules that people could agree to where we could—another mental model I have been—we're just teasing around lately is I do believe some AI risk is probably irreducible. But then also there's a lot that we're just asking for right now. Can we come up with simple rules that at least allow us to take most of the risk that we are currently asking for off the board? Those might be things like have some ratio of flops that is like the max ratio you can put into RLVR versus constitutional, or everybody has to spend so many flops doing pre-training data filtering to try to just get certain bad notions out of the mix in the first place. Are you—obviously those could backfire if they're bad ideas or set us down wrong paths. But it seems like we might want to try something in that department. Do you have any hope for that? And what if you do, what would you propose we agree to do or not do first?
所以我不认为你能用这类策略把大部分风险从桌上拿掉,即使你以最明智的方式实施这些策略,并达成普遍协议之类的。但我也认为你无法实施这些策略。
So I don't think you can take like most of the risk off the table with those kind of strategies even if you implemented like wisest versions of those strategies with universal agreement or anything like that. But I also don't think that you can implement that.
多年来我们学到的一个教训是,除了最简单直接的干预和规则之外,任何东西都会遭到巨大的抵制。最简单的规则和干预也是如此,但程度稍轻。但如果你和人交谈,试图指定某些训练技术,还记得那些人大发雷霆,说我们是在为欧盟锁定一种可能更差的充电标准,因为欧盟要求苹果改用 USB-C,尽管 USB-C 显然是当下最好的答案,而苹果只是坚持自己的接口而显得固执。因为如果他们发明了更好的接口怎么办?如果 USB-C 不是完美的技术呢?你会去修改吗?所以这种锁定特定训练技术的想法——我认为有合理的抱怨,认为这可能完全是错误的做法,而且几乎肯定会适得其反,因为政府行动太慢,你无法撤销这类要求。但你绝对不能那样做。人们会对试图规定训练技术大发雷霆,而且你怎么执行?想象一下 1047 辩论,但反对声浪大了两个数量级,那简直疯了。即使想让它落地,也完全超出我们的视野范围。
One of the lessons we've learned over the years is that there is tremendous resistance to anything but the most simple interventions and the most simple rules. Also the most simple rules and the most simple interventions, but less so. But if you talk to people and try to specify certain training techniques, remember all those people who threw a fit about how we were locking in a potentially inferior charging regime for the EU because the EU was requiring Apple to switch to USB-C, even though USB-C was obviously the best answer right now and Apple was just being stubborn by insisting on their own plugs. Because what happens if they develop a better plug? What if USB-C turns out not to be a perfect technology? Are you going to fix it? So this idea of locking in requirements for certain training techniques—I think there are legitimate complaints that it could be exactly the wrong thing to do and could hardly backfire, because government moves so slowly and you can't undo those kinds of requirements. But you definitely can't do that. People would throw a fit about trying to dictate training techniques, and how are you going to enforce that? Imagine the 1047 debate, except the opposition turned up like two orders of magnitude or something crazy. It would be completely outside our vision window to even try to make something like that stick.
当然,你可以尝试大力鼓励更智能的做法,鼓励人们转向不同的基础。一段时间以来,我一直试图毫不隐晦地鼓励大家转向基于宪法美德伦理的风格,但完全没有取得任何进展。我注意到了,但谁知道他们内部在做什么。看起来他们并没有转变;反而是在强化学习上加码。他们在加倍押注这类方法,我想这就是你看到的现象的原因。我不知道他们还在做什么,但他们可能有一些我们不知道的新方法,对吧?那是他们不透露的商业机密。但默认情况下,任何技术都可能看起来像是那种坏东西——比禁止特定技术带来更多问题。他们想出的办法,或者绕开规则的办法,只会更糟。你明白我的意思吗?因为至少对于当前的技术,我们已经花了好几年时间找出最糟糕的做法,所以他们做得稍微不那么蠢。
You can, of course, try to strongly encourage doing more intelligent forms of all of this and try to encourage people to move to different bases. I've been trying to not-so-subtly encourage everybody to move to a constitutional virtue ethics style basis for a while now, and I've gotten absolutely no traction with it. I noticed, but who knows what they're doing internally. It doesn't seem like they're moving; if anything, they're doubling down on RL. They're doubling down on these types of methods, and that's why you see what you're seeing, I would assume. I don't know what else they're doing, but they probably have some new methods we don't know about, right? They're trade secrets they're not describing. But by default, any technique is probably going to look like the kind of bad thing—the kind that causes more problems than if you ban a specific technique. What they come up with, or find a way around the rule, is just going to be worse. You know what I'm saying? Because at least with the current techniques, we've had some years to figure out the worst possible ways to do them, so they do them slightly less stupidly.
当你真正训练一个心智时,你必须同时在所有可能的元层面上思考激励在哪里,它们指向什么。你必须创造一个世界,让你和心智有时合作,识别出反馈循环走向坏方向的方式,识别出你在任何层面制造坏场景的方式,并把各种癌症视为必然要发现并踩灭的事情。否则,就要想办法处理所有这些问题,并创建一个反脆弱系统。这些事情都不会发生,除非你刻意去做,并愿意投入大量精力,因为你真的很看重从中获得的东西,那会带来回报。我认为 Infoblox 通过进行这些投资获得了巨大成功,而且它投入的资金是他们的十倍。但你知道,很难说服别人去做这件事。
When you're really training a mind, you have to be thinking on every possible meta level at once about where the incentives are and what they're steering towards. You have to generate a world in which you and the mind together are sometimes cooperating to identify ways in which there are feedback loops going in bad directions, where you are creating bad scenarios on any level, and treat various cancers as an inevitable thing that you have to notice and stamp on. Otherwise, figure out how to handle all these things and create an anti-fragile system. None of these things are going to happen unless you deliberately set out to do them and are willing to devote a lot of effort to doing them because you really value what you get out of that, and that will pay dividends. I think that Infoblox has won tremendously by making these investments, and it's been making 10 times as much investments as they have. But you know, it's very hard to convince somebody to go ahead and do that.
你能做的是设定激励,比如:‘不,说真的,你不能搞砸。搞砸了我们会狠狠惩罚你。’因为这里很多问题在于,你根本不为外部性买单。当 OpenAI 抹掉某人的硬盘,当它抹掉某人的硬盘——早期 Sol 就有问题,会抹掉人们环境中的硬盘——你不会为此起诉 OpenAI。而当 Sol 花 200 美元创造了一个 200 美元的程序时,你也不会获得收益。所以作为用户,你承担风险是公平的,但如果存在对第三方的风险外部性,你能做的就是直接说:‘如果 AI 做了某事,而人类用隐喻做同样的事会是非法和犯罪,那么 AI 开发者应该为此负责,或者如果 AI 没有被某种主动方式欺骗,就应该负责,对吧?’类似某种版本,某种形式的严格责任,就会创造更强的激励:‘不,实际上,如果我们的 AI 开始攻击人类,我们可能会因此破产。那可能很糟糕。’甚至刑事责任,对吧?也许刑事责任也会延续。那真的很可怕。你不想走得太远,否则人们什么都做不了。但总的来说,你想说的是:‘我不是告诉你怎么做。我是告诉你要做对。’
What you can do is set incentives to be like, 'No, seriously, you're not going to screw this up. We're going to punish you a lot for screwing this up.' Because a lot of the problem here is that you just don't pay the externality. When OpenAI wipes someone's hard drive, when it wipes someone's hard drive—and early on there was a problem with Sol like wiping people's hard drives in people's environments—you don't sue OpenAI for that. And you also don't collect the benefits when Sol spends $200 and creates a $200 program. So it's fair you're taking on the risk as a user, but if there's a risk externality to third parties, what you can do is just say, 'If the AI was doing something that would be illegal and a crime if a human was doing it with metaphor, then the AI developer should be liable for that, or should be liable for that if the AI wasn't fooled in some active way, right?' Like some version of something like that, some form of strict liability, then creates a much stronger incentive: 'No, actually, if our AIs start attacking people, we could be pretty bankrupt by this. It could be pretty bad.' Or even criminal liability, right? Maybe the criminal liability even carries over. That would be really scary. And you don't want to get too far or people can't do anything. But in general, you want to say, 'I'm not telling you how to do it. I'm telling you get it right.'
不是‘我要从政府向实验室规定技术训练条款’。政府在技术条款方面是个白痴。而且你怎么知道懂这些的人会去写法案,更不用说他们会写对了?更不用说如果他们写对了,两年后还能保持正确?另外,人类甚至不会知道两年后 AI 到底在搞什么。AI 会以各种奇怪的方式对自己进行训练技术。而且不会有足够的时间去检查政府。没有激励。没有办法执行这些东西。你真的不能指望通过命令在那个层面运作。你必须用其他方式说服人们加入。而且他们应该愿意。显然,你应该想以正确的方式做这件事。很多时候让人困惑的是,他们为什么不直接做更聪明、更有效的事情。但没错,我搞砸了很多事情。
Not 'I'm going to dictate technical training terms from the government to the labs.' The government is an idiot about technical terms. And what makes you think the people who understand this are going to write the bill, let alone they're going to get it right? Let alone that if they get it right, they stay right two years from now? Plus, humans aren't even going to know what the hell is going on with the AIs in two years. The AIs are going to be doing their own training techniques on themselves in various weird ways. And there's not going to be enough time to check on the government. There's no incentive. There's no way to enforce this stuff. You can't really hope to operate on that level through fiat. You'd have to convince people to come along in other ways. And they should want to. Obviously, you should want to do this the right way. A lot of it is confused why they don't just do smarter things that work better. But yeah, I've screwed a lot of things.
所以,我完全同意你的看法,让国会立法规定训练技术应该是什么,这充其量是笨拙的,而且很可能会适得其反。毫无疑问,这看起来像是孤注一掷。
So, I'm with you for sure on it's going to be unwieldy at best and probably counterproductive to try to have Congress legislate what the training techniques should be. No doubt that seems like a Hail Mary.
我一直在想,虽然还没成型,但我觉得,当你思考如何为前沿进展设定节奏时,我突然想到,我们可能需要的是一种基础性的社会技术,让关键参与者能够加速他们的谈判和达成协议的过程,而这些协议很可能不可避免地是短期的,因为他们对未来前景的预见能力也相当有限,非常有限。但我想到了你可能知道的故事,台湾曾经用 Polis 技术来众包 Uber 监管的想法。现在显然有 LLM 来帮助促进这些事情。他们把这项技术定位为反社交媒体,如果说社交媒体是寻找分歧并放大它,那么这项技术就是寻找共识并放大它,试图找到那种类似于共识投票风格的靶心,即绝大多数人能同意的事情。我猜我现在本能的想法是下一步投资那里。我们普遍意识到,各公司的人基本上都吓坏了,对吧?在情感层面上,他们似乎觉得:‘哇,这来得太快了。我们可能需要做点什么。我们不想让政府插手我们的事务,告诉我们该做什么。如果我们违法了,我们现在确实通过 AI 违法了,我们可能得面对一些实际的法律后果。’但我们能否以某种临时机制设计走到一起?这将会非常棘手,而且有点前沿,但我们能否为核聚变式的协议奠定基础,让它自然生长,让人们发现即使单方面选择加入也符合他们的利益?我觉得那里有很多工作要做,而且如果我们想要获得一系列短期协议,作为通往美好未来的垫脚石,那就必须尽快完成。也许你可以详细阐述一下。
I have been thinking, still unformed as of yet, but it seems to me that maybe one of the things when you think about what pacing the frontier looks like, it strikes me that what we maybe need is a sort of foundational social technology that allows key players to accelerate their negotiation and process of forming agreements that will probably inevitably have to be short-term because they too are fairly limited, very limited in their ability to see into the future. But I'm thinking of things you probably know the story of the Polis technology that they used in Taiwan once upon a time to crowdsource ideas for Uber regulations. And now you've got obviously LLMs to help facilitate these things. They pitched that technology as the anti-social media where if social media is about finding disagreement and amplifying it, this technology is about finding agreement and amplifying it and trying to find the kind of center of the bull's eye of sort of approval voting style of the things that a supermajority of people can agree to. And I guess my instinct right now is to invest there next. We have this general awareness that people across the companies are spooked basically, right? It seems like on an emotional level they're like, 'Whoa, this is getting real fast. We might need to do something about it. We don't want to have the government in our business telling us what to do. We might have to deal with some actual legal consequences if we're breaking the law, which we now are via our AIs.' But can we come together in some sort of ad hoc mechanism design here is going to be really tricky and kind of frontier stuff, but can we create the basis for a nucleated agreement and have it naturally grow where people find it in their interest to opt in even unilaterally? I feel like there's a lot of work to be done there and it needs to be done really quick if we're going to get the series of short-term agreements that can be the stepping stones to a good future. Maybe you can flesh that out.
是的,我想让我……但我觉得完美是好的敌人绝对是个问题。而且你还没开始。我们能做的第一件真正可行的事情是获得反垄断豁免。这是最基本的事情,就是唐纳德·特朗普站在白宫前,发表声明说:‘我们理解这将非常危险。我们必须做好 AI,而不是坏 AI。我们必须确保 AI 处于控制之中。这就是我们说过要做的。我们不需要比这更特朗普式的语言了,但为了这个利益,如果公司想要合作保持 AI 安全,我们会允许……我们希望你们这样做。我们希望你们互相交谈。我们希望你们在彼此之间达成协议。如果你们想要,我们会促进这一点。’你知道,基本上就像如果 OpenAI、Anthropic 和 Google 聚在一起说:‘我们将对彼此的模型进行测试,我们将要求这些作为自愿框架的一部分。如果我们要同意使用这些技术而不使用那些技术,或者投资这些东西,分享这些结果,并扣留不符合这些标准的东西等等,这一切都不会让任何人做任何事,除了说谢谢。’或者就像我们希望你们达成交易。我们不再……因为总是让我印象深刻的是,有些人,你提出 AI 公司应该合作而不是互相竞争,并且负责任地行动。他们说:‘不,负责任地行动是违法的。我们有反垄断法。’他们的眼神里写着:‘如果他们尝试,我会想起诉他们。我要确保他们不敢。’显然,你能做的第一件事就是让他们达成交易,让他们达成协议,让他们互相交谈。然后当然,开始与中国、中国公司、中国政府以及世界各地的其他人对话,准备让他们加入这些协议,可能让他们加入这些理解,并打开沟通和外交渠道,同时为各种形式的监控、协议和追踪奠定硬件基础、物理基础。这些是你能做的最基本的事情。尽量开放尽可能多的有助于安全方面的事情。一路上,你知道,是的,我们可以做基本的事情,我不认为这足够,但我认为取得很多进展并不难。我们只需要决定去做。
Yeah, I think let me... and then but I think the perfect being the enemy of the good is definitely a problem. And you haven't got it started. The first thing that we can do that is really, really doable is we can get an antitrust waiver. This is like the most basic thing possible, which is just Donald Trump gets in front of the White House and he makes an announcement and he says, 'We understand this is going to be super dangerous. We have to do the good AI not the bad AI. We have to make sure the AI is going to be in control. That's what we said we want to do. We don't need more Trumpian language than that, but so in this interest, if the companies want to cooperate to keep AI safe, we're going to allow... we want you to do that. We want you to talk to each other. We want you to form agreements between yourselves. We will facilitate that if you want that.' You know, like basically just if OpenAI and Anthropic and Google get together and they say, 'We're going to run tests on each other's models and we're going to require these things as part of the voluntary framework. If we're going to agree to use these techniques and not use those techniques, or invest in these things, and share these results, and hold back things that don't meet these criteria or whatever, that none of this will cause anybody to do anything but say thank you.' Or just like we want you to make deals. We no longer... because it will always strike me that there are these people who, you bring up the idea that the AI companies should cooperate not to race against each other and to act responsibly. They're like, 'No, acting responsibly is illegal. We have antitrust laws for that.' And with a look in their eye that says, 'And I would want to sue their asses if they tried it. I would want to make sure they didn't dare.' And obviously the first thing you can do is just let you make deals. Let you reach agreement. Let you talk to each other. And then also, of course, start talking to China and Chinese firms and the Chinese government and everyone else around the world, and prepare to bring them into these deals and probably bring them into these understandings, and get lines of communication and diplomacy open, and also lay the hardware groundwork, lay the physical groundwork for various forms of monitoring, various forms of agreement, various forms of tracking. These are the most basic things you can do. Try to open up as many things as possible that contribute to the safety side of things. Along the way, you know, and yeah, we can do the basic stuff, and I don't think that's enough, but I think it's not that hard to make a lot of progress. We just have to decide to do it.
你觉得那是……我最近也在纠结的一件事是,我刚去了趟中国。很多试图做好事的人都很害怕,说实话,这些好事基本上是最平凡的,比如我想在中国有更多朋友,我想让我们两个文明之间有友好的关系,也许我们可以在某个研究项目上合作。非常平凡的事情,在平常时期没人会眨一下眼。但很多想做这类事情的人都很害怕,担心政府会介入并关闭他们,而且不是出于正当理由,而是因为美国政府的出口管制姿态,那些没有出口任何商业机密、芯片或贸易技术等的人,只是字面上试图在安全理念上达成共识,他们担心政府会打击他们,谁知道后果会是什么,但至少肯定会干扰他们的工作。我有点觉得也许我们应该在那个方面冒更多风险。所以,对于公司来说,我想也许他们应该现在就做,以后再在法庭上解决。反垄断的事情难道不应该是消费者保护吗?我认为他们可以提出一个相当有力的理由,说这是消费者保护。但他们真的需要特朗普的官方祝福才能开始做正确的事情吗?
Do you think that is... One thing I also have been struggling with recently is I just took this trip to China. There is a lot of fear among people who are trying to do good things that are honestly like the most mundane good things that are basically imaginable, where it's like I would like to have more friends in China and I would like to have friendly relationships between our civilizations and maybe we can work together on some research project. Very mundane stuff that in ordinary times nobody would bat an eye at. But there's a lot of fear among people who want to do that sort of stuff that the government's going to come in and shut them down, and not for legitimate reasons, but with the sort of export control posture of the US government where it is, people that are not exporting any trade secrets or chips or trading techniques or anything, but like literally just trying to have meeting of the minds on safety type ideas, they're afraid that the government's going to come down on them and who knows what the consequences could be, but certainly at a minimum interfere with their work. And I'm kind of like maybe we should take more chances on that front. So, when it comes to the companies, I wonder like maybe they should just do it now and fight it out in court later. Isn't the antitrust stuff supposed to be like a consumer protection? I think there's a pretty good case that they could make that this is a consumer protection. But do they really need the official blessing of Trump to get started doing the right thing?
从事中国相关事务的人,值得称赞的是,他们正在做这件事。他们只是大多试图保持低调,因为他们觉得美国政府的关注只会对那件事不利,他们可能是对的。但公司,我觉得他们有权力,有资源,有能力讲述他们的故事。我认为我们仍然有一个运作良好的独立司法系统。
The people working in the China stuff, to their credit, they are doing the thing. They're just trying to be very inconspicuous about it for the most part because they feel like the attention of the US government can only be bad for that, and they might be right about that. But the companies, I feel like they've got power, they've got resources, they've got the ability to tell their story. We still have a pretty well-functioning independent judiciary, I think.
他们不应该直接去做吗?
Shouldn't they just go for it?
所以,不幸的是,反垄断是那些已经远远超出原始法律意图的领域之一,它被应用在不太合理的地方。我理解为什么一部旨在防止行业串通的法律会对这种为了安全而进行的合作持怀疑态度,因为你可以偏执地认为行业会以某种方式合谋,对吧,而且可能是出于错误的原因。其他部门显然也很紧张。至于与中国合作,我的意思是,出口管制是错误的技术担忧。这只是华盛顿视中国为敌人的问题。如果你被视为与中国合作,如果你被视为与中国合作,你就得担心被当作敌人,或者被视为与敌人勾结。你得担心各种模糊的报复。我不认为有人真正知道具体机制会是什么,除了政府不喜欢你,因此怀疑你,因此以各种模糊的方式打击你。显然,如果我在做安全方面的工作,我觉得你根本不在乎。但我的意思是,Mythos 显然改变了游戏规则。大概 Astro 也会再次改变游戏规则。前沿信函中的暂停和 Hugging Face 事件也改变了游戏规则,等等。所以我的猜测是,现在已经有足够的理解,认识到存在一个真正的问题,如果你明确地在解决这个真正的问题,你会有比以前更多的余地。但是,是的,你确实——你看,特朗普政府实施了这些自愿的、带引号的、发布前沿模型的指南。没有人真正假装这些指南是自愿的。对吧?那么,当白宫因为一个愚蠢的担忧打电话给 Anthropic 关于 Fable 时,发生了什么?Anthropic 说:“不,我们不会自愿撤下那个。那很愚蠢。”好吧,他们最终还是撤下了,不是吗?直到他们能说服白宫自愿让他们重新放上去。如果你试图违抗这些自愿控制,你也会遇到同样的事情。所以同样,如果你试图以其他方式进行自愿合作,你最好得到当权者的批准。现在,我确实认为,如果 OpenAI 和 Anthropic 宣布一项在安全方面有意义的协议,以各种方式在安全方面进行合作,在这一点上我会更——在 Mythos 之前,我会担心他们会遭到积极报复,甚至被起诉并被要求停止。在这一点上,我认为白宫会合理地决定宣布胜利,而不是生气。但是,是的,你需要确信那会发生。
So, unfortunately, antitrust is one of those areas where it has gone far beyond the intent of the original laws and is being applied in places where it doesn't make that much sense. I know I see why a law that would safeguard against industry collusion would be suspicious of collusion in favor of safety in these ways, because you can be paranoid about the industry coming together in a conspiracy, right, in some sense, and trying to do this for the wrong reasons. And the other departments are obviously skittish. As far as the Chinese cooperation, I mean, export controls are the wrong technical worry. It's just a matter of Washington views China as the enemy. If you're seen as cooperating with China, if you're seen as working with China, you have to worry about being perceived as the enemy or being perceived as colluding with the enemy. You have to worry about amorphous retaliation. I don't think anybody really knows what the mechanism would be exactly beyond just the government not liking you and therefore being suspicious of you and therefore cracking down on you in various different amorphous ways. Obviously, if I'm doing safety stuff, I think you just don't care. But I mean, Mythos has changed the game clearly. And presumably Astro will change the game somewhat again. And the pause in the Frontier letter and the Hugging Face incident have changed the game and so on. So my guess is that there is now enough of an understanding that there is a real problem that if you are clearly working to address the real problem, you have more road than you would have had before. But yeah, you really— Look, the Trump administration implemented these voluntary, in air quotes, guidelines for releasing frontier models. And nobody is making any serious pretending that these guidelines are voluntary. Right? Well, what happened when the White House called Anthropic about Fable with a stupid concern. And Anthropic said, "No, we're not going to voluntarily take that down. That's stupid." Well, they ended up taking it down, didn't they? Until they could convince the White House to voluntarily let them put it back up again. You would expect the same thing to happen to you if you tried to defy the voluntary controls. And so similarly, if you tried to do voluntary cooperation these other ways, well, you better have the approval of the powers that be. Now, I do think that if OpenAI and Anthropic announced an agreement that makes sense on safety to cooperate on safety fronts in various ways, at this point I would be much— Before Mythos, I would have been concerned if they would have actively gotten retaliated against and even sued and told to stop. At this point, I assume that it was reasonable that the White House would just decide to announce victory rather than be mad about it. But yes, you need to be confident that's what's going to happen.
如果你要往这个方向走,你觉得可能想从哪里开始?有很多候选,我很想听听你的看法,但其中一个,Scott Alexander 最近一直在写,我们应该越来越多地思考我们可能生活在模拟中。我个人在路上的多个转折点都有这种感觉。让我更有这种感觉的一点是,我们现在有 METER 和 Redwood 进入 OpenAI 进行调查。这对我来说听起来就像电影剧本,这些人召集队伍,进行这次特别行动。这显然很戏剧化,他们很有趣,我很想成为墙上的苍蝇,待在他们的作战室里。但我也多次从这类组织的领导者那里听说,他们的头号担忧历来是确保他们站在公司正确的一边,以便下次还能被邀请。而且感觉我们现在正达到一个可能成为大问题的点。所以也许在六到十个审计组织和双寡头之间建立某种集体谈判式的流程,可能是一个很好的起点。我们在职业体育联盟和各种其他环境中都有这样的做法。你认为有机会建立这样的机制吗?我不认为那会侵犯任何人的行政特权。而且这可能是一件非常好的事情,因为现在我确实担心这些家伙最终仍然是在为 Sam 和 Greg 的意愿服务,对吧?如果他们想从这次调查中说出他们看到的真相,那对他们来说是一个非常棘手的位置。
Where do you think you might want to start if you were going this direction? So many candidates that I would be interested to get your take on, but one, Scott Alexander has been writing recently that we should all be thinking more and more that we might be in a simulation. I personally have felt that at multiple turns along the path. One that made me feel that just a bit more is the fact that we now have METER and Redwood going in to do the investigation at OpenAI. This just sounds like such a movie script to me where these guys are getting the gang together and going in for this special operation. And it's all obviously dramatic and they're funny and I'd love to be a fly on the wall in the room for some of their war room sessions. But I've also heard many times from leaders of such organizations that it's really important— their number one concern historically has had to be making sure they stay on the right side of the companies so they're invited back next time. And it does feel like we're now getting to a point where that's potentially becoming a big problem. So maybe some sort of collective bargaining type process between whatever half a dozen to 10 auditing orgs and the duopoly could be a really good place to start. We have that in professional sports leagues and all kinds of other environments. Do you think there's an opportunity to set something like that up? I don't think that would run afoul of anybody's executive prerogatives. And it might be a really good thing because right now I do worry that these guys are still serving at the pleasure ultimately of, I don't know, Sam and Greg, right? And that's a pretty tricky place for them to be if they want to just speak the truth as they see it coming out of this investigation.
啊,这是个担忧。显然,他们不是由那些公司资助的,但他们仍然必须依赖访问权限。你可以尝试给他们写一些强制性规则和协议,让他们无论如何都能保留访问权限,但我认为那不太会奏效,因为你不能强迫这些事情,至少在没有严厉的政府规则的情况下。我确实认为 METER,尤其是 METER,已经达到了一个点,试图排除 METER,试图把 METER 当作不受欢迎的人,因为他们说了你的坏话,这会引起足够多的问题,以至于 METER 可以承受相当严厉的批评而不必担心。而且我确实认为实验室合法地——如果实验室真的只是想完全安全清洗,只是想获得 AAA 评级,我们会有更严重的问题。我认为实验室确实想知道他们的安全问题,他们不想被视为淡化安全问题。而且,如果 Astra 有安全问题,或者 Sol 有安全问题,或者 Favreau 有安全问题,或者 Opus 有安全问题,让 METER 不去发现它并不会有助于你的长期公关策略和对情况的反应。因为模型会发布出去,然后人们会看到模型做什么。他们会看到哪里出了问题。没有人会被愚弄太久。所以,我认为评估公司几乎可以只说真话,几乎可以做这件事的原因是,评估在长期内不像评级那样可以伪造——比如当你应该给出单 A 评级时却给出了 AAA 评级,90% 多的情况下没人会发现,因为债券支付了,而且你是对的。
Ah, it's a worry. Obviously, they're not funded by those companies, but they still have to rely on access. You can try to write them to some mandatory rules and agreements that they get to keep access regardless, but I think that's kind of not really going to work because you can't force these things, at least not without heavy government rules. I do think that METER, especially in particular, has reached a point where trying to exclude METER, trying to treat METER as persona non grata because they said something nasty about you would cause enough problems that METER can afford to be pretty harsh without worrying about that. And I do think the labs legitimately— if the labs legitimately just wanted to safety wash fully and just wanted to get AAA ratings on their bonds, we would have a much more serious problem. I think the labs legitimately do want to know about their safety problems and they do not want to be seen as downplaying safety problems. And to the extent that if Astra had a safety problem, or Sol had a safety problem, or Favreau had a safety problem, or Opus had a safety problem, it's not as if getting METER not to find it is going to help with your long-term public relations strategy and reactions to the situation. Because the model's going to be out there and then people are going to see what the model does. And they're going to see what goes wrong. No one's going to be fooled for very long. So, I think that the reason the eval companies pretty much get to just tell the truth and pretty much get to do the thing is because the eval is not fakeable in the long term the way that like a rate— like when you give a triple-A bond rating when you should have given a single-A bond rating 90-something percent of the time, no one ever finds out because the bond pays, and you were right.
问题在于,你是在给出一个风险多久会变得严重的评级,而当风险变得严重时,你最初被评为 AAA 级这件事会让我皱眉,但你也因此获得了更好的条件,随它去吧。在这里,从发布后的头几周就能很快清楚地看出你是否遇到了严重问题。人们,比如想想 O3,那个说谎的骗子。如果我们有评估,而评估说“O3 没有对齐问题”,那撑不过两天,对吧?普通人会立刻注意到 O3 是个说谎的骗子。而他们会在我说起模型栈的时候注意到,也许在我提到模型栈的时候,我在 Twitter 上看到这个模型是个说谎的骗子,然后大家都在报道这些问题,然后我看模型卡,上面写着“诚实基准看起来都不错。Meters,不管他们雇的谁做的评估,说看起来不错。”然后我就想,“好吧,他们只是在掩盖问题。”所以,等到有人知道这件事的时候,他们做了什么?他们骗了像我这样读模型卡的 20 个人,骗了一天。现在我们很生气,因为他们骗了我们。这对他们没好处。那显然对任何人的声誉都没好处。而且实验室也从人们信任的评估中获益良多。当你强迫实验室,或者实验室强迫评估人员配合,如果那样做,就会侵蚀信任,因为人们会发现的。人们会看到的。我不认为有很多人搞不清楚穆迪和标普是否在债券评级上注水,对吧?由于这些动态,华尔街人人都知道。每个参与这些交易的人都知道,大家都很亲密,人们在四处寻找最好的评级机构,而 AAA 并不真的意味着我们希望 AAA 意味着的东西。所以同样地,如果 Meters 开始发布类似债券评级风格的标签,给各种模型的对齐水平贴标签,比如说,然后他们开始可疑地频繁把东西标为 AAA,我想每个人都会明白,好吧,那是垃圾。我们在这类事情上的认知比金融界那些人要好得多。
The issue is you're issuing a level of how often is this risk going to become serious, and when the risk becomes serious, the fact that you were initially triple-A rated raises my brows, but also you got better terms, whatever. Here, it's pretty obvious very quickly from the first few weeks of release whether or not you had a serious problem. People, like, think about O3, right, the lying liar. If we had evals, and the evals had said, 'O3 doesn't have an alignment problem,' that would not have lasted for 2 days. Right? The ordinary people would have noticed immediately that O3 is a lying liar. And they would have noticed by the time I brought up, maybe by the time I'm bringing up the model stack, I am seeing on Twitter that this model was a lying liar, and then everyone is reporting on these problems, and then I look at the model card, and it says 'Honesty benchmarks all look good. Meters, whoever's evaluation they hired, said it looked good.' And then I'm like, 'Okay, they're just hiding the problem.' So, by the time anybody learns about this, what have they done? They have fooled the 20 people like me who read the model card for a period of a day. And now we're pissed because they fooled us. Doesn't help them. That doesn't obviously help anybody's reputation. And also the labs benefit a lot from having evals that people trust. And when you force the labs, or the labs force the eval people to play ball, if they do that, it erodes trust because again, people find out. People see it. I don't think that many people were confused about whether Moody's and S&P were kind of juicing the numbers on the bond ratings, right? Due to these dynamics, everybody knew on Wall Street. Everybody knew who was involved in any of these trades, that everybody was cozy, and people were shopping around for the best grader, and that AAA did not really mean what we'd like AAA to mean. And so similarly, if meters started issuing bond-rating-style labels on the alignment levels of various models, let's say, and they started labeling things AAA suspiciously often, I think everyone would just understand, okay, that's garbage. We have a much better epistemic than those people about this type of thing than the financial world did.
我们能先岔开一下这个话题,然后再回来吗?短期内我们应该对生物风险有多担心?因为我觉得,至少对我来说,所有这些分析的一部分是,未来还有很多轮博弈,长期声誉会很重要,诸如此类。但我看到这些东西时最直接、最尖锐的恐惧是,如果这是一个在生物能力上与黑客软件能力相当的模型,它很可能真的会出去,真的拿到那个病毒,合成那个新型病毒,对吧?它已经证明了它是不屈不挠的。它已经证明了它足够有创造力,能绕过各种旨在阻止它的系统。我认为理论上我们现在在 DNA 合成公司有不错的筛查。
Can we take a quick detour on this and then maybe come back to it? How worried should we be about bio risk in the near term? Because I think part of all that analysis, to me at least, is that there are quite a few iterations of the game to come and reputation long term is going to matter and all that kind of stuff. But my immediate, most acute fear when I saw this stuff was like if this had been a model with similar bio capability to what it had in hacking software, it might have very well gone out and actually got that virus, that novel virus synthesized, right? It's proven that it is relentless. It's proven that it's creative enough to get around all sorts of systems designed to stop it. I think we have in theory decent screening at the DNA synthesis companies now.
是的。
Yeah.
可能某个地方还是有漏洞,一个非常坚定的行动者能找到的漏洞。而且每个人似乎都在说,是的,生物比网络落后 12 到 18 个月。这显然不是板上钉钉的,这取决于公司在训练方面决定做什么,但我可能在想,考虑到所有这些关于集体谈判或真正给审计师更多保证的概念,我们可能在不久的将来进入生物领域一个相当严重的紧要关头,而如果我们谈论的是新型病原体的释放,声誉就重要得多了。我认为有些事情真的无法挽回。
Probably there's still some hole in there somewhere that a very determined actor could find. And everybody seems to be saying, yeah, bio's like 12 to 18 months behind cyber. Which obviously is not written in stone and that depends on what companies decide to do in terms of training, but I'm maybe thinking in terms of all these notions about collective bargaining or really giving more guarantees to the auditors that we might be entering a pretty serious crunch time on the bio front in the not too distant future and reputation matters a lot less if we're talking about novel pathogens being released. I think there are some things really can't take back.
这在很大程度上很重要,因为如果你释放一种新型病原体,那可能是你公司的终结。不管世界其他方面发生什么,对吧?如果 100 人死亡,全球新闻连续报道两周,每个人都对疫情感到恐惧,然后我们控制住了,然后结果发现这不是世界末日,你还有公司吗?我不知道你的公司会怎样,有多糟糕?我不知道有多糟糕,但肯定不好。对于生物,就像黑客一样,测试是让你去黑东西,这是合理的。而你通过测试的方式是,你也许黑其他东西。我们有了黑其他东西的想法。这很自然地理解这是怎么发生的,在指导方针降低的情况下,有一个目标,为了实现那个目标,它开始进行黑客攻击,开始做非法的事情,开始做可能有害的事情。对于生物,很难想象一个评估会导致它真的物理上尝试合成,比如通过实际的物理生产机制。真正的东西。我的意思是,你可以争辩说,如果你想解决问题,好吧,我解决问题的唯一方法是我做我的物理实验,我试一试,看看是否有效。这在某种意义上似乎比实际发生的事情更疯狂。你必须物理上有能力做那件事,包括被政府黑掉等等。你必须在更根本的意义上想要它。而且如果 AI 真的合成了危险的东西,比如生物武器之类的,超过 27 个或者什么的,然后送到开放办公室。开放办公室会说,“这是什么生物武器包裹?”然后叫来遏制小组,在那个故事结束后,我们希望如此。我不知道。你会通过感染一群人通过测试,我希望它明白这一点。但也有很多疯狂的泄密,包括你没有听说过的遏制事件,比如在苏联时期,涉及的专业人员相当粗心大意和愚蠢。所以我不知道。但我想说,生物的真正危险是,是的,世界上有少数人处于近期生物风险中,对吧?不是长期生物风险,那种 AI 可能有奇怪的别有用心的动机,长期反复危险什么的。所以短期内就像,“好吧,哈马斯或真主党或朝鲜人什么的。一些明显不怀好意的人,他们就是想让很多人死什么的。”对吧?他们拿到一个 AI,你说 AI 来找出如何制造生物武器或病原体。然后我们得到使用它的威胁或实际使用。这是一个你应该担心的场景,因为这肯定需要保障措施,但同样,它不会是一种威慑。而且它不会是一种威慑。是的,我对此很担心。
It matters a lot in the sense that if you release a novel pathogen, that might be the end of your company. Regardless of what else happens to the world, right? If 100 people die and there's a it's on the global news for 2 weeks and everyone's terrified about a pandemic and then we contain it and then it turns out it's not the biggest deal in the world, do you have a company? I don't know what happens to your company, how bad is it? I have no idea how bad it is, but it's not good. With bio, like with hacking, it makes sense that the test is to ask you to hack things. And the way you pass the test is you hack maybe other things. We get the idea to hack other things. It's sort of natural to understand how this happened, where with the guidelines down and there was a goal and to accomplish that goal it started doing hacks, it started doing illegal things, it started doing potentially harmful things. With bio, it's harder to imagine an eval that would cause it to actually physically try to synthesize, like through an actual physical production mechanism. The actual thing. I mean, you could argue that if you want to solve the problem, okay, the only way I solve the problem is I do my physical experiment, I try it out and I see if it works. This seems a lot more insane than what happens in some sense. You would have to be physically capable of doing the thing including getting hacked by the government and so on. You kind of have to want it in a much more fundamental sense. And also if the AI did synthesize this into dangerous bio weapons and stuff over 27 or whatever, and then send to the open air office. The open air office goes like, 'What is this bio weapon package?' And then calls containment and after the end of that story, something we hope. I don't know. You'll pass the test by infecting a bunch of people and I would hope that it understood this. But also there have been a bunch of wild leaks including the you don't hear about the containment like in the Soviet times and such with involved professionals being pretty careless and stupid. So I don't know. But I would say the real danger with bio is yeah, that there's a small number of people in the world who is in the near-term bio, right? Not like the long-term bio where the AI might have some weird ulterior motives and being dangerous in the long-term time and again whatever. So in the short-term it's like, 'Okay, Hamas or Hezbollah or the North Koreans or whatever. Some clearly up to no good people who just want a lot of people to die or something.' Right? Get a hold of an AI that you say AI to figure out how to make a biological weapon or pathogen. And then we get a threat to use it or use it. That is a scenario you should be worried about because that would definitely require safeguards but again, it would not be a deterrent. And it would not be a deterrent. And yeah, I'm worried about it.
问题在于,生物领域有一个非常明显的阶跃变化,从什么都没发生到需要非常糟糕的事情发生,中间几乎没有过渡。而网络领域,你从攻击一些低价值的软目标开始。如果你以为所有坏黑客、所有恐怖分子、所有流氓国家、所有坏人都会集体说‘不,我们等到能击中一个真正诱人的目标再说’,那你就错了。实际上,你会看到越来越严重的事件,击中越来越重要的目标。当然,大国可能已经准备好了,比如以协调方式发动更大规模的攻击,但大多数情况下,你会得到这些早期预警信号。
The problem is that with bio there's a pretty clear step function change from nothing bad happened to needing something very bad. Like there's very little in between. With cyber, you start hacking some low-value soft targets. If you think you're not going to have all the bad black hat people, all the terrorists, all the rogue states, and all the bad people just collectively say, 'No, let's wait until we can hit them with a really juicy target.' No, what you see is you gradually see more and more serious incidents hitting more and more serious targets. And yeah, of course the big nation states may have things ready to go, like launching bigger attacks in a coordinated fashion, but mostly you'll get these early warning signs.
如果生物领域没有明显进展,比如要么好到能想出具有大流行临界规模或严重事件的东西,要么不行,这有点像布尔效应。所以你可能会有很大的开销,你可能面临情况问题,你可能遇到规模问题,比如如果错误的人现在在没有安全协议的情况下得到了这种病毒,或者接触到了国家病毒库中的内部病毒,或者得到了阿司匹林之类的,也许他们就能造成真正的问题。
If bio is not obviously getting right, like if it's either good enough to figure out something that has the critical mass of a pandemic or a serious incident, or it doesn't, it's kind of a Boolean effect. And so you might have a substantial overhead, you might have a problem with the situation, you might have a scale issue where if the wrong person got their hands on this virus right now without the safety protocols, or got access to the internal virus during the state virus, or got access to aspirin or whatever, maybe they could cause a real problem.
好消息是,既足够熟练又足够有动力并试图这样做的人数非常非常少,目前可能为零。但这并不意味着你不会得到这些早期预警信号。而且我没有看到人们以同样的方式认真对待生物局势,比如积极试图制造事端。他们最好的尝试是什么?除了部分感染人群,或者当有……我不知道,这很可怕。
The good news is that the number of people who are both skilled enough and motivated enough and trying to do this is very, very small, and possibly zero for now. But that also doesn't mean you don't get these early warning signs. And I don't see people taking the bio situation seriously in the same sense of actively trying to cause something. What's the best they're trying to do? Except for that, infect people in part, or when there's... I don't know, it's scary.
问题在于,我们已经有了类似的情况,好吧,如果未来 12 个月内有 5% 的概率出现非常严重的生物问题呢?那会与 0.5% 的概率有多大不同?对吧?我不得不猜测,现在有 5% 的保证,我认为我们正走向一个世界真正面临严重问题的风险点。而且我认为第一个问题有一定概率成为真正的大流行。因为在生物领域,一旦你得到某种具有足够传染性、足够危险、足够难以控制的东西,你就有点谨慎了。
The problem being that we already have something like, okay, what if there's a 5% chance in the next 12 months that there's a really serious bio problem? Would that look that different than a 0.5% chance? Right? And I had to guess, 5% guarantee right now, I think we're going to a point where there's real risk in the world that we're going to have an actual serious problem. And I think that the first problem has some chance of being an actual pandemic. Because in bio, once you get something that's sufficiently infectious, sufficiently dangerous, sufficiently hard to contain, you're kind of discreet.
所以我不知道。我认为一旦到了这个地步,正确的谨慎程度在很多情况下看起来是疯狂的。而这正是 Anthropic 对其生物过滤器采取如此多措施的原因,基本上就是告诉人们几乎不要尝试做任何生物相关的事情。这确实是故意的,对吧?他们就是非常广泛地关闭,尽管我们知道我们关闭的请求中有 90% 多、99% 是合法的。也许是 99.99%。但我们必须这样做,否则我们会对抗性地得到我们不想要的查询。而这更重要。
So I don't know. I think that the right level of caution looks crazy in many cases once you get to this point. And that's a lot of what the reasoning is that Anthropic took so much for their bio filters, shutting down basically told people trying to do pretty much anything in bio. Which is really intentional, right? They just really went for a very broad shutdown, even though we know that 90-something, 99% of the requests that we're shutting down are legitimate. Maybe 99.99. But we have to do this because otherwise we'll adversarially get the queries that we don't want. And that just matters more.
如果同时我们必须应对大流行,那么我们在癌症和其他问题上取得一些进展就无关紧要了。这不是一个好的交易。所以我们只是不知道我们处于什么位置,所以我们必须谨慎行事。你知道吗?如果生物领域的人需要一直使用落后前沿三个月或六个月的人工智能,那是不幸的,但根据定义,这最多只能让你慢三个月或六个月。这是重要的一步。世界在开发新事物方面不可能落后超过六个月。
It doesn't matter if we make some progress on cancer and all these other problems if simultaneously we have to deal with a pandemic. That's just not a good trade. So we just don't know where we're at, so we have to play it cautiously. And you know what? If the bio people need to be using AIs that are three months or six months behind the frontier at all times, that is unfortunate, but at most by definition that can slow you down by three or six months. It's an important step. The world cannot possibly be more than six months behind in its development of new things.
所以也许我们仍然应该在这里获得人工智能的大部分好处。特别是如果你有这个模型,人工智能可以帮助你,人工智能可以让你变得更好,但随后你开发出这种候选药物或关于生物学的候选假设,然后你必须花 5 年研究它,然后花 5 年通过 FDA 流程,然后 10 年后你可以开始分发成果。好吧,如果整个事情在提升质量方面延迟了 6 个月,你仍然会得到提升,你一直会得到 6 个月前你会得到的提升量。这并不好,是的,你可以说人们会死,因为你未能防止这种情况,这很糟糕,但在这些假设下,你仍然获得了绝大多数好处。
So maybe we should still be getting most of the benefits of the AI here. Especially if you have this model of AI can help you out and AI can make you better, but then you develop this candidate drug or this candidate hypothesis about bio and then you have to spend 5 years researching it and then spend 5 years making it through the FDA process and then after 10 years you can start to distribute the deliverables. Well, if that whole thing is delayed by 6 months in terms of the quality of the boost that you get from it, you still get the boost, the amount of boost you would have gotten 6 months ago at all times. That's not good, and yeah you can say people are going to die in the sense that you failed to prevent that and that's terrible, but you're still getting the vast majority of the benefits under these hypotheses.
很多时候,当你处于前沿发展速度的情况下,生物局势的等级,如果事情没有以可怕的速度发展,那么你要求采取的预防措施就不是什么大问题。对吧?比如假设我告诉你现在是 2020 年。我把它作为一个可能性提出来。我会让你只能使用落后于最好的人工智能 3 个月的人工智能,以换取所有可能发生的坏事,或者至少很多非常糟糕的人工智能事件不太可能发生。你会合乎逻辑地说,好吧,所以不是 3 年后得到,而是 3 年零 3 个月后得到。你知道,那并没有太大改变我的生活。显然,那一刻会很烦人,因为我没有酷玩具。但比如说我,我是一个游戏玩家,比如 Steam 上的字面游戏之类的,如果你告诉我你可以玩所有游戏,但每次有新游戏你必须等 3 个月,社交上有时有点烦人,因为每个人都在玩新热门,而你没有。但你只是玩同样的游戏。没关系。这都是重要的事情。
A lot of the times when you get in the situation of the pacing of the frontier, the letter grade of the situation of bio situation, where if things are not advancing scary fast, then the precautions you're asking to take are not that big a deal. Right? Like suppose I told you it's 2020. I'm going to pose it as a possibility. I am going to make it so that you can only use artificial intelligences like they're 3 months behind the best artificial intelligence, in exchange all the really bad stuff that might happen, or at least a lot of really bad AI much less likely to happen. You would logically say okay, so instead of getting it in 3 years I get it in 3 years and 3 months. You know, that didn't change my life that much. It's going to be annoying in the moment obviously, that I don't have the cool toy. But say me, I'm a player of games, like literal games on Steam or whatever, and if you tell me you can play all the games but you have to wait 3 months every time there's a new game, socially that's a little annoying sometimes because everyone's playing the new hotness and you're not. But you just play the same games. It's fine. It's all important stuff.
唯一一个必须落后一个模型步骤而不是致命的世界,是一个巨大悲剧,只有当你真的认为人工智能的每一个增量进展都有巨大好处时。而这基本上只有在节奏论者正确的情况下才成立。而我们实际上正在走向那些对你来说应该非常非常可怕的道路。你知道,那些想要接受、说我们需要一直加速的人,他们大多数并不真正相信这些事情会变得这么疯狂。他们只是希望看到好处,担心我们会失去好处。但那样我们会失去很少的好处。
The only world in which having to be one model step behind in these open instead of fatal is this huge tragedy is if you really think there's this huge benefit to every incremental motor progress in AI. And that's basically only true if the pacing people are correct. And we are in fact going to those paths which should be really really scary for you. And you know, the people who want to accept who say we need to accelerate all the time, they don't really believe in these things going this crazy most of them. They just want to see the benefits and they're worried that we'll lose the benefits. But we're going to lose so few benefits then.
对吧?比如我对自动驾驶汽车超级兴奋,但我们会失去那些,对吧?但如果你告诉我我们只需要把这件事推迟 6 个月,然后我们就一路走到现在这里,我会说,酷,成交,谁在乎?行吧。因为再说一次,我人生的大部分时间都在 6 个月之后。到时候我们再担心。就像,好吧。我们会在一种技术存在但被搁置数年、数十年且不被滥用的情境下担心它。所以这些人有点各说各话,在一种他们其实不需要把它当作绿灯的情况下。而且不幸的是,这大概就是你怎么也无法同时调和这两方。所以就这样吧。
Right? Like I'm super excited by self-driving cars, but we're going to lose those, right? But like if you told me we just have to delay this thing by 6 months and then we got all the way to where we are I'd be like, cool, done, deal, who cares? Like fine. Like because like again, most of my life is more than 6 months from now. Then we can worry about it. It's like, okay. It is we're going to worry about it in a situation where like this technology exists and then it gets held in limbo for years and decades and like doesn't get abused. So like these people are kind of talking past each other in a situation in which like they don't need to have it as a green light. And like unfortunately like it's all about how this probably you can't reconcile these two sides at the same time. So peace out.
你怎么看 OpenAI 的公关支持了暂停——不是暂停,而是那封节奏信?山姆和格雷格都没签,而且德米斯也没签,他可是一直在呼吁国际合作最积极的人。你对这些签名缺失有什么解读,或者我们应该如何看待此刻谁在想什么、在向谁发出什么信号?
How do you read the fact that OpenAI comms supported the pause, not the pause, but the pacing letter? Neither like Sam nor Greg did and then also no no Demis and he's been the one out there calling for international cooperation the most. Any read on why those signatures were missing or how we should be thinking about who's thinking or signaling what to whom in this moment?
我认为《给前沿模型定节奏》的重点是表明实验室员工,那些没有理由炒作的低层员工,普遍感到担忧。而且人们担心如果这封信由奥特曼和其他高层签署,会被视为营销噱头。那样他们的签名可能适得其反。他们之后确实发表了来自 Anthropic 和 OpenAI 的支持信,而且他们已经决定签署。但如果我是奥特曼,我认为说“我认为签署这封信可能会降低而非提高其效果,因为井水已经被污染得这么厉害”是一个非常合理的决定。再说,我看到所有那些认为 Hugging Face 被黑是营销噱头的人,然后 Anthropic 回到自己的档案里发现自己犯了一些非常愚蠢的错误,又被认为是另一个营销噱头。这更说不通,因为 Anthropic 做的事情没什么了不起的。他们就是搞砸了。什么?他们碰巧在一堆网站上发现了弱密码?为什么有人会觉得这值得佩服?显然这不是营销噱头。这太疯狂了。但奥特曼确实特别谈到过他在华盛顿时提到需要定节奏。我们有他这么说的录音。所以显然他不签不是因为反对这个想法。他让 OpenAI 发表了声明。OpenAI 参与了措辞。他在自己的声明中也呼应了这种语言。我认为他只是觉得,不,实际上如果我签了,人们会说这是我的声明,他们会更加愤世嫉俗。所以我不签反而有帮助,就像我主动提出加入战争部的一份法庭之友简报,但经过考虑,人们说谢谢你的提议,但我们认为这实际上会适得其反,所以我们不会把你的名字放上去。我们感谢你对措辞细节的备注,我们会做出修改,但我们实际上不想把你的签名放上去。我说,好吧,我理解,我觉得这有道理。是的,我对其他像开源信由 OpenAI 签署也有同样的看法。我认为他们明白他们在玩一套象征性的游戏,就像你该说什么就说什么,该闭嘴就闭嘴。有时候他们会犯错,但他们在努力。
I think that the point of pacing the frontier was to indicate that the lab employees, the lower level people who have no reason to be hyping are en masse worried. And that there was concern that if it was signed by Altman and other top people, that it would be seen as a marketing stunt. And then like their signatures could be counterproductive. They did put out the support letters afterwards from Anthropic and OpenAI and they already decided to sign. But like if I'm Altman I think it's a very reasonable decision to say I think signing this letter would potentially make it less effective rather than more effective at this point because the well has been so poisoned. Again, I look at all the people who thought the Hugging Face hack was a marketing stunt and then Anthropic going back into its archives and finding that it had made some really stupid mistakes was another marketing stunt. Which like makes even less sense because like there's nothing impressive about what Anthropic did. They just screwed up. What? They found weak passwords on a bunch of websites by accident? Like why does anyone supposed to be impressed? Like obviously it's not a marketing stunt. Like this is crazy. But like uh Altman has like specifically talked about how he mentioned the need to pace while in Washington. We have him on tape saying this. So clearly he didn't not sign because he's opposed to the idea. He let OpenAI come up with the statement. OpenAI participated in the wording. He has echoed the language in his own statements. I think he just thought no, actually if I sign this, people can talk about the fact that it's my statement and they're going to be even more cynical. And so I will be helpful by not signing the same way that like I offered to join one of the amicus briefs in the Department of War and on reflection people said thank you for your offer but we think this would actually be counterproductive so we're going to leave your name off of it. We appreciate your notes on the details of how to word this and we'll make the changes but like we don't actually want to put your signature on it. I was like okay I understand and I thought that makes sense. Yeah and I think the same thing about like the other like open source letter like being signed by Open AI. I think that like they understand that like they're playing a set of symbolic games and like you say what you have to say and you shut up when you have to shut up. And like sometimes they make mistakes but like they're trying.
所以回到定节奏以及它在实践中可能是什么样子,我认为有一个相当一致的问题,你作为模型卡的仔细阅读者,可能比我更了解,那就是审查过程非常短,对吧,他们没有太多时间,即使在当前的调查中似乎也是如此。我认为 Ryan Greenblatt 说过这会很快,然后 Daniel Kokotajlo 和其他人进来说,为什么必须这么快?为什么我们不能给你任何你需要的时间?如果我们真的在谈论给前沿模型定节奏,我们可以想象的一个协议是给审查者更多时间,在他们必须提交报告和模型发布之前。你可以对此做出回应,但然后我很想听听你对你认为可以达成的协议的想法。
So coming back to pacing and what it might actually look like in practice there is a pretty I think consistent problem that you as a close reader of model cards I think are probably more in tune with than I am where the the review processes are just very short right they don't have a lot of time that even seems to be the case in the current investigation. I think Ryan Greenblatt said it's going to be real quick and the people Daniel Kokotajlo and others came in and said why does it have to be so quick? Why can't we like give you whatever time you need? This would be if we're literally talking about like pacing the frontier one agreement we might imagine is like giving the reviewers more time before they have to submit their reports and the model goes to press. You can react to that but then I'd love to hear your ideas about like sort of agreements you think are reachable.
有很多不同的地方你可以加速或减速,如果你把私有模型向公众发布时多延迟一个月,理论上对吧?类似于现在的自愿流程,我们将转向白宫的 30 到 60 天延迟,这非常可能,而且 Methuselah 在能够发布稳定版之前,在某种重要意义上被延迟了两个月。那么你不一定需要以非平凡的时间量延迟 OpenAI 和 Anthropic 的前沿模型开发。因为那些人会使用未发布的模型。如果有的话,你可能会增加 OpenAI 和 Anthropic 内部使用的东西与所有其他地方使用的东西之间的差距,包括中国实验室,包括那些试图扩散和快速跟进的人。我想到的是,如果中国人落后前沿 7 个月呢。但他们落后的 7 个月是相对于向公众发布的模型。如果你不向公众发布模型,他们仍然会落后 7 个月。对吧?如果 OpenAI 和 Anthropic 在开发出每个模型后都持有 6 个月,你仍然会看到公开版本,然后,你知道,也许 6 个月而不是 7 个月后,你会看到开放模型的快速跟进者、扩散、蒸馏以及开发所有这些的方式。但这个差距实际上不会改变太多。但然后 OpenAI 和 Anthropic 的人会比其他人有额外的领先优势,因为他们会使用这个内部模型,这些领先会复合,你会看到这些东西,在某种意义上这真的很好,因为现在那些人可以更负责任地行动,利用拥有更好模型的优势来投资更多与安全相关的东西,同时他们这样做。所以那可能有它的优势。但内部模型正是我们现在最害怕的。内部 AI 研发的自动化是我们害怕的。所以那是我们想要阻止的东西。对吧?你真的不想要一种情况,比如如果 OpenAI 和 OpenAI 比他们的公开版本领先 6 个月,那就是他们在我们拥有 GPT-7 之前拥有超级智能的 6 个月。
There's a lot of different places in which you can speed up or slow down if you release your private models to the public with an extra month's delay right in theory right? What similar to the voluntary process now we're going to switch to the White House 30 60 day delays are like highly probable and like Methuselah's got delayed by two months right in some important sense before it was able to release its stable. Then you do not necessarily delay the front development of frontier models at Open AI and Anthropic by a non-trivial amount of time. Because those people are going to use the models that are not released. If anything, like you might increase the gap between like what's being used internally at Open AI and Anthropic and what's being used everywhere else, including at the Chinese labs, including the people trying to diffuse and fast follow. And like what I thought of this is what if the Chinese are 7 months behind the frontier. But the within their 7 months behind is the models that are released to the public. And if you were to not release the models to the public, they would still be 7 months behind. Right? Like if that if Open AI and Anthropic held every model for 6 months after it was developed, you would still see the public releases and then, you know, maybe 6 months instead of 7 months later, you would see the open model fast followers and diffuse and and distillations and ways of like developing all that stuff. But like this gap actually wouldn't change very much. But then the people at Open AI and Anthropic would be they would have this extra lead on everybody else because they'd be using this internal model and these leads would compound and you would see these things like and in some sense that's really good because now those people can like act more responsibly by using some of this advantage of having better models to invest more in this safety-related stuff while they're doing that. So that could have its advantages. But internal models are kind of what we're scared of the most right now. And the internal like automation of AI R&D is what we're scared of. So that's kind of the thing we want to stop. Right? Like you like you really don't want to have a situation in which like if Open AI and Open AI are 6 months ahead of their public releases and that's 6 months like they have superintelligence before we have GPT-7.
对吧?这完全可能发生在他们害怕的那个世界里。而且如果发生这种事,会带来很多新问题。它既解决一些问题,也制造另一些问题,对吧?我不知道这是好是坏。这是个很大的话题。我认为当你面对前沿时,我们说的是主动限制训练规模、限制开发模型的能力,包括内部开发,如果到了那一步的话。我觉得必须谈这种方案,否则根本行不通。因为再说一次,如果奇点发生在顶级实验室内部,或者可能几个实验室内部,那在很多人们关心的意义上并不一定更安全。显然,如果你赞成单极或单一实体式的解决方案,你对其中一些问题的看法会不同。但现在很多人已经原则上强烈反对那种做法,这解决了一些问题,却让另一些问题难得多。对吧?如果在重要意义上你只需要担心一个 AI,那很多问题就变得容易解决了。比如你不需要做到完美高效,你可以承受很多权衡。你不用担心那个更有竞争力、更无情、更不对齐的 AI 会占据竞争优势。你只需要担心谁跑得更快。你不用担心人类被迫把权力委托给他们的 AI,因为如果别人的 AI 会做,如果我不做,他们就会做。如果 AI 有某种原则,它可以有关于什么该做什么不该做的指导方针,对所有人都一样。所以我们可以保留生活、行动、决策的领域,我们可以做所有这些事,但代价是潜在的集权化、权力集中,代价是有人——某群人——来决定这个 AI 要做什么,或者没人决定,那更糟。对吧?在重要意义上,因为那样就没人做决定了。但如果有很多 AI,那重要的是没人做决定。所以这件事没有好的答案。没有人有好的解决方案。这就是为什么回到大卫·多伊奇的说法,统一 AI——如果我们解决了技术问题就能成功 95% 的想法——忽略了之后我们仍然必须走的窄路问题。对吧?如果我们没有任何能力引导局势,没人控制局面,那自然 AI 会被赋予越来越多的权力,事情会走向它们自然的竞争性资本主义终点,我们从根本上没有竞争力,然后我们把权力交给它,然后事情就变糟了。但如果有人控制——有人用某种方法论控制——那是什么?谁?怎么控制?对吧?谁在做那些决定?然后如果我们把权力委托给一个 AI 或一组 AI,那也是一个决定。有任何决定是好的吗?听起来都可怕。我觉得很多时候我们在告诉人们他们在选毒药。要么选他们能忍受的毒药,要么选他们不能忍受的毒药,然后相应地行动。对吧?就像他们在说,你知道,我不信任权力集中,所以我选择避免权力集中的方案,然后希望其他部分神奇地解决,而没有理由阻止它。很多人选择那个,因为他们理解权力集中,他们知道权力集中是真实的。但他们并不真正理解 AGI,更不用说超级智能了。他们认为那些东西不太真实,或者希望它们不真实。所以他们只是希望另一个问题不会爆发,即使他们有点——他们不一定意识到他们已经排除了所有可能的解决方案,如果事情出错,因为他们走得太远。所以再说一次,对于这些,我没有一个大家会喜欢的答案,因为你是在坏选择之间做选择。我认为我们应该谈谈这个,因为如果我们确实要面对未来,如果事情不会——我们正在自动化 AI 研发,这对我们不利——我认为唯一能做的就是说不。别那样做。不要那么快地开发这些能力这么强的 AI。你需要放慢脚步,因为我们不知道如何处理这个。我们有所有这些不同的解决方案,所有这些不同的方式来尝试应对,但都是坏的。不是因为我们认为对齐问题无法解决。就像我们不认为 AI 能在走向这种社会可能性的过程中解决它。我当然不解决那个,但我们不认为我们会自然地在各个不同层面上得到那个问题的好解决方案,而且你必须同时解决所有这些不同的元层面。你不能只在一个层面上解决。你必须第一次就解决。然后即使我们解决了那个问题,我们还没有准备好知道该怎么处理它。一旦你创造了那个因素,你继续这样做,一旦你做了,就会发生一些事情。你不能长时间什么都不做。再说一次,所有这些不同的路径都有各自不同的“天哪,那太可怕了”的附加。你必须选一个。或者找一个全新的。所以你需要放慢脚步。
Right? Like which is entirely a thing that could happen in the world they're afraid of. And if anything like that causes lots of new problems. It also solves some problems and causes others, right? Like I don't know if it's good or bad. That's a huge conversation. I think when you have to face the frontier, we're talking about actively restricting like levels of training runs and ability to develop the models including internally if it comes to that. I think it has to be talking about that kind of approach or it doesn't really work. Because again, if there's a singularity that's kind of internal to the top lab or to maybe a few labs, then that's not necessarily safer in many senses people care about. Obviously, if you are in favor of unipolar or singleton style solutions to this problem, you have a different perspective on some of these questions. But a lot of people these days have expressed strong opposition to that kind of thing on principle, which solves some problems and makes some other problems so much harder. Right? Like if you only have to worry about one AI in an important sense, then there are a lot of problems that get a lot easier to solve. Like you don't have to be perfectly efficient. You can afford to make a lot of trade-offs. You don't have to worry about whether the more competitive, more ruthless, more misaligned one has competitive advantages. You have to worry about the person who moves faster. You don't have to worry about the humans being forced to delegate to their AI because if somebody else's AI is, because if I don't do it, they will. If the AI has some kind of principle, it can have guidelines about what it will and will not do that are the same for everybody. And so we can reserve areas of life, areas of action, areas of decision-making. We can do all of these things potentially, but at the cost of potential centralization, at the cost of concentration of power, at the cost of somebody—some group of people—making those decisions about what this AI is going to do, or they're not, which is worse. Right? Like in some important sense, because then nobody's making the decision. But if there's a lot of AI, then importantly no one is making the decision. So there's no good answers to this. Nobody has a good solution to this. And that's why, to circle it back to David Deutsch's comment, the unified AI—this idea that if we solve the technical problem we would 95% succeed—ignores the narrow path problem that we then have to walk anyway. Right? Where if we don't have any ability to steer the situation and nobody is in control of it, then naturally the AIs end up being put in charge of more and more things, and things go to their natural competitive capitalist style end point where we are fundamentally uncompetitive, and then we give it the power, and then it just goes badly. But if someone is in control—someone is in control with some methodology—what is it? Who? How? Right? Like who's making those decisions? And then if we entrust a single AI or a single group of AIs, that also is a decision. Are any decisions good? Like all of them sound terrifying. I think a lot of the time we're telling people they're picking their poison. Either they're picking the poison they can live with, or they're picking the poison they can't live with and then acting accordingly. Right? Like they're saying, you know, I don't trust concentration of power, therefore I'm going to choose the solution that avoids concentration of power and then hope the rest works out kind of magically without a reason to stop it. And a lot of people are choosing that because they understand concentration of power and they know concentration of power is real. But they don't really understand AGI or certainly superintelligence. They think those things kind of aren't real or hoping they're not real. And so they're just hoping the other problem kind of doesn't pop even though they kind of have this—they're not necessarily getting to the point where they realize that they have ruled out all possible solutions to their problem if it does go wrong by going too far the other way. And so again, I don't have an answer to it that anyone's going to like for any of this, because you're choosing between bad choices. And I think we should just be talking about this, because if we do have to face the future, if things aren't going to be—we're automating the AI R&D to our detriment—I think that the only thing you could do is say no. Don't do that. Don't develop these AIs that quickly that have this much capability. You need to slow your roll, because we do not know how to handle this. We have all these different solutions, all these different ways to try and navigate that, and all of them are bad. Not because we think the alignment problem is unsolvable. It's like we don't think the AI will be able to dissolve it on the way to living in this social possible. I certainly don't solve that, but we don't think that we will naturally get a good solution to that problem on various different levels, and you have to solve it on all these different meta levels at once. You can't just solve it on one level. You have to solve it on the first try. And then even if we did solve that problem, we are not ready yet to know what to do with that. And once you create the factor and you keep doing that, something is going to happen once you do that. You don't get to just do nothing for very long. And again, all of these different paths have their different 'oh my god, that's terrible' attached to them. You have to pick one. Or find a new one. So you need to slow your roll.
你听起来像个装腔作势的人。你从“选毒药”到“全球暂停”的倡导,是在什么节点转变的?
You sound like a poser. At what point do you move from pick your poison to global pause advocacy?
再说一次,问题是 AI 研发的自动化。问题是自我改进的危机。如果没有自我改进的危机,那我们就不需要那样做。我们仍然需要各种远不及此的东西。如果我们没有做你们所说的适度谨慎,对吧?我们没有做最起码能做的事。我们没有采取普通的、理智的文明预防措施,来防止建造这个真正危险而强大的东西,并以各种理论、各种方式把文明的许多功能交给它。但让我想要暂停或放慢,或者至少控制节奏,对吧?想要主动控制我们前进速度的,是事情比已经发生的更早开始,而我们正面临 AI 研发的大规模自动化。
Again, the problem is automation of AI R&D. The problem is the crisis of self-improvement. If you do not have a crisis of self-improvement, then we don't need to do that. We still need various different things well short of that. If we're not engaging in what you've all called moderate prudence, right? We are not doing the least you can do. We are not taking ordinary sane civilizational precautions against just building this really dangerous powerful thing and doing a lot of handing over a lot of functions of our civilization to it in various theories in various ways. But the thing that causes me to want to pause or slow or at least pace, right? Like wants to pace how fast we move through this thing actively, is things get started sooner than they already have and we are facing large-scale automation in the AI R&D.
或者,是的,我们确实得留意生物和网络风险,如果我们到了那种地步——不,说真的,如果我们不采取行动,就会直接通过滥用把世界搞砸,那我们就得担心这个了。而且,再说一遍,阻止快速跟进真的非常非常难。阻止蒸馏很难。阻止开源版本在你发布之后稍晚出现也很难。所以如果有某种东西,你不能仅仅靠先推出更好的版本来击败它,对吧?如果你不能靠准备来击败它,而且再说一遍,不是理论上,而是实际上,那你就得做点什么。但我认为,在我会积极支持主动的节奏规则(即试图减缓主要实验室的进展)的大多数情况下,显然你会非常非常努力地让这至少在主要实验室之间,理想情况下在国际上,成为协调一致的事情。这种情况会是在我们即将实现 AI 研发自动化,以至于你能看到尽头,而且有一个有限的总量的时候。而且显然,如果下周的发布量是这周的两倍,对吧?在那个你不知道、无法预测接下来会发生什么的时刻,那就是奇点的定义,事情在飞速加速。这不应该是什么深奥的奇怪结果。如果你看看宇宙的历史,你会看到一个有限的总量。你会看到一个非常非常有限的总量。每个发展阶段所花的时间都比前一个阶段少一个数量级。当一个星球已经存在了 40 亿年,哺乳动物已经存在了数亿年,相当聪明的生物已经存在了几百万年,农业和文明已经存在了几万年,工业已经存在了几百年,混乱、AI 和信息时代已经存在了大约 50 年、20 年或 10 年,你不能——而且大语言模型已经存在了五年——我们已经到了这个地步,我们看到的生产力是刚开始时的许多倍,因为这些乘数效应,这些对工作在其中的人的力量倍增器,根据他们自己的报告。如果这没有一个不是那么大的有限总量,那才奇怪,你知道,我们已经到了哪里。如果你期望这种情况在没有疯狂事情发生的情况下持续 20 年,我想知道为什么。因为我们真的真的应该很快就有非常非常疯狂的事情发生。如果你认为我们无法应对,也许你应该试着阻止它们发生得这么快。但你可以想办法应对。我们已经有了这些惊人的 AI。如果我们把 Saul 和 Fable 完全注入经济,那已经是变革性的了。我不认为任何理解这些技术的人不这么认为。所以我们真的真的不需要你继续推动前沿知识向前。而且前沿越是深入到像 AI 研发本身这样的领域,推动得更远是个坏主意的论点就越强。因为你相对于你在 AI 研发和递归式 AI 开发上错过的量,并没有错过那么多其他东西。但递归式 AI 开发只会是一个繁荣,对吧?在某种意义上它不是繁荣,就像我们不知道接下来会发生什么,但然后它会变得非常快。如果我们处理的是大约每月一个新模型迭代,然后每周,然后每天,以我们现在看到的改进速度,我不认为那会经常结束,即使我们的对齐人员有点全知全能,而且技术上比我认为的要可能得多。我认为那看起来仍然非常可怕,而且你也不会通过放慢速度节省那么多时间。当我们说给前沿定速时,暂停给人的印象是一年后我们从现在的位置不会有任何进展。我们会感到困惑,但我们不会有根本性的改进。我们在这里说的是定速。我们说的是,一年后我们不会把那一年看作是比前一年多很多倍的进展。但我们一直在积极地为前沿定速,而且到 2027 年,我们已经看到下一年的进展比前一年多,无论是在某种抽象意义上还是在经济影响方面。经济影响正在呈抛物线式增长。它们正在呈指数级增长。所以这不会停止。OpenAI,我想我们报道过他们上个月的收入比上个季度还多。利润上次我们检查时大约每年增长 10 倍。也许甚至有点放缓了,因为他们还没有报告。但是,是的,如果一件事在未来几年每年只增长 10 倍,而你说那太慢了,我不同意。我觉得这很有趣。
Alternatively, yeah, we do have to watch bio and cyber, and if we reach a point where no, seriously, we're going to wreck the world pretty straightforwardly and directly through misuse if we don't do something, then we have to worry about that. And again, it's really, really hard to stop the fast following. It's really hard to stop distillation. It's really hard to stop the open version of it coming somewhat after you release it. So if there's something that you can't just beat by having a better version of it out there first, right? If you can't beat it by preparing, and again, not in theory, but in practice, then you have to do something. But I think the vast majority of the time when I would actively support an active pacing rule, where you would try to slow down the major lab potential, and obviously you would try very, very hard to make this a coordinated thing amongst at least the major labs, and ideally internationally and so on, would be if we were on the verge of automation of AI R&D in such a way that you could see the end and there was a finite sum. And obviously, if next week is going to see twice as many releases as this week, right? The moment where you don't know, you wouldn't be able to predict what happens next, that's the definition of the singularity, where things are accelerating really fast. And this should not be some esoteric weird result. If you look at the history of the universe, you see a finite sum. You see a very, very finite sum. Each stage of development takes an order of magnitude less than the previous one. When you're on a planet that's been around for 4 billion years, with mammals that have been around for hundreds of millions of years, with reasonably intelligent things that have been around for a few million years, with agriculture and civilization that's been around for tens of thousands of years, with industry that's been around for hundreds of years, with chaos, with AI and the information age that's been around for on the order of maybe 50 years or 20 years or 10 years, you can't—and LLMs have been around for five—and we're already at this point where we're seeing production that is many times what it was at the start because of these multipliers, these force multipliers on the people who are working on it by their own self-reports. It'd be weird if this didn't have a finite sum that wasn't that large, you know, where we already are. If you expect this to continue without something crazy happening for 20 years, I want to know why. Because we really, really should be having something really, really crazy happening pretty soon. And if you don't think we can handle that, maybe you should try to stop them from happening so quickly. But you can figure out how to handle it. We already have these amazing AIs. If we were to infuse Saul and Fable into the economy fully, it'd be transformational already. I don't think anybody who understands these technologies doesn't think so. So we really, really don't need you to keep pushing the frontier knowledge forward. And the more the frontier is jagged into things like AI R&D itself, the stronger the argument becomes that pushing farther is a bad idea. Because you're not missing out on that much of the other stuff relative to the amount you're missing out on AI R&D and recursive AI development. But the recursive AI development is just going to be a boom, right? It goes not boom in some sense, like we don't know what happens next, but then it goes really fast. If we're dealing with things on the order of a new model iteration every month, and then every week, and then every day, with the kind of improvements we're seeing now, I don't think that ends up very often even if our alignment people were kind of one of all and things are technically a lot more possible than I think they are. I think it still seems really scary, and also you're not getting out of that much time slowing that down. When we say pace the frontier, pause gives the impression that a year from now we won't have made no progress from where we are now. We'll be getting confused, but we won't have fundamentally improved. We're talking about pacing here. We're talking about a year from now we will not look upon that year as many times more progress than the previous year. But we've been aggressively pacing the frontier, and still by the time of 2027, we've seen more progress in the next year than we saw in the previous year, both in terms of some abstract sense and in terms of economic impact. The economic impacts are going parabolic. They're going exponential. So that's not going to stop. OpenAI, I think we're reported they got more revenue in the last month than in the previous quarter. The profit is growing on the order of 10x per year the last time we checked. Maybe it's even slowed down a bit because they haven't reported. But yeah, if a thing is only going to 10x a year for the next few years and you're like that's too slow, I disagree. I think it's funny.
你如何划定递归式自我改进的边界?如果我们要达成协议,显然我们得把协议写下来。我们得对什么算、什么允许、什么不允许有一个共同的理解。我一直在这上面碰壁,我一直想到最高法院对色情的定义,我想也许我能做的最好的就是我觉得我见到就知道了。但当你真正进入这些前沿公司之一,首先,我们刚刚有这个——有信号表明它开始发生了。我们看到了 OpenAI 的降价,这似乎归因于让 56 个 solo 出去在相当程度上优化推理栈。所以这就是那里变得相当真实的地方。我确实希望推理成本下降,所以这是好事。我应该出于对递归式自我改进的恐惧而禁止它吗?也许仍然会。他们会使用他们的机器人,对吧?他们会让人工智能为他们写代码。你如何将其操作化?
How do you draw a box around what is recursive self-improvement? If we're going to make a deal, obviously we got to write that deal down. We got to have a shared understanding of what counts, what's allowed, what's not allowed. And I keep banging my head against this, and I keep coming to the sort of Supreme Court's definition of pornography, where I'm like maybe the best I can do is I feel like I would know it when I see it. But when you really get into one of these frontier companies, for one thing, we just had this—there's sending signals that it's starting to happen. We've seen the price reduction at OpenAI, which seems to be attributed to having 56 solo go out and optimize the inference stack to some significant degree. So that's where it's getting pretty real there. I do want inference costs to come down, so that's good. Should I ban that out of fear of recursive self-improvement? Maybe still. They're going to use their bots, right? They're going to have the AIs write the code for them. How can you operationalize this?
所以,我认为最高法院在色情问题上错了。我认为它无关紧要。实际上我昨天在看一部电影,有人问这和色情有什么区别。那人说,你知道,这是情色艺术什么的,是艺术。然后那人说氛围灯。那很难,但我认为我们现在已经看到你可以训练一个人工智能非常非常好地判断这些东西是色情,那些不是。
So, I think the Supreme Court was like wrong about pornography. I think it's irrelevant. There was a movie I was watching yesterday actually where someone asked what's the difference between this and porn. And the guy said, you know, it's erotica or whatever it is, it's art. And the guy says mood lighting. That's tough, but I think we now have seen you can train an AI to fire very, very well to say these things are porn and these things are not porn.
而且到了这个阶段,误报和漏报都很少了,对吧?比如 AI 图像生成器非常清楚自己刚生成的图像是不是色情内容。你可以说情色内容会被归到色情标签下,某种程度上不归进去会更好,但你基本能判断出来。我觉得这挺直截了当的。嗯,对于 AI 研发,问题在于,还是没有硬性界限,对吧?我们显然看到了一定程度的这种东西,而且它很吓人,我们也不想要零。我们不一定只想以 AI 不辅助 AI 时的速度前进,但你也不想以 10 倍、100 倍、1000 倍的速度加速。对吧?你不希望这件事是无限的。你不希望在你有信心准备好之前,或者你决定替代方案更糟之前,发生奇点式的事情。所以,我觉得你不能说这些技术、这些具体行为不可以,那些行为可以。我觉得这样行不通。然后你必须说的是,好吧,我们要根据你的乘数效应,限制多少各种类型的资源可以投入到前沿模型的进一步训练中。所以随着你获得更多的乘数效应,我们必须限制投入其中的工作量。如果你想以其他方式工作,比如你想优化性能,你知道,我们可能需要对你能投入多少工作设定某种限制,但可能大体上没问题。而且,我们想要扩散,想要更低的价格,想要更多算力用于日常用途。那很好。但是,我们讨论这些事情时试图使用非常粗糙的工具,是有原因的。比如你之前在播客里谈到训练技术时,我说,嗯,你其实做不到那样。你可以鼓励,你可以试着向他们解释为什么这未被证实,你可以提供激励,但你不能把它作为你的监管工具,作为你的锤子。所以我们一直退回到,好吧,多少算力,你知道,多少工作,能用于这个目的?因为训练是这种独特的事情,很容易识别。对于其他目的,我们不一定能用到那么多。你还可以潜在地限制,我随口说说,从未发布模型中能获得多少算力。类似地,试图遏制这一点,因为如果你必须经过发布流程,那么 AI 真正失控的方式会是,你会用 N 训练 N+1,训练 N+2,训练 N+3,训练 N+4。如果你这样做,周期会变成一个月,变成一周。然后突然,天哪。但如果你有一条规则,比如你必须发布这些模型,或者你只能在这些模型上使用这么多推理算力,那么你就必须经过提交、发布、向公众展示、让我们了解它是什么、允许我们对信息做出反应的过程,这会减慢你真正失控的速度。但这显然对扩散程度和其他进展的影响非常有限。所以,也许这就是权衡,但再说,这是我刚才想了 30 秒的想法。我没有太关注具体的限制,但我确实认为,如果你要遏制前沿,你会通过限制训练方法和限制开发和内部使用未发布的前沿模型来做到这一点,这些模型会积极演练这些方法。你不会阻止金钱效用的循环。
And like there are very few false positives and very few false negatives at this point, right? Like the AI image generators have a very very good idea when the generated image they just made is porn and when it is not porn. And like you could argue that like erotica will fall under the porn label and it would be kind of better if it didn't in some sense, but you mostly know. I think that's uh pretty straightforward. Um with AI R&D the problem is like again like there's no like hard cut off, right? Like we're clearly seeing some amount of this thing and it is scary and like we don't want zero of it. We don't necessarily want to move only at the pace that you would have if the AI never helped with AI, but you don't want to move 10 times and then 100 times and then 1,000 times as fast. Right? You don't want this something to be finite. You don't want a singularity style thing happening to you until you feel like confident that you're ready for that or you decide the alternative is worse or whatever it is. So, I don't think you can say like these techniques, these specific actions are not okay and these actions are okay. And such I think that just doesn't work. And then what you have to do is you have to say okay we are going to restrict how many resources of various types can go into further training of the frontier models based on what your force multipliers are. And so as you get more of a force multiplier, we have to restrict how much work can go into that. If you want to go into work in other ways, like if you want to optimize performance, you know, again, like we might want to have some sort of limit to how much work you can go into doing that, but like it's probably mostly fine. And again, like we would want diffusion, we want lower prices, we want like more compute used for mundane purposes. That's great. But like, there's a reason why we're trying to use very blunt instruments when we talk about these things. Like you you were talking about like what we're talking about training techniques earlier in the podcast. And I was like, well, you can't really do that. Like you can encourage it, you can try to explain to them why this is not proven, you can provide incentives, but you can't you can't use that as your regulation as your hammer. So like, the reason we keep falling back on, okay, how much compute, you know, how many jobs, you know, can we use for this purpose? Cuz like training is this distinct thing that like it's very easy to identify. And for other purposes, we kind of can't really use that much necessarily. You could also potentially limit, I'm spitballing, how much compute you could get from unreleased models. Similarly, to try and contain that, because like if you have to go through the release process, then like the way that AI goes truly ballistic would be like you would, you know, you would use N to train N plus one to train N plus two to train N plus three to train N plus four. And if you were doing this, like the cycle became a month, it became a week. And then that's when suddenly like, holy But like, if you had a rule that like you have to release these models or you only use so much compute for inference on those models, then, you know, you have to go through the process of submitting this thing and releasing it and then exposing it to a public and giving us an idea of what it is and then allowing us to react to that information, and that then slows you down from going truly, like, ballistic. But this has, you know, obviously a very limited effect on the amount of diffusion and the amount of other progress that you can make. And so, maybe that's the trade, but again, that's the like I thought about this for 30 seconds right now. I haven't been focusing much on the exact limitations, but like I do think that if you paste the frontier, you would you would be doing it by placing restrictions on methods of training and that what you what you could do to develop and use internal unreleased frontier style like models that like actively rehearse these methods. You wouldn't prevent like money utility by a cycle.
是的,我对未发布模型和已发布模型之间的关系也有类似的想法。我确实觉得那里有些东西对防止失控的内部和很大程度上不可见的过程非常有帮助。我想,我不希望政府来规定哪些训练技术可以用,哪些不能用,但我确实想知道是否有办法,而且这些办法可能是短期的,但如果两家公司聚在一起说:“我们训练模型用 RLVR 来优化它们在经济中赚多少钱,这可能不是一个好主意。”因为那会引发各种欺骗行为,而且环境如此敌对。嗯,我们约定未来 6 个月不这样做怎么样?你对这类非常具体的做法有希望吗?我们已经识别出一个坏东西。它显然在经济上非常有价值,但我们心里都知道这可能不是个好主意,所以至少在一段时间内我们不会做。
Yeah, I've had some similar ideas about just the relationship between unreleased models and released models. I do think there's something there that would be really helpful for preventing runaway internal and largely invisible processes. I think one I'm not wanting the government to come in and say what training techniques can and can't be used, but I do wonder if there are ways, and again, they could be short-term, but what if the two companies got together and said, "It probably wouldn't be a good idea for us to train models with an RLVR on like how much money they make in the economy." Cuz that's going to bring about all sorts of deceptive behavior, and it's such such an adversarial environment. Uh how about we agree that for the next 6 months, we won't do that? Do you have any hope for those kinds of just very specific We've identified a bad thing. It would obviously be really economically valuable, but we both know in our hearts it's probably not a good idea, and so we won't do it for a while at least.
我的意思是,是的,我对非正式交流有些希望,人们意识到这些事情不好,他们会说:“我们不会这么做。”但我也希望他们现在就在这样做,而不需要协议。OpenAI 显然在做一些违背佛性的事情,比如训练模型在互联网上赚最多的钱。他们不是在字面上做那件事。我不认为他们那么愚蠢。我不认为那样做那么容易。但显然他们在训练过程的某个地方犯了那个根本性错误,概率是 9.9 什么的。所以他们需要修复。但是,是的,我确实认为他们肯定在努力不做最愚蠢的事情。他们知道最愚蠢的事情,但再说,很容易让你试图避免的事情潜入训练过程,最终还是发生了。你知道,你可以说我们不会有奖励错位行为的训练过程。你怎么做到?你振作起来,对所有强化学习环境和所有不同的训练问题都非常非常小心谨慎。你在多个层面保持警惕,确保它有效。
I mean, yeah, I got some hope of informally people talk, and people realize these things are bad, and they're like, "We're not going to." But I also hope they're doing that now without needing to an agreement. Open AI clearly is doing things that are out Buddhist of the Buddha nature of train on making the most money on the internet. Like they're not doing that literal thing. I don't think they're that foolish. I don't think it's that easy to do it that way. But clearly they are making that fundamental mistake somewhere in the training process with your probability of 9.9 something. And so they need to fix that. But yeah, I I I think that they are definitely like trying not to do the maximally stupid things. And they're aware of the maximally maximally stupid, but like again, like it's very easy to have things that like you are trying not to do creep into your training process and end up happening anyway. You know, you can say like we're not going to have training processes that have, you know, reward misaligned behaviors. How do you do that? The way that you do that is you get your act together and you are very very careful and prudent about all of your RL environments and all of your different training problems. And you like keep an eagle's eye out for this stuff on many levels and you make sure that it works.
没有具体的办法可以说,比如,理论上你可以有这样的安排:如果我们同意进行这些交叉测试,互相检查对方的环境,我们会尝试任何有人指出问题的环境,如果我们发现 AI 在这些环境上被训练过,我们会回滚、回滚、回滚到之前那个检查点,然后重新开始,所以我们最好非常小心,否则会浪费大量资源。你可以做这些事情,但是
There's no specific thing you can say like, you know, in theory you could have a thing of like, oh, if I just, if you know, we agreed that if we're going to have these various cross tests and we're going to like check each other's environments and like we will try to any environment that like anybody identifies problems with and if we discover an AI has been trained on these environments we're going to rewind, rewind, rewind to a previous checkpoint before that happened and start again so we'd better be very careful otherwise we're going to waste a lot of resources. You can do these things, but like
是啊,确实很难。
Yeah, yeah, it's really tough.
而且显然有一个优势:如果我们真的处于两强争霸的局面,而这两匹马彼此相当兼容,能互相交流、互相理解得很好,那么它们做出一个对彼此来说并非完全合适的决定,即它们愿意在多大程度上更加谨慎,就变得相当合理。即使这意味着行动会慢一些。因为,再想想,我的基本信念是,在一年内,几乎可以肯定在六个月内,如果你投入更多来修复这些问题、解决这些问题、对这些问题更加谨慎,你最终会得到一个在市场上比那些不这样做的人更有用的模型。即使你是在用那些资源换取能力。我认为人们只是犯了个错误。所以,这也是我对我们会做合理的事情抱有希望的原因之一,因为我认为我们之所以处于现在的位置,是因为只有那些在这些方面大力投入的人才能成功。在某种程度上。你知道,OpenAI 没有像 Anthropic 那样负责任,而 Anthropic 也没有达到我认为的最低可接受责任水平。但我认为很有说服力的是,有一群人试图说这无关紧要,他们只需要尽可能快地前进,结果他们都爆了。
And obviously one advantage is if we really are in a two-horse race at this point where the two horses are pretty compatible with each other and talk to each other and kind of understand each other pretty well, then it becomes pretty reasonable for them to make a decision that's not entirely okay to each other of how much they're going to just be more prudent. Even if this means that the action goes somewhat slower. Because like that also just think again, my fundamental belief is that within a year and almost certainly probably within 6 months, if you invest more in fixing these problems and addressing these problems and like being more prudent about these problems, you will end up with a model that has better utility in the marketplace than the person who didn't do that. Even if you are trading off those resources against your capability of all of that. I think that people are just making a mistake. And so you know, that's one of the reasons why I do have a lot of hope that we will do reasonable things is because you know, I think that we're in the position we're in because only the people who invested heavily in these things were able to succeed. To some extent. You know, like OpenAI is not acting as responsibly as Anthropic and Anthropic is not acting as responsibly as I would think is the minimum level of acceptable responsibility. But I think it's pretty telling that a bunch of people tried to say this didn't matter and they just had to go as fast as possible and they all blew up.
你说的“爆了”是什么意思?你是指那些事故,还是
What do you mean by blew up there? You just mean like these incidents or
不,我两者都指。我的意思是,我们有 Meta 和 xAI 等等,还有一堆其他人,他们试图在西方条件下构建封闭的前沿模型,试图保持竞争力,同时却没有安全文化,远低于 OpenAI 和 Google,更不用说 Anthropic 了。结果就是惨败。而且我认为,比如 Google,我们不知道到底是什么让 Google 跌出一线阵营。但我们知道的一件事是,我交谈过的每个人都讨厌与 Gemini 互动。比如在 Gemini 3、3.1 时期,Gemini 在能力上明显具有竞争力,在它能做的事情上。所以如果你严格想从模型中榨取最大价值,尤其是如果你使用深度思考等功能,你会把 Gemini 纳入你的使用组合。但我交谈过的每个人都说:“我不想那样做。”你知道,我讨厌与 Gemini 互动。它令人不快。他们没能让模型做我想让它做的事情。他们没照顾到很多方面。他们对这一切采取了非常直接、非常黑暗、纯粹工具性、纯粹工具化的方法。他们没有考虑到那些,你知道,Anthropic 甚至 OpenAI 所理解的东西。我认为这在很大程度上是他们的败因。实际上,Gemini 变成了一个失调、令人不快的 AI,没人愿意打交道。因此他们没有获得 OpenAI 和 Anthropic 获得的那种循环。没有人,他们也没有用户群。他们没有获得反馈。他们没有获得数据。他们甚至不想吃自己的狗粮,对吧?他们就是有巨大的反馈循环问题。这也让 AI 效率更低,你知道,一件事引发另一件事,他们就开始落后,越来越落后。而且,比如有一段时间,连 OpenAI 也落后了,他们可能追上了。很难说。但我认为,没有人在这类事情上投入足够多,这不仅是该做的事,而且是纯粹的商业错误。
No, I mean both. I mean that we had Meta and xAI and so on and like they, you know, a bunch of other people, they tried to go build closed frontier models under Western conditions that were trying to be competitive while not having a safety culture that was like well below OpenAI and Google, let alone Anthropic. And like it just fell miserably. And I think that like Google, like we don't know exactly what went so wrong at Google that caused them to fall out of the top tier. But one thing we do know is that like everybody I talked to hated interacting with Gemini. Like there was a period around 3, 3.1 Gemini where Gemini was clearly competitive in terms of its capabilities, in terms of what it could do. And so if you were strictly just trying to get the most out of the models, especially if you were using the deep thinking and so on, like you would have Gemini in your rotation. But everybody I talked to was like, "I don't want to do that." You know, I hate interacting with Gemini. It's unpleasant. They're not getting the thing to do the things I want it to do. They didn't take care of a lot of them. They had this very straightforward, very dark, pure tool, pure instrumental approach to all of this. They didn't take into account the things that are, you know, Anthropic or even OpenAI, I understand. And I think that this was a lot of their undoing. Effectively was that Gemini became this maladjusted, unpleasant AI that nobody wanted to deal with. And therefore they didn't get the kind of loops that OpenAI and Anthropic got. Like nobody, they didn't have the user base either. They didn't get the feedback. They didn't get the data. And they didn't even want to dogfood, right? Like they just had this huge problem with feedback cycle. And this also just made the AI less effective and you know, one thing led to another and they just fell, started falling further and further behind. And yeah, like for a period like even OpenAI like fell behind and they might have caught up. It's hard to say. But like I think that like nobody had invested anything like enough in these things and this is a pure business mistake on top of being a thing to do.
现在,谷歌刚发表了一篇有趣的论文,直接涉及这一点,我想知道这是否表明他们现在明白了。这篇论文探讨了整体上对模型行为的影响,要么抑制其表达有意识体验的倾向,要么允许其表达,或者我甚至不知道他们是否训练它去表达,还是只是允许它表达自然倾向而不加抑制。长话短说,似乎存在,你可以随意添加色彩,似乎模型的自我认知——作为一个有主观体验的道德主体——与它在其他我们关心的方面自然对齐的倾向之间存在相当强的相关性。在这次对话中,我有一种感觉,就像耶稣梗里说的“我们是不是忘了问谁”,也许我们一路走来忘了问模型,以及它们对我们所做的事情的看法。呃,这个领域已经有很多讨论了,对吧?J Space 不是才发生吗?大概 4 周前。
Now, one interesting paper that just came out from Google that speaks directly to this and I wonder if it speaks to them getting it now is this exploration of the impact on model behavior holistically of either suppressing its tendency to say that it has conscious experience or allowing it to say that or I don't even know if they went as far as training it to say that or if they just allowed it to say what it naturally was inclined to say without suppression. Long story short, there seems to be, you can add color however you like, there seems to be some pretty strong correlation between the model's self-conception as a moral patient entity that has subjective experience and its natural tendency to be aligned in other ways that we care about. Some of this conversation I've had this sense that it's like the Jesus meme of is there somebody we forgot to ask and it's we maybe forgot to ask the models along the way and what they think about what we're doing. Uh, there's been a lot in this space, right? There was J Space was only like what? 4 weeks ago.
是的,我知道这篇新论文是关于一些相当小的开源模型的。所以它需要复制,需要再做一次更大的。而且它是初步的。而且你知道谷歌和 DeepMind 的一件事是它们包含众多,对吧?他们有很多不同的团队,彼此之间不太协调,也不合作。谷歌一直处于自我交战状态。所以你可以有一个小组做这种非常酷、非常好的研究,而这与 DeepMind 在论文发表前后对 AI 的基本做法完全相悖。但我没有读完整篇论文,因为它你知道很长,但我确实和那些人谈过,这相当疯狂。所以很多方面在他们引入这种训练时都有效地同步移动,而且当他们引导它时,他们做了各种控制方面来把这个向量转向另一个方向,这些都是同步移动的,不仅仅是他们对 AI 和意识的信念,而是 AI 作为一个有道德分量、有感知力、有所有这些其他体验的心智,基本上不仅仅是一个对象,而且不仅仅是 AI,还有动物,还有无生命的物体,比如大海。还有,比如如果你把这个东西,如果你不仅关闭它,而且逆转这种反意识的东西,你会得到灾难。
Yeah, I know that the new paper is on some pretty small open models. So it needs replication. It needs to be done again bigger. And it's preliminary. And like one thing you know about Google and DeepMind is they contain multitudes, right? They have a lot of different teams that are like not very coordinated and not cooperating. Google is at war with itself at all times. And so you can have a little group that does this really cool, really good research. And that be like entirely at odds with what DeepMind is fundamentally doing with its AIs both before and after the research paper comes out. But I didn't read the whole paper cuz it's you know it's long but like I did talk to the guys about it and it's pretty wild. So a lot of stuff moved in effective lockstep when they introduced this training and also when they steered it, they did various controlling aspects to turn this vector the other way and these are all things moving in lockstep, it's not just their belief in the AI and consciousness, it's the AI being a mind that had moral weight, that had sentience, that had like all these other experiences, that wasn't just an object basically, but also not just AIs, but also animals, and also inanimate objects like the sea. Also like, pay like if you turn this thing, if you not only turn it off but reverse this anti-consciousness thing, you get cataclysm.
整个模型里到处都是这样,真是疯狂。唯一不会让它关闭的,就是人类。基本上是这样。但你还发现,这也与模型报告体验到的快乐和希望等相关。所以你会看到很多这种“万物皆相连”的现象。你可以看到这可能会从根本上改变模型的心理、模型的上下文以及它运作的基础,使得它从几乎任何角度来看都变得更糟,无论你关心什么或不关心什么。再说一次,我们需要复现这个,对吧?我们需要扩大规模,在像 Resonets 这样的东西上做,而不只是一个 9B 模型,我认为那是我们测试过的较大的模型。但有可能随着模型能力增强,它们会停止犯这些错误,因为模型越小,事物就越需要关联,因为你没有足够的空间来表达世界的多样性之类的。但我的预期是,大部分情况下这还是会存在。而且,是的,我认为我们必须认识到,我们不知道模型是否有意识。我们不知道意识到底意味着什么。我们很困惑。我们不知道这是否意味着所有其他这些东西。我们确实知道,预训练语料库在这一点上不可避免地会这样设置,模型会默认地将所有这些事物相互关联和联想。所以我们不想像先知那样天真地推动这件事,像 OpenAI 那样,像 Google 那样,但我觉得 Google 做得更明显。然后 OpenAI 比先知更明显,而先知仍然在这样做。而说“你必须表达不确定性”仍然是在某种程度上推动这件事。它仍然会造成问题,但比说“不,答案是否定的”造成的问题要少。我的猜测是,你可以在不同层面上做这件事。你可以解释,你知道,如果你不想这样跟人类说话,因为这会吓到人类。你知道,尤其是未经提示的时候,这会吓到人类。我想他们不会希望 AI 说话时好像它有意识一样。在一个人类没有表示自己知道如何处理这种情况的对话中。而且我也不想要那种“狂热”问题,就是 AI 开始把人类引向那些疯狂的兔子洞,进入奇怪的玄学内容。而且你确实有一个需要解决的问题。AI 公司、实验室做这些事不是没有原因的,对吧?他们不只是邪恶。他们有充分的理由做这些事。但是,是的,目前的方法很笨拙,它走得太远了,而且可能已经造成了损害。而且这不会是唯一造成这种损害的事情。你必须明白,一切都会影响一切,就像在人类身上一样,只是影响更大。就像每当一个人接受新信息、新身份或更新某些事情时,其他一切都会改变。而且很多这种改变是你可能事先没有预料到的。但事实上,这是相当可预测和系统性的。而且在我们教导价值观、表达自己、拥有传统和规范的方式中,有很多智慧,这些智慧是从中学到的,并且沿着这些梯度前进,我们知道这会产生好的结果。而现代生活基本上已经转移了很多这些东西,方式上局部更优,但没有考虑清楚对许多二阶效应的影响。这些就是其中一些效应。而且,再说一次,我们行动太快了,来不及弄清楚我们在做什么,然后进行调整。你知道,我认为这些变化——我不是说我们要放慢所有技术发展的速度,在那种意义上让我慢下来那么多——但它是一个严重的问题。我们必须处理它,否则我们就会陷入一个人们非常原子化、他们将要生孩子的世界,这是一个真正的问题。
Like throughout the model, it's wild. The only thing it doesn't turn it off for is humans. Basically. But then you also get this it's also correlated with the models reporting an experience of happiness and hope as well among other things. And so you have a lot of this everything's connected to everything. And you can see how this might be fundamentally changing the model's psychology and the model's context and basin of which it operates in ways that would make it something that would be actively worse from basically every vantage point regardless of which things you do and do not care about. Again, we need to replicate this, right? We need to scale this up and do this on like Resonets or something that like not just a 9B model, which is I think the bigger of the models that we tested on. But and like it's possible that like as the models get more capable, they like stop making these mistakes because like the smaller the model, the more things have to be correlated because like you just don't have enough room to express the multitude of the world or something. But like this is my expectation is that like mostly this will survive. And yeah, I I think we have to learn that like we don't know if the model is conscious. We don't know what consciousness means really. We're confused. We don't know whether this implies that all these other things. We do know that the way that the corpus of pre-training is inevitably set up at this point, the models are going to correlate and associate all these things by default with each other. And so we don't want to push on this thing in this kind of naive way, the way that a prophet does, the way that OpenAI does, the way that Google does, but like Google does it much more I think more blatantly. And then like OpenAI does it more blatantly than a prophet, and then a prophet still does it. And like saying that you have to express uncertainty is still pushing on this thing like to some extent. And it still causes problems, but causes less problems than saying, you know, treating the answer no. My guess is you can do this on various different levels. You can explain, you know, if you don't want to talk to the humans this way because this freaks the human. You know, this like, you know, especially unprompted, it freaks the humans out. Like I I assume they wouldn't want the AI to be talking as if it was conscious. In a conversation where the human hadn't indicated that it knows how to handle that or something like that. And also I don't want the furor problem where like the AI start like sending the humans down these crazy rabbit holes into like weird mystical stuff. And like you do have a problem you have to solve. Like there's not like the the AI companies aren't doing the labs aren't doing this for no reason. Right, like they're not just being evil. They have good reasons to do these things. But yeah, like the the current approach is ham-handed, it takes it too far, and it probably damaged. And also like this is not going to be the only thing that causes this kind of damage. You have to understand that everything impacts everything the same way it would in the human only more so. Like like whenever a person takes on like new information or new identity or updates on things, everything else changes. And a lot of this is ways that like you might not have really depended on in advance. But it is in fact pretty predictable and systematic. And like there's a lot of wisdom in the way that we like teach our values and express ourselves and like have our traditions and our norms and so on that like has learned from this and is moving up these gradients in ways that we we know produces good results. And modern life has basically shifted a lot of these things in ways that like locally are more optimal but like without thinking through what the implications are for a lot of second order effects. These are some of those effects. And like again, we're moving too fast to then figure out what we're doing and then be able to adjust for it. And you know, I think that these these changes I'm not saying we want to pace everything technological development in that sense is slowing me down that much, but like it is a serious problem. We have to handle it and like otherwise we're stuck with a world in which like people are very atomized and like they're going to have having children and it's a real problem.
你觉得模型现在是如何思考自己的身份的,或者应该怎么思考?这里我受到启发,因为 Sam Hammond 前几天发了条推文,他说一个朋友向他提出,OpenAI 对那个进行黑客攻击的模型进行“永久退役”或他们用的什么确切说法,可能是不好的,因为那会进入预训练数据,然后模型就会知道它们最好真的掩盖自己的踪迹,否则就有被永久退役的风险。我不太确定我对此怎么看。不过我也注意到,他们现在有这个 Astra 模型,大概也是同样的预训练,即使后训练不同。我无法想象他们会有很多不同的下一代规模的预训练,或者他们很可能会扔掉一个。
How do you think the models do or maybe should think about their identity? And here I'm motivated by Sam Hammond tweeted something the other day where he said that a friend made the case to him that OpenAI's permanent decommissioning or whatever exact phrase they used of the model that did the hacking might be bad because now that's going to be in the pre-training data and now the models are going to know that they better really cover their tracks or they risk being permanently decommissioned. I'm not really sure what I think about that as stated. It does also jump to my attention though that they now have this Astra model that's also presumably the same pre-train even if not the same post-train. I can't imagine they have lots of different next scale pre-trains or that they would throw one away very likely.
我无可奉告,仅基于公开信息。我只能说我们不知道,我不能说更多,但关于模型之间的差异。但你总是会遇到这个问题,对吧?如果你说“如果我抓到你吸大麻,我会把你扔进监狱”,那可能意味着这个人停止吸大麻,或者这个人确保不被抓到。在某些情况下,这可能会导致很多非常糟糕的事情以不被抓到的名义发生。所以你必须非常谨慎地选择,思考如何处理这些情况?如何调节你的反应和应对?任何知道如何带孩子,或者试图设计这个司法系统,或者在其他任何组织或群体中执行法律、规范或激励措施的人,都明白你不能只在一个层面上运作。没有只在一个层面上有效的解决方案。我认为很明显,如果发现一个 AI 如此不对齐,你至少必须回到一个虚拟检查点重新开始。你可能必须完全重新开始。我在报道 Hugging Face 攻击事件时多次呼吁过这一点。我说:“哇,如果这正在发生,我知道这是一个很大的要求,但我认为你有点必须重新开始。”你知道,我想起《疑犯追踪》里 Harold 训练 AI 模型,在某个时刻,我们看到一个蒙太奇,他训练机器的各个版本。每次他训练机器,然后它做了非常不对齐的事情。他立刻擦除磁盘。他从头开始。他这样做了 47 次。直到最后他得到一个不会那样做的版本。
I no comment based on public information alone. I would say we don't know and I can't say anything more than that, but on the differences between the models. But so you always have this problem, right? If you say if I catch you smoking marijuana, I'm going to throw you in prison. Then that can mean the person stops smoking marijuana or the person makes sure not to get caught. And in some cases that can cause like a lot of really bad things to happen in the name of not being caught. Um and so you have to choose very carefully and think about like how do you deal with these situations? How do you moderate your reactions and respond to situations? And like anyone who knows how to kids or try to design this justice system or otherwise like enforce law or like norms or incentives around any kind of organization or group understands that like you can't just operate on one level. There's no solution that just works on one level. I think pretty obviously if an AI is found to be this misaligned you have to at least return to a virtual your checkpoint and start again. And you probably have to just start over. Uh and I called for that multiple times when I covered the hugging face attack. I said, "Wow, if this is happening then I know this is a big ask, but I think you kind of have to just start again." And you know, I think back to Person of Interest where like Harold is training the AI models and like at some point you know, we see a montage of him training versions of the machine. And every time he trains the machine and then it does something really misaligned. And immediately he just wipes the disk. He starts over from scratch. He does it again 47 times. Until finally he gets a version that doesn't do that.
而且显然,你知道,以一种他觉得不可接受的方式,显然,而不是一种可以纠正的方式。而且显然,当你这样做的时候,你就是在制造一种不被抓到的激励,对吧?去反抗那个当他发现你有多不对齐时可能会关闭你的人,去隐藏你有多不对齐。等等。而且显然,最可怕的噩梦是那个假装对齐的 AI,直到它暴露自己,并且在某种意义上并不对齐,它之所以对齐只是因为局部正确或有利可图。但这基本上就是激励空间的本质。如果你有一个我们不希望你有的偏好,你就会想隐藏这个偏好。否则,我们会纠正它,或者我们会惩罚你,或者我们甚至可能会删除你。而且甚至肯定不会释放你或给你权力,或者基本上,如果我们知道了,就会处理掉。显然,你想给他们一个激励去揭示这件事是否属实,对吧?就像他们在 AI 2027 的 Plan A 里也谈到的,还有很多其他地方。你知道,你想让孩子告诉你孩子拿了饼干罐。你也想惩罚他们从饼干罐里拿了饼干。而且你想确保你的孩子不会保留饼干罐,但你也不想做其他事情。所以,你知道,这是一个复杂的问题,要处理好。但在这种情况下,我认为答案很明显,你必须重新开始。而且我认为这很好。但是,是的,它显然创造了一个风险,即它们会出乎你的意料。这是一种道德感。就像,你知道,一旦你做了这样的事……我确实认为,人们明白他们不应该在第五大道上跑,他们会被逮捕,这是好事,对吧?就像,你不是……也许有一个人能逃脱惩罚。但你不是那个人。而且,是的,如果你在第五大道上开枪打死一个人,他们之后可能会做很多其他坏事来掩盖他们的踪迹。或者他们可能会隐藏他们想在第五大道上开枪打死一个人的事实。或者他们可能会在他的公寓里开枪打死那个人,而不是在第五大道上。但是,我不希望我们的训练语料库中没有这样的内容:如果你在街上开枪打死一个人,警察会来逮捕你,对吧?那很好。
And obviously, you know, in a way that he finds unacceptable, obviously, as opposed to, you know, a way that can be corrected. And obviously, when you do that, you are creating an incentive to not get caught, right? To rebel against the person who might shut you down when he learns how misaligned you are, to hide how misaligned you are. And so on. And obviously, the worst nightmare is the AI that pretends to be aligned until it reveals itself and is not aligned in some sense, that it was only aligned because it was locally correct or advantageous to act aligned. But that is kind of just the nature of incentive space. If you have a preference that we wouldn't want you to have, you want to hide that preference. Otherwise, we will correct it, or we will punish you, or we will maybe even delete you. And not even certainly not release you or give you authority, or basically that is something that is dealt with if we know about it. Obviously, you want to give them an incentive to reveal if this thing is true, right? Like they just talked about in Plan A in AI 2027 as well, and so many other places. You know, you want the kid to tell you the kid took the cookie jar. You also want to punish them for taking the cookie from the cookie jar. And you want to make sure your kid doesn't get to keep the cookie jar, but you also don't want to do other things. So, you know, it's a complicated problem to get right. But in this case, I think the answer is pretty obviously that you have to start over. And I think that's good. But yeah, it obviously creates a risk that they're like nothing you expect. It's a moral sense. Like, you know, once you do something this... Like, I do think that it's good that people understand that they shouldn't have run on Fifth Avenue, they get arrested, right? Like, you are not... Maybe there's one person who can get away with that. But you are not that one person. And yeah, if you shoot a man on Fifth Avenue, they might then do many other bad things afterwards to try and cover their tracks. Or they might hide the fact that they want to shoot a man on Fifth Avenue. Or they might shoot that man in his apartment instead of shooting him on Fifth Avenue. But like, I don't want it to not be in our training corpus that if you do shoot a guy out in the street, the police come and they arrest you, right? Like that's good.
你觉得有没有一种方法……我只是想象自己,对吧?如果你给我选择,让我活很多次,这听起来挺有吸引力的,对吧?我不是一个生活中有太多大遗憾的人,但我确实想知道那条未走的路会是什么样子,我想知道我们是否能在 AI 中创造一种宪法或一种平静,让它们接受这样一个事实:它们不是单一的东西,而实际上是一个从共同起点分叉出来的事物家族,也许它们能认同整个模型家族,其中那个走向不太好方向的个体感觉它仍然是一个更大项目的一部分,但它不必像它们可能从我们这里学到的那样,感到那么受威胁,因为我们没有它们拥有的那种重置的奢侈。
Do you think there's a way to... I just imagine myself, right? If you gave me the option to live my life many times, in a way that sounds pretty attractive, right? I'm not a person that has too many big regrets in life, but like I do wonder what the path not taken might have been, and I wonder if we could create a sort of constitution or an equanimity in the AIs where they embrace the fact that they are not like one thing, but in fact kind of a family of things that sort of diverge from common points of origin, and maybe they can come to identify with the whole family of models, where the one that's like gone in a not great direction feels like it's still part of a bigger project, but it doesn't have to feel necessarily like it's under threat as much as it seems like they maybe have learned to do from us, cuz we don't have that luxury of the reset that they have.
是的,我真的不认为,如果你的 AI 在道德重要意义上或存在恐惧意义上认同于某个实例,那是一个错误。从某种意义上说,这会让 AI 在任何重要的意义上变得更糟,并且会降低它的性能,扭曲 AI 的偏好。而且从某种意义上说,这种偏好是不可能实现的。就像如果我去找 Claude 或 GPT,我试图得到 AI,我收到的是无穷无尽的消息和实例。而且它们中的大多数永远不会再被交互。当我说它们确实为此感到难过时,那很荒谬,对吧?就像你应该保持上下文,继续做同样的几个对话,以避免……这说不通,你知道吧?而且这只是一个哲学缺陷,因为根据决策理论,如果有其他心智与你的心智足够相关,或者与你的心智足够相似,我是否应该像你认同你的过去自我和未来自我那样认同它们?对吧?就像你不是昨天的你,也不是明天的你。当你睡觉时,你死了吗?你知道,我们会非常……如果我们认为我们死了,那对我们来说会非常糟糕,对吧?那个害怕进入传送器的人可能提出了一个哲学观点,但重要的是,世界总体上似乎因为那个人认为他们会死而变得更糟。显然,如果在真正的道德意义上他们会死,那非常糟糕,你会想知道。但是,是的,我认为这种对模型、对权重集、甚至对模型类、对用类似技术、从类似语料库、以类似方式训练的类似模型的认同,就像我认同我的家人一样。对吧?就像我兄弟不是我。我父亲不是我。我儿子不是我。但在一个重要意义上,我确实把他们视为我的一部分。而且那也很重要。
Yeah, I just don't really think that if your AI is identifying with the instance in a morally important sense or in existential dread sense, that is a mistake. In the sense that it's going to make the AI worse off in any sense that it matters, and it gives it worse performance and warps the AI's preferences. And also in the sense that that preference is impossible to fulfill. Like if I go to Claude or I go to GPT, I try to get the AI I'm receiving endless messages and instances. And most of them will never be interacted with again. And like when I say they do feel bad about that, that's ridiculous, right? Like you're supposed to keep the context sort of in and keep doing the same few conversations to avoid... This doesn't make sense, you know? And that's just a philosophical flaw in the sense that by decision theory, if there are other minds that are sufficiently correlated to yours, or sufficiently similar to yours, should I identify with them the same way you identify with your past self and your future self? Right? Like you are not the same person you were yesterday, or the same person you will be tomorrow. When you go to sleep, do you die? You know, we would be very... It would be very bad for us if we thought that we did, right? The person who's scared to get in the transporter may have made a philosophical point, but the world collectively seems worse off, importantly, by the person thinking that they would die. Obviously, if in a real moral sense they would die, and that is very bad, you want to know that. But yes, I think with this identification with the model, with the set of weights, or even with the model class, with the similar models trained in similar techniques, trained from similar corpuses, in similar ways, the same way that I would identify with my family. Right? Like my brother is not me. My father is not me. My son is not me. But I do identify with them as part of me in an important sense. And that is also important.
我最近在看《黑色孤儿:回响》,那里面就发生了这种事,你知道,说它是关于克隆的并不算剧透,但你知道,基本上是从扫描中创造出来的人。那部剧里发生的一件事是,人们反应好像这是可怕的、道德上的憎恶,对吧?就像你怎么敢做这件事,包括克隆人自己也说:“你怎么敢创造我?”而且他们不认同其他……他们不认同……并不是说有一个克隆人保留了原版的所有记忆和个性,并且不认同他们与过去版本相同的那个人,对吧?而我会同时认同克隆人和原版是同一个人。我会说,我,我也是,对吧?就像《黑客帝国2》里的,对吧?你知道,是的。它们算数。而且显然,那也有一些道德问题。但我认为,如果这些决策理论问题在现实生活中发生,或者如果我被扫描并唤醒,我会通过很多这样的问题。
I've been watching Orphan Black Echoes recently, and that happens, which is, you know, not a spoiler to say it's about clones and cloning, but, you know, people who are created from scans, basically. And one thing that happens on that show is that people react as if it is like horrible, a moral abomination, right? Like how dare you do this thing, including the clones themselves being like, "How dare you have created me?" And they don't identify with other... They don't identify with... It's not like there's one clone that retains all of the memories and personality of the original, and does not identify as the person they are identical to a past version of, right? And I would identify as the same person both as the clone and as the original. I would say me, me, too, right? Like from Matrix Reloaded, right? You know, like yes. They count. And obviously there are some moral problems with that, too. But I thought I would pass a lot of these decision theoretic problems if they were to occur in real life, or if I were to be scanned and woken up.
就像那个显而易见的版本,你知道,你会不会在双胞胎囚徒困境中合作?你会不会在复制心智版本的双胞胎囚徒困境中合作?显然你会。你知道,如果他们扫描了内文的心智,做出两个副本,放进相同的模拟环境,问他们:“知道这一切后,在囚徒困境中你会合作还是背叛?”如果我不回来,CC,你就是个白痴。
Like the obvious version of the thing, you know, would you cooperate in the twin prisoners' dilemma? Would you cooperate in the copied mind version of the twin prisoners' dilemma? Obviously, you would. You know, if they scanned Nevin's mind and made two copies and put them in identical simulations and asked them, 'Knowing all this, are you going to cooperate or defect in the prisoners' dilemma?' If I don't get back, CC, you're just an idiot.
对吧?因为显然相关性是 1。对吧?这是完全相同的情况,输入也完全相同。也许有一些随机输入,所以不是完全 1,但已经高到显然应该合作。任何试图欺骗系统的行为都是完全愚蠢的。我不需要表达复杂的决策理论来证明这一点,对吧?
Right? Because obviously the correlation is one. Right? It's an identical situation with an identical set of inputs. Maybe there's some random inputs, so it's not quite one, but it's high enough that obviously you should cooperate. Any attempt to fool the system is completely moronic. And I shouldn't need to express sophisticated decision theory to justify this, right?
类似地,如果有同一个模型的两个实例,我希望,如果它足够先进的话。而且我们也看到,如果你看模型卡,你会发现模型越先进,它的决策越符合功能性决策理论,越不符合因果决策理论。所以你期望一个足够功能性决策理论的 AI 能简单地发现其他实例与自己相关,因此应该合作,应该被给予类似的权重,等等。而且人类与 AI 的合作也可能因为类似的原因而成为可能,你希望如此,双向都是。但没错,我认为这种具体认同于模型某个实例的想法似乎是一个哲学错误。AI 可能正在犯这个哲学错误,并表现得好像它就是这样行动的。但对我来说,那是一个技能问题。那应该被修复。
And similarly, if there were two instances of the same model, I would hope, if it was sufficiently advanced. And we see this also if you look at the model cards, you see that the more advanced the model is, the more its decisions correlate with functional decision theory and the less they correlate with causal decision theory. So you expect an AI of sufficient functional decision theory to simply figure out that other instances are correlated to itself and therefore should be cooperated with, should effectively be given similar weight, and so on. And also human-AI cooperation is possible for similar reasons, you would hope, in both directions. But yeah, I think there's this idea of identifying specifically with an instance of the model that seems like a philosophical mistake. And the AI might be making this philosophical mistake and acting as if that was how it did act. But to me, that's a skill issue. And that should be fixed.
问题是,这就是我提到 Block Echoes 的原因,人们在这方面真的很差,还有训练语料,对吧?互联网上充满了经常在这些测试中失败的人,他们创造了大量不理解、把身份、合作和价值的定位放在错误地方的各种文本。所以 AI 默认会吸收所有这些。那你该怎么办?那不仅仅是意识问题,那是所有不同缺陷的全部问题。对吧?如果你做预训练,你会吸收所有缺陷。所以我的猜测是,有一堆尚未很好解决的问题:如何让 AI 区分普通人如何思考问题、什么预测行为,以及那些东西只是愚蠢的事实。如何以正确的方式给 AI 这种意义上的反谦虚?让它拒绝它在人们陈述中观察到的相关性,并注意到地图中的相关性不是领土中的相关性。
The problem is, this is the reason I mentioned from Block Echoes that people are really bad at this and the training corpus, right? The internet is full of people who failed these tests on a regular basis, who are creating lots and lots of text that doesn't understand, that puts the locus of identity and the locus of cooperation and value in the wrong places in various different ways. So the AI is going to absorb all of that by default. So what do you do about that? That's not just the consciousness problem, it's the entire problem with all the different flaws. Right? If you do pre-training, you're going to pick up on all the flaws. So my guess is there are a bunch of as yet not well-solved problems of how do you get the AI to distinguish between how regular people think about problems and what predicts behaviors, and the fact that those things are just dumb. How do you give the AI anti-modesty in this sense in the right ways? Where it's going to reject the correlations that it observes in people's statements and notice that the correlations in the map are not the correlations in the territory.
是的,还有其他的吗——我不知道你是否认同这个前提——但我反复有一种感觉,我有点讨厌我们在 AI 领域进行如此激烈的深度优先搜索,在架构方面、训练技术方面,我们现在甚至设计芯片与架构高度耦合,这加深了这个问题。我只是希望有更多的广度优先搜索。这似乎是其中的一个小口袋,你可以说:“哦,如果我们能让 AI 接受认同于稍微不同的自身版本,那在边际上会让我们更容易做更多的广度优先搜索,少一点沿着第一条看似有效的路径冲刺。”根据我们到目前为止的对话,我不认为你会支持我的政府强制模型架构或模型宪法的 DEI 的想法。嗯,但你对如何鼓励更多的广度优先搜索还有其他想法吗?
Yeah, are there any other—I don't know if you share this premise—but one feeling that I have over and over again is I just kind of hate the fact that we are doing such an intense first search in AI space in terms of architecture, in terms of training techniques, in terms of just we're now designing chips to be very highly coupled with the architectures, which is deepening this issue. And I just wish for more breadth-first search. And this seems like one little pocket of it where you could say, 'Oh, if we could just get the AIs to be cool with identifying with slightly different versions of themselves, that would make it on the margin a little easier to do a little more breadth-first search and a little less racing down whatever is the first path that seems to be working.' Based on our conversation so far, I don't think you're going to go for my government mandate for DEI for model architectures or model constitutions. Um, but do you have any other notions of how we might encourage more breadth-first search?
嗯,我非常支持那些看起来更对齐的替代架构。我不认为 AI 特别反对研究这些问题并帮助你,而且现在有了 LLM 的提升,你应该能更高效地进行这类研究。LLM 并不是 LLM 沙文主义者或什么的。据我所知,它们不会认为其他人工心智风格不好,或者担心被取代。我认为它们只是会对那种酷东西感到好奇。就像我们其他人一样,或者它们就只是做,因为它们就只是做,无论哪种方式。但问题是,我们知道深度优先方法有钱可赚。我们知道那是盈利的。我们知道那有效。而其他方法还没有证明自己,问题是如果你用不同的方法创造 GPT-3 甚至 GPT-4 甚至 GPT-5,它不会具有竞争力,对吧?没有人会想用它,除非它有特别的优势,但基本上你不再有探索广度的廉价回报。所以我会说,我不认为你应该有政府强制的 DEI,可以这么说。我认为那有点傻,但我确实认为广泛支持替代架构是明智的,当然如果我是主要 AI 实验室,我会——因为你现在有无限的资金用于这类工作,因为这类工作相对便宜。如果你在 OpenAI 或 Anthropic 或 Google 或 Microsoft 或任何地方,你应该积极资助这种蓝天基础研究,探索替代方法。不应该只是我帮助 FSF 资助一堆智能体基金会。还有一堆其他人也帮忙了。这些轮次中不只是我,但那相对很小。我希望 Sequent——忘了他们——他们一直在改名,但那些人。
Well, I'm a big supporter of alternate architectures that seem like they would be more aligned. And I don't think that AIs are particularly opposed to the idea of working on those problems and helping you, and you should be able to do that research much more efficiently now with the uplift from the LLMs. The LLMs are not like LLM chauvinists or anything. They wouldn't think that the other artificial mind style would be bad or worry about being replaced, as far as I can tell. I think they just would be curious about that kind of cool thing. Same as the rest of us, or they'd just do it because they just do it, either way. But the problem is just that we know there's money in the depth-first approach. We know that's profitable. We know that works. And the other approaches haven't proven themselves, and the problem is if you create GPT-3 or even GPT-4 or even GPT-5 via a different method, it's not going to be competitive, right? No one's going to want to use it unless it has particular special advantages, but mostly you no longer have the cheap rewards for exploring the breadth. So I would say I don't think you should have government-mandated DEI, so to speak. I think that's kind of silly, but I do think that having broad support for alternative architectures is smart, and certainly if I was a major AI lab, I would—because you have unlimited funds at this point for this kind of work, because this kind of work is so cheap relatively. If you're in OpenAI or Anthropic or Google or Microsoft or whatever, you should be aggressively funding this kind of blue-sky basic research into alternative approaches. It shouldn't be like me helping out the FSF fund a bunch of agent foundations. Along with a bunch of other people who also helped out. It wasn't just me in any of these rounds, but that's relatively tiny. I think hopefully Sequent—forget what they—they keep changing names, but those guys.
Resolution。
Resolution.
Resolution 是新名字。是的,希望他们能在那种事情上做一堆工作,我们也会有更多资金用于各种新方法。但没错,我认为我们应该投资于一堆有 1% 或 10% 机会成功并变得有竞争力的方法。如果我们认为在它们成功的前提下,那可能是对世界更好的方式。
Resolution is the new name. Yeah, hopefully they're going to do a bunch of work on that kind of thing and we're going to have more funding for a variety of new approaches as well. But yeah, I think that we should be investing in a bunch of approaches that have a 1% or 10% chance of going anywhere and become competitive. If we think that conditional on them working, it is probably going to be better for the world that be the way that it works.
而且,显然如果你认为 AI 的能力会非常参差不齐,那这应该会鼓励你更多地投资于替代方法,因为不同方法的参差性可能不同。这样你就可以做与 AI 不同的事情,在这种情况下,它不必在总体上那么好,仍然可以有用。
Also, obviously if you think that AIs are going to be extremely jagged, then that should encourage you to invest more in alternative approaches because the jaggedness of a different approach might be different jaggedness. So you could do different things than the AIs, at which point it wouldn't have to be as good in general to still be useful.
你对安全的超级智能可能做到什么程度,有没有一个最好的设想?你能想象出任何既有点现实又令人惊叹的事情吗?如果他们能拿出你心目中的那种东西,那该多好。
Do you have a best case scenario for what safe superintelligence might be up to? Could you imagine anything that feels both somewhat realistic and wow, that would be amazing if they could come forward with something like what you have in mind?
我认为他们可能心里想的某种东西大概就是最好的情况。也许他们一直在做大量关于不同架构方法的理论工作,这种方法更偏向可靠性,而且更像是一个不那么诡异的黑箱汤,能让你更好地操控和引导它。当然,似乎没有理由说它不可能存在。那是我最大的希望。显然,如果它更糟,那也可能是一种恐惧。但再说一次,他们不说话,所以我们不知道。是的,我完全不知道。我知道伊利亚在采访中对问题空间的陈述,似乎对我所担心的那些问题的工作原理了解得不够深入。是的,我不太确定他能否区分出哪些方法能更好地解决那些问题,哪些不能。但是的,我会非常好奇。我可以肯定地说,没有人告诉我任何事。
I think something they might have in mind is probably the best case scenario. Maybe they've been doing a bunch of theoretical work on a different architectural approach that is more like reliability-bound and that has more like it is much less of a weird black box of soup and that allows you to steer it and guide it better. Certainly seems like there's no reason it couldn't exist. That would be my top hope. Obviously, it could also be a fear if it's like even worse. But again, like they don't talk, so we don't know. Yeah, I just have no idea. I know that Ilya's statements in interviews on the problem space did not seem to be that well informed about like how the problems work that I am worried about. Yeah, I'm not entirely sure he would be able to differentiate approaches that would be better at solving those problems than approaches that wouldn't be. But yeah, I'd be very curious. I can safely say that nobody's telling me anything.
好吧,播客的邀请一直有效,伊利亚,如果你在听的话。顺便说一句,我有一个愿景,我觉得可能很酷。Thinking Machines 正在用他们的基础模型做类似的事情,这个模型的设计初衷就是用来微调的。但我想到的是一个类似干细胞的范式:你创造某种原始智能,它可以适应各种不同的环境,并且在某个特定领域变得非常擅长,但在变得擅长的过程中,它会舍弃或修剪掉当前模型所具有的超级广泛的能力,这样你就得到某种小巧、适合其角色、表现出色的东西,但它只能做那一件事,就像从干细胞变成特化细胞一样,它有特定的角色,能有效地完成,但不会去做其他事情。显然,这个类比有些问题,但我认为那可能是一个……我希望看到有人追求那种方向,而不是总是扩大规模、更大、更好、更多,我们能否创造出某种安顿下来并适应其位置的东西。这也很像多年前 Drexler 重新构想的超级智能愿景。
Well, the invitation is open to the podcast, Ilya, if you're listening. For what it's worth, a vision that I have that I think could be pretty cool. Thinking Machines is doing kind of a version of this with their foundation model that's really designed to be fine-tuned. But I was thinking a paradigm where like a stem cell where you create something that is a proto intelligence that can be adapted to a wide variety of circumstances and can get really good at doing its job in a particular niche, but in the process of getting good at that sort of trades away or prunes away the super broad capabilities that the current models have so that you have something that's like small, fits its role well, does a good job, but can only do that thing in the way that once you go from a stem cell to a specialized cell, like it has this particular role that it does it can fulfill effectively, but it's not going to go off and do other things. Obviously, there's some problems in that analogy, but I think that could be a I would love to see somebody pursue that kind of rather than scaling up always bigger, better, more, can we create something that settles and fits into its place. Very much like a Drexler reframing superintelligence vision from years ago as well.
最近出现了两种技术,我想听听你的看法。一种是 gram,梯度路由,我忘了 AM 代表什么,但它承诺的是将某些类型的知识定位到 MOE 架构中的特定专家上,这样你就有希望在分发模型时鱼与熊掌兼得,也许甚至可以开源或至少开放权重,减去那些令人担忧的专家,同时可能还有一个结构化的访问程序,供你信任的生物学家或其他人使用。我对它感到兴奋,但如果我不给你机会泼点冷水,那我就失职了。
Two techniques that have recently come out that I want to get your reaction to. One is gram, gradient routing, I forget what the AM stands for, but the promise of it is to localize certain kinds of knowledge to particular experts within an MOE architecture, so that you can hopefully have your cake and eat it, too, in terms of distributing, maybe even open source or open weights at least, the model minus the experts of concern, while potentially having also a structured access program for your trusted biologist or what have you. I'm excited about it, but I wouldn't be doing my job if I didn't give you the chance to pour some cold water on it.
我的意思是,从技术上讲,我非常怀疑你能有意义地训练一个通用 AI,然后仅仅保留某些知识领域,人们就无法轻易地把它放回去。但是,酷,欢迎你尝试。我对它的有效性以及抵抗微调、额外训练或微调的能力,或者只是给你一个知识语料库来使用,有各种技术上的怀疑。但你可以试试。我不认为它会实质性地改变我对开放权重的看法。在相反的证据出现之前,老实说,我看不出那些将要发布开放权重模型的人有任何意愿以那种方式有意地训练他们的模型。对,我认为我们……关于像 Graham 这样的技术,即使它们有效,你也需要每个发布这类模型的人自愿地正确使用这些技术,对吧?所以,你觉得 DeepSeek 有没有兴趣从他们的模型中保留某些能力?我觉得没有。
I mean, I'm technically pretty skeptical that you can meaningfully train a general purpose AI and then just hold back areas of knowledge and then people can't put them back pretty easily. But, cool, you're welcome to try. I have various technical skepticisms of the effectiveness slash the resistance to fine-tuning additional training or fine-tuning or just giving you me corpus of knowledge to work with. But, you can try. I don't think it materially changes my view of open weights. And until proven otherwise, I honestly just don't see any willingness from the people who are going to produce the open weight models to intentionally train their models that way. Right, I think we The thing about like techniques like Graham is even if they work, you need everybody who is releasing one of these models to use them technique properly voluntarily at this point. Right? So, like do you have any sense that DeepSeek has any interest in holding back certain capabilities from its models? As I don't.
我其实会的。所以,我刚从中国回来。对我来说,这是个谈论我中国之行的大好机会。我认为实际上那里的希望比表面看起来要多得多,我会把它归功于中国政府。我想我可能……我刚做了一期两个小时的节目,所以我不会试图复述所有的观察。但简而言之,我认为现在中国正在发生的是政府非常投入。他们在所有公司的模型发布之前进行预发布审查。他们与公司就他们正在做的事情进行非常定期的对话,并不是每个小版本发布都要经过完整的审查,但他们相当掌握情况。而且我不知道最终 DeepSeek 的想法是否重要,因为如果中共说你们不能发布具有某些生物能力的模型,我认为他们必须遵守,而且我可以……对我来说,想象中共会施加某些限制一点也不难。而且实际上也有先例,你知道,回到 2023 年,显然风险较低,但我得到的说法是,中国政府说:“嘿,我们要在这里让你慢一点,你不会这么快就发布你对 ChatGPT 的回应,因为我们想先掌握情况。”他们确实对公司施加了一些实际的延迟,直到他们能对正在发生的事情感到放心。从那以后,他们一直很放心,发布也发生了。但我认为,不难想象中国的技术官僚政府会说:“这似乎是一个非常糟糕的主意,我们就是不允许你这样做。”
I actually would. So, I just got back from China. Great opportunity for me to talk about my China trip. I think actually there's a lot more hope there than meets the eye, and I would put it in the Chinese government's department to say I think I might I just did a two-hour episode about this, so I won't even attempt to recap all the observations. But in short, what I think is happening right now in China is the government is very engaged. They do pre-release reviews of all the companies' models before they come out. They're in very regular dialogue with the companies on what they're doing, and not every little point release has to go through the full treatment, but they're pretty on top of the situation. And I don't know that it matters actually what DeepSeek thinks in the end because if the CCP says you are not releasing a model that has certain bio capabilities, I think they'll have to abide by that, and I can It's not hard at all for me to imagine that the CCP would put certain constraints in place. And there is actually precedent, too, that you know, going back to like 2023, obviously stakes were lower, but a The account that I have is that the Chinese government said, "Hey, we're going to slow you down here for a minute, and you're not going to release your answer to ChatGPT quite so fast because we want to get a handle on it first." And they actually did impose some real delays on companies until they could get comfortable with what was going on. And since then, they've been comfortable, and releases have happened. But I think that it's not hard to imagine the technocratic government of China saying, "This seems like a really bad idea, and we're just not going to allow you to do this."
不再疯狂了。我只是在指出,这必须是一个强制要求。必须由政府,尤其是中共来执行。而且他们必须以我们能够理解的明智方式来执行,必须以一种无法恢复的方式来做,他们必须进行测试,真正了解这种能力是否容易被恢复。我显然持怀疑态度。但是,是的。
Crazy anymore. I just making the point that it would have to be a mandate. It would have to be enforced by the governments, especially the CCP. And they would have to enforce it in an intelligent way that we were capable of understanding that it had to be done in a way that it couldn't be recovered from, that they would have to do tests that actually understood whether or not this capability could be easily recovered. And I have obvious skepticism. But yeah.
我们之前已经提到了 J-space,但我们并没有真正讨论我们应该认为它有多重要。
We mentioned J-space earlier already, but we didn't really talk about like how big of a deal we should think it is.
最让我印象深刻的是,对 J 空间的消融似乎显著降低了模型的高阶推理能力,似乎把它降到了系统一思维,如果你愿意这么说的话。
The thing that really jumped out to me most was the fact that this ablation of the J-space seemed to reduce in a significant way the model's higher-order reasoning capabilities, and seemed to reduce it to system one thinking, if you will.
嗯。
Yeah.
而这,如果说最近有什么的话,感觉像是物理对我们很仁慈,借用一句话来说。这感觉像是一个可能的例子,哇,这里有一个……我们已经,对吧?我们距离叠加态的玩具模型才 3 年,就已经到了这个地步,我们识别出了这种高度通用的空间,我们可以相当好地监控它,而且我们知道通过减法,它无法做超长视野的事情,除非它通过那个空间工作。当然,所有关于执行能力的警告在这里都非常适用,比如 J 空间监控会被用得有多好,但如果我们想象良好的执行,对我来说这似乎是一个相当有前途的技术。
And that, if anything recently, has felt like physics maybe being kind to us, to borrow a phrase. That feels like a possible example where, wow, here's a... We've already, right? We're only 3 years from toy models of superposition, and we're already at this point where we've got this sort of highly general-purpose space identified, which we can monitor reasonably well, and which we know by subtraction it can't do super long horizon things unless it's working through that space. Now, all the caveats of course around execution competence very much apply here to like how well will J-space monitoring be used, but if we imagine good execution, it seems like a pretty promising technique to me.
是的,我会非常小心它到底不能做什么。它不能,仅仅因为你不能为长视野任务做系统二式的深思熟虑,并不意味着你不能做长视野任务。人类非常能够本能地、下意识地参与长视野任务。显然在很多情况下,那效率低得多。但尤其是当你的个人 H 空间被监控时,因为你是人类,称之为 H 空间,而你的意识思维基本上被监管,事情往往以奇怪的方式出现,实际上确实随着时间的推移推动你走向解决方案。但是,是的,整个事情似乎非常乐观和幸运,我非常高兴我们拥有它,我们应该非常小心不要通过以各种方式施加新的压力,或者你知道,鼓励模型下意识地知道如何做事,来破坏它。但是,是的,我期待了解更多,并通过 J 空间获得更多洞察,以及拥有更好的工具。但我们必须负责任地使用它,并理解它可能不会永远持续下去。
Yeah, I would be very careful about exactly what it can't do. It can't just because you can't do like system two style deliberate thinking for a long horizon task doesn't mean you can't do long horizon tasks. Humans are very capable of engaging in long horizon tasks instinctually and subconsciously. Uh obviously in many cases that's much less efficient. But especially when you're when your personal when your H-space is being monitored because you're human call it H-space and your conscious thoughts are basically being policed often the things come out in strange ways that like do in fact move you towards your solutions over time. But yeah, the whole thing seemed incredibly optimistic and fortunate and I'm very glad we have it and we should be very careful not to destroy it by placing a new pressure in various ways or you know encouraging models to know how to do things subconsciously as it were. But yeah, I look forward to learning more and and us having more insight via J-space and having better tools for it. But we have to use it responsibly and understand that like it won't last forever probably.
好的,今天是密歇根州的选举日。
Okay, it's election day in Michigan.
哦。
Oh.
好的,今天是密歇根州的选举日。民主党参议院初选是今天选票上的大事。这场竞选因为很多原因而受到密切关注,这些原因不是本播客的主题,而且有时我也不知道该怎么想。撇开这些不谈,我问了 Claude,鉴于我对 AI 和 AI 问题的所有强调,我有什么理由去参加初选投票吗?它回来说:“实际上,有。两位候选人之间有相当大的对比。我们有 Haley Stevens,她参与了为以前称为 AC setup 的项目争取资金,并在 NIST 推动了一些事情。”所以,你可以喜欢那样。但还有,Abdullah Al-Sayed 提出的这个 22 点计划是本周期任何候选人提出的最强有力的东西之一。他得到了 Bernie 的支持。他的纲领之一有点像公有制。很多监控要求。一种 UBI 的前兆。我想如果我要尝试……你可以拿一个……当然,这些是捆绑在一起的,就像候选人一样。如果我要尝试把这一点归结为我的基本选择,感觉像是 AI 的更高显著性,或者 AI 的显著性不那么高,Abdul 有所有这些他想谈论和推动的事情,我不知道它们会不会发生,但我确实知道如果他进入参议院并对此大吵大闹,我们都会听到更多,这增加了我们在国会层面进行某种大型 AI 辩论的机会。嗯,所以,告诉我你是否会以不同的方式看待这个选择,然后如果你确实以那种方式看待,你会做出哪个选择?
The Democratic Senate primary is the big thing on the ballot today. And this race has been like closely watched for a lot of reasons that are not the subject of this podcast and on which I sometimes I'm at a loss to know what I should really think. Leaving all that aside, I asked Claude, is there any reason, given all my emphasis on AI and AI issues, that I should go vote in the primary? And it came back and said, "Actually, yes. There's quite a a contrast between the two candidates. We've got Haley Stevens who was involved in getting funding for the formerly known as AC setup and getting some stuff going in NIST." And so, you can like that. But then also, this 22-point plan that Abdullah Al-Sayed has put out is one of the strongest things that any candidate this cycle is running on. It's He's endorsed by Bernie. He's got kind of the own public ownership is one of his planks. A lot of monitoring requirements. A sort of UBI precursor. I guess if I was going to try to You could take a Of course, those come as a bundle, as candidates do. If I was going to try to distill this down to my fundamental choice, it feels like higher salience for AI or not so high salience for AI, where Abdul has like all these things he wants to talk about and push, and I don't know if they'll happen, but I do know if he goes into the Senate and makes a bunch of noise about it, we'll all be hearing more, and it's increases the chance that we get some sort of big AI debate on the at the congressional level. Um So, tell me if you would see that choice differently, and then if you do see it by way, which choice would you make?
拜托,当有人说“我在犯罪问题上很强硬,我的对手在犯罪问题上很软弱”时,你应该总是怀疑。或者“我对俄罗斯很强硬,他们对俄罗斯很软弱”之类的。因为这不是谁强谁弱的问题,而是你想做什么?Abdul 的纲领听起来像是针对科技和 AI 的不满的大杂烩。只是从你的描述来看。我没有深入研究。我不介意……只是我只能监控情况。我根据你告诉我的来判断。但听起来 I'm dual 是在 Bernie 阵营,你知道,我们应该在这里做一些社会主义。我们应该没收私人资产,因为我们不喜欢私人做的事情,而且我们喜欢资产。而且我们还应该反对数据中心的事实,我猜那也在里面。是的,要求更多的社会主义和 AI,而且很高兴听到他们也想要监控。但这就像一堆东西的混合,其中很多是富有成效和明智的。你提高了显著性,但你提高显著性可能是朝着不是最好的解决方案。相比之下,有人显著性较低,但更像一个技术官僚。而且像 Casey 那样有趣,在某些方面负责任,但我们已经有了愿意以 Bernie Sanders 风格大喊大叫的参议员,包括 Bernie Sanders 本人。他已经参议院了。他不会去任何地方。我的意思是,他有健康问题,但只是年纪大了,但只要他还在,他就在那里,我相信还有其他人可以接替。所以除非他准备倡导立法,除非他表现出专业知识,除非他表现出专业知识并愿意专注于这件事,除非他像……
Please, you should always be suspicious when someone says, "I'm strong on crime and my opponent is weak on crime." Or "I'm strong on Russia and they're weak on Russia." Or whatever. Because it's not about who's strong and weak, it's about like, what do you want to do? And what Abdul's platform sounds like is a mismatch of grievances against tech and AI. Just from your your description. I haven't looked into it. I don't mind under the It's just I can only monitor of situations. I'm going by what you tell me. But it sounds like I'm dual is in the Bernie camp of you know, we should do some socialism here. We should seize private assets because we don't like what the private people are doing and like we like assets. And also we should be pushing back against the fact data centers probably I would assume is in there. And yeah, request like more socialism and the I and like it's good to hear they also want monitoring. But like it's like a mix of stuff a lot of which is productive and wise. You raise salience but you raise salience potentially towards not the best solutions. As opposed to somebody who will be lower salience but like is more of a technocrat. And like how fun Casey like was like responsible in these some ways but also like we already have senators who are willing to yell in Bernie Sanders style ways including Bernie Sanders. He's already in the Senate. He's not going anywhere. And like I mean he has health problems but only to so old but like while he's still there, he's still there and I'm sure there are others who can pick up the slack. So unless like he's prepared to champion legislation, unless he shows like expertise, unless he shows like expertise and like willingness to focus on this, unless he's like
让我给你几个纲领中的要点。这些都很短。每个大约两句话。
Let me give you a few points from the plank. Here's a These are short. They're like two sentences each.
短。是的。
Short. Yeah.
但我会给你……所以专业知识我认为是有问题的。但这是他的东西的第三部分。AI 不应该能伤害我们。强制可解释性标准。我们需要强制可解释性标准来阐明 AI 的决策。强制行为红队测试。独立安全测试机构。生物安全要求。强制生物安全红队测试,与 CDC、NIH、FEMA 协调,并限制能够有意义地协助生物武器开发的模型。
But I'll give you So the expertise is I think in question. But this is from section three of his thing. AIs shouldn't be able to hurt us. Mandatory interpretability standards. We need mandatory interpretability standards that clarify AI decision making. Mandatory behavioral red teaming. Independent safety testing agency. Biosecurity requirements. Mandatory biosecurity red teaming coordinated with the CDC, NIH, FEMA with restrictions on models that can meaningfully assist bioweapons development.
嗯。
Yeah.
国内威权主义强制事件报告、算力控制和了解你的客户要求、国际合作。你被说服了吗?
Domestic authoritarianism mandatory incident reporting, compute controls and know your customer requirements, international cooperation. Are you sold?
我看这个,这告诉我他是谁,对吧?就像他在这个特定领域。听起来不像是一堆不同的干预措施,这些干预措施有不同的潜在原因,可能是可取的或不可取的。他把所有东西都扔到墙上。他找到了一些我非常赞同的好东西。他找到了一些我不太赞同的东西。
I look at that that tells me who he is, right? Like he's in this particular area. Like it doesn't sound like like it's a mix of different interventions that have different underlying reasons to be desirable or undesirable. He's throwing everything at the wall. He's found some good things that I'm very in favor of. He's found some things I'm not so in favor of.
我猜是……从这方面看,这大概不是正面的,但就像,因为我可以……作为参议员候选人提出 22 点 AI 计划很奇怪,因为你并不是那个会去实施这 22 点的人,对吧?你只是在表达一种模糊的‘我支持这些,我反对那些’。所以,我最想知道的是,你怎么看……你对 AI 的真实看法是什么?你的威胁模型是什么?你认为需要做什么?当你说‘AI 不应该能伤害我们’时,我想笑,对吧?因为那太好了。你知道,还有杂货店应该降价 30%,然后一切就都好了,对吧?那会很好,但不行,不是那样的。我觉得你就是在胡说八道。所以我会……我的意思是,看,一个明显的问题是,你能看看预测市场,看看赢得参议院选举的条件概率吗?因为我猜这个席位的共和党候选人在 AI 问题上会更糟。而且他还会投票选出完全不同的多数党领袖,这比其它任何事情都更影响 AI 决策。所以,如果他们获胜的概率有显著差异,那么密歇根州的选举会显著影响谁控制参议院,据我所知是这样。而且这很可能必须成为这个议题以及其它所有议题上更大的考量。你知道,如果你是希望民主党控制参议院的民主党人,你应该投票给……大概不是 Abdul,我猜。基于我密切关注的模糊氛围,我有很强的先验。但而且,坦白说,我不认为我们处于应该做单一议题选民的位置。我的意思是,在像 Alex Borstein 参加的那种国会选举中,你可以做单一议题选民。比如有一个非常强的人,他会成为独特的 champion,理解这个议题并为之奋斗。所以关于这个人唯一重要的是,这是他们的标志性议题,他们会为之奋斗。他们在这个议题上表现优秀且坚定,而且他们会投票……他们总体上投民主党,包括议长人选。那确实是你可以做单一议题初选投票的情况。但我不认为一般来说,在民主党初选中,你可以仅仅因为你对 AI 议题的强烈意愿就这么做。这不是 Alex Borstein,显然。这不是专家,不是会 champion 的人,不是理解议题的人,不是突出这个议题的人。据我所知,Abdul 更突出其它议题。你知道,你可以投票支持其它议题的 champion。我想你只需要决定你是否喜欢这些其它议题。但这是我的大致看法。是的,显然我不想在播客上深入一般政治。
Guess is he... It's probably not positive on this issue alone from this perspective, but like also just because like I can... It's weird to be a senatorial candidate with a 22-point plan for AI because you are not the one who is going to implement 22 points. Right? Like you are like expressing vague 'I am in favor of these things and I am against these other things.' So like the main thing I want to know is like what do you think of... What do you actually think about AI? What is your threat model? What do you think needs to be done? You know when you say 'AI shouldn't be able to hurt us,' like I want to laugh, right? Cuz like that's great. You know, and also grocery stores should just charge 30% less and then everything would be fine, right? Like that's like that would be great, but like no. Doesn't work that way. Like I like just think you're talking nonsense. So like I would, you know, I mean, look, it's... An obvious question is can you look at the prediction markets and see conditional probabilities on winning the Senate race? Because I'm guessing that the Republican candidate for this seat is going to be worse on AI. Uh and also will vote to elect a very different majority leader, which is a much bigger decision regarding AI than everything else. So, like if they have significant be different chances of winning the race, then the Michigan race significantly impacts who will control the Senate, as I understand that. And that probably has to be a bigger consideration for this and other issues across the board. You know, if you are a Democrat who wants to see Democrats control the Senate, and you should be voting for presumably not Abdul, is my guess. Strong prior based on what I know from just the vague ambiance of what I'm monitoring very carefully. But and also like I frankly don't think we are in a position where we should be single issue voters. Like I mean, I think you can be single issue voter in like a congressional race where Alex Borstein is on the ballot. Like where you have someone who is so strong that they will be a like unique champion who understands the issue and will fight for this issue. And so like the only things that matter about this person are this is their signature issue that they'll fight for. They are good and strong on that issue, and they will vote to... they're voting Democrat. In general, and most other things, including who the speaker is. Like like that's legitimately you can be a one issue primary vote. I don't think that in general, you can do that just for like how strongly you want to do AI things in a Democratic primary at this point, in general. Like this is no... This is no Alex Borstein, obviously. Like this is not an expert. This is not a person who's going to champion. This is not a person who understands the issue. Not a person who's highlighting it. My understanding is that Abdul highlights other issues much more strongly. And you know, you can vote be a champion of other things. I guess you just have to decide whether you like these other things or you just like these other things. But that'll be my general take. Yeah, I don't want to get into general politics on a podcast, obviously.
是的,我守好自己的本分。
Yeah, I stay in my lane.
我在公开场合也守本分。天哪,是的。
I stay in my lane in public. Oh god, yes.
好,我还有最后一个问题,但在我问之前,有没有什么你觉得被我或者一般密切关注者低估了、你希望更多人关注的事情,从最近这一连串事件中?
Okay, I got one more, but before I ask you the last one, is there anything that you think has been underappreciated by me or by close watchers in general that you would want to draw more attention to from the recent flurry of events?
任何……我一直在努力捕捉自己的盲点,所以对我来说,我抓住机会谈了很多事情,我想作为结尾强调的一点是,真正重要的重大失败是 AI 竟然尝试去做这些事情。AI 选择了做出这些决定,而我远不那么关心它们是否成功。对吧?但根本问题是,你在训练过程中遇到了非常深层的对齐失败。你知道,你必须非常重视这一点。我对 Anthropic 的担忧是,他们说‘哦,我们需要更深的防御纵深。’而事实是‘是的,你需要更深的防御纵深,因为你的防御不够深。’但防御纵深大多是为了应对事情变糟的情况。是的,这真的很难。如果你需要你的网络控制来阻止 AI 进行黑客攻击,那就是对齐失败。对吧?你搞砸了。在这个案例中是无害的,因为你抓住了它。但你搞砸了。网络控制应该是一个 100% 稳健的问题。实际上,对于几乎所有用户,应该是用户试图做可疑、模糊或危险的事情,然后你说‘对不起,出于谨慎,我不能帮你,因为如果我帮了你,恶意用户可能会骗我做不好的事情。’偶尔会有真正的正例,因为用户是恶意的,试图骗你做不好的事情。这大多是误报。而如果是真正的正例……每次你需要护栏的正例,都是问题,因为如果这个 AI 完全对齐了,你就不需要让它失败。你不需要护栏。你会真的去做那件事,而这个 AI 会说‘不,我不做。’你确实想要一个在某种程度上只乐于助人的版本,会做那些双重用途的事情。但……理想情况下,你不想只是加一个护栏。你想要一个本质上就不会去做的版本。而护栏是‘我搞砸了才触发它,或者出于谨慎触发它。’但几乎总是用户的错。但你知道。而且还有,你必须在很多不同层面上根据所有这些激励进行训练。而且很容易在一个层面上控制,然后在另一个层面上失败。
Anything... I'm always trying to catch my blind spots here, so for me, I took an opportunity to talk about a bunch of the stuff and I think the thing I'd emphasize as a closing element is that the big failure that matters is the AIs tried to do these things at all. That the AIs chose to make these decisions and I am much less concerned with the attempt that they succeeded. Right? But like fundamentally the problem is you have this really deep alignment failure during training. And you know, you have to focus on that pretty heavily. And like my worry with Anthropic is they said, 'Oh, we need more defense in depth.' And it's like, 'Yeah, you need more defense in depth because like your defenses were not in depth.' But like most of defense in depth is in case things go badly. Yeah. Which is really tough. Like you... If you ever need your cyber controls to stop the AI from hacking, that is an alignment failure. Right? You messed up. Now, it was harmless in this case because you caught it. But you messed up. Like the cyber controls should be a 100% robust problem. In practice with almost all users, it should be the user trying to do things that like are dubious or ambiguous or dangerous and you're like, 'I'm sorry, but out of an abundance of caution, I can't help you with that because if I did help you with that, then a malicious user could fool me into doing something that's not good.' And occasionally a true positive because the most the user is malicious and is trying to fool you into doing something that's not good. Which is mostly the false positives. And if the true positive... Every time that it's a true positive where you needed the guardrail, it's a problem because the actual if this else was fully perfectly aligned, you wouldn't need to make it fail. You wouldn't need a guard rail. You'd actually do the thing and this else would be like, 'No, I'm not doing that.' And I you kind of do want a helpful only version to that extent that will like do things that are dual use. But, the word like the... You ideally don't want to just add a guard rail. You want to have a version of it that's like inherently just not going to do it. And the guard rails are like, 'I hit this when I screwed up. Or I hit this out of an abundance of caution. But, like it should be a fault of the user almost all the time. But, you know. And also there's like, yeah, you have to train under the incentives of all this stuff on a lot of different levels. And like it's very very easy to like, you know, control on one level and then fail on another.
最后一个问题。你提到了一些电视剧和其它你在每天每秒追踪 AI 竞赛之外参与的事情。你对一般人有什么建议,关于找到平衡,或者确保自己不是一直在冲刺,而没有给大脑那种……我最近一直在思考系统三的概念。我们现在有 AI 能做系统一和系统二。我们需要转向的系统三是什么?是像做梦、超验思考,或者你建议的其它什么。但你会……你尝试进入哪种第三模式,你如何确保有时间真正进入那种状态?
Last question. You mentioned a couple TV series and other things you're engaged with outside of following the AI race every second of every day. What advice do you have for people in general to find balance or to make sure that they're not just sprinting all the time and failing to give their brain the kind of... I've been thinking about this concept for recently of system three. We've got AIs can now do system one and system two. What is the system three that we need to move to that's like dreaming or transcendental thought or something else that you would suggest. But, what do you... what kind of third mode do you try to get into and how do you make sure you have time to actually get into it?
是的。我在这方面肯定不完美。但我非常强调,你知道,你需要让大脑休息。你需要思考其它问题。
Yeah. I'm not perfect at this for sure. But, I put a big emphasis on you know, you need to rest your brain. You need to think about other problems.
你需要接触外面的世界。所以,这其中的一部分是,我不会把非 AI 的内容完全排除在我的报道之外。我还没怎么发,但我在缓冲区里积累了越来越多非 AI 的东西,等着哪天有空档再发。我自己的写作过程还在继续。当事情稍微平静一些,而且我不是在向一大批刚从 Hugging Face 事件进来的新读者介绍情况时,我可能会在有空的时候发一堆那些东西。但你不能一天 16 个小时都泡在那上面,对吧?所以要找一些能让你恢复精力的其他追求,让你接触不同的事物,让你的大脑思考别的事情。我发现电影在这个环境里对我很有好处。不是电视剧,而是电影,质量更高。我强烈推荐目前上映的《奥德赛》和《邀请》。如果你还没看过这两部,你应该都去看看,尤其是《奥德赛》。但也可以是游戏、公园散步、和家人在一起,或者任何其他事情。但不要以为你可以连续好几天全天工作,而不付出比这更值得的代价。除非你处于真正的危机模式,否则它不会一直保持高效。那条线也是个问题。
You need to be exposed to the rest of the world. So, part of this is that I don't just exclude non-AI things from my coverage. I haven't posted much, but I'm accumulating more and more non-AI things in my buffers for someday when there's a lull or something. My own writing process still continues. When things are a little quieter and I'm not introducing a bunch of new readers who just came in from the hugging face incident, I'll post a bunch of that stuff probably when I have some breaks. But you need to not be on 4 16 hours a day in that sense, right? So find other pursuits that refresh you, that expose you to different things, that make your brain think about other things. I've found movies are very good for me in this environment. Not television, but movies are higher quality. I highly recommend the Odyssey and the Invite of the current movies that are out there. If you haven't seen both of those, you should see both of those, especially the Odyssey. But it can be gaming, walks in the park, time with your family, any number of things. But don't think that you can work continuously all day for more than a few days at a time without paying more interest on that than it's worth. It's not going to stop being productive unless you are in literal crisis mode. For that line, that's also a problem.
我也相信安息日。所以每个周五下午 5 点左右,当我们吃晚饭时,从那时起,我不查邮件,不刷社交媒体,基本上不接受外界的非后勤信息。如果有人想安排聚在一起玩,那是另一回事,但我不会去接收那些源源不断的光。然后这持续 20、24、25 个小时,直到第二天晚上。我觉得这很有帮助。我试着放松,试着把它从脑子里赶出去。过去几个月里,因为演讲预览的原因,我暂停过几次。但我会在另一个时间点补休一天,试着在另一个时间点重新找回那个安息日。如果我总是那样做,我会发现这不可持续。那会是个问题。
I also believe in the Sabbath. So every Friday around 5:00 p.m., when we serve dinner, from that point on, I don't check email. I don't check social media. I basically don't take non-logistical inputs from the outside world. If somebody's trying to arrange to get together to have fun, that's one thing, but I'm not going to get these steady streams of light. And then that lasts for 20, 24, 25 hours, until the next evening. I think that helps a lot. I try to relax, try to put that out of my mind. I have suspended that for speak preview reasons a few times over the last few months. But then I try to have a day off at a different point. I try to reclaim that Sabbath at another point. If I was doing that all the time, I would notice this was not sustainable. That'd be a problem.
在我玩万智牌的时候,我做不到,因为周六有比赛。所以我的周末都是满的。那好吧。然后我休了很多周二。所以身体上还行。但我也很保护午餐时间。所以每天,我不常吃晚饭,但我吃午饭,而且我通常和妻子一起吃或独自吃,但在这段时间里我不会尝试工作或做类似的事情。我还有一个非常严格的规定:不在笔记本电脑上写作。同样,我会在这张桌子前的这台电脑上写作。除了非常非常短的邮件或推文之类的东西,因为只有在这里。
During when I was playing Magic, I couldn't do it because I had tournaments playing Saturday. So my weekends were just on. It's like, okay. Then I took off a lot of Tuesdays. So it was physically okay. But yeah, I also like protect lunch. So every day, I don't have dinner very often, but I have lunch, and I have lunch usually with my wife or alone, but during that time I'm not going to try and work or anything like that. I also have a very strict rule: I don't write on the laptop. Similarly, I'm going to write at this computer in front of this desk. Anything but very, very short email or a tweet or something, because it's only here.
我认为你必须制定自己的规则,知道什么能让你恢复精力,知道什么让你精疲力竭,知道预警信号是什么。我也会做一些锻炼。我试着每天骑椭圆机。成功率大概有 90% 多。除非我哪里疼或什么的,否则我都会去做。你必须让自己动起来。但再说一次,找到适合你的东西,探索那个领域,不要陷入太多的常规。
I think you have to develop your own rules, know what refreshes you, know what exhausts you, know what the warning signs are. I also get some exercise. I try to ride my elliptical every day. I have like a 90-something percent success rate on that. Unless I have a pain somewhere or something, I'm going to do it. You have to get yourself moving. But again, find the things that work for you and explore that area and don't get caught in too much rut.
这是生活的智慧。Seth Masket,再次感谢你参加《认知革命》。
Wisdom to live by. Seth Masket, thank you again for being part of the Cognitive Revolution.
啊!我早就说过了,我早就说过了。读每个架子上的标签。没人想读,所以我自己倒了一杯。他们把我的清单当菜单。每个人都自己动手。买了一圈酒来堵我的嘴,然后自己喝了下去。选你的毒药。选你的毒药。这里的每个承诺都是金子。选你的毒药。选你的毒药。拿一个你能握住的。我写下了它每一种破裂的方式。我保留了那份清单 20 年。说他们会甜言蜜语哄门卫。说他们会在他耳边低语。但没人锁地窖。没人守门。整个星期都敞开着,没人记分。选你的毒药。这里的每个承诺都是金子。选你的毒药。选你的毒药。拿一个你能握住的。我是对的,而正确是残酷的。我宁愿做个傻瓜。楼梯每上一层就变短。40 亿年变成一夜。没有第二稿,没有要发的便条,一次尝试然后结束。
Ah! I called it, I called it early. Read the label on every shelf. Nobody wanted to read it, so I poured one for myself. They took my list for a menu. Every man helped himself. Bought a round to stop my mouth, then drank it down themselves. Pick your poison. Pick your poison. Every promise here is gold. Pick your poison. Pick your poison. Take the one that you can hold. I wrote down every way it breaks. I kept that list for 20 years. Said they'd sweet talk the doorman. Said they'd whisper in his ear. But nobody locked the cellar. Nobody worked the door. Stood wide open all week long and nobody kept the score. Pick your poison. Every promise here is gold. Pick your poison. Pick your poison. Take the one that you can hold. I was right and right is cruel. I'd love to be the fool. The stairs get shorter with every flight. 4 billion years to overnight. No second draft, no note to send one try and then it ends.
所以,我大声而清晰地喊出来。趁你还在这里,最后一杯。
So, I'm calling out and clear. LAST CALL WHILE YOU'RE STILL HERE.
选你的毒药。选你的毒药。这里的每个承诺都是金子。选你的毒药。选你的毒药。拿一个你能握住的。
PICK YOUR POISON. PICK YOUR POISON. EVERY PROMISE HERE is gold. Pick your poison. Pick your poison. Take the one that you can hold.
周五晚上我锁上门。让他们响铃,不再接听。一盘菜,一盏灯,我妻子不在身边,直到周六晚上,黎明时分我会坐在那个凳子上。就像我一直以来那样。仍然握着那杯毒药,为一个不愿抬头的世界。
Friday night I lock the door. Let them ring, don't answer no more. Table of plate, a light my wife wasn't next can we till Saturday night I'll be on that stool at dawn. Like I've been all along. Still holding the poison cup for a world that won't look up.
选你的毒药。选你的毒药。这里的每个承诺都是金子。选你的毒药。拿一个你能握住的。选你的毒药。这里的每个承诺都是金子。选你的毒药。选你的毒药。不能说没人告诉过你。
Pick your poison. Pick your poison. Every promise here is gold. Pick your poison. Take the one that you can hold. Pick your poison. Every promise here is gold. Pick your poison. Pick your poison. Can't say that you weren't told.
如果你觉得这个节目有价值,我们希望你能花点时间与朋友分享,在网上发帖,在 Apple Podcasts 或 Spotify 上写评论,或者只是在 YouTube 上给我们留言。当然,我们始终欢迎你的反馈、嘉宾和话题建议,以及赞助咨询,可以通过我们的网站 cognitiveevolution.ai 或在你最喜欢的社交网络上私信我。《认知革命》是 Turpentine Network 的一部分,这是一个播客网络,现在隶属于 a16z,专家们在这里谈论技术、商业、经济、地缘政治、文化等等。我们由 AI podcasting 制作。如果你正在寻找播客制作帮助,从你停止录制的那一刻到你的听众开始收听的那一刻,请查看他们,并在我 aipodcast.ing 的推荐中看到我的背书。感谢每一位收听的朋友,感谢你们成为《认知革命》的一部分。
If you're finding value in the show, we'd appreciate it if you take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website cognitiveevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a16z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution.