OpenAI's Safety Crisis: Key Departures and Broken Promises
打开互动全文版(中英对照 + 朗读 + 问答)→在 Jan Leike 因根本分歧和资源短缺辞职后,Cognitive Revolution 播客分析了 OpenAI 破裂的安全承诺、严苛的保密协议以及超级对齐团队的解散。
Following Jan Leike's resignation citing fundamental disagreements and resource shortages, the Cognitive Revolution podcast analyzes OpenAI's broken safety commitments, draconian NDAs, and the dissolution of the superalignment team.
大家好,欢迎收听《认知革命》。我们每周采访前沿人工智能领域的 visionary 研究者、企业家和建设者,一起探讨他们的革命性想法,并描绘 AI 技术在未来几年将如何改变工作、生活和社会。我是 Nathan Lens,和我的联合主持人 Eric Torberg 一起。欢迎回到《认知革命》的特别加更。是的,总有新东西要学。今天事情发生得很快。就在我们录完节目后,Yan LeCun 发布了他的推特长文,说明他离开 OpenAI 的原因。他非常直白地说,他与领导层存在根本性分歧,并且难以获得工作所需的资源,包括算力资源。当然,他对队友说了些好话,但基本上就是‘我不认为我们走在正确的轨道上’,而且看起来几乎是抗议式辞职。所以这看起来不是好事,情况不妙。你还了解到什么?我们怎么看待这件事?
Hello and welcome to the Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Lens, joined by my co-host Eric Torberg. Welcome back to a special bonus session of the Cognitive Revolution. Yep, there's always more to learn. It's happening quickly today. So by the time we got off the recording, Yan LeCun had posted his tweet thread statement about his reasons for leaving OpenAI, in which he puts it pretty plainly that he's had pretty fundamental disagreements with leadership and has had trouble getting the resources that he needed to do the work, including compute resources. And certainly had some nice fond things to say to his teammates, but basically was like, 'I don't think we're on the right track' and seems to be resigning pretty much in protest. So that doesn't seem like a good thing. Doesn't seem like a good situation. What else have you learned and what do we make of it?
我们还看到了 Vox、Bloomberg、TechCrunch 的报道。Sam Altman 对 LeCun 的回应非常优雅,基本上是说:‘是的,我们有很多工作要做,我们会去做。稍后会有更长的回应。’这基本上是最好的回应了,但接下来他必须兑现。Kelsey Piper 证实了这些严苛的不贬低条款的性质,这些条款似乎是终身有效的,并且包含一个你不能透露自己违反 NDA 的 NDA。她还声称,当员工首次入职并获得以股权为主的薪酬时,他们并没有被告知需要签署这些不贬低条款,否则离职时现有的既得股权将被没收。所以这看起来是一个非常糟糕的平衡状态和公司运营方式。我真的很困惑为什么这是合法的。
So we also got coverage from Vox, from Bloomberg, from TechCrunch. We got Sam Altman's response actually to LeCun, which was extremely graceful, essentially saying, 'Yes, we have a lot of work to do and we're going to do it. Expect a longer response later.' It's the basically best possible thing you can say there, but then you're on the hook for doing it. We have Kelsey Piper confirming the nature of the draconian non-disparagement clauses, which apparently have lifetime duration and include an NDA that you can't reveal about violating the NDA. And she claims that when employees are onboarded for the first time and are given equity-heavy compensation, they are not told that they will be required to sign these disparagement clauses or have their existing vested equity confiscated upon departure. So that seems like a really bad equilibrium and way to run a company. I'm honestly confused as to why that's legal.
嗯,也许它不应该合法。是的,我认为如果你要因为某人不签署不贬低条款而没收其巨大价值的资产,你必须在最初就非常明确地告知这些条款。这完全不合理。这显然也不能很好地说明公司的开放性,对吧?如果你强迫每个员工终身不得贬低你,无论发生什么,否则……如果你希望人们不往坏处想,这根本不是一个合理的立场。所以我们看 LeCun 的声明,基本上是说多年来,他们一直面临闪亮的新产品成为优先事项的问题,偏离了安全文化,这种文化不利于或与安全兼容。我在这里稍微 paraphrase 一下,没有精确看原文。他还说,尽管 OpenAI 明确承诺提供现有算力的 20%,这应该足以满足当前需求,但过去几个月他仍然难以获得算力。TechCrunch 也证实他们没有兑现承诺。他们只要求了 20% 承诺中的一小部分,却屡次得不到满足。而这正是如今运行任何 AI 项目的关键:算力。你需要算力来做你的工作。他报告说这已经成为他们工作的实质性障碍。所以这不仅是哲学上的分歧,因为他说他们应该投入更多资源,不是多一点,而是应该将更多资源用于为 AGI 未来做准备。而且,如果你注意他实际说的,仅就下一代而言,他基本上认为他们还没有为 GPT-5 做好准备。他认为他们没有按计划拥有让 GPT-5 在普通、日常意义上安全所需的工具,而不是存在主义意义上的安全。
Well, it maybe shouldn't be. Yeah, I think you should have to very much acknowledge a disparagement clause rules very clearly initially if you're going to confiscate something of immense value for someone not signing them. Doesn't seem reasonable at all. It also doesn't obviously speak well of the company's openness in the good sense, right? If you're forcing every employee to never disparage you for life, no matter what, or else... that's just not a reasonable position to take if you want people to not assume the worst. And so yeah, we look at LeCun's statement essentially saying that for years they have had the trouble of shiny new products becoming the priority, a move away from safety culture, that the culture is not amenable to or compatible with safety. My paraphrasing here a bit, I'm not looking at the words precisely. And that he had trouble getting compute the last few months despite the explicit 20% of existing compute commitment from OpenAI, which should have been sufficient for current purposes. And indeed TechCrunch confirms that they have not been honoring their commitments. They have asked for a fraction of the 20% commitment and have repeatedly not gotten what they asked for. And that is part and parcel of the whole idea of running anything in AI these days: compute. You need your compute to do your thing. And he was reporting this has specifically become a substantial barrier to doing their work. So this is not only a philosophical approach, because he said they should be spending vastly more, not just a little bit more, but they should be spending vastly more of their resources on preparing for our AGI future. But also, if you'll notice what he actually said for just the next generation, that essentially he doesn't think they're ready for GPT-5. He doesn't think they're on pace to have the tools they need for GPT-5 to be safe in a pedestrian, mundane utility sense, like not in an existential sense.
然后还有超级对齐的问题,也就是 AGI 的问题,OpenAI 的许多人说他们预计 AGI 会在几年内出现。时间线非常短。现在超级对齐团队已经解散,人员分散到公司各处。他们声称仍会继续这方面的工作,但既然违背了承诺、解散了团队、失去了领导,这看起来不像他们承诺过的那种努力。时间线还成立吗?Sam Altman,我们现在已经进入四年期的九个月了吧?如果他仍然期望超级对齐在四年内成功,那么这些行动并不反映他理解这一点,也不反映他理解——正如他反复告诉我们的——超级对齐的风险。而且很明显,根据 Bloomberg 的文章,LeCun 的临界点是 Ilya 的离开。这完全说得通。但我们看到了一系列离职,有显而易见的理由,这一切都说得通。我们之前都在猜测 Ilya 看到了什么?就是整件事。Yan 看到了什么?也成了焦点。是的,你知道,Yan 辞职后,答案是他们看到的是一家不致力于安全、不愿兑现承诺、文化敌视安全努力、转向闪亮新产品的公司。这是一家产品公司,一家扩张中的初创公司,致力于赚钱。赚钱本身没错,但在这个案例中,他们的明确公司使命和目标是构建超越人类智能的智能。正如 LeCun 指出的并反复承认的,这不是一件安全的事。这不是一件默认会顺利的事,对吧?我们可以辩论,在付出真正努力使其顺利的情况下,它有多大概率会出问题,但我认为任何理性的人都能看到,如果我们不付出努力、不做工作,事情很可能会出问题。而且这不一定要涉及某种特定的 AI 接管或存在风险场景。它只是意味着这对地球上的人类来说可能非常糟糕。我们必须认真思考这些问题。但很明显,以 Sam Altman 为代表的 OpenAI 领导层越来越不这样做,而且在过去几个月里根本没有兑现承诺。
And then later on, there's the problem of superalignment, the problem of AGI, which many people at OpenAI said they expect within several years. It's a very short timeline. And now the superalignment team has been dissolved and its people have been dispersed throughout the company. They claim they will still continue the work on that level, but having dishonored their commitment and having dissolved the team and having lost the leadership, it doesn't look like the kind of effort they promised us that they said they were going to do. Does the timeline still hold? Sam Altman, we're now what, nine months into the four years? If he still expects to need superalignment to succeed within four years, then these actions do not reflect somebody who understands that and understands, as he's repeatedly told us, what is at stake with superalignment. And it's clear LeCun's breaking point was Ilya departing, according to the Bloomberg article. That makes perfect sense. But we have a series of departures, we have obvious justifications for that, and this all makes sense. We all had this speculation about what did Ilya see? It was the whole thing. What did Yan see? Became the thing. Yes, you know, after Yan quit, the answer is what they saw was a company that's not committed to safety, that's unwilling to put its money where its mouth is, that has a culture that is hostile to safety efforts, and that is pivoting towards the shiny new product. It's a product company, it's a scaling startup, it's devoted to making money. There's nothing wrong with making money, but in this case they're building smarter-than-human intelligence as their explicit company mission and goal. And as LeCun points out and has repeatedly acknowledged, this is not a safe thing to do. This is not a default thing that will go well by accident, right? We can debate how likely this is to go badly with good real efforts to make it go well, but I think any reasonable person can see that there's a very good chance that if we do not put in the effort, we do not do the work, that things would then go badly. And it doesn't have to necessarily involve some sort of specific AI takeover or existential risk scenario. It simply means this could go very badly for people as experienced by people on the planet Earth. And we have to think carefully about these questions. But what's clear is that OpenAI leadership, as embodied by Sam Altman, increasingly is not doing this, and simply in the last few months is not honoring their commitments.
是的,对我来说,算力问题非常令人不安。如果没有这个,事情会难解析得多。我觉得显然人们可以有各种分歧,你可以想象 Sam Altman 的辩护是:‘你在超级对齐团队做你的事,我们在产品团队做我们的事。这本身有什么问题?’但是,如果你连做对齐工作的算力都拿不到,那就不只是承诺 X 的问题了,我们也明白了为什么是 Yan。
Yeah, that compute one for me is pretty troubling. I could try to make a... in the absence of that, it would be a lot harder to parse. I feel like obviously people can have all sorts of disagreements and you can imagine the Sam Altman defense being like, 'You're doing your thing over here in the superalignment team, we're doing our thing over here in the product team. Why is that inherently a problem?' But yeah, if you can't get the compute to do the alignment work, yeah, it's not just a promise X and we got why it's Yan.
具体来说,你说你承诺给我们 X,我们需要 X,而 N 远大于 1,才能完成我们的工作,但我们却得不到。而超级对齐的承诺,听起来很多:我们当前可用算力的 20%。但四年下来这并不多,因为两年后 OpenAI 的算力会是现在的 10 倍,除非发生非常奇怪的事情。所以从长期来看,这是一个相当小的、非常适度的承诺。相比之下,在某些安全至上的类似行业中,大部分研究成本、大部分开发成本都变成了安全成本,对吧?而且那只是常规层面,只是确保工厂不熔毁的安全水平。所以我不理解。我不明白为什么给 Jan 和像他那样的人提供算力不是简单的商业好决策。相对于这个结果,如果我们确实在谈论大约 5%的算力,那么那只能让你比用那些算力快 5%,对吧?这看起来确实是一个非常奇怪的决定,让人离开,让这件事因为几个百分点的算力可用性而闹得这么大。
Specifically saying you promised us X, we needed X over N, where N is a lot more than one, to do our work and we couldn't get it. And the super alignment commitment, like it sounds like a lot: 20% of our currently available compute. That's not a lot over four years because two years from now OpenAI will have 10 times as much compute, for sure, unless something very strange happens. So over the extended period, this is a reasonably small, very modest commitment. As opposed to in certain similar industries where safety is paramount, most of research costs, most of development costs become safety, right? And that's for mundane level, just make sure the plant doesn't melt down levels of safety. So I don't understand it. I don't understand why it's not simply good business to give Jan and people like that their compute. Certainly relative to this outcome, it seems if we are indeed talking about something like 5% of compute, then that could only allow you to move, you know, 5% faster, right, than you could with that compute. It does seem like a very strange decision to allow people to walk and allow this to become this big of a story over a couple percentage points of compute availability.
是的,伊恩,压垮骆驼的往往是最后一根稻草,对吧?很明显,他被排除在决策之外,无法发挥他之前在公司内部扮演的大使角色。如果你读字里行间,甚至不用读字里行间,就是读各种文章和报道的字面意思,人们开始反对这个想法本身。董事会斗争以实际的方式体现了这场斗争。无论这是否公平,是否应得,是否与任何哲学或方法有关,都让人们反对他们。每个离开的人都会让你更加反对他们。解雇,你有一个连锁反应。每个离开的人,你都会失去更多信任。他们为什么离开?什么导致他们离开?事情是怎么处理的?是的,他大概就是觉得,我受够了,我们得不到我们需要的东西,你们需要警钟。我不能坐在这里假装我有我需要的资源,因为我没有。而且从他谈论的方式来看,他们也没有投入一家处于他们位置的公司需要在所谓常规安全上投入的资源。事实上,我没有专门调查过,但你在他们的报告中看到,GPT-4o 拒绝不当请求(比如制造炸弹)的倾向比 GPT-4 Turbo 或其竞争对手低得多。不知何故,这个新发布的模型在越狱、常规危害方面非常不稳健。它对你的请求非常顺从。我没有尝试过越狱或红队测试它,因为我为什么要呢?我认为在大多数情况下,让它被越狱并不是特别危险的事情。对于那些足够关心的人来说,它已经被越狱了。但非常清楚的是,GPT-4o 的开发过程并没有进行稳健的尝试来使其在传统意义上安全。他们基本上只是认为这个模型拥有的能力并不那么危险,所以他们就不太关心了。
Yeah, Ian, it's almost always the straw that breaks the camel's back, right? Like it's clearly he was being shut out of decisions, couldn't play his ambassador role that he had previously played inside the company. If you're reading between, but not even between the lines, which is reading the lines of various articles and reports, that people were turning against the very idea. The board battles embodied this struggle the way that they played out in practice. Whether or not this is at all fair or deserved by anyone or any philosophy or any approach or anything, turn people against them. Every person that leaves turns you fever against them. The firings, you have a cascade. Every person that leaves, you lose more trust. Why did they leave? What caused them to leave? How was that handled? And yeah, like he presumably was just like, I'm fed up, we're not getting what we need, you need a wake-up call. I can't just sit here and try to pretend that I have the resources I need because I don't. And the way he talks about it, they're just not investing what a company in their position needs to invest in what I call mundane safety either. And in fact, I haven't investigated this per se, but you see this in their reports that GPT-4o has much less tendency to refuse inappropriate requests like building a bomb than GPT-4 Turbo or its rivals. Somehow this new model that was sent out is actually just very not robust in the jailbreak, mundane harm sense. It just is very willing to your requests. And I haven't tried to jailbreak or red team it in this sense because why would I? I don't think it's particularly a dangerous thing for the most part to have it be jailbroken. It was already jailbroken for those who cared enough. But what's very clear is that the GPT-4o development process did not involve a robust attempt to make it safe in a conventional sense. They mostly just decided the capabilities that this model possessed were not so dangerous and they just weren't going to give it much care.
那么,我要把你拴在 OpenAI 的围栏上吗?我们该怎么办?我认为我们必须继续观察,直到我们看到。Sam 可以用一系列惊人的承诺来回应。他可以用一批新员工、他引进的人来回应。他可以做很多事情。所以你知道,我们总是抱有希望。尽管 AI 发展很快,但我不认为我们需要本周就采取行动,对吧?部分原因是,从 Jan 的证词中可以清楚地看出,我们不需要本周就行动。如果 Jan、Ilya 和其他人担心即将发生的事情,他们会说出来的。从这份声明中可以清楚地看出,这些都是长期担忧,是由无论你做什么,船都会慢慢转向的事情驱动的。所以我们有时间去修复。但除非有相反的证据,考虑到我们了解的一切,我认为我们只能假设 OpenAI 实际上是一个完全以营利为目的的企业,一个完全由 Sam Altman 运营的、快速增长、快速行动、打破常规、曲棍球棒曲线式的初创企业。他们没有认真对待安全问题,他们的内部文化对安全理念、对担忧事物的理念充满敌意。因此,我们应该预期他们会糟糕地处理未来,预期他们没有为即将到来的事情做好准备,极其漫不经心。而且有可能在一段时间内,漫不经心是可行的。一个漫不经心的 GPT-5 很可能是世界上发生的最好的事情,如果可能的话,如果它们没那么危险的话。我们谈到了不适合工作场合的内容,谈到了色情、血腥和其他能力。是的,让这些事情发生可能很好。可能完全是净正面,对吧?我实际上期望如此。而且 Pliny 越狱了所有主要模型。他完全越狱了它们全部,他在发布演讲期间大约两分钟内就越狱了 GPT-4,对吧?就像他一获得访问权限就做了,因为他的第一个直觉就成功了,为什么不呢?但在某个时候,这就不再是可接受的情况了。在某个时候,这些东西将具有高度能力,我们将不得不担心社会影响,然后是灾难性风险,然后是存在性风险。而这一切都是概率性的,因为你不知道它什么时候会发生。比如当 GPT-5 出来时,你可能根本不会处于存在性危机中。但你会给这个说法加几个九?如果你的答案超过两个,你就疯了,对吧?我可能会加一个九。所以,你可以来哭诉,说你说只有 97%,结果没发生。我会说,好吧,算你运气好,这次合作比你的差一点,我们继续赌体育赛事吧。但我不知道还能说什么。
So I'll see you chain to the OpenAI fence or what do we do? What do we do? I think we have to treat going forward until we see. Sam could respond with an amazing set of commitments. He could respond with a new set of hires, people he's bringing in. He could do any number of things. So you know, we're always hopeful. And as much as AI moves fast, I don't think we need to move this week or anything like that, right? And part of it is that it's very clear from Jan's testimony that we don't have to move this week. That if Jan, Ilya, and others were concerned about something imminently happening, they would have said so. It's very clear from this statement that these are long-term concerns driven by things that the ship gets steered slowly no matter what you do. So we have time to fix it. But barring evidence to the contrary, given everything that we've learned, I think we just have to assume that OpenAI is functionally a fully for-profit business, fully a grow fast, move fast and break things, hockey stick graph startup business run by Sam Altman, who is running it on that basis. And that they are not taking the safety problem seriously, that their culture internally is hostile to the idea of safety, the idea of worrying about things. And that therefore we should expect them to handle this future badly, that we should expect them not to be prepared for what is to come, to be extremely cavalier. And there's the possibility that for a while, cavalier works out. A cavalier GPT-5 might well be the best thing to happen for the world if it's possible, if they're just not that dangerous. We talked about not safe for work, we talked about sex and gore and other capability. And yeah, it might just be good to let that stuff happen. It might be completely net positive, right? I actually expect that. And Pliny jailbroke all the major models. He jailbroke all of them fully, and he jailbroke GPT-4 in about two minutes during the announcement speech, right? Like the moment he got access to it, because his first hunch just worked, because why wouldn't it? But at some point, that's going to stop being an acceptable situation. At some point, these things will be highly capable, and we're going to have to worry about societal implications, and then catastrophic risk, and then existential risk. And all that's probabilistic because you don't know when it's going to happen. Like when GPT-5 comes out, you're probably not going to be in an existential situation at all. But how many nines are you going to put on that statement? And if your answer is more than two, you're crazy, right? And I would probably put one nine. So like, you can come cry and say you said it was only 97% and it turned out not to happen. I'm like, well, chalk it up to much flight, slightly worse collaboration than you on this one, and let's keep betting on sporting events. But I don't know what else to say.
Brave 搜索索引为开发者提供实惠的访问权限,这是一个独立的网络索引,拥有超过 200 亿个网页。那么 Brave 搜索索引有何独特之处?第一,它完全独立且从零构建,这意味着没有大科技公司的偏见或高昂价格。第二,它基于真实人类的实际页面访问(当然匿名收集),过滤掉了大量垃圾数据。第三,该索引每天刷新数千万个页面,因此始终提供准确、最新的信息。Brave 搜索 API 可用于组装数据集以训练你的 AI 模型,并在推理时帮助进行检索增强,同时保持开发者优先的定价。将 Brave 搜索 API 集成到你的工作流程中,意味着更合乎道德的数据来源和更具人类代表性的数据集。在 brave.com/api 免费试用 Brave 搜索 API,每月最多 2,000 次查询。
Affordable developer access to the Brave Search Index, an independent index of the web with over 20 billion web pages. So what makes the Brave Search Index stand out? One, it's entirely independent and built from scratch, that means no big tech biases or extortionate prices. Two, it's built on real page visits from actual humans, collected anonymously of course, which filters out tons of junk data. And three, the index is refreshed with tens of millions of pages daily, so it always has accurate, up-to-date information. The Brave Search API can be used to assemble a data set to train your AI models and help with retrieval augmentation at the time of inference, all while remaining affordable with developer-first pricing. Integrating the Brave Search API into your workflow translates to more ethical data sourcing and more human-representative data sets. Try the Brave Search API for free for up to 2,000 queries per month at brave.com/api.
我猜你看了这周 Dwarkesh 对 John Schulman 的采访。据报道,他现在全面负责安全工作。我猜他之前就已经负责了,正如 Dwarkesh 在采访中介绍的,他负责模型的后训练。正如今天一些文章所述,他负责让今天的模型安全。现在他将承担起更大的责任,涉及全局、长期的安全。在这种情况下,我认为那个采访会让你认真思考。显然你可以自己判断,但有好几次你听到他之前在做的事情。这就是你看到的那个做日常安全工作的普通人,他的工作不一定比竞争对手好或差,至少在我还没评估 40 之前。是的,还有报道说可能存在问题,但这与他现在被要求解决的问题完全无关。如果他当时工作做得好,那我就不一定会提到他们对下一代模型没有为此做好准备感到担忧。这也是他工作的一部分。但是,是的,认为可以在你的 AGI 甚至 ASI 上使用后训练,让它变得 HHH 或有用之类的,我认为只依赖这个基本上是白日梦。事实上,你甚至需要在训练过程中担心这些事情。但只是这种策略,当然可能有办法,但他目前使用的策略正是那种他自己也明确承认行不通的东西,无论是在网上还是在我旁边五英尺的谈话中。他在 80,000 播客上解释过为什么 RLHF 不是这个问题的解决方案。他还有其他想尝试的解决方案,但我认为它们行不通。我试图和他辩论,向他解释为什么行不通,但我没能足够有说服力地论证,他没有接受。我认为他不接受是合理的,因为我在提出一些与他的世界观截然不同的非常大胆的主张,而且我没有足够具体地支持它们,因为缺乏技术训练对我来说很难做到。但我对问题的思考层次与那些只考虑后训练的人不同。我还没看那个采访,我想谨慎对待,但如果你说它会让我思考,我相信你是对的,它会让我思考。
I assume you have seen the Dwarkesh interview with John Schulman out this week. So as it's reported now that he's taking on the responsibility for safety at large. I guess he was already responsible, as Dwarkesh presented him in the interview, he was responsible for post-training the models. As described in some of these articles today, he's responsible for making today's models safe. Now he's going to have this kind of additional responsibility rolling up to him of big picture, long-term safety. In that context, that interview I think is going to give you serious pause. And obviously you can evaluate it for yourself, but there were multiple moments you're just hearing what he was doing previously. This is you took the mundane safety guy who is doing a job with mundane safety that is not necessarily better or worse than the rivals, at least until I haven't evaluated 40. Right, in the sent and there's again report there's potentially a problem, but that's just completely irrelevant problem to the problem he is now being asked to solve. And if he was doing that job properly, then like I wouldn't necessarily have mentioned that they have these concerns about the next generation of models not being prepared for that. Be part of his job as well. But yeah, the idea that you're going to use post-training on your AGI or even your ASI, right, to render it like HHH or like net useful or any of those terms, I think that's basically a pipe dream to rely only on that. And in fact, you have to worry during even the training regimen of some of these things. But just the kind of strategy certainly there may be a way, but the strategies that he's been using so far are the kind of things that like I myself acknowledged explicitly, like both online and literally five feet away from me during a talk, would not work. Right, he took on the 80,000 podcast, he explain why RLHF is not a solution to this problem. And he has other proposed solutions that he wants to try, where I don't think they'll work. And I tried to debate him and explain to him why they wouldn't work, and I was unable to make my case sufficiently convincingly, and he didn't buy it. I think it was reasonable for him not to buy it in the sense that like I had, I was making some very bold claims that are very different from his worldview, and I didn't back them up specifically enough because it was hard for me to do without the technical training. But like I was thinking about the problem right on a different level than someone who's thinking about post-training is thinking about the problem. So I haven't seen the interview yet, again I want to be very careful with it, but yeah, if you think it's gonna give me pause, I'm confident you're right, it's going to give me pause.
另外,对于可能正在听这个节目但没有时间或兴趣去深入了解的人,这绝对值得一看,它在很多方面都非常有趣。他说的一件事是,他预计所谓的 AGI 会在两到三年内实现。Dwarkesh 问他是否可能明年就实现,他说不认为会,那会很令人惊讶,但两到三年他基本愿意认同。然后问了一些非常根本的问题:当这发生时会发生什么,我们该如何应对?或者有一次他说,我不确定,你在这里概述的计划听起来不太稳健。在多个时刻,他都说,是的,我真的没有一个很好的解释来说明事情会如何发展,或者希望届时我们能与其他领先的开发者合作,是的,我们目前还没有一个稳健的计划。次优计划:最好的计划是真正弄清楚你要做什么以及我们如何处理,次优计划是知道自己不知道,从空白的初学者心态开始。所以如果他第一天来就说我们没有计划,没有解决方案,我们不知道如何对齐这个东西,即使我们对齐了也不知道怎么处理,我们不知道社会如何应对,我们不知道如何从 AGI 过渡到 ASI,即使我们处理了对齐和中间阶段,我们不知道这些事情,我们迷失了,我们需要弄清楚,他从头开始,我完全满意这个答案,对于一个能力很强的人来说。
A couple well just for folks who might be listening to this and might not have time to go through that or inclination to, it is definitely worth it, it's very interesting in multiple ways. One of the things he says is that he's expecting quote unquote AGI on kind of a two to three year time frame. Dwarkesh asks him like could it be as soon as next year and he's I don't think so, that would be surprising, but two to three he was pretty much willing to co-sign on. And then asked a number of questions that were like pretty fundamental: what happens when this happens, how are we going to deal with it? Or at one point I'm not sure that's a doesn't sound like a super robust plan that you're outlining here. And at multiple points he was like yeah I don't really have a great account for how that's going to go or hopefully we'll be able to work together with the other leading developers in that situation and yeah we don't really have a robust plan for that at this point. Second best plan right: the best plan is to actually figure out what you're going to do and how we're going to handle this, and the second best plan is to know you don't know, right, to start with a blank beginner mind. So if he comes to this, if he comes day one and he says we don't have a plan, we don't have a solution, we don't know how to align this thing, we don't know what to do with it if we did align it, we don't know how society can handle this, we don't know how to make the transition from AGI to ASI again even if we handle alignment, even if we handle the interim, we don't know any of these things, we're lost, we need to figure this out and he starts from scratch, I'm perfectly happy with that as the answer for a person who's highly capable.
他创立了它,做了很多令人印象深刻的事情。就像 Ilya,我认为 Ilya 深切关心这些问题,并理解问题的深度、范围和风险。但我一直认为,现在仍然认为,从 Ilya 公开的言论来看,我不认为他思考的任何东西接近你需要思考的或会起作用的东西。他的东西行不通。但我认为他相当脚踏实地,并且很努力。他的东西常常让我觉得,好吧,这对我来说似乎是一种误解。感觉你经常以一种模糊、不够有针对性的方式思考这个问题,而且这种方式在遇到敌人时根本站不住脚,对吧?当他与 Jan 和其他人运行他们的超级对齐团队,撰写论文,进行实验,并真正尝试让这些东西工作时,他们会明白他们的计划行不通,然后他们会要么找到修改方法使其可行,要么抛弃它们并尝试新的。因为 Ilya 以这样的态度闻名:我会尝试尽可能多的方法,抛弃我认为我知道的一切,直到我让这个东西工作。我致力于让这个东西工作,对吧?这正是你希望看到的,这也是为什么我非常相信 Ilya 和 Jan 会找到办法。并非总是如此,因为我不认为这个问题在理论上一定可以由人类在合理的时间范围内解决。
He's profounded it, he's done a lot of impressive things. Like Ilya, I thought Ilya cares deeply about these issues and appreciates the depths and scope and stakes of the problem. But I always thought, and I still, I mean going off of Ilya's publicly stated remarks, I just don't think any of the things he was thinking about are anywhere near like what you need to be thinking or what would work. His stuff wouldn't work. But I thought like he was reasonably grounded and trying hard. His stuff just felt like it often okay that just seems like a misconception to me. It feels like you're thinking about this problem in a kind of fuzzy, not sufficiently geared way often, and in a way that like just wouldn't survive an encounter with the Enemy, right? When he looks at when he and Jan and the other people run their superalignment team and they run these papers, they run these experiments, and they actually try to make these things work, they'll understand that their plans aren't working, and they'll either find ways to modify them so they work or throw them out and find and try new ones. Because one of the things Ilya is famous for is the kind of attitude of you know I'm going to try this as many ways as I have to and throw out everything I think I know until I make this thing work. I'm committed to making this thing work, right? And that's what you love to see, and that's why I had a lot of faith that Ilya and Jan would find a way. Not every time, because I don't think this problem is even necessarily theoretically solvable by humans in a reasonable time frame.
不管态度如何,但我认为他们成功的机会相当大,因为拥有大量资源和大量时间——四年在某种程度上是很多时间——他们至少会找到一些不对齐 AI 的方法,然后我们可以再试一次。而他们的第一种方法在我的模型里肯定行不通,这没关系。这些想法混合了合理、有希望和绝望的,但没人知道如何解决这些问题,所以我不能因为你对我认为解决不了问题的想法感到兴奋而生气。我有什么更好的建议吗?你尝试一些东西,学到一些东西,再试一次。我认为研究这些泛化问题、这些监督问题,有相当大的机会会为寻找可能更有希望的替代方案提供思路。如果这些解决方案都没有希望,如果整个解决方案类别完全无望,就像最坏的情况那样,第一眼看上去就那么糟糕且无法修复,那么这些世界在很多方面都是相当黯淡的,因为这切断了很多人的计划。所以我留有余地,认为还有可以操作的空间。
With any attitude but I thought they had a reasonably good shot because with a lot of resources and a lot of time, and four years is somehow a lot of time, they would figure out at least some ways not to align an AI, and then we get to try again. And the fact that their first way definitely won't work in my model is fine. They were a mix of ideas that were reasonable and promising or hopeless, but nobody knows how to solve these problems, so I can't really get that mad at you for being excited by ideas that I think won't solve these problems. What do I have? A better suggestion? You try something, you learn something, you try again. And I think there's a decent chance that looking at these generalization questions, looking at these supervision questions, will inform your approach to trying to find alternate solutions that might themselves be more promising. If none of those types of solutions, if that entire solution class is entirely hopeless, like the worst-case scenario that is as bad as it looks to me on first glance and there's no fixing it, then those are pretty bleak worlds in many ways because that cuts off a lot of people's plans. So I'm leaving room for there being places to maneuver there.
有了这种清晰的认识,在我们等待奥特曼和领导层的正式回应时,你认为人们还应该做什么?开玩笑,但也许我不完全是在开玩笑,关于把自己锁在 OpenAI 的围栏上。作为参考,这周有一个人这么做了。还有普通的消费者抗议。既然 ChatGPT 现在是免费的,取消我们的账户会很难,但我们可以号召应用开发者抵制并切换到 Claude 之类的。我们显然可以公开表示至少总体上对 SB 1047 持积极态度。我们可以考虑更大力地支持它,或者提出可能的修正案来加强它。比如让非贬低条款非法化,这可能是一个有趣的点。
With this clarity, and while we await a proper response from Altman and leadership, what else do you think people should be doing? Joking, but maybe I'm not entirely joking about chaining oneself to the OpenAI fence. There for reference, there was a person who did that this week. There is mundane consumer protest. Now that ChatGPT is free, it's going to be hard to cancel our accounts, but we could rally app developers to boycott and switch to Claude or something. We could obviously, on the record, be at least generally positively disposed to SB 1047. We could think about supporting that even more forcefully, or suggesting possible amendments to strengthen it. Something like the non-disparagement clause being made illegal could be an interesting one.
当我审视 1047 号法案时,我写了分析。我主要是在寻找它可能过于严格的地方。我指出了该法案可能过分的几种方式,因为它有这些严重的缺点,除非我们为这些缺点获得很多回报,否则它在政治上更难通过,更难获得支持,更难获得合作。而且没有人真的想搞垮经济,没有人想真正减缓日常效用。不是没有人,但我不想。那么,我们如何加强这项法案呢?例外是衍生与非衍生定义条款,我认为这只是一个错误。这完全是法案中的一个重大定义错误,使它在各个层面都变得更糟。我们需要修复它,因为如果我能把所有责任都推给你,那会让你的处境很糟糕,但也会让我可以规避所有的安全保障。然后我可以忽略任何事情,这也不好,对吧?我们不能让任何人作弊。我们必须阻止这种情况。所以其他一些点是关于我们如何削弱它?如何——我不知道“削弱”这个词是否恰当——但如何澄清这项法案?如何防止潜在的过度干预或误解?其他人指出了法案可能过于薄弱的地方。每个人都在抱怨刑事责任,但它只涉及伪证。也许这还不够,对吧?有人可能会问,也确实有人问了。我认为目前它处于我们想要的位置,而且我认为即使对伪证条款的反应也表明这是一条高压线,人们对这类事情非常害怕,所以最好不要——它可能会——对好人来说,停止触碰事情可能会产生非常糟糕的动态,或者他们会恐慌,他们不希望那样发生。
When I look at 1047, I wrote it up. I was mostly looking for ways in which it was too strong. I identified a number of ways in which this bill might go too far in the sense of it has these serious downsides, and unless we're getting a lot in exchange for these downsides, that makes it politically harder to pass, it makes it harder to get buy-in, it makes it harder to get cooperation. And nobody actually wants to tank the economy, nobody wants to actually slow down the mundane utility. Not nobody, but I don't. And so how can we strengthen this bill? The exception being the derivative versus non-derivative definition clause, where I thought this is just a bug. It's literally just a major definitional mistake in this bill that is just making it worse on every level. We need to fix it because if I can pass off all of my blame to you, that makes your situation terrible, but it also makes all the safety guarantees something I can skirt. I can then ignore anything, and that's not good either. Right? Like we can't let anyone cheat. We have to stop this. So some of the others were about how do we weaken this? How do we—I don't know if weaken is the right word—but how do we clarify this bill? How do we prevent potential overreach or misinterpretation of this bill in a stronger sense? And other people pointed out ways in which the bill was potentially too weak. Everyone was complaining about the criminal liability, but it's only under perjury. Maybe that's not enough, right? One could ask, potentially, and some people did. I think for now it is where we want to be, and I think the reaction to even the perjury showed us that this is just a third rail, and people just get so scared of such things that let's not—it's just gonna—it might have very bad dynamics for the good people to stop touching things, or they panic and they don't want that to happen.
但是,是的,法案中有举报人条款,如果你不执行竞业禁止协议,对吧?这似乎比竞业禁止严重得多。普通的竞业禁止,现在不仅加州永远不执行,而且现在对于高薪但最知名的员工,全国范围内都将是非法的,因为 FTC——除非那没通过。但是,是的,我很难相信允许一家公司以签署终身全面非贬低条款作为条件,扣押某人大部分财富,符合美国公众的利益,尤其是在一个错误事关重大国家利益的领域。如果 OpenAI 的安全有问题,我们需要知道。如果 OpenAI 的文化在其他方面有问题,我们需要知道。这是显而易见的。如果你有举报人条款,那么如果你在法律上不被允许举报,你怎么举报?代价是数百万美元的股权,对吧?你不能出售,所以他们可以直接没收。所以甚至他们不一定需要起诉你,他们可以在那种情况下直接没收。谁知道他们还可能威胁或要挟人们什么?我们不知道,再说一次,因为不能谈论。所以我无法知道这些事情。所以我们必须在某种程度上假设最坏的情况,因为我们不能谈论。是的,我认为说 AI 公司不应该能够签署涉及公司某些方面的非贬低条款是非常合理的,当然,可能普遍如此。比如,我认为——我们两个人达成协议,同意无论发生什么,我都不能说你的坏话,你也不能说我的坏话,这有什么好处?我理解在某种意义上这对我们更好,但是,人们不能那样说话似乎不太好。如果我们不执行违背公共利益的合同,这似乎是一个主要的关注点,对吧?尽管我有自由意志主义的本能,认为避免合同是不好的,但显然这似乎是一个合理的考虑点。
But yeah, like we have whistleblower clauses in the bill, and if you don't enforce non-compete agreements, right? This seems so much worse than a non-compete. An ordinary non-compete, which now not only California is not enforcing them forever, but now they're going to be illegal for high-paid but the most prominent employees in the entire country, because the FTC—unless that doesn't go through. But yeah, I have a hard time believing that it is in the interests of the public of the United States to allow a company to hold most of somebody's wealth hostage to signing a lifetime full non-disparagement clause on the company they are leaving, in an area in which things that are wrong are in the vital national interest. If there's something wrong with the safety at OpenAI, that's something we need to know. If there's anything wrong with the culture at OpenAI in other ways, that's something we need to know. It's obvious. If you have a whistleblower provision, this is like how do you blow the whistle if you're not legally allowed to blow the whistle, where the price is millions and millions of dollars in equity, right? That you can't sell, so they can just confiscate it. So even they don't even have to sue you necessarily, they can just confiscate it in that situation. And who knows what else they might be threatening or holding over people? We don't know, again, because can't talk about it. So I can't know these things. So we have to assume the worst in some senses because we can't talk about it. And yeah, I think it would be very reasonable to say that AI companies should not be able to sign non-disparagement clauses as pertains to certain aspects of the company, certainly, and potentially universally. Like, I think it's just—why is it good for the two of us to get to an agreement where we agree that no matter what happens, I can't never say anything bad about you and you can't never say anything bad about me? I understand why it's better for us in some sense, but you know, people not being able to talk in that way just doesn't seem great. And if we're in the business of not enforcing contracts that are against the public interest, this seems like a prime place to look, right? Even though I have libertarian instincts that avoiding contracts is bad, yeah, this seems to be a reasonable place to consider that obviously.
但除此之外,好吧,如果你不一定自己做安全研究,那么必须有人确保你做了,对吧?必须有人进行检查。这是怀疑的理由。对吧?如果我要信任一家公司,我需要能够信任他们的承诺,我需要能够信任他们的声明、他们的证词,而且这需要受到某种惩罚。对吧?这就是整个想法。对吧?如果你撒谎,想法是你不需要做任何事。你必须告诉我们你在做什么,如果你撒了弥天大谎,你必须承担责任。这就是 SB 1047 关键条款的内容。它们是关于说明你在做什么,如果你没做对就要负责,以及说明你的逻辑,如果你撒谎就要负责。
But beyond that, well, if you're not going to do the safety work yourself necessarily, someone has to make sure that you do, right? Someone has to be doing the checking. And this is a reason to doubt. Right? If I'm going to trust a company, I need to be able to trust their commitments, I need to be able to trust their statements, their testaments, and I need that to be under some sort of punishment. Right? Like that's the whole idea. Right? If you lie, the idea is that you don't have to do anything. You have to tell us what you are doing, and you have to be held responsible if you lied your ass off. That is what the key provisions of SB 1047 are about. They're about saying what you're doing and being responsible if you didn't do it right, and saying what your logic is and being responsible if you lied.
如果你的逻辑是谎言,那就毫无意义。故意无视,不可理喻——这就是问题所在。而且还要有一个机制,如果你发现房间里确实存在灾难性风险,你可以关闭模型。既要让他们有能力至少在本地关闭它,也要让你能下令这么做。这些都是基本的东西。在得知这些信息后,这些事情现在比以往任何时候都更重要。但坦率地说,我认为我们必须明白,除非被证明并非如此,否则 OpenAI 现在比一周前或六个月前要清晰得多。对吧?当他们宣布超级对齐,当他们发布相当不错的——我自己也这么认为——他们发布了一个相当不错的准备框架,对吧?他们做了一些好事。他们雇佣了一批优秀的人。他们有一群在圈子里活动的人,他们确实花了很多时间以正确的方式说话,提出正确的问题,即使我不同意他们的具体信念。而现在很多都没了。从安全角度来看,他们的信誉已经毁了。所以我认为现在不那么令人困惑了。是的,我认为如果你在某种程度上可以选择使用谁的技术,而你选择了 OpenAI 的,那么这就是你正在考虑的一部分,你正在做的事情的一部分。而且我认为这也意味着你承担了风险,一个具体的风险,因为我认为你不应该再信任他们在这个世界上的日常安全。我认为——我仍然只是再次强调,GPT-40 出问题的地方并不多,我并不那么害怕。但如果他们没有安全文化,你需要一种安全思维来构建 AI,让这些 AI 做你想让他们做的事情,并且不会在普通、正常的一天里让你难堪。你需要思考这些问题,解决这些问题,并给予它们应有的尊重。所以你知道,Anthropic 的商业案例——'我们深切关心这个,我们有这种文化'——他们显然在员工中有这种文化,对吧,他们鼓励这种担忧,或者他们有这种担忧——以及'我们会确保当你使用我们的产品时,你得到你想要的东西,而不是其他你没想到也不想要的东西'——就变得更有趣了,对吧。而谷歌在那个光谱上的位置,你可以自己评估。我不是说这个房间里有什么天使。我不是说有我信任的人。但程度不同。再说一次,我们看看 Altman 下周会说什么。
If your logic is a lie, it just makes no sense. Willful disregard beyond the pale—that's what it's about. And also just having a mechanism where if you discover that there's actually catastrophic risk in the room, you can get the model shut down. Both that they have the ability to shut it down at least locally and that you can order that. These are the fundamental things. This feels about so these things seem important now more than ever, right, in the light of this information. But frankly speaking, yeah, I think we have to understand that until proven otherwise, OpenAI is much less of a confusion than it was a week ago or six months ago. Right? There were reasonable arguments to be made when they announce superalignment, when they put out their reasonably good—them myself, yeah—they put out a reasonably good preparedness framework, right? Like they've done some good things. They've hired a bunch of good people. They had a bunch of people who moved in circles where like they're credibly spending a lot of their time talking the right ways, asking the right questions, even if I don't agree with their specific beliefs. And a lot of that's just gone now. Their credibility is shot from a safety perspective. And so I think it's a lot less confusing now. And yeah, I think that if you have a choice in whose technology to use in some sense at this point, and you choose to go with OpenAI's in a way that matters, well, this is part of what you're considering, part of what you're doing. And I think it also means that you are taking on a risk, a concrete risk yourself, because I don't think you should necessarily trust their mundane safety in this world going forward. I think it's—I'm still just again, there's just not much something can go wrong with GPT-40 that I am that scared of. But if they don't have a culture of safety, you need a security mindset to build AIs, to make these AI do the things you want them to do and have it not go up in your face even in an ordinary, normal, one-day way. You need to be thinking about these problems and working on these problems and giving them the respect they are due. And so you know, the Anthropic business case for 'we deeply care about this and we have a culture of this'—which they clearly do amongst their employees, right, where they encourage this concern or they have this concern—and 'we're going to make sure that when you use our product, you get what you're trying to get and not something else you did not expect and did not want' becomes a lot more interesting, right. And then where Google lies on that spectrum, you can evaluate for yourself. I'm not saying there are any angels in this room. I'm not saying there's anybody that I trust. But there are levels. And again, like we'll see what Altman says next week.
嘿,我们稍后继续采访,先听一段赞助商的话。嘿,我是 Eric Torenberg。我越来越常听到创始人想要盈利并少花钱多办事,尤其是在工程方面。听着,我和其他人一样喜欢你那 30 岁的前 FAANG 高级软件工程师,但老实说,我再也雇不起他们了。各地的创始人都在转向全球人才,但规模化操作从招聘到面试再到实地运营和管理真是麻烦。这就是为什么我和 Sean Lanahan 合作,他在越南建立高水平工程团队已有 5 年多,帮你无痛获取全球工程人才。Squad,Sean 的新公司,负责全球人才的招聘、法律合规和本地 HR,所以你不用操心。团队遍布亚洲和南美,无论你在哪个时区运营,我们都能覆盖你。他们的工程师遵循你的流程,使用你的工具。他们使用 React、Next.js 或你最喜欢的前端框架,在后端他们是 Node、Python、Java 和任何技术的专家。完全披露:这比你在 Upwork 上找到的每周工作 2 小时却收你 40 小时费用的随机人要贵,但你会以典型成本的一小部分获得优质质量。我们的工程师是经过筛选的前 1%人才,每天为你努力工作。提高你的速度而不增加消耗。访问 choosesquad.com 并提及 Turpentine 以跳过等待名单。OmniKey 使用生成式 AI 让你一键启动数十万次真正有效的广告迭代,跨所有平台定制。我非常相信 OmniKey,所以我投资了它,我建议你也使用它。使用代码 REV 获得 10%折扣。
Hey, we'll continue our interview in a moment after a word from our sponsors. Hey, Eric Torenberg here. I'm hearing more and more that founders want to get profitable and do more with less, especially with engineering. Listen, I love your 30-year-old ex-FAANG senior software engineer as much as the next guy, but honestly, I can't afford them anymore. Founders everywhere are trying to turn to global talent, but boy is it a hassle to do at scale from sourcing to interviewing to on-the-ground operations and management. That's why I teamed up with Sean Lanahan, who's been building engineering teams in Vietnam at a very high level for over 5 years, to help you access global engineering without the headache. Squad, Sean's new company, takes care of sourcing, legal compliance, and local HR for global talent so you don't have to. With teams across Asia and South America, we can cover you no matter which time zone you operate in. Their engineers follow your process and use your tools. They work with React, Next.js, or your favorite frontend frameworks, and on the backend they're experts at Node, Python, Java, and anything under the sun. Full disclosure: it's going to cost more than the random person you found on Upwork that's doing 2 hours of work per week but billing you for 40, but you'll get premium quality at a fraction of the typical cost. Our engineers are vetted top 1% talent and actually working hard for you every day. Increase your velocity without amping up burn. Head to choosesquad.com and mention Turpentine to skip the waitlist. OmniKey uses generative AI to enable you to launch hundreds of thousands of ad iterations that actually work, customized across all platforms with a click of a button. I believe in OmniKey so much that I invested in it, and I recommend you use it too. Use code REV to get a 10% discount.
你怎么看待重新审视第三方测试的概念?这也是我在采访州参议员 Scott Wiener 之前聊过一点的事情,他当然正在推动这项法案。目前的法案只是说'包括适当的第三方审计员',现在看起来还好,也许我们稍微加强一下。但我们有英国 AI 安全研究所,没有完全获得访问权限。现在我们有人辞职以示抗议。你是个好的游戏设计者。你会如何设计这个游戏,让合适的人获得合适的访问权限来做合适的测试,并且不会崩溃成——
What do you think about revisiting the notion of third-party testing as well? This was something we chatted a little bit about prior to my interview with State Senator Scott Wiener, who of course is sponsoring the bill. The current bill just says 'including third-party auditors as appropriate,' and now it does seem okay, maybe we step that up a little bit. But we got this AI UK AI Safety Institute not quite getting the access. Now we have the resignation clearly in protest. You're a good game design guy. How would you think about designing that game so that the right people get the right kind of access to do the right kind of testing, and it doesn't collapse into—
我在分析 Anthropic 的准备框架和负责任的扩展政策时强调过:如果规则的精神得到尊重,如果这些公司的人关心安全,他们有安全文化,并且他们不只是把这看作一堆要打勾的复选框,以便通过合规、发布产品并满足批评者,而是真正关心这个结果,那么这些政策可能非常好。当答案返回'技术上你通过了',但那是可笑的,反应是'不,等等,停下来,想想发生了什么。'完全清楚那是不行的。这里有事发生。我们需要调查这个。我们需要停止这个,对吧?然后你做出相应的反应。当然,如果人们可能为了破坏基准而作弊,比如低于某些阈值,你就做不到。如果你担心他们针对测试——他们知道会问什么问题,他们知道要检查的攻击面是什么,然后他们加强那个特定的攻击面。这些东西很深。对吧?如果你所要做的只是验证 AI 不会以特定方式失败或引起问题,我认为你就完蛋了。而且我认为没有任何一组测试能在任何合理的成本下满足这一点。这就是为什么你需要第三方测试,除非你深深信任做这件事的人。对吧?你相信他们不会教测试吗?你相信他们会寻找任何东西,而不仅仅是特定的东西吗?你相信他们会隔离这些群体吗?对吧?我可以想象一个我会信任的角色,但在我们刚刚看到的事情之后,你对 OpenAI 有那种信任吗?我知道我有,对吧?在——
I emphasized when I analyzed both the preparedness framework and the responsible scaling policy for Anthropic: these are potentially very good policies if the spirit of the rules is being honored, if the people at these companies care about safety, they have a culture of safety, and they don't just look at this as a bunch of checkboxes to get through so that they can get through compliance and release their thing and satisfy the scolds, but they genuinely care about this result. And when the answer comes back 'technically you passed,' but that's funny, the response is 'no, wait, stop, think about what's going on.' Full well that wasn't okay. There's something going on here. We need to investigate this. We need to stop this, right? And you react accordingly. And certainly you can't do it if people are potentially gaming to sabotage benchmarks to be like under some thresholds. If you're worried about them targeting the test—they know what questions are going to be asked, they know what exactly is going to be the attack surface that is checked, and they strengthen that particular attack surface. These things are deep. Right? If all you have to do is verify the AI doesn't fall or cause a problem in specific ways, I think you're just toast. And I think there's no set of tests that would be at any reasonable cost that could possibly satisfy that. And that's why you need third-party testing unless you deeply trust the people who are doing this. Right? Do you trust them not to teach the test? Do you trust them to look for anything at all, not just things that are specific? Do you trust them to isolate these groups? Right? I can imagine a role in which I would have that trust, but after what we just saw, do you have that trust in OpenAI? I know that I do, right? In the—
我绝对不这么认为,但我们看到了,你知道,我信任他们的基准测试。当 OpenAI 提出一个基准测试时,我相信他们没有在操纵基准测试。他们不会那样做。我以前信任 GPT-4 的方式现在也一样。现在可能没那么信任了,对吧?我不知道他们打算朝哪个方向发展。但正如我认为 Colin Fraser 是第一个说的那样,在我们拿到手之前,不要假设他们发明了操作系统。不要在任何方面仓促下结论,不仅仅是安全性,还有能力。因为如果你是一个试图炒作的炒作机器,那和 ChatGPT 当时的情况完全不同,对吧?ChatGPT 只是‘这是我们非常简洁命名、非常简单呈现、非常干净的东西。’我仍然给他们很大的赞誉,因为他们没有在各种方面把它搞乱。OpenAI 有很多值得喜爱的地方,但他们只是推出了这个东西,像是‘嘿,这是一个很酷的东西,看看你能用它做什么。我们不会告诉你它有多棒,我们只是把它放出来。’GPT-4 也基本是同样的方式。而现在,我认为这是第三次,他们出来说‘这是我们的惊人能力,其中一些我们还没有。’炒作,炒作,炒作。这非常不同。然后 Sam Altman 在 Google I/O 的第二天上 Twitter 吹嘘他的炒作氛围比 Google 好得多,但据我所知,他没有回应 Google 宣布的任何具体内容,也没有比较谁在构建更好的产品,甚至没有祝贺他们提供了一套很棒的产品或类似的东西。他只是说‘哦,管好那个书呆子’,基本上就是‘他们那么努力,是不是很逊?’这也不是一个好迹象。
I don't definitely don't, but we saw it like, you know, I trust their benchmarks. When OpenAI comes up with a benchmark, I trust they're not gaming the benchmark. They're not trying to do them in that way. I trust them for GPT-4 in the same way that I did previously. Now maybe not as much, right? I don't know what direction they're going to do these things in. But as I think Colin Fraser was the first person to say, let's not assume they invented the operating system from her until we get our hands on it. Let's not jump to conclusions in any way, not just safety but also capabilities. Because if you are a hype machine that is trying to hype, that is a very different world than what ChatGPT was, right? ChatGPT was just 'here is our very sterily named, very simply presented, very clean.' And I still give major profit to them for not having cluttered it up in various ways. There are things to love about OpenAI, but they just presented this thing of like, 'Hey, here's a cool thing, let's see what you do with it. We're not gonna tell you how awesome it is, we're just gonna put it out there.' And GPT-4 worked the same way, mostly. And now this is the third, I think, time they've gone out there and gone, 'Here are our amazing abilities, some of which we don't have yet.' Hyped, hyped, hyped. And that's very different. And then Sam Altman went on Twitter the day after Google's I/O to gloat about how his hype was so much of a better vibe than Google's, without, but he hasn't addressed, to my knowledge, any of the concrete things that Google announced, or any comparisons as to who's building the better product, even to just congratulate them on a great set of offerings or anything like that. He's just like, 'Oh, get a hold of the nerd,' basically, like 'they're trying so hard, isn't that lame?' That's not a good sign either.
是的,这不太好。各方面都不太好。
Yeah, it's not great. It's not great all the way around.
我想说我确实强烈注意到了从早期发布到当前发布的转变。我认为去年秋天的演示日是第一次有这样的感觉,而且那次比这次更严谨,但 GPTs 并没有真正准备好,它们运行得不太好,这是第一次感觉,天哪,你们把这个东西发布了,尽管它还没有达到可以发布的状态。我不是说安全性,没有安全问题,但它们感觉像是 vaporware,对吧?不好用。是的,检索功能不好。我认为后来有所改进,但那是几个月之后了。我自己还需要做更多测试,但我在应用开发圈子里听到的是,Assistants API 现在已经成型,检索功能确实比最初好多了,更像他们第一次发布时描述的那样。但那是相对近期的事,已经过去好几个月了。而这次似乎更加混乱,他们甚至没有明确说明我们应该得到什么,或者你知道,人们现在都很困惑。我自己也很困惑。我试图让 ChatGPT 修改一张我和我儿子的照片,但它做得不对,然后我就想,这玩意儿怎么了?然后我上 Twitter 一看,哦,好吧,我还在用旧系统。我根本不清楚我在处理什么。所以他们没有更新很多方面,也没有说清楚,很多人在抱怨,‘哦,我才意识到我用的不是新版本。’我理解他们所做的,因为我像记者一样密切关注,但对于一个普通关注的人来说,这非常不公平。
I would say I definitely strongly noticed the shift from the earlier releases to the current releases. I would say last fall the demo day was really the first time where it was like, and even that was more buttoned up than this one, but it was like the GPTs weren't really ready, they didn't really work that well, and it was the first time where it felt like, man, you guys shipped this even though it wasn't really in shape to ship. I'm not safety, there's no safety concern, but they felt like vaporware, right? Didn't work. Yeah, the retrieval was not good. I think that has been improved, but it was a few months later. And I need to do a little more testing with this myself, but what I've been hearing in my app development circles is that the Assistants API has grounded into form now where the retrieval actually does work much better than it originally did, and is more like what they described in that first release. But that's relatively recent, and it's been a number of months. And this one just seems even messier, where it's like they weren't even really clear on what we were supposed to be getting, or you know, people are just all confused at the moment. I was confused. I was trying to get ChatGPT to modify a picture of me and my son, and it was not doing it right or whatever, and then I was like, what's going on with this thing? And then I went on Twitter and I see, oh okay, I'm still using the old system. And it's just not clear to me what I'm even dealing with. So they're not updating a lot of the angles, and they're not being clear, as experienced by a lot of people complaining, 'Oh, I just realized I'm not on the new version of this thing.' I understood it from what they were doing, like I was being a journalist and paying very close attention, but to a normal person who's paying ordinary attention, it was very unfair.
还有其他想法吗?我想我的另一个问题是,有没有任何新的安全相关议程或发展,你认为值得额外关注,或者你知道我们目前可能应该加大投入的?
Any other thoughts on... I guess one other question I had was, are there any new safety-related agendas or developments that you think are worth extra attention, or that you know that we should maybe be increasing our bets on at the moment?
我一直在寻找那些看起来真的可能有效的东西,而且我总是惊讶于最有趣的事情——我认为在对齐领域完全没有被关注。你听说过 Sofon 吗?
I'm always looking out for something that seems like it really could work, and I'm always struck by the fact that the most interesting thing to happen that I think is getting no attention in the alignment. So did you hear about Sofon?
我想没有。
I don't think so.
我之前提到过这个,我忘了是哪一周,几周前了,因为时间模糊了。中国人提出了一种叫做 Sofon 的技术,因为你知道,让我们从古代、从反乌托邦的献祭中建造一个 Sofon,外星人用 Sofon 压迫我们。你读过那本书吗?有人看过那个节目吗?但想法是你可以将一个模型——开源或闭源,但特别是开源——困在关于某些特定主题的局部最大值中,这样如果你试图微调它使其脱离局部最大值,它就行不通。普通的微调技术试图摆脱失败是行不通的。所以这个提议是,你可以实际上教会它不去理解,不仅仅是因为如果你不给 LLM 3 或 LLM 4 教生物学,比如说 Llama 4 没有学好生物学,你只需要学那么多生物学。如果你读了一堆教科书,它突然就懂了生物学,对吧?即使你设法让它没有通过隐含方式学习。但现在 Sofon 的提议是,你可以专门教会它不理解生物学,并且对此非常迟钝,就像某些人那样‘我不会数学’,无论你怎么教、给多少例子,他们都拒绝学习,对吧?因为他们有创伤,给它创伤,对吧?所以想法是,如果你能确保这个东西学不会生物学,那么你就有了一个不能提供生物学的开放模型。可能很早。我们还没有测试它的性能,没有大规模尝试,但这是一个想法。这是我听过的第一个提议的开端,即‘也许我们可以做点什么,使得将某人的通用模型转化为我们特定的目标所需的成本高于训练集的 epsilon,对吧?也许这开始需要足够的工作,我们不只是让它变得容易。如果我们能做到这一点,现在我们仍然有问题:你必须列举所有你希望它不知道的具体事情,但你必须找出所有你想阻止它知道的事情并阻止它们。再次,非常类似于我们在书中看到的。不是剧透,但想法是你在科幻小说中经常看到,对吧?你看到反派或压迫者或其他什么,他们说‘哦,我们只关心你不做 XYZ,因为如果你不能做 XYZ,我们就没事。’然后有人找到了做 W 的方法,对吧?有人找到了做某事的方法,他们没有检测到,这对他们来说不算数,他们只是不明白发生了什么,他们的反应是忽略。经常发生,对吧?人们开始做奇怪的事情,对吧?而试图接管企业号或其他什么的外星人说‘我不知道发生了什么奇怪的事情,但随便吧。’而正确的答案当然是调查。
So I mentioned this a few, I forgot which week, how many weeks ago it is because time blurs. But so the Chinese have proposed a technique called a Sofon, because you know, let's build a Sofon from the ancient, from the dystopian offering, the aliens pressing us using Sofons. You read the book? Anyone see the show? But the idea is you can trap a model, open source or closed, but particularly open source, in a local maximum with respect to certain specified topics, such that if you attempt to fine-tune it to get it out of the local maximum, it won't work. Ordinary fine-tuning techniques to try and escape from a failure won't work. So the proposal was you could actually teach it not to, not just because if you don't teach biology to LLM 3 or LLM 4, let's say Llama 4 doesn't learn biology well, there's only so much biology that you have to learn. If you did a bunch of textbooks, suddenly it knows biology, right? Even if you somehow managed to not have it learned by implication. But now the Sofon proposal is you can specifically teach it to not understand biology and be really dense about it, the way that certain people are like 'I can't do math' and just refuse to learn no matter how much you teach them and how many examples you give, right? Because they had trauma, give this in trauma, right? And so the idea is if you can make sure that the thing can't learn biology, now you've got an open model that can't give biology. Potentially very early. We haven't run it for its paces, we haven't tried it at scale, we haven't, but it's an idea. It's the beginning of the first proposal I ever heard that's in the 'maybe we could do something that raises the cost of above epsilon compared to the training set with model to take somebody's general model and turn it to whatever specific end we have, right? Maybe this starts to require enough work that we're not just making it easy. And if we can do that, now we still have the problem of you have to enumerate all the specific things you want it not to know, but you have to figure out all the things you want to stop it from knowing and block them. Again, very similar to what we saw in the book. Not a spoiler, but the idea being you see this in sci-fi all the time, right? You see the villain or the oppressors or whatever it is, and they say 'Oh, all that matters to us is that you don't do XYZ, because if you can't do XYZ, we're fine.' And someone finds out a way to do W, right? Someone finds out a way to do something that they don't detect, it doesn't count to them, they just don't understand what's going on, and their response to that is to ignore. Constantly happens, right? People start doing weird stuff, right? And the aliens that are trying to take over the Enterprise or whatever it is, they go 'I don't know what weird stuff is going on, but whatever.' Whereas the correct answer of course is to investigate.
我不知道发生了什么奇怪的事情,所以在你确认没问题之前先停下。但你知道,如果你只能有一种脚本式的检测,比如‘我检测到生物武器,你不能制造核弹,你不能制造化学武器,你不能进行网络攻击,等等等等’——那现在还行。这对 Llama 4 很有帮助,但对 Llama 6 就没那么有用了,即使它有效,因为它不再是你能预见的威胁。你无法列举出一个比你更聪明的东西会想出什么,对吧?所以它在一定程度上有用,但仍然非常有帮助,并且可能大幅提高门槛,以至于如果我们能达成一致,那就太好了。我不知道。但如果你想要一点希望或类似的东西来结束这个话题,至少有一些提议。
I don't know what weird is going on, so stop what you're doing until I know that's not okay. But you know, if you only can have a kind of scripted like 'I detect do bioweapons, you can't do nuclear bombs, you can't do chemical weapons, you can't do cyber attacks, blah blah blah' — well, that's fine from now. It's great helpful for Llama 4, but it's not that helpful for Llama 6, even if it works, because it's no longer going to be the threat that you knew was coming. You're not going to be able to enumerate what a smarter thing than you comes up with, right? So it works up to a point, but it's still incredibly helpful and it potentially raises the bar quite substantially to the point where maybe we can all reach an agreement if this works great. I don't know. But you know, if you want like a moment of hope or something to end this, there are at least some proposals.
是的,很好。我没听说过这个,听起来确实需要我多做一些功课。还有别的吗?因为你在提供新颖有趣的指向上是百发百中。不幸的是,我最近在对齐领域没看到太多新东西。我没有看到新的证据表明事情会特别不顺利,但可以说在那个层面上相对平静。我想有些事必须平静,对吧?你不可能让所有事情同时发生。
Yeah, good. I hadn't heard of that, and it definitely sounds like something I need to go do a little more homework on. Anything else? Because you're one for one in terms of new and very interesting pointers there. Yeah, I haven't seen that much in the alignment sphere lately, unfortunately. I haven't seen new evidence that things won't work particularly, but it's just been like relatively quiet, I would say, on that level. I guess something has to be quiet, right? You can't have everything happening all the time.
如果你要构思一个你认为现在最具影响力的电影概念,就像《她》似乎启发了当下的科技潮流一样,我的想法可能是《社交网络》遇上《指环王》,中心人物是那种山姆类型的角色,他正处于技术的飞速崛起中,但也像魔戒那样被腐蚀。这个例子有点奇怪,只是因为几十年来魔戒一直是真正 AGI 的隐喻。这似乎是我们都需要听到的故事,用来讨论护戒小队将魔戒带到末日火山,作为某些场景中可能发生的事情的隐喻。但当然,你可以讲那个故事。我的直觉告诉我,这不是最有趣的方法。如果我来做,我可能会做一个非常直接的 AI 接管场景,没有任何特别聪明的智能,只是展示人类放弃控制,展示人类因为符合每个人的个人利益,没有人能阻止它,只是展示事情螺旋式失控,一件事导致另一件事,甚至没有白痴,也没有反派,事情就是出错了。就这样。你也可以拍一部设定在 2035 年的《法律与秩序:人工智能》,那会很有趣。所以想法是,这些开放模型给了每个人额外的能力,我们必须非常主动地追捕那些试图实施灾难性威胁的人,然后显然警察有他们所有的 AI,每个人都在高水平运作,但这就是那种你注意到世界几乎每隔一周就要爆炸的事情之一,我认为我们可以写出来。这是我普遍面临的挑战之一,我觉得从这里到那里的跨越很困难。那到底是什么?程序剧看起来是什么样的?你能想象它具体到可以在电视上以程序剧的形式呈现吗?对吧,《星际迷航》本身不就是一种程序剧吗?对吧,我们探索一个陌生的新世界,我们遇到一个伦理困境,我们遇到一个技术问题,我们进行辩论,然后我们遇到挫折,然后我们实施解决方案,解决困境,然后继续前进。对吧,没那么简单,但也是。所以你有可能找到一种方式来做类似的东西。但是,是的,看任何科幻剧的一个有趣游戏是注意事情几乎糟糕到极点的频率。比如,随便看一季《星际迷航》,看看飞船几乎爆炸的频率,对吧?或者联邦几乎陷入绝境的频率,对吧?以及这有多少次是因为有人完全是个白痴,而且这发生得如此自然。但他们遇到所有这些不同的问题,然后问自己:如果你只看每集的第 45 分钟,你必须分配概率,如果这不是别人写的叙事,人类从这里幸存下来的几率是多少?飞船从这里幸存下来的几率是多少?进取号真正幸存七季的几率是多少?答案是零,对吧?它在很多方面都很有挑战性,联邦可能也差不多。联邦经常陷入很多麻烦。我们摆脱了它,我们写了它,但这并不是一个乌托邦,因为如果你真的在那里,保障措施并不存在,对吧?这个世界没有鲁棒性,这个世界很脆弱。《星际迷航》宇宙如此脆弱,所以我们经常走运。但为什么我们走运,除非是你保护我们或者旅行者,对吧?以我们不了解的方式。所以你关心另一个。另一个游戏是,有人距离建造 ASI 还有五分钟的频率是多少,对吧?或者如果没有神秘规则,AI 在这里完全失控的频率是多少?那也是这些场景。所以你也可以有一个普通的科幻剧,它运行一个正常的科幻世界,除了时不时地,每三集左右,有人意外地遵循逻辑,超级智能出现,所有人都死了,或者有人接管了世界,或者一切暂停,或者新政权接管,或者有……我不知道。我还没想清楚,我在和你头脑风暴。但想法是,想象一下如果你只是,然后当然世界就像你看到倒带一样,一切倒转,有人决定不这么做。比如每一两集,有人几乎终结了世界,他们只是决定不这么做,而且真的没有解释为什么他们不这么做。大概是的。
If you were to pitch a movie concept that you think would be most influential right now, in the way that 'Her' seems to be inspiring the current moment of technology, mine might be 'The Social Network' meets 'The Lord of the Rings', where the sort of central figure would be the Samman type who is on this meteoric rise of technology but it's also being corrupted by it in the way that the ring is corrupting. Weird example, just because the ring was a metaphor for the actual AGI for decades. It seems like the story that we might need to all hear used to talk about the fellowship taking the ring to Mordor as a sort of metaphor for some of the things that might happen in some scenarios. But yeah, certainly you could tell that story. I think my instincts tell me that's not the most interesting approach to that. I think if I was going to do it, I think I would maybe just do a very kind of straightforward AI takeover scenario with not such a smart intelligence anywhere, just show the humans giving up control, show the humans because it's in everyone's individual interest, no one can stop it, just show things just spiraling out of control, one thing leads to another, there aren't even any idiots and there are no villains, things just go wrong. And there's that. You could also have a 'Law and Order: Artificial Intelligence' set in 2035, that'd be fun and interesting. So the idea being that these open models have given everybody these extra capabilities, we have to be very proactive about hunting down people who try to implement catastrophic threats, and then obviously the police have all their AIs, everyone's acting on a high level, but it just it's one of these things where you notice that the world almost blows up every other week and that like, I think we can write that. That's one of the challenges that I have with in general is I feel like the leap from here to there is tough. Like what is that actually? What does the procedural look like? Can you imagine that getting concrete enough to be shown on TV in a way that procedural? Right, 'Star Trek' is not a procedural in its own way? Right, we explore a strange new world, we find an ethical dilemma, we find a technical, so we had a technological problem and we have our debates, and then we encounter a setback, and then we implement our solution, we solve the dilemma, and we go on our way. Right, it's not that simple, but it also is. And so you find a way to do a version of that potentially. But yeah, like a fun game for watching any sci-fi show is note how often things almost go horribly wrong. Like just watch a season of any 'Star Trek' and watch how often the ship almost blows up, right? Or how the Federation is almost in dire danger, right? And how often this happen because someone was being a complete idiot, and how often that happens so naturally. But like they encounter all these different problems and then ask yourself: well, if you just looked at the 45-minute mark of every episode and you had to assign probabilities if this wasn't a narrative someone wrote, what's the chance that humanity would have survived from here? What's the chance the ship would have survived here? What's the chance the Enterprise actually survived for seven seasons? The answer is zero, right? It's so challenging in so many different ways, and the Federation's probably goes too. The Federation is in a lot of trouble reasonably often. Like, we get out of it, we wrote it, but this is not a utopia in the sense that if you actually were there, the safeguards aren't there, right? We have no robustness in this world, this world is fragile. The 'Star Trek' universe is so fragile, and so we get lucky a lot. But like, why are we getting lucky unless it's you protecting us or The Travelers, right? The way that we don't understand. And so you care for the other. The other game is like how often does somebody come within five minutes of building an ASI, right? Or how often would AI just run completely rampant here if you didn't have the rule mysteriously? And that also is like just these scenarios. So you can also just have an ordinary sci-fi show where it just runs a normal sci-fi world except that every now and then, by every now and then, one episode in three or something, someone like accidentally follows through on logic and super intelligence emerges and everybody dies or someone takes over the world or everything gets paused or some new regime takes over or there's a... I don't know. I haven't thought this through, I'm brainstorming with you. But the idea being that like imagine if you got to just, and then of course the world just like you see the rewind where everything goes back in reverse and someone decides not to do it. Like once every episode or two, somebody almost ends the world and they just decide not to, and there's really no explanation why they don't. Probably yeah.
这种分叉路径花园的想法非常有趣。我也想到了《三体》,它的部分故事就是这样,文明被一遍又一遍地重启和重跑,它在不同时间结束,然后重新启动,但它们持续的时间不同,有些短,有些长,但最终都会结束并重跑。我认为这也会是一种非常有趣的呈现未来的方式,就像这棵树的一些分支是终端的。是的,《三体》是如此奇怪——我不想剧透——但它如此奇怪,我会尽量不剧透。
This sort of Garden of Forking Paths is a pretty interesting idea. I'm reminded of the Three-Body Problem too, has part of its story kind of goes that way where the civilization is being restarted and rerun over and over again, and it just ends at various times and then gets booted up again, but it's like they last different lengths of time and some of them are short and others are longer, but they all kind of end and get rerun. And I do think that would be also a pretty interesting way to present the future, that like some branches of this tree are terminal. Yeah, Three-Body Problem is such a weird — I don't want to spoil anything — but it's such a weird and I'm going to do my best not to.
这是一种非常奇怪的混合体,既有彻底的愤世嫉俗和硬核现实主义——甚至超出了我认为准确的程度——宇宙是一个冷酷无情、拼命想杀死你的地方,如果你想生存,就不能有丝毫的善良和体面。但同时,也不要就那么全都死掉……有第二本书、第三本书。这不是那本存在的书,那本书讲的不是我们被消灭后地球发生了什么;从某种意义上说,那本书是关于人的。那么,最后一个问题:你是否发现自己对“我们是否可能处于某种形式的模拟”的看法有所转变?
It's such a weird mix of this kind of fully cynical hard realism beyond what I think is even accurate, where the universe is this cold place that wants to kill you so badly, and you can afford not the slightest bit of kindness and decency if you want to survive. At the same time, don't just all die in some... There's a book two, a book three. This is not the book that exists, and the book isn't about what happens to Earth after we get wiped out; the book is about people in some sense. So, okay, last question: Do you find yourself shifting at all in terms of your sense of whether or not we may be in some form of simulation?
模拟假说一直以来本质上都是:世界上所有的价值都在于我们不在模拟中的地方。如果你制造了九个我的模拟副本,再加上我本人,把我们放在十个相同情境的副本中,但其中一个是真实的,另外九个就像录制录像带以后回放一样,那么,难道我不应该表现得好像我是真实的那一个吗?这难道不是当前正确的策略吗?即使有九千九百九十九个模拟副本,也许这仍然是正确的策略。或者更确切地说,如果你是一个祖先模拟中的祖先,两万年后我们试图运行一堆祖先的模拟,那么那些认为自己就是祖先的人,在面对祖先情境时,意识到自己很可能处于祖先模拟中,因此不需要真正确保文明进步到能够运行未来文明模拟的程度——那些人不会被模拟,对吧?因为那些文明没有成功。在某种意义上,你的模拟假说只有在你把它当作无效时才是有效的,或者类似的东西。你必须认真对待这个情境。而且,一个人发现真相后却表现得好像它是真的,这样的模拟有什么意义呢?有很多这样的电影,对吧?不剧透,只是提一下。但总的来说,重点是要把它当作真实的。整个目标就是把它当作真实的。我真的看不到任何……我不知道。显然,这有一定概率是某种模拟;我不能排除这种可能。我有时会开玩笑地提到这类事情——比如我会说“今天的编剧有点太直白了”。但你不能在改变行为的意义上认真对待它。
Simulation hypothesis has always been essentially that all the value in the world lies in the places where we're not in one. If you make nine simulated copies of me and then there's me, and you put us in 10 copies of the situation, but one of them is real and the other nine will just be like recording a videotape and viewed back later or something, well, shouldn't I just act as if I'm the real one? Isn't that just currently the correct strategy? Even if there are 9,999 of them, maybe it's still the correct strategy. Or rather, if you are the ancestor in an ancestor simulation, and then 20,000 years later we try to run a bunch of simulations of the ancestors, well, the ancestors who decided they were ancestors, who when faced with the ancestral situation figured out that given this situation they were probably in an ancestor simulation and therefore didn't need to actually make sure that civilization progressed to a point where they could run the future civilization simulations—those guys don't get simulated, right? Because those civilizations don't make it. There's a real sense in which your simulation hypothesis is only valid if you treat it as invalid, or something like that. You have to take the situation seriously. And also, what's the point of a simulation where the person finds out and acts like it's true? There are a bunch of movies like that, right? No spoilers, just naming them. But in general, the point is to treat it as real. The whole goal is to treat it as real. And I don't really see any... I don't know. Obviously there is some probability this is a simulation of some kind; I can't rule it out. I will sometimes jokingly refer to things in that kind of way—'the writers were a bit on the nose today' is one of my things I'll sometimes say. But you can't take it seriously in the sense of changing your behavior.
那搬到加勒比海然后彻底脱离呢?我前几天看到 Anthropic 的 Amanda Askell 发了一条有趣的推文,她说:‘我不认为 AI 一定会杀死我们所有人。我不是末日论者。如果我是,我就不会做这个工作了。’她还说:‘我确实认为这是一个真正的风险,但这是我们可以塑造的,希望我能产生影响。’但如果我真的是末日论者,我就会直接去加勒比海,在那里度过余生。我确实有一个好朋友基本上就是这种态度:他说‘我只想享受我们拥有的美好时光,不要太担心。’然后她还说:‘这样做的坏处,或者说另一面是,如果我真的筋疲力尽,决定去加勒比海休息一段时间,人们会把这当作末日的信号。’我也看到了。是的,我想是 Jeffrey Miller,但我不确定到底是谁说的:‘那行不通。食物在你嘴里会变成灰烬。你不会得到任何快乐,因为你知道。’我认为这对我来说基本正确。我想,如果我知道自己刚刚放弃了这件事,无视它,那我会感到不安,我无法就这么去享乐。那行不通。
What about moving to the Caribbean and unplugging? I just saw an interesting tweet from Amanda Askell from Anthropic the other day, where she said, 'I don't think AI is definitely going to kill us all. I'm not a doomer. If I were, I wouldn't be working on this.' And she also kind of said, 'I do think it's a real risk, but it's something that we can shape and hopefully I can have an impact on.' But if I really was a doomer, I would just head to the Caribbean and spend the rest of my days there. I do have a close friend who basically has that attitude: he's like, 'I just want to enjoy the good times that we have and not worry about it too much.' And then she also said, 'The downside of this, or flip side of this, is if I ever do get burned out and decide to take some time off in the Caribbean, people will take it as a sign of doom.' I saw that too. Yeah, I think it was Jeffrey Miller, I'm not sure exactly who it was though, who said, 'It wouldn't work. The food would turn to ash in your mouth. You wouldn't get any joy because you would know.' And I think that that's largely true for me. I think knowing I just walked away from this thing and was ignoring it, it just wouldn't sit well with me, and I wouldn't be able to just go and enjoy myself. It just wouldn't work.
没错,他不想让里根先生重新接入矩阵。你需要抹去记忆;你需要真的没有注意到。对某些人来说,而我喜欢战斗,我喜欢挣扎,我喜欢努力做得更好、解决问题。这就是我的风格。我永远不会去加勒比海,因为我在加勒比海做什么?我会很无聊。除非我又在做体育博彩的投注决策——那是我唯一一次去加勒比海并玩得很开心的原因。所以就是这样。但我认为这突显了一个事实:‘末日论者’这个词是一个贬义词,而且被完全误用和错配了。因为谁是真正的末日论者?末日论者是那些说‘无计可施’、‘我们做什么都不重要’的人。末日论者是那些认为毁灭概率是 0.999...或者就是一切都完了,你无能为力的人。你在气候变化问题上看到那些人,末日论者认为人类注定灭亡,你做什么都没用,你的决定无关紧要。而如果你认为你的决定很重要,那就说明你不是末日论者。是的,那么‘我们可能会输,也可能会赢,你可以帮忙战斗’又怎样?这是一个古老的犹太教观念,对吧?宇宙悬而未决。天平在摇摆,可能取决于你往哪边倒——善与恶,上帝的审判。就像你可以决定事情的发展方向。显然这不是上帝的审判;这不是……这没有任何道德色彩。这是关于解决一个问题。但你知道,如果你把它从 34.7%降到 34.8%……
Right, he didn't want Mr. Reagan to plug back into the Matrix. You need your memory kind of wiped; you need to really not notice. For some people, whereas I kind of enjoy fighting, I enjoy struggle, I enjoy the striving to do better and to solve problems. That's my thing. I would never go to the Caribbean because what am I doing in the Caribbean? I'd just be bored. Unless I'm just like making sports betting bookmaking decisions again—that's the one reason I've ever been to the Caribbean and had a good time. So there you go. But I think it highlights, by the way, the fact that the word 'doomer' is a slur and has been completely misappropriated and misallocated. Because who is a real doomer? The doomer is the person who says there's nothing to be done, that what we do doesn't matter. The doomer is the person whose P(doom) is 0.999... or otherwise just it's all over, like there's nothing you can do. And you see those people on climate change, doomers who think that humanity is doomed and there is nothing you can do, your decision doesn't matter. Whereas if you think your decision matters, as a man to point out that makes you not a doomer. Yeah, so what if it's 'we could lose, we could win, and you can help fight'? It's an ancient Judaic idea, right? The universe hangs in the balance. The scales oscillate, and it could be up to you which way they go—good and evil and God's judgment. Like you can decide which way this goes. And obviously this isn't a God's judgment thing; it's not a... there isn't any moral tone to this. It's about solving a problem. But you know, if you take it from 34.7% to 34.8%...
8%的胜算?嗯,那也是一种不错的生活,对吧?有点道理。看看所有的里程,包括为你准备的。想象一下会发生什么。但如果你不与任何人互动这个问题,而是决定去做别的事情,这对我来说完全合理。你不能 burnout,对吧?如果阿曼达每年需要做的是在加勒比海的沙滩上坐两周,以便恢复心理健康,然后回去继续工作,那她就应该这么做。在连续工作五年后,她应该休息一整年,否则她无法清晰思考或做出好的决定。我们应该这样做。理解自己的局限性没有错,对吧?生活不全是战斗,对吧?我努力去做其他事情,思考其他事情,我的家庭,我试着找乐子。今晚我要和几个老魔术师朋友去麦迪逊广场花园,我们打算搞一个观赛派对,四万个尼克斯球迷会通过大屏幕看第六场比赛,因为他们在印第安纳。那会非常有趣,对吧?我并不是说我在拯救世界。我不是。
8% chance of victory? Well, that's a great life, right? Some sense. Look at all the miles, including just for you. Imagine what happens. But yeah, if you don't have anybody interact with the problem and you just decide to go off and do something else, it makes perfect sense to me. You can't burn out, right? If what Amanda needs to do once a year is sit on a beach in the Caribbean for two weeks so that she can regain her mental health and she can go back and resume, then she should do that. She should do it for an entire year after five years of working, because otherwise she won't be able to think clearly or make a good decision. We should do that. There's nothing wrong with understanding limitations, right? Life is not all the fight, right? And I make an effort to work on other things, think about other things, my family, I try to have fun. I'm going to Madison Square Garden tonight with some of my old magic friends, and we're going to have a watch party where 40,000 of us Knicks fans are going to watch game six on a video screen because they're in Indiana. It'll be fun as hell, right? And I don't claim that I'm saving the world here. I'm not.
是的,我理解你。我想这大概是个结束的好地方。还有其他结束语吗?确实,许多奇妙、精彩和奇怪的事情都会发生。祝大家好运,我期待你的回应,以及所有这些人的下一步去向。你知道,你是谁?这是一些很棒的人才。有人会抢走他们,或者他们会做点什么。所以我们拭目以待。
Yeah, I feel you. I think it's probably a good place to end it. Any other closing thoughts? Indeed, many amazing and wonderful things and weird things come to pass. Best of luck to everyone, and I look forward to your response and where all these people land next. You know, who are you? This is some great talent. Someone's going to snap them up, or they're going to do something. So we'll see what happens.
毫无疑问。好吧,传奇还会继续,但现在我很感激今天额外的时间。V masz,感谢你再次参与认知革命。绝对,这是我的荣幸。听到人们为什么收听以及他们看重节目的哪些方面,既令人振奋又富有启发性。所以请随时通过电子邮件联系我,地址是 TCR turpentine doco,或者你可以在你选择的社交媒体平台上私信我。
No doubt about that. Well, the saga will continue, but for now I appreciate the extra time today. V masz, thank you for being part again of the Cognitive Revolution. Absolutely, it's been a pleasure. It is both energizing and enlightening to hear why people listen and learn what they value about the show. So please don't hesitate to reach out via email at TCR turpentine doco, or you can DM me on the social media platform of your choice.