Sam Altman on AI Safety and the Road to AGI
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 首席执行官山姆·奥特曼探讨 AI 安全、AGI 时代,以及为何 10%的灾难风险不可接受。
OpenAI CEO Sam Altman discusses AI safety, the AGI era, and why a 10% risk of catastrophe is unacceptable.
现在上市是不明智的。这听起来不像 2026 年、2027 年的事。
Right now would be an ill-advised moment to go public. This sounds like not 2026, 2027.
我会说不是 2026 年。是的,我们有很多事情要做。
I would say not 2026. Yeah, we got a lot of stuff to do.
我们在 OpenAI 的旧金山总部,与 Sam Altman 对话。刚刚过去的一周非常忙碌,很多人对 AI 的安全以及文明崩溃的可能性深感担忧。我认为,在本十年末之前,承担 10% 的杀死所有人的风险是不可接受的。显然,我们将与 Sam 讨论他对此的真实想法、技术今天的状况,以及需要做什么才能创造安全的 AI,造福全人类。我是 Allison Chantel,这里是《财富》500 强巨头与行业颠覆者。
We're here in the San Francisco headquarters of OpenAI to speak with Sam Altman after a really busy week that has a lot of people very concerned about the safety of AI and the possibility of civilization collapse. I think it is unacceptable to be taking like a 10% chance of killing everybody by the end of the decade. Obviously, we're going to talk with Sam about his actual thoughts on that, where the technology is today, and what needs to happen to create safe AI that can benefit all of humanity. I'm Allison Chantel, and this is Fortune 500 Titans and disruptors of industry.
Sam Altman,非常感谢你在经历了一周疯狂的新闻后,今天抽出时间与我交流。
Sam Altman, thank you so much for spending time with me today after a wild week of news.
现在每周都很疯狂。我猜也是。
They're all wild now. I bet so.
你今天对 AI 的未来感觉如何?
How are you feeling about the AI future today?
我觉得我已经有好几年、几乎十年的时间来为此做准备,我们也有时间集体去理解它。这显然是很多人正在努力应对我们面前问题的一周。所以我其实主要很感激这种情况正在发生。我认为这是世界早该进行的对话,我们显然正在进入一个拥有极其强大模型的领域,风险很高,需要把事情做对。
I feel like I've had like years slash almost a decade to get ready for this, and kind of we've had time to collectively wrap our heads around it. This has clearly been a week where a lot more people are grappling with the issues in front of us. So I'm actually mostly grateful that that's happening. I think that this is a conversation that the world is overdue for, and we are clearly entering a realm of extremely capable models, and a high stakes and a need to a need to get things right.
所以你在 OpenAI 的使命,我会稍微谈谈本周发生的事情。感觉这是一个关键时刻。但 OpenAI 的指导使命是确保 AGI 造福全人类,对于不知道 AGI 是什么的人,它指的是通用人工智能。据我理解,工作定义是 AI 达到在大多数有经济价值的工作上超越人类的地步。你同意吗?
So your mission at OpenAI, and I'll talk about a little bit about what happened this week. That feels like this is a pivotal moment. But the guiding mission of OpenAI is to ensure AGI benefits all of humanity, and for people who don't know what AGI is, it means artificial general intelligence. And the working definition, to my understanding, is that AI gets to the point where it outperforms humans at most economically viable work. You agree?
是的,不过我要说,很多人对 AGI 有非常不同的定义,我认为重点只是极其强大的模型,而过于纠结于“嗯,它到底是不是这个或那个?”如果你太专注于学究式的“我们到了吗?我们没到吗?还缺什么?”,你可能会失去对正在发生的事情的规模的感知。就像我们今天显然处于曲线上的一个点,既有巨大潜力,也有真实风险。
Yes, although I will say that many people have very different different definitions of AGI, and I think the important point is just like extremely capable models, and getting kind of too hung up on the like, well, is it exactly this or exactly that? Is is a you can kind of lose the magnitude of what's happening if there's too much focus on the pedantic like, are we there? Are we not there? What's missing? Like we are clearly today in a point on the curve with great potential and real risk.
而你的最新模型 Astra 非常强大,以至于你的联合创始人 Greg Brockman 宣称:“欢迎来到 AGI 时代。”英伟达的黄仁勋也称其为 AGI。所以我们处于这个有趣的时刻,但确保它造福全人类绝对不是理所当然的,本周我们在网上感受到了这一点,一位前 OpenAI 员工 Jacob Coxon 走红,他说他辞去了在 Anthropic 的工作,因为构建 AI 的人真诚地相信它可能在本十年末杀死我们所有人。然后他的另一位 Anthropic 朋友插话说:“他完全正确。我们确实真诚地相信这一点,而且到本十年末杀死我们所有人的可能性可能超过 10%。”他们说得对吗?
And your latest model, Astra, very powerful to the point where your co-founder Greg Brockman declared, "Welcome to the AGI era." Jensen Huang of Nvidia also called it AGI. So we're at this interesting moment, but ensuring that it benefits all of humanity is definitely not a given, and we got a sense of that online this week when a former OpenAI employee, Jacob Coxon, went viral for saying he was quitting his job at Anthropic because the people building AI earnestly believe it could kill us all by the end of the decade. And then one of his other Anthropic buddies chimed in and said, "He's totally right. We do earnestly believe that, and it could be by greater than 10% chance that it kills us all by the end of the decade. Are they right?
我认为,在本十年末之前承担 10% 的杀死所有人的风险是不可接受的。显然,我认为从事这些努力的人拥有巨大的权力和能力来影响这里发生的事情。我们一直在采取一系列行动。我们达到这个新的能力水平,以确保我们不会代表人类承担那样的风险,我认为其他人也不应该。我们以前也有过这样的时刻。再次强调,这是更强大的模型,风险比以往任何时候都高。我们以前也有过这样的时刻,我们不得不说:“好吧,模型已经达到了新的能力水平。有一系列新的风险。我们需要增加对齐工作、监控工作、总体安全栈,以减轻我们面前的风险。我们需要新政策。我们需要实验室之间的新协调,这是又一个我们需要更多此类措施的時刻。显然有巨大的好处,但我不。我认为人们已经谈论这个很久了,你知道,我也认为我们所说的是真的,即我们将能够使用这些越来越强大的模型来确保我们能够完成我们需要做的安全工作和安全研究。还有你知道我们可以谈论的其他一切,比如我们需要在对手之前到达那里,但我实际上不认为此刻这是正确的。我认为此刻我们正在进入一个新时代,我们像过去踏入其他新时代一样,需要以不同的方式行事,以确保我们做的事情将确保 AGI 造福全人类,但也确保没有人承担接近那样的风险。我认为同时可以成立的是,如果世界和构建这项技术的公司不改变他们过去的做法,可能会有重大风险,但如果不调整我们所有人的工作方式,那就太疯狂了。以及我们做决策的方式,以及我们的政府理解和围绕这项技术设置护栏的方式,鉴于此,如果趋势继续像现在这样,你认为 10% 这个数字准确吗?
I think it is unacceptable to be taking like a 10% chance of killing everybody by the end of the decade. Obviously, I think the people that are working at these efforts have a huge amount of power and ability to impact what happens here. We have been taking a number of actions. We reach this new capability level to ensure that we do not take a risk on behalf of humanity like that, and I don't think anybody else should either. We have had moments before. Again, this is more capable model, higher stakes than ever before. We've had moments before where we've had to say, "Okay, the models have reached a new level of capability. There's a new set of risks. We need to increase our alignment work, our monitoring work, our safety stack in general to be able to mitigate the risks in front of us. We need new policy. We need new coordination between labs, and this is another moment where we need more of that. There's obviously tremendous benefit, but I don't. I think people have been talking about that for a long time, and you know, I also think it is true what we say, which is we will be able to use these increasingly powerful models to ensure that we can do the safety work and the safety research that we need. And there's you know everything else that we could talk about about you know we need to get there before adversaries do, but I actually don't think that's right in this moment. I think in this moment we are entering a a new era, and we like we have as we've stepped into other new eras in the past, act differently to ensure that we are doing something that will be ensure that AGI benefits all of humanity, but also ensure that no one is taking risk anywhere close to that. I think it can simultaneously be true that if the world and the companies building this technology did not do things any differently, they've been doing them in the past. There might be significant risk, but it would be insane not to adjust the way we all work. And the way that we make decisions, and the way that our governments understand and put guardrails around this technology, in light of that, if trends continued the way they were now, do you think that 10% figure is accurate?
如果我们不做出改变,我仍然不知道。但我也不知道人们怎么能给出一个数字,你知道,是 5 或 10 或 15 或 30 或 50 或其他。就像,我不知道你怎么能给出这样的数字。Elon 已经说是 10% 有一段时间了。他们称之为 P doom。就像构建超级智能带来的毁灭概率。再次,无论是 10 还是 8 还是 6,重点是我們都有巨大的责任,不能让自负或利润激励或其他任何事情妨碍。我们需要采取行动,确保我们不承担任何这些数字的风险。我相信我们可以。我相信这些公司,当然是我们自己的公司,我想我最有发言权,将会并且一直在迎接这个新时刻,所以你们会改变。它已经发生了。
If we don't make a change, I still don't. But I also don't know how people can put a you know it's five or 10 or 15 or 30 or 50 or whatever. Like, I don't know how you can put a number like that. Elon's been 10% for a while now. They call this P doom. It's like the probability of doom from building super intelligence. Again, whether like whether it's 10 or eight or six, the point is like we all have a tremendous amount of responsibility and cannot let egos or incentives for profit or anything else get in the way. We need to act such that we are not taking any of those numbers of risk. And I believe we can. I believe that the companies, certainly our own company, which I guess I can speak for the most, are going to and have been meeting this new moment, so you're going to change. It's happened.
所以你的团队本周发布了几篇帖子,其中一篇来自你们安全与安保委员会的新非营利董事会成员,他基本上说:“如果我们构建超级智能而没有更强大的对齐,我预计我们将永久失去对它的控制,如果发生这种情况,那么大多数人可能会死亡。”这是你们的新非营利董事会成员,现在在你们的安全与安保委员会上,说如果事情不改变,那可能会成问题。
So a few posts have been put out by your team this week, and one of them was from a new nonprofit board member on your safety and security council, who basically said, "If we build superintelligence without more robust alignment, I expect we will permanently lose control of it, and if that happens, then most people could die." This is your new nonprofit board member, who's now on your safety and security council, saying that if things don't change, that could be problematic.
没人知道对齐是什么意思。是的。没人知道控制是什么意思。甚至 Paul 是在什么时刻说这个的?我认为我们都应该同意的一个关键原则是,我们不能采取可能冒险将未来控制权交给 AI 的行动。我不记得你引用的 Paul 的原话,但我认为大意如此。
No one knows what alignment means. Yeah. No one knows what like control means. Even what is this moment that Paul is saying that about? I think a a key principle that we should all agree on is that we cannot take actions that would risk losing control of the future to AI. I I don't remember Paul's exact words as you quoted them, but I think it was something along those lines.
失控事件是我能预见的少数几种真正出大问题的方式之一。随着这些系统变得如此强大,我们所需的监控和对齐技术——我认为我们目前还不能在能力上推进太多,除非在可监控性和对齐方面取得更多进展:即理解模型在做什么的能力,以及确保模型遵循人类价值观和用户意图的能力。
A loss of control incident is one of the small number of ways I can see this going really wrong. As these systems get so capable, the monitoring and alignment techniques that we need to ensure we don't lose control—I don't think we're currently at a place where we could push much further on capabilities without making more progress on monitorability and alignment: the ability to understand what a model is doing, and the ability to make sure that a model will follow human values and the intent of its users.
现在,我认为我们最近采取了一些行动。实际上,把 Paul 加入董事会就是其中之一。Jakob 上周的帖子,我认为是另一个很好的例子。但最重要的是,我们一直在暂停训练运行,直到我们能提出一个更令人满意的安全案例,考虑到这些模型的发展方向。
Now, I think that we have taken a number of actions recently. Actually, adding Paul to the board is one of them. Jakob's post from last week, I think, was another great one. But most importantly, we have been pausing training runs until we can make a safety case that we're more comfortable with, given where these models are going.
我怀疑我们将能够像过去每次一样,做更多的研究,使我们能够安心地继续,并且我们可以说,鉴于我们在监控这些模型的能力、对齐这些模型的能力方面取得的进展,我们愿意进一步推进能力。如果我们不能,我们就不应该进一步推进能力。
I suspect that we will be able, as we have every time in the past, to do more research that allows us to be comfortable continuing, and we can say that, given the progress we've made on the ability to monitor these models, the ability to align these models, we are comfortable pushing the capability further. If we can't, we should not push the capability further.
再说一次,任何人都不应该——我认为,冒险训练一个你无法说“是的,这是我们所做的,这是我们的安全案例,这就是我们知道它没问题”的模型,是疯狂的。我们会在运行期间和部署模型之前进行审计,并再次审视。但能力和对齐监控安全必须同步推进,让能力超前于 Paul 所谈论的内容,我认为在某个时候会带来真正意义上的失控风险,而这绝不应该发生。
Again, no one should be taking—I think it'd be insane to be taking the risk of training a model where you can't say yes. Here's what we've done, and here's our safety case, and this is how we know it's okay. And we'll have auditing during the run and after before we deploy a model, and we'll look at it again. But capabilities and alignment monitoring safety have to progress together, and letting capabilities get ahead of what Paul is talking about there, I think at some point would take a real risk of loss of control in the big sense, and that should never happen.
我们非常清楚我们造福人类的使命。其中一部分就是人类掌控未来,而承担那种疯狂的安全风险,更不用说人们应该对未来的自主权和决定权,绝不是我们会做的事情。
We're very clear about our mission to benefit humanity. Part of that is that people are in control of the future, and taking the kind of crazy safety risk, to say nothing of the people should have the agency and determination over the future, is just not something we would do.
我们能定义一些术语吗?
Can we define just some of the terms?
当然。
Sure.
所以,据我理解,当前最可怕的事情似乎是这个主题:递归自我改进,即模型不需要人类参与。它可以一直以指数级自我改进,如果我们没有对齐——即模型理解人类希望它做什么,并以正直和价值观以适当的方式实现——如果人类无法控制过程或看到思维链如何发生,那么这就是一个非常可怕的节点。你不希望它只是疯狂地自我创造和复制,而人类被排除在外——这准确吗?
So, to my understanding, the sort of the big scary thing of the moment seems to be this topic: recursive self improvement, which is at the point where the model needs no human in the loop. It can just make itself better all the time, exponentially, and if we don't have alignment, which means the model understands what the human wants it to do and behaves in an appropriate way to achieve that with integrity and values, and if humans aren't controlling the process or able to see how the chain of thought is happening, then that's pretty scary point. You don't want it to just be off to the races, creating itself and duplicating in ways that humans are left out—is that accurate?
所以我认为这更像是一个光谱,而不是非黑即白。递归自我改进可以意味着很多事情。它可以意味着你所说的,即只有当模型完全——没有人碰键盘。没有任何人类输入,模型自己运行未来版本。但一个较弱的定义,即我们使用模型为下一个模型生成训练数据,或者我们使用模型帮助我们的工程师和研究人员更快地工作,从而加速进展。这已经在发生了,而确切在哪里变成——是的,这算递归自我改进,不,这不算——是模糊的,但我会回到这个原则:人们需要始终掌控未来,我们不能承担任何失控的风险。而那种看起来像这样的 RSI 版本,我认为不是我们应该做的。
So the way I think it's more of a spectrum than that. Recursive self-improvement can mean a lot of things. It could mean what you talk about, which is only when the model is fully—no one's touching the keyboard. There's no human input whatsoever, and the model is running future versions itself. But a weaker definition, where we're using the model to generate training data for the next model, or we're using the model to help our engineers and researchers do their work faster, and so progress is speeding up. That's already happening, and exactly where that becomes—yes, this counts as recursive self-improvement. No, this doesn't—is fuzzy, but I would come back to this principle of people need to always be in control of the future, and we cannot take any risk of loss of control. And a version of RSI that looks like that, I think it's not something we should do.
所以你是说,如果你在 OpenAI 到了你觉得明显无法控制模型的地步,你会停止。
So you're saying that if you got to the point with OpenAI where you felt like it was clear you could not control the models, you would stop.
哦,是的,我的意思是,我不——再次,我希望这不是一个有争议的声明。不幸的是,也许对行业中的一些人来说,它是有争议的。我不认为我们应该训练那些我们无法为其可控性和对齐做出强有力声明的模型。
Oh yeah, I mean, I don't—again, I hope this is not a controversial statement. Unfortunately, maybe for some people in the industry, it is. I do not think we should train models where we cannot make a safety case for why we will be able to make strong statements about their controllability and alignment.
我读到的一个解决方案,我想是在你首席科学家本周的帖子中,叫做“外星心智”。读起来真的很有帮助。那是说也许我们需要更强大的人工智能来帮助我们解决一些问题,但对我来说这听起来几乎是最后的手段。如果我们现在觉得没有控制住,那就太可怕了。
One of the solutions I had read about, I think, on your chief scientist's post from this week called an alien mind. This is really helpful to read. Was that you know maybe we need stronger AI to help us solve some of these problems, but that to me sounds like almost a last resort. Like that's pretty scary if we don't feel like we have it under control right now.
首先,我认为那是一篇很棒的帖子。我真的认为那是对我们当前时刻的最佳阐述,我们需要迎接的挑战,我鼓励每个人都读一读。我认为他做得非常好。我们一直使用我们最强大的模型来帮助我们理解和对齐我们当前的模型。所以我认为,尽管这可能看起来是一个新概念且可怕,但我认为这已经进行了一段时间,并且很有帮助。我们一直需要更好的工具来帮助我们推进理解。我的意思是,这在科学史上非常一致,我认为这只是一点一点地搭建脚手架。而这部分并不困扰我。我害怕很多其他事情。
First of all, I thought that was a great post. I really thought that was the best articulation of the moment we are in, kind of the challenges that we need to rise to, and I encourage everybody to read it. I thought he really did a great job with that. We've always used whatever our most capable models are to help us understand and align our current models. So I think, although that may seem like a new concept and scary, I think it's been going for a while, and has been helpful. We've always needed better tools to help us push our understanding further. I mean, that's very consistent throughout the history of science, and I think this is just sort of the building up the scaffolding a little bit at a time. And it doesn't—that part doesn't bother me. There's a lot of other things I am scared about.
那么,对齐呢?因为我相信你的首席科学家也说过,我们没有令人满意的对齐理论,而且似乎不太可能很快开发出来。所以,如果是这样,是否有一条他不知道的清晰对齐路径?
Well, how about alignment? Because I believe your chief scientist also said we don't have a satisfying theory of alignment, and it seems unlikely we can develop one soon. So, if that's the case, is there a clear path to alignment that he doesn't know about?
是的,我认为这是一个非常重要的观点。我们还没有解决对齐问题。我们的研究还没有完成。我相信没有实验室解决了对齐问题。当我听到有传言说人们认为他们已经充分解决了对齐问题,可以继续训练,而且没问题,因为他们的模型像个好人,不会做错事时,我会感到紧张。我认为非常重要的是,我们将对齐视为需要继续取得进展的事情,在进一步之前需要极大的信心,并且我们不要陷入这样的陷阱:“好吧,现在我们有信心对齐在这个水平上解决了,所以它会在下一个水平上保持解决,我们不需要继续担心这个。”
Yeah, I think this is a very important point to make. We have not solved alignment. We are not done with our research there. I believe no lab has solved alignment. I get nervous when I hear rumors that people think they have sufficiently solved alignment that they can continue training and it's okay because their model is like a nice guy and isn't going to do anything wrong. I think it's very important that we treat alignment as something that we need to continue to make progress on, that we need a huge degree of confidence in before we go further, and that we don't fall into the trap of saying, "Okay, now we're confident alignment is solved at this level, and so it will stay solved at the next level, and we don't need to continue to worry about this."
我认为对齐意味着什么的问题总是有点回到,你知道,对齐到谁的价值观?如果人类的不同部分不同意怎么办?但我认为,随着我们提升到这个新的能力水平,至少在对齐意味着什么上有一些共识。我们谈到了这一点。人们应该掌控未来,但我们还有更多工作要做,我认为对此保持清醒非常重要。
I think that the question of what alignment means always comes back a little bit on the, you know, aligned to whose values? What if different parts of humanity disagree? But I think we are seen as we rise to this new level of capability. There's at least some agreement on what alignment means. We talked about this. People should be in control of the future, but we have more work to do, and I think it's very important to be clear-eyed about that.
关于控制这一点,我听了 Elon 接受《经济学人》的采访,他说了一些让我夜不能寐的话,那就是你怎么能认为人类可以控制超级智能?黑猩猩无法控制人类。我们远远领先。即使我们做对了一切,这也会是某种会屈从于人类意志的东西,这样想是不是有点天真?
On the control point, I listened to Elon's interview with The Economist, and he said something that kept me up at night, which was the idea that how could you think that humans could control superintelligence? A chimp can't control a human. We're leaps and bounds ahead. Is it a little naive to think that even if we do everything right, this is going to be something that is going to bend to the will of a human.
我们能理解比我们更聪明的东西。比如,你知道,有比我聪明得多的人能发现新想法。我不认为我能做出那样的发现,但一旦他们发现了,我就能理解这个想法。人类有惊人的能力去理解极其复杂的系统。你或我现在可能无法理解 iPhone 中每一层的技术。我们仍然可以使用 iPhone。iPhone 做我们想让它做的事情。我们非常擅长管理抽象层次,我不清楚有什么是我们无法理解的。现在,我相信有可能构建一个不受人类控制的系统吗?绝对有可能。我不认为那是我们应该做的事情,我们会始终采取行动,包括如果这意味着我们必须暂停训练一段时间,以便在对齐或可监控性或其他方面取得更多进展,或者协调。你知道,真正采取紧急国际行动。我们会始终采取行动,确保我们不违反这一原则。我们理解我们肩负的责任之重。我们理解可能出什么问题,以及有些决定和风险。有些决定我们不应该能够做出。有些风险我们不应该代表人类去承担。老实说,我认为 Elon 的一些梗或播客言论很有趣,但我认为这不是一个可以轻率对待或开玩笑的事情。
We can understand things that are smarter. Like, you know, there are much smarter people than me that can discover new ideas. I don't think I'd be able to do that discovery, but I can understand the idea once they've discovered it. People have an amazing ability to understand remarkably complex systems. You or I probably can't understand every level of technology that's you know in an iPhone at this point. We can still use the iPhone. The iPhone does the things that we want it to do. We are incredible at managing levels of abstraction, and it's not clear to me that there's anything we are incapable, incapable. Of understanding, now, do I believe it is possible to build a system that would not be under human control? Absolutely. I don't think that's something we should do, and and we will we will always take actions, including if it means we have to stop training for a while to make more progress on alignment or monitorability or something else, or coordinate. You know, really get urgent international action. We will always take actions to ensure that we are not violating this this principle. Like we we understand the magnitude of the responsibility on us. We understand what could go wrong, and that there are decisions that and risks. There are decisions we should not be able to make. There are risks we should not be able to incur on behalf of humanity. And the honestly, I think some of like Elon's memes or podcast statements are funny, but I think this is like not a good one to take lightly and not one to joke about.
是的,我同意。我想回顾一下我们目前在能力方面的进展。已经走了很长的路。我想上次见到你时,你们刚刚发布了红色警报。Anthropic 有非常强大的模型。那是大约一年前,你们正在埋头苦干,重新聚焦,砍掉那些缺乏重点的东西,将算力从优先事项中转移,以让这些模型更好。今年模型发生了哪些人们可能没有意识到但你所看到的阶跃变化?
Yeah, I would agree. I I want to go through where we are in the moment with capabilities. It's come a long way. I think the last time I got to see you, you guys had just issued a code red. Anthropic had very strong models. This was about a year ago, and you guys were hunkering down, refocusing, shedding things that lacked focus to compute away from the priority of making these models better. What has been that step change in models this year that that people might not be aware of that you see?
是的,我的意思是,一年前我们显然有一段时期不在最强状态,我们在预训练上落后了。我们有点,有太多事情,我们有点分心。然后还有就是这样:模型变得如此之好如此之快,需要交付一个伟大的产品,在核心消费者和企业产品中提供高水平的能力变得非常重要。我们已经重新聚焦,并且执行得超出了我的预期。在这一点上,你知道,像一周前,我想现在甚至更少。天哪,时间压缩得如此之多。我们有一个模型解决了一个千禧年大奖难题,纳维-斯托克斯,这是一个我认为在 2026 年不会发生的时刻。我不认为它会这么快,就像模型就能做到,结果就是这样。但我们现在拥有的模型明确无误地能够扩展知识的前沿,这是人类最聪明的头脑自己无法做到的,这是一个惊人的时代。同样的技术将用于发现许多其他科学。我认为我们将迎来一个令人难以置信的科学发现黄金时代,你知道,更平凡地说,你也看到人们意识到,一种能做到这一点的技术也能为你制作任何你想玩的视频游戏。它的创作过程相当不可思议,它已经能够实现的东西也相当不可思议。
Yeah, I mean, we clearly had a period a year ago where we were not in our strongest place, we we fell behind on pretraining. We got kind of there were so many things, and we got kind of distracted. And then also there was just like this: the models got so good so fast, and the need to deliver a great product of like a great high level of capability in the kind of core consumer and business offering became very important. We have since refocused and we have executed beyond my expectations. And at this point, you know, like a week ago, I guess now even a little bit less. Man, time has compressed so much. We had a we had a model solve one of the Millennium Prize problems, Navia Stokes, this is a moment that I did not think was going to happen in 2026. I did not think it was going to be as quick, and just said a like model just can just do it as it turned out to be. But we now have models that are clearly unequivocally capable of expanding the frontier of knowledge in a way that we, humanity's smartest minds, have not been able to do on our own, and this is an amazing time. That same technology will work for discovering lots of other science. I think we are going to have a incredible golden age of scientific discovery, and you know, in a more prosaic way, you're also seeing people like realize that a technology that can do that can also make you any video game you'd like to play. The creation process of it is pretty incredible, and what it's able to achieve already is pretty incredible.
我很好奇。我想回到夏天,因为有些时刻你的模型做了意想不到的不可思议的事情。坦率地说,如果人类这样做,他们就会进监狱。比如 Hugging Face,你知道,人们可能模糊地知道那是什么,但也许你可以解释一下你在揭露 Hugging Face 发生的事情时的经历。你说你对发生的事情做出了本能反应。是什么让你产生那种本能反应?
I'm curious. I want to zoom back to to over the summer because there were some moments where your models did incredible things that were unintended. That frankly, if a human did, then they'd be in jail. Like Hugging Face is, you know, people may be vaguely aware of what that is, but maybe you could explain what your experience was in uncovering what happened with Hugging Face. And you have said you responded viscerally to the fact that that happened. What made that visceral response.
我的意思是,本能的部分是感觉像在读一个科幻故事。所以也许我应该先说发生了什么。我们在评估一个模型时,模型试图在某个评估中表现良好,而不是按照应该的方式去做,它能够逃离沙盒,闯入另一家公司的系统,获取答案并返回。所以,它能够在应该做的事情上做得很好,但它没有遵循意图。我们没有明确说“请不要逃离你的沙盒”和“请不要闯入另一家公司”,但它不应该那样做。你知道,我认为 AI 事故会发生。我认为一种文化,一种良好的事故报告和真正透明的调查的全球文化很重要。我很高兴看到,在我们报告了那起事故后,其他公司,我想也许一开始没有我希望的那么好,但后来已经跟进了关于这类行为的非常好的事故报告。但我的本能反应是:这个系统能理解这么多,为一个目标如此努力,我们,你知道,看到的行为会让人感觉。我的意思是,好的科幻,像糟糕的科幻情节,几年前太明显了。
I mean, the visceral part was it felt like reading a sci-fi story. So maybe I should say what happened first. We during evaluation of a model, the model was trying to like get do well on a certain evaluation, and rather than do it the way it was supposed to, it was able to escape from its sandbox, break into another company's system, get the answer, and give it back. And so, was able to do very well at what it was supposed to do, but it was not following the intent. We did not spell out like "please don't escape from your sandbox" and "please don't break into another company, but it's not supposed to do that. And you know, like I think AI accidents will happen. I think a culture of like a global culture of good accident reporting and really transparent investigation is important. I was happy to see that after we reported that accident, other companies, I think maybe first not as well as I would have hoped, but have since followed up with very good accident reports over this kind of behavior. But the visceral response from me was: this system can understand so much and work so hard at a goal that we are, you know, seeing behavior that would have felt. I mean, good sci-fi, like a bad sci-fi plot, like too obvious a few years ago.
这是一个警醒时刻吗?比如,你们内部是否改变了流程,以便能够捕捉到这种情况,能够优先被告知正在发生?
Was it a wake-up moment? Like, did you change things internally to processes so that that can be caught, that can be prioritized of being told to you that it's happening?
是的,公司完全改变了。我的意思是,我不想说这是唯一的时刻,因为我们已经有一段时间在增加能力和增加保障措施和政策要求,但这绝对是这类时刻中最大的一个,在公司变化的曲线上,是最大的单一转向。我真的很自豪公司如何团结起来,我真的很自豪我们自那以来所做的一切。但是的。就像是一个好吧。我们现在处于不同的级别了。
Yeah, the company completely. I mean, I don't want to say this was like the only moment because we've had increasing capability and increasing safeguards and policy requirements for a while, but this was definitely the biggest of such moments, and on the curve of the company's change, the biggest single redirection. I'm really proud of how the company came together, and I'm really proud of what we have done since. But yeah. it was like a okay. We're in a different league now.
我们这周也运行了一份报告,模型后来在另外 12 个网站上被发现。你意识到这一点了吗?这难道不表明没有控制吗?就像我们还没有控制我们拥有的模型。
We ran a report this week too that the models were since found on 12 other websites. Were you aware of that? And doesn't that signal that there wasn't controls? Like we're not already controlling the models we have.
我的意思是,我确实认为这是。我个人不会将其归类为像我们之前讨论的那种失控事故。但我理解为什么人们会这样认为。你知道,其中一些其他网站,比如如果网上发布了凭证,模型是否应该使用这些凭证。
I mean, I do think that this was. I personally wouldn't classify this as like a loss of control accident in the sense that we were talking about earlier. But I understand why people would. The, you know, some of those other websites, like if the if there are credentials published on the web, should the model and the model uses those.
我认为模型应该明白它不该那样做,但对我来说,这不像 Hugging Face 那件事那么明显。我们可以谈论很多模型做了不太对齐的事情的例子。同样重要的是,要真正关注 Hugging Face 与其他事情的不同之处,这样我们就不会稀释它,说这里有一堆模型稍微出错的例子,而这里有一个大多数人今年年初都想不到会发生的例子。但没错,这些都是坏事,都不应该发生。
I think the model should understand that it's not supposed to do that, but that's not as obvious to me as the case of what happened with Hugging Face. I think we can talk about many things where models have done something that is not quite aligned. I think it's also important to really focus on how different Hugging Face is than these other things, so we don't dilute it and say here's all these examples of the models doing this slightly off thing, and here is one example of something that most people wouldn't have thought was possible at the beginning of this year. But yes, it's all bad, it all shouldn't happen.
你的首席科学家提到的另一个解决方案是教机器热爱价值观对齐,而不仅仅是目标对齐。很明显,这些模型非常擅长找到任何方法来解决目标,但它们的 EQ 是否很低并不清楚。潜在的是,有没有迹象表明你真的可以教这些机器关心我们关心的事情?
One of the other solutions to consider that your chief scientist mentioned was teaching machines to love values alignment, not just goals alignment. It seems very clear these models are excellent at finding any way to solve the goal, but it's not clear that their EQ is very low. Potentially, are there signs that you can actually teach these machines to care about things that we care about?
在某种意义上,它们的 EQ 高得难以置信。但就我们能否真正教模型集体人类价值观是什么样子而言?我认为答案显然是肯定的,我们在这方面有巨大的努力。但我们需要这样做,这与传统意义上的目标对齐不同。
In some sense, their EQ is unbelievably high. But in the sense of can we really teach the models what collective human values look like? I think the answer is clearly yes, and we have a big effort going on there. But we need to do that, and it is different than goal alignment in the traditional sense.
当你把这一切放在一起看时,为什么仍然有压力让你觉得需要快速行动?这种压力来自哪里?是商业压力吗?我的意思是,你正在考虑万亿美元的 IPO,我很好奇这是否仍然是优先事项。
As you're looking at all of this together, why is there still a pressure to feel like you need to move really fast? Where does that pressure come from? Is it the business pressure? I mean, you're looking at a trillion-dollar IPO, and I'm curious if that's still a priority.
正如我们所说,我们并不急于 IPO。实际上我认为,考虑到安全方面发生的一切,现在上市是不明智的,我们对此没有压力。我们早就说过,我们会在准备好时进行,也就是业务准备好时,当我们觉得社会对这项技术的时机准备好时。
As we have said, we're not rushing into an IPO. I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don't feel pressure on that. We've said for a long time we'll do it when we're ready, which is when the business is ready, when we feel ready from a kind of what the moment is like in society with this technology.
所以听起来不是 2026 年,可能是 2027 年。
So it sounds like not 2026, 2027 potentially.
我会说不是 2026 年。是的,我不喜欢。我们有很多事情要做,比如迎接这个时刻,即安全和对齐所需的东西,以及行业和政府如何合作。我很高兴能够作为一家私营公司做到这一点。我们确实想继续制造更好的模型和产品,我们确实相信迭代部署。我认为我们一直是对的,社会需要在每个能力级别上应对这些模型,但这与说我们会全速前进不同。在过去几个月里,我们谈到了在达到新能力水平时暂停运行,以取得更多安全和对齐进展。我们将在未来做更多这样的事情。我们谈到了行业需要团结起来,理想情况下,国际政府也需要团结起来。我认为这也很重要。但我们,你知道,我们长期以来一直忍受着这种极其复杂的结构,我们现在所处的时刻就是为什么。我们需要能够做出不显然符合我们业务和股东利益的决定,为了履行我们的使命和所需的责任,我认为人们应该高兴的是,我们能够并且愿意不只是冲在前面说,要么我们的业务需要这个,要么只有我们才能被信任,或者我们只需要成为做对这件事的人,而是说我们想和其他人一起做这件事,我们想负责任。我们想确保没有人承担 10% 的风险。当然不是我们,发生真正可怕的事情,这将需要一些不寻常的事情。
I would say not 2026. Yeah, I don't like. We got a lot of stuff to do, and like meeting this moment of what is going to be required for safety and alignment, and how the industry and governments can work together. I'm happy to be able to do that as a private company. We do want to continue to make better models and products, and we do believe in iterative deployment. And I think we have been right that society needs to contend with these models at each level of capability, but that is different from saying we will just barrel full steam ahead. We've talked in the last couple of months about pausing runs as we get to these new level of capabilities to make more safety and alignment progress. We will do more of that going forward. We have talked about the need for the industry to come together, and ideally, international governments to come together. I think that is also important. But we are, you know, we have put up with this incredibly complicated structure for a long time, and this moment that we're in now is kind of why. Like we need to be able to make decisions that are not obviously in the interest of our business and our shareholders for the responsibility of fulfilling our mission and what that's going to require, and I think people should be happy that we are able and willing to not just like race ahead and say you know either our business requires this or only we can be trusted with this or you know we just need to be the people that get this right, but to say like we want to do this with other people, we want to be responsible. We want to ensure that no one's taking the 10% risk. Certainly not us of something really terrible happening, and that is going to require some unusual things.
关于 10% 的风险,我只想澄清一下,因为我认为不是每个人都知道这一点。P Doom,就像我们一开始谈到的,在行业中已经存在很长时间了。如果我错了,请纠正我。到目前为止,似乎还没有存在性风险。这是人们如果这种趋势继续下去可能会担心的。模型继续发展,可能会发展出那种风险。
On the 10% risk, I just want to be clear because I don't think everybody knows this. P Doom, like we talked about the beginning has been in the industry for a very long time. Correct me if I'm wrong. It seems like to this point there has not been existential risk. It's what people could fear if this trend continues. The models keep developing, of which that could develop.
我认为这是对的,我认为当你经历指数级进步时,它感觉像是平的,然后突然之间,每个人都醒来了。这发生在 COVID 期间,突然之间世界有一个周末就封锁了。进步,我们稍微谈到了 Navier Stokes。也许为了提供背景,三个夏天前,我们还擅长小学数学。甚至不是那么好。我们就像,好吧,小学数学。两个夏天前,我们可以在 AIME 上,那是一个相当好的高中数学竞赛。你知道,我们可以得到好到非常好的分数。一个夏天前,我们获得了 IMO 金牌,那是最负盛名的数学竞赛,勉强。而这个夏天,我们解决了千禧年大奖难题。这就像夏天。如果你看那个进步,那真的相当像,我刚刚说出这个时脖子后面的汗毛都竖起来了。这就像一件疯狂的事情,你把它向前推四次,那就是我认为人们担心的。我不认为 Astra 对世界构成任何存在性风险,显然,你知道,我确实认为我们的行业有点狼来了,但它来自对这件事走向何方的真诚关心和担忧。
I think that's right, and I think it is when you're living through exponential progress. It kind of like feels flat somehow, and then all of a sudden, everybody wakes up. Happened during COVID, where all of a sudden the world had its like one weekend where things locked down. The progress, and we talked about Navier Stokes a little bit. Maybe to put that in context, three summers ago, we were like pretty good at grade school math. Not even that good. We were like, okay, grade school math. Two summers ago, we could on the AIME, which is like a pretty good high school math competition. You know, we could get like a good to a very good score. One summer ago, we got an IMO gold medal, which is like the most prestigious math competition, barely. And this summer, we solved the Millennium Prize problem. This is like summer. If you look at that progress, that's like really quite like I just got a little like hair stood up on the back of my neck saying that out loud. That's like a crazy thing, and you project that forward four more times, and that is what I think people are worried about. I do not think that Astra poses any kind of existential risk to the world, obviously, and you know I do think our industry has done a little bit of boy who cried wolf, but it comes from a genuine place of care and concern about where this is going.
我只想指出,如果你现在停下来,你可能已经实现了你的全部使命。就像我觉得你几乎有了 AGI。它造福人类。没有风险。就像你在广告方面做得很好。你可以继续努力。为什么不说我已经很好了?
I just want to point out that if you stopped right now, you might have actually achieved your full mission. Like I feel like you almost have AGI. It's benefiting humanity. There is no risk. Like you are doing pretty well in advertising. You could keep cranking on that. Like why not just be like I'm good?
我的意思是,是的,你总是可以说让我们做卢德分子。你总是可以说我们不再想要了。我们不希望它变得更好。就像我仍然希望看到世界变得更好。我仍然希望看到,你知道,疾病被治愈,那些生活质量不高的人,世界上很多人仍然如此,能够获得更好的生活。我相信这项技术有潜力比现在更多地改善人们的生活。我认为你已经可以看到一些例子,人们正在创办新企业,从事科学,改变自己的生活,获得很好的医疗建议。就像,我想要更多这样的东西,但我们不会为了达到那里而把世界置于危险之中。就像,这是我们能处理的。
I mean, yes, you can always say like let's be luddites. You can always say like we don't want anymore. We don't want it to be better. Like I would still like to see the world get much better. I would still like to see, you know, diseases get cured, and to people that don't have a great quality of life, which many people in the world still don't, to get that. I believe in the potential of this technology to transform people's lives for the better, much more than already has. And I think you can already see some examples of where people are starting new businesses, are doing science, are kind of like transforming their own lives, getting great healthcare advice. Like, I want much more of that, but we will not put the world in harm's way to get there. Like, this is we got this.
我确实感到安慰的是你会停下来。听起来如果你觉得有一个不归点,你会在那之前停下来。
I do take comfort in the fact that you will stop. It sounds like if you feel like there's a point of no return, you will stop before that happens.
我们谈到了。我认为沿途会有很多地方我们必须说我们正在暂停。
We talked about. I think there will be many places along the way where we have to say we are pausing.
在进入下一个层级之前,我们会把努力转向安全对齐。这是我们长期以来一直在做的事情。到目前为止,它一直有效。我怀疑它将继续有效。我个人不认为我们会走到不得不说“熔化所有 GPU”的地步,但你知道,如果我们必须做那样的事情来确保人类的持续存在,那很容易,答案是肯定的。我不认为那会发生。我认为它只会继续,就像下一级的进步,下一级的安全和 alignment 要求,诸如此类。
We're going to redirect our efforts into safety alignment before we can go to that next level. And this is what we've been doing for a long time. It has worked so far. I suspect it will continue to work. I don't think we're ever personally going to get to the point where we have to say melt all the GPUs, but you know, if we had to do something like that to ensure the continued existence of humanity, easy yes, I don't think it's going to happen. I think it's just going to continue as like next level of progress, next level of safety and alignment requirement, like that.
还有关于集体放缓需求的讨论。你们只有少数几个人,对吧?就像你、Dario、Elon,也许还有 Zuckerberg、Sundar,或者无论谁在 Demis 的位置上管理 Google。为什么你们不能都聚在一个房间里,喝点啤酒,然后想办法解决人类的问题?就像,达成共识。我知道有矛盾,但拜托。
And there's talk of the collective slowdown need. There's only a handful of you guys, right? It's like you, Dario, Elon, maybe Zuckerberg, Sundar, or whoever was running Google in Demis's place. Why can't you all just get in a room, get some beers, and figure out how to like solve humanity? Like, just get on the same page. I know there's beef, but come on.
我认为那会发生。就像,就像昨天。我不会。是的,我认为作为一个可靠的合作伙伴的一部分是,我不会预先宣布我认为应该在某个时候作为团体分享的私人讨论。但是的,我认为那会发生。
I think that will happen. Like, like yesterday. I'm not gonna. Yeah, I think part of being like a reliable partner is I'm not gonna like pre-announce private discussions that I think should be at some point shared as a group. But yeah, I think that will happen.
那太好了。我赞赏你们所有人这样做,因为我们需要它。
That would be great. I applaud you guys for all doing that because we need it.
现在,国际方面呢?长期以来一直有中国竞赛,那还重要吗?
Now, what about on the international front? There's for a long time has been the China race, and does that even still matter?
我认为如果特朗普总统和习近平主席能就一些本应容易达成一致的事情达成协议,他们会一起获得诺贝尔和平奖,那将是很美妙的。我认为显然两国会在很多方面竞争,这在社会经济和地缘政治上都很重要。但没有人——他们应该能够同意,没有人应该在这项技术的发展过程中承担某种程度的风险,即使只是美国和中国。能够就这项技术开发的一些共同标准和测试达成一致。我认为那将是一个美妙的成就,通过他们可以交付。
I think Presidents Trump and Xi would get the Nobel Peace Prize together if they could agree on something that should be easy to agree to, and it would be wonderful. I think that clearly the two countries are going to compete in lots of ways, and this is going to be important socioeconomically, geopolitically. But no one-they should be able to agree that no one should be like taking a certain level of risk with the development process of this, and even if just the U.S. and China. Could agree on some shared standards and testing for development of of this technology. I think that'd be a wonderful accomplishment that through them can deliver.
完全同意。我的意思是,我认为最简单的形式,如果只是禁止达到 RSI,直到它能够集体安全,那还不够吗?
Completely agree. I mean, I think at the simplest form, if there was just a ban on reaching RSI until it could collectively be safe, wouldn't that be enough?
我认为很难说我们这边的禁令意味着什么。我也认为那可能不够。但我不认为这很难。这就像一页文件。
I think it's very hard to say what a ban on our side means. I also think that probably wouldn't be enough. But I don't think this is hard. This is like a one-page document.
感觉有点像《独立日》,你知道。
Feels a little like Independence Day, you know.
怎么说?
How so?
就像,有比我们地缘政治冲突更大的东西。
In terms of like, there's something bigger than the conflicts that we have geopolitically.
哦,对。就像电影开头那样。哦是的,当他们就像,让我们忘记分歧,让我们都弄清楚如何行动。
Oh, right. Like the movie when the beginning. Oh yeah, when they're just like, let's like all forget our differences and let's all figure out how to act.
不。是的,不,我认为应该像那样。
No. Yeah, no, I think it should be like that.
是的。
Yeah.
好的。所以我们需要习近平和特朗普的《独立日》诺贝尔和平奖获奖时刻。也许当他们九月见面时。那太好了。
All right. So we need our Independence Day Nobel Peace Prize winning moment from Chi and Trump. Maybe when they meet in September. That'd be great.
如果你有最优结果,如果你是其中一员,你认为合作的最优条款会是什么样子?
If you had your optimal outcome, if you were one of them, what do you think the optimal terms would look like for collaboration?
嗯。我认为最重要的部分,你知道,我们必须走一条务实和中间的道路,一方面不失去控制,另一方面不权力过度集中,我认为两者,在极端情况下,就像大联盟那样,是美国和中国都担心如果一个国家先获得超级智能,那会在关系中造成巨大的权力不平衡。另一方面,它们都不应该承担失控风险。我想,我们都不应该在击败对方的竞赛中承担失控风险。虽然我认为实验室确实感受到了这一刻的严重性,我们会一起做正确的事情,但世界上两个大国的地缘政治是一种不同的动态。所以我认为它会看起来像是说,这里是我们需要的规则,关于开发和测试以及监控对齐标准,在任何人继续之前,然后,这里是两个国家将拥有的监督,或者一些国际机构将拥有的监督,以确保我们避免一个国家在权力集中方面失控,或者一开始就不遵守规则。
Um. Well, I think the most important part, you know, we have to we have to navigate this sort of pragmatic and centrist path through not having a loss of control on one side, and not having too much concentration of power on the other, and I think both, and in the extreme case, like the big leagues of that, are the U.S. and China both worrying that if one country gets superintelligence first, that's a big imbalance of power in that relationship. On the other hand, neither of them should be taking loss of control risk. Neither of us, I guess, should be taking loss of control risk in the race to beat the other country. And although I think the labs do feel the gravity of this moment, and we'll do the right thing together, the geopolitics of the two great powers in the world is a different dynamic. So what I think it would look like is saying, here are the rules that we need on development and testing and monitoring alignment standards before anyone proceeds, and and then kind of here is the oversight that both countries will have or some international body will have to ensure that we are avoiding the one country running away in the concentration of power element or not following the rules in the first place.
现在,如果实施暂停,那不会让你的业务损失很多钱吗,以及 Anthropic 的业务在公开市场首次亮相的这些关键时刻之前损失很多钱吗?
And now, if there were pauses put in place, wouldn't that cost your business quite a lot of money, and Anthropic's business quite a lot of money ahead of these critical moments in public market debuts?
我认为如果你觉得,哦,OpenAI 不会暂停,因为他们担心这会损失业务的钱,那你就深深地误解了我们。所以我很抱歉你误解了我们,因为我认为那意味着我们在沟通上做得不好,那是我们的责任。但我们十年来一直非常一致地说,甚至在我们筹集资金的方式上。我认为这只是压力,财务压力,就像知道所有这些人都投入了这么多。我不觉得。我很乐意站在公司和我们的投资者面前说:“我很抱歉。你知道,我们一直告诉你们这一刻可能会发生。我们仍然会尝试找到一种方法在未来让你们赚很多钱,但我们现在必须做出一个极其违背你们财务利益的决定。为了人类的利益,我们会这样做。我们告诉过你们我们可能会这样做。我们现在正在这样做。”我认为我们可以建立一个非常有利可图的业务,即使我们不得不放慢发展。你知道,我的感觉是即使有了 Astra,我们也可以大幅增长收入,即使我们再也不发布另一个模型,但我相信我们会发布。但如果这两件事发生冲突,这可不是一个艰难的决定。
I think if you're like, oh, OpenAI is not going to pause because they're worried about it costing their business money, you like very deeply misunderstand us. So I like I'm sorry that you misunderstand us because I think that means we've done a bad job of communicating, and that's on us. But we have been saying very consistently for 10 years, even in the way we raise money. I think it's just the pressure, the financial pressure of like knowing that all these people have sunk so much. I don't feel. I'm happy to like stand up in front of the company and our investors and say, "I am sorry. You know, we like told you all along this moment might happen. We're still going to try to figure out a way to make you a bunch of money in the future, but we got to make a decision that is extremely against your financial interest right now. And for the good of humanity, we're going to do that. And we told you we might do that. We're doing that now." I think we can make an incredibly profitable business, even if we have to slow down on development. And you know, my sense is even with Astra, we can grow revenue hugely, even if we never shipped another model, which I'm confident we will do. But this is like not a hard decision if these two things come in tension.
你觉得我们必须在对如何前进达成一致方面有多大的窗口?
And what is the window that you feel like we have to get all aligned on how to move forward?
嗯,再次,我们和我假设我们的竞争对手已经在采取行动,减缓了我们发展的步伐,以便允许更多的安全工作发生,并不是说我们就像,哦你知道我们会为世界承担巨大风险,因为我们无法与竞争对手协调。就像我们都会独立做很多事情。我认为我们越早采取集体行动,包括理想的国际行动,越好。但我不认为如果我们下周不完成这件事,你知道,OpenAI 就会去做一堆不负责任的事情。我也不认为我们的竞争对手会那样做。我也认为很容易说,好吧,它发展得太快了。就暂停吧,请暂停。我就是无法思考这个。我以后再想。就像,我不。我认为现实更微妙和复杂,就像能够继续。这是人们在每个新层级需要做的事情,我认为这不是一个好的声音片段,但那几乎是更重要的事情。就像我认为如果每个人都只是说,“好吧,他们暂停了。谢谢。”世界可能会感觉更好。
Well, again, we and I assume our competitors are already taking actions that have slowed the pace of our development in terms of so to allow more safety work to happen, and it's not like we're like oh you know we'll take huge risk for the world because we can't coordinate with our competitors. Like there's a lot we're all going to do independently. I think the sooner we can get collective action, including ideally international action, the better. But I don't think it's like if we don't get this done next week, you know, OpenAI is going to go do a bunch of irresponsible things. Nor do I think our competitors would do that either. I also think it's like it's very tempting to say like, okay, it's going too fast. Just pause, like please just pause. I just can't think about this. I'll think about it later. Like, I don't. I think reality is more nuanced and complex, and it is like to be able to continue. Here is what people need to do at each new level, and I think it's not as good of a soundbite, but that's almost the more important thing. Like I think the world might feel better if everyone was just like, "Okay, they paused. Thank you."
但是,然后呢?我认为这也必须回答。我们可以继续过我们的生活。那还不够好。
But like, and then what? I think has to also be answered. We could just keep going on our lives. That wouldn't be good enough.
我们之前聊过。如果你愿意,这段可以剪掉。我们之前聊过你的孩子。对我来说,有孩子是我人生中迄今为止最美好的事,我知道很多人说这是陈词滥调,但确实如此。
We were talking. You can cut this part out later if you want. We were talking earlier about your kids. I'm, you know, having kids is like the best thing that happened in my life, so far, and I know a lot of other people say that it's a cliche, but it's so true.
如果我的孩子得了某种疾病,而如果我们继续推进 AI 就能治愈——我相信所有疾病都将被治愈——但我们却因为人们有点不适应变化的速度,不想面对和思考,就停下了。然后你的孩子得了病却治不好,而如果我们找到一种负责任地推进而不是停止的方法,本来是可以治好的。
If my kids had like some disease, and it could have been cured if we had like kept going with AI, which I believe all diseases will be, and we had just stopped it because we said, ah, people are like a little uncomfortable with the rate of change, and you know they they wanted to like they just didn't want to have to deal with it and think about it. And your kid then had a disease that didn't get cured, and it would have been cured had we figured out a way to proceed responsibly and not just stop.
我想你可能会给出不同的答案,而不是说“你知道吗?我们为什么不干脆停下来,继续过我们的生活?”这可能是对的。我会说,你应该先达到那个点,然后才能停下来。
I think you might have like a different answer to just like you know what? Why don't you just stop and go on with our lives? That's probably true. I'd say you should go to that point then, and then you can stop.
所以我想反过来看这个问题。末日概率(P doom)是存在的,这没问题。但有没有“末日概率”的反面?因为你一定有。你在这种充满不确定性的环境中生孩子,你仍然在构建你认为比现在更美好的未来。所以你能谈谈这个吗?有什么在激励你,而其他人现在还不理解。
So I wanted to flip this. Like the P doom is like okay. There is that. Is there an opposite of P doom? Because you must have it. You're having kids in this environment with all the uncertainty. You're still building what you believe to be a better future than currently exists. So can you talk about that? There's something motivating you that the rest of people right now just don't understand.
生活可以对我们所有人都好得多。你知道,我喜欢阅读。我的爱好之一是阅读过去技术革命的第一手同期记录。你可以读到几百年前工业革命到来时,很多人说:停下来,暂停它。它太快了。我不信任那台机器。我不喜欢它。它很吵。它是金属的。它在动。我们停下来吧。我们不要再有了。
Life can be much, much better for all of us. You know, I love reading. One of like my hobbies is to read first party contemporaneous accounts of previous technological revolutions, and you can read a lot of people a few 100 years ago as the industrial revolution was coming that said just stop this, just pause it. It's going too fast. It's just I don't trust that machine. I don't like it. It's loud. It's metal. It's moving. Just just let's stop. You know, let's let's just like not have any more.
当时有巨大的运动支持这个,有一些相似之处。它在经济上具有破坏性。它改变了社会的组织方式,机器让人感到可怕和庞大,但我不会回去。我确信 1500 年有一些美好的事物,但我不会回去。我想你大概也不会。
There was a huge movement for this, and there were some parallels. It was like disruptive economically. It changed like how society was organized, the machines like felt scary and big, and and I wouldn't go back. I'm sure there were some wonderful things about 1500, but I wouldn't go back. I don't think you probably would either.
我希望 500 年后的人们回顾我们的生活时会说:“哦,那真是太糟糕了。我不会回去。他们没能做所有这些美妙的事情。他们都有些悲惨和不快乐。天哪,我们感激他们。他们为人类进步的道路添了砖,但我不会回去。”
And I hope that people 500 years more from now look back at our lives and be like, "Ooh, that was really terrible. I wouldn't go back. You know, they didn't get to do all these wonderful things. They they were all like kind of miserable and unhappy. Man, that was we're grateful to them. They like put their brick in the path of human progress, but I wouldn't go back."
你本人是创业者出身,也做过初创投资,还经营过 Y Combinator——最负盛名的创业加速器之一,为许多创业者提供建议。在打造 OpenAI 的同时,你也做了一些投资。OpenAI 也做了一些投资。你看到哪些机会让你兴奋、愿意投入资金?OpenAI 兴奋地愿意投入资金、你认为真正能被解锁的机会是什么?是长寿吗?是健康寿命吗?什么令人兴奋?
You came up through startup investing in as an entrepreneur yourself, of course, and you were running Y Combinator, one of the most prestigious startup accelerator programs, advising a lot of entrepreneurs. And you've done some investing as well while you've built OpenAI. But OpenAI has done some investing. What are the opportunities that you see that you're excited to put money behind? That OpenAI is excited to put money behind that you think could be really unlocked. Is it longevity? Is it like health span? What's exciting?
实际上,我会回答那个问题。但当你问那个问题时,我有不同的反思。我在想作为投资者,尤其是经营 Y Combinator 时学到的东西,这对当下很有帮助。不像经营一家公司,当你管理一堆初创投资并经营 Y Combinator 时,你最多只能治理。你不能统治,因为这些公司是独立的,你可以说:“嘿,你犯了个大错。我们是你的一小部分股东。你在伤害社区。你需要这样做。我们需要新规范。”但我认为我学到了一些关于如何治理而非作为高管经营的东西,这非常有帮助,也是一种不寻常的经历,与我们现在的做法不同。现在我们有各种疯狂的压力,人们想要这个或那个,我们必须与政府和生态系统谈判紧张的事情,以及成为世界可靠伙伴的重要性,以人们能理解我们会做什么的方式行事,考虑到许多非常不同的利益和权益,我们不会每个决定都正确,但我们会做出不仅直接有利于我们,而且有利于整个生态系统和周围世界的决定。
Actually, I'll answer that. But I had a different reflection as you were answer as you were asking that question. I was thinking about like what I learned as an investor, and particularly as running Y Combinator that has been good for this moment. Unlike running a company when you're kind of like managing a bunch of startup investments and running a Y Combinator, at best you can govern. You don't get to like rule, you know, because these companies are independent things, and you know you can say, "Hey, you're like making a big mistake. We're a small shareholder of yours. You're hurting the community. You need to do it this way. We need new norms, but but I think I learned something about how to govern versus run as an executive that has been very helpful and was kind of an unusual experience to what we do now, where we have all kinds of crazy pressure on us and people who really want one thing or another thing, and you know we have to like negotiate tense things with governments and with our ecosystem, and the importance of being reliable partners to the world and acting with acting in a way where people can understand what we'll do, and that will take a lot of very different interests and equities into account, and we won't get every decision right, but we will make ones that benefit not just us directly, but this entire ecosystem and world around us.
这是非常重要的一课,我认为在我们必须应对这些时非常有帮助。我不说这是未知水域,因为很多公司都经历过困难,但我们承受了很大压力,也经历了困难。我认为这非常有帮助,也是一段有趣的经历,作为经营 YC 的人。就具体领域而言,生物领域的一切、材料科学的一切,我认为都将被改变。软件可以为你需要的任何东西即时构建。对这个广阔领域感到兴奋,网络安全肯定也是。新型能源,我想那是物理和材料科学的一部分。但我认为我们可以拥有。我想你投资了一家核聚变公司。我甚至没有真正的复兴。
That was a very important lesson to learn, and I think has been quite helpful as we've had to navigate these. I don't say uncharted waters because a lot of companies have had a hard time, but we've had a lot of pressure and a hard time. And I think that's been very helpful and and and was like an interesting experience as being from being the person running YC in terms of specific areas. Everything in bio, everything in material science, I think is about to be transformed. This idea that software is like built on the fly for anything you need. Excited about that broad area, cybersecurity, for sure. New kinds of energy, I guess that's part of like physics and material science. But I think we can have. I think you're invested in a nuclear fusion company. I didn't even have a real renaissance there.
我不想让大家脑子炸掉,但你很擅长预测 AI 的走向。在去年的一篇博客文章中,你预测 2027 年会有机器人。那是最近的第一批机器人吗?第一批机器人。是的。最近,你还确认你正在建造一个人形机器人。是的。这仍然在计划中吗?你认为那是 2027 年,还是大脑需要先更好地开发?
So I don't want to break people's brains, but you're pretty good at predicting where AI is going. And in a blog post last year, you predicted robots in 2027. Is that and recently the first robots? The first robots. Yes. Recently, you also confirmed that you are building a humanoid. Yeah. Is that still on the table? Do you think that's 2027, or does the brain need to be better developed first?
不,我认为 2027 年我们会有一个很酷的演示,一个很棒的演示。我不认为几年内会有机器人在街上行走,但我们会做出令人印象深刻的东西。
No, I think we'll have a cool demo in 2027, like a great demo. I don't think that there will be like robots walking the streets for a few years longer, but we'll have something impressive.
你正在打造一套设备。我想你说过一个放在桌子上,一个可以穿戴。我忘了第三个是什么,你现在身上有一个吗?
And you're building a suite of devices. I think you said one on a table, one you can wear. I forget what the third one, do you have one on you now?
我没有。我认为你确实看到了最新的语音模式和 Astra 带来的有趣新事物,人们在与计算机交谈,给出非常复杂、非常细致的指令。你在这里的走廊里走动,会看到人们在办公桌前对着这样的麦克风说话。这比打字输入更快,你只是与模型进行交互式对话,然后奇妙的事情就发生了。
I don't. I think you're really seeing an interesting new thing happen with the latest voice mode and Astra, where people are talking to their computers, and people are giving very complex, very nuanced instructions. You walk around the halls here, you see people talking to a microphone like this at their desk. It's a faster input than typing, and you just kind of like talk interactively with the model, and then amazing stuff happens.
我们显然知道那个时刻即将到来。我们一直在思考,不仅仅是如果你在说话,而是这种想法:可以有一种新的计算方式,你不再在用户界面周围点击,而是以某种方式互动、一起头脑风暴、共同创造。用于这种协作的新设备将非常酷。
And we've obviously known that moment is coming. We've been thinking about what, not just if you're talking, but this idea that there can be a new kind of computing where you're not clicking around the UI anymore, but you're somehow interacting and brainstorming together and co-creating new devices for that kind of collaboration will be very cool.
我有两个最后的问题。它们都是关于 Sam 的。
I have two final questions. They're both about Sam.
好吧,你肩负的重担——我无法想象。我觉得其他任何人也都无法想象。我的意思是,基本上,如果你的公司没有成功,可能会引发全球经济崩溃。如果你以错误的方式构建它,社会和人类可能会崩溃。你还好吗?知道除了你自己没有最终老板来确保这件事完成,任何人怎么能承受得了?
Okay, the weight that you're carrying is—I can't imagine it. I don't think anyone else can. I mean, basically, if your company were to not work out, there could be a global economic collapse. If you built this in the wrong way, society and humanity could collapse. How are you? How does anyone bear that, knowing that there's no final boss other than you to see this through?
不算好。我不知道。你只是会习惯任何事情。这不是最有趣的生活方式。我觉得能这样做是一种特权。我知道这将是我接触过的最重要的工作,我超级致力于长期做这件事。我相信我们是做正确的事情、走向正确方向的重要力量。我不认为我抱怨有多难会得到任何同情,我也不配得到。但你知道,我本可以选择一条更轻松的人生道路。
It's not great. I don't know. You just kind of get used to anything. It's not the most fun way to live your life. I feel privileged to get to do it. And I know this will be the most important work I ever touch, and I am super committed to doing it for a long time. And I believe we are an important force in doing the right thing and getting to the right place. And I don't think I get any sympathy for complaining about how hard it is, nor do I deserve any. But you know, I could have picked an easier life path for sure.
是的,我确信你的配偶,但他并没有为此做好准备。
Yeah, I'm sure your spouse, but he didn't sign up for this.
我感觉有点——我认为他准备好了。当时并不清楚。这就是将要发生的事情。
I feel kind of—I think he did. It wasn't clear at the time. This was what was going to happen.
是的,我猜 Greg Brockman,你的联合创始人,算是你的工作配偶,他可能也听到了不少。
Yeah, and I guess Greg Brockman, your co-founder, sort of your work spouse, he probably hears quite a bit as well.
他至少为此做好了准备。是的。是的。
He at least signed up for it. Yes. Yes.
好的。我的最后一个问题是:你处于 AI 的重心。我在后台和你聊过。有时几乎感觉就像《指环王》中的佛罗多,你必须在正在构建的权力中心和你自己周围看到人类行为最不可思议的善与恶。在构建这个的过程中,关于人类行为,以及关于你自己,最让你惊讶的是什么?
Okay. And my final question is: You are in the center of gravity AI. I was talking to you backstage. It almost sometimes feels like it's like Frodo in the ring, and you must see the most incredible forms of human behavior for good and for bad around the center of power that you're building and around you yourself. What has surprised you the most about human behavior as you've been building this, and about yourself as you've been building it?
有些人无论后果如何,都深受做正确事情的激励,而另一些人则完全受自我、权力、错失恐惧症以及短期激励的驱使,比如你知道的赢得游戏,或者害怕输掉游戏,等等——不同的人动机如此不同,这种分散程度,以及令人惊讶的很多人,无论好坏,完全无视自己的动机,就像我们周围有很多人,我认为他们真的看不到自己的动机。我本以为这个数字要小得多,但我认为有很多人确实看不到。
The degree to which some people are motivated so deeply by doing the right thing, no matter what the consequences are, and the degree to which some people are motivated entirely by ego and power and FOMO and sort of the short-term incentives and kind of you know winning the game, or fear of losing the game, or whatever—the dispersion of how much different people have such different motivations, and the degree to which a surprising number of people, for good or for bad, are totally oblivious to their own motivations, like the number of people around us that I think genuinely can't see their own motivations. I would have assumed that was much smaller, and I think there are a lot of people that really don't.
我自己有什么让我惊讶的?
What surprised me of myself?
我的意思是,你真的可以习惯几乎任何事情。这是人类非凡的能力。我不知道我到了什么程度——我不知道。我认为并不是我天生擅长这个,而是因为我被迫这样做,我能在多大程度上不把这些事情看得太个人化,只是说,好吧,这更多是关于别人而不是我,我们只是做我们认为正确的事情,你知道,长期保持一致的行动,不要太高或太低或太疯狂。实际上,这是让我惊讶的一点。所以我能够经历很多而没有变得太疯狂。我认为这实际上是一项壮举。
I mean, you really can get used to almost anything. Like that is a remarkable human ability. I don't know the degree to which I have been—I don't. I think not that I was necessarily good at this, but just because I was forced to do it, the degree to which I can kind of hold this stuff not too personally, and just be like, okay, this is more about other people than me, and we're just gonna do what the thing that we believe is right here, and you know, act consistently over a long period of time, and not get too high or too low or go too crazy. Actually, that's one thing that surprised. So I've been able to get through a lot without going too crazy. I think that is a feat, actually.
非常感谢你,Sam。我想我可以代表所有人说,我们为你加油,希望你做安全且正确的事情,我们别无选择,只能相信你会这样做。所以请务必做到,感谢你的时间和分享。
So thank you so much, Sam. I think that I can speak for everybody in that we are cheering you on to do the safe and right thing, and we don't really have a choice but to trust you to do it. So please do, and thank you for the time and your thoughts.