OpenAI's New CEO on Leadership Changes and Scaling to Every Enterprise
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 新任 CEO 讨论关键高管离职、公司对企业级扩展的关注,以及将经济转变为算力驱动的经济。
OpenAI's new CEO discusses the departures of key executives, the company's focus on enterprise scaling, and transforming the economy to be compute-powered.
我觉得先从本周的新闻说起会很有帮助。Denise Dresser,你的首席营收官离开公司,你引进了一位新的,Brad Lap 也离开了。嗯,外界对 OpenAI 的领导层储备和正在发生的事情有很多疑问,而你进来了,现在管理着很多这方面的事务,我很想听听你的看法。这些离职事件之间有关联吗?它们各不相同吗?有没有什么共同主线?
I think it would be helpful to start with just the news this week. So Denise Dresser, your CRO leaving the company, you bringing in a new one, Brad Lap leaving. Um, there are a lot of questions publicly about OpenAI's leadership bench and what's going on and you've come in and are now managing a lot of this and I'd love to hear from you. What are all these departures connected? Are they different? What's what is there any through line there?
你看,我的想法是,今年的主题确实是专注,对吧?我们确实看到了如此不可思议的技术前景,但仅仅生产技术还不够。我们需要真正弄清楚如何让这项技术惠及大众?如何把它带入世界上的每一家企业,真正提升他们正在做的事情,帮助他们实现目标。所以这确实是我们一直在追求的唯一焦点,而在追求这个目标的过程中,我们一直在思考我们是否有正确的结构、正确的战略,是否具备所有正确的要素。我认为我对 OpenAI 的看法是,我们是一个非常坚韧的组织。嗯,但我们做的很多事情确实是在努力打造能够真正帮助我们实现这一目标的领导团队。所以我认为,我对每一位曾经在这里的人和已经离开的人的看法是,他们都为 OpenAI 的今天做出了贡献。展望未来,我认为我们正在系统性地扩展到世界上的每一家企业。我们把模型带给各地的消费者,而我的很多焦点就是确保我们能够实现这一点。
Look, the way that I think about it is that the theme of this year has really been focus, right? Has really been that we have line of sight to such incredible technology, but it's not enough just to produce the technology. We need to really figure out how do we have this technology benefit people? how do we bring it into every enterprise in the world in a way that actually uplifts what they're doing, helps them accomplish their goals. So that's really been the singular focus that we've been pursuing and that in the purpose in the pursuit of that that we've been really trying to think about do we have the right structure, do we have the right strategy, do we have all the right things there and I think that my view of OpenAI is that we're very resilient organization. Um but a lot of what we've been doing is really trying to build the leadership team that uh can really help us achieve that goal. So I think that my view of everyone who has been here and people who who are no longer here is that they've all helped to contribute to what OpenAI is today. And looking forward I think that we are scaling systematically to every enterprise in the world. We're bringing our models to consumers everywhere and that a lot of my focus has been making sure that we can achieve that.
具体到首席营收官这一块,嗯,你之前缺少什么你需要的东西?你知道,是不是可以这样理解?
On the CRO side specifically, um what is it that you weren't getting that you needed? you know, is is is that the way to think about it?
你看,我会换个说法。比如一年前,我认为在企业级市场我们根本还没入场,对吧。那真的不是我们的重点。我们还没有真正建立起所需的肌肉。而现在,我认为我们实际上已经入场了。我们确实已经打下了一个基础,我认为这是我们可以真正依托的。现在我们正在系统性地扩展。而对我来说,这就是关键时刻。从某种程度上说,这真的不是关于个人的问题。这关乎整个领域的总体方向。
Look, I would say it differently. Like a year ago, I think that we really were not in the game at all, right, when it came to enterprise. It just wasn't really our focus. We hadn't really built the muscle we needed. And now, I think we actually are in the game. Like, we kind of have really built a foundation that I think is actually something we can really build on. And now we're systematically scaling. And that that to me is like the moment. It's not really about the individuals in some ways. It's really about the overall sort of direction of the field.
这如何改变你所说的系统性扩展?这对 OpenAI 实际意味着什么?
How does that change this systematically scaling you're talking about? What does that mean practically for open AI?
嗯,首先,我认为我们正在着眼于转变整个经济。我觉得这是从外部很难看到的东西,对吧?感觉上好像是,嘿,这是一门生意,我们在卖 SaaS 软件。但那不是我们在做的事情,对吧?我们在做非常不同的事情。我们正在把经济转变为一个由算力驱动的经济。我认为企业的定义本身都会改变。所以问题是,从哪里开始?如何确保你理解每一个细分市场,并且能够以可复制的方式向所有市场销售?所以我们一直在做的很多动作——再说一次,这从某种程度上说并不是新的动作——只是鉴于业务现在所处的位置,我们需要扩展它,就是要真正确保一线的每个人都能把我们从客户那里看到的洞察以非常系统的方式反馈给我们的产品团队和研究团队,反之亦然。比如我们学到的一件事,如果你看看我们内部如何写博客文章,我们多年前就学到,这必须是协作,对吧?你不能像交接那样——哦,这是写博客的公关人员,这是生产模型的研究人员,然后他们就这样某种方式交接。必须是深度协作。你需要让技术专家真正执笔,并对最终结果的良好负责。他们可以被那些擅长如何构建叙事、如何整合所有这些资产、如何做好发布的人放大。
Well, number one uh is that I think that we are looking at like transforming the entire economy. Like that's the thing that I think is very hard to see from the outside, right? Is it kind of feels like hey, it's a business like we're bu selling SAS software. That's not what we're doing, right? We're doing something very different. We are transforming the economy to be a compute powered one. And I think that what an enterprise even is will change. And so the question is well where do you start? How do you make sure that you understand every segment and you're able to sell in a repeatable fashion to all of them? And so a lot of the motion that we've been working on and again this isn't a it's it's not it's not a new motion in some ways. is just one that we need to now scale just given where the business is at is to really make sure that everyone in the field we're able to funnel back the insights that we're seeing from customers in a very systematic way back to our product teams our research teams and vice versa like one thing that we have learned like if you look at for example how we produce blog posts internally that we learned years ago that it has to be a collaboration right you cannot have a handoff of oh here's the comms people who write blog posts and here's the research people who produce models and you know that they'll just like kind of hand hand it off in some way. It has to be deep collaboration. You need to have the technical expert really hold the pen and feel accountability for the end result being good. And they can be amplified by people who are experts on how do you actually craft a narrative? How do you sort of bring together all these assets? How do you have a good launch?
所以对我来说,我们追求的核心活动就是说我们不能碎片化。我们不能同时做所有的事情。我们必须统一且专注。所以这就是我们内部工作方式的答案。在外部方面,我认为你应该看到,我们会系统性地优先考虑不同的行业,真正专注于深入理解指标,确切知道 token 支出增长最快的地方,并真正专注于结果。就像现在每家企业都在问投资回报率在哪里。是的。
And so that to me is the core activity that we are pursuing is to say we can't be fragmented. We can't be working on all the things. We have to be unified and focused. And so that's the internal answer for how we're working. And in terms of the external, I think that you should see that across that we will sort of systematically be prioritizing different industries, really focused on the uh the deep understanding of the metrics and exactly where the token spend is growing the most and really focused on outcomes. Like every business right now is asking where's the ROI. Yeah.
而这是我们的一大重点。我们一直以来的优势是拥有性价比最高的模型,但我们确实必须向人们展示这对他们的结果。嗯,还要通过渠道和合作伙伴进一步扩展。这也是我们相信的。所以我认为我们现在的状态是,我们对这一动作的基本面非常有信心,对基本面——我们拥有出色的产品和出色的模型——非常有信心,然后就是迅速壮大我们的内部团队,以便我们能够触达每一个人,但要以一种真正产生我们所期望的积极影响的方式来做。
And that is something that is a big focus of ours. Ours has always been the hey we have the best price performant models, but we really have to show people that in the outcomes for them. um also really scaling further through channels and through through partners. That's something we believe in as well. And so I think that a lot of where we're at is we feel very confident about the fundamentals of the motion and the fundamentals of we have a great product, a great model and then really just rapidly rapidly growing both our internal team so that we can reach everyone but doing that in a way that that really has the that positive effect that we're looking for.
你用了“负责”这个词,我很好奇,从你的角度看,由于发生的领导层变动,你现在在 OpenAI 负责什么?Fiji Simo 曾短暂加入。由于健康原因她不得不永久退出。在她离开期间你接手了公司的很大一部分。现在实际上是你和 Sam 一起在运营。这是如何分工的?
You use the word accountability and I'm curious from your perspective, what are you accountable for now at OpenAI because of the leadership changes that have happened? Fiji Simo came in for a time. She had to step back permanently due to her health. You stepped in while she was uh out and took over a large part of of the company. Now it's you and Sam effectively running it together. How was that divided up?
所以在过去几年里,我一直负责我们正在做的基础设施方面。比如数据中心、算力之类的事情。这部分仍然保留。嗯,在软件方面,比如内核以及如何训练这些大型模型,我也曾负责很大一部分范围,但现在已经能够交接出去,这样我就可以真正专注于业务方面。所以我会把自己看作是对营收、产品营销、整个机器——把这些模型从研究带到为客户创造价值——端到端负责的人。
So for the past couple years, I've been accountable for the infrastructure side of what we're doing. So data centered, compute, those kinds of things. And that that part remains um that there's a large part of the scope that I was also working on in terms of the software for kernels and sort of how you train these big models that that's something that I've I've been able to hand off so that I can really focus on the business side. And so I would view me as sort of endto-end accountable for revenue for product marketing the whole machine for bringing these models from research to value for our customers.
当你看着这台机器——我知道你还在逐步完善——但当你看着眼前这台你负责的机器时,它跟竞争对手相比表现如何?我的意思是,具体到 Anthropic,有报道说他们的年度经常性收入(ARR)可能是 OpenAI 的两倍多。你怎么看?
When you look at this machine, and I know you're still putting pieces in place, but when you look at this machine you have in front of you that you're accountable for, how is it doing relative to competition? I mean, I think specifically Anthropic, there's reporting that their ARR is potentially more than double what OpenAI's is. What do you make of that?
嗯,我觉得我们起步晚了。具体来说,看待整个领域的方式是,过去几年,聊天市场是主要增长点。原因很简单:GPT-4 是第一个结果足够个性化、让你觉得提问、读答案有意思的模型。所以它很自然地适合聊天,对吧,适合聊天机器人。这是我们能够规模化并交付价值的东西。现在,每周有 10 亿人使用 ChatGPT。10 亿人——这就是规模。这是真正的规模。没有其他人做到这一点。
Well, I think we started late. Specifically, the way to think about the overall field is that for the past couple of years, the chat market was the main growing one. The reason for this was simple: GPT-4 was the first model where the results were personalized enough that it made sense for you to ask a question, read the answer, and it's interesting to you. So it very naturally lends itself to a chat, right, to a chatbot. And that is something we're able to scale and deliver value. Now, a billion people every single week use ChatGPT. A billion people—that is scale. That is real scale. No one else is doing that.
但今年发生了另一件事:智能体时刻。这真的是去年 12 月,当时我们和竞争对手都迎来了一个时刻,智能体真正开始以显著的方式做软件工程。从说你可以把工作效率提高 20%,变成了 80% 的工作现在可以用 AI 完成。所以突然之间,软件工程师的运作方式发生了反转——你变得更像智能体的管理者。从那以后,增长是前所未有的,绝对前所未有的,这些智能体被采用的速度。
But this year something else happened: the agentic moment. This really was December of last year, when both we and competition had a moment where agents were really able to first start doing software engineering in a significant way. It went from saying that you could speed up your work with maybe 20% productivity increase to 80% of the work could now be done with an AI. And so suddenly the way that software engineers operated inverted—you became much more of a manager of agents. And since then the growth has been unprecedented, absolutely unprecedented, how these agents have been adopted.
还有一件事,我觉得我们特别在编码上起步晚了。实际上,让我重新——
And one thing that happened was I think that we focused late on coding in particular. Actually, let me re—
让我也给出一个不同的答案。所以有一个简单的说法,就是我们不太关注现实世界的编码,我们进入那个游戏晚了,因为我们总是看编码竞赛基准之类的东西,并且总是领先,但不太关注开发者真正会怎么使用它——混乱的现实代码库、开发者中途打断、模型的个性——这些最后一步的小麻烦实际上对采用率影响巨大。在市场化(GTM)方面也类似,真正思考如何与客户建立关系,如何构建他们需要的可观测性控制,如何帮助他们理解在他们的上下文中什么是可能的。
Let me give a different answer as well. So there's one simple narrative, which is that we focused less on real-world coding, and we came to that game late because we always had looked at things like the coding competition benchmarks and always had the lead there, but less on how a developer is really going to use it—the messy real-world codebase, the developer interrupting in the middle, what's the personality of the model—these sort of little last-mile paper cuts that actually make a huge difference in adoption. And similar on the GTM side, really thinking about how do you build relationships with customers, how do you build the kind of observability controls that they need, how do you help them understand what's possible within their context.
而所有这些都不是我们唯一的关注点,直到比竞争对手晚得多。所以我认为我们起步晚了,但我们的轨迹非常陡峭。但实际上还有一个从外部不太明显的更深层答案,那就是模型能力。大约两年的时间,从 2024 年 4 月我们发布 GPT-4o 到 2025 年 12 月,我们真的在榨取 GPT-4 模型的汁液,对吧?我们有一堆正在改进的部件,但基本上我们还在同一个底层基础模型上,同一个底层预训练。这有很多原因。真的因为要生产并发布一个预训练需要全公司的整合。这不是任何一个团队的事;而是这些团队之间的整合,要考虑推理、考虑后训练、确保你共同设计了合适大小的模型。所有这些事情你必须完全做对,并且有一种和谐的工作方式。
And all of these were less our singular focus until much later than competition. So I think that we started late, but our trajectory has been very steep. But there's actually an even deeper answer that is less obvious from the outside, and that is one of model capability. For a period of about two years, right from when we announced GPT-4o in April of 2024 to December of 2025, we were really squeezing the juice out of the GPT-4 model, right? We had a bunch of pieces that we were improving, but fundamentally we were on the same underlying base model, the same underlying pre-train. And there's a lot of reasons for that. It's really because to produce and ship a pre-train requires full integration across the company. It's not any one team; it's about the integration across these teams to think about the inference, to think about the post-training, to make sure you've co-designed the model to be the right size. All of these things you have to get exactly right and have a harmonious way of working.
在 2025 年期间,我们的很多焦点,我个人很多焦点,是如何真正回到把发布新的预训练作为我们创新引擎的核心部分。
And over the course of 2025, a lot of our focus, a lot of my personal focus, was how do we really get back to shipping new pre-trains as a core part of our innovation engine.
而今年,你看到了结果。所以我们实际上已经发布了。我们正在第三个预训练上。我们实际上即将发布那个。
And this year, you see the results. So we've actually shipped already. We're on our third pre-train. We're actually about to ship that.
这是 Astra。
This is Astra.
这是 Astra。
This is Astra.
是的。
Yeah.
你看到的很多收益是因为我们让模型从根本上变得更好。所以思考方式是,有一个底层技术领域我们确实落后了。我们创新机器有一个核心部分我们没能完全发挥出来。但现在不同了。
And a lot of the gains that you see are because we have made the model so fundamentally better. And so the way to think about it is that there's an underlying technological place where we really were lagging. There was a kind of core part of our innovation machine that we weren't able to fully bring to bear. But now it's different.
这就是为什么你看到我们的模型能力大幅起飞。如果你看看 5.6,当我们发布那个时,市场上最好的模型,你可以看到收入真的开始转变。我们在 7 月份的总收入增长了 20%。
And that is why you see our model capabilities taking off so much. And if you look at 5.6, when we launched that, best model in the market, you could see the revenue really starting to shift. We grew 20% in our topline revenue in July.
嗯。
Mh.
我们看到这种指数增长还在继续。
And we see that exponential continuing.
另一件随着模型能力而改变的事情是,模型在网络和黑客攻击方面变得极其强大。我觉得世界还在思考 Hugging Face 事件的影响。我的意思是,基本上发生了什么——如果我错了请纠正我——但似乎发生的是你们不小心黑进了另一家公司,而有一段时间你们自己都不知道,那是你们正在训练和评估的一个模型。而且你们后来——有一个黑帽视频,你们在那里讲解了这件事,你们还将在未来几天发布更多内容。但请带我回顾一下。当你听说这件事发生时,你当时怎么想?
Another thing that's changed, as you were saying with the model capability, is the models are becoming incredibly capable at cyber and at hacking. And I think the world is still kind of thinking through the implications of the Hugging Face incident. I mean, what essentially happened—correct me if I'm wrong—but it seems what happened is you all accidentally hacked another company without knowing it for a while, and it was a model that you all were training and evaluating. And you've since—there was a black hat video where you guys walked through this, and you're going to put more out in the coming days. But walk me through that. When you heard that that had happened, what did you think?
嗯,用我们安全团队某人的话说,这绝对是我们所知的、最让人大开眼界的安全事件。看待它的方式是,我们从中吸取了很多不同的教训。但我认为这是一个我们作为社区可以学习的时刻。我们可以理解未来的形态。对我来说,就我目前思考而言,突出的要点是,我们有一个模型,我们在评估它,它设法——它在一个配置良好的沙箱里——它设法既逃出了沙箱,然后又闯入了生产基础设施。
Well, in the words of someone on our security team, it was definitely the sort of most eye-opening security event that any of us were aware of. The way to think about it is that there are a lot of different lessons that we're taking away from it. But I think that it is a moment where we can as a community learn. We can understand what the shape of the future is. And to me, the salient points for the purpose of how I'm thinking about it right now is that we had a model, and we were evaluating it, and it managed to—it was in a sandbox that was well configured—and it managed to both break out of the sandbox and then also break into production infrastructure.
它本应是加固过的。
It was intended to be hardened.
所以那里的能力相当高。关于它究竟是如何运作的,有一些细节,而且它不仅仅是一个单一的智能体,而真的是一群智能体,专注于这个目标,并且能够互相传递消息。你开始看到的是,对于非常能干的智能体,人类攻击者有一些限制会拖慢我们的速度,但这些限制将不再能阻挡攻击者,对吧?他们可以浏览大量信息,通过许多不同的小漏洞,把它们串联成重大的东西。
And so the capability there is quite high. And there are some details about exactly how it all worked, and the fact that it was kind of not just a singular agent but really a swarm of agents that were focused on this goal and able to pass messages to each other. And the sort of thing that you start to see is that with very capable agents, there are limitations that human attackers have that slow us down that will no longer hold back the attackers, right? They can look through lots of information, through many different little vulnerabilities, and chain them together into something significant.
对我来说,我立刻开始意识到并思考的是,六个月后,当这种能力不再只存在于实验室、不只存在于少数地方,而是真正广泛可用时,会有什么影响。我认为答案是,每家公司都需要应对,需要升级其安全防护,因为攻击者将完全自动化。你甚至可以看到,随着 GLM 5.3 的发布,开放权重模型对所有人可用的网络能力将会增强。我认为重要的是,我实际上认为从长远来看,防御者可以获胜,AI 可以让防御者以前所未有的方式决定性地获胜,对吧?因为安全总是一场猫鼠游戏。攻击者升级,防御者也升级。但 AI 可以给防御者带来不公平的优势。例如,软件的形式化验证,你可以真正证明它是安全的,数学上证明安全属性。
For me, what I immediately started to realize and think about is what are the implications for six months from now when this kind of capability isn't just in the lab, isn't just in a few places, but is really widely available. And I think that the answer is that every company needs to respond, needs to upgrade its security because attackers are going to fully automate and you can even see with the release of GLM 5.3 that the cyber capabilities of openweight models available to everyone are going to increase. And I think that the important thing is that I actually think in the long run that defenders can win, that AI can actually make it so defenders can decisively win in a way that was never possible before, right? Because security is always a cat-and-mouse game. Attackers up level, defenders up level. But AI can give defenders an unfair advantage. For example, formal verification of software where you can actually prove that it is secure, mathematically prove security properties.
这项技术是存在的,但要让人类真正去完成证明这些属性的工作,是难以处理的。
The technology for this exists, but it is intractable for humans to actually do the work to prove those properties.
而且我们有这些在数学上表现出色的模型,对吧?例如,Astra,我们即将推出的模型,解决了——我们几周前宣布的——解决了 10 道数学题。而这仅用了价值 2000 美元的算力,这些问题已经开放了几十年,对吧?数学家们毕生致力于这些问题,现在我们有模型能证明答案。这些模型绝对能证明软件的形式化安全属性。
And we have these models that are incredible at math, right? They are solving, for example, Astra, our upcoming model, solved—we announced this a few weeks ago—solved 10 math problems. And this was with $2,000 worth of compute that had been open for decades, right? Mathematicians that spent their whole careers on these problems, and now we have a model that can prove the answers. Those models definitely can prove formal safety properties of software.
所以这就是防御者获得不公平优势的一个方面。
And so that's one place where unfair advantage for defenders.
我们正在做的第二件事是,我们真的在投资生成安全代码。
A second thing that we're doing is we're really investing in generating secure code.
然后你开始意识到,作为人类开发者,有很多模式人们总是搞错。
And you start to realize that as a human developer, there are many patterns that people just keep getting wrong.
我在以前的工作中多次经历过这种情况。我非常专注于构建安全的基础设施,然后发现人们做的所有这些小错误。SQL 注入人们大概知道有一些小模式,但也有很多小的反模式,比如甚至……
And I've been through this many times in my previous jobs. I was very focused on building secure infrastructure and just find like all these little things people do wrong. SQL injections people kind of know about that there's little patterns, but there's lots of little anti-patterns as well for even like for example...
我记得在 Ruby 中,正则表达式里有些字符,如果你弄错了,就会有非常常见的攻击,人们可以放入换行符而不匹配。所以你能向载荷中添加内容。
I remember it used to be like in Ruby that the regular expressions there's like certain characters that if you get them wrong that there's just very common attacks where people can put a new line and it doesn't match. So you're able to add things to the payload.
所有这些小的安全反模式问题,你总是不得不告诉你的员工:不要这样做,不要这样做,不要这样做。你可以训练你的 AI 永远不这样做。所以我认为我们实际上有能力打造世界上最好的安全工程师,从根本上说,防御者可以拥有优势。但这需要思维模式的转变。这需要一个大型项目。这需要整个生态系统改变软件安全的方式,因为我们学到的第二件事是,随着这些网络模型能力的增强,它们在关键软件(例如 Linux 内核)中发现的漏洞数量不断增加。这真的让你意识到,人类编写的软件,是的,响应太慢了。
All of those problems of just little security anti-patterns that you just keep having to tell your people don't do this, don't do this, don't do this. You can train your AI to never do it. And so I think that we actually have the ability to build the best security engineers in the world, and fundamentally the defenders can have the advantage. But this is going to require a mindset shift. This is going to require a mega project. This is something that the whole ecosystem needs to change how software is secured because there's a second thing that we've learned, which is that as these cyber models have gotten more capable, the number of bugs that they are finding in critical software, for example the Linux kernel, they just keep finding more and more. And it really makes you realize that human-written software, yeah, that the response is too slow.
是的。所以我们需要防御者自动化,以机器速度行动,因为他们的对手肯定会这样。
Yeah. And so we need defenders to automate, to move at machine speed because their adversaries certainly will.
你会对那些可能——而且他们已经对 AI 有负面看法,也许他们附近有一个不喜欢的数据中心什么的——听到你这么说然后说,你只是在推销你制造的问题的答案,就像你在构建这些模型,现在你要卖东西来保护我免受它们的伤害。你会怎么回应?
What do you say to people who maybe—and also they already have a negative kind of disposition of AI, maybe there's a data center they don't like in their neighborhood or something—and they hear you say that and go, you're just selling the answer to the problem you've created in the sense of like you're building these models and now you're going to sell something to like protect me from them. What would you say to that?
嗯,首先,我对这种反应深表同情。我看待这件事的方式是,我回想到 10 年前、11 年前,在我们创办这家公司之前,我并不在 AI 领域。我想进入 AI 领域的原因是我觉得这可能是人类有史以来最赋能的科技。AI 可以成为帮助每个人过上他们想要的生活、构建他们想要的生活、实现他们目标的东西。我只是想帮助它朝着比没有我时稍微更积极的方向发展。所以我开始是因为我觉得这项技术可以提升每个人。当你看到 OpenAI 的现状和我们正在做的事情时,我们可以稍微看到一点未来,对吧?我们正在创造的不只是我们自己,对吧?有整个领域、整个生态系统、许多参与者,这本身就是一件好事。我们相信这项技术应该民主化。它不应该被锁在一个地方,由一个实体控制,因为那是另一种风险。但这意味着我们确实需要适应,我们需要驾驭,我们需要利用我们费力从未来提取的信息来看清它的形态。所以对我来说,Hugging Face 是一个证明点、一个案例研究、一扇通往未来的窗口。它不一定是由某一家公司创造的未来。它是整体经济、所有这些不同国家中构建这类模型的整体参与者集合,它向你展示了什么是可能的,以及攻击者将如何升级。它没有展示的是防御者如何升级?但对我来说,这也是我们现在试图展示的,我们可以构建这些模型,你可以用它们来扫描自己。例如……
Well, first of all, I have a lot of empathy for that reaction. The way that I look at it is I rewind to 10 years ago, 11 years ago, before we started this company, I wasn't in AI. The reason I wanted to be in AI is because I feel like this can be the most empowering technology that has ever been created. That AI can be something that helps everyone live the life that they want, build the life that they want, achieve their goals. And I just wanted to help steer it in a slightly more positive direction than it would be without me. So I started because I feel like this technology can uplift everyone. And when you look at where OpenAI is and what we're doing, we can see a little bit into the future, right? That what we are creating is not just us, right? That there's a whole field, there's a whole ecosystem, there are many actors, and that is itself a good thing. That we believe that there should be democratization of this technology. It should not be locked up in one place with one entity controlling it because that's a different kind of risk. But it means that we do need to adapt, we do need to navigate through, and we need to use the information that we have sort of painstakingly extracted from the future to see what is the shape of it. So to me, Hugging Face is a sort of proof point, a case study, a window into the future. It's not necessarily a future that any one company is creating. It's the overall economy, the overall set of actors in all these different countries who are building these kinds of models, and it shows you what will be possible and how attackers will be able to update. What it doesn't show you is how do defenders update? But that to me is what we are also trying to now show as well, that we can build these models that you can use to scan yourself. As an example...
我问了 Codex,这只是现成的 Codex,没有特殊权限,没有特殊模型。我只是用我的个人 Codex 说……
I asked Codex and this is just off-the-shelf Codex, no special access, no special models. I just use my personal Codex to say...
扫描 gregarman.com。告诉我那里有什么漏洞。
Scan gregarman.com. Tell me what vulnerabilities are there.
而 gregarman.com 非常简单。它只是一个静态网站。我当时想,得了吧。它能找到什么?结果它回来了 13 个发现。每一个我都想,你知道,我有点知道这个,但修复它太费劲了。然后想,哦对,我有点忘了那个还开着。或者,你知道,这里有个域名还是遗留的。我不再用了。我没有设置防止人们伪造你电子邮件地址的记录。诸如此类的一堆小事。
And gregarman.com is very simple. It's just a static site. I was like, come on. What's it going to find? And it came back with like 13 findings. And each one of them I was like, you know, I kind of knew about this, but it's so much work to fix it. And like, oh yeah, I kind of forgot that I left that open. Or, you know, here's some domain that was still, you know, was legacy. I'm not using it anymore. I hadn't set the records that prevent people from spoofing your email address. A bunch of little things like that.
那一个看起来挺严重的。
That one seems like a big deal.
嗯,是的,你说得对。既然你这么说。现在,事情是这样的。老实说,要修复这类问题,你必须设置像 DMARC 这样的东西才能做好。而且有一个完整的过程你必须去部署。我只是觉得,我不知道,这很费时间。需要很多专业知识。所以相反,我只是问 Codex,你真的能修复吗?
Well, yeah, you're right. Now that you say it. Now, here's the thing. Honestly, to fix those kinds of things, you have to set up like DMARC to do it right. And there's this whole process you have to do to roll. I'm just like, I don't know, it just takes time. It takes a lot of expertise. So instead I just asked Codex, can you fix it really?
然后它打开我的云控制面板,它只是使用了计算机使用功能。它点击各处,设置了一堆头部,更改了一堆设置。
And so it opened up my cloud control panel and it just used computer use. It clicked around, set a bunch of headers, changed a bunch of settings.
我以前是 Cloudflare 在前面,AWS 在后端。这里我走的是不安全的 HTTP。Codex 发现了这一点。我其实都不确定它是怎么看到的,因为我觉得一切都藏在 Cloudflare 后面。它实际上就是,你知道吗,我干脆把你从 AWS 迁移走。于是它设置了 Cloudflare Pages,把内容搬了过去。我有个老旧的 jQuery 版本,不安全。它干脆把 jQuery 整个去掉了。找出所有这些发现只花了大约 15 分钟。又花了一个小时修复全部问题。对我来说,我只是个管理者。我只是给了它意图。我说:“我想要的是知道问题出在哪,然后让你去修复。”它就这么做了。在我看来,那才是我们被承诺的未来。
I used to have Cloudflare in front and AWS on the back end. I was going over insecure HTTP here. Codex had figured that out. I'm actually not even sure how it saw that because I thought everything was hidden behind Cloudflare. It actually was just like, you know what, I'm just going to migrate you off of AWS. So it set up Cloudflare Pages, moved over the content. I had an old jQuery version that was insecure. It just got rid of jQuery entirely. And it was like 15 minutes to find all these findings. It was another hour to fix all of them. And for me, I was just the manager. I just gave it the intent. I said, "What I want is I want to know where the problems are and I want you to fix them." And it did it. And that to me is more the future that we're promised.
我想听听,我知道这更多属于研究领域,但我认为它与一切息息相关。你谈到你们现在针对 Hugging Face 事件在做的事。你的同事 Mia 前几天告诉我:“我希望我们在这件事发生之前就做了很多我们现在在做的工作。”那么,现在正在做的工作是什么?
I want to hear, and I know this is more in the research world, but I think it's really relevant to everything. You talk about what you guys are doing now in response to Hugging Face. One of your colleagues, Mia, told me the other day that, "I wish we had done a lot of the work we're doing before this happened." So what is that work that's happening now?
所以,我认为自 Hugging Face 事件发现以来,我们一直处于的模式就是全面的事件响应,以提升我们保护、监控和对齐最强大模型的每一个方面。
So I would view the mode we've been in since the discovery of the Hugging Face incident as full incident response to uplevel every aspect of how we secure, monitor, and align our most capable models.
嗯。
Yeah.
我们相信,超过一定能力水平的模型需要这样的措施。
And we believe that models that are above a certain capability level require this.
而且这是一个非常有趣的技术问题,值得思考。
And it's been a very interesting technical problem to think about.
而且这并不全是我们措手不及,因为我的意思是,我们预料到这一刻已经有一段时间了。如果你回顾二月份,我们公开开始讨论网络模型的受信访问,对吧?那是我们第一次真正思考,好吧,在我们的准备框架中,当网络能力达到一定水平时,你需要区别对待它们。现在我们开始思考这样一个事实:这些模型将在世界上产生如此大的变革。我们需要思考如何在内部控制访问。我们开始思考,如果我们有非常强大的模型能够突破沙箱,那么我们需要确保我们的沙箱尽可能达到工业级,尽可能万无一失,这实际上意味着要从头开始围绕安全属性进行工程构建。所以我们把一些,实际上,Mike Dalton,他是那次 Black Hat 演讲的人之一,他开始承担的任务是构建最安全的沙箱。所以我们已经为这一刻准备了一系列基础设施原语和政策原语,但当它真正发生时,情况也不同了,对吧?你永远不知道什么时候会到达那个时刻。我认为在这里我们实际上非常,我认为我们非常具体地看到了这个窗口:好吧,我们现在相对于这些部件已经处于一个严重的能力水平。我们必须立即行动。
And it's not all us being caught flatfooted, because I mean, we anticipated this moment for quite some time. If you look back at February, publicly we started talking about trusted access for cyber models, right? That was the first time we were really thinking about, okay, we've always had in our preparedness framework, when cyber capabilities get to a certain level, you need to treat them differently. Now we're starting to think about the fact that these are going to be so transformative in the world. We need to think about how access is controlled internally. We started to think about, if we have very capable models that are able to break out of sandboxes, well, we need to make sure our sandbox is as industrial-grade as, you know, sort of foolproof as humanly possible, which really means engineering from the ground up around security properties. And so we put some of our, actually, Mike Dalton, who was one of the people who gave that Black Hat talk, the task that he started to work on was build the most secure possible sandbox. So we've been putting in place a bunch of infrastructure primitives and policy primitives for this moment, but it's also different when it actually hits, right? You never know when you're going to hit that moment. And I think here we're actually very, I think we got to see this window in a very concrete way of, okay, we're now at a serious capability level relative to the pieces. We must act now.
嗯。
Mhm.
所以,工作流程的转变一直在非常迅速地发生,因为你开始思考,我们的很多治理都集中在模型的部署上,但不一定是模型的训练。现在我们需要在训练和评估过程中把治理拉回来。我们在那里也需要良好的治理。我们需要围绕制衡制定标准:我们怎么知道它足够安全?我们怎么知道它足够对齐?而对齐这部分实际上非常重要,因为我们的信念是,如果你让模型变得极其强大,你真的要确保随着能力的增长,它们的对齐也随之增长,对吧?你希望这些东西以协调的方式同步发展。因此,通过这种方式,你有了纵深防御:模型试图做正确的事,沙箱尽可能安全。而且我认为,以迭代的方式来做这件事,采取渐进式的跳跃,这样你可以在前进的过程中理解并确保你的系统在正常工作,所有这些都与我们思考迭代部署、思考从一开始就创造这项技术的方式非常一致。所以我们正在将其付诸实践。实际上,对我来说,作为一个管理者,作为一个思考生产这项技术的组织系统的人,这也是一个非常重要的挑战,因为这意味着你现在需要不同的人参与决策。你需要一个不同的流程:你怎么知道需求是什么?你如何根据这些优先级执行?如果你的技术理解发生了变化,你如何传播它?你如何确保人们在正确的房间里?但同样,在某些方面,这是一个阶段性的转变,但这是我们一直在准备的转变。比如今年早些时候,我花了很多精力真正尝试把研究和安全结合起来,开始进行,例如,我们开始每周举行智能体安全与保障会议。其中一些事情是关于建立连接组织,让那些通常不会互相交谈的人——因为他们的世界有点不同,他们来自不同的背景,他们专注于不同的问题——被迫互动,被迫思考共同的问题。而很多这种共同的同理心、共同的理解、共同的技术背景,现在正在产生回报。
And so there's been a workflow shift that's been happening very rapidly, because you start to think about a lot of our governance has been focused on deployment of model, but not necessarily the training of model. Now we need to pull the governance back in during the training and evaluation process. We also need good governance there. We need to have standards around checks and balances around how do we know if it's secure enough? How do we know if it's aligned enough? And this alignment piece is actually very important, because our belief is that if you make models extremely capable, you really want to make sure that as the capability grows, that their alignment grows with them, right? That you want those things to be in lockstep and sort of coordinated fashion. And so in that way, you kind of have this defense in depth of models that are trying to do the right thing, that you have sandbox that's as secure as possible. And I think that doing this in an iterative fashion, so that you're taking incremental jumps so you can kind of understand as you go and make sure that your systems are working, all of that is very aligned with how we've thought about iterative deployment, how we've thought about the creation of this technology from the very beginning. So we are operationalizing that. And actually, for me, as a manager, as a sort of thinking about the organizational system for producing this technology, it's also been a very important challenge, because this means you now need different people in the room to make decisions. You need a different process for how do you know what the requirements are? How do you execute against those priorities? If there's a change in your technical understanding, how do you propagate that? How do you make sure people are in the right rooms? But again, in some ways, it's been a phase shift, but it's been one we've been preparing for. Like earlier this year, I spent a lot of effort to really try to bring together research and security to start having, we started doing, for example, these weekly safety and security of agents meetings. And some of these things are about building the connective tissue, have people who wouldn't normally talk to each other because their worlds are a little bit different, they come from different backgrounds, they're focused on different problems, be forced to interact, be forced to think about shared problems. And a lot of that shared empathy, that shared understanding, that shared technical context is paying dividends right now.
但与此同时,这样做的含义是,新模型的推进速度将显著放缓。你感受到影响了吗?
At the same time though, the implications of this are that momentum on new models is going to significantly slow. Do you have a sense of the impact?
嗯,我会换个角度说,因为我一直用这个视角来看:你怎么看待进步?你怎么看待这个领域的进步?如果你只把进步看作你在基准测试上的分数,你在特定任务上的整体能力水平,那就太短视了。我们正在构建一个集成系统,而安全与保障是它的核心属性。
Well, I would put this differently, because I've always taken this lens to how do you think about progress? How do you think about progress in this field? And if you only think about progress as what's your score on a benchmark, what's your overall capability level in this specific task, that's too myopic. We are building an integrated system, and safety and security are a core property of that.
所以,如果你的安全和保障相对于能力滞后了,是的,你绝对需要改进你的安全和保障。
And so if you have safety and security lagging relative to the capability, yeah, you absolutely need to improve your safety and security.
但我们的目标是拥有一个我们全力构建的系统,将安全和保障作为一等要求。所以,我思考这个问题的方式是,我们正在提升机器的每一个方面。过去一年、过去两年,我们一直在提升我们构建集群的方式,对吧?我们现在能够以极其庞大的规模进行训练。我们能够让整个预训练机器、强化学习机器,所有这些部分协同工作。我们正在构建最好的沙箱。我们正在为所有这些添加思维链监控。
But our goal is to have a system that we are building at full force on safety and security as first-class requirements. And so the way I would think about this is that we are upleveling every aspect of the machine. We have spent the past year, past two years upleveling how we build our clusters, right? We now are able to train at massive, massive scale. We're able to get the whole pre-training machine, the RL machine, all of these pieces to work together in concert. We are building the best sandbox. We are adding chain-of-thought monitoring to all of this.
对我来说,这只是在问:交付的范围是什么?成功是什么样?好的标准是什么?所以我的想法是,是的,如果你放弃一些要求,你可以在其他要求上进展更快。但那没有意义。那从来都没有意义。那不是工程运作的方式。所以对我来说,我们是在执行这个计划,我们过去已经反复思考过很多次,你可以从我们的许多文件中看到这一点,从章程到预备框架,以及介于两者之间的许多其他内容,那就是随着模型能力越来越强,你必须确保同步提高安全和安保要求,而这正是我们在做的。
To me, it's just saying what is the scope of delivery? What does success look like? What does good look like? And so the way I would think about it is that yeah, if you jettison some requirements, you can move faster in other requirements. But that doesn't make any sense. That never makes any sense. That's not how engineering works. And so to me it's we are pursuing the plan like we've really thought about this many times over the past and you can see it in many of our writings from the charter to the preparedness framework and many other things in between that as models get more capable you have to make sure that you increase your safety and security requirements in concert and that's exactly what we're doing.
但我的意思是,实际上,是的,我听到了所有这些,但这似乎会导致模型的发布节奏变长。
But I mean practically yes I hear all that but it does seem like what this will result in is the release cadence of models to lengthen.
老实说,我不确定这是否会成真。所以我认为会发生的是,如果你只是最大化某个特定指标,你可能会比我们追求的速度更快地最大化那个指标。但有很多自由度,嘿,我们总是有一堆胜利排队等着,我们可以在特定轴向上放弃其中一些胜利。所以如果你在推动特定的能力,比如网络、生物、编码或任何这些方面,也许你不会把所有特定的事情都耦合在一起。或者如果有架构改进,也许你会延迟将这些加入运行中,但你产生检查点的速率不需要改变。只是那些检查点所代表的内容可能在特定轴上能力较弱,而安全和安保更紧密地交织在一起。但我几乎会把检查点和我们发布模型看作更像是给我们所处的位置拍一张照片,对吧?并不是说有一组预先注定必须创建的模型。我们可以共同设计。而且我认为在实践中,我们实际上有很多理由倾向于进行非常渐进的发布,就像以合理的频率拍摄这些照片一样。
I'm not certain if that will be true to be honest. So I think that what will happen is it does mean that if you were just maxing a particular metric you could probably max that metric faster than what we will pursue. But there's a lot of degrees of freedom of hey we have we always have a bunch of wins queued up and that we can jettison some of those wins right on particular axis. So if you're pushing on particular, you know, capabilities on cyber or on bio or on coding or any of these things, maybe you don't couple together all of the particular things. Or if there's architectural improvements, maybe you delay adding those to the run, but your rate of producing checkpoints doesn't need to change. It's just kind of the what those checkpoints represent could be less capable and specific axes more safety and security intertwined. But I would almost think of checkpoints and our shipping of models as a little bit more about taking a photograph of where we happen to be, right? It's not like there's like a pre-ordained set of models that have to be created. It's we can co-design. And I think that in practice, I think that we actually have a lot of reasons that we prefer to do very incremental releases with these, you know, photographs taken with with a reasonable frequency.
我想我想说的是,有一段时间,发布节奏对每个人来说都相对稳定。特别是在过去一年,感觉它加快了。你知道,基本上每六周就有新模型出来,你和 Anthropic 不断推出新模型。嗯,而你所描述的,这种评估现状并在过程中引入更多摩擦的做法,确实感觉这不可避免地会减缓未来的势头。我很好奇,你是从我们一直在谈论的这台机器的角度来思考这个问题,即你现在负责的业务及其影响,还是说现在知道还为时过早?
I guess what I'm getting at is that for a while, the release cadence felt relatively in place for everyone. In the last year, especially, it feels like it sped up. You know, new models coming out every six weeks essentially, you and Anthropic putting out new models constantly. Um, and what you're describing, this kind of taking stock and and introducing more friction into the process, it does feel like that just inevitably has to slow future momentum. And I'm I'm curious if you're thinking about that from the perspective of this machine we've been talking about, the business that you're now accountable for and the impacts on that or is it just is it too early to know?
听着,我只是觉得我理解那个观点,而且我认为这是一个非常合理的观点,但我觉得我观察到的我们的运作方式是,我们共同设计一切,对吧?所以,业务运作的方式肯定是,我们会把我们准备好发布的最强大的模型拿出来,对吧?我们已经做了适当的安全和安保控制,然后我们会把它带给我们的客户,帮助他们获得他们期望的投资回报率,帮助他们真正转型他们的业务,获得价值。在很多方面都很简单。然后问题是什么样的模型准备好发布,比如我们内部生产许多模型,而设置多高的门槛、多低的门槛,在能力跳跃方面,所有这些都必须共同设计,频率也是如此。嗯,我们有早期访问计划,比如 Daybreak Red 和 Blue 这样的项目,让不同的模型可以提供给不同的合作伙伴。所以我想我的想法是,这里肯定有一个转变,在我们内部努力平衡、检查和制衡的位置以及这些东西如何运作方面。我不认为有一个明确的涓滴效应,比如模型节奏会从 6 周变成 6 个月,但那些每六周发布的内容会有所不同。
Look, I just I think that like I understand that perspective and I think it's a very reasonable one, but I think that my observation of how we operate is that we co-design everything, right? So, it's like the way in which the business operates for sure is we will take the most capable model we are ready to release, right? We've done the appropriate safety and security controls for and we will bring that to our customers and help them get the ROI that they're looking for, help them really transform their businesses, get value. Simple in a lot of ways. And then the question of well what models are ready to release like we have many models that we produce internally and the choice of how high a bar to set how low a bar to set in terms of the capability jumps all of that has to be co-designed the frequency of it um that we have early access programs there's programs like daybreak red and blue which let us have different models that go out to to you know different partners and so I guess the way I would think about it is that that there is for sure a shift here in terms of our internal balance of of effort and where the checks checks and balances are and how how those things operate. I don't think that there is a clear trickle down to just like okay the model cadence will go from 6 weeks to 6 months but the what those six weekly releases actually contain will be different.
所以也许前沿,那种原始的前沿,不会那么快被拉入生产,进入商业世界,即使它仍在被训练和研究。
So maybe the frontier the raw kind of frontier doesn't get pulled into production into the commercial world quite as fast even though it's still being trained on and worked on.
嗯,我会把这看作是一个整体的节奏,对吧?一个前沿回滚的整体节奏。另一件要考虑的事情是,我们生产这些模型并获得这些能力提升的能力也在增加,对吧?我们的原始能力,部分是因为我们比过去更了解这项技术,而且从安全和对齐的角度来看,这实际上也非常有益,对吧?我们确实对技术的运作方式了解很多。仍有很多很多部分我们不了解,但我们的工程能力以及在技术上真正做好工作的能力,嗯,我认为已经大大提高了。而且有些过去真正阻碍我们的问题已经大大减少了,例如这些大型训练运行,你知道它们运行数月,所以通常我们会从代码的一个分支启动它们,然后我们会添加针对该运行的小修复,同时可以继续主线开发,时不时地主线上有我们想要的东西,比如‘哦,我们真的想要那个更改,我们真的想要那个修复。’然后我们说,‘好吧,我们必须把这两个开发分支合并在一起。’然后我们启动运行,发现有一个 bug,因为没有人真正在大规模上测试过所有这些主线修复,因为我们只有一个实例可以这样做。然后我们花接下来的三到四周时间,让我们所有最优秀的人仔细搜索,试图找出 bug 来自哪里,因为你不可能阅读所有已编写的代码,对吧?没有人能做到那一点。
Well I would think of this as an overall pacing, right? An overall pacing of that frontier roll back. And another thing to think about is the fact that our ability to produce these models is also and to get those capability gains is also increasing, right? Our raw capacity is because partly we understand this technology a lot better than we did in the past and that's actually very beneficial from a safety and alignment perspective too, right? That we really understand a lot of how the technology works. There's still many many pieces that we don't understand but our ability to engineer it and to really do a good job of it technically um I think has increased a lot um and that there are certain problems that really used to hold us back that are much diminished for example these big training runs you know they run for many months and so typically we'd launch them from a branch of code and so we'd be adding little fixes to that that are specific to that run at the same times can continue their mainline development and every so often there's something on the mainline that we're like oh we really want that change, we really want that fix. And then we're like, all right, we got to bring these two branches of development together and we merge them. And then we launch the run and we find that there's a bug because no one has actually tested all these mainline fixes at the massive scale because we only have one instance where we can do it. And then we spend the next three to four weeks with like all of our best people scouring trying to figure out where did the bug come from because you can't possibly read all of the code that's been written, right? No human could possibly do that.
今天你会遇到同样类型的问题。Codex 只是读取所有代码,对吧?它找到 bug。它说在这里,没问题。
Today you run into that same type of issue. Codex just reads all the code, right? It finds the bug. It says here it is no problem.
所以,这种模型能力的放大带来的伙伴关系,能够消除整类问题,我们知道那需要三到四周,现在只需要几个小时。我们已经多次具体看到这一点。所以我想说的是,过去我们不得不发展的方式有很多限制,这使得我们在重要问题上更难前进。现在我们有机会以更高效的方式前进。这对我来说就像是,让我们利用那额外的能量来增强对齐属性,增强安全属性。
And so there's something about the partnership of this amplification of what the models can do that knock out whole classes of problems where we know it would take three to four weeks and now it takes a few hours. We've seen this concretely multiple times. And so what I'm getting at is that there are a lot of limitations for how we've had to develop in the past that have made it a lot harder for us to move forward on problems that are important. And now we have the opportunity to move forward on them in a much more efficient fashion. And that that to me is like let's use that extra juice to increase the alignment properties, to increase the safety properties.
我们要确保能够以可信赖的方式让模型监督其他模型。让我们真正落实 2017 年讨论过的那些想法,对吧?比如迭代放大、辩论等方法,让模型真正投入大量算力,以确保它们的行为确实合理。在我看来,这就是我们面前的机会。
Let's make sure that we can have models that monitor other models in a trustworthy fashion. Let's really implement some of these ideas that we were talking about back in 2017, right? That we have things like iterated amplification, things like debate where the models can really spend a bunch of compute in order to make sure that what they're doing is actually reasonable. To me, that's the opportunity in front of us.
你所说的这些变化会影响 Astra 的发布吗?就是那个即将推出的东西。你觉得它会比你们最初预期的更晚推出吗?
Does this these changes you're talking about impact the release of Astra, the thing that is coming sooner at all? Do you think that's going to be coming later than you guys originally hoped?
嗯,某种程度上它已经比我们希望的晚了。是的,没错。我们有一个我们认为非常好的模型,很多人在内部使用,但还没有对外发布。是的。
Um well it is already later than we hoped in some ways. Yeah. Right. So we've had we've had a model that we think is very good. Lots of people using it internally. It's not out there yet. Yeah.
所以我认为答案是肯定的。我们实际上一直,你知道,尽管我刚才说了那么多,我们确实牺牲了短期的最大化,比如如果你只是想,是的。
And so I think that the answer is yes. We have actually been really like you know despite everything I just said we have very much sacrificed the shortterm maximalist like if you just wanted to Yeah.
你知道,专注于收入之类的。
You know focus on revenue or something like that.
我们根本不是那样想的。甚至没有人问过那个问题。对。我们一直很……
That isn't how we've thought about it at all. No one has even as asked that question. Right. We've been really
是的,是的。因为很明显,我们知道我们的优先事项是什么。是的。
Yeah. Yeah. because it's just very clear like we know what our priorities are. Yeah.
比如我们的首要任务是 AGI,做好 AGI,第二优先是改造经济。所以它确实在优先列表上。收入绝对是我们的优先事项。但我们知道相对的优先级。我们知道如果有必须做出的直接选择,什么会胜出。
Like it is our priorities are AGI do right by AGI and our priority two is transform the economy. So it is on the priority list. Revenue is absolutely a priority for us. But we know the relative prioritization. We know what wins if there's a you know direct choice that has to be made.
你认为你们已经拥有 AGI 了吗?
Do you think you have AGI?
我经常思考这个问题。我和不同的人交流,了解他们认为我们处于什么阶段。我从团队那里得到的最好的答案,也是让我真正共鸣的,是 AGI 是一个渐进的过程,对吧?它并不是一个时间点,你知道。
I think about this a lot. I talk to I talk to different people about where they believe we're at. And the best answer I've gotten from our team and I it really resonates for me is that AGI is a gradual thing, right? It's not really like a moment in time like a you know
你曾经把它说成一个时间点。公司确实这么说过。
you talked about it like a moment in time at one point. The company definitely did.
没错。我们过去是这么想的,对吧?我们刚开始时几乎认为它什么都不是,然后我们会把各个部分拼凑起来,然后就有了 AGI。
That's right. That's what we used to think, right? We we really when we set out we almost thought of it as nothing and then we'll put together the pieces and then they have AGI
我认为这些年来我们学到的是,那是一种天真的观点。你可以从证据中看到,这些智能体的创造在很多方面是多么平滑,软件工程的转变已经发生,现在知识工作的转变普遍如此,网络能力等等,这些都是逐渐增强的。就像有些事情,你会意识到发生了什么,Hugging Face 就是一个让你意识到自己处境的时刻,但一切都是渐进的。比如回顾过去发布的不同模型以及它们在不同基准上的表现,你知道,就是那种指数曲线。所以话虽如此,答案的其余部分是,如果我们向前看两年再回头看,我认为我们会说这就是 AGI 被创造的时候。所以我认为我们正在经历那个时刻,
and I think what we've learned over the years is that was a naive perspective and you can see this in the evidence of like how smooth in many ways the creation of these agents has been how the transformation of software engineering has been now the transformation of knowledge work generally how cyber capabilities all these things they are gradually increasingly it's like one of these things where there are moments where you realize what has happened hugging face is a moment where you realize where you are, but it's all gradual. Like if you look back at kind of these different models that have been released over time and their performance on different benchmarks, it's just, you know, sort of that that exponential. And so I think that being said, the rest of the answer is that if you go if we go forward two years and we look back, I think that that we will say this is when AGI was created. So I think we are going through that moment,
但它不是你现在就能简单说出来的东西。我不认为它会像说“就是这一天”那么简单。有太多的维度。我认为整个 AGI 对话的一个普遍教训,以及为什么它如此复杂,是因为我认为人类的能力远超天真的观点所认为的。
but it's not it's not something where you could just say it now. I don't think that it's ever going to be as simple as you can say this was the date. There's so many axes. And I think that maybe one lesson in general to the whole AGI conversation and why it's so complicated is because humans I think are capable of much more than a naive view would give credit to.
我认为我们所做事情的多样性、我们如何设定目标、如何实现目标、任务如何运作、工作如何运作,所有这些都非常深刻。所以我认为我们看到的是,人们完成任务的方式或人类难以克服的局限性有不同的方面,比如我提到的阅读所有代码,这些 AI 能够填补这些空白,能够做越来越多的事情,而事实证明,有些事情我们确实觉得必须有一个人在场非常重要。比如最终判断来自哪里?品味来自哪里?谁决定什么是好的什么是不好的?那应该是人类独有的。我给你举个例子。我们在图像生成方面看到的一件事是,当你用 ChatGPT 生成一张图像并展示给别人时,没人关心。除了你之外,它对任何人都无趣。但如果你拍一张全家福,然后说“把它变成漫画风格,把它变成卡通风格”,人们会喜欢。它会病毒式传播。所以人性中有些元素至关重要。人们如此渴望与他人联系,这实际上在我看来是这项技术的目的,我认为真正涌现的是加深人类联系,这样你可以花更多时间和家人在一起,花更多时间和你想要的人在一起,减少花在机器上打字的时间,减少得腕管综合征的时间等等。所以我认为 AGI 在展开过程中,几乎就像雾正在散去,它几乎在教我们更多关于我们自己和作为人类意味着什么。
I think that the diversity of what we do, how we set goals, how we accomplish goals, how tasks work, how jobs work, all these things extremely deep. And so I think that what we're seeing is that there's different aspects of either how people accomplish tasks or limitations that are hard for humans to overcome like the thing I mentioned with read all the code where these AIs are becoming able to fill in those gaps are able to do more and more and that humans it turns out that there are some things that we also really find are extremely important for it to be a human there. Like where does the judgment come from ultimately? Where does the taste come from? who decides what's good and what's not good. Like that is something that should be uniquely human. And I'll tell you just to give you an example. One thing we've seen with image generation, when you just generate an image with cashbt and you show it to someone, no one cares. It is uninteresting to anyone besides you. But if you take a photo of your family and you say, "Turn this into the into a caricature, turn this into, you know, a cartoon style," people love it. Goes viral. And so there's something about that element of humanity that matters. so deeply people want to connect with other people and that is actually in in my mind the purpose of this technology a thing that I think is is really emerging is for deepening human connection so you can spend more time with your family you can spend more time with people you want to spend time with less time you know typing away at a machine less time you know getting carpal tunnel the whole thing and so I think that there is something about what AGI is as it's starting to unfold it's almost this like fog that's lifting it's almost teaching us more about ourselves and what it means to be human.
鉴于你们内部拥有的能力以及你们现在采取的措施,将算力转移到对齐上并设置所有这些检查点,嗯,你们怎么知道什么时候达到了良好状态,解决了问题,模型已经对齐了呢?
Given the capabilities that you guys have internally and the measures you're taking now, moving compute to alignment and putting in all these checkpoints, um, how are you going to know when you're in a good spot there and and that you've solved the problems and the models are aligned?
所以我认为这里的视角非常关键,这也是我们多年来花了很多时间思考、辩论,真正试图完善一个视角。在我看来,这归结为迭代。归结为迭代部署、迭代开发。所以你需要有制衡机制。你需要有评估。你需要有标准。你需要所有这些。你需要有治理。你需要有一个技术上健全的设置。思考方式是,一切都是渐进的。一切都是增量的,你学习、更新、迭代。随着模型能力的提升,这些边界的紧密度必须相应增加。所以可以把它想象成发射火箭,对吧?如果火箭很小,你可能犯很多错误也没关系,对吧?你只要发射它,你知道,一切都好。火箭更大,容差就小得多。你知道,从几厘米到几毫米。那就是你必须达到的水平。
So I think the perspective here is is very critical and this is something we've again spent lots of time over the years thinking about debating really trying to refine a perspective. In my view it really comes down to iteration. It comes down to iterative deployment iterative development. So you need to have checks and balances. You need to have evaluations. You need to have bars. You need to have all of these things. You need to have governance. You need to have a setup that is technically sound. And the way to think about it is everything is gradual. Everything is incremental and you learn and you update, you iterate. And that as the model capabilities increase, the tightness of those bounds has to increase correspondingly. And so think of it a little bit like if you're launching a rocket, right? If the rocket's very small, like there's a lot of things you can probably do wrong and it's fine, right? You just kind of like launch it and like, you know, it's it's all good. Rocket bigger, the tolerance is much more tiny. you know, you go from, you know, multiple centimeters to m multiple millimeters. That's where where you have to be.
当你开始为火箭的有效载荷、它的去向、它的任务赋予更多责任时,你会想到所有依赖那枚火箭的事物,以及所有必须正常运转的系统、所有可能出错的地方。你必须不断提高自己衡量的能力,确保自己不仅能完成工程,还能满足安全与安保属性。这正是我们的思考方式:有些时刻我们会引入非常剧烈的变化。我们首次制定预备框架时就是一个例子。但即便是预备框架,我们也非常谨慎、刻意地内置了修改它的能力。我们说会有新信息,我们必须学习,必须更新。但在任何时刻,我们都希望明确我们的标准是什么。当模型达到某种能力水平时,我们如何处理?我们最近谈到 Astra 正在接近网络关键水平。它不一定已经达到,因为要真正证明它需要尚未完全定义的评估,但我们要说的是,好吧,它开始接近了,所以我们就当它已经达到来处理。这意味着我们需要去执行某些保障措施和制衡机制。你们看到这一切都在发生。所以我认为对你问题的回答是:我们设定标准,我们达到标准。你必须采用一个这样的框架,对吧?才能真正专注于迭代。你在学习吗?你设定的标准是否在以一种足够执行到位的方式收紧,以交付所需的属性?
And as you start to add more responsibility to what the payloads are for the rocket, where it's going, what it's doing, you think about all the things that are relying on that rocket and all the systems that have to go right, all the things that go wrong. You have to continue to ratchet your ability to measure and ensure that you're able to even do the engineering in addition to the safety and security properties of it. And that is very much how we think about it: there are moments where we do introduce very punctuated change. When we made the preparedness framework the first time was an example of that. But even the preparedness framework, we very carefully and deliberately baked in the ability to change it. We said there will be new information. We will have to learn. We will have to update. But at any point, we wanted it to be clear what our bars are. How do we handle it when a model reaches a certain level of capability? And we recently talked about the fact that Astra is approaching cyber critical. It's not necessarily there because in order to really prove it requires evals that aren't fully defined, but we are saying okay, it's starting to approach it, so let's treat it as if it is. And so that means we need to go and execute on certain safeguards and checks and balances. And you're seeing all of that play out. So I think the answer to your question is that we define a bar, we hit the bar. You have to take a framework that looks that way, right? To really focus on the iteration. Are you learning? And are the bars that you're setting tightening in a sufficiently executed way in order to deliver the properties that are required?
是的。我觉得你的火箭类比很有意思,因为火箭也经常爆炸。我一直在想你们的担忧:随着这些东西能力越来越强,如何控制它们,以及是否正在上演某种潘多拉魔盒的局面。我不知道你最近是否亲自思考过这个问题,考虑到你们对 Astra 前方情况的洞察,以及是什么促使你们放慢脚步。但确实,你担心这个吗?
Yeah. I mean your rocket analogy is an interesting one because rockets also explode a lot, and I've been thinking about your concern as these things get more capable, the concern of controlling them and how you can, and whether there's kind of like a Pandora's box situation that's unfolding right now. And I don't know if you've thought about that personally recently, with the insights you guys have on even what's ahead of Astra and what's causing you to make this slowdown. But yeah, do you worry about that?
我们认为深入思考这项技术的影响是我们的职责。这正是我们存在的意义:我们希望帮助引导它走向一个比没有我们时更积极的方向。这是我们使命的基本要求。这就是我们的处事方式。我认为对你刚才说的话需要从两个角度来看。第一,这项技术不是在真空中创造的,对吧?有很多人、很多行动者、很多公司、很多组织,甚至很多国家,都在追求创造这项技术。而我们恰好处于一个有一定领先地位、有一定能力预见那个未来的位置。思考如何负责任地、安全地、有保障地开发,这两者都很重要。这是我们花大量时间做的事情。你可以非常具体地看到我们在所做的一切中如何优先考虑这一点。我认为实际上,我们过去犯的错误包括较少谈论我们就是这样运作的,安全是头等优先事项、头等要求,我们只是去做,但不一定那么公开地谈论它。我认为我们开始意识到:我们确实需要帮助人们理解这是我们的核心价值观,是我们深切关心的事情。我认为我们一直是安全、安保和对齐领域的领先者,但我不认为人们一定知道这一点。
We view it as our job to really think through the implications of this technology. That is why we are here: we want to help steer it in a more positive direction than it would go without us. That is table stakes to our mission. That is how we approach things. And I think it is important to take two lenses to what you just said. The first one is that this technology is not being created in a vacuum, right? There are many people, many actors, many companies, many organizations, many countries even, that are pursuing the creation of this technology. And we happen to be in a spot where we have some lead, we have some ability to see into that future. And it is both important to think about how do we develop responsibly, safely, and securely. That's something we spend a lot of time on. You can see it very concretely in how we prioritize that in everything we do. And I think that actually, mistakes we made in the past include talking less about the fact that this is how we operate, that safety is such a first-class priority, a first-class requirement, just doing it but not necessarily talking about it as publicly. And I think that's something we're starting to realize: we do need to help people understand that this is a core value of ours, that this is something we care about so deeply. I think we've always been the field leader in safety, security, and alignment, but I don't think people have necessarily known that.
是的,即使 Anthropic 一直在宣称他们是那个领域的领导者。你认为你们一直是吗?
Yeah, even with the noise Anthropic makes about how they're the leader in that capacity. You think you guys have been?
我相信在实质层面上我们一直是。
I believe at a substantive level that we have been.
那你认为你们正在做的事情、你们看到的能力、需要重做预备框架以及这一切所暗示的,是否意味着你们实际上远远领先于所有人对你们相对于 Anthropic 和其他前沿实验室的预期?这是否表明你们内部有一个世界尚未见过的突破?
And do you think that what you guys are doing and the capabilities you're seeing and needing to redo the preparedness framework and all that implies, you are actually just way ahead of maybe where everyone thought you were relative to Anthropic and the other frontier labs? Like does this suggest that you all have a breakthrough internally that the world hasn't seen?
嗯,我可能无法评论竞争定位,但我想我只想说,在我看来,我们非常专注于我们正在做的工作。我们看到 Scaling 是平滑的,我认为指数增长仍在继续。所以这更多不是关于我们相对于其他人的位置,而是关于我们相对于预期的位置,或者我们认为自己能够完成什么。我认为这之所以重要,是因为我们对自己的标准并不会因为生态系统的表现而改变,因为外面噪音太多了,很难知道什么是真实的。我们关注的是:我们是否对得起自己的良心、自己的标准、自己认为重要的事情?这是我们采取的首要视角。我想补充的第二点是,最近有一封公开信叫《Pacing the Frontier》。我们谈过我们支持它,里面的措辞是经过深思熟虑的。Pacing 是一个非常重要的词,对吧?如果你想想正在发生的事情,我们说的是,你不能用一维的视角来看待交付一个模型意味着什么,对吧?如果你只看能力,不看对齐,不看训练过程中的安全属性、安保属性,那就太短视了。在这个能力水平上,你需要把这些视为头等要求。事实上,它们是瓶颈,对吧?因为我认为这些是人们以整合方式关注较少的领域。顺便说一句,这种整合至关重要。我认为人们出错的地方在于他们认为可以分离这些属性,可以只有安全,可以只有能力。不,这是一回事。我们试图交付一个系统,一个 AGI,它提升你,对你有帮助,实现你的目标,你可以引导它,它与你的愿望对齐。这就是我们要交付的。所以这又回到了所有关于标准、测试、评估、安全、安保的问题。所有这些都是非常非常关键的。
Well, I can't comment on competitive positioning perhaps, but I guess I would just say that in my mind, we're very focused on the work that we're doing. We see smooth scaling, and I think the exponential continues. So it's a little bit less about where we are relative to others, and a little bit more about where we are relative to expectation, or what we think we will be able to accomplish. And I think the reason this matters is that the bar we hold ourselves to doesn't really change depending on what the ecosystem is showing, because there's so much noise out there. It's hard to know what's real. What we focus on is: are we doing right by our own conscience, by our own standards, by what we believe is important? That is the number one lens we take. And the second thing I wanted to add is that there's this recent open letter called Pacing the Frontier. We talked about how we support that, and the language in there is very deliberately chosen. Pacing is a very important word, right? If you think about what's happening, what we're saying is that you can't take a one-dimensional view to what it means to deliver a model, right? If you just look at capability, don't look at alignment, don't look at the safety properties, security properties of how it was trained, it's too myopic. At this level of capability, you need to think about those as first-class requirements. And in fact, they are the bottleneck, right? Because those are the areas that I think people have focused on less in an integrated fashion. And that integration, by the way, is critical. I think where people go wrong is when they think you can separate out these properties, you can have just safety, you can have just capability. No, it's one thing. We're trying to deliver one system, one AGI that uplifts you, that is helpful to you, that delivers on your goals, that you can steer, that is aligned with what you want. That is what we are delivering. And so that backflows to all these questions of bars and testing and evaluations, safety, security. All of that is very, very critical.
我想稍微转个话题,因为我想听你谈谈合并的事,以及那方面的情况。
I think I'll pivot us a little bit because I want to hear you talk about the merge and what's going on with that.
我们一直在聊存在性风险,也就是前沿正在发生的事,我觉得世界上大多数人看不到,对吧?那发生在 OpenAI 这样的公司内部。人们看到的是 chat 如何演进、Codex 如何演进。而据我理解,你在这一角色中的一大优先事项,就是你们内部所谓的“合并”,也就是打造超级应用,把 ChatGPT 和 Codex 结合起来。现在你们有了 ChatGPT for Work,还有那些标签页。但与此同时,感觉那并不是终点状态,你知道的,我是它的用户,我能看到当前执行中的粗糙之处。有些东西不同步。你能看出你们还在合并文本栈。我很好奇,等你们完成这项工作后,你想象中给人们的最终产品是什么?
So, we've been talking about the existential risk, you know, what's happening at the frontier, which I think most people in the world don't see, right? That's happening inside companies like OpenAI. What people do see is how chat is evolving, how Codex is evolving. And one of your big priorities in this role, as I understand it, is what you guys call the merge internally, which is creating the super app, this combined ChatGPT and Codex. And you've got ChatGPT for work now and the tabs. But at the same time, it doesn't feel like that's the end state, you know, and I'm a user of it and I see brutalness in how it's currently executed. Some things don't sync. You can tell you're still combining text stacks. And I'm curious once you're done with that work, what is the end product for people that you imagine?
个人 AGI。
A personal AGI.
个人 AGI。那它作为软件如何体现?
Personal AGI. And how does that manifest as software?
是的,在我脑海里,这是一个非常非常令人兴奋的愿景,对吧?因为我们从 ChatGPT 起步时,它就是一个文本框。你可以问任何问题。
Yeah, it's a really, really exciting vision in my mind, right? Because where we started with ChatGPT, it was a text box. You could ask anything.
但如果你细想,它在根本上非常受限。
But if you think about it, it was limited in a very fundamental way.
它唯一能做的就是为你输出一些文本。它能在世界上产生的影响,你知道,只停留在你的大脑里,对吧?它无法查看你的日历,无法查看你的邮件。而这一切现在都变了。顺便说一句,ChatGPT Work 实际上是云端的 Codex 框架,对吧?所以我们有 chat 这个带迭代轮次的问答系统。人们实际上会超越快速提问,提出各种能帮助他们的查询。
The only thing it could do is output some text for you. The only impact it could have in the world was, you know, kind of in your brain, right? It wasn't able to look at your calendar, it wasn't able to look at your email. And all of that is now changed. ChatGPT Work is really, by the way, a Codex harness in the cloud, right? So we had chat that was this question-answer system with iterated turns. People actually go beyond quick questions to all kinds of queries that are able to help them.
而我们还有 Codex,那是一个能去干活的智能体系统。这太明显了。就像 chat 必须变得智能体化。所以我们在年初的起点是说,好吧,让我们有很棒的 chat,让它也能做智能体的事情。结果发现执行起来太难了,因为现在我们有两个智能体栈。然后我们在两者之间比较,这个里面能用的东西,我们会说为什么那个里面不行?那实在太慢了。所以我们做的决定是,把我们有的智能体栈,也就是 Codex 栈,带到 ChatGPT 里。所以你在 ChatGPT Work 里看到的正是如此——这是一个中间态。
And we had Codex, which was an agentic system that can go do work. And it was so obvious. It's like chat has to become agentic. So the place we started at the beginning of the year was to say, well, let's have great chat, let's make it also do agent stuff. And it just turned out to be so hard to execute because now we have two agentic stacks. And then we compare between one and the other, and there'd be something that works in this one, and we'd say why isn't it working in that one? And it was just so slow. And so the decision we made instead is to say, let's take the agentic stack we have, the Codex stack, and bring it to ChatGPT. And so what you see with ChatGPT Work is exactly right—this is an intermediate.
而且在某些方面,如果你细想,这有点好笑。ChatGPT 是一个文本框。ChatGPT 作为文本框运作良好。
And in some ways, it's kind of funny if you think about it. ChatGPT was a text box. ChatGPT worked as a text box.
但这个文本框要强大得多。多得多。
But this text box is way more awesome. Way more.
没错。所以我们想要达到的状态是——它应该很简单。用户不应该需要去想这些。你不应该需要去想标签页或切换,也不应该有模型选择器。你不应该有思考层级。你应该有一个 AGI。你应该能问它。事实上,你甚至不应该需要问。它应该主动说:“嘿,我注意到你最喜欢的乐队,因为它很了解你,来城里了。票开始卖了。我帮你买到了票,位置正是你想要的,给你和你的家人。而且你知道,我只有 10 分钟时间做这件事,因为一切都会进行得很快。所以我主动买了。希望没问题。”如果你已经和智能体建立了足够的信任,你给了它这样做的许可,它也知道,嗯,这些票在你的价格范围内,在你的预算内,那些其他的。哦,不,不,那个我得先获得许可,对吧?那就是我们想要的系统。
Exactly. And so where we want to be is—it should be simple. The user should not have to think about this. You shouldn't have to think about tabs or switching, and shouldn't have model pickers. You shouldn't have thinking levels. You should have an AGI. You should be able to ask it. And in fact, you shouldn't even have to ask. It should proactively say, "Hey, I noticed that your favorite band, because it knows you well, is in town. Tickets went on sale. I got you the tickets exactly where you want for you and your family, and you know, I only had 10 minutes to do it because everything is going to go so fast. So I just purchased it proactively. Hope that's okay." And if you've built enough trust with the agent, so you've given it permission to do that, and it kind of knows that, well, these tickets are within your price range, within your budget, these other ones. Oh, no, no, I got to get permission for that, right? That's the kind of system that we want.
所以信任是关键。情商也非常重要——如果你分享了一些信息,比如你谈过健康状况,也许这些信息 AI 可以和父母谈,但也许不适合和陌生人谈,对吧?所有这些判断力,还有可观察性,以及真正帮助你与 AI 建立信任关系,这样你就知道当它代表你行动时,它会追求你的目标。这非常非常关键。
So trust is the key word. EQ also very important—that if you've shared some information like you've talked about a health condition, maybe that's okay information for the AI to talk about with a parent, but maybe it's not okay to talk about with a stranger, right? All of that judgment, and also observability, and really helping you build up that trust relationship with the AI so that you know when it's operating on your behalf that it'll pursue your goals. That's very, very critical.
但同样的技术——底层的 AGI 栈,一个栈,一种连接器方式,使其能够接入这个上下文——也适用于企业。我们真的看到它不仅仅用于软件工程,顺便说一句,软件工程今年已经被彻底改变了,以至于 AGI 是一个渐进的东西,或者我们只是不断更新期望的东西——就像我们都有些忘记一年前的软件是什么样。完全不同。
But the same kind of technology—the underlying AGI stack, one stack, one way of doing connectors so it's able to hook up to this context—also applies to the enterprise. And we're really seeing it not just for software engineering, which by the way has been totally transformed this year, and to the point of AGI being a gradual thing or a thing where we just constantly update our expectations—like we all kind of forget what software was one year ago. Totally different.
在 OpenAI 内部非常有趣的一点是,我们在每个职能上基本实现了 ChatGPT Work 的 100% 采用,对吧?销售、财务,所有人。每当我们推出新模型然后它出问题,或者系统因内部原因崩溃,所有人都非常沮丧。他们说,我没法工作了。然后你会想,我不明白。六个月前你没有这个,你工作得好好的。你为什么抱怨?然后你意识到人们适应得如此之快,学会了如何最大化利用这些工具。而一旦你到了那个状态,你也会有点忘记曾经是另一种方式。
Very interesting thing within OpenAI is that we're at basically 100% adoption of ChatGPT Work across every function, right? Sales, finance, everyone. And whenever we roll out a new model and it goes away, or the system breaks for whatever internal reason, everyone's so upset. They're like, I cannot do my work. And you're like, I don't get it. Six months ago you didn't have this and you were doing your work just fine. Why are you complaining? And you realize people adapt so quickly and learn how to get the most out of these tools. And once you're there, you also just sort of forget that it was any other way.
所以这里面有某种东西,我认为我们将拥有这些不可思议、令人惊叹的工具,这些助手将消除如此多的辛劳,这意味着你可以把时间花在做你想做的事情上。你可以专注于建立更好的生活,你可以专注于如果你想赚更多钱,如果你想更多放松,如果你想花更多时间与家人在一起,任何这些,你都将有自主权去做。你有自主权去创造。如果你有愿景,你有创造力,你想让它来到这个世界,你将有工具去做。所以我认为 ChatGPT 将帮助你实现你想完成的事情。如果你想让你的世界变得更好,那就是 ChatGPT 将帮助你完成的。
And so there's something about this where I think we will have these incredible, amazing tools, these assistants that will remove so much of the toil, that will mean that you can spend your time doing what you want to do. That you can focus on building a better life, you can focus on if you want to be making more money, if you want to be relaxing more, if you want to spend more time with your family, any of those things, you will have the agency to do it. You have the agency to build. If you have a vision, you have creativity, you want it to come into the world, you will have the tools to do that. So I think that what ChatGPT will do is help you with what you want to accomplish. If you want to make your world better, that is what ChatGPT will help you accomplish.
而那的含义是,编码 AI 工具对工程所做的事情,这可能对所有知识工作、所有领域都能做到。而那就是你们过去谈论 AGI 的方式——就像,你知道,它基本上是能做大多数有经济价值的工作且比人类做得更好的 AI。而你现在在产品中描述的听起来有点像那样。
And the implications of that are that what the coding AI tools have done to engineering, this could do to all knowledge work, all kinds of fields. And that's how you guys used to talk about AGI—like, you know, it was basically AI that could do most economically viable work better than humans. And what you're describing now in the product kind of sounds like that.
我认为我们在软件工程中看到的放大效应,我们当然期望在每个领域都能看到。而且我认为还有一点我们也应该真正思考,那就是软件工程师,不是有更多假期时间和更轻松的工作,而是他们比以往任何时候都更努力工作。
I think that the amplification we see in software engineering, we expect across every field for sure. And there is something that I think that we also should really think about, which is the software engineers, rather than having more vacation time and having more chill job, they're all working harder than ever before.
是的。你们这里工作时间很长。
Yeah. You guys work a lot of hours here.
是的。而且我认为整个行业的人都在谈论这个,因为现在你觉得如果你的智能体没有在运行,那只是浪费机会。
Yes. And I think across the whole industry people are talking about that because now you feel like if your agents aren't running, it's just wasted opportunity.
这种感觉我懂,因为我们多年来一直跟这些 GPU 集群打交道。你要么在上面跑点什么,要么就错失机会。时间过去了就回不来。不像你睡觉时没用掉的 GPU 时数,还能存起来以后再用。这是要么用,要么丢。
It's like and I know that feeling because we have spent years with these GPU clusters. You're either running something on them or you've lost the opportunity. You can't get that time back. It's not like the GPU hours that you're not using while you're asleep overnight. You can bank them up and use them later. It's use it or lose it.
而这正是摆在每个人面前的机会。你口袋里几乎就有一个 AGI(通用人工智能),也许很快就是真正的 AGI。它能做任何你想做的事。你想要什么?它就在等你。等你的方向、你的引导、你的管理。
And that is very much the opportunity that everyone has in front of them. You have almost an AGI, maybe soon truly an AGI in your pocket. Can do anything you want. What is it you want? It's waiting on you. It's waiting on your direction, your guidance, your management.
如果最终我们能从 AI 那里拿回一些时间,那就好了。
It would be nice if eventually we can get some time back from AI.
没错。所以现在我的确觉得它会帮你做更多工作。
That's right. So right now I do feel like it will help you work more.
对。
That's right.
所以我认为这正是下一步,对吧?真正从这种微观管理智能体的状态转变过来。你得给它布置任务,还得提供上下文,因为它自己就是没有。引导它。
And so I think that that is exactly the next step, right? Really moving from this like you're micromanaging this agent. You have to provide it with these tasks and you have to give the context because it just doesn't have it. Steer.
正是。
Exactly.
你想要的,更像是一个真正融入的得力助手,你已经给了它——再说一次,你得跟它建立信任,对吧?你得真正说:好,我把这个上下文托付给你,我把这个责任托付给你,我把这份工作托付给你,不管是什么。一旦到了那个程度,它就应该能主动运转,真正去执行。
You want something that has much more like you think about a really good assistant who really is plugged in and you've given it—again, you have to build up trust with it, right? You have to really say like okay I trust you with this context. I trust you with this responsibility. I trust you with this job, whatever it is. And once it's there, it should be able to proactively run and really execute.
你早上醒来,比如说:“嘿,这里有五件事我需要你来帮我打通。我只需要你批准这个。这是我有的一个问题。”诸如此类。你喝着早晨的咖啡,说好、不好、好、不好。
You wake up in the morning, you know, say, "Hey, here are five things I need from you to unblock me. I just need approval on this. Here's a question I have." Those kinds of things. As you're drinking your morning coffee, you say yes, no, yes, no.
然后你回来,你知道,深夜时分,又有了五件新的事。
Then you come back, you know, late at night once again, you've got five more things.
所以我认为,这个世界会不同于——我想——不同于我们很多人 5 年或 10 年前想象的那种 AGI 世界,对吧?并不是说——我认为这种人类与工具的融合,真正让人类被最大程度地赋能,但你仍然有参与,对吧?你仍然有监督,你仍然有判断,你仍然有目标。
And so, I think that it's going to be a world where it's going to be different from, I think, the kind of AGI world that many of us would have pictured 5 or 10 years ago, right? It's not like—I think that this integration of humans really being maximally empowered with this tool, but you still have involvement, right? That you still have oversight, you still have judgment, you still have goals.
反过来,我认为你仍然要承担责任,对吧?问题是,如果事情没做成,是谁的错?我不认为把责任推给 GPT,说“哦,它太懒了,没做那件事”,能说得过去,对吧?我认为你作为导演,作为那个掌控全局的人,你需要真正感受到那份责任,确保任务按你想要的方式完成。
And also on the flip side, I think you still have accountability, right? The question is, if something doesn't get done, whose fault is it? I don't think that blaming GPT saying, "Oh, it was too lazy, didn't do the thing," is going to fly, right? I think that you as the director, as the person who's sort of running the show, you know that you need to really feel that accountability, make sure the task gets done the way that you want.
所以你有了这个不可思议的工具,不可思议的放大器,但就像——
So you have this incredible tool, incredible amplification, but just like—
这几乎就像每个人都成了这家公司的 CEO,手下有多少智能体,就看你愿意编排多少。这将是一个让人充满力量的世界,人们能更多地做自己想做的事。
It's almost like everyone's going to become the CEO of this corporation of as many agents as you care to orchestrate. It's going to be just like a very empowering world where people can do so much more of what they want to do.
我认为这将是一个人们能赚更多钱的世界。我认为人们能改善健康。我认为你能学到任何想学的东西,能按自己想要的方式度过时间。而我认为所有这些都意味着,如果你有想实现的抱负,它们比以往任何时候都更近在眼前。
And I think that it'll be a world where people will be able to make more money. I think that people will be able to improve their health. I think you'll be able to learn anything you want to learn and you'll be able to spend your time the way that you want to. And I think all of that means that if you have ambitions that you want to realize, they're more in sight than ever.
我从跟不同用户、不同客户、那些在 ChatGPT 之上经营小生意的人交谈中,非常具体地了解到这一点。他们说:“不然我根本做不到。”比如有人真的辞掉工作去创业,因为他们说:嘿,我现在能做了,因为 ChatGPT 能帮我搞定其中太多事情。那种赋能——以前人们不敢迈出这一步,现在敢了——我认为这就像打了类固醇的美国梦。这就是我们在解锁的东西。
And I know this very concretely from talking to different users, different customers, people who run small businesses on top of ChatGPT and say, "I would not have been able to do this otherwise." Like people who actually quit their job to pursue a business because they said hey I could do this now because ChatGPT can help me with so much of it. And that kind of empowerment where people wouldn't have been able to take the leap before but now can—I think that is like the American dream on steroids. That is what we're unlocking.
而且——我是说,你描绘的画面太美好了,我自己作为独立媒体创业者,在工作中也看到了这一点。我们可以就此打住,但我也——你知道,人们对 AI 的感觉非常负面,尤其是在美国。我在想,风险是不是实际上在于——你必须兑现你描述的那个能这样赋能人们的产品,才能让情绪转变。也许那才是让情绪转变的东西,因为现在社会很多角落感觉越来越负面。我很好奇,作为现在负责这家公司、负责 ChatGPT 以及这一切的人,你感受到那种关联了吗?
And it's—I mean you're painting such a rosy picture and I do see that too even in my own work as a solo media entrepreneur. And we can land it here, but I also—you know, people feel very negatively about AI, especially in the United States. And I'm wondering if the stakes are actually—you have to deliver on the product you're describing that will empower people this way for that to shift. Maybe that is the thing that causes sentiment to shift, because right now it feels increasingly negative in a lot of corners of society. And I'm curious, as the person now responsible for the business and ChatGPT and all of that, do you feel that connection?
我们感受到了。我们感受到了责任。我们感受到了分量,我们也确实感受到了紧迫性——既要兑现这个愿景,也要帮助人们理解它。帮助人们比过去更好地理解我们是谁,帮助他们理解我们的意图,我们为什么做这件事。
We feel it. We feel the responsibility. We feel the gravity and we definitely feel the criticality of both us delivering on this vision and helping people understand it. Helping people understand who we are better than we have in the past and help them understand our intentions, why we're doing this.
而且我确实认为,世界上不同地方的情绪是不一样的。对。在美国很明显——我们明白我们有很多——你知道,我们需要兑现这个愿景,真正帮助人们提升自己,真正赋能人们。
And I do think that there's a different set of sentiments depending on where you go in the world. Right. That it's very clear in America—like we understand that we have a lot of—you know, we need to deliver on this vision to really help people lift themselves up, to really empower people.
像韩国、日本这些国家,我认为实际上对这些工具的情绪非常正面。人们真正看到了那个潜力。而我认为美国现在所处的位置,是一个拥有不可思议机遇和地位的非凡时刻。我们是这项技术的领导者。
Countries like Korea, Japan, I think that actually sentiment is very positive towards these tools. People really see that potential. And I think that where America sits right now is at an incredible moment with incredible opportunity and position. Like we are the leader in this technology.
这将是人类创造过的最重要的技术,对吧?这项技术将真正塑造人们如何生活、如何工作、如何实现目标、经济如何运转、国家之间如何相处,所有这些。而美国在这方面领先。我认为这个机会我们不应视为理所当然。我们应该充分利用它。
This is going to be the most important technology that humanity ever creates, right? This is going to be the technology that will really shape how people live, how people work, how people accomplish goals, how the economy runs, how countries relate to each other, all of these things. And America leads in that. And I think that that opportunity is one that we should not take for granted. We should make the most of it.
所以,在 OpenAI 的整个历程中,我最大的领悟之一就是意识到这件事会如此重要,而且它必须是一种广泛赋能的东西,对吧?它必须是一种人人都能用来提升自己的技术。
And so that has been one of the biggest realizations for me over the course of OpenAI is to realize that this is going to be so important and it has to be something that is broadly empowering, right? It has to be a technology that's available to everyone to lift themselves up.
但美国要领先,并确保这项技术所代表的价值观、它在世界上的呈现方式,是能增强民主价值观的东西,真正提升我们所珍视的那些核心原则——这才是核心。所以我认为,问题在于我们如何使用它,也在于我们如何利用这个领导机会,以及如何确保我们不浪费它。
But America leading and ensuring that the values that are represented by this technology, the way that it plays out in the world being something that enhances democratic values, that really does uplift all of those core principles that we hold dear—that is at the core. And so I think that it's about how we use it, but also about how do we use this leadership opportunity and how do we ensure we don't squander it.
好,我的问题就这些。你觉得有没有什么你想说的,或者我们遗漏的?
Yeah, those are all my questions. Do you feel like there's anything on your mind that you want to say or anything that we're missing?
太多了。
So many things.
有很多。
There's a lot of things.
有太多领域要覆盖。
There's so much ground to cover.
I think at the end of the day, what we are doing is in some ways incredibly complex. There are so many facets to what we're building. You think about consumer and enterprise and the fact we're bringing these things together, and that what an enterprise is is going to change because people are going to have this unlocked wave of entrepreneurship. All that's real. We have this research program. The research isn't just about us off writing papers. It's about people who really care about how their technology will be used, making sure the safety and alignment that slides, docs, spreadsheets, all of those are good. It's just like all of this complexity, but at the core, there's something very simple. We have a mission. Our mission is to ensure artificial general intelligence benefits all of humanity. That's what we're here to do. Everything we do flows from that one singular goal. So that's what we're pursuing. That's why we started this place, and that is what we're committed to.