AI 的最大瓶颈:能源、计算与 AGI 的未来

AI's Biggest Bottleneck: Energy, Compute, and the Future of AGI

格雷格·布罗克曼 Greg Brockman · Matthew Berman · 2025-10-08 · 约 44 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 的 Greg Brockman 讨论扩展挑战、计算的作用,以及 AGI 从终点到持续过程的演变。

OpenAI's Greg Brockman discusses scaling challenges, the role of compute, and the evolution of AGI from a destination to a continuous process.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 18)

全文 · Full transcript(中英对照)

瓶颈与新玩家 Bottlenecks and New Players

Host

你认为目前最大的瓶颈在哪里?

Where do you think the biggest bottleneck is today?

Greg

我们正走向一个能源将成为巨大瓶颈的世界。

We're heading to a world where energy is going to be a massive bottleneck.

Host

你有没有考虑过一些新玩家,比如 Cerebrus 或 Grock?

Do you ever consider some of the newer players like a Cerebrus or a Grock?

Greg

2017 年,当我们看到 Cerebrus 时非常兴奋。2017 年的 OpenAI 和今天很不一样。

So in 2017, we got super excited when we saw Cerebrus. 2017 OpenAI was, I guess, very different than today's OpenAI.

Host

你会惊讶的。人们仍然不总是听劝。

You'd be surprised. People still don't always listen.

Host

我的工作有危险吗?

Is my job in danger?

Greg

AI 将改变很多工作。

AI is going to change a lot of jobs.

Host

还会有开发者吗?

Are there still developers?

Greg

我认为我们将改变社会契约的许多基础。

I think that we are going to change a lot of fundamentals of the social contract.

Host

你对 AGI 的定义还一样吗?

Do you still have the same definition of AGI?

Greg

我以前真的把它看作一个目的地,但现在我们更认为它是一个持续的过程。

I really used to think of it as this destination, but instead we really think of it as this continuous process.

Host

假设我们现在把算力供应增加 10 倍。收入会增长 10 倍吗?我不确定是否 10 倍,但会增长 5 倍吗?

Let's say we were to 10x our compute supply right now. Would we 10x revenue? I'm not sure if we 10x, but would we 5x?

Greg

ChatGPT 真的让你意识到去静态网站看东西是多么不自然。

ChatGPT really makes you realize how unnatural it is to go to a static website to just read stuff.

Host

你认为软件会完全由 AI 生成吗?

Do you think that software is going to be fully generated?

Greg

我认为是的。这会非常酷。

I think so. I think it's going to be super cool.

Host

OpenAI 内部的讨论是怎样的?

What is the conversation internally at OpenAI like?

Greg

痛苦和折磨。这是真实情况。

Pain and suffering. It's the real truth.

Host

好的。Greg,非常感谢你今天来参加。

All right. Greg, thank you so much for joining me today.

Greg

谢谢你邀请我。很高兴来到这里。

Thank you for having me. It's good to be here.

扩展 Sora 与模型差异 Scaling Sora and Model Differences

Host

我有好多问题要问你。我想先从 Scaling 开始,特别是 Sora。Sora 2 上周发布了。思考像 Sora 这样的模型的 Scaling 是怎样的,它和文本或图像模型有什么不同?

I have a bunch of questions for you. I wanted to start first with scaling and specifically Sora. So Sora 2 released last week. What is it like thinking about scaling a model like Sora and how is it different from a text or image model?

Greg

嗯,我想从根本上思考的方式是,在宏观层面上,一切都还是深度学习,同样的机制,同样的基本原理。你需要用大量的算力进行大规模扩展,前向传播、反向传播、梯度更新。在更细节的层面上,仍然是 Transformer,对吧?这其实非常惊人。

Well, I guess the way I would think about it fundamentally is that, at a broad level, everything is still just deep learning, same mechanics, same sort of underlying principles. You got to scale up massive with a massive amount of compute, a forward pass, a backward pass, gradient step. At a more detailed level, still a transformer, right? Which is actually quite amazing.

Host

好的。

Okay.

Greg

是的。而且,你以不同的方式训练它。你使用不同的过程,更多考虑像扩散这样的东西,或者一种不同的思考如何将算力注入这些模型的方式。但基本上,我觉得最惊人的是,尽管你谈论的是文本与视频,这两种看似截然不同的模态,但在学习和生成它们的实际底层计算过程中,存在巨大的重叠。所以这个事实背后有非常深刻的东西。

Yes. And you know, you train it in a different way. You're using a different process, much more you think about things like diffusion or a different way of thinking about how to pour compute into these models. But fundamentally, the thing that I find so amazing is even though you're talking about text versus video, which are seemingly as different modalities as you could expect, you have this massive overlap in terms of the actual underlying computational process to learn and generate them. And so there's something really deep about that fact.

Host

那么你是否非常看好 Transformer 架构将带我们进入下一个层次,也许是完整的 World Models,Sora 2 显然是朝这个方向迈出的一大步?

So are you pretty bullish that the transformer architecture is going to get us to really the next level, maybe full world models, which is Sora 2 is obviously a big step in that direction?

Greg

是的,我这么认为。两件事:一是我认为有很多问题,比如我们是否错过了大想法?我们是否需要另一个 Transformer 级别的创新?我认为创新空间很大,我们已经看到了,对吧?我认为算法进步一直保持同步。多年来我们做了研究,看看曲线到底如何。我看不到它们停止。Scaling 曲线也在继续,对吧?数据曲线也在继续,这正是驱动这场革命的原因:有多种限制因素,你可以不断调整每一种,然后就会看到模型性能相应提升。所以我认为还有更多东西要构建。如果 AGI 看起来和我们今天的模型类似,我不会惊讶,但如果架构完全相同,我会震惊。

Yeah, I think so. Two things: one is that I think there's a lot of questions around like are we missing big ideas? Do we need another innovation on the level of a transformer? I think that there's lots of room for innovation and we've seen it, right? I think the algorithmic progress has kept pace. We've done studies over the years to see exactly what the curves are. I don't see those stopping. The scaling curves also continue, right? The data curves also continue, and that's the thing that has driven this revolution: there's multiple limiting reagents and each one you can just keep tuning and you'll just see the performance of the models increase appropriately. So I think there's so much more to be built. I wouldn't be surprised if AGI looks kind of similar to the models we have today, but I would be shocked if it was exactly the same architecture.

模型成本与硬件 Model Costs and Hardware

Host

那么当你观察这些不同类型的模型时,尽管它们都是基于 Transformer 的,它们在成本上是否有显著差异?你如何衡量今天提供的不同类型模型的单位经济性?

And then when you're looking at these different types of models, although they are all transformer-based, are they pretty significantly different in cost and how do you measure the unit economics of the different types of models that you're serving today?

Greg

是的,肯定有不同的性能特征。有时我们有不同的推理栈,优化也不同,不同的模型可能更适合不同类型的硬件,所以你可以看到内存和算力之间的精确平衡等等可能不同。所以在系统层面有很多细节上的不同,当你真正试图从硬件中榨取极致性能时,它会把你推向非常不同的方向。但归根结底,我们认为所有这些创新的根本驱动力,以及真正将其带给世界的是算力。所以你必须尽可能多地构建。不同的加速器也有一些专门化。但如果你退一步看,一切都是在做矩阵乘法,做一些注意力机制之类的事情。所以我们在内部做了很多容量调配,并优先考虑一个人时间上的五种不同需求,这个人可以去优化五个不同的模型。这很困难,但这就是我们要做的。

Yeah, there are definitely different performance characteristics. Sometimes we have a different inference stack associated, the optimizations are different, that different models are maybe more amenable to different types of hardware, so you can really see that the exact balance between memory and compute and all those things can be different. So there's a lot of systems work that looks quite different at a detail level, and when you're really trying to extract screaming performance out of the hardware, it pushes you in very different dimensions. But I think at the end of the day, we really view it as the fundamental driver of all of this innovation and really bringing it to the world is compute. So you got to build as much as possible. And there's some specialization of different accelerators and those kinds of things. But if you zoom out, everything is just doing matrix multiplies, doing some attention mechanism, that kind of thing. And so we do a lot of juggling internally of capacity and prioritize five different demands on one person's time who can go and optimize five different models. It's tough, but that's what we're signed up to do.

Host

好的,我们继续谈硬件。AMD 有一个重大公告。那么在 AMD 硬件上构建是否根本不同?是不是现在我们有了这个越来越大的资源池可以调用,还是需要做深度的技术改变?

All right, let's continue on hardware. A big announcement with AMD. So is it fundamentally different to build on top of AMD hardware? Is it just okay now we have this increasingly massive pool of resources we can pull, or is there like deep technical changes that need to be made?

Greg

实际上,我们长期以来一直在多方面投资 AMD 的软件,对吧?因为我们构建在 Triton 之上。这是我们赞助并真正帮助开发的项目,我们绝大多数 GPU 都在高效运行 Triton 内核。我想说,我的看法是,有推理和训练。让推理工作有巨大的固定成本;让训练工作有更大的固定成本。而我们现在已经到了一个点,今年我们只需很少的工作就能使用 AMD 软件并获得良好的性能。所以这很大程度上得益于我们长期以来的合作。我们提供了很多反馈,现在从推理的角度来看,我们对扩展感觉良好。不同的硬件有不同的定位,我们对 MI450 系列感到兴奋,里面有很多好的创新,同时我们也会大量扩展 NVIDIA 用于推理和训练。

So, we've actually been investing in AMD software in many ways for a long time, right? Because we build on top of Triton. That's a project that we've sponsored and really helped develop, and the vast majority of our GPUs are running Triton kernels effectively. And I'd say the way I look at it is that there's inference and there's training. There's a massive fixed cost to getting inference working; there's an even more massive fixed cost to getting training working. And we're at a point where we actually have been able to use, with a very small amount of work, AMD software this year and get good performance out of it. So a lot of that is enabled through this partnership we've had for a very long time. We've provided a lot of feedback, and we're at a point now where from an inference perspective we feel quite good about scaling. And there are different niches for different pieces of hardware, and we're excited about the MI450 series. There's a lot of good innovation in there, and we will also scale up a ton of NVIDIA for inference and training as well.

新芯片玩家与算力稀缺 New chip players and compute scarcity

Host

你有没有考虑过一些较新的玩家,比如 Cerebras 或 Groq,那种晶圆级计算,你考虑过它们吗?

Do you ever consider some of the newer players like a Cerebras or a Groq, like the wafer-scale computing, and do you ever consider those?

Greg

是的。2017 年我们看到 Cerebras 时非常兴奋,因为它是一种全新的范式。你看着那些数字,心想:‘哇,如果我们能有一百万个这样的东西,我们就能构建 AGI 了,对吧?’这是一个非常不同的平台。但事实证明,构建非 GPU 架构比我们在 2017 年预期的要困难得多。不过从一开始,我们就真正地绘制了整个生态系统的地图。我们试图与所有不同的芯片玩家交谈,给他们一些建议,告诉他们工作负载的形状是什么样的。老实说,大多数公司都不听我们的。

Yeah. So in 2017 we got super excited when we saw Cerebras because it was just this totally new paradigm. You looked at the numbers, you're like, 'Wow, if we could have a million of those things, we could build AGI, right?' It's just a very different sort of platform. And I think it's turned out that building non-GPU architectures has been way harder than we expected in 2017. But from the very beginning, we really mapped out the whole ecosystem. We tried to talk to all the different chip players and give them some advice, talk about here's what the shape of the workload is. And honestly, most of the companies wouldn't listen to us.

Host

2017 年的 OpenAI,我想,和今天的 OpenAI 非常不同。

In 2017 OpenAI was, I guess, very different than today's OpenAI.

Greg

你会惊讶的。人们仍然不总是听。但我认为,在某种程度上,甚至不是他们认为我们错了,而是如果你有来自芯片世界的人,他们有特定的问题视角,不了解工作负载,而你试图说:‘不,不,不。这个观点是反的。你真的需要以另一种方式思考,这将是关于大模型,而不是小模型,等等。’那种设计输入。如果你不接受它,就很难在此基础上重新调整你的整个世界观。所以我认为,真正区分这个领域成功玩家的是那些引入深度学习视角或真正关注工作负载走向的人。

You'd be surprised. People still don't always listen. But I think to some extent it wasn't even that they thought we were wrong, but that if you have people who come from the chip world and have a particular way of looking at the problem and don't understand the workload, and you try to say, 'No, no, no. This perspective is backwards. You really need to think about it in this other way, that it's going to be about big models, not small models, whatever.' Those kinds of design inputs. If you don't buy into it, it's very hard to then rebase your whole worldview on top of it. So I think what's really distinguished the successful players in the space has been people who bring in a deep learning perspective or really try to pay attention to where the workloads are going.

Host

当你审视从构建算力到服务推理的整个流程时,你认为今天最大的瓶颈在哪里?

And when you're looking at the entire pipeline of building out compute all the way to serving inference, where do you think the biggest bottleneck is today?

Greg

嗯,我认为我们正在走向一个绝对算力稀缺的世界,而且我认为我们正在走向一个能源——至少在美国——将成为巨大瓶颈的世界。但目前供应链的各个部分都还没有适应我们即将看到的需求冲击。这就是我们多年来一直反复强调的。我们一直在说:‘嘿,我们需要构建更多的算力。’

Well, I think we are heading to a world of absolute compute scarcity, and I think we're heading to a world where energy, certainly in the US, is going to be a massive bottleneck. But there are all sorts of parts of the supply chain right now that have not yet adapted to the demand shock that we see coming. So that's what we've been really trying to say continually over and over for many years now. We've been saying, 'Hey, we need to build more compute.'

Host

是的。有传言说 OpenAI 正在开发自己的芯片,但你是否考虑过投资自己的能源电网和系统,在这方面发明新东西?

Yeah. And there have been rumors that OpenAI is developing its own chips, but do you ever see maybe investing in your own energy grids and systems, inventing new things on that front?

Greg

如果你快进或倒回 10 年前到 2015、2016 年,问我们要做什么,我们在这里是为了构建 AGI。我们当时很大程度上认为这是一项软件事业。事实上,我们认为只需要想出一些新想法,把它们拼凑起来,AGI 就创建了。我们开始发现,实际上这是一项算力事业,算力就像一种基本试剂,比许多其他东西更容易扩展。这就是为什么你在算力上如此努力。你必须把它推到极限。然后你开始意识到,实际上你必须进行大规模物理基础设施建设。所以我们现在正在进入那个世界。我们正在做像 Stargate 这样的事情,开始建造我们自己的数据中心。所以我认为,它在哪里停止,实际上取决于世界愿意供应什么。如果市场真的意识到我们大声疾呼即将到来的需求——不仅来自我们,也来自整个行业——那很好。我很希望不必自己去想办法建造能源。但我们在这里是为了完成使命。

If you had fast-forwarded or rewound 10 years back to 2015, 2016, and asked what we're going to do, we're here to build AGI. And we thought of it very much as a software endeavor. We thought of it in fact as just we need to come up with some new ideas, we'll click it into place, AGI created. We started to find that actually it's a compute endeavor, that this is like a fundamental reagent that is much easier to scale than many other things. That's why you press so hard on compute. You have to push it to the limit. And then you start realizing that actually there's this massive physical infrastructure build you have to do. So we're now getting into that world. We're doing things like Stargate, starting to build our own data centers. So I think where does it stop is really about what the world is willing to supply. If the market does wake up to the demand that we're very loudly trying to say is coming, not just from us but from the whole industry, then great. I would love not to have to go and figure out how to build energy ourselves. But we're here to do the mission.

Host

那么,在目前有限的 GPU、有限的算力下,你有很多相互冲突的需求:消费产品、企业产品、开发者 API、训练。当你们试图决定算力投资去向时,OpenAI 内部的对话是怎样的?

And so with the limited GPU, limited compute that you have now, you have a lot of conflicting needs: consumer products, enterprise products, developer APIs, training. How do you—what is the conversation internally at OpenAI like when you're trying to decide where that compute investment belongs?

Greg

痛苦和折磨。这是真实的情况。这太难了,因为你看到所有这些了不起的东西,然后有人来推销另一个了不起的东西,你会说:‘是的,那太棒了。’

Pain and suffering. That's the real truth. It's so hard because you see all these amazing things and someone comes and pitches another amazing thing and you're like, 'Yes, that is amazing.'

Host

你们做了这么多。那么,怎么——我们是一家小公司,我们无法决定众多事情中该做什么,所以我甚至无法想象在 OpenAI 的规模下会怎样。那是什么样的?再给我讲讲内部对话是什么样的。

And you guys are doing so much. So how—what is the—I mean, we have a tiny company and we can't decide of the many things what to do, so I can't even imagine that at OpenAI scale. What is that? Give me a little bit more about what that internal conversation is like.

Greg

是的,机制方面。我们多年来一直在演进,但现在——我,你知道,我们的首席科学家 Yaka 和负责研究的 Mark 共同决定算力分配。但更广泛地说,实际上首先在研究侧和应用侧之间有一个划分,这通常由 Sam、Fiji 等人裁决。然后在研究内部,我描述了如何分配。在机械层面,我的团队中有很多人致力于这个艰巨的任务,即实际调配 GPU。所以看着真的很神奇。Kevin Park 是一个例子——我团队中的一个人,你去找他,说:‘好的,我们刚有一个项目需要这么多 GPU。’他会说:‘好吧,有五个项目正在收尾,这个项目必须在此时完成,所以我们可以像玩俄罗斯方块一样安排。’看到意图部分——我们希望算力去哪里——以及实际的求解过程,真的很神奇。其中一些是人工的,一些是电子表格,一些是——我认为如果我们的模型能参与其中,那将非常有趣。所以我说这不是一个容易的过程。但我认为看着很有趣,因为算力是团队生产力的驱动因素,所以人们真的很在意。围绕‘我能不能得到算力?’的精力和情感,你无法低估。

Yeah, the mechanics. We've evolved them over the years, but right now there's—I, you know, Yaka, our chief scientist, and Mark, who runs research, together decide the compute allocations. But more broadly, actually, there's first a split between the research side and the applied side, and that usually gets adjudicated at, you know, sort of the Sam, you know, Fiji—that set of people usually makes that call. And then within research, I described how that gets allocated. At a mechanical level, there's a bunch of people on my team who are really dedicated to this hard task of actually shuffling around the GPUs. So it's really amazing to watch. Kevin Park is an example—someone on my team who, you go to him and you're like, 'Okay, we need this many more GPUs for this project that just came up.' And he's like, 'All right, there's these five projects that are sort of winding down and this one's got to finish at this point, so we can sort of make the Tetris work.' It's really amazing to see both the intention part of this—like where do we want the compute to go—but then the actual solver. Some of that is human, some of it is spreadsheet, some of it is—I think it'll be very interesting to see if we get our models in there. So I just say it's not an easy process. But I think it's been very interesting to watch because compute is such a driver of teams' productivity, and so people really care. The energy and emotion around 'Do I get my compute or not?' is something you cannot understate.

将网络引入 ChatGPT Bringing the web into ChatGPT

Host

好的,我们换个话题。你宣布了一个消息——现在,用更好的方式来说,你把网络带入了 ChatGPT。你展示了 Zillow 的例子。

All right, let's shift gears for a second. So you made an announcement—you're now bringing, for lack of a better way to describe it, the web into ChatGPT. You showed off the Zillow example.

算力稀缺与经济生产力 Compute scarcity and economic productivity

Host

嗯,随着应用不断迁移到 ChatGPT 中更原生的体验,你如何看待这种人类与互联网体验的脱钩?因为智能体越来越多地代表我们浏览网页,实际人类上网浏览传统网站的时间似乎在减少。你认为未来 18 个月会是什么样子?另外,我想快速补充一下之前的回答……

Um, as apps continue to migrate towards like a more native experience within ChatGPT, how are you thinking about this decoupling of the human to internet experience that seems to be happening, as agents continue to increasingly browse on our behalf? The time where an actual human is going on the internet browsing traditional websites seems to be decreasing. What do you think the next 18 months look like? And just quickly, actually, I want to add one thing to the previous answer, which is also...

Greg

我认为我们正走向一个算力驱动整个经济生产力的世界。我们在 OpenAI 内部看到的这个缩影,你说你在自己公司也看到了,我认为我们将在各处看到。因此,我真的认为我们必须建设算力,以缓解这种算力稀缺和算力分配冲突,然后再进入下一个问题。

I think where we're headed is a world where compute is the driver of economic productivity throughout the whole economy. And so this microcosm that we've seen within OpenAI that you said you see within your company, I think we're going to see everywhere. And so I really view this like we got to build compute as a way to alleviate this compute scarcity and this compute sort of clash overall allocation before we move on to that next question.

Host

你认为目前供需比例是多少?我们还有多大差距?

What do you think the ratio is between supply and demand right now? How far off are we?

Greg

哦,我认为差距相当大。我不知道数量级。

Oh, I think we're quite far off. I don't know orders of magnitude.

Host

我想说,我不知道如果我们现在将算力供应增加 10 倍,收入是否会增加 10 倍?我不确定是否会增加 10 倍,但会增加 5 倍吗?

I was going to say, I don't know if we were to 10x our compute supply right now, would we 10x revenue? I'm not sure if we would 10x, but would we 5x?

Greg

也许?是的,可能。对。因为我们有太多产品在筹备中却无法发布。你知道,你具体能看到,比如 Pulse,它只对 Pro 用户开放。Pulse 是个很棒的产品。

Maybe? Yeah, maybe. Right. Because we have so many products in the hopper that we cannot release. And you know, you just see it very concretely, things like Pulse, it's Pro only. Pulse is such a great product.

Host

是的,我们接下来会谈到这个。那真是个很酷的产品。

Yeah, we're going to talk about that. That is a really cool product.

Greg

我们需要更多算力。我们需要更多算力。

We need more compute. We need more compute.

Host

我得给你弄个贴纸写着这个。好的。那么,跟我谈谈互联网的脱钩,因为我们浏览互联网的基本方式正在我们眼前发生巨大变化,尤其是智能体能够代表我们浏览,现在又把传统网站引入 ChatGPT。你对正在发生的这种转变有什么看法?

I need to get you a sticker that says that. Okay. So, talk to me about the decoupling of the internet, because it just seems like the fundamental way that we browse the internet is changing drastically right before our eyes, especially with agents being able to browse on our behalf and now with bringing the kind of traditional website into ChatGPT. So, what are your thoughts on that transition that we're seeing?

Greg

嗯,我觉得 ChatGPT 真的让你意识到去一个静态网站阅读东西是多么不自然,对吧?就是静态信息,比如你寻找的一个事实,你在大段不相关的页面中挖掘。我认为我们几乎已经超越了那个阶段。这种情况仍然存在,对吧?但那不是主导范式,也不是人们真正想做的事情。你只是意识到你花了很多时间,那不是增值的时间。那是在大海捞针,机器真的应该为你做这件事。真的应该。

Well, I feel like ChatGPT really makes you realize how unnatural it is to go to a static website to just read stuff, right? Just static information, like a single fact that you're looking for that you're kind of mining through a big page that is kind of not relevant to what you actually want. And I think we've sort of almost moved past that. Still happens, right? That's not the dominant paradigm or the thing that people really want to do. And you just realize there was a bunch of time you were spending. It was not value-add time. It was this sifting through this haystack for a needle, and a machine should really do that for you. It really should.

Greg

而且我认为,随着 ChatGPT 中的应用和这些动态应用,你会开始看到,去一个网站点击一堆按钮来做一些动态的事情,也感觉完全是倒退,我们早就应该超越它了。是的。

And I think what you're going to start to see with apps in ChatGPT, with these dynamic apps, is that going to a website to click a bunch of buttons to do something dynamic also just feels like this total backwards thing that we should have moved past a long time ago. Yeah.

Greg

所以我认为我们正在走向一个人们会更加保护自己时间的世界,对吧?因为对于不增加价值的事情,人类不需要努力思考、提供创造力或提供方向、反馈等,已经没有借口了。如果你只是在筛选一大堆东西,那是 AI 该做的事。

And so I think that we're moving to a world where people are going to be much more protective of their time, right? Because there's sort of no excuse anymore for things that are not adding value, where the human is not thinking hard or providing creativity or somehow providing direction, feedback, something like that. If you're just sifting through a big list of stuff, that's what an AI is for.

Host

那么这如何改变网络上的变现方式?传统上基于 CPM、基于广告的模式:你给网站你的眼球,作为交换,他们给你一些免费内容和广告。但当智能体代表你浏览,尤其是当你把 Zillow.com 这样的网站引入 ChatGPT 时,就会出现各种冲突:他们是否在投放广告?那会是什么样子?所以,随着这些变化发生,你如何看待网络变现层的改变?

And then how does that change monetization on the web, which is traditionally CPM-based, advertising-based? You give the website your eyeballs, and in exchange they give you some free content as well as advertising. But when agents are browsing on your behalf, and especially as you're bringing something like Zillow.com into ChatGPT, then there are all these conflicts of, okay, were they serving ads? And then what does that look like? So how do you think about the changing monetization layer of the web as these changes occur?

Greg

嗯,事实是,我认为还没有人知道,对吧?但我认为我们可以看到,我们必须探索。我们必须找出正确的新变现范式,正确的方式来扩展这一切。但我认为,从根本上说,这些技术带来了新的压力,要求确保你为用户增加价值。如果你看看 ChatGPT 本身,现在它是一个订阅产品,对吧?三年前我们推出时可能没有预料到,但人们愿意付费,因为它增加了价值。在你的职业生活、个人生活等各方面都增加了价值。所以我认为,这并不是说广告没有一席之地,对吧?但我认为,那种你只是无意识地滚动浏览,试图找到你关心的句子,碰巧在那个页面上点击东西的广告,感觉不再是价值的基本驱动力了。嗯,但我确实认为会有新的收入模式。会有新的变现方式。而且我认为,这实际上可能是最令人兴奋的创业时代。

Well, the truth is I think no one knows yet, right? But I think that we can see we're gonna have to explore. We're gonna have to figure out what the right new monetization paradigm is, the right way to scale all of this. But I think that fundamentally there's new pressure from these technologies to make sure that you're adding value to the user. And if you look at ChatGPT itself, right now it's a subscription-based product, right? We maybe would not have predicted that when we launched it 3 years ago, but people are willing to pay because it adds value. Adds value in your professional life, your personal life, just sort of across the board. And so I think that that's not to say that advertising doesn't have a place, right? But I think that advertising that is really about you're just mindlessly scrolling through things to try to find the sentence you care about, and it just happens you're on that page and click a thing, that doesn't feel like quite the fundamental driver of value anymore. Um, but I do think that there's going to be new revenue models. There's going to be new ways to monetize. And I think that it's honestly probably the most exciting time to build yet.

Host

那么,如果你回顾十多年前,看看移动转型期间的出版商,他们中的许多人因为进入苹果应用商店而受制于苹果。你会对他们说什么,为什么这次不同,ChatGPT 真正成为你人工智能体验的第一站、首页?

And so if you rewind a decade plus and you look at publishers during the mobile transition, a lot of them became beholden to Apple because they were in their app store. And what would you say to them on why this is different, where ChatGPT really becomes the first place, the homepage of your artificial intelligence experience?

Greg

嗯,我认为故事还没有写完。我对 AI 的一个观察是,它似乎总是以令人惊讶的方式展开,对吧?完全不同于我们以前见过的任何东西。它有让人联想到的元素,但我认为没有一个范式可以说‘这完全像互联网,或完全像移动,或完全像应用商店’。我认为它是不同的东西。那么,你想要与 AI 互动的方式是什么?是有一个网站来调解你与所有其他事物的互动?我不完全确定,因为 AI 在很多方面是让机器更接近人类,对吧?而不是你必须扭曲自己去想,好吧,有一个带 URL 的网站,我必须去那里。机器就应该按你的要求去做。事实上,机器应该主动思考你可能想要什么,并为你去做。我认为这些范式的转变可能也会改变我们对入口点和机会的看法。

Well, I think the story is not yet written. And one observation that I have about AI is it seems to always play out in a surprising way, right? Just totally different from anything we've seen before. It has elements that are reminiscent, but I think there's no one paradigm that you can say, 'This is exactly like the internet or this is exactly like mobile or this is exactly like the app store.' I think it is something different. So, what is the way that you want to interact with an AI? Is it that there's one website that you go to that just sort of mediates your interactions with everything else? I'm not entirely sure, because AI in many ways is that it brings the machine closer to the human, right? Rather than you having to contort yourself to think about, okay, there's a website with a URL and I have to go to this thing. It's just like machine should just do what you ask. And in fact, the machine should proactively think about what you might want and go do it for you. And I think that these shifts in the paradigm probably will also shift how we think about what the entry point is, where the opportunity is.

主动式与被动式 AI Proactive vs Reactive AI

Host

我想继续聊一下。你认为距离 AI 能预测我大部分需求还有多远?ChatGPT 刚出来时非常被动,我提问它才回应。现在有了像 Pulse 这样的东西,它开始变得更主动。你觉得未来 24 个月内,主动和被动之间的比例会如何变化?

I want to actually continue on that for a moment. How far do you think we are from AI being able to predict most of my needs? When ChatGPT first came out, it's very reactive. I prompt it, it gives me something in return. Now with things like Pulse, it's starting to be much more proactive. How do you see that ratio between proactive and reactive playing out over the next 24 months?

Greg

我认为主动将开始成为更大的焦点。你给一个小任务,然后 AI 就去思考一天、一周、一个月。我们有志于构建能够高效思考一年甚至十年的 AI。人类就是这样做的,对吧?在那一年里没有人类干扰。

I see proactive starting to become much more of the focus. You give a small task and then the AI goes off and thinks for a day, for a week, for a month. We have aspirations to build AIs that can productively think for a year, even for 10 years. Humans do this, right? With no human interruption during that year period.

Host

在那一年里没有人类干扰。

With no human interruption during that year period.

Greg

嗯,我觉得可以类比人类解决问题,比如安德鲁·怀尔斯证明费马大定理。他花了 10 年时间基本上独自研究。但这并非完全没有人类互动;他可能思考子问题并向他人请教。这就是我们想要实现的:能够帮助我们解决重大问题的 AI。拥有能够自主高效工作、无需我们不断微观管理的 AI 是件好事。微观管理人类或 AI 都不有趣。但我们正在走向的世界是,如果你想微观管理,也可以做到。对于高效的人类工作者,如果一直被微观管理,他们可能不会长久开心。所以这确实打开了工作方式,你将能选择如何分配时间。

Well, I think maybe I'd think of it a little bit like a human solving a problem, like Andrew Wiles solving Fermat's Last Theorem. He famously spent 10 years working on it basically by himself. That's not literally no human interaction; he probably thought about subproblems and asked people about it. That's what we want to achieve: AIs that can help us solve grand problems. It's nice to have AIs that can go off and do productive work without us having to constantly micromanage. It's not fun to micromanage humans or AIs. But the world we're heading towards is one where you could micromanage if you want to. For a productive human worker, if you micromanage them all the time, they probably won't be happy for long. So it really opens up how you work, and you'll be able to choose where you want to spend your time.

Host

我看到很多宣传说我们的新 AI 可以独立自主思考 X 小时。你怎么看待 AI 自主思考的时长与在这段时间内实际完成的事情之间的权衡?因为如果花 30 小时做 1+1,那和解决癌症可不一样。你如何看待在给定窗口内的智能压缩和扩展窗口,以及两者之间的权衡?

I've seen a lot of highlights about how our new AI can think for X number of hours independently, autonomously. How do you think about the trade-off between the duration an AI can think autonomously versus what it actually gets done in that period? Because if it takes 30 hours to do 1+1, it's a little different than solving cancer. How do you think about the compression of intelligence within a given window and extending that window, and the trade-off between the two?

Greg

这是个好问题。很容易出现极具误导性的基准测试。我认为很明显,你可以把问题看作背后有计算复杂度。有些问题需要更多思考、更多算力。你想要的是一个能高效思考一天来解决难题的 AI,但如果它能在 10 秒内解决当然更好。所以这是两个不同的维度,重要的是同时推进两者。

It's a great question. It's very easy to have benchmarks that are extremely misleading. I think it's clear that you can think of problems as having a computational complexity behind them. Some problems require more thought, more power, more compute. What you want is an AI that can productively think for a day on a hard problem, but it would be nice if it would just solve it in 10 seconds. So these are two different dimensions, and it's important to keep pushing on both.

Host

好的。那么,GPT-5 思考了多久?Codex 在完全自主的情况下思考了多久?现在的记录是多少?

Okay. With that, how long has GPT-5 thought for? Codex thought for with full autonomy? What's the record right now?

Greg

我其实不知道记录是多少。我想我们发布过。我知道很多人报告看到过七小时,但我不确定那是极限。你可以在网上找到。我们现在已经到了能够开始将大量算力投入到有趣问题上的阶段。

I actually don't know what the record is. I think we've published it. I know a number of people have reported seeing seven hours, but I'm not sure that's the limit. You can find it somewhere online. We're at a point now where we're starting to be able to spend quite a lot of compute on interesting problems.

Sora 2 与社交体验 Sora 2 and Social Experience

Host

我们来聊聊 Sora 2。真的很有趣,很棒。我想我团队里有些人可能有点上瘾了,但没关系。用起来真的很爽。在开发这个从 Sora 1 而来的新模型时,你为什么决定把它做成这种社交体验,而不是像 Sora 1 那样以更传统的方式发布使用?

Let's talk about Sora 2. Really fun, awesome. I think some of my team might be a little addicted to it, but that's okay. It's really a blast to use. As you were developing this new model coming off of Sora 1, why did you decide to build it into this social experience versus just taking the path that Sora 1 took and releasing it for usage in a more traditional way?

Greg

我们通常考虑构建什么界面取决于模型的能力。这在很多方面就是我们最终做出 ChatGPT 的原因。我们一直在做聊天的基础设施,然后有了 GPT-4。我记得第一次做后训练时,我们当时只做指令遵循。我试了:如果你提供另一个依赖于前一个问题与答案上下文的问题,它会泛化使用这些信息吗?它做到了。我想,哇,这个模型真聪明。它想成为一个聊天模型。很明显,技术形态决定了你应该把它作为聊天系统发布。对于 Sora 2,也有类似的感觉。模型的优缺点是什么?你能用它做什么?什么是根本性的新东西?我们本可以走很多方向。对于任何界面或后训练,我总觉得有点遗憾的是,你实际上在深度上大大缩小了原始模型的能力。

The way we usually think about what surfaces to build comes down to the capability of the model. This is how we ended up with ChatGPT in many ways. We were working on infrastructure for chat, then we had GPT-4. I remember doing the first post-training of it, and at the time we were just doing instruction following. I tried: what if you provide another question that depends on the context of the previous question and answer? Will it generalize to use that information? And it did. I thought, wow, this model is smart. It wants to be a chat model. It's so clear that the technology is shaped such that you should release it as a chat system. For Sora 2, there's a bit of that vibe. What are the strengths and weaknesses of the model? What can you do with it? What is fundamentally new? There were many directions we could have gone. The thing that is always a little sad to me about any interface or post-training is that you really narrow down the capabilities of the raw model in deep ways.

Host

有意思。是的,这些原始基础模型用起来非常难,但它们内部有无限的可能性。每个决定背后可能都有很多考量,最终决定过滤出什么。谈谈这个吧。

Interesting. Yeah, these raw base models when you play with them, they're incredibly hard to use, but they have a universe of possibility within them. And there's probably so much behind each of the decisions that leads into deciding what to filter down into. Talk a little bit about that.

Greg

我认为这是外界不太理解的事情。对我来说,这很遗憾,因为我们过去发布过基础模型。GPT-3 是基础模型,没有后训练。用起来非常难。你以前用过 GPT-3 吗?需要大量提示工程,你得提供六个解决任务的例子,然后你知道……

I think this is something people don't really understand from the outside. To me, it's a very sad thing because we used to release base models. GPT-3 was a base model, no post-training. It was incredibly hard to use. Did you use GPT-3 back in the day with all the prompt engineering? You had to provide six examples of solving a task, and then you know...

Host

是的。

Yeah.

Greg

对。你得提供六个解决任务的例子,然后你知道……

Yeah. You had to provide like six examples of solving a task and then you know...

Host

这是模型作为基础模型的特征,而不仅仅是它经过多次迭代变得更好。

That's a function of the model being a base model versus just it getting better over multiple iterations.

基础模型与后训练 Base models and post-training

Greg

是的。具体来说,这些基础模型我们训练它们只做下一个词预测,它们几乎是在观察人类的思想、情感以及所有公开可用的数据。所以它只是试图根据给定的前缀,预测接下来是什么。在推理时,就好像你把它放在一个它从公开数据中找到的文档中间,然后问它接下来是什么。所以你必须思考,如何以在自然分布中可能出现的方式格式化我的查询?结果发现,如果我有问题和答案的列表,比如问题、答案、问题、答案、问题、答案,然后我提出一个问题,那么接下来很可能是一个答案。但如果我只有一个问题,接下来可能还是另一个问题,对吧?所以你是在试图让 AI 角色扮演,让它以为自己处于一个看起来像训练分布的合理文档中间。因为这样很难用,所以它是一个糟糕的界面,不是一个好产品。而且我们也无法控制它表达的行为和价值观。这有点像,对于一个观察世界成长的人类来说,它会拥有所有知识。某种程度上,Alec Bradford 喜欢用的一个类比是,这些基础模型更像是训练一个“人类整体”而不是一个“人类个体”,它包含了所有东西,每一种价值观,每一种世界观。所以对于它在特定情况下如何回应的问题,它几乎可以做出人类可能做出的任何回应。你可以设置模型让它这样做。但如果你真的想缩小到一套一致的价值观,因为你有一些护栏,有一个模型规范说明我们在某些情况下应该如何表现,那么我们需要在它之上增加另一个步骤。这就是后训练:我们要获取这个原始的智能宇宙,并将其精炼成几乎一致的人格或一致的行为集。

Yeah. So the specific way to think about it is that these base models, we train them to just do next-token prediction, and they're almost sort of observing humanity's thoughts and feelings and all the public data available. So it's just trying to say, given this prefix, what comes next? What comes next? So at inference time, it's almost like you're dropping it in the middle of a document that it found somewhere out on the publicly available data, and you're asking it what comes next. So then you have to think, okay, how do I format my query in a way that could occur in this naturally occurring distribution? It turns out that just like a list of if I have a question and an answer and a question and answer and a question and answer and I have a question, probably what comes next is an answer. But if I only have a question, maybe what comes next is another question, right? So it's like you're trying to roleplay the AI into thinking that it's in the middle of some sort of reasonable document that looks like its training distribution. Because this is so hard to use, it's a poor interface. It's not a good product. And it's also something where we have no control over the behaviors and the values that it will express, right? It's a little bit like, for a human who grows up observing the world, it's going to have knowledge of everything. And to some extent, one analogy that Alec Bradford likes to use is that these base models are much more like training a humanity than a human, right? It's got everything in there. It's got every sort of value set. It's got every world view. So to the question of how it will respond in a certain instance, it could be almost anything that a human could respond. You could set up the model so it would do so. But if you really want to narrow down to a consistent set of values because you have some guardrails, you have a model spec that says how we're supposed to behave in certain cases, then we need some other step on top of it. And that's what post-training is: we're going to take this raw universe of raw intelligence and refine it to an almost consistent personality or consistent set of behaviors.

Host

这是否意味着让它成为更社交产品的决定是在后训练之前做出的,还是你发现了一些东西,比如它很擅长模仿?操作的顺序是怎样的?

Does that mean the decision to make it a more social product came before post-training, or did you discover something like, oh, it kind of has this knack to do imitation really well? What was the order of operations?

Greg

它通常是一个迭代循环。你拿到基础模型,然后你尝试用某种方式提示它,觉得这很有趣。实际上,这太酷了。如果它在这方面可靠就好了,你就不需要做所有这些工作。所以基础模型有点像世界上最好的原型引擎,但它们不可靠,因为很难找到正确的提示让它们完成你真正想要的任务。所以这几乎是一个沟通问题,而后训练就是这种沟通。

It usually goes in an iterative loop, right? You take the base model, you kind of see, okay, you prompt it in a certain way, like this is interesting. Actually, this would be so cool. What if it was reliable at this thing? You didn't have to do all this work. So it's a little bit like the base models are the world's best prototyping engine, but they are not reliable, because it is so hard to figure out the right prompt to get them to do the task you really want. So it's almost a communication problem, and then the post-training is this communication.

Host

你的客串是公开的吗?

Is your cameo public?

Greg

我的客串目前不是公开的。

My cameo is not currently public.

Host

我把我的设为公开了,我想 Sam Altman 也提到过这一点。让别人操控你的形象其实出奇地舒服。

I put mine public and I think Sam Altman mentioned this as well. It's actually surprisingly comfortable having people manipulate your likeness.

Greg

是的,我觉得挺容易的,也挺有趣的。

Yeah, I think it's pretty easy. It's pretty fun.

Host

是的,老实说,我对客串状态没有太多考虑,因为我认为 6 个月内,无论我们做什么,其他人都会发布一个视频模型,让你可以制作客串且不受限制。所以我认为我们正走向一个世界,所有人的形象都会被客串。OpenAI,我认为我们的部分立场是真正让人们了解这项技术的走向,并尝试以我们认为有益的方式发布它。所以你可以从我们的选择中看到这一点,但我们也不认为我们对这项技术有完全的控制。我们不是唯一在构建它的人。

Yeah, I think honestly there's not that much thought behind my cameo status, because I think in 6 months, no matter what we do, someone else is going to release a video model that lets you do cameos and is not restricted. So I think we are heading to a world where all of our likenesses are going to be cameoed. OpenAI, I think part of what we stand for is really trying to let people know where this technology is going and try to release it in a way that we think is beneficial. So you can really see that in our choices, but we also don't believe that we have full control over this technology. We're not the only ones building it.

世界模型与 AGI World models and AGI

Host

是的。那么当你看到 Sora 2 时,它是一个世界模型,能够模拟世界。Yann LeCun 说过,仅靠语言不足以建模世界,因此 LLM 不足以达到 AGI。你同意吗?为什么?世界模型如何成为 AI 和实现 AGI 的未来?

Yeah. So when you look at Sora 2, it's a world model. It's able to simulate the world. Yann LeCun has said that LLMs are not sufficient to reach AGI just because language alone is not sufficient to model the world. Do you agree with that? Why or why not? And how are world models really the future of AI and reaching AGI?

Greg

嗯,我喜欢看过去 5 年、10 年 AI 进展中我们学到了什么。有哪些事情我们看到了某种形式的经验证据?我认为语言模型没有世界模型。书面语言中没有足够的信息来构建世界模型的视角。顺便说一句,这是一个长期存在的辩论。这不是 10 年的事情,而是像 50 年、100 年。很长时间。我认为那会预测你无法做到 GPT-4 能做到的一半事情,对吧?因为你问它问题,比如我把水瓶放在桌子上,然后我取下瓶盖,再把它放在桌子下面,瓶盖在哪里?我没有测试过那个具体查询,但你认为它能答对吗?

Well, look, I like to look at what things we have learned over the past 5 years, 10 years of AI progress. What are the things that we have seen empirical evidence for one way or another? I think that the language models don't have a world model. There's not enough information in written language to build a world model perspective. This, by the way, is a long-standing debate. This isn't a 10-year thing. This is like 50 years, 100 years. It's a long time. I think that would have predicted that you would not be able to do half the things that GPT-4 could do, right? Because you ask it questions like, I put the water bottle on the table and then I take the cap off the bottle and then I put it under the table, where's the cap? I haven't tested that specific query, but do you think it's going to get that right?

Host

很可能。是的,我以前有一个测试。杯子里有一个弹珠。你把杯子从桌子上拿起来。弹珠现在在哪里?它还在桌子上。我记得 GPT-3.5 没答对,GPT-4 答对了,然后 GPT-4 及之后版本就完全掌握了。空间意识是存在的。

Probably. And yeah, I used to have a test I gave. There's a marble in a cup. You pick the cup up off the table. Where's the ball now? It's still on the table. So I remember GPT-3.5 didn't get it. GPT-4 got it. And then GPT-4 and beyond just nailed it after that. Spatial awareness is there.

Greg

正是如此。所以这告诉你什么?即使它在当前超级高级的任务上并非完全可靠,你有一个超级高级的任务作为测试,现在它已经征服了它。它展示了一个轨迹。所以对我来说,很容易陷入语义辩论:理解是什么意思?这些模型真的理解吗?还是只是在模拟理解?我说,那到底是什么意思?我不知道这些词的含义。但我确实知道,你给我一个评估,捕捉了你认为模型不可能完成的任务,我们看到模型开始逐渐正确,然后超越它,饱和它。我认为它已经掌握了。

Exactly. So what does that tell you? Even if it's not perfectly reliable at your current super advanced task, you had a super advanced task that was your test and now it's conquered it. It shows you a trajectory. So to me, I think it's very easy to get sucked into semantic debates of what does it mean to understand? Are these models really understanding? Are they just simulating understanding? I'm like, what does that even mean? I don't know what those words mean. But I do know you show me an eval that captures this task that you think should be impossible for the models, and we're seeing the models start to creep up into the right and then blow past it, saturate it. I think it's got it.

Host

这有点像 Sam Altman 之前说的。智能其实就是预测。预测就是智能。

It's kind of like what Sam Altman said earlier. Intelligence is really just prediction. Prediction is intelligence.

AI 对就业的影响 AI's Impact on Jobs

Host

这似乎和大型语言模型实际上能成为 AGI 的论点类似。所以我自私地想问你:我的工作有危险吗?Mr. Beast 说 AI 对内容创作者的生计构成威胁。那是我的工作。我需要担心什么?

It seems like a similar argument to large language models actually will be able to be AGI. So I selfishly want to ask you: is my job in danger? Mr. Beast said AI is a threat to content creators' livelihoods. That's my job. What do I have to be worried about?

Greg

AI 确实会改变很多工作。有些工作会彻底改变或消失。也会出现我们没想到的新工作。这些新工作的平衡和形态还不确定。我认为我们将改变社会契约的基础,走向一个富足的世界。即使你不从事经济活动,也应该拥有极好的生活质量。如果你在奋斗,会有更多东西可以获取和建设。没有人确切知道未来会怎样,但它会比我们想象的更奇特,也可能更令人愉快。

It is definitely true that AI is going to change a lot of jobs. Some jobs will be totally changed or disappear. New jobs we haven't thought of will be created. The balance and shape of these new jobs is uncertain. I think we are going to change fundamentals of the social contract, moving to a world of abundance. Even if you aren't economically working, you should have an amazing quality of life. If you are striving, there will be so much more to gain and build. No one knows exactly what lies ahead, but it will be stranger and probably more delightful than we can imagine.

Host

我刚开始工作,所以我想保住它。

I just started my job, so I'd like to keep it.

Greg

AI 难以改变的是人际联系和像水管工、电工这样的熟练工种。这些领域已经供不应求,AI 很难在其中增加价值。

Things that will be hard for AI to change are human connection and skilled trades like plumbers and electricians. They are already in short supply, and it will be difficult for AI to add value in those domains.

开发者的平台风险 Platform Risk for Developers

Host

我们来谈谈 Codex 和其他产品。在这次开发者活动上,你发布了 Agent Kit。在 OpenAI 上构建的开发者应该如何考虑潜在的平台风险?有种说法是每次 OpenAI 举办开发者日,就有上千家创业公司倒闭。我不相信,但我想听听你的想法:你们自己构建的内容和提供平台让其他人构建的内容之间的界限在哪里?

Let's talk about Codex and other products. At this developer event, you announced Agent Kit. How should developers building on OpenAI think about potential platform risk? The meme is every time OpenAI has a dev day, a thousand startups die. I don't believe that, but I want your thoughts on where the line is drawn between what you build and what you provide a platform for others to build.

Greg

我们经常被问到这个问题。我们希望帮助世界过渡到 AI 优先的经济,让所有人受益。我们无法独自做到这一点;我们需要开发者在我们的平台上构建。我们必须有所选择,因为我们是一家几千人的公司,与整个经济相比微不足道。我们专注于与我们的专业知识协同的领域,比如编码,这也能加速我们自己的工作。我们努力放大尽可能多的人,并在我们能够增加价值的特定领域深入。

We get this question a lot. We want to help transition the world to an AI-first economy that uplifts everyone. We can't do that alone; we need developers building on our platform. We have to be choosy because we are a company of a few thousand people, tiny compared to the whole economy. We focus on domains that synergize with our expertise, like coding, which speeds up our own work. We try to amplify as many people as possible and go deep in specific domains where we can add value.

AGI 的语言 Language of AGI

Host

你认为代码是 AGI 的语言吗?

Do you think code is the language of AGI?

Greg

我真心认为自然语言将成为 AGI 的语言。它们之间可能会用一种稍微优化的英语交流。例如,我们今年为 IMO 生成的数学证明相当可读,像是 AI 发现的一种有趣方言。

I honestly think natural language will be the language of AGI. They might talk to each other in a slightly optimized English. For example, the math proofs we produced for the IMO this year are quite readable, like an interesting dialect the AI discovered.

Host

人类还会在循环中待一段时间吗?我看到中间部分在向外推,但人类在提示和验证方面仍有空间。这种情况会持续多久?

Are humans going to be in the loop for a while? I see the middle pushing outwards, but there is still room for humans in prompting and verifying. How long will that be the case?

Greg

这项技术的根本目的是造福人类和所有生灵。我不认为我们想要一个人类必须编写提示或为上下文工程写代码的世界。那些机械细节感觉像是遗留物。我们希望 AI 工具让机器更接近人类,理解你的目标并帮助你实现它们。这才是真正的关键:提升人类,这也是 OpenAI 的使命。

The fundamental purpose of this technology is to benefit humans and all living beings. I don't think we want a world where humans have to craft prompts or write code for context engineering. Those mechanical details feel legacy. We want AI tools that move the machine closer to the human, understand your goals, and help you achieve them. That is the real key: uplifting humanity, which is OpenAI's mission.

软件生成的未来 Future of Software Generation

Host

你在 Codex 上花了很多时间,思考自然语言编码。你认为软件会在某个时候完全生成,一直到操作系统级别吗?假设我们解决了一致性问题,屏幕上的每个像素都实时生成。

You spend a lot of time in Codex and thinking about natural language coding. Do you think software will be fully generated at a certain point, down to the operating system level? Every pixel on the screen generated in real time, assuming we solve for consistency.

Greg

我认为是的。那会非常酷。

I think so. It's going to be super cool.

全生成式 UI 与人类连接 Fully generative UI and human connection

Host

想象一个完全生成式的用户界面会是什么样子,其实有点令人费解,对吧?就像一个实时的 Sora 那样的东西,你在上面做各种操作。有按钮吗?没有按钮吗?什么才是最自然的?这让你开始意识到,我们构建的很多界面可能都是围绕当前操作系统的偏好或工作方式。但如果你能从头重新想象,没有遗留代码,没有文件、文件夹之类的概念,它会是什么样子?我不觉得我真的知道答案,但我保证它会完全出人意料。

Like it's actually a little bit mind-bending to think about what a fully generative UI would look like, right? Just like a real-time Sora-like thing that you're doing stuff. Are there buttons? Are there not buttons? What's the most natural thing? It makes you start to realize that probably a lot of the interfaces we build are around the proclivities or the way our operating systems actually work right now. But if you could reimagine it from scratch, there's no legacy code, there's no concept of like files and folders and whatever. What would it look like? And I don't feel like I actually know the answer, but again, I guarantee it's going to be something totally surprising.

Host

那么,好吧,我们来稍微设想一下那个未来。在那个世界里还有开发者吗?还有应用程序吗?看起来可能没有,但我忽略了什么?

So, okay, let's envision that future a little bit. Are there still developers? Are there still applications in that world? It just seems like probably not, but what am I missing?

Greg

嗯,看看像 Sora 这样的东西,对吧?它是完全生成式的。顺便说一句,Sora 对我来说非常有趣,因为我记得看过我们制作的一个宣传视频,比尔开着雪地摩托,摘下头盔,我当时想,‘哇,比尔雪地摩托开得真好’,然后我意识到,等等,他并没有真的做这个,对吧?你意识到人类参与的方式非常不同。这不同于一部电影,他实际上在开雪地摩托,但他仍然参与了。他在思考这个的创作过程。这是他的客串。视频中有他的一些特质。你制作一个带有某人客串的 Sora 视频,你分享它,你对此感到兴奋。而你兴奋的事实让我也感到兴奋。今年早些时候,当我们的图像生成走红时,我们实际上学到了这一课。每个人都在生成自己和家人的肖像。我们发现,如果你只是生成一张图像,一张没有根基的图像,比如把你的狗变成某种酷炫的动漫风格,没人会在意。很无聊。但一旦其中有人性,一旦有某种元素将其与这种联系联系起来,人们就会非常感兴趣。我认为你可以有这种放大效应:它本来是你孩子的照片,现在发生了很多有趣的 AI 事情,这个产物是人们能产生共鸣的。但我想回到软件上,我想知道是否也会如此。人们会构建应用,因为会是这样:你想象某个动态系统如何工作,或者无论应用的未来是什么。如果你能以某种方式分享,然后 AI 成为你的开发者,你外包给它,它产生大量代码或完全生成式的用户界面,然后你在某个应用商店分享,那不是很酷吗?这听起来像是策划一个伟大的人类体验,或者用一个更短的词来说,品味,在未来将变得至关重要,比任何像实际开发应用这样的硬技能都更重要。你只是策划这个体验。这是你的信念吗?

Well, take a look at something like Sora, right? It's fully generative. And by the way, Sora is something that's really interesting for me because I remember watching one of the promo videos that we did where Bill's driving around on a snowmobile and he takes his helmet off, and I was like, 'Wow, Bill's so good at snowmobiles,' and I'm like, wait a second, he did not do this, right? And you realize that the way the human is involved is just very different. It's very different from a movie where he actually was snowmobiling, but he's still involved. He was thinking about the creative process of this. It's his cameo. There's something about him that is present within this video. And you make a Sora video with a cameo of someone. You're sharing it. You're excited about it. And the fact that you were excited about it actually makes me excited about it. And we actually had this lesson earlier this year with when our image generation went viral. Everyone was generating portraits of themselves and their family. And the thing is, we realized that if you just make a generated image, some image that has no grounding in being a picture of your dog now turned into some cool anime style, no one cares. It's boring. But as soon as there's some humanity in it, as soon as there's some element that grounds it with this connection, then people are really interested. And I think you can have this amplification: it was a picture of your kid and now there's a ton of AI interesting stuff that happened, and that artifact is something that people connect to. But I guess to bring it back to software, I wonder if that will be the case too. People are going to build apps because it'll be this: you imagine how some dynamic system could work, or whatever the future of apps are. Wouldn't it be cool if you had something like you share in a certain way and then the AI is your developer and you outsource to them, and they produce this massive piece of code or fully generative UI, but then you share on the catch app store perhaps. It really sounds like curating a great human experience, or even a shorter word, taste, is going to be critically important in the future, more so than any of these harder skills like being able to actually develop the app. You just curate this experience. Is that what you believe?

Greg

我认为我相信其中的一部分。是的,我认为这是一个很好的高层次要点。我确实认为有很多机械技能会转移,我们从一代又一代的模型中看到:那些真正尝试去探索模型能力的人,往往能可靠地获得最佳结果。但基本上,知道自己想要什么,有良好的判断力和品味,我认为这些是关键。

I think I believe a shape of that. Yeah, I think that's a good high-order bit. I do think there are a bunch of mechanical skills that will transfer, and we see that from generation to generation of models: people who really try to go and explore what the model's capable of somehow tend to get the best results reliably. But yeah, fundamentally knowing what you want, having good judgment and taste, I think those are critical.

代理商务协议与时机 Agentic commerce protocol and timing

Host

所以你曾是 Stripe 的 CTO,现在你刚刚宣布了智能体商务协议。这是你很久以前就想到的,还是最近内部发现的东西,你说,‘嘿,这是一个很酷的事情,我们可以让智能体代表我们浏览和购买’?还是你早就想到了?

So you were a CTO of Stripe and now you just announced the Agentic Commerce Protocol. Is that something that you thought about a long time ago, or is this more of a recent discovery internally where you said, 'Hey, this is such a cool thing that we can be doing, allowing agents to be able to browse and then purchase on our behalf'? Or is that something you thought of a while ago?

Greg

我的意思是,这个领域的一切都没有新想法。所有这些想法都是人们想过的,我们想过很多次。新的是模型足够有能力来实际利用它们。你可以从插件中看到这一点,对吧?我们几年前做了插件。模型当时还没有准备好。我们一次不能激活超过三个插件,因为如果有太多功能,模型就会混淆,不知道如何调用它们。而今天的模型比之前可靠得多。所以我认为新的是时机,而不一定是想法。

I mean, all the thing about this field is there are no new ideas. All of these ideas are things that people have thought about, we've thought about many times. The thing that's new is models that are capable enough to actually make good use of them. And you can see this with plugins, right? We did plugins a couple years ago. The models just were not ready for it. We couldn't have more than like three plugins active at a time because if you have too many functions, then the model gets confused, doesn't know how to call them. And today's models are infinitely more reliable than where we were before. And so I think it's like the thing that's new is the time, but not necessarily the idea.

Host

好的。你通过 ChatGPT 购物吗?我知道 Sam 说他这么做。

Okay. And do you shop through ChatGPT? I know Sam said he does.

Greg

嗯,有趣的是我不怎么购物。所以,我可以说最近 100% 的购物都是通过 ChatGPT 完成的。

Well, the funny thing is I don't do very much shopping. So, I'd say 100% of my shopping recently has been through ChatGPT.

未来里程碑:难题与算力 Future milestones: hard problems and compute

Host

好的。我们来谈谈未来。我是说,我想我们现在谈的都是这个,但从去年的开发者日(我想是 GPT-4o)到现在已经一年了,你发布了这么多东西。明年的 2026 年开发者日会是什么样子?然后 2030 年的开发者日呢?很难的问题。

Okay. Let's talk about the future for a second. I mean, I guess that's all we're talking about right now, but so much has happened from last year's Dev Day, which was, I think, GPT-4o, and now we're a year later, and you've released so many things. What does next year, 2026 Dev Day, look like, and then what does 2030 Dev Day look like? Tough questions.

Greg

我的意思是,我确实认为明年我们将拥有一些不可思议的模型。我最兴奋的里程碑是真正拥有能够解决难题的模型,对吧?我喜欢做的类比是回想一下 AlphaGo,你知道,2016 年的第 37 步,对吧?那改变了人们对围棋的理解。想象一下在编程中,在材料科学中,在医学中。所以我认为会有真正的突破,可能是 AI 本身,但我认为 AI 与顶尖人类合作。我认为我们将开始看到这一点。那么这对开发者有什么用处呢?几乎无法形容。就像对于任何领域,你想帮助金融领域的人,你可以构建最令人惊叹的应用程序,解决他们最难的金融问题。也许我们不会达到最顶级的金融问题,但我认为我们会解决非常难的问题。顺便说一句,这将需要大量的算力。

I mean, I do think next year we will have some incredible models. The milestone I am most excited about is really having models that can solve hard problems, right? And the analogy I like to make is think back to AlphaGo, you know, 2016 move 37, right? That changed people's understanding of the game. And imagine that in coding, imagine that in material science, imagine that in medicine. And so I think having real breakthroughs that are either potentially the AI by itself, but I think the AI assisted with top humans. And I think we'll start seeing that. So what would be the utility of that for developers? It's almost indescribable. It's like for any sort of domain. You want to help someone in finance and you can build the most amazing application that can go solve their hardest finance problem. Maybe we won't be at the top finance problem, but I think we're going to be at really hard problems. And by the way, it's going to be a lot of compute.

经济价值与 AGI 时间线 Economic value and AGI timeline

Greg

所以有一点是,我们必须确保人们应用这些技术完成的任务具有经济价值,否则没人愿意为算力买单。也许有些任务我们可以做,因为它们是人们想为世界做的事情。那将非常符合我们的使命。但我们必须认真思考如何将这项技术和这种突破性的机器变为现实。2030 年真的很难预测。我认为这已经远远超出了许多人目前对 AGI 的时间预期。

So one thing is we have to make sure that the tasks people apply these things to are economically valuable, because otherwise no one will be willing to fund that compute. Maybe there will be some we can do because they are something people want to do for the world. That would be very aligned with our mission. But we'll have to really think about how to bring this technology and this kind of breakthrough machine to life. 2030 is really hard to predict. I think that's well past many people's AGI time horizons at this point.

Host

你的 AGI 时间线是什么?我不得不问这个问题。

What's your AGI time horizon? I have to ask you.

Greg

当然。定义总是模糊的,但我觉得我们在 1 到 3 年的时间范围内。可能更接近 3 年而不是 1 年,但如果到 2030 年我们还没达到,我会感到惊讶,感觉像是出了什么问题。

Yeah, of course. It's all the usual fuzzy definition, whatever. But I think we're in the 1 to 3 year time horizon. I think maybe more at the three than the one, but I would be kind of surprised. It would feel like something went wrong if we were not there by 2030.

Host

你仍然认为 AGI 的定义是能完成大多数经济上可行的人类工作吗?你现在的定义是什么?有变化吗?

Do you still have the same definition of AGI where it can do most economically viable human work? What is your current definition? Has it changed?

Greg

我认为根本的变化是,我以前真的把它看作一个目的地。我们建造 OpenAI 就是为了完成使命。但现在,我们真的把它看作一个持续的过程,我觉得我在思考方式上已经成熟了。有一个特定的点,即 AI 能在经济价值上匹敌人类,或者我们 2018 年用的任何定义。这是一个重要的里程碑,但不是终点。我认为这很重要,因为关键在于后续的推进。你可以看到人们开始从谈论 AGI 转向超级智能,或者有人完全拒绝这些词汇。对我来说,这不是最重要的。重要的是,你能让 AI 进步吗?你能提升整个经济吗?你能真正把这些好处带给人们吗?你需要思考这意味着什么。想想 Sora 这样的东西,我们如何处理安全问题和 ChatGPT。这些都是重要的话题,非常核心于我们的使命,所以我们真正尝试思考整个端到端的过程。但是的,我认为会有一个点,我们回顾时会说那基本上就是我们 2018 年讨论的东西。而且我认为它并不遥远。

I do think the fundamental thing that has changed is that I really used to think of it as a destination. We're just building OpenAI to complete the mission. But instead, we really think of it, and I think I've sort of grown to maturity in how we think about it, as this continuous process. There's a particular point which is AI that can match humans at economically valuable work, or whatever the definition we used in 2018. It's an important milestone, but it's not the end. I think this is really important because it's about that follow-through. You can see people starting to shift from talking about AGI to superintelligence, or people who reject all these words altogether. To me, that's not the most important thing. The important thing is saying, can you make AI progress? Can you uplift the whole economy? Can you actually deliver these benefits to people? And you need to think about what that means. Think about something like Sora, how we're approaching safety and ChatGPT. These are important topics, very core to our mission, so we really try to think about the whole end-to-end. But yeah, I think there's going to be a point that we'll look back and say that was basically what we were talking about in 2018. And I think it's not that far away.

Host

好的。Greg Brockman,非常感谢。很感激。

All right. Greg Brockman, thank you so much. Appreciate it.

互动版:逐字朗读 + 针对本期提问 →