OpenAI 总裁 Greg Brockman 谈 GPT-5.5(Spud)与 AI 未来

OpenAI President Greg Brockman on GPT-5.5 (Spud) and the Future of AI

格雷格·布罗克曼 Greg Brockman · Big Technology 播客 · 2026-04-23 · 约 26 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 总裁 Greg Brockman 讨论新模型 GPT-5.5 的能力及其对 OpenAI 竞争地位的意义。

OpenAI president Greg Brockman discusses the new GPT-5.5 model, its capabilities, and what it means for the company's competitive position.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 12)

全文 · Full transcript(中英对照)

引言与 GPT-5.5 概览 Introduction and GPT-5.5 Overview

Host

OpenAI 总裁兼联合创始人 Greg Brockman 加入我们,讨论 OpenAI 的最新模型 Spud,也就是 GPT 5.5,以及它让 OpenAI 在竞争中处于什么位置。接下来就是。打开 Big Technology 播客,今天我们有一期紧急节目,与 OpenAI 总裁兼联合创始人 Greg Brockman 一起,全面探讨 GPT 5.5,即著名的 Spud 模型,看看它能做什么,对 OpenAI 意味着什么。Greg,很高兴见到你。欢迎回到节目。

OpenAI president and co-founder Greg Brockman joins us to discuss OpenAI's newest model, Spud, aka GPT 5.5, and where it leaves OpenAI competitively. That's coming up right after this. Open the Big Technology podcast, today we have an emergency episode with OpenAI president and co-founder Greg Brockman all about GPT 5.5, the famous Spud model. Looking at what it does and what it means for OpenAI. Greg, great to see you. Welcome back to the show.

Greg Brockman

谢谢你邀请我。希望这不算太紧急。

Thank you for having me. Hope it's not too much of an emergency.

Host

嗯,我确实是在拉斯维加斯的酒店房间里录的,所以比我们上次对话更紧急,但我们有一些时间准备。所以,很高兴和你一起。那么,我们就从这开始吧。你能确认 GPT 5.5 就是 Spud 吗?

Well, I am definitely recording in a Vegas hotel room, so more emergency than our last conversation, but we had some time to prepare. So, it's great to be on with you. So, let's just start with this. Can you confirm GPT 5.5 is Spud?

Greg Brockman

是的。

Yes.

Host

好的。GPT 5.5 是什么?

Okay. What is GPT 5.5?

Greg Brockman

嗯,这是一个了不起的模型。我认为它在很多方面是朝着用计算机完成工作的新方式迈出的一步。这是一种新的智能类别。它在编程等方面非常有用,对吧?以及调试和解决非常棘手问题的各个方面,非常主动,真正能够用很少的指令端到端地解决问题。但对我来说,最引人注目的不是它在编码方面变得更好。我认为那是每个人都期望的。但事实上,它现在真的跨越了通用应用的有用性门槛。所以,它在创建幻灯片、电子表格方面好得多,在计算机使用、使用浏览器、能够点击那些本来难以让 AI 操作的应用程序方面也好得多。所以,我认为我们真的看到了这种使用计算机的新方式的出现,它始于这种以智能为核心的东西。

Well, it's an amazing model. I think in many ways it is a step towards a new way of getting work done with a computer. It's a new class of intelligence. It's extremely useful at things like programming, right? And all the different aspects of debugging and solving very hard and gnarly problems, just being very proactive, and really being able to solve problems end to end with little instruction. But, the thing that's to me most remarkable is not necessarily the fact that it got better at coding. Like that, I think is what everyone kind of expects. But, the fact that it's now really crossed the threshold of usefulness for general kinds of applications. And so, it's much better at creating slides, spreadsheets, much better at computer use, using your browser, being able to kind of click through applications that are otherwise hard to have an AI operate. And so, I think that we're really seeing the emergence of this new way of using a computer, and it starts with this kind of intelligence at the core.

两年研究历程与未来改进 Two-Year Research Process and Future Improvements

Host

我们上次交谈时,你提到这实际上是两年研究过程的结晶。那么,这是两年前计划好的吗?OpenAI 的计划能追溯到那么远吗?

When we spoke last, you mentioned that this was effectively the culmination of a two-year research process. So, was this planned two years ago? Is that how far back OpenAI plans?

Greg Brockman

我会说,是的,我们的规划时间跨度确实很长。但有一点要注意,我们叠加了许多不同时间尺度的研究想法和赌注。所以,可以这样理解:我们在整个技术栈的每一个部分都在不断取得进展。因此,GPT-5.5 代表的不是一个终点。在很多方面,它是一个起点。它确实是朝着我们即将在未来几个月内看到的模型迈出的一步。我认为你应该期待,在模型能力的各个方面,我们将有更大的改进。我认为这将非常令人兴奋,我们一直在思考如何让我们生产的东西对实际使用、真实用户和真实应用更有用。

I would say that yes, we do have very long horizons for how we plan. Now, one note is that we stack together many research ideas and bets on a variety of timescales. And so, the way to think about it is that we are making constant progress across every single part of the stack. And so, what GPT-5.5 represents is not an end point. In many ways, it's a beginning point. It's really a step towards the kinds of models that we see coming over even just upcoming months. And I think that you should expect that we are going to have even larger improvements in the capability across a wide variety of these aspects of what the model can do. And that's something that I think will be very exciting, and we're just always thinking about how can we make what we're producing more useful for real-world use, for real users, and real applications.

Host

你能具体分享一下未来几个月我们应该关注哪些方面吗?如果这是开始,那是什么的开始?

Can you share specifically what those aspects are that we should be looking out for over the next few months? If this is the beginning, what is it the beginning of?

Greg Brockman

嗯,我认为我们有一个大愿景,你可以在很多方面看到它,不仅仅是模型,还有那种,你知道,你把模型看作大脑,你可以把系统和工具比如 Codex 以及超级应用这样的应用看作围绕它的身体,使其成为有用的 AI。而这正是正在发生的事情:从语言模型是像我们这样的实验室生产的东西,转向真正有用的 AI。它实际上是一个助手,试图解决你的目标,真正按照你的指令运作。你现在可以看到 Codex 正在成为这个应用,不仅适用于程序员,实际上适用于任何使用计算机的人。它并不完美,对吧?仍然有一些任务它应该能做但做得不太对。有时个性不是你想要的,对吧?它不太,你知道,它非常强大,做了很多了不起的事情,但它与你沟通的方式让你仍然需要花时间仔细阅读,好吧,它到底是如何解决这个问题的?所以,这些方面,我们确切知道如何让它们变得更好。我们已经从 5.4 到 5.5 取得了相当显著的改进。我认为我们将在使这些模型有用的每一个方面取得更显著的改进。内部有一点要知道的是,我们非常关注最终应用。这实际上是我们在过去 12 到 18 个月左右发生变化的一件事:我们过去真的只专注于改进基准测试,让这些模型在智力上更强大,但现在我们真正专注于将它们带到实际应用中。让我们考虑金融、销售、营销,每一个人们使用计算机的功能,我们如何帮助他们完成计算机工作?我们如何让模型不仅具有理论上的帮助能力,而且实际上经历过这些任务,实际上能够看到好的样子?我认为我们正在走向一个地方,你作为做工作的人,你是监督者,你几乎是这个自主公司的 CEO,或者你知道,这支智能体舰队,也许这样说更合适。它们根据你的目标运作。现在,你仍然要负责,对吧?你仍然在驾驶座上。你仍然是那个思考的人:嗯,这真的是我想要的吗?这项工作达到标准了吗?但具体点击了哪些按钮、写了什么代码、电子表格中的公式如何工作这些细节,如果它们对你评估是否得到你想要的东西不重要,你可以从中抽离出来。所以,我认为这就像为每个工人增加杠杆。

Well, I think that the big vision we have, and you can see it reflected in many things, not just the models, but the kind of you know, you think about the models as the brain, you can think about the systems and the harnesses like Codex and the applications like the super app as almost the body around it to make it into a useful AI. And that's really what's happening is a shift from language models being the thing that is produced by labs like ourselves to an AI that's actually useful. It's actually an assistant that's out there trying to solve your goal, that's really operating according to your instruction. And you can see right now Codex is becoming this app that's not just for the coders. It's really for anyone using a computer. And that it's not perfect, right? That there are still some tasks where it should be able to do it and it doesn't quite get it right. Sometimes the personality isn't quite what you wanted, right? That it doesn't quite, you know, it's like extremely powerful and out there doing a lot of really amazing things, but the way it communicates back to you that you have to still spend some time really trying to read through, okay, exactly how did it solve this problem? And so, these aspects, we know exactly how to make them much better. And we've already had a pretty remarkable improvement from 5.4 to 5.5. I think we're going to have even more remarkable improvements across every single aspect of what makes these models useful. And one thing to know internally is that we think a lot about the end application. Like that is one thing that changed for us over the past 12, you know, 18 months, something like that, is that we used to really just be focused on let's improve on the benchmarks, let's make these models more cerebrally capable, but we now are really focused on let's bring them to real-world application. Let's think about finance, sales, marketing, every single function that someone uses a computer, how can we help with their computer work? How can we actually make the model have not just the theoretical capability to help, but is actually experienced those kinds of tasks, it's actually been able to see what good looks like. And I think that the place we're going is one where you as a person doing work that you are the overseer, you are the CEO of almost this autonomous corporation or you know, of this fleet of agents perhaps is more the way to say it. And that they are operating according to your goals. Now, you are still accountable, right? You're still in the driver's seat. You're still the person who thinks about, well, is this what I actually wanted? Was this work up to standard? But that the details of exactly what buttons were clicked and exactly the kind of code that was written or exactly how the formula in the spreadsheet works, that you can abstract yourself from those if they're not important to the evaluation of whether or not something was what you wanted. And so, I think it's like increasing leverage for every worker.

预训练与强化学习对比 Pre-training vs Reinforcement Learning

Host

好的,让我猜猜发生了什么,你告诉我我有多接近。我的意思是,我在想这个。这就像两年工作的结合。有两种不同类型的 AI 训练。有预训练,你只是让模型通过预测下一个词来变得普遍聪明,然后是强化训练,你让它出去实际尝试完成不同的任务,当它做得好时你奖励它,这有效地教会了它或让它学会了如何做这些任务。

Okay, let me take my best guess as to what's happening and you tell me how close I am. I mean, I'm thinking about this. This is like a combination of two years of work. There's two different types of AI training. There's the pre-training where you just make the model generally smart by having it predict the next word and then reinforcement training where you have it go out and actually try to accomplish different tasks and you reward it when it does a good job with those tasks and effectively it sort of teaches it or it learns how to do those tasks.

端到端协同设计 End-to-end co-design

Host

你的意思是不是说,这是 OpenAI 首次将大量针对特定任务的强化学习注入到这个模型中,从而产生了你所说的结果?

Is what you're saying basically that this is the first result where OpenAI has loaded a ton of reinforcement learning on task-specific stuff into this model, and that's what's producing the results you're talking about?

Greg Brockman

嗯,我其实想换个说法。我认为流程中有很多步骤:预训练、中期训练、强化学习、数据收集。所有这些共同作用才产生最终结果,而模型与世界的连接方式也是让它有用的关键。我真正想说的是,我们在每一个环节都有投入,并且拥有可重复的流程。我们有一个团队,从整个技术栈的角度来思考:'我们如何让这个对实际应用更有用?'这不是单一因素,而是整体努力。就像造车,不只是有个更好的引擎。你可以造出很棒的引擎,但如果车的其他部分质量跟不上,那也没用。所以真正的创新在于端到端的协同设计,所有环节以可重复的方式结合在一起,让我们的模型对用户越来越好。

Well, I would actually say it a little differently. I would say that there's many steps in the pipeline: pre-training, mid-training, reinforcement learning, data collection. All these things come together to produce the end result, and the way it's connected to the world is also key to making it useful. The thing I'm really saying is we have been investing in every single one of these and have a repeatable process. We have a team that looks across the whole stack to ask, 'How do we make this more useful for real-world applications?' It's not any one thing; it's the overall effort. If you're building a car, it's not just about having a better engine. You could build a great engine, but if the rest of the car isn't up to quality, it won't matter. So the real innovation is the end-to-end co-design, all coming together in a repeatable fashion to make these models better for our users.

直观使用与提示工程 Intuitive use and prompt engineering

Host

你今天早些时候参加了一个媒体电话会议,你说模型更直观地知道你想要什么,所以你不必像过去那样精确地说明。Rune 有一条推文:'有早期迹象表明 5.5 是一个称职的 AI 研究伙伴。几位研究人员只给了一个高层次的算法思路,就让 5.5 整夜运行实验变体,醒来后发现完成了扫描、仪表盘和样本,从未碰过代码或终端。'两个问题:你们是怎么做到的?这是否意味着提示工程已死?

You were on a media call earlier today and said the model more intuitively knows what you want, so you don't have to spell it out exactly as in the past. There's a tweet from Rune: 'There are early signs of 5.5 being a competent AI research partner. Several researchers let 5.5 run variations of experiments overnight given only a high-level algorithmic idea, waking up to find a completed sweep, dashboards, and samples, never having touched the code or terminal at all.' Two-part question: How do you do that? And does that mean prompt engineering is dead?

Greg Brockman

首先,我认为这归结于我们所说的新一类能力、新一类智能。模型变得更直观易用,因为它们对你所问的内容有更深的理解。它们会查看上下文,试图弄清楚自己被要求做什么。至于第二部分,提示工程是否已死?我实际上认为提示工程可能比以前更有活力。但现在,你花了很多时间向你的电脑解释你想要什么,塞入上下文,说'情况是这样的,我想要这个。'然后你想:'为什么我要向电脑解释这些?电脑应该帮我干活。我不想一步步分解任务;我想指向一个方向,让它处理细节,以我能观察并沿途提供反馈的方式给我结果。'所以我认为提示工程会演变:你可以用更少的努力从这些模型中获得更多,但用同样的努力你仍然有倍增效应。我们现在只是刚刚触及当今模型能力的天花板。

Number one, I think it really comes down to what we mean by a new class of capability, a new class of intelligence. The models are becoming much more intuitive to use because they have a deeper understanding of what you're asking. They look at the context, try to puzzle out what they're being asked to do. As for the second part, is prompt engineering dead? I actually think prompt engineering may be even more vibrant than before. But right now, you spend so much time explaining to your computer what you want, packing in context, saying, 'Here's the situation, here's what I want.' And you think, 'Why do I have to explain this to my computer? The computer should be doing the work to help me. I don't want to break down the task step by step; I want to point it in a direction and have it take care of the details, getting me the result in a way I can observe and provide feedback along the way.' So I think prompt engineering will evolve: you can get so much more out of these models with much less effort, but with the same effort you still have a multiplier. We're just at the leading edge of seeing the ceiling of what today's models are capable of.

经济学与防御性 Economics and defensibility

Host

让我简单谈谈构建这样一个模型的经济学。一直有一种模式:大型模型发布后,被开源模型制作者蒸馏,开源只落后几个月。当投资较小时,领先几个月很重要。但现在投资如此之大,能力提升如此显著,如果这种模式反复出现,长期来看如何保持防御性?

Let me briefly speak about the economics of building a model like this. There's been a pattern where big massive models come out, get distilled by open-source model makers, and open source is just a couple months behind. When investment was smaller, being a couple months ahead mattered a lot. But now that investment is so big and capabilities are increasing dramatically, how is this defensible in the long term if that pattern repeats?

Greg Brockman

我看法略有不同。我们真正的投资在于端到端的协同设计,拥有一个生产这项技术的人员系统,一种协作方式。其中一部分是利用大规模超级计算机来生产这些模型。但事情没那么简单,你不能直接拿输出蒸馏一下,就得到同样能力但更小更快的模型。如果真是这样,我们早就这么做了。蒸馏背后有很多艺术,但关键是,我们真正投资的是制造机器的机器。在部署方面,我们考虑了很多关于模型可能被滥用的防护和缓解措施。这是我们多年来一直在投资的领域,涵盖网络、生物等方面,正如我们公开的准备框架所示。因此,我们所做的每一件事都需要与如何继续取得进展以及如何广泛提供这些模型联系起来,因为我们相信这项技术能够赋能人们,提升所有人。

I look at it a little differently. The real investment we are making is into that end-to-end co-design, having a system of people producing this technology, a way of working together. Some of this is about leveraging massive supercomputers to produce these models. But it's not as simple as taking the output and distilling it to get exactly the same capability in a smaller, faster model. If that were the case, we would just do that. There's a lot of art behind distillation, but the point is that the real thing we are investing in is the machine that makes the machine. On the deployment side, we think a lot about safeguards and mitigations for many aspects of how these models could be misused. That's something we've been investing in for years, across areas like cyber and bio, as seen in our public preparedness framework. So every piece of what we do needs to connect to how we continue to make progress and how we make these models broadly available, because we believe this technology empowers people and lifts everyone up.

Host

是的。但回到这一点,这个模型的定价,我认为是上一个模型 GPT-5.4 的两倍。

Yeah. But, just to go back on that, the pricing on this model is, I think, double the last model, GPT-5.4.

竞争与定价策略 Competition and Pricing Strategy

Host

那么从经济学或商业角度来看,问题是,假设你不断进步,但由于训练模型投入了大量基础设施,如果开源能提供性能稍差但几乎相当、且成本更低的服务,你如何应对这种威胁?

Um and so, from an economics or a business standpoint, the question would be, you know, let's say you keep on progressing, but because there's been all this infrastructure that's been put towards training the models, if open source can deliver not as good performance, but almost as good, and do it cheaper, how do you handle that threat?

Greg Brockman

嗯,我对此看法略有不同。首先,回顾我们的历史,这并非由竞争驱动,而是源于我们自身的进步和渴望。我们每年都在降低同等智能水平的价格,有时甚至降低 100 倍。至少每年一个数量级,有时确实是 100 倍。但随之而来的是真正的杰文斯悖论:当你降低某样东西的成本时,活动量会大幅增加。我认为我们不断看到的是智能的回报:对于这些模型现在能够完成的任务,稍微多一点智能就能带来巨大收益。这正是 5-5 的故事,某种程度上你可以将其视为智能的渐进式改进,但我认为人们将用它做的事情会有巨大飞跃。顺便说一句,我实际上认为“渐进”这个词对于这个模型相对于 5-5 来说太轻描淡写了。你知道,在某些方面它只是 0.1 的改进,但我认为这实际上低估了我们在这个模型中看到的魔力,以及我们的早期测试者在实际工作中真正看到的魔力。

Well, again, I look at it a little differently. So, first of all, if you look at our history, which really is not driven by anything in competition. It's just like our own sort of progress and desire, we have dropped prices on the same level of intelligence year over year, sometimes by literally a factor of 100. Right? It's like at least an order of magnitude year over year, sometimes literally 100. But, the thing that keeps happening, it's real Jevons paradox, where it's like you lower the cost of something, way more activity happens, right? And I think that what we keep seeing is that there are returns to intelligence, right? That for the kinds of tasks that these models are now capable of doing, that a little bit more intelligence goes a long way. And I think that is the story of 5-5, that in some ways you can almost look at it as like, oh, there's just an incremental improvement in intelligence, but I think there's going to be a massive improvement in terms of what people use it for. And by the way, I actually think that incremental is actually very much an understatement for this model relative to 5-5. You know, it's a 0.1 improvement in some ways, but I think that that actually really undersells the magic that we see within this model and that our early testers have really seen in their practical work.

Host

那么,如果人们看到这些数字并说,OpenAI 面临 IPO 压力,因此我们一直享受的智能优惠价和免费午餐结束了。你会反驳这一点。

So, if people see these numbers and they say, there's IPO pressure on OpenAI, and therefore the you know, we've been getting a great deal on intelligence and the free ride is over. You would argue against that.

Greg Brockman

嗯,我的想法是,我们的业务在某些方面非常简单:我们租赁、建设、购买算力,然后以正利润率转售。只要运营利润率为正,并且对智能有可扩展的需求——我认为只要有待解决的问题,需求就存在,没人会缺问题,而且我们每一步都看到需求超过供应——那么我们就可以无限扩展算力。我认为这是我给团队的主要指令:我们需要在原始算力之上增加价值,并确保正运营利润率。这实际上与市场竞争无关,只关乎能否将算力转化为智能,使得产出价值略高于投入成本。我们一直在努力制造更高效的模型,但我们想要更多模型,更智能的模型,无论它们来自哪里,投入的算力本质相同。所以我认为竞争实际上很好,这个市场对创新非常有利,但它实际上推动了更多使用和生态系统的总体支出,你可以从我们和行业其他公司的收入数字中看到这一点。

I Yeah, look, the way I think about this is that we have a very simple business in some ways, right? We rent, build, buy compute, and we resell it with some positive margin. And as long as it's positive operating margin, and as long as there's scalable demand for intelligence, which I think is true as long as there's problems to solve, like no one's going to run out of problems to solve, and we've seen this at every step that the demand outstrips our supply, then we can scale that compute all day. And I think that in my mind, that's the main directive that I ask of the team. It's just like, just think about we need to add value on top of the raw compute and make sure that we are at positive operating margin on it, and that is something where it's actually not even about the different competition in the marketplace. It's just a question of can you have compute that gets turned into intelligence, and that's just how, you know, that it does that at a slightly improved value coming out relative to the cost going in. And I think that is something where again we're always trying to make more efficient models, but then we just want more of them, and then we want the more intelligent models, and regardless of where they're coming from, it's kind of all the same compute that's going in. And so, I think that it's actually a great competition, this marketplace has been great for innovation, but I think that it's actually something where it's driving more usage and more overall spend in the ecosystem, and you can see that in the revenue numbers of us and others in this industry.

网络安全与部署策略 Cybersecurity and Deployment Strategy

Host

好的,我想快速休息一下,然后回来和你聊聊网络安全、信任以及我们在这次紧急节目中能谈到的其他话题。我们马上回来。欢迎回到《大科技播客》,我们的嘉宾是 OpenAI 总裁兼联合创始人 Greg Brockman。Greg,让我问你关于网络安全的影响。OpenAI 和 Anthropic 采取了两种截然不同的方法。Anthropic 最新的巨型模型 Mythos 没有向公众发布。而这个,Spud 或 5.5,却向公众发布了。我的意思是,我直截了当地问你,将如此强大的模型在没有逐步实践的情况下发布到公众手中,是否有可能导致重大网络攻击?

Okay, I want to take a quick break and come back and talk with you about cybersecurity, trust, and whatever else we can get to in our time in this emergency show. We'll be back right after this. And we're back here on Big Technology Podcast with OpenAI president and co-founder Greg Brockman. Greg, let me ask you about the cybersecurity implications here. Two very different approaches between OpenAI and Anthropic. Anthropic's latest massive model, Mythos, is not released to the public. This one, you know, Spud or 5.5, is released to the public. I mean, let me just ask you straight up, is there a chance that releasing this powerful model into the public without this like step-by-step practice could lead to some major cyber attacks?

Greg Brockman

嗯,我实际上对问题的前提有不同看法。需要理解的是,多年来我们一直在投资网络安全保障和网络安全,作为我们准备框架的一部分,对吧?我们在看到这些能力出现之前就已经进行了大量投资。因此,我们一直采取非常审慎的逐步方法。就在过去几周,你可以看到我们扩展了网络安全可信访问计划。总的来说,我们相信生态系统的韧性。我们认为确实需要逐步推进,这些模型会不断改进。我们已经能看到更强大的模型,我们希望将这些模型交到防御者手中,以确保他们能够保护关键基础设施。我们相信这种韧性:当你将模型交到人们手中时,他们能够以其他方式无法实现的方式进行探索。因此,你需要一种渐进的方法,确保在推进过程中引入额外的保障措施,以最大化收益并降低风险。所以我们确实采取了审慎的方法。我认为我们的团队一直在非常努力地思考这个模型的网络安全影响。我们也相信迭代部署,这是随着模型不断改进而逐步发布的一部分。我们相信民主访问。我们认为创造这项技术的最终目标是赋予人们力量,确保它惠及全人类。因此,我们一直在努力解决如何安全、负责任地将这项技术广泛地应用于世界。

Well, I actually have a different view on the premise of the question. So, the thing to understand is that we have been investing in cyber safeguards and cybersecurity as a part of our preparedness framework for years, right? This is something we have invested in far ahead of having the kinds of capabilities we see coming. And so, we have been taking a very deliberate step-by-step approach. You can see even just over the past couple weeks where we've expanded our trusted access for cyber program, and in general, we believe in ecosystem resilience, right? That we think that you do want to go step-by-step, that these models are going to continuously get better. We have line of sight to even more capable ones, and that you want to be able to put these models in the hands of defenders to make sure that you're able to protect critical infrastructure. And we believe in that resilience of as you can bring these models into people's hands that then they're able to explore in ways that you would not be able to without that kind of access. And so, you kind of want this graduated approach and to make sure that you are moving down that pipeline as you can bring in additional safeguards in order to make sure that you can maximize the benefits and mitigate the risks. And so, we've really taken a deliberate approach. I think our team has been working incredibly hard to think through the cyber implications of this model. We also believe in iterative deployment. That's part of this really bringing the models as they continuously get better. And we believe in democratic access. We believe that ultimately the goal of creating this technology is to empower people to ensure that it does benefit all of humanity. And so, we are constantly trying to solve for how do we safely and responsibly bring this technology to bear in the world in a broad way.

Host

没错。我想可以说你的团队并不喜欢 Anthropic 部署 Mythos 的方式。这是 Sam 的一句原话:'这显然是令人难以置信的营销,说我们造了一颗炸弹,即将扔到你头上,我们会卖给你一个价值一亿美元的防弹掩体来运行你所有东西,但只有我们选你当客户才行。' 让我谈谈另一种情况,然后听听你的回应。

Right. And I think suffice it to say that your team hasn't been fans of the way that Anthropic's deployed mythos. This is a quote from Sam. It's clearly incredible marketing to say we have built a bomb, we're about to drop it on your head, we will sell you a bomb shelter for a hundred million to run all your stuff, but only if we pick you as a customer. Let me talk through the other case and then get your response.

部署与安全 Deployment and Security

Host

另一种情况是,你无法预料一切,显然会有一些漏洞只有部署并寻找它们的人或实体才能发现。所以,也许在广泛部署之前,先从一组可信的测试者开始是有意义的。你怎么看?

The other case would be there are you can't account for everything, and there are clearly going to be some vulnerabilities that can't that will only be found by people or entities deploying this and looking for them. So, maybe it makes sense to start with a trusted group of testers before you deploy it before you deploy it broadly. What do you think?

Greg Brockman

嗯,我认为正确答案是微妙的,它根植于你面对的技术细节和许多许多因素,对吧?你需要考虑模型进展如何,对吧?不仅是你自己的能力,还有生态系统中的其他人。你需要考虑让一个小群体访问并能够产生高杠杆效应——他们能否通过发现和生成补丁来发挥高杠杆作用——但随后你如何在整个行业中协调这些信息的披露?所以有很多因素。我认为真正的答案是,任何一种极端都不太对。有可以应用于特定情况的工具。而且我认为这不是我们第一次需要思考这个问题,也不会是最后一次。但有一点要注意的是,我们的模型已经在防御者手中有一段时间了,我们一直在建立我们的可信访问计划。我们发布的模型实际上并不是网络宽松的,对吧?它实际上内置了许多安全措施。这样你可以在私下分享、测试等事情之间留出差距。所以我认为我的简短回答是,在价值观上确实存在不同的思想流派:是希望将模型交到人们手中并赋予他们力量,还是希望集中控制而不让它们落入人们手中?这可能是这些辩论中潜在的紧张点。但我认为策略几乎是从细节中产生的,并且可以受这些价值观的启发。但任何一种极端,本能地,我认为都不会为世界带来最好的结果。

Well, I believe the correct answer here is subtle, and I think it is rooted in the technical specifics of what you have in front of you and many, many factors, right? You need to think about how are the models progressing, right? Not just your own capabilities, but others in the ecosystem. You need to think about what kind of benefit do you get from having a small group that has access and are able to have, you know, are they able to have high leverage by by being able to find and produce patches, but then how do you actually coordinate the disclosure of those across an industry? And so there's a lot of factors that go into it. I think that the true answer is like if either extreme is not quite right. There are tools that can be applied to a specific situation. And I think that this is not the first time we've had to think about this problem. It's not the last time we will have to think about it. But one thing to note is that we have had our model in the hands of defenders um for some time that we've been building up our trusted access program. That the model that we're releasing is actually not cyber-permissive, right? That it actually has a number of safeguards built into it. And that you can then have a gap between what you're privately sharing, testing, those kinds of things. And so I think my my short answer is like it's there's definitely these different schools of thought in terms of values of is the value that you want to get these models into people's hands and empower them or is the value that you kind of want them to be centralized and controlled and that you don't want them in people's hands? That is something that is a maybe underlying tension in some of these debates. But I think that the tactics, right, the you know, that those almost flow from the details and that they can be informed by these values. But either extreme, reflexively, I don't think will yield the best outcome for the world.

对智能体的信任 Trust in Agents

Host

好的,我想问你关于智能体的问题。如果可以的话,我们回到智能体。这些智能体如果你让它们有高度的自主性,它们工作得最好。我是说,这有点道理。所以我只是好奇想听听你的看法。随着我们拥有更多能做更多事情、访问更多文件并跨程序工作的智能体,现在应该对智能体给予多少信任?

Okay, I want to ask you about agents. Uh back to agents if we could. Um these agents work work the best if you sort of let them uh have a high degree of autonomy. I mean, sort of makes sense. Um so I'm just kind of curious to hear your perspective. As we get more agents that can do more things and access more files and work across programs, what is the proper amount of trust to put into agents right now?

Greg Brockman

所以,我认为实际上,目前智能体往往相当可靠。嗯,即使是像提示注入这样的问题,我认为仍然存在漏洞,但我们正在修补它们,模型也变得更加有韧性。但我也认为,另一面是,随着这些模型被赋予越来越多的责任和访问更重要的上下文,你需要对类似员工的情况有所应对——如果你有一个五人的团队,他们都值得信任,没问题。但如果你有 50 万个同样的员工,那么这些数字,就像大数定律一样,你会开始担心:“好吧,我如何实现良好的治理和监督?”对吧。所以,在我们投资这些能力并使超级应用不仅对程序员更易用,也对任何用电脑工作的人更易用的同时,我们也在投资治理和监督。你可以从我们最近发布的 Workspace 智能体中非常具体地看到这一点。在你的企业中,你现在可以定义智能体。你在云端获得一个托管的 Codex 框架。你可以连接工具,连接到你的 Slack,它就能工作。这太棒了。很多人使用它。看到它在一个组织内如何像病毒一样传播,当你看到别人使用另一个智能体时,你会想:“等等,我也可以构建一个。”然后你可以直接复制并做自己的事情。这是一个实现良好治理的机会,因为这已经内置于产品中:你的 IT 部门可以看到所有已创建的智能体,对于每个智能体,你可以看到它的对话记录,并且你可以精确地考虑它的护栏是什么。所以,我认为简短的回答是,你希望随着智能体被赋予的责任和所做事情的多样性增加,同时提升安全性、可观察性和监督。如果你不将这些同步进行,那么我认为这有点失衡。我认为考虑双方都很重要。

So, I think that right now, actually, agents tend to be quite reliable. Um and even things like prompt injections, I think that there's still holes there, but that we're patching them and that the models are becoming much more resilient. But, I also think that the flip side is that as these models be are given increasing responsibility and access to more important context, that you need to have some answer for just like if you have employees, you know, if you have a team of five employees, they're all kind of trustworthy, fine. But, if you have 500,000 of the same employees, that some somehow those numbers, right, just like that there's the law of large numbers, that you start to worry about, "Okay, how do I have good governance and oversight?" Right. And so, this is something where as we're investing in these capabilities and making the super app more accessible not just to coders, but to to any person doing work with a computer, also investing in governance and oversight. And you can see this very concretely in Workspace agents, which we released recently. So, that's within your enterprise, you can now define agents. So, you get a hosted Codex harness in the cloud. You can hook up tools, you can hook up to your Slack, and it's doing work. It's like awesome. A lot of people use it. It's been very cool to see how sort of viral it goes within an organization when you see you use someone else's agent, you're like, "Wait, I can build one of these, too." And you can just fork it and and do your own thing. And then that's an opportunity to have great governance in that you can see that's baked into the product where your IT organization can see all the agents that have been created that for an agent, you can see the conversations it's had, and that you can think about exactly what the guardrails are around it. So, I think that the short answer is like you want to ramp the responsibility entrusted with the agent and the diversity of things that agents are doing together with security, safety, observability, oversight. And if you're not doing those in hand-in-hand, then I think that that that that's a little bit out of balance. And I think it's important to to think about both sides.

Host

是的,基本上,放手去做,但要小心。

Yeah, basically, go ahead, but be careful.

Greg Brockman

但你要真正投入,对吧?我认为就像随着规模扩大,你可以进行原型设计,而规模的本质开始带来一个问题:你还有能力监督正在发生的事情吗?所以你需要确保每一步你都感觉校准了,你理解你的团队在做什么。

But, you and really lean in, right? I think it's like as you scale, like you can prototype and that that it's just the nature of scale that starts to bring in the do you still have the ability to to oversee what's going on? So, you need to to kind of make sure at each step do you feel like you're calibrated, you understand what your teams are up to.

算力驱动经济 Compute-Powered Economy

Host

Greg,让我们以此结束。你称之为算力驱动的经济。那是什么意思?

Greg, let's end with this. Um you called this a compute-powered economy. What does that mean?

Greg Brockman

嗯,我认为我们正在走向一个世界,其中投入问题的算力越多,问题解决得越快。而可解决问题的上限取决于可用算力。想想药物发现,对吧?能够解决复杂疾病。像阿尔茨海默症这样的复杂疾病目前超出了人类的能力范围。我们从未真正解决过它。但想象一个世界,你可以用一个吉瓦级的数据中心,让它思考如何解决阿尔茨海默症。一个月、一年,无论多久。它可能不仅仅是在大脑中解决这个问题,可能还需要咨询世界专家,也许需要建议在湿实验室进行实验。但如果你真的能解决这样一个问题,那对人类来说将是变革性的积极事件。我认为我们正在走向一个世界,重要问题就是这样被解决的。日常生活中的任务也可以这样解决。无论是拥有一个了解你、拥有你的个人背景、值得信赖的智能体,你可以向它咨询健康建议并得到可靠的信息。嗯,那只是一个放在口袋里的智能手机,对吧?你可以直接跟它说话,它就会在外面做事。它主动知道你的目标、你的兴趣以及如何帮助你。

Well, I think we are heading to a world where the more compute is poured into a problem, the faster that problem will be solved. And that the ceiling of problem that can be solved depends on how much compute is available. And you think about things like drug discovery, right? Being able to solve complex diseases. Like, those are Solving complex diseases like Alzheimer's is kind of outside of humanity's reach right now. We've never really done it. But, imagine a world where you can take a gigawatt data center and have it just think about how to solve Alzheimer's. For a month, for a year, however long it takes. And it may not be literally just cerebrally solving this problem, but it may have to consult with world experts, maybe it has to suggest experiments to get run in a wet lab. But, if you can actually solve such a problem, that would be such a transformatively positive thing for humanity. And I think we're heading to a world where that is how important problems get solved. And that is how tasks in your daily life can also be solved. Whether it's having an agent that knows you, that has your personal context, that is trustworthy, that you can ask for advice on health and you get back trustworthy information. Um And that's just a thing that's a smartphone that's that's in your pocket, right? You can just talk to it and it'll be out there doing things. It proactively knows what are your goals, what are your interest, and how it can help you.

算力稀缺性与基础设施 Compute scarcity and infrastructure

Greg Brockman

我认为,无论大小,算力将成为一种资源,它展示了计算机能在多大程度上帮助人们、替人们工作。我们正朝着那个世界前进,而这个世界需要我们共同建设。

And I think that big and small compute is going to be the resource that shows how much computers can be used to help people, to do work on behalf of people. And I think we're heading to that world and it's one that we're all building collectively.

Host

是的,这或许能解释你主导的那些大规模投资,押注于大型基础设施。

Yeah, and that I think would explain the massive investments that you've led making these big infrastructure bets.

Greg Brockman

还是不够。我们会感受到稀缺,我们已经感受到了。你现在就能感觉到,那些试图使用这些智能体的人却无法使用,因为遇到了速率限制。所以我们代表客户、代表所有想用这些智能体的人努力,确保有足够的算力,但我认为我们做不到。我们会尽力,但我觉得我们正走向一个算力稀缺的世界。而且我认为,我们都可以为此做出贡献,努力让世界上有更多的算力可用。

Still not enough. We're going to feel the scarcity. We're going to feel it. We're feeling it already. You can sense it right now on people who are trying to use these agents and just simply cannot, you know, hitting the rate limits. So we're working on behalf of our customers, on behalf of everyone who wants to use these agents to ensure that there is enough, and I don't think we're going to get there. We're going to do our best, but I think that we are headed to a world of compute scarcity. And again, I think this is something where we can all contribute to trying to help there just be more availability of this in the world.

Host

格雷格,今天很忙。总是感谢你抽出时间。和你聊天总是很愉快。再次感谢你来做客。

Greg, busy day. Always appreciate your time. Always great to speak with you. Thanks again for coming on.

Greg Brockman

我也是。聊得很开心。

Likewise. Great chatting.

互动版:逐字朗读 + 针对本期提问 →