Codex App: From Professional Tool to Mainstream Builder Platform
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 的 Codex 应用从专业开发者工具转向主流平台,超级碗广告激励数百万人使用 AI 代理进行创作。
OpenAI's Codex app shifts from a professional developer tool to a mainstream platform, with a Super Bowl ad inspiring millions to build with AI agents.
我第一次给别人看的时候,他们说:‘不可能。这就像个假演示。这不可能这么快。这会改变一切。’尤其是因为现在还不是我们能做到的最快速度。
The first time I showed it to someone, they were like, 'No way. This is like a fake demo. Like, this cannot be this fast. This will change everything.' Especially because it's not yet the fastest that we can actually get it to be.
我的体验是试用了这个应用。我真的不想再回到终端了。我意识到其实图形界面很棒。IDE 才是问题所在。有一种编程用的图形界面不是 IDE。你们似乎正在摸索这个,但我甚至不知道那叫什么。
My experience was trying the app. I didn't really want to go back to a terminal. What I realized is actually GUIs are great. IDEs are just the problem. There's something that's a GUI for programming that's not an IDE. And it seems like you're figuring that out, but I don't even know what that's called.
它叫 Codex。
It's called a Codex.
我是 Dan,我想暂时离开本期节目,告诉你关于 Granola 的事。Granola 是一个 AI 会议笔记工具,我几乎每天都在用。这听起来可能有点奇怪或有点 creepy,比如转录你所有的会议。但对我来说,作为领导者,它实际上不可或缺。Every 现在大约有 20 人,理解决策如何做出、我在会议中的表现以及如何最好地帮助团队,对我来说非常重要。Granola 对我来说就像一本领导力日志,我可以看到自己在会议中的表现、特定周出现了什么情况,以及下次如何做得更好。如果你想提升领导力并扩大公司规模,试试 Granola 作为你的 AI 会议记事本。访问 granola.ai/every,代码 every,即可获得三个月免费使用。现在回到节目。Tibo Andrew,欢迎来到节目。
Dan here and I want to take a second away from the episode to tell you about Granola. Granola is an AI notetaker for your meetings and I use it pretty much every day. That may sound a little bit weird or a little bit creepy like transcribe all your meetings. Well, for me, it's actually kind of indispensable as a leader. Every is about 20 people now, and it's really important to me that I understand how decisions get made, how I'm showing up in meetings, and how I can help my team the best way I can. Granola acts a little bit like a leadership log for me, so I can see how I've done in meetings, what situations came up in a particular week, and how I can do better next time. If you're trying to improve as a leader and scale your company, try Granola as your AI powered notepad for meetings. Head to granola.ai/every, code every to get three months free. And now back to the episode. Tibo Andrew, welcome to the show.
嘿,谢谢邀请我们。
Hey, thanks for having us.
谢谢邀请。
Thanks for having us.
太好了,很高兴能和你们聊天。对于那些不了解的人,Tibo,你是 OpenAI Codex 的负责人,Andrew,你是 OpenAI Codex 应用的技术团队成员,你们是当下最受关注的人。OpenAI 刚刚投放了一个关于 Codex 的超级碗广告。你们感觉如何?
Great, great to get to chat with you. So for people who don't know, Tibo, you are the head of Codex at OpenAI and Andrew, you are a member of the technical staff on the Codex app at OpenAI and you are the people of the moment. They just ran a Super Bowl commercial about Codex, OpenAI did. How are you feeling?
是啊,那个超级碗广告相当出人意料,不是吗?确实如此。我认为核心问题是,我觉得这是一个战略转变。你可能以为 OpenAI 会在超级碗投放 ChatGPT 广告,而不是 Codex,尤其是如果你看三四个月前 Codex 的定位是针对专业工程师的,可能不会投放面向更广泛受众的广告。很长一段时间以来,似乎存在这种分歧:Codex 是为专业工程师准备的,如果你想做 vibe coding,就在 ChatGPT 应用里做。但过去一两个月,这种情况似乎发生了很大变化。能谈谈这个吗?
Yeah, that Super Bowl ad was quite surprising, wasn't it? It really was. I think the core thing and the reason I want to start this conversation is it feels like that is a strategic shift. You would expect OpenAI to have run a ChatGPT commercial during the Super Bowl and maybe not, especially if you looked at Codex's positioning like three or four months ago for professional engineers, maybe not have run an ad targeted at a much broader audience. It felt like for a long time there was this divide where Codex was for professional engineers and if you wanted to do vibe coding you do that in the ChatGPT app and it seems like that has shifted a lot over the last like month or two. Can you tell me about that?
是的,我认为尤其是,你知道,我们可以谈谈上周。上周一我们发布了 Codex 应用,立刻看到大量下载,第一周就超过一百万次下载。然后我们知道我们将在周四发布一个非常强大的模型,比如 53 Codex,这让这一点非常明显:我们在这里提供令人难以置信的体验,我们非常致力于 Codex,而且智能体也确实开始发挥作用,能够创建这些东西,即使你技术能力稍弱。我认为这个应用真正展示了这一点:它更吸引人们去尝试,运行多个智能体,因为我们的模型非常擅长多任务处理,并且在长时间运行会话中很可靠。所以它让你能创造更多。因此,感觉也许我们可以激励更多人进行构建,并展示智能体已经到来,对吧?它不再是即将到来,而是将成为主流。为什么不尝试创造新东西,激励他人呢?我觉得这是我们想要强化的正确方向。
Yeah, I think especially like in you know we can talk about last week right so like last week on Monday we released the Codex app, immediately we saw like a ton of downloads, more than a million downloads in the first week, and then we knew that we were releasing like an extremely strong model like you know 53 Codex on Thursday that just made I think this it very visible that you know we're here to you know put incredible experiences out there, we're very committed to Codex, and like also agents are really starting to work and be able to create these things, you know, even if you're like a little bit less technical. I think like the app really showed that, you know, it's like it much more inviting for people to just try it and like know run multiple agents, you know, with our models being like very um very good at sort of like allowing for multitasking and being reliable for long-running sessions. So, like allows you to create a lot more. So, it just felt that, you know, maybe we can inspire more people to build and then show that agents are here, right? It's like it's not it's not coming. It's going to be mainstream. Um you know why don't you try and like create something new and you know inspire people. I felt like the right thing that we wanted to reinforce.
是的。在我们设计和开发这个应用的过程中,我们内部一直有一个自我要求:我们必须做出我们自己喜欢使用、并且用于所有工作的东西。如果我们做不到,我们就不会发布它。这从我们一开始就如此。我认为我们自己也感到惊讶,它竟然这么有趣。尤其是,你知道,我们在开始构建智能体技能之前就开始构建这个应用,一旦我们将它们结合起来,它就变成了一种非常丰富的交互体验,你可以打开浏览器,或者连接到各种服务。于是突然间,我们开始感受到这种真正互联的交互体验,并想要分享。我有点把这个广告看作是一封给构建者的情书,对吧?我从未在超级碗广告中看到过 Linux CD,所以那真的很酷。
Yeah. While we were designing and developing the app, one of our internal mandates to ourselves the whole time was that we had to make something that we love to use and that we used for like all of our work. And if we couldn't do that, then we weren't going to put this out. And this was back when we started. And I think that we surprised ourselves a lot with how fun it was. And especially as you know, we started to build this app before we started to build agent skills and then once we kind of paired them together, it became this really rich interactive experience where you could open the browser or you could connect to these various services. And so all of a sudden we started to feel this like really connected interactive experience and wanted to share. I kind of see the ad as a love letter to builders, right? I have never seen a Linux CD in a Super Bowl ad and so you know like that was really cool to watch.
这个广告有什么影响?
What was the impact of the ad?
我们还在衡量中。我们会看看长期效果如何,但我们确实看到了巨大的流量激增,非常显著,在太平洋标准时间下午 4 点广告播出后,流量激增,我们的系统承受了很大负载。所以这对我来说有点奇怪,就像人们正在看超级碗,然后就去安装应用并当场试用,但确实发生了。很多人联系我们说他们深受启发,之后就想动手构建,这正是我们的目标。
We're still to measure that. We'll see how it plays out over the long term but we saw a giant surge of traffic actually, remarkably, very quickly after 4:00 p.m. PST when it aired, the surge and our systems were under heavy load. So it felt kind of weird to me, like people are watching the Super Bowl and then going and installing the app and just trying it out right there and then, but that happened. And a lot of people reached out saying they were really inspired by it and just wanted to build afterwards, which is what we're aiming for.
我还是想谈谈战略转变。Codex 应用或者说 Codex 整体,从真正面向专业开发者的东西,转向面向更广泛受众,并且可能将一些 vibe coding 从 ChatGPT 转移到 Codex 应用。谈谈这个。
I still want to talk a little bit about the strategic shift. So Codex app moving from or Codex in general moving from something that is really for professional developers to moving to something that has a broader audience and maybe moving some of the vibe coding from ChatGPT into the Codex app. Tell me about that.
我不认为我们试图将 vibe coding 从 ChatGPT 转移到 Codex 应用。实际上有两件事在发生:一是我们在推动专业软件开发的边界,比如 53 Codex 在编码的顶级基准测试中击败了所有其他模型,所以它是一个非常强大的模型,而且在速度和成本上也是一流的表现。二是这个应用确实让事情更容易上手,因此吸引了更广泛的受众,但在内部,我们也看到这个应用在研究团队和我们自己的团队中被广泛使用,整个 Codex 团队都在使用它。它让人们更高效。
I don't think we're trying to move vibe coding from ChatGPT into the Codex app. We're very much two things happening: one, we're pushing the frontier on professional software development, like 53 Codex beats every other model on the top benchmarks for coding, so it is a very capable model, and it's also at the speed and cost it is a top performer. The second thing is the app does make things more accessible and so it does appeal to a wider audience, but internally we're also seeing the app is very much used within research, within our own team, the entire Codex team uses the app. It makes people more productive.
所以,这很大程度上是在深入研究我们认为智能体最佳使用方式,以及我们看到的那些让公司内外的人都非常高效的模式,然后全力投入其中。与此同时,就像,嘿,委托终于来了。它有效,而且更容易获得,我们将尝试看看如何打包并实际交付给更广泛的受众,但这可能不是 Codex 应用。
So, it's like very much leaning into how we think agents are best used, the patterns we were seeing that were making people very productive here at the company and outside, and then just sort of going all in on that. It does happen that at the same time, it's like, hey, delegation is finally here. It works, and it's much more accessible, and we're going to try and see how we can package that and actually ship this to a much wider audience, but that might not be the Codex app.
我的意思是,你一直在用。就像你就在里面构建。
I mean, you use that all the time. It's like you just build in there.
我写的 99% 的代码都是用 Codex 应用。
99% of the code that I write is using the Codex app.
一样。我的意思是,我现在就住在里面。
Same. I mean, I live in there now.
是的。好吧,这其实很有意思。我确实想特别谈谈这个应用,但我想回到你刚才说的,也许如果我理解正确的话,你有点像我们在推动前沿。我们看到很多人,可能不仅仅是高级工程师在使用这个。然而,关于谁在哪个应用里做什么的整体想法,也许你还没有完全搞清楚,而且界限并不像“不再在 ChatGPT 里 vibe coding,而是在 Codex 里 vibe coding”那么清晰。你可以在两者中都做,但我们还没有确切弄清楚你会在哪里做什么。
Yeah. Okay. Well, that's actually really interesting. I definitely want to talk about the app in particular, but I want to go back to the thing you just said, which is maybe if I'm reading you right, you're kind of like we're pushing the frontier. We're seeing lots of people who are maybe broader than just senior engineers using this. However, the overall idea of who is doing what in which app, maybe you haven't totally figured out yet, and it's not as clean of a line as like no longer vibe coding in ChatGPT or really vibe coding in Codex. You can do it in both, but we haven't figured out exactly which thing you're going to do where.
是的,我认为 Codex 是目前最强大的体验,所以你需要相当技术化,才能理解,嘿,代码实际上正在被编写。默认情况下它会在你的机器上执行;它在沙箱中执行。但你可能需要能够阅读代码才能充分利用 Codex。我们会在某个时候为 ChatGPT 带来类似的体验,它在沙箱和概念表示方式上会有不同的特性。也许我们不会显示,嘿,这个可怕的终端命令正在运行,你应该批准它。当然,你不应该对非技术人员这样做。而 Codex 确实是为了吸引所有程序员、构建者、技术人员或技术相关的人,比如数据科学这类人。是的。而且你知道,如果你使用 Codex 应用一段时间,你会看到来自 ChatGPT 的灵感。布局非常相似。我们自动命名你的对话。我们有上下文操作,但它相当简洁,对吧?Composer 看起来非常相似。你会在 ChatGPT 中看到一些类似的灵感用于其他类型的事情。但我们仍然相信,当我们开始为专业软件开发人员制作东西时,对我们来说,它值得一个专门的体验,能够真正展示模型的力量以及模型改变开发生命周期的方式。所以我们为此量身定制了一些东西。我们在内部与研究团队、产品团队取得了很大成功。所以我们会看得更远,但我觉得我们对这种量身定制的方法的结果非常满意。
Yeah, I think Codex is like the most powerful experience right out there, so you should be fairly technical so that you understand, hey, code is actually getting written. It's going to get executed on your machine by default; it's executed in the sandbox. But you should probably be able to read code in order to use Codex to its fullest. We will bring a similar experience to ChatGPT at some point, which will have different properties in terms of the sandbox and how concepts are represented. Maybe we won't be showing, hey, this scary terminal command thing is running and you should probably approve it. Of course, you shouldn't do that to someone who is not technical. And Codex is really there to appeal to all coders, builders, technical people, or technical adjacent like data science, these kinds of things. Yeah. And you know, if you use the Codex app for any amount of time, you can see the inspirations from ChatGPT. The layout's very similar. We auto-name your conversations. We've got contextual actions, but it's pretty clean, right? The composer looks very similar. And you'll see some of that inspiration back in ChatGPT for other types of things. But we still believe that when we set out to make something for the professional software developer, and for us, that deserved a dedicated experience that could really showcase the power of the models and the way the models could change the development life cycle. And so we made something very tailored to that. And we've had a lot of success internally with research teams, with product teams. And so we'll look beyond, but I think we're really happy with where we've ended up on that tailored approach.
你能告诉我决定投资 GUI 而不是 TUI 的原因吗?我觉得 TUI 现在很热门。显然你已经为 Codex 有了一个 TUI,你本可以说,好吧,我们要加倍投入,让终端体验比现在更好,并真正投资于那个,而不是,好吧,我们要做 GUI。我认为做 GUI 有点反直觉或反叙事。所以告诉我那个决策过程。
Can you tell me about the decision to invest in a GUI over a TUI? I feel like TUIs are so hot right now. And obviously you have one for Codex already and you could have said, okay, we're going to double down and just make the terminal experience even better than it is now and really invest in that versus, okay, we're going to go GUI. I think making a GUI is a little bit of a counterintuitive or counternarrative thing to do. So tell me about that decision process.
我认为这不是反直觉的。更可能的是它不主流,对吧?所以我们尝试了很多不同的方法。我非常认为我们仍处于实验阶段。我们主要负责两件事:构建最强大的、能够编码的实体,然后这逐渐会成为一个多智能体系统,它会变得越来越有能力,你将需要弄清楚如何引导和监督它的结果和行为。这是我们正在构建的一件事。然后我们也在构建你如何与它交互?什么是获得对这个非常强大的实体或实体系统正在做什么的可见性的最佳方式?你如何引导它们?你如何监督它们?所以我们仍在实验那是什么。当然,你可以在 TUI 中做到这一点。但在某个时候,它开始感觉非常有限,尤其是在多模态方面,比如模型可以画小图和生成图像,或者你可以用语音交谈。也许你同时运行很多个,然后你就开始失去跟踪。所以我们觉得我们需要开始尝试其他东西。直到我们在内部看到它变得超级流行,我们才觉得,我们必须把它发布到外部。这已经到了一个地步,它太好了,不能只留给自己。我的意思是,那就是你经历的旅程。你当时没有在应用中构建,尽管你是什么时候开始在应用中构建的?那实际上相当快,就像应用在自我构建的时候。
I think it wasn't counterintuitive. It's more maybe it's not mainstream, right? And so we experiment with a lot of different approaches. I very much consider that we're still in the experimentation phase. And we're responsible primarily for two things: building the most powerful entity out there that's capable of coding, and then increasingly this will become a multi-agent system and it will become more and more capable, and you will have to figure out how to steer and supervise its outcome and its behavior. That's one thing we're building. And then we're also building how do you even interact with this? What is the optimal way to have visibility into what this very capable entity or system of entities is doing? How do you steer them? How do you supervise them? And so we're very much still experimenting with what that is. Sure, you can do it in the TUI. At some point it starts to feel very limiting, especially on multimodal like the models can draw little diagrams and generate images, or you can talk over it using voice. Maybe you have many of them going in parallel and so you start to lose track. So we felt like we needed to start experimenting with something else. And it was only when we saw it become super popular internally that we were like, we have to ship this externally. This has come to a point where it's too good to just keep it to ourselves. I mean, that was like the journey you went through. You were not building in the app, although when did you start building in the app? That was actually fairly quickly, like when the app was building itself.
那相当快,是的,因为我从 TUI 和 IDE 扩展开始,我认为我个人的目标是如何尽快在应用上完全构建应用。
That was pretty quickly, and yeah, because I was starting with the TUI and with the IDE extension, and I think that my goal personally was how can I get to fully building the app on the app as fast as possible.
对。在构建这些东西时,很容易陷入一种模式,哦,这对某人会有好处。有人会喜欢这个,某种类型的人会喜欢这个。所以我们真的想尽快达到,我希望能够在应用上构建应用。我希望它能够通过技能自我运行。我希望它能在它生成的应用上点击操作。我希望这尽快成为我工作流程的一部分。
Right. It's really easy when building this stuff to slip into the mode of, oh, this will be good for somebody. Somebody will love this, a certain type of person will love this. So we really wanted to get quickly to, I want to be able to build the app on the app. I want it to be able to run itself with skills. I want it to click around on the app that it spawned. And I want this to be part of my workflow as soon as possible.
我有时还是会用 TUI 快速执行一些操作,但 UI 的灵活性——让一些面板持久、另一些临时——确实有它的好处。我们在应用里加入了语音功能,所以你可以用语音提示。我们还有 Mermaid 图表和完整的图像渲染。所有这些都只是我们想用专用 UI 做的事情的冰山一角。它很简单,而且是故意简单的,但我们会在动态内容上做很多文章。
I still use the TUI sometimes when I want to fire something quick, but there is something about the flexibility of controlling UI and being able to have some pains be persistent and others be ephemeral. We shipped voice with the app, so you can prompt with voice. We have mermaid diagrams in the app, we have full image rendering. All those things are the tip of the iceberg on what we want to do with a dedicated UI. It's pretty simple and simple intentionally, but we're going to do a lot with dynamic stuff there.
天花板高得多。
The ceiling is just much higher.
是的,很有趣。我的体验是,试用了这个应用后,我真的不想回到终端了。在那之前的几个月里,我主要用 Claude code 和终端里的 Codex 编程。我想我意识到的是,GUI 很棒,IDE 才是问题所在。有一种用于编程的 GUI 不是 IDE,你似乎正在摸索这个。我甚至不知道那叫什么。它叫 Codex 应用。
Yeah, it's interesting. My experience was trying the app. I didn't really want to go back to a terminal. I had been coding mostly in Claude code and some Codex in the terminal for several months before that. I think what I realized is that GUIs are great, IDEs are just the problem. There's something that's a GUI for programming that's not an IDE, and it seems like you're kind of figuring that out. I don't even know what that's called. It's called a Codex app.
在开发过程中有那么一刻,每个人都在 fork 同一个 IDE,我们面面相觑,说:‘嘿,我们是不是也该 fork VS Code?’我清楚地记得是哪一天。我不确定 IDE 是不是问题所在,但我有时会用卡车类比:我会偶尔打开 IDE。今天我就打开了一个,为了做一件非常具体的事,然后关掉它,回到 Codex 应用。我认为 Codex 应用作为日常主力工具很好,偶尔你需要 IDE 或非常复杂的终端设置,但这应该是你的大本营,是你运行中的智能体的指挥中心,一个你可以回来追踪所有事情的地方。关于是否允许像 IDE 那样自由形式的面板,我们做了很多设计决策。我们得出的结论是,这些模型擅长的是知道当前任务需要什么,所以我们希望对什么在什么时候显示有更全面的控制。你可以在计划模式中看到这一点:你不一定会得到一个编辑器,而是得到一个快速回答问题的方式。你有计划,可以编辑它,我们想在这方面做得更多。
There was a moment during development where everybody and their mother was forking the same IDE, and we looked at each other and said, 'Hey, should we have done a fork of VS Code as well?' I remember exactly which day it was. I don't know if I would say IDEs are the problem, but I go back to the truck analogy sometimes: I will open an IDE here and there. I opened one today for something very specific, then closed it and went back to using the Codex app. I think there is something there with the Codex app being a great daily driver, and occasionally you need an IDE or a really complex terminal setup, but this should be your home base, your command center for the agents that are running, a place you can come back to and track everything. There were a lot of design decisions around whether to allow free-form panels like an IDE. We concluded that a lot of what these models are great at is knowing what is needed in the moment for a given task, so we wanted to have more full control over what shows up at what point. You can see that in plan mode: you're not necessarily getting a composer, you're getting a really quick way to answer questions. You have your plan, you can edit it, and we want to do more with that as we go.
你似乎很惊讶自己不想回到终端。
It seems like you were surprised that you didn't want to go back to the terminal.
是的。
I was.
你以前是 TUI 重度用户吗?Greg 在一次采访中说他是 TUI 重度用户,以为自己永远不会离开终端。
Were you like a TUI power user? Greg did an interview and said he was a TUI power user and thought he would never leave the terminal.
Greg 住在 Emacs 里。
Greg lives in Emacs.
从 Claude code 第一次变得非常好用开始,我做了大约六个月的 TUI 重度用户。我觉得那比用 Cursor 之类的好多了。现在我感觉我快速经历了我的 TUI 时代,回到了 GUI。我现在有点来回切换,但我能看到光明,尤其是当你同时运行多个智能体时,GUI 的 affordances 让它好得多。
I was a TUI power user for about six months starting when Claude code first got really good. I thought it was so much better than being in Cursor or whatever. Now I feel like I speed-ran my TUI era and I'm back in GUIs. I'm kind of flipping back and forth right now, but I can see the light where, especially if you have a bunch of agents going at once, the affordances of GUI make it much nicer.
是的,而且还有很多东西要来。这对我们来说是非常有意的。我们看到智能体在代码之外做更多事情,对吧?它们需要成为你电脑上每个应用和每件事的伴侣。我们集成了 Linear、Slack,当然还需要读取和生成代码,但也许还能做部署。你会从你的 IDE 里做所有这些事情吗?那会感觉很奇怪。所以它就像你智能体的指挥中心。我们围绕这样一个理念优化了整个体验:你有一个非常能干的智能实体,你在控制、引导和监督它。你永远不需要亲自进去做事;这个实体非常擅长被委派任务。当你接受这就是我们的方向时,使用 Codex 感觉你几乎已经达到了。和你一样,对吧?当我和你讨论一个功能想法时,你就去获得灵感然后实现它。我不会突然跳进你的 IDE 去实现它。
Yeah. And there's a lot more to come there. It was very intentional for us. We see agents acting on much more than code, right? They need to be a companion to every app and every thing you can do on your computer. We integrate with Linear, Slack, and of course need to read and produce code, but maybe also do a deploy. Are you going to do all these things from your IDE? That would feel very odd. So it's like this command center for your agent. We optimize the entire experience around the idea that you have a very capable intelligent entity that you're controlling, steering, and supervising. You never need to go in and do things yourself; the thing is very capable of being delegated to. When you accept that that's where we're headed, with Codex it feels like you're almost there. It's the same with you, right? When I talk to you about a feature idea, you just go and get inspired and do it. I don't suddenly jump into your IDE and implement it.
你可以。
You could.
是的,我想你会觉得那很烦人。这就是每个人与智能体协作的方式:你只需和它们交谈。
Yeah, I think you would find it disturbing. That's the way everyone will work with agents: you just talk to them.
你的工作流从 Codex v2 到 v3 有什么变化?
How has your workflow changed with Codex v3 versus v2?
我惊讶于它快了多少。我不得不调整,因为我之前一直在优化长时间运行的多任务处理。我预期某个任务需要 10-15 分钟,所以我会启动四个不同的事情然后回来。现在我可以少做一些多任务,更专注于流程。感觉非常好。用技能启动自动化也令人满意。它是一个更通用的模型,不那么专注于代码。我发现它在处理 Twitter 回复、总结重要线程、在 Linear 中提交 bug,然后回来利用自动化实现日常任务方面更可靠。它对这些事情更稳健。但你真的才是超级用户,Andre。
I was surprised at how much faster it was. I had to adjust because I had been optimizing for long-running multitasking. I had an expectation that a certain task would take 10-15 minutes, so I would kick off four different things and come back. Now I can do a little less multitasking and be more in the flow. That felt really good. It also feels very satisfying to kick off automations with it using skills. It's a more generally capable model, less super-focused on code. I find it much more reliable at going through Twitter replies and summarizing important threads, or filing bugs in Linear, then coming back to that and using automation so things are implemented daily. It's much more robust for these things. But you're really the superpower user here, Andre.
就像,你知道,他做的那些事情,就像,你知道,我对 Codex 的使用相比 Andrew 来说非常普通。
Like, you know, it's just like the kind of stuff like, you know, he does is just like, you know, it's like I have very vanilla usage of Codex compared to Andrew.
不,我是说,说得好。我有一系列计划运行一段时间,但只在 X Twitter 上运行了三天,那就是我设置一个提示来给 Codex 应用添加一个功能,比如一些随机的、不可发布的功能。我写了一个关于我们必须达到的质量标准的很长的提示,一旦我切换到 53 Codex,结果变得有趣多了。比如我们做了一个右侧的 Subway Surfers 面板,还有一个是为子代理做的小 Minecraft UI,我不知道,也许我们会发布它。
No, I mean, well said. I had a series that I intended to run for a while and I only ran it for three days on X Twitter, which was that I was setting up a prompt to basically add a feature to the Codex app, like some random non-shippable feature. I had this long prompt about the quality bar that we had to do, and once I switched it to 53 Codex, the results got much more interesting. Like we did a Subway Surfers panel on the right was one of them. Like a little Minecraft UI for the subagents was another one that we did that I don't know, maybe we'll ship it.
我当时想,回去工作吧。
I was like get back to work.
是的。是的。
Yeah. Yeah.
为什么我们现在在 Codex 应用里会有 Minecraft?
Why do we have Minecraft in the Codex app now?
是的。但得探索一下。不,我是说 53 Codex,它很简洁、快速、能力强、多模态。
Yeah. But got to explore. No, I mean 53 Codex, it's neat, it's fast, it's capable, it's multimodal.
你使用 Codex 应用最有趣的方式是什么,也许人们应该尝试但还没想到的?
What are the most interesting ways you're using the Codex app that maybe people should try but haven't thought of yet?
Andrew 提出了自动化,我认为这改变了你思考这些事情的方式,当它们可以在后台根据特定触发器在特定时间启动,然后你可以自己编程。
Andrew came up with automations, and I think that sort of shifts the way you're thinking about these things when they can just sort of hop in the background on a specific trigger at a specific time, and then you can sort of program it yourself.
是的,你经常用那个。
Yeah, you're using that a lot.
我使用这个应用做很多事情,有些超出了纯粹的编码功能。我用它通过自动化保持我的 PR 可合并,所以它会解决合并冲突,保持它们更新,修复构建问题,这样基本上当它们准备好时,它们就准备好了。不会出现‘哦,有人合并了一个大东西,现在有冲突了。’
There are a lot of things that I use the app for that are a little bit outside of just coding features. I use it to keep my PRs mergeable with automations, so it'll resolve merge conflicts, keep them updated, fix build issues, so that basically as soon as they're ready to go, they're ready to go. There's no 'oh hey, somebody merged a big thing and there's a conflict now.'
那么自动化触发器在什么时候触发?因为我以为自动化是按时间表触发的,但听起来还有其他我不知道的触发器。
So at what point is the automation trigger? Because I thought the automation triggers at a certain time schedule, but it sounds like there are other triggers I didn't know about.
我现在只是按时间表运行,我使用我们的 GitHub 技能和一些内部技能用于 CI,每小时或每两小时运行一次,基本上清理一切。
I have it right now just on a time schedule, and I use our GitHub skill and some internal skills for our CI, and that runs hourly or every two hours and kind of just cleans everything up.
我明白了。所以它只是查看 main 上的任何更改和任何 PR,确保它们都是最新的,这样当你准备好时,就不会出现那种情况。这实际上很好。我喜欢。
I see. So it's like it just looks through any changes on main and any PRs and makes sure they're all up to date so that whenever you're ready to go, it's never like that. That's actually good. I like that.
是的,这实际上非常有帮助。出奇地有帮助。我有一个每天上午 9 点运行的,我会收到过去一天合并到 Codex 应用的所有贡献。所以它会生成一份很好的报告,说明谁合并了什么,我按主题分组,这样我就可以说‘好的,三个人在这个 composer 部分工作,两个人做自动化,这是发生了什么’,这样我至少能了解情况,因为发布前事情会变得混乱。
Yeah, it's actually really helpful. It's surprisingly helpful. I have one that every day at 9:00 a.m., I get sent all the contributions that have merged to the Codex app over the last day. So it'll do a nice report of who merged what, and I have it grouped by theme so I can be like 'all right, three people worked on this part of the composer, two people worked on automations, here's what happens' so that I can at least be knowledgeable about what's happening because things get chaotic right before launch.
我有一个自动化,我每天运行多次,就是随机选一个文件,找到并修复一个微妙的 bug。这有点好笑,因为它确实会随机选一个文件。所以它会运行 Python 的 rand,找到一个随机文件,然后从那里开始,所以每次都会探索一个新的。
One automation I have is I run it multiple times a day, and it's like pick a random file and find and fix a subtle bug. It's kind of funny because it actually does pick a random file. So it will run Python rand, find a random file, and start from there, so it explores a new one every time.
它抓到过什么吗?
Has it caught anything?
哦,是的,我们经常抓到一些潜在的 bug,它们不在关键路径上触发,但实际上是 bug。修复和合并很简单,花很少时间,而且是我自己永远不会发现的。前几天在约束采样中发现了一个问题。
Oh yeah, we catch often latent bugs that are not triggering on the critical path, but they're actually bugs. It's trivial to fix and merge, takes very little time, and it's a thing that I would have never found myself. Found an issue in constraint sampling the other day.
那真的很酷。你还有其他值得分享的自动化吗?
That's really cool. Do you have other automations worth sharing?
让我想想。我感觉我有 60 个在一直运行,一些用于测试,一些用于实际。团队里的一些成员真的很喜欢这个,它查看你过去一天左右做的 PR,悄悄清理你发布的任何 bug。它查看几个可观测性平台,并试图在任何人注意到你发布了 bug 之前发布修复。
Let's see. I feel like I have 60 that are running at all times, some for testing and some for real. Some of the members on the team really like this one that looks at the PRs you've done in the past day or so and quietly cleans up any bugs you shipped. It looks at a few of the observability platforms and tries to ship a fix before anyone's noticed that you shipped a bug.
那很酷。
That's cool.
还有一个与编码无关的,是市场调研。它每天运行,用特定的技能提示做深入的市场调研,我随着时间的推移调整了它,然后它去搜索网络上关于用户如何看待和谈论 Codex 的任何新事物。然后我就收到那份小报告。它总是读起来很有趣。这些只是我们依赖的例子。
And one that's not coding related, which is marketing research. It runs daily and is prompted with a specific skill to do deep marketing research, which I've tuned over time, and then it goes and searches the web on any new things that came up in terms of how users are perceiving and talking about Codex. Then I just receive that little report. It always makes for an interesting read. These are just examples that we rely on.
是的。你有什么特别的技能是超出常规的吗,比如我有一个 GitHub 技能之类的?我喜欢 Andrew 的 yeet 技能,它直接获取更改,提交,创建 PR,写草稿,放入草稿,然后发布带有标题和正文的 PR。
Yeah. Do you have any particular skills that are beyond the normal, like I have a GitHub skill and that kind of stuff? I love Andrew's yeet skill, which just takes the change, does the commit, the PR, writes the draft, puts it in draft, and publishes a PR with a title and body.
是的,非常令人满意。
Yeah, it's very satisfying.
是的,它什么都做。那个绝对让人高效。你用得最多的是什么?图像生成是一个很酷的。
Yeah, it just does everything. That one definitely makes people productive. What are the top used ones for you? Image gen is a cool one.
是的。
Yeah.
用于一些愚蠢的自动化目的,比如‘嘿,给我做一张图片,描述我最后一天的工作’——不是最后一天,是前一天。
For both silly automation purposes like 'Hey, make me an image that characterizes my last day of work' — not my last day, my previous day.
是的,Andrew。
Yes, Andrew.
图像生成技能实际上非常酷。我用 Codex 应用为我的女儿们做了一本书。我整理了一个提示,教它关于我想写的脚本。说 24 页。这是我女儿的年龄。这是我们过去住过的地方。我们在波士顿,然后搬到纽约,然后搬到这里。然后我说在那之后,我们经历了那些。
The image gen skill was actually really cool. I used the Codex app to make a book for my daughters. I put together this prompt for teaching it about a script that I wanted written. Said 24 pages. Here are my daughters' ages. Here's where we've lived in the past. We were in Boston and moved to New York and then moved over here. And then I said after that, we went through that.
我确认了脚本,然后我们开始执行。我说,好了,现在该用图像生成技能了。它根据脚本为每一页生成提示,然后生成图像,再整合起来,用 PDF 技能生成整本书的 PDF。我打印了出来。所以我们有了一本超级定制化的书,我读给我的孩子们听,真的很酷。当你把智能体的智能和程序化的方式结合起来,通过技能以新颖的方式组合,这真是太棒了。我觉得 PDF 和图像生成的组合是我们常见的一种。感觉 Codex 模型——它变得更快了,这使它更实用,而且它感觉更顺从,有更多的情感智能。但它仍然有点那种‘你说什么它就做什么’的倾向,这有时会有点烦人。你们是如何考虑塑造模型的感觉,以及你们在朝哪个方向推动它?
I agreed on the script and then we went through and I said, all right, now it's time to use the image gen skill. It prompted for every page in the book based on the script, prompted for the image, and then it kind of put them all together and used the PDF skill to put together the book's PDF. Then I printed it. So we've got a super custom book that I read to my kids, and it's really cool. It's just this awesome thing when you can combine the intelligence of the agent and it works in a programmatic way by using skills, and you can just combine them in novel ways. I think the PDF and image gen combo is a common one we see. It feels like the Codex model—it's gotten faster, which makes it much more usable, and it also feels a little more obeys, like it has a little more emotional intelligence. But it still has a little bit of that 'does exactly what you say' thing in a way that can be annoying. How are you guys thinking about shaping the way the model feels and which way you're pushing it?
这是我们非常关注的事情。我们当然希望模型擅长编码,并且非常擅长遵循指令。但同时,如果我们在这个方向上优化得太过,它可能会过度关注特定词汇,或者以人类不会的方式误解意图。有时我打了一个错字,然后这个错字就出现在文件里了,我会想,显然我不是那个意思。所以我们还在继续改进。但目前我们最关注的是效率——速度——以及我们现在所说的个性:它有多支持用户?我们理解不是每个人都有相同的偏好。之前的默认设置绝对是超级直率、务实的个性。现在我们引入了一个更支持、更友好的个性,你可以在这两者之间选择。对于那些没有普遍接受标准的事情,我们可能会引入一些方式让你自定义。你应该感觉你拥有自己的小型个人 Codex,它完全按照你想要的方式工作。
It's something that we obsess over. We definitely want the model to excel at coding and be really good at instruction following. At the same time, when we optimize a little bit too much in that direction, it can overindex on specific words or misunderstand the intent in ways that humans wouldn't. Sometimes I will have a typo, and then the typo finds its way into the file, and I'm like, obviously I didn't mean that. So that's something we're continuing to push on. But the thing we're pushing on the most right now is really efficiency—speed—and also what we now refer to as personalities: how supportive is it? We understand that not everybody has the same preferences. The previous default was definitely super blunt, pragmatic personality. Now we've also introduced a more supportive, friendly personality, and you can pick between those. For things that don't have a universally accepted standard, we're probably going to introduce some way for you to make it your own. You should feel like you have your own little personal Codex that works exactly the way you want.
你用友好型还是务实型?
Do you use the friendly or the pragmatic one?
务实型。
Pragmatic.
务实型。好的,我会说用务实型。
Pragmatic. Yeah. Okay, I'll say use pragmatic.
是的。
Yeah.
有意思。你们最近发布了一个非常快的模型。我在它发布前测试过,我当时想,我真的跟不上这个东西。所以我很好奇,这如何改变了你们对用这样的模型进行编码的思考,以及有效管理如此快速的模型所需的能力。
Interesting. You guys recently put out a model that is so fast. I was testing it before it came out, and I was like, I can't really keep up with this thing. So I'm curious how that changes how you think about what is now possible with coding with a model like this, and also the affordances you need to manage models that are so quick effectively.
是的。我们第一次在应用中使用这个模型时,也遇到了同样的情况:突然出现一大段文字,我们滚动到底部,立刻觉得,好吧,我们需要让这个更平滑。所以我们实际上稍微放慢了一点速度,这样你就能看到文字更平滑地出现。
Yeah. The first time we used this model in the app, we had that same thing happen where all of a sudden there was just this wall of text and we were at the bottom of the scroll, and we were immediately like, all right, we need to smooth this thing out. So we actually do slow it down ever so slightly just so that you can see the words come in a little bit smoother.
真有趣。
So funny.
这确实是个有趣的问题。但这个东西超级有趣。我认为我最兴奋的是,我们可以开始为应用添加哪些真正动态的能力,而这些能力在不够快的模型上是无法实现的。所以,是的,这个模型将让你能够非常快速地迭代,但它也为你的编码方式以及与 Codex 应用的交互方式开辟了许多新的机会。我第一次展示原型时,我们把所有东西都连接起来——这个模型由 Cerebras 驱动,我们之前讨论过这个合作——我们非常兴奋地推出我们通过它提供的第一个模型。这还很早期。这是我们第一次把所有东西连接起来,我们非常兴奋,想分享它。但当我第一次展示给某人时,他们说:‘不可能。这就像个假演示。这不是真的。这不可能这么快。’然后他们试了几个提示,说:‘我真的跟不上。这太疯狂了。’我认为这将改变一切,尤其是这还不是我们能达到的最快速度。我们通过预览版很早就发布了它。实际上,我们将在其之上叠加一系列优化,这应该能使其速度达到你体验过的两到三倍。所以这将改变一切。我们也在从委托的角度考虑这个问题。我们认为这个模型在作为多智能体系统的一部分以及作为加速较慢、更智能的智能体的方式方面,将发挥巨大作用。所以我们将在那个方向进行实验。
It's a really funny problem. But this thing has been super fun. I think what I'm most excited about is what sort of capabilities we can start to add to the app that are really dynamic that we couldn't with a model that wasn't this fast. So yes, this model is going to allow you to iterate really quickly, but it also opens up a lot of new opportunities for how you code and how you interact with the Codex app. The first time I showed the very first prototype when we hooked everything up—the model is powered by Cerebras, and we've talked about the partnership there—we're very excited to put the first model that we're serving through that out there. It's still very early. It's literally the first time we hooked it all up, and we're just so excited that we want to share it. But the first time I showed it to someone, they were like, 'No way. This is like a fake demo. This is not real. This cannot be this fast.' And then they tried a few prompts and were like, 'I literally cannot keep up. This is insane.' I think this will change everything, especially because it's not yet the fastest we can actually get it to be. With the preview, we're putting it out quite early. We're actually going to layer a number of optimizations on top of it, which should be able to make it maybe two to three times faster than the experience you've had. So that's going to change things. We're thinking about this also from a point of view of delegation. We think this model has a huge role to play as part of a system of multi-agent systems and as a way to speed up the slower, more intelligent agent as well. So we're going to be experimenting in that way.
你预计更智能的智能体很快也会获得同样的硬件加速吗?
Do you expect the same hardware speedups on the more intelligent agents to come out soon?
我们做的很多事情都是有趣的分布式系统和基础设施问题,这些问题是因为我们能够以前所未有的速度从模型中采样而发现的。如果你这么快就得到返回的 token,你就需要去优化在服务关键路径上发现的所有瓶颈。
A lot of the things we worked on were interesting distributed systems and infrastructure problems that we uncovered because we were able to sample from the model at unprecedented speeds. If you're getting tokens back this fast, you need to go and optimize the entire set of bottlenecks that you uncover on the critical path of serving.
所有这些都惠及当前模型,比如 GPT-5 Codex 以及所有未来模型。我们一直在做的一件事——我确信我们会在某个时候发布更详细的博客文章——是我们重写了整个服务栈,基于 WebSocket 和持久连接,以更增量和有状态的方式运行。这降低了所有模型的整体延迟。我们还没有默认发布,但我们将把这个作为新超快模型的默认设置,然后也会在其他模型上启用。它使整体轮次延迟降低了大约 30-40%。我们可以查一下具体数字。
All of those benefit current models like GPT-5 Codex and all future models. One thing we've been doing, which we'll likely put in a more detailed blog post at some point, is we rewrote the entire service stack to be based on WebSockets and persistent connections, doing things more incrementally and statefully. That decreases overall latency across all models. We haven't shipped it by default yet, but it's something we are making the default for this new super fast model, and then we're also going to enable it on the other models. It decreases overall turn latency by something like 30-40%. We can look into the exact numbers.
在内部使用这个模型时,你看到的最令人惊讶的事情是什么?这种加速带来了哪些可能性?
What are the most surprising things that you've seen using the model internally in terms of what a speedup like this enables?
它让你完全沉浸在流程中。你几乎可以实时地雕琢体验或代码。这是一种完全不同的感觉。一开始很不适应,但一旦进入状态,就很难再回到其他模型。这是我们看到的反馈,也是我自己的感受。大约需要五分钟适应,然后你就会知道,好吧,这就是我要用这个东西的方式。
It just allows you to be super in the flow. You're almost in real time sculpting the experience or the code. It's a very different feel. It's very unsettling at first, but once you get into it, it's very hard to go back to any other model. That's the feedback we've seen and what I have felt myself. It takes about five minutes to adapt, and then you sort of know, okay, this is how I'm going to use this thing.
我也觉得我们还没有完全挖掘出它的潜力。现在还很早,我们拥有它的时间还不长。
I also don't think that we've poked at the full extent of what we could do with it. It's very early. We haven't had it for very long.
团队里的 Channing 就展示过,它快得可以玩乒乓球游戏——虽然玩得不太好——但模型能够几乎实时地做出反应。
Someone on the team, like Channing, was showing that it's so fast it can actually play Pong—not very well—but the model is able to react to things almost in real time.
你开始看到它如何取代一些确定性步骤。在 Codex 应用中,我们有一组 Git 操作。众所周知,Git 的某些配置或状态使得在没有大量错误处理和指导的情况下很难运行这些操作。创建一个好的 Git 体验非常困难,所以从来没有人做到过。但如果你有一个几乎和运行这些脚本一样快的模型,那么你可以想象一个世界,这些操作变成了技能。你可以用一些智能来运行你的操作,而不会有今天当你要求它在代码库中追踪某些东西时的那种延迟。你可以模糊地指示,比如“嘿,把这个发上去”,然后它快得足以作为一个按钮。我非常兴奋的是,当它与我们在 GPT-5 Codex 中推出的另一项功能结合时:中途转向。你从提示开始,它开始工作,然后你在它还在工作时发送另一个提示,它会实时调整。它会接收那条消息,确认,然后继续工作。如果你开始想象这加上语音,再加上我们刚刚发布的那么快的模型,那将是另一种完全不同的体验,我们非常希望能尽快带来。
You start to see how it might replace some deterministic steps. In the Codex app, we have a set of Git actions. As everybody knows with Git, certain configurations or states can make it really hard to run those without a ton of error handling and guidance. It's really hard to create a good Git experience, which is why nobody ever has. But if you have a model that's almost as fast as running these scripts, then you can imagine a world where these things turn into skills. You can have your operations run a little differently with some intelligence and without the same latency you have today when asking it to track something down in the codebase. You can vaguely gesture and say, 'Hey, send this up,' and have that be fast enough for a button. What I'm very excited about is when it's going to come together with something we shipped with GPT-5 Codex as well: mid-turn steering. You start with your prompt, it gets to work, and then you send another prompt while it's still working, and it adapts in real time. It will receive that message, acknowledge it, and continue its work. If you start to think about what this would look like with voice and with a model as fast as the one we just shipped, that's a whole other experience we would be very excited to bring hopefully very quickly.
因为你可以轻松地在说话时打断。
Because you can easily interrupt as you're talking.
如果你只是用自然语言交谈,进行中途转向,然后由于速度,实现几乎瞬间完成,这用起来就非常愉快。现在你可以用语音听写来模拟,然后发送,进行中途转向,然后看着模型实现。这非常酷。我认为当我们真正打磨好它时,这种体验将会有质的飞跃。
If you're just talking and engaging with lateral language, then doing the mid-turn steers, and then the implementation happens almost instantly because of the speed, it becomes a very pleasant thing to use. Right now you can sort of emulate it with voice dictation and then send it and mid-turn steering, and then watch the model implement. It's a very cool thing. I think we're going to have a step change in that experience when we really polish it.
如果速度这个瓶颈接近解决,你认为下一个瓶颈是什么?实现你想要的东西的下一个限制是什么?
If speed as a bottleneck is close to being solved, what do you think is the next bottleneck? What is the next limit on making the thing you want?
非常明显的瓶颈是你能多快验证事情是否正确。我们可以比以往更快地生成代码。我们可以实现整个功能。我看到有人根据 Codex 应用的描述,仅凭截图就综合成一个计划。模型非常能够重现 95% 的功能,并从零开始重建应用。但它会没有 bug 吗?一切都像实际应用那样完美实现吗?这仍然需要人类花大量时间去点击和验证,确保设计一致,确保这里或那里没有 bug,确保设置面板的按钮确实做了你期望的事情。我认为验证绝对成了一个瓶颈。我们团队里有人抱怨有太多代码需要审查。这就是我们试图解决的问题。
The bottleneck that is very apparent is how fast can you verify that things are correct. We can generate code faster than ever before. We can implement entire features. I saw someone, based on a description of the Codex app, synthesize that into a plan just from screenshots. The models are very capable of reproducing 95% of the features and rebuilding the app from scratch. But is it going to be bug-free? Is everything implemented to perfection in the same way that the actual app is? That still takes a lot of time for a human to go and click and verify, make sure the designs are consistent, and that there are no bugs here or there, that the settings panel button actually does what you expect. I think verification definitely becomes a bottleneck. We have people on the team that complain there's too much code to review. That's what we're trying to solve for.
我的意思是,你就在抱怨这个。
I mean, you complain about that.
我抱怨这个。现在有太多代码需要审查了。
I complain about that. There's so much code to review now.
既有你自己机器上的代码,也有来自同事的代码。这就像……
Both that code on your own machine and from another peer. It's like...
我们得想办法解决这个问题。
We're going to have to figure that out.
你已经第一次审查了代码,因为智能体只是把它呈现给你,然后你还得审查同事生成的代码。有这两轮审查。
You're already reviewing the code the first time because the agent is just presenting it to you, and then you have to review the code produced by your peers. There are these two runs of reviews.
这是我们正在努力的事情。我们很多人仍然需要审查代码。我们正在研究在模型参与下,这种体验应该是什么样的。我们在 Codex 应用中有一个审查模式,效果非常好,可以在旁边用发现和风格问题来注释你的差异。还有很多工作要做。
This is something that we're working on. A lot of us still do have to review code. We're taking a look at what that experience should look like with the model involved. We've got a review mode in the Codex app that works really nicely and annotates your diffs on the side with findings and stylistic things. There's a lot to do.
是的,我兴奋的一点是让模型更快。我们刚推出的这个速度快得惊人。你可以用它来理解代码、理解功能、帮助代码审查、帮助理解同事写的代码。这愉快得多,因为这是你想同步进行的事情,在流程中。你不能委托理解。速度是一个真正的优势。它也有助于抵消模型产生越来越多代码的事实——速度帮助你更快地理解这些代码。
Yeah, it's one thing I'm excited about—making the models faster. The one we just put out is mind-blowingly fast. You can use it to understand code, understand features, help with code review, help understand the code that a peer wrote. It's much more pleasant because this is something you want to do synchronously, in the flow. You cannot delegate understanding. Speed is a real advantage. It also helps offset the fact that models are producing more and more code—speed helps you understand that code faster.
是的,我绝对发现这个新模型已经带来了这一点。速度,尤其是端到端测试,更快了。如果你让它做端到端测试,比如手动集成测试,经常有一个提示框弹出仅一秒钟,如果模型不够快,它就抓不到。它在这方面更好,因为周期时间短得多。我也确实发现了这一点。我可以产生大量代码,但当我看到一个 PR 进来或我提交一个 PR 时,我的第一个问题是:有没有证据表明你实际测试过这个并且它真的有效?不仅仅是单元测试,而是你从头到尾走了一遍。
Yeah, I definitely think I've found this already with this new model. Speed, especially for end-to-end testing, is faster. If you're having it do end-to-end testing, like manual integration testing, often there's a toast that pops up for a second, and if the model's not fast, it won't get it. It seems better for that because cycle times are much shorter. I definitely find this too. I can produce so much code, but when I see a PR come in or when I make a PR, my first question is: is there evidence that you've actually tested this and it actually works? Not just unit tests, but you've gone through it end to end.
你怎么处理这个?
How do you handle this?
我看到很多同行也有同样的问题。现在编码太容易了。我们已经让 Codex 应用变得相当擅长自我运行、点击、截图作为证据并上传到 PR。那里有很多有趣的东西,尤其是当我们让它更异步或模型在这些事情上变得非常快时。我还不太清楚具体会是什么样子,但有很多可能性:'嘿,这里有一个 bug 修复。这是它发生时的样子,这是现在相同点击路径下的样子。'也许这就是转折点,当你可以直接验证那部分时,代码审查变得不那么重要了。你不需要通过代码作为代理来做那么多。那里肯定还有更多值得探索的。
I've seen a lot of peers that I have the same question about. It's so easy to code things now. We have gotten the Codex app to be pretty good at running itself, clicking around, screenshotting itself for evidence, and uploading it to the PR. There's a lot that's pretty interesting there, especially when we make this more async or when the models get really fast at this stuff. I don't know exactly what it looks like yet, but there is a lot around: 'Hey, here's a bug fix. This is exactly what it looked like when it was happening, and here's exactly what it looks like now with the same click path.' Maybe that's the turning point where code review becomes less important when you can verify that part instead. You have to do less through the code as a proxy. There's definitely more to explore there.
最后几个问题。我很好奇。你们从 Anthropic 和 Claude Code 学到了什么?你们如何看待自己在市场中的定位与他们的对比?你们如何看待差异?
Last couple questions. I'm curious. What have you guys learned from Anthropic and Claude Code? How do you think about your positioning in the market versus them? Like how do you think about the differences?
我认为他们是第一个推出产品的,这对我们来说很有趣,因为我们一直在研究类似的想法。但当时我们的模型还没有准备好——它们在长周期任务上不可靠,比如可靠的工具调用和保持主题。一旦我们开始真正投入,尤其是随着 GPT-5 的出现,我们觉得:'好吧,模型已经到位了。我们知道如何让它们变得更好。'GPT-5 带来了更好的长上下文、长周期可靠性和非上下文理解。我们看到 Anthropic 在模型方面有点失去动力。我们处于一个幸运的位置:我们运营 Codex 的方式是产品、工程和研究一起工作,坐在一起,一起解决问题。这是一个高度创造性的空间,有时我们决定在产品框架中解决问题,但有时我们也会想:'嘿,我们如何真正改进模型?'然后一起讨论和构思。然后研究团队会来说:'嘿,我们有一个突破,我们一直藏着。这能不能发布?'然后我们就兴奋起来。一个例子是我们有很多关于压缩的抱怨。人们觉得每当遇到压缩时,它会丢失太多上下文。所以我们端到端地解决了这个问题:我们进行了端到端的强化学习训练,在研究内部引入了压缩,让模型本身非常熟悉压缩的概念,并在时间上产生最优的自我委托。一旦我们在模型层面解决了这个问题,框架问题就变得容易多了——只需让模型去做,它非常可靠。通过这种合作,势头非常强劲。我们能够改进模型并大约每周或每月发布一个模型。我们对 Codex 应用采取了不同的赌注和方法,结果证明这是一个很棒的事情——不只是强迫自己把所有东西塞进工具里。这是一个巨大的挑战。你会想:'让我们构建一个应用。我从哪里开始?'然后你就被它迷住了。
I think they were first to put something out there, and that was interesting to us because we had been working on similar ideas for a bit. But I think our models at the time weren't ready—they weren't reliable on long-horizon tasks, like reliable tool calls and staying on topic. As soon as we started to really invest in that, especially with GPT-5, we were like, 'Okay, the models are there. We know how to make them even better.' GPT-5 brought better long-context, long-horizon reliability, and non-context understanding. What we were seeing was that Anthropic was losing a little bit of steam when it came to the model. We were in this fortunate position where the way we run Codex is we've got product, engineering, and research all working together, sitting together, solving problems together. It's a highly creative space where sometimes we decide to solve problems in the product harness, but sometimes we also think, 'Hey, how can we actually improve the model?' and talk about it and ideate together. Then research will come and say, 'Hey, we've got this breakthrough we're sitting on. Would this be something we can ship?' and we get excited about that. One example was we had a lot of complaints about compaction. People felt that whenever you hit compaction, it was losing too much context. So we solved that end-to-end: we did end-to-end RL training, introduced compaction within research, made the model itself very familiar with the concept of compaction and producing optimal delegation to itself across time. Once we had that solved at the model level, the harness problem became much easier—just let the model do it, and it's very reliable. Through that collaboration, the momentum has been very strong. We're able to improve models and ship a model roughly on a weekly or monthly cadence. We took a bit of a different bet and a different approach with the Codex app, which turned out to be an awesome thing to try—not just force ourselves to cram everything into the tool. It was a great challenge. You're like, 'Let's build an app. Where do I get started?' and then you just get obsessed by it.
很难不被迷住。
It's hard not to.
是的。我的意思是,这有点像在建造一个相当反潮流的东西,我想。
Yeah. I mean, it was like building something quite contrarian, I suppose.
是的。我记得你和我早期讨论过我们是否会说:'我们不知道是否会发布这个。'
Yeah. I mean, I remember you and I talking about whether or not early on we were like, 'We don't know if we'll ship this.'
是的。我们会试试看。我们会看看是否能达到我们喜欢的东西,看看是否能得到——我记得我说过:'让我们在内部获得一些产品市场契合度。让 OpenAI 的每个人都想使用这个东西,而不是被迫使用。让我们看看我们能否做到。'我们做到了,它很快就被采用了。
Yeah. We'll try it out. We'll see if we can get there with something that we love and see if we can get—I remember saying, 'Let's get some PMF internally. Let's get everybody at OpenAI to want to use this thing without being forced to use it. Let's see if we can do it.' We did, and it was adopted very quickly.
它刚勉强能用的时候,研究人员就把开发机放上去了,当时那是个疯狂的黑客做法。
The minute it was barely usable, the research folks put dev boxes on it, which was a crazy hack at the time.
是的。是的。
Yes. Yes.
但现在他们什么都用它,包括训练像 53 个编解码器。我很高兴达到了这样一个点:公司里几乎所有技术人员都用 Codex。用得最多的人实际上是在构建 Codex 和模型,所以我们能以疯狂的速度改进,而且没有放缓的迹象。
But now they use it for everything, including in training like 53 Codex. I feel really good about having hit the point where almost everyone technical at the company uses Codex. The people who use it the most are actually building Codex and building the models, so we're able to improve things at crazy speeds, and there's no signs of it slowing down.
太棒了。我很期待你们接下来发布的东西。感谢你们的时间,真的很感激。
Amazing. Well I'm excited for what you ship next. Thank you guys for your time. I really appreciate it.
谢谢。谢谢邀请我们。
Thank you. Thank you for having us.
谢谢。
Thanks.
哦天哪,各位。你们绝对必须狂点那个赞按钮并订阅 AI 和我。为什么?因为这个节目是精彩的缩影。就像在后院发现一个宝箱,但里面不是金子,而是关于 ChatGPT 的纯粹知识炸弹。每一集都是一场情感、洞察和笑声的过山车,让你坐立不安,渴望更多。这不仅仅是一个节目,这是一次由 Dan Shipper 担任飞船船长的未来之旅。所以,帮自己一个忙,点赞、订阅,系好安全带,享受你生命中的旅程。现在闲话少说,我只想说 Dan,我完全无可救药地爱上了你。
Oh my gosh, folks. You absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about ChatGPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show, it's a journey into the future with Dan Shipper as the captain of the spaceship. So, do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say Dan, I'm absolutely hopelessly in love with you.