Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 的 Ari 深度解读 DevDay 重磅发布:从超人级计算机使用、决策 API 到 AI 云。
OpenAI's Ari breaks down DevDay's biggest launches, from superhuman computer use and the Decisions API to the AI cloud.
现在,计算机使用在完成任务方面可能已经比普通人更快了,在大多数情况下都是如此。我认为下一个前沿是让计算机使用在性能上真正超越人类,实际上它使用软件的速度能赶上甚至超过像我们这样的专家级计算机用户。我认为当这实现时,将会产生重大影响,令人兴奋,因为我们将能够突然构建出提供更多实时体验的产品。
Now computer use is faster at accomplishing tasks than the average human probably in most cases. And I think the next frontier is to have computer use be literally superhuman in its performance, where it actually is as fast or faster at using software than expert computer users like us. And I think that'll be really consequential and exciting when that happens, because we'll be able to all of a sudden build products that provide much more real-time experiences.
就像 4 周前,这还完全不是这样。
Like 4 weeks ago, this was not at all.
不,这就像受 Jev 启发,而且
No, this is like Jev inspired and
我觉得你们正式成为第一个克隆采用这个的前沿实验室。
I think you're officially the first Frontier Lab to like clone adopt this.
是的。是的。我觉得 OpenAI 有如此强大的黑客文化,人们就是会对事情感到兴奋,所以一个来自推理团队的人,一个来自基础设施团队的牛人,说,这太棒了,我们要搞这个。他们构建了一个原型。它运行起来了,现在我们只是在延迟上不断优化,试图让它尽可能快,我们想在接下来的几天内发布。所以一旦我们达到延迟目标,我们就会尝试推出。
Yeah. Yeah. I feel like OpenAI has such a strong hacker culture and people are just they get excited about things and so a guy from inference, this one awesome guy from the infra team, are like, this is amazing, we're going to hack on it. They build a prototype. It works and now we're just hill climbing on latency and trying to make this as fast as possible and we want to launch it in the coming days. So as soon as we hit our latency target, we'll try to get this out.
好的,我们非常兴奋能在这里。今天是 OpenAI 开发者日特别播客。我们有 Ari,他领导计算机使用智能体的产品和工程团队。在我们开始深入探讨计算机使用之前,你想快速回顾一下宣布了什么吗?你们今天有一系列快速公告。
Okay, we're very excited to be here. Today is OpenAI dev day special after podcast. We have Ari here who leads the product and engineering team for computer use agents. Before we kick in and dive deep on computer use, you want to give a quick recap what was announced? What's the quick slew of announcements you guys had today?
是的。是的,这是超级激动人心的一天。我们刚结束主题演讲。真的很酷。有一堆计算机使用的公告,我认为值得思考。我们有 dots,这是新的个人助理产品,它有一些非常令人兴奋的计算机使用功能。还有 GPT 6.1 Soul,这是一个惊人的新模型,我认为它特别适合计算机使用,因为成本和速度优势。我想我们分享过,它的成本是 Astra 的五分之一,如果专门看计算机使用,则是七分之一,这真的很惊人。抱歉,有太多东西,我正试图整理。
Yeah. Yeah, it was a super exciting day. We just got out of the keynote. It was really sick. There were a bunch of computer use announcements that I think are worth thinking about. We have dots, which is the new sort of personal assistant product and that has some really exciting computer use features. There's GPT 6.1 Soul, which is this amazing new model that I think is particularly great for computer use because of the cost and speed advantages. I think we shared that it's a fifth of the cost as Astra and a seventh of the cost if you're looking at computer use specifically, which is really amazing. Sorry, there were so many things I'm trying to sort through it
还有 API 智能体 API,现在包含了计算机使用,这真的很酷,因为现在开发者可以在与 codeex 和 chatbt 相同的计算机使用基础上构建。
and the API agents API which now has computer use in it which is really cool because now developers can build on the same computer use that is part of codeex and and chatbt.
然后还有一些我们现有计算机使用功能的演示,比如 app shot,你可以把你电脑上正在做的事情的上下文快速带入 codeex 和 chatbt。还有在 Mac 上的原生计算机使用,Roman 让它自动截取他的应用截图,他可以在计算机使用使用他的应用时在电脑上做其他事情。所以是的,非常激动人心的主题演讲。
And then there were some demos of our existing computer use features like app shot where you can take the context of something you're doing on your computer and bring it into codeex and chatbt really fast. And then like native computer use on your Mac where Roman had it taking screenshots of his app automatically and he could do other things on his computer while computer use was using his applications. So yeah, really exciting keynote
更不用说这个 decisions API。decisions API
and not to mention this decisions API. decisions API
一开始,它们都是同一个模型吗?比如这是同一个数据集蒸馏到不同模型,还是说计算机使用和 decisions API 是分开的?所以 decisions API 真正酷的地方在于,你知道,它有所有这些新能力,它并行进行推理,它没有推理能力,它是一个比我们用于计算机使用的模型更小的模型,所以这些能力让它非常快。它们也让它不太擅长做长期、复杂的任务。所以我认为,如何将这些方法结合起来仍然是一个开放的研究领域。但是的,我真的很兴奋看到人们用 decisions API 构建什么。
uh off the bat are they all the same model like this is or the same data set distilled to different models um like basically like is computer using decisions API or are they like kind of separate so what's really cool about the decisions API is it you know it it has all these new capabilities it does inference in parallel um it doesn't have reasoning it's a smaller model um than the ones we use for computer use um and so those those capabilities make it really fast. Um they also make it a little bit less good at doing like long horizon um sort of sophisticated tasks. And so I think it's I would say it's still an open area of research for how we like bring those approaches together. But uh yeah, I'm really excited to see what people build with the decisions API.
有趣的一点是,dots 现在附带了个人电脑。所以看起来它们更加持久。你已经使用了一段时间。人们应该如何突破界限?人们应该以什么为目标?他们应该尝试什么?就我个人而言,现在我用它来处理很多客户服务,比如哦这个错了。我不想登录。我不想验证。找到什么并修复它。我们应该如何进一步推进?人们应该尝试什么?
One of the interesting things is dots now have attached personal computers. So it seems like they're very much more persistent. You've been using them for a while. How should people push the bounds? It's like what should people aim for? What should they try? Personally, right now I use it for a lot of customer service like oh this was wrong. I don't want to sign in. I don't want to authenticate. Find whatever and just get it fixed. How how should we push further? What should people try?
Dots 是一个非常酷的产品,因为每个 dot 都可以访问它自己的云端 Linux 虚拟计算机。这与我们的其他产品不同。你知道,传统上我们可以访问云端的浏览器,或者它可以访问你自己的电脑,但现在你可以在云端拥有整个 Linux 电脑,所以它可以运行完整的桌面应用程序。它也可以使用网络浏览器。所以是的,你知道,我认为计算机使用的强大之处,以及我认为它如此令人兴奋的原因,是因为它使得智能体可以做任何你作为一个人可以做的事情。因为世界上所有的软件都是为人类设计的,现在智能体可以使用同样的软件,你可以委托给智能体。所以,是的,就像你在电脑上做的任何事情,你都可以让一个 dot 去做。是的,我认为特别有用的是什么,真的取决于最终用户是谁,以及他们生活中什么是有价值的。但是的,我会从思考你花时间在什么事情上,以及如何将这些委托给智能体开始。
Dots are really cool product because each dot has access to its own Linux virtual computer in the cloud. Which is different from our other products. You know, traditionally we've have access to a browser in the cloud or it has access to your own computer, but now you get your own entire Linux computer in the cloud and so it can run full desktop applications. And it can also use a web browser. And so yeah, you know, I think the powerful thing about computer use and the reason why I think it's so exciting is because it makes it so that the agent can do anything you as a person can do. Because all the software in the world was designed for humans and now agents can use that same software and you can delegate to the agent. So, yeah, like anything that you would do on a computer, you can ask a dot to do. Yeah, I think what particularly useful is going to really depend on who the end user is and what what's valuable in their life. But yeah, I would just start by thinking about like what are the things that you spend time on and how could you delegate those to an agent?
是的,很多航班预订和购物,老实说甚至像玩游戏之类的,对吧?
Yeah, a lot of flight booking and shopping and honestly even like playing a game or whatever, right?
完全正确。嗯,是的,我不知道。对我来说。嗯,我最近做的一件事,嗯,我一直在努力,我订阅了一个备餐服务,因为我想吃得健康,你知道,我真的很喜欢我找到的这个备餐服务,因为它让我可以高度精细地定制我订购的餐食。所以,我可以说,我想要这么多克鸡肉和这么多克米饭。嗯,但它太复杂了,我花了两个小时才完成一个订单。我发现我可以让计算机使用为我做这件事,它在 15 分钟内完成了。嗯,所以我实际上节省了两个小时。嗯,它既比我快八倍,又为我节省了两个小时,在 GP6.1 Soul 上。嗯,所以这些是我觉得非常强大的任务类型。
Totally. Um, yeah, I don't know. for me. Um, something I did recently, um, I've been working on I I've, um, subscribed to a meal prep service cuz I was trying to like eat healthy, you know, and I really like this meal prep service I found because it lets me customize the meals I order to like a high degree of granularity. So, I can say like, I want this many grams of chicken and this many grams of of of rice. Um, but it was so complicated, it took me two hours to do an order. And I found that I could ask computer use to do it for me and it did it in 15 minutes. Um, so I I actually saved two hours. Um, it both did it eight times faster than I could and it saved me two hours on on GP6.1 Soul. Um, so those are the kinds of tasks that I feel like uh are really powerful.
作为一个创作者,我可以立即告诉你,我的头号用例是自动化 YouTube。呃,因为 YouTube 不通过 API 暴露很多东西,你只能把它放在虚拟机里运行,比如他们的 AB 测试功能或制作社区帖子。这些都不能通过 API 获得,因为他们讨厌开发者。嗯,不管怎样,所以
As a creator, I can tell you automatically immediately my number one use case is automating YouTube. Uh, because uh YouTube doesn't expose a lot of things via API and you have to just put it in a VM and just like run it uh for like let's say let's say their AB testing feature or making community posts. None of this is available by API because they hate developers. Um anyway, so
我从我们的开发者体验团队那里听说过。他们经常用它来处理 YouTube。是的,这真的很棒。
I've heard that from our developer experience team. They use it with YouTube a lot. Yeah, it's really awesome.
嗯,所以我想画一下,你知道,呃,假设我想来点刺激的。
Um so I want to draw for you know uh let's say I want to get a little bit spicy.
有一档领先的 AI 播客,我们的朋友,以一句话出名,说计算机使用在过去两年里没有进步。
One of the leading AI podcasts, our friend, is famous for saying that computer use hasn't advanced in the last two years.
是的。
Yeah.
这个说法很有意思,我觉得你是全世界最适合聊这个话题的人之一。事情到底是怎么进步的?
Which is a very interesting statement, and I think you're one of the best people in the world to talk about this. How have things progressed?
是的,你知道,那是几个月前说的,我希望他们现在看法不一样了,因为计算机使用已经完全是 180 度的不同了。
Yeah, you know, they said that a few months ago, I think, and I hope they have a different perspective now, because computer use is like 180 degrees different.
他可是很难被打动的人。
He's a tough guy to impress.
好吧。
Okay.
你基本上整个职业生涯都在做某种计算机自动化,对吧?比如在 Apple 做 Shortcuts,然后是 Sky,再然后加入 OpenAI。你能讲讲你的主线是什么吗?是什么在驱动你?当年有什么是不可能做到的?你的里程碑又是什么?
You've basically spent your whole career working on some kind of computer automation, right? Like Shortcuts at Apple, and then Sky, and then joining OpenAI. Can you draw what your through line is, what's driving you, and what wasn't possible back then, and what your milestones were?
我想接着追问一个问题,从上周到今天的 Codex 计算机使用,最大的变化是什么?是模型?是 dots?还是 harness?所以既有整个历史,也有今天发布里真正改变的东西?
I guess to add on to that as a follow-up question, what's the major change from using Codex computer use from like last week through to today? Is it model? Is it dots? Is it harness? So all the history plus what really just changed in today's announcements?
是的,说到主线,我一直对自动化、帮助人们把任务自动化这件事很兴奋,因为这样你就能省下时间,把精力放在比精细操作电脑更重要的事情上。所以,是的,这就是我们做那些产品的原因。我之前在 Apple。我们创办了一家公司叫 Sky。后来我们加入了 OpenAI,这真的很令人兴奋。
Yeah, on the through line, I guess I've always been excited about automation and helping people automate tasks, because then you can save time in your life and focus on things that are more important to you than operating a computer very intricately. And so, yeah, that was why we worked on some of those products. I was at Apple before. We made a company called Sky. We ended up joining OpenAI, which is really exciting.
几乎就像你得在 Apple 周围折腾,直到 Apple 说,好吧,我们直接雇你,你就可以在里面干活了,对吧?
Almost like you have to hack around Apple until Apple was like, fine, we'll just hire you and you can just work on the inside, right?
那是个很酷的工作地方。回头看 Sky 真的很有意思,因为我们在那里也在做计算机使用,而当时的模型能力差太多了,现在模型仅仅在过去一年里就在计算机使用上变得极其强大。我觉得我看到的最大变化是,以前它们能可靠地启动任务,但随后就会遇到问题,而现在它们非常擅长调试。它们非常擅长重试,反思什么有效、什么无效。
It was a cool place to get to work. It was really interesting looking back at Sky, as we were working on computer use there as well, and the models were so much less capable, and now the models just in the last one year have become extraordinarily capable at computer use. I think the biggest delta that I see is before they could reliably start tasks but then they would run into problems, and now they're really good at debugging. They're really good at trying again, introspecting what is and isn't working.
而且我觉得我们也把计算机使用这个领域本身向前推进了。我们用了更多技术。现在计算机使用经常会写代码。所以如果你真的在 Codex 里看,手动展开工具调用,你会看到它不是一次只做一个动作。它实际上是在写 JavaScript 代码并执行,由计算机执行,有时一次完成很多动作,这是很大的加速,也是很强的能力。我们用了更多可访问性、某种多模态界面。所以模型可能用截图,可能用可访问性,可能用 Playwright。它可以根据手头任务使用很多不同机制。然后,是的,模型加速一直非常惊人。
And I think we've also brought the computer use field itself forward. I think we're using more techniques. Now computer use often writes code. So if you actually look at it in Codex and you expand the tool calls manually, you can see that it's not just doing one action at a time. It's actually writing JavaScript code that it executes, that the computer executes, to perform sometimes many actions at once, which is a great speed up and great capability. We use more accessibility, sort of multimodal interfaces. So the model may use screenshots, it may use accessibility, it may use Playwright. It can use a lot of different mechanisms based on the task at hand. And then yeah, the model acceleration has been just amazing.
那今天有什么不同?
So what's different today?
我觉得我们一直在改进。所以仅仅一天的差异,可能比过去一个月或两个月的差异还没那么重要。但是的,我觉得 dots 里的计算机使用真的很令人兴奋,还有我们推出的新模型。
I think we're making improvements all the time. So I think just like one day's difference is probably a little bit less consequential than even the past month or the past two months. But yeah, I think the computer use in dots is really exciting, as well as the new model that we came out with.
在主题演讲里,Tijall 提到计算机使用速度提升了 7 倍,在一些基准上也好很多。你们怎么考虑衡量它?计算机使用是那种,正如你说的,是随时间改进的东西。是 harness?是模型?是后训练?你们内部怎么看如何衡量它有多好,新模型的变化又是什么?
On the keynote, Tijall was mentioning 7x improvements in computer use speed, a lot better on a few benchmarks. How do you guys think about measuring it? Computer use is one of those things where, as you say, it's improvements over time. Is it harness? Is it model? Is it post-training? How do you guys look at it internally about measuring how good it is, and what were the changes with the new model?
我们其实有很多不同的衡量方式。其中一些是在 harness 的不同排列和配置上。情况有点复杂,因为我们的生产产品有更多安全检查,而这些会根据手头任务的需要做不同配置。所以有很多衡量方式,但我觉得不管怎么衡量,我们都发现相当一致的提升。这些提升有时在 harness,有时在模型。是的,我对这个结果非常兴奋,GPT 6.1 在计算机使用上的成本效益甚至比它相对 Astra 的基线成本改进还要好。看到这个真的很酷。
We actually have a bunch of different ways of measuring it. Some of which are on different permutations and configurations of the harness. It's a bit of a complicated story because our production products have some more safety checks, and those are configured differently based on the needs of the task at hand. So there's a lot of ways to measure it, but I think regardless of how we measure it, we find pretty consistent gains. And those gains are sometimes in the harness and sometimes in the model. And yeah, I was really excited by this result that GPT 6.1 is even more cost effective for computer use than its baseline cost improvement as compared to Astra. It's really cool to see.
是的。我是说,直播里我很喜欢的一个画面是,你们在改进曲线的帕累托前沿,而且有很多讨论是关于你们如何和 harness 一起改进它。你能举一些你们有过的顿悟时刻的例子吗?不管是模型驱动 harness,还是 harness 驱动模型,随便哪种。
Yeah. I mean, one of the visuals I really liked from the live stream was that you're sort of improving the Pareto frontier of your curve, and there was a lot of talking about how you're improving it together with the harness. Can you give some examples of aha moments that you had, whether it's on model driving the harness, harness driving the model, whatever?
我不想重复自己,但我觉得引入更多模态真的非常强大。一个更具体的例子是,过去我觉得很多计算机使用产品不得不花大量时间滚动。所以它会截个图,试着做点什么,然后会想,哦,我得往下滚到下一页结果,然后再截图,再试着做点什么,再往下滚。所以我觉得有了可访问性,以及直接访问 DOM 和其他类似东西,现在语言模型实际上能看到整个页面或整个应用,它能写代码一次完成多个步骤。所以我觉得这些可能是最大的单个顿悟时刻。还有很多相比之下没那么令人兴奋的小顿悟,但实际上我们也发现,很多速度提升是由很多小痛点驱动的,我们必须深入反思。
I don't mean to repeat myself, but I think introducing more modalities has been really powerful. One more specific example of that is, in the past I think we saw a lot of computer use products had to spend a lot of time scrolling. So it would take a screenshot, it would try to do something, it would be like, oh, I got to scroll down to the next page of results, and then it would take a screenshot, and then it would try to do something, it would scroll down again. And so I think with accessibility and other direct access to the DOM and other things like that, now the language model can actually see an entire page or an entire application, it can write code that can do multiple steps at once. And so I think those have been probably the biggest single aha moments. There's a lot of tiny ones that are less exciting in comparison, but actually we do find also that a lot of speed improvements are driven by a lot of little paper cuts that we got to go in and introspect.
很多非常艰难的工程。我是说,appshots 总体来说,对吧?我觉得人们不太理解区别,因为你在 Codex 里截图时会看到一个不错的视觉呈现,但他们可能没理解区别在于你实际上能驱动每个按钮,而且每个文本都有非常优化的表示。
A lot of really hard engineering. I mean, appshots in general, right? I think people don't quite get the difference, because there's a nice visual in Codex when you take a snapshot, but they don't maybe get the difference that you are able to actually drive each button and you have each text in a very optimal representation.
是的。是的。完全正确。是的。如果你想特别极客一点,其实还挺好玩的。你可以进入 Codex,按两个 Command 键截一个 app shot。这样你就从你正在用的任何应用里抓取内容,带进 Codex 或 ChatGPT,然后如果你点击附件,再点右上角那个小小的按钮,你就能看到原始文本,看到原始的可访问性表示。
Yeah. Yeah. Exactly. Yeah. It's kind of fun actually if you want to be really nerdy about it. You can go into Codex, take an app shot by hitting the two command keys. So you grab the content from whatever app you're working with, bring it into the Codex or ChatGPT, and then if you click on the attachment and you click on this little tiny button in the top right, you can see the raw text and you see the raw accessibility representation.
是的,我们投入了大量工作来把所有信息倾倒出来,同时还要做到 token 高效。这里面有点艺术。结果发现,最初为有 accessibility 需求、想用屏幕阅读器技术的人类发明的技术,对他们使用电脑很有帮助,对语言模型使用电脑也很有帮助。所以这做起来真的很有趣。
And yeah, we've put a lot of work into dumping everything out, but also making it token efficient. There's a bit of an art to it. And it turns out that the same technology that was invented for humans who maybe have accessibility needs, who want to use a screen reader technology, that technology is really helpful for them to be able to use computers. It's also really helpful for an LM to be able to use computers. So that's been really fun to get to work on.
补充一下背景,我觉得很多人不了解 appshots。他们甚至不知道这是个功能。就是你双击 command 键,它会拉进来一个看起来像截图的东西,你会想,哦,我为什么打开了一个截图扔进去?不,它实际上是在拉取所有元数据、所有代码、一切。
For context, I feel like a lot of people don't understand appshots. They don't even know it's a feature. It's when you double hit command, it pulls in what looks like a screenshot, and you're like, oh, why have I opened up just a screenshot and thrown in? No, it's actually pulling all the metadata, all the code, everything.
是的,没错。所以就像,如果你截取一个带链接的网页的截图,截图并不包含链接指向哪里。也不包含,比如你截取日历的截图,事件标题会被截断,但当你用 appshot 时,它给语言模型提供关于一切的全部上下文,这让它能做更多事情。
Yeah. Exactly. Yeah. So it's like, you know, if you take a screenshot of a web page that has a link, the screenshot doesn't include where the link goes. It doesn't include, you know, maybe you take a screenshot of your calendar, the event titles are truncated, you know, but when you take an appshot, it gives like the language model like full context about everything and that lets it just sort of like do much more.
是的。想了解更多的话,Jason Liu,我邀请他在 AI Engineer 上做了一个完整的 workshop。太棒了。
Yeah. For those who want to see more, Jason Liu, I invited him to do a full workshop on this at AI Engineer. Amazing.
他做得非常好。
He did a great job.
关于 computer use 智能体,我有一个更宏观的愿景问题。你举的例子,截图、滚动页面、再截图,就像我们今天看到的,它们能自动化很多。瓶颈在哪里?是模型?是 harness?你觉得两年后会怎样?你觉得它会连续运行几个小时吗?我们怎么达到那一步?对 computer use 的走向有什么预测吗?
I have a broader vision question on computer use agents. So your example of take a screenshot, scroll page, take a screenshot where we were today, they can automate a lot. What are the bottlenecks? Is it models? Is it harnesses? Where do you see it going in like 2 years? Do you see it just running for hours? How do we get there? Any predictions on where computer use goes?
是的,我觉得真正疯狂的是,团队在过去几个月里取得的成就是,现在 computer use 在完成任务上可能比普通人更快,在大多数情况下。我认为下一个前沿是让 computer use 的表现真正超越人类,实际上比我们这样的专家电脑用户使用软件一样快或更快。我认为当那发生时,会非常有影响力和令人兴奋,因为我觉得我们会突然能够构建提供更多实时体验的产品,而且我认为它也会降低使用 computer use 的门槛或激活能量。我想这会让我们开始默认用智能体做某些我们习惯手动做的事情。我觉得这也很令人兴奋,因为它会节省我们大量时间。我认为有很多不同的小麻烦和瓶颈阻碍着这一点。我觉得,是的,有模型方面的问题,有推理方面的问题,有 harness 方面的问题,有表示方面的问题。我们发现,随着 computer use 变快,我们越来越受限于操作本身的速度。比如,在我们的 computer use 任务基准测试中,相当一部分时间实际上是在等待 doordash.com 本身加载。
Yeah, I mean I think what's really crazy that I think the teams accomplished over the past couple months is that now computer use is like faster at accomplishing tasks than like the average human probably in most cases. And I think that the next frontier is to have computer use be like literally superhuman in its performance where it actually is as fast or faster at using software than like expert computer users like us. And I think that'll be really consequential and exciting when that happens because I think we'll be able to all of a sudden build products that provide just much more realtime experiences and I think it'll also lower the barrier to entry or the activation energy I suppose of using computer use. I think will make us start to default to doing certain things in agents that we've become accustomed to doing manually. And I think that's exciting also because it'll save us a ton of time. And I think there's a, you know, there are a lot of different little paper cuts and bottlenecks that are sort of standing in the way of that. I think that there's, yeah, there's things on the model side, there's things on the inference side, there's things on the harness side, there's things in the representation. You know, we find that as computer use gets faster, we're increasingly bottlenecked by just like the speed of doing an operation. Like for example, you know, a non-trivial amount of time in our benchmarks of computer use tasks is actually like, let's say you're automating a task on doordash.com, like a lot of the time is actually waiting for doordash.com itself to load, you know.
是的。
Yeah.
然后你就写一个等待,然后执行等待。
Then you just write a wait and then you execute the wait.
是的,完全正确。而且你想做到,实际上非常重要的是,你希望尽可能减少延迟,从它最终加载完成到你触发语言模型执行下一个动作之间,这本身其实是一门统计科学。
Yeah, totally. And you want to get actually it's actually really important that you get that de like you want as little delay as possible between when it finally finishes loading and when you go and trigger the LM do the next action which is actually itself a statistical science.
也许可以用事件驱动的方式来做。
Like an event driven way maybe to do that.
在可能的情况下,你希望它是事件驱动的,JavaScript 有加载事件,网页浏览器有网页导航的加载事件,但还有其他类型的事件实际上真的无法事件驱动,所以有很多复杂性。
When possible you want it to be event driven and JavaScript has load events for the web browser has load events for web navigation but there's other types of events that actually really can't be event driven so there's a lot of complexity.
我想到的一个例子是,和客服聊天,回复可能需要 30 秒,也可能需要 3 分钟。
The one that comes to mind is like chatting with customer service replies could take 30 seconds, could take 3 minutes.
哦,对。我用 Codex 处理过很多机器人。很好,但我也想知道对方是否知道他们在和机器人说话,因为我用完整的句子回答。我正确使用大写。我给出完整的参考编号等等。这太好了,明显太好了。我不在乎。我只是想得到我的
Oh, right. I have dealt with so many bots with Codex. It's great, but I also wonder if the other side knows that they're talking to a bot cuz I'm like answering in complete sentences. Like I'm capitalizing correctly. Like I'm giving full reference numbers and everything. Like it's too it's clearly too good. I don't care. Like I'm just like trying to get my
我提示它,你知道,别假装你是机器人。要像一个非常恼火的人类。简短的俏皮话,像催促它。做所有这些。我还告诉它,在等待回复时,使用子智能体研究更好的方法,弄清楚我们需要什么。不错。就像人类的小干预。
I prompted it to like you know don't pretend you're a bot. Be be very annoyed human. Short one-liners like push it. Do all this. I also tell it while you're waiting for responses like use sub agents to research better ways to figure out what we need. Nice. It's just like human little intervention.
太棒了。我也觉得有一半时间对方是个机器人。所以现在你让机器人互相交谈了。
That's awesome. I also feel like half the time it's a bot on the other end. So now you got the bots talking to each other.
是的。我还要说,你知道,我们现在达到的 computer use 的一个里程碑是,三四年前我们还害怕把语言模型连接到网络和我们的设备。现在我却让它为我配置 DNS。哇。我让它付账单。真的像数万美元的东西,我就直接交给它,用 computer use 大胆冒险,你知道,最坏能发生什么。所以这一切都很好。
Yeah. I will also say, you know, like one milestone of computer use that we're at now is, you know, 3-4 years ago we were scared of hooking up LMs to the web and to our devices. And now I'm having it configure DNS for me. Wow. I'm having it pay my bills. And like really like tens of thousands of dollars of like stuff I'm just sending it over and yoloing with computer use and like, you know, what's the worst thing that can happen. So that's all really good.
我想现在你显然也要吃自己的狗粮,所有这些。既然你已经通过 API 发布了这个,你想告诉开发者哪些陷阱或建议,因为他们即将第一次遇到所有这些?
I think now that you've obviously you also have to dog food your own products and all these things. Now that you've sort of released this in API, what are some pitfalls or tips that you want to tell developers because they're about to I guess encounter all this firsthand?
首先,我非常兴奋我们把 computer use 带入了 agents API。我觉得这真的很棒,因为显然很多开发者正在构建想要与第三方网站和服务协作的应用,所以 computer use 具有这种通用性。它可以与任何东西协作。所以现在开发者突然可以使用与我们构建时相同的 computer use 实现来构建。如果你想构建自己的 computer use harness,那有很棒的工作要做,但很难,而且我们在自己的 computer use harness 上训练模型。所以使用与模型分发版本一致的那个有优势。实际上可能有速度、成本和准确性优势。所以我觉得人们能在此基础上构建真的很棒。
First of all, I'm just really excited that we brought computer use into the agents API. I think this is really great because obviously a lot of developers are building applications that want to be able to work with third party websites and services and so computer use has this universality to it. It can work with anything. So now all of a sudden developers can build using the same computer use implementation that we're building on. I think there's great work to be done if you want to build your own computer use harness but it's hard and also we train our models on our computer use harness. So there is an advantage to using the one that's in distribution for the model. There actually might be a speed and cost and accuracy advantage. So I think it's really great for people to get to build on top of that.
是的,就你刚才说的那点,我觉得我们都还在这个过程中,也许我们中的一些人比世界上很多人更早地适应了这项技术并开始信任它。所以我认为我们有责任随着时间推移去建立这种信任:确保我们构建的东西是可靠的,构建正确的安全校验,在做付款这类有后果的事情之前征求用户同意,也许还要根据具体应用,确保只让它访问任务真正需要的网站或应用。所以这是值得思考的重要一点。不过,我真的很鼓励大家去试试新的智能体 API,在上面构建各种酷炫的东西。我们很想听听你们的反馈,看看效果如何。
And yeah, to the point that you were making, I think we're all still in the process, and maybe some of us are ahead of many people in the world in getting comfortable with this technology and trusting it. So I think it's incumbent on us to build that trust over time by making sure we're building things that are reliable, by building the right kinds of safety checks, by asking for the user's consent before doing something consequential like making a payment, by asking — maybe depending on the application — making sure you're only letting it access the websites or applications that it actually needs for the task. So that's something important to think about. But yeah, I'd really encourage people to try the new agents API, build all kinds of cool stuff on it. We'd love to hear your feedback depending on how it goes.
你有没有看到它影响开发工作流的方式有什么变化?关于 dots,有一点是你能在 Slack 里看到它,能看到人们用语音来构建。Roman 演示的那个例子——「改一下这个应用,并沿途给我发截图」之类的。在采用方面,你有没有看到人们如何把 computer use 用于编码工作流?大家能从中借鉴什么技巧吗?
Have you seen any changes in the way it affects dev workflows? So one of the things with dots is you're seeing it in Slack, you're seeing people use voice and build. The example Roman showed of "change this app and send me screenshots along the way" and all this. Is there anything you're seeing there in adoption about how people are using computer use for coding workflows? Any tips people should take from that?
computer use 里我最喜欢的用例之一,也是我们在实际中经常看到的,就是让智能体真正去测试它自己构建的软件,这比听起来要有意义得多。因为传统上你在 Codex 里构建东西,Codex 帮你构建好,然后你得去测试它,你现在就成了智能体的 QA,对吧?有了 computer use,你就能完成整个软件开发生命周期,智能体可以构建软件,也可以测试它。所以我构建东西、让智能体去测试它,玩得很开心——等它交到我手上时,它已经能正常工作了。我还有额外的乐趣,因为有时我在开发 computer use 本身,所以现在我有一个 computer use 智能体,它在用我的 computer use 智能体,而后者又在用别的东西。所以是的,我真的觉得这是一个极其强大的用例。
One of my favorite use cases for computer use, and one that we see a lot in the wild, is computer use letting the agent actually test the software that the agent has built, which is far more consequential than it sounds. Because traditionally you'd build something in Codex, and Codex builds it for you, and then you have to test it, and you are now like QA for the agent, right? So with computer use you can complete the software development life cycle, where the agent can build software, it can test it. So I have a lot of fun building stuff, having the agent test it — by the time it comes to me it's already working. I have extra fun because sometimes I'm developing computer use itself, and so now I have a computer use agent that's using my computer use agent that's using something else. So yeah, I really think this is a super powerful use case.
我开发了一个可视化试玩测试技能,它能真正抓出很多设计问题,而这些问题通常你只看代码是发现不了的。它也非常适合克隆应用。所以如果你在用某个糟糕的 SaaS,想干掉这个 SaaS,你就一屏一屏地克隆它。显然 computer use 可以完全驱动一切,截图、记录下来,然后用 Codex 克隆所有东西。不过,感谢你们所有的进展。我想时间到了。这不会是我们最后一次聊。
I have a visual play test skill that I've developed that really catches a lot of design issues that normally when you just look at code you wouldn't really pick it up. It's also really good for cloning apps. So if you're using a shitty SaaS and you want to kill the SaaS, you just clone it screen by screen by screen. And obviously computer use can completely drive everything, take screenshots, note it down, and then clone everything with Codex. But yeah, thanks for all your progress. I think that is our time. This is not the last that we're going to talk.
是的。太棒了。这次真的很有意思。谢谢你们邀请我。
Yeah. Cool. This has been really fun. Thank you guys for having me.
好的。
All right.
好。我们时间卡得很紧。直接开始吧。
Okay. We're strictly cut off. We're just going to dive right in.
来吧。好的。
Let's do it. Yeah.
好。那么,Nikunj,我们非常高兴能请到你。你在 API 这边发布了很多东西,就像我们刚才聊到的,现在你可以用 computer use 智能体来构建了。你想重点讲讲 API 这边的变化吗,也顺便介绍一下你自己和你的工作?
Okay. So, Nikunj, we're very excited to have you. You've shipped a lot on the API side, like we just talked about with re — you can now build with computer use agents. Anything you want to highlight on the API side of changes, and introduce yourself a little on what you do?
当然可以。我叫 Nick。我负责 API 团队的产品。在这里大概 3 年了。一直在做模型发布。我觉得这在我于 OpenAI 的这段时间里一直是个持续的事情。每出一个新模型,我们基本上都会和后训练团队、研究团队紧密合作,弄清楚它新在哪里。然后我们在 API 里把这些能力暴露出来。大致就是这么回事。如果你看看 GPT-6 所有的新东西,我们发布的酷炫新能力首先是异步函数调用。你在 Codex、dots 之类里看到的很多东西就是,工具调用耗时太长,所以你不必在工具运行时暂停模型的执行。你可以发起一个工具调用,继续运行、继续推理,然后再回来查看。所以我们发布了异步工具调用。我们还发布了轮次中途引导。现在你可以在模型推理中途注入消息。所以当你的工具调用完成时,你可以把这些指令放进去。
Yeah, for sure. My name is Nick. I lead product for the API team. Been here for roughly 3 years. Been working on launching models. I feel like that's just been a constant thing throughout my time here at OpenAI. And with every new model we try to basically work super closely with the post-training team, the research team, to figure out what's new in it. And then we expose those capabilities in the API. So that's the basic way of putting it. And if you just look at everything that's new with GPT-6, the cool new capabilities that we launched were firstly async function calling. So what you see with a lot of the things you're seeing in Codex and dots and everything is that tool calls take so long that you don't have to pause the model's execution while the tool is running. So you could just kick off a tool call, keep running, keep reasoning, and then check back in. So we launched async tool calling. We launched mid-turn steering. So now you can inject messages while the model is reasoning in the middle. So as your tool call finishes you can put in those instructions.
那某种程度上也是一种模型对齐能力,对吧?就像他们得把这种能力训练进去——
And that's also partially a model alignment capability, right? Like they have to train in the ability to train—
我觉得我们在应用里一直都有。在它推理时,你一直可以引导它。以前不是最好的,现在好多了。
I feel like we've had it in the app. You could always, as it's reasoning, you could steer it. It wasn't the best. It's gotten much better.
很期待看它表现如何。
Excited to see how it does this.
我们在 API 里的主要目标是,一旦某个能力被训练进 harness,就把它放进 API。所以我们会等那个时刻,直到它足够好。而这其中很多实际上是由 websockets 驱动的,我们大概几个月前发布了它。websockets 打开了与模型之间这种完整的双向通信。这不是 GPT Live 那个东西,我说的只是 GPT-6。你可以做所有这些异步工具调用、异步推理、注入消息。所以这是一个做起来很有趣的 API。我觉得我非常享受。
And our main goal in the API is to put things in the API once it's trained into the harness. So we kind of wait for that moment until it's good enough. And a lot of that is actually being powered by websockets, which we launched a few — I want to say months ago. And so websockets just opens this whole bidirectional communication thing with the model. This is not the GPT Live thing, I'm just talking about GPT-6. And you can do all these async tool calling, async reasoning, injecting messages. So it's a really fun API to work on. I think I'm really enjoying it.
这就是为什么我们是工程播客,因为我们可以聊 websockets。这也和 ultra fast 非常搭,对吧?我想这是它第一次在 API 里可用。
This is why we are the engineering podcast, because we get to talk about websockets. This also pairs very well with ultra fast, right? Like that is now, I think, for the first time ever available in the API.
是的。
Yes.
那基本上就是你能达到的理论最快速度。前沿智能。
Which is basically the theoretical fastest speed you can ever get. Frontier intelligence.
是的。是的。做那个项目真是太令人兴奋了。我想在讲 API 之前,UltraFast 最有意思的部分就是看着推理团队用 Astra 大展身手。他们一直在跑这些 Codex 智能体,试图榨出更多性能。我想说,至少有好几个月,很多精力都集中在效率上、把成本压下来,我们就是这样把 Luna 的价格降了大约 80%。
Yeah. Yeah. It's been so exciting to work on that project. I think before I go into the API, the most fun part of UltraFast has been just watching the inference team cook with Astra. Like they're just constantly having these Codex agents running, trying to squeeze out more performance. And I would say at least for a couple of months a lot of it was focused on efficiency and driving the cost down, which is how we were able to cut the Luna price by like 80%.
其中很大一部分是由他们落地的推理改进推动的,而现在他们又转向了另一个方向:怎么让它跑得尽可能快。UltraFast 在 Astra 这样的模型上表现真的令人惊叹——能跑到那么快真的很酷。还有 WebSockets——其实我们第一次上线 WebSockets 是为了 GPT-5.3 Codex Spark,我都不敢相信我们给模型起了这么个名字。但当时就是为了它上线的,显然它帮助很大,因为你必须要有工具调用。你必须真正降低与工具来回交互的开销。所以 WebSockets 在这方面非常棒。
A lot of that was driven by the inference improvements they landed, and then now they've shifted gears towards how can we make this run as fast as possible. UltraFast has just been amazing to see on a model like Astra — to go that fast has been really cool. And WebSockets — actually, the first time we launched WebSockets it was for GPT-5.3 Codex Spark, which I can't believe we named a model that. But that's what we launched it for, and obviously it helps so much because you've got to have the tool calls. You've got to really reduce the overhead of going back and forth with tools. So WebSockets is awesome for that.
是啊。看着还挺好玩的——我有我的重置用量额度,然后还有我从来没用过的 Spark 用量额度。就好像我想要的话它就在那儿。
Yeah. It's always cute to see — I have my reset usage limit and then I have my Spark usage limit that I never use. Like it's there if I want it.
我觉得它已经没了。终于没了。
I think it's gone. Finally it's gone.
是啊。你在慢慢把那些老模型都淘汰掉。
Yeah. You're slowly killing off all the old ones.
那是很棒的一周。我是说,那是我们第一次把前沿智能做到极致的速度。大家真的很喜欢。
It was a great week. I mean, it was the first time we had frontier intelligence at extreme speeds. People really liked it.
是啊。所以这是第一次它回归。
Yeah. So first time it comes back.
是的。
Yeah.
而 5.3 Spark 明确归功于 Cerebras。你们对 UltraFast 是否与 Cerebras 有关既不确认也不否认,但大家——我就直说吧,大家确实很在意,也很好奇。而且你们也有自己的芯片。
And for 5.3 Spark, it's explicitly attributed to Cerebras. You guys are not confirming or denying that UltraFast is related to Cerebras, but people are — I'll just say that people do care and are wondering about it. And you have your own silicon as well.
嗯,房间里的大象——决策模型。哦对。决策 API。我们是第一个和 Dogo 做 Jev 深度解析的播客,我也在 AI Engineer 上请他做过分享。你多快看到 Jev 然后就——
Um, elephant in the room — decision models. Oh yeah. Decision API. We were the first podcast to do a big Jev deep dive with Dogo, and I also featured him at AI Engineer. How quickly did you see Jev and go like —
是的,首先,要大大地致敬 Dogo 和 Jev 团队,他们真正启发了市场上这一整个细分领域。显然 Jev 一出来,所有人都为之疯狂。我们的用户来找我们,同时我们内部团队也说,我们需要一个快得多的分类系统。我不想抢在即将推出的一些 dots 功能前面透露太多,但你会看到一些很酷、非常敏捷快速的东西建立在决策 API 之上。不过,是的,致敬 Jev 启发了这整件事。显然 OpenAI 里有一堆人被勾起了极客兴趣,他们想,我们怎么才能把它做出来?我们不打算训练一个新模型,但是——
Yeah, firstly, huge props to Dogo and the Jev team for really inspiring this whole segment in the market. Obviously Jev comes out, everyone's losing their minds over it. Our users are hitting us up, but also our internal teams are like, we need a much faster classification system. I don't want to get ahead of some of the dots features that are going to come, but you're going to see some cool, really snappy fast things built on top of the decisions API. But yeah, props to Jev for inspiring this whole thing. Obviously a bunch of people at OpenAI get nerd sniped by that and they're like, how can we make this work? We're not going to train a new model, but —
大概四周前,这还根本不在计划里。
Like four weeks ago, this was not on.
完全没有。不,这就像是受 Jev 启发,然后——
Not at all. No, this is like Jev inspired and like —
我觉得你们正式成为第一个克隆并采用这个的前沿实验室。
I think you're officially the first frontier lab to like clone and adopt this.
是的。是的。我觉得 OpenAI 有非常浓厚的黑客文化,大家就是会对一些东西感到兴奋。所以一个做推理的人,还有一个基础设施团队里特别棒的人,他们说,这太棒了。我们要来搞它。他们做了个原型,能跑通,现在我们就在延迟上不断爬坡,试着把它做到尽可能快,我们想在未来几天内上线。所以一旦我们达到延迟目标,就会试着把它推出来。
Yeah. Yeah. I feel like OpenAI is such a strong hacker culture, and people are just like they get excited about things. So a guy from inference, this one awesome guy from the infra team are like, this is amazing. We're going to hack on it. They build a prototype, it works, and now we're just hill climbing on latency and trying to make this as fast as possible, and we want to launch it in the coming days. So as soon as we hit our latency target, we'll try to get this out.
有意思——在黑客文化的同时,你们也像 Sam 说的那样,是 99% 可用、最可靠的 API 之一,而且我想大概是使用量最大的,这直接归你们团队。大家应该怎么看待决策 API?我感觉很多人看到了 Jev,听到了热度,但还没用它构建过。你们正在让它变得非常主流。大家应该把它看作什么?应该怎么用它?
It's interesting — at the same time as hacker culture, you also, as Sam said, like 99% one of the most reliable APIs with I think probably the most usage, which is your team directly. How should people see decisions API? I feel like a lot of people saw Jev, heard the buzz, haven't built with it. You're making it very mainstream. What should people see it as? How should they use it?
是的,我觉得我们看到的主要用例就是非常快的分类。所有计算机使用的演示都很惊艳、很酷。当然我觉得会有局限,比如让 Astra 写一段 JavaScript 脚本来控制你的电脑,对比让 Luna 一次挑一个动作。我觉得它不会达到同样的智能水平,但也许有些计算机使用任务它已经足够好了。所以很期待看到它落地。我在内部看到的另一个很酷的原型,是人们把它和 GPT Live 接起来。GPT Live 是我们的双向实时 API,它建立在前端模型和后端模型这套模式上。GPT Live 就是那个超快的思考者—说话者。GPT Live 是说话的那个。超快,非常擅长委派。然后你有像 Astra 这样的东西坐在后面,但工具调用在 GPT Live 里一直感觉很慢。所以人们一直在做这些工具调用演示,让 GPT Live 控制电脑,感觉就敏捷自然多了。所以我挺期待看到大家用 Live 和决策 API 上的 Luna 做出什么。那会相当令人兴奋。是的。
Yeah, I think the main use cases we've seen is really fast classification. All the computer use demos have been amazing and really cool. I think there will be limitations of course, in terms of having Astra write a JavaScript script to control your computer versus having Luna pick one action at a time. I think it's not going to be at the same intelligence level, but maybe there's some computer use tasks that this is good enough for. So excited to see that come through. The other cool prototype I've seen internally is people hooking it up with GPT Live. So GPT Live is our bidirectional realtime API, and it's built on this model of front-end models and backend models. So GPT Live is this super fast thinker-talker thing. So GPT Live is the talker. Super fast, really good at delegation. And you have something like Astra sitting at the back, but tool calling has always felt really slow in GPT Live. And so people have been putting together these tool calling demos of GPT Live controlling a computer, and it just feels so much more snappy and natural. So I'm kind of excited to see what people do with Live and with Luna on decisions API. So that'll be pretty exciting. Yeah.
所以我想帮大家把这点理清楚,尤其是从产品角度,因为很多人一直在推出 Jev 克隆。过去两周大概有 100 个。
So I want to iron this out for people, especially from the product side, because a lot of people have been putting out Jev clones. There's been about 100 in the last two weeks.
哦真的吗?太厉害了。
Oh really? That's amazing.
第一个,但就像——他们能克隆 Jev API,说实话那就是结构化输出,而 OpenAI 是最先做的。
The first but like — they can clone the Jev API, which is honestly structured outputs, which OpenAI was first to.
是的。
Yeah.
对。所以我觉得咱们帮大家理清楚,什么才算决策模型,就重要的方面而言。它不只是延迟。也不只是结构化输出,对吧?因为我完全可以就用 Luna。决策模型定价和 Luna 一样,对吧?
Right. So like, I think let's iron out for people what is a decision model, as far as what is important. It is not just latency. It's not just structured output, right? Because I could just have Luna. It's the decision models priced the same as Luna, right?
嗯,把推理关掉然后加上结构化输出。我就有 Jev 了吗?你知道,并没有,对吧?而那才是真正的——
Um, have turned off reasoning and then have structured output. Do I have a Jev? You know, no, right? And that's the real —
那里有一种置信度。
There's a confidence there.
是的。是的。完全对。我觉得方式是——我们并没有为此训练一个新模型。我们纯粹是建立在已有的同一套 Luna 权重之上。所以是的,这其实就是 Luna,而在它之上,你要做的是约束。结构化输出是其中很大一部分。你真正在优化的是推理栈,让它在 TDF 上跑得非常快,而因为你可以有多个问题,你要做的就是基本上并行跑这些。是的,你把它们作为一个批次来跑。人们正在研究各种各样的推理技术来试着把它做到尽可能快,但至少我们起步时、在这个第一版里的实现,就是在 Luna 之上做这件事,看看效果如何。显然你想把它放出去——这是开放的、经典的迭代式部署——放出去,看看大家怎么想,然后我们按需做更多模型改进。所以是的,这就是决策 API。
Yeah. Yeah. Totally. I think the way that — so we haven't trained a new model for this. We're building this purely on top of the same Luna weights that we have. So yeah, this is really just Luna, and on top of that, what you're doing is you're constraining. So structured output is a big part of it. You're really optimizing the inference stack to get very fast on TDF, and because you can have multiple questions, what you do is you basically run those in parallel. Yeah, you run those as a batch. There are all sorts of inference techniques people are working on to try to make it as fast as possible, but I'd say at least our implementation of it at the start and in this first version is zeroing this on top of Luna to see how it goes. And obviously you want to put it out there — this is open, classic iterative deployment thing — put it out there, see what people think, and then we'll make more model improvements as needed. So yeah, that's the decisions API.
而且显然一个好处是你们有视觉能力,他们没有视觉能力,对吧?
And obviously as a benefit you have vision, they don't have vision, right?
确实。
True.
显然就像,我们用 Luna 免费就能获得。
Obviously comes like, we get it for free with Luna.
对。
Yeah.
对。我确实觉得有些创新,听起来还没到来,如果还是同样的 Luna 权重的话,那就是置信度和校准这些东西。这是我们在播客里聊过的话题,关于基准测试的校准,因为基本上关键在于,基于人类反馈的强化学习(RLHF)会把你坍缩到你想听到的东西上。
Yeah. I do think that some of the innovations, it sounds like it's still to come if it's still the same Luna weights, which is the confidence stuff and calibration. This is something that we've talked about on the podcast with benchmarking calibration, because basically the whole point is that RLHF kind of collapses you towards what you want to hear.
对。
Yeah.
但并不是实际的置信度是多少。
But like not actually like what the amount of confidence is.
对。对。完全同意。我很想看看结果如何。也许这些会是我们未来模型发布时需要爬坡的关键领域。
Yeah. Yeah. Totally. I'm eager to see how it pans out. Maybe these are going to be the key areas where we may have to hill climb with a future model release.
然后在架构方面,辩论中的另一件事,显然没人知道,因为 Jeff 不谈这个,但两种猜测是:一是可能是融合模型而不是自回归,但你能以自己的方式实现并行生成。另一种是某种机械可解释性类型的东西,你分析激活值然后直接输出权重,你们都在研究这个,人们也猜测过。这两者都有过演示。我想 Gemini 分享了 Gemini diffusion、Gemma diffusion 在 Jev 风格输出上,可解释性的人也从中层提取过可解释性,但这都是猜测。就像你到底想瞄准什么,对吧?因为你可以实现 API。每个人都能实现 API。这其实挺简单的。但还有速度,还有准确性,还有其他校准特性。我不知道还有什么。
And then architecture-wise, the other thing that's in the debate, obviously nobody knows because Jeff doesn't talk about it, but the two speculations are: one, maybe the fusion model instead of autoregressive, but you are able to achieve the parallel generation in your way. And then the other one is some mech interp type thing that you're analyzing the activations and then just outputting the weights, which you guys have all done the research on, people have speculated. There have been demos on both of these as well. I think Gemini shared a Gemini diffusion, Gemma diffusion on a Jev style output, and interp people have also pulled out interp from a middle layer, but this is all speculation. It's just like what are you trying to aim for, right? Because you can achieve the API. Everyone can achieve the API. It's actually pretty trivial. But then there's the speed, then there's the accuracy, then there's the other calibration features. I don't know what else.
对。对。不,完全同意。整个领域现在被启动起来真是太酷了,人们会做很多很酷的东西,大家会互相学习,是的,我很兴奋。
Yeah. Yeah. No, totally. It's so cool that this whole space has been kicked off now and people are going to do so much cool stuff and everyone's going to learn from each other and yeah, I'm excited about it.
我觉得作为平台团队的一员,你的很多工作就是赋能构建者。你觉得人们应该用 decisions API 以及计算机使用智能体来构建什么?有没有你内部一直在构建的东西,你觉得在新变化之后真正打开了局面?
I feel like being on the platform team, a lot of your job is to empower builders. What do you think people should build with decisions API and also computer use agents? Any stuff that you've been building with internally that you think really opens up after the new changes?
对。好,让我们想想。Decisions API 在内部的用例已经很明显了,比如用户运营团队就扑上去了。我们说我们得把所有支持工单分类。还有什么?显然有那些很酷的 GB 现场演示。我相信 Codex 应用团队可能会接手这个,试着做点酷的东西。所以,你知道,这整个东西大概一周前才启动,所以非常非常早期。很兴奋。
Yeah. Okay, let's think. Decisions API use cases internally have been pretty obvious, like the user ops team was jumping on it. We're like we got to classify all of our support tickets. What else came up? Obviously there were like the really cool GB live demos. I'm sure like the Codex app team might pick this up and try to do something cool with it. So, you know, like this whole thing started like a week ago, so it's very very early. Excited.
在评估方面有很大的推动,用语言模型作为评判者,在那里实现非常低的延迟,对吧?
There was a big push in evals, LM as a judge, having really low latency there, right?
对。对,那会很有意思。然后关于智能体 API,我们基本上有一堆第一方产品,比如在 OpenAI 完全构建在它之上。我们有刚发布的 COC 安全相关的东西,完全构建在智能体 API 之上。我们有那种,我们今天要发布一个会议类型的东西。我想有个演示。你记得 Sam 展示插件扩展的时候吗?有个演示,比如你在日历里,你可以让你的会议笔记直接落到 granola 之类的东西里。
Yeah. Yeah, that'll be interesting to see. And then with the agents API, we're basically having a bunch of first-party products like at OpenAI built fully on top of it. We've had the COC security stuff that just went out that's fully built on top of the agents API. We have sort of the we're having like a meetings type of thing launching today. I think there was like a demo. Do you remember like the plug-in extensions when Sam was showing it? There was like a demo for like you're in a calendar, you can sort of like have your meeting notes like drop into granola type thing.
对。
Yeah.
所以所有这些东西都完全构建在智能体 API 之上。是的,我很兴奋看到,我们刚把它推出去,看看人们会在上面构建什么。
And so all of that stuff is fully built on top of the agents API. And yeah, I'm just excited to see, we're just getting this out and let's see what people build on top of.
我觉得你展示得非常好。整个编辑空间、页面、协作、添加到你的文档。这很多。所以人们可以从中获得很多灵感。
I think you showed it off very well. The whole edit spaces, pages, collaborate, add in your doc. Like that's a lot. So there's a lot of inspiration people can build from.
对。对。用 Astra 一切皆有可能,你知道,现在事情进展得太快了,人们从想法到执行如此之快。太惊人了。
Yeah. Yeah. All possible with Astra, you know, like things just move so fast now that people go from idea to execution so quickly. It's amazing.
有没有什么你希望人们关注并给你反馈的,比如什么,也许你只是把它放出去,你想要,而且有个岔路口,你想让开发者帮你决定。
Is there something that you want people to focus on to give you feedback, like what, maybe you're just putting this out there and you want, and there's like a fork in the road and you want developers to help you decide.
所以我觉得智能体 API 和 decisions API,这些是我们最新的产品,非常希望得到任何和所有的反馈,以弄清楚把它们带向何方。我觉得在这边我们对 responses API 非常开放,它算是我们这边的主力。现在真的专注于性能,性能主要体现在两个方式。第一就是延迟。我们一直在重写整个 responses API 的路径,从首字延迟(TTFT)、字间延迟(TBT)的角度让它尽可能快。所以这继续是我们主要的关注领域。第二件我们一直在尝试做的是真正深入缓存,特别是这些个人智能体,基本上就是一个单线程永远进行下去。我们一直在努力提升缓存的水平。我们现在提供比如 30 分钟内缓存命中的保证。实际上,对于我们的一个用户,我们刚推出了长得多的缓存窗口。所以我们提供比如 12 小时的缓存保证,这样你就有保证的缓存命中,
So I think agents API and decisions API, they're like these are our newest products, would love any and all feedback on that to figure out where to take them. I think over here we're very open on responses API, which is sort of like our workhorse over here. Like really focused on performance right now and the performance comes in like two main ways. First is just like latency. We've been like rewriting the whole responses API track to like make it as fast as possible from a TTFT perspective, TBT perspective. So that's like continues to be like a main area of focus for us. The second thing we've been trying to do is like really go deep on caching, particularly with these like personal agents that are basically like a single thread that just goes on and on forever. We've been trying to like really up our game on caching. We provide now guarantees of like cache hits within like 30 minutes. We're actually like, for one of our users, we just launched like a much longer cache window. So we have like a 12-hour caching guarantee that we offer so that you have like guaranteed cache hits for
有公开的 API 吗?
Is there a public API?
还没有。那还在预览中。我们会尽快把它推给所有人。但就像为缓存多付一点点,对吧?我们保证缓存读取在长得多的时期内有效。所以即使比如你的 instinct 线程,你在上面做了些事,然后 3 到 4 小时后再回来,你仍然能获得缓存的性能。
Not yet. That's in preview. We're going to try to get that out to everyone as soon as possible. But like just pay a little bit more for the cache, right? And we like guarantee cache reads for like a much longer period. So even if like your instinct thread for example, like you just do something on it and then come back to it like 3 to 4 hours later, you're still getting the caching performance out of it.
用新模型也把成本降低了不少,对吧?
Cut the cost there quite a bit too, right, with the new model.
哦对。对。我们在压低缓存读取。对。
Oh yeah. Yeah. We're like driving down cache reads. Yeah.
对于构建者来说他们应该实现,因为便宜得多。
For builders they should implement because it's significantly cheaper.
对。对。就像构建你的应用时要非常有缓存意识,并使用我们的提示诊断或缓存诊断工具来弄清楚哪里在掉。所以缓存这部分真的很重要。对,我还想谈谈预热。我们现在在 API 里有了。所以如果你知道嘿我要获得一个缓存,抱歉,我要获得这个提示。
Yeah. Yeah. Just like building your apps with like to be very cache aware and sort of like use our prompt diagnostics or cache diagnostics tool to figure out like where things are dropping off. And so the caching part is like really important. Yeah, I also want to talk about pre-warming. We have that in the API now. So like if you know that hey I'm going to get a cache, sorry, I'm going to get this prompt.
我只想预先预热缓存,现在就把缓存费用付了,然后让它在接下来的 30 分钟里随时待命。
I just want to pre-warm the cache, pay the cache fee right now, and then have it ready to go for the next 30 minutes for whenever.
而且我可以生成那个线程的许多实例。
And I can spawn many instances of that thread.
没错。是的。你可以一直这样用下去。所以,我非常期待收到关于底层性能方面的反馈,这样我们就能继续让 Responses API 成为在 LLM 之上构建应用时性能最强、最可靠的方式。然后你们基本上就有了我们的新产品,我就是在寻求任何和所有的反馈。
Exactly. Yeah. You can just keep going. So yeah, I'm very excited about getting feedback on the low-level performance things so that we can keep making the Responses API the most performant and reliable way to build on top of an LLM. And then you basically have our new products where I'm just looking for any and all feedback.
对,直接用起来。所以,就告诉我们——请帮我们定义一下路线图吧。
Yeah, just use it. So yeah, just tell us what — just define our road map for us, please.
所以,我觉得缓存这件事很棒,对吧?显然非常需要,但归根结底,你还是会撞上百万 token 的上下文窗口,而且在可预见的未来这大概不会改变。你还是需要好的压缩,对吧?那方面的最佳实践是什么?
So yeah, I think for me the caching thing is great, right? Obviously very needed, but at the end of the day, you're still bumping up against a million token context and that's probably not going to change for the foreseeable future. Like you still need good compression, right? What is the best practice there?
是的。是的,完全同意。
Yeah. Yeah, totally.
首先,OpenAI 有自己的专有压缩,叫 comp。它在智能体 API 里。
So firstly, OpenAI has its own proprietary compression which is comp. It's in the agent API.
你替我们决定,对吧?
You decide for us, right?
没错。所以在智能体 API 里,它内建在 harness 中。如果你用的是 Responses API,有两种做法。一种是我们所说的服务端压缩,就是你告诉 Responses API,如果 token 达到这个阈值,就自动压缩,减少使用的上下文。第二种是 slash compact,如果你想要完全控制,就可以随时 slash compact,自己决定什么时候压缩。
Yeah, exactly. So in the agents API it comes built into the harness. And if you're in the Responses API there's two ways of doing it. One is what we call server-side compaction, which is you basically tell the Responses API that if you ever hit this threshold of tokens, just auto-compact it and reduce the context being used. And the second way is slash compact, which is if you want full control, so you can slash compact at any time, have your own logic on when to.
它不是——
It's not a —
是的,但我的意思是,它就是手动覆盖。
Yeah, but I mean it is the manual override.
对,它是手动方式。
Yeah, it is the manual way.
而且,我不知道,但很多大型编程智能体喜欢手动做这件事。我的意思是,如果你看开源 Codex harness 里的 Codex 实现,你能看到他们用 /compact 来做。我们还在研究新的压缩技术。其中一些你会在 Codex harness 里看到——它已经在 Codex harness 里实现了——还有一些我们正在试验的基于文件的系统。所以,围绕压缩也有很多很酷的东西在进行。
And I don't know, but a lot of the big coding agents like to do it manually. I mean, if you look at the Codex implementation of it in the open-source Codex harness, you can see that they use /compact and do it. And there's also new compaction techniques that we're working on. Some of them you will be able to see in the Codex harness — it's already implemented in the Codex harness — and they're some file-based systems that we are experimenting with. So yeah, lots of cool stuff going on around compaction as well.
酷。我们快没时间了。我想你已经谈了很多关于性能的内容,也谈了很多你们正在发布的新 API。关于平台的未来,你能再给点提示,说说你感兴趣的其他方向吗?
Cool. We are running out of time. I think you've talked a lot about performance and talked a lot about the new APIs that you're launching. Can you give us any other hints as to things that you're interested in as far as the future of the platform is concerned?
我们显然非常底层。我之前在 Stripe 工作,在 Stripe,很多工作就是在核心支付原语之上构建这些更高层的原语和产品。我一直很好奇在 AI 里做这件事的最佳方式是什么。我觉得我们已经尝试过几次。我们很久以前推出过 Assistants API,但它并不是很合适。我们现在在走这个智能体 API 的路子,它给你 Codex harness,但其中应该给多少灵活性?这是个开放问题。比如我们应该如何设置记忆库,以及所有这些更高层的 API 对象,来抽象掉更多存储概念?这是我非常好奇、想要弄清楚如何设计的一整个领域。我觉得 AI 里很多事情就是:有一个底层 API 原语,看一个示例 harness,然后让你的编程智能体去实现它。但其中有多少应该内建到 API 里,是我一直在思考的问题。
We're obviously very low level. I used to work at Stripe before this, and at Stripe a lot of the game was building these higher-level primitives and products on top of the core payments primitives. And I'm always curious about what the best way of doing that is in AI. And I think we've had a couple of attempts at that. We had launched the Assistants API way back in the day and it wasn't really the right fit. We're sort of going off with this agents API and it gives you the Codex harness, but where's the right amount of flexibility to give in that? That's an open question. Like how should we have memory vaults and all of these higher-level API objects to abstract away more storage concepts? This is a whole space that I'm very curious about figuring out how we design. I think a lot of things in AI are just: have a low-level API primitive and see an example harness and go and have your coding agent implement that. But how much that should be built into the API is a constant question that I'm thinking about.
所以我不知道大家对此有没有想法。如果有人有想法,会非常有意思,我很想听听。
So I don't know if folks have thoughts on that. If anyone has ideas, it'll be super interesting to hear.
是的,我一直会回到并用来收尾的类比是:你在构建一朵 AI 云,对吧?这是 Sam 一年前说过的话,这几乎就像你在做 AWS 的发明,你必须做——好,这是 EC2,这是 S3,这是——但你在做的是每一个的 AI 原生版本。有很多类比。所以,你在为你知道会——的东西预热缓存,而且很好的是它全都暴露给构建者,因为它打开了构建新东西的方式。
Yeah, the analogy I always bring back to and to end there is you're building an AI cloud, right? Which is something that Sam said a year ago, and it's almost like you're sort of doing the AWS invention and you have to do okay this is EC2 and this is S3 and this is like — but you're doing the AI-native versions of each of these. There are a lot of analogies. So, you're pre-warming caches for stuff that you know will — and it's nice that it's all exposed to builders cuz it just opens up ways that you can build new things.
是的,绝对如此。
Yeah, absolutely.
好的,太棒了。那么,一切。谢谢大家。谢谢。是的。
Okay, awesome. Well, everything. Thank you guys. Thank you. Yeah.