从无代码到 ChatGPT Work:自下而上的雄心时代

From No-Code to ChatGPT Work: The Era of Bottoms-Up Ambition

阿克谢·内森 Akshay Nathan · Latent Space · 2026-07-28 · 约 71 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

来自 OpenAI 的 Akshay 探讨了想法和品味如何成为新的瓶颈,以及 ChatGPT Work 如何体现将代码魔力带给每个人的使命。

Akshay from OpenAI discusses how ideas and taste are the new bottlenecks, and how ChatGPT Work embodies the mission of bringing code's magic to everyone.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 52)

全文 · Full transcript(中英对照)

瓶颈:想法与品味 Bottleneck: ideas and taste

Akshay

我觉得瓶颈会变成想法和品味。因为现在谁都能构建,这真的是一个自下而上的野心时代。可做的事太多了,所以任何时候,制约你的都是想法数量和你手头同时在做的项目数量。

I think the bottleneck becomes ideas and taste, I guess. Because anyone can build now, it really is the era of bottom-up ambition. And because there's so much to be built, you're always going to be bottlenecked by the amount of ideas and the amount of things you're doing at any given time.

Host

关于想法有一点很有意思:它们不会凭空出现,通常来自某个地方。在产品开发里,它们来自——

One interesting part about ideas is that they're not in a vacuum. They usually come from somewhere. In product development, they're coming from—

Akshay

来自和用户交流,或者是对你看到的摩擦、收到的反馈做出的反应。我们说的那种通才总会很有价值——把反馈回路闭环,再基于这些反馈或用户沟通去产生想法,无论具体形式是什么。

Talking to users, or reacting to friction or feedback. There will always be value in these generalists we talked about, closing that loop and coming up with ideas grounded in feedback or in talking to users, whatever it is.

赞助商插曲 Sponsor interlude

Host

在进入今天的正题之前,我有一小段话要对听众说。谢谢你们。如果不是你们选择点进来收听我们的内容,我们不可能把你们显然想要的 AI 工程、科学与娱乐内容带给大家。几乎每天都有赞助商来联系我们,但幸运的是,有足够多的听众真正订阅了我们,让这一切可以在没有广告的情况下持续下去,我们也想保持这种状态。但我只想请大家帮一个忙:你们能做的最有力、也完全免费的一件事,就是点击订阅按钮。这是我对你们唯一的请求,而它对我、以及我的团队来说意义重大——我们每周都拼命把《In Space》带到你们面前。如果你们订阅了,我保证我们永远不会停止努力,把节目做得更好。现在,让我们进入正题。

Before we get into today's episode, I just have this small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.

从无代码到ChatGPT Work From no-code to ChatGPT Work

Host

好,今天我们请到的是来自 OpenAI 的 Akshay。欢迎。

Okay, we're here in the studio with Akshay from OpenAI. Welcome.

Akshay

谢谢。

Thank you.

Host

还有我们可靠的联合主持人 Vibhu。你最近发布了 ChatGPT Work,你负责核心产品工程。这一路走来非常长。我觉得很有意思的是,你最早做的是无代码或低代码,包括 Walrus 和 Airtable。从某种程度上说,ChatGPT Work 有点像超级应用里的超级应用——这就是终极无代码:你只需要写一句提示词。

And with our trusty co-host Vibhu. So you recently launched ChatGPT Work. You lead core product engineering. It's been a long journey into all this. I find it very interesting that you started with no-code or low-code with Walrus and Airtable. To some extent, ChatGPT Work is kind of like the super app of super apps—well, here is the ultimate no-code: you just write a prompt.

Akshay

是啊,万物真是螺旋式回归。我的职业生涯有很长一段时间——我一开始做消费者金融科技——但后来有了一个假设:工程师用代码能做到的那些事,如果能以更容易上手的方式带给更多人,那会是真正的魔法。我们当时在做一家创业公司,实际上在 LLM 和视觉 LLM 出现之前,我就做过用 AI 做自动化测试。那时候很粗糙,但我们在力所能及的范围内做了。之后我在 Airtable 待了一段时间,围绕同一个假设:如果能把数据库、或者说数据库背后的参数能力带给普通人,会对他们非常有用。但等 LLM 登场之后,一切就清楚了:那就是缺失的那块拼图——把代码的魔法带给所有人,而他们不必了解底层原理所需的最后一项技术。所以我觉得,这次发布以及我们做的所有事,正是那个想法的体现。

Yeah, it's funny how things come full circle. For a long time in my career—I started in consumer fintech—but after that, there was this hypothesis: the things we could do with code as engineers, if we could bring that to many more people in a more accessible way, that would be truly magical. We were working on a startup, and actually before LLMs, before vision LLMs, I worked on doing automated testing with AI. It was kind of janky back then, but you know, doing what we could. Then I worked at Airtable for a while on the same thesis: if we could bring a database, or the parameters behind a database, to people, that would be really useful to them. But once LLMs came onto the scene, it became clear that this was the missing piece—the missing technology required to bring the magic of code to everyone without them having to know what's going on under the hood. So I think this launch, and all the stuff we've been up to, is the manifestation of that.

加入OpenAI与变化 Joining OpenAI and what changed

Host

你刚加入时情况怎么样?你是 2023 年加入 OpenAI 的。现在我们有了好多新东西:ChatGPT、Codex 应用、ChatGPT for work。事情发生了哪些变化?

How was stuff when you joined? You joined OpenAI in 2023. Now we've got so much more stuff: ChatGPT, Codex app, ChatGPT for work. How have things changed?

Akshay

其实我觉得更有趣的反而是那些没变的东西。我记得加入的时候公司大约 500 人。我之前担心的是:我一直在找更早阶段的公司,加入这家够不够像创业公司?结果我进去后发现,天哪,这比我能在想象中能想到的还要更像创业公司。这一点直到现在都没变。那种自下而上的野心,以及任何人都能做点事、产生想法并把它直接发布出来的能力,真的很酷。但就使命而言,最打动我的是这个使命:把前沿智能带给每个人,构建 AGI,然后把它交到每个人手里。而且我们当时就已经承认,这条路径不会是线性的。我们大概会尝试不同的产品,会有一些成功,一些不成功。但愿景没有变,使命也没有变,而现在我们开始看到各个部分拼合到一起了。这真的很酷。

I actually think the more interesting thing is how things haven't changed. When I joined, I remember it was about 500 people. One thing I was worried about was that I was looking for something more early-stage, and was it going to be startup-y enough? I joined and I was like, dude, this feels even more start-up-y than I could ever imagine. That really hasn't changed even now. The level of bottom-up ambition and the ability of anyone to do anything, have an idea, and ship it is really cool. But on the mission side, what was really compelling to me is this mission of bringing frontier intelligence to everyone, building AGI and then bringing it to everyone. And even then, I think we acknowledged that the vision wasn't going to be a linear progression. We were probably going to try different products and have different things that succeed and don't. But the vision has stayed the same, the mission has stayed the same, and we're starting to see the pieces fall together. That's really cool.

企业经验教训 Lessons from enterprise

Host

你之前做过企业业务。很多人从没碰过 ChatGPT 企业版。你从那里学到了什么,现在带进了你手头的工作?

You worked on enterprise. A lot of people never touched ChatGPT for enterprise. What is something you learned from there that you're bringing into your work now?

Akshay

我觉得企业场景不存在放之四海而皆准的解决方案。我记得 ChatGPT 企业版刚推出的时候,我们会和客户交流。当时大概是 ChatGPT 发布一年后,每个人都特别兴奋,想把 AI 带进企业。出现了很多团队,被设置为 AI 部署团队,预算非常庞大。如果你问任何人他们兴奋的是什么、想解决什么,一开始你会得到一些很基础的答案:嗯,我们有所有这些联系人和数据,等等。但如果你再问他们一个具体的用例,希望 AI 在他们的工作场所中实现什么,你会得到完全不同的答案分布——各种类型的答案像爆炸一样涌现。这很有意思,因为当你在使用这些模型和产品时,你面对的是一个盒子,你可以对它说任何话。这是它的魔法。但反过来,这也意味着你不知道该拿它做什么。在企业场景里,我认为很大一部分工作是要真正站在用户所在的位置:他们想解决什么用例,然后真正教他们如何用 AI 在那里获得杠杆。

I think there's no one-size-fits-all solution in enterprise. I remember in the early days of ChatGPT Enterprise, we would talk to customers. That was about a year after ChatGPT was released, and everyone was so excited to bring AI into their enterprise. There were all these teams being set up as AI deployment teams with enormous budgets. If you asked anyone what they were excited about, or what they were excited to solve, at first you'd get the baseline answers: yeah, we have all these contacts and data, and so on. But then if you asked them about a discrete use case they wanted AI to enable in the workplace, you'd get such a different variance—an explosion of different types of answers. It's interesting because when you use these models and products, you have this box, and you can say anything to it. That's the magic. But on the flip side, it also means you don't know what to do with it. In enterprise, I think a big part of that is meeting users where they are: what use cases are they trying to solve, and then actually teaching them how they can use AI to gain leverage there.

FDE与产品侧 FDE vs product side

Host

这在多大程度上区别于 forward deployed engineering(FDE)?还是说——

Do you meaningfully differentiate that from forward deployed engineering? Or—

Akshay

我觉得这里面有市场推广(go-to-market)的一面,也有产品的一面。

I think there's the go-to-market side of it and then there's the product side of it.

Host

我觉得你们得更多站在产品这一侧,对吧。

I think you need to be more on the product side, yeah.

Akshay

不管我们把 FDE 这种模式做得多好,归根结底,如果用户正看着自己的电脑或手机,那么产品里的工作就是要让他们能上手,并告诉他们该去哪里。

However good we get at FDE motion, I think at the end of the day, if we have a user looking at their computer or looking at their phone, it's our job in the product to be enabling them and showing them where to go.

采用与市场机会 Adoption and Market Opportunity

Akshay

所以我们对此感到非常兴奋。

So we're really excited about that.

Host

你觉得过去三年的采纳情况有什么变化吗?出现了阶跃性的变化,有了推理模型等等。企业仍然把 AI 看作黑盒,不知道拿它怎么办——还是说情况已经变了?

Do you think there have been changes over the past three years of adoption? There've been step function changes—you have reasoning models and whatnot. There's still the same problem: enterprise sees a black box and doesn't know what to do with it—or have things changed?

Akshay

我是说,我们看到采用量在大幅上升,对吧?每个人都超级兴奋。很多人——数以百万、千万计的人——都在用 ChatGPT,他们已经大致知道怎么和 AI 协作。但每当我们解锁一项新能力,比如现在看到的智能体,可能还是只有一小部分早期采用者真正理解:'你可以做任何事,只要确保有正确的上下文、连接正确的工具,并且你在监督它——任何事都是可能的。' 然后还有一个 10 倍甚至 100 倍更大的市场,他们还没理解、还没看到这一点。所以我觉得这就是下一阶段。回答你的问题,我认为采用已经发生并且增长很快,但我觉得机会远比这更大。这正是我们想在 ChatGPT Work 上发力的地方。

I mean, we're seeing a huge uptick, right? Everyone's extremely excited about it. Many people—millions, hundreds of millions—are using ChatGPT. They understand how to generally work with AI. But then every time a new capability gets unlocked—like now we're seeing with agents—there's probably still a contingent of early adopters who truly get it: 'You can do anything. You just have to make sure the right context is there, it's connected to the right tools, and that you're supervising it—but anything is possible.' Then there's this 10x or 100x bigger market that doesn't yet get that or doesn't yet see that. So I think that's the next stage. To answer your question, I think adoption is there and growing fast, but I think the opportunity is far, far bigger than that. That's where we want to play, especially with ChatGPT Work.

ChatGPT Work的起源 The Genesis of ChatGPT Work

Host

好,那我们直接跳到 ChatGPT Work 吧。大概一个月前才发布的。是什么决策过程促成了它?就是整体整合成超级应用——我们官方是这么叫的吗?你们还弃用了浏览器。就概括一下你过去几个月做这个东西的过程吧。

Yeah. Well, let's skip ahead to ChatGPT Work. It was only announced about a month ago. What was the decision process that led into it? There was this overall merging into the super app—is that what we're officially calling it? You deprecated the browser as well. Just summarize your last couple of months working on this thing.

Akshay

是的,感觉已经过了很久,但其实也就几个月。我觉得最显而易见的推动力是我们发布 Codex 的时候——甚至内部还在做 Codex 的时候。它真的让我们很惊讶。我们最近公布了一些数据,但当时 OpenAI 内部非开发者的采用出现了真正的拐点。在产品开发过程中,我们会去做用户体验访谈,跟内部员工聊。让我印象最深的是——比如你去找战略财务或市场的人,他们都在用 Codex 处理自己的工作,这部分很酷,但最让我印象深的是人们对自己用 Codex 是多么自豪,就好比……

Yeah, it feels like forever now, but I guess it's only been a few months. I think the most salient impetus was when we released Codex—or even internally had Codex. It was really surprising to us. We recently put out some stats on this, but there was this real inflection of adoption among non-developers at OpenAI. And through this product development process, we go to these UX sessions to talk to people internally. The thing that stuck out to me is—like you go talk to strategic finance or marketing or whatever—they're all using Codex for their use cases. That part's cool, but the thing that really stuck out to me is how proud people were that they were using Codex, like how...

Host

就好像,'我不该用它的,但我用了。'

It's like, 'I'm not supposed to be using it, but I am.'

Akshay

那意味着他们是最早接触这个新事物的人,但同时他们也觉得自己拥有了超能力,对吧?我们当时意识到,Codex 的力量、智能体的力量,再加上我们已有的庞大分发基础——那些已经了解并喜爱 ChatGPT 的人——怎么把这种力量展示给他们?怎么带给他们?这是个很难的产品问题,也很棘手,有很多种做法。所以我们后来称之为'融合',慢慢形成超级应用,最终在 ChatGPT Work 中推出:就是怎么做到这一点。但这源于一个最初的认知:这种能力不仅仅属于开发者——可能比我们想象的早得多——它可以扩展到所有人。

Was that they were early to this new thing, but it was also that they felt they had a superpower, right? And what we recognized then is that the power of Codex—the power of agents—combined with the massive distribution base of people who have come to know and love ChatGPT: how do we show that to them? How do we bring it to them? That's a hard product problem and a tricky thing—there are many ways to go about it. So that's what we called 'the merge' and the super app over time, and ultimately launched it in ChatGPT Work: how do we do that? It came from that initial realization that the power was not only for developers—much earlier than probably even we thought—it could be extended to everyone.

定位与生产力 Positioning and Productivity

Host

你怎么看待这些产品的不同定位?它是给谁用的,对吧?Codex 一开始甚至只是命令行,然后是应用。现在 ChatGPT 和 Codex 合并成了 ChatGPT Work。这是面向普通用户、企业,还是面向工作场景?你怎么定位它?

How do you see the products differently? Who is it for, right? Codex started out even as a CLI, then an app. Now there's a merge of ChatGPT and Codex into ChatGPT Work. Is it the opening for the average user, for enterprise, for work? How do you position it?

Akshay

我们想把它定位成:你做的是偏工作(worky)相关的事情。

I think we want to position it for when you're doing worky things.

Host

一时找不到更好的词,对吧?

Lack of a better word, right?

Akshay

我觉得生产力其实是我支持的支柱——这正是团队的名字。我们叫它生产力而不是企业或工作,是因为也有个人生产力,对吧?看 ChatGPT Work,我见过人们在个人生活中做那些严格来说不算工作的事,但智能体非常能干。最近有人在我们的 Slack 上发了一个例子:有人没收到包裹,然后拿到了亚马逊或快递公司拍的照片。他让 ChatGPT Work 去找包裹在哪。智能体非常执着:它拿了照片,查看了社区周边的大量房源信息,然后精确地找到了包裹所在的公寓楼,把信息告诉了他。所以我觉得有很多这种偏工作或生产力相关的事情。这就是我们想做的产品。你问到 Codex。我们认为 Codex 是一个持久的品牌,但我们有一个原则:我们不想让用户被困在一个无法获得产品全部力量的标签页或体验里。所以基本上你在桌面端 Codex 版产品里能做的,在 ChatGPT Work 里都能做,反之亦然。但我们做了一些有主见的产品决策,比如:如果要去拉一个仓库,你想把多少 git 状态暴露给最终用户?或者你想在多大程度上让用户看到智能体的思考过程,比如 diff 向前滚动,让你看到按钮背后的 diff 内容?然后安全方面,我们怎么考虑沙箱,以及确保我们在一种状态和另一种状态之间设置正确的默认值?所以这背后有我们的观点,但我们确实不希望用户需要选择自己身处哪种体验。

I think productivity is actually the pillar I support—that's the name of the team. The reason we call it productivity and not enterprise or work is because there's also personal productivity, right? With ChatGPT Work, I've seen people do things in their personal lives that you wouldn't technically classify as work, but these agents are super capable. One recent example someone posted on our Slack: someone had a missed package—they didn't receive it—and then got a photo of it from Amazon or whatever courier. They asked ChatGPT Work to find out where the package was. The agent was extremely tenacious: it took the image, looked at a bunch of listings around the neighborhood, and figured out exactly the apartment complex where the package was. It gave them the information. So I think there are all these things that are worky or productivity-related. That's what we want the product to be. You asked about Codex. We think Codex is a durable brand, but we have a principle: we don't want a user to get stuck in a tab or experience where they don't get the power of the product. So basically everything you can do in the Codex version of the product on desktop, you can do in ChatGPT Work, and vice versa. But we made some opinionated product decisions, like how much of the git state—if you're going to get a repo—do we expose to the end user? Or how much do we make the experience of seeing the agent's thinking, like diff forward, so you're exposed to the diff side of the button? And then on the safety side, how do we think about sandboxing and making sure we have the right defaults in one state versus the other? There are some opinions behind that, but we don't want the user to need to choose which experience they're in.

共享框架,不同体验 Shared Harness, Different UX

Host

这对 AGI 来说是个好目标,对吧?人们不想选择自己想要哪个版本的 AGI,他们只想让 AGI 替他们决定。能给个直接回答吗?我不太清楚:Codex 运行框架和 ChatGPT Work 运行框架是一样的吗?只是界面上的差异,还是实际上存在提示词层面甚至更深的差异?

That's a good goal for AGI, right? People don't want to choose what version of AGI they want; they just want the AGI to decide for them. Can I get a straight answer? It's not super clear to me: is the Codex harness and the ChatGPT Work harness the same? Is it just UI affordances, or are there actually prompt-level or even deeper differences?

Akshay

运行框架是一样的——运行框架是共享的。两个产品里我们都改进了运行框架,让它适合知识工作,尤其是涉及插件、计算机使用或产物(artifacts)的部分。无论你处在哪个体验中,都能获得这种能力。在界面方面,我们有自己明确的观点:当你处于 Codex 模式时,界面应该是什么样、应该如何表现。

The harness is the same—the harness is shared. In both products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins, computer use, or artifacts. You get that power regardless of which experience you're in. On the UX side, we have opinionated takes: when you're in Codex mode, what the UX should be and how it should behave.

比较Work与Codex模式 Comparing Work and Codex modes

Host

关于我提到的沙盒环境,有些东西底层的能力框架应该是一样的。其实我很好奇——我们能不能运行一个查询,看看它在两种模式下会有什么不同?

And some stuff around the sandbox that I mentioned, but the underlying harness of capabilities should be the same. Actually, I'm just kind of curious—maybe we can—is there a query that we can run that would look different in the two modes?

Akshay

是的,我试过让它创建一个退休计算器电子表格之类的,两种模式都试了。在 Codex 模式下,你可能需要在一个代码仓库里,但你会看到我正在创建的表格的差异,以及文件编辑。但在工作模式下,你看不到这些。

Yeah, I tried asking it to create a retirement calculator spreadsheet or something, in both modes. In Codex mode, you might have to be in a repo for this, but you'll see the diffs of the sheet that I was creating, and the file edits. But in Work mode, you won't be able to see that.

生产力团队结构 Productivity team structure

Host

我觉得这非常清晰。另外我还想深入了解一下你们的生产力团队。首先,除了生产力团队之外,高层团队还有哪些?生产力不是涵盖一切吗?

I think that's super clear. And then the other thing I wanted to dive into was the productivity team. First of all, what are the top-level teams other than productivity? Isn't productivity everything?

Akshay

我们有专注于 ChatGPT 的团队,也就是面向消费者的核心聊天体验。我不认为所有东西都是生产力。人们每天用 ChatGPT 来搜索、琢磨怎么给亲人写消息、思考如何学习一个新主题等等。里面还有太多东西:生成图片,以及数亿用户正在使用的许多聊天功能。显然这需要一个非常专注的团队。另外还有专注于企业、基础设施、API 之类的团队。

So we have a team focused on ChatGPT—the core chat experience for consumers. I don't think it's all productivity. People use ChatGPT every day for search, to figure out how to write messages to loved ones, to think about how to learn a new topic, and so on. There's so much more inside: creating images, and many other chat features that hundreds of millions of users are using. Obviously that warrants a very dedicated effort. And there are teams focused on enterprise, infrastructure, API, and things like that as well.

现场演示:双模式运行 Live demo: running both modes

Host

我来调出来。好的,我已经让两个都在运行了。

I'll bring it up. Yeah, I have both of them running.

Akshay

好。

Yeah.

Host

这是工作模式。这里有一个 Codex 版本。我选了 5.6,所以需要一点时间。我想我们让它先跑着,等完成了再看看一些差异。

This is Work. There's a Codex version here. I picked 5.6, so this will take a while. I think we'll just keep it in the background, and as they finish, we'll look at some of the differences.

Akshay

是的,但如果你马上切回 Codex 版本,你会发现它看起来很好。没错。它很动态,看起来就像在一个 Git 仓库里。你可能会漏掉一些东西,因为其中一部分是在真正的思维链里,包括那些改动以及我们的展示方式,但是——

Yeah, but immediately I think if you flip back to the Codex version, you'll see that it seems good. Exactly. It's dynamic; it seems like you're in a Git repo. You might miss some stuff because some of it is in the actual chain of thought with those changes and how we display that, but—

设计理念:为何融合体验 Design philosophy: why merge experiences

Host

有没有什么反直觉的决策?有没有一个东西你本来想发布,但收到反馈后觉得‘算了,我们不这么做’?这背后的思考是什么?

Is there an unintuitive [decision]? Is there a thing that you wanted to ship, and then you got feedback and you were like, no, let's not do it? What's the thinking behind that?

Akshay

你是指 ChatGPT 工作模式吗?

In ChatGPT Work?

Host

对。

Yeah.

Akshay

我觉得我们本来可以走的一个方向是,把这两种体验完全分开——比如做成不同的应用,甚至在同一个应用里做成完全不同的体验。那为什么要合并呢?大家显然都很喜欢 Codex,为什么要把这些产品放在一起?我觉得这里的直觉是:我们所有人的工作都在被 AI 剧烈地改变,坦白说每隔几个月就变一次。我觉得自己一觉醒来,做的已经和几个月前完全不是同一件事了。我的假设——或者说我们团队的假设——是,我们在这项技术中构建的一部分东西是给人们杠杆。比如你工作中比较平凡的部分,或者那些如果你能自动化,就能更快地分享更多想法、或者做到现在能做的事。正因为如此,这实际上可能会模糊人们之间的界限:有人只写代码,有人写战略文档,有人策划活动,有人做营销,有人做播客,等等,对吧?所以这些东西会随着时间推移而模糊。而试图根据‘你是谁’来划一条清晰的界线是会很困难的。我们应该让用户自由选择,但不要把他们框死。我们在这里做的很多工作——比如保持原语的一致性,例如插件在这个产品、ChatGPT 和云端是统一的——都是因为这个。这就是那个论点:最终这些东西会汇聚在一起。我们不想对何时进入哪种体验做太多规定,但我们也不想把任何人框住。

I think one direction we could have gone with this is keeping the experiences completely separate—like different apps, or even different experiences within the same app. Why merge it all? Codex is obviously loved. Why bring these products together? And I think the intuition here is that all of our jobs are changing dramatically with AI, frankly every few months. I feel like I wake up and I'm doing a completely different thing from what I was doing a few months ago. My hypothesis—or I should say our hypothesis—is that part of what we're building in this technology is giving people leverage. Things like the more mundane parts of your job, or parts that if you could automate, you'd be able to share more ideas faster, or whatever you're able to do now. And because of that, that might actually blur the lines between someone who's only writing code, or creating strategy docs, planning events, helping with marketing, doing podcasts, or whatever, right? These things are going to get blurred over time. Trying to draw a hard boundary based on who you are is going to be tough. We should enable users to choose, but we shouldn't box them in. A lot of the work that went in here—keeping the primitives the same, for example, plugins that are unified across this product, ChatGPT, and the cloud—was because of that. It's this thesis that eventually things are going to come together. We don't want to be prescriptive about when to be in either experience, but we don't want to box anyone in.

框架演进:ChatGPT对比Codex Harness evolution: ChatGPT vs Codex

Host

我在想,有没有一些用户非常习惯于旧的 ChatGPT 框架,而它现在实际上已经被 Codex 框架取代了。我无法想象那是什么样子,但可能他们更偏向对话那一侧。你能对比一下这两种框架吗?毕竟只有你见过两者。

I wonder if there are users who are very tuned to the old ChatGPT harness that is effectively now replaced by the Codex harness. I can't imagine what that was, but maybe they're on the more conversational side. Can you compare and contrast the two harnesses, since only you've seen it?

Akshay

是的,我觉得现有的 ChatGPT 框架今天仍然存在。它就在这个应用里。

Yeah, I mean, I think the existing ChatGPT harness still exists today. It exists in this app.

Host

就是经典模式,对吧?

The classic one, right?

Akshay

你只要新建一个聊天,不要进入工作模式,对吧?

You just start a new chat and you don't go under Work, right?

Host

对,如果你新建聊天并进入 Chat,你就是在用现有的——

Yeah, if you start a new chat and go to Chat, then you're talking to ChatGPT with the existing—

Akshay

——实例。对,我想就是那个意思。

—instance. Yeah, I guess, you know.

Host

对,所以这个不会去写代码,或者在行内执行。它不在行内沙盒里。

Yeah, so this one is not going to code. Or it's going to be inline. It's not in an inline sandbox.

Akshay

是的,实际上如果你的目标是创建电子表格,我们会尽量引导你进入工作模式。这是一个路由决策,抱歉。

Yeah, actually we do try to push you to go to Work if you're creating a spreadsheet. And this is a router decision, sorry.

Host

这是一个路由决策吗?

Is it a router decision?

Akshay

这是模型在做出的决策。它看到你能够或者试图做某件事,而这件事在工作模式下会得到更好的支持。不过我觉得你的问题是:ChatGPT 聊天框架的优势是什么?

This is the decision that the model is making. It sees that you're able to, or you're trying to do, something that would be better served in Work mode. But I think your question was: what are the advantages of the ChatGPT chat harness?

Host

更广泛地说——我基本上是想做一段关于框架工程的口述史。

It's more broadly—I basically wanted to do an oral history of harness engineering.

Akshay

嗯。

Mhm.

Host

对吧?ChatGPT 框架从——姑且称之为 o1 时代——一直陪我们到现在。而现在它实际上正在被 Codex 框架取代。两者有一定重叠,但我很好奇,如果说有变化,到底变了什么。

Right? The ChatGPT harness lasted us from, let's call it, the o1 era until now. And now it's being effectively replaced by the Codex harness. They're overlapping somewhat, but I'm curious what changed, if there is.

Akshay

我的看法是,这就像是一个持续的分化、融合、再分化、再融合的过程。在 Chat 里,对于我之前提到的很多用例——搜索或学习——我们其实是在优化延迟和个性化等不同方面。人们喜欢 ChatGPT 的原因,正是因为我们长期在优化这些东西。而在 Codex 上,我们学到的是:如果你给智能体一个像计算机一样无限灵活的环境,它能做出非常非常强大的事。所以当我们思考‘知识工作该选哪种模式’的时候,我们觉得把这个计算机环境带过来更自然,也许可以把计算机的某些细节对不习惯的用户抽象掉,但把同样的力量给他们。

My perspective on this is that there's a constant process of divergence, convergence, divergence, convergence. In Chat, for many of the use cases I was talking about—search or learning—we're really optimizing for latency and personality, among other things. The reason people love ChatGPT is because we've been optimizing for those things for so long. With Codex, what we learned is that if you give the agent access to an infinitely flexible environment as a computer, you can do really, really powerful things. So when we think about, okay, for knowledge work, which mode should we choose? It felt more natural to us to bring that computer environment, and maybe abstract some of the details of the computer away from users who aren't used to it, but give them that same power.

模型选择指南 Model selection guidance

Akshay

但归根结底,我觉得我们想把能力带到所有地方,对吧?我们希望在人所在的地方与他们相遇。所以我相信未来还会有工作要做,让各项能力在所有场景下同样强大。但这只是我们历史上在产品上一直关注什么、现在又在关注什么的问题。

But ultimately, I think that we want the power in all places, right? We want to meet people where they are. So, I'm sure there'll be work down the road in order to get things to be equally capable in all scenarios. But it's just a question of what we've been focusing on the product on historically, and what we're focusing on now.

Host

我想在这一点之外,除了 Harness 和工作中什么时候用 Codex、什么时候用 GPT,你还发布了新模型,对吧?有没有什么使用建议?人们喜欢精打细算,比如只在高度推理时用 Terra,而这种情况你也许会用 Soul,等等。

I think alongside that, outside of just Harness and when to use Codex and GPT at work, there's also the new models you've released, right? Any guidance there? People love to min-max what to use, like only use Terra on high reasoning versus, you know, you want to use Soul here, ignore...

Akshay

有 32 个选项。

There's 32 options.

Host

对,对,对。但话虽如此,对于那些正在扩展、在工作中尝试各种工具、却并不了解这些选项具体区别的人,你们有什么建议呢?

Yeah, yeah, yeah. But that being said, for people that are expanding, so productivity trying stuff for work that don't have the breakdown of what all this is, what's the advice, right?

Akshay

嗯,我是说,我觉得在给建议之前,首先一点是:没有这些模型,这一切都不可能实现。我想你之前问过工作的灵感是什么,早些时候我也提到我们在 Codex 上看到的东西,但那也是因为模型能力越来越强。现在这种情况再次发生,我觉得又是一个阶跃式的提升。回到建议的问题:我们希望默认设置就是最好的。我们希望默认值有明确的主张,所以我们选定了一个我们认为对所有人都最好的默认配置。同时我们也为高级用户提供了底层的选项。可以说现在选项可能太多了,我们正在努力简化。但你可以扩展推理级别,也可以在不同模型类别之间切换,如果确实需要的话。默认配置应该对大多数用例是最好的。所以我对大多数人的建议是坚持默认设置。然后,如果你遇到某个情况,想尝试不同的配置,如果你在成本端没看到效率,或在智能端没看到质量,那就可以改一下默认值,看看能不能得到更好的结果。但我们认为默认设置应该已经足够好。

Well, I mean, I think before the advice, the first thing is none of this would be possible without these models. I think you asked earlier what was the inspiration for work, and earlier I mentioned what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That's happening again. I think it's another step function jump now. And to answer the question on advice: we want the default to be the best possible. We want to be opinionated about the default, so we've chosen a default that we think is going to be the best for everyone. And we have options under the hood for power users. One could argue that there might be too many right now, and we're working on simplifying it. But you can extend the reasoning level, and you can change between the different model classes if you need to, but the default should be the best for most use cases. So my advice to most people would be to stick to that. And then, if you reach a situation where you want to try a different configuration, if you're not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think that the default should be good enough.

目标与超推理 Goal vs ultra reasoning

Host

我有……我想跟你探讨一下,因为你比我经验丰富得多。我最近一直在用 Soul Light,但开了 Goal。我的想法是,Goal 基本上是在增强推理力度,但会有更多的中途终止和来回。

I have... I'm just going to run something by you since you have way more experience than me. I've recently been doing Soul Light, but with Goal. With the idea that Goal basically augments the reasoning effort but with more terminations and turns.

Akshay

嗯。

Mhm.

Host

这么想对吗?相对于 Soul Ultra 或者 Soul 超高模式?

Is that a good way to think about it? As opposed to Soul Ultra or Soul, you know, extra high.

Akshay

是的,很难说。

Yeah. It's hard to say.

Host

因为这有点像交互效应。

Because it's like an interaction effect.

Akshay

没错。这里其实有种偏好,你作为个人,喜欢怎么和模型协作?你希望有多少次你说的那种“终止”,让你可以引导或确保它做的是对的?我觉得一般来说,大家应该尝试适合自己的方式。我认为用 Ultra 或类似多智能体配置,最适合那些极其复杂的任务,比如开放式探索,或者很适合并行化的任务。我觉得用 Goal 最适合那些你能够持续取得可验证进展的任务。但在我看来,大多数任务其实并不属于这两类,至少一开始是这样的。所以我认为最好的第一步是先用默认配置试试,然后再看你想往哪个方向走。

Exactly. There's a preference, you know, for you as an individual: how do you like to collaborate with the models? How many of those terminations, as you call them, do you want, where you can steer or make sure it's doing the right thing? I think generally people should try whatever works for them. I think using Ultra or multi-agent setups are best for tasks that are either incredibly complicated, like open explorations, or very parallelizable. I think using Goal is best for tasks where you'll be able to make consistent progress in a way that's verifiable over time. But I think most tasks don't actually fall into either of those buckets, at least when people are starting. So that's why I think the best first step is trying the default configuration and then seeing where you want to go from there.

UI滑块与默认模型 UI slider and model defaults

Host

对。你们做了一个滑块,实际上非常有助于减少那种不知所措的感觉。

Right. You guys worked on a slider which is actually super helpful for reducing the amount of panic.

Akshay

对,对。

Yeah, yeah.

Host

至少在移动端上挺好用。这里有个很漂亮的阶梯。

It's nice on mobile, at least. There's a nice ladder here.

Akshay

我还没试过。

I haven't tried it.

Host

所以你在那里可以看到高级视图,但如果点击高级视图……

So, you have the advanced view there, but if you click advanced view...

Akshay

就是一个简单的阶梯,对。

Just a simple ladder, yeah.

Host

非常漂亮,非常多彩。

Very pretty, very colorful.

Akshay

对,这里的想法是把它压缩到一个维度,哪怕实际上有多个维度,对吧?我们尝试为用户把它投射到单一的维度上。一端代表速度和效率,另一端代表质量和彻底性。

Yeah, the idea here was to reduce it to one dimension, even though there are multiple dimensions, right? Try to project it onto a single dimension for the user. Something that represents speed and efficiency on one side, and quality and thoroughness on the other side.

Host

我只是很困惑它这么常用 Soul。比如低段位(bronze)那档……

I am just puzzled that it uses Soul so much. Like, the lower bronze...

Akshay

我想那是 Spider,如果我没记错的话。哦,确实是。

I think it's Spider, if I'm not mistaken. Oh, it is.

Akshay

对,你看到我们这边了。所以他们把 Terra 预设为唯一的轻量级选项。

Yeah, you see our side. So, they preset Terra to only be the light one.

Host

我明白了。

I see.

Akshay

但我觉得其实很多人会……更多人应该用 Terra。一个原因是 Soul 总是算力不够用。

But I think a lot of people actually would—more people should use Terra. One, because Soul keeps running out of capacity.

Akshay

我就是原因,你知道吧。这里有我们 10 分钟的——

I'm the reason, you know. Here's 10 minutes of our—

Host

你看吧。

There you go.

Akshay

退休金计算器。

retirement calculator.

Host

哦,这就是那个 Excel 的东西——天哪,看这个。

Oh, that's the Excel thing we're—Oh my god, look at that.

Host

这是工作,然后 Codex 还在跑,所以我们稍后再回来。我觉得看看思维过程和推理会很有趣。而且,我想这大概用了 8 分钟在工作上,Codex 仍在运行。

That's work, and then Codex is still cooking, so we'll get back into it. I think it'll be interesting to actually see the thought process, the reasoning. And also, I guess this is 8 minutes on work. Codex is still cooking.

工件与发布协调 Artifacts and launch coordination

Host

对了,顺便问一下,你认识 Gabriel Chua 吗?他是 OpenAI 安全团队的。他给我看了这个,我特别震惊,这看起来就像 Excel。是的,它能编辑 Excel 文件。你从来没买过 Excel 授权,对吧?但不知怎么的,这个居然能用,而且是智能体式 Excel。

Yeah, and by the way, I... Do you know Gabriel Chua? He's part of the OpenAI safety team. He showed me this, and I was pretty shocked that this looks like Excel. Yeah. It edits Excel files. You never paid for an Excel license, right? But somehow this is workable, and it's agentic Excel.

Akshay

是的,我的意思是,我们这次发布大力推进的一件事就是工件(artifacts),对吧?在模型这边——我觉得如果你拿它和此前的 5.5、5.4 相比,你会看到这些工件的质量有了非常显著的提升。产品端也是如此。

Yeah, I mean, one of the big pushes that we made for this launch was artifacts, right? Both on the model side—I think if you compare this with 5.5 and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of these artifacts. And then also on the product side.

Host

用户体验端也很厉害。比如托管站点什么的,再也不需要自己去托管一个小网页了。

The UX side is also crazy. Like hosted sites and whatnot—no longer needing to host your own little webpage.

Akshay

哦,关于这个我有个故事。我可以单独讲。我需要先处理这里的画面,但我们可以稍后再切回来说这个。

Oh, I have a story about that. I can do a separate thing. I'll need to take the visuals here, but we'll cut to that later.

Host

我想,是因为你们在做出这一重大举动,而且你们在同一天发布了 5.6 和 ChatGPT Work,所以是有协调的吗?模型训练团队和产品框架(harness)团队之间互相有影响,还是说发布日期刚好撞在了同一天?

Was there coordination, I guess, because you were making this big move, and you launched 5.6 on the same day as ChatGPT Work? Was there influence between the model training teams and the harness teams, or did launch days just happen to line up on the same day?

Akshay

我想,你知道,我们与研究团队合作非常紧密。我的意思是,我觉得这是这份工作最神奇的地方之一——也是最有乐趣的部分。

I think, you know, we collaborate heavily with the research teams. And I mean, I think that's one of the most magical parts of the job—the most fun parts of the job.

工件与迭代 Artifacts and iteration

Akshay

但确实,我就以 artifact 为例:你在表面之下看到的很多工作,都是为了确保我们有合适的基础设施来训练模型,让它在这一点上做得更好;而在产品层面,则是让用户能与模型在 artifact 上进行协作的正确体验。实际上,这个整个查看器——这里的直觉是,你并不一定就不需要 Excel 许可证。这是第一阶段。这大概不是你在做退休计算器时想要的结果。你想要迭代,而且当你看到它时,如果这个东西与你实际看到的、或者你发给 Sean 后你的同事会看到的东西高度保真,我觉得那会让迭代变得容易得多,也会让你在迭代方面信任这个产品。

But yeah, just using artifacts as an example, a lot of what you're seeing underneath the hood involved work to make sure we had the right infrastructure to train the models to get better at this, and on the product side, the right experience for users to collaborate with the model on an artifact. In fact, this whole viewer—the intuition here is that it's not necessarily that you wouldn't need an Excel license. This is stage one. This is probably not what you meant when you're making a retirement calculator. You want to iterate, and when you're seeing it, if this thing is high fidelity to what you'd actually see, or what your co-workers would see if you were to send this to Sean, I think that makes it so much easier and makes you trust the product in terms of iteration.

多人共享与权限 Multiplayer sharing and permissions

Host

你说“同事会看到”的时候,你设想的是 artifact 上的多人、多团队协作吗?你们有什么想法吗?分享,对吧?是的。

When you say co-workers would see, do you see a multiplayer, multi-team collaboration with artifacts? Anything you guys think about? Share it, right? Yeah.

Akshay

是的,这是我们正在积极思考的事情。我们内部注意到一件事——不透露太多路线图的话——就是很多时候有人会给我发消息问问题,我会去问 ChatGPT Work,然后把答案发回给他们。

Yeah, it's something we're actively thinking about. One thing we've noticed internally—without talking too much about the roadmap—is that there are many times when someone will ping me about something, I'll ask ChatGPT Work the question, and then I'll ping them back the answer.

Host

最简单的就是,你知道,我们三个都在同一个托管会话里。

Like the simplest would be, you know, the three of us are just all on one hosted [session].

Akshay

没错。而且我会想,我是不是真的需要在这个环节里。也许只是重新措辞他们的问题,或者从某些语境里提取信息之类的。但当我给他们回复答案时,这个过程也会有信息损耗。我给的只是我对 ChatGPT Work 生成内容的理解。但在表面之下,有这么多上下文——比如现实世界的信息——可能会很有用。所以答案就是预先响应每一个进来的请求。

Exactly. And I'll think about whether I was even required in this loop. Maybe it was rephrasing what they were asking, or pulling from some context or whatever. But when I gave them back the answer, that process was also lossy. I gave them just my interpretation of what ChatGPT Work cooked up. But underneath the hood, there's so much context—like the real world and stuff—that could be interesting. So the answer was to preemptively respond to every inbound request.

Host

不,这确实就是我的工作内容,有时候我就是干这个的。你复制粘贴,然后你就成了 AI 到 AI 的消息转发服务。

No, it's just like literally this is what I do sometimes as my job. You copy-paste, and then you're just a message forwarding service from AI to AI.

Akshay

我觉得这挺有意思的,对吧?它帮助人们理解你可以提出什么样的请求、可以委派什么。很多时候人们只有试过才意识到,或者有人演示给你看,你才会觉得“哦,好吧,好吧”。

I think it's interesting, right? It helps people understand the capability of what you can ask and delegate. Oftentimes people don't realize until they try, or someone shows you, and then you're like, 'Oh, okay, okay.'

Akshay

我明白了。是的。我觉得这里还有一个轻微的安全问题:基本上你就是权限层。比如,我可以查询你查询的所有东西,也能得到自动化回复,但我可能不该看到它。而我根本不会知道,因为我不该知道我不知道的事。

I see. Yeah. I think there's also a light security issue where basically you're the permissions layer. Like yes, I could query everything that you query and get an automated response, but maybe I'm not supposed to see it. And there's no way I would know, because I'm not supposed to know what I don't know.

Akshay

尤其是,你知道,用 ChatGPT Work 时,当你让它连接你的插件,它会从你的本地文件之类的地方拉取信息。智能体能访问到的上下文是非常私密的。这是我们需要保护的东西。所以这绝对会是一个挑战。

Especially as, you know, with ChatGPT Work, when you're asking it to connect your plugins, and it's pulling from your local files and stuff like that, the amount of context that the agent has access to is deeply personal. That's something we need to preserve. So that'll definitely be a challenge.

站点与知识工作 Sites and knowledge work

Host

有 Excel、有 PowerPoint、有 Docs——工作上的三大件。你还想过哪些其他工作格式?你知道,你显然在 Airtable 工作过。未来会不会有类似 OpenAI Airtable 的东西?如果你们最终去做,那会是什么样子?

There's Excel, there's PowerPoint, there's Docs—the grand trio of work. What other formats of work do you think about? You know, obviously you worked on Airtable. Is there a future where there's like OpenAI Airtable? What does that look like if you ever ended up doing it?

Akshay

这是个非常好的问题。你没提到的一个是 Sites,我觉得那是这次发布的重要部分。Sites 有一个方面大家经常谈起,尤其是在 Twitter 或 X 上,就是那种原型工具。实际上,这次发布就体现了这一点。你们之前提到的模型滑块,几乎完全是在 Sites 里开发的。设计、工程和产品之间的协作就是在一个 Sites 上进行的,我们可以在里面摆弄各种操作方式,感受它。但另一个比较少被提到的方面是,Sites 是知识工作的 artifact。前几天我还在和我们公司财务团队的人聊天,我们提到,现在他们每个月团队协作的报告,过去都是在幻灯片和电子表格里,现在就直接放在 Sites 里。Sites 成了他们跨团队协作的机制。原因是它的带宽更高。像 PowerPoint 和 Excel 这些工具有一定的灵活性,但到某个地步你就会触到边界——要么是人不知道怎么用某个功能,要么是产品本身不支持。但用 Sites 的话,你可以做任何事情,提出任何要求,然后都能得到。一旦人们见识到这种神奇之处,我觉得它真的很有价值。

It's a really good question. One you didn't bring up was Sites, and I think that was a big part of this launch. There's one side of Sites that people commonly talk about, especially on Twitter or X, which is this sort of prototyping tool. Actually, we saw that happen with this launch. Even the model slider that you guys were referencing earlier was developed almost fully in a Sites. The collaboration between design, engineering, and product on that was on a Sites, where we could play with the affordances and figure out how it feels. But the other aspect that's a little less talked about is that Sites is an artifact for knowledge work. I was actually talking to someone the other day who was on our corporate finance team, and we were mentioning how, now when they have these reports they're working on as a team month-to-month, historically those were in slide decks and spreadsheets, and now they're just in Sites. Sites is the mechanism they collaborate across the team. And the reason is it's somewhat higher bandwidth. At some point, these tools like PowerPoint and Excel are infinitely flexible, but you reach the boundary of either a human not knowing how to use some feature, or the product itself not supporting it. But in the case of a Sites, you can do anything, you can ask for anything, and you can get that. Once people see that magic, I think it's been really valuable.

案例:构建Strata Case study: Building Strata

Host

是啊,我给你看看我的案例。这涉及所有热门话题:ChatGPT Work、5.6、token 亿万富翁、token maxing、Sites 还有自动研究。我是这个叫 Strata 的游戏的粉丝。它基本上就是一个小的桌面游戏,和那种叠在上面的物理积木一起玩。所以上周末我拍了 30 张照片,直接扔进 ChatGPT。1.7 亿 token 之后,出来一个 Sites,里面有完整可玩的东西,还有 3D 积木摆放等等。因为它需要物理积木,我需要朋友来训练它,让他们变厉害,这样我就能跟他们玩。但我也可以在上面训练 AI。那就是你的自动搜索——这进入自动研究。所以你想训练自己的 AI,然后让它们互相自我对弈。我需要设置两个 AI。所以这是 AI 对 AI,它们会自我对弈。显然 AI 一开始很烂,然后你想定义一个损失函数,让它们变好。我可不想监督这一切。我当时在圣马特奥参加一个会议。结果我做的是自动研究,并创建基准测试。参数太多了,我读不过来。所以我开始让它生成一个 Sites,它就创建了这个实验室面板。它创建的 Sites 有快捷方式吗?

Yeah, let me show you my case study. This involves all the hot topics: ChatGPT Work, 5.6, token billionaires, token maxing, Sites, and auto research. I'm a fan of this game called Strata. It's basically a little board game you play with physical blocks that come on top of it like that. So over the weekend I took 30 photos and just threw it into ChatGPT. 1.7 billion tokens later, out comes this Sites with a fully playable thing with 3D block placement and everything. Because it requires physical blocks, I needed friends to train on it so they can get better, so I can play against them. But also I could do things like train an AI on it. And that's your auto search—that gets into auto research. So you want to train your own AIs and then make sure they self-play against each other. I need to set both the AIs. So this is AI versus AI, and they're going to self-play. Obviously the AIs start out bad, and then you want to define a loss function and get good. I wasn't going to supervise all this. I was down in San Mateo attending a conference. What I ended up doing was auto researching on this and creating benchmarks. And there were just way too many parameters for me to read. So I started asking it for a Sites, and it created this lab panel. Where is there a shortcut for a Sites that it's created?

Akshay

呃,你应该能在侧边栏找到 Sites——侧边栏顶部,左侧边栏。

Uh, you should be able to go in the sidebar to Sites—top of the sidebar, left sidebar.

Host

这个吗?哦,左边?

This one? Oh, left?

Akshay

对,一直滚到最上面。

Yeah, just scroll all the way to the top.

Host

哦,哦,上面写着 Sites。

Oh, oh, it says Sites.

Akshay

对。

Yeah.

研究工件作网站 Research artifacts as websites

Host

哦,出来了。对。哦。

Oh, there you go. Yeah. Ooh.

Akshay

所以,它会生成网站。我不太确定这是不是我想要的,但让我给你看看它生成的东西。我认为作为研究产出,必须准确传达正在做的事情。这个东西最终被我发布了出来。后来我把它从 Sites 上移走,因为我想要更强大的数据库和基础设施,而 Sites 满足不了我。但这是可以拿来摆弄的研究产出,比如思考你在给 AI 训练调什么超参数。我在尝试做缩放定律之类的东西,还做了各种游戏优化。而且你可以直接把它作为研究产出丢出来——我不需要再读 ChatGPT 的输出,我读网站的输出。但随之而来的还有海量内容。你看这个东西多长,数字太多了,挺让人头疼的。所以我得从那里开始整理。但这是从 markdown 出发的一个有趣转变。

So, it creates the sites. I don't think this is exactly what I wanted, but let me show you what it popped up. I think as a research artifact, it's very important to communicate exactly what is being done. This outputs this thing which I eventually started publishing. So I moved it off of Sites because I wanted more database and infrastructure than Sites afforded me. But this is research output that you can start to mess with and think about, like, what hyperparameters you're tuning for training your AIs. I was trying to make scaling laws and everything, and doing all sorts of game optimization stuff. And the fact that you can just throw this up as a research artifact—I no longer need to read ChatGPT output, I read site output. But then there's also huge sprawl. Look at how long this thing is, there are so many numbers, it's pretty overwhelming. So then I have to start putting in from there. But it's an interesting transition from markdown.

Host

对,实际上你不是在输出 markdown,而是输出了一个完整的可用网站。

Yeah, actually, instead of putting out markdown, you're putting out a whole functional site.

Akshay

我觉得 markdown 并不是很适合阅读,对吧?还不如直接写个 HTML 网站。而且我觉得你可以在定制上做很多文章。你有技能(skills)来解释你想要什么。我注意到它们非常冗长,其实很多信息我并不需要。

I think markdown just isn't optimal for people to read, right? Might as well just write an HTML website. And I don't know, I think you can do a lot with customizing this. You have your skills that explain what you want. Like I noticed they're quite verbose. I don't need a lot of this information.

Host

冗长。

Verbose.

Akshay

把网站放在旁边的好处就在于,你可以不断迭代,决定哪些要、哪些不要,对吧?

The nice thing about having a site side by side is you just iterate on what you want and what you don't, right?

实践中的灵活工件 Flexible artifacts in practice

Host

对,我不知道这是否会唤起你关于内部运作方式的什么故事。我这样做对吗?

Yeah, I don't know if that triggers any stories for you about how it's run internally. Am I doing this right?

Akshay

是的,我觉得这是一种我们正在看到的、各种团队都在使用的工作流。以前的标准产物是幻灯片之类的东西,现在变成了网站。而且网站因为就是 HTML,所以具有无限的灵活性。所以如果你想让某样东西更突出,比如在幻灯片里会显得很刻意,你可以让它成为主视觉,对吧?我觉得人们开始意识到这一点了。显然,要让这些东西更容易协作,还有很多工作要做。你提到它们很长、很冗长,可以拆分。我相信我们肯定有办法解决。

Yeah, I mean, I think this is a workflow that we're seeing all different types of teams use. The canonical artifact that was previously a deck or something is now becoming a site. And with a site, because it's just HTML, it's infinitely flexible. So if you want to give more prominence to a certain thing—in a slide deck it would feel forced—you can have it be the hero image, right? I think people are starting to see that. There's obviously more work to be done to make these things easier to collaborate on. You mentioned that they're very long and verbose and could be broken up. I'm sure there's something we can do about that.

Host

长。

Long.

Akshay

对,对。但我们开始意识到,这是一种非常有趣的格式,而且比他们以前用的灵活多了。

Yeah, yeah. But I think we're starting to see that this is a really interesting format for people to use, and it's much more flexible than what they had before.

设计产品来构建产品 Designing a product to build products

Host

我觉得你的工作也变得有点元了。你不是在设计产品,而是在设计一个用来制造产品的产品。我很好奇你是怎么应对的。

I think your job also becomes kind of meta. You're not designing the products; you're designing a product to make products. I'm curious how you manage that.

Akshay

我觉得我们在审视用户体验时一直在思考的一件事,就是如何在简单性和能力之间取得平衡。如果我们设计的是一个像你说的用来构建东西的产品,你能构建的东西非常多,但我们不能把这一切都摆在你面前,因为你会被淹没。

I think one thing we've been thinking a lot about when we look at the UX is how to balance simplicity with capability. If we're designing a product, like you said, that is meant to build things, you can build so many different things, but we can't put all that in front of you because you'll get overwhelmed.

Host

嗯。

Yeah.

Akshay

所以我们在 ChatGPT 上也遇到过类似的问题或挑战,尤其是现在能做的事情太多了。我们一直在寻求的平衡是:如何给用户足够的界面空间,让他们能够表达自己,告诉智能体自己的需求,验证它是否使用了正确的工具、是否从正确的来源获取信息,然后界面又能及时让开。以及如何构建一个这样的系统,让我们能够展示能做什么,而不是告诉人们能做什么?因为这在很大程度上取决于他们如何发现下一个用例,然后再下一个,如果他们真的想让 AI 赋予自己超级能力的话。

And so we had similar problems or similar challenges even with ChatGPT, but especially now when there's so much that can be done. The balance we're constantly trying to strike is: how can we give the user enough of a UI surface where they can be expressive, they can tell the agent what they need, they can verify it's using the right tools and pulling from the right sources, but then it gets out of the way. And how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is going to be about how they discover the next use case and the next one after that, if they really want to be super powered by the AI.

Host

嗯。

Yeah.

沿途测试工具 Testing tools along the way

Host

有意思。我觉得每个人都有自己的做法,对吧?我也做了一个类似的版本,同一个游戏。我没有拍任何棋盘或规则书的照片。我直接丢进去目标。18 分 53 秒之后,用掉了大量 token,我得到了一个类似的版本。显然没有那些自动研究之类的功能,但你知道的。

It's interesting. I feel like everyone also has a different way to do it, right? I made a similar version of this, same game. I didn't take any pictures of the board or the rulebook. I just threw in the goal. 18 minutes 53 seconds later, a lot of tokens later, I got a similar version. Obviously not with all the auto research and whatnot, but you know.

Akshay

你总得赶上所有最新潮流。

You got to do all the latest trends.

Host

而且,对,我是用 Codex 做的。但挺有意思的,对吧?

And yeah, I did it with Codex. But it's interesting, right?

Akshay

对。而且这显然是 GPT Image 在生成头像。这对游戏设计非常有用。很多游戏设计师都很喜欢 GPT Image。

Yeah. And this is obviously GPT Image generating the avatars. Very good for game design. A lot of game designers are really into GPT Image.

Akshay

我觉得更大的收获可能是,我们做这些事的原因主要是为了测试工具,对吧?比如,这是在 5.6 发布之前的一次测试。我之前在 5.5 上做过这个游戏。让我不再需要喂它规则的能力,这是一个相当小众的游戏,它自己找不到怎么做。

I will say the broader takeaway probably is that the reason we do this is more so just to test the tools, right? Like, this was also a test before 5.6 came out. I had done the game on 5.5. The ability for me to no longer need to feed it the rules—it's a pretty niche game—it couldn't find how to do this on its own.

Host

哦,对。

Oh, yeah.

Host

呃,5.6?

Uh, 5.6?

Akshay

自动分发。所以我才特别想测试 5.6 的能力。

Auto distribution. That's why I was also very keen on testing the 5.6 capability.

Akshay

但你知道,随着新东西不断涌现,这些只是我们用来测试事物的支线任务,对吧?

But you know, as work comes out, as new things come out, these are just our side quests to test things, right?

展示而非讲述 Show, don't tell

Host

对。我觉得这有点像私人邮件,但也没那么私密。不过也很有价值,因为现在你可以把它发给朋友们。而且我是通过看到这个才知道这个游戏的。

Yeah. It's some kind of private email, I guess. Which is not all that private. But it's also valuable because now you can send this to your friends. And I mean, I learned about this game through seeing this.

Host

这游戏很难。他做得非常好。

Hard game. He's very good.

Akshay

呃,没人跟你竞争的时候赢是挺舒服的。但没错,这是一个经典的强化学习问题:用自对弈来引导你的游戏 AI。对,你看工作和个人生活多容易互相渗透。我做的东西虽然是私人的,但直接启发了我共事的人,因为我把它展示给他们看了。他们会说:“哦,GPT 还能做这个?”我猜这就是增长策略。

Uh, it's good to win when no one is competing with you. But yes, it's a classic reinforcement learning problem of self-play bootstrapping your game AI. Yeah, you see how easily work becomes personal and personal becomes work, because the thing I do for personal actually directly informs the people I work with, because I showed it to them. They were like, 'Oh, you can do that with GPT?' Which, I imagine, is the growth strategy.

Akshay

对。“展示而非讲述”是很重要的一环。我觉得我们还没有完全破解它:向人们展示他们能用这个产品做的一切,而不是试图通过文章或新手引导之类的东西来教他们。

Yeah. The 'show, don't tell' is a big piece. I think we still haven't fully cracked it—showing people all the things they can do with the product versus trying to teach that to them through articles or onboarding or whatever.

Host

对。所以,在他们需要的时刻出现。

Yeah. So, meeting them in the moment.

Akshay

这对我个人来说是个职业风险,因为我以前是做开发者关系的,对吧?那里的工作就是展示。然后你就会想:“什么叫你不需要?其实你的工作是告诉。”

It's a career risk for me, because I used to be in developer relations, right? Where your job is to show. And then you're like, 'What do you mean you don't need? Actually your job is to tell.'

Host

嗯嗯。

Mhm.

Akshay

然后产品的人会说:“如果我们的产品足够直观,那我们就不需要你了。”所以就这样。

And then the product people are like, 'Well, we don't need you if our product is intuitive enough.' So.

定制展示 Tailoring the showing

Akshay

是的,这就是模型的魔力。你可以根据用户的具体需求定制讲述或展示——他们关心什么、过去做过什么、在采用旅程中正处于哪个阶段。我认为这将是一个巨大的机会。

Yeah, I mean, that's the magic of the models. You can tailor the telling or the showing specifically to what the user needs—what they care about, what they've done in the past, exactly where they are on the adoption journey. I think that's going to be a huge opportunity.

Host

现在定制展示似乎越来越容易了,对吧?人们有不同的使用场景。虽然你说过你不想把人分到不同的桶里,但针对不同类别的人其实也不难。但问题是——你说你的团队更广泛地关注……你用的词是什么?生产力?

Seems easier and easier now to tailor custom showing, right? People have different use cases. As much as you said you don't want to segment people into different buckets, it's also not that hard for people in different categories. But the question is—you said your team is more broadly on... what was the term you used? Productivity?

Akshay

生产力。

Productivity.

Host

生产力,现在基本上就是 Work。

Productivity, which is now Work, basically.

从开发者到所有人 From developers to everyone

Host

是 Work 吗?还有我们没有触及的另一类用户群吗?会不会有一群人,他们用的不是 ChatGPT、Codex 或 Work?有没有更多大众还没被触达的?

Is it Work? Is there another distribution we're not hitting? Is there a group of people that will have something different from ChatGPT, Codex, or Work? Is there more that the mass isn't targeting?

Akshay

我把它看作一个分阶段的过程。愿景是给每个人带来有用的智能体。我们从开发者开始。开发者历来是早期采用者——他们愿意忍受更多摩擦,自己搭好环境等等。Codex 就是从这里起步的。我认为下一个机会是我们所说的通用知识工作——也就是开发者周围的所有其他职能。从开发者转向这个群体时,当然会面临固有的挑战——就是我们说的‘展示而非讲述’,让产品更容易理解,带来对这个群体比对开发者更重要的新能力,比如工件、计算机使用等等。然后同样的经验也适用:就像我们把从开发者那里学到的经验带到通用知识工作一样,下一阶段是把这些经验带给每一个人,无论他们在生活中做什么。我们已经看到一点端倪了——你举的那个游戏例子,就处在娱乐、个人生活和职业生活的边界上。我上班时全天使用 ChatGPT,在家也用它做所有事情。前几天我用它来制定一周的食谱,并保存在它的计算机环境中,方便我随时回去查看。是不是每个人都这么做了?可能还没有,因为我们还在努力,但最终我们希望能让大家都做到。

I see it as a sequencing. The vision is to bring useful agents to everyone. We started with developers. Developers are historically early adopters—they're willing to put up with friction, set things up, etc. That's where Codex started. I think the next opportunity is what we call general knowledge work—all the other functions around developers. When you go from developers to this segment, there are inherent challenges, obviously—this show-not-tell thing we're talking about, making the product more understandable, bringing in new capabilities that matter more for this cohort than they do for developers—things like artifacts, computer use, etc. Then the same learning applies: just like we took the learnings from developers and brought them to general knowledge work, the next stage is taking those learnings and bringing them to everyone, no matter what they're doing in their lives. We're already seeing a little of that—the game example you have is on the border between fun, personal life, and professional life. I use ChatGPT at work full-time, and at home for everything. The other day I used it to come up with a meal plan and saved it in its computer environment so I can keep going back to it. Is everyone doing that yet? Probably not, because of things we're working on, but eventually we want to get people there.

Host

ChatGPT 生活。

ChatGPT life.

Akshay

对,没错。ChatGPT 做饭。但我觉得那里有很多机会。我的看法是:我们在软件工程领域打下了基础,然后我们会把同样的经验从软件工程带到知识工作,再带给所有人。

Yeah, exactly. ChatGPT cooking. But I think there's a lot of opportunity there. I see it as: we built the foundation in software engineering, and we're going to take the same learnings from software engineering to knowledge work, and then to everyone.

高级用户建议 Power user advice

Host

你有什么高级用户建议吗?我觉得有一群人把 ChatGPT 当生活,用它做所有事,24/7 在线。而那批人和另一批人之间有一点差距——另一批人就像‘好吧,我上班会用,偶尔用用,有时问点问题’。你有没有什么建议、心得体会或推荐,或者你发现的、有助于弥合这个差距的要点?

Do you have any power user advice? I feel like there's a group of people that will live in it, use it for everything, stay on it 24/7. And then there's a bit of a gap between that crew and people who are like, okay, I use it for work, I use it occasionally, sometimes I ask questions. Any advice, any learnings, anything you recommend—takeaways that help bridge that gap?

Akshay

我看到了几点。一是,它确实能帮你拓宽对可能性的想象。这对我来说也是一个学习过程。技术发展得太快了——甚至三个月前,我还会说‘模型绝不可能做到’,现在却发现‘哇,真的可以了’。

I think a couple things I've seen. One is that it really helps to broaden your imagination of what's possible. This has been a learning even for me. The technology has progressed so fast—even three months ago, I'd say no way the models can do this. Now it's like, wow, they actually can.

Host

但你们有……

But you have...

Akshay

我们现在内部正在做绩效评审。人们总说这种事是模型的强项——有句老话:‘我不想写评审,就拿 AI 来写。’但写出来的东西也需要被评估。没错,确实如此。说正经的,以前那基本上就是垃圾内容,有点用,但没什么效率。现在我发现模型可以比我做得好得多,尤其是在这个环境里。它能获取背景信息,了解大家在忙什么、做了什么有影响的事,还能指出我可能都没看到的成绩。它可以接触到所有东西——代码、代码评审、Slack,应有尽有。所以在这个领域它非常强大。而六个月前,上一次评审周期时,我试着用它,一点用都没有。这次却出奇地有用。所以我认为,不断拓展‘什么是可能的’这个边界——即使你以前试过——可能是我最大的建议。另一件事是:你投入得越多,尤其是在这个模型可以访问你电脑上或 ChatGPT Work 里所有内容的环境里,你就能随着时间创建工件并保存在你的资料库中,让模型持续访问它们。你给它的信息越多,无论关于你的生活还是工作,它就越有价值。而且它会在让你意想不到的方面变得更有价值——它可能会主动调用上下文,以你可能没想到的方式。但它需要先能访问那些上下文。

We're going through our review cycle right now internally. People always talked about this as something the models are good at—there's a cliché: 'I don't want to write reviews, I just use AI to do it.' But then it needs to be evaluated as well. Yeah, exactly. In all seriousness, before it was basically slop—helpful, but not super productive. Now I've found the model can do a much, much better job than me, especially in this environment. It pulls context on what people are up to, the things they've done to make a difference, and highlights wins I might not even have seen. It has access to everything—the code, code reviews, Slack, everything. So it's incredibly powerful in that domain. Six months ago, the last time we did this cycle, I tried using it and it wasn't helpful at all. This time it's been incredibly helpful. So I think continuing to push the frontier of what's possible—even if you tried something before—is my biggest piece of advice. The other thing is: the more you put in, especially in this environment where the model has access to everything on your computer or in ChatGPT Work, you can create artifacts over time and save them in your library, and the model continues to have access to them. The more information you give it about whatever domain you're in—your life or your work—the more valuable it becomes. It'll become more valuable in ways that might surprise you; it might proactively pull context in ways you hadn't thought of. But it needs access to that context first.

审查礼仪与智能体搜索 Review etiquette and agentic search

Host

有一件事我想谈谈——就是评审这件事。因为我还是觉得那是一个非常敏感的事。你们是创始人,你们管理员工,也招聘过员工。作为管理者,我非常不愿意发布任何 LLM 生成的内容,尤其是涉及人的时候,因为那会让人觉得你不够用心。大概在 OpenAI,人们显然更能接受被 GPT 来评价。但这件事有什么非正式规则吗?比如有什么礼仪?

One thing I wanted to talk about—the review stuff. Because I still think that's a very sensitive thing. You're founders, you manage people, you've hired people. As a manager myself, I'm very reticent to put out any LLM-generated things, especially when it comes to people, because it feels like you don't care. Presumably at OpenAI, people are obviously more open to being basically rated by GPT. But are there any unofficial rules around this? What's the etiquette?

Akshay

哦,我觉得礼仪就是:我永远不会只用 AI 写东西,然后把它当作给某人的评审。我说的是更偏收集背景信息。

Oh, I think the etiquette is that I would never write something only via AI and present it as a review for someone. What I was talking about is more like gathering context.

Host

对。这正是它……

Yeah. That's the place where it's...

Akshay

所以它只是搜索。是智能体式搜索。

So it's just search. It's agentic search.

Host

就像智能体式搜索,但你可以比以往更灵活地定制和引导它。

It's like agentic search, but you can tailor and steer it much more capably than before.

AI工具飞轮 The Flywheel of AI Tools

Akshay

关键是,这里有一个飞轮效应。因为 Codex,人们能做很多事;因为 ChatGPT,现在能做事的人比以往任何时候都多得多。如果你能做更多的事,也就容易错过一些东西。所以我认为我们需要用同样的工具来跟上人们产生的影响,并了解我们在哪里能帮上忙。

And the thing is, there's a flywheel happening. Because of Codex, people are able to do things. Because of ChatGPT, more people are able to do so much more now than ever before. And if you're able to do so much more, it's easy to miss things as well. So I think we need to use these same tools to keep up with all the impact that people are having, and understand where we can be helpful.

大规模缺失之物 Missing Things at Scale

Host

我想说的是,显然我经营着一家小公司,所以搜索很容易。但在 OpenAI 这个规模,你们在 Slack 里发那么多消息,你觉得它会漏掉东西吗?

I think the thing is, obviously I run a small company, so it's easy to search. But at the scale of OpenAI, with the many messages you guys put in Slack, do you think that it misses things?

Akshay

可能会,但我觉得我自己也会漏掉东西。

Probably, but I think that I also miss things.

Host

那没关系,对吧?它需要达到这个水平——人类水平。

It doesn't matter, right? It's as it needs to be — human level.

Akshay

这是相对的,对吧?是的。有时候它能找到你不会找到的东西,这就很好。比如现在,我的 Codex 系统提示是这样设置的:我的每个项目都有一个独立的 notes 文件。它会把学到的东西写到那里。然后全局的提示可以从所有这些文件里获取。所以有时候它会说,“哦,你四个月前做过这个项目,这里有一条我们记过的笔记。”然后它会随机把它拉回上下文里,这是我从来不会做、也没想过的。我就会想,“好吧,这真是超人类。”而且它可以节省好几个小时去整理东西,或者找到已经做过的事。虽然它可能会漏掉我会做的事,但当它找到东西时非常有用。我的解决方案非常不复杂,就是 markdown 文件,随时会被拉取。

It's relative, right? Yeah. Sometimes it's nice when it finds things you wouldn't. Like right now, my Codex system prompts are set up in such a way that every project I have has a separate notes MD. It just writes learnings there. Then the global one can pull from all these. So sometimes it'll be like, “Oh, there's this project you did like four months ago. Here's a note that we had.” And it randomly pulls it back into context that I would never do, haven't thought about. And I'm like, “Okay, this is quite superhuman.” And you know, it'll save hours on chunking stuff or find something that's already been done. As much as it might miss stuff I would do, it's very useful when it finds stuff. And I have a very non-super-engineered solution to this. It's just markdown files that get pulled whenever they want.

梗图自动化与模型幽默 Meme Automation and Model Humor

Akshay

是的,我其实有一个关于这个的有趣轶事。最近在为这次发布做准备,团队已经努力了好几个月。在那段时间里,Slack、文档和其他地方有大量的对话和讨论。团队里有一位成员设置了一个定时自动化任务,去查看所有正在发生的事情,找出最好的梗,然后发到我们的共享频道里。这里有两件很酷的事。第一,我认为模型随着时间推移真的开始变得有趣了,而一年前完全不是这样。第二,就像你说的,它们会以你意想不到的方式找到东西,并建立你可能没想到的联系。这对生成梗很有帮助,因为你能看到真正让你惊讶的东西,而且因此很好笑。所以,显然这不是这项技术最富有成效的用途,但它确实展现了这种新兴能力:找到你原本不会知道的信息。

Yeah, I actually have a funny anecdote about this. Recently, while gearing up for this launch, the team has been really cooking on it for a couple months. Over that time, there's so much conversation and chatter going on in Slack, Docs, and elsewhere. One of the members of the team set up this scheduled automation to look at everything that's going on and come up with the best memes and post them in one of our shared channels. There are two cool things about this. First, I think the models are, over time, actually starting to become funny, whereas a year ago that was not at all the case. The second is, as you were saying, they find things in surprising ways that you may not have thought of, and create connections that you may not have thought of. That really helps with meme generation because you can see something that genuinely surprises you and is funny in that way. So yeah, obviously that's not the most productive use of this technology, but it does capture this emerging capability: finding information that you otherwise would not know.

成功发布与千万用户 Launch Success and 10 Million Users

Host

说到这次发布,我觉得我已经说过这是很长时间以来最成功的一次发布。我个人认为甚至比 5.0 还要成功。你现在宣布有 1000 万用户。感觉有什么不同吗?你经历过很多次发布。

Talking about the launch, I think I've pretty much said this is the most successful launch in a long time. I think even more successful personally than 5.0. And you're announcing 10 million users. Does it feel different? You've been through a lot of launches.

Akshay

我觉得这感觉像是一个顶点。嗯,我想说两点。第一,这感觉像是一个顶点,就像我前面提到的:这是我们长期以来的愿景。我们在内部看到了 Codex 的神奇,并非常兴奋地把它带给更多人。看到它发挥作用,看到我们达到你提到的那个数字所代表的分发目标——我认为这非常巨大,也超级令人兴奋。另一面是还有很多事情要做,这也非常令人兴奋。ChatGPT 作为一个整体,几乎人人都把它等同于 AI,并且喜欢它,拥有数亿用户。所以 1000 万确实很酷,但我们需要把它带给所有人。我们需要让每个人都感受到这种魔力。所以这是接下来的下一步。不过,是的,我对目前的进展和未来的机会感到非常兴奋。

I think it feels like a culmination. Well, I think two things. One, it feels like a culmination, like I was mentioning earlier: this vision that we've been on for a long time. We saw the magic of Codex internally, and we were extremely excited to bring it to many more people. To see it working, to see us reach the distribution goal in the numbers you mentioned—I think that's huge and super exciting. The flip side is there's so much more to do, and that's also really exciting. ChatGPT as a whole is a product that everyone almost equates to AI, and loves, and has hundreds of millions of users. So 10 million is really cool, but we need to get this to everyone. We need everyone to feel this magic. So that's the next step from here. But yeah, I'm extremely pumped about how it's going so far and the opportunities ahead.

ChatGPT Work对比Codex ChatGPT Work vs. Codex

Host

太棒了。我还想问一下——因为我一直在密切关注这个数字——它从某个时候开始从只是 Codex 用户转变成了 Codex 加 ChatGPT Work。显然,因为它们是同一个框架,重点是你无法把它们分开计数。你们大概有 10 亿 ChatGPT 用户吗?什么时候一下子跳到 10 亿的?这不是 ChatGPT 的默认设置吗?还是不是?

Awesome. I did want to also—since I've been tracking the number closely—it transitioned at some point from just Codex users to Codex plus ChatGPT Work. Obviously, because it's the same harness, the whole point is that you can't count them separately. Do you have roughly a billion ChatGPT users? When did it just jump to 1 billion right away? Isn't that the default on ChatGPT or no?

Akshay

如果你用的是免费版 ChatGPT,我们不会默认把你带到 ChatGPT Work。目前它也仅供付费用户使用。我们需要一个过程来教育用户了解这个产品的价值,让他们尝试,从他们的反馈中学习,并随着时间的推移把它做得更好。目标是让今天喜欢 ChatGPT 的尽可能多的人感受到 ChatGPT Work 的力量,但我认为这会是一个旅程。

We don't default you into ChatGPT Work if you're on ChatGPT free. It's also only available to paid users right now. There's a process of educating users on what the value of this product is, having them try it, learning from their feedback, and making it better over time. The goal is to get as many people who love ChatGPT today to feel the power of ChatGPT Work, but I think it'll be a journey.

Host

是的。而且 Codex 在可预见的将来仍会作为一个品牌存在。

Yeah. And Codex will still be alive as a brand for the foreseeable future.

Akshay

嗯。

Yeah.

Host

我们只需要在界面上根据需要在它们之间切换。

We'll just toggle between them as needed for UI stuff.

Akshay

是的,我认为这比那个观点更有力。我们完全打算像对待开发者一样——开发者长期以来一直是我们的核心市场——而且我们可以做更多的事情来让 Codex 特别适合软件开发,我们会继续这样做。这一点也不会削弱这一点。如果要说的话,它应该会增加 Codex 这类工具的实用性,因为现在你可以在写草稿、创建工件或搜索代码之间无缝切换。

Yeah, I think it's an even stronger point than that. We fully intend to treat developers—developers have been a core market for us for so long—and there's so much more we can do to make Codex great specifically for software development, and we'll continue to do that. This doesn't take away from that at all. If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a draft, creating an artifact, or doing a search over your code.

术语与OpenClaw Terminology and OpenClaw

Host

我确实想知道这些术语会多大程度地渗透到非技术用户。比如,如果我想用 artifacts,他们必须学会说 artifacts 吗?还是说,你懂的……

I do wonder how much this terminology leaks to the non-technical user. Like, do they have to learn to say artifacts if I want artifacts, or, you know,

Akshay

这很有趣,我们在内部叫 artifacts,因为团队都这么叫,但外部没人这么说。没人会叫它 artifact。不过我觉得人们会用他们习惯的说法来描述东西。所以如果 ChatGPT Work 擅长做幻灯片,他们就会说 ChatGPT Work 擅长做幻灯片,而这正是我们想要的。

It's funny, like we call artifacts internally because that's what the teams call them, but externally no one says that. No one calls it an artifact. But I think people describe things with whatever they're used to. So if ChatGPT Work is good at creating slides, they'll say ChatGPT Work is good at creating slides, and that's actually what we want.

Host

还有一件大事——我的意思是,现在是 2026 年 7 月。OpenAI 还发生了一件大事,就是 OpenClaw,我认为这是很多人第一次真正把智能体用于个人事务,但同时也以同样的方式跨越到工作中。

One big other thing—I mean, it's July of 2026. One big thing that also happened for OpenAI was OpenClaw, and I think that was a lot of people's first time really using agents for personal stuff, but also crossing over to work in that same way.

OpenClaw灵感 OpenClaw inspiration

Host

据我所知,OpenClaw 目前仍是独立项目。但你自己有没有经历过 OpenClaw 时刻?从 OpenClaw 到 Codex,你有没有吸取到什么经验?反过来呢?

As far as I understand, OpenClaw is still independent. But did you go through your own OpenClaw moments? Were there any lessons you took from OpenClaw to Codex or back, whatever?

Akshay

我觉得这里面有太多灵感了。我自己确实也经历过一个 OpenClaw 时刻。

I think there's a lot of inspiration. I did go through my own OpenClaw moment.

Host

对,讲讲这个故事。

Yeah, tell the story.

Akshay

我和我妻子搭了一个 OpenClaw,想用它管理家里所有的事情。倒不是家里有多少事,但它其实挺有用的。我们给了它一个日历,它就开始帮我们创建日程之类的。后来跑 OpenClaw 的那台笔记本挂了,我们再没机会把它捡起来。但那里面的灵感很多。比如在 ChatGPT Works 的 Web 和移动端,你可以访问一个持久化的电脑环境,能存文件,而且这些文件在会话之间会保留。我们的想法就是让这类用例变得可行。我们团队有个人现在就是用 ChatGPT Works 来做以前用 OpenClaw 做的事,我觉得已经完全迁移过来了——就是健身计划、饮食追踪这些。这算是个工作型的东西,对吧?不一定是正经工作,但属于个人效率这个范畴。而且它有同样的一套原语:有定时任务,能把文件存在文件系统里,能随时间推移引用这些内容。所以你会开始看到同样类型的用例出现,真的很酷。

My wife and I set up an OpenClaw to try to manage everything in our house. Not that there's a ton, but it was actually quite useful. We gave it a calendar, and it started creating events for us and stuff. At some point the laptop it was running on died, and we never got the chance to pick it back up. But there's a lot of inspiration there. Like, in ChatGPT Works on web and mobile, you get access to this persistent computer environment where you can store files, and those files stay around between sessions. The idea is to be able to enable use cases like this. One of the members of our team actually uses ChatGPT Works for what they used OpenClaw for before, and I feel like it has completely transitioned, which is workout planning and meal tracking. That's a worky thing, right? Not work necessarily, but it's in the personal productivity space. But it has all the same primitives. It has scheduled tasks. It has the ability to store files on a file system. It has the ability to reference those things over time. And so you start to see the same types of use cases emerge, which has been really cool.

ChatGPT Works对比OpenClaw ChatGPT Works vs OpenClaw

Host

有没有可能 ChatGPT Works 完全取代 OpenClaw?当然它们是独立的,所以……

Is there a point that ChatGPT Works completely replaces OpenClaw? Obviously, they're independent, so...

Akshay

是,我离 OpenClaw 有点远,所以我不太能谈它的路线图。但我不认为会完全取代。我觉得 OpenClaw 团队打造的开源技术非常了不起,这个需求会一直在。我们可以在产品中汲取灵感。而且你知道,听说过和使用过 ChatGPT 的人比用过 OpenClaw 的人多得多。如果我们能把 OpenClaw 的魔力带给这些人,那就是成功。在 ChatGPT Works 这边,我们非常看重的一个点是:核心体验是你来到这个产品,和这个智能体对话,开启一个会话——随便你怎么叫它。这个产品的魔力在于,在那一刻你可以做任何事。我们想打造这样一个产品:你不需要点击按钮或跳转到别的地方,就能在这个统一的地方获得你的财务应用或任何其他产品里的所有功能。这就是目标。我们希望做一个可扩展的、带插件的系统,让你能连接完成财务任务所需的工具。如果你在做科研,我们有办法扩展系统,让你写技术方案,而且它表现很好。总会有我们支持的产品在这些方面做到最好,但我们希望把尽可能多的魔力放在这个核心体验里。

Yeah, I mean, I'm not close to it, so I can't speak to the OpenClaw roadmap. But I don't think so. I think there's always a need for the incredible open source technology that that team has built. I think we can draw inspiration in the product. And you know, I think many more people have heard about and used ChatGPT than have used OpenClaw. If we can take the magic from OpenClaw and bring it to them, I think that'll be a success. I think that one thing on the ChatGPT Works side that we feel strongly about is that the core experience is that you come to this product and you have a conversation, start a session, whatever you want to call it, with this agent. And the magic of the product is that you can do anything in that moment. We would like to create a product where you don't have to click a button or go to a different place, whatever, and you can get whatever functionality exists in your finances app or any other product in this one place. So that's the goal. We want an extensible system with plugins where you can connect to the tools that you need in order to accomplish a financial task. If you're doing science work, we have an ability to extend the system so you can write the tech and it performs well. There'll always be products that we support that are best in class at those things, but we want as much of the magic as possible in that core experience, you know.

财务插件与数据架构 Finance plugin and data architecture

Host

你觉得现在用 ChatGPT 做财务,能把你以前用 Wolfram 能做的所有事都做掉吗?

Do you think that you can do everything you used to do with Wolfram's in ChatGPT finance?

Akshay

我其实试过。ChatGPT 现在还不能替你托管现金和资产,所以那部分还不行。但养老规划、财务规划、预算这类我们当时在关注的内容,用今天的 ChatGPT 配合财务插件全都能做到。所以我感觉至少那个部分对我已经被替代了。

I actually tried it. I mean, ChatGPT doesn't yet custody cash and assets for you, so that part no, not yet. But there's a whole component of retirement planning and financial planning and budgeting and stuff that we were looking into when I was there. And with the finance plugin, that's all possible with ChatGPT today. So I feel like at least that component is replaced for me.

Host

我其实还没接进去。我有点不敢看那个答案。

I haven't really plugged it in yet. I'm somewhat scared to look at the answer.

Akshay

说实话,我对健康和财务也是这样的心态。我就想着,不知道,不知道。不过它真的很好。我觉得很酷的一点在于,传统 UX 里,你想给用户越多的能力,就需要加越多的控件和花哨功能。比如财务和预算应用,总有一堆筛选器、搜索栏之类的东西。但现在只要有正确的数据连接,你就能拥有你想要的一切。你可以在这个框里问任何问题,然后得到答案。我觉得这非常强大。

Like, that's honestly the same reason for health and finances. I'm like, I don't know. I don't know. It's really good. I mean, it's really cool how in conventional UX, the more power you want to give to a user, the more knobs and bells and whistles you need to add. For finance and budgeting apps, there are always a bunch of filters and search bars and stuff like that. But now, with the right connectivity to the right data, you can have whatever you want. You can ask any question you want in that box and get the answer. I think that's super powerful.

Host

我觉得把它集中在一个地方也很好,对吧?你有各种健康应用,我有智能秤、手表这些不同的东西。能集中放在一起真的很好。这其实是 OpenClaw 理念的一部分,对吧?就是你会拥有一个个人 OS,大概 ChatGPT 也想成为那样。我确实觉得,仅仅依赖通过 MCP、CLI、API 之类的方式实时拉取数据还不够。我有点数据工程背景,你仍然需要一个数据仓库,或者某种缓存或语义层。你有这种感觉吗?还是你们已经有这个东西了?

I think it's also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, all these different things. It's just nice to centrally collocate it. Which is, you know, part of the whole thing of OpenClaw, right? That you would have a personal OS, which presumably ChatGPT wants to become. I do think that just relying on just-in-time pulling of data, let's say through MCP, CLI, API, whatever, is still not enough. I come from a bit of a data engineering background. You still want a data warehouse or some kind of caching or semantic layer. Do you feel that, or do you already have that?

Akshay

我没法把所有细节都讲清楚,但我认为这取决于访问模式,对吧?如果你要立刻给出答案,那就很难从所有这些来源拉数据。但我们在 ChatGPT Works 上想支持的很多用例,并不需要你立刻得到结果,更像是你想让智能体去执行的一个任务,那本身需要一定时间。而且现在有了程序化工具调用、子智能体这些,其中一部分时间还可以并行化。所以我认为——很有可能——通过 MCP 调用第三方服务能做到的上限已经被大幅拉高了。我们对此非常兴奋。

I can't speak to all the details on how everything works, but I think it depends on the access pattern, right? If you want to answer immediately, then yes, it's very difficult to do that if you need to pull from all of these sources. But a lot of the use cases that we want to enable on ChatGPT Works aren't necessarily something that you need immediately. It's more like a task that you want the agent to go and do. And that's going to take a certain amount of time. And with things like programmatic tool calling and sub agents now, some of that time is also parallelizable. And so it's possible — I think it's very possible — that the ceiling on what can be done with MCPs and calling out to these third-party services has been raised substantially. So we're really excited about that.

子代理与扩展性 Subagents and extensibility

Host

你提到了子智能体。我得展开讲讲。Ultra 是一个新模式。ChatGPT 本身有一些特别的交互设计来展示这些智能体。说实话,我对它们做不了太多,只能看着。你在这方面的体验如何?对正在用子智能体构建应用的开发者,你有没有想指出的设计问题?

You mentioned sub agents. I got to double click on that. Ultra is a new mode. You have special affordances in ChatGPT itself to show off the agents. Can't really do much with them, to be honest. I just watch. Um, what have been your experiences? Any design issues that you would call out to other builders building with sub agents?

Akshay

我觉得这又回到了我之前提到的那个平衡问题。

I think it sort of goes back to the balance that I was raising earlier.

子代理 Subagents

Akshay

我想关键在于让构建者看到工具的能力,同时提供足够的抽象层,不让他们感到不知所措。关于子智能体,我们想展示的是:你可以把一个有许多并行分支、或复杂到子智能体能处理的任务交出来,这个产品就是为你准备的。模型能够完成这些目标,或者说尝试去完成。所以这就是我们在产品中展示它们的原因,也是我们目前设计上的思路。其实还有一种迭代方式,你可以看到它们正在做什么之类的,我觉得那样可能会因为信息量过大而让人不知所措。所以这是我们目前所做的一个双刃剑式的取舍。

Felt like showing builders the power of the tool, but also creating enough abstraction to not overwhelm them. With subagents, the thing we wanted to show is that you can take a task that has many parallel tracks or is complicated in a way that subagents can handle, and this product is for you. The model can accomplish those goals, or try to accomplish them. So that's the point of showing them in the product, and that's where we've gone with the design. There's another iteration of this where you can see exactly what they're doing and things like that, which I think could verge on being overwhelming with information. So this is the double-edged trade-off we made for now.

Host

你们确实显示了大量的转录文本。

You do display quite a lot of transcripts.

Akshay

对,对。

Right, right.

Host

我觉得你是想展示更多内容吗?

I think you want to display more than that?

Akshay

不,不,这样就足够了。

No, no, it's fine.

Host

有些人可能想要看到更多。我就是那种会把很多事情丢给目标、几乎每个目标都会告诉它使用子智能体的人。听起来有点多余,对吧?但每次我都会说:“好,尽可能使用子智能体。”我有很多朋友也推荐并这么做。而有时候我会和另一些人聊,他们说:“好,这个子任务我希望你在这里用子智能体。”我敢肯定他们会希望看到子智能体是怎么被使用的。对我来说主要有两点:第一是净时间效率,所以分散到各个子智能体上;第二可能是成本——不用又大又贵的模型,而是分摊到许多更小、更便宜的模型上。有些人就想要这种程度的控制。所以如果你做的事情有重复性——比如我想建一个能每天早上持续执行的东西——我可能会想去微调这里的子智能体、那里的子智能体,这样两边都能看到。但我认为,如果我没记错的话,默认是隐藏的。有一个下拉菜单能展开很多内容,我当时想:“好吧,我就让它一直开着。”

Some people could want more. I'm one of those people who will basically throw a lot of stuff at the goal, and on pretty much every goal I'll tell it to use subagents. Seems redundant, right? But every time I say, “Okay, use subagents where possible.” And I have a lot of friends who recommend the same and do the same. Whereas sometimes I'll talk to people who say, “Okay, this is where I want you to use subagents for this subtask.” And I'm sure they would appreciate seeing into how they're being used. For me it's primarily two things. One is net time efficiency, so span out across subagents. Two is probably cost. Don't use the big expensive model; offload to a lot of smaller, cheaper models. And some people want that level of control. So, if you have repetition in what you're doing — say I want something built where I want it to consistently do this every day — I might want to go in and fine-tune subagents here, subagents there. So you can see both. But I think, if I'm not mistaken, it's hidden by default. There's a dropdown that goes pretty deep, and I'm like, “Okay, I'm just going to keep it on.”

Host

你可以改变它们使用的模型吗?

You can change the model that they use?

Akshay

我知道我会让它们被引导。我的意思是,我知道 Anthropic 在 Claude Code 里提供了这个功能。你可以告诉 Fable 用 Sonnet 或 Opus,用 Sonnet 作为子智能体。所以这是很简单的事,你告诉它用 Sonnet 去展开子智能体,它更便宜也更快。我想如果这个功能现在还没有,未来也可以加进去,但我觉得还有一方面……

I know I tell them to be steered. I'll say, you know, Anthropic offers this in Claude Code. You can tell Fable to use Sonnet or Opus to use Sonnet as a subagent. So it's a pretty trivial thing — you tell it to span out subagents with Sonnet, which is cheaper and faster. I would assume if it's not there, it could be built there, but I think there's a side of...

Host

太多开关了。

Too many toggles.

Akshay

嗯,其实这不是一个开关,就是在聊天里告诉它。我的做法是用提示词。我觉得这个东西会被抽象掉,除非你是为了重复性而构建的。所以如果我在构建某个东西——比如说播客准备——研究人物、做非常深入广泛的调研——我可能想把它配置成更便宜更快的模型,只用来做网络搜索。我可以想象一个两者都想要的世界。我觉得目前的默认状态其实挺好的——默认隐藏,但你可以下拉查看更多完成细节。我知道在 5.6 发布时很多人聊过这个。这个功能非常喜欢用大量子智能体,导致 ChatGPT 应用因为处理器负担太重而崩溃,不过嗯……

Mhm. It's not a toggle, actually. It's just — you tell it in chat. The way I do it is prompt it. And I think this is something that gets abstracted unless it's something you built for repetition. So, if I'm building something — say that's a podcast prep — researching people, doing very deep, extensive research — I might want to configure it to a cheaper, faster model just for web search. I can see a world in which you want both. I think the default is actually pretty good right now — where it's hidden, but you can drop down and get some more info into what's done. I know people talked a lot about it on the 5.6 launch. This thing loves to use a lot of subagents and causes the ChatGPT app to crash because it's so processor heavy, but um...

Host

那么,你的个人体验是怎样的?

So, what is your personal experience?

Akshay

是啊,我倒是没遇到过它因为子智能体崩溃。

Yeah, I haven't had it crash from subagents.

Host

我也没有。咱们俩都用的是大笔记本,但我知道有人提过这事。这是一个讨论话题,只是我们没遇到同样情况。但这又是一次“氛围评估”,对吧?人们会说:“天哪,这么多子智能体同时跑,太疯狂了。”而我会说:“我觉得这没问题,很好。”但这就是人们会拿出来说的事。

I haven't either. We both have big laptops, but I know people brought it up. It was a topic of discussion, though we didn't see the same. But it's another vibe eval, right? People are like, “Okay, the amount of subagents all running is crazy.” And I'm like, “I think this is okay. I think it's good.” But that's just stuff people bring up.

Akshay

我觉得当我们发布产品时,我们对于谁会用 Ultra 4、他们什么时候该用它,并没有那么明确的立场。后来我们做了一些改动,比如要求你主动开启,并且在高级设置里才能找到,因为它就是为那类人准备的。它是给那些理解接下来会发生什么的进阶用户用的,因为根据你的使用场景,它也可能消耗更多你的额度限制。

I think when we launched the product, we weren't as opinionated about who would use Ultra 4 and when they should use it. Since then, we made some changes, like require you to turn it on and find it in the advanced settings, because that's who it's for. It's for power users who understand what's going to happen, because it can also, depending on your use case, use more of your limits.

Host

是的。

Yes.

Akshay

所以我觉得我们的很多反馈就是来自这方面。

So that's where I think a lot of the feedback was coming for us.

Host

好吧,重置额度。

Okay, reset the limits.

Akshay

总是重置额度。

Always reset the limits.

记忆 Memory

Host

呃,你知道,今天我们要重置一下,因为我想换个话题,讲讲这个框架的最后一块:记忆。最近很多人在评论记忆功能。ChatGPT 的新记忆系统以前很烂,不怎么好。然后还有这个人——也是基本上同样的说法。还有 Samir——你应该和他一起工作——也在谈记忆。对此你能说些什么?

Well, you know, today we're resetting because I want to change topics to one last piece of the harness: memory. A lot of people are commenting on memory recently. ChatGPT's new memory system used to suck — not very good. And then this guy — also basically the same thing. And Samir, who you presumably work with, talking about memory. What can you say there?

Akshay

我觉得 Samir 和团队——以及研究团队——随着时间的推移做了大量更新和改进。当我和朋友、家人聊到他们喜欢 ChatGPT 的什么地方时,比如它了解他们、让他们觉得自己的 ChatGPT 就是自己的 ChatGPT,我觉得这个大概会排在第一位……

I think Samir and the team — and the research teams — have made a ton of updates and improvements over time. I think when I talk to friends and family members about what they love about ChatGPT, like the fact that it knows them, that they feel like their ChatGPT is their ChatGPT, I think comes up probably...

Host

嗯。

Yeah.

Akshay

第一。ChatGPT Work 在云端——默认所有对话都会继承 ChatGPT 的记忆,所以你会知道它们会了解你的背景信息。它们也能够把内容写回到这个记忆里。

Number one. And ChatGPT Work in the cloud — by default all conversations inherit from ChatGPT memory, so you'll know they'll know context about you. They'll also be able to write back to this memory.

Host

就像是有一段小文本,对吗?就像你会告诉我你正在写入,对吗?它是……

With a small text, right? Like you tell me when you're writing, right? Is it...

Akshay

不,它是我们发布的同一个记忆 V3 系统的一部分。

No, it's part of the same memory V3 system that we launched.

Host

你说 V3 是什么意思?

What do you mean V3?

Akshay

嗯。

Yeah.

Akshay

所以我觉得这非常强大,因为从 ChatGPT 到 ChatGPT Work,感觉就像是我已经使用这个产品多年的一个延伸。所以这太棒了。而且很高兴看到人们在这里认可了这些改进。

So I think that's been really powerful, because going from ChatGPT to ChatGPT Work feels like an extension of what I've already been doing with the product for sometimes many years. So that's been awesome. And it's awesome to see that people are recognizing the improvements here.

Host

所以这基本上是一个检索问题,对吧?比如,你检索到了正确的东西吗?你是不是过度关注了错误的东西?是假阳性更多还是假阴性更多,如果这说得通的话?哪个问题更大?

So it's basically a retrieval problem, right? Like, are you retrieving the right things? Are you over-focusing on the wrong things? Is there more false positive or false negative, if that makes sense? What's the bigger problem?

Akshay

我并不是直接研究记忆的,所以很难确定地说哪个问题更大,但我认为你说得对。这有两方面:一方面确保它了解关于你的事情;另一方面还要有情商,在恰当的时刻主动提起这些事,或者给你带来正面而非负面的惊喜。所以这是一个非常有挑战性的问题,但也是我们认为非常值得去解决的巨大机遇,这也是为什么我们在这方面投入很大。

I don't work on memory directly, so it's hard to say with certainty what the bigger problem is, but I think you're right. There are two sides to it: making sure it knows things about you, but then also having the EQ to bring those things up at the right moments — proactively or surprising you in ways that are positive, not negative. So it's a very challenging problem, but something we feel is a huge opportunity to get right, which is why we've made big investments in it.

跨聊天与工作的记忆 Memory across chat and work

Host

你怎么看在构建 ChatGPT for Work 时,与普通聊天应用、与 Codex 不同的那一面,跨项目的记忆管理、协作等等?怎么看那些与框架无关的部分?比如,如果我在一个项目里有四组会话,在构建记忆系统方面有什么经验吗?

How do you see the side of building ChatGPT for work differently from the regular chat app, different from Codex, managing memory across different projects, collaboration and whatnot? How do you see the side that's separate from the harness? So, if I have four threads on one project, any learnings on how to build memory systems there?

Host

补充一些背景,我想稍微引导一下:做聊天类应用时,你有很多一次性的对话,对吧?切换到工作场景后,可能你要做一个月,或者经常做。那么当我增加更多会话时,就远不止单线程了,而且那里可能就有记忆。

For background as well, I guess, to steer it a bit: when you do chat-style applications, you have a lot of one-offs, right? When you switch to work, it might be something you're doing for a month, something you do a lot. Now, as I add more sessions, there's a lot more than just single-threaded, and there might be memory there.

Akshay

首先,我会反驳一点:记忆的深度或价值,在聊天和工作之间并非本质不同。确实,聊天里有很多较短的会话,但我认为 ChatGPT 这个产品已经存在很久了,而且人们今天已经把它用于与工作、生产力相关的事情。所以我们发现有大量的价值。我在个人使用中也感受到了:这些一次性的对话随着时间累积成相当持久的东西,很好地代表了我这个人。我知道时不时会在 X 上流传一些帖子,说 ChatGPT 会告诉你它知道你的一切,而人们总是惊讶于它有多深。

I think first I'd challenge that the depth of the memory or the value of it is fundamentally different across chat and work. It is true that there are a lot of shorter sessions on chat, but I think the ChatGPT product has had a ton of longevity, as long as this technology has been around, and people already use it for worky, productivity-related things today. So I think we've found that there's a lot of value. I even found this in my personal usage: all of these one-offs add up over time into something quite durable, a quite good representation of who I am. I know from time to time something will go viral on X about ChatGPT telling you everything it knows about you, and people are always surprised by how deep that is.

Host

有趣的就是让它吐槽我,你懂吧。

The fun, roast me, you know.

Akshay

完全正确。所以这些都想说明,我认为现有 ChatGPT 产品里已经有很深的积累,这也是我们认为把它带入工作产品很有价值的原因。但另一个我提这个的原因,是希望我们可以用一些相同的基础元素和系统,在这里也扩展记忆。而且我知道专注这个的团队眼下正在推进。

Exactly. So that's all to say I think there's a lot of depth in the existing ChatGPT product, and that's why we think it's valuable to bring into the work product. But the other reason I brought that up is because hopefully we can use some of the same fundamental primitives and systems to extend memory here as well. And I know this is something the team that focuses on this is working through right now.

Chronicle深度解析 Chronicle deep dive

Host

我想提一个记忆功能,说实话我自己用得不多,我很好奇你用不用—— Chronicle,屏幕上现在正显示。它有点像超级记忆,还是它到底是什么?

I wanted to bring up one element of memory, which I honestly don't really use much, and I'm curious if you do: Chronicle, which is up on screen right now. It's kind of a super memory — or what is it?

Akshay

它的思路是,它可以学习你如何使用电脑,是记忆的另一个输入源。目前它还是实验性的,不是默认开启,但我建议你试试。我觉得有趣的是,它跟我们之前聊的呼应上了——你当时问:ChatGPT 会漏掉东西吗?比如在 Slack 里搜索时,它会漏掉信息吗?因为数据量太大了。你也可以对电脑上你做的一切提出同样的问题:它会知道你做的所有事吗?它能捕捉到意图之类的东西吗?可能不会。但它很可能会发现你自己都不知道的事情。然后如果它能在恰当的时候、以主动的方式把这些东西呈现给你,比如你在做任务的时候,至少我觉得它会相当有用。所以值得一试。

The idea is that it can learn from how you're using your computer, and it's another input source into memory. It's experimental right now, and it isn't default-on, but I'd recommend you try it. I think it's quite interesting how it goes back to a conversation we were having earlier. You were asking: can ChatGPT miss things? On Slack, when it's searching, does it miss things? Because there's such a volume of stuff. You can ask the same question about everything you're doing on your computer. Is it going to know everything you're doing? Is it going to capture the intent and stuff like that? Probably not. But it probably will find things you might not know about. And then, if it can surface those to you at relevant times in proactive ways, like when you're doing tasks, then I've found, at least, that it can be quite helpful. So it's worth trying.

Host

所以主要用于洞察和长期上下文?

So, mostly for insights and longer-term context?

Akshay

对,正是如此。它提供洞察,并建立上下文,让你在某些任务上更高效。但没亲身体会的话,真的很难描述。

Yeah, exactly. Insights, and it builds context that can make you more productive on certain tasks. But it's hard to describe without feeling it.

Host

我觉得你还是能明显感受到的。就像他们这里说的那个意思,对吧?直接翻看我的记忆,或者翻看我的日志,然后添加技能。

I will say you can feel it pretty well. Like the idea of what they're saying here, right? Just check through my memories, or check through my logs, and add skills.

Akshay

嗯。

Yeah.

Host

相当被低估了,对吧?

Pretty underrated, right?

Akshay

但那是自动化操作。你可以用计划任务(cron job)来重复。

But that's automations. You can repeat that using a cron job.

Host

查看你的记忆并创建技能。

Checking through your memories and creating skills.

Akshay

对。

Yeah.

Host

但我认为,Chronicle 本身对记忆的创建才是不同之处。因为开着 Chronicle,你的记忆会更深。

But I think the creation of the memories from Chronicle itself is what's different. You have much deeper memories because you have Chronicle on.

Akshay

它就在那里。我用得不多,但也许我只需要更多例子。我猜你们内部肯定大量使用,所以我总是想找些使用场景。

It's there. I don't use it much, but maybe I just need more examples. I imagine you guys use a lot of it internally, so I'm always fishing for use cases.

Host

是啊,我会直接打开它,然后它就会自动运行。

Yeah, I would just try turning it on and then it just auto works.

Akshay

对,然后看它会在哪些地方开始帮到你。我觉得你会惊讶的。

Yeah, and seeing where it might start helping you. I think you'll be surprised.

Host

太好了。

Yeah, amazing.

构建前AI与后AI Building pre-AI and post-AI

Host

我想关于 ChatGPT for Work 的整体覆盖就聊到这里吧。在构建以及所有这些方面,已经有了很多不错的进展和讨论。社区里以及 OpenAI 内部都有很多前创始人。你觉得事情变化大吗?我想听听你对打造 AI 之前和之后时代的整体反思。

I think that was about it in terms of the overall coverage of ChatGPT work. There's been a lot of good progress and discussion on building, and all these things. There are a lot of ex-founders in the community and at OpenAI as well. Do you think things have changed a lot? I guess, your overall reflection on building pre-AI and post-AI.

Akshay

我觉得事情变化非常大。看到今天从创意到真正的东西能多快完成,非常令人兴奋。即使是更早以前,比如五到十年前,如果你很拼、愿意做最小可行产品,速度也很快。但现在你能构建的东西范围要广得多。而且我们内部构建时所看到的是,这让你可以更快验证,与用户交流,与内部医生等交流,确保走在正确的轨道上。这个循环变得比以往任何时候都更紧密,这对产品开发来说是一大胜利。对消费者和用户也是胜利,因为理想情况下,这意味着他们一开始就能得到更好的产品。

I think things have changed a ton. It's super exciting to see how quickly you can go from idea to something real today. Even before, like 5 or 10 years ago, it was fast if you were scrappy and willing to build the minimal viable thing. But now the extent of what you can build is much, much broader. And what we've seen internally building is that it gives you an opportunity to validate much more quickly, to talk to users, to talk to internal doctors, etc., and make sure you're on the right track. That loop, I think, has become tighter than ever before, and that's a win for product development. It's a win for consumers and users too, because ideally that means they're getting much better products out of the gate.

Host

那是不是意味着你的团队更小了?

Does that mean your team is smaller?

Akshay

我觉得现在要做的事更多了。个人或小团队能完成的事比以前更多,以前可能需要更多人。但同时,也有更多的事要做。所以团队的野心也大得多。

I think there's much more to do now. People can accomplish more individually or in a small team than before, where it would have required more people. But at the same time, there's also more to do. So the teams are much more ambitious.

团队角色与团队建设 Team roles and team building

Host

你有没有看到角色范围、团队构建方式的变化?比如几年前我们常见的团队,与现在理想团队相比,有什么不同?

Have you seen any changes in the scope of roles and building teams, how we used to have teams a few years ago versus what ideal teams look like now?

Akshay

我认为我们看到典型产品开发职能之间的界限变得模糊——比如 EM(工程经理)、PM(产品经理)、工程师、设计师等等。

I think we've seen a blurring of the lines between the typical product development functions — between EM, PM, engineer, designer, etc.

Host

对,每当我引用这句话:科技行业最后只会剩下四种工作。一种是 AI 垃圾炮——就是那些只烧一堆 token 的人。然后是 SRE,也就是更负责任的人。还有负责卖东西的成年人,然后就是长得好看的人。

Yeah, when I bring up this quote, there will be only four jobs left in tech. There's AI slop cannon — the people who just burn a bunch of tokens. Then there's SRE — people who are more responsible. There are grown-ups who sell things, and then there are hot people.

T型技能与瓶颈 T-Shaped Skills and Bottlenecks

Akshay

这个观点很有意思。我觉得每个人都会在某种程度上成为 T 型人才。AI 并不会让每个人都变成通才。比如以前我根本不可能设计出东西,即使现在,我也许也没有所需的视觉品味,但我可以在 AI 的帮助下迭代出某个方案。但每个人还是会有一个专长,就像 T 的那一竖。你有自己感兴趣的专长,在 AI 的帮助下可以不断深入、越做越好,同时你也是通才。有了这个基础,你能做到的事情几乎是无限的。

It's an interesting take. I think everyone will be T-shaped in a way. It's not like AI will enable everyone to become a generalist. You know, things like I never would have been able to come up with a design before. Even now, I might not have the visual taste required, but I can iterate on something with the help of AI. But people will have a specialty — like the straight line in the T, or the upward line in the T. So you have a specialty you're interested in. With the help of AI, you can go deeper and become better over time, but then you'll also be a generalist. With that foundation, what you can accomplish is almost limitless.

Host

那你的专长方面,瓶颈是什么?你需要更多设计师吗?需要更多“垃圾内容大炮”吗?需要更多“颜值高的人”吗?

What are you bottlenecked by in terms of specialties? Do you need more designers? Do you need more slop cannons? Do you need more hot people?

Akshay

我觉得瓶颈会变成想法和品味。因为现在谁都能动手构建,我认为这真的是一个自下而上、凭野心驱动的时代。正因为有太多东西可以去做,你永远都会受制于想法数量,以及你在任何时间点正在做的事的数量。

I think the bottleneck becomes something like ideas and taste, I guess. Because anyone can build now, I think it really is the era of bottoms-up ambition. And because there's so much to be built, you're always going to be bottlenecked by the amount of ideas and the amount of things you're doing at any given time.

Host

你觉得模型能帮助解决这个问题吗?

Do you think models help solve that?

Akshay

模型?

Models?

Host

嗯。

Yeah.

Akshay

我的意思是,我有个例子,比如我在前端设计方面有点技能。它们会给我四个截然不同的示例,看效果如何。这确实很烧 token,但之后我基本上会把它们浓缩起来:“好,我喜欢这部分,我喜欢那部分,我们把它们合在一起。” 看起来,是,我心里是有个愿景,但怎么说呢。我想说,我特别希望实现、但还没实现的一个自动化就是“给我带来新想法”,对吧?不知怎么的,LLM 就是做不到。关于想法,有一点很有意思:它们不是凭空冒出来的,通常来自某个地方。在产品开发中,它们来自和用户交流,或者来自你感受到的摩擦,或者是在你已经规划好的某个基础上叠加反馈。所以我觉得,这正是我们说的这些通才永远有价值的地方——闭环反馈,从这些反馈或用户交流中产生有根基的想法,诸如此类。

I mean, I have the example of having a front-end design skill. They give me four drastically different examples of what this looks like. It sure burns a lot of tokens, but then I'll mostly just condense down: 'Okay, I like this part, I like this part. Let's draw these together.' And it's like, yeah, I had a vision, but I don't know. I would say the one automation that I would love to work, and it doesn't work, is, 'Bring me new ideas,' right? Somehow LLMs are just not it. One interesting thing about ideas is that they're not in a vacuum. They usually come from somewhere. In product development, they come from talking to users, or reacting to friction you're seeing, or feedback building on some foundation you already planned out. So I think that's where there will always be value in these generalists we talked about — closing that loop, coming up with those ideas that are grounded in that feedback, or talking to users, or whatever it is.

定义生产力 Defining Productivity

Host

好。你负责生产力团队。你怎么定义生产力?

Cool. You lead the productivity team. How do you define productivity?

Akshay

我认为我们的使命是让人们能够做到以前做不到的事。现在我们主要从知识工作的角度来思考。看知识工作时,我觉得人们不再被自己的角色、背景或训练所束缚。不管你在哪个职能岗位,你都能突然开始构建东西,接触到原本可能无法解读的数据。然后我觉得这也延伸到个人生活——最终我们希望给你杠杆。我们希望产品里的模型能给你杠杆,让你能为自己创造时间,去做你热爱的事。

I think our mission is to make it possible for people to do things that they weren't able to do before. Right now we're thinking about it from the perspective of knowledge work. When I look at knowledge work, I see people are no longer siloed by their roles, no longer siloed by their background or training. No matter what function you're in, you can suddenly build things and get access to data that you otherwise might not be able to interpret. And then I think that extends to your personal life — we want to give you leverage at the end of the day. We want the models in the product to give you leverage so you can create time for yourself to do the things you love.

Host

这是否也转化成了一种衡量生产力的方式?你怎么衡量杠杆?

Does that also translate to a way to measure productivity? How do you measure leverage?

Akshay

我觉得我们还没搞清楚。部分原因是它太多样了。每个人目标不同,真正要衡量的就是他们实现目标的能力:我们到底有没有帮到你?

I think we haven't figured this out yet. Part of the reason is that it's so diverse. Everyone has different goals, and really the true measurement is their ability to achieve that goal. Did we help you, or did we not?

Host

嗯。

Yeah.

Akshay

而且如果不预先知道那个目标,还要为每个人量身定制,这就非常难。

And it's very difficult without knowing what that goal is up front, and also tailoring it for every individual.

Host

ChatGPT 上的点赞和点踩不会给你任何信息,对吧?

And the thumbs up and thumbs down from ChatGPT doesn't give you anything, right?

Akshay

对。你不知道他们踩的是答案的内容、它的语气,还是它有没有帮助他们达成目标。这很难。但这个问题我们需要解决,整个行业也需要解决,因为如果这就是我们存在的意义,那就是我们衡量成功的方式。

Right. You don't know if they're thumbs-downing the content of the answer, the vibe of it, or whether it helped them with their goal. That's difficult. But it's something we will need to figure out, and the industry at large will need to figure out, because that's how we measure success — if this is what we're for.

Host

你觉得它改变了生产力以及你衡量生产力的方式吗?基本上,你说过现在能做的工作多得多,范围也大得多。它改变了吗?

Do you think it's changed productivity and how you measure it? Basically, you said there's a lot more work that can be done, a lot more scope. Has it changed?

Akshay

我觉得有一点是始终成立的:你真正想衡量的是,你的团队、个人、或者你自己,是否能够达成目标,或者是否离目标更近。但以前我们用代理指标,比如代码提交数、代码行数,诸如此类。

I think it was always true that what you really wanted to measure is whether your team, the individual, or you personally were able to hit the goal, or are closer to hitting whatever your goal is. But previously we used proxies for this — code commits, lines of code, whatever.

Host

故事点。

Story points.

Akshay

对,没错。故事点。还有——

Yeah, exactly. Story points. And like —

Host

顺便说一句,又绕回来了。

Coming back, by the way.

Akshay

也许吧。我是说,这也是变化的一部分。有了 AI 之后,这些代理指标开始失效了。你使用的 token 数量、或者你提交的 pull request 数量,不再和你的团队能不能达成目标、或者是否在轨道上高度相关。所以我们需要想出新的衡量方式。

Maybe. I mean, but that is part of the change. With AI now, those proxies are starting to fall apart. The number of tokens you use, or the number of pull requests you make, are no longer as hyper-correlated with whether your team is able to hit the goal or is on track to hit their goals. So we'll need to come up with new measurements.

衡量生产力与避坑 Measuring Productivity and Avoiding Traps

Host

给在听的经理们一个可以尝试的建议吧。

For the managers listening, give them one thing to try.

Akshay

对我来说,重要的是打席(at-bats)。我们作为团队,是否在锻炼一种能力——不仅要有打席的数量,更要有质量?我们能不能走完全程:产生想法、把它做出来、获得反馈、针对反馈作出回应、真正验证或推翻假设,然后再进入下一个想法?我们能不能高效地做到这一点?这关系到实际写出的代码、做出的设计、写出的规格,等等,但也关系到团队文化。我们是否有谦逊的心态,并能在整个过程中一次次重复,同时保持动力和兴奋?我觉得这在当下很重要,尤其是我们处在这项技术的前沿,有太多东西要建、太多事要做。这可能是我们最看重的东西。

For me, what's important is at-bats. Are we, as a team, building the muscle to have not just quantity of at-bats, but quality? Are we able to go all the way from generating an idea, building it out, getting feedback, reacting to that feedback, actually validating or invalidating the hypothesis, and going on to the next idea? Are we able to do that really efficiently? That goes to the actual code being written, the designs being made, the specs being written, whatever — but also the culture of the team. Do we have the humility and the ability to go through that process many times and stay motivated and excited throughout? That's the thing that I think is important now, especially when we're on the frontier of this technology and there's so much to build, so much to do. That's probably the most important thing we look at.

Host

在围绕团队工作衡量生产力时,人们会掉进哪些陷阱?我感觉有太多的情况是:“好,我们接入了一堆 LLM,我们有这个那个的看板”,但实际变化不大,对吧?

Any traps people fall into around measuring productivity with teamwork? I feel like there's a lot of 'okay, we added a lot of LLMs, we have dashboards for this and that,' but not much has changed, right?

Akshay

那就是陷阱,没错。

That is the trap, yes.

Host

而且这个问题的更广泛指向,是给正在搭建的经理和团队的——他们应该怎么处理这件事?

And the broader source of the question is for the managers and teams building — how should they approach this?

Akshay

我觉得陷阱可能就是:把动作和进展混为一谈。

I think maybe the trap is conflating motion and progress.

运动与进展 Motion vs progress

Akshay

我认为,由于我们现有的工具,行动比以往任何时候都更容易。但进展要求你对自己真正想要实现的目标非常明确和刻意。这又回到了我们的测量问题。我们之前谈到我们 OpenAI 能否找出衡量用户生产力的方法。由于多样性,这是个非常困难的问题,但作为一个团队,你应该对进展对你和你的团队意味着什么有一个非常明确和刻意的看法。如果没有这一点,你就很容易把这两者混为一谈。我觉得那个 batch 真是个好东西。

I think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be very prescriptive and deliberate about what you're actually trying to achieve. And it goes back to our question of measurement. We were talking about whether we at OpenAI can figure out how to measure productivity for our users. That's a very hard problem because of the diversity, but as a team you should have a really prescriptive and deliberate view on what progress looks like for you and for your team. And if you don't have that, then it's very easy to conflate these two things. I think that batch is a really great thing.

Host

我真的很高兴。我很喜欢关于行动与进展的讨论。我觉得这句话我们会写进文章里。你非常慷慨地付出了时间。非常感谢,并祝贺你获得 1000 万。

I'm really glad. I like the discussion between motion and progress. I think that's a quote that we're going to feature in the write-up. You've been very generous with your time. Thank you so much, and congrats on 10 million.

Akshay

是的,谢谢你的邀请。

Yeah, thank you for having me.

Host

下一个就是 100。两个月后。

The next one at 100. In two months.

Akshay

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →