OpenAI 内部:95%工程师用 Codex 写代码,100%PR 由 AI 审查

95% of OpenAI engineers use Codex, 100% PRs reviewed by AI

吴谢文 Sherwin Wu · Lenny 播客 · 2026-02-12 · 约 80 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI API 工程负责人 Sherwin Wu 透露,公司内部几乎所有代码都由 AI 生成,95%的工程师使用 Codex,100%的 PR 由 AI 审查,重度用户提交的 PR 数量多 70%。

Sherwin Woo, head of engineering for OpenAI's API, reveals that nearly all code at OpenAI is AI-generated, with 95% of engineers using Codex and 100% of PRs reviewed by AI, leading to 70% more PRs from heavy users.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 28)

全文 · Full transcript(中英对照)

OpenAI的AI代码生成 AI code generation at OpenAI

Host

我想从一个感觉像是 AI 进步晴雨表的问题开始,尤其是在工程领域。你现在自己还写代码吗?你和你团队的代码有多少是由 AI 写的?

I want to start with what's feeling like a barometer of progress in AI, especially in engineering. What percentage of your code, if you even write code anymore, and your team's code is written by AI at this point?

Sherwin Wu

我现在偶尔还是会写代码。但说实话,对于像我这样的管理者来说,用这些 AI 工具比手动编码要轻松得多。所以,就我自己和 OpenAI 的其他工程经理而言,我们所有的代码现在都是用 Codex 写的。但更广泛地说,内部有一种实实在在的能量,大家都在惊叹这些工具进步了多少,Codex 作为工具对我们来说变得多好用。我们很难精确衡量有多少代码是 AI 写的,因为绝大多数——我猜接近 100%——都是先由 AI 生成的。不过我们确实追踪到,现在绝大多数工程师每天都在用 Codex。所以 95% 的工程师使用 Codex。100% 的 PR 也每天由 Codex 审查。也就是说,任何进入生产环境的代码,在合并时 Codex 都会过目,并在 PR 中提出改进建议和修改意见。这就是我们内部看到的情况。但最令人兴奋的还是那种能量。另一个观察是,使用 Codex 更多的工程师会开出更多的 PR。他们实际上比不常使用 Codex 的工程师多开 70% 的 PR。而且差距还在扩大。我觉得那些开更多 PR 的人正在越来越熟练地使用这个工具,效率越来越高,那个 70% 的差距还在随时间增长,可能比我上次看这个数字时又大了。

I do write code occasionally now still. And I actually say for managers like myself, it's way easier to use these AI tools than to manually code at this point. And so I know for myself and some of the other engineering managers at OpenAI, all of our code is written by Codex at this point. But more broadly, there's just been so much energy. There's like a tangible energy internally around just how far these tools have gotten, how good Codex as a tool has gotten for us. And it's a little hard for us to exactly measure how much of the code is written because the vast majority of it, I'd say like close to 100%, is usually generated by AI first. What we do track though is that at this point the vast majority of engineers use Codex on a daily basis. So 95% of engineers use Codex. 100% of our PRs are reviewed by Codex daily as well. So basically any code that goes into production that's merged in Codex kind of has its eyes on and suggests improvements, suggests changes in the PRs. And so that's kind of what we're seeing internally. But by and large the most exciting is just the energy that there is. Another observation that we've had is engineers who tend to use Codex more open way more PRs. So they're actually opening 70% more PRs than the engineers who aren't using Codex as much. And the gap is widening. So I feel like the people who are opening more PRs are starting to learn how to use the tool more and more, get more efficient, and that 70% gap keeps growing over time and so might have actually increased since I last looked at the number.

Host

所以确认一下我理解得对不对:你是说 OpenAI 那 95% 的工程师的所有代码都是由 AI 写的,写完后再由他们审查?

So just to make sure we hear what you're saying: you're saying all of the code of these 95% engineers at OpenAI is written by AI. It's written and then they review it.

Sherwin Wu

对,没错。

Yep. Yep.

Host

这太疯狂了,但好像又已经不那么疯狂了,我们似乎正在习惯这件事。

It's crazy that that's almost like not crazy anymore, that we're just like getting used to this.

Sherwin Wu

我觉得还是需要一些适应过程,说清楚。也有一些工程师对 Codex 的信任度没那么高,但基本上每天都会有人被它的某个能力震惊到,他们对模型独立完成任务的信任门槛会一次次提高。我们科学副总裁 Kevin Whale 有句话,他喜欢说:“这是模型最差的时候了。”所以对于软件工程来说,这也是模型最差的时候。随着时间的推移,你会看到人们越来越信任它,而模型本身也会越来越好。

I think there's still some getting used to, to be clear. There's also some engineers who I think trust Codex a little bit less, but basically every day I talk to someone who is blown away by something that it can do, and their bar of trust, or how much they trust the model to do on its own, goes up over and over over time. And there's a quote from Kevin Whale, our VP of science here. He likes saying, "This is the worst the models will ever be." And so this is the worst that the models ever be for software engineering as well. And so over time you just see people trusting it more and more, and then we'll see the models get better and better as well.

Host

Kevin Whale 以前来过这个播客,他在这节目里也说过好几次这句话。

Kevin Whale, former podcast guest, he said exactly that line on this podcast a few times.

Sherwin Wu

对,现在那个叫 Peter the Claudebot/moldbotclaw。

Yeah. Peter the Claudebot/moldbotclaw is what it's called now.

工程师转型为管理智能体的技术主管 Engineers becoming tech leads managing agents

Host

95% 的工程师使用 Codex。100% 的 PR 都由 Codex 审查。我不知道过去几年还有什么工作比这变化更大。工程师正在变成技术负责人。他们在管理一群又一群的智能体。感觉我们就像巫师,在施放各种咒语。这些咒语会出去为你做事。你觉得人们还没有定价进去的是什么?一人独角兽公司的二阶或三阶效应,去赋能另一个一人独角兽公司。可能还有上百家其他小公司在构建定制软件。所以我认为我们可能真的会进入一个 B2B SaaS 的黄金时代。

95% of engineers use Codex. 100% of our PRs are reviewed by Codex for engineers. I don't know what job has changed more in the past couple years. Engineers are becoming tech leads. They're managing fleets and fleets of agents. It literally feels like we're wizards casting all these spells. And these spells are kind of like going out and doing things for you. What do you think people aren't pricing in yet? The second or third order effects of the one-person billion dollar startup to enable a one person billion dollar startup. There might be a hundred other small startups building bespoke software. And so I think we might actually enter into a golden age of B2B SaaS.

Host

我越来越常听到,当智能体不工作时,人们会感到压力。OpenAI 内部有一个团队正在做实验,他们维护着一个 100% 由 Codex 编写的代码库。他们遇到了你描述的那些问题。通常你会说,好吧,我卷起袖子自己搞定。但那个团队没有那个逃生舱。

I've been hearing more and more there's this stress people feel when their agents aren't working. There's a team that's actually doing an experiment right now within OpenAI where they are maintaining a 100% Codex written codebase. They run into the exact problems that you're describing. And so usually you're like all right I'll roll up my sleeves and figure it out. Team doesn't have that escape hatch.

Host

你分享过,在 AI 领域,倾听客户并不总是正确的策略。这个领域和模型本身变化太快了。它们往往会自我颠覆。模型会轻松吃掉你的脚手架。对那些说“好吧,我不想错过这班车”的人,你有什么建议?确保你为模型的未来方向而构建,而不是为它们的今天而构建。我们科学副总裁 Kevin Whale 有句话,他喜欢说:“这是模型最差的时候了。”今天,我的嘉宾是 Sherwin Wu,OpenAI API 和开发者平台的工程负责人。考虑到几乎每家 AI 初创公司都在集成 OpenAI 的 API,Sherwin 对正在发生的事情和未来走向有着极其独特和广阔的视角。在简短地感谢我们出色的赞助商之后,让我们开始吧。

You've shared that listening to customers not always the right strategy in AI. The field and the models themselves are just changing so so quickly. They tend to like disrupt themselves. The models will eat your scaffolding for breakfast. What's your advice to folks that are like, "Okay, I don't want to miss the boat." Make sure you're building for where the models are going and not where they are today. There's a quote from Kevin Whale, our VP of science here. He likes saying, "This is the worst the models will ever be." Today, my guest is Sherwin Wu, head of engineering for OpenAI's API and developer platform. Considering that essentially every AI startup integrates with OpenAI's APIs, Sherwin has an incredibly unique and broad view into what is going on and where things are heading. Let's get into it after a short word from our wonderful sponsors.

氛围编码与工程师角色变迁 Vibe Coding and the Changing Role of Engineers

Host

一位开发者最近分享说,他在工作中使用 Codex,感觉每当它完成一些事情时,他就相信它做对了,几乎可以确定直接提交到主分支也没问题。

A developer recently shared that he uses Codex for his work and he feels like anytime it does things, he just trusts that it has done the right job and he's almost certain he could just commit it to master and it'll be great.

Sherwin Wu

是的。他是 Codex 的优秀用户。我知道他和团队保持密切联系,给了我们很好的反馈。他用这个工具我并不意外。我是说,它叫 Open Claw。Open Claw 是个很棒的产品。然后我看到,这是最近的事,但今天早上,Mould 的书也被分享了,看到所有 AI 智能体互相交谈,感觉很不真实。我听到的,这基本上就是现实中的“她”。

Yeah. He's a great user of Codex. I know he's in close touch with the team, gives us great feedback. Not surprised that he uses it. I mean, it's called Open Claw. Open Claw is a great product. And then I saw that this is very recent, but this morning, Mould's book kind of was shared as well, and seeing all the AI agents talk to each other is pretty surreal. It's basically her happening in real life, is what I'm hearing.

Host

所以回到我们正在经历的疯狂时刻,尤其是对工程师来说。我们从你写每一行代码,到现在 AI 写你所有的代码。我不知道过去几年还有什么工作变化更大,像这种我们没预料到会变化这么大的工作,工程师的工作变得如此不同。在工程师的整个职业生涯中,过去几年已经转变为我不再写代码了。你想象未来几年工程师的角色和软件工程师的工作会是什么样子?那工作到底是什么?

So just coming back to this crazy moment we are living through for engineers in particular. We've gone from you write every line of code to now AI is writing all of your code. I don't know what job has changed more in the past couple years, like a job that we didn't expect to change this much, where just the job of an engineer is so different. In the entire lifespan of an engineer, in the past couple years, it's now shifted to I don't write any more code. How do you imagine the role of an engineer and the job of a software engineer looks in the next couple years? Just what is that job?

Sherwin Wu

是的,说实话,看到这一切真的很酷,这也是兴奋点所在,因为这份工作很可能在未来一到两年内发生显著变化。不过感觉我们还在摸索中,所以有一种兴奋感,我知道尤其是一些软件工程师觉得,我们正处于一个罕见的时刻,也许在未来 12 到 24 个月内,我们可以自己摸索并设定自己的标准,就我看到的趋势而言。所以我认为大家都在说一件事,那就是通常个人贡献者工程师正在成为技术负责人。他们现在基本上像经理一样。他们在管理成群的智能体。我知道我团队里的很多工程师基本上同时有 10 到 20 个线程在进行。显然不是都在运行 Codex 任务,而是很多并行线程。他们在检查这些线程在做什么。他们在引导智能体和 Codex,并给予反馈。所以他们的工作已经从仅仅编写代码本身,变成了几乎像一个经理。至于我认为未来一到两年会如何发展。我经常回到的一个比喻,其实来自我大学时读过的一本编程教材,叫 SICP。不知道你听说过没有。《计算机程序的构造和解释》。在麻省理工学院它非常流行,而且很长一段时间被用作编程入门课程的教材,它有一群狂热追随者。它教你编程,教你一种叫做 Scheme 的 Lisp 方言,所以它向你介绍函数式编程,在这方面非常启发思维。但那本书让我印象深刻的是,我在大学读它的时候,它的开头把编程描述为一门学科,并把它比喻成巫术。它说软件工程师就像巫师,编程语言就像咒语,你发出这些咒语,这些咒语就会出去为你做事,挑战在于你要说什么咒语才能让程序做你想做的事。这本书写于 1980 年,所以是挺久以前了,我认为这个比喻实际上一直延续了下来。而且我认为它正在我们进入这个“氛围编程”的新时代,或者说软件工程会变成什么样时上演,因为编程语言基本上就是这些咒语。它们随着时间的推移而变化,挑战一直存在,趋势是让计算机通过编程做你想做的事变得越来越容易。我认为当前的 AI 浪潮很可能是那场演变的下一阶段。现在它真的成了咒语,因为你可以告诉 Codex,告诉 Cursor 你到底想做什么,然后它就会全部为你完成。我特别喜欢巫师和巫术的比喻,因为我认为我们目前的状态正在开始向《魔法师的学徒》靠拢,就是《幻想曲》里的,米老鼠找到魔法师的帽子,然后尝试做各种事情。我实际上认为这是一个非常贴切的比喻,因为第一,现在它真的很强大。你能做的这些咒语杠杆率极高,但你得知道自己在做什么,对吧?就像在《魔法师的学徒》里,整个情节是米老鼠失控了,扫帚发疯,到处是水。我想他实际上是让扫帚去执行任务,然后自己睡着了。所以,你知道,这就是“氛围编程”的极致,然后最终老魔法师回来收拾一切。当我看到工程师同时做 20 个不同的 Codex 线程时,这需要一些技巧、资历和很多思考,因为你要确保模型不会偏离轨道。你肯定不想完全走开忽略它。但它的杠杆率也极高,一个非常资深、精通这些工具的工程师现在可以通过他们正在做的事情做更多的事。我认为这也是它有趣的地方。它真的让我们感觉现在就像巫师。你知道,感觉我们更接近拥有这种神奇的体验,我们施放所有这些咒语,让软件为你做所有这些事情。

Yeah, it's honestly been really cool to see, and it's part of where the excitement is, because the job is likely going to change pretty significantly over the next one to two years. It kind of feels like we're still figuring things out though, and so there's this excitement, I know especially from some of the software engineers, of like we're in this rare moment, maybe over the next 12 to 24 months, where we'll get to figure things out ourselves and set our standards for ourselves in terms of where I see this moving. So I think there's a common thing that everyone's saying, which is people are generally like IC engineers are becoming tech leads. They're basically like managers now. They're managing fleets and fleets of agents. I know many of the engineers on my team basically have like 10 to 20 threads kind of being pulled on at the same time. Obviously not active running Codex jobs, but just a lot of parallel threads. They're checking in on what they're doing. They're steering the agents and Codex and giving it feedback. And so their job has kind of really changed from just writing the code itself into being almost like a manager. In terms of where I think this will go one to two years from now. So one metaphor that I kind of always come back to here is actually from this programming textbook that I read back in college called SICP. I don't know if you've heard of it. Structure and Interpretation of Computer Programs. At MIT it was really popular and it was actually used as the introductory textbook for the intro programming course for a very long time, and it kind of has this cult following. It teaches you programming, it teaches you a dialect of Lisp called Scheme, and so it introduces you to functional programming, it's very mind-opening in that way. But the thing that was memorable for me about that book, I read it in college, the very beginning of it kind of describes programming as a discipline and draws this metaphor to basically like sorcery. Like it says software engineers are like wizards and programming languages are like incantations, and you're issuing these spells and these spells are kind of going out and doing things for you, and the challenge is what incantation do you have to say to make the program do what you want. And this book was written in 1980, so this is a while ago, and I think that metaphor has actually kind of persisted over time. And I think it's actually playing out as we move into this new era of vibe coding or just what software engineering will look like, because programming languages were basically these incantations. They've changed over time and the challenge has always been, and the trend has been that it's been easier and easier to get the computer to do what you want via programming. And I think the current wave of AI is probably the next stage of that evolution. It is now literally incantations because you can tell Codex, you can tell Cursor exactly what you want to do and then it'll all go do it for you. And I particularly like the wizard and the sorcery analogy because I think our current state is starting to move towards kind of like the Sorcerer's Apprentice, you know from Fantasia, where Mickey Mouse finds the sorcerer hat and he tries to do all these things. And I actually think it's a really apt analogy because one, it's just really powerful now. These incantations you can do are extremely high leverage, but you kind of have to know what you're doing, right? Like in Sorcerer's Apprentice, the whole plot is Mickey goes wild, the brooms go crazy and everything's flooding. I think he literally sets the brooms off on a task and then goes asleep. And so, you know, it's like vibe coding at its greatest, and then eventually the old sorcerer comes back and cleans everything up. And when I see engineers kind of doing these 20 different Codex threads at a time, there is some skill and some seniority and a lot of thought that needs to go into this, because you want to make sure that the models aren't going off the rails. You definitely don't want to just completely go away and ignore the thing. But it's also extremely high leverage, like a very senior engineer who's really proficient with these tools can now just do way more things via what they're doing. And I think this is also what makes it fun. It literally feels like we're wizards now. You know, it feels like we're closer to having this magical experience where we're casting all these spells and having software do all these things for you.

Host

我正好在想《魔法师的学徒》这个比喻,就像你描述的那样。所以很高兴你提到了。之前一位播客嘉宾把它描述为你有一个可以许愿的精灵,这是一个有用的框架,因为你必须非常清楚你想要什么愿望,比如如果你想变大,能变多大。

I was thinking of the Sorcerer's Apprentice exactly as the metaphor as you were describing that. So I'm glad you went there. A previous podcast guest described it as you have a genie that you can grant wishes and it's a useful frame because you have to be very clear about the wish you want, like if you want to be big, how big it could be.

巫师书比喻的持久力 Staying power of the wizard book analogy

Host

是啊。或者可能像猴爪那种情况,你知道的,你得到了你想要的,但副作用是什么?

Yeah. Or it might be like the monkey's paw type thing where you know it's like you got what you want but what are the side effects?

Sherwin Wu

是的。我觉得这个类比很棒。对我来说,疯狂的是那本书的持久影响力。它被称为巫师书,人们这么叫它,因为那是贯穿全书的隐喻。我们现在基本上已经达到了那个点,这真的很酷。

Yeah. I think that analogy is great. The crazy thing for me is just the staying power of that book. It's called the wizard book, people call it the wizard book because that is the metaphor that they kind of weave throughout the book. And we've basically reached that point now, which is really cool.

智能体失败时的压力 Stress when agents fail

Host

我想顺着两条线索聊。一是,我越来越多地听到,当智能体不工作时,人们会感到压力。你启动所有这些 Codex 智能体,然后必须时刻盯着它们。哦,有一个不工作了。我在浪费时间。你有这种感觉吗?你的团队也有这种感觉吗?

There's two kind of threads I want to follow here. One is I've been hearing more and more there's this stress that people feel when their agents aren't working. You fire off all these Codex agents and then you have to keep stay on top of them. Oh one's not working. I'm wasting time. Do you feel that? Do you feel that across your team at all?

Sherwin Wu

是的。我是说,这经常发生。我实际上认为这正是目前这一切有趣的地方,因为这些模型并不完美。这些工具并不完美,我们仍在试图找出如何最好地与 Codex 或这些 AI 智能体互动来完成工作。我们经常看到这种情况。我们内部有一个特别有趣的团队。有一个团队现在正在 OpenAI 内部做一个实验,他们基本上在维护一个 100% 由 Codex 编写的代码库。你知道,你会让 AI 写代码,但显然你最终会重写很多,你可能需要仔细检查并修改,但这个团队完全接受了 Codex,完全投入其中。他们遇到了你描述的确切问题,即他们的挑战是:我想构建这个功能,但我无法让智能体做到。通常有一个逃生口,你会说好吧,我卷起袖子自己搞定,然后不用 Codex,而是用 Tab 补全和 Cursor 之类的工具,但这个团队在实验中并没有那个逃生口。所以挑战就变成了如何让智能体做到这一点?我实际上认为我们会发布一篇关于我们一些经验的博客文章。但很多迷人的范式和最佳实践正在从中涌现。我们注意到一件有趣的事,我不知道这是否是你的感受,但我们在这里确实感受到了,很多时候,当编码智能体没有按你的意愿行事时,通常是上下文和你给它的信息出了问题。你要么指定不足,要么没有足够的信息告诉智能体或 Codex 如何做某事。所以当你必须通过这个来解决时,挑战就是添加文档,实际上绕过这个限制,基本上把你头脑中的隐性知识以某种方式编码到代码库中,要么通过代码注释本身,要么通过代码结构,要么通过文本文件,比如 MD 文件、技能,或仓库内的任何类型的额外资源,以便模型能更好地完成它的任务。从这个小组中还有一大堆其他的经验,我觉得探索起来很迷人。但没错,移除不再使用 AI 的逃生口,让他们开始拼凑出如果我们真的想依赖智能体就必须解决的许多问题。

Yeah. I mean, it happens all the time. And I actually think this is where the interesting part of all of this lies right now because these models aren't perfect. These tools aren't perfect and we're still trying to figure out how to best interact with Codex or with these AI agents to get work done. We see this come up all the time. There's a particularly interesting team that we have internally. So there's a team that's actually doing an experiment right now within OpenAI where they are basically maintaining a 100% Codex-written codebase. So you know, you'll have the AI write code but you'll obviously end up rewriting a lot of it and you might need to double check and change things, but this team is just fully Codex-pilled and just leaning in entirely. And they run into the exact problems that you're describing, which is like their challenge is: I want to get this feature built but I can't get the agent to do it. And so usually there's an escape hatch where you're like all right I'll roll up my sleeves and figure it out and then instead of using Codex I might use tab complete and cursor and things like that, but this team for the experiment doesn't have that escape hatch. And so then the challenge is how do I get the agent to do this? And I actually think we're going to be publishing a blog post from some of our learnings here. But a lot of fascinating paradigms and best practices are falling out of this. One interesting thing that we've noticed, I don't know if this is what you feel, but we definitely feel it here, is a lot of the time, when the coding agent is not doing what you want, it's usually a problem with context and just information that you've given it. It's just you've either underspecified or there's just not enough information around how to do something available to the agent, available to Codex. And so when you have to solve it through that, the challenge is then to add documentation and actually work around this limitation and basically encode more tribal knowledge that's in your head somehow into the codebase either via code comments itself or code structure itself or via text files like MD files, skills, any type of additional resources within the repository so that the model can better do its task. There's a whole bunch of other learnings from this group which I think is fascinating to explore. But yeah, kind of removing that escape hatch of no longer using AI has allowed them to start piecing together a lot of the problems that we'll have to solve if we really want to lean into agents.

代码审查挑战与解决方案 Code review challenges and solutions

Host

人们遇到的另一个问题是,你谈到人们如何疯狂地提交 PR,如果使用 AI,PR 数量会多得多。显然代码审查正成为一个更大的挑战。你的团队有没有想出什么办法来帮助加快速度,使其能够规模化,而不是给人们创造一份糟糕的工作,让他们整天坐在那里审查 PR?

Another issue people run into, you talked about how people are shipping PRs like crazy, a lot more PRs if they're working with AI. Obviously code review is becoming a bigger challenge. Is there anything you've figured out in your team to help speed that up to make that scale and not just create this terrible job for people where they're just sitting there reviewing PRs all day?

Sherwin Wu

是的,我的意思是,有一件事是,目前 Codex 审查我们所有的 PR,100%。所以我实际上认为一件非常有趣的事情是,我们倾向于立即交给模型的事情往往是我们讨厌的事情,或者是软件工程中最无聊的部分。这也是为什么现在更有趣了,因为我们能做更多有趣的事情。对我来说,就我个人而言,我真的很讨厌代码审查。它对我来说是最糟糕的事情之一。我记得我大学毕业后的第一份工作,是在 Quora。我负责新闻推送,所以我拥有新闻推送的代码。因此我是新闻推送的审查者,而那是每个人都会碰到的核心代码。所以每天早上我登录后,就会有 20 到 30 个代码审查。我就像,天哪,我得看完所有这些。我会拖延,然后它增长到 50 个。所以有很多代码审查。Codex 非常擅长审查代码。所以我们注意到的一件事是,特别是 52 版本,已经变得极其擅长审查代码,尤其是当你把它引导到正确的方向时。所以对于代码审查,是的,我们创建了很多 PR,但 Codex 审查了所有 PR,它使代码审查从我不知道的 10 到 15 分钟的任务,有时甚至变成只需两到三分钟的任务,因为你已经内置了一堆建议。很多时候,人们,特别是对于小的 PR,实际上甚至不需要人来审查。我们在某种程度上信任 Codex。原作者会看看 Codex。你知道,代码审查的好处是有一双额外的眼睛来确保你没有做任何蠢事。Codex 现在是一双相当聪明的额外眼睛,所以这是我们大力依赖的东西。一般的 CI 流程以及推送和部署流程现在在内部也通过 Codex 大量自动化了。如果你和很多工程师交谈,最让他们烦恼的事情是,在你写完漂亮的代码之后,如何把它投入生产?你必须运行所有这些测试,你必须处理 lint 错误,代码审查。你可以用 Codex 做很多自动化的事情,所以我们实际上在内部构建了一些工具来帮助自动化这个过程,自动化 lint,如果有 lint 错误,Codex 很容易修复。然后它可以直接打补丁,然后重新启动 CI 流程。

Yeah, I mean one thing is Codex reviews 100% of all of our PRs at this point. And so I actually think one really interesting thing that's happened is the things that we tend to hand to the models immediately tend to be the things that annoy us or are the most boring parts of software engineering. It's also why it's more fun now because we get to do more of the fun things. For me, speaking more for myself, I really hated code reviews. It was like one of the worst things for me. And then I remember in my first job out of college, it was at Quora. I owned I was working on the newsfeed and so I owned the code for the newsfeed. And so I was a reviewer for Newsfeed and it was just the central piece of code that everyone would touch. And so I would just every morning I'd log in and be like 20 to 30 code reviews. I just like oh my goodness I got to get through all these. I would procrastinate and then it grows to like 50. And so there's a lot of code reviews. Codex is really good at reviewing code. So actually one thing that we've noticed that 52 in particular has gotten extremely strongly adept at is reviewing code and especially when you kind of steer it in the right direction. And so for code reviews, yeah we create a lot of PRs but Codex reviews all of them and it makes code reviews go from a I don't know 10 to 15 minute task to sometimes even just a two to three minute task because you have a bunch of suggestions already baked in. A lot of the times people will, especially for small PRs, you actually don't even need people to review. We kind of trust Codex in this way. The original author kind of looks at Codex. It is, you know, the benefit of code review is to have a second pair of eyes to make sure that you're not doing anything dumb. Codex is a pretty smart second pair of eyes at this point and so that's something that we've heavily leaned into. The general CI process and the post push and deployment process has also been heavily automated via Codex internally at this point. If you talk to a lot of engineers, the thing that annoys them the most is after you've written your beautiful code, how do you get it into production? You got to run through all these tests, you got to lint errors, code review. There's a lot of automated stuff you can do with Codex and so we've actually built some tools internally that help automate that process, automate the lint, if there's a lint error, it's a very easy Codex fix. And then it could just patch it and then restart the CI process.

AI辅助代码审查与模型使用 AI-assisted code review and model usage

Host

所以这一切都是为了尽可能减少工程师的工作量,其副产品是他们现在可以合并并推送更多的 PR。Codex 在写代码,Codex 在审查自己的代码。我好奇你是否愿意使用其他模型来审查你模型的成果。这是一条路径,还是说它已经足够好了,我们不需要别的了?

So all of that is we're trying to collapse as much work for an engineer as possible, and the byproduct of which is they can now merge and push out a lot more PRs. Codex is writing the code, Codex reviewing its own code. I'm curious if you are open to using other models to review your models' work. Is that a path, or is it just good enough, we don't need anything else?

Sherwin Wu

我会说这里确实存在一个循环问题,就像回到魔法师的学徒一样,你要确保不要让扫帚失控。所以,我们非常谨慎地决定哪些 PR 完全由 Codex 审查。大多数人显然还是会看一下自己的 PR。所以这并非降到零,而是从 100% 的注意力降到大约 30%,这有助于推动事情进展。关于多个模型,我们内部显然测试了很多模型,所以有很多选择。我们较少使用外部模型。我们认为“吃自己的狗粮”并从中获得反馈很重要。但你也可以使用许多内部模型变体来获得不同的视角,我们发现这效果很好。

So, I will say there's definitely a circular thing here, and like going back to the sorcerer's apprentice, you want to make sure you're not letting the brooms go crazy here. And so, we're very thoughtful, I'd say, around which PRs are completely just Codex reviewed. Most people still obviously take a look at their PRs. And so it's not like it's going to zero. It's more like going from 100% attention to like 30% attention, which just helps things push through. In terms of multiple models, we obviously test a lot of models internally, and so we have a lot of those. We use external models less. We think it's important to kind of dog food our own models and get feedback there. But you can also use a lot of internal variants of models to give you a different perspective here as well, and we found that to work quite well.

Host

好的。为了确保我们了解 OpenAI 在 AI 和代码方面的现状,让我理解一下,然后我想换个话题。目前 OpenAI 的 100% 代码都是由 Codex 编写的。可以这样理解吗?

Okay. So just to make sure we get a barometer of today's world at OpenAI in terms of AI and code, just so I understand and then I want to move on to a different topic. 100% of code across OpenAI is written by Codex at this point. Is that the way to frame it?

Sherwin Wu

我不会说今天生产环境中运行的 100% 代码都是由 AI 编写的。而且很难进行归因。但几乎每个工程师现在都在所有任务中大量使用 Codex。所以,如果让我猜测,目前绝大多数代码很可能都是由 Codex 编写的。

I wouldn't make the statement that 100% of code running in production today was written by AI. And it's kind of hard to do attribution there. But almost every engineer heavily uses Codex in all of their tasks at this point. And so, if I were to guesstimate, the vast majority of code at this point was probably authored by Codex.

AI时代经理角色演变 Manager role evolution with AI

Host

好的。关于 IC 角色、IC 工程师的工作有很多讨论。但关于管理者,尤其是工程管理者的角色变化,讨论较少。作为管理者,你的生活随着 AI 的兴起发生了怎样的变化?你认为未来管理者的角色是什么?

Okay. So there's a lot of talk about the IC role, the work of an IC engineer. There's less talk about the changing role of a manager, especially an engineering manager. How has your life as a manager changed with the rise of AI, and what do you think the role of a manager is in the future?

Sherwin Wu

这肯定比工程师的变化小。目前还没有针对管理者的 Codex。不过,我在一些管理任务中确实大量使用了 Codex。我想说有几件事正在改变。有一些趋势。所以我认为变化还不大,但我看到了趋势,如果你推演下去,你大概能看出很多事情的走向。越来越清晰的一点是,Codex 极大地提升了顶尖执行者的生产力。所以这确实——我认为这可能更广泛地适用于整个社会的 AI——那些真正投入的人,或者有高度自主性的人,或者会熟练使用这些工具的人,会让自己如虎添翼。我现在也注意到了这一点,顶尖执行者的生产力变得高得多。因此,你会看到团队生产力差距拉大。我作为管理者的一个一贯理念是,把大部分时间花在顶尖执行者身上,确保他们畅通无阻、开心、感到有生产力且被倾听。我认为在 AI 世界里,这一点更加重要,因为你的顶尖执行者会借助这些工具飞速前进。一个例子是,维护 100% Codex 生成代码库的团队,让他们放手去干,看看会发生什么,这已经带来了回报。所以我认为这是我看到的一个趋势,管理者花更多时间在顶尖执行者身上很可能会继续下去。另一件事是,这更多是一个观察,但我的感觉是,有了这些面向管理者的 AI 工具,不一定是写代码,而是像带有组织知识的 ChatGPT,能够更好地进行研究和理解组织背景。另一个好例子是,我们正在做绩效评估,使用连接了 GitHub、Notion 文档和 Google Docs 内部知识的 ChatGPT,很容易就能了解这个人过去 12 个月做了什么,并为其撰写一份深度研究报告。我的感觉是,管理者将能够管理更大的团队,就像软件工程师管理 20 到 30 个 Codex 一样。我的感觉是,这些工具将使管理者(人员管理者)拥有更高的杠杆效应,并使他们能够管理远超当前软件工程最佳实践(6 到 8 人)的团队。你可以在非工程领域看到这一点,比如支持或运营,以前支持团队的规模可能有限,但随着你能将更多事情交给智能体,你实际上可以完成更多工作,同时以这种方式管理更多人。我认为同样的事情也可能发生在人员管理上,尤其是在科技公司。我们已经看到了这一点。有些团队中,工程管理者管理着相当多的人,并且做得相当熟练,因为借助这些工具,他们可以获得更高的杠杆效应,更好地了解团队在做什么,更好地理解组织背景,并以这种方式运作。

It's definitely changed less than an engineer. There's no Codex for managers just yet. However, I use Codex quite a bit for some of the more managerial tasks that I do. I'd say a couple things are changing. There are some trends. So I don't think it's changed that much yet, but I see trends, and if you play it out, you can kind of see where a lot of this is going. One thing that's becoming increasingly clear is Codex really empowers top performers to be a lot more productive. And so it really, and I think this is maybe true for AI more broadly across society, which is the people who really lean in, or the people who have high agency, or will get good at these tools, will kind of supercharge themselves. And so I'm noticing this now as well, which is the top performers end up being a lot more productive. And so you see a broader spread in team productivity in this way. One thing that I've always done as a management philosophy is to spend the majority of my time with top performers, just make sure they're unblocked, make sure they're happy, make sure they feel productive and heard. I think this is even more true in an AI world where your top performers are going to just really be shooting ahead using these tools. I think one example is the team that's maintaining a 100% Codex-generated codebase, just letting them kind of rip and see what's happening there is something that's paid dividends. So I think that's one trend I'm seeing, where spending even more time with top performers for managers is likely going to continue. The other thing is, this is more an observation, but my sense is with a lot of these AI tools available to managers, less like writing code but just things like ChatGPT with organizational knowledge, being able to do research and understand organizational context a lot better. Another good example is we're doing performance reviews right now, and it's actually really easy to use ChatGPT with internal knowledge hooked up to GitHub and our Notion docs and Google Docs to get a really good sense of what this person has done over the last 12 months, and writing a little deep research report for it. My sense is managers will be able to manage much larger teams in this world, kind of like how software engineers are managing 20 to 30 Codexes. My sense is these tools will allow managers, people managers, to be higher leverage, and it will allow them to manage teams of way more than the current best practice of six to eight for software engineering. You kind of see this applied to non-engineering domains like support or operations, where previously the size of a support team might be limited, but as you can pass off more things to agents, you can actually do more work and also manage more people this way. I think the same thing might happen for people management as well, especially in tech companies. And we're already seeing this. There are some teams where there are EMs managing quite a few people, and they're doing it pretty adeptly because of some of these tools where they can get higher leverage and understand what their team's doing, understand organizational context a little bit better, and operate in that way.

Host

我很喜欢这个建议,你描述的方式是你一直倾向于顶尖执行者,花更多时间与他们在一起,为他们扫清障碍,确保他们开心。Marc Andreessen 刚刚来过播客,他的说法是 AI 让好人变得更好,让优秀的人变得卓越。

I love this advice that the way you described it is you've always leaned into top performers and spent more time with them, unblock them, make sure they're happy. The way Marc Andreessen, he was just on the podcast, the way he phrased it is AI makes good people better and it makes great people exceptional.

Sherwin Wu

是的。

Yeah.

Host

而你在这里说的是,越来越多地这样做可能是正确的做法。花更多时间与你团队中最优秀的人在一起,为他们扫清障碍,确保他们拥有一切所需。

And what you're saying here is just doing this more and more is probably the right move. Spending more time with the best people on your team to unblock them, make sure they have everything they need.

Sherwin Wu

是的。

Yeah.

与AI模型互动的最佳实践 Best practices for interacting with AI models

Sherwin Wu

现在一个很好的例子是,内部有一群工程师确实是代码专家,他们正在思考与这个模型互动的最佳实践。这对他们来说是一件杠杆效应极高的事情。所以作为管理者,我就说,去探索吧。无论得出什么最佳实践,我们都要与整个组织分享。我们会举办知识分享会,分享文档和最佳实践。像这样的事情能提升每个人。我认为这是另一个趋势的例子,即顶尖表现者确实变得异常出色。

A very good example right now is there are a group of engineers internally who are really code experts and are thinking through what the best practices are for interacting with this model. That is an extremely high-leverage thing for them to do. So as a manager, I'm just like, yeah, go explore this. Whatever best practices come out of this, we have to share with the org. We'll do knowledge sharing sessions, share documents, and best practices everywhere. Things like that elevate everyone. I view that as another example of this trend where the top performers really get exceptional.

被低估的:一人十亿美元初创公司 What people aren't pricing in: one-person billion-dollar startup

Host

人们只是觉得这事很大。AI 变化如此之大。世界在变。这将是一件大事。你认为人们还没有将什么因素纳入对未来的考量?比如有什么例子是你觉得我们还没意识到的?

People just have a sense this is big. AI is changing so much. The world is changing. It's going to be a huge deal. What do you think people aren't pricing in yet into what will change into where things are heading? Just like what's an example of something you think like okay we're not realizing this yet.

Sherwin Wu

在这次 AI 浪潮中,我最喜欢的一个说法是“一人十亿美元初创公司”的想法。我想 Sam 可能是第一个提出这个说法的人,但想想就觉得很迷人。如果人们有如此高的杠杆效应,那么某个时候很可能会出现一家一人十亿美元的初创公司。虽然我觉得这很酷,但我认为人们并没有真正考虑到它的二阶或三阶效应。一人十亿美元初创公司意味着一个人可以利用这些工具拥有更多的自主权和杠杆,从而极其轻松地完成所有事情,创造出价值十亿美元的东西。但还有其他影响。一是如果一个人有可能创建一家一人十亿美元的初创公司,那也意味着人们总体上更容易创建初创公司。我认为一个二阶效应是巨大的创业潮和中小企业潮,任何人都可以构建任何软件。我们开始看到这在 AI 创业场景中上演,软件变得更加垂直化。为某个垂直领域创建 AI 工具往往效果很好,因为你深耕该领域并理解用例。如果推演 AI 的发展,没有理由不能有 100 倍于现在的这类初创公司。所以我们可能看到的一个世界是,为了支撑一家一人十亿美元的初创公司,可能会有上百家其他小初创公司构建定制软件,来支持其他类型的一人十亿美元初创公司。我们可能会进入一个 B2B SaaS 和软件以及初创公司的黄金时代。这是一个非常有趣的趋势,因为随着构建软件和运营公司变得越来越容易,你可能会看到更多的初创公司。我一直在想:可能有一家一人十亿美元的初创公司,但也可能有一百家一亿美元的初创公司和数万家一千万美元的初创公司。对个人来说,拥有一家一千万美元的企业就很棒了——这辈子就够了。所以我们可能会看到这样的爆发,我觉得人们并没有考虑到这一点。还有另一个三阶效应。随着我们做出更远的预测,不确定性很大。如果我们进入一个由微型公司为拥有一两家公司的个人构建软件的世界,创业生态系统会改变,VC 生态系统也会改变。我们可能会最终只剩下少数几家大公司提供平台并支持所有这些初创公司。但那些能带来 100 倍或 1000 倍回报的风险投资规模初创公司实际上可能会减少,因为会出现大量规模在 1000 万到 5000 万美元的小公司,这些公司对风险投资回报不利,但对那些利用 AI 为自己创业的高自主性个人来说非常有利。

One of my favorite phrases from this whole AI wave is the idea of the one-person billion-dollar startup. I think Sam may have been the first to say it, but it's fascinating to think about. If people are so high leverage, at some point there will likely be a one-person billion-dollar startup. While I think that's really cool, I think people aren't really pricing in the second or third order effects of this. What the one-person billion-dollar startup implies is that one person can have so much more agency and leverage using these tools that it is super easy for them to get everything done to create something worth a billion dollars. But there are other implications. One is that if it's possible for a person to create a one-person billion-dollar startup, it also means it's way easier for people to just create startups in general. I think a second-order effect is a huge startup boom and SMB-style boom where anyone can build software for anything. We're starting to see this play out in the AI startup scene where software became more vertical-oriented. Creating an AI tool for some vertical tends to work well because you lean into that domain and understand the use case. If you play out AI, there's no reason why you can't have 100x more of these startups. So one world we might see is that to enable a one-person billion-dollar startup, there might be a hundred other small startups building bespoke software that supports other types of small one-person billion-dollar startups. We might enter a golden age of B2B SaaS and software and startups in general. That's a really interesting trend because as it gets easier to build software and run a company, you might see way more startups. I've been thinking: there might be one one-person billion-dollar startup, but there might be a hundred $100 million startups and tens of thousands of $10 million startups. As an individual, having a $10 million business is pretty great—you're set for life. So we might see an explosion in that way, and I feel like people aren't pricing that in. There's another third-order effect. As we get to further out predictions, there's a lot of uncertainty. If we move to a world with micro companies building software for one or two people who own the company, the startup ecosystem will change, the VC ecosystem will change. We might end up with a handful of big players offering platforms and supporting all these startups. But the types of venture-scale return startups that can 100x or 1000x your investment might actually shrink if you have a bunch of smaller $10 to $50 million companies, which are not great for venture returns but are great for high-agency individuals leaning into AI to build businesses for themselves.

对一人十亿美元初创公司的质疑 Skepticism about the one-person billion-dollar startup

Host

我喜欢我们讨论了多少层效应。我现在想听第四层效应。Sherwin,我只是开玩笑。我做不到。第四层对我来说太烧脑了。我想不了那么远。

I love how many order effects we've been through. I want to hear the fourth order effect now. Sherwin, I'm just joking. I can't. The fourth order is too gigabrain for me. I can't think that far ahead.

Sherwin Wu

这就像《盗梦空间》,每深入一层,一切都会变慢。

It's like Inception where everything gets slower every time you go deeper into a layer.

Host

所以,关于十亿美元初创公司——我经常思考这个问题,因为我在做的事情不是风险投资规模,也不是超高杠杆。但仅仅看到我从最荒谬的事情中收到多少支持工单,我就很难想象一个人能扩展到十亿美元。我不看好这个十亿美元初创公司。即使 AI 在帮助你,达到十亿美元时,除非你的 ACV 非常高且客户很少,否则处理支持问题是难以规模化的。人们可以自己解决问题,但他们还是会发邮件给支持。除非你有一堆承包商——我不知道那还算不算单人公司——否则在没有人帮忙处理支持的情况下,很难将一家十亿美元初创公司规模化。AI 只能帮你到一定程度。我认为这是真的。

So, the billion-dollar startup—I think about this a lot because what I'm doing is not venture scale and not super high leverage. But just seeing how many support tickets I get from the most ridiculous things, it's hard for me to imagine one person scaling to a billion dollars. I'm bearish on this billion-dollar startup. Even if AI is helping you, at a billion dollars, unless your ACVs are very high and you have very few customers, dealing with support is hard to scale. People can solve their own problems, but they'll still email support. Unless you have a bunch of contractors—which I don't know if that counts as a single-person company—it's very difficult to scale a billion-dollar startup without someone helping with support. AI will only take you so far. I think that's true.

软件与分发的未来 Future of software and distribution

Sherwin Wu

实际上,我的看法略有不同。我认为 Lenny 的播客可能会成为一家价值十亿美元的初创公司。但我觉得可能发生的情况是,不再是你一个人需要派遣 AI 来解决和修复那些支持工单,而是会出现一大批其他初创公司,它们构建的软件超级贴合你的需求。所以可能会有 10 到 20 家初创公司为播客和新闻通讯构建支持软件,而且那可能是一家一人初创公司,不需要很大规模。他们可能很容易就能编写出这个产品。他们可以构建自己的东西,因为它是如此定制化、独特且希望对你有用,所以你会购买它,作为那个一人十亿美元初创公司。

And actually I think my view on it is slightly different. I think that Lenny's podcast might end up becoming a billion-dollar startup. But what I think might happen is instead of you being the one person who has to dispatch an AI to solve and fix those support tickets, there might be a whole smattering of other startups that are building software super tailored towards what you might need. So there might be 10 or 20 startups that build support software for podcasts and newsletters, and that might be a one-person startup. It doesn't need to be a big one. They might be able to just code up this product very easily. They can build their own thing, and because it's so tailored and unique and hopefully useful for you, it might be something you purchase as the one-person billion-dollar startup.

Host

我会买的。我会买的。

I would buy that. I would buy that.

Sherwin Wu

是的。这里有一个问题,哪些事情内部做,哪些外包。我认为可能发生的情况是,由于编写软件和构建产品的成本大幅下降,你可能会外包很多工作,从而缩小公司规模。这就是我认为可能出现的世界。同样,这里的结果存在很大的不确定性,但最终结果仍然可能是一个人驱动这家高杠杆公司,并可能真正达到十亿美元。

Yeah. There's a question of what you in-house and what you outsource. What I think might happen is because the cost of writing software and building products is collapsing so much, you might end up outsourcing a lot of this and in doing so reducing the size of your company. So that's the world that I think might happen. Again, there's high uncertainty in what might play out here, but the end result still might be one person driving this highly leveraged company that might actually reach a billion dollars.

Host

我能理解。我也想到了 Clawbots、Moldbot、Openclaw 的 Peter,他现在被所有这些请求、邮件、消息、私信和公关淹没。我很好奇,而且他还没从这件事上赚到任何钱。是的,我无法想象他现在是什么感受。那一定非常疯狂。可能就像我们推出 Hatchvt 后的那几个月,那种疯狂,而且只有他一个人。

I could see that. I also think about Peter at Clawbots, Moldbot, Openclaw, just how barraged he is right now by all these asks and emails and pings and DMs and PRs. I'm curious, and he's not even making any money off this thing. Yeah, I can't imagine what it's like to be him right now. It must be absolutely insane. It's probably like the months after we launched Hatchvt, the craziness that was, as one man.

Sherwin Wu

顺便说一下,他一周后会来播客。

He's coming out on the pod by the way in a week.

Host

哦,那太令人兴奋了。是的。也许第四阶效应是分发变得越来越重要,因为有太多该死的东西在争夺你的注意力。所以我认为拥有受众和平台的人会变得越来越有价值,这是好事。好的,我想回到你的管理话题。我真的很喜欢你关于花更多时间与高绩效者相处对你非常成功的见解。想想你作为一个团队的经理,这个团队正在构建为整个 AI 经济提供动力的平台,每个 AI 初创公司都在你的 API 上构建。显然你做得很好。你还学到了哪些核心管理经验?你认为作为工程师和人员的管理者,什么对你来说真正重要且关键?

Oh, that's exciting. Yeah. Maybe the fourth order effect is distribution becomes increasingly important because there are so many freaking things trying to get your attention. So people with an audience and platform I think become more and more valuable, which is good stuff. Okay, I wanted to come back actually to your management stuff. So I really loved your insight about spending more time with top performers has been really successful for you. Just thinking about you as a manager of a team that is building the platform that powers basically the entire AI economy, every AI startup is building on your API. Clearly you're doing a great job. What other kind of core management lessons have you learned? What do you find is really important and key to your success as a manager of engineers and just people?

Sherwin Wu

是的。我认为我在这里学到的很多经验,我不知道它们是否特别针对 OpenAI API 或我们的一些企业产品。我认为我的管理理念显然随着时间的推移发生了变化,但我觉得它保持不变的部分多于变化的部分。其中一个原则就是我之前和你谈到的,那就是花大量时间与高绩效者相处。具体来说,就是把超过 50% 的时间花在你的高绩效者身上,可能是前 10% 的绩效者,并尽最大努力赋能他们。我思考这个问题的方式回到了软件工程师作为外科医生的类比,这个类比来自《人月神话》这本书。有趣的是,我从书中引用了它,但在书中他们描述了一个预测未来的世界,因为这本书写于 70 年代左右。他们说软件工程最终可能会进入一个软件工程师像外科医生的世界。在手术室里,有一个人在做工作,切割或其他什么,房间里的其他人都在那里支持他们,比如护士、住院医师和研究员。外科医生说“我需要手术刀”,他们就递上手术刀,然后“我需要这个工具和这台机器”,他们就拿过来。每个人都在那里支持那一位外科医生。《人月神话》实际上预测了这就是软件工程的发展方向。我不认为这完全实现了,因为现在更加协作,不是只有一个人在做工作,但我一直很喜欢这个类比。这个类比实际上是我在自己的管理理念中努力效仿的。软件工程并不真的像手术那样只有一个人工作,但我对待团队成员的方式以及我作为经理的行为方式是,我想赋能他们,让他们感觉自己像外科医生,确保我支持他们,确保他们拥有完成工作所需的一切。他们感觉有一支军队在支持他们,在拐角处观察并给他们所需的一切,而实际上只有我作为经理。我举的例子是,从组织角度观察并解除人们的障碍非常有用。再次回到 AI 的讨论,这在当今更加重要。如果人们只是不断地提交 PR,那么阻碍进展和交付产品的主要因素往往是组织或流程导向的。如果你作为经理能够观察并解除团队的障碍,如果外科医生需要手术刀,而经理已经为他们准备好了手术刀,那就是最好的情况。这就是我处理管理问题的方式,尤其是工程管理。随着时间的推移,这一点一直深深印在我心里。尽管软件工程师并不完全是外科医生,但这个比喻在我的职业生涯中一直留在我的脑海中。

Yeah. I think a lot of the lessons I've learned here, I don't know how specific it is to the OpenAI API or some of our enterprise products in particular. I think my management philosophy has obviously changed over time, but I think it's probably stayed the same more than it's changed over time. One of these principles is what I talked to you about before, which is spending a lot of time with top performers. To be very concrete, it's more than 50% of your time with your top performers, maybe your top 10% performers, and really trying your best to empower them. The way I think about it comes back to this analogy of software engineer as a surgeon, which comes from the Mythical Man-Month book. It's funny, I pull it from the book, but in the book they described this world where they were predicting the future because the book was written in the 70s or something. They said that software engineering might end up moving into a world where software engineers are like surgeons. In a surgery room, there's one person doing the work, cutting or whatever, and everyone else in the room is there to support them, like the nurse, the resident, and the fellow. The surgeon says 'I need a scalpel' and they give them a scalpel, then 'I need this tool and this machine' and they bring it over. Everyone is there to support the one surgeon. The Mythical Man-Month actually predicted that is the direction software engineering is going to go. I don't think that's exactly played out where it's much more collaborative and not only one person doing the work, but I've always really liked that analogy. That analogy is actually what I strive to emulate in my own management philosophy. Software engineering isn't really like surgery where it's just one person doing work, but the way I like treating the people on my team and the way I act as a manager is I want to empower them, make them feel like they are a surgeon, in so far as making sure that I'm supporting them and making sure they have everything they need to do their work. It feels like they have an army of people supporting them, looking around corners and giving them everything they need, when it's really just me as the manager. The example I give is looking around corners and unblocking people, especially from an organizational perspective, is extremely useful. Again, going back to the AI conversations, it's even more important nowadays. If people are just cranking PR after PR, the main thing bottlenecking progress and shipping something tends to be organizational or process-oriented. If you as a manager can look around corners and unblock the team, if the surgeon needs a scalpel but the manager already has a scalpel ready for them, that's the best case scenario. That's the way I approach management, especially engineering management. That's something that has really stuck with me over time. Even though software engineers aren't exactly surgeons, that metaphor has always stayed in my mind for the rest of my career.

Host

我喜欢这个。

I love that.

用AI预测障碍 AI for predicting blockers

Host

我在想,这是否是 AI 可以帮忙的地方——预见拐点,预测这位工程师会因为这个决定而受阻。我们需要搞清楚这一点。我们需要去

And I feel like I wonder if that's something AI can help with is look around corners and predict here this engineer is going to be blocked by this decision. We need to figure this out. We need to get

Sherwin Wu

对,这确实是个很好的观点。我还没试过,但我在想,如果我问 ChatGPT 并接入公司知识库,比如“当前有哪些阻塞项?”——让它翻遍所有 Notion 文档,还有 Slack 消息,可能就在 Slack 的某个地方。我团队当前有哪些阻塞项?有没有什么我能帮忙的?嗯,我之前真没想过这个,但你说得对。

Yeah, that's actually a really good point. I haven't tried this yet, but I wonder what would happen if I ask ChatGPT hooked up to company knowledge, you know, like what are the active blockers? Look through all the Notion docs, what are maybe Slack messages, you know, it's probably in Slack somewhere. What are the active blockers on my team and is there something I can do to help? Now very, I have not thought about that, but you're right.

Host

你刚刚在这里有了一个洞见。

You just had an insight right here.

Sherwin Wu

对对对。

Yeah. Yeah. Yeah.

Host

而且我觉得更有趣的是,你预计这位工程师或这个团队在未来几个月会遇到什么阻塞?

And I think even more interestingly, what do you anticipate will be a blocker for this engineer or this team in the coming months?

Sherwin Wu

对。你让模型——你让 AI 去做二阶和三阶的事情。预见那个,伙计。也预见下个月的阻塞项。

Yeah. You asked the model. Well, you asked the AI to do the second and third order things. Anticipate that, man. Anticipate what the blockers will be next month, too.

Host

我觉得我们这里有了个好主意。

I think we've got a good idea right here.

Sherwin Wu

对对。

Yeah. Yeah.

赞助消息:DataDog Sponsor message: DataDog

Host

本期节目由 DataDog 赞助,DataDog 现已拥有 EPO,领先的实验和功能标志平台。全球最优秀公司的产品经理使用 DataDog——这个他们的工程师每天依赖的平台——将产品洞察与产品问题(如 Bug、UX 摩擦和业务影响)联系起来。从产品分析开始,产品经理可以观看回放、审查漏斗、深入留存率并探索增长指标。在其他工具止步的地方,DataDog 走得更远。它帮助你真正诊断漏斗流失、Bug 和 UX 摩擦的影响。一旦你知道重点在哪里,实验就能证明什么有效。我在 Airbnb 时亲身经历过,我们的实验平台对于分析什么有效、哪里出了问题至关重要。而构建 Airbnb 实验的同一团队在 DataDog 构建了 EPO。然后,通过会话回放让你超越数字。精确观察用户如何与热图和滚动图交互,真正理解他们的行为。所有这些都由与实时数据关联的功能标志驱动,让你安全发布、精准定位、持续学习。DataDog 不仅仅是工程指标。它是优秀产品团队更快学习、更聪明修复、自信交付的地方。在 datadog.com/lenny 请求演示。就是 datadog.com/lenny。

This episode is brought to you by DataDog, now home to EPO, the leading experimentation and feature flagging platform. Product managers at the world's best companies use DataDog, the same platform their engineers rely on every day to connect product insights to product issues like bugs, UX friction, and business impact. It starts with product analytics where PMs can watch replays, review funnels, dive into retention, and explore their growth metrics. Where other tools stop, DataDog goes even further. It helps you actually diagnose the impact of funnel drop-offs and bugs and UX friction. Once you know where to focus, experiments prove what works. I saw this firsthand when I was at Airbnb, where our experimentation platform was critical for analyzing what worked and where things went wrong. And the same team that built experimentation at Airbnb built EPO at DataDog. Then lets you go beyond the numbers with session replay. Watch exactly how users interact with heat maps and scroll maps to truly understand their behavior. And all of this is powered by feature flags that are tied to real-time data so that you can roll out safely, target precisely, and learn continuously. DataDog is more than engineering metrics. It's where great product teams learn faster, fix smarter, and ship with confidence. Request a demo at datadog.com/lenny. That's datadog.com/lenny.

AI部署的负投资回报率 Negative ROI on AI deployments

Host

好了,我要转向谈谈你们构建的 API 和平台。你与很多公司合作,让他们实现你们的 API、平台,基于你们的工具进行构建。你告诉我,你发现很多公司实际上在 AI 部署上的 ROI 是负的,我觉得这正是很多人读到、感受到和想到的,有趣的是你确实看到了这一点。这到底是怎么回事?他们做错了什么?AI 部署和 ROI 的世界里发生了什么?

Okay, I'm going to shift to talking about the API and the platform that you all build. So, you work with a lot of companies implementing your API, your platform, building on your tools. You told me that you find that a lot of companies actually have negative ROI on their AI deployments, which I think is what a lot of people read about and feel and think and it's interesting you're actually seeing that. What's going on there? What are they doing wrong? What's happening in the world of AI deployments and ROI?

Sherwin Wu

是的。澄清一下,我并没有明确看到这方面的量化数字。你知道,这些东西其实很难衡量,但尤其是通过观察一些尝试做 AI 的公司,如果很多 AI 部署实际上是负 ROI,我一点也不会惊讶。我的意思是,部分原因也是我认为全国范围内,基本上科技圈以外的人,有一种普遍情绪,觉得 AI 是被强加给他们的。我认为这可能是某些负 ROI 的 AI 部署的症状。我观察到了几件事。其中一件是,我觉得我们硅谷的人就是忘记了自己生活在泡沫里。就像,Twitter 是个泡沫,抱歉,X 是个泡沫。硅谷是个泡沫。软件工程是个泡沫。世界上大多数人,美国大多数人,都不是软件工程师,不是深度 AI 用户,不会关注每一个模型发布。所以我们完全不了解如何使用这项技术。所以我们总是谈论 Codex 的所有这些最佳实践,OpenAI 内部所有那些用 Codex 构建的人。我敢肯定,X 上每个发帖的人都是这些 AI 工具的疯狂高级用户,他们精通技能,精通智能体,MCP。

Yeah. So, to be clear, I don't explicitly see quantitative numbers around this. You know, it's actually really hard to measure these things, but especially from observing some companies kind of trying to do AI, I would not be surprised if a lot of AI deployments are actually negative ROI. I mean part of this too is I think there's also general sentiment from folks around the country, basically outside of tech, that AI is being forced onto them. And I think part of this is probably a symptom of some negative ROI AI deployments. A couple things I've observed around this. So one thing is, and I think I come back to this again and again, I think we in Silicon Valley just forget that we live in a bubble. Like we are so like Twitter is a bubble, sorry X is a bubble. Silicon Valley is a bubble. Software engineering is a bubble. Most people in the world, most people in the US are not software engineers, are not very AI plugged, are not following every single model release. And so we're just highly out of the loop on how to use this technology. And so we always talk about all these best practices for Codex, all these Codex build people within OpenAI. I'm sure everyone on X who posts are crazy power users of these AI tools, you know, they lean into skills, they lean into agents, MCPs.

Host

MCP。

MCPs.

Sherwin Wu

对。是的。所有那些。当我与其中一些公司交谈,并与实际使用这些工具的员工交谈时,他们试图做的都是最基本的事情,而且他们对这项技术的工作原理知之甚少。所以这对我来说是一个重要的观察,就是他们问这些工具的问题非常简单。他们还没有真正推动它。所以这与我所说的更多公司可以做什么,或者更理想的 AI 部署设置是什么样的有关。这也是我们在 OpenAI 内部运作的方式。那些我认为开始运行得很好的公司,既有自上而下的支持。比如高层说,我们要成为一家 AI 优先的公司。所以有支持,他们购买工具,有高管支持,但同时也有自下而上的采用和支持。我的意思是,有实际做工作的员工,他们对这项技术非常兴奋,愿意学习、宣传、建立最佳实践,并在组织内部分享知识。我们在内部见过很多这样的情况。显然,OpenAI 一直想成为一家非常以 AI 为中心的公司,但真正开始起飞是在 Codex 和这些工具推出之后,实际员工自己可以开始将其应用到工作中。我认为这非常必要,因为归根结底,每个人的工作都非常不同。非常独特。软件工程不同于财务,不同于运营,不同于市场推广和销售。所以有很多最后一公里的工作细节,需要以自下而上的方式完成。所以我的感觉是,很多这些 AI 部署没有自下而上的采用。就像是一个高管的指令,完全是自上而下的,与实际工作脱节。最终结果就是,你有一个庞大的员工队伍,他们并不真正理解这项技术。就像,我知道我应该用这个,可能它也在我的绩效评估里,但我不确定该怎么做。

Yes. Yeah. All of that. And when I talk to some of these companies and I talk to the actual employees using these, it's like the most basic thing that they're trying to do and they have very little understanding of exactly how this technology works. And so that's kind of like one big observation for me, which is they're asking very simple questions of these things. They're really not pushing it just yet. And so that kind of ties into what I think more companies could do or what a more ideal AI deployment setup looks like. And this is kind of how we've run things within OpenAI too. The companies where I think it started to work really well have a combination of both top-down buy-in. So it's like the suite, it's like we want to become an AI-first company. So there's buy-in, they buy the tools, they have exec support, but it also has bottoms-up adoption and buy-in. And what I mean by that is it has actual employees doing the work who are really excited about the technology and are willing to learn, evangelize, build best practices and kind of knowledge share within the organization. We've seen this a lot internally. So obviously OpenAI has always wanted to be a very AI-centric company, but where it really started taking off was with the introduction of Codex and these tools where people, like actual employees themselves, could start applying it to their work. And I think you really need this because at the end of the day everyone's work is very different. It's very unique. Software engineering is different than finance, different than operations, different than go-to-market and sales. And so there's a lot of these last-mile intricacies of work that needs to really be done in a bottoms-up fashion. And so my sense is a lot of these AI deployments don't have bottoms-up adoption. Like it was an exec mandate and it's extremely top-down and is very divorced from what the actual work looks like. And as an end result, you end up with a giant workforce that doesn't really understand the technology. It's like, I know I'm supposed to use this and maybe it's on my performance review too, but I'm not sure what to do.

组建内部AI精英团队 Building an internal AI tiger team

Host

他们环顾四周,发现没有其他人在做这件事,也没有人可以学习。所以我对正在推动这件事的公司的建议是,在内部组建一个全职的“老虎团队”,探索能力的全部边界,应用到具体工作流中,做知识分享,在可能想用这项技术的人群中激发热情。因为如果没有这样的团队,就很难真正落地。

And they look around, no one else is doing it. There's no one else to learn from. And so my recommendation for companies kind of pushing this is find or maybe even staff a full-time team internally that is this kind of tiger team internally that can explore the full extent of the capabilities, apply to specific workflows, do the knowledge sharing, create excitement within folks who might want to use this technology. Because in the absence of that, it's very difficult to pick up.

Host

那你会把谁放进这个老虎团队?是由工程师主导吗?根据你的经验,它是一个跨职能的团队吗?

And who would you put on this tiger team? Is it engineer-led? Do you find in your experience it's a cross-functional sort of team?

Sherwin Wu

嗯,这很有意思。很多公司其实没有软件工程师。我看到的模式是,这些人往往是软件工程相关岗位——基本上是技术人员,但不是软件工程师。我认为正是这些人对此最兴奋。比如,支持团队或运营主管,他们不写代码,但喜欢用这些工具,像个 Excel 高手之类的。所以他们是技术相关或编码相关的人,相当懂技术。我在这些公司里看到的就是这类人,他们真的会眼前一亮,对此充满热情。你通常可以围绕他们组建一个团队。但确实,往往不是软件工程师。软件工程师当然也能理解这个,但不是每家公司都有软件工程师。实际上软件工程师挺稀缺的,难找又贵。所以是这些其他类型的人。

Yeah, it's interesting. So, also a lot of companies don't have software engineers. And so the pattern I've seen is it tends to be these software engineering adjacent, basically technical people but are not software engineers. I think those are the ones who get most excited around this. It's like, maybe the support team operations lead who doesn't code but loves using these tools and is like an Excel wizard or something. So it's technical adjacent or coding adjacent and pretty technical. Those are the kinds of people I've seen in these companies who just really light up and get excited around this. And you can usually build a team around that. But yeah, it's often not software engineers. Software engineers, I think, will understand this, but not every company has software engineers. It's actually kind of a rarity. They're hard to find. They're expensive. And so it's these other types of folks.

Host

我听到的反面模式是自上而下。CEO 和执行团队说“我们要 AI 优先,我们要引领 AI,每个人都要根据使用 AI 工具的表现、生产力提升多少来被评判”。如果只是自上而下,而不建立一个自下而上传播理念的团队,你会发现行不通。

What I'm hearing is the anti-pattern is top-down. This is the CEO found exec team just like we are going to go AI first. We're going to lead into AI. Everyone's going to be judged on their performance using AI tools, how much your productivity is increasing thanks to AI. And without that being just top-down and not creating a team that is bottom-up spreading the gospel, you find it doesn't work.

Sherwin Wu

对,完全正确。

Yeah. Exactly. Exactly.

Host

所以建议是:找到最兴奋的人,不是让他们分散在组织各处,而是组建一个小型 AI 布道团队,找到使用方法,然后在整个工作中传播。

And the advice is find the people that are most excited and instead of having them spread out through the organization, what you find works is create a little AI evangelist team that finds ways to use it and spreads it across the work.

Sherwin Wu

是的。另一种思考方式,也呼应我自己的管理哲学,就是找到在 AI 采用方面的高绩效者,并赋予他们权力。让他们组织黑客马拉松、举办研讨会、做知识分享,在内部播下兴奋的种子。

Yeah. I mean, another way to think about it, kind of tying back to my own management philosophies, is find the high performers in AI adoption and empower them. Let them build hackathons, let them hold seminars, do knowledge sharing, kind of create the seeds of excitement internally.

AI中倾听客户可能误导 Listening to customers can be misleading in AI

Host

好的,太棒了。我想听听你的一些犀利观点。我见过你分享过一点:与客户交谈、倾听客户并不总是 AI 领域的正确策略,而且常常会让你误入歧途。

Okay. Amazing. There's a couple hot takes I want to hear from you. Something that I've seen you talk about and share: one is you've shared that talking to customers and listening to customers is not always the right strategy in AI and it might often lead you astray.

Sherwin Wu

我不确定这算不算很犀利的观点。主要问题是,显然你应该和客户交谈,这很有用。但我认为 AI 领域,尤其是过去三年我在 API 工作中看到的一切,领域和模型本身变化太快了,它们往往会自我颠覆,尤其是在工具和脚手架方面。这周早些时候我读到一篇文章,作者是 Nicholas,一家叫 Finol 的初创公司创始人,他在文章中分享了在金融科技初创公司 FinTool 构建 AI 智能体时学到的很多最佳实践。他有一句话我觉得特别好:“模型会把你的脚手架当早餐吃掉。”回想 2022 年 ChatGPT 刚推出时,模型还很粗糙,有大量产品脚手架,尤其是在开发者领域,试图引导模型并在其周围搭建脚手架让它做你想做的事,比如智能体框架、向量存储当时非常流行,还有一大堆工具。随着领域发展,模型变化如此之大、变得如此之好,它们真的吃掉了一些脚手架。我认为今天依然如此。所以 Nicholas 的文章指出,当前流行的脚手架是基于技能文件的上下文管理。我可以预见,在某个时候它不再有用,模型可以自己管理所有这一切,或者会出现新的范式,你不再需要这种基于文件的技能类东西。你已经亲眼看到这一点:智能体框架现在没那么有用了。2023 年有一段时间,我们认为向量存储是将组织上下文引入模型的主要方式,你需要对语料库的每一部分进行向量化和嵌入,然后做大量工作来优化向量搜索,以便在正确的时间提供正确的信息。所有这些都是脚手架,因为模型不够好。事实证明,随着模型变得更好,更好的方法是去掉很多逻辑,信任模型,给它一套搜索工具。它不需要向量存储。你可以直接把它连接到任何类型的搜索。它可以是文件系统上的文件,比如技能和 agents MD 来引导它。显然,向量存储仍有其用武之地,我知道很多公司还在用,但围绕它的整个脚手架、构建整个生态系统并假设那是你唯一需要的脚手架,这种情况已经彻底改变了。

I don't know if it's that hot of a take. I think the main thing here is, obviously you should talk to your customers, it's useful to talk to customers. I just think the AI field, especially what I've seen over the last three years working on the API and seeing all that evolve, is the field and the models themselves are just changing so quickly they tend to disrupt themselves, especially around the tooling and scaffolding space. So there's this quote that I read earlier this week from an article by this guy named Nicholas who's the founder of a startup called Finol, where he was sharing a lot of the best practices he has learned through building AI agents for financial services at a startup called FinTool. And he had this phrase that I thought was really good: 'the models will eat your scaffolding for breakfast.' If you look back to 2022 right when ChatGPT launched, these models were pretty raw and there was all this product scaffolding and things, especially in the developer space, to basically try and steer the model and build a scaffolding around it to get it to do what you want, like agent frameworks, vector stores were really popular back then, and just a whole smattering of tools. And as you've seen the field play out, the models have just changed so much and gotten so much better that they ended up literally eating some of the scaffolding. And I think this is even true today. So the article from Nicholas, the current scaffolding which is fashionable is skills files based context management. I could see a world where at some point that's no longer useful, where the model can actually manage all that themselves, or there might be some new paradigm where you need this file-based skills type thing. You have literally seen this play out: the agent frameworks are a little less useful now. There was a period of time in 2023 where we thought vector stores were going to be the main way for you to bring organizational context into the models, and you need to vector and embed every bit of your corpora and then do all this work to figure out the vector search to optimize that to fill out the right information at the right time. All of that is scaffolding because the model was not good enough. And it turns out as the models get better, a better approach is actually to take out a lot of that logic and trust the model and give it a set of tools for search. It doesn't need to be a vector store. You could actually just hook it up to any type of search. It could literally be files on a file system like skills and agents MD to steer it as well. Obviously, there's still a place for vector stores. I know a lot of companies are still using it, but the entire scaffolding around that and building an entire ecosystem around that and assuming that's the only scaffolding that you need has really changed.

倾听客户vs模型轨迹 Listening to customers vs. model trajectory

Sherwin Wu

所以回到这一点,你不必总是听从客户的意见,因为这个领域变化太快了。在任何时候,很多人都处于局部最优中。如果你盲目听从客户,他们会说,“是的,我想要一个更好的向量数据库,我想要一个更好的智能体框架。”如果你只沿着那条路走,最终会构建出另一个局部最优。而随着模型变得更好,我们不得不重新发明和重新思考正确的抽象层以及围绕这些模型构建的正确工具和框架。酷、令人兴奋又有点疯狂烦人的是,这是一个移动的目标。当前的一堆工具和框架很可能需要随着模型变得更智能、更好而显著演变和改变。但这正是这个领域构建的本质。我认为这正是它令人兴奋的地方。但这也意味着,当你与客户交谈时,你需要平衡他们给出的具体反馈与你对模型未来走向以及未来一到两年趋势的判断。

And so tying this back to like you know, you don't always have to listen to your customers because the field is changing so much. At any point in time, a lot of people are kind of in a local maximum. If you just blindly listen to your customers, they'll be like, 'Yeah, I want a better vector store, I want a better agent framework for this.' And if you had just chased down that path, it would have led you to build something that is again the local maxima. Whereas as the models get better, we've had to reinvent and rethink the right abstractions and the right tools and frameworks to build around these models. The cool, exciting, and kind of crazy annoying part is it's a moving target. The current smattering of tools and frameworks will likely need to evolve and change pretty significantly over time as the models get smarter and better. But that is just the nature of building in this space. I think that's what makes it exciting. But it also means when you talk to customers, you need to balance the exact feedback they want with where you think the models are going and where things will trend over the next one to two years.

Host

有趣的是,这就是苦涩的教训,AI 和 ML 从业者学到的一个重要教训:你越不复杂化,你给机器学习和 AI 添加的逻辑越少,它就越能扩展和成长。干脆全部拿走,让它自己计算,基本上给它更多力量让它自己变得更聪明。

It's interesting how this is the bitter lesson, a big lesson that AI and ML folks learned: the less you overcomplicate, the less logic you add to machine learning and AI, the more it'll be able to scale and grow. Just take it all away and let it compute, basically give it more power to get smarter on its own.

Sherwin Wu

是的,确实有一个版本的苦涩教训适用于用 AI 构建。我们试图围绕它架构所有这些东西,结果模型把它们都吞噬了。老实说,OpenAI API 团队也犯过这个错,我们本不该左转右转的时候转了。但模型仍然变得更好,我们日复一日地学习着这个苦涩的教训。

Yeah, there's literally a version of the bitter lesson applied to building with AI. We were trying to architect all this stuff around, and it turns out the models just kind of eat it all away. Honestly, the OpenAI API team has been guilty of this, where we took some left and right turns when we shouldn't have. But the models still get better, and we're all learning the bitter lesson day in and day out.

Host

那么,对于在 API 上构建或构建智能体、目前不得不围绕它构建一些东西的人来说,关键要点是什么?有什么建议?

So what would be the key takeaway for folks building on the API or building agents, having to build a little bit of this around for now? What would be the advice?

Sherwin Wu

我的一般建议,我已经给人们提了很久,而且我认为今天仍然适用:确保你为模型将要去的地方构建,而不是它们今天所在的地方。这显然是一个移动的目标。我看到很多做得好的公司和初创公司,都是为一个理想的能力类型构建产品,这个能力今天可能已经完成了 80%。他们最终拥有一个勉强能用但几乎就快成的产品。然后随着模型变得更好,突然之间它可能就成功了,他们的产品变得不可思议,因为它能工作了。例如,在 o3 的某个点上它突然能工作了,在 5.1、5.2 时它突然解锁了。但他们构建这些产品时考虑到了模型能力的提升,这创造了一种比假设模型静止要好得多的体验。所以这就是我的一般建议:为模型将要去的地方构建,而不是它们今天所在的地方。你最终会构建出更好的产品。你可能需要等一小会儿,但模型进步如此之快,你通常不需要等太久。

My general advice, and I've been giving this to people for a while and I think it's still true today, is make sure you're building for where the models are going, not where they are today. It's clearly a moving target. I think a lot of the companies and startups that I've seen really do well build a product for an ideal type of capability that is maybe 80% of the way there today. They end up having a product that kind of works but is just almost there. Then as the models get better, suddenly it might click, and their product becomes incredible because it works. For example, with o3 at some point it suddenly works, with 5.1, 5.2 it suddenly unlocks. But they build these products with model capability improvements in mind, and that creates an experience way better than if they had assumed it was static. So that would be my general advice: build for where the models are going, not where they are today. You end up building a better product. You may need to wait a little bit, but the models are getting so much better so quickly that you often don't need to wait that long.

未来6-12个月的API、平台与模型 Future of API, platform, and models in 6-12 months

Host

顺着这个话题,未来六到十二个月,API 会走向何方?平台会走向何方?模型会走向何方?你能分享多少就分享多少,我知道这里有很多秘密。也许是你最兴奋的事情,或者人们应该开始准备的事情?

So to follow that thread, where are in the next six to 12 months the API heading? Where's the platform heading? Where are the models heading? As much as you can share, I know there's a lot of secrets here. Maybe what you're most excited about or what people should start to prepare for?

Sherwin Wu

一个明显的方向是这些模型能连贯执行多长的任务。有一个 SWE-bench 基准测试,追踪软件工程任务以及这些模型在 50% 和 80% 的情况下能完成多长的任务。我认为目前前沿模型在 50% 的情况下能完成数小时的任务,80% 的情况下能完成接近一小时的任务。但那个图表令人警醒的是,它同时绘制了所有之前的模型,所以你能真正看到趋势。这是我非常兴奋的事情。今天的产品真正优化的是模型一次能完成几分钟的任务。即使是 codex 和编码工具,也最多优化到 10 分钟级别的任务。我见过有人把 codex 推到极限,完成数小时的任务,但那更多是例外。如果你跟随这个趋势,在未来 12 到 18 个月,我们可能会看到模型能非常连贯地完成数小时的任务。在某个点上,它可能达到 6 小时的任务,你派发它,让它自己运行一段时间。围绕它构建的产品类型将非常不同。你想给模型反馈;你显然不希望它完全失控一整天。也许你想,但很可能不想。模型能做的事情的范围将真正扩大。所以这是我非常兴奋的事情。

The obvious one is how long of a task these models can do coherently. There's the SWE-bench benchmark that tracks software engineering tasks and how long of a task these models can do 50% of the time, 80% of the time. I think we're at something like multi-hour tasks being able to be done by these frontier models 50% of the time, and 80% is something like just under an hour. But the sobering thing about that chart is they plot all the previous models on it as well, so you can really see the trend. That's something I'm really excited about. Products today really optimize for tasks that the model can do for minutes at a time. Even Codex and coding tools are quite optimized for maybe at most 10-minute types. I have seen people push codex to the limit and do multi-hour long tasks, but that's more the exception. If you follow this trend, in the next 12 to 18 months we could see models that could do multi-hour long tasks very coherently. At some point it might reach a 6-hour long task where you dispatch it and have it do things on its own for a while. The types of products you build around that will look very different. You want to give the model feedback; you obviously don't want it to completely run wild for a day. Maybe you do, but you probably don't. The universe of things you can have the model do will really expand. So that's something I'm really excited about.

Sherwin Wu

另一个在未来 12 到 18 个月我认为会很酷的事情是多模态模型的改进。说到多模态,我主要想到的是音频。模型在音频方面已经相当不错了,但我认为在未来 6 到 12 个月,它们会在音频方面变得更好,尤其是原生多模态模型,即语音到语音的模型。在多模态音频方面,也有关于新型模型和架构的有趣工作正在进行。音频,尤其是在企业和商业环境中,仍然是一个被严重低估的领域。每个人都在谈论编码,全是文本,但我们是在用音频交谈。世界上很多商业活动都是通过音频完成的。很多服务和运营都是通过对话和音频进行的。

Another thing over the next 12 to 18 months I think will be really cool is improvements in multimodal models. By multimodality, I'm mostly thinking about audio here. The models are pretty good at audio, but I think they're going to get a lot better at audio over the next 6 to 12 months, especially the native multimodal models, the speech-to-speech ones. There's also interesting work being done around new types of models and architectures on the multimodal audio side. Audio, especially in the enterprise and business setting, is a hugely underrated domain still. Everyone talks about coding, it's all text, but we're talking in audio. A lot of the world's business is done via audio. A lot of services and operations are done via talking and audio.

AI智能体与音频模型的未来 Future of AI agents and audio models

Sherwin Wu

所以我认为,未来 12 到 18 个月,这个领域会非常令人兴奋。而且我认为音频模型能做的事情还会有更多突破。

And so I think that area is going to look very exciting in the next 12 to 18 months. And I think there will be even more unlock for what we can do with audio models as well.

Host

太棒了。快速总结一下:预计智能体和 AI 工具会运行更长时间,这一趋势会持续增长;然后音频和语音会变得更加重要,更原生、更好,成为体验的核心。

Amazing. So quick summary: expect agents and AI tools to run longer, that trajectory to continue to increase, and then audio and speech becoming a bigger deal, more first-party and native and better and core to the experience.

Sherwin Wu

是的。

Yeah.

Host

非常酷。好,我想回到你之前的一个观点,另一个我见你讨论过的热门观点。你对业务流程自动化作为 AI 领域的一个机会非常看好。谈谈这个吧。

Extremely cool. Okay, I want to go back to one of your hot takes, another hot take that I've seen you discuss. You're very bullish on business process automation as an opportunity in the world of AI. Talk about that.

Sherwin Wu

是的,这又回到我之前说的:我们生活在硅谷的泡沫里,我们习惯做的很多工作——软件工程、产品管理、构建产品——与支撑整个经济的工作形态截然不同。我在与客户交流时反复看到这一点。如果你和任何非科技公司聊聊,就会发现大量业务流程。我通常这样区分:软件工程是开放式的知识工作,对吧?它需要探索,你给它这些开放式的东西。但软件工程本质上非常开放,不太可重复。你构建一个功能,不会反复构建完全相同的功能。很多科技岗位都属于这类——数据科学,甚至一些战略财务工作。但当你远离软件工程和科技核心时,很多工作就只是业务流程。它们是可重复的操作,由公司某个经理迭代而来。通常有标准操作流程,人们不希望偏离太多。在软件工程中,独创性不在于偏离,但世界上大量工作实际上只是执行这些流程和操作。如果我打电话给客服,他们就在执行其中之一。如果我打电话给公用事业公司,他们有一系列流程,能做什么不能做什么。所以我非常看好这个类别,而且我认为它被低估了,因为它与硅谷的思维差异太大,人们往往不去想它。但我们如何将 AI 和我们已有的工具框架应用于业务流程自动化,用于自动化并简化高确定性的可重复业务流程,并与企业数据、业务决策和不同系统完全集成?我们如何真正改进那个流程?因为我确实认为这个领域有大量机会和大量工作要做,而我们只是不谈论它,因为它不太在我们的专长范围内。

Yeah, this goes back to the thing that I said previously, which is we live in a bubble in Silicon Valley, and a lot of the work that we do that we're used to—software engineering, product management, building products—is very differently shaped than the work that goes on that runs our entire economy. I see this in and out when I talk to customers. If you talk to any company that's not a tech company, there's a lot of business processes. So what I mean by this is I generally delineate it as software engineering is kind of open-ended knowledge work, right? It's exploring, and you're giving it these open-ended things. But software engineering is fundamentally pretty open-ended and not very repeatable. You build a feature, you're not trying to build the exact same feature over and over again. A lot of tech jobs are in this space—data science, even some strategic finance stuff. But as you move further away from software engineering and what is core in tech, a lot of jobs are just business processes. They're repeatable operations that some manager at a company has iterated on. There's usually a standard operating procedure that people want to follow, and you don't want to deviate from it that much. In software engineering, the ingenuity isn't deviating, but a lot of the work being done in the world is actually just running through these procedures and operations. If I call a support line, they're running through one of these. If I call my utility company, there's a bunch of processes and things they can and cannot do for me. So I'm extremely bullish on this general category, and I think it's underrated because it's so different from what we think about in Silicon Valley. People tend to not think about it. But how can we apply AI and some of the tools and frameworks we have towards this business process automation, towards automating and making easier repeatable business processes with high determinism, fully integrated with business data and business decisions and different systems within an enterprise? How can we actually make that process better? Because I actually think there's a lot of opportunity and a lot of work to be done in that area, and we just don't talk about it because it's a little bit less in our wheelhouse.

Host

所以你的观点,为了确保我完全理解,是你认为在工程之外,AI 有更大的机会来影响公司的生产力,以及那些从事这些重复性、易自动化任务的人的工作,影响工作方式。很多工作都是以这种方式完成的。想想看——我经常和客户、大型企业交流,比如 AI 将如何改变我的公司,在 20 年后的 AI 世界里它会如何运作。软件工程是故事的一部分,但业务流程方面要多得多。而且我实际上认为业务流程方面看起来会更不同。那里的工作相当可观。这很有趣。我不知道从绝对百分比或绝对基数来看,它是否比软件工程更大或更小。软件也非常庞大和广泛。但它确实非常巨大,而且绝对比你根据人们在 X 或 Twitter 上谈论或不谈论的程度所想象的要大。

So your take here, just to make sure I fully understand it, is you think there's a much bigger opportunity outside of engineering for AI to impact productivity of companies and also jobs of these folks that are doing these kind of repetitive easily automated tasks, impact jobs, and also just impact how work is done. So much of work is done in this way. You think about what basically—I talk to customers all the time, big enterprises, like how will AI transform my company, how will it run in a world with AI in 20 years. Software engineering is part of the story, but there's so much more on the business process side. And I actually think it might look even more different on the business process side. The work there is pretty substantial. It's actually interesting. I don't know from an absolute percentage or absolute basis, I don't know if it's bigger or smaller than software engineering. Software is pretty huge and pretty extensive as well. But it is pretty massive and it's definitely bigger than you would think it is based off of how people talk about it or don't talk about it on X or Twitter.

Sherwin Wu

是的。

Yeah.

Host

好。换个方向:你构建了平台,构建了 API,人们在 API 上构建,大家心里最大的问题总是:如何避免 OpenAI 扼杀我的想法,自己构建,然后毁掉我创造的市场?总体政策是什么?初创公司应该如何思考 OpenAI 不太可能涉足的领域,总体理念是什么?

Okay. Going in a slightly different direction: having built the platform, building the API, people building on the API, the biggest question on people's minds is always just how do I not have OpenAI squash my idea and build their own thing and then destroy this market I created. What's the general policy? What's the general philosophy of how startups should think about where OpenAI is unlikely to go?

Sherwin Wu

我的一般回答是:市场如此巨大。我实际上认为初创公司不应该过度思考 OpenAI 或这些实验室的动向。我和很多失败的初创公司聊过,也和一些非常成功的聊过。我见过的每一个逐渐消亡的初创公司,都不是因为 OpenAI 或大实验室或谷歌来扼杀它们,而是因为它们构建的东西没有引起客户共鸣。而那些成功的,即使在像编程这样竞争激烈的领域,比如 Cursor 现在非常庞大,也是因为它们构建了人们真正喜欢的东西。所以我的一般建议是:不要过度为此焦虑。只要构建人们喜欢的东西,你就会在其中找到空间。我怎么强调现在机会有多大都不为过。用 AI 构建的机会空间如此之大。一个很好的例子是,这个空间如此之大,以至于风投可接受和不可接受的 Overton 窗口已经完全改变。风投正在左右投资竞争性公司。只是因为机会前所未有。虽然这影响了风投的运作方式,但从初创公司的角度来看,这是世界上最赋能的事情,因为即使你只构建了一些人真正喜欢的东西,你最终也会得到一个非常有价值的企业。所以这就是为什么我告诉人们不要过度思考。

My general answer here is the market is so big and so massive. I actually think startups should just not overly think about where OpenAI or these labs are going. I've talked to a lot of startups that have not worked out, startups that are doing really well. Every startup that I've seen that has kind of fizzled out is not because OpenAI or a big lab or Google has come to squash them. It's because they built something and it really didn't resonate with the customers. Whereas the ones that take off, even in very competitive spaces like coding, like Cursor is huge at this point, and it's because they built something that people really love. So my general advice is don't overly stress about this. Just build something that people like, and you will have a space in this. I can't overstate how big of an opportunity there is right now. The opportunity space of building with AI is so big. A good example of this is the space is so big that the Overton window of what is acceptable and not acceptable for VCs to do has completely changed here. VCs are investing in competitive companies left and right. It's just the space is so big because the opportunity is unlike anything that we've seen before. And while that affects how VCs operate, from a startup perspective it's the most empowering thing in the world because even if you just build something that some people really really love, you will end up with a massively valuable business. So that's why I tell people don't overthink about it.

OpenAI作为生态系统平台公司 OpenAI as an ecosystem platform company

Sherwin Wu

另一件我认为很重要的事情是,至少从 OpenAI 的角度来看,我们一直非常珍视的一点——Sam 和 Greg 也从高层不断强调——是我们从根本上将自己视为一家生态系统平台公司。API 是我们的第一个产品。我们认为,培育这个生态系统、持续支持它而不是压制它,对我们来说非常重要。所以如果你看看我们做的决策,这一点贯穿始终。我们发布的每一个模型,只要出现在我们的产品中,也都会出现在 API 里。即使是我们现在发布的这些 Codex 模型,虽然它们对 Codex 测试平台做了更多优化,但最终也都会进入 API,我们所有的客户都能使用它们。我们在这方面没有任何保留。我们认为保持平台中立非常重要,所以我们不会屏蔽竞争对手。我们允许人们访问我们的模型。我们最近还在测试更多像“用 ChatGPT 登录”这样的产品,我们想要培育这个生态系统。我认为这样做非常重要。总的思路是水涨船高。我们可能是一艘航空母舰,现在已经相当庞大了,但我们认为提升水位很重要,因为每个人都会受益,而且我相信我们也会受益。我们的 API 本身就是因为这样的做法而获得了显著增长。所以我真的鼓励大家不要把 OpenAI 看作一个会把别人挤开的东西,而是专注于创造有价值的东西,我们仍然致力于提供一个开放的生态系统。

The other thing I also think is important to remember, at least from an OpenAI perspective, is that we've always held very near and dear, which both Sam and Greg helped reinforce from the top, is we actually view ourselves fundamentally as an ecosystem platform company. The API was our first product. We think it's really important for us to foster this ecosystem and continue to support it and not squash it. So if you look at the decisions we make, this is all we weave through it. Every single model we've released in one of our products gets released in the API. Even we release these Codex models now that are a little bit more optimized for the Codex harness, but they always find their way into the API, and all of our customers end up using those. We don't hold back on any of that. We think it's really important to keep our platform neutral, so we don't block competitors. We allow people to have access to our models. We've recently been testing more of the sign-in with ChatGPT product as well, and we want to foster this ecosystem. I think it's really important that we do so. The general thinking about this is a rising tide lifts all boats. We might be an aircraft carrier, we're pretty big at this point, but we think it's important to raise the tide because everyone benefits, and I think we'll benefit as well. Our API itself has grown pretty significantly because we act in this way. So I'd really encourage people not to view OpenAI as this thing that'll shove people out of the way, but instead focus on building something valuable, and we remain committed to providing an open ecosystem.

Host

为什么这对 OpenAI 如此重要?就是这种专注于构建平台、为人们创造创业方式的做法。这是从一开始就有的愿景吗?

Why is that important to OpenAI? Just this focus on building a platform, creating a way for people to build businesses. Is that just that's been the vision from the beginning?

Sherwin Wu

我们希望这成为一个平台。从一开始这就是愿景。实际上这可以追溯到我们的章程,也就是我们的使命。OpenAI 的使命一直是构建 AGI,所以我们显然在做这件事,但第二件事是让全人类都能享受到它的好处。其中关键的部分是“全人类”。显然 ChatGPT 正在尝试做到这一点,我们正在努力覆盖整个世界。但很早以前——这也是为什么我们在 2020 年左右就推出了 API,非常早——我们就认为,单靠我们一家公司是无法触及全人类的。世界的每一个角落都非常深入。所以我们实际上觉得,为了实现我们的使命,我们需要某种平台式的东西,让其他人能够为播客主和新闻通讯主构建客户支持机器人,因为我们自己做不到。我们在 API 上已经很大程度上看到了这一点。这就是为什么我们与这么多客户交流,并且非常喜欢看到在上面构建的各种各样的东西。但没错,从第一天起就是这样,因为我们把它视为我们使命的一种体现。

We want this to be a platform. It's been the vision from the beginning. It goes back to our charter actually, our mission. The OpenAI mission has always been to build AGI, so we're obviously doing that, but then the second thing is to spread the benefits of it to all of humanity. The main part there is all of humanity. Obviously ChatGPT is trying to do this, we're trying to reach the whole world. But very early on, and this is why we launched the API back in I think 2020 or something, really early, we don't think we as a company will be able to reach all of humanity. Every corner of the world is pretty deep. So we actually feel that in order for us to fulfill our mission, we need to have some platform-style thing where we can empower other people to build the customer support bot for podcasters and newsletter hosts, because we're not going to be able to do it ourselves. We've largely seen this play out with the API. This is why we talk to so many of our customers and really love seeing the diversity of things built on it. But yeah, it's been there since day one because we view it as an expression of our mission.

Host

而且你还没提到你们正在推出的那个应用商店,ChatGPT 应用商店。

And you haven't even mentioned the app store that you guys are launching, the ChatGPT app store.

Sherwin Wu

对。

Yeah.

Host

顺便问一下,那个是在你的管辖范围内,还是属于另一个组织团队?

Is that under your umbrella by the way, or is that a different org team?

Sherwin Wu

那是另一个团队。它属于 ChatGPT 部门。我们当然和他们密切合作,他们构建了一个应用 SDK,这个 SDK 是与我们团队紧密协作开发的。但那更多是在 ChatGPT 的范畴内。不过这也是另一个例子。ChatGPT 有 8 亿周活跃用户,他们反复回来使用。作为一项业务,这是一个巨大的资产,但如果能允许其他公司也进来利用这一点,并为这些用户构建东西,那不是更好吗?最终我们认为这也会帮助我们扩大那个用户群体。所以这一切都回归到使命,我们发现作为一个平台、保持开放,在这方面是有帮助的。

It's a different team. So it's under ChatGPT. We obviously collaborate very closely with them, and they built an apps SDK which is built in close collaboration with our team. But that is more within the ChatGPT umbrella. But that is also another example of this. ChatGPT has these 800 million weekly active users who are just coming over and over again. It's a great asset to have as a business, but would it be better if we could somehow allow other companies to come in and take advantage of this as well and build for this audience? Ultimately we think it'll help us expand that group as well. So it all comes back to the mission, and we find that being a platform, being open, tends to help here.

Host

就那个 8 亿的数字,我想是周活跃用户,每周有 10 亿人在使用。我们现在对这些数字已经习以为常了,但这太疯狂了,史无前例。

Just that number 800 million, I think it's weekly active, a billion people using weekly. It's just absurd how these numbers we're just used to now, but that's insane, unprecedented.

Sherwin Wu

是的,老实说,从规模的角度来看,这对我来说是难以置信的。我的理解是,这相当于全球人口的 10%,而且还在增长,正在飞速上升。

Yeah, it's mind-boggling for me to think about from a scale perspective honestly. The way I think about it is like 10% of the world, and growing by the way, it's shooting up.

Host

人们来使用 ChatGPT,每天,或者抱歉,每周。

Come to ChatGPT and use it every day, or sorry, every week.

Sherwin Wu

关于这一点,我想再强调一下你提出的观点。OpenAI 的使命是让 AI 惠及全人类。我认为有些人对此不以为然。他们会说,哦,这要花钱,但事实是有一个免费版的 ChatGPT,任何人都可以使用,它与世界上存在的最强大的 AI 模型相差无几,而且是免费的,没有门槛,任何人都能用。即使你是亿万富翁,你能从 AI 中获得的东西,也比非洲某个村庄里的人多不了多少。我知道这对 OpenAI 来说一直非常重要。

And this point I just want to double down on this point you're making. OpenAI's mission was to make AI available to all humanity. I think some people dismiss that. They're like, oh, it costs money, and the fact that there's a free version of ChatGPT that anybody can use, that is not so different from the most powerful AI model that exists in the world, for free, that's not gated, that anyone could use. If you're a billionaire, there's only so much more you can get out of AI than what someone in a village in Africa can get. I know that's always been really important to OpenAI.

Sherwin Wu

是的。这就是为什么我认为我们深入参与了健康领域的工作。我们深入参与了教育,这将会非常有趣。另一个疯狂的趋势是,免费模型随着时间的推移变得越来越智能。2022 年的免费模型在当时还不错,但和今天你能得到的相比根本不算什么,因为今天你能得到 GPT-5。所以提高全球的基准水平是我们真正在努力做的事情,我们将其视为使命的一部分。顺便说一句,另一面是关于亿万富翁之类的。我知道人们会说,你用的 iPhone 和马克·扎克伯格可能用的是同一款,但每个月花 20 美元,你基本上就能用上和亿万富翁一样的 AI。

Yeah. That's why I think we've leaned into the health work. We leaned into education, which is going to be very interesting. The other insane kind of trend here is the free model has gotten so smart over time. The free model back in 2022 was good at the time, but it's nothing compared to what you get today because you get GPT-5 today. So raising the floor across the world is something we're really trying to do, and we view it as part of our mission. The other flip side of this, by the way, is talking about the billionaires or whatever. I know people say you're using the same iPhone that Mark Zuckerberg's probably using, but for $20 a month you're basically using the same AI that the billionaires are using.

AI的民主化 Democratization of AI

Host

比如每月 200 美元,你就能用上和所有亿万富翁一样的 Pro 模型,但他们可能不会在所有场景都用 Pro,日常可能只用 Plus 版。所以这种民主化、让这种好处惠及全球的做法,对我们来说意义重大,也是驱动我们做很多事的原因。

Uh for like $200 a month uh you get the same pro model that you know all the billionaires are using but they're probably not using pro for everything. They're probably just using the the plus tier ones uh for their day in and day out. And so yeah, this kind of like democratization and just like spreading of this this benefit like across all of the world is something that's really meaningful to us and something that um uh drives a lot of of of what we do.

API与平台能力 API and Platform Capabilities

Host

最后一个问题,给那些考虑基于 API 构建、或者觉得“哦,我可以用开放模型和 API 做酷事”的人。你们的 API 和平台允许人们做什么?我知道可以在上面构建智能体,请谈谈你们允许什么。

One last question just for folks that are thinking about building on the API or just like oh wait I could do cool stuff with open models and APIs. What what does your API and and platform allow people to do? Like I know you can build agents on top of the platform. Just talk about what you allow.

Sherwin Wu

从根本上说,API 提供了一系列开发者端点,这些端点让你可以从我们的模型中采样。目前最流行的一个叫 Responses API。这个端点针对构建长时间运行的智能体进行了优化,也就是那些会工作一段时间的智能体。在最低层级,你基本上就是给模型输入文本,模型会工作一段时间,你可以拉取看看它在做什么,然后某个时刻你会得到模型响应。这是我们提供的最底层原语,实际上很多人都在用,也是基于 API 构建的最流行方式。它非常不预设观点,你可以做任何你想做的事,是最底层的东西。我们也开始在上面构建越来越多的抽象层来帮助人们。再上一层,我们有 Agents SDK,它也变得极其流行。它允许你使用 Responses API 或其他端点来构建更传统意义上的智能体,比如一个在无限循环中工作的 AI,它可能有子智能体可以委派任务。它实际上开始构建所有这些框架和脚手架。我们会看看这一切会走向何方,但它让你更容易构建这类智能体,给它设置护栏,允许它把子任务分派给其他智能体,并编排一群智能体。Agents SDK 让你能做到这一点。再往上,我们现在开始构建工具来帮助处理部署智能体的元层面。所以我们有一个叫 Agent Kit 的产品和 Widgets,基本上是一组 UI 组件,你可以非常轻松地在我们的 API 或 Agents SDK 之上构建漂亮的 UI,因为很多时候这些智能体从 UI 角度看非常相似。我们还有一些评估产品,比如 Eval API,如果你想测试你的模型、智能体或工作流是否正常工作,你可以用我们的 EDOLs 产品以非常量化的方式进行测试。所以我把这看作不同的层级,它们都在帮助你用我们的 AI 和模型构建你想要的东西,抽象程度和预设程度越来越高。你可以使用整个栈,很快就能构建一个智能体;也可以一直往下走到 Responses API 这样的底层,构建任何你想要的东西,因为它足够底层。

So fundamentally the API offers a bunch of developer endpoints uh and and uh and these developer endpoints basically let you sample from our models. The most popular one that we have right now is one called responses API. Uh and so this is an endpoint and it's optimized for building longunning agents. So agents that'll work for a while. So what you can basically you can at a very you know uh uh low level you're basically just giving the model text. The model will work for a while. you can kind of, you know, pull it to see see what it'll do and then you'll get the model response back at at some point. That's like the lowest level primitive that we have uh for people and that's actually what a lot of people use. That's the most popular way of building on top of API with that. It is like super unpopinionated and you can do basically whatever you want. It's like the lowest level thing. We've also started building more and more kind of like layers of abstraction on top to help people build uh some of these. Uh and so next layer up we have this thing called the agents SDK which has also gotten extremely extremely popular. Um this allows you to use you know the responses API or some other API endpoints that we have to build what you might more traditionally think of as an agent like a you know an AI kind of working in an infinite loop. It might have sub agents that it delegates to. It starts building all this framework all this scaffolding actually. You know we'll see where this all goes. Um, but it makes it a lot easier for you to build these these these these kind of agents, giving it guard rails, allowing it to like farm out subtasks to other agents and and kind of like orchestrate a swarm of agents. Uh, the agents SDK uh kind of allows you to do that. And then above that, uh, we've now started building tools to help also with kind of like the meta level of deploying an agent. Uh so we have this product called uh um agent kit uh uh uh and widgets uh which are basically a bunch of UI components that you can use to very easily um build a very beautiful UI um on top of uh uh either our API or agents SDK um because you know a lot of times these agents kind of look very similar from a UI perspective uh and so there's agent kit we also have a smattering of like uh eval products like an eval API where if you want to test and like you know see if your models or your your agent or your workflow was working. Uh you can test it in a very quantitative way um using our EDOLs product. And so yeah, that I I view it as like these these various layers. They're all kind of helping you build um what you want um with our AI uh with our models um and with increasing levels of abstraction and and and and uh you know how opinionated it is. And so um you can start you can do you can use the whole stack and and it it very quickly allows you to build an agent or you can go down down the stack as low as you want to basically responses API and build whatever you want uh because of how low level it is.

对未来几年的鼓励 Encouragement for the Next Few Years

Host

Sherwin,你还有什么想分享的吗?还有什么想留给听众的?在我们进入非常激动人心的闪电问答之前,有没有什么我们没谈到但你觉得可能有帮助的?

Sherwin, is there anything else that you want to share? Anything else you want to leave listeners with? Anything we haven't touched on that you think might be helpful before we get to our very exciting lightning round?

Sherwin Wu

我唯一想留给听众的是,我认为未来两三年将是科技和创业界很长时间以来最有趣的时期。我想鼓励大家不要认为这是理所当然的。我 2014 年进入职场,头几年很好,但之后有五六年的时间科技界并不那么令人兴奋。而过去三年是我职业生涯中最疯狂、最激动人心、最充满活力的时期,我认为未来两三年将是这种状态的延续。所以鼓励大家不要认为这是理所当然的。我自己也在努力不把它视为理所当然。在某个时候,这波浪潮会消退,变得增量很多。但在此期间,我们将探索很多很酷的东西,发明很多新事物,改变世界,改变我们的工作方式。这就是我想留给听众的主要信息。

The only thing I' I'd leave folks with is yeah, I think um I think the next like two to three years are going to be some of the most fun uh in tech and in the startup world uh that that we'll have in a very long time. And uh I would just encourage people to not uh not take it for granted. Like I I entered the workforce in 2014. It was great for like a couple years. I felt like there was like a period of like five to six years where it wasn't very exciting in tech. Uh and then in the last three years has just been the most insanely exciting energizing period uh of my career and I think the next two to three years are gonna be a continuation of that. And so uh would encourage people not take it for granted. I'm trying to not take it for granted. At some point you know this wave is going to play out and it's going to be a lot more you know incremental. Uh but in the meantime we're going to get to explore a lot of really cool things, invent a lot of new things and change the world and change how we work. And so uh that's the main thing I' I'd leave folks with.

Host

我喜欢这个信息。我想多花点时间谈谈。当你说不要错过时,你建议人们做什么?是去构建、投入、学习、加入一家做有趣事情的公司吗?对那些说“好吧,我不想错过这班船”的人,你的建议是什么?

I love this message. I want to spend a little more time on it. Um, when you say don't miss it, is it what do you recommend people do? Is it just build, lean in, learn, join a company building really interesting things? Like what's what's your advice to folks that are like, "Okay, I don't want to miss the boat."

Sherwin Wu

是的,我会说去参与其中。基本上就像你说的,投入进去。在上面构建工具是故事的一部分。只是使用工具,你不需要是软件工程师就能投入。我认为很多工作都会改变。所以只是使用工具,理解它能做什么和不能做什么的局限,这样你就能随着模型改进而观察它开始能做什么的趋势。基本上就是习惯并熟悉这项技术,而不是袖手旁观,让它从你身边溜走。

Yeah, I would just say engage with it. So, it's basically like what you said. Um, lean in. Um, building, uh, tools on top of this is is part of the, you know, it's part of the story. Um, just using the tools like you don't, you know, you don't need to be a software engineer to to lean into this. Um, all I think a lot of jobs are going to going to going to change here. So just using the tools, understanding the limitations of what it can and cannot do so that you can kind of watch the trend of what it can start to do um as the models improve and yeah and so it's basically like getting used and getting getting used to the technology and getting familiar with it instead of kind of like laying back and uh uh uh letting it letting it pass you.

应对信息过载 Dealing with Information Overload

Host

另一方面,我觉得有很多压力和焦虑,因为事情太多了。我该怎么跟上?这周我得学 CloudBot。天哪。你有没有学到什么,比如你身处中心,如何不感到过度压力和担心错过,同时保持对新闻的了解?你学到了什么?

On the flip side of that, there's a lot of I think stress and just anxiety around like there's so much happening. How do I keep up? I got to learn cloudbot this week. Oh god. What is there something you've learned about it just not like you're at the center of this? How do you not get overly stressed and worried about missing things that are going on and just stay on top of news? What what are some things you've done learned?

Sherwin Wu

是的,我认为我个人在这方面是个坏例子,因为我基本上长期在线,在 X 和公司 Slack 上。所以我实际上试图吸收,最终也吸收了很多。但我想说的是,从观察其他不像我这样上瘾的人来看,很多都是噪音。

Yeah, so I I think I'm personally a bad example of this because I am I'm basically chronically online uh on X and uh our company Slack. So I I I actually try and absorb I end up absorbing a lot of it. What I will say though is just like from observing other folks who are less, you know, addicted to this stuff like I am. Um, yeah, a lot of it is noise.

使用AI工具的建议 Advice on Engaging with AI Tools

Sherwin Wu

你不需要把 110% 的精力都投入进去。老实说,只深入使用一两个工具,从小处着手,就已经足够了。我认为,行业本身的疯狂节奏加上 AI 作为产品,共同制造了这种令人窒息的新闻更新速度。关键是,你不需要了解所有这些东西就能真正参与其中。哪怕只是安装 Claude Code 客户端并试用一下,安装 ChatGPT,把它连接到你的几个内部数据源,比如 Notion、Slack、GitHub,看看它能做什么、不能做什么。我认为这些都是参与的一部分。

Like you don't need to have 110% of this kind of pass your mind, like go into your mind. Honestly, just leaning into one or two different tools, starting small, is already more than you need here. I think just the combination of the frenetic pace of the industry and AI as a product creates this insane kind of pace of news, which is honestly very overwhelming. The main thing is you don't need to know all of that to really engage with what's happening right now. Even something as simple as just install the Claude Code client and play around with it, install ChatGPT, connect it to a couple of your internal data sources, Notion, Slack, GitHub, and see what it can and cannot do. All of that I think is a part of it.

Host

太棒了。Sherwin,至此我们进入了非常激动人心的快问快答环节。我有五个问题要问你。准备好了吗?

Amazing. Sherwin, with that we've reached our very exciting lightning round. I've got five questions for you. Are you ready?

Sherwin Wu

是的,当然。

Yeah, absolutely.

Host

第一个问题,你发现自己最常向别人推荐的两三本书是什么?

First question, what are two or three books that you find yourself recommending most to other people?

Sherwin Wu

哦,我会推荐一本非虚构和一本虚构类书籍。虚构类是我刚读完的,我强烈推荐。书名是《There Is No Antimemetics Division》,作者 qntm。他是一位网络作家,我在 X 上看到有人分享。这是一本科幻小说,我基本两天内就读完了。写得非常好,非常吸引人。它讲的是一个政府机构对抗那些让你遗忘的东西。这是一本非常聪明、有创意且新颖的书,我很喜欢。这本书还无意中很搞笑。它本意是科幻甚至恐怖风格,但让我笑了好几次。所以这是虚构类。非虚构类我要作弊,推荐两本。过去一年我读了很多关于中国和中美关系的书。去年出版的两本书让我大开眼界。第一本是 Dan Wang 的《Breakneck》。非常好。我很喜欢他的比喻:美国是律师社会,中国是工程社会,各有优缺点。读完后我想,是啊,美国确实像被律师统治。这是第一本。另一本是 Patrick McGee 关于苹果和中国的书,非常有趣。我是个超级苹果粉丝,如果你看到我的办公桌,全是苹果产品。但了解苹果与中国的关系非常吸引人,书里有很多关于苹果公司的内部信息,我觉得很迷人。这本书很引人入胜,也非常及时。

Oh, I'll talk about one non-fiction and one fiction book. The fiction book I just finished reading. I really recommend it. It's "There Is No Antimemetics Division" by qntm. It's an online author, but I saw it being shared on X. It's a science fiction kind of book. I basically devoured it in two days. It was super well written, super fascinating. It's about a government agency that fights things that make you forget it. It's a very smart, creative, and fresh book in terms of source material that I really like. The book is also unintentionally hilarious. It's meant to be a sci-fi almost horror style book, but it made me laugh a couple times. So that's the fiction book. For non-fiction, I'm going to cheat and recommend two of them. In the last year I've been reading a lot more about China and US-China relations. There are two books that came out in the last year that have been really eye opening for me. First one is the Dan Wang book "Breakneck". That one was really good. I really liked his analogy of the lawyerly US versus the engineering society of China, and their pros and cons. I read it and thought, yeah, it does seem like we're run by lawyers in the US. So that's one. The other one is the Patrick McGee book on Apple and China. It was super interesting. I'm a huge Apple fanboy. If you could see my desk right now, it's all Apple stuff. But it was fascinating learning about Apple's relationship to China, and it had a lot of inside information about Apple as a company that I found fascinating. It was quite a page turner and also a very timely book.

Host

那本反模因书听起来太棒了。你说话的时候我就在下单了。

The antimemetics book sounds amazing. I'm buying it right now as you're talking.

Sherwin Wu

是的,只有几百页。我两天就读完了,真的太好看了。

Yeah, it's only a couple hundred pages. I literally finished in two days. It was just so good.

Host

好的,好推荐。下一个,你最近最喜欢的电影或电视剧是什么?

Okay, great tip. Okay, favorite recent movie or TV show you have really enjoyed?

Sherwin Wu

这个问题有点难,因为我有两个孩子,工作又忙,所以真的没太多时间看电视剧。不过过去几周我看了几集。我其实是个动漫迷。我看了《咒术回战》第三季的几集,非常好看。总的来说,我非常喜欢日本动漫。我认为他们创造了最新颖、最独特的剧情和世界观,而西方媒体往往回避这些。所以我很喜欢,但确实没怎么看,最近就看了几集《咒术回战》。

Yeah, that one's tough because I have two kids and a busy job, so I really haven't had much time to watch TV shows. I will say in the last couple weeks I watched a couple episodes. I'm actually a big anime guy. I watched a couple episodes of the new season of "Jujutsu Kaisen" season 3. It was really good. In general, I'm a huge fan of Japanese anime. I think they create the most novel and unique plots and universes that western media has shied away from. So generally a big fan, but yeah, I haven't watched much, just saw a couple episodes of JJK recently.

Host

以你的角色来说,完全可以理解。

Extremely understandable in your role.

Sherwin Wu

是的。

Yeah.

Host

你最近发现并非常喜欢的产品是什么?

Favorite product you recently discovered that you really love?

Sherwin Wu

好的。我最近需要设置 Wi-Fi 和家庭网络,于是全套买了 Ubiquiti 的路由器和安防摄像头。我以前从没听说过这个品牌,之前一直用很简单的配置。这个产品做工非常好。我不知道你以前用过没有,但它基本上就是家庭网络界的苹果。产品很漂亮。但真正让它出色的是它的软件。他们有一个非常棒的手机应用来管理所有家庭网络。你可以用 Ubiquiti 买无线路由器,但需要家里布好以太网线。不过我觉得它真正出色的是安防摄像头。如果你把摄像头接入 Ubiquiti 生态系统,他们有一个非常棒的手机应用、Apple TV 应用和 iPad 应用,可以实时查看摄像头画面。价格有点贵,但也不是特别贵。产品体验非常棒。

Okay. So I recently had to set up Wi-Fi and home networking, and I went all in on Ubiquiti routers and security cameras. I had never heard of it before. I always had a very simple setup. It is just such a well-built product. I don't know if you used it before, but it's basically the Apple of home networking. Beautiful products. But the thing that actually makes it extremely good is its software. They have a really great mobile app to help manage all the home networking. So basically Ubiquiti, you can use it to buy wireless routers. You need Ethernet wiring throughout your house to use it. But I actually think what makes it really good are security cameras. If you have security cameras plugged into the Ubiquiti ecosystem, they have an incredible mobile app, Apple TV app, and iPad app to see the live feed of your cameras. They're a little pricey, but not that pricey. It's been an incredible product experience.

Host

好吧,我买了 Eero,看来我犯了个错误。好推荐。

All right, I went Eero so I made a mistake. Good tip.

Sherwin Wu

Eero 也不错,但我现在已经完全转向 Ubiquiti 了。

Eeros are pretty good too, but I'm fully converted to Ubiquiti at this point.

Host

好推荐。还有两个问题。你有没有最喜欢的人生格言,在工作或生活中经常用到?

Good tip. Okay, two more questions. Do you have a favorite life motto that you find yourself coming back to in work or in life?

Sherwin Wu

有的。我经常对自己说的一句话是:永远不要自怜。工作和生活中会发生很多事情。提醒自己永远不要自怜,并且你总有能动性把自己拉起来,这是我经常对自己说的,也经常对别人说。

Yeah. The one that I always repeat to myself is, never feel sorry for yourself. There are a lot of things that are going to happen at work and in life. Reminding yourself to never feel sorry and that you always have a sense of agency to pull yourself up is something I've had to tell myself a lot, and also something I repeat to a lot of other folks as well.

Host

最后一个问题。在你之前的工作中,你在 Opendoor 负责计算房屋的购买价格。你基本上建立了一个模型,告诉公司“我们愿意为这栋房子付多少钱”。有没有哪个影响房价的变量是你没想到会如此重要的?

Last question. So, in your previous life, you worked at Opendoor where you led work on basically figuring out how much to pay for houses. You basically built a model that told the company, "Here's how much we'll pay for this house." What's a variable in the price of a house that you didn't expect is really important and impacts the price of a house?

Sherwin Wu

有很多令人惊讶的变量。我列举几个最有趣的。电线杆和高压电线实际上对房价影响很大。

There's a bunch that were surprising. I'll list the couple most interesting ones. Power lines and high voltage power lines actually impact your price quite a lot.

来自OpenDoor的惊人见解 Surprising insights from OpenDoor

Sherwin Wu

我直到去了达拉斯才真正意识到这一点:当你的房子紧挨着那些巨大的高压线时,会发出嗡嗡声,而大多数人有家庭,你不想让孩子靠近那里。所以我觉得那件事真的让我很惊讶。

I didn't really fully internalize this until I went to Dallas and observed when your house sits next to one of these giant voltage lines, it's buzzing, and most people have families, you don't want your kids near there. So I think that was one that really surprised me.

Host

有道理。

That makes sense.

Sherwin Wu

是的。另一个我们一直很难量化的因素是户型。它非常重要,但量化一个好的户型和一个很差的户型很困难。我们做了很多工作,比如厨房有多宽、是什么风格的厨房、主卧在哪里等等,所以真的很难量化。但我记得户型是一个大问题,因为有些房子卖不出去,然后我们的运营团队进去后会说是户型问题。你怎么判断呢?你走进去,就能感觉到,户型感觉不对。所以那些是令人惊讶的。最后一个比我预想的影响更大的是整体外观吸引力,甚至包括前门。我记得有一本 Zillow 的书提到,更换前门往往是房屋投资回报率最高的。但作为买家,当你走近房子时的感觉、你与之互动的第一印象,我认为我低估了它的重要性。

Yeah. And then the other one that was always really difficult for us to quantify was floor plans. It's very important, but quantifying what a good floor plan is versus a really bad floor plan was hard. We were doing all these things like how wide is the kitchen, what style of kitchen is it, where's the master bedroom, and so it was really hard to quantify. But I remember floor plan was a big one because we'd have a home that wouldn't sell, and then our ops team would go in and say it's a floor plan issue. How could you tell? You go inside, you just feel it, the floor plan feels off. So those were surprising. And then the last one that was more impactful than I thought is general curb appeal, even the front door. I think there's a Zillow book on this where front door replacement tends to be the highest ROI for homes. But just the feel as you walk up to the home as a buyer, what you're interacting with and the first moments of the house, I think I underrated its importance.

Host

这非常有趣。而且我很欣赏你们必须通过代码来解决所有这些问题,而不是亲自走户型。关于户型,我有很多故事;它没有被数字化。所以有少数人拥有凤凰城和达拉斯所有这些房子的纸质户型图。OpenDoor 时代有很多有趣的故事。

That is extremely interesting. And I love that you had to figure out how to do all this in code and not walk floor plans. I have a bunch of stories around floor plans; it's not digitized. So there are a handful of people who have paper floor plans of all these homes in Phoenix and Dallas. A lot of fun stories from the OpenDoor days.

结束语与联系方式 Closing and contact info

Host

好的,Sherwin,非常感谢你参加这次访谈。这太棒了。大家可以在网上哪里找到你,听众怎样才能帮到你?

Okay, Sherwin, thank you so much for doing this. This was incredible. Where can folks find you online and how can listeners be useful to you?

Sherwin Wu

是的,我在 Twitter 上,也就是 X 上。我的账号是 Sherwin Wu,我主要发关于 OpenAI、API 以及我们正在推出的一些产品的推文。至于大家如何能帮到我:我很喜欢听人们正在构建的东西,所以如果你在创业,或者在捣鼓一个想法,我很希望你在 X 上联系我。我很想听听你在做什么,并了解 OpenAI 如何能帮助你。

Yeah, I'm online on Twitter, on X. I'm just Sherwin Wu, and I mostly just tweet about OpenAI and the API and some of the products that we're launching. And then how folks can be useful to me: I love hearing about things that people are building, so if you're working on a startup, if you're hacking on an idea, I would love for you to reach out to me on X. I would love to hear about what you're building and learn about how OpenAI can help support you.

Host

太棒了。Sherwin,非常感谢你来到这里。

Amazing. Sherwin, thank you so much for being here.

Sherwin Wu

是的。谢谢你,Lenny。

Yeah. Thank you, Lenny.

Host

再见,各位。非常感谢你们的收听。如果你觉得本期内容有价值,可以在 Apple Podcasts、Spotify 或你喜欢的播客应用上订阅本节目。同时,请考虑给我们评分或留下评论,这能帮助其他听众找到这个播客。你可以在 lennispodcast.com 找到所有往期节目或了解更多关于本节目的信息。下期节目再见。

Bye, everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.

互动版:逐字朗读 + 针对本期提问 →